End of Summer Flash Sale: Up to 40% off Kubernetes courses, certifications, and bundles!

Automate Network Health Checks with the network.healthchecks Ansible Collection

Automate Network Health Checks with the network.healthchecks Ansible Collection

Introduction

Network health checks are something every operations team runs — but often manually, inconsistently, or through one-off scripts tied to a specific platform. The network.healthchecks collection from Red Hat changes that by providing a set of platform-agnostic Ansible roles that collect, parse, and evaluate device health in a structured, repeatable way.

Whether you are a network engineer just getting started with Ansible automation, or an automation engineer adding network observability to your existing playbooks, this collection gives you a validated, production-ready starting point. It covers BGP, OSPF, interfaces, CPU, memory, filesystem, uptime, crash files, and environmental sensors — all from the same automation framework.

Why Does This Matter?

  • Manual health checks do not scale. Logging into 50 routers to run show ip bgp summary and eyeballing the output is toil. It is slow, error-prone, and impossible to audit consistently.
  • Platform diversity makes custom scripts expensive to maintain. An IOS script does not work on NX-OS. A NX-OS script does not work on EOS. The network.healthchecks collection abstracts the platform layer so the same playbook task works across your mixed-vendor fleet.
  • Structured output enables automation. Instead of parsing raw CLI text, the collection returns a structured dict with a top-level result: PASS/FAIL and per-check breakdowns you can act on programmatically — in a pipeline, a report, or an EDA rulebook.
  • It is Red Hat validated content. This collection ships through Automation Hub as validated content, which means it goes through a formal review process and carries Red Hat support.

What Is in the Collection

The collection ships 10 roles, each focused on a specific health check domain:

RoleWhat It Checks
network.healthchecks.bgpBGP neighbor states and session health
network.healthchecks.ospfOSPF neighbor states (IPv4 and IPv6)
network.healthchecks.interfacesInterface admin and operational state
network.healthchecks.cpuCPU utilization against configurable thresholds
network.healthchecks.memoryMemory utilization, free memory, buffers, and cache
network.healthchecks.filesystemDisk free percentage against a configurable threshold
network.healthchecks.uptimeSystem uptime for stability monitoring
network.healthchecks.crashfilesPresence of crash files indicating instability
network.healthchecks.environmentTemperature, fans, and power supply (NX-OS only)
network.healthchecks.ip_reachabilityPing reachability to a list of IP targets from the device

Info

The ip_reachability role is fully implemented and documented inside the collection but is not listed in the collection’s main README as of v2.0.0. It is available to use — the role README under roles/ip_reachability/ covers the variables and output format.

Platform Support

The routing protocol roles (BGP and OSPF) have the broadest platform coverage — Cisco IOS, IOS-XR, NX-OS, Arista EOS, Junos, and VyOS. Resource monitoring roles (CPU, memory, filesystem, uptime, crashfiles) support IOS, IOS-XR, NX-OS, and EOS. The environment role is currently NX-OS only.

Platform dispatch is automatic. The role reads ansible_network_os — e.g. cisco.ios.ios or cisco.nxos.nxos — extracts the OS identifier, and includes the correct OS-specific tasks and parser template. You write one playbook; the collection handles the rest.

How It Works Under the Hood

Each role follows the same four-step pipeline:

  1. Validate — confirm the target platform is in the supported list; fail gracefully if not.
  2. Collect — run a show command via ansible.utils.cli_parse with a TTP template that converts raw CLI output into a structured Ansible fact.
  3. Evaluate — pass the parsed fact through a filter plugin (health_check_view, interfaces_health_check_view, or ospf_health_check_view) that applies your named checks and thresholds, returning a health_checks dict.
  4. Assert — the final task uses failed_when: "'FAIL' == health_checks.result", giving you native Ansible failure behavior you can catch, notify on, or escalate.

The TTP templates in each role’s templates/ directory are the key to multi-platform support. There is one template per OS per command — for example ios_show_ip_bgp_summary.yaml and nxos_show_ip_bgp_summary.yaml — and the role selects the right one automatically.

Getting Started

Install from Automation Hub

You can browse the collection on the Red Hat Content Catalog (no login required) or go directly to the Automation Hub collection page (Red Hat login required) to check versions and dependencies.

Add the validated content server to your ansible.cfg:

[galaxy]
server_list = automation_hub

[galaxy_server.automation_hub]
url=https://console.redhat.com/api/automation-hub/content/validated/
auth_url=https://sso.redhat.com/auth/realms/redhat-external/protocol/openid-connect/token
token=<your-token>

Get your token from the Automation Hub web UI, then install the collection:

ansible-galaxy collection install network.healthchecks

Install Platform and Parsing Dependencies

You also need ansible.utils and ansible.netcommon for CLI parsing, plus the platform-specific collection(s) for your devices:

ansible-galaxy collection install ansible.utils ansible.netcommon
ansible-galaxy collection install cisco.ios cisco.nxos   # add others as needed

Info

The collection requires Ansible 2.15 or later. Check meta/runtime.yml in the installed collection for the exact version constraint.

Running Health Checks

All roles share the same invocation pattern — include the role and pass your check definitions as variables. The failed_when inside the role handles failure propagation automatically.

BGP Health Check

- name: Check BGP health
  ansible.builtin.include_role:
    name: network.healthchecks.bgp
  vars:
    bgp_health_check:
      name: health_check
      vars:
        details: false
        checks:
          - name: all_neighbors_up
            ignore_errors: false
          - name: min_neighbors_up
            min_count: 2

This task runs identically against a Cisco IOS router and a Cisco NX-OS switch. On IOS, the role runs show ip bgp summary and uses ios_show_ip_bgp_summary.yaml to parse it. On NX-OS, it uses nxos_show_ip_bgp_summary.yaml. The check logic and variable interface are identical.

The ignore_errors: true option on a check lets you include informational checks — like bgp_status_summary — without letting them drive the overall result to FAIL.

CPU Health Check — IOS vs NX-OS

The CPU role uses a two-tier threshold model: a warning level and a critical level.

- name: Check CPU health
  ansible.builtin.include_role:
    name: network.healthchecks.cpu
  vars:
    cpu_utilization:
      warning_threshold: 60    # results in WARNING status
      critical_threshold: 90   # results in FAIL status
      details: false

On IOS, the role parses show processes cpu. On NX-OS, it parses the same command in NX-OS format (which has a different output structure). The role normalises both into the same health_checks dict, so you get identical output regardless of platform.

The CPU role is the only one that returns a three-level result: PASS, WARNING, and FAIL. This makes it practical to alert at 60% and page at 90% from the same playbook, rather than picking a single cutoff point.

Interfaces Health Check

- name: Check interface states
  ansible.builtin.include_role:
    name: network.healthchecks.interfaces
  vars:
    interfaces_health_check:
      name: health_check
      vars:
        details: true
        checks:
          - name: all_operational_state_up
          - name: all_admin_state_up
          - name: min_operational_state_up
            min_count: 4

Setting details: true adds a per-interface breakdown to the output, listing each interface by name with its admin and operational state — useful for post-change validation reports and audit trails.

OSPF Health Check

OSPF checks follow the same pattern as BGP. The role collects both IPv4 and IPv6 OSPF neighbor data and evaluates them independently:

- name: Check OSPF health
  ansible.builtin.include_role:
    name: network.healthchecks.ospf
  vars:
    ospf_health_check:
      name: health_check
      vars:
        details: false
        checks:
          - name: all_neighbors_up
            ignore_errors: false
          - name: min_neighbors_up
            min_count: 1

Understanding the Output

Every role sets a health_checks fact with a consistent structure. The top-level result is the overall verdict; each named check has its own status and check-specific fields.

BGP example output:

{
  "result": "FAIL",
  "all_neighbors_up": {
    "status": "FAIL",
    "up": 1,
    "down": 1,
    "total": 2
  },
  "min_neighbors_up": {
    "status": "PASS",
    "up": 1,
    "down": 1,
    "total": 2
  }
}

CPU example output:

{
  "result": "WARNING",
  "cpu_utilization": {
    "status": "WARNING",
    "message": "CPU utilization is above the threshold",
    "1_min_avg": 72,
    "5_min_avg": 68,
    "threshold": 60
  }
}

Interfaces example output (with details):

{
  "result": "FAIL",
  "all_operational_state_up": {
    "status": "FAIL",
    "interfaces_status_summery": {
      "admin_down": 1, "admin_up": 5,
      "down": 1, "up": 5, "total": 6
    }
  },
  "min_operational_state_up": {
    "check_status": "PASS",
    "interfaces_status_summery": {
      "admin_down": 1, "admin_up": 5,
      "down": 1, "up": 5, "total": 6
    }
  }
}

A few things worth noting about how the output behaves:

  • ignore_errors: true on a check prevents that check’s FAIL from rolling up to the top-level result. Use this for informational checks — like a summary — that you want in the output without driving a pipeline failure.
  • details: true adds raw neighbor or interface data to the output, useful for post-change validation reports or audit trails.
  • The failed_when condition in each role task means Ansible failure integrates naturally with your existing playbook error handling, block/rescue patterns, and AAP job template failure notifications.

Practical Patterns

Pre/post change validation — run a set of health check roles before a maintenance window, register results, run the same roles after the change, and compare. If BGP neighbors that were Established before the change are now down, the post-check fails and you know immediately.

Scheduled monitoring from AAP — wrap the roles in a playbook, publish it as an AAP job template, and run it on a schedule. Set up a notification on job failure to route FAIL results to your alerting or ITSM tool.

Event-Driven Automation integration — pair with Ansible EDA. When a monitoring platform fires an alert, an EDA rulebook triggers a targeted health check playbook against the affected device. The structured health_checks output feeds back into the event payload for context-rich incident tickets or automated remediation decisions.

Composite health check playbook — include multiple roles in sequence to build a single pre/post health report across BGP, interfaces, CPU, and memory in one run. Because every role uses the same health_checks fact name, register each role result to a unique variable with register: and assemble a summary at the end.

Conclusion

The network.healthchecks collection removes the platform-specific scripting tax that makes network health checking expensive to build and maintain. You define your checks once — in Ansible variables — and the collection handles CLI collection, output parsing, and threshold evaluation across IOS, NX-OS, IOS-XR, EOS, Junos, and VyOS. The consistent result: PASS/FAIL structure means health checks integrate cleanly into pipelines, job templates, and event-driven workflows without any custom glue code.

If you manage a mixed-vendor network and want consistent, auditable health checks that run the same way everywhere, this collection is worth adding to your automation toolkit.

Happy Engineering!

Gineesh Madapparambath

Gineesh Madapparambath

Gineesh Madapparambath is the founder of techbeatly. He is the co-author of The Kubernetes Bible, Second Edition and the author of Ansible for Real Life Automation. He has worked as a Systems Engineer, Automation Specialist, and content author. His primary focus is on AI, Ansible Automation, Containerization (OpenShift & Kubernetes), and Infrastructure as Code (Terraform). (Read more: gineesh.com)


Note

Disclaimer: The views expressed and the content shared in all published articles on this website are solely those of the respective authors, and they do not necessarily reflect the views of the author’s employer or the platform. We strive to ensure the accuracy and validity of the content published on our website. However, we cannot guarantee the absolute correctness or completeness of the information provided. It is the responsibility of the readers and users of this website to verify the accuracy and appropriateness of any information or opinions expressed within the articles. If you come across any content that you believe to be incorrect or invalid, please contact us immediately so that we can address the issue promptly.

Share :

Related Posts

Learn Ansible: A Comprehensive Guide to Courses, Certifications, and Exams (2026 Update)

Learn Ansible: A Comprehensive Guide to Courses, Certifications, and Exams (2026 Update)

Introduction Ansible is one of the most widely used open-source automation tools, letting teams automate everything from Linux and Windows servers to …

Best Ways to Set Up an Ansible Development Environment (Linux, Windows, macOS, and Windows WSL)

Best Ways to Set Up an Ansible Development Environment (Linux, Windows, macOS, and Windows WSL)

Best Ways to Set Up an Ansible Development Environment (Linux, Windows, and macOS) If you want to write Ansible playbooks, roles, modules, or …

Ansible Capacity Planning: Ansible contol node and automation controller

Ansible Capacity Planning: Ansible contol node and automation controller

When you’re running automation at scale, whether it’s using Ansible contol node or Ansible Automation Platform (AAP) with its automation controller , …

Install Ansible AWX on Kubernetes

Install Ansible AWX on Kubernetes

Ansible AWX is the upstream project for the Ansible automation controller (Part of Red Hat Ansible Automation Platform). It offers a web-based user …

Where Should You Keep Your Ansible Collection?

Where Should You Keep Your Ansible Collection?

Simplifying Automation with Organized Collections Introduction Ansible Content Collections revolutionize automation content management by offering a …

AWS Dynamic Inventory in Ansible Automation Platform: aws_ec2 Plugin Best Practices

AWS Dynamic Inventory in Ansible Automation Platform: aws_ec2 Plugin Best Practices

Table of Contents The problem with a static AWS inventory Quick CLI demo: the plugin in 30 seconds Anatomy of the aws_ec2.yml file Getting hostnames …