Operations 27 min read

Node Exporter Monitoring: Key Metrics, Alert Rules & Deployment Guide

This comprehensive guide covers deploying Node Exporter for Linux server monitoring, detailing key CPU, memory, disk, and network metrics with PromQL queries, alert rule examples for Prometheus Alertmanager, Grafana dashboard setup, and best practices for reducing false positives.

MaGe Linux Operations
MaGe Linux Operations
MaGe Linux Operations
Node Exporter Monitoring: Key Metrics, Alert Rules & Deployment Guide

Introduction

Node Exporter is the core Prometheus component for collecting Linux server metrics. Operations engineers must understand its exposed metrics and configure effective alert rules to detect and locate failures early.

Architecture

Node Exporter runs as a standalone binary on each target server, listening on port 9100. Prometheus pulls metrics via HTTP every 15 seconds (default). Node Exporter reads /proc and /sys in real time and exposes metrics at /metrics. Prometheus stores the time‑series data; Alertmanager evaluates alert rules and sends notifications.

┌─────────────────┐
│ Prometheus      │
│                 │
│  pull metrics   │
└────────┬────────┘
         │ HTTP
         │ :9100/metrics
         ↓
┌─────────────────┐
│ Node Exporter   │
│                 │
│  read /proc     │
│  read /sys      │
│  expose metrics │
└─────────────────┘

Metric Naming & Types

Prometheus metrics follow <namespace>_<name>_<unit>_<suffix> (e.g., node_cpu_seconds_total, node_memory_MemTotal_bytes). Node Exporter mainly uses Counter (monotonically increasing) and Gauge (can go up or down).

PromQL Basics

Essential queries include label filtering ( node_cpu_seconds_total{cpu="0"}), regex matching ( mode=~"user|system"), aggregation ( sum by (mode), avg), range vectors ( [5m]), rate/increase ( rate(node_cpu_seconds_total[5m])), and arithmetic for utilization percentages.

Key Metrics & Alert Rules

CPU

node_cpu_seconds_total

(Counter) with labels cpu, mode (user, system, idle, iowait, irq, softirq, steal, nice, guest).

Usage:

(1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m]))) * 100

.

Alert: HighCPUUsage > 80% for 5m; HighIOWait > 30% for 5m.

Load

node_load1

, node_load5, node_load15 (Gauge). Normalized:

node_load5 / count(node_cpu_seconds_total{mode="idle"}) by (instance)

.

Alert: HighSystemLoad > 2 (normalized) for 5m.

Memory

Total: node_memory_MemTotal_bytes; Available: node_memory_MemAvailable_bytes (includes reclaimable cache/buffers); Free: node_memory_MemFree_bytes; Buffers/Cached; Swap total/free.

Usage: (MemTotal - MemAvailable) / MemTotal * 100.

Alerts: HighMemoryUsage > 90% (5m), LowMemoryAvailable < 1GB (5m, critical), HighSwapUsage > 50% (10m).

Disk

Filesystem size/avail/free ( node_filesystem_size_bytes, node_filesystem_avail_bytes, node_filesystem_free_bytes) with labels device, fstype, mountpoint.

I/O: node_disk_io_time_seconds_total, read/write bytes, completed I/O operations.

Usage: (size - avail) / size * 100 (exclude tmpfs/devtmpfs).

Alerts: HighDiskUsage > 85% (5m), DiskWillFillIn4Hours via predict_linear (5m), HighInodeUsage > 85% (5m), HighDiskIO > 90% (10m).

Network

Receive/transmit bytes/packets/errors/drops ( node_network_receive_bytes_total, etc.) with device label.

Rate queries exclude loopback ( device!~"lo").

Alerts: HighNetworkTraffic > 800 MB/s (5m), NetworkErrors > 10/s (5m), NetworkDrops > 10/s (5m).

System

Boot time ( node_boot_time_seconds), current time ( node_time_seconds), forks, context switches, running/blocked processes.

Alerts: ServerRebooted if uptime < 600s (1m, info), ClockSkew > 30s (5m), HighContextSwitches > 100k/s (10m), TooManyBlockedProcesses > 10 (5m).

Filesystem (inodes)

node_filesystem_files

(total inodes), node_filesystem_files_free.

Usage: (files - files_free) / files * 100.

Alert: HighInodeUsage > 85% (5m).

Deployment Steps

1. Deploy Node Exporter

Download latest release (e.g., v1.6.1), extract, create symlink. Create systemd service with collectors excluding docker/kubelet mount points and virtual network devices. Start and enable.

# systemd unit example
ExecStart=/opt/node_exporter/node_exporter \
  --collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/) \
  --collector.netclass.ignored-devices=^(veth.*|br.*|docker.*|virbr.*|lo)$ \
  --collector.netdev.device-exclude=^(veth.*|br.*|docker.*|virbr.*|lo)$

Verify with curl localhost:9100/metrics | head -20.

2. Configure Prometheus Scraping

Add job_name: 'node_exporter' with static targets (e.g., 192.168.1.101:9100) and labels env, group. Reload via curl -X POST localhost:9090/-/reload. Check Targets UI for UP status.

3. Configure Alert Rules

Create /etc/prometheus/rules/node_exporter_alerts.yml with groups for CPU, load, memory, disk, network, system. Each rule includes expr, for duration, severity label (critical/warning/info), and annotations with summary / description using {{ $labels.instance }} and {{ $value | humanize }}. Validate with promtool check rules and reload.

4. Configure Alertmanager

Install Alertmanager, define routes by severity (critical immediate, warning grouped, info business hours). Receivers via webhook (e.g., prometheus-webhook-dingtalk), WeChat, email. Add inhibit rules (e.g., NodeDown suppresses other alerts for same instance). Start service.

5. Integrate Notification Channels

Deploy prometheus-webhook-dingtalk for DingTalk; add WeChat and email receivers in Alertmanager config with templates.

6. Optimize Alert Rules

Adjust thresholds per environment (e.g., CPU 80% → 85%, extend for to 10m).

Filter by labels (e.g., env="production") to exclude test/dev.

Add inhibit rules (e.g., HighMemoryUsage suppresses HighSwapUsage).

Use silences for planned maintenance via Alertmanager UI.

7. Grafana Dashboards

Import official dashboards (ID 1860, 11074) or build custom panels with PromQL: CPU usage, memory usage, disk usage, network traffic (receive/transmit rates).

Common Commands

Node Exporter

systemctl start|stop|restart|status node_exporter
journalctl -u node_exporter -f
curl http://localhost:9100/metrics
ss -tulnp | grep 9100

Prometheus

promtool check config /etc/prometheus/prometheus.yml
promtool check rules /etc/prometheus/rules/*.yml
curl -X POST http://localhost:9090/-/reload
curl http://localhost:9090/api/v1/targets
curl 'http://localhost:9090/api/v1/query?query=up'

Alertmanager

amtool check-config /etc/alertmanager/alertmanager.yml
amtool alert
amtool silence add alertname=HighCPUUsage instance=192.168.1.101:9100
amtool silence query

Troubleshooting Path

Check alert details (server, metric).

View Grafana dashboard for that server.

SSH and run top, free, df.

Apply metric‑specific diagnostics.

Monitor after fix; if false positive, tune rule.

Risks & Best Practices

Deployment

Node Exporter needs access to /proc, /sys; do not run inside containers.

Open firewall port 9100; use TLS in production.

Alert Rules

Avoid too‑low thresholds; set adequate for durations.

Separate rules per environment; use multi‑channel for critical alerts.

Notifications

Prevent alert storms with grouping and inhibition.

Different policies for work vs. off hours.

Include context and runbook links in alerts.

Verification

Node Exporter: curl localhost:9100/metrics | grep node_cpu.

Prometheus: query up{job="node_exporter"} → should return 1.

Alert rules: check Alerts page in Prometheus UI.

Notifications: simulate load with stress --cpu 8 --timeout 600s and confirm delivery.

Rollback Procedures

Node Exporter: stop service, restore old symlink, restart.

Prometheus: restore prometheus.yml from backup, reload.

Alertmanager: restore alertmanager.yml, restart service.

Production Considerations

High Availability

Prometheus: federation or remote storage.

Alertmanager: cluster mode.

Node Exporter: per‑node, no HA needed.

Retention

--storage.tsdb.retention.time=30d
--storage.tsdb.retention.size=100GB

Security

Firewall restrictions; Basic Auth or TLS; regular updates.

Performance

Tune scrape_interval; drop unneeded metrics; use remote_write.

Summary

Node Exporter provides comprehensive Linux metrics (CPU, memory, disk, network, system). Engineers must master key metrics and PromQL, configure sensible alerts with proper thresholds and durations, integrate notification channels, visualize with Grafana, and continuously refine rules to minimize noise. A robust monitoring foundation enables early detection, rapid diagnosis, and timely response.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AlertingPrometheusServer MonitoringPromQLGrafanaAlertmanagerNode ExporterLinux Metrics
MaGe Linux Operations
Written by

MaGe Linux Operations

Founded in 2009, MaGe Education is a top Chinese high‑end IT training brand. Its graduates earn 12K+ RMB salaries, and the school has trained tens of thousands of students. It offers high‑pay courses in Linux cloud operations, Python full‑stack, automation, data analysis, AI, and Go high‑concurrency architecture. Thanks to quality courses and a solid reputation, it has talent partnerships with numerous internet firms.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.