Node Exporter Monitoring: Key Metrics, Alert Rules & Deployment Guide
This comprehensive guide covers deploying Node Exporter for Linux server monitoring, detailing key CPU, memory, disk, and network metrics with PromQL queries, alert rule examples for Prometheus Alertmanager, Grafana dashboard setup, and best practices for reducing false positives.
Introduction
Node Exporter is the core Prometheus component for collecting Linux server metrics. Operations engineers must understand its exposed metrics and configure effective alert rules to detect and locate failures early.
Architecture
Node Exporter runs as a standalone binary on each target server, listening on port 9100. Prometheus pulls metrics via HTTP every 15 seconds (default). Node Exporter reads /proc and /sys in real time and exposes metrics at /metrics. Prometheus stores the time‑series data; Alertmanager evaluates alert rules and sends notifications.
┌─────────────────┐
│ Prometheus │
│ │
│ pull metrics │
└────────┬────────┘
│ HTTP
│ :9100/metrics
↓
┌─────────────────┐
│ Node Exporter │
│ │
│ read /proc │
│ read /sys │
│ expose metrics │
└─────────────────┘Metric Naming & Types
Prometheus metrics follow <namespace>_<name>_<unit>_<suffix> (e.g., node_cpu_seconds_total, node_memory_MemTotal_bytes). Node Exporter mainly uses Counter (monotonically increasing) and Gauge (can go up or down).
PromQL Basics
Essential queries include label filtering ( node_cpu_seconds_total{cpu="0"}), regex matching ( mode=~"user|system"), aggregation ( sum by (mode), avg), range vectors ( [5m]), rate/increase ( rate(node_cpu_seconds_total[5m])), and arithmetic for utilization percentages.
Key Metrics & Alert Rules
CPU
node_cpu_seconds_total(Counter) with labels cpu, mode (user, system, idle, iowait, irq, softirq, steal, nice, guest).
Usage:
(1 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m]))) * 100.
Alert: HighCPUUsage > 80% for 5m; HighIOWait > 30% for 5m.
Load
node_load1, node_load5, node_load15 (Gauge). Normalized:
node_load5 / count(node_cpu_seconds_total{mode="idle"}) by (instance).
Alert: HighSystemLoad > 2 (normalized) for 5m.
Memory
Total: node_memory_MemTotal_bytes; Available: node_memory_MemAvailable_bytes (includes reclaimable cache/buffers); Free: node_memory_MemFree_bytes; Buffers/Cached; Swap total/free.
Usage: (MemTotal - MemAvailable) / MemTotal * 100.
Alerts: HighMemoryUsage > 90% (5m), LowMemoryAvailable < 1GB (5m, critical), HighSwapUsage > 50% (10m).
Disk
Filesystem size/avail/free ( node_filesystem_size_bytes, node_filesystem_avail_bytes, node_filesystem_free_bytes) with labels device, fstype, mountpoint.
I/O: node_disk_io_time_seconds_total, read/write bytes, completed I/O operations.
Usage: (size - avail) / size * 100 (exclude tmpfs/devtmpfs).
Alerts: HighDiskUsage > 85% (5m), DiskWillFillIn4Hours via predict_linear (5m), HighInodeUsage > 85% (5m), HighDiskIO > 90% (10m).
Network
Receive/transmit bytes/packets/errors/drops ( node_network_receive_bytes_total, etc.) with device label.
Rate queries exclude loopback ( device!~"lo").
Alerts: HighNetworkTraffic > 800 MB/s (5m), NetworkErrors > 10/s (5m), NetworkDrops > 10/s (5m).
System
Boot time ( node_boot_time_seconds), current time ( node_time_seconds), forks, context switches, running/blocked processes.
Alerts: ServerRebooted if uptime < 600s (1m, info), ClockSkew > 30s (5m), HighContextSwitches > 100k/s (10m), TooManyBlockedProcesses > 10 (5m).
Filesystem (inodes)
node_filesystem_files(total inodes), node_filesystem_files_free.
Usage: (files - files_free) / files * 100.
Alert: HighInodeUsage > 85% (5m).
Deployment Steps
1. Deploy Node Exporter
Download latest release (e.g., v1.6.1), extract, create symlink. Create systemd service with collectors excluding docker/kubelet mount points and virtual network devices. Start and enable.
# systemd unit example
ExecStart=/opt/node_exporter/node_exporter \
--collector.filesystem.mount-points-exclude=^/(dev|proc|sys|var/lib/docker/.+|var/lib/kubelet/.+)($|/) \
--collector.netclass.ignored-devices=^(veth.*|br.*|docker.*|virbr.*|lo)$ \
--collector.netdev.device-exclude=^(veth.*|br.*|docker.*|virbr.*|lo)$Verify with curl localhost:9100/metrics | head -20.
2. Configure Prometheus Scraping
Add job_name: 'node_exporter' with static targets (e.g., 192.168.1.101:9100) and labels env, group. Reload via curl -X POST localhost:9090/-/reload. Check Targets UI for UP status.
3. Configure Alert Rules
Create /etc/prometheus/rules/node_exporter_alerts.yml with groups for CPU, load, memory, disk, network, system. Each rule includes expr, for duration, severity label (critical/warning/info), and annotations with summary / description using {{ $labels.instance }} and {{ $value | humanize }}. Validate with promtool check rules and reload.
4. Configure Alertmanager
Install Alertmanager, define routes by severity (critical immediate, warning grouped, info business hours). Receivers via webhook (e.g., prometheus-webhook-dingtalk), WeChat, email. Add inhibit rules (e.g., NodeDown suppresses other alerts for same instance). Start service.
5. Integrate Notification Channels
Deploy prometheus-webhook-dingtalk for DingTalk; add WeChat and email receivers in Alertmanager config with templates.
6. Optimize Alert Rules
Adjust thresholds per environment (e.g., CPU 80% → 85%, extend for to 10m).
Filter by labels (e.g., env="production") to exclude test/dev.
Add inhibit rules (e.g., HighMemoryUsage suppresses HighSwapUsage).
Use silences for planned maintenance via Alertmanager UI.
7. Grafana Dashboards
Import official dashboards (ID 1860, 11074) or build custom panels with PromQL: CPU usage, memory usage, disk usage, network traffic (receive/transmit rates).
Common Commands
Node Exporter
systemctl start|stop|restart|status node_exporter
journalctl -u node_exporter -f
curl http://localhost:9100/metrics
ss -tulnp | grep 9100Prometheus
promtool check config /etc/prometheus/prometheus.yml
promtool check rules /etc/prometheus/rules/*.yml
curl -X POST http://localhost:9090/-/reload
curl http://localhost:9090/api/v1/targets
curl 'http://localhost:9090/api/v1/query?query=up'Alertmanager
amtool check-config /etc/alertmanager/alertmanager.yml
amtool alert
amtool silence add alertname=HighCPUUsage instance=192.168.1.101:9100
amtool silence queryTroubleshooting Path
Check alert details (server, metric).
View Grafana dashboard for that server.
SSH and run top, free, df.
Apply metric‑specific diagnostics.
Monitor after fix; if false positive, tune rule.
Risks & Best Practices
Deployment
Node Exporter needs access to /proc, /sys; do not run inside containers.
Open firewall port 9100; use TLS in production.
Alert Rules
Avoid too‑low thresholds; set adequate for durations.
Separate rules per environment; use multi‑channel for critical alerts.
Notifications
Prevent alert storms with grouping and inhibition.
Different policies for work vs. off hours.
Include context and runbook links in alerts.
Verification
Node Exporter: curl localhost:9100/metrics | grep node_cpu.
Prometheus: query up{job="node_exporter"} → should return 1.
Alert rules: check Alerts page in Prometheus UI.
Notifications: simulate load with stress --cpu 8 --timeout 600s and confirm delivery.
Rollback Procedures
Node Exporter: stop service, restore old symlink, restart.
Prometheus: restore prometheus.yml from backup, reload.
Alertmanager: restore alertmanager.yml, restart service.
Production Considerations
High Availability
Prometheus: federation or remote storage.
Alertmanager: cluster mode.
Node Exporter: per‑node, no HA needed.
Retention
--storage.tsdb.retention.time=30d
--storage.tsdb.retention.size=100GBSecurity
Firewall restrictions; Basic Auth or TLS; regular updates.
Performance
Tune scrape_interval; drop unneeded metrics; use remote_write.
Summary
Node Exporter provides comprehensive Linux metrics (CPU, memory, disk, network, system). Engineers must master key metrics and PromQL, configure sensible alerts with proper thresholds and durations, integrate notification channels, visualize with Grafana, and continuously refine rules to minimize noise. A robust monitoring foundation enables early detection, rapid diagnosis, and timely response.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
MaGe Linux Operations
Founded in 2009, MaGe Education is a top Chinese high‑end IT training brand. Its graduates earn 12K+ RMB salaries, and the school has trained tens of thousands of students. It offers high‑pay courses in Linux cloud operations, Python full‑stack, automation, data analysis, AI, and Go high‑concurrency architecture. Thanks to quality courses and a solid reputation, it has talent partnerships with numerous internet firms.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
