Skip to main content
Comprehensive guide for monitoring Stable nodes and performing routine maintenance tasks.

Monitoring stack overview

  • Prometheus: Metrics collection
  • Grafana: Visualization and dashboards
  • AlertManager: Alert routing and management
  • Node Exporter: System metrics
  • Loki: Log aggregation (optional)

Quick monitoring setup

Step 1: enable Prometheus metrics

Restart node:

Step 2: install Prometheus

Step 3: install Grafana

Key metrics to monitor

Node health metrics

System metrics

Grafana dashboard setup

Import Stable dashboard

Custom dashboard import

Import dashboards via Grafana UI:

AlertManager configuration

Install AlertManager

Alert rules

Log monitoring

Systemd logs

Log analysis scripts

Loki setup (optional)

Health check endpoints

HTTP endpoints

Health check script

Maintenance tasks

Daily maintenance

Weekly maintenance

Database maintenance

Performance monitoring

Resource usage tracking

Query performance

Monitoring best practices

  1. Set up redundant monitoring
    • Use external monitoring services
    • Implement cross-node monitoring
    • Set up dead man’s switch alerts
  2. Alert fatigue prevention
    • Tune alert thresholds based on baseline
    • Use alert grouping and inhibition
    • Implement escalation policies
  3. Data retention
    • Keep metrics for 30 days minimum
    • Archive important logs
    • Regular backup of monitoring configs
  4. Security
    • Secure Grafana with strong passwords
    • Use HTTPS for all endpoints
    • Restrict prometheus access
  5. Documentation
    • Document all custom metrics
    • Maintain runbooks for alerts
    • Keep dashboard descriptions updated

Next steps

Last modified on April 23, 2026