Monitoring, Runbooks & DR

Prometheus and Grafana visibility, 30 alert-driven runbooks, and quarterly DR drills with defined RTO/RPO targets.

  • Every service exposes Prometheus metrics; Grafana ships dashboards for cluster overview and per-node resource usage.
  • Thirty alert-driven runbooks cover critical failure modes (audit chain integrity, quorum loss, SPIRE identity rotation, biome signature verification) through warnings (MTU mismatch, certificate mismatch, SMART warnings) and external integration health.
  • gough dr drill exercises the full recovery process — secondary connectivity, backup integrity, database replication lag, Vault state replication, SPIRE federation, and DNS update capability — without impacting production; RTO/RPO targets are documented per cluster size.
  • Capacity forecasting is backed by WaddleAI and degrades gracefully if that service is unreachable, rather than blocking cluster operations.

← Back to all features

Full technical documentation →