Monitoring, Runbooks & DR
Prometheus and Grafana visibility, 30 alert-driven runbooks, and quarterly DR drills with defined RTO/RPO targets.
- Every service exposes Prometheus metrics; Grafana ships dashboards for cluster overview and per-node resource usage.
- Thirty alert-driven runbooks cover critical failure modes (audit chain integrity, quorum loss, SPIRE identity rotation, biome signature verification) through warnings (MTU mismatch, certificate mismatch, SMART warnings) and external integration health.
- gough dr drill exercises the full recovery process — secondary connectivity, backup integrity, database replication lag, Vault state replication, SPIRE federation, and DNS update capability — without impacting production; RTO/RPO targets are documented per cluster size.
- Capacity forecasting is backed by WaddleAI and degrades gracefully if that service is unreachable, rather than blocking cluster operations.