Skip to content

Production support process

Defines how Chronacta production incidents are triaged and escalated.

Tier Scope Examples
L1 First response, gather facts, run read-only checks Acknowledge page, capture version/cluster status, open incident channel
L2 Diagnose from dashboards, execute runbooks Interpret Grafana panels, chronacta verify, rolling restart, follower lag
L3 Storage/HA engineering, code or data surgery Corruption, replication divergence, migration failures, patch releases

Escalate L1 → L2 within 15 minutes for SEV-1/SEV-2; L2 → L3 when runbooks do not resolve within SLA.

Level Definition Response target
SEV-1 Data loss risk, prolonged write outage, split-brain Immediate page; continuous work until mitigated
SEV-2 Degraded HA (no quorum / high lag), backup pipeline down Respond within 30 minutes
SEV-3 Single feature impaired (projection, subscription, Admin UI) Business hours
SEV-4 Docs, cosmetic, non-prod Backlog

Use Grafana overview before shell access:

Symptom Panel / metric Next step
Writes failing Engine ready (chronacta_engine_ready) Check /readyz, disk free space, troubleshooting.md
HA unstable Cluster leader, Replication lag chronacta cluster status; confirm single leader
Slow appends Append latency Compare with baseline; check disk sync, load
Projections stale Projection health, Lag by component chronacta projection list; resume/rebuild
Auth storm Auth and rate limits Review chronacta_auth_failures_total; check IdP / tokens
Disk filling On-disk storage Run scavenge dry-run; expand volume

Prometheus alerts: prometheus/chronacta-alerts.yaml. SLO table: monitoring.md.

  1. Acknowledge the page and open an incident channel.
  2. Capture: version (chronacta-server --version), cluster status, recent deploy, symptoms.
  3. Open Grafana Chronacta Overview dashboard; note firing alerts.
  4. Follow troubleshooting.md; prefer dry-run before destructive ops.
  5. Collect support bundle: chronacta support collect (see support-bundle.md).
  6. Take a backup before scavenge, restore, or membership surgery when data risk exists.
  1. On-call engineer (L2) → platform owner (L3) if SEV-1 not mitigated in 30 minutes.
  2. Involve storage/HA specialists for corruption or replication divergence.
  3. Customer-facing status updates for SEV-1/SEV-2 every 30 minutes until mitigated.
  • Timeline, root cause, blast radius, follow-ups.
  • Update runbooks if a gap blocked recovery.
  • File ops tickets for systemic fixes.
  • Prefer rolling upgrade (ha-rolling-upgrade.md).
  • Storage migrations: keep CHRONACTA_AUTO_MIGRATE=false unless the target release documents an automatic migration path.
  • Run pre-release checks before promote: make proto test test-race vet build.

Scenario: replication lag alert fires during business hours.

  1. L1 acknowledges and posts chronacta-server --version + link to Grafana.
  2. L2 confirms ChronactaReplicationLagHigh, checks leader panel and chronacta cluster status.
  3. L2 decides: transient load vs follower down; follows HA runbook.
  4. Document outcome; verify alert clears within SLO window.