Production support process
Это содержимое пока не доступно на вашем языке.
Defines how Chronacta production incidents are triaged and escalated.
Support tiers
Section titled “Support tiers”| Tier | Scope | Examples |
|---|---|---|
| L1 | First response, gather facts, run read-only checks | Acknowledge page, capture version/cluster status, open incident channel |
| L2 | Diagnose from dashboards, execute runbooks | Interpret Grafana panels, chronacta verify, rolling restart, follower lag |
| L3 | Storage/HA engineering, code or data surgery | Corruption, replication divergence, migration failures, patch releases |
Escalate L1 → L2 within 15 minutes for SEV-1/SEV-2; L2 → L3 when runbooks do not resolve within SLA.
Severity
Section titled “Severity”| Level | Definition | Response target |
|---|---|---|
| SEV-1 | Data loss risk, prolonged write outage, split-brain | Immediate page; continuous work until mitigated |
| SEV-2 | Degraded HA (no quorum / high lag), backup pipeline down | Respond within 30 minutes |
| SEV-3 | Single feature impaired (projection, subscription, Admin UI) | Business hours |
| SEV-4 | Docs, cosmetic, non-prod | Backlog |
Dashboard-first triage (SEV-2)
Section titled “Dashboard-first triage (SEV-2)”Use Grafana overview before shell access:
| Symptom | Panel / metric | Next step |
|---|---|---|
| Writes failing | Engine ready (chronacta_engine_ready) |
Check /readyz, disk free space, troubleshooting.md |
| HA unstable | Cluster leader, Replication lag | chronacta cluster status; confirm single leader |
| Slow appends | Append latency | Compare with baseline; check disk sync, load |
| Projections stale | Projection health, Lag by component | chronacta projection list; resume/rebuild |
| Auth storm | Auth and rate limits | Review chronacta_auth_failures_total; check IdP / tokens |
| Disk filling | On-disk storage | Run scavenge dry-run; expand volume |
Prometheus alerts: prometheus/chronacta-alerts.yaml. SLO table: monitoring.md.
On-call
Section titled “On-call”- Acknowledge the page and open an incident channel.
- Capture: version (
chronacta-server --version), cluster status, recent deploy, symptoms. - Open Grafana Chronacta Overview dashboard; note firing alerts.
- Follow troubleshooting.md; prefer dry-run before destructive ops.
- Collect support bundle:
chronacta support collect(see support-bundle.md). - Take a backup before scavenge, restore, or membership surgery when data risk exists.
Escalation
Section titled “Escalation”- On-call engineer (L2) → platform owner (L3) if SEV-1 not mitigated in 30 minutes.
- Involve storage/HA specialists for corruption or replication divergence.
- Customer-facing status updates for SEV-1/SEV-2 every 30 minutes until mitigated.
Post-incident
Section titled “Post-incident”- Timeline, root cause, blast radius, follow-ups.
- Update runbooks if a gap blocked recovery.
- File ops tickets for systemic fixes.
Change management
Section titled “Change management”- Prefer rolling upgrade (ha-rolling-upgrade.md).
- Storage migrations: keep
CHRONACTA_AUTO_MIGRATE=falseunless the target release documents an automatic migration path. - Run pre-release checks before promote:
make proto test test-race vet build.
Tabletop exercise
Section titled “Tabletop exercise”Scenario: replication lag alert fires during business hours.
- L1 acknowledges and posts
chronacta-server --version+ link to Grafana. - L2 confirms
ChronactaReplicationLagHigh, checks leader panel andchronacta cluster status. - L2 decides: transient load vs follower down; follows HA runbook.
- Document outcome; verify alert clears within SLO window.

