Troubleshooting runbooks
Это содержимое пока не доступно на вашем языке.
Operational playbooks for Chronacta single-node and HA deployments.
Quick triage
Section titled “Quick triage”| Symptom | First checks |
|---|---|
| gRPC unavailable | Process up? Port open? chronacta health -server HOST:PORT |
| Append rejected | Expected version? Auth/RBAC? Disk space / min free bytes? Quotas? |
| Read lag / empty | Wrong stream ID? Scavenge removed indexes? Wrong from version? |
| Admin UI login fails | Auth enabled? Token store path? Development mode only for labs |
| Backup upload fails | Remote credentials, bucket, network; list-remote |
| Follower lag | Leader healthy? Network partition? cluster status lag fields |
Disk pressure
Section titled “Disk pressure”chronacta storage-status— notedata_bytes, readiness reason.- Dry-run scavenge:
chronacta scavenge -max-events-per-stream N -dry-run. - Review plan; execute with
-confirmonly after backup. - Create backup:
chronacta backup create …(+ auto-upload if configured).
HA: leader unavailable
Section titled “HA: leader unavailable”- Confirm process and host health on each node.
chronacta cluster statusfrom a reachable node.- If quorum lost, restore connectivity before membership changes.
- Promote only with documented procedure (ha-rolling-upgrade.md).
Split-brain recovery
Section titled “Split-brain recovery”- Stop writes on contested nodes (take offline or fence).
- Identify the node with the highest committed global position / epoch from durable storage + cluster metadata.
- Bring that node up as sole writer; rejoin followers from snapshot or catch-up.
- Never force two leaders with overlapping epochs — discard or reseed the divergent replica after backup.
Subscriptions stuck
Section titled “Subscriptions stuck”chronacta subscription list/ Admin UI Subscriptions.- Inspect lag, in-flight, dead letters.
- Pause → replay from a safe checkpoint → resume.
- Ack stuck in-flight only when safe for the consumer.
Transient subscribe dropped (legacy)
Section titled “Transient subscribe dropped (legacy)”Default (v1.1+): slow live subscribers enter catch-up mode; the stream stays open and SubscribeResponse.control may report CATCHING_UP. Tune client processing or use pkg/client/resilience SubscribeLive for auto-resubscribe after server restarts.
Legacy disconnect: set CHRONACTA_SUBSCRIPTION_CATCHUP_MODE=false to restore pre-v1.1 behavior (ErrSlowSubscriber disconnect).
Connector lag
Section titled “Connector lag”- Check
chronacta-connectorlogs and job handler HTTP status. - Verify subscription checkpoint advances (
subscription get -id JOB_SUB). - Handlers must be idempotent on
event_id(at-least-once delivery). - See connector.md.
Projections stalled
Section titled “Projections stalled”- Admin UI Projections → Detect stalled / status panel.
- Check errors and dead letters.
- Resume or rebuild from a known-good checkpoint after fixing the program.
Metrics
Section titled “Metrics”Prometheus: chronacta_* including tenant labels chronacta_tenant_events_total{tenant_id=…}. See monitoring.md.

