Replication and HA architecture
Это содержимое пока не доступно на вашем языке.
HA implementation for Chronacta clustering. Enabled when mode=cluster in data/cluster/cluster.json. Single-node deployments use mode=standalone and skip replication networking.
Problem
Section titled “Problem”Single-node Chronacta is correct and durable on one machine. HA adds tolerated node loss without:
- losing committed events;
- accepting writes on split-brain;
- breaking stream version / global position monotonicity;
- silently diverging replicas.
Topology
Section titled “Topology” ┌─────────────┐ Clients ─────────►│ Leader │───┐ └─────────────┘ │ │ │ AppendEntries (WAL batches) ┌──────▼──────┐ │ │ Follower │◄──┘ └─────────────┘ ┌─────────────┐ │ Follower │◄── … └─────────────┘- Leader: sole writer; assigns global positions (same as today on one node).
- Followers: apply replicated log; serve reads optionally (eventual consistency).
- Learners (optional later): catch up without vote in quorum.
Membership is persisted in cluster metadata; nodes can be added or removed through the cluster API.
Consensus
Section titled “Consensus”Raft (or compatible library) for:
- leader election;
- cluster metadata (members, epoch);
- replication commit index.
Event payload bytes replicate as already-encoded WAL frames after local leader commit — not a second ad-hoc format.
Replication unit
Section titled “Replication unit”- Leader commits to local WAL + segment (existing path).
- Leader sends AppendEntries with batch id (existing record batch UUID), global position range, payload bytes.
- Follower validates checksum, applies via recovery/append path, advances match index.
- Leader advances commit index when majority acks; then ACKs client.
Not allowed: rsync/scp of segment directories as the primary protocol (backup restore remains separate ops).
Consistency semantics
Section titled “Consistency semantics”| Operation | Single-node | HA mode |
|---|---|---|
| Append | Linearizable on node | Linearizable via leader + quorum persist |
| Read stream | Strong (local log) | Leader: strong; follower: bounded staleness |
Read $all |
Strong | Same as stream read policy |
| Subscribe | Local dispatcher | Leader-only or forwarded according to deployment |
Client rules (HA)
Section titled “Client rules (HA)”- Discover leader via
ClusterServiceor gRPCFailedPreconditionredirect. - Retries: use idempotency keys on append.
- Follower reads: document max lag; never claim linearizability.
Failure modes
Section titled “Failure modes”| Scenario | Expected behavior |
|---|---|
| Leader crash | Raft elects new leader; uncommitted client writes fail/timeout |
| Follower crash | Quorum may persist; follower catches up on rejoin |
| Network partition | Minority partition stops accepting writes |
| Duplicate AppendEntries | Idempotent by batch id + global position |
| Reordered delivery | Reject / buffer until contiguous |
| Full disk on follower | Stop replication; alert via metrics |
Cluster metadata
Section titled “Cluster metadata”Persisted under data/cluster/ (see pkg/cluster):
cluster.json— cluster id, epoch, members, current leadernode.json— this node’s id and advertise address
Single-node deployments use DefaultSingleNode() without enabling cluster networking.
Observability
Section titled “Observability”Cluster metrics:
chronacta_cluster_leader(gauge) — implementedchronacta_replication_lag_positions— implementedchronacta_replication_append_entries_total— implementedchronacta_cluster_epoch— implemented
Wire format: pkg/replication (EVFR frames wrapping committed WAL batches).
Related
Section titled “Related”HA production criteria
Section titled “HA production criteria”| Criterion | Target |
|---|---|
| Quorum write ack | Leader + majority of voting followers |
| Lag after catch-up | Near zero under steady load (cluster status) |
| Rolling upgrade | Followers first; promote optional; no dual writers |
| Chaos smoke | make ha-chaos green |
| Chaos full | make ha-chaos-full / CHRONACTA_CHAOS_FULL=1 — 30m race budget + compaction-under-load engine smoke |
| Split-brain | Fence contested writers; recover from highest durable epoch/position |
HA production profile
Section titled “HA production profile”SLO / RPO / RTO
Section titled “SLO / RPO / RTO”| Metric | Target (single-DC cluster) |
|---|---|
| RPO | 0 for acknowledged writes (quorum persisted before client ACK) |
| RTO | < 30s leader failover under default election timeouts (lab); tune election_timeout for production |
| Replication lag (steady) | 0 positions after catch-up (replication_lag in cluster status) |
| Readiness | /readyz on metrics port: engine disk/WAL healthy; leader not ready while replication_lag > ReadinessMaxReplicationLag (default 0) |
Partition matrix (automated)
Section titled “Partition matrix (automated)”| Scenario | Test |
|---|---|
| Follower rejects client write | TestFollowerRejectsAppend |
| Leader alone cannot satisfy quorum (3-node) | TestThreeNodeQuorumBlocksWhenOnlyLeader |
| Isolated follower does not self-elect | TestIsolatedFollowerDoesNotSelfElect |
| Majority partition elects new leader | TestThreeNodeMajorityFailover |
| Committed events survive leader kill | TestLeaderKillCommittedEventsSurvive |
Commercial chaos harness
Section titled “Commercial chaos harness”| Mode | Command | Duration |
|---|---|---|
| Smoke | make ha-chaos |
Go HA tests (~seconds) |
| Full | CHRONACTA_CHAOS_FULL=1 make ha-chaos-full |
30m race budget + HACommercial tests |
| Custom duration | CHRONACTA_CHAOS_DURATION=10m CHRONACTA_CHAOS_FULL=1 make ha-chaos |
Sustained append under load |
HA tests: pkg/server/ha_commercial_test.go (TestHACommercialSustainedAppendTwoNode, TestHACommercialLongChaosUnderLoad, TestHACommercialRollingUpgradeWithMigrateThreeNode).
Explicit non-goals
Section titled “Explicit non-goals”- Multi-datacenter active/active writes
- Automatic conflict resolution across divergent streams
- Replacing NATS as external fan-out
Consumer groups are documented separately in consumer-groups.md.

