Table of Contents

Automatic drift remediation (operator playbook)

This page documents Orleans.Lattice.Replication 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at automatic-drift-remediation.md, and llms.txt lists every page.

Cross-cluster replication in Orleans.Lattice.Replication is eventually consistent: every mutation rides the per-tree write-ahead log to each peer, the receiver applies it HLC-monotonically, and concurrent edits converge through the per-tree LatticeMergeMode. In the steady state every cluster eventually holds the same data for a given shard. Silent divergence - two clusters that have applied different effective state for the same shard and stay that way - should never happen, but a transport bug, a partial garbage-collection of the WAL, or an operator mistake can produce it.

The anti-entropy stack is the safety net for that case. It is a layered pipeline that detects divergence, localises it, repairs it, and guards the repair behind operator controls. This page is the operator playbook for the stack as a whole: what each stage does, the single posture they all share (off by default), how to opt in end to end, the metrics they expose, and how to read the failure modes you will see in telemetry.

Each stage has its own reference page; this playbook links them rather than repeating their detail:

Default-off posture

Every stage ships dark. With defaults unchanged, a host runs ordinary replication and nothing else: no probe scheduler, no new RPC traffic, no automatic repair, and no behaviour change. The whole stack is opt-in, stage by stage, so you can enable detection and watch drift telemetry for as long as you like before ever enabling automatic repair.

Stage Master flag Default
Digest probe (detection) DigestProbeEnabled false
Merkle walk (localisation) MerkleWalkEnabled false
Targeted leaf re-replay (repair) LeafReReplayEnabled false
Bootstrap-snapshot fallback (repair) BootstrapFallbackEnabled false
Automatic repair master gate AutoRemediateOnDigestMismatch false

The flags are layered AND-gates. Localisation only runs on a detected mismatch; repair only runs on a localised leaf; the bootstrap fallback only runs when leaf re-replay could not reach the localised divergence (see the bootstrap-snapshot fallback for the exact trigger conditions); and AutoRemediateOnDigestMismatch is an additional master gate in front of both repair stages. Detection and localisation are never gated by the repair controls, so you can observe drift without sending any repair traffic.

One case is not left to operator discretion: enabling a WAL retention ceiling (WalRetention) on a replicated tree lets the sender garbage-collect entries a lagging cross-cluster shipper has not shipped yet, which - unlike a local consumer's fall-off - is invisible to the fall-off detector and would diverge the receiver silently. The silo therefore refuses to start with that combination on a tree declared in ReplicatedTrees unless the detection stage (DigestProbeEnabled) is enabled or the risk is explicitly acknowledged via AllowWalRetentionWithoutAntiEntropy - both read from the cluster-wide options, not a per-tree override. A tree enabled only at runtime is not checked. See WalRetention for the full rule.

The pipeline, stage by stage

  1. Detect. The digest probe is a low-frequency, read-only background pass that compares the local content digest of each shard the tree's live shard map routes to against every peer's digest. A sustained Mismatch for a (tree, shard, peer) triple is the signal to localise that shard; because the digest also folds each cluster's own checkpoint offset, it is a conservative trigger rather than proof of divergence. The probe never mutates data and never advances a replication cursor.
  2. Localise. On a mismatch, the Merkle walk descends the local B+ tree top-down and narrows the divergence to a single leaf or a small set of leaves, using the clusters' one shared coordinate - separator-key ranges. It is strictly read-only.
  3. Repair from the WAL. Targeted leaf re-replay re-ships the retained WAL entries covering the localised ranges to the diverged peer. The repair travels the same TX-aware, causal-stable apply path as ordinary replication and is idempotent at the receiver: a recently applied (originClusterId, hlc, key, op) identity is suppressed, and anything else re-applies to the same state.
  4. Repair when re-replay cannot reach the divergence. When re-replay cannot supply the missing writes, the bootstrap-snapshot fallback re-derives the committed projection of only the divergent leaf range from the live tree and re-ships those committed rows. See the bootstrap-snapshot fallback for the conditions that trigger it and the bounds it respects.
  5. Guard. The remediation guards wrap the repair stages with an operator opt-in gate, a per-(tree, peer) rate cap, and a per-(tree, peer) circuit breaker, so automatic repair is opt-in, bounded, and self-fencing.

Opting in end to end

To enable the full pipeline including automatic repair, opt into each stage and the master gate. The byte and entry caps below bound the cost of a single pass so a pathological tree cannot turn a background safety net into an expensive scan.

siloBuilder.AddLatticeReplication(o =>
{
    o.ClusterId = "cluster-a";
    o.ReplicatedTrees = new Dictionary<string, LatticeMergeMode>
    {
        ["orders"] = LatticeMergeMode.LwwRegister,
    };

    // 1. Detection (off by default).
    o.DigestProbeEnabled = true;
    o.DigestProbeInterval = TimeSpan.FromMinutes(5);
    o.DigestProbeJitter = 0.2;

    // 2. Localisation (off by default). Runs only on a detected mismatch.
    o.MerkleWalkEnabled = true;
    o.MerkleWalkMaxDepth = 16;
    o.MerkleWalkMaxBytes = 1024 * 1024;

    // 3. Repair from the WAL (off by default). Runs only after a localised leaf.
    o.LeafReReplayEnabled = true;
    o.LeafReReplayMaxEntries = 4096;
    o.LeafReReplayMaxBytes = 1024 * 1024;

    // 4. Repair when re-replay cannot reach the divergence (off by default).
    o.BootstrapFallbackEnabled = true;
    o.BootstrapFallbackMaxEntries = 4096;
    o.BootstrapFallbackMaxBytes = 1024 * 1024;

    // 5. Guards: master gate, rate cap, circuit breaker (gate off by default).
    o.AutoRemediateOnDigestMismatch = true;
    o.RemediationTrafficBudgetFraction = 0.01;
    o.RemediationTrafficWindow = TimeSpan.FromMinutes(1);
    o.RemediationFailureThreshold = 3;
    o.RemediationCircuitResetInterval = TimeSpan.FromMinutes(5);
});

Two cross-cutting prerequisites apply to the repair stages:

  • A real transport. The repair re-ship goes through IReplicationTransport; the default no-op transport delivers nothing and returns an unaccepted ack, so every repair pass reports zero entries shipped and counts toward the remediation circuit breaker. Wire the gRPC binding (or a custom transport) for genuine cross-cluster repair.
  • Projection-digest maintenance must be on. Detection reads the core library's leaf-projection digest, which only exists when MaintainProjectionDigest is true (the default for user trees). A tree that opts out of digest maintenance has no digest to compare, so the entire stack is inert for it - see the projection-rebuild digest opt-out for the cross-cluster impact.

Metrics surface

Every stage emits on the single orleans.lattice.replication meter. The table below is the operator's at-a-glance index; each stage's reference page documents its tags and emission semantics in full.

Stage Metric Read it as
Detection digest_probe.compared Every shard/peer comparison, tagged with its outcome.
Detection digest_probe.mismatch A digest mismatch for a (tree, shard, peer) triple - divergent content, or checkpoint offsets that differ between the clusters.
Localisation merkle_walk.localised A pass narrowed the mismatch to one or more leaves.
Localisation merkle_walk.aborted A pass stopped before localising, tagged with its reason.
Repair (WAL) leaf_rereplay.entries WAL entries re-shipped to the peer.
Repair (WAL) leaf_rereplay.skipped A pass skipped without re-shipping, tagged with its reason.
Repair (snapshot) bootstrap_fallback.triggered A fallback pass began.
Repair (snapshot) bootstrap_fallback.entries Committed entries re-shipped to the peer.
Repair (snapshot) bootstrap_fallback.skipped A fallback pass skipped without re-shipping, tagged with its reason.
Guards digest_remediation.disabled An observable gauge of every (tree, peer) whose repair is currently disabled.
Guards digest_remediation.skipped A repair pass skipped before sending traffic, tagged with its reason.

All metric names are exposed as constants on LatticeReplicationMetrics for dashboards built from the public surface. The instrument catalogue is in observability; the shipped Grafana dashboards come from the separate Orleans.Lattice.Dashboards package, whose metric-to-panel map records which panel charts each instrument.

Failure-mode matrix

The stack is designed to fail safe and to make why it is not repairing legible in telemetry. The three failure modes an operator most often needs to recognise:

Failure mode What you see What it means Operator action
Version skew digest_probe.compared{outcome=version_skew} and merkle_walk.aborted{reason=version_skew} The two clusters carry different contribution-function versions, so their digests are not comparable - typically a rolling upgrade in flight. No divergence is asserted and no repair is attempted. Expected during an upgrade; it self-clears once both sides run the same version. Investigate only if it persists after the rollout completes.
Re-replay cannot reach the divergence leaf_rereplay.skipped{reason=wal_trimmed} or {reason=range_empty}, then either bootstrap_fallback.triggered (fallback on) or bootstrap_fallback.skipped{reason=disabled} (fallback off) Re-replay could not supply the missing writes for the localised range (a trimmed WAL or a below-cursor gap). Enable BootstrapFallbackEnabled so the snapshot fallback can re-derive the committed projection of the divergent range. While it is off, the divergence is detected and localised but not repaired.
Circuit-breaker tripped digest_remediation.disabled{reason=circuit_open} for a (tree, peer), with digest_remediation.skipped{reason=circuit_open} per skipped pass Repair failed RemediationFailureThreshold times in a row for that pair, so the breaker opened and is fencing further repair for RemediationCircuitResetInterval. Investigate the underlying repair failures (transport, peer health). The breaker half-opens after the cooldown and closes itself on a successful trial pass; no manual reset is required.

Two further skip reasons are normal background noise rather than failures: digest_probe.compared{outcome=remote_unavailable} (the peer has digesting turned off for that tree, or no real probe transport is registered) and digest_remediation.skipped{reason=opt_out} / digest_remediation.disabled{reason=opt_out} (you have not set AutoRemediateOnDigestMismatch, so detection runs but repair is intentionally off). A spent rate cap surfaces as digest_remediation.skipped{reason=budget_exhausted}; the skips stop once the RemediationTrafficWindow rolls over, while the matching digest_remediation.disabled{reason=budget_exhausted} series clears only on the pair's next remediation pass that completes without failing.

  1. Enable detection alone (DigestProbeEnabled) on a representative tree and watch digest_probe.mismatch to learn its steady-state baseline. Because the digest folds each cluster's own checkpoint offset, a non-zero baseline alone does not prove divergence.
  2. Add localisation (MerkleWalkEnabled) and confirm walks complete or abort with an understood reason.
  3. Wire a real transport and enable the repair stages (LeafReReplayEnabled, then BootstrapFallbackEnabled) with the guards in place, but leave AutoRemediateOnDigestMismatch off so you can rehearse the telemetry.
  4. Flip AutoRemediateOnDigestMismatch last, starting with conservative RemediationTrafficBudgetFraction and RemediationFailureThreshold values, and watch digest_remediation.disabled to confirm the guards behave as expected under load.