Table of Contents

Leaf-Projection Rebuild & Digest

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at projection-rebuild.md, and llms.txt lists every page.

Orleans.Lattice's partitioned write-ahead log (WAL) is, in a fully replicated deployment, the canonical durable record of every leaf mutation. Each leaf grain materialises that log into a per-activation in-memory projection (the entry cache - a sorted dictionary owned by the leaf grain for the lifetime of the activation, not persisted). The persisted leaf state row carries no entries: it holds topology, the per-partition projection checkpoint offsets, the HLC clock and version vectors, a 16-byte projection-digest XOR fold, and a little replay bookkeeping. On every activation the cache is rebuilt: the leaf reloads its latest snapshot where it has a usable one and replays each WAL partition from just past the offset that snapshot covers, or from the start of the partition's readable window when no snapshot covers it (see Snapshot-on-fall-off safety net). Two operational concerns naturally arise:

  1. Drift detection. If a silo's leaf projection diverges from the WAL prefix it claims to have applied - a cosmic-ray bit flip, a storage-provider read-after-write anomaly, a bug in ILeafProjection.Apply - how does an operator notice before downstream readers do?
  2. Recovery from WAL trim. If a leaf has been cold long enough that the WAL has been trimmed past its last persisted checkpoint, the leaf cannot resume by tail-replay alone. What does activation do?

This document covers the two surfaces that answer those questions: ILattice.GetLeafProjectionDigestAsync (drift detection) and ProjectionRebuildPolicy (recovery from genuine loss), together with the MaxLeafReplayEntries / LeafProjectionRetention cost thresholds that govern when a long or stale cold replay is flagged.

Drift detection: GetLeafProjectionDigestAsync

LeafProjectionDigest digest = await tree.GetLeafProjectionDigestAsync(
    shardIndex: 0,
    cancellationToken);

// digest.Hash             - 16-byte XxHash128 fingerprint of the shard's projection
// digest.EntryCount       - entries (live + tombstoned) folded into the hash
// digest.CheckpointOffset - highest per-leaf projection-checkpoint offset
// digest.Version          - contribution-function version; compare hashes only
//                           between digests that carry the same Version

GetLeafProjectionDigestAsync reads the requested physical shard's root and returns a pre-folded digest in O(1) grain hops at the shard level. The shard's internal-node root maintains a running XOR-fold over every descendant leaf's running per-leaf hash. Each internal node persists the latest digest snapshot (hash, entry count, checkpoint offset) of each of its children and re-derives its subtree aggregate from that table - the XOR of the child hashes, the sum of their entry counts, and the maximum of their checkpoint offsets - whenever a child publishes a fresh snapshot upward. A leaf publishes after its mutations persist, and mutations that land within DigestCoalescingWindowMs (default 5 ms) share one publish; structural changes such as splits and tombstone reaps publish immediately. The public shard hash is

XxHash128( xor_subtree_hash || subtree_entry_count || subtree_max_checkpoint )

where xor_subtree_hash is the bitwise XOR of every descendant leaf's 16-byte running hash. A single-byte difference at any leaf - a stale tombstone, a missing TTL stamp, a divergent vector clock - surfaces as a different shard hash. Operators running multiple silos against the same WAL can poll the digest from each silo and compare bytes; equality is the strongest possible cross-silo state-equivalence check the library provides.

When the shard's root is a single leaf (flat-tree case, no internal node yet exists), the digest is read directly from the root leaf. A never-written shard is no exception: the read first materialises its root leaf, so an empty shard reports EntryCount = 0 and CheckpointOffset = 0 with a hash equal to the XxHash128 of 32 zero bytes (the all-zero 16-byte fold followed by the two zero counters).

XxHash128 is a non-cryptographic hash: it is chosen for ~10x lower CPU cost than SHA-256 on the per-mutation hot path and for its uniformly distributed output (which the XOR-fold algebra requires). The digest is a drift-detection fingerprint, not an authentication tag - a malicious operator with write access to the projection state could craft a collision, but the digest's job is to catch silent corruption, not to defend against forgery.

What is folded into the leaf hash

For every entry in the leaf's in-memory entry cache (a sorted dictionary keyed with StringComparer.Ordinal, rebuilt from the WAL on every activation) the implementation computes a 16-byte XxHash128 contribution over the following fields, in this order:

  1. key (length-prefixed UTF-8)
  2. lww.Timestamp.WallClockTicks (Int64, little-endian)
  3. lww.Timestamp.Counter (Int32, little-endian)
  4. lww.IsTombstone (byte, 0x00 or 0x01)
  5. lww.ExpiresAtTicks (Int64, little-endian - 0 when unset)
  6. lww.OriginClusterId (length-prefixed UTF-8, -1 sentinel for null)
  7. lww.VectorClock (a deterministic ordinal-sorted feed of every (replicaId, hlc.WallClockTicks, hlc.Counter) triple, or -1 sentinel when null/empty)
  8. lww.Value (length-prefixed bytes - -1 sentinel for tombstones)

The per-entry contributions are XOR-folded into a 16-byte running hash that is maintained incrementally on every mutation and persisted on the leaf state row as LeafNodeState.ProjectionHash. Insert XORs the new contribution in; replace XORs the old contribution out and the new one in (the old contribution cancels under self-inverse XOR); delete XORs the contribution out. Because XOR is commutative, associative, and self-inverse, the running hash is independent of insertion order and idempotent re-application of the same mutation is a no-op - exactly the algebra LWW already provides for entry state.

The public per-leaf digest is the XxHash128 of (running_xor || entryCount || checkpointOffset), where checkpointOffset is the leaf's partition-0 projection checkpoint (the legacy scalar slot), so two silos at different partition-0 replay positions report distinct digests even if their post-state happens to coincide; on a multi-partition tree the other partitions' positions are not folded in. The shard-level aggregate then XOR-folds each descendant leaf's running_xor directly (no per-leaf XxHash chaining step) and applies the same (xor_subtree_hash || subtree_entry_count || subtree_max_checkpoint) framing at the root. The XOR-fold makes the shard aggregate commutative and self-inverse, which is what lets each internal node maintain it incrementally as children publish updates upward.

Determinism contract

The digest is byte-stable across silos because every input is canonicalised:

  • The leaf's entry cache is a SortedDictionary<string, LwwValue<byte[]>> built with StringComparer.Ordinal, so the per-entry contributions are identical on every silo regardless of insertion order.
  • All numeric fields use little-endian framing via BinaryPrimitives.
  • All strings use Encoding.UTF8, length-prefixed with an Int32.
  • Length-prefix sentinels (-1) distinguish tombstone from empty value and null-string from empty-string so adjacent variable-length fields cannot collide.
  • VersionVector keys are sorted with StringComparer.Ordinal before feeding so dictionary insertion order does not perturb the output.

Topology changes and the aggregate

The internal-node aggregate is maintained incrementally as children publish ChildDigestSnapshot updates upward, so the aggregate's correctness depends on a single invariant: each child contributes to exactly one parent at any instant. A B+ tree split moves a contiguous half of a node's children to a new sibling, which transiently violates that invariant if the moved children's per-child digest rows are left behind on the donor or if a moved child keeps publishing to its former parent. Both would double-count the moved subtree's entries in the shard total.

The split path preserves the one-parent invariant in two steps:

  1. Prune on the donor. When an internal node splits, it removes the moved children's rows from its persisted per-child digest table and recomputes its SubtreeProjectionHash, SubtreeEntryCount, and SubtreeHighestCheckpointOffset from the remaining rows before publishing the corrected aggregate upward. The XOR fold's self-inverse algebra makes the recompute exact - the moved rows cancel cleanly out of the running hash.
  2. Reject stale publishes. Each internal node folds a child's digest snapshot only from a child it currently owns. A snapshot arriving from a child that has already been re-parented to the new sibling is rejected and its stale row (if any) dropped, so a moved child that races a publish against its re-parenting cannot reintroduce a double count. Every publisher also stamps a monotonic sequence, and a snapshot older than the one already folded for that child is dropped, so a late coalesced publish carrying a pre-split count cannot overwrite a fresher one. A donor likewise re-seeds a child's parent pointer only for children it still owns, so a moved child is never pointed back at the node it left.

The net effect is that EntryCount stays exactly equal to the number of distinct entries (live plus tombstoned) under the shard across an arbitrary sequence of internal-node splits, with no transient over- or under-count visible to a quiescent digest read.

The upward publish that maintains the aggregate is a cross-grain RPC that recurses up the internal-node chain. An internal node never makes it while holding its own non-reentrant split gate, and a parent whose gate is busy - for example mid-split - parks the incoming snapshot, keeping only the freshest one per child, and folds it before it releases the gate, so a publish never waits on its parent's split (issue #3523). A parent that is itself mid-mutation can still leave the await neither completing nor faulting, so every upward publish, from a leaf or an internal node, is bounded by LatticeOptions.DigestPublishTimeout (default 15 s): on the deadline the publish is abandoned and a TimeoutException is raised, with no count drift - the abandoned publish never partially applied at the parent, the publisher keeps its snapshot marked pending, and the next mutation's publish re-drives convergence. Where an internal node has already made its own change durable - accepting a split, or removing a reclaimed child - the timeout is logged and contained rather than surfaced, so that change is not lost. Set the option to InfiniteTimeSpan to restore the historical unbounded await. A non-zero orleans.lattice.internal.digest_publish.timeouts counter surfaces the condition; it counts internal-node publishes only.

Cost and where to call it

Because the per-entry XOR fold is maintained incrementally on every mutation, GetLeafProjectionDigestAsync does not re-walk the leaf's entry cache on each call - the running hash is already on the leaf's persisted state, so the per-leaf computation collapses to a single fixed-size XxHash128 over (running_xor || entryCount || checkpointOffset). The shard root delegates to the root internal node, which returns its persisted subtree aggregate in a single grain hop without re-visiting any descendant. The leaves themselves are not activated by the digest poll: each leaf published its contribution upward after its last mutation persisted (within one DigestCoalescingWindowMs), and the internal-node aggregate is the source of truth at read time. A whole-tree poll therefore costs O(shardCount) grain hops, regardless of how many leaves each shard owns or how many entries each leaf holds.

The cold-start path remains correct: if the shard root or any internal ancestor is activated for the first time, its persisted state is loaded from storage along with the aggregate it already stamped on the previous shutdown - no leaf walk is required to reconstruct it.

Heap allocations on the hot path are bounded:

Allocation Per call
XxHash128 for per-entry contributions (one cached per leaf grain activation) reused via TryGetHashAndReset
XxHash128 for the outer digest framing one per digest read, at the leaf or the internal-node root
byte[16] XxHash128 hash from GetHashAndReset() unavoidable (the result)
byte[16] hash clone carried by each upward digest publish bounded by tree height; cloned so subsequent XOR updates do not retroactively mutate the parent's captured bytes
String / VC scratch buffers pooled (stackalloc 256 fast path; ArrayPool<byte>.Shared and ArrayPool<string>.Shared for the rare overflow)

The O(shardCount) per-tree cost makes the digest cheap enough for steady-state monitoring - including periodic cross-silo equality canaries - not just on-demand diagnostics. It is safe to call against a live shard under load: it observes the current in-memory projection without taking any kind of consistency freeze. The result is necessarily a snapshot at one wall-clock instant, however, so two calls under sustained writes will report different digests; equality is meaningful only between quiescent observations (no in-flight writes to the shard between the two reads being compared, and at least one DigestCoalescingWindowMs elapsed since the last write so every coalesced publish has landed).

Cross-silo divergence example

// On every silo hosting the cluster, schedule a periodic poll
// over every shard and compare digests. A mismatch is a conservative
// trigger to investigate, not proof of drift: compare only quiescent
// reads carrying the same Version, because digests also differ when
// two reads were taken at different replay positions.
var routing = await tree.GetRoutingAsync();
foreach (var shardIndex in routing.Map.GetPhysicalShardIndices())
{
    LeafProjectionDigest digest = await tree.GetLeafProjectionDigestAsync(
        shardIndex,
        cancellationToken);
    // emit (silo, treeId, shardIndex, digest.Version, digest.Hash, digest.EntryCount,
    // digest.CheckpointOffset) to your telemetry pipeline.
}

Error surface

Condition Exception
shardIndex is not a physical shard of the per-tree map ArgumentOutOfRangeException
The tree id starts with the reserved system prefix _lattice_ LatticeReservedTreeNamespaceException (an InvalidOperationException subclass)
cancellationToken was already cancelled OperationCanceledException
Tree has LatticeOptions.MaintainProjectionDigest = false InvalidOperationException
An access gate is configured and does not authorise the caller to read the whole tree uniformly (a digest cannot be narrowed per key, so a partial allow is refused too) LatticeAuthorizationDeniedException

Opting out of digest maintenance

The digest's maintenance cost is small in absolute terms, but it recurs with the write load. Each leaf mutation costs one in-memory XOR fold over the entry's contribution. When the leaf's digest has changed, the leaf publishes it upward: the parent internal node rewrites its persisted subtree aggregate and publishes to its own parent in turn, up to the shard root, so each publish costs O(treeHeight) writes. With DigestCoalescingWindowMs at 0 every such mutation publishes. By default (5 ms) the foreground writes that land within one window - sets, deletes, range deletes and typed CRDT delta applies - share a single publish, while merge traffic and structural changes such as splits, tombstone reaps and saga terminals publish immediately. For trees that do not poll the digest - workloads that rely exclusively on audit logs, integration tests, application-level checksums, or external reconciliation, and never call GetLeafProjectionDigestAsync - the maintenance cost is pure write amplification.

LatticeOptions.MaintainProjectionDigest (default true) flips the behaviour off:

siloBuilder.ConfigureLattice(opts =>
{
    // Turn off digest maintenance globally - leaf mutations stop
    // updating the running XOR fold and stop publishing
    // ChildDigestSnapshot upward to internal-node ancestors.
    opts.MaintainProjectionDigest = false;
});

// Or per-tree:
siloBuilder.ConfigureLattice("audited-tree", opts =>
{
    opts.MaintainProjectionDigest = false;
});

When the opt-out is in effect:

  • Leaf-mutation funnels (StoreEntry / RemoveEntry) take a trimmed path that LWW-merges the value, bumps the delivery sequence, and returns without touching the persisted ProjectionHash.
  • The leaf does not publish ChildDigestSnapshot upward, so no internal-node ancestor updates its SubtreeProjectionHash for that mutation. The whole upward chain is quiescent.
  • ILattice.GetLeafProjectionDigestAsync throws InvalidOperationException. When the opt-out comes from the configured options it fails fast at the public surface, before any routing-table fetch or grain hop. That check does not read the registry, so an opt-out set only through the tree's registry override or the latch below is caught after the shard hop, by the same check in the leaf and internal grains, which also stops a direct grain-handle caller.
  • Persisted state is not rewritten. Any ProjectionHash already on disk from a previous-enabled period remains untouched.

Cross-cluster impact: anti-entropy drift detection

In a deployment running Orleans.Lattice.Replication, the anti-entropy peer digest probe reads this same leaf-projection digest to detect silent divergence between clusters. A tree with MaintainProjectionDigest = false (or one whose registry latch has disabled it permanently) has no digest to compare, so the probe skips that tree and classifies a peer in the same state as RemoteUnavailable rather than a mismatch. The whole automatic drift-detection-and- remediation stack - the probe, the Merkle-walk localisation, targeted leaf re-replay, and the bootstrap-snapshot fallback - is therefore inert for any tree that opts out of digest maintenance. Disable the digest only for trees you do not need cross-cluster drift telemetry on; see the automatic drift-remediation playbook for what the stack provides and how to opt in.

Disabling is a one-way operation per tree

The first mutation that lands while maintenance is disabled stamps an irreversible registry latch on the tree, reported as TreeConfigurationReport.ProjectionDigestPermanentlyDisabled by ILatticeTreeAdmin.GetTreeConfigAsync (see Orleans.Lattice.Api.TreeAdmin). The stamp is best-effort: a registry failure never fails the mutation, and the leaf retries the stamp on its next mutation while maintenance is still disabled, so the latch is set only once a stamp succeeds. Once the latch is set, every subsequent activation resolves MaintainProjectionDigest as false regardless of the per-tree override or the silo-wide default, and ILattice.GetLeafProjectionDigestAsync keeps throwing.

The latch exists because the digest is an XOR-fold aggregate over every mutation: any mutation accepted while maintenance was off permanently invalidates the persisted aggregate, and silently re-engaging maintenance would publish a known-stale digest as if it were authoritative. Once stamped, the one-way latch makes this impossible to mis-configure: an operator who turns the option back on for a tree that has already accepted writes under the disabled setting will see the resolved value stay at false and the digest API stay broken, rather than producing a digest that disagrees silently with the ground-truth entries.

The only way to re-engage digest maintenance for a latched tree is to rebuild the tree (or its leaf range) from scratch under a fresh registry entry. If you anticipate needing the digest later, leave it enabled.

Per-tree precedence and system trees

Resolution order for MaintainProjectionDigest:

  1. System-tree prefix override. Trees whose id begins with the reserved system prefix _lattice_ (e.g. the internal registry tree) always resolve as false regardless of configuration. System trees are not replicated and have no cross-silo drift-detection consumer, so the maintenance work is pure overhead.
  2. Registry latch. If ProjectionDigestPermanentlyDisabled is set, the resolved value is false.
  3. Per-tree override. If the tree's registry override (TreeConfigurationReport.MaintainProjectionDigest, written through ILatticeTreeAdmin.SetTreeConfigAsync) is set, that value wins over the configured options of step 4. Operators can opt an individual tree out (or, while the latch is not yet set, back in) without flipping the silo-wide default.
  4. Configured options. Falls back to LatticeOptions.MaintainProjectionDigest as configured for the tree: a named ConfigureLattice(treeName, ...) override when the host set one, otherwise the silo-wide value.

Disabling the digest is recommended for write-amplification-sensitive deployments that do not need cross-silo drift telemetry. Keep it enabled when you operate multiple silos against the same WAL and rely on the digest as a state-equivalence canary, or when chaos / soak tests use the digest as a post-condition oracle.

Why not store the digest in the WAL?

Moving the digest aggregate into the WAL would not eliminate the write amplification: the per-leaf XOR fold is already negligible (it lives inline in the leaf's persisted state - there is no extra WAL append for it today). The real amplification is the upward chain of internal-node updates: every publish from a leaf - one per coalesced group of foreground writes by default, or one per mutation when DigestCoalescingWindowMs is 0 - rewrites the persisted subtree aggregate on each ancestor up to the shard root. That cost lives in internal-node grain state, not in the WAL, and is the whole point of the incremental aggregate - readers need to find the pre-folded shard hash in O(1). Reconstructing it by replaying the WAL on every digest poll would defeat the optimisation and produce a per-call cost proportional to WAL size, which is strictly worse than the per-leaf walk it replaced. The opt-out is the correct knob for deployments that do not need the aggregate at all.

Recovery: fall-off-log triggers and ProjectionRebuildPolicy

When a leaf grain reactivates it consults its persisted per-partition ProjectionCheckpointOffsetsByPartition[p] (and the legacy scalar ProjectionCheckpointOffset for the partition-0 back-compat slot) and decides how to recover. The classifier runs once per partition in [0, WalPartitions); the leaf is refused as fall-off-log if any partition's classifier reports genuine loss, while a cost signal or the snapshot advisory below still tail-replays. Three triggers classify an individual partition - but only the first indicates missing data, and only the first is fatal:

  1. WAL trimmed past checkpoint. Partition p's WAL has GC'd entries the leaf still considers unapplied. A tail replay would skip those entries and converge to the wrong state. Skipped unless the partition's checkpoint is positive: the -1 "nothing applied" sentinel means the leaf has no in-memory state to lose to a trimmed prefix on that partition, and a checkpoint of 0 is skipped as well. This is the only trigger that routes to ProjectionRebuildPolicy. The exact loss boundary is tail > checkpoint + 1: the entry at the checkpoint is already applied, so trimming it loses nothing. A cold activation that finds no covering snapshot hands the classifier the -1 sentinel so that it replays the whole readable window, which blinds this trigger; that path is guarded separately against the leaf's durable checkpoint with the same boundary, and a trim past it surfaces LeafProjectionStaleException without consulting the policy.
  2. Replay budget. The gap walHead[p] - checkpoint[p] exceeds LatticeOptions.MaxLeafReplayEntries (default 10 000). This is a cost signal, not a loss signal, and the gap is not the leaf's work: it is measured across the whole WAL partition, which every leaf pinned to that partition shares, while MaxLeafReplayEntries is a per-leaf, post-range-filter budget. The two are in different units, and on a partition carrying ~1,350 leaves the gap overstates a leaf's real work by up to that fan-out. What the comparison does establish is a sound upper bound: the head is the next offset the partition will assign, so every entry a leaf applies lies inside (checkpoint, head) and applied <= gap always holds. Since issue #2275 nothing acts on the comparison (it was once a candidate the replay confirmed): an over-budget gap yields the decision TailReplayOverBudget and the leaf tail-replays exactly as for TailReplay, and its only remaining effect is that it suppresses the snapshot advisory described under Snapshot-on-fall-off safety net. Confirming the gap in the classifier would mean reading (checkpoint, head) before the replay reads it again - doubling the most expensive part of activation - so the verdict is taken on every replay path, during the replay that happens anyway, by counting the entries that actually pass the per-leaf range filter (ShouldApplyDuringReplay). The warning and the orleans.lattice.leaf.activation_replays_over_budget counter are emitted at that exact count, whatever the comparison said. The comparison is skipped for the -1 sentinel, whose gap would charge a fresh leaf for every sibling's WAL; the count is not, because it charges a fresh leaf only for its own range. The per-slice WalReplaySliceBudget still bounds individual coordinator reads on this path.
  3. Cold past retention. The persisted projection age exceeds LatticeOptions.LeafProjectionRetention (default 7 days). Also a cost signal only - an old checkpoint does not imply a trimmed WAL, so it degrades to the same non-fatal TailReplayOverBudget replay. Age is a property of the leaf, not of any one partition, so every partition's classification sees the same age. Note the activation path currently supplies TimeSpan.Zero as the age, so this trigger does not fire from activation today (tracked in #1738).

Why triggers 2 and 3 are not fatal (issue #1738). They were, until a tree holding fully intact data was permanently bricked by a replay gap of 10,648 against the 10,000 default - 648 entries, 6.5% over budget - while every offset it needed was still readable in the WAL. A cost guardrail must never be more destructive than the cost it guards against: a slow activation is recoverable, a tree that refuses to activate is not. Replay cost is bounded on the read side instead, by WalReplayMaxRecordsPerTurn (which yields between turns) and WalMaterialiserMaxConcurrentReplays.

Reading the over-budget warning (issue #2023). The warning names the tree, the leaf grain id, and the WAL partition ordinal, and it states its own fault criterion: a persisted checkpoint that does not advance across repeats is a fault, not a slow replay. That comparison is only valid between lines naming the same leaf and the same partition. partition is iterated [0, WalPartitions) inside every leaf's activation, so it does not identify a leaf; before the leaf id was added, consecutive lines were one-per-minute samples of arbitrary different leaves, and their checkpoints appearing to repeat or move backwards was an artifact of that sampling rather than a stalled replay. The log is throttled per (tree, leaf, partition), starting at one line a minute and doubling to a ceiling of one an hour the longer that leaf keeps reporting, so a leaf that is genuinely stuck still reports an unchanging checkpoint but at a decaying rate: compare consecutive lines naming that leaf, rather than expecting a fixed cadence. The backoff advances only when a line is actually emitted: a repeat the per-tree cap withholds keeps its place in the queue instead of backing off having said nothing, so a leaf on a busy tree still rotates into the budget and still yields the two comparable lines the criterion needs. A clean in-budget activation retires the backoff outright, so a leaf that misbehaves, recovers, and regresses hours later reports at the base interval rather than inheriting the accumulated ceiling. A per-tree cap additionally bounds how many repeat lines one tree may emit in a window, because a tree with L leaves and P partitions has L x P throttle keys and so L x P times the per-key rate - which is how this warning reached 46% of one deployment's container log, rolling away the older entries that were the evidence needed to diagnose it (issue #2100). Any repeats the cap withholds are reported as a summary line, so the cap is not silent while that tree keeps replaying, and a leaf partition reporting over budget for the FIRST time is exempt from it, so a newly appearing condition still surfaces promptly. "First time" is a property of the key's own history and not of what the gate happens to have retained: internal housekeeping never restores the exemption, so the exemption cannot be re-earned by churn on a large estate. It is restored only after that leaf partition has been silent for a full ceiling interval, at which point its return really is new information. That summary is carried by a later occurrence on the tree, so a tree's final withheld tally goes unreported once the condition resolves or the leaf deactivates; the counter below is the exact census for that case. The orleans.lattice.leaf.activation_replays_over_budget counter is tagged tree, partition, and the derived tenant, never by leaf - leaf count is unbounded, so it cannot be a metric dimension - which means the counter measures the rate and the log makes the per-leaf call.

What the warning reports, and the stall fault (issue #2149). The cost warning now carries the quantities it actually compared: the leaf's applied entry count (post-range-filter), the budget, and - so the two can be reconciled against the classifier - the partition head and gap. Before this, the line carried neither head nor gap, so the over-budget factor could not be read off the instrument at all; a dimensionless figure obtained by dividing an absolute WAL offset by the budget survived across three issues before a measurement run caught it. applied <= gap holds by construction, so a line violating it is reporting two different windows.

Evaluating "did the checkpoint advance?" by hand across a log is what found the livelocked leaf of issue #2165, and it is exactly what stopped working once benign lines outnumbered it 560:1. That evaluation is now performed in-process: a leaf that re-enters replay for the same partition from an unchanged persisted checkpoint is reported as a distinct stalled-replay fault, on its own throttle so cost noise can never suppress it. The first activation stays silent - one cold activation is not a stall - and every repeat at a frozen checkpoint warns. The two conditions are independent: a stalled leaf whose own work is small warns as a fault and not as a slow replay, which is the #2165 shape exactly.

On the healthy multi-partition path every partition's classifier returns TailReplay, and the leaf executes a two-pass replay across all partitions (per-partition Set / Delete absorption with TxCommit / TxAbort / DeleteRange deferred until every partition has populated its pending-tx record, then drained; a DeleteRange whose range cannot overlap the leaf's key range is instead consumed in the first pass, applying nothing - issue #3601) followed by a post-pass per-partition checkpoint reconciliation that advances each partition's ProjectionCheckpointOffsetsByPartition[p] to the highest applied offset once the saga-prepare clamp lifts.

The ProjectionRebuildPolicy enum on LatticeOptions is consulted only when trigger 1 fires, and every value currently fails closed:

Policy Behaviour
SnapshotThenWal (default) Intended to recover from the leaf's snapshot and then the WAL. The snapshot half is not specific to this policy: the per-leaf snapshot rehydrate runs at the start of every activation under every policy, before the classifier, covering the prefix and letting the tail replay handle the remainder. What is not yet integrated is a recovery for the case where that rehydrate has already declined and the WAL is genuinely short: there the leaf surfaces LeafProjectionStaleException rather than reconstructing the lost prefix.
FullRebuildFromWal Diagnostic. Intended to replay from the absolute tail of the WAL, but the policy is reached only when the WAL has been trimmed and a complete history is unavailable, so the leaf surfaces LeafProjectionStaleException here too; no full-rebuild recovery path is integrated.
Fail Surfaces a LeafProjectionStaleException at activation time and waits for an operator-driven rebuild.

This policy is reached only on genuine loss (the WAL trimmed past checkpoint + 1). A replay-budget or projection-age overrun against an intact WAL never consults it - see the trigger list above.

Starvation-drive admission

Background starvation drives share the same process-wide replay permits as leaf activations, but never queue for one. Across all trees, drives may hold at most half the replay permits in circulation - the configured replay ceiling less any that memory-pressure withholding is holding back - rounded down with a minimum of one (issue #3610). This leaves capacity for foreground and maintenance activations whenever more than one permit circulates. While only one does - a ceiling of one, or a larger ceiling at the withholding floor - drives and activations share that single permit.

Two callers request drives, and they do not compete for that share on equal terms (issue #3575):

  • the WAL GC's blocked-leaf sweep, the only caller that lifts a pin holding a tree's cursor floor, may use the whole share, and on a pass over the floor holders it has classified it touches the one nearest the floor first, alone, so a free permit goes to the floor rather than to whichever touch reaches the gate first (issue #3610);
  • a leaf's own coverage-lag timer, which drives a leaf that has never checkpointed or whose checkpoint has stopped advancing, never takes the last free slot: it is admitted only while at least two slots of the share are free, so however many drives hold the others, one slot is always free for the sweep. Where the share is a single slot there is nothing to reserve, so the timer instead leaves that slot to a sweep drive that was refused it, until the sweep is admitted again or five minutes have passed.

The timer reaches every stalled leaf on a fixed cadence, so without this it won the share by volume: on one deployment the sweep was refused about nine touches in ten, and the WAL of the trees it was trying to clear grew without being reclaimed.

When no permit is immediately available, or the drive's part of the GC share is occupied, the drive is refused before replay starts. The refusal is a result rather than an exception (issue #3761): the drive returns an admission-refused verdict, and the refusal is counted on orleans.lattice.saturation.refusals with source=replay_permit_admission and arm=gc_share. It neither advances nor retires the leaf's retention pin, and neither caller treats it as a fault:

  • the sweep records the try as outcome=admission_refused on orleans.lattice.wal.gc.blocked_leaf_reactivations (beside the drive's own drove_admission_refused verdict when the drive returned one), logs it at Debug without a stack, and does not count it as attempted or charge it against the consumer's attempt budget, so a consumer the sweep never managed to drive is never abandoned. A pass keeps no more touches in flight than this silo's part of the GC share, so its own touches do not refuse each other, and re-drives a refused touch once a sibling touch of the same pass frees a slot (issue #3761). A consumer still refused when the pass ends retries after a jittered delay of one to one and a half minutes that doubles with each consecutive refusal, up to its ordinary fifteen-minute cooldown;
  • the timer counts it as reason=recheck_drive_refused on orleans.lattice.leaf.snapshot.driver.declines and backs off: it skips its next drive opportunities, counting each as reason=recheck_drive_deferred - one after a first refusal, rising to four to seven after repeated ones, jittered per leaf. A drive that is admitted ends the backoff.

The activation queue's depth and drain policy are unchanged. Per-tree touch limits alone cannot bound the aggregate load of many trees on this process-wide gate (issue #3480).

A live leaf whose projection has gone stale

A leaf can find its projection stale while it is still activated. Its in-memory cache was built before the WAL was trimmed, so it keeps serving reads and taking writes, but its persisted checkpoint needs an offset the trim removed. Before latching this fault, the starvation drive attempts a conservative warm-cache rescue. This is not a cold rebuild: it writes a snapshot of the surviving cache only when that activation can prove the entire claimed prefix is present.

Eligibility starts with a successfully hydrated snapshot covering every configured partition. A successful replay retains its independently scanned, contiguous per-partition frontier; foreground checkpoint hints are never evidence for that frontier. Later hydration, reset, topology changes, failed or overlapping replay, and - for an unfiltered replay - gaps in the observed WAL sequence invalidate the proof. A filtered replay's slices omit other owners' records by design (issue #3565), so there only non-ascending offsets invalidate it, and the window is proven instead by checking, once the partition is read, that the oldest surviving WAL offset is still at or below the first offset the window needed. Each partition's activation anchor must be at or behind its persisted checkpoint, and its current checkpoint must not exceed the proven frontier. The live WAL must still cover everything after the proposed snapshot claim.

The drive also excludes unresolved transactions, active mutations, another capture, split/merge activity, retirement, and moved-away seals. It holds the topology gate across snapshot storage and checkpoint persistence; sealing and unsealing use that same gate. The snapshot claim is fixed before capture, and eligibility is rechecked immediately before copying the rows. Only a store-confirmed kept snapshot permits the checkpoint persist; the drive then awaits durable-pin publication before reporting success. A failure after the snapshot was kept propagates rather than claiming that nothing was saved. Storage is addressed using the leaf's bound tree identity, never an identity taken from a replayed mutation.

Every rescue decline emits one structured warning per reason per activation, with LeafId, TreeId, and typed DeclineReason fields:

Reason Evidence that is missing or blocks rescue
UnprovenBaseline No fully covering snapshot-seeded successful replay frontier.
CacheRehydratedOrReset The original cache was replaced or its replay barrier retired.
TopologyChanged This activation attempted a topology change after its baseline.
ReplayIncomplete Replay failed, overlapped, remains active, or has not satisfied its barrier.
UnknownPartition The stale partition or complete partition width is not established.
CheckpointUnproven A current checkpoint exceeds its independently proven frontier.
ActivationAnchorAhead A persisted checkpoint is behind the activation's snapshot anchor.
WalGapBeyondCache The oldest surviving WAL offset is beyond the claimed prefix plus one.
PendingTransactions Prepared, shadowed, or unresolved replay work remains.
MutationInFlight An admitted mutation has not finished.
CaptureInFlight Another snapshot capture owns the capture slot.
SplitInFlight A split is interrupted or the topology gate is occupied.
RetiredOrSealed The leaf is retired or holds moved-away slots.
CaptureDeclined Snapshot storage did not acknowledge a kept capture.
StorageFailure A probe or capture failed before kept coverage was acknowledged.

Cold, already-lost, partially snapshot-covered, or otherwise unproven caches still decline. This tree-agnostic path also covers derived trees, but does not trigger their re-derivers and cannot reconstruct lost data. The proof is activation-local and is never serialized or reconstructed from a persisted checkpoint. A cache already warm before this code was deployed has no recorded proof and declines with UnprovenBaseline. Restarting to deploy the change discards that cache; this path cannot recover it.

Two drivers still reach such a leaf: the coverage-lag timer, which drives a leaf whose checkpoint has stopped advancing, and the WAL GC's blocked-leaf sweep. If the rescue declines, the first starvation drive logs one Error naming the leaf and tree, and latches that verdict for the activation (issue #3450). While the persisted checkpoints are unchanged:

  • the timer skips the drive entirely, so a stale leaf no longer takes a replay permit or raises a timer fault once per stall window;
  • a drive from the WAL GC's blocked-leaf sweep still receives LeafProjectionStaleException, without another replay. The sweep treats that verdict as terminal (issue #3478): it records it once as outcome=latched_stale on orleans.lattice.wal.gc.blocked_leaf_reactivations, logs it at Information, and stops driving the leaf for the rest of the tree's blocked episode. The leaf's pins stay in the WAL GC cursor floor, so no WAL it still needs is trimmed, and the tree is reported once per sweep pass with a Warning naming how many latched stale leaves hold it.

Any change a genuine apply or an operator reset makes to a persisted checkpoint clears the latch, and a new activation starts without it. A split's checkpoint hint does not: a hint for the partition the drive found stale is refused and logged as a Warning, because stamping it would persist a checkpoint past the trimmed range the leaf never applied, and would clear the latch with nothing repaired (issue #3477). The leaf keeps its WAL retention pin.

Treat the Error as data at risk rather than as noise. The live activation may hold the only copy of writes in the trimmed range. Once it is recycled or the silo restarts, the next cold activation refuses the leaf with the same exception. RebuildLeafProjectionAsync does not recover those writes: it resets the checkpoint, and the next activation reloads the leaf's snapshot where it has one and replays only the WAL that survives. So, before the activation is lost:

  1. Capture a logical backup or export of the tree while the leaf is still serving.
  2. Deploy a build that fixes the cause of the over-trim.
  3. Restore the tree from that backup. For a tree that is derived from another source, delete it and re-derive it instead.

Configuration

siloBuilder.ConfigureLattice(o =>
{
    // Expected replay size before a cold activation is flagged as
    // over-budget (a warning plus a counter - never a failure):
    o.MaxLeafReplayEntries = 100_000;

    // Age at which a cold projection would be flagged as stale. The
    // activation path does not yet supply an age, so this does not
    // fire today (see trigger 3 above):
    o.LeafProjectionRetention = TimeSpan.FromDays(30);

    // How to recover when the WAL is genuinely short:
    o.ProjectionRebuildPolicy = ProjectionRebuildPolicy.SnapshotThenWal;
});

Snapshot-on-fall-off safety net

The three activation-time triggers above react to a fall-off-log condition after it has already happened. The snapshot-on-fall-off path is the preventative safety net: while a leaf is still healthy, it captures a canonical-row image of its in-memory cache to a dedicated snapshot grain whenever any partition's WAL tail approaches that partition's persisted checkpoint. On the next activation the leaf rehydrates its cache from the blob rows and tail-replays each partition forward from its captured offset. It does so when the snapshot is newer than the persisted partition-0 checkpoint, and also when the snapshot is at or behind it, provided the cache starts empty or any partition's WAL prefix has been trimmed: a snapshot may then be the only durable copy of a prefix the WAL GC trimmed, and each partition's checkpoint is lowered to what the reloaded cache holds before the tail replay. Activation then proceeds without ever needing to fall back into SnapshotThenWal / FullRebuildFromWal / Fail recovery, even when the WAL has been trimmed past the original checkpoint.

The capture path is leaf-driven, not maintenance-driven:

  • At activation, the leaf runs the fall-off-log detector once per partition. When no trigger has fired - not even one of the two cost signals - but a partition's persisted checkpoint sits within the oldest LeafSnapshotMargin fraction (default 0.30) of that partition's readable WAL window, the detector returns a non-fatal snapshot advisory. The leaf latches the advisory, finishes its tail replay across every partition, and then captures a single snapshot before yielding the activation turn.
  • While the leaf remains hot, every LeafSnapshotReClassifyEveryNCheckpoints (default 64) successful checkpoint persist re-runs the classifier and drives another capture on advisory. Pass 0 to disable the periodic recheck entirely.
  • Independently of the advisory and of that setting, a leaf holding a checkpointed partition that no snapshot covers yet captures one at activation and on its later checkpoint persists and coverage-lag checks, until a capture succeeds, within a per-activation attempt budget that re-arms after a backoff (issue #2692). A capture covers every partition checkpointed at the time, so a partition that checkpoints later brings the leaf back once more, and an attempt that covered at least one partition is not charged against the budget: only captures that make no progress spend it (issue #3576). This is what gives a tree that has stopped taking writes snapshot coverage on its next activation.
  • A graceful deactivation also captures one for any checkpointed partition that no snapshot covers yet, before it publishes its final durable pin, so a short-lived activation that never reached the periodic cadence still leaves coverage behind.
  • A checkpoint persist whose snapshot recheck lands a capture publishes the durable pin again after that recheck, so the capture reaches the pin at once instead of being left to the later frontier_pin barrier, which a deactivation deadline can skip. The final persist of a graceful deactivation also publishes its pin before the recheck, so a recheck that overruns the deadline cannot cost the pin. The coverage-lag check also republishes a pin that has fallen below min(persisted checkpoint, coverage), first committing a pending checkpoint advance once MaterialiserCheckpointInterval has elapsed so a write-idle leaf's advance cannot stay pending (issue #3608), and the WAL GC's blocked-leaf sweep asks a floor-holding leaf for the same step before it spends a replay permit on a drive. That step takes no replay permit and replays nothing, and a capture that fails or is declined leaves coverage, and so the pin, where they were (issue #3599).
  • A single-flight guard suppresses overlapping captures: a slow SaveAsync does not pin a follow-on capture behind it; the follow-on is dropped and the next cadence tick re-evaluates.

Each capture overwrites the previous blob; only the most recent snapshot is retained per leaf. The WAL remains the long-term audit trail.

Cold-leaf limitation

A leaf that never activates while drifting below the WAL retention window will not be captured by this path. Such a leaf also holds no in-memory state to lose; if the WAL has been trimmed past its checkpoint with no covering snapshot, its next activation refuses it with LeafProjectionStaleException, whatever the configured ProjectionRebuildPolicy. The snapshot-on- fall-off path is a safety net for active leaves.

Resumable cold replay

The activation-time WAL replay that rebuilds a leaf's projection cache is resumable. A leaf rebuilds its cache from the WAL on every activation, but the persisted ProjectionCheckpointOffset is what survives across activations. Historically the checkpoint was advanced only by a single reconciliation step at the very end of the replay, so a leaf whose un-snapshotted prefix could not be drained inside one activation window - a large WAL relative to Orleans' ~30s RuntimeRequested deactivation budget - made no durable progress: a mid-replay deactivation discarded every applied entry, the next activation restarted from the same offset, and because the coverage-gated WAL GC (correctly) refuses to trim an un-snapshotted prefix, the WAL grew without bound while the leaf never converged.

The replay now flushes the checkpoint incrementally, at each replay slice boundary, over the strictly contiguous, fully-applied prefix. Because a checkpoint persist drives the periodic snapshot recheck above, this also captures an incremental snapshot as the replay progresses. A mid-replay teardown therefore loses at most one flush interval; the next activation rehydrates from the incremental snapshot and resumes from the last durable offset instead of replaying from zero, and the now snapshot-covered prefix becomes trimmable so retention stays bounded.

The incremental advance can never license a checkpoint (or the materialiser pin) past work that is not yet durable. The replay defers two kinds of record:

  • a deferred saga terminal (TxCommit / TxAbort) or DeleteRange, which is applied only in the replay's second pass (a range delete that cannot overlap the leaf's key range is not deferred: the first pass consumes it, so it takes no ledger slot); and
  • an unresolved saga prepare, whose pending-transaction bucket no snapshot captures because the matching terminal is itself deferred.

Since issue #2165 the leaf writes each such record verbatim into a ledger on its persisted state row before the flush that advances past it, so the record and the checkpoint it licenses land in the same state write, and a resumed activation reconstructs the work from the ledger instead of re-reading it from the WAL. For deferred terminals the ledger is bounded by MaxDurableUnresolvedReplayWork (default 1 024); past that bound the flush falls back to holding the checkpoint below the deferred entry until the second pass applies it. An unresolved prepare is always recorded (issue #2183), because nothing drains a prepare whose saga never terminates.

Steady-state activations are unaffected: a single-slice replay with no deferred terminals still flushes exactly once at the end of the first pass, the same net timing as the old end-of-replay reconciliation. Only a multi-slice replay of a large prefix sees the extra intermediate persists - which is exactly the case the resumability guarantees.

Configuration

siloBuilder.ConfigureLattice(o =>
{
    // Trigger a proactive snapshot capture when the leaf's
    // persisted checkpoint is within 30% of the WAL tail. Lower
    // values reduce snapshot frequency; raise to capture earlier.
    o.LeafSnapshotMargin = 0.30;

    // While a leaf stays hot, re-run the fall-off classifier every
    // N successful checkpoint persists and re-capture on advisory.
    // Set to 0 to disable the periodic recheck entirely.
    o.LeafSnapshotReClassifyEveryNCheckpoints = 64;
});
  • ILattice.GetLeafProjectionDigestAsync - the public surface.
  • ILattice.GetLeafProjectionDigestForRangeAsync - the range-scoped analogue (null bounds give the whole-shard digest) that backs the cross-cluster Merkle-walk localisation.
  • LeafProjectionDigest - the returned readonly record struct.
  • LatticeOptions.MaintainProjectionDigest - opt out of the per-mutation XOR fold and upward publication for digest-indifferent workloads.
  • ProjectionRebuildPolicy - the activation-time recovery policy.
  • LatticeOptions.MaxLeafReplayEntries, LatticeOptions.LeafProjectionRetention, LatticeOptions.MaterialiserCheckpointInterval, LatticeOptions.MaterialiserCheckpointEntries, LatticeOptions.LeafSnapshotMargin, LatticeOptions.LeafSnapshotReClassifyEveryNCheckpoints - see Configuration.
  • LeafProjectionStaleException - surfaced on genuine loss under every ProjectionRebuildPolicy, including the default SnapshotThenWal, whose post-rehydrate recovery is not yet integrated; also rethrown by a starvation drive against a leaf latched stale.
  • ILattice.RebuildLeafProjectionAsync and ILattice.GetMaterialiserLagAsync - see Operator tooling: rebuild and lag.

Operator tooling: rebuild and lag

Activation-time ProjectionRebuildPolicy recovers a leaf when it cold-starts and a fall-off-log trigger fires. Two complementary surfaces let an operator drive recovery and observe materialiser health without waiting for an activation:

  • ILattice.RebuildLeafProjectionAsync(int shardIndex, CancellationToken)
  • ILattice.GetMaterialiserLagAsync(CancellationToken)

Rebuild a shard's projection from the WAL

// GetLeafProjectionDigestAsync reported a mismatch, or an
// integrity check flagged a leaf's projection as suspect. Force every
// leaf in the shard to re-materialise its projection on its next
// activation: from its snapshot where it has one, then from the WAL.
await tree.RebuildLeafProjectionAsync(shardIndex: 0, cancellationToken);

RebuildLeafProjectionAsync walks every leaf in the named physical shard via the sibling chain and, for each leaf, clears only the projection-state slots that the materialiser owns:

  • The per-activation entry cache is dropped when the grain deactivates (the cache is never persisted, so there is nothing to clear on the state row itself).
  • The persisted projection checkpoint is reset to the -1 "nothing applied" sentinel (matching the WAL reader's fromOffsetExclusive = -1 start-of-log convention) and the per-partition checkpoints are dropped, so a partition that no snapshot covers is replayed from offset 0 inclusive on the next activation. Setting 0 instead would cause the materialiser to skip offset 0, because replay reads strictly past the persisted checkpoint.
  • The persisted running projection hash is cleared.
  • In-memory pending-saga, pending-tx-offset, recently-terminal, and backstopped-terminal dedup buffers are dropped, together with the destination-side shadow markers.
  • The leaf grain is deactivated. The next activation re-materialises the projection through the standard activation-time path, which under every ProjectionRebuildPolicy value first attempts to rehydrate from the leaf's captured snapshot. The rebuild neither clears nor bypasses that snapshot: when the leaf has a usable one, its cache is reloaded from it, and each partition the snapshot covers replays only the WAL entries after the snapshot's captured offset.

Topology-bearing state is preserved: TreeId, ShardIndex, the leaf's key range, and the sibling pointers stay intact. The rebuild does not re-shape the tree - it only re-derives the materialised projection for the range the leaf already owns. Only the part that no snapshot covers is re-derived from the WAL: a snapshot-covered prefix is restored from the snapshot rather than re-applied, so a row the snapshot captured wrongly - after a corruption or a projection bug - survives the rebuild unless a later WAL entry for that key replaces it.

RebuildLeafProjectionAsync does not take a tree-wide consistency lock. Readers and writers continue to land on the shard during the rebuild; in-flight writes hit the leaf's standard write path (which goes through the WAL) and will be visible after the next activation re-replays them. A read that reaches a rebuilt leaf waits for that leaf's activation replay to finish rather than seeing a partial projection, so operators rebuilding under load should expect a window of higher read latency rather than missing entries; a replay that fails fails the waiting read, and the next request retries it. Pair the rebuild with a digest re-poll after replay stabilises to confirm the projection converged.

Error surface:

Condition Exception
shardIndex is not a physical shard of the per-tree map ArgumentOutOfRangeException
Tree id starts with the reserved system prefix _lattice_ LatticeReservedTreeNamespaceException (an InvalidOperationException subclass)
cancellationToken was already cancelled OperationCanceledException
An access gate is configured and does not authorise the caller for whole-tree admin LatticeAuthorizationDeniedException

Observe materialiser lag

long lag = await tree.GetMaterialiserLagAsync(cancellationToken);
// An estimate of the WAL entries the worst shard's leaves have not yet
// folded into their projections. Each WAL head is the next offset to be
// assigned and each checkpoint the last offset applied, so even a
// caught-up shard with a non-empty WAL reports a small positive value:
// read a steady value as caught up and a growing one as falling behind.

GetMaterialiserLagAsync returns the maximum lag across all physical shards of the tree. For each shard the lag is the sum, over the tree's WAL partitions, of

walHead[p] - min(checkpointOffset across leaves in the shard)

with each term clamped at zero so a checkpoint that has temporarily raced ahead of the head observation (e.g. between the head fetch and the per-leaf checkpoint fetch) cannot contribute a negative value. walHead[p] is the next offset partition p will assign - one past its newest entry - while a checkpoint is the offset of the last entry a leaf applied, so a partition whose newest entry the checkpoint has reached still contributes 1; the result is 0 only when every term clamps to zero, as on an empty WAL. The per-leaf checkpoint read is each leaf's partition-0 checkpoint, which the reduction applies to every partition's head as an approximation, so on a multi-partition tree the figure is an estimate rather than an exact count of unapplied entries; a shard with no leaves reports the sum of its partition heads. The result is the worst-shard lag because a single slow shard is the SLO-relevant signal - averaging it would mask the actual problem.

The intended monitoring shape is a periodic poll (e.g. every 5-30 s) feeding a gauge in the telemetry pipeline:

// Pseudo-code for an operator polling loop. The interval and
// thresholds are deployment-specific - they scale with WAL
// ingestion rate and the SLO for read-after-write recency.
long lag = await tree.GetMaterialiserLagAsync(cancellationToken);
if (lag > 10_000)
{
    // Materialiser is falling more than 10 000 entries behind the
    // WAL on at least one shard. Investigate slow leaf activation,
    // a stuck replay coordinator, or backpressure on the storage
    // provider.
}

A growing lag indicates the materialiser is not keeping up with WAL ingestion. Common causes are slow leaf activation under storage-provider backpressure, a stuck ILeafReplayCoordinatorGrain, or a deactivation storm cycling leaves faster than they can replay. A persistent lag at a small positive value (a few entries) is expected: each non-empty partition contributes at least 1 even when caught up, and under sustained write load the materialiser checkpoints in batches, so the most recently published WAL entries naturally lag the head briefly.

Error surface:

Condition Exception
Tree id starts with the reserved system prefix _lattice_ LatticeReservedTreeNamespaceException (an InvalidOperationException subclass)
cancellationToken was already cancelled OperationCanceledException
An access gate is configured and does not authorise the caller to read the whole tree LatticeAuthorizationDeniedException

When to use rebuild vs. activation-time recovery

Scenario Surface
Leaf cold-starts and the WAL has been trimmed past its checkpoint with no covering snapshot (genuine loss) Activation refuses the leaf with LeafProjectionStaleException under every ProjectionRebuildPolicy (the automatic recovery paths are not yet integrated): restore the tree from a backup, or accept the loss of the trimmed range and run RebuildLeafProjectionAsync
Leaf cold-starts with a replay gap over MaxLeafReplayEntries with the WAL intact Non-fatal: the leaf tail-replays, and warns only when the entries it actually applies exceed the budget; no operator action needed. The LeafProjectionRetention age trigger does not fire from activation today
Operator detects a digest mismatch across silos, or an integrity check flagged a corrupted projection, or a fix to how WAL entries are applied to the projection requires re-materialisation RebuildLeafProjectionAsync (manual, while live). It re-applies only the WAL no snapshot covers: a prefix the leaf's snapshot covers is restored from the snapshot, so a row the snapshot captured wrongly survives unless a later WAL entry for that key replaces it
Operator wants a steady-state gauge to know whether the materialiser is keeping up GetMaterialiserLagAsync

The two paths share the same replay seam: RebuildLeafProjectionAsync clears state and lets the standard activation-time path do the re-materialisation. There is no second, parallel rebuild code path to maintain - the operator surface is a controlled trigger for the existing recovery logic.