---
title: "Leaf-Projection Rebuild & Digest"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/projection-rebuild.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/projection-rebuild.md"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Leaf-Projection Rebuild & Digest

Part of the [Orleans.Lattice documentation](architecture.md).

Orleans.Lattice's partitioned write-ahead log (WAL) is, in a fully replicated
deployment, the canonical durable record of every leaf mutation. Each leaf
grain materialises that log into a per-activation in-memory projection
(the entry cache - a sorted dictionary owned by the leaf grain for the
lifetime of the activation, not persisted). The persisted leaf state row
carries no entries: it holds topology, the per-partition projection
checkpoint offsets, the HLC clock and version vectors, a 16-byte
projection-digest XOR fold, and a little replay bookkeeping. On every
activation the cache is rebuilt: the leaf reloads its latest snapshot where
it has a usable one and replays each WAL partition from just past the offset
that snapshot covers, or from the start of the partition's readable window
when no snapshot covers it (see
[Snapshot-on-fall-off safety net](#snapshot-on-fall-off-safety-net)). Two
operational concerns naturally arise:

1. **Drift detection.** If a silo's leaf projection diverges from the
   WAL prefix it claims to have applied - a cosmic-ray bit flip, a
   storage-provider read-after-write anomaly, a bug in `ILeafProjection.Apply` -
   how does an operator notice before downstream readers do?
2. **Recovery from WAL trim.** If a leaf has been cold long enough that the
   WAL has been trimmed past its last persisted checkpoint, the leaf cannot
   resume by tail-replay alone. What does activation do?

This document covers the two surfaces that answer those questions:
`ILattice.GetLeafProjectionDigestAsync` (drift detection) and
`ProjectionRebuildPolicy` (recovery from genuine loss), together with
the `MaxLeafReplayEntries` / `LeafProjectionRetention` cost thresholds
that govern when a long or stale cold replay is flagged.

## Drift detection: `GetLeafProjectionDigestAsync`

```csharp verify
LeafProjectionDigest digest = await tree.GetLeafProjectionDigestAsync(
    shardIndex: 0,
    cancellationToken);

// digest.Hash             - 16-byte XxHash128 fingerprint of the shard's projection
// digest.EntryCount       - entries (live + tombstoned) folded into the hash
// digest.CheckpointOffset - highest per-leaf projection-checkpoint offset
// digest.Version          - contribution-function version; compare hashes only
//                           between digests that carry the same Version
```

`GetLeafProjectionDigestAsync` reads the requested **physical shard**'s
root and returns a pre-folded digest in `O(1)` grain hops at the shard
level. The shard's internal-node root maintains a running
**XOR-fold** over every descendant leaf's running per-leaf hash. Each
internal node persists the latest digest snapshot (hash, entry count,
checkpoint offset) of each of its children and re-derives its subtree
aggregate from that table - the XOR of the child hashes, the sum of
their entry counts, and the maximum of their checkpoint offsets -
whenever a child publishes a fresh snapshot upward. A leaf publishes
after its mutations persist, and mutations that land within
[`DigestCoalescingWindowMs`](configuration/options-reference-1.md#digestcoalescingwindowms)
(default 5 ms) share one publish; structural changes such as splits and
tombstone reaps publish immediately. The public shard hash
is

```text
XxHash128( xor_subtree_hash || subtree_entry_count || subtree_max_checkpoint )
```

where `xor_subtree_hash` is the bitwise XOR of every descendant leaf's
16-byte running hash. A single-byte difference at any leaf - a stale
tombstone, a missing TTL stamp, a divergent vector clock - surfaces as
a different shard hash. Operators running multiple silos against the
same WAL can poll the digest from each silo and compare bytes;
equality is the strongest possible cross-silo state-equivalence check
the library provides.

When the shard's root is a single leaf (flat-tree case, no internal
node yet exists), the digest is read directly from the root leaf. A
never-written shard is no exception: the read first materialises its
root leaf, so an empty shard reports `EntryCount = 0` and
`CheckpointOffset = 0` with a hash equal to the XxHash128 of 32 zero
bytes (the all-zero 16-byte fold followed by the two zero counters).

XxHash128 is a non-cryptographic hash: it is chosen for ~10x lower CPU
cost than SHA-256 on the per-mutation hot path and for its uniformly
distributed output (which the XOR-fold algebra requires). The digest is
a drift-detection fingerprint, not an authentication tag - a malicious
operator with write access to the projection state could craft a
collision, but the digest's job is to catch silent corruption, not to
defend against forgery.

### What is folded into the leaf hash

For every entry in the leaf's in-memory entry cache (a sorted dictionary keyed
with `StringComparer.Ordinal`, rebuilt from the WAL on every activation) the
implementation computes a 16-byte XxHash128 contribution over the following
fields, in this order:

1. `key` (length-prefixed UTF-8)
2. `lww.Timestamp.WallClockTicks` (`Int64`, little-endian)
3. `lww.Timestamp.Counter` (`Int32`, little-endian)
4. `lww.IsTombstone` (`byte`, `0x00` or `0x01`)
5. `lww.ExpiresAtTicks` (`Int64`, little-endian - `0` when unset)
6. `lww.OriginClusterId` (length-prefixed UTF-8, `-1` sentinel for null)
7. `lww.VectorClock` (a deterministic ordinal-sorted feed of every
   `(replicaId, hlc.WallClockTicks, hlc.Counter)` triple, or `-1` sentinel
   when null/empty)
8. `lww.Value` (length-prefixed bytes - `-1` sentinel for tombstones)

The per-entry contributions are XOR-folded into a 16-byte running hash
that is **maintained incrementally on every mutation** and persisted on
the leaf state row as `LeafNodeState.ProjectionHash`. Insert XORs the new contribution in; replace XORs the
old contribution out and the new one in (the old contribution cancels
under self-inverse XOR); delete XORs the contribution out. Because XOR
is commutative, associative, and self-inverse, the running hash is
independent of insertion order and idempotent re-application of the
same mutation is a no-op - exactly the algebra LWW already provides
for entry state.

The public per-leaf digest is the XxHash128 of `(running_xor ||
entryCount || checkpointOffset)`, where `checkpointOffset` is the leaf's
partition-0 projection checkpoint (the legacy scalar slot), so two silos
at different partition-0 replay positions report distinct digests even
if their post-state happens to coincide; on a multi-partition tree the
other partitions' positions are not folded in. The shard-level aggregate then XOR-folds each descendant
leaf's `running_xor` directly (no per-leaf XxHash chaining step) and
applies the same `(xor_subtree_hash || subtree_entry_count ||
subtree_max_checkpoint)` framing at the root. The XOR-fold makes the
shard aggregate commutative and self-inverse, which is what lets each
internal node maintain it incrementally as children publish updates
upward.

### Determinism contract

The digest is byte-stable across silos because every input is canonicalised:

- The leaf's entry cache is a `SortedDictionary<string, LwwValue<byte[]>>`
  built with `StringComparer.Ordinal`, so the per-entry contributions
  are identical on every silo regardless of insertion order.
- All numeric fields use little-endian framing via `BinaryPrimitives`.
- All strings use `Encoding.UTF8`, length-prefixed with an `Int32`.
- Length-prefix sentinels (`-1`) distinguish tombstone from empty value
  and null-string from empty-string so adjacent variable-length fields
  cannot collide.
- `VersionVector` keys are sorted with `StringComparer.Ordinal` before
  feeding so dictionary insertion order does not perturb the output.

### Topology changes and the aggregate

The internal-node aggregate is maintained incrementally as children
publish `ChildDigestSnapshot` updates upward, so the aggregate's
correctness depends on a single invariant: **each child contributes to
exactly one parent at any instant**. A B+ tree split moves a contiguous
half of a node's children to a new sibling, which transiently violates
that invariant if the moved children's per-child digest rows are left
behind on the donor or if a moved child keeps publishing to its former
parent. Both would double-count the moved subtree's entries in the
shard total.

The split path preserves the one-parent invariant in two steps:

1. **Prune on the donor.** When an internal node splits, it removes the
   moved children's rows from its persisted per-child digest table and
   recomputes its `SubtreeProjectionHash`, `SubtreeEntryCount`, and
   `SubtreeHighestCheckpointOffset` from the remaining rows before
   publishing the corrected aggregate upward. The XOR fold's
   self-inverse algebra makes the recompute exact - the moved rows
   cancel cleanly out of the running hash.
2. **Reject stale publishes.** Each internal node folds a child's
   digest snapshot only from a child it currently owns. A
   snapshot arriving from a child that has already been re-parented to
   the new sibling is rejected and its stale row (if any) dropped, so a
   moved child that races a publish against its re-parenting cannot
   reintroduce a double count. Every publisher also stamps a
   monotonic sequence, and a snapshot older than the one already folded
   for that child is dropped, so a late coalesced publish carrying a
   pre-split count cannot overwrite a fresher one. A donor likewise
   re-seeds a child's
   parent pointer only for children it still owns, so a moved child is
   never pointed back at the node it left.

The net effect is that `EntryCount` stays exactly equal to the number
of distinct entries (live plus tombstoned) under the shard across an
arbitrary sequence of internal-node splits, with no transient
over- or under-count visible to a quiescent digest read.

The upward publish that maintains the aggregate is a cross-grain RPC
that recurses up the internal-node chain. An internal node never makes
it while holding its own non-reentrant split gate, and a parent whose
gate is busy - for example mid-split - parks the incoming snapshot,
keeping only the freshest one per child, and folds it before it
releases the gate, so a publish never waits on its parent's split
(issue #3523). A parent that is itself mid-mutation can still leave the
await neither completing nor faulting, so every upward publish, from a
leaf or an internal node, is bounded by
`LatticeOptions.DigestPublishTimeout` (default 15 s): on the deadline
the publish is abandoned and a `TimeoutException` is raised, with no
count drift - the abandoned publish never partially applied at the
parent, the publisher keeps its snapshot marked pending, and the next
mutation's publish re-drives convergence. Where an internal node has
already made its own change durable - accepting a split, or removing a
reclaimed child - the timeout is logged and contained rather than
surfaced, so that change is not lost. Set the option to
`InfiniteTimeSpan` to restore the historical unbounded await. A
non-zero `orleans.lattice.internal.digest_publish.timeouts` counter
surfaces the condition; it counts internal-node publishes only.

### Cost and where to call it

Because the per-entry XOR fold is maintained incrementally on every
mutation, `GetLeafProjectionDigestAsync` does **not** re-walk the
leaf's entry cache on each call - the running hash is already on the
leaf's persisted state, so the per-leaf computation collapses to a
single fixed-size XxHash128 over `(running_xor || entryCount ||
checkpointOffset)`. The shard root delegates to the root internal
node, which returns its persisted subtree aggregate in a single grain
hop without re-visiting any descendant. The leaves themselves are not
activated by the digest poll: each leaf published its contribution
upward after its last mutation persisted (within one
`DigestCoalescingWindowMs`), and the internal-node aggregate is the
source of truth at read time.
A whole-tree poll therefore costs `O(shardCount)` grain hops,
regardless of how many leaves each shard owns or how many entries
each leaf holds.

The cold-start path remains correct: if the shard root or any internal
ancestor is activated for the first time, its persisted state is
loaded from storage along with the aggregate it already stamped on the
previous shutdown - no leaf walk is required to reconstruct it.

Heap allocations on the hot path are bounded:

| Allocation                              | Per call    |
|-----------------------------------------|-------------|
| `XxHash128` for per-entry contributions (one cached per leaf grain activation) | reused via `TryGetHashAndReset` |
| `XxHash128` for the outer digest framing | one per digest read, at the leaf or the internal-node root |
| `byte[16]` XxHash128 hash from `GetHashAndReset()` | unavoidable (the result) |
| `byte[16]` hash clone carried by each upward digest publish | bounded by tree height; cloned so subsequent XOR updates do not retroactively mutate the parent's captured bytes |
| String / VC scratch buffers             | pooled (`stackalloc 256` fast path; `ArrayPool<byte>.Shared` and `ArrayPool<string>.Shared` for the rare overflow) |

The `O(shardCount)` per-tree cost makes the digest cheap enough for
steady-state monitoring - including periodic cross-silo equality
canaries - not just on-demand diagnostics. It is safe to call against
a live shard under load: it observes the current in-memory projection
without taking any kind of consistency freeze. The result is
necessarily a snapshot at one wall-clock instant, however, so two
calls under sustained writes will report different digests; equality
is meaningful only between **quiescent observations** (no in-flight
writes to the shard between the two reads being compared, and at least
one `DigestCoalescingWindowMs` elapsed since the last write so every
coalesced publish has landed).

### Cross-silo divergence example

```csharp verify
// On every silo hosting the cluster, schedule a periodic poll
// over every shard and compare digests. A mismatch is a conservative
// trigger to investigate, not proof of drift: compare only quiescent
// reads carrying the same Version, because digests also differ when
// two reads were taken at different replay positions.
var routing = await tree.GetRoutingAsync();
foreach (var shardIndex in routing.Map.GetPhysicalShardIndices())
{
    LeafProjectionDigest digest = await tree.GetLeafProjectionDigestAsync(
        shardIndex,
        cancellationToken);
    // emit (silo, treeId, shardIndex, digest.Version, digest.Hash, digest.EntryCount,
    // digest.CheckpointOffset) to your telemetry pipeline.
}
```

### Error surface

| Condition                                                    | Exception                          |
|--------------------------------------------------------------|------------------------------------|
| `shardIndex` is not a physical shard of the per-tree map     | `ArgumentOutOfRangeException`      |
| The tree id starts with the reserved system prefix `_lattice_` | `LatticeReservedTreeNamespaceException` (an `InvalidOperationException` subclass) |
| `cancellationToken` was already cancelled                    | `OperationCanceledException`       |
| Tree has `LatticeOptions.MaintainProjectionDigest = false`   | `InvalidOperationException`        |
| An access gate is configured and does not authorise the caller to read the whole tree uniformly (a digest cannot be narrowed per key, so a partial allow is refused too) | `LatticeAuthorizationDeniedException` |

### Opting out of digest maintenance

The digest's maintenance cost is small in absolute terms, but it recurs
with the write load. Each leaf mutation costs one in-memory XOR fold over
the entry's contribution. When the leaf's digest has changed, the leaf
publishes it upward: the parent internal node rewrites its persisted
subtree aggregate and publishes to its own parent in turn, up to the
shard root, so each publish costs `O(treeHeight)` writes. With
[`DigestCoalescingWindowMs`](configuration/options-reference-1.md#digestcoalescingwindowms)
at `0` every such mutation publishes. By default (5 ms) the foreground
writes that land within one window - sets, deletes, range deletes and
typed CRDT delta applies - share a single publish, while merge traffic
and structural changes such as splits, tombstone reaps and saga
terminals publish immediately. For trees
that do not poll the digest - workloads that rely exclusively on
audit logs, integration tests, application-level checksums, or
external reconciliation, and never call
`GetLeafProjectionDigestAsync` - the maintenance cost is pure write
amplification.

`LatticeOptions.MaintainProjectionDigest` (default `true`) flips the
behaviour off:

```csharp verify
siloBuilder.ConfigureLattice(opts =>
{
    // Turn off digest maintenance globally - leaf mutations stop
    // updating the running XOR fold and stop publishing
    // ChildDigestSnapshot upward to internal-node ancestors.
    opts.MaintainProjectionDigest = false;
});

// Or per-tree:
siloBuilder.ConfigureLattice("audited-tree", opts =>
{
    opts.MaintainProjectionDigest = false;
});
```

When the opt-out is in effect:

- Leaf-mutation funnels (`StoreEntry` / `RemoveEntry`) take a trimmed
  path that LWW-merges the value, bumps the delivery sequence, and
  returns without touching the persisted `ProjectionHash`.
- The leaf does not publish `ChildDigestSnapshot` upward, so no
  internal-node ancestor updates its `SubtreeProjectionHash` for that
  mutation. The whole upward chain is quiescent.
- `ILattice.GetLeafProjectionDigestAsync` throws
  `InvalidOperationException`. When the opt-out comes from the configured
  options it fails fast at the public surface, before any routing-table
  fetch or grain hop. That check does not read the registry, so an opt-out
  set only through the tree's registry override or the latch below is
  caught after the shard hop, by the same check in the leaf and internal
  grains, which also stops a direct grain-handle caller.
- Persisted state is **not** rewritten. Any `ProjectionHash` already on
  disk from a previous-enabled period remains untouched.

#### Cross-cluster impact: anti-entropy drift detection

In a deployment running `Orleans.Lattice.Replication`, the anti-entropy
peer digest probe reads this same leaf-projection digest to detect
silent divergence between clusters. A tree with
`MaintainProjectionDigest = false` (or one whose registry latch has
disabled it permanently) has no digest to compare, so the probe skips
that tree and classifies a peer in the same state as `RemoteUnavailable`
rather than a mismatch. The whole automatic drift-detection-and-
remediation stack - the probe, the Merkle-walk localisation, targeted
leaf re-replay, and the bootstrap-snapshot fallback - is therefore inert
for any tree that opts out of digest maintenance. Disable the digest
only for trees you do not need cross-cluster drift telemetry on; see the
[automatic drift-remediation playbook](../lattice.replication/automatic-drift-remediation.md)
for what the stack provides and how to opt in.

#### Disabling is a one-way operation per tree

The first mutation that lands while maintenance is disabled stamps an
irreversible registry latch on the tree, reported as
`TreeConfigurationReport.ProjectionDigestPermanentlyDisabled` by
`ILatticeTreeAdmin.GetTreeConfigAsync` (see
[Orleans.Lattice.Api.TreeAdmin](../lattice.api.treeadmin/README.md)).
The stamp is best-effort: a registry failure never fails the mutation,
and the leaf retries the stamp on its next mutation while maintenance is
still disabled, so the latch is set only once a stamp succeeds.
Once the latch is set, every subsequent activation resolves
`MaintainProjectionDigest` as `false` regardless of the per-tree
override or the silo-wide default, and
`ILattice.GetLeafProjectionDigestAsync` keeps throwing.

The latch exists because the digest is an XOR-fold aggregate over
**every** mutation: any mutation accepted while maintenance was off
permanently invalidates the persisted aggregate, and silently
re-engaging maintenance would publish a known-stale digest as if it
were authoritative. Once stamped, the one-way latch makes this impossible
to mis-configure: an operator who turns the option back on for a tree
that has already accepted writes under the disabled setting will see
the resolved value stay at `false` and the digest API stay broken,
rather than producing a digest that disagrees silently with the
ground-truth entries.

The only way to re-engage digest maintenance for a latched tree is to
rebuild the tree (or its leaf range) from scratch under a fresh
registry entry. If you anticipate needing the digest later, leave it
enabled.

#### Per-tree precedence and system trees

Resolution order for `MaintainProjectionDigest`:

1. **System-tree prefix override.** Trees whose id begins with the
   reserved system prefix `_lattice_` (e.g. the internal registry
   tree) always resolve as `false` regardless of configuration.
   System trees are not replicated and have no cross-silo
   drift-detection consumer, so the maintenance work is pure
   overhead.
2. **Registry latch.** If `ProjectionDigestPermanentlyDisabled` is
   set, the resolved value is `false`.
3. **Per-tree override.** If the tree's registry override
   (`TreeConfigurationReport.MaintainProjectionDigest`, written through
   `ILatticeTreeAdmin.SetTreeConfigAsync`) is set, that value wins over the
   configured options of step 4. Operators can
   opt an individual tree out (or, while the latch is not yet set,
   back in) without flipping the silo-wide default.
4. **Configured options.** Falls back to `LatticeOptions.MaintainProjectionDigest`
   as configured for the tree: a named `ConfigureLattice(treeName, ...)`
   override when the host set one, otherwise the silo-wide value.

Disabling the digest is recommended for **write-amplification-sensitive
deployments that do not need cross-silo drift telemetry**. Keep it
enabled when you operate multiple silos against the same WAL and rely
on the digest as a state-equivalence canary, or when chaos / soak
tests use the digest as a post-condition oracle.

#### Why not store the digest in the WAL?

Moving the digest aggregate into the WAL would not eliminate the
write amplification: the per-leaf XOR fold is already negligible
(it lives inline in the leaf's persisted state - there is no extra
WAL append for it today). The real amplification is the upward
chain of internal-node updates: every publish from a leaf - one per
coalesced group of foreground writes by default, or one per mutation when
`DigestCoalescingWindowMs` is `0` - rewrites the persisted subtree
aggregate on each ancestor up to the shard root.
That cost lives in internal-node grain state, not in the WAL, and is
the *whole point* of the incremental aggregate - readers need to
find the pre-folded shard hash in `O(1)`. Reconstructing it by
replaying the WAL on every digest poll would defeat the optimisation
and produce a per-call cost proportional to WAL size, which is
strictly worse than the per-leaf walk it replaced. The opt-out is
the correct knob for deployments that do not need the aggregate at
all.

## Recovery: fall-off-log triggers and `ProjectionRebuildPolicy`

When a leaf grain reactivates it consults its persisted
per-partition `ProjectionCheckpointOffsetsByPartition[p]` (and the
legacy scalar `ProjectionCheckpointOffset` for the partition-0
back-compat slot) and decides how to recover. The classifier runs
**once per partition** in `[0, WalPartitions)`; the leaf is refused
as fall-off-log if **any** partition's classifier reports genuine loss,
while a cost signal or the snapshot advisory below still tail-replays.
Three triggers classify an individual
partition - but only the **first** indicates missing data, and only
the first is fatal:

1. **WAL trimmed past checkpoint.** Partition `p`'s WAL
   has GC'd entries the leaf still considers unapplied. A tail
   replay would skip those entries and converge to the wrong state.
   Skipped unless the partition's checkpoint is positive: the -1
   "nothing applied" sentinel means the leaf has no in-memory state
   to lose to a trimmed prefix on that partition, and a checkpoint of
   `0` is skipped as well.
   **This is the only trigger that routes to `ProjectionRebuildPolicy`.**
   The exact loss boundary is `tail > checkpoint + 1`: the entry *at*
   the checkpoint is already applied, so trimming it loses nothing.
   A cold activation that finds no covering snapshot hands the classifier
   the -1 sentinel so that it replays the whole readable window, which
   blinds this trigger; that path is guarded separately against the
   leaf's durable checkpoint with the same boundary, and a trim past it
   surfaces `LeafProjectionStaleException` without consulting the policy.
2. **Replay budget.** The gap `walHead[p] - checkpoint[p]`
   exceeds `LatticeOptions.MaxLeafReplayEntries` (default `10 000`).
   This is a **cost** signal, not a loss signal, and the gap is not the
   leaf's work: it is measured across the whole WAL partition, which
   every leaf pinned to that partition shares, while
   `MaxLeafReplayEntries` is a **per-leaf, post-range-filter** budget.
   The two are in different units, and on a partition carrying ~1,350
   leaves the gap overstates a leaf's real work by up to that fan-out.
   What the comparison does establish is a sound **upper bound**: the
   head is the next offset the partition will assign, so every entry a
   leaf applies lies inside `(checkpoint, head)` and `applied <= gap`
   always holds. Since issue #2275 nothing acts on the
   comparison (it was once a candidate the replay confirmed): an
   over-budget gap yields the decision `TailReplayOverBudget` and the
   leaf tail-replays exactly as for `TailReplay`, and its only remaining
   effect is that it suppresses the snapshot advisory described under
   [Snapshot-on-fall-off safety net](#snapshot-on-fall-off-safety-net).
   Confirming the gap in the classifier would mean reading
   `(checkpoint, head)` before the replay reads it again - doubling the
   most expensive part of activation - so the **verdict** is taken on
   every replay path, during the replay that happens anyway, by counting
   the entries that actually pass the per-leaf range filter
   (`ShouldApplyDuringReplay`). The warning and the
   `orleans.lattice.leaf.activation_replays_over_budget` counter are
   emitted at that exact count, whatever the comparison said. The
   comparison is skipped for the -1 sentinel, whose gap would charge a
   fresh leaf for every sibling's WAL; the count is not, because it
   charges a fresh leaf only for its own range. The per-slice
   `WalReplaySliceBudget` still bounds individual coordinator reads on
   this path.
3. **Cold past retention.** The persisted projection age exceeds
   `LatticeOptions.LeafProjectionRetention` (default 7 days). Also a
   cost signal only - an old checkpoint does not imply a trimmed WAL,
   so it degrades to the same non-fatal `TailReplayOverBudget` replay.
   Age is a property of the leaf, not of any one partition, so every
   partition's classification sees the same age. Note the activation path currently
   supplies `TimeSpan.Zero` as the age, so this trigger does not fire
   from activation today (tracked in #1738).

> **Why triggers 2 and 3 are not fatal (issue #1738).** They were,
> until a tree holding fully intact data was permanently bricked by a
> replay gap of 10,648 against the 10,000 default - 648 entries, 6.5%
> over budget - while every offset it needed was still readable in the
> WAL. A cost guardrail must never be more destructive than the cost it
> guards against: a slow activation is recoverable, a tree that refuses
> to activate is not. Replay cost is bounded on the read side instead,
> by `WalReplayMaxRecordsPerTurn` (which yields between turns) and
> `WalMaterialiserMaxConcurrentReplays`.

> **Reading the over-budget warning (issue #2023).** The warning names the
> tree, the **leaf grain id**, and the **WAL partition ordinal**, and it
> states its own fault criterion: a persisted checkpoint that does not
> advance across repeats is a fault, not a slow replay. That comparison is
> only valid between lines naming the **same leaf and the same partition**.
> `partition` is iterated `[0, WalPartitions)` inside *every* leaf's
> activation, so it does not identify a leaf; before the leaf id was added,
> consecutive lines were one-per-minute samples of arbitrary different
> leaves, and their checkpoints appearing to repeat or move backwards was an
> artifact of that sampling rather than a stalled replay. The log is
> throttled per (tree, leaf, partition), starting at one line a minute and
> doubling to a ceiling of one an hour the longer that leaf keeps reporting,
> so a leaf that is genuinely stuck still reports an unchanging checkpoint but
> at a decaying rate: compare consecutive lines naming that leaf, rather than
> expecting a fixed cadence. The backoff advances only when a line is actually
> emitted: a repeat the per-tree cap withholds keeps its place in the queue
> instead of backing off having said nothing, so a leaf on a busy tree still
> rotates into the budget and still yields the two comparable lines the
> criterion needs. A clean in-budget activation retires the backoff outright,
> so a leaf that misbehaves, recovers, and regresses hours later reports at the
> base interval rather than inheriting the accumulated ceiling. A per-tree cap additionally bounds how many
> *repeat* lines one tree may emit in a window, because a tree with L leaves
> and P partitions has L x P throttle keys and so L x P times the per-key
> rate - which is how this warning reached 46% of one deployment's container
> log, rolling away the older entries that were the evidence needed to
> diagnose it (issue #2100). Any repeats the cap withholds are reported as a
> summary line, so the cap is not silent while that tree keeps replaying, and
> a leaf partition reporting over budget for the FIRST time is exempt from it,
> so a newly appearing condition still surfaces promptly. "First time" is a
> property of the key's own history and not of what the gate happens to have
> retained: internal housekeeping never restores the exemption, so the
> exemption cannot be re-earned by churn on a large estate. It is restored only
> after that leaf partition has been silent for a full ceiling interval, at
> which point its return really is new information. That summary is
> carried by a later occurrence on the tree, so a tree's final withheld tally
> goes unreported once the condition resolves or the leaf deactivates; the
> counter below is the exact census for that case. The
> `orleans.lattice.leaf.activation_replays_over_budget` counter is tagged
> `tree`, `partition`, and the derived `tenant`, never by leaf - leaf count
> is unbounded, so it cannot be a metric dimension - which means the counter
> measures the rate and the log makes the per-leaf call.

> **What the warning reports, and the stall fault (issue #2149).** The cost
> warning now carries the quantities it actually compared: the leaf's
> **applied entry count** (post-range-filter), the **budget**, and - so the
> two can be reconciled against the classifier - the partition **head** and
> **gap**. Before this, the line carried neither `head` nor `gap`, so the
> over-budget *factor* could not be read off the instrument at all; a
> dimensionless figure obtained by dividing an absolute WAL offset by the
> budget survived across three issues before a measurement run caught it.
> `applied <= gap` holds by construction, so a line violating it is
> reporting two different windows.
>
> Evaluating "did the checkpoint advance?" by hand across a log is what
> found the livelocked leaf of issue #2165, and it is exactly what stopped
> working once benign lines outnumbered it 560:1. That evaluation is now
> performed in-process: a leaf that re-enters replay for the same partition
> from an **unchanged** persisted checkpoint is reported as a distinct
> **stalled-replay fault**, on its own throttle so cost noise can never
> suppress it. The first activation stays silent - one cold activation is
> not a stall - and every repeat at a frozen checkpoint warns. The two
> conditions are independent: a stalled leaf whose own work is small warns
> as a fault and not as a slow replay, which is the #2165 shape exactly.

On the healthy multi-partition path every partition's classifier
returns `TailReplay`, and the leaf executes a two-pass replay across
all partitions (per-partition Set / Delete absorption with
`TxCommit` / `TxAbort` / `DeleteRange` deferred until every partition
has populated its pending-tx record, then drained; a `DeleteRange` whose
range cannot overlap the leaf's key range is instead consumed in the
first pass, applying nothing - issue #3601) followed by a
post-pass per-partition checkpoint reconciliation that advances each
partition's `ProjectionCheckpointOffsetsByPartition[p]` to the
highest applied offset once the saga-prepare clamp lifts.

The `ProjectionRebuildPolicy` enum on `LatticeOptions` is consulted only
when trigger 1 fires, and every value currently fails closed:

| Policy | Behaviour |
|---|---|
| `SnapshotThenWal` *(default)* | Intended to recover from the leaf's snapshot and then the WAL. The snapshot half is not specific to this policy: the per-leaf snapshot rehydrate runs at the start of every activation under every policy, before the classifier, covering the prefix and letting the tail replay handle the remainder. What is not yet integrated is a recovery for the case where that rehydrate has *already declined* and the WAL is genuinely short: there the leaf surfaces `LeafProjectionStaleException` rather than reconstructing the lost prefix. |
| `FullRebuildFromWal` | Diagnostic. Intended to replay from the absolute tail of the WAL, but the policy is reached only when the WAL has been trimmed and a complete history is unavailable, so the leaf surfaces `LeafProjectionStaleException` here too; no full-rebuild recovery path is integrated. |
| `Fail` | Surfaces a `LeafProjectionStaleException` at activation time and waits for an operator-driven rebuild. |

> This policy is reached **only** on genuine loss (the WAL trimmed past
> `checkpoint + 1`). A replay-budget or projection-age overrun against an
> intact WAL never consults it - see the trigger list above.

### Starvation-drive admission

Background starvation drives share the same process-wide replay permits as
leaf activations, but never queue for one. Across all trees, drives may hold
at most half the replay permits in circulation - the configured replay
ceiling less any that memory-pressure withholding is holding back - rounded
down with a minimum of one (issue #3610). This leaves capacity for foreground
and maintenance activations whenever more than one permit circulates. While
only one does - a ceiling of one, or a larger ceiling at the withholding
floor - drives and activations share that single permit.

Two callers request drives, and they do not compete for that share on
equal terms (issue #3575):

- the WAL GC's blocked-leaf sweep, the only caller that lifts a pin holding
  a tree's cursor floor, may use the whole share, and on a pass over the
  floor holders it has classified it touches the one nearest the floor
  first, alone, so a free permit goes to the floor rather than to whichever
  touch reaches the gate first (issue #3610);
- a leaf's own coverage-lag timer, which drives a leaf that has never
  checkpointed or whose checkpoint has stopped advancing, never takes the
  last free slot: it is admitted only while at least two slots of the share
  are free, so however many drives hold the others, one slot is always free
  for the sweep. Where the share is a single slot there is nothing to
  reserve, so the timer instead leaves that slot to a sweep drive that was
  refused it, until the sweep is admitted again or five minutes have passed.

The timer reaches every stalled leaf on a fixed cadence, so without this it
won the share by volume: on one deployment the sweep was refused about nine
touches in ten, and the WAL of the trees it was trying to clear grew without
being reclaimed.

When no permit is immediately available, or the drive's part of the GC
share is occupied, the drive is refused before replay starts. The refusal
is a result rather than an exception (issue #3761): the drive returns an
admission-refused verdict, and the refusal is counted on
`orleans.lattice.saturation.refusals` with `source=replay_permit_admission`
and `arm=gc_share`.
It neither advances nor retires the leaf's retention pin, and neither
caller treats it as a fault:

- the sweep records the try as `outcome=admission_refused` on
  `orleans.lattice.wal.gc.blocked_leaf_reactivations` (beside the drive's
  own `drove_admission_refused` verdict when the drive returned one), logs
  it at `Debug` without a stack, and does not count it as `attempted` or
  charge it
  against the consumer's attempt budget, so a consumer the sweep never
  managed to drive is never abandoned. A pass keeps no more touches in
  flight than this silo's part of the GC share, so its own touches do not
  refuse each other, and re-drives a refused touch once a sibling touch of
  the same pass frees a slot (issue #3761). A consumer still refused when
  the pass ends retries after a jittered delay of one to one and a half
  minutes that doubles with each consecutive refusal, up to its ordinary
  fifteen-minute cooldown;
- the timer counts it as `reason=recheck_drive_refused` on
  `orleans.lattice.leaf.snapshot.driver.declines` and backs off: it skips
  its next drive opportunities, counting each as
  `reason=recheck_drive_deferred` - one
  after a first refusal, rising to four to seven after repeated ones,
  jittered per leaf. A drive that is admitted ends the backoff.

The activation queue's depth and drain policy are unchanged. Per-tree touch
limits alone cannot bound the aggregate load of many trees on this
process-wide gate (issue #3480).

### A live leaf whose projection has gone stale

A leaf can find its projection stale while it is still activated. Its
in-memory cache was built before the WAL was trimmed, so it keeps
serving reads and taking writes, but its *persisted* checkpoint needs
an offset the trim removed. Before latching this fault, the starvation
drive attempts a conservative **warm-cache rescue**. This is not a cold
rebuild: it writes a snapshot of the surviving cache only when that
activation can prove the entire claimed prefix is present.

Eligibility starts with a successfully hydrated snapshot covering every
configured partition. A successful replay retains its independently
scanned, contiguous per-partition frontier; foreground checkpoint hints
are never evidence for that frontier. Later hydration, reset, topology
changes, failed or overlapping replay, and - for an unfiltered replay -
gaps in the observed WAL sequence invalidate the proof. A filtered
replay's slices omit other owners' records by design (issue #3565), so
there only non-ascending offsets invalidate it, and the window is proven
instead by checking, once the partition is read, that the oldest
surviving WAL offset is still at or below the first offset the window
needed. Each partition's activation anchor must
be at or behind its persisted checkpoint, and its current checkpoint
must not exceed the proven frontier. The live WAL must still cover
everything after the proposed snapshot claim.

The drive also excludes unresolved transactions, active mutations,
another capture, split/merge activity, retirement, and moved-away seals.
It holds the topology gate across snapshot storage and checkpoint
persistence; sealing and unsealing use that same gate. The snapshot
claim is fixed before capture, and eligibility is rechecked immediately
before copying the rows. Only a store-confirmed kept snapshot permits
the checkpoint persist; the drive then awaits durable-pin publication
before reporting success. A failure after the snapshot was kept propagates
rather than claiming that nothing was saved. Storage is addressed using
the leaf's bound tree identity, never an identity taken from a replayed
mutation.

Every rescue decline emits one structured warning per reason per
activation, with `LeafId`, `TreeId`, and typed `DeclineReason` fields:

| Reason | Evidence that is missing or blocks rescue |
|---|---|
| `UnprovenBaseline` | No fully covering snapshot-seeded successful replay frontier. |
| `CacheRehydratedOrReset` | The original cache was replaced or its replay barrier retired. |
| `TopologyChanged` | This activation attempted a topology change after its baseline. |
| `ReplayIncomplete` | Replay failed, overlapped, remains active, or has not satisfied its barrier. |
| `UnknownPartition` | The stale partition or complete partition width is not established. |
| `CheckpointUnproven` | A current checkpoint exceeds its independently proven frontier. |
| `ActivationAnchorAhead` | A persisted checkpoint is behind the activation's snapshot anchor. |
| `WalGapBeyondCache` | The oldest surviving WAL offset is beyond the claimed prefix plus one. |
| `PendingTransactions` | Prepared, shadowed, or unresolved replay work remains. |
| `MutationInFlight` | An admitted mutation has not finished. |
| `CaptureInFlight` | Another snapshot capture owns the capture slot. |
| `SplitInFlight` | A split is interrupted or the topology gate is occupied. |
| `RetiredOrSealed` | The leaf is retired or holds moved-away slots. |
| `CaptureDeclined` | Snapshot storage did not acknowledge a kept capture. |
| `StorageFailure` | A probe or capture failed before kept coverage was acknowledged. |

Cold, already-lost, partially snapshot-covered, or otherwise unproven
caches still decline. This tree-agnostic path also covers derived trees,
but does not trigger their re-derivers and cannot reconstruct lost data.
The proof is activation-local and is never serialized or reconstructed
from a persisted checkpoint. A cache already warm before this code was
deployed has no recorded proof and declines with `UnprovenBaseline`.
Restarting to deploy the change discards that cache; this path cannot
recover it.

Two drivers still reach such a leaf: the coverage-lag timer, which
drives a leaf whose checkpoint has stopped advancing, and the WAL GC's
blocked-leaf sweep. If the rescue declines, the first starvation drive
logs one `Error` naming the leaf and tree, and latches
that verdict for the activation (issue #3450). While the
persisted checkpoints are unchanged:

- the timer skips the drive entirely, so a stale leaf no longer takes
  a replay permit or raises a timer fault once per stall window;
- a drive from the WAL GC's blocked-leaf sweep still receives
  `LeafProjectionStaleException`, without another replay. The sweep
  treats that verdict as terminal (issue #3478): it records it once as
  `outcome=latched_stale` on
  `orleans.lattice.wal.gc.blocked_leaf_reactivations`, logs it at
  `Information`, and stops driving the leaf for the rest of the tree's
  blocked episode. The leaf's pins stay in the WAL GC cursor floor, so
  no WAL it still needs is trimmed, and the tree is reported once per
  sweep pass with a `Warning` naming how many latched stale leaves hold
  it.

Any change a genuine apply or an operator reset makes to a persisted
checkpoint clears the latch, and a new activation starts without it. A
split's checkpoint hint does not: a hint for the partition the drive
found stale is refused and logged as a `Warning`, because stamping it
would persist a checkpoint past the trimmed range the leaf never
applied, and would clear the latch with nothing repaired (issue #3477).
The leaf keeps its WAL retention pin.

Treat the `Error` as data at risk rather than as noise. The live
activation may hold the only copy of writes in the trimmed range. Once
it is recycled or the silo restarts, the next cold activation refuses
the leaf with the same exception. `RebuildLeafProjectionAsync` does not
recover those writes: it resets the checkpoint, and the next activation
reloads the leaf's snapshot where it has one and replays only the WAL that
survives. So, before the activation is lost:

1. Capture a logical backup or export of the tree while the leaf is
   still serving.
2. Deploy a build that fixes the cause of the over-trim.
3. Restore the tree from that backup. For a tree that is derived from
   another source, delete it and re-derive it instead.

### Configuration

```csharp verify
siloBuilder.ConfigureLattice(o =>
{
    // Expected replay size before a cold activation is flagged as
    // over-budget (a warning plus a counter - never a failure):
    o.MaxLeafReplayEntries = 100_000;

    // Age at which a cold projection would be flagged as stale. The
    // activation path does not yet supply an age, so this does not
    // fire today (see trigger 3 above):
    o.LeafProjectionRetention = TimeSpan.FromDays(30);

    // How to recover when the WAL is genuinely short:
    o.ProjectionRebuildPolicy = ProjectionRebuildPolicy.SnapshotThenWal;
});
```

## Snapshot-on-fall-off safety net

The three activation-time triggers above react to a fall-off-log
condition *after* it has already happened. The snapshot-on-fall-off
path is the preventative safety net: while a leaf is still healthy,
it captures a canonical-row image of its in-memory cache to a
dedicated snapshot grain whenever any partition's WAL tail
approaches that partition's persisted checkpoint. On the next
activation the leaf rehydrates its cache from the blob rows and
tail-replays each partition forward from its captured offset. It does
so when the snapshot is newer than the persisted partition-0
checkpoint, and also when the snapshot is at or behind it, provided the
cache starts empty or any partition's WAL prefix has been trimmed: a
snapshot may then be the only durable copy of a prefix the WAL GC
trimmed, and each partition's checkpoint is lowered to what the
reloaded cache holds before the tail replay. Activation then
proceeds without ever needing to fall back into
`SnapshotThenWal` / `FullRebuildFromWal` / `Fail` recovery, even
when the WAL has been trimmed past the original checkpoint.

The capture path is **leaf-driven**, not maintenance-driven:

- At activation, the leaf runs the fall-off-log detector once per
  partition. When no trigger has fired - not even one of the two cost
  signals - but a partition's persisted checkpoint sits within the
  oldest `LeafSnapshotMargin` fraction (default `0.30`) of that
  partition's readable WAL window, the detector returns a non-fatal
  snapshot advisory. The leaf latches the advisory, finishes its tail
  replay across every partition, and then captures a single snapshot
  before yielding the activation turn.
- While the leaf remains hot, every
  `LeafSnapshotReClassifyEveryNCheckpoints` (default `64`)
  successful checkpoint persist re-runs the classifier and drives
  another capture on advisory. Pass `0` to disable the periodic
  recheck entirely.
- Independently of the advisory and of that setting, a leaf holding a
  checkpointed partition that no snapshot covers yet captures one at
  activation and on its later checkpoint persists and coverage-lag
  checks, until a capture succeeds, within a per-activation attempt
  budget that re-arms after a backoff (issue #2692). A capture covers
  every partition checkpointed at the time, so a partition that
  checkpoints later brings the leaf back once more, and an attempt that
  covered at least one partition is not charged against the budget:
  only captures that make no progress spend it (issue #3576). This is what gives
  a tree that has stopped taking writes snapshot coverage on its next
  activation.
- A graceful deactivation also captures one for any checkpointed
  partition that no snapshot covers yet, before it publishes its final
  durable pin, so a short-lived activation that never reached the
  periodic cadence still leaves coverage behind.
- A checkpoint persist whose snapshot recheck lands a capture publishes
  the durable pin again after that recheck, so the capture reaches the pin
  at once instead of being left to the later `frontier_pin` barrier, which
  a deactivation deadline can skip. The final persist of a graceful
  deactivation also publishes its pin before the recheck, so a recheck
  that overruns the deadline cannot cost the pin. The coverage-lag check also republishes a pin that has fallen
  below `min(persisted checkpoint, coverage)`, first committing a pending
  checkpoint advance once `MaterialiserCheckpointInterval` has elapsed so a
  write-idle leaf's advance cannot stay pending (issue #3608), and the WAL GC's
  blocked-leaf sweep asks a floor-holding leaf for the same step before it
  spends a replay permit on a drive. That step takes no replay permit and
  replays nothing, and a capture that fails or is declined leaves coverage,
  and so the pin, where they were (issue #3599).
- A single-flight guard suppresses overlapping captures: a slow
  `SaveAsync` does not pin a follow-on capture behind it; the
  follow-on is dropped and the next cadence tick re-evaluates.

Each capture overwrites the previous blob; only the most recent
snapshot is retained per leaf. The WAL remains the long-term audit
trail.

### Cold-leaf limitation

A leaf that never activates while drifting below the WAL retention
window will not be captured by this path. Such a leaf also holds no
in-memory state to lose; if the WAL has been trimmed past its
checkpoint with no covering snapshot, its next activation refuses it
with `LeafProjectionStaleException`, whatever the configured
`ProjectionRebuildPolicy`. The snapshot-on-
fall-off path is a safety net for **active** leaves.

### Resumable cold replay

The activation-time WAL replay that rebuilds a leaf's projection cache
is **resumable**. A leaf rebuilds its cache from the WAL on every
activation, but the persisted `ProjectionCheckpointOffset` is what
survives across activations. Historically the checkpoint was advanced
only by a single reconciliation step at the very *end* of the replay,
so a leaf whose un-snapshotted prefix could not be drained inside one
activation window - a large WAL relative to Orleans' ~30s
`RuntimeRequested` deactivation budget - made no durable progress: a
mid-replay deactivation discarded every applied entry, the next
activation restarted from the same offset, and because the
coverage-gated WAL GC (correctly) refuses to trim an un-snapshotted
prefix, the WAL grew without bound while the leaf never converged.

The replay now flushes the checkpoint **incrementally**, at each
replay slice boundary, over the strictly contiguous, fully-applied
prefix. Because a checkpoint persist drives the periodic snapshot
recheck above, this also captures an incremental snapshot as the
replay progresses. A mid-replay teardown therefore loses at most one
flush interval; the next activation rehydrates from the incremental
snapshot and resumes from the last durable offset instead of replaying
from zero, and the now snapshot-covered prefix becomes trimmable so
retention stays bounded.

The incremental advance can never license a checkpoint (or the
materialiser pin) past work that is not yet durable. The replay defers
two kinds of record:

- a **deferred** saga terminal (`TxCommit` / `TxAbort`) or
  `DeleteRange`, which is applied only in the replay's second pass (a
  range delete that cannot overlap the leaf's key range is not deferred:
  the first pass consumes it, so it takes no ledger slot); and
- an **unresolved saga prepare**, whose pending-transaction bucket no
  snapshot captures because the matching terminal is itself deferred.

Since issue #2165 the leaf writes each such record verbatim into a
ledger on its persisted state row before the flush that advances past
it, so the record and the checkpoint it licenses land in the same state
write, and a resumed activation reconstructs the work from the ledger
instead of re-reading it from the WAL. For deferred terminals the ledger
is bounded by `MaxDurableUnresolvedReplayWork` (default `1 024`); past
that bound the flush falls back to holding the checkpoint below the
deferred entry until the second pass applies it. An unresolved prepare
is always recorded (issue #2183), because nothing drains a prepare whose
saga never terminates.

Steady-state activations are unaffected: a single-slice replay with no
deferred terminals still flushes exactly once at the end of the first
pass, the same net timing as the old end-of-replay reconciliation.
Only a multi-slice replay of a large prefix sees the extra intermediate
persists - which is exactly the case the resumability guarantees.

### Configuration

```csharp verify
siloBuilder.ConfigureLattice(o =>
{
    // Trigger a proactive snapshot capture when the leaf's
    // persisted checkpoint is within 30% of the WAL tail. Lower
    // values reduce snapshot frequency; raise to capture earlier.
    o.LeafSnapshotMargin = 0.30;

    // While a leaf stays hot, re-run the fall-off classifier every
    // N successful checkpoint persists and re-capture on advisory.
    // Set to 0 to disable the periodic recheck entirely.
    o.LeafSnapshotReClassifyEveryNCheckpoints = 64;
});
```

## Related surfaces

- `ILattice.GetLeafProjectionDigestAsync` - the public surface.
- `ILattice.GetLeafProjectionDigestForRangeAsync` - the range-scoped
  analogue (`null` bounds give the whole-shard digest) that backs the
  cross-cluster [Merkle-walk localisation](../lattice.replication/anti-entropy-merkle-walk.md).
- `LeafProjectionDigest` - the returned `readonly record struct`.
- `LatticeOptions.MaintainProjectionDigest` - opt out of the
  per-mutation XOR fold and upward publication for digest-indifferent
  workloads.
- `ProjectionRebuildPolicy` - the activation-time recovery policy.
- `LatticeOptions.MaxLeafReplayEntries`, `LatticeOptions.LeafProjectionRetention`,
  `LatticeOptions.MaterialiserCheckpointInterval`,
  `LatticeOptions.MaterialiserCheckpointEntries`,
  `LatticeOptions.LeafSnapshotMargin`,
  `LatticeOptions.LeafSnapshotReClassifyEveryNCheckpoints` - see [Configuration](configuration.md).
- `LeafProjectionStaleException` - surfaced on genuine loss under every
  `ProjectionRebuildPolicy`, including the default `SnapshotThenWal`,
  whose post-rehydrate recovery is not yet integrated; also rethrown by
  a starvation drive against a leaf latched stale.
- `ILattice.RebuildLeafProjectionAsync` and `ILattice.GetMaterialiserLagAsync` -
  see [Operator tooling: rebuild and lag](#operator-tooling-rebuild-and-lag).

## Operator tooling: rebuild and lag

Activation-time `ProjectionRebuildPolicy` recovers a leaf when it
cold-starts and a fall-off-log trigger fires. Two complementary
surfaces let an **operator** drive recovery and observe materialiser
health without waiting for an activation:

- `ILattice.RebuildLeafProjectionAsync(int shardIndex, CancellationToken)`
- `ILattice.GetMaterialiserLagAsync(CancellationToken)`

### Rebuild a shard's projection from the WAL

```csharp verify
// GetLeafProjectionDigestAsync reported a mismatch, or an
// integrity check flagged a leaf's projection as suspect. Force every
// leaf in the shard to re-materialise its projection on its next
// activation: from its snapshot where it has one, then from the WAL.
await tree.RebuildLeafProjectionAsync(shardIndex: 0, cancellationToken);
```

`RebuildLeafProjectionAsync` walks every leaf in the named physical
shard via the sibling chain and, for each leaf, **clears only the
projection-state slots** that the materialiser owns:

- The per-activation entry cache is dropped when the grain
  deactivates (the cache is never persisted, so there is nothing to
  clear on the state row itself).
- The persisted projection checkpoint is reset to the `-1` "nothing
  applied" sentinel (matching the WAL reader's
  `fromOffsetExclusive = -1` start-of-log convention) and the
  per-partition checkpoints are dropped, so a partition that no
  snapshot covers is replayed from offset `0` inclusive on the next
  activation. Setting `0` instead would cause the materialiser to skip
  offset `0`, because replay reads strictly past the persisted
  checkpoint.
- The persisted running projection hash is cleared.
- In-memory pending-saga, pending-tx-offset, recently-terminal, and
  backstopped-terminal dedup buffers are dropped, together with the
  destination-side shadow markers.
- The leaf grain is deactivated. The next activation re-materialises
  the projection through the standard activation-time path, which
  under every `ProjectionRebuildPolicy` value first attempts to
  rehydrate from the leaf's captured snapshot. The rebuild neither
  clears nor bypasses that snapshot: when the leaf has a usable one,
  its cache is reloaded from it, and each partition the snapshot
  covers replays only the WAL entries after the snapshot's captured
  offset.

**Topology-bearing state is preserved**: `TreeId`, `ShardIndex`, the
leaf's key range, and the sibling pointers stay intact. The rebuild
does not re-shape the tree - it only re-derives the materialised
projection for the range the leaf already owns. Only the part that no
snapshot covers is re-derived from the WAL: a snapshot-covered prefix
is restored from the snapshot rather than re-applied, so a row the
snapshot captured wrongly - after a corruption or a projection bug -
survives the rebuild unless a later WAL entry for that key replaces it.

`RebuildLeafProjectionAsync` does **not** take a tree-wide
consistency lock. Readers and writers continue to land on the shard
during the rebuild; in-flight writes hit the leaf's standard write
path (which goes through the WAL) and will be visible after the
next activation re-replays them. A read that reaches a rebuilt leaf
waits for that leaf's activation replay to finish rather than seeing
a partial projection, so operators rebuilding under load should
expect a window of higher read latency rather than missing entries;
a replay that fails fails the waiting read, and the next request
retries it. Pair the rebuild
with a digest re-poll after replay stabilises to confirm the
projection converged.

Error surface:

| Condition | Exception |
|---|---|
| `shardIndex` is not a physical shard of the per-tree map | `ArgumentOutOfRangeException` |
| Tree id starts with the reserved system prefix `_lattice_` | `LatticeReservedTreeNamespaceException` (an `InvalidOperationException` subclass) |
| `cancellationToken` was already cancelled | `OperationCanceledException` |
| An access gate is configured and does not authorise the caller for whole-tree admin | `LatticeAuthorizationDeniedException` |

### Observe materialiser lag

```csharp verify
long lag = await tree.GetMaterialiserLagAsync(cancellationToken);
// An estimate of the WAL entries the worst shard's leaves have not yet
// folded into their projections. Each WAL head is the next offset to be
// assigned and each checkpoint the last offset applied, so even a
// caught-up shard with a non-empty WAL reports a small positive value:
// read a steady value as caught up and a growing one as falling behind.
```

`GetMaterialiserLagAsync` returns the **maximum lag across all
physical shards** of the tree. For each shard the lag is the sum, over
the tree's WAL partitions, of

```text
walHead[p] - min(checkpointOffset across leaves in the shard)
```

with each term clamped at zero so a checkpoint that has temporarily
raced ahead of the head observation (e.g. between the head fetch and
the per-leaf checkpoint fetch) cannot contribute a negative value.
`walHead[p]` is the next offset partition `p` will assign - one past its
newest entry - while a checkpoint is the offset of the last entry a leaf
applied, so a partition whose newest entry the checkpoint has reached
still contributes `1`; the result is `0` only when every term clamps to
zero, as on an empty WAL. The
per-leaf checkpoint read is each leaf's partition-0 checkpoint, which
the reduction applies to every partition's head as an approximation, so
on a multi-partition tree the figure is an estimate rather than an exact
count of unapplied entries; a shard with no leaves reports the sum of
its partition heads. The result is the worst-shard lag because a single
slow shard is the SLO-relevant
signal - averaging it would mask the actual problem.

The intended monitoring shape is a periodic poll (e.g. every
5-30 s) feeding a gauge in the telemetry pipeline:

```csharp verify
// Pseudo-code for an operator polling loop. The interval and
// thresholds are deployment-specific - they scale with WAL
// ingestion rate and the SLO for read-after-write recency.
long lag = await tree.GetMaterialiserLagAsync(cancellationToken);
if (lag > 10_000)
{
    // Materialiser is falling more than 10 000 entries behind the
    // WAL on at least one shard. Investigate slow leaf activation,
    // a stuck replay coordinator, or backpressure on the storage
    // provider.
}
```

A growing lag indicates the materialiser is not keeping up with WAL
ingestion. Common causes are slow leaf activation under
storage-provider backpressure, a stuck `ILeafReplayCoordinatorGrain`,
or a deactivation storm cycling leaves faster than they can replay.
A persistent lag at a small positive value (a few entries) is
expected: each non-empty partition contributes at least `1` even when
caught up, and under sustained write load the materialiser checkpoints
in batches, so the most recently published WAL entries naturally lag
the head briefly.

Error surface:

| Condition | Exception |
|---|---|
| Tree id starts with the reserved system prefix `_lattice_` | `LatticeReservedTreeNamespaceException` (an `InvalidOperationException` subclass) |
| `cancellationToken` was already cancelled | `OperationCanceledException` |
| An access gate is configured and does not authorise the caller to read the whole tree | `LatticeAuthorizationDeniedException` |

### When to use rebuild vs. activation-time recovery

| Scenario | Surface |
|---|---|
| Leaf cold-starts and the WAL has been trimmed past its checkpoint with no covering snapshot (genuine loss) | Activation refuses the leaf with `LeafProjectionStaleException` under every `ProjectionRebuildPolicy` (the automatic recovery paths are not yet integrated): restore the tree from a backup, or accept the loss of the trimmed range and run `RebuildLeafProjectionAsync` |
| Leaf cold-starts with a replay gap over `MaxLeafReplayEntries` with the WAL intact | Non-fatal: the leaf tail-replays, and warns only when the entries it actually applies exceed the budget; no operator action needed. The `LeafProjectionRetention` age trigger does not fire from activation today |
| Operator detects a digest mismatch across silos, or an integrity check flagged a corrupted projection, or a fix to how WAL entries are applied to the projection requires re-materialisation | `RebuildLeafProjectionAsync` (manual, while live). It re-applies only the WAL no snapshot covers: a prefix the leaf's snapshot covers is restored from the snapshot, so a row the snapshot captured wrongly survives unless a later WAL entry for that key replaces it |
| Operator wants a steady-state gauge to know whether the materialiser is keeping up | `GetMaterialiserLagAsync` |

The two paths share the same replay seam: `RebuildLeafProjectionAsync`
clears state and lets the standard activation-time path do the
re-materialisation. There is no second, parallel rebuild code path
to maintain - the operator surface is a controlled trigger for the
existing recovery logic.
