---
title: "Options Reference: ShardHealingEnabled to WalSaturationMaterialiserLagThreshold - Configuration"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/configuration/options-reference-3.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/configuration.md?plain=1#L1225-L1602"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Options Reference: `ShardHealingEnabled` to `WalSaturationMaterialiserLagThreshold`

Part of [Options Reference](options-reference.md), in [Configuration](../configuration.md).

## `ShardHealingEnabled`

Whether each tree runs a steady-state observer that automatically folds an over-split tree back towards its registry-pinned base shard count (default: `true`).

A tree that a bulk ingest shattered stays shattered forever without this: shard count only ever grew, and every subsequent cold start paid for the excess. The observer heals it with no operator action, which is the whole point - the deployments that most need healing are the ones nobody reconfigures.

**Cost: a healthy tree polls no shard at all.** The sweep's first phase reads the shard map and the pinned base shard count, both of which the orchestrator already holds, and a tree at or below its base is settled there. Only a tree that really is over-split goes on to poll shard hotness and the tree's quiescence verbs. That is what makes it safe to run forever on every tree.

Healing is deliberately **not** gated on [`AutoSplitEnabled`](options-reference-1.md#autosplitenabled). Disabling the splitter on an already-shattered deployment is precisely the configuration that most needs healing, and coupling them would leave such a tree damaged forever.

**When off:** no fold is admitted and no shard is polled. A tree whose healing orchestrator has never started registers no reminder and starts no timer. One that is already running keeps sweeping on its timer until it deactivates, but each sweep now only reads the tree's shard map and records a `disabled` decision (still publishing the backlog), and its once-a-minute keepalive reminder stays registered - every later tick of it returns without doing anything. An in-flight fold is left to its own resumable coordinator and is never cancelled, so the tree stays consistent.

It is worth knowing how fast each direction is. Turning it **off** takes effect on the next sweep, because every sweep re-reads the option. Turning it back **on** resumes an orchestrator that has run before on its next sweep if its timer is still armed, otherwise on its next once-a-minute keepalive reminder, which stays registered while healing is off - the same shape as [`AutoSplitEnabled`](options-reference-1.md#autosplitenabled). Only a tree whose orchestrator has never started, because healing was off every time the tree bootstrapped it, waits longer: the orchestrator does not latch itself while disabled, so it starts on the next bootstrap call, and the tree issues that call when it activates - so in practice re-enabling is picked up the next time the tree activates, whether or not anything writes to it. No redeploy is needed either way, and off is never slower than on, which is the right direction for a kill switch.

**When an operator would turn it off:** to isolate healing while diagnosing shard-routing or consolidation faults, or on a tree that is deliberately over-provisioned above its registry pin and should stay that way. To pause healing while still measuring the backlog, prefer `MaxConcurrentShardConsolidations = 0` instead.

```csharp verify
// Stop automatic healing on one tree (takes effect on the next sweep):
siloBuilder.ConfigureLattice("hand-tuned-tree", o => o.ShardHealingEnabled = false);

// Or keep observing the backlog while admitting no folds:
siloBuilder.ConfigureLattice("hand-tuned-tree", o => o.MaxConcurrentShardConsolidations = 0);
```

## `ShardHealingInterval`

How often each tree's healing orchestrator observes the tree (default: 30 seconds). The value is the orchestrator's grain-timer period, so options validation rejects one above about 49.7 days; see [Timeout and budget ceiling](options-reference-1.md#timeout-and-budget-ceiling).

**Cost:** one sweep per tree per interval. On a healthy tree that is a map read and a comparison, which is why the cadence can be this short.

**When off:** there is no "off". Options validation **rejects** a non-positive interval rather than treating it as disabled, because an interval that could never schedule an observation is far more likely to be a mistake than an intent; the failure message points at [`ShardHealingEnabled`](#shardhealingenabled) instead.

**When an operator would change it:** lengthen it on a box with very many trees, where the aggregate sweep cost matters more than healing latency.

## `EmptyTreeProbeBudget`

Ceiling on the emptiness probe that reshard and resize initiation run before taking their empty-tree fast paths (default: 10 seconds).

`ReshardAsync` and `ResizeAsync` each need only a boolean - "does this tree hold any live key?" - to decide whether they can repin registry state directly instead of starting a migration coordinator. Both used to answer it with `CountAsync`, a strongly-consistent whole-tree fan-out that walks every leaf chain, then discards its result and retries whenever the shard map moves under it, giving up only once [`MaxScanRetries`](options-reference-2.md#maxscanretries) is exhausted. Initiation is exactly when that map is most likely to be churning, so an unbounded exact count could consume the caller's whole response budget and time the operation out before it had started.

Both now probe existence directly, OR-ing a short-circuiting per-shard check that stops at the first non-empty leaf. That needs no reconciliation against a moving shard map: a count must reconcile because a key migrating between shards is briefly visible on both the source and the destination, which double-counts, whereas a split only ever *moves* keys - never creating, destroying, or leaving one present on neither side - so a key that exists is seen by at least one shard wherever the split has got to, and seeing it twice still just means "a key exists".

This budget is the remaining backstop for a probe that parks rather than returns. The answer is deliberately one-sided: it may report non-empty while the last keys migrate away, but never empty while a key exists anywhere, and only "empty" unlocks a fast path - so the one consequential direction cannot be wrong. Every inconclusive outcome (the budget elapsing, or a shard faulting) is reported as "not empty", so initiation simply proceeds down the normal coordinator path.

Set to `InfiniteTimeSpan` to wait indefinitely; the options validator rejects any other non-positive value.

This option can be changed freely at any time. The new value takes effect on the next reshard or resize.

## `ActivationReadyTimeout`

Hard ceiling on how long a `ShardRootGrain`'s one-time activation-readiness seed may run before it is abandoned and surfaced to the preparing turn as a `TimeoutException` (default: 15 seconds). The seed is the chain of cross-grain awaits a brand-new or freshly-reactivated shard runs the first time it prepares for an operation: the defensive `state.ReadStateAsync` re-read, the tree-registry `RegisterAsync`, the deterministic root-leaf init pair, and the initial shard-state write.

This seed runs while the shard holds its non-reentrant activation gate. During a startup reshard or a membership change Orleans can reject or park one of those messages (the target registry or leaf activation is not yet visible) and leave the caller-side `await` neither completing nor faulting. Without a ceiling the parked seed pins the gate, every interleaved read/write turn on the activation stalls behind it, the lattice grain's per-shard fan-out saturates at its in-flight limit, and the whole write pipeline wedges with no fault and no activation recycle until the caller-side Orleans response deadline (30 seconds by default) expires. With the ceiling the parked seed is abandoned and the turn faults cleanly with a `TimeoutException`, which the existing transient-exception retry envelope on every mutation path catches and re-runs against refreshed routing or registration once the startup reshard has settled. Abandoning a parked seed never loses data or double-registers: every cross-grain step is idempotent on retry, and a failed shard-state write reverts the in-memory seed so the retry re-runs cleanly.

Set to `InfiniteTimeSpan` to disable the ceiling and restore the historical unbounded-await behaviour; the options validator rejects any other non-positive value.

When the deadline fires, the seed throws `ShardActivationTimeoutException` (publicly visible, derived from `TimeoutException`). The exception is retriable by construction - every cross-grain step in the seed is idempotent on retry - and every public `ILattice` operator that drives the seed transparently absorbs up to two consecutive occurrences before propagating to the caller, so external code generally does not need to special-case the cold-start race. Coverage spans the per-key read / write surface (via the central stale-routing envelope), multi-shard fan-outs (with per-shard wraps so a single shard's seed-timeout retries only that shard, not every sibling), per-tree coordinator entry points (resize / reshard / snapshot / merge / delete / recover / purge / bulk-load / compaction / projection-rebuild), the saga path (`SetManyAtomicAsync`), the digest path (`GetLeafProjectionDigestAsync`), and the per-shard warmup probe. Callers that want to detect or instrument the absorbed retries explicitly can catch the typed exception (it carries `TreeId`, `ShardIndex`, and `TimeoutSeconds` slots for per-occurrence attribution).

This option can be changed freely at any time. The new value takes effect on the next seed.

## `DigestPublishTimeout`

Hard ceiling on how long a single upward digest publish may run before it is abandoned and surfaced to the publishing turn as a `TimeoutException` (default: 15 seconds). It bounds the child-digest propagation an internal node issues to its parent after folding a child's digest, and the same publish from a leaf to its parent.

The publish is a cross-grain RPC that recurses up the internal-node chain toward the shard root. An internal node sends it only after releasing its non-reentrant split gate, and a parent whose gate is busy parks the incoming snapshot for the gate holder to fold rather than waiting for the gate, but an upward await can still be left neither completing nor faulting - a parent that is itself mid-mutation can hold it until the Orleans response timeout. With the ceiling the parked publish is abandoned and the turn faults with a `TimeoutException`. Abandoning a publish never drifts the digest count: the publish never partially applied at the parent, the digest is staleness-tolerant, and the next mutation's dirty-flag publish re-drives convergence. A non-zero `orleans.lattice.internal.digest_publish.timeouts` counter surfaces the condition for internal-node publishes.

Set to `InfiniteTimeSpan` to disable the ceiling and restore the historical unbounded-await behaviour; the options validator rejects any other non-positive value.

This option can be changed freely at any time. Leaves and internal nodes resolve their tree's options once per activation, so a new value takes effect on each node's next activation.

## `SoftDeleteDuration`

How long a soft-deleted tree's data is retained in storage before being permanently purged (default: 72 hours). During this window the shards stored under the tree's own id are marked deleted - reads and writes that reach them throw `InvalidOperationException` - but their grain state still exists in the storage provider. On a resized, restored or remediated tree whose data lives behind an alias, the delete marks and later purges the live copy the alias targets; see [Deleting an aliased tree](../tree-deletion.md#deleting-an-aliased-tree). After the duration elapses, a grain reminder triggers a full purge that walks every shard stored under the id, clears all leaf and internal node state, and deactivates each grain.

A resize uses the same window for the tree's old physical copy. It soft-deletes that copy silently - the tree does not read as deleted and no tree-lifecycle event is published - and `ILattice.UndoResizeAsync` can reverse the resize until the copy is purged. See [Retiring a resized tree's original copy](../tree-deletion.md#retiring-a-resized-trees-original-copy). The copy an undone resize discards is purged after the same window, but its write-ahead-log retention is released at once; see [Discarding an undone resize's copy](../tree-deletion.md#discarding-an-undone-resizes-copy).

Set to `TimeSpan.Zero` to purge on the first reminder tick, which fires one minute after the delete (the reminder period is clamped to a 1-minute minimum by the Orleans reminder floor).

```csharp verify
// Retain deleted trees for 7 days
siloBuilder.ConfigureLattice(o => o.SoftDeleteDuration = TimeSpan.FromDays(7));

// Immediate purge for a specific tree
siloBuilder.ConfigureLattice("ephemeral-tree", o =>
{
    o.SoftDeleteDuration = TimeSpan.Zero;
});
```

This option can be changed freely at any time. A deleted tree's purge reminder fires on the cadence of the value in force when it was deleted, but each tick compares the elapsed time against the current value, and the reported recovery deadline is computed from the current value too. A change therefore also reaches trees that are already soft-deleted: lengthening it defers their purge to the first reminder tick at or after the new deadline, while shortening it cannot purge one before its first reminder tick.

With `Orleans.Lattice.Apps`, an installed app's manifest can declare this value for each tree the app creates, and `AddLatticeApps` applies it to those trees' named options as a configure step. Under the registration-order rule above, any `ConfigureLattice` call that sets `SoftDeleteDuration` and is registered after `AddLatticeApps` - global or per-tree - still wins. See [Orleans.Lattice.Apps](../../lattice.apps/README.md).

## `SplitDrainBatchSize`

Number of entries per batch during the shadow-write drain phase of an adaptive split (default: 1024). Larger batches reduce the number of drain rounds but increase per-round memory and storage I/O. The options validator does not check it; a value of `0` or less falls back to the 1024 default.

This option can be changed freely at any time.

## `StarvationDriveBudget`

Hard ceiling on how long a single WAL GC starved-leaf checkpoint drive may run while holding a permit on the per-silo WAL replay concurrency gate, before it abandons its replay and releases that permit (default: 5 minutes = `4 * WalDrainBudget`).

**What it defends against (issue #3065).** The retention sweep reactivates a dormant leaf whose unusable durable pin is blocking its tree's WAL cursor floor, and drives that leaf's outstanding replay under the same replay permit gate the activation path uses. Before this budget existed, every await inside the permit-guarded region was passed `CancellationToken.None`: no timeout, no cancellation, anywhere. A drive whose commit-log read never returned therefore held one of a small number of per-silo permits **indefinitely**, and could not even be cancelled. The gate drained, every subsequent leaf activation on that silo queued behind it, and the silo presented as an activation outage. Measured on a frozen production container: gate ceiling 2, available 0, 345 activations queued, the oldest for 66.8 minutes.

A caller-side timeout does not help and is not what this option is. The sweep's grain call already times out at the Orleans response-timeout default and records `undelivered`; abandoning the caller's wait does nothing to the grain-side method, which keeps running and keeps holding its permit. The budget has to be enforced **inside** the permit-guarded region, which is what this option does.

**How the budget is enforced.** The drive creates a `CancellationTokenSource` for the budget and passes its token to every await in the region, and additionally bounds its own wait on the work with that token. Both halves are needed and they fail differently: the token genuinely terminates work that honours cancellation, while the wait bound covers a host-supplied storage call that ignores it. In the second case the storage call keeps running detached - but **without a permit**, which is the property that matters. The permit is acquired and released in the outer frame, the only frame guaranteed to run its `finally`, so abandonment cannot skip the release.

**Sizing.** The default is `4 * WalDrainBudget`. A full starved-leaf replay reads at most `MaxLeafReplayEntries` entries in slices of `WalReplaySliceBudget`, which at the shipped defaults is 40 slice reads; 5 minutes leaves 7.5 seconds per read against a healthy sub-millisecond read and a 15-second `WalFlushTimeout`, so a legitimately slow replay finishes comfortably inside it. Raise it if your storage tier is slow enough that real drives are being abandoned - the `orleans.lattice.wal.replay.starvation_drive_abandonments` counter tells you, and it is the only series that can. Lower it only if you are willing to trade drive completion for faster permit recovery.

**`InfiniteTimeSpan` is rejected, unlike most other timeout options here.** An infinite budget restores exactly the outage this option exists to bound, so it is refused at validation rather than honoured at runtime. The validator also rejects zero and negative values.

**Observability.** An abandoned drive increments `orleans.lattice.wal.replay.starvation_drive_abandonments` (tagged `tree`), zero-primed per tree beside the sweep's `attempted` arm so an absent series proves the build is not deployed rather than that no drive has run, and emits a warning log carrying the leaf, the tree, the budget and how long the drive actually ran. The warning takes one of two forms, split at the moment the drive acquired its replay permit (issue #3479). A drive that **was replaying** reports the time it spent being admitted to the replay gate, the time it spent replaying, and how far the replay advanced the leaf's checkpoint (in memory and persisted) before it was abandoned; only this form points at the storage provider and the leaf's replay gap. A drive that was abandoned **before it was admitted** never read storage, so its warning says so and says to raise the budget. GC drives never queue for a replay permit - one that finds no capacity is refused at once, counted on `orleans.lattice.saturation.refusals` under the `replay_permit_admission` source and returned as the `drove_admission_refused` verdict rather than abandoned - so this form arises only when admission itself outlasts the budget: a budget far too small, or a stall of the drive's thread while it is admitted. It also records the `drove_timed_out` verdict on `orleans.lattice.wal.gc.blocked_leaf_reactivations` - but that arm is near-silent in practice, because the caller has usually already timed out, so read the counter and not the verdict. See [Metrics](../metrics.md).

This option can be changed freely at any time. A drive reads the budget from its leaf's resolved options, which a leaf resolves when it activates, so a new value reaches each leaf on its next activation.

## `StorageUsageCacheTtl`

Cache lifetime for `ILattice.GetStorageUsageAsync` reports (default: 10 seconds). The per-tree storage-usage aggregator fans out across the tree's shards and WAL partitions to assemble a byte-accurate `TreeStorageUsageReport`; this TTL coalesces repeat callers (dashboard scrapes, the background poller, and direct API calls) so a single fan-out serves a whole window. Set to `TimeSpan.Zero` to disable caching - every call fans out fresh. See [Tree Storage](../tree-storage.md#measuring-retained-storage-at-runtime).

```csharp verify
// Hold storage reports for 30 s to cut dashboard fan-out cost
siloBuilder.ConfigureLattice("metrics-tree", o => o.StorageUsageCacheTtl = TimeSpan.FromSeconds(30));
```

This option can be changed freely at any time.

## `StorageUsagePollInterval`

Cadence at which every silo's background storage-usage poller calls `ILatticeAdmin.PollWalUsageAsync` so the WAL-bytes and over-threshold storage gauges (`orleans.lattice.storage.wal_bytes`, `orleans.lattice.storage.policy.over_threshold`) populate automatically, without any caller having to invoke `ILattice.GetStorageUsageAsync`. Default 15 seconds. The poll path is **leaf-free**: it activates only WAL partition grains, so idle trees stay cold. The poll fans out to every registered tree's WAL-only aggregator; each is a single cluster-wide activation, so its publish lands on its own host silo's metrics sink however many silos poll. The tree's deep aggregator publishes the same WAL-bytes series on its own host silo and is placed independently, so more than one silo can export a series for one tree: aggregate the storage gauges across silos with `max by (tree)`, not `sum by (tree)` (see [The `tree` dimension across aliasing](../metrics/tag-conventions.md#the-tree-dimension-across-aliasing)). Running the poller on every silo is intentional - it needs no leader election, and the aggregator's `StorageUsageCacheTtl` coalesces redundant polls from sibling silos. When a tree's aggregator migrates to another silo, the old silo stops refreshing that series and it expires from its sink after a staleness horizon (four times the slowest active poll interval, floored at 60 seconds), so a migration does not leave a duplicate series behind for good. This poll does not measure the snapshot-bytes and leaf-state-bytes gauges, and the total-bytes gauge publishes nothing for a tree until a deep path has measured it (see [`StorageUsageDeepPollInterval`](#storageusagedeeppollinterval)); after that, each poll recomputes the total from the freshly polled retained WAL payload plus the last deep snapshot and leaf-state figures, so between deep refreshes its WAL term is the retained payload rather than the physical occupancy a deep refresh uses.

This is a **global** knob read from the default (unnamed) options; per-tree overrides do not apply. Set to `TimeSpan.Zero` or a negative value to stop this WAL-bytes poll; while [`StorageUsageDeepPollInterval`](#storageusagedeeppollinterval) is also disabled (its default) the poller then does nothing, and the gauges populate only when the public storage-usage API is called.

```csharp verify
// Poll every 5 s for tighter dashboard freshness
siloBuilder.ConfigureLattice(o => o.StorageUsagePollInterval = TimeSpan.FromSeconds(5));

// Disable the poller (gauges populate only on explicit API calls)
siloBuilder.ConfigureLattice(o => o.StorageUsagePollInterval = TimeSpan.Zero);
```

This option is read once when the poller starts on each silo.

## `StorageUsageDeepPollInterval`

Optional cadence at which the same background poller *also* drives the **deep** storage gauges - `orleans.lattice.storage.snapshot_bytes`, `orleans.lattice.storage.leaf_state_bytes`, and `orleans.lattice.storage.total_bytes` - by calling the non-force `ILatticeAdmin.GetTotalStorageUsageAsync`. The faster [`StorageUsagePollInterval`](#storageusagepollinterval) poll refreshes only the WAL-bytes surface (it touches only WAL partition grains); this deep poll additionally reads each shard root's incrementally-maintained byte totals. That read is **O(1) per shard root** - it never walks the leaf chain or activates per-leaf snapshot grains - so it activates only the shard roots and never pins idle leaves resident. It never invokes the operator-only force-refresh (`ILatticeAdmin.RefreshStorageUsageAsync`) that re-walks every leaf.

Defaults to `TimeSpan.Zero`, which **disables** the deep poll: the snapshot / leaf-state / total-bytes gauges then populate only on demand via `ILattice.GetStorageUsageAsync` or the operator-driven `ILatticeAdmin.RefreshStorageUsageAsync`. Set a positive value - typically a small multiple of `StorageUsagePollInterval` - to keep the deep gauges live on a dashboard. A value at or below `TimeSpan.Zero` disables it. Like `StorageUsagePollInterval`, this is a **global** knob read from the default (unnamed) options; per-tree overrides do not apply. The sink's staleness horizon is sized off the slower of the two cadences, so a deep series survives a few missed deep polls before expiring after a real migration. While the deep poll is disabled, the deep gauges report *no measurement at all* for a tree rather than a synthesised zero, and the companion `orleans.lattice.storage.usage_deep_published` gauge reads `0` for that tree so the absence is explained rather than merely silent (see [Metrics](../metrics.md) and issue #2693).

```csharp verify
// Refresh the deep storage gauges once a minute (WAL bytes still refresh
// on the faster StorageUsagePollInterval cadence).
siloBuilder.ConfigureLattice(o => o.StorageUsageDeepPollInterval = TimeSpan.FromSeconds(60));

// Leave disabled (default): deep gauges populate only on explicit API calls.
siloBuilder.ConfigureLattice(o => o.StorageUsageDeepPollInterval = TimeSpan.Zero);
```

This option is read once when the poller starts on each silo.

## `StorageUsageRollupBudget`

Wall-clock budget a cluster-wide storage-usage roll-up may spend sampling trees before it stops dispatching and returns what it has (default: 20 seconds). Applies to `ILatticeAdmin.GetTotalStorageUsageAsync` and `ILatticeAdmin.RefreshStorageUsageAsync`.

Bounding the fan-out with [`MaxConcurrentStorageUsageTrees`](options-reference-2.md#maxconcurrentstorageusagetrees) and [`MaxConcurrentStorageUsageSurfaces`](options-reference-2.md#maxconcurrentstorageusagesurfaces) caps the *burst* a roll-up imposes, but it cannot cap the *total* work: a deep refresh re-walks every shard of every tree, so a large enough catalogue cannot be sampled inside one Orleans response deadline however gently it is dispatched. Without a budget the whole call then fails on the deadline and the caller learns nothing at all. With one, the trees sampled so far report real figures, the remainder report as not-answered, and `ClusterStorageUsageReport.Partial` is set - the same "an honest flagged lower bound beats a silently wrong or absent answer" rule the per-surface reporting follows.

Set it comfortably below the response deadline of the transport carrying the call, so the truncated report can still be returned. A non-positive value **disables** the budget, restoring run-to-completion behaviour.

This is a cluster-wide knob read from the default (unnamed) options by the admin grain that drives the roll-up; per-tree overrides do not apply, because that grain is not keyed by tree.

```csharp verify
// Allow a larger catalogue longer to sample before the roll-up truncates.
siloBuilder.ConfigureLattice(o => o.StorageUsageRollupBudget = TimeSpan.FromSeconds(45));

// Disable the budget: sample every tree however long it takes.
siloBuilder.ConfigureLattice(o => o.StorageUsageRollupBudget = TimeSpan.Zero);
```

This option can be changed freely at any time; a new value applies to the next roll-up.

## `TombstoneGracePeriod`

How long a deleted key's tombstone is retained before it becomes eligible for permanent removal by the compaction process (default: 24 hours). The grace period exists so that all cache replicas (`LeafCacheGrain` activations across silos) have time to observe the delete via delta replication before the tombstone disappears.

Set to `Timeout.InfiniteTimeSpan` to disable tombstone compaction entirely. This is useful for trees where deletes are rare or where tombstone accumulation is acceptable.

```csharp verify
// Compact aggressively (12 hours)
siloBuilder.ConfigureLattice(o => o.TombstoneGracePeriod = TimeSpan.FromHours(12));

// Disable compaction for a specific tree
siloBuilder.ConfigureLattice("archive-tree", o =>
{
    o.TombstoneGracePeriod = Timeout.InfiniteTimeSpan;
});
```

This option can be changed freely at any time. The new grace period takes effect on the next compaction reminder tick. The reminder interval is automatically set to match the grace period (clamped to a minimum of 1 minute, the Orleans reminder floor).

## `TxDecisionRetention`

Retention window for a completed saga's commit/abort decision in the per-tree transaction registry after the saga asks the registry to forget it (default: 60 seconds). The registry stamps a forgotten-at tombstone instead of evicting the decision; for the duration of the window the registry's status and snapshot reads continue to surface the decision so that a process which installs a *new* pending bucket on that txid *after* the saga's terminal fan-out can still resolve the verdict and apply the terminal directly.

The primary race the window guards is the retroactive shadow-forward sweep at the start of an adaptive shard split: the split coordinator replays every in-flight prepared mutation from the source leaves into the destination shard's pending-saga buckets, and its post-sweep cleanup pass resolves any orphan bucket whose terminal has already broadcast by reading the retained verdict. Without retention, a saga that completed microseconds before the sweep installed its pending bucket would leave a destination-shard orphan with no recoverable outcome.

Once the window elapses the decision is *masked* rather than deleted: the registry's status and snapshot reads report the saga as indeterminate ("a decision exists, and this tree is no longer entitled to report it") rather than as in flight. A txid the tree never recorded at all still reports in flight. Read paths hide a prepared key whose saga is indeterminate instead of falling through to its pre-saga value, because absence asserts nothing while in flight would be an affirmative and possibly wrong claim that the saga did not commit.

Expired tombstones are physically purged by the next forget call against the registry (an inline prune pass), or by a new saga's admission check when the registry shard is at its [`TxRegistryAdmissionBudgetBytes`](#txregistryadmissionbudgetbytes) budget; a repeated commit or abort mark carrying the *same* outcome is recognised as idempotent and deliberately leaves the tombstone in place, so it can never resurrect a decision the tree already retired. A tombstone held by a live point-in-time cursor pin is skipped by both the purge and the read-side mask, so a pinned snapshot keeps reading its decisions for as long as the pin lives. Set `TimeSpan.Zero` to restore the pre-tombstone immediate-evict semantic (legacy behaviour; reintroduces the orphan risk - reserved for tests or trees with `AutoSplitEnabled = false`). Increase beyond 60 s only if your operational profile produces sweep durations longer than that (very large shards under sustained write load, cascading split storms).

This option can be changed freely at any time.

## `TxRegistryAdmissionBudgetBytes`

Fail-safe admission bound on each saga decision registry shard's row (default: 768 KiB, `null` disables it). A tree's registry is split into [`TxRegistryShardCount`](#txregistryshardcount) shards, and each shard persists its whole state as one grain-state row, and every `ForgetAsync` leaves a tombstone in that row for [`TxDecisionRetention`](#txdecisionretention). Group commit (issue #3475) lets one tree run far more sagas per second than before. So the tombstone count, which is roughly throughput multiplied by retention, can now grow the row past what a storage provider accepts in a single row: about 1 MB for Azure Table grain state. Once the row is that large, every registry write fails, and with it every saga on the tree.

The bound refuses **new** sagas before they do any work, and never refuses work already under way. Before each new saga registers, the registry computes an O(1) estimate of its row size from dictionary counts. The estimate is deliberately weighted to over-count against the real JSON row. If the estimate is at or above the budget, the registry first purges any tombstones that have aged out of the retention window. If the estimate is still at or above the budget, the saga fails with a `LatticeSaturatedException` whose `SaturationSource` is `LatticeSaturationSource.TxRegistryCapacity`.

The refusal is retryable. A refused saga discards the transaction id it drew, so a retry mints a fresh id and, with more than one [`TxRegistryShardCount`](#txregistryshardcount) shard, may land on a shard with room; a single shard's capacity returns only as its tombstones age out, which takes seconds, so on an unsharded tree retry after a back-off. The library never retries a refused single-tree write. A cross-tree write has already recorded its preparing phase and armed its one-minute keepalive reminder before it dispatches the sub-sagas, so while the caller still receives the refusal, the keepalive re-dispatches the prepare phase about once a minute until the sub-sagas are admitted. `MarkCommittedAsync`, `MarkAbortedAsync`, `ForgetAsync`, recovery, and every status read are never refused, so in-flight sagas always complete, and completing them is what frees room.

Ceiling maths. One tombstone costs about 116 bytes of JSON, so a 1 MB row holds about 8,600 tombstones. At the default 60 s retention, that is about 143 sagas/s sustained per registry shard. The estimate weights a tombstone at 128 bytes (its decision entry plus its forgotten-at stamp), so the 768 KiB default admits about 6,100 tombstones, or about 100 sagas/s per shard at 60 s retention, and leaves headroom for concurrent admissions to overshoot, since each check is a probe and not a reservation. These figures are **per registry shard**: each shard admits against its own row and its own budget, so a tree with eight shards sustains about 800 sagas/s at 60 s retention. To raise the ceiling further, raise [`TxRegistryShardCount`](#txregistryshardcount), which is the preferred lever because it changes no correctness window. Shortening `TxDecisionRetention` also raises it, subject to that option's own safety guidance.

```csharp verify
siloBuilder.ConfigureLattice("orders", o =>
{
    o.TxRegistryAdmissionBudgetBytes = 512 * 1024;
});
```

Set `null` only on a storage provider with no practical per-row limit. The value must be at least 1 when set. This option can be changed freely at any time.

## `TxRegistryShardCount`

Number of saga decision registry shards that new sagas are minted across, per tree (default: 1, range 1 to 256). Each shard is a separate registry grain activation keyed `_lattice_txshard_{n}_{treeId}`, with its own persisted row, its own [`TxRegistryAdmissionBudgetBytes`](#txregistryadmissionbudgetbytes) budget, and its own decisions revision. The sustained atomic-saga rate a tree can retain within [`TxDecisionRetention`](#txdecisionretention) therefore scales linearly with the shard count: about 100 sagas/s per shard at the default budget and retention.

A saga's shard is chosen when its transaction id is minted, and the shard index is stamped into the id itself (a version-8 UUID). Every later registry call for that saga - from the saga coordinator, a leaf resolving a pending intent, a shard root, a split, or a replication receiver - routes to the owning shard from the stamped index alone, never from this setting. Tree-wide reads (multi-key reads, scans, cursors, backups, and replication snapshots) fan out over every shard up to the tree's durable shard high-water mark, which a shard raises before its first write, plus the legacy registry, and sum the per-shard revisions; a per-silo coalescer shares one fan-out across the concurrent reads of the same tree.

**Opt-in.** The default of `1` keeps the pre-sharding layout: one registry per tree, keyed by the bare tree id, and ordinary version-4 transaction ids that route to it. A silo running an older version resolves every txid against that legacy registry, so raise the value only once every silo, and every replication peer that applies this cluster's sagas, runs a version that understands sharded ids. Transaction ids minted before the change keep routing to the legacy registry, which tree-wide reads always include, so enabling sharding needs no migration and the legacy row drains within one retention window.

**Changing the value.** Because routing never reads the setting and tree-wide reads cover the durable high-water mark, silos configured with different values still agree on every read, and changing the value on a live cluster is safe in either direction. Raising it spreads new sagas across more shards; lowering it, to `1` included, reroutes nothing and still reads every shard already written to.

Eight shards lift the per-tree ceiling from about 100 to about 800 sagas/s while keeping a tree-wide read at nine registry calls (eight shards plus the legacy registry) and one high-water read.

```csharp verify
siloBuilder.ConfigureLattice(o =>
{
    o.TxRegistryShardCount = 16;
});
```

This option is read from the global (unnamed) options, so a per-tree override has no effect.

## `VersionVectorRetention`

Declared retention window for `VersionVector` replica entries (default: `InfiniteTimeSpan`, no pruning). **No Lattice component currently reads this option, so setting it has no effect**: nothing in the library prunes a version vector with a cutoff derived from it, and no cache expunges vectors on its schedule. The pruning primitive it describes is the public `VersionVector.PruneOlderThan(long minRetainedUtcTicks)`, which removes every replica entry whose wall-clock tick is older than the cutoff. A host that prunes vectors itself must apply the same cutoff on every replica that merges against them, because a pruned entry is reinstated by the next merge with a replica that still holds it.

There is an advisory lower bound, `LatticeOptions.DefaultMinVersionVectorRetention` (1 hour). It is **not enforced** - it is provided only as a reference constant - but a cutoff below it is typically unsafe on networks where clock skew exceeds the window, because pruning may then drop entries that are still causally relevant.

## `WalBytePressureReclaimTarget`

Low-water fraction of `WalMaxRetainedBytes` that disarms the advisory byte-pressure policy (default: `0.8`), providing hysteresis so a tree hovering near the ceiling is not trimmed on every GC pass. The policy *arms* when the WAL's sampled on-disk occupancy (see [`WalMaxRetainedBytes`](options-reference-4.md#walmaxretainedbytes)) crosses the full ceiling (high-water) and re-triggers a byte-pressure trim on each pass until a sample, taken before or after a pass's trim, finds that occupancy at or below `WalMaxRetainedBytes * WalBytePressureReclaimTarget` (low-water), at which point it disarms. While disarmed, growth that stays inside the `(low-water, ceiling]` band does not re-trigger. The value is not validated at startup but is sanitised when a pass evaluates it: a value above `1` is treated as `1`, and a non-positive or `NaN` value falls back to the default `0.8`. Ignored when `WalMaxRetainedBytes` is `null`. See [WAL](../wal.md) and [Tree Storage](../tree-storage.md).

This option can be changed freely at any time. The new value takes effect on the next GC tick.

## `WalMaxBatchBytes`

Maximum byte budget the WAL partition grain accumulates into a single storage flush (default: **4 MiB**, `4L * 1024 * 1024`). Whichever of `WalMaxBatchBytes` or `WalMaxBatchEntries` is reached first triggers the flush. Measured against the *exact* serialised size of each `WalRecord` under the WAL grain's wire format - the per-entry encoder walks every field through the same Orleans-binary codec the storage provider sees, and the bytes it produces are handed straight to `IWalStorageProvider.AppendEncodedBatchAsync`, so the grain pays exactly one encode per append and the budget is an exact ceiling for any batch of more than one entry (a single entry larger than the whole budget is flushed alone rather than refused), suitable for sizing against the Azure Table Storage 4 MB transactional-batch limit (which has zero tolerance for under-counts).

This option can be changed freely at any time. The new value takes effect on the next batch boundary.

## `WalMaxBatchEntries`

Maximum number of WAL entries the partition grain coalesces into a single storage flush (default: 100). Lower values reduce per-entry flush latency at the cost of throughput. Whichever of `WalMaxBatchEntries` or `WalMaxBatchBytes` is reached first triggers the flush.

This option can be changed freely at any time. The new value takes effect on the next batch boundary.

## `WalFlushTimeout`

Hard ceiling on how long a single WAL partition flush may run before it is cancelled and surfaced to callers as a `TimeoutException` (default: 15 seconds). The ceiling covers both the storage-provider append and the post-failure tail resync.

Bounding the flush is what keeps a provider call that hangs indefinitely - for example against a partition left half-activated by a placement/reshard race - from pinning its in-flight slot forever. Without a ceiling the hung slot is never removed from the in-flight chain, the chain saturates at `WalMaxPendingBatches`, and every subsequent append back-pressures behind a flush that can never settle (a steady-state stall with no fault and no activation recycle). With the ceiling the hung flush faults cleanly, the existing failure handler resynchronises the dense-offset tail from the provider, drains the chain, and callers retry.

The default of 15 seconds sits above the Azure Tables SDK's worst-case legitimate retry envelope under sustained throttling (~10 seconds: three exponential backoffs plus the call times), so a healthy flush never trips it, yet well below the SDK's per-try network timeout so a true hang is still caught and the wedged shard self-heals promptly. Set to `InfiniteTimeSpan` to disable the ceiling and restore the historical unbounded-await behaviour; the options validator rejects any other non-positive value.

This option can be changed freely at any time. The new value takes effect on the next flush.

## `WalAppendDispatchTimeout`

Hard ceiling on how long the per-tree WAL writer will wait on a single outbound append dispatch to a WAL partition grain before abandoning the await and surfacing a `TimeoutException` to the caller (default: 30 seconds).

The dispatch is the writer-side cross-grain RPC into the WAL partition grain - it is the outermost observable seam on the write pipeline and was historically unbounded on the writer side, so a wedged partition activation would hold every caller's dispatch parked until the Orleans response deadline (30 seconds by default) expired - a blind hang with no per-shard attribution. Bounding the dispatch converts that blind hang into a structured fault with per-shard counter attribution (the `orleans.lattice.wal.append_dispatch.timeouts` counter, tagged `tree` and `shard`), so a wedged shard surfaces immediately and the request pipeline releases its slot rather than back-filling behind the wedge.

This option does **not** fix the wedge mechanism itself - the grain-side flush / activation deadlines already bound their own regions - it bounds the symptom on the writer side and makes every wedge attributable to a specific `(tree, shard)` in O(timeout) instead of O(response timeout) time.

The default of 30 seconds sits above the legitimate envelope of a fully-saturated dispatch (one healthy flush + headroom). It is the same figure as the Orleans default response timeout, so on a stock silo the two deadlines coincide and what this bound adds is the typed, per-shard-attributed fault; it surfaces a true park ahead of the Orleans deadline only on a silo whose `ResponseTimeout` is raised above it. Set to `InfiniteTimeSpan` to disable the ceiling and restore the historical unbounded-await behaviour; the options validator rejects any other non-positive value.

This option can be changed freely at any time. The new value takes effect on the next dispatch.

## `WalFlushPreflightTimeout`

Hard ceiling on how long a WAL partition's `FlushAsync` may spend in its preflight region (the synchronous setup and initial scheduler yield that precede the bounded provider call) before the flush is abandoned and the slot drains (default: 5 seconds).

The preflight region is normally microseconds, but if the activation's grain scheduler never resumes the post-yield continuation (e.g. a startup reshard / membership change parked the activation, a non-cooperative work item is hogging the scheduler, or the activation is being torn down mid-flush) the slot sits in `_inFlight` with no deadline armed - the existing `WalFlushTimeout` only covers the provider call, which has not yet been issued - and the chain saturates at `WalMaxPendingBatches` with no fault and no activation recycle. With the ceiling the parked preflight faults cleanly as a `TimeoutException` routed through the normal failure handler, the slot drains, and the `orleans.lattice.wal.flush.preflight.timeouts` counter (tagged `tree` and `shard`) attributes the trip to the affected partition.

The default of 5 seconds is orders of magnitude above the legitimate microsecond envelope, yet small enough that a genuinely stalled scheduler is caught before the writer-side dispatch deadline (`WalAppendDispatchTimeout`) trips. Set to `InfiniteTimeSpan` to disable the ceiling and restore the historical unbounded-await behaviour; the options validator rejects any other non-positive value.

This option can be changed freely at any time. The new value takes effect on the next flush.

## `WalDrainBudget`

Hard ceiling on how long a WAL partition grain's `OnDeactivateAsync` drain may run before the remaining in-flight slots are force-faulted and the chain is released so the activation can finish tearing down (default: 75 seconds = `5 * WalFlushTimeout`). Bounds the host-level SIGTERM drain so the silo's shutdown accounting (the benchmark host's `FINAL` line, an `IHostApplicationLifetime.ApplicationStopping` cancellation source) always settles within bounded time of the SIGTERM, regardless of whether the underlying storage provider is healthy.

Defends against the saturating-storage-account wedge: when the provider call's await is parked behind an SDK retry loop in pre-attempt back-off, the existing per-flush `WalFlushTimeout` may not fire promptly (the SDK observes cancellation only between attempts, not during back-off), so a chain with N in-flight slots can hold the deactivation indefinitely. With this budget the drain:

1. Signals every in-flight flush's cancellation token at entry (the per-activation drain `CancellationTokenSource` is linked into each per-flush deadline at flush construction, so a single `Cancel()` cancels every in-flight provider call in one shot);
2. Awaits the chain to settle naturally for up to `WalDrainBudget`;
3. Force-faults any slot that has not unlinked when the budget expires, with a typed `TimeoutException` faulted onto every parked ack `TaskCompletionSource` so callers parked on `AppendAsync` / `AppendBatchAsync` are released rather than parking through the rest of host shutdown.

The `orleans.lattice.wal.shard.drain.budget.expirations` counter and `orleans.lattice.wal.shard.drain.budget.force_faulted_slots` histogram (both tagged `tree` and `shard`) attribute every budget-driven force-fault per partition. A zero counter on a healthy drain; any non-zero rate identifies a shard whose provider call could not be cancelled inside the budget.

The default of 75 seconds is `5 * WalFlushTimeout` - sized so a healthy chain with cap = 16 in-flight flushes has time to drain naturally (each flush is itself bounded by `WalFlushTimeout`) while a wedged chain still surfaces within a bounded window of the SIGTERM. Set to `InfiniteTimeSpan` to disable the ceiling and restore the historical unbounded-drain behaviour; the options validator rejects any other non-positive value.

This option can be changed freely at any time. The new value takes effect on the next deactivation.

## `WalSaturationSampleInterval`

Cadence at which the silo-scoped sampler that backs `IWalSaturationSignal` and `IWalSaturationObserver` recomputes the per-tree saturation state from the writer-side admission gate and the recent dispatch-timeout-trip rate (default: 200 ms). A shorter interval lowers the worst-case transition latency observers see (the bound is one sample interval beyond the underlying signal crossing the threshold) at the cost of slightly more timer-driven sampler work. The 200 ms default keeps subscribers well within the one-second transition-latency promise on the public surface while keeping the sampler at a negligible CPU footprint on an idle silo. Set to `InfiniteTimeSpan` to disable the sampler entirely - every tree's signal stays `Healthy` forever, the observable `orleans.lattice.wal.saturation.state` gauge publishes no series (a tree appears in it only once the sampler has observed it), and `IWalSaturationObserver` callbacks never fire. The options validator rejects any other non-positive value. See [WAL Saturation Signal](../wal-saturation-signal.md).

This is a **global** knob read from the default (unnamed) options, and it is read once, when the sampler starts: a new value takes effect on the next silo start, and per-tree overrides do not apply. The sampler reads every other classifier option it consults - the `WalSaturation*` thresholds, windows and switches below, and `WalDrainLagConsumerFreshness` / `WalDrainLagHolderLogInterval` - from the same default instance on each tick, so a per-tree override of any of them is ignored by the classifier. (A per-tree `WalSaturationFlushLatencyThreshold` still sets the threshold at which that tree's WAL shards count slow flushes, but the classifier acts on those trips only when the default instance sets a threshold too.)

## `WalSaturationThrottledRatio`

Per-partition admission-depth ratio (in `[0.0, 1.0]`) at or above which the saturation signal raises a tree to `WalSaturationState.Throttled` (default: 0.75). Computed as `in_flight / WalMaxPendingBatches` on each partition; the tree's state is the worst case across its partitions. Below the ratio the tree stays `Healthy`; at or above the ratio - including at the cap itself - it advances to `Throttled`. `Saturated` is reserved for the acute inputs (the dispatch-timeout rate crossing `WalSaturationDispatchTimeoutThreshold`, provider failures, sustained flush latency); a partition at the cap with a non-empty wait queue advances to `Saturated` only when [`WalSaturationAcuteOnly`](options-reference-4.md#walsaturationacuteonly) is `false`. The 0.75 default leaves a 25%-of-cap headroom for callers to slow down before the cap pins. Must be in the inclusive range `[0.0, 1.0]`; `NaN` is rejected.

This option can be changed freely at any time. The new value takes effect on the next sampler tick.

## `WalSaturationDispatchTimeoutThreshold`

Minimum number of `orleans.lattice.wal.append_dispatch.timeouts` trips observed within a single `WalSaturationSampleInterval` window that raises a tree to `WalSaturationState.Saturated` regardless of admission-semaphore depth (default: 1). Captures the dispatch-deadline failure-tail of the saturation regime (parked dispatches abandoned because a downstream shard wedged) in addition to the admission-depth fast signal. Raise it on dashboards where occasional single trips are expected without operator concern; the value is per-window, so `WalSaturationDispatchTimeoutThreshold = 3` with a 200 ms sample interval permits up to 2 trips per window - 10 trips/second - steady-state before flagging Saturated. Must be greater than or equal to 1.

This option can be changed freely at any time. The new value takes effect on the next sampler tick.

## `WalSaturationProviderFailureRateThreshold`

Minimum number of provider-side commit failures (any exception surfaced from a downstream WAL shard append dispatch other than the writer-side `TimeoutException` already captured by `WalSaturationDispatchTimeoutThreshold`, and other than the cancellations and shutdown refusals excluded below) observed within a single `WalSaturationSampleInterval` window that raises a tree to `WalSaturationState.Saturated` regardless of admission-semaphore depth and dispatch-timeout trips (default: 1). Captures the third saturation regime the writer side cannot otherwise surface: a downstream storage provider whose commit calls return quickly (so neither the admission depth nor the dispatch deadline crosses the threshold) but terminally fail at a high rate, e.g. an Azure Tables single-account 409-Conflict burst where the SDK retry races a server-side-already-committed transaction.

Without this input the caller saw the failure tail (a `SetAsync` / `SetManyAsync` faulted) but the per-tree saturation signal stayed `Healthy` and any back-pressure consumer (the bench TCP reader, an upstream load balancer) had no leading-edge surface to slow down before the leak became visible at the operator level. Cancellations and shutdown refusals are excluded from the counter: every `OperationCanceledException`, whichever token it carries, and every `LatticeShuttingDownException` from this writer's or a downstream peer's drain gate, so neither a caller-side abandonment nor a drain release inflates the saturation signal.

Set to `0` to disable the trigger entirely (matches the `Timeout.InfiniteTimeSpan` sentinel on the sample-interval option). The validator rejects any other negative value. This option can be changed freely at any time; the new value takes effect on the next sampler tick.

## `WalSaturationFlushLatencyThreshold`

Per-provider-flush wall-clock latency at or above which a WAL partition, after each provider flush, increments a per-(tree, shard) flush-latency trip counter that feeds the saturation classifier (default: `null`, disabled). When the threshold is set and the classifier observes a non-zero delta on the trip counter in each of the last `WalSaturationFlushLatencySampleWindows` sample windows in a row, the tree is upgraded to `WalSaturationState.Saturated` regardless of admission-semaphore depth, dispatch-timeout trips, or provider-failure trips. The sampler decides whether the input is on from the default (unnamed) options, like every other classifier setting (see [`WalSaturationSampleInterval`](#walsaturationsampleinterval)); a per-tree value sets only the latency at which that tree's WAL partitions count a trip, and the classifier acts on those trips only while the default instance sets a threshold too.

Captures the **small-batch blind spot** the existing three Saturated inputs cannot see: a workload that issues many small `SetAsync` calls against a saturating storage account never fills the per-partition admission semaphore (every batch is one entry, in-flight is rarely above 1-2), never trips `WalAppendDispatchTimeout` (the dispatch returns quickly with a slow but successful provider flush), and never tallies a `WalSaturationProviderFailureRateThreshold` trip (the flush succeeds, just slowly). The flush-latency input observes the same regime via the per-flush wall-clock cost. Sizing guidance: pick a threshold a few times your steady-state p99 provider flush latency (e.g. 500 ms when steady-state p99 is ~80 ms) so the trip counter stays at zero during healthy traffic and only ticks under genuine saturation.

The input is purely additive. Leaving the threshold at its default `null` is a zero-cost no-op: the WAL partition skips the trip-counter increment entirely and the classifier behaves exactly as it shipped before the input was introduced. Must be positive when set; the validator rejects `TimeSpan.Zero` and any negative value.

This option can be changed freely at any time. The new value takes effect on the next provider flush.

## `WalSaturationFlushLatencySampleWindows`

Number of consecutive `WalSaturationSampleInterval` sample windows that must each observe a non-zero `WalSaturationFlushLatencyThreshold` trip-counter delta before the classifier upgrades the tree to `WalSaturationState.Saturated` (default: 3). Acts as the noise floor for the flush-latency input - a single slow provider flush in an otherwise healthy window is normal jitter; three consecutive windows each containing at least one slow flush is the leading edge of a saturation regime. At the default 200 ms `WalSaturationSampleInterval` the minimum sustained-slow-flush duration that triggers `Saturated` is ~`3 * 200 ms = 600 ms`.

Set lower (minimum 1) to make the input more sensitive at the cost of more transient classifier flaps; set higher to lengthen the sustained-slow regime the classifier requires before flagging. Has no effect when `WalSaturationFlushLatencyThreshold` is left at its default `null`. The validator rejects values less than 1.

This option can be changed freely at any time. The new value takes effect on the next sampler tick.

## `WalSaturationMaterialiserLagThreshold`

Materialiser drain-lag duration above which the saturation sampler records a pressure level that feeds the classifier (default: 30 seconds; set to `null` to disable the input entirely). On every sampler tick (default 200 ms) the sampler measures the drain lag directly from in-memory state as the WAL head wall-clock timestamp (the newest routed entry's HLC, tracked by the commit-log writer) minus the slowest eligible cursor in the in-memory cursor registry (its lag-plane minimum over every registered consumer, leaf materialisers and tree-wide consumers alike, after the exclusions described under [`WalDrainLagConsumerFreshness`](options-reference-4.md#waldrainlagconsumerfreshness)), clamped at zero - a head-relative measure, so an idle but caught-up tree (the frontier reaches the head) reads zero and never trips, and a never-checkpointed block pin is treated as zero lag rather than the full head age. Because the lag is recomputed live each tick rather than at a WAL GC pass, the input engages immediately on a write spike instead of waiting for the next collection. When the lag stays strictly above this threshold for `WalSaturationMaterialiserLagSampleWindows` consecutive sampler windows the tree is held at `WalSaturationState.Throttled`. Unlike the dispatch-timeout, provider-failure, and flush-latency inputs, the drain-lag input never escalates to `Saturated`: it is a pure back-off that slows callers without ever tripping the writer admission gate's `LatticeSaturatedException` fast-fail. The resulting `Throttled` state engages the local WAL writer's per-append pacing (`WalThrottledAdmissionPace`) on the single-silo write path and the replication receiver flow control when a replicating peer exists (except on an aliased tree; see [Replication options](../configuration.md#replication-options)), so upstream writers slow before the drain backlog grows unbounded.

Captures the **drain-path blind spot** the flush-latency and admission inputs cannot see: a write burst can be accepted and flushed quickly (healthy flush latency, shallow admission depth) yet outrun the rate at which leaf materialisers project committed WAL entries into the tree, so the durable backlog and its pin floor grow while every other saturation input reads healthy. Sizing guidance: pick a threshold a few times your steady-state materialiser drain lag so the observation stays quiet during healthy traffic and only fires when projection genuinely falls behind ingest.

The input is on by default, at the 30-second threshold above. Setting the threshold to `null` disables it at zero cost: the sampler skips the lag computation and the classifier behaves exactly as it did before the input existed. Must be positive when set; the validator rejects `TimeSpan.Zero` and any negative value.

This option can be changed freely at any time. The new value takes effect on the next sampler tick.

Previous: [Options Reference: MaterialiserCheckpointInterval to ShardHealingCooldown](options-reference-2.md). Next: [Options Reference: WalDrainLagConsumerFreshness to WalMaterialiserPinShards](options-reference-4.md). Contents: [Configuration](../configuration.md).
