Table of Contents

Options Reference: WalDrainLagConsumerFreshness to WalMaterialiserPinShards

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at options-reference-4.md, and llms.txt lists every page.

Part of Options Reference, in Configuration.

WalDrainLagConsumerFreshness

Freshness window a WAL cursor report must fall inside to contribute to the materialiser drain-lag input of the saturation classifier (default: 5 minutes; TimeSpan.Zero disables the exclusion and restores the historical all-consumers behaviour). A consumer whose latest report is older stays fully registered - it still pins the WAL GC trim floor - but is left out of the lag-plane minimum, so an idle leaf in a live tree cannot hold the tree permanently Throttled. The same window also judges a leaf materialiser's position: a leaf whose cursor has not advanced within the window, and whose cursor itself predates the window, is left out of the lag plane even when its report is fresh (issue #3131). A leaf's cursor is the newest change applied to its own key range, while the WAL head is tree-wide, and a leaf re-reports that persisted position on every activation - so a leaf whose range simply received no writes would otherwise read as hours of phantom lag on a freshly started silo. A leaf that is draining advances its cursor and stays counted however far behind it is, and tree-wide consumers (view maintainers, WAL subscribers, replication shippers) are judged on report age alone, so a stalled tailer still registers as lag. When a tree crosses WalSaturationMaterialiserLagThreshold the sampler logs up to three eligible minimum holders, their report ages, and how long ago their positions last advanced; WalDrainLagHolderLogInterval controls repeat observations during a standing breach. The classifier gates only the advisory WalThrottledAdmissionPace, so the exclusion cannot permit trimming or lose data. The validator rejects a negative value.

WalDrainLagHolderLogInterval

Minimum interval between repeat holder-log observations while a tree remains strictly above WalSaturationMaterialiserLagThreshold (default: 10 minutes; null keeps edge-only behaviour). The first observation is immediate on the crossing tick, before the consecutive-window classifier need have reached Throttled. Each observation emits up to three Warning lines from the silo's saturation sampler, one per eligible consumer in ascending cursor order, under a shared ObservationId and ObservedAtUtc. The interval limits observation groups, not individual lines. Recovery at or below the threshold clears the tree's last-log state, so the next crossing logs immediately; disabling the drain-lag input also clears it. No periodic observations are emitted when that input or Warning logging is disabled.

This global option is read from the default (unnamed) options on every sampler tick; per-tree overrides are ignored. Rate limiting uses monotonic elapsed time, independently per tree and per silo. Must be positive when set. See drain-lag holder observations for the structured fields, identity disclosure and interpretation. The holder selection reuses the existing over-threshold snapshot and does not change the materialiser drain loop or add metric tags.

WalSaturationMaterialiserLagSampleWindows

Number of consecutive saturation-sampler windows that must each observe a materialiser drain-lag level strictly above WalSaturationMaterialiserLagThreshold before the classifier holds the tree at WalSaturationState.Throttled (default: 3). Acts as the noise floor for the drain-lag input, mirroring WalSaturationFlushLatencySampleWindows, so a single sampler tick cannot flip the regime.

Set lower (minimum 1) to make the input more sensitive at the cost of more transient classifier flaps; set higher to lengthen the sustained-lag regime the classifier requires before flagging. Has no effect when WalSaturationMaterialiserLagThreshold is set to null. The validator rejects values less than 1.

This option can be changed freely at any time. The new value takes effect on the next sampler tick.

WalSaturationMaterialiserPinLatencyThreshold

Duration at or above which a durable leaf-materialiser pin write counts as a saturation pressure trip (default: null, which disables the input entirely). The reporting leaf measures each pin write at its own call site and counts a trip when the call reaches this threshold or faults. When trips occur in WalSaturationMaterialiserPinLatencySampleWindows consecutive sampler windows the tree is held at WalSaturationState.Throttled.

Closes the durable-floor blind spot (issue #2015). Every other saturation input, the drain-lag input included, is derived from in-memory state: drain lag measures the WAL head against the in-memory cursor registry, which keeps advancing perfectly well while the durable pin store is stalled. But the floor the WAL garbage collector actually trims against is the durable one, so a stalled pin store leaves the WAL growing without bound while every existing input reads Healthy - which is exactly what happened in issue #2012. This input measures the durable write itself, so that condition becomes observable and actionable.

Like the drain-lag input, it holds the tree at Throttled and never escalates to Saturated. That is deliberate: a stalled pin store is a retention-floor maintenance problem, not an inability to accept writes. Escalating would engage the writer admission gate's LatticeSaturatedException fast-fail and convert slow WAL trimming into user-visible write failures, making the very incident it detects worse. Slowing producers is the correct response - it gives the pin store room to drain.

Measured at the call site, so the reading includes time a report spent queued ahead of the shard's non-reentrant activation, which is what the reporting leaf actually experiences and what a saturation signal should reflect. Sizing guidance: set the threshold well above the duration of a healthy durable pin write, so healthy traffic stays quiet. orleans.lattice.materialiser.pin.durable_write_latency records only the writes that met the threshold or faulted (and nothing while the input is disabled), so it shows the trips rather than a steady-state baseline to size from.

The input is purely additive. Leaving it at its default null is a zero-cost no-op: nothing is recorded, the sampler's map stays empty, and the classifier behaves exactly as before. Must be positive when set; the validator rejects TimeSpan.Zero and any negative value.

This option can be changed freely at any time. The new value takes effect on the next pin report. It is read from the default (unnamed) options - by the reporting leaf as well as by the sampler - so a per-tree override has no effect.

WalSaturationMaterialiserPinLatencySampleWindows

Number of consecutive saturation-sampler windows that must each observe at least one durable pin-write latency trip before the classifier holds the tree at WalSaturationState.Throttled (default: 3). Acts as the noise floor for the pin-latency input, mirroring WalSaturationFlushLatencySampleWindows, so a single slow write cannot flip the regime.

Set lower (minimum 1) to make the input more sensitive at the cost of more transient classifier flaps; set higher to lengthen the sustained-slow regime the classifier requires before flagging. Has no effect when WalSaturationMaterialiserPinLatencyThreshold is left at its default null. The validator rejects values less than 1.

This option can be changed freely at any time. The new value takes effect on the next sampler tick.

WalSaturationRecoveryWindow

Window after the most-recently observed WalSaturationState.Saturated transition during which the classifier holds a tree at or above Throttled even if the current sampler tick's per-partition depth observation would otherwise classify it as Healthy (default: 1 second). Defends against bursty per-partition WAL drain where one partition fills to cap, drains entirely in the next tick, and the next partition fills - the per-tick max(depth_ratio) across partitions oscillates between ~1.0 and ~0.0 within a single sampler period and the classifier would otherwise flap Healthy <-> Saturated at the sampler cadence with Throttled never observed as a stable state. Under the default WalSaturationAcuteOnly classification an at-cap partition already reads Throttled, so that depth-driven flap arises only with WalSaturationAcuteOnly = false; the window still holds a tree at Throttled after the acute causes that raise Saturated. With the window in effect, callers see Throttled persist as the natural lead-up and fall-back regime around saturation episodes; the canonical TCP / queue ingest reader pattern can take the advisory Throttled action (yield-per-line, lower-priority dispatch) for measurable durations rather than seeing only the binary pause-or-go pattern.

The window does NOT affect the Healthy -> Saturated transition latency - Saturated still fires immediately on the current tick's saturating condition, so the public saturation-signal surface's bound (transition latency under one WalSaturationSampleInterval) is preserved. It does NOT affect the recovery path either: once the window elapses AND the current tick observes no saturation pressure, the tree drops to Healthy and any pending IWalSaturationSignal.WaitForHealthyAsync completes; the window only delays the recovery by the configured value.

Set to TimeSpan.Zero to disable the upgrade entirely and restore the per-tick classifier behaviour the sampler shipped with. Set to Timeout.InfiniteTimeSpan to hold Throttled forever after the first Saturated observation - useful for tests that want a sticky Throttled floor without arming a wall-clock dependency, or for defensive deployments that prefer the saturation regime to be sticky. The validator rejects any other negative value.

This option can be changed freely at any time. The new value takes effect on the next sampler tick (or the next tick that would otherwise upgrade a tree, if the tree was Saturated more than WalSaturationRecoveryWindow ago).

WalSaturationRecoveryReleaseBatch

Maximum number of parked WAL-admission waiters a recovered partition admits per sampler tick (default: 16). Bounds the burst a recovery hands back to the admission pipeline, so a partition that has just dropped to Healthy is not immediately re-saturated by the entire population that was waiting on it.

The defect this closes is a thundering herd. IWalSaturationSignal.WaitForHealthyAsync parks one waiter per blocked WAL dispatch, and a recovery used to complete every one of them in a single pass. Once the parked population exceeds the partition's admission capacity (WalMaxPendingBatches, default 16), the released herd re-fills the pipeline to its cap before any meaningful drain has occurred, the classifier flips straight back to Saturated (in this depth-driven form, when WalSaturationAcuteOnly is false, under which an at-cap partition reads Saturated), and the cycle repeats with no net progress. Each failed cycle costs its callers a full WalAdmissionSaturationWaitBudget while holding a concurrency slot, so at scale the gate stops behaving as a back-pressure valve and becomes an absorbing state: recovery is observed, but no caller ever benefits from it.

Pacing the release turns that into a metered drain. The oldest-parked waiters are released first, so a paced drain cannot starve the callers closest to exhausting their wait budget, and the residue stays parked under the same key. Release is level-triggered: every tick that observes the partition Healthy releases a further batch, not only the tick that observes the Saturated -> Healthy transition. That is what guarantees the residue drains - the later ticks see no transition, so an edge-triggered release would strand every waiter beyond the first batch until the next saturation episode.

The default (16) equals WalMaxPendingBatches, so one tick admits exactly enough work to refill the partition's pipeline once. At the default WalSaturationSampleInterval of 200 ms that is roughly 80 admissions per second per partition, which comfortably exceeds the offered per-partition append rate of a saturating multi-silo workload, so pacing bounds the recovery burst without becoming the new bottleneck. Raise it if recovery is observed to be the limiting factor on a workload whose partitions drain far faster than the sampler cadence; lower it to meter recovery harder on a provider with a long tail.

Set to 0 to disable pacing entirely and restore the previous release-every-parked-waiter behaviour. The validator rejects negative values.

This option can be changed freely at any time. The new value takes effect on the next sampler tick.

WalSaturationAcuteOnly

When true, only acute causes - dispatch-timeout trips, provider failures, and sustained flush latency - classify a WAL partition as Saturated (default: true). A partition whose admission semaphore is merely at its cap then reads Throttled, and a caller parked at the writer's admission gate resumes as soon as its partition leaves Saturated instead of waiting for Healthy.

At-cap is the steady state of a well-pipelined partition, not a fault: the admission semaphore already bounds in-flight work at the cap, and that bound is the back-pressure. Classifying it Saturated closes the admission gate on healthy traffic, and the gate then holds each parked append until Healthy - a condition the WalSaturationRecoveryWindow hysteresis and the paced WalSaturationRecoveryReleaseBatch release both defer. A caller arriving during that same Throttled window passes the gate without parking - it pays only the WalThrottledAdmissionPace delay - so the parked caller was held on a stricter condition than the one that admits a newcomer. On the eight-silo set-many rig this gate wait was the dominant per-append cost (see issue #3348).

The historical at-cap classification is that defect, so the corrected verdict is the default. It changes when IWalSaturationSignal reports Saturated to every consumer (replication flow control, the atomic-write quiesce, cursors, view back-pressure, scaling pressure and the UnhealthyOnWalSaturated health check, dashboards): an at-cap partition now reads Throttled. It only ever reports Saturated less often, so callers observe LatticeSaturatedException less often and never in a new place. Set it to false to restore the historical classification exactly. The sampler reads this value silo-wide, like the other WalSaturation* classifier options, so set it globally rather than per tree; the WAL writer's admission gate reads the tree's own value only to decide whether a parked append resumes when its partition leaves Saturated or waits for Healthy.

This option can be changed freely at any time. The new value takes effect on the next sampler tick and the next gated append.

WalAdmissionSaturationWaitBudget

Wall-clock budget the WAL writer's pre-admission saturation gate spends parked, before an append enters its WAL partition's admission semaphore, while that partition's saturation verdict stays Saturated; when it expires the dispatch is refused with LatticeSaturatedException (default: 5 seconds). Closes the consumer-coverage gap where the admission semaphore was previously signal-blind: under the storage-account 409-Conflict regime the classifier raised Saturated many times before the first observable failure, but every new dispatch still admitted into the semaphore and parked at the cap, taking the full WalAppendDispatchTimeout (default 30 seconds) to surface as TimeoutException instead of the configured shorter budget.

Mechanically: before each append acquires its partition's admission slot, the writer reads the saturation verdict for that WAL partition rather than the tree-wide roll-up (issue #3348), so one partition at its cap cannot refuse appends routed to its idle siblings; genuinely tree-wide causes still reach every partition's verdict. (A replacement IWalSaturationSignal that cannot answer per partition falls back to the tree-wide verdict and to waiting for Healthy.) On Healthy / Throttled the gate check is a single concurrent-dictionary lookup and the gate never parks: a Healthy caller proceeds directly into the semaphore (no allocation, no extra await), while a Throttled caller first pays the separate local pace, WalThrottledAdmissionPace (25 ms by default). On Saturated the writer parks for at most this budget, or for whatever remains of the enclosing call's WalAdmissionSaturationCallBudget if that is smaller. With WalSaturationAcuteOnly at its default true the parked caller resumes as soon as its partition leaves Saturated; with it false the caller waits for Healthy. If the budget expires with the partition still Saturated, the writer throws LatticeSaturatedException carrying the id of the tree whose write-ahead log refused the append (on an aliased tree, the physical copy's id rather than the id the caller addressed), so the caller can detect the saturation regime via a single is check instead of waiting out the full WalAppendDispatchTimeout. A borderline-recovery race (the wait expires AND the partition recovered between the wait expiring and the re-check firing) is suppressed: the writer re-reads the verdict once after budget expiry and proceeds without refusal when the partition is no longer Saturated.

The budget should be shorter than WalAppendDispatchTimeout (so the saturation refusal wins over the dispatch timeout) and longer than one WalSaturationSampleInterval (so a transient classifier flap does not surface as a refusal). The default (5 seconds) leaves WalAppendDispatchTimeout's 30-second default as a strict outer bound and gives the storage account a realistic recovery window for the canonical 409-Conflict burst (typical recovery 1-3 seconds once offered load drops). Refusals are counted on the orleans.lattice.wal.writer.append.admission_saturation_refusals counter (tagged tree, partition), distinct from the dispatch-timeout counter (admission_timeouts) and the drain-release counter (drain.releases).

Set to TimeSpan.Zero to disable the admission-gate saturation check entirely (the historical pre-admission-gate behaviour). Set to Timeout.InfiniteTimeSpan to wait without limit for the partition to recover. The validator rejects any other negative value.

This option can be changed freely at any time. The new value takes effect on the next admission acquire (which is per-dispatch on the WAL writer hot path).

SetManyFanOutBudget

Wall-clock budget ILattice.SetManyAsync spends awaiting its shard fan-out before refusing the call with LatticeSaturatedException carrying LatticeSaturationSource.SetManyFanOut (default: Timeout.InfiniteTimeSpan, i.e. unbounded - the bound is opt-in).

SetManyAsync splits a batch across the shards its keys route to and awaits every branch, so the call costs the slowest branch rather than the typical one. As the shard count rises the probability that at least one branch is in its slow tail rises with it, so the call tracks the branch p99 and can degrade even as branch-level latency improves. That is the collapse measured in #3348: widening a cluster from four silos to eight improved the branch median 7.7x to 386 ms while the branch p99 degraded 11x to 94 s, and the batch tracked the p99.

Without this budget the wait is unbounded, which means the fan-out queues load it cannot drain instead of shedding it. Bounding it turns the fan-out into a back-pressure seam: once the budget expires with branches still outstanding, the caller is refused in budget time and can back off.

Refusal sheds the caller; it does not roll anything back. SetManyAsync is not atomic across shards, so branches that already committed stay committed, and outstanding branches keep running to completion. The durable outcome is identical to the unbounded wait - only the moment the caller is told changes. This is the same contract the fan-out already had for a faulted branch. Callers that need all-or-nothing semantics across shards should use the atomic-write saga instead. The saga's own batched dispatch also goes through SetManyAsync, so this budget bounds it too: a refusal there surfaces as LatticeSaturatedException with source AtomicWriteSaga and leaves the saga resumable on the caller's retry with the same operation id.

Sizing. Set the budget above the fan-out latency a healthy cluster actually exhibits and below the latency that characterises collapse, so it discriminates rather than fires indiscriminately. Thirty seconds is the recommended starting point, taken from the multi-silo measurements in #3348 on the reference rig: healthy four- and six-silo cohorts observed a per-call p99 of 5.5 s and 11.5 s, while collapsed eight-silo cohorts observed a per-call p50 of 34.7-60.2 s. Thirty seconds sits in the gap - roughly 2.6x above the healthy ceiling and below the collapsed floor - so it is inert on a healthy cluster and engages on a collapsed one. It also matches the WalAppendDispatchTimeout default, so a single stuck branch surfaces as a saturation refusal rather than an opaque dispatch timeout. Re-measure for your own cluster rather than porting that figure blindly: a deployment that sends larger SetManyAsync batches, has slower storage, or runs a much wider shard map has a different healthy ceiling.

The default is unbounded, so this option is opt-in on the 9.x line. Orleans.Lattice has shipped release tags, and a finite default is a behaviour change a conforming caller can be caught by: a batch that legitimately takes longer than the budget - a wide shard map, a large batch, or slow storage - would newly throw where it previously blocked and then succeeded. The repository has shipped a breaking change in a minor only where no conforming deployment could regress, which this does not clear, so the default stays at Timeout.InfiniteTimeSpan and the flip to a finite default is deferred to the next major (#3386). Until then, an unbounded fan-out remains the out-of-the-box behaviour and the collapse in #3348 is only mitigated on deployments that set a finite budget. Set one.

Unlike WalAdmissionSaturationWaitBudget, TimeSpan.Zero is not a disable sentinel and is rejected by the validator: a zero budget would refuse every batch immediately, which is never a useful configuration and is far more likely to be a mistake than an intention. Timeout.InfiniteTimeSpan is the default and awaits every branch however long it takes. The validator rejects every other non-positive value.

This option can be changed freely at any time. The new value takes effect on the next SetManyAsync call that fans out across more than one shard; single-shard batches never consult it, because there is no slowest branch to bound.

SetManyEnvelopeBudget

Wall-clock budget for the whole batched write an ILattice.SetManyAsync call performs - its gate, route, bucket and fan-out stages together, not one of them - before refusing it with LatticeSaturatedException carrying LatticeSaturationSource.SetManyEnvelope (default: Timeout.InfiniteTimeSpan, i.e. unbounded - the bound is opt-in). The clock starts where the orleans.lattice.set_many.duration histogram's does, after the call's argument, admission, authorization and write-interceptor checks, and it is enforced on the fan-out wait: stages that overrun before the fan-out are charged against the budget and refused there, and the post-commit events stage that publishes change notifications falls outside it.

What it bounds that SetManyFanOutBudget cannot. A batched write runs four write stages in sequence: gate (registering the tree's tombstone-compaction reminder and arming its autonomic loops), route, bucket, then fanout. Each carries its own stage timer, and a per-stage budget can only ever observe one of them. #2685 measured a call that breached by summing: a gate of 4,108.96 ms - 12,085x its 0.34 ms healthy baseline, because an uncached arming path ran on every write - plus a fanout of 26,709.17 ms, totalling 30,818 ms against a 30,000 ms Orleans response timeout. Neither stage breached on its own. A fan-out budget sized for the fan-out (30 s is the figure #3348 supports) never fires at 26.7 s, so the caller received an anonymous TimeoutException from Orleans naming no stage at all.

That has a consequence worth stating plainly, because it determines what a useful guard looks like: any assertion made about a single stage against the deadline passes both before and after a regression of this kind, since no single stage is ever the thing that breaches. Only the running total moves.

It composes with SetManyFanOutBudget rather than replacing it. The fan-out waits for the narrower of the two, so a deployment that sets both keeps its per-fan-out ceiling and stops the fan-out being granted a fresh full window by a call that already spent most of the caller's patience upstream. Setting only this one is the simpler configuration.

It applies to single-shard batches, which SetManyFanOutBudget deliberately does not. A single-shard tree has no branch dispersion to bound, which is why the fan-out budget skips it - but it does have an envelope, and before this option a single-shard tree could not have its batch writes bounded by anything. That is the shape the incident was measured on.

The refusal carries a per-stage breakdown of the stages that had finished, next to the share of the envelope that was left for the fan-out it abandoned. The breakdown is taken as the refusal is raised, before the abandoned fan-out's own time is recorded, so for the incident above under a 25-second budget it reads, for example, gate=4109.0ms, route=0.6ms, bucket=0.0ms, fan-out=0.0ms; elapsed=25000.6ms of a 25000ms envelope budget - a non-zero fan-out figure is time spent in earlier attempts that a routing change retried. Absolute magnitude says where time is spent and the ratio against a healthy baseline says what moved, and only the second is diagnostic for a system that was healthy hours earlier - ranking these stages by magnitude finds the fan-out, which had not changed much, while the gate is the stage that degraded by four orders of magnitude.

Sizing. Set it below the response timeout governing the call - the silo's SiloMessagingOptions.ResponseTimeout for a silo-to-silo write, or the client's ClientMessagingOptions.ResponseTimeout for an external one - with enough margin for the refusal to be built and marshalled back while the caller is still listening. Above that deadline it is dead configuration: the caller's own RPC deadline expires first and it sees the generic Orleans timeout this option exists to replace. Against the 30-second Orleans default, 25 seconds leaves a 5-second margin, mirroring DefaultMaxScanPageStallHeadroom.

Refusal sheds the caller; it does not roll anything back. SetManyAsync is not atomic across shards, so branches that already committed stay committed and outstanding branches run to completion. The durable outcome is identical to the unbounded wait - only the moment the caller is told, and what it is told, changes. Callers needing all-or-nothing semantics across shards should use the atomic-write saga instead. Its batched dispatch goes through SetManyAsync as well, so this envelope bounds it too: a refusal there surfaces as LatticeSaturatedException with source AtomicWriteSaga and leaves the saga resumable on the caller's retry with the same operation id.

The default is unbounded, so this option is opt-in, for the same reason as SetManyFanOutBudget: a finite default would change when LatticeSaturatedException first surfaces for a conforming caller on a released package. TimeSpan.Zero is rejected by the validator - it would refuse every batch write immediately - as is every other non-positive value except Timeout.InfiniteTimeSpan.

This option can be changed freely at any time. The new value takes effect on the next SetManyAsync call.

WalAdmissionSaturationCallBudget

Wall-clock budget one top-level call may spend waiting at the WAL admission saturation gate, summed across every append that call makes and every retry layer it passes through (default: Timeout.InfiniteTimeSpan, i.e. unbounded - the bound is opt-in).

WalAdmissionSaturationWaitBudget bounds one wait. It does not bound a call, because the write path holds three nested retry layers - the tree grain's stale-routing retry, the shard-activation retry, and the shard root's leaf-batch dispatch retry - and each re-dispatch opens a fresh per-append budget. A call can therefore accumulate a multiple of the configured budget while every individual wait is correctly bounded, which is why a per-append assertion never caught it. The Layer 3 cohort logs for #3348 record the consequence directly: flush of 4096 failed after 5 retry attempts against LatticeSaturatedException; 10488ms of that was saturation back-off, against a 5-second per-append budget. This option is the issue's own remedy 3 ("cap the total gate wait per top-level call"), and it subsumes remedy 4 ("reconsider whether all three retry layers may wait at the gate") by making the answer yes, but out of one shared allowance.

How the call is identified. Every public write entry point establishes its ambient transaction context, which also stamps a call-start instant into RequestContext if one is not already present. Nested entry points inherit the outermost stamp rather than re-stamping, so the budget measures the whole call and not each layer of it. The instant is stored as DateTimeOffset.UtcNow.UtcTicks rather than a Stopwatch timestamp because RequestContext values cross silos, where monotonic tick origins are not comparable; NTP skew is small against a multi-second budget, and this bounds back-pressure rather than enforcing a correctness invariant, so the wall-clock reading is sufficient.

What happens when it is spent. The gate takes the smaller of the remaining call allowance and WalAdmissionSaturationWaitBudget, so a call never waits past its allowance and a single append never waits past the per-append bound. Once the allowance is exhausted the gate refuses immediately without opening another wait, throwing LatticeSaturatedException with LatticeSaturationSource.WalAdmission and a message naming WalAdmissionSaturationCallBudget, so an operator can tell a per-call refusal from a per-append one.

Writes with no ambient call. Convergence-only and background writes never pass through a public entry point, so they carry no call-start stamp. They keep the per-append bound unchanged rather than being refused instantly against an epoch they never set.

Sizing. Set it to a small multiple of WalAdmissionSaturationWaitBudget - 15 seconds (3x the 5-second per-append default) is the recommended starting point. Below 1x it would pre-empt the per-append budget and make that option unreachable; far above 3x it stops discriminating, because the multiplication observed in #3348 was roughly 2x the per-append budget in the worst logged case. Keep it below SetManyFanOutBudget, so a batch whose branches are stuck at the gate surfaces as a WAL-admission refusal naming the real seam rather than as a fan-out expiry naming the symptom.

TimeSpan.Zero is accepted and means "never wait at the gate within a call" - a coherent fail-fast posture, unlike SetManyFanOutBudget where zero would refuse every batch outright. Timeout.InfiniteTimeSpan is the default and leaves only the per-append bound in force. The validator rejects every other negative value.

The default is unbounded, so this option is opt-in on the 9.x line, for the same reason as SetManyFanOutBudget: a finite default would change when LatticeSaturatedException first surfaces for a conforming caller on a released package, which is a breaking behavioural change. The flip to a finite default is deferred to the next major (#3390).

This option can be changed freely at any time. The new value takes effect at the next admission gate check.

WalThrottledAdmissionPace

Per-append pacing delay the WAL writer applies on the local admission path while the saturation verdict for the append's WAL partition is WalSaturationState.Throttled (default: 25 milliseconds; set to TimeSpan.Zero to disable local pacing). This is what gives the drain-lag (and any other Throttled-mapped) back-pressure teeth on the single-silo local-write path, where there is no remote replication sender to drip-feed and the Saturated-only WalAdmissionSaturationWaitBudget gate never engages. Before each dispatch admits into the per-partition admission semaphore the writer reads that partition's verdict once; on Throttled it awaits a single bounded Task.Delay of this duration, pacing the local producer so the materialiser drain can catch up. A partition's verdict carries the tree-wide Throttled causes (drain lag, pin latency, and the recovery-window hold) but only its own admission depth, so one busy partition does not pace appends routed to its idle siblings.

It is a pure back-off: it never throws, and it never escalates to LatticeSaturatedException - a Throttled tree slows callers, it does not fault them. The pacing is a no-op when no saturation signal is registered (single-node / unit-test writers), when the signal reports Healthy (a single concurrent-dictionary lookup, no await), and on Saturated (the separate admission gate already governs the dispatch, so the pace is skipped to avoid double-charging the caller). Caller-supplied cancellation surfaces as OperationCanceledException; a writer drain request short-circuits the pace silently so shutdown is never slowed.

Sizing guidance: the delay is added to every append admitted while its partition is Throttled, and concurrent appends wait out their delays in parallel, so it is a per-append latency charge rather than a partition-wide rate cap: one sequential producer is held to roughly 1 / WalThrottledAdmissionPace appends per second (about 40 per second at the default 25 ms), while concurrent producers are each slowed by the same delay without sharing a budget. Raise it to slow producers harder when the materialiser drain is the bottleneck; lower it (or set TimeSpan.Zero) when you would rather rely solely on the replication-side flow control. The validator rejects negative values.

This option can be changed freely at any time. The new value takes effect on the next admission acquire (per-dispatch on the WAL writer hot path).

WalMaxPendingBatches

Maximum number of in-flight storage-provider flushes the partition grain admits concurrently (default: 16, the measured Azure Tables Standard sweet spot at 4,000 keys/s offered load on Standard_D4as_v5). Raising this value increases pipeline depth against the storage provider - the next caller can enqueue a new flush as soon as the in-flight count drops below the cap, rather than waiting for the head of the in-flight chain to settle. The previous default was 8; raising it to 16 produced a +57% increase in steady-state silo throughput at the 4k:5 rung with no reliability regression.

Set to 1 to restore the historical single-in-flight shape (strict ordering against the provider; no pipeline depth). The options validator rejects a value below 1. Most workloads on durable backing stores benefit from the default; the strict-ordering shape is useful only when targeting a provider whose ordering guarantees are weaker than per-request linearisability.

The flush-cap-reached cutover backs off the calling task by awaiting the in-flight head, so the cap also acts as the natural back-pressure ceiling against caller fan-in. Raising the cap above what the storage provider can usefully serve in parallel degrades latency without improving throughput - more concurrent flushes compete for the same provider budget and grow each flush's slow-tail wait. At the canonical WalPartitions = 8 the combined fan-out is 8 * 16 = 128 concurrent flushes against the provider, which is at the edge of a single Azure Tables Standard storage account's sustained throughput budget; see WAL Tuning for the envelope above which the storage account becomes the binding constraint and the recovery path (WalPartitions fan-out across accounts, not a higher per-partition cap).

This option can be changed freely at any time. The new value takes effect on the next batch boundary.

WalAppendCoalescingInFlightThreshold

In-flight flush depth at or above which the partition grain stops kicking a flush for the final entry of an arriving append batch, letting that batch instead accumulate into the next flush window (default: 4; 0 disables coalescing and restores the historical unconditional kick).

The problem it solves is that a batched write arrives at any one partition already divided twice. ILattice.SetManyAsync fans a batch out over the tree's shards, and each shard's slice then fans out over the tree's WAL partitions, which every shard shares, so a 4,096-key batch over 64 shards and 16 partitions reaches each partition as 4096 / 64 / 16 = 4 entries per arrival. Before this option existed, the flush-kick predicate treated the final entry of every arriving batch as a reason to flush immediately, so each of those four-entry slices paid a full storage round trip. Append cost therefore never amortised with load: offering more work bought more concurrent small appends rather than fuller ones, and measured per-append entry counts fell as offered load rose (see #3396).

Suppressing that kick cannot strand a batch, which is why no timer is involved. Suppression requires the in-flight count to be at or above the threshold, and the threshold is at least 1 whenever coalescing is enabled, so a flush is necessarily outstanding at the moment of suppression; the flush-completion path already re-kicks whenever pending segments remain and the in-flight count is below WalMaxPendingBatches. The accumulating batch therefore has a guaranteed later drain. A threshold of 0 would be the one unsafe value - suppressing with nothing in flight - and is not expressible, because 0 means disabled.

The default is deliberately well below WalMaxPendingBatches so it engages before the in-flight cap does; a threshold at or above that cap could never fire, since admission already requires in_flight < WalMaxPendingBatches. A threshold of 1 is maximum coalescing (every arrival behind an outstanding flush accumulates) but overrides the measured pipeline-depth tuning for partitions whose batches are already well filled.

The option is self-disabling below its threshold: until that many flushes are concurrently in flight, the predicate is identical to the historical one, so a quiet or moderately loaded partition behaves exactly as before. It changes no method signature, wire format, ordering, durability, or offset assignment - offsets are still assigned under the state gate and each flush window is still strictly above every in-flight window - only how many entries share a window. Accumulation stays bounded by WalMaxBatchEntries and WalMaxBatchBytes, which cut a flush over regardless of the threshold.

This option can be changed freely at any time. The new value takes effect on the next batch boundary.

WalBatchedSingleEntryAppends

When true (the default), every single-entry WAL append - a point append (leaf SetAsync / DeleteAsync and their conditional forms, CRDT merge apply, pending-transaction staging, inline saga terminals) as well as a bulk append carrying exactly one entry - is dispatched through the interleaving batched grain method instead of the exclusive-turn per-entry overload. Point appends were routed in issue #812: before that they always took the exclusive turn, so each WAL partition admitted one point write per provider round trip and a set-point workload was capped at roughly WalPartitions divided by that round trip regardless of shard count. Under a wide fan-out whose per-leaf slices are one entry each - the dominant shape for uniformly distributed keys - the exclusive turn serialised every append on a partition for its whole provider round trip, collapsing concurrency to one and keeping WalAppendCoalescingInFlightThreshold from ever being reached. Ordering, durability, and offset density are unchanged, because the shard's internal state gate, not turn exclusivity, serialises offset assignment and the pending list. The orleans.lattice.wal.append.turn_wait and orleans.lattice.wal.append.queue_depth histograms are recorded only by the per-entry overload, so with this option on they no longer observe single-entry appends. Set to false to restore the per-entry overload.

WalMaxRetainedBytes

Optional advisory ceiling on a tree's WAL bytes (default: null, disabled). Despite the name it bounds on-disk occupancy, not live payload: when set, each ILatticeWalGc.RunOnceAsync pass samples the WAL's physical size - every byte it occupies, including dead bytes trimmed but not yet reclaimed by compaction, falling back per partition to the retained figure when a provider cannot report physical size - before and after its safe trim; if the pre-trim total exceeds the ceiling the policy schedules a byte-pressure trim (surfaced as the orleans.lattice.storage.policy.trim_triggered counter and LatticeWalGcReport.BytePressureTriggered), trimming toward WalMaxRetainedBytes * WalBytePressureReclaimTarget. The policy is advisory only: the GC never trims past the safe frontier (the slowest consumer's cursor, or the WalRetention window where set, overruled where available by the durable materialiser offset floor, intersected with the causal-stable frontier and held below the blocked floor) to honour it, so a tree can remain over the ceiling - LatticeWalGcReport.BytePressureOverThreshold and the orleans.lattice.storage.policy.over_threshold gauge report that occupancy, not its cause. A lagging consumer pinning the frontier is one cause, but occupancy counts dead bytes awaiting compaction, so a tree that reclaimed everything it was entitled to sets the flag with nothing pinning it; read LatticeWalGcReport.CursorFloorState and LatticeWalGcReport.RetainedBacklog to tell those apart. null disables the policy. See WAL and Tree Storage.

This option can be changed freely at any time. The new value takes effect on the next GC tick.

Size it above LatticeOptions.WalMaxRetainedBytesWorkingSetMultiple (2) times the largest tree's logical working set. A log-structured provider reclaims dead bytes only by rewriting a segment once they reach its compaction threshold (half the file at the file provider's default ratio), so a healthy tree's steady-state occupancy is about twice its live set. A ceiling below that multiple is unsatisfiable by construction, and with the default WalBytePressureReclaimTarget of 0.8 its disarm point sits below the tree's natural floor, so the advisory alarm arms and never clears. Because the working set grows, nothing validates the rule at startup; every pass evaluates it against the working set it just measured and reports a breach on orleans.lattice.wal.gc.ceiling_unsatisfiable, which is zero-primed per tree.

A per-tree override can be set at runtime, with no silo restart (issue #3333). The silo-wide value above is the default for every tree; an individual tree can pin its own ceiling through the tree-admin facade (lattice_treeadmin_tree_set_config with applyWalMaxRetainedBytes: true, or ILatticeTreeAdmin.SetTreeConfigAsync with TreeConfigurationUpdate.ApplyWalMaxRetainedBytes). The override is persisted on the tree's registry entry and re-read on every GC pass, so it takes effect on that tree's next pass rather than at the next silo start; a null value clears the override and restores the silo-wide default. This matters because the silo-wide value is usually sourced from environment or file configuration that a running host cannot change, so before this existed, correcting a ceiling that had become too small for a grown tree cost a restart - and a restart is exactly what an operator wants to avoid on a host whose WAL is already under pressure. The ceiling is advisory in both forms: an override can never cause the GC to trim past the safe frontier, so lowering one cannot lose data, it only moves the point at which byte-pressure trimming and the cadence floor engage.

This default was re-examined during the bounded-cold-start work and deliberately left disabled. It is a capacity quota rather than a retention mechanism: a correct value is a fraction of the volume the WAL lives on, which the library cannot know, and any value shipped as a default would be wrong for most deployments in one direction or the other. Enabling it adds no provider call while the default-on durability hold (WalDurabilityHoldCeilingBytes) is in force, because every collection pass already samples each WAL partition's byte size, before and after its trim, to bound that hold; only where the hold is disabled does enabling it add that per-partition probe to every pass.

Leaving it off does not blind an operator. orleans.lattice.wal.gc.passes is emitted unconditionally with a reclaimed | blocked | no_partitions | no_consumer | idle | over_ceiling | stranded | unclassified | failed outcome tag, and reclaimed volume is visible through orleans.lattice.wal.entries_trimmed; every pass still samples the WAL's occupancy for the default-on durability hold and records the post-trim sample on orleans.lattice.wal.gc.backlog_bytes wherever the provider accounts bytes, and a pass with no sample increments orleans.lattice.wal.gc.backlog_bytes_unavailable instead (reason="policy_disabled" while this option is unset, reason="provider_unsupported" when it is set), so "not measured" is reported rather than inferred from a missing series. The stranded arm in particular is reachable with this option off (issue #3213): it is decided by the trim scan's stop reason rather than by a byte sample, so a tree that reclaims nothing while its scan keeps stopping on WAL it must retain is distinguishable from a quiet one even where nothing accounts bytes. Set WalMaxRetainedBytes when the WAL volume has a hard size budget and the WAL provider accounts bytes; where the durability hold is disabled, it is also what makes the backlog histogram emit.

Setting it also changes the collection cadence, not only the reporting (issue #3119). A tree that is over the ceiling and reclaims nothing holds the adaptive interval at WalGcMinInterval instead of relaxing toward WalGcInterval, and reports the over_ceiling outcome rather than idle. This is deliberate and is what makes the ceiling enforceable at all: the trim itself can never cross the safe frontier, so when a lagging consumer pins that frontier, pass frequency is the only lever the policy has left, and the previous rule removed it at exactly the moment it was needed. Budget for it as a cost: the per-partition byte probe noted above is then paid at the floor rate for as long as the breach lasts, which is the same rate a reclaiming tree already sustains. It clears on its own once the tree drops back under the ceiling, so a ceiling set below what the tree can ever reach will hold that tree at the floor indefinitely; size the ceiling against a footprint the tree can actually return to. Leaving it off no longer forfeits the cadence correction entirely (issue #3213). A tree that reclaims nothing while retaining WAL is now held below the fault-retry ceiling rather than relaxing all the way to WalGcInterval, on every deployment and with no option set. Setting WalMaxRetainedBytes still buys the stronger guarantee - the cadence floor rather than a capped backoff - and it buys it against a budget the operator chose rather than against the mere presence of retained bytes.

WalDurabilityHoldCeilingBytes

Byte ceiling up to which the WAL garbage collector holds back a tree whose durable materialiser offset floor is absent - that is, when it cannot establish that any leaf has durably applied anything (default: 256 MiB, so the hold is on; null, 0, or any other non-positive value disables it). The ceiling is compared with the tree's WAL occupancy, sampled as WalMaxRetainedBytes samples it (physical bytes where the provider reports them, dead bytes included): below it such a partition is not trimmed; at or past it the collector trims anyway, because an unbounded WAL is the worse outage, and records each forced trim on orleans.lattice.wal.gc.durability_hold_forced. The hold engages only where the floor is absent, never where a floor exists but has stopped advancing, and only while every consumer cursor that admits trims is a leaf materialiser cursor the durable floor does not cover (or the cursor registry cannot be read), so a tree with a non-materialiser consumer such as a replication shipper that has advanced its cursor is never held, and a deployment with a materialiser wired enters it only transiently, while its floor is absent during a rolling upgrade or leaf churn, until the leaves re-pin (reason="pin_regressed" on orleans.lattice.wal.gc.durability_hold_engaged); and where the provider reports no byte size at all the hold declines rather than retaining without a bound. It is deliberately separate from the advisory WalMaxRetainedBytes, which never blocks a trim. See Metrics for durability_hold_forced and durability_hold_engaged.

WalPartitions

Per-tree, pinned at first registration. A tree pins the value in force for it when it is first registered in the tree registry - on its first use, through ILatticeTreeAdmin.CreateTreeAsync, or when an installed app registers it, in which case a manifest-declared walPartitions is pinned instead - so a silo-wide default change is non-breaking for already-registered trees - they continue to fan across whatever partition count they were created with. New trees pick up the current default unless an operator override is configured.

Activation-time replay is partition-aware. The leaf grain's activation-time materialiser iterates [0, WalPartitions) and runs an independent fall-off-log classification, slice read, and projection-checkpoint advance per partition. Per-partition checkpoints persist in the leaf's durable state as a nullable per-partition offset array (an additive serialization slot), with partition 0 also mirrored into the legacy scalar checkpoint slot so a downgrade to a host that has never observed multi-partition state still reads a valid single-partition shape. Per-partition cursor consumer ids take the form _lattice_materialiser_{treeId}_{leafGrainId}_{partition} so the WAL GC trims each partition independently against its own slowest consumer; on WalPartitions = 1 the legacy unsuffixed shape _lattice_materialiser_{treeId}_{leafGrainId} is preserved for wire compatibility with hosts that have never enabled multi-partition replay.

Must be >= 1. Values below 1 fail options validation.

WalMaterialiserPinShards

Number of durable-pin grain activations the per-tree leaf-materialiser checkpoint floor is spread across (default: 8). Each active leaf reports its per-WAL-partition projection frontier as a durable pin so the WAL garbage collector never trims past an entry a leaf has not yet projected. Historically every leaf in a tree funnelled its pin into a single per-tree grain, so a leaf-birth or split storm serialized O(leaves x partitions) durable writes through one activation and that activation became the bottleneck that wedged the drain path. Sharding spreads the load: each consumerId deterministically maps to one shard so the monotonic-max merge stays correct, and the WAL GC fans its read across every shard (plus the legacy single-key shape) so no trim floor is lost.

Set to 1 to restore the historical single-activation shape. Changing this value is a durable-store migration, and the two directions are not symmetric. The WAL garbage collector reads every shard under the current count (under both the current and the earlier shard-key separator) plus the legacy unsharded key, so raising the count keeps every existing pin readable. Lowering it stops the collector reading pins held by shards at or above the new count - and 1 reads only the legacy key - until each affected consumer re-pins under the new layout on its next report. A deliberate rollout (drain, change, redeploy) is therefore recommended rather than flipping it on a hot cluster. The value is read from the default (unnamed) options, so per-tree overrides do not apply, and it is not validated at startup: a value below 1 is treated as 1.