Table of Contents

Options Reference: WalMaterialiserPinBuckets to WalRetention

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at options-reference-5.md, and llms.txt lists every page.

Part of Options Reference, in Configuration.

WalMaterialiserPinBuckets

The minimum number of durable state slots a single pin shard's blob is split across (default: 1, which is the historical single-slot layout and is byte-for-byte what every pre-bucketing build wrote while the shard is small). Orleans persists grain state as one whole blob, so a shard holding N consumer pins rewrites all N of them to record one leaf's advance. On a large tree that is a multi-megabyte write per flush, which is what made the pin store the bottleneck in issue #2012. Bucketing splits the persistence of one shard across several slots and rewrites only the slots whose contents changed, turning an O(consumers on the shard) write into an O(consumers in the bucket) one.

Buckets are the write dimension; WalMaterialiserPinShards is the read dimension. They are orthogonal, and that is the point of having both. Raising the shard count also shrinks each blob, but it widens the WAL garbage collector's per-pass grain fan-in by the same factor, because the collector must read every shard to compute the trim floor. Raising the bucket count shrinks the blob without touching that fan-in at all: one activation still answers for the whole shard and unions its buckets in memory, so GetPinsAsync still costs one grain call per shard. Reach for buckets when pin writes are slow, and for shards when pin calls are queueing.

The configured value is a floor, not the layout. The shard estimates its serialised size as pins change, and when a slot would exceed a 64 KiB byte budget it widens its layout by powers of two (up to 1,024 slots) until each slot is at most half that budget; the width it settled on is recorded in bucket zero. This is what keeps the default safe on a large tree: Orleans' Azure Table grain storage rejects any blob over 983,040 bytes, and before issue #3576 a shard at the default of 1 grew linearly with its consumer count (about 1 MB per 128k preseeded keys) until every flush and birth seed was rejected with Data too large. A provider rejection that reports an oversize payload also doubles the size estimate, so an estimate that under-counts drives a split rather than an identical retry. A failing write is retried after a backoff that doubles from 1 s to 60 s rather than on every flush tick; after three consecutive failures a birth seed or removal fails fast without touching the store (the pin is already merged in memory, which is what the WAL garbage collector reads, and it is made durable by the first successful flush). One coalesced flush writes at most eight slots, so a wide shard drains across several ticks instead of monopolising the non-reentrant pin grain.

Changing this value changes the durable layout but needs no migration step, and both directions are safe. Raising it: existing pins stay in the legacy slot, which every activation keeps reading, and each consumer's pin moves to its bucket the next time that consumer reports. Lowering it: an activation reads the wider layout it finds recorded in bucket zero, merges it, and consolidates into the narrower one only when the smaller layout still leaves each slot within half the byte budget; otherwise it stays at the recorded width, so no pin is stranded and no slot is pushed over the provider limit. An automatic widening follows the same crash-safe order as a configured one: the new slots are written before bucket zero records the new width, and the old slots are rewritten last, so an activation that dies part way through reads either the old layout or the new one in full. The direct-store write a leaf falls back to during silo teardown also routes by the width recorded in bucket zero. After a move out of the legacy slot, every activation keeps reading the legacy slot until every bucket that receives its pins has been written; the legacy slot is then emptied, so a consumer removed after the split is not resurrected from it at its pre-split frontier by the next activation. An activation that finds a non-empty legacy slot under a bucketed layout, including one a pre-#3576 build left behind at a bucket count of 2 or more, treats it as an unfinished move: it keeps every one of those pins, rewrites the buckets they route to, and then empties the legacy slot. Every read at activation is fail-closed: if bucket zero or any slot of the recorded layout cannot be read, activation fails and the WAL garbage collector reports the pin census as unavailable and retries, instead of computing a floor from a partial map, which could only raise it. A pass that cannot read the pin or offset census of any shard fails closed: it trims nothing, TTL included, and retries on the next pass.

Rolling upgrades and rollbacks. A build that predates #3576 does not read the width recorded in bucket zero, and once the legacy slot has been retired it holds no pins. If a pin shard reactivates on such a silo at a bucket count of 1 while a rollout is in progress, that silo sees none of the split shard's pins. During the rollout window, either drain the older silos before the new build starts serving, or run the whole cluster (old and new silos) at a bucket count of at least 2, ideally at or above the widest recorded width, so an older silo reads the buckets. The same applies to a rollback to an older build after a shard has split itself. As with the shard count, a deliberate rollout (drain, change, redeploy) is still preferable to flipping it on a hot cluster. It is read from the default (unnamed) options, so per-tree overrides do not apply. Must be >= 1; the validator rejects values below 1.

Pair with orleans.lattice.materialiser.pin.durable_write_latency and orleans.lattice.materialiser.pin.reports_shed when sizing: a sustained non-zero shed rate means the pin store is the bottleneck and the bucket count is the knob for it.

WalMaterialiserPinFlushIntervalMs

Debounce window, in milliseconds, over which a shard's durable pin writes are coalesced behind a grain-timer flush (default: 250 ms). Within the window the shard advances its in-memory monotonic-max frontier on every advancing report but persists at most one durable WriteStateAsync, collapsing a report burst into one durable write per shard per window. The coalescing only ever retains more WAL than an immediate write would (the persisted floor lags the in-memory floor by at most one window), so it is always GC-safe.

Set to 0 to disable coalescing so every advancing report persists synchronously, matching the historical shape. The value is not validated at startup; a negative value behaves like 0. It is read from the default (unnamed) options, so per-tree overrides do not apply.

This option can be changed freely at any time. The new value takes effect on the next pin report.

WalMaterialiserPinShedCeiling

Maximum time a durable leaf-materialiser pin shard may shed coalescible reports continuously before one is forced through (default: null, disarmed, which preserves the historical shedding exactly). Under sustained leaf-activation churn the non-sheddable pin writes can hold a shed window open indefinitely and starve the path that restamps materialiser coverage, so the durable pin stops advancing and the retained WAL grows. With a ceiling set, once a shard has shed for that long the next report is issued regardless and orleans.lattice.materialiser.pin.shed_forced records it, bounding pin staleness at the cost of at most one enqueued write per ceiling period per shard. Arming it cannot lose data. It is read from the default (unnamed) options, so per-tree overrides do not apply. The validator requires a positive value, or null.

WalMaterialiserMaxConcurrentReplays

Per-silo ceiling on the number of leaf grains that may run their activation-time WAL replay concurrently (default: 0, which resolves at runtime to the lesser of Environment.ProcessorCount and the container's enforced CPU grant). A mass reactivation (for example after a docker restart or a silo rejoin) can otherwise stampede the scheduler as every reactivating leaf replays its WAL backlog at once; the ceiling makes the surplus queue on a process-wide gate and drain in waves instead. A no-op activation (a leaf with no tree binding) consumes no permit.

Why the grant and not the processor count alone. The ceiling is CPU-derived, so it is only meaningful against the CPU the process can actually obtain - a bound expressed in CPUs that is read off a figure the kernel does not enforce bounds nothing. Environment.ProcessorCount reflects the cgroup quota only while DOTNET_PROCESSOR_COUNT is unset - that variable overrides it, and overriding it is a legitimate thing to do for the thread pool's sake. A container granted 6 CPUs whose environment carries DOTNET_PROCESSOR_COUNT=16 therefore used to size this gate at 16, a 2.67x oversubscription of the CPU-derived ceiling, with nothing in the process able to tell (issue #2816). The default now reads /sys/fs/cgroup/cpu.max (and the cgroup v1 pair) directly and takes the minimum of the two figures, so a quota the kernel actually enforces can lower the ceiling but never raise it. An operator who lowers DOTNET_PROCESSOR_COUNT is still obeyed exactly, and every other subsystem sized from the processor count is untouched. An unreadable or unlimited quota - a non-Linux host, a container with no CPU limit - is treated as unknown, not as zero, and leaves the processor count as the ceiling.

The resolved ceiling, the configured option, the processor count, the grant, and the resolved managed heap ceiling are all written to the log once per process at silo start, so the effective figure can be read off the log rather than inferred from the host's vCPU count. The heap ceiling is reported there because the first four figures are all CPU quantities while the resource a concurrent replay exhausts first is the managed heap (issue #2784); a record carrying only the CPU side reads as a complete account of the sizing and is not one. It is reported as unknown when the runtime declines to supply it, which is a finding in its own right - heap occupancy then yields no verdict and the adaptive layer described below cannot engage on it.

The ceiling is sized once, but concurrency adapts beneath it (issue #2862). The sizing above is CPU-derived and correct as far as it goes, but on a mass reactivation the binding constraint is usually the heap, not the CPU: a whole-window replay buffers the window, so N concurrent replays hold N such buffers against one fixed managed ceiling. The gate therefore carries a memory dimension on top of the CPU-derived figure. A replay that finishes while managed heap occupancy is at or above 75% of the GC hard limit does not return its permit, and a replay that fails for memory pressure does not return its permit whatever the occupancy; a withheld permit is returned only when a replay completes cleanly and occupancy has receded below 60%. Those two figures are a hysteresis band, not one edge, so the gate settles at the concurrency the heap can afford rather than oscillating across a single threshold. The occupancy denominator is the GC hard limit (GC.GetGCMemoryInfo().TotalAvailableMemoryBytes, floored by the cgroup memory limit when one is readable) and not the container memory grant: under a container limit .NET applies its default GCHeapHardLimitPercent to the grant, so a 12 GiB grant yields a 9 GiB managed ceiling and a threshold written against the grant would sit 33% above the limit that actually throws. Backpressure never withholds the last permit, so the effective ceiling has a floor of one and the mechanism cannot latch, and it works by declining to return a permit already taken rather than by resizing, so it can never raise concurrency above the configured ceiling. Both arms of orleans.lattice.wal.replay.permit_adaptations are primed at zero when the gate is sized, so a flat zero is a measured zero rather than an absent series; the withheld arm additionally carries a trigger tag (occupancy or fault) naming which of the two mechanisms above withheld the permit, and each trigger value is primed separately (issue #2883). The counters are monotonic totals, so neither of them - nor their difference across two scrapes - can express what a process restart destroyed: the adaptation is held entirely in process memory, and a silo that dies at a withheld level of five starts again at zero with nothing in either series marking the discontinuity. orleans.lattice.wal.replay.permits_withheld is an observable gauge of that live level, published so the boundary is visible where a delta over the counters cannot show it (issue #2784). It is not a substitute for the counters: being sampled, a withhold and its restore that both fall between two scrapes leave it unchanged while the counters record both.

Set to a positive value to pin the ceiling explicitly; an explicit value supersedes both figures. Must be >= 0; the validator rejects negative values. The gate is sized once per silo process, from the resolved options of whichever leaf first takes a replay permit, and is never resized: a new value takes effect on the next silo start. Set it globally, because a per-tree override sizes the whole silo's gate only when a leaf of that tree happens to be first.

The CPU-derived default is an initial guess, and on a store-bound deployment it is the wrong one (issue #3921). Both inputs to the default are CPU quantities, which is correct only while a replay is CPU-bound. When the replay waits on the grain store instead - a store that serialises its writes, such as a single SQLite file behind ADO.NET - more permits do not finish more replays per second: each concurrent replay queues behind the same writes, so raising the ceiling only lengthens every hold. A long hold is what trips the admission gate's no_progress arm (see WalReplayPermitMaxQueueWait), so the gate then refuses more work, not less. Measured on one such deployment, pinning this option from the derived 6 to 16 raised saturation refusals by nearly half, raised timed-out index loads six-fold, took a cold open that had completed in 12 and in 26 minutes to one that did not complete in 45, and left five of six granted cores idle throughout. On a store-bound deployment tune this option to the store's write parallelism. Two instruments tell the regimes apart: orleans.lattice.wal.replay.permit_hold is how long each replay held its permit, and the rate of orleans.lattice.wal.replay.permits_served is the gate's service rate. If raising the ceiling raises the service rate, replays are CPU-bound and the ceiling was the constraint; if the service rate stays flat while hold time grows, the store is the constraint and the ceiling should come down. Sizing the ceiling automatically from measured permit service time is tracked as issue #3933. See Metrics.

Starvation drives take a bounded, non-queueing share of the gate (issues #3480, #3575). The WAL GC sweep's starved-leaf drives and a leaf's own coverage-lag timer drives replay under this same gate, but they never queue for a permit, and together they hold at most half the replay permits in circulation - the resolved ceiling less any the heap backpressure described above is withholding - rounded down, but never below one, leaving the rest for ordinary activations. A drive that cannot take a free permit immediately within that share is refused at once, before it replays anything; the refusal is returned as a result rather than thrown, and is counted on orleans.lattice.saturation.refusals under the replay_permit_admission source. Within the share a coverage-lag timer drive never takes the last free slot, which is kept for the sweep - the only drive that clears a pin holding a tree's WAL cursor floor - and on a single-slot share the timer yields that slot while a refused sweep drive is outstanding. The sweep keeps no more of its touches in flight than the share it can use, so its own drives do not refuse each other, and re-drives a refused touch once a sibling touch of the same pass frees a slot (issue #3761); it records each refusal as the admission_refused outcome on orleans.lattice.wal.gc.blocked_leaf_reactivations, outside attempted, and retries a consumer still refused at the end of the pass after a jittered delay that doubles with each consecutive refusal; the timer counts its refusals, and the drive opportunities its backoff then skips, as the recheck_drive_refused and recheck_drive_deferred reasons on orleans.lattice.leaf.snapshot.driver.declines. See Metrics.

This ceiling is one of two factors (issue #2867). Peak replay memory is the product of how many replays run concurrently and how much each one buffers, and the paragraph above is careful about the first while being silent about the second - which is how the ceiling came to be treated as the control. It is not. The per-replay factor is the slice width the replay reads the commit log in, and it is settable in its own right as WalReplaySliceBudget (issue #2898). Until that option existed the width was a private constant with no option, no environment variable and no overlay entry behind it, so lowering peak memory by lowering the width was not something configuration could express. Both terms of the product are now settable.

That matters because both ends of the one dial you do have are failure modes, so there is not always a setting that works. Set the ceiling too high and the concurrent whole-window buffers sum past the managed heap hard limit, an OutOfMemoryException cancels the activation, and the replay banks nothing. Set it too low and cold leaves queue for a permit past the request timeout, the runtime cancels the activation, and the replay again banks nothing. Either way no snapshot is banked, the durable materialiser pin stays unusable, and WAL garbage collection reports blocked - so the next replay window is larger than the one that just failed.

The configured width is a starting point, not a floor: a replay whose slice read fails for memory pressure narrows its own width to a quarter and retries the same range (issue #2742), a reactive, per-replay, per-activation adaptation that widens back towards the configured width on success and starts afresh at it on the next activation. Setting the option lowers the ceiling that adaptation works down from; it does not disable it. Whether it is engaging at all is visible in orleans.lattice.wal.replay.slice_narrowings, tagged by tree and partition and primed at zero, which is the instrument to read before concluding that a memory-pressure failure was caused by the slice width - a flat zero alongside climbing activation failures means the failing allocation was somewhere else and the width is not the lever. See Metrics.

The narrowing reached only the activation-time replay when it was introduced. Two further replay sites - the snapshot-cursor rebuild and the frozen-baseline tail fold - kept a constant of the same name and the same value and passed it straight through, so for a period the recovery path was the one without the resilience, and every check that compared the two values passed while the behaviour had diverged. All three sites now read through one shared reader that owns the width, the retry, the widening and the counter's priming together (issue #2899), and all three honour this option.

WalReplayPermitQueueDepthPerPermit

How many activations may queue for a WAL replay permit, per permit the replay gate was sized to (default: 4; 0 admits an unbounded queue, the historical shape). The admitted-waiter bound is this figure multiplied by the resolved ceiling (see WalMaterialiserMaxConcurrentReplays), so it scales with the deployment's own CPU grant. An activation refused admission gets a fast, attributable LatticeSaturatedException to retry after a backoff instead of a silent wait that would outlive its request deadline. A refusal is expected backpressure rather than a fault, so it is not logged per occurrence and carries no stack trace: each silo writes at most one warning a minute stating how many replays it refused since the previous one, and the exact count is on orleans.lattice.saturation.refusals under the replay_permit_admission source (issue #3906). Admission is refused only when this depth bound and WalReplayPermitMaxQueueWait both say the queue is unhealthy. The validator rejects a negative value.

WalReplayPermitMaxQueueWait

The longest smoothed WAL replay-permit queue wait treated as healthy (default: 5 seconds; TimeSpan.Zero disables this half of the admission decision and restores the pure depth bound). It is the demand-side half of the admission rule: depth alone says the queue is long, while this says it is not draining, which is what separates a wide but healthy fan-out from a self-sustaining backlog. Both halves must hold before an activation is refused. The validator rejects a negative value.

The not-draining half has two arms, and a refusal names the one that fired, in its message and as the arm tag on orleans.lattice.saturation.refusals (issue #3921). wait_exceeded: the smoothed wait of recently terminated permit waits is at or above this bound, so waits are completing but slowly and too much is queued; WalReplayPermitQueueDepthPerPermit is a relevant lever. no_progress: no queued activation has acquired a permit for at least this bound, so the permits already issued are not coming back. That is a hold-time condition, not a queue-depth one: the replays holding the permits are slow, and the queue-depth option does nothing for it. Its refusal message reports how long it has been since a permit was released to the queue and points at replay and store throughput; if the replays are bound on the grain store, lowering WalMaterialiserMaxConcurrentReplays toward the store's write parallelism is the remedy and raising it makes the refusals more frequent. A permit withheld by memory backpressure also does not come back, which orleans.lattice.wal.replay.permits_withheld shows.

WalReplayMaxRecordsPerTurn

Number of WAL records a single activation-time replay projects before yielding the Orleans turn cooperatively (await Task.Yield()), so a long replay does not monopolise the activation's turn and starve other grain calls on the same activation (default: 256). This is distinct from the cross-RPC slice width, WalReplaySliceBudget, which bounds how many entries a single slice read returns; this option bounds the synchronous run length within a single replay turn. They default to the same number and mean different things.

Set to 0 to disable the cooperative yield so replay runs to completion without voluntarily yielding (the historical shape). Must be >= 0; the validator rejects negative values.

WalReplaySliceBudget

Number of WAL entries a single replay requests per commit-log slice read, and the width it widens back towards after a memory-pressure narrowing (default: 256).

This is the second of the two factors that set peak replay memory (issue #2898). Peak draw is the product of how many replays run at once and how much each one buffers; the first factor is WalMaterialiserMaxConcurrentReplays and this is the second. Lowering it trades round trips for a smaller resident slice on a host that cannot afford the default width.

Do not confuse it with WalReplayMaxRecordsPerTurn, which defaults to the same number. That option bounds how many records a replay applies within one scheduler turn before yielding cooperatively, and so governs silo responsiveness. This one bounds how many entries a single cross-RPC slice read returns, and so governs allocation. Changing one does not change the other.

The value is the starting width, not a floor. A read refused for memory pressure is retried at a quarter of the current width, floored at a single entry, and widens back towards this value on success - so configuring it lowers the ceiling the replay works down from rather than disabling the adaptation. Whether the adaptation is engaging is visible in orleans.lattice.wal.replay.slice_narrowings; see Metrics.

All three replay sites honour it: the activation-time replay, the snapshot-cursor rebuild, and the frozen-baseline tail fold.

Must be >= 1; the validator rejects zero and negative values, because a single entry is the narrowest legal read and a width of zero would request nothing and never advance.

WalGcInterval

Upper bound on the cadence at which the per-silo core WAL garbage-collection scheduler runs a ILatticeWalGc.RunOnceAsync pass over every registered tree (default: 1 hour, enabled). This is the quiet-path ceiling of an adaptive band whose floor is WalGcMinInterval: a tree that is reclaiming entries is collected far more often, and relaxes back to this interval once it has nothing left to reclaim. The core library ships the WAL garbage collector, but historically only drove it for replicated trees (via the replication package's per-tree maintenance grain). That left two retention gaps: a durable-WAL host that runs without the replication package never trimmed its WAL at all, and every non-replicated tree in a replicated host was never collected - both grew without bound, and WalRetention was inert for them. The built-in scheduler closes the gap by collecting every registered tree, replicated or not, so WalRetention is effective with no further configuration wherever the scheduler runs. It runs only where the WAL garbage collector is registered: AddLatticeWalGc installs both, and the durable WAL provider registrations (AddAzureTableWalStorage, AddFileWalStorage) and AddLatticeReplication call it for you, but AddLattice alone does not, so a host that supplies its own durable provider through AddWalStorage(factory) calls AddLatticeWalGc itself.

A pass is retention housekeeping, not a latency-sensitive operation. Its cost scales with trees x WalPartitions storage reads (a head scan plus a trim per partition) and runs on every silo, so this ceiling is deliberately coarse to keep the idle storage cost low. Lowering it would only make a quiet tree poll more often, paying for passes that reclaim nothing, while buying no responsiveness on a tree that has work to do - that is what the floor is for. A host that needs a tighter disk bound - a high write rate paired with a small WalRetention - can lower it; TimeSpan.Zero (or any non-positive value) disables the scheduler entirely, restoring the historical caller-driven behaviour.

The first pass is not run at silo start: it is staggered by WalGcStartupDelay, a random offset of half to one full stagger window (15 to 30 seconds by default), so the silo finishes activating before the scheduler adds scan/trim I/O and a rolling cluster restart does not align every silo's fan-out into a correlated I/O storm.

// Tighten the cadence on a high-write durable-WAL host, or disable it.
siloBuilder.ConfigureLattice(o => o.WalGcInterval = TimeSpan.FromMinutes(5));
siloBuilder.ConfigureLattice(o => o.WalGcInterval = TimeSpan.Zero); // disable

The scheduler composes with the replication maintenance grain (which collects replicated trees on its own faster cadence): RunOnceAsync and the underlying WAL TrimAsync are idempotent, and the pass never trims past the minimum consumer cursor or the leaf-materialiser checkpoint floor, so a tree collected by both drivers is trimmed safely and it never over-trims. A per-tree GC failure is logged and skipped without stalling the rest of the pass. This is a global knob read from the default (unnamed) options; per-tree overrides do not apply. It is read once when the scheduler starts; change it before silo start to take effect.

WalGcMinInterval

Floor of the WAL garbage collector's per-tree adaptive cadence band (default: 30 seconds). WalGcInterval is the ceiling.

The scheduler tracks a cadence per registered tree. A pass that reclaims at least one entry sets that tree's next interval to this floor; a pass that reclaims nothing, or throws, doubles it geometrically up to the ceiling. A tree with a growing log is therefore collected every 30 seconds while it has work to do, and a quiet tree relaxes back to the hourly ceiling. The orleans.lattice.wal.gc.interval histogram publishes the cadence chosen for each tree, so a series pinned at the floor means a log that is still growing.

A blocked tree is one exception, and is also held at this floor. A pass that reclaimed nothing because an unusable durable leaf-materialiser pin disabled the consumer-cursor branch (outcome="blocked" on orleans.lattice.wal.gc.passes) is not a quiet tree: it cannot reclaim at all, and its WAL grows without bound. Backing it off would give it the fewest passes precisely when it needs the most, and because the same observation drove both the label and the cadence, the backoff was previously self-reinforcing - being unable to reclaim was itself the evidence used to decide to look less often (issue #2702). Holding it at the floor also bounds recovery: whatever repairs the pin, the stranded bytes do not return until a pass runs and trims them, so this floor is the time-to-reclaim after a repair. The floor introduces no new load level, because a reclaiming tree already runs at it indefinitely, and it is not a latch - a tree that stops being blocked relaxes exactly like any other quiet tree. The other two exceptions are described under WalMaxRetainedBytes: a tree over its byte ceiling is also held at this floor, and a tree that reclaims nothing while its scan keeps stopping on WAL it must retain (outcome="stranded") relaxes only as far as the faulted ladder's cap described below (five minutes at the defaults) rather than to WalGcInterval.

Cost: a pass on a tree with a backlog is exactly the pass that would have run later anyway; the adaptivity moves the work earlier rather than adding it. A tree with nothing to reclaim relaxes off the floor within a few passes. A tree that is blocked and cannot recover polls at the floor indefinitely; that is the alarm state, and its cost is bounded by the floor while the unbounded WAL growth it signals is not.

Separately from the per-tree cadence, the scheduler runs two silo-wide ladders, for the two conditions in which a pass has no per-tree cadence to set at all. Both start at this floor and double geometrically, and both are published as orleans.lattice.wal.gc.scheduler_backoff, tagged cause.

  • cause="empty" - the tree registry was read successfully and holds no collectable tree. That is a correct observation of an idle silo rather than a fault, so this ladder relaxes all the way to WalGcInterval: an empty host costs nothing, while one whose first tree is about to register still picks it up promptly.
  • cause="faulted" - the registry could not be read at all. This ladder is capped far lower, at five minutes clamped into [WalGcMinInterval, WalGcInterval], because the wait a faulted pass picks is the operator's blindness window: the scheduler has learned nothing, so it cannot notice the fault clearing until it next tries. Sharing the quiet ceiling made that window up to a full WalGcInterval - an hour at stock defaults - during which a silo that had already recovered was indistinguishable from one that was dead (issue #3064). The cap bounds recovery without retrying hard: the enumeration that failed is a full key-range scan across every registry shard, so retrying it at the floor forever is how a transient fault becomes a storm.

A pass that finds trees and schedules their ordinary per-tree cadences records cause="scheduled" at the floor instead, meaning no scheduler-wide backoff is in force. That arm is recorded on every healthy pass, so it is also what shows the series is deployed at all.

Either ladder resets to the floor on the next pass that finds work, and the faulted ladder additionally resets on any pass whose enumeration merely succeeded - including one that found nothing. A registry that answered "nothing here" has proved it can be read, which is the only thing that ladder measures. Pair the backoff series with orleans.lattice.wal.gc.scheduler_consecutive_faults, which reports the current fault streak and returns to a measured zero on the first success.

When off: a non-positive value, or any value above WalGcInterval, collapses the band to WalGcInterval, which reproduces the historical fixed-interval tick exactly. Both spellings are honoured; neither is a validation error.

When an operator would turn it off: if the responsive cadence is producing more storage traffic than the retention it buys is worth - typically a host with very many quiet trees whose logs never grow.

Like WalGcInterval, this is a global knob read from the default (unnamed) options once, when the scheduler starts; per-tree overrides do not apply.

WalGcStartupDelay

Upper bound on the randomized delay before a silo's WAL garbage collector runs its first pass (default: 30 seconds). The scheduler draws a uniform offset in [delay / 2, delay), so the first pass lands 15 to 30 seconds after start. The value is capped at WalGcInterval, so a host never waits longer for its first pass than its own ceiling.

The floor of half a window keeps the first pass out of the silo's activation storm; the random component de-correlates silos so a rolling restart does not align every silo's fan-out into one I/O spike.

Before this existed the first pass was drawn from [WalGcInterval / 2, WalGcInterval) - 30 to 60 minutes at the shipped ceiling - which meant a freshly restarted silo did no WAL collection at all for up to an hour. On a box whose leaves cannot resume replay until the WAL prefix is actually trimmed, that hour is a cold start that never converges.

Cost: one fan-out earlier in the silo's life than before.

When off: TimeSpan.Zero (or any non-positive value) means no stagger at all and runs the first pass immediately, forfeiting both the activation-window guard and the de-correlation. To restore the historical deferral instead, set it equal to WalGcInterval: the draw is then [interval / 2, interval), exactly as it was. These are two different settings and it is worth being deliberate about which one you want.

When an operator would change it: lengthen it on a silo whose activation storm is long enough that 15 seconds is still inside it; set it to WalGcInterval when the early first pass itself is the problem.

This is a global knob read from the default (unnamed) options when the scheduler starts.

WalRetention

Optional wall-clock hard ceiling on WAL retention (default: null, disabled). When set, the WAL garbage collector trims entries whose HLC wall-clock is older than now - WalRetention regardless of consumer cursor position, bounding worst-case disk usage even when a registered consumer is hopelessly behind. The lagging consumer then "falls off the log" on its next read, surfacing the gap to the auto-bootstrap trigger (replication-side concern). When null, the GC predicate is purely min(consumer cursors), and a lagging consumer pins the WAL until it catches up. Must be strictly greater than TimeSpan.Zero when set. The core options validator does not enforce that (the replication package's LatticeReplicationOptions.WalRetention, which is copied into this option when this one is unset, is validated): a zero or negative value puts the TTL ceiling at or after the current time, so the TTL clause then admits effectively every entry for trimming.

This option can be changed freely at any time. The new value takes effect on the next GC tick.