Options Reference: Timeout and budget ceiling to MaterialiserCheckpointEntries
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at options-reference-1.md, and llms.txt lists every page.
Part of Options Reference, in Configuration.
Timeout and budget ceiling
The timeout, budget and cadence options the runtime arms as timers - ActivationReadyTimeout, DigestPublishTimeout, EmptyTreeProbeBudget, HotShardSampleInterval, MaxScanPageStallDuration, SetManyEnvelopeBudget, SetManyFanOutBudget, ShardForwardTimeout, ShardHealingInterval, StarvationDriveBudget, WalAdmissionSaturationCallBudget, WalAdmissionSaturationWaitBudget, WalAppendDispatchTimeout, WalDrainBudget, WalFlushPreflightTimeout, WalFlushTimeout, WalSaturationSampleInterval and WalThrottledAdmissionPace - must be at most 0xFFFFFFFE milliseconds (about 49.7 days), the longest wait a .NET timer accepts and the longest period an Orleans grain timer accepts (HotShardSampleInterval and ShardHealingInterval are grain-timer periods). Options validation rejects a longer finite value such as TimeSpan.MaxValue, which would otherwise pass and then fail every operation that armed it. Where an option documents Timeout.InfiniteTimeSpan, use that to remove the bound instead.
That list is not every cadence the runtime arms. The storage-usage poller and the WAL garbage collector clamp a longer StorageUsagePollInterval, StorageUsageDeepPollInterval, WalGcInterval, WalGcMinInterval or WalGcStartupDelay to the same ceiling instead of rejecting it, so each wakes at least that often. CompactionShardTickInterval (the compactor's per-shard timer period) and StorageUsageRollupBudget (a cluster roll-up's cancellation deadline) are not checked against the ceiling at all: a longer finite value passes validation and then throws each time a compaction pass or a roll-up arms it.
Structural sizing (registry-pinned)
MaxLeafKeys, MaxInternalChildren, and ShardCount used to live on LatticeOptions but are now pinned per-tree on the TreeRegistryEntry. They are seeded from LatticeConstants on first tree use (defaults 128 / 128 / 64) and can be changed through:
ILattice.ResizeAsync(newMaxLeafKeys, newMaxInternalChildren)- see Tree Sizing. Runs online; empty-tree fast-path if no data exists.ILattice.ReshardAsync(newShardCount)- see Online Reshard. Grows or shrinks the physical shard count online; an empty tree is re-pinned directly (fast-path).- Pre-pinning the sizing explicitly before the tree's first use with
ILatticeTreeAdmin.CreateTreeAsync(treeId, shardCount, maxLeafKeys, maxInternalChildren)from Orleans.Lattice.Api.TreeAdmin. Creation is idempotent, and the supplied sizing is honoured only when the call registers the tree for the first time. - Declaring it in an installed app's manifest (Orleans.Lattice.Apps), whose per-tree
shardCount,maxLeafKeys,maxInternalChildren,walPartitionsandvirtualShardCountpins are applied when the install first registers each tree the app creates; a tree that already exists keeps its pinned structure.
Virtual shard space (constant)
The virtual shard space is 4096 virtual slots for every tree except one created by an installed app whose manifest declares a virtualShardCount (see Orleans.Lattice.Apps): that tree's persisted ShardMap is created with the declared slot count when the tree is first registered, and routing always hashes over the slot count the tree's map holds. A resize carries that map over to the resized copy and a reshard keeps its slot count, so the declared count is not lost to either, and a reshard target cannot exceed it. At the default, keys hash into [0, 4096), and the per-tree ShardMap collapses ranges of virtual slots onto physical shards. This indirection decouples logical key routing from the physical shard count, enabling adaptive shard splitting without rehashing existing keys.
The default identity map (virtual slot i routed to physical shard i % ShardCount) reproduces the legacy hash % ShardCount routing exactly only when the pinned ShardCount divides 4096 evenly. That is a property of the arithmetic, not a checked invariant: ShardMap.CreateDefault rejects only a non-positive shard count or one larger than the virtual space, so a count that does not divide 4096 still routes every slot to a shard but no longer reproduces hash % ShardCount. The 4096 default is a compile-time constant - a tree with no persisted ShardMap routes over it, so changing it in source would re-route that tree's keys, and it is treated as a breaking wire-format change.
AdmissionAdvisoryBytes
Optional non-enforcing advisory ceiling, in bytes, on a tree's estimated storage footprint (the same total MaxEstimatedBytes is checked against), used to right-size MaxEstimatedBytes before turning enforcement on. null (the default) disables the byte advisory dry-run signal. When set it must be at least 1 (enforced by the options validator). A tree over this ceiling is flagged by the orleans.lattice.admission.over_advisory gauge, and each call this ceiling would have rejected increments orleans.lattice.admission.would_reject (dimension bytes) once - counting only the write calls that check the enforcing caps, listed under MaxEstimatedBytes - but the advisory itself never rejects a write. Resolvable per tree.
// Dry-run a 512 MiB byte ceiling: no rejections, just the would-reject signal.
siloBuilder.ConfigureLattice("bulk-ingest", o => o.AdmissionAdvisoryBytes = 512L * 1024 * 1024);
See Metrics for the advisory-first-then-enforce adoption workflow.
AdmissionAdvisoryLiveKeys
Optional non-enforcing advisory ceiling on a tree's live (non-tombstone) key count, used to right-size MaxLiveKeys before turning enforcement on. null (the default) disables the live-key advisory dry-run signal. When set it must be at least 1 (enforced by the options validator). Drives the same non-rejecting orleans.lattice.admission.over_advisory and orleans.lattice.admission.would_reject (dimension keys) signals as AdmissionAdvisoryBytes, for the key dimension. Resolvable per tree.
// Dry-run a 1,000,000 live-key ceiling before enforcing it.
siloBuilder.ConfigureLattice("bulk-ingest", o => o.AdmissionAdvisoryLiveKeys = 1_000_000);
AtomicActionRetention
Retention window for a terminal atomic-action saga's coordinator state (default: 48 hours). Once an IAtomicActionGrain.ExecuteAsync plan reaches a terminal outcome, the coordinator retains its persisted progress for this window so that re-issuing the same operation id returns the memoized outcome rather than re-running the plan. After the window expires the coordinator clears its state and a re-issue starts a fresh saga. Minimum effective interval is 1 minute (Orleans reminder granularity). Set Timeout.InfiniteTimeSpan to retain saga state indefinitely. This is the atomic-action analogue of the atomic-write retention below and is configured independently of it. It is read from the default (unnamed) options, so per-tree overrides do not apply. See Atomic Actions.
AtomicWriteRetention
Retention window for completed SetManyAtomicAsync saga state (default: 48 hours). After a saga reaches a terminal state, its coordinator grain retains its persisted progress for this window so duplicate submissions with the same operation ID are idempotent. A retention reminder fires at the end of the window and clears the state. Minimum effective interval is 1 minute (Orleans reminder granularity); smaller non-infinite values are effectively floored at that granularity. Set Timeout.InfiniteTimeSpan to disable automatic cleanup. A single-tree saga resolves the window per tree; everything else that retains state for this window - the cross-tree transaction's coordinator and receiver, the materialised-view cross-tree coordinator, and, with Orleans.Lattice.Replication, the cross-cluster saga coordinator, participant and write fence - reads it from the default (unnamed) options. See Atomic Writes.
This option can be changed freely at any time.
AutoSplitEnabled
Master switch for adaptive shard splitting. When true (the default), the tree's hot-shard monitor periodically polls shard hotness counters and triggers splits when a shard's ops/sec exceeds HotShardOpsPerSecondThreshold. When false, no autonomic splits occur; the physical shard count then changes only through an explicit ILattice.ReshardAsync or automatic over-split healing, which is deliberately not gated on this switch.
This option can be changed freely at any time. Turning it off takes effect on the monitor's next sampling pass. Turning it back on resumes a monitor that has run before - on its next sampling pass if its timer is still armed, otherwise on its next one-minute keepalive reminder - but a tree whose monitor has never started, because the tree's grains have only activated while the switch was off, does not start one until those grains next activate.
AutoSplitMinTreeAge
Minimum tree age before the hot-shard monitor begins sampling (default: 60 seconds). Prevents splits during initial bulk-load bursts that would otherwise appear as sustained hot-shard traffic.
This option can be changed freely at any time.
CacheTtl
Minimum time between consecutive delta refreshes from the primary leaf in a leaf's read-through cache, when that primary leaf is activated on another silo. When set to TimeSpan.Zero (the default), every such read triggers a delta refresh - the version-vector comparison on the primary is cheap but the RPC overhead remains. Setting a non-zero value allows the cache to serve reads from its local dictionary without contacting the primary, trading freshness for lower read latency.
The TTL does not govern a cache whose primary leaf is activated on the same silo. After its first refresh such a cache compares a revision cookie the primary advances on every state change: while the revision is unchanged it serves from its local dictionary with no refresh at all, and when the revision has moved it refreshes immediately, whatever the TTL. The TTL is therefore a bound on cross-silo refresh traffic.
// Allow up to 100 ms of staleness for lower read latency
siloBuilder.ConfigureLattice(o => o.CacheTtl = TimeSpan.FromMilliseconds(100));
// Per-tree: aggressive freshness for a real-time tree
siloBuilder.ConfigureLattice("realtime", o =>
{
o.CacheTtl = TimeSpan.Zero; // refresh on every cross-silo read (default)
});
This option can be changed freely at any time. The new TTL takes effect on the next read. A value of TimeSpan.Zero preserves the original cross-silo behaviour (refresh on every read).
BackgroundDrainLeavesPerPass
Maximum number of source leaves a background coordinator - the split drain, the cross-tree merge drain, or the online snapshot copy - visits in a single pass before persisting its resume cursor and yielding (default: 64 leaves).
These walks do not hold a shard root, so unlike an unbounded read walk they cannot head-of-line-block user traffic. What they do hold is their own non-reentrant coordinator, so an unbounded pass over a thousand-leaf shard makes that coordinator unable to answer a progress query or honour a cancellation for the whole sweep, buffers the sweep's working set for its whole duration, and puts the entire sweep at risk from a single interruption. Bounding the pass turns all three into steady background work resumable from a persisted key cursor.
Raising it drains faster at the cost of longer individual turns; setting it to 0 or less disables the bound and restores the unbounded walk. It does not govern the authoritative post-freeze sweeps in a split's Swap / Complete phases or a consolidation's Swap / Complete phases, which are deliberately unbounded - see Bounded background leaf walks below.
The tombstone compactor, the shard consolidator and the empty-leaf reclaim walk keep their own long-standing per-pass leaf caps (CompactionLeafBatchSize governs the first and the third, ConsolidationDrainLeavesPerPass the second) and inherit only the wall-clock net from BackgroundDrainMaxDuration. The tombstone compactor inherits that net at two levels: once for its own leaf walk, and once inside each leaf's compaction turn. The shard consolidator also bounds how many donor entries it accumulates into each merge call to the survivor shard; see ConsolidationDrainBatchSize.
// Yield more often on a tree whose coordinators must stay responsive to
// progress queries and cancellation while a large shard drains.
siloBuilder.ConfigureLattice("large-shard-tree", o => o.BackgroundDrainLeavesPerPass = 16);
This option can be changed freely at any time. It takes effect on the next pass.
BackgroundDrainMaxDuration
Wall-clock safety net for a single background coordinator drain pass (default: 10 seconds). When a pass has spent this long it persists its resume cursor and yields at the next leaf boundary that offers one.
BackgroundDrainLeavesPerPass is the primary, deterministic bound; this covers the case a leaf count cannot, where a small number of leaves are individually very slow - cold activations rehydrating large snapshots, or a fan-out merge into many target shards. Set to TimeSpan.Zero to disable it and rely on the leaf count alone.
The empty-leaf reclaim walk inherits this same net. Reclaim is the one walk here that does hold a shard root's turn, so it is the walk where the distinction matters most: its leaf bound counts probes, and the cost of a probe varies by orders of magnitude between a warm activation and a cold one rehydrating from storage, so a probe budget tuned for a warm shard is not a time bound on a cold one. A pass has been observed holding a shard root for 59.2 seconds with a user scan enqueued behind it, which is head-of-line blocking on that shard for the whole of it. Because queue wait falls outside the scan's own page-fill ceiling, that surfaces to the caller as an Orleans TimeoutException rather than as a stall the scan can report on itself. See How fast a shard actually heals.
Every pass logs its elapsed time, probe count, fold count and stop reason at information level, including passes that fold nothing, so the bound can be seen binding in the field rather than only asserted in a test.
A single leaf's tombstone-compaction turn runs under this net as well, which is the one place it bounds work inside a leaf rather than a walk across leaves. A leaf's reap loop awaits one WAL append per condemned entry, so its cost scales with the tombstone count and an unbounded turn overran the request timeout on exactly the leaves that most needed compacting. Reusing this net rather than adding a knob keeps the budgets on that path reconcilable: at the default a turn stops issuing new appends after 10 seconds and the append already in flight can block for the whole 15-second storage busy window, giving a 25-second worst case against Orleans' 30-second timeout. Raising this much past 15 seconds spends that margin. A truncated turn still reaps at least one condemned entry, and every entry it reaps stays reaped, so each pass finds strictly fewer condemned entries than the one before until the leaf drains; the turn reports itself incomplete so the coordinator keeps the leaf nominated; see Tombstone Compaction.
The orphaned-leaf passes - ILattice.InspectOrphanedLeavesAsync, SurveyOrphanedLeavesAsync and RepairOrphanedLeavesAsync - run under this net too. Each shard batch holds the shard root's turn and yields once it has spent this long, and the tree-level call checks the same budget between shard batches, so one call is bounded by roughly twice this value (about 20 seconds at the default). See Driving a pass to completion.
It also sizes the slice an online resize drives its snapshot drain in. Each slice copies for this long, capped at 10 seconds so it returns well inside Orleans' 30-second response timeout, then persists every shard's resume key and hands the snapshot's turn back so its keepalive reminder is not starved. Because that bound is load-bearing for the caller, TimeSpan.Zero does not disable it: a zero value falls back to the 10-second cap. See Tree Sizing.
This option can be changed freely at any time.
Bounded background leaf walks
The background walks below traverse a shard's leaf chain. Each is either work-bounded and resumable or deliberately atomic, and which one it is follows from what the surrounding protocol needs rather than from how long the walk happens to be. Most run on a background coordinator. Empty-leaf reclaim, the moved-away seal and the tree purge instead run on the shard root itself, so a call to that shard that does not interleave with them - a range-scan page, for example - waits behind them. Of those three only reclaim is bounded, which is why its bound is the one a user request feels directly:
| Walk | Bounded? | Why |
|---|---|---|
Split Drain phase |
Yes | Routing still points every moved slot at the source, and the walk already flushed to the target in batches, so a pass boundary makes nothing newly observable. |
Split Swap / Complete final sweeps |
No | Each makes the target provably equal to the source's final committed state immediately around the routing flip. The source is frozen for the moved slots, so extra time cannot change what the sweep reads - only how long routing is in flux. |
| Split retroactive prepared-mutation sweep | No | The Drain phase that follows imports pre-saga values into the target; the destination-side shadow markers this sweep installs are what stop a reader surfacing one of them. A half-finished sweep would leave markers installed for some in-flight sagas and not others. Its recovery contract is also "re-run the whole sweep", which is what makes its per-snapshot status re-check sound for every leaf. |
| Cross-tree merge drain | Yes | The merge has never been atomic - it merges each leaf's delta as it goes - and its contract is convergence, not a point-in-time image. |
| Online snapshot copy | Yes | An online snapshot is a converging mirror, not a point-in-time image: shadow-forwarding mirrors concurrent writes onto the destination throughout, and shards are already copied one at a time across ticks. |
| Offline snapshot copy | No | The destination shard is assembled bottom-up by a single bulk load, which by contract needs the complete sorted entry set and an empty destination, so there is no intermediate position to resume from. The source is quiesced for the copy's duration. |
| Tombstone compaction | Yes | Per-leaf compaction is idempotent and independent. |
| Empty-leaf reclaim | Yes | A fold is a complete, independently correct unlink; a pass boundary between two folds leaves the chain in a state no reader can distinguish from one where the pass simply had less work. Unlike every other bounded row, this walk holds the shard root, so its bound is the only configurable limit on how long a user read or write waits behind a background walk. |
Consolidation Drain phase |
Yes | The donor still serves traffic and still shadow-forwards; the survivor's copy only becomes authoritative in Swap. |
Consolidation Swap / Complete final sweeps |
No | Both run inside the freeze window, where nothing they read can change and holding the window open across ticks is a real availability cost. |
| Moved-away seal | No | In a split's or a consolidation's Swap, before the shard giving up the moved slots starts rejecting them, its shard root marks every leaf in its chain as having moved those slots away. In a consolidation the surviving shard, which usually sealed the same slots itself when the donor was split out of it, then lifts that seal from its own leaves the same way before the routing flip. Each runs in one turn from the first leaf to the last, because a half-installed seal errs in the dangerous direction: a leaf that is already sealed reports a key its shard still owns as absent. Its recovery contract is to re-run the whole walk, which is idempotent. The per-leaf writes go out a fixed 32 at a time, and no option bounds the walk. |
| Tree purge | No | It runs only on a deleted tree, whose reads and writes already throw, so a request it delays could not have succeeded anyway. It is also destructive, so it has no position to resume from: clearing a leaf removes the sibling pointer that leads to the next one, which is why the walk reads that pointer first. |
Every resumable walk records its position as a key, never a leaf grain id. Orleans grains are virtual, so an id persisted across a pass boundary can activate a fresh, empty grain whose sibling pointer is null - a resumed walk would conclude it had reached the end of the chain and stop, silently leaving the rest of the shard unvisited with no exception, metric or log line to show for it. A key is always re-descended onto whichever leaf now owns it, so a leaf that has been split, or reclaimed, between two passes cannot truncate the sweep. (Empty-leaf reclaim holds its key cursor in the shard root's activation rather than in persisted state, so it restarts at the head of the chain when the activation recycles. That is correctness-neutral: the cursor only decides where a pass begins, never what it is willing to fold.)
A walk yields only where it can name that position - the visited leaf's exclusive high bound, which is exactly where the next leaf begins. A leaf that declares no usable high bound is not a stopping point: the walk keeps going rather than stop without a resume position, so a misconfigured bound degrades to the unbounded walk rather than to a truncated sweep.
Each deliberately-atomic walk is instead made attributable: when it holds its grain for more than ten seconds it logs a warning naming the operation and the leaf count, so a stalled grain explains itself rather than surfacing only as a flood of Orleans long-request warnings.
CompactionLeafBatchSize
Maximum number of leaves the tombstone-compaction coordinator visits within a single shard before yielding for one CompactionShardTickInterval. Default 64 leaves, floor 1. The leaf walk resumes on the next timer tick from a persisted in-shard cursor, so progress survives silo crashes the same way the shard cursor does. This is the dominant control on peak concurrent leaf activations during a pass: with batching, peak activations are bounded by CompactionLeafBatchSize * (CollectionAge / CompactionShardTickInterval) regardless of tree size. The default 64 reproduces pre-batching behaviour exactly on shards with <= 64 leaves (the common case). Values below 1 are clamped to 1 with a one-shot warning per tree per process. Snapshotted at pass start; changing the option mid-pass does not reshape the in-flight pass.
// Cut peak concurrent leaf activations by yielding more aggressively
// within each shard. Trades pass wall-clock for activation headroom.
siloBuilder.ConfigureLattice("activation-sensitive-tree", o => o.CompactionLeafBatchSize = 16);
For the relationship between CompactionLeafBatchSize and CompactionShardTickInterval, the multiplicative activation-pressure bound, and the worked-example table, see Tombstone Compaction - CompactionLeafBatchSize.
CompactionShardTickInterval
Gap inserted between consecutive per-shard ticks during a tombstone-compaction pass. Default 500 milliseconds, floor 100 milliseconds. The cadence is a scheduler-fairness knob and the dominant control on activation pressure during a pass - lowering it shortens the pass but raises the peak concurrent leaf activation count. The cadence is snapshotted at the start of each pass and can be changed freely at any time; the next pass picks up the new value. The default was lowered from 2 s to 500 ms once the dirty-leaves fast path landed - on a tree with no recent deletes a pass activates only the shard root grains, so the tighter cadence is safe to ship by default.
For the worked-example trade-off table, the activation-pressure model, the relationship to GrainCollectionOptions.CollectionAge, and operator-triage guidance (use ILattice.CompactShardAsync for "compact one shard fast"), see Tombstone Compaction - CompactionShardTickInterval.
ConsolidationDrainBatchSize
Maximum number of donor entries the online shard consolidator accumulates into a single merge call to the survivor shard while it drains a donor (default: 1024 entries). Larger values reduce per-call overhead; smaller values bound peak memory on the coordinator silo and the size of the Orleans grain message. The drain is idempotent under any chunking - re-running converges by CRDT last-writer-wins - so this is purely a cost knob and never a correctness input. The options validator rejects a value below 1, because a fold that flushes no entries could never drain its donor.
This option can be changed freely at any time. The consolidator reads it at the start of each drain pass, so a new value applies from the next pass.
ConsolidationDrainLeavesPerPass
Maximum number of donor leaves the online shard consolidator visits in a single background drain pass before persisting its resume cursor and yielding (default: 16 leaves). This is what turns consolidating a thousand-leaf donor from one unbounded stall into steady background work: each pass does a bounded amount of drain, records the key it reached, and lets the next timer tick continue, and a crash resumes from that persisted key instead of restarting the sweep. Set it higher to consolidate faster at the cost of longer individual turns on the donor's leaves. Each pass also inherits the wall-clock net of BackgroundDrainMaxDuration.
The options validator rejects a value below 1. Unlike BackgroundDrainLeavesPerPass, where 0 or less disables the bound, a pass that visits no leaf would leave a fold sitting in its drain phase indefinitely rather than disabling anything. To stop automatic healing admitting new folds, set MaxConcurrentShardConsolidations to 0 or ShardHealingEnabled to false.
This option can be changed freely at any time; a new value applies to the next pass.
CursorIdleTtl
Sliding idle timeout for every stateful cursor - one opened via OpenKeyCursorAsync, OpenEntryCursorAsync, OpenDeleteRangeCursorAsync, OpenSnapshotKeyCursorAsync or OpenSnapshotEntryCursorAsync, or the WherePredicate variant of any of them (default: 48 hours). Each successful cursor step refreshes the reminder; if it fires without intervening activity the cursor grain clears its persisted state, unregisters the reminder, and deactivates. Minimum effective interval is 1 minute (Orleans reminder granularity); smaller values are clamped to the floor. Set Timeout.InfiniteTimeSpan to disable automatic cleanup - cursors then live until CloseCursorAsync is called. See Durable Cursors.
This option can be changed freely at any time.
DefaultLockLeaseDuration
The lease ILatticeLockGrain grants when an acquire supplies a non-positive LockAcquireRequest.LeaseDuration, or passes a non-positive duration to TryAcquireAsync - that is, when the caller defers to the server default (default: 30 seconds). A holder that neither renews nor releases before its lease elapses has the lock reclaimed and handed to the next FIFO waiter, so this value bounds how long a crashed holder can wedge a lock. The validator requires it to be positive. Both lock options are read from the default (unnamed) options, so per-tree overrides do not apply. See Distributed lock.
MaxLockLeaseDuration
The ceiling every granted or renewed lease is clamped to (default: 5 minutes). A caller cannot pin a lock for longer than this even by requesting a larger duration - the grant is silently capped - so a misconfigured client cannot hold a contended lock for hours. The validator requires it to be positive and at least DefaultLockLeaseDuration.
DiagnosticsCacheTtl
How long the internal diagnostics grain caches a TreeDiagnosticReport before assembling a fresh sample (default: 5 seconds). ILattice.DiagnoseAsync is an admin-rate API; caching coalesces repeat callers (e.g. dashboards polling every few seconds) so that a single fan-out walks every shard rather than one per call.
Shallow (deep: false) and deep (deep: true) reports are cached independently. The cache is invalidated immediately when an adaptive split commits, so the next call after a topology change always returns a fresh report.
Set to TimeSpan.Zero to disable caching entirely - every call assembles a new report. This is useful in tests or for tight polling scenarios where staleness is unacceptable.
// Disable caching for a debug tree
siloBuilder.ConfigureLattice("debug-tree", o => o.DiagnosticsCacheTtl = TimeSpan.Zero);
This option can be changed freely at any time. The new TTL takes effect on the next DiagnoseAsync call.
DigestCoalescingWindowMs
How long (in milliseconds) a BPlusLeafGrain defers a pending cross-grain projection-digest publish to its parent internal node, coalescing multiple per-mutation publishes into a single hop (default: 5 - the c2-xxviii measured sweet spot at the c2-iii operating point, a 27% drop in caller-visible SetAsync p50). When set to a positive value, the first mutation that changes the leaf's digest arms a one-shot grain timer; subsequent mutations arriving within the window observe the pending timer and skip rescheduling, so N writes share one cross-grain publish to the parent, which the timer sends after the window, outside any write's commit. Only that publish is deferred: the leaf's own running hash still advances in memory on every mutation. It is not written to storage by each write, though; it is persisted with the rest of the leaf's state, normally at the leaf's next checkpoint flush (see MaterialiserCheckpointInterval).
Coalescing is scoped to the per-write hot path: sets, deletes and range deletes (SetAsync, SetManyAsync, DeleteAsync, DeleteRangeAsync, with their conditional and predicate forms) and typed CRDT delta applies (ApplyCrdtDeltaAsync, ApplyCrdtDeltaManyAsync), replicated deltas included. Whole-entry merges - the migration drains and shadow-forwarding of splits and consolidations, replicated last-writer-wins writes, tree merges, snapshot copies, restores and bulk loads - and structural events (leaf split, projection rebuild, saga terminal apply, tombstone-reap compaction, checkpoint flush) bypass the window and publish synchronously so operator-tooling oracles (e.g. RebuildLeafProjectionAsync followed by GetLeafProjectionDigestAsync) observe post-publish state without a settle delay.
The window is ignored on a leaf whose resolved MaintainProjectionDigest is false, because there is then no upward publish to coalesce.
Set to 0 to restore the synchronous-publish shape on every path - useful for tests that issue read-after-write digest oracles against a parent internal node within the same task continuation, or for operators whose downstream consumers depend on bit-exact synchronous publish timing.
// Restore the synchronous-publish shape for a test tree
siloBuilder.ConfigureLattice("test-tree", o => o.DigestCoalescingWindowMs = 0);
This option can be changed freely at any time. A leaf reads the window with the rest of its resolved options when it activates, so a new value takes effect on each leaf's next activation.
EventStreamProviderName
Name of the Orleans stream provider Lattice publishes LatticeTreeEvent notifications onto (default: "Default"). The same name must be configured on every silo (publishers) and on the client (subscribers); register the provider via the standard siloBuilder.AddMemoryStreams("Default") / equivalent durable-stream extension. Only consulted when PublishEvents is true.
This option can be changed freely at any time. The new value takes effect on the next publish.
HotShardConsolidationSkewRatio
The load-skew ratio at or below which a tree counts as uniformly loaded, and is therefore a candidate for automatic over-split healing rather than for a split (default: 1.15). Skew is maxShardRate / medianShardRate, so a perfectly even tree scores exactly 1.0.
This is the lower edge of the split/heal hysteresis band; HotShardMinSkewRatio is the upper edge. The interval between them is a dead band in which neither loop acts, and it is what stops a tree being split, folded, and split again forever. Options validation rejects a configuration in which the two regions could overlap, so the separation is structural rather than a convention two loops politely observe.
Cost: none. It is compared against a statistic the healing sweep already computes.
When off: there is no "off". Raising it towards HotShardMinSkewRatio narrows the dead band and makes healing act on less uniform trees; lowering it makes healing more conservative. To stop healing entirely use ShardHealingEnabled.
When an operator would change it: on a tree whose steady-state load is genuinely a little uneven and which therefore never quite qualifies as uniform, so an over-split tree sits at a permanent backlog. Raise it, but keep it strictly below HotShardMinSkewRatio.
HotShardMinShardEntries
Minimum live-entry count a shard must hold before an autonomic split is admitted (default: 1024). Splitting a nearly empty shard multiplies grain activations and relieves nothing.
Cost: one CountAsync leaf-chain walk per candidate shard, and only for candidates that already cleared every cheaper admission clause. A uniformly loaded tree produces no candidates and pays for no probes at all.
When off: 0 disables the occupancy floor and its probe, restoring the pre-#1834 behaviour in which a hot but nearly empty shard could be split. Negative values are rejected by options validation.
When an operator would turn it off: if the occupancy probe is measurably expensive on a very wide tree, or when deliberately pre-splitting a tree that is about to be loaded and is legitimately empty at admission time.
HotShardMinSkewRatio
The load-skew ratio at or above which a tree's load counts as concentrated enough for a split to relieve anything (default: 1.5).
This is the clause that distinguishes a genuine hot shard from a bulk ingest. Rate alone cannot: a bulk write drives every shard far above HotShardOpsPerSecondThreshold at once, and splitting under that shape relieves nothing while permanently multiplying grain activations - which is exactly how a tree ends up with an order of magnitude more shards than it needs. The gate is applied once per pass at tree level, not per candidate, because a per-candidate comparison would refuse the second-hottest shard of a genuinely skewed tree and silently regress multi-shard relief.
Cost: none beyond a median over the per-shard rates the pass already collected, computed in a pooled buffer.
When off: a value at or below 1.0 disables the clause entirely, restoring pure rate-based admission. Note that reverting the shape gate is not the same as reverting to the pre-#1834 behaviour: MaxPhysicalShardsPerTree and HotShardMinShardEntries are separate clauses that did not exist before, and both must be zeroed as well.
When an operator would turn it off: rarely, and only after confirming the refusals really are hot shards rather than bulk ingest - the orleans.lattice.split.admission.deferred counter carries a uniform_load reason tag that says which. Do not raise the value above roughly 1.6: an existing regression test splits a two-shard tree at rates 500/800 (a skew of 1.6), so a higher default would break legitimate hot-shard relief.
HotShardOpsPerSecondThreshold
The ops/sec threshold on a single shard that triggers an adaptive split (default: 200). Lowering this value makes the system more aggressive about splitting; raising it allows shards to absorb more load before splitting.
This option can be changed freely at any time.
HotShardSampleInterval
How often each tree's hot-shard monitor - the per-tree sampler that decides when to trigger an adaptive split - polls every shard's hotness counters (default: 30 seconds). Shorter intervals increase detection responsiveness at the cost of more grain calls. A non-positive value falls back to the 30-second default. The value is the monitor's grain-timer period, so options validation rejects one above about 49.7 days; see Timeout and budget ceiling.
This option can be changed freely at any time. The monitor arms its sampling timer with this value when it starts, so a new cadence takes effect when the monitor next activates.
HotShardSplitCooldown
Minimum time between consecutive splits of the same shard (default: 2 minutes). Prevents rapid re-splitting before the post-split load distribution has stabilised.
This option can be changed freely at any time.
KeysPageSize
The number of keys (or entries) requested from each shard per page during ordered key scans (ScanKeysAsync) and entry scans (ScanEntriesAsync) (default: 512). Larger pages reduce the number of grain calls at the cost of larger messages. This is a performance tuning knob and does not affect tree structure. It must be greater than 0; the options validator rejects zero and negative values.
This option can be changed freely at any time. It takes effect on the next scan.
LeafCachePreWarmCount
Number of LeafCacheGrain activations each shard root primes when
ILattice.WarmUpAsync runs (default: 8, enabled). Paired with
LeafAccessModelFlushIntervalMs, which is the coalescing window (in
milliseconds, default 30000) for persisting the ranking model.
The default matches the shard root's own pre-warm fan-out concurrency, so warm-up issues exactly one bounded wave of priming calls per shard and never queues behind its own semaphore: warm-up costs one round trip regardless of how many shards the tree has. It is one eighth of the hard ceiling of 64, leaving ample headroom for a deployment with a wider hot set.
LeafCacheGrain is a [StatelessWorker] read-through cache. After a silo
restart every activation is cold, so the first read of each leaf pays activation
plus a full delta pull from its primary leaf - a latency spike concentrated on
exactly the leaves a skewed workload reads most. Setting a positive count asks
each shard root to pay that cost up front, at warm-up, off the critical path of
the first real read.
The ranking is not a recency list. Each shard root keeps a bounded histogram of the leaves its reads route to: every routed cache read increments the target leaf's visit count. At warm-up the histogram is ranked by observed read frequency - "what fraction of reads land here", rather than "what happened to be touched last". Under a skewed or cyclic key distribution the two disagree sharply: a leaf touched once just before shutdown outranks a genuinely hot leaf under recency, but not under frequency. Measured on held-out synthetic traces (train on the first half, score against the hot set of the second), frequency recovered 96% of the true hot set on a Zipf-skewed trace and 98% on a cyclic/sequential one, against 56% and 53% for a recency list of the same size.
A first-order Markov chain over leaf identities, ranked by personalised PageRank, was implemented and measured first. It never beat the plain histogram on any trace - it lost by 12.5 points on the skewed trace and 3.1 on the cyclic one - because for a chain fitted to a single observed trajectory the empirical visit vector is already stationary for the fitted matrix, so the ranking pass reproduced its own input. The transition rows were removed rather than carried at roughly 100 KB resident per shard-root activation for no ranking benefit.
Bounds, all fixed and independent of the tree's key space:
| Bound | Value | Effect |
|---|---|---|
| Resident tracked leaves | 256 leaves | Coldest 25% pruned when exceeded. |
| Persisted leaves | 64 leaves | Upper bound on LeafCachePreWarmCount. |
| Persisted snapshot size | roughly 3 KB | Rides inside the shard root's own state. |
| Pre-warm fan-out concurrency | 8 in flight per shard | Bounded like the shard fan-out it nests inside. |
Operational characteristics:
- On by default, with
0as the kill switch. At0no access is tracked, nothing is persisted, and warm-up behaves exactly as it did before the option existed. Valid values are0to64inclusive; a value outside that range fails options validation. An operator would turn it off to remove the per-shard-root model write from a read-saturated box, or to isolate the feature while diagnosing a warm-up fault. - Zero read-path cost when disabled, and an O(1) allocation-free record when
enabled. The read path never awaits a storage write: a grain timer persists the
model at most once per
LeafAccessModelFlushIntervalMs, and only when it has changed. Clean deactivation flushes once more. - Best-effort. A failure to prime any individual leaf is swallowed, so
pre-warm can never fail
WarmUpAsync.orleans.lattice.warmup.leaf_cache.prewarmedcounts the leaves primed successfully, so a failure shows only as a shortfall there. - Correct silo locality. The shard root is the only caller of the leaf cache, so the stateless-worker activations warm-up creates land on the silo that will serve the subsequent reads.
- Loss bound. An ungraceful silo kill loses at most one flush window of observations. A missing or stale model simply pre-warms fewer (or less useful) leaves - it can never produce an incorrect read.
Set LeafAccessModelFlushIntervalMs to 0 to persist the model only on clean
deactivation: free under read load, but the model is lost entirely on an
ungraceful kill.
See Metrics for the three instruments this feature publishes.
Both options can be changed freely at any time; they take effect on each shard root's next activation.
LeafHydrationResidentBytes
Ceiling on the resident footprint of a partially hydrated leaf's snapshot, in bytes (default: 1 MiB). Once a leaf's materialised ranges exceed it, the least recently used clean hydrated blocks are evicted. Only applies when LeafPartialHydrationEnabled is on.
The default sits above the mean payload leaf measured on the reference deployment and roughly eleven times its mean metadata leaf, so an ordinary leaf hydrates on demand and never evicts; the bound bites only on leaves materially larger than the measured shape.
Cost: an eviction re-reads the evicted range from the frame if it is needed again. A block that has taken a mutation is pinned for the rest of the activation and is never evicted, so eviction can never lose a write or resurrect a removal.
When off: 0 means unbounded - hydrate on demand and never evict. That is not a kill switch for the mechanism, only for its eviction half. Negative values are rejected by options validation.
When an operator would change it: raise it on a read-heavy tree with very large leaves whose working set is being evicted and re-read; set it to 0 when memory is plentiful and the re-read cost matters more than the footprint.
Interaction with MaxLeafBytes: the leaf read, digest and freeze seams walk bounded key windows rather than the whole cache, and that bounds peak residency only because eviction can shed a window once the walk has moved past it. Eviction only runs while the resident footprint exceeds this budget, so a budget that is unbounded (0) or is not materially smaller than MaxLeafBytes sheds nothing: those walks stay windowed in shape and become whole-leaf in cost. The lazy hydration frame is retained across a completed ranged hydration rather than released, so in that regime the frame's bytes sit on top of a fully resident leaf. The degradation is silent - every seam still returns the same answer - so the silo logs a one-shot advisory per tree naming both configured values when this budget is unbounded or is at or above a tenth of MaxLeafBytes. It is a performance advisory, not a rejection: the configuration is legitimate on a host with ample memory, and nothing is clamped. To restore the bound, set this budget well below MaxLeafBytes.
LeafPartialHydrationEnabled
Whether a leaf activating from a binary snapshot attaches the frame as a lazily hydrated backing store instead of decoding every row up front (default: true).
This is what stops activation cost being a function of blob size. The cache reports the whole snapshot's row count, footprint, and live count immediately, so nothing observable distinguishes it from a fully decoded cache; entry ranges are materialised out of the frame only as reads require them. A key-addressed read materialises the one block holding the key; ContainsKey answers from the frame's index table and materialises nothing.
Cost: a seek and a bounded decode on the first read of each range, instead of one whole-blob decode at activation. A leaf that takes a write while maintaining its projection digest still materialises in full, because the canonical full-walk hash cannot be folded over rows the leaf has not read - so the saving is on the read-dominated cold-start path, which is the path that matters here.
When off: the rehydrate decodes every row up front, exactly as it did before the mechanism existed. The switch is read once per rehydrate and touches no persisted shape, so it is safe to flip in either direction on a running deployment and needs no migration.
When an operator would turn it off: to isolate the mechanism while diagnosing an activation-path fault, or on a workload whose every activation reads the entire leaf anyway, where the seeks are pure overhead.
LeafProjectionRetention
Age beyond which a leaf's persisted projection checkpoint counts as stale at activation (default: 7 days). It is an advisory cost signal, like MaxLeafReplayEntries, not a recovery trigger: an old checkpoint does not imply the WAL has been trimmed, so when the log still covers the window the leaf needs it tail-replays and converges normally, and only a genuine trim past the checkpoint consults ProjectionRebuildPolicy. The activation path currently supplies no checkpoint age - the persisted checkpoint carries no capture timestamp - so this trigger never fires from activation today; the option is retained because it is public API and the detector honours it for any caller that does supply an age. Set to Timeout.InfiniteTimeSpan to disable the age-based trigger; any other non-positive value is rejected by options validation.
This option can be changed freely at any time.
LeafRetirementRetryDeadline
How long a write blocked by a leaf-retirement latch keeps retrying before it fails (default: 2 seconds).
Empty-leaf reclaim latches a leaf closed for the handful of grain calls between deciding to fold it and clearing its state. A write routed to that leaf inside the window is refused, never applied: applying it would put a record into state that is about to be destroyed. Refusing is what makes the fold safe, but on its own it would turn every write unlucky enough to arrive during an ordinary fold into a caller-visible error. This deadline is what closes that gap - the write retries with exponential backoff and jitter until the latch clears and the range is served by the leaf that now owns it.
The default is derived from two regimes rather than picked:
- Transient regime. The latch is held across the compare-and-swap and the routing retirement, and the latter makes up to three attempts at its parent call, so the window's ceiling is four grain calls: sub-millisecond co-located, single-digit milliseconds cross-silo, plus scheduling jitter. Tens of milliseconds is a generous reading, and 2 seconds clears that by about two orders of magnitude, so jitter alone can never exhaust it.
- Stuck regime. If those calls are instead hitting the Orleans response timeout (30 seconds by default), the same arithmetic gives about two minutes. The default deliberately does not cover that. A fold blocked for minutes is an outage, not a transient overlap, and a write that inherited its wait would exceed any caller's own timeout while reporting nothing.
Exhausting the deadline throws. That is the intended behaviour, and it is the entire point of the interlock: it converts a silent loss of an acknowledged write into a loud failure the caller can retry or surface.
One caveat worth stating plainly, because the safety argument rests on it: waiting is only safe because a retirement latch clears, and the call that reopens an abandoned fold can itself fail. When it does, the latch is held until the leaf deactivates - far beyond any sane value here - and the write fails loudly. That is the correct outcome, but "the latch is guaranteed to clear" is a conditional guarantee, not an unconditional one.
// Fail fast instead of waiting out a fold: useful in a test suite, where a
// blocked write should surface immediately rather than retry for seconds.
siloBuilder.ConfigureLattice("latency-sensitive-tree", o => o.LeafRetirementRetryDeadline = TimeSpan.FromMilliseconds(500));
Lower this only to make a test suite fail fast; raising it trades caller latency under a slow fold for fewer surfaced faults. This option can be changed freely at any time.
LeafSnapshotBinaryEncodingEnabled
Whether a leaf snapshot capture persists its rows as a compact binary frame rather than as the legacy object graph (default: true).
The legacy shape persists each row as an object, which under the default JSON grain-storage serializer costs a property-name envelope and a base64 string for every row. The frame collapses that to one length-prefixed buffer of raw value bytes and lets the decode path run without materialising a key string, a base64 buffer, or a scratch array per row. Measured persisted-size reduction ranges from 64% on a leaf of many small rows to 3% on one packed with 4 KiB float32 payloads (where base64 of the payload dominates either way); the decode-allocation and decode-CPU saving is uniform across all shapes and is the larger lever for cold start.
This switch gates the write side only. Reading is always dual: a blob is sniffed for the frame's magic prefix on load, so a legacy blob is decoded exactly as it always was. Adoption is therefore by lazy rewrite, never by migration - a leaf whose durable blob is still legacy simply persists a frame the next time it captures naturally. Nothing rewrites a blob eagerly and nothing discards one.
Cost: none in steady state. The encode is one buffer allocation per capture.
When off: captures persist the legacy row graph again, byte for byte. Blobs already written as frames stay readable, because the read path does not consult this switch.
When an operator would turn it off - and this is the case that matters. Set it to false before rolling back to a build that predates the frame. Such a build has no dual-read and would see a frame-carrying blob as an empty row set, which is data loss rather than a slow start, because the coverage-gated WAL garbage collector has been authorised to trim the prefix that snapshot covers. Turning the switch off pins new captures to the legacy shape so the rollback target can read them; leave it off until every leaf has captured at least once, then roll back.
LeafSnapshotMaxCoverageLagSeconds
Upper bound, in seconds, on how long an active leaf may leave its durable snapshot coverage lagging behind its projection checkpoint (default: 300; 0 disables). A leaf that has been active this long with any partition's checkpoint ahead of the coverage its durable snapshot records drives a capture, closing the gap. The same check also publishes the leaf's durable pin when it has fallen below min(persisted checkpoint, coverage), without taking a replay permit, so a write-idle leaf does not hold the WAL GC offset floor at a stale pin (issue #3599). It exists for the leaf that serves only reads: every other capture driver is activation-scoped or write-driven, and reads keep the grain from deactivating, so without this bound its coverage - and with it the WAL GC offset floor for the whole tree - could lag without limit. The validator accepts 0 to 86400.
LeafSnapshotSegmentBytes
Largest encoded snapshot frame, in bytes, persisted as a single BLOB column (default: 4 MiB). A capture whose frame exceeds it is split into row-aligned segments of at most this size, each in its own leaf-snapshot-segment grain-state row, and hydration decodes one segment at a time. This bounds the contiguous allocation on the hydration read path - the storage provider materialises a BLOB column as one array before lattice code runs - rather than the total bytes a leaf holds. The default sits above the Large Object Heap threshold, so ordinary leaves are never segmented. Values below 64 KiB are clamped up rather than rejected. A leaf plans a large capture's segments against its tree's resolved value and stages them itself, but a capture the leaf hands over whole is split by the snapshot store, which is addressed by leaf and carries no tree name and so reads this value from the default (unnamed) options. A per-tree override therefore governs only the leaf-side plan; set it globally so both sides use one window. See Tree storage.
MaintainProjectionDigest
Controls whether each leaf maintains the per-mutation XOR fold and publishes its digest upward to its internal-node ancestors after writing (default: true). On the per-write hot path the writes that land within one DigestCoalescingWindowMs share a single publish.
When true (the default), ILattice.GetLeafProjectionDigestAsync returns a pre-folded O(1) shard aggregate that operators and chaos tests can poll to detect cross-silo drift. Each leaf mutation costs one in-memory XOR over the entry's contribution. A leaf publishes upward only when its digest has changed, but every internal ancestor that folds a publish rewrites its persisted subtree hash and publishes on to its own parent whether or not that hash changed, so each publish costs O(treeHeight) writes. With DigestCoalescingWindowMs at 0 every digest-changing mutation publishes; otherwise the foreground writes within one window share a publish, while merges and structural events publish immediately (see DigestCoalescingWindowMs).
When false, leaf mutations take a trimmed path: they LWW-merge the value and bump the delivery sequence but skip both the XOR fold and the upward publication. The persisted ProjectionHash is left untouched. ILattice.GetLeafProjectionDigestAsync then fast-fails with InvalidOperationException at the public surface rather than returning a stale aggregate. Recommended for write-amplification-sensitive deployments that rely on audit logs or external reconciliation for cross-silo state-equivalence and do not poll the digest.
// Global opt-out:
siloBuilder.ConfigureLattice(opts => opts.MaintainProjectionDigest = false);
// Or per-tree:
siloBuilder.ConfigureLattice("audited-tree", opts =>
{
opts.MaintainProjectionDigest = false;
});
Disabling is a one-way operation per tree. The first mutation that lands while maintenance is disabled stamps an irreversible latch on the tree's registry entry (reported as TreeConfigurationReport.ProjectionDigestPermanentlyDisabled by the tree-admin configuration read). Once the latch is set, every subsequent activation resolves MaintainProjectionDigest as false regardless of the per-tree override or the silo-wide default, and ILattice.GetLeafProjectionDigestAsync keeps throwing. The latch exists because the digest is an XOR-fold aggregate: any mutation accepted while maintenance was off permanently invalidates the persisted aggregate, and silently re-engaging maintenance would publish a known-stale digest as if it were authoritative. The only way to re-engage digest maintenance for a latched tree is to rebuild it (or its leaf range) from scratch under a fresh registry entry.
System trees (those whose id begins with _lattice_) are always resolved as false regardless of configuration, because system trees are not replicated and have no cross-silo drift-detection consumer.
Per-tree precedence. A runtime override persisted on the tree's registry entry - set through ILatticeTreeAdmin.SetTreeConfigAsync with TreeConfigurationUpdate.ApplyMaintainProjectionDigest - overrides the configured value, whether that comes from a per-tree ConfigureLattice override or the silo-wide default; the latch overrides all of them.
See Projection Rebuild for the cost model and the WAL-storage rationale.
MaterialiserCheckpointEntries
Entry-count threshold above which a pending materialiser checkpoint is force-flushed to durable storage even if MaterialiserCheckpointInterval has not elapsed (default: 5 000). Together with MaterialiserCheckpointInterval this bounds replay cost on a worst-case crash: at most MaterialiserCheckpointEntries mutations have to be replayed against the projection on activation. It must be at least 1 (enforced by the options validator).
This option can be changed freely at any time.