Table of Contents

Options Reference: MaterialiserCheckpointInterval to ShardHealingCooldown

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at options-reference-2.md, and llms.txt lists every page.

Part of Options Reference, in Configuration.

MaterialiserCheckpointInterval

How long the leaf-projection materialiser may defer persisting an advancing checkpoint offset before flushing it to durable storage (default: 5 seconds). Combined with MaterialiserCheckpointEntries, this controls coalescing of materialiser-side high-water-mark writes: the checkpoint is persisted as soon as either threshold is met. Set to TimeSpan.Zero to persist on every advance (every-entry mode - strict RTO at the cost of one extra storage write per commit). Set to Timeout.InfiniteTimeSpan to disable time-based flushing and rely solely on the entry-count threshold. Any other negative value is rejected by options validation.

A graceful deactivation always force-flushes a pending checkpoint, so a clean silo shutdown loses no progress regardless of interval. A worst-case crash loses up to MaterialiserCheckpointInterval x steady-state apply rate of replay work on restart.

The interval is also re-checked on each leaf's coverage-lag tick (LeafSnapshotMaxCoverageLagSeconds), so a pending advance on a leaf that receives no further advances is still persisted once the interval has elapsed, at the next tick, rather than held until the leaf deactivates.

// Strict RTO: checkpoint on every advance.
siloBuilder.ConfigureLattice("strict-tree", o => o.MaterialiserCheckpointInterval = TimeSpan.Zero);

This option can be changed freely at any time.

MaxAtomicActionArgsBytes

The maximum size, in bytes, of a single custom step's argument payload within an atomic-action plan (default: 32 KiB). A step whose payload exceeds the bound is rejected before the saga starts, so a wire- or storage-supplied payload cannot bloat persisted saga state without bound. Must be positive; the options validator does not check it. It is read from the default (unnamed) options, so per-tree overrides do not apply. See Atomic Actions.

MaxAtomicActionSteps

The maximum number of steps an atomic-action plan submitted to IAtomicActionGrain.ExecuteAsync may contain (default: 64). A plan exceeding the bound is rejected before the saga starts, so a pathological plan cannot pin an activation for an unbounded time. Must be positive; the options validator does not check it. It is read from the default (unnamed) options, so per-tree overrides do not apply. See Atomic Actions.

MaxCacheValueBytes

Optional upper bound, in bytes, on the resident value-payload memory a single LeafCacheGrain activation may hold in its read-through mirror. null (the default) leaves the mirror unbounded - it grows to a faithful 1:1 copy of the primary leaf's live entry set, which is the lowest-latency configuration but scales per-silo per-tree memory linearly with the touched-leaf entry count.

When set to a positive value, the cache evicts value payloads only (never whole rows) in least-recently-used order once the sum of resident byte[] payload lengths would exceed the budget. The per-row metadata envelope (timestamp, delivery-sequence position, tombstone / migration flags, expiry) is always retained, so eviction cannot violate the cursor-based delta-refresh, pending-key, moved-away, or migrated-entry contracts described in Read Caching. A value read that lands on an evicted payload transparently delegates to the primary leaf for the authoritative bytes (one RPC) and is counted as a cache miss; existence checks are answered from the retained metadata with no RPC; hot keys stay resident and continue to serve from memory. Only the value payload is bounded - the retained envelope metadata (tens of bytes per row) is not counted against this budget, so plan for a small fixed overhead per live key on top of the configured budget.

// Cap each cache activation at 256 MiB of resident value payloads.
siloBuilder.ConfigureLattice(o => o.MaxCacheValueBytes = 256L * 1024 * 1024);

// Per-tree: leave a low-cardinality hot tree unbounded (default) but
// bound a large cold-scan tree so a full sweep cannot pin its whole
// value set in every silo's cache.
siloBuilder.ConfigureLattice("cold-archive", o => o.MaxCacheValueBytes = 32L * 1024 * 1024);

Intended as deploy-time configuration; the budget is re-read on each cache refresh so a running silo honours option changes, but toggling it on a warm activation only bounds payloads merged after the change. A null budget preserves the original unbounded behaviour. When set it must be at least 1 (enforced by the options validator).

MaxConcurrentAutoSplits

Maximum number of in-flight adaptive splits per tree (default: 2). Because HotShardMonitorGrain is keyed per tree, this limit is enforced independently per tree in a multi-tree cluster.

This option can be changed freely at any time.

MaxClusterConcurrentAutoSplits

Optional cluster-wide ceiling on the total number of autonomic splits that may be in flight concurrently across all trees (default: null - disabled). Because HotShardMonitorGrain is keyed per tree, MaxConcurrentAutoSplits only bounds one tree's splits; in a multi-tenant or many-tree cluster the summed drain I/O from every tree splitting at once can saturate the storage provider even though no single tree exceeds its own cap. Set a positive value to opt in to a singleton admission gate that caps the aggregate concurrent split count; a value below 1 is rejected by options validation.

The cluster ceiling is enforced in addition to each tree's MaxConcurrentAutoSplits and can only ever lower the number of splits a tree triggers, never raise it. When left at its null default the gate makes no admission decision, so splitting behaves exactly as it would without the option. The monitor still reports its tree's in-flight split count to the cluster gate while splits are in flight, plus one final call to clear that count once they finish, because the gate is also the cluster's split-activity source - the one ILatticeAdmin.GetSplitActivityAsync reads; a tree with no split in flight issues no extra RPC. Admission uses a per-tree heartbeat model: each monitor re-reports its tree's authoritative in-flight split count every pass and the gate expires any footprint that stops being refreshed, so a silo that crashes mid-split has its share of the ceiling reclaimed at expiry instead of wedging splitting cluster-wide.

Per-group tuning composes naturally: low-traffic tree groups clamp their own MaxConcurrentAutoSplits down through named options, while a single global MaxClusterConcurrentAutoSplits bounds the aggregate.

// Opt in to a cluster-wide ceiling of 4 concurrent autonomic splits,
// regardless of how many trees are hot at once.
siloBuilder.ConfigureLattice(o => o.MaxClusterConcurrentAutoSplits = 4);

// A low-traffic tree group additionally clamps its own per-tree cap to 1;
// all such trees still share the single global ceiling above.
siloBuilder.ConfigureLattice("cold-archive", o => o.MaxConcurrentAutoSplits = 1);

Watch orleans.lattice.split.in_flight (summed across the tree tag) to size the ceiling, and orleans.lattice.split.admission.deferred{reason=cluster_cap} to see whether it is binding (the counter's other reasons fire whether or not the gate is on). This option can be changed freely at any time.

MaxConcurrentDrains

Maximum number of per-shard drains an online snapshot (ILattice.SnapshotAsync in SnapshotMode.Online) may run concurrently (default: 4). Each drain reads one source shard's leaf chain and bulk-loads it into the corresponding destination shard while live writes keep mirroring onto the destination through shadow-forwarding. Higher values shorten the snapshot at the cost of proportionally more background drain I/O and coordinator memory; values below 1 are treated as 1. The snapshot stays crash-safe and idempotent under any cap, because re-running it converges by last-writer-wins. An online resize drains its snapshot at this concurrency inside each wall-clock-bounded slice (see BackgroundDrainMaxDuration), banking every in-flight shard's resume key when the slice ends.

This option can be changed freely at any time.

MaxConcurrentMigrations

Maximum number of shard splits (grow) or shard consolidations (shrink) an online reshard (ILattice.ReshardAsync) may drive concurrently (default: 4). Each split drains one physical shard's upper-half virtual slots into a newly allocated target shard, and each fold drains one shard into an adjacent survivor; running several in parallel shortens the reshard at the cost of proportionally more background drain I/O. Values below 1 are treated as 1. Splits driven by a reshard are independent of autonomic splits, so this cap and MaxConcurrentAutoSplits compose additively. A shrink's folds are bounded by this option, not by MaxConcurrentShardConsolidations, which governs automatic healing only.

This option can be changed freely at any time.

MaxConcurrentShardConsolidations

How many online shard consolidations (folds) may be in flight against one tree at a time (default: 1).

Automatic over-split healing admits at most one new fold per sweep and never more than this many concurrently. The cadence, not a burst, is the rate limiter: a badly over-split tree heals steadily rather than converting its damage into a thundering herd of concurrent drains.

Cost: each admitted fold drains a donor shard's entries into its survivor in bounded background passes. Raising the cap multiplies that background traffic.

When off: 0 is legal and admits nothing (negative values are rejected by options validation). It pauses admission while leaving the observer running, so the tree keeps publishing its healing backlog and an operator can watch the damage without acting on it. That is a different question from ShardHealingEnabled, which switches the mechanism off: no fold is admitted and no shard is polled, and a tree whose healing never started arms no reminder or timer at all.

When an operator would change it: raise it to drain a large backlog faster on a box with spare I/O; set it to 0 to freeze healing while keeping the backlog measurement.

MaxConcurrentSnapshotBaselineFolds

Maximum number of per-leaf WAL tail folds that a single shard's baseline capture may have in flight at once (default: 4). Applies inside a shard root's baseline capture, the per-shard step of a snapshot-isolated cursor open.

The capture runs in two passes over the shard's leaf chain. The first freezes each leaf's committed projection and must stay sequential, because the chain is discovered one sibling hop at a time. The second folds each frozen leaf's (leaf_frontier, capturedHead] WAL tail, and is the dominant cost on a shard whose leaves carry a deep tail. That second pass is safe to overlap: point-in-time consistency comes from the uniform capturedHead dominating every leaf's frozen frontier, not from the order the folds run in, so folding several leaves at once cannot change what is captured.

The captured baseline is byte-identical under any value, including 1. Results are consumed in strict leaf-chain order regardless of the order the folds complete, so a key present on more than one leaf resolves exactly as it would under a serial fold. Only the dispatch schedule changes.

The bound is a sliding window over unconsumed results, not merely over in-flight calls: a leaf's slot is re-dispatched only once its rows have been merged into the union, so a slow fold early in the chain cannot let every later fold complete and pile its rows in memory. Raising the value shortens the shard's non-reentrant hold on a deep chain at the cost of that much more concurrent fold memory and that many more simultaneous calls into the leaves; lowering it to 1 restores a strictly serial fold. Values below 1 are clamped to 1.

This knob is a different dimension from MaxConcurrentSnapshotCaptures, which bounds how many shards capture at once. The two multiply into the peak concurrent leaf folds one snapshot open can dispatch (16 at both defaults), so consider them together when tuning an open on a wide tree.

This option can be changed freely at any time; a new value applies to the next baseline capture.

// A tree whose leaves carry deep WAL tails: widen the per-shard fold
// window but narrow the shard fan-out, holding the peak at 4 x 4 = 16
// concurrent folds while shortening each individual shard's hold.
siloBuilder.ConfigureLattice("deep-tail", o =>
{
    o.MaxConcurrentSnapshotBaselineFolds = 8;
    o.MaxConcurrentSnapshotCaptures = 2;
});

// Restore a strictly serial fold on a memory-constrained silo.
siloBuilder.ConfigureLattice(o => o.MaxConcurrentSnapshotBaselineFolds = 1);

MaxConcurrentSnapshotCaptures

Maximum number of shard roots that opening a snapshot-isolated (point-in-time) cursor may block on their per-shard baseline capture at once (default: 4). Opening such a cursor freezes a baseline on every physical shard root; each capture walks that shard's whole leaf chain and materialises its rows on the shard root's non-reentrant turn, so fanning the capture out to every shard simultaneously blocks every shard root at once - starving cross-cluster replication applies and reads queued on those same roots. Bounding the fan-out keeps all but this many shard roots free while the open proceeds in waves. Lower values reduce the per-open blast radius at the cost of a longer open; higher values open faster but block more shard roots at once. The captured baseline and its point-in-time consistency are identical under any cap - only the dispatch schedule changes. Values below 1 are clamped to 1.

This option can be changed freely at any time; a new value applies to the next snapshot-cursor open.

MaxConcurrentStorageUsageTrees

Maximum number of trees a cluster-wide storage-usage roll-up samples concurrently (default: 8). Applies to ILatticeAdmin.GetTotalStorageUsageAsync, ILatticeAdmin.RefreshStorageUsageAsync, and the background poller's ILatticeAdmin.PollWalUsageAsync.

The roll-up is a two-level fan-out and the levels multiply: every tree sampled concurrently fans out again to its own shard roots and WAL partitions, bounded by MaxConcurrentStorageUsageSurfaces. Left unbounded, a cluster of 90 trees at the default 64 shards and 8 WAL partitions dispatches roughly 90 x (64 + 8) = 6,480 grain calls in a single burst that all race one 30 s Orleans response deadline, so the roll-up fails wholesale with response timeouts instead of merely taking longer. Bounding both levels caps the peak at MaxConcurrentStorageUsageTrees x MaxConcurrentStorageUsageSurfaces (128 by default) and makes the roll-up degrade in latency instead.

Raising it shortens a roll-up on a large, healthy cluster; lowering it further reduces the burst a roll-up imposes on silos serving live traffic. The aggregated figures are identical under any bound - only the dispatch schedule changes - and the per-tree ordering in ClusterStorageUsageReport.Trees follows the registry's sort order regardless. Values below 1 are clamped to 1.

This is a cluster-wide knob read from the default (unnamed) options by the admin grain that drives the roll-up; per-tree overrides do not apply, because the grain is not keyed by tree. It can be changed freely at any time; a new value applies to the next roll-up.

MaxConcurrentStorageUsageSurfaces

Maximum number of per-tree storage surfaces - shard roots plus WAL partitions - that a single tree's storage-usage aggregator queries concurrently (default: 16). Applies to ILattice.GetStorageUsageAsync and every path that reaches it, including the cluster roll-up.

The bound spans both surface kinds jointly, so a tree never has more than this many usage reads outstanding regardless of how its shard count and WalPartitions divide. A wide tree (the default shard count is 64) would otherwise dispatch every shard-root read at once even for a single-tree report. This is the inner level of the two-level fan-out described under MaxConcurrentStorageUsageTrees.

The report is byte-for-byte identical under any bound - only the dispatch schedule changes. Values below 1 are clamped to 1.

This option can be changed freely at any time; a new value applies to the next storage-usage fan-out.

// Halve the cluster-wide roll-up burst on a silo that also serves
// latency-sensitive traffic: 4 x 8 = 32 concurrent calls at peak.
siloBuilder.ConfigureLattice(o =>
{
    o.MaxConcurrentStorageUsageTrees = 4;
    o.MaxConcurrentStorageUsageSurfaces = 8;
});

// A single very wide tree can narrow its own surface fan-out further
// without changing the cluster-wide roll-up bound.
siloBuilder.ConfigureLattice("wide-archive", o => o.MaxConcurrentStorageUsageSurfaces = 4);

ShedSnapshotOpensWhenSaturated

Whether opening a snapshot-isolated (point-in-time) cursor is shed fast with a retryable LatticeSaturatedException when the tree's per-silo WAL saturation signal reports Saturated at the moment of the open, before the per-shard baseline capture is fanned out (default: true). A snapshot open freezes and materialises every shard's leaf chain on the non-reentrant shard roots - heavier than a single write - so admitting one into an already-saturated tree piles that work onto roots collapsing under write back-pressure, starving replication applies and reads queued on the same roots, and a client that retries on the resulting timeout sustains a scan storm. With the option enabled the open is refused at admission: the caller receives a typed, retryable back-pressure error and the fan-out never starts. Only Saturated (the "pause new appends" regime) sheds; a Throttled tree is unaffected and stays browsable, mirroring the atomic-write saga's quiesce gate. The open reads the signal under the id the tree was addressed by, while the signal is sampled under the id of the write-ahead log the tree's writes land in, so on an aliased tree - after a resize, a shadow-cutover restore, a schema remediation or an operator-set alias, whose log belongs to the physical copy - the open never sees Saturated and is never shed (see Resolution and scope). Over the state-API / Explorer surface the refusal is mapped to gRPC ResourceExhausted, and the Explorer shows a plain-language message that the table is very busy and to try again in a few seconds, rather than the raw fault. Set to false to restore the prior behaviour where a snapshot open always proceeds regardless of the saturation regime. See Snapshot Cursors.

This option can be changed freely at any time; a new value applies to the next snapshot-cursor open.

MaxCursorSnapshotPinTtl

Hard upper bound on how long the per-tree transaction registry will retain the saga-decision snapshot captured by a point-in-time durable cursor (default: 7 days). A live point-in-time cursor slides this TTL on every Next*Async; a stalled cursor that misses the slide will eventually have its pin reaped by the registry, after which the next call surfaces LatticeCursorSnapshotExpiredException and the cursor must be reopened.

The cap exists so a forgotten point-in-time cursor cannot stall registry-tombstone pruning forever. A cap shorter than TxDecisionRetention is floored to it, because the registry's own tombstone retention already covers anything shorter. Any non-positive value (for example Timeout.InfiniteTimeSpan) disables the registry-side cap entirely: the pin then never expires on its own and is released only when its cursor closes - through CloseCursorAsync, or when CursorIdleTtl reaps an idle cursor - and it counts toward MaxPinnedSagaDecisions until then. See Durable Cursors - Point-in-time cursors.

This option can be changed freely at any time.

MaxDurableUnresolvedReplayWork

Maximum number of unresolved replay-work records a leaf carries in its durable state so that its incremental flush ceiling may advance past them (default: 1 024). A leaf's flush ceiling is clamped below every unresolved saga prepare and every undrained deferred terminal, because neither survives an activation teardown in memory. Recording that work durably removes the need to re-read it, which is what lets a WAL partition that never wins the single first-pass drain slot bank forward progress instead of replaying the identical range on every activation. Records are struck off as the work resolves, so in the steady state the list is empty and this bound is never approached.

The bound applies to deferred terminals only. Dropping a deferred terminal at the bound is safe because the second replay pass re-reads and drains it, so the fall-back to the older clamping behaviour is transient. The one exception is a deferred terminal at the head of its partition's replay window, whose refusal would freeze that partition's checkpoint rather than slow it: it is admitted past the bound, and at most one offer per partition per activation can qualify. A resident unresolved prepare is recorded unconditionally and is never dropped at the bound: nothing drains a prepare whose saga never terminates, so dropping one would pin the flush ceiling one below the prepare offset permanently and the leaf would bank no durable forward progress at all. Past the bound the persisted row is therefore allowed to grow; rather than capping, every prepare recorded that leaves the ledger at or beyond the bound - deferred terminals and prepares counted together - increments the orleans.lattice.leaf.unresolved_prepare_ledger_beyond_cap counter, and the first in each activation also logs a warning. Persist risk: Azure Table grain storage rejects writes above its ~960 KB grain-state limit (see Storage provider per-row limits), bounding persisted row growth. Read risk: the larger SQLite limit on the repository-context host's default local durability profile permits growth that can exhaust memory or the read budget during activation, before grain-level repair can run. A successful persist is not proof of a safe activation read. Alert on that counter on every profile. See Metrics.

Setting this to 0 disables the mechanism entirely, restoring the behaviour in which the in-memory clamp is the only thing keeping unresolved work alive across a teardown.

This option can be changed freely at any time.

MaxEstimatedBytes

Optional enforcing cap, in bytes, on a tree's estimated storage footprint: TreeStorageUsageReport.TotalBytes, which sums the WAL's on-disk occupancy (physical bytes, dead bytes not yet compacted included, or the retained payload for a provider that cannot report a physical size), snapshot blobs and leaf/shard-root state. The orleans.lattice.storage.total_bytes gauge reports that figure after a deep storage-usage refresh, but the WAL-only poll of StorageUsagePollInterval replaces its WAL term with the retained payload, so between deep refreshes the gauge leaves out the WAL's dead bytes, which the cap counts. null (the default) leaves estimated bytes unbounded; enforcement is strictly opt-in. When set it must be at least 1 (enforced by the options validator). Once the tree's cached estimated-byte footprint reaches the cap, a locally-authored call to SetAsync (and its TTL overload), SetIfVersionAsync, GetOrSetAsync, SetManyAsync, SetManyWherePredicateAsync, SetManyAtomicAsync (every overload, except a delete-only batch, which can only shrink the tree), SetManyAtomicWhereAsync (both overloads), ApplyCrdtDeltaAsync (both overloads, which the typed CRDT accessors call) or ApplyCrdtDeltaManyAsync is rejected with a LatticeQuotaExceededException carrying the bytes dimension; an atomic batch is refused before its saga starts. The cross-tree SetManyAtomicAsync extension checks it too, for every participating tree whose slice carries an upsert, before any tree is staged. The cap is checked only by those calls, once per call, so a batch is admitted or refused as a whole.

The cap is best-effort and approximate: it is evaluated against a cached, eventually-consistent per-tree aggregate (the same TTL-coalesced aggregator that backs the storage-usage gauges), never a per-write fan-out, so every call admitted before a refreshed sample lands - including the whole of a batch admitted on one check - can carry the tree past the cap, and each new activation of the tree's grain (a stateless worker, so a silo can hold several) fails open (accepts writes) until its own first sample lands. The cross-tree extension reads that cached aggregate when it admits the batch, and fails open only when the aggregate cannot be read. Replication and atomic-write-saga apply paths bypass the cap, so an incoming replicated write is never rejected. Resolvable per tree.

// Cap the "bulk-ingest" tree at 1 GiB of estimated storage footprint.
siloBuilder.ConfigureLattice("bulk-ingest", o => o.MaxEstimatedBytes = 1024L * 1024 * 1024);

Prefer the advisory-first workflow: dry-run with AdmissionAdvisoryBytes and watch orleans.lattice.admission.would_reject before promoting to this enforcing cap. See Metrics. This option can be changed freely at any time; a new value takes effect on the next checked call.

MaxKeyLength

Optional upper bound on the number of characters in a key (default: null, unbounded). When set, it is checked on entry by SetAsync (and its TTL overload), SetIfVersionAsync, GetOrSetAsync, SetManyAsync (every entry), SetManyWherePredicateAsync (every entry, before its predicate is evaluated), SetManyAtomicAsync (every upsert, all three overloads; a delete key passed to the upsert-and-delete overload is checked only when the saga writes it, so an over-long one fails the saga with an InvalidOperationException once compensation completes rather than being refused up front), SetManyAtomicWhereAsync (every entry, before its precondition is evaluated), ApplyCrdtDeltaAsync (both overloads, which the typed CRDT accessors call) and ApplyCrdtDeltaManyAsync (every entry): a call carrying a longer key is rejected with an ArgumentException before any shard work, so an atomic batch is refused before its saga starts. The cross-tree SetManyAtomicAsync extension, which the BeginAtomicWrite builder's CommitAsync also goes through, checks every entry of every participating tree, deletes included, against that tree's bound the same way, before any tree is staged. The bound is checked only by the calls named here. Leaving it null preserves the historical unbounded behaviour; when set it must be at least 1.

siloBuilder.ConfigureLattice(o => o.MaxKeyLength = 1024);

This option can be changed freely at any time. It is read on each checked call, so a new value takes effect on the next one.

MaxLeafBytes

Maximum live state size, in bytes, a single leaf may hold before it splits (default: 64 MiB). The size is the leaf's running total of UTF-8 key length plus stored value length per entry (a tombstone counts its key only). It complements the structural MaxLeafKeys bound: the key count bounds how many entries a leaf holds, and this bounds how large they may be in aggregate.

A key-count bound alone cannot keep a leaf snapshottable. A tree with large values can grow a multi-hundred-megabyte leaf while still holding fewer keys than MaxLeafKeys, so it never splits on count. Capturing that leaf's snapshot has to materialise its payload in one contiguous buffer, which fails under heap pressure, and a leaf that cannot capture holds its tree's WAL trim floor down, because the floor is a minimum across every leaf - so one oversized leaf is enough to stop the whole tree reclaiming WAL. The bound is therefore checked on the write path and again on the snapshot-capture path, which divides an over-bound leaf before its payload is materialised (splitting repeatedly, up to eight times per pass), so an already-oversized leaf is repaired even on a tree that has stopped taking writes.

A leaf holding a single entry larger than the bound is irreducible - a split pivots on a median key and cannot divide one entry - so it is left intact and reported on orleans.lattice.leaf.byte.overflow with outcome=irreducible rather than split into an empty donor forever. See Metrics.

When off: 0 disables the byte bound and restores pure key-count splitting. The options validator rejects a negative value.

This option can be changed freely at any time. A leaf reads it with the rest of its resolved options when it activates, so a new value takes effect on each leaf's next activation. Keep LeafHydrationResidentBytes well below it; that section describes the advisory the silo logs when it is not.

MaxLeafEntriesBeforeForcedCompaction

Maximum total entry count (live plus tombstones) on a single leaf before the leaf requests an out-of-cycle compaction pass for its shard (default: 0, disabled). It complements MinTombstoneRatioForCompaction: a small leaf at a high tombstone ratio is reaped through the ratio trigger, while a large leaf that has accumulated tombstones at a low ratio is reaped through this one. It only fires when the leaf actually holds at least one tombstone or expired entry, and the compaction grain enforces a per-shard CompactionTriggerCooldown so a hot leaf cannot monopolise the compactor.

This default was re-examined during the bounded-cold-start work and deliberately left disabled, for a reason a host is better placed than the library to accept: the compaction grain opens the trigger metric-tag scope only when this knob or MinTombstoneRatioForCompaction is non-default, so arming it cluster-wide would start tagging previously untagged per-leaf instruments and silently break any dashboard filtering on the empty trigger label.

Evaluating the trigger is cheap. The leaf evaluates it only after a point or range delete, and reads its entry and live counts in O(1) without materialising any rows, so it does not force a partially hydrated leaf (see LeafPartialHydrationEnabled) to hydrate in full.

Nothing is lost by leaving it off: the reminder-driven compaction pass still reaps tombstones on its regular cadence.

When an operator would turn it on: on a tree that churns - repeated re-writes, prunes, or expiries over the same key space - where waiting for the reminder cadence lets tombstones accumulate. Set it per tree rather than globally. The repository-context host does exactly this, arming it at 2000 entries for its churn trees and leaving its write-once payload tree at the library default.

See Tombstone Compaction for the full trigger model.

MaxLeafReplayEntries

Soft budget on the number of WAL entries a leaf expects to replay against its projection at activation time (default: 10 000). Exceeding it is not an error: when the WAL still covers every offset the leaf needs, the leaf replays anyway - the result is identical, just slower - and the overrun is reported as a warning plus the orleans.lattice.leaf.activation_replays_over_budget counter. The budget is counted per leaf, after the per-leaf key-range filter, and it never consults ProjectionRebuildPolicy; only a WAL trimmed past the leaf's checkpoint does. Persistent overruns mean the materialiser is checkpointing too slowly for the write rate - tune MaterialiserCheckpointInterval / MaterialiserCheckpointEntries, or raise this budget. It must be at least 1 (enforced by the options validator). See Projection Rebuild for the full trigger set.

siloBuilder.ConfigureLattice(o => o.MaxLeafReplayEntries = 100_000);

This option can be changed freely at any time. The new value takes effect on the next leaf activation.

MaxLiveKeys

Optional enforcing cap on the number of live (non-tombstone) keys a single tree may hold. null (the default) leaves the live-key count unbounded; enforcement is strictly opt-in. When set it must be at least 1 (enforced by the options validator). Once the tree's cached live-key count reaches the cap, a locally-authored call to one of the write calls listed under MaxEstimatedBytes is rejected with a LatticeQuotaExceededException carrying the keys dimension; as there, the cap is checked only by those calls, once per call.

Shares the best-effort / approximate, fail-open, replication-bypassing semantics of MaxEstimatedBytes: the cap is compared against the cached, eventually-consistent per-tree aggregate (never a per-write fan-out), so calls admitted before a refreshed sample lands (a whole batch on one check) can carry the tree past the cap, each new activation of the tree's grain accepts writes until its own first sample lands, and replicated / saga-applied writes are never rejected. A time-expired entry that compaction has not yet reaped counts as live until the next deep re-anchor, so this effect can only make the cap bite early, never late. Resolvable per tree.

// Cap the "sessions" tree at 5,000,000 live keys.
siloBuilder.ConfigureLattice("sessions", o => o.MaxLiveKeys = 5_000_000);

Prefer the advisory-first workflow: dry-run with AdmissionAdvisoryLiveKeys and watch orleans.lattice.admission.would_reject before promoting to this enforcing cap. See Metrics. This option can be changed freely at any time; a new value takes effect on the next checked call.

MaxPhysicalShardsPerTree

Absolute ceiling on how many physical shards autonomic growth may give one tree (default: 256). It bounds adaptive splitting only; an explicit ILattice.ReshardAsync is an operator decision and is not capped by it.

The ceiling is a backstop, not a tuning knob. Without one, a pathological admission sequence has no terminating condition. Note the deployment that motivated this work turned out not to have hit it: measurement on a restored copy found every tree, including the three the embedder writes, sitting at exactly the base 64 physical shards. The thousand-plus figure reported for those trees counts leaf grains - B+ tree nodes inside a shard - not shards, and a leaf count is not what adaptive split creates or what consolidation folds. The ceiling therefore guards against a failure mode that is real in principle but was not the one observed; the observed cold-start cost was a grain activation per leaf, which bounded hydration and the approximate index address directly.

Cost: none. It is one comparison per admission.

When off: 0 or any negative value means no ceiling, which is the pre-#1834 behaviour.

When an operator would change it: raise it for a genuinely enormous tree whose splits are all justified, after checking that the refusals really are hitting the ceiling - the orleans.lattice.split.admission.deferred counter carries a shard_ceiling reason tag. Turning it off entirely restores the unbounded growth this ceiling exists to prevent.

MaxPinnedSagaDecisions

Footprint cap on the number of saga decisions that may be pinned across all live point-in-time cursors, enforced by each saga decision registry shard of a tree (default: 100 000). With the default single registry shard (TxRegistryShardCount = 1) that is one cap for the whole tree; with more shards, each shard and the legacy registry holds its own. OpenKeyCursorAsync / OpenEntryCursorAsync, or their WherePredicate variants, opened with pointInTime: true consult the registry: if accepting the new snapshot would push the pinned-decision count past this cap on any registry shard the snapshot touches, the open call throws LatticeCursorRegistryPinExhaustedException. With a single registry shard no pin is installed; when the snapshot spans several, a shard that did accept its part of the pin keeps it until the pin's TTL (MaxCursorSnapshotPinTtl) lapses. Existing pinned cursors continue paging.

Sized for a tree carrying a steady-state in-flight-saga set in the low thousands plus a handful of overlapping long-running point-in-time cursors. Raise if a workload routinely opens many concurrent multi-day point-in-time cursors against a saga-heavy tree; lower if a single tree must keep registry footprint tightly bounded.

This option can be changed freely at any time.

MaxScanRetries

Maximum bounded-retry passes for CountAsync, CountPerShardAsync, GetManyAsync, ScanKeysAsync, and ScanEntriesAsync when the shard topology changes mid-operation (default: 3); GetManyAsync also spends a pass when a saga commits while its fan-out is in flight. If the topology keeps mutating after every reconciliation step, the operation throws InvalidOperationException rather than returning a silently incomplete result. Under the default split rate-limits (MaxConcurrentAutoSplits = 2, HotShardSplitCooldown = 2 minutes), exhausting 3 retries is not a realistic operational concern. See Scan reliability.

This option can be changed freely at any time.

MaxLeavesPerScanPage

Maximum number of leaves one shard range-scan page fill may visit before returning a partial page with HasMore set (default: 64).

A page fill used to be bounded only by its output - it walked the sibling chain until the page was full. A leaf can cost a grain call (plus a snapshot rehydration or WAL replay when cold) and still contribute nothing to the page, because its entries were filtered as moved-away by an adaptive split, tombstoned and still inside TombstoneGracePeriod, TTL-expired, or rejected by a pushed-down predicate. That made the worst case O(leaves in range) per call, and because shard reads are deliberately non-reentrant such a call head-of-line-blocks every other request to that shard - point reads, writes, splits and health probes alike - for its whole duration.

Bounding the work turns one long call into several short ones, releasing the shard between them so other traffic interleaves. The bound applies whether or not the page has collected anything, because the run of leaves it exists to bound is exactly the run that contributes nothing. A page fill stops on it only where it can name a resume position - the key boundary of the leaf it stopped on - so a page can come back empty while still reporting more, and the scan continues from that position rather than from a returned key. Forward progress is preserved without disarming the bound.

Correctness is unaffected. Atomic visibility for a scan comes from the saga-decision snapshot pinned for the lifetime of each uninterrupted underlying enumeration, not from the granularity of an individual shard call, so every page of that enumeration - including an extra boundary introduced by this bound - observes the identical view. If the enumeration is interrupted and the scan transparently reconnects - after an enumeration abort, or to resume a stalled page fill - it resumes from the last key it yielded under a freshly captured saga-decision view. See Consistency.

Raising it trades shard fairness for fewer round trips. Setting it to 0 or less disables the bound and restores the historical unbounded walk. The default is deliberately well above what a dense scan needs, so a healthy tree never reaches it.

This option can be changed freely at any time.

MaxScanPageDuration

Wall-clock safety net for a single shard range-scan page fill (default: 5 seconds). The clock starts when the grain call begins, so it measures the whole time the shard is held - preparing the grain and traversing down to the start leaf included - rather than only the leaf loop. Once it expires the page fill stops at the first leaf boundary it can name a resume position for, returning a partial page rather than continuing to hold the non-reentrant shard.

MaxLeavesPerScanPage is the primary, deterministic bound; this covers the case a leaf count cannot, where a small number of leaves are individually very slow - typically cold activations rehydrating a large snapshot or replaying a WAL projection. It is sampled between leaf reads, so it bounds how many slow leaves one page fill chains together; it cannot preempt a leaf read already in flight. Set to TimeSpan.Zero to disable it and rely on the leaf count alone.

Because it is cooperative it is a graceful bound: it can only stop the walk somewhere the walk can name a resume position, which by construction means between leaf reads. MaxScanPageStallDuration is the hard ceiling that covers the two cases it structurally cannot.

This option can be changed freely at any time.

MaxScanPageStallDuration

Hard end-to-end ceiling on a single shard range-scan page fill (default: derived, see below). Unlike MaxScanPageDuration, which the walk samples cooperatively between leaf reads, this one bounds the whole grain call: when it elapses the call stops waiting, so the deliberately non-reentrant shard is released and its queue drains. Work the walk has already completed - the rows it has read, or a finished partial page for a walk that accumulates an aggregate - is returned as a short page the caller resumes from; only a fire that catches the walk with nothing to show faults with a ScanPageStalledException.

It exists because a cooperative budget can only stop a walk at a point the walk reaches. Two shapes never reach one:

  • the prologue or descent parks - preparing the shard, or traversing down to the start leaf, before the leaf loop is entered at all;
  • a single leaf read already in flight never returns, so the budget is never sampled again.

Either shape holds the shard for as long as the underlying call takes, which in the field has meant minutes: a page fill has been observed holding a shard for 576 seconds against a 5 second budget, head-of-line-blocking every point read, write, split and health probe on that shard for the whole hold.

The fault is retriable and loses no work. A page fill only reads a key range, so nothing is half-applied, and the caller resumes from the continuation token it already holds. Every ceiling fire is counted by orleans.lattice.shard_root.scan_page.ceiling_outcomes, tagged outcome = banked (a short page was returned) or discarded (the call faulted). The faulted fires are also counted by orleans.lattice.shard_root.scan_page.stalls, tagged with the phase (prologue, descent, or leaf-walk) that names which of the shapes above occurred, or baseline-fold for the fold pass of a snapshot cursor open's per-shard baseline capture (OpenSnapshotKeyCursorAsync / OpenSnapshotEntryCursorAsync), which runs under the same ceiling.

The default is derived, not fixed

The ceiling is only useful if it fires before the Orleans response timeout that governs the call. Past that deadline the caller has already given up, so it sees an anonymous Orleans timeout instead of the typed ScanPageStalledException - and, worse, nothing has released the shard, which is the whole point of the ceiling.

That is a constraint relative to another setting, not an absolute number, so the option is left null by default and derived:

effective ceiling = SiloMessagingOptions.ResponseTimeout - 5 seconds

floored so it never drops below MaxScanPageDuration. With Orleans' own 30 second default response timeout that yields 25 seconds. A deployment that tightens ResponseTimeout to 12 seconds gets a 7 second ceiling automatically, where a hardcoded default would have been dead configuration - never firing before the caller's deadline, and so never releasing the shard.

The derivation reads the local silo's response timeout, which is the deadline that applies to a silo-to-silo page fill. An external client configured with a different ClientMessagingOptions.ResponseTimeout is not visible from the silo; a cluster that sets the two differently should configure this option explicitly.

Set it explicitly to override the derivation:

siloBuilder.ConfigureLattice(o => o.MaxScanPageStallDuration = TimeSpan.FromSeconds(45));

Set it to Timeout.InfiniteTimeSpan to disable the ceiling and restore the historical behaviour where a stalled page fill holds the shard until the Orleans response deadline expires. When set explicitly it must be strictly greater than MaxScanPageDuration, so the graceful bound always gets the chance to return a partial page before the hard ceiling faults the call.

This option can be changed freely at any time. It is armed per call, so a new value takes effect on the next page fill.

MaxValueSizeBytes

Optional upper bound, in bytes, on the size of a value or CRDT delta (default: null, unbounded). It is checked by the same calls, at the same points, as MaxKeyLength: on entry by SetAsync (and its TTL overload), SetIfVersionAsync, GetOrSetAsync, SetManyAsync (every entry), SetManyWherePredicateAsync (every entry), SetManyAtomicAsync (every upsert), SetManyAtomicWhereAsync (every entry), ApplyCrdtDeltaAsync (both overloads, measuring the delta) and ApplyCrdtDeltaManyAsync (every entry's delta), where a larger value or delta is rejected with an ArgumentException before any shard work. The cross-tree SetManyAtomicAsync extension, which the BeginAtomicWrite builder's CommitAsync also goes through, checks each entry's value - for a write staged with Set(LatticeStagedCrdtWrite), the merged CRDT state it writes rather than its delta - against its tree's bound the same way, for every participating tree, before any tree is staged. The bound is checked only by the calls named here. Leaving it null preserves the historical unbounded behaviour; when set it must be at least 1.

siloBuilder.ConfigureLattice(o => o.MaxValueSizeBytes = 1024 * 1024);

This option can be changed freely at any time. It is read on each checked call, so a new value takes effect on the next one.

OptimisticShardRootPointReads

When enabled (the default), a point read (GetAsync) first tries an optimistic read that is allowed to interleave with other reads on the same shard root, rather than queueing behind them. Without it, each shard root serves one point read per full leaf round trip, which caps point-read throughput at roughly the shard count divided by the leaf round-trip time.

The optimistic read always validates the shard root's routing epoch. Routing-changing calls bump the root epoch at entry and exit; point Sets bracket only prepare, split-link/promotion and retired-leaf retry work. An uncontended present read uses the primary leaf's raw byte reply without computing, transporting or comparing an ownership stamp. A point-write admission epoch plus the in-flight count detect writes that start and finish during the await. Reads overlapping point Sets, and raw misses, instead require a versioned reply with a leaf ownership stamp. The stamp combines a fresh activation identity with a generation bumped around splits, seals, move-away, consolidation and retirement. The leaf returns it only when it owns the key's half-open range and can observe ownership and value together without awaiting. A read overlapping root routing changes or returning a missing/different required stamp retries serially, including splits below an internal node that leave the root epoch unchanged.

The optimistic read resolves its leaf only from routing tables the serial path has already cached, and reads the primary leaf grain directly rather than through the leaf cache, because a cache replica refreshed mid-split can briefly disagree with the shard root's routing. Serial reads also warm leaf ownership stamps. A matching stamp proves genuine absence as well as a present value; a null without that proof is always re-read serially. Old-wire leaves omit the additive ownership fields and therefore cannot validate when proof is required. Pending-transaction and shadowed-migration replies cannot supply ownership proof because their visibility may require an awaited registry check.

Non-splitting point Sets do not suppress optimistic reads, so present and absent reads can validate during continuous point writes. Other mutations remain conservatively bracketed for their full duration. Disable the option to restore the fully serial read path:

siloBuilder.ConfigureLattice(o => o.OptimisticShardRootPointReads = false);

This option can be changed freely at any time. It is resolved when a grain activation first reads the tree's options, so a new value takes effect as activations are recycled.

PrefetchEntriesScan

When enabled (default: false), ScanEntriesAsync pre-fetches the next page from each shard in the background while the current page is being consumed by the k-way merge. This hides per-shard grain-call latency and can significantly reduce wall-clock time for large scans across many shards.

// Enable globally
siloBuilder.ConfigureLattice(o => o.PrefetchEntriesScan = true);

Pre-fetch can also be controlled per-call via the prefetch parameter on ScanEntriesAsync, which overrides the global option:

// Override for a single call regardless of global setting
await foreach (var entry in tree.ScanEntriesAsync(prefetch: true))
{
    // ...
}

Because each pre-fetched page is held in memory until consumed, callers that abort iteration early (e.g. Take(n)) pay for pages they never read. For bounded scans, leave this disabled or pass prefetch: false explicitly.

This option can be changed freely at any time.

PrefetchKeysScan

When enabled (default: false), ScanKeysAsync pre-fetches the next page from each shard in the background while the current page is being consumed by the k-way merge. This hides per-shard grain-call latency and can significantly reduce wall-clock time for large scans across many shards.

// Enable globally
siloBuilder.ConfigureLattice(o => o.PrefetchKeysScan = true);

Pre-fetch can also be controlled per-call via the prefetch parameter on ScanKeysAsync, which overrides the global option:

// Override for a single call regardless of global setting
await foreach (var key in tree.ScanKeysAsync(prefetch: true))
{
    // ...
}

Because each pre-fetched page is held in memory until consumed, callers that abort iteration early (e.g. Take(n)) pay for pages they never read. For bounded scans, leave this disabled or pass prefetch: false explicitly.

This option can be changed freely at any time.

ProjectionRebuildPolicy

Selects the recovery strategy a leaf takes on genuine loss: when the WAL has been trimmed past the leaf's persisted projection checkpoint and no snapshot covers the gap (default: SnapshotThenWal). It is never consulted for the cost signals MaxLeafReplayEntries and LeafProjectionRetention, which indicate a long or stale replay rather than missing data; on those the leaf tail-replays and converges normally.

Value Behaviour
SnapshotThenWal Uses the per-leaf snapshot as the recovery base, then tail-replays the WAL since it. The snapshot-rehydrate half is live and runs at every activation before this policy is ever consulted. What is not integrated is a recovery after that rehydrate has declined, so on genuine loss the leaf currently surfaces LeafProjectionStaleException rather than rebuild over the lost prefix.
FullRebuildFromWal Diagnostic. Intended to replay from the absolute tail of the WAL, but the policy is consulted only when the WAL has been trimmed and a complete history is unavailable, so it fails closed with LeafProjectionStaleException like the other values; no full-rebuild recovery path is integrated.
Fail Surfaces LeafProjectionStaleException at activation and waits for an operator-driven rebuild.

Because the policy is consulted only when the missing prefix is genuinely gone, every value currently fails closed with LeafProjectionStaleException in that case: replaying only the surviving suffix would rebuild the leaf over the lost prefix and advance the materialiser pin past unrecoverable data. A value that is not a defined ProjectionRebuildPolicy member fails options validation.

This option can be changed freely at any time.

PublishEvents

When true, Lattice publishes LatticeTreeEvent notifications on the Orleans stream namespace orleans.lattice.events covering per-key writes, atomic-write completions, splits, compactions, snapshots, resizes, reshards, and tree-lifecycle transitions (default: false, opt-in per tree). Consumers subscribe via LatticeExtensions.SubscribeToEventsAsync. Publication is fire-and-forget and log-and-swallow, so a missing or misconfigured stream provider never breaks the write path. Per-tree overrides applied via ILattice.SetPublishEventsEnabledAsync are persisted on the tree's registry entry and override the silo-wide default. The one exception is TreePurged: a purge removes the tree's registry entry, and the override with it, before it publishes, so the tree's configured PublishEvents decides whether that event is published. See Events.

This option can be changed freely at any time. Per-tree overrides take effect on the publishing activation immediately; other activations refresh within a few seconds.

ShardForwardTimeout

Hard ceiling on how long a single outbound shard-to-shard write forward may run before it is cancelled and surfaced to callers as a TimeoutException (default: 15 seconds). It bounds both the online-resize shadow forward and the adaptive-split migration forward.

During a reshard swap the destination shard's ownership is changing, and Orleans can reject the outbound forward message and leave the caller-side await neither completing nor faulting. Without a ceiling the forwarding turn never returns, the lattice grain's per-shard fan-out saturates at its in-flight limit, and the whole write pipeline wedges with no fault and no activation recycle. With the ceiling the parked forward is abandoned and the turn faults cleanly with a TimeoutException, which the existing transient-exception retry envelope on every mutation path catches and re-runs against refreshed routing once the swap has settled. Abandoning a forward never loses data: convergence on the destination shard is independently guaranteed by last-writer-wins plus the split coordinator's authoritative leaf-chain drain (the Drain phase and the Complete-phase final drain).

Set to InfiniteTimeSpan to disable the ceiling and restore the historical unbounded-await behaviour; the options validator rejects any other non-positive value.

This option can be changed freely at any time. A shard root resolves its tree's options once per activation, so a new value takes effect on each shard root's next activation.

ShardHealingBackpressureOpsPerSecond

The tree's median shard rate at or above which automatic healing yields to foreground traffic (default: 200.0).

The median rather than the summed tree rate, deliberately: a sum scales with the shard count, so the thousand-shard tree that most needs healing would look like the busiest tree on the box and would never heal, exactly inverting the intent. Yielding is total - no new fold is admitted and no in-flight fold is driven - so a loaded tree costs no consolidation traffic at all.

The default deliberately equals HotShardOpsPerSecondThreshold. That separates the split and heal loops in the load domain as well as the skew domain: at any load where a split is even conceivable, healing has already yielded. If you retune either value, keep backpressure at or below the split threshold or you open a load band in which both control loops are live at once.

Cost: none; the median is already computed for the skew ratio.

When off: 0 is legal and heals regardless of load. NaN and negative values are rejected by options validation.

When an operator would change it: lower it on a latency-sensitive tree that should never compete with healing; set it to 0 on a tree that is idle by design and must heal promptly.

ShardHealingCooldown

How long automatic healing stands off a tree after a sweep observes a split in flight on it (default: 5 minutes).

This is the time-domain half of the split/heal hysteresis; the skew dead band is the space-domain half. It deliberately does not fire after each completed fold: the sweep interval and the concurrency cap already pace healing, and a per-fold cooldown would make a thousand-fold heal take weeks.

Cost: none.

When off: 0 is legal and disables the post-split stand-off, leaving only the skew dead band to prevent oscillation. Negative values are rejected by options validation.

When an operator would change it: lengthen it on a tree whose post-split load takes a long time to settle; set it to 0 only on a tree where splitting is disabled entirely, so there is no split for healing to stand off from.