LatticeOptions
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at latticeoptions.md, and llms.txt lists every page.
Part of Lattice Public API Reference.
See Configuration for detailed guidance, mutability constraints, and per-tree overrides via the tree registry.
Structural sizing is pinned per-tree in the registry, not in
LatticeOptions.MaxLeafKeys,MaxInternalChildren, andShardCountare seeded into the tree's registry entry on first tree use from the built-in defaults (128 / 128 / 64). After seeding they are mutable only throughILattice.ResizeAsyncandILattice.ReshardAsync. Callers who want non-default sizing should either callResizeAsync/ReshardAsyncon a freshly-created tree (empty-tree fast-path) or create the tree with explicit sizing through the tree-administration facade'sILatticeTreeAdmin.CreateTreeAsync(seeOrleans.Lattice.Api.TreeAdmin); a tree an installed app creates takes the sizing pins its manifest declares when the install first registers it (seeOrleans.Lattice.Apps).LatticeOptionscarries no sizing properties, soConfigureLattice("tree-name", ...)cannot seed them; see tree registry for how the pin is resolved.
The virtual shard space is not a runtime option. A tree routes over 4096 virtual slots, a compile-time default, unless its persisted
ShardMaprecords a different slot count - which only a tree created by an installed app whoseOrleans.Lattice.Appsmanifest declares avirtualShardCountdoes. Routing always hashes over the slot count the tree's map holds, a map cannot route to more physical shards than it has slots, and the default identity map assigns virtual slotito physical shardi % ShardCount.
The table below covers commonly tuned options; several option families
are documented next to the feature they tune (for example the
WalSaturation* knobs under WAL saturation back-pressure), and
Configuration is the complete reference.
| Property | Type | Default | Description |
|---|---|---|---|
KeysPageSize |
int |
512 | Keys per page in enumeration pagination. |
MaxKeyLength |
int? |
null (unbounded) |
Optional upper bound on a key's length in characters (string.Length). When set, it is checked on entry by SetAsync (both overloads), SetIfVersionAsync, GetOrSetAsync, SetManyAsync, SetManyWherePredicateAsync, ApplyCrdtDeltaAsync (both overloads), ApplyCrdtDeltaManyAsync, every upsert of SetManyAtomicAsync (all three overloads) and every entry of SetManyAtomicWhereAsync, and a longer key is rejected with ArgumentException before any shard work - an atomic batch before its saga starts. The cross-tree SetManyAtomicAsync extension checks every entry of every participating tree against that tree's bound, delete keys included, before any tree is staged. The delete keys of a mixed set-and-delete SetManyAtomicAsync batch are checked only when the saga writes the batch, after it has started: a longer key then aborts the whole batch and the call throws InvalidOperationException rather than ArgumentException. DeleteAsync, the range deletes and the bulk-load calls never check it. Must be at least 1 when set. See Configuration. |
MaxValueSizeBytes |
int? |
null (unbounded) |
Optional upper bound, in bytes, on a written value or CRDT delta, checked by the same calls at the same points as MaxKeyLength: a larger value is rejected with ArgumentException on entry - an atomic batch before its saga starts, and a cross-tree SetManyAtomicAsync batch before any tree is staged. ApplyCrdtDeltaAsync and ApplyCrdtDeltaManyAsync (and the typed accessors built on them) measure the delta; the cross-tree batch measures each entry's stored value, which for a staged CRDT write is the merged state it stores, not its delta. Must be at least 1 when set. For the per-tree MaxLiveKeys / MaxEstimatedBytes caps see Admission back-pressure. See Configuration. |
TombstoneGracePeriod |
TimeSpan |
24 h | Minimum age before a tombstone is eligible for compaction. InfiniteTimeSpan disables compaction. |
CompactionShardTickInterval |
TimeSpan |
500 ms | Gap between consecutive per-shard ticks during a compaction pass. Floor 100 ms; values below are clamped with a one-shot warning. Snapshotted at pass start. Scheduler-fairness knob, independent of leaf activation lifetime. See Tombstone Compaction. |
CompactionLeafBatchSize |
int |
64 | Maximum number of leaves the coordinator visits within a single shard before yielding for one CompactionShardTickInterval. The leaf walk resumes on the next timer tick from a persisted in-shard key cursor, so peak concurrent leaf activations during a pass are capped regardless of tree size. Also sizes the empty-leaf reclaim walk, at sixteen probes per foldable leaf, which is separately bounded in wall clock by BackgroundDrainMaxDuration. Floor 1; values below are clamped with a one-shot warning. Snapshotted at pass start. See Tombstone Compaction. |
SoftDeleteDuration |
TimeSpan |
72 h | Retention window after soft-delete before purge. |
CacheTtl |
TimeSpan |
TimeSpan.Zero |
Minimum time between read-cache refreshes. Zero means refresh on every read. |
PrefetchKeysScan |
bool |
false |
When true, ScanKeysAsync overlaps the next page fetch with consumption. Overridable per call via prefetch. |
PrefetchEntriesScan |
bool |
false |
When true, ScanEntriesAsync overlaps the next page fetch with consumption. Gated separately from PrefetchKeysScan because entry pages carry values. Overridable per call. |
AutoSplitEnabled |
bool |
true |
Master switch for autonomic shard splitting. |
HotShardOpsPerSecondThreshold |
int |
200 | Ops/sec on a single shard that triggers an adaptive split. |
HotShardSampleInterval |
TimeSpan |
30 s | How often the hot-shard monitor polls hotness counters. |
HotShardSplitCooldown |
TimeSpan |
2 min | Minimum time between consecutive splits of the same shard. |
MaxConcurrentAutoSplits |
int |
2 | Per-tree cap the hot-shard monitor fills with adaptive splits: a pass starts new splits only while the tree's in-flight shard migrations are below it, and every shard that is the source of an unfinished split or the donor of a consolidation fold counts, so an automatic healing fold in flight leaves one fewer slot for a new split. |
MaxConcurrentMigrations |
int |
4 | Maximum parallel shard splits (grow) or shard consolidations (shrink) an online reshard (ReshardAsync) drives concurrently. A grow counts every shard already migrating against it, including an adaptive split or an automatic healing fold that was running when the reshard began; a shrink counts only its own folds and starts none while a healing fold is still running. The hot-shard monitor and automatic shard healing start no new split or fold while a reshard runs. |
MaxConcurrentDrains |
int |
4 | Maximum parallel per-shard drains an online SnapshotAsync (SnapshotMode.Online) dispatches, while live writes keep mirroring to the destination by shadow forwarding. |
SplitDrainBatchSize |
int |
1024 | Entries per batch during the drain phase of a split. |
ShardForwardTimeout |
TimeSpan |
15 s | Hard ceiling on a single outbound shard-to-shard write forward (the online-resize shadow forward and the adaptive-split migration forward). A forward that exceeds it is cancelled and surfaced as a TimeoutException, which the normal stale-routing retry envelope re-runs against refreshed routing - preventing a forward parked against a shard whose ownership is changing during a reshard swap from pinning the foreground write turn and wedging the per-shard fan-out. InfiniteTimeSpan restores the historical unbounded await. |
ActivationReadyTimeout |
TimeSpan |
15 s | Hard ceiling on a shard root's one-time activation-readiness seed (the first-touch cross-grain awaits a brand-new or freshly-reactivated shard runs while holding its non-reentrant activation gate: the defensive state re-read, the tree-registry registration, the deterministic root-leaf init, and the initial shard-state write). A seed that exceeds it is abandoned and surfaced as a TimeoutException, which the normal transient-exception retry envelope re-runs against refreshed routing once the dependency recovers - preventing a registry or leaf RPC parked against a not-yet-visible activation during a startup reshard or membership change from pinning the gate and wedging every interleaved read/write on the shard. Each seed step is idempotent on retry, so abandoning a parked seed never loses data or double-registers. InfiniteTimeSpan restores the historical unbounded await. |
DigestPublishTimeout |
TimeSpan |
15 s | Hard ceiling on a single internal-node upward digest publish (the child-digest propagation an internal node sends to its parent after folding a child's digest, recursing up toward the shard root). The publish is sent only after the node has released its non-reentrant split gate, and a parent whose gate is busy parks the incoming snapshot for the gate holder to fold rather than waiting on the gate, but the upward await can still be left neither completing nor faulting. A parked publish is abandoned and faulted as a TimeoutException; the digest is staleness-tolerant, so the next mutation's publish re-drives convergence with no count drift. InfiniteTimeSpan restores the historical unbounded await. |
AutoSplitMinTreeAge |
TimeSpan |
60 s | Minimum tree age before the hot-shard monitor begins sampling. |
MaxScanRetries |
int |
3 | Maximum bounded-retry passes for CountAsync / CountPerShardAsync / GetManyAsync / ScanKeysAsync / ScanEntriesAsync when topology changes (or, for GetManyAsync, a saga commits) mid-call; exhaustion throws InvalidOperationException. See Scan reliability. |
CursorIdleTtl |
TimeSpan |
48 h | Sliding idle timeout for stateful cursors. InfiniteTimeSpan disables auto-cleanup. |
MaxCursorSnapshotPinTtl |
TimeSpan |
7 d | Hard cap on the registry-side lifetime of a point-in-time cursor's snapshot pin, re-armed on every step and never shorter than TxDecisionRetention. A non-positive value (Timeout.InfiniteTimeSpan, for example) disables the cap: the pin then does not expire and is released when the cursor closes. |
MaxPinnedSagaDecisions |
int |
100 000 | Cap on the saga decisions pinned across all live point-in-time cursors, enforced by each saga decision registry shard of a tree (one cap for the whole tree with the default single shard). See Point-in-time cursors. |
AtomicWriteRetention |
TimeSpan |
48 h | Retention window for completed SetManyAtomicAsync saga state (idempotency window). InfiniteTimeSpan disables auto-cleanup. |
TxDecisionRetention |
TimeSpan |
60 s | Retention window for a completed saga's commit/abort decision in the per-tree registry after ForgetAsync. TimeSpan.Zero restores legacy immediate-evict semantics. See Configuration. |
TxRegistryAdmissionBudgetBytes |
long? |
768 KiB | Row-size budget, per registry shard, at or above which the transaction registry refuses new atomic-write sagas with a retryable LatticeSaturatedException (TxRegistryCapacity). null disables the bound. See Configuration. |
TxRegistryShardCount |
int |
1 | Saga decision registry shards new sagas are minted across, per tree (1 to 256). Each shard has its own row and admission budget, so the per-tree saga ceiling scales with it. Opt-in: raise only once every silo and replication peer runs a sharding-aware version. Safe to change live in either direction. Global (unnamed options) only. See Configuration. |
VersionVectorRetention |
TimeSpan |
InfiniteTimeSpan |
Declared but currently has no effect: no code reads this option, so setting it prunes nothing. The pruning primitive it describes, VersionVector.PruneOlderThan(long minRetainedUtcTicks), has no production caller but can be called directly. See Configuration. |
DiagnosticsCacheTtl |
TimeSpan |
5 s | Cache lifetime for DiagnoseAsync reports. TimeSpan.Zero disables caching. |
MaterialiserCheckpointInterval |
TimeSpan |
5 s | Time-based threshold for flushing a leaf-projection checkpoint. |
MaterialiserCheckpointEntries |
int |
5 000 | Entry-count threshold for flushing a leaf-projection checkpoint. |
MaxLeafReplayEntries |
int |
10 000 | Advisory per-leaf budget on the WAL entries a cold leaf expects to replay at activation. Exceeding it is not an error: while the WAL still covers every offset the leaf needs it replays anyway and reports the overrun as a warning and the orleans.lattice.leaf.activation_replays_over_budget counter; ProjectionRebuildPolicy is not consulted. |
LeafProjectionRetention |
TimeSpan |
7 d | Advisory age beyond which a leaf's persisted projection checkpoint warrants a warning at activation; an old checkpoint still tail-replays and converges, and never routes to ProjectionRebuildPolicy. The activation path currently passes a zero age (the checkpoint carries no capture timestamp), so this trigger does not fire from activation today. InfiniteTimeSpan disables the age trigger. |
ProjectionRebuildPolicy |
enum | SnapshotThenWal |
Consulted only on genuine loss at activation - the WAL has been trimmed past a leaf's persisted checkpoint and no snapshot covers the gap - and every value currently surfaces LeafProjectionStaleException there. See Configuration. |
MaintainProjectionDigest |
bool |
true |
When false, GetLeafProjectionDigestAsync fast-fails with InvalidOperationException. Disabling is a one-way operation per tree - the first mutation under the disabled setting stamps an irreversible registry latch. System trees (_lattice_*) always resolve as false. |
PublishEvents |
bool |
false |
Opt-in publication of LatticeTreeEvent notifications onto an Orleans stream. See Events. |
EventStreamProviderName |
string |
"Default" |
Stream provider Lattice publishes events onto. |
WalPartitions |
int |
8 | Number of independent WAL partitions per tree. Pinned per-tree in the tree registry when the tree is first registered (on first use, before any data is written), from the value this option resolves to for that tree at that moment; changing the option later does not affect an existing tree. |
WalMaxBatchEntries |
int |
100 | Maximum WAL entries coalesced into a single flush. |
WalMaxBatchBytes |
long |
4 MiB | Maximum byte budget coalesced into a single flush. |
WalFlushTimeout |
TimeSpan |
15 s | Hard ceiling on a single WAL partition flush (the provider append plus the post-failure tail resync). A flush that exceeds it is cancelled and surfaced as a TimeoutException routed through the normal failure handler, preventing a hung provider call from pinning its in-flight slot and wedging the in-flight chain. InfiniteTimeSpan restores the historical unbounded await. |
WalFlushPreflightTimeout |
TimeSpan |
5 s | Hard ceiling on a WAL partition's FlushAsync preflight region (the synchronous setup and initial scheduler yield that precede the bounded provider call). If the activation's grain scheduler never resumes the post-yield continuation within the deadline, the slot would sit in _inFlight with no provider-call deadline armed (WalFlushTimeout only covers the provider call, which has not been issued yet). The faulted preflight surfaces as a TimeoutException routed through the normal failure handler, the slot drains, and the orleans.lattice.wal.flush.preflight.timeouts counter attributes the trip per (tree, shard). InfiniteTimeSpan restores the historical unbounded await. |
WalAppendDispatchTimeout |
TimeSpan |
30 s | Hard ceiling on a single writer-side outbound WAL shard append-batch / append dispatch. A dispatch that exceeds it is abandoned and surfaced as a TimeoutException so the request pipeline releases its slot rather than back-filling behind a wedged shard until the Orleans response timeout expires (MessagingOptions.ResponseTimeout, 30 s by default in Orleans - the same as this option's default - so the dispatch deadline shortens the wait only on hosts that raise the response timeout). Does not fix any wedge mechanism - the grain-side flush / activation deadlines already bound their own regions - it bounds the symptom on the writer side and makes every wedge it catches attributable to a specific (tree, shard) via the orleans.lattice.wal.append_dispatch.timeouts counter, which counts only this deadline's trips. InfiniteTimeSpan restores the historical unbounded await. |
WalDrainBudget |
TimeSpan |
75 s | Hard ceiling on how long a WAL partition grain's OnDeactivateAsync drain may run before the remaining in-flight slots are force-faulted and the chain is released so the activation can finish tearing down. Bounds the host-level SIGTERM drain so the silo's shutdown accounting always settles within bounded time of the SIGTERM, regardless of whether the storage provider is healthy. The drain signals every in-flight flush's linked cancellation token at drain entry (so a co-operative provider gives up promptly), waits for the chain to settle naturally for up to this budget, and then force-faults any slot that has not unlinked with a typed TimeoutException so callers parked on AppendAsync / AppendBatchAsync are released. The matching orleans.lattice.wal.shard.drain.budget.expirations counter and orleans.lattice.wal.shard.drain.budget.force_faulted_slots histogram attribute the trip per (tree, shard). InfiniteTimeSpan restores the historical unbounded-drain behaviour. |
StarvationDriveBudget |
TimeSpan |
5 min | Hard ceiling on how long a single WAL GC starved-leaf checkpoint drive may run while holding a permit on the per-silo WAL replay concurrency gate, before it abandons its replay and releases that permit (issue #3065). Before this budget existed every await inside the permit-guarded region was passed CancellationToken.None, so a drive whose commit-log read never returned held one of a small number of per-silo permits indefinitely and could not be cancelled; the gate drained and the silo presented as an activation outage. A caller-side timeout does not address this - the sweep's grain call already times out at the Orleans response-timeout default while the grain-side method keeps running and keeps holding its permit - so the budget is enforced inside the region, with a real CancellationTokenSource for work that honours cancellation and a bound on the drive's own wait for host-supplied storage that does not. The permit is acquired and released in the outer frame so abandonment cannot skip the release. Default is 4 * WalDrainBudget, comfortably above a legitimately slow full replay. Unlike most timeout options here, InfiniteTimeSpan is rejected rather than honoured, because an infinite budget restores exactly the outage this option bounds; zero and negative values are rejected too. An abandoned drive increments orleans.lattice.wal.replay.starvation_drive_abandonments and records the drove_timed_out drive verdict on orleans.lattice.wal.gc.blocked_leaf_reactivations. |
WalRetention |
TimeSpan? |
null |
Optional wall-clock hard ceiling for WAL retention. null means retention is bounded purely by consumer cursors. Trimmed by a WAL GC driver: the built-in scheduler that runs every WalGcInterval wherever the WAL collector is registered (AddLatticeWalGc, which the durable WAL providers and AddLatticeReplication call for you; AddLattice alone does not), or the replication maintenance grain for replicated trees. |
WalGcInterval |
TimeSpan |
1 hour (enabled) | Cadence at which the per-silo core WAL garbage-collection scheduler runs ILatticeWalGc.RunOnceAsync over every registered tree, so a durable-WAL host gets bounded WAL retention without the replication package and for non-replicated trees. The scheduler exists only where the WAL collector is registered (AddLatticeWalGc - the durable WAL providers and AddLatticeReplication call it; AddLattice alone does not); there, the hourly default makes WalRetention effective with no further configuration; a pass is retention housekeeping, so the coarse default keeps the storage cost low (cost scales with trees x WalPartitions per silo). Composes with the replication maintenance grain - RunOnceAsync and the underlying WAL TrimAsync are idempotent, and the pass honours the minimum consumer cursor and leaf-materialiser checkpoint floor, so it never over-trims. Global knob read from the default (unnamed) options; per-tree overrides do not apply. TimeSpan.Zero or a negative value disables the scheduler. |
WalMaxRetainedBytes |
long? |
null |
Optional advisory ceiling on retained WAL bytes per tree. When set, each ILatticeWalGc.RunOnceAsync pass samples retained bytes before and after its safe trim; if the pre-trim total exceeds the ceiling the policy schedules a byte-pressure trim (BytePressureTriggered), and BytePressureOverThreshold reports whether the tree is still over after the trim. Advisory only - the GC never trims past the safe frontier to honour it. null disables the policy. |
WalBytePressureReclaimTarget |
double |
0.8 | Fraction of WalMaxRetainedBytes a byte-pressure trim aims to reclaim toward. Ignored when WalMaxRetainedBytes is null. |
StorageUsageCacheTtl |
TimeSpan |
10 s | Cache lifetime for ILattice.GetStorageUsageAsync reports. TimeSpan.Zero disables caching. |
MaxConcurrentStorageUsageTrees |
int |
8 | Maximum trees a cluster-wide storage-usage roll-up (GetTotalStorageUsageAsync, RefreshStorageUsageAsync, PollWalUsageAsync) samples at once. The roll-up is a two-level fan-out and the levels multiply, so this times MaxConcurrentStorageUsageSurfaces is the peak in-flight grain-call count (128 by default). Unbounded, a 90-tree cluster at 64 shards and 8 WAL partitions dispatched roughly 6,500 concurrent calls that all raced one response deadline. Bounding makes the roll-up degrade in latency instead; the aggregated figures and the registry-sorted tree ordering are identical under any bound. Cluster-wide knob read from the default (unnamed) options; per-tree overrides do not apply. Values below 1 are clamped to 1. |
MaxConcurrentStorageUsageSurfaces |
int |
16 | Maximum per-tree storage surfaces - shard roots plus WAL partitions, bounded jointly - that one tree's ILattice.GetStorageUsageAsync fan-out queries at once. The inner level of the two-level roll-up fan-out, and per-tree overridable like MaxConcurrentSnapshotCaptures. The report is byte-for-byte identical under any bound; only the dispatch schedule changes. Values below 1 are clamped to 1. |
StorageUsagePollInterval |
TimeSpan |
15 s | Cadence at which every silo's background poller calls ILatticeAdmin.PollWalUsageAsync so the storage.wal_bytes and storage.policy.over_threshold gauges populate without any caller invoking the public API. The poll path is leaf-free: it activates only WAL partition grains, so idle trees stay cold. Snapshot / leaf-state / total-bytes gauges populate on demand via ILattice.GetStorageUsageAsync and ILatticeAdmin.RefreshStorageUsageAsync, or on the optional StorageUsageDeepPollInterval cadence. Global knob read from the default (unnamed) options; per-tree overrides do not apply. TimeSpan.Zero or a negative value disables the poller. |
StorageUsageDeepPollInterval |
TimeSpan |
TimeSpan.Zero (disabled) |
Optional cadence at which the same poller also drives the deep storage.snapshot_bytes / storage.leaf_state_bytes / storage.total_bytes gauges by calling the non-force ILatticeAdmin.GetTotalStorageUsageAsync. The deep read is O(1) per shard root (it never walks the leaf chain or activates per-leaf snapshot grains), so it activates only shard roots and never pins idle leaves resident; it never invokes the force-refresh path. Defaults to TimeSpan.Zero (disabled), preserving the activation-light WAL-only poll. Global knob read from the default (unnamed) options; per-tree overrides do not apply. |
RetryPolicy |
ILatticeRetryPolicy? |
null |
Optional opt-in retry policy applied at the boundary of the single-key ILattice mutators (SetAsync, SetIfVersionAsync, GetOrSetAsync, ApplyCrdtDeltaAsync, DeleteAsync) and the range deletes (DeleteRangeAsync, DeleteRangeWherePredicateAsync); the batch, atomic, and bulk-load entry points do not route through it. Only consulted under an active LatticeIdempotencyContext scope. null preserves the throw-and-revert default. See Retry Policy. |