ILattice: Maintenance operations
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at ilattice-2.md, and llms.txt lists every page.
Part of ILattice, in Lattice Public API Reference.
Maintenance operations
These methods manage tree structure and lifecycle. Several of them
take the tree offline - reads and writes throw
InvalidOperationException while the operation is in progress. Plan
maintenance windows accordingly.
Bulk loading
| Method |
Signature |
Description |
BulkLoadAsync |
Task BulkLoadAsync(IReadOnlyList<KeyValuePair<string, byte[]>> entries) |
One-shot bottom-up bulk load into an empty tree. Entries are sorted internally. Throws InvalidOperationException on the second and subsequent calls (every shard must still be empty). Not safe to use as a streaming-append primitive - for continuous ingestion, use SetAsync or the streaming BulkLoadAsync extension on LatticeExtensions. See Bulk Loading. Throws LatticeReplicationModeMismatchException when the tree is declared for cross-cluster replication as a typed CRDT mode (see Replication modes - Single shape per tree). |
BulkAppendChunkAsync |
Task<int> BulkAppendChunkAsync(string operationId, IReadOnlyList<KeyValuePair<string, byte[]>> sortedEntries) |
Appends one chunk of a bulk load and returns the number of entries appended after any write interception. The idempotent, resumable primitive underneath the streaming BulkLoadAsync extension: a per-shard operation id of the form "{operationId}-{shardIndex}" is derived and each shard records its last completed id, so re-driving the same chunk is a no-op. A shard remembers only its most-recently-completed id, so resume from the last un-acknowledged chunk and never re-drive a chunk a later chunk has superseded on the same shard. Keys must ascend within a chunk and across the stream. Enforces the whole-tree LatticeOperation.BulkLoad gate. Throws ArgumentException when operationId is null or empty and ArgumentNullException when sortedEntries is null; an empty chunk returns 0. Throws LatticeReplicationModeMismatchException when the tree is declared for cross-cluster replication as a typed CRDT mode (see Replication modes - Single shape per tree). See Bulk Loading. |
Tree lifecycle
| Method |
Signature |
Description |
TreeExistsAsync |
Task<bool> TreeExistsAsync() |
Returns true if this tree is registered. A tree is registered on first use (its first options resolution or shard-root operation, reads included) and unregistered when its purge completes. Authorized as a whole-tree Read: a caller the access gate denies is told false rather than refused, so existence is never disclosed to an unauthorized caller. |
GetAllTreeIdsAsync |
Task<IReadOnlyList<string>> GetAllTreeIdsAsync() |
Returns all registered tree IDs in sorted order. System trees (_lattice_*) are excluded. Physical trees created by ResizeAsync / SnapshotAsync are included. Authorized as a whole-tree Read of the tree the call is addressed to. When a tenancy add-on's enumeration filter is active and the caller has asserted an active tenant, the list is pruned to the trees that tenant may observe. |
DeleteTreeAsync |
Task DeleteTreeAsync() |
Soft-deletes the tree. Data is retained for LatticeOptions.SoftDeleteDuration before purge. Idempotent: a call on an id already recorded as deleted is a no-op. A tree created again under the id of a purged one is a new, live tree, and deleting it deletes it (see Reusing a purged tree ID). WARNING: Takes the tree offline - reads and writes throw InvalidOperationException until RecoverTreeAsync. Acts on the logical tree: on a tree a resize, a shadow-cutover restore or a schema remediation has aliased to a physical copy, it deletes the live copy the alias targets and pins that copy for the recovery and purge that follow (see Deleting an aliased tree). Throws InvalidOperationException when the alias targets a copy that was not created for this tree or that another tree also aliases, when another tree aliases this one, while a resize, restore or remediation holds the tree's alias, or when one or more materialised views derive from this tree; tear those views down first via ILatticeViewFactory.DeleteAsync (see Materialised views). See Tree Deletion. |
RecoverTreeAsync |
Task RecoverTreeAsync() |
Recovers a soft-deleted tree before purge completes. On an aliased tree it recovers the live copy the delete pinned. Throws InvalidOperationException when the tree has not been deleted (a live resized tree is not deleted, so recovering it throws), while a purge is in progress, or when its purge has already completed. On a tree created again under a purged tree's id it clears the stale deletion record and returns, since that tree is live (see Reusing a purged tree ID). |
PurgeTreeAsync |
Task PurgeTreeAsync() |
Immediately purges a soft-deleted tree without waiting for the retention window. WARNING: Permanently destroys all data. On an aliased tree it purges the live copy the delete pinned and unregisters both that copy and the logical tree. Accept-then-poll: the purge is recorded as in progress and its shard walk runs on the tree's deletion coordinator, where no response timeout can stop it part-way; the call waits at most 15 s (or half the silo response timeout, if shorter) and returns once the purge completes or with it still running - follow it through the tree-admin deletion status. A call while the purge runs, or after it completed while the id is still unregistered, returns without error. Throws InvalidOperationException when the tree has not been deleted - including a tree created again under a purged tree's id, which is live (see Reusing a purged tree ID). See Manual Purge. |
Resize and reshard
| Method |
Signature |
Description |
ResizeAsync |
Task ResizeAsync(int newMaxLeafKeys, int newMaxInternalChildren) |
Online - changes the tree's node fan-out. newMaxLeafKeys must be greater than 1 and newMaxInternalChildren greater than 2, or ArgumentOutOfRangeException is thrown. Idempotent for the same parameters while running; throws InvalidOperationException when a resize with different parameters, or a reshard, is already in progress, while an undo of an earlier resize is still unwinding (IsResizeUndoPendingAsync reports true), while the tree is deleted, or while a shadow-cutover restore or schema remediation holds the tree's alias. Reads and writes remain available throughout, but the copy has the limits of an online SnapshotAsync (below): a typed CRDT delta apply or bulk append that reaches a shard after the copy has passed its key is not carried into the resized tree. The copy follows the tree's shard map, so a shard an adaptive split added is copied too, and the swap carries the map over to the resized tree. Undoable within LatticeOptions.SoftDeleteDuration. Returns once the intent is persisted; use IsResizeCompleteAsync to poll for completion. Crash-safe. See Tree Sizing. |
UndoResizeAsync |
Task UndoResizeAsync() |
Undoes the most recent resize. Available at every phase - before the swap it aborts cleanly, and after the swap it restores the old tree, recovering it from soft-delete only if the resize had already retired it. Once the resize has completed, valid while the old tree is still within LatticeOptions.SoftDeleteDuration. Accept-then-poll: the undo is admitted even while a resize phase is in flight, the coordinator unwinds at its next phase or snapshot-slice boundary, and the call waits only a bounded time (inside the default response timeout) before returning; poll IsResizeUndoPendingAsync for an unwind that outlasts it. A retry while the undo is pending is acknowledged again. Throws InvalidOperationException when there is no resize to undo (naming the resize already undone, when there is one), when the unwind cannot be applied because the old tree has already been purged, while the tree is deleted, or while a shadow-cutover restore or schema remediation holds the tree's alias. See Tree Sizing. |
ReshardAsync |
Task ReshardAsync(int newShardCount, CancellationToken cancellationToken = default) |
Online - grows or shrinks the tree's physical shard count to newShardCount: a larger count splits shards, a smaller one folds adjacent shards together and releases the retired shards' storage before the reshard reports complete. newShardCount must be at least 2 and <= the smaller of 4096 and the tree's virtual-slot count (the number of slots in its shard map: 4096 unless an installed app's manifest declared another virtualShardCount), or ArgumentOutOfRangeException is thrown. The current count is a no-op, while a tree observed to be empty is re-pinned to any count in that range without a migration, its shard map rebuilt as the identity map over the same slot count. Idempotent for the same target while running; throws InvalidOperationException when a different target, or a resize, is already in progress. Returns once the intent is persisted; use IsReshardCompleteAsync to poll. Crash-safe. Transparently absorbs up to two ShardActivationTimeoutExceptions from the coordinator's shard-root activation-readiness seed (a cold-start race during startup reshards); a third consecutive seed timeout surfaces the typed exception to the caller. See Online Reshard. |
Merge
| Method |
Signature |
Description |
MergeAsync |
Task MergeAsync(string sourceTreeId) |
Merges every entry from sourceTreeId into this tree using last-writer-wins by HybridLogicalClock timestamp. Tombstones are preserved. The source tree is unmodified. Source and target trees may have different shard counts. If a shard consolidation on the source tree (a shrinking reshard or an automatic healing fold) retires a source shard while the merge runs, the merge re-drains the source's current shards before it completes. Throws ArgumentException when sourceTreeId equals this tree, and InvalidOperationException when a user-origin call names a source in a reserved namespace (_lattice_, sys-, or t/), when the source does not exist, or when a merge from a different source is already in progress (re-issuing the same source is idempotent). The caller must be authorized to read the whole source tree - a key-filtered grant is refused rather than narrowed - as well as to administer this one (LatticeAuthorizationDeniedException otherwise). See Architecture. |
Snapshots
| Method |
Signature |
Description |
SnapshotAsync |
Task SnapshotAsync(string destinationTreeId, SnapshotMode mode, int? maxLeafKeys = null, int? maxInternalChildren = null) |
Copies the tree's live (not deleted, not expired) entries into destinationTreeId; each copied entry keeps its source HLC version and absolute TTL expiry. In Offline mode the source is locked during the copy, so the destination is a point-in-time image. In Online mode the source remains available throughout and the writes it accepts while the copy runs are shadow-forwarded, so the destination is a mirror maintained during the copy rather than a point-in-time image; typed CRDT delta applies and bulk appends are not forwarded. Either mode registers the destination with the source's shard map and copies every physical shard that map routes to, a shard an adaptive split added included, keeping each entry only on the shard the map routes its key to, so the stale copies a split leaves behind are not carried over. The destination must not already exist: the snapshot creates it with the source tree's shard count, and the optional sizing overrides apply to it. A null sizing override takes the library default rather than the source tree's sizing. Throws ArgumentOutOfRangeException when a supplied maxLeafKeys is not greater than 1 or maxInternalChildren not greater than 2, ArgumentException when destinationTreeId names this tree, LatticeReservedTreeNamespaceException when a user-origin call names a destination in the _lattice_ or sys- namespace or in another tenant's t/ namespace, and InvalidOperationException when the destination already exists or a snapshot with different parameters is already in progress (re-issuing identical parameters is idempotent). WARNING: Offline mode takes the tree offline. See Snapshots. |
Operation status
| Method |
Signature |
Description |
IsMergeCompleteAsync |
Task<bool> IsMergeCompleteAsync() |
true once no merge is in progress (vacuously true when none has ever been initiated). Monotonic: once true for a given operation, never returns false again. |
IsSnapshotCompleteAsync |
Task<bool> IsSnapshotCompleteAsync() |
Same semantics for SnapshotAsync. |
IsResizeCompleteAsync |
Task<bool> IsResizeCompleteAsync(CancellationToken cancellationToken = default) |
Same semantics for ResizeAsync. Answers without waiting for an in-flight resize phase. |
IsResizeUndoPendingAsync |
Task<bool> IsResizeUndoPendingAsync(CancellationToken cancellationToken = default) |
true while an undo accepted by UndoResizeAsync is still unwinding; false once it has finished or when none was requested. Takes precedence over IsResizeCompleteAsync when telling a running resize, an unwinding undo, and no resize apart (an undo of an already completed resize leaves IsResizeCompleteAsync true). Answers without waiting for an in-flight resize phase. |
IsReshardCompleteAsync |
Task<bool> IsReshardCompleteAsync(CancellationToken cancellationToken = default) |
Same semantics for ReshardAsync. |
Diagnostics
| Method |
Signature |
Description |
DiagnoseAsync |
Task<TreeDiagnosticReport> DiagnoseAsync(bool deep = false, CancellationToken cancellationToken = default) |
Returns a per-shard health snapshot - depth, root-is-leaf, live-key count, tombstone count (deep only), hotness counters, ops/sec, split/bulk state - plus a bounded ring buffer of recent adaptive-split events. Repeated calls within LatticeOptions.DiagnosticsCacheTtl (default 5 s) are served from cache; shallow and deep reports are cached independently. Not for hot-path or correctness-critical decisions - use the operation-specific APIs (CountAsync, IsResizeCompleteAsync, etc.) instead. See Diagnostics. |
RebuildLeafProjectionAsync |
Task RebuildLeafProjectionAsync(int shardIndex, CancellationToken cancellationToken = default) |
Operator-driven recovery: clears the projection state of every leaf in the specified physical shard, so each leaf's next activation re-materialises its projection - from the leaf's snapshot where one exists, then the WAL after it. Topology-bearing state is preserved. See Operator tooling and Projection Rebuild. |
GetMaterialiserLagAsync |
Task<long> GetMaterialiserLagAsync(CancellationToken cancellationToken = default) |
Returns the largest per-shard materialiser lag, in WAL entries: for each physical shard, how far its WAL partition heads run ahead of the lowest leaf-projection checkpoint across that shard's leaves (summed over the shard's partitions; a shard with no leaves reports its heads). Authorized as a whole-tree Read. The figure is an estimate: each head is the next offset its partition will assign and each checkpoint the last offset a leaf applied, and every partition's head is measured against the leaf's partition-0 checkpoint, so a caught-up shard with a non-empty WAL still reports a small positive value rather than 0. Read a steady value as caught up and a growing one as the materialiser falling behind WAL ingestion. See Operator tooling. |
CompactShardAsync |
Task<bool> CompactShardAsync(int shardIndex, CancellationToken cancellationToken = default) |
Operator-tooling tombstone-compaction request: schedules an out-of-cycle compaction pass scoped to a single physical shard, bypassing the per-shard cooldown gate. Returns false when compaction is disabled (TombstoneGracePeriod = Timeout.InfiniteTimeSpan) or when a pass is already in flight. Throws ArgumentOutOfRangeException when shardIndex is not a physical shard of the tree. See Tombstone Compaction. |
var report = await tree.DiagnoseAsync(deep: true, cancellationToken);
Console.WriteLine($"Tree {report.TreeId}: {report.TotalLiveKeys} live, {report.TotalTombstones} tombstones across {report.ShardCount} shards.");
foreach (var shard in report.Shards)
{
Console.WriteLine($" shard {shard.ShardIndex}: depth={shard.Depth}, live={shard.LiveKeys}, ops/s={shard.OpsPerSecond:F1}");
}
Storage usage
| Method |
Signature |
Description |
GetStorageUsageAsync |
Task<TreeStorageUsageReport> GetStorageUsageAsync(CancellationToken cancellationToken = default) |
Returns a byte-accurate breakdown of the tree's on-disk footprint: retained WAL bytes, physical WAL bytes (what the WAL backend actually occupies, including framing and trimmed-but-unreclaimed space), captured leaf-snapshot bytes, and leaf-state bytes, plus the storage total (TotalBytes - the physical WAL, snapshot, and leaf-state figures summed). The aggregator fans out to every physical shard and WAL partition - bounded by LatticeOptions.MaxConcurrentStorageUsageSurfaces (default 16) so a wide tree queries its shard roots in waves - then caches the assembled report for LatticeOptions.StorageUsageCacheTtl (default 10 s). Partial is true when a surface could not be accounted: either the configured IWalStorageProvider does not support byte accounting (the in-memory and Azure Table providers both do), or a shard root / WAL partition failed or timed out. A surface that did not answer contributes nothing rather than a zero, so the total is a flagged lower bound rather than a silently understated figure. Diagnostic / capacity-planning use only - not a hot-path API. |
var usage = await tree.GetStorageUsageAsync(cancellationToken);
Console.WriteLine($"Tree {usage.TreeId}: {usage.TotalBytes} bytes total " +
$"(WAL={usage.WalPhysicalBytes}, snapshots={usage.SnapshotBytes}, leaf-state={usage.LeafStateBytes})");
The TreeStorageUsageReport fields are:
| Field |
Type |
Meaning |
TreeId |
string |
The tree the report covers. |
WalRetainedBytes |
long |
Retained (un-trimmed) WAL payload bytes summed across every partition. |
WalPhysicalBytes |
long |
Physical WAL bytes summed across every partition: every byte the backend occupies for the tree, including per-record framing and dead (trimmed but not yet reclaimed) payload, so it can exceed WalRetainedBytes. Falls back to the retained figure for a provider without physical accounting. |
SnapshotBytes |
long |
Captured leaf-snapshot key + value bytes summed across every leaf. |
LeafStateBytes |
long |
Leaf-state key + value bytes - UTF-8 key length plus stored value length over every row a leaf holds, a tombstone counting its key only - summed across every leaf in every shard. |
TotalBytes |
long |
WalPhysicalBytes + SnapshotBytes + LeafStateBytes - the WAL term is the physical figure, not WalRetainedBytes. |
Partial |
bool |
true when a storage surface could not be accounted - a WAL provider without byte accounting, or a shard root / WAL partition that failed or timed out. The unaccounted surface contributes nothing rather than a zero, so the report is a flagged lower bound. |
SampledAt |
DateTimeOffset |
When the underlying fan-out was sampled. |
LiveKeys |
long |
Summed live (non-tombstone) key count across every shard - the figure per-tree admission control compares against LatticeOptions.MaxLiveKeys. Best-effort / eventually-consistent (see Metrics - admission control). A lower bound when Partial is true due to a shard that did not answer; a WAL-only cause of Partial leaves it exact. |
WalMaxRetainedBytes |
long? |
The effective retained-WAL ceiling for this tree - the registry runtime override if one is set, else the named LatticeOptions for the tree, else the silo-wide default - or null when no ceiling applies. A diagnostic echo of the resolved value; it never changes trim behaviour. Always exact, so Partial says nothing about it. Consumed by Orleans.Lattice.Scaling, which has no options monitor of its own and would otherwise have to read the silo-wide default for every tree. |
For a cluster-wide roll-up across every registered tree, resolve the
ILatticeAdmin grain (see ILatticeAdmin).
Lifecycle
| Method |
Signature |
Description |
WarmUpAsync |
Task WarmUpAsync(CancellationToken cancellationToken = default) |
Pre-activates every physical shard root and each shard's current root-node grain (root leaf when the tree is flat, root internal node otherwise) for this tree using a bounded-concurrency fan-out. On a brand-new empty shard, warm-up runs the same root-materialisation path the first traffic write would take - it materializes the deterministic root leaf at startup instead of under hot-path load. Designed to be called once at host startup, after the lattice has been resolved and before the first hot-path write lands, so the Orleans placement-directory + grain-storage first-touch cost (and root-materialization persistence cost on an empty tree) is absorbed while the silo is idle rather than against producer-driven flush concurrency. The fan-out is capped at min(physicalShardCount, 32) simultaneous probes. Throws LatticeReservedTreeNamespaceException (an InvalidOperationException subclass) when invoked on an internal system tree. Requires whole-tree Read authorization - the same authority as DiagnoseAsync - so a denied caller receives LatticeAuthorizationDeniedException. Records the orleans.lattice.warmup.invocations counter and the orleans.lattice.warmup.duration histogram. |
// Run once during host startup, after the lattice has been
// resolved but before any producer traffic is accepted.
await tree.WarmUpAsync(cancellationToken);
Optional leaf-cache pre-warm
Warm-up pre-activates shard roots and root nodes unconditionally. It can
additionally prime the read-through leaf caches that were hottest before the
silo went down, which is what removes the cold-start read-latency spike rather
than just the first-write spike. It is on by default:
LeafCachePreWarmCount defaults to 8
leaves per shard, and 0 switches the whole feature off (no access tracking,
nothing persisted, nothing primed).
When enabled, each shard root maintains a bounded histogram of the leaves its
reads route to and persists a compact snapshot of it alongside its own durable
state (on a coalescing timer, LeafAccessModelFlushIntervalMs, and on
deactivation). A reactivated shard root restores that histogram, and when
WarmUpAsync next reaches it, ranks the histogram by observed read frequency
and primes the top N leaf caches. Ranking by frequency rather than
recency is what makes the selection useful: an LRU list ranks a leaf touched
once immediately before shutdown above a leaf that is read constantly, whereas
frequency ranks by how much of the shard's read traffic actually lands on each
leaf. Because the shard root is the only caller of the [StatelessWorker]
cache, the primed activations land on the silo that will serve the subsequent
reads.
Pre-warm is strictly best-effort: a leaf that has since been merged away, or a
transient storage fault, is swallowed and never fails WarmUpAsync. See
LeafCachePreWarmCount for the memory
and persistence bounds, and Metrics for the
orleans.lattice.warmup.leaf_cache.* instruments.
Events
| Method |
Signature |
Description |
SetPublishEventsEnabledAsync |
Task SetPublishEventsEnabledAsync(bool? enabled, CancellationToken cancellationToken = default) |
Sets or clears the per-tree override for event publication. true forces on, false forces off, null removes the override so the tree inherits LatticeOptions.PublishEvents. The override is persisted and survives silo restarts. Propagation is best-effort. See Events - Per-tree override. |
// Force events on for this tree regardless of the silo default:
await tree.SetPublishEventsEnabledAsync(true, cancellationToken);
// Clear the override and inherit the silo default:
await tree.SetPublishEventsEnabledAsync(null, cancellationToken);
For subscribing to published events on the cluster client, see
SubscribeToEventsAsync under
LatticeExtensions.
Change history
| Method |
Signature |
Description |
ScanEntryHistoryAsync |
Task<EntryHistoryPage> ScanEntryHistoryAsync(string key, HybridLogicalClock? fromHlc, HybridLogicalClock? toHlc, int limit, string? continuation) |
Reads one page of a key's revision timeline between optional inclusive HLC bounds. Served from the durable history view when one is enabled (EntryHistoryPage.Source is View); otherwise falls back, best-effort, to the retained WAL window, reporting truncation through Truncated / EarliestAvailable (WalWindow), or an empty None page when the commit-log read seam is not registered. The page carries its Revisions oldest-first and a Continuation token that is non-null while more remain. limit must be positive (ArgumentOutOfRangeException) and is capped at 1,024 revisions per page; a key the access gate does not let the caller read returns an empty None page rather than throwing. Non-mutating. See Change history. |
SetHistoryRetentionAsync |
Task SetHistoryRetentionAsync(HistoryRetentionMode? mode, TimeSpan? window) |
Sets this tree's durable-history retention: the HistoryRetentionMode applied to LWW value bytes in revision rows (MetadataOnly, FullValue, or Hybrid; null clears the override, falling back to MetadataOnly) and the age after which a row expires (null removes the age bound; a supplied window must be positive). Persisted on the registry entry; absorbed forward without a view rebuild. See History views. |
GetHistoryRetentionAsync |
Task<HistoryRetentionSettings> GetHistoryRetentionAsync() |
Returns the effective retention policy - the persisted override, or the defaults (MetadataOnly, no age bound). Authorized as a whole-tree Read. |