Consistency Guarantees
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at consistency.md, and llms.txt lists every page.This document is the contract for what a caller of ILattice is
guaranteed to observe. It states each guarantee in caller-visible terms
only and does not describe how Lattice delivers them.
For the implementation of any guarantee, follow the cross-references in each section:
- Topology changes - Shard Splitting, Online Reshard
- Atomic batches - Atomic Writes
- Multi-page enumerations - Durable Cursors, Snapshot Cursors
- Durable copies - Snapshots, Tree Sizing, Tree Storage
- Read path - Read Caching
- State merge - State Primitives
- TTL expiry - TTL
Consistency levels
The tables below state a guarantee for each operation they list, mostly in
terms of the following four levels. They do not cover every ILattice
method: a method that no table lists is not classified here.
| Level | What the caller observes |
|---|---|
| Linearizable | The call takes effect at a single point between its invocation and its return. After the call returns, any subsequent read on any client at any silo observes the new value (subject to the cache-staleness note below). |
| Strongly consistent | The call observes a result consistent with some real-time point during its execution - no entry is missed, double-counted, or misattributed even when the underlying topology is changing concurrently. For scans this means no phantom or missing keys; for counts it means the exact live key count. |
| Snapshot (online) | The call observes a best-effort point-in-time view that is correct per-key but is not guaranteed to be a single global instant. Equivalent to a non-repeatable read isolation level. |
| Eventually consistent | The call may reflect a bounded staleness window (read-cache staleness, replication lag), but converges to the authoritative state within a configurable interval. |
An additional property - atomicity - is called out for batch operations and means all-or-nothing commit. Atomicity is orthogonal to the visibility model of concurrent readers, and both are stated where they apply.
A separate property - atomic visibility - is called out for atomic batches and snapshot reads, and means a concurrent reader observes either the full effect of a saga or none of it, never a partial view.
Single-key operations
| Operation | Guarantee | Notes |
|---|---|---|
GetAsync |
Linearizable under default CacheTtl = TimeSpan.Zero; eventually consistent when CacheTtl > 0 |
The default cache configuration refreshes on every read, so the call observes the latest committed value. Raising CacheTtl trades freshness for fewer round-trips; staleness is then bounded by CacheTtl + one cache refresh. |
GetWithVersionAsync |
Linearizable | Returns the value along with its authoritative HLC version for use in CAS loops. Bypasses the read cache. |
ExistsAsync |
Linearizable under default CacheTtl = TimeSpan.Zero; eventually consistent when CacheTtl > 0 |
Same dual classification as GetAsync. |
SetAsync (with or without TTL) |
Linearizable | The write is durably persisted before the call returns. Continues to hold across shard splits, resize, and reshard - topology changes are retried transparently (see below). |
SetIfVersionAsync |
Linearizable CAS | Atomic compare-and-set against the HLC version returned by GetWithVersionAsync. |
GetOrSetAsync |
Linearizable | No read-then-write race. |
DeleteAsync |
Linearizable | The deletion is visible to subsequent reads under the same guarantee as any other write. |
Single-key operations transparently retry on any topology-change exception, within a 60-second wall-clock budget per call. A caller sees a topology-change fault only when the topology keeps changing for that whole budget, in which case the library's internal stale-routing exception surfaces rather than a silently wrong result.
Batch operations
| Operation | Guarantee | Notes |
|---|---|---|
GetManyAsync |
Per-key linearizable (default CacheTtl); per-key eventually consistent (CacheTtl > 0) |
Atomic visibility tree-wide: a concurrent SetManyAtomicAsync is observed either entirely or not at all across the requested key set. The batch is not a global snapshot across non-saga keys - two unrelated keys may reflect different real-time points. When the transaction registry is unreachable and a requested key carries a prepared saga mutation, the call throws the retryable LatticeTransactionOutcomeUnavailableException rather than risk a torn result. |
SetManyAsync |
Per-key linearizable, batch non-atomic | Each key is written under its own linearization point. A partial failure leaves the batch half-applied with no rollback, and because the fan-out fails fast without cancelling its siblings, writes for other keys may still be in flight when the exception is thrown - the half-applied state is settled only once those branches finish, so re-read rather than assume. The same holds when an opt-in finite LatticeOptions.SetManyFanOutBudget (unbounded by default) expires: the call is refused with LatticeSaturatedException while the outstanding branches run to completion. An opt-in finite LatticeOptions.SetManyEnvelopeBudget behaves the same way and additionally covers a single-shard batch, which the fan-out budget does not bound. |
SetManyAtomicAsync |
Per-key linearizable, batch atomic, atomic-visible tree-wide, atomic-visible across clusters | All-or-nothing: on success every key holds its new value; on failure every key holds its pre-saga value. A concurrent reader observes the saga atomically - the post-decision visibility flip is a single tree-wide point. The atomic-visibility guarantee extends across every cluster the tree replicates to. See Atomic Writes. |
DeleteRangeAsync |
Strongly consistent | Every key in the range is tombstoned. Robust against sparse multi-shard distributions. For resumable or crash-safe range deletes use OpenDeleteRangeCursorAsync instead. |
CountAsync |
Strongly consistent, atomic-visible tree-wide | Exact live key count under the topology snapshot the call observes. A concurrent SetManyAtomicAsync is observed atomically (included entirely or excluded entirely). Throws InvalidOperationException if topology changes outrun the retry budget (LatticeOptions.MaxScanRetries, default 3), or LatticeTransactionOutcomeUnavailableException if the transaction registry stays unreachable while a counted key carries a prepared saga mutation (see When the registry cannot be reached). |
CountPerShardAsync |
Strongly consistent, atomic-visible tree-wide | Per-shard counts are topology-consistent with the observed shard layout. Same atomic-visibility guarantee as CountAsync. |
Enumeration
| Operation | Guarantee | Notes |
|---|---|---|
ScanKeysAsync |
Strongly consistent, strictly ordered, atomic-visible across each uninterrupted enumeration | Keys are yielded in lexicographic order with no duplicates and no gaps, even when shard splits or rebalances run concurrently. A concurrent SetManyAtomicAsync is observed identically across every page: either all of its keys appear or none. Bounded by LatticeOptions.MaxScanRetries (default 3); throws InvalidOperationException if the retry budget is exhausted. Transparently recovers from server-side enumeration aborts up to the wrapper's maxAttempts parameter (default 8), and from stalled page fills on a smaller budget (see Retry exhaustion), but each such reopen captures a fresh saga-decision view, so a saga that commits between the fault and the reopen can appear pre-commit for keys yielded before the reopen and post-commit for keys after it. |
ScanEntriesAsync |
Strongly consistent, strictly ordered, atomic-visible across each uninterrupted enumeration | Same key ordering and atomic-visibility guarantees as ScanKeysAsync. Values reflect the authoritative state at the moment each key is yielded. |
Durable cursor steps - live mode (NextKeysAsync, NextEntriesAsync, DeleteRangeStepAsync) |
Per-step strongly consistent (key and entry steps also atomic-visible), cross-step snapshot | Each key or entry step is a strongly consistent scan, atomic-visible tree-wide within that step. A DeleteRangeStepAsync step deletes through DeleteRangeAsync and so carries its per-key-only visibility. Across steps, a key updated between two pages is observed at its newest value when it is next visited, but once yielded by a cursor it is never re-yielded. A saga that commits between page i and page i+1 may have its keys split across the two pages - use point-in-time mode (below) for cross-step atomicity. See Durable Cursors. |
Durable cursor steps - point-in-time mode (opened with pointInTime: true) |
Strongly consistent, strictly ordered, atomic-visible for the cursor's lifetime | Every page reads against the saga-decision view captured at OpenAsync time. A SetManyAtomicAsync that commits between two pages is observed identically on every page (either all of its keys, or none). A stalled cursor whose pin lifetime is exceeded surfaces LatticeCursorSnapshotExpiredException on its next call and must be reopened. Not available for DeleteRangeStepAsync. See Durable Cursors - Point-in-time cursors. |
Snapshot cursor steps - zero-observable-writes mode (OpenSnapshotKeyCursorAsync, OpenSnapshotEntryCursorAsync) |
Snapshot-isolated, strictly ordered, atomic-visible for the cursor's lifetime | Every page reflects the tree state captured at open time. No write committed after open - foreground SetAsync / DeleteAsync, saga SetManyAtomicAsync, DeleteRangeAsync, or replication apply - is ever visible to the cursor on any page. The captured LatticeSnapshotCoordinate is deterministic across silo failover. Open-time capture cost (the materialised per-shard baseline row count) is bounded by LatticeOptions.MaxSnapshotReplayEntries; exceeding the budget throws LatticeSnapshotReplayBudgetExceededException. Pages are served from a per-cursor frozen baseline captured at open - held in memory, and persisted durably before the first page that reports more results - so a later WAL GC that trims the committed prefix cannot empty or partial-fill the snapshot. Across an adaptive shard split the snapshot pins the ShardMap at open and resolves each key's owning shard by virtual slot under that pinned map, so every key is surfaced exactly once at its last-writer-wins value as of the pinned point in time - a donor shard's retained orphan copies of moved keys are not re-surfaced, and post-split writes shadow-forwarded through the donor are still observed. See Snapshot Cursors. |
Retry exhaustion
CountAsync, CountPerShardAsync, ScanKeysAsync, and
ScanEntriesAsync use a bounded retry budget
(LatticeOptions.MaxScanRetries, default 3) to reconcile against
concurrent topology changes. If the topology continues to mutate beyond
the budget the call throws InvalidOperationException rather than
returning a silently incomplete result. This is not a realistic concern
under default settings; see
API Reference - Scan reliability for tuning
guidance.
Independently of topology, a multi-key read (GetManyAsync, the counts,
and the key and entry scans) that reaches a key carrying a prepared saga
mutation while the transaction registry cannot be reached throws the
retryable LatticeTransactionOutcomeUnavailableException instead of
guessing that key's value. A read whose keys carry no prepared mutation
completes normally while the registry is unreachable.
The streaming scan wrappers (ScanKeysAsync, ScanEntriesAsync)
additionally recover from mid-scan enumeration aborts (silo failover,
idle expiry, cold start, scale-down, and - in proportion to concurrent load
rather than to any incident - a page request that reaches a stateless-worker
activation other than the one holding the enumerator) up to the wrapper's
maxAttempts parameter (default 8). On reconnect the stream resumes with no
duplicates, no gaps, and ordering preserved, under a freshly captured
saga-decision view.
They also resume a scan whose shard page fill stalled past its hard ceiling
(ScanPageStalledException), on a separate, smaller budget: at most
min(maxAttempts, 2) consecutive stalls with no key or entry yielded in
between, and at most 64 stall resumes over the whole scan. A stall resume
reopens the scan the same way, under a freshly captured saga-decision view.
Once either budget is spent the ScanPageStalledException is rethrown, so a
scan that could not finish never ends as if it had.
Maintenance operations
| Operation | Guarantee | Notes |
|---|---|---|
BulkLoadAsync |
Linearizable on an empty tree | Throws if any shard already has data. After return, all entries are visible under the guarantees above. |
SnapshotAsync(Offline) |
Linearizable point-in-time copy, within the limits noted | The source's physical shards - 0 to ShardCount - 1 of its pinned shard count and every shard its shard map routes to, split-added shards included - are locked (reads and writes throw InvalidOperationException) when the copy starts, and each is unlocked as soon as its own copy completes, so the source becomes available again shard by shard. The destination holds the live entries those shards held when they were locked, and routes them through the source's shard map (explained below the table); tombstones and expired entries are not copied. |
SnapshotAsync(Online) |
Strongly consistent, within the limits noted | The source tree remains available for linearizable point traffic and strongly-consistent scans throughout. The source's shards - the same set Offline locks - mirror their writes to the destination shard of the same index while the drain copies their live entries, and the destination resolves each key last-writer-wins between a mirrored write and the drained copy. A mirrored write other than a merge does not carry the source's timestamp - the destination stamps it from its own clock - while a drained copy keeps its source timestamp, so the order in which the two arrive can decide which survives: a drained copy merged after a newer mirrored write of the same key can outrank it when the source's clock runs ahead of the destination's, leaving the destination at the older value. The mirror does not carry typed CRDT deltas (ApplyCrdtDeltaAsync, ApplyCrdtDeltaManyAsync, and every typed CRDT accessor, all of which write through them), so a delta applied to a key after the drain has copied it does not reach the destination. Nor does it carry bulk-load appends (BulkAppendChunkAsync and the streaming BulkLoadAsync extension), so an entry appended where the drain has already passed does not reach the destination either. Tombstones and expired entries are not copied: a key deleted before the snapshot began, or already expired when the drain reached it, has no entry on the destination, so a mirrored merge with an older timestamp - a replicated write, for example - is applied there, whereas the source, still holding the delete or the expired entry, discards it. |
ResizeAsync / UndoResizeAsync |
Linearizable (online), within the limits noted | Point operations and strongly-consistent scans continue throughout. Callers observe at most a single transparent retry at the alias swap. The rebuild is an Online snapshot into the resized copy, and after the swap the tree routes by the shard map that snapshot followed, so the Online limits carry into the tree itself: a typed CRDT delta applied between the drain copying its key and the swap is lost, as is a bulk-load append that lands where the drain has already passed; a write made while the copy runs can lose to an older drained copy of its key, as for Online; and deletes made before the resize, like entries already expired when the drain reached them, are not carried, so a merge with an older timestamp - a replicated write, for example - that arrives during or after the resize is applied instead of discarded. UndoResizeAsync after the swap restores the old physical tree and deletes the resized copy, so every write the resized tree accepted after the swap is discarded. |
ReshardAsync |
Linearizable (online) | Reads and writes remain linearizable across every concurrent shard split. |
MergeAsync(sourceTreeId) |
Eventually convergent (LWW) | For each key present in both trees, the entry with the higher HLC wins. On completion the destination is strongly consistent with the LWW merge of both inputs. The source tree is unmodified. |
DeleteTreeAsync |
Linearizable (takes tree offline) | After return, every subsequent read or write throws InvalidOperationException until RecoverTreeAsync. Data is retained for SoftDeleteDuration before purge. The delete acts on the logical tree, so on a tree a resize, shadow-cutover restore or schema remediation has aliased to a physical copy it takes that copy offline, and RecoverTreeAsync and PurgeTreeAsync act on the same copy; a resize's retirement of its old copy is not a delete and leaves the tree readable. On such an aliased tree a shard that an adaptive split added after the alias was set is not taken offline, so keys routed to it stay readable and writable after the delete. |
RecoverTreeAsync |
Linearizable | Restores full availability. |
PurgeTreeAsync |
Destructive, accept-then-poll | Permanently removes all data. The shard walk runs in the background: the call waits a bounded time (15 seconds, or half the silo response timeout when that is shorter) and can return with the purge still running, so follow it through the tree-admin deletion status until it reports the purge complete. |
TreeExistsAsync, GetAllTreeIdsAsync |
Eventually consistent (registry read) | May briefly lag a concurrent registration or deletion observed by another client. |
IsMergeCompleteAsync, IsSnapshotCompleteAsync, IsResizeCompleteAsync, IsReshardCompleteAsync |
Monotonic | Once true for a given operation, never returns false again. Vacuously true when no operation of that kind has ever been initiated. |
DiagnoseAsync |
Point-in-time snapshot (non-linearizable) | A best-effort per-shard health sample for dashboards and post-mortems. Repeat calls within LatticeOptions.DiagnosticsCacheTtl return the same cached result. Not for hot-path or correctness-critical decisions - use the operation-specific APIs instead. See Diagnostics. |
How a copy places keys: SnapshotAsync registers the destination with the
source's ShardMap (or none, for a source still on its default routing) and
split allocation mark, then locks (in Offline mode), mirrors (in Online
mode) and copies every physical shard in the union of 0 to
ShardCount - 1 and the indices that map routes to, each into the
destination shard of the same index; ResizeAsync rebuilds a tree through the
same copy and carries that map onto the tree at the swap. Every slot therefore
routes to the same physical index on both trees, whether the source's map was
reshaped by an adaptive split (which moves slots to a shard numbered from
ShardCount upward without changing the pinned count), an online reshard, a
shard consolidation, or an installed app's virtualShardCount. A shard that
gives up a slot keeps its own, now out-of-date, copy of that slot's keys, so
the copy keeps an entry only when the map routes its key to the shard it was
read from. The map is captured when the copy starts; the autonomic split
monitor starts no split while a snapshot or resize is in flight.
Atomic visibility
SetManyAtomicAsync delivers strict atomic visibility tree-wide and
across clusters: no reader, on any silo, in any cluster, ever
observes a partial view of an in-flight saga. Concretely, once
SetManyAtomicAsync returns without throwing, every key in the batch
holds its target value across every silo in the local cluster with no
intermediate partial-visibility window observable to any reader. As
the saga's writes propagate to each peer cluster, the same
all-or-nothing window holds on every remote tree.
Atomic visibility is observed by every read path that does not stream across multiple grain calls:
| Read path | Atomic visibility |
|---|---|
GetAsync, ExistsAsync, GetWithVersionAsync, GetOrSetAsync, SetIfVersionAsync |
Per-key linearizable; an in-flight saga's writes stay invisible until the saga commits, at which point all of its keys flip atomically. Until then a read sees each key's pre-saga value, while GetOrSetAsync and SetIfVersionAsync treat a key carrying a pending saga write as absent. When the transaction registry reports a saga's outcome as indeterminate (for example a decision that aged out of TxDecisionRetention), GetAsync, ExistsAsync and GetWithVersionAsync read the key as absent rather than at either value; when the registry cannot be reached they throw the retryable LatticeTransactionOutcomeUnavailableException. |
GetManyAsync, CountAsync, CountPerShardAsync |
Tree-wide for the call. |
ScanKeysAsync, ScanEntriesAsync |
Tree-wide for each uninterrupted underlying enumeration; a transparent reconnect after an enumeration abort, or a resume after a stalled page fill, resumes under a freshly captured saga-decision view. |
| Durable key/entry cursor (point-in-time mode) | Tree-wide for the lifetime of the cursor. |
| Snapshot key/entry cursor (zero-observable-writes mode) | Tree-wide for the lifetime of the cursor. Stricter than point-in-time mode: hides every concurrent write, not only sagas. |
| Durable key/entry cursor (live mode) | Tree-wide within each step; not preserved across steps. A delete-range cursor step deletes with the per-key-only visibility of a one-shot DeleteRangeAsync (next row). |
DeleteRangeAsync (one-shot) |
Per-key only; a concurrent saga may be observed as committed for some keys and pending for others. Use SetManyAtomicAsync to layer atomic deletion semantics on top. |
See Atomic Writes for the saga primitive and Durable Cursors for the point-in-time cursor mode.
Cross-tree (multi-tree) atomic visibility
IGrainFactory.SetManyAtomicAsync (and the BeginAtomicWrite
builder) extend the single-tree atomic-visibility contract to a batch
spanning two or more distinct ILattice trees: either every targeted
key across every participating tree becomes visible, or none of them do.
A two-level saga drives this - a coordinator grain keyed by the
operationId writes a single global commit/abort decision, and each
participating tree's saga decision registry delegates the status of its
prepared txid to that coordinator until the decision lands. Before the
decision, every tree returns InFlight for the saga (prepared keys are
invisible, indistinguishable from pre-saga); after it, every tree
returns the same global verdict. A tree that cannot reach the
coordinator at all returns Indeterminate instead, which keeps the
prepared keys invisible without asserting that the saga did not commit.
The coordinator's single decision write
is the cross-tree linearization point.
This is atomicity, not cross-tree read isolation, and the two are easy
to conflate. The atomic guarantee is anchored to the coordinator's single
decision write as one global linearization point: at any single instant the
saga is either undecided (no participating tree reports a verdict - InFlight
when the coordinator answers that it is still preparing, Indeterminate when
the coordinator cannot be reached - and every
prepared key invisible) or decided (every
tree returns the same global verdict). There is no instant at which the commit
is durably half-applied, so an observer that could sample every participating
tree at one instant always sees all-pre or all-post. What the design does
not add is a cross-tree read snapshot (serializable isolation across
trees): Lattice has no single read operation spanning multiple trees, so a
reader that issues separate per-tree reads at different instants which
straddle the linearization point can compose a mixed view - tree A's pre-saga
value read before the flip alongside tree B's post-saga value read after it -
exactly as two independent SELECTs under read-committed isolation can in a
relational store. That is a property of the reader taking several
point-in-time samples, not a hole in the write's atomicity: within any
single read, every saga key on that tree resolves against one global
verdict, so a tree is never seen showing a partial slice of one cross-tree
saga. See
Atomic Writes: Cross-tree atomic writes.
The same all-or-nothing visibility holds across clusters. Each
participating tree's terminals replicate on their own per-tree WAL feed,
so a receiver can apply one tree's terminal before a sibling tree's. A
receiver-side coordinator grain (ILatticeCrossTreeReceiverGrain), keyed
by (originClusterId, operationId), holds every participating tree
invisible until a terminal has arrived for each tree it expects, then
flips them together - mirroring the authoring cluster's single-write
flip. The set of trees it waits for is scoped to the trees actually
replicated on that receiver (the batch's participant set intersected with
LatticeReplicationOptions.ReplicatedTrees), so a cross-tree batch
spanning a mix of replicated and non-replicated trees remains valid: the
receiver preserves cross-tree atomic visibility across exactly the subset
of trees it hosts. See
Atomic Writes: Cross-cluster cross-tree visibility.
Topology and durability notes
Shard splits and reshards
Every guarantee in this document holds during an active shard split
or reshard. Callers do not normally observe topology-change
exceptions: point operations transparently retry within their
wall-clock budget, and scans use bounded reconciliation. In the rare
case of retry exhaustion, a point operation surfaces the original
stale-routing exception and a scan or count throws
InvalidOperationException, rather than returning a silently
incomplete result. See Shard Splitting.
Linearizability of point reads is also preserved across a leaf
reactivation that happens after a split. A post-split write routed to
the donor for an already-moved slot is shadow-forwarded into the target
shard's WAL but keeps the donor's source stamp; when the target's live
leaf reactivates cold and replays the WAL past a checkpoint that
pre-dates that forward, it resolves each replayed mutation's ownership
by the key's virtual slot under the current ShardMap rather than by
the stamped ShardIndex, so the forwarded record (for example a
post-split delete) is applied rather than dropped - the moved key reads
at its true last-writer-wins value and is never resurrected with a
drained pre-forward copy. See Shard Splitting - Live leaf
reactivation.
Read-cache staleness
GetAsync, ExistsAsync, and GetManyAsync are the only methods on
ILattice that may be served from the per-silo read cache. Staleness
is bounded by LatticeOptions.CacheTtl (default TimeSpan.Zero -
refresh on every read). Raising CacheTtl trades freshness for fewer
round-trips. GetWithVersionAsync bypasses the cache for CAS safety.
With LatticeOptions.OptimisticShardRootPointReads on (the default), a
GetAsync whose optimistic read validates is answered by the owning
leaf itself, not the cache; only a read that falls back to the serial
path can be served from the cache, so CacheTtl still bounds the
staleness a GetAsync can observe.
Cache staleness never weakens atomic visibility: keys covered by
an in-flight saga always observe the registry-coordinated outcome
regardless of CacheTtl.
Operations that hold a shard for their whole duration
A shard's serial calls - its non-optimistic reads, scan pages and
maintenance walks - are deliberately non-reentrant, so while one runs on
a shard every other serial request to that shard queues behind it, and
point writes are held back until it finishes (a serial call also waits
for the point writes already in flight before it starts). Only
always-interleaved calls overlap it: batch writes, a few diagnostic
probes, and optimistic point reads, which fall back to the queued
serial path whenever the running call could change the shard's
routing. Most leaf-chain
walks are bounded by
MaxLeavesPerScanPage and return a
partial page rather than holding the shard indefinitely.
A small set of operations is deliberately exempt, because whole-walk atomicity is the mechanism delivering a guarantee rather than an accident of implementation. Releasing the shard part-way through would open exactly the window each one exists to close:
| Operation | Why it stays atomic |
|---|---|
| Marking moved-away slots during a shard split | Runs immediately before the source enters Reject phase so no read crosses the Swap boundary observing an unmarked leaf. A partially-marked shard would serve a stale orphan value for a moved key. |
| Unmarking them during a consolidation | The mirror image: the survivor's leaves are unsealed only after the donor has been frozen and drained onto it, so no read crosses the freeze observing an unsealed leaf. |
| Capturing a snapshot-cursor baseline | A baseline covering only part of the leaf chain is not a baseline, so there is no partial result to bank and resume from. The captured WAL head is read only after every leaf has frozen, which is what makes the baseline a single instant: a write that lands mid-walk precedes that head and is folded into the baseline, and every later write is excluded. The hold is bounded instead by the hard stall ceiling (MaxScanPageStallDuration); abandoning the capture is safe because nothing is observable until the baseline is seeded. |
| Purging a shard on tree deletion | The tree is already offline, so nothing is waiting on it; the walk is also destructive and cannot be resumed once a leaf's sibling pointer is cleared. |
| Applying an atomic-write terminal when the shard has no record of the leaves the saga touched | This rare fallback walks the shard's whole leaf chain inside the terminal's turn. The terminal's timestamp must sort above every prepare the shard stamped, which a maximum taken over part of the chain cannot guarantee, and releasing the shard between leaves would let a reader see the saga applied on some leaves and still pending on others. |
DeleteRangeAsync used to be on this list and is not any more: its per-shard
walk is now work-bounded, driven to completion by the tree-level call. That
weakens nothing, because the fan-out across physical shards was already
parallel with no cross-shard atomicity and the documented visibility is per-key
only, so the whole-shard atomicity of the old walk was an implementation
artifact rather than a guarantee.
The cost of the ones that remain is that a shard holding one of them for a long
time is unavailable to every other caller meanwhile. Each logs a warning naming
the operation and the number of leaves visited when it exceeds a short internal
threshold, so a burst of Orleans long-request warnings on a shard can be
attributed to the walk causing it rather than only to the requests it blocked.
Three of them have had their hold shortened rather than removed. Marking and
unmarking moved-away slots discover the leaf chain sequentially but fan the
per-leaf writes out with bounded concurrency, so the hold costs roughly the
chain walk plus the slowest batch rather than the sum of every write. The
snapshot-baseline capture fans its per-leaf fold pass out
(MaxConcurrentSnapshotBaselineFolds)
and is bounded end to end by the hard stall ceiling. Reading the baseline's WAL
head before freezing is not a safe way to shorten it further: a leaf whose own
head had already moved past that point when it froze would bake later writes
into the baseline, and the folds cannot be rewound.
For a range delete that must be resumable or crash-safe across a process
restart, use OpenDeleteRangeCursorAsync; see
Durable Cursors.
TTL expiry
Expired entries are filtered on every user-facing read path. See TTL.
Clock skew
The guarantees in this document assume reasonably synchronised silo
clocks (not client clocks). Two concurrent writes resolve by HLC: the
write with the later wall-clock tick wins. In pathological drift
scenarios the tombstone-grace window
(LatticeOptions.TombstoneGracePeriod, default 24 h) gives a lagging
replica time to converge before physical reclamation.
Cancellation
Every ILattice method accepts a CancellationToken. Cancellation
before a mutation has committed leaves the operation as if it had never
been attempted. Cancellation after a long-running coordinator (saga,
resize, snapshot, reshard, merge) has accepted the request does not
roll back: the coordinator drives itself to a terminal state on its
own, and the matching Is*CompleteAsync eventually returns true.
What Lattice does not guarantee
- Global transaction ordering. Two concurrent
SetManyAtomicAsynccalls touching overlapping keys resolve pairwise by LWW; there is no serializable global order across sagas. - Reader isolation during one-shot range deletes.
DeleteRangeAsyncis per-key linearizable but is not registry-coordinated, so a reader concurrent with a long-running range delete may observe some keys tombstoned and others still live. For an isolated range delete, stage it viaSetManyAtomicAsyncor gate visibility with an application-level marker key. (SetManyAtomicAsyncitself does guarantee reader isolation tree-wide.) - Cross-tree atomicity of incidental multi-tree operations. A single
ILatticetree's per-operation atomic visibility does not automatically extend to operations that touch more than one tree as a side effect (e.g.MergeAsync,SnapshotAsync): those are LWW-convergent on the destination but not atomic-visible, so readers of both trees may observe the in-flight state. For an explicit all-or-nothing write spanning multiple trees, use the cross-tree saga primitive (SetManyAtomicAsync/ theBeginAtomicWritebuilder), which extends atomic visibility across every participating tree - see Cross-tree (multi-tree) atomic visibility above. - Cross-tree causality. The causal+ guarantees shipped in the replication package are scoped to single-tree writes; a multi-tree operation does not establish a causal edge between those trees on remote peers. Each tree converges independently under its own per-tree vector clock.