Durable Cursors
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at durable-cursors.md, and llms.txt lists every page.Durable cursors are server-side, checkpointed iterators for long-running key
scans and resumable range deletes. Unlike the stateless
ScanKeysAsync / ScanEntriesAsync scans - which are
bounded by LatticeOptions.MaxScanRetries and die with the client process -
and a one-shot DeleteRangeAsync, which reports no progress until it returns,
a cursor grain persists its position to Orleans storage
after every page. A new activation reads that checkpoint and continues exactly
where the previous one stopped, making export jobs, ETL pipelines, and
range-delete sweeps transparent to silo failovers, client restarts, and
topology changes (shard splits).
See API Reference - Stateful Cursors for the full method signatures, return types, error surface, and code examples. This document covers the design, grain lifecycle, and performance characteristics.
When to use a durable cursor
| Scenario | Recommendation |
|---|---|
| Short-lived scan, client stays up, few thousand keys | Stateless ScanKeysAsync / ScanEntriesAsync - lower overhead, no grain state |
| Export or migration that may span minutes | Durable cursor - survives failover, no client retry code needed |
| Range delete that must survive interruption | OpenDeleteRangeCursorAsync - tracks tombstoning progress across steps |
Topology under aggressive splitting, MaxScanRetries exhaustion |
Durable cursor - each step has its own retry budget; topology churn only affects one step at a time |
| Cursor ID must be handed off between processes or services | Durable cursor - any client that knows the opaque ID can resume |
| Multi-page scan that must observe the same saga-decision view across every page | Durable key/entry cursor opened with pointInTime: true - see Point-in-time cursors |
Architecture
Grain model
Each cursor is a single ILatticeCursorGrain activation keyed
{treeId}/{cursorId}, where cursorId is a server-assigned opaque GUID
returned by the Open*Async call. The grain is an internal implementation
detail - its interface is declared internal, so application code cannot
reference it. Callers interact exclusively through the ILattice facade.
sequenceDiagram
participant C as Client
participant L as ILattice
participant G as ILatticeCursorGrain
participant P as Persistent State
participant S as Shards
rect rgb(235,245,255)
Note over C,P: Open
C->>L: Open*Async(spec)
L->>G: OpenAsync(treeId, spec)
G->>P: WriteStateAsync() - phase=Open, spec frozen
L-->>C: cursorId (opaque GUID)
end
rect rgb(235,255,240)
Note over C,S: Next (repeated per page)
C->>L: Next*Async(cursorId, pageSize)
L->>G: Next*Async(pageSize)
G->>S: ScanKeysAsync(effStart, effEnd)
S-->>G: keys (split-ordering preserved)
G->>P: WriteStateAsync() - advance LastYieldedKey
G->>G: SlideTtlAsync() - refresh idle-TTL reminder
L-->>C: LatticeCursorKeysPage
end
rect rgb(255,245,235)
Note over C,P: Close
C->>L: CloseCursorAsync(cursorId)
L->>G: CloseAsync()
G->>G: UnregisterTtlAsync()
G->>G: release pins and snapshot baselines
G->>P: ClearStateAsync()
G->>G: DeactivateOnIdle()
L-->>C: ok
end
Cursor phase state machine
sequenceDiagram
participant C as Client
participant G as ILatticeCursorGrain
Note over G: NotStarted
C->>G: OpenAsync(treeId, spec)
Note over G: Open - spec and treeId persisted
loop Until HasMore = false
C->>G: Next*Async(pageSize)
G-->>C: LatticeCursorKeysPage / LatticeCursorEntriesPage (HasMore = true / false)
end
Note over G: Exhausted
opt Further Next*Async calls are idempotent
C->>G: Next*Async(pageSize)
G-->>C: empty page (HasMore = false)
end
C->>G: CloseAsync()
Note over G: Closed - state cleared, grain deactivates
Persisted state
LatticeCursorState is intentionally minimal - a silo restart only needs to
replay one page of work, and the checkpoint must be cheap to write on every
step.
| Field | Type | Purpose |
|---|---|---|
TreeId |
string |
Target tree grain key |
Spec |
LatticeCursorSpec |
Kind, start/end bounds, direction, PointInTime, ZeroObservableWrites (snapshot cursors), and the optional server-side Predicate of a *WherePredicateAsync cursor - frozen when the cursor opens |
Phase |
LatticeCursorPhase |
NotStarted / Open / Exhausted / Closed |
LastYieldedKey |
string? |
Last key returned or tombstoned. null before the first step. |
DeletedTotal |
int |
Cumulative tombstone count (delete-range cursors only) |
PointInTimeSnapshot |
Dictionary<Guid, TxStatus>? |
The per-tree transaction-registry snapshot captured at OpenAsync time. Persisted only for point-in-time cursors; null for live-mode cursors. |
SnapshotPinId |
Guid |
The pin id the cursor mints at open and registers with the tree's transaction registry so the decisions it captured are retained. Empty for live-mode cursors and for point-in-time cursors whose snapshot captured no decided saga (in-flight entries are never pinned). A snapshot cursor carries a minted id that is never registered, because it pins no registry decisions. |
SnapshotCoordinate |
LatticeSnapshotCoordinate? |
Tree-wide coordinate captured when a zero-observable-writes snapshot cursor opens: the pinned shard-map version (and, for a multi-shard tree, the map itself), the per-shard, per-partition WAL heads the frozen baselines were captured at, and the per-open baseline token that keys those baselines, so a reactivated cursor keeps serving the view it opened with. A snapshot cursor persists no PointInTimeSnapshot: its frozen baselines already fix saga visibility at the captured heads. null for non-snapshot cursors. |
SnapshotBaselinePersisted |
bool |
true once a snapshot cursor has durably flushed its per-shard frozen baselines. The baselines seed the transient snapshot leaves in memory at open and are flushed lazily only the first time a page reports more results (the cursor must now survive past page 1 across failover or eviction); a cursor that drains in a single page never sets it and can skip the durable baseline delete on close. false for non-snapshot cursors. |
Step sequence
Every Next*Async / DeleteRangeStepAsync call follows the same pattern:
sequenceDiagram
participant C as Client
participant L as ILattice
participant G as ILatticeCursorGrain
participant S as Shards
C->>L: NextKeysAsync(cursorId, pageSize)
L->>G: NextKeysAsync(pageSize)
G->>G: EnsureOpenFor(Keys)
G->>G: ComputeEffectiveRange() - advance bounds past LastYieldedKey
G->>S: ScanKeysAsync(effStart, effEnd, reverse)
S-->>G: streamed keys (up to pageSize, split-ordering preserved)
G->>G: WriteStateAsync() - checkpoint LastYieldedKey
G->>G: SlideTtlAsync() - refresh idle-TTL reminder
G-->>L: LatticeCursorKeysPage{Keys, HasMore}
L-->>C: LatticeCursorKeysPage
Effective range computation and resumption
Resumption after a silo failover requires no replay of prior pages - the grain
recomputes effStart / effEnd from the persisted LastYieldedKey and
issues the next bounded ScanKeysAsync / ScanEntriesAsync call:
- Forward scan:
effStart = LastYieldedKey + "\0"- the lexicographically first key strictly after the last yielded one.effEndis unchanged. - Reverse scan:
effEnd = LastYieldedKey- the last yielded key becomes the exclusive upper bound, so it is not re-yielded.effStartis unchanged.
A key yielded by step i is therefore never re-yielded by step i+1 or later, regardless of any shard splits that occur between steps.
Ordering under concurrent shard splits
Because each step delegates to the normal ScanKeysAsync /
ScanEntriesAsync / DeleteRangeAsync path, ordering preservation under
concurrent shard splits applies within each step: concurrent shard splits are reconciled via
in-line cursor injection into the k-way merge priority queue, bounded by
LatticeOptions.MaxScanRetries. See Shard Splitting
for the full reconciliation design.
Across steps, global ordering is preserved by the effective-range logic: the continuation bound strictly excludes every previously-yielded key, so a split that moves keys between steps is naturally handled by the next step's sharded range query.
Point-in-time cursors
A durable key or entry cursor opened with pointInTime: true extends the
tree-wide atomic-visibility guarantee that already covers single-call reads
(GetManyAsync, CountAsync, CountPerShardAsync) to a multi-page
enumeration. The cursor freezes the per-tree saga-decision view at open time
and every subsequent page reads against that frozen view - a
SetManyAtomicAsync saga that commits between two pages is observed
identically on every page (either all of its keys, or none), never as a torn
pre/post split.
var cursorId = await tree.OpenEntryCursorAsync(
startInclusive: null,
endExclusive: null,
reverse: false,
pointInTime: true);
try
{
while (true)
{
var page = await tree.NextEntriesAsync(cursorId, pageSize: 500);
foreach (var (k, v) in page.Entries)
{
// Every page sees the same in-flight-saga view captured at
// OpenEntryCursorAsync time. Sagas that commit between pages
// are atomically visible across the cursor.
}
if (!page.HasMore) break;
}
}
finally
{
await tree.CloseCursorAsync(cursorId);
}
How it works
- Capture at open. The open takes one snapshot of the tree's
saga-decision registry and persists the resulting
Dictionary<Guid, TxStatus>as the cursor'sPointInTimeSnapshot(see Persisted state). Once the tree's registry is sharded (sagas minted whileLatticeOptions.TxRegistryShardCountis above1), the snapshot unions every registry shard the tree has written to plus the legacy registry, and re-reads the shards' revisions to confirm the union is a consistent cut, settling after a bounded number of attempts for the latest union - which is still exact for every individual saga. - Pin retention. If the snapshot recorded a decision for any txid
(entries the snapshot read as
InFlightare excluded - they have no tombstone to protect yet), the cursor mints a pin id and pins those txids on each registry shard that owns at least one of them, so the registry retains every observed decision (including one it has already tombstoned) for the cursor's lifetime. The pin is requested forLatticeOptions.MaxCursorSnapshotPinTtl(default 7 days), which is also the registry's cap on the pin's lifetime (a positive value shorter thanLatticeOptions.TxDecisionRetentionis raised to it); a non-positive value disables the cap, so the pin does not expire and is released only when the cursor closes or its idle-TTL reminder fires. The pin id is persisted as the cursor'sSnapshotPinId. Entries the snapshot read asIndeterminateare pinned too: an indeterminate reading is an aged-out tombstone whose decision row is still stored, which is exactly the row a pin exists to protect, and because the retention mask is pin-aware the pin also restores the recorded outcome to this cursor. - Per-step replay. Every
NextKeysAsync/NextEntriesAsyncre-enters the captured snapshot viaLatticeRegistrySnapshotContext.BeginScope(...)before fanning out to leaves. Every leaf RPC for the step reads the same registry view- identical to the steady-state behaviour of
GetManyAsync/CountAsync/CountPerShardAsync, just held across multiple pages. Note what carries the per-step guarantee here, because it is easy to credit to the wrong mechanism: inside the scopeBPlusLeafGrain.ResolvePendingStatusAsyncanswers from the captured dictionary and returns without contacting the registry, so a step's per-key readings were fixed at open and no amount of registry-side pruning can move them. The pin in step 2 is not what makes that true and could not be - there is no registry lookup on this path for it to affect. The pin matters for the reads that do reach the registry: an unscoped read of the same keys while the cursor is open, and (per step 2) an entry the snapshot captured asIndeterminate.
- identical to the steady-state behaviour of
- Pin refresh. Each step also refreshes the pin on every registry
shard that holds it, sliding the registry-side TTL; the refresh succeeds
only while every one of those shards still holds the pin. A cursor that
pages actively never runs out the pin TTL; a stalled cursor that misses
the slide will eventually be reaped by the registry, and its next step
throws
LatticeCursorSnapshotExpiredExceptionand closes the cursor. - Release on close / TTL expiry.
CloseCursorAsyncand the cursor's own idle-TTL reminder both release the pin from every registry shard, freeing the retained decisions so registry tombstone-prune can resume.
Caps and failure modes
Three independent caps bound the registry footprint a forgotten or stalled point-in-time cursor can occupy:
| Cap | Default | Effect |
|---|---|---|
LatticeOptions.CursorIdleTtl |
48 h | Cursor-grain idle reminder releases the pin on inactivity. |
LatticeOptions.MaxCursorSnapshotPinTtl |
7 d | Registry-side hard cap on a single pin's lifetime, never shorter than TxDecisionRetention (a shorter positive value is raised to it). A live cursor slides this on every Next*Async; a stalled cursor that misses the slide surfaces LatticeCursorSnapshotExpiredException on its next call and the cursor must be reopened. A non-positive value (for example Timeout.InfiniteTimeSpan) disables the cap: the pin never expires and is released only by close or the idle-TTL reminder. |
LatticeOptions.MaxPinnedSagaDecisions |
100 000 | Cap on the union of saga decisions pinned by every live point-in-time cursor on a tree's saga-decision registry; once the registry is sharded, each registry shard enforces it over the decisions it owns. Opening a point-in-time cursor (OpenKeyCursorAsync / OpenEntryCursorAsync with pointInTime: true) throws LatticeCursorRegistryPinExhaustedException when accepting the new snapshot would breach the cap; existing pinned cursors continue paging. |
| Condition | Exception |
|---|---|
Opening a point-in-time cursor would push the tree's pinned-decision count past MaxPinnedSagaDecisions |
LatticeCursorRegistryPinExhaustedException |
NextKeysAsync / NextEntriesAsync on a point-in-time cursor whose pin has been evicted (TTL elapsed or registry reaper ran) |
LatticeCursorSnapshotExpiredException |
A delete-range cursor cannot be opened in point-in-time mode:
OpenDeleteRangeCursorAsync exposes no pointInTime option at all, because
range deletes are mutations rather than snapshot reads. Point-in-time pinning
applies only to key and entry cursors (OpenKeyCursorAsync /
OpenEntryCursorAsync and their WherePredicate variants).
Cost vs. live mode
Live-mode and point-in-time cursors share the same per-step checkpoint and shard fan-out cost. Point-in-time mode adds:
- One saga-decision registry snapshot read and one pin registration at open (the registration is skipped when the snapshot captured no decided saga). On a sharded registry the read fans out across the tree's registry shards and the registration goes to each shard that owns a pinned decision.
- One pin refresh per step, issued before the step's scan, on each shard holding the pin.
- One pin release at close or idle-TTL expiry, on every registry shard.
The persisted PointInTimeSnapshot adds one dictionary entry per
in-flight or recently-completed saga at open time to the cursor's
storage row; it does not grow as the cursor pages.
Self-cleanup (idle-TTL reminder)
To prevent cursor state leaking when a client forgets CloseCursorAsync,
every cursor grain registers a sliding idle-TTL reminder (cursor-ttl) after
every successful call. If the reminder fires with no intervening activity, the
grain clears its persisted state and deactivates.
sequenceDiagram
participant R as Orleans Reminders
participant G as ILatticeCursorGrain
Note over G: After every Open / Next / Step
G->>R: RegisterOrUpdateReminder("cursor-ttl", dueTime=CursorIdleTtl, period=CursorIdleTtl)
Note over R: CursorIdleTtl elapses with no calls
R->>G: ReceiveReminder("cursor-ttl")
G->>G: OnTtlExpiredAsync() -> state.ClearStateAsync()
G->>R: UnregisterReminder("cursor-ttl")
G->>G: DeactivateOnIdle()
LatticeCursorGrain inherits this machinery from the internal TtlGrain
abstract base class, which also backs AtomicWriteGrain. Each grain overrides
TtlReminderName, ResolveTtl, and OnTtlExpiredAsync independently -
CursorIdleTtl and AtomicWriteRetention are separate options and do not
share a value.
Configuration:
// Per-tree
siloBuilder.ConfigureLattice("my-tree", o =>
o.CursorIdleTtl = TimeSpan.FromHours(6));
// Global default
siloBuilder.ConfigureLattice(o =>
o.CursorIdleTtl = TimeSpan.FromHours(6));
Set CursorIdleTtl = Timeout.InfiniteTimeSpan to disable automatic cleanup.
The minimum effective interval is 1 minute (Orleans reminder granularity);
smaller values are clamped to that floor.
Performance characteristics
Per-step overhead
Every Next*Async / DeleteRangeStepAsync call incurs two additional I/O
round-trips above a direct stateless scan call:
| Cost component | Magnitude | Notes |
|---|---|---|
WriteStateAsync - checkpoint |
1 x storage write per step | Serialises LatticeCursorState (< 10 KB for a live-mode cursor; see Grain state size). On memory provider: negligible. On Azure Table / SQL: ~1-5 ms. |
RegisterOrUpdateReminder - TTL slide |
1 x reminder-table write per step | ~1-5 ms round-trip. See debounce below. |
| Extra grain round-trip | +1 Orleans call per step | ILatticeCursorGrain sits between ILattice and the shard fan-out. Typically < 1 ms on a local cluster. |
| Shard fan-out | Same as ScanKeysAsync / ScanEntriesAsync |
Each step is a normal sharded scan - no additional shard calls. |
Total per-step overhead: ~2-10 ms, dominated by the storage provider round-trip.
Large-export scenario
For a 10 million key export with pageSize = 500 (20 000 steps):
| Metric | Value |
|---|---|
| Steps | 20 000 |
| Checkpoint writes | 20 000 |
| Reminder slides | 20 000 |
| Extra wall-clock time at 2 ms/step | ~= 40 s |
| Extra wall-clock time at 5 ms/step | ~= 100 s |
This overhead is typically small relative to the actual I/O cost of streaming 10 M keys across the network, but it is not zero.
Reducing reminder write frequency
The internal TtlGrain base exposes a virtual SlideDebounce property
(default TimeSpan.Zero, meaning slide on every call). Overriding it in a
subclass throttles RegisterOrUpdateReminder calls to at most one per
interval, accepting a slightly stale TTL window in exchange for lower
reminder-table pressure. This is an internal extension point; it is not
surfaced on LatticeOptions.
The simpler alternative is to increase pageSize: halving the step
count halves the reminder and checkpoint write count proportionally.
Grain state size
LatticeCursorState is intentionally minimal. For a live-mode cursor, even
with a 4 KB LastYieldedKey and a 1 KB spec, the checkpoint is < 10 KB. A
point-in-time cursor adds its captured decision map (one entry per in-flight
or recently completed saga at open), and a snapshot cursor adds its
coordinate - including, on a multi-shard tree, the pinned routing map.
Aggregate reminder-table storage for a typical fleet of concurrent cursors
is negligible.
Concurrent cursors
Each cursor grain is a single-threaded Orleans activation. Concurrent pages from different cursors run in parallel with no cross-cursor coordination - N concurrent cursors are equivalent in throughput to N independent stateless scans, plus the per-step overhead per cursor.
Stateless vs. durable - decision guide
| Dimension | Stateless (ScanKeysAsync / ScanEntriesAsync) |
Durable cursor (live mode) | Durable cursor (point-in-time) |
|---|---|---|---|
| Survives silo failover | No - stream terminates | Yes - resumes from checkpoint | Yes - resumes from checkpoint and retained registry pin |
| Survives client restart | No | Yes - cursor ID is the resume token | Yes |
| Caller retry code needed | Required for robustness under splits | None needed | None needed |
| Per-page overhead | Zero | ~2-10 ms (checkpoint + reminder slide) | ~2-10 ms (adds one registry refresh RPC) |
| Ordering under splits | Per-call reconciliation (see Shard Splitting) | Per-step reconciliation | Per-step reconciliation |
| Atomic visibility | Scan-lifetime tree-wide | Per-step tree-wide; not preserved across pages | Cursor-lifetime tree-wide - identical saga view on every page |
| Max scan duration | Bounded by MaxScanRetries |
Unbounded - each step has its own budget | Unbounded while it pages - each step slides the registry pin; see Caps and failure modes for what ends an idle one |
| Idle cleanup | No state to clean up | Automatic via idle-TTL reminder | Idle-TTL reminder + registry-pin TTL |
| Cursor ID transferable across processes | No | Yes | Yes |
| Available for range delete | No | Yes via OpenDeleteRangeCursorAsync |
No - range deletes are mutations, not snapshot reads |
| Best for | Interactive queries, short scans | Long exports, ETL, background sweeps | Long exports that must observe a single saga-decision view across every page |
See also
- API Reference - Stateful Cursors - full method signatures, return types, error surface, and code examples.
- Consistency - Enumeration - the formal consistency classification of live-mode and point-in-time cursor steps.
- Atomic Writes - the saga primitive whose visibility flip point-in-time cursors freeze for the cursor's lifetime.
- Shard Splitting - ordering preservation under concurrent topology changes (applied within each cursor step).
- Configuration -
CursorIdleTtl,MaxCursorSnapshotPinTtl,MaxPinnedSagaDecisions,MaxScanRetries, and other tunables.