Table of Contents

Durable Cursors

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at durable-cursors.md, and llms.txt lists every page.

Durable cursors are server-side, checkpointed iterators for long-running key scans and resumable range deletes. Unlike the stateless ScanKeysAsync / ScanEntriesAsync scans - which are bounded by LatticeOptions.MaxScanRetries and die with the client process - and a one-shot DeleteRangeAsync, which reports no progress until it returns, a cursor grain persists its position to Orleans storage after every page. A new activation reads that checkpoint and continues exactly where the previous one stopped, making export jobs, ETL pipelines, and range-delete sweeps transparent to silo failovers, client restarts, and topology changes (shard splits).

See API Reference - Stateful Cursors for the full method signatures, return types, error surface, and code examples. This document covers the design, grain lifecycle, and performance characteristics.

When to use a durable cursor

Scenario Recommendation
Short-lived scan, client stays up, few thousand keys Stateless ScanKeysAsync / ScanEntriesAsync - lower overhead, no grain state
Export or migration that may span minutes Durable cursor - survives failover, no client retry code needed
Range delete that must survive interruption OpenDeleteRangeCursorAsync - tracks tombstoning progress across steps
Topology under aggressive splitting, MaxScanRetries exhaustion Durable cursor - each step has its own retry budget; topology churn only affects one step at a time
Cursor ID must be handed off between processes or services Durable cursor - any client that knows the opaque ID can resume
Multi-page scan that must observe the same saga-decision view across every page Durable key/entry cursor opened with pointInTime: true - see Point-in-time cursors

Architecture

Grain model

Each cursor is a single ILatticeCursorGrain activation keyed {treeId}/{cursorId}, where cursorId is a server-assigned opaque GUID returned by the Open*Async call. The grain is an internal implementation detail - its interface is declared internal, so application code cannot reference it. Callers interact exclusively through the ILattice facade.

sequenceDiagram
    participant C as Client
    participant L as ILattice
    participant G as ILatticeCursorGrain
    participant P as Persistent State
    participant S as Shards

    rect rgb(235,245,255)
        Note over C,P: Open
        C->>L: Open*Async(spec)
        L->>G: OpenAsync(treeId, spec)
        G->>P: WriteStateAsync() - phase=Open, spec frozen
        L-->>C: cursorId (opaque GUID)
    end

    rect rgb(235,255,240)
        Note over C,S: Next (repeated per page)
        C->>L: Next*Async(cursorId, pageSize)
        L->>G: Next*Async(pageSize)
        G->>S: ScanKeysAsync(effStart, effEnd)
        S-->>G: keys (split-ordering preserved)
        G->>P: WriteStateAsync() - advance LastYieldedKey
        G->>G: SlideTtlAsync() - refresh idle-TTL reminder
        L-->>C: LatticeCursorKeysPage
    end

    rect rgb(255,245,235)
        Note over C,P: Close
        C->>L: CloseCursorAsync(cursorId)
        L->>G: CloseAsync()
        G->>G: UnregisterTtlAsync()
        G->>G: release pins and snapshot baselines
        G->>P: ClearStateAsync()
        G->>G: DeactivateOnIdle()
        L-->>C: ok
    end

Cursor phase state machine

sequenceDiagram
    participant C as Client
    participant G as ILatticeCursorGrain

    Note over G: NotStarted

    C->>G: OpenAsync(treeId, spec)
    Note over G: Open - spec and treeId persisted

    loop Until HasMore = false
        C->>G: Next*Async(pageSize)
        G-->>C: LatticeCursorKeysPage / LatticeCursorEntriesPage (HasMore = true / false)
    end

    Note over G: Exhausted

    opt Further Next*Async calls are idempotent
        C->>G: Next*Async(pageSize)
        G-->>C: empty page (HasMore = false)
    end

    C->>G: CloseAsync()
    Note over G: Closed - state cleared, grain deactivates

Persisted state

LatticeCursorState is intentionally minimal - a silo restart only needs to replay one page of work, and the checkpoint must be cheap to write on every step.

Field Type Purpose
TreeId string Target tree grain key
Spec LatticeCursorSpec Kind, start/end bounds, direction, PointInTime, ZeroObservableWrites (snapshot cursors), and the optional server-side Predicate of a *WherePredicateAsync cursor - frozen when the cursor opens
Phase LatticeCursorPhase NotStarted / Open / Exhausted / Closed
LastYieldedKey string? Last key returned or tombstoned. null before the first step.
DeletedTotal int Cumulative tombstone count (delete-range cursors only)
PointInTimeSnapshot Dictionary<Guid, TxStatus>? The per-tree transaction-registry snapshot captured at OpenAsync time. Persisted only for point-in-time cursors; null for live-mode cursors.
SnapshotPinId Guid The pin id the cursor mints at open and registers with the tree's transaction registry so the decisions it captured are retained. Empty for live-mode cursors and for point-in-time cursors whose snapshot captured no decided saga (in-flight entries are never pinned). A snapshot cursor carries a minted id that is never registered, because it pins no registry decisions.
SnapshotCoordinate LatticeSnapshotCoordinate? Tree-wide coordinate captured when a zero-observable-writes snapshot cursor opens: the pinned shard-map version (and, for a multi-shard tree, the map itself), the per-shard, per-partition WAL heads the frozen baselines were captured at, and the per-open baseline token that keys those baselines, so a reactivated cursor keeps serving the view it opened with. A snapshot cursor persists no PointInTimeSnapshot: its frozen baselines already fix saga visibility at the captured heads. null for non-snapshot cursors.
SnapshotBaselinePersisted bool true once a snapshot cursor has durably flushed its per-shard frozen baselines. The baselines seed the transient snapshot leaves in memory at open and are flushed lazily only the first time a page reports more results (the cursor must now survive past page 1 across failover or eviction); a cursor that drains in a single page never sets it and can skip the durable baseline delete on close. false for non-snapshot cursors.

Step sequence

Every Next*Async / DeleteRangeStepAsync call follows the same pattern:

sequenceDiagram
    participant C as Client
    participant L as ILattice
    participant G as ILatticeCursorGrain
    participant S as Shards

    C->>L: NextKeysAsync(cursorId, pageSize)
    L->>G: NextKeysAsync(pageSize)
    G->>G: EnsureOpenFor(Keys)
    G->>G: ComputeEffectiveRange() - advance bounds past LastYieldedKey
    G->>S: ScanKeysAsync(effStart, effEnd, reverse)
    S-->>G: streamed keys (up to pageSize, split-ordering preserved)
    G->>G: WriteStateAsync() - checkpoint LastYieldedKey
    G->>G: SlideTtlAsync() - refresh idle-TTL reminder
    G-->>L: LatticeCursorKeysPage{Keys, HasMore}
    L-->>C: LatticeCursorKeysPage

Effective range computation and resumption

Resumption after a silo failover requires no replay of prior pages - the grain recomputes effStart / effEnd from the persisted LastYieldedKey and issues the next bounded ScanKeysAsync / ScanEntriesAsync call:

  • Forward scan: effStart = LastYieldedKey + "\0" - the lexicographically first key strictly after the last yielded one. effEnd is unchanged.
  • Reverse scan: effEnd = LastYieldedKey - the last yielded key becomes the exclusive upper bound, so it is not re-yielded. effStart is unchanged.

A key yielded by step i is therefore never re-yielded by step i+1 or later, regardless of any shard splits that occur between steps.

Ordering under concurrent shard splits

Because each step delegates to the normal ScanKeysAsync / ScanEntriesAsync / DeleteRangeAsync path, ordering preservation under concurrent shard splits applies within each step: concurrent shard splits are reconciled via in-line cursor injection into the k-way merge priority queue, bounded by LatticeOptions.MaxScanRetries. See Shard Splitting for the full reconciliation design.

Across steps, global ordering is preserved by the effective-range logic: the continuation bound strictly excludes every previously-yielded key, so a split that moves keys between steps is naturally handled by the next step's sharded range query.

Point-in-time cursors

A durable key or entry cursor opened with pointInTime: true extends the tree-wide atomic-visibility guarantee that already covers single-call reads (GetManyAsync, CountAsync, CountPerShardAsync) to a multi-page enumeration. The cursor freezes the per-tree saga-decision view at open time and every subsequent page reads against that frozen view - a SetManyAtomicAsync saga that commits between two pages is observed identically on every page (either all of its keys, or none), never as a torn pre/post split.

var cursorId = await tree.OpenEntryCursorAsync(
    startInclusive: null,
    endExclusive: null,
    reverse: false,
    pointInTime: true);
try
{
    while (true)
    {
        var page = await tree.NextEntriesAsync(cursorId, pageSize: 500);
        foreach (var (k, v) in page.Entries)
        {
            // Every page sees the same in-flight-saga view captured at
            // OpenEntryCursorAsync time. Sagas that commit between pages
            // are atomically visible across the cursor.
        }
        if (!page.HasMore) break;
    }
}
finally
{
    await tree.CloseCursorAsync(cursorId);
}

How it works

  1. Capture at open. The open takes one snapshot of the tree's saga-decision registry and persists the resulting Dictionary<Guid, TxStatus> as the cursor's PointInTimeSnapshot (see Persisted state). Once the tree's registry is sharded (sagas minted while LatticeOptions.TxRegistryShardCount is above 1), the snapshot unions every registry shard the tree has written to plus the legacy registry, and re-reads the shards' revisions to confirm the union is a consistent cut, settling after a bounded number of attempts for the latest union - which is still exact for every individual saga.
  2. Pin retention. If the snapshot recorded a decision for any txid (entries the snapshot read as InFlight are excluded - they have no tombstone to protect yet), the cursor mints a pin id and pins those txids on each registry shard that owns at least one of them, so the registry retains every observed decision (including one it has already tombstoned) for the cursor's lifetime. The pin is requested for LatticeOptions.MaxCursorSnapshotPinTtl (default 7 days), which is also the registry's cap on the pin's lifetime (a positive value shorter than LatticeOptions.TxDecisionRetention is raised to it); a non-positive value disables the cap, so the pin does not expire and is released only when the cursor closes or its idle-TTL reminder fires. The pin id is persisted as the cursor's SnapshotPinId. Entries the snapshot read as Indeterminate are pinned too: an indeterminate reading is an aged-out tombstone whose decision row is still stored, which is exactly the row a pin exists to protect, and because the retention mask is pin-aware the pin also restores the recorded outcome to this cursor.
  3. Per-step replay. Every NextKeysAsync / NextEntriesAsync re-enters the captured snapshot via LatticeRegistrySnapshotContext.BeginScope(...) before fanning out to leaves. Every leaf RPC for the step reads the same registry view
    • identical to the steady-state behaviour of GetManyAsync / CountAsync / CountPerShardAsync, just held across multiple pages. Note what carries the per-step guarantee here, because it is easy to credit to the wrong mechanism: inside the scope BPlusLeafGrain.ResolvePendingStatusAsync answers from the captured dictionary and returns without contacting the registry, so a step's per-key readings were fixed at open and no amount of registry-side pruning can move them. The pin in step 2 is not what makes that true and could not be - there is no registry lookup on this path for it to affect. The pin matters for the reads that do reach the registry: an unscoped read of the same keys while the cursor is open, and (per step 2) an entry the snapshot captured as Indeterminate.
  4. Pin refresh. Each step also refreshes the pin on every registry shard that holds it, sliding the registry-side TTL; the refresh succeeds only while every one of those shards still holds the pin. A cursor that pages actively never runs out the pin TTL; a stalled cursor that misses the slide will eventually be reaped by the registry, and its next step throws LatticeCursorSnapshotExpiredException and closes the cursor.
  5. Release on close / TTL expiry. CloseCursorAsync and the cursor's own idle-TTL reminder both release the pin from every registry shard, freeing the retained decisions so registry tombstone-prune can resume.

Caps and failure modes

Three independent caps bound the registry footprint a forgotten or stalled point-in-time cursor can occupy:

Cap Default Effect
LatticeOptions.CursorIdleTtl 48 h Cursor-grain idle reminder releases the pin on inactivity.
LatticeOptions.MaxCursorSnapshotPinTtl 7 d Registry-side hard cap on a single pin's lifetime, never shorter than TxDecisionRetention (a shorter positive value is raised to it). A live cursor slides this on every Next*Async; a stalled cursor that misses the slide surfaces LatticeCursorSnapshotExpiredException on its next call and the cursor must be reopened. A non-positive value (for example Timeout.InfiniteTimeSpan) disables the cap: the pin never expires and is released only by close or the idle-TTL reminder.
LatticeOptions.MaxPinnedSagaDecisions 100 000 Cap on the union of saga decisions pinned by every live point-in-time cursor on a tree's saga-decision registry; once the registry is sharded, each registry shard enforces it over the decisions it owns. Opening a point-in-time cursor (OpenKeyCursorAsync / OpenEntryCursorAsync with pointInTime: true) throws LatticeCursorRegistryPinExhaustedException when accepting the new snapshot would breach the cap; existing pinned cursors continue paging.
Condition Exception
Opening a point-in-time cursor would push the tree's pinned-decision count past MaxPinnedSagaDecisions LatticeCursorRegistryPinExhaustedException
NextKeysAsync / NextEntriesAsync on a point-in-time cursor whose pin has been evicted (TTL elapsed or registry reaper ran) LatticeCursorSnapshotExpiredException

A delete-range cursor cannot be opened in point-in-time mode: OpenDeleteRangeCursorAsync exposes no pointInTime option at all, because range deletes are mutations rather than snapshot reads. Point-in-time pinning applies only to key and entry cursors (OpenKeyCursorAsync / OpenEntryCursorAsync and their WherePredicate variants).

Cost vs. live mode

Live-mode and point-in-time cursors share the same per-step checkpoint and shard fan-out cost. Point-in-time mode adds:

  • One saga-decision registry snapshot read and one pin registration at open (the registration is skipped when the snapshot captured no decided saga). On a sharded registry the read fans out across the tree's registry shards and the registration goes to each shard that owns a pinned decision.
  • One pin refresh per step, issued before the step's scan, on each shard holding the pin.
  • One pin release at close or idle-TTL expiry, on every registry shard.

The persisted PointInTimeSnapshot adds one dictionary entry per in-flight or recently-completed saga at open time to the cursor's storage row; it does not grow as the cursor pages.

Self-cleanup (idle-TTL reminder)

To prevent cursor state leaking when a client forgets CloseCursorAsync, every cursor grain registers a sliding idle-TTL reminder (cursor-ttl) after every successful call. If the reminder fires with no intervening activity, the grain clears its persisted state and deactivates.

sequenceDiagram
    participant R as Orleans Reminders
    participant G as ILatticeCursorGrain

    Note over G: After every Open / Next / Step
    G->>R: RegisterOrUpdateReminder("cursor-ttl", dueTime=CursorIdleTtl, period=CursorIdleTtl)

    Note over R: CursorIdleTtl elapses with no calls
    R->>G: ReceiveReminder("cursor-ttl")
    G->>G: OnTtlExpiredAsync() -> state.ClearStateAsync()
    G->>R: UnregisterReminder("cursor-ttl")
    G->>G: DeactivateOnIdle()

LatticeCursorGrain inherits this machinery from the internal TtlGrain abstract base class, which also backs AtomicWriteGrain. Each grain overrides TtlReminderName, ResolveTtl, and OnTtlExpiredAsync independently - CursorIdleTtl and AtomicWriteRetention are separate options and do not share a value.

Configuration:

// Per-tree
siloBuilder.ConfigureLattice("my-tree", o =>
    o.CursorIdleTtl = TimeSpan.FromHours(6));

// Global default
siloBuilder.ConfigureLattice(o =>
    o.CursorIdleTtl = TimeSpan.FromHours(6));

Set CursorIdleTtl = Timeout.InfiniteTimeSpan to disable automatic cleanup. The minimum effective interval is 1 minute (Orleans reminder granularity); smaller values are clamped to that floor.

Performance characteristics

Per-step overhead

Every Next*Async / DeleteRangeStepAsync call incurs two additional I/O round-trips above a direct stateless scan call:

Cost component Magnitude Notes
WriteStateAsync - checkpoint 1 x storage write per step Serialises LatticeCursorState (< 10 KB for a live-mode cursor; see Grain state size). On memory provider: negligible. On Azure Table / SQL: ~1-5 ms.
RegisterOrUpdateReminder - TTL slide 1 x reminder-table write per step ~1-5 ms round-trip. See debounce below.
Extra grain round-trip +1 Orleans call per step ILatticeCursorGrain sits between ILattice and the shard fan-out. Typically < 1 ms on a local cluster.
Shard fan-out Same as ScanKeysAsync / ScanEntriesAsync Each step is a normal sharded scan - no additional shard calls.

Total per-step overhead: ~2-10 ms, dominated by the storage provider round-trip.

Large-export scenario

For a 10 million key export with pageSize = 500 (20 000 steps):

Metric Value
Steps 20 000
Checkpoint writes 20 000
Reminder slides 20 000
Extra wall-clock time at 2 ms/step ~= 40 s
Extra wall-clock time at 5 ms/step ~= 100 s

This overhead is typically small relative to the actual I/O cost of streaming 10 M keys across the network, but it is not zero.

Reducing reminder write frequency

The internal TtlGrain base exposes a virtual SlideDebounce property (default TimeSpan.Zero, meaning slide on every call). Overriding it in a subclass throttles RegisterOrUpdateReminder calls to at most one per interval, accepting a slightly stale TTL window in exchange for lower reminder-table pressure. This is an internal extension point; it is not surfaced on LatticeOptions.

The simpler alternative is to increase pageSize: halving the step count halves the reminder and checkpoint write count proportionally.

Grain state size

LatticeCursorState is intentionally minimal. For a live-mode cursor, even with a 4 KB LastYieldedKey and a 1 KB spec, the checkpoint is < 10 KB. A point-in-time cursor adds its captured decision map (one entry per in-flight or recently completed saga at open), and a snapshot cursor adds its coordinate - including, on a multi-shard tree, the pinned routing map. Aggregate reminder-table storage for a typical fleet of concurrent cursors is negligible.

Concurrent cursors

Each cursor grain is a single-threaded Orleans activation. Concurrent pages from different cursors run in parallel with no cross-cursor coordination - N concurrent cursors are equivalent in throughput to N independent stateless scans, plus the per-step overhead per cursor.

Stateless vs. durable - decision guide

Dimension Stateless (ScanKeysAsync / ScanEntriesAsync) Durable cursor (live mode) Durable cursor (point-in-time)
Survives silo failover No - stream terminates Yes - resumes from checkpoint Yes - resumes from checkpoint and retained registry pin
Survives client restart No Yes - cursor ID is the resume token Yes
Caller retry code needed Required for robustness under splits None needed None needed
Per-page overhead Zero ~2-10 ms (checkpoint + reminder slide) ~2-10 ms (adds one registry refresh RPC)
Ordering under splits Per-call reconciliation (see Shard Splitting) Per-step reconciliation Per-step reconciliation
Atomic visibility Scan-lifetime tree-wide Per-step tree-wide; not preserved across pages Cursor-lifetime tree-wide - identical saga view on every page
Max scan duration Bounded by MaxScanRetries Unbounded - each step has its own budget Unbounded while it pages - each step slides the registry pin; see Caps and failure modes for what ends an idle one
Idle cleanup No state to clean up Automatic via idle-TTL reminder Idle-TTL reminder + registry-pin TTL
Cursor ID transferable across processes No Yes Yes
Available for range delete No Yes via OpenDeleteRangeCursorAsync No - range deletes are mutations, not snapshot reads
Best for Interactive queries, short scans Long exports, ETL, background sweeps Long exports that must observe a single saga-decision view across every page

See also

  • API Reference - Stateful Cursors - full method signatures, return types, error surface, and code examples.
  • Consistency - Enumeration - the formal consistency classification of live-mode and point-in-time cursor steps.
  • Atomic Writes - the saga primitive whose visibility flip point-in-time cursors freeze for the cursor's lifetime.
  • Shard Splitting - ordering preservation under concurrent topology changes (applied within each cursor step).
  • Configuration - CursorIdleTtl, MaxCursorSnapshotPinTtl, MaxPinnedSagaDecisions, MaxScanRetries, and other tunables.