---
title: "Snapshot cursors (zero observable writes)"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/snapshot-cursors.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/snapshot-cursors.md"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Snapshot cursors (zero observable writes)

Part of the [Orleans.Lattice documentation](architecture.md).

`ILattice.OpenSnapshotKeyCursorAsync(...)` and
`ILattice.OpenSnapshotEntryCursorAsync(...)` open a **strict
snapshot-isolation** cursor: every page returned by the cursor reflects
the tree state captured at open time, and no concurrent write -
foreground `SetAsync` / `DeleteAsync`, saga `SetManyAtomicAsync`, range
delete, or replication apply - is ever visible to the cursor for the
remainder of its lifetime.

Snapshot cursors compose with the live cursors documented in
[Durable Cursors](durable-cursors.md): the cursor ID, pagination
contract, and `CloseCursorAsync` lifecycle are identical. Only the
visibility semantics differ.

## When to use a snapshot cursor

| Scenario | Cursor flavour |
|---|---|
| Long-running export, audit, or report that must reflect a single instant | Snapshot (`OpenSnapshot*CursorAsync`) |
| Pagination where the latest writes should appear on later pages | Live (`OpenKeyCursorAsync` / `OpenEntryCursorAsync`) |
| Stable saga-decision view across pages, but mid-page foreground writes are fine | Point-in-time live (`pointInTime: true`) - see [api.md](api/ilattice-1.md#point-in-time-cursors) |
| Strict isolation against every concurrent write *and* every concurrent saga | Snapshot |

## How the snapshot is captured

`OpenSnapshot*CursorAsync` performs these deterministic steps before
returning the cursor ID:

1. **Routing capture.** The tree's routing map is re-read from the
   registry (not taken from a cached copy) and its version pinned on the
   coordinate - with the full map as well when the tree has more than one
   physical shard - so all paging fan-outs target the same shard layout.
2. **Per-shard frozen-baseline capture.** Every shard root walks its
   leaf chain through `IShardRootGrain.CaptureSnapshotBaselineAsync`,
   freezing each `BPlusLeafGrain`'s committed projection and folding its
   own `(leaf_frontier, capturedHead]` WAL tail exactly once (CRDT folds
   are not idempotent, so each record is applied to a single leaf a
   single time). The per-leaf results are unioned into one fully
   materialised per-shard projection and **seeded in memory** into the
   transient snapshot leaf for that shard, keyed by
   `{physicalTreeId}/{shardIndex}/{baselineToken:N}` - the physical tree
   the capture resolved, so a later alias cutover of the logical tree
   still reaches the seeded leaf. The seed is *not* written
   durably at capture - see [Lazy baseline persistence](#lazy-baseline-persistence)
   below. The uniform `capturedHead` (each shard's per-partition WAL
   head, read after every leaf has frozen) is the bound recorded on the
   coordinate. To avoid blocking every shard root at once (each capture
   runs on the shard root's non-reentrant turn, contending with
   replication applies and reads), the fan-out is bounded to
   [`LatticeOptions.MaxConcurrentSnapshotCaptures`](configuration/options-reference-2.md#maxconcurrentsnapshotcaptures)
   shard captures at a time (default 4); the open then proceeds in waves.
   The captured baseline is identical under any cap - only the dispatch
   schedule changes. Within a single shard, the fold pass is itself
   fanned out over a sliding window of
   [`LatticeOptions.MaxConcurrentSnapshotBaselineFolds`](configuration/options-reference-2.md#maxconcurrentsnapshotbaselinefolds)
   leaves (default 4) to cut that shard's hold; the freeze pass stays
   sequential because the chain is discovered one sibling hop at a time.
   Fold results are consumed in strict leaf-chain order, so the baseline
   is byte-identical to a serial fold. The whole capture is additionally
   covered by the hard end-to-end stall ceiling
   ([`MaxScanPageStallDuration`](configuration/options-reference-2.md#maxscanpagestallduration)),
   so a shard root can never be held indefinitely by an unresponsive
   leaf: the capture is abandoned instead, which is safe because it
   performs no observable writes until the baseline is seeded.
3. **Saga-decision snapshot.** The tree's transaction registry is read, as
   for a point-in-time cursor, but its decisions are not carried into the
   cursor: the coordinate records only a diagnostic registry clock, which
   is currently always zero, and a transport failure of that read does not
   fail the open (any other registry fault does). Saga
   visibility is fixed by the frozen baseline itself - a batch whose commit
   terminal lands after the captured head stays pending, and so invisible,
   on every leaf it touched.

The captured values are packaged as a
`LatticeSnapshotCoordinate` (Orleans-serializable; alias `ol.lsc`) and
persisted on `LatticeCursorState.SnapshotCoordinate`. The coordinate
carries a fresh per-open `SnapshotBaselineToken` that identifies the
durable baseline rows. The coordinate is deterministic - replaying with
the same coordinate yields the same page sequence, even after silo
failover.

> **Why a frozen baseline rather than open-time WAL replay?** A snapshot
> scan is an ephemeral reader, not a registered WAL cursor. Earlier
> versions replayed each shard's WAL from offset `0` to the captured head
> at every page; once `LatticeWalGc` trimmed the committed prefix the
> reader depended on, that from-zero replay silently returned empty or
> partial results (and restarted any CRDT counter fold from zero). Freezing
> and materialising the projection once at open, then serving those durable
> rows, removes the dependency on the WAL prefix entirely: a later trim
> cannot perturb an already-frozen baseline, and a leaf rebuilt after
> eviction reloads the same rows for a stable point-in-time view across
> failover.

## How pages are materialised

On every `NextKeysAsync` / `NextEntriesAsync` call, the cursor grain:

1. Resolves the per-page sub-range from the cursor's persisted
   bookmark.
2. Fans out to the per-shard transient snapshot leaves addressed
   by `{physicalTreeId}/{shardIndex}/{baselineToken:N}`. Each snapshot leaf was
   seeded in memory at capture from the materialised per-shard
   projection. Pages are served straight from that in-memory baseline -
   on a multi-shard tree through a donor-orphan / virtual-slot ownership
   filter that resolves each key's owner against the coordinate's pinned
   routing map (the point-in-time counterpart of the live read path's
   moved-away guard) - with **no WAL replay**. If the leaf
   was evicted and later reloads, it reloads from its durable baseline
   row (written lazily; see below). A coordinate persisted before the
   frozen-baseline store existed (empty `SnapshotBaselineToken`) falls
   back to the legacy from-zero WAL replay for wire compatibility.
3. Performs a k-way merge of the per-shard pages back into the
   cursor's scan order, advancing the persisted bookmark.

## Lazy baseline persistence

A snapshot whose entire range drains in a **single page** is the common
case (small range, or a generous page size). For that case the durable
baseline write at capture plus the durable delete at close are pure
write-amplification: the snapshot never survives past the first page, so
the in-memory seed is sufficient.

The frozen baseline is therefore seeded **in memory** at capture and
persisted **lazily**: the cursor flushes every shard's baseline durably
(`ISnapshotLeafGrain.EnsurePersistedAsync`) the first time a page reports
`HasMore == true`, *before* it returns that page or writes its
continuation bookmark. The client only ever observes a continuation
token after every shard baseline is durable, so any cursor that survives
past page 1 - and can therefore fail over to another silo - is always
backed by a durable baseline.

| Outcome | Durable writes |
|---|---|
| Snapshot drains in one page (`HasMore == false`) | none - served from the in-memory seed, close skips the durable delete |
| Snapshot spans multiple pages | one baseline write per shard on the first `HasMore == true` page, deleted at close |

**Eviction before the first read.** Because nothing is persisted until
the first multi-page boundary, a seeded leaf that idle-evicts *before its
first page is read* loses the unread snapshot. The next read fails fast
with `LatticeSnapshotExpiredException` (an `InvalidOperationException`),
and the caller simply reopens the snapshot. This is an availability edge
only - a wrong or partial point-in-time view is never returned.

The snapshot leaves are activation-cached and idle-evict under the
silo's ordinary Orleans grain-collection policy
(`GrainCollectionOptions.CollectionAge`).
`LatticeOptions.SnapshotLeafIdleTtl` (default 30 minutes) is described
as this eviction window, but no runtime code currently reads it, so
changing it has no effect. A
subsequent page after eviction transparently rebuilds the leaf by
reloading the same durable frozen baseline, so the view stays stable
regardless of any WAL trimming that happened in the meantime.

## WAL retention

A snapshot cursor still registers itself with the WAL cursor registry
through `IWalCursorRegistry.ReportCursorAsync(...)` for the lifetime of the
cursor, but it reports the coordinate's registry clock (always zero for a
coordinate captured by the current build) with no blocked floor - a
registration the WAL garbage collector excludes from its trim floor - so it
does not hold back WAL trimming. That is safe because pages are served from
the frozen baseline rather than the WAL: the registration is a diagnostic
anchor (it is what `orleans.lattice.snapshot.pins` counts), not a
correctness dependency, and trimming the WAL cannot perturb an
already-captured baseline.
Any durable per-shard baselines (written only once a snapshot spans
multiple pages) are deleted when the cursor is closed
(`CloseCursorAsync`) or evicted by the idle-TTL reminder; a single-page
snapshot persists nothing and so deletes nothing.

### Baseline TTL leak-guard

A durable baseline is normally deleted at close, but an interrupted close
(silo crash, lost client) can orphan a baseline row keyed by a per-open
token no other cursor reuses. To bound that leak, every persisted
baseline carries a sliding time-to-live governed by
`LatticeOptions.SnapshotBaselineTtl` (default 6 hours; set to
`Timeout.InfiniteTimeSpan` to disable). `SnapshotBaselineStorageGrain`
arms a self-clearing Orleans reminder when a baseline is written and
slides it forward while the snapshot is actively served, so a long-running
scan keeps its baseline alive while an abandoned one is reclaimed once the
TTL lapses. Both the slide on write and the serving leaf's keep-alive
touch are throttled (no more than once per `SnapshotBaselineTtl / 2`,
floored at one minute) so the reminder table is not rewritten on every
page.

A reclaimed baseline is gone for good. A cursor left idle past
`SnapshotBaselineTtl` keeps serving pages only while its snapshot leaves
stay activated; once a leaf has to reload its baseline (after idle eviction
or a failover) and finds the row reclaimed, the page throws
`LatticeSnapshotExpiredException`, exactly as for a baseline that was never
persisted, and the caller reopens the snapshot.

## Bounding the cost

Open-time cost is gated by
`LatticeOptions.MaxSnapshotReplayEntries` (default 10 million entries
per shard). With the frozen-baseline store the per-shard cost is the
**materialised baseline row count** - what the snapshot leaf seeds into
memory - rather than the captured WAL head: after a GC trim the head can
be arbitrarily large while the real projection is small. The gate is
checked once every shard's capture has returned: if the deepest shard's
baseline exceeds this budget, `OpenSnapshot*CursorAsync` throws
`LatticeSnapshotReplayBudgetExceededException` and no cursor is created.
The budget does not bound the capture itself - by the time it is
compared, every shard's baseline has already been built and seeded in
memory into its snapshot leaf. Those seeds are never persisted, and
they are released when the orphaned leaves idle-evict.

## Admission control under saturation

A snapshot open is heavier than a single write: it freezes and materialises every shard's leaf chain on the non-reentrant shard roots. Admitting one into a tree that is already WAL-saturated piles that work onto roots collapsing under write back-pressure, starving replication applies and reads queued on the same roots, and a client that retries on the resulting timeout sustains a scan storm.

To prevent that, when [`LatticeOptions.ShedSnapshotOpensWhenSaturated`](configuration/options-reference-2.md#shedsnapshotopenswhensaturated) is enabled (the default) and the tree's per-silo WAL saturation signal reports `Saturated` at the moment of the open, `OpenSnapshot*CursorAsync` sheds the open at admission - before any per-shard baseline capture is fanned out - by throwing a [`LatticeSaturatedException`](https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/src/lattice/LatticeSaturatedException.cs) carrying the tree id, with `SaturationSource` `SnapshotCursorOpen` (counted on `orleans.lattice.saturation.refusals` under `source=snapshot_cursor_open`). That source is not automatically retryable: an immediate retry would re-fan the capture onto the saturated shard roots the shed protects. Only `Saturated` (the "pause new appends" regime) sheds; a `Throttled` tree is unaffected and stays browsable. The check reads the signal under the id the tree was addressed by, while the signal is sampled under the id of the write-ahead log the tree's writes land in, so on an aliased tree - after a resize, a shadow-cutover restore or a schema remediation, whose writes land in the physical copy's log - the check does not see that log's saturation and does not shed the open. The caller contract is the same as every other `LatticeSaturatedException` source: back off briefly and retry once the tree drains.

Over the read-only state API this refusal is mapped to gRPC `ResourceExhausted`; the Explorer catches that code, leaves the connection connected (other trees stay browsable) and does not auto-retry (which would amplify the storm). The Explorer browses a tree live by default and opens a snapshot cursor only when a scan is switched to its Snapshot mode; a shed open then shows as keys that did not load in time, with a **Try again** action. Set the option to `false` to restore the prior always-open behaviour.

## Observability

| Instrument | Kind | Tags | Description |
|---|---|---|---|
| `orleans.lattice.snapshot.replay.duration` | Histogram (ms) | `tree`, `shard`, `tenant` | Per-shard wall-clock WAL replay time when a snapshot leaf opens from a legacy coordinate (one persisted before the frozen-baseline store, with no baseline token). A frozen-baseline snapshot never replays the WAL, so it never records this. |
| `orleans.lattice.snapshot.replay.entries` | Counter (`{entry}`) | `tree`, `shard`, `tenant` | WAL entries consumed by that legacy snapshot-leaf replay. Not recorded for a frozen-baseline snapshot. |
| `orleans.lattice.snapshot.pins` | ObservableGauge (`{pin}`) | `tree`, `tenant` | Live WAL cursor-registry pins held by snapshot cursors (registrations that do not hold back trimming - see [WAL retention](#wal-retention)), derived from the WAL cursor registry rather than accumulated. |

`orleans.lattice.snapshot.pins` reports the pins held **now**, derived from the WAL
cursor registry, rather than accumulating a `+1` on open and a `-1` on close. That is
what makes it self-healing: an activation lost with its silo never emits a `-1`, but
there is no compensating write to lose, so the series falls to zero the moment the
registry stops holding the pin. A value that climbs and stays up is a real pin leak.

Two properties of that derivation are worth knowing before you trust a reading:

- **The value is as of the tree's last WAL GC pass, not as of the scrape.** Opening and
  closing a cursor updates the series immediately, but the re-derivation that drops a pin
  no one could report the release of runs once per GC pass. So a pin lost with its silo
  clears within that tree's *effective* GC interval, which the scheduler varies between
  [`WalGcMinInterval`](configuration/options-reference-5.md#walgcmininterval) and
  [`WalGcInterval`](configuration/options-reference-5.md#walgcinterval) (30 s to 1 h by default). A quiet tree
  relaxes toward the ceiling, so the worst case on a healthy tree is the ceiling; a tree
  blocked by an unusable pin holds at the floor, so the tree most likely to be under
  investigation is also the one that reconciles most often.
- **Setting `WalGcInterval` to zero or less disables the scheduler, and with it the
  self-healing.** Neither the per-tree zero-priming nor the re-derivation runs. Opens and
  closes still keep the series correct, and it still cannot ratchet, so it remains more
  accurate than a plain accumulating counter - but a pin lost with its silo will never age
  out of the gauge, because nothing re-reads the registry. Re-enable WAL GC if you need
  this series to self-heal.

## Examples

### Manual lifecycle (durable cursor shape)

This shape is the one to use when the cursor must outlive the local
scope - the `cursorId` is a `string` so it can be persisted to a
database, sent to a queue, or resumed after a process restart.

```csharp verify
var cursorId = await lattice.OpenSnapshotEntryCursorAsync();
try
{
    while (true)
    {
        var page = await lattice.NextEntriesAsync(cursorId, pageSize: 500);
        foreach (var kv in page.Entries)
        {
            // Process kv.Key / kv.Value. Values reflect the tree state at
            // the moment OpenSnapshotEntryCursorAsync returned, regardless
            // of concurrent writes elsewhere.
        }
        if (!page.HasMore) break;
    }
}
finally
{
    await lattice.CloseCursorAsync(cursorId);
}
```

### Scoped lifecycle (`await using`)

For the common case where the cursor's lifetime is bounded by a
single stack frame, the `LatticeExtensions.Open*CursorScopeAsync`
family returns a `LatticeScopedCursor` that implements
`IAsyncDisposable`. Disposing the scope calls `CloseCursorAsync`
exactly once, even if the body throws. The scope is implicitly
convertible to its underlying `string` cursor ID, so it can be passed
directly to `NextKeysAsync` / `NextEntriesAsync` /
`DeleteRangeStepAsync`.

```csharp verify
await using var scope = await lattice.OpenSnapshotEntryCursorScopeAsync();
while (true)
{
    var page = await lattice.NextEntriesAsync(scope, pageSize: 500);
    foreach (var kv in page.Entries)
    {
        // Same snapshot semantics as the manual shape - the only
        // difference is who calls CloseCursorAsync.
    }
    if (!page.HasMore) break;
}
// scope.DisposeAsync() runs here and closes the cursor.
```

The `*Scope` family covers the five unfiltered cursor flavours - `OpenKeyCursorScopeAsync`,
`OpenEntryCursorScopeAsync`, `OpenSnapshotKeyCursorScopeAsync`,
`OpenSnapshotEntryCursorScopeAsync`, and `OpenDeleteRangeCursorScopeAsync` -
so the choice between scoped and manual is independent of the
cursor's semantics. A predicate-filtered cursor (the
`Open*CursorWherePredicateAsync` methods) has no scope helper; wrap its id
in `new LatticeScopedCursor(lattice, cursorId)` for the same close-on-dispose
behaviour. Pick the scoped shape when the cursor lives and
dies inside one method; pick the manual shape when the cursor ID
must survive a serialization or process boundary.

## Out of scope

- **Cross-cluster snapshot rendezvous.** A snapshot is local to the
  cluster that opened it.
- **Writable snapshots.** There is no snapshot variant of
  `OpenDeleteRangeCursorAsync`; a delete-range cursor spec is rejected under
  zero-observable-writes.
- **Snapshot reuse across cursors.** Two cursors opened at logically
  identical coordinates do not share their materialised snapshot leaves:
  every open mints a fresh baseline token, and the snapshot leaves and their
  baseline rows are keyed by it.
- **Live-vs-snapshot diff.** Compute the diff in the caller by running
  a live cursor against the same range.
