---
title: "Orleans.Lattice.Backup architecture"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.backup/architecture.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice.backup/architecture.md"
package: "Orleans.Lattice.Backup"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.backup/llms-full.txt"
---
# Orleans.Lattice.Backup architecture

Part of the [Backup documentation](README.md).

This page describes the end-to-end capture, incremental, restore, scheduling, and sink pipelines by behaviour, and the core Lattice seams they attach to. Public types are named; the engine, coordination grains, collectors, authorizer, and inventory registry are internal and are described by their effect.

## Where it attaches

The backup engine is a set of silo-singleton services and per-scope coordination grains layered strictly above the core data plane. It reuses the core primitives unchanged rather than reaching around them:

- **The snapshot cursor** gives a capture its point-in-time isolation. A capture opens through the public snapshot-cursor surface, so it inherits the core zero-observable-writes read, the saturation shedding, and the core replay-budget guard (a snapshot open is refused when the deepest shard's baseline would exceed `LatticeOptions.MaxSnapshotReplayEntries`). Before it opens the cursor, a capture also compares the scope's live entry count against the same budget and fails fast, before any snapshot is pinned or data is read, when the scope would exceed it.
- **The write-ahead log** gives an incremental capture its resume points. The base backup's manifest records the per-partition WAL offsets of its consistency cut; an incremental reads forward from those offsets.
- **The last-writer-wins merge and bulk-load shard seams** give a restore its causal fidelity. Entries replay through the same HLC-preserving merge and bulk-load entry points the data plane uses, so every entry's hybrid-logical-clock, version vector, origin cluster id, expiry, and tombstone flag land bit-identical to the capture.
- **The access gate** gives every operation its authorization. Backup and restore are dedicated capabilities evaluated against the same gate the data path consults.
- **The tree registry** gives the sink, catalog, and shadow-cutover restore their storage. The reserved `sys-backup-*` trees are ordinary dogfooded `ILattice` trees that self-register and carry the core `sys-` catalog-hiding prefix.
- **The view infrastructure** gives the catalog its audit trail. `AddLatticeBackup` ensures views are present so a durable per-key history view over the catalog tree records every manifest catalogued and removed.

## Full capture pipeline

`ILatticeBackupCaptureService.CaptureAsync` runs a capture in ordered steps (the phase names after the first step are the values of the `phase` metric tag; see [Observability](observability.md)):

1. **Authorize.** The scope is authorized fail-closed at its root against the `Backup` capability before any data is touched. A whole-tree scope is a whole-tree check; a prefix or key scope is a point check at the prefix or key. A failure here is tagged `snapshot-open`.
2. **Snapshot open.** The in-scope size is checked against the replay budget, the per-partition WAL head frontier is recorded, and a point-in-time cursor is opened over the scope. This is the phase that sheds under saturation or is rejected by the replay-budget guard.
3. **Export.** In-scope entries are streamed out of the pinned snapshot, page by page, with their full last-writer-wins / CRDT metadata, straight into an artifact in the configured sink under a per-capture artifact id. The stream is hashed as it flows, so a large tree is never materialized whole.
4. **Sink write.** The self-describing manifest - whose id is the payload's SHA-256 content address - is written to the sink.
5. **Manifest commit.** The manifest is registered in the catalog, keyed by the backup id. Committing the manifest is the point at which the backup becomes enumerable through the catalog; a restore by id can already resolve it once step 4 has written the manifest, because a restore that misses the catalog reads the manifest from the sink.

The manifest that results records the consistency cut (WAL sequence, HLC timestamp, per-origin frontier, and per-partition WAL offsets; on a full capture the HLC timestamp is the highest HLC over the captured entries - see [`BackupConsistencyCut`](api.md#backupconsistencycut)), the shard topology, a per-key shape and merge-mode map, per-origin provenance high-water marks, the capturing cluster's id, and an optional compression-dictionary reference. The id is the content address of the backup, so an identical retry derives the same id and registers once, keeping the first registration's capture time; the artifact id is not content-addressed, so the retry's payload is written again under a fresh artifact id that the re-registered manifest references.

## Backup set and the cross-tree fence

`CaptureSetAsync` captures one full backup per scope under a single `BackupSetManifest`. A single-tree set, or a set with `CrossTreeConsistent` left `false`, issues no extra coordination and captures each member with the cheap per-tree cut. Set membership is not stored as a set: the manifest is returned but never persisted, and the only durable trace is the `SetId` / `SetName` / `SetCreatedAtUtc` stamp pushed onto each member's own manifest. A single-tree set is deliberately left unstamped, so it is also given no `SetId` - an id minted for it would match no catalog row and resolve to no member trees.

When the flag is set over more than one tree, the set is captured inside a single shared cross-tree causal fence. The fence is selected only after every in-flight cross-tree atomic saga touching the set has drained to a terminal decision, so a cross-tree atomic write is never torn across the set boundary: for each such batch, either all members are present at or under the fence or none are. Each attempt's drain wait is bounded by `LatticeBackupOptions.CrossTreeFenceDrainTimeout` and polled at `CrossTreeFencePollInterval`, and the capture makes at most `MaxCrossTreeFenceAttempts` attempts: an attempt is discarded and retried when a new cross-tree saga registers on the set during the capture window, or one is still in flight when the attempt re-observes. A drain that times out, or a final attempt that is discarded, fails the capture with `LatticeBackupCrossTreeFenceException`. A discarded attempt's member backups are not removed: each was already written and catalogued as an ordinary, unstamped backup, so it stays listable (an unchanged tree re-derives the same content-addressed id on the next attempt and simply re-registers it). The selected fence and its drain statistics are recorded on the set manifest as a `BackupSetFence`, whose `HlcTimestamp` is the wall-clock tick count at which the fence was selected. The fence is a quiescence window rather than a shared snapshot: each member is still captured by its own point-in-time snapshot, one tree after another, so it keeps cross-tree atomic writes whole but is not a single point-in-time cut across the trees for writes that touch only one of them.

## Incremental capture pipeline

`ILatticeBackupIncrementalCaptureService.CaptureIncrementalAsync` emits a forward-WAL differential layered on a base backup. It reads the base manifest from the sink, inherits the base's scope, resumes from the base backup's per-partition WAL offsets, and folds the entries that changed since the base cut into a delta artifact. The delta is written as the same uniform entry-array artifact shape a full backup uses, so the restore chain decodes a base and its increments through one path. The manifest records the base id as `BaseBackupId`, and the increment's id folds that base id into the content address of its delta payload.

The incremental **falls back to a full capture** when it cannot produce a sound delta: when the base resume point has been trimmed off the WAL (WAL fall-off), when a range delete surfaces in the delta window (a range delete cannot be expressed as a forward-entry delta without risking a missed removal), or when the base chain was captured on a different cluster (an incremental chain is bound to its base's capturing cluster, so a request to extend it elsewhere starts a fresh chain rather than forking the lineage). A fallback is recorded as a capture retry with an incremental-fallback reason.

## Restore pipeline

`ILatticeBackupRestoreService.RestoreAsync` replays a manifest chain into a target tree (the bold names are the values of the restore `phase` metric tag: `read`, `verify`, and `merge`). It offers the request to the coordinated-restore seam (see below) only once the target manifest is resolved and the restore is authorized, so a replicated target is handed to the coordinated path only on behalf of an authorized caller:

1. **Read.** The target manifest is resolved - catalog first, then the sink - and the restore is authorized fail-closed on both of the trees it names, before anything is installed: the Restore capability over the target tree, and - only when the restore retargets the backup onto a different tree - the Backup capability over the tree that manifest was captured from. Only then is the request offered to the coordinated-restore seam. When the seam declines, the rest of the manifest chain is read, base-first - a full backup, or a base plus ordered increments up to the chosen point - and every distinct source tree in it is authorized too. See [Authorization](#authorization).
2. **Verify.** Every referenced artifact is validated against its recorded content digest. Any mismatch aborts the restore with `LatticeRestoreValidationException` before anything is installed.
3. **Merge.** The entries are applied through the HLC-preserving seams, either **in place** - a bottom-up bulk-load when the target tree has never been registered and a single full whole-tree backup is restored with no narrower sub-scope, otherwise a last-writer-wins merge that converges with whatever the target already holds - or via an **atomic shadow-cutover** that builds a fresh physical tree (bulk-loaded from a single full backup, merged from a chain with increments) and swaps the target's registry alias to it in one step.

A restore is idempotent: re-running the same request converges to the same state. A shadow-cutover restore records the physical tree it built (`ShadowPhysicalTreeId`) and the physical tree the alias resolved to beforehand (`PreviousPhysicalTreeId`).

`RevertRestoreAsync` undoes a shadow-cutover by swapping the registry alias back to `PreviousPhysicalTreeId`, restoring the pre-restore state. It is idempotent and rejects a result that did not come from a shadow-cutover restore. Authorization covers the logical target tree, while the physical tree ids on the result are caller-supplied, so they are re-resolved against the registry before the alias moves: a physical tree that is neither the target, nor a shadow the engine stamped for that target, nor the target's current physical tree is refused. That keeps a caller authorized to restore one tree from redirecting it onto another tree's shards, which the data-plane gate - bound to the logical tree id - would then serve under the redirected tree's own policy.

A shadow-cutover restore and a revert each hold the target tree's alias reservation - a restore from the moment it registers its shadow tree, a revert from just before it moves the alias - and release it when they complete. While it is held, a delete of the tree is refused with `InvalidOperationException`, as is a resize, a schema remediation, or a different restore or revert of it, each of which takes the same reservation. A restore or revert that fails part-way keeps the reservation: it is released only when the same request is retried to completion or, for a restore, when its aborted shadow is garbage-collected through `ILatticeCoordinatedRestoreEngine.DeleteShadowAsync` (which a coordinated restore runs when it aborts), so a revert, or a local restore whose shadow is never garbage-collected, that is refused again on every retry - by the ownership guard, say - leaves the tree's delete and its other alias changes refused. The shadow tree outlives a failure too: a shadow-cutover restore registers it before it takes the reservation, so a restore that fails or is refused after that point - including one refused outright because the tree is deleted or another operation holds the reservation - leaves it registered, empty when nothing was built, until the same request is retried to completion, which reuses it, or the aborted shadow is garbage-collected (as a coordinated restore does when it aborts). A restore or revert is likewise refused with `InvalidOperationException` while the tree is deleted or a delete of it is pending. The restore's alias swap, and a revert that re-points the alias rather than removing it, also consult the registered [tree ownership guard](../lattice/tree-registry.md#ownership-bounded-aliasing), whose refusal throws `LatticeTreeOwnershipDeniedException`. See [Deleting an aliased tree](../lattice/tree-deletion.md#deleting-an-aliased-tree).

Because a restore preserves each entry's origin cluster id and provenance, a restored tree replays its captured per-origin history faithfully rather than presenting as a single new origin. That preservation is necessary but, for a replicated tree, not sufficient on its own: a single cluster restoring in isolation can have its restored cut re-advanced by the cross-cluster union of the other peers. Restoring a tree that is currently replicated is therefore promoted to an all-or-nothing coordinated restore across the tree's current peers - every cluster builds its shadow, then commits the cutover together under a shared write fence, or all clusters compensate back to their pre-restore state. Restoring an unreplicated tree needs no such coordination and takes the plain local path above. A backup set with at least one replicated member restores every member tree as one group under the same coordinated cutover; a set with no replicated member restores its members one after another, each through its own local shadow-cutover.

## Scheduling and retention

`ILatticeBackupScheduler` is the public front door; the actual coordination is a single per-scope grain keyed by `BackupScopeKey.For(scope)`. Keying the grain by the scope key means on-demand triggers, scheduled captures, and retention for the same scope are serialized through one coordinator and never overlap - a trigger issued while a capture is in flight returns `null` rather than starting a second one.

`EnsureScheduleAsync` registers Orleans reminders for the scope's configured full and incremental cadences (clamped up to the one-minute reminder minimum). Registering a schedule - through `EnsureScheduleAsync` or `ScheduleRecurringBackupAsync` - authorizes the caller's `Backup` capability over the scope, because the reminder that later fires carries no caller: each reminder-fired cycle runs system-origin, so the trust decision is taken at registration. Each scheduled cycle runs the capture and, when retention is enabled, a retention pass afterwards. `PruneAsync` evaluates the chain against `RetentionKeepLast` and `RetentionMaxAge` and prunes only backups that fail every enabled rule, always preserving the base chain of a retained increment; it returns a `BackupRetentionReport`.

Per-scope schedule registration and last-run status are tracked so the control facade and the observable gauges can report a scope's health (see `BackupSchedulerRuntimeStatus`).

## The sink seam

`ILatticeBackupSink` is the storage boundary. It stores two kinds of content - streamed artifacts and self-describing, content-addressed manifests - and its artifact surface is chunk-streaming on both write and read so a large payload never buffers whole. `AddLatticeBackup` installs the default in-cluster sink, which dogfoods the reserved `sys-backup-store` tree, storing manifests and streamed artifact chunks as ordinary rows. A durable external sink (for example [Azure Blob Storage](../lattice.backup.azureblob/architecture.md)) implements the same interface and replaces the registration; because the engine talks only to the seam, it stays unaware of the sink's backend.

Content addressing (`BackupContentHash`, lowercase hex SHA-256) is what makes the pipeline idempotent at the backup level: identical payload bytes derive an identical backup (manifest) id, so a retried capture re-registers the same backup rather than a duplicate, and re-registering a manifest is harmless. Artifact ids are deliberately not content-addressed - each capture names its artifact `{treeId}-{scopeKind}-{ticks}-{guid}`, where the scope kind is `WholeTree`, `Prefix`, or `Key` - so a retried capture writes its payload again under a fresh artifact id, and the artifact the earlier attempt wrote is left orphaned in the sink. The recorded artifact digest is what a restore and a health check verify the bytes against.

## The catalog

`ILatticeBackupCatalogStore` is the in-cluster index of manifests, persisted into the reserved `sys-backup-catalog` tree keyed by backup id. Every mutation runs through the standard write path, so it is captured by the durable per-key history view enabled by default - the record of what was catalogued and removed stays auditable beyond the source WAL window. Because the catalog tree carries the `sys-` prefix it is hidden from the default state-catalog surface, so the backup control API is the sole enumeration point for backups.

## Authorization

Backup and restore are dedicated capabilities (`Backup` for capture, `Restore` for author / bulk-load) evaluated against the registered core access gate through the shared enforcement helper the data plane already uses. This means the backup path inherits the gate's behaviour exactly: the system-origin bypass for infrastructure-authored turns, the zero-cost short-circuit when only the no-op core gate is registered (a cluster with no authorization add-on pays nothing), caller-subject resolution through the membership seam, and the bootstrap-administrator break-glass. A scope is authorized at its root: a partial or filtered allow is refused fail-closed, exactly as a bulk-load or admin operation is.

A restore names **two** trees, and both are governed. The tree written into is authorized for `Restore`. The tree the backup was captured from - the manifest's own recorded scope - is authorized for `Backup`, the same capability every other manifest-consuming verb (describe, delete, export artifact, health) gates on, and the authority the caller would have needed to capture that manifest itself. Without it a caller holding `Restore` over a tree it owns could have the contents of a tree it holds nothing on replayed into a tree it controls, gated only on knowing a backup id. Two properties keep the rule tight. The captured-source check is **skipped when source and target are the same tree**, so the ordinary same-tree point-in-time rollback needs no `Backup` grant and only a genuine cross-tree retarget - clone, disaster recovery into a fresh id, environment seeding - requires one. And the scope put to the gate is the **effective restore scope retargeted back onto the source tree**, so a sub-scoped restore authorizes only the range actually replayed, never the whole captured scope. Every distinct source tree in the manifest chain is checked, not just the tip's: an increment inherits its base's scope on the capture path, but the sink is a trust boundary, so the chain is gated rather than inferred. `RevertRestoreAsync` and `CommitShadowAsync` consult no manifest - they move a registry alias between physical trees whose provenance is separately asserted against the target - so they gate the target alone.

These checks run before a restore into a replicated tree is handed to the coordinated path, and `RestoreSetAsync` authorizes the Restore capability over every member tree before it dispatches anything, so a refused caller never reaches the coordinated path's admission probe, a peer, or its coordinator. That admission probe (`ILatticeCoordinatedRestoreEngine.ProbeAdmissionAsync`) makes the same Restore check over the effective restore scope as the shadow build does. See [Coordinated restore](../lattice.replication/coordinated-restore.md) for how the coordinated path refuses or aborts.

## Reserved namespace

The `sys-backup-store`, `sys-backup-catalog`, and `sys-backup-health` trees, and the catalog's history and index views (`sys-backup-catalog-history`, `sys-backup-catalog-index`), are reserved under the `sys-backup-` prefix. An application tree that shadowed that namespace could corrupt the catalog, so `LatticeBackupReservedTrees` lets an application validate its own tree ids against the reserved prefix (mirroring the guards the membership and authorization packages enforce on their own reserved namespaces).
