Disaster recovery
This page documents Orleans.Lattice.Backup 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at disaster-recovery.md, and llms.txt lists every page.How to recover Orleans.Lattice backups after catastrophic loss of the cluster that took them - and how the backup surface keeps itself recoverable so that recovery is possible in the first place.
This guide covers the operator-facing model and runbook. For the per-member API,
see the API reference and the
Orleans.Lattice.Api.Backup API reference.
The problem: two stores that can drift
A backup has two distinct pieces of state:
- Payload - the self-describing, content-addressed
BackupManifestand the artifacts it references (named per capture, not by content), written to the external sink. This is the backup itself. - Discovery index - the per-cluster catalog in the reserved
sys-backup-catalogtree, which lists the manifests a cluster knows about so they can be enumerated, described, and selected as restore points.
The catalog lives inside the very cluster whose data the backup protects. If that cluster is lost (corrupted grain storage, a wiped or rebuilt silo, a lost storage account), a catalog-only view of the world would report no backups even though the payload is sitting intact in the durable sink. Conversely, a catalog can retain a row whose sink payload has been deleted out from under it, so it offers a restore point that will fail the moment it is used.
The model: the sink is the single source of truth
Orleans.Lattice resolves this by treating the durable sink as authoritative
and the catalog as a rebuildable projection over it. A BackupManifest is
fully self-describing: it records the consistency cut, shard topology, per-key
shape and merge-mode map, per-origin provenance, content descriptors (with
SHA-256 digests), and - for an increment - its BaseBackupId. Nothing the catalog
holds is unique to the catalog; every row can be re-derived from the sink.
Four capabilities follow from that model, each exposed through the
Orleans.Lattice.Api.Backup control facade. Long restores use the
accept-then-poll ILatticeBackupOperations start verbs; catalog maintenance and
health operations remain on ILatticeBackupControl. Every operation below
authorizes fail-closed except the advisory IsHealthMonitoringAvailableAsync
flag:
| Capability | Operation | What it does |
|---|---|---|
| Rebuild catalog | RebuildCatalogFromSinkAsync |
Re-derives the catalog by enumerating every manifest in the sink and re-registering it. |
| Scrub catalog | ScrubCatalogAgainstSinkAsync |
Flags (and optionally prunes) catalog rows whose sink payload is gone. |
| Cold restore | StartColdRestoreAsync, then poll GetOperationStatusAsync |
Restores a backup into a fresh cluster from the sink alone, with no surviving catalog. |
| Health monitoring | IsHealthMonitoringAvailableAsync, CheckBackupHealthAsync, GetBackupHealthAsync, ConfigureBackupHealthAsync |
Periodically re-verifies that each backup's sink payload is present and intact. |
Rebuild the catalog from the sink
RebuildCatalogFromSinkAsync enumerates every manifest the sink holds (via
ILatticeBackupSink.ListManifestsAsync, which returns manifests in backup-id
order) and re-registers each into sys-backup-catalog under system origin. It is
idempotent: a manifest already catalogued is reconciled in place, keeping its
immutable capture timestamp, rather than duplicated; a catalog missing rows the
sink has is repopulated. It returns a BackupCatalogRebuildReport whose invariant
is ScannedCount == RegisteredCount + ReconciledCount.
Use it whenever the catalog has drifted from the sink - after a non-clean cluster restart, a partial storage loss, or any time the enumerated backups look incomplete.
Scrub the catalog against the sink
ScrubCatalogAgainstSinkAsync is the reconcile pass in the other direction: it
enumerates every catalog row and probes the sink for its resolvability (manifest
present, and every referenced artifact present and committed), reporting the
orphans - rows whose payload is gone. It is non-destructive by default:
with pruneOrphans: false it only flags orphans in the returned
BackupCatalogScrubReport; with pruneOrphans: true it removes each orphan row.
It is idempotent, and shares the same high-privilege, fail-closed Restore grant as
the rebuild. Scrubbing keeps a dead backup from ever being listed or offered as an
incremental base or restore point only to fail later.
Cold restore into a fresh cluster
StartColdRestoreAsync on the backup operations facade is the operator-facing acid test that a backup is genuinely useful after cluster
loss. It restores a backup into a brand-new, independent cluster whose only
shared state with the original is the sink:
- It bootstraps the reserved
sys-trees if they are absent, so a cluster whose catalog has never existed can proceed. - It resolves the target (tip) manifest directly from the sink, never the catalog.
- It hands the request to the existing HLC-preserving restore engine, which walks
the
BaseBackupIdchain and verifies every referenced artifact against its recorded digest before applying anything. That engine reads catalog-first with a sink fallback, so on a cluster whose catalog is gone the whole chain resolves from the sink. - It replays the chain, so the recovered tree keeps every entry's hybrid-logical-clock, version vector, origin cluster id, expiry, and tombstone flag exactly as captured.
- It re-projects the catalog from the sink, so the recovered cluster ends up with a correct catalog.
It reuses LatticeRestoreRequest / LatticeRestoreResult, authorizes fail-closed
against the target scope (the request's target tree, or the tree the sink-held
manifest was captured from; when neither resolves, the Restore capability over
the reserved catalog tree instead) - and, when it recovers a backup into a tree other than
the one it was captured from, additionally against the Backup capability over
that captured source scope - and is idempotent. Recovering into a fresh tree id is
a supported workflow, so plan the recovering principal's grants accordingly: it
needs Restore on the new id and Backup on the id the backup was taken from.
It throws
LatticeRestoreValidationException when the backup is absent from the sink, the
base chain is broken, or an artifact is missing or tampered.
Because a cold restore depends on nothing but the sink, the same call recovers a single tree or an incremental chain (pass the tip's backup id) - as long as the sink is reachable. It takes one backup id and restores one tree, so a backup set is recovered one member at a time: each member is an ordinary per-tree backup in the sink.
Recovery runbook
When a cluster is lost and you are standing up a replacement that points at the same durable sink:
- Stand up the replacement cluster with the backup package registered and the
same durable sink configured (for example the
Azure Blob sink pointed at the
surviving storage account). The reserved
sys-trees start empty. - Cold-restore each tree you need with
StartColdRestoreAsync, targeting the tree id you want and the backup id (or the tip of an incremental chain), then pollGetOperationStatusAsyncuntil the operation is terminal. Each accepted operation bootstraps thesys-trees on first use and re-projects the catalog as it goes. - Verify the catalog by listing backups through the control facade; every
cold restore re-projects every manifest the sink holds, so after the first one
the catalog reflects the whole sink, including backups of trees you have not
restored. To re-project the catalog for discovery before, or without, any cold
restore, run
RebuildCatalogFromSinkAsync. - Scrub with
ScrubCatalogAgainstSinkAsyncif you suspect the sink itself lost some payload, so the catalog only advertises resolvable restore points.
The recovered cluster is causally faithful to the source: entries replay through the HLC-preserving merge and bulk-load seams, so a restored tree converges identically to the original.
Keeping backups recoverable: health monitoring
Recovery only works if the sink payload is actually present and intact when you need it. A blob can be deleted out of band, a lifecycle policy can expire it, or an artifact can bit-rot. Health monitoring surfaces these faults before a disaster rather than at restore time.
Every catalogued backup is auto-enrolled in a periodic health check. Each check
resolves the backup's manifest, confirms every referenced artifact is present and
committed, and re-hashes each artifact against the digest the manifest recorded at
capture time, so silent corruption is caught, not just deletion. The result is a
BackupHealthReport (status Healthy / Warning / Missing / Unknown, the
missing artifact ids, the hash-mismatched artifact ids, when it was last checked,
and a human-readable explanation) persisted in the reserved sys-backup-health
tree.
Key properties:
- Gated on a durable sink. The monitor is inert, and the Explorer's health
views hidden, when the registered sink is not durable (
ILatticeBackupSink.IsDurableisfalse). Verifying payload that shares the fate of the cluster it protects - the ephemeral in-cluster sink - proves nothing about disaster recovery, so there is no point running it there. - Periodic and configurable. The monitor runs as a reminder-driven grain that
mirrors the backup scheduler. It sweeps on a cluster-wide cadence (default every
6 hours) and re-verifies each enrolled backup whose interval has elapsed. The
MultiSiteManufacturingsample runs it every 5 minutes. - Per-backup overrides. An operator can enable or disable monitoring and set a
custom interval per backup with
ConfigureBackupHealthAsync, trigger an on-demand check withCheckBackupHealthAsync, and read the last stored report withGetBackupHealthAsync. A backup is re-verified only by a sweep that finds its interval fully elapsed since its last check, and that check was timestamped part-way through an earlier sweep. So an interval shorter than the sweep cadence takes effect as the cadence, and an interval equal to a whole number of sweep periods - including the default, which equals the cadence - can take effect one sweep period later than configured. - Peer visibility for replicated trees. For a backup of a replicated tree, each
sweep also refreshes the cross-cluster sink-sharing verdict, and the report carries
it as
PeerVisibilityplus thePeerUnconfirmedClusterIdsthat could not see the sink. A backup that is locally intact but whose sink is provably not readable from a peer cluster is reportedWarning, notHealthy, with an explanation naming the peers - because a coordinated restore would abort on it. See Un-restorable backups: a sink that is not shared.
Un-restorable backups: a sink that is not shared
A coordinated restore of a replicated tree is all-or-nothing across every cluster, and each cluster resolves the manifest chain from its own configured sink. Point each region at an isolated sink and every capture succeeds, every local health check passes, and the restore aborts - at the worst possible moment.
That failure mode is now caught well before any restore - at silo start and on
every backup-health sweep. When at least one tree is
replicated and the deployment has at least one peer, each cluster writes a tiny
self-naming marker into its own sink at start and reads every peer's marker back out
of that same sink. A marker that is missing while its peer is reachable proves the
sinks are separate. The verdict is logged loudly at start, annotated onto every
affected backup's health report, and - if SinkSharingEnforcement is set to
FailFast - blocks the silo from starting at all. A missing marker from a peer that
is itself unreachable is reported Unverified and never fails anything; it is
re-probed on the next sweep.
A replicated tree backed by the default in-cluster sink is rejected outright at start regardless of that setting, since a per-cluster reserved tree is provably invisible to a peer.
Nothing runs, and nothing changes, for a non-replicated tree, a single-cluster deployment, or a host that does not wire the replication package. See Configuration for the enforcement modes and the probe timeout.
Health in the Explorer
In the Explorer's Backups area, each row
of the backup catalogue shows the latest stored health report as a health pill, and
the Health page (/backups/health) lists backups with their latest reports. A
focused address, /backups/health?backup={id}, shows one backup's report - its
status, when it was checked, the missing or uncommitted artifacts, the hash
mismatches and, for a replicated tree, which peer clusters can see the sink -
with Check now and the backup's periodic health schedule. When no durable sink
is configured, the Health page and the health pills are not shown.
Sink durability posture
Because the sink is the single source of truth for disaster recovery, its durability is the durability of your backups. Recommendations:
- Use a durable, external sink (the Azure Blob sink, or another
ILatticeBackupSinkwhoseIsDurableistrue) for any backup you expect to survive cluster loss. The in-cluster sink is a convenience for development and dogfooding; it shares the fate of the cluster and is not a disaster-recovery target. - Give the sink storage account its own redundancy and geo-replication posture appropriate to your recovery objectives; the backup surface treats whatever the sink returns as truth.
- Keep health monitoring enabled so payload loss is detected on a cadence rather
than discovered during an outage, and act on any
WarningorMissingreport. - Guard the sink against out-of-band deletion (lifecycle policies, retention locks) so a backup you still depend on is not expired underneath the catalog.
See also
- API reference - the health service, store, and cold-restore engine seams.
Orleans.Lattice.Api.BackupAPI reference - the rebuild, scrub, cold-restore, and health control-facade operations.- Architecture - the capture, incremental, restore, and sink pipelines.
Orleans.Lattice.Backup.AzureBlob- the durable Azure Blob Storage sink.