Table of Contents

Disaster recovery

This page documents Orleans.Lattice.Backup 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at disaster-recovery.md, and llms.txt lists every page.

How to recover Orleans.Lattice backups after catastrophic loss of the cluster that took them - and how the backup surface keeps itself recoverable so that recovery is possible in the first place.

This guide covers the operator-facing model and runbook. For the per-member API, see the API reference and the Orleans.Lattice.Api.Backup API reference.

The problem: two stores that can drift

A backup has two distinct pieces of state:

  • Payload - the self-describing, content-addressed BackupManifest and the artifacts it references (named per capture, not by content), written to the external sink. This is the backup itself.
  • Discovery index - the per-cluster catalog in the reserved sys-backup-catalog tree, which lists the manifests a cluster knows about so they can be enumerated, described, and selected as restore points.

The catalog lives inside the very cluster whose data the backup protects. If that cluster is lost (corrupted grain storage, a wiped or rebuilt silo, a lost storage account), a catalog-only view of the world would report no backups even though the payload is sitting intact in the durable sink. Conversely, a catalog can retain a row whose sink payload has been deleted out from under it, so it offers a restore point that will fail the moment it is used.

The model: the sink is the single source of truth

Orleans.Lattice resolves this by treating the durable sink as authoritative and the catalog as a rebuildable projection over it. A BackupManifest is fully self-describing: it records the consistency cut, shard topology, per-key shape and merge-mode map, per-origin provenance, content descriptors (with SHA-256 digests), and - for an increment - its BaseBackupId. Nothing the catalog holds is unique to the catalog; every row can be re-derived from the sink.

Four capabilities follow from that model, each exposed through the Orleans.Lattice.Api.Backup control facade. Long restores use the accept-then-poll ILatticeBackupOperations start verbs; catalog maintenance and health operations remain on ILatticeBackupControl. Every operation below authorizes fail-closed except the advisory IsHealthMonitoringAvailableAsync flag:

Capability Operation What it does
Rebuild catalog RebuildCatalogFromSinkAsync Re-derives the catalog by enumerating every manifest in the sink and re-registering it.
Scrub catalog ScrubCatalogAgainstSinkAsync Flags (and optionally prunes) catalog rows whose sink payload is gone.
Cold restore StartColdRestoreAsync, then poll GetOperationStatusAsync Restores a backup into a fresh cluster from the sink alone, with no surviving catalog.
Health monitoring IsHealthMonitoringAvailableAsync, CheckBackupHealthAsync, GetBackupHealthAsync, ConfigureBackupHealthAsync Periodically re-verifies that each backup's sink payload is present and intact.

Rebuild the catalog from the sink

RebuildCatalogFromSinkAsync enumerates every manifest the sink holds (via ILatticeBackupSink.ListManifestsAsync, which returns manifests in backup-id order) and re-registers each into sys-backup-catalog under system origin. It is idempotent: a manifest already catalogued is reconciled in place, keeping its immutable capture timestamp, rather than duplicated; a catalog missing rows the sink has is repopulated. It returns a BackupCatalogRebuildReport whose invariant is ScannedCount == RegisteredCount + ReconciledCount.

Use it whenever the catalog has drifted from the sink - after a non-clean cluster restart, a partial storage loss, or any time the enumerated backups look incomplete.

Scrub the catalog against the sink

ScrubCatalogAgainstSinkAsync is the reconcile pass in the other direction: it enumerates every catalog row and probes the sink for its resolvability (manifest present, and every referenced artifact present and committed), reporting the orphans - rows whose payload is gone. It is non-destructive by default: with pruneOrphans: false it only flags orphans in the returned BackupCatalogScrubReport; with pruneOrphans: true it removes each orphan row. It is idempotent, and shares the same high-privilege, fail-closed Restore grant as the rebuild. Scrubbing keeps a dead backup from ever being listed or offered as an incremental base or restore point only to fail later.

Cold restore into a fresh cluster

StartColdRestoreAsync on the backup operations facade is the operator-facing acid test that a backup is genuinely useful after cluster loss. It restores a backup into a brand-new, independent cluster whose only shared state with the original is the sink:

  1. It bootstraps the reserved sys- trees if they are absent, so a cluster whose catalog has never existed can proceed.
  2. It resolves the target (tip) manifest directly from the sink, never the catalog.
  3. It hands the request to the existing HLC-preserving restore engine, which walks the BaseBackupId chain and verifies every referenced artifact against its recorded digest before applying anything. That engine reads catalog-first with a sink fallback, so on a cluster whose catalog is gone the whole chain resolves from the sink.
  4. It replays the chain, so the recovered tree keeps every entry's hybrid-logical-clock, version vector, origin cluster id, expiry, and tombstone flag exactly as captured.
  5. It re-projects the catalog from the sink, so the recovered cluster ends up with a correct catalog.

It reuses LatticeRestoreRequest / LatticeRestoreResult, authorizes fail-closed against the target scope (the request's target tree, or the tree the sink-held manifest was captured from; when neither resolves, the Restore capability over the reserved catalog tree instead) - and, when it recovers a backup into a tree other than the one it was captured from, additionally against the Backup capability over that captured source scope - and is idempotent. Recovering into a fresh tree id is a supported workflow, so plan the recovering principal's grants accordingly: it needs Restore on the new id and Backup on the id the backup was taken from. It throws LatticeRestoreValidationException when the backup is absent from the sink, the base chain is broken, or an artifact is missing or tampered.

Because a cold restore depends on nothing but the sink, the same call recovers a single tree or an incremental chain (pass the tip's backup id) - as long as the sink is reachable. It takes one backup id and restores one tree, so a backup set is recovered one member at a time: each member is an ordinary per-tree backup in the sink.

Recovery runbook

When a cluster is lost and you are standing up a replacement that points at the same durable sink:

  1. Stand up the replacement cluster with the backup package registered and the same durable sink configured (for example the Azure Blob sink pointed at the surviving storage account). The reserved sys- trees start empty.
  2. Cold-restore each tree you need with StartColdRestoreAsync, targeting the tree id you want and the backup id (or the tip of an incremental chain), then poll GetOperationStatusAsync until the operation is terminal. Each accepted operation bootstraps the sys- trees on first use and re-projects the catalog as it goes.
  3. Verify the catalog by listing backups through the control facade; every cold restore re-projects every manifest the sink holds, so after the first one the catalog reflects the whole sink, including backups of trees you have not restored. To re-project the catalog for discovery before, or without, any cold restore, run RebuildCatalogFromSinkAsync.
  4. Scrub with ScrubCatalogAgainstSinkAsync if you suspect the sink itself lost some payload, so the catalog only advertises resolvable restore points.

The recovered cluster is causally faithful to the source: entries replay through the HLC-preserving merge and bulk-load seams, so a restored tree converges identically to the original.

Keeping backups recoverable: health monitoring

Recovery only works if the sink payload is actually present and intact when you need it. A blob can be deleted out of band, a lifecycle policy can expire it, or an artifact can bit-rot. Health monitoring surfaces these faults before a disaster rather than at restore time.

Every catalogued backup is auto-enrolled in a periodic health check. Each check resolves the backup's manifest, confirms every referenced artifact is present and committed, and re-hashes each artifact against the digest the manifest recorded at capture time, so silent corruption is caught, not just deletion. The result is a BackupHealthReport (status Healthy / Warning / Missing / Unknown, the missing artifact ids, the hash-mismatched artifact ids, when it was last checked, and a human-readable explanation) persisted in the reserved sys-backup-health tree.

Key properties:

  • Gated on a durable sink. The monitor is inert, and the Explorer's health views hidden, when the registered sink is not durable (ILatticeBackupSink.IsDurable is false). Verifying payload that shares the fate of the cluster it protects - the ephemeral in-cluster sink - proves nothing about disaster recovery, so there is no point running it there.
  • Periodic and configurable. The monitor runs as a reminder-driven grain that mirrors the backup scheduler. It sweeps on a cluster-wide cadence (default every 6 hours) and re-verifies each enrolled backup whose interval has elapsed. The MultiSiteManufacturing sample runs it every 5 minutes.
  • Per-backup overrides. An operator can enable or disable monitoring and set a custom interval per backup with ConfigureBackupHealthAsync, trigger an on-demand check with CheckBackupHealthAsync, and read the last stored report with GetBackupHealthAsync. A backup is re-verified only by a sweep that finds its interval fully elapsed since its last check, and that check was timestamped part-way through an earlier sweep. So an interval shorter than the sweep cadence takes effect as the cadence, and an interval equal to a whole number of sweep periods - including the default, which equals the cadence - can take effect one sweep period later than configured.
  • Peer visibility for replicated trees. For a backup of a replicated tree, each sweep also refreshes the cross-cluster sink-sharing verdict, and the report carries it as PeerVisibility plus the PeerUnconfirmedClusterIds that could not see the sink. A backup that is locally intact but whose sink is provably not readable from a peer cluster is reported Warning, not Healthy, with an explanation naming the peers - because a coordinated restore would abort on it. See Un-restorable backups: a sink that is not shared.

Un-restorable backups: a sink that is not shared

A coordinated restore of a replicated tree is all-or-nothing across every cluster, and each cluster resolves the manifest chain from its own configured sink. Point each region at an isolated sink and every capture succeeds, every local health check passes, and the restore aborts - at the worst possible moment.

That failure mode is now caught well before any restore - at silo start and on every backup-health sweep. When at least one tree is replicated and the deployment has at least one peer, each cluster writes a tiny self-naming marker into its own sink at start and reads every peer's marker back out of that same sink. A marker that is missing while its peer is reachable proves the sinks are separate. The verdict is logged loudly at start, annotated onto every affected backup's health report, and - if SinkSharingEnforcement is set to FailFast - blocks the silo from starting at all. A missing marker from a peer that is itself unreachable is reported Unverified and never fails anything; it is re-probed on the next sweep.

A replicated tree backed by the default in-cluster sink is rejected outright at start regardless of that setting, since a per-cluster reserved tree is provably invisible to a peer.

Nothing runs, and nothing changes, for a non-replicated tree, a single-cluster deployment, or a host that does not wire the replication package. See Configuration for the enforcement modes and the probe timeout.

Health in the Explorer

In the Explorer's Backups area, each row of the backup catalogue shows the latest stored health report as a health pill, and the Health page (/backups/health) lists backups with their latest reports. A focused address, /backups/health?backup={id}, shows one backup's report - its status, when it was checked, the missing or uncommitted artifacts, the hash mismatches and, for a replicated tree, which peer clusters can see the sink - with Check now and the backup's periodic health schedule. When no durable sink is configured, the Health page and the health pills are not shown.

Sink durability posture

Because the sink is the single source of truth for disaster recovery, its durability is the durability of your backups. Recommendations:

  • Use a durable, external sink (the Azure Blob sink, or another ILatticeBackupSink whose IsDurable is true) for any backup you expect to survive cluster loss. The in-cluster sink is a convenience for development and dogfooding; it shares the fate of the cluster and is not a disaster-recovery target.
  • Give the sink storage account its own redundancy and geo-replication posture appropriate to your recovery objectives; the backup surface treats whatever the sink returns as truth.
  • Keep health monitoring enabled so payload loss is detected on a cadence rather than discovered during an outage, and act on any Warning or Missing report.
  • Guard the sink against out-of-band deletion (lifecycle policies, retention locks) so a backup you still depend on is not expired underneath the catalog.

See also