---
title: "Chaos Tests"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/chaos-tests.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice/chaos-tests.md"
package: "Orleans.Lattice"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice/llms-full.txt"
---
# Chaos Tests

Part of the [Orleans.Lattice documentation](architecture.md).

Orleans.Lattice ships a suite of integration tests that bombard a running
cluster (single-cluster and multi-site) with concurrent reads, writes,
scans, atomic-write sagas, topology mutations, and inter-site network
partitions, then assert that the system's public correctness guarantees
hold. They act as the end-to-end contract for the properties described
in [Consistency](consistency.md) and [Replication](../lattice.replication/README.md) -
specifically that the public `ILattice` API is strongly consistent
across arbitrary concurrent shard splits, online resizes, and online
reshards; that point-mutation public-API calls (`DeleteRangeAsync`,
`SetIfVersionAsync` (the public compare-and-swap entry point), `ScanKeysAsync` / `ScanEntriesAsync` cancellation) hold their stated
invariants under concurrent contention; that `SetManyAtomicAsync`
remains atomically visible (zero-or-all keys per poll) on the
authoring site and on every receiver site; that
`SetManyAtomicAsync` keeps a multi-tree saga's keys
all-or-nothing across every participating tree under concurrent shard
splits; and that the per-merge-mode CRDT dispatch paths converge across
partitioned sites. The single-cluster suite also exercises the
recovery protocols (resumable splits, two-phase root promotion,
shadow-write atomicity, shadow-forwarding, and registry version stamping)
under random storage-write faults. The
replication suite extends those guarantees to the production shipper,
WAL trim, per-peer liveness, tombstone-reap filtering, and the gRPC
transport; the Azure Table WAL suite pins append-batch atomicity and
offset monotonicity against the local Azurite emulator. The
`lattice.schema` companion suite additionally pins that
`SetManyAtomicAsync` stays all-or-nothing while a tree's schema
enforcement policy or target schema version is changed concurrently, and
that every stored value stays self-describing and decodable across a
version advance and an eager background migration.

Every fixture is tagged `[Category("Chaos")]`, so any test filter carrying `TestCategory!=Chaos` - the Tier 2 filter of the tiered test workflow among them - skips them. Most are also marked `[NonParallelizable]` so each has its cluster to itself; `AdmissionControlChaosTests`, `BoundedCacheEvictionChaosTests`, `ClusterSplitConcurrencyChaosTests`, `ShardConsolidationChaosTests`, `LatticePredicatePushdownChaosTests` and the replication suite's `AntiEntropyRemediationGuardChaosTests` are not.

### Core chaos suite (`test/lattice/BPlusTree/`, plus one fixture in `test/lattice/Predicates/`)

| Test class | File | Purpose |
|---|---|---|
| Happy-path chaos | `ChaosIntegrationTests.cs` | Strong invariants *during* heavy concurrent load with manually-triggered splits. |
| Chaos + storage faults | `ChaosWithFaultsIntegrationTests.cs` | Parametrized theory that injects random storage faults; asserts eventual convergence after the fault window closes. |
| Chaos + online resize | `ChaosResizeIntegrationTests.cs` | Full-workload chaos while `ResizeAsync` changes fan-out in the background under `SnapshotMode.Online`. Exercises the `TreeResizeGrain` phase machine (Snapshot → Swap → Reject → Cleanup), shadow-forwarding on every source shard, and the alias swap. |
| Chaos + online reshard | `ChaosReshardIntegrationTests.cs` | Full-workload chaos while `ReshardAsync` grows the physical shard count (4 to 8) in the background. Exercises the reshard coordinator's migration loop and its dispatch-budget clamping (`MaxConcurrentMigrations`), and `ShardMap` convergence when reshard-dispatched splits race with workload writes. The same fixture also shrinks a tree (4 to 2) under the same workload, exercising fold planning, every fold's drain, freeze, swap and finalise racing live writes, and the release of each retired donor's storage while scans and counts may still hold a pre-fold shard map. |
| Atomic-write reader isolation | `AtomicVisibilityChaosTests.cs` | Strict reader isolation: a continuous reader concurrent with `SetManyAtomicAsync` always observes either the full pre-saga snapshot, the full post-saga snapshot, or all keys hidden - never a partial view. Quiescent topology (no concurrent split/resize/reshard). |
| Atomic-write reader isolation across shard split | `ShardSplitTopologyTests.cs` | Same zero-or-all visibility invariant as `AtomicVisibilityChaosTests`, but the topology mutator drives a manual `SplitAsync` on shard 0 concurrently with a chain of `SetManyAtomicAsync` sagas. Exercises shadow-forward of saga prepares onto the destination shard and the saga terminal-broadcast retry onto the new owner via `StaleShardRoutingException`. |
| Atomic-write reader isolation across online resize | `ResizeTopologyTests.cs` | Same zero-or-all visibility invariant, but the topology mutator runs an online `ResizeAsync` (`MaxLeafKeys` / `MaxInternalChildren` to 8) concurrently with the saga chain. Exercises shadow-forwarding from the source physical tree to the destination, the alias swap, and the saga terminal-broadcast retry onto the new owner via `StaleTreeRoutingException`. |
| Atomic-write reader isolation across online reshard | `ReshardTopologyTests.cs` | Same zero-or-all visibility invariant, but the topology mutator runs a 4-shard → 8-shard `ReshardAsync` concurrently with the saga chain. Exercises the retroactive prepared-mutation sweep at `BeginShadowWrite`, the registry's `TxDecisionRetention` tombstone window, and the saga terminal-fan-out shadow-forward fallback that mirrors `TxCommit` / `TxAbort` marks onto the destination shard via the post-Complete `MovedAwaySlots` lookup. |
| Digest determinism under load | `ChaosDigestIntegrationTests.cs` | `ILattice.GetLeafProjectionDigestAsync` is byte-stable across repeated calls in a write-quiescent window after concurrent writer / scanner load, and per-shard `EntryCount` sums equal `CountAsync`. |
| Range delete under load | `ChaosRangeDeleteIntegrationTests.cs` | A worker repeatedly issues `DeleteRangeAsync` over the middle band of a 600-key universe (`rd-000200` to `rd-000400`) while point writers, a refill writer re-inserting delete-band keys, scanners and a split coordinator run. Post-window, every key in the two protected bands is present and envelope-valid - the delete never strays outside its range - and scanners never observe a malformed value. The live count inside the delete band is deliberately not pinned, because deletes and refills race. |
| Compare-and-swap under contention | `CompareAndSwapChaosTests.cs` | Four writers increment an eight-key counter universe through a read-then-`SetIfVersionAsync` (CAS) loop while a split coordinator churns shards. A lost CAS returns `false` and the caller re-reads with `GetWithVersionAsync` and retries; post-window, every stored counter equals the number of successful CAS calls made against it, so no update was lost, and no caller saw an exception outside the documented transient class. |
| Scan cancellation under load | `ScanCancellationChaosTests.cs` | Scanner workers repeatedly open `ScanKeysAsync` / `ScanEntriesAsync`, cancel 5-24 ms in, and re-open while writers churn the universe. Cancellation must surface as `OperationCanceledException` (an enumeration abort is tolerated), a partial scan must never yield an unknown key or a malformed value, and a fresh full scan after the window must return exactly the pinned universe - the check that no cancelled enumerator left state behind. |
| Multi-silo restart under load | `MultiSiloRestartChaosTests.cs` | Two-silo `TestCluster`, sustained write/read load on an `ILattice` tree, secondary silo restarted every ~2.5 s via `TestCluster.RestartSiloAsync`. Post-window invariants: pinned `CountAsync` and an envelope-valid value on every key, read back with a bounded retry (up to 60 s) that absorbs only silo-reactivation faults; during the window every exception is tolerated and counted, and only a read that returns a malformed value fails the run. Uses `ProcessScopeMemoryGrainStorage` (a static-dictionary-backed `IGrainStorage` shared across every silo in the test process) so secondary-silo restart does not wipe the shard-root / registry topology - per-silo Orleans memory storage would otherwise let a re-placed shard-root activation read empty state and overwrite the live topology with a fresh leaf root (the underlying split-brain that previously surfaced as `InvalidCastException`). It also registers one process-shared `InMemoryWalStorageProvider` on every silo through `AddWalStorage`, because a reactivated leaf rebuilds its entries from the write-ahead log and the default per-silo provider would vanish with the restarted silo, silently emptying every WAL partition it hosted. |
| Cross-tree atomic write under shard churn | `ChaosCrossTreeAtomicWriteIntegrationTests.cs` | A commit worker drives one all-or-nothing `SetManyAtomicAsync` saga per generation into the same logical slot of three distinct trees, while per-tree split coordinators churn shards and reader workers probe every settled committed generation. Asserts cross-tree all-or-nothing: a settled committed key is present in **all three trees or none**, and every committed generation is durably present in all trees post-window, even as splits move keys between shards mid-saga on every tree. |
| Admission cap under cross-shard pressure | `AdmissionControlChaosTests.cs` | Concurrent cross-shard writers drive one tree with an enforcing `LatticeOptions.MaxLiveKeys` cap and a second tree with only an advisory `LatticeOptions.AdmissionAdvisoryLiveKeys` ceiling. The enforcing cap must reject some writes with `LatticeQuotaExceededException` and keep rejecting once its aggregate settles above the cap - overshoot past the configured value is allowed - while the advisory tree never rejects. |
| Bounded read-through cache eviction | `BoundedCacheEvictionChaosTests.cs` | Pins `LatticeOptions.MaxCacheValueBytes` so small that almost every cached payload is evicted down to its metadata, so nearly every read takes the eviction path back to the primary leaf. Under concurrent overwrite churn, eviction must never turn a live key into a false miss or surface a stale or cross-key payload. |
| Guarded atomic write under split churn | `ChaosPredicateAtomicSetManyIntegrationTests.cs` | Guarded `SetManyAtomicAsync<T>` batches (`Score >= Guard`, evaluated server-side against each key's pre-saga value) under split churn and concurrent point writes: a batch over an always-matching band always commits and stamps every key, and a batch holding a permanently failing key always returns `PreconditionFailed` and writes nothing. |
| Conditional batch write under split churn | `ChaosPredicateConditionalSetManyIntegrationTests.cs` | The guarded `SetManyAsync<T>` overload stamps a marker onto a band of keys while splits move them and point writers rewrite them: a key whose current value fails the guard is never written, a matching key is, and only submitted keys are ever considered. After the window it also audits the quiesced tree for leaves that no descent from their shard root reaches, and reads every key back through the routing path, so a split that lost its separator fails the run. |
| Predicate-filtered cursors under split churn | `ChaosPredicateCursorIntegrationTests.cs` | Predicate-filtered key, entry and snapshot-entry cursors paged across a mid-paging split: pages stay in ascending key order, every surfaced item satisfies the predicate, and a stable band returns exactly its matching keys. |
| Conditional range-delete cursor under split churn | `ChaosPredicateDeleteRangeCursorIntegrationTests.cs` | A resumable conditional range-delete cursor stepped to completion in bounded pages under split churn and conflicting writes tombstones every in-range key that matches the predicate and nothing else. |
| Predicate `GetManyAsync` under split churn | `ChaosPredicateGetManyIntegrationTests.cs` | The predicate `GetManyAsync<T>` overload, evaluated server-side on the owning leaf: every returned value satisfies the predicate, and a stable band returns exactly its matching keys however splits move them. |
| Conditional range delete under split churn | `ChaosPredicateRangeDeleteIntegrationTests.cs` | The conditional `DeleteRangeAsync<T>` overload tombstones exactly the in-range keys whose value matches the predicate, and never touches a key outside the range, while splits and point writers race it. |
| Predicate scans under split churn | `ChaosPredicateScanIntegrationTests.cs` | The predicate `ScanEntriesAsync<T>` / `ScanKeysAsync<T>` / `ScanValuesAsync<T>` overloads across a mid-scan split: output stays in ascending key order, every surfaced value satisfies the predicate, and a stable band is returned in full. |
| Cluster-wide split admission after a crash | `ClusterSplitConcurrencyChaosTests.cs` | Models a silo crashing mid-split by reporting split footprints that saturate the `LatticeOptions.MaxClusterConcurrentAutoSplits` ceiling and are then never refreshed; once their time-to-live lapses a fresh grant must succeed, so a crash cannot wedge splitting cluster-wide. |
| Retry policy masks storage faults | `RetryPolicyChaosTests.cs` | Parametrized theory (5%, 15% and 30% fault probability) that arms one-shot storage write faults and requires every caller-side `SetAsync` to succeed through `BoundedExponentialRetryPolicy` under an ambient `LatticeIdempotencyContext`; a companion test arms a fault on every 15 ms tick while incrementing a `PnCounter` and requires the counter to equal the number of increments that succeeded, so a retry never double-counts. |
| Shard consolidation under churn | `ShardConsolidationChaosTests.cs` | Online shard-consolidation folds run while a split driver shatters the tree, writers ingest, and shard roots are force-deactivated: no key ever becomes unreachable, no acknowledged write is lost, and no virtual slot is left unrouted. |
| Predicate translator and evaluator agree | `Predicates/LatticePredicatePushdownChaosTests.cs` | Eight workers evaluate a storm of structurally random predicates against a shared pool of 256 encoded documents, both through the server-side predicate evaluator and as the compiled lambda; any disagreement fails the run. |

### Cross-cluster, gRPC, and Azure Table WAL suites

The replication packages and the Azure Table WAL provider ship their own chaos suites - cross-cluster convergence
and atomic visibility (`test/lattice.replication/Chaos/`), the gRPC transport
suite (`test/lattice.replication.grpc/Chaos/`), and the Azure Table WAL suite
(`test/lattice.storage.azuretable/Chaos/`). They drive the real replication
pipeline using in-process test clusters, a simulated delivery pump, an in-memory
test server, and the Azurite emulator, and are documented in
[the replication chaos tests](../lattice.replication/chaos-tests.md).

### Schema enforcement and versioning atomicity suite (`test/lattice.schema/Chaos/`)

The `lattice.schema` companion package ships its own chaos fixtures that prove the
core atomic-write guarantee is preserved *while the schema control plane mutates
underneath it* - a policy is set or cleared, or a target schema version is advanced
and eagerly migrated - concurrently with a chain of atomic sagas. They run on a
single-silo `TestCluster` with core lattice, schema enforcement, and schema
versioning all registered (`SchemaAtomicChaosClusterFixture`), and are tagged
`[Category("Chaos")]` `[NonParallelizable]` like most other suites.

| Test class | File | Purpose |
|---|---|---|
| Atomic write under policy churn | `AtomicWriteUnderPolicyChurnChaosTests.cs` | A committer drives cross-tree `SetManyAtomicAsync` sagas (alternating a compliant and a non-compliant leg) while a churner flips each participating tree's enforcement policy between "require JSON" and "no policy". Asserts every saga is decided as a unit against the policy current at admission - it either commits every leg or rejects the whole batch and mutates no tree - and a concurrent reader never observes a torn (cross-generation) snapshot. |
| Atomic write under version advance | `AtomicWriteUnderVersionAdvanceChaosTests.cs` | A committer drives single-tree `SetManyAtomicAsync` batches into one versioned tree while a churner advances the target schema version v1 -> v2 -> v3 concurrently, then a quiesced eager background migration re-stamps the tree. Asserts every batch lands all-or-nothing, that every read is always envelope-stripped and upcast to the current target, and that the eager migration preserves the last committed snapshot. |

## The workload

The four full-workload single-cluster fixtures (Tests 1-4 below) run a
parallel workload over a fixed key *universe*. The resize and reshard
fixtures pre-register their tree at 4 shards with `MaxLeafKeys = 4`; the
happy-path and faults fixtures seed a fresh tree id without pre-registering
it, so their trees take the library defaults (64 shards, 128 keys per leaf,
128 children per internal node) and their topology churn comes from the
split driver. Writers only rewrite existing keys with
monotonically-increasing values of the form `v-{keyIndex}-{writerId}-{seq}`.
Any value matching that envelope proves the byte array is internally
consistent.

The atomic-visibility fixtures (Tests 5-6, 8), the per-mode
convergence fixtures (Test 9), the multi-site smoke (Test 10), the
range-delete / CAS / scan-cancel public-API fixtures, the
production-shipper fixtures (WAL trim, liveness + inbound stats,
compaction + shipping), and the downstream-package fixtures (gRPC
transport, Azure Table WAL) do not follow this exact shape - each
defines its own universe and worker mix appropriate to the invariant
it targets. See the per-test sections, the suite tables at the top of
this document, or [the replication chaos tests](../lattice.replication/chaos-tests.md)
for the replication and storage fixtures.

Fixture and parameter differences:

| Test | Fixture | `MaxLeafKeys` | `MaxInternalChildren` | Key prefix | Universe |
|---|---|---|---|---|---|
| Happy-path | `FourShardClusterFixture` (tree not pre-registered: 64 shards) | 128 (library default) | 128 (library default) | `chaos-{i:D5}` | 500 |
| Chaos + faults | `MultiShardFaultInjectionClusterFixture` (tree not pre-registered: 64 shards) | 128 (library default) | 128 (library default) | `fchaos-{i:D5}` | 200 |
| Chaos + resize | `FourShardClusterFixture` (4 shards) | 4 -> `16` mid-run | 128 (library default) -> `16` mid-run | `resize-chaos-{i:D5}` | 200 |
| Chaos + reshard | `FourShardClusterFixture` (4 shards -> 8, or -> 2 in the shrink run) | 4 | 128 (library default) | `reshard-chaos-{i:D5}` | 200 |

```mermaid
flowchart LR
    Seed[Seed universe<br/>N keys] --> Chaos

    subgraph Chaos[Chaos window]
      direction TB
      PW[Point writers] --> Tree
      BW[Bulk writers] --> Tree
      AW[Atomic writers] --> Tree
      PR[Point readers] --> Tree
      BR[Bulk readers] --> Tree
      SC[Scanners] --> Tree
      CT[Counters] --> Tree
      TM[Topology mutator<br/>split / resize / reshard<br/>± fault injector] --> Tree
      Tree[(ILattice)]
    end

    Chaos --> Assert[Assert invariants]
```

Worker categories (the exact mix and counts vary per test; each fixture declares them as constants at the top of its source):

* **Point writers** - `SetAsync` on random universe keys.
* **Bulk writers** - `SetManyAsync` with batches of 8 random keys
  (happy-path / faults only).
* **Atomic writers** - `SetManyAtomicAsync` with batches of 2 random keys
  (happy-path / faults only).
* **Point readers** - `GetAsync`; validates envelope if a value is returned.
* **Bulk readers** - `GetManyAsync` for 16 random keys (happy-path /
  faults only).
* **Scanners** - happy-path: rotating `ScanKeysAsync`, `ScanEntriesAsync`,
  reverse scan and range scan, each of which must yield exactly the
  universe (or the range) in order, with no duplicates and no unknown
  keys; resize and reshard: full-tree `ScanKeysAsync` with no duplicates
  and no unknown keys; faults: `ScanEntriesAsync`, checking order,
  unknown keys and envelopes.
* **Counters** - `CountAsync` must always equal the pinned universe size
  (the faults fixture lets the count drift during its fault window).
* **Topology mutator** - test-specific:
  * happy-path: every ~200 ms drives a manual split, and then a split
    pass, on a randomly chosen non-empty shard.
  * faults: the same split driver at a ~500 ms cadence, plus a fault
    injector that arms random `WriteStateAsync` faults.
  * resize: initiates `ResizeAsync` once at the window start and pumps
    the coordinator to completion.
  * reshard: initiates `ReshardAsync(8)` once at the window start and
    pumps the reshard coordinator and the per-shard split passes to
    completion; the shrink run initiates `ReshardAsync(2)` instead and
    pumps the coordinator and every fold it starts.

## Test 1 - Happy-path chaos (`ChaosIntegrationTests`)

This test establishes that `ILattice`'s consistency guarantees hold
*during* the chaos window, not just after it closes. Every operation
observes a fully consistent view of the tree.

### What it proves

| Invariant | Mechanism under test |
|---|---|
| `CountAsync` returns the exact universe size, always | Per-slot count routing against the authoritative `ShardMap`, each shard counted in work-bounded batches, plus version stability check |
| `ScanKeysAsync` / `ScanEntriesAsync` yield exactly the universe, no duplicates, no unknowns, in strict sorted order | In-line reconciliation-cursor injection into the k-way merge + `HashSet` dedup |
| `ScanKeysAsync(null, null, reverse: true)` yields the full universe in reverse | Reverse-scan path also reconciles |
| `ScanKeysAsync(start, end)` yields exactly the in-range slice | Range pruning is slot-aware |
| `GetAsync` / `GetManyAsync` never return a corrupt value | Writes are atomic per-shard; CRDT LWW resolves concurrent rewrites |
| No public-API call throws an unhandled exception | Stale routing retries and enumeration aborts are transparent |
| Splits during a scan never cause data loss, duplication, or out-of-order output | `MovedAwaySlots` + version stamping + in-line reconciliation |

### Tolerated transients

The test counts these as transient rather than as failures:

* `EnumerationAbortedException` - a stream cursor grain deactivated
  mid-iteration. The caller re-issues the scan.
* The stale shard-routing signal - a routing activation used a cached
  shard map after a concurrent split committed its swap. The routing
  tier invalidates its map and retries within a bounded wall-clock
  budget, so the signal reaches the caller only once that budget is spent.
* An `InvalidOperationException` reporting that an atomic write failed and
  was rolled back (a saga aborted by a transient routing fault mid-split),
  or that a count, scan or `GetManyAsync` exhausted
  `LatticeOptions.MaxScanRetries` while the topology or the saga rate kept
  changing.
* `TimeoutException` - an Orleans call timeout under saturated load.

Any other exception, or any observed envelope/duplicate/missing-key
violation, fails the test.

### Pass criteria

After the chaos window closes:

* `CountAsync` matches the pinned universe size exactly.
* `ScanKeysAsync` yields exactly the pinned universe (no gaps, no extras).
* Every worker category performed at least one operation (proves the
  workload ran under real concurrency, not a degenerate single-thread
  schedule); splits and atomic writes count attempts.
* Once the atomic writer's committed and transiently failed sagas total
  at least three, at least half of them committed.
* Zero envelope violations were observed *during* the window.

## Test 2 - Chaos + storage faults theory (`ChaosWithFaultsIntegrationTests`)

This parametrized theory layers random storage faults on top of the same
workload. Unlike the happy-path test, per-operation invariants are
*weakened* during the fault window - arbitrary exceptions are tolerated
because a failed `WriteStateAsync` legitimately cascades into split
aborts, stale routing, and count drift. Instead, the test asserts
**eventual convergence**: once faults stop and the cluster quiesces,
the tree must recover to the exact same pinned universe with every
value still matching its envelope.

`faultProbability` is the probability, per 20 ms tick, that the fault
injector arms a fresh one-shot `WriteStateAsync` fault on a randomly
chosen target grain (the initial leaf and the shard-root grain of shards
0-3, a subset of the tree's 64 shards).
Orleans' `FaultInjectionGrainStorage` consumes each armed fault on the
next write for that grain, so the injector re-arms continuously to
approximate a steady-state failure rate.

> Note: Orleans' one-shot fault API caps concurrent armed faults at
> ≈ `|targets|`. Higher `faultProbability` primarily drives faster
> re-arm latency rather than a linear increase in fault count. The
> gradient is still meaningful for exercising recovery paths under
> progressively heavier disruption.

### Test phases

```mermaid
sequenceDiagram
    participant Test
    participant Tree as ILattice (64 shards)
    participant Injector
    participant Workers
    Test->>Tree: Seed universe (faults off)
    Test->>Injector: Start at p=faultProbability
    Test->>Workers: Start 13 role workers + split coordinator
    loop Chaos window (8 s)
        Injector-->>Tree: AddFaultOnWrite(random target)
        Workers-->>Tree: mixed reads/writes/scans/splits
        Note over Workers: exceptions tolerated<br/>envelope-check values if observed
    end
    Test->>Injector: Stop (cts fires)
    Test->>Tree: DrainAndHealAsync (up to 15 s)
    Note over Tree: retry writes over universe<br/>until 3 consecutive clean passes
    Test->>Tree: Assert strong invariants
```

### Tolerated during faults

Every exception type is tolerated and counted (`tolerated-write-errors`,
`tolerated-read-errors`, `tolerated-scan-errors`, etc.). A single
storage fault cascades into many observable shapes:

* Direct `InvalidOperationException` from the faulted write.
* `OrleansException` wrappers when a faulted grain deactivates.
* `EnumerationAbortedException` if a stream cursor was on the
  deactivated grain.
* `StaleShardRoutingException` after a shard map swap when the split
  coordinator crashed and resumed mid-phase.
* `ArgumentException` from the injector itself when a target already
  has an armed fault pending (skipped).

Envelope violations (a value that doesn't start with `v-{index}-`) are
**not** tolerated - CRDT LWW is supposed to preserve atomicity of the
value payload even when the wrapping write fails.

### Healing phase (`DrainAndHealAsync`)

After the fault injector stops, lingering armed faults remain on
whichever targets weren't hit during the chaos window. The test drains
them by replaying writes over the entire universe until **3 consecutive
passes complete exception-free**, bounded by a 15 s timeout. This loop:

* Consumes any remaining one-shot faults (each fires once on its next
  write, clearing itself).
* Gives resumable splits and pending root promotions time to reach
  their `RunSplitPassAsync` keepalive tick and replay.
* Rewrites every universe key with a fresh value on each pass
  (`v-{i}-heal-{pass}`), so after three clean passes every key holds a
  value a fault-free pass wrote, and that value matches its envelope.

### Pass criteria (post-quiescence)

After healing:

* `CountAsync == UniverseSize` exactly.
* `ScanKeysAsync` yields exactly the pinned universe.
* `ScanEntriesAsync` yields exactly the pinned universe with every value
  matching its envelope.
* Every universe key appears in the post-heal `ScanKeysAsync`.
* Zero envelope violations were observed during the whole run.
* Point writes, point reads, scans, counts and splits each completed at
  least once and at least one atomic write was attempted; the injector
  armed at least one fault (for `p > 0`); in the no-fault baseline at
  least 70% of atomic writes committed (with fewer than ten attempts: at
  least one committed and at most one failed).

## Test 3 - Chaos + online resize (`ChaosResizeIntegrationTests`)

This test targets the online resize path. A full concurrent workload
runs against a seeded tree while `ResizeAsync` changes the B+ fan-out
to `MaxLeafKeys = 16` / `MaxInternalChildren = 16` under
`SnapshotMode.Online`. The resize starts inside the chaos window and its
phases - snapshot drain, alias swap, per-shard reject phase, cleanup - are
pumped there under live traffic; any phase still running when the window
closes is finished in a post-window drain of up to 30 s before the
invariants are checked.

### Recovery surfaces exercised

* `TreeResizeGrain` phase machine (Snapshot → Swap → Reject → Cleanup)
  under sustained traffic.
* Shadow-forwarding on every source shard - the workload's live point
  writes during the drain must be mirrored to the destination with their
  original HLCs (the forward carries last-writer-wins writes only; typed
  CRDT deltas and bulk appends are not mirrored).
* Alias swap - mid-flight `GetAsync` / `SetAsync` on a stateless-worker
  `LatticeGrain` activation holding a stale alias must transparently
  re-resolve and retry.
* Strongly-consistent `CountAsync` / `ScanKeysAsync` during the online
  snapshot drain and Rejecting phase.

### Tolerated transients

The same set as the happy-path test, plus `StaleTreeRoutingException`
raised during the alias swap window.

### Pass criteria

After the chaos window closes:

* `CountAsync` matches the pinned universe size exactly.
* `ScanKeysAsync` yields exactly the pinned universe.
* The resize coordinator reports itself idle, so the resize ran to completion.
* Every worker category performed at least one operation; the resize
  was driven to completion.
* Zero envelope violations observed during the window.

## Test 4 - Chaos + online reshard (`ChaosReshardIntegrationTests`)

This test targets the online reshard path - growing the physical shard
count from 4 to 8 while the tree continues to serve traffic. The
reshard is kicked off synchronously before the chaos timer starts (so
cold-activation cost doesn't burn the window on slow Release CI
runners); the in-window driver only pumps the migration loop to
completion. A second test in the fixture runs the same workload while
`ReshardAsync(2)` shrinks a fresh 4-shard tree to 2 shards, its driver
pumping the coordinator and every fold it starts. Both runs finish any
residual work in a post-window drain of up to 30 s.

### Recovery surfaces exercised

* `TreeReshardGrain` migration loop under sustained traffic -
  eligibility filtering, dispatch-budget clamping
  (`MaxConcurrentMigrations`), re-evaluation across ticks.
* `ShardMap` convergence when reshard-dispatched splits race with
  workload writes (shadow-write, drain, swap, reject, permanent
  `MovedAwaySlots`).
* No invariant drift across the full reshard window.
* Shrink run: fold planning and concurrency under sustained traffic,
  every fold's drain, freeze, swap and finalise racing live writes, and
  the release of each retired donor's storage while scans and counts may
  still hold a pre-fold shard map.

Neither run exercises the autonomic hot-shard monitor's rule that
suppresses its passes while a reshard is in flight: the fixture keeps
the default `HotShardSampleInterval` and `AutoSplitMinTreeAge`, so every
monitor pass returns inside the minimum-tree-age grace period, before it
reaches that rule, until well after a reshard normally completes, and no
assertion reads the monitor.

### Tolerated transients

`EnumerationAbortedException`, the stale shard-routing signal,
`TimeoutException`, and the same saga-rollback and
`MaxScanRetries`-exhaustion `InvalidOperationException` messages as the
happy-path test.

### Pass criteria

After the chaos window closes:

* `CountAsync` matches the pinned universe size exactly.
* `ScanKeysAsync` yields exactly the pinned universe.
* The reshard coordinator reports itself idle, so the reshard ran to completion.
* The post-reshard `ShardMap` has at least `ReshardTarget` distinct
  physical shards.
* Every worker category performed at least one operation.
* Zero envelope violations observed during the window.

The shrink run checks the same count, scan, idle-coordinator, worker and
envelope criteria, and additionally that the final `ShardMap` holds exactly
two distinct physical shards and that each of the two shards that left the
map reports itself retired and holds no leaves.

## Test 5 - Atomic-write reader isolation (`AtomicVisibilityChaosTests`)

This test asserts the **universal reader-isolation invariant** for
`SetManyAtomicAsync`: every poll of a continuous reader concurrent
with an in-flight saga must observe either the full pre-saga snapshot,
the full post-saga snapshot, or all keys hidden - never a partial view.
The invariant holds **per poll**, with no bounded-window caveat, across
50 sequential saga rounds at a 10 ms reader cadence.

### What it proves

| Invariant | Mechanism under test |
|---|---|
| Continuous reader observes zero-or-all keys at every poll | The leaf's prepared-write commit path stages every saga write in a per-transaction pending bucket, and every key a read reaches resolves against one transaction-registry decision, the tree-wide visibility flip |
| Saga drives 16 keys spanning multiple leaves through the full prepare → terminal pipeline | `AtomicWriteGrain` per-shard terminal broadcast, idempotent under concurrent retry |
| Final post-round value is preserved across 50 iterations | LWW resolution under saga commit ordering |

### Workload

* **Seed phase** - `SetAsync` for each of 16 keys (`atomic-00` … `atomic-15`) at round 0, followed by a single `SetManyAtomicAsync` at round 0 to land all keys through the saga path before reader rounds begin.
* **Saga rounds** - 50 sequential rounds. Each round starts a continuous reader task that polls all 16 keys via `GetManyAsync` every 10 ms, and concurrently issues `SetManyAtomicAsync` with the post-round value envelope.
* **Reader classification** - every poll is bucketed: `fullPre` (every key at the previous round's value), `fullPost` (every key at the new round's value), `fullHidden` (every key missing during the prepare → terminal window), or **split** (any mixed observation, which fails the test).

### Pass criteria

* Zero split-view failures across all 50 rounds.
* `totalPolls > 0` and `fullPostPolls > 0` (proves the reader and saga ran under real concurrency).
* Final `GetManyAsync` of all 16 keys yields the round-50 envelope on every key.

### Tolerated transients

The reader's `GetManyAsync` may observe `OperationCanceledException` at the round boundary; saga writes are not expected to surface any transient - the saga's own retry-on-stale-routing logic absorbs split / resize / reshard activity at the API layer.

### Companion observability

See [Metrics](metrics/instrument-catalog-7.md#saga--coordinator--lifecycle).

## Test 6

Three sibling fixtures extend Test 5's reader-isolation invariant
across each of the three online topology mutations (shard split,
online resize, online reshard). Every fixture seeds the same 16-key
universe, then drives 15 sequential `SetManyAtomicAsync` rounds while
the topology mutator runs in parallel; a continuous reader polls all
16 keys every 10 ms and every poll must observe either the full
pre-round value, the full post-round value, or all 16 keys hidden -
never a partial subset.

### What they prove

| Invariant | Mechanism under test |
|---|---|
| Saga prepares survive a mid-flight shard split | Source shard's shadow-forward pipeline mirrors prepared entries to the destination during the split's drain phase; saga terminal-broadcast retries onto the new owner via `StaleShardRoutingException` |
| Saga prepares survive a mid-flight online resize | Source physical tree shadow-forwards every last-writer-wins write (including saga prepares; typed CRDT deltas and bulk appends are not mirrored) to the destination physical tree during the snapshot drain; the alias swap is observed via the stale tree-routing signal and the saga terminal-broadcast retries onto the new owner |
| Saga prepares survive a mid-flight online reshard | `TreeReshardGrain` migration loop dispatches per-shard splits; the retroactive prepared-mutation sweep at `BeginShadowWrite` and the terminal-fan-out shadow-forward fallback together mirror prepares and the saga's `TxCommit` / `TxAbort` marks onto the destination shard via the post-Complete `MovedAwaySlots` lookup; the registry's `TxDecisionRetention` tombstone window absorbs duplicate terminals |

### Workload (per fixture)

* **Seed phase** - `SetAsync` for each of 16 keys at round 0, followed by a single `SetManyAtomicAsync` at round 0 so the universe is pinned through the saga path before the topology mutator starts.
* **Topology kick-off** - exactly once before the saga loop: `SplitAsync(shard 0)` for `ShardSplitTopologyTests`, `ResizeAsync(MaxLeafKeys=8, MaxInternalChildren=8)` for `ResizeTopologyTests`, or `ReshardAsync(8)` for `ReshardTopologyTests`. A background driver pumps the coordinator's `RunSplitPassAsync` / `RunResizePassAsync` / `RunReshardPassAsync` to completion while the saga loop runs.
* **Saga loop** - 15 sequential rounds. Each round starts a continuous reader task that polls all 16 keys via `GetManyAsync` every 10 ms and concurrently issues `SetManyAtomicAsync` with the post-round value envelope.
* **Reader classification** - identical to Test 5: `fullPre`, `fullPost`, `fullHidden`, or **split** (mixed observation, fails the test).
* **Drain phase** - after the saga loop, the test pumps the coordinator to idle and asserts the final post-round value is present on every key.

### Pass criteria (per fixture)

* Zero split-view failures across all 15 rounds.
* `totalPolls > 0` and `fullPostPolls > 0`.
* Final `GetManyAsync` of all 16 keys yields the round-15 envelope on every key.
* The split, resize or reshard coordinator reports itself idle before the test exits.

### Tolerated transients

* The stale shard-routing and stale tree-routing signals on the reader's `GetManyAsync` - every fixture's reader loop catches both and retries.
* `OperationCanceledException` at the round boundary when the reader's CTS fires.

### Companion observability

Same `orleans.lattice.atomic_write.*` histograms as Test 5, plus the topology-mutator-side counters that fire when the split / resize / reshard actually mutates the tree (`orleans.lattice.leaf.splits`, `orleans.lattice.shard.splits_committed`, and the per-split `orleans.lattice.split.retroactive_forward.duration` / `.entries` pair on shadow-forward). See [Metrics](metrics.md).

## Test 7 - Digest determinism under load (`ChaosDigestIntegrationTests`)

This test exercises `ILattice.GetLeafProjectionDigestAsync` under
sustained concurrent load and asserts two determinism invariants that
gate the digest's value as a cross-silo divergence detector:
**byte-identical repeated calls** in a write-quiescent window, and
**per-shard `EntryCount` sums equal `CountAsync`**.

### What it proves

| Invariant | Mechanism under test |
|---|---|
| Digest hash is byte-stable across repeated calls when no writes occur in between | The digest is an order-independent XOR fold of per-entry XxHash128 contributions, maintained at each mutation and aggregated up the internal nodes, so once the last coalesced publish has landed, repeated reads in a write-quiescent window hash the same persisted aggregate |
| Sum of per-shard `EntryCount` equals `CountAsync` | Shard-level digest counts are accountable against the tree's own population view |
| Digest computation is safe to call concurrently with foreground writer / scanner traffic | No exception is observed on any worker - digest rendering does not block or interfere with the read / write path |

### Workload

* **Seed phase** - 200 keys (`chaos-digest-{i:D5}`) preloaded with `SetAsync`.
* **Chaos window (~8 s)** - 4 writer tasks rewriting random keys, 2 scanner tasks calling `KeysAsync`, 1 digest poller calling `GetLeafProjectionDigestAsync` for every shard in a tight loop.
* **Quiesce phase** - after the chaos window closes, the digest is sampled twice in succession with no intervening writes.

### Pass criteria

* No exception observed on any worker during the chaos window (`EnumerationAbortedException` is tolerated on the scanner; everything else is fatal).
* For every shard, `secondPass[s].Hash` equals `firstPass[s].Hash` and `EntryCount` is equal.
* Sum of `firstPass[s].EntryCount` equals `tree.CountAsync()`.
* `tree.CountAsync()` equals 200 (writers only update existing keys; no inserts or deletes).

### Tolerated transients

* `EnumerationAbortedException` from the scanner's `KeysAsync` enumeration when a stream cursor grain deactivates mid-iteration.

## Test 8 - Cross-cluster atomic-visibility chaos (`CrossClusterAtomicVisibilityChaosTests`)

This test asserts the **cross-cluster receiver-side reader-isolation
invariant** for `SetManyAtomicAsync`: a saga authored on one site and
shipped via the WAL replication transport to two peer sites must be
observed all-or-nothing on every receiver, even when the inter-site
delivery topology is partitioned and healed mid-workload. It is the
cross-cluster sibling of [Test 5](#test-5---atomic-write-reader-isolation-atomicvisibilitychaostests)
and exercises the same reader-isolation mechanism - prepared writes
staged in per-transaction pending buckets and made visible by one
transaction-registry decision - through the receiver-side
prepared/terminal apply seam.

### What it proves

| Invariant | Mechanism under test |
|---|---|
| Every saga's keys land all-or-nothing on every receiver site | Receiver-side prepared/terminal apply seam staging prepared writes in the leaf's per-tx pending bucket, then flipping them on terminal arrival |
| Source HLC rides through the wire verbatim | `LatticeHlcOverrideContext` wrapping the apply call so the receiver does not stamp a fresh local HLC |
| Repeated terminal delivery is idempotent | The per-tree transaction registry classifies an arriving terminal against the decision already recorded and returns without a write when the outcome matches, so a redelivery inside the retention window changes nothing; the leaf additionally absorbs a second delivery within one activation from its recently-terminal set |
| Mid-workload partition does not produce partial-saga visibility on any site | Prepares queued behind the partition and the matching terminal both ship after heal; the receiver-side staging buffer holds prepared entries off the visible projection until the terminal arrives |
| Producer-side per-key WAL filter does not strand terminals | `ReplicationShipperGrain.ShouldShip` bypasses `KeyFilter` / `KeyPrefixes` for `TxCommit` / `TxAbort` records |

### Workload

* **Topology** - three independent `TestCluster` instances (`site-0`, `site-1`, `site-2`) wired through `MultiSiteClusterFixture`, each with its own `MemoryGrainStorage`-backed Lattice and a per-site `ReplicationApplier` driven by the chaos delivery pump.
* **Author phase** - every site concurrently runs 6 local `SetManyAtomicAsync` sagas (4 keys per saga, deterministic per-saga key prefix), so each saga emits 4 prepared `Set` records + 1 `TxCommit` per touched shard onto its source site's WAL.
* **Partition cycle** - `site-0`'s loop isolates `site-2` after the first third of its workload and heals it after the second third. The `ChaosDeliveryPump` continues polling but does not advance the cursor on the partitioned edges, so the prepares and terminals authored during the outage queue at the source and ship en bloc after heal.
* **Drain phase** - after every author task completes, `pump.HealAllAndDrainAsync` heals every edge and waits for every per-edge cursor to catch up to its sender's WAL tail, with a 60 s timeout.

### Pass criteria

* On every receiver site, for every authored saga: the count of visible keys is either `0` (saga not yet shipped or aborted) or `KeysPerSaga` (saga fully visible). Any partial-visibility count is a saga-atomicity violation and fails the test.
* On every receiver site, every saga's keys are present after the drain - every authored saga is a local commit, so universal visibility is the strong post-drain assertion.
* `pump.PumpErrors` is empty - a transient grain failure during the run surfaces here without aborting the loop, but the convergence assertion remains the source of truth.

### Tolerated transients

The chaos pump's per-edge loop catches and queues transient grain exceptions onto `PumpErrors`; only sustained faults that prevent convergence within the drain timeout fail the test. The universe is small (3 sites x 6 sagas x 4 keys = 72 keys) so authoring completes in seconds and the drain typically settles in under a second after heal.

### Companion observability

Saga writes emit `orleans.lattice.atomic_write.duration` / `orleans.lattice.atomic_write.batch_size` on the authoring site (see [Metrics](metrics/instrument-catalog-7.md#saga--coordinator--lifecycle)). On the receiver side every inbound entry's apply attempt records `orleans.lattice.replication.apply.duration`, tagged with the tree, the source peer's cluster id (`peer`), the apply outcome and the owning tenant.

## Test 9

Four of the replication suite's per-merge-mode convergence fixtures -
last-writer-wins, OR-Set, PN-Counter and MV-Register; the suite covers
every other `LatticeMergeMode` value too, catalogued in
[the replication chaos tests](../lattice.replication/chaos-tests.md) -
prove that the producer-side change-feed → shipper → receiver-side
applier pipeline converges every site to the same final state under
concurrent multi-site writes and mid-workload partitions. Every
fixture wires three sites through `MultiSiteClusterFixture` /
`ChaosDeliveryPump`, declares the test tree under the relevant merge
mode, lets every site author a disjoint workload while one site is
isolated and re-healed mid-flight, drains the pump, then asserts the
mode-specific convergence invariant pointwise across sites.

### What they prove

| Fixture | Mode | Convergence invariant | Mechanism under test |
|---|---|---|---|
| `LwwRegisterConvergenceChaosTests` | `LwwRegister` | Every site reads the same `VersionedValue` after drain - the lexicographic `(HLC, originClusterId)` winner across all authored writes | The receiver applies each shipped write through its replication-apply seam with the source HLC, so LWW resolution picks the same winner on every site; per-edge change-feed cursors do not skip entries across partition heal |
| `OrSetConvergenceChaosTests` | `OrSet` | Every site's `OrSet(key).GetAsync()` yields exactly the union of authored adds (test 1), or the union of authored adds minus the union of authored removes (test 2) | The receiver folds each shipped typed delta into its local state under `LatticeOriginContext` (a bootstrap entry that carries full state instead of a delta is merged and written back by compare-and-swap); OR-Set's commutative-monoid merge absorbs out-of-order receive |
| `PnCounterConvergenceChaosTests` | `PnCounter` | Every site's `PnCounter(key).ValueAsync()` returns the same algebraic sum of authored deltas | The same receiver-side typed-delta fold as OR-Set; per-replica P/N maps merge by component-wise max |
| `MvRegisterConvergenceChaosTests` | `MvRegister` | Every site's `MvRegister<T>(key).ValuesAsync()` yields exactly the dot-frontier expected from the authored history: concurrent writes survive as a multi-value set, and any write whose dot is causally dominated by a later writer's observed context is superseded on every replica | The same receiver-side typed-delta fold under `LatticeOriginContext`; dot-context merge drops dominated entries and pointwise-maxes the per-replica context maps |

### Workload (per fixture)

* **Topology** - 3 `TestCluster` instances wired through `MultiSiteClusterFixture` declared with the fixture's merge mode; `ChaosDeliveryPump` drives every inter-site edge.
* **Author phase** - every site authors a disjoint family of writes against a single key (`k`):
  * LWW: 40 sequential `SetAsync` calls per site.
  * OR-Set test 1: 25 sequential `AddAsync` calls per site.
  * OR-Set test 2: 15 adds + 2 observed-removes per site.
  * PN-Counter: 30 increments + 10 decrements per site.
  * MV-Register test 1: every site is isolated before any write, each writes one value concurrently, and all three values survive the heal as a multi-value set.
  * MV-Register test 2: two-phase scenario - site 0 issues two sequential `SetAsync` calls and drains so every peer observes its dot context; then sites 1 and 2 write concurrently behind a partition that isolates site 2 from site 1, producing two surviving concurrent dots that both dominate the site-0 entry.
* **Partition cycle** - one site's loop isolates a target site for part of its workload, and both the window and the (driver, target) pair vary per fixture: site 2 isolates site 1 over the middle half of the LWW writes; site 0 isolates site 2 from half-way to its last add in OR-Set test 1, and over the middle third in test 2; site 1 isolates site 0 over the middle third for PN-Counter; in MV-Register every site is isolated for test 1's concurrent writes, and site 2 for test 2's concurrent phase.
* **Drain phase** - after every author task completes, `pump.HealAllAndDrainAsync(30 s)` heals every edge and waits for every per-edge cursor to catch up.

### Pass criteria

* LWW: pointwise equality of `(Value, Version)` across all 3 sites after drain.
* OR-Set: pointwise set-equivalence of `Elements()` across all 3 sites against the closed-form expected union.
* PN-Counter: pointwise equality of `ValueAsync()` across all 3 sites against `SiteCount * (IncrementsPerSite - DecrementsPerSite)`.
* MV-Register: pointwise set-equivalence of `ValuesAsync()` across all 3 sites against the closed-form expected frontier (surviving concurrent dots only).

### Tolerated transients

* PN-Counter and MV-Register wrap `IncrementAsync` / `DecrementAsync` and `SetAsync` in a call-site retry (8 attempts with linear backoff, then one final unretried call) that catches an `InvalidOperationException` carrying `CAS budget exhausted`. The wrapper is defensive only: both accessors now apply one producer-side delta per call - a single read, then one `ApplyCrdtDeltaAsync`, with no compare-and-swap loop - so neither raises that exception, and a foreign-origin merge that lands mid-call cannot abort the local write.
* OR-Set / LWW: no per-call retry needed - the pipeline is non-blocking from the producer side; transient pump errors queue onto `pump.PumpErrors` and surface in the post-drain assertion if convergence stalls.

### Companion observability

Producer-side: the `orleans.lattice.replication.ship.*` family (see [Metrics](metrics.md)) - the `ship.duration`, `ship.effective_batch_size` and `ship.ack_latency` histograms plus its payload, manifest and dictionary counters. Receiver-side: `orleans.lattice.replication.apply.duration`, tagged with the tree, the source peer's cluster id (`peer`), the apply outcome and the owning tenant - it carries no merge-mode tag, so each fixture's streams are told apart by its tree - and a 3-site run records two `peer` values on each receiver site (2 inbound edges each), six in all.

## Test 10 - Multi-site fixture smoke (`MultiSiteClusterFixtureSmokeTests`)

These are not chaos tests in the bombardment sense - they are
deterministic single-write smoke tests that pin the simpler
invariants the cross-cluster convergence suite relies on. They ship
under `[Category("Chaos")]` so they run in the chaos tier alongside the
suites they diagnose; a failure here
short-circuits failure analysis on the larger convergence /
atomic-visibility fixtures by isolating which seam broke.

### What they prove

| Test | Invariant |
|---|---|
| `Site_change_feed_yields_locally_authored_lww_entry` | A locally-authored `SetAsync` lands on the per-site producer-side WAL, the local `IChangeFeed` yields the captured `WalRecord`, and the record carries the site's `OriginClusterId` and declared `LatticeMergeMode` |
| `Delivery_pump_ships_lww_entry_from_site_0_to_site_1` | A locally-authored `SetAsync` on site 0 propagates end-to-end through the chaos delivery pump to site 1, where a subsequent `GetAsync` returns the authored value |

### Tolerated transients

None - these tests run on a quiescent 2-site fixture with no partition / concurrency. Any exception is fatal.

## Test 11 - Cross-tree atomic write under shard churn (`ChaosCrossTreeAtomicWriteIntegrationTests`)

This test asserts the **cross-tree all-or-nothing visibility invariant**
for `SetManyAtomicAsync`: a key committed by a cross-tree saga
is present in **every** participating tree or in **none** of them - a
reader must never observe a partial cross-tree commit (some trees
flipped, others not), even while shard splits move keys between shards
mid-saga on every tree.

A single commit worker writes a fresh per-generation key into the same
logical slot of three distinct trees as one all-or-nothing cross-tree
saga; three reader workers continuously probe committed generations
across all three trees; and one split coordinator per tree churns shards
continuously throughout the ~15 s window.

### What it proves

| Invariant | Mechanism under test |
|---|---|
| A settled committed generation key is present in all three trees | Coordinator's single global decision flips visibility uniformly across every participating tree's `ITxRegistryGrain` via delegation |
| A settled committed key never vanishes under concurrent split churn | Per-tree finalize promotes prepared entries into the main store before settlement; splits are read-transparent for promoted committed keys |
| Every committed generation is durably present in all trees post-window | Cross-tree all-or-nothing completeness across the whole run |

### Workload

* **Seed phase** - 200 `SetAsync` keys per tree across three trees (`xct-{run}-0..2`), each on a four-shard fixture (`FourShardClusterFixture`) so splits are enabled.
* **Commit worker** - one cross-tree saga per generation at a 50 ms cadence: a `LatticeTreeBatch` per tree writing `x-{gen}` into all three trees under operation id `xctop-{run}-{gen}`. Asserts each saga returns `Committed`.
* **Reader workers** - three readers repeatedly probe the six most recently **settled** generations (ending at the commit frontier minus a `SettleMargin` of 5) in all three trees; a settled generation that reads absent in any tree is a failure.
* **Split coordinators** - one per tree at a 250 ms cadence, picking a random non-empty shard and driving `SplitAsync` + `RunSplitPassAsync`.

### Pass criteria

* Zero invariant violations (no settled committed key absent from any tree).
* Post-window: every committed generation `0..lastCommitted` present in all three trees.
* The commit worker committed at least one saga, readers performed at least one probe, and split coordinators attempted at least one split (proves real concurrency).

### Why the invariant is checked only on settled generations

The cross-tree feature guarantees all-or-nothing *commit*, not a
simultaneous cross-tree read snapshot. Each per-tree `GetAsync` is an
independent point-in-time read, so the global decision flip landing
*between* a reader's three samples yields a benign transient skew that is
not a partial-commit violation. The reader therefore checks monotonicity
only on generations the commit frontier has advanced a safety margin
past, by which point every participant's finalize has promoted the entry
into its main store. Post-window completeness then closes the gap for the
in-flight tail.

### Tolerated transients

The stale shard-routing signal, `EnumerationAbortedException`, cursor
snapshot expiry or pin exhaustion, `TimeoutException`, and the saga's own
rollback / topology-churn / fan-out-saturation / "fewer than 2 virtual
slots" `InvalidOperationException` messages are bucketed as transient -
the commit worker moves on to its next generation, and the readers and
split coordinators keep looping; any other exception fails the test.

## Test 12 - Cross-cluster cross-tree atomic visibility (`CrossClusterCrossTreeAtomicVisibilityChaosTests`)

This test (in the replication suite) asserts the **receiver-side
cross-tree all-or-nothing invariant**: when a cross-tree
`SetManyAtomicAsync` batch spanning two replicated trees ships to a remote
cluster, a reader on the receiver must never observe one tree committed
while a sibling tree is still pre-saga. Because each tree's terminals
replicate on their own WAL feed, the receiver routinely applies one tree's
terminal before the other's; the receiver coordinator grain
(`ILatticeCrossTreeReceiverGrain`) must hold every participating tree
invisible until all of the batch's replicated terminals have arrived, then
flip them together.

### Workload

* Two sites, each authoring `SetManyAtomicAsync` batches spanning two
  replicated trees (`chaos-xt-a`, `chaos-xt-b`) under deterministic
  per-batch key namespaces and a stable `operationId`.
* Two independent `ChaosDeliveryPump` instances - one per tree - ship the
  two trees' change feeds with **independent** partition cycles, so the
  receiver routinely sees one tree's terminal before the other's.
* A mid-workload partition isolates each tree's edges at staggered points
  and heals them later, forcing terminal skew across the two trees.

### Pass criteria

* On every receiver site, for every authored batch, tree A's keys and
  tree B's keys have **identical** presence - both fully visible or both
  fully absent. Tree A visible while tree B is absent (or vice versa) is a
  cross-tree atomicity violation.
* No single-tree partial view (every batch's per-tree key set is wholly
  present or wholly absent).
* Post-drain: every batch's keys are present on both trees on every site.

## Test 13 - Atomic write under enforcement-policy churn (`AtomicWriteUnderPolicyChurnChaosTests`)

This test (in the `lattice.schema` suite) asserts that **cross-tree
`SetManyAtomicAsync` stays all-or-nothing while a tree's schema
enforcement policy is changed live**. Schema enforcement validates the
whole batch once, up front at the coordinator, against whatever policy is
current at admission - so concurrent policy churn may change *which*
outcome a given saga gets, but must never split a batch.

A committer drives one cross-tree saga per generation across two trees -
one leg always compliant, the other alternating compliant (odd
generations) and deliberately non-compliant (even generations) - while a
churner flips each tree's policy between "require JSON" and "no policy" at
a 5 ms cadence, and a reader polls both trees on every generation.

### What it proves

| Invariant | Mechanism under test |
|---|---|
| A committed saga installs every leg on every tree | Admission-time whole-batch validation decides the saga atomically before any prepare |
| A rejected saga mutates no tree | A single admission failure rejects the whole batch; no leg is applied |
| A reader never sees a torn per-tree snapshot | Atomic-write reader isolation holds under concurrent policy mutation |

### Workload

* **Committer** - one cross-tree `SetManyAtomicAsync` saga per generation (60 generations) over `treeA` (always compliant) and `treeB` (compliant on odd, non-compliant on even generations), each value embedding its generation.
* **Policy churner** - flips each tree's policy between a `LatticeSchemaRule.Json()` policy and no policy at a 5 ms cadence.
* **Reader** - a concurrent per-generation read of both trees, asserting each tree's key set is at one coherent generation (never a torn cross-generation mix).

### Pass criteria

* Zero invariant violations: every committed generation is present on both trees on every key; every rejected generation advanced neither tree.
* At least one generation committed and at least one was rejected by the churned policy (proves the race actually exercised both outcomes).

### Tolerated transients

The committer treats a `LatticeSchemaViolationException` (or a non-committed outcome) as the reject branch; any other exception fails the test.

## Test 14 - Atomic write under schema-version advance and eager migration (`AtomicWriteUnderVersionAdvanceChaosTests`)

This test (in the `lattice.schema` suite) asserts that **single-tree
`SetManyAtomicAsync` stays all-or-nothing and every stored value stays
self-describing and decodable while a versioned tree's target schema
version is advanced concurrently, and that the one-call eager background
migration preserves that atomic snapshot when it re-stamps the tree**. It
runs in two phases matching the two distinct version-change operations and
their concurrency contracts.

A target-version advance is a config-only change (it moves no data), so it
is safe to run concurrently with live writes: new writes stamp at the new
target and existing values are upcast on read. The eager data migration is
a shadow-build with alias cutover whose v1 contract requires the tree be
write-quiescent, so it is validated in a quiesced window - not hammered
under live writes.

### What it proves

| Invariant | Mechanism under test |
|---|---|
| Every atomic batch lands as a unit under a concurrent version advance | The write interceptor stamps each delta at admission; the advance never tears the saga |
| A read never returns raw enveloped bytes and never throws for a reachable version | The read decoder strips the envelope and upcasts to the current target through the registered hop chain |
| Eager migration preserves the last committed snapshot | The quiesced shadow-build re-stamps each value from its own version to the target and cuts the alias over atomically |

### Workload

* **Committer** - single-tree `SetManyAtomicAsync` batches (60 generations) of 12 keys into one versioned tree, each value embedding its generation.
* **Version churner** - advances the target schema version v1 -> v2 -> v3 (once each) concurrently with the write loop.
* **Reader** - a concurrent per-generation read that fails on any raw enveloped byte leak or torn batch.
* **Quiesced migration** - after the write loop drains, one `MigrateToTargetVersionAsync` re-stamps every value (written across v1..v3) to the current target, followed by an idempotent second call.

### Pass criteria

* Zero invariant violations across the concurrent phase (all-or-nothing, decodable, no raw-envelope leak).
* The target advanced exactly twice (v1->v2->v3).
* The eager migration succeeded, preserved the last committed generation on every key, and the idempotent re-migration also succeeded.

### Tolerated transients

None beyond the churner's own cancellation; any unexpected exception during the advance, the write loop, or the migration fails the test.

## Observed recovery surfaces

Between them, the chaos tests exercise every recovery path documented
in [shard-splitting.md](shard-splitting.md),
[online-reshard.md](online-reshard.md),
[tree-sizing.md](tree-sizing.md),
[tombstone-compaction.md](tombstone-compaction.md),
[wal.md](wal.md),
[wal-storage-providers.md](wal-storage-providers.md),
[state-primitives.md](state-primitives.md),
[../lattice.replication/README.md](../lattice.replication/README.md),
[../lattice.replication/transport.md](../lattice.replication/transport.md),
[../lattice.replication.grpc/README.md](../lattice.replication.grpc/README.md),
and the architecture notes:

The table below covers the four full-workload topology-mutation chaos fixtures (Tests 1-4). The atomic-write reader-isolation (Test 5), atomic-visibility-across-topology siblings (Test 6), digest-determinism (Test 7), cross-cluster atomic-visibility (Test 8), per-mode convergence chaos (Test 9), multi-site smoke (Test 10), and the per-test invariant fixtures (range delete, CAS, scan cancel, multi-silo restart, WAL trim, liveness + inbound stats, OR-Map convergence, compaction + shipping, gRPC transport, Azure Table WAL) target orthogonal invariants and are documented in their own grids below.

| Surface | Happy path | Faults | Resize | Reshard |
|---|:---:|:---:|:---:|:---:|
| Concurrent reads/writes during split shadow phase | ✅ | ✅ | - | ✅ |
| `ScanKeysAsync` / `ScanEntriesAsync` in-line reconciliation | ✅ | ✅ | ✅ | ✅ |
| `CountAsync` per-slot routing + version stability + bounded retry | ✅ | ✅ | ✅ | ✅ |
| `StaleShardRoutingException` transparent retry | ✅ | ✅ | - | ✅ |
| `StaleTreeRoutingException` transparent retry across alias swap | - | - | ✅ | - |
| Permanent `MovedAwaySlots` rejection after split completion | ✅ | ✅ | - | ✅ |
| Resumable `SplitInProgress` intent replay across crashes | - | ✅ | - | - |
| Two-phase root promotion (`PendingPromotion`) replay | - | ✅ | - | - |
| Shadow `MergeManyAsync` atomicity under failed source write | - | ✅ | - | - |
| Registry `ShardMap.Version` stamping under retry | - | ✅ | - | ✅ |
| Idempotent drain chunks | - | ✅ | ✅ | ✅ |
| `TreeResizeGrain` phase machine under live traffic | - | - | ✅ | - |
| Per-source-shard shadow-forwarding under live traffic | - | - | ✅ | - |
| `TreeReshardGrain` migration loop + dispatch-budget clamping | - | - | - | ✅ |

### Saga-atomicity surfaces (Tests 5, 6, 8)

| Surface | Atomic vis. quiescent | Split + saga | Resize + saga | Reshard + saga | Cross-cluster |
|---|:---:|:---:|:---:|:---:|:---:|
| Continuous reader observes zero-or-all keys per poll | ✅ | ✅ | ✅ | ✅ | ✅ |
| `AtomicWriteGrain` per-shard terminal broadcast idempotent under retry | ✅ | ✅ | ✅ | ✅ | ✅ |
| Prepared-write commit path holds keys hidden until terminal arrives | ✅ | ✅ | ✅ | ✅ | ✅ |
| Saga prepares shadow-forwarded onto destination shard mid-split | - | ✅ | - | ✅ | - |
| Saga prepares shadow-forwarded onto destination physical tree mid-resize | - | - | ✅ | - | - |
| Retroactive prepared-mutation sweep at `BeginShadowWrite` | - | - | - | ✅ | - |
| Saga terminal-fan-out shadow-forward fallback via post-Complete `MovedAwaySlots` | - | - | - | ✅ | - |
| Registry `TxDecisionRetention` tombstone absorbs duplicate terminals | - | - | - | ✅ | ✅ |
| Receiver-side prepared/terminal apply seam holds prepares off projection | - | - | - | - | ✅ |
| `LatticeHlcOverrideContext` preserves source HLC on receiver | - | - | - | - | ✅ |
| `ShouldShip` bypass for `TxCommit` / `TxAbort` records | - | - | - | - | ✅ |

### Per-mode convergence surfaces (Test 9)

| Surface | LWW | OR-Set | PN-Counter | MV-Register |
|---|:---:|:---:|:---:|:---:|
| Producer-side change feed yields locally-authored mutations with origin id | ✅ | ✅ | ✅ | ✅ |
| Shipper drives delivery to every peer site under partition-cycled topology | ✅ | ✅ | ✅ | ✅ |
| Receiver-side mode-specific dispatch (source-HLC LWW apply / typed-delta fold) | ✅ | ✅ | ✅ | ✅ |
| `LatticeOriginContext` flagging foreign-origin writes during apply | ✅ | ✅ | ✅ | ✅ |
| Commutative-monoid CRDT merge absorbs out-of-order receive | - | ✅ | ✅ | ✅ |
| Dot-context supersession of causally-dominated entries on merge | - | - | - | ✅ |

### Per-test invariant surfaces (range delete, CAS, scan cancel, restart, replication, storage)

The chaos tests below target surfaces that the four full-workload
topology grids above do not cover: per-call public-API invariants
(range delete, CAS, scan cancellation), Orleans membership churn
(multi-silo restart), the production replication pipeline (WAL trim,
liveness + inbound stats, OR-Map convergence,
compaction + shipping), and the two downstream-package suites (gRPC
transport, Azure Table WAL). The range-delete, CAS, scan-cancel and
multi-silo-restart columns map to rows of the core suite table at the top
of this document; the WAL-trim, liveness, OR-Map, compaction, gRPC and
Azure Table WAL columns map to fixtures catalogued in
[the replication chaos tests](../lattice.replication/chaos-tests.md).

| Surface | Range delete | CAS | Scan cancel | Multi-silo restart | WAL trim | Liveness + inbound | OR-Map | Compaction + shipping | gRPC transport | Azure Table WAL |
|---|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|
| `DeleteRangeAsync` range exclusivity under concurrent writers | ✅ | - | - | - | - | - | - | - | - | - |
| `DeleteRangeAsync` cross-shard fan-out + tombstone scope | ✅ | - | - | - | - | - | - | - | - | - |
| `SetIfVersionAsync` (CAS) linearisable winner under contention | - | ✅ | - | - | - | - | - | - | - | - |
| `SetIfVersionAsync` (CAS) lost update returns `false`, and the re-read-and-retry loop loses no increment | - | ✅ | - | - | - | - | - | - | - | - |
| `ScanKeysAsync` / `ScanEntriesAsync` cooperative cancellation surfaces `OperationCanceledException` | - | - | ✅ | - | - | - | - | - | - | - |
| Cancelled scan leaks no grain-side enumerator state | - | - | ✅ | - | - | - | - | - | - | - |
| `TestCluster.RestartSiloAsync` mid-workload preserves universe (count, envelope) | - | - | - | ✅ | - | - | - | - | - | - |
| Process-shared `IGrainStorage` isolates membership churn from storage disappearance | - | - | - | ✅ | - | - | - | - | - | - |
| Producer-side WAL trim cannot prune un-acked entries | - | - | - | - | ✅ | - | - | - | - | - |
| Real `AddLatticeReplication` + loopback transport under sustained writes | - | - | - | - | ✅ | ✅ | - | ✅ | - | - |
| Outbound liveness probe: `peer.last_contact_seconds` climbs while a peer edge is isolated and falls below half a probe interval within ten probe intervals of the heal | - | - | - | - | - | ✅ | - | - | - | - |
| Budget-driven receiver-side apply faults drain in full, the receiver still converges, and the inbound peer-stats row records contact with no outbound-only backlog or in-flight counts | - | - | - | - | - | ✅ | - | - | - | - |
| OR-Map convergence under concurrent multi-site mutation + partition | - | - | - | - | - | - | ✅ | - | - | - |
| Per-tree typed-CRDT shape resolution end-to-end on producer + receiver dispatch | - | - | - | - | - | - | ✅ | - | - | - |
| `ReplicationShipperGrain.ShouldShip` keeps `MutationKind.Tombstone` envelopes off the wire | - | - | - | - | - | - | - | ✅ | - | - |
| `ITombstoneCompactionGrain.RunCompactionPassAsync` mid-shipment preserves receiver convergence | - | - | - | - | - | - | - | ✅ | - | - |
| gRPC sender retry loop under 15% and 30% per-call channel faults delivers every attempted entry | - | - | - | - | - | - | - | - | ✅ | - |
| Re-deliveries the channel faults induce are absorbed as no-ops by the receiver's exact-identity dedupe cache or the idempotent leaf apply | - | - | - | - | - | - | - | - | ✅ | - |
| Azurite-backed WAL keeps every shard's offsets dense, gap-free and duplicate-free under sustained concurrent appends across shards | - | - | - | - | - | - | - | - | - | ✅ |

Legend: ✅ = covered by a live test.

## Runtime characteristics

Every chaos suite shares the same shape: a short chaos window (a few seconds of
sustained workload with topology churn and / or fault injection), a bounded
heal / drain phase, and a post-quiescence invariant assertion. Per-suite cost
scales with the workload described in each suite's "What it proves" / "Purpose"
column rather than being fixed, so this section stays accurate as suites are
added without a per-suite runtime table to maintain.

- Single-cluster suites (`test/lattice/BPlusTree/`) run key universes of 8 to
  600 keys (16 keys for the saga-atomicity suites) across up to ~16 parallel
  workers - writers, scanners, and a topology or saga driver - with a chaos
  window of roughly 6-20 s (up to 90 s for the shard-consolidation suite) and
  a heal / drain budget up to ~60 s for the saga-atomicity suites. Most
  trees start at 4 shards (64 where a fixture does not pre-register its
  tree; the multi-silo restart and bounded-cache fixtures pin 2 and 1) and
  grow under split or reshard; the shrinking reshard run folds its tree
  from 4 shards to 2.
- Cross-cluster suites (`test/lattice.replication/Chaos/`) run two or three
  in-process clusters over a fault-injectable inter-site delivery layer, with a
  single-key or single-tree universe (the atomic-visibility suite uses ~72
  keys), a chaos window of a few seconds with a site or edge partitioned
  mid-workload, and a drain budget up to ~60 s. See
  [the replication chaos tests](../lattice.replication/chaos-tests.md) for the
  per-suite catalog.

The cross-cluster, gRPC, and Azure Table suites drive the real replication
pipeline (and the real gRPC transport service and Azure Table provider where
applicable) rather than test doubles of that logic, but they run in-process:
in-process test clusters with a simulated delivery pump, an in-memory test
server, and the Azurite emulator stand in for networked silos and a real cloud
storage account.

## Derived-state recovery across an identity swap

Derived state that tails a source tree's write-ahead log - a folded /
materialised view, and a tag index maintained in a sibling index tree - must
recover when the source's PHYSICAL identity is repointed under its logical
registry alias (a restore-style cutover, possibly repeated). Two chaos suites,
each hosted in the sibling package whose fixture already stands up the machinery,
cover this and are tracked by
[issue #1167](https://github.com/NSTA1/Orleans.Lattice/issues/1167):

| Suite | Where | What it proves |
|---|---|---|
| Materialised-view recovery across an identity swap | `test/lattice.replication/Chaos/DerivedStateRecoveryAcrossIdentitySwapChaosTests.cs` | A folded view tails a logical source tree whose physical identity is repointed under its registry alias repeatedly, under a sustained mutation workload that accretes backlog between drains. The maintainer rebinds to the source's current physical id event-driven when the tree registry pushes an alias change (with a coarse backstop re-resolve as a safety net), and on a change rebuilds against the new physical source and rebinds its tail. After the workload quiesces the view reflects exactly the final identity's contents: keys dropped by a cutover are retracted (a WAL tail alone can never retract a key the restored source never had), changed values win, and a large abandoned backlog never survives. Complements the deterministic single-swap regression that landed with the maintainer heal. |
| Tag-index reconcile under repeated restore | `test/lattice.backup/Chaos/BackupRestoreReconcileChaosTests.cs` | A tag index over a live subject tree is driven through a real shadow-cutover restore under a large tagged working set with a concurrent reader hammering the tag query. The restore fires a prompt reconcile that drops every membership row absent from the restored point-in-time; the reader never observes an out-of-universe key or a torn membership. A second case injects one aborted catalogue enumeration into the restore's tag-index discovery and requires the restore to recover that scan rather than swallow the abort and leave every orphan row in place (issue #3233). Repeated restores to successively earlier point-in-times narrow membership monotonically to each restore's subject with no row stranded from a superseded restore. |

## Routing and cross-tree recovery across a shadow-cutover restore

A shadow-cutover restore swaps a logical tree's alias to a freshly loaded shadow
tree while retaining the previous physical tree so the restore can be reverted.
Two consumers must recover across that swap: stateless-worker routing
activations that cache the logical-to-physical alias, and an in-flight
cross-tree atomic-write saga that drives a participant through that alias. Two
chaos suites, hosted in the backup package whose fixture stands up both the
restore machinery and the core lattice, cover this and are tracked by
[issue #1167](https://github.com/NSTA1/Orleans.Lattice/issues/1167):

| Suite | Where | What it proves |
|---|---|---|
| Routing self-heal across a shadow-cutover | `test/lattice.backup/Chaos/ShadowCutoverRoutingSelfHealChaosTests.cs` | Many stateless-worker routing activations are warmed against a tree's pre-cutover identity under sustained concurrent reads, then the tree is cut over to a restored shadow while the readers keep hammering. The retained tree refuses logical-alias traffic with the internal staleness signal, which the routing tier catches and re-resolves onto the shadow. No reader ever surfaces that signal or observes an out-of-universe value, and after the cutover every activation converges onto the restored snapshot with no key left resolving to pre-cutover content. The repeated-cutover case restores to successively earlier point-in-times and proves the heal re-arms for each. |
| Cross-tree atomic write across a participant cutover | `test/lattice.backup/Chaos/CrossTreeAtomicWriteAcrossCutoverChaosTests.cs` | A sustained stream of cross-tree atomic writes spans two trees while one participant is repeatedly cut over to a restored shadow underneath the saga. The saga's deadline-bounded retry loops must absorb the staleness signal the retired physical tree raises, re-resolve onto the shadow, and reach a clean terminal decision - no call leaks the signal to the caller. After the cutover storm a fresh all-or-nothing batch still commits and is atomically visible on both trees, proving the cut-over participant self-healed and is not wedged. Complements the mocked prepare-side and terminal-broadcast stale-routing retry regressions by exercising the whole coordinator over a real restore. |

## See also

* [Consistency](consistency.md) - the per-operation guarantees these
  tests verify against the public API under topology mutation.
* [Replication](../lattice.replication/README.md) - the producer-side change-feed →
  shipper → receiver-side applier pipeline that the cross-cluster
  suite exercises.
* [Adaptive Shard Splitting](shard-splitting.md) - the split protocol
  exercised by the happy-path, faults, reshard, and split + saga
  tests.
* [Online Reshard](online-reshard.md) - the reshard coordinator and
  its interaction with autonomic splits.
* [Tree Sizing](tree-sizing.md#resizing-an-existing-tree) - the online
  resize path exercised by `ChaosResizeIntegrationTests` and
  `ResizeTopologyTests`.
* [Architecture](architecture.md) - grain layers, root promotion,
  bounded retry, and the invariants the chaos tests verify.
* [State Primitives](state-primitives.md) - HLC and LWW, which
  guarantee the value envelope holds even under concurrent rewrites.
