Release 2026-09-05: Added to Changed
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at 2026-09-05-1.md, and llms.txt lists every page.Part of Release 2026-09-05, in Changelog.
Added
The documentation is now a browsable website, published to GitHub Pages. The whole documentation corpus, the runnable samples and their source, and the root README, feature catalogue and package inventory are rendered as one searchable site: full-text search, navigation grouped exactly as
PACKAGES.mdgroups packages by seam andFEATURES.mdgroups capabilities by concern, syntax-highlighted snippets, and rendered Mermaid diagrams. Every sample gains a generated Source page, so a sample'sProgram.csis one click from its README exactly as it is when browsing the repository, and the two largest samples list their remaining files rather than inlining a page nobody can read. Nothing about the in-repo Markdown changes: the site is generated from it, so the repository stays the source of truth and the docs remain readable on github.com. The published release notes carry only what has shipped: the changelog's rolling Unreleased section is filtered out of the site, so a visitor is never told about a change they cannot yet install. The build lives indocs-site/and is published by the newDocsworkflow on the corelattice-v*release tag - once per release wave, from the released commit - or on manual dispatch. The same build runs on every documentation pull request as a link-integrity gate that fails closed on any broken relative link or in-page anchor. (Orleans.Lattice9.6.0)The durable WAL materialiser pin store is now bucketed, and its write latency is a first-class saturation input. The pin store keeps one grain-state blob per shard holding an entry per routed leaf consumer, and Orleans writes grain state whole, so a single flush was O(consumers per shard) - 1.33 MB on an observed 7,841-file tree.
WalMaterialiserPinGrainnow splits that state acrossLatticeOptions.WalMaterialiserPinBucketCountbuckets, routing each consumer through a MurmurHash3fmix32avalanche before the bucket modulo so the bits selecting the bucket carry no information about the bits selecting the shard. That avalanche is load-bearing rather than cosmetic: a shard activation only writes the consumers routed to it, so with the naturalbucketCount == shardCountconfiguration an unmixed hash selected both ordinals identically and bucketing degenerated into renaming the state slot - one observed 8-shard, 8-bucket deployment held 1,093,492 bytes in one bucket and 484 in another, with bucket membership tracking the shard ordinal perfectly. The default stays 1, so no existing deployment changes on upgrade, and rebucketing is safe whenever it is raised: an activation reads every bucket and merges monotonic-max, so a copy stranded in an old bucket loses to the newer one and the worst case is dead bytes and over-retention, never over-trim. Alongside it the durable pin path becomes observable and sheddable.WalSaturationSamplermeasures pin-write latency and feeds it into the saturation signal, reporting through the newWalSaturationCause; only the steady-state per-checkpoint write is sheddable, since the birth block-pin seed and the deactivation flush are last-chance writes and dropping the flush that produces a leaf's first real frontier would leave no durable floor at all. A shed window requires aShedTriggerFloorMsfloor so that a cheap synchronous fault - whose duration is no evidence about the store's service rate - cannot shed the very retry the debounce rollback contract exists to guarantee, and a shed report now rolls its debounce back so shedding is a pure deferral rather than a silent loss. Both new instruments carry the derived tenant dimension, are documented in the dashboards panel map, and have Grafana panels on the commit-path dashboard. (Orleans.Lattice9.6.0,Orleans.Lattice.Dashboards9.6.0)
Changed
Four unreachable empty-shard guards are removed from the shard root's read path.
GetShardProjectionDigestAsync,GetShardProjectionDigestForRangeAsync,GetTopologySnapshotAsyncandWarmUpAsynceach awaitedPrepareForOperationAsyncand then re-checked whether the shard still had no root. That check could never be true: the prepare either takes its fast path, which requires a non-null root, or runsEnsureRootAsync, which seeds and persists a deterministic root leaf, or throws - and nothing in the library ever assigns a null root back. The guards were dead from the commit that introduced them, and two of them documented an "empty shard" result the surrounding code cannot produce, which in turn left a deadsnapshot is nullbranch inLatticeStateQuery's per-shard structure scan. All five sites are gone, the shard-root topology read is now typed non-nullable to match what it always returned, and a new unit fixture pins the real behaviour against a never-written shard - all four entry points seed a root through the prepare and report that seeded root leaf - so the removed assumption cannot be reintroduced silently. No behavioural change: every removed branch was unreachable, and the three near-identical guards that are reachable (PurgeAsync,ReseedNodeBindingsAsync,GetRootNodeRefAsync, none of which prepares first) are untouched. (#1996) (Orleans.Lattice9.6.0,Orleans.Lattice.Api.State9.6.0)The six background-coordinator leaf walks now share one bounded-walk implementation, and every resume cursor is a key. The split drain, the cross-tree merge drain, the online snapshot copy, tombstone compaction and the shard-consolidation drain each walked a shard's leaf chain their own way: four were unbounded and two hand-rolled their own budget, so the same known-good pattern existed in three divergent copies and half the sites had no bound at all. An unbounded pass here does not head-of-line-block a shard root the way an unbounded read walk does, but it does hold the coordinator's own non-reentrant activation for the whole sweep, so a thousand-leaf shard left the coordinator unable to answer a progress query or honour a cancellation, buffered the sweep's working set throughout, and put the entire sweep at risk from a single interruption. All six now route through one helper that owns the budget, the stop condition and the resume position, governed by two new options,
BackgroundDrainLeavesPerPass(default 64) andBackgroundDrainMaxDuration(default 10 seconds), while the compactor and the consolidator keep their own long-standing per-pass leaf caps. Every resumable walk persists its position as a key, never a leaf grain id: Orleans grains are virtual, so an id persisted across a pass boundary can activate a fresh, empty grain whose null sibling pointer a resumed walk reads as a clean end of chain, silently leaving the rest of the shard unvisited - which is why the compactor previously had to defend that cursor with a liveness probe (#1970); a key is instead re-descended onto whichever leaf now owns it, so a leaf split or reclaimed between two passes cannot truncate the sweep, and the probe is gone. A walk yields only where it can name that position, so a misconfigured bound degrades to the unbounded walk rather than to a truncated one. Atomicity is preserved wherever it was load-bearing, and the reasoning is recorded at each site rather than left implicit: the split's two authoritative post-freeze sweeps, its retroactive prepared-mutation sweep, the consolidation's two post-freeze sweeps, and the offline snapshot copy all stay deliberately unbounded, because each either runs against a source that can no longer change while routing flips onto its target, or - for the retroactive sweep - installs the destination-side shadow markers the phase that follows depends on and carries a whole-sweep re-run recovery contract, or - for the offline copy - assembles the destination bottom-up in a single bulk load that by contract needs the complete sorted entry set. Each of those is now attributable through the existing slow-atomic-walk warning instead. A leaf-id cursor persisted by an older build is discarded on load and its shard restarts, which is safe because every one of these walks is a fixed point under LWW merge. (#1973) (Orleans.Lattice9.6.0)The README quick start now writes a typed value instead of a byte array. The first write a new reader saw handed
SetAsyncabyte[], which reads as though encoding every value by hand is the normal way to use the store. The quick start now stores and reads back a record through the typedSetAsync<T>/GetAsync<T>extensions, which default to JSON and need no serializer argument, and presents the rawbyte[]surface afterwards as the opt-in it actually is - for callers who want to own the encoding. (Orleans.Lattice9.6.0)Three steady-state allocations are trimmed from the view-maintainer drain path and the receiver-side causal-apply buffer. All three are output-identical and touch no public API. (1)
ViewKeyCollisionDetector.Detect, which runs unconditionally on every view drain, now presizes its first-source map to the batch count instead of growing it from empty, removing the grow/rehash chain; the colliding list and set stay unpresized because a well-configured injective re-key never collides and so never allocates a backing store. (2)ViewMaintainerGrain.FlushCompletedFilterBatchesAsyncnow coalesces once into a local and presizes theupsertsanddeleteslists to the coalesced-survivor count, which they partition, removing their regrowth churn per completed atomic batch. (3)CausalApplyBuffer.TryAddno longer allocates a fresh empty eviction list on every add - steady state never evicts - and instead hands back a shared empty sentinel, materialising a real list only when an eviction actually occurs; the sole caller reads the list only on the eviction outcome. A new[MemoryDiagnoser]benchmark suite (draintrims) isolates each shape against its prior form: on a 256-entry drain the detector map falls from 22,408 B to 8,432 B, the upsert/delete partition from 8,936 B to 6,256 B, and a 256-add no-eviction burst from 8,192 B to 0 B. (Orleans.Lattice9.6.0,Orleans.Lattice.Replication9.6.0)The view-maintainer drain now folds collision detection and write coalescing in a single pass, and the shared metrics tick stops paying for LINQ and duplicate hash probes. All three changes are output-identical and touch no public API. (1) The drain's steady-state fast path ran
ViewKeyCollisionDetector.DetectandViewWriteCoalescer.Coalesceas two consecutive passes over the same batch on adjacent lines. Both group by the same view key under the same ordinal comparer, so the second pass rebuilt per-key state the first had just discarded: twoDictionaryinstances, a collidingHashSet, and two hash probes for every write. A new internalViewBatchFold.Foldfolds both into one pass over one per-key slot map, carrying the first-seen source key and the leading write together, so each write is hashed once and the colliding set is replaced by a per-slot flag. Both public helpers are unchanged and still serve the range-delete path, which cannot coalesce by view key. First-seen ordering, last-writer-wins tie-breaking (a tie leaves the incumbent in place), and report-each-colliding-key-once are all preserved, including the case where an unattributed write precedes the attributable writes that actually collide; a randomised equivalence test asserts the fold agrees element-for-element with the two helpers across 400 generated batches. (2)SharedMetricsSampler.SampleAllAsyncde-duplicated an explicitly requested tree-id scope withDistinct(StringComparer.Ordinal).ToList()on every tick, which allocates the LINQ iterator, its internal set, and a list grown from empty because a deferred iterator givesToListno count hint; it now folds directly into a list presized to the request, scanning ordinally for the short scopes a dashboard actually names and keeping a set for larger ones so a pathological request cannot go quadratic. First-seen order is preserved, which matters because the tree-id order drives insertion order into the returned metrics dictionary. (3) The same tick probed two dictionaries twice where once suffices:SampleViewLagAsyncaccumulated its per-tree rollup with aTryGetValuefollowed by an indexer set, andResolveViewLagcalledTryGetValuetwice with the same key to read two fields off one struct. The first now usesCollectionsMarshal.GetValueRefOrAddDefault- a missing key yields a defaulted rollup, exactly what the failedTryGetValuehanded back - and the second reads both fields from one lookup. A new[MemoryDiagnoser]benchmark suite (fusiontrims) isolates each shape against its prior form, calling the real production code on both sides of the drain pair: on a 256-write drain over 64 view keys the fold falls from 37,688 B to 28,264 B and from 17,338 ns to 6,990 ns, a 16-id scope dedup falls from 712 B to 184 B at unchanged time, and a 512-row metrics tick falls from 11,430 ns to 6,165 ns at identical allocation. (Orleans.Lattice9.6.0,Orleans.Lattice.Api.State9.6.0)Three hot-path hash-probe reductions: slot routing, view-maintainer staging, and CRDT delta combine. All three are output-identical and touch no public API. (1)
LatticeGrain.BuildOwnedSlotMap, which partitions the virtual slot space by owning physical shard on everyCountAsyncandCountPerShardAsyncfan-out, hashed each virtual slot five times across two intermediateDictionary<int, int>instances: a countingTryGetValueplus indexer set, then a cursor read, a bucket read and a cursor write on the fill pass. Physical shard indices are small, dense and non-negative while the slot array is sized to the virtual slot count (4096 by default), so the partitioner now buckets into owner-indexed arrays and hashes nothing per slot, reusing the count array as a decrementing write cursor and filling top-down so every bucket stays ascending without a sort. This is the same virtual-versus-physical index insightShardMap.GetPhysicalShardIndicesalready applies. A pathologically sparse or negative owner value falls back to the prior hashed form, so no input regresses. (2) The view maintainer's three per-mutation staging accumulators -StagedTransaction.NoteOffset(a min-fold),StagePrepare(a read-to-discount-superseded-bytes followed by an overwrite) andRecordOrdinaryOverStagedKey(a max-fold) - each ran aTryGetValuefollowed by an indexer set on the same key, hashing twice per staged entry; each now resolves in one probe viaCollectionsMarshal.GetValueRefOrAddDefault, whose defaulted miss is exactly what the failedTryGetValuehanded back. (3)CrdtShapeRegistry's pointwise-max delta combine (PointwiseMaxLongandPointwiseMaxHlc), which backs the PnCounter, GCounter and VersionVector shipping-side delta coalesce, carried the same double-probe fold on the branch that raises the running max, which on a coalesce is most entries; both now resolve in one probe. A new[MemoryDiagnoser]benchmark suite (slotfolds) isolates each shape against its prior form, calling the real production code on the optimized side of the slot-partition and delta-combine pairs: partitioning 4096 slots over 16 shards falls from 46,589 ns to 9,890 ns and from 18.03 KB to 17.20 KB, a 256-entry staging batch falls from 14,827 ns to 11,679 ns at identical allocation, and a 32-replica PnCounter delta combine falls from 1,239 ns to 854 ns. (Orleans.Lattice9.6.0)Three more hot-path hash-probe reductions: reshard slot accounting, the tenancy record merge, and the WAL GC pin union. All three are output-identical and touch no public API. (1)
TreeReshardGrain's migrating pass rebuilds a virtual-slot ownership histogram on every tick of an online reshard to decide which physical shards are hot enough to split. It counted into aDictionary<int, int>seeded per physical shard, so every virtual slot in the space (4096 by default) paid a hash read plus a hash write for its increment, and the eligibility scan that followed then hashed a second and third time per physical shard to read the count back and project it into the eligible list. Physical shard indices are small, dense and non-negative while the slot array spans the virtual space, so the histogram now counts into an owner-indexed array and returns the counts aligned to the caller's physical-shard ordinals, which removes the downstream reads as well: nothing on the path hashes at all. This is the same virtual-versus-physical index insightLatticeGrain.BuildOwnedSlotMapandShardMap.GetPhysicalShardIndicesalready apply, and a pathologically sparse index set falls back to a binary search over the ascending index list rather than over-allocating, so no input regresses. (2)TenantRecord.MergeFrom, the composite CRDT join run on every replica-to-replica tenancy reconcile, folded its four slot maps - admin subjects, cross-tenant grants, allowed regions and region statuses - with aTryGetValuefollowed by an indexer set on the same key, hashing twice per merged entry in all four loops; each now resolves in one probe viaCollectionsMarshal.GetValueRefOrAddDefault, whose defaulted miss is exactly what the failedTryGetValuehanded back. (3) The WAL garbage collector's two durable-pin unions (ReadDurablePinsAsyncandReadDurablePinOffsetsAsync), which fold the per-shard pin maps read from everyIWalMaterialiserPinGraininto one lowest-wins map on each GC sweep, carried the same double-probe shape on both the hit and the miss branch; both are now single-probe, and both unions are presized to the widest shard's consumer count, which is a lower bound on the union because the pin routing shards a tree's pins without partitioning its consumers - so the union no longer grows and rehashes from empty. A new[MemoryDiagnoser]benchmark suite (reshardfolds) isolates each shape against its prior form, calling the real production code on the optimized side of the histogram and tenancy-merge pairs: a 4096-slot histogram over 16 physical shards falls from 25,080 ns to 2,641 ns and from 760 B to 600 B, a tenancy merge over 48 subjects, 48 grants and 16 regions falls from 4,637 ns to 2,933 ns at identical (zero) allocation, and an 8-shard, 64-consumer pin and offset union falls from 43,281 ns to 23,766 ns and from 11,432 B to 5,768 B. (Orleans.Lattice9.6.0,Orleans.Lattice.Tenancy9.6.0)Three dense physical-shard fan-out reductions on the write saga, the bulk-load extension and the streaming restore. All three are output-identical and touch no public API. (1)
AtomicWriteGrain.PrepareAsync, run once per transactional multi-key write, bucketed every entry into an unsizedDictionary<int, List<(string Key, int Index)>>: a probe-then-store hash pair per entry, a growth chain per bucket, and a second reference to a key the saga already held, carried only so the per-shard capture could read it back. It now buckets through a new sharedShardFanout.BucketIndiceshelper, which indexes buckets by owning physical shard directly, gives each the shard-fairmin(max(4, n/shards), 256)capacity the other batch paths already use, and buckets bareintindices - a quarter of the element width, and no second key reference, because the capture reads each key back off the entry list it is already given. (2) The publicBulkLoadAsyncextension maintained four parallel dictionaries keyed by the same physical shard index (the per-shard buffer, its in-flight append task, its chunk counter and its shard grain), costing one hash per buffered entry and six more on every chunk flush. All four collapse into one owner-indexed slot object per shard held in a new sharedShardSlots<T>, so the per-entry hash becomes an array read and the per-flush six disappear. (3)LatticeBackupRestoreService's two streaming restore accumulators did the same double-probe per record over the whole restore stream, andMergeApplyAsyncadditionally grew each per-shard merge batch from empty despite flushing it at a known, caller-supplied size - so every batch abandoned its whole bucket and entry array at each doubling on the way there. Both now accumulate throughShardSlots<T>, and each merge batch is presized to its flush bound (clamped, so a caller-configured batch size cannot size an allocation unbounded). Both new primitives keep the existing hashed fallback for a hand-built map carrying a negative or pathologically large physical shard index, so no input regresses, and differential tests assert each new shape against the shape it replaced, on the same input, rather than against hand-written expectations. A new[MemoryDiagnoser]benchmark suite (fanoutslots) isolates each shape against its prior form, calling the real production code on the optimized side of every pair: bucketing a 1000-entry saga batch over 8 shards falls from 45.91 KB to 7.46 KB (and from a 35.0 us to a 22.1 us median), buffering a 10,000-entry bulk load falls from a 69.4 us to a 44.8 us median, and accumulating a 20,000-record restore stream in 4096-record batches falls from 1,685.66 KB to 1,064.56 KB (and from a 1,326.7 us to a 1,231.4 us median). (Orleans.Lattice9.6.0,Orleans.Lattice.Backup9.6.0)Three dense-partitioning reductions on the ordered-scan, cursor and bulk-append paths. All three are output-identical and touch no public API. (1)
LatticeGrain.GroupSlotsByOwner, which every ordered scan and key drain calls to split the virtual slots it still needs across the physical shards that own them, hashed each requested slot twice - once to probe for its owner's bucket and once to store into it - grew aList<int>per owner, then copied each list into a fresh array and sorted the copy. The owner index lives in a tiny dense domain (one entry per shard root, typically 1-16) even though the virtual slot space it is projected from is 4096 wide, so the grouping now counts per owner into an owner-indexed array, allocates each bucket at its exact final size, reuses that count array as the per-bucket write cursor, and sorts each bucket in place: it hashes nothing and copies nothing, and the separateToSortedArraycopy step is gone. (2)LatticeCursorGrain'sBuildOwnedSlotsByShard, rebuilt once per cursor activation over the whole 4096-wide slot array, was an open-coded duplicate of theLatticeGrain.BuildOwnedSlotMappartition that had already been made dense, so it kept paying the hashed-lists-then-copy cost the shared helper had shed; it now delegates to that helper, which both inherits the dense form and removes the duplicate that could drift. (3)BulkAppendChunkAsyncbucketed each chunk into an unsizedDictionary<int, List<KeyValuePair<string, byte[]>>>with a probe-then-store pair per entry and a growth chain per bucket; it now buckets through the sharedShardFanout.BucketEntrieshelper under the same shard-fairmin(max(4, n/shards), 256)capacity bound the other batch paths already use. All three keep the existing hashed fallback for a hand-built map carrying a negative or pathologically large physical shard index, so no input regresses, and the ascending-slot order the batch readers require is preserved - differential tests assert each new shape against the shape it replaced, on the same input, rather than against hand-written expectations. A new[MemoryDiagnoser]benchmark suite (slotgroupfolds) isolates each shape against its prior form, calling the real production code on the optimized side of every pair: grouping 1024 requested slots over 8 shards falls from 18.52 KB to 4.77 KB (and from an 11.86 us to a 9.84 us median), a 4096-slot cursor partition falls from 51.21 KB to 16.76 KB (and from a 17.14 us to an 8.81 us median), and bucketing a 1000-entry append chunk falls from 45.91 KB to 27.94 KB (and from a 46.32 us to a 23.96 us median). (Orleans.Lattice9.6.0)Three batch-path reductions: shard fan-out bucketing, the snapshot baseline union, and WAL batch grouping. All three are output-identical and touch no public API. (1) Every multi-key read and write -
GetManyAsync,SetManyAsync,SetManyWhereAsyncand the CRDT delta batch - groups its batch by the physical shard that owns each key before fanning out. That grouping key lives in a tiny dense domain (one entry per shard root, typically 1-16) even though the virtual slot space it is projected from is 4096 wide, so hashing it once per batch entry was pure overhead; all four sites now bucket into an owner-indexed array and hash nothing. The read path additionally presized every bucket to the whole batch size, so a 1000-key read over 8 shards reserved 8000 slots to store 1000 keys - it now uses the same shard-fairmin(max(4, n/shards), 256)bound the write path already used.ShardMap.GetPhysicalShardIndicesreturns distinct ascending indices and every valueShardMap.Resolvecan return is drawn from that set, so the last element bounds the dense array exactly; a hand-built map carrying a negative or pathologically large index falls back to the prior hashed form, mirroring the guardGetPhysicalShardIndicesandLatticeGrain.BuildOwnedSlotMapalready keep, so no input regresses. (2)ShardRootGrain.CaptureSnapshotBaselineAsync's cross-leaf union, rebuilt on every snapshot capture, kept aSortedDictionaryof values - a red-black node per key, and two ordinal-string tree walks per row, one to probe and one to store - plus a parallelDictionaryof per-key merge modes written per row, plus a third keyed read per key while materialising. Value and mode now fold into one flat map written in a single probe viaCollectionsMarshal.GetValueRefOrAddDefault, with the ordinal order imposed by one final key sort; both orders are ascending underStringComparer.Ordinalover a distinct key set, so the materialised baseline is unchanged, including the donor-orphan collision case where the per-key mode must follow whichever value the LWW merge kept (a winning row with a null mode still clears the stored one). Because leaves own disjoint key ranges, the union is sized from the first leaf's row count, so the flat map lands its backing store in one shot rather than trading the red-black tree's allocate-once nodes for a grow-and-rehash chain over a wide value tuple. (3)WalCommitLogWriter.AppendManyAsyncbuilt the same$"{treeId}/{partition}"grain key three times over three parallel dictionaries, costing every entry a string interpolation, a hash probe for its entry list and a second full hash lookup to append its reverse index. It now keeps one batch object per partition and memoizes the last one, so the dominant single-partition batch formats and hashes its grain key exactly once for the whole batch instead of once per entry, and both per-partition lists are presized to their shard-fair share. A new[MemoryDiagnoser]benchmark suite (batchfolds) isolates each shape against its prior form, calling the real production code on the optimized side of the fan-out and snapshot-union pairs: bucketing a 1000-key batch over 8 shards falls from 63.91 KB to 14.27 KB (and from a 46.5 us to a 31.2 us mean), the entry-bucketing variant falls from 28.82 KB to 27.94 KB, a 16-leaf 6400-row snapshot union falls from 1,305.85 KB to 1,073.33 KB (and from a 3,835 us to a 1,479 us median), and a 512-record WAL batch falls from 67.75 KB to 36.63 KB (and from a 61.8 us to a 30.6 us median). (Orleans.Lattice9.6.0)The snapshot baseline capture now sits under a hard stall ceiling and folds its tail in a bounded fan-out.
CaptureSnapshotBaselineAsyncheld a deliberately non-reentrant shard root for the whole of a two-pass walk over its leaf chain, with no ceiling on how long that hold could last. The walk cannot be made resumable the way the issue originally proposed:FreezeProjectionAsyncalways freezes at the leaf's own frontier, andFoldTailOntoFrozenAsyncfolds only forward over(frontier, capturedHead]and skips any partition whose frontier already exceeds it, so a leaf whose head advanced pastcapturedHeadbefore being frozen would bake post-capture writes into its frozen cache and have the forward fold skip that partition - a silent zero-observable-writes violation with no rewind, since CRDT and LWW folds are not invertible. The two mitigations that do not require a resumable walk are applied instead. The capture is armed with the hard stall ceiling, which is safe to abandon because point-in-time consistency comes fromcapturedHeaddominating every leaf frontier rather than from the exclusive hold, and the capture is read-only right up to its closing seed, so a stalled capture merely fails the open and is retried with a fresh baseline token. The fold pass, which dominates the cost and, unlike the chain walk, is order-independent because every fold targets an already-captured head, now fans out under a ring-buffer window that bounds completed-but-unconsumed results as well as in-flight ones - a plain semaphore would let folds run far ahead of a slow chain-head leaf and pile their row sets in memory alongside the union. Results are still consumed in strict leaf-chain order, so the union, including the tie-sensitive merge-mode adoption rule on a donor-orphan collision, is byte-identical to a serial fold, and abandoned folds have their faults observed so they cannot resurface as unobserved task exceptions. The fan-out width isLatticeOptions.MaxConcurrentSnapshotBaselineFolds(default 4), resolved through a synchronous resolver seam so the fan-out costs no registry round trip inside the window the ceiling bounds, and the pass reports under its ownScanPagePhase.BaselineFoldbecause the serial walk's "the read in flight was leaf N + 1" no longer describes it. (Orleans.Lattice9.6.0)