Release 2026-08-29: Fixed
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at 2026-08-29-2.md, and llms.txt lists every page.Part of Release 2026-08-29, in Changelog.
Fixed
A tree left with a routable-but-unbound leaf now heals itself on its next typed CRDT write, with no operator action, no recover call and no restart. The
9.4.1repair for this issue ran only insideRecoverTreeAsync, which throws on a tree that is not currently soft-deleted, so the only remediation available to a deployment already stuck in this state was to callDeleteTreeAsync()on a live production tree and complete aRecoverTreeAsync()inside the 72-hourSoftDeleteDurationwindow - a remedy that risks destroying the data it is meant to rescue, on precisely the tree whose purge timeout caused the damage in the first place. The binding is now re-asserted at the point of use instead: when a typed CRDT write faults because its target leaf has no bound tree id, the owning shard root - which always knows the tree id and shard index the leaf is missing, and is the only party that does, since a cleared leaf retains nothing to resolve its own tree from - re-asserts the binding and retries the write once. Driving the repair from the fault rather than from a lifecycle hook is what lets it reach a deployment whose delete and recover already happened hours ago; the healthy path pays nothing, because an exception filter is only evaluated once a fault is already in flight. The repair is deliberately narrow so it cannot mask a genuine fault: it matches only onLatticeCrdtShapeNotRegisteredExceptioncarrying an emptyTreeId, which is raised on exactly one branch (RequireBoundTreeId), so a real "no shape registered for this tree" fault - which carries the tree id it could not resolve - propagates untouched and is never retried; and the retry is single-shot, so a leaf whose binding cannot be restored still fails closed exactly as before rather than silently accepting an unbound write. The repair logs a warning, because a leaf losing its binding remains an anomaly worth surfacing even though it is now self-correcting. Also fixed alongside it: a splitting leaf seeds its new sibling with its own tree id verbatim, so an unbound leaf minted another unbound leaf on every split and the damage spread across the key range as the tree grew - plainSetAsyncwrites resolve noCrdtShape, so an unbound leaf keeps accepting them, keeps filling and keeps splitting while only its typed CRDT writes fail, which is why the reported incident's stuck node was a split-created leaf deep in the tree rather than the shard's deterministic root leaf. The donor now logs a warning naming itself and the sibling it is about to mint; it deliberately does not throw, since that would break splits on exactly the trees that need to keep serving writes while they heal, and the write-path repair re-binds both nodes on the next typed CRDT write routed to them. Finally, recovery's own re-assert no longer infers a whole shard's health from its leftmost leaf:9.4.1probed that leaf and returned immediately when it was still bound, on the reasoning thatPurgeAsyncclears the leftmost leaf first so a bound leftmost leaf proves nothing in the shard was cleared, but that inference does not hold once a split can inherit an unbound donor's binding, and the same short-circuit also stranded a shard whose previous repair had failed part-way through its fan-out, since the retry then probed an already-re-bound leftmost leaf and returned.ReseedNodeBindingsAsyncnow walks the topology unconditionally, which is affordable because both setters are idempotent and short-circuit inside the callee, so an already-bound node costs one round trip and no storage write and recovery is a rare operator action; the walk stays bounded at 4,096 nodes at a fan-out of 16, and nodes beyond that budget are now re-bound by the write path on their next typed CRDT write rather than being reported as permanently unwritable. Five regression tests pin the new contract, including self-heal of an unbound leaf with no recover call at all, self-heal of a non-leftmost split-created leaf, self-heal after a full delete / interrupted purge / recover cycle, and negative controls that a genuinely unregistered shape still throws un-retried and that an unrepairable binding still fails closed. (#1744) (Orleans.Lattice)A cluster storage-usage roll-up now reports an honest lower bound instead of a silently understated total, and no longer collapses into response timeouts on a many-tree cluster. Three defects compounded in the same path. First,
LatticeStorageUsageGrainhandled a failed shard-root usage read by contributing a zeroedShardStorageUsageinto the tree's totals and leavingTreeStorageUsageReport.Partialunset, so one shard that failed or timed out understatedLeafStateBytes,SnapshotBytes,LiveKeys, andTotalByteswhile still presenting the report as complete - a wrong answer dressed as an authoritative one. The sibling WAL path had always handled the identical case correctly, returning a sentinel the caller turned intoPartial; the shard path simply never adopted the pattern. A surface that does not answer now contributes nothing and setsPartial, so the figures are a flagged lower bound and the existing "one bad shard does not abort the tree" resilience is unchanged. Second, the roll-up was a two-level uncapped fan-out - one unboundedTask.WhenAllper tree inLatticeAdminGrain, each fanning out again over every shard root and WAL partition - and the levels multiplied: a 90-tree cluster at the default 64 shards and 8 WAL partitions dispatched roughly 6,500 concurrent grain calls in one burst, all racing a single 30 s Orleans response deadline, which was observed failing wholesale on a local harness and stopping the moment the tree count was lowered. Both levels are now bounded by new options -MaxConcurrentStorageUsageTrees(default 8) andMaxConcurrentStorageUsageSurfaces(default 16, spanning shard roots and WAL partitions jointly) - capping the peak at 128 in flight regardless of cluster size, so the roll-up degrades in latency rather than collapsing. Bounding the burst is necessary but not sufficient, because it caps concurrency and not total work: a deep refresh re-walks every shard of every tree, so a large enough catalogue still cannot be sampled inside one response deadline however gently it is dispatched, and the whole call then fails and tells the caller nothing. A newStorageUsageRollupBudget(default 20 seconds, cluster-wide, non-positive to disable) caps the total: once it lapses the roll-up stops dispatching, the trees sampled so far keep their real figures, every remaining tree short-circuits to not-answered without dialing its grain, and the report comes back flaggedPartial- the same "an honest flagged lower bound beats an absent answer" rule the per-surface reporting follows. A caller-driven cancellation is deliberately still distinguished from a budget lapse and continues to abort the roll-up outright rather than being laundered into a confident-looking partial. The registry-sorted tree ordering thatClusterStorageUsageReport.Treesguarantees is preserved under the bound, and the aggregated figures are identical to the unbounded ones. Third, a cancellation mid-roll-up could be swallowed into a confident-looking zero by the same catch-all that hid shard failures, and the two sequential fan-out joins could abandon the second batch's tasks unobserved if the first threw; cancellation now propagates promptly, every dispatched slot is settled before the fan-out returns, and the backgroundLatticeStorageUsagePollerclamps an out-of-range poll interval and wraps its loop so neither a bad interval nor an options-validation failure can fault the hosted service and take the silo down with it. The byte-pressure over-threshold gauge is deliberately still gated on WAL completeness alone rather than onPartial, so an unrelated shard failure does not silence it. (#1728) (Orleans.Lattice9.4.0)An
Rgamerge no longer drops the incoming per-replica counter maxima when the other side carries no nodes.Rgacarries two independent pieces of merge state - the node union, and aContextmap of the highest counter seen per replica that letsNextCountermint the next dot in O(1) - butMergeFromshort-circuited onother.Nodes.Count == 0before foldingContext, so an incoming sequence with a populated context and no nodes had its maxima silently discarded. The early return dates to the original RGA primitive and was simply never revisited whenContextwas introduced several releases later, at which point the guard stopped covering everything the merge is responsible for. The consequence is the same-dot class this release has been closing: an under-counted context lets the next localInsertAfterre-mint an already-authored dot, so two distinct nodes share a dot identity and the sequence stops converging. It is not reachable by local authoring - every path that bumpsContextalso adds a node - but it is reachable from any decoded wire or storage payload, sinceNodesandContextare both public settable serialized properties, and it becomes routine the moment tombstone compaction is added, which is the standard RGA maintenance operation and precisely whatContextexists to survive. The fold now happens ahead of the short-circuit, and deliberately still belowEnsureContextRebuilt, whose own early return on a non-empty context would otherwise suppress the rebuild of the receiver's own maxima. (#1724) (Orleans.Lattice9.4.0)Rga.Clonenow copies each node's value bytes instead of sharing them, closing one leg of the CRDT buffer-ownership contract.ICrdt<TSelf>.Clonepromises a deep, independent copy, and composites lean on it:OrMap.Gethands backClone()precisely so a caller receives something safe to mutate.Rgaduplicated every node but aliased each node'sbyte[], so a caller reading a sequence out of anOrMap<string, Rga>got a live handle on the map's durable state and could write through a returnedRgaNode.Valuewithout passing any mutation API - the same defect already fixed one level up inOrMap.CloneandOrMap.Get, left behind in this sibling. The rule the family follows is now written down onICrdt<TSelf>and in the primitives instructions, and it is decided by provenance rather than by type: ingress from the caller (Set/Add/InsertAfter) is a documented hand-off and copies nothing, a fold from a peer or a delta copies the winning candidate, and egress to a caller copies. The copy uses a span copy rather thanArray.Clone(identical allocation, but roughly 3-4x faster on theordedupmicrobench), and an empty span'sToArray()returns the sharedArray.Empty<byte>()singleton, so a tombstoned node - which is what an aged RGA is mostly made of - adds no allocation at all. Documented under Who owns the bytes. (#1724) (Orleans.Lattice9.4.0)The CRDT buffer-ownership contract is now actually honoured across the whole primitive family, and enforced structurally rather than by review. That contract, written down on
ICrdt<TSelf>and in the primitives instructions, decides ownership by provenance: ingress from a caller is a hand-off, a fold from a peer or a delta copies the winning candidate, and egress to a caller copies. Four legs were still open, in two primitives.Rga.MergeFromandRga.MergeDeltaadopted a peer's or a delta'sRgaNode.Valueby reference - both when creating a node the receiver had not seen and when an incoming value won a same-dot collision - so a fold left the receiver's durable state aliased to an array somebody else still owned.Rga.ToList, the sequence's materialised projection and its primary read API, handed out the live node buffers; because the resolved order is cached and nothing invalidates it, one write through a returned value corrupted every later read - the exact hazard that method's own documentation described, which it had closed for the list shape (viaAsReadOnly) while leaving the element buffers open, so no downcast was even needed.MvRegister- one of the four primitivesICrdt<TSelf>names in its own remarks - violated all three legs at once: itsMergeFromandMergeDeltaadopted the peer's and the delta's entry buffers on the same two paths, itsClonewas a shallow entry copy, andValuesreturned the stored arrays and cached them, handing two readers the same buffer. TheClonecase had a comment justifying itself on the grounds that the value bytes were "treated as immutable by every production call site" - the same reasoning already rejected forRga.Clone, and the reason this class of defect kept recurring: each previous fix was pinned only by a hand-written test for the one type that had just been fixed, leaving the next sibling free to repeat it. The observable consequence surfaced throughOrMap<TKey, TValue>.Get, which hands back aClone()of the sole contributor precisely so a read is safe to mutate, and folds later contributors in viaMergeFrom: over anMvRegistervalue both paths leaked, including the single-contributor fast path that is the dominant steady-state read, so a caller that mutated what it read wrote into the map.LatticeReplicationConfigEntryinherited the same leak through its mode register without a line of its own being at fault. Every leg now copies with a span copy (AsSpan().ToArray(), matchingBoundedRegister), and the cost is kept where it belongs: only a winning fold candidate is copied so an idempotent steady-state merge still allocates nothing, an empty or tombstoned payload reuses the sharedArray.Empty<byte>()singleton, the cached projections now cache the ordering (the traversal and sort, which is the expensive part) and pay only the per-value copy, and the internal call sites that immediately consume the bytes - the typedRgaAccessorandMvRegisterAccessor, which deserialise each value on the next line, and the index-resolution paths that read only the dots - are routed at newinternalaliasing views (Rga.MaterializeShared,MvRegister.ValuesShared) so they pay nothing at all. No public signature changed. Reachability is deliberately not overstated: every production delta path takes serializedbyte[](ApplyCrdtDeltaAsync) and every accessor decodes per read, so these were latent public-API contract violations rather than in-flight replication corruption. The enforcement is the substantive half of the change:CrdtBufferOwnershipContractTestsBase(shared testing library, bound inOrleans.LatticeandOrleans.Lattice.Replication) walks each registered CRDT's real object graph and comparesbyte[]instances by reference identity across clone, state fold, delta fold, and every public projection, and additionally fails when a CRDT type declared in the package has no specimen, when a type grows a publicbyte[]-bearing projection nothing covers, and when a specimen's declared payload shape is wrong - which also pins the set primitives (GSet,OrSet,RwSet) as payload-free by construction, since they encode elements to base64 string keys and retain no caller array. A future primitive therefore cannot join the family without picking the contract up. (#1724) (Orleans.Lattice9.4.0,Orleans.Lattice.Replication9.4.0)A bounded register decoded from storage now takes its fold direction from the registered merge mode, not from the stored payload. A
MaxRegisterand aMinRegisterare one primitive pointed in opposite directions, and that direction lived in two places that were never reconciled: theLatticeMergeModeregistered for the key (which is what the tree dispatches on, and what the documentation already described as authoritative) and anIsMinbit carried on the state. Neither decode seam stamped the bit from the mode - theCrdtShapeRegistryshape decode andBoundedRegisterAccessorHelper.ReadAsyncboth took whatever the persisted bytes said - so a payload that reached the store without passing through a directional accessor (a rawSetAsyncof hand-authored JSON, a client-supplied state, an older or foreign writer) fixed the wrong direction for the rest of that key's life. Nothing detected it:MergeFromresolves under the receiver's direction and never inspects the other side's, so aMinkey whose payload claimedMaxsimply kept answering with the greatest value ever written, silently and forever, including through the publicGetRegisterAsyncprojection. Both seams now re-stamp on decode, so a disagreeing payload self-heals on read. The stamp is an in-place write on the just-decoded instance, so it allocates nothing, andIsMinkeeps its[Id(3)]slot and its setter - there is no wire-format or public API change, and a correctly-authored payload decodes exactly as before. A direction mismatch is deliberately not raised as an error: throwing on a CRDT fold path can wedge replication on one bad payload and would still leave the stored state wrong. (#1724) (Orleans.Lattice9.4.0)The ambient vector-clock frontier is now copied at the seams where it becomes durable state or escapes to a caller, so a co-located sender's
VersionVectorcan no longer be rewritten under committed entries.VersionVectoris a mutable CRDT, andLatticeVectorClockContext.Currentwas a plain cast out ofRequestContext- one shared instance, assigned straight into the persistedLwwValue<T>.VectorClockof every entry written inside the scope, at write sites acrossBPlusLeafGrain,AtomicWriteGrain, andShardRootGrainand onward through the WAL record, snapshot, and replication-apply paths. On the inbound replication path that instance arrives inside an[Immutable]carrier (ApplyCrdtDeltaItem.SourceVectorClock), and Orleans elides the deep copy for an[Immutable]payload when the callee is co-located - so the object promoted into the durable frontier of many entries at once was the sender's own, still reachable and still mutable on both sides. The same aliasing ran outward throughLwwEntry, handing a caller a live handle on stored state. Because the elision only happens under co-location, no cross-silo test would ever have shown it. The context setter now copies on the way in andLwwEntrycopies on the way out, fixing the whole class at the two narrowest seams rather than at the dozens of sites that read the frontier back; the scope-restore path skips the second copy because its captured value was already copied on entry. Both are?.Clone(), and a purely local write leaves the frontiernull, so the dominant path allocates nothing and the cost elsewhere is one clone per scope rather than per write. Found by the new grain-boundary contract guard below. (#1725) (Orleans.Lattice9.4.0)The
[Immutable]same-silo copy elision is now guarded by a contract test instead of a one-off audit. Orleans skips the deep copy for an[Immutable]payload on a co-located call and hands the receiver the sender's own instance, which is sound only while nobody mutates it - and roughly 140 tracked types carry[Immutable]alongside a mutablebyte[]or collection. Re-reading them all was rejected as the wrong shape of work: the risk is structurally bounded, because CRDT payloads cross grain boundaries as opaquebyte[]and are decoded into a fresh object graph before being folded, so no shared buffer is ever folded in place. That is a reachability argument about today's signatures rather than an invariant, so it is now enforced by the reflection-drivenImmutableGrainBoundaryContractTestsBase(bound in the core test project, alongside the established serializable-exception and grain-key guards): one test pins that no CRDT state type is shared across a boundary by an elided copy - counting a CRDT reached through an[Immutable]carrier, which is what surfaced theVersionVectoraliasing above - and a second requires every[Immutable]payload on the boundary that carries a mutable buffer to be listed with a written justification, with a companion test failing any listed entry that has gone stale. A future grain method that puts a typed CRDT, or any other in-place-folded payload, on the boundary now failsbuild-and-testrather than reaching production. (#1724) (Orleans.Lattice9.4.0)Fourteen tools documented "Read-only" no longer create the tree they were asked to read, so a read-only identity can no longer provision durable storage. Every CRDT read verb on the data facade (
GCounterGetAsync,PnCounterGetAsync,GSetGetAsync,OrSetGetAsync,RwSetGetAsync,OrFlagGetAsync,RwFlagGetAsync,MvRegisterGetAsync,MaxRegisterGetAsync,MinRegisterGetAsync,OrMapGetAsync,SequenceGetAsync,VersionVectorGetAsync) andLatticeTreeAdmin.GetTreeStatsAsyncrouted straight into a shard-root grain, whose activation resolves per-tree options throughLatticeOptionsResolver- and that resolver seeds a missing registry entry on first use. Reading an id that had never existed therefore registered it as anActivetree with the full default shard fan-out, whichTreeExistsAsyncand the tree listing then confirmed. The same seeding ran onDeleteAsyncandDeleteRangeAsync, whose contracts describe an unknown tree as a routine no-op. The authorization consequence is the sharp edge: a caller holding onlyReadandRangeReadis not even offered the tree-creation verb, yet could mint an unbounded number of durable trees through the read path, so a privilege refused at the front door was reachable from behind it. The thirteen CRDT reads and the statistics snapshot now probe the tree registry first and short-circuit exactly asGetAsyncandReadRangeAsyncalready did, returning the documented empty answer (0, an empty sequence,false,null, an empty map, a zeroed statistics snapshot) without provisioning anything. The two delete verbs answer the same way but from inside the core tree grain, sequenced after the delete gate they already applied, so a caller who may not delete still receives an authorization denial rather than a no-op-shaped success. The probe costs one extra grain call only on the miss path, needs no privilege the operation itself did not already require, and every documented empty-read and no-op-delete result is unchanged. (Orleans.Lattice9.4.0,Orleans.Lattice.Api.Data9.4.0,Orleans.Lattice.Api.TreeAdmin9.4.0)MetadataOnlyandHybriddurable-history retention are now actually enforced, so an operator who asks for hash-and-length only no longer keeps full plaintext.MetadataOnlyis the default retention mode and is documented as storing a revision's content hash and byte length but not its bytes, withHybridkeeping bytes only inside a recent window. The view-backed history path honoured the policy, but the write-ahead-log fallback that answers a history query for revisions not yet folded into a materialised view mapped every mutation with the value preview unconditionally attached and stamped the rowFullValueregardless of the tree's configured mode - so all three modes behaved identically toFullValue, andGetEntryHistoryAsyncreturned readable plaintext for a tree whose policy said it was not retained. The fallback now resolves the tree's effective policy once per page and applies the same retention shaping the view path uses: the hash and byte length are always reported, the bytes are attached only when the resolved mode keeps them, and the row is stamped with the mode that actually applied, so thevalueRetainedflag reported to callers is truthful. A CRDT delta is unaffected, since a delta carries no last-writer-wins value bytes. (Orleans.Lattice9.4.0,Orleans.Lattice.Api.State9.4.0)Setting a history-retention policy on a reserved system tree is now refused, as its contract already stated.
SetHistoryRetentionAsyncdocuments rejection both for a non-positive window and for a reserved system tree id, but only the window half was enforced: a caller could set an age-bounded retention window onsys-auth-policyorsys-membership-groupsand have it persist, which would age out the authorization and membership history the cluster's own audit trail depends on. The verb now rejects a reserved system data tree alongside the identifiers it already refused. The guard is scoped to this one verb, so the first-party add-ons that legitimately configure retention on their own system trees from inside the silo are unaffected. (Orleans.Lattice.Api.TreeAdmin9.4.0)Dropping a materialised view is now idempotent, and a library-owned view is now protected from a runtime drop.
DropViewAsyncdocuments both properties and held neither. A second drop of the same view, or a drop of a name that never existed, failed with a not-found error instead of succeeding quietly; and the four system history views that the auth, membership, and tenancy packages declare from their own initializers dropped successfully, only for the next silo start to re-declare them - which is precisely the reason the contract gives for refusing the drop. The startup-declaration guard only ever inspected views registered throughAddLatticeViews, so it never saw them. An absent view is now the documented no-op, and a view whose source tree lives in a reserved namespace is refused with a clear error, so the two cases stay distinct rather than one masking the other. (Orleans.Lattice9.4.0,Orleans.Lattice.Api.TreeAdmin9.4.0)Reading the tag catalogue for an index that does not exist now reports not-found instead of silently materialising the index.
ListIndexTagsAsyncandScanTagMembersAsyncresolved an index handle by name and enumerated it, which both answered an empty catalogue and created the backingtag-{indexName}membership tree. That destroyed a distinctionScanEntriesAsyncexplicitly documents and that a caller needs in order to tell a typo from a real result: a tag-filtered scan reportsIndexNotFoundfor an index that was never materialised andFoundwith zero entries for a real-but-empty one, but a single catalogue read of the mistyped name flipped it permanently toFound. Both verbs now resolve without materialising and report the same not-found the tree-administration tag verbs already do, leaving a real-but-empty index answering an empty page as before. (Orleans.Lattice.Api.State9.4.0)MCP tool discovery no longer reports a backend timeout as a valid, tiny permission set. Resolving a caller's facade-group access catches a failure and fails closed, which is correct for a denial - but the same catch swallowed a cancelled, deadline-exceeded, or unavailable transport fault and returned an empty grant set, so
tools/listanswered successfully advertising a single meta-tool. A platform administrator entitled to the full tool surface was observed advertised one tool with no error anywhere on the wire, and a client had no way to tell "your permissions were revoked" from "the silo did not answer in time". A fault that means no authoritative answer arrived - a transport status ofCancelled,DeadlineExceeded,Unavailable,Internal,ResourceExhaustedorAborted, an Orleans response deadline, or silo churn - is now surfaced to the client as a retryable discovery error rather than an answer, while a status that is an answer (PermissionDenied,Unauthenticated,NotFound, and every argument-shaped code) still fails closed exactly as before. The fail-closed guarantee is strictly preserved: raising the error advertises nothing at all, so the new behaviour is never wider than the behaviour it replaces. (Orleans.Lattice.Api.Mcp9.4.0)Reshard and resize initiation no longer answer "is this tree empty?" with a strongly-consistent whole-tree count, which could time the operation out under concurrent write load.
TreeReshardGrain.ReshardAsyncandTreeResizeGrain.ResizeAsynceach decide whether to take an empty-tree fast path by askingILattice.CountAsyncfor a live-key count, but that count walks every leaf chain of every shard, discards its result and retries whenever the shard map moves under it, and gives up only onceMaxScanRetriesis exhausted - after whichCountAsync's own stale-alias envelope runs the whole loop a second time. Initiation is precisely when the map is most likely to be churning: a caller may be writing concurrently, and a small leaf fan-out splits continuously. The probe could therefore consume the caller's entire response budget and surface as aTimeoutExceptionbefore the operation had started - observed in CI asReshard_under_concurrent_write_load_does_not_wedge_with_short_forward_deadlinetiming out at 30 s with the initiation turn still awaiting the count and the coordinator never reached. Both now probe existence directly through the new internalIShardRootGrain.AnyAsync, which short-circuits at the first non-empty leaf, OR-ed across the tree's physical shards. That needs no reconciliation against a moving shard map, which is the crux of the difference: a count must reconcile because a key migrating between shards is briefly visible on both the source and the destination and would be double-counted, whereas a split only ever moves keys - never creating one, destroying one, or leaving one present on neither side - so a key that exists is observed by at least one shard wherever the split has got to, and observing it twice still just means "a key exists". The answer is deliberately one-sided: it may report non-empty while the last keys migrate away, but never empty while a key exists anywhere, and only "empty" unlocks a fast path, so the one consequential direction cannot be wrong. A newLatticeOptions.EmptyTreeProbeBudget(default 10 seconds,InfiniteTimeSpanto wait indefinitely) bounds a probe that parks rather than returns, and every inconclusive outcome is read as "not empty" so initiation proceeds down the normal coordinator path. The same count-to-answer-a-boolean shape is also removed from the replication config authority: the internalILatticeTreeContentProbenow exposesHasContentAsyncinstead ofCountAsync, implemented by taking the first key from the tree's key stream, so enabling replication on a large tree no longer pays a whole-tree count to learn whether a bootstrap snapshot is needed. Documented underEmptyTreeProbeBudget. (Orleans.Lattice9.4.0,Orleans.Lattice.Replication9.4.0)The leaf delivery epoch is now unique across processes, so a cache can no longer be silently denied its full resync.
BPlusLeafGrain's delivery epoch identifies one leaf activation and is what tells aLeafCacheGrainholding a stale cursor to fall back to a full-snapshot delivery, but it was minted from astatic longwith no initialiser - zero at the start of every silo process - so every silo independently handed out1, 2, 3, .... The epoch is compared across processes (a cache outlives the silo whose activation minted the cursor it holds), so two activations in different silos could mint the same epoch, and a collision is not benign: the epoch-mismatch branch is skipped, the caller's higher sequence then satisfies thealready at headcheck, and the leaf ships an empty delta. The cache adopts the new cursor and resumes incremental delivery, so fresh writes still flow and the fault looks healthy - but the one-time resync never happens, every key that diverged while the cache was disconnected stays stale indefinitely, and becauseLeafCacheGraingates its own cache eviction on the same epoch flip, keys the leaf has since deleted stay visible in the cache's read view. Collisions cluster precisely around silo restarts, when a surviving cache holding a low-numbered epoch is most likely, and the failure is silent - no exception, no log, no metric. The seed is now randomised per process into[1, long.MaxValue / 2](0stays reserved forLeafDeliveryCursor.Empty, and per-activation monotonicity within a process is unchanged), and as defence in depth the leaf now also treats a cursor whose sequence is ahead of the activation's as stale - impossible to have been issued by that activation - and answers it with a snapshot, withLeafCacheGrainapplying the matching condition so the snapshot still triggers its eviction pass. No wire-format change:LeafDeliveryCursorkeeps its[Id(0)] long Epoch; only the value's provenance changes. (#1711) (Orleans.Lattice9.4.0)The WAL materialiser pin grain key is now storage-safe and unambiguous, healing itself across the change.
WalMaterialiserPinGrainis persistent and its key was composed as{treeName}#s{shard}. Azure Table grain storage carries a grain key into the Partition/Row key columns and the request URL, both of which reject#, so the key depended on the backend sanitising it - and such sanitisation is lossy, which can collapse two distinct grain identities onto one persisted row. A materialiser pin holds WAL retention state, so a confusion there is durability-relevant rather than cosmetic. The shard suffix now uses the storage-safe~s, and the composer is marked[GrainKeyBuilder]so the existing reflection-driven storage-safety guard audits it from now on instead of relying on review. The change strands no WAL and needs no operator action: the GC's read fan-in already performed a dual read to pick up pre-sharding keys, and now also reads every shard under the legacy separator, so a pin written by an earlier build keeps holding the trim floor while new pins are written under the safe key as their consumers re-pin. Suffix parsing is additionally anchored at the last separator and accepts only an all-digit suffix, fixing an ambiguity in the previous leading-IndexOfparse that truncated a tree whose own name contained the separator. (#1701) (Orleans.Lattice9.4.0)Ten convergence and contract defects across the CRDT primitives, found by a consolidated audit of
src/lattice/Primitives/after #1705. That change fixedOrMap.Clone's nested-value aliasing but stopped there, and the same defect class was still live on four furtherOrMapseams and on two sibling primitives it never touched.OrMap.MergeFrom,OrMap.MergeDelta,OrMap.SetandOrMap.Getall retained a caller's, peer's, or delta's value-CRDT instance by reference, so a map's durable state stayed aliased to an object the other side still owned:a.MergeFrom(b)leftaandbsharing nested CRDTs (violating purity of merge in its argument, and letting a later mutation ofasilently rewriteb),MergeDeltaadopted the producer's delta objects (so applying a delta that is legitimately retried or fanned out to several peers corrupted it),Setstored the caller's instance, and - worst, because it is a read -Getfolded the first contributing entry into a fresh accumulator by mutating it, which for a single-entry key meant everyGetmutated the very state it was reporting; aBoundedRegistervalue read twice returned different answers.BoundedRegister.Cloneshared its value and order-key byte arrays, so a register handed out ofOrMap.Getgave the caller a live handle on the map's persisted bytes; the deep copyICrdt.Clonepromises is now made (measured cost below), and the merge/delta-apply fold likewise copies a winning candidate rather than adopting the other side's buffer, whileSetkeeps its documented hand-off.OrMap.Keys()sorted with the ambient culture, so two replicas under differentCurrentCulturesettings enumerated the same converged map in different orders - underen-USthe keysaandBorder asa, B, under the ordinal comparison every other enumeration surface in the library uses they order asB, a; keys are now ordinal-sorted.CrdtShapeRegistrymapped a null OR-Set delta element to the empty string when building its dedup key, so a null-element dot and a legitimately empty-element dot with the same(replicaId, counter)collided and one was silently dropped during delta combination; null elements are now skipped, matching the sibling G-Set path.LwwValue<T>.Mergewas not commutative: its tie-break ran out after(Timestamp, OriginClusterId, IsTombstone)and then returned the left operand unconditionally, so two writes agreeing on all three but differing inExpiresAtTicks,Value, orIsMigratedresolved to whichever arrived first and two replicas never converged - reachable becauseOriginClusterIdis null for every purely local write andHybridLogicalClockcarries no node id, making a same-wall-tick collision between two leaves ordinary. The order now extends through those three fields (Valuelexicographically forbyte[], the only instantiation the library persists) and the XML doc states the residual honestly instead of claiming a commutativity it cannot deliver for an arbitrary unorderedT.HybridLogicalClockoverflowed its logical counter:Tickand all threeMergebranches incrementedCounterunguarded, so a clock pinned atint.MaxValueby a peer's malformed or malicious stamp wrapped toint.MinValueand went backwards, breaking the monotonicity every last-writer-wins decision in the library rests on; the counter now saturates.Rga.ToListcached and handed back the mutable backing list, so a caller who downcast the returnedIReadOnlyList<>could rewrite - or empty - the RGA's cached materialisation for every subsequent reader; the cache now holds a read-only view. AndMvRegisterEntry,OrMapDelta<,>andOrMapDeltaEntry<,>were marked[Immutable]while carrying a mutablebyte[]or a mutable value-CRDT the receiver folds in place, which told Orleans to skip the deep copy on a same-silo grain call and hand the receiver the producer's own objects; the attribute is removed from the three, with the reason recorded on each type. Every one of the ten was proved by a regression test that fails against the prior code before the fix was written; two further reported findings were investigated and disproved (VersionVector's handling of an all-zero clock slot, which is a correct join whoseIsBottomis documented as a structural hint, andLwwValue.Merge's associativity, which the left bias in fact preserved - only commutativity was broken) and are unchanged. (Orleans.Lattice9.4.0)A materialised view's tree id is now usable and unambiguous as a grain key, and the change heals itself. A view tree id is an Orleans grain primary key and is carried into
ShardRootGrain's composite key ({treeId}/{shardIndex}) - a persistent grain - but neither property that.github/instructions/grains.instructions.mdrequires of such a key was held. The shadow-swap generation suffix used#, one of the characters (/,\,#,?, plus the control characters0x00-0x1F/0x7F-0x9F) that Azure Table grain storage rejects because it carries a grain key into the Partition/Row key columns and the request URL; the historical failure mode is an opaque HTTP 400 that no in-memory test storage reproduces, which is how #1529 reached production. And because a view name was validated only for null/empty, a view nameda/bproduced the persistent shard-root keyview-a/b/0, while a view namedorders#g2collided with generation 2 of a view namedorders- two grain identities, one state row. View names are now validated at every creation seam (the storage-unsafe set, plus the reserved generation separator), and the generation suffix uses the storage-safe~; the composer is marked[GrainKeyBuilder]so the existing reflection-driven storage-safety guard audits it automatically from now on. Rejecting/also closes a tenancy hazard for free: without it a caller could name a viewt/other/ordersand have the tree planted in another tenant's reserved namespace. The separator change strands no data and forces no rebuild. The maintainer records a legacy-generation ceiling, pinned once at activation to the generation then active: generations at or below it keep resolving through the old separator, and the next rebuild allocates a higher generation under the new one, so a view converges on its own with no operator action and the superseded generation is still reclaimed on the normal grace cadence. Pinning happens at activation rather than inside a particular verb precisely because a rebuild can be driven without the activation path having run first - a ceiling pinned after a generation was already allocated would misclassify it and send every read to a tree that was never written. A view created after the upgrade pins0, and generation 0 carries no suffix, so every generation it allocates is storage-safe. A name that an older build accepted is still restored, and reported in the log rather than refused, so adopting the rule cannot strand an existing view. (#1696) (Orleans.Lattice9.4.0,Orleans.Lattice.Api.TreeAdmin9.4.0)A reserved-namespace rejection on the data gRPC binding is now a typed client error instead of an opaque server fault. The core reserved-namespace guards threw a bare
InvalidOperationException, which fell throughLatticeDataApiGrpcService's typed catch arms to the generic handler and surfaced asStatusCode.Internalwith "The data-API request failed; see the cluster logs for the cause". Naming a tree in a reserved, internally-composed namespace is a deterministic caller-side precondition - the id is not addressable through the public surface for anyone - so reporting it as a server fault sent callers to the cluster logs for something they could see and fix themselves. The guards now throw a newLatticeReservedTreeNamespaceException(deriving fromInvalidOperationException, so existing broader catches are unaffected, and carrying the required no-op[RegisterCopier]for same-silo deep copies), which the binding maps toInvalidArgumentcarrying the self-contained message. Covered by new regressions on both the tenant and system-data namespaces. (Orleans.Lattice9.4.0,Orleans.Lattice.Api.Data.Grpc9.4.0)A single-scope
CaptureSetAsyncno longer hands back a set id that matches nothing. Set membership is not stored as a set:BackupSetManifestis returned to the caller but never persisted, and the only durable trace is theSetId/SetName/SetCreatedAtUtcstampStampSetMembershipAsyncpushes onto each member's own manifest. That stamping is deliberately skipped for a one-member set - it is indistinguishable from a plain backup and lists as one, and a one-tree set has no cross-tree atomicity to preserve - butBuildSetManifestran unconditionally and still minted a content-addressed SHA-256 id for it. The result was a 64-hex identifier indistinguishable from a valid one that appeared nowhere else:ListBackupsAsyncreportedSetId = nullon the member's catalog row, so a remote consumer grouping catalog rows by the id the create response had just returned found nothing, with no error to catch - a "backup sets" view reasonably concluded the set had been lost or pruned. The reach was not in-process only; the value crosses the backup control facade, its gRPC binding, and the MCP remote adapter, where grouping catalog rows is the only thing a remote consumer can do with a set id (there is no remoteRestoreSetoperation).BackupSetManifest.SetIdis nowstring?and a one-member set reportsnull, so the create response agrees with the catalog row by construction; an empty id is still rejected, because that is a malformed id rather than the absence of one. The deliberate< 2stamping rule is unchanged - the skipped stamp was correct and the phantom id was the defect - and both halves of the decision now derive from one sharedBackupSetIdentity(which also owns the content-address algorithm), with the stamping pass keying off the minted id rather than re-deriving the threshold, so the id and the stamp cannot drift apart. In-cluster,RestoreSetAsync's unresolved-set failure no longer reads as though the set once existed: it re-derives the single-member content address of each catalogued backup id and, on a match, names the exactbackupIdto pass toRestoreAsyncinstead, so anyone who persisted an id from an older build is rescued rather than told the set does not exist; a genuinely unknown id keeps a distinct message. The extra catalog walk runs only on that failure path, which throws either way. Covered by new capture regressions asserting the create response and the catalog rows agree for both arities (including that grouping rows by a reported id resolves exactly the set's members), a rescue regression pinning the message content andParamName, wire and MCP-adapter round trips of the absent id, and a newBackupSetIdentitysuite pinning the content address against an independent computation so previously-issued ids cannot be orphaned. Breaking for a consumer that readsSetManifest.SetIdas non-nullable. (#1687) (Orleans.Lattice.Backup9.4.0,Orleans.Lattice.Api.Backup9.4.0,Orleans.Lattice.Api.Backup.Grpc9.4.0,Orleans.Lattice.Api.Mcp9.4.0,Orleans.Lattice.Explorer9.4.0)The RGA sequence CRDT no longer mints a colliding dot after a legacy-payload merge.
Rga.MergeFromandRga.MergeDeltafolded the incoming side's per-replica maxima into the serializedContextcounter cache but never rebuilt this sequence's own maxima first, so a sequence loaded from a payload that predates theContextfield (nodes present, cache deserialized empty) that merged a peer before its first local insert was left with aContextreflecting only the incoming side. The nextInsertAfterthen re-minted an already-authored dot -NextCountersees a now-non-emptyContextand skips its lazy rebuild - producing two nodes sharing one(replicaId, counter)identity, which breaks convergence: replicas that applied the same operations in a different order no longer agreed on the sequence. Both merge entry points now callEnsureContextRebuilt()before folding, exactly as the siblingOrMap.MergeFrom/OrMap.MergeDeltaalready do; the pure-localInsertAfterpath was already guarded. No public API change. Covered by newRgaContextTestsregressions on both the full-state (MergeFrom) and delta (MergeDelta) merge paths. (Orleans.Lattice9.4.0)The RGA sequence CRDT's full-state merge now reattaches a tombstone-before-insert placeholder's parent, so the merge stays commutative. When a tombstone is delivered before its matching insert (out-of-order or partial delivery),
Rga.MergeDeltarecords a tombstoned placeholder node whoseParentDotis a stand-inRootso a later insert can reattach the real parent.MergeDelta's insert path performs that reattachment, butRga.MergeFrom- the full-state merge reached by the publicRga.Mergeand the replication applier's registered CRDT merger - folded only the incoming side's tombstone flag and value for a dot already present locally and never reattachedParentDot. A replica that recorded a placeholder via a delta and then learned the dot's authoritative structure through a full-state anti-entropy / catch-up merge kept the placeholder's live children mis-rooted underRoot, soMerge(a, b)andMerge(b, a)linearised those children in different positions and the sequence stopped converging.MergeFromnow foldsParentDotby the same deterministic max rule it already applies to the value (a real parent'sCounter >= 1dominates theRootplaceholder'sCounter 0), making the reattachment order-independent. No public API or wire-format change. Covered by newRgaMergeConvergenceTestsregressions. (Orleans.Lattice9.4.0)OrMap<TKey, TValue>.Clonenow deep-copies each key's nested value CRDT, so a map merge no longer mutates its left operand.ICrdt<T>.Cloneis contracted to return "a deep, independent copy ... mutating the returned value must never affect the receiver", and the staticOrMap.Merge(a, b)isa.Clone().MergeFrom(b). ButCloneduplicated each key's entry list while sharing theOrMapEntryobjects - and their nested value CRDTs - with the source by reference.MergeFromresolves a same-dot value collision in place viaexisting.Value.MergeFrom(other.Value), so when two replicas authored the same(replicaId, counter)dot under one key with divergent values,Merge(a, b)mutateda's nested value in place: the merge was not pure, and a caller still holdingasaw it change under an operation documented to leave both operands untouched.Clonenow copies each entry and clones its value (new OrMapEntry<TValue>(e.ReplicaId, e.Counter, e.Value.Clone())), mirroringRga.Clone's per-node deep copy, so a merge folds only into the throwaway clone. No public API or wire-format change. Covered by newOrMapCloneIsolationTestsregressions. (Orleans.Lattice9.4.0)The multi-value register's merge is now commutative when two replicas disagree on the value carried under one dot.
MvRegister.MergeFrom(and the equivalentMergeDelta) keeps a dot still present on both sides, but on such a collision it kept the local side's value unconditionally, soMerge(a, b)andMerge(b, a)disagreed whenever the same(replicaId, counter)dot carried different bytes on each side - a reachable divergence, because areplicaIdis a caller-supplied string with no minted-once guarantee, yet the type documents itself "commutative, associative, and idempotent". Both merge paths now break a same-dot value tie deterministically, keeping the lexicographically-greater value bytes (the same byte-order ruleRgaalready applies), so the two merge orders converge. The lazy survivors-list allocation is preserved, so an idempotent re-merge still copies nothing and touches no state. No public API or wire-format change. Covered by newMvRegisterMergeConvergenceTestsregressions on both merge paths. (Orleans.Lattice9.4.0)The RGA sequence CRDT's delta merge now folds a same-dot parent/value collision by the same deterministic max rule as its full-state merge, so delta-fed replicas converge.
Rga.MergeFromresolves a dot already present locally by a deterministic max -CompareDotfor the structural parent,CompareBytesfor the value - butRga.MergeDelta's insert path unconditionally overwroteexisting.ParentDotandexisting.Valuewith the arriving delta's (last-arrival-wins). Two replicas that applied the same colliding-dot inserts as deltas in different orders therefore disagreed with each other, and with a full-state-fed replica, even though the type documents itself "commutative, associative, and idempotent".MergeDeltanow applies the identicalCompareDot/CompareBytesmax rules, so an insert refresh is order-independent; a real parent (Counter >= 1) still dominates a tombstone-before-insert placeholder'sRootparent (Counter 0), so placeholder reattachment is unchanged, and a re-delivered identical insert stays a no-op. No public API or wire-format change. Covered by newRgaMergeDeltaConvergenceTestsregressions. (Orleans.Lattice9.4.0)The default replication framing decoder's per-entry bounds check no longer integer-overflows on a forged length.
IReplicationBatchEncoder.TryDecodeFraming- the default interface implementation inherited by any encoder that does not override it - validated each entry body with a narrowingcursor + lengthsum. Both operands areintandlengthis read straight off the wire (up toint.MaxValue), so the sum could overflow to a negative value and slip past the truncation guard, downgrading the fail-closedArgumentExceptionframing rejection into a rawArraySegmentout-of-bounds throw. The sum is now widened tolong, mirroring the sibling fix already applied toOrleansBinaryReplicationBatchEncoder. Covered by a new regression. (Orleans.Lattice.Replication9.4.0)The default replication framing decoder's routing-field bounds check no longer integer-overflows. The same narrowing
cursor + lengthdefect affectedReadLengthPrefixedUtf8, which decodes thetreeNameandoriginClusterIdrouting fields; a forged length nearint.MaxValueoverflowed past the guard into a rawSpan.Slice/Encoding.UTF8.GetStringthrow instead of the descriptive framing rejection. Widened tolongas above. Covered by a new regression. (Orleans.Lattice.Replication9.4.0)The auth-admin prefix-scope upper bound no longer wraps a trailing
U+FFFFcode unit.LatticeAuthAdmin.PrefixUpperBound- reached fromExplainAsyncviaTranslateScope, which turns aPrefix-scoped introspection request into the exclusiveRangeEndhanded to the access gate - incremented the prefix's final code unit unconditionally. A prefix ending inU+FFFFtherefore wrapped toU+0000, producing an upper bound that sorts below the prefix and inverts the half-open[prefix, bound)range, so the explained verdict could silently drop matching rules. It now scans back to the last code unit belowchar.MaxValue, increments that, drops the trailing max units, and returnsnull(unbounded above) when none exists - mirroring the canonicalBackupConstants.PrefixUpperBound. Covered by a new regression. (Orleans.Lattice.Api.Auth9.4.0)Three more prefix-scope upper-bound helpers no longer wrap a trailing
U+FFFFcode unit. The same non-canonical unconditional-incrementPrefixUpperBoundthat was fixed inLatticeAuthAdminsurvived in three sibling range scans over security-relevant data:LatticeAuthorizationPolicyStore.PrefixUpperBoundbounds the authorization-rule scan behindListRulesForTreeAsync,LatticeMembershipDirectory.PrefixUpperBoundbounds both the forward group-closure walk and the reverse members-of scan over membership edges, and the repo-context bootstrap indexer's local copy bounds its stored-file-meta reconcile scan. In each case a prefix ending inU+FFFFwrapped toU+0000, producing an upper bound that sorts below the prefix and inverts the half-open[prefix, bound)range, so an authorization-rule read or membership-edge enumeration could silently return an empty or under-reported set. The first two now scan back to the last code unit belowchar.MaxValue, increment it, drop the trailing max units, and returnnull(unbounded above) when none exists - mirroring the canonicalBackupConstants.PrefixUpperBound; the bootstrap copy is removed in favour of the already-canonicalRepoContextPortability.PrefixUpperBound. Covered by new regressions. (Orleans.Lattice.Auth9.4.0,Orleans.Lattice.Membership9.4.0,Orleans.Lattice.Api.Mcp.RepoContext)A quota refusal now reaches a remote gRPC caller as
ResourceExhaustedcarrying the breached dimension, instead of an opaqueInternalfault.LatticeQuotaExceededExceptionderives fromInvalidOperationExceptiondeliberately, so an in-process caller that does not know the type still absorbs it - but that same choice meant no gRPC binding ever caught it: a repo-wide search found 36 references across the source tree and not one was acatch. Every admission refusal therefore fell through to the generic handler and surfaced asStatusCode.Internalwith a message telling the caller to read the cluster logs. That is wrong on three counts. It is not a server fault but a deterministic capacity outcome the caller can act on;Internalis the one status a well-behaved gRPC client must not retry, so a transientops-per-secondrate breach that would clear on the next tick was reported as a permanent failure; and the exception's whole payload (TreeId,Dimension,Current,Limit) was discarded, leaving a remote caller unable to tell a rate breach from a live-key ceiling. Both reachable bindings now map it to the canonicalResourceExhausted, the same code each already uses for the sibling storage-saturation refusal, and attach the non-sensitive fields as response trailers (lattice-quota-dimension,lattice-quota-tree, and, where the dimension carries a numeric ceiling,lattice-quota-currentandlattice-quota-limit) following the existing trailer convention, so a client branches on the dimension without parsing prose. No key, value, or tenant id is echoed back. On the schema binding the new arm is deliberately ordered ahead of theInvalidOperationExceptionarm, which would otherwise shadow the subclass. Two neighbouring defects are corrected alongside: the data binding'sLatticeTenantAccessDeniedExceptioncomment claimed to cover the quota-breach case it never saw, andITenantRateLimiter's shipped XML still described the limiter as wiring no enforcement, untrue since the admission controller began consulting it. Covered by new per-binding regressions on a persistent (keys/bytes) and the transientops-per-seconddimension. (#1695) (Orleans.Lattice.Api.Data.Grpc9.4.0,Orleans.Lattice.Api.Schema.Grpc9.4.0,Orleans.Lattice.Tenancy9.4.0)Undoing a tree resize now works while the resize is still running, instead of failing with
Cannot recover a tree that has not been deleted.UndoResizeAsync- surfaced aslattice_treeadmin_tree_resize_undo- documented an after-swap undo window covering theSwap,Reject, andCleanupphases, but its first compensation step unconditionally recovered the old physical tree from soft-delete. Only theCleanupphase ever soft-deletes that tree, and it does so at the very end of the pipeline, so throughoutSwapandReject- and inCleanupitself until the delete lands - the old tree was still live andITreeDeletionGrain.RecoverAsyncrejected the recovery with an internalInvalidOperationExceptionraised by an unrelated grain. Because the recovery was step 1, the undo aborted before clearing shadow-forward on the old shards, removing the alias, deleting the destination tree, restoring the registry entry, or resetting the resize state: nothing was compensated, and every retry failed identically, leaving the tree wedged mid-resize with no working way to undo it. The recovery is now conditional onITreeDeletionGrain.IsDeletedAsync, so undo runs its full compensation in every phase and stays retryable. The probe is deliberately at the call site rather than softeningRecoverAsyncinto a no-op, because on the public tree-recover path "not deleted" genuinely is a caller error and must keep throwing. No data was ever at risk - the defect was confined to the compensation path and the tree stayed readable throughout. Seven regression tests pin undo entered at each ofSnapshot,Swap,Reject, andCleanup(both before and after the soft delete), on a completed resize, and across a repeated call, asserting the full compensation actually ran rather than merely that undo stopped throwing. An eighth pins that a genuinely impossible recovery - the pre-resize tree already purged - still surfaces to the caller rather than being swallowed by the new probe. The undo contract is corrected in Tree sizing, the API reference, and the tree-admin facade README. (#1742) (Orleans.Lattice9.4.1,Orleans.Lattice.Api.TreeAdmin9.4.1)A leaf that receives a CRDT write before the owning shard root has attached it now surfaces the actionable
LatticeCrdtShapeNotRegisteredExceptioninstead of an opaqueArgumentException.BPlusLeafGrain.ApplyCrdtDeltaAsyncdefensively coalesced an unsetTreeIdtostring.Emptybefore callingCrdtShapeRegistry.TryGet, whose openingArgumentException.ThrowIfNullOrEmpty(treeId)guard therefore always threw "The value cannot be an empty string. (Parameter 'treeId')" - a message naming neither the tree, the key, nor the grain - and made the informativeLatticeCrdtShapeNotRegisteredExceptionwritten directly beneath it unreachable, so an operator saw only a bare argument fault where a self-describing one had already been authored. Both leaf CRDT paths now resolve the bound tree id through a single guard before consulting the registry: the producer-side apply, and the prepared-atomic-write terminal-commit fold (FoldPreparedCrdtDelta), which carried the identical?? string.Emptyshape and the same unreachable throw. The typed exception is raised with an emptyTreeIdand a message naming the grain, the key, the merge mode, and the calling path, and states that the leaf received a CRDT write before the shard root seeded it viaSetTreeIdAsync- which points at the question the opaque fault masked (a routing or lifecycle race, for example a tree deleted and recovered underneath in-flight writes) rather than at a missing shape registration. The failure stays fail-closed, so nothing is folded, stored, or written to the WAL before it is raised, and it stays on the same typed exception the API bindings already map to a client-side precondition status rather than an opaque server fault, so no caller-visible catch shape changes and a normally-attached leaf is byte-for-byte unaffected. An audit of the sibling?? string.Emptycall sites in the leaf confirms these two were the only ones reaching an empty-rejecting guard; the rest feed metric tags, WAL record fields, and resolvers that guard againstnullalone. Four regression tests pin the new behaviour on both paths and two positive controls pin the attached-leaf paths. (#1740) (Orleans.Lattice9.4.1)A cold leaf whose replay gap exceeds
MaxLeafReplayEntriesno longer bricks its tree while the write-ahead log is fully intact.LatticeFallOffLogDetector.ClassifyAsyncfunnelled all three of its fall-off triggers into the sameProjectionRebuildPolicyswitch, every branch of which throwsLeafProjectionStaleException("operator-driven rebuild is required"), so two cost signals were reported as unrecoverable corruption: a replay gap wider thanLatticeOptions.MaxLeafReplayEntries(default10_000), and a checkpoint older thanLatticeOptions.LeafProjectionRetention. Neither implies data loss. The affected leaf could then never activate, and because the documented remedyRebuildLeafProjectionAsyncactivates the leaf in order to rebuild it, the prescribed recovery re-threw the very exception it exists to clear, leaving no supported way out - which is what drove a downstream consumer to adopt a delete-and-re-derive workaround for what was never a corruption. Observed in the field on a repository-context deployment: the WAL scanned clean to end of file withtail=13031,persistedCheckpoint=17288, andhead=27936, so the trim trigger was false (nothing the leaf still needed had been trimmed) while the budget trigger was true by 648 entries, 6.5% over budget; a tree holding 35,826 embeddings was left permanently un-activatable while every record it needed was readable, and raising the budget alone replayed all 10,648 entries and recovered it intact. Only genuine loss now reaches the rebuild policy - the WAL trimmed pastcheckpoint + 1, the same boundary the activation-time #945 guard uses - and its exception now states why it is fatal rather than only naming the decision. A cost trigger against a covering WAL returns the new non-fatalFallOffLogDecision.TailReplayOverBudget: the leaf tail-replays as it otherwise would, converging to an identical projection, and reports the overrun as a warning plus a neworleans.lattice.leaf.activation_replays_over_budgetcounter (taggedtree, gated onILogger.IsEnabledso the templated warning allocates nothing when filtered, and surfaced on the Commit Path dashboard). Replay cost is bounded by limiting the work rather than refusing it -WalReplayMaxRecordsPerTurnalready yields between turns andWalMaterialiserMaxConcurrentReplayscaps concurrent replays - because a slow activation is recoverable where a tree that refuses to activate is not. The change is strictly more permissive: a deployment that never trips a trigger takes the identical path, there is no wire-format, persisted-state, or configuration change, and a tree already stuck this way heals itself on its next activation with no operator action.MaxLeafReplayEntriesandLeafProjectionRetentionare retained and keep a tuning role as the threshold at which an over-budget replay is warned and metered. Their XML docs are corrected (the former promised a "fall back to a full projection rebuild" that never existed), as isProjectionRebuildPolicy's, which claimedSnapshotThenWal"works even when the WAL has been trimmed below the leaf's previous checkpoint" while the implementation threw - the per-leaf snapshot rehydrate half is live in core at activation Step 0, but no recovery runs once it has declined - along with Projection rebuild and State model. Also recorded: the activation path passesTimeSpan.Zeroas the checkpoint age, so the retention trigger cannot fire from activation today. Six regression tests pin the contract, including the exact production offsets and the non-regression that genuine loss still fails closed when the budget is also exceeded. (#1738) (Orleans.Lattice9.4.1,Orleans.Lattice.Dashboards9.4.1)A
DeleteTreeAsync/RecoverTreeAsynccycle over a partially purged tree no longer leaves a routable-but-unbound leaf that fails every typed CRDT write to its key range forever.ShardRootGrain.PurgeAsyncclears a shard's leaves and internal nodes before it clears the shard root itself, so an interruption part-way through leaves the shard root'sRootNodeIdandRootIsLeafintact while the nodes they address have had their state wiped - the reported incident was aPurgeTreeAsynccall that exceeded the 30-second grain-call timeout on a tree holding roughly 42,500 vectors, whose out-of-band reminder purge was then abandoned, and purge is explicitly best-effort so this is a reachable state rather than a corruption.EnsureRootSlowAsyncfast-paths onRootNodeId is not nulland callsSetTreeIdAsync/SetShardIndexAsynconly on the branch that creates a node, so a node's tree-id binding was established once at creation and never re-asserted; the recovered shard root therefore kept routing writes to a leaf with no bound tree id, and every typed CRDT write to that key range failed withLatticeCrdtShapeNotRegisteredExceptionpermanently. The field evidence matched that shape exactly: always the same leaf id rather than churn across leaves, recurring roughly every 90 seconds for over ten minutes, surviving a full cold container restart - so persisted damage, not an in-flight activation race - while the rest of the tree read fine.RecoverTreeAsyncnow asks every shard root to re-assert its node bindings, placed between clearingIsDeletedon the shards and persisting the recovered deletion state so that a failed repair leaves the tree still marked deleted and an operator retry stays clean rather than stranding a half-repaired topology behind a "cannot recover a tree that has not been deleted" precondition. The re-assert is near-free on a healthy tree: purge always clears the leftmost leaf first, so a leftmost leaf whose tree id is still bound proves nothing in that shard was cleared and the shard root returns after a single probe, needing no new persisted state and no new serializer field. Only a shard that is actually damaged pays a topology walk, which descends internal nodes rather than the leaf sibling chain - clearing a leaf wipes its sibling pointer, so walking the chain would stop at the first cleared leaf, whereas internal nodes keep their child ids until the internal sweep that only begins after every leaf is cleared - and re-assertsSetTreeIdAsyncon each internal node andSetTreeIdAsyncplusSetShardIndexAsyncon each leaf that routing can still deliver to, bounded at 4,096 nodes per call at a fan-out of 16 so the repair cannot reproduce the timeout that caused the damage. Both setters are already idempotent, so re-asserting a healthy binding is a no-op. The probe is best-effort and logs a warning rather than throwing when the descent itself fails - a shard whose internal root was also cleared has nothing to descend, and recovery already succeeded in that case before this change, so throwing would make recovery newly fragile for no gain. The fail-closed guard that makes an unbound CRDT write throw is deliberately left intact rather than having the leaf resolve its own tree id lazily, because that would mask genuine routing faults and the current fault is a useful signal. Also recorded:TreeDeletionGrain.PurgeNowAsync, the synchronous path behind the publicPurgeTreeAsync, never setsPurgeInProgressbefore its shard walk, which is whyRecoverAsync's existing purge guards all passed on a tree the reminder-driven path would have refused to recover. A tree already damaged by an earlier build heals on its next delete/recover cycle, because the probe reads live state rather than a persisted flag. Six regression tests pin the contract, including end-to-end delete / interrupted purge / recover / CRDT write coverage over both a single-leaf root and an internal root, and a positive control that an undamaged tree stays bound. (#1744) (Orleans.Lattice9.4.1)