Table of Contents

Release 2026-08-29: Fixed

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at 2026-08-29-2.md, and llms.txt lists every page.

Part of Release 2026-08-29, in Changelog.

Fixed

  • A tree left with a routable-but-unbound leaf now heals itself on its next typed CRDT write, with no operator action, no recover call and no restart. The 9.4.1 repair for this issue ran only inside RecoverTreeAsync, which throws on a tree that is not currently soft-deleted, so the only remediation available to a deployment already stuck in this state was to call DeleteTreeAsync() on a live production tree and complete a RecoverTreeAsync() inside the 72-hour SoftDeleteDuration window - a remedy that risks destroying the data it is meant to rescue, on precisely the tree whose purge timeout caused the damage in the first place. The binding is now re-asserted at the point of use instead: when a typed CRDT write faults because its target leaf has no bound tree id, the owning shard root - which always knows the tree id and shard index the leaf is missing, and is the only party that does, since a cleared leaf retains nothing to resolve its own tree from - re-asserts the binding and retries the write once. Driving the repair from the fault rather than from a lifecycle hook is what lets it reach a deployment whose delete and recover already happened hours ago; the healthy path pays nothing, because an exception filter is only evaluated once a fault is already in flight. The repair is deliberately narrow so it cannot mask a genuine fault: it matches only on LatticeCrdtShapeNotRegisteredException carrying an empty TreeId, which is raised on exactly one branch (RequireBoundTreeId), so a real "no shape registered for this tree" fault - which carries the tree id it could not resolve - propagates untouched and is never retried; and the retry is single-shot, so a leaf whose binding cannot be restored still fails closed exactly as before rather than silently accepting an unbound write. The repair logs a warning, because a leaf losing its binding remains an anomaly worth surfacing even though it is now self-correcting. Also fixed alongside it: a splitting leaf seeds its new sibling with its own tree id verbatim, so an unbound leaf minted another unbound leaf on every split and the damage spread across the key range as the tree grew - plain SetAsync writes resolve no CrdtShape, so an unbound leaf keeps accepting them, keeps filling and keeps splitting while only its typed CRDT writes fail, which is why the reported incident's stuck node was a split-created leaf deep in the tree rather than the shard's deterministic root leaf. The donor now logs a warning naming itself and the sibling it is about to mint; it deliberately does not throw, since that would break splits on exactly the trees that need to keep serving writes while they heal, and the write-path repair re-binds both nodes on the next typed CRDT write routed to them. Finally, recovery's own re-assert no longer infers a whole shard's health from its leftmost leaf: 9.4.1 probed that leaf and returned immediately when it was still bound, on the reasoning that PurgeAsync clears the leftmost leaf first so a bound leftmost leaf proves nothing in the shard was cleared, but that inference does not hold once a split can inherit an unbound donor's binding, and the same short-circuit also stranded a shard whose previous repair had failed part-way through its fan-out, since the retry then probed an already-re-bound leftmost leaf and returned. ReseedNodeBindingsAsync now walks the topology unconditionally, which is affordable because both setters are idempotent and short-circuit inside the callee, so an already-bound node costs one round trip and no storage write and recovery is a rare operator action; the walk stays bounded at 4,096 nodes at a fan-out of 16, and nodes beyond that budget are now re-bound by the write path on their next typed CRDT write rather than being reported as permanently unwritable. Five regression tests pin the new contract, including self-heal of an unbound leaf with no recover call at all, self-heal of a non-leftmost split-created leaf, self-heal after a full delete / interrupted purge / recover cycle, and negative controls that a genuinely unregistered shape still throws un-retried and that an unrepairable binding still fails closed. (#1744) (Orleans.Lattice)

  • A cluster storage-usage roll-up now reports an honest lower bound instead of a silently understated total, and no longer collapses into response timeouts on a many-tree cluster. Three defects compounded in the same path. First, LatticeStorageUsageGrain handled a failed shard-root usage read by contributing a zeroed ShardStorageUsage into the tree's totals and leaving TreeStorageUsageReport.Partial unset, so one shard that failed or timed out understated LeafStateBytes, SnapshotBytes, LiveKeys, and TotalBytes while still presenting the report as complete - a wrong answer dressed as an authoritative one. The sibling WAL path had always handled the identical case correctly, returning a sentinel the caller turned into Partial; the shard path simply never adopted the pattern. A surface that does not answer now contributes nothing and sets Partial, so the figures are a flagged lower bound and the existing "one bad shard does not abort the tree" resilience is unchanged. Second, the roll-up was a two-level uncapped fan-out - one unbounded Task.WhenAll per tree in LatticeAdminGrain, each fanning out again over every shard root and WAL partition - and the levels multiplied: a 90-tree cluster at the default 64 shards and 8 WAL partitions dispatched roughly 6,500 concurrent grain calls in one burst, all racing a single 30 s Orleans response deadline, which was observed failing wholesale on a local harness and stopping the moment the tree count was lowered. Both levels are now bounded by new options - MaxConcurrentStorageUsageTrees (default 8) and MaxConcurrentStorageUsageSurfaces (default 16, spanning shard roots and WAL partitions jointly) - capping the peak at 128 in flight regardless of cluster size, so the roll-up degrades in latency rather than collapsing. Bounding the burst is necessary but not sufficient, because it caps concurrency and not total work: a deep refresh re-walks every shard of every tree, so a large enough catalogue still cannot be sampled inside one response deadline however gently it is dispatched, and the whole call then fails and tells the caller nothing. A new StorageUsageRollupBudget (default 20 seconds, cluster-wide, non-positive to disable) caps the total: once it lapses the roll-up stops dispatching, the trees sampled so far keep their real figures, every remaining tree short-circuits to not-answered without dialing its grain, and the report comes back flagged Partial - the same "an honest flagged lower bound beats an absent answer" rule the per-surface reporting follows. A caller-driven cancellation is deliberately still distinguished from a budget lapse and continues to abort the roll-up outright rather than being laundered into a confident-looking partial. The registry-sorted tree ordering that ClusterStorageUsageReport.Trees guarantees is preserved under the bound, and the aggregated figures are identical to the unbounded ones. Third, a cancellation mid-roll-up could be swallowed into a confident-looking zero by the same catch-all that hid shard failures, and the two sequential fan-out joins could abandon the second batch's tasks unobserved if the first threw; cancellation now propagates promptly, every dispatched slot is settled before the fan-out returns, and the background LatticeStorageUsagePoller clamps an out-of-range poll interval and wraps its loop so neither a bad interval nor an options-validation failure can fault the hosted service and take the silo down with it. The byte-pressure over-threshold gauge is deliberately still gated on WAL completeness alone rather than on Partial, so an unrelated shard failure does not silence it. (#1728) (Orleans.Lattice 9.4.0)

  • An Rga merge no longer drops the incoming per-replica counter maxima when the other side carries no nodes. Rga carries two independent pieces of merge state - the node union, and a Context map of the highest counter seen per replica that lets NextCounter mint the next dot in O(1) - but MergeFrom short-circuited on other.Nodes.Count == 0 before folding Context, so an incoming sequence with a populated context and no nodes had its maxima silently discarded. The early return dates to the original RGA primitive and was simply never revisited when Context was introduced several releases later, at which point the guard stopped covering everything the merge is responsible for. The consequence is the same-dot class this release has been closing: an under-counted context lets the next local InsertAfter re-mint an already-authored dot, so two distinct nodes share a dot identity and the sequence stops converging. It is not reachable by local authoring - every path that bumps Context also adds a node - but it is reachable from any decoded wire or storage payload, since Nodes and Context are both public settable serialized properties, and it becomes routine the moment tombstone compaction is added, which is the standard RGA maintenance operation and precisely what Context exists to survive. The fold now happens ahead of the short-circuit, and deliberately still below EnsureContextRebuilt, whose own early return on a non-empty context would otherwise suppress the rebuild of the receiver's own maxima. (#1724) (Orleans.Lattice 9.4.0)

  • Rga.Clone now copies each node's value bytes instead of sharing them, closing one leg of the CRDT buffer-ownership contract. ICrdt<TSelf>.Clone promises a deep, independent copy, and composites lean on it: OrMap.Get hands back Clone() precisely so a caller receives something safe to mutate. Rga duplicated every node but aliased each node's byte[], so a caller reading a sequence out of an OrMap<string, Rga> got a live handle on the map's durable state and could write through a returned RgaNode.Value without passing any mutation API - the same defect already fixed one level up in OrMap.Clone and OrMap.Get, left behind in this sibling. The rule the family follows is now written down on ICrdt<TSelf> and in the primitives instructions, and it is decided by provenance rather than by type: ingress from the caller (Set/Add/InsertAfter) is a documented hand-off and copies nothing, a fold from a peer or a delta copies the winning candidate, and egress to a caller copies. The copy uses a span copy rather than Array.Clone (identical allocation, but roughly 3-4x faster on the ordedup microbench), and an empty span's ToArray() returns the shared Array.Empty<byte>() singleton, so a tombstoned node - which is what an aged RGA is mostly made of - adds no allocation at all. Documented under Who owns the bytes. (#1724) (Orleans.Lattice 9.4.0)

  • The CRDT buffer-ownership contract is now actually honoured across the whole primitive family, and enforced structurally rather than by review. That contract, written down on ICrdt<TSelf> and in the primitives instructions, decides ownership by provenance: ingress from a caller is a hand-off, a fold from a peer or a delta copies the winning candidate, and egress to a caller copies. Four legs were still open, in two primitives. Rga.MergeFrom and Rga.MergeDelta adopted a peer's or a delta's RgaNode.Value by reference - both when creating a node the receiver had not seen and when an incoming value won a same-dot collision - so a fold left the receiver's durable state aliased to an array somebody else still owned. Rga.ToList, the sequence's materialised projection and its primary read API, handed out the live node buffers; because the resolved order is cached and nothing invalidates it, one write through a returned value corrupted every later read - the exact hazard that method's own documentation described, which it had closed for the list shape (via AsReadOnly) while leaving the element buffers open, so no downcast was even needed. MvRegister - one of the four primitives ICrdt<TSelf> names in its own remarks - violated all three legs at once: its MergeFrom and MergeDelta adopted the peer's and the delta's entry buffers on the same two paths, its Clone was a shallow entry copy, and Values returned the stored arrays and cached them, handing two readers the same buffer. The Clone case had a comment justifying itself on the grounds that the value bytes were "treated as immutable by every production call site" - the same reasoning already rejected for Rga.Clone, and the reason this class of defect kept recurring: each previous fix was pinned only by a hand-written test for the one type that had just been fixed, leaving the next sibling free to repeat it. The observable consequence surfaced through OrMap<TKey, TValue>.Get, which hands back a Clone() of the sole contributor precisely so a read is safe to mutate, and folds later contributors in via MergeFrom: over an MvRegister value both paths leaked, including the single-contributor fast path that is the dominant steady-state read, so a caller that mutated what it read wrote into the map. LatticeReplicationConfigEntry inherited the same leak through its mode register without a line of its own being at fault. Every leg now copies with a span copy (AsSpan().ToArray(), matching BoundedRegister), and the cost is kept where it belongs: only a winning fold candidate is copied so an idempotent steady-state merge still allocates nothing, an empty or tombstoned payload reuses the shared Array.Empty<byte>() singleton, the cached projections now cache the ordering (the traversal and sort, which is the expensive part) and pay only the per-value copy, and the internal call sites that immediately consume the bytes - the typed RgaAccessor and MvRegisterAccessor, which deserialise each value on the next line, and the index-resolution paths that read only the dots - are routed at new internal aliasing views (Rga.MaterializeShared, MvRegister.ValuesShared) so they pay nothing at all. No public signature changed. Reachability is deliberately not overstated: every production delta path takes serialized byte[] (ApplyCrdtDeltaAsync) and every accessor decodes per read, so these were latent public-API contract violations rather than in-flight replication corruption. The enforcement is the substantive half of the change: CrdtBufferOwnershipContractTestsBase (shared testing library, bound in Orleans.Lattice and Orleans.Lattice.Replication) walks each registered CRDT's real object graph and compares byte[] instances by reference identity across clone, state fold, delta fold, and every public projection, and additionally fails when a CRDT type declared in the package has no specimen, when a type grows a public byte[]-bearing projection nothing covers, and when a specimen's declared payload shape is wrong - which also pins the set primitives (GSet, OrSet, RwSet) as payload-free by construction, since they encode elements to base64 string keys and retain no caller array. A future primitive therefore cannot join the family without picking the contract up. (#1724) (Orleans.Lattice 9.4.0, Orleans.Lattice.Replication 9.4.0)

  • A bounded register decoded from storage now takes its fold direction from the registered merge mode, not from the stored payload. A MaxRegister and a MinRegister are one primitive pointed in opposite directions, and that direction lived in two places that were never reconciled: the LatticeMergeMode registered for the key (which is what the tree dispatches on, and what the documentation already described as authoritative) and an IsMin bit carried on the state. Neither decode seam stamped the bit from the mode - the CrdtShapeRegistry shape decode and BoundedRegisterAccessorHelper.ReadAsync both took whatever the persisted bytes said - so a payload that reached the store without passing through a directional accessor (a raw SetAsync of hand-authored JSON, a client-supplied state, an older or foreign writer) fixed the wrong direction for the rest of that key's life. Nothing detected it: MergeFrom resolves under the receiver's direction and never inspects the other side's, so a Min key whose payload claimed Max simply kept answering with the greatest value ever written, silently and forever, including through the public GetRegisterAsync projection. Both seams now re-stamp on decode, so a disagreeing payload self-heals on read. The stamp is an in-place write on the just-decoded instance, so it allocates nothing, and IsMin keeps its [Id(3)] slot and its setter - there is no wire-format or public API change, and a correctly-authored payload decodes exactly as before. A direction mismatch is deliberately not raised as an error: throwing on a CRDT fold path can wedge replication on one bad payload and would still leave the stored state wrong. (#1724) (Orleans.Lattice 9.4.0)

  • The ambient vector-clock frontier is now copied at the seams where it becomes durable state or escapes to a caller, so a co-located sender's VersionVector can no longer be rewritten under committed entries. VersionVector is a mutable CRDT, and LatticeVectorClockContext.Current was a plain cast out of RequestContext - one shared instance, assigned straight into the persisted LwwValue<T>.VectorClock of every entry written inside the scope, at write sites across BPlusLeafGrain, AtomicWriteGrain, and ShardRootGrain and onward through the WAL record, snapshot, and replication-apply paths. On the inbound replication path that instance arrives inside an [Immutable] carrier (ApplyCrdtDeltaItem.SourceVectorClock), and Orleans elides the deep copy for an [Immutable] payload when the callee is co-located - so the object promoted into the durable frontier of many entries at once was the sender's own, still reachable and still mutable on both sides. The same aliasing ran outward through LwwEntry, handing a caller a live handle on stored state. Because the elision only happens under co-location, no cross-silo test would ever have shown it. The context setter now copies on the way in and LwwEntry copies on the way out, fixing the whole class at the two narrowest seams rather than at the dozens of sites that read the frontier back; the scope-restore path skips the second copy because its captured value was already copied on entry. Both are ?.Clone(), and a purely local write leaves the frontier null, so the dominant path allocates nothing and the cost elsewhere is one clone per scope rather than per write. Found by the new grain-boundary contract guard below. (#1725) (Orleans.Lattice 9.4.0)

  • The [Immutable] same-silo copy elision is now guarded by a contract test instead of a one-off audit. Orleans skips the deep copy for an [Immutable] payload on a co-located call and hands the receiver the sender's own instance, which is sound only while nobody mutates it - and roughly 140 tracked types carry [Immutable] alongside a mutable byte[] or collection. Re-reading them all was rejected as the wrong shape of work: the risk is structurally bounded, because CRDT payloads cross grain boundaries as opaque byte[] and are decoded into a fresh object graph before being folded, so no shared buffer is ever folded in place. That is a reachability argument about today's signatures rather than an invariant, so it is now enforced by the reflection-driven ImmutableGrainBoundaryContractTestsBase (bound in the core test project, alongside the established serializable-exception and grain-key guards): one test pins that no CRDT state type is shared across a boundary by an elided copy - counting a CRDT reached through an [Immutable] carrier, which is what surfaced the VersionVector aliasing above - and a second requires every [Immutable] payload on the boundary that carries a mutable buffer to be listed with a written justification, with a companion test failing any listed entry that has gone stale. A future grain method that puts a typed CRDT, or any other in-place-folded payload, on the boundary now fails build-and-test rather than reaching production. (#1724) (Orleans.Lattice 9.4.0)

  • Fourteen tools documented "Read-only" no longer create the tree they were asked to read, so a read-only identity can no longer provision durable storage. Every CRDT read verb on the data facade (GCounterGetAsync, PnCounterGetAsync, GSetGetAsync, OrSetGetAsync, RwSetGetAsync, OrFlagGetAsync, RwFlagGetAsync, MvRegisterGetAsync, MaxRegisterGetAsync, MinRegisterGetAsync, OrMapGetAsync, SequenceGetAsync, VersionVectorGetAsync) and LatticeTreeAdmin.GetTreeStatsAsync routed straight into a shard-root grain, whose activation resolves per-tree options through LatticeOptionsResolver - and that resolver seeds a missing registry entry on first use. Reading an id that had never existed therefore registered it as an Active tree with the full default shard fan-out, which TreeExistsAsync and the tree listing then confirmed. The same seeding ran on DeleteAsync and DeleteRangeAsync, whose contracts describe an unknown tree as a routine no-op. The authorization consequence is the sharp edge: a caller holding only Read and RangeRead is not even offered the tree-creation verb, yet could mint an unbounded number of durable trees through the read path, so a privilege refused at the front door was reachable from behind it. The thirteen CRDT reads and the statistics snapshot now probe the tree registry first and short-circuit exactly as GetAsync and ReadRangeAsync already did, returning the documented empty answer (0, an empty sequence, false, null, an empty map, a zeroed statistics snapshot) without provisioning anything. The two delete verbs answer the same way but from inside the core tree grain, sequenced after the delete gate they already applied, so a caller who may not delete still receives an authorization denial rather than a no-op-shaped success. The probe costs one extra grain call only on the miss path, needs no privilege the operation itself did not already require, and every documented empty-read and no-op-delete result is unchanged. (Orleans.Lattice 9.4.0, Orleans.Lattice.Api.Data 9.4.0, Orleans.Lattice.Api.TreeAdmin 9.4.0)

  • MetadataOnly and Hybrid durable-history retention are now actually enforced, so an operator who asks for hash-and-length only no longer keeps full plaintext. MetadataOnly is the default retention mode and is documented as storing a revision's content hash and byte length but not its bytes, with Hybrid keeping bytes only inside a recent window. The view-backed history path honoured the policy, but the write-ahead-log fallback that answers a history query for revisions not yet folded into a materialised view mapped every mutation with the value preview unconditionally attached and stamped the row FullValue regardless of the tree's configured mode - so all three modes behaved identically to FullValue, and GetEntryHistoryAsync returned readable plaintext for a tree whose policy said it was not retained. The fallback now resolves the tree's effective policy once per page and applies the same retention shaping the view path uses: the hash and byte length are always reported, the bytes are attached only when the resolved mode keeps them, and the row is stamped with the mode that actually applied, so the valueRetained flag reported to callers is truthful. A CRDT delta is unaffected, since a delta carries no last-writer-wins value bytes. (Orleans.Lattice 9.4.0, Orleans.Lattice.Api.State 9.4.0)

  • Setting a history-retention policy on a reserved system tree is now refused, as its contract already stated. SetHistoryRetentionAsync documents rejection both for a non-positive window and for a reserved system tree id, but only the window half was enforced: a caller could set an age-bounded retention window on sys-auth-policy or sys-membership-groups and have it persist, which would age out the authorization and membership history the cluster's own audit trail depends on. The verb now rejects a reserved system data tree alongside the identifiers it already refused. The guard is scoped to this one verb, so the first-party add-ons that legitimately configure retention on their own system trees from inside the silo are unaffected. (Orleans.Lattice.Api.TreeAdmin 9.4.0)

  • Dropping a materialised view is now idempotent, and a library-owned view is now protected from a runtime drop. DropViewAsync documents both properties and held neither. A second drop of the same view, or a drop of a name that never existed, failed with a not-found error instead of succeeding quietly; and the four system history views that the auth, membership, and tenancy packages declare from their own initializers dropped successfully, only for the next silo start to re-declare them - which is precisely the reason the contract gives for refusing the drop. The startup-declaration guard only ever inspected views registered through AddLatticeViews, so it never saw them. An absent view is now the documented no-op, and a view whose source tree lives in a reserved namespace is refused with a clear error, so the two cases stay distinct rather than one masking the other. (Orleans.Lattice 9.4.0, Orleans.Lattice.Api.TreeAdmin 9.4.0)

  • Reading the tag catalogue for an index that does not exist now reports not-found instead of silently materialising the index. ListIndexTagsAsync and ScanTagMembersAsync resolved an index handle by name and enumerated it, which both answered an empty catalogue and created the backing tag-{indexName} membership tree. That destroyed a distinction ScanEntriesAsync explicitly documents and that a caller needs in order to tell a typo from a real result: a tag-filtered scan reports IndexNotFound for an index that was never materialised and Found with zero entries for a real-but-empty one, but a single catalogue read of the mistyped name flipped it permanently to Found. Both verbs now resolve without materialising and report the same not-found the tree-administration tag verbs already do, leaving a real-but-empty index answering an empty page as before. (Orleans.Lattice.Api.State 9.4.0)

  • MCP tool discovery no longer reports a backend timeout as a valid, tiny permission set. Resolving a caller's facade-group access catches a failure and fails closed, which is correct for a denial - but the same catch swallowed a cancelled, deadline-exceeded, or unavailable transport fault and returned an empty grant set, so tools/list answered successfully advertising a single meta-tool. A platform administrator entitled to the full tool surface was observed advertised one tool with no error anywhere on the wire, and a client had no way to tell "your permissions were revoked" from "the silo did not answer in time". A fault that means no authoritative answer arrived - a transport status of Cancelled, DeadlineExceeded, Unavailable, Internal, ResourceExhausted or Aborted, an Orleans response deadline, or silo churn - is now surfaced to the client as a retryable discovery error rather than an answer, while a status that is an answer (PermissionDenied, Unauthenticated, NotFound, and every argument-shaped code) still fails closed exactly as before. The fail-closed guarantee is strictly preserved: raising the error advertises nothing at all, so the new behaviour is never wider than the behaviour it replaces. (Orleans.Lattice.Api.Mcp 9.4.0)

  • Reshard and resize initiation no longer answer "is this tree empty?" with a strongly-consistent whole-tree count, which could time the operation out under concurrent write load. TreeReshardGrain.ReshardAsync and TreeResizeGrain.ResizeAsync each decide whether to take an empty-tree fast path by asking ILattice.CountAsync for a live-key count, but that count walks every leaf chain of every shard, discards its result and retries whenever the shard map moves under it, and gives up only once MaxScanRetries is exhausted - after which CountAsync's own stale-alias envelope runs the whole loop a second time. Initiation is precisely when the map is most likely to be churning: a caller may be writing concurrently, and a small leaf fan-out splits continuously. The probe could therefore consume the caller's entire response budget and surface as a TimeoutException before the operation had started - observed in CI as Reshard_under_concurrent_write_load_does_not_wedge_with_short_forward_deadline timing out at 30 s with the initiation turn still awaiting the count and the coordinator never reached. Both now probe existence directly through the new internal IShardRootGrain.AnyAsync, which short-circuits at the first non-empty leaf, OR-ed across the tree's physical shards. That needs no reconciliation against a moving shard map, which is the crux of the difference: a count must reconcile because a key migrating between shards is briefly visible on both the source and the destination and would be double-counted, whereas a split only ever moves keys - never creating one, destroying one, or leaving one present on neither side - so a key that exists is observed by at least one shard wherever the split has got to, and observing it twice still just means "a key exists". The answer is deliberately one-sided: it may report non-empty while the last keys migrate away, but never empty while a key exists anywhere, and only "empty" unlocks a fast path, so the one consequential direction cannot be wrong. A new LatticeOptions.EmptyTreeProbeBudget (default 10 seconds, InfiniteTimeSpan to wait indefinitely) bounds a probe that parks rather than returns, and every inconclusive outcome is read as "not empty" so initiation proceeds down the normal coordinator path. The same count-to-answer-a-boolean shape is also removed from the replication config authority: the internal ILatticeTreeContentProbe now exposes HasContentAsync instead of CountAsync, implemented by taking the first key from the tree's key stream, so enabling replication on a large tree no longer pays a whole-tree count to learn whether a bootstrap snapshot is needed. Documented under EmptyTreeProbeBudget. (Orleans.Lattice 9.4.0, Orleans.Lattice.Replication 9.4.0)

  • The leaf delivery epoch is now unique across processes, so a cache can no longer be silently denied its full resync. BPlusLeafGrain's delivery epoch identifies one leaf activation and is what tells a LeafCacheGrain holding a stale cursor to fall back to a full-snapshot delivery, but it was minted from a static long with no initialiser - zero at the start of every silo process - so every silo independently handed out 1, 2, 3, .... The epoch is compared across processes (a cache outlives the silo whose activation minted the cursor it holds), so two activations in different silos could mint the same epoch, and a collision is not benign: the epoch-mismatch branch is skipped, the caller's higher sequence then satisfies the already at head check, and the leaf ships an empty delta. The cache adopts the new cursor and resumes incremental delivery, so fresh writes still flow and the fault looks healthy - but the one-time resync never happens, every key that diverged while the cache was disconnected stays stale indefinitely, and because LeafCacheGrain gates its own cache eviction on the same epoch flip, keys the leaf has since deleted stay visible in the cache's read view. Collisions cluster precisely around silo restarts, when a surviving cache holding a low-numbered epoch is most likely, and the failure is silent - no exception, no log, no metric. The seed is now randomised per process into [1, long.MaxValue / 2] (0 stays reserved for LeafDeliveryCursor.Empty, and per-activation monotonicity within a process is unchanged), and as defence in depth the leaf now also treats a cursor whose sequence is ahead of the activation's as stale - impossible to have been issued by that activation - and answers it with a snapshot, with LeafCacheGrain applying the matching condition so the snapshot still triggers its eviction pass. No wire-format change: LeafDeliveryCursor keeps its [Id(0)] long Epoch; only the value's provenance changes. (#1711) (Orleans.Lattice 9.4.0)

  • The WAL materialiser pin grain key is now storage-safe and unambiguous, healing itself across the change. WalMaterialiserPinGrain is persistent and its key was composed as {treeName}#s{shard}. Azure Table grain storage carries a grain key into the Partition/Row key columns and the request URL, both of which reject #, so the key depended on the backend sanitising it - and such sanitisation is lossy, which can collapse two distinct grain identities onto one persisted row. A materialiser pin holds WAL retention state, so a confusion there is durability-relevant rather than cosmetic. The shard suffix now uses the storage-safe ~s, and the composer is marked [GrainKeyBuilder] so the existing reflection-driven storage-safety guard audits it from now on instead of relying on review. The change strands no WAL and needs no operator action: the GC's read fan-in already performed a dual read to pick up pre-sharding keys, and now also reads every shard under the legacy separator, so a pin written by an earlier build keeps holding the trim floor while new pins are written under the safe key as their consumers re-pin. Suffix parsing is additionally anchored at the last separator and accepts only an all-digit suffix, fixing an ambiguity in the previous leading-IndexOf parse that truncated a tree whose own name contained the separator. (#1701) (Orleans.Lattice 9.4.0)

  • Ten convergence and contract defects across the CRDT primitives, found by a consolidated audit of src/lattice/Primitives/ after #1705. That change fixed OrMap.Clone's nested-value aliasing but stopped there, and the same defect class was still live on four further OrMap seams and on two sibling primitives it never touched. OrMap.MergeFrom, OrMap.MergeDelta, OrMap.Set and OrMap.Get all retained a caller's, peer's, or delta's value-CRDT instance by reference, so a map's durable state stayed aliased to an object the other side still owned: a.MergeFrom(b) left a and b sharing nested CRDTs (violating purity of merge in its argument, and letting a later mutation of a silently rewrite b), MergeDelta adopted the producer's delta objects (so applying a delta that is legitimately retried or fanned out to several peers corrupted it), Set stored the caller's instance, and - worst, because it is a read - Get folded the first contributing entry into a fresh accumulator by mutating it, which for a single-entry key meant every Get mutated the very state it was reporting; a BoundedRegister value read twice returned different answers. BoundedRegister.Clone shared its value and order-key byte arrays, so a register handed out of OrMap.Get gave the caller a live handle on the map's persisted bytes; the deep copy ICrdt.Clone promises is now made (measured cost below), and the merge/delta-apply fold likewise copies a winning candidate rather than adopting the other side's buffer, while Set keeps its documented hand-off. OrMap.Keys() sorted with the ambient culture, so two replicas under different CurrentCulture settings enumerated the same converged map in different orders - under en-US the keys a and B order as a, B, under the ordinal comparison every other enumeration surface in the library uses they order as B, a; keys are now ordinal-sorted. CrdtShapeRegistry mapped a null OR-Set delta element to the empty string when building its dedup key, so a null-element dot and a legitimately empty-element dot with the same (replicaId, counter) collided and one was silently dropped during delta combination; null elements are now skipped, matching the sibling G-Set path. LwwValue<T>.Merge was not commutative: its tie-break ran out after (Timestamp, OriginClusterId, IsTombstone) and then returned the left operand unconditionally, so two writes agreeing on all three but differing in ExpiresAtTicks, Value, or IsMigrated resolved to whichever arrived first and two replicas never converged - reachable because OriginClusterId is null for every purely local write and HybridLogicalClock carries no node id, making a same-wall-tick collision between two leaves ordinary. The order now extends through those three fields (Value lexicographically for byte[], the only instantiation the library persists) and the XML doc states the residual honestly instead of claiming a commutativity it cannot deliver for an arbitrary unordered T. HybridLogicalClock overflowed its logical counter: Tick and all three Merge branches incremented Counter unguarded, so a clock pinned at int.MaxValue by a peer's malformed or malicious stamp wrapped to int.MinValue and went backwards, breaking the monotonicity every last-writer-wins decision in the library rests on; the counter now saturates. Rga.ToList cached and handed back the mutable backing list, so a caller who downcast the returned IReadOnlyList<> could rewrite - or empty - the RGA's cached materialisation for every subsequent reader; the cache now holds a read-only view. And MvRegisterEntry, OrMapDelta<,> and OrMapDeltaEntry<,> were marked [Immutable] while carrying a mutable byte[] or a mutable value-CRDT the receiver folds in place, which told Orleans to skip the deep copy on a same-silo grain call and hand the receiver the producer's own objects; the attribute is removed from the three, with the reason recorded on each type. Every one of the ten was proved by a regression test that fails against the prior code before the fix was written; two further reported findings were investigated and disproved (VersionVector's handling of an all-zero clock slot, which is a correct join whose IsBottom is documented as a structural hint, and LwwValue.Merge's associativity, which the left bias in fact preserved - only commutativity was broken) and are unchanged. (Orleans.Lattice 9.4.0)

  • A materialised view's tree id is now usable and unambiguous as a grain key, and the change heals itself. A view tree id is an Orleans grain primary key and is carried into ShardRootGrain's composite key ({treeId}/{shardIndex}) - a persistent grain - but neither property that .github/instructions/grains.instructions.md requires of such a key was held. The shadow-swap generation suffix used #, one of the characters (/, \, #, ?, plus the control characters 0x00-0x1F / 0x7F-0x9F) that Azure Table grain storage rejects because it carries a grain key into the Partition/Row key columns and the request URL; the historical failure mode is an opaque HTTP 400 that no in-memory test storage reproduces, which is how #1529 reached production. And because a view name was validated only for null/empty, a view named a/b produced the persistent shard-root key view-a/b/0, while a view named orders#g2 collided with generation 2 of a view named orders - two grain identities, one state row. View names are now validated at every creation seam (the storage-unsafe set, plus the reserved generation separator), and the generation suffix uses the storage-safe ~; the composer is marked [GrainKeyBuilder] so the existing reflection-driven storage-safety guard audits it automatically from now on. Rejecting / also closes a tenancy hazard for free: without it a caller could name a view t/other/orders and have the tree planted in another tenant's reserved namespace. The separator change strands no data and forces no rebuild. The maintainer records a legacy-generation ceiling, pinned once at activation to the generation then active: generations at or below it keep resolving through the old separator, and the next rebuild allocates a higher generation under the new one, so a view converges on its own with no operator action and the superseded generation is still reclaimed on the normal grace cadence. Pinning happens at activation rather than inside a particular verb precisely because a rebuild can be driven without the activation path having run first - a ceiling pinned after a generation was already allocated would misclassify it and send every read to a tree that was never written. A view created after the upgrade pins 0, and generation 0 carries no suffix, so every generation it allocates is storage-safe. A name that an older build accepted is still restored, and reported in the log rather than refused, so adopting the rule cannot strand an existing view. (#1696) (Orleans.Lattice 9.4.0, Orleans.Lattice.Api.TreeAdmin 9.4.0)

  • A reserved-namespace rejection on the data gRPC binding is now a typed client error instead of an opaque server fault. The core reserved-namespace guards threw a bare InvalidOperationException, which fell through LatticeDataApiGrpcService's typed catch arms to the generic handler and surfaced as StatusCode.Internal with "The data-API request failed; see the cluster logs for the cause". Naming a tree in a reserved, internally-composed namespace is a deterministic caller-side precondition - the id is not addressable through the public surface for anyone - so reporting it as a server fault sent callers to the cluster logs for something they could see and fix themselves. The guards now throw a new LatticeReservedTreeNamespaceException (deriving from InvalidOperationException, so existing broader catches are unaffected, and carrying the required no-op [RegisterCopier] for same-silo deep copies), which the binding maps to InvalidArgument carrying the self-contained message. Covered by new regressions on both the tenant and system-data namespaces. (Orleans.Lattice 9.4.0, Orleans.Lattice.Api.Data.Grpc 9.4.0)

  • A single-scope CaptureSetAsync no longer hands back a set id that matches nothing. Set membership is not stored as a set: BackupSetManifest is returned to the caller but never persisted, and the only durable trace is the SetId / SetName / SetCreatedAtUtc stamp StampSetMembershipAsync pushes onto each member's own manifest. That stamping is deliberately skipped for a one-member set - it is indistinguishable from a plain backup and lists as one, and a one-tree set has no cross-tree atomicity to preserve - but BuildSetManifest ran unconditionally and still minted a content-addressed SHA-256 id for it. The result was a 64-hex identifier indistinguishable from a valid one that appeared nowhere else: ListBackupsAsync reported SetId = null on the member's catalog row, so a remote consumer grouping catalog rows by the id the create response had just returned found nothing, with no error to catch - a "backup sets" view reasonably concluded the set had been lost or pruned. The reach was not in-process only; the value crosses the backup control facade, its gRPC binding, and the MCP remote adapter, where grouping catalog rows is the only thing a remote consumer can do with a set id (there is no remote RestoreSet operation). BackupSetManifest.SetId is now string? and a one-member set reports null, so the create response agrees with the catalog row by construction; an empty id is still rejected, because that is a malformed id rather than the absence of one. The deliberate < 2 stamping rule is unchanged - the skipped stamp was correct and the phantom id was the defect - and both halves of the decision now derive from one shared BackupSetIdentity (which also owns the content-address algorithm), with the stamping pass keying off the minted id rather than re-deriving the threshold, so the id and the stamp cannot drift apart. In-cluster, RestoreSetAsync's unresolved-set failure no longer reads as though the set once existed: it re-derives the single-member content address of each catalogued backup id and, on a match, names the exact backupId to pass to RestoreAsync instead, so anyone who persisted an id from an older build is rescued rather than told the set does not exist; a genuinely unknown id keeps a distinct message. The extra catalog walk runs only on that failure path, which throws either way. Covered by new capture regressions asserting the create response and the catalog rows agree for both arities (including that grouping rows by a reported id resolves exactly the set's members), a rescue regression pinning the message content and ParamName, wire and MCP-adapter round trips of the absent id, and a new BackupSetIdentity suite pinning the content address against an independent computation so previously-issued ids cannot be orphaned. Breaking for a consumer that reads SetManifest.SetId as non-nullable. (#1687) (Orleans.Lattice.Backup 9.4.0, Orleans.Lattice.Api.Backup 9.4.0, Orleans.Lattice.Api.Backup.Grpc 9.4.0, Orleans.Lattice.Api.Mcp 9.4.0, Orleans.Lattice.Explorer 9.4.0)

  • The RGA sequence CRDT no longer mints a colliding dot after a legacy-payload merge. Rga.MergeFrom and Rga.MergeDelta folded the incoming side's per-replica maxima into the serialized Context counter cache but never rebuilt this sequence's own maxima first, so a sequence loaded from a payload that predates the Context field (nodes present, cache deserialized empty) that merged a peer before its first local insert was left with a Context reflecting only the incoming side. The next InsertAfter then re-minted an already-authored dot - NextCounter sees a now-non-empty Context and skips its lazy rebuild - producing two nodes sharing one (replicaId, counter) identity, which breaks convergence: replicas that applied the same operations in a different order no longer agreed on the sequence. Both merge entry points now call EnsureContextRebuilt() before folding, exactly as the sibling OrMap.MergeFrom / OrMap.MergeDelta already do; the pure-local InsertAfter path was already guarded. No public API change. Covered by new RgaContextTests regressions on both the full-state (MergeFrom) and delta (MergeDelta) merge paths. (Orleans.Lattice 9.4.0)

  • The RGA sequence CRDT's full-state merge now reattaches a tombstone-before-insert placeholder's parent, so the merge stays commutative. When a tombstone is delivered before its matching insert (out-of-order or partial delivery), Rga.MergeDelta records a tombstoned placeholder node whose ParentDot is a stand-in Root so a later insert can reattach the real parent. MergeDelta's insert path performs that reattachment, but Rga.MergeFrom - the full-state merge reached by the public Rga.Merge and the replication applier's registered CRDT merger - folded only the incoming side's tombstone flag and value for a dot already present locally and never reattached ParentDot. A replica that recorded a placeholder via a delta and then learned the dot's authoritative structure through a full-state anti-entropy / catch-up merge kept the placeholder's live children mis-rooted under Root, so Merge(a, b) and Merge(b, a) linearised those children in different positions and the sequence stopped converging. MergeFrom now folds ParentDot by the same deterministic max rule it already applies to the value (a real parent's Counter >= 1 dominates the Root placeholder's Counter 0), making the reattachment order-independent. No public API or wire-format change. Covered by new RgaMergeConvergenceTests regressions. (Orleans.Lattice 9.4.0)

  • OrMap<TKey, TValue>.Clone now deep-copies each key's nested value CRDT, so a map merge no longer mutates its left operand. ICrdt<T>.Clone is contracted to return "a deep, independent copy ... mutating the returned value must never affect the receiver", and the static OrMap.Merge(a, b) is a.Clone().MergeFrom(b). But Clone duplicated each key's entry list while sharing the OrMapEntry objects - and their nested value CRDTs - with the source by reference. MergeFrom resolves a same-dot value collision in place via existing.Value.MergeFrom(other.Value), so when two replicas authored the same (replicaId, counter) dot under one key with divergent values, Merge(a, b) mutated a's nested value in place: the merge was not pure, and a caller still holding a saw it change under an operation documented to leave both operands untouched. Clone now copies each entry and clones its value (new OrMapEntry<TValue>(e.ReplicaId, e.Counter, e.Value.Clone())), mirroring Rga.Clone's per-node deep copy, so a merge folds only into the throwaway clone. No public API or wire-format change. Covered by new OrMapCloneIsolationTests regressions. (Orleans.Lattice 9.4.0)

  • The multi-value register's merge is now commutative when two replicas disagree on the value carried under one dot. MvRegister.MergeFrom (and the equivalent MergeDelta) keeps a dot still present on both sides, but on such a collision it kept the local side's value unconditionally, so Merge(a, b) and Merge(b, a) disagreed whenever the same (replicaId, counter) dot carried different bytes on each side - a reachable divergence, because a replicaId is a caller-supplied string with no minted-once guarantee, yet the type documents itself "commutative, associative, and idempotent". Both merge paths now break a same-dot value tie deterministically, keeping the lexicographically-greater value bytes (the same byte-order rule Rga already applies), so the two merge orders converge. The lazy survivors-list allocation is preserved, so an idempotent re-merge still copies nothing and touches no state. No public API or wire-format change. Covered by new MvRegisterMergeConvergenceTests regressions on both merge paths. (Orleans.Lattice 9.4.0)

  • The RGA sequence CRDT's delta merge now folds a same-dot parent/value collision by the same deterministic max rule as its full-state merge, so delta-fed replicas converge. Rga.MergeFrom resolves a dot already present locally by a deterministic max - CompareDot for the structural parent, CompareBytes for the value - but Rga.MergeDelta's insert path unconditionally overwrote existing.ParentDot and existing.Value with the arriving delta's (last-arrival-wins). Two replicas that applied the same colliding-dot inserts as deltas in different orders therefore disagreed with each other, and with a full-state-fed replica, even though the type documents itself "commutative, associative, and idempotent". MergeDelta now applies the identical CompareDot / CompareBytes max rules, so an insert refresh is order-independent; a real parent (Counter >= 1) still dominates a tombstone-before-insert placeholder's Root parent (Counter 0), so placeholder reattachment is unchanged, and a re-delivered identical insert stays a no-op. No public API or wire-format change. Covered by new RgaMergeDeltaConvergenceTests regressions. (Orleans.Lattice 9.4.0)

  • The default replication framing decoder's per-entry bounds check no longer integer-overflows on a forged length. IReplicationBatchEncoder.TryDecodeFraming - the default interface implementation inherited by any encoder that does not override it - validated each entry body with a narrowing cursor + length sum. Both operands are int and length is read straight off the wire (up to int.MaxValue), so the sum could overflow to a negative value and slip past the truncation guard, downgrading the fail-closed ArgumentException framing rejection into a raw ArraySegment out-of-bounds throw. The sum is now widened to long, mirroring the sibling fix already applied to OrleansBinaryReplicationBatchEncoder. Covered by a new regression. (Orleans.Lattice.Replication 9.4.0)

  • The default replication framing decoder's routing-field bounds check no longer integer-overflows. The same narrowing cursor + length defect affected ReadLengthPrefixedUtf8, which decodes the treeName and originClusterId routing fields; a forged length near int.MaxValue overflowed past the guard into a raw Span.Slice / Encoding.UTF8.GetString throw instead of the descriptive framing rejection. Widened to long as above. Covered by a new regression. (Orleans.Lattice.Replication 9.4.0)

  • The auth-admin prefix-scope upper bound no longer wraps a trailing U+FFFF code unit. LatticeAuthAdmin.PrefixUpperBound - reached from ExplainAsync via TranslateScope, which turns a Prefix-scoped introspection request into the exclusive RangeEnd handed to the access gate - incremented the prefix's final code unit unconditionally. A prefix ending in U+FFFF therefore wrapped to U+0000, producing an upper bound that sorts below the prefix and inverts the half-open [prefix, bound) range, so the explained verdict could silently drop matching rules. It now scans back to the last code unit below char.MaxValue, increments that, drops the trailing max units, and returns null (unbounded above) when none exists - mirroring the canonical BackupConstants.PrefixUpperBound. Covered by a new regression. (Orleans.Lattice.Api.Auth 9.4.0)

  • Three more prefix-scope upper-bound helpers no longer wrap a trailing U+FFFF code unit. The same non-canonical unconditional-increment PrefixUpperBound that was fixed in LatticeAuthAdmin survived in three sibling range scans over security-relevant data: LatticeAuthorizationPolicyStore.PrefixUpperBound bounds the authorization-rule scan behind ListRulesForTreeAsync, LatticeMembershipDirectory.PrefixUpperBound bounds both the forward group-closure walk and the reverse members-of scan over membership edges, and the repo-context bootstrap indexer's local copy bounds its stored-file-meta reconcile scan. In each case a prefix ending in U+FFFF wrapped to U+0000, producing an upper bound that sorts below the prefix and inverts the half-open [prefix, bound) range, so an authorization-rule read or membership-edge enumeration could silently return an empty or under-reported set. The first two now scan back to the last code unit below char.MaxValue, increment it, drop the trailing max units, and return null (unbounded above) when none exists - mirroring the canonical BackupConstants.PrefixUpperBound; the bootstrap copy is removed in favour of the already-canonical RepoContextPortability.PrefixUpperBound. Covered by new regressions. (Orleans.Lattice.Auth 9.4.0, Orleans.Lattice.Membership 9.4.0, Orleans.Lattice.Api.Mcp.RepoContext)

  • A quota refusal now reaches a remote gRPC caller as ResourceExhausted carrying the breached dimension, instead of an opaque Internal fault. LatticeQuotaExceededException derives from InvalidOperationException deliberately, so an in-process caller that does not know the type still absorbs it - but that same choice meant no gRPC binding ever caught it: a repo-wide search found 36 references across the source tree and not one was a catch. Every admission refusal therefore fell through to the generic handler and surfaced as StatusCode.Internal with a message telling the caller to read the cluster logs. That is wrong on three counts. It is not a server fault but a deterministic capacity outcome the caller can act on; Internal is the one status a well-behaved gRPC client must not retry, so a transient ops-per-second rate breach that would clear on the next tick was reported as a permanent failure; and the exception's whole payload (TreeId, Dimension, Current, Limit) was discarded, leaving a remote caller unable to tell a rate breach from a live-key ceiling. Both reachable bindings now map it to the canonical ResourceExhausted, the same code each already uses for the sibling storage-saturation refusal, and attach the non-sensitive fields as response trailers (lattice-quota-dimension, lattice-quota-tree, and, where the dimension carries a numeric ceiling, lattice-quota-current and lattice-quota-limit) following the existing trailer convention, so a client branches on the dimension without parsing prose. No key, value, or tenant id is echoed back. On the schema binding the new arm is deliberately ordered ahead of the InvalidOperationException arm, which would otherwise shadow the subclass. Two neighbouring defects are corrected alongside: the data binding's LatticeTenantAccessDeniedException comment claimed to cover the quota-breach case it never saw, and ITenantRateLimiter's shipped XML still described the limiter as wiring no enforcement, untrue since the admission controller began consulting it. Covered by new per-binding regressions on a persistent (keys / bytes) and the transient ops-per-second dimension. (#1695) (Orleans.Lattice.Api.Data.Grpc 9.4.0, Orleans.Lattice.Api.Schema.Grpc 9.4.0, Orleans.Lattice.Tenancy 9.4.0)

  • Undoing a tree resize now works while the resize is still running, instead of failing with Cannot recover a tree that has not been deleted. UndoResizeAsync - surfaced as lattice_treeadmin_tree_resize_undo - documented an after-swap undo window covering the Swap, Reject, and Cleanup phases, but its first compensation step unconditionally recovered the old physical tree from soft-delete. Only the Cleanup phase ever soft-deletes that tree, and it does so at the very end of the pipeline, so throughout Swap and Reject - and in Cleanup itself until the delete lands - the old tree was still live and ITreeDeletionGrain.RecoverAsync rejected the recovery with an internal InvalidOperationException raised by an unrelated grain. Because the recovery was step 1, the undo aborted before clearing shadow-forward on the old shards, removing the alias, deleting the destination tree, restoring the registry entry, or resetting the resize state: nothing was compensated, and every retry failed identically, leaving the tree wedged mid-resize with no working way to undo it. The recovery is now conditional on ITreeDeletionGrain.IsDeletedAsync, so undo runs its full compensation in every phase and stays retryable. The probe is deliberately at the call site rather than softening RecoverAsync into a no-op, because on the public tree-recover path "not deleted" genuinely is a caller error and must keep throwing. No data was ever at risk - the defect was confined to the compensation path and the tree stayed readable throughout. Seven regression tests pin undo entered at each of Snapshot, Swap, Reject, and Cleanup (both before and after the soft delete), on a completed resize, and across a repeated call, asserting the full compensation actually ran rather than merely that undo stopped throwing. An eighth pins that a genuinely impossible recovery - the pre-resize tree already purged - still surfaces to the caller rather than being swallowed by the new probe. The undo contract is corrected in Tree sizing, the API reference, and the tree-admin facade README. (#1742) (Orleans.Lattice 9.4.1, Orleans.Lattice.Api.TreeAdmin 9.4.1)

  • A leaf that receives a CRDT write before the owning shard root has attached it now surfaces the actionable LatticeCrdtShapeNotRegisteredException instead of an opaque ArgumentException. BPlusLeafGrain.ApplyCrdtDeltaAsync defensively coalesced an unset TreeId to string.Empty before calling CrdtShapeRegistry.TryGet, whose opening ArgumentException.ThrowIfNullOrEmpty(treeId) guard therefore always threw "The value cannot be an empty string. (Parameter 'treeId')" - a message naming neither the tree, the key, nor the grain - and made the informative LatticeCrdtShapeNotRegisteredException written directly beneath it unreachable, so an operator saw only a bare argument fault where a self-describing one had already been authored. Both leaf CRDT paths now resolve the bound tree id through a single guard before consulting the registry: the producer-side apply, and the prepared-atomic-write terminal-commit fold (FoldPreparedCrdtDelta), which carried the identical ?? string.Empty shape and the same unreachable throw. The typed exception is raised with an empty TreeId and a message naming the grain, the key, the merge mode, and the calling path, and states that the leaf received a CRDT write before the shard root seeded it via SetTreeIdAsync - which points at the question the opaque fault masked (a routing or lifecycle race, for example a tree deleted and recovered underneath in-flight writes) rather than at a missing shape registration. The failure stays fail-closed, so nothing is folded, stored, or written to the WAL before it is raised, and it stays on the same typed exception the API bindings already map to a client-side precondition status rather than an opaque server fault, so no caller-visible catch shape changes and a normally-attached leaf is byte-for-byte unaffected. An audit of the sibling ?? string.Empty call sites in the leaf confirms these two were the only ones reaching an empty-rejecting guard; the rest feed metric tags, WAL record fields, and resolvers that guard against null alone. Four regression tests pin the new behaviour on both paths and two positive controls pin the attached-leaf paths. (#1740) (Orleans.Lattice 9.4.1)

  • A cold leaf whose replay gap exceeds MaxLeafReplayEntries no longer bricks its tree while the write-ahead log is fully intact. LatticeFallOffLogDetector.ClassifyAsync funnelled all three of its fall-off triggers into the same ProjectionRebuildPolicy switch, every branch of which throws LeafProjectionStaleException ("operator-driven rebuild is required"), so two cost signals were reported as unrecoverable corruption: a replay gap wider than LatticeOptions.MaxLeafReplayEntries (default 10_000), and a checkpoint older than LatticeOptions.LeafProjectionRetention. Neither implies data loss. The affected leaf could then never activate, and because the documented remedy RebuildLeafProjectionAsync activates the leaf in order to rebuild it, the prescribed recovery re-threw the very exception it exists to clear, leaving no supported way out - which is what drove a downstream consumer to adopt a delete-and-re-derive workaround for what was never a corruption. Observed in the field on a repository-context deployment: the WAL scanned clean to end of file with tail=13031, persistedCheckpoint=17288, and head=27936, so the trim trigger was false (nothing the leaf still needed had been trimmed) while the budget trigger was true by 648 entries, 6.5% over budget; a tree holding 35,826 embeddings was left permanently un-activatable while every record it needed was readable, and raising the budget alone replayed all 10,648 entries and recovered it intact. Only genuine loss now reaches the rebuild policy - the WAL trimmed past checkpoint + 1, the same boundary the activation-time #945 guard uses - and its exception now states why it is fatal rather than only naming the decision. A cost trigger against a covering WAL returns the new non-fatal FallOffLogDecision.TailReplayOverBudget: the leaf tail-replays as it otherwise would, converging to an identical projection, and reports the overrun as a warning plus a new orleans.lattice.leaf.activation_replays_over_budget counter (tagged tree, gated on ILogger.IsEnabled so the templated warning allocates nothing when filtered, and surfaced on the Commit Path dashboard). Replay cost is bounded by limiting the work rather than refusing it - WalReplayMaxRecordsPerTurn already yields between turns and WalMaterialiserMaxConcurrentReplays caps concurrent replays - because a slow activation is recoverable where a tree that refuses to activate is not. The change is strictly more permissive: a deployment that never trips a trigger takes the identical path, there is no wire-format, persisted-state, or configuration change, and a tree already stuck this way heals itself on its next activation with no operator action. MaxLeafReplayEntries and LeafProjectionRetention are retained and keep a tuning role as the threshold at which an over-budget replay is warned and metered. Their XML docs are corrected (the former promised a "fall back to a full projection rebuild" that never existed), as is ProjectionRebuildPolicy's, which claimed SnapshotThenWal "works even when the WAL has been trimmed below the leaf's previous checkpoint" while the implementation threw - the per-leaf snapshot rehydrate half is live in core at activation Step 0, but no recovery runs once it has declined - along with Projection rebuild and State model. Also recorded: the activation path passes TimeSpan.Zero as the checkpoint age, so the retention trigger cannot fire from activation today. Six regression tests pin the contract, including the exact production offsets and the non-regression that genuine loss still fails closed when the budget is also exceeded. (#1738) (Orleans.Lattice 9.4.1, Orleans.Lattice.Dashboards 9.4.1)

  • A DeleteTreeAsync / RecoverTreeAsync cycle over a partially purged tree no longer leaves a routable-but-unbound leaf that fails every typed CRDT write to its key range forever. ShardRootGrain.PurgeAsync clears a shard's leaves and internal nodes before it clears the shard root itself, so an interruption part-way through leaves the shard root's RootNodeId and RootIsLeaf intact while the nodes they address have had their state wiped - the reported incident was a PurgeTreeAsync call that exceeded the 30-second grain-call timeout on a tree holding roughly 42,500 vectors, whose out-of-band reminder purge was then abandoned, and purge is explicitly best-effort so this is a reachable state rather than a corruption. EnsureRootSlowAsync fast-paths on RootNodeId is not null and calls SetTreeIdAsync / SetShardIndexAsync only on the branch that creates a node, so a node's tree-id binding was established once at creation and never re-asserted; the recovered shard root therefore kept routing writes to a leaf with no bound tree id, and every typed CRDT write to that key range failed with LatticeCrdtShapeNotRegisteredException permanently. The field evidence matched that shape exactly: always the same leaf id rather than churn across leaves, recurring roughly every 90 seconds for over ten minutes, surviving a full cold container restart - so persisted damage, not an in-flight activation race - while the rest of the tree read fine. RecoverTreeAsync now asks every shard root to re-assert its node bindings, placed between clearing IsDeleted on the shards and persisting the recovered deletion state so that a failed repair leaves the tree still marked deleted and an operator retry stays clean rather than stranding a half-repaired topology behind a "cannot recover a tree that has not been deleted" precondition. The re-assert is near-free on a healthy tree: purge always clears the leftmost leaf first, so a leftmost leaf whose tree id is still bound proves nothing in that shard was cleared and the shard root returns after a single probe, needing no new persisted state and no new serializer field. Only a shard that is actually damaged pays a topology walk, which descends internal nodes rather than the leaf sibling chain - clearing a leaf wipes its sibling pointer, so walking the chain would stop at the first cleared leaf, whereas internal nodes keep their child ids until the internal sweep that only begins after every leaf is cleared - and re-asserts SetTreeIdAsync on each internal node and SetTreeIdAsync plus SetShardIndexAsync on each leaf that routing can still deliver to, bounded at 4,096 nodes per call at a fan-out of 16 so the repair cannot reproduce the timeout that caused the damage. Both setters are already idempotent, so re-asserting a healthy binding is a no-op. The probe is best-effort and logs a warning rather than throwing when the descent itself fails - a shard whose internal root was also cleared has nothing to descend, and recovery already succeeded in that case before this change, so throwing would make recovery newly fragile for no gain. The fail-closed guard that makes an unbound CRDT write throw is deliberately left intact rather than having the leaf resolve its own tree id lazily, because that would mask genuine routing faults and the current fault is a useful signal. Also recorded: TreeDeletionGrain.PurgeNowAsync, the synchronous path behind the public PurgeTreeAsync, never sets PurgeInProgress before its shard walk, which is why RecoverAsync's existing purge guards all passed on a tree the reminder-driven path would have refused to recover. A tree already damaged by an earlier build heals on its next delete/recover cycle, because the probe reads live state rather than a persisted flag. Six regression tests pin the contract, including end-to-end delete / interrupted purge / recover / CRDT write coverage over both a single-leaf root and an internal root, and a positive control that an undamaged tree stays bound. (#1744) (Orleans.Lattice 9.4.1)