Table of Contents

Release 2026-09-09: Fixed

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at 2026-09-09-2.md, and llms.txt lists every page.

Part of Release 2026-09-09, in Changelog.

Fixed

  • Three more extreme-but-permitted interval and lease settings that silently inverted into their own opposite are now saturated, completing the tick-overflow sweep begun in #2221 and #2342. Each site added a configured span to a monotonic Stopwatch or tick clock and stored the result as a deadline, and each did the add unguarded, so a span large enough to carry the sum past long.MaxValue wrapped it to a past instant - turning "wait a very long time" into "the deadline has already passed" and firing the guarded work immediately and repeatedly, the exact opposite of what the operator asked for. (1) LeafWalkBudget, the per-turn bound on a shard scan-page and background-drain leaf walk, computed start + (long)(duration.TotalSeconds * Stopwatch.Frequency) from MaxScanPageDuration / BackgroundDrainMaxDuration, neither of which carries an upper bound; an extreme value wrapped the deadline negative so ShouldYield returned true on the very first leaf, truncating every page to a single leaf and defeating the practically-unbounded budget the value expressed. (2) WalShardGrain's placement-move fence computed the same product from the coordinator-supplied quiesce lease, so an extreme lease wrapped _fenceDeadlineTicks negative and ThrowIfMoveFenced read the fence as already expired, self-deactivating and re-resolving placement on the first fenced append instead of holding the shard still for the coordinator to copy a stable tail - defeating the quiesce. (3) RepoContextSelfIndexGrain's reconcile scheduler computed nowTicks + interval.Ticks + jitterTicks from ReconcileInterval (settable per second through LATTICE_RECONCILE_INTERVAL_SECONDS, with no upper bound), so an extreme interval wrapped NextReconcileAfterTicks negative and the tick's nowTicks >= NextReconcileAfterTicks gate then re-drove the reconcile - a re-walk for a mounted repository, an outbound git fetch for a git-sourced one - on every tick, the busiest possible cadence in place of the rarest one configured. All three now saturate at the ceiling instead of wrapping: the WAL drain helper is generalised from SaturatingDrainDeadlineTicks to SaturatingStopwatchDeadlineTicks and shared with the fence, LeafWalkBudget gains a private saturating deadline, and the reconcile scheduler gains a SaturatingAddTicks mirroring the one TxRegistryGrain already carries. Ordinary configurations compute an identical deadline and are unchanged; every affected type is internal, so no public surface changes. (Orleans.Lattice)

  • A registry that cannot determine a saga's outcome now says so, instead of answering "still in flight" and letting a reader see the pre-saga value it was told to hide. Three separate paths in TxRegistryGrain reported TxStatus.InFlight for a saga whose outcome was unknown to them rather than known to be undecided: the tombstone-retention mask on GetStatusAsync and GetStatusManyAsync, and both cross-tree delegation paths when the dial to the owning coordinator threw. InFlight is not a neutral reading. AtomicVisibilityGate.ResolveKey treats it as a licence to fall through to the pre-saga value, and the leaf's replay-time self-heal treats it as a reason to leave a prepare resident, so a single failed grain call surfaced superseded data as current, and an aged-out decision re-pinned the very projection checkpoint the self-heal exists to release. All three now answer Indeterminate, the case introduced for issue #2328, which the gate already hides. The distinction the retention path draws is exact rather than blanket: a txid with a stored decision row that is merely no longer reportable is Indeterminate, while a txid the registry has genuinely never seen stays InFlight, so the widening costs nothing for the overwhelmingly common unknown-txid read. The dial-failure paths are narrowed to the outer call: the inner catch that rolls back a failed cache write still returns the verdict it successfully obtained, because there the answer is in hand and only its caching failed. A masked-but-recorded decision would otherwise strand the leaf's self-heal permanently, so ITxRegistryGrain gains GetRecordedStatusAsync, a deliberately narrow bypass that reads the stored row past the retention mask and is called only from the activation-time sweep - never from a read path, which is the alternative issue #2318 explicitly rejects, since a read must not resurrect an outcome the registry has stopped vouching for. IsTombstoneExpiredAt is also made pin-aware, matching the prune path that has always skipped pinned tombstones, so a snapshot pin no longer holds a row in storage while the status read reports it masked; the pinned union is memoised on the earliest instant its answer could change and short-circuits entirely when no pins exist, so the reader hot path pays one count check in the common case. Making the mask pin-aware also makes the revision token a function of the pin set, because its live-expired term now counts only unpinned rows: a pin lapsing re-masks its rows and raises that term on its own, but a pin taking cover of already-masked rows lowers it, so the token is extended to stay sound in that direction too, as described in the tombstone-expiry entry below. Finally, CrossTreeInFlightObservation gains an additive UnresolvableCount so the backup fence can distinguish "sagas still preparing" from "sagas whose coordinators I could not reach"; the existing two-argument constructor is retained and chains to the new one, so source and wire compatibility both hold. (#2318) (Orleans.Lattice)

  • The registry's decision snapshot now says "I cannot currently determine this" instead of saying nothing, so an aged-out commit is no longer indistinguishable from a saga that never happened. A decision whose tombstone outlived TxDecisionRetention was dropped from the dictionary that SnapshotAsync returns, which folded two very different facts onto one reading: "no decision was ever recorded" and "a decision was recorded and the registry is no longer entitled to report it". Every same-process reader treats both identically, so the collapse was invisible in the core - but that dictionary is also the frozen registry view a cross-cluster bootstrap export is built from, and there it was load-bearing in the wrong direction. An aged-out Committed saga reached a bootstrapping peer as absence, was read as "still preparing", and could never be corrected, because the terminal that would have flipped it was already behind the incremental stream the receiver drains after the snapshot: the source held "committed", the receiver held "preparing", permanently, with no repair path on either side. TxStatus therefore gains a fourth case, Indeterminate, carried in place of the recorded outcome for exactly the rows the retention mask used to delete, so absence in a snapshot once again means only what it says. The read path is unchanged in effect and strictly safer in derivation: AtomicVisibilityGate.ResolveKey hides an indeterminate saga's prepared keys rather than falling through to the pre-saga value, because falling through is an affirmative claim that the saga did not commit - the one thing an indeterminate reading says nobody knows - and the branch is placed ahead of the already-terminal orphan rule so it cannot inherit that rule's fall-through. The bootstrap export now skips only genuinely decided sagas, so an indeterminate saga's prepared rows ship and the receiver holds exactly what the source holds. Three consumers that derived a decision by elimination are corrected in the same pass: the shard-split drain and the leaf's ambient-miss defensive path both tested for InFlight and treated everything else as decided, which under the widening would have silently read an indeterminate saga as an abort; and ShadowedMigrationReadGuard.ResolveSaga, which gates reads routed to a leaf still shadowed by an in-progress migration, tested for equality with Committed and passed everything else straight through to the migrated pre-saga value - for an indeterminate saga, the same affirmative "it did not commit" claim the visibility gate was changed to stop making. It now routes Indeterminate down the committed arm, which is strictly more available than gating unconditionally: once the terminal has landed on this leaf the projected entry is post-saga whichever way the saga went, so that read is served rather than deferred, and only the stale-routing case is gated. The new case is additive by value and no existing member is renumbered, so a mixed-version cluster stays wire-compatible; an older node that has not learned the case degrades to its existing default, which is the conservative reading. (#2328) (Orleans.Lattice, Orleans.Lattice.Replication)

  • A cross-tree transaction whose participating trees disagree on their configured cluster id is now rejected at admission, instead of proceeding on a premise nothing checked. Several arguments in the cross-tree protocol take the form "the guard verdict holds on T1, therefore the entry is safe on T2", and that step across the tree boundary is licensed only by ClusterId(T1) == ClusterId(T2). LatticeReplicationOptions.ClusterId is a per-tree named option, so nothing in the type system or the call graph makes the agreement hold; the options validator asserts it in prose and is structurally unable to check it, because a relation between two trees' configurations is not observable from the single options instance a validator is handed. The premise is now enforced at the two points where a participant set first exists and the relation therefore first becomes expressible: the authoring coordinator's admission of a cross-tree write, and the replicated cross-tree barrier's wait-set freeze on the receiver. Both fail loudly with the two disagreeing trees and their resolved ids named, before anything is staged, persisted, or dispatched, so a rejected saga leaves no state behind. Resolution goes through the existing ILatticeOriginClusterIdResolver seam rather than reading options directly, which keeps the core library free of a replication dependency and costs one cached lookup per participant. Existing deployments are unaffected: the check is on agreement, not on any particular value, so the core default resolver's uniform empty cluster id satisfies it for every tree - and that is the correct verdict rather than a vacuous pass, because a host with no replication configured genuinely has one cluster identity. Neither check re-runs on a resumed saga or a later terminal, because a participant set past admission already has writes staged in hidden buckets and must still be driven to a terminal decision; stranding it would be worse than the drift the check would report. A tree whose per-tree slice is empty is dropped from the participant set before the check, so a write that provably carries no verdict to or from it is not failed. The premise is also now stated on the option itself and on the validator that could not discharge it, including the accurate account of what holds uniformity up today - an empty default that fails the validator's own non-empty check, plus AddLatticeReplication registering cluster-wide - so that divergence requires a deliberate per-tree override rather than arriving by accident. (#2354) (Orleans.Lattice, Orleans.Lattice.Replication)

  • The premise that a cross-tree transaction id is never delegated to two coordinators at once is now enforced where it is created, instead of being assumed by five separate consumers and established by none of them. TxRegistryState carries two delegation maps - ExternalAuthorities for a saga this tree authored and ReceiverDecisionAuthorities for one it received - and every consumer that probes them relies on a txid appearing in at most one: ResolveAnyDelegatedAsync probes the authoring map first and returns whatever it finds, the terminal-decision paths drop from both without checking which was populated, and the backup fence's in-flight count sums them. A txid present in both would therefore have two coordinators for one decision, and the answer a reader received would depend on which map its call site happened to probe first - the two coordinators being under no obligation to agree, and nothing anywhere reconciling them. Tracing the five sites that read the premise as given showed each one deferring to the next, correctly, with the chain terminating in a subsumption rather than an enforcer. Both Register*DecisionAuthorityAsync methods now reject a registration that would produce coexistence, before any mutation, so a rejected call needs no unwind and leaves neither map, the registration epoch, nor the persisted state touched. The check is deliberately on the consequence (two rows for one txid) rather than on the cause (a terminal arriving stamped with the receiver's own cluster id): a self-origin comparison is not expressible in operands the core library holds, because DefaultLatticeOriginClusterIdResolver.Resolve returns the empty string on any host that has not registered replication, so such a check would pass vacuously in exactly the single-cluster deployment it was meant to protect. A missing operand stops you; a present-but-vacuous one lets you ship. Failing closed is safe for existing deployments because coexistence is currently unreachable: a replicated terminal's origin is re-stamped verbatim across every rehop, so no path produces the two registrations. InvalidOperationException is used rather than a new serializable exception type so no alias is added and no wire format changes. The premise is also now stated on both state fields, on both interface members, on ResolveAnyDelegatedAsync (whose docstring previously asserted disjointness with no warrant), and at the receiving cluster's local-origin guard in ReplicationApplier, which is that cluster's only enforcement rather than defence-in-depth behind the sender-side outbound filter it was documented as backing. The affected types are internal, so no public surface changes. (#2353) (Orleans.Lattice, Orleans.Lattice.Replication)

  • A saga decision that ages out of the registry's retention window no longer leaves a mid-flight multi-key reader holding a snapshot it will accept as authoritative. The registry's readable decision surface is Decisions masked at read time by ForgottenAt against the configured TxDecisionRetention, so a tombstone crossing that boundary removes a row from every reader's view because the clock advanced, with no write anywhere to hang a counter bump on. DecisionsRevision was a function of the Decisions map alone, so it could not announce that change - and the reader-side fast path in LatticeGrain.IsSnap2StableAsync short-circuits on revision equality and never reaches the IsSnapshotStable rule that might otherwise have caught it. The probe therefore returned "nothing moved" across a window in which a saga's outcome had become unobservable, and the reader's snap1 dictionary was reused as if still authoritative. The revision is now a composite comparison token: the persisted decisions counter, plus the number of tombstones currently past their retention boundary, plus a new persisted TombstoneRetirementEpoch counting tombstones already retired from that live set. The third term is not decoration - without it the sum can fall, because ForgetAsync advances the decisions counter once per batch while its inline prune removes k rows from the live-expired term, letting the token revisit a value it previously carried under a different surface; the same holds when a terminal clears an already-expired tombstone. Both retirement paths now raise the epoch and both unwind paths lower it, so the token is non-decreasing across every sequence. A fourth term, the persisted TombstonePinUnmaskEpoch, closes the same hole on the pin axis. Making the retention mask pin-aware makes the live-expired term a function of SnapshotPins as well as ForgottenAt, and a snapshot pin taking cover of already-masked rows lowers it with no compensating write - the same falling-token aliasing reached by a different route. PinSnapshotAsync therefore counts the rows the pin is about to un-mask and raises the epoch by that count plus one: the count restores monotonicity, and the extra unit makes the token move strictly across a pin that genuinely changed the readable surface, which compensating by the count alone would not. A pin that un-masks nothing leaves the token exactly where it was, so opening a cursor over live decisions does not spuriously invalidate every concurrent reader. Because that term is pin-derived, invalidating the pinned-union cache now also drops the live-expired memo: the memo's validity horizon is clamped to the first pin lapse and so covers the clock-driven direction, but an explicit pin, unpin, or refresh moves the surface with no clock advance for the horizon to catch, and a stale count there would hold the token still across a real change. A second, adversarial defect is fixed alongside it: the tombstone-clearing prologue in MarkCommittedAsync / MarkAbortedAsync ran before the write-once terminal guard, so a duplicate same-outcome terminal was classified Record rather than Idempotent - resurrecting a forgotten decision, restarting its retention window, and bumping the revision for a surface change that never happened. Classification now runs first: a same-outcome repeat on a tombstoned txid is genuinely inert, while a conflicting outcome still clears the tombstone, preserving the "a fresh terminal after a forget is a new authoritative outcome" semantic exactly. The live-expired scan sits on the [AlwaysInterleave] reader hot path, so it is memoised on the earliest instant at which its answer could change, keyed on the retention window it was computed against; the cache is derived from persisted state and in-memory only, so losing it can never change an answer. GetDecisionsRevisionAsync deliberately keeps its long return type - folding the composite into the existing token rather than widening the signature - so a mixed-version cluster mid-rolling-upgrade continues to interoperate; the new state field takes the next free [Id(10)] and decodes to 0 on legacy state, which is exactly its value on a registry that has never retired a tombstone; the pin term takes [Id(11)] on the same reasoning. Three docstrings that asserted an unconditional "repeated same-outcome calls are no-ops" idempotence guarantee are corrected to state its real bound (it holds only while the decision is still recorded) and to name the ordering argument that carries the local saga path, which the replication apply path does not share. (#2332) (Orleans.Lattice)

  • A failed persist on a saga's terminal decision no longer strands the cross-tree delegation maps ahead of disk, which had turned a still-preparing cross-tree saga into a confidently wrong "in-flight" answer and let a backup capture window certify itself as quiescent. TxRegistryGrain.MarkCommittedAsync and MarkAbortedAsync mutate four in-memory maps before their WriteStateAsync: the tombstone-clearing prologue removes a ForgottenAt row and the retained Decisions row beneath it, an unconditional drop removes the txid's rows from both ExternalAuthorities and ReceiverDecisionAuthorities, and the decision core then records the outcome and bumps the revision. Their catch arms restored only the last pair. A persist that threw therefore left the two delegation maps, and any cleared tombstone, permanently ahead of the persisted state for the remainder of the activation - and those rows are not bookkeeping. The delegation row is the only coordinator pointer GetStatusAsync holds, so losing it makes an unknown txid fall through to InFlight: not a degraded answer but a confident wrong one, indistinguishable at the call site from a saga that is genuinely still preparing. Four consumers read the difference. The snapshot provider ships a committed cross-tree saga as prepared; the backup drain gate starts a capture early; the backup post-capture fence certifies a window as stable; and ResolveAnyDelegatedAsync has nothing left to resolve against. Both catches now restore every map the call mutated, following the save-and-restore-or-remove precedent already used by RegisterExternalDecisionAuthorityAsync, and the delegation drop moves from above the write-once terminal guard to below it. That reordering is the structural half of the fix: above the guard, the Idempotent and Conflict exits mutated both maps and then returned without ever reaching a persist at all, so the maps were changed on a path that writes nothing and has no catch to unwind. Below it, every path that reaches the drop also reaches the write, so the drop is either persisted or unwound. The unwind deliberately restores the rows rather than bumping CrossTreeRegistrationEpoch: the backup fence compares the epoch as a delta across the capture window but the in-flight count absolutely, so a compensating epoch bump would be absorbed into the baseline and observed by nothing, while a restored row is observed by the clause that actually matters. That asymmetry is now documented at the fence, along with the population each of its two clauses covers - the epoch clause covers exactly the sagas that registered during the window and says nothing about those that registered before it, which is the drain gate's job. Ordinary (non-throwing) operation is bit-identical; the affected types are internal, so no public surface changes. (#2352) (Orleans.Lattice, Orleans.Lattice.Backup)

  • Three more unbounded TimeSpan options that are added directly to a tick counter now saturate instead of silently overflowing to a past instant, completing the sweep begun in #2221. Each of these options is validated only against a lower bound, so an extreme (multi-millennia) value makes now.Ticks + option.Ticks wrap negative, and the resulting "past" deadline inverts the option's intent rather than extending it. (1) The materialised-view shadow swap computed the swapped-out generation's reclaim-eligibility instant as DateTime.UtcNow.Ticks + OldGenerationReclaimGrace.Ticks; on overflow the gate now < ReclaimEligibleAtTicks reads false immediately, so an extreme grace reclaims (deletes) the old-generation tree at once - while it may still be serving the very readers the grace exists to protect - instead of deferring reclaim. (2) The WAL shard deactivation drain computed its Stopwatch-tick deadline as startTicks + (long)(WalDrainBudget.TotalSeconds * Stopwatch.Frequency); on overflow the cast yields a negative deadline, so the drain loop force-faults every in-flight flush slot on its first check instead of granting the configured (effectively unbounded) grace. (3) The replication secret cache computed each entry's expiry as now + SecretRefreshInterval.Ticks; on overflow every entry reads as already expired, so the cache is defeated and the remote secret source is hit on every call - the opposite of the interval's load-shedding purpose. All three now clamp at their domain ceiling (DateTime.MaxValue.Ticks, long.MaxValue, and DateTimeOffset.MaxValue.UtcTicks respectively) through a small saturating helper, mirroring LockAdmissionCore.SaturatingLeaseExpiry and LatticeView.ComputeCacheExpiry from #2221. Ordinary configurations are bit-identical; the helpers are internal, so no public surface changes. Each helper carries a red-before / green-after regression test, and the secret-cache fix additionally asserts end-to-end that an extreme interval keeps the entry cached rather than re-fetching on every call. (Orleans.Lattice, Orleans.Lattice.Replication)

  • A shadow-cutover or cold restore of an incremental backup chain no longer corrupts the restored tree. LatticeBackupRestoreService.BuildShadowCoreAsync populated the shadow tree with BulkLoadRawAsync unconditionally, but that bottom-up B+ tree loader requires globally ascending, duplicate-free per-shard input: it chunks the entry list into leaves and takes each leaf separator from sortedEntries[i].Key. A backup chain streams base-first then each increment in WAL-drain order, so a chain carrying any increment is non-monotonic at the manifest boundary and repeats every key an increment rewrote. Loaded across more than one leaf this produced non-increasing leaf separators, so the internal routing mis-directed reads - an overwritten key resolved to the stale leaf and an untouched key became unreachable - silently returning wrong values or losing data, with no exception raised. The in-place restore path already gated this by chain shape (bulk-load only for a single full manifest, otherwise the HLC-converging merge path); the three shadow-cutover reachability paths - local ShadowCutover, the coordinated BuildShadowAsync seam, and LatticeBackupColdRestoreService - all funnel through BuildShadowCoreAsync, which lacked the gate. It now mirrors the in-place gate, taking the bulk-load fast path only for a single full manifest and otherwise routing through MergeApplyAsync, which reconciles duplicate keys by HLC and converges correctly into the always-empty shadow tree. The affected types are internal, so no public surface changes. (#2283) (Orleans.Lattice.Backup)

  • ShardMap.GetPhysicalShardIndices() no longer throws IndexOutOfRangeException on a negative slot value. The method dedups slot values through a bool bitmap sized to max + 1 and indexed directly by the slot value, but its first pass computed only max and never guarded against a negative slot, so bitmap[v] with v < 0 indexed the span out of bounds and threw. ShardMap is a public type whose Slots property is publicly settable and carried on the Orleans wire format, so a negative slot is reachable both through the public API and through a corrupt or hand-built map deserialised off the wire. The first pass now also tracks min and the bitmap fast path is guarded on min >= 0; a negative slot falls through to the general HashSet fallback that already exists in the method, which dedups and sorts arbitrary integers faithfully (returning, for example, [-1, 0]). Production maps are always non-negative - CreateDefault emits i % physicalShardCount and split logic emits small non-negative indices - so the change is inert for every valid input and only converts the crash on invalid input into a faithful result. (#2284) (Orleans.Lattice)

  • The shared metrics feed no longer coalesces two subscribers with genuinely different tree scopes onto one sampling loop when their tree-id lists alias under a comma join. SharedMetricsSampler funnels every subscription with a matching request signature onto one shared SamplerLoop, and that loop samples using only the first attacher's request, so two requests that build the same signature share one loop and the later subscriber silently receives the first request's tree set. BuildSignature rendered the requested tree-id set with a bare string.Join(',', ...), which is not injective: one tree literally named a,b and two trees a and b both render a,b, so with all other request fields equal (and, by default, visibility disabled so the signature is identity-free) the two requests wrongly coalesce and the second subscriber is served metrics for a tree set it never asked for. The sibling identity component already length-prefixes group ids and claims against exactly this aliasing class (#971); the tree-id component now uses the same AppendLengthPrefixed framing, so distinct tree sets can no longer collide while genuinely identical sets still coalesce and the empty-set signature is unchanged. The affected type is internal, so no public surface changes. (#2288) (Orleans.Lattice.Api.State)

  • A leaf whose replay exceeds its per-activation budget now banks fresh snapshot coverage on the activation that discovers the shortfall, instead of rolling its own persisted checkpoint backward on every rehydrate forever. The rehydrate path clears the whole cache and then loads every WAL partition from the leaf snapshot unconditionally, so a partition whose persisted per-partition checkpoint had advanced past the snapshot's coverage was rolled back to that coverage, reopening the replay gap the previous activation had just closed. This is a regression of #1819 through a mechanism its fix does not cover: the incremental checkpoint flush added by #1831 genuinely does advance the durable grain state, and the rehydrate then throws that progress away. Measured live on the deployed repository-context container, where a single leaf on tree repo-context-vector-membership held two durable rows that disagreed in the same second: its grain state had been written 4,504 times and carried checkpoint 160,972, while its snapshot row had been written 3 times in two days and carried 155,852 - and 155,852, the number absent from the store the checkpoint accessor reads, is the one the fault warning printed. A per-partition coverage-deficit capture already existed in MaybeRunPeriodicSnapshotRecheckAsync, but it ran only on the 64-checkpoint cadence, whose counter resets on every activation, so an over-budget leaf torn down before it persists 64 checkpoints never reached that capture and re-armed the identical rollback indefinitely. A one-shot off-cadence deficit capture now runs at most once per activation: it is latched during rehydrate, fires only when a rehydrate genuinely lowered a partition, is gated on the same no-loss precondition the deactivate-time capture already trusts, re-checks the deficit across the partitions before it acts, and banks strictly higher coverage. What it banks is the offset that replay actually re-applied, never the known-durable grain-state frontier, so a teardown-truncated activation converges monotonically across activations rather than in a single jump - a bare max(snapshot, persisted) would have claimed coverage the cache does not hold, which is silent loss, and a second discriminator in the fixture pins the distinction by asserting the re-applied offset and not the frontier. The escape sits above the cadence gate rather than below it, because LatticeOptions.LeafSnapshotReClassifyEveryNCheckpoints is documented to govern the periodic re-classification only: placed below the threshold <= 0 early return, setting that tuning knob to 0 would have left a frozen leaf with no escape at all, its WAL pin never lifted and its WAL growing without bound, which is a consequence out of all proportion to a tuning value. Nothing needed hoisting for that placement, so every other property of the block is unchanged - the single-flight guard, the no-loss precondition, the deficit re-check, clearing the latch before awaiting the capture, and the latch-retire fall-through - and a placement guard asserts the escape still banks fresh coverage with the cadence disabled, verified as a discriminator by reinstating the early return above it. The option's own documentation is corrected to name both activation-scoped captures. The affected types are internal, so no public surface changes. (#2220) (Orleans.Lattice)

  • A leaf that has just committed a checkpoint advance to durable storage no longer loses that activation to a digest publish that faults or parks, which was a second route into the same never-converging leaf. FlushPendingCheckpointAsync completes its durable PersistAsync and then runs an inline notification tail - the cursor report, the upward digest publish, and the snapshot recheck - with nothing between the durable write and that tail to stop a failure in it propagating out of the flush. An activation that had already banked its checkpoint advance was therefore destroyed by work running after its point of no return, and the next activation re-read the identical range, looping cold replay. The post-persist tail is now contained so that no failure after the durable write escapes the flush. Each failure is logged rather than swallowed, so a genuine upward cascade stays observable, and the digest stays dirty so convergence re-drives out of band on the next mutation. Separately, the leaf's outbound OnChildDigestPublishedAsync call is now bounded by LatticeOptions.DigestPublishTimeout, mirroring BPlusInternalGrain.PublishUpwardAsync: that publish is a cross-grain RPC recursing toward the shard root, and a parent that is itself mid-mutation can leave the await neither completing nor faulting until Orleans' fixed response timeout, consuming the leaf's whole activation budget for one notification. On the deadline the publish is abandoned, its eventual completion harmlessly unobserved, and a TimeoutException naming the leaf, the tree, the parent and the option is thrown for the caller to record and re-drive; a fresh one-shot cancellation source per publish keeps this reentrancy-trivial, because the leaf's turns are serialised so there is never a second publish in flight to bound, no pooled source to disarm, and no fired-in-the-race-window hazard to reason about. Timeout.InfiniteTimeSpan restores the historical unbounded await. The regression fixture is control-armed and asserts the flushed branch by its durable side effect - the persisted checkpoint having advanced - rather than by observing the publish, which is the thing being contained. DigestPublishTimeout is a pre-existing public option and the affected types are internal, so no public surface changes. (#2220) (Orleans.Lattice)

  • A leaf whose durable unresolved-replay-work ledger has saturated can now still record a resident saga prepare, so it stops freezing at an unchanged checkpoint and losing every write routed to it. The ledger introduced for #2165 is bounded by LatticeOptions.MaxDurableUnresolvedReplayWork, and at the cap the next resident prepare could not be recorded at all. It fell back to the in-memory _pendingTxOffsets clamp, which pins the incremental flush ceiling one below the prepare offset permanently, so the leaf re-entered replay from an unchanged persisted checkpoint on every activation, banked zero durable forward progress, and lost every write routed to it. The asymmetry that makes this different from the deferred terminals the cap was written for is that a deferred terminal dropped at the cap self-heals in pass 2 when it is re-read, whereas an unresolved prepare has no terminal to drain, so dropping it is a permanent freeze; dropping it is also unsafe for the aged-out-commit reason the self-terminalisation fix below preserves InFlight prepares for. A resident prepare is therefore recorded unconditionally whenever the ledger is enabled, and the cap continues to bound only the deferred terminals that can be safely re-read. With the ledger disabled at a cap of 0 the in-memory clamp remains the sole protection, unchanged. A two-arm regression pins it: the fix arm advances the persisted checkpoint strictly across successive interrupted activations and converges on the head, while the control arm with the fix disabled pins the checkpoint one below the prepare forever. Because recording past the cap is a correctness requirement rather than a free choice, it ships with its own signal: a new orleans.lattice.leaf.unresolved_prepare_ledger_beyond_cap counter tagged by tree and WAL partition, paired with a once-per-activation warning, fires when the row grows past the cap. That guard exists for a provider-dependent hazard rather than for the profile this repository runs, and it is observability only - nothing here caps or drops a prepare, because a behavioural cap would reintroduce the exact drop-and-freeze defect this removes. On the default local durability profile the row is backed by SQLite and the growth is write amplification rather than a correctness problem; on an Orleans.Lattice.Storage.AzureTable deployment the 1 MB entity cap makes an unbounded row a genuine persist hazard for which the operator previously had no signal before the write failed. It is deliberately left off every bundled dashboard and recorded as intentionally unpaneled, because on the local profile those dashboards target it sits flat at zero and a panel would show every default-deployment operator a permanently empty graph; the metric-to-panel map carries it instead, with the Azure Table alerting note. This adds one public member, LatticeMetrics.LeafUnresolvedPrepareLedgerBeyondCap. (#2183) (Orleans.Lattice, Orleans.Lattice.Dashboards)

  • A leaf holding a saga prepare whose terminal never arrived now resolves it against the transaction registry during replay, instead of pinning its projection checkpoint, and its WAL, forever. A saga prepare can stay resident in a leaf's pending-transaction map after the saga has been terminally decided in the registry but the terminal never reached the leaf. Once the durable unresolved-replay ledger is at capacity that resident prepare clamps the incremental flush ceiling one below its offset, the projection checkpoint cannot advance, the checkpoint pins the coverage-gated WAL collection, and every subsequent activation re-reads the same prefix and banks no forward progress. Nothing removes a resident prepare but the two terminal paths, ApplyTxCommit and ApplyTxAbort, which by hypothesis never fire, so the pin is permanent rather than slow. During activation replay, after the deferred terminals drain and before the final checkpoint reconciliation, each still-resident prepare is now resolved against the per-tree registry and any terminal decision is applied locally through that same ordinary path. The clamp then lifts as a consequence of the effect landing, so no acknowledged write is dropped - relaxing the clamp directly is the obvious alternative and is precisely the wrong one, because the clamp is what protects the unapplied effect. An undecided prepare is left resident and still clamps, including a decision aged out of retention, which reads back as InFlight and is therefore indistinguishable from one still in flight. The per-transaction registry resolution is contained: a timeout or fault leaves the prepare resident with its clamp intact and defers the heal to a later activation, rather than failing activation and having the host retry into the same timeout under exactly the load that produces the pin. Only the resolution is contained - the ApplyTxCommit and ApplyTxAbort effects still propagate, and cooperative cancellation is never swallowed. A two-arm regression control with two discriminators and three guards covers committed and aborted self-terminalisation, the undecided-clamp safety property, and the registry-failure containment. The affected types are internal, so no public surface changes. (#2190) (Orleans.Lattice)

  • A quiet tree no longer deadlocks its own activation against the autonomic loops that activation exists to arm, expiring both at the 30-second response deadline and arming neither. LatticeGrain.OnActivateAsync arms two autonomic loops by awaiting EnsureRunningAsync on the hot-shard monitor and on the shard-healing orchestrator, which is the activation-time arming introduced so that a tree which has stopped taking writes can still heal. Each loop, mid-pass, calls the tree's ILattice.Is*CompleteAsync status verbs. LatticeGrain is a [StatelessWorker], so on a quiet, idle-collected tree that status call must birth a fresh LatticeGrain activation, whose OnActivateAsync arms the same loop by awaiting the same EnsureRunningAsync. Neither grain is reentrant, so the loop, busy inside its pass, cannot admit the arming call and the activation cannot complete; both expire at the response deadline. Only quiet trees cycle, which is exactly the field signature: a busy tree serves the status verb from a resident stateless-worker activation with no birth at all, and warns zero times. Both EnsureRunningAsync verbs are now [AlwaysInterleave], so the idempotent, synchronous arming no-op is admitted while a pass holds the turn. That is safe because each implementation runs its _running guard and its _running = true set with no await between them, so at most one caller ever runs the body, and a call admitted mid-pass observes _running == true and returns synchronously having mutated nothing. Both heads are fixed together rather than only the one observed in the field: the healing head is latent there only because a structurally settled tree short-circuits its sweep before reaching the status verbs, and fixing the monitor alone would have made the healing head more reachable, not less. The two implementations differ deliberately and both remain no-ops under interleaving - the monitor sets _running before its enabled check, so a disabled monitor latches and the existing arming behaviour is preserved, while the orchestrator checks its switch between guard and set and does not latch. The activation-time try/catch that keeps arming from failing activation is unchanged. A two-head discriminator-and-guard integration test ships with it: each discriminator fails without the attribute and passes with it, and the resident-tree guards, which observe no birth, pass on both arms as positive controls. The affected types are internal, so no public surface changes. (#2218) (Orleans.Lattice)

  • A configured LatticeOptions property can no longer be silently dropped on its way to a tree, and six options that were inert in every deployment now take effect. LatticeOptionsResolver built ResolvedLatticeOptions from a hand-maintained per-property initializer, so any LatticeOptions property missing from that list was ignored and the compiled default ran in its place. The failure is silent in the worst direction: the host accepts the configuration and logs it, and it then has no effect, so it presents to whoever set it as "the setting did not help" rather than "the setting did not apply". Six options were live instances of the defect - the four admission gauges MaxLiveKeys, MaxEstimatedBytes, AdmissionAdvisoryLiveKeys and AdmissionAdvisoryBytes, and two WAL replay bounds: WalMaterialiserMaxConcurrentReplays, whose 0 default sized the concurrency gate to ProcessorCount regardless of what was configured, and WalReplayMaxRecordsPerTurn, whose configured per-turn budget was silently replaced by the 256 default. Any past tuning attempt that set one of these was inert, and should be discarded rather than merely revisited. Fixing only the six would have left the defect factory running, because the seventh dropped option arrives with the next option anyone adds, so the copy list is eliminated instead: ResolvedLatticeOptions.CopyConfigurableBaseOptionsFrom reflectively copies every public read/write instance property declared on LatticeOptions through a cached property list, and the resolver applies its small set of derived and registry-pinned overrides after that copy so those transformed values still win. The three structural pins MaxLeafKeys, MaxInternalChildren and ShardCount remain required init-only members on the derived type and are excluded from the copy by construction. The deliverable that keeps it fixed is the guard rather than the copy: the reflection test that audits propagation now covers the full surface with both of its exemption allowlists removed, and its sentinel picker learned to mint distinct instances for the interface and delegate options (ILatticeRetryPolicy and the WAL storage provider factory), so a type it cannot sentinel is now a loud failure instead of a silent skip - which is the property the allowlists had quietly traded away, and the reason six inert options could be documented as known-inert and stay that way. Verified as a discriminator by temporarily excluding MaxLiveKeys from the reflective copy, which fails the guard by name and lists the option's real consumer sites, and passes again once restored. LatticeOptionsResolver and ResolvedLatticeOptions are internal, so no public surface changes; the six options were already public and are now honoured. (#2182, #2203) (Orleans.Lattice)

  • A leaf on a partition that never wins the replay drain slot now banks its checkpoint, instead of re-reading the same WAL range on every activation forever. Issue #2089 narrowed this replay livelock by letting the last-swept partition drain its deferred terminals in place, and its own comment recorded the residual it left behind. That residual is the common case rather than a tail: exactly one partition per activation is drain-eligible, so with the deployed eight pin buckets seven of every eight are ineligible in any activation. On such a partition an unresolved saga prepare or an undrained deferred terminal clamps the incremental flush ceiling to offset - 1, TryFlushRecoveredCeilingAsync declines to flush because the clamped ceiling is not above the current checkpoint, and the activation banks nothing - so the next activation re-reads the identical range and repeats. Observed in production as one leaf sitting at a single distinct checkpoint across 69 activations spanning 6h51m52s, checkpoint delta zero. The drain slot is deliberately not widened, because the cross-partition dependency behind it holds: when absorbing a partition that is not swept last, a TxCommit in it may have prepares in an unabsorbed partition that are not yet in _pendingTx, and a DeleteRange may target keys not yet in the cache, neither of which is resolvable at that moment. The defect is that pass 2 is unreachable inside the activation budget, not that its precondition is too strict, so the ceiling is instead made independent of reaching pass 2. LeafNodeState gains a bounded list of unresolved replay work - partition, offset and mutation per record - carried in the same state row as the checkpoint, so the single WriteStateAsync that persists the checkpoint persists the record atomically with it and there is no window in which a checkpoint is durable and the work it skipped is not. Deferred terminals and applied prepares are recorded durably, the minimum-unresolved-prepare-offset lookup skips ledgered offsets in one edit that covers both clamp sites, activation rebuilds _pendingTx and the deferred terminal list from state rather than from the WAL, and records are struck off when their terminal drains or their transaction resolves - the strike-off deliberately running before the empty-offset-map early-out, so a terminal replaying in a later activation still clears the ledger. Records above the effective checkpoint are dropped because replay will re-read them; the invariant is that the ledger covers exactly the offsets this replay will not. Two tests were classified before running: against unmodified source both fail, with the observed checkpoint sequence 1,1,1,1,1,1,1,1,1,1,1,1 - twelve activations, one distinct checkpoint, delta zero, the production shape - and both pass with the change. This adds public surface: LatticeOptions.MaxDurableUnresolvedReplayWork bounds the list and is on by default at 1 024, with DefaultMaxDurableUnresolvedReplayWork alongside it, and at the cap recording fails so the pre-existing clamping behaviour applies - the degradation is to the previous build rather than to data loss. The option was also missing from LatticeOptionsResolver's explicit field-by-field copy, without which it would always have taken its compiled default and been inert as a per-tree knob, configurable in name only. The new state record and its wire alias are internal. (#2165) (Orleans.Lattice)

  • Shard healing is now armed when a tree activates, not only when it is written, so a tree that has stopped taking writes can still heal. Every EnsureMonitorAsync call site sat on a write path, which made arming a function of traffic shape rather than of the tree existing: a tree that a bulk ingest over-split and that then served only reads never armed the healing orchestrator, so the trees most in need of healing were precisely the ones that never got it. LatticeGrain now implements IGrainBase and arms from OnActivateAsync, activation being the one seam a new entry point cannot forget to annotate - which is how the gap arose in the first place, since the arming call was added to the write verbs and the read verbs were never revisited. The eight operation-path call sites are kept and are not redundant: both helpers return without latching their flag when arming loses the race with reminder-service startup, so the repeated path is the retry that recovers from it, and once armed it costs a bool test rather than a call. Arming never fails activation, and the catch at that seam is deliberately wider than the reminder-transient filter the helpers apply to themselves, because arming reaches grain storage and the reminder table - both transiently unavailable during silo start, which is exactly when activations happen. Propagating would take the tree offline for reads and writes, trading a missing background loop for an outage of the data it exists to maintain; nothing is lost, because neither helper latches on failure so the next operation re-attempts, and the operation-path catch stays narrow so a non-transient failure still surfaces to the next writer with its original shape. The AutoSplitEnabled asymmetry - the hot-shard monitor short-circuits on it, healing deliberately does not - and the system-tree short-circuit are both untouched. The affected types are internal, so no public surface changes. (#1877) (Orleans.Lattice)

  • Shard healing is now armed even when the hot-shard monitor faults, closing a third route to an already-shattered tree that never heals. EnsureMonitorAsync awaited the two autonomic loops in sequence, so any exception from the monitor arming that was not the reminder-service transient propagated before the healing arming was ever reached. The severity comes from composing that with the write-gated arming corrected above: all eight operation-path arming call sites are on write paths and no read path arms, so on a tree that has stopped taking writes activation is the only arming opportunity, and the usual "neither helper latches its flag, so the next operation re-attempts" mitigation is unavailable because there is no next operation. A single transient monitor fault was therefore enough to leave such a tree unhealed until it next deactivated, which is the same outcome that the two preceding healing fixes exist to prevent, reached by a third route. A monitor arming fault is now captured, the healing arming attempted regardless, and the captured fault rethrown with its original stack trace, so the write path still surfaces it exactly as before rather than having it swallowed. When both loops fault the monitor fault keeps precedence, because that is what a caller sees today when healing is never reached, and the displaced healing fault is logged rather than dropped; a try/finally would have been shorter but lets the second fault replace the first, changing the exception a caller sees. Both helpers are unmodified, so the AutoSplitEnabled asymmetry, the system-tree short-circuit and the filtered catches are preserved by construction, and a deferral returns normally without entering the fault path. Control 3 failed / 125 passed against treatment 0 failed / 128 passed. The affected types are internal, so no public surface changes. (#2187) (Orleans.Lattice)

  • Three absolute-tick expiry computations that overflowed for an extreme-but-permitted option value now saturate at DateTime.MaxValue, so a maximal-but-valid duration retains "nearly forever" instead of silently flipping to "expire immediately" or throwing on every operation. Each site validated its option only as strictly positive (no upper bound) and then formed now + window in absolute ticks, so a large enough value wrapped long to a negative tick. (1) LockAdmissionCore.Grant/Renew computed nowTicks + leaseTicks, which a large MaxLockLeaseDuration overflows to a negative LeaseExpiresAtTicks that IsLeaseExpired reads as an already-expired fresh lease - a mutual-exclusion violation - and that the distributed-lock grain then cannot materialise as a DateTimeOffset. (2) HistoryRetentionShaper.Shape computed drainNowTicks + Window.Ticks for a durable-history row, so a huge HistoryRetentionWindow stamped a negative expiry that the next TTL sweep treats as already-expired and drops, losing the very history the operator configured the window to keep longest. (3) LatticeView computed DateTime.UtcNow + ReadHandleCacheTtl with the checked DateTime + TimeSpan operator, so a huge ReadHandleCacheTtl threw ArgumentOutOfRangeException while resolving the active tree on every view read. All three now saturate at the maximum representable tick. Every affected member is internal, so no public surface changes. (Orleans.Lattice)

  • A leaf's same-silo revision cookie is now unique over the leaf's lifetime and is published on the activation path, closing two unbounded stale-read windows in LeafCacheGrain. The cookie was a per-activation bump count: the registry entry is removed on deactivation and re-seeded at zero on the next activation, so values were reused across activations and a cache still holding a stamp from an earlier activation could compare EQUAL to a different activation's state and take RefreshAsync's "provably fresh" early return - staleness on the revision branch with no TTL to bound it. Each activation's counter is now seeded from a process-wide monotonic ticket shifted left by 24 bits, and deactivation raises that ticket source past the counter value the activation actually reached, so uniqueness holds unconditionally instead of resting on the shift width as a spacing hint: a bump is one per state-advancing operation, so a leaf sustaining a few thousand writes a second reaches 2^24 in under an hour, and the workload that makes the overrun reachable is the same one that makes the adjacent-ticket collision land. Separately, WAL replay and snapshot rehydrate rebuild the in-memory projection without passing through any bumping site, so a leaf returning from a projection rebuild left the registry with no entry at all; RefreshAsync takes its revision branch only when an entry is present, so a co-located cache fell through to the TTL gate and kept serving a pre-replay snapshot for the whole TTL window while the leaf itself held the rows. The activation path now publishes unconditionally once the projection is rebuilt, which is safe only on top of the ticket seed - publishing while activations still restarted from zero would arm the collision by making the entry present with a value a previous activation had already published. Every affected member is internal, so no public surface changes. (#2166) (Orleans.Lattice)

  • A B+ leaf whose range shrinks to nothing is now reclaimed, and seven windows in which a concurrent write could be acknowledged and then lost are closed. A leaf left holding an empty range after its keys moved away stayed activated, chain-linked and routable forever, so a long-lived tree accumulated leaves that served no keys, and the reclaim pass that should have folded them away did not exist. Adding that pass surfaced the harder half of the problem: reclaim is the first operation in this codebase that WIDENS an existing leaf's range, which falsified assumptions elsewhere. The blocking defects were a probe-to-clear window in which a write acknowledged against a leaf being folded was destroyed, a routing entry retired before the predecessor had widened to cover the vacated span (so keys in that span resolved to nothing), and a chain walk bounded by folds but not by steps, which could spin. Two further defects came out of the tests written to prove those fixes: LeafCacheGrain's SplitKey prune assumed ranges only ever narrow, leaving a folded leaf publishing a boundary inside its own span, and the ABA guard added for the routing fix changed the revision cookie's meaning from unique-within-an-activation to unique-over-the-leaf's-lifetime, which the consumer's pre-TTL equality check shows is the correct meaning. The substance of the fix is the fixture that pins it: fourteen interleaved write-versus-reclaim concurrency tests, a split-versus-fold race test, and two falsification pairs each verified red against a real pre-fix build and green against the fix. The root cause of the original item returning was that no test drove a write concurrently with a reclaim pass. Every affected type is internal, so no public surface changes. (#2099) (Orleans.Lattice)

  • Leaf reclaim now declines to unlink the sibling of an in-flight split, instead of unlinking a leaf that is about to receive rows. SplitAsync persists the split intent and the new NextSibling pointer in one atomic block, so the new sibling S becomes chain-reachable the instant the intent lands, but CompleteSplitAsync seeds S's key range BEFORE awaiting the row merge. In that window S has a declared range, zero rows, a SplitState of Unsplit and no moved-away slots, so every reclaim-blocking condition evaluated ON S is legitimately false, and the predecessor's compare-and-swap passes because the observed next pointer really is S. Reclaim therefore unlinked S, and the merge then landed the split's rows on a leaf no walk could reach: acknowledged writes lost, with nothing thrown, because the retirement latch is a bare instance field that a later deactivation clears. The evidence that S is dangerous exists only on the splitting predecessor, which is where the guard now lives: TryUnlinkSuccessorAsync declines when the caller is mid-split and the successor it is being asked to unlink is that split's own sibling. The two declinations cover opposite orderings - the existing compare-and-swap catches a split that lands after the walk read the chain, and the new guard catches one that lands before - and the second conjunct is load-bearing rather than defensive, because on the split ABORT path the next pointer is restored while the split state is still in progress, and it is that conjunct which still permits ordinary reclaim of the original successor. Both callers release the retirement latch on a declination, so a decline cannot livelock the pass. The affected types are internal, so no public surface changes. (#2160) (Orleans.Lattice)

  • The over-budget leaf replay warning no longer scales its log volume with the size of the tree. The warning was throttled per throttle key, and a tree with L leaves and P WAL partitions has L x P of them, so aggregate volume grew with the tree even though each individual key was well behaved. In one deployment this single warning reached 46% of the container log and rolled away the evidence needed to diagnose the condition producing it. A per-tree, per-window cap on REPEAT detail lines now sits on top of the existing per-key throttle. What is deliberately not traded away is the signal: a leaf partition reporting over budget for the FIRST time is exempt from the cap, so a genuinely novel condition still surfaces promptly, everything the cap withholds is reported as a summary line rather than dropped silently, and the orleans.lattice.leaf.activation_replays_over_budget counter remains an exact census regardless of what is logged. The gate's own state is committed only when a line is really emitted, so a capped repeat does not pay backoff for a line it never printed; the stamp sweep retires rather than deletes, so housekeeping cannot manufacture false novelty; a clean in-budget activation retires the backoff; and both capacity sweeps are rate-limited. The affected types are internal, so no public surface changes. (#2100) (Orleans.Lattice)

  • The over-budget leaf-replay warning no longer fires for leaves that are one to two orders of magnitude under budget, and a genuinely stuck leaf is now reported as a distinct fault rather than left to be recovered by hand from the noise. LatticeFallOffLogDetector compared walHead[p] - checkpoint[p] - the extent of a whole WAL partition, shared by every leaf pinned to it, measured before the per-leaf range filter - against MaxLeafReplayEntries, which is documented and consumed as a per-leaf, post-filter budget on the entries a leaf replays through ILeafProjection.Apply. Those are different quantities in different units, and the comment directly above the test already contained the argument that condemns it: the apparent gap "overstates the actual work by the full WAL contribution of every sibling leaf in the same shard partition". That reasoning is fully general, but the guard derived from it exempted only the -1 sentinel, leaving the identical overstatement uncorrected for every checkpoint at or above zero. Measured on a repository-context deployment carrying ~1,350 leaves per partition, real per-leaf work was order 10^2 against a 10^4 budget while the warning fired continuously: 19,639 lines in 6.26 hours, of which 19,604 were benign. The comparison, not the threshold, was wrong - raising the budget would have suppressed the symptom and kept the defect. The classifier's cheap pre-check is retained and re-cast as what it soundly is: because every entry a leaf applies lies within (checkpoint, head], applied <= gap always holds, so a gap within budget proves the leaf is under budget and is dismissed for free, while a gap over budget proves nothing and now only nominates a candidate. Confirming it in the classifier would require reading (checkpoint, head] before the replay reads it again, doubling the most expensive part of activation - and that read is precisely the one whose overrun constitutes the livelock this warning exists to expose - so the verdict is taken during the replay that happens anyway, by counting the entries that actually pass ShouldApplyDuringReplay, at a cost of one increment per entry. The warning and the orleans.lattice.leaf.activation_replays_over_budget counter are emitted on that exact count, so the counter is now in the same units as the budget it is named after. Separately, and load-bearing: the criterion that a checkpoint failing to advance across repeats "is a fault, not a slow replay" was stated in the message and left for an operator to evaluate across a log - the evaluation that found a leaf livelocked for 6h52m at an unchanging checkpoint losing ~15.8 writes/hour, and the evaluation that stops working when the lines that matter are outnumbered 560:1. It is now performed in-process: a leaf re-entering replay for the same partition from an unchanged persisted checkpoint raises a distinct stalled-replay warning on an independent throttle, so cost noise can never suppress a fault, and the first activation stays silent because one cold activation is not a stall. On the measured deployment the two changes together yield zero lines for the 19,604 benign leaves and about 68 for the frozen one. Both warnings keep the per-leaf throttle and leaf qualification, and the cost line now also carries the head and gap it compared against - it previously carried neither, which is why the over-budget factor was not readable from the instrument at all and a dimensionless figure obtained by dividing an absolute WAL offset by the budget survived across three issues before measurement caught it. The affected types are internal, so no public surface changes. (#2149) (Orleans.Lattice, Orleans.Lattice.Dashboards)

  • A floating-point index range query against zero no longer silently drops a grain whose indexed value is the opposite-signed zero. GrainIndexRangeBuilder routed only == and != against a zero literal through the signed-zero handler that treats -0.0 and +0.0 as one equivalence class; every relational operator against zero fell through to the ordinary path, which keys off the literal's own slot. Because -0.0 and +0.0 compare equal yet occupy adjacent, distinct slots in the order-preserving key encoding, x.Score >= 0.0 built a range that started at the +0.0 slot and, being an exact range, dropped the residual predicate, so a stored -0.0 that satisfies >= 0.0 was excluded with no error. The mirror cases > -0.0 (which wrongly included +0.0) and <= -0.0 (which wrongly dropped +0.0) were wrong the same way. All six operators against a zero literal now share the "zero band" spanning both slots, so a comparison includes or excludes both zeros together. The affected type is internal, so no public surface changes. (Orleans.Lattice.GrainIndex)

  • Restoring a backup chain that passes through an empty increment no longer fails with a spurious missing-artifact error. An increment captured with no intervening writes is a first-class chained manifest carrying one content descriptor with ChunkCount == 0 that streams zero chunks. LatticeBackupRestoreService.ValidateManifestAsync treated "streamed zero chunks" as "artifact absent from the sink" and threw LatticeRestoreValidationException before the integrity check, so a valid chain ending in or passing through an empty increment could not be restored, even though the empty artifact hashes to SHA-256 of the empty input and validates. The presence check now fires only when the descriptor claims chunks (ChunkCount > 0); an empty descriptor falls through to the digest comparison, which validates it. The affected type is internal, so no public surface changes. (Orleans.Lattice.Backup)

  • A valid empty backup is no longer misreported as an unresolvable orphan and pruned from the catalog. InClusterLatticeBackupSink.ProbeAsync classified a content descriptor as a missing artifact whenever no chunk rows existed for it, but an empty increment or empty full backup legitimately writes no chunk rows. The empty artifact was therefore reported missing, which made BackupSinkResolution.IsResolvable false, so LatticeBackupCatalogScrubService pruned the manifest as an orphan and LatticeBackupHealthService flagged it degraded, silently discarding a good restore point during a routine sweep and independent of any restore attempt. The probe now skips the presence check for a descriptor with ChunkCount == 0, so an empty artifact resolves instead of registering as missing. The affected type is internal, so no public surface changes. (Orleans.Lattice.Backup)

Published releases, newest first. Each section is keyed by its publish date; within a date, packages advance on their own patch digits per docs/RELEASING.md.