Release 2026-08-25
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at 2026-08-25.md, and llms.txt lists every page.Part of the changelog.
A coordinated lockstep release advances the whole published package family to 9.3.0: every Orleans.Lattice and Orleans.Lattice.* package on the v9 line moves to 9.3.0 together. The minor bump is driven by two additive public primitives in the core library - a FIFO-fair distributed lock (ILatticeLockGrain, #1608) and a generic atomic-action / TCC saga coordinator (IAtomicActionGrain, #1609); the verified-WAL assurance work, the warm-authorization and batch / hot-path allocation optimisations, and a set of correctness and security fixes (two Entra Explorer credential-isolation fixes, an MCP telemetry allow-list-bypass fix, the OrSet duplicate-delivery convergence fix, the prefix-scoped backup bound fix, and the replication framing-overflow hardening) ride along. The not-yet-published Orleans.Lattice.Api.Mcp.RepoContext, Orleans.Lattice.Api.Mcp.RepoContext.Replication, and Orleans.Lattice.Storage.File packages remain unreleased and are unaffected.
Added
- A cluster-wide, FIFO-fair distributed lock / lease primitive. A new public grain
ILatticeLockGrain, keyed by lock name, packages the single-threaded-grain FIFO mutual-exclusion pattern as a first-class primitive that also gets the failure modes right:AcquireAsyncenqueues the caller FIFO and completes when the lock is granted (or faults with aTimeoutExceptionwhen the caller'sMaxWaitelapses), without ever blocking the grain's activation turn;TryAcquireAsyncis the non-blocking variant;RenewAsyncandReleaseAsyncare honoured only for the current holder's fencing token; andGetStatusAsyncreturns a diagnostic snapshot. Every grant carries a strictly-increasing, persisted, never-reused fencing token (the standard Kleppmann fencing guarantee), so a superseded holder is detectable by the resource it guards, and a bounded lease is reclaimed and handed to the next FIFO waiter if the holder neither renews nor releases before expiry, so a crashed holder cannot wedge the lock forever. The lock's fencing and admission decisions are extracted into a pure, deterministic core (LockAdmissionCore) that both the production grain and a Coyote concurrency model execute, so its safety properties - monotonic fencing, stale-token rejection, mutual exclusion, and expired-lease reclamation - are machine-checked against every adversarial interleaving (with a non-vacuous guard test), not just asserted by integration tests. Documented in Distributed Lock and Verified Distributed Lock. (#1608) (Orleans.Lattice9.3.0) - A public, generic atomic-action saga / TCC coordinator grain. A new public grain
IAtomicActionGrain, keyed by a caller-supplied operation id (its idempotency key), runs an ordered plan of steps all-or-nothing: each step pairs a forward effect with a compensating effect, and if a later forward step faults, every already-committed step is compensated in strict reverse order so the action leaves no partial effect behind. It generalizes the key-only atomic write to arbitrary caller-defined effects. A step never carries a serialized delegate; a custom step names a pre-registered handler by a stable id (registered at silo start viaAddLatticeAtomicAction, with a per-handler version tag) and carries a size-bounded, Orleans-serializable args payload, so a persisted plan is crash-recoverable and safe: resolving a handler id fails closed (an unknown or unregistered id never executes), and a crash-resume that finds a handler's version tag changed underneath it parks rather than replaying a changed effect. A fluentAtomicActionPlanBuildermixes custom.Step(handlerId, args)steps with a built-in.TreeWrite(treeId, ...)step whose forward write delegates to the existing atomic-write machinery (IAtomicWriteGrain) - inheriting the verified single-tree-atomic / cross-tree-2PC guarantee - and whose compensation the library synthesizes by capturing and restoring key pre-images. The saga persists its per-step status vector after every transition and resumes reminder-driven to a single terminal outcome (Committed,Compensated, or terminalCompensationFailed, which surfacesCompensationFailedExceptionfor operator intervention); re-issuing the same plan under a terminal operation id returns the memoized outcome without re-running effects. The step-sequencing and crash-resume decisions are extracted into a pure, deterministic core (AtomicActionPlanCore) that both the production grain and a Coyote concurrency model execute, so its all-or-nothing-or-compensated, reverse-order-once, and resume-exactly-once safety is machine-checked against every adversarial crash interleaving (with a non-vacuous guard test), not just asserted by integration tests. Documented in Atomic Action and Verified Atomic Action. (#1609) (Orleans.Lattice9.3.0) - The WAL's concurrency seams are now driven by verified cores and machine-checked for their durability-safety properties. Extending the proven-core pattern from the atomic-commit protocol to the write-ahead log, six WAL decision points are extracted into pure, deterministic cores that both the production grains and a Coyote verification layer execute, so each safety property is proven of the exact production logic rather than only exercised by integration tests: the shipping watermark (
WalShippingWatermark) never advances past a gap in the acked set; the GC per-entry trim predicate (WalGcTrimCore) trims only past the minimum acked cursor across all consumers, never stranding a laggard; the WAL cursor registry (InMemoryWalCursorRegistry) max-merges each consumer's cursor so a stale re-delivery never regresses it; the shard-move fence (WalMoveFenceCore) keeps the fence check atomic with the offset assignment so a quiesce can never fence an already-numbered append; and the commit-log writer's shutdown drain (WalAdmissionGateCore) releases every parked admission caller by observing the drain token in the wait set; and the per-shard offset allocator (WalOffsetAllocationCore) reads and advances the log-offset counter in a single step so two concurrent appends never share an offset and the assigned sequence stays dense and strictly ascending. A Coyote concurrency tier model-checks each core under adversarial interleavings, and every model ships a non-vacuous guard test that removes the one fix the property depends on and proves Coyote then finds the violation. This is an assurance change with no public API or runtime-behaviour change (the extractions are behaviour-preserving delegations; the existing WAL, GC, and cursor-registry tests are unchanged and pass); the verified guarantees are documented in Verified WAL and demonstrated by the VerifiedWalDurability sample. (Orleans.Lattice9.3.0) - The verified-WAL Coyote tier now also covers the buffer-pin blocked-floor lifecycle and the resumable placement-move tail copy. Two further WAL decision points are extracted into pure, deterministic cores that both the production path and a Coyote model execute: the blocked-floor meet (
WalBlockedFloorCore) folds each consumer's live buffer pin into the minimum across consumers, so through every interleaving of pin-take, pin-raise, and pin-clear the GC's floor never rises above a live pin and never trims an entry a buffering receiver still needs; and the placement-move resume arithmetic (WalMoveResumeCore) resumes an interrupted tail copy just past what the target already holds, so a coordinator that crashes and re-drives at any offset boundary copies every retained offset exactly once, with no duplicate and no gap. Each ships a Coyote model plus a non-vacuous guard test (a maximum-pin join that strands a lagging buffer; a fixed-floor resume that re-copies a landed prefix). This is an assurance change with no public API or runtime-behaviour change (both extractions are behaviour-preserving delegations fromInMemoryWalCursorRegistryandLatticeAdminGrain; the existing cursor-registry and WAL-move tests are unchanged and pass). Documented in Verified WAL. (Orleans.Lattice9.3.0)
Changed
Three time optimisations on the warm authorization decision hot path, with no behaviour, ordering, or public API change. Three independent, benchmark-verified changes to the synchronous, in-memory decision path the enforcement gate pays on every gated operation once the compiled policy snapshot is warm, none of which touches the public API: (1)
CompiledPolicynow indexes its governed trees in aFrozenDictionary<string, CompiledTree>rather than aDictionary<string, CompiledTree>held behind anIReadOnlyDictionaryfield, so the once-per-decisionTryGetTreelookup becomes a direct call on a read-optimised frozen map instead of an interface-dispatchedDictionary.TryGetValue; (2)CompiledTreelikewise indexes its exact-key rule buckets in aFrozenDictionary<string, CompiledRule[]>, so the most-specific (exact-key) tier lookup inResolvePointis a direct frozen lookup; and (3)CompiledTree.TryBestInBucketwalks each scope tier'sCompiledRule[]byref readonlyover aReadOnlySpan<CompiledRule>instead of aforeachthat copied every ~40-byteCompiledRulestruct into a loop variable per iteration, copying a rule only when a new best is found. Both compiled structures are built once at snapshot-compile time and read many times per decision, so freezing them is a pure read-side win; ordinal (case-sensitive) key comparison is preserved. A new cluster-freeAuthDecisionBenchmarksmicrobench (opt-in suiteBENCH_MICROBENCH_SUITE=authdecision) drivesPolicyEvaluator.Evaluatedirectly against a representative ten-tree default-deny snapshot; atfullfidelity the decisionMeanfalls on every scope tier - exact-key 34.13 ns to 28.70 ns (-15.9%), tree-wide 29.79 ns to 26.05 ns (-12.6%), prefix 37.73 ns to 35.04 ns (-7.1%) - and a mixed twelve-decision batch 442.25 ns to 338.07 ns (-23.6%); every path stays zero-allocation on the single-decision tiers. No public API change (all three types are internal in-process snapshot state that never crosses a grain boundary or serializes). (Orleans.Lattice.Auth9.3.0)Three allocation trims on the multi-key
GetManyAsync/SetManyAsyncbatch fan-out hot paths, with no behaviour, ordering, or wire-format change. Three independent,MemoryDiagnoser-verified optimisations on the batch read and write fan-out, none of which touches the public API: (1)LatticeGrain.GetManyAsyncCoretakes a single-shard fast path whenever every requested key routes to one physical shard - the only case for a single-shard tree, and the dominant case generally. It skips the shard-bucketingDictionary<int, List<string>>and its per-shard bucketList(routing the caller's key list straight through), the per-callConcurrentDictionarymerge target, the per-shard task fan-out, and the finalConcurrentDictionary->Dictionarycopy, issuing oneIShardRootGrain.GetManyAsyncunder the identical registry-snapshot scope + topology-version + snap2 stability checks and returning the shard's own result dictionary directly. This mirrors the shard grain's existing single-leaf fast path one layer down, so atomic visibility is unchanged. (2)ShardRootGrain.TraverseForBatchReadAsyncpresizes its multi-leaf merge dictionary to the requested key count (the returned count is bounded by it), eliminating the geometric grow/rehash chain on a batch that spans more than one leaf. (3)LatticeGrain.SetManyAsyncCoretakes the symmetric single-shard fast path on the bulk-write side: when every entry routes to one physical shard it skips the shard-bucketingDictionary<int, List<KeyValuePair<string, byte[]>>>and its whole-batch per-shardListcopy (routing the caller's entry list straight through), and the per-callList<Task>+Task.WhenAllwrapper, issuing one write to the shard directly; the caller's list is only read (never mutated) after the fan-out, so reusing it is safe. DeterministicMemoryDiagnoserAllocateddeltas on the microbenchmarks: single-shardPoint get many(exercises (1)) 9,056 B to 7,864 B (-13.2%), thePoint get many sweepbatch-size lanes (exercise (1)) at batch 64 16,512 B to 9,656 B (-41.5%), at batch 32 12,544 B to 8,704 B (-30.6%), at batch 16 10,368 B to 8,144 B (-21.5%); the new multi-leafPoint get many deep tree(exercises (1) and (2)) 58,744 B to 53,992 B (-8.1%, of which (2) contributes 912 B); andBulk load(exercises (3)) 102,840 B to 73,785 B (-28.3%) - every affected benchmark strictly lower, none regressed. No public API change (all three are internal implementation details; theILattice/IShardRootGrainsignatures and results are unchanged). (Orleans.Lattice9.3.0)Three steady-state allocation trims on the shard-root batch-write, batched CRDT receiver, and atomic-write saga hot paths, with no behaviour, ordering, or wire-format change. Three independent,
MemoryDiagnoser-verified optimisations, none of which touches the public API: (1)ShardRootGrain's batch-write rejection guard gains aList<KeyValuePair<string, byte[]>>overload that iterates entry keys by index instead of the priorThrowIfRejectedForAnyKey(entries.Select(e => e.Key))projection, so each guardedSetManyAsync/SetManyWherePredicateAsynccall (one per shard the write fans out to) no longer allocates aSelectListIteratorthat the steady-state (no split, no moved slots) guard early-returns without ever enumerating. (2)LatticeGrain.ApplyCrdtDeltaManyAsync's per-item receiver fold now enters a singlereadonly struct CrdtReceiverAmbientScopethat saves and restores the three ambient origin / vector-clock / HLC-overrideRequestContextslots on the stack, replacing three nestedLatticeOriginContext.With/LatticeVectorClockContext.With/LatticeHlcOverrideContext.Withscopes that each allocated a heapIDisposable- removing3Nscope objects per batch ofN. (3)AtomicWriteGrain's saga prepare derives its sorted touched-shard set directly from the shard-bucket dictionary keys it already builds, removing a separateHashSet<int>plus a full secondrouting.Map.Resolvepass over every key on every prepare. DeterministicMemoryDiagnoserAllocateddeltas from the new cluster-freeHotPathAllocationBenchmarksmicrobench (opt-in suiteBENCH_MICROBENCH_SUITE=hotpath), which isolates each technique because the full end-to-end cluster benchmarks allocate on the order of megabytes per op and sit the trims below their run-to-run noise floor: guard projection (exercises (1)) 80 B to 0 B per guarded call, size-independent; receiver ambient scope (exercises (2)) 112 B per item - atN=1618,816 B to 17,024 B, atN=6475,264 B to 68,096 B, atN=256301,056 B to 272,384 B (all -9.5%); atomic touched shards (exercises (3)) 568 B to 424 B at 1 key (-25.4%), 688 B to 544 B at 2 keys (-20.9%), 4,016 B to 3,840 B at 64 keys (-4.4%) - every affected benchmark strictly lower, none regressed. No public API change (all three are internal implementation details; theILattice,IShardRootGrain,IReplicationApplyGrain, andIAtomicWriteGrainsignatures and results are unchanged). (Orleans.Lattice9.3.0)
Fixed
The Entra Explorer login provider now binds silent token renewal to the account that actually signed in, preventing cross-operator credential confusion.
MsalEntraInteractiveTokenAcquirer.AcquireSilentAsyncrenewed withaccounts.FirstOrDefault(), an arbitrary cached account never correlated to the operator who opened the connection. Whenever the MSAL token cache held more than one account, a connection could silently renew with a different operator's token and issue cluster calls under the wrong identity.EntraTokenRequestgains an optionalUsername(additive; existing constructions are unaffected) that the auth method sets from the interactive sign-in result and threads into renewal, and the acquirer now selects the matching cached account, failing closed to a re-challenge when it is absent rather than grabbing a different one. Covered by new acquirer and auth-method regressions. (Orleans.Lattice.Explorer.Entra9.3.0)The Entra Explorer login provider is now registered per Blazor circuit rather than as a process-global singleton, restoring per-operator credential isolation.
AddExplorerEntraAuthregistered both the MSAL-backed token acquirer (which owns an in-memory token cache) and the Entra auth method as singletons, so every circuit shared one operator's token cache - the exact process-global auth-session leak the Explorer credential-isolation invariant forbids, and which the sibling web-head provider already avoids. Both are nowScoped, matching the scopedIExplorerAuthSessionthat consumes them (no captive dependency). This is a DI-lifetime change only; the public types andIEntraInteractiveTokenAcquirerinterface are unchanged. Covered by a new registration-lifetime regression. (Orleans.Lattice.Explorer.Entra9.3.0)The MCP telemetry deny-all metric-access gate now rejects an unconstrained label selector, closing an allow-list bypass that leaked denied metrics. In the
DenyAllExceptAllowedposture the PromQL authorization gate extracts the metric names a query references and admits it only when every name is allow-listed, failing closed when no name can be extracted (for example a bare{job="api"}label selector, which matches series across every metric name). That empty-name guard was bypassable: padding the bare selector with any one admitted metric -up or {job="api"}- left the extracted name set non-empty (["up"]), so the gate admitted the whole expression and the backend returned every series carryingjob="api", including denied metrics. The extractor now flags a top-level{...}selector that is neither anchored to a metric name in name position nor pinned by an exact__name__="..."matcher, and the gate fails closed on it, on both the instant-query and range-query paths. The anchored (up{job="api"}), exact-name ({__name__="up"}), and read-all paths are unchanged. Both the extractor type and its reference struct are internal, so there is no public API change. Covered by newPromQlMetricExtractorandTelemetryToolHandlersregressions. (Orleans.Lattice.Api.Mcp.Telemetry9.3.0)OrSet.MergeDeltais now idempotent under duplicate delivery. The observed-remove set's delta-merge add path appended each incoming add dot unconditionally (viaAdd) instead of unioning it, so replaying the sameOrSetDeltatwice - the norm under at-least-once delta replication and cross-tree stage replay - grew the per-element dot list without bound and broke convergence, contradicting the method's own documented idempotency contract. The add path now de-duplicates each(element, dot)pair exactly as the removes path already did, and a newOrSetMergeDeltaTestsfixture closes the coverage gap (every sibling CRDT already had a duplicate-delivery idempotency test;OrSetwas the sole exception). No public API change. (Orleans.Lattice9.3.0)Prefix-scoped backup and restore no longer silently skip keys whose prefix ends in
U+FFFF.BackupConstants.PrefixUpperBoundadvanced the prefix's final code unit unconditionally, so a trailingU+FFFFwrapped toU+0000, producing an exclusive upper bound that sorts below the prefix and inverting the half-open scan range - a prefix-scoped capture or restore then matched nothing - while an empty prefix threwIndexOutOfRangeException. It now rolls over trailing maximum code units, incrementing the last unit belowchar.MaxValueand dropping the max tail, and returns an unbounded (null) upper bound when none exists, matching the canonicalDataReader.PrefixUpperBound/RepoContextPortability.PrefixUpperBoundcontract. No public API change. (Orleans.Lattice.Backup9.3.0)The binary replication framing decoder rejects a forged field or entry length near
int.MaxValuewith its precise truncation error instead of a raw slice exception. Four bounds checks inOrleansBinaryReplicationBatchEncoder- the uncompressed and inflated-tail routing-string (treeName/originClusterId) readers and the per-entry body readers - summed the read cursor and an attacker-controllable declared length in 32-bit arithmetic, which overflows to a negative value and slips past the guard, the same overflow the compressed-body length check was already hardened against. All four now widen the sum to 64 bits, so a truncated or hostile payload fails closed with the descriptive framing-corruptionArgumentExceptionrather than an opaquespan.Slice/ArraySegmentexception. No public API change. (Orleans.Lattice.Replication9.3.0)The
azure-throughputbenchmark rig no longer mis-grades the Orleans 10.2.2 startup manifest-convergence burst as a cohort regression (benchmark tooling only; no library or package change). After the family-wide Orleans10.2.0->10.2.2bump (#1598), a fresh single silo converges its cluster-manifest / grain-directory view lazily per grain type: the first hot-path message addressed to a grain type the local silo has not yet resolved is rejected by the placement service with anOrleansException: No active nodes are compatible with grain <type> ... Known nodes with grain type: none, logged atfail:by theOrleans.Messagingcategory. Orleans re-addresses the waiting message once the manifest converges, so the caller's operation still succeeds (the cohort reaches a clean FINAL withfailed=0) and the burst lands in the pre-measurement warm window that thet>=15ssteady-state filter trims - steady-state throughput is unaffected. On the single-silo rig the affected types are the write-path grains the warm-up probe does not pre-activate (walmaterialiserpin,leafsnapshotstorage) plus lateshardrootactivations. The Layer 2 cohort-verdict grader counted these ~617 benign lines per cohort toward its DEGRADED exception tally, so every cohort graded DEGRADED and the report's HEALTHY-only aggregation excluded all of them (no HEALTHY cohorts ... row not updated), which looked like a total throughput collapse even though the measured steady-state means matched or slightly beat the prior baseline. The grader now subtracts current-cohort-attributable placement-convergence lines from the verdict-relevant exception count (mirroring the existing benign shutdown-race and warmup-retry exclusions), surfacing the excluded count as a diagnostic reason. The match is anchored on both theNo active nodes are compatible with grainseam and theKnown nodes with grain type: nonecold-manifest phenotype, so a genuine placement fault (nodes known but version-incompatible, or an unsatisfiable placement filter) still counts, and a convergence retry that ever failed for real would surfacefailed>0and still grade FAILED. Covered by newTest-CohortVerdict.ps1regressions. (#1605)