---
title: "Release 2026-09-02: Breaking to Changed - Changelog"
url: "https://nsta1.github.io/Orleans.Lattice/changelog/2026-09-02-1.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/CHANGELOG.md?plain=1#L1307-L1354"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
---
# Release 2026-09-02: Breaking to Changed

Part of [Release 2026-09-02](../changelog/2026-09-02.md), in [Changelog](../CHANGELOG.md).

## Breaking

- **A cluster-wide `Telemetry` grant is now authorized over the all-trees sentinel `"*"` rather than the reserved auth policy tree, so an operator who authored the grant against the old scope must re-author it.** This is the fix for ([#1795](https://github.com/NSTA1/Orleans.Lattice/issues/1795)) and is described in full under Fixed. The blast radius is deliberately small, which is why it ships in a minor rather than waiting for the next major: the **documented** authoring path, `LatticeScope.ClusterWide()`, produced a `"*"` scope that the authorizers never consulted, so a grant written the documented way was inert and no conforming deployment can regress. Only a deployment that discovered the mismatch and worked around it by granting `Telemetry` on the reserved `sys-auth-policy` tree is affected. No public API signature changed, and the constants involved are internal, so this is a behavioural change with no compile break. The failure direction is fail-closed: an affected caller is **denied** cluster telemetry rather than over-granted, so the risk is the availability of one telemetry surface, never an escalation. **Migration:** re-author the grant with `LatticeScope.ClusterWide()` (equivalently, scope `Tree:*`) for `LatticeOperation.Telemetry`; the reserved-tree grant can then be withdrawn. (`Orleans.Lattice.Auth` 9.5.0, `Orleans.Lattice.Api.TreeAdmin` 9.5.0)

## Added

- **You can now apply many CRDT deltas in one call, instead of paying a full round trip per key.** `ILattice.ApplyCrdtDeltaManyAsync` is the CRDT counterpart of `SetManyAsync`: it fans out to shards in parallel and collapses each leaf's slice into a single write-ahead-log dispatch, so an N-key batch costs one commit-log round trip per leaf rather than N. Marking presence or membership flags - a loop of `OrFlag(key).EnableAsync(...)` - was the motivating workload, and it now has a typed one-liner, `EnableManyAsync`, which reads every current flag in one batched call and mints all the deltas from that snapshot, halving the round trips again. The merge mode is declared once for the batch rather than per entry, because a tree resolves exactly one CRDT shape; mixing shapes inside a tree converges locally but diverges at replication, so the API does not offer that footgun. The batch is **not atomic**, like `SetManyAsync` - but because every entry folds a delta rather than overwriting a value, a retry converges instead of clobbering a concurrent writer. When the batch must land all-or-nothing, `StageEnableManyAsync` mints the staging tokens from the same single read and `LatticeAtomicWriteBuilder.SetMany` couples them into one existing cross-tree saga - no new transactional machinery, so the all-or-nothing and cross-cluster guarantees are unchanged. ([#1921](https://github.com/NSTA1/Orleans.Lattice/issues/1921)) (`Orleans.Lattice` 9.5.0)

- **Adaptive shard split gained its missing inverse, leaf snapshots gained a compact binary encoding, and a cold leaf no longer pays an all-or-nothing load.** Five defects and two missing capabilities on the core tree's activation and maintenance paths are addressed together. **Shard consolidation** is the missing inverse of adaptive split: a split was previously a one-way door, so a tree that split under transient load stayed split forever. Consolidation folds an over-split tree back down fully online, and a new healing orchestrator observes tree shape and heals automatically rather than waiting to be invoked. **Split admission now requires load to be genuinely skewed rather than merely high**, so a bulk ingest no longer shatters a tree while a truly hot shard still splits exactly as before (`HotShardMinSkewRatio`, `HotShardMinShardEntries`, and a `MaxPhysicalShardsPerTree` ceiling). **Leaf snapshots gained a compact binary codec** in place of the previous JSON-plus-base64 encoding, with transparent dual-read and lazy rewrite, so blobs already persisted stay readable indefinitely and are re-encoded on their next natural capture with no migration step (`LeafSnapshotBinaryEncodingEnabled`). **Leaf activation now hydrates incrementally**, so its cost follows what is actually read rather than how large the leaf happens to be, instead of loading the entire snapshot up front (`LeafPartialHydrationEnabled`, `LeafHydrationResidentBytes`). Two durability defects are fixed with them: write-ahead-log garbage collection deferred its first pass by up to an hour, so a box recreated more often than that never trimmed at all - the delay is now a short startup window with a floor between passes (`WalGcStartupDelay`, `WalGcMinInterval`); and a single deferred mutation pinned the incremental checkpoint flush for an entire replay, so a long replay made no durable progress and re-ran the identical window on its next activation. Every mechanism is default-on with its own documented kill switch, so an operator who hits trouble disables one behaviour rather than reverting the image, and the consolidation and healing surfaces are internal - the new configuration knobs on `LatticeOptions` are the whole of the public surface. New instruments on the `orleans.lattice` meter report split admission, consolidation progress and healing decisions, charted by new panels on the bundled Overview and Replication dashboards. ([#1830](https://github.com/NSTA1/Orleans.Lattice/issues/1830)) (`Orleans.Lattice` 9.5.0, `Orleans.Lattice.Dashboards` 9.5.0)

- **You can now authenticate lattice callers against any standards-conformant OpenID Connect provider - Okta, Auth0, Keycloak, Ping, Google - without writing an authenticator.** The new `Orleans.Lattice.Membership.Oidc` companion package adds `siloBuilder.AddLatticeOidc(...)`, a generic bearer-credential authenticator driven entirely by the provider's discovery document. Point it at an `Authority` (or an explicit `MetadataAddress`), name the exact `Issuer` you trust and the `Audiences` you accept, and it discovers the JWKS, rotates signing keys on its own refresh schedule with last-known-good fallback, and maps `sub` to the lattice subject and `groups` / `roles` / `role` to asserted groups - each configurable. It is an additive sibling to the Entra authenticator rather than a replacement: register both, plus the JWT and anonymous authenticators, on the same silo, and call `AddLatticeOidc(...)` once per issuer. Selection is unambiguous by construction because it is an exact ordinal issuer match (or an explicit `SchemeHint`), never a prefix, wildcard, or catch-all, so an authenticator never claims another provider's token regardless of registration order. Three fail-closed guards are load-bearing: signature algorithms are pinned from the discovery document's `id_token_signing_alg_values_supported` (or from an explicit `Algorithms` list, which wins), and a provider that advertises none rejects every token rather than accepting any - so the algorithm-confusion gap an unpinned allow-list leaves open is closed, including the classic case of an `HS256` forgery signed with the provider's published public key; an empty `Audiences` list throws at construction instead of silently accepting any validly-signed token minted for a different relying party; and a token whose subject collides with a reserved sentinel resolves to the anonymous subject rather than to an anonymous-labelled principal that still carries the token's groups. The issuer is also pinned to the configured value rather than widened to whatever the discovery document advertises. Because authenticator selection reads the token's issuer, and the credential bridges stamp the scheme from operator configuration rather than from the caller, that selection parse runs on a pre-authentication path; it is bounded by the validating handler's own maximum token size, so an unauthenticated caller cannot use an oversized credential to amplify work once per registered authenticator, and no credential that could have authenticated is ever declined. ([#1804](https://github.com/NSTA1/Orleans.Lattice/issues/1804)) (`Orleans.Lattice.Membership.Oidc` 9.5.0)

- **A backup sink that is not actually shared across your clusters is now caught when a silo starts, instead of when a coordinated restore fails.** A coordinated restore of a replicated tree only works if every cluster can read the *same* backup, but until now a deployment where each cluster pointed at its own isolated sink looked perfectly healthy right up until the restore saga aborted all-or-nothing - by which point you had been capturing backups you could never restore. Each cluster now writes a tiny self-cleaning canary marker into its configured sink at startup and checks whether it can read every peer's marker back, which is a real reachability proof rather than the previous assumption that any external sink must be shared. The verdict is three-valued so a cold start cannot cry wolf: seeing a peer's marker is `Shared`, not seeing it while that peer is also unreachable on the control channel is `Unverified` and stays quiet, and not seeing it while the peer *is* reachable is `NotShared` - the genuine misconfiguration. Detection defaults to a loud startup warning (`LatticeBackupOptions.SinkSharingEnforcement`, default `Warn`), so upgrading cannot brick a deployment that is merely transiently unreachable; set `FailFast` to refuse to start on a refuted sink, or `Disabled` to skip the probe entirely - though `Disabled` deliberately does not relax the pre-existing hard failure on an in-cluster sink. The same signal now also feeds the periodic backup health sweep, so a backup that verifies locally but cannot be seen by a peer reports `Warning` with the unreachable cluster named, and the problem surfaces in the Explorer Backups tab long before anyone attempts a restore. This closes the blind spot in the old startup guard, which only rejected the in-cluster sink and assumed every external sink was shared. Single-cluster and non-replicated deployments are untouched: with no replicated trees or no peers the probe never runs, writes nothing, and reports exactly what it did before. ([#1669](https://github.com/NSTA1/Orleans.Lattice/issues/1669)) (`Orleans.Lattice.Backup` 9.5.0, `Orleans.Lattice.Replication` 9.5.0)

- **An Orleans grain's typed state can now be indexed and queried, without a hand-maintained secondary index.** The new `Orleans.Lattice.GrainIndex` companion package lets you declare that a grain's persistent state should be tracked in a lattice tree, then ask typed predicate questions over it - "which `User` grains are 18 or over?" - and stream back grain references, keys, or matches. An index is opt-in twice: the silo declares it with `AddGrainIndex<TGrain, TState>(...)`, naming each projected property explicitly with `Include(u => u.Age)` so write amplification is always a deliberate choice, and the grain marks its state with the `[Indexed]` facet. Grains enrol on activation and on every state write, while a resumable, rate-limited, reminder-driven backfill onboards the dormant population from an application-supplied `IGrainKeySource`; both routes write through one projection path, so they converge on a single duplicate-free index. Queries reuse the core predicate dialect, so filtering runs in the tree shards: a single-property comparison becomes a contiguous range scan over an order-preserving key encoding, a conjunction becomes one scan per property whose grain keys are intersected, and a disjunction unions and de-duplicates its branches. Entry updates for a grain are all-or-nothing, a failed index write is surfaced to the caller *and* retried from a durable outbox marker that carries the whole plan (so a repair never has to wake the grain), index trees live under a reserved `__grainindex/` namespace and are cluster-local unless you opt in, and a declaration change that would invalidate already-written entries is rejected at startup rather than silently answering wrong. Eight instruments on the shared `orleans.lattice` meter, a Grafana dashboard, and an `IGrainIndexAdmin` surface (status, progress, pause/resume/rebuild, single-pass) report and control it. ([#657](https://github.com/NSTA1/Orleans.Lattice/issues/657)) (`Orleans.Lattice.GrainIndex` 9.5.0)

- **A silo restart can now pre-warm the leaf caches a shard was actually reading, instead of serving its first reads cold.** Set `LatticeOptions.LeafCachePreWarmCount` to a non-zero value and each shard root keeps a small bounded histogram of the leaf `GrainId`s its read traversals land on, persists a compact snapshot of it into its own durable state, and - on the next `ILattice.WarmUpAsync` after reactivation - primes that many `LeafCacheGrain` activations for the leaves it read most. Ranking by observed read frequency rather than by recency means a leaf touched once just before shutdown does not displace a leaf that was being read constantly, which is the failure mode a recency list has on the skewed and cyclic access patterns a B+ tree produces under a real key distribution: on held-out synthetic traces (train on the first half, score against the hot set of the second) frequency recovered 96% of the true hot set on a Zipf-skewed trace and 98% on a cyclic one, against 56% and 53% for a recency list of the same size. Because the pre-warm is driven from the shard root, the `[StatelessWorker]` cache activations land on the silo hosting that shard root - the same silo that will serve the subsequent reads - so no silo-lifecycle infrastructure is involved. The feature is **off by default** (`0` disables it), the model is bounded to 256 tracked leaves resident and 64 persisted (a snapshot under ~3 KB), the read hot path costs two predictable branches and no allocation when disabled, the snapshot is written on a coalescing timer and at deactivation rather than per read, and a leaf that fails to warm is logged and skipped rather than failing `WarmUpAsync`. Three new metrics (`orleans.lattice.leaf_access.model.leaves`, `orleans.lattice.warmup.leaf_cache.prewarmed`, `orleans.lattice.warmup.leaf_cache.duration`) and a Grafana panel report what the model holds and what each warm-up primed. ([#332](https://github.com/NSTA1/Orleans.Lattice/issues/332)) (`Orleans.Lattice` 9.5.0)

- **A troubleshooting guide now walks you from a symptom to a fix.** `docs/lattice/troubleshooting.md` shows how to read a `DiagnoseAsync` report field by field - including a worked example and the traps that lead to a wrong conclusion, such as a shallow report always reporting zero tombstones and an all-zero shard entry meaning the diagnostics fan-out failed - and then covers storage-provider exceptions on write, concurrent shard-split activity, slow scans, and stale reads, each as symptom, likely cause, how to confirm, and how to fix. ([#342](https://github.com/NSTA1/Orleans.Lattice/issues/342))

- **A migration guide for importing an existing Redis, relational, or Cosmos DB dataset into a tree.** [Migrating from an External Store](../docs/lattice/external-store-migration.md) is the operational playbook that sits on top of the bulk-loading reference: it compares the three ingest paths (the one-shot `ILattice.BulkLoadAsync`, the streaming `IAsyncEnumerable` overload, and the resumable tree-administration Begin/Append/Commit session) against the ordering guarantee each one asks you to own, then documents the contract a migration actually has to satisfy - which paths enforce an empty target and which silently do not, ordinal key comparison, what `BulkLoadOrderException` does and does not catch, and where the path is idempotent rather than transactional. Sections on key design and value encoding cover the choices that cannot be revised after the load (zero-padding numeric segments, separator selection, why a key prefix is a logical range and not a placement hint, choosing an `ILatticeSerializer<T>`, and planning schema evolution up front), and per-store playbooks for Redis, relational databases, and Cosmos DB cover stable resumable enumeration and each store's ordering pitfalls. It closes with how to verify what landed before cutting over. ([#341](https://github.com/NSTA1/Orleans.Lattice/issues/341))

- **You can now see exactly how much write latency each `IMutationObserver` is costing you.** Observers run inline on the grain write path, so a slow one is paid for by every caller - but until now the only signal was the aggregate `observer` step on `orleans.lattice.leaf.commit.duration`, which says a cost was incurred without saying by whom. A new `orleans.lattice.observer.duration` histogram (milliseconds) on the existing `orleans.lattice` meter times each observer callback individually and tags it `observer` (the observer's CLR type name) and `tree`, so a misbehaving observer doing synchronous I/O or blocking on a downstream call is named on the same OpenTelemetry pipeline as the traffic it slows down. The measurement is taken on the faulting path too, so an observer that throws slowly is as visible as one that returns slowly, and the dispatcher continues to swallow and log the exception. The zero-cost-when-unused guarantee is unchanged: the dispatcher's no-observer fast path returns before any timing work, and a registered observer with no metrics listener attached costs one boolean read per publish. The sample spans only the callback, so the warning the dispatcher logs for a faulting observer is not billed to the observer that faulted - a slow log sink can no longer read as a slow observer. Charted by a new by-observer p50/p95 panel on the bundled Commit Path dashboard. A new `DashboardJsonTests` drift guard also fails the build on duplicate Grafana panel ids, which is how two concurrently-authored panels silently collide on the conventional `max(id) + 1`; one such pre-existing collision on the Overview dashboard is corrected with it. No public API changed. ([#430](https://github.com/NSTA1/Orleans.Lattice/issues/430)) (`Orleans.Lattice` 9.5.0, `Orleans.Lattice.Dashboards` 9.5.0)

## Changed

- **Three more warm read paths now presize their result collection to a tight in-hand bound, trimming redundant heap from leaf range-entry reads, tree-structure topology scans, and the causal-apply drain, with no change to behaviour or public API.** All three edits preserve output byte-for-byte and touch only internal/private method bodies: (1) `BPlusLeafGrain.GetEntriesAsync` grew its result list from empty while scanning the leaf's ordered cache rows - the single unpresized method among its already-presized read siblings; it now presizes to `Math.Min(Cache.Count, 256)`, the exact tight upper bound the sibling `GetKeysAsync` already uses (capped at 256 so a range/split-truncated scan never over-allocates against a full leaf). (2) `LatticeStateQuery.GetTreeStructureAsync` grew its top-level `rootNodes` list from empty across the scanned shard span while every nested topology list was already presized; it now presizes to that span (`endShard - startShard`, one root summary per shard - the budget-capped upper bound), which required only reordering the two lines that compute the span. (3) `CausalApplyBuffer.DrainSatisfied` grew its drained `ready` list from empty while walking the parked-entry list; it now presizes to the parked-entry count read under the buffer's own lock (the drain yields at most one record per parked entry), which collapses to a zero-capacity list on the empty steady-state buffer so the common in-order delivery path is unaffected. Measured on the new `readpathpresize` microbench suite (`BENCH_MICROBENCH_SUITE=readpathpresize`, `MemoryDiagnoser`, in-process; each lane rebuilds the production shape per call so the `Allocated` delta is exactly the removed heap): leaf range-entry read (512 cached rows) 16.16 KB -> 12.08 KB (-25%, the removed list regrowth); tree-structure scan (64 shards) 3.12 KB -> 2.55 KB (-18%, the removed list regrowth); causal-apply drain (128 parked entries) 2.14 KB -> 1.05 KB (-51%, the removed list regrowth). The `BPlusLeafGrain` suite in `Orleans.Lattice.Tests`, the `Orleans.Lattice.Api.State.Tests` suite, and the `CausalApplyBuffer` suite in `Orleans.Lattice.Replication.Tests` cover the preserved read and drain semantics. (`Orleans.Lattice` 9.5.0, `Orleans.Lattice.Api.State` 9.5.0, `Orleans.Lattice.Replication` 9.5.0)

- **Three warm read paths now presize (or stop copying) their result collection, trimming redundant heap from view-catalog listing, live-entry reads, and snapshot-shard scans, with no change to behaviour or public API.** All three edits preserve output byte-for-byte and touch only internal/private method bodies: (1) `LatticeStateQuery.ListViewsAsync` built its authenticated `visible` list already presized to the candidate count, then copied it into a throwaway array via `visible.ToArray()` on every authenticated request; it now hands the list back as an `IReadOnlyList<ViewListing>` with no copy and iterates it index-based (not `foreach`) so reading through the interface does not box `List<T>`'s struct enumerator - a hazard that would otherwise offset the saved array on small lists. (2) `BPlusLeafGrain.GetLiveEntriesAsync` grew its result dictionary from empty while folding every live cached row, reallocating its bucket and entry arrays; it now presizes to the cached row count (the fold's upper bound), exactly as the sibling `GetLiveRawEntriesAsync` already did. (3) `SnapshotLeafGrain.GetKeysAsync` / `GetEntriesAsync` / `GetRawEntriesAsync` grew their result list from empty while scanning the pinned folder's ordered rows; they now presize to `Math.Min(limit, entryCount)`, a tight upper bound on the emitted count that stays small on a narrow page. Measured on the new `readpathtrims` microbench suite (`BENCH_MICROBENCH_SUITE=readpathtrims`, `MemoryDiagnoser`, in-process; each lane rebuilds the production shape per call so the `Allocated` delta is exactly the removed heap): view-catalog listing (256 visible views) 4.08 KB -> 2.05 KB (-50%, the removed array copy) and 1.56 us -> 1.11 us; live-entry read (512 cached rows) 46.97 KB -> 14.38 KB (-69%, the removed dictionary regrowth) and 12.31 us -> 6.50 us; snapshot-shard scan (512 rows) 16.16 KB -> 8.05 KB (-50%, the removed list regrowth) and 4.42 us -> 3.13 us. The `Orleans.Lattice.Api.State.Tests` suite covers the view-listing behaviour, and the `SnapshotLeafGrain`, `BPlusLeafGrain`, and `LatticeCursorGrain` suites in `Orleans.Lattice.Tests` cover the preserved read semantics. (`Orleans.Lattice` 9.5.0, `Orleans.Lattice.Api.State` 9.5.0)

- **The grain-index query executor and the state-API metrics observer now trim three redundant heap allocations from their warm query-intersect and metrics-delta paths, with no change to behaviour or public API.** All three edits preserve output byte-for-byte and touch only internal/private method bodies: (1) `GrainIndexQueryExecutor.IntersectAsync` grew its per-clause `survivors` dictionary from empty as it folded each later clause's scan into a multi-clause AND query, reallocating its bucket and entry arrays; because every survivor is a key already present in the driving `candidates` set, that set's current count is a tight upper bound, so it now presizes `survivors` to `candidates.Count`, removing the grow-from-empty rehash churn on each intersect pass. (2) `LatticeStateMetricsObserver.ObserveAsync` grew its `changed` list from empty on every delta tick of a metrics subscription; it now presizes the list to the current sample's tree count (an upper bound on the changed set), removing that regrowth. (3) The same method built its removed-tree ids via `.Where(...).ToList().OrderBy(...).ToArray()`, whose intermediate `.ToList()` is a throwaway list because `OrderBy` already materialises and sorts its source; it now sorts the filtered key sequence directly into the result array. Measured on the new `queryproj` microbench suite (`BENCH_MICROBENCH_SUITE=queryproj`, `MemoryDiagnoser`, in-process; each lane rebuilds the production shape per call so the `Allocated` delta is exactly the removed heap): multi-clause intersect (512 candidates, all surviving a second clause) 48,096 B -> 14,720 B (-69%, the removed survivor-dictionary regrowth), metrics changed list (256 changed trees) 4,264 B -> 2,104 B (-51%, the removed list regrowth), and removed-id projection (16 removed of 272) 1,056 B -> 872 B (-17%, the removed intermediate `List`). The `Orleans.Lattice.GrainIndex.Tests` (1,216 tests) and `Orleans.Lattice.Api.State.Tests` (492 tests) suites cover the preserved behaviour. (`Orleans.Lattice.GrainIndex` 9.5.0, `Orleans.Lattice.Api.State` 9.5.0)

- **The materialised-view maintainer now trims three redundant heap allocations from its warm cross-tree and batch-coalesce paths, with no change to behaviour or public API.** All three edits are in `Orleans.Lattice`'s view layer and preserve output byte-for-byte: (1) `ViewCatalog.All()` returned `_views.Values.ToArray()`, copying the `ConcurrentDictionary`'s moment-in-time snapshot (already an immutable `ReadOnlyCollection` over a freshly built `List`) into a second array that every caller only enumerates; it now returns that snapshot directly, removing one array allocation and copy per call (the warm caller is `ComputeViewWaitSet` on the cross-tree drain path). (2) `ViewMaintainerGrain.ComputeViewWaitSet` allocated a `HashSet` (plus its bucket and entry arrays) for the participant membership test on every cross-tree atomic batch; because such a batch spans only a handful of participant source trees in the common case, it now uses a threshold-gated (`<= 8`) ordinal linear scan over the participant list and keeps the `HashSet` only above the threshold - the membership answer is identical either way. (3) `ViewWriteCoalescer.Coalesce` grew its survivor `List` and key-index `Dictionary` from empty as it folded each WAL batch, reallocating their backing arrays; it now presizes both to the batch's known count (via `TryGetNonEnumeratedCount`), which is bounded by the actual input size and so presizes small on a sparse drain - unlike a fixed-`batchSize` presize, which over-allocates there. Measured on the new `viewmaint` microbench suite (`BENCH_MICROBENCH_SUITE=viewmaint`, `MemoryDiagnoser`, in-process; each lane rebuilds the production shape per call so the `Allocated` delta is exactly the removed heap): catalog snapshot 1,096 B -> 592 B (-46%), cross-tree wait set 1,240 B -> 984 B (-21%, the removed `HashSet` and its arrays), and batch coalesce of 256 writes 55,024 B -> 24,776 B (-55%, the removed list/dict regrowth). The 498-test `Orleans.Lattice.Tests.Views` suite (including `ViewWriteCoalescerTests`, `ViewCrossTreeCoordinatorGrainTests`, and the cross-tree maintainer integration tests) covers the preserved behaviour. (`Orleans.Lattice` 9.5.0)

- **The autoscaling scale-in safety gate is now genuinely split-aware: a silo can no longer be scaled in underneath an adaptive shard split.** `Orleans.Lattice.Scaling` has always listed "no shard split in flight" among its scale-in preconditions, but the probe behind it was a no-op that always answered "no splits", so the suppression never actually fired - scale-in relied on the WAL-healthy and all-dimensions-low preconditions plus the gate window alone. The reason it was stubbed is that shard-split progress was only ever exposed as an OpenTelemetry histogram, and metrics are write-only in-process, so there was nothing to read back. This adds the missing readable surface and wires it up. The cluster's split-admission singleton - which already held every tree's in-flight footprint, but only when the opt-in `MaxClusterConcurrentAutoSplits` ceiling was configured - is now fed by every autonomic monitor and reduced into a new public `ILatticeAdmin.GetSplitActivityAsync`, returning a `SplitActivityReport` (`InFlight`, `ReportingTrees`, `ObservedAt`, `AnyInFlight`). A tree with no ceiling configured contributes an **observation-only** footprint held in its own ledger, which the readable queries sum but admission never counts: such a tree never agreed to share a cluster budget, so it must not be able to throttle - or in the limit starve - a tree that opted in, and the number of splits a capped tree may start is unchanged. Publication is **edge-triggered**: a tree reports only while it actually has splits in flight, plus one call to clear its footprint when they finish, so an idle cluster issues no extra grain call; splits a monitor triggers are published in the same pass that starts them, so the gate never misses a split it caused, and a pass aborted by a suppressor (a resize, a bulk graft, or auto-split being switched off mid-drain) keeps its outstanding footprint alive rather than letting the activity source lapse to zero while splits are still draining. The heartbeat is also non-fatal: it sits upstream of the split triggers, so a transient gate failure is logged and skipped rather than aborting the pass and costing the tree its elasticity. Reading it costs one call to that singleton per sample tick and never fans out across trees or shards. Footprints carry a time-to-live, so a silo lost mid-split has its share reclaimed on expiry instead of suppressing scale-in indefinitely. Degradation is deliberately **fail-open** - an unreachable admin surface reports "no split in flight" and logs, because reporting the opposite would let a persistently unreachable surface suppress scale-in forever, turning a small self-correcting risk (a split doing some rework) into an unbounded cost ceiling. Split awareness is default-on with a kill switch, `LatticeScalingSignalOptions.SplitAwareScaleIn`, for a deployment with autonomic splitting disabled where the query would be pure overhead. Scale-**out** is unaffected either way. ([#1224](https://github.com/NSTA1/Orleans.Lattice/issues/1224)) (`Orleans.Lattice` 9.5.0, `Orleans.Lattice.Scaling` 9.5.0)

- **The aggregation applier now walks its freshly-materialised dictionaries through the struct enumerator instead of `.Keys`/`.Values`, removing a throwaway collection wrapper from three warm view-reduce paths.** `AggregationApplier` reduces every view contribution against dictionaries that live for a single call - the per-contribution accumulator slot map, and the per-materialise inverse/fold-inverse shard maps decoded from the store - and a `.Keys`/`.Values` access on such a fresh dictionary allocates a `KeyCollection`/`ValueCollection` wrapper on first touch (a long-lived dictionary caches the wrapper and pays nothing, but these do not). Three sites still did so, completing the direct-iteration pass the sibling `WorstKey`/`LargestSourceKey` helpers already started: (1) `ContributeNumericAsync`'s opportunistic cleanup loop walked `slots.Keys` over the fresh accumulator map; (2) `MaterialiseInverseAsync`'s shard gather walked `shards.Values` over the fresh batched-get result *and*, nested inside, `DecodeInverse(bytes).Values` over each fresh per-shard decode, so it allocated one wrapper for the outer map plus one per shard; (3) `MaterialiseFoldAsync`'s shard gather walked `shards.Values` over the fresh batched-get result. All three now iterate the dictionary directly, visiting exactly the same entries in the same order, so min/max/set-union/fold results and the emptied-slot cleanup are byte-for-byte unchanged. The public API is unchanged - every edit is to a `private` method on the internal applier. Measured on the new `aggiter` microbench suite (`BENCH_MICROBENCH_SUITE=aggiter`, `MemoryDiagnoser`, in-process; each lane rebuilds the fresh maps as production does per call, so the `Allocated` delta is exactly the removed wrapper(s)): numeric cleanup drops from 264 B to 240 B (-24 B, the removed `KeyCollection`); inverse materialisation drops from 4,880 B to 4,664 B (-216 B, the outer `ValueCollection` plus one per shard across an 8-shard group); fold materialisation drops from 4,688 B to 4,664 B (-24 B, the outer `ValueCollection`). The 498-test `Orleans.Lattice.Tests.Views` suite (including the 64 `AggregationApplierTests` covering numeric contribute/retract, inverse min/max/set-union, fold re-fold, and approximate-mode eviction) covers the preserved behaviour. (`Orleans.Lattice` 9.5.0)

- **The aggregation row codec now decodes view rows through a stack-only span reader instead of a per-decode `MemoryStream` + `BinaryReader`, mirroring the exact-size-array encode path already in place.** The codec serialises and deserialises the membership, inverse, and fold-inverse rows that the materialised-view maintainer reads back on every drain and query. The encode side was already allocation-trimmed to write into an exact-size `byte[]` via a `RowWriter` ref struct, but the three decode methods (`DecodeMembership`, `DecodeInverse`, `DecodeFoldInverse`) still wrapped each incoming buffer in a `MemoryStream` and a `BinaryReader`, allocating both throwaway objects on every row read on the view read path. They now read straight out of the `ReadOnlySpan<byte>` through a new stack-only `RowReader` ref struct that decodes the identical wire format - little-endian primitives, a 1-byte bool, `BinaryWriter`'s 7-bit length-prefixed UTF-8 strings, and raw byte runs - so the bytes read are byte-for-byte the legacy format and rows already persisted in the view tree continue to decode unchanged. The public API is unchanged - the codec is internal and every edit is to a private decode method. Measured on the extended `rowcodec` microbench suite (`BENCH_MICROBENCH_SUITE=rowcodec`, `MemoryDiagnoser`, in-process): each decode sheds the fixed 120 B `MemoryStream` + `BinaryReader` overhead per call - membership 280 B to 160 B (-43%), inverse (32 members) 3,888 B to 3,768 B, fold-inverse (32 members) 4,440 B to 4,320 B (the inverse and fold decodes still allocate the decoded strings, byte runs, and result dictionary, which are the payload and unavoidable) - while wall-clock drops from 147.2 ns to 59.8 ns for membership (-59%), 2,888.6 ns to 2,119.4 ns for inverse (-27%), and 3,178.8 ns to 1,946.9 ns for fold-inverse (-39%). New byte-compatibility tests in `AggregationRowCodecTests` decode reference-`BinaryWriter`-produced buffers to hold the wire format closed, and the existing round-trip suite covers the decoded values. (`Orleans.Lattice` 9.5.0)

- **The view maintainer now classifies each freshly drained batch inline during its one mandatory fold, instead of re-walking the buffer with up to three post-hoc `List.Exists(...)` passes per drain.** The maintainer tails its source write-ahead log continuously; every drain folds a batch of projected writes (or aggregation contributions) into a buffer, and the prior code then re-scanned that whole buffer purely to answer "did this batch carry a range operation?". The filter drain did this twice - `collected.Exists(w => w.Kind == RangeReconcile)` in `ViewMaintainerGrain.DrainAsync` and `collected.Exists(w => w.Kind == RangeDelete)` in `ApplySurvivorsAsync` - and the aggregation drain once, `contributions.Exists(c => c.Kind == RangeReconcile)` in `DrainAggregationAsync`. Because the fold loop already visits every element, each classification is now recorded in a `bool` set inside the fold and the separate `.Exists` passes are removed, so a freshly drained batch is walked once instead of two or three times. The `RangeDelete` fast-path flag is proven safe to compute at fold time: the only mutations between population and its use (`ConvertRangeReconcilesToMarkers`, which only rewrites `RangeReconcile` entries to `Upsert`, and `ShapeHistoryWritesAsync`, which reshapes only `Upsert` rows and passes every non-`Upsert` through untouched) never add or remove a `RangeDelete`, so the flag equals the removed scan on every path including accumulative views. Behaviour is identical and the public API is unchanged - every edit is to a `private` grain method - and the change is allocation-neutral by construction (the buffer is built identically either way and the removed `.Exists` scans used a cached static delegate), so this is a pure CPU-cycle trim on a continuously running maintenance path. Measured on the new `viewdrain` microbench suite (`BENCH_MICROBENCH_SUITE=viewdrain`, `MemoryDiagnoser`; each lane performs the mandatory fold so the delta is exactly the removed pass(es), and `Allocated` is equal within each pair): the two-scan filter classification drops from 2,128.7 ns to 1,080.9 ns at a 256-write batch (-49%) and 7,640.0 ns to 3,622.7 ns at 1,024 (-53%); the one-scan aggregation classification drops from 1,302.9 ns to 957.9 ns at 256 (-26%) and 4,567.4 ns to 3,951.0 ns at 1,024 (-13%). The 495-test `Orleans.Lattice.Tests.Views` suite (including the maintainer drain integration tests) covers the preserved behaviour. (`Orleans.Lattice` 9.5.0)

- **The cross-tree and view coordination barriers now canonicalise their wait/participant sets with a HashSet-plus-in-place-sort helper instead of `Distinct().OrderBy()`, removing LINQ's ordering machinery from three warm coordination paths.** Three sites froze or stamped a string set with `source.Distinct(StringComparer.Ordinal).OrderBy(v => v, StringComparer.Ordinal).ToList()/.ToArray()`: `ViewCrossTreeCoordinatorGrain` and `LatticeCrossTreeReceiverGrain` freezing their wait set on the first registration/terminal, and `LatticeCrossTreeTxGrain` stamping the sorted participant tree-id set onto every sub-saga's terminal for the receiver-side visibility barrier. Per call, that LINQ form allocates an `OrderedEnumerable` wrapper, a materialised element buffer, a projected key array, and an integer sort-index map, all on top of the result collection. A new internal `CanonicalStringSet` helper de-duplicates through a single `HashSet<string>` pass and sorts the result in place, so only the result (and the transient dedup set) is allocated; the output is byte-for-byte identical (a de-duplicated set rendered in a total ordinal order is deterministic regardless of input order), so the exact-match barrier comparison and the stamped participant set are unchanged. The public API is unchanged - the helper is internal. Measured on the new `crosstree` microbench suite (`BENCH_MICROBENCH_SUITE=crosstree`, `ShortRun`, `MemoryDiagnoser`; allocation is deterministic and is the win, wall-clock for these sub-microsecond operations sits within run-to-run noise): the wait-set freeze drops from 744 B to 392 B at width 2 (-47%), 1,112 B to 640 B at width 8 (-42%), and 4,960 B to 3,368 B at width 64 (-32%); the participant set drops from 648 B to 328 B at width 2 (-49%), 1,400 B to 960 B at width 8 (-31%), and 5,592 B to 4,032 B at width 64 (-28%). The prior-versus-new equivalence of the helper is held closed by new tests in `CanonicalStringSetTests`. (`Orleans.Lattice` 9.5.0)

- **Three warm dictionary/set maintenance paths now trim their per-call heap: the tag-index reconcile baseline prune, the aggregation approximate-mode eviction key scan, and small-tag normalisation.** (1) `TagIndexReconcileGrain.PruneStaleBaselines` ran two LINQ passes over `Baselines.Keys` on every reconcile begin - a `.All(keep.Contains)` method-group delegate on the no-drift path and a `.Where(t => !keep.Contains(t)).ToList()` capturing closure plus `Where` iterator on the drift path; it now collects stale tree-ids in a single manual pass over the dictionary, deferring the list until the first stale key so the common no-drift case allocates nothing beyond the membership set. (2) `AggregationApplier.LargestSourceKey` iterated the freshly decoded per-mutation inverse-shard map through `map.Keys`, allocating a throwaway `KeyCollection` on each approximate-mode min/max eviction; it now iterates the map's entries directly with the struct enumerator, mirroring the sibling `WorstKey`. (3) `LatticeTagIndexContext.NormalizeTags` always built a `HashSet<string>` to dedup, even for the handful of tags a tag write or query carries in practice; it now dedups small tag sets (<= 16) with an ordinal linear scan against the accumulating result and keeps the set only for larger inputs (preserving `O(n)` membership there). All three preserve exact behaviour - first-seen order, per-element validation including duplicates, and the same stale/eviction/dedup results - and the public API is unchanged, since every edit is to a private method. Measured on the new `alloctrims` microbench suite (`BENCH_MICROBENCH_SUITE=alloctrims`, `ShortRun`, `MemoryDiagnoser`; allocation is deterministic and is the win, wall-clock for these sub-microsecond operations sits within run-to-run noise): baseline prune drops from 960 B to 896 B on the no-drift path (-64 B, the removed delegate) and 1,040 B to 912 B on the drift path (-128 B, the removed closure and iterator); the eviction key scan drops by a size-independent 24 B per call (352 B to 328 B at 4 entries, 632 B to 608 B at 16, 2,144 B to 2,120 B at 64); and small-tag normalisation drops from 560 B to 192 B (-66%). The dedup contract is held closed by a new duplicate-tag regression test in `LatticeTagIndexIntegrationTests`, and the existing reconcile and aggregation suites cover the other two. (`Orleans.Lattice` 9.5.0)

- **Three structural read fan-out sites now issue one batched multi-get instead of a sequential read per item, cutting per-operation index-tree grain calls on the tag-index AND query, the atomic-write pre-image capture, and materialised-view group materialisation.** Each site awaited one index-tree read per sibling tag, per written key, or per view shard, so its caller-visible read round-trips grew linearly with the fan-out width. (1) `LatticeTagIndexContext.QueryAsync` (all-tags branch) probed each candidate key's `(T-1)` sibling-tag membership rows with one sequential `RowLiveAsync` apiece; it now issues a single batched `GetManyAsync` per candidate, collapsing the AND query's index reads from `O(candidates x tags)` to `O(candidates)` while preserving the exact per-row live/flag-mode semantics and duplicate-tag correctness. (2) `AtomicActionGrain.RunTreeWriteForwardAsync` captured each written key's pre-image with one `GetAsync` per key; it now reads the whole step's pre-images with one `tree.GetManyAsync`. (3) `AggregationApplier.MaterialiseInverseAsync` and `MaterialiseFoldAsync` gathered a group's shards with one `store.GetAsync` per slot; they now read all slots with one `store.GetManyAsync` on a new internal `IAggregationViewStore.GetManyAsync` (implemented over the batching view tree and the buffering overlay). The public API is unchanged - the new multi-get lives on the internal view-store interface, and the batched calls reuse the existing public `ILattice.GetManyAsync` fan-out semantics (absent/tombstoned keys omitted), so results are byte-for-byte identical. Measured on the new `fanout` microbench suite (`BENCH_MICROBENCH_SUITE=fanout`). Exact deterministic round-trip census (`BENCH_FANOUT_ROUNDTRIPS_ONLY=true`): atomic pre-image 128 keys 128 -> 1 read (0.01x), view inverse 64 shards 64 -> 1 (0.02x), tag-index AND 8 tags x 100 keys 700 -> 100 (0.14x). BDN wall-clock (`ShortRun`, Width 64): atomic pre-image 38.6 us -> 29.6 us (-23%) and view inverse 41.0 us -> 28.7 us (-30%); the tag-index probe's in-process wall-clock is dominated by the trivial per-read cost and the loss of the baseline's per-candidate short-circuit, so its win is the index-tree grain-call reduction the census measures, not in-process latency. (`Orleans.Lattice` 9.5.0)

Newer release: [Release 2026-09-02](../changelog/2026-09-02.md). Older release: [Release 2026-09-02: Fixed to Security](../changelog/2026-09-02-2.md). Contents: [Changelog](../CHANGELOG.md).
