Release 2026-09-02: Breaking to Changed
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at 2026-09-02-1.md, and llms.txt lists every page.Part of Release 2026-09-02, in Changelog.
Breaking
- A cluster-wide
Telemetrygrant is now authorized over the all-trees sentinel"*"rather than the reserved auth policy tree, so an operator who authored the grant against the old scope must re-author it. This is the fix for (#1795) and is described in full under Fixed. The blast radius is deliberately small, which is why it ships in a minor rather than waiting for the next major: the documented authoring path,LatticeScope.ClusterWide(), produced a"*"scope that the authorizers never consulted, so a grant written the documented way was inert and no conforming deployment can regress. Only a deployment that discovered the mismatch and worked around it by grantingTelemetryon the reservedsys-auth-policytree is affected. No public API signature changed, and the constants involved are internal, so this is a behavioural change with no compile break. The failure direction is fail-closed: an affected caller is denied cluster telemetry rather than over-granted, so the risk is the availability of one telemetry surface, never an escalation. Migration: re-author the grant withLatticeScope.ClusterWide()(equivalently, scopeTree:*) forLatticeOperation.Telemetry; the reserved-tree grant can then be withdrawn. (Orleans.Lattice.Auth9.5.0,Orleans.Lattice.Api.TreeAdmin9.5.0)
Added
You can now apply many CRDT deltas in one call, instead of paying a full round trip per key.
ILattice.ApplyCrdtDeltaManyAsyncis the CRDT counterpart ofSetManyAsync: it fans out to shards in parallel and collapses each leaf's slice into a single write-ahead-log dispatch, so an N-key batch costs one commit-log round trip per leaf rather than N. Marking presence or membership flags - a loop ofOrFlag(key).EnableAsync(...)- was the motivating workload, and it now has a typed one-liner,EnableManyAsync, which reads every current flag in one batched call and mints all the deltas from that snapshot, halving the round trips again. The merge mode is declared once for the batch rather than per entry, because a tree resolves exactly one CRDT shape; mixing shapes inside a tree converges locally but diverges at replication, so the API does not offer that footgun. The batch is not atomic, likeSetManyAsync- but because every entry folds a delta rather than overwriting a value, a retry converges instead of clobbering a concurrent writer. When the batch must land all-or-nothing,StageEnableManyAsyncmints the staging tokens from the same single read andLatticeAtomicWriteBuilder.SetManycouples them into one existing cross-tree saga - no new transactional machinery, so the all-or-nothing and cross-cluster guarantees are unchanged. (#1921) (Orleans.Lattice9.5.0)Adaptive shard split gained its missing inverse, leaf snapshots gained a compact binary encoding, and a cold leaf no longer pays an all-or-nothing load. Five defects and two missing capabilities on the core tree's activation and maintenance paths are addressed together. Shard consolidation is the missing inverse of adaptive split: a split was previously a one-way door, so a tree that split under transient load stayed split forever. Consolidation folds an over-split tree back down fully online, and a new healing orchestrator observes tree shape and heals automatically rather than waiting to be invoked. Split admission now requires load to be genuinely skewed rather than merely high, so a bulk ingest no longer shatters a tree while a truly hot shard still splits exactly as before (
HotShardMinSkewRatio,HotShardMinShardEntries, and aMaxPhysicalShardsPerTreeceiling). Leaf snapshots gained a compact binary codec in place of the previous JSON-plus-base64 encoding, with transparent dual-read and lazy rewrite, so blobs already persisted stay readable indefinitely and are re-encoded on their next natural capture with no migration step (LeafSnapshotBinaryEncodingEnabled). Leaf activation now hydrates incrementally, so its cost follows what is actually read rather than how large the leaf happens to be, instead of loading the entire snapshot up front (LeafPartialHydrationEnabled,LeafHydrationResidentBytes). Two durability defects are fixed with them: write-ahead-log garbage collection deferred its first pass by up to an hour, so a box recreated more often than that never trimmed at all - the delay is now a short startup window with a floor between passes (WalGcStartupDelay,WalGcMinInterval); and a single deferred mutation pinned the incremental checkpoint flush for an entire replay, so a long replay made no durable progress and re-ran the identical window on its next activation. Every mechanism is default-on with its own documented kill switch, so an operator who hits trouble disables one behaviour rather than reverting the image, and the consolidation and healing surfaces are internal - the new configuration knobs onLatticeOptionsare the whole of the public surface. New instruments on theorleans.latticemeter report split admission, consolidation progress and healing decisions, charted by new panels on the bundled Overview and Replication dashboards. (#1830) (Orleans.Lattice9.5.0,Orleans.Lattice.Dashboards9.5.0)You can now authenticate lattice callers against any standards-conformant OpenID Connect provider - Okta, Auth0, Keycloak, Ping, Google - without writing an authenticator. The new
Orleans.Lattice.Membership.Oidccompanion package addssiloBuilder.AddLatticeOidc(...), a generic bearer-credential authenticator driven entirely by the provider's discovery document. Point it at anAuthority(or an explicitMetadataAddress), name the exactIssueryou trust and theAudiencesyou accept, and it discovers the JWKS, rotates signing keys on its own refresh schedule with last-known-good fallback, and mapssubto the lattice subject andgroups/roles/roleto asserted groups - each configurable. It is an additive sibling to the Entra authenticator rather than a replacement: register both, plus the JWT and anonymous authenticators, on the same silo, and callAddLatticeOidc(...)once per issuer. Selection is unambiguous by construction because it is an exact ordinal issuer match (or an explicitSchemeHint), never a prefix, wildcard, or catch-all, so an authenticator never claims another provider's token regardless of registration order. Three fail-closed guards are load-bearing: signature algorithms are pinned from the discovery document'sid_token_signing_alg_values_supported(or from an explicitAlgorithmslist, which wins), and a provider that advertises none rejects every token rather than accepting any - so the algorithm-confusion gap an unpinned allow-list leaves open is closed, including the classic case of anHS256forgery signed with the provider's published public key; an emptyAudienceslist throws at construction instead of silently accepting any validly-signed token minted for a different relying party; and a token whose subject collides with a reserved sentinel resolves to the anonymous subject rather than to an anonymous-labelled principal that still carries the token's groups. The issuer is also pinned to the configured value rather than widened to whatever the discovery document advertises. Because authenticator selection reads the token's issuer, and the credential bridges stamp the scheme from operator configuration rather than from the caller, that selection parse runs on a pre-authentication path; it is bounded by the validating handler's own maximum token size, so an unauthenticated caller cannot use an oversized credential to amplify work once per registered authenticator, and no credential that could have authenticated is ever declined. (#1804) (Orleans.Lattice.Membership.Oidc9.5.0)A backup sink that is not actually shared across your clusters is now caught when a silo starts, instead of when a coordinated restore fails. A coordinated restore of a replicated tree only works if every cluster can read the same backup, but until now a deployment where each cluster pointed at its own isolated sink looked perfectly healthy right up until the restore saga aborted all-or-nothing - by which point you had been capturing backups you could never restore. Each cluster now writes a tiny self-cleaning canary marker into its configured sink at startup and checks whether it can read every peer's marker back, which is a real reachability proof rather than the previous assumption that any external sink must be shared. The verdict is three-valued so a cold start cannot cry wolf: seeing a peer's marker is
Shared, not seeing it while that peer is also unreachable on the control channel isUnverifiedand stays quiet, and not seeing it while the peer is reachable isNotShared- the genuine misconfiguration. Detection defaults to a loud startup warning (LatticeBackupOptions.SinkSharingEnforcement, defaultWarn), so upgrading cannot brick a deployment that is merely transiently unreachable; setFailFastto refuse to start on a refuted sink, orDisabledto skip the probe entirely - thoughDisableddeliberately does not relax the pre-existing hard failure on an in-cluster sink. The same signal now also feeds the periodic backup health sweep, so a backup that verifies locally but cannot be seen by a peer reportsWarningwith the unreachable cluster named, and the problem surfaces in the Explorer Backups tab long before anyone attempts a restore. This closes the blind spot in the old startup guard, which only rejected the in-cluster sink and assumed every external sink was shared. Single-cluster and non-replicated deployments are untouched: with no replicated trees or no peers the probe never runs, writes nothing, and reports exactly what it did before. (#1669) (Orleans.Lattice.Backup9.5.0,Orleans.Lattice.Replication9.5.0)An Orleans grain's typed state can now be indexed and queried, without a hand-maintained secondary index. The new
Orleans.Lattice.GrainIndexcompanion package lets you declare that a grain's persistent state should be tracked in a lattice tree, then ask typed predicate questions over it - "whichUsergrains are 18 or over?" - and stream back grain references, keys, or matches. An index is opt-in twice: the silo declares it withAddGrainIndex<TGrain, TState>(...), naming each projected property explicitly withInclude(u => u.Age)so write amplification is always a deliberate choice, and the grain marks its state with the[Indexed]facet. Grains enrol on activation and on every state write, while a resumable, rate-limited, reminder-driven backfill onboards the dormant population from an application-suppliedIGrainKeySource; both routes write through one projection path, so they converge on a single duplicate-free index. Queries reuse the core predicate dialect, so filtering runs in the tree shards: a single-property comparison becomes a contiguous range scan over an order-preserving key encoding, a conjunction becomes one scan per property whose grain keys are intersected, and a disjunction unions and de-duplicates its branches. Entry updates for a grain are all-or-nothing, a failed index write is surfaced to the caller and retried from a durable outbox marker that carries the whole plan (so a repair never has to wake the grain), index trees live under a reserved__grainindex/namespace and are cluster-local unless you opt in, and a declaration change that would invalidate already-written entries is rejected at startup rather than silently answering wrong. Eight instruments on the sharedorleans.latticemeter, a Grafana dashboard, and anIGrainIndexAdminsurface (status, progress, pause/resume/rebuild, single-pass) report and control it. (#657) (Orleans.Lattice.GrainIndex9.5.0)A silo restart can now pre-warm the leaf caches a shard was actually reading, instead of serving its first reads cold. Set
LatticeOptions.LeafCachePreWarmCountto a non-zero value and each shard root keeps a small bounded histogram of the leafGrainIds its read traversals land on, persists a compact snapshot of it into its own durable state, and - on the nextILattice.WarmUpAsyncafter reactivation - primes that manyLeafCacheGrainactivations for the leaves it read most. Ranking by observed read frequency rather than by recency means a leaf touched once just before shutdown does not displace a leaf that was being read constantly, which is the failure mode a recency list has on the skewed and cyclic access patterns a B+ tree produces under a real key distribution: on held-out synthetic traces (train on the first half, score against the hot set of the second) frequency recovered 96% of the true hot set on a Zipf-skewed trace and 98% on a cyclic one, against 56% and 53% for a recency list of the same size. Because the pre-warm is driven from the shard root, the[StatelessWorker]cache activations land on the silo hosting that shard root - the same silo that will serve the subsequent reads - so no silo-lifecycle infrastructure is involved. The feature is off by default (0disables it), the model is bounded to 256 tracked leaves resident and 64 persisted (a snapshot under ~3 KB), the read hot path costs two predictable branches and no allocation when disabled, the snapshot is written on a coalescing timer and at deactivation rather than per read, and a leaf that fails to warm is logged and skipped rather than failingWarmUpAsync. Three new metrics (orleans.lattice.leaf_access.model.leaves,orleans.lattice.warmup.leaf_cache.prewarmed,orleans.lattice.warmup.leaf_cache.duration) and a Grafana panel report what the model holds and what each warm-up primed. (#332) (Orleans.Lattice9.5.0)A troubleshooting guide now walks you from a symptom to a fix.
docs/lattice/troubleshooting.mdshows how to read aDiagnoseAsyncreport field by field - including a worked example and the traps that lead to a wrong conclusion, such as a shallow report always reporting zero tombstones and an all-zero shard entry meaning the diagnostics fan-out failed - and then covers storage-provider exceptions on write, concurrent shard-split activity, slow scans, and stale reads, each as symptom, likely cause, how to confirm, and how to fix. (#342)A migration guide for importing an existing Redis, relational, or Cosmos DB dataset into a tree. Migrating from an External Store is the operational playbook that sits on top of the bulk-loading reference: it compares the three ingest paths (the one-shot
ILattice.BulkLoadAsync, the streamingIAsyncEnumerableoverload, and the resumable tree-administration Begin/Append/Commit session) against the ordering guarantee each one asks you to own, then documents the contract a migration actually has to satisfy - which paths enforce an empty target and which silently do not, ordinal key comparison, whatBulkLoadOrderExceptiondoes and does not catch, and where the path is idempotent rather than transactional. Sections on key design and value encoding cover the choices that cannot be revised after the load (zero-padding numeric segments, separator selection, why a key prefix is a logical range and not a placement hint, choosing anILatticeSerializer<T>, and planning schema evolution up front), and per-store playbooks for Redis, relational databases, and Cosmos DB cover stable resumable enumeration and each store's ordering pitfalls. It closes with how to verify what landed before cutting over. (#341)You can now see exactly how much write latency each
IMutationObserveris costing you. Observers run inline on the grain write path, so a slow one is paid for by every caller - but until now the only signal was the aggregateobserverstep onorleans.lattice.leaf.commit.duration, which says a cost was incurred without saying by whom. A neworleans.lattice.observer.durationhistogram (milliseconds) on the existingorleans.latticemeter times each observer callback individually and tags itobserver(the observer's CLR type name) andtree, so a misbehaving observer doing synchronous I/O or blocking on a downstream call is named on the same OpenTelemetry pipeline as the traffic it slows down. The measurement is taken on the faulting path too, so an observer that throws slowly is as visible as one that returns slowly, and the dispatcher continues to swallow and log the exception. The zero-cost-when-unused guarantee is unchanged: the dispatcher's no-observer fast path returns before any timing work, and a registered observer with no metrics listener attached costs one boolean read per publish. The sample spans only the callback, so the warning the dispatcher logs for a faulting observer is not billed to the observer that faulted - a slow log sink can no longer read as a slow observer. Charted by a new by-observer p50/p95 panel on the bundled Commit Path dashboard. A newDashboardJsonTestsdrift guard also fails the build on duplicate Grafana panel ids, which is how two concurrently-authored panels silently collide on the conventionalmax(id) + 1; one such pre-existing collision on the Overview dashboard is corrected with it. No public API changed. (#430) (Orleans.Lattice9.5.0,Orleans.Lattice.Dashboards9.5.0)
Changed
Three more warm read paths now presize their result collection to a tight in-hand bound, trimming redundant heap from leaf range-entry reads, tree-structure topology scans, and the causal-apply drain, with no change to behaviour or public API. All three edits preserve output byte-for-byte and touch only internal/private method bodies: (1)
BPlusLeafGrain.GetEntriesAsyncgrew its result list from empty while scanning the leaf's ordered cache rows - the single unpresized method among its already-presized read siblings; it now presizes toMath.Min(Cache.Count, 256), the exact tight upper bound the siblingGetKeysAsyncalready uses (capped at 256 so a range/split-truncated scan never over-allocates against a full leaf). (2)LatticeStateQuery.GetTreeStructureAsyncgrew its top-levelrootNodeslist from empty across the scanned shard span while every nested topology list was already presized; it now presizes to that span (endShard - startShard, one root summary per shard - the budget-capped upper bound), which required only reordering the two lines that compute the span. (3)CausalApplyBuffer.DrainSatisfiedgrew its drainedreadylist from empty while walking the parked-entry list; it now presizes to the parked-entry count read under the buffer's own lock (the drain yields at most one record per parked entry), which collapses to a zero-capacity list on the empty steady-state buffer so the common in-order delivery path is unaffected. Measured on the newreadpathpresizemicrobench suite (BENCH_MICROBENCH_SUITE=readpathpresize,MemoryDiagnoser, in-process; each lane rebuilds the production shape per call so theAllocateddelta is exactly the removed heap): leaf range-entry read (512 cached rows) 16.16 KB -> 12.08 KB (-25%, the removed list regrowth); tree-structure scan (64 shards) 3.12 KB -> 2.55 KB (-18%, the removed list regrowth); causal-apply drain (128 parked entries) 2.14 KB -> 1.05 KB (-51%, the removed list regrowth). TheBPlusLeafGrainsuite inOrleans.Lattice.Tests, theOrleans.Lattice.Api.State.Testssuite, and theCausalApplyBuffersuite inOrleans.Lattice.Replication.Testscover the preserved read and drain semantics. (Orleans.Lattice9.5.0,Orleans.Lattice.Api.State9.5.0,Orleans.Lattice.Replication9.5.0)Three warm read paths now presize (or stop copying) their result collection, trimming redundant heap from view-catalog listing, live-entry reads, and snapshot-shard scans, with no change to behaviour or public API. All three edits preserve output byte-for-byte and touch only internal/private method bodies: (1)
LatticeStateQuery.ListViewsAsyncbuilt its authenticatedvisiblelist already presized to the candidate count, then copied it into a throwaway array viavisible.ToArray()on every authenticated request; it now hands the list back as anIReadOnlyList<ViewListing>with no copy and iterates it index-based (notforeach) so reading through the interface does not boxList<T>'s struct enumerator - a hazard that would otherwise offset the saved array on small lists. (2)BPlusLeafGrain.GetLiveEntriesAsyncgrew its result dictionary from empty while folding every live cached row, reallocating its bucket and entry arrays; it now presizes to the cached row count (the fold's upper bound), exactly as the siblingGetLiveRawEntriesAsyncalready did. (3)SnapshotLeafGrain.GetKeysAsync/GetEntriesAsync/GetRawEntriesAsyncgrew their result list from empty while scanning the pinned folder's ordered rows; they now presize toMath.Min(limit, entryCount), a tight upper bound on the emitted count that stays small on a narrow page. Measured on the newreadpathtrimsmicrobench suite (BENCH_MICROBENCH_SUITE=readpathtrims,MemoryDiagnoser, in-process; each lane rebuilds the production shape per call so theAllocateddelta is exactly the removed heap): view-catalog listing (256 visible views) 4.08 KB -> 2.05 KB (-50%, the removed array copy) and 1.56 us -> 1.11 us; live-entry read (512 cached rows) 46.97 KB -> 14.38 KB (-69%, the removed dictionary regrowth) and 12.31 us -> 6.50 us; snapshot-shard scan (512 rows) 16.16 KB -> 8.05 KB (-50%, the removed list regrowth) and 4.42 us -> 3.13 us. TheOrleans.Lattice.Api.State.Testssuite covers the view-listing behaviour, and theSnapshotLeafGrain,BPlusLeafGrain, andLatticeCursorGrainsuites inOrleans.Lattice.Testscover the preserved read semantics. (Orleans.Lattice9.5.0,Orleans.Lattice.Api.State9.5.0)The grain-index query executor and the state-API metrics observer now trim three redundant heap allocations from their warm query-intersect and metrics-delta paths, with no change to behaviour or public API. All three edits preserve output byte-for-byte and touch only internal/private method bodies: (1)
GrainIndexQueryExecutor.IntersectAsyncgrew its per-clausesurvivorsdictionary from empty as it folded each later clause's scan into a multi-clause AND query, reallocating its bucket and entry arrays; because every survivor is a key already present in the drivingcandidatesset, that set's current count is a tight upper bound, so it now presizessurvivorstocandidates.Count, removing the grow-from-empty rehash churn on each intersect pass. (2)LatticeStateMetricsObserver.ObserveAsyncgrew itschangedlist from empty on every delta tick of a metrics subscription; it now presizes the list to the current sample's tree count (an upper bound on the changed set), removing that regrowth. (3) The same method built its removed-tree ids via.Where(...).ToList().OrderBy(...).ToArray(), whose intermediate.ToList()is a throwaway list becauseOrderByalready materialises and sorts its source; it now sorts the filtered key sequence directly into the result array. Measured on the newqueryprojmicrobench suite (BENCH_MICROBENCH_SUITE=queryproj,MemoryDiagnoser, in-process; each lane rebuilds the production shape per call so theAllocateddelta is exactly the removed heap): multi-clause intersect (512 candidates, all surviving a second clause) 48,096 B -> 14,720 B (-69%, the removed survivor-dictionary regrowth), metrics changed list (256 changed trees) 4,264 B -> 2,104 B (-51%, the removed list regrowth), and removed-id projection (16 removed of 272) 1,056 B -> 872 B (-17%, the removed intermediateList). TheOrleans.Lattice.GrainIndex.Tests(1,216 tests) andOrleans.Lattice.Api.State.Tests(492 tests) suites cover the preserved behaviour. (Orleans.Lattice.GrainIndex9.5.0,Orleans.Lattice.Api.State9.5.0)The materialised-view maintainer now trims three redundant heap allocations from its warm cross-tree and batch-coalesce paths, with no change to behaviour or public API. All three edits are in
Orleans.Lattice's view layer and preserve output byte-for-byte: (1)ViewCatalog.All()returned_views.Values.ToArray(), copying theConcurrentDictionary's moment-in-time snapshot (already an immutableReadOnlyCollectionover a freshly builtList) into a second array that every caller only enumerates; it now returns that snapshot directly, removing one array allocation and copy per call (the warm caller isComputeViewWaitSeton the cross-tree drain path). (2)ViewMaintainerGrain.ComputeViewWaitSetallocated aHashSet(plus its bucket and entry arrays) for the participant membership test on every cross-tree atomic batch; because such a batch spans only a handful of participant source trees in the common case, it now uses a threshold-gated (<= 8) ordinal linear scan over the participant list and keeps theHashSetonly above the threshold - the membership answer is identical either way. (3)ViewWriteCoalescer.Coalescegrew its survivorListand key-indexDictionaryfrom empty as it folded each WAL batch, reallocating their backing arrays; it now presizes both to the batch's known count (viaTryGetNonEnumeratedCount), which is bounded by the actual input size and so presizes small on a sparse drain - unlike a fixed-batchSizepresize, which over-allocates there. Measured on the newviewmaintmicrobench suite (BENCH_MICROBENCH_SUITE=viewmaint,MemoryDiagnoser, in-process; each lane rebuilds the production shape per call so theAllocateddelta is exactly the removed heap): catalog snapshot 1,096 B -> 592 B (-46%), cross-tree wait set 1,240 B -> 984 B (-21%, the removedHashSetand its arrays), and batch coalesce of 256 writes 55,024 B -> 24,776 B (-55%, the removed list/dict regrowth). The 498-testOrleans.Lattice.Tests.Viewssuite (includingViewWriteCoalescerTests,ViewCrossTreeCoordinatorGrainTests, and the cross-tree maintainer integration tests) covers the preserved behaviour. (Orleans.Lattice9.5.0)The autoscaling scale-in safety gate is now genuinely split-aware: a silo can no longer be scaled in underneath an adaptive shard split.
Orleans.Lattice.Scalinghas always listed "no shard split in flight" among its scale-in preconditions, but the probe behind it was a no-op that always answered "no splits", so the suppression never actually fired - scale-in relied on the WAL-healthy and all-dimensions-low preconditions plus the gate window alone. The reason it was stubbed is that shard-split progress was only ever exposed as an OpenTelemetry histogram, and metrics are write-only in-process, so there was nothing to read back. This adds the missing readable surface and wires it up. The cluster's split-admission singleton - which already held every tree's in-flight footprint, but only when the opt-inMaxClusterConcurrentAutoSplitsceiling was configured - is now fed by every autonomic monitor and reduced into a new publicILatticeAdmin.GetSplitActivityAsync, returning aSplitActivityReport(InFlight,ReportingTrees,ObservedAt,AnyInFlight). A tree with no ceiling configured contributes an observation-only footprint held in its own ledger, which the readable queries sum but admission never counts: such a tree never agreed to share a cluster budget, so it must not be able to throttle - or in the limit starve - a tree that opted in, and the number of splits a capped tree may start is unchanged. Publication is edge-triggered: a tree reports only while it actually has splits in flight, plus one call to clear its footprint when they finish, so an idle cluster issues no extra grain call; splits a monitor triggers are published in the same pass that starts them, so the gate never misses a split it caused, and a pass aborted by a suppressor (a resize, a bulk graft, or auto-split being switched off mid-drain) keeps its outstanding footprint alive rather than letting the activity source lapse to zero while splits are still draining. The heartbeat is also non-fatal: it sits upstream of the split triggers, so a transient gate failure is logged and skipped rather than aborting the pass and costing the tree its elasticity. Reading it costs one call to that singleton per sample tick and never fans out across trees or shards. Footprints carry a time-to-live, so a silo lost mid-split has its share reclaimed on expiry instead of suppressing scale-in indefinitely. Degradation is deliberately fail-open - an unreachable admin surface reports "no split in flight" and logs, because reporting the opposite would let a persistently unreachable surface suppress scale-in forever, turning a small self-correcting risk (a split doing some rework) into an unbounded cost ceiling. Split awareness is default-on with a kill switch,LatticeScalingSignalOptions.SplitAwareScaleIn, for a deployment with autonomic splitting disabled where the query would be pure overhead. Scale-out is unaffected either way. (#1224) (Orleans.Lattice9.5.0,Orleans.Lattice.Scaling9.5.0)The aggregation applier now walks its freshly-materialised dictionaries through the struct enumerator instead of
.Keys/.Values, removing a throwaway collection wrapper from three warm view-reduce paths.AggregationApplierreduces every view contribution against dictionaries that live for a single call - the per-contribution accumulator slot map, and the per-materialise inverse/fold-inverse shard maps decoded from the store - and a.Keys/.Valuesaccess on such a fresh dictionary allocates aKeyCollection/ValueCollectionwrapper on first touch (a long-lived dictionary caches the wrapper and pays nothing, but these do not). Three sites still did so, completing the direct-iteration pass the siblingWorstKey/LargestSourceKeyhelpers already started: (1)ContributeNumericAsync's opportunistic cleanup loop walkedslots.Keysover the fresh accumulator map; (2)MaterialiseInverseAsync's shard gather walkedshards.Valuesover the fresh batched-get result and, nested inside,DecodeInverse(bytes).Valuesover each fresh per-shard decode, so it allocated one wrapper for the outer map plus one per shard; (3)MaterialiseFoldAsync's shard gather walkedshards.Valuesover the fresh batched-get result. All three now iterate the dictionary directly, visiting exactly the same entries in the same order, so min/max/set-union/fold results and the emptied-slot cleanup are byte-for-byte unchanged. The public API is unchanged - every edit is to aprivatemethod on the internal applier. Measured on the newaggitermicrobench suite (BENCH_MICROBENCH_SUITE=aggiter,MemoryDiagnoser, in-process; each lane rebuilds the fresh maps as production does per call, so theAllocateddelta is exactly the removed wrapper(s)): numeric cleanup drops from 264 B to 240 B (-24 B, the removedKeyCollection); inverse materialisation drops from 4,880 B to 4,664 B (-216 B, the outerValueCollectionplus one per shard across an 8-shard group); fold materialisation drops from 4,688 B to 4,664 B (-24 B, the outerValueCollection). The 498-testOrleans.Lattice.Tests.Viewssuite (including the 64AggregationApplierTestscovering numeric contribute/retract, inverse min/max/set-union, fold re-fold, and approximate-mode eviction) covers the preserved behaviour. (Orleans.Lattice9.5.0)The aggregation row codec now decodes view rows through a stack-only span reader instead of a per-decode
MemoryStream+BinaryReader, mirroring the exact-size-array encode path already in place. The codec serialises and deserialises the membership, inverse, and fold-inverse rows that the materialised-view maintainer reads back on every drain and query. The encode side was already allocation-trimmed to write into an exact-sizebyte[]via aRowWriterref struct, but the three decode methods (DecodeMembership,DecodeInverse,DecodeFoldInverse) still wrapped each incoming buffer in aMemoryStreamand aBinaryReader, allocating both throwaway objects on every row read on the view read path. They now read straight out of theReadOnlySpan<byte>through a new stack-onlyRowReaderref struct that decodes the identical wire format - little-endian primitives, a 1-byte bool,BinaryWriter's 7-bit length-prefixed UTF-8 strings, and raw byte runs - so the bytes read are byte-for-byte the legacy format and rows already persisted in the view tree continue to decode unchanged. The public API is unchanged - the codec is internal and every edit is to a private decode method. Measured on the extendedrowcodecmicrobench suite (BENCH_MICROBENCH_SUITE=rowcodec,MemoryDiagnoser, in-process): each decode sheds the fixed 120 BMemoryStream+BinaryReaderoverhead per call - membership 280 B to 160 B (-43%), inverse (32 members) 3,888 B to 3,768 B, fold-inverse (32 members) 4,440 B to 4,320 B (the inverse and fold decodes still allocate the decoded strings, byte runs, and result dictionary, which are the payload and unavoidable) - while wall-clock drops from 147.2 ns to 59.8 ns for membership (-59%), 2,888.6 ns to 2,119.4 ns for inverse (-27%), and 3,178.8 ns to 1,946.9 ns for fold-inverse (-39%). New byte-compatibility tests inAggregationRowCodecTestsdecode reference-BinaryWriter-produced buffers to hold the wire format closed, and the existing round-trip suite covers the decoded values. (Orleans.Lattice9.5.0)The view maintainer now classifies each freshly drained batch inline during its one mandatory fold, instead of re-walking the buffer with up to three post-hoc
List.Exists(...)passes per drain. The maintainer tails its source write-ahead log continuously; every drain folds a batch of projected writes (or aggregation contributions) into a buffer, and the prior code then re-scanned that whole buffer purely to answer "did this batch carry a range operation?". The filter drain did this twice -collected.Exists(w => w.Kind == RangeReconcile)inViewMaintainerGrain.DrainAsyncandcollected.Exists(w => w.Kind == RangeDelete)inApplySurvivorsAsync- and the aggregation drain once,contributions.Exists(c => c.Kind == RangeReconcile)inDrainAggregationAsync. Because the fold loop already visits every element, each classification is now recorded in aboolset inside the fold and the separate.Existspasses are removed, so a freshly drained batch is walked once instead of two or three times. TheRangeDeletefast-path flag is proven safe to compute at fold time: the only mutations between population and its use (ConvertRangeReconcilesToMarkers, which only rewritesRangeReconcileentries toUpsert, andShapeHistoryWritesAsync, which reshapes onlyUpsertrows and passes every non-Upsertthrough untouched) never add or remove aRangeDelete, so the flag equals the removed scan on every path including accumulative views. Behaviour is identical and the public API is unchanged - every edit is to aprivategrain method - and the change is allocation-neutral by construction (the buffer is built identically either way and the removed.Existsscans used a cached static delegate), so this is a pure CPU-cycle trim on a continuously running maintenance path. Measured on the newviewdrainmicrobench suite (BENCH_MICROBENCH_SUITE=viewdrain,MemoryDiagnoser; each lane performs the mandatory fold so the delta is exactly the removed pass(es), andAllocatedis equal within each pair): the two-scan filter classification drops from 2,128.7 ns to 1,080.9 ns at a 256-write batch (-49%) and 7,640.0 ns to 3,622.7 ns at 1,024 (-53%); the one-scan aggregation classification drops from 1,302.9 ns to 957.9 ns at 256 (-26%) and 4,567.4 ns to 3,951.0 ns at 1,024 (-13%). The 495-testOrleans.Lattice.Tests.Viewssuite (including the maintainer drain integration tests) covers the preserved behaviour. (Orleans.Lattice9.5.0)The cross-tree and view coordination barriers now canonicalise their wait/participant sets with a HashSet-plus-in-place-sort helper instead of
Distinct().OrderBy(), removing LINQ's ordering machinery from three warm coordination paths. Three sites froze or stamped a string set withsource.Distinct(StringComparer.Ordinal).OrderBy(v => v, StringComparer.Ordinal).ToList()/.ToArray():ViewCrossTreeCoordinatorGrainandLatticeCrossTreeReceiverGrainfreezing their wait set on the first registration/terminal, andLatticeCrossTreeTxGrainstamping the sorted participant tree-id set onto every sub-saga's terminal for the receiver-side visibility barrier. Per call, that LINQ form allocates anOrderedEnumerablewrapper, a materialised element buffer, a projected key array, and an integer sort-index map, all on top of the result collection. A new internalCanonicalStringSethelper de-duplicates through a singleHashSet<string>pass and sorts the result in place, so only the result (and the transient dedup set) is allocated; the output is byte-for-byte identical (a de-duplicated set rendered in a total ordinal order is deterministic regardless of input order), so the exact-match barrier comparison and the stamped participant set are unchanged. The public API is unchanged - the helper is internal. Measured on the newcrosstreemicrobench suite (BENCH_MICROBENCH_SUITE=crosstree,ShortRun,MemoryDiagnoser; allocation is deterministic and is the win, wall-clock for these sub-microsecond operations sits within run-to-run noise): the wait-set freeze drops from 744 B to 392 B at width 2 (-47%), 1,112 B to 640 B at width 8 (-42%), and 4,960 B to 3,368 B at width 64 (-32%); the participant set drops from 648 B to 328 B at width 2 (-49%), 1,400 B to 960 B at width 8 (-31%), and 5,592 B to 4,032 B at width 64 (-28%). The prior-versus-new equivalence of the helper is held closed by new tests inCanonicalStringSetTests. (Orleans.Lattice9.5.0)Three warm dictionary/set maintenance paths now trim their per-call heap: the tag-index reconcile baseline prune, the aggregation approximate-mode eviction key scan, and small-tag normalisation. (1)
TagIndexReconcileGrain.PruneStaleBaselinesran two LINQ passes overBaselines.Keyson every reconcile begin - a.All(keep.Contains)method-group delegate on the no-drift path and a.Where(t => !keep.Contains(t)).ToList()capturing closure plusWhereiterator on the drift path; it now collects stale tree-ids in a single manual pass over the dictionary, deferring the list until the first stale key so the common no-drift case allocates nothing beyond the membership set. (2)AggregationApplier.LargestSourceKeyiterated the freshly decoded per-mutation inverse-shard map throughmap.Keys, allocating a throwawayKeyCollectionon each approximate-mode min/max eviction; it now iterates the map's entries directly with the struct enumerator, mirroring the siblingWorstKey. (3)LatticeTagIndexContext.NormalizeTagsalways built aHashSet<string>to dedup, even for the handful of tags a tag write or query carries in practice; it now dedups small tag sets (<= 16) with an ordinal linear scan against the accumulating result and keeps the set only for larger inputs (preservingO(n)membership there). All three preserve exact behaviour - first-seen order, per-element validation including duplicates, and the same stale/eviction/dedup results - and the public API is unchanged, since every edit is to a private method. Measured on the newalloctrimsmicrobench suite (BENCH_MICROBENCH_SUITE=alloctrims,ShortRun,MemoryDiagnoser; allocation is deterministic and is the win, wall-clock for these sub-microsecond operations sits within run-to-run noise): baseline prune drops from 960 B to 896 B on the no-drift path (-64 B, the removed delegate) and 1,040 B to 912 B on the drift path (-128 B, the removed closure and iterator); the eviction key scan drops by a size-independent 24 B per call (352 B to 328 B at 4 entries, 632 B to 608 B at 16, 2,144 B to 2,120 B at 64); and small-tag normalisation drops from 560 B to 192 B (-66%). The dedup contract is held closed by a new duplicate-tag regression test inLatticeTagIndexIntegrationTests, and the existing reconcile and aggregation suites cover the other two. (Orleans.Lattice9.5.0)Three structural read fan-out sites now issue one batched multi-get instead of a sequential read per item, cutting per-operation index-tree grain calls on the tag-index AND query, the atomic-write pre-image capture, and materialised-view group materialisation. Each site awaited one index-tree read per sibling tag, per written key, or per view shard, so its caller-visible read round-trips grew linearly with the fan-out width. (1)
LatticeTagIndexContext.QueryAsync(all-tags branch) probed each candidate key's(T-1)sibling-tag membership rows with one sequentialRowLiveAsyncapiece; it now issues a single batchedGetManyAsyncper candidate, collapsing the AND query's index reads fromO(candidates x tags)toO(candidates)while preserving the exact per-row live/flag-mode semantics and duplicate-tag correctness. (2)AtomicActionGrain.RunTreeWriteForwardAsynccaptured each written key's pre-image with oneGetAsyncper key; it now reads the whole step's pre-images with onetree.GetManyAsync. (3)AggregationApplier.MaterialiseInverseAsyncandMaterialiseFoldAsyncgathered a group's shards with onestore.GetAsyncper slot; they now read all slots with onestore.GetManyAsyncon a new internalIAggregationViewStore.GetManyAsync(implemented over the batching view tree and the buffering overlay). The public API is unchanged - the new multi-get lives on the internal view-store interface, and the batched calls reuse the existing publicILattice.GetManyAsyncfan-out semantics (absent/tombstoned keys omitted), so results are byte-for-byte identical. Measured on the newfanoutmicrobench suite (BENCH_MICROBENCH_SUITE=fanout). Exact deterministic round-trip census (BENCH_FANOUT_ROUNDTRIPS_ONLY=true): atomic pre-image 128 keys 128 -> 1 read (0.01x), view inverse 64 shards 64 -> 1 (0.02x), tag-index AND 8 tags x 100 keys 700 -> 100 (0.14x). BDN wall-clock (ShortRun, Width 64): atomic pre-image 38.6 us -> 29.6 us (-23%) and view inverse 41.0 us -> 28.7 us (-30%); the tag-index probe's in-process wall-clock is dominated by the trivial per-read cost and the loss of the baseline's per-candidate short-circuit, so its win is the index-tree grain-call reduction the census measures, not in-process latency. (Orleans.Lattice9.5.0)