Release 2026-08-29: Added to Changed
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at 2026-08-29-1.md, and llms.txt lists every page.Part of Release 2026-08-29, in Changelog.
Added
- Opt-in multi-tenancy for an Orleans.Lattice deployment, single-cluster or multi-cluster. A new companion package family partitions the keyspace into isolated tenants: each tenant's trees live under a reserved
t/{tenant}/prefix that the data path enforces fail-closed, so one tenant can never read, write, or enumerate another's data, and a host that does not reference the packages is byte-for-byte unchanged.Orleans.Lattice.Tenancyadds a durable, CRDT-backed tenant registry with a create / suspend / resume / delete lifecycle (delete cascading every tree the tenant owns), per-tenant quotas (bytes, keys, memory, tree count, and ops per second) with usage metering, demand-apportioned rate limiting, and billable overage accounting, converged or per-region enforcement scopes, optional per-tenant region residency governing where a tenant's data may live and replicate, and per-tenant observability gauges on theorleans.lattice.tenancymeter. It fills in the inert tenancy seams core declares - tenant-scoped tree naming, enumeration pruning, and region visibility - so isolation holds at the data plane, at every catalog choke point, and in region discovery alike; the reservedsys-tenant-*registry trees are read-isolated on the control plane so no data-plane grant can scan them.Orleans.Lattice.Api.TenantAdminexposes the operator control plane as a transport-agnostic facade authorized fail-closed through the shared access gate:ILatticeTenantAdminfor the tenant lifecycle and quota authoring,ILatticeTenantRegionAdminfor region authorization and residency under a two-tier operator / tenant-admin model,ILatticeTenantSelfServicefor the tenant-scoped reads any authenticated caller may make, and an optionalILatticeTenantScopedTreeAdminletting a tenant administer its own trees by unqualified name.Orleans.Lattice.Api.TenantAdmin.Grpcbinds all of it over code-first gRPC with public clients. The external data plane, state, tree-administration, schema, replication, and backup facades resolve a caller-supplied tree name into the caller's tenant namespace and carry the caller's asserted active tenant across the split-head gRPC boundary, so scoping and per-tenant quotas hold end to end rather than only in-process. Tenant awareness reaches operators and agents too: the Explorer gains a signed-in tenant crumb and an operator tenant selector, and the MCP server exposes read-only tenant self-awareness tools (lattice_tenant_current,lattice_tenant_list,lattice_tenant_get), tenant-admin control tools, and a region catalog scoped to the tenant's actionable regions - all keyed purely on whether tenancy is enabled, with no new opt-in flag, so a non-tenancy deployment's Explorer UI and MCP tool surface are byte-for-byte unchanged. Documented in Multi-tenancy, Tenant-Admin API, and Tenant-Admin gRPC, and demonstrated by the MultiTenancy sample. (#1616) (Orleans.Lattice.Tenancy9.4.0,Orleans.Lattice.Api.TenantAdmin9.4.0,Orleans.Lattice.Api.TenantAdmin.Grpc9.4.0,Orleans.Lattice9.4.0,Orleans.Lattice.Auth9.4.0,Orleans.Lattice.Api.Abstractions9.4.0,Orleans.Lattice.Explorer9.4.0,Orleans.Lattice.Api.Mcp9.4.0) - A
catalogmicrobench suite that measures what it costs to page a tree catalog end to end. Selected withBENCH_MICROBENCH_SUITE=catalog(or--suite catalog), it reports two independent things. First an exact, deterministic, host-independent census of the grain round-trips a full pagination needs, sweeping tenant counts 1 / 8 / 64 / 256 against both an unscoped enumeration and a tenant-scoped one, with visibility enforcement on and off, so the benefit of scoping a catalog to a tenant is visible as tenant count rises; it is written to acatalog-roundtrips.jsonsidecar next to the run'sresults.json, andBENCH_CATALOG_ROUNDTRIPS_ONLY=trueprints it and skips the rest. Then a BenchmarkDotNet latency pass comparing the per-entry projection shape against the batched one over identical captured page partitions, so the measured delta is the call shape alone, plus an end-to-end arm driving the realLatticeStateQuery.ListTreesAsyncas a production-code anchor. Every modelled grain read yields to the scheduler rather than completing synchronously, which is a conservative lower bound on a real grain call, and the census counts the registry-internal reads separately from the caller-visible round-trips so the report cannot be read as claiming reads disappeared. Documented in Benchmarks. (#1686) (Orleans.Lattice9.4.0) - Resilient streaming scans for materialised views. New
LatticeViewExtensions.ScanKeysAsyncandScanEntriesAsync(plus the typedTypedLatticeViewExtensions.ScanEntriesAsync<T>) wrapILatticeView.KeysAsync/EntriesAsyncand transparently recover fromOrleans.Runtime.EnumerationAbortedException- raised when the remote enumerator is reclaimed mid-walk by grain deactivation, idle-expiry, or a rebuild's shadow-swap - by resuming from the successor of the last-yielded key, so a long-running view scan has no gaps or duplicates and preserves ordering, the range bounds, the view's reserved floor, and cancellation. They mirror the existingILattice.ScanKeysAsync/ScanEntriesAsyncwrappers and share their bounded reconnect budget (maxAttempts, defaultLatticeExtensions.DefaultScanReconnectAttempts = 8); the rawKeysAsync/EntriesAsyncprimitives are retained for deliberate low-level use. Documented in API surface - Materialised views. (#1650) (Orleans.Lattice9.4.0) - A single public helper for the exclusive upper bound of a key-prefix range.
LatticeKeyRange.PrefixUpperBound(string), in the rootOrleans.Latticenamespace, computes the smallest string that sorts strictly after every key beginning with a given prefix under ordinal comparison - the exclusive end of the half-open[prefix, bound)range a prefix scan walks - and returnsnullwhen the range is unbounded above (an empty prefix, or one consisting solely ofU+FFFF). It replaces nine ad-hoc per-package copies of the same operation, five of which used a rollover-unsafe increment (see the corresponding Changed entry). The bound is computed at the UTF-16 code-unit level, the exact granularity of ordinal comparison, so it is correct even for keys containing surrogate pairs, and it allocates only the single result string (viastring.Create, with no intermediate array, closure, or boxing) or nothing at all on the unbounded path. Covered by a new property-based and boundary test suite. (Orleans.Lattice9.4.0)
Changed
- The remaining observed-remove dot-reconciliation sites now take the allocation-free path when the live side is small, byte comparison vectorises, and dot equality tests the cheap member first. Four related trims to the CRDT primitives, each measured on the
ordedupMemoryDiagnoser microbench suite (OrCrdtReconcileBenchmarks) rather than assumed. First, the small-side guard shipped forOrSetandOrMapin the entry above was missing from six sibling sites -RwSet.LiveDotCountandRwSet.AddObservedTombstones,OrFlag.DisableandOrFlag.LiveEnableCount,RwFlag.EnableandRwFlag.LiveDisableCount- which kept building aHashSet<OrSetDot>sized to a flag's or key's whole tombstone history even when the live side was one or two dots, the dominant case for a repeatedly toggled flag. The flag dedup path (the one shape in the family that scans a list it is simultaneously appending to, so it was measured rather than assumed to behave like its siblings) falls from 913 ns to 192 ns at a 16-dot history, 2,740 ns to 641 ns at 64, and 11,693 ns to 1,854 ns at 256, dropping 1,728 B / 2,152 B / 8,368 B of per-call heap. Second,HashSet<OrSetDot>construction is now presized-and-filled everywhere instead of going through the collection constructor at sixteen sites: that constructor takesIEnumerable<T>, so seeding it from aList<T>boxes the list's struct enumerator and dispatches through an interface per element, measuring roughly 3x slower (325 ns to 995 ns at 16 dots, 5,159 ns to 13,662 ns at 256) and allocating 48 B more - the opposite of what the shorter spelling suggests. The one shape is now behindOrSetDotSet.Build. Third, the byte-sequence comparison that breaks a same-dot value collision inRgaandMvRegisterwas a scalar byte-at-a-time loop while the siblingBoundedRegisteralready used the vectorised span compare; only the sign of the result is load-bearing, so the two are interchangeable, andSequenceCompareTomeasures 2.2 ns / 5.3 ns / 14.2 ns against 7.4 ns / 125.7 ns / 384.4 ns at 16 / 256 / 1024 bytes. Fourth,OrSetDot's synthesised equality comparedReplicaIdbeforeCounter, but dots within one list overwhelmingly share a replica id, so the string comparison almost never discriminated before reaching the singlelongcomparison that does; an explicit counter-firstEqualsmeasures 22.6 ns against 27.6 ns over 16 dots and 63.0 ns against 100.2 ns over 64 (at 256 the delta falls inside the run's noise band). All four are behaviour-preserving, and the sign-equivalence of the vectorised comparison and the equality/hash consistency of the reorderedEqualsare both pinned by new property tests. RestoringBoundedRegister.Clone's deep copy (see the Fixed entry) costs, measured in the same suite, 25.4 ns / 34.5 ns / 47.8 ns against 12.9 ns / 14.3 ns / 11.5 ns for the prior shared-array clone, and 120 B / 168 B / 360 B against 40 B, at 16 / 64 / 256-byte payloads; expressing the copy as a span copy rather thanArray.Clonerecovers roughly 3-4x of that (Array.Clonemeasured 102.9 ns / 125.4 ns / 138.8 ns for the identical allocation). Three further reported candidates were measured and dropped: hoisting aGetAlternateLookupout of twoOrSetloops (no measurable hypothesis - it saves a field read per iteration), makingOrSet.Elements/RwSet.Elements/GSet.Valueslazy (all four call sites materialise the result, so laziness would add an iterator state machine on top of the caller's list rather than remove an allocation), andRga's cache invalidation on append (out of scope this round). (Orleans.Lattice9.4.0) - Catalog paging now reads a page's registry entries in one batched call and probes deletion state as one bounded concurrent wave, instead of two sequential grain round-trips per emitted entry.
LatticeStateQuery.ListTreesAsyncandListTagIndexesAsyncpaid anILatticeRegistry.GetEntryAsyncand anITreeDeletionGrain.IsDeletedAsyncper entry, each awaited in sequence, so a default 100-entry page cost 200 sequential awaits before it could be returned. Both methods now run in two passes: the existing filters and the per-entryIsCatalogEntryVisibleAsynccheck run first, unchanged and in their current positions, accumulating the ids that survive until the page is full; then exactly that page is projected with one newILatticeRegistry.GetEntriesAsynccall plus oneTask.WhenAllfan-out of the deletion probes, both started and awaited inside a singleEnterSystemOriginscope (the marker isRequestContext-backed, so the scope must outlive the fan-out). The filter-first ordering is a correctness requirement, not a preference: the visibility check thins the candidate set, so batching ahead of it would both over-fetch and read entries the per-entry path would have dropped.GetEntriesAsyncis documented as a call-shape optimisation and never an authorization boundary - it grants nothing, applies no filtering of its own, and returns only entries the caller could already have read one at a time. Caller-facing behaviour is unchanged: same entries, same order, sameNextPageToken, same visibility outcomes, across visibility on/off, anonymous subjects, tenant set/unset/default,IncludeSystemTrees, and a supplied page token; concurrency stays bounded byEffectivePageSizebecause only a filled page is ever batched.ListViewsAsyncandListCoveredTreesAsyncare deliberately left as they are - neither has a per-entry registry read to batch, andListViewsAsync's opt-in per-entry cost is aRequestContext-backedViewReadContextscope that cannot be fanned out concurrently without a design change. Measured on the newcatalogmicrobench suite: page projection for 100 entries falls from ~114-133 us to ~46-51 us (roughly 2.6x) and full pagination of an 8192-tree 64-tenant catalog from ~8.8-9.5 ms to ~4.1-4.2 ms (roughly 2.2x), with the grain message count essentially unchanged - the win is sequential depth, not fewer messages - at a bounded cost of about 77 additional bytes per entry for the fan-out arrays. (#1686) (Orleans.Lattice9.4.0,Orleans.Lattice.Api.State9.4.0) - Tree-id enumerations are now scoped at the registry instead of fetching the whole catalog and filtering client-side. Every enumeration choke point read every registered tree id across the grain boundary and then discarded most of it, materialising a second full list to do so; for a tenant caller roughly
(N-1)/Nof the transferred ids were dropped, and the tag-index catalogs fetched the entire cluster catalog only to keep thetag-prefixed ids.ILatticeRegistrygains a prefix-scopedGetAllTreeIdsAsync(string? prefix)overload (the parameterless one delegates to it with anullprefix and is exactly equivalent, so no caller changes behaviour). Because the registry is itself an ordinally-sorted Lattice tree, a prefix occupies one contiguous key range, so the scoped call is a bounded range scan over[prefix, LatticeKeyRange.PrefixUpperBound(prefix))- the registry stops touching pages outside the range at all, which makes this an I/O reduction and not merely a wire-size one, and the benefit grows with the number of tenants. Pushed down at the four sites where it is provably equivalent to the previous shape: the tenant delete cascade (t/{tenant}/, which also lets it dial the registry directly and drops the placeholder "probe tree id" it previously needed to reach the public surface), both tag-index catalogs and the tag-index reconcile trigger (tag-), and the state-API tree catalog (the active tenant's prefix, but only when a non-default tenant is asserted and the request excludes system trees, because the ids a prefix scan skips are exactly the ones theIncludeSystemTreesswitch already drops). The newLatticeTenantTrees.ComposePrefix(TenantId)supplies the tenant prefix. The prefix is a performance hint and never an authorization boundary - it can only ever return a subset of what the caller could already enumerate, the reserved system-tree exclusion is applied inside the scan so a caller-supplied prefix can never widen it, and the tenant-enumeration filter and per-entry visibility checks all still run unchanged. Covered by new registry range-scan regressions (including theU+FFFFunbounded-above case and the parameterless/null-prefix equivalence) and a new pushdown-contract suite asserting narrowing only ever shrinks the result, never narrows a system-tree request, and never narrows under the default tenant. (#1682) (Orleans.Lattice9.4.0,Orleans.Lattice.Api.State9.4.0,Orleans.Lattice.Api.TreeAdmin9.4.0,Orleans.Lattice.Api.TenantAdmin9.4.0) - Backup scheduling now absorbs transient reminder-registry failures and the gRPC backup control handler surfaces them as retryable. The backup scheduler grain's reminder operations (register / read / unregister, reached from
ScheduleRecurringAsync,CancelScheduleAsync, and schedule-status reads) are now wrapped in a small bounded retry with backoff that classifies a transient reminder-registry failure - the reminder service still initializing, or a reminder-table read/write timing out under load - and rides it out instead of propagating it to the caller; a genuinely stuck reminder service still surfaces its original exception once the budget is exhausted. TheOrleans.Lattice.Api.Backup.Grpccontrol handler's catch-all now distinguishes those transient/infrastructure failures (mapped to the retryableUnavailablestatus) from genuine internal faults (Internal), and attaches a correlation id to both opaque surfaced messages so an operator can tie a client-side failure to the fully-logged server exception without the message leaking internal detail. This removes an intermittent Coverage-pipeline flake inLatticeBackupGrpcClientE2ETests(an opaqueInternal"The backup control-API request failed" at theScheduleBackupAsyncstep) at its source and hardens production scheduling. (#1654) (Orleans.Lattice.Backup9.4.0,Orleans.Lattice.Api.Backup.Grpc9.4.0) - Internal long-running scans now use the resilient reconnect wrappers. Remaining production scan sites that walked a tree or materialised view with the raw
EntriesAsync/EntriesAsync<T>/KeysAsyncprimitives have been converted to the resilientScanEntriesAsync/ScanEntriesAsync<T>wrappers appropriate to their receiver:ILattice-backed scans (schema remediation and compliance diffing, membership group enumeration, and tenant usage / overage metering) now useLatticeExtensions/TypedLatticeExtensions, andILatticeView-backed scans (backup catalog-index queries) now use the newLatticeViewExtensions. The reserved internalISystemLatticesurface gains a matching internalScanEntriesAsyncwrapper, now used by the queue cold-start backlog scan and the view-rebuild digest walk. These long scans no longer surfaceOrleans.Runtime.EnumerationAbortedExceptionwhen the remote enumerator is reclaimed mid-walk by grain deactivation, idle-expiry, silo failover, or a view shadow-swap; they transparently reconnect and resume from the successor of the last-yielded key with no gaps or duplicates. This is a read-side reconnect hardening only and does not alter any apply-side atomic-visibility guarantee. (#1651) (Orleans.Lattice9.4.0,Orleans.Lattice.Schema9.4.0,Orleans.Lattice.Membership9.4.0,Orleans.Lattice.Tenancy9.4.0,Orleans.Lattice.Api.Backup9.4.0) - The key-prefix upper-bound operation is now defined once and shared. Nine call sites across the package family each carried their own "advance the prefix to the exclusive scan bound" helper, several using the naive
chars[^1]++form that wraps a trailingU+FFFFtoU+0000and inverts the half-open scan range. The reachable and latent rollover bugs in the auth-admin, authorization-policy, membership, and repo-context copies were fixed directly in their own changes (see the Fixed entries above); this change removes the remaining duplication by routing every in-family call site - including the still-unconvertedOrleans.Latticetag-index and queue cores and theOrleans.Lattice.Schemadead-letter key, whose latent copies (guarded only by an undocumented separator-terminated-prefix invariant) it hardens by construction - through the new publicLatticeKeyRange.PrefixUpperBound, so the rollover-safe algorithm has a single definition; the two internal helpers that must return a non-null bound (SchemaDeadLetterKey.PrefixEnd,LatticeQueueCore.PrefixEnd) now assert their separator-terminated invariant explicitly. The two already-correct copies outside the core dependency graph (the ExplorerDataReaderand the RepoContext MCP portability layer) are deliberately left in place. No public behaviour changes. (Orleans.Lattice9.4.0,Orleans.Lattice.Membership9.4.0,Orleans.Lattice.Auth9.4.0,Orleans.Lattice.Schema9.4.0,Orleans.Lattice.Backup9.4.0,Orleans.Lattice.Api.Auth9.4.0) - The observed-remove CRDT read and remove hot paths no longer allocate a tombstone-sized
HashSetwhen the live side is small. Three dot-reconciliation sites -OrSet's per-keyLiveDotCount(behind every OR-SetIsEmpty/Count/Elements/Containsread),OrSet.Remove's tombstone-dedup, andOrMap's per-keyLiveEntryCount(behind every OR-MapIsEmpty/Count/Contains/Keysread) - built aHashSet<OrSetDot>sized to a key's whole observed-remove history on every call once that history crossed the linear-scan threshold, even though the live side (an element's surviving add dots, or a key's live entries) is 1-2 in the overwhelming common case of a churned key with a long tombstone history. Each guard now also diverts to the existing allocation-free linear membership scan when the live side is at or below the threshold, exactly as the sibling merge paths (OrSet.MergeMap,OrMap's merge reconciliation) already guarded both sides; the scan isO(liveCount * tombCount), the same asymptotic as building the tomb-sized set whenliveCountis a small constant, but heap-free. Results are identical and no public API changes. Verified by the newordedupMemoryDiagnoser microbench suite (OrCrdtReconcileBenchmarks), which measures the read/remove reconciliation dropping from 640 B - 14,704 B per call (scaling with tombstone-history length) to 0 B across tombstone histories of 16 - 512 dots. (Orleans.Lattice9.4.0) - Every OR-family merge fold now takes the allocation-free linear path whenever the incoming delta is small, whatever the accumulated dot list has grown to.
OrSet.MergeMap,RwSet.MergeMap, and the two flags' (OrFlag/RwFlag)UnionInto/UnionDotsunion helpers all fold an incoming dot list into an accumulated one, but took the allocation-free linear path only when both sides were at or below the linear-scan threshold (4). A steady-state 1-2-dot replication delta folded into a key or flag that had accumulated a long dot history therefore seeded aHashSet<OrSetDot>from the whole accumulated list on every merge, hashing every accumulated dot's replica id to answer a couple of membership questions. The guard now diverts on the incoming side alone. Because the linear branch appends at most threshold-many dots, the scan staysO(delta * accumulated)- bounded byO(4 * accumulated), the same asymptotic as building the accumulated-sized set - but allocates nothing. This completes the small-side trim made on the add-winsOrSet/OrMapread and remove paths and then extended to the remove-wins and flag read paths, leaving no OR-family reconciliation site that sizes a transient set to a history it only probes a couple of times. Results are identical and there are no public API changes. Verified by the newmergefoldMemoryDiagnoser microbench suite (CrdtMergeFoldBenchmarks), whose baseline reproduces the set-building shape shipped today so the delta is attributable to this change alone: the fold drops from 760 B / 2,104 B / 8,320 B / 31,000 B per merge to 0 B across accumulated dot histories of 16 / 64 / 256 / 1,024 dots, and stays 14x - 25x faster throughout - including at 1,024 dots, far past any realistic key, confirming the linear branch does not degrade at length. (Orleans.Lattice9.4.0) - The cross-cluster replication shipper now learns of a source-tree identity swap from a pushed event instead of polling the registry every tick. When a shadow-cutover restore, resize, or reshard repoints a logical source tree's registry alias to a new physical WAL, the tree registry now fires a new core extensibility hook (
ITreeAliasObserver, carrying aTreeAliasChange) once per effective-physical-id change, and the replication package fans it to the affected per-(tree, peer)shipper grains, which rebind immediately and re-ship from the new physical log start. Combined with caching the peer wire-version and shared-dictionary negotiation results and recomputing them only on an advertised-capability change, this removes all three steady-state per-tick_lattice_treesregistry/metadata resolutions an idle shipper previously performed, so a quiet replication link does zero idle registry reads. Because the notification is pushed synchronously with the swap (and reaches the shipper even while the inter-site edge is partitioned), it also closes the former detection-window sharp edge in which a still-live shipper could keep tailing the retired identity's WAL and ship keys confined to the abandoned identity - keys plain last-writer-wins shipping never retracts. A new coarse backstop,LatticeReplicationOptions.ShipSourceIdentityBackstopInterval(default 30 s), re-resolves only as a safety net against a missed notification. Fully backward compatible - no wire-format change - and subsumes the earlier idle-read measurement. Documented in Source-identity rebind and Tree alias observers. (#1665) (Orleans.Lattice.Replication9.4.0,Orleans.Lattice9.4.0) - The materialised-view maintenance hot path no longer allocates a throwaway UTF-8
byte[]per hashed key. Three sites on the view write / rebuild path each hashed a key or payload by first heap-allocating the encode viaEncoding.UTF8.GetBytes(string); all three now encode into a stack (or, for a long input, a pooled) UTF-8 span before hashing - the same idiomLatticeShardingandShardMap.GetVirtualSlotalready use for core shard routing. The hash inputs are byte-for-byte identical, so slot indices, operation ids, and view digests are unchanged and the change is wire- and format-compatible. (1)AggregationRowCodec.Slot- the per-contribution accumulator-shard routing hash (XxHash32), run once or twice per source mutation feeding a count / sum view - drops from 72 B to zero allocation (30.7 ns to 17.5 ns). (2)AggregationApplier.OperationId- the per-contribution atomic-flip idempotency-id hash (XxHash64) - removes the 96 B encodebyte[](376 B to 280 B; the residual 280 B is the interpolated id-payload string, unchanged and common to both shapes) at 144.9 ns to 103.8 ns. (3)ViewMaintainerGrain.ComputeTreeDigestAsync- the per-entry key encode inside the order-independent view-tree drift digest (XxHash128) - reuses one pooled buffer across the whole scan (astackallocspan cannot be hoisted across theawait foreach, so the buffer is a single rented array returned in afinally), collapsing N per-entry allocations to one rental: at 1,024 entries it falls from 41,560 B to a flat 600 B and 60.5 us to 48.7 us, and the allocation is now constant in entry count rather than linear. All three are covered by the existingViewstest suite (behaviour is identical) and measured on a newhashallocMemoryDiagnoser microbench suite (HashingAllocationBenchmarks, selected withBENCH_MICROBENCH_SUITE=hashalloc). (Orleans.Lattice9.4.1)