Table of Contents

Release 2026-09-09: Security

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at 2026-09-09-1.md, and llms.txt lists every page.

Part of Release 2026-09-09, in Changelog.

Security

  • A gRPC tree-administration caller could act on a second tree it held no grant on, because only the first tree id in the request was ever authorized. LatticeTreeAdminApiGrpcAuthInterceptor extracts one target tree id per call and hands that single id to the registered ILatticeTreeAdminApiAuthorizer, which is correct for the great majority of the surface because the great majority of its requests name exactly one tree. Two do not. SetTreeAlias carries a PhysicalTreeId alongside its TreeId, and SnapshotTree carries a DestinationTreeId; in both cases the facade acts on that second id just as materially as on the first - SetTreeAliasAsync(req.TreeId, req.PhysicalTreeId, ...) re-points a logical name at the physical tree named second, and SnapshotTreeAsync(req.TreeId, req.DestinationTreeId, ...) writes a snapshot into the tree named second. The interceptor never showed either to the authorizer, so a caller granted administration of a single tree it legitimately owned could name any other tree in the estate as the second id and have the call admitted on the strength of the grant it held on the first (CWE-863, an incorrect authorization decision rather than a missing one - the check ran, it just adjudicated the wrong subject). The two RPCs are the worst possible pair to have missed: one re-points a name the victim's own clients resolve through, and the other writes into the victim's tree. Authorization now runs once per distinct tree a request acts on. The interceptor extracts a secondary target where one exists and authorizes it through the same seam as the primary, so an authorizer written against the existing contract needs no change and sees the second tree exactly as it already sees the first, with the same PermissionDenied mapping and the same cancellation semantics. A call naming one tree, or naming the same tree twice, still costs exactly one authorizer round trip, which is asserted directly so the common path cannot silently regress into two. The extraction helper is internal static and DescribeCall keeps its existing signature, so no public API signature changed. (Orleans.Lattice.Api.TreeAdmin.Grpc)

  • Suspending a tenant no longer stops its writes locally while its peers keep applying them. LatticeTenantPolicyEngine.ValidateActiveTenant refuses any tenant whose status is not Active, so on the authoring path suspension does what an operator expects and the tenant's writes stop. The inbound replication path asked a different and strictly weaker question: ReplicationTenantIsolationGate decided admission on whether the tenant exists and whether it resides in this region, and never on its status. Suspension was therefore a one-sided control, which is the failure mode that makes it dangerous rather than merely incomplete - an operator suspending a tenant in response to an incident, a compliance hold, or non-payment observes local writes stop and reasonably concludes the tenant is contained, while every other region in an active-active estate continues to accept and apply that tenant's mutations into the very trees the suspension was meant to freeze. The gap widens with the size of the estate and is invisible from the region the operator is working in. Both halves of the gate now enforce status. The compiled-snapshot fast path checks the status carried on the snapshot, and the registry fallback moves from an existence probe to a full record read so it can check the record's own status; a suspended tenant is refused with a new, distinct RejectSuspendedTenant decision rather than being folded into the existing unknown-tenant or wrong-region arms, so an operator reading a dead letter or a metric sees why the entry was refused instead of being told the tenant does not exist. The distinction is carried through to the dead-letter reason and the metric outcome tag, both of which gain a corresponding value. The registry fallback's cost is unchanged in the shape that matters - it was already one grain call, and reading the record rather than probing for it returns the status in the same round trip. Platform and legacy trees are untouched, and both new checks are additive enum members and constants on types that gate nothing publicly, so no public API signature changed. (Orleans.Lattice.Tenancy, Orleans.Lattice.Replication)

  • A change-feed subscription to a materialised view is now authorized against the view's source tree, not against the view id the caller supplied. The state API's read facade already states this rule and implements it: reading a view-* tree opens a ViewReadContext scope that makes the data-plane access gate bypass itself, so the readability of the source is the only authorization boundary a view read has, and GetEntry / ScanEntries / GetEntryHistory each resolve the source and gate on it. The change feed did not. LatticeStateObserver passed the caller-supplied tree id straight to ResolveTreeReadAccessAsync, so a subscription to view-orders was adjudicated against a grant on view-orders - a question whose answer cannot protect orders. Nothing downstream could compensate, because the feed tails the write-ahead log directly rather than flowing through the gated ILattice surface; the file's own comment already identified it as the sole point that must honour the read policy for the live stream. The consequence is a live, unpruned stream of every key, change kind, HLC timestamp and category of the source data to a subject with no grant on that source (CWE-863). It did not require a hand-authored view grant to reach: view- is deliberately not a reserved prefix - only sys-auth- is - so a whole-estate Tree:* allow reaches a view tree, and because the specific deny on the source was never the tree being evaluated, an explicit Deny on orders was bypassed rather than merely absent. A subscription now resolves a view to its source before deciding, through the local view catalog first and the durable cluster registry second, and fails closed with the same not-found result every other read surface returns when the source cannot be resolved - an unresolvable view is refused, never tailed ungated. Resolving the source also yields the source's key filter, so a prefix-granted subscriber is pruned to the keys it may read on the source rather than being handed the whole view. The resolution is lifted into one shared internal seam that the read facade now calls too, so the two surfaces cannot drift apart again; an ordinary tree is still decided against itself, which is asserted directly. Every changed type is internal, so no public API signature changed. (Orleans.Lattice.Api.State)

  • An unauthenticated caller can no longer grow silo memory without limit by presenting a stream of invalid credentials. MembershipResolutionCache memoizes resolved subjects in a ConcurrentDictionary keyed by the caller-supplied credential token, so that a warm caller is served without re-authenticating or re-reading the directory. The insert was unconditional. MembershipContext.ResolveUncachedAsync returns the well-known anonymous subject both when no registered authenticator claims a credential and when the selected authenticator rejects it, and that verdict carries no token expiry, so it was cached for the full ResolutionCacheTtl - five minutes by default, and the cache is on by default. Every distinct token a caller presented therefore minted a permanent map entry keyed on bytes the caller chose and no authenticator ever accepted, before proving anything at all, making the cheapest possible request the one that consumes memory (CWE-770 / CWE-400). Nothing reclaimed the entries: an entry expires only logically, compared against ExpiresAt on read, and the map is cleared only when a sys-membership- tree happens to mutate, so on a silo whose membership is stable the map never shrank. The fix removes the amplifier at its root by making the cache positive-only: an anonymous verdict is never stored, so the pre-authentication path can no longer insert at all, and re-resolving one costs nothing that matters because a rejected credential returns before the directory is ever consulted. The remaining positive cache is bounded in its own right, because a legitimate population that rotates its tokens also mints a fresh key per token; admitting a new key at the ceiling first sweeps whatever has already expired, so a full-but-stale cache reclaims itself rather than going permanently cold, and refreshing a key already present is always allowed so an in-use entry is never dropped from under the warm path. A refused insert costs one re-authentication and never changes an answer, since a miss re-runs the authoritative resolution. MembershipResolutionCache is internal sealed and ILatticeMembershipContext is untouched, so no public API signature changed. (Orleans.Lattice.Membership)

  • The EnvVarCredentialAuthorizer failed-attempt lockout can no longer be bypassed, or turned into an unbounded memory leak, by varying the case of a username. The authorizer resolves a credential by env var name, <prefix><username>, through Environment.GetEnvironmentVariable, which on Windows is case-insensitive - so admin, ADMIN and aDmIn all resolve to one and the same configured password hash. The per-username failed-attempt map that throttles guessing was keyed StringComparer.Ordinal, so those spellings were three different keys with three independent AttemptRecords, each carrying its own MaxFailedAttempts budget and its own LockedUntil. The lockout therefore did not lock anything: for a username with c cased letters there are 2^c spellings of the same credential, so the effective budget was MaxFailedAttempts * 2^c (160 guesses for a five-letter name at the default of 5), and because an attacker can rotate to a fresh spelling the instant one locks, guessing was in practice continuous and unthrottled rather than merely widened. That defeats a control the class documents as load-bearing precisely because "username/password credentials are low-entropy and guessable", and a correct password is accepted under any spelling, so nothing about the bypass costs the attacker anything. The same mismatch broke a second, separately-documented invariant: the attempt record is created before the password is verified, and the unknown-user branch deliberately declines to track a caller it cannot resolve specifically so that a stream of invented usernames cannot grow the map (CWE-770). Because every case variant of a known username does resolve, an unauthenticated caller who knows one valid username - admin, root - could insert a new permanent record per spelling without ever presenting a valid password, growing the map by 2^c entries and falsifying the code's own claim that "its working set is bounded by the number of configured users". One root cause, two impacts: CWE-307 and CWE-770. The attempt map is now keyed by a comparer that matches the identity semantics of the lookup it guards - OrdinalIgnoreCase where the platform resolves environment variables case-insensitively, Ordinal elsewhere - so every spelling of one credential shares one budget and one record on Windows, while a genuinely case-sensitive host keeps distinct names independent and is byte-for-byte unchanged. A legitimate Windows login that uses a different spelling from the configured variable still succeeds. The change is to a private field and a new internal constructor overload that lets the case-insensitive shape be exercised deterministically on any host, so no public API signature changed. (Orleans.Lattice.Api.State.Grpc)

  • The schema policy and version caches are now bounded, so a caller can no longer grow silo memory without limit by naming distinct trees. LatticeSchemaPolicyProvider and LatticeSchemaVersionProvider each memoize per-tree resolution in a ConcurrentDictionary keyed by tree id, and both cache a null sentinel on a miss so that an ungoverned or unversioned tree does not re-hit the store on every write. The key is caller-supplied and the map had no size bound and no expiry: the only eviction was Invalidate and the mutation observer, which fire solely for a write to the corresponding sys-schema- tree keyed by that exact id. Deleting a tree evicted nothing, and resolving a tree that never existed still left a permanent entry behind, so the map grew monotonically with the cumulative set of every distinct tree id ever resolved on that silo and never shrank (CWE-770). The compliance-scan path is the cheapest trigger, because it resolves and caches a sentinel for any id at all without requiring the tree to exist, though it is authorization-gated; the write path is reachable at lower privilege but each new id there also implies a durable tree. Both caches now stop admitting new keys at a bound, while a refresh of a key already present is always allowed so an in-use entry is never evicted out from under the hot path. The bound cannot change an answer: a refused insert only costs a store round-trip on the next miss, and a governed tree still resolves its policy correctly with the cache full, which is asserted directly. Both providers are internal sealed and ILatticeSchemaPolicyProvider / ILatticeSchemaVersionProvider are untouched, so no public API signature changed. (Orleans.Lattice.Schema)

  • The replication receive gate and the per-tree option resolvers no longer retain an entry for every tree id they have ever seen. ReplicationReceiveGate caches the inbound per-tree pause answer for a deliberately short window so the apply hot path pays at most one grain call per tree per window. The window was enforced only on the value - each entry carries an ExpiresAtUtc that is compared on read - and never on the map, so an expired entry was re-read, found stale, and overwritten, but an entry for a tree that stopped being written was never removed at all. The dictionary therefore grew by one permanent entry per distinct tree id ever applied, including trees since deleted, and a peer shipping an unlimited stream of distinct tree ids could grow it without bound on the inbound replication path (CWE-770). ConfiguredLatticeMergeModeResolver and ConfiguredLatticeOriginClusterIdResolver had the same shape without even the value-level expiry: both memoize a per-tree lookup with GetOrAdd keyed by a caller-supplied tree id, both cache a result for an id with no replication configuration at all, and both are cleared only when the options monitor fires a change. All three are now bounded. The receive gate first sweeps entries whose window has already closed and admits the new key if that freed room, falling back to skipping the insert only when the map is genuinely full of live entries; the two resolvers refuse a new key at the bound while leaving the hit path a single dictionary read, as documented. In every case a refused insert costs one extra grain call or options read and never changes the answer - in particular the gate still reports an engaged fence with the cache full, so a bound can never let a deferred entry apply during a pause, which is asserted directly. All three types are internal sealed, so no public API signature changed. (Orleans.Lattice.Replication)

  • A tenant could read without limit, create trees without limit, degrade its neighbours, and address a shared tree outside its own namespace; six gaps in tenant isolation are closed. They are reported together because they are one story - the quota, isolation and noisy-neighbour strands of multi-tenancy each had a seam that enforcement had never been wired to - and because five of the six share a single root cause: a mechanism that was registered, documented and configurable, but that nothing on the live path actually consulted. (1) The entire read plane sat outside quota and rate admission. MaxOpsPerSecond was charged on writes only, so a tenant could saturate a shared silo with reads that cost it nothing, starving co-resident tenants of the very system trees they depend on - the ceiling whose purpose is to bound a noisy neighbour was bounding half the traffic. Reads are now charged, and covering them properly meant covering three distinct enforcement shapes rather than the two that are obvious: the local per-operation gate call, three sites that delegate to the static enforcement helper instead, and the range-read path, which classifies itself as RangeRead rather than Read and so escapes any predicate that matches only the latter. The charge is taken strictly after the gate authorizes the read and only when it is allowed, never before - the tenant is a caller assertion that only the gate validates, so charging first would let an unauthorized caller drain a named victim's budget and then read the victim's usage and ceiling back out of the refusal, turning the quota system into the denial-of-service vector it exists to prevent. The footprint dimensions are deliberately not applied to reads, because refusing reads to a tenant over its storage quota would trap it, unable to read the data it must delete to get back under. (2) MaxTreeCount never bound at the point of creation. It was consulted only by the tenant-scoped administration facade, one layer above the facade that actually creates trees - and that inner facade is itself tenant-aware and directly reachable, so a tenant reaching it created trees inside its own namespace with no ceiling applied at all. Enforcement moves to the inner seam every create funnels through, and is deliberately removed from the outer facade rather than kept in both: admission consumes a rate token, so a duplicated check bills a single create twice. The distinction generalises - a stateless shape check may be duplicated for defence in depth, a budget-consuming one may not. (3) The overage meter was registered but never driven, so a tenant's quota breach was recorded nowhere and the billing-ready overage signal was always zero. (4) The replication isolation gate performed an uncached tenant-registry grain read on every inbound apply, so one tenant's replication stream could degrade every other tenant's; it now answers from the compiled policy snapshot and falls back to the registry only on a miss, still fail-closed. (5) WarmUpAsync was ungated, exposing an unauthorized, unmetered fan-out of real activation work across every shard of any named tree. (6) A tenant could escape its own namespace entirely through the reserved sys- prefix. Tenant scoping composes the active tenant into a tree name, and deliberately passes an already-qualified name through uncomposed so it is never double-composed - correct for the first-party add-ons that own the sys- system-data trees, but it meant a confined tenant naming one had the id returned uncomposed and therefore global. Such a tree sits outside the t/{tenant}/ prefix that per-tenant tree-count and footprint accounting enumerates, so it is invisible to both ceilings meant to bound it; it is shared with every other tenant that picks the same name, which is a direct cross-tenant read and write channel between parties that should not be able to observe one another at all; and it can collide with the identity, authorization and tenant-registry stores that namespace holds. A non-default tenant addressing it outside a system-origin scope is now refused. The guard sits on the single tenant-resolution seam every facade passes through rather than in each facade, so a facade added later inherits it; an earlier per-facade guard was written and then deleted precisely because the seam fires first and two copies would drift. Both hard constraints are preserved throughout: a host that has not registered the tenancy add-on resolves the inert null controller and the default tenant, so every new path short-circuits and behaviour is byte-for-byte what it was, and the admit path allocates nothing when tenancy is enabled - only a refusal allocates, because only a refusal throws. An adversarial review of the six fixes then found eight further gaps, all closed here rather than deferred, because most of them are the same defect class the six were: an enforcement point that exists but is not reached. The read charge fired on turns the gate never adjudicated. The enforcement helpers return a successfully-completed result both when the gate ran and allowed and when the gate was skipped entirely, and those two are indistinguishable to the caller; the charge tested only for a system-origin turn, which is strictly narrower than the skip condition, so a read under an authorised materialised-view scope was billed to a tenant nobody had validated. That is externally reachable through the State API, and it is precisely the attack the authorize-then-account ordering exists to prevent, arrived at from the other side. The charge now skips on exactly the bypass condition, and deliberately not on the default null gate: a null gate is a permissive policy, and "every request is allowed" is a decision taken about an ordinary tenant caller who must still be metered, whereas a bypassed gate means the turn is infrastructure and not a tenant request at all. Conflating the two would have turned an unrelated composition choice - tenancy registered without an auth add-on - into a silent quota bypass. The tree-count ceiling was still evadable, because the create check read a metering-cadence snapshot that fails open until a tenant's first sample lands and lags by up to one interval thereafter, so a burst of concurrent creates all saw the same stale count. The count is now read authoritatively from the tree registry at the decision point, and only when a ceiling actually exists, so the common unbounded case still costs nothing. Two whole-keyspace read paths were outside the new charge entirely: a backup capture, which is the largest tenant-triggerable read the platform offers, and the snapshot cursor, which reads snapshot leaf grains directly and so never crosses the seam that charges the live cursors. Both are now charged at their own seam, after authorization; the operation classifier that decides what counts as a read was also an exact match against a [Flags] enum, silently classifying every composite operation as a non-read, and is now a mask test that agrees with the tenancy gate's own read mask. One tenant could stop quota enforcement for every other tenant, because an exception raised while accruing its overage propagated out of the shared metering pass and skipped every tenant enumerated after it; since admission is driven by those published samples, a tenant able to induce contention on its own overage record could stall its neighbours' ceilings. The pass now contains a per-tenant failure and continues. The replication snapshot fast path could keep a deleted tenant admitted indefinitely, because rebuild failures are logged and swallowed with the previous snapshot left in effect, so a deny had silently become an allow; the fast path is now trusted only while the snapshot is authoritative, falling back to the registry whenever a rebuild is outstanding or failing. Bounding it by snapshot age was considered and rejected: rebuilds are mutation-driven rather than periodic, so on a quiet estate an hours-old snapshot is exactly correct and an age bound would have reintroduced the per-apply registry call that fix (4) removed. The t/ structural namespace had the same escape shape as sys-, but only for a malformed id carrying no tenant segment, which resolves to platform ownership and is then allowed unconditionally; that one shape is now refused. A well-formed foreign id is deliberately not refused at this seam, because cross-tenant grants are real and only the gate can adjudicate them - refusing here would have broken granting altogether, which a pre-existing test caught. Finally, the outer tenant facade still injected and null-checked an admission controller it no longer reads, which is dead security config: a reader sees enforcement wired in and reasonably assumes this layer enforces. It is removed, along with the class documentation that still described the check. (Orleans.Lattice, Orleans.Lattice.Tenancy, Orleans.Lattice.Api.TreeAdmin, Orleans.Lattice.Api.TenantAdmin, Orleans.Lattice.Backup)

  • A replication peer can no longer name the tree a coordinated restore overwrites. The inbound cross-cluster saga control channel authorizes the origin cluster and nothing else: LatticeSagaGrpcService gates every call through ISagaPeerAuthorizer, which answers only "is this a configured peer". The tree the restore lands in travels in the same message as SagaControlRequest.TargetTree, and on the single-tree path RestoreParticipant passed it straight into LatticeRestoreRequest.TargetTreeId - a field whose documented purpose is to redirect a restore into a tree other than the one that was captured - and then fenced it and performed the atomic alias swap. Nothing between the peer and the swap re-derived that tree against local replication enrollment, so authenticating as any configured peer was sufficient to overwrite any tree on the receiving cluster, including a tree deliberately kept cluster-local, another tenant's tree, and the reserved sys- authorization and membership trees that carry the cluster's own policy and identity state. The only namespace guard on the path, BackupConstants.ThrowIfReservedTree, blocks the sys-backup- prefix alone and does not cover them, and the backup access gate is a no-op in any host that does not wire the authorization add-on. Rolling a cluster's authorization tree back to an earlier captured state is a privilege-escalation primitive, not merely a data-integrity one, which is what lifts this above the "a mesh peer is already trusted" reading. The gate was missing only here, and every sibling path already had it: the coordinator refuses to promote a restore to a saga at all unless the target is replicated, the set restore path filters every member through IReplicatedTreeMembership.IsReplicated on prepare, commit and abort alike, the push data path classifies every inbound entry against enrollment, and the snapshot export path checks it before serving. The single-tree path now re-derives the same seam it already had injected, refusing a target this cluster does not replicate before anything is built, swapped, or reverted. The check is applied independently in all three phases rather than only in prepare, so a replayed or forged commit cannot reach the alias swap on the strength of never having prepared, and an abort cannot revert an alias on a tree the prepare gate would have refused; a refusal votes abort with the tree named, lifts any defensive fence, and is counted under a new not-replicated reason on the participant vote, commit and abort counters, so a peer probing outside its enrollment is visible rather than silent. Because a legitimate single-tree saga can only ever name a replicated tree, no valid coordinated restore changes behaviour, and a host that does not wire the membership seam is unaffected. RestoreParticipant is internal sealed and no public API signature changed. (Orleans.Lattice.Replication)

  • Every server-side gRPC authorization interceptor now reaches an allow/deny decision on all four call shapes, closing a latent authorization bypass on streaming RPCs. Grpc.Core.Interceptors.Interceptor implements each handler as a pass-through to the continuation, so a shape an interceptor does not override is not gated leniently - it is not gated at all. All ten Lattice auth interceptors overrode UnaryServerHandler, four also overrode ServerStreamingServerHandler, and none overrode ClientStreamingServerHandler or DuplexStreamingServerHandler. Today's surface is not exploitable, because the only non-unary methods that exist are server-streaming ones on the four services that already covered that shape, so this is a repair to a latent hole rather than a live one. It is still a direct violation of the repository invariant that a security gate must reach an explicit deny/allow decision on every branch with deny as the default arm, and the failure mode it sets up is the worst kind: nothing breaks when the gap is introduced, and it becomes live the moment someone adds a streaming RPC to an already-gated service - the change least likely to prompt a review of the interceptor, on a service whose authorization everyone reasonably assumes is settled. The replication interceptor's own comment claimed its service-prefix scoping was written "to keep the interceptor resilient to future RPC additions on the same services", which is precisely the resilience the missing overrides denied it. All ten interceptors now override all four handlers, enforcing through the same authorizer with the same deny semantics; a streaming call is decided before the first message is read, and the existing precedent of scoping streaming handlers by service prefix alone (the unauthenticated-method exemption covers a unary-only RPC) is preserved. All ten types are internal sealed, so no public surface changes. A new reflection-driven contract test runs per package and holds any interceptor that gates one server call shape to gating all four, so a future interceptor is audited automatically and a regression fails build-and-test instead of shipping. (Orleans.Lattice.Api.Auth.Grpc, Orleans.Lattice.Api.Backup.Grpc, Orleans.Lattice.Api.Data.Grpc, Orleans.Lattice.Api.Replication.Grpc, Orleans.Lattice.Api.Schema.Grpc, Orleans.Lattice.Api.State.Grpc, Orleans.Lattice.Api.Telemetry.Grpc, Orleans.Lattice.Api.TenantAdmin.Grpc, Orleans.Lattice.Api.TreeAdmin.Grpc, Orleans.Lattice.Replication.Grpc)