Orleans.Lattice.Tenancy
This page documents Orleans.Lattice.Tenancy 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.Opt-in multi-tenancy for Orleans.Lattice: complete tenant isolation and runtime resource governance, layered over a small number of generic core seams.
What is it?
Orleans.Lattice.Tenancy makes a tenant a first-class citizen of a Lattice
deployment. Tenants that share a cluster (or a set of replicated clusters) are:
- Completely isolated in what they can access - a subject's active tenant may only read, write, enumerate, administer, back up, restore, or replicate trees inside its own tenant's namespace; and
- Governed at runtime - each tenant carries aggregate quotas across all of its trees (durable bytes, live keys, resident memory, tree count, request rate) and an optional burst allowance whose overage is explicitly metered, both adjustable at runtime through the control plane. Its tenant record also carries a physical placement binding (a dedicated WAL provider and/or a silo placement filter - the filter is recorded but not acted on); the control plane creates every tenant on the shared placement, and a binding is immutable in effect once the tenant's trees are placed.
It is a companion package, following the same model as lattice.auth and
lattice.schema: the tenancy logic (registry, compiled quota/isolation policy,
admission metering, tenant-aware enforcement, overage metering, physical-placement
binding) lives here, and core gains only thin, generic null seams. When this
package is not registered those seams resolve to null implementations and core
keeps its exact current path, so a tree that never opts in pays zero overhead.
The tenant lifecycle and governance control plane ships as a sibling facade
family - see Orleans.Lattice.Api.TenantAdmin
and its gRPC binding.
Core properties
- Opt-in and non-destructive. Enabling the feature on an existing cluster
preserves all existing configuration and data. Every pre-tenancy tree keeps its
bare, unsegmented id and is adopted into a reserved
defaulttenant that owns the entire legacy namespace. Existing per-tree options, registry entries, aliases, shard maps, and data are untouched. - Zero-cost when absent. With the package unregistered, tree-id derivation, enumeration, and the access gate behave byte-for-byte as before, because the core seams resolve to their null defaults.
- Fail-closed isolation. The tenant boundary is a hard default-deny wall. Every access path - data plane, enumeration/catalog, control plane, backup/restore, replication apply, observability, and Explorer - is tenant-scoped. Cross-tenant access exists only where an explicit grant or a platform-operator scope authorizes it.
- Hard dependency on identity. The tenant is a membership attribute, so
AddLatticeTenancyis guarded: it throws anInvalidOperationExceptionat registration time - not at silo start, and never as a silent downgrade to an unenforced state - unlessAddLattice,AddLatticeMembership, andAddLatticeAuthhave all already run on the same builder. - Coordination-free multi-cluster. Tenant definitions and usage are convergent
CRDT state - a registry record merges field by field, and usage enforcement reads a
convergent sum (no locks, no consensus) with bounded, quantified overshoot - so they
converge across every cluster the
sys-tenant-*trees replicate to. The package does not enroll those trees for replication itself, andReplicateLatticeSystemTreescovers only the membership and authorization trees.
Quick start
Register the package on the silo, alongside the auth and membership packages it depends on:
using Orleans.Lattice.Tenancy;
siloBuilder.AddLatticeTenancy(options =>
{
// Seed the reserved `default` tenant (unbounded quota) so an existing
// cluster's legacy trees are adopted non-destructively. Default: true.
options.SeedDefaultTenant = true;
// Materialise the durable tenant-definition history view so tenant changes
// are queryable without a process restart. Default: true.
options.EnableDurableHistoryView = true;
});
Tune the durable history retention on the sys-tenant-registry tree (the usage and
overage trees keep no per-key history):
using Orleans.Lattice;
using Orleans.Lattice.Tenancy;
siloBuilder.ConfigureLatticeTenancy(options =>
{
options.HistoryRetentionMode = HistoryRetentionMode.FullValue;
options.HistoryRetentionWindow = TimeSpan.FromDays(30);
});
Isolation model
- Structural tenant-segment prefix. A tenant owns a namespace of trees
addressed by an unqualified name; the tenancy layer injects a reserved tenant
segment so a tree id self-describes its owner. The composed id has the shape
t/{tenantId}/{name}, and the gate enforces ownership with a cheap ordinal prefix check - the same shape as the existing_lattice_/sys-reserved namespaces. The tenant prefix is a third reserved namespace with its own user-write guard on theILatticemutation surface: a user-origin write there may name at/id only when the id's structural owner is the caller's own active tenant - which is exactly what the facades compose. Reads are not guarded: a read naming another tenant'st/id reaches the access gate, which refuses the crossing unless a cross-tenant grant (or a platform-operator scope) authorizes it. A grant's scope is matched against that fullt/{owner}/...id, so a grant offered for an unqualified name such asordersis accepted but covers no tree. The app-tree prefixa/used by installable apps is deliberately not reserved or treated as qualified: an app treea/{app}/{tree}is an ordinary unqualified name, so it composes tot/{tenantId}/a/{app}/{tree}and each tenant gets its own copy of an installed app (with tenancy off it stays the barea/{app}/{tree}). An app that declares replication enrols each tenant's composed trees per install, through the runtime replication configuration (see Replication intent), and they are admitted by the tenant replication isolation gate like any other tenant tree. Compose and inspect tenant tree ids with the coreLatticeTenantTreeshelper:
using Orleans.Lattice;
TenantId acme = TenantId.Parse("acme");
// "t/acme/orders"
string treeId = LatticeTenantTrees.Compose(acme, "orders");
bool isScoped = LatticeTenantTrees.IsTenantScoped(treeId);
if (LatticeTenantTrees.TryGetTenant(treeId, out TenantId owner))
{
// owner == acme
}
- Derived trees are scoped through their name. A materialised view is not a
caller-supplied tree id but a tree the maintainer derives from the view's name,
so it is the name that is resolved to the active tenant: a view created as
ordersmaterialises ast/{tenant}/view-orders. Placing the tenant segment outermost is what makes ownership, enumeration filtering, and the tenant delete cascade apply to a view tree exactly as they do to any other tree, and it lets two tenants use the same unqualified view name over their own same-named sources while each reads back only its own. Because the maintainer, the view catalog, and the durable view registry are all keyed by the view name, scoping the name is also what makes the isolation survive a silo restart. Tag-index trees are not partitioned today and remain cluster-global. - The asserted tenant must reach the silo on every transport. Tenant scoping
is applied inside each API facade, so it only takes effect when the caller's
asserted tenant has been lifted onto the ambient context. In a co-hosted head
that happens in-process and flows to the grain on the Orleans request context.
On a split head - an API head in its own process reaching the silo over gRPC
- each binding must lift the
lattice-active-tenantheader itself. The control-plane bindings do it through the sharedLatticeActiveTenantAssertionhelper, and the data binding through its own replaceableILatticeDataApiActiveTenantBridgeseam. A binding that did not would not fault: its facade would resolve the reserved default tenant and serve the caller the shared cluster-global namespace, so the behaviour is covered by a contract guard rather than left to review.
- each binding must lift the
- A refused assertion is reported as a refusal, on every surface. The resolver
denies a caller by resolving the uninitialised
default(TenantId)"no tenant" value - anullTenantId.Value, deliberately distinct from the reservedTenantId.Default, whose value isdefault. Every surface that reads a resolved tenant honours that sentinel: the data plane refuses the operation, and the tenant self-awareness surface refuses too rather than reporting a live descriptor, so "which tenant am I acting as" can never answer with a tenant the caller was denied. The distinction matters when reading this page: the reserved default tenant is a real tenant that a caller legitimately resolves when it asserts nothing, whereas the sentinel means the assertion was rejected. - A denial is an authorization outcome, not a fault. A call refused by
fail-closed tenant resolution surfaces as
PermissionDeniedon every gRPC binding, carrying the reason (the apps bindings send a fixed message instead). It is deliberately notInternal: that is a retryable status, so a client would back off and retry a decision that can never change, and the refusal would be counted against the server-fault rate operators alert on. A call that resolves cleanly but breaches the tenant's quota is a different outcome again - capacity, not authorization - and surfaces asResourceExhaustedon the data and schema gRPC bindings. The tree-administration binding (where a create can breach the tree-count ceiling) maps it, as theInvalidOperationExceptionthatLatticeQuotaExceededExceptionderives from, toFailedPrecondition. An enumeration for a denied caller returns an empty page rather than an error, so listing never leaks the cluster-global catalog. - Identity-derived, enforced at the auth gate. The active tenant is carried in
the Orleans
RequestContextunder a single well-known key, stamped at the edge from the caller's asserted tenant and validated against the caller's membership at the auth seam. The fail-closed access gate behindILatticeAccessGateis made tenant-aware: a request is denied unless the subject's active tenant owns the target tree (prefix match), or an explicit cross-tenant grant or platform-operator scope authorizes it. - Membership, status and grant changes take effect at once. The gate answers
from a compiled snapshot of the tenant registry, and a registry write (adding or
removing a tenant admin, suspending or deleting a tenant, approving, rejecting, or
revoking a grant) only schedules a background rebuild of it. While that rebuild
is outstanding, or while rebuilds are failing, the gate confirms every request
that acts as an asserted active tenant against the registry itself: the active
tenant's record must exist, be
Active, and list the subject as an admin (the same rule the snapshot applies), and a cross-tenant crossing must also find an active grant in the owning tenant's record. Both checks must pass on their own. So a removed admin, or a subject acting as a just-suspended or deleted tenant, is refused on that tenant's own trees as soon as the write commits, a revoked grant stops admitting access, and an added admin or an approved grant admits the next request. An owned-tree request reads one record and a crossing reads its two records concurrently. A request that cannot be confirmed (for example because the registry read fails) is denied. Only requests in that window pay the registry read; the steady state stays an in-memory decision. - Every silo sees the change. The change feed that drives the rebuild fires only
on the silo whose grain committed the registry write, so each silo's snapshot is
also kept current by a cluster-wide tenant-policy epoch. Before a registry write
returns, the committing silo advances the epoch and pushes it to every silo, which
marks its snapshot out of date and rebuilds; the write completes once every silo
has acknowledged, or has had its lease lapse. Each silo holds its snapshot
authoritative only while it holds a live lease from the epoch (renewed every third
of
PolicySnapshotLeaseDuration) and has compiled the latest epoch it has seen. A silo that cannot know it is current - its lease has lapsed, it has been told of a change it has not compiled, or it could not publish a write of its own - confirms active-tenant requests and crossings against the registry (or denies) exactly as above, and the inbound replication isolation gate falls back to the registry for tenant existence and status in the same windows. The steady state pays only a few field reads and a timestamp read. A restarted epoch holds each write open for about 1.1 times the lease (one lease plus a tenth), or until every silo cluster membership does not report dead has leased from it, so no silo leased by its previous incarnation stays authoritative. One window is bounded rather than closed: a silo that crashes after committing a registry write but before publishing it leaves the other silos unaware of that write until cluster membership declares it dead, at which point every surviving silo rebuilds. Each silo also builds its snapshot at start-up, and on its first decision if that has not happened yet, so a new silo never reports a registered tenant as unregistered. The same epoch and lease keep each silo's residency and placement views of the registry current, with the same fail-closed fallbacks: see Every silo applies a residency change and Placement follows the registry on every silo. - Active-tenant assertion. A subject carries a set of tenant memberships, but
the active tenant is always a caller-supplied assertion, never inferred from
that set - there is no implicit "sole membership" default. Every branch that
consumes the assertion re-validates it against the caller's own membership, and
it is denied unless the named tenant is registered,
Active, and lists the caller as an admin subject. A request that asserts nothing resolves the reserveddefaulttenant, which is what keeps legacy adoption non-destructive; on a tenant-owned (t/...) tree that unasserted request is denied instead, because the uninitialised "no tenant" value can never be an active tenant. Assertingdefaultexplicitly is not the same as asserting nothing: it is validated like any other assertion, and because the reserved tenant is seeded with no admin subjects (and the control plane refuses to add any) that assertion fails validation. - Tenant id grammar. A
TenantIdmatches^[a-z0-9](https://github.com/NSTA1/Orleans.Lattice/tree/release/9.9/docs/lattice.tenancy/[a-z0-9-]{0,61}[a-z0-9])?$(lower-case alphanumeric and hyphen, 1-63 chars). This guarantees a tenant id can never contain the/segment separator and never begins with_, so it cannot collide with or spoof the_lattice_namespace or thet/{tenant}/{name}segment structure. The grammar alone does not exclude asys-prefix (those are all legal characters), so tenant creation additionally rejects an id beginning withsys-or_lattice_: a tenant id travels into tree ids, metric labels, and log lines beside real tree ids, and one shadowing a reserved namespace is an avoidable confusion trap. The check is applied at create only, so a tenant registered before the guard existed stays readable and deletable. The iddefaultis reserved for the legacy-adoption tenant: it can never be suspended, deleted, given quotas, have its admin-subject set changed, or be named on either side of a cross-tenant grant offer (each fails closed with aReservedTenantOperationException), while a resume of it is an allowed no-op. Tenant ids are immutable once created. - Tenant-scoped tree naming.
AddLatticeTenancyreplaces the core's no-opITenantContextResolverwith one that reads the caller's active tenant and re-validates it against that caller's own membership before it is allowed to scope a name. This is what makesservices.GetLatticeAsync("orders")addresst/acme/ordersfor a caller acting asacmeandt/globex/ordersfor one acting asglobex, rather than handing both the same physical tree. The tree-addressing API facades do the same: the data, state, tree-administration, schema, replication, and backup facades resolve the caller-supplied name through theResolveEffectiveTreeIdAsyncextensionLatticeTenantExtensionsadds overITenantContextResolver(the interface itself carries onlyResolveCurrentAsyncand its synchronousTryResolveCurrentfast path) at their entry point and use that one effective id for both the authorization check and the operation, so a verb can never authorize one tree and act on another. Without that, an unqualified name would stay a shared default-tenant tree. A caller that asserts no tenant resolves the reserveddefaulttenant and keeps its bare tree ids (non-destructive adoption); a caller asserting a tenant it may not act as resolves the uninitialised "no tenant" value, which fails closed with aLatticeTenantAccessDeniedExceptionrather than silently defaulting. Because the effective id is tenant-owned, usage metering and quota admission attribute the traffic to the acting tenant - attribution only lines up with the acting tenant once the name is actually scoped. - Enumeration pruning.
AddLatticeTenancylikewise replaces the core's no-opITenantEnumerationFilter, so a tree-id enumeration (the cluster-state tree catalog, the tag-index catalog, the view catalog, the in-cluster all-tree-ids read) is pruned to the trees the active tenant owns; platform-owned_lattice_andsys-ids are left in, for the catalog's system-tree switch and the per-entry authorization check to govern. Pruning is defence in depth rather than the boundary: a caller that asserts no tenant is not pruned, and is confined instead by the per-entry authorization check, which composes the same tenant enforcer the write path uses and denies a tenant-scoped tree outright when no active tenant is selected. That check is the durable guarantee - an existence probe can never out-reach the enforcement decision, so no broad grant orDefaultEffect = Allowposture can surface another tenant's tree names. - Enumeration is scoped at the source. Because a tenant's trees all begin
t/{tenant}/and the tree registry is itself an ordinally-sorted Lattice tree, a tenant's trees occupy one contiguous key range. Where it is provably equivalent to the unscoped read, an enumeration pushes that prefix down to the registry (LatticeTenantTrees.ComposePrefixsupplies it), so the scan is bounded to the tenant's own range and no other tenant's ids cross the grain boundary at all - rather than transferring the whole catalog for the caller to discard most of. The tenant delete cascade, which enumerates exactly one tenant's trees, is scoped this way, as is the tree catalog when a non-default tenant is active and the request excludes system trees (the ids a prefix scan skips are the ones that switch already drops). The prefix is a performance hint, never an authorization boundary: it can only ever return a subset of what the caller could already enumerate, and the pruning filter and per-entry authorization check still run unchanged. - Registry-store read isolation. The
sys-tenant-*registry, usage, and overage trees hold the cross-tenant registry itself - every tenant's admin subjects, quotas, region residency, and cross-tenant grants. They live in thesys-system-data namespace, so first-party access runs system-origin and short-circuits the gate; every external request is governed with control-plane read isolation, exactly like the reservedsys-auth-*policy store. A data-plane read or scan is denied independently ofDefaultEffect, and a cluster-wide all-trees (Tree:*) wildcard grant never reaches them, so no broad data-plane role can enumerate one tenant's metadata from another. Only a bootstrap administrator, a system-origin caller, or an explicit rule an operator deliberately scopes at a registry tree may read them.
Placement follows the registry on every silo
A tenant tree's WAL placement is resolved once, when the tree is first registered,
from an in-memory placement view of the registry, and the resulting pin is
immutable. The view is kept current by the same epoch and lease, so a placement
change made through one silo reaches every silo before the write returns. Placement
cannot be confirmed against the registry instead - it is resolved inside the tree
registry's own turn, which a registry read could re-enter - so while a silo's view
is not authoritative a tenant tree's registration waits up to a fifth of
PolicySnapshotLeaseDuration (2 seconds by default) for the view to catch up, and
is otherwise refused with a retryable TimeoutException rather than pinned to a
placement that may be stale. Creating a tenant and then its trees works without a
retry: the wait covers the rebuild the tenant write triggers. Non-tenant trees are
never affected.
Resource governance
Each tenant carries aggregate quotas across all of its trees, expressed by
TenantQuotas:
| Dimension | Property | Meaning |
|---|---|---|
| Durable bytes | MaxBytes |
Aggregate durable size across the tenant's trees (WAL, snapshot, and leaf-state bytes). |
| Live keys | MaxKeys |
Aggregate live-key count. |
| Resident memory | MaxMemoryBytes |
Aggregate resident memory, metered as the summed serialized leaf and shard-root grain-state bytes of the tenant's trees. |
| Tree count | MaxTreeCount |
Number of trees the tenant may own. |
| Request rate | MaxOpsPerSecond |
Cluster-wide ops/sec ceiling. |
| Burst | BurstPercent |
Percentage overage above the steady-state caps. |
A null cap on a dimension means unlimited on that dimension; a bounded cap, like
BurstPercent, must be non-negative, and a negative one is rejected with an
ArgumentException when the quotas are authored. The reserved
default tenant is permanently unbounded (it can never be given quotas), and every
newly created tenant starts with no caps until an operator sets them, so opt-in
never suddenly throttles an existing workload.
- Compiled quota policy. Steady-state enforcement uses a compiled policy
snapshot with a monotonic epoch, refreshed off the
sys-tenant-*change feed and evaluated synchronously in-memory with no I/O once warm - mirroring how the auth package answersILatticeDecisionEnginefrom a compiled snapshot behind its own fail-closed access gate. - Burst and metering. Usage at or below the steady-state cap is ordinary; usage
above the cap and at or below
cap x (1 + burst%)is admitted and metered as overage - a first-class, billing-ready signal distinct from ordinary usage; usage abovecap x (1 + burst%)is refused withLatticeQuotaExceededExceptioncarrying the tenant id and dimension (Dimensionisbytes,keys,memory, ortrees, andops-per-secondfor the request-rate budget). A tenant with burst0refuses as soon as usage exceeds the cap. The refusal gates new writes only: usage that already sits above the cap - for example after a cap is lowered - keeps accruing overage on every metering tick, in whichever band it sits. Over the data gRPC binding the refusal reaches a remote caller asResourceExhaustedcarrying the breached dimension as a trailer; the tenant id is not added as a trailer (only the status message names it), because the caller asserted its own active tenant on the request. - Metering drives enforcement, on a cadence. A footprint quota (bytes, keys,
memory, and the tree count an ordinary write is checked against) is admitted
against the tenant's metered usage, so it binds only once a usage sample lands;
the request rate and the tree-count check at creation do not wait for one (both
are covered below). Each silo
runs a background metering cycle every
TenantUsageAccountingOptions.MeterInterval(default 30 seconds) that walks each tenant's own trees - a bounded range scan over the tenant'st/{tenant}/key range, not a read of the whole catalog - samples their footprint, and rolls the result up into that tenant's per-cluster usage slot. Every id registered in that range is sampled, including the physical copy that a resize, a shadow-cutover restore or a schema remediation registers beside the tree it aliases; the logical id's report resolves the alias to that same copy, so an aliased tree's footprint, and its tree count, are counted twice. The reserveddefaulttenant is skipped: it can carry no quotas, so it is never metered. Admission deliberately fails open for a tenant with no landed sample yet, so a cold silo never spuriously refuses; that means enforcement arms one cycle after a tenant first has usage. SettingMeterIntervalto zero disables metering entirely and leaves footprint admission permanently open, which is only appropriate for a deployment running tenancy without resource governance. - A tenant's first non-empty sample always publishes. Republishing a usage slot is gated
by a hysteresis band (
PublishMinAbsoluteDelta/PublishMinRelativeDelta) so a stream of negligible movements does not churn the registry. That band damps churn between successive samples, so it is deliberately not applied to a tenant's first publish: until the slot exists admission is fail-open and no quota binds at all, so a tenant whose whole footprint sits below the absolute floor (default 65,536) would otherwise never be governed. Establishing the slot costs one write per tenant per publisher lifetime; every movement after it is damped as normal. - A stale footprint is re-anchored, not trusted. The key and memory figures a
tree reports are activation-scoped: a shard root rebuilds them as its leaves
republish on commit boundaries, so they read zero after a reactivation until
writes resume - and Orleans collects idle grains, so that needs no restart or
fault. The byte figure is unaffected because it adds durable WAL retention. A
tree that reports no keys and no leaf bytes yet a non-zero total is therefore
showing a cold cache rather than an empty tree, and metering re-anchors it with
a deep walk instead of publishing the zero. Without that,
MaxKeysandMaxMemoryByteswould fail open - admitting a tenant well over quota - whileMaxBytesandMaxTreeCountkept binding. The cost is self-limiting: a large tree re-anchors once and then reports non-zero, so only a genuinely empty tree that still retains WAL is re-walked, and walking an empty tree is cheap. - Request rate is enforced with the footprint dimensions.
MaxOpsPerSecondis applied by the same admission seam, ahead of the footprint checks, from the tenant's silo-local token budget. A breach surfaces asLatticeQuotaExceededExceptionon theops-per-seconddimension and is explicitly transient: the budget refills continuously, so an immediate retry after a short backoff succeeds, unlike a footprint breach which persists until the tenant's usage drops. - Reads are rate-admitted too, but never footprint-admitted. A read is charged
against
MaxOpsPerSecondat the read plane, so a tenant cannot saturate a shared silo with reads that cost it nothing - the rate ceiling is what bounds a noisy neighbour, and leaving the entire read plane outside it left that ceiling bounding only half the traffic. The footprint dimensions are deliberately not applied to reads: refusing reads because a tenant is over its storage quota would trap it, unable to read the data it must delete to get back under. The charge is taken strictly after the access gate authorizes the read, never before, because the tenant is a caller assertion that only the gate validates - charging first would let an unauthorized caller drain a named victim's budget and read the victim's usage and ceiling back out of the refusal. The charge covers every read shape a tenant can drive, including the two that do not cross the data-plane seam: a backup capture, which is the largest tenant-triggerable read the platform offers, is charged once per capture at its own seam after the backup authorizer allows it; and a snapshot cursor page, which reads snapshot leaf grains directly, is charged per page. A turn the access gate never adjudicated - a system-origin turn, or authorised view-maintenance traffic - is never charged, because there is no validated tenant to charge it to. - Tree creation is admitted.
MaxTreeCountis charged where a tree is explicitly created, so the one dimension whose whole purpose is to bound tree creation binds at the point of creation. It is enforced once, at the tree-administration facade every explicit create funnels through, rather than additionally at the tenant-scoped facade above it: admission consumes a rate token, so evaluating it at both layers would bill a single create twice. The ceiling is checked against an authoritative count of the tenant's registered trees read at the moment of the create (every id registered under itst/{tenant}/prefix, so the physical copy beside an aliased tree counts as a tree of its own), not against the metered sample, so it binds even for a tenant that has never been metered; creates that read the count concurrently can each be admitted, so the cap can be overshot by at most the number of creates in flight. A tree the data plane registers implicitly on first use - for example the first write to a new name - never passes that check: it is bounded only by the metered tree count an ordinary write is admitted against, so implicit creation can overshoot the cap until the next metering sample lands. - The reserved
sys-namespace is closed to tenants. Tenant scoping composes the active tenant into a tree name, and deliberately passes an already-qualified name through uncomposed so it is never double-composed. The reservedsys-system-data namespace counts as already-qualified, which is right for the first-party add-ons that own those trees but meant a tenant naming one had the id returned uncomposed - and therefore global. Such a tree sits outside thet/{tenant}/prefix that per-tenant tree-count and footprint accounting enumerates, so it is invisible to the quotas meant to bound it; it is shared with every other tenant that picks the same name; and it can collide with an add-on's own store. A non-default tenant addressing that namespace outside a system-origin scope is now refused withLatticeTenantAccessDeniedException. A malformedt/-prefixed id carrying no tenant segment is refused on the same seam, for the same reason: it resolves to platform ownership, which the tenancy gate allows unconditionally. A well-formed foreign id such ast/other/ordersis deliberately not refused here, because cross-tenant grants are real and only the gate can adjudicate them - the resolution layer cannot see grants, so it must not decide crossings. - Apply-path admission bypass, never isolation bypass. As in core, the replication-apply and saga-apply paths bypass quota admission (they re-enter under a foreign/prepared scope) but never bypass the tenant isolation boundary.
Enforcement scope (multi-cluster)
Quota admission runs under a TenantEnforcementScope. Today the scope is
cluster-wide rather than per tenant: every tenant is admitted under
TenantUsageAccountingOptions.DefaultEnforcementScope, read live, and the tenant
record carries no scope of its own (the resolver is a seam for a future per-tenant
override):
GlobalConverged(default). For the slow-moving storage gauges (bytes, keys, memory, tree count) each cluster contributes its current local usage for the tenant to a per-cluster-slot state CRDT (a map fromClusterIdto that cluster's latest sample). A cluster writes only its own slot and reads the whole map, so global usage is the sum-fold over every slot published so far - the fold is not filtered by the tenant's region statuses. The map holds another cluster's slot only once thesys-tenant-usagetree replicates between them; until then the global fold equals this cluster's own sample. Enforcement admits against the global fold, giving a single global budget rather thanlimit x clusters, with bounded transient overshoot. The monotonic overage tallies use grow-onlyGCounters (one per bytes, keys, memory, and tree-count dimension), but they are not metered from the global fold: whatever the scope, each cluster accrues the overage of its own local usage above the tenant's whole steady-state cap into its own component, and the converged tally sums the components. Slots are republished on a cadence with hysteresis so continuous usage does not flood the replication path.PerCluster(fallback). Each cluster admits against only its own local usage slot, so effective global capacity islimit x clusters. Selectable, cluster-wide, by operators who prefer hard-partitioned capacity. The scope changes only which figure admission reads: every cluster still meters and publishes its slot on the same cadence, so it does not remove the usage-publishing traffic.
No enforcement scope introduces cross-cluster coordination or consensus -
GlobalConverged reads a convergent CRDT sum, it never locks or votes.
Rate limiting
The ops/sec limit is always enforced per-cluster, whichever enforcement scope is
configured (a rate window is too short relative to replication lag for a
converged global count to be meaningful). It is enforced by silo-local, in-process token buckets - a per-silo
singleton limiter (not a grain) the data-plane entry path consults with a lock-free
token decrement - so the per-op hot path takes zero grain hops. A low-frequency
per-(tenant, cluster) budget coordinator divides the cluster rate across the live
silos at lease cadence (O(silos), never O(ops)). LatticeTenantRateLimiterOptions
tunes that coordinator; none of its knobs touch the per-op hot path, so a
misconfiguration changes only how the cluster rate is split, never whether
enforcement stays lock-free. Only an active tenant with a positive
MaxOpsPerSecond gets a bucket: a MaxOpsPerSecond of 0 leaves the tenant as
unthrottled as null does. Each silo's share is floored at one operation per
second, and BurstPercent applies here too: a bucket may run about BurstPercent
percent of the silo's share (at least one operation when the percent is positive)
ahead of the steady rate, while a burst of 0 admits operations no closer together
than the share's steady spacing.
| Option | Type | Default | Meaning |
|---|---|---|---|
LeaseInterval |
TimeSpan |
30s |
How often the coordinator re-apportions each tenant's cluster rate across the live silos. A longer interval lowers coordination cost but widens the transient overshoot bound (lease interval times cluster rate); the default is sized for work backed by a whole-tree registry scan. A non-positive value falls back to the default, and the tick period is held to about 49.7 days (the longest period a timer accepts). |
LeaseCycleTimeout |
TimeSpan |
20s |
The bound on a single lease cycle. A cycle that exceeds it is cancelled and retried on a later tick, so a stalled tenant-registry read can never occupy the loop for longer than one interval. Clamped down to LeaseInterval if set at or above it, so the duty cycle stays bounded, and held to about 49.7 days (the longest delay a timer accepts). A non-positive value falls back to the default. |
MaxLeaseBackoff |
TimeSpan |
5m |
The ceiling the lease interval backs off to after consecutive cycle failures. The effective interval doubles per consecutive failure and resets to LeaseInterval on the first success, so a persistently unhealthy registry is probed at a decaying rate rather than hammered every tick. A value below LeaseInterval disables backoff; a non-positive value falls back to the default. |
RateSnapshotTtl |
TimeSpan |
2m |
How long a read of the registry's configured rates stays usable before the next cycle re-reads it. Configured rates change at administrative cadence, so caching them decouples the frequent re-apportionment of token buckets from the expensive whole-tree registry scan. The snapshot is stale-if-error, so a failed refresh apportions from the previous snapshot rather than pruning every tenant's bucket. A non-positive value falls back to the default. |
Apportionment |
TenantRateApportionmentStrategy |
Demand |
Demand leases demand-proportionally and degrades to static-even when no cluster-wide demand aggregate is available; StaticEven is the zero-coordination fallback that splits the rate evenly. The package ships no cluster-wide demand aggregator - its in-process demand exchange always reports none - so Demand apportions exactly as StaticEven does. |
DemandReserveFraction |
double |
0.2 |
The fraction of the cluster rate that demand-proportional leasing reserves and splits evenly, guaranteeing an idle silo a non-zero floor so it can never be starved out of building demand. In [0, 1] (a value outside is clamped to that range); ignored under StaticEven, and whenever no cluster-wide demand aggregate is available (see Apportionment). |
A breach surfaces as a LatticeQuotaExceededException on the ops-per-second
dimension. Unlike the footprint dimensions it is transient: the same call
generally succeeds once the bucket refills, so a client should treat it as a
back-pressure signal to retry rather than as a durable capacity failure. Over the
data gRPC binding it reaches
a remote caller as a ResourceExhausted RpcException carrying the breached
dimension as a trailer, so a client can tell a retryable rate breach from a
footprint breach that will not clear on its own.
Store write contention
Each of the three sys-tenant-* stores' write paths - ITenantRegistry.PutAsync, the
usage-slot publish, and the overage accrual - is an optimistic read-merge-write:
the store reads the tenant's record with its version, folds the change in with the
record's CRDT join, and writes back only if the version has not moved. A write that
loses that race re-reads (now seeing the competing write) and merges again, at once
and with no backoff, so a concurrent change is never dropped. After a small, fixed
number of lost races on the same tenant's record the store gives up and throws one
of three public exceptions, each carrying the Tenant and the number of Attempts
it made. The retries absorb ordinary contention; each exception signals sustained
write contention on one tenant.
| Exception | Raised by | What happens |
|---|---|---|
TenantRegistryConcurrencyException |
ITenantRegistry.PutAsync, which every record change the tenant-administration facades make is written through (DeleteAsync removes a record outright and never raises it) |
The change is not applied and the exception reaches the caller, which may retry. The tenant-administration gRPC binding has no arm for it, so a remote caller sees Internal. |
TenantUsageConcurrencyException |
The metering cycle's usage-slot publish | Caught and logged for that tenant: its overage accrual is skipped for the tick too, the rest of the pass continues, and the next tick retries. |
TenantOverageConcurrencyException |
The metering cycle's overage accrual | Caught and logged for that tenant: that tick's overage is not recorded - the tally is a per-tick sum, so it is not recovered later - and the next tick accrues as normal. |
Region residency
Which regions a tenant lives in is a per-tenant, runtime-mutable choice layered on top of the replication topology.
The region sets
Most confusion about region residency comes from collapsing distinct sets into a single notion of "where a tenant is". Each has a different owner and a different surface:
| Set | Who controls it | Surface | What it means |
|---|---|---|---|
| Physical / routable | Operator (deployment topology) | lattice_list_regions in the MCP binding |
Every region the deployment actually has a route to. |
| Allowed | Operator only - a tenant admin cannot change it | ILatticeTenantRegionAdmin.AuthorizeAllowedRegionsAsync |
The regions a tenant is permitted to place residency in. |
| Resident | Tenant admin, but only within the allowed set | ILatticeTenantRegionAdmin.SetResidencyAsync |
The regions the tenant actually replicates to and is served from. |
The union of allowed and resident is the tenant's actionable set: the regions
it is in, plus the regions it may move into. A tenant admin never needs the physical
list - SetResidencyAsync refuses anything outside the allowed set - so tenant-facing
discovery never shows a tenant caller more than its actionable set plus the region
serving the call. See
MCP security for
how that scoping applies to region discovery, and
Tenant-aware surfaces for what a tenant-asserting caller is
shown today.
- Allowed vs resident. A platform operator authorizes, per tenant, the allowed region set; the tenant's delegated admin selects its residency set (the subset it actually replicates to and is served from) within that allowed set. A tenant that has never configured residency - every newly created tenant - is treated as online in every region, the pre-residency admit-all behaviour, until it does.
- Metadata everywhere, data to the residency set. Tenant definitions
converge to every region the registry tree replicates to, so any such region can
fail-closed answer "is this tenant resident here?". A tenant's data is shipped to peers like any other replicated
tree; the receiving region refuses (and dead-letters) a replicated write for a
tenant that is not
Onlinethere, so the data lands only where the tenant is online. The same gate refuses and dead-letters a replicated write for a tenant the receiving region does not know or holds suspended. - Symmetric multi-master. An
Onlineregion is a full read-write replica; there is no primary or leader. Enforcement ties in at the gate (a tenant notOnlinein the serving region is refused) and the replication apply path (a tenant's replicated writes land only in a region where it isOnline).
Every silo applies a residency change
Each silo answers "is this tenant online here?" from an in-memory residency view of the registry, and the change feed that refreshes it fires only on the silo that committed the write. So the view is kept current by the same cluster-wide epoch and lease as the tenant-policy snapshot (see Isolation model, "Every silo sees the change"): a residency change made through one silo reaches every silo before the write returns. While a silo's view is not authoritative - it has been told of a change it has not yet compiled, or its lease has lapsed - the tenant gate and the replication isolation gate confirm the tenant's residency against its registry record instead, so a tenant drained or taken offline through any silo is refused on every silo at once, and one brought online is admitted at once. A residency check that cannot be confirmed is refused. The steady state stays an in-memory lookup.
Lifecycle states
Each region carries one TenantRegionStatus per tenant, readable through
GetTenantRegionStatusAsync:
| Status | Resident? | Meaning |
|---|---|---|
None |
No | No relationship. An allowed but not yet entered region reports None. |
Provisioning |
Yes | The region has been added to the residency set and is not yet serving. |
Backfilling |
Yes | The step between Provisioning and Online, reserved for copying existing data into the region; not yet serving. |
Online |
Yes | A full read-write replica, and the only status in which this region serves the tenant. |
Draining |
No | The region has been dropped from residency and no longer serves. |
Offline |
No | The step after Draining: drained, and no longer serving. |
Removed |
No | Terminal: the removal is complete. |
The resident set is exactly the rows whose status is Provisioning,
Backfilling, or Online. A region does not serve the tenant until it reaches
Online, and once any region status is set the tenant is served only in a region
where its status is exactly Online.
SetResidencyAsync applies only the first step of each path: Provisioning for an
added region and Draining for a dropped one. The later steps (Provisioning ->
Backfilling -> Online, and Draining -> Offline -> Removed) are single-step
promotions, and the two paths are completed differently.
The remove path completes on its own. With the
tenant-admin control API
registered, each silo watches its own serving region (its cluster id) and, when a
tenant's status there becomes Draining, advances it to Offline and then to
Removed without any caller. Nothing needs to be waited for first: a region stops
serving a tenant and stops admitting its replicated writes the moment its status
leaves Online, and outbound shipping of the writes it accepted while online does not
depend on the status. A dropped region whose silos never run again (a decommissioned
region) keeps the Draining it was given, which is harmless - it is already neither
resident nor serving.
The add path needs an operator step. No shipped component backfills a region a
tenant is added to. While the tenant is not Online in a region, that region refuses
and dead-letters every replicated write for it - including anything a backfill would
apply, because dead-letter replay and snapshot re-seed go through the same gate - so
an added region lacks whatever was written while it was not admitting the tenant, and
cannot be filled in until it is Online. Promoting it automatically would declare an
incomplete replica online without anyone deciding to, so nothing does: an added region
stays at Provisioning, and a tenant whose residency has been set is served in no
region until an operator advances one. Where the region was admitting the tenant's
writes up to the add - residency being configured for the first time, so every region
was admit-all until then - it misses at most the writes shipped since, and advancing
it is enough. Otherwise advance it and then recover the gap: replay the region's
dead-lettered writes for the tenant's trees (see the
dead-letter queue), or let the
anti-entropy digest probe
repair it where that is enabled. Advance the region one step at a time,
Provisioning -> Backfilling -> Online, from host code on a silo:
using Orleans.Lattice.Tenancy;
// Advances a region one legal lifecycle step and returns its committed status.
// Call it twice to take an added region from Provisioning to Online.
static async Task<TenantRegionStatus> PromoteRegionAsync(
ITenantRegistry registry, TenantId tenant, string regionId, string clusterId, CancellationToken cancellationToken)
{
var record = await registry.GetAsync(tenant, cancellationToken)
?? throw new InvalidOperationException($"Tenant '{tenant}' is not registered.");
if (!record.TryPromoteRegionStatus(regionId, clusterId, out _))
{
return record.GetRegionStatus(regionId);
}
var committed = await registry.PutAsync(record, cancellationToken);
return committed.GetRegionStatus(regionId);
}
TenantRecord.TryPromoteRegionStatus applies only the next legal step
(TenantRegionLifecycle.TryNextPromotion) and is a no-op at Online, Removed, and
None. It stamps the promotion as the immediate successor of the status it read, not
at wall-clock now, so it never overwrites a residency change a tenant admin commits
after that read. Prefer it to TenantRecord.SetRegionStatus, which applies any status
whose stamp supersedes the current one and checks neither the step nor that ordering.
Quota accounting does not follow these statuses: the GlobalConverged
fold sums every cluster slot the tenant has published, whatever the status of that
slot's region.
Invariants
Enforced by ILatticeTenantRegionAdmin and never bypassed by a transport binding:
- Residency is always a subset of allowed. Setting residency to a region outside
the allowed set is refused with
TenantRegionNotAllowedException. - The last resident region can never be removed. Narrowing residency to the empty
set is refused with
TenantLastRegionException. - A region a tenant is still resident in cannot be revoked. Drain it with
SetResidencyAsyncfirst, then revoke it withAuthorizeAllowedRegionsAsync. - An unknown tenant fails closed on every operation: with
TenantNotFoundExceptionfor a platform operator, and with the sameLatticeAuthorizationDeniedExceptionas any other refusal for a non-operator caller, so a tenant admin cannot probe for another tenant's existence.
Region residency is administered through the
ILatticeTenantRegionAdmin
control-plane facade, which is reachable
over gRPC and
through MCP tools.
Observability
TenantObservabilityOptions (default PublishGauges = true, PublishInterval 30
seconds) publishes per-tenant gauges - current usage against each quota dimension and
the metered overage tallies - on a fixed cadence, so an operator can see per-tenant
consumption and headroom. Separately, every registered
ITenantRegionStatusChangeListener is notified of each tenant's local-region status
transition - any change of TenantRegionStatus, such as a region entering
Provisioning or Draining - once the residency snapshot has observed it.
Every instrument is an observable gauge on the orleans.lattice.tenancy meter
(LatticeTenantMetrics.MeterName). Each series carries a single tenant tag
(LatticeTenantMetrics.TagTenant): the owning tenant's id on every per-tenant
series (default for the reserved legacy-adoption tenant), and the reserved
_platform_ sentinel on the one cluster-aggregate series,
orleans.lattice.tenancy.tenants. The per-tenant series cover every tenant in the
registry. Set PublishGauges = false to publish none of them.
| Instrument | Unit | Meaning |
|---|---|---|
orleans.lattice.tenancy.tenants |
{tenant} |
Cluster-aggregate count of tenants in the warm usage index. It belongs to the platform rather than to any tenant, so it carries the reserved _platform_ sentinel as its tenant value rather than being left untagged - a tenant-scoped matcher then excludes it by stating so, not by accident of absence. |
orleans.lattice.tenancy.usage.bytes |
By |
The tenant's current aggregate durable bytes. |
orleans.lattice.tenancy.usage.keys |
{key} |
The tenant's current aggregate live-key count. |
orleans.lattice.tenancy.usage.memory_bytes |
By |
The tenant's current aggregate resident memory. |
orleans.lattice.tenancy.usage.trees |
{tree} |
The number of trees the tenant currently owns. |
orleans.lattice.tenancy.quota.bytes |
By |
The tenant's steady-state MaxBytes ceiling. |
orleans.lattice.tenancy.quota.keys |
{key} |
The tenant's steady-state MaxKeys ceiling. |
orleans.lattice.tenancy.quota.memory_bytes |
By |
The tenant's steady-state MaxMemoryBytes ceiling. |
orleans.lattice.tenancy.quota.trees |
{tree} |
The tenant's steady-state MaxTreeCount ceiling. |
orleans.lattice.tenancy.quota.burst_percent |
% |
The tenant's BurstPercent headroom above its bounded ceilings. |
orleans.lattice.tenancy.overage.bytes |
By |
Converged, durable metered byte overage accrued above the byte ceiling. |
orleans.lattice.tenancy.overage.keys |
{key} |
Converged, durable metered key overage. |
orleans.lattice.tenancy.overage.memory_bytes |
By |
Converged, durable metered resident-memory overage. |
orleans.lattice.tenancy.overage.trees |
{tree} |
Converged, durable metered owned-tree overage. |
Each instrument name is also a public constant on LatticeTenantMetrics
(TenantsName, UsageBytesName, UsageKeysName, UsageMemoryBytesName,
UsageTreesName, QuotaBytesName, QuotaKeysName, QuotaMemoryBytesName,
QuotaTreesName, QuotaBurstPercentName, OverageBytesName, OverageKeysName,
OverageMemoryBytesName, and OverageTreesName), and the LatticeTenantMetrics.Meter
instance is public, so a listener can subscribe by reference rather than by name.
The four ceiling gauges (quota.bytes, quota.keys, quota.memory_bytes, and
quota.trees) emit a measurement only for a tenant whose corresponding dimension
is bounded - an unbounded (null) ceiling contributes no series at all, so "no
series" reads as "unlimited on that dimension" rather than "zero".
quota.burst_percent is emitted for every tenant, 0 when it has no burst
allowance. Usage gauges reflect the last landed metering sample (see
MeterInterval above), folded across every published cluster slot whatever the
enforcement scope; a registered tenant with no sample yet reports zero usage
rather than no series, so a zero reading can also mean "not yet metered" - and the
reserved default tenant, which is never metered, always reads zero usage. The
overage.* gauges are the billing-ready tallies: they are grow-only converged sums,
not instantaneous readings.
The same figures are readable in-process through public seams the package
registers: ITenantOverageBilling returns the converged metered overage for one
tenant or every tenant, for a billing consumer to poll; ITenantObservabilityView
returns the caller's own validated active tenant's usage, quota, burst, and overage
snapshot, and every tenant's only under an explicit
TenantObservabilityScope.ClusterWide(subject) scope whose subject validates as a
platform operator (anything else falls back to the active tenant); and ITenantUsageReader reads one tenant's usage by id
with no visibility check of its own, so its consumer must authorize the caller
against that tenant first.
MaxOpsPerSecond has no gauge: the rate budget is enforced from silo-local token
buckets rather than from a published aggregate, so a breach is observed through the
ops-per-second LatticeQuotaExceededException rather than a series.
These instruments are charted by the bundled Per-Tenant Observability Grafana
dashboard (LatticeDashboardKind.Tenancy), which offers a templated tenant
variable so a panel can be scoped to a single tenant or to every tenant. See
docs/lattice.dashboards/metrics-to-panel-map.md for the instrument-to-panel
mapping. To consume them directly instead, subscribe to the
orleans.lattice.tenancy meter from your OpenTelemetry exporter.
The derived tenant dimension
Beyond this meter, every instrument that Orleans.Lattice and its add-on packages
publish carries the same derived tenant tag (LatticeTenantLabel.TagTenant). It
is emitted on tenancy-on and tenancy-off clusters alike, so a dashboard query is
byte-identical in both deployment modes, and it is derived from the tree id rather
than read from the caller:
| Tree id | tenant value |
|---|---|
A well-formed t/{tenantId}/{name} id |
The owning tenant's id. |
| A bare, unsegmented id (every id on a tenancy-off cluster) | default (LatticeTenantLabel.DefaultTenant), the legacy-adoption tenant. |
A _lattice_ or sys- id, or a malformed t/ id |
_platform_ (LatticeTenantLabel.PlatformTenant). |
An instrument with no tree dimension carries _platform_ too, unless its
measurement is attributable to a tenant some other way - the per-tenant gauges on
this meter carry the tenant id directly. default is a real, queryable tenant;
_platform_ names platform state that no tenant may see. The sentinel is a tag
value rather than an absent label, and it opens with an underscore - which the
tenant-id grammar forbids - so it can never collide with a real tenant. A
tenant-scoped matcher such as tenant="acme" therefore excludes platform series
by construction, and a tree-id regex is not a substitute: tree!~"^t/.*" would
also match the _lattice_ and sys- platform trees.
Security
- Fail-closed everywhere. Every enforcement seam denies on an unmatched request.
Tenant data isolation and tenant lifecycle administration are each independent of
the data-plane
DefaultEffect, so an unmatched request always resolves to deny even underDefaultEffect = Allow. - Registry confidentiality. The
sys-tenant-*registry, usage, and overage trees are control-plane read-isolated: a data-plane read or scan is denied independently ofDefaultEffect, and no cluster-wide all-trees (Tree:*) grant can reach them, so the cross-tenant registry can never be enumerated through a broad data-plane read role. See "Registry-store read isolation" above. - A tenant admin cannot self-grant cross-tenant read. The registry escape hatch -
"an explicit rule an operator deliberately scopes at a registry tree" - is an
operator act by construction, and here operator has a narrow, specific
meaning: a bootstrap administrator (the break-glass root of trust configured on
the silo) or a subject explicitly promoted to access-administrator through the
access-administration delegation - by a bootstrap administrator, or by an existing
delegate, since a delegate may delegate further. It does not
mean "any authenticated caller," "any caller with a broad data-plane grant," or "a
tenant admin." The reason a tenant admin cannot perform this act is structural, not a
matter of degree: authoring any authorization rule is a write to the reserved
sys-auth-*policy store, which requires whole-treeAdminon that store - a control-plane capability held only by the operators just defined. A tenant admin's authority is only its membership of a tenant's admin-subject set, which the tenant-tier facades check on that tenant's own record: it confers nothing on the policy store and nothing over a tenant whose set does not name it, so a tenant admin can neither reach the policy store to author a registry-read rule (for itself or anyone else) nor act for another tenant. Direct writes tosys-tenant-*are likewise refused off the system-origin path by the reserved-prefix write guard. Consequently, granting cross-tenant registry visibility always requires a deliberate operator decision to author (or delegate the authority to author) that rule; a caller acting purely as a tenant admin has no path to it. - Two-tier governance. A platform-operator capability (cluster-wide) performs tenant lifecycle, quota and burst changes, and allowed-region authorization, and may act on any tenant's behalf. A delegated per-tenant admin capability (scoped to one tenant) manages that tenant's trees, admin subjects, schema, region residency, and its own side of a cross-tenant grant (offering its data, or approving, rejecting, or revoking a grant it is party to) strictly within the granted quota and allowed-region set - it can neither raise its own caps, widen its allowed regions, nor reach another tenant.
- Enable-gated. No tenant can be created unless the feature is enabled, and every
tenant-administration control tool is contributed only when the host opts control
in (
EnableTenantAdminControlTools).
Tenant-aware surfaces
Tenancy also reaches the operator- and agent-facing surfaces. Each is a module the host registers (below); once registered, it keys its behaviour off whether tenancy is actually present rather than off a separate opt-in flag, so a deployment without tenancy keeps a byte-for-byte-unchanged UI and tool surface.
- Explorer. The Explorer's web head (
AddLatticeExplorerWeb) always registers its tenant view (AddExplorerTenantView()). With tenancy on, the active tenant becomes the root node of every tenant-scoped address (/t/{tenant}/...), typingt/in the address line re-roots the current address at another tenant the caller may reach, and the Tenancy area serves the operator's tenant directory and each tenant's own pages. An address for another tenant is a switch request, authorized fail-closed through the operator gate; a refusal redirects to the active tenant with a warning. A deployment without tenancy shows no tenant root anywhere. See Tenant scope. - MCP. When tenancy is wired and the self-awareness module is registered
(
AddTenantSelfAwarenessTools(), which the split-head remote registration calls itself whenever a tenant-admin endpoint is configured; the module then self-gates on the tenant self-service facade and needs no flag), the MCP server contributes a read-only tenant self-awareness tool group -lattice_tenant_current(the tenant the caller is operating as),lattice_tenant_list(the tenants the caller may access), andlattice_tenant_get(one accessible tenant's lifecycle and per-region residency). Each is scoped fail-closed to the caller's subject: an anonymous caller lists nothing, and an inaccessible tenant is indistinguishable from an absent one. The tenant-admin control tools - the lifecycle and quota toolslattice_tenant_create,lattice_tenant_suspend,lattice_tenant_resume,lattice_tenant_delete, andlattice_tenant_set_quotas, plus the region-residency toolslattice_tenant_authorize_regions,lattice_tenant_set_residency, andlattice_tenant_region_status- remain separately gated behindEnableTenantAdminControlTools. Every tool in that group is annotated destructive and non-read-only exceptlattice_tenant_region_status, which is a read. Region discovery is tenant-scoped too, and fails closed:lattice_list_regionshonours a non-default tenant assertion only after re-validating it against the caller's own membership, and then advertises that tenant's actionable set plus the region serving the call, annotated with its standing; an assertion that does not validate is shown the serving region alone, with no tenant annotation. In the shipped registrations the discovery tool cannot validate an assertion - a co-hosted head carries no caller credential into it, so the caller resolves as anonymous, and a split head has no resolver able to validate one - so a tenant-asserting caller is currently shown only the serving region. SeeOrleans.Lattice.Api.Mcp.
Configuration reference
Only LatticeTenancyOptions is bound by the registration delegate:
AddLatticeTenancy(Action<LatticeTenancyOptions>?) and
ConfigureLatticeTenancy(Action<LatticeTenancyOptions>) each accept that type and no
other. Every other options type below is a plain registered option, so configure it on
the service collection directly - for example
services.Configure<TenantUsageAccountingOptions>(o => o.MeterInterval = TimeSpan.FromSeconds(10)).
LatticeTenancyOptions
| Property | Type | Default | Meaning |
|---|---|---|---|
HistoryRetentionMode |
HistoryRetentionMode |
MetadataOnly |
Retention mode for the durable per-key history captured on the sys-tenant-registry tree; the usage and overage trees keep none. History is never disabled by default. |
HistoryRetentionWindow |
TimeSpan? |
null |
Age after which a registry history revision row expires; null means no age bound. Must be strictly positive when supplied. |
EnableDurableHistoryView |
bool |
true |
Whether to create the durable history materialised view (sys-tenant-registry-history) over the sys-tenant-registry tree. |
SeedDefaultTenant |
bool |
true |
Whether to seed the reserved default tenant (unbounded quota) at startup when absent. The seed is create-if-absent, so it never clobbers an operator's later edits. |
PolicySnapshotLeaseDuration |
TimeSpan |
10s |
How long a silo may treat its compiled tenant-policy, residency and placement snapshots as authoritative without renewing its lease from the cluster-wide tenant-policy epoch; renewed every third of this. While the lease is lapsed, every request that acts as an asserted active tenant (on its own trees or across a cross-tenant grant), residency checks and inbound-replication tenant checks are confirmed against the registry or denied, and a tenant tree's registration waits up to a fifth of this for the placement view to become authoritative before it is refused. It is also the most a registry write can be held open (about 1.1 times this) when a silo cannot be reached, or just after the epoch restarts, so keep it well below the Orleans response timeout. Must be strictly positive and at most 0xFFFFFFFE milliseconds (about 49.7 days). |
TenantUsageAccountingOptions
Governs usage metering and the quota-enforcement scope every tenant is admitted under.
| Property | Type | Default | Meaning |
|---|---|---|---|
DefaultEnforcementScope |
TenantEnforcementScope |
GlobalConverged |
The enforcement scope every tenant's quota admission runs under, read live; there is no per-tenant override yet. |
PublishMinAbsoluteDelta |
long |
65536 (64 * 1024) |
Absolute movement, in the sampled unit, below which a usage republish is damped. A tenant's first non-empty publish is never damped. |
PublishMinRelativeDelta |
double |
0.05 |
Relative movement, as a fraction of the last published value, below which a usage republish is damped. Per dimension the effective threshold is the larger of PublishMinAbsoluteDelta and this fraction of that dimension's last published value, and the slot republishes when any one dimension moves by at least its threshold. A negative value of either knob is treated as zero. |
MeterInterval |
TimeSpan |
30s |
The per-silo metering cycle that samples each tenant's footprint and rolls it into that tenant's per-cluster usage slot. Zero or a negative value disables metering entirely, which pins footprint admission in its documented fail-open branch so an authored footprint quota never binds (the request rate and the tree-count check at creation still apply). Re-read before every cycle, so a reload to zero or a negative value stops a running loop. A value above about 49.7 days (0xFFFFFFFE milliseconds, the longest delay a timer accepts) is clamped to it. |
TenantObservabilityOptions
Governs the per-tenant gauges described under Observability.
| Property | Type | Default | Meaning |
|---|---|---|---|
PublishGauges |
bool |
true |
Whether to publish the per-tenant observable gauges on the orleans.lattice.tenancy meter. false leaves the meter inert and skips the periodic overage scan. |
PublishInterval |
TimeSpan |
30s (DefaultPublishInterval) |
How often the publisher re-samples the warm usage index and the overage billing seam. A non-positive value is treated as the default, and a value above about 49.7 days (the longest period a timer accepts) is clamped to it. |
LatticeTenantRateLimiterOptions
Governs how a tenant's cluster-wide MaxOpsPerSecond is divided across live
silos. See Rate limiting for the full table.
See also
Orleans.Lattice.Api.TenantAdmin- the transport-agnostic tenant-administration and region-residency control facades.Orleans.Lattice.Api.TenantAdmin.Grpc- the code-first gRPC binding and remote client for the tenant-administration facade.Orleans.Lattice.Explorer- the web UI whose tenant view roots each tenant-scoped address at the active tenant and offers a platform operator a tenant switcher.Orleans.Lattice.Api.Mcp- the MCP server that contributes the read-only tenant self-awareness tools.Orleans.Lattice.Auth- the authorization gate the tenant boundary is enforced at.Orleans.Lattice.Membership- the identity layer that resolves the caller subject a tenant's admin-subject set is matched against.- MultiTenancy sample - a runnable end-to-end walkthrough of opt-in wire-up, tenant tree naming, the tenant lifecycle, and the fail-closed control plane.