Remote hosting
This page documents Orleans.Lattice.Api.Mcp 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at remote.md, and llms.txt lists every page.The MCP server can run in the silo (co-hosted with the facades it binds, resolving them in-process) or out of the silo as a standalone host that reaches the cluster over the network. AddLatticeMcpRemote(...) wires the out-of-silo topology: the same built-in tool modules, bound over the Orleans.Lattice.Api.*.Grpc clients instead of the in-process facades.
When to use it
Use remote hosting when the MCP endpoint cannot live on a cluster silo - for example a dedicated agent-gateway process, a host in a different trust zone, or a single MCP front door fronting a cluster it is not a member of. When the MCP server can co-host on a silo, prefer the in-process topology (Setup): it avoids a network hop and the extra credential-forwarding configuration below.
Wiring
AddLatticeMcpRemote(...) registers the MCP infrastructure (it calls AddLatticeMcp internally), the gRPC-backed facade adapters, the credential-forwarding interceptor, and, for each configured group, the matching tool module (telemetry excepted - see the Telemetry option below). Configure an endpoint only for the groups you want to serve:
var builder = WebApplication.CreateBuilder();
builder.Services.AddLatticeMcpRemote(o =>
{
o.State = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster-a.internal:5001" };
o.Data = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster-a.internal:5001" };
o.EnableDataWrites = false;
// Required for any caller's tools to be discovered remotely (see below).
o.Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster-a.internal:5001" };
// Runtime per-tree replication control (inspect always; enable/disable gated).
o.Replication = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster-a.internal:5001" };
o.EnableReplicationControl = true;
});
var app = builder.Build();
app.MapLatticeMcp();
Each LatticeApiMcpRemoteEndpoint names the served Endpoint (surfaced verbatim in the lattice_capabilities report) and optionally supplies a pre-built CallInvoker so a host that already owns a tuned gRPC channel (custom TLS, retries, deadlines) can reuse it instead of the address-derived default.
The remote binding keeps the same coarse authorizer seam as the in-silo one: the default authorizer denies every group tool, so register a permissive or custom ILatticeApiMcpAuthorizer before serving (see Security).
Options
| Option | Purpose |
|---|---|
State / Data / Auth / Backup / Replication / TreeAdmin / TenantAdmin / Telemetry |
The per-group remote endpoint, or null to not serve that group. The tree-administration endpoint backs every tree-administration tool - the read-only diagnostics tools, the lifecycle and control tools (lattice_treeadmin_tree_* plus the bulk-load, WAL placement / move, orphaned-leaf, view, tag-index, compaction, and retention tools), and the schema tools (lattice_treeadmin_schema_*) - since the tree-administration-API and schema-API gRPC services are co-hosted on the same silo address. The TenantAdmin endpoint backs both the read-only tenant self-awareness tools (lattice_tenant_current / lattice_tenant_list / lattice_tenant_get) and, when EnableTenantControl is set, the tenant-admin control tools (lattice_tenant_create / lattice_tenant_suspend / lattice_tenant_resume / lattice_tenant_delete / lattice_tenant_set_quotas / lattice_tenant_authorize_regions / lattice_tenant_set_residency / lattice_tenant_region_status), since the self-service read RPCs are co-hosted on the tenant-administration gRPC service address. The Telemetry endpoint is the one group endpoint the binding wires no gRPC adapter or tool module for: it names where the region's telemetry facade is served and feeds only discovery and region routing (lattice_capabilities, lattice_list_regions). The lattice_telemetry_* tools themselves come solely from the co-located telemetry tool module (AddTelemetryTools), which queries this head's own configured metrics backend. The current region advertises the telemetry group when either the endpoint or that module is present - at the configured endpoint when one is set, because it is the one another process can route to - and a peer region never inherits this head's co-located module. |
CredentialHeaderName |
Header the resolved caller credential is stamped onto for the outbound call. Defaults to authorization. |
CredentialScheme |
Scheme prefix prepended to the outbound token ("{scheme} {token}"). Defaults to Bearer; empty sends the bare token. |
ActiveTenantHeaderName |
Header the caller's ambient active tenant is forwarded on for the outbound call, so the remote cluster's per-tenant write admission and quota enforcement reach the caller's tenant. Defaults to lattice-active-tenant; empty disables forwarding. See Credential flow over the wire. |
AdministratorCredential |
The static admin service credential used for trusted, read-only permission introspection of each caller. See discovery below. For a long-lived server prefer a self-refreshing managed-identity token (see Refreshing administrator token). |
EnableDataWrites / EnableBackupControl / EnableAuthAdministration / EnableReplicationControl / EnableSchemaControl / EnableLifecycleControl / EnableTenantControl |
Forward the destructive-verb opt-in to the corresponding tool module; each defaults to false. Ignored when that group's endpoint is unset. EnableSchemaControl gates the mutating lattice_treeadmin_schema_* tools and EnableLifecycleControl gates the tree-administration lifecycle/control mutation tools; both are ignored when TreeAdmin is unset. EnableTenantControl gates the mutating lattice_tenant_* admin tools and is ignored when TenantAdmin is unset; the read-only tenant self-awareness tools are served whenever TenantAdmin is set, with no flag. |
RegionId |
The id of the current (default) region a call targets when no region selector is supplied. Defaults to current. |
ClusterId |
The Orleans cluster id of the current region, surfaced in lattice_list_regions. Optional advertisement metadata; when unset, the discovery tool resolves it from the state facade at read time. |
Regions |
Additional peer regions a caller may target with the optional per-call region argument; empty by default. Each is a LatticeApiMcpRemoteRegionOptions with its own required RegionId, optional ClusterId, and the same eight optional per-group endpoints as the top level (State / Data / Auth / Backup / Replication / TreeAdmin / TenantAdmin / Telemetry). Keep each RegionId unique: the binding does not reject a duplicate. |
VerifyRegionIdentity |
When true, each peer region's endpoint is probed once and its reported cluster id checked against the region's advertised ClusterId before any call is routed there; a region that reaches the wrong cluster is omitted from lattice_list_regions and rejected fail-closed. A probe that cannot reach the region fails that listing and call closed but is not cached, so a later call re-probes; a peer with no advertised ClusterId or no State endpoint cannot be verified and stays routable, and so does every peer when this head configures no top-level State endpoint, because the probe runs through the head's own state client. Defaults to false. See Region targeting behind a global load balancer. |
Multi-region routing
By default the endpoints above define a single region - the current cluster. To let one MCP head front more than one region, add a LatticeApiMcpRemoteRegionOptions per peer to Regions. The top-level endpoints stay the default (current) region, so an existing single-region configuration is unchanged; the peers are additive and opt-in per call.
var builder = WebApplication.CreateBuilder();
builder.Services.AddLatticeMcpRemote(o =>
{
// The current (default) region, targeted when no `region` is supplied.
o.RegionId = "us-east";
o.ClusterId = "cluster-us-east";
o.Data = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://us-east.internal:5001" };
o.Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://us-east.internal:5001" };
// A reachable peer region a caller may target with `"region": "eu-west"`.
// Each peer endpoint is that region's OWN, region-pinned silo FQDN - never a
// shared/anycast endpoint - so the `region` selector is deterministic.
o.Regions.Add(new LatticeApiMcpRemoteRegionOptions
{
RegionId = "eu-west",
ClusterId = "cluster-eu-west",
Data = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://eu-west.internal:5001" },
Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://eu-west.internal:5001" },
});
// Prove each peer reaches the cluster it advertises before routing to it.
o.VerifyRegionIdentity = true;
});
var app = builder.Build();
app.MapLatticeMcp();
The region list lattice_list_regions reports, and the routing a region argument drives, are both built once at startup from this single configuration, so discovery and routing can never disagree. A region serves a facade group only when that group's per-region endpoint is set (the sole exception is Telemetry on the current region, which is also served by a co-located tool module when no endpoint is configured); an unset group is reported unavailable for the region and is not routable there (fail-closed discovery). A cross-region call forwards the same caller credential to the target region's gRPC binding via the same interceptor described below, so the target authorizes it independently. See Tools for the caller-facing surface.
Region targeting interacts with tenant residency
On a cluster running the tenancy add-on, the configured topology above is the physical region set. It is not what a tenant caller may target. Once a tenant's residency has been set, it is served only from the regions where its status is exactly Online - a status no shipped component advances a region to (see Tools) - while a tenant that has never set residency is served in every region. So:
- A region a tenant is not online in refuses the call. Targeting it with a
regionargument reaches the peer, and the peer's residency gate refuses it. Configuration reachability and tenant reachability are different questions, and only the data owner can answer the second. lattice_list_regionsdoes not currently return a tenant-scoped list. Its tenant-scoped answer - the calling tenant's actionable set (allowedunionresident) plus the current region, each annotated with the tenant's standing - is served only for a tenant assertion the head validates, and no shipped registration validates one: the discovery tool runs without the caller's credential, and a remote head registers no validating resolver. A tenant-asserting call is therefore answered with the current region alone and no annotation, or, by a head without the tenant-admin endpoint (next bullet), with the full configured topology. In a scoped answer anisResident: falseentry is a validlattice_tenant_set_residencydestination but not yet a valid routing destination. See Tools.- A remote head needs the tenant-admin endpoint to scope discovery. Set
LatticeApiMcpRemoteOptions.TenantAdminand the composition registers anITenantRegionVisibilityResolver(replacing any core default) that resolves the caller's standing over the region-residency RPC; when that lookup cannot establish the tenant's standing, a tenant-assertedlattice_list_regionsfails closed to the current region alone. The lookup is never made for an assertion the head cannot validate: a non-default tenant is honoured only when the head'sITenantContextResolverresolves it as the caller's own, and a refused assertion - or one made to a head with no validating resolver - is answered with the current region alone and notenantScopeannotation. The remote composition registers no validating resolver, so today every non-default assertion to a head with the endpoint is answered that way and the standing lookup is not made. Without the endpoint the head has no way to resolve standing, so it answers every call unscoped, exactly as a non-tenancy head does: a tenant-assertedlattice_list_regionsreturns the full configured topology. On a tenancy estate, configureTenantAdminon every remote head tenant callers reach. The lookup is made only when a call asserts a non-default tenant, so an operator call and a tenancy-off head never pay the round trip.
Region targeting behind a global load balancer
Region targeting is deterministic - a region selector must reach that exact region. That is fundamentally at odds with a global anycast load balancer (for example Azure Front Door), whose job is to hide which region serves a request by latency-routing to the nearest healthy origin. So a peer region must be addressed by its own, region-pinned endpoint (its silo's direct gRPC FQDN), never a shared Front Door endpoint. Point a region at an anycast endpoint and a call targeting it lands on whichever region the load balancer picks, and the served-region annotation becomes untrustworthy.
Two consequences when the deployment fronts its regions with a load balancer that enforces an origin lock (a required header such as X-Azure-FDID):
- Stamp the origin-lock header on the direct dial yourself. Reach the peer's internal gRPC ingress directly and supply a pre-built
CallInvoker(viaLatticeApiMcpRemoteEndpoint.CallInvoker) that adds the required header - one global Front Door id typically validates every regional origin, so the same invoker works for every peer. Do not route gRPC through the load balancer: it would inject the header itself (a duplicate the origin lock rejects) and gRPC-through-Front-Door is not generally supported. - Turn on
VerifyRegionIdentity. It probes each peer's state facade once and checks the reported cluster id against the region's advertisedClusterId. A region whose endpoint reaches the wrong cluster - the exact symptom of a mis-pointed or anycast endpoint - is omitted fromlattice_list_regionsand rejected fail-closed when targeted, so a misconfiguration surfaces as a clean discovery gap rather than a call silently answered by the wrong region.
The reference architecture already builds one origin-lock CallInvoker per silo FQDN that stamps X-Azure-FDID on every outbound gRPC call (its MCP head dials the silo directly, not through Front Door). That same per-endpoint invoker is the hook a multi-region deployment hands to each peer region's LatticeApiMcpRemoteEndpoint.CallInvoker, alongside VerifyRegionIdentity, to target regions deterministically behind a Front Door estate.
Credential flow over the wire
The remote binding's credential-forwarding interceptor stamps the resolved caller credential onto each outbound gRPC call as {scheme} {token}. It resolves the credential in order: an administrator service credential for a system-origin introspection call, then the ambient LatticeCredentialContext stamped by the tool, then the HttpContext credential bridge, then anonymous. The remote cluster's gRPC binding then authenticates that credential and applies its own per-tree / per-key access gate - so enforcement still happens at the data owner, not at the MCP host.
The same interceptor also forwards the caller's ambient active tenant. When the tool invocation stamped an active tenant (from the inbound ActiveTenantHeaderName header), it is re-emitted on the outbound call as the ActiveTenantHeaderName metadata header (lattice-active-tenant by default), so the remote cluster's data binding lifts it back onto its own ambient scope and per-tenant write admission / quota enforcement runs at the data owner. The tenant is an assertion the remote tenancy add-on re-validates against the caller's membership; when no tenant is asserted no tenant header is added and the outbound call is byte-for-byte unchanged.
Discovery requires the auth endpoint
The in-silo permission-scoped discovery relies on a system-origin bypass to introspect a caller's effective permissions. That bypass does not cross the wire. Remotely, the discovery core must authenticate as an administrator to introspect a non-administrator caller, so serving any group's tools to non-administrator callers requires both the Auth endpoint and an administrator credential to be configured - the static AdministratorCredential, or the self-refreshing managed-identity source described below. Without the Auth endpoint no caller - administrator or not - is offered any group's tools; with it but without an administrator credential, only an administrator caller can enumerate tools remotely.
Refreshing the administrator token
AdministratorCredential is a static token. When acquired from Entra it typically carries a ~1h lifetime, so a long-lived remote MCP head silently loses its introspection capability once it expires (discovery then advertises no group tools to any caller - administrators included, since every introspection still forwards the expired token - until the process is restarted or the value is rotated by hand). For an always-on server, register the managed-identity administrator source instead: it acquires the silo-audience token from an Azure.Core TokenCredential, caches it, and refreshes it a configurable skew before expiry.
using Azure.Core;
var builder = WebApplication.CreateBuilder();
TokenCredential credential = CreateManagedIdentityCredential();
builder.Services.AddLatticeMcpRemote(o =>
{
o.Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster.internal:5001" };
// No static o.AdministratorCredential needed.
});
static TokenCredential CreateManagedIdentityCredential()
=> throw new NotImplementedException(
"Use ManagedIdentityCredential or DefaultAzureCredential from Azure.Identity.");
builder.Services.AddLatticeMcpManagedIdentityAdministrator(o =>
{
o.Credential = credential;
o.Scope = "api://<silo-app-id>/.default"; // the remote silo audience
o.RefreshSkew = TimeSpan.FromMinutes(5); // optional; defaults to 5 minutes
});
The managed-identity source takes precedence over AdministratorCredential regardless of registration order. It is fail-closed: if token acquisition fails it forwards no administrator credential - the introspection call falls back to the caller's own credential, so only an administrator caller can still enumerate tools - and it self-heals on the next successful acquisition rather than forwarding a stale token.
LatticeApiMcpManagedIdentityAdministratorOptions (populated through the AddLatticeMcpManagedIdentityAdministrator delegate):
| Option | Type | Default | Purpose |
|---|---|---|---|
Credential |
TokenCredential? |
none (required) | The Azure.Core credential the administrator token is acquired from, for example new ManagedIdentityCredential() or new DefaultAzureCredential(). Only the bearer token it produces is forwarded to the remote cluster. |
Scope |
string |
"" (required) |
The scope the token is requested for - the remote silo's audience, for example api://<silo-app-id>/.default. Must be non-empty and non-whitespace. |
RefreshSkew |
TimeSpan |
5 minutes | How long before a cached token's expiry the source proactively acquires a fresh one. Must not be negative. |
A missing Credential, a blank Scope, or a negative RefreshSkew fails options validation.
OAuth discovery (RFC 9728)
A remote head is the common place to advertise OAuth discovery, because a client that connects to it over the internet has no pre-shared token. AddLatticeMcpRemote wires the base MCP binding (including the discovery endpoint and challenge hint), so opt in by layering the ProtectedResourceMetadata option onto the shared LatticeApiMcpOptions with an additive AddLatticeMcp call. See Setup for what each field means and Security for why the metadata endpoint is anonymous.
var builder = WebApplication.CreateBuilder();
builder.Services.AddLatticeMcpRemote(o =>
{
o.State = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster.internal:5001" };
o.Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster.internal:5001" };
});
// Additive: layer discovery onto the same options the remote binding registered.
builder.Services.AddLatticeMcp(o =>
{
o.ProtectedResourceMetadata = new LatticeApiMcpProtectedResourceMetadata
{
Resource = new Uri("https://mcp.example.com"),
AuthorizationServers = { new Uri("https://login.microsoftonline.com/<tenant>/v2.0") },
ScopesSupported = { "api://<server-app-id>/.default" },
};
});
var app = builder.Build();
app.MapLatticeMcp();
Deferred tools
A few tools back facade operations that have no gRPC method yet, so they cannot be served remotely. The remote host defers them - they are simply omitted from the remote tool list rather than advertised and then failing. Currently deferred: lattice_state_get_tree_summary, lattice_state_get_shard_summaries, lattice_state_get_physical_shard_count, and lattice_backup_inventory. They remain fully available in the in-silo topology; each becomes discoverable remotely with no other change once its gRPC method is bound.