---
title: "Remote hosting"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.api.mcp/remote.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice.api.mcp/remote.md"
package: "Orleans.Lattice.Api.Mcp"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.api.mcp/llms-full.txt"
---
# Remote hosting

Part of the [Api.Mcp documentation](README.md).

The MCP server can run **in the silo** (co-hosted with the facades it binds, resolving them in-process) or **out of the silo** as a standalone host that reaches the cluster over the network. `AddLatticeMcpRemote(...)` wires the out-of-silo topology: the same built-in tool modules, bound over the `Orleans.Lattice.Api.*.Grpc` clients instead of the in-process facades.

## When to use it

Use remote hosting when the MCP endpoint cannot live on a cluster silo - for example a dedicated agent-gateway process, a host in a different trust zone, or a single MCP front door fronting a cluster it is not a member of. When the MCP server can co-host on a silo, prefer the in-process topology ([Setup](setup.md)): it avoids a network hop and the extra credential-forwarding configuration below.

## Wiring

`AddLatticeMcpRemote(...)` registers the MCP infrastructure (it calls `AddLatticeMcp` internally), the gRPC-backed facade adapters, the credential-forwarding interceptor, and, for each configured group, the matching tool module (telemetry excepted - see the `Telemetry` option below). Configure an endpoint only for the groups you want to serve:

```csharp verify
var builder = WebApplication.CreateBuilder();

builder.Services.AddLatticeMcpRemote(o =>
{
    o.State = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster-a.internal:5001" };
    o.Data = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster-a.internal:5001" };
    o.EnableDataWrites = false;

    // Required for any caller's tools to be discovered remotely (see below).
    o.Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster-a.internal:5001" };

    // Runtime per-tree replication control (inspect always; enable/disable gated).
    o.Replication = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster-a.internal:5001" };
    o.EnableReplicationControl = true;
});

var app = builder.Build();
app.MapLatticeMcp();
```

Each `LatticeApiMcpRemoteEndpoint` names the served `Endpoint` (surfaced verbatim in the `lattice_capabilities` report) and optionally supplies a pre-built `CallInvoker` so a host that already owns a tuned gRPC channel (custom TLS, retries, deadlines) can reuse it instead of the address-derived default.

The remote binding keeps the same coarse authorizer seam as the in-silo one: the default authorizer denies every group tool, so register a permissive or custom `ILatticeApiMcpAuthorizer` before serving (see [Security](security.md#2-the-coarse-authorizer-seam)).

## Options

| Option | Purpose |
|---|---|
| `State` / `Data` / `Auth` / `Backup` / `Replication` / `TreeAdmin` / `TenantAdmin` / `Telemetry` | The per-group remote endpoint, or `null` to not serve that group. The tree-administration endpoint backs every tree-administration tool - the read-only diagnostics tools, the lifecycle and control tools (`lattice_treeadmin_tree_*` plus the bulk-load, WAL placement / move, orphaned-leaf, view, tag-index, compaction, and retention tools), and the schema tools (`lattice_treeadmin_schema_*`) - since the tree-administration-API and schema-API gRPC services are co-hosted on the same silo address. The `TenantAdmin` endpoint backs both the read-only tenant self-awareness tools (`lattice_tenant_current` / `lattice_tenant_list` / `lattice_tenant_get`) and, when `EnableTenantControl` is set, the tenant-admin control tools (`lattice_tenant_create` / `lattice_tenant_suspend` / `lattice_tenant_resume` / `lattice_tenant_delete` / `lattice_tenant_set_quotas` / `lattice_tenant_authorize_regions` / `lattice_tenant_set_residency` / `lattice_tenant_region_status`), since the self-service read RPCs are co-hosted on the tenant-administration gRPC service address. The `Telemetry` endpoint is the one group endpoint the binding wires no gRPC adapter or tool module for: it names where the region's telemetry facade is served and feeds only discovery and region routing (`lattice_capabilities`, `lattice_list_regions`). The `lattice_telemetry_*` tools themselves come solely from the co-located telemetry tool module (`AddTelemetryTools`), which queries this head's own configured metrics backend. The current region advertises the telemetry group when either the endpoint or that module is present - at the configured endpoint when one is set, because it is the one another process can route to - and a peer region never inherits this head's co-located module. |
| `CredentialHeaderName` | Header the resolved caller credential is stamped onto for the outbound call. Defaults to `authorization`. |
| `CredentialScheme` | Scheme prefix prepended to the outbound token (`"{scheme} {token}"`). Defaults to `Bearer`; empty sends the bare token. |
| `ActiveTenantHeaderName` | Header the caller's ambient active tenant is forwarded on for the outbound call, so the remote cluster's per-tenant write admission and quota enforcement reach the caller's tenant. Defaults to `lattice-active-tenant`; empty disables forwarding. See [Credential flow over the wire](#credential-flow-over-the-wire). |
| `AdministratorCredential` | The **static** admin service credential used for trusted, read-only permission introspection of each caller. See [discovery](#discovery-requires-the-auth-endpoint) below. For a long-lived server prefer a self-refreshing managed-identity token (see [Refreshing administrator token](#refreshing-the-administrator-token)). |
| `EnableDataWrites` / `EnableBackupControl` / `EnableAuthAdministration` / `EnableReplicationControl` / `EnableSchemaControl` / `EnableLifecycleControl` / `EnableTenantControl` | Forward the destructive-verb opt-in to the corresponding tool module; each defaults to `false`. Ignored when that group's endpoint is unset. `EnableSchemaControl` gates the mutating `lattice_treeadmin_schema_*` tools and `EnableLifecycleControl` gates the tree-administration lifecycle/control mutation tools; both are ignored when `TreeAdmin` is unset. `EnableTenantControl` gates the mutating `lattice_tenant_*` admin tools and is ignored when `TenantAdmin` is unset; the read-only tenant self-awareness tools are served whenever `TenantAdmin` is set, with no flag. |
| `RegionId` | The id of the current (default) region a call targets when no `region` selector is supplied. Defaults to `current`. |
| `ClusterId` | The Orleans cluster id of the current region, surfaced in `lattice_list_regions`. Optional advertisement metadata; when unset, the discovery tool resolves it from the state facade at read time. |
| `Regions` | Additional peer regions a caller may target with the optional per-call `region` argument; empty by default. Each is a `LatticeApiMcpRemoteRegionOptions` with its own required `RegionId`, optional `ClusterId`, and the same eight optional per-group endpoints as the top level (`State` / `Data` / `Auth` / `Backup` / `Replication` / `TreeAdmin` / `TenantAdmin` / `Telemetry`). Keep each `RegionId` unique: the binding does not reject a duplicate. |
| `VerifyRegionIdentity` | When `true`, each peer region's endpoint is probed once and its reported cluster id checked against the region's advertised `ClusterId` before any call is routed there; a region that reaches the wrong cluster is omitted from `lattice_list_regions` and rejected fail-closed. A probe that cannot reach the region fails that listing and call closed but is not cached, so a later call re-probes; a peer with no advertised `ClusterId` or no `State` endpoint cannot be verified and stays routable, and so does every peer when this head configures no top-level `State` endpoint, because the probe runs through the head's own state client. Defaults to `false`. See [Region targeting behind a global load balancer](#region-targeting-behind-a-global-load-balancer). |

## Multi-region routing

By default the endpoints above define a single region - the current cluster. To let one MCP head front more than one region, add a `LatticeApiMcpRemoteRegionOptions` per peer to `Regions`. The top-level endpoints stay the default (current) region, so an existing single-region configuration is unchanged; the peers are additive and opt-in per call.

```csharp verify
var builder = WebApplication.CreateBuilder();

builder.Services.AddLatticeMcpRemote(o =>
{
    // The current (default) region, targeted when no `region` is supplied.
    o.RegionId = "us-east";
    o.ClusterId = "cluster-us-east";
    o.Data = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://us-east.internal:5001" };
    o.Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://us-east.internal:5001" };

    // A reachable peer region a caller may target with `"region": "eu-west"`.
    // Each peer endpoint is that region's OWN, region-pinned silo FQDN - never a
    // shared/anycast endpoint - so the `region` selector is deterministic.
    o.Regions.Add(new LatticeApiMcpRemoteRegionOptions
    {
        RegionId = "eu-west",
        ClusterId = "cluster-eu-west",
        Data = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://eu-west.internal:5001" },
        Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://eu-west.internal:5001" },
    });

    // Prove each peer reaches the cluster it advertises before routing to it.
    o.VerifyRegionIdentity = true;
});

var app = builder.Build();
app.MapLatticeMcp();
```

The region list `lattice_list_regions` reports, and the routing a `region` argument drives, are both built once at startup from this single configuration, so discovery and routing can never disagree. A region serves a facade group only when that group's per-region endpoint is set (the sole exception is `Telemetry` on the current region, which is also served by a co-located tool module when no endpoint is configured); an unset group is reported unavailable for the region and is not routable there (fail-closed discovery). A cross-region call forwards the same caller credential to the target region's gRPC binding via the same interceptor described below, so the target authorizes it independently. See [Tools](tools.md#region-targeting) for the caller-facing surface.

### Region targeting interacts with tenant residency

On a cluster running the tenancy add-on, the configured topology above is the **physical** region set. It is not what a tenant caller may target. Once a tenant's residency has been set, it is served only from the regions where its status is exactly `Online` - a status no shipped component advances a region to (see [Tools](tools.md#tenant-region-residency-lattice_tenant_authorize_regions-lattice_tenant_set_residency-lattice_tenant_region_status)) - while a tenant that has never set residency is served in every region. So:

- **A region a tenant is not online in refuses the call.** Targeting it with a `region` argument reaches the peer, and the peer's residency gate refuses it. Configuration reachability and tenant reachability are different questions, and only the data owner can answer the second.
- **`lattice_list_regions` does not currently return a tenant-scoped list.** Its tenant-scoped answer - the calling tenant's actionable set (`allowed` union `resident`) plus the current region, each annotated with the tenant's standing - is served only for a tenant assertion the head validates, and no shipped registration validates one: the discovery tool runs without the caller's credential, and a remote head registers no validating resolver. A tenant-asserting call is therefore answered with the current region alone and no annotation, or, by a head without the tenant-admin endpoint (next bullet), with the full configured topology. In a scoped answer an `isResident: false` entry is a valid `lattice_tenant_set_residency` destination but not yet a valid routing destination. See [Tools](tools.md#tenant-scoped-region-discovery).
- **A remote head needs the tenant-admin endpoint to scope discovery.** Set `LatticeApiMcpRemoteOptions.TenantAdmin` and the composition registers an `ITenantRegionVisibilityResolver` (replacing any core default) that resolves the caller's standing over the region-residency RPC; when that lookup cannot establish the tenant's standing, a tenant-asserted `lattice_list_regions` fails closed to the current region alone. The lookup is never made for an assertion the head cannot validate: a non-default tenant is honoured only when the head's `ITenantContextResolver` resolves it as the caller's own, and a refused assertion - or one made to a head with no validating resolver - is answered with the current region alone and no `tenantScope` annotation. The remote composition registers no validating resolver, so today every non-default assertion to a head with the endpoint is answered that way and the standing lookup is not made. Without the endpoint the head has no way to resolve standing, so it answers every call unscoped, exactly as a non-tenancy head does: a tenant-asserted `lattice_list_regions` returns the full configured topology. On a tenancy estate, configure `TenantAdmin` on every remote head tenant callers reach. The lookup is made only when a call asserts a non-default tenant, so an operator call and a tenancy-off head never pay the round trip.

## Region targeting behind a global load balancer

Region targeting is **deterministic** - a `region` selector must reach that exact region. That is fundamentally at odds with a global anycast load balancer (for example Azure Front Door), whose job is to *hide* which region serves a request by latency-routing to the nearest healthy origin. So a peer region must be addressed by its **own, region-pinned endpoint** (its silo's direct gRPC FQDN), never a shared Front Door endpoint. Point a region at an anycast endpoint and a call targeting it lands on whichever region the load balancer picks, and the served-region annotation becomes untrustworthy.

Two consequences when the deployment fronts its regions with a load balancer that enforces an origin lock (a required header such as `X-Azure-FDID`):

- **Stamp the origin-lock header on the direct dial yourself.** Reach the peer's internal gRPC ingress directly and supply a pre-built `CallInvoker` (via `LatticeApiMcpRemoteEndpoint.CallInvoker`) that adds the required header - one global Front Door id typically validates every regional origin, so the same invoker works for every peer. Do **not** route gRPC *through* the load balancer: it would inject the header itself (a duplicate the origin lock rejects) and gRPC-through-Front-Door is not generally supported.
- **Turn on `VerifyRegionIdentity`.** It probes each peer's state facade once and checks the reported cluster id against the region's advertised `ClusterId`. A region whose endpoint reaches the wrong cluster - the exact symptom of a mis-pointed or anycast endpoint - is omitted from `lattice_list_regions` and rejected fail-closed when targeted, so a misconfiguration surfaces as a clean discovery gap rather than a call silently answered by the wrong region.

The reference architecture already builds one origin-lock `CallInvoker` per silo FQDN that stamps `X-Azure-FDID` on every outbound gRPC call (its MCP head dials the silo directly, not through Front Door). That same per-endpoint invoker is the hook a multi-region deployment hands to each peer region's `LatticeApiMcpRemoteEndpoint.CallInvoker`, alongside `VerifyRegionIdentity`, to target regions deterministically behind a Front Door estate.

## Credential flow over the wire

The remote binding's credential-forwarding interceptor stamps the resolved caller credential onto each outbound gRPC call as `{scheme} {token}`. It resolves the credential in order: an administrator service credential for a system-origin introspection call, then the ambient `LatticeCredentialContext` stamped by the tool, then the `HttpContext` credential bridge, then anonymous. The remote cluster's gRPC binding then authenticates that credential and applies its own per-tree / per-key access gate - so enforcement still happens at the data owner, not at the MCP host.

The same interceptor also forwards the caller's ambient active tenant. When the tool invocation stamped an active tenant (from the inbound `ActiveTenantHeaderName` header), it is re-emitted on the outbound call as the `ActiveTenantHeaderName` metadata header (`lattice-active-tenant` by default), so the remote cluster's data binding lifts it back onto its own ambient scope and per-tenant write admission / quota enforcement runs at the data owner. The tenant is an assertion the remote tenancy add-on re-validates against the caller's membership; when no tenant is asserted no tenant header is added and the outbound call is byte-for-byte unchanged.

## Discovery requires the auth endpoint

The in-silo permission-scoped discovery relies on a **system-origin bypass** to introspect a caller's effective permissions. That bypass does not cross the wire. Remotely, the discovery core must authenticate as an administrator to introspect a non-administrator caller, so serving any group's tools to non-administrator callers requires both the `Auth` endpoint and an administrator credential to be configured - the static `AdministratorCredential`, or the self-refreshing managed-identity source described [below](#refreshing-the-administrator-token). Without the `Auth` endpoint no caller - administrator or not - is offered any group's tools; with it but without an administrator credential, only an administrator caller can enumerate tools remotely.

## Refreshing the administrator token

`AdministratorCredential` is a **static** token. When acquired from Entra it typically carries a ~1h lifetime, so a long-lived remote MCP head silently loses its introspection capability once it expires (discovery then advertises no group tools to any caller - administrators included, since every introspection still forwards the expired token - until the process is restarted or the value is rotated by hand). For an always-on server, register the managed-identity administrator source instead: it acquires the silo-audience token from an `Azure.Core` `TokenCredential`, caches it, and refreshes it a configurable skew before expiry.

```csharp verify
using Azure.Core;

var builder = WebApplication.CreateBuilder();
TokenCredential credential = CreateManagedIdentityCredential();

builder.Services.AddLatticeMcpRemote(o =>
{
    o.Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster.internal:5001" };
    // No static o.AdministratorCredential needed.
});

static TokenCredential CreateManagedIdentityCredential()
    => throw new NotImplementedException(
        "Use ManagedIdentityCredential or DefaultAzureCredential from Azure.Identity.");

builder.Services.AddLatticeMcpManagedIdentityAdministrator(o =>
{
    o.Credential = credential;
    o.Scope = "api://<silo-app-id>/.default";            // the remote silo audience
    o.RefreshSkew = TimeSpan.FromMinutes(5);             // optional; defaults to 5 minutes
});
```

The managed-identity source takes precedence over `AdministratorCredential` regardless of registration order. It is **fail-closed**: if token acquisition fails it forwards no administrator credential - the introspection call falls back to the caller's own credential, so only an administrator caller can still enumerate tools - and it self-heals on the next successful acquisition rather than forwarding a stale token.

`LatticeApiMcpManagedIdentityAdministratorOptions` (populated through the `AddLatticeMcpManagedIdentityAdministrator` delegate):

| Option | Type | Default | Purpose |
|---|---|---|---|
| `Credential` | `TokenCredential?` | none (required) | The `Azure.Core` credential the administrator token is acquired from, for example `new ManagedIdentityCredential()` or `new DefaultAzureCredential()`. Only the bearer token it produces is forwarded to the remote cluster. |
| `Scope` | `string` | `""` (required) | The scope the token is requested for - the remote silo's audience, for example `api://<silo-app-id>/.default`. Must be non-empty and non-whitespace. |
| `RefreshSkew` | `TimeSpan` | 5 minutes | How long before a cached token's expiry the source proactively acquires a fresh one. Must not be negative. |

A missing `Credential`, a blank `Scope`, or a negative `RefreshSkew` fails options validation.

## OAuth discovery (RFC 9728)

A remote head is the common place to advertise OAuth discovery, because a client that connects to it over the internet has no pre-shared token. `AddLatticeMcpRemote` wires the base MCP binding (including the discovery endpoint and challenge hint), so opt in by layering the `ProtectedResourceMetadata` option onto the shared `LatticeApiMcpOptions` with an additive `AddLatticeMcp` call. See [Setup](setup.md#oauth-discovery-rfc-9728) for what each field means and [Security](security.md#oauth-discovery-is-anonymous-by-design) for why the metadata endpoint is anonymous.

```csharp verify
var builder = WebApplication.CreateBuilder();

builder.Services.AddLatticeMcpRemote(o =>
{
    o.State = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster.internal:5001" };
    o.Auth = new LatticeApiMcpRemoteEndpoint { Endpoint = "https://cluster.internal:5001" };
});

// Additive: layer discovery onto the same options the remote binding registered.
builder.Services.AddLatticeMcp(o =>
{
    o.ProtectedResourceMetadata = new LatticeApiMcpProtectedResourceMetadata
    {
        Resource = new Uri("https://mcp.example.com"),
        AuthorizationServers = { new Uri("https://login.microsoftonline.com/<tenant>/v2.0") },
        ScopesSupported = { "api://<server-app-id>/.default" },
    };
});

var app = builder.Build();
app.MapLatticeMcp();
```

## Deferred tools

A few tools back facade operations that have no gRPC method yet, so they cannot be served remotely. The remote host defers them - they are simply omitted from the remote tool list rather than advertised and then failing. Currently deferred: `lattice_state_get_tree_summary`, `lattice_state_get_shard_summaries`, `lattice_state_get_physical_shard_count`, and `lattice_backup_inventory`. They remain fully available in the in-silo topology; each becomes discoverable remotely with no other change once its gRPC method is bound.

## Next

- [Setup](setup.md) - the in-silo topology.
- [Security](security.md) - the fail-closed posture the remote host preserves.
