Orleans.Lattice reference architecture: deploy and configuration guide

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.

This is the operator guide for the active-active, cross-region Orleans.Lattice reference estate on Azure Container Apps (ACA). It covers prerequisites, the parameter reference, deploy/verify/teardown for both network options, day-2 operations, the optional hardening and upgrade path, and the real-Azure validation runbook.

For the design, the topology diagrams, the consistency scoping, and the rationale behind every choice, read the design document first: ../reference-architecture.md. This guide does not duplicate the design; it tells you how to run the kit.

What is in the kit

Folder Contents
bicep/ main.bicep orchestrator, the per-concern modules (compute, storage, networking, vnet, privatedns, observability, frontdoor), bootstrap.bicep (registry pre-build seam), and entra/ (the Microsoft Graph extension module + its scoped bicepconfig.json), plus example parameter sets (main.bicepparam and params/).
hosts/ The three container host projects - Silo, MCP, and Explorer - each with a chiselled, non-root Dockerfile. They reference the published Orleans.Lattice NuGet packages, plus the shared Common/ hosting library (the Front Door origin lock and probe helpers, tested in Common.Tests/). See hosts/README.md for the full host configuration surface.
deploy/ Deploy-ReferenceArchitecture.ps1, the single idempotent orchestrator; deployment-sample.ps1, a zero-decision wrapper that deploys a three-region sample estate from a deployment name; and deploy/README.md documenting both.
local/ A Docker Compose harness that stands the whole estate up on one machine for development. See local/README.md.
local-dev/ A two-region variant of the local harness that builds every head straight from src/** by project reference (no published Orleans.Lattice packages), runs two network-isolated clusters with per-region primary storage and one shared backup sink, and swaps Entra for per-request dev identities under real deny-by-default. See local-dev/README.md.

Prerequisites

  • An Azure subscription, and an identity with rights to create resource groups, Container Apps, storage accounts, Key Vaults, Azure Monitor workspaces, an Azure Container Registry, and an Azure Front Door profile in that subscription.
  • For Entra provisioning: rights to create app registrations and service principals, and a Privileged Role Administrator (or Global Administrator) role: the Bicep grants tenant admin consent for the silo's Microsoft Graph application permission declaratively, and that grant needs the privileged role.
  • Tooling on the operator workstation:
    • PowerShell 7.0 or later.
    • The Azure CLI (az), signed in (az login) to the target tenant. The deployer calls only core az command groups (account, group, deployment, acr, ad), so no CLI extension is needed.
    • Docker is not required on the operator workstation: images are built server-side with az acr build.

Quick start

$key = Read-Host -AsSecureString 'Replication key'
$gpw = Read-Host -AsSecureString 'Grafana admin password'

./deploy/Deploy-ReferenceArchitecture.ps1 `
    -SubscriptionId 00000000-0000-0000-0000-000000000000 `
    -ResourceGroup rg-lattice `
    -Location eastus `
    -BaseName lattice `
    -Regions @(
        @{ regionCode = 'use'; location = 'eastus' },
        @{ regionCode = 'euw'; location = 'westeurope' }
    ) `
    -ImageTag 2025.07.29 `
    -ReplicationTrees 'orders=LwwRegister,inventory=OrSet' `
    -ReplicationKey $key `
    -GrafanaAdminPassword $gpw

One invocation converges the whole estate and prints the resulting endpoints. Re-running it converges again; it never duplicates resources. Add -WhatIf to preview without mutating Azure: it prints each az command up to and including the pass-1 deployment, then stops, because the Entra deployment and pass 2 need pass 1's outputs.

Parameter reference

Deployment script

Deploy-ReferenceArchitecture.ps1 is the operator entry point. Its exhaustive internals (the two-pass sequence, the secret handling, the idempotency guarantees) are documented in deploy/README.md.

Parameter Required Notes
-SubscriptionId yes Target subscription.
-ResourceGroup yes Created if absent (idempotent).
-Location yes Resource-group location.
-BaseName yes 3-16 lowercase alphanumerics, shared estate-wide.
-Regions yes Array of @{ regionCode = '...'; location = '...' }, or the compact string form 'regionCode=location' (for example 'use=eastus'). One or many.
-ImageTag yes Tag applied to all three built images.
-SiloImageRepository / -McpImageRepository / -ExplorerImageRepository no Registry repository names for the three built images (defaults lattice-silo / lattice-mcp / lattice-explorer).
-DeploymentOption no public (default) or private. See below.
-ZoneRedundant no $true (default) or $false. Zone-redundant compute; applies to both options (both are VNet-injected).
-ReplicationTrees no Estate-wide treeName=MergeMode,... map.
-BackupPrimaryRegionCode no Defaults to the first region.
-IngressAllowedCidrs no Ingress allow-list seam (public option). Currently only echoed as a networking module output; no ingress applies it yet.
-SiloMinReplicas / -SiloMaxReplicas no Silo scale floor (default 1) and ceiling (default 3). The floor is never zero.
-AuthDefaultEffect no Deny (default, secure) or Allow (throwaway dev only).
-RequireApiAuthorization no Default $true.
-EnableDataApi no Default $true. Exposes the read-write Data API (write surface); set -EnableDataApi:$false to withhold it.
-EnableReplicationControl no Default $true. Co-hosts the runtime per-tree replication control plane (the sys-replication-config tree, the silo replication-control gRPC binding, and the MCP lattice_replication_* tools), fail-closed behind an explicitly authored Replication grant. -EnableReplicationControl:$false withholds it.
-EnableBackupControl no Default $true. Makes the MCP head advertise the backup tool group (read plus capture / restore / delete); the silo's backup facade is always co-hosted and every call needs an authored Backup grant. -EnableBackupControl:$false withholds the MCP group.
-EnableDigestAntiEntropy / -DigestProbeIntervalSeconds no Default $false / 0. Cross-cluster anti-entropy (digest probe, Merkle-walk drift localisation, bounded automatic repair), applied to every region; the interval optionally shortens the probe cadence (0 keeps the package default).
-ReplicationKey yes SecureString, byte-identical across every run and region. Required by both options.
-GrafanaAdminPassword yes SecureString.
-EntraEnabled / -EntraTenantId Entra Enable Entra and target the tenant.
-EntraClientId no Use a pre-existing audience app instead of deploying entra/entra.bicep.
-ExplorerWebClientId no Explorer console web-app (client) id, used only with -EntraClientId (when entra/entra.bicep is skipped); otherwise read from its explorerClientId output.
-EntraAudiences no Extra accepted token audiences.
-SecurityAdmin no The single Entra security administrator seeded as the sole initial-access principal (root of trust). An object id (GUID) or a UPN / email (resolved to its object id). Defaults to the deploying user when Entra is enabled. Further administrators are granted at runtime via the Explorer Access area.
-ExplorerRedirectUris no Defaults derived from the deployed FQDNs.
-SkipImageBuild no Reuse images already present at -ImageTag.
-WhatIf no Preview without mutating Azure. Prints each az command through the pass-1 deployment, then stops (the Entra deployment and pass 2 need pass 1's outputs).

Bicep top-level parameters

bicep/main.bicep is the all-at-once template the script drives. The parameters an operator overrides directly (when deploying the template by hand rather than through the script) are:

Parameter Default Notes
baseName (required) 3-16 lowercase alphanumerics.
regions (required) Array of { regionCode, location }.
imageTag (required) Host image tag.
siloImageRepository / mcpImageRepository / explorerImageRepository lattice-silo / lattice-mcp / lattice-explorer Registry repository names for the three built images.
registryLocation resource-group location Location of the shared registry; the script pins it to the first region to match bootstrap.bicep.
orleansServiceId baseName Orleans service id, estate-wide (each region's cluster id is <baseName>-<regionCode>).
logAnalyticsDailyQuotaGb / logAnalyticsRetentionInDays 1 / 30 Per-region Log Analytics ingestion cap and retention.
deploymentOption public public or private.
zoneRedundant true Zone-redundant compute (replicas spread across availability zones). Applies to both options - both are VNet-injected.
siloMinReplicas / siloMaxReplicas 1 / 3 Silo autoscale bounds.
backupPrimaryRegionCode first region The single backup-primary region.
replicationKey '' @secure(); the per-cluster replication key (both options - authenticates replication over public ingress, or over the private VNet mesh as defense in depth).
grafanaAdminPassword required @secure(); per-region Grafana admin password (no default; must be non-empty).
ingressAllowedCidrs [] Ingress allow-list seam (public option); echoed as a networking module output only - no ingress applies it yet.
authDefaultEffect Deny Authorization default effect estate-wide.
requireApiAuthorization true Whether the facades and MCP require authorization.
enableReplicationControl false Runtime per-tree replication control plane. The template defaults it off; the script's -EnableReplicationControl defaults on.
enableBackupControl false Whether the MCP head advertises the backup tool group. The template defaults it off; the script's -EnableBackupControl defaults on.
enableDigestAntiEntropy / digestProbeIntervalSeconds false / 0 Cross-cluster anti-entropy, applied to every region; 0 keeps the package's probe cadence.
entraEnabled / entraTenantId / entraClientId / entraAudiences off / '' Entra authentication.
explorerWebClientId / explorerAuthScope '' Explorer hosted-web OIDC: its own web-app client id and the delegated silo scope it requests on-behalf-of the operator. Threaded from the entra deployment on a later pass.
prometheusQueryEndpoint / frontDoorId '' Forward-threaded seams; empty on pass 1, activated on pass 2 (compile-cycle avoidance). Managed by the script.
explorerPublicOrigin / mcpPublicUrl / mcpAuthScope '' Explorer and MCP public Front Door URLs (OIDC redirect host, OAuth discovery resource) and the silo scope MCP clients request; Azure-assigned values threaded on a later pass.

The read-write Data API switch is not a main.bicep parameter: it is the compute module's dataApiEnabled (default true), which the script sets from -EnableDataApi on pass 2, so a hand-deployed main.bicep always exposes the Data API (every mutation is still subject-gated).

The per-region module parameters (bicep/modules/*.bicep) are internal seams the orchestrator wires; you do not set them by hand. Each module header documents its own inputs and outputs.

Deploy, verify, and teardown

Public option (default)

The public option exposes each head over ACA external ingress (server TLS, HTTP/2) fronted by a single global Azure Front Door Standard profile, and stores the per-cluster replication key in a per-region Key Vault. Deploy it with the Quick start command above (-DeploymentOption public, the default). The environment is still VNet-injected (each region gets a per-region VNet with a delegated ACA infrastructure subnet) so it is zone-redundant; it simply keeps an external ingress and no cross-region VNet peering. Each region therefore consumes a /23 infrastructure subnet from a non-overlapping per-region address plan.

Private option

The private option puts every regional ACA environment on an internal-only, VNet-integrated ingress with full-mesh global VNet peering, so cross-region replication travels private address space. Select it with -DeploymentOption private. (Both options are VNet-injected; the private option adds internal-only ingress plus the peering, on top of the per-region VNets the public option already provisions.) Replication is still authenticated by the per-cluster replication key - held in a per-region Key Vault and read via managed identity, exactly as in the public option - layered on top of the private transport as defense in depth, so -ReplicationKey is required here too.

The private option provisions no Front Door: main.bicep deploys the frontdoor module only when deploymentOption is public, because Front Door Standard has no Private Link origins. Its heads are reachable only on their internal ingress from inside the peered VNets, they run without the X-Azure-FDID origin lock (there is no Front Door id to assert), and the deployer prints only the per-region head FQDNs and skips the warm-up it gives the public option's scale-to-zero heads.

Scope of "private". This closes the ingress and the inter-region replication path, not the entire data plane. Silos still reach Azure Storage (WAL tables, backup blob) and the container registry over public PaaS endpoints (managed-identity authenticated), and the replication Key Vault keeps publicNetworkAccess enabled but firewalled to the region workload subnet. Private endpoints for Storage, ACR, and Key Vault are a documented further-hardening step this reference architecture does not yet implement.

./deploy/Deploy-ReferenceArchitecture.ps1 `
    -SubscriptionId ... -ResourceGroup rg-lattice-private `
    -Location eastus -BaseName lattice `
    -Regions @(@{ regionCode='use'; location='eastus' }, @{ regionCode='euw'; location='westeurope' }) `
    -ImageTag 2025.07.29 -DeploymentOption private `
    -ReplicationKey $key -GrafanaAdminPassword $gpw

Private-option cross-region name resolution is handled automatically by the kit. An internal ACA environment injected into a customer VNet gets no automatic private DNS zone, so bicep/modules/privatedns.bicep (deployed by main.bicep only when -DeploymentOption private) provisions one customer-managed private DNS zone per environment default domain, publishes a wildcard A record pointing every *.<defaultDomain> at that environment's static inbound IP, and links every zone to every region VNet. Each region can then resolve its peers' internal head FQDNs over the peered private network, so cross-region replication converges with no manual DNS step. The per-region VNet foundation and its full-mesh peering live in bicep/modules/vnet.bicep.

Every managed environment is zone-redundant by default (zoneRedundant, default true) under both options, because both are VNet-injected. Once the silo runs more than one replica (see Scaling behaviour) those replicas are spread across availability zones - matching the zone-redundant durability of the WAL storage tier. Set zoneRedundant to false to opt an estate back out (for example a single-zone dev estate).

Verify

After the script prints the endpoints (on the private option, use the per-region head FQDNs from inside a peered VNet in place of the Front Door hostnames):

  • Open the Explorer Front Door hostname in a browser; sign in (Entra, when enabled) and confirm the operator console loads and lists the cluster.
  • Point an MCP client at the MCP Front Door hostname and confirm the tool list is returned.
  • Write a key in one region and read it back from another to confirm active-active convergence (see the validation runbook below for the exact procedure).

Teardown

The estate is contained in a single resource group (plus its Entra app registrations). Tear it down with:

az group delete --name rg-lattice --yes
# Remove the three Entra app registrations the kit created (by display name).
# The kit names them "<BaseName> Lattice silo facade", "<BaseName> Lattice MCP
# endpoint", and "<BaseName> Lattice Explorer console" (here BaseName = lattice):
foreach ($app in 'lattice Lattice silo facade','lattice Lattice MCP endpoint','lattice Lattice Explorer console') {
    $id = az ad app list --display-name "$app" --query '[0].appId' -o tsv
    if ($id) { az ad app delete --id $id }
}

Deleting the resource group removes the container apps, storage, Key Vaults, Azure Monitor workspaces, registry, and (public option) Front Door profile. The Key Vaults have soft delete (90-day retention) and purge protection enabled, so a deleted vault keeps its name for the whole retention window and cannot be purged early. A vault's name derives from the resource group id and the region code, not from -BaseName, so to redeploy immediately use a different -ResourceGroup (or different region codes).

Day-2 operations

Scaling behaviour

  • The silo scales on the lattice.scaling compute-axis signal through a KEDA Prometheus scale rule (lattice-scaling-wal-pressure) that queries the region's managed Prometheus. The compute module's siloScaleQuery default, max(orleans_lattice_scaling_scale_value{lattice_head="silo"}), reads the series the in-environment collector stamps with lattice_head="silo"; each region's Azure Monitor workspace holds only that region's silo series, so no app-name label is needed. The default siloScaleThreshold of 0.5 is below 1 on purpose: KEDA asks for ceil(value / threshold) replicas and the scale value never exceeds the current replica count, so a threshold of 1 could only hold or shrink the pool, while 0.5 asks for twice the current count at full saturation. The same division applies at rest: the scale value never reads below its own floor (Scaling:MinReplicas, which compute sets from siloMinReplicas), so an idle region asks for twice that floor, capped at the ceiling - two replicas at the defaults. The floor is pinned at or above one replica (never zero) so the cluster always has a membership quorum; the ceiling defaults to three. A stopped replica gets the termination grace period (120 seconds) to drain: it refuses new writes with LatticeShuttingDownException and settles in-flight WAL flushes. An interrupted shard split or reshard is not handed off; it resumes from its persisted phase when its coordinator reactivates on another silo.
  • The MCP and Explorer heads scale to zero and wake on HTTP concurrency; they are stateless (MCP) or session-isolated (Explorer) admin surfaces and cost nothing while idle.

Backup and restore

  • The single backup-primary region's silo runs the backup scheduler and writes full and incremental backup chains to the shared global Azure Blob backup sink. Standby regions have restore-only (read) access to the sink.
  • A restore is fleet-wide for a replicated tree - every tree in -ReplicationTrees, plus any tree whose replication is enabled at runtime. It is promoted to an all-or-nothing coordinated restore across every current replication peer, whatever restore mode is requested: it refuses to start unless every peer is reachable, and on commit every region's tree cuts over together to the backup's point in time, replacing live data written after the backup (each region's pre-restore physical tree is retained). Only an in-place restore into a tree that is not replicated on the restoring cluster merges by per-key HLC/LWW, where a restored value never overwrites a causally newer live value. See Coordinated multi-cluster restore and the backup package's disaster recovery guide.

Failover and disaster recovery

  • Every region is a full read-write peer, so a regional outage is absorbed by the surviving regions with no promotion step: on the public option Azure Front Door latency-routes clients to the nearest healthy region and fails over to the next-nearest (the private option has no Front Door, so its clients fail over through their own private connectivity). The Explorer console is the one exception: its Blazor Server circuit must stay on one replica, so Front Door pins every operator to the first region's Explorer and fails over to a standby region (with a fresh circuit) only if that region goes down.
  • The only single-region role is the backup primary. If that region is lost, designate a new primary by re-running the deployer with a different -BackupPrimaryRegionCode; the replication key and data are unaffected.

Observability

  • Each region has a managed Prometheus (Azure Monitor workspace) and a self-hosted Grafana head pre-provisioned with the bundled Orleans.Lattice dashboards. Reach Grafana at its per-region ingress; sign in with the -GrafanaAdminPassword you supplied (Prometheus is queried through the region's managed identity, no scraped secret).

  • Metrics reach that workspace through an in-environment OpenTelemetry collector container app (one per region). A Container Apps environment cannot natively scrape a container app into an Azure Monitor workspace, so the collector scrapes the silo /metrics endpoint over the environment's internal network and remote-writes to the region's data collection endpoint. A co-located aad-auth-proxy sidecar mints the managed-identity token (the region identity holds Monitoring Metrics Publisher on the data collection rule) so the write carries no static secret. The KEDA scaler and the MCP telemetry tools then read the same workspace back.

Connect an MCP client

The MCP head exposes the Lattice control surface (state, auth-admin, tree-administration, and telemetry tool groups, plus data, backup and replication while those surfaces are enabled - the deployer's default) as a Model Context Protocol server over streamable HTTP. It runs stateless behind Front Door, is authenticated with a Microsoft Entra bearer token that a spec-compliant client acquires automatically via OAuth discovery (below), and is origin-locked: Front Door injects the X-Azure-FDID header on the client's behalf, so a client that reaches the head through the Front Door hostname supplies only an Authorization header. (A client that bypasses Front Door and dials a region's container-app FQDN directly must add the matching X-Azure-FDID header itself.) This describes the public option. The private option has no Front Door, so its heads carry no origin lock, and because the deployer threads Mcp:PublicUrl from the Front Door MCP hostname, its heads serve no OAuth discovery document either: use the bearer-token fallback below against a region's internal MCP FQDN, from inside a peered VNet.

Preferred: automatic OAuth discovery (RFC 9728)

When Entra is enabled the head advertises OAuth 2.0 Protected Resource Metadata (RFC 9728), so a spec-compliant MCP client acquires its own token and nothing is pasted by hand. On the first unauthenticated request the head returns 401 with a WWW-Authenticate challenge whose resource_metadata parameter points at the anonymous metadata document served at /.well-known/oauth-protected-resource. The client fetches that document, reads the Entra authorization_servers issuer and the silo scopes_supported scope from it, runs the sign-in flow itself, and retries with the token it obtained. Point the client at the head URL and let it discover the rest - no Authorization header:

{
  "mcpServers": {
    "lattice-ra": {
      "type": "http",
      "url": "https://<mcp-front-door-hostname>/",
      "tools": ["*"]
    }
  }
}

bicep/entra/entra.bicep pre-authorizes the Visual Studio Code, Visual Studio, and Azure CLI first-party clients for the silo scope by default (its preAuthorizedMcpClientIds parameter), so those clients sign in with no client id to supply and no consent prompt; add another client's id to that parameter to pre-authorize it too (for example a GitHub Copilot app id captured from a real sign-in). A client that prompts for a client id can use the Visual Studio Code id aebc6443-996d-45c2-90f0-388ff96faa56. A signed-in caller sees no tool groups - only the lattice_capabilities meta-tool - until the security administrator grants their Entra object id (oid) access in the Explorer console's Access area; discovery then advertises only the tool groups they hold, and every forwarded call is re-authorized at the silo.

Fallback: supply a bearer token by hand

A client that does not support OAuth discovery - and the raw JSON-RPC smoke-test below - authenticates instead by minting an Entra token and sending it as a static Authorization header.

1. Mint an access token for the silo facade. The MCP tools call through to the region silo, so the token's audience is the silo facade app, not the MCP head. For an interactive operator, the Azure CLI mints one against the silo App ID URI:

$token = az account get-access-token `
  --resource "api://<tenantId>/<BaseName>-silo" `
  --query accessToken -o tsv

For unattended automation, register a service principal, assign it the silo app role, and use the client-credentials grant instead. Either way the silo resolves the caller's subject from the token's stable oid claim and enforces the deny-by-default per-tree access model against it, so the principal must be granted the rules (or bootstrap-administrator status) for the trees and groups it will use.

2. Point an MCP client at the Front Door MCP hostname. For GitHub Copilot CLI, add an http server to ~/.copilot/mcp-config.json:

{
  "mcpServers": {
    "lattice-ra": {
      "type": "http",
      "url": "https://<mcp-front-door-hostname>/",
      "headers": { "Authorization": "Bearer <token>" },
      "tools": ["*"]
    }
  }
}

The same two inputs (the Front Door URL and the Authorization: Bearer <token> header) drive any MCP client that speaks streamable HTTP.

3. Smoke-test the endpoint. A raw JSON-RPC initialize + tools/list confirms discovery without a full client:

$url = "https://<mcp-front-door-hostname>/"
$h = @{
  Authorization  = "Bearer $token"
  "Content-Type" = "application/json"
  Accept         = "application/json, text/event-stream"
}
$init = '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"probe","version":"1"}}}'
Invoke-WebRequest -Method Post -Uri $url -Headers $h -Body $init -UseBasicParsing | Out-Null
$list = '{"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}}'
(Invoke-WebRequest -Method Post -Uri $url -Headers $h -Body $list -UseBasicParsing).Content

Notes and gotchas:

  • Tool arguments use treeId. The data and state tools name their tree parameter treeId (not treeName). The read-range tool's (lattice_data_read_range) bounds (startInclusive, endExclusive), pageSize, and continuationToken are optional parameters, left out of the schema's required list, so a hand-written call may omit them.
  • Token lifetime. An Entra access token expires in about an hour. Re-mint and refresh the Authorization header before it lapses (automation should acquire a fresh token per session).
  • Per-request routing and explicit targeting. Front Door routes each request independently, with no session affinity: to the lowest-latency healthy region, spreading requests across every region within 50 ms of it, so consecutive calls can land in different regions. The config-plane grants are symmetric across regions (the auth tree is replicated), but telemetry is per region (each head queries its own region's managed Prometheus), and a data-plane tree written in one region is not visible to a read routed to another until replication converges it. The MCP head is wired for cross-region targeting, so you do not need to bypass Front Door for a deterministic single-region check: call lattice_list_regions to enumerate the estate's regions (each region's id, cluster id, and per-facade reachability), then pass an optional region argument on any tool call to pin it to that region. The served region is echoed back in the result's region metadata. A call with no region targets the head's own (current) region. See Target a specific region below.

Target a specific region

Every region's MCP head fronts its co-located silo and holds a direct, region-pinned route to every peer region's silo (the same FQDN cross-region replication uses - never the anycast Front Door hostname). This lets one Front Door MCP endpoint serve the whole estate while still letting a caller address a single region deterministically:

  • lattice_list_regions returns every routable region with its regionId, clusterId, isCurrent flag, and per-facade-group reachability. A region that cannot be reached (or, with identity verification on, whose endpoint does not actually reach its advertised cluster) is omitted, so the list is fail-closed.
  • Any tool call accepts an optional region argument set to a regionId from that list. The call is routed to that region's silo, re-authorized there against the same forwarded token (fail-closed - a region is never an authorization bypass), and the result's region metadata reports which region served it.

Because the estate is fronted by one global Front Door, the head enables region-identity verification: before a peer is routed to, its state facade is probed once and its reported cluster id compared to the advertised one, so a region accidentally pointed at an anycast endpoint that latency-routes to the wrong cluster is rejected rather than silently answered by the wrong region. A region value that is unknown, unreachable, or fails verification returns a typed MCP error, never a call served by the wrong cluster.

For example, write a key through one region and read it back from another to confirm convergence, both through the same Front Door endpoint:

# List regions, then target each explicitly by its regionId on the tool call.
$list = '{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"lattice_list_regions","arguments":{}}}'
(Invoke-WebRequest -Method Post -Uri $url -Headers $h -Body $list -UseBasicParsing).Content

# Read a key pinned to a specific region (regionId from the list above).
$get = '{"jsonrpc":"2.0","id":4,"method":"tools/call","params":{"name":"lattice_data_get","arguments":{"treeId":"orders","key":"o-1","region":"euw"}}}'
(Invoke-WebRequest -Method Post -Uri $url -Headers $h -Body $get -UseBasicParsing).Content

Optional hardening and upgrade path

The baseline ships secure but with the Front Door Web Application Firewall (WAF) off and on the Standard SKU. Two opt-in steps harden it further.

Enable a Front Door WAF custom-rule policy (Standard)

The Standard profile supports custom WAF rules (rate limiting, geo-filtering, IP allow/deny). Provision a Microsoft.Network/FrontDoorWebApplicationFirewallPolicies policy and attach it to the profile with a securityPolicies association over the endpoint domains. The frontdoor module header carries the exact snippet and the enableWaf seam. Custom rules add no SKU cost but are billed per policy and per rule evaluation.

Upgrade Azure Front Door Standard to Premium

Premium adds Azure-managed WAF rule sets (OWASP core + bot protection) and Private Link private origins. Upgrading:

  • Changes the profile SKU from Standard_AzureFrontDoor to Premium_AzureFrontDoor and lets you attach the managed rule sets in addition to (or instead of) custom rules.
  • Enables Private Link origins, so the heads can be reached privately rather than over public ingress. This is what would let the private network option, which provisions no Front Door today, gain one: with Premium + Private Link the client-facing origins never need public ingress at all. Note that the private option's internal ingress already keeps replication traffic off the public internet; Premium extends that to the client-facing path.
  • Carries a higher base monthly cost than Standard plus managed-rule request charges. Weigh it against the estate's exposure and compliance requirements.

Close the remaining public data-plane surfaces (private endpoints)

Even under the private network option, "private" today means private ingress and a private inter-region replication path - not a fully private data plane. Two public PaaS surfaces remain, both authenticated by managed identity:

  • The silos reach Azure Storage (WAL tables, backup blob) and the container registry over their public service endpoints.
  • The replication Key Vault keeps publicNetworkAccess enabled (firewalled to the region workload subnet via a service endpoint), rather than publicNetworkAccess: Disabled behind a private endpoint.

To reach a zero-public-surface posture, add private endpoints for Storage, the registry, and Key Vault (with publicNetworkAccess: Disabled and private DNS zone links per region), and switch the Key Vault firewall from a service-endpoint virtualNetworkRule to a private endpoint. This is a deliberate, separate hardening effort and is not implemented in the baseline.

Local development

To exercise the whole estate on one machine before touching Azure, use the Docker Compose harness under local/. It runs the three heads against Azurite and a local Prometheus/Grafana, with the security bypass toggles (Entra off, plaintext h2c) documented and defaulted for development only. See local/README.md.

For a two-region topology - to exercise cross-cluster replication and differentiated per-identity authorization on one machine - use the local-dev/ harness instead. It mirrors local/ but differs in four deliberate ways: every head builds directly from src/** by project reference, so the stack always reflects your working tree with no pack or publish step (third-party dependencies still restore from NuGet); it stands up two network-isolated regions bridged only by two silo-only seams, one for replication and one to a single shared backup sink; each region gets its own isolated Azurite primary storage, while both share that one backup sink (a coordinated restore of a replicated tree needs every cluster to read the same backup); and it replaces Entra with hand-crafted per-request dev identities enforced under real deny-by-default, so an agent can act as any of several differentiated identities by setting a bearer token. See local-dev/README.md.

Real-Azure validation runbook

Status: recorded. The kit has been deployed to and operated against live Azure two-region estates during development. The public-network, Entra-enabled path (Front Door reachability, Explorer OIDC sign-in, MCP tool discovery, and the backup and replication wiring) was validated on the public reference estates; a dedicated two-region private estate (internal ingress, full-mesh VNet peering, customer-managed private DNS) was then stood up to validate the private-network deployment option and the availability (health-probe) reliability work. The evidence recorded below is drawn from those runs and labelled by the topology it was captured on.

Canonical validation topology: two regions (for example eastus + westeurope), public network option, Entra enabled. A private-network variant (internal ingress, VNet peering, customer-managed private DNS) was also validated for the private deployment option.

  1. Deploy. Run the Quick start command for two regions. Record the printed Front Door and per-region head endpoints.
    • Evidence: recorded. A two-region private estate converged with both regions on internal ingress (eastus2 and westus2 managed environments, each with its own default domain and static inbound IP, all three heads Running); the public estates recorded the Front Door and per-region head endpoints during development.
  2. Reachability. Open the Explorer Front Door hostname and confirm sign-in and the cluster view; call the MCP Front Door hostname and confirm the tool list.
    • Evidence: recorded. From an in-VNet client in eastus2, the westus2 environment internal silo and MCP FQDNs resolved over the customer-managed private DNS zones and returned HTTP 200 on /health cross-region; Explorer Front Door OIDC sign-in and MCP tool discovery were validated on the public Entra estates during development.
  3. Active-active convergence. Write a key through region A's State API and read it back through region B's State API; then write the same key concurrently in both regions and confirm the CRDT merge result is identical on both sides.
    • Evidence: recorded (transport and control plane). Cross-region replication is wired end to end: peers are addressed by internal FQDN and, on the private option, those FQDNs now resolve across regions via the private DNS zones, so the peered network path is usable. The peer Push path was observed returning HTTP 200 with entries shipping, and the config and policy replication plane was observed converging between regions (a grant authored on one cluster appeared on the peer within seconds). The prior data-convergence blocker (a non-serializable leaf-projection exception, issue #1336) is fixed and released in the core library this estate references. A formal concurrent-write / merged-read capture was not taken on the private test estate.
  4. Autoscale. Drive load at one region's silo and confirm the KEDA scaler raises the replica count above the floor, then scales back down after the load stops.
    • Evidence: recorded. The silo carries a KEDA Prometheus scale rule (named lattice-scaling-wal-pressure, over the lattice.scaling scale value; minReplicas 1, maxReplicas 3; the template sets no polling interval or cooldown, so the Container Apps defaults of 30 s and 300 s apply), verified present on the live estate; the MCP and Explorer heads scale to zero and back (minReplicas 0, maxReplicas 3) on an HTTP concurrency rule. Availability under change was validated directly: a forced single-revision silo cutover served continuous HTTP 200 with zero dropped requests across the roll, as the readiness and startup probes gate traffic to warm replicas only. A sustained load-driven scale-out / scale-in timeline was not captured.
  5. Backup and restore. Confirm the primary region writes a backup chain to the sink, then perform a restore into a standby and confirm the restored value is present and causally consistent.
    • Evidence: recorded (wiring). The primary region is flagged for backup and a dedicated backup blob account is provisioned and wired to the silo, with the backup gRPC facade co-hosted and the MCP backup tool group gated fail-closed. An end-to-end backup-chain write and restore-into-standby was not driven on the private test estate.
  6. Teardown. Run the teardown commands and confirm the resource group and the three Entra apps are removed.
    • Evidence: recorded. The private test estate resource group was deleted after validation.

Cost note

The validated two-region public topology's steady-state cost is dominated by the always-on silo replicas (two per region at rest with the default floor and scale threshold, up to the ceiling while compute pressure scales a region out; see Scaling behaviour), the two managed Prometheus workspaces and Grafana heads, the two Standard storage accounts plus the shared backup blob account, the two Key Vaults, the shared container registry, and the single Front Door Standard profile. The scale-to-zero MCP and Explorer heads add negligible idle cost. Record the actual monthly figure from Azure Cost Management after the validation run. Enabling Front Door Premium is the single largest cost lever (see the upgrade path above).