ClusterScaling - autoscaling on Azure Container Apps
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.A deployable, end-to-end sample that runs a real multi-silo Orleans.Lattice
cluster on Azure Container Apps (ACA) and proves the
Orleans.Lattice.Scaling autoscaling signal drives KEDA replica scale-out on
the compute axis. A bundled .NET load driver generates compute-axis pressure,
and the deploy tooling wires an ACA KEDA metrics-api scale rule to the
/lattice/scale signal so ACA adds replicas under load and removes them
afterwards.
This is the "does it work in anger" capstone for the autoscaling-signal work
(issue #1216, epic #1190). It consumes the shipped scaling surface
(MapLatticeScalingSignal, AddLatticeScalingSignal, and the health check from
Orleans.Lattice.Scaling); it does not redefine them. Its scale rule is written
inline in main.bicep in the shape of the package's reference ACA scale rule,
including its targetValue: 0.5 (see "Why targetValue: 0.5 and not 1?"
under Deploy).
The two axes (read this first)
Orleans.Lattice's scaling signal reports two axes:
- Compute axis - grain-activation, host-resource (CPU and memory) and
WAL-dispatch pressure. This is the only axis wired to replica count.
When compute pressure rises,
scaleValueclimbs and the autoscaler adds replicas. - Storage axis - retained WAL bytes. This is advisory: it feeds observability and the health check, and it never inflates replica count. Relieving storage pressure is an operational action (rebalancing WAL partitions), not an autoscaling one.
This sample drives the compute axis. The load driver issues a high op rate
across many distinct trees and keys with a small fixed payload (256 bytes by
default), so it grows activation and dispatch pressure while keeping retained-byte
growth small. Bulk-loading large values at a low op rate would move mostly the
storage axis, which never feeds scaleValue, so KEDA would not scale - which is
why the driver keeps payloads deliberately small.
Architecture
flowchart TB
subgraph WS["Your workstation"]
DEPLOY["deploy.ps1<br/>(provision + build + push)"]
DRIVE["drive-load.ps1<br/>LoadDriver (compute load)"]
POLL["az poll<br/>(replica count)"]
end
subgraph RG["Azure resource group"]
ACR["Azure Container Registry<br/>(Basic)"]
STORE[("Azure Storage - Tables<br/>clustering, reminders, grain state, WAL")]
subgraph ACA["Container Apps environment"]
KEDA["KEDA metrics-api rule"]
subgraph APP["Container app (1..N replicas)"]
SILO["Orleans silo<br/>(Azure clustering)"]
API["data API gRPC<br/>(Basic-gated)"]
SCALE["/lattice/scale<br/>(scrape target)"]
end
end
end
DEPLOY -->|"1. az acr build: build + push image"| ACR
DEPLOY -->|"2. deploy bicep"| ACA
APP -->|"pull image (managed identity, AcrPull)"| ACR
SILO -->|"clustering, state, WAL (managed identity)"| STORE
DRIVE -->|"gRPC + Basic over managed TLS"| API
POLL -->|"reads replica count"| APP
KEDA -->|"reads scaleValue"| SCALE
KEDA -->|"sets replica count"| APP
- Silo host (
src/ClusterScaling.Silo) - one container image, run as many ACA replicas. Each replica joins a genuine Orleans cluster over Azure Storage clustering (managed identity), persists grain state and the Lattice WAL to Azure Table storage (managed identity), and co-hosts:- the write-capable data API gRPC surface, gated by a Basic admin credential whose salted PBKDF2 hash arrives as an ACA secret; and
- the
/lattice/scaleHTTP signal endpoint the KEDA scale rule scrapes, plus/healthzand/readyzhealth endpoints.
- Load driver (
src/ClusterScaling.LoadDriver) - a small .NET console that speaks gRPC to the data API over TLS, presents the admin Basic credential, and drives sustained compute-axis load, printing offered-load throughput. - Deploy tooling (
deploy/) -main.bicepandregistry.bicepplusdeploy.ps1,drive-load.ps1, andteardown.ps1.
Credential and TLS posture
- The operator supplies a plaintext admin password to
deploy.ps1(as aSecureString). The script hashes it with the repository'stools/New-LatticeStateCredential.ps1helper (salted PBKDF2-SHA256) and passes only the hash to the bicep template. - The bicep injects the hash as a container-app secret, surfaced through the
LATTICE_DATA_USER_<admin>environment variable the data-API authorizer reads.deploy.ps1never stores the plaintext, never bakes it into the image, and never passes it on a command line (the hashing helper reads it from an inherited environment variable).drive-load.ps1does hand it to the local LoadDriver as a--passwordargument, because the driver presents it as the Basic credential. - The data-API
BasicAdminDataApiAuthorizerverifies the inboundauthorization: Basic base64(user:pass)header against that hash in constant time. An anonymous or wrong-password call is rejected withPermissionDenied. - Basic-over-cleartext would be unsafe on its own. It is legitimate here because
ACA terminates TLS at its managed ingress: the credential rides an
encrypted HTTP/2 channel from the driver to the ingress, and the container is
reachable only through that ingress. This is the upgrade over the localhost
PasswordProtectionsample, which has no transport encryption.
Prerequisites
- An Azure subscription and
az login, on an account that can create role assignments (the deploy assigns Storage Table Data Contributor and AcrPull to the app's managed identity). - Azure CLI with the
containerappextension (deploy.ps1installs/updates it). - The .NET SDK (net10.0) to run the load driver.
- PowerShell 7.2 or later:
deploy.ps1hashes the admin password withtools/New-LatticeStateCredential.ps1, which requires 7.2.
No container registry, image, or local Docker daemon is required up front:
deploy.ps1 provisions a Basic Azure Container Registry as part of the
deployment and builds the silo image into it for you (see below).
The silo image
You don't build the image, provision a registry, or run Docker by hand -
deploy.ps1 orchestrates all of it. By default it:
- provisions a Basic Azure Container Registry (
registry.bicep) into the resource group; - stages a clean copy of the build inputs (
Directory.Build.targets,src/andsamples/ClusterScaling/src/, withoutbin,objor.vs) in a temporary folder and runsaz acr buildthere againstsrc/ClusterScaling.Silo/Dockerfile, which builds and pushes the image server-side in that registry (no local Docker daemon) and streams the build log to your console; and - wires the container app to pull the image using the managed identity
(
AcrPull), so no registry username or key is ever used.
Two overrides are available:
-Registry <name>- build+push into an existing ACR you already own instead of provisioning a new one. The managed-identityAcrPullpull path is still wired for you.-ContainerImage <ref>- deploy a pre-built external image and skip both the registry provisioning and the build. NoAcrPullis wired; you own that image's pull access. Build one yourself with:az acr build --registry <myregistry> --image clusterscaling-silo:latest ` --file samples/ClusterScaling/src/ClusterScaling.Silo/Dockerfile .
Deploy
cd samples/ClusterScaling/deploy
$pw = Read-Host -AsSecureString -Prompt 'Admin password'
./deploy.ps1 `
-ResourceGroup rg-clusterscaling `
-Location eastus `
-AdminPassword $pw `
-MinReplicas 1 -MaxReplicas 10
deploy.ps1 is idempotent. It provisions the Basic container registry, builds
and pushes the silo image into it (see above), then provisions the managed
identity, the Tables-only storage account (shared-key access disabled), the
Storage Table Data Contributor and AcrPull role assignments, the Log Analytics
workspace, the Container Apps environment, and the container app with the KEDA
metrics-api scale rule (valueLocation: scaleValue, targetValue: 0.5,
minReplicas/maxReplicas). It prints the ingress FQDN and the exact
drive-load.ps1 command to run next.
Why
targetValue: 0.5and not1? KEDA computesdesiredReplicas = ceil(scaleValue / targetValue).scaleValueis the dominant per-replica utilisation (0..1) times the current replica count, so a single replica caps it at1.0. WithtargetValue: 1that yieldsceil(1.0 / 1) = 1even at full saturation, so the cluster can never bootstrap its first scale-out.targetValue: 0.5targets 50% per-replica utilisation and leaves headroom: a saturated single replica (scaleValuenear1.0) givesceil(1.0 / 0.5) = 2, while an idle one (scaleValuenear0.1) stays atceil(0.1 / 0.5) = 1. Lower the value for a more aggressive (earlier) scale-out; raise it toward1to require heavier saturation first.
Drive load and observe scale-out
$pw = Read-Host -AsSecureString -Prompt 'Admin password'
./drive-load.ps1 `
-ResourceGroup rg-clusterscaling `
-AdminPassword $pw `
-Rate 2000 -Duration 300
drive-load.ps1 resolves the ingress FQDN, waits for the app's /healthz to
answer 200, launches the bundled LoadDriver (compute-axis load), and - while it
runs - polls az containerapp replica list
to print a replica-count timeline interleaved with the driver's continuous
offered-load throughput. Example shape:
Replica-count timeline (offered-load lines come from the driver):
[t= 10s] replicas = 1
t= 10.0s offered= 20,000 offered/s= 2,000 completed= 19,880 ...
[t= 40s] replicas = 3
[t= 70s] replicas = 6
...
Timing expectations. Scale-out lags the load by tens of seconds. Three delays stack between offered load and a new replica:
- the signal's sample interval (
SampleInterval, 5 seconds by default) - the scalar itself snaps up immediately, with no smoothing on the way up; - the KEDA polling interval (ACA default 30s); and
- the time a new replica takes to start and join the cluster.
Scale-in is deliberately slower (it prevents replica thrashing): the signal
holds a falling scalar until every scale-in precondition has held for
ScaleInGateWindow (2 minutes by default) and then lets it decay through the
EWMA (EwmaHalfLife, 30 seconds by default), and ACA applies its own scale-in
cooldown on top. Sustain the load for a while - the 5 minute default is
comfortable - then watch the count settle back toward minReplicas after the
driver stops:
az containerapp replica list -g rg-clusterscaling -n <app> --query 'length(@)' -o tsv
The minReplicas floor keeps the scrape target reachable; the maxReplicas
ceiling is the hard cap the autoscaler never exceeds regardless of how high
scaleValue climbs.
Verify the credential gate
An anonymous or wrong-password call to the data API is rejected. The load driver
fails fast with a clear message if -AdminPassword does not match what
deploy.ps1 hashed into the secret, so a mismatch surfaces immediately rather
than as silent zero throughput.
Teardown
./teardown.ps1 -ResourceGroup rg-clusterscaling
Deletes the whole resource group, after a confirmation prompt (-Yes skips it;
-NoWait returns without waiting for the deletion to finish). An idle
deployment is not free even at
minReplicas=1: the always-on replica bills vCPU + memory per second, Log
Analytics bills for ingested logs, and the storage account bills for the tables
it retains. Tear down as soon as an experiment finishes.
When to use / when not to use
Use this sample when you want to:
- See the compute-axis scaling signal drive real horizontal scale-out on live infrastructure, end to end.
- Copy a correct managed-identity ACA wiring for a multi-silo Lattice cluster
(Azure clustering + reminders + grain state + WAL, no keys or connection
strings) and a KEDA
metrics-apiscale rule against/lattice/scale. - Understand the Basic-over-managed-TLS credential posture for the write-capable data API and the ACA-secret hash injection.
Do not use this sample when:
- You want to scale on the storage axis. It is advisory and never wired to replica count; relieving WAL pressure is a rebalancing action, not an autoscaling one.
- You need a local, dependency-free demo. Start with
HelloWorldor, for the credential mechanism alone,PasswordProtection(single in-process silo, no Azure). - You want a production deployment blueprint verbatim. This is a teaching sample: it uses external ingress so you can drive load from your workstation, a single storage account, and permissive storage network ACLs. A production deployment would scope ingress (IP allow-list or internal), separate the WAL account from clustering, and lock down the storage firewall.
Layout
samples/ClusterScaling/
README.md
src/
ClusterScaling.Silo/ # multi-silo host: clustering + WAL + data API gRPC + scale signal
Program.cs
BasicAdminDataApiAuthorizer.cs
ClusterScaling.Silo.csproj
Dockerfile # the silo image deploy.ps1 builds with az acr build
ClusterScaling.LoadDriver/ # compute-axis gRPC load generator
Program.cs
LoadDriverOptions.cs
ClusterScaling.LoadDriver.csproj
deploy/
main.bicep # identity, storage, role, ACA env + app, KEDA scale rule
registry.bicep # the Basic container registry deploy.ps1 provisions by default
deploy.ps1 # hash password, provision, build + push image, print FQDN (idempotent)
drive-load.ps1 # run LoadDriver + poll replica timeline
teardown.ps1 # delete the resource group