KEDA on Azure Container Apps
This page documents Orleans.Lattice.Scaling 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at keda-aca.md, and llms.txt lists every page.An end-to-end walkthrough for autoscaling an Orleans.Lattice cluster on Azure
Container Apps (ACA) using the scaling signal as a KEDA custom metric. The
ClusterScaling sample is a complete,
deployable implementation of everything below.
The shape
ACA scales container-app replicas with KEDA under the hood. The
metrics-api KEDA scaler polls an HTTP endpoint, reads a numeric value out of the
JSON response, and drives the replica count toward a target. The scaling
endpoint is exactly that: a GET endpoint returning a JSON body with a top-level
scaleValue.
KEDA metrics-api scaler Orleans.Lattice silo (each replica)
---------------------- -----------------------------------
GET https://<app>/lattice/scale ---> MapLatticeScalingSignal()
reads $.scaleValue <--- { "scaleValue": 3.4, ... }
targetValue: 0.5
desiredReplicas = ceil(scaleValue / targetValue)
Because the activation and resource dimensions are cluster aggregates, KEDA can poll any replica and read a whole-cluster demand. The WAL-dispatch dimension and the smoothing state are the answering replica's own, so two replicas can answer slightly differently - see cluster-aggregate answering.
Host wiring
Register the signal on the silo, map the endpoint, and listen on the ACA ingress target port:
using Orleans.Lattice.Scaling;
siloBuilder.AddLatticeScalingSignal(options =>
{
// Never recommend fewer than the minimum viable cluster.
options.MinReplicas = 2;
// Refresh well inside KEDA's poll interval.
options.SampleInterval = TimeSpan.FromSeconds(5);
});
using Orleans.Lattice.Scaling;
var app = WebApplication.Create();
app.MapLatticeScalingSignal(); // GET /lattice/scale
The custom scale rule
Add a custom scale rule of type metrics-api to the container app. In bicep:
// Built from the app name and the environment's default domain: referencing the
// container app's own ingress FQDN inside its own declaration is a circular
// self-reference, which Bicep rejects.
var scaleSignalUrl = 'https://${appName}.${environment.properties.defaultDomain}/lattice/scale'
// Inside the Microsoft.App/containerApps resource's properties.template:
scale: {
minReplicas: 2
maxReplicas: 20
rules: [
{
name: 'lattice-scale'
custom: {
type: 'metrics-api'
metadata: {
url: scaleSignalUrl
valueLocation: 'scaleValue'
targetValue: '0.5'
activationTargetValue: '1'
}
}
}
]
}
valueLocation: 'scaleValue'- the JSON path KEDA reads. The endpoint emits it as a stable, camelCase top-level property.targetValue: '0.5'- the per-replica pressure the rule holds the pool at.scaleValueis the dominant compute pressure (0.0to1.0) times the current replica count, so it never exceeds that count except at theMinReplicasfloor or while the scale-in gate holds - or the smoothing releases - an earlier, higher value. An ACA custom scale rule exposes no metric type, so KEDA's defaultAverageValueapplies and it computesdesiredReplicas = ceil(scaleValue / targetValue): the pool grows when the dominant pressure rises abovetargetValue, and shrinks - once the scale-in gate lets the value fall - when fewer replicas would still carry the same load at or belowtargetValue. AtargetValueof1therefore never adds a replica - at full saturation it asks for exactly the current count - so scale-out needs a value below1.0.5is the value theClusterScalingsample deploys: a saturated pool asks for twice its current size.activationTargetValue: '1'- the threshold above which KEDA activates the app from zero (if you allow scale-to-zero). Keep it aligned with yourMinReplicas.minReplicas/maxReplicas- the ACA replica envelope. SetminReplicasto your quorum floor andmaxReplicasto your capacity ceiling. The rule divides the signal'sMinReplicasfloor bytargetValuelike any otherscaleValue, so once the answering replica has taken its first sample it never asks for fewer thanceil(MinReplicas / targetValue)replicas. With the host wiring above (MinReplicas = 2) an idle pool'sscaleValuesits at the floor, 2, and the rule asks forceil(2 / 0.5)= 4 replicas, so the pool rests at four (capped atmaxReplicas) rather than atminReplicas: 2. For the pool to rest atminReplicas, keepMinReplicasat or belowminReplicastimestargetValue- at most 1 here; its default of0leaves the floor entirely tominReplicas.
Polling and stabilization versus EWMA
KEDA polls on its own interval (pollingInterval, default 30s) and ACA applies
its own cooldown before scaling in. These stack on top of the signal's own
smoothing:
- Scale-out is fast on both sides: the signal snaps up immediately, and KEDA scales out on its next poll.
- Scale-in is deliberately slow: the signal only lets the scalar fall after
every scale-in precondition has held for
ScaleInGateWindow(default 2m), and KEDA/ACA then apply their own cooldown on top. TuneEwmaHalfLifeandScaleInGateWindowfor how conservative you want scale-in, and leave the KEDA/ACA cooldown to guard against poll-to-poll flapping.
Keep SampleInterval (signal freshness) well below pollingInterval (KEDA read
cadence) so KEDA never reads a stale sample.
Health probes
The scaling health check reports the cached snapshot's verdict, and its
activation and resource inputs are the cluster's worst-silo values, so every
replica reports the same verdict for them (only the WAL inputs are the replica's
own). A readiness probe on it therefore takes every replica out of rotation
together when the hottest silo crosses the Unhealthy bound - see
readiness on AKS. Register it with a tag so you can choose
which probe endpoint includes it:
using Microsoft.Extensions.DependencyInjection;
using Orleans.Lattice.Scaling;
var services = new ServiceCollection();
services.AddHealthChecks().AddLatticeScalingHealthCheck(tags: new[] { "ready" });
Map an endpoint filtered to the ready tag and set it as the ACA readiness probe
path only if that cluster-wide drain is what you want; otherwise expose it on a
separate endpoint for alerting.
See also
- AKS for the Kubernetes
ScaledObjectand HPA alternatives. - Configuration for every knob the walkthrough tunes.
ClusterScalingsample for the full deployable implementation.