KEDA and HPA on AKS
This page documents Orleans.Lattice.Scaling 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at aks.md, and llms.txt lists every page.Autoscaling an Orleans.Lattice cluster on Azure Kubernetes Service (AKS), or any
Kubernetes cluster, using the scaling signal. Two options: a KEDA ScaledObject
(recommended) or a native Horizontal Pod Autoscaler (HPA) against a custom metric.
The host wiring is identical to the ACA walkthrough:
AddLatticeScalingSignal on the silo, MapLatticeScalingSignal on the web host,
and the endpoint served on the pod's container port.
Option 1: KEDA ScaledObject (recommended)
Install KEDA in the cluster, then apply a ScaledObject with a metrics-api
trigger pointed at the in-cluster service:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: lattice-scaledobject
spec:
scaleTargetRef:
name: lattice-silo # the Deployment to scale
minReplicaCount: 2
maxReplicaCount: 20
pollingInterval: 15 # seconds; keep above SampleInterval
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleDown:
stabilizationWindowSeconds: 120 # scale-in pacing on top of the gate
triggers:
- type: metrics-api
metadata:
url: "http://lattice-silo.default.svc.cluster.local/lattice/scale"
valueLocation: "scaleValue"
targetValue: "0.5"
urltargets the headless or ClusterIP service in front of the silo pods; KEDA polls one pod and reads its snapshot, whose activation and resource dimensions are cluster aggregates (the WAL-dispatch dimension and the smoothing state are that pod's own - see cluster-aggregate answering).valueLocation: "scaleValue"andtargetValue: "0.5"behave exactly as in the ACA rule: with the trigger's defaultAverageValuemetric type KEDA asks forceil(scaleValue / targetValue)pods, and becausescaleValuenever exceeds the current pod count (apart from theMinReplicasfloor and an earlier, higher value the scale-in gate holds or the smoothing is still releasing), atargetValueof1could never add a pod. The floor is divided too: with the ACA host wiring (MinReplicas = 2) an idle cluster asks forceil(2 / 0.5)= 4 pods and rests at four rather than atminReplicaCount: 2- see theminReplicasnote under the ACA rule for choosing the two floors.pollingIntervalshould stay aboveLatticeScalingSignalOptions.SampleIntervalso KEDA never reads a stale sample. Scale-in betweenminReplicaCountandmaxReplicaCountis paced by the managed HPA's scale-down stabilization window (advanced.horizontalPodAutoscalerConfig.behavior.scaleDown.stabilizationWindowSeconds, 300 seconds when unset), which stacks on the signal's ownScaleInGateWindow. KEDA'scooldownPeriodapplies only to scaling to zero, so it has no effect with aminReplicaCountof 2.
KEDA creates and manages the underlying HPA for you.
Option 2: HPA against a custom metric
If you prefer a native HPA, expose the scale value as an external metric through
the Prometheus adapter (scrape the orleans.lattice.scaling meter
via the OpenTelemetry Prometheus exporter, so scaleValue is available as
orleans_lattice_scaling_scale_value), then:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: lattice-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: lattice-silo
minReplicas: 2
maxReplicas: 20
metrics:
- type: External
external:
metric:
name: orleans_lattice_scaling_scale_value
target:
type: AverageValue
averageValue: "0.5"
behavior:
scaleDown:
stabilizationWindowSeconds: 120
Use an External metric, not a Pods one. The value is a cluster-wide demand
in replica-units that every silo exports, so an External metric with an
AverageValue target of 0.5 makes the HPA divide it by the current pod count
and settle on ceil(scaleValue / 0.5) replicas - the same arithmetic as the KEDA
rule, including its inability to add a pod at a target of 1 and its rest at
twice the signal's MinReplicas floor (the exported gauge carries the floored
value).
A Pods metric would instead average the near-identical per-pod values and
multiply by the current pod count, overshooting by roughly that factor. For the
same reason, have the adapter's external-metric query aggregate the per-pod
series with max or avg rather than sum.
The KEDA route is preferred because the metrics-api trigger reads the endpoint
directly and needs no Prometheus-adapter plumbing; the HPA route is useful when
you already run the Prometheus adapter and want a single autoscaling mechanism.
Readiness
The scaling health check
can back a probe, but it is not a per-pod signal: it reads the same cached
snapshot the endpoint serves, so its activation and resource inputs are the
cluster's worst-silo values and every pod reports the same verdict for them
(only the WAL inputs are the pod's own). Wired into a readinessProbe, it
therefore takes every pod out of rotation together once the cluster's hottest
silo crosses the Unhealthy bound - a Degraded result still answers 200 on
the default ASP.NET Core status mapping. If that is the behaviour you want, wire
it like this; otherwise map it on its own endpoint (or keep it out of the
readiness tag group) and use it for alerting:
readinessProbe:
httpGet:
path: /readyz
port: 8080
periodSeconds: 10
See also
- KEDA on Azure Container Apps for the managed-ACA equivalent.
- Observability for the meter the HPA route scrapes.