Table of Contents

KEDA and HPA on AKS

This page documents Orleans.Lattice.Scaling 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at aks.md, and llms.txt lists every page.

Autoscaling an Orleans.Lattice cluster on Azure Kubernetes Service (AKS), or any Kubernetes cluster, using the scaling signal. Two options: a KEDA ScaledObject (recommended) or a native Horizontal Pod Autoscaler (HPA) against a custom metric.

The host wiring is identical to the ACA walkthrough: AddLatticeScalingSignal on the silo, MapLatticeScalingSignal on the web host, and the endpoint served on the pod's container port.

Install KEDA in the cluster, then apply a ScaledObject with a metrics-api trigger pointed at the in-cluster service:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: lattice-scaledobject
spec:
  scaleTargetRef:
    name: lattice-silo            # the Deployment to scale
  minReplicaCount: 2
  maxReplicaCount: 20
  pollingInterval: 15             # seconds; keep above SampleInterval
  advanced:
    horizontalPodAutoscalerConfig:
      behavior:
        scaleDown:
          stabilizationWindowSeconds: 120   # scale-in pacing on top of the gate
  triggers:
    - type: metrics-api
      metadata:
        url: "http://lattice-silo.default.svc.cluster.local/lattice/scale"
        valueLocation: "scaleValue"
        targetValue: "0.5"
  • url targets the headless or ClusterIP service in front of the silo pods; KEDA polls one pod and reads its snapshot, whose activation and resource dimensions are cluster aggregates (the WAL-dispatch dimension and the smoothing state are that pod's own - see cluster-aggregate answering).
  • valueLocation: "scaleValue" and targetValue: "0.5" behave exactly as in the ACA rule: with the trigger's default AverageValue metric type KEDA asks for ceil(scaleValue / targetValue) pods, and because scaleValue never exceeds the current pod count (apart from the MinReplicas floor and an earlier, higher value the scale-in gate holds or the smoothing is still releasing), a targetValue of 1 could never add a pod. The floor is divided too: with the ACA host wiring (MinReplicas = 2) an idle cluster asks for ceil(2 / 0.5) = 4 pods and rests at four rather than at minReplicaCount: 2 - see the minReplicas note under the ACA rule for choosing the two floors.
  • pollingInterval should stay above LatticeScalingSignalOptions.SampleInterval so KEDA never reads a stale sample. Scale-in between minReplicaCount and maxReplicaCount is paced by the managed HPA's scale-down stabilization window (advanced.horizontalPodAutoscalerConfig.behavior.scaleDown.stabilizationWindowSeconds, 300 seconds when unset), which stacks on the signal's own ScaleInGateWindow. KEDA's cooldownPeriod applies only to scaling to zero, so it has no effect with a minReplicaCount of 2.

KEDA creates and manages the underlying HPA for you.

Option 2: HPA against a custom metric

If you prefer a native HPA, expose the scale value as an external metric through the Prometheus adapter (scrape the orleans.lattice.scaling meter via the OpenTelemetry Prometheus exporter, so scaleValue is available as orleans_lattice_scaling_scale_value), then:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: lattice-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: lattice-silo
  minReplicas: 2
  maxReplicas: 20
  metrics:
    - type: External
      external:
        metric:
          name: orleans_lattice_scaling_scale_value
        target:
          type: AverageValue
          averageValue: "0.5"
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 120

Use an External metric, not a Pods one. The value is a cluster-wide demand in replica-units that every silo exports, so an External metric with an AverageValue target of 0.5 makes the HPA divide it by the current pod count and settle on ceil(scaleValue / 0.5) replicas - the same arithmetic as the KEDA rule, including its inability to add a pod at a target of 1 and its rest at twice the signal's MinReplicas floor (the exported gauge carries the floored value). A Pods metric would instead average the near-identical per-pod values and multiply by the current pod count, overshooting by roughly that factor. For the same reason, have the adapter's external-metric query aggregate the per-pod series with max or avg rather than sum.

The KEDA route is preferred because the metrics-api trigger reads the endpoint directly and needs no Prometheus-adapter plumbing; the HPA route is useful when you already run the Prometheus adapter and want a single autoscaling mechanism.

Readiness

The scaling health check can back a probe, but it is not a per-pod signal: it reads the same cached snapshot the endpoint serves, so its activation and resource inputs are the cluster's worst-silo values and every pod reports the same verdict for them (only the WAL inputs are the pod's own). Wired into a readinessProbe, it therefore takes every pod out of rotation together once the cluster's hottest silo crosses the Unhealthy bound - a Degraded result still answers 200 on the default ASP.NET Core status mapping. If that is the behaviour you want, wire it like this; otherwise map it on its own endpoint (or keep it out of the readiness tag group) and use it for alerting:

readinessProbe:
  httpGet:
    path: /readyz
    port: 8080
  periodSeconds: 10

See also