Table of Contents

KEDA on Azure Container Apps

This page documents Orleans.Lattice.Scaling 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at keda-aca.md, and llms.txt lists every page.

An end-to-end walkthrough for autoscaling an Orleans.Lattice cluster on Azure Container Apps (ACA) using the scaling signal as a KEDA custom metric. The ClusterScaling sample is a complete, deployable implementation of everything below.

The shape

ACA scales container-app replicas with KEDA under the hood. The metrics-api KEDA scaler polls an HTTP endpoint, reads a numeric value out of the JSON response, and drives the replica count toward a target. The scaling endpoint is exactly that: a GET endpoint returning a JSON body with a top-level scaleValue.

   KEDA metrics-api scaler                Orleans.Lattice silo (each replica)
   ----------------------                 -----------------------------------
   GET https://<app>/lattice/scale  --->  MapLatticeScalingSignal()
        reads $.scaleValue          <---  { "scaleValue": 3.4, ... }
        targetValue: 0.5
        desiredReplicas = ceil(scaleValue / targetValue)

Because the activation and resource dimensions are cluster aggregates, KEDA can poll any replica and read a whole-cluster demand. The WAL-dispatch dimension and the smoothing state are the answering replica's own, so two replicas can answer slightly differently - see cluster-aggregate answering.

Host wiring

Register the signal on the silo, map the endpoint, and listen on the ACA ingress target port:

using Orleans.Lattice.Scaling;

siloBuilder.AddLatticeScalingSignal(options =>
{
    // Never recommend fewer than the minimum viable cluster.
    options.MinReplicas = 2;
    // Refresh well inside KEDA's poll interval.
    options.SampleInterval = TimeSpan.FromSeconds(5);
});
using Orleans.Lattice.Scaling;

var app = WebApplication.Create();
app.MapLatticeScalingSignal(); // GET /lattice/scale

The custom scale rule

Add a custom scale rule of type metrics-api to the container app. In bicep:

// Built from the app name and the environment's default domain: referencing the
// container app's own ingress FQDN inside its own declaration is a circular
// self-reference, which Bicep rejects.
var scaleSignalUrl = 'https://${appName}.${environment.properties.defaultDomain}/lattice/scale'

// Inside the Microsoft.App/containerApps resource's properties.template:
scale: {
  minReplicas: 2
  maxReplicas: 20
  rules: [
    {
      name: 'lattice-scale'
      custom: {
        type: 'metrics-api'
        metadata: {
          url: scaleSignalUrl
          valueLocation: 'scaleValue'
          targetValue: '0.5'
          activationTargetValue: '1'
        }
      }
    }
  ]
}
  • valueLocation: 'scaleValue' - the JSON path KEDA reads. The endpoint emits it as a stable, camelCase top-level property.
  • targetValue: '0.5' - the per-replica pressure the rule holds the pool at. scaleValue is the dominant compute pressure (0.0 to 1.0) times the current replica count, so it never exceeds that count except at the MinReplicas floor or while the scale-in gate holds - or the smoothing releases - an earlier, higher value. An ACA custom scale rule exposes no metric type, so KEDA's default AverageValue applies and it computes desiredReplicas = ceil(scaleValue / targetValue): the pool grows when the dominant pressure rises above targetValue, and shrinks - once the scale-in gate lets the value fall - when fewer replicas would still carry the same load at or below targetValue. A targetValue of 1 therefore never adds a replica - at full saturation it asks for exactly the current count - so scale-out needs a value below 1. 0.5 is the value the ClusterScaling sample deploys: a saturated pool asks for twice its current size.
  • activationTargetValue: '1' - the threshold above which KEDA activates the app from zero (if you allow scale-to-zero). Keep it aligned with your MinReplicas.
  • minReplicas / maxReplicas - the ACA replica envelope. Set minReplicas to your quorum floor and maxReplicas to your capacity ceiling. The rule divides the signal's MinReplicas floor by targetValue like any other scaleValue, so once the answering replica has taken its first sample it never asks for fewer than ceil(MinReplicas / targetValue) replicas. With the host wiring above (MinReplicas = 2) an idle pool's scaleValue sits at the floor, 2, and the rule asks for ceil(2 / 0.5) = 4 replicas, so the pool rests at four (capped at maxReplicas) rather than at minReplicas: 2. For the pool to rest at minReplicas, keep MinReplicas at or below minReplicas times targetValue - at most 1 here; its default of 0 leaves the floor entirely to minReplicas.

Polling and stabilization versus EWMA

KEDA polls on its own interval (pollingInterval, default 30s) and ACA applies its own cooldown before scaling in. These stack on top of the signal's own smoothing:

  • Scale-out is fast on both sides: the signal snaps up immediately, and KEDA scales out on its next poll.
  • Scale-in is deliberately slow: the signal only lets the scalar fall after every scale-in precondition has held for ScaleInGateWindow (default 2m), and KEDA/ACA then apply their own cooldown on top. Tune EwmaHalfLife and ScaleInGateWindow for how conservative you want scale-in, and leave the KEDA/ACA cooldown to guard against poll-to-poll flapping.

Keep SampleInterval (signal freshness) well below pollingInterval (KEDA read cadence) so KEDA never reads a stale sample.

Health probes

The scaling health check reports the cached snapshot's verdict, and its activation and resource inputs are the cluster's worst-silo values, so every replica reports the same verdict for them (only the WAL inputs are the replica's own). A readiness probe on it therefore takes every replica out of rotation together when the hottest silo crosses the Unhealthy bound - see readiness on AKS. Register it with a tag so you can choose which probe endpoint includes it:

using Microsoft.Extensions.DependencyInjection;
using Orleans.Lattice.Scaling;

var services = new ServiceCollection();
services.AddHealthChecks().AddLatticeScalingHealthCheck(tags: new[] { "ready" });

Map an endpoint filtered to the ready tag and set it as the ACA readiness probe path only if that cluster-wide drain is what you want; otherwise expose it on a separate endpoint for alerting.

See also