---
title: "KEDA on Azure Container Apps"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.scaling/keda-aca.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice.scaling/keda-aca.md"
package: "Orleans.Lattice.Scaling"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.scaling/llms-full.txt"
---
# KEDA on Azure Container Apps

Part of the [Scaling documentation](README.md).

An end-to-end walkthrough for autoscaling an Orleans.Lattice cluster on Azure
Container Apps (ACA) using the scaling signal as a KEDA custom metric. The
[`ClusterScaling` sample](../../samples/ClusterScaling/README.md) is a complete,
deployable implementation of everything below.

## The shape

ACA scales container-app replicas with KEDA under the hood. The
`metrics-api` KEDA scaler polls an HTTP endpoint, reads a numeric value out of the
JSON response, and drives the replica count toward a target. The scaling
endpoint is exactly that: a GET endpoint returning a JSON body with a top-level
`scaleValue`.

```
   KEDA metrics-api scaler                Orleans.Lattice silo (each replica)
   ----------------------                 -----------------------------------
   GET https://<app>/lattice/scale  --->  MapLatticeScalingSignal()
        reads $.scaleValue          <---  { "scaleValue": 3.4, ... }
        targetValue: 0.5
        desiredReplicas = ceil(scaleValue / targetValue)
```

Because the activation and resource dimensions are cluster aggregates, KEDA can
poll any replica and read a whole-cluster demand. The WAL-dispatch dimension and
the smoothing state are the answering replica's own, so two replicas can answer
slightly differently - see
[cluster-aggregate answering](architecture.md#cluster-aggregate-answering).

## Host wiring

Register the signal on the silo, map the endpoint, and listen on the ACA ingress
target port:

```csharp verify
using Orleans.Lattice.Scaling;

siloBuilder.AddLatticeScalingSignal(options =>
{
    // Never recommend fewer than the minimum viable cluster.
    options.MinReplicas = 2;
    // Refresh well inside KEDA's poll interval.
    options.SampleInterval = TimeSpan.FromSeconds(5);
});
```

```csharp verify
using Orleans.Lattice.Scaling;

var app = WebApplication.Create();
app.MapLatticeScalingSignal(); // GET /lattice/scale
```

## The custom scale rule

Add a `custom` scale rule of type `metrics-api` to the container app. In bicep:

```bicep
// Built from the app name and the environment's default domain: referencing the
// container app's own ingress FQDN inside its own declaration is a circular
// self-reference, which Bicep rejects.
var scaleSignalUrl = 'https://${appName}.${environment.properties.defaultDomain}/lattice/scale'

// Inside the Microsoft.App/containerApps resource's properties.template:
scale: {
  minReplicas: 2
  maxReplicas: 20
  rules: [
    {
      name: 'lattice-scale'
      custom: {
        type: 'metrics-api'
        metadata: {
          url: scaleSignalUrl
          valueLocation: 'scaleValue'
          targetValue: '0.5'
          activationTargetValue: '1'
        }
      }
    }
  ]
}
```

- `valueLocation: 'scaleValue'` - the JSON path KEDA reads. The endpoint emits it
  as a stable, camelCase top-level property.
- `targetValue: '0.5'` - the per-replica pressure the rule holds the pool at.
  `scaleValue` is the dominant compute pressure (`0.0` to `1.0`) times the current
  replica count, so it never exceeds that count except at the `MinReplicas` floor
  or while the scale-in gate holds - or the smoothing releases - an earlier,
  higher value. An ACA custom scale rule
  exposes no metric type, so KEDA's default `AverageValue` applies and it computes
  `desiredReplicas = ceil(scaleValue / targetValue)`: the pool grows when the
  dominant pressure rises above `targetValue`, and shrinks - once the scale-in
  gate lets the value fall - when fewer replicas would still carry the same load
  at or below `targetValue`. A `targetValue` of `1`
  therefore never adds a replica - at full saturation it asks for exactly the
  current count - so scale-out needs a value below `1`. `0.5` is the value the
  [`ClusterScaling` sample](../../samples/ClusterScaling/README.md) deploys: a
  saturated pool asks for twice its current size.
- `activationTargetValue: '1'` - the threshold above which KEDA activates the app
  from zero (if you allow scale-to-zero). Keep it aligned with your `MinReplicas`.
- `minReplicas` / `maxReplicas` - the ACA replica envelope. Set `minReplicas` to
  your quorum floor and `maxReplicas` to your capacity ceiling. The rule divides
  the signal's `MinReplicas` floor by `targetValue` like any other `scaleValue`,
  so once the answering replica has taken its first sample it never asks for
  fewer than `ceil(MinReplicas / targetValue)` replicas. With the host wiring
  above (`MinReplicas = 2`) an idle pool's `scaleValue` sits at the floor, 2, and
  the rule asks for `ceil(2 / 0.5)` = 4 replicas, so the pool rests at four
  (capped at `maxReplicas`) rather than at `minReplicas: 2`. For the pool to rest
  at `minReplicas`, keep `MinReplicas` at or below `minReplicas` times
  `targetValue` - at most 1 here; its default of `0` leaves the floor entirely to
  `minReplicas`.

## Polling and stabilization versus EWMA

KEDA polls on its own interval (`pollingInterval`, default 30s) and ACA applies
its own cooldown before scaling in. These stack on top of the signal's own
smoothing:

- **Scale-out** is fast on both sides: the signal snaps up immediately, and KEDA
  scales out on its next poll.
- **Scale-in** is deliberately slow: the signal only lets the scalar fall after
  every scale-in precondition has held for `ScaleInGateWindow` (default 2m), and
  KEDA/ACA then apply their own cooldown on top. Tune `EwmaHalfLife` and
  `ScaleInGateWindow` for how conservative you want scale-in, and leave the
  KEDA/ACA cooldown to guard against poll-to-poll flapping.

Keep `SampleInterval` (signal freshness) well below `pollingInterval` (KEDA read
cadence) so KEDA never reads a stale sample.

## Health probes

The scaling health check reports the cached snapshot's verdict, and its
activation and resource inputs are the cluster's worst-silo values, so every
replica reports the same verdict for them (only the WAL inputs are the replica's
own). A readiness probe on it therefore takes every replica out of rotation
together when the hottest silo crosses the `Unhealthy` bound - see
[readiness on AKS](aks.md#readiness). Register it with a tag so you can choose
which probe endpoint includes it:

```csharp verify
using Microsoft.Extensions.DependencyInjection;
using Orleans.Lattice.Scaling;

var services = new ServiceCollection();
services.AddHealthChecks().AddLatticeScalingHealthCheck(tags: new[] { "ready" });
```

Map an endpoint filtered to the `ready` tag and set it as the ACA readiness probe
path only if that cluster-wide drain is what you want; otherwise expose it on a
separate endpoint for alerting.

## See also

- [AKS](aks.md) for the Kubernetes `ScaledObject` and HPA alternatives.
- [Configuration](configuration.md) for every knob the walkthrough tunes.
- [`ClusterScaling` sample](../../samples/ClusterScaling/README.md) for the full deployable implementation.
