Table of Contents

Orleans.Lattice.Scaling

This page documents Orleans.Lattice.Scaling 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.

Opt-in autoscaling signal for Orleans.Lattice: a read-only, cluster-aggregate, two-axis (compute and storage) pressure snapshot that an external autoscaler can scrape to size the silo pool.

The package answers one question - how many silo replicas does this cluster need right now? - and answers a second, advisory one alongside it - is any WAL storage account hot enough that its partitions should be spread across more accounts? The first drives replica autoscaling on the compute axis; the second is signal-only and never moves a replica.

Why it exists

Orleans clusters that back a B+ tree store scale on two independent axes:

  • Compute: grain-activation working set, host CPU and memory, and the write-ahead-log dispatch pipeline. When these run hot the fix is more silos.
  • Storage: retained WAL bytes and per-account backend write throughput. When a single storage account tops out, adding silos does not help; the fix is to spread WAL partitions across more accounts.

A generic CPU-based autoscaler conflates the two and either over-provisions compute to relieve a storage bottleneck or ignores a genuine compute shortfall. Orleans.Lattice.Scaling separates them: it publishes a single compute-axis replica-demand scalar that an external autoscaler (KEDA, an HPA custom metric, or an Azure Container Apps custom scale rule) consumes directly, and it reports storage-axis pressure as an operator recommendation that maps onto the ILatticeAdmin WAL-move workflow.

The two axes

  • Compute axis (ComputePressure) - normalised activation, host-resource, and WAL-dispatch pressure (each 0.0 idle to 1.0 saturated), plus the worst-case WalSaturationState the answering silo has observed across its trees. The dominant dimension times the current replica count is the raw replica demand.
  • Storage axis (StoragePressure) - whether the retained WAL bytes of the trees that declare a WalMaxRetainedBytes ceiling reached the configured fraction of those ceilings, the aggregate retained bytes, a per-catalogue-key breakdown (WalAccountPressure) that classifies each account as throughput-bound or capacity-bound, and an optional WalRebalanceRecommendation.

Both axes roll up into a single ScalingSignal carrying the smoothed, scale-in-gated ScaleValue (in replica-units) an autoscaler should act on, a concrete RecommendedReplicas count, the raw RawScaleValue before smoothing, a human-readable Reason, and a SampledAt timestamp. The storage axis never contributes to ScaleValue - it is reported for operator action only.

Quick start

Register the signal on the silo builder, then expose it over an HTTP endpoint on the co-hosted web host (the optional ASP.NET Core health check is covered in Configuration):

using Orleans.Lattice.Scaling;

siloBuilder.AddLatticeScalingSignal(options =>
{
    options.MinReplicas = 2;
    options.SampleInterval = TimeSpan.FromSeconds(5);
});
using Orleans.Lattice.Scaling;

var app = WebApplication.Create();

// KEDA / ACA scrape target - serves the ScalingSignal as JSON with a top-level
// scaleValue property the autoscaler reads.
app.MapLatticeScalingSignal();

Resolve ILatticeScalingSignal anywhere in the container to read the current snapshot directly:

using System.Threading;
using Orleans.Lattice.Scaling;

async Task InspectAsync(ILatticeScalingSignal signal, CancellationToken cancellationToken)
{
    ScalingSignal current = await signal.GetScalingSignalAsync(cancellationToken);
    Console.WriteLine($"scaleValue={current.ScaleValue} replicas={current.RecommendedReplicas} reason={current.Reason}");
}

Documentation

Document What it covers
Architecture Collectors, aggregation, EWMA smoothing, asymmetric scale-in gating, cluster-aggregate answering, and the storage-axis-never-scales-replicas invariant.
Configuration Every LatticeScalingSignalOptions and LatticeScalingHealthCheckOptions knob, its default, and guidance.
API ILatticeScalingSignal, the snapshot DTOs, and the registration and endpoint extension methods.
KEDA on Azure Container Apps End-to-end ACA walkthrough: the metrics-api custom scale rule, targetValue, min/max replicas, and polling versus EWMA.
KEDA and HPA on AKS A KEDA ScaledObject and the HPA custom-metric alternative on Kubernetes.
Storage pressure How per-account WAL pressure maps to the multi-account fan-out remediation and the ILatticeAdmin move workflow (signal-only in this release).
Observability The orleans.lattice.scaling meter instruments and the bundled Grafana dashboard.

Sample

samples/ClusterScaling is a deployable Azure Container Apps sample: a multi-silo Orleans cluster on real Azure Storage clustering and WAL via managed identity, co-hosting the gRPC data API and the scaling endpoint, with a bundled load driver that drives the compute axis so KEDA scales the replica count out.