Table of Contents

Configuration

This page documents Orleans.Lattice.Scaling 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at configuration.md, and llms.txt lists every page.

Every knob on LatticeScalingSignalOptions (bound through AddLatticeScalingSignal) and LatticeScalingHealthCheckOptions (bound through AddLatticeScalingHealthCheck), with its default and guidance.

LatticeScalingSignalOptions

Configure it with the Action<LatticeScalingSignalOptions> overload of AddLatticeScalingSignal:

using Orleans.Lattice.Scaling;

siloBuilder.AddLatticeScalingSignal(options =>
{
    options.EndpointPath = "/lattice/scale";
    options.MinReplicas = 2;
    options.SampleInterval = TimeSpan.FromSeconds(5);
    options.EwmaHalfLife = TimeSpan.FromSeconds(30);
    options.ScaleInGateWindow = TimeSpan.FromMinutes(2);
    options.ActivationScaleInThreshold = 0.25;
    options.ResourceScaleInThreshold = 0.25;
    options.WalDispatchScaleInThreshold = 0.25;
    options.ActivationWorkingSetTarget = 100_000;
    options.RetainedBytesAdvisoryRatio = 0.8;
    options.AccountSaturationWindow = TimeSpan.FromSeconds(30);
    options.StorageRecommendationsEnabled = true;
});

Endpoint and floor

Option Type Default Guidance
EndpointPath string /lattice/scale The HTTP path MapLatticeScalingSignal serves the signal from. MapLatticeScalingSignal(path) can override it per-call; keep this and the mapped path in sync.
MinReplicas int 0 Lower bound applied to both ScaleValue and RecommendedReplicas - neither is reported below this floor once the first sample lands. Set it to your cluster's minimum viable silo count so scale-in never suggests dropping below quorum. An autoscaler acting on ScaleValue divides this floor by its targetValue like any other value, so under the shipped targetValue of 0.5 a cluster at rest holds ceil(MinReplicas / 0.5) replicas - twice the floor - capped at the autoscaler's maximum. For the pool to rest at quorum instead, set the quorum as the autoscaler's own minimum (minReplicas / minReplicaCount) and keep MinReplicas at or below that minimum times targetValue (see the custom scale rule).

Compute axis

Option Type Default Guidance
SampleInterval TimeSpan 5s How often the silo recomputes the signal. The per-scrape facade reads the cached result, so this is the freshness bound, not the scrape cost. Keep it well below your autoscaler's polling interval. A non-positive value falls back to the 5-second default, and a value above the timer ceiling (about 49.7 days) is clamped to it.
EwmaHalfLife TimeSpan 30s Half-life of the exponentially-weighted moving average applied to the scalar on the scale-in (release) side. Longer damps noise and makes scale-in more conservative; scale-out reacts immediately regardless. A non-positive value disables the smoothing, so a falling scalar is adopted directly once the gate allows it.
ScaleInGateWindow TimeSpan 2m How long every scale-in precondition (all compute dimensions low, WAL healthy, no shard split in flight) must hold continuously before the scalar is allowed to fall. Any break resets the window.
ActivationScaleInThreshold double 0.25 Activation-pressure level (0..1) at or above which the activation dimension is too hot to permit scale-in.
ResourceScaleInThreshold double 0.25 Resource-pressure level (0..1) at or above which the resource dimension is too hot to permit scale-in.
WalDispatchScaleInThreshold double 0.25 WAL-dispatch-pressure level (0..1) at or above which the WAL-dispatch dimension is too hot to permit scale-in.
ActivationWorkingSetTarget int 100000 Per-silo grain-activation count treated as full activation saturation. Activation pressure is activationCount / target, clamped to 0..1, and the worst silo sets the cluster value. Size it to the activation count at which a silo's memory or scheduler starts to strain. A non-positive value disables the activation dimension (it reads 0.0).
SplitAwareScaleIn bool true Whether scale-in is suppressed while any adaptive shard split is in flight cluster-wide. Reads ILatticeAdmin.GetSplitActivityAsync once per SampleInterval - a single call to the split-admission singleton, never a fan-out. Set to false to make the axis inert (a deployment with autonomic splitting disabled, where the query is pure overhead). Scale-out is never influenced either way. See split-aware scale-in.

Storage axis

The storage axis is report-only: none of these knobs affect the compute ScaleValue.

Option Type Default Guidance
RetainedBytesAdvisoryRatio double 0.8 Fraction of an account's retention budget at or above which its retained WAL bytes count as capacity pressure. The budget is the sum of the effective per-tree WalMaxRetainedBytes ceilings (runtime override, then named options, then the silo default) of the trees holding partitions on that account, attributed across their partitions. A value above 1 is treated as 1, and a non-positive or NaN value falls back to the 0.8 default. A tree whose ceiling is null contributes neither bytes nor budget, so an account backing only such trees never reports capacity pressure.
AccountSaturationWindow TimeSpan 30s How long a provider key must be continuously observed saturated before the collector classifies it throughput-bound and recommends a move. Debounces a transient blip. A non-positive value classifies on the first saturated sample.
StorageRecommendationsEnabled bool true Master switch for emitting a WalRebalanceRecommendation. When false the collector still reports OverThreshold and the per-account breakdown but leaves Recommendation null.

Every default above is also a public constant or static field on LatticeScalingSignalOptions - DefaultEndpointPath, DefaultMinReplicas, DefaultSampleInterval, DefaultEwmaHalfLife, DefaultScaleInGateWindow, DefaultScaleInThreshold (shared by the three scale-in thresholds), DefaultActivationWorkingSetTarget, DefaultSplitAwareScaleIn, DefaultRetainedBytesAdvisoryRatio, DefaultAccountSaturationWindow, and DefaultStorageRecommendationsEnabled - so a host can reference them rather than repeat the values.

LatticeScalingHealthCheckOptions

AddLatticeScalingHealthCheck registers an ASP.NET Core health check that projects the signal onto a single HealthStatus. Bind the named options under the check's registered name (default orleans.lattice.scaling):

using Microsoft.Extensions.DependencyInjection;
using Orleans.Lattice.Scaling;

var services = new ServiceCollection();

services.AddHealthChecks().AddLatticeScalingHealthCheck(tags: new[] { "ready" });

services.Configure<LatticeScalingHealthCheckOptions>(
    LatticeScalingHealthCheckOptions.DefaultName,
    options =>
    {
        options.ComputePressure = new LatticeScalingHealthCheckOptions.DoubleTier(0.85, 0.95);
        options.UnhealthyOnWalSaturated = true;
        options.DegradeOnWalThrottled = true;
        options.DegradeOnStorageOverThreshold = true;
    });
Option Type Default Guidance
ComputePressure DoubleTier? 0.85 / 0.95 (DefaultComputePressure) Tiered bound on the worst normalised compute dimension: at or above the soft bound reports Degraded, at or above the hard bound reports Unhealthy. Set to null to disable the tiered compute signal.
UnhealthyOnWalSaturated bool true When true, a Saturated worst-case WAL state reports Unhealthy regardless of the compute ratios.
DegradeOnWalThrottled bool true When true, a Throttled worst-case WAL state contributes Degraded.
DegradeOnStorageOverThreshold bool true When true, an over-threshold storage axis contributes Degraded. The storage axis never escalates past Degraded because it is advisory and not wired to the replica recommendation.

The failureStatus and tags arguments to AddLatticeScalingHealthCheck are the standard ASP.NET Core health-check registration parameters: failureStatus is the status reported when the check throws (defaults to Unhealthy), and tags let a host filter the check into a readiness or liveness probe group.

See also