Table of Contents

Back-pressure health check

This page documents Orleans.Lattice.Replication 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at health-check.md, and llms.txt lists every page.

Orleans.Lattice.Replication ships an ASP.NET Core IHealthCheck that turns the per-peer telemetry maintained by ReplicationPeerStats (see observability) into a single Healthy / Degraded / Unhealthy verdict suitable for a Kubernetes readiness probe, an Azure App Service health endpoint, or any other host that consumes Microsoft.Extensions.Diagnostics.HealthChecks.

The check is purely a consumer of the existing peer telemetry surface. It does not poll, schedule, or invoke RPCs; every probe walks the in-memory ReplicationPeerStats.Snapshot() once and returns. A high-frequency probe (1 Hz or faster) costs roughly an O(peers) dictionary scan and is safe to run on the same cadence as the host's other readiness checks.

Registration

siloBuilder.AddLatticeReplication(o =>
{
    o.ClusterId = "cluster-a";
    o.ReplicatedTrees = new Dictionary<string, LatticeMergeMode>
    {
        ["orders"] = LatticeMergeMode.LwwRegister,
    };
});

siloBuilder.Services
    .AddHealthChecks()
    .AddLatticeReplicationHealthCheck();

AddLatticeReplicationHealthCheck needs AddLatticeReplication on the same service collection: the check reads the cluster-wide singleton ReplicationPeerStats that AddLatticeReplication registers, and resolves it when the check is first constructed at probe time, so the relative order of the two calls does not matter but omitting AddLatticeReplication makes every probe fail. The default registered name is "orleans.lattice.replication" (the same string as LatticeReplicationHealthCheckOptions.DefaultName). Override the name and tags when the host has more than one health check (for example, to expose the replication probe under a ready ASP.NET Core tag):

siloBuilder.Services
    .AddHealthChecks()
    .AddLatticeReplicationHealthCheck(name: "replication", tags: new[] { "ready" });

The extension also takes an optional failureStatus, the status reported when the check itself throws (default Unhealthy); it does not change the threshold-derived Degraded / Unhealthy verdict of a probe that completes.

Threshold tiers

The check classifies every (tree, peer) pair captured in telemetry against three orthogonal signals. Each signal has a soft (degraded) and hard (unhealthy) bound; the worst per-peer classification across all signals becomes the peer's verdict, and the worst per-peer verdict becomes the aggregate probe result. Set a tier to null to disable that signal entirely.

Signal Source Default soft Default hard
EntriesBehind ReplicationPeerSnapshot.EntriesBehind (a per-tick lower bound on the unshipped backlog: the size of the just-shipped batch when the drain filled the batch cap, otherwise 0) 1 000 10 000
LastContactSeconds ReplicationPeerSnapshot.LastContactSeconds (age of last successful contact) 30 s 300 s
ConsecutiveErrors ReplicationPeerSnapshot.ConsecutiveErrors (failure streak since last success) 5 50

The EntriesBehind, LastContactSeconds, and ConsecutiveErrors tiers above classify outbound snapshot rows only - inbound rows carry zero EntriesBehind by construction, and no tier classifies an inbound row's ConsecutiveErrors streak. The inbound counterpart of the contact tier is exposed separately as the inbound silence signal:

Property Description Default
InboundDegradedAfter Duration of inbound silence after which the row contributes Degraded to the aggregate verdict. Timeout.InfiniteTimeSpan (disabled)
InboundCriticalAfter Duration of inbound silence after which the row contributes Unhealthy to the aggregate verdict. Timeout.InfiniteTimeSpan (disabled)

The inbound signal is opt-in - a host that wants readiness gating on inbound liveness configures finite thresholds. A peer that this silo only ships to (and never receives from) produces no inbound rows and is excluded from this signal regardless of the configured thresholds. Inbound rows appear in the degradedPeers / unhealthyPeers arrays with the label suffix " (inbound)" so dashboards can distinguish them from outbound rows.

An inbound row's silence is the time since this silo last finished applying inbound entries the peer authored. Every receive path records it for a run the receiver admits (an enrolled tree whose wire merge mode matches; a dropped run records nothing) - a multi-entry batch once per per-origin run, and a single-entry push (the usual shape from a low-rate sender) or a batch applied entry by entry once per entry. An empty liveness-probe push applies nothing and does not refresh it, so on a link that carries no writes the inbound silence keeps growing even while the peer is healthy: size finite thresholds above the longest write gap you expect from each peer. A row that has only ever recorded failed applies has no successful contact (NaN) and is skipped by this signal.

Because the shipper records EntriesBehind from a single drain, the reading never exceeds the effective ship batch size (at most ShipBatchSize, default 256). With the default ShipBatchSize the default 1 000 / 10 000 bounds therefore cannot trip; a host that relies on this signal sets bounds below its ship batch size.

Defaults are exposed as public static readonly fields on LatticeReplicationHealthCheckOptions (DefaultEntriesBehind, DefaultLastContactSeconds, DefaultConsecutiveErrors, DefaultUnhealthyAfter, DefaultInboundDegradedAfter, DefaultInboundCriticalAfter). The check reads the named options instance that matches its registered name (LatticeReplicationHealthCheckOptions.DefaultName unless overridden), so bind overrides under that name - an unnamed Configure<LatticeReplicationHealthCheckOptions>(...) targets the default options instance, which the check never reads. Each tier is a (Degraded, Unhealthy) pair: LatticeReplicationHealthCheckOptions.LongTier for EntriesBehind and ConsecutiveErrors, and LatticeReplicationHealthCheckOptions.DoubleTier for LastContactSeconds. A host overrides any subset:

siloBuilder.Services.Configure<LatticeReplicationHealthCheckOptions>(LatticeReplicationHealthCheckOptions.DefaultName, o =>
{
    // Tighter back-pressure bound for an interactive workload.
    o.EntriesBehind = new LatticeReplicationHealthCheckOptions.LongTier(200, 2_000);
    // Disable the contact-age signal entirely - rely on the error streak alone.
    o.LastContactSeconds = null;
});

A peer whose observed signal is strictly greater than the soft bound classifies as at least Degraded; strictly greater than the hard bound classifies as Unhealthy immediately - no sustained-degraded grace window applies on the hard path.

Sustained-degraded escalation

A transient degraded blip (one or two probes) is usually noise; a sustained degraded state is a real back-pressure event. The check escalates a peer that has remained Degraded for longer than UnhealthyAfter to Unhealthy:

  • The escalation timer starts on the transition from Healthy to Degraded.
  • It is reset to null the moment the peer drops back below every soft bound. A subsequent re-degradation starts a fresh timer.
  • A peer that hits the hard (Unhealthy) bound on any signal always reports Unhealthy immediately and drops any prior degraded-since record, so a future recovery starts cleanly.
  • A non-positive UnhealthyAfter (TimeSpan.Zero, Timeout.InfiniteTimeSpan, or any negative) disables sustained-degraded escalation entirely - it does not mean "escalate immediately" - leaving a hard (unhealthy) bound on some signal as the only path to Unhealthy. For the strictest gating, use a small positive value: escalation then fires on the first probe after the one that entered the degraded tier.

Default is 60 seconds, sized to absorb one or two probe-cadence blips while escalating within an interactive operator-response window.

The per-peer "first-degraded-at" map lives on the health-check instance itself, so AddLatticeReplicationHealthCheck registers the check as a singleton on the underlying ServiceCollection (the default IHealthChecksBuilder.AddCheck<T> lifetime is transient, which would discard the escalation map on every probe). A custom registration that wants to replace the check must respect that lifetime, otherwise sustained-degraded escalation will silently stop firing.

NaN contact samples

ReplicationPeerSnapshot.LastContactSeconds is double.NaN for a peer that has never had a successful contact recorded - either because it has only ever failed (RecordError without a paired RecordSuccess) or because it has only ever had backlog recorded against it. The check excludes NaN samples from the LastContactSeconds tier, on the principle that "we have never tried" and "we tried and failed" are operationally distinct conditions. The ConsecutiveErrors tier covers the latter so the two signals are orthogonal: a peer that fails its first ship attempt reports ConsecutiveErrors = 1 and LastContactSeconds = NaN, contributes to the error tier, and is silent on the contact tier.

Probe result shape

The check returns a HealthCheckResult whose Data dictionary populates a standard set of keys for a dashboard or alert rule to pivot on:

Key Type Description
peers int Total number of telemetry rows - one per (tree, peer, direction), so a pair this silo both ships to and receives from counts once per direction.
degraded int Count of peers in the degraded tier (after sustained-degraded escalation has been applied).
unhealthy int Count of peers in the unhealthy tier.
degradedPeers string[] tree/peer labels for every degraded peer. Present only when degraded > 0.
unhealthyPeers string[] tree/peer labels for every unhealthy peer. Present only when unhealthy > 0.

The aggregate Status is the worst per-peer classification across the snapshot. An empty snapshot returns Healthy with peers = 0 - a fresh silo that has not yet recorded any peer telemetry is by definition not in back-pressure.

Telemetry-drop garbage collection

A peer that drops out of ReplicationPeerStats.Snapshot() between probes (e.g. a cross-cluster peer the operator removed from the topology) is cleared from the degraded-since map on the next probe. A future re-appearance starts a fresh grace window; the implementation does not retain stale "first-degraded-at" records past the point where the underlying telemetry stops including the peer.

The metrics in observability are the raw inputs the health check classifies:

  • orleans.lattice.replication.peer.entries_behind feeds the EntriesBehind tier.
  • orleans.lattice.replication.peer.last_contact_seconds feeds the LastContactSeconds tier.
  • orleans.lattice.replication.peer.consecutive_errors feeds the ConsecutiveErrors tier.

A host that exports the meter to OpenTelemetry can recreate the threshold tiers as alert rules and have the health check act purely as a gate for the orchestrator's readiness probe. The meter and the health check are independent surfaces; either, both, or neither may be subscribed without impacting the other.