Observability
This page documents Orleans.Lattice.GrainIndex 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at observability.md, and llms.txt lists every page.Every grain index publishes metrics on the shared orleans.lattice meter and
exposes an administrative surface, IGrainIndexAdmin, for status and control.
Metrics
The package adds no meter of its own: its instruments sit on orleans.lattice
alongside the core's, so a host that already collects Lattice metrics picks these
up with no configuration change. GrainIndexMetrics.Meter is that same core meter
instance (LatticeMetrics.Meter), exposed so a listener or custom exporter can
subscribe by reference.
Tags
| Tag | Values | Applies to |
|---|---|---|
index |
the logical index name | every instrument |
tenant |
always _platform_ |
every instrument |
path |
activation, backfill, outbox |
write_failures (all three); grains_enrolled (activation and backfill only) |
The path tag names the route that did the work: activation is the activation
and mutation path that physically writes a grain's entries, backfill is the
background crawl, and outbox is a deferred or retried index write.
The two grains_enrolled series are not additive. Every enrolment is performed -
and counted - by the grain's own activation path, the first time the grain
crosses into the index. The backfill series instead counts every grain whose
activation the crawl drove to completion, so a crawl-driven grain normally
appears on both - except a key whose grain has no persisted state, which projects
nothing, is never enrolled, and so appears on backfill only. Chart the series
side by side rather than summing them. The outbox drain records no enrolments,
only its write_failures.
Every series also carries the repository-wide derived tenant dimension, which
is emitted on tenancy-on and tenancy-off clusters alike so an index telemetry
query is byte-identical in both deployment modes. For grain-index series it is
always the constant platform sentinel _platform_: an index is declared per
grain type and materialises into its own cluster-local tree under the reserved
__grainindex/ namespace, spanning every grain of that type, so a measurement is
a property of a grain type's cluster-wide index rather than of any one tenant's
traffic. That means grain-index series are, by design, invisible to a
tenant-scoped telemetry query.
Instruments
| Instrument | Type | Unit | What it reports |
|---|---|---|---|
orleans.lattice.grainindex.grains_enrolled |
Counter<long> |
{grain} |
Grains onboarded into an index, by the route that onboarded them. |
orleans.lattice.grainindex.entries |
UpDownCounter<long> |
{entry} |
Net change in the number of entries an index holds, so the running sum is the current entry count. |
orleans.lattice.grainindex.write_failures |
Counter<long> |
{failure} |
Failures to publish a grain's index entries, by the route that failed. |
orleans.lattice.grainindex.projection.duration |
Histogram<double> |
ms |
Time to project one grain's state into index entries and diff it against the stored projection. |
orleans.lattice.grainindex.backfill.processed |
ObservableGauge<long> |
{grain} |
Keys a background backfill has taken from its key source. |
orleans.lattice.grainindex.backfill.total |
ObservableGauge<long> |
{grain} |
Best-effort size of the population a backfill has to cover. |
orleans.lattice.grainindex.backfill.percent_complete |
ObservableGauge<double> |
% |
How far through its population a backfill has reached. |
orleans.lattice.grainindex.backfill.state |
ObservableGauge<int> |
{state} |
Lifecycle state of a backfill, as the numeric GrainIndexBackfillState. |
The four observable gauges read a frozen snapshot published by the backfill grain, so a scrape never recomputes progress and only the silo hosting a crawl's activation - the one that knows where the crawl has reached - publishes its series.
backfill.total publishes a series only for an index whose key source returns
an approximate count from TryGetApproximateCountAsync (the default
implementation returns none), and backfill.percent_complete only when that
count is positive - or once the crawl has completed, which reports 100 whatever
its key source. Without a denominator they stay silent rather than reporting a
misleading figure, and backfill.processed is the progress signal. See
The optional count.
Backfill state values
backfill.state reports the numeric GrainIndexBackfillState:
| Value | State |
|---|---|
0 |
NotStarted |
1 |
Running |
2 |
Paused |
3 |
Completed |
4 |
Failed |
What to alert on
write_failuresrising on theactivationpath means state writes are succeeding but their index projection is not. The entries are not lost - they are queued in the outbox - but the index is stale for those grains until the drain catches up. A sustained rise on theoutboxpath means the drain itself is failing.backfill.stateat4(Failed) means a pass faulted as a whole (the error is in the status'sFailureMessage). The crawl is held at its checkpoint and the one-minute reminder heartbeat returns it toRunningon its next tick, so an index that keeps returning to4has a persistent fault to fix rather than a crawl that merely needs resuming.backfill.processedorbackfill.percent_completeflat while the state isRunningmeans the crawl is not advancing. Either each pass is a no-op - the silo hosting the crawl has no key source registered for the index, or does not declare it, so the pass logs a warning and returns - or nothing is driving the passes at all (BackfillEnabled = falsewith no caller ofRunBackfillPassAsync). A key source that throws moves the state toFailedinstead, and one that yields nothing completes the crawl.projection.durationp99 climbing means the projection path is becoming a latency contributor to the grain's own write path under the default synchronous projection mode.
Dashboards
A Grafana dashboard covering these instruments ships in
src/lattice.dashboards/Grafana, and each
instrument's panel is listed in the
metrics-to-panel map.
Prometheus mangles the OTel names. Under .AddPrometheusExporter()
(OpenTelemetry.Exporter.Prometheus.AspNetCore) dots become underscores, a unit
other than a {...} annotation is appended as a word unless the name already
ends with it, and counters gain a _total suffix: so
orleans.lattice.grainindex.grains_enrolled is scraped as
orleans_lattice_grainindex_grains_enrolled_total, projection.duration (unit
ms) as orleans_lattice_grainindex_projection_duration_milliseconds_bucket,
_sum and _count, and backfill.percent_complete (unit %) as
orleans_lattice_grainindex_backfill_percent_complete_percent. An exposition that
appends no unit word, such as the repository-context container's, drops those
unit segments; see
How an instrument name becomes a PromQL series name.
IGrainIndexAdmin
Resolve it from the silo's service provider to inspect and control the declared indexes.
| Member | What it does |
|---|---|
IReadOnlyList<string> DeclaredIndexes |
The names of every index declared in this silo. |
Task<GrainIndexStatus> GetStatusAsync(string indexName, CancellationToken) |
One index's status: its definition, entry count, drift state, and backfill progress. |
Task<IReadOnlyList<GrainIndexStatus>> ListStatusAsync(CancellationToken) |
The same for every declared index. |
Task<GrainIndexBackfillStatus> PauseBackfillAsync(string indexName, CancellationToken) |
Stops scheduling passes, keeping the checkpoint. |
Task<GrainIndexBackfillStatus> ResumeBackfillAsync(string indexName, CancellationToken) |
Resumes from the checkpoint. |
Task<GrainIndexBackfillStatus> RebuildAsync(string indexName, CancellationToken) |
Restarts the crawl from the beginning of the key range. |
Task<GrainIndexBackfillBatchResult> RunBackfillPassAsync(string indexName, CancellationToken) |
Runs exactly one pass now, whatever the schedule says. |
using Orleans.Lattice.GrainIndex;
public static class IndexHealth
{
public static async Task ReportAsync(
IGrainIndexAdmin admin,
CancellationToken cancellationToken)
{
foreach (var status in await admin.ListStatusAsync(cancellationToken))
{
Console.WriteLine($"{status.IndexName}: {status.EntryCount} entries");
}
}
}
An unknown index name throws GrainIndexNotDeclaredException.
RunBackfillPassAsync is the primitive that makes a backfill testable: combined
with BackfillEnabled = false it lets a test drive the crawl one pass at a time
instead of waiting on a schedule.
See also
- Backfill - the crawl these gauges report on.
- Configuration - the options the admin surface reflects.
- Architecture - what
write_failuresimplies.