---
title: "Observability"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.grainindex/observability.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice.grainindex/observability.md"
package: "Orleans.Lattice.GrainIndex"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.grainindex/llms-full.txt"
---
# Observability

Part of the [GrainIndex documentation](README.md).

Every grain index publishes metrics on the shared `orleans.lattice` meter and
exposes an administrative surface, `IGrainIndexAdmin`, for status and control.

## Metrics

The package adds no meter of its own: its instruments sit on `orleans.lattice`
alongside the core's, so a host that already collects Lattice metrics picks these
up with no configuration change. `GrainIndexMetrics.Meter` is that same core meter
instance (`LatticeMetrics.Meter`), exposed so a listener or custom exporter can
subscribe by reference.

### Tags

| Tag | Values | Applies to |
|---|---|---|
| `index` | the logical index name | every instrument |
| `tenant` | always `_platform_` | every instrument |
| `path` | `activation`, `backfill`, `outbox` | `write_failures` (all three); `grains_enrolled` (`activation` and `backfill` only) |

The `path` tag names the route that did the work: `activation` is the activation
and mutation path that physically writes a grain's entries, `backfill` is the
background crawl, and `outbox` is a deferred or retried index write.

The two `grains_enrolled` series are not additive. Every enrolment is performed -
and counted - by the grain's own activation path, the first time the grain
crosses into the index. The `backfill` series instead counts every grain whose
activation the crawl drove to completion, so a crawl-driven grain normally
appears on both - except a key whose grain has no persisted state, which projects
nothing, is never enrolled, and so appears on `backfill` only. Chart the series
side by side rather than summing them. The outbox drain records no enrolments,
only its `write_failures`.

Every series also carries the repository-wide derived `tenant` dimension, which
is emitted on tenancy-on and tenancy-off clusters alike so an index telemetry
query is byte-identical in both deployment modes. For grain-index series it is
always the constant platform sentinel `_platform_`: an index is declared per
grain *type* and materialises into its own cluster-local tree under the reserved
`__grainindex/` namespace, spanning every grain of that type, so a measurement is
a property of a grain type's cluster-wide index rather than of any one tenant's
traffic. That means grain-index series are, by design, invisible to a
tenant-scoped telemetry query.

### Instruments

| Instrument | Type | Unit | What it reports |
|---|---|---|---|
| `orleans.lattice.grainindex.grains_enrolled` | `Counter<long>` | `{grain}` | Grains onboarded into an index, by the route that onboarded them. |
| `orleans.lattice.grainindex.entries` | `UpDownCounter<long>` | `{entry}` | Net change in the number of entries an index holds, so the running sum is the current entry count. |
| `orleans.lattice.grainindex.write_failures` | `Counter<long>` | `{failure}` | Failures to publish a grain's index entries, by the route that failed. |
| `orleans.lattice.grainindex.projection.duration` | `Histogram<double>` | `ms` | Time to project one grain's state into index entries and diff it against the stored projection. |
| `orleans.lattice.grainindex.backfill.processed` | `ObservableGauge<long>` | `{grain}` | Keys a background backfill has taken from its key source. |
| `orleans.lattice.grainindex.backfill.total` | `ObservableGauge<long>` | `{grain}` | Best-effort size of the population a backfill has to cover. |
| `orleans.lattice.grainindex.backfill.percent_complete` | `ObservableGauge<double>` | `%` | How far through its population a backfill has reached. |
| `orleans.lattice.grainindex.backfill.state` | `ObservableGauge<int>` | `{state}` | Lifecycle state of a backfill, as the numeric `GrainIndexBackfillState`. |

The four observable gauges read a frozen snapshot published by the backfill
grain, so a scrape never recomputes progress and only the silo hosting a crawl's
activation - the one that knows where the crawl has reached - publishes its
series.

`backfill.total` publishes a series only for an index whose key source returns
an approximate count from `TryGetApproximateCountAsync` (the default
implementation returns none), and `backfill.percent_complete` only when that
count is positive - or once the crawl has completed, which reports `100` whatever
its key source. Without a denominator they stay silent rather than reporting a
misleading figure, and `backfill.processed` is the progress signal. See
[The optional count](backfill.md#the-optional-count).

### Backfill state values

`backfill.state` reports the numeric `GrainIndexBackfillState`:

| Value | State |
|---|---|
| `0` | `NotStarted` |
| `1` | `Running` |
| `2` | `Paused` |
| `3` | `Completed` |
| `4` | `Failed` |

### What to alert on

- **`write_failures` rising** on the `activation` path means state writes are
  succeeding but their index projection is not. The entries are not lost - they
  are queued in the [outbox](architecture.md#the-outbox) - but the index is
  stale for those grains until the drain catches up. A sustained rise on the
  `outbox` path means the drain itself is failing.
- **`backfill.state` at `4` (`Failed`)** means a pass faulted as a whole (the
  error is in the status's `FailureMessage`). The crawl is held at its checkpoint
  and the one-minute reminder heartbeat returns it to `Running` on its next tick,
  so an index that keeps returning to `4` has a persistent fault to fix rather
  than a crawl that merely needs resuming.
- **`backfill.processed` or `backfill.percent_complete` flat** while the state
  is `Running` means the crawl is not advancing. Either each pass is a no-op -
  the silo hosting the crawl has no key source registered for the index, or does
  not declare it, so the pass logs a warning and returns - or nothing is driving
  the passes at all (`BackfillEnabled = false` with no caller of
  `RunBackfillPassAsync`). A key source that throws moves the state to `Failed`
  instead, and one that yields nothing completes the crawl.
- **`projection.duration` p99 climbing** means the projection path is becoming a
  latency contributor to the grain's own write path under the default
  synchronous projection mode.

### Dashboards

A Grafana dashboard covering these instruments ships in
[`src/lattice.dashboards/Grafana`](../lattice.dashboards/README.md), and each
instrument's panel is listed in the
[metrics-to-panel map](../lattice.dashboards/metrics-to-panel-map.md).

Prometheus mangles the OTel names. Under `.AddPrometheusExporter()`
(`OpenTelemetry.Exporter.Prometheus.AspNetCore`) dots become underscores, a unit
other than a `{...}` annotation is appended as a word unless the name already
ends with it, and counters gain a `_total` suffix: so
`orleans.lattice.grainindex.grains_enrolled` is scraped as
`orleans_lattice_grainindex_grains_enrolled_total`, `projection.duration` (unit
`ms`) as `orleans_lattice_grainindex_projection_duration_milliseconds_bucket`,
`_sum` and `_count`, and `backfill.percent_complete` (unit `%`) as
`orleans_lattice_grainindex_backfill_percent_complete_percent`. An exposition that
appends no unit word, such as the repository-context container's, drops those
unit segments; see
[How an instrument name becomes a PromQL series name](../lattice.dashboards/metrics-to-panel-map.md#how-an-instrument-name-becomes-a-promql-series-name).

## `IGrainIndexAdmin`

Resolve it from the silo's service provider to inspect and control the declared
indexes.

| Member | What it does |
|---|---|
| `IReadOnlyList<string> DeclaredIndexes` | The names of every index declared in this silo. |
| `Task<GrainIndexStatus> GetStatusAsync(string indexName, CancellationToken)` | One index's status: its definition, entry count, drift state, and backfill progress. |
| `Task<IReadOnlyList<GrainIndexStatus>> ListStatusAsync(CancellationToken)` | The same for every declared index. |
| `Task<GrainIndexBackfillStatus> PauseBackfillAsync(string indexName, CancellationToken)` | Stops scheduling passes, keeping the checkpoint. |
| `Task<GrainIndexBackfillStatus> ResumeBackfillAsync(string indexName, CancellationToken)` | Resumes from the checkpoint. |
| `Task<GrainIndexBackfillStatus> RebuildAsync(string indexName, CancellationToken)` | Restarts the crawl from the beginning of the key range. |
| `Task<GrainIndexBackfillBatchResult> RunBackfillPassAsync(string indexName, CancellationToken)` | Runs exactly one pass now, whatever the schedule says. |

```csharp verify
using Orleans.Lattice.GrainIndex;

public static class IndexHealth
{
    public static async Task ReportAsync(
        IGrainIndexAdmin admin,
        CancellationToken cancellationToken)
    {
        foreach (var status in await admin.ListStatusAsync(cancellationToken))
        {
            Console.WriteLine($"{status.IndexName}: {status.EntryCount} entries");
        }
    }
}
```

An unknown index name throws `GrainIndexNotDeclaredException`.

`RunBackfillPassAsync` is the primitive that makes a backfill testable: combined
with `BackfillEnabled = false` it lets a test drive the crawl one pass at a time
instead of waiting on a schedule.

## See also

- [Backfill](backfill.md) - the crawl these gauges report on.
- [Configuration](configuration.md) - the options the admin surface reflects.
- [Architecture](architecture.md#the-outbox) - what `write_failures` implies.
