Metrics
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at metrics.md, and llms.txt lists every page.Orleans.Lattice publishes runtime telemetry through System.Diagnostics.Metrics, so any
OpenTelemetry-compatible exporter (Prometheus, OTLP, Azure Monitor, Datadog,
etc.) can subscribe once to the library's meter and receive every instrument
without per-metric wiring.
Contents
- Meter: All instruments are owned by a single static
Meterexposed viaOrleans.Lattice.LatticeMetrics. - Tag conventions: Every Lattice instrument draws its tags from one consistent, low-cardinality vocabulary.
- Instrument catalog: To derive ops/sec, compute the rate of
shard.reads + shard.writesat the collector; the same underlying counters back the internal hotness monitor that drives autonomic splitting.- Process identity to Shard-level and tree registry
- Leaf-level (sourced from BPlusLeafGrain)
- Snapshot cursors (sourced from SnapshotLeafGrain / snapshot-cursor open path)
- WAL garbage collector (sourced from LatticeWalGc)
- WAL compaction (sourced from FileWalShard) to Foreground read envelopes (sourced from LatticeGrain)
- Shard-root SetManyAsync decomposition (sourced from ShardRootGrain) to WAL append pipeline (sourced from WalShardGrain and WalCommitLogWriter)
- Storage-provider commit pipeline to Grain-call observation (opt-in)
- Replication meter: The replication package (
Orleans.Lattice.Replication) publishes its own meter,orleans.lattice.replication, so an operator can subscribe to cross-cluster telemetry independently of the core lattice surface. - OpenTelemetry registration: Register the meter by name - this is the same pattern used for any other
System.Diagnostics.Metricssource. - Exposing metrics to an AI agent: Once the
orleans.latticemeter is scraped into a Prometheus-compatible backend, the opt-inOrleans.Lattice.Api.Mcp.Telemetrycompanion package can expose that backend to an AI agent as read-only Model Context Protocol tools, behind... - Bundled Grafana dashboards: The companion
Orleans.Lattice.Dashboardspackage ships ready-to-import dashboards keyed byLatticeDashboardKind. - Performance: All instruments are zero-allocation on the hot path: counters use primitive
Add, and hot-path histograms useStopwatch.GetTimestamp()deltas rather thanStopwatchinstances (off the hot path,orleans.lattice.warmup.duration... - Relationship to
DiagnoseAsync:ILattice.DiagnoseAsync(docs/lattice/diagnostics.md) returns a point-in-time snapshot intended for operator inspection and troubleshooting.
Meter
All instruments are owned by a single static Meter exposed via
Orleans.Lattice.LatticeMetrics:
| Member | Value |
|---|---|
LatticeMetrics.MeterName |
orleans.lattice |
LatticeMetrics.Meter |
the Meter instance (exposed for reference-based subscription in tests) |
The name is pinned by a regression test (LatticeMetrics_meter_name_is_orleans_lattice) so it cannot drift.
Reference-based subscription is why LatticeMetrics declares its Meter field
above every instrument, and builds every instrument from that field. Static field
initialisers run in declaration order, and a MeterListener publishes
already-existing instruments from inside the callback that may itself be the first
code to touch LatticeMetrics. An instrument declared above the Meter field
would therefore be published while that field is still null, so a listener
matching on ReferenceEquals(instrument.Meter, LatticeMetrics.Meter) would never
enable it and would record nothing without throwing. The ordering is enforced by
MeterFieldDeclarationOrderTests and demonstrated by MeterListeningTests.
The grain-index package (Orleans.Lattice.GrainIndex) publishes its instruments
(orleans.lattice.grainindex.*) on this same meter - GrainIndexMetrics.Meter is
the same instance - rather than on a meter of its own, so they reach any pipeline
that subscribes to orleans.lattice. They are catalogued in that package's
Observability document rather than below.
Replication meter
The replication package (Orleans.Lattice.Replication) publishes its own
meter, orleans.lattice.replication, so an operator can subscribe to
cross-cluster telemetry independently of the core lattice surface. The meter
name is pinned by a regression test in the replication package.
| Member | Value |
|---|---|
LatticeReplicationMetrics.MeterName |
orleans.lattice.replication |
LatticeReplicationMetrics.Meter |
the Meter instance |
The full per-instrument catalogue - ship / apply / lag, causal-apply buffer,
dead-letter, per-peer health, and bootstrap instruments, plus the replication
tag conventions - lives in the replication package's
Observability document, which is the
source of truth for every orleans.lattice.replication.* instrument.
To subscribe to both meters, register them by name in the OpenTelemetry pipeline:
using OpenTelemetry.Metrics;
builder.Services.AddOpenTelemetry()
.WithMetrics(metrics => metrics
.AddMeter("orleans.lattice")
.AddMeter("orleans.lattice.replication")
.AddPrometheusExporter());
OpenTelemetry registration
Register the meter by name - this is the same pattern used for any other
System.Diagnostics.Metrics source:
// In your silo host's Program.cs or similar composition root.
using OpenTelemetry.Metrics;
builder.Services.AddOpenTelemetry()
.WithMetrics(metrics => metrics
.AddMeter("orleans.lattice")
.AddPrometheusExporter()); // or AddOtlpExporter, AddAzureMonitorMetricExporter, etc.
The meter and its instruments are created when LatticeMetrics is first used, and
a pipeline that subscribes by name picks instruments up whenever they are
published, so adding it before the silo starts is sufficient - every
subsequently-activated grain publishes into the already-subscribed pipeline.
AddMeter with an exact name matches that meter only and does not cascade, so
each lattice meter a host registers packages for needs its own call or a
wildcard: AddMeter("orleans.lattice.*") subscribes every child meter
(orleans.lattice.replication, orleans.lattice.auth, and so on) but not the
core orleans.lattice meter itself, whose name has no trailing segment. A deployment that will
be operated, rather than only demonstrated, should also register the two runtime
meters - Microsoft.Orleans (grain activations and activation latency) and
System.Runtime (GC heap, process working set, thread pool). Neither is in the
lattice family, and without them the endpoint carries no heap, process-memory,
or activation-latency series at all. See
Dashboards configuration.
Exposing metrics to an AI agent
Once the orleans.lattice meter is scraped into a Prometheus-compatible
backend, the opt-in Orleans.Lattice.Api.Mcp.Telemetry companion package can
expose that backend to an AI agent as read-only Model Context Protocol tools,
behind a dedicated cluster-wide Telemetry authorization grant and a backend
credential the agent never sees. See
MCP Telemetry.
Performance
All instruments are zero-allocation on the hot path: counters use primitive
Add, and hot-path histograms use Stopwatch.GetTimestamp() deltas rather than
Stopwatch instances (off the hot path, orleans.lattice.warmup.duration and
orleans.lattice.snapshot.replay.duration, which run once per warm-up or
snapshot open rather than per operation, do time with a Stopwatch). When no
listener is attached, the measurement callbacks are elided by the runtime.
Relationship to DiagnoseAsync
ILattice.DiagnoseAsync (docs/lattice/diagnostics.md) returns a
point-in-time snapshot intended for operator inspection and troubleshooting.
The metrics pipeline described here is the continuous telemetry feed for
dashboards and alerting. Both are sourced from the same underlying grain state
(shard hotness counters, leaf statistics) - they are complementary, not
redundant.