Orleans.Lattice.Dashboards
This page documents Orleans.Lattice.Dashboards 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.Orleans.Lattice.Dashboards is a sibling package that ships pre-built Grafana dashboards and provisioning templates for the orleans.lattice, orleans.lattice.replication, orleans.lattice.replication.grpc, orleans.lattice.auth, orleans.lattice.membership, orleans.lattice.backup, orleans.lattice.scaling, and orleans.lattice.tenancy meters, and its Overview dashboard also charts the exact-KNN gather instruments of the repository-context Orleans.Lattice.Api.Mcp.RepoContext meter. Install it when you want operator dashboards bundled with the library version - the core library has no dependency on it.
What is it?
The package is a thin, dependency-light delivery vehicle for operator dashboards. Each dashboard is an embedded Grafana JSON resource, retrieved by a typed kind through a single public accessor, so the dashboards travel with the exact library version that emits the metrics they chart.
- Bundled, version-pinned dashboards. Operator dashboards for the overview, commit path, replication, replication gRPC transport security, atomic writes, materialised views, identity/authorization, backup/restore, autoscaling-signal, per-tenant observability, and grain-index surfaces ship as embedded Grafana JSON, each retrieved by a typed kind. Each package version is built and drift-guarded against the library source it ships with, so a dashboard never references an instrument that library version does not emit.
- No replication dependency. The package takes a runtime dependency only on
Orleans.Lattice(the core library). The replication dashboard's queries reference instruments on theorleans.lattice.replicationmeter, but the package does not link againstOrleans.Lattice.Replication; that meter is only emitted when the replication package is registered on the silo separately. Local-only deployments install the dashboards without pulling in replication and simply omit the replication dashboard. - Drift-guarded coverage. Every metric name a dashboard references resolves to an instrument declared in the library source, and every instrument the guard can discover is referenced by at least one panel, apart from a short allow-list left uncharted on purpose - both directions are enforced by CI tests so a rename or a new unpaneled instrument fails the build before it ships stale. The core guard's forward direction reaches only instruments published by, or named by instrument-name constants on, the two static metric classes; each add-on package whose meter has its own dashboard, and the grain-index package that publishes on the core meter, enforces the forward direction for its own instruments from its own tests. The repository-context meter is held to its metric-to-panel map rows instead - every instrument has a row, and each row's charted or not-charted claim must agree with the panels - and the
orleans.lattice.api.mcpmeter is deliberately left uncharted. See Architecture for the one gap this leaves and the instruments currently outside it.
Core Properties
- Self-contained. Dashboards are embedded resources; retrieving one is a synchronous in-process call with no I/O and no external service.
- Import-ready. Each accessor returns a complete Grafana dashboard model (panels, templating, time range) suitable for direct import or file-system provisioning.
- Operator-oriented. The kinds map to focused operator workflows - for example at-a-glance overview, commit-path latency triage, cross-cluster replication lag, atomic-write saga deep-dive, materialised-view freshness, identity/authorization enforcement, per-tenant usage against quota, and grain-index backfill progress.
Features
| Dashboard | Source meter | Focus |
|---|---|---|
Overview |
orleans.lattice, plus Orleans.Lattice.Api.Mcp.RepoContext for its exact-KNN panels |
Throughput and active shards, leaf-write percentiles and leaf-scan p95, cache hit-rate, tombstone churn, splits, autonomic split admission, retroactive split-forward, shard consolidation and healing, compaction (duration, pass duration, leaves visited, shard retries / skips, dirty leaves, tombstone ratio), atomic-write outcomes, coordinator completions and phase-tick failures, tree-lifecycle, events, runtime config changes, and top-of-stack read-path latency envelopes and shard-root optimistic point-read outcomes, plus storage footprint and byte-pressure trim, WAL compression and dictionary training, the WAL saturation regime, snapshot replay and pins, write admission, the distributed-lock, atomic-action (saga / TCC) and grain-call contention rows, tree-registry fan-in, leaf-division and split-completion health, the repository-context exact-KNN gather work, wall time and budget outcomes, and the deployed build. Includes three side-by-side atomic-write panels (saga duration p50/p95/p99, batch size p50/p95/p99, and a saga-failure-rate panel with 1% / 5% threshold lines). |
CommitPath |
orleans.lattice |
WAL-first commit path: per-step commit latency (wal / apply / observer / digest), leaf commit concurrency and per-observer inline latency, SetAsync / SetManyAsync envelope and stage breakdowns, the WAL append, writer-admission and shard-dispatch pipeline, storage-provider write latency, phase-2 commit, retries and timeouts, LatticeSaturatedException refusals by refusing seam and the replay-admission refusals by arm, compaction latency and TTL tombstone churn, digest publish, reshard and warm-up activity, scan-page leaf-read coalescing and the shard-root wedge guards, and leaf lifecycle diagnostics - the materialiser pin path, activation and deactivation outcomes, snapshot capture and hydration, the WAL replay permit gate, deferred-terminal ledger refusals, the resident leaf working set, span fail-open commits, and the WAL GC blocked-consumer population. |
Replication |
orleans.lattice.replication, plus orleans.lattice for its WAL, WAL-compaction and WAL GC panels |
Ship / apply / lag percentiles, ship ack latency and adaptive batch size, apply parallelism, WAL ship vs trim throughput and the log-tailing producer's append vs ship rate, dead-letter queue churn, apply FIFO and causal violations, causal apply-buffer occupancy and the dependency-wait histogram, fell-off-log and suppressed fall-off events, per-peer entries / bytes behind, batches in flight, last contact, consecutive errors, negotiated wire version and down-stamps, anti-entropy digest probes, Merkle walks, leaf re-replay and drift remediation, bootstrap and bootstrap-fallback activity, shipping optimisations (coalescing, doorbell, content-hash elision, shared-dictionary compression), coordinated-restore saga panels, and the core meter's WAL GC, WAL compaction and WAL recovery-discard panels. |
AtomicWrites |
orleans.lattice |
Dedicated SetManyAtomicAsync saga deep-dive: outcome rate, saga duration p50/p95/p99 and p95 by outcome, batch size p50/p95/p99 and p95 by outcome, per-tree committed throughput, range-window non-committed saga count, a separate saga-failure-rate panel, saga phase-duration and broadcast sub-attribution breakdowns, per-key saga work and fan-out size, cross-tree atomic-write outcomes, failure rate, coordinator duration and participant fan-out, and the saga decision registry's write rate by outcome, group-commit coalescing factor and write duration. The right home for incident triage and SLO drill-down; the Overview row is the at-a-glance teaser. |
MaterialisedViews |
orleans.lattice |
Cluster-wide materialised-view health: apply-lag and drain-backlog-depth percentiles, filter / re-project and aggregation apply throughput, and warning panels for lag-budget evictions, re-key collisions, atomic-staging backstop fall-backs, cross-tree joint-atomicity violations, source back-pressure self-throttling, and aggregation reserved-key rejections. Keyed by view name (and cluster); deliberately offers no per-silo filter, because a view's maintainer is a single grain activation that migrates between silos, so the dashboard aggregates across the whole cluster. Needs only the core meter, not the replication package. |
Authorization |
orleans.lattice.auth, orleans.lattice.membership |
Identity and authorization operator view: enforcement-gate decision throughput (by effect and operation), decision-latency percentiles, compiled-snapshot rebuild rate and the snapshot epoch / age / subjects gauges, alongside the subject-resolution cache hit-ratio and hit / miss throughput and the identity-directory search latency percentiles, hit / miss throughput and hit ratio. Useful only when the authentication / authorization packages are registered on the silo. |
Backup |
orleans.lattice.backup |
Backup / restore operator view: capture / restore throughput and duration percentiles, per-backup size / artifact / entry distributions, cumulative processed throughput, retention reclaim and prune rates, incremental lag (entries and age behind the base cut), capture / restore failure rates by reason, capture retries, scheduler skipped-run, overrun and capture-failure counters, the cross-tree-consistent fence selection / drain / retry counters and drain-wait percentiles, and the inventory gauges (tracked count, max chain depth, catalog bytes, oldest / newest age, and per-scope last-run status and last-success age). Useful only when the backup package is registered on the silo. Its scope selector narrows only the retention and scheduler panels, whose instruments carry a scope tag; the other panels apply no scope matcher - see the backup meter rows. |
Scaling |
orleans.lattice.scaling |
Autoscaling-signal operator view: the two scale-value gauges (the smoothed, scale-in-gated value an autoscaler acts on and the raw, un-smoothed instantaneous demand), the three normalised compute-pressure dimensions (activation / host-resource / WAL-dispatch), the recommended silo replica count, and the storage-axis stats (WAL catalogue keys over the advisory threshold and whether a WAL rebalance is recommended). Useful only when the scaling package is registered on the silo. |
ReplicationGrpc |
orleans.lattice.replication.grpc |
Replication gRPC transport-security view: the insecure (plaintext) channel construction counter as a cluster-wide cumulative total and as a per-second rate broken out by peer cluster id and transport (push / saga_control / snapshot), so an accidental production plaintext downgrade under AllowPlaintextEndpoints is visible rather than silent. Useful only when the gRPC replication transport is registered on the silo. |
Tenancy |
orleans.lattice.tenancy |
Per-tenant observability operator view: the registered-tenant count (cluster aggregate) and, dimensioned by tenant, the usage series (stored bytes, live keys, resident memory, owned trees), the quota ceilings and burst-headroom percentage, and the durable metered overage series (bytes / keys / memory / trees), so a burst or a sustained overage is attributable to a tenant. A templated tenant variable scopes every per-tenant panel to one tenant (a tenant's own view) or to all tenants (the platform-operator view); the registered-tenant count is a cluster aggregate and ignores it. Useful only when the tenancy package is registered on the silo. |
GrainIndex |
orleans.lattice |
Grain-index operator view: each index's backfill lifecycle state and percent complete, its processed-versus-total crawl progress, its live entry count, onboarding throughput split by route (activation versus backfill), projection-latency percentiles, and index-write failure rates by route. A templated index variable scopes every panel to one index or to all of them. Sources the shared core meter - the grain-index package publishes no meter of its own - so the series appear under an existing lattice subscription, but only once the grain-index package is registered on the silo. |
Quick Start
Install the package:
<PackageReference Include="Orleans.Lattice.Dashboards" Version="<X.Y.Z>" />
Wire up the meters the core dashboards read. AddMeter matches a meter name exactly and does not cascade, so each add-on dashboard (ReplicationGrpc, Authorization, Backup, Scaling, Tenancy) also needs its own meter registered by name, as do the Overview dashboard's exact-KNN panels (Orleans.Lattice.Api.Mcp.RepoContext) - see Configuration:
builder.Services.AddOpenTelemetry()
.WithMetrics(b => b
.AddMeter("orleans.lattice")
.AddMeter("orleans.lattice.replication") // omit if no replication
.AddMeter("Microsoft.Orleans") // Orleans runtime: activations, activation latency, directory
.AddMeter("System.Runtime") // .NET runtime: GC heap, working set, thread pool
.AddPrometheusExporter());
The two runtime meters back no bundled panel, but without them the endpoint carries no heap, process-memory, or activation-latency series at all - see Configuration for why they are registered and what they cost.
Retrieve the dashboard JSON by kind:
using Orleans.Lattice.Dashboards;
public static class DashboardJsonLoader
{
public static IReadOnlyDictionary<LatticeDashboardKind, string> LoadDashboards()
=> new Dictionary<LatticeDashboardKind, string>
{
[LatticeDashboardKind.Overview] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.Overview),
[LatticeDashboardKind.CommitPath] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.CommitPath),
[LatticeDashboardKind.Replication] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.Replication),
[LatticeDashboardKind.AtomicWrites] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.AtomicWrites),
[LatticeDashboardKind.MaterialisedViews] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.MaterialisedViews),
[LatticeDashboardKind.Authorization] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.Authorization),
[LatticeDashboardKind.Backup] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.Backup),
[LatticeDashboardKind.Scaling] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.Scaling),
[LatticeDashboardKind.ReplicationGrpc] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.ReplicationGrpc),
[LatticeDashboardKind.Tenancy] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.Tenancy),
[LatticeDashboardKind.GrainIndex] = LatticeDashboards.GetGrafanaDashboardJson(LatticeDashboardKind.GrainIndex),
};
}
LatticeDashboards.All enumerates every kind in declaration order, so a provisioning loop stays complete as new dashboards are added.
Either import each JSON via Grafana's Dashboards -> New -> Import UI, or write the strings to a provisioning directory referenced by Provisioning/dashboards.yaml.
What these dashboards assume about the exposition
The bundled panels are written against a host exporting through
.AddPrometheusExporter(), as in the Quick Start above. Two properties of that
exporter are load-bearing, and a scrape that lacks either renders the affected
panels empty rather than wrong, which is the harder failure to notice.
Unit suffixes are part of the series name. The exporter derives a suffix from
the instrument's declared unit and appends it to the family name - unless the
name already ends with it, so orleans.lattice.storage.wal.stored_bytes stays
orleans_lattice_storage_wal_stored_bytes_total - and so a histogram
declared unit: "ms" is exported as <name>_milliseconds, and one declared
unit: "s" as <name>_seconds. That is why the panels read, for example,
orleans_lattice_atomic_write_duration_milliseconds_bucket and not
orleans_lattice_atomic_write_duration_bucket. A panel that omits the suffix
names a series that exporter does not emit; it does not fall back to the
unsuffixed family.
An annotation unit such as {entry} is documentation rather than a dimension and
contributes no suffix. DashboardBucketUnitSuffixTests in test/lattice.dashboards
asserts every _bucket token a bundled panel reads against this rule, so a
bucket panel and its instrument cannot drift apart silently. Every other token
is held to the same rule by the dashboards drift guard, which resolves only the
exact family .AddPrometheusExporter() emits for an instrument's declared unit
and kind - see
How an instrument name becomes a PromQL series name.
Quantile panels need a real histogram exporter. Every bundled quantile panel
reads _bucket series through histogram_quantile, which only an exporter that
publishes bucket boundaries can satisfy.
Not every Lattice endpoint is such an exporter, and the difference is easy to miss
because both return 200. The repository-context container serves its own
hand-rolled Prometheus exposition built over a MeterListener, which reports each
measurement as a raw value and never surfaces the bucket boundaries an exporter
would need to publish. Having no honest bounds to publish, it renders every
Histogram<T> as a Prometheus summary carrying only _sum and _count, and it
appends no unit suffix at all. Against that endpoint every quantile panel here is
unavailable, not zero - and a panel carrying the common or vector(0) repair
displays a literal zero that is indistinguishable from a genuine sustained-zero
fault. Chart those series as rate(_sum) / rate(_count) instead. The missing unit
word empties a unit-suffixed panel there even when it is not a quantile panel: the
Overview dashboard's exact-KNN gather wall-time panel reads
repocontext_retrieval_exact_scan_duration_seconds_total, which the container -
the only host in this repository that publishes that instrument - exposes as
repocontext_retrieval_exact_scan_duration_total. The full account
is in the container guide.
Two template variables read scrape-side labels, not instrument tags. No
instrument emits a cluster tag: the cluster selector on the Overview,
CommitPath, Replication, AtomicWrites and MaterialisedViews dashboards
reads a label your scrape configuration attaches (for example a labels: entry
on a Prometheus static_configs target), and its All value is .*, so those
dashboards still populate when no such label exists. The silo selector on
Overview, CommitPath, Replication and AtomicWrites reads the instance
label Prometheus attaches to every scrape target. Every other selector (tree,
peer, scope, view, index, tenant) reads an instrument tag, and a
selector whose All value is .+ also excludes every series that lacks that
label. A panel therefore applies a selector only to instruments that carry its
label, so a selector does not narrow a panel over an instrument without it;
DashboardSelectorLabelPresenceTests checks every matcher and every variable
source against the tag sets in the metric-to-panel map.
Reference
For day-to-day use:
- API Reference - the public
LatticeDashboardsaccessor andLatticeDashboardKindkinds. - Configuration - meter registration, dashboard selection, and the Grafana provisioning templates.
- Metric-to-panel map - the per-instrument coverage table across every meter the dashboards chart.
For internals (the "how"):
- Architecture - embedded JSON resources, the bidirectional drift guard, and the provisioning-template layout.