---
title: "Surfaces"
url: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.api.state/surfaces.html"
source: "https://github.com/NSTA1/Orleans.Lattice/blob/release/9.9/docs/lattice.api.state/surfaces.md"
package: "Orleans.Lattice.Api.State"
version: "9.9.0"
documents: "Orleans.Lattice 9.9.0 (release line 9.9)"
built: "2026-10-04"
all-pages: "https://nsta1.github.io/Orleans.Lattice/llms.txt"
bundle: "https://nsta1.github.io/Orleans.Lattice/docs/lattice.api.state/llms-full.txt"
---
# Surfaces

Part of the [Api.State documentation](README.md).

The state API exposes read-only surfaces: discovery, structure, entry inspection, change history, dead letters, change observation, metrics, and cluster info. The gRPC binding exposes RPCs for the remotely supported read surfaces; facade-only summary helpers such as `GetTreeSummaryAsync`, `GetShardSummariesAsync`, and `GetPhysicalShardCountAsync` are available only to in-process consumers. The examples below drive the remote surface through `LatticeStateApiGrpcClient`; the facade DTOs flow directly or inside gRPC wrapper records when you consume the facade remotely.

Every example assumes a client built like this (see [Client](client.md)):

```csharp verify
using Grpc.Net.Client;
using Microsoft.Extensions.DependencyInjection;
using Orleans.Serialization;

var serializerProvider = new ServiceCollection().AddSerializer().BuildServiceProvider();
using var channel = GrpcChannel.ForAddress("https://cluster.example:5001");
var stateClient = LatticeStateApiGrpcClient.Create(channel.CreateCallInvoker(), serializerProvider);
```

## Discovery

`ListTreesAsync` returns a deterministic, paged catalog of the registered trees. `ListViewsAsync` does the same for materialised views. Paging is driven by `CatalogRequest.PageSize` and the opaque `PageToken` carried forward from the previous page.

```csharp verify
using Grpc.Net.Client;
using Microsoft.Extensions.DependencyInjection;
using Orleans.Serialization;

var serializerProvider = new ServiceCollection().AddSerializer().BuildServiceProvider();
using var channel = GrpcChannel.ForAddress("https://cluster.example:5001");
var stateClient = LatticeStateApiGrpcClient.Create(channel.CreateCallInvoker(), serializerProvider);

string? pageToken = null;
do
{
    var page = await stateClient.ListTreesAsync(
        new CatalogRequest { PageSize = 100, PageToken = pageToken, IncludeSystemTrees = false },
        cancellationToken);

    foreach (var tree in page.Entries)
    {
        Console.WriteLine($"{tree.TreeId}  shards={tree.ShardCount}  {tree.Lifecycle}");
    }

    pageToken = page.NextPageToken;
}
while (pageToken is not null);
```

Each `TreeCatalogEntry` carries the tree id, its shard count, its lifecycle, and its effective read-only `Config` (except that a catalog entry's `Config.VirtualShardCount` always reads the 4,096 default, even for a tree declared with a smaller slot space; the facade-only `GetTreeSummaryAsync` reports the observed value). `IsAlias` reports whether the logical id resolves to a different physical tree (set by a resize, a shadow-cutover restore, a schema-remediation cutover, or an explicit alias assignment), with `PhysicalTreeId` carrying the resolved id in that case and `null` when the logical id is its own physical id. It also carries `RestoreShadowOfTreeId`: when non-null, the entry is the shadow physical tree of a shadow-cutover restore, and the value is the logical tree the restore was performed for (the alias that now resolves to it). It is stamped from the registry when the shadow tree is created, so a client can classify and group restore shadows as a first-class fact rather than parsing the tree name; it is null for every ordinary tree and for the logical alias itself. Set `IncludeSystemTrees` to surface the trees the catalog hides by default - the internal `_lattice_` trees, the materialised-view backing trees, and the `sys-` system data trees a Lattice add-on owns (on a `ListViewsAsync` request it likewise surfaces the `sys-` system views) - and `IncludeViewStats` on a `ListViewsAsync` request to populate each view's `Lag` (source write-ahead-log entries committed but not yet applied) and `EntryCount` - both `null` when not sampled. Each `ViewStateSummary` carries the view name, its `SourceTreeId`, two classification flags - `IsAggregation` (a grouped-reduce view) and `IsHistory` (a change-history / accumulative view whose rows are serialized history records backing the source tree's history timeline rather than directly inspectable value data) - and, for a view created at runtime from a host-registered projection provider, its `ProjectionProviderKey` and `ProjectionVersion` (`null` for a startup-declared view). Its `LastDigest` is never populated and always reads `null`.

`ListTagIndexesAsync` lists the tag-index membership trees as their own category; set `SourceTreeId` to restrict it to the indexes that cover one tree. `ListTagValuesAsync` then enumerates the distinct tag values carried by a single index over its subject tree, in ascending ordinal order - pass both `SourceTreeId` (the subject tree) and `IndexName` (the index). Both are paged like the other catalogs, so a client can populate a tag picker without scanning the index tree itself.

Three index-wide read methods browse a multi-tree tag index without decoding any internal membership-tree naming convention. Each is scoped to a single `IndexName` and spans every tree the index covers:

- `ListCoveredTreesAsync(CatalogRequest { IndexName })` returns a `CoveredTreeCatalogPage` of the subject-tree ids the index covers, in ascending ordinal order.
- `ListIndexTagsAsync(CatalogRequest { IndexName })` returns a `TagValueCatalogPage` of the distinct tag values across the whole index (the index-wide union of `ListTagValuesAsync` over every covered tree).
- `ScanTagMembersAsync(TagMemberScanRequest { IndexName, Tag })` returns a `TagMemberScanPage` of the live `TagMember` rows (each a `{ TreeId, Key }` pair) carrying that tag, ordered by `(tree id, key)`. Membership rows whose primary key no longer exists (stale until the next reconcile) are filtered out, so the page reflects only live members. `PageSize` defaults to 100 (a value below 1 takes the default) and is clamped to 1000; resume with the returned `NextPageToken`.

`ListIndexTagsAsync` and `ScanTagMembersAsync` first check that the index exists: an `IndexName` with no registered tag index fails with `KeyNotFoundException` in-process rather than answering with an empty page (the gRPC binding does not map that exception, so it arrives as a generic `Internal`).

Because these three span many trees, they present no single target tree to a host authorizer - they authorize like the cluster-wide `ListTrees` / `ListViews` rather than the subject-tree-scoped `ListTagValues`.

## Structure

`GetTreeStructureAsync` returns the structural node graph of a tree as a `StructureResponse`. `Roots` holds one `NodeStateSummary` per shard root the tree's shard map routes keys to - including a shard an adaptive split added above the pinned shard count - unless `ShardIndex` names a single shard to read; each node reports its kind (leaf or internal), child fan-out, and `SubtreeKeyCount` - the live-key count under that subtree. Summing the roots' subtree counts gives the tree's total live-key count when the response carries every shard root. The walk expands each shard root's subtree before it moves to the next shard, so once the `MaxNodes` budget runs out the remaining shard roots are omitted entirely, with no per-node marker. `Truncated` is set whenever any part of the graph was cut short - by the node budget or the depth limit - so a `false` value means every shard root is present, while under `true` the sum is only a lower bound; request a single shard with `ShardIndex` to read one the budget left out. Within an expanded subtree, a node whose children were cut short carries `HasMoreChildren`, the cue to drill into it with `SubPathNodeId`. A node's `SplitInProgress` always reads `false`: the structure walk does not report in-flight splits. An unknown tree is part of the typed contract, not a fault: the response carries `Status = TreeNotFound` with empty `Roots` rather than raising a gRPC `NotFound`, matching the found/absent convention `GetEntryAsync` uses.

```csharp verify
using Grpc.Net.Client;
using Microsoft.Extensions.DependencyInjection;
using Orleans.Serialization;

var serializerProvider = new ServiceCollection().AddSerializer().BuildServiceProvider();
using var channel = GrpcChannel.ForAddress("https://cluster.example:5001");
var stateClient = LatticeStateApiGrpcClient.Create(channel.CreateCallInvoker(), serializerProvider);

var structure = await stateClient.GetTreeStructureAsync(
    new StructureRequest { TreeId = "factory-floor", DepthLimit = 3 },
    cancellationToken);

long liveKeys = 0;
foreach (var root in structure.Roots)
{
    liveKeys += root.SubtreeKeyCount;
}
Console.WriteLine($"shards={structure.Roots.Count}  liveKeys={liveKeys}");
```

The walk is bounded by `DepthLimit` (default 4, a negative value takes the default, capped at 64) and `MaxNodes` (default 1000, a value below 1 takes the default, capped at 100,000), and can be focused on a single shard (`ShardIndex`) or rooted at a sub-path node (`SubPathNodeId`) to drill in without materialising the whole graph.

## Entries

`ScanEntriesAsync` returns a key-ordered page of entries. The `Mode` selects the cursor isolation. The default `EntryScanMode.Snapshot` opens a **snapshot-isolated cursor**: every page reflects the tree state captured when the scan opened, isolated from concurrent writes, at the cost of an all-shard baseline capture at open. `EntryScanMode.Live` opens a baseline-free **live cursor** whose paging is keyed on the last yielded key - it never duplicates an already-returned key, but writes committed after the open can appear on later pages and a value reflects its state at read time. `EntryScanMode.LivePointInTime` is a live cursor that additionally pins the in-flight-saga decision view at open without the per-shard baseline capture. Prefer a live mode for casual browsing so a scan does not fan a baseline capture out to every shard root; the Explorer defaults its browse scan to `Live`. The scan can run forward or reverse, be bounded by `StartInclusive` / `EndExclusive`, carry a `ValuePreviewBudget` for inline value previews, and apply a server-side `Predicate`; `PageSize` and `ValuePreviewBudget` fall back to `LatticeApiStateOptions.DefaultScanPageSize` / `DefaultScanValuePreviewBytes` when unset and are clamped to `MaxScanPageSize` / `MaxScanValuePreviewBytes` ([Configuration](configuration.md)). The mode is fixed when the cursor opens, so it is ignored on a continuation request. Set `IndexName` together with `Tag` to restrict the scan to the keys of the tree that carry that tag in the index: a tag-filtered scan opens no cursor and pages on the last returned key, so `Mode`, the key range, `Reverse`, and `Predicate` do not apply to it. The response `Status` distinguishes outcomes that would otherwise be an indistinguishable empty page: an unknown tree returns `TreeNotFound`, and a tag-filtered scan (`IndexName` and `Tag` set) against a tag index that has never been materialised returns `IndexNotFound` - so a mistyped `IndexName` is not silently reported as a real-but-empty `Found` result. Both are typed statuses on a normal response, not a gRPC fault.

```csharp verify
using Grpc.Net.Client;
using Microsoft.Extensions.DependencyInjection;
using Orleans.Serialization;

var serializerProvider = new ServiceCollection().AddSerializer().BuildServiceProvider();
using var channel = GrpcChannel.ForAddress("https://cluster.example:5001");
var stateClient = LatticeStateApiGrpcClient.Create(channel.CreateCallInvoker(), serializerProvider);

string? continuation = null;
do
{
    var page = await stateClient.ScanEntriesAsync(
        new EntryScanRequest { TreeId = "factory-floor", PageSize = 200, ContinuationToken = continuation },
        cancellationToken);

    foreach (var entry in page.Entries)
    {
        Console.WriteLine($"{entry.Key}  ({entry.ValueLength} bytes)");
    }

    continuation = page.ContinuationToken;
}
while (continuation is not null);
```

Because a **snapshot-isolated** scan opens a per-shard baseline capture that is heavier than a single read, its open can be shed under saturation. When the target tree is WAL-saturated (writes are being back-pressured), the server sheds the snapshot open at admission rather than piling that capture onto shard roots that are already collapsing: `ScanEntriesAsync` then fails with gRPC status `ResourceExhausted` and no cursor is created. This is a transient, retryable back-pressure signal (mapped from the core `LatticeSaturatedException` and gated by [`LatticeOptions.ShedSnapshotOpensWhenSaturated`](../lattice/configuration/options-reference-2.md#shedsnapshotopenswhensaturated), on by default) - back off briefly and retry once the tree drains. The shed looks the saturation signal up under the tree id the scan addresses, while the signal is recorded per physical tree, so a tree whose name is aliased onto another physical copy (after an online resize or a shadow-cutover restore) is not shed even while that copy is saturated. A `Live` scan opens no baseline and is not subject to this shed, which is another reason to prefer it for casual browsing. The Explorer surfaces the shed as a plain "this table is very busy, try again" notice and leaves the rest of the connection usable. A snapshot open is also refused, with no cursor created, when the deepest shard's frozen baseline would exceed `LatticeOptions.MaxSnapshotReplayEntries` rows (default 10,000,000) - `LatticeSnapshotReplayBudgetExceededException` in-process; the gRPC binding does not map that refusal, so it arrives as a generic `Internal` - use a `Live` scan on such a tree.

A drained scan releases its cursor itself. A continuation token that names an unknown, drained, closed, or expired cursor fails the page with `ArgumentException` (gRPC `InvalidArgument`), and a client that abandons a multi-page scan before draining it should release the cursor promptly with `CancelScanAsync`, the best-effort, idempotent cleanup verb described in the [gRPC Contract](grpc-contract.md#service).

GetEntryAsync returns the full record for a single key as an EntryGetResponse. Its `Status` reports the outcome: `Found` (the `Entry` carries the full record), `KeyNotFound` (the tree exists but the key does not), or `TreeNotFound` (no such tree). A not-found outcome is part of the typed contract, not a fault - it returns the structured response with a null `Entry` rather than throwing - so a caller distinguishes an unknown tree from a missing key by status. An unauthorised reader sees the same typed not-found a genuine miss returns, so existence is never leaked.

Every `EntryRecord` carries a `CrdtShape` tag: the name of the entry's CRDT merge mode (for example `"OrSet"`) when the value is a typed CRDT, or `null` for an opaque last-writer-wins value. The shape is resolved per key: the merge mode the leaf recorded for that key wins, and a key with no recorded mode (a plain last-writer-wins key, or one written before the leaf stamped modes) falls back to the tree's declared merge mode, so a mixed-mode tree can report different shapes for different keys. A materialised-view tree resolves its shape by view kind instead (see below). The record also carries the per-key `MergeMode` the leaf recorded (`null` when it recorded none) and a `Raw` flag that is `true` whenever `ValuePreview` is the raw stored bytes rather than a decoded CRDT projection - an opaque value, or a typed CRDT whose value could not be decoded on this silo - so a consumer can tell a CRDT entry apart from opaque bytes without decoding the value.

A CRDT entry (`CrdtShape` is non-null) additionally carries `CurrentMembers`: the decoded **current, complete, live members** of the key's folded CRDT state, produced server-side via the registered shape decoder's value-level projection (`ICrdtProvenanceDecoder.DecodeCurrentValue`). This is a point-in-time snapshot of the materialised value, not a per-revision change timeline - that timeline lives on `GetEntryHistoryAsync`. It contains only the shape-specific live members presently in the value, such as set/map entries, register values, sequence nodes, counter totals, or flag state. Each member is a `CrdtMemberValue` (element bytes, the contributing replica id where one exists, and a shape-specific ordinal). `CurrentMembers` is populated on both the single-key `GetEntryAsync` detail and the `ScanEntriesAsync` entry list. An opaque last-writer-wins entry leaves `CurrentMembers` empty and is rendered from its raw value bytes unchanged; the field also degrades to empty when no decoder (or shape registry) is registered.

Materialised-view and tag-index trees reuse the same rendering. A predicate / key-preserving view stores its source tree's value verbatim, so it mirrors the **source** tree's CRDT shape and decodes the same live members; a tag-index membership tree declared as a flag renders its current boolean state. An aggregation view, a history (accumulative) view, and a default last-writer-wins tag index are not member CRDTs and stay opaque blobs without crashing. The Explorer's entry detail (opened from a tree's **Keys** tab in its **Data** area) uses `CurrentMembers` to render a CRDT key's (or CRDT view / tag entry's) current state instead of its raw serialized blob. A scan of an aggregation view (a grouped-reduce or custom-fold view) returns only its materialised group values: the view's internal accumulator / inverse / membership rows, kept under a reserved key prefix, are excluded from the scan exactly as they are from the canonical `ILatticeView` read surface.

## Change history

`GetEntryHistoryAsync` returns a continuation-paged page of a single key's **change-history timeline** as an `EntryHistoryResponse`. Each `EntryRevisionRecord` carries the revision's `Hlc` (the timeline order key), its `Kind` (set, delete, CRDT delta, or range tombstone), the authoring `OriginClusterId`, a `Category` (always `User`: the history substrate does not record whether a revision was a maintenance rewrite), the revision's merge `Mode`, and a value-or-metadata view bounded by the tree's retention mode: a size-bounded `ValuePreview` (plus full `ValueLength`) when values are retained, or a `ValueHash` and length only under metadata-only retention. A CRDT-delta revision carries its author delta as a size-bounded `Delta` preview instead. The revision's own key is `SourceKey`; on a range-tombstone marker that is the inclusive start of the swept range and `EndKey` carries its exclusive upper bound, so a consumer reads the swept interval off the record rather than inferring it (`EndKey` is `null` for every point revision). A CRDT revision whose bytes were retained in full also carries the decoded element-level `MemberChanges` (added / removed element, replica, and causal ordinal).

Each revision carries a per-row `Retention` descriptor (the mode applied when the row was written, and whether its value bytes were retained), so a consumer can detect a retention-config transition by diffing adjacent revisions of the same key. The top-level `Bound` reports how the timeline is bounded: `BoundedByAge` when sourced from the durable per-key history view (clean, age-bounded, never truncated), `Truncated` when the retained write-ahead-log window has lost its oldest revisions (with `EarliestAvailable` carrying the clock of the oldest entry still readable on the key's write-ahead-log partition - the trim floor, which need not be a revision of this key), or `WalWindowFallback` when no history view is enabled. The write-ahead-log fallback lists every retained log record that names the key, so it can include records that are not logical changes to it: the staged writes of an atomic batch whose saga later aborted (reported as set revisions) and the reap marks tombstone compaction logs (reported as delete revisions). `FromHlc` / `ToHlc` optionally bound the window (both inclusive), and `Limit` / `ValuePreviewBudget` fall back to `LatticeApiStateOptions.DefaultHistoryPageSize` / `DefaultHistoryValuePreviewBytes` when unset and are clamped to `MaxHistoryPageSize` / `MaxHistoryValuePreviewBytes`. Set `Reverse` to order revisions newest-first within each page; paging still advances from oldest to newest. A fresh read (no `ContinuationToken`) that returns no revisions reports `Status = KeyNotFound` unless the bound is `Truncated`: a truncated write-ahead-log window may have aged out the revisions of a key that still exists, so that case stays `Found` with an empty page and the `Truncated` bound. Under `BoundedByAge` or `WalWindowFallback`, an empty fresh read - including one whose HLC window excludes every revision - reports `KeyNotFound`.

```csharp verify
using Grpc.Net.Client;
using Microsoft.Extensions.DependencyInjection;
using Orleans.Serialization;

var serializerProvider = new ServiceCollection().AddSerializer().BuildServiceProvider();
using var channel = GrpcChannel.ForAddress("https://cluster.example:5001");
var stateClient = LatticeStateApiGrpcClient.Create(channel.CreateCallInvoker(), serializerProvider);

string? continuation = null;
do
{
    var page = await stateClient.GetEntryHistoryAsync(
        new EntryHistoryRequest { TreeId = "factory-floor", Key = "press-7", Limit = 100, ContinuationToken = continuation },
        cancellationToken);

    foreach (var revision in page.Revisions)
    {
        var retained = revision.Retention.ValueRetained ? $"{revision.ValueLength} bytes" : $"hash {revision.ValueHash}";
        Console.WriteLine($"{revision.Hlc}  {revision.Kind}  {retained}");
    }

    continuation = page.ContinuationToken;
}
while (continuation is not null);
```

## Dead letters

The state facade exposes two read-only strict-mode schema-enforcement dead-letter operations:

- `Task<int> GetDeadLetterCountAsync(string treeId, CancellationToken cancellationToken = default)` counts retained dead-letter entries for a tree.
- `Task<DeadLetterQueuePage> ListDeadLettersAsync(DeadLetterQueueRequest request, CancellationToken cancellationToken = default)` lists the queue in append (time) order.

The gRPC client exposes the corresponding RPCs:

- `Task<DeadLetterCountResponse> GetDeadLetterCountAsync(DeadLetterCountRequest request, CancellationToken cancellationToken = default)` over the `GetDeadLetterCount` unary RPC.
- `Task<DeadLetterQueuePage> ListDeadLettersAsync(DeadLetterQueueRequest request, CancellationToken cancellationToken = default)` over the `ListDeadLetters` unary RPC.

`DeadLetterQueueRequest` carries `TreeId`, `PageSize` (default `100`, clamped to `1000`), and `PageToken`; `DeadLetterQueuePage` returns `Entries` and `NextPageToken`. Each `DeadLetterEntryRecord` carries the diverted key, size-bounded value preview, full value length, validation reason, ingest source, UTC timestamp, and whether the preview was truncated. The listing never mutates tree data and never replays or requeues diverted items. A hidden tree, a cluster without schema enforcement, or an empty queue returns `0` or an empty page rather than disclosing more.

## Change observation

`ObserveChangesAsync` subscribes to a tree's live mutation stream - optionally narrowed to a key range with `StartInclusive` / `EndExclusive` - and yields a `StateChangeNotification` per mutation until the call is cancelled or the server ends the stream. Each notification carries the `TreeId`, the affected `Key`, the change `Kind` (set, delete, or range delete), the mutation's `Hlc` stamp, and its `Category` (a user-driven write or library maintenance). For a range delete, `Key` is the inclusive lower bound and `EndExclusiveKey` the exclusive upper bound of the deleted range; `EndExclusiveKey` is `null` for a single-key change. Set `IncludeMaintenance` on the request to also observe maintenance rewrites.

Every notification also carries `Position`, an opaque monotonic resume cursor. Supply the last successfully-processed value as the request's `ContinuationToken` when re-subscribing to resume immediately after that notification. Delivery is at-least-once, so a resumed stream may redeliver a notification that was already processed - apply changes idempotently, and do not parse the `Position` token. Notifications are ordered by write-ahead-log sequence within each WAL partition; with several partitions each drain cycle emits every partition's new entries in turn, so there is no global order across partitions - use each notification's `Hlc` for a stable client-side order. The feed tails the durable write-ahead log rather than buffering notifications, so a slow consumer simply reads more slowly; a resume `ContinuationToken` that has fallen outside the WAL retention window fails the subscription with gRPC `FailedPrecondition` (`LatticeStateCursorExpiredException` in-process) rather than silently skipping the gap - restart from the live tail. A token minted under a different WAL partition count is treated as expired too, a malformed token fails with `InvalidArgument` (`ArgumentException` in-process), and a subscription to an unknown, system, or unreadable tree fails with `NotFound` (`KeyNotFoundException` in-process). With auth-backed read visibility on, a subscriber whose read grant covers only part of the tree (a prefix grant) receives only the point changes to keys it may read, and never a range-delete notification. The subscription resolves the tree's physical copy once, when it opens, and tails that copy's log for its lifetime: if the tree's alias moves to another physical copy while it runs (an online resize or a shadow-cutover restore), the stream keeps reading the old copy and delivers nothing further, so re-subscribe from the live tail after such a change.

```csharp verify
using System.Threading;
using Grpc.Net.Client;
using Microsoft.Extensions.DependencyInjection;
using Orleans.Serialization;

var serializerProvider = new ServiceCollection().AddSerializer().BuildServiceProvider();
using var channel = GrpcChannel.ForAddress("https://cluster.example:5001");
var stateClient = LatticeStateApiGrpcClient.Create(channel.CreateCallInvoker(), serializerProvider);

using var cts = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken);
cts.CancelAfter(TimeSpan.FromSeconds(30));

await foreach (var change in stateClient.ObserveChangesAsync(
    new StateObserveRequest { TreeId = "factory-floor" }, cts.Token))
{
    Console.WriteLine($"{change.Kind} {change.Key}");
}
```

Because the feed is server-streamed, cancel the call (via the `CancellationToken`) to unsubscribe; the server tears the subscription down when the stream ends.

## Metrics

`GetMetricsSnapshotAsync` returns a one-shot `TreeMetricsSnapshot` for the requested trees - live keys and shard count per tree, with optional shard hotness (`IncludeShardHotness`) and view lag (`IncludeViewLag`). `ObserveMetricsAsync` subscribes to a live feed: the server emits the initial full snapshot, then **delta-coalesced** snapshots as the metrics move, on the cadence set by `SampleInterval` (per subscription; `null` takes the configured `LatticeApiStateOptions.MetricsSampleInterval`, and the one-shot poll ignores it). `TreeIds` names the trees to sample - `null` or empty samples every visible tree - and `IncludeSystemTrees` adds the materialised-view backing trees and the `sys-` system data trees (the internal `_lattice_` trees are never sampled). Each `TreeMetricsSnapshot` carries its `SampledAt` time and `IsInitial`, which is `true` on a subscription's first tick and on a one-shot poll, both of which carry every visible tree; a later tick carries only the trees whose aggregates changed, plus `RemovedTreeIds` naming trees that have disappeared since the previous tick. A `TreeMetrics` row's `Lifecycle` always reads `Active` - the sampler does not look it up - so take a tree's lifecycle from the `ListTreesAsync` catalog.

```csharp verify
using Grpc.Net.Client;
using Microsoft.Extensions.DependencyInjection;
using Orleans.Serialization;

var serializerProvider = new ServiceCollection().AddSerializer().BuildServiceProvider();
using var channel = GrpcChannel.ForAddress("https://cluster.example:5001");
var stateClient = LatticeStateApiGrpcClient.Create(channel.CreateCallInvoker(), serializerProvider);

var snapshot = await stateClient.GetMetricsSnapshotAsync(
    new TreeMetricsRequest { TreeIds = new[] { "factory-floor" }, IncludeShardHotness = true },
    cancellationToken);

foreach (var tree in snapshot.Trees)
{
    Console.WriteLine($"{tree.TreeId}  liveKeys={tree.LiveKeys}  shards={tree.ShardCount}");
}
```

Many subscribers requesting the same metrics share a single underlying sampling loop, and a cluster with no metric subscribers samples nothing - see [Efficiency](efficiency.md).

### Detail paused under saturation

Each per-tree snapshot is assembled from a single per-shard diagnostics walk that backs both the tile aggregates and the per-shard hotness rows. When a tree is reporting WAL saturation, the sampler deliberately **skips that walk** so the metrics surface never piles read load onto shard roots that are already contended by the write backlog. Such a snapshot sets `DetailPaused = true`: `ShardCount` (from a single fan-out-free routing read) and any requested view lag are still populated, but the live counts (`LiveKeys`, `Tombstones`, `MinDepth`/`MaxDepth`, `ShardsSplitting`) are reported as zero and `ShardHotness` is empty. This is a transient, best-effort state, not an error - the detail returns automatically on the next sample once the tree settles. A consumer should surface it as a "paused - tree is busy" note (the Explorer's tree metrics panel shows a paused note) rather than treating the zeros as real counts. The sampler looks the saturation signal up under the id it is sampling - the name the request supplies, or the catalog id - while the signal is recorded per physical tree, so a tree whose name resolves to a different physical copy (an aliased tree after an online resize or a shadow-cutover restore, or a tenant-local name a request supplies under an asserted tenant) is checked against the wrong id and is walked even while its own copy is saturated.

## Cluster info

`GetClusterInfoAsync` returns a single `ClusterInfo` record identifying the cluster the client is connected to - its Orleans `ClusterId` (the deployment's logical cluster identity) and `ServiceId` (stable across rolling deployments of the same logical service). Either field is empty when the host did not configure it. The request envelope (`ClusterInfoRequest`) carries no fields today; it exists so the RPC can grow additive projection options later without changing the method signature. A consumer such as the Explorer's **Cluster** page uses it to show which cluster it is looking at.

```csharp verify
using Grpc.Net.Client;
using Microsoft.Extensions.DependencyInjection;
using Orleans.Serialization;

var serializerProvider = new ServiceCollection().AddSerializer().BuildServiceProvider();
using var channel = GrpcChannel.ForAddress("https://cluster.example:5001");
var stateClient = LatticeStateApiGrpcClient.Create(channel.CreateCallInvoker(), serializerProvider);

var info = await stateClient.GetClusterInfoAsync(new ClusterInfoRequest(), cancellationToken);
Console.WriteLine($"cluster={info.ClusterId}  service={info.ServiceId}");
```
