Table of Contents

Security

This page documents Orleans.Lattice.Api.Mcp.Telemetry 9.9.0, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at security.md, and llms.txt lists every page.

The telemetry add-on sits on a dual-credential trust boundary: who may ask a telemetry question (MCP-side authorization) and how the proxy authenticates to the metrics backend (the backend credential) are two independent halves. Neither leaks into the other.

The two halves

flowchart LR
    agent[AI agent] -- "Lattice credential" --> mcp[MCP telemetry tool]
    mcp -- "LatticeOperation.Telemetry grant?" --> gate{authorized}
    gate -- "backend credential" --> backend[(Prometheus backend)]
  1. MCP-side authorization. A caller sees and can invoke the lattice_telemetry_* tools only if its effective authorization includes a cluster-wide LatticeOperation.Telemetry grant. Two checks enforce it. Discovery runs through the same permission-scoped filter the rest of the MCP surface uses, so an ungranted caller never sees the group. Every tool then re-checks the capability at call time through TelemetryAccessAuthorizer - the same cluster-wide seam the transport-neutral telemetry facade consults - before its range guardrails and metric-access checks and before the backend is called. A refused caller gets Success = false with the fixed Error "Reading cluster telemetry requires the Telemetry capability granted cluster-wide.", which is distinct from a metric-access denial and echoes neither the caller's subject id nor the gate's reason.
  2. Backend credential. The proxy stamps the configured backend credential (bearer, basic, mutual-TLS, or a rotating dynamic bearer token) on every backend request. Static credentials are supplied by the host in LatticeApiMcpTelemetryOptions.Credential; the dynamic-bearer mode instead resolves a token per request from a registered ITelemetryBackendTokenProvider. Either way the backend credential is entirely separate from any caller identity.

The caller's Lattice credential is never forwarded to the backend. The backend client's only collaborators are an HttpClient, the telemetry options, the backend-token seam, and a server-side logging sink; it holds no reference to any Lattice credential source, so there is no path by which the caller's identity could reach the backend.

Conversely, the backend credential is never exposed to the caller - including on the error path. A backend transport, timeout, or payload fault is reported to the caller with a fixed message that interpolates nothing from the caught exception. This matters because the proxy does not own every HttpMessageHandler in its own pipeline: a host may insert delegating handlers, and diagnostic or retry handlers commonly include request headers in their messages, so exception text is an uncontrolled channel that can carry the outbound Authorization value. The fault detail is not discarded - it is logged server-side by the backend proxy, which is the trusted side of the boundary - so an outage stays diagnosable without the credential crossing back.

The Telemetry capability

LatticeOperation.Telemetry is a cluster-wide capability, deliberately distinct from the data-plane operations:

  • It is not part of the data-plane aggregate and is conferred by no other operation, including Admin. A caller must be granted Telemetry explicitly.
  • It is granted over the all-trees sentinel scope, LatticeScope.ClusterWide(), because telemetry is a cluster-wide concern rather than a per-tree one. The grant is an ordinary Allow rule that the existing policy pipeline compiles and evaluates with no special-casing.
using Orleans.Lattice.Auth;

// Grant an automation agent read access to cluster telemetry, and nothing else.
var rule = new LatticeAuthorizationRule(
    "agent-telemetry",
    LatticeSubjectSelector.User("agent"),
    LatticeScope.ClusterWide(),
    LatticeOperation.Telemetry,
    LatticeEffect.Allow);

Because the capability is a distinct bit, a Telemetry grant never widens a caller's data-plane reach, and a data-plane grant never confers telemetry access. A Telemetry bit carried on a tree-scoped rule confers nothing either: discovery counts the capability only from a rule written at cluster-wide tree scope - a key- or prefix-scoped rule on the all-trees sentinel is not counted - and the call-time check authorizes over the cluster-wide sentinel, which a rule scoped to a real tree never matches.

Metric-access allow-list

Beyond the yes/no capability, the host can restrict which metrics a granted caller may read. LatticeApiMcpTelemetryOptions.MetricAccess selects the posture:

  • ReadAll (default) - any metric the backend exposes is readable.
  • DenyAllExceptAllowed - only the exact names and *-wildcard patterns in AllowedMetrics are readable; everything else is denied.

Matching is whole-name and case-sensitive. A wildcard entry is anchored at both ends, its * never matches a newline, and it is matched without backtracking, so a caller-supplied name - the metadata tool's metric argument, or a name lifted out of a query - cannot carry a trailing newline past an entry, and the check stays linear in the name's length.

The allow-list is enforced consistently across all four tools:

  • lattice_telemetry_query and lattice_telemetry_query_range extract the metric names referenced by the PromQL expression and reject the call if any referenced name is not admitted, if the expression names a metric through a matcher that cannot be reduced to an exact name, if it carries a label-only selector not anchored to a metric name (the right-hand side of up or {job="api"}, which would select series across every metric), or if no metric name can be extracted from it at all - before the backend is called.
  • lattice_telemetry_list_metrics filters the returned names to the admitted set.
  • lattice_telemetry_metric_metadata rejects a non-admitted named metric and, for an unnamed call, returns only admitted metrics.

The PromQL metric-name extraction is an allow-list gate, not a full PromQL parser, and it is fail-closed. It scans for identifiers in metric-name position and for the reserved __name__ label matcher inside a {...} label set, and gates every name it resolves against the allow-list - whether the metric is written directly (up) or through an exact __name__ matcher ({__name__="up"}, which contributes up as a referenced name). A selector that references a metric through a form that cannot be reduced to a fixed name under DenyAllExceptAllowed is rejected, not allowed through: a regex __name__=~"..." matcher, a negated __name__!="..." / __name__!~"..." matcher, and any malformed or unterminated __name__ matcher all fail the call closed, as does a top-level {...} label selector - terminated or not - that is neither anchored to a metric name nor pinned by an exact __name__ matcher. Function and aggregation calls, PromQL keywords and operators, grouping-modifier label lists, other {...} label names, quoted strings, and numeric or duration literals are skipped. A keyword is skipped only where Prometheus reads it as one: Prometheus also accepts the aggregation operators, and / or / unless, by, without, offset, start, and end as a bare metric name wherever an operand is expected, so up or min evaluates the metric min and the gate checks min against the allow-list. This closes the allow-list bypass where a caller named a denied series only through __name__; naming metrics directly in the expression remains the clearest way to write an admitted query.

Range guardrails

A range query is bounded so a single call cannot ask the backend for an unbounded scan: end - start may not exceed MaxRange (default 24h) and step may not exceed MaxStep (default 1h). An over-budget request is rejected with a clean Success = false result and never reaches the backend.

Fail-clean surfacing

Every fault path returns a structured result rather than throwing: a capability refusal, backend timeout, HTTP failure, non-success backend status, malformed payload, guardrail rejection, or metric-access denial arrives as Success = false with a human-readable Error. A guardrail rejection, a non-success backend status, and a metric-access denial each carry their own specific message, because callers act on that difference. Two paths report a fixed message instead: a capability refusal, because the underlying denial carries the caller's subject id and the gate's reason, and a backend transport or payload fault, because its free-text detail is the channel that could disclose the backend credential. The one deliberate exception is the metadata tool: a 404 from the backend metadata endpoint is treated as an absent metadata surface and degrades to Success = true with an empty Metrics list, not a failure. Only a genuine caller cancellation propagates as a cancellation. An agent therefore never sees a raw transport exception, and a denied metric is reported as a clear, actionable message.

Next

  • Tools - the four telemetry tools and their results.
  • Setup - configuring the backend credential and the metric-access allow-list.
  • MCP security - the fail-closed discovery model this capability plugs into.