Health probing
This page documents Orleans.Lattice.Api.Mcp.RepoContext, which is unreleased, in the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at health-probing.md, and llms.txt lists every page.Part of Container quickstart.
The runtime image is distroless and shell-less, so probing is HTTP-only - there is no shell-exec healthcheck:
GET /health/live- process and silo host alive (liveness).GET /health/ready- readiness (routing), and it is the conjunction of two independent components on the local durability profile, three on Azure: the lifecycle phase (silo joined, activation-time WAL replay done, durable stores reachable, MCP serving), the vector plane having demonstrated that semantic retrieval works, and - underDurabilityProfile.Azureonly - the scaling-signal health check, which the local profile never wires. A deployment with no embedder bound, and a host with no repository registered yet, both count as ready on the vector-plane component - there is no vector plane to wait for in the first case and nothing to serve in the second. This image always binds its embedding provider, though, so an unreachable companion is not the first case: it reportskeyword.vector_plane_unavailableand holds the component down once a repository is registered. A plane whose open an admission gate keeps refusing past the declared refusal bounds (LATTICE_REPOCONTEXT_ANN_OPEN_MAX_CONSECUTIVE_REFUSALS,LATTICE_REPOCONTEXT_ANN_OPEN_REFUSAL_TERMINAL_SECONDS) is equally not ready but needs capacity rather than time;repocontext_healthreports it asretrievalPhasesaturated_unavailablerather thanbuilding, while the/health/readycomponent body currently prints the same not-ready line for both.GET /health/backup- whether the durable agent-memory tree is actually being captured. It is served on its own path, and its component carries neither the liveness nor the readiness tag, deliberately: a failing backup must not restart the container or stop MCP traffic reaching it, because that would turn a durability fault into an availability outage and would stop the very traffic that makes the memory worth protecting. It is three-valued rather than two-valued, which is the whole point -Healthymeans a capture demonstrably happened (or that backup is switched off, which is stated in the body rather than implied by silence),Degradedmeans nothing is known to be protected yet, andUnhealthymeans an attempt failed. A container that has captured nothing reportsDegraded, neverHealthy: before issue #2640 no health component read the backup status at all, so a deployment on which every capture threw answered/health/liveand/health/readygreen and the only evidence was a log line. The response body carries the full positive statement - which tree, how many entries, when, and the last failure text - so a probe does not have to be followed by a log search. Before issue #2980 it carried none of that: the description was computed on every probe and then discarded at the HTTP boundary, exactly as/health/readydid before #2962, so the body was a single status word. Two layers below that,Describe()itself dropped both the last failure text and the sink inventory once any capture had succeeded, soUnhealthystated that a capture had failed without ever stating why, and a container whose local counters claimed successful captures while the sink held nothing suppressed that warning in exactly the case that warrants it.DisabledandProtectedare bothHealthyat 200, andDegradedanswers 200 as well: the endpoint keeps the framework's default status mapping, under which onlyUnhealthyis a 503. A probe reading only the status code therefore cannot tell a container with a verified backup from one that has captured nothing yet, or from one capturing nothing anywhere. The three-valued verdict separates nothing captured yet (Degraded) and an attempt failed (Unhealthy) from success, but it deliberately does not separate switched off from protected - only the body does that, and before #2980 the body could not. Alert on the body, not on the code.GET /health/silo- grain liveness: silo membership is active and the grain layer answered a trivial call, re-checked on every probe. Its component carries neither the liveness nor the readiness tag. It is the endpoint the image's exec-formHEALTHCHECK(--healthcheck) probes, sodocker pshealth reports it rather than readiness, and a still-starting (Degraded) silo answers 503 just as anUnhealthyone does.
Readiness is therefore not-ready during startup replay and during drain, but those are not the only causes: a box whose vector plane cannot serve reports 503 indefinitely while remaining alive and answering MCP calls. The response body names each component on its own line - <component>: <verdict>: <description> beneath the aggregate status word - so the component holding readiness down is readable straight from the probe (issue #2962). Before that it returned a bare Unhealthy, and every component description was computed on each probe and then discarded at the HTTP boundary, so the breakdown had to be reconstructed from four other surfaces. Naming the component is still not the same as naming the cause, so a sustained 503 must not be used by itself as a rollback signal. Narrow it with /health/live (200 means the process is fine), then an MCP call (an answer means the lifecycle component is satisfied), then a repocontext_search whose retrievalPath of keyword.vector_plane_unavailable corroborates the component the body already named, then docker compose ps to establish which side of the vector plane is at fault: an embedder that is missing, exited, or (unhealthy) is itself the cause and is directly actionable, whereas an embedder reporting (healthy) alongside a 503 rules the embedder out and places the fault host-side. Finally, /metrics separates a plane that has never been ready from one that was ready and lost it: repocontext_retrieval_ready_seconds_count is stamped once per process on the first transition into a ready phase (tagged with the phase it first reached), so its absence means the plane has never been ready in this container's lifetime, while repocontext_retrieval_unavailable_total counts fault episodes under a closed cause label: the three capability-loss retrievalPath values (keyword.vector_plane_unavailable, keyword.index_degraded, keyword.exact_fallback_suppressed) plus probe (a readiness probe rather than a real query saw the plane unable to serve), saturated (an admission gate refused the plane's open past its bound) and unknown. See Interpreting a persistent 503 for the same procedure written as a walkthrough.
Two properties of that state are worth stating because both are deliberate and both are easy to misread. Issuing a query by hand does not clear a persistent 503, and the host is already trying: a warmup service issues the same semantic query from application start and retries with backoff (waits of 2, 4, 8, 16 and 32 seconds, then every 30 seconds) until the plane answers or shutdown begins, so a persistent 503 is the warmup failing repeatedly rather than an absence of traffic. Once the plane is ready the service keeps supervising it for the life of the host, re-checking the readiness phase every 30 seconds and re-driving the warmup if readiness is revoked. A box with a repository registered but no vectors for it stays not-ready by design, because the search reports keyword.vector_plane_unavailable; a box with no repository registered reports ready, because there is nothing it could be asked to serve. And readiness lags a fault on purpose: once the plane has served, a fault must persist for a 30-second hold-down before readiness is revoked, and any successful retrieval inside that window clears the episode outright.
A 200 from /metrics is not evidence of a serving container
Issue #2868 recorded 43 minutes during which this container answered /metrics with HTTP 200 and a complete scrape while every MCP call needing authorization returned 500. The reserved sys-auth-policy tree had wedged, so tool discovery could not resolve a caller's effective permissions and failed with the MCP server's transient discovery-unavailable error, which it raises instead of advertising a falsely narrow tool set. The health signal and the service signal disagreed, and the reassuring one was the only one anything could read.
Three facts about that outage are worth separating, because two of them are properties worth keeping and only the third was the defect.
- Detection already worked.
/health/siloreportedHealth=unhealthyfor the entire window. The grain-liveness probe point-reads the reserved policy tree throughILatticeAuthorizationPolicyStore, so it exercises the exact seam that was wedged. Nothing needed to be added to make the check notice. /metricsis genuinely independent of the grain layer, not merely not exercising it. It is an in-process render of the collector's aggregated meter state and makes no grain call at all, which is why it stayed up. That independence is a feature: a scrape that failed during the wedge would delete the telemetry at the exact moment it is needed to explain the wedge. The endpoint therefore still answers 200 when the container is wedged, deliberately.- The verdict reached no consumer. It existed only in the container's health log. Prometheus, dashboards, alert rules, and any orchestrator probe pointed at the single exposed port all read the scrape, and the scrape carried no health series whatsoever. The correct answer was computed every fifteen seconds and nothing could read it.
The fix is therefore not a second endpoint, which is something somebody has to wire up and the deployment that most needs it is the one that will not. The verdict is put on the endpoint that is already scraped:
| Instrument | Meaning |
|---|---|
lattice_repocontext_health_status |
The current verdict, one series per component and status, carrying 1 on the component's current verdict and 0 on the other two. |
lattice_repocontext_health_evaluations_total |
Verdicts published per component since process start. |
lattice_repocontext_silo_probe_faults_total |
Grain-liveness probe failures attributed to a bounded cause. |
Alert on lattice_repocontext_health_status{component="silo",status="unhealthy"} == 1.
Read the status gauge against the evaluation counter, always. The gauge cannot express "unknown": a component that has never been evaluated carries zero on all three arms, which is byte-identical to a container that is healthy on two arms and zero on the third only because it is healthy. The counter disambiguates them. A zero there means no verdict has been published yet, so the status block is uninformative rather than green; a positive count beside an all-zero status block could only be an export defect. Without that denominator the change would reproduce, one level up, the exact false green it exists to remove.
A rising evaluation count is also positive proof that the authorization seam is being exercised, because the grain-liveness component reads the reserved policy tree on every evaluation. That matters because health components are evaluated on demand: before this, the seam was touched only when something probed /health/silo over HTTP, so a deployment whose liveness probe was pointed at /metrics never exercised it at all. A background publisher now runs the checks on a fixed cadence (first run at 15s, then every 30s), so the seam is exercised whether or not anything asks and the answer is on the scrape either way.
The fault counter attributes a failure to one of five causes, and the taxonomy is drawn where the remedies differ rather than where the exceptions do:
cause |
Meaning | Remedy |
|---|---|---|
probe-deadline |
The grain call neither returned nor threw within the probe's own deadline. This is the wedge shape of issue #2868. | Only a restart has been observed to clear it. |
grain-timeout |
A TimeoutException surfaced from the call itself, so a downstream deadline expired first. |
Investigate the shard named in the exception; often transient. |
access-denied |
The tree answered and refused the grant. | A grant defect on a healthy tree. A restart does nothing for it. |
drain-hung |
A graceful shutdown outran its stop-grace window. | The container is already trying to stop; see Graceful shutdown. |
unexpected |
A failure this taxonomy does not name. | Read the health body; a sustained non-zero here means the taxonomy needs an arm. |
All five arms, and the full component-by-status product, are published from process start, so a zero on any of these is a measurement rather than a series nobody has created yet. The cause dimension is bounded and every arm is primed, and the five arms are asserted to sum to an independently maintained total, so the per-cause breakdown accounts for every fault rather than for the subset something remembered to export.
Two readings are deliberately allowed to disagree, and knowing why saves a wasted investigation. The fault counter counts probe failures; the status gauge carries the verdict. During startup a failing probe is graded degraded rather than unhealthy, so a container joining slowly and one wedged from its first second are distinguishable - the first raises no faults, the second raises one per evaluation while both read degraded. Grading alone cannot tell those apart, which is why the cause is attached to the degraded arm too.
Recovery: assessed, and deliberately not automated
Self-recovery of a wedged authorization tree is not safely automatable from inside this process, and nothing here attempts it.
The only remedy observed to clear a probe-deadline wedge is a process restart, and a process cannot reliably restart itself while the fault it is reacting to is a hung grain call: the shutdown path drains the silo, the drain issues grain calls, and those are exactly the calls that are not answering. That is how a wedge becomes a drain-hung, which is strictly worse than the wedge because the container then stops serving the traffic it could still have served on paths that do not need authorization. Worse still, access-denied presents on the same health surface and a restart does nothing for it, so a blanket self-restart would loop a container whose configuration, not whose state, is wrong.
Acting on the verdict is therefore a decision for the deployment, and the mechanism is this: Docker does not restart a container on Health=unhealthy. The sample compose file sets restart: unless-stopped, which acts on process exit and never reads health, so an unhealthy container is left running indefinitely. Health-triggered restart is a Swarm feature (an unhealthy task is rescheduled) and a Kubernetes one (a failing livenessProbe restarts the container); in neither case is it restart:.
The recovery posture is an accepted one, not an unclosed gap
Issue #2906 settled this deliberately: no autoheal sidecar is added, and no healthcheck timing is changed. The reasoning is worth recording, because "add something that restarts it" is the intuitive answer and it is wrong here on the evidence this page already contains.
An automatic restart would have to fire on the health verdict, and the verdict does not distinguish the causes. access-denied presents on exactly the same surface as probe-deadline, and a restart does nothing for it, because what is wrong is the configuration and not the state. So a blanket auto-restart would fire on a fault it cannot fix, and the outcome is worse than doing nothing: a visible stuck container becomes an invisible restart loop. A stuck container at least holds its evidence - its logs, its metrics, and its docker inspect health log; a looping one discards that evidence every cycle and reads as ordinary churn.
What makes the acceptance defensible is that the verdict now leaves the container. Before issue #2905 it was computed correctly and reached nobody: /health/silo graded the container unhealthy for the full 43 minutes of the issue #2868 outage while /metrics answered 200 and every dashboard read green. Since #2905 the verdict is published onto the /metrics scrape, so the wedge is visible to the same scraper that already watches everything else, and an alert brings a human. That is the intended recovery path: detection is automatic, remediation is human.
The remedy itself is still a process restart, and the caveat above applies when performing it: a graceful drain issues grain calls, and under a wedge those are precisely the calls that are not answering, so prefer docker restart over an in-process shutdown.