Multi-Site Manufacturing - an Orleans.Lattice sample
This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.A working thin slice of a regulated process-engineering traceability
system (turbine-blade lifecycle: forge → heat-treat → machining → NDT
→ MRB → FAI) that uses Orleans.Lattice as the fact store and
convergent state layer for an inventory system running across two
independent Orleans clusters.
It is not a scripted replay. Operators create parts, advance them through process stages, record inspections, raise non-conformances, issue MRB dispositions, and sign off FAI from a Blazor dashboard - and the system behaves like a minimal MES/QMS slice backed by Lattice.
See
architecture.mdfor the structural view (topology, component graph, grain interdependencies, Lattice trees, replication sequence) andapproach.mdfor the implementation rationale and gotchas. Unfamiliar term? Check theglossary.md.
Lattice capabilities demonstrated
| Capability | How it shows up |
|---|---|
| Ordered fact log per entity | Every domain event (ProcessStepCompleted, InspectionRecorded, NonConformanceRaised, MrbDisposition, ReworkCompleted, FinalAcceptance) is an immutable key in the mfg-facts tree, keyed {serial}/{wallTicks:D20}/{counter:D10}/{factId} so a forward range scan yields HLC-ascending history. |
| HLC-ordered fold → convergent state | ComplianceFold.Fold sorts facts by (WallClockTicks, Counter, FactId) before applying them, so concurrent producers across sites converge on the same ComplianceState. Contrasted live in the UI against a naïve arrival-order baseline running over the same fact stream. |
| Divergence visible under chaos | Two backends (baseline, lattice) receive the same facts via a fan-out router. Chaos-induced reorder causes the arrival-order baseline to drift; the HLC-ordered lattice fold does not. Divergent rows surface in the dashboard organically - no scripted saga. |
| Folded materialised view + read-time join | The dashboard summary is not a sample-owned read model. A library-maintained folded view (mfg-compliance, registered via AddLatticeViews/AddFoldedView) folds each part's mfg-facts in business-HLC order into an accumulator carrying the lattice compliance state, latest process stage, and fact count. The snapshot scans that view and joins each part's arrival-order BaselineState from the baseline backend at read time - the divergence between the two independently-maintained halves cannot be reproduced by any fold over mfg-facts, so it is joined per part rather than materialised. The library keeps the folded half current directly off the write-ahead log, so no application-side summary tree remains. See docs/lattice/materialised-views.md. |
| Tag-index secondary view | mfg-site-activity keys facts part-major as {serial}/{site} and the built-in Orleans.Lattice tag index (opened through the injected ILatticeTagIndexFactory, membership tree tag-mfg-site) tags each key with its site. ListAtSiteAsync answers "parts at site X" via WithAnyTags(site) - the site is deliberately not a key prefix, so the tag index is the genuine access path. A worked example of the built-in tag index replacing a hand-rolled secondary-index tree. |
| Typed CRDT delta shipping | mfg-part-labels is one OR-Set per serial, accessed through lattice.OrSet(serial) and replicated cross-cluster as LatticeMergeMode.OrSet - the package ships typed add / remove / merge deltas instead of raw byte writes. The companion mfg-part-operator tree is a per-serial LWW register kept cluster-local - see Per-tree replication policy below for the rationale. |
| Partition tolerance via shadow prefixes | During a simulated intra-cluster partition, PartCrdtStore writes to a shadow key prefix; PartitionHealHostedService promotes shadows back onto the canonical keys on heal. |
| Range scans as primitives | The per-part fact-history fold and the partition-heal sweep are plain half-open range scans over lex-ordered keys - no custom indexing layer. The per-site view instead uses the built-in tag index (see above). |
| Cross-cluster replication via the shipped package | Orleans.Lattice.Replication provides the shipper, applier, and dead-letter handling over the core write-ahead log; Orleans.Lattice.Replication.Grpc provides the push transport. Each tree opts in with a LatticeMergeMode on the replicated-tree map (LwwRegister for mfg-facts and mfg-site-activity; OrFlag for its tag-mfg-site membership tree; OrSet for mfg-part-labels); see docs/lattice.replication/ for the wire format and bootstrap protocol. |
| Receiver-side applier decoration | BaselineReplicationApplier decorates the package's IReplicationApplier singleton; on every cross-cluster apply it mirrors mfg-facts writes into the local naive BaselineFactBackend and raises FederationRouter.FactReplicated, so the side-by-side divergence visualisation and the dashboard activity feed both update without polling. |
| Durable operational state via Orleans grains | Chaos configuration (IProcessSiteGrain, IBackendChaosGrain, IPartitionChaosGrain, IReplicationDisconnectGrain) persists to Azure Table Storage - restart the host and the system resumes exactly where it left off. The lattice tree write-ahead log persists to the same storage account via Orleans.Lattice.Storage.AzureTable (Azurite locally), so tree state survives silo restarts. Replication's per-peer shipping cursors persist to the same storage account. |
| Idempotent bulk-load on startup | InventorySeeder emits 5 representative parts covering every ComplianceState the facts can reach (two Nominal, then one each of FlaggedForReview, Rework, and Scrap) through the same router operators use. A singleton IInventorySeedStateGrain gates the seed so re-running against the same storage account preserves inventory and operator mutations. |
| Coordinated multi-cluster restore | Orleans.Lattice.Backup captures the replicated mfg-facts tree to a shared external sink that every cluster can read - under docker compose a dedicated Azurite blob account (azurite-backup) reachable from both clusters. The CoordinatedRestoreOperator facade restores it; because mfg-facts is a replicated tree, the backup package promotes the restore into an all-or-nothing coordinated saga across the participating clusters. See Coordinated multi-cluster restore below. |
Per-tree replication policy
Each Lattice tree opts into the replication mode that matches the convergence semantics of the data it stores. The choice is per-tree, not global, because not every CRDT has meaningful cross-cluster behaviour - and demonstrating that explicitly is part of why the sample exists.
| Tree | Replication mode | Rationale |
|---|---|---|
mfg-facts |
LwwRegister |
Each fact key is {serial}/{wallTicks:D20}/{counter:D10}/{factId} - globally unique, so LWW per key never collides. |
mfg-site-activity |
LwwRegister |
Part-major activity rows keyed {serial}/{site}; the newest fact for a part at a site wins per key, so LWW converges. |
tag-mfg-site |
OrFlag |
Tag-index membership rows for the per-site view. Under active-active replication both clusters tag keys, so the index authors flag-CRDT (enable-wins) membership dots that converge without a single-writer assumption; an LWW membership tree would silently drop a posting written concurrently in the other cluster. |
mfg-part-labels |
OrSet |
Process labels are an additive set; typed OR-Set deltas (add / remove / merge) reconcile concurrent writes from any silo or cluster without conflict. |
mfg-part-operator |
not replicated (cluster-local by design) | A bare LWW register over a key both clusters would write to - and HLCs from disjoint cluster ID spaces have no meaningful global order, so "last write wins" between US and EU is semantically arbitrary. The sample keeps the register cluster-local; a production system that genuinely needs cross-cluster operator handover would model it as an OR-Set of (replica, operator) tuples or as an explicit acquire/release token, and the part-detail UI labels the assign button "Assign (local-only)" so the choice is visible to the operator. |
This is a design choice, not a limitation: opting
mfg-part-operator out of replication is the right answer for a bare
LWW register across disjoint HLC namespaces. Within a single cluster
the register still converges across silos via the lattice's internal
HLC, which is what the per-tree replication mode is selecting against.
Coordinated multi-cluster restore
The sample wires Orleans.Lattice.Backup so a replicated tree can be
backed up and restored across every participating cluster as a single
all-or-nothing operation. This closes the loop on the replication story:
replication keeps clusters converged during normal operation; a
coordinated restore rewinds every cluster to the same captured cut
together, with no peer re-advancing past it and no reader observing a
torn / partial restore.
Shared external sink. All clusters read and write one external sink
instead of a per-cluster in-cluster sink, registered before
AddLatticeBackup() so it wins over the package's in-cluster default.
This is mandatory, not cosmetic: the backup package's startup guard
rejects the in-cluster sink for a replicated tree, because a backup
captured on one cluster must be resolvable from any peer. A manifest
written by the cluster that captured is read back by any cluster that
restores, purely through the shared sink.
The sink must be genuinely shared across every cluster. Under
docker compose each cluster runs in its own container with its own
Azurite, so a filesystem directory is not shared between them - a
backup captured on us would be invisible to eu and a coordinated
restore would abort. The sample therefore runs a dedicated Azurite
account, azurite-backup, multi-homed onto both cluster networks, and
points every silo at it through the Azure Blob sink
(ConnectionStrings:BackupBlobStorage ->
AddLatticeBackupAzureBlob, one shared msmfg-shared-backup
container). When no shared blob account is configured (the legacy
single-machine host-process path where all silos share one host; the
in-memory quick-start registers no backup sink) the
host falls back to FileSystemBackupSink
(src/MultiSiteManufacturing.Host/Backup/), an ILatticeBackupSink
backed by a shared filesystem directory (Backup:SharedSinkPath,
defaulting to a shared temp directory).
Operator trigger. CoordinatedRestoreOperator
(src/MultiSiteManufacturing.Host/Backup/) is a DI-registered facade -
the same seam OperatorActions uses, drivable by the UI, a gRPC service,
or a test. CaptureFactTreeAsync(name) captures the whole mfg-facts
tree to the shared sink and returns a content-addressed backup id;
RestoreFactTreeAsync(backupId) restores it with an atomic shadow
cutover. The facade calls the backup package's public RestoreAsync
entry point and holds no saga wiring of its own: because mfg-facts is
declared replicated, the package's IRestoreSagaDispatcher (installed by
AddLatticeReplication) promotes the restore into a coordinated
multi-cluster saga automatically, decided by the target tree's current
replication membership. Where the tree is not replicated - as in the
sample's in-process tests, which construct the facade directly - the
identical call runs as a plain local restore; the host itself registers
the facade only on the replicated (two-cluster) path.
What the sample test proves. CoordinatedRestoreSampleTests
(test/MultiSiteManufacturing.Tests/Backup/) stands up two in-memory
clusters that share one sink directory and asserts the cross-cluster
consistency guarantee a coordinated restore provides: a backup captured
on one cluster restores identically onto both (all-or-nothing - every
captured key lands, byte for byte, on every cluster), and a repeated
restore converges to the same content without re-advancing the cut. The
full saga transport (the coordinator and participant write fence that
AddLatticeReplication activates in the host, and the gRPC control
channel that AddLatticeReplicationGrpc registers) is
covered end to end by the replication package's own coordinated-restore
suites; the sample test focuses on the shared-sink-enabled cross-cluster
property in process.
Fault-injection surface
The sample layers five tiers of fault injection, each modelling a distinct real-world failure class. The dashboard's chaos fly-out drives tiers 1 to 4b (tiers 4 and 4b through its Cluster split and Replication disconnect presets); tier 5 is driven from the Docker CLI:
| Tier | Models | Toggle |
|---|---|---|
| 1 | Site unavailable / WAN latency | IsPaused, DelayMs on IProcessSiteGrain |
| 2 | Per-backend storage jitter, transient failure, write amplification, ingress reordering | IBackendChaosGrain per backend, read by the ChaosFactBackend decorator wrapping that backend |
| 3 | Cross-site out-of-order arrival (the site grain releases admitted facts four at a time, shuffled) | ReorderEnabled on IProcessSiteGrain |
| 4 | Simulated intra-cluster silo partition | IPartitionChaosGrain + router hash filter |
| 4b | App-level cross-cluster replication pause | IReplicationDisconnectGrain |
| 5 | Genuine cross-cluster transport partition | docker network disconnect against the peer Traefik |
Running
The supported local topology is Docker Compose: three Azurite containers
(one per cluster plus the shared azurite-backup account), four silos (two
per cluster), and a Traefik proxy per cluster - host ports 5001 (US) and
5002 (EU).
./run.ps1
Once the stack is up, open the two dashboards side-by-side and the shared observability pane:
| URL | What it is |
|---|---|
| http://localhost:5001 | US-cluster Blazor dashboard (sticky-cookie pinned to silo-us-a or silo-us-b). |
| http://localhost:5002 | EU-cluster Blazor dashboard (sticky-cookie pinned to silo-eu-a or silo-eu-b). |
| http://localhost:3000 | Grafana - anonymous Viewer access, or admin/admin for edit rights. Every dashboard Orleans.Lattice.Dashboards ships lives under Dashboards -> Orleans.Lattice (see Observability below). |
See architecture.md for the full network and
port layout and the Tier-5 partition commands.
Observability
The compose topology also includes a single Prometheus + Grafana pair giving cross-cluster visibility into both regions:
| Service | Host port | Purpose |
|---|---|---|
prometheus |
- | Scrapes silo-{us,eu}-{a,b}:8080/metrics (multi-homed onto both cluster networks). |
grafana |
3000 |
Renders every dashboard shipped by Orleans.Lattice.Dashboards. |
Open http://localhost:3000 (anonymous Viewer access - admin/admin if
you want edit rights). Under Dashboards -> Orleans.Lattice you'll find
every dashboard the package ships (catalogued in
docs/lattice.dashboards/); the
three this sample exercises most are:
- Orleans.Lattice - Overview - throughput, leaf-write percentiles, cache hit-rate, splits, atomic-write outcomes.
- Orleans.Lattice - Commit Path - WAL-first per-step commit latency
(
wal/apply/observer/digest), WAL append and writer admission, leaf activation replays. - Orleans.Lattice - Replication - ship/apply/lag percentiles, dead-letter churn, per-peer entries/bytes behind.
Note - every replicated tree in the sample ships through
Orleans.Lattice.Replicationover theOrleans.Lattice.Replication.Grpcpush transport:mfg-factsandmfg-site-activityasLwwRegister, itstag-mfg-sitemembership tree asOrFlag(enable-wins flag-CRDT membership),mfg-part-labelsasOrSet(typed CRDT delta shipping). The sample-specific seams on top of it areBaselineReplicationApplier, a decorator on the package'sIReplicationApplierthat mirrors cross-clustermfg-factswrites into the divergence-visualisation backend; the Tier 4b chaos decoratorsChaosReplicationTransportandChaosReplicationApplier; andReplicationActivityTracker, which bridges the replication meter into the per-peer ship/recv strip (seeapproach.md). Seedocs/lattice.replication/for the wire format and bootstrap protocol.
The JSON for these dashboards is bind-mounted read-only from
src/lattice.dashboards/Grafana/ - a CI test in the package keeps
them in sync with the live meter instruments, and the
OpenTelemetry.Exporter.Prometheus.AspNetCore exporter in
src/MultiSiteManufacturing.Host/Program.cs exposes the
orleans.lattice and orleans.lattice.replication meters at
/metrics on each silo.
Exploring the cluster with Orleans.Lattice.Explorer
Each silo co-hosts the read-only Orleans.Lattice.Api.State gRPC surface on
its dedicated :8081 h2c listener - the same one the cross-cluster replication
service uses. Each cluster's Traefik exposes it through the existing published
endpoint (5001 US, 5002 EU) via a non-sticky, round-robin,
PathPrefix(/orleans.lattice.api.state/) router with an active health check,
so the
Orleans.Lattice.Explorer can browse the running
cluster's trees, views, metrics, topology, and data with no new host ports. The
health check probes each silo's :8080 HTTP port and evicts a stopped silo
within a few seconds (the probe runs every 5s), so the explorer
transparently fails over to the surviving silo instead of flickering. The
sticky Blazor / router that pins each browser tab's
SignalR circuit is untouched; the state-API router just has a higher-priority,
more-specific prefix.
run-explorer.ps1 launches the Blazor web explorer
(src/lattice.explorer/Web/Orleans.Lattice.Explorer.WebHost.csproj) pointed at a
cluster. It seeds the endpoint through the explorer's launcher-friendly
environment bootstrap, so nothing in your per-user explorer config is
hand-edited. The script also exports any -Username / -Password it is given,
but the web explorer honours the endpoint seed alone (its environment credential
seed is off by default, and this head does not turn it on), so you sign in from
the explorer itself.
Anonymous (default)
./run.ps1 # state-API authorization is OFF by default
./run-explorer.ps1 # Blazor web explorer -> US cluster (http://localhost:5001)
./run-explorer.ps1 -Cluster eu
The explorer connects anonymously over loopback h2c (insecure-loopback-dev
transport mode). Open the printed http://localhost:5290 once the web head
starts.
With state-API authentication
Supply a username and password to run.ps1. It generates the salted PBKDF2 hash
with tools/New-LatticeStateCredential.ps1 and delivers it to every silo
container as LATTICE_STATE_USER_<username> through a git-ignored .env file;
the plaintext password never reaches a container env, a command line, or the
compose file. The host then enables RequireAuthorization = true with the
reference EnvVarCredentialAuthorizer, so an anonymous explorer is rejected and
a signed-in one succeeds: sign in at the explorer's sign-in dialog with the same
username and password.
./run.ps1 -Username alice -Password 'Sup3rSecret'
./run-explorer.ps1 # then sign in as alice in the explorer
./run.ps1 -Down deletes the generated .env; every run that brings the stack
up (-Clean included, which deletes it first) rewrites it from the switches
passed to that run, so a credential from an earlier run never lingers.
Backup and restore from the Explorer
./run.ps1 -Backup also registers the backup control API and its gRPC binding
on every silo, served through each Traefik's /orleans.lattice.api.backup/
router, so the Explorer's backup UI can capture and restore against the stack.
The binding runs with authorization off - a demo-grade posture - and the
switch combines with -Username / -Password. On the replicated stack the
host also tightens the backup-health monitor to re-verify every catalogued
backup against the shared sink every 5 minutes (ConfigureLatticeBackupHealth;
the default is six hours), so the Explorer's backup health column reflects
sink faults quickly.
Inspecting change history
The sample enables a durable change-history view (with full-value retention) over
two CRDT trees on startup and then seeds a multi-revision timeline into them, so the
Explorer's per-key History view has something non-trivial - and durable - to show
out of the box (see docs/lattice/change-history.md):
mfg-part-operator(last-writer-wins register) gets a sequence of operator handoffs on one part's key, so the History view renders successive values plus diffs.mfg-part-labels(process-label OR-Set) gets interleaved label adds and removes on the same part's key, so the History view renders element-level member changes.
Both are seeded for part HPT-BLD-S1-2028-00002. The History view is the History
tab of a tree's workspace in the Explorer's Data area, and an entry's detail panel
links to it for that key. To see it:
- Start the cluster and explorer:
./run.ps1then./run-explorer.ps1. - In the explorer, open tree
mfg-part-operator(ormfg-part-labels) in the Data area, select keyHPT-BLD-S1-2028-00002on the Keys tab, and follow History of this key in its entry panel. - The timeline follows live changes by itself once it has loaded. On that part's
detail page in the
uscluster's sample UI (the cluster./run-explorer.ps1opens by default), add a process label (formfg-part-labels) or assign an operator (formfg-part-operator, which is cluster-local) and watch the new revision appear at the top of the timeline; Newest first is on by default. A live revision shows its kind, time and origin first and gains its value diff or member changes once the durable view records it.
The durable view is enabled by a small startup activator (HistoryShowcaseActivator)
that sets a value-retaining retention mode and creates a history view on each tree;
the change-history doc explains the retention modes and the truncation caveats.
Project layout
samples/MultiSiteManufacturing/
|-- README.md (this document - capabilities)
|-- architecture.md (topology, components, grains, trees, replication)
|-- approach.md (rationale, semantics, gotchas)
|-- glossary.md (domain + technical terms)
|-- MultiSiteManufacturing.slnx (Contracts + Host + Tests)
|-- Dockerfile (the one silo image all four silos run)
|-- docker-compose.yml (two-cluster topology + Prometheus/Grafana)
|-- run.ps1 (docker compose wrapper)
|-- run-explorer.ps1 (launches Orleans.Lattice.Explorer at a cluster)
|-- observability/ (Prometheus scrape config + Grafana provisioning)
|-- traefik/ (per-cluster routing: us.yml, eu.yml)
|-- src/
| |-- MultiSiteManufacturing.Contracts/ (gRPC .proto surface)
| `-- MultiSiteManufacturing.Host/ (ASP.NET Core + Orleans + Blazor)
|-- test/
| `-- MultiSiteManufacturing.Tests/ (NUnit)
`-- tools/
`-- SeedParts/ (dev aid that bulk-inserts synthetic parts; not in the slnx)
Scope
The sample is deliberately narrow: one product family (HPT blade), a
five-state severity lattice, no operator sign-in (the only authentication
is the replication shared secret and the optional state-API credential
above), no grpc-web, no Kubernetes manifests, no operator CLI (tools/SeedParts
is a dev aid only). It exists to
exercise Orleans.Lattice under realistic ordering and partition
scenarios - not to be a production MES.