Table of Contents

Multi-Site Manufacturing - an Orleans.Lattice sample

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at README.md, and llms.txt lists every page.

A working thin slice of a regulated process-engineering traceability system (turbine-blade lifecycle: forge → heat-treat → machining → NDT → MRB → FAI) that uses Orleans.Lattice as the fact store and convergent state layer for an inventory system running across two independent Orleans clusters.

It is not a scripted replay. Operators create parts, advance them through process stages, record inspections, raise non-conformances, issue MRB dispositions, and sign off FAI from a Blazor dashboard - and the system behaves like a minimal MES/QMS slice backed by Lattice.

See architecture.md for the structural view (topology, component graph, grain interdependencies, Lattice trees, replication sequence) and approach.md for the implementation rationale and gotchas. Unfamiliar term? Check the glossary.md.

Lattice capabilities demonstrated

Capability How it shows up
Ordered fact log per entity Every domain event (ProcessStepCompleted, InspectionRecorded, NonConformanceRaised, MrbDisposition, ReworkCompleted, FinalAcceptance) is an immutable key in the mfg-facts tree, keyed {serial}/{wallTicks:D20}/{counter:D10}/{factId} so a forward range scan yields HLC-ascending history.
HLC-ordered fold → convergent state ComplianceFold.Fold sorts facts by (WallClockTicks, Counter, FactId) before applying them, so concurrent producers across sites converge on the same ComplianceState. Contrasted live in the UI against a naïve arrival-order baseline running over the same fact stream.
Divergence visible under chaos Two backends (baseline, lattice) receive the same facts via a fan-out router. Chaos-induced reorder causes the arrival-order baseline to drift; the HLC-ordered lattice fold does not. Divergent rows surface in the dashboard organically - no scripted saga.
Folded materialised view + read-time join The dashboard summary is not a sample-owned read model. A library-maintained folded view (mfg-compliance, registered via AddLatticeViews/AddFoldedView) folds each part's mfg-facts in business-HLC order into an accumulator carrying the lattice compliance state, latest process stage, and fact count. The snapshot scans that view and joins each part's arrival-order BaselineState from the baseline backend at read time - the divergence between the two independently-maintained halves cannot be reproduced by any fold over mfg-facts, so it is joined per part rather than materialised. The library keeps the folded half current directly off the write-ahead log, so no application-side summary tree remains. See docs/lattice/materialised-views.md.
Tag-index secondary view mfg-site-activity keys facts part-major as {serial}/{site} and the built-in Orleans.Lattice tag index (opened through the injected ILatticeTagIndexFactory, membership tree tag-mfg-site) tags each key with its site. ListAtSiteAsync answers "parts at site X" via WithAnyTags(site) - the site is deliberately not a key prefix, so the tag index is the genuine access path. A worked example of the built-in tag index replacing a hand-rolled secondary-index tree.
Typed CRDT delta shipping mfg-part-labels is one OR-Set per serial, accessed through lattice.OrSet(serial) and replicated cross-cluster as LatticeMergeMode.OrSet - the package ships typed add / remove / merge deltas instead of raw byte writes. The companion mfg-part-operator tree is a per-serial LWW register kept cluster-local - see Per-tree replication policy below for the rationale.
Partition tolerance via shadow prefixes During a simulated intra-cluster partition, PartCrdtStore writes to a shadow key prefix; PartitionHealHostedService promotes shadows back onto the canonical keys on heal.
Range scans as primitives The per-part fact-history fold and the partition-heal sweep are plain half-open range scans over lex-ordered keys - no custom indexing layer. The per-site view instead uses the built-in tag index (see above).
Cross-cluster replication via the shipped package Orleans.Lattice.Replication provides the shipper, applier, and dead-letter handling over the core write-ahead log; Orleans.Lattice.Replication.Grpc provides the push transport. Each tree opts in with a LatticeMergeMode on the replicated-tree map (LwwRegister for mfg-facts and mfg-site-activity; OrFlag for its tag-mfg-site membership tree; OrSet for mfg-part-labels); see docs/lattice.replication/ for the wire format and bootstrap protocol.
Receiver-side applier decoration BaselineReplicationApplier decorates the package's IReplicationApplier singleton; on every cross-cluster apply it mirrors mfg-facts writes into the local naive BaselineFactBackend and raises FederationRouter.FactReplicated, so the side-by-side divergence visualisation and the dashboard activity feed both update without polling.
Durable operational state via Orleans grains Chaos configuration (IProcessSiteGrain, IBackendChaosGrain, IPartitionChaosGrain, IReplicationDisconnectGrain) persists to Azure Table Storage - restart the host and the system resumes exactly where it left off. The lattice tree write-ahead log persists to the same storage account via Orleans.Lattice.Storage.AzureTable (Azurite locally), so tree state survives silo restarts. Replication's per-peer shipping cursors persist to the same storage account.
Idempotent bulk-load on startup InventorySeeder emits 5 representative parts covering every ComplianceState the facts can reach (two Nominal, then one each of FlaggedForReview, Rework, and Scrap) through the same router operators use. A singleton IInventorySeedStateGrain gates the seed so re-running against the same storage account preserves inventory and operator mutations.
Coordinated multi-cluster restore Orleans.Lattice.Backup captures the replicated mfg-facts tree to a shared external sink that every cluster can read - under docker compose a dedicated Azurite blob account (azurite-backup) reachable from both clusters. The CoordinatedRestoreOperator facade restores it; because mfg-facts is a replicated tree, the backup package promotes the restore into an all-or-nothing coordinated saga across the participating clusters. See Coordinated multi-cluster restore below.

Per-tree replication policy

Each Lattice tree opts into the replication mode that matches the convergence semantics of the data it stores. The choice is per-tree, not global, because not every CRDT has meaningful cross-cluster behaviour - and demonstrating that explicitly is part of why the sample exists.

Tree Replication mode Rationale
mfg-facts LwwRegister Each fact key is {serial}/{wallTicks:D20}/{counter:D10}/{factId} - globally unique, so LWW per key never collides.
mfg-site-activity LwwRegister Part-major activity rows keyed {serial}/{site}; the newest fact for a part at a site wins per key, so LWW converges.
tag-mfg-site OrFlag Tag-index membership rows for the per-site view. Under active-active replication both clusters tag keys, so the index authors flag-CRDT (enable-wins) membership dots that converge without a single-writer assumption; an LWW membership tree would silently drop a posting written concurrently in the other cluster.
mfg-part-labels OrSet Process labels are an additive set; typed OR-Set deltas (add / remove / merge) reconcile concurrent writes from any silo or cluster without conflict.
mfg-part-operator not replicated (cluster-local by design) A bare LWW register over a key both clusters would write to - and HLCs from disjoint cluster ID spaces have no meaningful global order, so "last write wins" between US and EU is semantically arbitrary. The sample keeps the register cluster-local; a production system that genuinely needs cross-cluster operator handover would model it as an OR-Set of (replica, operator) tuples or as an explicit acquire/release token, and the part-detail UI labels the assign button "Assign (local-only)" so the choice is visible to the operator.

This is a design choice, not a limitation: opting mfg-part-operator out of replication is the right answer for a bare LWW register across disjoint HLC namespaces. Within a single cluster the register still converges across silos via the lattice's internal HLC, which is what the per-tree replication mode is selecting against.

Coordinated multi-cluster restore

The sample wires Orleans.Lattice.Backup so a replicated tree can be backed up and restored across every participating cluster as a single all-or-nothing operation. This closes the loop on the replication story: replication keeps clusters converged during normal operation; a coordinated restore rewinds every cluster to the same captured cut together, with no peer re-advancing past it and no reader observing a torn / partial restore.

Shared external sink. All clusters read and write one external sink instead of a per-cluster in-cluster sink, registered before AddLatticeBackup() so it wins over the package's in-cluster default. This is mandatory, not cosmetic: the backup package's startup guard rejects the in-cluster sink for a replicated tree, because a backup captured on one cluster must be resolvable from any peer. A manifest written by the cluster that captured is read back by any cluster that restores, purely through the shared sink.

The sink must be genuinely shared across every cluster. Under docker compose each cluster runs in its own container with its own Azurite, so a filesystem directory is not shared between them - a backup captured on us would be invisible to eu and a coordinated restore would abort. The sample therefore runs a dedicated Azurite account, azurite-backup, multi-homed onto both cluster networks, and points every silo at it through the Azure Blob sink (ConnectionStrings:BackupBlobStorage -> AddLatticeBackupAzureBlob, one shared msmfg-shared-backup container). When no shared blob account is configured (the legacy single-machine host-process path where all silos share one host; the in-memory quick-start registers no backup sink) the host falls back to FileSystemBackupSink (src/MultiSiteManufacturing.Host/Backup/), an ILatticeBackupSink backed by a shared filesystem directory (Backup:SharedSinkPath, defaulting to a shared temp directory).

Operator trigger. CoordinatedRestoreOperator (src/MultiSiteManufacturing.Host/Backup/) is a DI-registered facade - the same seam OperatorActions uses, drivable by the UI, a gRPC service, or a test. CaptureFactTreeAsync(name) captures the whole mfg-facts tree to the shared sink and returns a content-addressed backup id; RestoreFactTreeAsync(backupId) restores it with an atomic shadow cutover. The facade calls the backup package's public RestoreAsync entry point and holds no saga wiring of its own: because mfg-facts is declared replicated, the package's IRestoreSagaDispatcher (installed by AddLatticeReplication) promotes the restore into a coordinated multi-cluster saga automatically, decided by the target tree's current replication membership. Where the tree is not replicated - as in the sample's in-process tests, which construct the facade directly - the identical call runs as a plain local restore; the host itself registers the facade only on the replicated (two-cluster) path.

What the sample test proves. CoordinatedRestoreSampleTests (test/MultiSiteManufacturing.Tests/Backup/) stands up two in-memory clusters that share one sink directory and asserts the cross-cluster consistency guarantee a coordinated restore provides: a backup captured on one cluster restores identically onto both (all-or-nothing - every captured key lands, byte for byte, on every cluster), and a repeated restore converges to the same content without re-advancing the cut. The full saga transport (the coordinator and participant write fence that AddLatticeReplication activates in the host, and the gRPC control channel that AddLatticeReplicationGrpc registers) is covered end to end by the replication package's own coordinated-restore suites; the sample test focuses on the shared-sink-enabled cross-cluster property in process.

Fault-injection surface

The sample layers five tiers of fault injection, each modelling a distinct real-world failure class. The dashboard's chaos fly-out drives tiers 1 to 4b (tiers 4 and 4b through its Cluster split and Replication disconnect presets); tier 5 is driven from the Docker CLI:

Tier Models Toggle
1 Site unavailable / WAN latency IsPaused, DelayMs on IProcessSiteGrain
2 Per-backend storage jitter, transient failure, write amplification, ingress reordering IBackendChaosGrain per backend, read by the ChaosFactBackend decorator wrapping that backend
3 Cross-site out-of-order arrival (the site grain releases admitted facts four at a time, shuffled) ReorderEnabled on IProcessSiteGrain
4 Simulated intra-cluster silo partition IPartitionChaosGrain + router hash filter
4b App-level cross-cluster replication pause IReplicationDisconnectGrain
5 Genuine cross-cluster transport partition docker network disconnect against the peer Traefik

Running

The supported local topology is Docker Compose: three Azurite containers (one per cluster plus the shared azurite-backup account), four silos (two per cluster), and a Traefik proxy per cluster - host ports 5001 (US) and 5002 (EU).

./run.ps1

Once the stack is up, open the two dashboards side-by-side and the shared observability pane:

URL What it is
http://localhost:5001 US-cluster Blazor dashboard (sticky-cookie pinned to silo-us-a or silo-us-b).
http://localhost:5002 EU-cluster Blazor dashboard (sticky-cookie pinned to silo-eu-a or silo-eu-b).
http://localhost:3000 Grafana - anonymous Viewer access, or admin/admin for edit rights. Every dashboard Orleans.Lattice.Dashboards ships lives under Dashboards -> Orleans.Lattice (see Observability below).

See architecture.md for the full network and port layout and the Tier-5 partition commands.

Observability

The compose topology also includes a single Prometheus + Grafana pair giving cross-cluster visibility into both regions:

Service Host port Purpose
prometheus - Scrapes silo-{us,eu}-{a,b}:8080/metrics (multi-homed onto both cluster networks).
grafana 3000 Renders every dashboard shipped by Orleans.Lattice.Dashboards.

Open http://localhost:3000 (anonymous Viewer access - admin/admin if you want edit rights). Under Dashboards -> Orleans.Lattice you'll find every dashboard the package ships (catalogued in docs/lattice.dashboards/); the three this sample exercises most are:

  • Orleans.Lattice - Overview - throughput, leaf-write percentiles, cache hit-rate, splits, atomic-write outcomes.
  • Orleans.Lattice - Commit Path - WAL-first per-step commit latency (wal / apply / observer / digest), WAL append and writer admission, leaf activation replays.
  • Orleans.Lattice - Replication - ship/apply/lag percentiles, dead-letter churn, per-peer entries/bytes behind.

Note - every replicated tree in the sample ships through Orleans.Lattice.Replication over the Orleans.Lattice.Replication.Grpc push transport: mfg-facts and mfg-site-activity as LwwRegister, its tag-mfg-site membership tree as OrFlag (enable-wins flag-CRDT membership), mfg-part-labels as OrSet (typed CRDT delta shipping). The sample-specific seams on top of it are BaselineReplicationApplier, a decorator on the package's IReplicationApplier that mirrors cross-cluster mfg-facts writes into the divergence-visualisation backend; the Tier 4b chaos decorators ChaosReplicationTransport and ChaosReplicationApplier; and ReplicationActivityTracker, which bridges the replication meter into the per-peer ship/recv strip (see approach.md). See docs/lattice.replication/ for the wire format and bootstrap protocol.

The JSON for these dashboards is bind-mounted read-only from src/lattice.dashboards/Grafana/ - a CI test in the package keeps them in sync with the live meter instruments, and the OpenTelemetry.Exporter.Prometheus.AspNetCore exporter in src/MultiSiteManufacturing.Host/Program.cs exposes the orleans.lattice and orleans.lattice.replication meters at /metrics on each silo.

Exploring the cluster with Orleans.Lattice.Explorer

Each silo co-hosts the read-only Orleans.Lattice.Api.State gRPC surface on its dedicated :8081 h2c listener - the same one the cross-cluster replication service uses. Each cluster's Traefik exposes it through the existing published endpoint (5001 US, 5002 EU) via a non-sticky, round-robin, PathPrefix(/orleans.lattice.api.state/) router with an active health check, so the Orleans.Lattice.Explorer can browse the running cluster's trees, views, metrics, topology, and data with no new host ports. The health check probes each silo's :8080 HTTP port and evicts a stopped silo within a few seconds (the probe runs every 5s), so the explorer transparently fails over to the surviving silo instead of flickering. The sticky Blazor / router that pins each browser tab's SignalR circuit is untouched; the state-API router just has a higher-priority, more-specific prefix.

run-explorer.ps1 launches the Blazor web explorer (src/lattice.explorer/Web/Orleans.Lattice.Explorer.WebHost.csproj) pointed at a cluster. It seeds the endpoint through the explorer's launcher-friendly environment bootstrap, so nothing in your per-user explorer config is hand-edited. The script also exports any -Username / -Password it is given, but the web explorer honours the endpoint seed alone (its environment credential seed is off by default, and this head does not turn it on), so you sign in from the explorer itself.

Anonymous (default)

./run.ps1                 # state-API authorization is OFF by default
./run-explorer.ps1        # Blazor web explorer -> US cluster (http://localhost:5001)
./run-explorer.ps1 -Cluster eu

The explorer connects anonymously over loopback h2c (insecure-loopback-dev transport mode). Open the printed http://localhost:5290 once the web head starts.

With state-API authentication

Supply a username and password to run.ps1. It generates the salted PBKDF2 hash with tools/New-LatticeStateCredential.ps1 and delivers it to every silo container as LATTICE_STATE_USER_<username> through a git-ignored .env file; the plaintext password never reaches a container env, a command line, or the compose file. The host then enables RequireAuthorization = true with the reference EnvVarCredentialAuthorizer, so an anonymous explorer is rejected and a signed-in one succeeds: sign in at the explorer's sign-in dialog with the same username and password.

./run.ps1 -Username alice -Password 'Sup3rSecret'
./run-explorer.ps1        # then sign in as alice in the explorer

./run.ps1 -Down deletes the generated .env; every run that brings the stack up (-Clean included, which deletes it first) rewrites it from the switches passed to that run, so a credential from an earlier run never lingers.

Backup and restore from the Explorer

./run.ps1 -Backup also registers the backup control API and its gRPC binding on every silo, served through each Traefik's /orleans.lattice.api.backup/ router, so the Explorer's backup UI can capture and restore against the stack. The binding runs with authorization off - a demo-grade posture - and the switch combines with -Username / -Password. On the replicated stack the host also tightens the backup-health monitor to re-verify every catalogued backup against the shared sink every 5 minutes (ConfigureLatticeBackupHealth; the default is six hours), so the Explorer's backup health column reflects sink faults quickly.

Inspecting change history

The sample enables a durable change-history view (with full-value retention) over two CRDT trees on startup and then seeds a multi-revision timeline into them, so the Explorer's per-key History view has something non-trivial - and durable - to show out of the box (see docs/lattice/change-history.md):

  • mfg-part-operator (last-writer-wins register) gets a sequence of operator handoffs on one part's key, so the History view renders successive values plus diffs.
  • mfg-part-labels (process-label OR-Set) gets interleaved label adds and removes on the same part's key, so the History view renders element-level member changes.

Both are seeded for part HPT-BLD-S1-2028-00002. The History view is the History tab of a tree's workspace in the Explorer's Data area, and an entry's detail panel links to it for that key. To see it:

  1. Start the cluster and explorer: ./run.ps1 then ./run-explorer.ps1.
  2. In the explorer, open tree mfg-part-operator (or mfg-part-labels) in the Data area, select key HPT-BLD-S1-2028-00002 on the Keys tab, and follow History of this key in its entry panel.
  3. The timeline follows live changes by itself once it has loaded. On that part's detail page in the us cluster's sample UI (the cluster ./run-explorer.ps1 opens by default), add a process label (for mfg-part-labels) or assign an operator (for mfg-part-operator, which is cluster-local) and watch the new revision appear at the top of the timeline; Newest first is on by default. A live revision shows its kind, time and origin first and gains its value diff or member changes once the durable view records it.

The durable view is enabled by a small startup activator (HistoryShowcaseActivator) that sets a value-retaining retention mode and creates a history view on each tree; the change-history doc explains the retention modes and the truncation caveats.

Project layout

samples/MultiSiteManufacturing/
|-- README.md                         (this document - capabilities)
|-- architecture.md                   (topology, components, grains, trees, replication)
|-- approach.md                       (rationale, semantics, gotchas)
|-- glossary.md                       (domain + technical terms)
|-- MultiSiteManufacturing.slnx       (Contracts + Host + Tests)
|-- Dockerfile                        (the one silo image all four silos run)
|-- docker-compose.yml                (two-cluster topology + Prometheus/Grafana)
|-- run.ps1                           (docker compose wrapper)
|-- run-explorer.ps1                  (launches Orleans.Lattice.Explorer at a cluster)
|-- observability/                    (Prometheus scrape config + Grafana provisioning)
|-- traefik/                          (per-cluster routing: us.yml, eu.yml)
|-- src/
|   |-- MultiSiteManufacturing.Contracts/   (gRPC .proto surface)
|   `-- MultiSiteManufacturing.Host/        (ASP.NET Core + Orleans + Blazor)
|-- test/
|   `-- MultiSiteManufacturing.Tests/       (NUnit)
`-- tools/
    `-- SeedParts/                          (dev aid that bulk-inserts synthetic parts; not in the slnx)

Scope

The sample is deliberately narrow: one product family (HPT blade), a five-state severity lattice, no operator sign-in (the only authentication is the replication shared secret and the optional state-API credential above), no grpc-web, no Kubernetes manifests, no operator CLI (tools/SeedParts is a dev aid only). It exists to exercise Orleans.Lattice under realistic ordering and partition scenarios - not to be a production MES.