Table of Contents

Configuration

This page is part of the documentation for Orleans.Lattice 9.9.0 (release line 9.9), built 2026-10-04. It is also published as markdown, with every table and list, at configuration.md, and llms.txt lists every page.

Compression has no core LatticeOptions knobs. The seam itself - the ILatticeCompressor contract, the registration helpers, the tag-space partitioning, and the shared-dictionary opt-in - is documented in compression.md; the knobs of its opt-in auto-trained dictionary are listed under Compression dictionary training options below. The per-consumer option keys live in their owning project's configuration doc: replication framing-tail compression in Orleans.Lattice.Replication configuration, and stored WAL payload compression in Orleans.Lattice.Storage.AzureTable configuration. The compression algorithm and Zstd level are safe to change after data already exists: stored payloads are self-describing and read back by their own per-row tag, so a level/algorithm change applies only to newly written data while existing rows decode unchanged.

Contents

Registering Lattice

Every silo must call AddLattice to register the grain storage provider that Lattice grains use internally:

siloBuilder.AddLattice((silo, name) => silo.AddMemoryGrainStorage(name));

The name parameter is the storage provider name ("lattice"). Replace AddMemoryGrainStorage with any Orleans storage provider (Azure Blob, Azure Table, ADO.NET, etc.) for durable grain state. The write-ahead log is a separate seam: AddLattice registers an in-memory WAL provider, so a durable deployment also registers a durable WAL provider - see WAL storage providers.

The grain storage provider must enforce ETags

The provider configureStorage registers must enforce ETags (optimistic concurrency) on write: a write that presents an ETag other than the stored row's current one must fail with Orleans' InconsistentStateException. Orleans' memory, Azure Table, Azure Blob, Cosmos DB and ADO.NET providers all do. Lattice depends on it in two ways:

  • Direct writes. The durable write-ahead-log pin store writes its bucketed slots straight through the provider rather than through IPersistentState, and when that grain cannot be reached during shutdown the leaf cursor reporter writes the same slots itself, from outside any grain. Both read a slot and write it back with the ETag they read, so only the provider's ETag check stops one writer from overwriting the other.
  • Duplicate activations. Orleans' default grain directory is eventually consistent, so during a membership change two activations of the same grain can briefly coexist. The provider's ETag check is what rejects the stale activation's write.

With a provider that accepts a stale write, a stale activation can persist an older leaf checkpoint while the durable write-ahead-log pin, which merges by maximum, keeps the newer offset. The write-ahead-log garbage collector can then trim entries a later cold rebuild of that leaf needs. Nothing fails at the time, so the loss is silent.

Each silo checks the requirement once as it becomes active. It writes a reserved probe row (grain type and state name _lattice_grain-storage-fencing-probe, one row shared by the cluster) twice, then writes it again presenting the first, now stale, ETag, and expects InconsistentStateException. What happens next is set by LatticeGrainStorageFencingOptions.Mode:

Mode Provider rejects the stale write Provider accepts the stale write Probe cannot decide
Warn (default) Logs the posture at Information. Logs a Warning; the silo starts. Logs a Warning; the silo starts.
Reject Logs the posture at Information. Logs an Error and fails silo start. Logs a Warning; the silo starts.
Disabled No probe runs; the silo logs once that the check is disabled.

The probe cannot decide when the provider faults, throws something other than InconsistentStateException on the stale write, keeps losing races to other silos' probes, is not registered, or does not finish within ProbeTimeout (30 seconds by default). Those are not evidence the provider is unsafe, so even Reject only warns on them rather than turning a transient storage fault into a failed start.

Use Reject to make the requirement fail closed:

siloBuilder.ConfigureLatticeGrainStorageFencing(o => o.Mode = LatticeGrainStorageFencingMode.Reject);

Warn is the default so that upgrading never turns a running deployment's next restart into a failed start. Use Disabled only for a deliberately non-durable provider, such as a benchmark's no-op storage, where the probe's writes are wasted.

Setting Options

Lattice uses the standard .NET named options pattern. Each tree resolves its options by name (the tree ID passed to GetGrain<ILattice>(treeId)).

Global defaults

Use ConfigureLattice without a tree name to set defaults that apply to every tree unless overridden:

siloBuilder.ConfigureLattice(o =>
{
    o.CacheTtl = TimeSpan.FromMilliseconds(100);
    o.TombstoneGracePeriod = TimeSpan.FromHours(6);
});

Per-tree overrides

Pass a tree name to override specific options for a single tree:

siloBuilder.ConfigureLattice("high-throughput-tree", o =>
{
    o.HotShardOpsPerSecondThreshold = 500;
    o.PrefetchKeysScan = true;
});

siloBuilder.ConfigureLattice("archive-tree", o =>
{
    o.TombstoneGracePeriod = Timeout.InfiniteTimeSpan; // disable compaction
});

Per-tree overrides are layered on top of the global defaults. Only the properties you set in the override are changed; everything else inherits from the global configuration. Configure actions run in registration order, so register the global ConfigureLattice call before the per-tree ones: a global call registered later overwrites every property it sets on every tree, per-tree overrides included.

Options are validated when an instance is first built. The global (unnamed) instance is built while the silo starts, so an invalid global value fails silo start. A per-tree instance is built the first time that tree's options are resolved, so an invalid per-tree value does not fail silo start: that tree's operations fail with an OptionsValidationException instead.

Not every option is read per tree, and not every option is re-read on every use. Some are read only from the global (unnamed) instance, so a per-tree override of them has no effect, and many are resolved once by the grain that uses them when it activates - a leaf, for example, resolves its tree's options when it activates and keeps them for the life of that activation, so a change reaches it on its next activation. Where it matters, an option's section below says which applies.

A per-tree override is matched to the tree id it names, exactly. A resize, a shadow-cutover restore or a schema remediation moves a tree's data behind an alias onto a physical copy that has an id of its own, and the shards, leaves and write-ahead-log partitions that hold that copy resolve their options under the copy's id rather than the tree's. An override registered for the tree's name therefore does not reach the options those parts read - MaxLeafBytes, LeafSnapshotSegmentBytes and WalMaxBatchEntries among them - which take the global values instead. The options the tree's own entry point reads on each call, such as the write bounds and the admission caps, still honour it.

Structural sizing is pinned per-tree in the registry, not in LatticeOptions. MaxLeafKeys, MaxInternalChildren, and ShardCount are seeded into the TreeRegistryEntry on first tree use from canonical defaults in LatticeConstants (128 / 128 / 64) and are mutable only through ILattice.ResizeAsync and ILattice.ReshardAsync. This prevents accidental divergence between the layout a tree was built with and a later configuration change. For capacity-planning guidance and per-provider limits see Tree Storage.

The virtual shard space is not a LatticeOptions property. A tree routes over 4096 virtual slots - a compile-time constant - unless its persisted ShardMap records a different slot count, which only a tree created by an installed app whose Orleans.Lattice.Apps manifest declares a virtualShardCount does. Slots are referenced by integer index, so changing a tree's slot count would re-route its keys. The virtual space is deliberately generous; the real ceiling on useful shard counts is scan fan-out and activation cost.

Cold-start and healing defaults

Lattice ships a set of mechanisms that bound cold-start cost and repair a tree that autonomic splitting has over-grown. All of them are on by default. That is deliberate: they exist to heal deployments that are already damaged, and a deployment that is already damaged is exactly the one nobody is going to reconfigure. A capability that has to be switched on heals nothing.

Each mechanism has its own escape hatch, so an operator who hits trouble disables one behaviour rather than rolling the image back. There is no single master switch, by design: a global flag would make the safe response to a problem in one mechanism the loss of all of them.

Mechanism What it does Kill switch What turning it off restores
Responsive WAL collection Collects a tree with a growing WAL every 30 seconds instead of hourly, relaxing back when it is quiet WalGcMinInterval = TimeSpan.Zero (or above WalGcInterval) The historical fixed-interval tick at WalGcInterval
Early first WAL pass Runs the first collection pass 15 to 30 seconds after silo start WalGcStartupDelay = WalGcInterval The historical [interval / 2, interval) deferral (30 to 60 minutes)
Binary leaf-snapshot codec Persists leaf snapshots as a compact frame instead of an object graph LeafSnapshotBinaryEncodingEnabled = false Captures persist the legacy row graph; reads stay dual either way
Shape-aware split admission Refuses to shatter a tree under a uniform bulk ingest, and caps autonomic growth HotShardMinSkewRatio = 1.0 and MaxPhysicalShardsPerTree = 0 and HotShardMinShardEntries = 0 Pure rate-based admission with no ceiling and no occupancy floor
Bounded leaf hydration Materialises snapshot rows on demand instead of decoding the whole blob at activation LeafPartialHydrationEnabled = false A full decode at every rehydrate
Automatic over-split healing Folds an over-split tree back towards its pinned base shard count ShardHealingEnabled = false An over-split tree stays over-split forever
Leaf-cache pre-warm Primes the hottest leaf caches at warm-up from a persisted access histogram LeafCachePreWarmCount = 0 No access tracking, nothing persisted, and a warm-up that primes nothing

The recovering replay flush ceiling is the exception with no switch. It is a defect repair in the leaf's deferred-offset ledger, not a behaviour to opt into, and there is nothing to configure.

Two things are worth knowing before you reach for any of these.

Reverting shape-aware split admission takes three settings, not one. The physical-shard ceiling and the occupancy floor are separate clauses that did not exist before the mechanism. An operator who zeroes only the skew ratio still cannot grow a tree past 256 shards and will conclude the revert did not work.

The snapshot codec switch is what you set before a rollback. Reads are always dual, so the switch is safe to flip in either direction on a running cluster. But a build that predates the frame has no dual-read and would see a frame-carrying blob as an empty row set. Set LeafSnapshotBinaryEncodingEnabled = false, leave it off until every leaf has captured at least once, and only then roll back.

Storage Provider Name

Lattice grains use the storage provider named "lattice" (exposed as LatticeOptions.StorageProviderName). The AddLattice extension method passes this name to your storage registration delegate. In advanced scenarios where you register storage directly, use this constant to ensure the provider name matches:

siloBuilder.AddMemoryGrainStorage(LatticeOptions.StorageProviderName);

Replication options

Cross-cluster replication is configured by the Orleans.Lattice.Replication package, not by LatticeOptions. The full options reference - including ReplicatedTrees, ReplicationPeers, ShipDoorbellEnabled, the backoff triple, and the maintenance cadence knobs - lives on LatticeReplicationOptions and is documented in Orleans.Lattice.Replication configuration; the shipper and maintenance drivers those knobs tune are described in Replication drivers. Peer membership in particular has its own resolution model (topology seam vs. ReplicationPeers projection) covered in Peer configuration.

AddLatticeReplication also supplies values for some options in this file: it mirrors the replication-side ReplogPartitions (onto WalPartitions), WalMaxBatchEntries, WalMaxBatchBytes, WalMaxPendingBatches, WalStorageProvider, and WalRetention onto the tree's LatticeOptions, but only when the replication-side value is set away from its own default and the core value is still at the core default, so a direct LatticeOptions override wins whenever it sets a value other than the core default (an override that sets the core default itself cannot be told apart from no override, and the replication-side value replaces it). See WAL and replog.

The replication receiver also consumes this file's WAL-saturation options indirectly: with AddLatticeReplication, a receiver translates its local WAL-saturation state (driven by WalSaturationThrottledRatio and the other WalSaturation* knobs above) into sender backoff hints by default. The receiver reads the signal under the tree name the batch carries, while the signal is sampled under the id of the write-ahead log the tree's writes land in, so on an aliased tree - after a resize, a shadow-cutover restore, a schema remediation or an operator-set alias, whose log belongs to the physical copy - the receiver never sees that saturation and sends no backoff hint (see Resolution and scope). That mapping is tuned with the separate WalSaturationReceiverFlowControlOptions, documented in Receiver flow control.

Tag-index reconciliation options

The background sweep that keeps each tag index consistent with the trees it covers is tuned on a separate options type, LatticeTagIndexReconciliationOptions, which each index resolves by its index name. Set a default for every index with ConfigureLatticeTagIndexReconciliation(configure) and a per-index override with ConfigureLatticeTagIndexReconciliation(indexName, configure). The default form applies through ConfigureAll, so, as with ConfigureLattice, register it before any per-index override that should win over it:

Option Type Default Meaning
Enabled bool true Whether the scheduled background sweep runs for the index; false unregisters its schedule. The on-demand reconcile call is unaffected.
Interval TimeSpan 1 hour Cadence between digest-gated sweeps. Must be positive (enforced by the options validator); a value below the 1-minute Orleans reminder minimum is clamped up to it.
ChunkSize int 16 Covered trees processed per phase-timer tick, bounding the work one tick performs. Must be at least 1 (enforced by the options validator).
ProbeOnly bool false Audit-only: the sweep probes and reports digest mismatches but never repairs them or advances the digest baseline.

See Background reconciliation for how the digest-gated sweep works.

Retry policy options

LatticeOptions.RetryPolicy accepts any ILatticeRetryPolicy; the shipped BoundedExponentialRetryPolicy is tuned by BoundedExponentialRetryPolicyOptions. That type is not resolved through the options pattern. AddLatticeRetryPolicy(configure) builds one policy from it at the moment of the call and assigns that single instance to RetryPolicy for every tree, and new BoundedExponentialRetryPolicy(options) builds one for you to assign yourself. The values are captured when the policy is built, so later changes to the options instance have no effect, and an out-of-range value throws ArgumentOutOfRangeException from that call rather than at first use.

Option Type Default Meaning
MaxAttempts int 4 Total attempts, including the first. Must be at least 1.
InitialDelay TimeSpan 50 ms Delay after the first failed attempt; each later delay doubles, up to MaxDelay. Must not be negative; TimeSpan.Zero retries without waiting.
MaxDelay TimeSpan 2 seconds Upper bound on any single delay. Must be at least InitialDelay.
RetryableExceptionClassifier Func<Exception, bool>? null (every exception is retried) When set, an exception the classifier rejects is rethrown immediately. A cancellation of the caller's token is never retried.

The policy wraps only the single-key and range mutations - SetAsync (with and without a TTL), SetIfVersionAsync, GetOrSetAsync, DeleteAsync, ApplyCrdtDeltaAsync, DeleteRangeAsync and DeleteRangeWherePredicateAsync; batch, atomic and bulk-load writes never enter it. It runs only for a mutation made inside a LatticeIdempotencyContext scope - without one there is no retry, whatever RetryPolicy holds - and when its budget is exhausted the original failure is rethrown with its stack trace. Because AddLatticeRetryPolicy assigns through ConfigureAll, a per-tree ConfigureLattice(treeName, o => o.RetryPolicy = ...) override wins only when it is registered after AddLatticeRetryPolicy (see Per-tree overrides). See Retry policy for the full model.

Compression dictionary training options

Compression has no LatticeOptions knobs (see the note at the top of this page), but its opt-in auto-trained shared dictionary is tuned by CompressionDictionaryTrainingOptions. Like the retry options, this type is not resolved through the options pattern: AddLatticeAutoTrainingCompressionDictionary(configure) on the service collection captures one silo-wide instance at the moment of the call, validates it then - throwing ArgumentOutOfRangeException for a value outside the ranges below - and registers a single provider built from it. The replication package's AddLatticeAutoSharedDictionary(configure) takes the same options and forces Enabled on.

Option Type Default Meaning
Enabled bool false Whether auto-training is active. While false the provider ignores observed payloads, trains nothing, resolves no dictionary id and emits no telemetry.
MaxSampleCount int 1024 Reservoir cap in samples; the oldest sample is evicted to admit a new one. Must be at least 1.
MaxReservoirBytes long 8 MiB Reservoir cap in total bytes; the oldest samples are evicted until the reservoir fits. Must be at least 1 and at least MaxSampleBytes.
MaxSampleBytes int 64 KiB Largest payload admitted to the reservoir; a larger one is ignored rather than truncated. Must be at least 1.
SamplingRate double 1.0 Probability that an observed payload is admitted. Must be in [0, 1]; NaN is rejected.
DictionaryCapacityBytes int 112 KiB Largest trained dictionary; also caps the footprint of every retained version. Must be at least 1.
MinSamplesToTrain int 100 Samples required before a training pass runs; a pass requested below it is skipped rather than throwing. Must be at least 1.
MinTrainingInterval TimeSpan 5 minutes Minimum interval since the previous training attempt; a pass requested inside it is skipped. TimeSpan.Zero disables the cadence gate. Must not be negative.
RetainedVersionCount int 4 Dictionary versions kept resolvable after a roll-over, including the current one. Must be at least 1.
FirstDictionaryId uint 1 Id of the first trained dictionary; later versions count up from it. Must not be 0, which is reserved for "no dictionary".

See Auto-trained dictionaries for how sampling, training and roll-over behave.

Per-call option types

WalMoveOptions is passed to each ILatticeAdmin.ExecuteWalMoveAsync call rather than registered as silo configuration; a null argument uses WalMoveOptions.Default.

Option Type Default Meaning
QuiesceLease TimeSpan 30 seconds How long the source partition stays fenced while the move copies its tail and flips the placement pin; if the move fails, the fence self-heals after this lease. A non-positive value uses the default.
CopyPageSize int 256 Entries copied per page from source to target. A non-positive value uses the default.
VerifyAfterCopy bool true in WalMoveOptions.Default Whether the move checks that the target tail matches the copied source range before flipping the pin. It is a plain bool, so new WalMoveOptions { ... } leaves it false; start from WalMoveOptions.Default with { ... } instead.
MaxConcurrentPartitionMoves int 1 Partitions a batch move copies in parallel; ignored by the single-partition overload. A non-positive value uses the default.

See Moving a partition to another account for the move itself.

Full Example

using Azure.Storage.Blobs;

var connectionString = "UseDevelopmentStorage=true";
var builder = WebApplication.CreateBuilder();

builder.UseOrleans(silo =>
{
    silo.UseLocalhostClustering();

    // Register Lattice with Azure Blob grain storage. The write-ahead log stays
    // in memory unless a durable WAL provider is registered as well - see
    // wal-storage-providers.md.
    silo.AddLattice((silo, name) =>
        silo.AddAzureBlobGrainStorage(name, options =>
        {
            options.BlobServiceClient = new BlobServiceClient(connectionString);
        }));

    // Global defaults
    silo.ConfigureLattice(o =>
    {
        o.KeysPageSize = 1024;
        o.TombstoneGracePeriod = TimeSpan.FromHours(12);
        o.SoftDeleteDuration = TimeSpan.FromHours(72);
    });

    // Per-tree: enable prefetch for a scan-heavy tree
    silo.ConfigureLattice("events", o =>
    {
        o.PrefetchKeysScan = true;
    });

    // Per-tree: disable compaction for an append-only tree
    silo.ConfigureLattice("audit-log", o =>
    {
        o.TombstoneGracePeriod = Timeout.InfiniteTimeSpan;
    });
});