Expanse turns ordinary hardware — the design target is a decade-old laptop — into a node of a self-healing cluster: replicated storage, scheduled services, VIP failover and a live web console. Built on NixOS from proven parts, not hand-rolled distributed systems.
All clear — nothing is degraded, unreachable or failing.
Recent events
17:41:52block/blocks/default/webput
17:41:50volume/volumes/pgdataput
17:40:09node/nodes/n3/statusput
Nodes
Node
Role
State
Raft address
API address
n1Leader
voter
Online
10.0.0.11:7444
10.0.0.11:7443
n2
voter
Online
10.0.0.12:7444
10.0.0.12:7443
n3
voter
Online
10.0.0.13:7444
10.0.0.13:7443
tty1 · expanse-console.service
EXPANSE 1.1.9 2026-09-28 17:42:10 UTC ┌─ n1 ─────────────────────────────────────────────────────────────────┐│ Hostname n1 Uptime 3d 4h 12m ││ Node ID 7d2f4c1e-9b3a-4e61-a0c7-5f8e21d3b940 ││ Health HEALTHY│├─ Management ─────────────────────────────────────────────────────────┤│ eno1 10.0.0.11 Web UI https://10.0.0.11:8443││ exp0 10.42.1.1 │├─ Cluster ────────────────────────────────────────────────────────────┤│ Cluster expanse Role leader ││ Quorum 3/2 (3 nodes) │├─ Hardware ───────────────────────────────────────────────────────────┤│ CPU Intel(R) Core(TM) i5-3320M (4 threads) ││ Memory 5.9 GiB free of 7.7 GiB ││ Disks sda 500 GB ST500LM012, sdb 500 GB ST500LM012 ││ Mirror system md127 raid1 [UU]││ / 38 GiB free of 40 GiB │└──────────────────────────────────────────────────────────────────────┘Alt+F2 for a login shell (this screen is read-only)
Left: the live web console every node serves on :8443. Right: the read-only host console every installed node shows on tty1.Illustrative mock-ups rendered in HTML, not screenshots.
One bootable ISO installs a node. Nodes find each other, join with single-use tokens, and share one strongly consistent control plane. Everything below ships in 1.1.9 and is exercised by NixOS VM tests that kill nodes, partition networks and pull disks.
Install from a USB stick
A hybrid UEFI/BIOS ISO with a full-screen installer TUI or a one-file unattended mode. Declarative btrfs + LVM layout, an impermanent root that is wiped on every boot, and a two-disk mirror layout that boots from either disk alone.
Raft-replicated state over mTLS, linearizable reads by default, single-writer leases as the anti-split-brain primitive, witness nodes for 2+1 topologies, and a read-only degraded mode when a node cannot reach a leader.
DRBD 9 on LVM thin pools: create, resize, snapshot and restore, online repair and verified split-brain recovery. A volume starts on one node and grows to its replication target as nodes join, one fully synced replica at a time.
Declare a service, its replicas, resources and placement; the scheduler runs it as hardened systemd units. Rolling updates, anti-affinity, singleton and daemonset placement, cross-node logs, and explain when something will not place.
A WireGuard overlay between every node, lease-fenced VIPs that move with their service, L4 and L7 balancing, cluster DNS for <block>.<ns>.expanse.local, and a default-deny nftables firewall per node.
Every fleet-wide configuration change is a generation: diff any two, roll back behind a confirm. Upgrade one node at a time with no quorum loss; a switch watchdog rolls a node back automatically if a switch goes wrong.
Every node serves the UI on :8443: dashboard, cluster, nodes, health, blocks, volumes, generations and events, updating over Server-Sent Events. Password login plus optional OIDC single sign-on with a fail-closed allow-list.
A Prometheus /metrics endpoint on every node, a shipped alert-rule file and Grafana dashboard, and an in-UI health page that derives the same critical alerts with no external tooling at all.
A per-cluster CA with zero-downtime rotation, HMAC single-use join tokens, mTLS on every internal port, age-encrypted secrets, and a browser-facing UI CA that is separate from the cluster's own.
One installed node is already a usable cluster. Each step below is the real command, not a sketch.
1
Install a node
Write the ISO to a USB stick and boot the target. The installer picks disks, network and an SSH key, shows what it will destroy, then installs. Headless? SSH in and run it there.
$ nix build github:team-expanse/expanse/v1.1.9#iso$ sudo dd if=result/iso/*.iso of=/dev/sdX bs=4M# boot the target from the stick, then:$ expanse install --tui
2
Form a cluster
Bootstrap the first node as the sole voter. It prints a single-use join command; run it on each further node. Volumes and blocks gain replicas automatically as nodes join.
Open the web console on any node, or use the CLI: apply a block, create a volume, watch health. Everything the UI does, the CLI does too — it is one control plane with two ways in.
Writes are forwarded to the Raft leader, reads are linearizable by default, and every action the web console offers is a CLI verb underneath. When a block will not place, explain tells you why each node was rejected.
Declarative and idempotent.block apply is a no-op when nothing changed and a rolling update when something did.
Cross-node logs.block logs default/web --replica 1 streams from whichever node runs it.
Doctors for storage and network. PASS/WARN/FAIL checks with a remediation hint on every row.
root@n1 — ssh
$ expanse cluster statuscluster: expanse (5c0b3e2a-…)version: 1.1.9generation: 12leader: n1quorum: 3/2nodes: 3 ID ROLE STATE RAFT API n1 voter leader 10.0.0.11:7444 10.0.0.11:7443 n2 voter voter 10.0.0.12:7444 10.0.0.12:7443 n3 voter voter 10.0.0.13:7444 10.0.0.13:7443$ expanse ctl block apply -f web.yaml --dry-run# validated, nothing written$ expanse ctl block apply -f web.yamlcreated default/web$ expanse ctl block explain webweb: 3/3 placed (phase RUNNING)$ expanse ctl volume listID NAME SIZE STATE REPLICAS NODES (PRIMARY)vol-… pgdata 10Gi healthy 3/3 n1,n2,n3 (n1)$ expanse doctor storage --vg expanseCHECK STAT DETAILdrbd-module PASS loaded, version 9.2.14volume-group PASS expanse: 41.3% free of 943718400000 bytesthin-pools PASS 1 pool(s) healthy: poolsystem-mirror PASS md md127 RAID1 across 2 devicesdrbd-resources PASS 1 resource(s) healthy
110+
NixOS VM tests that kill nodes, partition networks and pull disks — not just unit tests.
0
Acknowledged writes lost across leader failover, partitions and rolling upgrades in those tests.
≤15s
VIP failover budget after a hard power-off of the holder, measured in VM tests from an external client (typically ~10 s).
~2.5%
Of one core for the idle Raft leader: three containers on one bare-metal host, no volumes; followers 1.1–1.2%.
Why Expanse
Built the boring way, on purpose.
The design rule is simple: adopt a proven component for the data plane, build only the control plane that ties them together. That keeps the whole system small enough to reason about — and to test with real failures.
Expanse is
A host OS and cluster manager in one image. There is nothing to install on top; the node boots into the agent, the web console and the tty1 host console.
Sized for the machines you have. 2 cores, 2 GB RAM and a 20 GB disk is the design target, and a single node is a valid cluster from day one.
Strongly consistent. One Raft log holds cluster state, leases, generations and even join-token consumption, so two nodes can never both believe they own something.
Reproducible. Every node is a NixOS configuration; an upgrade is a new generation with a boot entry to fall back to.
Tested with real failures. Every feature has VM-test coverage of node kills, partitions and restarts, plus a chaos suite with a split-brain invariant checker.
Expanse is not
Kubernetes. Blocks are systemd-supervised units described in YAML, not pods. There is no Kubernetes API, no container runtime and no operator ecosystem.
A hypervisor first. VMs are one block type (QEMU/KVM on a replicated disk). Failover is a cold restart elsewhere; there is no live migration.
A file server yet. SMB/NFS shares are paused: the Samba block deploys but its failover has an open bug. Use volumes, Postgres, iSCSI or VMs today.
TPM-sealed. Secrets are age-encrypted, keyed from the cluster secret. TPM sealing is a named, deferred feature, not a shipped one.
Validated on every box. Automated coverage is on QEMU/KVM; the laptop, mini-PC, server and ARM targets in the hardware matrix are still marked untested.
Concern
Adopted component
What Expanse adds
Host OS and upgrades
NixOS, systemd-boot
Generations, rollback, a switch watchdog, an impermanent root
Cluster state
hashicorp/raft, memberlist
Leases, generations, join protocol, mTLS CA with rotation
A scheduler with rolling updates, anti-affinity, singleton leases
Network
WireGuard, nftables, miekg/dns
Mesh reconciliation, lease-fenced VIPs, L4/L7 balancing, cluster DNS
Backup and metrics
restic, Prometheus, Grafana
Consistent snapshot paths, one-command restore, shipped rules and dashboard
How Expanse is made
A joint human–AI project.
Expanse's code, tests and documentation, and this website, were built by a human maintainer working with AI coding assistants. The AI wrote much of the code and prose. The maintainer directs the work and decides what ships.
It is in the history. Most commits in both repositories credit an AI co-author in a Co-Authored-By trailer. Read the commit log.
Claims come with evidence. What this site says is drawn from the project's tests, docs and changelog, and what has not been verified is marked as such. Real hardware, for example, is still untested.
Judge it on its record. The VM tests exercise failures, not just the happy path, but automated tests are no substitute for your own evaluation. Read the known issues before trusting Expanse with data you care about.
Boot a node tonight. Add the second one whenever.
The quickstart takes a blank machine to a one-node cluster with a web console, and shows how each further node joins with a single command.