TWINSTOR
Tech Preview · v0.5.0 · July 2026

TWINSTOR
Two hosts. Zero compromise. Full HA.

Hyperconverged storage for two XCP-ng hosts. TWINSTOR turns their local disks into fully redundant, self-healing shared storage: live migration, automatic VM restart, rolling updates. No SAN, no witness node, no third host.

Host A always active VMs Local disk Host B always active VMs Local disk synchronous replication every byte on both hosts, reads always local
Measured, not estimated

The numbers from our lab

Every figure below was measured on real hardware: power cables pulled, a real power-strip cut, disks removed live.

~3 minfrom a total site power cut to all VMs running again, zero human intervention
0dropped pings on the survivor when the other host loses power
22 sfrom OS boot to storage serving after a power outage
234automated tests on the decision logic, plus 14 functional scenarios on real hardware
200+randomized chaos-hammer iterations, each passing only on fully unaided recovery
~9,000lines of production Rust orchestrating battle-tested components (DRBD, LIO, multipath, XHA)
The problem

Two servers alone can't give you HA

Every small site faces the same dilemma: you need two servers for redundancy, but high availability needs shared storage, and every traditional answer adds cost and complexity.

  • Hardware SAN / NAS: a third box to buy, power, and maintain (and a new single point of failure)
  • VMware vSAN 2-node: per-core licensing plus a witness appliance that must run somewhere else
  • Other storage clusters: a third node or quorum device, and a steep learning curve
  • No shared storage: no live migration, no HA, manual disaster recovery
The solution

Shared storage from the disks you already have

Both hosts hold an identical, continuously synchronized copy of every byte. Both serve storage at all times: there is no passive node and no failover delay, because there is no failover.

  • Reads are local. Every VM reads from its own host's disk: zero network hop, full NVMe/SSD performance.
  • Writes are safe. Acknowledged only once stored on both hosts. Lose a server at any instant: not a single acknowledged write is lost.
  • No tiebreaker infrastructure. Where competitors need a witness, TWINSTOR uses the network gateway you already have.
  • Zero patches to XCP-ng. Works with any XCP-ng version; your support path stays standard.
One command

Everything, at a glance

twinstor status opens with a one-line health verdict and colors only the values that carry state — green while the pair is whole, amber or red the moment something needs you.

twinstor status

health:  ● healthy  — serving, peer redundant

     version:     0.5.0
     daemon:      running
     host:        twinstor-1 (this node), peer twinstor-2
     pool master: twinstor-1 (this node)
     pool HA:     enabled

   drbd
       disk:       UpToDate
       connection: Connected
       peer disk:  UpToDate

   readiness:   ready (local serving, peer redundant)
   iscsi:       active
   tiebreaker:  192.168.1.1 (gateway)
The data promise

Four commitments, in plain terms

Every failure behavior in TWINSTOR reduces to these four commitments. They are what actually protects your VMs when hardware and networks misbehave.

1

A write your VM was told is safe is never silently destroyed

If the two copies ever diverge, TWINSTOR only ever discards writes that no VM received a completion for. The losing side of a partition freezes instead of serving, so diverged history never accumulates.

2

When the evidence is ambiguous, it stops and asks

No verdict means no action: both copies are kept, an alert names the decision that is yours, and waiting costs nothing. It never silently picks a side it cannot justify.

3

Decisions are made on evidence, not a single glance

A winner verdict requires the pool's answer to be stable across a 45-second sampling window, and a host that is actively running your VMs is never auto-discarded, whatever the control plane claims.

4

Everything is accountable

Every incident is a numbered episode in the logs from open to close, every resolution names the VMs it touched, and every state that waits for a human is listed in one table in the FAQ.

Fire and forget

Every failure heals itself or fails closed

Designed to run unattended for years at sites with no IT staff. Every scenario below was exercised on real hardware; the numbers are measured, not estimated.

FailureWhat happensMeasured result
Host power lossSurvivor keeps serving, VMs auto-restart0 dropped pings; victim's VMs back in ~2 min
Total site power outageBoth hosts boot, storage converges, VMs restart~3 min from power cut to all VMs running, unattended
Local disk / RAID array failureVMs keep running, transparently served from the other host's copyWorkloads never missed a beat; full redundancy 23 s after the disk returned
Network cut between hostsIsolated node steps aside, survivor continuesAutomatic recovery and resync on rejoin
Software crashStorage keeps serving while the service restarts itself0 s VM downtime
Split-brain (both sides writing)Fails closed: one authoritative copy wins, loud alert; if undecidable, it refuses to discard anythingValidated end-to-end on a real crash
Rolling Pool UpdateFully transparent, no special commandsReal RPU via Xen Orchestra: ~7 min, VMs uninterrupted
TWINSTOR upgradeOne command re-run of the installerZero downtime: storage keeps serving during the update

Recovery times include the servers' own boot time. TWINSTOR's share is tiny: in the power-outage test, storage was serving 22 seconds after the OS booted. Safe even on consumer SSDs: TWINSTOR detects volatile write caches and disables them, so an acknowledged write is on persistent media before the ack.

TWINSTOR vs the alternatives

Nothing besides the two hosts

Where TWINSTOR stands apart: no witness anywhere, no async data-loss window, no storage VM eating host RAM, no storage licensing layer, one command per node.

TWINSTORvSAN 2-nodeProxmox VE 2-nodeStorMagic SvSAN
Third box / witnessNoneWitness appliance, hosted off-site (required)QDevice on a third machine (required for HA quorum)Witness VM or cloud service (required for automated failover)
Synchronous replicationYes: every acknowledged write is on both hostsYesNo: async ZFS replication at intervals (data-loss window on failover)Yes
Split-brain tiebreakerYour existing network gatewayWitness quorumQDevice quorum voteWitness quorum
Per-host footprintOne tiny program orchestrating stock componentsIn-kernel, licensed per coreBuilt-in, no storage VMStorage VM per host (reserved RAM/CPU)
Setup1 command per nodeCluster + witness deploymentManual: cluster + QDevice + per-VM replication jobs; DRBD/ZFS knowledge expectedStorage VM + witness deployment
HypervisorXCP-ngESXi onlyProxmox VE onlyESXi / Hyper-V / SvHCI

Need 3+ nodes, 3-way replication, or storage tiers? That's XOSTOR's territory, not TWINSTOR's. vSphere Essentials-class deployments (the tier Broadcom discontinued) map one-to-one onto two XCP-ng hosts with TWINSTOR.

Where TWINSTOR shines

Built for the millions of sites that need exactly two servers

Retail & franchise networks

POS, inventory, surveillance on two small hosts per store. A direct replacement for 2-node vSAN hit by Broadcom licensing: no witness per store, no per-core tax, one golden script fleet-wide.

Remote & branch offices

File services, AD replicas, local apps with no IT staff on site. The ~3-minute unattended power-outage recovery means the site is back before anyone arrives in the morning.

Industrial & manufacturing edge

MES, SCADA historians, line supervision on the factory floor. Local reads give predictable latency; synchronous replication means a host failure never loses an acknowledged write.

Hospitality & healthcare

PMS/EMR systems that must stay up, in buildings with a network closet rather than a server room. Two short-depth hosts and a switch is the entire footprint.

Maritime, energy, disconnected sites

Ships, wind farms, substations: places where "send a technician" means days. No remote hands needed for recovery; the whole stack is observable from shore.

MSPs & VMware leavers

A standardized, scriptable 2-node building block: identical setup at every client, alerts centralized in Xen Orchestra, zero-downtime updates during business hours.

Two ways to deploy

Pick your topology at setup; both keep your data safe

Fate-shared (default)

Everything rides one network path. Ideal for ultra-simple, low-power edge boxes: a pair of mini-PCs on one switch. When the shared path fails, both halves fail together cleanly: no way to diverge, no split-brain to untangle. The most-tested path.

Needs: one link between the hosts. That's it.

Arbitrated (direct link)

Replication gets a dedicated back-to-back cable as its default path: synchronous writes and resync storms stay off the switch, which becomes the automatic, alerted fallback. Rides through switch-side gray failures that fate-shared cannot.

Needs: a second NIC per host and a crossover / DAC cable. Validated in the physical torture campaign.

Honest limits

TWINSTOR keeps exactly two synchronized copies on exactly two nodes. If you need 3+ nodes, 3-way replication, or multiple storage tiers, use XOSTOR or a SAN. The two hosts should share a reliable network (bonding and/or stacked switches recommended), and capacity exhaustion / silent bit-rot monitoring are on the documented backlog. Tech Preview status: it's a 0.x series, so read the release notes before upgrading a live pool.

Get started

Two commands. That's the whole setup.

The interactive wizard auto-detects the pool peer, lists candidate disks with model and serial, generates credentials, creates the storage, and enables HA. twinstor uninstall reverses everything.

Requirements: 2 XCP-ng hosts in a pool, a free local disk on each, a network link between them (1 Gbps works, 10 Gbps recommended), BIOS set to power on after AC loss. No SAN, no witness, no special wiring.

# Node A — run the wizard, follow the prompts
twinstor setup

# Node B — same wizard; it detects node A
twinstor setup

# Live status anytime
twinstor status

Two commands later you have shared storage, live migration, and automatic VM restart on host failure. Full docs ship with the preview: README, PRODUCT, FAQ, BEHAVIOR, TEST_PLAN, ARCHITECTURE.