Business → High availability (business.ha, Warda Business, beta;
internal/ha, internal/admin/ha.go): two boxes paired, a primary and
a secondary, share one IPv4 address with VRRP (keepalived); the devices
are given that address as their DNS server, and whichever box holds it
answers them. When the box that holds it stops, or its DNS stops
answering, the other takes it within seconds.
Network requirements:
- both boxes on the same network, each with a fixed IPv4 address (a DHCP reservation or a static address);
- a shared address free on that network, with the length of its
network (
192.168.1.2/24): outside the ranges of every DHCP server, given to no device; - a virtual router (1 to 255) the same on both boxes and different from every other VRRP router of the network (derived from the secret of the pair by default; change it when the network already has VRRP);
- VRRP (IP protocol 112) allowed between the two boxes: Warda configures it in unicast to the address of the other box, so no multicast is needed, but a firewall between them must let it through;
- HTTPS (443, or the port of Warda) from the secondary to the primary and back, for the pairing and the synchronisation;
- the clocks of both boxes on time (NTP): the signed calls between them are refused beyond 5 minutes of difference;
- the devices given the shared address as DNS server, by the DHCP server of the network (the box of the provider, the router), or by the DHCP server of Warda on the primary (it then gives the shared address as DNS server; it stays off on the secondary, which refuses to turn it on (409): two servers would give addresses on the same network; it moves with a promotion, see below).
Pairing. On the future primary, Create a pair → Make a pairing
code: a code of one use WH-XXXX-XXXX-XXXX-XXXX-XXXX (five groups of
four letters and digits, no 0/O nor 1/I: 100 bits), valid 15 minutes,
dropped after 5 wrong tries, with the address of this box to type on the
other one. On the future secondary, Join a pair: Address of the
primary (its IPv4 address, with its HTTPS port when it is not 443) and
Pairing code (any case; a code of another form is refused at once,
400), then Join (confirmed: "The rules, devices, household, settings
and accounts of this box are replaced by those of the primary within a
minute. Its administration log is kept."). Each box proves it knows the
code over the fingerprints of both HTTPS certificates (HMAC-SHA256 keyed
with a key derived from the code by Argon2id — 32 MiB, 2 passes, about a
sixth of a second on a Raspberry Pi 4 —, salted with both fingerprints and
a nonce): a machine in the middle, which cannot show the certificate of
the primary, is refused, and one that records a proof cannot find the
code from it within its 15 minutes. A box of 0.6.8 and a box of a later
version do not pair (update both first). The primary then gives a
secret shared by both boxes; every later call between them
(/api/v1/ha/peer/…, HTTPS only, without an account) is signed with it
(X-Warda-HA-Date, X-Warda-HA-Signature) and goes to the certificate
pinned at the pairing. Making a code and joining need backups.restore:
they give or replace the whole configuration. The page shows the
Certificate of this box (its fingerprint) to compare.
Synchronisation. The secondary asks the primary for its configuration
every minute (task ha.sync): a backup of the store without the history,
applied only when it changed, by restarting Warda on the secondary (a few
seconds, like a restore: its journal starts again), after checking it.
Everything comes from the primary, the accounts included (the same
administrators on both boxes), except what belongs to each box: its
administration log, its identifier, the name and certificate of its
encrypted DNS, the addresses it gives for its own names, its site of
Warda Business, the pair itself; the DHCP servers of Warda stay off on the
secondary, which learns at each exchange those the primary runs (IPv4,
IPv6). The secondary says so ("This box is the secondary: its
configuration comes from the primary. Make your changes on the primary; a
change made here is replaced at the next change of the primary.").
Without an exchange for 3 minutes, the other box is shown unreachable.
keepalived. On the primary, Shared address: the address, the
Virtual router, the Announcements (seconds) (1 to 10), and on each
box its Interface and its Priority (1 to 254; 150 for the primary,
100 for the secondary: the box of the highest priority holds the address
when both are healthy). Warda makes the configuration of keepalived from
them (Configuration of keepalived, previewed before saving: GET /api/v1/business/ha/keepalived?vip=&interface=&router_id=&advert=&priority=&peer=):
both boxes start as BACKUP, the primary takes the address back 30
seconds after it is healthy again (preempt_delay), a password of VRRP
derived from the secret of the pair (VRRP sends it in clear: it only keeps
a mistaken router out), unicast_peer the other box, and a check every 5
seconds, warda healthcheck -url http://127.0.0.1:80/healthz -dns 127.0.0.1:53 (2 failures: the box goes to FAULT and leaves the
address; the DNS of Warda must answer a name it answers itself); each
change of state runs warda ha-notify -data-dir /var/lib/warda, which
writes the state for the page (State of VRRP: Held by this box,
Held by the other box, Fault, keepalived stopped).
- Debian package and Raspberry Pi image: the service never writes the
configuration of keepalived itself; it writes its parameters in its data
directory and asks
warda-system.service(root), which checks each of them, installs keepalived when missing, writes/etc/keepalived/keepalived.conffrom the template of Warda (only a file marked as written by Warda is ever replaced) and restarts it. - Docker (and an archive): Warda writes
keepalived.confin its data directory; keepalived runs in its own container on the network of the host, withNET_ADMIN,NET_BROADCAST,NET_RAWandDAC_OVERRIDE, the program of Warda for its checks and the data directory: an example is inpackaging/docker-compose.yml(commented out). Restart it after a change of the shared address (docker compose restart keepalived).
Make this box primary (on the secondary; the other box must answer:
it becomes the secondary and takes the configuration of this one; the DHCP
servers of Warda move too: the old primary turns its own off first, then
the new one turns on those the old one ran, each restarting when its DHCP
servers change, so that they never run on both — confirmed: "… If the
other box runs the DHCP server of Warda, it moves to this one.") and
End the pair (both boxes forget each other and keepalived stops: the
shared address answers no more; each box keeps its configuration). To
replace a primary that is gone for good: end the pair on the secondary,
then pair again. The administration log records ha.pairing,
ha.pairing.cancel, ha.paired, ha.settings, ha.sync.failed,
ha.promoted, ha.demoted, ha.unpaired.
APIs (administrators; system unless said): GET /api/v1/business/ha
(role, fingerprint, port, interfaces, pairing — its code only
for the accounts with backups.restore —, peer, peer_fingerprint,
since, peer_seen, peer_reachable, shared, local, sync, vrrp,
managed, keepalived, config, config_error, dhcp; in config,
the password of VRRP is ******** for the accounts that do not change the
system); POST and
DELETE /api/v1/business/ha/pairing (make or drop a code); POST /api/v1/business/ha/join {"address": "192.168.1.10", "code": "WH-…"}; PUT /api/v1/business/ha/settings {"vip", "router_id", "advert", "interface", "priority"}; GET /api/v1/business/ha/keepalived (the password of VRRP hidden as in
config); POST /api/v1/business/ha/promote;
DELETE /api/v1/business/ha (end the pair). A pair is two boxes on one
network; later phases are in the design note.