Warda-DNSDocs Warda: from ward — to protect, guardian
v0.6.9

High availability

Business → High availability (business.ha, Warda Business, beta; internal/ha, internal/admin/ha.go): two boxes paired, a primary and a secondary, share one IPv4 address with VRRP (keepalived); the devices are given that address as their DNS server, and whichever box holds it answers them. When the box that holds it stops, or its DNS stops answering, the other takes it within seconds.

Network requirements:

  • both boxes on the same network, each with a fixed IPv4 address (a DHCP reservation or a static address);
  • a shared address free on that network, with the length of its network (192.168.1.2/24): outside the ranges of every DHCP server, given to no device;
  • a virtual router (1 to 255) the same on both boxes and different from every other VRRP router of the network (derived from the secret of the pair by default; change it when the network already has VRRP);
  • VRRP (IP protocol 112) allowed between the two boxes: Warda configures it in unicast to the address of the other box, so no multicast is needed, but a firewall between them must let it through;
  • HTTPS (443, or the port of Warda) from the secondary to the primary and back, for the pairing and the synchronisation;
  • the clocks of both boxes on time (NTP): the signed calls between them are refused beyond 5 minutes of difference;
  • the devices given the shared address as DNS server, by the DHCP server of the network (the box of the provider, the router), or by the DHCP server of Warda on the primary (it then gives the shared address as DNS server; it stays off on the secondary, which refuses to turn it on (409): two servers would give addresses on the same network; it moves with a promotion, see below).

Pairing. On the future primary, Create a pair → Make a pairing code: a code of one use WH-XXXX-XXXX-XXXX-XXXX-XXXX (five groups of four letters and digits, no 0/O nor 1/I: 100 bits), valid 15 minutes, dropped after 5 wrong tries, with the address of this box to type on the other one. On the future secondary, Join a pair: Address of the primary (its IPv4 address, with its HTTPS port when it is not 443) and Pairing code (any case; a code of another form is refused at once, 400), then Join (confirmed: "The rules, devices, household, settings and accounts of this box are replaced by those of the primary within a minute. Its administration log is kept."). Each box proves it knows the code over the fingerprints of both HTTPS certificates (HMAC-SHA256 keyed with a key derived from the code by Argon2id — 32 MiB, 2 passes, about a sixth of a second on a Raspberry Pi 4 —, salted with both fingerprints and a nonce): a machine in the middle, which cannot show the certificate of the primary, is refused, and one that records a proof cannot find the code from it within its 15 minutes. A box of 0.6.8 and a box of a later version do not pair (update both first). The primary then gives a secret shared by both boxes; every later call between them (/api/v1/ha/peer/…, HTTPS only, without an account) is signed with it (X-Warda-HA-Date, X-Warda-HA-Signature) and goes to the certificate pinned at the pairing. Making a code and joining need backups.restore: they give or replace the whole configuration. The page shows the Certificate of this box (its fingerprint) to compare.

Synchronisation. The secondary asks the primary for its configuration every minute (task ha.sync): a backup of the store without the history, applied only when it changed, by restarting Warda on the secondary (a few seconds, like a restore: its journal starts again), after checking it. Everything comes from the primary, the accounts included (the same administrators on both boxes), except what belongs to each box: its administration log, its identifier, the name and certificate of its encrypted DNS, the addresses it gives for its own names, its site of Warda Business, the pair itself; the DHCP servers of Warda stay off on the secondary, which learns at each exchange those the primary runs (IPv4, IPv6). The secondary says so ("This box is the secondary: its configuration comes from the primary. Make your changes on the primary; a change made here is replaced at the next change of the primary."). Without an exchange for 3 minutes, the other box is shown unreachable.

keepalived. On the primary, Shared address: the address, the Virtual router, the Announcements (seconds) (1 to 10), and on each box its Interface and its Priority (1 to 254; 150 for the primary, 100 for the secondary: the box of the highest priority holds the address when both are healthy). Warda makes the configuration of keepalived from them (Configuration of keepalived, previewed before saving: GET /api/v1/business/ha/keepalived?vip=&interface=&router_id=&advert=&priority=&peer=): both boxes start as BACKUP, the primary takes the address back 30 seconds after it is healthy again (preempt_delay), a password of VRRP derived from the secret of the pair (VRRP sends it in clear: it only keeps a mistaken router out), unicast_peer the other box, and a check every 5 seconds, warda healthcheck -url http://127.0.0.1:80/healthz -dns 127.0.0.1:53 (2 failures: the box goes to FAULT and leaves the address; the DNS of Warda must answer a name it answers itself); each change of state runs warda ha-notify -data-dir /var/lib/warda, which writes the state for the page (State of VRRP: Held by this box, Held by the other box, Fault, keepalived stopped).

  • Debian package and Raspberry Pi image: the service never writes the configuration of keepalived itself; it writes its parameters in its data directory and asks warda-system.service (root), which checks each of them, installs keepalived when missing, writes /etc/keepalived/keepalived.conf from the template of Warda (only a file marked as written by Warda is ever replaced) and restarts it.
  • Docker (and an archive): Warda writes keepalived.conf in its data directory; keepalived runs in its own container on the network of the host, with NET_ADMIN, NET_BROADCAST, NET_RAW and DAC_OVERRIDE, the program of Warda for its checks and the data directory: an example is in packaging/docker-compose.yml (commented out). Restart it after a change of the shared address (docker compose restart keepalived).

Make this box primary (on the secondary; the other box must answer: it becomes the secondary and takes the configuration of this one; the DHCP servers of Warda move too: the old primary turns its own off first, then the new one turns on those the old one ran, each restarting when its DHCP servers change, so that they never run on both — confirmed: "… If the other box runs the DHCP server of Warda, it moves to this one.") and End the pair (both boxes forget each other and keepalived stops: the shared address answers no more; each box keeps its configuration). To replace a primary that is gone for good: end the pair on the secondary, then pair again. The administration log records ha.pairing, ha.pairing.cancel, ha.paired, ha.settings, ha.sync.failed, ha.promoted, ha.demoted, ha.unpaired.

APIs (administrators; system unless said): GET /api/v1/business/ha (role, fingerprint, port, interfaces, pairing — its code only for the accounts with backups.restore —, peer, peer_fingerprint, since, peer_seen, peer_reachable, shared, local, sync, vrrp, managed, keepalived, config, config_error, dhcp; in config, the password of VRRP is ******** for the accounts that do not change the system); POST and DELETE /api/v1/business/ha/pairing (make or drop a code); POST /api/v1/business/ha/join {"address": "192.168.1.10", "code": "WH-…"}; PUT /api/v1/business/ha/settings {"vip", "router_id", "advert", "interface", "priority"}; GET /api/v1/business/ha/keepalived (the password of VRRP hidden as in config); POST /api/v1/business/ha/promote; DELETE /api/v1/business/ha (end the pair). A pair is two boxes on one network; later phases are in the design note.