Docs / Losing the controller
Losing the controller, and getting it back
The controller is one machine. Its database holds every node, customer, address and machine record, and its keys include the CA every node trusts. The machines themselves do not need the controller to keep running: a guest on a node stays up while the controller is gone. What stops is the panel, the API, scheduled backups and anything that changes a machine.
What is kept, and where
| Copy | Where | Made | Opens with |
|---|---|---|---|
controller-<time>.db and .keys.tar.gz | the controller's backup_dir | nightly, 14 kept | nothing: it is plain, and it holds the CA key |
controller-<time>.recovery | the same directory, and /var/lib/klyrn-vm/controller-recovery on every managed node | nightly once a recovery key exists, 3 kept | the recovery key |
The first row is on the controller's own machine. If that machine is lost, so is that copy. The second row is the one that is somewhere else.
Before anything goes wrong
- Settings, Backups, "If this server is lost": Make a recovery key.
- Store the key it shows (
klyrn-recovery-1:...) somewhere that is not the controller and not a node. It is shown once and the controller does not keep it. Without it nobody can open a sealed copy. - Check the same panel a day later: it names each node that holds the newest copy. An alert is raised if no node has fetched one in two days.
A node can store a sealed copy and cannot read it. That is deliberate: the archive holds the CA's private key, and a node that could read it could give itself any identity.
Rebuilding on a new machine
- Install the same or a newer
klyrn-vmbinary on the new machine. Do not start the service. - Take the newest
controller-*.recoveryfrom any node (/var/lib/klyrn-vm/controller-recovery/). - Run:
klyrn-vm restore --recovery controller-<time>.recovery --key klyrn-recovery-1:...
--key also accepts the path of a file holding the key. If you still have the plain pair from the old machine, --db and --keys do the same without a key.
- The command checks the database before it writes anything, refuses to replace a controller that is already there unless
--forceis given, and prints how many nodes and machines it restored. It starts nothing. - Point the controller's hostname at the new machine. The nodes dial that name; none of them needs touching. Their certificates chain to the CA that was just restored.
systemctl start klyrn-vm-controller, and watch the nodes come back.
What a restore does not bring back
- Anything that happened after the copy was sealed: at most a day.
- Backups of machines that were stored on the lost controller's own disks, when the controller was also the backup server. Those are gone with it. A backup server on a different machine from the controller avoids this.
- A node whose certificate ran out while the controller was away needs to be enrolled again.
Proving it
scripts/restore-drill.sh boots the newest plain copy on the controller's own machine, with no network, and compares its nodes, machines and CA with the live controller. It proves the copy is restorable; it does not prove the copy survives the machine. The drill for that is the procedure above on a second machine, and it should be run once after the recovery key is made.
This page is generated from the guide that ships in the product's repository, so it describes the release it was built from. What changed in each release.