Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Watchdog + last-known-good rollback

What this gives you. Three recovery mechanisms — health-gated last-known-good (LKG) config promotion, boot-time rollback for invalid configs and counted crash loops, and a systemd watchdog — so a broken config or a wedged daemon does not strand the panel.

When to use it. You see a rollback banner, the daemon booted from an LKG config, or you want to know when to set watchdog.lkg_rollback_enabled = false (during rapid debug restarts). All three mechanisms are on by default with sensible bounds.

Quick setup. Leave everything at defaults:

[watchdog]
lkg_enabled = true
lkg_rollback_enabled = true
stability_window = "5m"
dormantctl status    # reports rollback state and running config fingerprint

dormant combines three recovery mechanisms:

  1. a health-gated last-known-good (LKG) config snapshot;
  2. boot-time rollback for invalid configs and counted crash loops;
  3. a systemd watchdog that restarts a process whose engine stops responding.

None of them changes the running rule engine in place. Rollback chooses the config for a new process; the watchdog asks systemd to replace a wedged one.

Last-known-good promotion

With watchdog.lkg_enabled = true, the daemon promotes the running config to $XDG_STATE_HOME/dormant/last-known-good.toml after watchdog.stability_window (default 5m) when all of these remain true:

  • every engine-liveness probe succeeds;
  • no reload occurs during the window;
  • the config bytes on disk still match the installed config;
  • no configured display reports every controller unhealthy.

An uncommanded display has no health evidence and does not block promotion. Repeated display-health deferrals are capped; after three, the candidate is promoted with lkg_promoted_with_unhealthy_display rather than leaving the daemon with no LKG.

The LKG is separate from <config_dir>/backups/. The web UI writes those backups before each browser apply. The daemon writes the LKG only after a healthy stability window, regardless of whether the config arrived through the web UI, file watch, SIGHUP, IPC reload, or boot.

Boot-time rollback

Rollback has three forms:

  • Immediate rollback: the current config fails to build, its bytes differ from the LKG, and the LKG builds successfully. The daemon logs config_rollback_boot and starts from the LKG. If the LKG also fails, startup fails normally with startup_failed.
  • Counted rollback: the same config fingerprint crashes or wedges at least three times in a 6-minute window. The next boot selects the LKG and records an active rollback.
  • Sticky continuation: while the original config remains unchanged, later boots in that crash storm keep selecting the LKG and log config_rollback_continued. After the quiet window, the daemon gives the original config one retry and logs config_rollback_retry.

The original config file is never overwritten. dormantctl status and the web dashboard show the pending rollback state.

Recovery after a rollback boot

Fix the intended config, validate it, then either let the file watcher pick it up or reload explicitly:

dormantctl validate
dormantctl reload

The running daemon's file watcher follows the OPERATOR config path (the file you edit) even while it is booted from the LKG substitute, so saving the fix from an editor triggers the same recovery without running dormantctl reload at all. Either way this is a live, in-place reload — no restart, no dropped engine state.

A reload that validates and builds successfully while a rollback is active clears the pending-reload banner, logs config_rollback_recovered, and atomically clears the persisted crash-loop state (rollback_active back to false, rolled_back_from back to null in crash-loop.json). Confirm recovery with:

dormantctl status

The banner disappearing (and status no longer reporting a rollback) confirms the fix took effect. A reload that still fails validation changes nothing: the daemon keeps running from the LKG, the banner and crash-loop state stay exactly as they were, and no LKG candidate is armed from the still-broken bytes.

Restarting the daemon (systemctl --user restart dormant) still works and remains a safe fallback — for example if the file watcher isn't running for some other reason — but it is no longer the required recovery step.

Counted-rollback gate

Set watchdog.lkg_rollback_enabled = false while intentionally restarting the daemon rapidly during debugging. This disables the counted trigger for a new storm. It does not disable immediate rollback for an invalid config or cancel an already-active sticky rollback.

Emergency wake

If displays are dark while the daemon is unavailable:

dormantctl emergency-wake

The command tries daemon IPC first with a bounded timeout. If IPC fails, it loads the config and drives the display controllers directly.

systemd watchdog

The shipped unit uses Type=notify and WatchdogSec=150. dormantd sends READY=1 after startup and sends WATCHDOG=1 only after the engine answers a liveness probe through its control channel. A scheduled but wedged process therefore does not keep itself alive with superficial watchdog pings.

Reload sends watchdog pings at internal step boundaries so a slow controller chain is not killed during a healthy generation swap. If a probe fails, the daemon logs watchdog_probe_failed, withholds the ping, and resets any pending LKG stability window.

Without NOTIFY_SOCKET / WATCHDOG_USEC, the same engine probe runs every 30 seconds for LKG health accounting, but no systemd notification is sent.

Upgrade order

Install the new dormantd binary before reloading a unit that changes from Type=simple to Type=notify:

install -Dm755 target/release/dormantd ~/.local/bin/dormantd
systemctl --user daemon-reload
systemctl --user restart dormant

Reloading the new unit while an old binary is still installed makes systemd wait for a READY=1 message that binary cannot send.

Configuration

[watchdog]
lkg_enabled = true
lkg_rollback_enabled = true
stability_window = "5m"
KeyDefaultNotes
watchdog.lkg_enabledtrueEnable LKG promotion
watchdog.lkg_rollback_enabledtrueEnable counted crash-loop rollback
watchdog.stability_window"5m"Healthy runtime required before promotion; minimum 30s

WatchdogSec, crash-count thresholds, and the quiet window are fixed recovery mechanics rather than config policy.

State and privacy

Recovery state lives under $XDG_STATE_HOME/dormant (normally ~/.local/state/dormant):

FilePurpose
last-known-good.tomlProven-stable config snapshot
last-known-good.meta.jsonAdvisory fingerprint and promotion timestamp
crash-loop.jsonBounded crash/restart history
discount-<nonce>One-shot clean-exit marker used around startup races

Directories are mode 0700; files are mode 0600. The files stay local and contain no telemetry.