Skip to content

The false alarm that became the watchdog

The first run of the health watchdog reported pbs-backup-local as inactive, a CRITICAL on a backup path that was perfectly fine. The storage was healthy the whole time. What the watchdog actually caught was a momentary auth handshake, and it treated one bad sample as ground truth.

That false alarm is why the watchdog works the way it does now.

What the inactive really was

Here's the sequence. I added the pbs-backup-local storage entry to the Powerspec node's storage.cfg. The first check came back inactive with a 401, because PVE authenticates to PBS with a password file (/etc/pve/priv/storage/pbs-backup-local.pw), and that node didn't have it yet. I copied the credential over, re-checked, and it went active.

Notice what the inactive status means. It's PVE saying "I can't authenticate to my backup server." That's exactly the thing you'd want to know about before a backup silently fails. So the storage check isn't validating the USB drive itself, it's validating that the backup path from each node to PBS actually works. The check was doing its job. The problem was the alerting.

The structural problem

A single bad sample as ground truth can't work for a state check. A pool that flaps to inactive for a moment — an auth refresh, the PBS VM bouncing, a ten-second SSH hiccup, a VM caught mid-restart — fires a CRITICAL even though the path is fine and self-corrected.

The fix is debounce: re-sample a bad state finding before alerting, and only page if the problem persists across the retry window.

  • Transient blip, sample 1 bad, sample 2 cleared, stays silent
  • Real outage, bad bad bad, still alerts

Config is debounce: {retries: 2, delay: 10}. Set retries: 0 to turn it off.

What I deliberately did not debounce

The staleness checks, Kopia snapshot age and vzdump log age, are not debounced, and that's on purpose. A stale backup is stale. Re-sampling won't fix it, it'll only delay a real alert. State checks describe a point in time that can be transient; staleness checks describe history. Different rules.

Verified behavior

  • pbs-backup-local on both nodes, 6 samples 3 seconds apart, active all six.
  • Debounce unit tests: transient to SILENT, persistent to ALERTED, retries: 0 to ALERTED.
  • Recent cron runs: all empty, all silent.

The honest caveat is that this suppresses transient false alarms, not real ones. If the PBS VM were genuinely down, the watchdog pages after the ~20 second retry window, because that's what you'd want to know.

The full design is in the pve-health-watchdog project page and the notifications page.

Comments