pve-health-watchdog¶
A silent-when-healthy homelab health watchdog. It probes hosts over SSH, runs the checks you've configured, and stays completely quiet when everything is fine. When something is wrong it prints a compact alert to stdout and optionally pushes it to Gotify.
Built for Proxmox (VE + PBS) but the check registry makes it work for anything you can reach over SSH.
The false alarm that started it¶
The first run reported pbs-backup-local as inactive, a CRITICAL on a backup path that was perfectly fine. The storage was healthy the whole time. What the watchdog caught was a momentary auth handshake, and it treated one bad sample as ground truth.
That single false alarm is why the debounce design exists. The full story is in The false alarm that became the watchdog.
Why¶
Proxmox can alert you when a backup job runs, but it can't tell you "VM 102 stopped at 3 a.m.", "your ZFS mirror degraded", "your external backup drive is mounted but no longer writable", or "Kopia hasn't produced a snapshot in 4 days." This fills that gap.
Checks¶
| Check | What it verifies |
|---|---|
pve_vms |
Expected VMs are running (parses qm list) |
pve_pools |
Every storage pool is online/active/ready (pvesm status) |
zfs |
ZFS pools are ONLINE with no degraded/offline/faulted members |
mount |
A path is mounted (and optionally provably writable) |
kopia |
Newest Kopia snapshot is younger than max_age_days |
vzdump_staleness |
Newest vzdump log is younger than max_age_days |
Findings carry severity, critical (wrong state / unreachable) or warning (staleness thresholds). Gotify priority is configurable per level.
Debounce¶
A bad state sample is re-checked a few times before alerting, so a transient blip doesn't page you. Config is debounce: {retries: 2, delay: 10}; retries: 0 disables it.
- Transient blip, sample 1 bad, sample 2 cleared, stays silent
- Real outage, bad bad bad, still alerts
Staleness checks are deliberately not debounced. A stale backup is stale, and re-sampling only delays a real alert. State checks describe a point in time that can be transient; staleness checks describe history.
The false alarm that produced this design is in the blog post The false alarm that became the watchdog.
Running it¶
Stdlib + PyYAML only, no framework, no daemon, no state. It uses your ~/.ssh/config aliases (ssh pve0), so it inherits whatever identity you've set up, and BatchMode=yes is forced so password prompts are impossible.
cp config.example.yaml config.yaml
cp .env.example .env
python3 watchdog.py --notify-test # test Gotify
python3 watchdog.py --verbose # healthy → one-line summary
python3 watchdog.py # healthy → prints nothing
config.yaml holds your specifics, .env holds only secrets (referenced as ${GOTIFY_TOKEN}). Both are gitignored, so the repo is safe to share as-is.
Schedule it however you like. In my setup it runs every 30 minutes as a no_agent cron script — empty stdout means nothing is delivered, so a healthy run is silent.
See also¶
- Notifications — the Gotify side
- Repo: Wildium/pve-health-watchdog
- Blog post: The false alarm that became the watchdog