Skip to content

pve-health-watchdog

A silent-when-healthy homelab health watchdog. It probes hosts over SSH, runs the checks you've configured, and stays completely quiet when everything is fine. When something is wrong it prints a compact alert to stdout and optionally pushes it to Gotify.

Built for Proxmox (VE + PBS) but the check registry makes it work for anything you can reach over SSH.

The false alarm that started it

The first run reported pbs-backup-local as inactive, a CRITICAL on a backup path that was perfectly fine. The storage was healthy the whole time. What the watchdog caught was a momentary auth handshake, and it treated one bad sample as ground truth.

That single false alarm is why the debounce design exists. The full story is in The false alarm that became the watchdog.

Why

Proxmox can alert you when a backup job runs, but it can't tell you "VM 102 stopped at 3 a.m.", "your ZFS mirror degraded", "your external backup drive is mounted but no longer writable", or "Kopia hasn't produced a snapshot in 4 days." This fills that gap.

Checks

Check What it verifies
pve_vms Expected VMs are running (parses qm list)
pve_pools Every storage pool is online/active/ready (pvesm status)
zfs ZFS pools are ONLINE with no degraded/offline/faulted members
mount A path is mounted (and optionally provably writable)
kopia Newest Kopia snapshot is younger than max_age_days
vzdump_staleness Newest vzdump log is younger than max_age_days

Findings carry severity, critical (wrong state / unreachable) or warning (staleness thresholds). Gotify priority is configurable per level.

Debounce

A bad state sample is re-checked a few times before alerting, so a transient blip doesn't page you. Config is debounce: {retries: 2, delay: 10}; retries: 0 disables it.

  • Transient blip, sample 1 bad, sample 2 cleared, stays silent
  • Real outage, bad bad bad, still alerts

Staleness checks are deliberately not debounced. A stale backup is stale, and re-sampling only delays a real alert. State checks describe a point in time that can be transient; staleness checks describe history.

The false alarm that produced this design is in the blog post The false alarm that became the watchdog.

Running it

Stdlib + PyYAML only, no framework, no daemon, no state. It uses your ~/.ssh/config aliases (ssh pve0), so it inherits whatever identity you've set up, and BatchMode=yes is forced so password prompts are impossible.

cp config.example.yaml config.yaml
cp .env.example .env
python3 watchdog.py --notify-test   # test Gotify
python3 watchdog.py --verbose       # healthy → one-line summary
python3 watchdog.py                 # healthy → prints nothing

config.yaml holds your specifics, .env holds only secrets (referenced as ${GOTIFY_TOKEN}). Both are gitignored, so the repo is safe to share as-is.

Schedule it however you like. In my setup it runs every 30 minutes as a no_agent cron script — empty stdout means nothing is delivered, so a healthy run is silent.

See also