Skip to content

Self Hosted

The manual snapshot habit that ate 400 GB

I run two Hermes agent VMs on my Fedora workstation — libvirt/KVM, not Proxmox. For months the backup story for those VMs was me, clicking "snapshot" in virt-manager every so often when I was about to do something risky. That worked, in the sense that the snapshots existed. What I didn't notice is what they were doing to my disk.

The 8.5 GB that showed up as free

The fix in the last post was two more matching 32 GB sticks, and the sticks arrived. Four slots filled, all four the same part number, negotiated speed 4800 MT/s at 2DPC — right in the predicted range. With 128 GB installed I figured the GPU VM could take 114 and still leave the host its 2 GB. The host OOM'd.

The 401 spam that wasn't HA

Gotify kept pinging with 401s. The natural read is "an app's token is broken" or "gotify is down." Neither was true. The gotify server is healthy, and the 401 was content inside a forwarded log line. The thing posting the alert was loggifly, the log-monitor container, matching a keyword in a container's access log and forwarding it through Apprise.

This is the class of work: gotify ping, who actually posted it, what container logged it, who the real client is, what the true cause is.

The 56 GB that wasn't there

The GPU VM on the Powerspec box runs a 27B model that keeps its experts resident in RAM, so I wanted more RAM for it. The host has 64 GB, the VM had 50, and I figured I had room to grow. I bumped the VM to 56 GB and the host OOM'd. Twice.

This is the story of why, and it's mostly a story about assumptions I didn't check.

The false alarm that became the watchdog

The first run of the health watchdog reported pbs-backup-local as inactive, a CRITICAL on a backup path that was perfectly fine. The storage was healthy the whole time. What the watchdog actually caught was a momentary auth handshake, and it treated one bad sample as ground truth.

That false alarm is why the watchdog works the way it does now.

The one-byte bug

The voice assistant on the HA box has three legs: speech-to-text on the GPU box, an LLM brain, and text-to-speech back out. For weeks, a full pipeline run over the websocket died in about 30 milliseconds. The whisper bridge log showed a client connecting and immediately disconnecting with zero audio. Read that as "HA can't reach the bridge" and you'll chase it for days.

It wasn't HA. It wasn't the GPU. It was my own test client, and the whole thing came down to one byte.

The stale mount that ate half the backup windows

The PBS datastore on the external drive was shared into the VM over a PVE virtiofs mount. Backups failed roughly half the time, always with the same error, os error 23 — ENFILE, too many open files. Half the 03:00 UTC vzdump window was going to waste, and the failure looked random.

It wasn't random, and the fix wasn't a better mount.