Skip to content

The stale mount that ate half the backup windows

The PBS datastore on the external drive was shared into the VM over a PVE virtiofs mount. Backups failed roughly half the time, always with the same error, os error 23 — ENFILE, too many open files. Half the 03:00 UTC vzdump window was going to waste, and the failure looked random.

It wasn't random, and the fix wasn't a better mount.

Why virtiofs is the wrong tool for backup storage

A virtiofs share is a FUSE daemon on the host (virtiofsd) talking to the guest over a vhost-user socket. Every file operation the guest does on that share goes through the FUSE layer. For a backup datastore that opens a lot of files and holds them open while writing chunks, that layer is the problem. The daemon hits its file-descriptor ceiling and the guest gets ENFILE.

There's no tuning that makes this right for backup-critical storage. The rule I landed on: prefer raw-block passthrough over virtiofs for backup storage. A raw block device passthrough gives the guest the device directly, no FUSE in the middle, no descriptor ceiling. I migrated the PBS datastore off virtiofs to raw-block passthrough on 2026-08-29 and the ENFILE failures stopped.

The second half of the story — the mount that went stale

Raw-block passthrough fixed the descriptor problem, but the remaining virtiofs share on the personal VM (the Kopia repo on the 4 TB external drive) taught me a different lesson.

A drive that has run clean for a long time and suddenly throws I/O errors is, on first check, almost always a transient mount/transport fault, not a dying disk. I had a huge journalctl | grep -c 'Input/output error' count and nearly reported a failing-drive signature from that number alone. That count is the dead mount being hammered, not sector rot.

What actually happened: a power outage on the hypervisor. The USB drive re-enumerated (sdf to sdg), which is a boot signature, not a cable flake. Two things broke independently. The host FUSE daemon was stale, started on the old device node. And the guest's virtio-fs transport wedged at "Transport endpoint is not connected."

The guest side is the hard part. A guest umount plus a fresh mount does not re-attach. A remount that returns rc=0 but still reports the dead endpoint is the signature that the vhost-user channel is broken at kernel level. qm reset (the ACPI power-button reset) was insufficient, the error advanced to "Connection refused" and stayed there. Only qm stop plus qm start worked, because a fresh QEMU start spawns a new virtiofsd and re-establishes the channel.

The part that actually hurt

The unclean shutdown left the NTFS volume dirty and killed the Kopia container, which never came back. The hourly maintenance heartbeat stops at Aug 28 and does not resume until a manual restart on Sep 26. That's a month of dark backups, invisible until I looked.

The cheapest answer to "how long has this backup path been dark" is kopia maintenance info — the gap in the advance-epoch run history pins exactly when the server last ran. And read-clean is not writable. After the transport was fixed, existing files read fine but creating a new file failed ENENT on specific directories with zero raw I/O errors. That's NTFS directory-index damage from the drive being yanked mid-write, not a transport problem. Always prove the write path by actually running a backup write before calling it recovered.

The root-cause class here is a power event hitting an unmonitored backup path. The flaky-USB framing was a red herring. The month-dark window was a monitoring gap, not hardware. If last reboot shows a pattern of unclean shutdowns, get a UPS.

The reference setup is on the PBS page and the backups page.

Comments