The manual snapshot habit that ate 400 GB¶
I run two Hermes agent VMs on my Fedora workstation — libvirt/KVM, not Proxmox. For months the backup story for those VMs was me, clicking "snapshot" in virt-manager every so often when I was about to do something risky. That worked, in the sense that the snapshots existed. What I didn't notice is what they were doing to my disk.
What virt-manager snapshots actually cost¶
When you snapshot a running VM through virt-manager, you get an internal snapshot: a new qcow2 layer in the disk chain plus a .save file holding the VM's entire RAM state. My agent VMs are allocated 4 GB and 16 GB of RAM, and those .save files came in at 15–16 GB each, every snapshot, every time.
The OpenClaw VM had seven snapshots going back to May. The chain looked like this:
2026-05-06 before cc-connect
+- 2026-05-23 Before hermes
+- 2026-06-08 Disk full
+- 2026-06-19 awesome
+- 2026-08-24
+- 2026-09-15
+- 2026-09-30
Yes, one of them is named "Disk full". Past me was apparently foreshadowing. The layers plus the RAM saves added up to roughly 425 GB on that one VM, on a 1.9 TB root filesystem that was sitting at 62% used.
External snapshots are the right shape for automation¶
libvirt also does external snapshots: disk-only, no RAM image, each one a small qcow2 overlay. The VM has to stay running, and restoring one is the disk equivalent of a power loss — the guest reboots rather than resuming mid-thought. For a rollback safety net on agent VMs, that's exactly what I want, and they're cheap enough to create and delete on a schedule.
The whole automation is one script and a systemd user timer. The full reference with the script, the timer units, and the gotchas is on the Libvirt VM Snapshots page. The short version:
~/bin/vm-snapshot.shtakes a--disk-only --atomicsnapshot of each VM viavirsh -c qemu:///system snapshot-create-as, names itauto-<timestamp>, then deletes the oldestauto-*snapshots past a retention count of 4.- A systemd user timer runs it at 03:30 every night, with
Persistent=trueso a powered-off box catches up at next boot.loginctl enable-lingeris what makes user timers survive logout. - The
auto-prefix is the safety line: the script only ever prunes snapshots it made itself. Manual ones are invisible to it.
The gotchas I hit¶
Three things that aren't obvious and cost me a debugging detour each:
virsh listshowed nothing. On a desktop Fedora install the default URI is the session libvirt instance, which doesn't own your VMs. Every command needs-c qemu:///system.snapshot-list --namepads its output with leading spaces. My first prune loop grepped for^auto-against that output, matched nothing, and "pruned" zero snapshots while logging success. The fix is ased 's/^[[:space:]]*//'before the grep.- Deleting a mid-chain snapshot is a data migration, not an unlink. libvirt blockcommits the child layer into the parent first. Deleting my 67 GB June snapshot took several minutes of disk churn with the VM running the whole time. That's fine on a schedule at 3:30 AM; it's not fine if you're expecting
snapshot-deleteto be instant.
One UI note: the auto- snapshots show up in Cockpit's VM → Snapshots tab like anything else, since Cockpit just reads libvirt. But Cockpit's own "Create snapshot" button makes internal snapshots — RAM .save files, the exact thing that cost me 400 GB — with no disk-only option in the dialog. So the UI is good for looking, and the timer is the only thing allowed to create.
One thing I'm still honest about: neither guest has qemu-guest-agent installed, so these snapshots are crash-consistent — filesystem-wise it's like pulling the plug. The script already tries --quiesce first and falls back, so the moment I install the agent package in both guests every snapshot upgrades itself to a fs-consistent one. That's a sudo dnf install qemu-guest-agent away, which is the one step I keep putting off because it needs me, not the script.
Cleaning out the old chain¶
With the automation in place I finally deleted the whole manual chain except the most recent snapshot on each VM. Seven deletions, each one a blockcommit of gigabytes, both VMs live the entire time. When it finished, df went from 1.2T used to 816G — about 400 GB back — and the trees collapsed to one manual snapshot plus the managed auto- ones per VM.
If you're running plain libvirt and your snapshot plan is "me, remembering", the script is about sixty lines and the timer is nine. The part that takes real time is the first prune — turns out moving four hundred gigabytes of qcow2 layers is just slow.