Libvirt VM Snapshots (Fedora KVM Host)¶
My Fedora workstation runs KVM VMs via libvirt, including the two VMs that host my Hermes agents. For a long time the only safety net was manual snapshots from virt-manager — which worked right up until I ran out of disk. This page is the automatic snapshot + pruning setup that replaced that.
This is not Proxmox
Proxmox has snapshot scheduling built into the GUI. Plain libvirt on Fedora does not — you script it with virsh and schedule it with a systemd timer. That's what this page covers. For Proxmox VMs, use Proxmox Backup Server instead.
The two kinds of libvirt snapshots¶
This distinction is the whole ballgame, and it's the reason my manual snapshots ate so much disk.
| Internal (what virt-manager makes) | External disk-only (what the script makes) | |
|---|---|---|
| RAM state | Saved to a .save file (8–16 GB per snapshot) |
Not saved — VM must be running, snapshot is disk only |
| Disk | New qcow2 layer inside the chain | New qcow2 overlay, chain managed by libvirt |
| Restore | Resumes the VM exactly as it was | Reverts disk; VM reboots like a power loss |
| Cost per snapshot | Huge | Small diff |
External snapshots are the right shape for automation: cheap, scriptable, and pruning one of them is a single virsh command.
Use qemu:///system
virsh list with no URI often shows an empty list on a desktop Fedora install — it's talking to the session instance of libvirt, not the system one that actually owns your VMs. Every command below uses -c qemu:///system explicitly.
The script¶
~/bin/vm-snapshot.sh. It takes a disk-only snapshot of each VM, then deletes the oldest ones past a retention count. Snapshots it creates are named auto-<ISO timestamp>; anything without that prefix (your manual snapshots) is never touched.
#!/usr/bin/env bash
# Automatic libvirt external snapshots + pruning.
# Disk-only (no RAM image) so each snapshot is a small qcow2 diff, not a 15GB .save.
# Only touches snapshots named auto-* — manual snapshots are never pruned.
set -uo pipefail
URI="qemu:///system"
VMS=("OpenClaw native ubuntu24.04" "ubuntu24.04 Hermes Home")
KEEP="${KEEP:-4}" # keep the newest N auto-snapshots per VM
PREFIX="auto-"
LOG="$HOME/.local/state/vm-snapshots.log"
TS="$(date +%Y-%m-%dT%H:%M:%S)"
log() { echo "$(date '+%F %T') $*" >> "$LOG"; }
for vm in "${VMS[@]}"; do
state=$(virsh -c "$URI" domstate "$vm" 2>/dev/null) || { log "SKIP $vm (not found)"; continue; }
if [ "$state" != "running" ]; then
log "SKIP $vm (state=$state; external snapshots need a running domain)"
continue
fi
name="${PREFIX}${TS}"
# Try fs-quiesced (needs qemu-guest-agent in the guest); fall back to plain.
if virsh -c "$URI" snapshot-create-as --domain "$vm" --name "$name" \
--disk-only --quiesce --atomic >>"$LOG" 2>&1; then
log "OK $vm -> $name (quiesced)"
elif virsh -c "$URI" snapshot-create-as --domain "$vm" --name "$name" \
--disk-only --atomic >>"$LOG" 2>&1; then
log "OK $vm -> $name (crash-consistent, no guest agent)"
else
log "FAIL $vm -> $name"
continue
fi
# Prune: delete oldest auto-* snapshots beyond KEEP (lexicographic == chronological
# thanks to the ISO timestamp in the name). Deleting an in-chain external snapshot
# makes libvirt blockcommit the child into it first — safe, just IO-heavy.
mapfile -t autos < <(virsh -c "$URI" snapshot-list --domain "$vm" --name 2>/dev/null \
| sed 's/^[[:space:]]*//' | grep "^${PREFIX}" | sort)
if [ "${#autos[@]}" -gt "$KEEP" ]; then
for old in "${autos[@]:0:${#autos[@]}-$KEEP}"; do
if virsh -c "$URI" snapshot-delete --domain "$vm" --snapshotname "$old" >>"$LOG" 2>&1; then
log "PRUNE $vm deleted $old"
else
log "PRUNE-FAIL $vm $old"
fi
done
fi
done
Test it once by hand before scheduling anything:
~/bin/vm-snapshot.sh
tail ~/.local/state/vm-snapshots.log
virsh -c qemu:///system snapshot-list --domain "your vm name" --tree
The schedule¶
A systemd user timer, so no root and no crontab. User timers only survive logout if lingering is on:
~/.config/systemd/user/vm-snapshot.service:
[Unit]
Description=Automatic libvirt external snapshots for agent VMs
[Service]
Type=oneshot
ExecStart=%h/bin/vm-snapshot.sh
~/.config/systemd/user/vm-snapshot.timer:
[Unit]
Description=Daily automatic VM snapshots + pruning
[Timer]
OnCalendar=*-*-* 03:30:00
RandomizedDelaySec=900
Persistent=true
[Install]
WantedBy=timers.target
Persistent=true matters on a desktop — if the box was off at 03:30, the job runs at next boot instead of silently skipping.
systemctl --user daemon-reload
systemctl --user enable --now vm-snapshot.timer
systemctl --user list-timers vm-snapshot.timer
Verifying¶
# what exists, and in what order (source of truth)
virsh -c qemu:///system snapshot-list --domain "your vm name" --tree
# which disk file the running VM is actually writing to
virsh -c qemu:///system domblklist "your vm name"
# the job's own log
tail ~/.local/state/vm-snapshots.log
# force a run right now
systemctl --user start vm-snapshot.service
vol-list doesn't show everything
virsh vol-list does not reliably list the active external-snapshot overlay files. Use snapshot-list --tree and domblklist to see the real chain.
Gotchas that bit me¶
snapshot-list --namepads names with leading spaces. Grepping that output without stripping whitespace silently matches nothing, and pruning becomes a no-op that looks like it worked. The script pipes throughsed 's/^[[:space:]]*//'first.- Timestamps need seconds. With minute granularity, a manual test run in the same minute as a scheduled one collides with the existing overlay file and fails with
external snapshot file already exists. It fails safe (no corruption), but it fails. - Deleting a mid-chain snapshot is not instant. libvirt blockcommits the child layer into the parent first — that's copying gigabytes. On my box a 22 GB layer took about three minutes; the VMs keep running the whole time.
--quiesceneedsqemu-guest-agentinside the guest. Without it the snapshot is crash-consistent (the filesystem equivalent of pulling the plug). The script tries quiesced first and falls back automatically; installqemu-guest-agentin the guests to upgrade every future snapshot to a fs-consistent one.
Cockpit: visibility yes, creation no¶
The snapshots this script makes show up in Cockpit's Virtual Machines page (VM → Snapshots tab) like any other libvirt snapshot — Cockpit just asks libvirt, and libvirt doesn't care who made it. On Fedora that view needs cockpit-machines and libvirt-dbus installed, and your user in the libvirt group to see the system VMs.
But there's a difference worth knowing before you click anything in that UI:
- Cockpit's "Create snapshot" button makes internal snapshots — the kind that save the VM's RAM to a
.savefile. That's the exact snapshot type that cost me 400 GB. The Cockpit dialog has no disk-only/external option. Look in Cockpit all you want; let the timer be the one creating snapshots. - Revert and delete work from the UI, with the same rules as
virsh— reverting discards everything newer in the chain, and deleting a mid-chain snapshot starts a blockcommit that can churn for minutes on a big layer. The UI gives you no warning about either; it just starts doing it.
Restoring¶
# revert a running VM to a snapshot (disk reverts, VM reboots)
virsh -c qemu:///system snapshot-revert --domain "your vm name" --snapshotname auto-2026-10-08T03:30:00
Reverting deletes newer snapshots
snapshot-revert on an external snapshot chain discards everything after the snapshot you revert to. That includes the current active overlay. Know which auto- you're reverting to before you run it.