Skip to content

Libvirt VM Snapshots (Fedora KVM Host)

My Fedora workstation runs KVM VMs via libvirt, including the two VMs that host my Hermes agents. For a long time the only safety net was manual snapshots from virt-manager — which worked right up until I ran out of disk. This page is the automatic snapshot + pruning setup that replaced that.

This is not Proxmox

Proxmox has snapshot scheduling built into the GUI. Plain libvirt on Fedora does not — you script it with virsh and schedule it with a systemd timer. That's what this page covers. For Proxmox VMs, use Proxmox Backup Server instead.

The two kinds of libvirt snapshots

This distinction is the whole ballgame, and it's the reason my manual snapshots ate so much disk.

Internal (what virt-manager makes) External disk-only (what the script makes)
RAM state Saved to a .save file (8–16 GB per snapshot) Not saved — VM must be running, snapshot is disk only
Disk New qcow2 layer inside the chain New qcow2 overlay, chain managed by libvirt
Restore Resumes the VM exactly as it was Reverts disk; VM reboots like a power loss
Cost per snapshot Huge Small diff

External snapshots are the right shape for automation: cheap, scriptable, and pruning one of them is a single virsh command.

Use qemu:///system

virsh list with no URI often shows an empty list on a desktop Fedora install — it's talking to the session instance of libvirt, not the system one that actually owns your VMs. Every command below uses -c qemu:///system explicitly.

The script

~/bin/vm-snapshot.sh. It takes a disk-only snapshot of each VM, then deletes the oldest ones past a retention count. Snapshots it creates are named auto-<ISO timestamp>; anything without that prefix (your manual snapshots) is never touched.

#!/usr/bin/env bash
# Automatic libvirt external snapshots + pruning.
# Disk-only (no RAM image) so each snapshot is a small qcow2 diff, not a 15GB .save.
# Only touches snapshots named auto-* — manual snapshots are never pruned.
set -uo pipefail

URI="qemu:///system"
VMS=("OpenClaw native ubuntu24.04" "ubuntu24.04 Hermes Home")
KEEP="${KEEP:-4}"   # keep the newest N auto-snapshots per VM
PREFIX="auto-"
LOG="$HOME/.local/state/vm-snapshots.log"
TS="$(date +%Y-%m-%dT%H:%M:%S)"

log() { echo "$(date '+%F %T') $*" >> "$LOG"; }

for vm in "${VMS[@]}"; do
    state=$(virsh -c "$URI" domstate "$vm" 2>/dev/null) || { log "SKIP $vm (not found)"; continue; }
    if [ "$state" != "running" ]; then
        log "SKIP $vm (state=$state; external snapshots need a running domain)"
        continue
    fi

    name="${PREFIX}${TS}"
    # Try fs-quiesced (needs qemu-guest-agent in the guest); fall back to plain.
    if virsh -c "$URI" snapshot-create-as --domain "$vm" --name "$name" \
            --disk-only --quiesce --atomic >>"$LOG" 2>&1; then
        log "OK   $vm -> $name (quiesced)"
    elif virsh -c "$URI" snapshot-create-as --domain "$vm" --name "$name" \
            --disk-only --atomic >>"$LOG" 2>&1; then
        log "OK   $vm -> $name (crash-consistent, no guest agent)"
    else
        log "FAIL $vm -> $name"
        continue
    fi

    # Prune: delete oldest auto-* snapshots beyond KEEP (lexicographic == chronological
    # thanks to the ISO timestamp in the name). Deleting an in-chain external snapshot
    # makes libvirt blockcommit the child into it first — safe, just IO-heavy.
    mapfile -t autos < <(virsh -c "$URI" snapshot-list --domain "$vm" --name 2>/dev/null \
                         | sed 's/^[[:space:]]*//' | grep "^${PREFIX}" | sort)
    if [ "${#autos[@]}" -gt "$KEEP" ]; then
        for old in "${autos[@]:0:${#autos[@]}-$KEEP}"; do
            if virsh -c "$URI" snapshot-delete --domain "$vm" --snapshotname "$old" >>"$LOG" 2>&1; then
                log "PRUNE $vm deleted $old"
            else
                log "PRUNE-FAIL $vm $old"
            fi
        done
    fi
done
chmod +x ~/bin/vm-snapshot.sh

Test it once by hand before scheduling anything:

~/bin/vm-snapshot.sh
tail ~/.local/state/vm-snapshots.log
virsh -c qemu:///system snapshot-list --domain "your vm name" --tree

The schedule

A systemd user timer, so no root and no crontab. User timers only survive logout if lingering is on:

loginctl enable-linger $USER

~/.config/systemd/user/vm-snapshot.service:

[Unit]
Description=Automatic libvirt external snapshots for agent VMs

[Service]
Type=oneshot
ExecStart=%h/bin/vm-snapshot.sh

~/.config/systemd/user/vm-snapshot.timer:

[Unit]
Description=Daily automatic VM snapshots + pruning

[Timer]
OnCalendar=*-*-* 03:30:00
RandomizedDelaySec=900
Persistent=true

[Install]
WantedBy=timers.target

Persistent=true matters on a desktop — if the box was off at 03:30, the job runs at next boot instead of silently skipping.

systemctl --user daemon-reload
systemctl --user enable --now vm-snapshot.timer
systemctl --user list-timers vm-snapshot.timer

Verifying

# what exists, and in what order (source of truth)
virsh -c qemu:///system snapshot-list --domain "your vm name" --tree

# which disk file the running VM is actually writing to
virsh -c qemu:///system domblklist "your vm name"

# the job's own log
tail ~/.local/state/vm-snapshots.log

# force a run right now
systemctl --user start vm-snapshot.service

vol-list doesn't show everything

virsh vol-list does not reliably list the active external-snapshot overlay files. Use snapshot-list --tree and domblklist to see the real chain.

Gotchas that bit me

  • snapshot-list --name pads names with leading spaces. Grepping that output without stripping whitespace silently matches nothing, and pruning becomes a no-op that looks like it worked. The script pipes through sed 's/^[[:space:]]*//' first.
  • Timestamps need seconds. With minute granularity, a manual test run in the same minute as a scheduled one collides with the existing overlay file and fails with external snapshot file already exists. It fails safe (no corruption), but it fails.
  • Deleting a mid-chain snapshot is not instant. libvirt blockcommits the child layer into the parent first — that's copying gigabytes. On my box a 22 GB layer took about three minutes; the VMs keep running the whole time.
  • --quiesce needs qemu-guest-agent inside the guest. Without it the snapshot is crash-consistent (the filesystem equivalent of pulling the plug). The script tries quiesced first and falls back automatically; install qemu-guest-agent in the guests to upgrade every future snapshot to a fs-consistent one.

Cockpit: visibility yes, creation no

The snapshots this script makes show up in Cockpit's Virtual Machines page (VM → Snapshots tab) like any other libvirt snapshot — Cockpit just asks libvirt, and libvirt doesn't care who made it. On Fedora that view needs cockpit-machines and libvirt-dbus installed, and your user in the libvirt group to see the system VMs.

But there's a difference worth knowing before you click anything in that UI:

  • Cockpit's "Create snapshot" button makes internal snapshots — the kind that save the VM's RAM to a .save file. That's the exact snapshot type that cost me 400 GB. The Cockpit dialog has no disk-only/external option. Look in Cockpit all you want; let the timer be the one creating snapshots.
  • Revert and delete work from the UI, with the same rules as virsh — reverting discards everything newer in the chain, and deleting a mid-chain snapshot starts a blockcommit that can churn for minutes on a big layer. The UI gives you no warning about either; it just starts doing it.

Restoring

# revert a running VM to a snapshot (disk reverts, VM reboots)
virsh -c qemu:///system snapshot-revert --domain "your vm name" --snapshotname auto-2026-10-08T03:30:00

Reverting deletes newer snapshots

snapshot-revert on an external snapshot chain discards everything after the snapshot you revert to. That includes the current active overlay. Know which auto- you're reverting to before you run it.