The 56 GB that wasn't there¶
The GPU VM on the Powerspec box runs a 27B model that keeps its experts resident in RAM, so I wanted more RAM for it. The host has 64 GB, the VM had 50, and I figured I had room to grow. I bumped the VM to 56 GB and the host OOM'd. Twice.
This is the story of why, and it's mostly a story about assumptions I didn't check.
The assumption I had wrong¶
I believed the passed-through GPU carved out a chunk of host RAM for itself, something like 4.8 GB of CMA reserved for the passthrough device. That's a common thing to read in passthrough guides, and it sounded right, so I built my sizing math on top of it.
It was false. On this host:
dmesg says "0 pages cma reserved". The GPU holds no host RAM at all. The real cost is just the normal host OS baseline — kernel, Proxmox services, the QEMU process itself.
What the numbers actually said¶
I stopped estimating and started measuring.
- Kernel sees 63.38 GiB physical.
MemTotal(the usable pool) is 62.18 GiB. The ~1.2 GiB gap is kernel/BIOS/firmware reservation, and that happens on bare Ubuntu too, it's not a Proxmox thing. - The host baseline (MemTotal − MemAvailable − kvm RSS) measured about 3.28 GiB. Squarely in the normal 2–4 GiB range for a host OS.
- Guest RAM is charged one-to-one against host RSS. The
kvmprocess RSS equals the full guest allocation, not what the guest happens to be using, because ballooning was off and the guest keeps its page cache resident. A 50 GiB guest means a 50 GiB RSS, full stop.
So the sizing rule, measured:
56 GB fails that rule before you even start. The guest itself only sees ~53 GiB usable, and the host has nothing left to breathe with. The practical ceiling on a 64 GiB board is 51200 MB (50 GiB), which is exactly where the VM had been running stable.
Ballooning is a polite request¶
The VM has a virtio-balloon device with free-page-reporting enabled, and I'd been thinking of it as OOM insurance. It isn't. The balloon is a request — the host asks the guest for memory back, and the guest can say no. Nothing was triggering it. A balloon sitting in the config does not stop a 56 GB guest from OOMing a 62 GB host.
The fix¶
The model needs about 49 GB for the experts plus ~6 GB for the rest, plus a guest OS on top, call it 52 GB. So the target wasn't crazy, the hardware was the wall.
The right answer was two more matching 32 GB sticks, reusing what we already own, which takes the ceiling problem away entirely. The honest catch is that filling all four slots means 2 DIMMs per channel. Arrow Lake does DDR5-6400 at one DIMM per channel; at two per channel the negotiated speed typically drops to around 4800–5600, and XMP often won't apply with four sticks. Worth it here because the GPU holds most of the experts resident anyway, and the RAM bandwidth only serves the CPU-side minority.
The reference numbers live on the Proxmox page. The story lives here.