Two of my three VMs freeze after roughly 3 days of uptime. SSH,
console and qemu-guest-agent all stop responding. The host stays
up and responsive the whole time. Only qm stop + qm start recovers
them.
Setup:
- proxmox-ve 9.2.0, pve-manager 9.2.21, kernel 7.0.2-6-pve (pinned)
- AMD host, single NVMe (LVM-thin), no RAID
- Realtek RTL8168h (r8169) in a VLAN-aware bridge vmbr1
- Host on VLAN 101, guests on VLAN 104
- VM 101: Debian, VirtFusion with nested KVM guests
- VM 102: Debian 13, kernel 6.12.111, Pterodactyl panel
- VM 100: same host, same bridge, never freezes
What I observed during a hang:
- qm status reports running, qm monitor responds, screendump works
- guest kernel log ends mid-stream at 11:02:02 with no panic, no
oops, no hung task, no OOM
- the kernel timestamp on the guest console stops advancing
entirely
- alt-sysrq-s/u/b via qm monitor had no effect
- qm agent ping times out
Already ruled out:
- split lock detection: lscpu reports no split_lock_detect,
journalctl --grep split_lock is empty
- the 7.0.6 to 7.0.14 freeze reports: I run 7.0.2-6 and have it
pinned
- memory pressure: guest had 3GB of 3.8GB free
- fail2ban: 0 banned, 0 failed
- disk space: 73% used
- fstrim: last run 2 days before the hang
- NIC: ethtool -S shows 0 errors, 0 drops, 0 missed, link stable
since boot, no r8169 messages in dmesg
Already configured for the next occurrence:
- ufw logging off and kernel.printk limited (console was flooded
with UFW BLOCK lines that buried any earlier output)
- serial console on both VMs with socat logging to a file on the
host
Questions:
1. Is there anything on the host side that can stop a guest this
abruptly while leaving the qemu process healthy and leaving no
trace in the guest?
2. Are there known issues with nested virtualization on AMD in
this kernel branch that could affect a sibling VM?
3. Anything else worth capturing before I hard stop the VM next
time?
console and qemu-guest-agent all stop responding. The host stays
up and responsive the whole time. Only qm stop + qm start recovers
them.
Setup:
- proxmox-ve 9.2.0, pve-manager 9.2.21, kernel 7.0.2-6-pve (pinned)
- AMD host, single NVMe (LVM-thin), no RAID
- Realtek RTL8168h (r8169) in a VLAN-aware bridge vmbr1
- Host on VLAN 101, guests on VLAN 104
- VM 101: Debian, VirtFusion with nested KVM guests
- VM 102: Debian 13, kernel 6.12.111, Pterodactyl panel
- VM 100: same host, same bridge, never freezes
What I observed during a hang:
- qm status reports running, qm monitor responds, screendump works
- guest kernel log ends mid-stream at 11:02:02 with no panic, no
oops, no hung task, no OOM
- the kernel timestamp on the guest console stops advancing
entirely
- alt-sysrq-s/u/b via qm monitor had no effect
- qm agent ping times out
Already ruled out:
- split lock detection: lscpu reports no split_lock_detect,
journalctl --grep split_lock is empty
- the 7.0.6 to 7.0.14 freeze reports: I run 7.0.2-6 and have it
pinned
- memory pressure: guest had 3GB of 3.8GB free
- fail2ban: 0 banned, 0 failed
- disk space: 73% used
- fstrim: last run 2 days before the hang
- NIC: ethtool -S shows 0 errors, 0 drops, 0 missed, link stable
since boot, no r8169 messages in dmesg
Already configured for the next occurrence:
- ufw logging off and kernel.printk limited (console was flooded
with UFW BLOCK lines that buried any earlier output)
- serial console on both VMs with socat logging to a file on the
host
Questions:
1. Is there anything on the host side that can stop a guest this
abruptly while leaving the qemu process healthy and leaving no
trace in the guest?
2. Are there known issues with nested virtualization on AMD in
this kernel branch that could affect a sibling VM?
3. Anything else worth capturing before I hard stop the VM next
time?