Debian guests freeze after ~3 days on PVE 9.2.21 / kernel 7.0.2-6-pve, host unaffected, no logs before hang

kevin2001

Member
Dec 29, 2021
8
0
21
25
Two of my three VMs freeze after roughly 3 days of uptime. SSH,
console and qemu-guest-agent all stop responding. The host stays
up and responsive the whole time. Only qm stop + qm start recovers
them.

Setup:
- proxmox-ve 9.2.0, pve-manager 9.2.21, kernel 7.0.2-6-pve (pinned)
- AMD host, single NVMe (LVM-thin), no RAID
- Realtek RTL8168h (r8169) in a VLAN-aware bridge vmbr1
- Host on VLAN 101, guests on VLAN 104
- VM 101: Debian, VirtFusion with nested KVM guests
- VM 102: Debian 13, kernel 6.12.111, Pterodactyl panel
- VM 100: same host, same bridge, never freezes

What I observed during a hang:
- qm status reports running, qm monitor responds, screendump works
- guest kernel log ends mid-stream at 11:02:02 with no panic, no
oops, no hung task, no OOM
- the kernel timestamp on the guest console stops advancing
entirely
- alt-sysrq-s/u/b via qm monitor had no effect
- qm agent ping times out

Already ruled out:
- split lock detection: lscpu reports no split_lock_detect,
journalctl --grep split_lock is empty
- the 7.0.6 to 7.0.14 freeze reports: I run 7.0.2-6 and have it
pinned
- memory pressure: guest had 3GB of 3.8GB free
- fail2ban: 0 banned, 0 failed
- disk space: 73% used
- fstrim: last run 2 days before the hang
- NIC: ethtool -S shows 0 errors, 0 drops, 0 missed, link stable
since boot, no r8169 messages in dmesg

Already configured for the next occurrence:
- ufw logging off and kernel.printk limited (console was flooded
with UFW BLOCK lines that buried any earlier output)
- serial console on both VMs with socat logging to a file on the
host

Questions:
1. Is there anything on the host side that can stop a guest this
abruptly while leaving the qemu process healthy and leaving no
trace in the guest?
2. Are there known issues with nested virtualization on AMD in
this kernel branch that could affect a sibling VM?
3. Anything else worth capturing before I hard stop the VM next
time?
 
1. please try with the most recent kernel version
2. provide pversion -v
3. provide the VM configs
 
  • Like
Reactions: fiona
Hi,
additionally to what @fabian said, the host's journal from around the time of the issue might also be interesting.
 
1. please try with the most recent kernel version
2. provide pversion -v
3. provide the VM configs
pveversion -v and the three VM configs below (SSH keys and
cloud-init usernames removed).

Note on memory: the host has 62GB total and 58GB is allocated to
guests (50GB to VM 101, 4GB each to 100 and 102). VM 101 runs
VirtFusion with nested KVM guests. All three use cpu: host.

Regarding the kernel: I'm currently pinned to 7.0.2-6-pve because
of the reports of VM freezes on 7.0.6 through 7.0.14. 7.0.14-22 is
installed. Is that version expected to contain a fix for those
reports, or would moving there risk introducing a second failure
mode? This is a production host.
 
You mentioned no OOM killer event entries, but I still would be curious if reducing the allocated memory to VM 101 a bit would make any difference. 58GB out of 62GB seems a bit tight.