Hey folks - I am experiencing a complete server.. "hang?" about every other month, with the machine not even responding to pings and an HDMI display attached to the box not showing the usual proxmox "configure via website" banner (no signal at all). A power cycle clears it right up.
I have very little evidence to back it up, but I believe it's related to the fact that one of the VMs is not running locally on LVM but over ISCSI to an attached NAS (the rest are local LVM). The reason I'm leaning this way is that there are 2 messages roughly around that time - a proxmox package update and a network error with the NAS unreachable (though the NAS seems to think it's OK and the unifi dashboard claims the same)
While I think neither should actually cause the host to become unresponsive, what techniques/logs would I go looking in to root cause a relatively rare event like this? There's nothing critical on that VM so I will probably migrate it anyway at some point, but want to see if I'm leaning in the right direction
Environment information:
I have very little evidence to back it up, but I believe it's related to the fact that one of the VMs is not running locally on LVM but over ISCSI to an attached NAS (the rest are local LVM). The reason I'm leaning this way is that there are 2 messages roughly around that time - a proxmox package update and a network error with the NAS unreachable (though the NAS seems to think it's OK and the unifi dashboard claims the same)
While I think neither should actually cause the host to become unresponsive, what techniques/logs would I go looking in to root cause a relatively rare event like this? There's nothing critical on that VM so I will probably migrate it anyway at some point, but want to see if I'm leaning in the right direction
Environment information:
Kernel Version Linux 6.2.16-15-pve #1 SMP PREEMPT_DYNAMIC PMX 6.2.16-15 (2023-09-28T13:53Z) |
PVE Manager Version pve-manager/8.0.4/d258a813cfa6b390 |