Hi all,
I have a standalone PVE 9 node that warm-resets without warning roughly every
1-4 weeks, and the interval is getting shorter. There are several VMs and the server is in production.
After three weeks of digging I have ruled out everything I can think of and would appreciate ideas before I
start swapping hardware.
1. Server Config:
Supermicro X13DEI
2x INTEL(R) XEON(R) SILVER 4509Y - family 6, model 143, stepping 8 (CPUID 06-8f-08)
Microcode 0x2b000650 (from BIOS; intel-microcode package not installed)
BIOS 2.8 (10/08/2025), BMC 01.05.19 - current Supermicro bundle, reflashed 28 Sep 2026
Dual PSU, SSD: ZFS with NVMe boot (blk-mq), Broadcom tg3 + others
Standalone node, no cluster, no HA. 5 KVM guests.
2. The host performs a warm reset. No shutdown sequence, no panic, no oops. The
journal stops mid-sentence during completely routine activity and resumes at
the next boot about 2 minutes later.
24 Aug 2026 16:24:46 EEST -> up 16:26:55 (kernel 7.0.14-8-pve)
19 Sep 2026 00:22:16 EEST -> up 00:24:19 (kernel 7.0.14-12-pve)
29 Sep 2026 12:44:16 EEST -> up 12:46:38 (kernel 7.0.14-12-pve)
Intervals: 26 days, then 10 days. No pattern in time of day (afternoon, after
midnight, midday), no correlation with load, backups or any scheduled job.
For reference, a confirmed mains outage on this machine took ~3:13 to come
back, so the ~2:03 gap looks like a warm reset rather than a cold boot.
Diagnosis so far:
No Power loss
The BMC stayed powered through all three (30+ days BMC uptime across the
Sep 19 event). The SEL correctly logs "First AC Power on" for two genuine
outages (30 Apr, 20 Aug) but recorded NOTHING for these three resets.
SEL is healthy: 22 entries, Overflow false, 4074 free units.
Kernel panic / oops
kernel.panic=10 and kernel.panic_on_oops=1 were both active for the 29 Sep
event. ERST pstore is registered and working, and was EMPTY afterwards.
So the kernel apparently never got as far as panicking.
Watchdog / fencing
ipmitool mc watchdog get -> Timer Stopped, Expiration Flags: None
softdog loaded but never armed (no cluster, no HA)
Hardware errors
i10nm_edac loaded on all 8 controllers, rasdaemon running, zero MCE/EDAC
errors. No thermal events. Sensors nominal (12V 12.14, VCCIN 1.784 both
sockets, CPU1 43C / CPU2 56C).
ipmitool chassis restart_cause -> "unknown" for all three.
The node ran 35 days with no incident (27 May - 1 Jul 2026). Resets began
around 1 Jul. I am checking whether a kernel update landed in that window.
There are two kernel messages I am unsure about:
1. x86/CPU: Running old microcode
2. x86/split lock detection: #AC: crashing the kernel on kernel split_locks and warning on user-space split_locks
Microcode 0x2b000650 appears to be the revision from Intel's 20251111 release,
so it is not ancient, but the kernel's minimum-revision table is unhappy with
it. intel-microcode is not installed - I am about to add it.
I also have a case open with Supermicro about the SEL recording nothing.
Full diagnostic archive available on request - journals of both crashed boots,
IPMI SEL/sensors, lspci -nnk, DMI, module/taint details.
Any advice will be appreciated.
Thanks
I have a standalone PVE 9 node that warm-resets without warning roughly every
1-4 weeks, and the interval is getting shorter. There are several VMs and the server is in production.
After three weeks of digging I have ruled out everything I can think of and would appreciate ideas before I
start swapping hardware.
1. Server Config:
Supermicro X13DEI
2x INTEL(R) XEON(R) SILVER 4509Y - family 6, model 143, stepping 8 (CPUID 06-8f-08)
Microcode 0x2b000650 (from BIOS; intel-microcode package not installed)
BIOS 2.8 (10/08/2025), BMC 01.05.19 - current Supermicro bundle, reflashed 28 Sep 2026
Dual PSU, SSD: ZFS with NVMe boot (blk-mq), Broadcom tg3 + others
Standalone node, no cluster, no HA. 5 KVM guests.
2. The host performs a warm reset. No shutdown sequence, no panic, no oops. The
journal stops mid-sentence during completely routine activity and resumes at
the next boot about 2 minutes later.
24 Aug 2026 16:24:46 EEST -> up 16:26:55 (kernel 7.0.14-8-pve)
19 Sep 2026 00:22:16 EEST -> up 00:24:19 (kernel 7.0.14-12-pve)
29 Sep 2026 12:44:16 EEST -> up 12:46:38 (kernel 7.0.14-12-pve)
Intervals: 26 days, then 10 days. No pattern in time of day (afternoon, after
midnight, midday), no correlation with load, backups or any scheduled job.
For reference, a confirmed mains outage on this machine took ~3:13 to come
back, so the ~2:03 gap looks like a warm reset rather than a cold boot.
Diagnosis so far:
No Power loss
The BMC stayed powered through all three (30+ days BMC uptime across the
Sep 19 event). The SEL correctly logs "First AC Power on" for two genuine
outages (30 Apr, 20 Aug) but recorded NOTHING for these three resets.
SEL is healthy: 22 entries, Overflow false, 4074 free units.
Kernel panic / oops
kernel.panic=10 and kernel.panic_on_oops=1 were both active for the 29 Sep
event. ERST pstore is registered and working, and was EMPTY afterwards.
So the kernel apparently never got as far as panicking.
Watchdog / fencing
ipmitool mc watchdog get -> Timer Stopped, Expiration Flags: None
softdog loaded but never armed (no cluster, no HA)
Hardware errors
i10nm_edac loaded on all 8 controllers, rasdaemon running, zero MCE/EDAC
errors. No thermal events. Sensors nominal (12V 12.14, VCCIN 1.784 both
sockets, CPU1 43C / CPU2 56C).
ipmitool chassis restart_cause -> "unknown" for all three.
The node ran 35 days with no incident (27 May - 1 Jul 2026). Resets began
around 1 Jul. I am checking whether a kernel update landed in that window.
There are two kernel messages I am unsure about:
1. x86/CPU: Running old microcode
2. x86/split lock detection: #AC: crashing the kernel on kernel split_locks and warning on user-space split_locks
Microcode 0x2b000650 appears to be the revision from Intel's 20251111 release,
so it is not ancient, but the kernel's minimum-revision table is unhappy with
it. intel-microcode is not installed - I am about to add it.
I also have a case open with Supermicro about the SEL recording nothing.
Full diagnostic archive available on request - journals of both crashed boots,
IPMI SEL/sensors, lspci -nnk, DMI, module/taint details.
Any advice will be appreciated.
Thanks
Code:
root@DR-CVETY-SRV-01:/# journalctl -b -1 -o short-precise | grep -v timedate1 | tail -25
Sep 29 12:37:46.597652 DR-CVETY-SRV-01 systemd[1]: systemd-timedated.service: Deactivated successfully.
Sep 29 12:38:16.518059 DR-CVETY-SRV-01 systemd[1]: Starting systemd-timedated.service - Time & Date Service...
Sep 29 12:38:16.560780 DR-CVETY-SRV-01 systemd[1]: Started systemd-timedated.service - Time & Date Service.
Sep 29 12:38:46.573636 DR-CVETY-SRV-01 systemd[1]: systemd-timedated.service: Deactivated successfully.
Sep 29 12:39:16.519660 DR-CVETY-SRV-01 systemd[1]: Starting systemd-timedated.service - Time & Date Service...
Sep 29 12:39:16.561552 DR-CVETY-SRV-01 systemd[1]: Started systemd-timedated.service - Time & Date Service.
Sep 29 12:39:46.594890 DR-CVETY-SRV-01 systemd[1]: systemd-timedated.service: Deactivated successfully.
Sep 29 12:40:16.516778 DR-CVETY-SRV-01 systemd[1]: Starting systemd-timedated.service - Time & Date Service...
Sep 29 12:40:16.555065 DR-CVETY-SRV-01 systemd[1]: Started systemd-timedated.service - Time & Date Service.
Sep 29 12:40:46.559035 DR-CVETY-SRV-01 systemd[1]: systemd-timedated.service: Deactivated successfully.
Sep 29 12:41:16.518837 DR-CVETY-SRV-01 systemd[1]: Starting systemd-timedated.service - Time & Date Service...
Sep 29 12:41:16.558831 DR-CVETY-SRV-01 systemd[1]: Started systemd-timedated.service - Time & Date Service.
Sep 29 12:41:46.563557 DR-CVETY-SRV-01 systemd[1]: systemd-timedated.service: Deactivated successfully.
Sep 29 12:42:16.582721 DR-CVETY-SRV-01 systemd[1]: Starting systemd-timedated.service - Time & Date Service...
Sep 29 12:42:16.625817 DR-CVETY-SRV-01 systemd[1]: Started systemd-timedated.service - Time & Date Service.
Sep 29 12:42:46.631038 DR-CVETY-SRV-01 systemd[1]: systemd-timedated.service: Deactivated successfully.
Sep 29 12:42:53.962197 DR-CVETY-SRV-01 pveproxy[1007626]: worker exit
Sep 29 12:42:54.003674 DR-CVETY-SRV-01 pveproxy[8225]: worker 1007626 finished
Sep 29 12:42:54.003722 DR-CVETY-SRV-01 pveproxy[8225]: starting 1 worker(s)
Sep 29 12:42:54.007863 DR-CVETY-SRV-01 pveproxy[8225]: worker 1041684 started
Sep 29 12:43:16.530851 DR-CVETY-SRV-01 systemd[1]: Starting systemd-timedated.service - Time & Date Service...
Sep 29 12:43:16.578499 DR-CVETY-SRV-01 systemd[1]: Started systemd-timedated.service - Time & Date Service.
Sep 29 12:43:46.583258 DR-CVETY-SRV-01 systemd[1]: systemd-timedated.service: Deactivated successfully.
Sep 29 12:44:16.520102 DR-CVETY-SRV-01 systemd[1]: Starting systemd-timedated.service - Time & Date Service...
Sep 29 12:44:16.563185 DR-CVETY-SRV-01 systemd[1]: Started systemd-timedated.service - Time & Date Service.
root@DR-CVETY-SRV-01:/# grep -H . /sys/module/*/taint 2>/dev/null | grep -v ':$'
/sys/module/file_protector/taint:OE
/sys/module/snapapi26/taint:OE
/sys/module/spl/taint:O
/sys/module/zfs/taint:PO
root@DR-CVETY-SRV-01:/# cat /proc/sys/kernel/tainted
12293
root@DR-CVETY-SRV-01:/# cat /proc/cmdline
initrd=\EFI\proxmox\7.0.14-12-pve\initrd.img-7.0.14-12-pve root=ZFS=rpool/ROOT/pve-1 boot=zfs
root@DR-CVETY-SRV-01:/# grep -m1 microcode /proc/cpuinfo
microcode : 0x2b000650
root@DR-CVETY-SRV-01:/# pveversion -v
proxmox-ve: 9.2.0 (running kernel: 7.0.14-12-pve)
pve-manager: 9.2.10 (running version: 9.2.10/43df2e01f27a1a19)
proxmox-kernel-helper: 9.2.0
proxmox-kernel-7.0: 7.0.14-12
proxmox-kernel-7.0.14-12-pve-signed: 7.0.14-12
proxmox-kernel-7.0.14-8-pve-signed: 7.0.14-8
proxmox-kernel-7.0.2-6-pve-signed: 7.0.2-6
proxmox-kernel-7.0.2-4-pve-signed: 7.0.2-4
proxmox-kernel-6.17: 6.17.13-21
proxmox-kernel-6.17.13-21-pve-signed: 6.17.13-21
proxmox-kernel-6.17.13-11-pve-signed: 6.17.13-11
proxmox-kernel-6.17.13-9-pve-signed: 6.17.13-9
proxmox-kernel-6.17.2-1-pve-signed: 6.17.2-1
ceph-fuse: 19.2.3-pve4
corosync: 3.1.10-pve3
criu: 4.1.1-1
frr-pythontools: 10.6.1-1+pve3
ifupdown2: 3.3.0-1+pmx12
intel-microcode: 3.20251111.1~deb13u1
ksm-control-daemon: 1.5-1
libjs-extjs: 7.0.0-7
libproxmox-acme-perl: 1.7.2
libproxmox-backup-qemu0: 2.0.2
libproxmox-rs-perl: 0.4.1
libpve-access-control: 9.1.1
libpve-apiclient-perl: 3.4.2
libpve-cluster-api-perl: 9.1.6
libpve-cluster-perl: 9.1.6
libpve-common-perl: 9.2.1
libpve-guest-common-perl: 6.0.5
libpve-http-server-perl: 6.0.5
libpve-network-perl: 1.6.7
libpve-notify-perl: 9.1.6
libpve-rs-perl: 0.15.3
libpve-storage-perl: 9.1.8
libspice-server1: 0.15.2-1+b1
lvm2: 2.03.31-2+pmx1
lxc-pve: 7.0.0-2
lxcfs: 7.0.0-pve1
novnc-pve: 1.7.0-2
proxmox-backup-client: 4.2.5-1
proxmox-backup-file-restore: 4.2.5-1
proxmox-backup-restore-image: 1.0.0
proxmox-enterprise-support-keyring: 1.1
proxmox-firewall: 1.2.3
proxmox-kernel-helper: 9.2.0
proxmox-mail-forward: 1.0.3
proxmox-mini-journalreader: 1.7
proxmox-offline-mirror-helper: 0.7.4
proxmox-widget-toolkit: 5.2.7
pve-cluster: 9.1.6
pve-container: 6.1.13
pve-docs: 9.2.4
pve-edk2-firmware: 4.2025.05-3
pve-esxi-import-tools: 1.0.1
pve-firewall: 6.0.5
pve-firmware: 3.18-5
pve-ha-manager: 5.2.5
pve-i18n: 3.10.0
pve-qemu-kvm: 11.0.3-2
pve-xtermjs: 6.0.0-2
qemu-server: 9.2.4
smartmontools: 7.5-pve2
spiceterm: 3.4.2
swtpm: 0.8.0+pve3
vncterm: 1.9.2
zfsutils-linux: 2.4.3-pve1