Hi,
We are troubleshooting an intermittent stability issue on a production Proxmox VE host and would appreciate some advice before changing hardware or firmware.
System:
We are troubleshooting an intermittent stability issue on a production Proxmox VE host and would appreciate some advice before changing hardware or firmware.
System:
- Proxmox VE 9.2.5
- Current kernel: 6.11.11-2-pve
- Newer 7.0.14-6-pve is installed, but 6.11 is currently pinned because we previously experienced a kernel panic
- CPU: Intel Core i7-14700K
- Motherboard: ASUS TUF GAMING B760-PLUS WIFI
- BIOS: 1604 (15/12/2023)
- Linux microcode: 0x12b
- RAM: 128 GB DDR5, 4x32 GB Corsair CMK64GX5M2B5600C40, running at 4800 MT/s
- 3x Kingston NV2 2 TB NVMe
- 8 Windows VMs
- Storage for VMs is LVM-thin
- Previous kernel panic
- Previous Bad page state
- A KVM process once became stuck in D state inside io_uring_del_tctx_node
- Repeated pvesh / pveproxy segfaults
- Kernel warnings involving intel_idle, cpuidle and RCU
- High I/O wait, usually around 20–25%
- Some recent segfaults are:
pvesh[...] segfault ... in perl ... likely on CPU 1 (core 0, socket 0)
pvesh[...] segfault ... in perl ... likely on CPU 0 (core 0, socket 0)
pveproxy worker[...] segfault ... in perl ... likely on CPU 0 (core 0, socket 0)
Several of them fail at the same Perl location:
Perl_sv_setsv_flags + 0x12f
19872f: 41 8b 57 08 mov 0x8(%r15),%edxÇ
We verified the Perl/PVE packages with dpkg -V and did not find corruption in Perl or the main PVE packages.
We also captured this kernel warning:
intel_idle leaked IRQ state
WARNING: CPU: 0 PID: 0 at drivers/cpuidle/cpuidle.c:269 cpuidle_enter_state+0x3c3/0x460
CPU: 0 PID: 0 Comm: swapper/0
Hardware name: ASUS System Product Name/TUF GAMING B760-PLUS WIFI
BIOS 1604 12/15/2023
Call Trace:
cpuidle_enter_state
cpuidle_enter
call_cpuidle
do_idle
cpu_startup_entry
We have also seen warnings from:
rcu_sched_clock_irq
on CPU0.
Earlier, many faults appeared around logical CPUs 10/11, which are siblings of the same physical core, so as a diagnostic measure we disabled both. They remain offline:
CPU10: offline
CPU11: offline
However, the host later continued producing segfaults on CPU0/CPU1, so disabling CPU10/11 did not completely solve the issue.
We tested CPU0 against CPU2 by forcing 300 pvesh queries on each:
CPU0 OK=300 FAIL=0
CPU2 OK=300 FAIL=0
So the issue is intermittent and cannot currently be reproduced on demand.
Storage checks
All three NVMe drives report:
- SMART PASSED
- 0 media/data integrity errors
- 0 NVMe error log entries
- Normal temperatures (~33–36°C)
There are no obvious:
NVMe timeout/reset
PCIe AER errors
MCE
Machine Check
EDAC errors
The LVM-thin pools are active and healthy, with plenty of free data and metadata space.
However, I/O wait is consistently high. We captured threads in D state waiting in functions such as:
rq_qos_wait
submit_bio_wait
folio_wait_bit_common
block_write_begin_int
dm_thin_find_block
The affected processes include:
iou-wrk
dm-thin
kcopyd
The block devices currently have:
scheduler: none
wbt_lat_usec: 2000
The physical NVMe latency itself is generally low, so we suspect the I/O wait is occurring somewhere in the QEMU/io_uring/LVM-thin/device-mapper path rather than because the NVMe devices are failing.
Separate VM issue
VM107 was recently cloned from VM106.
All VMs answer QEMU Guest Agent except VM107:
VM101 QGA OK
VM102 QGA OK
VM103 QGA OK
VM104 QGA OK
VM105 QGA OK
VM106 QGA OK
VM107 QGA FAIL
VM201 QGA OK
This is a separate problem that we plan to fix inside Windows when the user is disconnected.
We also run an external monitoring system that polls the Proxmox API regularly, so there are many API requests. We do not currently believe this is the root cause of the host instability.
Current thoughts / next steps
Our main suspects are currently:
- BIOS / firmware / Intel 14th-gen platform stability
- CPU / IMC
- RAM
- Kernel-specific issue
- QEMU/io_uring/LVM-thin interaction for the I/O wait
We are planning a maintenance window to:
- Update the ASUS BIOS to a current stable release
- Apply Intel Default Settings
- Keep the same 6.11.11-2-pve kernel initially, so we only change one variable
- Run Memtest86+ for at least 4 passes
- Configure kdump/crashkernel so that a future kernel panic can be properly captured
Before doing that, does this pattern look familiar to anyone?
In particular:
- Have you seen similar random Perl/PVE segfaults with Intel 13th/14th-gen CPUs?
- Could the intel_idle leaked IRQ state / RCU warnings indicate a firmware/CPU issue rather than a normal kernel bug?
- Would you prioritize BIOS/microcode + Memtest before testing another kernel?
- Regarding the I/O wait, does the combination of rq_qos_wait, dm-thin, io_uring and wbt_lat_usec=2000 suggest anything specific worth testing?
Any advice on the safest next diagnostic step would be appreciated.