PVE 9.2.5 – Intermittent segfaults, cpuidle/RCU warnings and previous kernel panic on i7-14700K / ASUS B760

Infotelecom

Member
Jul 20, 2022
8
0
21
Hi,

We are troubleshooting an intermittent stability issue on a production Proxmox VE host and would appreciate some advice before changing hardware or firmware.

System:

  • Proxmox VE 9.2.5
  • Current kernel: 6.11.11-2-pve
  • Newer 7.0.14-6-pve is installed, but 6.11 is currently pinned because we previously experienced a kernel panic
  • CPU: Intel Core i7-14700K
  • Motherboard: ASUS TUF GAMING B760-PLUS WIFI
  • BIOS: 1604 (15/12/2023)
  • Linux microcode: 0x12b
  • RAM: 128 GB DDR5, 4x32 GB Corsair CMK64GX5M2B5600C40, running at 4800 MT/s
  • 3x Kingston NV2 2 TB NVMe
  • 8 Windows VMs
  • Storage for VMs is LVM-thin
The host is currently running, but over the last weeks we have seen several different symptoms:

  • Previous kernel panic
  • Previous Bad page state
  • A KVM process once became stuck in D state inside io_uring_del_tctx_node
  • Repeated pvesh / pveproxy segfaults
  • Kernel warnings involving intel_idle, cpuidle and RCU
  • High I/O wait, usually around 20–25%



  • Some recent segfaults are:

    pvesh[...] segfault ... in perl ... likely on CPU 1 (core 0, socket 0)
    pvesh[...] segfault ... in perl ... likely on CPU 0 (core 0, socket 0)
    pveproxy worker[...] segfault ... in perl ... likely on CPU 0 (core 0, socket 0)



    Several of them fail at the same Perl location:

    Perl_sv_setsv_flags + 0x12f

    19872f: 41 8b 57 08 mov 0x8(%r15),%edxÇ






    We verified the Perl/PVE packages with dpkg -V and did not find corruption in Perl or the main PVE packages.

    We also captured this kernel warning:

    intel_idle leaked IRQ state
    WARNING: CPU: 0 PID: 0 at drivers/cpuidle/cpuidle.c:269 cpuidle_enter_state+0x3c3/0x460

    CPU: 0 PID: 0 Comm: swapper/0
    Hardware name: ASUS System Product Name/TUF GAMING B760-PLUS WIFI
    BIOS 1604 12/15/2023

    Call Trace:
    cpuidle_enter_state
    cpuidle_enter
    call_cpuidle
    do_idle
    cpu_startup_entry
    We have also seen warnings from:

    rcu_sched_clock_irq
    on CPU0.

    Earlier, many faults appeared around logical CPUs 10/11, which are siblings of the same physical core, so as a diagnostic measure we disabled both. They remain offline:

    CPU10: offline
    CPU11: offline
    However, the host later continued producing segfaults on CPU0/CPU1, so disabling CPU10/11 did not completely solve the issue.

    We tested CPU0 against CPU2 by forcing 300 pvesh queries on each:

    CPU0 OK=300 FAIL=0
    CPU2 OK=300 FAIL=0
    So the issue is intermittent and cannot currently be reproduced on demand.

    Storage checks​

    All three NVMe drives report:
    • SMART PASSED
    • 0 media/data integrity errors
    • 0 NVMe error log entries
    • Normal temperatures (~33–36°C)
      There are no obvious:

    NVMe timeout/reset
    PCIe AER errors
    MCE
    Machine Check
    EDAC errors
    The LVM-thin pools are active and healthy, with plenty of free data and metadata space.

    However, I/O wait is consistently high. We captured threads in D state waiting in functions such as:

    rq_qos_wait
    submit_bio_wait
    folio_wait_bit_common
    block_write_begin_int
    dm_thin_find_block
    The affected processes include:

    iou-wrk
    dm-thin
    kcopyd
    The block devices currently have:

    scheduler: none
    wbt_lat_usec: 2000
    The physical NVMe latency itself is generally low, so we suspect the I/O wait is occurring somewhere in the QEMU/io_uring/LVM-thin/device-mapper path rather than because the NVMe devices are failing.

    Separate VM issue​

    VM107 was recently cloned from VM106.

    All VMs answer QEMU Guest Agent except VM107:

    VM101 QGA OK
    VM102 QGA OK
    VM103 QGA OK
    VM104 QGA OK
    VM105 QGA OK
    VM106 QGA OK
    VM107 QGA FAIL
    VM201 QGA OK
    This is a separate problem that we plan to fix inside Windows when the user is disconnected.

    We also run an external monitoring system that polls the Proxmox API regularly, so there are many API requests. We do not currently believe this is the root cause of the host instability.

    Current thoughts / next steps​

    Our main suspects are currently:
    1. BIOS / firmware / Intel 14th-gen platform stability
    2. CPU / IMC
    3. RAM
    4. Kernel-specific issue
    5. QEMU/io_uring/LVM-thin interaction for the I/O wait
      We are planning a maintenance window to:
    • Update the ASUS BIOS to a current stable release
    • Apply Intel Default Settings
    • Keep the same 6.11.11-2-pve kernel initially, so we only change one variable
    • Run Memtest86+ for at least 4 passes
    • Configure kdump/crashkernel so that a future kernel panic can be properly captured
      Before doing that, does this pattern look familiar to anyone?

    In particular:
    • Have you seen similar random Perl/PVE segfaults with Intel 13th/14th-gen CPUs?
    • Could the intel_idle leaked IRQ state / RCU warnings indicate a firmware/CPU issue rather than a normal kernel bug?
    • Would you prioritize BIOS/microcode + Memtest before testing another kernel?
    • Regarding the I/O wait, does the combination of rq_qos_wait, dm-thin, io_uring and wbt_lat_usec=2000 suggest anything specific worth testing?
      Any advice on the safest next diagnostic step would be appreciated.