Increased latency on disks since 9.2.6 and associated 7.0.x kernel

weppa

Renowned Member
Nov 2, 2015
19
7
68
I have a couple of proxmox VE hosts, running 9.2.4 with 7.0 kernel.
These systems are updated regularly (like once a month) with proper reboot

For the last update (9.2.4 to 9.2.6 - on the 7.0 kernel) I have something weird happening that trigger alerts in MUNIN (yes i'm using munin)
The different hosts run on comptely different hardware and datacenter, and all of them display the same behavior, which never happened in the past

Screenshot_3Kz1zKCTPn.png

Basically it's like the disks are lagging, on dev/loop, WRITE IO
The screenshots are on "loop" but the graphs are the same on the real nvme (physical disks) graphs, or any other disk graph.

This behaviour can be seen on
- the host itself (screenshot)
- linux vms
- LXCs

So it propagates on all subsystems

All hosts run on MDADM SSD raid (enterprise) arrays, all systems are not loaded at all (less than n10% CPU usage) on modern and very fast hardware (EPYC)

I'll dig deeper in the hosts logs to try to pinpoint a root cause, but clearly this is specific to 9.2.6 (possible 9.2.5 whihc I skipped) and started happening on a variety of hardware the minut I rebooted the hosts.

Any advice ?
 
Last edited:

Commonalities:
  • mdraid
  • Munin
 
Thanks @Neobin !

In the meantime
proxmox-boot-tool kernel pin 7.0.14-5-pve
proxmox-boot-tool refresh

UPDATE : I can confirm the problem DISAPPEARS with kernel 7.0.14-5

So it's specific to kernels after this one.
 
Last edited: