as response to myself with the higher power consumtion. Fixed it by installing the latest revision and a couple of reboots. No idea what cause it, but one core stayed stuck on Max Freq. Now its back to normal. (7.0.2-4-pve)
May 19 22:44:52 px1 QEMU[3776]: kvm: vfio_container_dma_map(0x595d88370dd0, 0xe10bb000, 0x1000, 0x7834682db000) = -28 (No space left on device)
May 19 22:44:52 px1 kernel: DMAR: DRHD: handling fault status reg 2
May 19 22:44:52 px1 kernel: DMAR: [DMA Read NO_PASID] Request device [04:10.0] fault addr 0xe2162000 [fault reason 0x06] PTE Read access is not set
May 19 22:44:52 px1 kernel: DMAR: DRHD: handling fault status reg 102
May 19 22:44:52 px1 kernel: DMAR: [DMA Read NO_PASID] Request device [04:10.0] fault addr 0xe20bb000 [fault reason 0x06] PTE Read access is not set
May 19 22:44:52 px1 QEMU[3776]: kvm: vfio_container_dma_map(0x595d88370dd0, 0xe10ba000, 0x1000, 0x7834682da000) = -28 (No space left on device)
as response to myself with the higher power consumtion. Fixed it by installing the latest revision and a couple of reboots. No idea what cause it, but one core stayed stuck on Max Freq. Now its back to normal. (7.0.2-4-pve)
+ 41.10% 0.05% [kernel] [k] entry_SYSCALL_64_after_hwframe
+ 41.05% 0.10% [kernel] [k] do_syscall_64
+ 40.59% 0.12% [kernel] [k] x64_sys_call
+ 27.84% 0.58% [kernel] [k] kvm_arch_vcpu_ioctl_run
+ 16.87% 1.50% [kernel] [k] __schedule
+ 15.73% 0.14% [kernel] [k] schedule
+ 12.36% 0.01% [kernel] [k] kvm_vcpu_halt
+ 12.29% 0.05% [kernel] [k] kvm_vcpu_block
+ 8.76% 0.09% [kernel] [k] do_idle
+ 7.49% 0.45% [kernel] [k] try_to_wake_up
+ 7.24% 0.08% [kernel] [k] __kvm_vcpu_kick
+ 7.21% 0.02% [kernel] [k] wake_up_process
+ 7.14% 0.11% [kernel] [k] rcuwait_wake_up
+ 5.68% 0.00% [kernel] [k] call_cpuidle
+ 5.65% 0.22% [kernel] [k] ttwu_do_activate
+ 5.60% 0.09% [kernel] [k] cpuidle_enter_state
+ 5.53% 0.00% libc.so.6 [.] ioctl
+ 5.51% 0.00% [unknown] [k] 0x70632d34365f3638
+ 5.51% 0.00% [unknown] [k] 0x00000000000001b8
+ 5.51% 0.00% [unknown] [k] 0x0000633e38f33140
+ 5.51% 0.00% libglib-2.0.so.0.8400.4 [.] g_free
+ 5.39% 0.00% libc.so.6 [.] 0x000075cf23379b7b
+ 5.36% 0.01% [kernel] [k] asm_exc_page_fault
+ 5.27% 0.00% [kernel] [k] __x64_sys_ioctl
+ 5.27% 0.00% [kernel] [k] kvm_vcpu_ioctl
+ 5.18% 0.03% [kernel] [k] exc_page_fault
+ 4.96% 0.11% [kernel] [k] do_user_addr_fault
+ 4.81% 0.02% [kernel] [k] acpi_idle_enter
+ 4.80% 0.01% [kernel] [k] cpuidle_enter
+ 4.72% 0.65% [kernel] [k] pick_next_task_fair
+ 4.65% 0.07% [kernel] [k] handle_mm_fault
+ 4.64% 0.21% [kernel] [k] enqueue_task
+ 4.56% 0.19% [kernel] [k] __handle_mm_fault
+ 4.55% 0.15% [kernel] [k] dequeue_task
+ 4.50% 0.00% perf [.] 0x000056d1ebedc3f4
+ 4.47% 0.00% perf [.] 0x000056d1ec0538b3
+ 4.36% 0.21% [kernel] [k] dequeue_task_fair
+ 4.14% 0.00% perf [.] 0x000056d1ebedcbe8
+ 4.11% 1.20% [kernel] [k] dequeue_entities
+ 4.09% 1.12% [kernel] [k] pv_native_safe_halt
root@p1 ~ # cat /etc/kernel/cmdline
root=ZFS=rpool/ROOT/pve-1 boot=zfs mitigations=off
root@p1 ~ # cat /proc/cmdline
initrd=\EFI\proxmox\7.0.2-6-pve\initrd.img-7.0.2-6-pve root=ZFS=rpool/ROOT/pve-1 boot=zfs mitigations=off
This looks like the same issue as Post #97Kernel: 7.0.2-6-pve
HW CPU: AMD GX-424CC
I am experiencing very high CPU usage after a kernel update.
There is a huge overhead in CPU usage in SY time.
perf output:
Code:+ 41.10% 0.05% [kernel] [k] entry_SYSCALL_64_after_hwframe + 41.05% 0.10% [kernel] [k] do_syscall_64 + 40.59% 0.12% [kernel] [k] x64_sys_call + 27.84% 0.58% [kernel] [k] kvm_arch_vcpu_ioctl_run + 16.87% 1.50% [kernel] [k] __schedule + 15.73% 0.14% [kernel] [k] schedule + 12.36% 0.01% [kernel] [k] kvm_vcpu_halt + 12.29% 0.05% [kernel] [k] kvm_vcpu_block + 8.76% 0.09% [kernel] [k] do_idle + 7.49% 0.45% [kernel] [k] try_to_wake_up + 7.24% 0.08% [kernel] [k] __kvm_vcpu_kick + 7.21% 0.02% [kernel] [k] wake_up_process + 7.14% 0.11% [kernel] [k] rcuwait_wake_up + 5.68% 0.00% [kernel] [k] call_cpuidle + 5.65% 0.22% [kernel] [k] ttwu_do_activate + 5.60% 0.09% [kernel] [k] cpuidle_enter_state + 5.53% 0.00% libc.so.6 [.] ioctl + 5.51% 0.00% [unknown] [k] 0x70632d34365f3638 + 5.51% 0.00% [unknown] [k] 0x00000000000001b8 + 5.51% 0.00% [unknown] [k] 0x0000633e38f33140 + 5.51% 0.00% libglib-2.0.so.0.8400.4 [.] g_free + 5.39% 0.00% libc.so.6 [.] 0x000075cf23379b7b + 5.36% 0.01% [kernel] [k] asm_exc_page_fault + 5.27% 0.00% [kernel] [k] __x64_sys_ioctl + 5.27% 0.00% [kernel] [k] kvm_vcpu_ioctl + 5.18% 0.03% [kernel] [k] exc_page_fault + 4.96% 0.11% [kernel] [k] do_user_addr_fault + 4.81% 0.02% [kernel] [k] acpi_idle_enter + 4.80% 0.01% [kernel] [k] cpuidle_enter + 4.72% 0.65% [kernel] [k] pick_next_task_fair + 4.65% 0.07% [kernel] [k] handle_mm_fault + 4.64% 0.21% [kernel] [k] enqueue_task + 4.56% 0.19% [kernel] [k] __handle_mm_fault + 4.55% 0.15% [kernel] [k] dequeue_task + 4.50% 0.00% perf [.] 0x000056d1ebedc3f4 + 4.47% 0.00% perf [.] 0x000056d1ec0538b3 + 4.36% 0.21% [kernel] [k] dequeue_task_fair + 4.14% 0.00% perf [.] 0x000056d1ebedcbe8 + 4.11% 1.20% [kernel] [k] dequeue_entities + 4.09% 1.12% [kernel] [k] pv_native_safe_halt
mitigations=off also does not resolve the issue.
Code:root@p1 ~ # cat /etc/kernel/cmdline root=ZFS=rpool/ROOT/pve-1 boot=zfs mitigations=off root@p1 ~ # cat /proc/cmdline initrd=\EFI\proxmox\7.0.2-6-pve\initrd.img-7.0.2-6-pve root=ZFS=rpool/ROOT/pve-1 boot=zfs mitigations=off
The issue does not occur on kernel 6.17.
cat /sys/devices/system/clocksource/clocksource0/current_clocksource is hpet, and not tsc, then that is your problem, and the fix.no, read_hpet has 1% usage.This looks like the same issue as Post #97
Ifcat /sys/devices/system/clocksource/clocksource0/current_clocksourceishpet, and nottsc, then that is your problem, and the fix.
Is this the correct way to set it to force TSC?Update: Root cause found — HPET clocksource, not kernel 7.0
After further investigation, the high CPU usage was not caused by kernel 7.0 itself, but by the HPET (High Precision Event Timer) being used as the system clocksource.
Using perf top we identified that read_hpet was consuming ~60% of kernel CPU time. HPET is an older hardware timer that is expensive to read, and with multiple VMs polling the clocksource constantly (especially chronyd/NTP and PFSense), the kernel was overwhelmed with HPET reads.
The underlying issue was that TSC (Time Stamp Counter) was being flagged as unstable and the kernel fell back to HPET. This was confirmed by:
cat /sys/devices/system/clocksource/clocksource0/current_clocksource
hpet
The fix was to force TSC as reliable in /etc/kernel/cmdline:
root=ZFS=rpool/ROOT/pve-1 boot=zfs tsc=reliable
Followed by running proxmox-boot-tool kernel pin 7.0.0-3-pve
After rebooting with TSC as the active clocksource, CPU usage dropped from ~85% system time to ~7% system time, and kernel 7.0.0-3-pve is now running perfectly with all VMs at normal CPU levels.
Hardware: AMD Ryzen 3 PRO 2200GE
Note: This issue was present on kernel 6.17 as well, but kernel 7.0 made it significantly worse, which is what triggered the investigation.
agent: 1,fstrim_cloned_disks=1
balloon: 0
bios: ovmf
boot: order=scsi0;ide0
cores: 4
cpu: x86-64-v3
machine: pc-q35-11.0
memory: 10240
name: prod-rs2v-srv2
net0: virtio=removed,bridge=vlan0,queues=4
numa: 1
ostype: win11
scsi0: vm_data:vm-108-disk-1,cache=writeback,discard=on,iothread=1,size=65G,ssd=1
scsi1: vm_data:vm-108-disk-3,cache=writeback,discard=on,iothread=1,size=65G,ssd=1
scsihw: virtio-scsi-single
smbios1: uuid=removed
sockets: 2
vmgenid: removed
root@prox001:~# qm config 108
agent: 1,fstrim_cloned_disks=1
balloon: 0
bios: ovmf
boot: order=scsi0;ide0
cores: 4
cpu: x86-64-v3
efidisk0: vm_data:vm-108-disk-0,efitype=4m,ms-cert=2023k,pre-enrolled-keys=1,size=1M
ide0: none,media=cdrom
machine: pc-q35-11.0
memory: 10240
meta: creation-qemu=9.2.0,ctime=1745503025
name: prod-rs2v-srv2
net0: virtio=removed,bridge=vlan0,queues=4
numa: 1
ostype: win11
scsi0: vm_data:vm-108-disk-1,cache=writeback,discard=on,iothread=1,size=65G,ssd=1
scsi1: vm_data:vm-108-disk-3,cache=writeback,discard=on,iothread=1,size=65G,ssd=1
scsihw: virtio-scsi-single
sockets: 2
tpmstate0: vm_data:vm-108-disk-2,size=4M,version=v2.0
vga: qxl

I'm waiting for a response from the maintainers.@fiona: is there an estimated schedule for the upstream patch in core.c?
with QEMU 10.2, there was a switch to using io_uring for the IO thread event loops and the IO pressure/wait accounting is set via the io_uring subsystem now. It's a different kernel subsystem from before, so it's not unexpected if it's different.Everything seems to be working fine. However, since the reboot required to apply the update, I have noticed an increase in “IO Pressure Stall”.
The update correspond to the begin of the "red mountain"![]()
no, read_hpet has 1% usage.
The heavy load here is generated by syscalls related to virtualization and context switching.

We use essential cookies to make this site work, and optional cookies to enhance your experience.