Slow/hanging Proxmox consoles after upgrade to 9.2.20 (from 9.2.4)

baas54

Member
Nov 21, 2023
9
2
8
Symptom
After upgrading to Proxmox VE 9.2.20 (September 24, from 9.2.4), LXC/VM consoles in the web interface sometimes failed to start (termproxy ... failed: exit code 1, failed waiting for client: timed out), and overall interaction (console use, text highlighting/copy-paste, even SSH connection setup at times) was noticeably slower than before the upgrade.

Confirmed contributing causes
  • PBS VM (705, Proxmox Backup Server,aka pbs1) temporarily unreachable: pvestatd/pvedaemon were periodically checking pbs1's datastores, and those checks timed out (No route to host / Connection timed out), tying up worker processes — so other API tasks (like console requests) had to wait for a free worker. Resolved by fully stopping/starting VM 705.
  • Proxmox firewall (nftables) active at node level: deliberately enabled a week earlier for a strict rule on one LXC, with ACCEPT/ACCEPT policy at datacenter level. Disabling it gave a noticeable improvement, though not fully back to the pre-upgrade level.
Investigated and ruled out
  • Orphaned dtach sockets — present but not the core issue.
  • DNS/reverse-DNS resolution — fast, not the cause.
  • GSSAPI/SSH auth delay — disabled, not the cause.
  • Cluster leftovers (old corosync config, /etc/pve/nodes/) — clean; node is correctly single-node.
  • QEMU monitor socket lock on VM 705 — no other processes holding the socket.
  • Cron jobs / vzdump backup config — VM 705 is not in the backup job's VMID list.
  • PCIe correctable errors (r8169 NIC) — present since January, unrelated to this upgrade.
  • CIFS shares (TNAS, ProtonDrive) — now responding in milliseconds, no hang.
  • Ballooning setting on VM 705 — disabling had no effect.

Open loose end (unexplained, doesn't appear to be blocking)
VM 705 keeps logging qga command failed - unable to open monitor socket roughly every 60 seconds, despite agent: 0 in its config and a full stop/start plus host reboot. However, an strace at the moment of the call showed a fast, successful connect() — so this looks more like cosmetic log noise than an actual hang.


Conclusion
Two concrete, confirmed contributing factors were resolved (pbs1 reachability, nftables firewall off), with noticeable improvement. The remaining slowness likely matches a broader "sluggish webUI after upgrade to PVE 9" pattern
 
Last edited:
* please share `pveversion -v`
* please check (and optionally share) the journal after booting when the system is sluggish for any errors (check if the errors were present when the system was not sluggish as well - then that's likely not the cause)
 
An additional observation:
When logged in to an LXC container via the Proxmox web interface (8006) the console input lags sometimes a second. It is displayed after a while. This behavior is not the case when logged in via ssh.

When sluggish this message is in the logs every minute:
Sep 27 17:50:26 pve1 pvedaemon[1623]: VM 705 qga command failed - unable to open monitor socket

This is also new:
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.5: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.5: device [8086:a335] error status/mask=00000001/00002000
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.5: [ 0] RxErr (First)
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.5: AER: Error of this Agent is reported first
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.6: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.6: device [8086:a336] error status/mask=00000001/00002000
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.6: [ 0] RxErr (First)
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.6: AER: Error of this Agent is reported first
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
Sep 27 17:56:29 pve1 kernel: r8169 0000:03:00.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Transmitter ID)
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.0: device [8086:a334] error status/mask=00000001/00002000
Sep 27 17:56:29 pve1 kernel: r8169 0000:03:00.0: device [10ec:8168] error status/mask=00001001/00006000
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.0: [ 0] RxErr (First)
Sep 27 17:56:29 pve1 kernel: pcieport 0000:00:1d.0: AER: Error of this Agent is reported first
Sep 27 17:56:29 pve1 kernel: r8169 0000:03:00.0: [ 0] RxErr (First)
Sep 27 17:56:29 pve1 kernel: iwlwifi 0000:04:00.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
Sep 27 17:56:29 pve1 kernel: r8169 0000:03:00.0: [12] Timeout
Sep 27 17:56:29 pve1 kernel: iwlwifi 0000:04:00.0: device [8086:3165] error status/mask=00000001/00002000
Sep 27 17:56:29 pve1 kernel: iwlwifi 0000:04:00.0: [ 0] RxErr (First)
Sep 27 17:56:29 pve1 kernel: r8169 0000:02:00.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
Sep 27 17:56:29 pve1 kernel: r8169 0000:02:00.0: device [10ec:8168] error status/mask=00000001/00006000
Sep 27 17:56:29 pve1 kernel: r8169 0000:02:00.0: [ 0] RxErr (First)

A new kernel stricter with PCIe devices perhaps?
These messages occur a few times per 24 hours.


pveversion -v
proxmox-ve: 9.2.0 (running kernel: 7.0.14-19-pve)
pve-manager: 9.2.20 (running version: 9.2.20/49318c671b82f31e)
proxmox-kernel-helper: 9.2.0
proxmox-kernel-7.0.14-19-pve-signed: 7.0.14-19
proxmox-kernel-7.0: 7.0.14-19
proxmox-kernel-7.0.14-4-pve-signed: 7.0.14-4
proxmox-kernel-6.8.12-33-pve-signed: 6.8.12-33
proxmox-kernel-6.8: 6.8.12-33
ceph-fuse: 19.2.6-pve4
corosync: 3.1.10-pve3
criu: 4.1.1-1
dnsmasq: 2.91-1+deb13u2
frr-pythontools: 10.6.1-1+pve3
ifupdown2: 3.3.0-1+pmx12
intel-microcode: 3.20251111.1~deb13u1
ksm-control-daemon: 1.5-1
libjs-extjs: 7.0.0-7
libproxmox-acme-perl: 1.7.2
libproxmox-backup-qemu0: 2.0.3
libproxmox-rs-perl: 0.4.1
libpve-access-control: 9.1.2
libpve-apiclient-perl: 3.4.3
libpve-cluster-api-perl: 9.1.6
libpve-cluster-perl: 9.1.6
libpve-common-perl: 9.2.2
libpve-guest-common-perl: 6.0.5
libpve-http-server-perl: 6.0.5
libpve-network-perl: 1.6.7
libpve-notify-perl: 9.1.6
libpve-rs-perl: 0.15.3
libpve-storage-perl: 9.1.10
libspice-server1: 0.15.2-1+b1
lvm2: 2.03.31-2+pmx1
lxc-pve: 7.0.0-2
lxcfs: 7.0.0-pve1
novnc-pve: 1.7.0-2
openvswitch-switch: 3.5.0-1+b1
proxmox-backup-client: 4.2.6-1
proxmox-backup-file-restore: 4.2.6-1
proxmox-backup-restore-image: 1.0.0
proxmox-enterprise-support-keyring: 1.1
proxmox-firewall: 1.2.3
proxmox-kernel-helper: 9.2.0
proxmox-mail-forward: 1.0.3
proxmox-mini-journalreader: 1.7
proxmox-offline-mirror-helper: 0.7.4
proxmox-widget-toolkit: 5.2.10
pve-cluster: 9.1.6
pve-container: 6.1.14
pve-docs: 9.2.12
pve-edk2-firmware: 4.2026.08-1
pve-esxi-import-tools: 1.0.1
pve-firewall: 6.0.6
pve-firmware: 3.18-6
pve-ha-manager: 5.2.5
pve-i18n: 3.10.0
pve-qemu-kvm: 11.0.3-3
pve-xtermjs: 6.0.0-2
qemu-server: 9.2.8
smartmontools: 7.5-pve2
spiceterm: 3.4.2
swtpm: 0.8.0+pve3
vncterm: 1.9.2
zfsutils-linux: 2.4.4-pve1
 
Try with "pcie_aspm=off" for a quick test. My guess would be, that what you are seeing are classic ASPM issues. Kernel tries to use ASPM, device/platform cannot correctly handle it ==> kernel thinks, that a transfer issue took place, while in reality the device was saving power and didn't respond.

Let's see what happens with disabled ASPM. Could be a platform Erratum.
 
Last edited:
pcei_aspm=off has no effect.
I did some additional checking; the PCIe messages were sporadically present before the upgrade to 9.2.20,

Next I will try to pin my kernel to 6.8.12-33-pve

to be continued..
 
No luck,same issues on 6.8.12-33-pve and pcie_aspm=off

After running for a couple of hours first impression is that kernel 6.8.12-33-pve and pcie_aspm=off does not suffer from these hickups
 
Last edited:
You are on RaptorLake-S (judging by the PCI IDs) and are seeing issues where such a relatively old Kernel (in relation to the platform itself) is performing better, than a newer one?
What motherboard are we talking about? This sounds a bit like a BIOS issue. Already on the latest one, that is available for your board?
 
  • Like
Reactions: news