Updating Proxmox led to NVMe-Bug

coffee_engine

New Member
Nov 30, 2024
1
0
1
Hi all,

I have a server running on Proxmox, which uses four NVMe-Drives in a ZFS-Raid-Z2. Since I have recently updated my Proxmox, since then I have the Issue that the NVMe-Drives periodically go down, and the VMs running on that datastore are crashing. Typically the Issue occurs during a Backup of the VMs with the integrated backup-solution to a Proxmox-Backup-Server, but it also happens when another IO-Intensive task runs, such as a zpool scrub. Journalctl lists the following:

Code:
Nov 30 14:54:18 mars pmxcfs[2212]: [dcdb] notice: data verification successful
Nov 30 14:57:14 mars pvestatd[2328]: ocean: error fetching datastores - 500 read timeout
Nov 30 14:57:14 mars pvestatd[2328]: status update time (7.070 seconds)
Nov 30 14:57:24 mars pvestatd[2328]: ocean: error fetching datastores - 500 read timeout
Nov 30 14:57:24 mars pvestatd[2328]: status update time (7.068 seconds)
Nov 30 14:57:34 mars pvestatd[2328]: ocean: error fetching datastores - 500 read timeout
Nov 30 14:57:34 mars pvestatd[2328]: status update time (7.065 seconds)
Nov 30 14:57:35 mars kernel: nvme nvme3: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0xffff
Nov 30 14:57:35 mars kernel: nvme nvme3: Does your device have a faulty power saving mode enabled?
Nov 30 14:57:35 mars kernel: nvme nvme3: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug
Nov 30 14:57:35 mars kernel: nvme nvme0: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0xffff
Nov 30 14:57:35 mars kernel: nvme nvme0: Does your device have a faulty power saving mode enabled?
Nov 30 14:57:35 mars kernel: nvme nvme0: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug
Nov 30 14:57:35 mars kernel: nvme nvme1: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0xffff
Nov 30 14:57:35 mars kernel: nvme nvme1: Does your device have a faulty power saving mode enabled?
Nov 30 14:57:35 mars kernel: nvme nvme1: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug
Nov 30 14:57:35 mars kernel: nvme nvme2: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0xffff
Nov 30 14:57:35 mars kernel: nvme nvme2: Does your device have a faulty power saving mode enabled?
Nov 30 14:57:35 mars kernel: nvme nvme2: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug
Nov 30 14:57:35 mars kernel: nvme 0000:05:00.0: Unable to change power state from D3cold to D0, device inaccessible
Nov 30 14:57:35 mars kernel: nvme 0000:06:00.0: Unable to change power state from D3cold to D0, device inaccessible
Nov 30 14:57:35 mars kernel: nvme nvme1: Disabling device after reset failure: -19
Nov 30 14:57:35 mars kernel: nvme nvme2: Disabling device after reset failure: -19
Nov 30 14:57:35 mars kernel: nvme 0000:04:00.0: Unable to change power state from D3cold to D0, device inaccessible
Nov 30 14:57:35 mars kernel: nvme nvme0: Disabling device after reset failure: -19
Nov 30 14:57:35 mars kernel: nvme 0000:07:00.0: Unable to change power state from D3cold to D0, device inaccessible
Nov 30 14:57:35 mars kernel: nvme nvme3: Disabling device after reset failure: -19
Nov 30 14:57:35 mars kernel: I/O error, dev nvme2n1, sector 2064186456 op 0x1:(WRITE) flags 0x0 phys_seg 1 prio class 0
Nov 30 14:57:35 mars kernel: I/O error, dev nvme0n1, sector 1066834128 op 0x0:(READ) flags 0x4000 phys_seg 4 prio class 0
Nov 30 14:57:35 mars kernel: I/O error, dev nvme0n1, sector 1066844160 op 0x0:(READ) flags 0x0 phys_seg 18 prio class 0
Nov 30 14:57:35 mars kernel: zio pool=NVMePool vdev=/dev/disk/by-id/nvme-INTEL_SSDPE2KX020T8_PHLJ326100VP2P0BGN_1-part1 error=5 type=2 offset=1056855076864 size=4096 flags=1572992
Nov 30 14:57:35 mars kernel: zio pool=NVMePool vdev=/dev/disk/by-id/nvme-INTEL_SSDPE2KX020T8_PHLJ326100VP2P0BGN_1-part1 error=5 type=1 offset=546211336192 size=126976 flags=1074267264

I have already added "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" to the kernel-boot-parameters accordingly, but it does not help at all. I have also disabled ASPM in the BIOS, and tried updating the Kernel to 6.11.x, but to no avail. The only thing that seems to bring my server back to a useable state is periodically setting the following values to the PCI-Devices:

echo "on" > /sys/bus/pci/devices/<device_id>/power/control
echo 0 > /sys/bus/pci/devices/<device_id>/d3cold_allowed

. I have made sure that my BIOS and my Drive-Firmwares are Up to Date. I have also tried to update again to see if there are any fixes for this, but this is unfortunately not the case.

The Server is using the following hardware:

Mainboard: Supermicro X13SCH-F
CPU: Intel Xeon E-2478
Chipset: Intel C266
The affected Drives: Intel DC-P4510 (4 TB)
RAM: 128GB Kingston ECC

Proxmox-Version: 8.3.0

I would be really grateful if somebody could look into that, since it renders my Proxmox-Setup nearly unusable.

EDIT: Corrected the issue report.
 
Last edited:
I am using Supermicro X13SAE-F and I have a pair of nvme drives using zfs configuration. By running nvme_core.default_ps_max_latency_us=0 resolved my issue.

Looking at your log, it looks as though Proxmox is having issues with:
Nov 30 14:57:35 mars kernel: nvme nvme3: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug

I believe you can disable aspm from the bios iirc. Otherwise, I would suggest shutting down the system, unplug the power for 30 secs and then restart everything again (I have seen that actually helped to reset some of the hardware in the past).
 
Hi all,

I have a server running on Proxmox, which uses four NVMe-Drives in a ZFS-Raid-Z2. Since I have recently updated my Proxmox, since then I have the Issue that the NVMe-Drives periodically go down, and the VMs running on that datastore are crashing. Typically the Issue occurs during a Backup of the VMs with the integrated backup-solution to a Proxmox-Backup-Server, but it also happens when another IO-Intensive task runs, such as a zpool scrub. Journalctl lists the following:

Code:
Nov 30 14:54:18 mars pmxcfs[2212]: [dcdb] notice: data verification successful
Nov 30 14:57:14 mars pvestatd[2328]: ocean: error fetching datastores - 500 read timeout
Nov 30 14:57:14 mars pvestatd[2328]: status update time (7.070 seconds)
Nov 30 14:57:24 mars pvestatd[2328]: ocean: error fetching datastores - 500 read timeout
Nov 30 14:57:24 mars pvestatd[2328]: status update time (7.068 seconds)
Nov 30 14:57:34 mars pvestatd[2328]: ocean: error fetching datastores - 500 read timeout
Nov 30 14:57:34 mars pvestatd[2328]: status update time (7.065 seconds)
Nov 30 14:57:35 mars kernel: nvme nvme3: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0xffff
Nov 30 14:57:35 mars kernel: nvme nvme3: Does your device have a faulty power saving mode enabled?
Nov 30 14:57:35 mars kernel: nvme nvme3: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug
Nov 30 14:57:35 mars kernel: nvme nvme0: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0xffff
Nov 30 14:57:35 mars kernel: nvme nvme0: Does your device have a faulty power saving mode enabled?
Nov 30 14:57:35 mars kernel: nvme nvme0: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug
Nov 30 14:57:35 mars kernel: nvme nvme1: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0xffff
Nov 30 14:57:35 mars kernel: nvme nvme1: Does your device have a faulty power saving mode enabled?
Nov 30 14:57:35 mars kernel: nvme nvme1: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug
Nov 30 14:57:35 mars kernel: nvme nvme2: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0xffff
Nov 30 14:57:35 mars kernel: nvme nvme2: Does your device have a faulty power saving mode enabled?
Nov 30 14:57:35 mars kernel: nvme nvme2: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" and report a bug
Nov 30 14:57:35 mars kernel: nvme 0000:05:00.0: Unable to change power state from D3cold to D0, device inaccessible
Nov 30 14:57:35 mars kernel: nvme 0000:06:00.0: Unable to change power state from D3cold to D0, device inaccessible
Nov 30 14:57:35 mars kernel: nvme nvme1: Disabling device after reset failure: -19
Nov 30 14:57:35 mars kernel: nvme nvme2: Disabling device after reset failure: -19
Nov 30 14:57:35 mars kernel: nvme 0000:04:00.0: Unable to change power state from D3cold to D0, device inaccessible
Nov 30 14:57:35 mars kernel: nvme nvme0: Disabling device after reset failure: -19
Nov 30 14:57:35 mars kernel: nvme 0000:07:00.0: Unable to change power state from D3cold to D0, device inaccessible
Nov 30 14:57:35 mars kernel: nvme nvme3: Disabling device after reset failure: -19
Nov 30 14:57:35 mars kernel: I/O error, dev nvme2n1, sector 2064186456 op 0x1:(WRITE) flags 0x0 phys_seg 1 prio class 0
Nov 30 14:57:35 mars kernel: I/O error, dev nvme0n1, sector 1066834128 op 0x0:(READ) flags 0x4000 phys_seg 4 prio class 0
Nov 30 14:57:35 mars kernel: I/O error, dev nvme0n1, sector 1066844160 op 0x0:(READ) flags 0x0 phys_seg 18 prio class 0
Nov 30 14:57:35 mars kernel: zio pool=NVMePool vdev=/dev/disk/by-id/nvme-INTEL_SSDPE2KX020T8_PHLJ326100VP2P0BGN_1-part1 error=5 type=2 offset=1056855076864 size=4096 flags=1572992
Nov 30 14:57:35 mars kernel: zio pool=NVMePool vdev=/dev/disk/by-id/nvme-INTEL_SSDPE2KX020T8_PHLJ326100VP2P0BGN_1-part1 error=5 type=1 offset=546211336192 size=126976 flags=1074267264

I have already added "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off" to the kernel-boot-parameters accordingly, but it does not help at all. I have also disabled ASPM in the BIOS, and tried updating the Kernel to 6.11.x, but to no avail. The only thing that seems to bring my server back to a useable state is periodically setting the following values to the PCI-Devices:

echo "on" > /sys/bus/pci/devices/<device_id>/power/control
echo 0 > /sys/bus/pci/devices/<device_id>/d3cold_allowed

. I have made sure that my BIOS and my Drive-Firmwares are Up to Date. I have also tried to update again to see if there are any fixes for this, but this is unfortunately not the case.

The Server is using the following hardware:

Mainboard: Supermicro X13SCH-F
CPU: Intel Xeon E-2478
Chipset: Intel C266
The affected Drives: Intel DC-P4510 (4 TB)
RAM: 128GB Kingston ECC

Proxmox-Version: 8.3.0

I would be really grateful if somebody could look into that, since it renders my Proxmox-Setup nearly unusable.

EDIT: Corrected the issue report.
Hi, I have the same problem with two different Intel DC P4610 on two motherboards (Gigabyte C246N-WU2 and Asus WS C246 Pro) with Xeon E-2288G. I'm on Ubuntu 24.04.5, I tried to disable every power saving option in the bios and applying the settings you suggested but the problem is still here. This is what I see when the problem appens:

Code:
Oct  2 10:44:49 www [173018.878174] pcieport 0000:00:1c.0: AER: Correctable error message received from 0000:00:1c.0
Oct  2 10:44:49 www [173018.879934] pcieport 0000:00:1c.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
Oct  2 10:44:49 www [173018.881738] pcieport 0000:00:1c.0:   device [8086:a33c] error status/mask=00000001/00002000
Oct  2 10:44:49 www [173018.883490] pcieport 0000:00:1c.0:    [ 0] RxErr                  (First)
Oct  2 10:44:49 www [173019.045668] pcieport 0000:00:1c.0: AER: Correctable error message received from 0000:00:1c.0
Oct  2 10:44:49 www [173019.047438] pcieport 0000:00:1c.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
Oct  2 10:44:49 www [173019.049196] pcieport 0000:00:1c.0:   device [8086:a33c] error status/mask=00000001/00002000
Oct  2 10:44:49 www [173019.050906] pcieport 0000:00:1c.0:    [ 0] RxErr                  (First)
Oct  2 10:45:21 www [173051.017701] nvme nvme0: controller is down; will reset: CSTS=0xffffffff, PCI_STATUS=0x10
Oct  2 10:45:21 www [173051.018398] nvme nvme0: Does your device have a faulty power saving mode enabled?
Oct  2 10:45:21 www [173051.019160] nvme nvme0: Try "nvme_core.default_ps_max_latency_us=0 pcie_aspm=off pcie_port_pm=off" and report a bug
Oct  2 10:45:21 www [173051.047767] nvme 0000:03:00.0: enabling device (0000 -> 0002)
Oct  2 10:45:21 www [173051.048292] nvme nvme0: Disabling device after reset failure: -19
Oct  2 10:45:21 www [173051.056732] buffer_io_error: 2868598 callbacks suppressed
Oct  2 10:45:21 www [173051.056732] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 70254669 starting block 282169335)
Oct  2 10:45:21 www [173051.056739] Buffer I/O error on device nvme0n1p2, logical block 282044151
Oct  2 10:45:21 www [173051.056784] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 111150387 starting block 503146453)
Oct  2 10:45:21 www [173051.056790] Buffer I/O error on device nvme0n1p2, logical block 503021269
Oct  2 10:45:21 www [173051.056797] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 111150387 starting block 503146457)
Oct  2 10:45:21 www [173051.056797] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 24930525 starting block 716252231)
Oct  2 10:45:21 www [173051.056799] Buffer I/O error on device nvme0n1p2, logical block 503021273
Oct  2 10:45:21 www [173051.056801] Buffer I/O error on device nvme0n1p2, logical block 503021274
Oct  2 10:45:21 www [173051.056802] Aborting journal on device nvme0n1p2-8.
Oct  2 10:45:21 www [173051.056802] Buffer I/O error on device nvme0n1p2, logical block 503021275
Oct  2 10:45:21 www [173051.056804] Buffer I/O error on device nvme0n1p2, logical block 503021276
Oct  2 10:45:21 www [173051.056804] EXT4-fs (nvme0n1p2): failed to convert unwritten extents to written extents -- potential data loss!  (inode 24930525, error -5)
Oct  2 10:45:21 www [173051.056806] Buffer I/O error on dev nvme0n1p2, logical block 243826688, lost sync page write
Oct  2 10:45:21 www [173051.056807] Buffer I/O error on device nvme0n1p2, logical block 716127047
Oct  2 10:45:21 www [173051.056808] JBD2: I/O error when updating journal superblock for nvme0n1p2-8.
Oct  2 10:45:21 www [173051.056809] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 111149064 starting block 716172320)
Oct  2 10:45:21 www [173051.056811] Buffer I/O error on device nvme0n1p2, logical block 716046889
Oct  2 10:45:21 www [173051.056813] Buffer I/O error on device nvme0n1p2, logical block 716046890
Oct  2 10:45:21 www [173051.056814] Buffer I/O error on device nvme0n1p2, logical block 716046891
Oct  2 10:45:21 www [173051.056815] EXT4-fs error (device nvme0n1p2) in ext4_reserve_inode_write:6291: Journal has aborted
Oct  2 10:45:21 www [173051.056821] Buffer I/O error on dev nvme0n1p2, logical block 0, lost sync page write
Oct  2 10:45:21 www [173051.056822] EXT4-fs (nvme0n1p2): I/O error while writing superblock
Oct  2 10:45:21 www [173051.056824] EXT4-fs (nvme0n1p2): Remounting filesystem read-only
Oct  2 10:45:21 www [173051.056909] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 111149064 starting block 716172352)
Oct  2 10:45:21 www [173051.056916] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 111149064 starting block 716172072)
Oct  2 10:45:21 www [173051.056919] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 24930525 starting block 716252230)
Oct  2 10:45:21 www [173051.056922] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 107515325 starting block 344806789)
Oct  2 10:45:21 www [173051.057249] Buffer I/O error on dev nvme0n1p2, logical block 132, lost async page write
Oct  2 10:45:21 www [173051.057820] EXT4-fs warning (device nvme0n1p2): ext4_end_bio:369: I/O error 10 writing to inode 70254669 starting block 282169331)
Oct  2 10:45:21 www [173051.069427] Buffer I/O error on dev nvme0n1p2, logical block 135, lost async page write
Oct  2 10:45:21 www [173051.069910] Buffer I/O error on dev nvme0n1p2, logical block 140, lost async page write
Oct  2 10:45:21 www [173051.070362] Buffer I/O error on dev nvme0n1p2, logical block 162, lost async page write
Oct  2 10:45:21 www [173051.070805] Buffer I/O error on dev nvme0n1p2, logical block 188, lost async page write
Oct  2 10:45:21 www [173051.071250] Buffer I/O error on dev nvme0n1p2, logical block 210, lost async page write
Oct  2 10:45:21 www [173051.071706] Buffer I/O error on dev nvme0n1p2, logical block 214, lost async page write
Oct  2 10:45:21 www [173051.072141] Buffer I/O error on dev nvme0n1p2, logical block 231, lost async page write
Oct  2 10:45:26 www [173055.995412] EXT4-fs (nvme0n1p2): shut down requested (2)

Did you solve it? Thanks.