Regression in 7.0.14-20-pve: Realtek RTL8168h/r8169 fails with PCIe AER/DPC error – 7.0.14-19-pve works

Dennigma

Member
Jun 22, 2023
21
9
8
Hi,

I have found what appears to be a reproducible network regression between Proxmox kernels 7.0.14-19-pve and 7.0.14-20-pve.

The affected NIC is:

02:00.0 Ethernet controller [0200]: Realtek Semiconductor Co., Ltd.
RTL8111/8168/8211/8411 PCI Express Gigabit Ethernet Controller
[10ec:8168] (rev 15)

Subsystem: Realtek Semiconductor Co., Ltd. Device [10ec:0123]
Kernel driver in use: r8169
The driver identifies the controller as:

RTL8168h/8111h
Firmware:

rtl8168h-2_0.0.2 02/26/15

System​

proxmox-ve: 9.2.0
pve-manager: 9.2.21

proxmox-kernel-7.0: 7.0.14-20
proxmox-kernel-7.0.14-20-pve-signed: 7.0.14-20
proxmox-kernel-7.0.14-19-pve-signed: 7.0.14-19

Kernel 7.0.14-20-pve – broken​

When booting 7.0.14-20-pve, the NIC fails while the PHY is being initialized:

r8169 0000:02:00.0 enp2s0: PHY [r8169-0-200:00] driver [Generic FE-GE Realtek PHY] (irq=MAC)
r8169 0000:02:00.0 enp2s0: configuring for phy/gmii link mode
pcieport 0000:00:1c.0: DPC: containment event, status:0x1f01: unmasked uncorrectable error detected
pcieport 0000:00:1c.0: PCIe Bus Error: severity=Uncorrectable (Non-Fatal), type=Transaction Layer, (Receiver ID)
r8169 0000:02:00.0 enp2s0: rtl_ocp_gphy_cond == 1 (loop: 10, delay: 25).
pcieport 0000:00:1c.0: device [8086:a33d] error status/mask=00100000/00010000
pcieport 0000:00:1c.0: [20] UnsupReq (First)
pcieport 0000:00:1c.0: AER: TLP Header: 0x34000000 0x02000010 0x00000000 0x00000000
r8169 0000:02:00.0: AER: can't recover (no error_detected callback)
pcieport 0000:00:1c.0: AER: device recovery failed
Afterwards:

enp2s0: <NO-CARRIER,BROADCAST,MULTICAST,UP>
and:

Speed: Unknown!
Duplex: Unknown! (255)
Auto-negotiation: on
Link detected: no
The interface never recovers and consequently vmbr0 stays down.

Kernel 7.0.14-19-pve – working​

I then booted the same machine without making any hardware, cabling, switch or network configuration changes into 7.0.14-19-pve.

The same NIC initializes normally:

Generic FE-GE Realtek PHY r8169-0-200:00: attached PHY driver (mii_bus:phy_addr=r8169-0-200:00, irq=MAC)
r8169 0000:02:00.0 enp2s0: Link is Down
vmbr0: port 1(enp2s0) entered blocking state
vmbr0: port 1(enp2s0) entered forwarding state
vmbr0: port 1(enp2s0) entered disabled state
r8169 0000:02:00.0 enp2s0: Link is Up - 1Gbps/Full - flow control off
vmbr0: port 1(enp2s0) entered blocking state
vmbr0: port 1(enp2s0) entered forwarding state
ethtool enp2s0 then reports:

Speed: 1000Mb/s
Duplex: Full
Auto-negotiation: on
master-slave status: slave
Link detected: yes
enp2s0 and vmbr0 are both UP,LOWER_UP, and networking works normally.

Direct comparison​

7.0.14-19-pve7.0.14-20-pve
NICRTL8168h/8111hRTL8168h/8111h
PCI ID10ec:8168 rev 1510ec:8168 rev 15
Driverr8169r8169
Firmwarertl8168h-2_0.0.2rtl8168h-2_0.0.2
Physical link1 Gbps / FullNo carrier
AER/DPC errorNoYes
vmbr0UPDOWN
NetworkingWorkingBroken
PCIe ASPM is disabled in both cases:

LnkCtl: ASPM Disabled
The PCIe link itself is reported as:

LnkSta: Speed 2.5GT/s, Width x1
This is reproducible by switching between the two installed kernels:

7.0.14-19-pve → working

7.0.14-20-pve → NIC fails during initialization with AER/DPC Unsupported Request

7.0.14-19-pve → working again


I have a secondary Intel NIC in this host, so I can keep SSH access while running the affected 7.0.14-20-pve kernel and can provide additional logs, register dumps or perform tests if needed.

Please let me know if any additional debugging output would be useful.
 
Last edited:
I have the same (works in -19, stays down in -20), but with the bnxt_en driver for an Broadcom Inc. and subsidiaries BCM57414 NetXtreme-E 10Gb/25Gb RDMA Ethernet Controller:

Code:
[   11.380325] bnxt_en 0000:01:00.0 nic6: renamed from enp1s0f0np0
[   13.608525] bnxt_en 0000:01:00.0 nic6: NIC Link is Up, 25000 Mbps (NRZ) full duplex, Flow control: none
[   13.608749] bnxt_en 0000:01:00.0 nic6: FEC autoneg off encoding: Clause 91 RS(528,514)
[   13.615436] bond1: (slave nic6): Enslaving as a backup interface with an up link
[   14.049829] bnxt_en 0000:01:00.0 nic6: entered allmulticast mode
[   14.050450] bnxt_en 0000:01:00.0 nic6: entered promiscuous mode
[  116.710312] bnxt_en 0000:01:00.0 nic6: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms
[  139.177347] bond1: (slave nic6): link status definitely down, disabling slave
[  139.430682] bnxt_en 0000:01:00.0 nic6: Device requests max timeout of 100 seconds, may trigger hung task watchdog (kernel default 120s)
[  139.599364] bond1: (slave nic6): link status definitely up, 25000 Mbps full duplex
[  425.191895] bnxt_en 0000:01:00.0 nic6: hwrm req_type 0x23 seq id 0x705 error 0xf
[  425.193279] bnxt_en 0000:01:00.0 nic6: hwrm req_type 0xb4 seq id 0x706 error 0xf
[  426.215911] bnxt_en 0000:01:00.0 nic6: Abandoning msg {0x23 0x707} len: 0 due to firmware status: 0x2000001
[  427.239929] bnxt_en 0000:01:00.0 nic6: Abandoning msg {0x23 0x70a} len: 0 due to firmware status: 0x2000001
[  428.263946] bnxt_en 0000:01:00.0 nic6: Abandoning msg {0x23 0x70d} len: 0 due to firmware status: 0x2000001
[  429.287934] bnxt_en 0000:01:00.0 nic6: Abandoning msg {0x23 0x710} len: 0 due to firmware status: 0x2000001
[  430.311928] bnxt_en 0000:01:00.0 nic6: Abandoning msg {0x23 0x713} len: 0 due to firmware status: 0x2000001
[  431.335936] bnxt_en 0000:01:00.0 nic6: Abandoning msg {0x23 0x716} len: 0 due to firmware status: 0x2000001
[  432.359943] bnxt_en 0000:01:00.0 nic6: Abandoning msg {0x23 0x719} len: 0 due to firmware status: 0x2000001
[  433.383950] bnxt_en 0000:01:00.0 nic6: Abandoning msg {0x23 0x71c} len: 0 due to firmware status: 0x2000001

Interestingly only two out of six controllers show this behaviour. A good one gives me:

Code:
[   11.367400] bnxt_en 0000:85:00.0 nic2: renamed from ens27f0np0
[   12.655564] bnxt_en 0000:85:00.0 nic2: NIC Link is Up, 25000 Mbps (NRZ) full duplex, Flow control: none
[   12.655846] bnxt_en 0000:85:00.0 nic2: FEC autoneg off encoding: Clause 91 RS(528,514)
[   12.663238] bond2: (slave nic2): Enslaving as a backup interface with an up link
[   13.521540] bnxt_en 0000:85:00.0 nic2: entered allmulticast mode
[   13.522695] bnxt_en 0000:85:00.0 nic2: entered promiscuous mode

Regards,
Robert
 
I tried some things, apparently it helps to use the grub parameter

GRUB_CMDLINE_LINUX_DEFAULT="quiet nosplash debug pcie_aspm=off"

It seems to be a bug with the energy saving of the kernel, the default drivers don't support this fast silence / wake up calls.
I did NOT try to install the vendor drivers, which might've helped, too.
 
see https://forum.proxmox.com/threads/p...after-upgrade-to-kernel-7-0-14-20-pve.186749/ for an issue that might be related.

if possible try booting with `pci=notph` (see the linked thread for other tests) - thanks!

Thanks, I tested pci=notph on the affected pve3 host with 7.0.14-20-pve. Unfortunately, it does not resolve the issue.

The parameter is definitely active:

BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve ... pci=notph
PCIe TPH is disabled
However, the RTL8168h/8111h (10ec:8168, XID 541) still fails to establish a link:

Speed: Unknown!
Duplex: Unknown! (255)
Link detected: no
The same DPC/AER errors occur during PHY initialization:

pcieport 0000:00:1c.0: DPC: containment event, status:0x1f01: unmasked uncorrectable error detected
pcieport 0000:00:1c.0: [20] UnsupReq (First)
pcieport 0000:00:1c.0: AER: TLP Header: 0x34000000 0x02000010 0x00000000 0x00000000
r8169 0000:02:00.0: AER: can't recover (no error_detected callback)
r8169 0000:02:00.0 enp2s0: rtl_ocp_gphy_cond == 1 (loop: 10, delay: 25).
pcieport 0000:00:1c.0: AER: device recovery failed
For comparison, on the same host, same NIC and same 7.0.14-20-pve kernel, booting with pcie_aspm=off makes the NIC work normally:

Speed: 1000Mb/s
Duplex: Full
Link detected: yes
r8169 0000:02:00.0 enp2s0: Link is Up - 1Gbps/Full
The DPC/AER/UnsupReq errors are also absent with pcie_aspm=off.

So far my results on this host are:

7.0.14-19-pve -> works
7.0.14-20-pve -> fails
7.0.14-20-pve + pci=notph -> fails
7.0.14-20-pve + pcie_aspm=off -> works
I can run additional tests from the linked thread on this host if useful.
 
Thanks, I tested pci=notph on the affected pve3 host with 7.0.14-20-pve. Unfortunately, it does not resolve the issue.
Your NIC does not use the bnxt_en driver (I overlooked that (as the second post mentioned bnxt) - so I think the issue is probaly unrelated.
https://bugzilla.proxmox.com/show_bug.cgi?id=8110 - this issue was opened recently for realtek nic issues with 7.0.14-20

7.0.14-20-pve + pcie_aspm=off -> works
I can run additional tests from the linked thread on this host if useful.
it would be great if you could test 7.0.14-21-pve (which is available on pve-test) without pcie_aspm=off on the command line

Thanks!
 
Your NIC does not use the bnxt_en driver (I overlooked that (as the second post mentioned bnxt) - so I think the issue is probaly unrelated.
https://bugzilla.proxmox.com/show_bug.cgi?id=8110 - this issue was opened recently for realtek nic issues with 7.0.14-20


it would be great if you could test 7.0.14-21-pve (which is available on pve-test) without pcie_aspm=off on the command line

Thanks!
Hey, yeah i did read that post. it's not directly connected to my issue. and i can not use the pve-test mirror.
 
  • Like
Reactions: Stoiko Ivanov
Your NIC does not use the bnxt_en driver (I overlooked that (as the second post mentioned bnxt) - so I think the issue is probaly unrelated.
https://bugzilla.proxmox.com/show_bug.cgi?id=8110 - this issue was opened recently for realtek nic issues with 7.0.14-20


it would be great if you could test 7.0.14-21-pve (which is available on pve-test) without pcie_aspm=off on the command line

Thanks!
I was able to free one host for testing:

tested the new 7.0.14-21-pve kernel on the affected host (pve3).

Unfortunately, 7.0.14-21-pve still reproduces the issue without any workaround.

Hardware:

Realtek RTL8168h/8111h
PCI ID: 10ec:8168
r8169 XID: 541
Interface: enp2s0

7.0.14-21-pve without workaround​

The kernel was booted without pcie_aspm=off and without pci=notph:

7.0.14-21-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.14-21-pve root=/dev/mapper/pve-root ro quiet nosplash debug
The NIC does not establish a link:

Speed: Unknown!
Duplex: Unknown! (255)
Link detected: no

LnkCtl: ASPM Disabled; RCB 64 bytes, LnkDisable- CommClk+
LnkSta: Speed 2.5GT/s, Width x1
The same DPC/AER failure seen with 7.0.14-20-pve occurs:

pcieport 0000:00:1c.0: DPC: containment event, status:0x1f01: unmasked uncorrectable error detected
pcieport 0000:00:1c.0: [20] UnsupReq (First)
pcieport 0000:00:1c.0: AER: TLP Header: 0x34000000 0x02000010 0x00000000 0x00000000
r8169 0000:02:00.0: AER: can't recover (no error_detected callback)
r8169 0000:02:00.0 enp2s0: rtl_ocp_gphy_cond == 1 (loop: 10, delay: 25).
pcieport 0000:00:1c.0: AER: device recovery failed

7.0.14-21-pve with​

I then booted the same host, NIC and kernel with only pcie_aspm=off added:

7.0.14-21-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.14-21-pve root=/dev/mapper/pve-root ro quiet nosplash debug pcie_aspm=off
The NIC works normally:

Speed: 1000Mb/s
Duplex: Full
Link detected: yes
PCIe state:

MaxPayload 128 bytes, MaxReadReq 4096 bytes
LnkCtl: ASPM L1 Enabled; RCB 64 bytes, LnkDisable- CommClk+
LnkSta: Speed 2.5GT/s, Width x1
The driver initializes successfully:

r8169 0000:02:00.0 eth0: RTL8168h/8111h, 00:0a:cd:00:02:31, XID 541
r8169 0000:02:00.0 enp2s0: PHY [r8169-0-200:00] driver [Generic FE-GE Realtek PHY]
r8169 0000:02:00.0 enp2s0: configuring for phy/gmii link mode
r8169 0000:02:00.0 enp2s0: Link is Up - 1Gbps/Full - flow control off
There are no DPC/AER/UnsupReq/rtl_ocp errors in this boot.

Current test matrix:

7.0.14-19-pve -> works
7.0.14-20-pve -> fails
7.0.14-20-pve + pci=notph -> fails
7.0.14-20-pve + pcie_aspm=off -> works
7.0.14-21-pve -> fails
7.0.14-21-pve + pcie_aspm=off -> works
So 7.0.14-21-pve still reproduces the regression, while pcie_aspm=off continues to work around it.

One additional observation: in the working 7.0.14-21-pve + pcie_aspm=off boot, the device reports MaxReadReq 4096 bytes. I previously observed MaxReadReq 512 bytes in the failing 7.0.14-20-pve state. I don't know whether this is related to the root cause, but I mention it in case it is useful.

I can run further tests or provide full boot logs if needed.
 
  • Like
Reactions: carsten_h
7.0.14-21-pve + pcie_aspm=off -> works
Thanks for the tests! much appreciated!

We had a report about kernel 7.0.14-20-pve and realtek nics in our bugzilla:
https://bugzilla.proxmox.com/show_bug.cgi?id=8110

@t.lamprecht reverted a few patches that seem relevant in:
https://git.proxmox.com/?p=pve-kernel.git;a=commitdiff;h=803fe655550733d9b7c36669194fcb8fcd795751
but we're currently also considering a different fix

In any case I think this should be fixed in kernel 7.0.14-22-pve when it becomes available
 
  • Like
Reactions: Dennigma
Thanks for the tests! much appreciated!

We had a report about kernel 7.0.14-20-pve and realtek nics in our bugzilla:
https://bugzilla.proxmox.com/show_bug.cgi?id=8110

@t.lamprecht reverted a few patches that seem relevant in:
https://git.proxmox.com/?p=pve-kernel.git;a=commitdiff;h=803fe655550733d9b7c36669194fcb8fcd795751
but we're currently also considering a different fix

In any case I think this should be fixed in kernel 7.0.14-22-pve when it becomes available
Sounds good. If you need tests, just let me know, I'm at the moment able to free one host and install test kernels.
 
  • Like
Reactions: Stoiko Ivanov
  • Like
Reactions: Dennigma
A test-kernel for another regression (in bnxt_en) just got uploaded to:

http://download.proxmox.com/temp/kernel-7.0.14-22_bnxt-test1/

this should contain the current fix mentioned above - if you can - please test booting that kerne (without pcie_aspm=off) - and let us know
if it works

see https://forum.proxmox.com/threads/p...de-to-kernel-7-0-14-20-pve.186749/post-871933 for more details

Thanks!
Thanks, I tested the 7.0.14-22~bnxt-test1 test kernel on the affected pve3 host.

The test kernel fixes the issue for me without pcie_aspm=off.

The test was performed with:

7.0.14-22-pve
BOOT_IMAGE=/boot/vmlinuz-7.0.14-22-pve root=/dev/mapper/pve-root ro quiet nosplash debug
So neither pcie_aspm=off nor pci=notph was used.

Hardware:

Realtek RTL8168h/8111h
PCI ID: 10ec:8168
r8169 XID: 541
Interface: enp2s0
With the test kernel the NIC comes up normally:

Speed: 1000Mb/s
Duplex: Full
Link detected: yes
PCIe state:

MaxPayload 128 bytes, MaxReadReq 4096 bytes
LnkCtl: ASPM Disabled; RCB 64 bytes, LnkDisable- CommClk+
LnkSta: Speed 2.5GT/s, Width x1
Driver initialization also completes normally:

r8169 0000:02:00.0 eth0: RTL8168h/8111h, 00:0a:cd:00:02:31, XID 541
Generic FE-GE Realtek PHY r8169-0-200:00: attached PHY driver
r8169 0000:02:00.0 enp2s0: Link is Up - 1Gbps/Full - flow control off
Most importantly, the previous failure is no longer present. There is no DPC containment event, no UnsupReq, no rtl_ocp_gphy_cond failure and no failed AER recovery.

Updated test matrix:

7.0.14-19-pve -> works
7.0.14-20-pve -> fails
7.0.14-20-pve + pci=notph -> fails
7.0.14-20-pve + pcie_aspm=off -> works
7.0.14-21-pve -> fails
7.0.14-21-pve + pcie_aspm=off -> works
7.0.14-22~bnxt-test1 -> works without workaround
So whatever relevant fix/change is included in this test kernel appears to resolve the regression on my RTL8168h/8111h system as well.

Thanks!
 
Update: It's still running fine, I even migrated some VMs back to this host. Seems to be stable for over 4 hours now. (Time wasn't an issue since it always happened while booting)