Proxmox VE 9.2.21 - BCM57412 bnxt_en network failure after upgrade to kernel 7.0.14-20-pve

could you (and others affected):
* share the journal from booting up the machine until the issue occurred
* try disabling some offloading functionality: `ethtool -K <nic> rx-gro-hw off lro off`
Joined journal from boot until network is down and some more informations.
Then I reboot server and disabled offloading with your command on botn nic 2 and 3, but after 8 minutes network goes down again.
 

Attachments

Thanks for the logs!
could you try booting with `pci=notph` added to the command-line (might have a detrimental effect on performance)

also - as most logs here (for the broadcom nics) look like the NICs are in a bond - which bond-mode is used? (the stanza from /etc/network/interfaces would help)
 
Thanks for the logs!
could you try booting with `pci=notph` added to the command-line (might have a detrimental effect on performance)

also - as most logs here (for the broadcom nics) look like the NICs are in a bond - which bond-mode is used? (the stanza from /etc/network/interfaces would help)

Code:
auto bond0
iface bond0 inet manual
        bond-slaves eno1np0 eno2np1
        bond-miimon 100
        bond-mode 802.3ad
        bond-xmit-hash-policy layer3+4
        
auto vmbr0
iface vmbr0 inet manual
        bridge-ports bond0
        bridge-stp off
        bridge-fd 0
        bridge-vlan-aware yes
        bridge-vids 2-4094

auto vmbr0.9
iface vmbr0.9 inet static
        address 10.0.9.240/24
        gateway 10.0.9.254
 
  • Like
Reactions: Stoiko Ivanov
Thanks for confirming! To clarify — when you say it landed in 7.2.x, do you mean upstream Linux 7.2, or is there already a timeline for when a Proxmox kernel build containing this fix (opt-in or otherwise) will be available? I couldn't find an official Proxmox announcement for a 7.2-based kernel yet, just community speculation referencing Ubuntu's 7.2 kernel release for 26.10.
I meant upstream 7.2.x. our current 7.0.14 kernels already contain a lot of patches cherry-picked from both 7.1.x and 7.2.x
 
Thanks for the logs!
could you try booting with `pci=notph` added to the command-line (might have a detrimental effect on performance)

also - as most logs here (for the broadcom nics) look like the NICs are in a bond - which bond-mode is used? (the stanza from /etc/network/interfaces would help)
extract from /etc/network/interfaces :
auto bond0
iface bond0 inet manual
bond-slaves nic2 nic3
bond-miimon 100
bond-mode balance-tlb
The bond is attached to vmbr1, which is VLAN-aware and carries several VLAN interfaces including the management network.

I will now test with "pci=notph" added to the kernel command line and collect a new set of logs.
 
  • Like
Reactions: Stoiko Ivanov
Hello Stoiko,

I performed the test with:

pci=notph

and booted again on kernel 7.0.14-20-pve.

The system remained operational longer than before, however the issue eventually reoccurred and the host finally rebooted.

The most interesting sequence I found in the logs is:

DMAR: DRHD: handling fault status reg 2
DMAR: [DMA Read NO_PASID] Request device [31:00.0] fault addr ...
[fault reason 0x06] PTE Read access is not set

Immediately afterwards:

bnxt_en 0000:31:00.1 nic3:
Fatal firmware reset event

followed by:

NETDEV WATCHDOG
TX timeout detected

and finally:

bnxt_init_nic err: fffffff0
nic open fail (rc: fffffff0)

Device 31:00.x corresponds to the BCM57412 adapter used by nic2 / nic3.

This seems to suggest that the root cause may still be related to DMA/IOMMU handling leading to a Broadcom firmware reset.

The bond configuration remains:

auto bond0
iface bond0 inet manual
bond-slaves nic2 nic3
bond-miimon 100
bond-mode balance-tlb

Please let me know if you would like me to test additional kernel parameters or a different bond mode.

Best regards,
 
  • Like
Reactions: Stoiko Ivanov
I can basically confirm your findings, we have the following trouble:

7.0.14-20

Network connectivity on an empty host fails after a short time with
================
NETDEV WATCHDOG: CPU x; transmit queue 1 timed out 5700 ms
bnxt_en: 0000:01:00.1 NETDEV WATCHDOG CPU y: transmit queue 2 timed out 5184 ms

Reverting to 7.0.14-19 and it runs fine.
HW: Lenovo Server with
OCP3 Card
Broadcom
Broadcom 57454 10/25GbE SFP28 4-port OCP Ethernet Adapter

Second finding:
We had problems with a Broadcom Firmware-Update
- Running Kernel 7.0.14-19 with
10/25Gb 2-port SFP28 BCM57414 OCP3 Adapter 235.1.164.14 OCP 3.0
runs fine (with a 10 Gbit LR transeiver)

Updating to
10/25Gb 2-port SFP28 BCM57414 OCP3 Adapter 237.1.148.0 OCP 3.0
broke the connection. Downgrading and it was fine.

Any recommendations on this ?
 
Just chiming in and wanting to follow, we also had problems with this. Everything was working great with 7.0.14-19, but not with 7.0.14-20.
Dell Server
Broadcom 57414
Firmware: 23.11.16.22
 
Hello Stoiko,

A quick update regarding the pci=notph testing.

I started a new test using the same setup:

- kernel 7.0.14-20-pve
- pci=notph enabled

I can confirm that the parameter is active:

BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet pci=notph

This is a new test run following the previous crash.

The host has now been running for more than 40 minutes and network connectivity remains fully operational. This is significantly longer than the previous failures, which typically occurred after approximately 5 to 20 minutes.

I still observe DMAR fault messages in the current boot:

DMAR: DRHD: handling fault status reg 2
DMAR: [DMA Read NO_PASID] Request device [31:00.0] fault addr ...
[fault reason 0x06] PTE Read access is not set

Interestingly, despite these DMAR messages being present, the BCM57412 interfaces have remained stable so far and the bond is still operational.

At this stage, pci=notph appears to improve system stability considerably, although I am continuing to monitor the host to determine whether the issue is fully resolved or only delayed.

I will provide another update if the problem reoccurs.

Best regards,
Emmanuel
 
College told me to try to add
iommu=pt
in
/etc/kernel/cmdline
A similar bug seems to arise in current debian kernels (6.12.111) with crashes after arising after 5-30 minutes. I am checking.
 
A further update regarding the pci=notph testing.

I started a new independent test using:

- kernel 7.0.14-20-pve
- pci=notph enabled

I can confirm the parameter was active:

BOOT_IMAGE=/boot/vmlinuz-7.0.14-20-pve root=/dev/mapper/pve-root ro quiet pci=notph

This test remained stable much longer than previous runs.

Without pci=notph, the issue typically occurred after approximately 5 to 20 minutes.

With pci=notph, the host remained operational for approximately 1 hour and 11 minutes before failing again and rebooting.

I still observed the same DMAR fault during the current boot:

DMAR: DRHD: handling fault status reg 2
DMAR: [DMA Read NO_PASID] Request device [31:00.0]
fault reason 0x06
PTE Read access is not set

Device 31:00.0 corresponds to one port of the BCM57412 adapter.

The previous failing boot again showed the same sequence:

DMAR fault
→ Broadcom firmware/driver errors
→ bnxt_init_nic failure
→ bond0 losing all active interfaces
→ I/O errors on dm-7
→ host reboot

Therefore, pci=notph appears to significantly improve stability but does not completely resolve the issue.

Please let me know if you would like me to test additional kernel parameters, BIOS settings (VT-d ?), or a different bond mode.
Best Regards
 
I must admit that iommu=pt did run longer (about 32 minutes, but crashed also in the end)
Unfortunately i made an error configuring the kernel-parameters, i am rechecking.
 
Last edited:
Note sure if this helps or not, but in our cluster of 4 nodes one of the nodes is running on 7.0.14-20-pve has been running for over an hour with active VM's appears to be working, for us only when the entire cluster went to 7.0.14-20-pve that we had issues.
 
I also observed that a node with a vm was running longer. I decided to downgrade as fast as possible as you might just try your luck.
I downgraded/pined the kernel to 7.0.14-19 and everything now remains rock solid.
 
Last edited:
For me it helped to use "pcie_aspm=off". Does anyone want to try / confirm that? Using "pci=notph" didn't do anything for me, right after booting, the connection was gone. Same without "pci=notph". But the other command helped in my case.
 
Today same problem , after reboot with new Kernel 7.0.14-19 no connection with node ( Cluster with 3 Nodes ). You have to upgrade niccli and firmware of your Broadcom NICs , I have N225P - 2 x 25/10GbE OCP 3.0 Adapter and N210P - 2 x 10GbE PCIe OCP 3.0 Adapter. Now is Cluster with Ceph working properly again with the last Enterprise Kernel.
 
I had the same problem with 4 network cards and 3 nodes.

Nodes lost connection between them, and systems reboots after a few minutes.

01:00.0 Ethernet controller: Broadcom Inc. and subsidiaries BCM57504 NetXtreme-E 10Gb/25Gb/40Gb/50Gb/100Gb Ethernet (rev 12)
01:00.1 Ethernet controller: Broadcom Inc. and subsidiaries BCM57504 NetXtreme-E 10Gb/25Gb/40Gb/50Gb/100Gb Ethernet (rev 12)
01:00.2 Ethernet controller: Broadcom Inc. and subsidiaries BCM57504 NetXtreme-E 10Gb/25Gb/40Gb/50Gb/100Gb Ethernet (rev 12)
01:00.3 Ethernet controller: Broadcom Inc. and subsidiaries BCM57504 NetXtreme-E 10Gb/25Gb/40Gb/50Gb/100Gb Ethernet (rev 12)

I already had the have the latest firmware in the network cards.

I had to boot and select the previous kerne 7.0.14-19-pve, to continue working.

Hope they can fix it in the new kernel release.

Regards.
 
Hi,

Same issue here with Broadcom BCM57504 NIC (bnxt_en driver) on 7.0.14-20-pve, working fine on 7.0.14-19-pve and previous.

Tell me what logs you need if you need any.
 
Having the same issue with Broadcom BCM57414 Dual 25Gb adapter, on a DELL PowerEdge R6515 with an AMD EPYC CPU and PVE 9 kernel 7.0.14-20-pve.

Also using Linux bond, in active-backup mode. There are two bonds on each server - only the second bond stopped working (luckily it's not the one that carries management VLAN), but that bond is using the secondary 10G port from the same Broadcom adapter. The primary bond, that continued to work was using the first 10G port on that same adapter. Primary bond (the one that worked) is connected to a switch. The second bond is actually a direct connection between the servers themselves. Don't know if that is important.

Journal log attached.

My Broadcom firmware is 38.11.38.06 and iDRAC is not offering anything newer.

Using the 7.0.14-19-pve kernel solved the issue for me as well.
 

Attachments