Proxmox VE 9.2.21 - BCM57412 bnxt_en network failure after upgrade to kernel 7.0.14-20-pve

Just to clarify, how long does it take from the test kernel release until it becomes available on the production/stable channel for download and update?
 
Same problem whit my BCM57502
Linux amd2 7.0.14-20-pve #1 SMP PREEMPT_DYNAMIC PMX 7.0.14-20 (2026-09-24T10:43Z) x86_64 GNU/Linux

root@amd2:~# lspci -nn -s 02:00.002:00.0 Ethernet controller [0200]: Broadcom Inc. and subsidiaries BCM57502 NetXtreme-E 10Gb/25Gb/40Gb/50Gb Ethernet [14e4:1752] (rev 12)root@amd2:~# ethtool -i nic0driver: bnxt_enversion: 7.0.14-20-pvefirmware-version: 229.0.121.0/pkg N/Aexpansion-rom-version: bus-info: 0000:02:00.0supports-statistics: yessupports-test: yessupports-eeprom-access: yessupports-register-dump: yessupports-priv-flags: noroot@amd2:~# devlink dev info pci/0000:02:00.0pci/0000:02:00.0: driver bnxt_en serial_number 9C-6B-00-FF-FE-CD-8B-D8 board.serial_number 0123456789ABCDEFGH versions: fixed: board.id B650D4U3-2L2Q/BCM asic.id 1752 asic.rev B2 running: fw.psid 0.0.28 fw N/A fw.mgmt 229.0.121.0 fw.mgmt.api 1.10.3 stored: fw.psid 0.0.28 fw N/A fw.mgmt 229.0.121.0

Attached txt file is kernel msg for bnxt_en
And png file is this NIC BCM57502 connection to backup server - and random drops.

Have a nice day!
 

Attachments

Last edited:
Same issue here, on different hardware and with newer NIC firmware, so in my case a firmware update alone does not seem to be the fix.

Setup
  • HPE ProLiant DL360 Gen11, 3-node PVE 9.2 cluster
  • Broadcom BCM57412 NetXtreme-E Dual 10Gb SFP+ OCP 3.0 (bnxt_en)
  • NIC firmware: 237.1.141.0/pkg 237.1.148.0 (newer than the 236.x mentioned above)
  • Both ports in an active-backup bond, MTU 9000, used for DRBD/LINSTOR storage and migration

Observations on 7.0.14-20-pve

The firmware crashes on every boot, about 20–25 s after driver init and without any load. The event codes are identical each time:
Code:
07:23:24 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms
07:23:29 bnxt_en 0000:3c:00.1 nic5: TX timeout detected, starting reset task!

The first time, the card recovered by itself. About 10 minutes later, during a live migration, it hung for good:
Code:
07:33:08 bnxt_en 0000:3c:00.0 nic4: TX timeout detected, starting reset task!
07:33:08 bnxt_en 0000:3c:00.0 nic4: Abandoning msg {0xb1 0xd0b} len: 0 due to firmware status: 0x2000001
07:33:08 bnxt_en 0000:3c:00.0 nic4: hwrm stat ctx alloc failure rc: fffffff0
07:33:08 bnxt_en 0000:3c:00.0 nic4: nic open fail (rc: fffffff0)
Both ports went down, and the node lost its storage network.

modprobe -r bnxt_en && modprobe bnxt_en brought the link back temporarily. After a full power off/on via iLO, still on -20, the same thing happened again 20 s after boot:
Code:
08:02:27 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms
08:02:29 bnxt_en 0000:3c:00.1 nic5: Abandoning msg {0x23 0x437} len: 0 due to firmware status: 0x2000001
08:02:33 bnxt_en 0000:3c:00.1 nic5: NETDEV WATCHDOG: CPU: 50: transmit queue 1 timed out 5632 ms

IOMMU faults: every firmware crash on -20 coincides with a DMAR fault from the same NIC (Intel platform, IOMMU in translated mode). There are no DMAR faults on 7.0.14-17:
Code:
07:23:24 DMAR: [DMA Read NO_PASID] Request device [3c:00.1] fault addr 0xdc669000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
07:23:24 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27
07:33:02 DMAR: [DMA Read NO_PASID] Request device [3c:00.0] fault addr 0xe2e2a000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
07:33:08 bnxt_en 0000:3c:00.0 nic4: TX timeout detected, starting reset task!
08:02:27 DMAR: [DMA Read NO_PASID] Request device [3c:00.1] fault addr 0xdc698000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
08:02:27 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27
So it looks like the NIC is doing DMA to an address that is no longer (or not yet) mapped in the IOMMU, which then triggers the firmware fatal reset. That points to a DMA mapping change in bnxt_en or intel-iommu between -17/-19 and -20, rather than to the NIC firmware itself.

Workaround: booted 7.0.14-17-pve (the previous kernel on this host). There have been no firmware events since, including heavy live migrations. The other two nodes (same hardware) stay on 7.0.14-17 without issues.

Attached: bnxt-related kernel messages from both failing boots (filtered and deduplicated) and pveversion -v. Full logs are available on request.
 

Attachments

Same issue here, on different hardware and with newer NIC firmware, so in my case a firmware update alone does not seem to be the fix.

Setup
  • HPE ProLiant DL360 Gen11, 3-node PVE 9.2 cluster
  • Broadcom BCM57412 NetXtreme-E Dual 10Gb SFP+ OCP 3.0 (bnxt_en)
  • NIC firmware: 237.1.141.0/pkg 237.1.148.0 (newer than the 236.x mentioned above)
  • Both ports in an active-backup bond, MTU 9000, used for DRBD/LINSTOR storage and migration

Observations on 7.0.14-20-pve

The firmware crashes on every boot, about 20–25 s after driver init and without any load. The event codes are identical each time:
Code:
07:23:24 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms
07:23:29 bnxt_en 0000:3c:00.1 nic5: TX timeout detected, starting reset task!

The first time, the card recovered by itself. About 10 minutes later, during a live migration, it hung for good:
Code:
07:33:08 bnxt_en 0000:3c:00.0 nic4: TX timeout detected, starting reset task!
07:33:08 bnxt_en 0000:3c:00.0 nic4: Abandoning msg {0xb1 0xd0b} len: 0 due to firmware status: 0x2000001
07:33:08 bnxt_en 0000:3c:00.0 nic4: hwrm stat ctx alloc failure rc: fffffff0
07:33:08 bnxt_en 0000:3c:00.0 nic4: nic open fail (rc: fffffff0)
Both ports went down, and the node lost its storage network.

modprobe -r bnxt_en && modprobe bnxt_en brought the link back temporarily. After a full power off/on via iLO, still on -20, the same thing happened again 20 s after boot:
Code:
08:02:27 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms
08:02:29 bnxt_en 0000:3c:00.1 nic5: Abandoning msg {0x23 0x437} len: 0 due to firmware status: 0x2000001
08:02:33 bnxt_en 0000:3c:00.1 nic5: NETDEV WATCHDOG: CPU: 50: transmit queue 1 timed out 5632 ms

IOMMU faults: every firmware crash on -20 coincides with a DMAR fault from the same NIC (Intel platform, IOMMU in translated mode). There are no DMAR faults on 7.0.14-17:
Code:
07:23:24 DMAR: [DMA Read NO_PASID] Request device [3c:00.1] fault addr 0xdc669000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
07:23:24 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27
07:33:02 DMAR: [DMA Read NO_PASID] Request device [3c:00.0] fault addr 0xe2e2a000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
07:33:08 bnxt_en 0000:3c:00.0 nic4: TX timeout detected, starting reset task!
08:02:27 DMAR: [DMA Read NO_PASID] Request device [3c:00.1] fault addr 0xdc698000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
08:02:27 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27
So it looks like the NIC is doing DMA to an address that is no longer (or not yet) mapped in the IOMMU, which then triggers the firmware fatal reset. That points to a DMA mapping change in bnxt_en or intel-iommu between -17/-19 and -20, rather than to the NIC firmware itself.

Workaround: booted 7.0.14-17-pve (the previous kernel on this host). There have been no firmware events since, including heavy live migrations. The other two nodes (same hardware) stay on 7.0.14-17 without issues.

Attached: bnxt-related kernel messages from both failing boots (filtered and deduplicated) and pveversion -v. Full logs are available on request.
Can you try test kernel found here https://forum.proxmox.com/threads/p...de-to-kernel-7-0-14-20-pve.186749/post-871933 so far it seems to fix the issue
 
Not directly related to this Kernel bug but HPE released new firmware files (238.1.168.0) for serveral Broadcom NIC's.
Code:
This product fixes an issue where retransmitted NC-SI responses could use an incorrect MCTP tag, causing the response to be discarded or misrouted.
This product fixes an issue where ports without pluggable optical modules could incorrectly report the module status as "Not Inserted" instead of "Not Applicable."
 
We experienced the very same issue with kernel-7.0.14-20-pve and Broadcom network adapters (Broadcom Adv. Quad 25Gb Ethernet) in a Dell PowerEdge 770 server.

System is booting, but network connectivity is extremely instable and causes frequent reboots.

Firmware 38.11.38.06 (latest firmware that iDRAC offered to install)

Switching back to kernel 7.0.14-3-pve solved the problem.
 
Can you try test kernel found here https://forum.proxmox.com/threads/p...de-to-kernel-7-0-14-20-pve.186749/post-871933 so far it seems to fix the issue
Test result for 7.0.14-22~bnxt-test1 on Intel (HPE DL360 Gen11, BCM57412, NIC firmware 237.1.141.0):

  • No Fatal firmware reset event, no TX timeouts and no DMAR faults since boot. On -20 the first crash came 20–25 s after driver init, on every boot.
  • Load test: live migrations of 24 GB and 32 GB VMs to and from this node over the bnxt_en bond, at ~1 GB/s. This is exactly the scenario that killed the NIC on -20. No issues.
  • DRBD/LINSTOR storage traffic over the same NICs is stable.

So the fix works on Intel / intel-iommu as well. Thanks for the quick turnaround!
 
Feedback on kernel 7.0.14-22-pve

We experienced instability with previous kernel releases on our Proxmox cluster, particularly after upgrading from 7.0.14-16-pve.

Today we tested 7.0.14-22-pve on one node. After approximately 30 minutes of uptime, the node appears stable:

  • No unexpected reboots
  • No kernel panics
  • No obvious issues in dmesg/journal logs
  • Cluster services and Ceph operating normally
It is obviously too early to draw definitive conclusions, but so far the new kernel looks promising compared to the previous versions that caused problems in our environment.

We will continue monitoring and provide an update after a longer uptime period.
 
Same issue here. With kernel 7.0.14-20-pve and Supermicro AOC-STG-b2T-NI22 (based on BCM57416) updated to firmware 237.1.148.0, the system is not stable and network comes and goes every few minutes. Using kernel 7.0.14-16-pve is fine. I am gonna test 7.0.14-22-pve shortly.
 
In our lab we were hit with this issue due to running BCM57412 bnxt_en Nics when we updated to the .14-20 kernel. We pinned to .14-19 as a workaround.
Today I updated to the .14.22 kernel and so far no issues seen.
 
Last edited:
  • Like
Reactions: lucius_the
Having the same issue with Broadcom BCM57414 Dual 25Gb adapter, on a DELL PowerEdge R6515 with an AMD EPYC CPU and PVE 9 kernel 7.0.14-20-pve.

Also using Linux bond, in active-backup mode. There are two bonds on each server - only the second bond stopped working (luckily it's not the one that carries management VLAN), but that bond is using the secondary 10G port from the same Broadcom adapter. The primary bond, that continued to work was using the first 10G port on that same adapter. Primary bond (the one that worked) is connected to a switch. The second bond is actually a direct connection between the servers themselves. Don't know if that is important.

Journal log attached.

My Broadcom firmware is 38.11.38.06 and iDRAC is not offering anything newer.

Using the 7.0.14-19-pve kernel solved the issue for me as well.

I can confirm that 7.0.14-22-pve kernel has resolved the issue.
iDRAC is also offering newer firmware for the Broadcom network card on my servers that have cards of that brand, new version is 38.11.70.00.