Proxmox VE 9.2.21 - BCM57412 bnxt_en network failure after upgrade to kernel 7.0.14-20-pve

Same here with 7.0.14-20-pve (7.0.14-19-pve works including switch back to it).

Broadcom BCM57416 NetXtreme-E 10GBase-T Ethernet on a Supermicro Server.

Logs are attached.
 

Attachments

Same on the DELL R640 and Broadcom BCM57404 Dual 25Gb SFP28 Ethernet (firmware 21.60.22.11) with 7.0.14-20-pve (7.0.14-19-pve works).
'pci=notph' and 'pcie_aspm=off' doesn't help.

Code:
Oct 02 09:34:17 host kernel: bnxt_en 0000:3b:00.0 eth0: Broadcom BCM57404 NetXtreme-E 10Gb/25Gb Ethernet found at mem ab020000, node addr 00:0a:f7:ab:82:6e
Oct 02 09:34:17 host kernel: bnxt_en 0000:3b:00.0: 63.008 Gb/s available PCIe bandwidth (8.0 GT/s PCIe x8 link)
Oct 02 09:34:17 host kernel: bnxt_en 0000:3b:00.1 eth1: Broadcom BCM57404 NetXtreme-E 10Gb/25Gb Ethernet found at mem ab000000, node addr 00:0a:f7:ab:82:6f
Oct 02 09:34:17 host kernel: bnxt_en 0000:3b:00.1: 63.008 Gb/s available PCIe bandwidth (8.0 GT/s PCIe x8 link)
Oct 02 09:34:17 host kernel: bnxt_en 0000:d8:00.0 eth2: Broadcom BCM57404 NetXtreme-E 10Gb/25Gb Ethernet found at mem ee820000, node addr 00:0a:f7:ab:82:1c
Oct 02 09:34:17 host kernel: bnxt_en 0000:d8:00.0: 63.008 Gb/s available PCIe bandwidth (8.0 GT/s PCIe x8 link)
Oct 02 09:34:17 host kernel: bnxt_en 0000:d8:00.1 eth3: Broadcom BCM57404 NetXtreme-E 10Gb/25Gb Ethernet found at mem ee800000, node addr 00:0a:f7:ab:82:1d
Oct 02 09:34:17 host kernel: bnxt_en 0000:d8:00.1: 63.008 Gb/s available PCIe bandwidth (8.0 GT/s PCIe x8 link)
Oct 02 09:34:17 host kernel: bnxt_en 0000:3b:00.1 ens1f1np1: renamed from eth1
Oct 02 09:34:17 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: renamed from eth0
Oct 02 09:34:17 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: renamed from eth2
Oct 02 09:34:17 host kernel: bnxt_en 0000:d8:00.1 ens3f1np1: renamed from eth3
Oct 02 09:34:25 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: NIC Link is Up, 25000 Mbps full duplex, Flow control: ON - receive & transmit
Oct 02 09:34:25 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: FEC autoneg off encoding: None
Oct 02 09:34:25 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: NIC Link is Up, 25000 Mbps full duplex, Flow control: none
Oct 02 09:34:25 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: FEC autoneg off encoding: None
Oct 02 09:34:25 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Unqualified SFP+ module detected on port 0
Oct 02 09:34:25 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Module part number W4GPP
Oct 02 09:34:26 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: NIC Link is Up, 25000 Mbps full duplex, Flow control: ON - receive & transmit
Oct 02 09:34:26 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: FEC autoneg off encoding: None
Oct 02 09:34:26 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: NIC Link is Up, 25000 Mbps full duplex, Flow control: none
Oct 02 09:34:26 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: FEC autoneg off encoding: None
Oct 02 09:34:26 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: Unqualified SFP+ module detected on port 0
Oct 02 09:34:26 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: Module part number 25GSFP28100BR
Oct 02 09:34:28 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: Unqualified SFP+ module detected on port 0
Oct 02 09:34:28 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: Module part number 25GSFP28100BR
Oct 02 09:34:30 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Unqualified SFP+ module detected on port 0
Oct 02 09:34:30 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Module part number W4GPP
Oct 02 09:34:31 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: Unqualified SFP+ module detected on port 0
Oct 02 09:34:31 host kernel: bnxt_en 0000:d8:00.0 ens3f0np0: Module part number 25GSFP28100BR
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: NETDEV WATCHDOG: CPU: 1: transmit queue 25 timed out 5120 ms
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: TX timeout detected, starting reset task!
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [0.0]: tx{fw_ring: 0 prod: 9a cons: 9a}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [0]: rx{fw_ring: 1 prod: 5f3} rx_agg{fw_ring: 29 agg_prod: 1000 sw_agg_prod: 1000}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [0]: cp{fw_ring: 0 raw_cons: 41d}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [1.0]: tx{fw_ring: 1 prod: 173 cons: 173}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [1]: rx{fw_ring: 2 prod: 401} rx_agg{fw_ring: 30 agg_prod: 1000 sw_agg_prod: 1000}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [1]: cp{fw_ring: 16 raw_cons: 7e}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [2.0]: tx{fw_ring: 2 prod: 11b cons: 117}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [2]: rx{fw_ring: 3 prod: 400} rx_agg{fw_ring: 31 agg_prod: 1000 sw_agg_prod: 1000}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [2]: cp{fw_ring: 17 raw_cons: 65}
....
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [26.0]: tx{fw_ring: 26 prod: f4 cons: ee}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [26]: rx{fw_ring: 27 prod: 400} rx_agg{fw_ring: 55 agg_prod: 1000 sw_agg_prod: 1000}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [26]: cp{fw_ring: 41 raw_cons: 56}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [27.0]: tx{fw_ring: 27 prod: b cons: b}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [27]: rx{fw_ring: 28 prod: 400} rx_agg{fw_ring: 56 agg_prod: 1000 sw_agg_prod: 1000}
Oct 02 09:35:29 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [27]: cp{fw_ring: 42 raw_cons: 4}
Oct 02 09:35:30 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Resp cmpl intr err msg: 0x51
Oct 02 09:35:30 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: hwrm_ring_free type 1 failed. rc:fffffff0 err:0
Oct 02 09:35:31 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Resp cmpl intr err msg: 0x51
Oct 02 09:35:31 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: hwrm_ring_free type 1 failed. rc:fffffff0 err:0
...
Oct 02 09:36:40 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Resp cmpl intr err msg: 0x51
Oct 02 09:36:40 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:0
Oct 02 09:36:41 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Resp cmpl intr err msg: 0x51
Oct 02 09:36:41 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:0
Oct 02 09:36:42 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Resp cmpl intr err msg: 0x51
Oct 02 09:36:42 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:0
Oct 02 09:36:43 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Error (timeout: 500015) msg {0x51 0x338} len:0
...
Oct 02 09:37:07 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Error (timeout: 500015) msg {0x51 0x353} len:0
Oct 02 09:37:07 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: hwrm_ring_free type 0 failed. rc:fffffff0 err:0
Oct 02 09:37:08 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Error (timeout: 500015) msg {0x61 0x354} len:0
Oct 02 09:37:08 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Error (timeout: 500015) msg {0x61 0x355} len:0
Oct 02 09:37:09 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: Error (timeout: 500015) msg {0x61 0x356} len:0
...
 
proxmox-kernel-7.0.12-21 has just been pushed to the pve-test repository.

The changelog is as follows:
Code:
proxmox-kernel-7.0 (7.0.14-21) trixie; urgency=medium

  * apparmor: fix a NULL pointer dereference on connected unix datagram sockets
    that any unprivileged user could trigger. The fix is taken from the mailing
    list, as it is not merged upstream yet.

  * block: fix possible silent data corruption on NVMe controllers with SGL
    support. When a large request whose memory segments are not page aligned,
    for example from direct I/O, got split, its first part could be sent in a
    format that cannot describe those segments.

  * cherry-picks for the following CVEs from stable 7.2 and 6.18:
    - CVE-2026-100073: ext4: fix transaction overflow during writeback
    - CVE-2026-93781: scsi: core: Do not block on tag allocation in scsi_eh_lock_door()
    - CVE-2026-97530: scsi: qla2xxx: Fix soft lockup polling continuation IOCB signature
    - CVE-2026-97537: scsi: qla2xxx: Fix queue teardown NULL dma_free and bitmap locking
    - CVE-2026-97536: scsi: qla2xxx: Fix use-after-free of qpair work on queue teardown
    - CVE-2026-93829: smb: client: fix races in cifsd thread creation
    - CVE-2026-93794: smb/client: flush dirty data before punching a hole
    - CVE-2026-90053: net/sched: sch_htb: limit htb_classify inner-class filter hops
    - CVE-2026-97611: net: openvswitch: fix use-after-free of the flow table mask array
    - CVE-2026-98023: vxlan: reject dynamic fdb entries that reference a nexthop id
    - CVE-2026-97615: net: bridge: use option bits for CFM/MRP frame handlers
    - CVE-2026-93288: netfilter: nfnetlink_log: wait for rcu grace period before freeing pernet state
    - CVE-2026-97608: netfilter: nf_log: unregister loggers before per-net teardown
    - CVE-2026-97523: mptcp: close race between scheduler and state change
    - CVE-2026-97522: mptcp: fix bad accounting in __mptcp_subflow_push_pending()
    - CVE-2026-98123: sctp: fix soft lockup from unpadded ASCONF-ACK parameter iteration
    - CVE-2026-98130: sctp: fix a TOCTOU race in SCTP_CMD_TIMER_START
    - CVE-2026-97903: exit: hold a reference to thread_pid across proc_flush_pid
    - CVE-2026-98163: cgroup: Avoid iteration of dying tasks with zero refcount
    - CVE-2026-97971: nstree: check listing permission before taking a namespace reference
    - CVE-2026-97594: landlock: Fix use-after-free of the source's parent directory
    - CVE-2026-97942: x86/alternatives: Exclude text poking against change_page_attr()

  * cherry-pick fixes from stable 7.2 and 6.18 and mainline 7.3 that have no CVE assigned yet:
    - xfs: fix reclaimed page accounting in xfs_buf_free
    - NFS: fix eof updates after NFSv4.2 fallocate/zero-range
    - smb: client: cancel reconnect work in clean_demultiplex_info()
    - cifs: Fix server use-after-free in cifs_chan_skip_or_disable()
    - smb: client: fix rlist race and missing initialization
    - smb: client: fix next_buffer UAF and NextCommand bounds in compound PDUs
    - smb: client: fix server->total_read for compound encrypted PDUs
    - net: lock the socket in sock_gettstamp()
    - netlink: do not free nlk->groups while lockless readers can use it
    - tcp: exclude old ACKs from tcp fast path
    - tcp: do not let tcp_rmem be set below 4096
    - tcp: fix use-after-free of retransmit_skb_hint in tcp_send_synack()
    - net/sched: codel: bound the dropping loop per dequeue call
    - net/sched: hhf: cap hh_flows_limit at change time
    - net/sched: sch_hfsc: bound the classify inner-filter walk with a drift budget
    - net/sched: reject IDR error pointers when deleting actions
    - net/sched: cls_u32: fix manual hash table handle IDR aliasing
    - openvswitch: avoid reallocating confirmed conntrack labels
    - net: openvswitch: conntrack: avoid modifying shared unconfirmed ct entry
    - net: openvswitch: conntrack: fix helper UAF due to extensions realloc
    - net/sched: act_ct: avoid modifying shared unconfirmed ct entry
    - net/sched: act_ct: fix helper UAF due to extensions realloc
    - vxlan: use one headroom snapshot for neighbour replies
    - bonding: use skb_cow_head() in bond_do_alb_xmit() and rlb_arp_xmit()
    - net: bridge: mdb: restart port group walk after deletion
    - net: bridge: mcast: don't truncate the port group walk on teardown
    - netfilter: flowtable: hold reference on ct until flow is released
    - netfilter: ip6t_rpfilter: reject routes without inet6_dev
    - ipv6: do not let ipv6_find_hdr() return an offset past the packet end
    - ipv6: Prevent rt6_insert_exception() for dying fib6_info.
    - ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink().
    - ip_gre: Reject enabling collect metadata through changelink
    - sctp: hold asoc or transport before mod_timer() in timer handlers
    - xfrm: fix compat ALLOCSPI request use-after-free
    - xfrm: iptfs: fix runt reassembly panic from short inner tot_len
    - xfrm: iptfs: fix stack OOB read in iptfs_skb_reset_frag_walk()
    - esp: downgrade zerocopy managed frags before mutating skb frags
    - xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
    - xfrm: hold net_device reference under RCU in bundle creation
    - xfrm: save input state data before secpath resets
    - xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
    - ipv6: xfrm: use full sockets in local error paths
    - exec: Cleanup POSIX timers right after de_thread()
    - signal: Prevent exec() race
    - posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
    - futex: Also allocate private hash on vfork()
    - workqueue: Fix NULL current_pwq deref in flush dependency check
    - KVM: x86: Re-pend GET_NESTED_STATE_PAGES if getting said pages fails

 -- Proxmox Support Team <support@proxmox.com>  Wed, 30 Sep 2026 18:32:36 +0200

Please ensure that acceleration / offloading is disabled on affected Broadcom nics via:
Code:
ethtool -K <nicX> gso off gro off lro off rx-gro-hw off rx-gro-list off rx-udp-gro-forwarding off

Please post any feedback on this thread to see if this helps with these network cards.

EDIT: Additionally, if problems persist with kernel -21, please retest with:
Code:
ethtool -K <nicX> rx-vlan-offload off tx-vlan-offload off rx-vlan-filter off
 
Last edited:
we uploaded a kernel with the changes for bnxt between 7.0.14-19-pve and 7.0.14-20-pve reverted:
http://download.proxmox.com/temp/kernel-7.0.14-22~test-bnxt-revert1/

It would be great if people experiencing the issues with bnxt_en nics could test this kernel - as it would help us in narrowing down the issue!

Additionally (and especially if this kernel still causes issues) - testing the mainline kernel builds from Ubuntu would be appreciated:
https://kernel.ubuntu.com/mainline/
(the latest 7.2 is probably best)
keep in mind that mainline builds by Ubuntu don't ship ZFS - so this only works if you don't have / on ZFS
 
proxmox-kernel-7.0.14-20-pve-signed and kernel-7.0.14-22~test-bnxt-revert1 and Ubuntu linux-image-unsigned-7.2.6-070206-generic_7.2.6-070206.202609141300 don't work in my situation.

Code:
Oct 02 16:04:40 host kernel: Linux version 7.0.14-22-pve (build@proxmox) (gcc (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44) #1 SMP PREEMPT_DYNAMIC PMX 7.0.14-22~test-bnxt-revert1 (2026-10- ()
Oct 02 16:04:40 host kernel: bnxt_en 0000:3b:00.0 eth0: Broadcom BCM57404 NetXtreme-E 10Gb/25Gb Ethernet found at mem ab020000, node addr
Oct 02 16:04:40 host kernel: bnxt_en 0000:3b:00.1 eth1: Broadcom BCM57404 NetXtreme-E 10Gb/25Gb Ethernet found at mem ab000000, node addr
Oct 02 16:04:40 host kernel: bnxt_en 0000:d8:00.0 eth2: Broadcom BCM57404 NetXtreme-E 10Gb/25Gb Ethernet found at mem ee820000, node addr
Oct 02 16:04:40 host kernel: bnxt_en 0000:d8:00.1 eth3: Broadcom BCM57404 NetXtreme-E 10Gb/25Gb Ethernet found at mem ee800000, node addr
...
Oct 02 16:05:01 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: NETDEV WATCHDOG: CPU: 0: transmit queue 11 timed out 5001 ms
Oct 02 16:05:01 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: TX timeout detected, starting reset task!
Oct 02 16:05:01 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [0.0]: tx{fw_ring: 0 prod: 23e cons: 230}
Oct 02 16:05:01 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [0]: rx{fw_ring: 1 prod: 436} rx_agg{fw_ring: 29 agg_prod: 1000 sw_agg_prod: 1000}
Oct 02 16:05:01 host kernel: bnxt_en 0000:3b:00.0 ens1f0np0: [0]: cp{fw_ring: 0 raw_cons: 129}
...

The 'cmdline' line contains only:
Code:
... root=/dev/mapper/pve-root ro quiet mitigations=off split_lock_detect=off

After adding "intel_iommu=on iommu=pt" to the 'cmdline', at the very least, the driver didn't crash immediately, but after about 5 minutes there was a brief "reset" (bnxt_en 0000:3b:00.0 ens1f0np0: NIC Link is Down) for 1 sec.

7.0.14-12 and 7.0.14-19 work fine.
 
  • Like
Reactions: Stoiko Ivanov
After adding "intel_iommu=on iommu=pt" to the 'cmdline', at the very least, the driver didn't crash immediately, but after about 5 minutes there was a brief "reset" (bnxt_en 0000:3b:00.0 ens1f0np0: NIC Link is Down) for 1 sec.

7.0.14-12 and 7.0.14-19 work fine.
Thanks for the test! so it seems reverting the 2 patches does not resolve the issue (and it's not related to our diff to mainline) :/
does the system have an Intel or AMD CPU (most reporters seem to have AMD)?
 
Thanks for the test! so it seems reverting the 2 patches does not resolve the issue (and it's not related to our diff to mainline) :/
does the system have an Intel or AMD CPU (most reporters seem to have AMD)?
Intel Xeon Gold 6248R CPU in a HPE DL360 Gen10

Indeed 7.0.14-22 didn't help, issue happened just after 10 minutes boot time. Still having 7.0.14-19 working fine
 
  • Like
Reactions: Stoiko Ivanov