I tested kernel-7.0.14-22_bnxt-test1, in my case, the problem is gone, there are no issues in the logs regarding the bnxt_en driver, and the network is working.
237.1.141.0/pkg 237.1.148.0 (newer than the 236.x mentioned above)07:23:24 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms
07:23:29 bnxt_en 0000:3c:00.1 nic5: TX timeout detected, starting reset task!
07:33:08 bnxt_en 0000:3c:00.0 nic4: TX timeout detected, starting reset task!
07:33:08 bnxt_en 0000:3c:00.0 nic4: Abandoning msg {0xb1 0xd0b} len: 0 due to firmware status: 0x2000001
07:33:08 bnxt_en 0000:3c:00.0 nic4: hwrm stat ctx alloc failure rc: fffffff0
07:33:08 bnxt_en 0000:3c:00.0 nic4: nic open fail (rc: fffffff0)
modprobe -r bnxt_en && modprobe bnxt_en brought the link back temporarily. After a full power off/on via iLO, still on -20, the same thing happened again 20 s after boot:08:02:27 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms
08:02:29 bnxt_en 0000:3c:00.1 nic5: Abandoning msg {0x23 0x437} len: 0 due to firmware status: 0x2000001
08:02:33 bnxt_en 0000:3c:00.1 nic5: NETDEV WATCHDOG: CPU: 50: transmit queue 1 timed out 5632 ms
07:23:24 DMAR: [DMA Read NO_PASID] Request device [3c:00.1] fault addr 0xdc669000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
07:23:24 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27
07:33:02 DMAR: [DMA Read NO_PASID] Request device [3c:00.0] fault addr 0xe2e2a000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
07:33:08 bnxt_en 0000:3c:00.0 nic4: TX timeout detected, starting reset task!
08:02:27 DMAR: [DMA Read NO_PASID] Request device [3c:00.1] fault addr 0xdc698000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear
08:02:27 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27
pveversion -v. Full logs are available on request.Can you try test kernel found here https://forum.proxmox.com/threads/p...de-to-kernel-7-0-14-20-pve.186749/post-871933 so far it seems to fix the issueSame issue here, on different hardware and with newer NIC firmware, so in my case a firmware update alone does not seem to be the fix.
Setup
- HPE ProLiant DL360 Gen11, 3-node PVE 9.2 cluster
- Broadcom BCM57412 NetXtreme-E Dual 10Gb SFP+ OCP 3.0 (bnxt_en)
- NIC firmware:
237.1.141.0/pkg 237.1.148.0(newer than the 236.x mentioned above)- Both ports in an active-backup bond, MTU 9000, used for DRBD/LINSTOR storage and migration
Observations on 7.0.14-20-pve
The firmware crashes on every boot, about 20–25 s after driver init and without any load. The event codes are identical each time:
Code:07:23:24 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms 07:23:29 bnxt_en 0000:3c:00.1 nic5: TX timeout detected, starting reset task!
The first time, the card recovered by itself. About 10 minutes later, during a live migration, it hung for good:
Both ports went down, and the node lost its storage network.Code:07:33:08 bnxt_en 0000:3c:00.0 nic4: TX timeout detected, starting reset task! 07:33:08 bnxt_en 0000:3c:00.0 nic4: Abandoning msg {0xb1 0xd0b} len: 0 due to firmware status: 0x2000001 07:33:08 bnxt_en 0000:3c:00.0 nic4: hwrm stat ctx alloc failure rc: fffffff0 07:33:08 bnxt_en 0000:3c:00.0 nic4: nic open fail (rc: fffffff0)
modprobe -r bnxt_en && modprobe bnxt_enbrought the link back temporarily. After a full power off/on via iLO, still on -20, the same thing happened again 20 s after boot:
Code:08:02:27 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27, min wait 1300 ms, max wait 4200 ms 08:02:29 bnxt_en 0000:3c:00.1 nic5: Abandoning msg {0x23 0x437} len: 0 due to firmware status: 0x2000001 08:02:33 bnxt_en 0000:3c:00.1 nic5: NETDEV WATCHDOG: CPU: 50: transmit queue 1 timed out 5632 ms
IOMMU faults: every firmware crash on -20 coincides with a DMAR fault from the same NIC (Intel platform, IOMMU in translated mode). There are no DMAR faults on 7.0.14-17:
So it looks like the NIC is doing DMA to an address that is no longer (or not yet) mapped in the IOMMU, which then triggers the firmware fatal reset. That points to a DMA mapping change in bnxt_en or intel-iommu between -17/-19 and -20, rather than to the NIC firmware itself.Code:07:23:24 DMAR: [DMA Read NO_PASID] Request device [3c:00.1] fault addr 0xdc669000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear 07:23:24 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27 07:33:02 DMAR: [DMA Read NO_PASID] Request device [3c:00.0] fault addr 0xe2e2a000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear 07:33:08 bnxt_en 0000:3c:00.0 nic4: TX timeout detected, starting reset task! 08:02:27 DMAR: [DMA Read NO_PASID] Request device [3c:00.1] fault addr 0xdc698000 [fault reason 0x71] SM: Present bit in first-level paging entry is clear 08:02:27 bnxt_en 0000:3c:00.0 nic4: Fatal firmware reset event, data1: 0x201, data2: 0xda27
Workaround: booted 7.0.14-17-pve (the previous kernel on this host). There have been no firmware events since, including heavy live migrations. The other two nodes (same hardware) stay on 7.0.14-17 without issues.
Attached: bnxt-related kernel messages from both failing boots (filtered and deduplicated) andpveversion -v. Full logs are available on request.
This product fixes an issue where retransmitted NC-SI responses could use an incorrect MCTP tag, causing the response to be discarded or misrouted.
This product fixes an issue where ports without pluggable optical modules could incorrectly report the module status as "Not Inserted" instead of "Not Applicable."
Test result for 7.0.14-22~bnxt-test1 on Intel (HPE DL360 Gen11, BCM57412, NIC firmware 237.1.141.0):Can you try test kernel found here https://forum.proxmox.com/threads/p...de-to-kernel-7-0-14-20-pve.186749/post-871933 so far it seems to fix the issue
Fatal firmware reset event, no TX timeouts and no DMAR faults since boot. On -20 the first crash came 20–25 s after driver init, on every boot.Yeah i want to know also, atm i am on kernel .19 pinned everywhere.Just to clarify, how long does it take from the test kernel release until it becomes available on the production/stable channel for download and update?
kernel -22 was released earlier today. it worked for my (not directly related to this) problemYeah i want to know also, atm i am on kernel .19 pinned everywhere.
Having the same issue with Broadcom BCM57414 Dual 25Gb adapter, on a DELL PowerEdge R6515 with an AMD EPYC CPU and PVE 9 kernel 7.0.14-20-pve.
Also using Linux bond, in active-backup mode. There are two bonds on each server - only the second bond stopped working (luckily it's not the one that carries management VLAN), but that bond is using the secondary 10G port from the same Broadcom adapter. The primary bond, that continued to work was using the first 10G port on that same adapter. Primary bond (the one that worked) is connected to a switch. The second bond is actually a direct connection between the servers themselves. Don't know if that is important.
Journal log attached.
My Broadcom firmware is 38.11.38.06 and iDRAC is not offering anything newer.
Using the 7.0.14-19-pve kernel solved the issue for me as well.
We use essential cookies to make this site work, and optional cookies to enhance your experience.