Thanks
@sightkick, his post saved me from a bad upgrade night of my Dell T440 from pve version 8 to 9 and I think I can add one piece that might help:
a way to check whether your host is already suffering from this before you upgrade anything.
Background
The known issue: on Dell PowerEdge hardware with onboard
Broadcom BCM5720 NICs (tg3 driver), kernel
6.17+ hits a fatal PCIe AER error during the ACPI reboot transition. The mechanism is a conditional power-down workaround in tg3 (active since 6.6.2) combining badly with an AER-disabling DMI whitelist that covers specific Dell
R-series servers but not others — so non-whitelisted models are exposed.
It matters beyond the T440. BCM5720 is used widely across the PowerEdge 14G line (R440, R640, R740, T340, T440, T640) and appears in other vendors' hardware too.
One reported case initially looked exactly like a
PERC H730P failure before being correctly traced to the NIC. If something looks RAID-controller-shaped after an upgrade on this hardware, check dmesg for tg3 and AER lines before concluding it's the controller.
The part I want to add: you may already be affected, and you can check today
This was the surprise for me. I assumed it was a future risk tied to the new kernel. It wasn't — my host was already faulting on
kernel 6.8, and had been for two weeks, silently.
Open iDRAC → Maintenance → Lifecycle Log and search for PCI1318.
You're looking for entries like:
Code:
PCI1318 — A fatal error was detected on a component at bus X device 0 function 0
PCI1318 — A fatal error was detected on a component at bus X device 0 function 1
Cross-check the bus/device against your onboard NIC with
lspci | grep -i ethernet. On mine,
bus 4 device 0 function 0/1 was exactly the two onboard Broadcom ports.
I had these recurring from
21 August through 4 September, on PVE 8.4 with kernel 6.8. Nothing in the OS complained. I only found them because I went looking for something else.
Here's the before/after that convinced me:
- 4 September: copying a ~34 GB Clonezilla image from a USB disk to a CIFS share — the copy crawled and repeatedly froze. I assumed the Windows target was slow and moved on.
- 14 September, after replacing the NIC: identical operation, same source disk, same destination share, same data. Finished in a few minutes with no stalls at all.
The only variable that changed was the NIC. That first copy happened on the
last day of the recorded PCI1318 window. A fatal AER error resets the device, the link drops and re-establishes, TCP stalls and retransmits — which from the operator's chair looks exactly like a large copy randomly freezing.
So:
if large sustained transfers on one of these hosts feel inexplicably slow or stall, check the Lifecycle Log. You may already be losing to this and not know it.
What I did
Replaced the onboard NICs with an
Intel I350-T2 (dual-port 1GbE, igb driver, in-tree, no driver install needed), then
disabled the Broadcom NICs in BIOS:
System BIOS Settings → Integrated Devices → Embedded NIC1 and NIC2 → Disabled (OS)
Unplugging the cables is
not enough. The bug is driver-level, not traffic-level — tg3 loading at all is the problem. Disabled (OS) removes them from OS enumeration while keeping them available to BIOS/UEFI pre-boot.
iDRAC stages BIOS changes. If the page shows Current Value: Enabled with Pending Value: Disabled (OS), nothing has been committed and rebooting from the OS will not apply it — the pending value just persists. Either use
Apply And Reboot at the bottom of the page and confirm a job appears under Maintenance → Job Queue, or set it via
F2 at POST, which writes directly.
Confirmation on kernel 7.0
Upgraded 8.4.21 →
9.2.20 on
kernel 7.0.14-17-pve — well past 6.17, i.e. the kernel where this would fire. After the upgrade and reboot:
dmesg -T | grep -iE "aer|pcie bus error|tg3"
No tg3 lines. No PCIe bus errors. Only four ACPI _OSC messages about firmware not delegating AER control to the OS, which are informational and appear on plenty of systems.
pve8to9 --full: 0 failures before and after. PERC H730P enumerated cleanly on the new kernel (megasas 07.734.00.00-rc1, FW Ready, both VDs). Whole upgrade took under 90 minutes.
Two traps if you do the same NIC swap
1. A brand-new NIC with no link light is probably not dead. New interfaces have no stanza in /etc/network/interfaces, and ifupdown never touches an interface not listed with auto/allow-hotplug — so it sits administratively down, which on igb
disables the PHY entirely. Completely dark port, no light at all, indistinguishable from dead hardware in every cable/switch/slot test. Before suspecting the card:
Code:
ip link set <iface> up
ethtool <iface> | grep -i "link detected"
2. Interface-name collisions via alternative names — and they're invisible. The T440's firmware reports the
onboard NIC as being in "slot 1", so udev generates ens1f0/ens1f1 for the Broadcom pair and registers them as
altnames on eno1/eno2. An add-in card in physical PCIe slot 1 wants exactly those names, and loses.
BIOS-disabling the onboard NICs removes them and their altnames entirely, which fixes it as a side effect.
Hope it helps.