Encountered a issue reboot when upgrading pve version 8 to 9.

phumstory

New Member
May 22, 2025
1
0
1

PowerEdge T440 Tower Server - Log Error​

When rebooting it will hang, check the log error:

Category: System Health​

Severity: Critical
Message Id: PCI1318
Description: a fatal error was detected on a component at bus 4 device 0 function 1
(Log PCI1318 ที่เจอ (Bus 4, Dev 0, Func 1) → ตรงกับ PERC H730P ใน Slot 4)

**Version 8 does not encounter this issue.**
 
T440 was introduced in 2017, PVE v9 kernel may have dropped support for the PERC card.

https://search.brave.com/search?q=d...summary=1&conversation=6ec7f411cd70ed85cc6a3d

You may want to replace the card with a newer model, check for firmware upgrades, or look for an updated version of perccli64

Also have the option to roll back to / restore PVE 8, it is supported until next year - and will give you some time to test upgrade to V9 on different hardware. You might be due for a hardware refresh as well, enterprise usually does it ~every 5 years or when support ends
 
Last edited:
I have the same issue. Exchanged with a second "same" controller for troubleshooting purpose. Changes to HBA Mode with only one harddrive. Proxmox 9 fresh installed. But no success. From the first reboot on the error occurs.
I think the error appears since nearly one year!? But first on later Proxmox 8 versions.

Does anybody got a fix for it?
Which controller seems to be a good hardware refresh?
 
The solution for me is found! It wasn´t the controller. The problem ist the Kerner >6.5 with the Broadcom network interface cards. The are not compatible. If you disable them in system Bios the machine ist booting fine (PVE9) without any error AND with the Perc H730p working. Check out the reddit thread: https://www.reddit.com/r/Proxmox/comments/1o6lrh2/server_not_rebooting_properly/

Now I put in another network card and everything is fine!

Hope this helps every DELL T440 User to give their server a new job
 
Thanks @sightkick, his post saved me from a bad upgrade night of my Dell T440 from pve version 8 to 9 and I think I can add one piece that might help: a way to check whether your host is already suffering from this before you upgrade anything.

Background​


The known issue: on Dell PowerEdge hardware with onboard Broadcom BCM5720 NICs (tg3 driver), kernel 6.17+ hits a fatal PCIe AER error during the ACPI reboot transition. The mechanism is a conditional power-down workaround in tg3 (active since 6.6.2) combining badly with an AER-disabling DMI whitelist that covers specific Dell R-series servers but not others — so non-whitelisted models are exposed.

It matters beyond the T440. BCM5720 is used widely across the PowerEdge 14G line (R440, R640, R740, T340, T440, T640) and appears in other vendors' hardware too.

One reported case initially looked exactly like a PERC H730P failure before being correctly traced to the NIC. If something looks RAID-controller-shaped after an upgrade on this hardware, check dmesg for tg3 and AER lines before concluding it's the controller.

The part I want to add: you may already be affected, and you can check today​


This was the surprise for me. I assumed it was a future risk tied to the new kernel. It wasn't — my host was already faulting on kernel 6.8, and had been for two weeks, silently.

Open iDRAC → Maintenance → Lifecycle Log and search for PCI1318.

You're looking for entries like:

Code:
PCI1318 — A fatal error was detected on a component at bus X device 0 function 0
PCI1318 — A fatal error was detected on a component at bus X device 0 function 1

Cross-check the bus/device against your onboard NIC with lspci | grep -i ethernet. On mine, bus 4 device 0 function 0/1 was exactly the two onboard Broadcom ports.

I had these recurring from 21 August through 4 September, on PVE 8.4 with kernel 6.8. Nothing in the OS complained. I only found them because I went looking for something else.

Here's the before/after that convinced me:
  • 4 September: copying a ~34 GB Clonezilla image from a USB disk to a CIFS share — the copy crawled and repeatedly froze. I assumed the Windows target was slow and moved on.
  • 14 September, after replacing the NIC: identical operation, same source disk, same destination share, same data. Finished in a few minutes with no stalls at all.
The only variable that changed was the NIC. That first copy happened on the last day of the recorded PCI1318 window. A fatal AER error resets the device, the link drops and re-establishes, TCP stalls and retransmits — which from the operator's chair looks exactly like a large copy randomly freezing.

So: if large sustained transfers on one of these hosts feel inexplicably slow or stall, check the Lifecycle Log. You may already be losing to this and not know it.

What I did​

Replaced the onboard NICs with an Intel I350-T2 (dual-port 1GbE, igb driver, in-tree, no driver install needed), then disabled the Broadcom NICs in BIOS:

System BIOS Settings → Integrated Devices → Embedded NIC1 and NIC2 → Disabled (OS)

Unplugging the cables is not enough. The bug is driver-level, not traffic-level — tg3 loading at all is the problem. Disabled (OS) removes them from OS enumeration while keeping them available to BIOS/UEFI pre-boot.

iDRAC stages BIOS changes. If the page shows Current Value: Enabled with Pending Value: Disabled (OS), nothing has been committed and rebooting from the OS will not apply it — the pending value just persists. Either use Apply And Reboot at the bottom of the page and confirm a job appears under Maintenance → Job Queue, or set it via F2 at POST, which writes directly.

Confirmation on kernel 7.0​


Upgraded 8.4.21 → 9.2.20 on kernel 7.0.14-17-pve — well past 6.17, i.e. the kernel where this would fire. After the upgrade and reboot:

dmesg -T | grep -iE "aer|pcie bus error|tg3"

No tg3 lines. No PCIe bus errors. Only four ACPI _OSC messages about firmware not delegating AER control to the OS, which are informational and appear on plenty of systems.

pve8to9 --full: 0 failures before and after. PERC H730P enumerated cleanly on the new kernel (megasas 07.734.00.00-rc1, FW Ready, both VDs). Whole upgrade took under 90 minutes.

Two traps if you do the same NIC swap​


1. A brand-new NIC with no link light is probably not dead. New interfaces have no stanza in /etc/network/interfaces, and ifupdown never touches an interface not listed with auto/allow-hotplug — so it sits administratively down, which on igb disables the PHY entirely. Completely dark port, no light at all, indistinguishable from dead hardware in every cable/switch/slot test. Before suspecting the card:

Code:
ip link set <iface> up
ethtool <iface> | grep -i "link detected"

2. Interface-name collisions via alternative names — and they're invisible. The T440's firmware reports the onboard NIC as being in "slot 1", so udev generates ens1f0/ens1f1 for the Broadcom pair and registers them as altnames on eno1/eno2. An add-in card in physical PCIe slot 1 wants exactly those names, and loses.

BIOS-disabling the onboard NICs removes them and their altnames entirely, which fixes it as a side effect.

Hope it helps.
 
Last edited:
  • Like
Reactions: Onslow