Hi,
Possibly related thread: https://forum.proxmox.com/threads/r...e-aer-dpc-error-–-7-0-14-19-pve-works.186785/
I can reproduce an AMD Radeon PCI passthrough failure when switching a Proxmox host from 7.0.14-19-pve to 7.0.14-20-pve. Booting the previous kernel restores GPU initialization inside the same Debian VM, and booting -20 again reproduces the failure.
TEST RESULTS
2026-10-05: 7.0.14-20-pve -> fails
2026-10-08: 7.0.14-19-pve -> works
2026-10-08: 7.0.14-20-pve -> fails again
2026-10-08: 7.0.14-19-pve -> works
2026-10-08: 7.0.14-22-pve -> fails again
2026-10-08: 7.0.14-19-pve -> works
For the October 8 tests:
- VM configuration and generated QEMU command were identical before redaction.
- QEMU remained pve-qemu-kvm 11.0.3-4.
- The Debian guest kernel remained 6.12.111+deb13-amd64.
- No VFIO binding, VM PCI settings, or BIOS configuration changes were made as part of these tests.
On -19, the guest enumerates the Radeon and HDMI audio, binds amdgpu and snd_hda_intel, initializes 8176 MiB VRAM, and creates /dev/dri/renderD128. A rendering workload was not benchmarked/tested.
On -20 or -22, the VM still starts and the guest agent responds, but the guest has only its emulated display and no Radeon render device. Both GPU functions on the host show "Unknown header type 7f". The guest kernel sees invalid PCI configuration for the GPU and does not enumerate it as a usable device.
FAILURE OUTPUT
VM startup on -20:
Important: the initial reset ioctl warning ALSO appears on the working -19 boot. The distinguishing symptoms on -20 are the D3 errors and inaccessible PCI configuration, not that initial warning alone.
Host kernel log on the failing boot includes:
There are also repeated vfio_bar_restore messages. Host lspci -vv shows "Unknown header type 7f" for both 03:00.0 and 03:00.1.
The repeat-failure guest kernel reports:
ENVIRONMENT
- Proxmox VE 9.2; pve-manager 9.2.21
- qemu-server 9.2.10
- pve-qemu-kvm 11.0.3-4
- pve-firmware 3.18-6
- pve-edk2-firmware 4.2026.08-1
- CPU: AMD Ryzen 5 5600G with Radeon Graphics
- System: HP Pavilion Gaming Desktop TG01-2xxx
- Motherboard: HP 8906, version SMVB
- BIOS: AMI F.40, reported release date 04/08/2026
- GPU: AMD Navi 23 [Radeon RX 6600/6600 XT/6600M], PCI ID 1002:73ff, revision c1, subsystem 103c:8a33
- HDMI audio: PCI ID 1002:ab28
- Host PCI addresses: 0000:03:00.0 and 0000:03:00.1
- Separate IOMMU groups: 11 and 12; AMD interrupt remapping enabled
- A separate integrated Cezanne GPU remains on the host
PASSTHROUGH SETUP
Both GPU functions are assigned through the same hostpci0 entry. The current configuration does not specify pcie=1 or a custom ROM file. The fuller redacted VM configuration and relevant QEMU device arguments are attached.
The host amdgpu driver initializes the discrete GPU during boot, then Proxmox detaches it and binds both functions to vfio-pci at VM startup. This same handoff occurs on both tested host kernels. Early VFIO binding has not been tested as a workaround.
The GPU's reset_method is "bus". vfio_pci disable_idle_d3 is N. Runtime power/control is auto. D3hot alone is not a reliable failure indicator here: it was also observed while the successfully initialized GPU was idle on -19.
UPGRADE CONTEXT
The October 5 transaction installed kernel -20 and also updated QEMU from 11.0.3-3 to 11.0.3-4. However, the October 8 comparison kept QEMU at 11.0.3-4 while changing the host kernel, and passthrough worked on -19 and failed again on -20.
The original October 5 guest used Debian kernel 6.12.107+deb13-amd64. Both October 8 tests used 6.12.111+deb13-amd64, so that guest update does not explain the October 8 working/failing comparison.
CURRENT WORKAROUND / QUESTION
The host is now persistently pinned to boot 7.0.14-19-pve with proxmox-boot-tool, and that kernel package is marked manually installed to retain it as a fallback. The pin applies on the next reboot; it does not change the running kernel.
Is there a known amdgpu, PCI power-management, or VFIO regression between these two Proxmox kernel builds? Is there a specific fix or newer kernel worth testing? I have not tested newer kernels or identified an offending commit.
Attached are redacted diagnostic excerpts for the working boot and both failures. Host/guest names, VM identifiers, network identifiers, UUIDs, storage names, and deployment-specific paths have been removed or replaced. Hardware IDs, PCI addresses, versions, timestamps, and relevant errors are preserved.
The attached files only contain the -20 reboots.
Thanks for any help!
Possibly related thread: https://forum.proxmox.com/threads/r...e-aer-dpc-error-–-7-0-14-19-pve-works.186785/
I can reproduce an AMD Radeon PCI passthrough failure when switching a Proxmox host from 7.0.14-19-pve to 7.0.14-20-pve. Booting the previous kernel restores GPU initialization inside the same Debian VM, and booting -20 again reproduces the failure.
TEST RESULTS
2026-10-05: 7.0.14-20-pve -> fails
2026-10-08: 7.0.14-19-pve -> works
2026-10-08: 7.0.14-20-pve -> fails again
2026-10-08: 7.0.14-19-pve -> works
2026-10-08: 7.0.14-22-pve -> fails again
2026-10-08: 7.0.14-19-pve -> works
For the October 8 tests:
- VM configuration and generated QEMU command were identical before redaction.
- QEMU remained pve-qemu-kvm 11.0.3-4.
- The Debian guest kernel remained 6.12.111+deb13-amd64.
- No VFIO binding, VM PCI settings, or BIOS configuration changes were made as part of these tests.
On -19, the guest enumerates the Radeon and HDMI audio, binds amdgpu and snd_hda_intel, initializes 8176 MiB VRAM, and creates /dev/dri/renderD128. A rendering workload was not benchmarked/tested.
On -20 or -22, the VM still starts and the guest agent responds, but the guest has only its emulated display and no Radeon render device. Both GPU functions on the host show "Unknown header type 7f". The guest kernel sees invalid PCI configuration for the GPU and does not enumerate it as a usable device.
FAILURE OUTPUT
VM startup on -20:
Code:
error writing '1' to '/sys/bus/pci/devices/0000:03:00.0/reset': Inappropriate ioctl for device
failed to reset PCI device '0000:03:00.0', but trying to continue as not all devices need a reset
kvm: vfio: Unable to power on device, stuck in D3
kvm: vfio: Unable to power on device, stuck in D3
TASK OK
Important: the initial reset ioctl warning ALSO appears on the working -19 boot. The distinguishing symptoms on -20 are the D3 errors and inaccessible PCI configuration, not that initial warning alone.
Host kernel log on the failing boot includes:
Code:
vfio-pci 0000:03:00.0: Invalid PCI ROM header signature: expecting 0xaa55, got 0xffff
There are also repeated vfio_bar_restore messages. Host lspci -vv shows "Unknown header type 7f" for both 03:00.0 and 03:00.1.
The repeat-failure guest kernel reports:
Code:
pci 0000:06:10.0: [1002:73ff] type 7f class 0xffffff conventional PCI
ENVIRONMENT
- Proxmox VE 9.2; pve-manager 9.2.21
- qemu-server 9.2.10
- pve-qemu-kvm 11.0.3-4
- pve-firmware 3.18-6
- pve-edk2-firmware 4.2026.08-1
- CPU: AMD Ryzen 5 5600G with Radeon Graphics
- System: HP Pavilion Gaming Desktop TG01-2xxx
- Motherboard: HP 8906, version SMVB
- BIOS: AMI F.40, reported release date 04/08/2026
- GPU: AMD Navi 23 [Radeon RX 6600/6600 XT/6600M], PCI ID 1002:73ff, revision c1, subsystem 103c:8a33
- HDMI audio: PCI ID 1002:ab28
- Host PCI addresses: 0000:03:00.0 and 0000:03:00.1
- Separate IOMMU groups: 11 and 12; AMD interrupt remapping enabled
- A separate integrated Cezanne GPU remains on the host
PASSTHROUGH SETUP
Code:
bios: ovmf
machine: q35
cpu: host
cores: 8
memory: 24576
hostpci0: 0000:03:00
vga: std
Both GPU functions are assigned through the same hostpci0 entry. The current configuration does not specify pcie=1 or a custom ROM file. The fuller redacted VM configuration and relevant QEMU device arguments are attached.
The host amdgpu driver initializes the discrete GPU during boot, then Proxmox detaches it and binds both functions to vfio-pci at VM startup. This same handoff occurs on both tested host kernels. Early VFIO binding has not been tested as a workaround.
The GPU's reset_method is "bus". vfio_pci disable_idle_d3 is N. Runtime power/control is auto. D3hot alone is not a reliable failure indicator here: it was also observed while the successfully initialized GPU was idle on -19.
UPGRADE CONTEXT
The October 5 transaction installed kernel -20 and also updated QEMU from 11.0.3-3 to 11.0.3-4. However, the October 8 comparison kept QEMU at 11.0.3-4 while changing the host kernel, and passthrough worked on -19 and failed again on -20.
The original October 5 guest used Debian kernel 6.12.107+deb13-amd64. Both October 8 tests used 6.12.111+deb13-amd64, so that guest update does not explain the October 8 working/failing comparison.
CURRENT WORKAROUND / QUESTION
The host is now persistently pinned to boot 7.0.14-19-pve with proxmox-boot-tool, and that kernel package is marked manually installed to retain it as a fallback. The pin applies on the next reboot; it does not change the running kernel.
Is there a known amdgpu, PCI power-management, or VFIO regression between these two Proxmox kernel builds? Is there a specific fix or newer kernel worth testing? I have not tested newer kernels or identified an offending commit.
Attached are redacted diagnostic excerpts for the working boot and both failures. Host/guest names, VM identifiers, network identifiers, UUIDs, storage names, and deployment-specific paths have been removed or replaced. Hardware IDs, PCI addresses, versions, timestamps, and relevant errors are preserved.
The attached files only contain the -20 reboots.
Thanks for any help!