Reproducible Radeon Navi 23 passthrough regression: 7.0.14-19 works; 7.0.14-20 and -22 leaves GPU stuck in D3

nblomquist

New Member
Aug 7, 2025
6
3
3
Hi,

Possibly related thread: https://forum.proxmox.com/threads/r...e-aer-dpc-error-–-7-0-14-19-pve-works.186785/

I can reproduce an AMD Radeon PCI passthrough failure when switching a Proxmox host from 7.0.14-19-pve to 7.0.14-20-pve. Booting the previous kernel restores GPU initialization inside the same Debian VM, and booting -20 again reproduces the failure.

TEST RESULTS

2026-10-05: 7.0.14-20-pve -> fails
2026-10-08: 7.0.14-19-pve -> works
2026-10-08: 7.0.14-20-pve -> fails again
2026-10-08: 7.0.14-19-pve -> works
2026-10-08: 7.0.14-22-pve -> fails again
2026-10-08: 7.0.14-19-pve -> works

For the October 8 tests:
- VM configuration and generated QEMU command were identical before redaction.
- QEMU remained pve-qemu-kvm 11.0.3-4.
- The Debian guest kernel remained 6.12.111+deb13-amd64.
- No VFIO binding, VM PCI settings, or BIOS configuration changes were made as part of these tests.

On -19, the guest enumerates the Radeon and HDMI audio, binds amdgpu and snd_hda_intel, initializes 8176 MiB VRAM, and creates /dev/dri/renderD128. A rendering workload was not benchmarked/tested.

On -20 or -22, the VM still starts and the guest agent responds, but the guest has only its emulated display and no Radeon render device. Both GPU functions on the host show "Unknown header type 7f". The guest kernel sees invalid PCI configuration for the GPU and does not enumerate it as a usable device.

FAILURE OUTPUT

VM startup on -20:

Code:
error writing '1' to '/sys/bus/pci/devices/0000:03:00.0/reset': Inappropriate ioctl for device
failed to reset PCI device '0000:03:00.0', but trying to continue as not all devices need a reset
kvm: vfio: Unable to power on device, stuck in D3
kvm: vfio: Unable to power on device, stuck in D3
TASK OK

Important: the initial reset ioctl warning ALSO appears on the working -19 boot. The distinguishing symptoms on -20 are the D3 errors and inaccessible PCI configuration, not that initial warning alone.

Host kernel log on the failing boot includes:

Code:
vfio-pci 0000:03:00.0: Invalid PCI ROM header signature: expecting 0xaa55, got 0xffff

There are also repeated vfio_bar_restore messages. Host lspci -vv shows "Unknown header type 7f" for both 03:00.0 and 03:00.1.

The repeat-failure guest kernel reports:

Code:
pci 0000:06:10.0: [1002:73ff] type 7f class 0xffffff conventional PCI

ENVIRONMENT

- Proxmox VE 9.2; pve-manager 9.2.21
- qemu-server 9.2.10
- pve-qemu-kvm 11.0.3-4
- pve-firmware 3.18-6
- pve-edk2-firmware 4.2026.08-1
- CPU: AMD Ryzen 5 5600G with Radeon Graphics
- System: HP Pavilion Gaming Desktop TG01-2xxx
- Motherboard: HP 8906, version SMVB
- BIOS: AMI F.40, reported release date 04/08/2026
- GPU: AMD Navi 23 [Radeon RX 6600/6600 XT/6600M], PCI ID 1002:73ff, revision c1, subsystem 103c:8a33
- HDMI audio: PCI ID 1002:ab28
- Host PCI addresses: 0000:03:00.0 and 0000:03:00.1
- Separate IOMMU groups: 11 and 12; AMD interrupt remapping enabled
- A separate integrated Cezanne GPU remains on the host

PASSTHROUGH SETUP

Code:
bios: ovmf
machine: q35
cpu: host
cores: 8
memory: 24576
hostpci0: 0000:03:00
vga: std

Both GPU functions are assigned through the same hostpci0 entry. The current configuration does not specify pcie=1 or a custom ROM file. The fuller redacted VM configuration and relevant QEMU device arguments are attached.

The host amdgpu driver initializes the discrete GPU during boot, then Proxmox detaches it and binds both functions to vfio-pci at VM startup. This same handoff occurs on both tested host kernels. Early VFIO binding has not been tested as a workaround.

The GPU's reset_method is "bus". vfio_pci disable_idle_d3 is N. Runtime power/control is auto. D3hot alone is not a reliable failure indicator here: it was also observed while the successfully initialized GPU was idle on -19.

UPGRADE CONTEXT

The October 5 transaction installed kernel -20 and also updated QEMU from 11.0.3-3 to 11.0.3-4. However, the October 8 comparison kept QEMU at 11.0.3-4 while changing the host kernel, and passthrough worked on -19 and failed again on -20.

The original October 5 guest used Debian kernel 6.12.107+deb13-amd64. Both October 8 tests used 6.12.111+deb13-amd64, so that guest update does not explain the October 8 working/failing comparison.

CURRENT WORKAROUND / QUESTION

The host is now persistently pinned to boot 7.0.14-19-pve with proxmox-boot-tool, and that kernel package is marked manually installed to retain it as a fallback. The pin applies on the next reboot; it does not change the running kernel.

Is there a known amdgpu, PCI power-management, or VFIO regression between these two Proxmox kernel builds? Is there a specific fix or newer kernel worth testing? I have not tested newer kernels or identified an offending commit.

Attached are redacted diagnostic excerpts for the working boot and both failures. Host/guest names, VM identifiers, network identifiers, UUIDs, storage names, and deployment-specific paths have been removed or replaced. Hardware IDs, PCI addresses, versions, timestamps, and relevant errors are preserved.

The attached files only contain the -20 reboots.

Thanks for any help!
 

Attachments

Last edited:
Thanks, I have now tested both suggestions on official 7.0.14-22-pve and isolated the kernel change causing the regression.

Summary:
- amdgpu.aspm=0: still fails on -22.
- Binding the GPU to VFIO before amdgpu: works on -22 for guest GPU initialization.
- Removing one amdgpu autosuspend-cleanup change fixes -20; adding it to -19 reproduces the failure. Both directions were confirmed with fresh clean builds.

IDENTIFIED COMMIT

Ubuntu kernel commit:
af6fecb3558161c483db9399723321e63347f4f0

Upstream commit:
ef5fcf2a6c320676bf8be2dadac93d9023b468b7

Title: drm/amdgpu: fix autosuspend cleanup during removal

The change adds this line after pm_runtime_forbid() in amdgpu_pci_remove():

pm_runtime_dont_use_autosuspend(dev->dev);

Source-isolation results, using the original host-amdgpu-to-VFIO handoff:

1. Official 7.0.14-19-pve works; official -20 and -22 fail.
2. Clean local rebuild of the exact Proxmox -19 endpoint works.
3. Clean local rebuild of the exact Proxmox -20 endpoint fails.
4. -20 with only the candidate line removed works.
5. -19 with only the candidate line added fails with the same signature.
6. Repeated both modified builds from fresh source/output trees, without reused compiled objects: clean -20 revert works; clean -19 add-back fails.