Intel 12400 iGPU Passthough suddenly stops to work.

HariSeldon

Member
Nov 6, 2023
19
0
6
Hi everyone,

I’ve my pve working perfectly for 3 years or so.

After I decided to update to proxmox 9, everything was working good at a first glance.

The only issue I have is that my Debian 13 VM, after few days it’s running and using the igpu with pass through the igpu disappear in the VM. If I reboot pve everything come back to work.

The gpu pass though stop to work by itself after 1 week or so… I need to reboot to make appear again the GPU in the guest.

Any help how to debug this?

I was trying to read dmesg on pve without anything helpful found.
 
In the log I find a lot of this:


root@pve:~# journalctl -kp warning
Jul 19 11:30:18 pve kernel: r8169 0000:06:00.0: can't disable ASPM; OS doesn't have ASPM control
Jul 19 11:30:18 pve kernel: usb: port power management may be unreliable
Jul 19 11:30:18 pve kernel: spl: loading out-of-tree module taints kernel.
Jul 19 11:30:18 pve kernel: zfs: module license 'CDDL' taints kernel.
Jul 19 11:30:18 pve kernel: Disabling lock debugging due to kernel taint
Jul 19 11:30:18 pve kernel: zfs: module license taints kernel.
Jul 19 11:30:18 pve kernel: spi-nor spi0.0: supply vcc not found, using dummy regulator
Jul 19 11:30:18 pve kernel: faux_driver regulatory: Direct firmware load for regulatory.db failed with error -2
Jul 19 11:30:19 pve kernel: nvme nvme0: using unchecked data buffer
Jul 19 11:30:24 pve kernel: vfio-pci 0000:00:02.0: resetting
Jul 19 11:30:24 pve kernel: vfio-pci 0000:00:02.0: reset done
Jul 19 11:30:26 pve kernel: vfio-pci 0000:00:02.0: resetting
Jul 19 11:30:26 pve kernel: vfio-pci 0000:00:02.0: reset done
Jul 19 11:30:26 pve kernel: vfio-pci 0000:00:02.0: resetting
Jul 19 11:30:26 pve kernel: vfio-pci 0000:00:02.0: reset done
Jul 19 11:30:27 pve kernel: vfio-pci 0000:03:00.0: resetting
Jul 19 11:30:27 pve kernel: vfio-pci 0000:03:00.0: reset done
Jul 19 11:30:27 pve kernel: vfio-pci 0000:05:00.0: resetting
Jul 19 11:30:28 pve kernel: vfio-pci 0000:05:00.0: reset done
Jul 19 11:30:29 pve kernel: platform INT3515:01: deferred probe pending: Serial bus multi instantiate pseudo device driver: Error creating i2c-client, idx 0
Jul 19 11:30:29 pve kernel: vfio-pci 0000:03:00.0: resetting
Jul 19 11:30:29 pve kernel: vfio-pci 0000:03:00.0: reset done
Jul 19 11:30:29 pve kernel: vfio-pci 0000:05:00.0: resetting
Jul 19 11:30:29 pve kernel: vfio-pci 0000:05:00.0: reset done
Jul 19 11:30:29 pve kernel: vfio-pci 0000:03:00.0: resetting
Jul 19 11:30:29 pve kernel: vfio-pci 0000:03:00.1: resetting
Jul 19 11:30:30 pve kernel: vfio-pci 0000:03:00.0: reset done
Jul 19 11:30:30 pve kernel: vfio-pci 0000:03:00.1: reset done
Jul 19 11:30:30 pve kernel: vfio-pci 0000:05:00.0: resetting
Jul 19 11:30:30 pve kernel: vfio-pci 0000:05:00.0: reset done
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored rdmsr: 0x492 data 0x0
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored rdmsr: 0x1a6 data 0x0
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored wrmsr: 0x1a6 data 0x11
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored rdmsr: 0x1a6 data 0x0
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored rdmsr: 0x1a7 data 0x0
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored wrmsr: 0x1a7 data 0x11
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored rdmsr: 0x1a7 data 0x0
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored rdmsr: 0x3f6 data 0x0
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored wrmsr: 0x3f6 data 0x11
Jul 19 11:30:37 pve kernel: kvm: kvm [1270]: ignored rdmsr: 0x3f6 data 0x0
Jul 19 11:30:43 pve kernel: kvm_do_msr_access: 61 callbacks suppressed
Jul 19 11:30:43 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0x64d data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0x3f1 data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0x600 data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0x3f2 data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0xc0010200 data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0xc0010202 data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0xc0010204 data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0xc0010206 data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0xc0010208 data 0x0
Jul 19 11:30:44 pve kernel: kvm: kvm [1481]: ignored rdmsr: 0xc001020a data 0x0
Jul 19 11:31:01 pve kernel: x86/split lock detection: #AC: CPU 0/KVM/1405 took a split_lock trap at address: 0x704ccf4d
Jul 19 11:31:01 pve kernel: x86/split lock detection: #AC: CPU 1/KVM/1406 took a split_lock trap at address: 0x704ccf4d

I suppose that once In a week this can create an issue and the reset is not done correctly...how to fix this?
 
Last edited:
I've disabled the split lock detection and the idle state of the iGPU, with this configuration:


Code:
GRUB_DEFAULT=0
GRUB_TIMEOUT=5
GRUB_DISTRIBUTOR=`( . /etc/os-release && echo ${NAME} )`
GRUB_CMDLINE_LINUX_DEFAULT="quiet intel_iommu=on iommu=pt initcall_blacklist=sysfb_init split_lock_detect=off"
GRUB_CMDLINE_LINUX=""

#/etc/moprobe.d/vfio.conf
options vfio-pci ids=8086:4682,8086:7a50,8086:7a62,1002:743f,1002:1478,1002:1479,1002:ab28,8086:125c disable_vga=1 disable_idle_d3=1
options vfio_iommu_type1 allow_unsafe_interrupts=1


#etc/modprobe.d/pve-blacklist.conf
blacklist nvidiafb
blacklist radeon
blacklist amdgpu
blacklist nouveau
blacklist nvidia
blacklist nvidia-gpu
blacklist snd_hda_intel
blacklist snd_hda_codec_hdmi
blacklist snd_sof_pci_intel_tgl
blacklist snd_soc_avs
blacklist i915
blacklist ahci
blacklist igc
blacklist xe


#/etc/modules
vfio
vfio_iommu_type1
vfio_pci
vfio_virqfd

and now the output is much cleaner, even some reset still there, all the warning are only related to about 1 hour ago during the host boot, then is clean:


Code:
root@pve:~# journalctl -kp warning
Jul 19 11:50:44 pve kernel: r8169 0000:06:00.0: can't disable ASPM; OS doesn't have ASPM control
Jul 19 11:50:44 pve kernel: usb: port power management may be unreliable
Jul 19 11:50:44 pve kernel: spl: loading out-of-tree module taints kernel.
Jul 19 11:50:44 pve kernel: zfs: module license 'CDDL' taints kernel.
Jul 19 11:50:44 pve kernel: Disabling lock debugging due to kernel taint
Jul 19 11:50:44 pve kernel: zfs: module license taints kernel.
Jul 19 11:50:44 pve kernel: faux_driver regulatory: Direct firmware load for regulatory.db failed with error -2
Jul 19 11:50:44 pve kernel: spi-nor spi0.0: supply vcc not found, using dummy regulator
Jul 19 11:50:45 pve kernel: nvme nvme0: using unchecked data buffer
Jul 19 11:50:50 pve kernel: vfio-pci 0000:00:02.0: resetting
Jul 19 11:50:50 pve kernel: vfio-pci 0000:00:02.0: reset done
Jul 19 11:50:50 pve kernel: kvm: Running KVM with ignore_msrs=1 and report_ignored_msrs=0 is not a
                            a supported configuration.  Lying to the guest about the existence of MSRs
                            may cause the guest operating system to hang or produce errors.  If a guest
                            does not run without ignore_msrs=1, please report it to kvm@vger.kernel.org.
Jul 19 11:50:52 pve kernel: vfio-pci 0000:00:02.0: resetting
Jul 19 11:50:52 pve kernel: vfio-pci 0000:00:02.0: reset done
Jul 19 11:50:52 pve kernel: vfio-pci 0000:00:02.0: resetting
Jul 19 11:50:52 pve kernel: vfio-pci 0000:00:02.0: reset done
Jul 19 11:50:53 pve kernel: vfio-pci 0000:03:00.0: resetting
Jul 19 11:50:53 pve kernel: vfio-pci 0000:03:00.0: reset done
Jul 19 11:50:53 pve kernel: vfio-pci 0000:05:00.0: resetting
Jul 19 11:50:54 pve kernel: vfio-pci 0000:05:00.0: reset done
Jul 19 11:50:55 pve kernel: platform INT3515:01: deferred probe pending: Serial bus multi instantiate pseudo device driver: Error creating i2c-client, idx 0
Jul 19 11:50:55 pve kernel: vfio-pci 0000:03:00.0: resetting
Jul 19 11:50:55 pve kernel: vfio-pci 0000:03:00.0: reset done
Jul 19 11:50:55 pve kernel: vfio-pci 0000:05:00.0: resetting
Jul 19 11:50:55 pve kernel: vfio-pci 0000:05:00.0: reset done
Jul 19 11:50:55 pve kernel: vfio-pci 0000:03:00.0: resetting
Jul 19 11:50:55 pve kernel: vfio-pci 0000:03:00.1: resetting
Jul 19 11:50:56 pve kernel: vfio-pci 0000:03:00.0: reset done
Jul 19 11:50:56 pve kernel: vfio-pci 0000:03:00.1: reset done
Jul 19 11:50:56 pve kernel: vfio-pci 0000:05:00.0: resetting
Jul 19 11:50:56 pve kernel: vfio-pci 0000:05:00.0: reset done
 
and the winner is: your unknown hardware or your old bios code?

Unless we can determine whether it’s just the GPU that’s missing or if the virtual machine is hanging, there’s nothing we can do.



Like he said, I’m also guessing this might be a BIOS (or, more precisely, VBIOS) issue.
*If they say, “I didn’t specify that,” I suppose I’ll have no choice but to reply, “I haven’t been told anything about that, so I don’t know what to do.”

Since we don’t know what version he updated from to Proxmox 9 or the virtual machine settings, all we can really say is, “I wonder what everyone else did?”

If it’s running fine after upgrading from Proxmox 8 to 9, it’s possible that some fixes were implemented. Previously, with the pre-update settings, the screen wouldn’t even display after the update.

If you've been tinkering with various iGPU settings in your virtual machine configuration, you should search the forum for threads regarding iGPU issues in PVE9, review your settings, and then check to see if the problem still occurs.
What I mean in the first post is that ONLY the gpu is missing in the VM after one week or so that is working perfectly.

And this is obviously NOT related to any vbios as it was perfectly working for 3+ years with pve8.0.

What I discovered is that maybe is related to the VM startup: it seems triggered from the weekly backup done in “stop” method. If you turn off and turn on again the machine (sometimes not always) the GPU is not passed through…
 
Since I'm not sure if you're actually willing to fix this, I'm going to step back. Sorry. I deleted my previous post.
I've never experienced anything like this with the Intel N100, which is a CPU from a similar generation.
I hope you'll be able to fix this.
 
  • Like
Reactions: news
Since I'm not sure if you're actually willing to fix this, I'm going to step back. Sorry. I deleted my previous post.
I've never experienced anything like this with the Intel N100, which is a CPU from a similar generation.
I hope you'll be able to fix this.
Obviously that I’m willing to fix this, just I really don’t understand why is pointing out the vbios if I said from the beginning that is something software related to update from proxmox 8 to 9…due to the fact that for 3+ years everything was working correctly and config didn’t changed, just updated to pve9.

It’s like asking for color of car seat when you have an issue with the engine…
 
There have been threads where a problems after a recent Linux kernel update and a motherboard BIOS update fixed it. Passthrough is often influenced by both kernel and BIOS changes.
You could say the software update caused the issue but since it appears to trigger a problem in the BIOS and a BIOS fix resolved the issue, it's unlikely that the kernel will be changed. And sometimes the color of the seat can tell which specific engine is involved (because of limited edition fabric only for a specific car type or something but I feel like I'm stretching this analogy beyond it's breaking point now).
Maybe you can try different kernel versions between the 6.8.x (PVE 8) and the current 7.0.y one to pin down when the system becomes unstable? You can install older proxmox-kernel packages on PVE 9.
 
  • Like
Reactions: news
There have been threads where a problems after a recent Linux kernel update and a motherboard BIOS update fixed it. Passthrough is often influenced by both kernel and BIOS changes.
You could say the software update caused the issue but since it appears to trigger a problem in the BIOS and a BIOS fix resolved the issue, it's unlikely that the kernel will be changed. And sometimes the color of the seat can tell which specific engine is involved (because of limited edition fabric only for a specific car type or something but I feel like I'm stretching this analogy beyond it's breaking point now).
Maybe you can try different kernel versions between the 6.8.x (PVE 8) and the current 7.0.y one to pin down when the system becomes unstable? You can install older proxmox-kernel packages on PVE 9.
Thanks for the reply, I understand that I can perform some kind of “trial & error” process, changing this and that.

However Initial point was related to understand if someone knows which are the possibile causes of a iGPU passed though and working, “suddenly” disappear from the guest.

Is a known problem? Someone knows if it’s already happened? Which can be the causes ?
Kernel update: ok but what? Power saving management that doesn’t wake up the gpu sometimes? Kernel modules not load correctly? Just guessing here and there.

Saying that I couldn’t be helped because of VBIOS not provided…it’s really unbelievable.

It’s like you ask on some forum support on why your AMD Radeon card crash with some game configuration, and you get answered “Buy Nvidia”…
 
Last edited: