NVIDIA vGPU passthrough - PCI AER and ReBAR not working properly

tiopz

Member
Jan 31, 2024
11
3
8
Hi,

I have some issues with NVIDIA vGPU functionality. I followed this guide and everything is working fine : I tested both Linux and Windows VM.

However, when starting the VM, these errors are shown
Code:
kvm: -device vfio-pci,host=0000:91:00.2,id=hostpci0,bus=ich9-pcie-port-1,addr=0x0: warning: vfio 0000:91:00.2: Could not enable error recovery for the device
kvm: -device vfio-pci,host=0000:91:00.2,id=hostpci0,bus=ich9-pcie-port-1,addr=0x0: warning: VFIO dma-buf not supported in kernel: PCI BAR IOMMU mappings may fail

I exchanged with Gigabyte support (the manufacturer of the server) and confirmed that BIOS settings are correct (BIOS is up to date as well) :
- PCI AER and ARI are enabled
- SR-IOV and IOMMU are enabled
- ReBAR is enabled

Here is the configuration of the VM
Code:
agent: 1
bios: ovmf
boot: order=scsi0
cores: 16
cpu: host
efidisk0: nvme:vm-9998-disk-0,efitype=4m,ms-cert=2023k,pre-enrolled-keys=1,size=1M
hostpci0: mapping=rtx-pro-6000,mdev=nvidia-1561,pcie=1
ipconfig0: xxx
machine: q35
memory: 32768
meta: creation-qemu=9.2.0,ctime=1766391077
name: vgpu-linux
net0: xxx
numa: 0
ostype: l26
scsi0: nvme:vm-9998-disk-1,iothread=1,size=50G
scsihw: virtio-scsi-single
smbios1: uuid=xxx
sockets: 1
startup: up=120
vmgenid: xxx

Here are the packages version
Code:
proxmox-ve: 9.2.0 (running kernel: 7.0.14-5-pve)
pve-manager: 9.2.4 (running version: 9.2.4/5e5ae681198514d4)
proxmox-kernel-helper: 9.2.0
proxmox-kernel-7.0.14-5-pve-signed: 7.0.14-5
proxmox-kernel-7.0: 7.0.14-5
proxmox-kernel-7.0.12-1-pve-signed: 7.0.12-1
proxmox-kernel-7.0.6-2-pve-signed: 7.0.6-2
proxmox-kernel-7.0.2-4-pve-signed: 7.0.2-4
proxmox-kernel-6.17.13-18-pve-signed: 6.17.13-18
proxmox-kernel-6.17: 6.17.13-18
proxmox-kernel-6.17.13-13-pve-signed: 6.17.13-13
proxmox-kernel-6.17.13-9-pve-signed: 6.17.13-9
proxmox-kernel-6.17.2-1-pve-signed: 6.17.2-1
amd64-microcode: 3.20251202.1~bpo13+1
ceph-fuse: 19.2.3-pve2
corosync: 3.1.10-pve2
criu: 4.1.1-1
frr-pythontools: 10.6.1-1+pve2
ifupdown2: 3.3.0-1+pmx12
ksm-control-daemon: 1.5-1
libjs-extjs: 7.0.0-5
libproxmox-acme-perl: 1.7.1
libproxmox-backup-qemu0: 2.0.2
libproxmox-rs-perl: 0.4.1
libpve-access-control: 9.1.1
libpve-apiclient-perl: 3.4.2
libpve-cluster-api-perl: 9.1.6
libpve-cluster-perl: 9.1.6
libpve-common-perl: 9.1.17
libpve-guest-common-perl: 6.0.4
libpve-http-server-perl: 6.0.5
libpve-network-perl: 1.6.6
libpve-notify-perl: 9.1.6
libpve-rs-perl: 0.15.3
libpve-storage-perl: 9.1.6
libspice-server1: 0.15.2-1+b1
lvm2: 2.03.31-2+pmx1
lxc-pve: 7.0.0-2
lxcfs: 7.0.0-pve1
novnc-pve: 1.7.0-2
proxmox-backup-client: 4.2.3-1
proxmox-backup-file-restore: 4.2.3-1
proxmox-backup-restore-image: 1.0.0
proxmox-firewall: 1.2.3
proxmox-kernel-helper: 9.2.0
proxmox-mail-forward: 1.0.3
proxmox-mini-journalreader: 1.7
proxmox-offline-mirror-helper: 0.7.4
proxmox-widget-toolkit: 5.2.6
pve-cluster: 9.1.6
pve-container: 6.1.11
pve-docs: 9.2.3
pve-edk2-firmware: 4.2025.05-2
pve-esxi-import-tools: 1.0.1
pve-firewall: 6.0.5
pve-firmware: 3.18-5
pve-ha-manager: 5.2.4
pve-i18n: 3.9.0
pve-qemu-kvm: 11.0.2-1
pve-xtermjs: 6.0.0-2
qemu-server: 9.2.0
smartmontools: 7.5-pve2
spiceterm: 3.4.2
swtpm: 0.8.0+pve3
vncterm: 1.9.2
zfsutils-linux: 2.4.3-pve1

The NVIDIA vGPU driver version is 595.71.03
The card is an NVIDIA RTX Pro 6000

Running kernel version is Linux server 7.0.14-5-pve #1 SMP PREEMPT_DYNAMIC PMX 7.0.14-5 (2026-07-14T12:32Z) x86_64 GNU/Linux

How can I fix the error mentionned above ?
 
Looks more like a VFIO/kernel compatibility issue than a BIOS/ReBAR problem.

Since the vGPU itself works, I would first verify the physical GPU BAR sizes with

lspci -vv

and then boot the same setup with the installed 6.17 kernel.

If the „VFIO dma-buf not supported in kernel“ warning disappears or the BAR behavior changes, that strongly points to a 7.0/VFIO regression or compatibility issue rather than ReBAR being misconfigured.
 
Hi,

Thank you, here is the output of lspci -vv
Code:
91:00.0 3D controller: NVIDIA Corporation GB202GL [RTX PRO 6000 Blackwell Server Edition] (rev a1)
        Subsystem: NVIDIA Corporation Device 204e
        Control: I/O- Mem+ BusMaster+ SpecCycle- MemWINV- VGASnoop- ParErr- Stepping- SERR- FastB2B- DisINTx+
        Status: Cap+ 66MHz- UDF- FastB2B- ParErr- DEVSEL=fast >TAbort- <TAbort- <MAbort- >SERR- <PERR- INTx-
        Latency: 0
        Interrupt: pin A routed to IRQ 523
        NUMA node: 6
        IOMMU group: 120
        Region 0: Memory at 1ef000000000 (64-bit, prefetchable) [size=64M]
        Region 2: Memory at 1ea000000000 (64-bit, prefetchable) [size=128G]
        Region 4: Memory at 1ef064000000 (64-bit, prefetchable) [size=32M]
        Capabilities: [40] Power Management version 3
                Flags: PMEClk- DSI- D1- D2- AuxCurrent=0mA PME(D0+,D1-,D2-,D3hot+,D3cold-)
                Status: D0 NoSoftRst+ PME-Enable- DSel=0 DScale=0 PME-
        Capabilities: [48] MSI: Enable- Count=1/16 Maskable+ 64bit+
                Address: 0000000000000000  Data: 0000
                Masking: 00000000  Pending: 00000000
        Capabilities: [60] Express (v2) Legacy Endpoint, IntMsgNum 0
                DevCap: MaxPayload 256 bytes, PhantFunc 0, Latency L0s <64ns, L1 unlimited
                        ExtTag+ AttnBtn- AttnInd- PwrInd- RBE+ FLReset+ TEE-IO-
                DevCtl: CorrErr+ NonFatalErr+ FatalErr+ UnsupReq-
                        RlxdOrd+ ExtTag+ PhantFunc- AuxPwr- NoSnoop- FLReset-
                        MaxPayload 256 bytes, MaxReadReq 512 bytes
                DevSta: CorrErr- NonFatalErr- FatalErr- UnsupReq- AuxPwr- TransPend-
                LnkCap: Port #0, Speed 32GT/s, Width x16, ASPM L1, Exit Latency L1 unlimited
                        ClockPM+ Surprise- LLActRep- BwNot- ASPMOptComp+
                LnkCtl: ASPM Disabled; RCB 64 bytes, LnkDisable- CommClk+
                        ExtSynch- ClockPM- AutWidDis- BWInt- AutBWInt-
                LnkSta: Speed 2.5GT/s (downgraded), Width x16
                        TrErr- Train- SlotClk+ DLActive- BWMgmt- ABWMgmt-
                DevCap2: Completion Timeout: Range AB, TimeoutDis+ NROPrPrP- LTR+
                         10BitTagComp+ 10BitTagReq+ OBFF Via message, ExtFmt- EETLPPrefix-
                         EmergencyPowerReduction Not Supported, EmergencyPowerReductionInit-
                         FRS-
                         AtomicOpsCap: 32bit+ 64bit+ 128bitCAS-
                DevCtl2: Completion Timeout: 50us to 50ms, TimeoutDis-
                         AtomicOpsCtl: ReqEn+
                         IDOReq- IDOCompl- LTR+ EmergencyPowerReductionReq-
                         10BitTagReq+ OBFF Disabled, EETLPPrefixBlk-
                LnkCap2: Supported Link Speeds: 2.5-32GT/s, Crosslink- Retimer+ 2Retimers+ DRS-
                LnkCtl2: Target Link Speed: 32GT/s, EnterCompliance- SpeedDis-
                         Transmit Margin: Normal Operating Range, EnterModifiedCompliance- ComplianceSOS-
                         Compliance Preset/De-emphasis: -6dB de-emphasis, 0dB preshoot
                LnkSta2: Current De-emphasis Level: -6dB, EqualizationComplete+ EqualizationPhase1+
                         EqualizationPhase2+ EqualizationPhase3+ LinkEqualizationRequest-
                         Retimer- 2Retimers- CrosslinkRes: unsupported
        Capabilities: [9c] Vendor Specific Information: Len=14 <?>
        Capabilities: [b0] MSI-X: Enable+ Count=9 Masked-
                Vector table: BAR=0 offset=00b90000
                PBA: BAR=0 offset=00ba0000
        Capabilities: [100 v1] Secondary PCI Express
                LnkCtl3: LnkEquIntrruptEn- PerformEqu-
                LaneErrStat: 0
        Capabilities: [12c v1] Latency Tolerance Reporting
                Max snoop latency: 1048576ns
                Max no snoop latency: 1048576ns
        Capabilities: [134 v1] Physical Resizable BAR
                BAR 2: current size: 128GB, supported: 64MB 128MB 256MB 512MB 1GB 2GB 4GB 8GB 16GB 32GB 64GB 128GB
        Capabilities: [140 v1] Virtual Resizable BAR
                BAR 2: current size: 4GB, supported: 64MB 128MB 256MB 512MB 1GB 2GB 4GB 8GB 16GB 32GB 64GB 128GB 256TB 512TB 1PB 2PB 4PB 8PB 16PB 32PB 64PB 128PB 256PB 512PB 1EB 2EB 4EB 8EB
        Capabilities: [14c v1] Data Link Feature <?>
        Capabilities: [158 v1] Physical Layer 16.0 GT/s <?>
        Capabilities: [188 v1] Physical Layer 32.0 GT/s <?>
        Capabilities: [1b8 v2] Advanced Error Reporting
                UESta:  DLP- SDES- TLP- FCP- CmpltTO- CmpltAbrt- UnxCmplt- RxOF- MalfTLP-
                        ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-
                        PoisonTLPBlocked- DMWrReqBlocked- IDECheck- MisIDETLP- PCRC_CHECK- TLPXlatBlocked-
                UEMsk:  DLP- SDES- TLP- FCP- CmpltTO- CmpltAbrt- UnxCmplt- RxOF- MalfTLP-
                        ECRC- UnsupReq- ACSViol- UncorrIntErr- BlockedTLP- AtomicOpBlocked- TLPBlockedErr-
                        PoisonTLPBlocked- DMWrReqBlocked- IDECheck- MisIDETLP- PCRC_CHECK- TLPXlatBlocked-
                UESvrt: DLP+ SDES+ TLP- FCP+ CmpltTO+ CmpltAbrt- UnxCmplt+ RxOF+ MalfTLP+
                        ECRC+ UnsupReq- ACSViol- UncorrIntErr+ BlockedTLP- AtomicOpBlocked- TLPBlockedErr-
                        PoisonTLPBlocked- DMWrReqBlocked- IDECheck- MisIDETLP- PCRC_CHECK- TLPXlatBlocked-
                CESta:  RxErr- BadTLP- BadDLLP- Rollover- Timeout- AdvNonFatalErr- CorrIntErr- HeaderOF-
                CEMsk:  RxErr+ BadTLP+ BadDLLP+ Rollover+ Timeout+ AdvNonFatalErr+ CorrIntErr+ HeaderOF+
                AERCap: First Error Pointer: 00, ECRCGenCap+ ECRCGenEn+ ECRCChkCap+ ECRCChkEn+
                        MultHdrRecCap- MultHdrRecEn- TLPPfxPres- HdrLogCap-
                HeaderLog: 00000000 00000000 00000000 00000000
        Capabilities: [200 v1] Lane Margining at the Receiver
                PortCap: Uses Driver+
                PortSta: MargReady- MargSoftReady-
        Capabilities: [248 v1] Alternative Routing-ID Interpretation (ARI)
                ARICap: MFVC- ACS-, Next Function: 0
                ARICtl: MFVC- ACS-, Function Group: 0
        Capabilities: [250 v1] Single Root I/O Virtualization (SR-IOV)
                IOVCap: Migration- 10BitTagReq+ IntMsgNum 0
                IOVCtl: Enable+ Migration- Interrupt- MSE+ ARIHierarchy+ 10BitTagReq-
                IOVSta: Migration-
                Initial VFs: 48, Total VFs: 48, Number of VFs: 48, Function Dependency Link: 00
                VF offset: 2, stride: 1, Device ID: 2bb5
                Supported Page Size: 00000573, System Page Size: 00000001
                Region 0: Memory at 00001ef066000000 (64-bit, prefetchable)
                Region 2: Memory at 00001ec000000000 (64-bit, prefetchable)
                Region 4: Memory at 00001ef004000000 (64-bit, prefetchable)
                VF Migration: offset: 00000000, BIR: 0
        Capabilities: [2a4 v1] Vendor Specific Information: ID=0001 Rev=1 Len=014 <?>
        Capabilities: [2bc v1] Power Budgeting <?>
        Capabilities: [2f4 v1] Device Serial Number 10-c2-5a-73-53-2d-b0-48
        Kernel driver in use: nvidia
        Kernel modules: nvidiafb, nouveau, nvidia_vgpu_vfio, nvidia

Will try 6.17 kernel when I can and keep you updated as well.
 
ReBAR looks fine on the host:

Physical Resizable BAR
BAR 2: current size: 128GB
So I think we can rule out the BIOS/ReBAR configuration.

Interesting thing though: the PCIe link is currently at 2.5 GT/s x16 while 32 GT/s x16 is supported. This may just be power management/idle behavior, but I would check lspci -vv again under GPU load.

The 6.17 test would still be useful to narrow down the VFIO dma-buf warning.
 
Hi,

I've tested with 6.17 (6.17.13-21) kernel but same issue. I updated the server as well with these package versions, but even with kernel 7 (7.0.14-11-pve) the error persists

Code:
proxmox-ve: 9.2.0 (running kernel: 7.0.14-11-pve)
pve-manager: 9.2.10 (running version: 9.2.10/43df2e01f27a1a19)
proxmox-kernel-helper: 9.2.0
proxmox-kernel-7.0: 7.0.14-11
proxmox-kernel-7.0.14-11-pve-signed: 7.0.14-11
proxmox-kernel-7.0.14-5-pve-signed: 7.0.14-5
proxmox-kernel-7.0.12-1-pve-signed: 7.0.12-1
proxmox-kernel-7.0.6-2-pve-signed: 7.0.6-2
proxmox-kernel-7.0.2-4-pve-signed: 7.0.2-4
proxmox-kernel-6.17: 6.17.13-21
proxmox-kernel-6.17.13-21-pve-signed: 6.17.13-21
proxmox-kernel-6.17.13-18-pve-signed: 6.17.13-18
proxmox-kernel-6.17.13-13-pve-signed: 6.17.13-13
proxmox-kernel-6.17.13-9-pve-signed: 6.17.13-9
proxmox-kernel-6.17.2-1-pve-signed: 6.17.2-1
amd64-microcode: 3.20251202.1~bpo13+1
ceph-fuse: 19.2.3-pve2
corosync: 3.1.10-pve3
criu: 4.1.1-1
frr-pythontools: 10.6.1-1+pve3
ifupdown2: 3.3.0-1+pmx12
ksm-control-daemon: 1.5-1
libjs-extjs: 7.0.0-7
libproxmox-acme-perl: 1.7.2
libproxmox-backup-qemu0: 2.0.2
libproxmox-rs-perl: 0.4.1
libpve-access-control: 9.1.1
libpve-apiclient-perl: 3.4.2
libpve-cluster-api-perl: 9.1.6
libpve-cluster-perl: 9.1.6
libpve-common-perl: 9.2.1
libpve-guest-common-perl: 6.0.5
libpve-http-server-perl: 6.0.5
libpve-network-perl: 1.6.7
libpve-notify-perl: 9.1.6
libpve-rs-perl: 0.15.3
libpve-storage-perl: 9.1.8
libspice-server1: 0.15.2-1+b1
lvm2: 2.03.31-2+pmx1
lxc-pve: 7.0.0-2
lxcfs: 7.0.0-pve1
novnc-pve: 1.7.0-2
proxmox-backup-client: 4.2.5-1
proxmox-backup-file-restore: 4.2.5-1
proxmox-backup-restore-image: 1.0.0
proxmox-enterprise-support-keyring: 1.1
proxmox-firewall: 1.2.3
proxmox-kernel-helper: 9.2.0
proxmox-mail-forward: 1.0.3
proxmox-mini-journalreader: 1.7
proxmox-offline-mirror-helper: 0.7.4
proxmox-widget-toolkit: 5.2.7
pve-cluster: 9.1.6
pve-container: 6.1.13
pve-docs: 9.2.4
pve-edk2-firmware: 4.2025.05-3
pve-esxi-import-tools: 1.0.1
pve-firewall: 6.0.5
pve-firmware: 3.18-5
pve-ha-manager: 5.2.5
pve-i18n: 3.10.0
pve-qemu-kvm: 11.0.3-2
pve-xtermjs: 6.0.0-2
qemu-server: 9.2.5
smartmontools: 7.5-pve2
spiceterm: 3.4.2
swtpm: 0.8.0+pve3
vncterm: 1.9.2
zfsutils-linux: 2.4.3-pve1

Will try to check when the GPU is underload if the speed of the link is 32 GT/s x16 under load as well.
 
As long as you're using QEMU 11, which supports VFIO_DEVICE_FEATURE_DMA_BUF, you can check whether it works, and any errors will be logged.
Even if it is lowered to 6.17 (6.17.13-21), Qemu will check whether that feature is available, regardless of the version.

https://patchew.org/QEMU/2026031719...260317195323.776669-3-john.levon@nutanix.com/

Code:
VFIO_DEVICE_FEATURE_DMA_BUF should be available in Linux v6.19 and QEMU 11.0.

Code:
proxmox-ve: 9.2.0 (running kernel: 7.0.14-12-pve)
pve-manager: 9.2.10 (running version: 9.2.10/43df2e01f27a1a19)
pve-qemu-kvm: 11.0.3-2

cat /boot/config-$(uname -r) | grep "CONFIG_VFIO_PCI_DMABUF"
CONFIG_VFIO_PCI_DMABUF=y

Isn't the fact that it's still being logged in kernel 7 due to a driver or something?
 
Last edited:
If CONFIG_VFIO_PCI_DMABUF=y is enabled on the 7.0 kernel and QEMU 11 still reports the feature as unavailable, this probably isn’t about kernel version anymore.

The interesting part may be the NVIDIA vGPU VFIO backend itself. This isn’t a regular PCI device bound to vfio-pci; the physical GPU is still using the NVIDIA driver and the VM gets an NVIDIA vGPU/mdev.

So QEMU may support

VFIO_DEVICE_FEATURE_DMA_BUF

and the kernel may support it for vfio-pci, while nvidia_vgpu_vfio simply does not expose that device feature. That would also explain why the warning persists on 7.0 despite CONFIG_VFIO_PCI_DMABUF=y.
 
The first error is benign, we've been getting these for years, it has to do with PCIe bus error recovery. The second error is also benign, you typically see it when your QEMU and kernel versions aren't matched. Are you using the enterprise repo or using custom/dev/test kernels? You must reboot between kernel version changes and also make sure you're 'actually' booting into the correct kernel (eg. kernel pinning etc could cause issues).

But in general, both of them can be safely ignored for vGPU purposes.
 
Last edited:
You should already know which module is currently in effect.

proxmox-ve: 9.2.0 (running kernel: 7.0.14-11-pve)
pve-manager: 9.2.10 (running version: 9.2.10/43df2e01f27a1a19)
         :
pve-qemu-kvm: 11.0.3-2

If I had used QEMU 10, I probably wouldn't have encountered any errors, but since there were no errors, I'm sure I would have been satisfied even without the new features.

Anyway, if you want to know about the errors that are being logged because you can't use the new features in QEMU 11, I think you should ask NVIDIA instead of posting here on this forum.

Even if you ask a question about NVIDIA here, while you might get an answer, no one can guarantee the accuracy of that answer.

edit

https://pve.proxmox.com/wiki/NVIDIA_vGPU_on_Proxmox_VE

Code:
To be eligible for support tickets, you must have an active and valid NVIDIA vGPU entitlement as well as an active and valid Proxmox VE subscription on your cluster, with level Basic, Standard or Premium. See the Proxmox VE Subscription Agreement[4] and the Proxmox Support Offerings[5] for more details.

To get support, open a ticket on the Proxmox Server Solutions Support Portal.

Since they seem to offer support, they might respond if you open a paid support ticket.
 
Last edited:
Hi,

cat /boot/config-$(uname -r) | grep "CONFIG_VFIO_PCI_DMABUF"

I can confirm that I have
Code:
cat /boot/config-7.0.14-11-pve | grep "CONFIG_VFIO_PCI_DMABUF"
CONFIG_VFIO_PCI_DMABUF=y

cat /boot/config-6.17.13-21-pve | grep "CONFIG_VFIO_PCI_DMABUF"

> returns nothing (which seems ok based on what you said)
VFIO_DEVICE_FEATURE_DMA_BUF should be available in Linux v6.19 and QEMU 11.0.

---
The first error is benign, we've been getting these for years, it has to do with PCIe bus error recovery. The second error is also benign, you typically see it when your QEMU and kernel versions aren't matched. Are you using the enterprise repo or using custom/dev/test kernels? You must reboot between kernel version changes and also make sure you're 'actually' booting into the correct kernel (eg. kernel pinning etc could cause issues).

That's strange that even with PCIe AER and ARI enabled in BIOS these errors are showing up ? Or maybe there is a reason that I don't understand ?

I am using the no-subscription repo.
I've rebooted to do my tests and make sure I boot inside the correct kernel version by selecting it in the Advanced boot options in Grub : i've also checked inside the Proxmox web interface the kernel version to make sure I selected the version I wanted to test.

---
Anyway, if you want to know about the errors that are being logged because you can't use the new features in QEMU 11, I think you should ask NVIDIA instead of posting here on this forum.

Even if you ask a question about NVIDIA here, while you might get an answer, no one can guarantee the accuracy of that answer.
Thank you for all of your answers, even if "it is not accurate", it's accurate enough to help me figure out and understand "what is happening" ! Especially trying to understand why these errors are happening despite the fact that the vGPU inside the VM is working properly.

I'm asking these questions here since I am doing some tests with the vGPU feature of the NVIDIA cards and because of the errors mentionned in my first post, I've asked Gigabyte to help me with these : they confirmed the settings in the BIOS are correct and that they couldn't reproduce these issues in RHEL 8 (which is on their supported OS list). Since the server is in production, I couldn't test with this OS, hence the post here.

I'm trying to figure out from where the issue might be coming : not BIOS settings (confirmed by Gigabyte), but is it proxmox or the nvidia driver ?

But in general, both of them can be safely ignored for vGPU purposes.
I've done some tests despite these errors and indeed this didn't lead to any issues inside the VM.

the kernel may support it for vfio-pci, while nvidia_vgpu_vfio simply does not expose that device feature. That would also explain why the warning persists on 7.0 despite CONFIG_VFIO_PCI_DMABUF=y.

I think that this is the last part that I had not "touched" yet. The NVIDIA vGPU driver is at version 595.71.03 on the server. The latest version available is 595.91.04
Maybe updating to the latest version, or going to drivers version 6xx (not available as of today as a vGPU driver) might resolve some issues as well.
I can't update it now but will keep this post updated to tell you if updating changed something.
 
That's strange that even with PCIe AER and ARI enabled in BIOS these errors are showing up ? Or maybe there is a reason that I don't understand ?

Linux v6.19 and QEMU 11.0 simply have the capability to check for the presence or absence of certain features; not all drivers or devices support those features. Therefore, I don't think the display is acting strangely.

they confirmed the settings in the BIOS are correct and that they couldn't reproduce these issues in RHEL 8 (which is on their supported OS list). Since the server is in production, I couldn't test with this OS, hence the post here.

Has RHEL 8 reached the version listed below? If not, I don't think it will even display an error because there's no feature to check for it—or perhaps the feature doesn't even exist. Sorry if I'm wrong?

Code:
VFIO_DEVICE_FEATURE_DMA_BUF should be available in Linux v6.19 and QEMU 11.0.

https://access.redhat.com/articles/red-hat-enterprise-linux-release-dates

RHEL 8.102024-05-222024-05-22 RHBA-2024:31354.18.0-553.el8_10

6.19 > 4.18

I can't find Qemu, but even on PVE, it probably wouldn't even throw an error unless it's Qemu 11.

I think the version update is a good way to narrow down the issue, but I think it’s best to check with NVIDIA to see if they’ve made any changes to make it work. That is, unless you want to end up complaining that the problem isn’t fixed even after repeatedly updating to the new version that doesn’t work.

Since it's just a warning, if NVIDIA doesn't take it seriously, I don't think they'll release an update or anything.
 
Last edited:
Therefore, I don't think the display is acting strangely.
Ok !

Has RHEL 8 reached the version listed below? If not, I don't think it will even display an error because there's no feature to check for it—or perhaps the feature doesn't even exist. Sorry if I'm wrong?
I don't know to be honest since I didn't check/tried or asked what version they used, sorry. But it would be logical that the error would not be reported if the feature is not available on their system.

I think the version update is a good way to narrow down the issue, but I think it’s best to check with NVIDIA to see if they’ve made any changes to make it work. That is, unless you want to end up complaining that the problem isn’t fixed even after repeatedly updating to the new version that doesn’t work.
Yes since I am still doing some tests, I will try to update the driver. If this persists, I will try to ask NVIDIA to check with them if these issues are to be expected or not.

I will update this thread once I've done proxmox/nvidia updates on the server and/or contacted NVIDIA.

Thank you for your time !
 
  • Like
Reactions: uzumo