[TTM] Buffer eviction failed

A user in the Ubuntu bugtracker reports that Ubuntu 25.04 (w/ Linux 6.14) does not exhibit the problematic behavior any more, so I guess upgrading your Debian 13 guests to trixie's backports kernel (6.16 at the time of posting) also has a chance of fixing this bug for you.
Thx for this - i checked the source of that kernel-version on and it seems that the "breaking code" is still there.
If I understood Post #22 (https://forum.proxmox.com/threads/ttm-buffer-eviction-failed.152720/post-721493) correctly, every kernel since 6.8.10 has it (still). So it should actually not help here - not sure why it might for the guy from the post you mentioned.
My symptoms are also a little different (Gui-Lock after some hours only) - but I will give that a try as well. I really would prefer having that fixed in an official kernel-package rather than patching and self-compiling etc. ;-)
 
Last edited:
  • Like
Reactions: jtru
Oh man, I just ran into this one yesterday and it cost me hours(!) until I nailed it down. Its a TrueNAS Scale VM running 6.12.15-production+truenas and it was having trouble getting zvols provisioned by democratic-csi. The zfs commands it fires off over ssh sometimes just hung. And then stuff broke..

What a mess. I too had the symptom with the GUI-lock. Switched to VirtIO-GPU for now. There is a thread and a bug report on the TrueNAS side about this.
 
what's the proxmox bugzilla and upstream issue tracker links for this bug/issue ?

who can confirm that this is still an issue?

is there a reliable way to reproduce the issue?
 
Last edited:
i can reproduce the problem with recent debian sid on pve 9.1.5 very easily

just add some qxl / spice display, install debian , boot up gnome desktop, open some console and run "while true;do find /;done" in a loop and watch dmesg with "dmesg -w" in another window.

getting screen freezes and "[TTM] Buffer eviction failed" messages very soon
 
This just started biting me for the first time ever in the past week.

First event was in a newly provisioned virtual machine running LXQt on Debian 13. Okay, I haven't used that software combo on Proxmox before, not too surprising.

Second event, just now, was on a Debian 13 VM on a different physical host running a web app with no GUI installed at all. It had 89 days of uptime.

Timing seems a bit suspicious... I guess I'll see if others start wedging up now, or what. And then when they do, change their video card to VirtIO-GPU.
 
Subject: [PATCH 1/1] drm/qxl: fixes qxl_fence_wait
Date: Fri, 8 Mar 2024 01:08:51 +0000 [thread overview]
Message-ID: <20240308010851.17104-2-dreaming.about.electric.sheep@gmail.com> (raw)
In-Reply-To: <20240308010851.17104-1-dreaming.about.electric.sheep@gmail.com>

Fix OOM scenario by doing multiple notifications to the OOM handler through
a busy wait logic.
Changes from commit 5a838e5d5825 ("drm/qxl: simplify qxl_fence_wait") would
result in a '[TTM] Buffer eviction failed' exception whenever it reached a
timeout.

Fixes: 5a838e5d5825 ("drm/qxl: simplify qxl_fence_wait")
Link: https://lore.kernel.org/regressions/fb0fda6a-3750-4e1b-893f-97a3e402b9af@leemhuis.info
Reported-by: Timo Lindfors <timo.lindfors@iki.fi>
Closes: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1054514
Signed-off-by: Alex Constantino <dreaming.about.electric.sheep@gmail.com>
---
drivers/gpu/drm/qxl/qxl_release.c | 20 ++++++++++++++------
1 file changed, 14 insertions(+), 6 deletions(-)

diff --git a/drivers/gpu/drm/qxl/qxl_release.c b/drivers/gpu/drm/qxl/qxl_release.c
index 368d26da0d6a..51c22e7f9647 100644
--- a/drivers/gpu/drm/qxl/qxl_release.c
+++ b/drivers/gpu/drm/qxl/qxl_release.c
@@ -20,8 +20,6 @@
* CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
*/

-#include <linux/delay.h>
-
#include <trace/events/dma_fence.h>

#include "qxl_drv.h"
@@ -59,14 +57,24 @@ static long qxl_fence_wait(struct dma_fence *fence, bool intr,
{
struct qxl_device *qdev;
unsigned long cur, end = jiffies + timeout;
+ signed long iterations = 1;
+ signed long timeout_fraction = timeout;

qdev = container_of(fence->lock, struct qxl_device, release_lock);

- if (!wait_event_timeout(qdev->release_event,
+ // using HZ as a factor since it is used in ttm_bo_wait_ctx too
+ if (timeout_fraction > HZ) {
+ iterations = timeout_fraction / HZ;
+ timeout_fraction = HZ;
+ }
+ for (int i = 0; i < iterations; i++) {
+ if (wait_event_timeout(
+ qdev->release_event,
(dma_fence_is_signaled(fence) ||
- (qxl_io_notify_oom(qdev), 0)),
- timeout))
- return 0;
+ (qxl_io_notify_oom(qdev), 0)),
+ timeout_fraction))
+ break;
+ }

cur = jiffies;
if (time_after(cur, end))
--
2.39.2

from https://lore.kernel.org/regressions/20240308010851.17104-2-dreaming.about.electric.sheep@gmail.com/
 
  • Like
Reactions: RolandK
As this was still bothering me and all the "workarounds" were not really working well for me, i started to research again and recently found this:

https://lists.freedesktop.org/archives/dri-devel/2026-July/584771.html

(This seems to be a more modern fix than the one "floating around" before as far as i understand it.
If it ever makes it into official kernel - who knows. As far as i know/understood the kernel devs decided NOT to fix that officially, because QXL-drivers are in maintenance-mode only and the fixes that exist were some strange "hacks" that don´t follow modern linux-kernel principles etc.)

Even if beeing a linux-veteran for decades :cool:, I am not a developer or any expert in fiddeling with kernel-sources etc.
--> So I used some "AI-Power" (Credits go out to Gemini ;)) to (let) create a script that automatically patches the QXL-driver with that and integrates with DKMS etc, so that it should survive minor kernel-updates directly with DKMS and even work with new major releases (unless other bigger changes break it).
If DKMS fails after kernel-updates - re-running the script should help (fetches new QXL-sources etc...)

--> For me this seems to work: I can use SPICE as Display again (not the horribly lagging virtio-gpu) and there are no freezes (not during boot and not after some time with "[TTM] Buffer eviction failed" - at least so far.

DISCLAIMER: This is definately not perfectly optimized and it also took me some iterations with AI to make it work porperly, but it seems to do the job now and is quite "lightweight" in my opinion.

The script is made on/for current Debian 13/Trixie (and MIGHT work on other Debian-based distros - UNTESTED!)

(... and some comments in the script are in German, but i was to lazy to translate)

USE AT YOUR OWN RISK and MAKE BACKUPS :cool:

But i thought I´ll share it for everyone to try and hopefully enjoy

USAGE:
  1. Save script on the VM
    1. optionally/recommended: change extenstion to .sh (upload was not allowed with .sh extension)
  2. make it executable (chmod +x <FILE>)
  3. run it
  4. shutdown VM
  5. change display (back) to SPICE (if you were using the virtio-gpu workaround as i was)
  6. boot VM
  7. ENJOY
 

Attachments

  • Like
Reactions: leesteken
the patch got submitted upstream, and already got some feedback - provided the author continues to iterate on it, it should eventually find its way into our kernel builds ;)
 
  • Like
Reactions: ucholak
Thx for the update about the patch being discussed any maybe officially getting into the kernel.
I hope i won´t be "your" (PVE) kernels only, but mainline/debian as well ;-)