Windows Server 2025 + FSLogix Fileserver: sporadic vioscsi Event ID 129 ("Reset to device \Device\RaidPortX") causing inaccessible volume

bzr

New Member
Jul 7, 2025
18
0
1
Hello,

we have observed a very rare but serious issue on two different Windows Server 2025 VMs running on Proxmox VE with Ceph RBD storage.
Both VMs are FSLogix profile container file servers with a nearly identical workload.

Windows log shows:
Source: vioscsi
Event ID: 129
Reset to device, \Device\RaidPort5, was issued.

After the first event appears, additional Event 129 warnings continue to occur.

The affected volume:

  • remains visible in Windows Explorer
  • drive letter remains present
  • shares remain visible
  • volume becomes inaccessible
  • all access to the volume hangs or fails
A reboot of the VM immediately restores normal functionality.

No filesystem repair is required afterwards.

We did not observe:
  • NTFS errors
  • ReFS errors
  • disk corruption
  • storage errors inside Windows
The issue affects only the volume hosting the FSLogix VHDX containers.

Environment:
Proxmox VE: 9.2.10
Storage: Ceph RBD
Ceph: 19.2.3
Windows Server 2025

VirtIO SCSI
Cache = Write Back
Discard = On
IO Thread = On

originally we ran on virtio-win 0.1.285. After finding the note in the Proxmox Wiki regarding Windows Server 2025 and IO-heavy workloads, we downgraded to: virtio-win 0.1.271


Has anybody seen similar behaviour with:
Windows Server 2025
FSLogix profile container storage
VirtIO SCSI
Ceph RBD

where the "Event ID 129 (vioscsi) Reset to device \Device\RaidPortX" eventually leads to a volume becoming inaccessible although the disk remains visible in Windows?

Any known relation to:

  • VirtIO 0.1.285
  • VirtIO 0.1.271
  • Windows Server 2025
  • FSLogix workloads
  • VirtIO-SCSI controller type
would be appreciated.

Thanks!
 
We found GitHub issue #756. Our symptoms are very similar, especially:

  • Event ID 129
  • vioscsi/viostor
  • affected volume becomes unresponsive
  • VM reboot restores operation
  • Ceph-backed storage
However, the github issue is marked as resolved and we are seeing this on Windows Server 2025 and virtio-win 0.1.285/0.1.271.
 
On both FSLogix VMs, line up the Event 129 timestamps with Ceph latency on the host for the same minutes. Run ceph osd perf, check the RBD image latency in the Proxmox graphs, and look for slow ops warnings in the Ceph log. If the resets match latency spikes, fix the Ceph side first. If they don't, install 0.1.302 on one of the two VMs only and keep the other on 0.1.271, so you can compare them under the same workload. Also test one disk with cache set to none instead of Write Back.

Drafted with AI, reviewed by me.
 
Last edited:
On both FSLogix VMs, line up the Event 129 timestamps with Ceph latency on the host for the same minutes. Run ceph osd perf, check the RBD image latency in the Proxmox graphs, and look for slow ops warnings in the Ceph log. If the resets match latency spikes, fix the Ceph side first. If they don't, install 0.1.302 on one of the two VMs only and keep the other on 0.1.271, so you can compare them under the same workload. Also test one disk with cache set to none instead of Write Back.
i cant see any ceph related performance issues. i have multiple VMs running on the same Ceph Cluster. Even the SQL Server runs flawlessly on virtio 0.1.285, but it is still on Windows 2016...

i set Cache to none on both machines and updated one to 0.1.302... all i can do is wait for now because the error occurs weeks or months appart...
 
i cant see any ceph related performance issues. i have multiple VMs running on the same Ceph Cluster. Even the SQL Server runs flawlessly on virtio 0.1.285, but it is still on Windows 2016...

i set Cache to none on both machines and updated one to 0.1.302... all i can do is wait for now because the error occurs weeks or months appart...
Do the Event 129 timestamps line up with a scheduled vzdump backup or snapshot of those two VMs? Sometimes I feel a little load off when I say "all I can do is wait" because I have so much stuff to do that kicking the can down the road is a relief.

Drafted with AI, reviewed by me.
 
Last edited:
Do the Event 129 timestamps line up with a scheduled vzdump backup or snapshot of those two VMs? Sometimes I feel a little load off when I say "all I can do is wait" because I have so much stuff to do that kicking the can down the road is a relief.
no they dont line up. Last time the backup finished at 21:45 and first 129 event was thrown at 23:26.
i hope that i can kick the can quite some time and the error never occurs again... that would be my relief
 
no they dont line up. Last time the backup finished at 21:45 and first 129 event was thrown at 23:26.
i hope that i can kick the can quite some time and the error never occurs again... that would be my relief
Any defrag event 258's from the storage optimizer near the 129 event?

Drafted with AI, reviewed by me.
 
Last edited:
Any defrag event 258's from the storage optimizer near the 129 event?

Drafted with AI, reviewed by me.
on one Server there is a Defrag 264 Event 1hour before the first vioscsi 129 event. But for another Volume. On the other there are no Defrag events.
 
Have the 129 resets on the two servers ever happened within 5 minutes of each other? It clearly isn't retrim. It is most likely the shared Ceph/host path or the per-VM vioscsi driver.

Drafted with AI, reviewed by me.
 
Have the 129 resets on the two servers ever happened within 5 minutes of each other? It clearly isn't retrim. It is most likely the shared Ceph/host path or the per-VM vioscsi driver.

Drafted with AI, reviewed by me.
no, not yet. the issue between the two vms occured days to weeks apart...
 
On both VMs I would enable the StorPort operational log:
wevtutil sl Microsoft-Windows-Storage-Storport/Operational /e:true
It logs IO latency and timeout detail before a reset, which shows whether requests slowed down first or just stopped completing.

When you get the next 129 check the host (qm monitor → info block, ceph health detail, stuck ops on that RBD image) and take a guest memory dump if you can before rebooting. The memory dump should be enough for maintainers to diagnose the virtio-win issue.

Drafted with AI, reviewed by me.
 
we got the error on another server. For the first time on that one. Similar Configuration:
-Windows 2025
-2 Volumes: C: 100GB and M: 4TB

the vioscsi 129 events have been thrown yesterday for the first time on that one. I updated virtio to .302. In the night the vioscsi 129 events came up again... on .302. I made a memory dump of the vm and rebooted it... the dump ist 4GB in size...
 
im pretty sure i did, yes. At least the driver Version under device Manager was displayed correctly
 
It isn't something the .302 driver fixes, then. With a second VM now affected, it is most likely the shared Ceph/host path.

Did `ceph health detail` or the OSD logs show slow ops at the time of last night's 129 events?

Drafted with AI.
 
too high latency would be displayed in ceph cluster health i suppose?
in total there are 3 VMs affected by the same problem until now.
all 3 VMs are Windows Server 2025 with now different versions of virtio-drivers.
All VMs are on differet Nodes. No other VMs on the same Nodes have problems. There are no health issues on th ceph cluster.
there are no Events in the OSD log except from something like this:
2026-10-06T17:36:28.685553+0200 mgr.bzr-gn-pve01 (mgr.85757044) 1477091 : cluster [DBG] pgmap v1477821: 673 pgs: 673 active+clean; 21 TiB data, 60 TiB used, 86 TiB / 146 TiB avail; 4.7 MiB/s rd, 13 MiB/s wr, 1.25k op/s
2026-10-06T17:36:30.686662+0200 mgr.bzr-gn-pve01 (mgr.85757044) 1477092 : cluster [DBG] pgmap v1477822: 673 pgs: 673 active+clean; 21 TiB data, 60 TiB used, 86 TiB / 146 TiB avail; 5.3 MiB/s rd, 18 MiB/s wr, 1.49k op/s
2026-10-06T17:36:31.732356+0200 osd.11 (osd.11) 1261 : cluster [DBG] 2.16e scrub starts
2026-10-06T17:36:32.687631+0200 mgr.bzr-gn-pve01 (mgr.85757044) 1477093 : cluster [DBG] pgmap v1477823: 673 pgs: 673 active+clean; 21 TiB data, 60 TiB used, 86 TiB / 146 TiB avail; 4.3 MiB/s rd, 18 MiB/s wr, 1.41k op/s
2026-10-06T17:36:33.522345+0200 osd.11 (osd.11) 1262 : cluster [DBG] 2.16e scrub ok
2026-10-06T17:36:34.688153+0200 mgr.bzr-gn-pve01 (mgr.85757044) 1477094 : cluster [DBG] pgmap v1477824: 673 pgs: 673 active+clean; 21 TiB data, 60 TiB used, 86 TiB / 146 TiB avail; 3.2 MiB/s rd, 17 MiB/s wr, 1.24k op/s
2026-10-06T17:36:36.689628+0200 mgr.bzr-gn-pve01 (mgr.85757044) 1477095 : cluster [DBG] pgmap v1477825: 673 pgs: 673 active+clean; 21 TiB data, 60 TiB used, 86 TiB / 146 TiB avail; 4.8 MiB/s rd, 23 MiB/s wr, 1.70k op/s

ceph health detail
HEALTH_OK

we have >85VMs on the Cluster, 17VMs on this particular Node. All Nodes share the same CEPH Cluster. Wouldn't there be more VMs affected when it was a ceph problem?