Here is a link that I've recently came across that may be useful: https://pve.proxmox.com/wiki/Windows_VirtIO_Drivers#0.1.285:_VirtIO_SCSI/VirtIO_Block:_Read_errors_and_performance_issues_with_IO-heavy_Windows_Server_2025_VMs
Blockbridge ...
This section of Wiki may be useful:
https://pve.proxmox.com/wiki/Windows_VirtIO_Drivers#0.1.285:_VirtIO_SCSI/VirtIO_Block:_Read_errors_and_performance_issues_with_IO-heavy_Windows_Server_2025_VMs
Blockbridge : Ultra low latency all-NVME shared...
This thread was started and stopped 2 years ago, I am not sure why it was necromancer'ed :-) I am sure OP has moved on since 2024.
Joined May 31, 2024
Last seen May 31, 2024
Blockbridge : Ultra low latency all-NVME shared storage for Proxmox...
Right, when you know you have a degraded cluster - any little blip is potentially a second and fatal failure. You either let the HA system do its job, which sounds like it performed as expected, or take over recovery by disabling HA...
We do this for all upgrades. We probably should have done this on all nodes straight away when it was in this state, but you live and learn right.
Disarming HA should have probably been our first step.
You may want to investigate the "cluster maintenance" option as a tool for the future. If it acts by suspending automatic HA decisions, giving you time to reboot/recover nodes - that may be something to keep in mind.
ha-manager crm-command...
We nailed it down to a failed dbus, causing a cascading effect. Just a perfect storm of things that are making us question HA a bit. i.e. we could have recovered this had HA not fenced and rebooted everything, but hindsight is wonderful right...
Hi @Daxcor,
NFS is a good choice for your use case. It is used for shared storage by a great number of entities, from home users to multi-billion-dollar corporations.
NFS has been around for 42 years. The good, bad, and ugly have probably been...
When you perform a backup with PBS, PVE/QEMU integrates the backup operation into the VM's I/O path. Data that is being modified while the backup is running has to be coordinated with the backup process. Depending on the storage, network, PBS...
Hi @Jianming , you can now use PVE with iSCSI storage with LVM overlay and snapshots.
However, thin provisioning will need to come from the array at this point in time.
These articles may be helpful to you...
If you have not done so yet, save system journal from each host as far in the past as it allows you. At this point its probably your only lifeline to RCA. I know the hindsight is 20/20, but the best time for data collection was, unfortunately...
I missed your recent update about dbus finding. That is a good lead. Nowhere in this thread (unless I missed it) did you post what version of PVE you are running. If I recall correctly, there were several fixes to dbus handling, in the recent 1-2...
You've already received good pointers on troubleshooting and should continue digging into the logs. That said, the above description is very worrisome...
You had 10-node cluster where 8 of them were failing to start services. We don't have...
Hi @choctaw, welcome to the forum.
Did you save any command-line output from the time you were attempting the upgrade?
Frankly, based on the iterations of troubleshooting you described, but without the corresponding command output, it is...
Here's a slight refinement to my live restore description above, after more time in the source. There are three different sizes at work. I had them blurred together.
The guest read is whatever size the guest asks for. The fetch from PBS is...
Hi @krisshurley , we covered behavior related to your findings in our KB article a while back: https://kb.blockbridge.com/technote/proxmox-qcow-snapshots-on-lvm/#data-consistency-and-information-leak-risks
One approach would be to wipe the...
Hey @abhijithiaaxin, welcome to the forum.
Whether it is a single-DC deployment or a split-DC deployment, you generally want an odd number of nodes in your cluster. But, as others have pointed out, a split-DC installation requires three...
It has been ready for years. A high-friction CLI/UX workflow is not the same thing as a system being "unusable" or "not production-grade." Thousands of enterprise nodes run multipath iSCSI in production every day—it just requires Linux systems...
Thanks for the tracker link @robertlukan . It also affects testing - cloning from the secondary is a common way to practice a recovery without touching replication.
For the 50TB question, @alexskysilk'` math is the way to figure it. Two things...