I have a cluster of three hosts with HA enabled in a homelab environment that typically stores disks on an NFS share from a storage appliance that recently had a hardware failure.
A few weeks ago I temporarily moved the root disk of one of my CTs to local-lvm of one of the node (pve0) as part of a minimum runtime while I replaced the underlying storage appliance that usually hosts my VM and CT disks. Tonight I installed the replacement storage appliance that just arrived and started up my typical VMs and CTs without any real issues using HA (which had been set to disabled while I waited for the replacement).
The underlying storage has had some IO pressure issues that I am still working to diagnose but when everything first starts up, for a few hours, things can be slow. This is not a common occurrence so I don't worry about it too much. However tonight it caused me an issue since HA decided to try and migrate the CT that was on local-lvm (since I hadn't moved it back to the network storage yet, I was waiting for the environment to stabilize first), like to balance resources. This migration failed during the disk transfer l(ikely due to the network traffic and overall activity on the nodes) and an error was logged in the task history. Typically this isn't an issue since PVE will gracefully fallback to the stopped CT image on the originating host and put the container in an error state that I need to clear from the HA manager. Tonight though it seems that it had actually moved the storage to the new host, then failed for some reason and deleted the disk image from the new host as a cleanup action, after it had already removed the image from the originating host. This leaves me in a state where the container config references a disk on local-lvm that doesn't exist anywhere. I don't have a backup of the container, which is a docker host for a few critical services that I really don't want to recreate (I know, I should have a backup, but that ship has sailed at this point and I have learned that lesson) so hoping there is a way to recover from this.
Is there any way to locate the disk on the original host in this case? or do I need to resign myself to recreating the container and the services it hosted?
The migration log for all of the HA migrations of this guest look like:
I don't see the original migration task in the task history so its possible the original issue is something else entirely. the 'NFS_Share:109/vm-109-disk-o.raw' is the original network share copy of the disk that I moved three weeks ago and then left since I planned to move back to it once everything was stable. Perhaps I can point the CT config at that disk as the root instead since it should be a valid copy that just hasn't been running for a few weeks (but all the underlying data should be just as it was three weeks ago right? as if I turned the container off and left it sit for a while?)
A few weeks ago I temporarily moved the root disk of one of my CTs to local-lvm of one of the node (pve0) as part of a minimum runtime while I replaced the underlying storage appliance that usually hosts my VM and CT disks. Tonight I installed the replacement storage appliance that just arrived and started up my typical VMs and CTs without any real issues using HA (which had been set to disabled while I waited for the replacement).
The underlying storage has had some IO pressure issues that I am still working to diagnose but when everything first starts up, for a few hours, things can be slow. This is not a common occurrence so I don't worry about it too much. However tonight it caused me an issue since HA decided to try and migrate the CT that was on local-lvm (since I hadn't moved it back to the network storage yet, I was waiting for the environment to stabilize first), like to balance resources. This migration failed during the disk transfer l(ikely due to the network traffic and overall activity on the nodes) and an error was logged in the task history. Typically this isn't an issue since PVE will gracefully fallback to the stopped CT image on the originating host and put the container in an error state that I need to clear from the HA manager. Tonight though it seems that it had actually moved the storage to the new host, then failed for some reason and deleted the disk image from the new host as a cleanup action, after it had already removed the image from the originating host. This leaves me in a state where the container config references a disk on local-lvm that doesn't exist anywhere. I don't have a backup of the container, which is a docker host for a few critical services that I really don't want to recreate (I know, I should have a backup, but that ship has sailed at this point and I have learned that lesson) so hoping there is a way to recover from this.
Is there any way to locate the disk on the original host in this case? or do I need to resign myself to recreating the container and the services it hosted?
The migration log for all of the HA migrations of this guest look like:
Code:
ask started by HA resource agent
2026-09-28 19:51:47 starting migration of CT 109 to node 'pve2' (10.x.x.x)
2026-09-28 19:51:47 volume 'NFS_Share:109/vm-109-disk-0.raw' is on shared storage 'NFS_Share'
2026-09-28 19:51:47 found local volume 'local-lvm:vm-109-disk-0' (in current VM config)
2026-09-28 19:51:47 using a bandwidth limit of 209715200 bytes per second for transferring 'local-lvm:vm-109-disk-0'
2026-09-28 19:51:47 ERROR: storage migration for 'local-lvm:vm-109-disk-0' to storage 'local-lvm' failed - no such logical volume pve/vm-109-disk-0
2026-09-28 19:51:47 aborting phase 1 - cleanup resources
2026-09-28 19:51:47 ERROR: found stale volume copy 'local-lvm:vm-109-disk-0' on node 'pve2'
2026-09-28 19:51:47 start final cleanup
2026-09-28 19:51:47 ERROR: migration aborted (duration 00:00:00): storage migration for 'local-lvm:vm-109-disk-0' to storage 'local-lvm' failed - no such logical volume pve/vm-109-disk-0
TASK ERROR: migration aborted
I don't see the original migration task in the task history so its possible the original issue is something else entirely. the 'NFS_Share:109/vm-109-disk-o.raw' is the original network share copy of the disk that I moved three weeks ago and then left since I planned to move back to it once everything was stable. Perhaps I can point the CT config at that disk as the root instead since it should be a valid copy that just hasn't been running for a few weeks (but all the underlying data should be just as it was three weeks ago right? as if I turned the container off and left it sit for a while?)