Why healthy SSD's LVM corrupts after taking it out of host and inserting back?

niall.keller

New Member
Aug 20, 2026
22
1
3
What may be the nature of the failure (beside unnoticed events I could not notice in the process)? The SSD is healthy.

1. Take the only SSD (LVM thin) out of proxmox host
2. Place it into another host
3. Boot new host into some Live USB environment, like Clonezilla
4. Clone the disk (deny fschk on source disk) into new/target disk (!the topic is about old/source disk, not this new one)
5. Try to put new disk into proxmox host, observe failures and give up trying to fix (root boots, but ext4 is not found on some partitions, complains about superblocks, etc)
6. Move original source disk into proxmox host, observe same failures

Is it the taking SSD out of host and re-inserting it back by itself changes/corrupts LVM thing?
Or is it inserting it into another host as secondary disk affects it?
Or is it clonezilla "silently" changing anything?
 
Last edited:
lvm-thin is an abstraction, and is not meant to be cloned. For ext4, if you changed sector sizes for example, between 512 and 4k that could maybe cause some of the issues you're seeing bc the underlying topology has changed.

https://search.brave.com/search?q=d...ersation=09a2da554a07044f36502e818e5c6cea7fa4

Hope you have backups, bc you're looking at re-installing at this point. Implement PBS soonest if you haven't already, preferably on separate hardware.

https://github.com/kneutron/ansitest/blob/master/proxmox/bkpcrit-proxmox.sh

https://github.com/kneutron/ansitest/blob/master/proxmox/proxmox-BULK-RESTORE-VMS--PARALLEL.sh
 
  • Like
Reactions: Johannes S
lvm-thin is an abstraction, and is not meant to be cloned. For ext4, if you changed sector sizes for example, between 512 and 4k that could maybe cause some of the issues you're seeing bc the underlying topology has changed.

https://search.brave.com/search?q=d...ersation=09a2da554a07044f36502e818e5c6cea7fa4

Hope you have backups, bc you're looking at re-installing at this point. Implement PBS soonest if you haven't already, preferably on separate hardware.

https://github.com/kneutron/ansitest/blob/master/proxmox/bkpcrit-proxmox.sh

https://github.com/kneutron/ansitest/blob/master/proxmox/proxmox-BULK-RESTORE-VMS--PARALLEL.sh
But the question is not about "new" drive, it's about the source disk.
 
Hello, it is very strange. It shall not do that.
Is it possible that the source disk has problem you didn't saw before ?
Make a SMART test, look at %use, ...
When you stopped the initial computer, was it a clean shutdown ?
 
Hello, it is very strange. It shall not do that.
Is it possible that the source disk has problem you didn't saw before ?
Make a SMART test, look at %use, ...
When you stopped the initial computer, was it a clean shutdown ?
It is a relatively new disk, proxmox webgui showed no issues. Will retest it anyway soon.
Shutdown has clean with 5-min cooldown.
 
  • Like
Reactions: ghusson
There is at least one thread mentioning Clonezilla potentially damaging source structure, but not explicitly mentioning moving the drive.

I'm wondering if it's Clonezilla destroying source disk structure. Or if in my case moving the drive could have caused the damage.

I'm fine CZ not supporting LVM thin, but destroying the source without explicit notification is too much. I have confirmed "no fsck" and expected no intervention from CZ while cloning.
 
Last edited:
Ok I have an idea of what could be arriving here.
Clonezilla classically reproduce disk structure and do things inside partitions.
I didn't paid a close look but what could be happening here :
- clonezilla reproduce LVM structure with same UUID
- clonezilla begins to copy data but LVM autodiscovery is lost because LVM UUIDs are present on both disks
- disks IOs are done sometimes on destination disk and sometime on source disk
- source disk is corrupted.
Try do clone the disk with a simple raw dd (dd /dev/disk_in /dev/disk_out bs=32M) ?
 
Last edited:
Ok I have an idea of what could be arriving here.
Clonezilla classically reproduce disk structure and do things inside partitions.
I didn't paid a close look but what could be happening here :
- clonezilla reproduce LVM structure with same UUID
- clonezilla begins to copy data but LVM autodiscovery is lost because LVM UUIDs are present on both disks
- disks IOs are done sometimes on destination disk and sometime on source disk
- source disk is corrupted.
Try do clone the disk with a simple raw dd (dd /dev/disk_in /dev/disk_out bs=32M) ?
What Claude says :

Short answer: pulling the SSD out and putting it back, or attaching it as a secondary disk on another machine, does not change anything by itself. A disk only gets altered when a system writes to it. The main suspect is therefore the Clonezilla step, and two known mechanisms match the symptoms. Which one actually happened can't be confirmed without examining the disk.

1. Clonezilla and LVM thin: support was missing for a long time

For years, the Clonezilla maintainer told users, including Proxmox users, that for technical reasons Clonezilla did not support LVM thin provisioning. Reported symptoms look similar. One user found that only the /boot partition outside LVM and the virtual hard drive inside the thin pool were backed up, with Clonezilla falling back to dd mode. Another ticket reports that Clonezilla failed even when forced into dd mode (-q1): it still tried to parse LVM and failed, even with LVM filters restricting device access.
Clonezilla / Discussion / Open Discussion: How to restore disk with thin-lvm volumes? +2

Support then changed over time. An intermediate release added a check that detects LVM thin provisioning and quits the program. Later, release 3.3.2-31 added support for LVM thin provisioning. With any earlier version, behavior on a Proxmox thin pool was unreliable. An option added earlier also matters: "--force" for vgcfgrestore, to force a metadata restore even with thin pool LVs. That operation writes LVM metadata.
Clonezilla / News +2

2. The real "mix-up": duplicate PVs

This hypothesis would explain why the source disk is damaged too. During and after a disk-to-disk clone, both disks carry the same PV UUID and the same pve VG. But LVM identifies physical volumes by UUID, not by device path. When both are visible, LVM picks one, and that choice can change. On the Proxmox forum, a user found after cloning that /dev/pve/root had become nvme1n1p3 (the clone) rather than the original nvme0n1p3, and not consistently.
netdata
proxmox

If Clonezilla rebuilt the LVM layout on the target while the same-UUID source was still visible, some of those writes could have landed on the source. The rebuild steps involved are pvcreate, vgcfgrestore, activation, then writing LVs with partclone or dd. Three effects are possible:

the source's LVM metadata gets rewritten;
a write to an LV that was mapped onto the wrong PV;
activating and deactivating the thin pool modifies its metadata (tmeta) on the source.

The symptoms fit this scenario: root boots, but ext4 is missing elsewhere and superblocks are invalid. On Proxmox, root is a regular LV, while VM and container disks live in the data thin pool. That pool is what breaks first if its metadata and data diverge.

A third possible cause is more mundane: source and target were swapped when the disks were selected. With two SSDs of the same model, that's an easy mistake.

How to verify it factually

Do this with only the source disk attached, from a live environment, and without repairing anything.

First take a raw image of the source (dd or ddrescue to another device). lvconvert --repair and fsck are destructive if you get it wrong.
Run pvck --dump metadata_all /dev/sdX3. The LVM metadata area keeps a circular history of versions. Each version has a seqno, a creation_host and a timestamp. A recent version written by the Clonezilla host proves the source was modified.
Compare against /etc/lvm/archive/ and /etc/lvm/backup/ on the Proxmox root, which still boots. They show the last state Proxmox itself wrote.
Run lvs -a pve to check the state of the pool, [data_tmeta] and [data_tdata]. Then run thin_check on the tmeta, after exporting it or activating it read-only.

If the original LVM metadata is still intact in the history, restoring it with vgcfgrestore --force to that seqno, working on the image, is often feasible. It's not guaranteed if thin pool data was overwritten.

Information needed to go further: which Clonezilla version was used, and in which mode (device-device beginner or expert, or device-image).