I recently upgraded two machines with big ZFS pools: one to PVE9.2.4, the other to PBS4.2.2, so both machines have Kernel 7.0 and ZFS 2.4.3 now.
One pool is raid-z1 (12x18TB), had one drive faulted, the other raid-z2 (12x18TB), with two drives faulted after reboot. zpool-online didn't help, I had to zpool-labelclear and zpool-replace the disks, with the consequence of several days of resilver.
With previous system versions, I can't remember having such problems, so I suspect that newer kernels/zfs might have different timings, with disks not ready in-time from the SAS2008/3008 adapters. storcli will show all disk as good, smartctl didn't show any problems on the faulted disks.
Are there any best-practices to avoid such reboot trouble?
Regards,
Andreas
One pool is raid-z1 (12x18TB), had one drive faulted, the other raid-z2 (12x18TB), with two drives faulted after reboot. zpool-online didn't help, I had to zpool-labelclear and zpool-replace the disks, with the consequence of several days of resilver.
With previous system versions, I can't remember having such problems, so I suspect that newer kernels/zfs might have different timings, with disks not ready in-time from the SAS2008/3008 adapters. storcli will show all disk as good, smartctl didn't show any problems on the faulted disks.
Are there any best-practices to avoid such reboot trouble?
Regards,
Andreas