So,
I do run a 3-Cluster in my homelab, using ZFS replikation for HA. Node names are bert, ernie and grobi.
Everything is running fine most of the time, but now and then for a period of 15-30min all replikation jobs fail like this.
How can i identify the real cause?
I do run a 3-Cluster in my homelab, using ZFS replikation for HA. Node names are bert, ernie and grobi.
Everything is running fine most of the time, but now and then for a period of 15-30min all replikation jobs fail like this.
There is - at the same time a „spike“ in CPU usage (up to 8-9%) but aside from that nothing unusual.Replication job '104-0' with target 'bert' and schedule '0/15' failed!
Last successful sync: 2026-10-11 01:45:01
Next sync try: ERROR
Failure count: 1
Error:
command 'zfs snapshot local-zfs/subvol-104-disk-0@__replicate_104-0_1791676801__' failed: got timeout
How can i identify the real cause?