Replication Jobs failing with timeout now and then

maxim.webster

Active Member
Nov 12, 2024
292
141
43
Germany
So,

I do run a 3-Cluster in my homelab, using ZFS replikation for HA. Node names are bert, ernie and grobi.

Everything is running fine most of the time, but now and then for a period of 15-30min all replikation jobs fail like this.

Replication job '104-0' with target 'bert' and schedule '0/15' failed!

Last successful sync: 2026-10-11 01:45:01
Next sync try: ERROR
Failure count: 1

Error:
command 'zfs snapshot local-zfs/subvol-104-disk-0@__replicate_104-0_1791676801__' failed: got timeout
There is - at the same time a „spike“ in CPU usage (up to 8-9%) but aside from that nothing unusual.

How can i identify the real cause?
 
I’d like to check for any storage-related delays or errors. Could you please share the output of the following commands?

Code:
zpool status -v
zpool events -v