Single flapping corosync link caused a node to fence itself. Is knet link_mode active the right fix?

luka0x73

New Member
Sep 29, 2026
2
1
3
We run a 6 node PVE 9.2.20 cluster with HA and two dedicated corosync links on separate NICs (Intel E810, link 0 and link 1 each on their own port, nothing else on those interfaces). Default config, so link_mode passive.
Today a single port started flapping on one node. The NIC counts mac_local_faults, the switch side shows zero CRC errors, physical link only went down once, but knet kept marking link 0 down and up every few seconds towards all peers.

[KNET ] link: host: 4 link: 0 is down
[KNET ] host: host: 4 (passive) best link: 1 (pri: 1)
[KNET ] rx: host: 4 link: 0 is up
[KNET ] host: host: 4 (passive) best link: 0 (pri: 1)
pve-ha-lrm: loop took too long (40 seconds)
pvescheduler: cfs-lock 'file-jobs_cfg' error: got lock request timeout
watchdog-mux: client watchdog expired - disable watchdog updates

Questions for people running larger HA clusters

1. ⁠Is anyone running link_mode: active in production?Does it actually make a single flapping link harmless?
2. ⁠Or do you stay on passive and raise knet_pong_count (or ping timeout/interval) to get some hysteresis?
3. ⁠Any other best practice to make sure one bad port or DAC can never cause a node to fence?

Root cause on the hardware side is being handled separately (DAC and switch). I am mainly interested in how to make corosync resilient against this in general.
Thanks!
 
link_mode: active means all corosync traffic- checksums, encrypt, decrypt, etc- happens twice. It is not insignificant. You CAN utilize that to protect against the symptom you describe, but its a rather heavy solution for a transient problem that you are correcting.

What might work well enough without the overhead IF you have no trust in your network stabiility is to increase knet_pong_count as long as you also specify a knet_ping_interval that will not exceed the corosync token timeout (default 3000ms) accounting for ping count. If you can see in the logs what the total length of time a flap takes to return, set the values
token: > value of (knet_pong_count x knet_ping_interval) > flap interval.

Obviously if this value starts climbing in the a high enough value I'd just disable the flapping nic.