We run a 6 node PVE 9.2.20 cluster with HA and two dedicated corosync links on separate NICs (Intel E810, link 0 and link 1 each on their own port, nothing else on those interfaces). Default config, so link_mode passive.
Today a single port started flapping on one node. The NIC counts mac_local_faults, the switch side shows zero CRC errors, physical link only went down once, but knet kept marking link 0 down and up every few seconds towards all peers.
[KNET ] link: host: 4 link: 0 is down
[KNET ] host: host: 4 (passive) best link: 1 (pri: 1)
[KNET ] rx: host: 4 link: 0 is up
[KNET ] host: host: 4 (passive) best link: 0 (pri: 1)
pve-ha-lrm: loop took too long (40 seconds)
pvescheduler: cfs-lock 'file-jobs_cfg' error: got lock request timeout
watchdog-mux: client watchdog expired - disable watchdog updates
Questions for people running larger HA clusters
1. Is anyone running link_mode: active in production?Does it actually make a single flapping link harmless?
2. Or do you stay on passive and raise knet_pong_count (or ping timeout/interval) to get some hysteresis?
3. Any other best practice to make sure one bad port or DAC can never cause a node to fence?
Root cause on the hardware side is being handled separately (DAC and switch). I am mainly interested in how to make corosync resilient against this in general.
Thanks!
Today a single port started flapping on one node. The NIC counts mac_local_faults, the switch side shows zero CRC errors, physical link only went down once, but knet kept marking link 0 down and up every few seconds towards all peers.
[KNET ] link: host: 4 link: 0 is down
[KNET ] host: host: 4 (passive) best link: 1 (pri: 1)
[KNET ] rx: host: 4 link: 0 is up
[KNET ] host: host: 4 (passive) best link: 0 (pri: 1)
pve-ha-lrm: loop took too long (40 seconds)
pvescheduler: cfs-lock 'file-jobs_cfg' error: got lock request timeout
watchdog-mux: client watchdog expired - disable watchdog updates
Questions for people running larger HA clusters
1. Is anyone running link_mode: active in production?Does it actually make a single flapping link harmless?
2. Or do you stay on passive and raise knet_pong_count (or ping timeout/interval) to get some hysteresis?
3. Any other best practice to make sure one bad port or DAC can never cause a node to fence?
Root cause on the hardware side is being handled separately (DAC and switch). I am mainly interested in how to make corosync resilient against this in general.
Thanks!