Hi,
I am troubleshooting an intermittent Corosync/KNET connectivity problem in a new 3-node Proxmox VE cluster.
The issue is reproducible when starting a VM on node 1. At that point, nodes 2 and 3 become unavailable in the Proxmox GUI, and pvecm status eventually shows only node 1 as a cluster member.
What makes the problem interesting is that the underlying Linux bond remains UP, LACP remains established, and there are no physical link errors or carrier loss reported by the host.
I would appreciate any advice from the Proxmox/Corosync/KNET community about what could cause this behavior or what additional diagnostics would be useful.
Environment
Three-node cluster:- tfe-proxmox-1
- tfe-proxmox-2
- tfe-proxmox-3
Proxmox VE 9.2.0
pve-manager 9.2.21
Linux kernel 7.0.14-22-pve
corosync 3.1.10-pve3
libknet1t64 1.35-pve2
ifupdown2 3.3.0-1+pmx12
Cluster transport:
Transport: knet
Secure auth: on
The cluster currently uses a single Corosync link.
Network topology
Each Proxmox node has two 25 GbE interfaces in a Linux LACP bond:nic4
nic6
|
+---- bond0 ---- vmbr0
The bond is configured as:
bond-mode: 802.3ad
bond-xmit-hash: layer2+3
LACP: active
LACP rate: fast
MTU: 1500
miimon: 100
The Corosync network is VLAN 720:
tfe-proxmox-1 192.168.20.1/29
tfe-proxmox-2 192.168.20.2/29
tfe-proxmox-3 192.168.20.3/29
The relevant configuration on node 1 is:
auto bond0
iface bond0 inet manual
bond-slaves nic4 nic6
bond-miimon 100
bond-mode 802.3ad
bond-xmit-hash-policy layer2+3
mtu 1500
bond-lacp-rate fast
auto bond0.720
iface bond0.720 inet static
address 192.168.20.1/29
mtu 1500
auto vmbr0
iface vmbr0 inet manual
bridge-ports bond0
bridge-stp off
bridge-fd 0
bridge-vlan-aware yes
bridge-vids 2-4094
mtu 1500
The same VLAN 720 is used for the Corosync addresses on nodes 2 and 3.
The second Corosync link was intentionally removed during troubleshooting. The current configuration therefore has only link 0.
Switch infrastructure
The Proxmox nodes are connected to a pair of Extreme Networks 7520-48Y switches.The switches are running:
Switch Engine 33.7.1.6
image: 33.7.1.6-patch1-26
The two switches use MLAG.
The Proxmox nodes are connected as follows:
ExtremeTF1
/ \
port 7 port 7
\ /
\ /
tfe-proxmox-1
ExtremeTF2
The same arrangement exists for node 2 on port 9 and node 3 on port 11.
The corresponding switch-side LAGs are configured as:
enable sharing 7 grouping 7 algorithm address-based L3_L4 lacp
enable sharing 9 grouping 9 algorithm address-based L3_L4 lacp
enable sharing 11 grouping 11 algorithm address-based L3_L4 lacp
The MLAG peer-link is:
enable sharing 49 grouping 49-50 algorithm address-based L2 lacp
The MLAG peer itself remains healthy:
- MLAG peer: Up
- Checkpoint: Up
- Peer connection failures: 0
- Checkpoint errors: 0
For example, node 1 has both physical links in the same LACP aggregator, with:
- both links UP
- 25 Gb/full duplex
- no link failures
- no actor churn
- no partner churn
- same aggregator ID
- partner MAC corresponding to the MLAG pair
Problem
The problem is intermittent but reproducible when starting a VM on tfe-proxmox-1.When the problem occurs, nodes 2 and 3 become red in the Proxmox GUI.
This is not only a GUI problem. Corosync actually loses cluster membership.
For example, during one failure:
root@tfe-proxmox-1:~# corosync-cfgtool -s
Local node ID 1, transport knet
LINK ID 0 udp
addr = 192.168.20.1
status:
nodeid: 1: localhost
nodeid: 2: disconnected
nodeid: 3: disconnected
And:
root@tfe-proxmox-1:~# pvecm status
Cluster information
-------------------
Name: Cluster-TFE-01
Config Version: 4
Transport: knet
Secure auth: on
Quorum information
------------------
Date: Thu Oct 8 10:18:28 2026
Quorum provider: corosync_votequorum
Nodes: 1
Node ID: 0x00000001
Quorate: No
Votequorum information
----------------------
Expected votes: 3
Highest expected: 3
Total votes: 1
Quorum: 2 Activity blocked
Membership information
----------------------
Nodeid Votes Name
0x00000001 1 192.168.20.1 (local)
Corosync/KNET log during the failure
The most interesting part is the sequence of events.First KNET loses node 3:
Oct 08 10:17:44 tfe-proxmox-1 corosync[2772]: [KNET ] link: host: 3 link: 0 is down
Oct 08 10:17:44 tfe-proxmox-1 corosync[2772]: [KNET ] host: host: 3 (passive) best link: 0 (pri: 1)
Oct 08 10:17:44 tfe-proxmox-1 corosync[2772]: [KNET ] host: host: 3 has no active links
Then Corosync loses the token:
Oct 08 10:17:45 tfe-proxmox-1 corosync[2772]: [TOTEM ] Token has not been received in 2343 ms
Oct 08 10:17:45 tfe-proxmox-1 corosync[2772]: [TOTEM ] A processor failed, forming new configuration