Zweiter Knoten nicht administrierbar

Oct 2, 2026
4
1
3
Hallo beisammen.

bei einem Proxmox zwei Knoten mit Quorum Cluster haben ich das Problem, dass ich nach einer Zeit von der URL des ersten Knoten den zweiten Knoten nicht mehr administrieren kann. Es sieht alles normal aus, VMs sind allen online, allerdings bekomme ich ein Timeout, wenn ich VMs oder den Knoten anklicke.

Wenn ich allerdings auf beiden Knoten das Cluster mit folgendem Befehl neu starte, dann geht es wieder eine Zeit lang:
systemctl restart pve-cluster

Wie kann ich herausfinden woran das liegt und es permanent lösen?

Vielen Dank im Voraus.

Grüße
Stefan
 
Direkt vor dem Neustart habe ich auf dem ersten Knoten folgendes gefunden:
Oct 01 16:51:24 proxmox1 pmxcfs[1075619]: [confdb] crit: cmap_dispatch failed: 2
Oct 01 16:51:24 proxmox1 pmxcfs[1075619]: [quorum] crit: quorum_dispatch failed: CS_ERR_LIBRARY
Oct 01 16:51:24 proxmox1 pmxcfs[1075619]: [status] notice: node lost quorum
Oct 01 16:51:24 proxmox1 pmxcfs[1075619]: [dcdb] crit: cpg_dispatch failed: CS_ERR_LIBRARY
Oct 01 16:51:24 proxmox1 pmxcfs[1075619]: [dcdb] crit: cpg_leave failed: CS_ERR_LIBRARY
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [status] crit: cpg_dispatch failed: CS_ERR_LIBRARY
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [status] crit: cpg_leave failed: CS_ERR_LIBRARY
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [quorum] crit: quorum_initialize failed: CS_ERR_LIBRARY (failed to connect to corosync)
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [quorum] crit: can't initialize service
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [confdb] crit: cmap_initialize failed: CS_ERR_LIBRARY (failed to connect to corosync)
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [confdb] crit: can't initialize service
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [dcdb] notice: start cluster connection
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [dcdb] crit: cpg_initialize failed: CS_ERR_LIBRARY (failed to connect to corosync)
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [dcdb] crit: can't initialize service
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [status] notice: start cluster connection
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [status] crit: cpg_initialize failed: CS_ERR_LIBRARY (failed to connect to corosync)
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [status] crit: can't initialize service
Oct 01 16:51:25 proxmox1 systemd[1]: Stopping pve-cluster.service - The Proxmox VE cluster filesystem...
Oct 01 16:51:25 proxmox1 pmxcfs[1075619]: [main] notice: teardown filesystem
Oct 01 16:51:26 proxmox1 pmxcfs[1075619]: [quorum] crit: quorum_finalize failed: CS_ERR_BAD_HANDLE
Oct 01 16:51:26 proxmox1 pmxcfs[1075619]: [confdb] crit: cmap_track_delete nodelist failed: CS_ERR_BAD_HANDLE
Oct 01 16:51:26 proxmox1 pmxcfs[1075619]: [confdb] crit: cmap_track_delete version failed: CS_ERR_BAD_HANDLE
Oct 01 16:51:26 proxmox1 pmxcfs[1075619]: [confdb] crit: cmap_finalize failed: CS_ERR_BAD_HANDLE

Der Quorum Server wurde zu dieser Zeit nicht neu gestartet.
 
Die pmxcfs-Meldungen sind eher eine Folge als die Ursache. failed to connect to corosync heißt, dass pmxcfs in dem Moment seine Verbindung zum lokalen corosync verloren hat. Mit dem QDevice hat das erstmal nichts zu tun, corosync auf proxmox1 selbst war weg oder hat sich neu gestartet. Spannender ist also das corosync-Log ab dem Zeitpunkt, wo der Knoten nicht mehr administrierbar wurde:
Code:
journalctl -u corosync --since "2026-10-01 12:00" --until "2026-10-01 17:00"
Schau da auf knet-Meldungen wie "link down", "host has no active links", Token-Timeouts oder neue Memberships. Wenn die zwei Corosync-Links regelmäßig flappen, hast du deine Ursache. Schreib mir auch mal die Ausgabe von pvecm status und corosync-cfgtool -s, solange das Problem noch besteht. Laufen beide Knoten auf demselben Stand? Poste pveversion -v von beiden.
 
Hi.

Heute habe ich das Problem beim Administrieren nicht. Aber in dem besagten Zeitraum haben ich solche Einträge gefunden:

Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] link: Resetting MTU for link 0 because host 1 joined
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 (passive) best link: 0 (pri: 0)
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 has no active links
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 (passive) best link: 0 (pri: 1)
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 has no active links
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 (passive) best link: 0 (pri: 1)
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 has no active links
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 (passive) best link: 0 (pri: 1)
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 has no active links
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 (passive) best link: 0 (pri: 1)
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 has no active links
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 (passive) best link: 0 (pri: 1)
Oct 01 16:53:37 proxmox1 corosync[2040]: [KNET ] host: host: 2 has no active links
Oct 01 16:53:37 proxmox1 corosync[2040]: [QUORUM] Sync members[1]: 1
Oct 01 16:53:37 proxmox1 corosync[2040]: [QUORUM] Sync joined[1]: 1
Oct 01 16:53:37 proxmox1 corosync[2040]: [TOTEM ] A new membership (1.94) was formed. Members joined: 1
Oct 01 16:53:37 proxmox1 corosync[2040]: [QUORUM] Members[1]: 1
Oct 01 16:53:37 proxmox1 corosync[2040]: [MAIN ] Completed service synchronization, ready to provide service.

Alle Verbindungen (Corosync 1 / 2, Storage, VM-Network, DMZ, usw.) sind jeweils an eigenen Switchen angeschlossen.
Geht es bei diesem Fall um das erste Corosync Netzwerk?

Die anderen Ausgaben sehen aktuell normal aus. Beider Knoten sind aktuell auf 9.2.20.
 
Die Zeilen von 16:53:37 sind nicht die Ursache. Das ist corosync beim Neustart nach deinem Restart. "host 2 has no active links" ist nach dem Start normal, bis knet die Verbindung wieder aufgebaut hat. Link 0 ist der Link, den du in der corosync.conf als ring0_addr eingetragen hast, dein erstes Corosync-Netz. Daraus würde ich aber nicht schließen, dass dieser Link das Problem ist. Die pmxcfs-Meldungen von 16:51:24 kommen wahrscheinlich ne Sekunde vor dem Stop durch systemd, das war also auch schon dein Restart.

Interessanter ist, was davor passiert, ab dem Zeitpunkt wo die GUI Timeout lief. Such im corosync-Log über mehrere Stunden vorher nach link: 0 is down, Token has not been received oder Retransmit List. Wenn da nichts steht, hängt eher pmxcfs auf proxmox2 als das Netz. Das kannst du beim nächsten Mal testen: auf proxmox2 ls /etc/pve/nodes ausführen. Hängt der Befehl, weißt du Bescheid. Poste deine /etc/pve/corosync.conf, dann kann man sehen, ob beide Links drin sind.