Host shell fails to start, cluster with failed nodes

andreisrr

Member
Feb 2, 2024
71
8
13
I had a 3 node mini cluster that became a 2 node cluster after having to give away one host for other purposes.
Today there is a hardware failure with one of the nodes, still working on it.
No VM would start on the remaining working node because of lack of quorum.

While attempting the temporary workaround described in this thread: https://forum.proxmox.com/threads/start-a-vm-without-turning-on-all-nodes.54470/#post-250881 I hit another major problem:
- tried to access the node (host) shell
got error:
failed waiting for client: timed out
TASK ERROR: command '/usr/bin/termproxy 5900 --path /nodes/pve03 --perm Sys.Console -- /bin/login -f root' failed: exit code 1
This happens while the shell console shows a bubble notification "connecting..." then shows "undefined ... code 1006".

Rebooting this node doesn't help.
Since at this time I can only access it remotely and only through the web interface, I can't seem to be able to do more.

All nodes running PVE-8.4.17, upgrade to PVE9 is being considered but not immediate.

EDIT: Finally, after intervening at hardware level, the second node started and I regained VM startability and access to shell on each node.
 
Last edited:
he second node started and I regained VM startability and access to shell on each node.
Great!

But nevertheless: search for "Quorum Device" and set one up asap. --> https://pve.proxmox.com/pve-docs/pve-admin-guide.html#_corosync_external_vote_support

And with the same importance: you won't get any (possibly important!) updates for 8.x as it is EOL'ed - which is unacceptable for a commercial environment... --> https://pve.proxmox.com/pve-docs/chapter-pve-faq.html#faq-support-table
 
I can understand why VMs wouldn't start.
But ...
- Is it normal behaviour for a PVE node in a degraded cluster without quorum to not have a functioning shell (via web ui) with such error?
- Is it a known issue with 8.x PVE?
- without a quorum device, within the existing frame, how can I regain shell access via web ui to that node? (assuming without dismantling the cluster setup)
 
connect through SSH
issue this command : pvecm expected 1
now you can operate, but warning : there is no split brain mecanism
example : move vm configuration files to the remaining host directories in order to boot them (if storage is available).
 
  • Like
Reactions: UdoB