No web UI randomly

FiltroMan

Member
Dec 30, 2022
15
12
8
Hi there! I've searched far and wide on the forums as well as using Gemini but I've been encountering a weird issue in the past couple of days: out of the blue (perhaps I'm not looking at the right logs) my web UI dies, and connecting straight to the host I notice that there are issues with pve-cluster.service which goes into error saying

Start request repeated too quickly
Failed with result 'exit-code'

Another error I get is also about fuse mountpoint not empty: with trial and error I manage to get access back to the web UI after deleting everything in /etc/pve and restarting pveproxy along with pve-cluster, any way to understand what's happening and how to fix it once and for all?

I can provide more details if needed.
 
with trial and error I manage to get access back to the web UI after deleting everything in /etc/pve
This in itself requires a reinstallation. I don't see any way around it.

and connecting straight to the host I notice that there are issues with pve-cluster.service which goes into error saying

Start request repeated too quickly
Failed with result 'exit-code'
A systemctl status pve-cluster would have helped or showing a relevenant part of journalctl -b 0. But you need to reinstall Proxmox 9.2 after removing everything under /etc/pve.
 
This in itself requires a reinstallation. I don't see any way around it.


A systemctl status pve-cluster would have helped or showing a relevenant part of journalctl -b 0. But you need to reinstall Proxmox 9.2 after removing everything under /etc/pve.
Well... I'd much rather avoid having to reinstall, but heck if there's no other way around it, I'm glad that a couple of days ago I finally managed to get PBS up and running, and made backups of all my LXCs and VMs

~
â—Ź pve-cluster.service - The Proxmox VE cluster filesystem
Loaded: loaded (/usr/lib/systemd/system/pve-cluster.service; >
Active: active (running) since Sat 2026-10-10 20:40:57 CEST; >
Invocation: f434eed6ea584479b3c5301f644aa048
Process: 1655 ExecStart=/usr/bin/pmxcfs (code=exited, status=0>
Main PID: 1656 (pmxcfs)
Tasks: 6 (limit: 25790)
Memory: 19.9M (peak: 20.5M)
CPU: 867ms
CGroup: /system.slice/pve-cluster.service
└─1656 /usr/bin/pmxcfs

Oct 10 20:40:56 proxmox systemd[1]: Starting pve-cluster.service ->
Oct 10 20:40:56 proxmox pmxcfs[1655]: [main] notice: resolved node>
Oct 10 20:40:56 proxmox pmxcfs[1655]: [main] notice: resolved node>
Oct 10 20:40:57 proxmox systemd[1]: Started pve-cluster.service
Then I guess I can trigger the issue again and see journalctl -b 0
 
Well... I'd much rather avoid having to reinstall, but heck if there's no other way around it, I'm glad that a couple of days ago I finally managed to get PBS up and running, and made backups of all my LXCs and VMs
I understand but deleting /etc/pve is much like uninstalling Proxmox.
Then I guess I can trigger the issue again and see journalctl -b 0
To help troubleshooting an issue people need concrete error messages and logs.
 
Weirdly enough, the two times I had to do it after a wee bit, all my hosts came back up
I did not realize you had a cluster; how many hosts and do you have a quorum device?
I rebooted the system and here's attached the result of journalctl -b 0
I can't see it, sorry. If the cluster worked as normal then -b 0 won't help. We need a system log when it's failing.
 
I do not have a cluster, it's a single machine
The why do you talk about multiple hosts?
It is currently failing though, I have no web UI: what other logs can I provide?
Can you login with SSH and get the journalctl -b 0? I see an attached file now (see end of my post) but not before.
Weirdly enough, the two times I had to do it after a wee bit, all my hosts came back up
Does this mean they work normally? Or what do you want to say with this message which suggest multiple hosts when you have only one?

I rebooted the system and here's attached the result of journalctl -b 0
After this everything look Chinese:
Oct 10 22:42:10 proxmox pveproxy[1810]: /etc/pve/local/pve-ssl.key: failed to load local private key (key_file or key) at /usr/share/perl5/PVE/APIServer/AnyEvent.pm line 2150.
Maybe run a memtest to check if memory is alright? Maybe check your hardware, reinstall Proxmox and restore your VM/CT from backups?
 
Can you login with SSH and get the journalctl -b 0? I see an attached file now (see end of my post) but not before.
I have attached it in the previous message, spotty internet on my side.
Does this mean they work normally? Or what do you want to say with this message which suggest multiple hosts when you have only one?
I was referring to my LXCs and VMs by saying "my hosts", sorry for the confusion: the first time around after deleting the folder "nodes" contained in /etc/pve and restarting all the services every LXC and VM came back up nicely, shame that it decided to pull the same stunt today.
After this everything look Chinese:
Oct 10 22:42:10 proxmox pveproxy[1810]: /etc/pve/local/pve-ssl.key: failed to load local private key (key_file or key) at /usr/share/perl5/PVE/APIServer/AnyEvent.pm line 2150.
Maybe run a memtest to check if memory is alright? Maybe check your hardware, reinstall Proxmox and restore your VM/CT from backups?
I was genuinely wondering if I should run memtest and a check on the drives, after all... IMHO it kinda looks more something software related but I might as well be completely wrong about this.