Monitoring the health/status of an LXC container

SysMik

Member
Feb 8, 2023
8
0
6
Hello,


Is there any feature in Proxmox VE that allows monitoring the health/status of an LXC container and automatically restarting it if it becomes unresponsive or hangs?
For virtual machines, this can be achieved using a virtual watchdog such as the i6300esb watchdog, allowing the Proxmox host to detect a problem and automatically restart the VM.
Is there an equivalent mechanism available for LXC containers?
If not, is there a recommended way to implement such a health check and automatic restart for LXC containers?

Thanks in advance for your help!
 
Do you actually have any LXCs that are failing and need this watchdog/restart functionality? Or is it more of a theoretical question?

From what I'm seeing with search, there is nothing for LXCs specifically but you could utilize systemd to restart failed in-lxc daemons.

If you have an actual problematic LXC, you're probably looking at a homegrown script that runs once a minute at the host level, doing something like:

Code:
PATH=yourpath_here     # put a good path in if running from cron to avoid failures
for vmid in $(pct list |grep running |awk '{print $1}'); do
  timeout 20s pct exec $vmid date
  rc=$?
  if [ $rc -ne 0 ]; then
    pct reboot $vmid
    logger "ERROR - watchdog $0 rebooted $vmid"
    sleep 30 # adjust up as needed to allow for pct reboot to settle
  fi
done
# you could do a sleep 60 here and exec $0 if you just want to schedule it to run @reboot in cron instead of firing off every minute

You might have to add to it some. If you hatch a workable homegrown solution, kindly put it on github or something and lemme know.
 
  • Like
Reactions: Johannes S
Hello,

Thanks for your reply Kingneutron !

No, it's purely theoretical. After experiencing a VM crash, I added an i6300esd watchdog as a safeguard.
Since I couldn't find anything about this in the official documentation, I preferred to ask the community before going down the route of writing a custom script.
By the way, can an LXC actually "crash" given that it uses the host server's kernel? Wouldn't it be more accurate to say that Proxmox itself would have to crash, or that the container's processes/services could fail?
In your opinion, does it actually make sense to have a watchdog for an LXC?
 
This is what I've got so far, still needs work. I've been testing it between loop iterations with kill -19 (sigstop) on an LXC PID but it hangs.

Code:
#!/bin/bash

# 2026.Sep kneutron

PATH=/usr/local/bin:/usr/local/sbin:/usr/bin:/usr/sbin:/bin:/sbin:/usr/local/games:/usr/games:/root/bin:/root/bin/boojum:/usr/X11R6/bin:/usr/NX/bin:
# put a good path in if running from cron to avoid failures

# NOTE pct list     hangs if a pct is kill -19 sigstop, and will affect the webui display updates

echo "$(date) - starting LXC watchdog - running copies: $(ps ax| egrep -v 'grep|jstar' |grep -c lxc-watchdog)"
#for vmid in $(pct list |grep running |awk '{print $1}'); do

# basic parse
function getlxcnpid () {
  lxcvmid=$1
  lxcpid=$2
  echo "VMID: $lxcvmid"
  echo "lxcpid: $lxcpid"
  [ "$lxcpid" = "" ] && exit 44;
}

# vmid, pid
while :; do # forever
  mapfile -t my_array < <( ps ax |grep lxc-start |grep -v grep |awk '{print $NF,$1}')

  for line in "${my_array[@]}"; do 
    echo "linein: $line"
    getlxcnpid $line
    echo "$(date) - Processing running lxc: $lxcvmid"

# basic check lxc response to date request
set -x
    timeout -p 20s pct exec $lxcvmid date
    rc=$?
 
    if [ $rc -ne 0 ]; then
      msg="ERROR - watchdog $0 rebooted $lxcvmid"
      logger "$msg"
      echo "$(date) - $msg"
   
# needs work here - hangs unless ^C, if vmid is kill -19 sigstop
      timeout 45s pct reboot $lxcvmid
      rc=$?
   
      if [ $rc -eq 124 ]; then
        msg="$(date) - Taking drastic measures for LXC $lxcvmid"
        echo "$msg"; logger "$msg"
        kill -9 $lxcpid
        sleep 15
        pct start $lxcvmid
      fi # lxc needed killed
set +x   
      sleep 30 # adjust up as needed to allow for pct reboot to settle
    fi # lxc failed basic date check
  done # line / this vmid
 
set +x
sleep 60 # only run check once a minute
done # forever
 
Last edited: