Hello all,
we have a server running Proxmox Backup Server here, along with a Proxmox Virtual Environment running a single guest.
The system initially ran without any issues with PBS 3 and PVE 8 and was updated to PBS 4 / PVE 8 in May of this year, following the instructions at https://pbs.proxmox.com/wiki/Upgrade_from_3_to_4 and https://pve.proxmox.com/wiki/Upgrade_from_8_to_9. All packages are continuously updated; that is, we started in May with kernel 7.0.2-x and are now currently at 7.0.14-19.pve.
Unfortunately, since switching to the new major versions, we have been noticing system instabilities, which in most cases manifest during the (daily) verify job (though this is also the only job that runs regularly for an extended period of time).
These instabilities look in journalctl like this:
2–3 times a month:
1-2 times a month:
Just once:
There have also been two occasions since May where the verify job flagged a chunk as ".bad", even though the system runs entirely on ZFS and has never reported an error during any of the scrubs performed (including those started manually). In both cases, however, it was successfully verified that the supposedly corrupt chunk file matched, byte for byte, the version that was then restored from tape, and a subsequent verification of the file (by removing the “.bad” extension from the filename) accepted the chunk as correct again.
Actually, these issues seem more like a hardware failure or a memory problem. However, the hardware (AMD EPYC 9015 CPU, Supermicro H13SSL-NT motherboard, 64 GB ECC Reg Micron DDR5 4800 RAM) is only about 18 months old and we can't find any warnings or error messages in either the kernel log or the IPMI health event log. Even a Memtest86+ run with all patterns for over 16 hours and more than 20 passes didn't detect any errors.
Our hardware supplier rules out a hardware defect and asked when the problems started. Well, unfortunately, we have to admit: Since the Proxmox update... It would be quite a coincidence if hardware issues had occurred precisely at the same time as our Proxmox release updates. On the other hand, we also updated several other PVEs from version 8 to 9 around the same time, and they're all still working without any issues.
Does anyone have any suggestions on what else we could try?
we have a server running Proxmox Backup Server here, along with a Proxmox Virtual Environment running a single guest.
The system initially ran without any issues with PBS 3 and PVE 8 and was updated to PBS 4 / PVE 8 in May of this year, following the instructions at https://pbs.proxmox.com/wiki/Upgrade_from_3_to_4 and https://pve.proxmox.com/wiki/Upgrade_from_8_to_9. All packages are continuously updated; that is, we started in May with kernel 7.0.2-x and are now currently at 7.0.14-19.pve.
Unfortunately, since switching to the new major versions, we have been noticing system instabilities, which in most cases manifest during the (daily) verify job (though this is also the only job that runs regularly for an extended period of time).
These instabilities look in journalctl like this:
2–3 times a month:
proxmox-backup-proxy[1954]: Fatal glibc error: malloc.c:3091 (mremap_chunk): assertion failed: prev_size (p) == offset
systemd[1]: proxmox-backup-proxy.service: Main process exited, code=killed, status=6/ABRT
1-2 times a month:
proxmox-backup-proxy[413261]: free(): invalid pointer
systemd[1]: proxmox-backup-proxy.service: Main process exited, code=killed, status=6/ABRT
Just once:
kernel: traps: pvestatd[2453] general protection fault ip:5bd4935bc605 sp:7ffdc4fa3530 error:0 in perl[df605,5bd493521000+1ae000]
systemd[1]: pvestatd.service: Main process exited, code=killed, status=11/SEGV
There have also been two occasions since May where the verify job flagged a chunk as ".bad", even though the system runs entirely on ZFS and has never reported an error during any of the scrubs performed (including those started manually). In both cases, however, it was successfully verified that the supposedly corrupt chunk file matched, byte for byte, the version that was then restored from tape, and a subsequent verification of the file (by removing the “.bad” extension from the filename) accepted the chunk as correct again.
Actually, these issues seem more like a hardware failure or a memory problem. However, the hardware (AMD EPYC 9015 CPU, Supermicro H13SSL-NT motherboard, 64 GB ECC Reg Micron DDR5 4800 RAM) is only about 18 months old and we can't find any warnings or error messages in either the kernel log or the IPMI health event log. Even a Memtest86+ run with all patterns for over 16 hours and more than 20 passes didn't detect any errors.
Our hardware supplier rules out a hardware defect and asked when the problems started. Well, unfortunately, we have to admit: Since the Proxmox update... It would be quite a coincidence if hardware issues had occurred precisely at the same time as our Proxmox release updates. On the other hand, we also updated several other PVEs from version 8 to 9 around the same time, and they're all still working without any issues.
Does anyone have any suggestions on what else we could try?