[SOLVED] Proxmox crashing, hardware issue?

NevadaTech

Member
Nov 29, 2021
24
2
23
61
Hello,

I have a server crashing and I'd like some help. Please! Over the last couple weekend I've had a small maintenance window where I could take it offline for a couple hours. This weekend is a longer window so I'd like to get as much testing in as possible.

So far
  • I've taken the server offline and run Memtest+ on the hardware, about 140mins later the first full pass was finished with no errors, so maybe not the RAM?
  • the BIOS installed is the same as BIOS available v4.20
  • compared working server BIOS settings to this one, no differences that seemed likely (i.e. all of the SRV-IO+IOMMU+etc settings were the same; the BMC were not but that's a diff issue)
At this moment I'm thinking the AMD-Microcode?
  • from webGUI shell you can run grep -i microcode /proc/cpuinfo
  • virt07 microcode: 0x8701034 (working server/no errors)
  • virt09 microcode: 0x8701021 (errors)

Any other ideas of things to check/update while I have some down time?





============cut from wrong forum paste

Please review this excerpt from my Proxmox log. Prox is 8.4.19, the mobo is Asrock X470D4U, the RAM is 128GB (4x32GB) Nemix.

I have 5 other servers similar to this one working fine. The exception is one other server with the same HARDWARE ERROR that pops up on the console. I believe that one also has Nemix RAM. I believe the other four servers have Kingston KSM26ED8/16ME RAM but no errors.

12:37 Hardware Error
13:21 Hardware Error
14:53 Hardware Error
15:45 server reboot

I also see a SMART thermal message but that seems like more of a 'notice'.

Under the Prox log is an output from 'dmidecode -t 17'. While it lists specs it doesn't list actual manufacturer part number. I believe the RAM is actually 3200 speed but running at a lower 2666 speed. I tried an 'lshw -C memory' but lshw is not installed.


------------------------------------------ start some Proxmox log dump
Jun 30 12:37:21 virt09b kernel: mce: [Hardware Error]: Machine check events logged
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: Corrected error, no action required.
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: CPU:0 (17:71:0) MC17_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|Scrub]: 0x9c2041000000011b
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: Error Addr: 0x0000000bbf588300
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: IPID: 0x0000009600050f00, Syndrome: 0x000000040a801101
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
Jun 30 12:37:21 virt09b kernel: EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#1channel#0 (csrow:1 channel:0 page:0x0 offset:0x0 grain:64 syndrome:0x4)
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
Jun 30 12:43:12 virt09b smartd[1543]: Device: /dev/sdb [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 67 to 66
Jun 30 13:04:37 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:08:53 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:17:01 virt09b CRON[151843]: pam_unix(cron:session): session opened for user root(uid=0) by (uid=0)
Jun 30 13:17:01 virt09b CRON[151844]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
Jun 30 13:17:01 virt09b CRON[151843]: pam_unix(cron:session): session closed for user root
Jun 30 13:19:41 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:21:03 virt09b kernel: mce: [Hardware Error]: Machine check events logged
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: Corrected error, no action required.
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: CPU:0 (17:71:0) MC17_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|Scrub]: 0x9c2041000000011b
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: Error Addr: 0x0000000bbf588300
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: IPID: 0x0000009600050f00, Syndrome: 0x000000040a801101
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
Jun 30 13:21:03 virt09b kernel: EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#1channel#0 (csrow:1 channel:0 page:0x0 offset:0x0 grain:64 syndrome:0x4)
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
Jun 30 13:43:12 virt09b smartd[1543]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 67 to 66
Jun 30 13:45:16 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:47:15 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:50:08 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:54:31 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:57:41 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:05:23 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:06:44 virt09b systemd[1]: Starting systemd-tmpfiles-clean.service - Cleanup of Temporary Directories...
Jun 30 14:06:44 virt09b systemd[1]: systemd-tmpfiles-clean.service: Deactivated successfully.
Jun 30 14:06:44 virt09b systemd[1]: Finished systemd-tmpfiles-clean.service - Cleanup of Temporary Directories.
Jun 30 14:06:44 virt09b systemd[1]: run-credentials-systemd\x2dtmpfiles\x2dclean.service.mount: Deactivated successfully.
Jun 30 14:17:01 virt09b CRON[172546]: pam_unix(cron:session): session opened for user root(uid=0) by (uid=0)
Jun 30 14:17:01 virt09b CRON[172547]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
Jun 30 14:17:01 virt09b CRON[172546]: pam_unix(cron:session): session closed for user root
Jun 30 14:27:29 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:28:37 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:35:25 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:43:12 virt09b smartd[1543]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 66 to 67
Jun 30 14:44:19 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:45:10 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:53:09 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:53:53 virt09b kernel: mce: [Hardware Error]: Machine check events logged
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: Corrected error, no action required.
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: CPU:0 (17:71:0) MC17_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|Scrub]: 0x9c2041000000011b
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: Error Addr: 0x0000000bbf520300
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: IPID: 0x0000009600050f00, Syndrome: 0x000000040a801101
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
Jun 30 14:53:53 virt09b kernel: EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#1channel#0 (csrow:1 channel:0 page:0x0 offset:0x0 grain:64 syndrome:0x4)
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
Jun 30 14:56:46 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:06:35 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:07:25 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:08:03 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:13:12 virt09b smartd[1543]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 67 to 66
Jun 30 15:15:21 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:17:01 virt09b CRON[193198]: pam_unix(cron:session): session opened for user root(uid=0) by (uid=0)
Jun 30 15:17:01 virt09b CRON[193199]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
Jun 30 15:17:01 virt09b CRON[193198]: pam_unix(cron:session): session closed for user root
Jun 30 15:38:08 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:45:39 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
-- Reboot --
Jun 30 15:47:33 virt09b kernel: Linux version 6.8.12-30-pve (build@proxmox) (gcc (Debian 12.2.0-14+deb12u1) 12.2.0, GNU ld (GNU Binutils for Debian) 2.40) #1 SMP PREEMPT_DYNAMIC PMX 6.8.12-30 (2026-06-11T10:10Z) ()
Jun 30 15:47:33 virt09b kernel: Command line: BOOT_IMAGE=/boot/vmlinuz-6.8.12-30-pve root=/dev/mapper/pve-root ro quiet
Jun 30 15:47:33 virt09b kernel: KERNEL supported cpus:
Jun 30 15:47:33 virt09b kernel: Intel GenuineIntel



------------------------------------------ end some Proxmox log dump



------------------------------------------
dmidecode -t 17 show

Memory Device
Array Handle: 0x0014
Error Information Handle: 0x0021
Total Width: 72 bits
Data Width: 64 bits
Size: 32 GB
Form Factor: DIMM
Set: None
Locator: DIMM 0
Bank Locator: P0 CHANNEL B
Type: DDR4
Type Detail: Synchronous Unbuffered (Unregistered)
Speed: 2666 MT/s
Manufacturer: Unknown
Serial Number: 5D270016
Asset Tag: Not Specified
Part Number: Unknown
Rank: 2
Configured Memory Speed: 2666 MT/s
Minimum Voltage: 1.2 V
Maximum Voltage: 1.2 V
Configured Voltage: 1.2 V
Memory Technology: DRAM
Memory Operating Mode Capability: Volatile memory
Firmware Version: Unknown
Module Manufacturer ID: Unknown
Module Product ID: Unknown
Memory Subsystem Controller Manufacturer ID: Unknown
Memory Subsystem Controller Product ID: Unknown
Non-Volatile Size: None
Volatile Size: 32 GB
Cache Size: None
Logical Size: None

------------------------------------------


============end cut from wrong forum paste
 
Why is there a Proxmox VE v8.4.19 running?
What CPU Type do you use?

Under the Prox log is an output from 'dmidecode -t 17'. While it lists specs it doesn't list actual manufacturer part number. I believe the RAM is actually 3200 speed but running at a lower 2666 speed. I tried an 'lshw -C memory' but lshw is not installed.
So get the lshw programm with apt install lshw.

for the Asrock X470D4U you can download the English manual.
and on Page 22 you find:

So your 2666 MT/s is normal with 4x DDR4 DIMM.
 

Attachments

  • Bildschirmfoto vom 2026-07-17 23-45-04.png
    Bildschirmfoto vom 2026-07-17 23-45-04.png
    156.4 KB · Views: 4
  • Like
Reactions: UdoB
You can swap the RAM with one of your known-good servers first and see if the errors follow the DIMMs. The logs make me think it's more likely a memory issue than the microcode, especially with another system using the same Nemix RAM showing similar errors.
 
  • Like
Reactions: cwt
The same syndrome and nearly the same physical address occur multiple times, so this looks more like a persistent problem with a DIMM, memory channel, socket contact, memory controller or unstable RAM configuration than a random single-bit event.

“Corrected error, no action required” only means ECC was able to correct the data. Repeated corrected errors should still be treated as a hardware warning, because they may eventually become uncorrectable errors.
 
Here's another cut from today's console messages dumped via dmesg.

I'm leaning to the microcode because the error

cache level: L3/GEN, tx: GEN, mem-tx: RD

seems to reference the L3 cache of the CPU. And the AMD-Vi error is related to IO-MMU. To me, with the other things tested, that leaves microcode is not updated enough to process this stuff.

I'll probably try davidlynch's idea of swapping RAM-to-servers to see if the issue follows. I should have enough time tomorrow for a test like that. This was my first attempt at buying Nemix RAM so I'm jaded into thinking that was/is the cause.

<code>
[536970.519936] AMD-Vi: Completion-Wait loop timed out
[536974.203495] AMD-Vi: Completion-Wait loop timed out
[536975.603850] AMD-Vi: Completion-Wait loop timed out
[537044.229264] AMD-Vi: Completion-Wait loop timed out
[537085.047449] AMD-Vi: Completion-Wait loop timed out
[537203.453094] AMD-Vi: Completion-Wait loop timed out
[537210.329004] AMD-Vi: Completion-Wait loop timed out
[537241.554355] AMD-Vi: Completion-Wait loop timed out
[537244.521378] AMD-Vi: Completion-Wait loop timed out
[537324.197687] AMD-Vi: Completion-Wait loop timed out
[537700.683364] AMD-Vi: Completion-Wait loop timed out
[538354.266397] AMD-Vi: Completion-Wait loop timed out
[538700.540208] mce: [Hardware Error]: Machine check events logged
[538700.540214] [Hardware Error]: Corrected error, no action required.
[538700.540403] [Hardware Error]: CPU:0 (17:71:0) MC18_STATUS[Over|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|-]: 0xdc2040000000011b
[538700.540570] [Hardware Error]: Error Addr: 0x0000000fbfcf7280
[538700.540713] [Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x000000040a801103
[538700.540858] [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
[538700.540869] EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#3channel#1 (csrow:3 channel:1 page:0x0 offset:0x0 grain:64 syndrome:0x4)
[538700.541035] [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
[539028.213586] mce: [Hardware Error]: Machine check events logged
[539028.213591] [Hardware Error]: Corrected error, no action required.
[539028.213817] [Hardware Error]: CPU:0 (17:71:0) MC18_STATUS[Over|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|-]: 0xdc2040000000011b
[539028.214005] [Hardware Error]: Error Addr: 0x0000000bbf561f80
[539028.214176] [Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x040004000a801203
[539028.214356] [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
[539028.214368] EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#3channel#1 (csrow:3 channel:1 page:0x0 offset:0x0 grain:64 syndrome:0x400)
[539028.214552] [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
root@virt09b:~#

</code>
 
About 2 days ago I
  • updated the AMD-microcode
    • waited for a couple hours and no CPU/memory errors showed up on the console, the AMD-Vi wait loop still appeared but it seemed a little less frequent (seemed or reality?)
  • shutdown
  • I then went into the BIOS and set the IOMMU to Enabled instead of Auto; on all of my other in-house Asrock X470D4U based serves it is set to Auto
    • waited another couple hours and no AMD-Vi wait loop (or CPU/memory) messages appeared on the console
  • since Sunday usage is light I wanted to see what happened on a Monday
    • so far, Proxmox> Shell> dmesg shows neither error
Obviously I'll keep monitoring this but I believe it is fixed. I really was leaning into Nemix RAM as being the culprit but that was probably prejudice because I was not familiar with the brand name.

Now I can continue on with my Proxmox v9 migration. My initial lightly used production server suffered the 'no boot/straight to BIOS' issue - see --- https://forum.proxmox.com/threads/no-boot-after-pve-8-to-9-upgrade.170658/
 
The same syndrome and nearly the same physical address occur multiple times, so this looks more like a persistent problem with a DIMM, memory channel, socket contact, memory controller or unstable RAM configuration than a random single-bit event.

“Corrected error, no action required” only means ECC was able to correct the data. Repeated corrected errors should still be treated as a hardware warning, because they may eventually become uncorrectable errors.
Yes, I agree completely. The similarity+constant is why I thought the DIMM was bad. In one maintenance window I pulled the RAM and installed it reverse order. I also let the mobo dictate the default memory settings. I sure there are ways to eek out 10% more performance if I tweaked settings but server stability is more important.

And while the server was saying it was correct-able I felt it was an impending disaster. And I'm blaming that on the 2-3 crashes over the last month. I think those times is was 'un-correctable'.
 
NOT SOLVED - I just don't see a way to change the title back :(

Blah. After a few days the CPU/unified memory error has returned.

I have not seen the AMD-Vi error return - I'm sticking with the IOMMU set to Enabled (vs Auto) was the fix for that.

Current plan is to

  • swap the memory from the physical server virt09 (Nemix RAM) to physical server virt07 (Kingston RAM) next weekend and see if the error follows
 
There are muiltiple possibilities wich can cause the error:
Motherboard, Ram, unstable PSU or even things like thermal paste which has lost his thermal conductivity due to aging.
Also Bios and OS settings, or damaged files on OS disk.
You need to change things one by one to find the faulty component.

I sure there are ways to eek out 10% more performance if I tweaked settings but server stability is more important.
If you want more peformance, and shorten searching - change the mainparts and upgrade:
- new motherbaord (.i.e. AsrockRack x570d4u)
- new CPU AMD Ryzen 9 5950X, 16C/32T, 3.40-4.90GHz, its only ~320 Euro

If you have small Maintanance Windows, think about a Cluster Setup with CEPH.