[SOLVED] Proxmox crashing, hardware issue?

NevadaTech

Member
Nov 29, 2021
23
2
23
61
Hello,

I have a server crashing and I'd like some help. Please! Over the last couple weekend I've had a small maintenance window where I could take it offline for a couple hours. This weekend is a longer window so I'd like to get as much testing in as possible.

So far
  • I've taken the server offline and run Memtest+ on the hardware, about 140mins later the first full pass was finished with no errors, so maybe not the RAM?
  • the BIOS installed is the same as BIOS available v4.20
  • compared working server BIOS settings to this one, no differences that seemed likely (i.e. all of the SRV-IO+IOMMU+etc settings were the same; the BMC were not but that's a diff issue)
At this moment I'm thinking the AMD-Microcode?
  • from webGUI shell you can run grep -i microcode /proc/cpuinfo
  • virt07 microcode: 0x8701034 (working server/no errors)
  • virt09 microcode: 0x8701021 (errors)

Any other ideas of things to check/update while I have some down time?





============cut from wrong forum paste

Please review this excerpt from my Proxmox log. Prox is 8.4.19, the mobo is Asrock X470D4U, the RAM is 128GB (4x32GB) Nemix.

I have 5 other servers similar to this one working fine. The exception is one other server with the same HARDWARE ERROR that pops up on the console. I believe that one also has Nemix RAM. I believe the other four servers have Kingston KSM26ED8/16ME RAM but no errors.

12:37 Hardware Error
13:21 Hardware Error
14:53 Hardware Error
15:45 server reboot

I also see a SMART thermal message but that seems like more of a 'notice'.

Under the Prox log is an output from 'dmidecode -t 17'. While it lists specs it doesn't list actual manufacturer part number. I believe the RAM is actually 3200 speed but running at a lower 2666 speed. I tried an 'lshw -C memory' but lshw is not installed.


------------------------------------------ start some Proxmox log dump
Jun 30 12:37:21 virt09b kernel: mce: [Hardware Error]: Machine check events logged
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: Corrected error, no action required.
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: CPU:0 (17:71:0) MC17_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|Scrub]: 0x9c2041000000011b
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: Error Addr: 0x0000000bbf588300
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: IPID: 0x0000009600050f00, Syndrome: 0x000000040a801101
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
Jun 30 12:37:21 virt09b kernel: EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#1channel#0 (csrow:1 channel:0 page:0x0 offset:0x0 grain:64 syndrome:0x4)
Jun 30 12:37:21 virt09b kernel: [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
Jun 30 12:43:12 virt09b smartd[1543]: Device: /dev/sdb [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 67 to 66
Jun 30 13:04:37 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:08:53 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:17:01 virt09b CRON[151843]: pam_unix(cron:session): session opened for user root(uid=0) by (uid=0)
Jun 30 13:17:01 virt09b CRON[151844]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
Jun 30 13:17:01 virt09b CRON[151843]: pam_unix(cron:session): session closed for user root
Jun 30 13:19:41 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:21:03 virt09b kernel: mce: [Hardware Error]: Machine check events logged
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: Corrected error, no action required.
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: CPU:0 (17:71:0) MC17_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|Scrub]: 0x9c2041000000011b
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: Error Addr: 0x0000000bbf588300
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: IPID: 0x0000009600050f00, Syndrome: 0x000000040a801101
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
Jun 30 13:21:03 virt09b kernel: EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#1channel#0 (csrow:1 channel:0 page:0x0 offset:0x0 grain:64 syndrome:0x4)
Jun 30 13:21:03 virt09b kernel: [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
Jun 30 13:43:12 virt09b smartd[1543]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 67 to 66
Jun 30 13:45:16 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:47:15 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:50:08 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:54:31 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 13:57:41 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:05:23 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:06:44 virt09b systemd[1]: Starting systemd-tmpfiles-clean.service - Cleanup of Temporary Directories...
Jun 30 14:06:44 virt09b systemd[1]: systemd-tmpfiles-clean.service: Deactivated successfully.
Jun 30 14:06:44 virt09b systemd[1]: Finished systemd-tmpfiles-clean.service - Cleanup of Temporary Directories.
Jun 30 14:06:44 virt09b systemd[1]: run-credentials-systemd\x2dtmpfiles\x2dclean.service.mount: Deactivated successfully.
Jun 30 14:17:01 virt09b CRON[172546]: pam_unix(cron:session): session opened for user root(uid=0) by (uid=0)
Jun 30 14:17:01 virt09b CRON[172547]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
Jun 30 14:17:01 virt09b CRON[172546]: pam_unix(cron:session): session closed for user root
Jun 30 14:27:29 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:28:37 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:35:25 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:43:12 virt09b smartd[1543]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 66 to 67
Jun 30 14:44:19 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:45:10 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:53:09 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 14:53:53 virt09b kernel: mce: [Hardware Error]: Machine check events logged
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: Corrected error, no action required.
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: CPU:0 (17:71:0) MC17_STATUS[-|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|Scrub]: 0x9c2041000000011b
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: Error Addr: 0x0000000bbf520300
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: IPID: 0x0000009600050f00, Syndrome: 0x000000040a801101
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
Jun 30 14:53:53 virt09b kernel: EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#1channel#0 (csrow:1 channel:0 page:0x0 offset:0x0 grain:64 syndrome:0x4)
Jun 30 14:53:53 virt09b kernel: [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
Jun 30 14:56:46 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:06:35 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:07:25 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:08:03 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:13:12 virt09b smartd[1543]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 67 to 66
Jun 30 15:15:21 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:17:01 virt09b CRON[193198]: pam_unix(cron:session): session opened for user root(uid=0) by (uid=0)
Jun 30 15:17:01 virt09b CRON[193199]: (root) CMD (cd / && run-parts --report /etc/cron.hourly)
Jun 30 15:17:01 virt09b CRON[193198]: pam_unix(cron:session): session closed for user root
Jun 30 15:38:08 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
Jun 30 15:45:39 virt09b kernel: AMD-Vi: Completion-Wait loop timed out
-- Reboot --
Jun 30 15:47:33 virt09b kernel: Linux version 6.8.12-30-pve (build@proxmox) (gcc (Debian 12.2.0-14+deb12u1) 12.2.0, GNU ld (GNU Binutils for Debian) 2.40) #1 SMP PREEMPT_DYNAMIC PMX 6.8.12-30 (2026-06-11T10:10Z) ()
Jun 30 15:47:33 virt09b kernel: Command line: BOOT_IMAGE=/boot/vmlinuz-6.8.12-30-pve root=/dev/mapper/pve-root ro quiet
Jun 30 15:47:33 virt09b kernel: KERNEL supported cpus:
Jun 30 15:47:33 virt09b kernel: Intel GenuineIntel



------------------------------------------ end some Proxmox log dump



------------------------------------------
dmidecode -t 17 show

Memory Device
Array Handle: 0x0014
Error Information Handle: 0x0021
Total Width: 72 bits
Data Width: 64 bits
Size: 32 GB
Form Factor: DIMM
Set: None
Locator: DIMM 0
Bank Locator: P0 CHANNEL B
Type: DDR4
Type Detail: Synchronous Unbuffered (Unregistered)
Speed: 2666 MT/s
Manufacturer: Unknown
Serial Number: 5D270016
Asset Tag: Not Specified
Part Number: Unknown
Rank: 2
Configured Memory Speed: 2666 MT/s
Minimum Voltage: 1.2 V
Maximum Voltage: 1.2 V
Configured Voltage: 1.2 V
Memory Technology: DRAM
Memory Operating Mode Capability: Volatile memory
Firmware Version: Unknown
Module Manufacturer ID: Unknown
Module Product ID: Unknown
Memory Subsystem Controller Manufacturer ID: Unknown
Memory Subsystem Controller Product ID: Unknown
Non-Volatile Size: None
Volatile Size: 32 GB
Cache Size: None
Logical Size: None

------------------------------------------


============end cut from wrong forum paste
 
Why is there a Proxmox VE v8.4.19 running?
What CPU Type do you use?

Under the Prox log is an output from 'dmidecode -t 17'. While it lists specs it doesn't list actual manufacturer part number. I believe the RAM is actually 3200 speed but running at a lower 2666 speed. I tried an 'lshw -C memory' but lshw is not installed.
So get the lshw programm with apt install lshw.

for the Asrock X470D4U you can download the English manual.
and on Page 22 you find:

So your 2666 MT/s is normal with 4x DDR4 DIMM.
 

Attachments

  • Bildschirmfoto vom 2026-07-17 23-45-04.png
    Bildschirmfoto vom 2026-07-17 23-45-04.png
    156.4 KB · Views: 4
  • Like
Reactions: UdoB
You can swap the RAM with one of your known-good servers first and see if the errors follow the DIMMs. The logs make me think it's more likely a memory issue than the microcode, especially with another system using the same Nemix RAM showing similar errors.
 
  • Like
Reactions: cwt
The same syndrome and nearly the same physical address occur multiple times, so this looks more like a persistent problem with a DIMM, memory channel, socket contact, memory controller or unstable RAM configuration than a random single-bit event.

“Corrected error, no action required” only means ECC was able to correct the data. Repeated corrected errors should still be treated as a hardware warning, because they may eventually become uncorrectable errors.
 
Here's another cut from today's console messages dumped via dmesg.

I'm leaning to the microcode because the error

cache level: L3/GEN, tx: GEN, mem-tx: RD

seems to reference the L3 cache of the CPU. And the AMD-Vi error is related to IO-MMU. To me, with the other things tested, that leaves microcode is not updated enough to process this stuff.

I'll probably try davidlynch's idea of swapping RAM-to-servers to see if the issue follows. I should have enough time tomorrow for a test like that. This was my first attempt at buying Nemix RAM so I'm jaded into thinking that was/is the cause.

<code>
[536970.519936] AMD-Vi: Completion-Wait loop timed out
[536974.203495] AMD-Vi: Completion-Wait loop timed out
[536975.603850] AMD-Vi: Completion-Wait loop timed out
[537044.229264] AMD-Vi: Completion-Wait loop timed out
[537085.047449] AMD-Vi: Completion-Wait loop timed out
[537203.453094] AMD-Vi: Completion-Wait loop timed out
[537210.329004] AMD-Vi: Completion-Wait loop timed out
[537241.554355] AMD-Vi: Completion-Wait loop timed out
[537244.521378] AMD-Vi: Completion-Wait loop timed out
[537324.197687] AMD-Vi: Completion-Wait loop timed out
[537700.683364] AMD-Vi: Completion-Wait loop timed out
[538354.266397] AMD-Vi: Completion-Wait loop timed out
[538700.540208] mce: [Hardware Error]: Machine check events logged
[538700.540214] [Hardware Error]: Corrected error, no action required.
[538700.540403] [Hardware Error]: CPU:0 (17:71:0) MC18_STATUS[Over|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|-]: 0xdc2040000000011b
[538700.540570] [Hardware Error]: Error Addr: 0x0000000fbfcf7280
[538700.540713] [Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x000000040a801103
[538700.540858] [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
[538700.540869] EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#3channel#1 (csrow:3 channel:1 page:0x0 offset:0x0 grain:64 syndrome:0x4)
[538700.541035] [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
[539028.213586] mce: [Hardware Error]: Machine check events logged
[539028.213591] [Hardware Error]: Corrected error, no action required.
[539028.213817] [Hardware Error]: CPU:0 (17:71:0) MC18_STATUS[Over|CE|MiscV|AddrV|-|-|SyndV|CECC|-|-|-]: 0xdc2040000000011b
[539028.214005] [Hardware Error]: Error Addr: 0x0000000bbf561f80
[539028.214176] [Hardware Error]: IPID: 0x0000009600150f00, Syndrome: 0x040004000a801203
[539028.214356] [Hardware Error]: Unified Memory Controller Ext. Error Code: 0
[539028.214368] EDAC MC0: 1 CE Cannot decode normalized address on mc#0csrow#3channel#1 (csrow:3 channel:1 page:0x0 offset:0x0 grain:64 syndrome:0x400)
[539028.214552] [Hardware Error]: cache level: L3/GEN, tx: GEN, mem-tx: RD
root@virt09b:~#

</code>
 
About 2 days ago I
  • updated the AMD-microcode
    • waited for a couple hours and no CPU/memory errors showed up on the console, the AMD-Vi wait loop still appeared but it seemed a little less frequent (seemed or reality?)
  • shutdown
  • I then went into the BIOS and set the IOMMU to Enabled instead of Auto; on all of my other in-house Asrock X470D4U based serves it is set to Auto
    • waited another couple hours and no AMD-Vi wait loop (or CPU/memory) messages appeared on the console
  • since Sunday usage is light I wanted to see what happened on a Monday
    • so far, Proxmox> Shell> dmesg shows neither error
Obviously I'll keep monitoring this but I believe it is fixed. I really was leaning into Nemix RAM as being the culprit but that was probably prejudice because I was not familiar with the brand name.

Now I can continue on with my Proxmox v9 migration. My initial lightly used production server suffered the 'no boot/straight to BIOS' issue - see --- https://forum.proxmox.com/threads/no-boot-after-pve-8-to-9-upgrade.170658/
 
The same syndrome and nearly the same physical address occur multiple times, so this looks more like a persistent problem with a DIMM, memory channel, socket contact, memory controller or unstable RAM configuration than a random single-bit event.

“Corrected error, no action required” only means ECC was able to correct the data. Repeated corrected errors should still be treated as a hardware warning, because they may eventually become uncorrectable errors.
Yes, I agree completely. The similarity+constant is why I thought the DIMM was bad. In one maintenance window I pulled the RAM and installed it reverse order. I also let the mobo dictate the default memory settings. I sure there are ways to eek out 10% more performance if I tweaked settings but server stability is more important.

And while the server was saying it was correct-able I felt it was an impending disaster. And I'm blaming that on the 2-3 crashes over the last month. I think those times is was 'un-correctable'.