Proxmox host becomes partially unresponsive after random periods of uptime

taerarenai

New Member
Aug 25, 2026
6
1
1


Full disclosure - the below text is generated by AI since I've spent a LOT of hours trying to troubleshoot these issues with chatgpt/gemini and no luck. If there's any other information that I can provide please let me know. I'm trying to gather everything I can (after the reboot since its totally unreachable right now) and get back to the thread here.

Also, there's no monitor available for me to access the console physically and it happens in the worst scenarios ever (mostly at night or when Im busy and need to give up troubleshooting and just restart...)​

Hardware​

  • Host: Dell OptiPlex 7060
  • CPU: Intel 8th-generation CPU
  • RAM: DDR4
  • Network: Onboard Intel Ethernet adapter using the e1000e Linux driver
  • Storage: SSD
  • Proxmox: Proxmox VE
  • Guest: Home Assistant OS running as a VM
  • Network: Ethernet
  • Proxmox host IP: <PROXMOX_IP>

Problem​

The Proxmox host becomes partially unresponsive after an unpredictable amount of uptime.
The interval is highly variable — it can happen after approximately 3 days, or it can run normally for 3 months.

When the problem occurs:
  • The Proxmox host still responds to ICMP/ping.
  • TCP port 8006 is reachable.
  • TCP port 22 is reachable.
  • However, the Proxmox web interface does not load.
  • The Proxmox mobile app cannot connect.
  • SSH establishes the TCP connection but then the host closes the connection before completing the SSH handshake.
  • HTTPS connections to port 8006 do not complete and eventually time out.
  • The Home Assistant VM behaves inconsistently: sometimes it remains partially accessible, while other times it becomes completely inaccessible.
  • A complete reboot of the Proxmox host reliably restores normal operation.
For example, during the latest failure:

SSH:

Connection established.
...
kex_exchange_identification: Connection closed by remote host
Connection closed by <PROXMOX_IP> port 22
For HTTPS:
curl.exe -k -v --connect-timeout 5 --max-time 10 https://<PROXMOX_IP>:8006/api2/json/version
Connection timeout after 5001 ms
A basic TCP test to port 8006 reports the port as reachable, but an actual HTTPS connection cannot complete.

Previous observations​

I have previously seen e1000e / Intel NIC-related messages on this system, including NIC/link-related errors.
Various troubleshooting/configuration changes have been attempted previously, but none have permanently resolved the issue.
The failure is completely intermittent and does not appear to follow a predictable workload or uptime period.
There are no similar issues with other devices on the network.

Important detail​

The problem is not simply that the network connection goes completely down.

During the failure, the host can still respond to ping and accept TCP connections, but higher-level services such as SSH and the Proxmox API/web interface stop functioning correctly.
A reboot immediately restores normal operation.

I would like to determine whether this is likely related to:

  • e1000e / Intel NIC or driver
  • Linux kernel
  • Proxmox services
  • storage/I/O
  • hardware/firmware
  • another host-level issue

Non-AI: The above behaviour is simply what happened right now, however, there's been times in which everything became totally unresponsive. No pings, no guest VMs, nothing worked. There were also times when unplugging/plugging the ethernet cable got everything back up (when totally unreachable from what I could observe), but other times where only a restart would fix the issue.

I'm not sure if I'm fighting a hardware issue or a proxmox issue at this point and I dont want to spend money on AI suggestions blindly (such as new router, new NICs, new whatever).

Any help/suggestion would be highly appreciated. Thanks in advance !

P.S: I hope this is the correct section for this issue, cant find anything better.
 
I'd like to start with this
Bash:
journalctl -b0 -krp warning
ethtool -k YOURNICHERE
I'd recommend a memory test. Also try to grab a monitor and keep it attached. You need it for doing a memory test anyways. Sometimes things are printed but not logged.
Is the node's ip outside of the DHCP range? Is the microcode package installed? What modifications did you do to the node?
 
Last edited:
Tried delaying this as much as possible until I could get a monitor in, but unfortunately it couldnt wait anymore.
Home assistant VM (HAOS) became totally unresponsive (not just partially as before) so I had to restart it. Maybe I'll be luckier next time and I'll have a bit more time to get a monitor to see whats happening.

IP is in the DHCP range, but it is reserved for the proxmox host ( DHCP List manual assignment on TUF-AX5400).
I remember doing a memory test at some point and it passed with no issues.

Summary:
- DHCP in range, static
- Previous mem test ok
- microcode package installed - yes
- modifications - not sure I follow, theres no special changes made to the proxmox host - simple installation + HAOS vm (vmbr0)


The logs:

Code:
Aug 28 11:30:45 pve kernel: faux_driver regulatory: Direct firmware load for regulatory.db failed with error -2
Aug 28 11:30:01 pve kernel: EXT4-fs warning (device dm-8): ext4_multi_mount_protect:324: MMP interval 42 higher than expected, please wait.
Aug 28 11:30:01 pve kernel: kauditd_printk_skb: 117 callbacks suppressed
Aug 28 11:29:52 pve kernel: block nvme0n1: No UUID available providing old NGUID
Aug 28 11:29:51 pve kernel: nvme nvme0: using unchecked data buffer
Aug 28 11:29:50 pve kernel: spi-nor spi0.0: supply vcc not found, using dummy regulator
Aug 28 11:29:49 pve systemd-journald[360]: File /var/log/journal/2b8b15811a9d44579737cc4df52d9398/system.journal corrupted or uncleanly shut down, renaming and replacing.
Aug 28 11:29:49 pve kernel: zfs: module license taints kernel.
Aug 28 11:29:49 pve kernel: Disabling lock debugging due to kernel taint
Aug 28 11:29:49 pve kernel: zfs: module license 'CDDL' taints kernel.
Aug 28 11:29:49 pve kernel: spl: loading out-of-tree module taints kernel.
Aug 28 11:29:49 pve kernel: e1000e: unknown parameter 'EEE' ignored
Aug 28 11:29:49 pve kernel: wmi_bus wmi_bus-PNP0C14:04: [Firmware Bug]: WQBC data block query control method not found
Aug 28 11:29:49 pve kernel: ENERGY_PERF_BIAS: Set to 'normal', was 'performance'
Aug 28 11:29:49 pve kernel: hpet_acpi_add: no address or irqs in _CRS
Aug 28 11:29:49 pve kernel: x86/cpu: SGX disabled or unsupported by BIOS.

And


Code:
Features for nic0:
rx-checksumming: on
tx-checksumming: on
        tx-checksum-ipv4: off [fixed]
        tx-checksum-ip-generic: on
        tx-checksum-ipv6: off [fixed]
        tx-checksum-fcoe-crc: off [fixed]
        tx-checksum-sctp: off [fixed]
scatter-gather: on
        tx-scatter-gather: on
        tx-scatter-gather-fraglist: off [fixed]
tcp-segmentation-offload: off
        tx-tcp-segmentation: off
        tx-tcp-ecn-segmentation: off [fixed]
        tx-tcp-mangleid-segmentation: off
        tx-tcp6-segmentation: off
        tx-tcp-accecn-segmentation: off [fixed]
generic-segmentation-offload: off
generic-receive-offload: off
large-receive-offload: off [fixed]
rx-vlan-offload: on
tx-vlan-offload: on
ntuple-filters: off [fixed]
receive-hashing: on
highdma: on [fixed]
rx-vlan-filter: off [fixed]
vlan-challenged: off [fixed]
tx-gso-robust: off [fixed]
tx-fcoe-segmentation: off [fixed]
tx-gre-segmentation: off [fixed]
tx-gre-csum-segmentation: off [fixed]
tx-ipxip4-segmentation: off [fixed]
tx-ipxip6-segmentation: off [fixed]
tx-udp_tnl-segmentation: off [fixed]
tx-udp_tnl-csum-segmentation: off [fixed]
tx-gso-partial: off [fixed]
tx-tunnel-remcsum-segmentation: off [fixed]
tx-sctp-segmentation: off [fixed]
tx-esp-segmentation: off [fixed]
tx-udp-segmentation: off [fixed]
tx-gso-list: off [fixed]
tx-nocache-copy: off
loopback: off [fixed]
rx-fcs: off
rx-all: off
tx-vlan-stag-hw-insert: off [fixed]
rx-vlan-stag-hw-parse: off [fixed]
rx-vlan-stag-filter: off [fixed]
l2-fwd-offload: off [fixed]
hw-tc-offload: off [fixed]
esp-hw-offload: off [fixed]
esp-tx-csum-hw-offload: off [fixed]
rx-udp_tunnel-port-offload: off [fixed]
tls-hw-tx-offload: off [fixed]
tls-hw-rx-offload: off [fixed]
rx-gro-hw: off [fixed]
tls-hw-record: off [fixed]
rx-gro-list: off
macsec-hw-offload: off [fixed]
rx-udp-gro-forwarding: off
hsr-tag-ins-offload: off [fixed]
hsr-tag-rm-offload: off [fixed]
hsr-fwd-offload: off [fixed]
hsr-dup-offload: off [fixed]
 
Last edited:
Dear god (or what- or whoever you believe in the most), when is this going to stop?!?
See here:
And now should be the point, when you realize, that the Optiplex 7060 uses an Intel i219-LM.
Have fun!
 
  • Like
Reactions: Johannes S
Not sure what I should be getting out of that thread since its 6 pages long with a lot of back and forth.
I've tried those fixes at some point with no luck (like many others I'm seeing), but other than the fact that its an e1000 driver.....theres nothing saying its the same issue (or I'm missing it).

Are you suggesting changing the NIC is the only solution? I'm fine with that - I just dont wanna spend money on a hunch, hence me asking if there's something else that I need to check beforehand.

Appreciate the info.
 
I've tried those fixes at some point with no luck (like many others I'm seeing), but other than the fact that its an e1000 driver.....theres nothing saying its the same issue (or I'm missing it).

Do you have Detected Hardware Unit Hang entries regarding the e1000e in your logs or not?
 
  • Like
Reactions: Johannes S
Do you have Detected Hardware Unit Hang entries regarding the e1000e in your logs or not?

As of right now I'm unable to find any logs with that entry.
I DID see it show up at some point and applied some fixes mentioned in the above thread.

Not sure if that was fixed or I'm simply not correctly searching for the entry. Please see below:

Code:
root@pve:~# journalctl -k --no-pager | grep -i "Hardware Unit Hang"
root@pve:~# journalctl -k --no-pager | grep -iE "e1000e.*(error|reset|watchdog|hang|timeout)"
root@pve:~# journalctl -k --no-pager | grep -iE "e1000e.*(Link is|link down|link up)"
Apr 13 22:38:11 pve kernel: e1000e 0000:00:1f.6 nic0: NIC Link is Up 1000 Mbps Full Duplex, Flow Control: Rx/Tx
root@pve:~# ethtool -i nic0
driver: e1000e
version: 7.0.0-3-pve
firmware-version: 0.5-4
expansion-rom-version:
bus-info: 0000:00:1f.6
supports-statistics: yes
supports-test: yes
supports-eeprom-access: yes
supports-register-dump: yes
supports-priv-flags: yes
root@pve:~#
 
I find it highly entertaining that prompting AIs for hours and hours about something, you have no clue about, is no problem at all. Whereas at the same time, reading 6 pages of very likely relevant information is unacceptable. Fascinating stuff!
 
Last edited:
  • Like
Reactions: Johannes S
I find it highly entertaining that prompting AIs for hours and hours about something you have no clue about is no problem at all. Whereas at the same time, reading 6 pages of very likely relevant information is unacceptable. Fascinating stuff!

There's probably a saying that exists in every country - if you have nothing good to say, dont say anything.
Ive read the pages there a long time ago and, as mentioned, tried the fixes. There's still people in that thread encountering the issue.

I'm trying things ON TOP of that to ensure that changing the NIC is actually necessary.
Regardless, I'm not gonna entertain this discussion further.
 
  • Like
Reactions: VictorSTS