hi,
i added 2x micron 7400 (m.2) as zfs special device in kernel 6.8.x
and had random one of the drives drop/offline after 1-2 days.
i think this fixed it for me:
/etc/kernel/cmdline
pcie_aspm=off pcie_port_pm=off
now with 6 days of uptime on...
We have many points in common:
The hardware worked well in a previous software configuration
Ceph is used for SSDs, is it a CephRBD volume?
Several SSDs and several nodes are affected
The failure is random but the more the system is loaded, the...
Happy to give some background and thx for offering to give your thoughts…
Nodes: 3
Drive details per node: 1&2) 1x Samsung PM1733a 15.36TB, 1x Optane 905P 380GB as db; 3) 2x Intel D4502 7.68TB (dual link running on single link), 2x Optane P1600X...
I also want to add that I’ve been having problems with 6.8.12-9 and even 6.11.11-2. When I was using 6.8.12-4 it was somewhat ok. For reference my disks are Intel D4502 7.68TB U.2 NVMe. Other NVMe drives are fine however. Maybe I’ll revert to...
Hi,
installed a new node today, but when trying to update, I am encountering ssl issues.
Could it be that proxmox has certificate problems on their repo server?
root@pve:~# curl -v https://download.proxmox.com/debian/pve/dists/trixie/InRelease...
Hi Team,
Before connecting the SSDs with an M.2 adapter in a PCIe slot, I noticed that the disconnection of the SSDs occurs only under very specific conditions.
My configuration is based on 4 nodes and 1 qdevice with the latest PVE community...
I don't know how you have those NVMEs connected to the system, but I would try a different controller/slot/connection/bus/solution with these drives.
If the problem still persists, I would try changing the drives themselves.
@marcio79
I still have the issue of NVME SSDs disappearing after a few days of operation. This happens during intensive use such as backup.
I also updated the BIOS to the latest version including AGESA 1.2.0.C. I get the same errors.
I'm...
As discussed, I have updated the kernel to version 6.11.11-1-pve on each node.
So far, everything is working.
I will keep you informed after a few days if this kernel version no longer causes bugs with the NVMe SSDs.
It does indeed seem like the parameter nvme_core.default_ps_max_latency_us=0 is not being taken into account with the new version of grub, yet after booting we do have 0.
cat /sys/module/nvme_core/parameters/default_ps_max_latency_us
returns...
I agree, the bug appeared at some point in kernel version 6.8.12-x and it is still present in version 6.8.12-8.
I wonder if it is better to downgrade the kernel before 6.8.12-x, such as kernel-6.8.8-2, or upgrade to kernel 6.11.11-1 which is...
I am using arch Linux and i had drive disappearing after waking up from sleep. Previously i fixed it with settings `nvme_core.default_ps_max_latency_us=0` But after updating to latest kernel it started to occur again. it seems like...
I've had nothing but trouble with certain datacenter-class NVMe u.2 drives (drop out usually within 24 hours):
Intel P4600 6.4TB
Micron 9200 MAX 6.4TB
Both of them work great under ESXi, but randomly drop out in Proxmox. Kernel revs during the...
Hi Team,
At the end of 2023, several users reported issues with unexplained loss of access to NVMe SSDs, particularly Samsung 990 Pro NVMe SSDs, which I have. One or more NVMe SSDs suddenly disconnected and were no longer detected by Linux.
The...