How to change the number of cores (and memory) allocated to Ceph OSDs?

J-Rod

Active Member
Jun 29, 2026
108
65
28
I'm experimenting with Ceph compression and am happy with the reduction in space usage, but the OSDs seem to be CPU bottlenecked with what appears to be only two cores allocated per OSD. I can't seem to find the information in the PVE or Ceph documentation on changing this, and some online resources that mention specific commands in this area don't seem to apply to Ceph 20.

Also, I've seen it reported various times that PVE makes some tweaks to the default Ceph CPU & memory allocations from the Ceph default presumably to assist with resource utilization on hyper-converged installs. Where are these specific defaults documented?
 
I'm experimenting with Ceph compression and am happy with the reduction in space usage, but the OSDs seem to be CPU bottlenecked with what appears to be only two cores allocated per OSD. I can't seem to find the information in the PVE or Ceph documentation on changing this, and some online resources that mention specific commands in this area don't seem to apply to Ceph 20.

My guess is that you are not using KRBD, is that correct? KRBD uses the Kernel's page cache instead of caching things in user space, so you might want to give that a shot. You can enable it in your storage settings. A simple live-migration is enough to make a VM use KRBD once you've changed your config, IIRC.

Do note however that KRBD doesn't always have the newest features that librbd has, so if there's something in librbd you're relying on, you probably should not switch. If you don't know, you can try enabling it, live-migrating (or power-cycling) a VM for testing purposes, and see if everything still behaves as expected, including any operations you usually do. Just to be safe.

I'm also fairly certain that librbd is single-threaded by default. So that's most likely why there are only two cores per OSD, one of them is probably the IO thread.

If switching to KRBD isn't an option for you, you can also try to tune librbd's settings; setting rbd_read_from_replica_policy to localized in particular may yield to substantial results, depending on your CRUSH map.

Hope this helps! Please lemme know how it goes.

Also, I've seen it reported various times that PVE makes some tweaks to the default Ceph CPU & memory allocations from the Ceph default presumably to assist with resource utilization on hyper-converged installs. Where are these specific defaults documented?

That would honestly be new to me; I have not seen us patching any defaults in that regard. The configuration you see in Datacenter > Ceph > Configuration should tell you everything that's been set, the rest are stock settings from upstream directly.

The only thing I can think of that's being set automatically is [URL='https://docs.ceph.com/en/tentacle/rados/configuration/mclock-config-ref/#osd-capacity-determination-automated']osd_mclock_max_capacity_iops_ssd[/URL], but that's also from stock Ceph.
 
Thanks for the reply.
I did experiment with krbd, as well as partitioning the PCIe SSDs (Intel DC P3700) for multiple OSDs, and also disabled compression, among a host of other tweaks. I really never could exceed around 160MBytes/sec sustained writes per VM. It does seem that Ceph Crimson is designed to resolve some of the potential causes of this issue in my case.
Ultimately, I just wasn't happy with the overall extra resources and power consumed by the P3700s given the performance I was hoping for, so went back to SATA SSDs with ZFS+Replication for now. This was a home lab cluster, after all.