I'm experimenting with Ceph compression and am happy with the reduction in space usage, but the OSDs seem to be CPU bottlenecked with what appears to be only two cores allocated per OSD. I can't seem to find the information in the PVE or Ceph documentation on changing this, and some online resources that mention specific commands in this area don't seem to apply to Ceph 20.
My guess is that you are not using KRBD, is that correct? KRBD uses the Kernel's page cache instead of caching things in user space, so you might want to give that a shot. You can enable it in your storage settings. A simple live-migration is enough to make a VM use KRBD once you've changed your config, IIRC.
Do note however that KRBD doesn't always have the newest features that librbd has, so if there's something in librbd you're relying on, you probably should not switch. If you don't know, you can try enabling it, live-migrating (or power-cycling) a VM for testing purposes, and see if everything still behaves as expected, including any operations you usually do. Just to be safe.
I'm also fairly certain that librbd is single-threaded by default. So that's most likely why there are only two cores per OSD, one of them is probably the IO thread.
If switching to KRBD isn't an option for you, you can also try to
tune librbd's settings; setting
rbd_read_from_replica_policy to localized in particular may yield to substantial results, depending on your CRUSH map.
Hope this helps! Please lemme know how it goes.
Also, I've seen it reported various times that PVE makes some tweaks to the default Ceph CPU & memory allocations from the Ceph default presumably to assist with resource utilization on hyper-converged installs. Where are these specific defaults documented?
That would honestly be new to me; I have not seen us patching any defaults in that regard. The configuration you see in
Datacenter > Ceph > Configuration should tell you everything that's been set, the rest are stock settings from upstream directly.
The
only thing I can think of that's being set automatically is
[URL='https://docs.ceph.com/en/tentacle/rados/configuration/mclock-config-ref/#osd-capacity-determination-automated']osd_mclock_max_capacity_iops_ssd[/URL], but that's also from stock Ceph.