CEPH setup

PmUserZFS

Renowned Member
Feb 2, 2018
167
10
83
I have 6 nodes with the following drives

4x HP PM9A3 480GB(M.2 NVMe PCIe 4.0) slow due to 480gb, not fast as the u.2 and larger drives, benchmarks hard to find on the m.2 480gb.
5x PM1655 800GB(SAS 24 Gbps)
3x PM897(SATA 6 Gbps)

This will be a proxmox cluster with ceph, running in a lab, so often not that bussy, but we all hate slow performance ...
I will use FastEC not replica 3, sas and sata in the same pool.

any input/ideas how to optimize this hardware?
 
You can rise osd_memory_target but HP PM9A3 is more for read intensive. Try not to use them to see how will Ceph performance.
 
Are you using Ceph 20.2.x as EC got major improvement with 20.2.x ?

If you are mixing drives, expect the storage to be as "fast" as the slowest disk.

Maybe instead of mixing disks you could create different pools with different crush rules. It really depends on your needs.

In general, expect replica 3 to be the fastest.
 
Ceph is RAM hungry, not really great at high ram cost right now. As we are lucky and have enough ram(still) we are using 2gb per 1tb of NVMe storage.

The default setting in Proxmox is 2GB(flat) and that is waaaay too low for large NVMe disks. Dont even try to use 2gb for large NVMe disks, we have tried and it is not good.
 
  • Like
Reactions: Johannes S
Are you using Ceph 20.2.x as EC got major improvement with 20.2.x ?

If you are mixing drives, expect the storage to be as "fast" as the slowest disk.

Maybe instead of mixing disks you could create different pools with different crush rules. It really depends on your needs.

In general, expect replica 3 to be the fastest.
I will use 20.2.3 yes, with fast EC, but its still slower, but I hope not too noticable, only in benchmarks.
I will use affinity, maybe 0.5 on the sata drive to offset the somewhat slowness, they will get much less iops at least,

Estimated pm9a3 numbers, since there are no benchmarks I have found on them, they are basically boot drives so ..

1787232271168.png
 
Last edited:
Ceph is RAM hungry, not really great at high ram cost right now. As we are lucky and have enough ram(still) we are using 2gb per 1tb of NVMe storage.

The default setting in Proxmox is 2GB(flat) and that is waaaay too low for large NVMe disks. Dont even try to use 2gb for large NVMe disks, we have tried and it is not good.
osd_memory_target is 4GB by default? no? I will increase it to 8GB per OSD so 64GB will be rad cache per host, when I have build the cluster.

For now Im gonna setup one host and bench the drives, im curious about he pm9a3, how bad is it...
 
Last edited:
fio benchmark here:

fio --name=test --filename=/mnt/$mnt/testfile --size=2G --rw=randwrite --bs=4k --iodepth=$qd --ioengine=libaio --direct=1 --time_based --runtime=10 --group_reporting

1787248480334.png


Hm really dont know what to do with these drives now-- hm

separate ceph pools? dbwal on the nvme ?


NAME SIZE MODEL ROTA
sda 745.2G MO000800PZWSF 0
sdb 745.2G MO000800PZWSF 0
sdc 745.2G MO000800PZWSF 0
sdd 745.2G MO000800PZWSF 0
sde 745.2G MO000800PZWSF 0
sdf 894.3G MK000960GZXRB 0
sdg 894.3G MK000960GZXRB 0
sdh 894.3G MK000960GZXRB 0
sdi 14.5G STORE N GO 1
nvme0n1 447.1G VR000480KXLXF 0
nvme3n1 447.1G VR000480KXLXF 0
nvme2n1 447.1G VR000480KXLXF 0
nvme1n1 447.1G VR000480KXLXF 0

hm ...
 
You told it is slow but did not told the numbers.

With ceph you can create CRUSH rules where pool must save data. You can create new class, new root to distribute the data.
 
You told it is slow but did not told the numbers.

With ceph you can create CRUSH rules where pool must save data. You can create new class, new root to distribute the data.
Above is the write numbers which is most important for db/wal , since I couldnt find any review I had to guess prior. But as you can see its decent and good enough.

I currentl leaning on a SAS vm pool with replica and db/wal on the nvmes and thw sata drives with FastEC 3+2, and maybe also db/wal on 30ish GB on the nvmes
 
SSD as DB/WAL for SSD may not give expected performance. What was Ceph slow performance at the beginning ( not disk itself )?

As of small/lab Ceph cluster I do not see much lose of using EC vs replica.
 
osd_memory_target is 4GB by default? no? I will increase it to 8GB per OSD so 64GB will be rad cache per host, when I have build the cluster.

For now Im gonna setup one host and bench the drives, im curious about he pm9a3, how bad is it...
Maybe you are right, it is 4, but still way too low for large NVMe disk (7.68 or 15TB). Fine for smaller drives like 1 to 4TB I would say.
 
Old cluster (4K write)New vm-hot (4K write)Improvement
Avg latency2.49ms1.52ms~39% lower
Avg IOPS6,41510,504~64% higher
Max latency (tail)65.7ms7.55ms~8.7x lower
 
# ceph osd tree
ID CLASS WEIGHT TYPE NAME STATUS REWEIGHT PRI-AFF
-1 18.14612 root default
-3 3.02435 host pm1
0 sas 0.75609 osd.0 up 1.00000 1.00000
1 sas 0.75609 osd.1 up 1.00000 1.00000
2 sas 0.75609 osd.2 up 1.00000 1.00000
3 sas 0.75609 osd.3 up 1.00000 1.00000
-7 3.02435 host pm2
4 sas 0.75609 osd.4 up 1.00000 1.00000
5 sas 0.75609 osd.5 up 1.00000 1.00000
6 sas 0.75609 osd.6 up 1.00000 1.00000
7 sas 0.75609 osd.7 up 1.00000 1.00000
-10 3.02435 host pm3
8 sas 0.75609 osd.8 up 1.00000 1.00000
9 sas 0.75609 osd.9 up 1.00000 1.00000
10 sas 0.75609 osd.10 up 1.00000 1.00000
11 sas 0.75609 osd.11 up 1.00000 1.00000
-13 3.02435 host pm4
12 sas 0.75609 osd.12 up 1.00000 1.00000
13 sas 0.75609 osd.13 up 1.00000 1.00000
14 sas 0.75609 osd.14 up 1.00000 1.00000
15 sas 0.75609 osd.15 up 1.00000 1.00000
-16 3.02435 host pm5
16 sas 0.75609 osd.16 up 1.00000 1.00000
17 sas 0.75609 osd.17 up 1.00000 1.00000
18 sas 0.75609 osd.18 up 1.00000 1.00000
19 sas 0.75609 osd.19 up 1.00000 1.00000
-19 3.02435 host pm6
20 sas 0.75609 osd.20 up 1.00000 1.00000
21 sas 0.75609 osd.21 up 1.00000 1.00000
22 sas 0.75609 osd.22 up 1.00000 1.00000
23 sas 0.75609 osd.23 up 1.00000 1.00000

~# lsblk
NAME MAJ:MIN RM SIZE RO TYPE MOUNTPOINTS
sda 8:0 0 745.2G 0 disk
└─ceph--e2a0c09b--6dd9--4f3a--b73e--8cf188af0fb9-osd--block--ccc3da5f--2c37--490d--b1d3--16bb7dac74b2 252:1 0 745.2G 0 lvm
sdb 8:16 0 745.2G 0 disk
└─ceph--adf8e20d--8914--4607--a4a4--956a1bb48f7c-osd--block--0202c53c--1458--422d--937a--6c15fdac6c3c 252:3 0 745.2G 0 lvm
sdc 8:32 0 745.2G 0 disk
└─ceph--9a724381--4ea6--43e4--a253--3e3b3307ccdc-osd--block--cb0053b3--cc41--4395--abf6--effff4185ed5 252:5 0 745.2G 0 lvm
sdd 8:48 0 745.2G 0 disk
└─ceph--33a4878e--9e19--4538--9afe--61e8b941acf8-osd--block--3be28483--8d63--401c--8412--0c04314a0ec0 252:7 0 745.2G 0 lvm
sde 8:64 0 894.3G 0 disk
sdf 8:80 0 894.3G 0 disk
├─sdf1 8:81 0 2M 0 part
└─sdf2 8:82 0 894.3G 0 part
sdg 8:96 0 894.3G 0 disk
├─sdg1 8:97 0 2M 0 part
└─sdg2 8:98 0 894.3G 0 part
sdh 8:112 0 894.3G 0 disk
├─sdh1 8:113 0 2M 0 part
└─sdh2 8:114 0 894.3G 0 part
nvme3n1 259:0 0 447.1G 0 disk
├─nvme3n1p1 259:4 0 1007K 0 part
├─nvme3n1p2 259:5 0 1G 0 part
├─nvme3n1p3 259:6 0 63G 0 part
├─nvme3n1p4 259:25 0 30G 0 part
│ └─ceph--99a058a7--66cf--4672--ac34--144e8270be68-osd--db--17c96d6a--69ce--4afc--8739--3b53cd6822be 252:6 0 29G 0 lvm
├─nvme3n1p5 259:26 0 36G 0 part
└─nvme3n1p6 259:27 0 225G 0 part
nvme0n1 259:1 0 447.1G 0 disk
├─nvme0n1p1 259:7 0 1007K 0 part
├─nvme0n1p2 259:8 0 1G 0 part
├─nvme0n1p3 259:9 0 63G 0 part
├─nvme0n1p4 259:16 0 30G 0 part
│ └─ceph--d82fa4fa--9088--41fd--a0bd--4fae646d9883-osd--db--f27519e0--1110--4409--a33f--f7385c7c4c27 252:0 0 29G 0 lvm
├─nvme0n1p5 259:17 0 36G 0 part
└─nvme0n1p6 259:18 0 225G 0 part
nvme1n1 259:2 0 447.1G 0 disk
├─nvme1n1p1 259:10 0 1007K 0 part
├─nvme1n1p2 259:11 0 1G 0 part
├─nvme1n1p3 259:12 0 63G 0 part
├─nvme1n1p4 259:19 0 30G 0 part
│ └─ceph--091f2dda--b801--4627--9b5e--944607592fc1-osd--db--e7df45ee--1548--4a82--9167--ac55129a0721 252:2 0 29G 0 lvm
├─nvme1n1p5 259:20 0 36G 0 part
└─nvme1n1p6 259:21 0 225G 0 part
nvme2n1 259:3 0 447.1G 0 disk
├─nvme2n1p1 259:13 0 1007K 0 part
├─nvme2n1p2 259:14 0 1G 0 part
├─nvme2n1p3 259:15 0 63G 0 part
├─nvme2n1p4 259:22 0 30G 0 part
│ └─ceph--d6fc7298--f3d3--4e94--a963--47b76bbd6c14-osd--db--6aa95997--edcf--48e4--bb97--2cec19389d25 252:4 0 29G 0 lvm
├─nvme2n1p5 259:23 0 36G 0 part
└─nvme2n1p6 259:24 0 225G 0 part
 
# ceph osd tree
ID CLASS WEIGHT TYPE NAME STATUS REWEIGHT PRI-AFF
-1 18.14612 root default
-3 3.02435 host pm1
0 sas 0.75609 osd.0 up 1.00000 1.00000
1 sas 0.75609 osd.1 up 1.00000 1.00000
2 sas 0.75609 osd.2 up 1.00000 1.00000
3 sas 0.75609 osd.3 up 1.00000 1.00000
-7 3.02435 host pm2
4 sas 0.75609 osd.4 up 1.00000 1.00000
5 sas 0.75609 osd.5 up 1.00000 1.00000
6 sas 0.75609 osd.6 up 1.00000 1.00000
7 sas 0.75609 osd.7 up 1.00000 1.00000
-10 3.02435 host pm3
8 sas 0.75609 osd.8 up 1.00000 1.00000
9 sas 0.75609 osd.9 up 1.00000 1.00000
10 sas 0.75609 osd.10 up 1.00000 1.00000
11 sas 0.75609 osd.11 up 1.00000 1.00000
-13 3.02435 host pm4
12 sas 0.75609 osd.12 up 1.00000 1.00000
13 sas 0.75609 osd.13 up 1.00000 1.00000
14 sas 0.75609 osd.14 up 1.00000 1.00000
15 sas 0.75609 osd.15 up 1.00000 1.00000
-16 3.02435 host pm5
16 sas 0.75609 osd.16 up 1.00000 1.00000
17 sas 0.75609 osd.17 up 1.00000 1.00000
18 sas 0.75609 osd.18 up 1.00000 1.00000
19 sas 0.75609 osd.19 up 1.00000 1.00000
-19 3.02435 host pm6
20 sas 0.75609 osd.20 up 1.00000 1.00000
21 sas 0.75609 osd.21 up 1.00000 1.00000
22 sas 0.75609 osd.22 up 1.00000 1.00000
23 sas 0.75609 osd.23 up 1.00000 1.00000
 
Im not really impressed with the performance of ceph,

The above benchmark with this decent hardware ought to yield much better result, ie reference storpool.
 
Skimming through this thread i realized you never wrote any information about your networking.
Do you use at least 2x 100Gbit MLAG for ceph storage and public network?
If not then your ceph cluster is most likely limited by your networking.
 
  • Like
Reactions: uzumo
Skimming through this thread i realized you never wrote any information about your networking.
Do you use at least 2x 100Gbit MLAG for ceph storage and public network?
If not then your ceph cluster is most likely limited by your networking.

Doesnt affect latency and QD1 heavy workloads.

Running 2x25GB for both public and cluster.
Connectx5 for public optimized for latency
Intel XXV710 for cluster

tso on gso on gro on rx-checksum on tx-checksum-ip-generic on
adaptive-rx off adaptive-tx off rx-usecs 8
rx 4096 tx 4096
NUMA IRQ pinning
removed c states
etc
 
rados bench -p vm-hot 10 write -b 4096 -t 64 --no-cleanup
rados -p vm-hot cleanup

rados bench -p bulk-cold-data 10 write -b 4096 -t 64 --no-cleanup
rados -p bulk-cold-data cleanup

Results at t=64​


vm-hot (SAS, replica-3)bulk-cold-data (SATA, Fast EC 3+2)
Avg latency1.94ms1.15ms (~1.7x lower)
Avg IOPS32,98455,779 (~1.7x higher)
Bandwidth128.8 MB/s217.9 MB/s (~1.7x higher)
Max latency15.99ms19.02ms (only outlier where vm-hot wins)