failure of a network card in a ring cluster

silvered.dragon

Renowned Member
Nov 4, 2015
123
4
83
Good morning my friends,
let's say I have a 3 node proxmox/ceph cluster in a full mesh network with a dedicated dual port 10Gb in bond broadcast for ceph and other dedicated 1Gb cards for cluster and lan as showed in this tutorial.
https://pve.proxmox.com/wiki/Full_Mesh_Network_for_Ceph_Server

I have a question about this.
what happens if for example I unplug one 10Gb cable from the ring during the cluster acticity? for example the one that connects the first node with the fird node?
I'm not sure if in this case the 2 nodes goes down cause the bond is broken or maybe only one.
I can't test this in my office now cause I'm in a production environment.
 
If only one link is down then the cluster will have three possible truths (split brain). https://en.wikipedia.org/wiki/Split-brain_(computing) With that said, it is very unlikely that a cable breaks when left alone. It will be more common that one of the hosts is down or its NIC broke.
 
yes you are right is very difficoult but maybe it can happens accidentaly during a maintenance procedure, when you have so many servers and cables in a rack things like this can happens.. but now that I know I wil be more carefull. but if a similar split brain happens which is the best procedure to be implemented??
 
but if a similar split brain happens which is the best procedure to be implemented??
What do you mean? Check cables and plug them back in or shutdown one of the hosts.
 
I mean that if a split brain happens I will probably have some data inconsistency, do you think that simply replacing the cable or shutting down one node it's enough? How is split brain situations handled in ceph after restoring the hardware failure?
 
Ceph will stop to work in such a situation and you will need to shutdown a node so ceph gets back and starts recovering. You can build a virtual cluster and test that and other behaviors.
 
thank you, yes I really want to try but how I can create a ring in virtualbox? i think it's impossible because an internal virtual network is like a virtual switch I cannot decide where to put the cables.
 
I don't know about virtualbox. I thought about using Proxmox with two linux bridges that are not connected with any physical links.
 
thank you my friend for your suggestions. now everything is clear, you are right in a split brain situation ceph will stop to work and you cannot write anything, after systems return to normality everithing is working fine with no data loss only a little downtime.
thank you again
 

About

The Proxmox community has been around for many years and offers help and support for Proxmox VE, Proxmox Backup Server, and Proxmox Mail Gateway.
We think our community is one of the best thanks to people like you!

Get your subscription!

The Proxmox team works very hard to make sure you are running the best software and getting stable updates and security enhancements, as well as quick enterprise support. Tens of thousands of happy customers have a Proxmox subscription. Get yours easily in our online shop.

Buy now!