Hi everyone
Setup: In a 4-node Proxmox cluster, shared storage via iSCSI is connected to 3 of the 4 nodes simultaneously. The fourth node does not have the shared LUNs and acts purely as a compute node, relying solely on Ceph for storage.
In this setup, we simulated an HA event by hard-powering off one node.
The cluster restarted all the VMs. However, one VM—which had its disks on the shared LUNs—was registered on the compute node that lacked access to those LUNs. Consequently, it could not be started. Getting this VM back into a proper state was quite a hassle.
Is this a bug, given that registration occurred on a node that did not meet the necessary requirements?
Or is there an implicit requirement that all cluster members *must* have the identical storage configuration, simply because no validation check is performed?
Greetings
clabo
Setup: In a 4-node Proxmox cluster, shared storage via iSCSI is connected to 3 of the 4 nodes simultaneously. The fourth node does not have the shared LUNs and acts purely as a compute node, relying solely on Ceph for storage.
In this setup, we simulated an HA event by hard-powering off one node.
The cluster restarted all the VMs. However, one VM—which had its disks on the shared LUNs—was registered on the compute node that lacked access to those LUNs. Consequently, it could not be started. Getting this VM back into a proper state was quite a hassle.
Is this a bug, given that registration occurred on a node that did not meet the necessary requirements?
Or is there an implicit requirement that all cluster members *must* have the identical storage configuration, simply because no validation check is performed?
Greetings
clabo