Ceph Cipers Migration - Windows TPM States "can't map" - Cannot Start / Migrate Windows Guests

MN7675

Renowned Member
Dec 22, 2015
121
8
83
41
CEPHX Disabled.

pveversion -v

Code:
proxmox-ve: 9.2.0 (running kernel: 7.0.14-16-pve)
pve-manager: 9.2.18 (running version: 9.2.18/614bede5d65599c6)
proxmox-kernel-helper: 9.2.0
proxmox-kernel-7.0: 7.0.14-16
proxmox-kernel-7.0.14-16-pve-signed: 7.0.14-16
proxmox-kernel-7.0.14-6-pve-signed: 7.0.14-6
proxmox-kernel-7.0.2-6-pve-signed: 7.0.2-6
ceph: 20.2.4-pve4
ceph-fuse: 20.2.4-pve4
corosync: 3.1.10-pve3
criu: 4.1.1-1
frr-pythontools: 10.6.1-1+pve3
ifupdown2: 3.3.0-1+pmx12
intel-microcode: 3.20251111.1~deb13u1
ksm-control-daemon: 1.5-1
libjs-extjs: 7.0.0-7
libproxmox-acme-perl: 1.7.2
libproxmox-backup-qemu0: 2.0.2
libproxmox-rs-perl: 0.4.1
libpve-access-control: 9.1.1
libpve-apiclient-perl: 3.4.3
libpve-cluster-api-perl: 9.1.6
libpve-cluster-perl: 9.1.6
libpve-common-perl: 9.2.1
libpve-guest-common-perl: 6.0.5
libpve-http-server-perl: 6.0.5
libpve-network-perl: 1.6.7
libpve-notify-perl: 9.1.6
libpve-rs-perl: 0.15.3
libpve-storage-perl: 9.1.10
libspice-server1: 0.15.2-1+b1
lvm2: 2.03.31-2+pmx1
lxc-pve: 7.0.0-2
lxcfs: 7.0.0-pve1
novnc-pve: 1.7.0-2
proxmox-backup-client: 4.2.5-1
proxmox-backup-file-restore: 4.2.5-1
proxmox-backup-restore-image: 1.0.0
proxmox-enterprise-support-keyring: 1.1
proxmox-firewall: 1.2.3
proxmox-kernel-helper: 9.2.0
proxmox-mail-forward: 1.0.3
proxmox-mini-journalreader: 1.7
proxmox-widget-toolkit: 5.2.8
pve-cluster: 9.1.6
pve-container: 6.1.14
pve-docs: 9.2.10
pve-edk2-firmware: 4.2026.08-1
pve-esxi-import-tools: 1.0.1
pve-firewall: 6.0.5
pve-firmware: 3.18-6
pve-ha-manager: 5.2.5
pve-i18n: 3.10.0
pve-qemu-kvm: 11.0.3-3
pve-xtermjs: 6.0.0-2
qemu-server: 9.2.7
smartmontools: 7.5-pve2
spiceterm: 3.4.2
swtpm: 0.8.0+pve3
vncterm: 1.9.2
zfsutils-linux: 2.4.4-pve1

ceph.conf

Code:
[global]
        auth_client_required = none
        auth_cluster_required = none
        auth_service_required = none
        auth_supported = none
        cluster_network = 10.10.1.0/24
        fsid = 54da8900-a9db-4a57-923c-a62dbec8c82a
        mon_allow_pool_delete = true
        mon_host = 10.10.1.18 10.10.1.19 10.10.1.21
        ms_bind_ipv4 = true
        osd_journal_size = 5120
        osd_pool_default_min_size = 2
        osd_pool_default_size = 3
        public_network = 10.10.1.0/24

[client.crash]
        keyring = /etc/pve/ceph/$cluster.$name.keyring

[mds]
        debug_mds = 1
        debug_mds_balancer = 1
        keyring = /var/lib/ceph/mds/ceph-$id/keyring

[mds.vmhost6]
        host = vmhost6
        mds_standby_for_name = pve

[mds.vmhost7]
        host = vmhost7
        mds_standby_for_name = pve

[mds.vmhost8]
        host = vmhost8
        mds_standby_for_name = pve

[mon.vmhost6]
        public_addr = 10.10.1.18

[mon.vmhost7]
        public_addr = 10.10.1.19

[mon.vmhost8]
        public_addr = 10.10.1.21

[mgr]
        mgr_modules = zabbix


I figured I should start a new post since it's a new issue, but it is also a continuation of this other thread.


After migrating the ciphers, I am now unable to start any of the Windows guests with the following errors:

Code:
task started by HA resource agent
In some cases useful info is found in syslog - try "dmesg | tail".
rbd: sysfs write failed
TASK ERROR: start failed: can't map rbd volume vm-111-disk-2: rbd: sysfs write failed

running dmesg | tail shows the following:

Code:
[1719011.576938] libceph: auth protocol 'cephx' not allowed
[1719022.102919] libceph: auth protocol 'cephx' not allowed
[1719051.749248] libceph: auth protocol 'cephx' not allowed
[1719062.289476] libceph: auth protocol 'cephx' not allowed
[1719271.804244] libceph: auth protocol 'cephx' not allowed

Deleting and re-creating the TPM State disk does not resolve this issue.

P.S. I am not having any issues with Linux guests.
 
Last edited:
I temporarily enabled local storage on the hosts and moved the TPM State disks to that storage in .raw format.

I was able to successfully start up the guests that way and will run them in that manner until I can find a solution.

The problem still persists though.