[SOLVED] Root on encrypted ZFS boot issue suddenly

Mar 23, 2023
11
2
8
Boise, ID
I rebooted one of my nodes to get a new kernel, and on startup I was dropped into initramfs busybox. It is similar to this: https://pve.proxmox.com/wiki/ZFS:_Tips_and_Tricks#Boot_fails_and_goes_into_busybox

I am running (unsupported, I know) an encrypted ZFS rpool:
Code:
root@ThreadReaper:~# zfs get encryption,encryptionroot,keylocation,keyformat,keystatus rpool/ROOT rpool/ROOT/pve-1
NAME              PROPERTY        VALUE        SOURCE
rpool/ROOT        encryption      aes-256-gcm  -
rpool/ROOT        encryptionroot  rpool/ROOT   -
rpool/ROOT        keylocation     prompt       local
rpool/ROOT        keyformat       passphrase   -
rpool/ROOT        keystatus       available    -
rpool/ROOT/pve-1  encryption      aes-256-gcm  -
rpool/ROOT/pve-1  encryptionroot  rpool/ROOT   -
rpool/ROOT/pve-1  keylocation     none         default
rpool/ROOT/pve-1  keyformat       passphrase   -
rpool/ROOT/pve-1  keystatus       available    -

Up until this reboot, so somewhere between mid-July and today, it would simply prompt me for the passphrase Enter passphrase for 'rpool/ROOT': then continue booting. But now I get the (initramfs) prompt.

I was able to figure out how to get the system booted:

Bash:
zpool import -l -R /root rpool
# Key load error: Failed to open key material file: No such file or directory
# Enter passphrase for 'rpool/ROOT': ******
# 1 / 2 keys successfully loaded
exit
I believe the key failure is for rpool/data, which does automatically unlock just fine later in the boot with a keyfile stored in rpool/ROOT/pve-1. I could probably just mount rpool/ROOT/pve-1 and avoid the error message, but I don't want to reboot again to find out at 00:26 in the morning.

While debugging, I tried the following with no change in this new behavior:
  • Add rootdelay=30 to the kernel command line. This just had it sit there for ~30 seconds before dropping me to the busybox shell.
  • Boot the previous working kernel (proxmox-kernel-7.0.14-4-pve-signed). This did the same strange new behavior, so not a kernel issue.
  • Install an even fresher kernel (proxmox-kernel-7.0.14-17-pve-signed). Same behavior.
I have not rebooted the other nodes yet, so I don't know if it is just this one server or all of them. Since I have to manually interact with the servers to boot them no matter what, this is not a huge problem, but it is way less convenient to have to remember the commands above to get my server to boot. Did something change that broke the previously-working-but-unsupported encrypted ZFS root? Or did something in my setup get corrupted?
 
One strange thing I noticed is in /boot/grub/grub.cfg, the kernel argument is root=ZFS=/ROOT/pve-1 instead of root=ZFS=rpool/ROOT/pve-1:
Code:
menuentry 'Proxmox VE GNU/Linux' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_option 'gnulinux-simple-/dev/nvme4n1p3_/dev/nvme5n1p3' {
        load_video
        insmod gzio
        if [ x$grub_platform = xxen ]; then insmod xzio; insmod lzopio; fi
        insmod part_gpt
        insmod part_gpt
        echo    'Loading Linux 7.0.14-17-pve ...'
        linux   /ROOT/pve-1@/boot/vmlinuz-7.0.14-17-pve root=ZFS=/ROOT/pve-1 ro
        echo    'Loading initial ramdisk ...'
        initrd  /ROOT/pve-1@/boot/initrd.img-7.0.14-17-pve
}

It seems to have dropped the rpool out of the command line. Interestingly, rpool is not missing in the same file on a different node, implying something is wrong with how grub.cfg is generated on this particular node. It actually looks like it is missing all the command line arguments. This is how the file looks on a different node:
Code:
menuentry 'Proxmox VE GNU/Linux' --class proxmox --class gnu-linux --class gnu --class os $menuentry_id_option 'gnulinux-simple-/dev/sda3_/dev/sdb3' {
        load_video
        insmod gzio
        if [ x$grub_platform = xxen ]; then insmod xzio; insmod lzopio; fi
        insmod part_gpt
        insmod part_gpt
        echo    'Loading Linux 7.0.14-19-pve ...'
        linux   /ROOT/pve-1@/boot/vmlinuz-7.0.14-19-pve root=ZFS=/ROOT/pve-1 ro  root=ZFS=rpool/ROOT/pve-1 boot=zfs quiet
        echo    'Loading initial ramdisk ...'
        initrd  /ROOT/pve-1@/boot/initrd.img-7.0.14-19-pve
}
 
Last edited:
could you post "pveversion -v" and the contents of /var/log/apt/term.log referencing the last refresh of the grub config (from both nodes)?
 
Last edited:
  • Like
Reactions: Stoiko Ivanov
ThreadReaper (node with issue)
Code:
root@ThreadReaper:~# pveversion -v
proxmox-ve: 9.2.0 (running kernel: 7.0.14-17-pve)
pve-manager: 9.2.20 (running version: 9.2.20/49318c671b82f31e)
proxmox-kernel-helper: 9.2.0
proxmox-kernel-7.0.14-17-pve-signed: 7.0.14-17
proxmox-kernel-7.0: 7.0.14-17
proxmox-kernel-7.0.14-16-pve-signed: 7.0.14-16
proxmox-kernel-7.0.14-4-pve-signed: 7.0.14-4
proxmox-kernel-6.17: 6.17.13-21
proxmox-kernel-6.17.13-21-pve-signed: 6.17.13-21
amd64-microcode: 3.20251202.1~bpo13+1
ceph: 19.2.6-pve4
ceph-fuse: 19.2.6-pve4
corosync: 3.1.10-pve3
criu: 4.1.1-1
frr-pythontools: 10.6.1-1+pve3
ifupdown2: 3.3.0-1+pmx12
ksm-control-daemon: 1.5-1
libjs-extjs: 7.0.0-7
libproxmox-acme-perl: 1.7.2
libproxmox-backup-qemu0: 2.0.2
libproxmox-rs-perl: 0.4.1
libpve-access-control: 9.1.2
libpve-apiclient-perl: 3.4.3
libpve-cluster-api-perl: 9.1.6
libpve-cluster-perl: 9.1.6
libpve-common-perl: 9.2.2
libpve-guest-common-perl: 6.0.5
libpve-http-server-perl: 6.0.5
libpve-network-perl: 1.6.7
libpve-notify-perl: 9.1.6
libpve-rs-perl: 0.15.3
libpve-storage-perl: 9.1.10
libspice-server1: 0.15.2-1+b1
lvm2: 2.03.31-2+pmx1
lxc-pve: 7.0.0-2
lxcfs: 7.0.0-pve1
novnc-pve: 1.7.0-2
proxmox-backup-client: 4.2.5-1
proxmox-backup-file-restore: 4.2.5-1
proxmox-backup-restore-image: 1.0.0
proxmox-enterprise-support-keyring: 1.1
proxmox-firewall: 1.2.3
proxmox-kernel-helper: 9.2.0
proxmox-mail-forward: 1.0.3
proxmox-mini-journalreader: 1.7
proxmox-offline-mirror-helper: 0.7.4
proxmox-widget-toolkit: 5.2.10
pve-cluster: 9.1.6
pve-container: 6.1.14
pve-docs: 9.2.12
pve-edk2-firmware: 4.2026.08-1
pve-esxi-import-tools: 1.0.1
pve-firewall: 6.0.6
pve-firmware: 3.18-6
pve-ha-manager: 5.2.5
pve-i18n: 3.10.0
pve-qemu-kvm: 11.0.3-3
pve-xtermjs: 6.0.0-2
qemu-server: 9.2.8
smartmontools: 7.5-pve2
spiceterm: 3.4.2
swtpm: 0.8.0+pve3
vncterm: 1.9.2
zfsutils-linux: 2.4.4-pve1

Wrex (alternate node):
Code:
root@Wrex:~# pveversion -v
proxmox-ve: 9.2.0 (running kernel: 7.0.2-6-pve)
pve-manager: 9.2.20 (running version: 9.2.20/49318c671b82f31e)
proxmox-kernel-helper: 9.2.0
proxmox-kernel-7.0.14-19-pve-signed: 7.0.14-19
proxmox-kernel-7.0: 7.0.14-19
proxmox-kernel-7.0.14-12-pve-signed: 7.0.14-12
proxmox-kernel-7.0.14-8-pve-signed: 7.0.14-8
proxmox-kernel-7.0.2-6-pve-signed: 7.0.2-6
proxmox-kernel-6.17: 6.17.13-21
proxmox-kernel-6.17.13-21-pve-signed: 6.17.13-21
ceph: 19.2.6-pve4
ceph-fuse: 19.2.6-pve4
corosync: 3.1.10-pve3
criu: 4.1.1-1
frr-pythontools: 10.6.1-1+pve3
ifupdown2: 3.3.0-1+pmx12
intel-microcode: 3.20251111.1~deb13u1
ksm-control-daemon: 1.5-1
libjs-extjs: 7.0.0-7
libproxmox-acme-perl: 1.7.2
libproxmox-backup-qemu0: 2.0.3
libproxmox-rs-perl: 0.4.1
libpve-access-control: 9.1.2
libpve-apiclient-perl: 3.4.3
libpve-cluster-api-perl: 9.1.6
libpve-cluster-perl: 9.1.6
libpve-common-perl: 9.2.2
libpve-guest-common-perl: 6.0.5
libpve-http-server-perl: 6.0.5
libpve-network-perl: 1.6.7
libpve-notify-perl: 9.1.6
libpve-rs-perl: 0.15.3
libpve-storage-perl: 9.1.10
libspice-server1: 0.15.2-1+b1
lvm2: 2.03.31-2+pmx1
lxc-pve: 7.0.0-2
lxcfs: 7.0.0-pve1
novnc-pve: 1.7.0-2
proxmox-backup-client: 4.2.6-1
proxmox-backup-file-restore: 4.2.6-1
proxmox-backup-restore-image: 1.0.0
proxmox-enterprise-support-keyring: 1.1
proxmox-firewall: 1.2.3
proxmox-kernel-helper: 9.2.0
proxmox-mail-forward: 1.0.3
proxmox-mini-journalreader: 1.7
proxmox-offline-mirror-helper: 0.7.4
proxmox-widget-toolkit: 5.2.10
pve-cluster: 9.1.6
pve-container: 6.1.14
pve-docs: 9.2.12
pve-edk2-firmware: 4.2026.08-1
pve-esxi-import-tools: 1.0.1
pve-firewall: 6.0.6
pve-firmware: 3.18-6
pve-ha-manager: 5.2.5
pve-i18n: 3.10.0
pve-qemu-kvm: 11.0.3-3
pve-xtermjs: 6.0.0-2
qemu-server: 9.2.8
smartmontools: 7.5-pve2
spiceterm: 3.4.2
swtpm: 0.8.0+pve3
vncterm: 1.9.2
zfsutils-linux: 2.4.4-pve1
 

Attachments

could you also post (again for both nodes)

- proxmox-boot-tool status
- cat /etc/kernel/cmdline
- cat /etc/default/grub

thanks!
 
Well that is interesting:
Code:
root@ThreadReaper:~# proxmox-boot-tool status
Re-executing '/usr/sbin/proxmox-boot-tool' in new private mount namespace..
System currently booted with uefi
4F94-D7E3 is configured with: uefi (versions: BOOTX64.CSV, fbx64.efi, grub.cfg, grubx64.efi, mmx64.efi, shimx64.efi), grub (versions: 6.17.13-21-pve, 7.0.14-16-pve, 7.0.14-17-pve)
4F94-F13B is configured with: uefi (versions: BOOTX64.CSV, fbx64.efi, grub.cfg, grubx64.efi, mmx64.efi, shimx64.efi), grub (versions: 6.17.13-21-pve, 7.0.14-16-pve, 7.0.14-17-pve)
root@ThreadReaper:~# cat /etc/kernel/cmdline
root=ZFS=rpool/ROOT/pve-1 boot=zfs mem_encrypt=on kvm_amd.sev=1
root@ThreadReaper:~# cat /etc/default/grub
cat: /etc/default/grub: No such file or directory

Code:
root@Wrex:~# proxmox-boot-tool status
Re-executing '/usr/sbin/proxmox-boot-tool' in new private mount namespace..
System currently booted with uefi
3A35-D50F is configured with: uefi (versions: 6.17.13-21-pve, 7.0.14-12-pve, 7.0.14-19-pve, 7.0.2-6-pve)
3A36-6239 is configured with: uefi (versions: 6.17.13-21-pve, 7.0.14-12-pve, 7.0.14-19-pve, 7.0.2-6-pve)
root@Wrex:~# cat /etc/kernel/cmdline
root=ZFS=rpool/ROOT/pve-1 boot=zfs intel_iommu=on iommu=pt mitigations=auto,nosmt
root@Wrex:~# cat /etc/default/grub
# If you change this file or any /etc/default/grub.d/*.cfg file,
# run 'update-grub' afterwards to update /boot/grub/grub.cfg.
# For full documentation of the options in these files, see:
#   info -f grub -n 'Simple configuration'

GRUB_DEFAULT=0
GRUB_TIMEOUT=5
GRUB_DISTRIBUTOR=`( . /etc/os-release && echo ${NAME} )`
GRUB_CMDLINE_LINUX_DEFAULT="quiet"
GRUB_CMDLINE_LINUX=""

# If your computer has multiple operating systems installed, then you
# probably want to run os-prober. However, if your computer is a host
# for guest OSes installed via LVM or raw disk devices, running
# os-prober can cause damage to those guest OSes as it mounts
# filesystems to look for things.
#GRUB_DISABLE_OS_PROBER=false

# Uncomment to enable BadRAM filtering, modify to suit your needs
# This works with Linux (no patch required) and with any kernel that obtains
# the memory map information from GRUB (GNU Mach, kernel of FreeBSD ...)
#GRUB_BADRAM="0x01234567,0xfefefefe,0x89abcdef,0xefefefef"

# Uncomment to disable graphical terminal
#GRUB_TERMINAL=console

# The resolution used on graphical terminal
# note that you can use only modes which your graphic card supports via VBE/GOP/UGA
# you can see them in real GRUB with the command `videoinfo'
#GRUB_GFXMODE=640x480

# Uncomment if you don't want GRUB to pass "root=UUID=xxx" parameter to Linux
#GRUB_DISABLE_LINUX_UUID=true

# Uncomment to disable generation of recovery mode menu entries
#GRUB_DISABLE_RECOVERY="true"

# Uncomment to get a beep at grub start
#GRUB_INIT_TUNE="480 440 1"

Can I just copy the missing /etc/default/grub file from the working node?

Not sure if it is relevant or not, but ThreadReaper (node with issue) is using secure boot, and Wrex (working node) is too old and doesn't use secure boot.
 
Last edited:
the output on ThreadReaper sounds very wrong - it states both grub and systemd-boot are installed on the ESPs, yet there is no Grub config file?

could you please post "apt list --installed | grep -e systemd -e grub"?
 
  • Like
Reactions: DAVe3283
Code:
root@ThreadReaper:~# apt list --installed | grep -e systemd -e grub

WARNING: apt does not have a stable CLI interface. Use with caution in scripts.

grub-common/stable,now 2.12-9+pmx2 amd64 [installed]
grub-efi-amd64-bin/stable,now 2.12-9+pmx2 amd64 [installed]
grub-efi-amd64-signed/stable,now 1+2.12+9+pmx2 amd64 [installed]
grub-efi-amd64-unsigned/stable,now 2.12-9+pmx2 amd64 [installed,automatic]
grub-efi-amd64/stable,now 2.12-9+pmx2 amd64 [installed,automatic]
grub-pc-bin/stable,now 2.12-9+pmx2 amd64 [installed]
grub2-common/stable,now 2.12-9+pmx2 amd64 [installed]
libnss-systemd/stable,stable,now 257.13-1~deb13u1 amd64 [installed]
libpam-systemd/stable,stable,now 257.13-1~deb13u1 amd64 [installed]
libsystemd-shared/stable,stable,now 257.13-1~deb13u1 amd64 [installed,automatic]
libsystemd0/stable,stable,now 257.13-1~deb13u1 amd64 [installed]
proxmox-grub/stable,now 2.12-9+pmx2 amd64 [installed,automatic]
python3-systemd/stable,now 235-1+b6 amd64 [installed,automatic]
systemd-boot-efi/stable,stable,now 257.13-1~deb13u1 amd64 [installed,automatic]
systemd-boot-tools/stable,stable,now 257.13-1~deb13u1 amd64 [installed,automatic]
systemd-boot/stable,now 257.13-1+pmx1 amd64 [installed]
systemd-cryptsetup/stable,stable,now 257.13-1~deb13u1 amd64 [installed,automatic]
systemd-sysv/stable,stable,now 257.13-1~deb13u1 amd64 [installed]
systemd/stable,stable,now 257.13-1~deb13u1 amd64 [installed]

I switched ThreadReaper from EFI boot (systemd) to secure boot (grub) a long time ago, shortly after upgrading to Proxmox 9 if I recall. I can't find/remember the guide I used to enable secure boot, but it looks like I missed a step or followed a bad guide. Interesting that it worked fine up until now.

It appears that I should remove the `systemd-boot` metapackage. Should I `apt-get install --reinstall` any of the grub packages to get it to recreate the appropriate files?
 
I would recommend the following (on ThreadReaper)

1. remove the systemd-boot package
2. fill in /etc/default/grub using the one from the working node, but adapt "GRUB_CMDLINE_LINUX_DEFAULT" to get the correct kernel cmdline
3. clear out /etc/kernel/proxmox-boot-uuids
4. reformat your ESPs using "proxmox-boot-tool format /dev/something2 --force" (make sure to pick the right partition! it is usually /dev/sdX2 or /dev/nvmeXnYp2, but please double check!)
5. reinit your ESPs using grub: "proxmox-boot-tool init /dev/something2 grub"

step 4 is dangerous, please verify you run it using the correct device path, otherwise you will lose data!
 
Thanks for all the help @fabian!

One interesting thing I noticed is Wrex (working node) has a file /etc/default/grub.d/zfs.cfg containing GRUB_CMDLINE_LINUX="$GRUB_CMDLINE_LINUX root=ZFS=rpool/ROOT/pve-1 boot=zfs". That file is missing on ThreadReaper.

For step #2 I copied /etc/default/grub and /etc/default/grub.d/zfs.cfg to ThreadReaper, but I didn't manually add the other parameters from /etc/kernel/cmdline to /etc/default/grub. As you implied, the /etc/kernel/cmdline parameters did NOT end up in the boot partition grub.cfg files on ThreadReaper. Do I have to add them manually in both /etc/default/grub & /etc/kernel/cmdline from now on? I wonder why Wrex seems to copy them in to grub.cfg but ThreadReaper does not? Maybe I am just missing something obvious, since it is late here.

I have not rebooted yet, but the boot partitions now contain the ZFS root parameters in the command line so I expect it will boot correctly now, aside from the quirk mentioned above about /etc/kernel/cmdline.

Thanks again!