Hardware clock access missing in kernel 7.0.14-20 (and reboot issue)

leesteken

Distinguished Member
May 31, 2020
8,246
3,005
278
I use rtcwake to shutdown my Proxmox host (and run hwclock --systohc as root just before that) and this worked fine for many years including kernel 7.0.14-19.

The recent (no-subscription) proxmox-kernel-7.0.14-20-pve-signed breaks this and shows the following in the system log:
Code:
hwclock: Cannot access the Hardware Clock via any known method.
hwclock: Use the --verbose option to see the details of our search for an access method.
rtcwake: /dev/rtc0: unable to find device: No such file or directory

Is this intentional? Is there a work-around to get /dev/rtc0 back?

PS: The host does not complete a reboot with 7.0.14-20, it seems to hang after the shut down part without actually shutting down or rebooting.
I appreciate all the (urgent security) fixes in the new kernel version but this one doesn't really work for me.

EDIT: cat /sys/devices/system/clocksource/clocksource0/available_clocksource gives tsc hpet acpi_pm on both kernel versions.
 
Last edited:
I can reproduce this on a test node (a PVE VM, q35 + OVMF), so it's not just your hardware.

On 7.0.14-19 (and -17) everything is fine. rtc_cmos picks up the clock and hwclock works:

1790926272236.png

On 7.0.14-20 /dev/rtc0 is gone and I get the same hwclock error as you. The RTC driver is still built into the kernel, but it never gets loaded:

1790926287618.png

(Just checking: where your post says "7.0.14-19 breaks this", you mean -20, right?)

From what I can see, the kernel config is basically the same between the two versions, so this looks like a code change in -20. The clock device is still there, but -20 creates it in a different place than before (as a "platform" device instead of a "PnP" device). The RTC driver still looks for it in the old place, doesn't find it, and gives up. So nothing creates /dev/rtc0.

If you want to check it on your box, run this on -20:

Code:
cat /sys/bus/pnp/devices/*/id
ls -l /sys/bus/platform/devices/PNP0B00:00/driver

On mine the first one doesn't list PNP0B00 any more, and the second one says the driver link doesn't exist.

Btw, the clocksource thing in your edit isn't related. That's about the kernel's own timekeeping, not the hardware clock.

For now the easiest fix is to stay on -19:

Code:
proxmox-boot-tool kernel pin 7.0.14-19-pve

Once a fixed kernel comes out, run proxmox-boot-tool kernel unpin.

I couldn't reproduce the reboot hang though, my VM reboots fine on -20. That part might be specific to your hardware. If you can grab the last lines on the console when it hangs, that would help.
 
  • Like
Reactions: leesteken
(Just checking: where your post says "7.0.14-19 breaks this", you mean -20, right?)
Indeed (and fixed), thanks for reproducing the issue.
If you want to check it on your box, run this on -20:

Code:
cat /sys/bus/pnp/devices/*/id
ls -l /sys/bus/platform/devices/PNP0B00:00/driver

On mine the first one doesn't list PNP0B00 any more, and the second one says the driver link doesn't exist.
Maybe later I can test this on -20 but on -19 those give PNP0b00 and there is not such directory.
For now the easiest fix is to stay on -19:
Yes, this was my plan also until something better comes along, as I don't want to run outdated kernel versions for too long. I was mostly wondering whether this was intended as a security fix and whether there was a work-around.

I couldn't reproduce the reboot hang though, my VM reboots fine on -20. That part might be specific to your hardware. If you can grab the last lines on the console when it hangs, that would help.
The last lines on the console are a i915 driver crash because it does not handle unloading very well. This is the same for a successful reboot, so nothing related there. The system logs ends normally:
Code:
Reached target reboot.target - System Reboot.
Shutting down.
Syncing filesystems and block devices.
Sending SIGTERM to remaining processes...
Received SIGTERM from PID 1 (systemd-shutdow).
Journal stopped
I've used kernel parameters to chose a different reboot method in the past on other motherboards. This is usually fixed with a BIOS update. In this case it felt as just another thing with this particular kernel version.
 
I tried 7.0.14-21-pve from pve-test but that behaves the same as -20 except that my AMD GPU no longer shows the Linux boot or shutdown messages, which is very inconvenient when the system hangs with a black screen. Maybe an amdgpu driver or firmware missing as is did not upgrade pve-firmware? Anyway, I'll patiently wait for the next kernel version before testing again.