What is proper way to make PVE host a dhcp client in 9.2 ?

shodan

Well-Known Member
Sep 1, 2022
317
77
48
Hi,
I'm upgrading to 9.2
I need to update my procedure to make the PVE host a dhcp client for its web userinterface.

There are the following concerns to address

editing /etc/network/interfaces to change to dhcp client mode
maybe install a dhcpclient (isc-dhcp-client)
deal with apparmor complaints if any
make /etc/hosts update automatically on IP address changes
make pvebanner reload automatically on IP address changes

and automate this entire process so it's as easy as flipping dhcp client on or off

And how it should act is that

If the dhcp server is unavailable, it should fallback to a sane address
and this sane address should be the last given dhcp address or a specified fallback for the very first time (the one determined at proxmox install time)
after that, when the dhcp server is unavailable at boot, or become unavailable during operation,
it should not prevent booting and it should not cease to function when the dhcp server becomes unavailable
and when the dhcp server becomes available again, the system should seamlessly switch to the new (or old) provided address
it should not fail then give up forever
disconnecting and reconnecting the network cable from the host should be as minimally disruptive as possible.
it should not produce spurious error or warning messages in logs
I did research previous questions about this issues in the following URLs

Code:
https://forum.level1techs.com/t/dhcp-for-proxmox-host-server/196733/5
https://forum.proxmox.com/threads/configured-proxmox-as-dhcp.180572/
https://forum.proxmox.com/threads/dhcp-in-proxmox.185112/
https://forum.proxmox.com/threads/dhcp-option.160785/
https://forum.proxmox.com/threads/my-dream-dhcp-server-built-into-proxmox-how-do-you-users-workarround-it.98900/
https://forum.proxmox.com/threads/proxmox-9-0-4-doesnt-get-ip-by-dhcp.169721/
https://forum.proxmox.com/threads/proxmox-9-0-4-doesnt-get-ip-by-dhcp.169721/page-2
https://forum.proxmox.com/threads/proxmox-automated-installation-dhcp-issue.151844/
https://forum.proxmox.com/threads/pve-nodes-with-dhcp-assigned-ips-and-hostnames.135952/
https://forum.proxmox.com/threads/set-a-dynamic-address-to-pve.119847/
https://forum.proxmox.com/threads/set-dhcp-on-proxmox.168358/
https://forum.proxmox.com/threads/static-ip-to-dhcp.154422/
https://forum.proxmox.com/threads/ve_host-web-interface-setup-for-dhcp.27481/
https://www.reddit.com/r/Proxmox/comments/1lplwkx/is_there_a_definitive_way_to_get_proxmox_to/

The general attitude is "no, create manual IP lease exceptions instead"

I took all these threads and consults with an LLM.

Here are the general concerns it identified

Code:
1. Basic network configuration
2. DHCP client implementation
3. DHCP unavailable during normal operation
4. DHCP unavailable during boot
5. Address-collision handling
6. DHCP becomes available again
7. Cable disconnect and reconnect
8. /etc/hosts
9. Node hostname
10. DNS and /etc/resolv.conf
11. Default gateway and DHCP routes
12. pvebanner and /etc/issue
13. pveproxy
14. AppArmor
15. DHCP event architecture
16. Persistent state
17. First DHCP conversion
18. IPv6
19. Firewall, storage, monitoring, APIs and external dependencies
20. Standalone versus clustered hosts
21. Logging
22. Upgrade safety and idempotence
23. Disabling DHCP again

It provided a look at the "dhcp state machine" once completed

1790484307768.png

And there is even a checklist of what "done" would look like

1790484324556.png

So, before I do all of that, I mostly wanted to ask, if anyone had already done it ?

And I did find these

https://free-pmx.org/guides/dhcp-single/
https://free-pmx.org/guides/dhcp-cluster/

But I'm not really sure it addresses all of the above.
 
I believe your scheme to fall back to using an IP from an old lease would technically violate the DHCP standard. If I remember correctly, the spec says that after a change (like a reboot, or disconnect and reconnect) you can continue with the old address only if the lease hasn't expired. If it has you have to get a new lease from DHCP, and can't have an IP until the DHCP server gives you one (unless you use one from the link-local range). That doesn't mean you can't do it if you want, just that I don't think you will find an already existing easy option that works the way you outlined.

I get that you want it to work even if the DHCP server is misbehaving (or just missing), but I would like to know what problem you are trying to solve this way that wouldn't be better addressed with a static IP? My best guess is that you want one simple image you can deploy somewhere just by cloning it to new hardware and booting without needing to take a few seconds to set a static IP, but that doesn't really make sense. So, if you don't mind, why do you want this?
 
  • Like
Reactions: Johannes S
For the why, yes, I don't want to hardcode anything, each hardcoded setting is another maintenance item I have to keep track of, I want everything automated.

As for violating the dhcp standard for keeping the install time network stats, yes, however that is not more violating that using stating IP addresses.
And that would only occur when the DHCP server is failing to respond plus I'm putting it collision checking on top so I'm fine with this, I have an IDS running on the internal network so any collision would also be very easy to detect and fix, but since this would only happen when the dhcp server has failed, I would already have bigger problems to deal with.

I think the only problem I left out is cluster issues related with proxmox hosts changing addesses and also, if the server's IP address gets hardcoded into anything else, for instance iscsi targets and the like, but I only every use the .lan hostname so I think it's all good. In those places where I can't use the hostname I'll put some background process that updates the IP based on whatever the hostname is.

I'm also considering letting the DHCP server itself name the proxmox host, I'm not sure if hostnamectl really will let the dhcp server set the hostname but I think that's be great to also remove another hardcoded value. Then proxmox server is, whatever host the dhcp decide is "proxmox.lan" which is going to be important for the other aspect I proposed in a separate thread, this is to suspend and shutdown the proxmox server is idle and also including automatic wake on lan on demand
 
each hardcoded setting is another maintenance item I have to keep track of
I see it exactly the other way: each non hardcoded setting is a moving target.

Usually servers should have a stable configuration, nothing dynamic. DHCP is only for clients like laptops/mobile/desktops...

Just my two 2 €¢...
 
Yes I understand the orthodox network position, I too was there 10 thousands years ago before dhcp even existed.
It feels scary to relinquish manual control to an unreliable untested machine.
It feels empowering to give numbers to things and know they won't change.

In aerospace there is a similar concept with fuel control units, old system designers used to route most of the levels of the FCU into the cockpit so the pilots could tune and tweak all the parameters of the fuel mixture but eventually they found out that while it feels like giving the user control and empowering them, it does not.

It is just extra meaningless busywork that makes the whole system inflexible and brittle.
It only feels solid because it cannot change anymore without manual intervention.
They are now numbers "set in stone" but there's not really any upsides to this, it's just extra work when you need to change it.

And I understand the reply to that is "I'll never change it anyway".

But my stuff needs to change and I don't want to have this minutiae standing in my way.

My working system should never ask me to type in IP addresses, if I am having to type IP addresses anywhere other than the DHCP lease definition, then I consider the system to be in a failed state.

Host should only have one identity, their name, the actual IP numbers should be irrelevant and not hardcoded anywhere.

And yes I'm having the same argument with the firewall people, they sure do love their IP numbers too !

I remembers in the dark past, we even had to enter memory address numbers, IRQ numbers, DMA numbers, manually configure PPP protocols, we needed straight and crossover cables even baud rates and parity bits.

All of that is work for machines, the only time we should be concerned with those is when the system has failed. Working systems should not be designed in a way that makes these concerns necessary.

I don't even want my default gateway to have an hard coded address.

I think these are bad habits we need to lose.
And I'm pretty sure these habits are why dns is always breaking everything, as the memes go on social media.

My dns always works, because if it doesn't then nothing works, so it always works.
 
I'll be honest - I can see the idea for using DHCP to configure everything - but I can also see it being a gateway for pain.

I'm not sure if you're looking at running more than one PVE host in a cluster - or this is a single server operation - but there are many more services that need a static layout.

You'll also need to manually update / reconfigure your corosync configuration to make sure that continues to work - risking entire cluster stability if you end up losing quorum. I'd also be weary of what would break in migrations if any IP information changed.

If you use correct ACME style certs, you'll need to make sure DNS also updates to the correct set of hosts. Without this, if you use passkeys, this functionality will also break.

Once you also use IPv6, you'll also have to start hooking into RA / DHCPv6 renewals and acting on those as well to update the running system - especially if you use temporary addressing within IPv6.

Personally, I use static IPv6 only on my PVE nodes at home, and static dual stack for standalone physical machines in data centres. While I can understand the ideology of not using static addressing for anything - I can also see it causing many more issues than it solves rather than manually assigning a single address to your most critical infrastructure.

As for VMs, LXC installs, absolutely, go ham and have all of those configured via your DHCP infrastructure.
 

CRCinAU, you get it​

Yep, corosync is far too brittle, similar to ip based firewalling, would only continue working with an hostname process that keeps the hardcoded addresses aligned with hostnames.

I also don't really have a use for the cluster system since it doesn't allow to run one VM that uses the CPU, ram & hardware resources of multiple proxmox host, I haven't investigated it further, but the way it breaks even logging in when it fails is just too much

Code:
https://forum.proxmox.com/threads/no-quorum-error.113459/
https://forum.proxmox.com/threads/how-to-disable-quorum-mechanism-in-a-cluster.159498/

I have no working system for handling certificate distribution on the LAN, that will be its own set of headaches when it gets there. But the certificates will be issues to .lan domains, not their internal IP addresses, so I hope that will sidestep the problem entirely.

For IPv6 I don't know yet what it's going to look like, but I also would prefer to have everything be ipv6 with ipv4 disabled on the LAN entirely.
And never, ever look at or type a single IPv6 address in the same way that I've never had to type a memory address to use a program.
But that is still far off, I have a bunch of hardware that doesn't even support ipv6 to replace before I get there.

Making proxmox work off a dhcp client for the web interface isn't hard, I've done it for years. Latest version introduced issues with apparmor but I think that is resolved now. As long as you use hostname for everyhing, it doesn't seem to be an issue.

I just wanted to do it in a very clean way that continues to work in case of dhcp server outages.
 
I think I found a real problem with AppArmor,
The problem is that even when configured, AppArmor is complaining about dhclient creating sockets
Here is a summary of the entire thing

-----------------------------------
Hi,

I am testing DHCP operation for the management bridge on a standalone Proxmox VE 9.2 host and have encountered what appears to be an AppArmor compatibility/policy issue affecting isc-dhclient.

DHCP itself works correctly: the host receives and maintains its IPv4 address. However, AppArmor repeatedly denies normal AF_UNIX socket creation by dhclient.

I have reproduced the problem, confirmed that the local AppArmor include is being loaded, tried several increasingly explicit AppArmor rules, and finally traced dhclient in both enforce and complain mode.

The denied sockets turn out to be normal local sockets used for:


  • []/dev/log — dhclient syslog output
    [
    ]/var/run/nscd/socket — normal glibc NSCD probing

The DHCP protocol sockets themselves continue to work.

Environment

Code:
Proxmox VE:       9.2
pve-manager:      9.2.2
Debian:           13 (trixie)
Kernel:           7.0.2-6-pve
Architecture:     x86_64

ifupdown2:        3.3.0-1+pmx12
isc-dhcp-client:  4.4.3-P1-8

apparmor:         4.1.1-pmx1
libapparmor1:     4.1.1-pmx1
apparmor_parser:  4.1.1

apparmor-utils:   4.1.0-1
python3-apparmor: 4.1.0-1

This is a standalone PVE node.

Management bridge:

Code:
vmbr0

Physical bridge port:

Code:
nic0

DHCP successfully gives the host:

Code:
10.0.0.188/16

1. Production dhclient is running normally

Command:

Code:
ps -ef | grep '[d]hclient.*vmbr0'

Output:

Code:
root         829       1  0 15:21 ?        00:00:00 /sbin/dhclient -pf /run/dhclient.vmbr0.pid -lf /var/lib/dhcp/dhclient.vmbr0.leases vmbr0 -nw

DHCP itself is functional and the host receives 10.0.0.188/16.

2. AppArmor is enforcing the dhclient profile

Command:

Code:
aa-status --verbose

Relevant output:

Code:
apparmor module is loaded.

20 profiles are in enforce mode.

...

/{,usr/}sbin/dhclient

The running production process is confined by that profile.

3. Stock dhclient AppArmor profile

Command:

Code:
cat /etc/apparmor.d/usr.sbin.dhclient

Relevant parts:

Code:
#include <tunables/global>

/{,usr/}sbin/dhclient flags=(attach_disconnected) {
#include <abstractions/base>
#include <abstractions/nameservice>
#include <abstractions/openssl>

capability net_bind_service,
capability net_raw,
capability dac_override,
capability net_admin,

network packet,
network raw,

...

/etc/dhcp/ r,
/etc/dhcp/** r,

...

/{,usr/}sbin/dhclient-script Uxr,

...

[HEADING=2]Site-specific additions and overrides. See local/README for details.[/HEADING]
#include <local/usr.sbin.dhclient>
}

4. AppArmor denials

Command:

Code:
journalctl -k -b --no-pager |
grep 'apparmor="DENIED".*profile="/{,usr/}sbin/dhclient"'

Example output:

Code:
Sep 28 15:21:55 aria kernel: audit: type=1400 audit(1790623315.574:131): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=829 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:21:55 aria kernel: audit: type=1400 audit(1790623315.578:132): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=829 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:21:55 aria kernel: audit: type=1400 audit(1790623315.578:133): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=829 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:21:55 aria kernel: audit: type=1400 audit(1790623315.582:134): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=829 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:21:55 aria kernel: audit: type=1400 audit(1790623315.873:135): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=829 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none

On other runs I also get the same error for sock_type="stream".

The part that seems especially relevant is:

Code:
operation="create"
class="net"
info="failed protocol match"
family="unix"
protocol=0
requested="create"
denied="create"

5. First attempted workaround

I first tried adding family/type-specific network rules to:

Code:
/etc/apparmor.d/local/usr.sbin.dhclient

Rules:

Code:
network unix dgram,
network unix stream,

The profile was reloaded.

This did not change the result.

The same denial continued:

Code:
operation="create"
class="net"
info="failed protocol match"
family="unix"
protocol=0

6. Second attempted workaround: broad network/unix rules

I then tried:

Code:
network,
unix,

I verified that the local include really was being expanded into the profile.

Command:

Code:
apparmor_parser -p /etc/apparmor.d/usr.sbin.dhclient 2>/dev/null |
grep -nE '(^|[[:space:]])(deny|allow|audit|network|unix)|local/usr.sbin.dhclient' |
tail -100

Relevant output:

Code:
1293:##included <local/usr.sbin.dhclient>
1296:# class=net with protocol=0. Family-specific network rules did not match on
1297:# the tested PVE 9.2 image. Allow network mediation generically, but only
1299:# also allowed for subsequent AF_UNIX operations.
1300,
1301,

So the local include definitely was being loaded.

Nevertheless, the same denials continued:

Code:
apparmor="DENIED"
operation="create"
class="net"
info="failed protocol match"
profile="/{,usr/}sbin/dhclient"
family="unix"
sock_type="dgram"
protocol=0
requested="create"
denied="create"

7. Third attempted workaround: explicit create permission

I then tried matching the exact operation being reported:

Code:
network (create) unix stream,
network (create) unix dgram,

The local override was:

Code:
[HEADING=2]BEGIN DHCPCLIENTMODEFORPROXMOXHOSTS[/HEADING]
[HEADING=2]Debian 13 / AppArmor 4 network-v8 mediates socket creation separately.[/HEADING]
[HEADING=2]The observed dhclient denials are specifically AF_UNIX create operations:[/HEADING]
[HEADING=2]family=unix sock_type=stream protocol=0 requested=create[/HEADING]
[HEADING=2]family=unix sock_type=dgram  protocol=0 requested=create[/HEADING]
[HEADING=2]Grant only those create operations.[/HEADING]
network (create) unix stream,
network (create) unix dgram,

[HEADING=2]END DHCPCLIENTMODEFORPROXMOXHOSTS[/HEADING]

Again, the profile reloaded successfully.

Again, the same denial remained:

Code:
apparmor="DENIED"
operation="create"
class="net"
info="failed protocol match"
profile="/{,usr/}sbin/dhclient"
family="unix"
sock_type="dgram"
protocol=0
requested="create"
denied="create"

So explicitly granting the reported create operation did not solve it either.

8. Kernel AppArmor network features

Commands:

Code:
find /sys/kernel/security/apparmor/features/network_v9 
-type f -maxdepth 2 
-exec sh -c 'echo "--- $1"; cat "$1"' _ {} ;

find /sys/kernel/security/apparmor/features/network_v8 
-type f -maxdepth 2 
-exec sh -c 'echo "--- $1"; cat "$1"' _ {} ;

Output:

Code:
===== KERNEL NETWORK V9 FEATURES =====

--- /sys/kernel/security/apparmor/features/network_v9/af_unix
yes

--- /sys/kernel/security/apparmor/features/network_v9/af_mask
unspec unix inet ax25 ipx appletalk netrom bridge atmpvc x25 inet6 rose netbeui security key netlink packet ash econet atmsvc rds sna irda pppox wanpipe llc ib mpls can tipc bluetooth iucv rxrpc isdn phonet ieee802154 caif alg nfc vsock kcm qipcrtr smc xdp mctp

and:

Code:
===== NETWORK V8 FEATURES =====

--- /sys/kernel/security/apparmor/features/network_v8/af_inet
yes

--- /sys/kernel/security/apparmor/features/network_v8/af_mask
unspec unix inet ax25 ipx appletalk netrom bridge atmpvc x25 inet6 rose netbeui security key netlink packet ash econet atmsvc rds sna irda pppox wanpipe llc ib mpls can tipc bluetooth iucv rxrpc isdn phonet ieee802154 caif alg nfc vsock kcm qipcrtr smc xdp mctp

The running kernel therefore exposes network_v9.

9. Installed AppArmor ABI descriptions

Command:

Code:
grep -n -A12 -B2 'network' 
/etc/apparmor.d/abi/4.0 
/etc/apparmor.d/abi/3.0 2>/dev/null

Relevant output:

Code:
/etc/apparmor.d/abi/4.0:47 {af_mask {unspec unix inet ...
...
/etc/apparmor.d/abi/4.0:52 {af_mask {unspec unix inet ...

The installed ABI data has network_v8, but I do not see network_v9.

So, at least from what I can observe:

Code:
running kernel feature set: network_v9
installed ABI description:  network_v8

I do not know whether this is the cause, but it seems potentially relevant to the failed protocol match error.

10. AppArmor package versions

Command:

Code:
dpkg-query -W -f='${Package}\t${Version}\n' 
apparmor apparmor-utils libapparmor1 python3-apparmor 2>/dev/null

Output:

Code:
apparmor        4.1.1-pmx1
apparmor-utils  4.1.0-1
libapparmor1    4.1.1-pmx1
python3-apparmor        4.1.0-1

The PVE AppArmor core/library packages are 4.1.1-pmx1, while the Debian utility/Python packages are 4.1.0-1.

11. aa-logprof sees the events but proposes no rule

Initially aa-logprof was not installed.

I installed:

Code:
apt install apparmor-utils

Because this host has no /var/log/syslog, I exported the kernel journal manually.

Commands:

Code:
STATE=/var/lib/DHCPclientModeForProxmoxHosts
BASELINE=$(cat "$STATE/apparmor.baseline.epoch" 2>/dev/null || printf '0\n')
LOG=/root/dhclient-apparmor-audit.log

journalctl -k --since="@$BASELINE" --no-pager > "$LOG"

echo "Audit log: $LOG"
grep 'apparmor="DENIED".*profile="/{,usr/}sbin/dhclient"' "$LOG"

aa-logprof -f "$LOG"

aa-logprof output:

Code:
Updating AppArmor profiles in /etc/apparmor.d.
Reading log entries from /root/dhclient-apparmor-audit.log.
Complain-mode changes:
Enforce-mode changes:

It did not propose any rule for the dhclient denials.

-----------------------------------
Hit the 16'000 character limit, see next post
 
continued in next post during to character limit
-----------------------------------
12. strace while AppArmor was enforcing

I traced a second test dhclient while the normal production client remained active.

Command:

Code:
strace -ff
-o /root/dhclient-strace
-e trace=socket,connect,sendto,recvfrom,openat
/sbin/dhclient
-1
-pf /run/dhclient-test.pid
-lf /var/lib/dhcp/dhclient.vmbr0.leases
vmbr0

Because the production client was already managing vmbr0, the test invocation reported:

Code:
Error: ipv4: Address already assigned.

However, the trace captured the socket failures before that.

Command:

Code:
grep -HnE 'socket(AF_UNIX|connect(.(dev/log|journal|nscd|systemd|run/)' 
/root/dhclient-strace

Relevant output:

Code:
/root/dhclient-strace.5244:4(AF_UNIX, SOCK_DGRAM|SOCK_CLOEXEC, 0) = -1 EACCES (Permission denied)
/root/dhclient-strace.5245:9(AF_UNIX, SOCK_STREAM|SOCK_CLOEXEC|SOCK_NONBLOCK, 0) = -1 EACCES (Permission denied)
/root/dhclient-strace.5245:10(AF_UNIX, SOCK_STREAM|SOCK_CLOEXEC|SOCK_NONBLOCK, 0) = -1 EACCES (Permission denied)
/root/dhclient-strace.5245:38(AF_UNIX, SOCK_DGRAM|SOCK_CLOEXEC, 0) = -1 EACCES (Permission denied)
/root/dhclient-strace.5245:39(AF_UNIX, SOCK_DGRAM|SOCK_CLOEXEC, 0) = -1 EACCES (Permission denied)
/root/dhclient-strace.5245:41(AF_UNIX, SOCK_DGRAM|SOCK_CLOEXEC, 0) = -1 EACCES (Permission denied)

At the same time, the kernel recorded matching AppArmor events:

Code:
Sep 28 15:42:30 aria kernel: audit: type=1400 audit(1790624550.740:136): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=5244 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:42:30 aria kernel: audit: type=1400 audit(1790624550.743:137): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=5245 comm="dhclient" family="unix" sock_type="stream" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:42:30 aria kernel: audit: type=1400 audit(1790624550.743:138): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=5245 comm="dhclient" family="unix" sock_type="stream" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:42:30 aria kernel: audit: type=1400 audit(1790624550.867:139): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=5245 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:42:30 aria kernel: audit: type=1400 audit(1790624550.872:140): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=5245 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none
Sep 28 15:42:31 aria kernel: audit: type=1400 audit(1790624551.109:141): apparmor="DENIED" operation="create" class="net" info="failed protocol match" error=-13 profile="/{,usr/}sbin/dhclient" pid=5245 comm="dhclient" family="unix" sock_type="dgram" protocol=0 requested="create" denied="create" addr=none

13. Same test with dhclient temporarily in AppArmor complain mode

I temporarily changed only the dhclient profile to complain mode:

Code:
aa-complain /sbin/dhclient

Status confirmed that /{,usr/}sbin/dhclient was in complain mode.

Then I traced the same operation:

Code:
rm -f /root/dhclient-complain-strace.*

strace -ff -yy
-o /root/dhclient-complain-strace
-e trace=socket,connect,sendto,recvfrom,openat
/sbin/dhclient
-1
-pf /run/dhclient-test.pid
-lf /var/lib/dhcp/dhclient.vmbr0.leases
vmbr0

I immediately restored enforcement afterward:

Code:
aa-enforce /sbin/dhclient

14. In complain mode the previously denied AF_UNIX datagram socket succeeds and connects to /dev/log

The trace showed:

Code:
===== /root/dhclient-complain-strace.6859 =====
1  openat(AT_FDCWD</root>, "/etc/ld.so.cache", O_RDONLY|O_CLOEXEC) = 3</etc/ld.so.cache>
2  openat(AT_FDCWD</root>, "/lib/x86_64-linux-gnu/libc.so.6", O_RDONLY|O_CLOEXEC) = 3</usr/lib/x86_64-linux-gnu/libc.so.6>
3  openat(AT_FDCWD</root>, "/dev/null", O_RDWR) = 3</dev/null<char 1:3>>
4 socket(AF_UNIX, SOCK_DGRAM|SOCK_CLOEXEC, 0) = 3UNIX:[38024]
5 connect(3UNIX:[38024], {sa_family=AF_UNIX, sun_path="/dev/log"}, 110) = 0
6  +++ exited with 0 +++

So at least one of the denied sockets is simply the normal syslog socket.

15. dhclient actively uses that socket for logging

Later in the trace:

Code:
40 sendto(3UNIX:[38024-4303]>, "<30>Sep 28 15:50:47 dhclient[686"..., 98, MSG_NOSIGNAL, NULL, 0) = 98
41 sendto(3UNIX:[38024-4303]>, "<30>Sep 28 15:50:47 dhclient[686"..., 71, MSG_NOSIGNAL, NULL, 0) = 71
43 sendto(3UNIX:[38024-4303]>, "<30>Sep 28 15:50:47 dhclient[686"..., 84, MSG_NOSIGNAL, NULL, 0) = 84

Therefore the AF_UNIX datagram socket is definitively being used for dhclient logging.

16. The stream sockets are normal NSCD probes

Relevant trace:

Code:
===== /root/dhclient-complain-strace.6862 =====

9  socket(AF_UNIX, SOCK_STREAM|SOCK_CLOEXEC|SOCK_NONBLOCK, 0) = 4UNIX-STREAM:[37281]
10 connect(4UNIX-STREAM:[37281], {sa_family=AF_UNIX, sun_path="/var/run/nscd/socket"}, 110) = -1 ENOENT (No such file or directory)

11 socket(AF_UNIX, SOCK_STREAM|SOCK_CLOEXEC|SOCK_NONBLOCK, 0) = 4UNIX-STREAM:[37286]
12 connect(4UNIX-STREAM:[37286], {sa_family=AF_UNIX, sun_path="/var/run/nscd/socket"}, 110) = -1 ENOENT (No such file or directory)

There is no NSCD daemon/socket on this host, so ENOENT appears to be the expected result.

Under AppArmor enforcement these calls fail earlier at socket() with EACCES.

17. DHCP networking itself works

The same trace shows normal DHCP-related sockets being created successfully:

Code:
33 socket(AF_NETLINK, SOCK_RAW|SOCK_CLOEXEC, NETLINK_ROUTE) = 6
36 socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL)) = 6
37 socket(AF_INET, SOCK_DGRAM, IPPROTO_IP) = 7
38 socket(AF_INET, SOCK_DGRAM, IPPROTO_UDP) = 7

The client then receives DHCP traffic from the DHCP server:

Code:
48 recvfrom(7UDP:[0.0.0.0:68], ...,
{sa_family=AF_INET,
sin_port=htons(67),
sin_addr=inet_addr("10.0.0.1")},
...) = 303

This is why DHCP still works despite the AF_UNIX denials.

18. Helper processes can use AF_UNIX sockets normally

Other traced processes successfully use systemd:

Code:
socket(AF_UNIX, SOCK_STREAM|SOCK_CLOEXEC|SOCK_NONBLOCK, 0) = 3
connect(3, {sa_family=AF_UNIX, sun_path="/run/systemd/private"}, 23) = 0

Chrony also succeeds:

Code:
socket(AF_UNIX, SOCK_DGRAM|SOCK_CLOEXEC|SOCK_NONBLOCK, 0) = 3
connect(3, {sa_family=AF_UNIX, sun_path="/run/chrony/chronyd.sock"}, 110) = 0

This seems consistent with the profile containing:

Code:
/{,usr/}sbin/dhclient-script Uxr,

so these later helper processes are not subject to the main dhclient confinement in the same way.

19. Cleanup after testing

Some of the test dhclient processes daemonized, so I cleaned them up explicitly.

Before cleanup:

Code:
ps -eo pid,ppid,stat,cmd | grep -E '([s]trace|[d]hclient)'

Output:

Code:
829   1    Ss   /sbin/dhclient ... vmbr0 -nw
5241  1169 S+   strace ...
5245  1    Ss   /sbin/dhclient -1 ... vmbr0
6856  5348 S+   strace ...
6862  1    Ss   /sbin/dhclient -1 ... vmbr0

After killing only the test processes and test strace instances:

Code:
ps -eo pid,ppid,stat,cmd | grep -E '([s]trace|[d]hclient)'

Final output:

Code:
829       1 Ss   /sbin/dhclient -pf /run/dhclient.vmbr0.pid -lf /var/lib/dhcp/dhclient.vmbr0.leases vmbr0 -nw

So only the normal production DHCP client remains.

Expected behavior

I would expect the stock AppArmor policy to allow normal AF_UNIX socket creation used by the stock Debian/PVE dhclient, particularly:

Code:
socket(AF_UNIX, SOCK_DGRAM, 0)

for:

Code:
/dev/log

and the ordinary libc probe:

Code:
socket(AF_UNIX, SOCK_STREAM, 0)
connect(... "/var/run/nscd/socket")

Alternatively, if these operations are intentionally meant to be denied, I would expect the profile/policy to handle them without repeated unexpected DENIED messages during otherwise normal DHCP operation.

Actual behavior

In enforce mode:

Code:
socket(AF_UNIX, ..., 0)

returns:

Code:
EACCES

and AppArmor records:

Code:
class="net"
info="failed protocol match"
family="unix"
protocol=0

This happens even though:


  • []AF_UNIX support is present in the kernel AppArmor feature set
    [
    ]the profile includes the standard AppArmor base/name-service abstractions
    []multiple explicit local AF_UNIX/network rules were tested
    [
    ]the local rules were confirmed to be included in the expanded profile
  • aa-logprof sees the events but proposes no corresponding rule

Impact

DHCP itself currently works despite the denials.

The observable effects are:


  1. []dhclient cannot create its normal AF_UNIX syslog socket to /dev/log while confined.
    [
    ]glibc's normal NSCD probes are rejected with EACCES instead of naturally reaching ENOENT.
    []The kernel emits repeated AppArmor DENIED messages during normal dhclient operation/startup.
    [
    ]This makes an otherwise working DHCP configuration look like a security or networking failure.
  2. Apparently reasonable local AppArmor rules do not resolve the denials.

Possible ABI/feature mismatch

I do not want to claim this is the root cause without confirmation from the maintainers, but this seems suspicious:

Code:
PVE kernel exposes:       network_v9
installed AppArmor ABI:   network_v8

The actual denial specifically says:

Code:
info="failed protocol match"
protocol=0

and aa-logprof is unable to suggest a rule.

There is also this package-version combination:

Code:
apparmor        4.1.1-pmx1
libapparmor1    4.1.1-pmx1
apparmor-utils  4.1.0-1
python3-apparmor        4.1.0-1

Again, I do not know whether either difference is causative.

Questions

Could someone from Proxmox/AppArmor confirm:

  1. Should socket(AF_UNIX, SOCK_DGRAM, 0) and socket(AF_UNIX, SOCK_STREAM, 0) be permitted by the current PVE/Debian dhclient AppArmor profile?
  2. Why does the kernel report:
    Code:
    info="failed protocol match"
    protocol=0
    despite explicit AF_UNIX/network rules being present?
  3. Is it expected that the PVE kernel exposes AppArmor network_v9 while /etc/apparmor.d/abi/4.0 only appears to describe network_v8?
  4. Is the combination of PVE apparmor/libapparmor1 4.1.1-pmx1 with Debian apparmor-utils/python3-apparmor 4.1.0-1 expected?
  5. Does the shipped /etc/apparmor.d/usr.sbin.dhclient profile need an update for PVE 9.2 / Debian 13?
  6. Is there a recommended local AppArmor rule that correctly matches these protocol=0 AF_UNIX socket creations?
  7. If these socket operations are intentionally considered unnecessary, is there a recommended way to suppress the repeated audit noise without putting dhclient into complain mode or broadly weakening its confinement?

At this point I have restored the dhclient profile to enforce mode and removed the test clients. The production dhclient remains running normally and DHCP is functional.

Thanks.

-----------------------------------