Setup
- PVE 9.2.21, FRR 10.6.1-1+pve3, 3-node cluster (Ceph), EVPN controller (ASN 65000), EVPN underlay on a dedicated 10G network (MTU 9000).
- Firewall/edge (OPNsense) on a separate host, connected via a static VXLAN transit (1 Gbit). Exit nodes peer eBGP with it.
- Multi-tenant: one shared-services network (DNS, NTP, proxy, update mirror, monitoring, Ansible) and one network per customer. Each is a /24 split into 16 VNets (/28, anycast gateway).
Goal
- Tenant isolation by architecture: one EVPN zone (= VRF) per tenant, so there is no route between tenants. Then a firewall mistake (e.g. firewall=0 on a NIC) cannot open a path to another tenant.
- Controlled traffic tenant ↔ shared services (DNS, proxy, updates, Ansible), ideally routed inside the cluster at 10G, not hair-pinned via the 1G edge firewall.
What we observed
1. One zone for everything (current production) works well, but every node routes between all tenants. Isolation depends only on the PVE firewall (guest + datacenter FORWARD).
2. Multiple EVPN zones with exit nodes: on exit nodes PVE puts import vrf vrf_<zone> into the default VRF, and in each exit zone's VRF it adds ip route <subnet-of-other-zone> null0 for the subnets of all other zones. We read the null routes as deliberate protection against inter-zone forwarding via the exit node. So PVE apparently does not intend routing between zones.
3. Custom VRF route leaking via frr.conf.local (import vrf … route-map, match source-vrf, prefix lists) to leak only the shared-services prefixes into tenant VRFs and back.
- It works with a single additional zone.
- With several zones we saw re-advertised leaked routes in EVPN (type-5 /32s with the foreign L3VNI / route target). Nodes then forwarded via the wrong L3VNI, and the kernel dropped replies (ip route get … iif vrfbr_<zone> → RTNETLINK answers: Invalid argument, fib_validate_source).
- Restricting advertise ipv4 unicast route-map to own prefixes did not resolve it.
- After removing the zones, stale router bgp … vrf instances kept advertising routes until FRR was restarted on every node.
4. VNet zone change: pvesh set /cluster/sdn/vnets/<vnet> --zone <new> exists, but is refused with "can't change zone if subnets exist". The workaround is delete subnet → change zone → recreate subnet → one apply.
5. Inter-zone via the external firewall works, but every shared-services flow then goes through the 1 Gbit edge link and a single firewall.
Questions
- Is there a supported way to have multiple EVPN zones with policy-controlled inter-zone routing inside the cluster (e.g. tenant VRF → shared-services VRF, only selected prefixes/ports)?
- Is VRF route leaking between EVPN zones (import vrf via frr.conf.local) supported or planned? Is there a recommended way to prevent leaked routes from being re-advertised as EVPN type-5 with the importing VRF's RT?
- Are the per-zone null0 routes on exit nodes intended as inter-zone protection, and can they be controlled?
- Is the recommended design "one zone per tenant + router/firewall VM(s) inside the cluster attached to each zone"? Are there best practices (NIC limits, HA)?
Our interim solution: one zone per tenant, plus a router VM pair inside the cluster (nftables, keepalived) with one NIC per zone. Each zone has only a static default route to the router VIP in its own VRF, with no leaking.
- PVE 9.2.21, FRR 10.6.1-1+pve3, 3-node cluster (Ceph), EVPN controller (ASN 65000), EVPN underlay on a dedicated 10G network (MTU 9000).
- Firewall/edge (OPNsense) on a separate host, connected via a static VXLAN transit (1 Gbit). Exit nodes peer eBGP with it.
- Multi-tenant: one shared-services network (DNS, NTP, proxy, update mirror, monitoring, Ansible) and one network per customer. Each is a /24 split into 16 VNets (/28, anycast gateway).
Goal
- Tenant isolation by architecture: one EVPN zone (= VRF) per tenant, so there is no route between tenants. Then a firewall mistake (e.g. firewall=0 on a NIC) cannot open a path to another tenant.
- Controlled traffic tenant ↔ shared services (DNS, proxy, updates, Ansible), ideally routed inside the cluster at 10G, not hair-pinned via the 1G edge firewall.
What we observed
1. One zone for everything (current production) works well, but every node routes between all tenants. Isolation depends only on the PVE firewall (guest + datacenter FORWARD).
2. Multiple EVPN zones with exit nodes: on exit nodes PVE puts import vrf vrf_<zone> into the default VRF, and in each exit zone's VRF it adds ip route <subnet-of-other-zone> null0 for the subnets of all other zones. We read the null routes as deliberate protection against inter-zone forwarding via the exit node. So PVE apparently does not intend routing between zones.
3. Custom VRF route leaking via frr.conf.local (import vrf … route-map, match source-vrf, prefix lists) to leak only the shared-services prefixes into tenant VRFs and back.
- It works with a single additional zone.
- With several zones we saw re-advertised leaked routes in EVPN (type-5 /32s with the foreign L3VNI / route target). Nodes then forwarded via the wrong L3VNI, and the kernel dropped replies (ip route get … iif vrfbr_<zone> → RTNETLINK answers: Invalid argument, fib_validate_source).
- Restricting advertise ipv4 unicast route-map to own prefixes did not resolve it.
- After removing the zones, stale router bgp … vrf instances kept advertising routes until FRR was restarted on every node.
4. VNet zone change: pvesh set /cluster/sdn/vnets/<vnet> --zone <new> exists, but is refused with "can't change zone if subnets exist". The workaround is delete subnet → change zone → recreate subnet → one apply.
5. Inter-zone via the external firewall works, but every shared-services flow then goes through the 1 Gbit edge link and a single firewall.
Questions
- Is there a supported way to have multiple EVPN zones with policy-controlled inter-zone routing inside the cluster (e.g. tenant VRF → shared-services VRF, only selected prefixes/ports)?
- Is VRF route leaking between EVPN zones (import vrf via frr.conf.local) supported or planned? Is there a recommended way to prevent leaked routes from being re-advertised as EVPN type-5 with the importing VRF's RT?
- Are the per-zone null0 routes on exit nodes intended as inter-zone protection, and can they be controlled?
- Is the recommended design "one zone per tenant + router/firewall VM(s) inside the cluster attached to each zone"? Are there best practices (NIC limits, HA)?
Our interim solution: one zone per tenant, plus a router VM pair inside the cluster (nftables, keepalived) with one NIC per zone. Each zone has only a static default route to the router VIP in its own VRF, with no leaking.
