Hi everyone,
while planning the migration of lots of VMs running on hypervisors with software from a well known US company to Proxmox VE, our storage people requested that we implement some dynamic resource scheduling to keep the load on our storages balanced. So we built it. To be honest - with the help of AI - writing so much code in a small time frame and testing it would not be possible otherwise. Its running on production for about a week now, not only balancing disks, but it also allows to do storage migrations (which got on the todo list right after migrating to Proxmox VE).
proxmox-storage-drs (command name: pve-storage-drs — deliberately not “pve-drs”, to keep it out of the Dynamic Load
Balancer's way in both name and function) balances disk I/O load across configurable groups of
shared storages — FC/LVM, Ceph RBD, or anything else shared that PVE supports — by
live-migrating individual VM disks between the storages of a group. It never moves a guest.
In short:
It is implemented, dogfooded against a production cluster (ours), and released as a Debian
package targeting trixie — so it installs on a PVE 9 node with a plain
of the
operator manual, manpage and internals document all ship in the package.
One prerequisite: Prometheus.
The tool collects nothing itself and ships no exporter. It reads six per-disk I/O counters from
a Prometheus-compatible backend that you already run. The data exists on every cluster already —
pvestatd gathers the QEMU blockstats and PVE's InfluxDB external metric server exports
them. Plain Prometheus, VictoriaMetrics and gigapipe on ClickHouse all work;
checks the whole pipeline before you plan anything.
The one thing to read up front: what survives the trip between PVE and your backend matters,
because some transports silently drop one of the six counters for individual disks — your
dashboards keep drawing, and the balancer reads the gap as “this disk does nothing”. If you have
not set up an external metric server yet, this is the one piece of preparation you need.
How it was written
In the interest of transparency: due to time constraints, this project was written with
substantial AI assistance at every stage — planning and architecture as well as most of the
code, tests and documentation. Every commit names the model that wrote it, and the result went
through two independent review rounds (one human, one automated), whose findings are tracked
publicly in the repository. The test suite and the review record are there precisely so you do
not have to take the outcome on faith.
Test data welcome
Arguably the most valuable contribution is not code but real-cluster test data.
captures a fully replayable, pseudonymized bundle of your cluster (read-only by construction,
identifiers replaced, free-text fields never captured) — and
entire engine against such a bundle with zero network connections, offline, forever. Synthetic
fixtures can prove the solver optimal on six disks; only a real cluster can prove the solver and
the heuristic still agree at three hundred. The submission guide is in the repository under
If you have two shared storages that always feel unbalanced, give it a dry run:
Repository and releases:
https://github.com/bzed/proxmox-storage-drs
Documentation (installation, configuration, every command, safety properties):
https://github.com/bzed/proxmox-storage-drs/tree/main/docs/manual
while planning the migration of lots of VMs running on hypervisors with software from a well known US company to Proxmox VE, our storage people requested that we implement some dynamic resource scheduling to keep the load on our storages balanced. So we built it. To be honest - with the help of AI - writing so much code in a small time frame and testing it would not be possible otherwise. Its running on production for about a week now, not only balancing disks, but it also allows to do storage migrations (which got on the todo list right after migrating to Proxmox VE).
proxmox-storage-drs (command name: pve-storage-drs — deliberately not “pve-drs”, to keep it out of the Dynamic Load
Balancer's way in both name and function) balances disk I/O load across configurable groups of
shared storages — FC/LVM, Ceph RBD, or anything else shared that PVE supports — by
live-migrating individual VM disks between the storages of a group. It never moves a guest.
In short:
- Reads per-disk I/O load from Prometheus and the cluster inventory from the PVE API, then
computes a migration plan that evens out the load across a storage group while performing as
few migrations as possible. - The plan is computed by a real MILP solver (CBC), with a dependency-free heuristic as
fallback — and the two are cross-checked against each other in the test suite so they cannot
drift apart. - A snapshot free-space reserve is guaranteed on every storage, at every instant of a move —
and it is never traded away for a better balance, by construction. - Dry-run is the default. Nothing migrates until you explicitly select a less safe execution
mode, and there is a per-move confirmation mode before the unattended one. explain
narrates the plan in plain language: what is pinned and why, the payback arithmetic, and which
close move the solver rejected.
It is implemented, dogfooded against a production cluster (ours), and released as a Debian
package targeting trixie — so it installs on a PVE 9 node with a plain
apt installof the
.deb from the GitHub releases page. Licensed under AGPL-3.0-or-later. Theoperator manual, manpage and internals document all ship in the package.
One prerequisite: Prometheus.
The tool collects nothing itself and ships no exporter. It reads six per-disk I/O counters from
a Prometheus-compatible backend that you already run. The data exists on every cluster already —
pvestatd gathers the QEMU blockstats and PVE's InfluxDB external metric server exports
them. Plain Prometheus, VictoriaMetrics and gigapipe on ClickHouse all work;
verify-metricschecks the whole pipeline before you plan anything.
The one thing to read up front: what survives the trip between PVE and your backend matters,
because some transports silently drop one of the six counters for individual disks — your
dashboards keep drawing, and the balancer reads the gap as “this disk does nothing”. If you have
not set up an external metric server yet, this is the one piece of preparation you need.
How it was written
In the interest of transparency: due to time constraints, this project was written with
substantial AI assistance at every stage — planning and architecture as well as most of the
code, tests and documentation. Every commit names the model that wrote it, and the result went
through two independent review rounds (one human, one automated), whose findings are tracked
publicly in the repository. The test suite and the review record are there precisely so you do
not have to take the outcome on faith.
Test data welcome
Arguably the most valuable contribution is not code but real-cluster test data.
collect-testdatacaptures a fully replayable, pseudonymized bundle of your cluster (read-only by construction,
identifiers replaced, free-text fields never captured) — and
--replay runs theentire engine against such a bundle with zero network connections, offline, forever. Synthetic
fixtures can prove the solver optimal on six disks; only a real cluster can prove the solver and
the heuristic still agree at three hundred. The submission guide is in the repository under
tests/corpus/README.md — open the bundle and read it before sending it anywhere.If you have two shared storages that always feel unbalanced, give it a dry run:
pve-storage-drs -c /etc/pve/drs.yaml planRepository and releases:
https://github.com/bzed/proxmox-storage-drs
Documentation (installation, configuration, every command, safety properties):
https://github.com/bzed/proxmox-storage-drs/tree/main/docs/manual