Rule-database sync triggers false Net::SMTP 421/451 after successful reinjection, causing duplicate delivery

Marro A.

New Member
Oct 1, 2026
1
0
1
Hello,

we are investigating recurring duplicate deliveries on a Proxmox Mail Gateway cluster node. The evidence indicates that a rule-database synchronization triggers SIGUSR1 while a filter worker is reinjecting a message through local Postfix. This can cause PMG’s Perl Net::SMTP client to report an immediate, synthetic 421 [Net::SMTP] Timeout, although Postfix successfully accepts and forwards the message.

Environment:

  • Proxmox Mail Gateway: 8.0.1
  • pmg-api: 8.0.7
  • pmg-gui: 4.0.2
  • Postfix: 3.7.6
  • Perl Net::SMTP/Net::Cmd: 3.14
  • Kernel: 6.2.16-6-pve
  • Clustered PMG installation
  • Affected node: pmg-voda
The problem occurs in the following sequence:

  1. pmgmirror finishes synchronizing the rule database.
  2. PMG::Cluster calls PMG::DBTools::reload_ruledb().
  3. This calls PMG::Utils::reload_smtp_filter().
  4. reload_smtp_filter() sends SIGUSR1 to the pmg-smtp-filter parent.
  5. The parent forwards SIGUSR1 to its worker processes.
  6. A worker that is waiting for the SMTP response from local Postfix can immediately report:
sending data failed - got: 421 [Net::SMTP] Timeout : ERROR
  1. PMG consequently returns this temporary failure to the original sending server:
451 4.4.0 detected undelivered mail
  1. Local Postfix has nevertheless already accepted and forwarded the message.
  2. Because the original sender received 451, it legitimately retries the message, producing duplicate delivery.
Example from 2026-10-01:

2026-10-01T10:06:30.151365+0200 pmgmirror[838]:
finished rule database sync from host '10.254.254.49'

2026-10-01T10:06:30.153330+0200 pmg-smtp-filter[1337374]:
sending data failed - got: 421 [Net::SMTP] Timeout : ERROR
The interval between these two entries is approximately 1.97 milliseconds.

Four such errors occurred on 2026-10-01. Every one immediately followed completion of a rule-database synchronization:

10:06:30.151365 sync completed
10:06:30.153330 Net::SMTP timeout

10:38:34.730330 sync completed
10:38:34.742281 Net::SMTP timeout

12:34:34.086721 sync completed
12:34:34.087273 Net::SMTP timeout
12:34:34.087376 Net::SMTP timeout
We observed the same correlation during previous monitoring. The errors are not genuine expiration of the configured timeout:

  • Net::SMTP timeout: 120 seconds
  • smtpd_proxy_timeout: 400 seconds
  • lmtp_data_done_timeout: 600 seconds
  • Observed failure: milliseconds after the synchronization signal
A packet capture of traffic originating from local Postfix on port 10025 confirmed that Postfix emitted successful queue acknowledgements for affected transactions:

2026-09-17 17:12:29.005730 pmg-smtp-filter:
sending data failed - got: 421 [Net::SMTP] Timeout : ERROR

2026-09-17 17:12:29.007001 localhost:10025:
250 2.0.0 Ok: queued as CB96444E59
A second simultaneous transaction showed the same result:

2026-09-17 17:12:29.005733 pmg-smtp-filter:
sending data failed - got: 421 [Net::SMTP] Timeout : ERROR

2026-09-17 17:12:29.011294 localhost:10025:
250 2.0.0 Ok: queued as D14D244E54
Both messages were subsequently accepted by their next-hop servers, despite PMG returning 451 to the original SMTP clients.

The installed PMG source confirms the reload path:

Code:
sub reload_ruledb {
...
PMG::Utils::reload_smtp_filter();
}
sub reload_smtp_filter {
...
return kill(10, $pid); # send SIGUSR1
}
The filter parent forwards that signal to its children:

$SIG{'USR1'} = sub {
if (defined $prop->{children}) {
foreach my $pid (keys %{$prop->{children}}) {
kill(10, $pid); # SIGUSR1 children
}
}
};
The worker installs a handler that requests configuration reload:

Code:
$SIG{'USR1'} = sub {
    $self->{reload_config} = 1;
};
We also traced a filter worker and confirmed that a reload signal can interrupt a blocking system call:

Code:
--- SIGUSR1 {si_signo=SIGUSR1, si_code=SI_USER,
             si_pid=<pmg-smtp-filter-parent>, si_uid=0} ---

rt_sigreturn({mask=[]}) = -1 EINTR (Interrupted system call)
The response-reading path in the installed /usr/share/perl/5.36/Net/Cmd.pm uses:

Code:
my $select_ret = select($rout = $rin, undef, undef, $timeout);

if ($select_ret > 0) {
...
} else {
$cmd->_set_status_timeout;
return;
}
Consequently, both a real timeout (select() returns 0) and a system-call error such as EINTR (select() returns -1) are converted into the same synthetic 421 [Net::SMTP] Timeout status.

The write-side implementation in Net::Cmd handles EINTR separately, but this response-reading path does not appear to do so.


We understand that this installation is outdated and plan to update first to the latest PMG 8.2 release before migrating to PMG 9. However, reviewing newer PMG and Perl libnet source suggests that the same local reinjection and Net::Cmd response-reading behavior may still be present.

Could Proxmox please advise:

  1. Is this interaction between pmgmirror, SIGUSR1, Net::SMTP and filter-worker reinjection already known?
  2. Has it been fixed in a newer pmg-api or PMG release?
  3. Should configuration reload signals be deferred or blocked while a worker is processing/reinjecting a message?
  4. Is there a recommended safe workaround until the system is upgraded or a fix is available?