* [BUG] net/sched: netem: delay correlation has no effect
@ 2026-10-02 16:50 Christophe Burdinat
2026-10-04 17:45 ` Stephen Hemminger
0 siblings, 1 reply; 2+ messages in thread
From: Christophe Burdinat @ 2026-10-02 16:50 UTC (permalink / raw)
To: stephen@networkplumber.org; +Cc: netdev@vger.kernel.org
Hello,
Context: we found the following while evaluating a QUIC-based protocol (Media over QUIC) with netem-based network profiles from a 3GPP study (TR 26.934), where the reordering caused by per-packet jitter makes QUIC declare many packets lost and changes the results substantially. The following report has been generated with Claude code, but humanly reviewed.
netem's delay correlation ("delay TIME JITTER CORRELATION") has no measurable effect: consecutive packets get independent delays whatever the correlation, with or without a distribution table. tc-netem(8) says the correlation "isn't a true statistical correlation, but an approximation"; in practice it is not applied at all.
How it was measured
-------------------
Two network namespaces joined by a veth pair; on one end:
tc qdisc add dev va root netem limit 100000 \
delay 100ms 20ms <corr>% [distribution normal|pareto]
5000 UDP packets, 1 ms apart, each carrying its send timestamp; the receiver (same host clock) computes the one-way delay of each packet and the lag-1 autocorrelation of consecutive delays (first 100 packets left out: they wait together for ARP resolution). Script below (repro.sh; needs root, iproute2, python3).
distribution correlation delay median std lag-1 autocorr.
normal 0 % 100.4 ms 19.6 ms +0.013
normal 50 % 100.2 ms 19.8 ms +0.010
normal 99 % 99.4 ms 20.2 ms -0.008
pareto 0 % 93.6 ms 16.7 ms +0.021
pareto 50 % 93.5 ms 15.3 ms -0.003
pareto 99 % 93.6 ms 15.5 ms -0.003
(none) 0 % 99.6 ms 11.5 ms -0.019
(none) 50 % 100.2 ms 11.4 ms -0.002
(none) 99 % 99.9 ms 11.5 ms -0.025
With 4900 samples, |r| < 0.03 is noise. The delay spread doesn't change either. A consequence users notice: with jitter larger than the packet spacing, netem reorders packets heavily (in our tests, 90 % of packets arrived after a later-sent one, normal table, 1 ms spacing), and raising the delay correlation, the obvious knob to make delays vary smoothly, does not reduce it.
Cause
-----
get_crandom() blends the new 32-bit random value with the previous result:
answer = (value * ((1ull<<32) - rho) + state->last * rho) >> 32;
so the correlation is carried by the high-order bits of answer.
tabledist() then keeps only the low-order bits:
t = dist->table[rnd % dist->size]; /* size 4096 */
...
return ((rnd % (2 * (u32)sigma)) + mu) - sigma; /* no table */
Between two consecutive packets, answer moves by a step on the order of (1 - rho/2^32) * 2^32, e.g. ~4e7 at 99 %. That is far more than 4096, so the table index is effectively uniformly random and the correlation is lost. Without a table the step is comparable to or larger than 2 * sigma (4e7 for 20 ms), with the same result.
A model of the two functions in Python (below, model_crandom.py, using netem's normal.dist), 30,000 draws per row, mu 100 ms, sigma 20 ms:
current code high bits instead of modulo
table corr lag-1 std lag-1 std
normal 0 % -0.008 19.98 ms -0.003 19.99 ms
normal 25 % +0.005 19.97 ms +0.270 12.95 ms
normal 50 % +0.007 19.98 ms +0.510 9.08 ms
normal 90 % -0.003 20.00 ms +0.901 3.40 ms
normal 99 % -0.004 19.97 ms +0.995 1.63 ms
(none) 50 % -0.007 11.55 ms +0.500 6.67 ms
(none) 99 % -0.036 11.57 ms +0.995 1.11 ms
The model's current-code column matches the measurements.
Possible fixes
--------------
1. Index with the high bits:
t = dist->table[((u64)rnd * dist->size) >> 32];
return (((u64)rnd * (2 * (u32)sigma)) >> 32) + mu - sigma;
This restores the correlation (lag-1 ~ the configured value), but it also shrinks the jitter as the correlation grows (20 ms -> 1.6 ms at 99 % in the model): get_crandom()'s output is an AR(1) blend of uniform values, so its variance falls by about (1 - p)/(1 + p). The
"approximation" in the man page would become visible, and existing configurations that set a correlation would see much smaller jitter.
2. Keep the configured marginal distribution while correlating, e.g. rescale get_crandom()'s output to full range before indexing, or correlate in a different domain. That changes the meaning of an existing parameter too, so I suspect it needs discussion.
3. If neither is wanted, document in tc-netem(8) that delay correlation currently has no effect, so that users (and test specifications tha rely on it) stop expecting it to.
I'm happy to test patches.
Christophe Burdinat
^ permalink raw reply [flat|nested] 2+ messages in thread* Re: [BUG] net/sched: netem: delay correlation has no effect
2026-10-02 16:50 [BUG] net/sched: netem: delay correlation has no effect Christophe Burdinat
@ 2026-10-04 17:45 ` Stephen Hemminger
0 siblings, 0 replies; 2+ messages in thread
From: Stephen Hemminger @ 2026-10-04 17:45 UTC (permalink / raw)
To: Christophe Burdinat; +Cc: netdev@vger.kernel.org
On Fri, 2 Oct 2026 16:50:07 +0000
Christophe Burdinat <c.burdinat@ateme.com> wrote:
> Hello,
>
> Context: we found the following while evaluating a QUIC-based protocol (Media over QUIC) with netem-based network profiles from a 3GPP study (TR 26.934), where the reordering caused by per-packet jitter makes QUIC declare many packets lost and changes the results substantially. The following report has been generated with Claude code, but humanly reviewed.
Thanks for the bug report. Will look into it but do not expect anything soon.
This code is over 20 years old, and it may just be you are trying to use
it in a way that was never intended.
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-04 17:45 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-02 16:50 [BUG] net/sched: netem: delay correlation has no effect Christophe Burdinat
2026-10-04 17:45 ` Stephen Hemminger
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox