From: Adrian Tomasov <atomasov@redhat.com>
To: Eric Dumazet <edumazet@google.com>,
kernel test robot <oliver.sang@intel.com>
Cc: Adrian Tomasov <atomasov@redhat.com>,
Jason Xing <kerneljasonxing@gmail.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Paolo Abeni <pabeni@redhat.com>, Luke Yang <luyang@redhat.com>,
netdev@vger.kernel.org, linux-kernel@vger.kernel.org,
oe-lkp@lists.linux.dev, lkp@intel.com,
linuxppc-dev@lists.ozlabs.org
Subject: Re: [linus:master] [net] 5628f3fe3b: sockperf.throughput.UDP.msg_per_sec 20.1% regression
Date: Thu, 10 Sep 2026 16:30:40 +0200 [thread overview]
Message-ID: <20260910143041.18106-1-atomasov@redhat.com> (raw)
In-Reply-To: <202512112119.5b9829a-lkp@intel.com>
On Thu, Dec 11, 2025, kernel test robot <oliver.sang@intel.com> wrote:
> kernel test robot noticed a 20.1% regression of
> sockperf.throughput.UDP.msg_per_sec on:
> commit: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()")
We independently reproduced this on a second architecture, ppc64le (POWER10),
and can add data that both confirms the commit and shows what amplifies it.
Commit: 5628f3fe3b16 ("net: add NUMA awareness to skb_attempt_defer_free()")
Parent: 844c9db7f7f5 ("net: use llist for sd->defer_list")
Two factors govern the magnitude, and your x86 report and our ppc64le data
agree once both are considered:
(a) How much of the workload runs the net_rx softirq deferred-free path.
Tight loopback UDP (sockperf, iperf3 loopback) is softirq-dominated and is
hit hard. A real-NIC test bounded by other costs is barely affected.
(b) How many *possible* NUMA nodes skb_defer_free_flush() now walks. The new
for_each_node() loop iterates node_possible_map; on a POWER10 LPAR that is
0..31 (32 possible) with only 1 node online, so 31 of every 32 iterations
touch a cold, always-empty per-node list on every softirq RX pass. This
shows up as _find_next_bit(): ~2.76% on POWER10 vs <1% in your x86 profile.
--- Data point 1: stress-ng UDP, isolated commit A/B (POWER10) ---
POWER10, 8 CPUs, 62 GiB, NUMA 1 online / 32 possible. Two upstream v7.2-rc4
kernels, same compiler (gcc 16.1.1) and config; the only difference is the
presence of 5628f3fe3b16. stress-ng 0.19.03, --udp 1 -t 23, 3 runs,
bogo-ops/sec (higher is better):
Kernel Avg Delta
v7.2-rc4 + revert (good) 265145 baseline
v7.2-rc4 stock (bad) 204854 -22.7%
perf (stock vs revert): skb_defer_free_flush 2.75% -> 0%, _find_next_bit
2.76% -> 0% -- both eliminated by the revert; consistent with your x86 profile
(skb_defer_free_flush 0 -> ~4.1%, _find_next_bit -> ~1%).
--- Data point 2: iperf3 UDP loopback A/B (POWER10, second host) ---
Independent reproduction on another POWER10 host (32 possible / 1 online),
iperf3 -u -b 0 -l 16k -t 20, both ends pinned to the online node, 5 runs,
Gbit/s (higher is better):
Kernel (has 5628f3fe3b16?) median mean
no (6.12-based) 6.34 6.31 baseline
yes (7.1-based) 4.85 4.88 -22.7%
(These two kernels differ in base version, so this delta is not a single-commit
isolation on its own -- Data point 1 is the clean isolation. Both land at
-22.7%, matching your 20.1% sockperf figure.)
--- Data point 3: workload dependence (real-NIC vs loopback) ---
The same commit-vs-parent A/B measured with a *real NIC* iperf3 UDP stream
(200G, 3 reboots/kernel) on a 2-possible-node ppc64le host shows only -2.6%
(7515.8 -> 7313.8 Mb/s, zero overlap across reboots). Where the deferred-free
softirq path is a small fraction of the work, the regression is small -- the
mechanism is the same, the exposure differs. This also matches our observation
that a same-CPU-heavy stress-ng udp run barely triggers the cross-CPU
skb_attempt_defer_free() path and shows ~no change.
--- On Jason's question ---
On Sun, Dec 14, 2025, Jason Xing <kerneljasonxing@gmail.com> wrote:
> from what I've known, commit e20dfbad8aab2 and commit 21664814b89e altogether
> can lead to a similar regression ... could you also launch some experiments
> just around those two commits?
On POWER10 our isolation is a direct A/B of 5628f3fe3b16 against its parent
844c9db7f7f5 with identical compiler and config (Data point 1), and a full
revert of only 5628f3fe3b16 recovers the throughput and removes
skb_defer_free_flush()/_find_next_bit() from the profile. So on this hardware
the regression is attributable to 5628f3fe3b16 specifically; e20dfbad8aab2 and
21664814b89e are not in the delta. Happy to test those two separately if useful.
--- Suggested fix ---
Use for_each_online_node() instead of for_each_node() in skb_defer_free_flush()
(and cap what skb_attempt_defer_free() populates accordingly), reducing the loop
from N_possible to N_online -- 32 -> 1 on these POWER10 LPARs. We can build and
benchmark a candidate patch on the POWER10 hosts and report back on this thread.
If a fix is posted, the appropriate tags are:
Reported-by: kernel test robot <oliver.sang@intel.com>
Closes: https://lore.kernel.org/oe-lkp/202512112119.5b9829a-lkp@intel.com
Thanks,
Adrian Tomasov
Red Hat -- Kernel Performance QE
prev parent reply other threads:[~2026-09-10 14:31 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-12-11 14:31 [linus:master] [net] 5628f3fe3b: sockperf.throughput.UDP.msg_per_sec 20.1% regression kernel test robot
2025-12-11 16:15 ` Eric Dumazet
2025-12-12 3:59 ` Oliver Sang
2025-12-15 2:47 ` Jason Xing
2025-12-16 2:15 ` Oliver Sang
2026-09-10 14:30 ` Adrian Tomasov [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260910143041.18106-1-atomasov@redhat.com \
--to=atomasov@redhat.com \
--cc=edumazet@google.com \
--cc=kerneljasonxing@gmail.com \
--cc=kuniyu@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=lkp@intel.com \
--cc=luyang@redhat.com \
--cc=netdev@vger.kernel.org \
--cc=oe-lkp@lists.linux.dev \
--cc=oliver.sang@intel.com \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.