From: Tim Menninger <tmenninger@everpuredata.com>
To: linux-mm@kvack.org
Cc: Chuck Lever <cel@kernel.org>,
linux-nfs@vger.kernel.org, Jon Curley <jcurley@everpuredata.com>,
Eric Badger <ebadger@everpuredata.com>
Subject: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA
Date: Wed, 16 Sep 2026 23:22:27 +0000 [thread overview]
Message-ID: <20260916232227.4098143-1-tmenninger@everpuredata.com> (raw)
Chuck Lever suggested I bring this to linux-mm after we found some SLUB
measurements we could not account for while investigating an NFS/RDMA
throughput regression.
The original NFS discussion is here for context:
https://lore.kernel.org/linux-nfs/0e2f9688-097c-4bb9-a7ec-82b41bdb3653@slotpi15m67/
The NFS regression itself has been separated from this issue. What
remains interesting here is the behavior of the nfs_page slab cache on
this machine, particularly the measurements with and without perf lock.
The system is:
Intel Xeon Silver 4516Y+
2 sockets
24 cores/socket
2 threads/core
96 logical CPUs
NUMA node0 CPUs: 0-23,48-71
NUMA node1 CPUs: 24-47,72-95
The workload is a high-throughput NFS/RDMA direct-read workload using
1 MiB I/O, 160 threads, and iodepth 64.
The relevant debug options are all disabled:
# CONFIG_KASAN is not set
# CONFIG_PROVE_LOCKING is not set
# CONFIG_LOCK_STAT is not set
# CONFIG_DEBUG_SPINLOCK is not set
# CONFIG_DEBUG_LIST is not set
All measurements below were collected on the unpatched base kernel:
$ git rev-parse HEAD
940de590b839f71d6dc846160534bf202401b8b7
$ uname -r
7.3.0-rc1-mainline+
The initial observation was a high apparent contention rate on the
nfs_page slab's list_lock. In 10-second perf-lock captures I was seeing
roughly 4.5M contended acquisitions and about 170-185 us of average
reported wait.
Chuck reproduced a similar acquisition rate on a single-node EPYC
system, but saw only about 7 ns average wait and fewer than 100
cmpxchg_double_fail events over a corresponding interval. He suggested
checking cmpxchg_double_fail because __slab_free() drops list_lock and
retries when the freelist cmpxchg fails.
I repeated the measurements in three placement configurations:
A. workload unpinned, CQs all on node0
B. workload pinned to node0, CQs all on node0
C. workload unpinned, CQs balanced across the nodes
Without perf lock, throughput is similar in all three:
A. unpinned / CQs node0: ~45.5 GB/s
B. node0 pinned / CQs node0: ~45.7 GB/s
C. unpinned / balanced CQs: ~45.5 GB/s
During the perf-lock captures, throughput is approximately 25 GB/s.
For each instrumented 10-second window I ran:
sudo perf lock record -a -g -o "$D/perf-locks.data" -- sleep 10
mpstat -P ALL 1 10
Before and after the same window I sampled the counters under:
/sys/kernel/slab/nfs_page/
I also collected separate 10-second counter and mpstat windows under
the same workload configurations without perf lock.
The resulting slab counter deltas were:
A B C
unpinned/node0 node0/node0 unpinned/balanced
free_fastpath
instrumented 655,917,745 457,936,140 328,670,030
uninstrumented 34,906,233 119,984,226 59,469,931
free_slowpath
instrumented 236,735,724 7,720,917 270,716,506
uninstrumented 85,039,753 2,814 60,145,870
sheaf_flush
instrumented 39,842,700 42,706,800 8,341,440
uninstrumented 1,286,400 13,487,700 887,700
barn_put_fail
instrumented 663,994 711,745 139,037
uninstrumented 21,464 224,762 14,765
barn_get_fail
instrumented 4,609,925 840,344 4,651,225
uninstrumented 1,438,643 224,787 1,017,006
cmpxchg_double_fail
instrumented 70,524 8,077 28,062
uninstrumented 7,471 713 1,503
alloc_slowpath
all cases 0 0 0
The SLUB counter profile changes substantially with perf lock despite
the lower NFS throughput, and the exact mix depends strongly on
placement.
The uninstrumented node0/node0 run also reproduces the barn/sheaf
relationship Chuck pointed out earlier. The nfs_page sheaf capacity is
60:
sheaf_flush / 60 = 13,487,700 / 60 = 224,795
barn_put_fail = 224,762
The placement dependence of free_slowpath is also large. It falls from
about 85M events per 10 seconds in the unpinned/node0 case to 2,814 in
the node0/node0 case.
The reported perf-lock result, however, is similar across all three
placements:
contentions total wait average wait
unpinned/node0 4,782,839 14.39 min 180.47 us
node0/node0 4,582,012 14.13 min 185.05 us
unpinned/balanced 4,572,196 12.81 min 168.14 us
I also revisited an inconsistency Chuck noticed in my earlier
measurements. Previously I had compared aggregate perf-lock wait from
one 10-second capture with CPU utilization measured during a different
window.
I now have paired 10-second mpstat samples for each placement, with and
without perf lock. The node values below are averages of the per-CPU
%idle values for the CPUs in each NUMA node:
system-wide node0 node1
%idle %idle %idle
A. unpinned / CQs node0
uninstrumented 31.46 8.0 54.6
instrumented 3.60 0.06 7.1
B. node0 pinned / CQs node0
uninstrumented 84.26 69.2 99.4
instrumented 12.07 10.4 13.8
C. unpinned / balanced CQs
uninstrumented 66.86 66.7 66.9
instrumented 14.83 19.0 10.6
This resolves the accounting inconsistency in my earlier measurements.
The large aggregate perf-lock wait and high idle percentage had come
from different windows. In the aligned samples, the system is much
busier during the perf-lock capture than in the corresponding
uninstrumented run.
I am still unsure how representative the reported ~170-185 us average
wait is of the uninstrumented workload.
The remaining number I am less sure how to interpret is
cmpxchg_double_fail. In the uninstrumented windows I see:
unpinned / CQs node0: 7,471 / 10 sec
node0 pinned / CQs node0: 713 / 10 sec
unpinned / balanced CQs: 1,503 / 10 sec
compared with fewer than 100 in Chuck's test.
I understand that cmpxchg_double_fail counts failed slab freelist
updates rather than failed logical frees, so I am not sure what the
appropriate denominator is here. In particular, the node0/node0 case
still has 713 failures while sustaining full throughput and only 2,814
free_slowpath events over the interval.
My questions are:
1. Do the uninstrumented cmpxchg_double_fail rates above look abnormal
for this workload/topology, or are they within the range one would
expect from this degree of concurrency and NUMA placement?
2. Is there a less invasive way you would recommend measuring the
nfs_page list_lock/freelist contention? I would like to distinguish
the steady-state behavior from what is observed during the perf-lock
capture.
Thanks,
Tim
next reply other threads:[~2026-09-16 23:22 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 23:22 Tim Menninger [this message]
2026-09-17 13:46 ` [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA Harry Yoo
2026-09-17 14:10 ` Harry Yoo
2026-09-17 15:29 ` Peter Zijlstra
2026-09-18 7:04 ` Namhyung Kim
2026-09-17 16:25 ` Tim Menninger
2026-09-18 7:08 ` Vlastimil Babka (SUSE)
2026-09-21 15:31 ` Harry Yoo
2026-09-21 23:27 ` Tim Menninger
2026-09-18 7:14 ` Namhyung Kim
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260916232227.4098143-1-tmenninger@everpuredata.com \
--to=tmenninger@everpuredata.com \
--cc=cel@kernel.org \
--cc=ebadger@everpuredata.com \
--cc=jcurley@everpuredata.com \
--cc=linux-mm@kvack.org \
--cc=linux-nfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox