Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA
@ 2026-09-16 23:22 Tim Menninger
  2026-09-17 13:46 ` Harry Yoo
  0 siblings, 1 reply; 10+ messages in thread
From: Tim Menninger @ 2026-09-16 23:22 UTC (permalink / raw)
  To: linux-mm; +Cc: Chuck Lever, linux-nfs, Jon Curley, Eric Badger

Chuck Lever suggested I bring this to linux-mm after we found some SLUB
measurements we could not account for while investigating an NFS/RDMA
throughput regression.

The original NFS discussion is here for context:

    https://lore.kernel.org/linux-nfs/0e2f9688-097c-4bb9-a7ec-82b41bdb3653@slotpi15m67/

The NFS regression itself has been separated from this issue. What
remains interesting here is the behavior of the nfs_page slab cache on
this machine, particularly the measurements with and without perf lock.

The system is:

    Intel Xeon Silver 4516Y+
    2 sockets
    24 cores/socket
    2 threads/core
    96 logical CPUs

    NUMA node0 CPUs: 0-23,48-71
    NUMA node1 CPUs: 24-47,72-95

The workload is a high-throughput NFS/RDMA direct-read workload using
1 MiB I/O, 160 threads, and iodepth 64.

The relevant debug options are all disabled:

    # CONFIG_KASAN is not set
    # CONFIG_PROVE_LOCKING is not set
    # CONFIG_LOCK_STAT is not set
    # CONFIG_DEBUG_SPINLOCK is not set
    # CONFIG_DEBUG_LIST is not set

All measurements below were collected on the unpatched base kernel:

    $ git rev-parse HEAD
    940de590b839f71d6dc846160534bf202401b8b7

    $ uname -r
    7.3.0-rc1-mainline+

The initial observation was a high apparent contention rate on the
nfs_page slab's list_lock. In 10-second perf-lock captures I was seeing
roughly 4.5M contended acquisitions and about 170-185 us of average
reported wait.

Chuck reproduced a similar acquisition rate on a single-node EPYC
system, but saw only about 7 ns average wait and fewer than 100
cmpxchg_double_fail events over a corresponding interval. He suggested
checking cmpxchg_double_fail because __slab_free() drops list_lock and
retries when the freelist cmpxchg fails.

I repeated the measurements in three placement configurations:

    A. workload unpinned, CQs all on node0
    B. workload pinned to node0, CQs all on node0
    C. workload unpinned, CQs balanced across the nodes

Without perf lock, throughput is similar in all three:

    A. unpinned / CQs node0:       ~45.5 GB/s
    B. node0 pinned / CQs node0:   ~45.7 GB/s
    C. unpinned / balanced CQs:    ~45.5 GB/s

During the perf-lock captures, throughput is approximately 25 GB/s.

For each instrumented 10-second window I ran:

    sudo perf lock record -a -g -o "$D/perf-locks.data" -- sleep 10
    mpstat -P ALL 1 10

Before and after the same window I sampled the counters under:

    /sys/kernel/slab/nfs_page/

I also collected separate 10-second counter and mpstat windows under
the same workload configurations without perf lock.

The resulting slab counter deltas were:

                            A               B                 C
                     unpinned/node0    node0/node0   unpinned/balanced

    free_fastpath
    instrumented        655,917,745    457,936,140       328,670,030
    uninstrumented       34,906,233    119,984,226        59,469,931

    free_slowpath
    instrumented        236,735,724      7,720,917       270,716,506
    uninstrumented       85,039,753          2,814        60,145,870

    sheaf_flush
    instrumented         39,842,700     42,706,800         8,341,440
    uninstrumented        1,286,400     13,487,700           887,700

    barn_put_fail
    instrumented            663,994        711,745           139,037
    uninstrumented           21,464        224,762            14,765

    barn_get_fail
    instrumented          4,609,925        840,344         4,651,225
    uninstrumented        1,438,643        224,787         1,017,006

    cmpxchg_double_fail
    instrumented             70,524          8,077            28,062
    uninstrumented            7,471            713             1,503

    alloc_slowpath
    all cases                     0              0                 0

The SLUB counter profile changes substantially with perf lock despite
the lower NFS throughput, and the exact mix depends strongly on
placement.

The uninstrumented node0/node0 run also reproduces the barn/sheaf
relationship Chuck pointed out earlier. The nfs_page sheaf capacity is
60:

    sheaf_flush / 60 = 13,487,700 / 60 = 224,795
    barn_put_fail                      = 224,762

The placement dependence of free_slowpath is also large. It falls from
about 85M events per 10 seconds in the unpinned/node0 case to 2,814 in
the node0/node0 case.

The reported perf-lock result, however, is similar across all three
placements:

                         contentions    total wait    average wait
    unpinned/node0         4,782,839      14.39 min      180.47 us
    node0/node0            4,582,012      14.13 min      185.05 us
    unpinned/balanced      4,572,196      12.81 min      168.14 us

I also revisited an inconsistency Chuck noticed in my earlier
measurements. Previously I had compared aggregate perf-lock wait from
one 10-second capture with CPU utilization measured during a different
window.

I now have paired 10-second mpstat samples for each placement, with and
without perf lock. The node values below are averages of the per-CPU
%idle values for the CPUs in each NUMA node:

                                  system-wide       node0       node1
                                  %idle             %idle       %idle

    A. unpinned / CQs node0
       uninstrumented                 31.46           8.0        54.6
       instrumented                    3.60           0.06        7.1

    B. node0 pinned / CQs node0
       uninstrumented                 84.26          69.2        99.4
       instrumented                   12.07          10.4        13.8

    C. unpinned / balanced CQs
       uninstrumented                 66.86          66.7        66.9
       instrumented                   14.83          19.0        10.6

This resolves the accounting inconsistency in my earlier measurements.
The large aggregate perf-lock wait and high idle percentage had come
from different windows. In the aligned samples, the system is much
busier during the perf-lock capture than in the corresponding
uninstrumented run.

I am still unsure how representative the reported ~170-185 us average
wait is of the uninstrumented workload.

The remaining number I am less sure how to interpret is
cmpxchg_double_fail. In the uninstrumented windows I see:

    unpinned / CQs node0:       7,471 / 10 sec
    node0 pinned / CQs node0:     713 / 10 sec
    unpinned / balanced CQs:    1,503 / 10 sec

compared with fewer than 100 in Chuck's test.

I understand that cmpxchg_double_fail counts failed slab freelist
updates rather than failed logical frees, so I am not sure what the
appropriate denominator is here. In particular, the node0/node0 case
still has 713 failures while sustaining full throughput and only 2,814
free_slowpath events over the interval.

My questions are:

1. Do the uninstrumented cmpxchg_double_fail rates above look abnormal
   for this workload/topology, or are they within the range one would
   expect from this degree of concurrency and NUMA placement?

2. Is there a less invasive way you would recommend measuring the
   nfs_page list_lock/freelist contention? I would like to distinguish
   the steady-state behavior from what is observed during the perf-lock
   capture.

Thanks,
Tim


^ permalink raw reply	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-09-21 23:27 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-16 23:22 [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA Tim Menninger
2026-09-17 13:46 ` Harry Yoo
2026-09-17 14:10   ` Harry Yoo
2026-09-17 15:29     ` Peter Zijlstra
2026-09-18  7:04       ` Namhyung Kim
2026-09-17 16:25     ` Tim Menninger
2026-09-18  7:08       ` Vlastimil Babka (SUSE)
2026-09-21 15:31         ` Harry Yoo
2026-09-21 23:27           ` Tim Menninger
2026-09-18  7:14     ` Namhyung Kim

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox