Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Harry Yoo <harry@kernel.org>
To: Tim Menninger <tmenninger@everpuredata.com>
Cc: linux-mm@kvack.org, Chuck Lever <cel@kernel.org>,
	 linux-nfs@vger.kernel.org, Jon Curley <jcurley@everpuredata.com>,
	 Eric Badger <ebadger@everpuredata.com>,
	Vlastimil Babka <vbabka@kernel.org>,
	 Andrew Morton <akpm@linux-foundation.org>,
	Hao Li <hao.li@linux.dev>, Christoph Lameter <cl@gentwo.org>,
	 David Rientjes <rientjes@google.com>,
	Roman Gushchin <roman.gushchin@linux.dev>
Subject: Re: [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA
Date: Thu, 17 Sep 2026 14:46:21 +0100	[thread overview]
Message-ID: <aqvq8-6IJBDer90O@thinkstation> (raw)
In-Reply-To: <20260916232227.4098143-1-tmenninger@everpuredata.com>

Hi Tim and Chuck, thanks for reporting this to linux-mm.
Will take a look at this but let me Cc SLAB ALLOCATOR folks here first.

-- 
Cheers,
Harry / Hyeonggon

On Wed, Sep 16, 2026 at 11:22:27PM +0000, Tim Menninger wrote:
> Chuck Lever suggested I bring this to linux-mm after we found some SLUB
> measurements we could not account for while investigating an NFS/RDMA
> throughput regression.
> 
> The original NFS discussion is here for context:
> 
>     https://lore.kernel.org/linux-nfs/0e2f9688-097c-4bb9-a7ec-82b41bdb3653@slotpi15m67/
> 
> The NFS regression itself has been separated from this issue. What
> remains interesting here is the behavior of the nfs_page slab cache on
> this machine, particularly the measurements with and without perf lock.
> 
> The system is:
> 
>     Intel Xeon Silver 4516Y+
>     2 sockets
>     24 cores/socket
>     2 threads/core
>     96 logical CPUs
> 
>     NUMA node0 CPUs: 0-23,48-71
>     NUMA node1 CPUs: 24-47,72-95
>
> The workload is a high-throughput NFS/RDMA direct-read workload using
> 1 MiB I/O, 160 threads, and iodepth 64.
> 
> The relevant debug options are all disabled:
> 
>     # CONFIG_KASAN is not set
>     # CONFIG_PROVE_LOCKING is not set
>     # CONFIG_LOCK_STAT is not set
>     # CONFIG_DEBUG_SPINLOCK is not set
>     # CONFIG_DEBUG_LIST is not set
> 
> All measurements below were collected on the unpatched base kernel:
> 
>     $ git rev-parse HEAD
>     940de590b839f71d6dc846160534bf202401b8b7
> 
>     $ uname -r
>     7.3.0-rc1-mainline+
> 
> The initial observation was a high apparent contention rate on the
> nfs_page slab's list_lock. In 10-second perf-lock captures I was seeing
> roughly 4.5M contended acquisitions and about 170-185 us of average
> reported wait.
> 
> Chuck reproduced a similar acquisition rate on a single-node EPYC
> system, but saw only about 7 ns average wait and fewer than 100
> cmpxchg_double_fail events over a corresponding interval. He suggested
> checking cmpxchg_double_fail because __slab_free() drops list_lock and
> retries when the freelist cmpxchg fails.
>
> I repeated the measurements in three placement configurations:
> 
>     A. workload unpinned, CQs all on node0
>     B. workload pinned to node0, CQs all on node0
>     C. workload unpinned, CQs balanced across the nodes
> 
> Without perf lock, throughput is similar in all three:
> 
>     A. unpinned / CQs node0:       ~45.5 GB/s
>     B. node0 pinned / CQs node0:   ~45.7 GB/s
>     C. unpinned / balanced CQs:    ~45.5 GB/s
> 
> During the perf-lock captures, throughput is approximately 25 GB/s.
> 
> For each instrumented 10-second window I ran:
> 
>     sudo perf lock record -a -g -o "$D/perf-locks.data" -- sleep 10
>     mpstat -P ALL 1 10
> 
> Before and after the same window I sampled the counters under:
> 
>     /sys/kernel/slab/nfs_page/
> 
> I also collected separate 10-second counter and mpstat windows under
> the same workload configurations without perf lock.
> 
> The resulting slab counter deltas were:
> 
>                             A               B                 C
>                      unpinned/node0    node0/node0   unpinned/balanced
> 
>     free_fastpath
>     instrumented        655,917,745    457,936,140       328,670,030
>     uninstrumented       34,906,233    119,984,226        59,469,931
> 
>     free_slowpath
>     instrumented        236,735,724      7,720,917       270,716,506
>     uninstrumented       85,039,753          2,814        60,145,870
> 
>     sheaf_flush
>     instrumented         39,842,700     42,706,800         8,341,440
>     uninstrumented        1,286,400     13,487,700           887,700
> 
>     barn_put_fail
>     instrumented            663,994        711,745           139,037
>     uninstrumented           21,464        224,762            14,765
> 
>     barn_get_fail
>     instrumented          4,609,925        840,344         4,651,225
>     uninstrumented        1,438,643        224,787         1,017,006
> 
>     cmpxchg_double_fail
>     instrumented             70,524          8,077            28,062
>     uninstrumented            7,471            713             1,503
> 
>     alloc_slowpath
>     all cases                     0              0                 0
> 
> The SLUB counter profile changes substantially with perf lock despite
> the lower NFS throughput, and the exact mix depends strongly on
> placement.
> 
> The uninstrumented node0/node0 run also reproduces the barn/sheaf
> relationship Chuck pointed out earlier. The nfs_page sheaf capacity is
> 60:
> 
>     sheaf_flush / 60 = 13,487,700 / 60 = 224,795
>     barn_put_fail                      = 224,762
> 
> The placement dependence of free_slowpath is also large. It falls from
> about 85M events per 10 seconds in the unpinned/node0 case to 2,814 in
> the node0/node0 case.
> 
> The reported perf-lock result, however, is similar across all three
> placements:
> 
>                          contentions    total wait    average wait
>     unpinned/node0         4,782,839      14.39 min      180.47 us
>     node0/node0            4,582,012      14.13 min      185.05 us
>     unpinned/balanced      4,572,196      12.81 min      168.14 us
> 
> I also revisited an inconsistency Chuck noticed in my earlier
> measurements. Previously I had compared aggregate perf-lock wait from
> one 10-second capture with CPU utilization measured during a different
> window.
> 
> I now have paired 10-second mpstat samples for each placement, with and
> without perf lock. The node values below are averages of the per-CPU
> %idle values for the CPUs in each NUMA node:
> 
>                                   system-wide       node0       node1
>                                   %idle             %idle       %idle
> 
>     A. unpinned / CQs node0
>        uninstrumented                 31.46           8.0        54.6
>        instrumented                    3.60           0.06        7.1
> 
>     B. node0 pinned / CQs node0
>        uninstrumented                 84.26          69.2        99.4
>        instrumented                   12.07          10.4        13.8
> 
>     C. unpinned / balanced CQs
>        uninstrumented                 66.86          66.7        66.9
>        instrumented                   14.83          19.0        10.6
> 
> This resolves the accounting inconsistency in my earlier measurements.
> The large aggregate perf-lock wait and high idle percentage had come
> from different windows. In the aligned samples, the system is much
> busier during the perf-lock capture than in the corresponding
> uninstrumented run.
> 
> I am still unsure how representative the reported ~170-185 us average
> wait is of the uninstrumented workload.
> 
> The remaining number I am less sure how to interpret is
> cmpxchg_double_fail. In the uninstrumented windows I see:
> 
>     unpinned / CQs node0:       7,471 / 10 sec
>     node0 pinned / CQs node0:     713 / 10 sec
>     unpinned / balanced CQs:    1,503 / 10 sec
> 
> compared with fewer than 100 in Chuck's test.
> 
> I understand that cmpxchg_double_fail counts failed slab freelist
> updates rather than failed logical frees, so I am not sure what the
> appropriate denominator is here. In particular, the node0/node0 case
> still has 713 failures while sustaining full throughput and only 2,814
> free_slowpath events over the interval.
> 
> My questions are:
> 
> 1. Do the uninstrumented cmpxchg_double_fail rates above look abnormal
>    for this workload/topology, or are they within the range one would
>    expect from this degree of concurrency and NUMA placement?
> 
> 2. Is there a less invasive way you would recommend measuring the
>    nfs_page list_lock/freelist contention? I would like to distinguish
>    the steady-state behavior from what is observed during the perf-lock
>    capture.
> 
> Thanks,
> Tim


  reply	other threads:[~2026-09-17 13:46 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-16 23:22 [SLUB] nfs_page cmpxchg_double_fail and perf lock perturbation on dual-socket NFS/RDMA Tim Menninger
2026-09-17 13:46 ` Harry Yoo [this message]
2026-09-17 14:10   ` Harry Yoo
2026-09-17 15:29     ` Peter Zijlstra
2026-09-18  7:04       ` Namhyung Kim
2026-09-17 16:25     ` Tim Menninger
2026-09-18  7:08       ` Vlastimil Babka (SUSE)
2026-09-21 15:31         ` Harry Yoo
2026-09-21 23:27           ` Tim Menninger
2026-09-18  7:14     ` Namhyung Kim

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqvq8-6IJBDer90O@thinkstation \
    --to=harry@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=cel@kernel.org \
    --cc=cl@gentwo.org \
    --cc=ebadger@everpuredata.com \
    --cc=hao.li@linux.dev \
    --cc=jcurley@everpuredata.com \
    --cc=linux-mm@kvack.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=rientjes@google.com \
    --cc=roman.gushchin@linux.dev \
    --cc=tmenninger@everpuredata.com \
    --cc=vbabka@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox