Linux NFS development
 help / color / mirror / Atom feed
From: Tim Menninger <tmenninger@everpuredata.com>
To: Chuck Lever <cel@kernel.org>
Cc: Trond Myklebust <trondmy@kernel.org>,
	Anna Schumaker <anna@kernel.org>, Tejun Heo <tj@kernel.org>,
	Lai Jiangshan <jiangshanlai@gmail.com>,
	linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org,
	Eric Badger <ebadger@everpuredata.com>,
	Jon Curley <jcurley@everpuredata.com>
Subject: Re: [PATCH RFC 6/8] SUNRPC: Reduce rpciod workqueue contention
Date: Fri,  4 Sep 2026 23:23:42 +0000	[thread overview]
Message-ID: <20260904232342.1907377-1-tmenninger@everpuredata.com> (raw)
In-Reply-To: <336623df-ead4-4802-91c5-aa3e785f2971@slotpi15m67>

I realized the bad state correlates with where the RDMA CQs' completion
work runs.

This client has 16 RDMA transports and 32 CQs. The CQs receive 32
consecutive completion vectors out of the mlx5 device's 63-vector ring.
Depending on the global round-robin starting point, all 32 CQ workers can
execute on NUMA node 0, or they can split 17/15 between the two NUMA nodes.

On unpatched mainline (940de590b839 without your patch set), I tested five
consecutive allocation windows. On three of those, all 32 CQs executed on
NUMA node 0. On the other two they fell with a 17/15 split. All five
sustained ~46 GB/s.

With your v2 applied, all-local windows fall to roughly 24-28 GB/s while
split windows sustain ~46 GB/s.

I also tested this by changing only mlx5 IRQ affinity. I redirected
completion vectors 20-34 from node 0 to CPUs 24-38 on node 1. CQ allocation
windows 5-36 and 6-37, which would otherwise be entirely local and slow,
then sustained ~46 GB/s with their CQs split 17/15 across the nodes.

So the CQ allocation determines whether the regression is exposed, but
mainline is insensitive to that placement. The series introduces the
performance sensitivity. This also explains why v2 appeared
nondeterministic with respect to the rpciod affinity scope alone.

  reply	other threads:[~2026-09-04 23:23 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31 18:21 [PATCH RFC 0/8] Reduce lock contention in the NFS client Chuck Lever
2026-08-31 18:21 ` [PATCH RFC 1/8] SUNRPC: Use atomic_t for XID allocation Chuck Lever
2026-08-31 18:21 ` [PATCH RFC 2/8] SUNRPC: Execute initial async RPC states in caller's context Chuck Lever
2026-08-31 18:21 ` [PATCH RFC 3/8] SUNRPC: Split recv_lock out of xprt->queue_lock Chuck Lever
2026-08-31 18:22 ` [PATCH RFC 4/8] Set WQ_SYSFS on key NFS-related workqueues Chuck Lever
2026-08-31 18:22 ` [PATCH RFC 5/8] workqueue: Export the functions needed for WQ attribute modification Chuck Lever
2026-08-31 18:22 ` [PATCH RFC 6/8] SUNRPC: Reduce rpciod workqueue contention Chuck Lever
2026-09-02 20:40   ` Tim Menninger
2026-09-03 13:33     ` Chuck Lever
2026-09-03 23:50       ` Tim Menninger
2026-09-04 14:13         ` Chuck Lever
2026-09-04 23:23           ` Tim Menninger [this message]
2026-08-31 18:22 ` [PATCH RFC 7/8] NFS: Reduce nfsiod " Chuck Lever
2026-08-31 18:22 ` [PATCH RFC 8/8] SUNRPC: Reduce xprtiod " Chuck Lever

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260904232342.1907377-1-tmenninger@everpuredata.com \
    --to=tmenninger@everpuredata.com \
    --cc=anna@kernel.org \
    --cc=cel@kernel.org \
    --cc=ebadger@everpuredata.com \
    --cc=jcurley@everpuredata.com \
    --cc=jiangshanlai@gmail.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=tj@kernel.org \
    --cc=trondmy@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox