Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: "Chuck Lever" <cel@kernel.org>
To: "J Louis Kaplan" <Louis.Kaplan@arm.com>,
	"Dai Ngo" <dai.ngo@oracle.com>,
	"Jeff Layton" <jlayton@kernel.org>, NeilBrown <neil@brown.name>,
	"Olga Kornievskaia" <okorniev@redhat.com>,
	"Tom Talpey" <tom@talpey.com>
Cc: linux-nfs@vger.kernel.org, linux-rdma@vger.kernel.org,
	"Anna Schumaker" <anna@kernel.org>,
	"Jason Gunthorpe" <jgg@ziepe.ca>,
	"Leon Romanovsky" <leon@kernel.org>,
	"Trond Myklebust" <trondmy@kernel.org>
Subject: Re: [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled
Date: Tue, 06 Oct 2026 09:47:25 -0400	[thread overview]
Message-ID: <cbc3ff79-4866-4433-a01b-3d5eeab1962b@app.fastmail.com> (raw)
In-Reply-To: <20261006090028.3412544-1-Louis.Kaplan@arm.com>



On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> For NFSoRDMA using RoCE, and with IOMMU passthrough enabled, use of a
> 64KiB base page size was found to provide higher NFS read throughput
> than a 4KiB base page size. This patch series recovers some of that
> throughput with a 4k base page size by opportunistic use of large folios
> for the svc reply buffer.

> LLM Usage
> =========
>
> The idea of folio usage here was human-generated from observations of
> performance differences between 4KiB and 64KiB page size kernels.

Actually a similar approach has been tried before, but on x86_64:

https://lore.kernel.org/linux-nfs/20260319133610.2556826-1-cel@kernel.org/

And it had to be reverted:

- Mike Snitzer's report (4 Jun 2026):
    https://lore.kernel.org/linux-nfs/aiHlPmeZq3WgMwoJ@kernel.org/

- Jonathan Flynn's benchmark and perf data (5 Jun 2026), showing server
  CPU going from 8.54% to 76.35% with the commit present:
    https://lore.kernel.org/linux-nfs/3cb119b4b2a8aada30c0c60286778a54@mail.gmail.com/

These two are the Closes: targets in the revert, which is a39f0ce0c9da
upstream and 284d5ba931a5 in stable.

Lock contention analysis and the decision to revert

- "[PATCH] svcrdma: Cap Read sink allocations at PAGE_ALLOC_COSTLY_ORDER"
  (5 Jun 2026), whose description carries the zone->lock analysis:
    https://lore.kernel.org/linux-nfs/20260606035722.83175-1-cel@kernel.org/
- Flynn's results for that fix (25.4 GiB/s, against 30.3 regressed and
  73.9 reverted):
    https://lore.kernel.org/linux-nfs/65a2cdb132b0c28e69a29955e3bd37e7@mail.gmail.com/
- Your reply saying the two failed fixes mean 18755b8c2f24 has to be
  reverted (6 Jun 2026):
    https://lore.kernel.org/linux-nfs/096a2b91-7a19-48da-a06a-dc60e7150956@app.fastmail.com/

The earlier failed fix, "[PATCH] svcrdma: Avoid direct reclaim when
allocating Read sink buffers", is at
    https://lore.kernel.org/linux-nfs/20260605223118.75092-1-cel@kernel.org/.

So: Yes, we would like to ensure that svcrdma's RDMA Read buffers (at
least, and maybe RDMA Write) are contiguous so that the overhead of a
DMA mapping operation will be low. I'm not expert enough with the page
allocator's free path to address the lock contention problem, however.


-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

  parent reply	other threads:[~2026-10-06 13:47 UTC|newest]

Thread overview: 16+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
2026-10-06 13:49   ` Chuck Lever
2026-10-08 22:44     ` J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
2026-10-06 13:55   ` Chuck Lever
2026-10-08 22:54     ` J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
2026-10-06 14:00   ` Chuck Lever
2026-10-08 22:58     ` J Louis Kaplan
2026-10-09 15:28       ` Chuck Lever
2026-10-06 13:47 ` Chuck Lever [this message]
2026-10-08 22:50   ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=cbc3ff79-4866-4433-a01b-3d5eeab1962b@app.fastmail.com \
    --to=cel@kernel.org \
    --cc=Louis.Kaplan@arm.com \
    --cc=anna@kernel.org \
    --cc=dai.ngo@oracle.com \
    --cc=jgg@ziepe.ca \
    --cc=jlayton@kernel.org \
    --cc=leon@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=neil@brown.name \
    --cc=okorniev@redhat.com \
    --cc=tom@talpey.com \
    --cc=trondmy@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox