Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
* [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled
@ 2026-10-06  9:00 J Louis Kaplan
  2026-10-06  9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
                   ` (6 more replies)
  0 siblings, 7 replies; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-06  9:00 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan

For NFSoRDMA using RoCE, and with IOMMU passthrough enabled, use of a
64KiB base page size was found to provide higher NFS read throughput
than a 4KiB base page size. This patch series recovers some of that
throughput with a 4k base page size by opportunistic use of large folios
for the svc reply buffer.

To better understand the performance implications of this patch series,
an ablation test was run which added a single line to limit bio_vec
coalescing to a single 4 KiB page. Large-folio allocation was kept in
place. This brought direct-read throughput in configuration A (see
below) back to approximately plain-kernel levels. This indicates that
coalescing strongly contributes to the observed throughput gains with
passthrough enabled.

Tests with patch series on Linux 7.2 commit 8d3ae59288f1
========================================================

(The below tests were done before minor internal revisions.)

Performance:
fio read benchmark results on an NFSv4 mount are shown below with
various tested boot commandlines on the NFS server. NFSoRDMA was used
for all tests.

Configurations B-D explicitly disable passthrough as I wanted to confirm
that no significant regressions to the passthrough=0 are introduced by
this prototype. I plan to test more passthrough=1 configurations ASAP.

- Throughput and CPU usage were averaged over 60s runs.
- Request size was 1 MiB in all tests.
- Except where noted, 64 fio jobs and 64 nfsd threads were used.
- Clients used O_DIRECT in all tests.
- `Mode` in tables below is NFS server's access mode.
- Server throughput/core is mean throughput divided by mean measured
  server core equivalents, where 100% processor usage represents one core.
- Changes are relative to the plain kernel:
  100 * (patched - plain) / plain. Positive values mean increases.
- Percentages are calculated from means rounded to two decimal places.
- Results from the first buffered runs were discarded so as to
  only compare filled page cache performance.
- Direct results are average of two runs, except direct
  configuration E which is the average of four runs.
- Hardware used: Nvidia Grace (144 Arm Neoverse V2 cores),
  Broadcom 400G NICs connected to NUMA 1 on client and server

NFS server command lines tested:
A: cpufreq.off=1 cpuidle.off=1 iommu.passthrough=1 iommu.strict=0
   kpti=off maxcpus=72 mem=480g rcu_nocb_poll selinux=0
B: cpufreq.off=1 cpuidle.off=1 iommu.passthrough=0 iommu.strict=0
   kpti=off maxcpus=72 mem=480g rcu_nocb_poll selinux=0
C: cpufreq.off=1 cpuidle.off=1 iommu.passthrough=0 iommu.strict=1
   kpti=off maxcpus=72 mem=480g rcu_nocb_poll selinux=0
D: cpufreq.off=0 cpuidle.off=0 iommu.passthrough=0 iommu.strict=1
            maxcpus=72 mem=480g
E: <empty command line>
F: <empty command line> with only 8 fio jobs and 8 nfsd threads.

Configuration   Mode      Throughput change   Server throughput/core change
--------------  --------  ------------------  -----------------------------
A               Direct               +136.8%                         +95.5%
B               Direct                 +6.7%                         +74.9%
C               Direct                +11.3%                         +62.4%
D               Direct                 -1.9%                         +27.6%
E               Direct                 -2.6%                         +28.2%
F               Direct                +18.8%                         +21.8%

More benchmarks are planned for a better signal/noise ratio.

--------------  --------  ------------------  -----------------------------
A               Buffered               +0.2%                         +33.6%
B               Buffered                0.0%                         +28.7%
C               Buffered                0.0%                         +11.4%
D               Buffered                0.0%                          -8.0%
E               Buffered                0.0%                          +7.4%
F               Buffered                0.0%                         -12.0%

Buffered mode saw no loss of throughput but there was an increased
CPU usage observed in some configurations.

Functional testing:
Seven hour NFSv4 over RDMA stress run exercised buffered and direct NFSD
modes using fsstress, fsx, and fio, with concurrency up to 64 workers.
All 879 completed workloads passed.

Tests with patch series on nfsd-testing commit 32eb1a60b456
===========================================================

Performance:

Set-up as above.

Configuration   Mode      Throughput change   Server throughput/core change
--------------  --------  ------------------  -----------------------------
A               Direct               +104.5%                         +70.2%
E               Direct                 +3.4%                         +20.8%
A               Buffered               +1.0%                         +35.6%
E               Buffered                0.0%                         +15.5%

Functional testing:
A planned seven-hour NFSv4 over RDMA stress run stopped after
approximately one hour because the test harness stops on server kernel
warnings. It detected two correctable PCIe AER reports (RxErr and
BadTLP) from the root port upstream of the server's RoCE NIC.

Before the stop, 315 test workloads completed with exit status zero,
including fio checksum readbacks. The cause of the AER reports and any
relationship to this series have not been established.

Tests with patch series on nfsd-testing commit 56589cdb58819
============================================================

Functional testing:
On the rebased nfsd-testing tree, 490 workloads completed successfully.
They covered buffered and direct NFSD modes with 1, 8, 32, and 64
workers, using fsstress, fsx, and fio write/checksum readback tests.
The planned seven-hour run stopped after 1h37 when the harness detected
two USB hub descriptor timeouts that appear unrelated to the nfs stack.
A full 7 hour test will be re-attempted again as soon as practicable.

Known issues and follow-up
==========================

- Complete a seven-hour stress run on nfsd-testing and investigate
  whether the PCIe AER reports recur.
- Address memory-registration capacity limits with coalesced vectors
  in the force_mr and Xen paths.
- Gather more comprehensive performance data, including write data

LLM Usage
=========

The idea of folio usage here was human-generated from observations of
performance differences between 4KiB and 64KiB page size kernels. An
LLM was used for familiarization with the NFS and related subsystems
and for generating some of the code after lengthy prompted technical
discussions. An LLM was used for multiple rounds of review, including
for the cover letter; suggested changes were manually verified. It also
helped develop scripts to invoke the various test suites mentioned above
and analyze resulting logs.

Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM

J Louis Kaplan (6):
  sunrpc: Add helpers to build bvecs from contiguous pages
  nfsd: Coalesce contiguous pages for direct reads
  svcrdma: Coalesce contiguous pages when mapping replies
  svcrdma: Coalesce contiguous pages in RDMA Write chunks
  svcrdma: Coalesce contiguous pages in RDMA Read chunks
  sunrpc: Allocate svc request pages from large folios

 drivers/infiniband/core/rw.c          |   7 +-
 fs/nfsd/vfs.c                         |  11 +-
 include/linux/sunrpc/svc.h            |  90 +++++++++++++
 include/linux/sunrpc/svc_rdma.h       |  15 +++
 net/sunrpc/svc_xprt.c                 |  44 ++++++-
 net/sunrpc/xprtrdma/svc_rdma_rw.c     | 178 +++++++++++++++-----------
 net/sunrpc/xprtrdma/svc_rdma_sendto.c | 112 +++++++++-------
 7 files changed, 327 insertions(+), 130 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 15+ messages in thread

* [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages
  2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
@ 2026-10-06  9:00 ` J Louis Kaplan
  2026-10-06 13:49   ` Chuck Lever
  2026-10-06  9:00 ` [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads J Louis Kaplan
                   ` (5 subsequent siblings)
  6 siblings, 1 reply; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-06  9:00 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan

Add helpers to describe contiguous runs with multipage bvecs. These
functions have no callers. They are used in later commits.

Accept a caller-supplied maximum bvec length so transport callers
can enforce direct memory access (DMA) mapping and device segment
length limits. Also support counting entries without populating
an array so callers can size their allocations using the same
coalescing rules.

Coalescing requires contiguous physical memory and page descriptors,
and is disabled under KMSAN. Xen constraint handling remains WIP.

Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
 include/linux/sunrpc/svc.h | 90 ++++++++++++++++++++++++++++++++++++++
 1 file changed, 90 insertions(+)

diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index dfadea50e0c6d..e651c2312a912 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -11,6 +11,7 @@
 #ifndef SUNRPC_SVC_H
 #define SUNRPC_SVC_H
 
+#include <linux/bvec.h>
 #include <linux/in.h>
 #include <linux/in6.h>
 #include <linux/sunrpc/types.h>
@@ -376,6 +377,95 @@ static inline void svc_thread_init_status(struct svc_rqst *rqstp, int err)
 		kthread_exit(1);
 }
 
+/**
+ * svc_pages_to_bvec - build one bvec from physically contiguous pages
+ * @bv: bio_vec to initialize, or NULL to only measure the next extent
+ * @pages: first page-array slot in the range
+ * @nr_pages: number of page-array slots available from @pages
+ * @offset: byte offset in @pages[0]
+ * @len: maximum byte count to describe
+ * @max_bvec_len: maximum byte count in one bvec
+ *
+ * Return: The number of bytes described by @bv. The scan stops at a missing
+ * or physically non-contiguous page.
+ */
+static inline unsigned int
+svc_pages_to_bvec(struct bio_vec *bv, struct page *const *pages,
+		  unsigned int nr_pages, unsigned int offset,
+		  unsigned int len, unsigned int max_bvec_len)
+{
+	unsigned long first_pfn;
+	unsigned int bytes, used = 1;
+
+	if (WARN_ON_ONCE(!nr_pages || offset >= PAGE_SIZE || !len ||
+			 !max_bvec_len || !pages[0]))
+		return 0;
+	first_pfn = page_to_pfn(pages[0]);
+
+	bytes = min3(len, max_bvec_len,
+		     (unsigned int)(PAGE_SIZE - offset));
+	while (bytes < len && bytes < max_bvec_len && used < nr_pages) {
+		// TODO: handle Xen merge constraints, see `bvec_try_merge_page`
+		//	 for reference, or unify helper functionality
+		if (IS_ENABLED(CONFIG_KMSAN))
+			break;
+
+		if (!pages[used] ||
+		    page_to_pfn(pages[used]) != first_pfn + used ||
+		    pages[0] + used != pages[used])
+			break;
+		used++;
+		bytes = min3(len, max_bvec_len,
+			     (used << PAGE_SHIFT) - offset);
+	}
+
+	if (bv)
+		bvec_set_page(bv, pages[0], bytes, offset);
+	return bytes;
+}
+
+/**
+ * svc_pages_to_bvecs - build bvecs for a service page-array range
+ * @bvecs: bio_vec array to populate, or NULL to only count the entries
+ * @pages: first page-array slot in the range
+ * @nr_pages: number of page-array slots available from @pages
+ * @offset: byte offset in @pages[0]
+ * @len: byte count to describe
+ * @max_bvec_len: maximum byte count in one bvec
+ *
+ * If @bvecs is not NULL, it must have room for the count returned by a
+ * preceding count-only call.
+ *
+ * Return: The number of populated or required bvecs, or zero if the range
+ * cannot be described from the supplied page-array slots.
+ */
+static inline unsigned int
+svc_pages_to_bvecs(struct bio_vec *bvecs, struct page *const *pages,
+		   unsigned int nr_pages, unsigned int offset,
+		   unsigned int len, unsigned int max_bvec_len)
+{
+	unsigned int remaining = len;
+	unsigned int nents = 0;
+
+	while (remaining && nr_pages) {
+		struct bio_vec *bv = bvecs ? &bvecs[nents] : NULL;
+		unsigned int advanced;
+		unsigned int bytes;
+
+		bytes = svc_pages_to_bvec(bv, pages, nr_pages, offset,
+					  remaining, max_bvec_len);
+		if (!bytes)
+			break;
+		advanced = (offset + bytes) >> PAGE_SHIFT;
+		offset = offset_in_page(offset + bytes);
+		pages += advanced;
+		nr_pages -= advanced;
+		remaining -= bytes;
+		nents++;
+	}
+	return remaining ? 0 : nents;
+}
+
 struct svc_deferred_req {
 	u32			prot;	/* protocol (UDP or TCP) */
 	bool			secure;	/* RQ_SECURE of the original request */
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 15+ messages in thread

* [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads
  2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
  2026-10-06  9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
@ 2026-10-06  9:00 ` J Louis Kaplan
  2026-10-06  9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
                   ` (4 subsequent siblings)
  6 siblings, 0 replies; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-06  9:00 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan

Use the previously added helper functions for nfsd direct read buffers.

Before this change, the direct-read path builds one bvec per base page,
even when the destination pages are physically contiguous.

Coalesce these pages into multipage bvecs to reduce the number of
entries traversed when advancing the filesystem's read iterator.

Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
 fs/nfsd/vfs.c | 11 ++++++++---
 1 file changed, 8 insertions(+), 3 deletions(-)

diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index c7dd94697f2ff..329bb7d410765 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1124,11 +1124,16 @@ nfsd_direct_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
 	v = 0;
 	total = dio_end - dio_start;
 	while (total && v < rqstp->rq_maxpages && page < rqstp->rq_page_end) {
-		len = min_t(size_t, total, PAGE_SIZE);
-		bvec_set_page(&rqstp->rq_bvec[v], *page, len, 0);
+		len = svc_pages_to_bvec(&rqstp->rq_bvec[v],
+					page,
+					rqstp->rq_page_end - page,
+					0, min_t(size_t, total, UINT_MAX),
+					UINT_MAX);
+		if (!len)
+			break;
 
 		total -= len;
-		++page;
+		page += DIV_ROUND_UP(len, PAGE_SIZE);
 		++v;
 	}
 	rqstp->rq_next_page = page;
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 15+ messages in thread

* [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies
  2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
  2026-10-06  9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
  2026-10-06  9:00 ` [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads J Louis Kaplan
@ 2026-10-06  9:00 ` J Louis Kaplan
  2026-10-06 13:55   ` Chuck Lever
  2026-10-06  9:00 ` [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks J Louis Kaplan
                   ` (3 subsequent siblings)
  6 siblings, 1 reply; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-06  9:00 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan

Mapping reply pagelists one page at a time uses a separate direct
memory access (DMA) mapping and scatter/gather entry (SGE) for
each base page, even when those pages are physically contiguous.

Instead, map contiguous runs as multipage bvecs, capped by the device's
mapping and segment length limits. Use the same coalescing rules
when counting SGEs for the pull-up decision, avoiding unnecessary
copying when the coalesced reply fits the Send SGE limit.

Virtual-mapping devices bypass device-limit check.

Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
 include/linux/sunrpc/svc_rdma.h       |  15 ++++
 net/sunrpc/xprtrdma/svc_rdma_sendto.c | 112 +++++++++++++++-----------
 2 files changed, 82 insertions(+), 45 deletions(-)

diff --git a/include/linux/sunrpc/svc_rdma.h b/include/linux/sunrpc/svc_rdma.h
index 71bd5bcc5ce2a..c296bf8b8b1e3 100644
--- a/include/linux/sunrpc/svc_rdma.h
+++ b/include/linux/sunrpc/svc_rdma.h
@@ -42,6 +42,7 @@
 
 #ifndef SVC_RDMA_H
 #define SVC_RDMA_H
+#include <linux/dma-mapping.h>
 #include <linux/llist.h>
 #include <linux/sunrpc/xdr.h>
 #include <linux/sunrpc/svcsock.h>
@@ -134,6 +135,20 @@ static inline struct svcxprt_rdma *svc_rdma_rqst_rdma(struct svc_rqst *rqstp)
 	return container_of(xprt, struct svcxprt_rdma, sc_xprt);
 }
 
+static inline unsigned int
+svc_rdma_max_bvec_len(const struct svcxprt_rdma *rdma)
+{
+	struct ib_device *device = rdma->sc_cm_id->device;
+
+	if (ib_uses_virt_dma(device))
+		return UINT_MAX;
+
+	/* Respect both DMA mapping and RDMA device segment limits. */
+	return min_t(size_t,
+		     dma_max_mapping_size(device->dma_device),
+		     ib_dma_max_seg_size(device));
+}
+
 /*
  * Default connection parameters
  */
diff --git a/net/sunrpc/xprtrdma/svc_rdma_sendto.c b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
index efa352ddf9d71..9f2e3ee3c2936 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_sendto.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
@@ -260,10 +260,8 @@ static void svc_rdma_send_ctxt_unmap(struct svcxprt_rdma *rdma,
 		trace_svcrdma_dma_unmap_page(&ctxt->sc_cid,
 					     ctxt->sc_sges[i].addr,
 					     ctxt->sc_sges[i].length);
-		ib_dma_unmap_page(device,
-				  ctxt->sc_sges[i].addr,
-				  ctxt->sc_sges[i].length,
-				  DMA_TO_DEVICE);
+		ib_dma_unmap_bvec(device, ctxt->sc_sges[i].addr,
+				  ctxt->sc_sges[i].length, DMA_TO_DEVICE);
 	}
 }
 
@@ -730,44 +728,45 @@ svc_rdma_encode_reply_chunk(struct svc_rdma_recv_ctxt *rctxt,
 }
 
 struct svc_rdma_map_data {
-	struct svcxprt_rdma		*md_rdma;
 	struct svc_rdma_send_ctxt	*md_ctxt;
+	unsigned int			md_max_bvec_len;
 };
 
 /**
- * svc_rdma_page_dma_map - DMA map one page
+ * svc_rdma_bvec_dma_map - DMA map one physically contiguous bvec
  * @data: pointer to arguments
- * @page: struct page to DMA map
- * @offset: offset into the page
- * @len: number of bytes to map
+ * @bv: contiguous range to DMA map
  *
  * Returns:
  *   %0 if DMA mapping was successful
- *   %-EIO if the page cannot be DMA mapped
+ *   %-EIO if the range cannot be DMA mapped
  */
-static int svc_rdma_page_dma_map(void *data, struct page *page,
-				 unsigned long offset, unsigned int len)
+static int svc_rdma_bvec_dma_map(void *data, struct bio_vec *bv)
 {
 	struct svc_rdma_map_data *args = data;
-	struct svcxprt_rdma *rdma = args->md_rdma;
 	struct svc_rdma_send_ctxt *ctxt = args->md_ctxt;
+	struct svcxprt_rdma *rdma = ctxt->sc_rdma;
 	struct ib_device *dev = rdma->sc_cm_id->device;
 	dma_addr_t dma_addr;
 
+	if (bv->bv_len > args->md_max_bvec_len)
+		return -EIO;
+	if (WARN_ON_ONCE(ctxt->sc_send_wr.num_sge >= rdma->sc_max_send_sges))
+		return -EIO;
 	++ctxt->sc_cur_sge_no;
 
-	dma_addr = ib_dma_map_page(dev, page, offset, len, DMA_TO_DEVICE);
+	dma_addr = ib_dma_map_bvec(dev, bv, DMA_TO_DEVICE);
 	if (ib_dma_mapping_error(dev, dma_addr))
 		goto out_maperr;
 
-	trace_svcrdma_dma_map_page(&ctxt->sc_cid, dma_addr, len);
+	trace_svcrdma_dma_map_page(&ctxt->sc_cid, dma_addr, bv->bv_len);
 	ctxt->sc_sges[ctxt->sc_cur_sge_no].addr = dma_addr;
-	ctxt->sc_sges[ctxt->sc_cur_sge_no].length = len;
+	ctxt->sc_sges[ctxt->sc_cur_sge_no].length = bv->bv_len;
 	ctxt->sc_send_wr.num_sge++;
 	return 0;
 
 out_maperr:
-	trace_svcrdma_dma_map_err(&ctxt->sc_cid, dma_addr, len);
+	trace_svcrdma_dma_map_err(&ctxt->sc_cid, dma_addr, bv->bv_len);
 	return -EIO;
 }
 
@@ -776,20 +775,18 @@ static int svc_rdma_page_dma_map(void *data, struct page *page,
  * @data: pointer to arguments
  * @iov: kvec to DMA map
  *
- * ib_dma_map_page() is used here because svc_rdma_dma_unmap()
- * handles DMA-unmap and it uses ib_dma_unmap_page() exclusively.
- *
  * Returns:
  *   %0 if DMA mapping was successful
  *   %-EIO if the iovec cannot be DMA mapped
  */
 static int svc_rdma_iov_dma_map(void *data, const struct kvec *iov)
 {
+	struct bio_vec bv;
+
 	if (!iov->iov_len)
 		return 0;
-	return svc_rdma_page_dma_map(data, virt_to_page(iov->iov_base),
-				     offset_in_page(iov->iov_base),
-				     iov->iov_len);
+	bvec_set_virt(&bv, iov->iov_base, iov->iov_len);
+	return svc_rdma_bvec_dma_map(data, &bv);
 }
 
 /**
@@ -798,7 +795,7 @@ static int svc_rdma_iov_dma_map(void *data, const struct kvec *iov)
  * @data: pointer to arguments
  *
  * Returns:
- *   %0 if DMA mapping was successful
+ *   The number of mapped bytes if DMA mapping was successful
  *   %-EIO if DMA mapping failed
  *
  * On failure, any DMA mappings that have been already done must be
@@ -806,9 +803,11 @@ static int svc_rdma_iov_dma_map(void *data, const struct kvec *iov)
  */
 static int svc_rdma_xb_dma_map(const struct xdr_buf *xdr, void *data)
 {
-	unsigned int len, remaining;
+	struct svc_rdma_map_data *args = data;
+	unsigned int remaining;
 	unsigned long pageoff;
 	struct page **ppages;
+	unsigned int nr_pages;
 	int ret;
 
 	ret = svc_rdma_iov_dma_map(data, &xdr->head[0]);
@@ -818,15 +817,25 @@ static int svc_rdma_xb_dma_map(const struct xdr_buf *xdr, void *data)
 	ppages = xdr->pages + (xdr->page_base >> PAGE_SHIFT);
 	pageoff = offset_in_page(xdr->page_base);
 	remaining = xdr->page_len;
+	nr_pages = DIV_ROUND_UP(pageoff + remaining, PAGE_SIZE);
 	while (remaining) {
-		len = min_t(u32, PAGE_SIZE - pageoff, remaining);
-
-		ret = svc_rdma_page_dma_map(data, *ppages++, pageoff, len);
+		struct bio_vec bv;
+		unsigned int bytes = svc_pages_to_bvec(&bv, ppages,
+				nr_pages, pageoff, remaining,
+				args->md_max_bvec_len);
+		unsigned int advanced =
+				(pageoff + bytes) >> PAGE_SHIFT;
+
+		if (!bytes)
+			return -EIO;
+		ret = svc_rdma_bvec_dma_map(data, &bv);
 		if (ret < 0)
 			return ret;
 
-		remaining -= len;
-		pageoff = 0;
+		pageoff = offset_in_page(pageoff + bytes);
+		ppages += advanced;
+		nr_pages -= advanced;
+		remaining -= bytes;
 	}
 
 	ret = svc_rdma_iov_dma_map(data, &xdr->tail[0]);
@@ -836,10 +845,27 @@ static int svc_rdma_xb_dma_map(const struct xdr_buf *xdr, void *data)
 	return xdr->len;
 }
 
+static unsigned int
+svc_rdma_xb_page_sges(const struct xdr_buf *xdr,
+		      unsigned int max_bvec_len)
+{
+	const unsigned long offset = offset_in_page(xdr->page_base);
+	const unsigned int nr_pages =
+		DIV_ROUND_UP(offset + xdr->page_len, PAGE_SIZE);
+	struct page *const *ppages;
+
+	if (!xdr->page_len)
+		return 0;
+	ppages = xdr->pages + (xdr->page_base >> PAGE_SHIFT);
+	return svc_pages_to_bvecs(NULL, ppages, nr_pages, offset,
+				  xdr->page_len, max_bvec_len);
+}
+
 struct svc_rdma_pullup_data {
 	u8		*pd_dest;
 	unsigned int	pd_length;
 	unsigned int	pd_num_sges;
+	unsigned int	pd_max_bvec_len;
 };
 
 /**
@@ -848,25 +874,22 @@ struct svc_rdma_pullup_data {
  * @data: pointer to arguments
  *
  * Returns:
- *   Number of SGEs needed to Send the contents of @xdr inline
+ *   %0 if the SGE count was updated
+ *   %-EIO if the page list cannot be described
  */
 static int svc_rdma_xb_count_sges(const struct xdr_buf *xdr,
 				  void *data)
 {
 	struct svc_rdma_pullup_data *args = data;
-	unsigned int remaining;
-	unsigned long offset;
+	unsigned int page_sges =
+		svc_rdma_xb_page_sges(xdr, args->pd_max_bvec_len);
 
 	if (xdr->head[0].iov_len)
 		++args->pd_num_sges;
 
-	offset = offset_in_page(xdr->page_base);
-	remaining = xdr->page_len;
-	while (remaining) {
-		++args->pd_num_sges;
-		remaining -= min_t(u32, PAGE_SIZE - offset, remaining);
-		offset = 0;
-	}
+	if (xdr->page_len && !page_sges)
+		return -EIO;
+	args->pd_num_sges += page_sges;
 
 	if (xdr->tail[0].iov_len)
 		++args->pd_num_sges;
@@ -896,6 +919,7 @@ static int svc_rdma_check_pull_up(const struct svcxprt_rdma *rdma,
 	struct svc_rdma_pullup_data args = {
 		.pd_length	= sctxt->sc_hdrbuf.len,
 		.pd_num_sges	= 1,
+		.pd_max_bvec_len = svc_rdma_max_bvec_len(rdma),
 	};
 	int ret;
 
@@ -964,7 +988,6 @@ static int svc_rdma_xb_linearize(const struct xdr_buf *xdr,
 
 /**
  * svc_rdma_pull_up_reply_msg - Copy Reply into a single buffer
- * @rdma: controlling transport
  * @sctxt: send_ctxt for the Send WR; xprt hdr is already prepared
  * @write_pcl: Write chunk list provided by client
  * @xdr: prepared xdr_buf containing RPC message
@@ -979,8 +1002,7 @@ static int svc_rdma_xb_linearize(const struct xdr_buf *xdr,
  *   %0 if pull-up was successful
  *   %-EMSGSIZE if a buffer manipulation problem occurred
  */
-static int svc_rdma_pull_up_reply_msg(const struct svcxprt_rdma *rdma,
-				      struct svc_rdma_send_ctxt *sctxt,
+static int svc_rdma_pull_up_reply_msg(struct svc_rdma_send_ctxt *sctxt,
 				      const struct svc_rdma_pcl *write_pcl,
 				      const struct xdr_buf *xdr)
 {
@@ -1021,8 +1043,8 @@ int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
 			   const struct xdr_buf *xdr)
 {
 	struct svc_rdma_map_data args = {
-		.md_rdma	= rdma,
 		.md_ctxt	= sctxt,
+		.md_max_bvec_len = svc_rdma_max_bvec_len(rdma),
 	};
 	int ret;
 
@@ -1043,7 +1065,7 @@ int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
 	if (ret < 0)
 		return ret;
 	if (ret)
-		return svc_rdma_pull_up_reply_msg(rdma, sctxt, write_pcl, xdr);
+		return svc_rdma_pull_up_reply_msg(sctxt, write_pcl, xdr);
 
 	return pcl_process_nonpayloads(write_pcl, xdr,
 				       svc_rdma_xb_dma_map, &args);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 15+ messages in thread

* [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks
  2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
                   ` (2 preceding siblings ...)
  2026-10-06  9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
@ 2026-10-06  9:00 ` J Louis Kaplan
  2026-10-06  9:00 ` [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks J Louis Kaplan
                   ` (2 subsequent siblings)
  6 siblings, 0 replies; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-06  9:00 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan

When sending replies with remote direct memory access (RDMA)
Writes, the pagelist is represented by one bvec per base page.
This prevents contiguous pages from sharing a buffer descriptor.

Coalesce contiguous runs within each remote segment, subject to
device length limits. Count the resulting bvecs before acquiring
the context so its allocation reflects the coalesced entry count.
The counting pass leaves the payload cursor untouched.

Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
 net/sunrpc/xprtrdma/svc_rdma_rw.c | 77 +++++++++++++++++--------------
 1 file changed, 43 insertions(+), 34 deletions(-)

diff --git a/net/sunrpc/xprtrdma/svc_rdma_rw.c b/net/sunrpc/xprtrdma/svc_rdma_rw.c
index 63d00f0cc1dbd..a2e1c693b1bc1 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_rw.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_rw.c
@@ -445,46 +445,48 @@ static int svc_rdma_post_chunk_ctxt(struct svcxprt_rdma *rdma,
 	return 0;
 }
 
-/* Build a bvec that covers one kvec in an xdr_buf.
+/*
+ * With a NULL @ctxt, bvec constructors count entries without advancing
+ * @info. Otherwise they populate the context and advance @info.
  */
-static void svc_rdma_vec_to_bvec(struct svc_rdma_write_info *info,
-				 unsigned int len,
-				 struct svc_rdma_rw_ctxt *ctxt)
+static unsigned int
+svc_rdma_vec_to_bvec(struct svc_rdma_write_info *info, unsigned int len,
+		     struct svc_rdma_rw_ctxt *ctxt)
 {
+	if (!ctxt)
+		return 1;
+
 	bvec_set_virt(&ctxt->rw_bvec[0], info->wi_base, len);
 	info->wi_base += len;
 
 	ctxt->rw_nents = 1;
+	return 1;
 }
 
 /* Build a bvec array that covers part of an xdr_buf's pagelist.
+ * If @ctxt is NULL, only count the required entries.
  */
-static void svc_rdma_pagelist_to_bvec(struct svc_rdma_write_info *info,
-				      unsigned int remaining,
-				      struct svc_rdma_rw_ctxt *ctxt)
+static unsigned int
+svc_rdma_pagelist_to_bvec(struct svc_rdma_write_info *info,
+			  unsigned int remaining,
+			  struct svc_rdma_rw_ctxt *ctxt)
 {
-	unsigned int bvec_idx, bvec_len, page_off, page_no;
 	const struct xdr_buf *xdr = info->wi_xdr;
-	struct page **page;
-
-	page_off = info->wi_next_off + xdr->page_base;
-	page_no = page_off >> PAGE_SHIFT;
-	page_off = offset_in_page(page_off);
-	page = xdr->pages + page_no;
-	info->wi_next_off += remaining;
-	bvec_idx = 0;
-	do {
-		bvec_len = min_t(unsigned int, remaining,
-				 PAGE_SIZE - page_off);
-		bvec_set_page(&ctxt->rw_bvec[bvec_idx], *page, bvec_len,
-			      page_off);
-		remaining -= bvec_len;
-		page_off = 0;
-		bvec_idx++;
-		page++;
-	} while (remaining);
-
-	ctxt->rw_nents = bvec_idx;
+	unsigned int page_pos = info->wi_next_off + xdr->page_base;
+	unsigned int page_no = page_pos >> PAGE_SHIFT;
+	unsigned int max_bvec_len = svc_rdma_max_bvec_len(info->wi_rdma);
+	unsigned int page_off = offset_in_page(page_pos);
+	unsigned int nr_pages = DIV_ROUND_UP(page_off + remaining, PAGE_SIZE);
+	struct bio_vec *bvecs = ctxt ? ctxt->rw_bvec : NULL;
+	unsigned int nents;
+
+	nents = svc_pages_to_bvecs(bvecs, xdr->pages + page_no, nr_pages,
+				   page_off, remaining, max_bvec_len);
+	if (ctxt) {
+		ctxt->rw_nents = nents;
+		info->wi_next_off += remaining;
+	}
+	return nents;
 }
 
 /* Construct RDMA Write WRs to send a portion of an xdr_buf containing
@@ -492,9 +494,9 @@ static void svc_rdma_pagelist_to_bvec(struct svc_rdma_write_info *info,
  */
 static int
 svc_rdma_build_writes(struct svc_rdma_write_info *info,
-		      void (*constructor)(struct svc_rdma_write_info *info,
-					  unsigned int len,
-					  struct svc_rdma_rw_ctxt *ctxt),
+		      unsigned int (*constructor)(struct svc_rdma_write_info *,
+						  unsigned int,
+						  struct svc_rdma_rw_ctxt *),
 		      unsigned int remaining)
 {
 	struct svc_rdma_chunk_ctxt *cc = &info->wi_cc;
@@ -504,6 +506,7 @@ svc_rdma_build_writes(struct svc_rdma_write_info *info,
 	int ret;
 
 	do {
+		unsigned int nr_bvec;
 		unsigned int write_len;
 		u64 offset;
 
@@ -514,12 +517,18 @@ svc_rdma_build_writes(struct svc_rdma_write_info *info,
 		write_len = min(remaining, seg->rs_length - info->wi_seg_off);
 		if (!write_len)
 			goto out_overflow;
-		ctxt = svc_rdma_get_rw_ctxt(rdma,
-					    (write_len >> PAGE_SHIFT) + 2);
+		nr_bvec = constructor(info, write_len, NULL);
+		if (!nr_bvec)
+			return -EIO;
+
+		ctxt = svc_rdma_get_rw_ctxt(rdma, nr_bvec);
 		if (!ctxt)
 			return -ENOMEM;
 
-		constructor(info, write_len, ctxt);
+		if (WARN_ON_ONCE(constructor(info, write_len, ctxt) != nr_bvec)) {
+			svc_rdma_put_rw_ctxt(rdma, ctxt);
+			return -EIO;
+		}
 		offset = seg->rs_offset + info->wi_seg_off;
 		ret = svc_rdma_rw_ctx_init(rdma, ctxt, offset, seg->rs_handle,
 					   write_len, DMA_TO_DEVICE);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 15+ messages in thread

* [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks
  2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
                   ` (3 preceding siblings ...)
  2026-10-06  9:00 ` [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks J Louis Kaplan
@ 2026-10-06  9:00 ` J Louis Kaplan
  2026-10-06  9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
  2026-10-06 13:47 ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled Chuck Lever
  6 siblings, 0 replies; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-06  9:00 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan

Read chunks currently describe their destination buffers with one
bvec per base page. Coalesce contiguous destination pages to reduce
the number of entries passed to the transport mapping code.

Derive the number of consumed pages from the final page cursor,
since one bvec can now cover several pages.

Limit each iWARP Read context to one memory region: the core
registration code cannot split a coalesced entry between regions.
The force_mr case on other transports remains unresolved.

Update comment for multipage bvecs.

Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
 drivers/infiniband/core/rw.c      |   7 +--
 net/sunrpc/xprtrdma/svc_rdma_rw.c | 101 ++++++++++++++++++------------
 2 files changed, 65 insertions(+), 43 deletions(-)

diff --git a/drivers/infiniband/core/rw.c b/drivers/infiniband/core/rw.c
index 4fafe393a48c7..f9e0a5b8ea806 100644
--- a/drivers/infiniband/core/rw.c
+++ b/drivers/infiniband/core/rw.c
@@ -706,10 +706,9 @@ int rdma_rw_ctx_init_bvec(struct rdma_rw_ctx *ctx, struct ib_qp *qp,
 	 * is a throughput optimization, not a correctness requirement.
 	 * (iWARP, which does require MRs, is handled by the check above.)
 	 *
-	 * The rdma_rw_io_needs_mr() gate is not used here because nr_bvec
-	 * is a raw page count that overstates DMA entry demand -- the bvec
-	 * caller has no post-DMA-coalescing segment count, and feeding the
-	 * inflated count into the MR path exhausts the pool on RDMA READs.
+	 * We do not use max_sgl_rd to switch non-iWARP READs to MRs. Doing so
+	 * would require the MR builder to split multipage bvecs across MRs,
+	 * which it does not yet support. The force_mr case is handled above.
 	 */
 	return rdma_rw_init_map_wrs_bvec(ctx, qp, bvecs, nr_bvec, &iter,
 			remote_addr, rkey, dir);
diff --git a/net/sunrpc/xprtrdma/svc_rdma_rw.c b/net/sunrpc/xprtrdma/svc_rdma_rw.c
index a2e1c693b1bc1..074a75d251c96 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_rw.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_rw.c
@@ -789,54 +789,78 @@ static int svc_rdma_build_read_segment(struct svc_rqst *rqstp,
 {
 	struct svcxprt_rdma *rdma = svc_rdma_rqst_rdma(rqstp);
 	struct svc_rdma_chunk_ctxt *cc = &head->rc_cc;
-	unsigned int bvec_idx, nr_bvec, seg_len, len, total;
+	struct ib_device *dev = rdma->sc_cm_id->device;
+	unsigned int max_bvec_len = svc_rdma_max_bvec_len(rdma);
+	u64 remote_offset = segment->rs_offset;
+	unsigned int nr_bvec, len, remaining, total;
+	unsigned int base_pages;
 	struct svc_rdma_rw_ctxt *ctxt;
 	int ret;
 
-	len = segment->rs_length;
-	if (check_add_overflow(head->rc_pageoff, len, &total))
+	remaining = segment->rs_length;
+	if (check_add_overflow(head->rc_pageoff, remaining, &total))
 		return -EINVAL;
-	nr_bvec = PAGE_ALIGN(total) >> PAGE_SHIFT;
-	ctxt = svc_rdma_get_rw_ctxt(rdma, nr_bvec);
-	if (!ctxt)
-		return -ENOMEM;
-	ctxt->rw_nents = nr_bvec;
-
-	for (bvec_idx = 0; bvec_idx < ctxt->rw_nents; bvec_idx++) {
-		seg_len = min_t(unsigned int, len,
-				PAGE_SIZE - head->rc_pageoff);
-
-		if (!head->rc_pageoff)
-			head->rc_page_count++;
+	base_pages = PAGE_ALIGN(total) >> PAGE_SHIFT;
+	if (!remaining || head->rc_curpage >= rqstp->rq_maxpages ||
+	    base_pages > rqstp->rq_maxpages - head->rc_curpage)
+		goto out_overrun;
+
+	while (remaining) {
+		base_pages = DIV_ROUND_UP(head->rc_pageoff + remaining,
+					  PAGE_SIZE);
+
+		/*
+		 * The bvec MR builder cannot split a coalesced SG entry between
+		 * MRs. iWARP always uses MRs for RDMA Reads, so keep each context
+		 * within one MR. force_mr on other transports remains subject to
+		 * the core limitation.
+		 */
+		if (rdma_protocol_iwarp(dev, rdma->sc_port_num)) {
+			unsigned int nr_mrs = rdma_rw_mr_factor(dev,
+							rdma->sc_port_num,
+							base_pages);
 
-		bvec_set_page(&ctxt->rw_bvec[bvec_idx],
-			      rqstp->rq_pages[head->rc_curpage],
-			      seg_len, head->rc_pageoff);
+			base_pages = DIV_ROUND_UP(base_pages, nr_mrs);
+		}
 
-		head->rc_pageoff += seg_len;
-		if (head->rc_pageoff == PAGE_SIZE) {
-			head->rc_curpage++;
-			head->rc_pageoff = 0;
+		len = min_t(unsigned int, remaining,
+			    (base_pages << PAGE_SHIFT) - head->rc_pageoff);
+		nr_bvec = svc_pages_to_bvecs(NULL,
+					     rqstp->rq_pages + head->rc_curpage,
+					     base_pages, head->rc_pageoff, len,
+					     max_bvec_len);
+		if (!nr_bvec)
+			goto out_overrun;
+		ctxt = svc_rdma_get_rw_ctxt(rdma, nr_bvec);
+		if (!ctxt)
+			return -ENOMEM;
+		ctxt->rw_nents = svc_pages_to_bvecs(ctxt->rw_bvec,
+						    rqstp->rq_pages + head->rc_curpage,
+						    base_pages, head->rc_pageoff, len,
+						    max_bvec_len);
+		if (WARN_ON_ONCE(ctxt->rw_nents != nr_bvec)) {
+			svc_rdma_put_rw_ctxt(rdma, ctxt);
+			goto out_overrun;
 		}
-		len -= seg_len;
+		total = head->rc_pageoff + len;
+		head->rc_curpage += total >> PAGE_SHIFT;
+		head->rc_pageoff = offset_in_page(total);
 
-		if (len && ((head->rc_curpage + 1) > rqstp->rq_maxpages))
-			goto out_put;
-	}
+		ret = svc_rdma_rw_ctx_init(rdma, ctxt, remote_offset,
+					   segment->rs_handle, len,
+					   DMA_FROM_DEVICE);
+		if (ret < 0)
+			return -EIO;
 
-	ret = svc_rdma_rw_ctx_init(rdma, ctxt, segment->rs_offset,
-				   segment->rs_handle, segment->rs_length,
-				   DMA_FROM_DEVICE);
-	if (ret < 0)
-		return -EIO;
+		list_add(&ctxt->rw_list, &cc->cc_rwctxts);
+		cc->cc_sqecount += ret;
+		remote_offset += len;
+		remaining -= len;
+	}
 	percpu_counter_inc(&svcrdma_stat_read);
-
-	list_add(&ctxt->rw_list, &cc->cc_rwctxts);
-	cc->cc_sqecount += ret;
 	return 0;
 
-out_put:
-	svc_rdma_put_rw_ctxt(rdma, ctxt);
+out_overrun:
 	trace_svcrdma_page_overrun_err(&cc->cc_cid, head->rc_curpage);
 	return -EINVAL;
 }
@@ -908,9 +932,6 @@ static int svc_rdma_copy_inline_range(struct svc_rqst *rqstp,
 		page_len = min_t(unsigned int, remaining,
 				 PAGE_SIZE - head->rc_pageoff);
 
-		if (!head->rc_pageoff)
-			head->rc_page_count++;
-
 		dst = page_address(rqstp->rq_pages[head->rc_curpage]);
 		memcpy((unsigned char *)dst + head->rc_pageoff, src + offset, page_len);
 
@@ -1170,6 +1191,8 @@ static void svc_rdma_clear_rqst_pages(struct svc_rqst *rqstp,
 {
 	unsigned int i;
 
+	/* The cursor identifies every page touched while rebuilding the call. */
+	head->rc_page_count = head->rc_curpage + !!head->rc_pageoff;
 	for (i = 0; i < head->rc_page_count; i++) {
 		head->rc_pages[i] = rqstp->rq_pages[i];
 		rqstp->rq_pages[i] = NULL;
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 15+ messages in thread

* [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios
  2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
                   ` (4 preceding siblings ...)
  2026-10-06  9:00 ` [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks J Louis Kaplan
@ 2026-10-06  9:00 ` J Louis Kaplan
  2026-10-06 14:00   ` Chuck Lever
  2026-10-06 13:47 ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled Chuck Lever
  6 siblings, 1 reply; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-06  9:00 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan

Bulk allocation of base pages does not guarantee physical
contiguity, limiting opportunities to coalesce service buffers.

Allocate large folios opportunistically and populate the service
page arrays with their constituent pages. Cap folio size at
64 KiB and allocation order at PAGE_ALLOC_COSTLY_ORDER, falling
back through smaller orders to bulk base-page allocation.

Retain one reference per populated page-array slot so existing
page-based ownership and release paths remain valid.

Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
 net/sunrpc/svc_xprt.c | 44 ++++++++++++++++++++++++++++++++++++++-----
 1 file changed, 39 insertions(+), 5 deletions(-)

diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 9858dfcb846a8..90ec75a848da9 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -19,6 +19,7 @@
 #include <linux/sunrpc/bc_xprt.h>
 #include <linux/module.h>
 #include <linux/netdevice.h>
+#include <linux/sizes.h>
 #include <trace/events/sunrpc.h>
 
 #define RPCDBG_FACILITY	RPCDBG_SVCXPRT
@@ -708,20 +709,53 @@ static void svc_check_conn_limits(struct svc_serv *serv)
 static bool svc_fill_pages(struct svc_rqst *rqstp, struct page **pages,
 			   unsigned long npages)
 {
-	unsigned long filled, ret;
+	unsigned int order = min_t(unsigned int, PAGE_ALLOC_COSTLY_ORDER,
+				   get_order(SZ_64K));
+	unsigned long filled = 0;
 
-	for (filled = 0; filled < npages; filled = ret) {
-		ret = alloc_pages_bulk(GFP_KERNEL, npages, pages);
-		if (ret > filled)
+	/*
+	 * Prefer folios no larger than 64 KiB or the page allocator's
+	 * costly-order threshold. After a failure, use smaller orders for
+	 * the rest of this fill; order zero is handled by the bulk allocator.
+	 * All callers provide a range containing only NULL entries.
+	 */
+	while (filled < npages) {
+		struct folio *folio;
+		unsigned int nr_pages;
+		unsigned long ret;
+
+		while (order && (1UL << order) > npages - filled)
+			order--;
+
+		if (order) {
+			folio = folio_alloc(GFP_NOWAIT, order);
+			if (!folio) {
+				order--;
+				continue;
+			}
+
+			nr_pages = folio_nr_pages(folio);
+			folio_ref_add(folio, nr_pages - 1);
+			for (unsigned int i = 0; i < nr_pages; i++)
+				pages[filled + i] = folio_page(folio, i);
+			filled += nr_pages;
+			continue;
+		}
+
+		ret = alloc_pages_bulk(GFP_KERNEL, npages - filled,
+				       pages + filled);
+		if (ret) {
+			filled += ret;
 			/* Made progress, don't sleep yet */
 			continue;
+		}
 
 		set_current_state(TASK_IDLE);
 		if (svc_thread_should_stop(rqstp)) {
 			set_current_state(TASK_RUNNING);
 			return false;
 		}
-		trace_svc_alloc_arg_err(npages, ret);
+		trace_svc_alloc_arg_err(npages, filled);
 		memalloc_retry_wait(GFP_KERNEL);
 	}
 	return true;
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 15+ messages in thread

* Re: [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled
  2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
                   ` (5 preceding siblings ...)
  2026-10-06  9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
@ 2026-10-06 13:47 ` Chuck Lever
  2026-10-08 22:50   ` J Louis Kaplan
  6 siblings, 1 reply; 15+ messages in thread
From: Chuck Lever @ 2026-10-06 13:47 UTC (permalink / raw)
  To: J Louis Kaplan, Dai Ngo, Jeff Layton, NeilBrown,
	Olga Kornievskaia, Tom Talpey
  Cc: linux-nfs, linux-rdma, Anna Schumaker, Jason Gunthorpe,
	Leon Romanovsky, Trond Myklebust



On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> For NFSoRDMA using RoCE, and with IOMMU passthrough enabled, use of a
> 64KiB base page size was found to provide higher NFS read throughput
> than a 4KiB base page size. This patch series recovers some of that
> throughput with a 4k base page size by opportunistic use of large folios
> for the svc reply buffer.

> LLM Usage
> =========
>
> The idea of folio usage here was human-generated from observations of
> performance differences between 4KiB and 64KiB page size kernels.

Actually a similar approach has been tried before, but on x86_64:

https://lore.kernel.org/linux-nfs/20260319133610.2556826-1-cel@kernel.org/

And it had to be reverted:

- Mike Snitzer's report (4 Jun 2026):
    https://lore.kernel.org/linux-nfs/aiHlPmeZq3WgMwoJ@kernel.org/

- Jonathan Flynn's benchmark and perf data (5 Jun 2026), showing server
  CPU going from 8.54% to 76.35% with the commit present:
    https://lore.kernel.org/linux-nfs/3cb119b4b2a8aada30c0c60286778a54@mail.gmail.com/

These two are the Closes: targets in the revert, which is a39f0ce0c9da
upstream and 284d5ba931a5 in stable.

Lock contention analysis and the decision to revert

- "[PATCH] svcrdma: Cap Read sink allocations at PAGE_ALLOC_COSTLY_ORDER"
  (5 Jun 2026), whose description carries the zone->lock analysis:
    https://lore.kernel.org/linux-nfs/20260606035722.83175-1-cel@kernel.org/
- Flynn's results for that fix (25.4 GiB/s, against 30.3 regressed and
  73.9 reverted):
    https://lore.kernel.org/linux-nfs/65a2cdb132b0c28e69a29955e3bd37e7@mail.gmail.com/
- Your reply saying the two failed fixes mean 18755b8c2f24 has to be
  reverted (6 Jun 2026):
    https://lore.kernel.org/linux-nfs/096a2b91-7a19-48da-a06a-dc60e7150956@app.fastmail.com/

The earlier failed fix, "[PATCH] svcrdma: Avoid direct reclaim when
allocating Read sink buffers", is at
    https://lore.kernel.org/linux-nfs/20260605223118.75092-1-cel@kernel.org/.

So: Yes, we would like to ensure that svcrdma's RDMA Read buffers (at
least, and maybe RDMA Write) are contiguous so that the overhead of a
DMA mapping operation will be low. I'm not expert enough with the page
allocator's free path to address the lock contention problem, however.


-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages
  2026-10-06  9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
@ 2026-10-06 13:49   ` Chuck Lever
  2026-10-08 22:44     ` J Louis Kaplan
  0 siblings, 1 reply; 15+ messages in thread
From: Chuck Lever @ 2026-10-06 13:49 UTC (permalink / raw)
  To: J Louis Kaplan, Dai Ngo, Jeff Layton, NeilBrown,
	Olga Kornievskaia, Tom Talpey
  Cc: linux-nfs, linux-rdma, Anna Schumaker, Jason Gunthorpe,
	Leon Romanovsky, Trond Myklebust



On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> Add helpers to describe contiguous runs with multipage bvecs. These
> functions have no callers. They are used in later commits.
>
> Accept a caller-supplied maximum bvec length so transport callers
> can enforce direct memory access (DMA) mapping and device segment
> length limits. Also support counting entries without populating
> an array so callers can size their allocations using the same
> coalescing rules.
>
> Coalescing requires contiguous physical memory and page descriptors,
> and is disabled under KMSAN. Xen constraint handling remains WIP.
>
> Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
> Assisted-by: LLM
> ---
>  include/linux/sunrpc/svc.h | 90 ++++++++++++++++++++++++++++++++++++++
>  1 file changed, 90 insertions(+)
>
> diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
> index dfadea50e0c6d..e651c2312a912 100644
> --- a/include/linux/sunrpc/svc.h
> +++ b/include/linux/sunrpc/svc.h
> @@ -11,6 +11,7 @@
>  #ifndef SUNRPC_SVC_H
>  #define SUNRPC_SVC_H
> 
> +#include <linux/bvec.h>
>  #include <linux/in.h>
>  #include <linux/in6.h>
>  #include <linux/sunrpc/types.h>
> @@ -376,6 +377,95 @@ static inline void svc_thread_init_status(struct 
> svc_rqst *rqstp, int err)
>  		kthread_exit(1);
>  }
> 
> +/**
> + * svc_pages_to_bvec - build one bvec from physically contiguous pages
> + * @bv: bio_vec to initialize, or NULL to only measure the next extent
> + * @pages: first page-array slot in the range
> + * @nr_pages: number of page-array slots available from @pages
> + * @offset: byte offset in @pages[0]
> + * @len: maximum byte count to describe
> + * @max_bvec_len: maximum byte count in one bvec
> + *
> + * Return: The number of bytes described by @bv. The scan stops at a missing
> + * or physically non-contiguous page.
> + */
> +static inline unsigned int
> +svc_pages_to_bvec(struct bio_vec *bv, struct page *const *pages,
> +		  unsigned int nr_pages, unsigned int offset,
> +		  unsigned int len, unsigned int max_bvec_len)

Generally any function larger than two or three lines is to be
kept out of line. net/sunrpc/svc.c, perhaps, would be appropriate
here.


> +{
> +	unsigned long first_pfn;
> +	unsigned int bytes, used = 1;
> +
> +	if (WARN_ON_ONCE(!nr_pages || offset >= PAGE_SIZE || !len ||
> +			 !max_bvec_len || !pages[0]))
> +		return 0;
> +	first_pfn = page_to_pfn(pages[0]);
> +
> +	bytes = min3(len, max_bvec_len,
> +		     (unsigned int)(PAGE_SIZE - offset));
> +	while (bytes < len && bytes < max_bvec_len && used < nr_pages) {
> +		// TODO: handle Xen merge constraints, see `bvec_try_merge_page`
> +		//	 for reference, or unify helper functionality
> +		if (IS_ENABLED(CONFIG_KMSAN))
> +			break;
> +
> +		if (!pages[used] ||
> +		    page_to_pfn(pages[used]) != first_pfn + used ||
> +		    pages[0] + used != pages[used])
> +			break;
> +		used++;
> +		bytes = min3(len, max_bvec_len,
> +			     (used << PAGE_SHIFT) - offset);
> +	}
> +
> +	if (bv)
> +		bvec_set_page(bv, pages[0], bytes, offset);
> +	return bytes;
> +}
> +
> +/**
> + * svc_pages_to_bvecs - build bvecs for a service page-array range
> + * @bvecs: bio_vec array to populate, or NULL to only count the entries
> + * @pages: first page-array slot in the range
> + * @nr_pages: number of page-array slots available from @pages
> + * @offset: byte offset in @pages[0]
> + * @len: byte count to describe
> + * @max_bvec_len: maximum byte count in one bvec
> + *
> + * If @bvecs is not NULL, it must have room for the count returned by a
> + * preceding count-only call.
> + *
> + * Return: The number of populated or required bvecs, or zero if the range
> + * cannot be described from the supplied page-array slots.
> + */
> +static inline unsigned int
> +svc_pages_to_bvecs(struct bio_vec *bvecs, struct page *const *pages,
> +		   unsigned int nr_pages, unsigned int offset,
> +		   unsigned int len, unsigned int max_bvec_len)
> +{
> +	unsigned int remaining = len;
> +	unsigned int nents = 0;
> +
> +	while (remaining && nr_pages) {
> +		struct bio_vec *bv = bvecs ? &bvecs[nents] : NULL;
> +		unsigned int advanced;
> +		unsigned int bytes;
> +
> +		bytes = svc_pages_to_bvec(bv, pages, nr_pages, offset,
> +					  remaining, max_bvec_len);
> +		if (!bytes)
> +			break;
> +		advanced = (offset + bytes) >> PAGE_SHIFT;
> +		offset = offset_in_page(offset + bytes);
> +		pages += advanced;
> +		nr_pages -= advanced;
> +		remaining -= bytes;
> +		nents++;
> +	}
> +	return remaining ? 0 : nents;
> +}
> +
>  struct svc_deferred_req {
>  	u32			prot;	/* protocol (UDP or TCP) */
>  	bool			secure;	/* RQ_SECURE of the original request */
> -- 
> 2.43.0

-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies
  2026-10-06  9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
@ 2026-10-06 13:55   ` Chuck Lever
  2026-10-08 22:54     ` J Louis Kaplan
  0 siblings, 1 reply; 15+ messages in thread
From: Chuck Lever @ 2026-10-06 13:55 UTC (permalink / raw)
  To: J Louis Kaplan, Dai Ngo, Jeff Layton, NeilBrown,
	Olga Kornievskaia, Tom Talpey
  Cc: linux-nfs, linux-rdma, Anna Schumaker, Jason Gunthorpe,
	Leon Romanovsky, Trond Myklebust



On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> Mapping reply pagelists one page at a time uses a separate direct
> memory access (DMA) mapping and scatter/gather entry (SGE) for
> each base page, even when those pages are physically contiguous.
>
> Instead, map contiguous runs as multipage bvecs, capped by the device's
> mapping and segment length limits. Use the same coalescing rules
> when counting SGEs for the pull-up decision, avoiding unnecessary
> copying when the coalesced reply fits the Send SGE limit.
>
> Virtual-mapping devices bypass device-limit check.
>
> Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
> Assisted-by: LLM
> ---
>  include/linux/sunrpc/svc_rdma.h       |  15 ++++
>  net/sunrpc/xprtrdma/svc_rdma_sendto.c | 112 +++++++++++++++-----------
>  2 files changed, 82 insertions(+), 45 deletions(-)
>
> diff --git a/include/linux/sunrpc/svc_rdma.h b/include/linux/sunrpc/svc_rdma.h
> index 71bd5bcc5ce2a..c296bf8b8b1e3 100644
> --- a/include/linux/sunrpc/svc_rdma.h
> +++ b/include/linux/sunrpc/svc_rdma.h
> @@ -42,6 +42,7 @@
> 
>  #ifndef SVC_RDMA_H
>  #define SVC_RDMA_H
> +#include <linux/dma-mapping.h>
>  #include <linux/llist.h>
>  #include <linux/sunrpc/xdr.h>
>  #include <linux/sunrpc/svcsock.h>
> @@ -134,6 +135,20 @@ static inline struct svcxprt_rdma 
> *svc_rdma_rqst_rdma(struct svc_rqst *rqstp)
>  	return container_of(xprt, struct svcxprt_rdma, sc_xprt);
>  }
> 
> +static inline unsigned int
> +svc_rdma_max_bvec_len(const struct svcxprt_rdma *rdma)

There are some interesting infrastructural changes in this series.
Using bvecs through the svcrdma code has been on my to-do list
for quite some time.


> +{
> +	struct ib_device *device = rdma->sc_cm_id->device;
> +
> +	if (ib_uses_virt_dma(device))
> +		return UINT_MAX;
> +
> +	/* Respect both DMA mapping and RDMA device segment limits. */
> +	return min_t(size_t,
> +		     dma_max_mapping_size(device->dma_device),
> +		     ib_dma_max_seg_size(device));
> +}
> +
>  /*
>   * Default connection parameters
>   */
> diff --git a/net/sunrpc/xprtrdma/svc_rdma_sendto.c 
> b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
> index efa352ddf9d71..9f2e3ee3c2936 100644
> --- a/net/sunrpc/xprtrdma/svc_rdma_sendto.c
> +++ b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
> @@ -260,10 +260,8 @@ static void svc_rdma_send_ctxt_unmap(struct 
> svcxprt_rdma *rdma,
>  		trace_svcrdma_dma_unmap_page(&ctxt->sc_cid,
>  					     ctxt->sc_sges[i].addr,
>  					     ctxt->sc_sges[i].length);
> -		ib_dma_unmap_page(device,
> -				  ctxt->sc_sges[i].addr,
> -				  ctxt->sc_sges[i].length,
> -				  DMA_TO_DEVICE);
> +		ib_dma_unmap_bvec(device, ctxt->sc_sges[i].addr,
> +				  ctxt->sc_sges[i].length, DMA_TO_DEVICE);
>  	}
>  }
> 
> @@ -730,44 +728,45 @@ svc_rdma_encode_reply_chunk(struct 
> svc_rdma_recv_ctxt *rctxt,
>  }
> 
>  struct svc_rdma_map_data {
> -	struct svcxprt_rdma		*md_rdma;
>  	struct svc_rdma_send_ctxt	*md_ctxt;
> +	unsigned int			md_max_bvec_len;
>  };
> 
>  /**
> - * svc_rdma_page_dma_map - DMA map one page
> + * svc_rdma_bvec_dma_map - DMA map one physically contiguous bvec
>   * @data: pointer to arguments
> - * @page: struct page to DMA map
> - * @offset: offset into the page
> - * @len: number of bytes to map
> + * @bv: contiguous range to DMA map
>   *
>   * Returns:
>   *   %0 if DMA mapping was successful
> - *   %-EIO if the page cannot be DMA mapped
> + *   %-EIO if the range cannot be DMA mapped
>   */
> -static int svc_rdma_page_dma_map(void *data, struct page *page,
> -				 unsigned long offset, unsigned int len)
> +static int svc_rdma_bvec_dma_map(void *data, struct bio_vec *bv)
>  {
>  	struct svc_rdma_map_data *args = data;
> -	struct svcxprt_rdma *rdma = args->md_rdma;
>  	struct svc_rdma_send_ctxt *ctxt = args->md_ctxt;
> +	struct svcxprt_rdma *rdma = ctxt->sc_rdma;
>  	struct ib_device *dev = rdma->sc_cm_id->device;
>  	dma_addr_t dma_addr;
> 
> +	if (bv->bv_len > args->md_max_bvec_len)
> +		return -EIO;
> +	if (WARN_ON_ONCE(ctxt->sc_send_wr.num_sge >= rdma->sc_max_send_sges))
> +		return -EIO;
>  	++ctxt->sc_cur_sge_no;
> 
> -	dma_addr = ib_dma_map_page(dev, page, offset, len, DMA_TO_DEVICE);
> +	dma_addr = ib_dma_map_bvec(dev, bv, DMA_TO_DEVICE);
>  	if (ib_dma_mapping_error(dev, dma_addr))
>  		goto out_maperr;
> 
> -	trace_svcrdma_dma_map_page(&ctxt->sc_cid, dma_addr, len);
> +	trace_svcrdma_dma_map_page(&ctxt->sc_cid, dma_addr, bv->bv_len);
>  	ctxt->sc_sges[ctxt->sc_cur_sge_no].addr = dma_addr;
> -	ctxt->sc_sges[ctxt->sc_cur_sge_no].length = len;
> +	ctxt->sc_sges[ctxt->sc_cur_sge_no].length = bv->bv_len;
>  	ctxt->sc_send_wr.num_sge++;
>  	return 0;
> 
>  out_maperr:
> -	trace_svcrdma_dma_map_err(&ctxt->sc_cid, dma_addr, len);
> +	trace_svcrdma_dma_map_err(&ctxt->sc_cid, dma_addr, bv->bv_len);
>  	return -EIO;
>  }
> 
> @@ -776,20 +775,18 @@ static int svc_rdma_page_dma_map(void *data, 
> struct page *page,
>   * @data: pointer to arguments
>   * @iov: kvec to DMA map
>   *
> - * ib_dma_map_page() is used here because svc_rdma_dma_unmap()
> - * handles DMA-unmap and it uses ib_dma_unmap_page() exclusively.
> - *
>   * Returns:
>   *   %0 if DMA mapping was successful
>   *   %-EIO if the iovec cannot be DMA mapped
>   */
>  static int svc_rdma_iov_dma_map(void *data, const struct kvec *iov)
>  {
> +	struct bio_vec bv;
> +
>  	if (!iov->iov_len)
>  		return 0;
> -	return svc_rdma_page_dma_map(data, virt_to_page(iov->iov_base),
> -				     offset_in_page(iov->iov_base),
> -				     iov->iov_len);
> +	bvec_set_virt(&bv, iov->iov_base, iov->iov_len);
> +	return svc_rdma_bvec_dma_map(data, &bv);
>  }
> 
>  /**
> @@ -798,7 +795,7 @@ static int svc_rdma_iov_dma_map(void *data, const 
> struct kvec *iov)
>   * @data: pointer to arguments
>   *
>   * Returns:
> - *   %0 if DMA mapping was successful
> + *   The number of mapped bytes if DMA mapping was successful
>   *   %-EIO if DMA mapping failed
>   *
>   * On failure, any DMA mappings that have been already done must be
> @@ -806,9 +803,11 @@ static int svc_rdma_iov_dma_map(void *data, const 
> struct kvec *iov)
>   */
>  static int svc_rdma_xb_dma_map(const struct xdr_buf *xdr, void *data)
>  {
> -	unsigned int len, remaining;
> +	struct svc_rdma_map_data *args = data;
> +	unsigned int remaining;
>  	unsigned long pageoff;
>  	struct page **ppages;
> +	unsigned int nr_pages;
>  	int ret;
> 
>  	ret = svc_rdma_iov_dma_map(data, &xdr->head[0]);
> @@ -818,15 +817,25 @@ static int svc_rdma_xb_dma_map(const struct 
> xdr_buf *xdr, void *data)
>  	ppages = xdr->pages + (xdr->page_base >> PAGE_SHIFT);
>  	pageoff = offset_in_page(xdr->page_base);
>  	remaining = xdr->page_len;
> +	nr_pages = DIV_ROUND_UP(pageoff + remaining, PAGE_SIZE);
>  	while (remaining) {
> -		len = min_t(u32, PAGE_SIZE - pageoff, remaining);
> -
> -		ret = svc_rdma_page_dma_map(data, *ppages++, pageoff, len);
> +		struct bio_vec bv;
> +		unsigned int bytes = svc_pages_to_bvec(&bv, ppages,
> +				nr_pages, pageoff, remaining,
> +				args->md_max_bvec_len);
> +		unsigned int advanced =
> +				(pageoff + bytes) >> PAGE_SHIFT;
> +
> +		if (!bytes)
> +			return -EIO;
> +		ret = svc_rdma_bvec_dma_map(data, &bv);
>  		if (ret < 0)
>  			return ret;
> 
> -		remaining -= len;
> -		pageoff = 0;
> +		pageoff = offset_in_page(pageoff + bytes);
> +		ppages += advanced;
> +		nr_pages -= advanced;
> +		remaining -= bytes;
>  	}
> 
>  	ret = svc_rdma_iov_dma_map(data, &xdr->tail[0]);
> @@ -836,10 +845,27 @@ static int svc_rdma_xb_dma_map(const struct 
> xdr_buf *xdr, void *data)
>  	return xdr->len;
>  }
> 
> +static unsigned int
> +svc_rdma_xb_page_sges(const struct xdr_buf *xdr,
> +		      unsigned int max_bvec_len)
> +{
> +	const unsigned long offset = offset_in_page(xdr->page_base);
> +	const unsigned int nr_pages =
> +		DIV_ROUND_UP(offset + xdr->page_len, PAGE_SIZE);
> +	struct page *const *ppages;
> +
> +	if (!xdr->page_len)
> +		return 0;
> +	ppages = xdr->pages + (xdr->page_base >> PAGE_SHIFT);
> +	return svc_pages_to_bvecs(NULL, ppages, nr_pages, offset,
> +				  xdr->page_len, max_bvec_len);
> +}
> +
>  struct svc_rdma_pullup_data {
>  	u8		*pd_dest;
>  	unsigned int	pd_length;
>  	unsigned int	pd_num_sges;
> +	unsigned int	pd_max_bvec_len;
>  };
> 
>  /**
> @@ -848,25 +874,22 @@ struct svc_rdma_pullup_data {
>   * @data: pointer to arguments
>   *
>   * Returns:
> - *   Number of SGEs needed to Send the contents of @xdr inline
> + *   %0 if the SGE count was updated
> + *   %-EIO if the page list cannot be described
>   */
>  static int svc_rdma_xb_count_sges(const struct xdr_buf *xdr,
>  				  void *data)
>  {
>  	struct svc_rdma_pullup_data *args = data;
> -	unsigned int remaining;
> -	unsigned long offset;
> +	unsigned int page_sges =
> +		svc_rdma_xb_page_sges(xdr, args->pd_max_bvec_len);
> 
>  	if (xdr->head[0].iov_len)
>  		++args->pd_num_sges;
> 
> -	offset = offset_in_page(xdr->page_base);
> -	remaining = xdr->page_len;
> -	while (remaining) {
> -		++args->pd_num_sges;
> -		remaining -= min_t(u32, PAGE_SIZE - offset, remaining);
> -		offset = 0;
> -	}
> +	if (xdr->page_len && !page_sges)
> +		return -EIO;
> +	args->pd_num_sges += page_sges;
> 
>  	if (xdr->tail[0].iov_len)
>  		++args->pd_num_sges;
> @@ -896,6 +919,7 @@ static int svc_rdma_check_pull_up(const struct 
> svcxprt_rdma *rdma,
>  	struct svc_rdma_pullup_data args = {
>  		.pd_length	= sctxt->sc_hdrbuf.len,
>  		.pd_num_sges	= 1,
> +		.pd_max_bvec_len = svc_rdma_max_bvec_len(rdma),
>  	};
>  	int ret;
> 
> @@ -964,7 +988,6 @@ static int svc_rdma_xb_linearize(const struct xdr_buf *xdr,
> 
>  /**
>   * svc_rdma_pull_up_reply_msg - Copy Reply into a single buffer
> - * @rdma: controlling transport
>   * @sctxt: send_ctxt for the Send WR; xprt hdr is already prepared
>   * @write_pcl: Write chunk list provided by client
>   * @xdr: prepared xdr_buf containing RPC message
> @@ -979,8 +1002,7 @@ static int svc_rdma_xb_linearize(const struct xdr_buf *xdr,
>   *   %0 if pull-up was successful
>   *   %-EMSGSIZE if a buffer manipulation problem occurred
>   */
> -static int svc_rdma_pull_up_reply_msg(const struct svcxprt_rdma *rdma,
> -				      struct svc_rdma_send_ctxt *sctxt,
> +static int svc_rdma_pull_up_reply_msg(struct svc_rdma_send_ctxt *sctxt,
>  				      const struct svc_rdma_pcl *write_pcl,
>  				      const struct xdr_buf *xdr)
>  {
> @@ -1021,8 +1043,8 @@ int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
>  			   const struct xdr_buf *xdr)
>  {
>  	struct svc_rdma_map_data args = {
> -		.md_rdma	= rdma,
>  		.md_ctxt	= sctxt,
> +		.md_max_bvec_len = svc_rdma_max_bvec_len(rdma),
>  	};
>  	int ret;
> 
> @@ -1043,7 +1065,7 @@ int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
>  	if (ret < 0)
>  		return ret;
>  	if (ret)
> -		return svc_rdma_pull_up_reply_msg(rdma, sctxt, write_pcl, xdr);
> +		return svc_rdma_pull_up_reply_msg(sctxt, write_pcl, xdr);
> 
>  	return pcl_process_nonpayloads(write_pcl, xdr,
>  				       svc_rdma_xb_dma_map, &args);
> -- 
> 2.43.0

-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios
  2026-10-06  9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
@ 2026-10-06 14:00   ` Chuck Lever
  2026-10-08 22:58     ` J Louis Kaplan
  0 siblings, 1 reply; 15+ messages in thread
From: Chuck Lever @ 2026-10-06 14:00 UTC (permalink / raw)
  To: J Louis Kaplan, dai.ngo, jlayton, neil, okorniev, tom
  Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy



On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> Bulk allocation of base pages does not guarantee physical
> contiguity, limiting opportunities to coalesce service buffers.
>
> Allocate large folios opportunistically and populate the service
> page arrays with their constituent pages. Cap folio size at
> 64 KiB and allocation order at PAGE_ALLOC_COSTLY_ORDER, falling
> back through smaller orders to bulk base-page allocation.
>
> Retain one reference per populated page-array slot so existing
> page-based ownership and release paths remain valid.
>
> Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
> Assisted-by: LLM

Given past experience with contention in the page free path, IMO
this one needs careful macro and micro benchmarking.

I wonder if this change should get merged first instead of last.
It's probably going to have broad impact and high risk, and will
need testing across all hardware platforms and on large- and small-
memory configurations.


-- 
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages
  2026-10-06 13:49   ` Chuck Lever
@ 2026-10-08 22:44     ` J Louis Kaplan
  0 siblings, 0 replies; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-08 22:44 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: anna, jgg, leon, linux-nfs, linux-rdma, trondmy

> Generally any function larger than two or three lines is to be kept
> out of line.

Thanks for explaining that, Chuck. I will fix as suggested in any
subsequent patch series.

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled
  2026-10-06 13:47 ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled Chuck Lever
@ 2026-10-08 22:50   ` J Louis Kaplan
  0 siblings, 0 replies; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-08 22:50 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: anna, jgg, leon, linux-nfs, linux-rdma, trondmy

> Actually a similar approach has been tried before... And it had to be
> reverted

Thanks for sharing these useful links. They provide important context
that I was not aware of.

I'll have another read through the investigations and commit dates
to better understand the change timeline, but I wanted to share some
contention data I gathered prompted by your feedback.

During development of this patch series, I did notice lock contention
issues. I suspected the unbound workqueue but my analysis was shallow
and I didn't have a good idea how to improve the situation.

After a rebase I noticed a decrease in contention and found your commit
58202c29de936 in the git log which looked likely to be the beneficial
change.

I set up two test branches after reading your feedback. Both apply this
patch series (with minor conflict resolutions) to the older nfsd-testing
commit a39f0ce0c9da2 which reverts `Use contiguous pages for RDMA Read
sink buffers`. `ubwq` also reverts your patch 58202c29de936 whereas
`no-ubwq` does not.

For direct IO on the server and NFS reads in configuration A I mentioned
above (passthrough enabled), here are representative top lines of perf
contention data (spinlocks only).

ubwq (without 58202c29de936)
contended   total wait    max wait   avg wait   caller
   202787    588.67 ms   334.01 us    2.90 us   __rmqueue_pcplist+0x320
   281818    586.91 ms   267.99 us    2.08 us   free_pcppages_bulk+0x48
    88218    100.65 ms     5.57 us    1.14 us   rcu_report_qs_rdp+0x44
     1124     11.07 ms    28.13 us    9.85 us   tick_do_update_jiffies64+0x48
     5773      4.92 ms     3.42 us     851 ns   tmigr_update_events+0x184
     1274      2.38 ms    66.01 us    1.87 us   nfsd4_sequence_done+0x70
     1272      2.36 ms    51.23 us    1.86 us   nfsd4_sequence+0x114
     1074      2.05 ms     5.57 us    1.90 us   acpi_os_wait_semaphore+0x80
     1074      1.68 ms    15.68 us    1.56 us   free_pcppages_bulk+0x48
      739      1.04 ms     4.80 us    1.41 us   __queue_work+0xd4
      424    745.28 us     8.99 us    1.76 us   raw_spin_rq_lock_nested+0x3c
      362    649.74 us    18.78 us    1.79 us   __rmqueue_pcplist+0x320
    ... <truncated>

no-ubwq (with 58202c29de936):
contended   total wait   max wait   avg wait   caller
    68113     82.67 ms    5.92 us    1.21 us   rcu_report_qs_rdp+0x44
     1604      2.41 ms   21.73 us    1.50 us   nfsd4_sequence_done+0x70
     1689      2.31 ms   26.21 us    1.37 us   nfsd4_sequence+0x114
      884      1.66 ms   15.87 us    1.88 us   raw_spin_rq_lock_nested+0x3c
      153      1.21 ms   16.48 us    7.93 us   tick_do_update_jiffies64+0x48
      311    566.99 us   14.43 us    1.82 us   raw_spin_rq_lock_nested+0x3c
      423    459.48 us    2.98 us    1.09 us   force_qs_rnp+0x104
      185    335.12 us   27.14 us    1.81 us   __lwq_dequeue+0x34
       40    288.12 us   20.35 us    7.20 us   free_pcppages_bulk+0x48
      220    285.37 us    3.97 us    1.30 us   __queue_work+0xd4
      284    242.78 us    2.14 us     854 ns   tmigr_update_events+0x184
      153    216.64 us    3.10 us    1.41 us   acpi_os_wait_semaphore+0x80
      164    216.03 us   26.43 us    1.32 us   find_stateid_by_type+0x40
      143    206.07 us   13.66 us    1.44 us   nfsd4_sequence+0x238
       95    118.24 us    2.40 us    1.24 us   dma_pool_alloc+0x50
       85    116.29 us    2.75 us    1.37 us   acpi_os_signal_semaphore+0x88
       86    101.37 us    5.54 us    1.18 us   raw_spin_rq_lock_nested+0x3c
       34     84.18 us    8.45 us    2.47 us   mix_interrupt_randomness+0x104
       72     77.31 us    1.95 us    1.07 us   raw_spin_rq_lock_nested+0x3c
       43     74.94 us    4.29 us    1.74 us   raw_spin_rq_lock_nested+0x3c
       51     72.09 us    3.33 us    1.41 us   acpi_os_wait_semaphore+0x80
       62     67.26 us    1.92 us    1.08 us   dma_pool_free+0x3c
       42     61.66 us    2.46 us    1.47 us   acpi_os_wait_semaphore+0x80
       39     55.87 us    2.94 us    1.43 us   __queue_work+0xd4
       21     50.65 us   15.23 us    2.41 us   raw_spin_rq_lock_nested+0x3c
       35     45.76 us    2.05 us    1.31 us   raw_spin_rq_lock_nested+0x3c
       15     24.16 us    3.93 us    1.61 us   free_pcppages_bulk+0x48
    ... <truncated>

CPU usage was 7% lower in the no-ubwq branch. Throughput remained similar
in both branches for this test.

I'm still digesting the above (and more) data and trying to understand
how all the information fits together, but I wanted to share the above
in the meantime.

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies
  2026-10-06 13:55   ` Chuck Lever
@ 2026-10-08 22:54     ` J Louis Kaplan
  0 siblings, 0 replies; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-08 22:54 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: anna, jgg, leon, linux-nfs, linux-rdma, trondmy

> There are some interesting infrastructural changes in this series.
> Using bvecs through the svcrdma code has been on my to-do list
> for quite some time.

Glad to hear that. I'm happy to split out smaller parts of work from
this patch series if helpful.

^ permalink raw reply	[flat|nested] 15+ messages in thread

* Re: [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios
  2026-10-06 14:00   ` Chuck Lever
@ 2026-10-08 22:58     ` J Louis Kaplan
  0 siblings, 0 replies; 15+ messages in thread
From: J Louis Kaplan @ 2026-10-08 22:58 UTC (permalink / raw)
  To: cel, dai.ngo, jlayton, neil, okorniev, tom
  Cc: anna, jgg, leon, linux-nfs, linux-rdma, trondmy

> Given past experience with contention in the page free path, IMO
> this one needs careful macro and micro benchmarking.

Makes sense to me after reading through the links you posted in the
other reply.

> I wonder if this change should get merged first instead of last.
> It's probably going to have broad impact and high risk, and will
> need testing across all hardware platforms and on large- and small-
> memory configurations.

Happy to do that. Should I consider a separate single-commit submission
for this change?

^ permalink raw reply	[flat|nested] 15+ messages in thread

end of thread, other threads:[~2026-10-08 22:58 UTC | newest]

Thread overview: 15+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-06  9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
2026-10-06 13:49   ` Chuck Lever
2026-10-08 22:44     ` J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
2026-10-06 13:55   ` Chuck Lever
2026-10-08 22:54     ` J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks J Louis Kaplan
2026-10-06  9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
2026-10-06 14:00   ` Chuck Lever
2026-10-08 22:58     ` J Louis Kaplan
2026-10-06 13:47 ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled Chuck Lever
2026-10-08 22:50   ` J Louis Kaplan

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox