* [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages
2026-10-06 9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
@ 2026-10-06 9:00 ` J Louis Kaplan
2026-10-06 13:49 ` Chuck Lever
2026-10-06 9:00 ` [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads J Louis Kaplan
` (5 subsequent siblings)
6 siblings, 1 reply; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-06 9:00 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan
Add helpers to describe contiguous runs with multipage bvecs. These
functions have no callers. They are used in later commits.
Accept a caller-supplied maximum bvec length so transport callers
can enforce direct memory access (DMA) mapping and device segment
length limits. Also support counting entries without populating
an array so callers can size their allocations using the same
coalescing rules.
Coalescing requires contiguous physical memory and page descriptors,
and is disabled under KMSAN. Xen constraint handling remains WIP.
Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
include/linux/sunrpc/svc.h | 90 ++++++++++++++++++++++++++++++++++++++
1 file changed, 90 insertions(+)
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index dfadea50e0c6d..e651c2312a912 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -11,6 +11,7 @@
#ifndef SUNRPC_SVC_H
#define SUNRPC_SVC_H
+#include <linux/bvec.h>
#include <linux/in.h>
#include <linux/in6.h>
#include <linux/sunrpc/types.h>
@@ -376,6 +377,95 @@ static inline void svc_thread_init_status(struct svc_rqst *rqstp, int err)
kthread_exit(1);
}
+/**
+ * svc_pages_to_bvec - build one bvec from physically contiguous pages
+ * @bv: bio_vec to initialize, or NULL to only measure the next extent
+ * @pages: first page-array slot in the range
+ * @nr_pages: number of page-array slots available from @pages
+ * @offset: byte offset in @pages[0]
+ * @len: maximum byte count to describe
+ * @max_bvec_len: maximum byte count in one bvec
+ *
+ * Return: The number of bytes described by @bv. The scan stops at a missing
+ * or physically non-contiguous page.
+ */
+static inline unsigned int
+svc_pages_to_bvec(struct bio_vec *bv, struct page *const *pages,
+ unsigned int nr_pages, unsigned int offset,
+ unsigned int len, unsigned int max_bvec_len)
+{
+ unsigned long first_pfn;
+ unsigned int bytes, used = 1;
+
+ if (WARN_ON_ONCE(!nr_pages || offset >= PAGE_SIZE || !len ||
+ !max_bvec_len || !pages[0]))
+ return 0;
+ first_pfn = page_to_pfn(pages[0]);
+
+ bytes = min3(len, max_bvec_len,
+ (unsigned int)(PAGE_SIZE - offset));
+ while (bytes < len && bytes < max_bvec_len && used < nr_pages) {
+ // TODO: handle Xen merge constraints, see `bvec_try_merge_page`
+ // for reference, or unify helper functionality
+ if (IS_ENABLED(CONFIG_KMSAN))
+ break;
+
+ if (!pages[used] ||
+ page_to_pfn(pages[used]) != first_pfn + used ||
+ pages[0] + used != pages[used])
+ break;
+ used++;
+ bytes = min3(len, max_bvec_len,
+ (used << PAGE_SHIFT) - offset);
+ }
+
+ if (bv)
+ bvec_set_page(bv, pages[0], bytes, offset);
+ return bytes;
+}
+
+/**
+ * svc_pages_to_bvecs - build bvecs for a service page-array range
+ * @bvecs: bio_vec array to populate, or NULL to only count the entries
+ * @pages: first page-array slot in the range
+ * @nr_pages: number of page-array slots available from @pages
+ * @offset: byte offset in @pages[0]
+ * @len: byte count to describe
+ * @max_bvec_len: maximum byte count in one bvec
+ *
+ * If @bvecs is not NULL, it must have room for the count returned by a
+ * preceding count-only call.
+ *
+ * Return: The number of populated or required bvecs, or zero if the range
+ * cannot be described from the supplied page-array slots.
+ */
+static inline unsigned int
+svc_pages_to_bvecs(struct bio_vec *bvecs, struct page *const *pages,
+ unsigned int nr_pages, unsigned int offset,
+ unsigned int len, unsigned int max_bvec_len)
+{
+ unsigned int remaining = len;
+ unsigned int nents = 0;
+
+ while (remaining && nr_pages) {
+ struct bio_vec *bv = bvecs ? &bvecs[nents] : NULL;
+ unsigned int advanced;
+ unsigned int bytes;
+
+ bytes = svc_pages_to_bvec(bv, pages, nr_pages, offset,
+ remaining, max_bvec_len);
+ if (!bytes)
+ break;
+ advanced = (offset + bytes) >> PAGE_SHIFT;
+ offset = offset_in_page(offset + bytes);
+ pages += advanced;
+ nr_pages -= advanced;
+ remaining -= bytes;
+ nents++;
+ }
+ return remaining ? 0 : nents;
+}
+
struct svc_deferred_req {
u32 prot; /* protocol (UDP or TCP) */
bool secure; /* RQ_SECURE of the original request */
--
2.43.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* Re: [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages
2026-10-06 9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
@ 2026-10-06 13:49 ` Chuck Lever
2026-10-08 22:44 ` J Louis Kaplan
0 siblings, 1 reply; 16+ messages in thread
From: Chuck Lever @ 2026-10-06 13:49 UTC (permalink / raw)
To: J Louis Kaplan, Dai Ngo, Jeff Layton, NeilBrown,
Olga Kornievskaia, Tom Talpey
Cc: linux-nfs, linux-rdma, Anna Schumaker, Jason Gunthorpe,
Leon Romanovsky, Trond Myklebust
On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> Add helpers to describe contiguous runs with multipage bvecs. These
> functions have no callers. They are used in later commits.
>
> Accept a caller-supplied maximum bvec length so transport callers
> can enforce direct memory access (DMA) mapping and device segment
> length limits. Also support counting entries without populating
> an array so callers can size their allocations using the same
> coalescing rules.
>
> Coalescing requires contiguous physical memory and page descriptors,
> and is disabled under KMSAN. Xen constraint handling remains WIP.
>
> Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
> Assisted-by: LLM
> ---
> include/linux/sunrpc/svc.h | 90 ++++++++++++++++++++++++++++++++++++++
> 1 file changed, 90 insertions(+)
>
> diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
> index dfadea50e0c6d..e651c2312a912 100644
> --- a/include/linux/sunrpc/svc.h
> +++ b/include/linux/sunrpc/svc.h
> @@ -11,6 +11,7 @@
> #ifndef SUNRPC_SVC_H
> #define SUNRPC_SVC_H
>
> +#include <linux/bvec.h>
> #include <linux/in.h>
> #include <linux/in6.h>
> #include <linux/sunrpc/types.h>
> @@ -376,6 +377,95 @@ static inline void svc_thread_init_status(struct
> svc_rqst *rqstp, int err)
> kthread_exit(1);
> }
>
> +/**
> + * svc_pages_to_bvec - build one bvec from physically contiguous pages
> + * @bv: bio_vec to initialize, or NULL to only measure the next extent
> + * @pages: first page-array slot in the range
> + * @nr_pages: number of page-array slots available from @pages
> + * @offset: byte offset in @pages[0]
> + * @len: maximum byte count to describe
> + * @max_bvec_len: maximum byte count in one bvec
> + *
> + * Return: The number of bytes described by @bv. The scan stops at a missing
> + * or physically non-contiguous page.
> + */
> +static inline unsigned int
> +svc_pages_to_bvec(struct bio_vec *bv, struct page *const *pages,
> + unsigned int nr_pages, unsigned int offset,
> + unsigned int len, unsigned int max_bvec_len)
Generally any function larger than two or three lines is to be
kept out of line. net/sunrpc/svc.c, perhaps, would be appropriate
here.
> +{
> + unsigned long first_pfn;
> + unsigned int bytes, used = 1;
> +
> + if (WARN_ON_ONCE(!nr_pages || offset >= PAGE_SIZE || !len ||
> + !max_bvec_len || !pages[0]))
> + return 0;
> + first_pfn = page_to_pfn(pages[0]);
> +
> + bytes = min3(len, max_bvec_len,
> + (unsigned int)(PAGE_SIZE - offset));
> + while (bytes < len && bytes < max_bvec_len && used < nr_pages) {
> + // TODO: handle Xen merge constraints, see `bvec_try_merge_page`
> + // for reference, or unify helper functionality
> + if (IS_ENABLED(CONFIG_KMSAN))
> + break;
> +
> + if (!pages[used] ||
> + page_to_pfn(pages[used]) != first_pfn + used ||
> + pages[0] + used != pages[used])
> + break;
> + used++;
> + bytes = min3(len, max_bvec_len,
> + (used << PAGE_SHIFT) - offset);
> + }
> +
> + if (bv)
> + bvec_set_page(bv, pages[0], bytes, offset);
> + return bytes;
> +}
> +
> +/**
> + * svc_pages_to_bvecs - build bvecs for a service page-array range
> + * @bvecs: bio_vec array to populate, or NULL to only count the entries
> + * @pages: first page-array slot in the range
> + * @nr_pages: number of page-array slots available from @pages
> + * @offset: byte offset in @pages[0]
> + * @len: byte count to describe
> + * @max_bvec_len: maximum byte count in one bvec
> + *
> + * If @bvecs is not NULL, it must have room for the count returned by a
> + * preceding count-only call.
> + *
> + * Return: The number of populated or required bvecs, or zero if the range
> + * cannot be described from the supplied page-array slots.
> + */
> +static inline unsigned int
> +svc_pages_to_bvecs(struct bio_vec *bvecs, struct page *const *pages,
> + unsigned int nr_pages, unsigned int offset,
> + unsigned int len, unsigned int max_bvec_len)
> +{
> + unsigned int remaining = len;
> + unsigned int nents = 0;
> +
> + while (remaining && nr_pages) {
> + struct bio_vec *bv = bvecs ? &bvecs[nents] : NULL;
> + unsigned int advanced;
> + unsigned int bytes;
> +
> + bytes = svc_pages_to_bvec(bv, pages, nr_pages, offset,
> + remaining, max_bvec_len);
> + if (!bytes)
> + break;
> + advanced = (offset + bytes) >> PAGE_SHIFT;
> + offset = offset_in_page(offset + bytes);
> + pages += advanced;
> + nr_pages -= advanced;
> + remaining -= bytes;
> + nents++;
> + }
> + return remaining ? 0 : nents;
> +}
> +
> struct svc_deferred_req {
> u32 prot; /* protocol (UDP or TCP) */
> bool secure; /* RQ_SECURE of the original request */
> --
> 2.43.0
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 16+ messages in thread* Re: [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages
2026-10-06 13:49 ` Chuck Lever
@ 2026-10-08 22:44 ` J Louis Kaplan
0 siblings, 0 replies; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-08 22:44 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: anna, jgg, leon, linux-nfs, linux-rdma, trondmy
> Generally any function larger than two or three lines is to be kept
> out of line.
Thanks for explaining that, Chuck. I will fix as suggested in any
subsequent patch series.
^ permalink raw reply [flat|nested] 16+ messages in thread
* [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads
2026-10-06 9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
@ 2026-10-06 9:00 ` J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
` (4 subsequent siblings)
6 siblings, 0 replies; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-06 9:00 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan
Use the previously added helper functions for nfsd direct read buffers.
Before this change, the direct-read path builds one bvec per base page,
even when the destination pages are physically contiguous.
Coalesce these pages into multipage bvecs to reduce the number of
entries traversed when advancing the filesystem's read iterator.
Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
fs/nfsd/vfs.c | 11 ++++++++---
1 file changed, 8 insertions(+), 3 deletions(-)
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index c7dd94697f2ff..329bb7d410765 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1124,11 +1124,16 @@ nfsd_direct_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
v = 0;
total = dio_end - dio_start;
while (total && v < rqstp->rq_maxpages && page < rqstp->rq_page_end) {
- len = min_t(size_t, total, PAGE_SIZE);
- bvec_set_page(&rqstp->rq_bvec[v], *page, len, 0);
+ len = svc_pages_to_bvec(&rqstp->rq_bvec[v],
+ page,
+ rqstp->rq_page_end - page,
+ 0, min_t(size_t, total, UINT_MAX),
+ UINT_MAX);
+ if (!len)
+ break;
total -= len;
- ++page;
+ page += DIV_ROUND_UP(len, PAGE_SIZE);
++v;
}
rqstp->rq_next_page = page;
--
2.43.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies
2026-10-06 9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads J Louis Kaplan
@ 2026-10-06 9:00 ` J Louis Kaplan
2026-10-06 13:55 ` Chuck Lever
2026-10-06 9:00 ` [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks J Louis Kaplan
` (3 subsequent siblings)
6 siblings, 1 reply; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-06 9:00 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan
Mapping reply pagelists one page at a time uses a separate direct
memory access (DMA) mapping and scatter/gather entry (SGE) for
each base page, even when those pages are physically contiguous.
Instead, map contiguous runs as multipage bvecs, capped by the device's
mapping and segment length limits. Use the same coalescing rules
when counting SGEs for the pull-up decision, avoiding unnecessary
copying when the coalesced reply fits the Send SGE limit.
Virtual-mapping devices bypass device-limit check.
Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
include/linux/sunrpc/svc_rdma.h | 15 ++++
net/sunrpc/xprtrdma/svc_rdma_sendto.c | 112 +++++++++++++++-----------
2 files changed, 82 insertions(+), 45 deletions(-)
diff --git a/include/linux/sunrpc/svc_rdma.h b/include/linux/sunrpc/svc_rdma.h
index 71bd5bcc5ce2a..c296bf8b8b1e3 100644
--- a/include/linux/sunrpc/svc_rdma.h
+++ b/include/linux/sunrpc/svc_rdma.h
@@ -42,6 +42,7 @@
#ifndef SVC_RDMA_H
#define SVC_RDMA_H
+#include <linux/dma-mapping.h>
#include <linux/llist.h>
#include <linux/sunrpc/xdr.h>
#include <linux/sunrpc/svcsock.h>
@@ -134,6 +135,20 @@ static inline struct svcxprt_rdma *svc_rdma_rqst_rdma(struct svc_rqst *rqstp)
return container_of(xprt, struct svcxprt_rdma, sc_xprt);
}
+static inline unsigned int
+svc_rdma_max_bvec_len(const struct svcxprt_rdma *rdma)
+{
+ struct ib_device *device = rdma->sc_cm_id->device;
+
+ if (ib_uses_virt_dma(device))
+ return UINT_MAX;
+
+ /* Respect both DMA mapping and RDMA device segment limits. */
+ return min_t(size_t,
+ dma_max_mapping_size(device->dma_device),
+ ib_dma_max_seg_size(device));
+}
+
/*
* Default connection parameters
*/
diff --git a/net/sunrpc/xprtrdma/svc_rdma_sendto.c b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
index efa352ddf9d71..9f2e3ee3c2936 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_sendto.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
@@ -260,10 +260,8 @@ static void svc_rdma_send_ctxt_unmap(struct svcxprt_rdma *rdma,
trace_svcrdma_dma_unmap_page(&ctxt->sc_cid,
ctxt->sc_sges[i].addr,
ctxt->sc_sges[i].length);
- ib_dma_unmap_page(device,
- ctxt->sc_sges[i].addr,
- ctxt->sc_sges[i].length,
- DMA_TO_DEVICE);
+ ib_dma_unmap_bvec(device, ctxt->sc_sges[i].addr,
+ ctxt->sc_sges[i].length, DMA_TO_DEVICE);
}
}
@@ -730,44 +728,45 @@ svc_rdma_encode_reply_chunk(struct svc_rdma_recv_ctxt *rctxt,
}
struct svc_rdma_map_data {
- struct svcxprt_rdma *md_rdma;
struct svc_rdma_send_ctxt *md_ctxt;
+ unsigned int md_max_bvec_len;
};
/**
- * svc_rdma_page_dma_map - DMA map one page
+ * svc_rdma_bvec_dma_map - DMA map one physically contiguous bvec
* @data: pointer to arguments
- * @page: struct page to DMA map
- * @offset: offset into the page
- * @len: number of bytes to map
+ * @bv: contiguous range to DMA map
*
* Returns:
* %0 if DMA mapping was successful
- * %-EIO if the page cannot be DMA mapped
+ * %-EIO if the range cannot be DMA mapped
*/
-static int svc_rdma_page_dma_map(void *data, struct page *page,
- unsigned long offset, unsigned int len)
+static int svc_rdma_bvec_dma_map(void *data, struct bio_vec *bv)
{
struct svc_rdma_map_data *args = data;
- struct svcxprt_rdma *rdma = args->md_rdma;
struct svc_rdma_send_ctxt *ctxt = args->md_ctxt;
+ struct svcxprt_rdma *rdma = ctxt->sc_rdma;
struct ib_device *dev = rdma->sc_cm_id->device;
dma_addr_t dma_addr;
+ if (bv->bv_len > args->md_max_bvec_len)
+ return -EIO;
+ if (WARN_ON_ONCE(ctxt->sc_send_wr.num_sge >= rdma->sc_max_send_sges))
+ return -EIO;
++ctxt->sc_cur_sge_no;
- dma_addr = ib_dma_map_page(dev, page, offset, len, DMA_TO_DEVICE);
+ dma_addr = ib_dma_map_bvec(dev, bv, DMA_TO_DEVICE);
if (ib_dma_mapping_error(dev, dma_addr))
goto out_maperr;
- trace_svcrdma_dma_map_page(&ctxt->sc_cid, dma_addr, len);
+ trace_svcrdma_dma_map_page(&ctxt->sc_cid, dma_addr, bv->bv_len);
ctxt->sc_sges[ctxt->sc_cur_sge_no].addr = dma_addr;
- ctxt->sc_sges[ctxt->sc_cur_sge_no].length = len;
+ ctxt->sc_sges[ctxt->sc_cur_sge_no].length = bv->bv_len;
ctxt->sc_send_wr.num_sge++;
return 0;
out_maperr:
- trace_svcrdma_dma_map_err(&ctxt->sc_cid, dma_addr, len);
+ trace_svcrdma_dma_map_err(&ctxt->sc_cid, dma_addr, bv->bv_len);
return -EIO;
}
@@ -776,20 +775,18 @@ static int svc_rdma_page_dma_map(void *data, struct page *page,
* @data: pointer to arguments
* @iov: kvec to DMA map
*
- * ib_dma_map_page() is used here because svc_rdma_dma_unmap()
- * handles DMA-unmap and it uses ib_dma_unmap_page() exclusively.
- *
* Returns:
* %0 if DMA mapping was successful
* %-EIO if the iovec cannot be DMA mapped
*/
static int svc_rdma_iov_dma_map(void *data, const struct kvec *iov)
{
+ struct bio_vec bv;
+
if (!iov->iov_len)
return 0;
- return svc_rdma_page_dma_map(data, virt_to_page(iov->iov_base),
- offset_in_page(iov->iov_base),
- iov->iov_len);
+ bvec_set_virt(&bv, iov->iov_base, iov->iov_len);
+ return svc_rdma_bvec_dma_map(data, &bv);
}
/**
@@ -798,7 +795,7 @@ static int svc_rdma_iov_dma_map(void *data, const struct kvec *iov)
* @data: pointer to arguments
*
* Returns:
- * %0 if DMA mapping was successful
+ * The number of mapped bytes if DMA mapping was successful
* %-EIO if DMA mapping failed
*
* On failure, any DMA mappings that have been already done must be
@@ -806,9 +803,11 @@ static int svc_rdma_iov_dma_map(void *data, const struct kvec *iov)
*/
static int svc_rdma_xb_dma_map(const struct xdr_buf *xdr, void *data)
{
- unsigned int len, remaining;
+ struct svc_rdma_map_data *args = data;
+ unsigned int remaining;
unsigned long pageoff;
struct page **ppages;
+ unsigned int nr_pages;
int ret;
ret = svc_rdma_iov_dma_map(data, &xdr->head[0]);
@@ -818,15 +817,25 @@ static int svc_rdma_xb_dma_map(const struct xdr_buf *xdr, void *data)
ppages = xdr->pages + (xdr->page_base >> PAGE_SHIFT);
pageoff = offset_in_page(xdr->page_base);
remaining = xdr->page_len;
+ nr_pages = DIV_ROUND_UP(pageoff + remaining, PAGE_SIZE);
while (remaining) {
- len = min_t(u32, PAGE_SIZE - pageoff, remaining);
-
- ret = svc_rdma_page_dma_map(data, *ppages++, pageoff, len);
+ struct bio_vec bv;
+ unsigned int bytes = svc_pages_to_bvec(&bv, ppages,
+ nr_pages, pageoff, remaining,
+ args->md_max_bvec_len);
+ unsigned int advanced =
+ (pageoff + bytes) >> PAGE_SHIFT;
+
+ if (!bytes)
+ return -EIO;
+ ret = svc_rdma_bvec_dma_map(data, &bv);
if (ret < 0)
return ret;
- remaining -= len;
- pageoff = 0;
+ pageoff = offset_in_page(pageoff + bytes);
+ ppages += advanced;
+ nr_pages -= advanced;
+ remaining -= bytes;
}
ret = svc_rdma_iov_dma_map(data, &xdr->tail[0]);
@@ -836,10 +845,27 @@ static int svc_rdma_xb_dma_map(const struct xdr_buf *xdr, void *data)
return xdr->len;
}
+static unsigned int
+svc_rdma_xb_page_sges(const struct xdr_buf *xdr,
+ unsigned int max_bvec_len)
+{
+ const unsigned long offset = offset_in_page(xdr->page_base);
+ const unsigned int nr_pages =
+ DIV_ROUND_UP(offset + xdr->page_len, PAGE_SIZE);
+ struct page *const *ppages;
+
+ if (!xdr->page_len)
+ return 0;
+ ppages = xdr->pages + (xdr->page_base >> PAGE_SHIFT);
+ return svc_pages_to_bvecs(NULL, ppages, nr_pages, offset,
+ xdr->page_len, max_bvec_len);
+}
+
struct svc_rdma_pullup_data {
u8 *pd_dest;
unsigned int pd_length;
unsigned int pd_num_sges;
+ unsigned int pd_max_bvec_len;
};
/**
@@ -848,25 +874,22 @@ struct svc_rdma_pullup_data {
* @data: pointer to arguments
*
* Returns:
- * Number of SGEs needed to Send the contents of @xdr inline
+ * %0 if the SGE count was updated
+ * %-EIO if the page list cannot be described
*/
static int svc_rdma_xb_count_sges(const struct xdr_buf *xdr,
void *data)
{
struct svc_rdma_pullup_data *args = data;
- unsigned int remaining;
- unsigned long offset;
+ unsigned int page_sges =
+ svc_rdma_xb_page_sges(xdr, args->pd_max_bvec_len);
if (xdr->head[0].iov_len)
++args->pd_num_sges;
- offset = offset_in_page(xdr->page_base);
- remaining = xdr->page_len;
- while (remaining) {
- ++args->pd_num_sges;
- remaining -= min_t(u32, PAGE_SIZE - offset, remaining);
- offset = 0;
- }
+ if (xdr->page_len && !page_sges)
+ return -EIO;
+ args->pd_num_sges += page_sges;
if (xdr->tail[0].iov_len)
++args->pd_num_sges;
@@ -896,6 +919,7 @@ static int svc_rdma_check_pull_up(const struct svcxprt_rdma *rdma,
struct svc_rdma_pullup_data args = {
.pd_length = sctxt->sc_hdrbuf.len,
.pd_num_sges = 1,
+ .pd_max_bvec_len = svc_rdma_max_bvec_len(rdma),
};
int ret;
@@ -964,7 +988,6 @@ static int svc_rdma_xb_linearize(const struct xdr_buf *xdr,
/**
* svc_rdma_pull_up_reply_msg - Copy Reply into a single buffer
- * @rdma: controlling transport
* @sctxt: send_ctxt for the Send WR; xprt hdr is already prepared
* @write_pcl: Write chunk list provided by client
* @xdr: prepared xdr_buf containing RPC message
@@ -979,8 +1002,7 @@ static int svc_rdma_xb_linearize(const struct xdr_buf *xdr,
* %0 if pull-up was successful
* %-EMSGSIZE if a buffer manipulation problem occurred
*/
-static int svc_rdma_pull_up_reply_msg(const struct svcxprt_rdma *rdma,
- struct svc_rdma_send_ctxt *sctxt,
+static int svc_rdma_pull_up_reply_msg(struct svc_rdma_send_ctxt *sctxt,
const struct svc_rdma_pcl *write_pcl,
const struct xdr_buf *xdr)
{
@@ -1021,8 +1043,8 @@ int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
const struct xdr_buf *xdr)
{
struct svc_rdma_map_data args = {
- .md_rdma = rdma,
.md_ctxt = sctxt,
+ .md_max_bvec_len = svc_rdma_max_bvec_len(rdma),
};
int ret;
@@ -1043,7 +1065,7 @@ int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
if (ret < 0)
return ret;
if (ret)
- return svc_rdma_pull_up_reply_msg(rdma, sctxt, write_pcl, xdr);
+ return svc_rdma_pull_up_reply_msg(sctxt, write_pcl, xdr);
return pcl_process_nonpayloads(write_pcl, xdr,
svc_rdma_xb_dma_map, &args);
--
2.43.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* Re: [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies
2026-10-06 9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
@ 2026-10-06 13:55 ` Chuck Lever
2026-10-08 22:54 ` J Louis Kaplan
0 siblings, 1 reply; 16+ messages in thread
From: Chuck Lever @ 2026-10-06 13:55 UTC (permalink / raw)
To: J Louis Kaplan, Dai Ngo, Jeff Layton, NeilBrown,
Olga Kornievskaia, Tom Talpey
Cc: linux-nfs, linux-rdma, Anna Schumaker, Jason Gunthorpe,
Leon Romanovsky, Trond Myklebust
On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> Mapping reply pagelists one page at a time uses a separate direct
> memory access (DMA) mapping and scatter/gather entry (SGE) for
> each base page, even when those pages are physically contiguous.
>
> Instead, map contiguous runs as multipage bvecs, capped by the device's
> mapping and segment length limits. Use the same coalescing rules
> when counting SGEs for the pull-up decision, avoiding unnecessary
> copying when the coalesced reply fits the Send SGE limit.
>
> Virtual-mapping devices bypass device-limit check.
>
> Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
> Assisted-by: LLM
> ---
> include/linux/sunrpc/svc_rdma.h | 15 ++++
> net/sunrpc/xprtrdma/svc_rdma_sendto.c | 112 +++++++++++++++-----------
> 2 files changed, 82 insertions(+), 45 deletions(-)
>
> diff --git a/include/linux/sunrpc/svc_rdma.h b/include/linux/sunrpc/svc_rdma.h
> index 71bd5bcc5ce2a..c296bf8b8b1e3 100644
> --- a/include/linux/sunrpc/svc_rdma.h
> +++ b/include/linux/sunrpc/svc_rdma.h
> @@ -42,6 +42,7 @@
>
> #ifndef SVC_RDMA_H
> #define SVC_RDMA_H
> +#include <linux/dma-mapping.h>
> #include <linux/llist.h>
> #include <linux/sunrpc/xdr.h>
> #include <linux/sunrpc/svcsock.h>
> @@ -134,6 +135,20 @@ static inline struct svcxprt_rdma
> *svc_rdma_rqst_rdma(struct svc_rqst *rqstp)
> return container_of(xprt, struct svcxprt_rdma, sc_xprt);
> }
>
> +static inline unsigned int
> +svc_rdma_max_bvec_len(const struct svcxprt_rdma *rdma)
There are some interesting infrastructural changes in this series.
Using bvecs through the svcrdma code has been on my to-do list
for quite some time.
> +{
> + struct ib_device *device = rdma->sc_cm_id->device;
> +
> + if (ib_uses_virt_dma(device))
> + return UINT_MAX;
> +
> + /* Respect both DMA mapping and RDMA device segment limits. */
> + return min_t(size_t,
> + dma_max_mapping_size(device->dma_device),
> + ib_dma_max_seg_size(device));
> +}
> +
> /*
> * Default connection parameters
> */
> diff --git a/net/sunrpc/xprtrdma/svc_rdma_sendto.c
> b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
> index efa352ddf9d71..9f2e3ee3c2936 100644
> --- a/net/sunrpc/xprtrdma/svc_rdma_sendto.c
> +++ b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
> @@ -260,10 +260,8 @@ static void svc_rdma_send_ctxt_unmap(struct
> svcxprt_rdma *rdma,
> trace_svcrdma_dma_unmap_page(&ctxt->sc_cid,
> ctxt->sc_sges[i].addr,
> ctxt->sc_sges[i].length);
> - ib_dma_unmap_page(device,
> - ctxt->sc_sges[i].addr,
> - ctxt->sc_sges[i].length,
> - DMA_TO_DEVICE);
> + ib_dma_unmap_bvec(device, ctxt->sc_sges[i].addr,
> + ctxt->sc_sges[i].length, DMA_TO_DEVICE);
> }
> }
>
> @@ -730,44 +728,45 @@ svc_rdma_encode_reply_chunk(struct
> svc_rdma_recv_ctxt *rctxt,
> }
>
> struct svc_rdma_map_data {
> - struct svcxprt_rdma *md_rdma;
> struct svc_rdma_send_ctxt *md_ctxt;
> + unsigned int md_max_bvec_len;
> };
>
> /**
> - * svc_rdma_page_dma_map - DMA map one page
> + * svc_rdma_bvec_dma_map - DMA map one physically contiguous bvec
> * @data: pointer to arguments
> - * @page: struct page to DMA map
> - * @offset: offset into the page
> - * @len: number of bytes to map
> + * @bv: contiguous range to DMA map
> *
> * Returns:
> * %0 if DMA mapping was successful
> - * %-EIO if the page cannot be DMA mapped
> + * %-EIO if the range cannot be DMA mapped
> */
> -static int svc_rdma_page_dma_map(void *data, struct page *page,
> - unsigned long offset, unsigned int len)
> +static int svc_rdma_bvec_dma_map(void *data, struct bio_vec *bv)
> {
> struct svc_rdma_map_data *args = data;
> - struct svcxprt_rdma *rdma = args->md_rdma;
> struct svc_rdma_send_ctxt *ctxt = args->md_ctxt;
> + struct svcxprt_rdma *rdma = ctxt->sc_rdma;
> struct ib_device *dev = rdma->sc_cm_id->device;
> dma_addr_t dma_addr;
>
> + if (bv->bv_len > args->md_max_bvec_len)
> + return -EIO;
> + if (WARN_ON_ONCE(ctxt->sc_send_wr.num_sge >= rdma->sc_max_send_sges))
> + return -EIO;
> ++ctxt->sc_cur_sge_no;
>
> - dma_addr = ib_dma_map_page(dev, page, offset, len, DMA_TO_DEVICE);
> + dma_addr = ib_dma_map_bvec(dev, bv, DMA_TO_DEVICE);
> if (ib_dma_mapping_error(dev, dma_addr))
> goto out_maperr;
>
> - trace_svcrdma_dma_map_page(&ctxt->sc_cid, dma_addr, len);
> + trace_svcrdma_dma_map_page(&ctxt->sc_cid, dma_addr, bv->bv_len);
> ctxt->sc_sges[ctxt->sc_cur_sge_no].addr = dma_addr;
> - ctxt->sc_sges[ctxt->sc_cur_sge_no].length = len;
> + ctxt->sc_sges[ctxt->sc_cur_sge_no].length = bv->bv_len;
> ctxt->sc_send_wr.num_sge++;
> return 0;
>
> out_maperr:
> - trace_svcrdma_dma_map_err(&ctxt->sc_cid, dma_addr, len);
> + trace_svcrdma_dma_map_err(&ctxt->sc_cid, dma_addr, bv->bv_len);
> return -EIO;
> }
>
> @@ -776,20 +775,18 @@ static int svc_rdma_page_dma_map(void *data,
> struct page *page,
> * @data: pointer to arguments
> * @iov: kvec to DMA map
> *
> - * ib_dma_map_page() is used here because svc_rdma_dma_unmap()
> - * handles DMA-unmap and it uses ib_dma_unmap_page() exclusively.
> - *
> * Returns:
> * %0 if DMA mapping was successful
> * %-EIO if the iovec cannot be DMA mapped
> */
> static int svc_rdma_iov_dma_map(void *data, const struct kvec *iov)
> {
> + struct bio_vec bv;
> +
> if (!iov->iov_len)
> return 0;
> - return svc_rdma_page_dma_map(data, virt_to_page(iov->iov_base),
> - offset_in_page(iov->iov_base),
> - iov->iov_len);
> + bvec_set_virt(&bv, iov->iov_base, iov->iov_len);
> + return svc_rdma_bvec_dma_map(data, &bv);
> }
>
> /**
> @@ -798,7 +795,7 @@ static int svc_rdma_iov_dma_map(void *data, const
> struct kvec *iov)
> * @data: pointer to arguments
> *
> * Returns:
> - * %0 if DMA mapping was successful
> + * The number of mapped bytes if DMA mapping was successful
> * %-EIO if DMA mapping failed
> *
> * On failure, any DMA mappings that have been already done must be
> @@ -806,9 +803,11 @@ static int svc_rdma_iov_dma_map(void *data, const
> struct kvec *iov)
> */
> static int svc_rdma_xb_dma_map(const struct xdr_buf *xdr, void *data)
> {
> - unsigned int len, remaining;
> + struct svc_rdma_map_data *args = data;
> + unsigned int remaining;
> unsigned long pageoff;
> struct page **ppages;
> + unsigned int nr_pages;
> int ret;
>
> ret = svc_rdma_iov_dma_map(data, &xdr->head[0]);
> @@ -818,15 +817,25 @@ static int svc_rdma_xb_dma_map(const struct
> xdr_buf *xdr, void *data)
> ppages = xdr->pages + (xdr->page_base >> PAGE_SHIFT);
> pageoff = offset_in_page(xdr->page_base);
> remaining = xdr->page_len;
> + nr_pages = DIV_ROUND_UP(pageoff + remaining, PAGE_SIZE);
> while (remaining) {
> - len = min_t(u32, PAGE_SIZE - pageoff, remaining);
> -
> - ret = svc_rdma_page_dma_map(data, *ppages++, pageoff, len);
> + struct bio_vec bv;
> + unsigned int bytes = svc_pages_to_bvec(&bv, ppages,
> + nr_pages, pageoff, remaining,
> + args->md_max_bvec_len);
> + unsigned int advanced =
> + (pageoff + bytes) >> PAGE_SHIFT;
> +
> + if (!bytes)
> + return -EIO;
> + ret = svc_rdma_bvec_dma_map(data, &bv);
> if (ret < 0)
> return ret;
>
> - remaining -= len;
> - pageoff = 0;
> + pageoff = offset_in_page(pageoff + bytes);
> + ppages += advanced;
> + nr_pages -= advanced;
> + remaining -= bytes;
> }
>
> ret = svc_rdma_iov_dma_map(data, &xdr->tail[0]);
> @@ -836,10 +845,27 @@ static int svc_rdma_xb_dma_map(const struct
> xdr_buf *xdr, void *data)
> return xdr->len;
> }
>
> +static unsigned int
> +svc_rdma_xb_page_sges(const struct xdr_buf *xdr,
> + unsigned int max_bvec_len)
> +{
> + const unsigned long offset = offset_in_page(xdr->page_base);
> + const unsigned int nr_pages =
> + DIV_ROUND_UP(offset + xdr->page_len, PAGE_SIZE);
> + struct page *const *ppages;
> +
> + if (!xdr->page_len)
> + return 0;
> + ppages = xdr->pages + (xdr->page_base >> PAGE_SHIFT);
> + return svc_pages_to_bvecs(NULL, ppages, nr_pages, offset,
> + xdr->page_len, max_bvec_len);
> +}
> +
> struct svc_rdma_pullup_data {
> u8 *pd_dest;
> unsigned int pd_length;
> unsigned int pd_num_sges;
> + unsigned int pd_max_bvec_len;
> };
>
> /**
> @@ -848,25 +874,22 @@ struct svc_rdma_pullup_data {
> * @data: pointer to arguments
> *
> * Returns:
> - * Number of SGEs needed to Send the contents of @xdr inline
> + * %0 if the SGE count was updated
> + * %-EIO if the page list cannot be described
> */
> static int svc_rdma_xb_count_sges(const struct xdr_buf *xdr,
> void *data)
> {
> struct svc_rdma_pullup_data *args = data;
> - unsigned int remaining;
> - unsigned long offset;
> + unsigned int page_sges =
> + svc_rdma_xb_page_sges(xdr, args->pd_max_bvec_len);
>
> if (xdr->head[0].iov_len)
> ++args->pd_num_sges;
>
> - offset = offset_in_page(xdr->page_base);
> - remaining = xdr->page_len;
> - while (remaining) {
> - ++args->pd_num_sges;
> - remaining -= min_t(u32, PAGE_SIZE - offset, remaining);
> - offset = 0;
> - }
> + if (xdr->page_len && !page_sges)
> + return -EIO;
> + args->pd_num_sges += page_sges;
>
> if (xdr->tail[0].iov_len)
> ++args->pd_num_sges;
> @@ -896,6 +919,7 @@ static int svc_rdma_check_pull_up(const struct
> svcxprt_rdma *rdma,
> struct svc_rdma_pullup_data args = {
> .pd_length = sctxt->sc_hdrbuf.len,
> .pd_num_sges = 1,
> + .pd_max_bvec_len = svc_rdma_max_bvec_len(rdma),
> };
> int ret;
>
> @@ -964,7 +988,6 @@ static int svc_rdma_xb_linearize(const struct xdr_buf *xdr,
>
> /**
> * svc_rdma_pull_up_reply_msg - Copy Reply into a single buffer
> - * @rdma: controlling transport
> * @sctxt: send_ctxt for the Send WR; xprt hdr is already prepared
> * @write_pcl: Write chunk list provided by client
> * @xdr: prepared xdr_buf containing RPC message
> @@ -979,8 +1002,7 @@ static int svc_rdma_xb_linearize(const struct xdr_buf *xdr,
> * %0 if pull-up was successful
> * %-EMSGSIZE if a buffer manipulation problem occurred
> */
> -static int svc_rdma_pull_up_reply_msg(const struct svcxprt_rdma *rdma,
> - struct svc_rdma_send_ctxt *sctxt,
> +static int svc_rdma_pull_up_reply_msg(struct svc_rdma_send_ctxt *sctxt,
> const struct svc_rdma_pcl *write_pcl,
> const struct xdr_buf *xdr)
> {
> @@ -1021,8 +1043,8 @@ int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
> const struct xdr_buf *xdr)
> {
> struct svc_rdma_map_data args = {
> - .md_rdma = rdma,
> .md_ctxt = sctxt,
> + .md_max_bvec_len = svc_rdma_max_bvec_len(rdma),
> };
> int ret;
>
> @@ -1043,7 +1065,7 @@ int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
> if (ret < 0)
> return ret;
> if (ret)
> - return svc_rdma_pull_up_reply_msg(rdma, sctxt, write_pcl, xdr);
> + return svc_rdma_pull_up_reply_msg(sctxt, write_pcl, xdr);
>
> return pcl_process_nonpayloads(write_pcl, xdr,
> svc_rdma_xb_dma_map, &args);
> --
> 2.43.0
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 16+ messages in thread* Re: [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies
2026-10-06 13:55 ` Chuck Lever
@ 2026-10-08 22:54 ` J Louis Kaplan
0 siblings, 0 replies; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-08 22:54 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: anna, jgg, leon, linux-nfs, linux-rdma, trondmy
> There are some interesting infrastructural changes in this series.
> Using bvecs through the svcrdma code has been on my to-do list
> for quite some time.
Glad to hear that. I'm happy to split out smaller parts of work from
this patch series if helpful.
^ permalink raw reply [flat|nested] 16+ messages in thread
* [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks
2026-10-06 9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
` (2 preceding siblings ...)
2026-10-06 9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
@ 2026-10-06 9:00 ` J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks J Louis Kaplan
` (2 subsequent siblings)
6 siblings, 0 replies; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-06 9:00 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan
When sending replies with remote direct memory access (RDMA)
Writes, the pagelist is represented by one bvec per base page.
This prevents contiguous pages from sharing a buffer descriptor.
Coalesce contiguous runs within each remote segment, subject to
device length limits. Count the resulting bvecs before acquiring
the context so its allocation reflects the coalesced entry count.
The counting pass leaves the payload cursor untouched.
Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
net/sunrpc/xprtrdma/svc_rdma_rw.c | 77 +++++++++++++++++--------------
1 file changed, 43 insertions(+), 34 deletions(-)
diff --git a/net/sunrpc/xprtrdma/svc_rdma_rw.c b/net/sunrpc/xprtrdma/svc_rdma_rw.c
index 63d00f0cc1dbd..a2e1c693b1bc1 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_rw.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_rw.c
@@ -445,46 +445,48 @@ static int svc_rdma_post_chunk_ctxt(struct svcxprt_rdma *rdma,
return 0;
}
-/* Build a bvec that covers one kvec in an xdr_buf.
+/*
+ * With a NULL @ctxt, bvec constructors count entries without advancing
+ * @info. Otherwise they populate the context and advance @info.
*/
-static void svc_rdma_vec_to_bvec(struct svc_rdma_write_info *info,
- unsigned int len,
- struct svc_rdma_rw_ctxt *ctxt)
+static unsigned int
+svc_rdma_vec_to_bvec(struct svc_rdma_write_info *info, unsigned int len,
+ struct svc_rdma_rw_ctxt *ctxt)
{
+ if (!ctxt)
+ return 1;
+
bvec_set_virt(&ctxt->rw_bvec[0], info->wi_base, len);
info->wi_base += len;
ctxt->rw_nents = 1;
+ return 1;
}
/* Build a bvec array that covers part of an xdr_buf's pagelist.
+ * If @ctxt is NULL, only count the required entries.
*/
-static void svc_rdma_pagelist_to_bvec(struct svc_rdma_write_info *info,
- unsigned int remaining,
- struct svc_rdma_rw_ctxt *ctxt)
+static unsigned int
+svc_rdma_pagelist_to_bvec(struct svc_rdma_write_info *info,
+ unsigned int remaining,
+ struct svc_rdma_rw_ctxt *ctxt)
{
- unsigned int bvec_idx, bvec_len, page_off, page_no;
const struct xdr_buf *xdr = info->wi_xdr;
- struct page **page;
-
- page_off = info->wi_next_off + xdr->page_base;
- page_no = page_off >> PAGE_SHIFT;
- page_off = offset_in_page(page_off);
- page = xdr->pages + page_no;
- info->wi_next_off += remaining;
- bvec_idx = 0;
- do {
- bvec_len = min_t(unsigned int, remaining,
- PAGE_SIZE - page_off);
- bvec_set_page(&ctxt->rw_bvec[bvec_idx], *page, bvec_len,
- page_off);
- remaining -= bvec_len;
- page_off = 0;
- bvec_idx++;
- page++;
- } while (remaining);
-
- ctxt->rw_nents = bvec_idx;
+ unsigned int page_pos = info->wi_next_off + xdr->page_base;
+ unsigned int page_no = page_pos >> PAGE_SHIFT;
+ unsigned int max_bvec_len = svc_rdma_max_bvec_len(info->wi_rdma);
+ unsigned int page_off = offset_in_page(page_pos);
+ unsigned int nr_pages = DIV_ROUND_UP(page_off + remaining, PAGE_SIZE);
+ struct bio_vec *bvecs = ctxt ? ctxt->rw_bvec : NULL;
+ unsigned int nents;
+
+ nents = svc_pages_to_bvecs(bvecs, xdr->pages + page_no, nr_pages,
+ page_off, remaining, max_bvec_len);
+ if (ctxt) {
+ ctxt->rw_nents = nents;
+ info->wi_next_off += remaining;
+ }
+ return nents;
}
/* Construct RDMA Write WRs to send a portion of an xdr_buf containing
@@ -492,9 +494,9 @@ static void svc_rdma_pagelist_to_bvec(struct svc_rdma_write_info *info,
*/
static int
svc_rdma_build_writes(struct svc_rdma_write_info *info,
- void (*constructor)(struct svc_rdma_write_info *info,
- unsigned int len,
- struct svc_rdma_rw_ctxt *ctxt),
+ unsigned int (*constructor)(struct svc_rdma_write_info *,
+ unsigned int,
+ struct svc_rdma_rw_ctxt *),
unsigned int remaining)
{
struct svc_rdma_chunk_ctxt *cc = &info->wi_cc;
@@ -504,6 +506,7 @@ svc_rdma_build_writes(struct svc_rdma_write_info *info,
int ret;
do {
+ unsigned int nr_bvec;
unsigned int write_len;
u64 offset;
@@ -514,12 +517,18 @@ svc_rdma_build_writes(struct svc_rdma_write_info *info,
write_len = min(remaining, seg->rs_length - info->wi_seg_off);
if (!write_len)
goto out_overflow;
- ctxt = svc_rdma_get_rw_ctxt(rdma,
- (write_len >> PAGE_SHIFT) + 2);
+ nr_bvec = constructor(info, write_len, NULL);
+ if (!nr_bvec)
+ return -EIO;
+
+ ctxt = svc_rdma_get_rw_ctxt(rdma, nr_bvec);
if (!ctxt)
return -ENOMEM;
- constructor(info, write_len, ctxt);
+ if (WARN_ON_ONCE(constructor(info, write_len, ctxt) != nr_bvec)) {
+ svc_rdma_put_rw_ctxt(rdma, ctxt);
+ return -EIO;
+ }
offset = seg->rs_offset + info->wi_seg_off;
ret = svc_rdma_rw_ctx_init(rdma, ctxt, offset, seg->rs_handle,
write_len, DMA_TO_DEVICE);
--
2.43.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks
2026-10-06 9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
` (3 preceding siblings ...)
2026-10-06 9:00 ` [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks J Louis Kaplan
@ 2026-10-06 9:00 ` J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
2026-10-06 13:47 ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled Chuck Lever
6 siblings, 0 replies; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-06 9:00 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan
Read chunks currently describe their destination buffers with one
bvec per base page. Coalesce contiguous destination pages to reduce
the number of entries passed to the transport mapping code.
Derive the number of consumed pages from the final page cursor,
since one bvec can now cover several pages.
Limit each iWARP Read context to one memory region: the core
registration code cannot split a coalesced entry between regions.
The force_mr case on other transports remains unresolved.
Update comment for multipage bvecs.
Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
drivers/infiniband/core/rw.c | 7 +--
net/sunrpc/xprtrdma/svc_rdma_rw.c | 101 ++++++++++++++++++------------
2 files changed, 65 insertions(+), 43 deletions(-)
diff --git a/drivers/infiniband/core/rw.c b/drivers/infiniband/core/rw.c
index 4fafe393a48c7..f9e0a5b8ea806 100644
--- a/drivers/infiniband/core/rw.c
+++ b/drivers/infiniband/core/rw.c
@@ -706,10 +706,9 @@ int rdma_rw_ctx_init_bvec(struct rdma_rw_ctx *ctx, struct ib_qp *qp,
* is a throughput optimization, not a correctness requirement.
* (iWARP, which does require MRs, is handled by the check above.)
*
- * The rdma_rw_io_needs_mr() gate is not used here because nr_bvec
- * is a raw page count that overstates DMA entry demand -- the bvec
- * caller has no post-DMA-coalescing segment count, and feeding the
- * inflated count into the MR path exhausts the pool on RDMA READs.
+ * We do not use max_sgl_rd to switch non-iWARP READs to MRs. Doing so
+ * would require the MR builder to split multipage bvecs across MRs,
+ * which it does not yet support. The force_mr case is handled above.
*/
return rdma_rw_init_map_wrs_bvec(ctx, qp, bvecs, nr_bvec, &iter,
remote_addr, rkey, dir);
diff --git a/net/sunrpc/xprtrdma/svc_rdma_rw.c b/net/sunrpc/xprtrdma/svc_rdma_rw.c
index a2e1c693b1bc1..074a75d251c96 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_rw.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_rw.c
@@ -789,54 +789,78 @@ static int svc_rdma_build_read_segment(struct svc_rqst *rqstp,
{
struct svcxprt_rdma *rdma = svc_rdma_rqst_rdma(rqstp);
struct svc_rdma_chunk_ctxt *cc = &head->rc_cc;
- unsigned int bvec_idx, nr_bvec, seg_len, len, total;
+ struct ib_device *dev = rdma->sc_cm_id->device;
+ unsigned int max_bvec_len = svc_rdma_max_bvec_len(rdma);
+ u64 remote_offset = segment->rs_offset;
+ unsigned int nr_bvec, len, remaining, total;
+ unsigned int base_pages;
struct svc_rdma_rw_ctxt *ctxt;
int ret;
- len = segment->rs_length;
- if (check_add_overflow(head->rc_pageoff, len, &total))
+ remaining = segment->rs_length;
+ if (check_add_overflow(head->rc_pageoff, remaining, &total))
return -EINVAL;
- nr_bvec = PAGE_ALIGN(total) >> PAGE_SHIFT;
- ctxt = svc_rdma_get_rw_ctxt(rdma, nr_bvec);
- if (!ctxt)
- return -ENOMEM;
- ctxt->rw_nents = nr_bvec;
-
- for (bvec_idx = 0; bvec_idx < ctxt->rw_nents; bvec_idx++) {
- seg_len = min_t(unsigned int, len,
- PAGE_SIZE - head->rc_pageoff);
-
- if (!head->rc_pageoff)
- head->rc_page_count++;
+ base_pages = PAGE_ALIGN(total) >> PAGE_SHIFT;
+ if (!remaining || head->rc_curpage >= rqstp->rq_maxpages ||
+ base_pages > rqstp->rq_maxpages - head->rc_curpage)
+ goto out_overrun;
+
+ while (remaining) {
+ base_pages = DIV_ROUND_UP(head->rc_pageoff + remaining,
+ PAGE_SIZE);
+
+ /*
+ * The bvec MR builder cannot split a coalesced SG entry between
+ * MRs. iWARP always uses MRs for RDMA Reads, so keep each context
+ * within one MR. force_mr on other transports remains subject to
+ * the core limitation.
+ */
+ if (rdma_protocol_iwarp(dev, rdma->sc_port_num)) {
+ unsigned int nr_mrs = rdma_rw_mr_factor(dev,
+ rdma->sc_port_num,
+ base_pages);
- bvec_set_page(&ctxt->rw_bvec[bvec_idx],
- rqstp->rq_pages[head->rc_curpage],
- seg_len, head->rc_pageoff);
+ base_pages = DIV_ROUND_UP(base_pages, nr_mrs);
+ }
- head->rc_pageoff += seg_len;
- if (head->rc_pageoff == PAGE_SIZE) {
- head->rc_curpage++;
- head->rc_pageoff = 0;
+ len = min_t(unsigned int, remaining,
+ (base_pages << PAGE_SHIFT) - head->rc_pageoff);
+ nr_bvec = svc_pages_to_bvecs(NULL,
+ rqstp->rq_pages + head->rc_curpage,
+ base_pages, head->rc_pageoff, len,
+ max_bvec_len);
+ if (!nr_bvec)
+ goto out_overrun;
+ ctxt = svc_rdma_get_rw_ctxt(rdma, nr_bvec);
+ if (!ctxt)
+ return -ENOMEM;
+ ctxt->rw_nents = svc_pages_to_bvecs(ctxt->rw_bvec,
+ rqstp->rq_pages + head->rc_curpage,
+ base_pages, head->rc_pageoff, len,
+ max_bvec_len);
+ if (WARN_ON_ONCE(ctxt->rw_nents != nr_bvec)) {
+ svc_rdma_put_rw_ctxt(rdma, ctxt);
+ goto out_overrun;
}
- len -= seg_len;
+ total = head->rc_pageoff + len;
+ head->rc_curpage += total >> PAGE_SHIFT;
+ head->rc_pageoff = offset_in_page(total);
- if (len && ((head->rc_curpage + 1) > rqstp->rq_maxpages))
- goto out_put;
- }
+ ret = svc_rdma_rw_ctx_init(rdma, ctxt, remote_offset,
+ segment->rs_handle, len,
+ DMA_FROM_DEVICE);
+ if (ret < 0)
+ return -EIO;
- ret = svc_rdma_rw_ctx_init(rdma, ctxt, segment->rs_offset,
- segment->rs_handle, segment->rs_length,
- DMA_FROM_DEVICE);
- if (ret < 0)
- return -EIO;
+ list_add(&ctxt->rw_list, &cc->cc_rwctxts);
+ cc->cc_sqecount += ret;
+ remote_offset += len;
+ remaining -= len;
+ }
percpu_counter_inc(&svcrdma_stat_read);
-
- list_add(&ctxt->rw_list, &cc->cc_rwctxts);
- cc->cc_sqecount += ret;
return 0;
-out_put:
- svc_rdma_put_rw_ctxt(rdma, ctxt);
+out_overrun:
trace_svcrdma_page_overrun_err(&cc->cc_cid, head->rc_curpage);
return -EINVAL;
}
@@ -908,9 +932,6 @@ static int svc_rdma_copy_inline_range(struct svc_rqst *rqstp,
page_len = min_t(unsigned int, remaining,
PAGE_SIZE - head->rc_pageoff);
- if (!head->rc_pageoff)
- head->rc_page_count++;
-
dst = page_address(rqstp->rq_pages[head->rc_curpage]);
memcpy((unsigned char *)dst + head->rc_pageoff, src + offset, page_len);
@@ -1170,6 +1191,8 @@ static void svc_rdma_clear_rqst_pages(struct svc_rqst *rqstp,
{
unsigned int i;
+ /* The cursor identifies every page touched while rebuilding the call. */
+ head->rc_page_count = head->rc_curpage + !!head->rc_pageoff;
for (i = 0; i < head->rc_page_count; i++) {
head->rc_pages[i] = rqstp->rq_pages[i];
rqstp->rq_pages[i] = NULL;
--
2.43.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios
2026-10-06 9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
` (4 preceding siblings ...)
2026-10-06 9:00 ` [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks J Louis Kaplan
@ 2026-10-06 9:00 ` J Louis Kaplan
2026-10-06 14:00 ` Chuck Lever
2026-10-06 13:47 ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled Chuck Lever
6 siblings, 1 reply; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-06 9:00 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy, J Louis Kaplan
Bulk allocation of base pages does not guarantee physical
contiguity, limiting opportunities to coalesce service buffers.
Allocate large folios opportunistically and populate the service
page arrays with their constituent pages. Cap folio size at
64 KiB and allocation order at PAGE_ALLOC_COSTLY_ORDER, falling
back through smaller orders to bulk base-page allocation.
Retain one reference per populated page-array slot so existing
page-based ownership and release paths remain valid.
Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
net/sunrpc/svc_xprt.c | 44 ++++++++++++++++++++++++++++++++++++++-----
1 file changed, 39 insertions(+), 5 deletions(-)
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 9858dfcb846a8..90ec75a848da9 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -19,6 +19,7 @@
#include <linux/sunrpc/bc_xprt.h>
#include <linux/module.h>
#include <linux/netdevice.h>
+#include <linux/sizes.h>
#include <trace/events/sunrpc.h>
#define RPCDBG_FACILITY RPCDBG_SVCXPRT
@@ -708,20 +709,53 @@ static void svc_check_conn_limits(struct svc_serv *serv)
static bool svc_fill_pages(struct svc_rqst *rqstp, struct page **pages,
unsigned long npages)
{
- unsigned long filled, ret;
+ unsigned int order = min_t(unsigned int, PAGE_ALLOC_COSTLY_ORDER,
+ get_order(SZ_64K));
+ unsigned long filled = 0;
- for (filled = 0; filled < npages; filled = ret) {
- ret = alloc_pages_bulk(GFP_KERNEL, npages, pages);
- if (ret > filled)
+ /*
+ * Prefer folios no larger than 64 KiB or the page allocator's
+ * costly-order threshold. After a failure, use smaller orders for
+ * the rest of this fill; order zero is handled by the bulk allocator.
+ * All callers provide a range containing only NULL entries.
+ */
+ while (filled < npages) {
+ struct folio *folio;
+ unsigned int nr_pages;
+ unsigned long ret;
+
+ while (order && (1UL << order) > npages - filled)
+ order--;
+
+ if (order) {
+ folio = folio_alloc(GFP_NOWAIT, order);
+ if (!folio) {
+ order--;
+ continue;
+ }
+
+ nr_pages = folio_nr_pages(folio);
+ folio_ref_add(folio, nr_pages - 1);
+ for (unsigned int i = 0; i < nr_pages; i++)
+ pages[filled + i] = folio_page(folio, i);
+ filled += nr_pages;
+ continue;
+ }
+
+ ret = alloc_pages_bulk(GFP_KERNEL, npages - filled,
+ pages + filled);
+ if (ret) {
+ filled += ret;
/* Made progress, don't sleep yet */
continue;
+ }
set_current_state(TASK_IDLE);
if (svc_thread_should_stop(rqstp)) {
set_current_state(TASK_RUNNING);
return false;
}
- trace_svc_alloc_arg_err(npages, ret);
+ trace_svc_alloc_arg_err(npages, filled);
memalloc_retry_wait(GFP_KERNEL);
}
return true;
--
2.43.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* Re: [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios
2026-10-06 9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
@ 2026-10-06 14:00 ` Chuck Lever
2026-10-08 22:58 ` J Louis Kaplan
0 siblings, 1 reply; 16+ messages in thread
From: Chuck Lever @ 2026-10-06 14:00 UTC (permalink / raw)
To: J Louis Kaplan, dai.ngo, jlayton, neil, okorniev, tom
Cc: linux-nfs, linux-rdma, anna, jgg, leon, trondmy
On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> Bulk allocation of base pages does not guarantee physical
> contiguity, limiting opportunities to coalesce service buffers.
>
> Allocate large folios opportunistically and populate the service
> page arrays with their constituent pages. Cap folio size at
> 64 KiB and allocation order at PAGE_ALLOC_COSTLY_ORDER, falling
> back through smaller orders to bulk base-page allocation.
>
> Retain one reference per populated page-array slot so existing
> page-based ownership and release paths remain valid.
>
> Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
> Assisted-by: LLM
Given past experience with contention in the page free path, IMO
this one needs careful macro and micro benchmarking.
I wonder if this change should get merged first instead of last.
It's probably going to have broad impact and high risk, and will
need testing across all hardware platforms and on large- and small-
memory configurations.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios
2026-10-06 14:00 ` Chuck Lever
@ 2026-10-08 22:58 ` J Louis Kaplan
2026-10-09 15:28 ` Chuck Lever
0 siblings, 1 reply; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-08 22:58 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: anna, jgg, leon, linux-nfs, linux-rdma, trondmy
> Given past experience with contention in the page free path, IMO
> this one needs careful macro and micro benchmarking.
Makes sense to me after reading through the links you posted in the
other reply.
> I wonder if this change should get merged first instead of last.
> It's probably going to have broad impact and high risk, and will
> need testing across all hardware platforms and on large- and small-
> memory configurations.
Happy to do that. Should I consider a separate single-commit submission
for this change?
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios
2026-10-08 22:58 ` J Louis Kaplan
@ 2026-10-09 15:28 ` Chuck Lever
0 siblings, 0 replies; 16+ messages in thread
From: Chuck Lever @ 2026-10-09 15:28 UTC (permalink / raw)
To: J Louis Kaplan, Dai Ngo, Jeff Layton, NeilBrown,
Olga Kornievskaia, Tom Talpey
Cc: Anna Schumaker, Jason Gunthorpe, Leon Romanovsky, linux-nfs,
linux-rdma, Trond Myklebust
On Thu, Oct 8, 2026, at 6:58 PM, J Louis Kaplan wrote:
>> Given past experience with contention in the page free path, IMO
>> this one needs careful macro and micro benchmarking.
>
> Makes sense to me after reading through the links you posted in the
> other reply.
>
>> I wonder if this change should get merged first instead of last.
>> It's probably going to have broad impact and high risk, and will
>> need testing across all hardware platforms and on large- and small-
>> memory configurations.
>
> Happy to do that. Should I consider a separate single-commit submission
> for this change?
I'm rethinking the patch ordering suggestion. The biovec changes
by themselves are desirable refactoring. This change, while also
desirable in concept, needs guidance from the MM folks and might
take some time and careful thought.
The lock contention in page free is because of the way NFSD handles
pages in its receive and send buffers:
During the initial fill of rq_pages when each thread is created,
all of the array elements are empty, and a folio allocation is
likely to be efficient.
However, when handling a small NFS request, NFSD frees and
reallocates only one or two pages at the front of the thread's
receive and buffers. Refilling those using folio allocation is
likely to get a small folio that is not contiguous with the
pages that remain in rq_pages from previous allocations.
So over time, the front of the buffer tends to become a
handful of individual pages.
Allocating from a large folio and then freeing just a few
of those pages is an anti-pattern for the page allocator.
It's optimized for the case where an entire folio is
allocated and then freed together.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 16+ messages in thread
* Re: [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled
2026-10-06 9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
` (5 preceding siblings ...)
2026-10-06 9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
@ 2026-10-06 13:47 ` Chuck Lever
2026-10-08 22:50 ` J Louis Kaplan
6 siblings, 1 reply; 16+ messages in thread
From: Chuck Lever @ 2026-10-06 13:47 UTC (permalink / raw)
To: J Louis Kaplan, Dai Ngo, Jeff Layton, NeilBrown,
Olga Kornievskaia, Tom Talpey
Cc: linux-nfs, linux-rdma, Anna Schumaker, Jason Gunthorpe,
Leon Romanovsky, Trond Myklebust
On Tue, Oct 6, 2026, at 5:00 AM, J Louis Kaplan wrote:
> For NFSoRDMA using RoCE, and with IOMMU passthrough enabled, use of a
> 64KiB base page size was found to provide higher NFS read throughput
> than a 4KiB base page size. This patch series recovers some of that
> throughput with a 4k base page size by opportunistic use of large folios
> for the svc reply buffer.
> LLM Usage
> =========
>
> The idea of folio usage here was human-generated from observations of
> performance differences between 4KiB and 64KiB page size kernels.
Actually a similar approach has been tried before, but on x86_64:
https://lore.kernel.org/linux-nfs/20260319133610.2556826-1-cel@kernel.org/
And it had to be reverted:
- Mike Snitzer's report (4 Jun 2026):
https://lore.kernel.org/linux-nfs/aiHlPmeZq3WgMwoJ@kernel.org/
- Jonathan Flynn's benchmark and perf data (5 Jun 2026), showing server
CPU going from 8.54% to 76.35% with the commit present:
https://lore.kernel.org/linux-nfs/3cb119b4b2a8aada30c0c60286778a54@mail.gmail.com/
These two are the Closes: targets in the revert, which is a39f0ce0c9da
upstream and 284d5ba931a5 in stable.
Lock contention analysis and the decision to revert
- "[PATCH] svcrdma: Cap Read sink allocations at PAGE_ALLOC_COSTLY_ORDER"
(5 Jun 2026), whose description carries the zone->lock analysis:
https://lore.kernel.org/linux-nfs/20260606035722.83175-1-cel@kernel.org/
- Flynn's results for that fix (25.4 GiB/s, against 30.3 regressed and
73.9 reverted):
https://lore.kernel.org/linux-nfs/65a2cdb132b0c28e69a29955e3bd37e7@mail.gmail.com/
- Your reply saying the two failed fixes mean 18755b8c2f24 has to be
reverted (6 Jun 2026):
https://lore.kernel.org/linux-nfs/096a2b91-7a19-48da-a06a-dc60e7150956@app.fastmail.com/
The earlier failed fix, "[PATCH] svcrdma: Avoid direct reclaim when
allocating Read sink buffers", is at
https://lore.kernel.org/linux-nfs/20260605223118.75092-1-cel@kernel.org/.
So: Yes, we would like to ensure that svcrdma's RDMA Read buffers (at
least, and maybe RDMA Write) are contiguous so that the overhead of a
DMA mapping operation will be low. I'm not expert enough with the page
allocator's free path to address the lock contention problem, however.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 16+ messages in thread* Re: [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled
2026-10-06 13:47 ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled Chuck Lever
@ 2026-10-08 22:50 ` J Louis Kaplan
0 siblings, 0 replies; 16+ messages in thread
From: J Louis Kaplan @ 2026-10-08 22:50 UTC (permalink / raw)
To: cel, dai.ngo, jlayton, neil, okorniev, tom
Cc: anna, jgg, leon, linux-nfs, linux-rdma, trondmy
> Actually a similar approach has been tried before... And it had to be
> reverted
Thanks for sharing these useful links. They provide important context
that I was not aware of.
I'll have another read through the investigations and commit dates
to better understand the change timeline, but I wanted to share some
contention data I gathered prompted by your feedback.
During development of this patch series, I did notice lock contention
issues. I suspected the unbound workqueue but my analysis was shallow
and I didn't have a good idea how to improve the situation.
After a rebase I noticed a decrease in contention and found your commit
58202c29de936 in the git log which looked likely to be the beneficial
change.
I set up two test branches after reading your feedback. Both apply this
patch series (with minor conflict resolutions) to the older nfsd-testing
commit a39f0ce0c9da2 which reverts `Use contiguous pages for RDMA Read
sink buffers`. `ubwq` also reverts your patch 58202c29de936 whereas
`no-ubwq` does not.
For direct IO on the server and NFS reads in configuration A I mentioned
above (passthrough enabled), here are representative top lines of perf
contention data (spinlocks only).
ubwq (without 58202c29de936)
contended total wait max wait avg wait caller
202787 588.67 ms 334.01 us 2.90 us __rmqueue_pcplist+0x320
281818 586.91 ms 267.99 us 2.08 us free_pcppages_bulk+0x48
88218 100.65 ms 5.57 us 1.14 us rcu_report_qs_rdp+0x44
1124 11.07 ms 28.13 us 9.85 us tick_do_update_jiffies64+0x48
5773 4.92 ms 3.42 us 851 ns tmigr_update_events+0x184
1274 2.38 ms 66.01 us 1.87 us nfsd4_sequence_done+0x70
1272 2.36 ms 51.23 us 1.86 us nfsd4_sequence+0x114
1074 2.05 ms 5.57 us 1.90 us acpi_os_wait_semaphore+0x80
1074 1.68 ms 15.68 us 1.56 us free_pcppages_bulk+0x48
739 1.04 ms 4.80 us 1.41 us __queue_work+0xd4
424 745.28 us 8.99 us 1.76 us raw_spin_rq_lock_nested+0x3c
362 649.74 us 18.78 us 1.79 us __rmqueue_pcplist+0x320
... <truncated>
no-ubwq (with 58202c29de936):
contended total wait max wait avg wait caller
68113 82.67 ms 5.92 us 1.21 us rcu_report_qs_rdp+0x44
1604 2.41 ms 21.73 us 1.50 us nfsd4_sequence_done+0x70
1689 2.31 ms 26.21 us 1.37 us nfsd4_sequence+0x114
884 1.66 ms 15.87 us 1.88 us raw_spin_rq_lock_nested+0x3c
153 1.21 ms 16.48 us 7.93 us tick_do_update_jiffies64+0x48
311 566.99 us 14.43 us 1.82 us raw_spin_rq_lock_nested+0x3c
423 459.48 us 2.98 us 1.09 us force_qs_rnp+0x104
185 335.12 us 27.14 us 1.81 us __lwq_dequeue+0x34
40 288.12 us 20.35 us 7.20 us free_pcppages_bulk+0x48
220 285.37 us 3.97 us 1.30 us __queue_work+0xd4
284 242.78 us 2.14 us 854 ns tmigr_update_events+0x184
153 216.64 us 3.10 us 1.41 us acpi_os_wait_semaphore+0x80
164 216.03 us 26.43 us 1.32 us find_stateid_by_type+0x40
143 206.07 us 13.66 us 1.44 us nfsd4_sequence+0x238
95 118.24 us 2.40 us 1.24 us dma_pool_alloc+0x50
85 116.29 us 2.75 us 1.37 us acpi_os_signal_semaphore+0x88
86 101.37 us 5.54 us 1.18 us raw_spin_rq_lock_nested+0x3c
34 84.18 us 8.45 us 2.47 us mix_interrupt_randomness+0x104
72 77.31 us 1.95 us 1.07 us raw_spin_rq_lock_nested+0x3c
43 74.94 us 4.29 us 1.74 us raw_spin_rq_lock_nested+0x3c
51 72.09 us 3.33 us 1.41 us acpi_os_wait_semaphore+0x80
62 67.26 us 1.92 us 1.08 us dma_pool_free+0x3c
42 61.66 us 2.46 us 1.47 us acpi_os_wait_semaphore+0x80
39 55.87 us 2.94 us 1.43 us __queue_work+0xd4
21 50.65 us 15.23 us 2.41 us raw_spin_rq_lock_nested+0x3c
35 45.76 us 2.05 us 1.31 us raw_spin_rq_lock_nested+0x3c
15 24.16 us 3.93 us 1.61 us free_pcppages_bulk+0x48
... <truncated>
CPU usage was 7% lower in the no-ubwq branch. Throughput remained similar
in both branches for this test.
I'm still digesting the above (and more) data and trying to understand
how all the information fits together, but I wanted to share the above
in the meantime.
^ permalink raw reply [flat|nested] 16+ messages in thread