From: J Louis Kaplan <Louis.Kaplan@arm.com>
To: cel@kernel.org, dai.ngo@oracle.com, jlayton@kernel.org,
neil@brown.name, okorniev@redhat.com, tom@talpey.com
Cc: linux-nfs@vger.kernel.org, linux-rdma@vger.kernel.org,
anna@kernel.org, jgg@ziepe.ca, leon@kernel.org,
trondmy@kernel.org, J Louis Kaplan <Louis.Kaplan@arm.com>
Subject: [RFC PATCH 5/6] svcrdma: Coalesce contiguous pages in RDMA Read chunks
Date: Tue, 6 Oct 2026 10:00:27 +0100 [thread overview]
Message-ID: <20261006090028.3412544-6-Louis.Kaplan@arm.com> (raw)
In-Reply-To: <20261006090028.3412544-1-Louis.Kaplan@arm.com>
Read chunks currently describe their destination buffers with one
bvec per base page. Coalesce contiguous destination pages to reduce
the number of entries passed to the transport mapping code.
Derive the number of consumed pages from the final page cursor,
since one bvec can now cover several pages.
Limit each iWARP Read context to one memory region: the core
registration code cannot split a coalesced entry between regions.
The force_mr case on other transports remains unresolved.
Update comment for multipage bvecs.
Signed-off-by: J Louis Kaplan <Louis.Kaplan@arm.com>
Assisted-by: LLM
---
drivers/infiniband/core/rw.c | 7 +--
net/sunrpc/xprtrdma/svc_rdma_rw.c | 101 ++++++++++++++++++------------
2 files changed, 65 insertions(+), 43 deletions(-)
diff --git a/drivers/infiniband/core/rw.c b/drivers/infiniband/core/rw.c
index 4fafe393a48c7..f9e0a5b8ea806 100644
--- a/drivers/infiniband/core/rw.c
+++ b/drivers/infiniband/core/rw.c
@@ -706,10 +706,9 @@ int rdma_rw_ctx_init_bvec(struct rdma_rw_ctx *ctx, struct ib_qp *qp,
* is a throughput optimization, not a correctness requirement.
* (iWARP, which does require MRs, is handled by the check above.)
*
- * The rdma_rw_io_needs_mr() gate is not used here because nr_bvec
- * is a raw page count that overstates DMA entry demand -- the bvec
- * caller has no post-DMA-coalescing segment count, and feeding the
- * inflated count into the MR path exhausts the pool on RDMA READs.
+ * We do not use max_sgl_rd to switch non-iWARP READs to MRs. Doing so
+ * would require the MR builder to split multipage bvecs across MRs,
+ * which it does not yet support. The force_mr case is handled above.
*/
return rdma_rw_init_map_wrs_bvec(ctx, qp, bvecs, nr_bvec, &iter,
remote_addr, rkey, dir);
diff --git a/net/sunrpc/xprtrdma/svc_rdma_rw.c b/net/sunrpc/xprtrdma/svc_rdma_rw.c
index a2e1c693b1bc1..074a75d251c96 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_rw.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_rw.c
@@ -789,54 +789,78 @@ static int svc_rdma_build_read_segment(struct svc_rqst *rqstp,
{
struct svcxprt_rdma *rdma = svc_rdma_rqst_rdma(rqstp);
struct svc_rdma_chunk_ctxt *cc = &head->rc_cc;
- unsigned int bvec_idx, nr_bvec, seg_len, len, total;
+ struct ib_device *dev = rdma->sc_cm_id->device;
+ unsigned int max_bvec_len = svc_rdma_max_bvec_len(rdma);
+ u64 remote_offset = segment->rs_offset;
+ unsigned int nr_bvec, len, remaining, total;
+ unsigned int base_pages;
struct svc_rdma_rw_ctxt *ctxt;
int ret;
- len = segment->rs_length;
- if (check_add_overflow(head->rc_pageoff, len, &total))
+ remaining = segment->rs_length;
+ if (check_add_overflow(head->rc_pageoff, remaining, &total))
return -EINVAL;
- nr_bvec = PAGE_ALIGN(total) >> PAGE_SHIFT;
- ctxt = svc_rdma_get_rw_ctxt(rdma, nr_bvec);
- if (!ctxt)
- return -ENOMEM;
- ctxt->rw_nents = nr_bvec;
-
- for (bvec_idx = 0; bvec_idx < ctxt->rw_nents; bvec_idx++) {
- seg_len = min_t(unsigned int, len,
- PAGE_SIZE - head->rc_pageoff);
-
- if (!head->rc_pageoff)
- head->rc_page_count++;
+ base_pages = PAGE_ALIGN(total) >> PAGE_SHIFT;
+ if (!remaining || head->rc_curpage >= rqstp->rq_maxpages ||
+ base_pages > rqstp->rq_maxpages - head->rc_curpage)
+ goto out_overrun;
+
+ while (remaining) {
+ base_pages = DIV_ROUND_UP(head->rc_pageoff + remaining,
+ PAGE_SIZE);
+
+ /*
+ * The bvec MR builder cannot split a coalesced SG entry between
+ * MRs. iWARP always uses MRs for RDMA Reads, so keep each context
+ * within one MR. force_mr on other transports remains subject to
+ * the core limitation.
+ */
+ if (rdma_protocol_iwarp(dev, rdma->sc_port_num)) {
+ unsigned int nr_mrs = rdma_rw_mr_factor(dev,
+ rdma->sc_port_num,
+ base_pages);
- bvec_set_page(&ctxt->rw_bvec[bvec_idx],
- rqstp->rq_pages[head->rc_curpage],
- seg_len, head->rc_pageoff);
+ base_pages = DIV_ROUND_UP(base_pages, nr_mrs);
+ }
- head->rc_pageoff += seg_len;
- if (head->rc_pageoff == PAGE_SIZE) {
- head->rc_curpage++;
- head->rc_pageoff = 0;
+ len = min_t(unsigned int, remaining,
+ (base_pages << PAGE_SHIFT) - head->rc_pageoff);
+ nr_bvec = svc_pages_to_bvecs(NULL,
+ rqstp->rq_pages + head->rc_curpage,
+ base_pages, head->rc_pageoff, len,
+ max_bvec_len);
+ if (!nr_bvec)
+ goto out_overrun;
+ ctxt = svc_rdma_get_rw_ctxt(rdma, nr_bvec);
+ if (!ctxt)
+ return -ENOMEM;
+ ctxt->rw_nents = svc_pages_to_bvecs(ctxt->rw_bvec,
+ rqstp->rq_pages + head->rc_curpage,
+ base_pages, head->rc_pageoff, len,
+ max_bvec_len);
+ if (WARN_ON_ONCE(ctxt->rw_nents != nr_bvec)) {
+ svc_rdma_put_rw_ctxt(rdma, ctxt);
+ goto out_overrun;
}
- len -= seg_len;
+ total = head->rc_pageoff + len;
+ head->rc_curpage += total >> PAGE_SHIFT;
+ head->rc_pageoff = offset_in_page(total);
- if (len && ((head->rc_curpage + 1) > rqstp->rq_maxpages))
- goto out_put;
- }
+ ret = svc_rdma_rw_ctx_init(rdma, ctxt, remote_offset,
+ segment->rs_handle, len,
+ DMA_FROM_DEVICE);
+ if (ret < 0)
+ return -EIO;
- ret = svc_rdma_rw_ctx_init(rdma, ctxt, segment->rs_offset,
- segment->rs_handle, segment->rs_length,
- DMA_FROM_DEVICE);
- if (ret < 0)
- return -EIO;
+ list_add(&ctxt->rw_list, &cc->cc_rwctxts);
+ cc->cc_sqecount += ret;
+ remote_offset += len;
+ remaining -= len;
+ }
percpu_counter_inc(&svcrdma_stat_read);
-
- list_add(&ctxt->rw_list, &cc->cc_rwctxts);
- cc->cc_sqecount += ret;
return 0;
-out_put:
- svc_rdma_put_rw_ctxt(rdma, ctxt);
+out_overrun:
trace_svcrdma_page_overrun_err(&cc->cc_cid, head->rc_curpage);
return -EINVAL;
}
@@ -908,9 +932,6 @@ static int svc_rdma_copy_inline_range(struct svc_rqst *rqstp,
page_len = min_t(unsigned int, remaining,
PAGE_SIZE - head->rc_pageoff);
- if (!head->rc_pageoff)
- head->rc_page_count++;
-
dst = page_address(rqstp->rq_pages[head->rc_curpage]);
memcpy((unsigned char *)dst + head->rc_pageoff, src + offset, page_len);
@@ -1170,6 +1191,8 @@ static void svc_rdma_clear_rqst_pages(struct svc_rqst *rqstp,
{
unsigned int i;
+ /* The cursor identifies every page touched while rebuilding the call. */
+ head->rc_page_count = head->rc_curpage + !!head->rc_pageoff;
for (i = 0; i < head->rc_page_count; i++) {
head->rc_pages[i] = rqstp->rq_pages[i];
rqstp->rq_pages[i] = NULL;
--
2.43.0
next prev parent reply other threads:[~2026-10-06 9:01 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-06 9:00 [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 1/6] sunrpc: Add helpers to build bvecs from contiguous pages J Louis Kaplan
2026-10-06 13:49 ` Chuck Lever
2026-10-08 22:44 ` J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 2/6] nfsd: Coalesce contiguous pages for direct reads J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 3/6] svcrdma: Coalesce contiguous pages when mapping replies J Louis Kaplan
2026-10-06 13:55 ` Chuck Lever
2026-10-08 22:54 ` J Louis Kaplan
2026-10-06 9:00 ` [RFC PATCH 4/6] svcrdma: Coalesce contiguous pages in RDMA Write chunks J Louis Kaplan
2026-10-06 9:00 ` J Louis Kaplan [this message]
2026-10-06 9:00 ` [RFC PATCH 6/6] sunrpc: Allocate svc request pages from large folios J Louis Kaplan
2026-10-06 14:00 ` Chuck Lever
2026-10-08 22:58 ` J Louis Kaplan
2026-10-09 15:28 ` Chuck Lever
2026-10-06 13:47 ` [RFC PATCH 0/6] Improve NFS server direct throughput with passthrough enabled Chuck Lever
2026-10-08 22:50 ` J Louis Kaplan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261006090028.3412544-6-Louis.Kaplan@arm.com \
--to=louis.kaplan@arm.com \
--cc=anna@kernel.org \
--cc=cel@kernel.org \
--cc=dai.ngo@oracle.com \
--cc=jgg@ziepe.ca \
--cc=jlayton@kernel.org \
--cc=leon@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=neil@brown.name \
--cc=okorniev@redhat.com \
--cc=tom@talpey.com \
--cc=trondmy@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox