Linux NFS development
 help / color / mirror / Atom feed
* [PATCH v1 1/5] SUNRPC: fire svc_xprt_free tracepoint before dropping the netns
@ 2026-09-22  1:51 Chuck Lever
  2026-09-22  1:51 ` [PATCH v1 2/5] SUNRPC: fire svc_defer_queue tracepoint before publishing the request Chuck Lever
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Chuck Lever @ 2026-09-22  1:51 UTC (permalink / raw)
  To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey
  Cc: linux-nfs, linux-rdma

svc_xprt_free() releases the transport's network namespace
reference and then fires trace_svc_xprt_free(), which reads
xpt_net->ns.inum through SVC_XPRT_ENDPOINT_ASSIGNMENTS. When the
transport holds the last reference, the namespace is already on
the cleanup list by the time the tracepoint reads it. The window
is narrow, since cleanup_net() runs from a workqueue, but nothing
holds the namespace alive across it.

Fire the tracepoint before any of the teardown steps, while every
field it records is still valid.

Fixes: 11bbb0f76e99 ("SUNRPC: Trace a few more generic svc_xprt events")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
 net/sunrpc/svc_xprt.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 77e28dcc4d2a..908569035e57 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -170,6 +170,8 @@ static void svc_xprt_free(struct kref *kref)
 	struct svc_xprt *xprt =
 		container_of(kref, struct svc_xprt, xpt_ref);
 	struct module *owner = xprt->xpt_class->xcl_owner;
+
+	trace_svc_xprt_free(xprt);
 	if (test_bit(XPT_CACHE_AUTH, &xprt->xpt_flags))
 		svcauth_unix_info_release(xprt);
 	put_cred(xprt->xpt_cred);
@@ -179,7 +181,6 @@ static void svc_xprt_free(struct kref *kref)
 		xprt_put(xprt->xpt_bc_xprt);
 	if (xprt->xpt_bc_xps)
 		xprt_switch_put(xprt->xpt_bc_xps);
-	trace_svc_xprt_free(xprt);
 	xprt->xpt_ops->xpo_free(xprt);
 	module_put(owner);
 }
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v1 2/5] SUNRPC: fire svc_defer_queue tracepoint before publishing the request
  2026-09-22  1:51 [PATCH v1 1/5] SUNRPC: fire svc_xprt_free tracepoint before dropping the netns Chuck Lever
@ 2026-09-22  1:51 ` Chuck Lever
  2026-09-22  1:51 ` [PATCH v1 3/5] svcrdma: fix a page leak in backchannel sends Chuck Lever
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 5+ messages in thread
From: Chuck Lever @ 2026-09-22  1:51 UTC (permalink / raw)
  To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey
  Cc: linux-nfs, linux-rdma

svc_revisit() adds the deferred request to xpt_deferred, drops
xpt_lock, and then fires trace_svc_defer_queue(), which reads the
request's XID and source address. Once the request is on the list,
an nfsd thread already servicing that transport can dequeue it in
svc_deferred_dequeue(), process it, and free it in
svc_xprt_release() before the tracepoint reads it. The window is
narrow, but nothing keeps the request alive across it.

Fire the tracepoint before the list_add(), while svc_revisit()
still owns the request.

Fixes: 8954c5c212d3 ("SUNRPC: Clean up request deferral tracepoints")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
 net/sunrpc/svc_xprt.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 908569035e57..713f4ba19ba1 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -1271,9 +1271,9 @@ static void svc_revisit(struct cache_deferred_req *dreq, int too_many)
 		return;
 	}
 	dr->xprt = NULL;
+	trace_svc_defer_queue(dr);
 	list_add(&dr->handle.recent, &xprt->xpt_deferred);
 	spin_unlock(&xprt->xpt_lock);
-	trace_svc_defer_queue(dr);
 	svc_xprt_enqueue(xprt);
 	svc_xprt_put(xprt);
 }
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v1 3/5] svcrdma: fix a page leak in backchannel sends
  2026-09-22  1:51 [PATCH v1 1/5] SUNRPC: fire svc_xprt_free tracepoint before dropping the netns Chuck Lever
  2026-09-22  1:51 ` [PATCH v1 2/5] SUNRPC: fire svc_defer_queue tracepoint before publishing the request Chuck Lever
@ 2026-09-22  1:51 ` Chuck Lever
  2026-09-22  1:51 ` [PATCH v1 4/5] svcrdma: release a send context stranded by a partial post Chuck Lever
  2026-09-22  1:51 ` [PATCH v1 5/5] svcrdma: release a receive context stranded by a partial Read post Chuck Lever
  3 siblings, 0 replies; 5+ messages in thread
From: Chuck Lever @ 2026-09-22  1:51 UTC (permalink / raw)
  To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey
  Cc: linux-nfs, linux-rdma

svc_rdma_bc_sendto() takes an extra reference on the page that
holds the backchannel Call message, so that the page is not
returned to the allocator while the Send that reads it is still
posted. Send completion once dropped that reference, back when the
page sat in the send context's page array. Commit 99722fe4d5a6
("svcrdma: Persistently allocate and DMA-map Send buffers")
switched the backchannel to svc_rdma_map_reply_msg(), which
DMA-maps the buffer and records no pages. Send completion no
longer touches the page, but the get_page() stayed.

xprt_rdma_bc_free() drops the allocation reference and nothing
drops the extra one, so every backchannel Call sent over RPC/RDMA
leaks its send buffer page.

The extra reference is still needed. A Call at or above
RPCRDMA_PULLUP_THRESH is DMA-mapped in place. The RPC client frees
rq_buffer when the callback task times out or is killed, whether
or not the Send has completed. Record the page in the send
context's page array so that svc_rdma_send_ctxt_release() drops
the reference after Send completion.

Fixes: 99722fe4d5a6 ("svcrdma: Persistently allocate and DMA-map Send buffers")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
 net/sunrpc/xprtrdma/svc_rdma_backchannel.c | 10 +++++++---
 1 file changed, 7 insertions(+), 3 deletions(-)

diff --git a/net/sunrpc/xprtrdma/svc_rdma_backchannel.c b/net/sunrpc/xprtrdma/svc_rdma_backchannel.c
index e5a78b761012..e0768d1ee556 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_backchannel.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_backchannel.c
@@ -85,10 +85,14 @@ static int svc_rdma_bc_sendto(struct svcxprt_rdma *rdma,
 	if (ret < 0)
 		return -EIO;
 
-	/* Bump page refcnt so Send completion doesn't release
-	 * the rq_buffer before all retransmits are complete.
+	/* The RPC client frees rq_buffer when the callback task ends,
+	 * whether or not the Send has completed. Hold the page until the
+	 * send context is released after Send completion.
 	 */
-	get_page(virt_to_page(rqst->rq_buffer));
+	sctxt->sc_pages[0] = virt_to_page(rqst->rq_buffer);
+	get_page(sctxt->sc_pages[0]);
+	sctxt->sc_page_count = 1;
+
 	sctxt->sc_send_wr.opcode = IB_WR_SEND;
 	return svc_rdma_post_send(rdma, sctxt);
 }
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v1 4/5] svcrdma: release a send context stranded by a partial post
  2026-09-22  1:51 [PATCH v1 1/5] SUNRPC: fire svc_xprt_free tracepoint before dropping the netns Chuck Lever
  2026-09-22  1:51 ` [PATCH v1 2/5] SUNRPC: fire svc_defer_queue tracepoint before publishing the request Chuck Lever
  2026-09-22  1:51 ` [PATCH v1 3/5] svcrdma: fix a page leak in backchannel sends Chuck Lever
@ 2026-09-22  1:51 ` Chuck Lever
  2026-09-22  1:51 ` [PATCH v1 5/5] svcrdma: release a receive context stranded by a partial Read post Chuck Lever
  3 siblings, 0 replies; 5+ messages in thread
From: Chuck Lever @ 2026-09-22  1:51 UTC (permalink / raw)
  To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey
  Cc: linux-nfs, linux-rdma

When ib_post_send() rejects a WR mid-chain, svc_rdma_post_send_err()
returns zero to say that a completion will clean up. The Write and
Reply chunk WRs in a reply's chain do carry completions, but those
only trace or close the transport. The Send at the chain's tail is
the one whose completion puts the send context, so a chain cut
before the Send leaves the context unreleased. svc_rdma_sendto()
reports success and never puts it either. No list the transport
destructor walks holds the context, so it leaks past teardown along
with the DMA mappings, page references, and rw contexts it owns.

The error path cannot release the context itself. The posted WRs
still reference its rw contexts and pages, and the HCA may still be
reading them. Nor can it drain the Send Queue to wait for them. A
drain needs a free SQ slot for its marker WR. A provider that
rejects a WR after accepting earlier ones has most likely diverged
from svcrdma's SQ accounting, so no slot svcrdma believes it holds
can be trusted.

Transfer the context to the transport instead. On a partial post,
svc_rdma_post_send() parks the context on a per-transport list and
reports success, the same ownership a fully posted chain has.
svc_rdma_free() releases the list after its own drain and before the
rw contexts are destroyed. The SQ reservation stays debited, since
the Send that refunds it never completes and the transport is
closing.

Reaching this path requires a provider to reject a WR after
accepting earlier ones in the same chain. That has not been
observed. The Fixes: tag names the commit that wrote the mistaken
contract. The chain became long enough to cut only with commit
10e6fc1054d9 ("svcrdma: Post the Reply chunk and Send WR together").

Fixes: 71b43531ee0b ("svcrdma: Post Send WR chain")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
 include/linux/sunrpc/svc_rdma.h          |  2 +
 net/sunrpc/xprtrdma/svc_rdma_rw.c        |  1 +
 net/sunrpc/xprtrdma/svc_rdma_sendto.c    | 47 +++++++++++++++++++-----
 net/sunrpc/xprtrdma/svc_rdma_transport.c |  2 +
 4 files changed, 43 insertions(+), 9 deletions(-)

diff --git a/include/linux/sunrpc/svc_rdma.h b/include/linux/sunrpc/svc_rdma.h
index 76aa5ec4ab40..b52f83f97d64 100644
--- a/include/linux/sunrpc/svc_rdma.h
+++ b/include/linux/sunrpc/svc_rdma.h
@@ -117,6 +117,7 @@ struct svcxprt_rdma {
 	struct llist_head    sc_recv_ctxts;
 
 	struct llist_head    sc_send_release_list;
+	struct llist_head    sc_send_stranded_ctxts;
 
 	atomic_t	     sc_completion_ids;
 };
@@ -300,6 +301,7 @@ extern int svc_rdma_process_read_list(struct svcxprt_rdma *rdma,
 /* svc_rdma_sendto.c */
 extern void svc_rdma_send_ctxts_destroy(struct svcxprt_rdma *rdma);
 extern void svc_rdma_send_ctxts_drain(struct svcxprt_rdma *rdma);
+extern void svc_rdma_send_ctxts_stranded_release(struct svcxprt_rdma *rdma);
 extern struct svc_rdma_send_ctxt *
 		svc_rdma_send_ctxt_get(struct svcxprt_rdma *rdma);
 extern void svc_rdma_send_ctxt_put(struct svcxprt_rdma *rdma,
diff --git a/net/sunrpc/xprtrdma/svc_rdma_rw.c b/net/sunrpc/xprtrdma/svc_rdma_rw.c
index 9aaaade99e6e..7b7879b2cc88 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_rw.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_rw.c
@@ -420,6 +420,7 @@ static int svc_rdma_post_chunk_ctxt(struct svcxprt_rdma *rdma,
 		return ret;
 
 	cc->cc_posttime = ktime_get();
+	bad_wr = first_wr;
 	ret = ib_post_send(rdma->sc_qp, first_wr, &bad_wr);
 	if (ret)
 		return svc_rdma_post_send_err(rdma, &cc->cc_cid, bad_wr,
diff --git a/net/sunrpc/xprtrdma/svc_rdma_sendto.c b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
index c09659b17351..05dcdb6c007a 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_sendto.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
@@ -292,6 +292,23 @@ void svc_rdma_send_ctxts_drain(struct svcxprt_rdma *rdma)
 		svc_rdma_send_ctxt_release(rdma, ctxt);
 }
 
+/**
+ * svc_rdma_send_ctxts_stranded_release - Release stranded send_ctxts
+ * @rdma: svcxprt_rdma being torn down
+ *
+ * Context: transport destructor only, after the QP has been drained
+ * and before its rw contexts are destroyed.
+ */
+void svc_rdma_send_ctxts_stranded_release(struct svcxprt_rdma *rdma)
+{
+	struct svc_rdma_send_ctxt *ctxt, *next;
+	struct llist_node *node;
+
+	node = llist_del_all(&rdma->sc_send_stranded_ctxts);
+	llist_for_each_entry_safe(ctxt, next, node, sc_node)
+		svc_rdma_send_ctxt_release(rdma, ctxt);
+}
+
 /**
  * svc_rdma_send_ctxt_put - Queue send_ctxt for deferred release
  * @rdma: controlling svcxprt_rdma
@@ -427,9 +444,15 @@ int svc_rdma_sq_wait(struct svcxprt_rdma *rdma,
  * @sqecount: number of SQ entries that were reserved
  * @ret: error code from ib_post_send
  *
+ * The transport is closing on return. Who owns the caller's context
+ * depends on how much of the chain was posted.
+ *
  * Return values:
- *   %0: At least one WR was posted; a completion handles cleanup
- *   %-ENOTCONN: No WRs were posted; SQ slots are released
+ *   %0: A prefix of the chain was posted. Its signaled tail was not,
+ *       so no completion will release the caller's context. The
+ *       caller keeps ownership. The SQ reservation stays debited.
+ *   %-ENOTCONN: No WR was posted; SQ slots are released and the
+ *       caller owns its context.
  */
 int svc_rdma_post_send_err(struct svcxprt_rdma *rdma,
 			   const struct rpc_rdma_cid *cid,
@@ -440,9 +463,6 @@ int svc_rdma_post_send_err(struct svcxprt_rdma *rdma,
 	trace_svcrdma_sq_post_err(rdma, cid, ret);
 	svc_rdma_xprt_deferred_close(rdma);
 
-	/* If even one WR was posted, a Send completion will
-	 * return the reserved SQ slots.
-	 */
 	if (bad_wr != first_wr)
 		return 0;
 
@@ -493,7 +513,8 @@ static void svc_rdma_wc_send(struct ib_cq *cq, struct ib_wc *wc)
  * In some error flow cases, svc_rdma_wc_send() releases @ctxt.
  *
  * Return values:
- *   %0: @ctxt's WR chain was posted successfully
+ *   %0: The transport owns @ctxt; a Send completion or the
+ *       transport destructor releases it
  *   %-ENOTCONN: The connection was lost
  */
 int svc_rdma_post_send(struct svcxprt_rdma *rdma,
@@ -519,9 +540,17 @@ int svc_rdma_post_send(struct svcxprt_rdma *rdma,
 
 	trace_svcrdma_post_send(ctxt);
 	ret = ib_post_send(rdma->sc_qp, first_wr, &bad_wr);
-	if (ret)
-		return svc_rdma_post_send_err(rdma, &cid, bad_wr,
-					      first_wr, sqecount, ret);
+	if (ret) {
+		ret = svc_rdma_post_send_err(rdma, &cid, bad_wr,
+					     first_wr, sqecount, ret);
+		if (ret)
+			return ret;
+
+		/* The posted prefix still references @ctxt. Only the
+		 * transport destructor may release it.
+		 */
+		llist_add(&ctxt->sc_node, &rdma->sc_send_stranded_ctxts);
+	}
 	return 0;
 }
 
diff --git a/net/sunrpc/xprtrdma/svc_rdma_transport.c b/net/sunrpc/xprtrdma/svc_rdma_transport.c
index f949601b2144..d449460e8f1e 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_transport.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_transport.c
@@ -197,6 +197,7 @@ static struct svcxprt_rdma *svc_rdma_create_xprt(struct svc_serv *serv,
 	init_llist_head(&cma_xprt->sc_recv_ctxts);
 	init_llist_head(&cma_xprt->sc_rw_ctxts);
 	init_llist_head(&cma_xprt->sc_send_release_list);
+	init_llist_head(&cma_xprt->sc_send_stranded_ctxts);
 	init_waitqueue_head(&cma_xprt->sc_send_wait);
 	init_waitqueue_head(&cma_xprt->sc_sq_ticket_wait);
 
@@ -673,6 +674,7 @@ static void svc_rdma_free(struct svc_xprt *xprt)
 	if (rdma->sc_qp && !IS_ERR(rdma->sc_qp))
 		ib_drain_qp(rdma->sc_qp);
 	svc_rdma_send_ctxts_drain(rdma);
+	svc_rdma_send_ctxts_stranded_release(rdma);
 
 	svc_rdma_flush_recv_queues(rdma);
 
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v1 5/5] svcrdma: release a receive context stranded by a partial Read post
  2026-09-22  1:51 [PATCH v1 1/5] SUNRPC: fire svc_xprt_free tracepoint before dropping the netns Chuck Lever
                   ` (2 preceding siblings ...)
  2026-09-22  1:51 ` [PATCH v1 4/5] svcrdma: release a send context stranded by a partial post Chuck Lever
@ 2026-09-22  1:51 ` Chuck Lever
  3 siblings, 0 replies; 5+ messages in thread
From: Chuck Lever @ 2026-09-22  1:51 UTC (permalink / raw)
  To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey
  Cc: linux-nfs, linux-rdma

When ib_post_send() rejects a WR mid-chain, svc_rdma_post_send_err()
returns zero and svc_rdma_post_chunk_ctxt() reports the Read chain
as posted. The only signaled WR in the chain is its tail, so a chain
cut before the tail posts nothing that completes.
svc_rdma_process_read_list() then reports the Read in progress, and
the receive context waits for svc_rdma_wc_read_done(), which never
runs. No list the transport destructor walks holds the context, so
it leaks past teardown along with its rw contexts and their DMA
mappings.

The caller cannot release the context on the error, because the
posted Read WRs still write into pages it owns. Nor can the error
path drain the Send Queue. The drain's marker WR needs an SQ slot,
and a provider that has just rejected a WR cannot be trusted to
have one.

Park the receive context on a per-transport list, as the Send path
does, and release it from svc_rdma_free() after the drain and before
the rw contexts are destroyed. svc_rdma_post_chunk_ctxt() posts only
Read chains, so pass it the receive context itself rather than the
embedded chunk context.

Reaching this path needs a provider to reject a WR after accepting
earlier ones in the same chain. That has not been observed.

Fixes: f13193f50b64 ("svcrdma: Introduce local rdma_rw API helpers")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
 include/linux/sunrpc/svc_rdma.h          |  2 ++
 net/sunrpc/xprtrdma/svc_rdma_recvfrom.c  | 18 ++++++++++++++
 net/sunrpc/xprtrdma/svc_rdma_rw.c        | 31 ++++++++++++++++++------
 net/sunrpc/xprtrdma/svc_rdma_transport.c |  2 ++
 4 files changed, 46 insertions(+), 7 deletions(-)

diff --git a/include/linux/sunrpc/svc_rdma.h b/include/linux/sunrpc/svc_rdma.h
index b52f83f97d64..87af38878b37 100644
--- a/include/linux/sunrpc/svc_rdma.h
+++ b/include/linux/sunrpc/svc_rdma.h
@@ -118,6 +118,7 @@ struct svcxprt_rdma {
 
 	struct llist_head    sc_send_release_list;
 	struct llist_head    sc_send_stranded_ctxts;
+	struct llist_head    sc_recv_stranded_ctxts;
 
 	atomic_t	     sc_completion_ids;
 };
@@ -263,6 +264,7 @@ extern void svc_rdma_handle_bc_reply(struct svc_rqst *rqstp,
 
 /* svc_rdma_recvfrom.c */
 extern void svc_rdma_recv_ctxts_destroy(struct svcxprt_rdma *rdma);
+extern void svc_rdma_recv_ctxts_stranded_release(struct svcxprt_rdma *rdma);
 extern bool svc_rdma_post_recvs(struct svcxprt_rdma *rdma);
 extern struct svc_rdma_recv_ctxt *
 		svc_rdma_recv_ctxt_get(struct svcxprt_rdma *rdma);
diff --git a/net/sunrpc/xprtrdma/svc_rdma_recvfrom.c b/net/sunrpc/xprtrdma/svc_rdma_recvfrom.c
index d029bcb7a5c0..ed5a83e9edd2 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_recvfrom.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_recvfrom.c
@@ -234,6 +234,24 @@ void svc_rdma_recv_ctxt_put(struct svcxprt_rdma *rdma,
 	llist_add(&ctxt->rc_node, &rdma->sc_recv_ctxts);
 }
 
+/**
+ * svc_rdma_recv_ctxts_stranded_release - Release stranded recv_ctxts
+ * @rdma: svcxprt_rdma being torn down
+ *
+ * Context: transport destructor only, after the QP has been drained
+ * and before its rw contexts and recv_ctxts are destroyed.
+ */
+void svc_rdma_recv_ctxts_stranded_release(struct svcxprt_rdma *rdma)
+{
+	struct svc_rdma_recv_ctxt *ctxt;
+	struct llist_node *node;
+
+	while ((node = llist_del_first(&rdma->sc_recv_stranded_ctxts)) != NULL) {
+		ctxt = llist_entry(node, struct svc_rdma_recv_ctxt, rc_node);
+		svc_rdma_recv_ctxt_put(rdma, ctxt);
+	}
+}
+
 /**
  * svc_rdma_release_ctxt - Release transport-specific per-rqst resources
  * @xprt: the transport which owned the context
diff --git a/net/sunrpc/xprtrdma/svc_rdma_rw.c b/net/sunrpc/xprtrdma/svc_rdma_rw.c
index 7b7879b2cc88..63d00f0cc1db 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_rw.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_rw.c
@@ -389,10 +389,18 @@ static void svc_rdma_wc_read_done(struct ib_cq *cq, struct ib_wc *wc)
  * - If ib_post_send() succeeds, only one completion is expected,
  *   even if one or more WRs are flushed. This is true when posting
  *   an rdma_rw_ctx or when posting a single signaled WR.
+ *
+ * Return values:
+ *   %0: The chain was posted; svc_rdma_wc_read_done() releases
+ *       @head. Or a prefix of the chain was posted; no completion
+ *       follows, and the transport destructor releases @head.
+ *   %-ENOTCONN: No WR was posted; the caller still owns @head.
+ *   %-EINVAL: @head's Read chain needs more SQ entries than exist.
  */
 static int svc_rdma_post_chunk_ctxt(struct svcxprt_rdma *rdma,
-				    struct svc_rdma_chunk_ctxt *cc)
+				    struct svc_rdma_recv_ctxt *head)
 {
+	struct svc_rdma_chunk_ctxt *cc = &head->rc_cc;
 	struct ib_send_wr *first_wr;
 	const struct ib_send_wr *bad_wr;
 	struct list_head *tmp;
@@ -422,10 +430,18 @@ static int svc_rdma_post_chunk_ctxt(struct svcxprt_rdma *rdma,
 	cc->cc_posttime = ktime_get();
 	bad_wr = first_wr;
 	ret = ib_post_send(rdma->sc_qp, first_wr, &bad_wr);
-	if (ret)
-		return svc_rdma_post_send_err(rdma, &cc->cc_cid, bad_wr,
-					      first_wr, cc->cc_sqecount,
-					      ret);
+	if (ret) {
+		ret = svc_rdma_post_send_err(rdma, &cc->cc_cid, bad_wr,
+					     first_wr, cc->cc_sqecount,
+					     ret);
+		if (ret)
+			return ret;
+
+		/* The posted prefix still references @head. Only the
+		 * transport destructor may release it.
+		 */
+		llist_add(&head->rc_node, &rdma->sc_recv_stranded_ctxts);
+	}
 	return 0;
 }
 
@@ -1169,7 +1185,8 @@ static void svc_rdma_clear_rqst_pages(struct svc_rqst *rqstp,
  * RDMA Reads have completed.
  *
  * Return values:
- *   %1: all needed RDMA Reads were posted successfully,
+ *   %1: RDMA Reads were posted; the transport now owns @head, and
+ *       a Read completion or the transport destructor releases it,
  *   %-EINVAL: client provided too many chunks or segments,
  *   %-ENOMEM: rdma_rw context pool was exhausted,
  *   %-ENOTCONN: posting failed (connection is lost),
@@ -1200,6 +1217,6 @@ int svc_rdma_process_read_list(struct svcxprt_rdma *rdma,
 		return ret;
 
 	trace_svcrdma_post_read_chunk(&cc->cc_cid, cc->cc_sqecount);
-	ret = svc_rdma_post_chunk_ctxt(rdma, cc);
+	ret = svc_rdma_post_chunk_ctxt(rdma, head);
 	return ret < 0 ? ret : 1;
 }
diff --git a/net/sunrpc/xprtrdma/svc_rdma_transport.c b/net/sunrpc/xprtrdma/svc_rdma_transport.c
index d449460e8f1e..df6c3501f97c 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_transport.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_transport.c
@@ -198,6 +198,7 @@ static struct svcxprt_rdma *svc_rdma_create_xprt(struct svc_serv *serv,
 	init_llist_head(&cma_xprt->sc_rw_ctxts);
 	init_llist_head(&cma_xprt->sc_send_release_list);
 	init_llist_head(&cma_xprt->sc_send_stranded_ctxts);
+	init_llist_head(&cma_xprt->sc_recv_stranded_ctxts);
 	init_waitqueue_head(&cma_xprt->sc_send_wait);
 	init_waitqueue_head(&cma_xprt->sc_sq_ticket_wait);
 
@@ -677,6 +678,7 @@ static void svc_rdma_free(struct svc_xprt *xprt)
 	svc_rdma_send_ctxts_stranded_release(rdma);
 
 	svc_rdma_flush_recv_queues(rdma);
+	svc_rdma_recv_ctxts_stranded_release(rdma);
 
 	svc_rdma_destroy_rw_ctxts(rdma);
 	svc_rdma_send_ctxts_destroy(rdma);
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-22  1:51 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-22  1:51 [PATCH v1 1/5] SUNRPC: fire svc_xprt_free tracepoint before dropping the netns Chuck Lever
2026-09-22  1:51 ` [PATCH v1 2/5] SUNRPC: fire svc_defer_queue tracepoint before publishing the request Chuck Lever
2026-09-22  1:51 ` [PATCH v1 3/5] svcrdma: fix a page leak in backchannel sends Chuck Lever
2026-09-22  1:51 ` [PATCH v1 4/5] svcrdma: release a send context stranded by a partial post Chuck Lever
2026-09-22  1:51 ` [PATCH v1 5/5] svcrdma: release a receive context stranded by a partial Read post Chuck Lever

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox