* [PATCH net-next v6 03/12] net/rds: guard every work-requeueing site with rds_destroy_pending()
2026-09-22 8:53 Allison Henderson
@ 2026-09-22 8:54 ` Allison Henderson
0 siblings, 0 replies; 7+ messages in thread
From: Allison Henderson @ 2026-09-22 8:54 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms; +Cc: achender
rds_conn_destroy() cancels the path works and then destroys the
per-path workqueue. The sites that can re-arm those works are
supposed to test rds_destroy_pending() under rcu_read_lock() first,
paired with the synchronize_rcu() in the destroy path, so that no new
work can be queued once the cancellation has begun.
Five arming sites never got that guard:
- rds_ib_send_cqe_handler() and rds_ib_send_add_credits() re-arm
cp_send_w when a send completion or a credit update clears
RDS_LL_SEND_FULL,
- rds_ib_recv_refill() re-arms cp_recv_w when the recv ring runs
low,
- rds_tcp_accept_one() kicks cp_recv_w on the freshly accepted
socket, and
- rds_sendmsg() arms cp_conn_w for a multipath connection whose
path 0 is not up yet.
The IB completion sites are reachable from soft-irq at any point
before the QP is drained, so a completion landing in the window
between the cancel and destroy_workqueue() in rds_conn_path_destroy()
re-arms a work on a workqueue that is about to be destroyed: with
delay 0 the work is queued directly on the freed workqueue, and with
delay 1 the timer survives destroy_workqueue() unseen and fires
afterwards, queueing from a timer_list that lives in the freed c_path
array.
Wrap all five sites in the same rcu_read_lock() +
rds_destroy_pending() pattern the other arming sites already use.
The four self-requeues in rds_send_worker() and rds_recv_worker() are
left alone on purpose: they run from inside the work item itself, and
cancel_delayed_work_sync() disables the work for the duration of the
cancel, so a requeue issued by the still-running callback is dropped
and none can follow once the cancel has returned.
With the predicate as it stands the guards cover the netns teardown and
module unload cases; the following patch extends it to the destroy of a
single connection.
Fixes: ebeeb1ad9b8a ("rds: tcp: use rds_destroy_pending() to synchronize netns/module teardown and rds connection/workq management")
Assisted-by: Claude-Code:claude-fable-5
Signed-off-by: Allison Henderson <achender@kernel.org>
---
net/rds/ib_recv.c | 6 +++++-
net/rds/ib_send.c | 18 ++++++++++++++----
net/rds/send.c | 11 ++++++++---
net/rds/tcp_listen.c | 10 +++++++---
4 files changed, 34 insertions(+), 11 deletions(-)
diff --git a/net/rds/ib_recv.c b/net/rds/ib_recv.c
index bd6cb3ffaa57..7d45808544a0 100644
--- a/net/rds/ib_recv.c
+++ b/net/rds/ib_recv.c
@@ -458,7 +458,11 @@ void rds_ib_recv_refill(struct rds_connection *conn, int prefill, gfp_t gfp)
(must_wake ||
(can_wait && rds_ib_ring_low(&ic->i_recv_ring)) ||
rds_ib_ring_empty(&ic->i_recv_ring))) {
- queue_delayed_work(conn->c_path->cp_wq, &conn->c_recv_w, 1);
+ rcu_read_lock();
+ if (!rds_destroy_pending(conn))
+ queue_delayed_work(conn->c_path->cp_wq,
+ &conn->c_recv_w, 1);
+ rcu_read_unlock();
}
if (can_wait)
cond_resched();
diff --git a/net/rds/ib_send.c b/net/rds/ib_send.c
index d6be95542119..bc411e96ad12 100644
--- a/net/rds/ib_send.c
+++ b/net/rds/ib_send.c
@@ -298,8 +298,13 @@ void rds_ib_send_cqe_handler(struct rds_ib_connection *ic, struct ib_wc *wc)
rds_ib_sub_signaled(ic, nr_sig);
if (test_and_clear_bit(RDS_LL_SEND_FULL, &conn->c_flags) ||
- test_bit(0, &conn->c_map_queued))
- queue_delayed_work(conn->c_path->cp_wq, &conn->c_send_w, 0);
+ test_bit(0, &conn->c_map_queued)) {
+ rcu_read_lock();
+ if (!rds_destroy_pending(conn))
+ queue_delayed_work(conn->c_path->cp_wq,
+ &conn->c_send_w, 0);
+ rcu_read_unlock();
+ }
/* We expect errors as the qp is drained during shutdown */
if (wc->status != IB_WC_SUCCESS && rds_conn_up(conn)) {
@@ -420,8 +425,13 @@ void rds_ib_send_add_credits(struct rds_connection *conn, unsigned int credits)
test_bit(RDS_LL_SEND_FULL, &conn->c_flags) ? ", ll_send_full" : "");
atomic_add(IB_SET_SEND_CREDITS(credits), &ic->i_credits);
- if (test_and_clear_bit(RDS_LL_SEND_FULL, &conn->c_flags))
- queue_delayed_work(conn->c_path->cp_wq, &conn->c_send_w, 0);
+ if (test_and_clear_bit(RDS_LL_SEND_FULL, &conn->c_flags)) {
+ rcu_read_lock();
+ if (!rds_destroy_pending(conn))
+ queue_delayed_work(conn->c_path->cp_wq,
+ &conn->c_send_w, 0);
+ rcu_read_unlock();
+ }
WARN_ON(IB_GET_SEND_CREDITS(credits) >= 16384);
diff --git a/net/rds/send.c b/net/rds/send.c
index 1afa981e5c06..32c411d10e3e 100644
--- a/net/rds/send.c
+++ b/net/rds/send.c
@@ -1378,9 +1378,14 @@ int rds_sendmsg(struct socket *sock, struct msghdr *msg, size_t payload_len)
* outstanding.
*/
if (!test_and_set_bit(RDS_RECONNECT_PENDING,
- &conn->c_path[0].cp_flags))
- queue_delayed_work(conn->c_path[0].cp_wq,
- &conn->c_path[0].cp_conn_w, 0);
+ &conn->c_path[0].cp_flags)) {
+ rcu_read_lock();
+ if (!rds_destroy_pending(conn))
+ queue_delayed_work(conn->c_path[0].cp_wq,
+ &conn->c_path[0].cp_conn_w,
+ 0);
+ rcu_read_unlock();
+ }
rds_send_ping(conn, 0);
}
diff --git a/net/rds/tcp_listen.c b/net/rds/tcp_listen.c
index 13fa60c1985b..8a0c54aced5e 100644
--- a/net/rds/tcp_listen.c
+++ b/net/rds/tcp_listen.c
@@ -316,10 +316,14 @@ int rds_tcp_accept_one(struct rds_tcp_net *rtn)
*/
if (READ_ONCE(sk->sk_state) == TCP_CLOSE_WAIT ||
READ_ONCE(sk->sk_state) == TCP_LAST_ACK ||
- READ_ONCE(sk->sk_state) == TCP_CLOSE)
+ READ_ONCE(sk->sk_state) == TCP_CLOSE) {
rds_conn_path_drop(cp, 0);
- else
- queue_delayed_work(cp->cp_wq, &cp->cp_recv_w, 0);
+ } else {
+ rcu_read_lock();
+ if (!rds_destroy_pending(cp->cp_conn))
+ queue_delayed_work(cp->cp_wq, &cp->cp_recv_w, 0);
+ rcu_read_unlock();
+ }
sock_put(sk);
--
2.25.1
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted
@ 2026-09-22 16:43 Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 01/12] net/rds: ib: don't enable interrupts in rds_ib_conn_free() Allison Henderson
` (4 more replies)
0 siblings, 5 replies; 7+ messages in thread
From: Allison Henderson @ 2026-09-22 16:43 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms; +Cc: achender
Hi all,
This is v6 of the connection-lifetime set (v1 at [1], v2 at [2],
v3 at [3], v4 at [6], v5 at [7]), following
"net/rds: own the fastpath locks across connection teardown", which is
in net-next. It is targeted at net-next: although the series fixes
real use-after-frees (one syzbot report and one report from Chengfeng
Ye below), it does so by reworking connection lifetime rather than
patching the individual crash sites, and that rework is too invasive
for net.
rds_conn_destroy() frees the connection, its paths and its workqueues
on the spot, relying on the documented assumption that "no one else is
referencing the connection". That assumption stopped being true a
long time ago: connections are destroyed on netns teardown and on a
protocol-version mismatch as well as on rmmod, while pointers to them
still live in socket rs_conn caches, congestion-map conn lists, CM
callbacks and workers, and - for as long as an application leaves data
unread - in every rds_incoming sitting on a receive queue. No single
Fixes: commit covers the rot, so the series carries Reported-by tags
where there are concrete reports instead.
Patch 1 fixes rds_ib_conn_free() re-enabling interrupts under
a caller's irqsave lock.
Patch 2 makes the passive-connection exits of __rds_conn_create()
undo conn_alloc() the same way the lost-race exit does, through one
helper (a refactor: the passive twin only exists for IB, so nothing
leaked there). Both are first in the series because the
reference-counted teardown calls conn_free() from more places.
Patch 3 guards the five work-arming sites that never tested
rds_destroy_pending(): two IB completion paths, the IB recv refill,
the TCP accept kick and the multipath reconnect in sendmsg.
Patch 4 gives the connection its own destroy-in-progress marker.
With f97d8c7bab78 in net-next there is no single-connection destroy
left for the old predicate to miss, but the later patches need the
precise answer: the IB unload re-sweep hands a connection to
rds_conn_destroy() more than once, and the passive-twin creation
must refuse a parent whose destroy has begun.
Based on UEK commits:
e2f5005adf63 net/rds: Add krefs to struct rds_connection
https://github.com/oracle/linux-uek/commit/e2f5005adf63
6c53ef92f46e net/rds: Merge uses of conn->c_destroy_in_prog & RDS_DESTROY_PENDING
https://github.com/oracle/linux-uek/commit/6c53ef92f46e
Patch 5 splits rds_conn_destroy() into a synchronous quiesce and a
kref-governed free.
Based on UEK commits:
2c8569e4c880 ("net/rds: Add krefs to struct rds_connection").
https://github.com/oracle/linux-uek/commit/2c8569e4c880
Patch 6 makes each transport's exit path wait for its connections to
actually be freed before the module goes away. The wait is
unbounded and warns every ten seconds: a connection reference can be
held for an application-controlled time (unread data), so a timeout
would only move the use-after-free from freed memory to unloaded
module text. rmmod blocking while data is queued and unread is the
historical RDS contract. For IB the wait re-sweeps the nodev list,
since device connections migrate to it asynchronously.
Based on UEK commits:
ece4b4e39afa ("net/rds: wait_event_timeout until zero connections during rmmod")
https://github.com/oracle/linux-uek/commit/ece4b4e39afa
905ec90e6166 ("net/rds: Each RDS transport should keep its own connection count")
https://github.com/oracle/linux-uek/commit/905ec90e6166
Patch 7 unlinks each transport node before its destroy, so a
free deferred past the teardown loop cannot write into the loop's
stack-local list head. IB marks a gathered node as claimed by the
sweep, since its connect and shutdown paths move the node too.
Patch 8 hands out real references everywhere a connection pointer
previously escaped bare, RCU-annotates parent->c_passive, refuses to
revive a passive connection whose destroy has begun, and serializes
the SIOCRDSSETTOS check with the rs_conn cache.
Based on UEK commits:
2c8569e4c880 ("net/rds: Add krefs to struct rds_connection")
https://github.com/oracle/linux-uek/commit/2c8569e4c880
0e9e3a72b7f7 ("net/rds: rds_sendmsg must use rs_conn only when not being destroyed").
https://github.com/oracle/linux-uek/commit/0e9e3a72b7f7
Patch 9 has the quiesce purge cp_send_queue under cp_lock, as every
adder to that queue holds it. No sender can be in flight at any of
today's destroy triggers (netns teardown, module unload), so this is
hygiene rather than a race fix, and the v5 refusals in the senders
that guarded against that impossible state are gone.
Patch 10 pins the connection across the RDMA-CM event handler
and rejects a connect request for a connection whose destroy has
already quiesced it.
Patch 11 drops the now-unused global rds_conn_count.
Based on UEK commits:
905ec90e6166 ("net/rds: Each RDS transport should keep its own connection count").
https://github.com/oracle/linux-uek/commit/905ec90e6166
Patch 12 makes struct rds_incoming hold a reference on i_conn, the
fix for the KASAN use-after-free Chengfeng Ye reported [4].
Based on UEK commits:
99b9a3715419 ("net/rds: fix crash by expanding kref coverage to rds_incoming.i_conn").
https://github.com/oracle/linux-uek/commit/99b9a3715419
Patches 5, 8, 6 and 12 are ports of the connection kref work Sharath
Srinivasan did for Oracle UEK, adapted to the upstream code.
The series has been validated with the RDS selftests over loopback-TCP
and RXE-RDMA, plus churn tests that delete network namespaces and
unload the modules under live rds-stress traffic.
Changes since v5 [7]:
- Rebased; the version-mismatch destroy is now an rds_conn_drop() in
net-next (f97d8c7bab78), so patch 4's Fixes: tag is gone and its
changelog describes what still needs the per-connection flag, and
the CM-callback deadlock reports against earlier versions no longer
apply.
- Patch 2: described as the refactor it is (the passive twin is
IB-only, single path); Fixes: tag dropped.
- Patch 9: reduced to the cp_lock purge; the send-side refusals and
the negative *queued signalling are dropped, since no sender can be
in flight at a destroy trigger.
- Patch 6: wait comment describes the initial-sweep-plus-resweep
contract; changelog notes the guarantee is complete only once incs
hold references.
- Patch 8: c_destroy_in_prog read through rds_destroy_pending() in
__rds_conn_create(); the c_passive puts documented as possibly the
last; READ_ONCE() on the unlocked rs_tos sample; rs_conn comment
notes the rds_release() exception.
- Patch 10: stale forward-reference comment fixed.
Changes since v4 [6]:
- Patch 7: the IB sweep marks each gathered node with a new
i_ib_node_detached flag instead of relying on list emptiness. A
node parked on the sweep's stack list is not empty, so a concurrent
rds_ib_add_conn() would have unlinked it from under the sweep's
lockless walk; rds_ib_add_conn(), rds_ib_remove_conn() and
rds_ib_conn_free() now leave a claimed node alone.
Changes since v3 [3]:
- Rebased; the version-mismatch destroy that patch 10 also had to
cope with is now an rds_conn_drop() in net (f97d8c7bab78), and
patch 10's changelog says what remains for it to cover.
- The uninitialised transport pointer in the CM event handler that
the second v2 review pass raised turned out to be the bug Aohan
Mei already posted a fix for (v2 at [5], stalled after review); it
is carried forward separately for net rather than added here.
- Patch 7: the teardown walks take a reference on each gathered
connection, so a connection destroyed earlier and freed by a
pending holder cannot vanish under the iterator, and the two IB
list movers tolerate a node the sweep already unlinked instead of
BUG_ON()ing; the nodev sweep gathers entry by entry, so the
resweep from the unload wait no longer relies on list_splice_init().
- Patch 8: SIOCRDSSETTOS check-then-act closed - the install in
rds_sendmsg() re-checks the socket's ToS under rs_lock; rs_conn
documented as a referenced, rs_lock-serialised cache; contract
comment restored to rds_conn_lookup().
- Patch 9: rds_send_probe() gets the same rds_destroy_pending()
test under cp_lock as rds_send_queue_rm() (a probe queued after the
purge pinned the connection for good).
- v3 patch 10 (rds_tcp_accept_one() destroy check) dropped: the check
was not serialised against the destroy, and the race it aimed at is
not reachable - every TCP destroy path stops the listener, flushing
the accept work, first. A comment now records that ordering.
- Changelog and comment corrections from the second v2 review pass
(self-requeue exemption in patch 3, c_refcount comment in patch 5,
the uninterruptible wait and the transport-text wake in patch 6,
"lock-free" wording in patch 11, cross-netns comment in patch 12).
Changes since v2 [2]:
- New patches 1 and 2: pre-existing rds_ib_conn_free() interrupt
state clobber and passive-path transport data leak, surfaced by
review of the teardown changes.
- Patch 6: rds_ib_destroy_nodev_conns() uses list_splice_init(), so
the resweep from the unload wait cannot splice a stale list head.
- Patch 8: SIOCRDSGETTOS reads rs_tos under rs_lock like SETTOS.
- New patch 9: rds_send_queue_rm() refuses a connection whose destroy
has begun, under cp_lock, and the quiesce purges cp_send_queue
under cp_lock (list corruption against an in-flight sender, and a
message that would pin the connection forever).
- New patch 10: rds_tcp_accept_one() does not install a socket on a
connection whose destroy has begun (socket left pointing at a freed
path).
Changes since v1 [1]:
- New patch 1: guard the five work-arming sites that never tested
rds_destroy_pending(); patch 2's changelog and comments narrowed
to what it actually newly covers.
- Patch 4 moved ahead of the reference holders, so no bisect point
has references without the unload wait; wait made unbounded with a
periodic warning instead of a 10 s timeout; IB exit re-sweeps the
nodev list for connections that detach from their device late.
- New patch 5: transport nodes unlinked before destroy (stack list
head use-after-free from a deferred conn_free).
- Patch 6: c_passive RCU-annotated; a destroyed passive child is
refused by __rds_conn_create() and clears the parent's pointer
itself; SIOCRDSSETTOS uses rs_lock; lookup comment reworded;
explicit not-for-stable note.
- New patch 7: reference across the CM event handler, and a
destroy-pending re-check in rds_ib_cm_handle_connect().
- Patch 9: changelog states what a lingering inc keeps alive and
that it blocks module unload.
- Changelog corrections throughout (netns teardown paths named as the
non-rmmod destroyers, stale rds_conn_path_destroy() reference).
[1] https://lore.kernel.org/netdev/20260904070248.160384-1-achender@kernel.org/
[2] https://lore.kernel.org/netdev/20260912035027.27447-1-achender@kernel.org/
[3] https://lore.kernel.org/netdev/20260914033719.138057-1-achender@kernel.org/
[4] https://lore.kernel.org/netdev/20260720184955.3008978-1-nicoyip.dev@gmail.com/
[5] https://lore.kernel.org/netdev/20260825021223.3483044-1-ljp1205831794@gmail.com/
[6] https://lore.kernel.org/netdev/20260917073958.174056-1-achender@kernel.org/
[7] https://lore.kernel.org/netdev/20260919061149.250658-1-achender@kernel.org/
Allison
Allison Henderson (8):
net/rds: ib: don't enable interrupts in rds_ib_conn_free()
net/rds: free every path's transport data on the passive create paths
net/rds: guard every work-requeueing site with rds_destroy_pending()
net/rds: make rds_destroy_pending() report a connection's own destroy
net/rds: unlink transport nodes before a possibly deferred connection
free
net/rds: take cp_lock to purge cp_send_queue in the quiesce
net/rds: pin the connection across RDMA-CM event handling
net/rds: drop rds_conn_count in favor of t_conn_count
Sharath Srinivasan (4):
net/rds: split connection destroy into quiesce and kref-governed free
net/rds: wait for connections to be freed on transport unload
net/rds: hold connection references in lookup, sockets and c_passive
net/rds: hold a connection reference from struct rds_incoming
net/rds/af_rds.c | 24 ++-
net/rds/connection.c | 330 +++++++++++++++++++++++++++++++++------
net/rds/ib.c | 22 ++-
net/rds/ib.h | 4 +
net/rds/ib_cm.c | 29 +++-
net/rds/ib_rdma.c | 67 ++++++--
net/rds/ib_recv.c | 6 +-
net/rds/ib_send.c | 18 ++-
net/rds/loop.c | 61 ++++++--
net/rds/message.c | 16 +-
net/rds/rdma_transport.c | 16 +-
net/rds/rds.h | 46 +++++-
net/rds/recv.c | 26 ++-
net/rds/send.c | 68 +++++++-
net/rds/tcp.c | 53 ++++++-
net/rds/tcp_listen.c | 21 ++-
16 files changed, 690 insertions(+), 117 deletions(-)
--
2.25.1
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH net-next v6 01/12] net/rds: ib: don't enable interrupts in rds_ib_conn_free()
2026-09-22 16:43 [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
@ 2026-09-22 16:43 ` Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 02/12] net/rds: undo conn_alloc() the same way on every __rds_conn_create() exit Allison Henderson
` (3 subsequent siblings)
4 siblings, 0 replies; 7+ messages in thread
From: Allison Henderson @ 2026-09-22 16:43 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms; +Cc: achender
rds_ib_conn_free() unlinks the connection from its device or nodev
list under spin_lock_irq()/spin_unlock_irq(). It is not only called
from the rmmod path, though: __rds_conn_create() calls
trans->conn_free() to undo a lost creation race while it still holds
rds_conn_lock, taken with spin_lock_irqsave(). The unconditional
spin_unlock_irq() then re-enables interrupts with rds_conn_lock held
and leaves them enabled when the caller's spin_unlock_irqrestore()
runs, defeating the irqsave the caller relied on.
Use the irqsave/irqrestore pair, as rds_tcp_conn_free() and
rds_loop_conn_free() already do.
Fixes: 745cbccac3fe ("RDS: Rewrite connection cleanup, fixing oops on rmmod")
Assisted-by: Claude-Code:claude-fable-5
Signed-off-by: Allison Henderson <achender@kernel.org>
---
net/rds/ib_cm.c | 9 +++++++--
1 file changed, 7 insertions(+), 2 deletions(-)
diff --git a/net/rds/ib_cm.c b/net/rds/ib_cm.c
index 6e3110a04ae6..53147793d44b 100644
--- a/net/rds/ib_cm.c
+++ b/net/rds/ib_cm.c
@@ -1271,6 +1271,7 @@ void rds_ib_conn_free(void *arg)
{
struct rds_ib_connection *ic = arg;
spinlock_t *lock_ptr;
+ unsigned long flags;
rdsdebug("ic %p\n", ic);
@@ -1278,12 +1279,16 @@ void rds_ib_conn_free(void *arg)
* Conn is either on a dev's list or on the nodev list.
* A race with shutdown() or connect() would cause problems
* (since rds_ibdev would change) but that should never happen.
+ *
+ * Callers may hold rds_conn_lock with interrupts disabled
+ * (__rds_conn_create() undoing a lost creation race), so do not
+ * re-enable interrupts unconditionally here.
*/
lock_ptr = ic->rds_ibdev ? &ic->rds_ibdev->spinlock : &ib_nodev_conns_lock;
- spin_lock_irq(lock_ptr);
+ spin_lock_irqsave(lock_ptr, flags);
list_del(&ic->ib_node);
- spin_unlock_irq(lock_ptr);
+ spin_unlock_irqrestore(lock_ptr, flags);
rds_ib_recv_free_caches(ic);
--
2.25.1
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH net-next v6 02/12] net/rds: undo conn_alloc() the same way on every __rds_conn_create() exit
2026-09-22 16:43 [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 01/12] net/rds: ib: don't enable interrupts in rds_ib_conn_free() Allison Henderson
@ 2026-09-22 16:43 ` Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 03/12] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
` (2 subsequent siblings)
4 siblings, 0 replies; 7+ messages in thread
From: Allison Henderson @ 2026-09-22 16:43 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms; +Cc: achender
trans->conn_alloc() may allocate transport data for every path of a
multipath connection - rds_tcp_conn_alloc() does - which is why the
lost-creation-race exit of __rds_conn_create() loops over all npaths
when it frees the connection it just built. The passive-connection
exit right above it frees only path 0.
That is not a leak today: a passive twin is only created for an IB
loopback connection (an incoming TCP connect to a local address is
refused with -EOPNOTSUPP before it gets here), and the IB transport is
not multipath, so npaths is 1 on that exit. But the two exits express
the same "undo conn_alloc()" step in two different ways, and the
following patches add another exit of the same kind. Move the loop
into a helper and use it everywhere, so that the step cannot silently
diverge if a multipath transport ever grows a passive twin.
Assisted-by: Claude-Code:claude-fable-5
Signed-off-by: Allison Henderson <achender@kernel.org>
---
net/rds/connection.c | 31 ++++++++++++++++++-------------
1 file changed, 18 insertions(+), 13 deletions(-)
diff --git a/net/rds/connection.c b/net/rds/connection.c
index b6c4beb50eaf..a96569a3ee9a 100644
--- a/net/rds/connection.c
+++ b/net/rds/connection.c
@@ -161,6 +161,22 @@ static void __rds_conn_path_init(struct rds_connection *conn,
cp->cp_flags = 0;
}
+/* Undo trans->conn_alloc(): it may have allocated transport data for
+ * every path of a multipath connection, not just for path 0.
+ */
+static void rds_conn_free_transport_data(struct rds_connection *conn,
+ int npaths)
+{
+ struct rds_conn_path *cp;
+ int i;
+
+ for (i = 0; i < npaths; i++) {
+ cp = &conn->c_path[i];
+ if (cp->cp_transport_data)
+ conn->c_trans->conn_free(cp->cp_transport_data);
+ }
+}
+
/*
* There is only every one 'conn' for a given pair of addresses in the
* system at a time. They contain messages to be retransmitted and so
@@ -316,7 +332,7 @@ static struct rds_connection *__rds_conn_create(struct net *net,
if (parent) {
/* Creating passive conn */
if (parent->c_passive) {
- trans->conn_free(conn->c_path[0].cp_transport_data);
+ rds_conn_free_transport_data(conn, npaths);
free_cp = conn->c_path;
kmem_cache_free(rds_conn_slab, conn);
conn = parent->c_passive;
@@ -332,18 +348,7 @@ static struct rds_connection *__rds_conn_create(struct net *net,
found = rds_conn_lookup(net, head, laddr, faddr, trans,
tos, dev_if);
if (found) {
- struct rds_conn_path *cp;
- int i;
-
- for (i = 0; i < npaths; i++) {
- cp = &conn->c_path[i];
- /* The ->conn_alloc invocation may have
- * allocated resource for all paths, so all
- * of them may have to be freed here.
- */
- if (cp->cp_transport_data)
- trans->conn_free(cp->cp_transport_data);
- }
+ rds_conn_free_transport_data(conn, npaths);
free_cp = conn->c_path;
kmem_cache_free(rds_conn_slab, conn);
conn = found;
--
2.25.1
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH net-next v6 03/12] net/rds: guard every work-requeueing site with rds_destroy_pending()
2026-09-22 16:43 [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 01/12] net/rds: ib: don't enable interrupts in rds_ib_conn_free() Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 02/12] net/rds: undo conn_alloc() the same way on every __rds_conn_create() exit Allison Henderson
@ 2026-09-22 16:43 ` Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 04/12] net/rds: make rds_destroy_pending() report a connection's own destroy Allison Henderson
2026-09-22 16:47 ` [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
4 siblings, 0 replies; 7+ messages in thread
From: Allison Henderson @ 2026-09-22 16:43 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms; +Cc: achender
rds_conn_destroy() cancels the path works and then destroys the
per-path workqueue. The sites that can re-arm those works are
supposed to test rds_destroy_pending() under rcu_read_lock() first,
paired with the synchronize_rcu() in the destroy path, so that no new
work can be queued once the cancellation has begun.
Five arming sites never got that guard:
- rds_ib_send_cqe_handler() and rds_ib_send_add_credits() re-arm
cp_send_w when a send completion or a credit update clears
RDS_LL_SEND_FULL,
- rds_ib_recv_refill() re-arms cp_recv_w when the recv ring runs
low,
- rds_tcp_accept_one() kicks cp_recv_w on the freshly accepted
socket, and
- rds_sendmsg() arms cp_conn_w for a multipath connection whose
path 0 is not up yet.
The IB completion sites are reachable from soft-irq at any point
before the QP is drained, so a completion landing in the window
between the cancel and destroy_workqueue() in rds_conn_path_destroy()
re-arms a work on a workqueue that is about to be destroyed: with
delay 0 the work is queued directly on the freed workqueue, and with
delay 1 the timer survives destroy_workqueue() unseen and fires
afterwards, queueing from a timer_list that lives in the freed c_path
array.
Wrap all five sites in the same rcu_read_lock() +
rds_destroy_pending() pattern the other arming sites already use.
The four self-requeues in rds_send_worker() and rds_recv_worker() are
left alone on purpose: they run from inside the work item itself, and
cancel_delayed_work_sync() disables the work for the duration of the
cancel, so a requeue issued by the still-running callback is dropped
and none can follow once the cancel has returned.
With the predicate as it stands the guards cover the netns teardown and
module unload cases; the following patch extends it to the destroy of a
single connection.
Fixes: ebeeb1ad9b8a ("rds: tcp: use rds_destroy_pending() to synchronize netns/module teardown and rds connection/workq management")
Assisted-by: Claude-Code:claude-fable-5
Signed-off-by: Allison Henderson <achender@kernel.org>
---
net/rds/ib_recv.c | 6 +++++-
net/rds/ib_send.c | 18 ++++++++++++++----
net/rds/send.c | 11 ++++++++---
net/rds/tcp_listen.c | 10 +++++++---
4 files changed, 34 insertions(+), 11 deletions(-)
diff --git a/net/rds/ib_recv.c b/net/rds/ib_recv.c
index bd6cb3ffaa57..7d45808544a0 100644
--- a/net/rds/ib_recv.c
+++ b/net/rds/ib_recv.c
@@ -458,7 +458,11 @@ void rds_ib_recv_refill(struct rds_connection *conn, int prefill, gfp_t gfp)
(must_wake ||
(can_wait && rds_ib_ring_low(&ic->i_recv_ring)) ||
rds_ib_ring_empty(&ic->i_recv_ring))) {
- queue_delayed_work(conn->c_path->cp_wq, &conn->c_recv_w, 1);
+ rcu_read_lock();
+ if (!rds_destroy_pending(conn))
+ queue_delayed_work(conn->c_path->cp_wq,
+ &conn->c_recv_w, 1);
+ rcu_read_unlock();
}
if (can_wait)
cond_resched();
diff --git a/net/rds/ib_send.c b/net/rds/ib_send.c
index d6be95542119..bc411e96ad12 100644
--- a/net/rds/ib_send.c
+++ b/net/rds/ib_send.c
@@ -298,8 +298,13 @@ void rds_ib_send_cqe_handler(struct rds_ib_connection *ic, struct ib_wc *wc)
rds_ib_sub_signaled(ic, nr_sig);
if (test_and_clear_bit(RDS_LL_SEND_FULL, &conn->c_flags) ||
- test_bit(0, &conn->c_map_queued))
- queue_delayed_work(conn->c_path->cp_wq, &conn->c_send_w, 0);
+ test_bit(0, &conn->c_map_queued)) {
+ rcu_read_lock();
+ if (!rds_destroy_pending(conn))
+ queue_delayed_work(conn->c_path->cp_wq,
+ &conn->c_send_w, 0);
+ rcu_read_unlock();
+ }
/* We expect errors as the qp is drained during shutdown */
if (wc->status != IB_WC_SUCCESS && rds_conn_up(conn)) {
@@ -420,8 +425,13 @@ void rds_ib_send_add_credits(struct rds_connection *conn, unsigned int credits)
test_bit(RDS_LL_SEND_FULL, &conn->c_flags) ? ", ll_send_full" : "");
atomic_add(IB_SET_SEND_CREDITS(credits), &ic->i_credits);
- if (test_and_clear_bit(RDS_LL_SEND_FULL, &conn->c_flags))
- queue_delayed_work(conn->c_path->cp_wq, &conn->c_send_w, 0);
+ if (test_and_clear_bit(RDS_LL_SEND_FULL, &conn->c_flags)) {
+ rcu_read_lock();
+ if (!rds_destroy_pending(conn))
+ queue_delayed_work(conn->c_path->cp_wq,
+ &conn->c_send_w, 0);
+ rcu_read_unlock();
+ }
WARN_ON(IB_GET_SEND_CREDITS(credits) >= 16384);
diff --git a/net/rds/send.c b/net/rds/send.c
index 1afa981e5c06..32c411d10e3e 100644
--- a/net/rds/send.c
+++ b/net/rds/send.c
@@ -1378,9 +1378,14 @@ int rds_sendmsg(struct socket *sock, struct msghdr *msg, size_t payload_len)
* outstanding.
*/
if (!test_and_set_bit(RDS_RECONNECT_PENDING,
- &conn->c_path[0].cp_flags))
- queue_delayed_work(conn->c_path[0].cp_wq,
- &conn->c_path[0].cp_conn_w, 0);
+ &conn->c_path[0].cp_flags)) {
+ rcu_read_lock();
+ if (!rds_destroy_pending(conn))
+ queue_delayed_work(conn->c_path[0].cp_wq,
+ &conn->c_path[0].cp_conn_w,
+ 0);
+ rcu_read_unlock();
+ }
rds_send_ping(conn, 0);
}
diff --git a/net/rds/tcp_listen.c b/net/rds/tcp_listen.c
index 13fa60c1985b..8a0c54aced5e 100644
--- a/net/rds/tcp_listen.c
+++ b/net/rds/tcp_listen.c
@@ -316,10 +316,14 @@ int rds_tcp_accept_one(struct rds_tcp_net *rtn)
*/
if (READ_ONCE(sk->sk_state) == TCP_CLOSE_WAIT ||
READ_ONCE(sk->sk_state) == TCP_LAST_ACK ||
- READ_ONCE(sk->sk_state) == TCP_CLOSE)
+ READ_ONCE(sk->sk_state) == TCP_CLOSE) {
rds_conn_path_drop(cp, 0);
- else
- queue_delayed_work(cp->cp_wq, &cp->cp_recv_w, 0);
+ } else {
+ rcu_read_lock();
+ if (!rds_destroy_pending(cp->cp_conn))
+ queue_delayed_work(cp->cp_wq, &cp->cp_recv_w, 0);
+ rcu_read_unlock();
+ }
sock_put(sk);
--
2.25.1
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH net-next v6 04/12] net/rds: make rds_destroy_pending() report a connection's own destroy
2026-09-22 16:43 [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
` (2 preceding siblings ...)
2026-09-22 16:43 ` [PATCH net-next v6 03/12] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
@ 2026-09-22 16:43 ` Allison Henderson
2026-09-22 16:47 ` [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
4 siblings, 0 replies; 7+ messages in thread
From: Allison Henderson @ 2026-09-22 16:43 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms; +Cc: achender
rds_conn_destroy() cancels the path works and then destroys the
per-path workqueue. However, nothing currently stops the
work-requeueing sites from queueing new work on the connection while
that happens. The existing code would suggest that this protection
is supposed to come from rds_destroy_pending(), since all of those
sites - apart from the workers' own self-requeues, which the sync
cancel in the destroy path already rejects - guard the queueing with
rds_destroy_pending() under rcu_read_lock() (the last stragglers were
converted by the previous patch), and rds_conn_destroy() already issues a
synchronize_rcu() after unhashing the connection. But the predicate
only tests for the two global teardown cases (netns destruction via
check_net(), module unload via ->t_unloading). Because the conn
itself lacks any indication that a destroy is in progress, the
predicate does not cover the destruction of a single connection
outside these two cases.
Today every rds_conn_destroy() does happen on one of those two global
paths - the last single-connection caller, the protocol-version
mismatch in rds_ib_cm_connect_complete(), was turned into a drop by
commit f97d8c7bab78 ("rds: ib: use rds_conn_drop() on protocol
version mismatch") - so the predicate is currently never wrong. It is
also never precise: it answers "is this connection's world going
away", not "is this connection being destroyed", and the following
patches need the second answer. Once the free is deferred to the
last reference, a connection can be handed to rds_conn_destroy() more
than once (the IB unload path re-sweeps its list until every
connection is gone) and must recognise its own destroy in progress;
the passive-twin creation must refuse a parent whose destroy has
begun; and a future single-connection destroy - a hot-unplugged IB
device, or the asynchronous teardown the Oracle UEK kernel has -
would re-open the window below. While a single-connection destroy
runs, a concurrent rds_cong_queue_updates() can
still find the connection on the congestion map's m_conn_list (the
conn is only removed from it after the paths are torn down) and call
queue_delayed_work() on a cp_wq that destroy_workqueue() has already
freed. Additionally, the other requeueing sites can likewise re-arm
works that live in the about-to-be-freed connection unless
rds_destroy_pending() has something to guard it with.
The version-mismatch path used to be covered: commit c90ecbfaf50d2
("rds: Use atomic flag to track connections being destroyed")
introduced the RDS_DESTROY_PENDING cp_flags bit for exactly this, and
after commit ebeeb1ad9b8ad ("rds: tcp: use rds_destroy_pending() to
synchronize netns/module teardown and rds connection/workq management")
it was set right before that rds_conn_destroy() call and
tested via rds_ib_is_unloading(). Commit cdc306a5c9cd3 ("rds: make
v3.1 as compat version") then removed the last set_bit while leaving
the test behind, so the bit has been dead ever since and per-conn
destroy has run unguarded.
Record the destroy on the connection itself: set
conn->c_destroy_in_prog before the unhash + synchronize_rcu() sequence
in rds_conn_destroy() and test it first in rds_destroy_pending(). The
existing rcu_read_lock() around every check-and-queue site pairs with
that synchronize_rcu(): once it returns, every new reader observes the
flag and refuses to queue, and anything queued before it is flushed or
cancelled by the existing teardown. Drop the now-unreferenced
RDS_DESTROY_PENDING bit and its dead test.
In the Oracle UEK kernel the equivalent conn->c_destroy_in_prog flag
is part of the larger connection refcounting rework ("net/rds: Add
krefs to struct rds_connection"), including ("net/rds: Merge uses of
conn->c_destroy_in_prog & RDS_DESTROY_PENDING"). This ports the
missing pieces of the requeue guard, which stand on their own.
Suggested-by: Sharath Srinivasan <sharath.srinivasan@oracle.com>
Assisted-by: Claude-Code:claude-fable-5
Signed-off-by: Allison Henderson <achender@kernel.org>
---
net/rds/connection.c | 8 ++++++++
net/rds/ib.c | 5 +----
net/rds/rds.h | 12 ++++++++++--
3 files changed, 19 insertions(+), 6 deletions(-)
diff --git a/net/rds/connection.c b/net/rds/connection.c
index a96569a3ee9a..242ca0570a47 100644
--- a/net/rds/connection.c
+++ b/net/rds/connection.c
@@ -579,6 +579,14 @@ void rds_conn_destroy(struct rds_connection *conn)
"%pI4\n", conn, &conn->c_laddr,
&conn->c_faddr);
+ /* Make rds_destroy_pending() true for this conn. Together with
+ * the synchronize_rcu() below this stops the work-requeueing
+ * sites (which all test rds_destroy_pending() under
+ * rcu_read_lock()) from queueing new work on the path
+ * workqueues once we start cancelling and destroying them.
+ */
+ WRITE_ONCE(conn->c_destroy_in_prog, true);
+
/* Ensure conn will not be scheduled for reconnect */
spin_lock_irq(&rds_conn_lock);
hlist_del_init_rcu(&conn->c_hash_node);
diff --git a/net/rds/ib.c b/net/rds/ib.c
index 786f39169bc1..9fe3b9951bd3 100644
--- a/net/rds/ib.c
+++ b/net/rds/ib.c
@@ -525,10 +525,7 @@ static void rds_ib_set_unloading(void)
static bool rds_ib_is_unloading(struct rds_connection *conn)
{
- struct rds_conn_path *cp = &conn->c_path[0];
-
- return (test_bit(RDS_DESTROY_PENDING, &cp->cp_flags) ||
- atomic_read(&rds_ib_unloading) != 0);
+ return atomic_read(&rds_ib_unloading) != 0;
}
void rds_ib_exit(void)
diff --git a/net/rds/rds.h b/net/rds/rds.h
index 2db49573dacd..50b08c28ab86 100644
--- a/net/rds/rds.h
+++ b/net/rds/rds.h
@@ -89,7 +89,6 @@ enum {
#define RDS_RECONNECT_PENDING 1
#define RDS_IN_XMIT 2
#define RDS_RECV_REFILL 3
-#define RDS_DESTROY_PENDING 4
/* Max number of multipaths per RDS connection. Must be a power of 2 */
#define RDS_MPATH_WORKERS 8
@@ -148,6 +147,14 @@ struct rds_connection {
c_pad_to_32:29;
int c_npaths;
bool c_with_sport_idx;
+ /* Set once, by rds_conn_destroy(), before it cancels the path
+ * works; read through rds_destroy_pending(). A site that arms
+ * a path work must test the predicate and queue the work inside
+ * one rcu_read_lock() section: the synchronize_rcu() that
+ * follows the store is what keeps a queue issued after the
+ * cancellation from landing on a destroyed workqueue.
+ */
+ bool c_destroy_in_prog;
struct rds_connection *c_passive;
struct rds_transport *c_trans;
@@ -994,7 +1001,8 @@ void __rds_put_mr_final(struct kref *kref);
static inline bool rds_destroy_pending(struct rds_connection *conn)
{
- return !check_net(rds_conn_net(conn)) ||
+ return READ_ONCE(conn->c_destroy_in_prog) ||
+ !check_net(rds_conn_net(conn)) ||
(conn->c_trans->t_unloading && conn->c_trans->t_unloading(conn));
}
--
2.25.1
^ permalink raw reply related [flat|nested] 7+ messages in thread
* Re: [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted
2026-09-22 16:43 [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
` (3 preceding siblings ...)
2026-09-22 16:43 ` [PATCH net-next v6 04/12] net/rds: make rds_destroy_pending() report a connection's own destroy Allison Henderson
@ 2026-09-22 16:47 ` Allison Henderson
4 siblings, 0 replies; 7+ messages in thread
From: Allison Henderson @ 2026-09-22 16:47 UTC (permalink / raw)
To: netdev, linux-rdma, pabeni, edumazet, kuba, horms
On Tue, 2026-09-22 at 09:43 -0700, Allison Henderson wrote:
> Hi all,
>
Resend of v6 sent by mistake, please ignore the resend.
Apologies for the spam!
Thanks!
Allison
> This is v6 of the connection-lifetime set (v1 at [1], v2 at [2],
> v3 at [3], v4 at [6], v5 at [7]), following
> "net/rds: own the fastpath locks across connection teardown", which is
> in net-next. It is targeted at net-next: although the series fixes
> real use-after-frees (one syzbot report and one report from Chengfeng
> Ye below), it does so by reworking connection lifetime rather than
> patching the individual crash sites, and that rework is too invasive
> for net.
>
> rds_conn_destroy() frees the connection, its paths and its workqueues
> on the spot, relying on the documented assumption that "no one else is
> referencing the connection". That assumption stopped being true a
> long time ago: connections are destroyed on netns teardown and on a
> protocol-version mismatch as well as on rmmod, while pointers to them
> still live in socket rs_conn caches, congestion-map conn lists, CM
> callbacks and workers, and - for as long as an application leaves data
> unread - in every rds_incoming sitting on a receive queue. No single
> Fixes: commit covers the rot, so the series carries Reported-by tags
> where there are concrete reports instead.
>
> Patch 1 fixes rds_ib_conn_free() re-enabling interrupts under
> a caller's irqsave lock.
>
> Patch 2 makes the passive-connection exits of __rds_conn_create()
> undo conn_alloc() the same way the lost-race exit does, through one
> helper (a refactor: the passive twin only exists for IB, so nothing
> leaked there). Both are first in the series because the
> reference-counted teardown calls conn_free() from more places.
>
> Patch 3 guards the five work-arming sites that never tested
> rds_destroy_pending(): two IB completion paths, the IB recv refill,
> the TCP accept kick and the multipath reconnect in sendmsg.
>
> Patch 4 gives the connection its own destroy-in-progress marker.
> With f97d8c7bab78 in net-next there is no single-connection destroy
> left for the old predicate to miss, but the later patches need the
> precise answer: the IB unload re-sweep hands a connection to
> rds_conn_destroy() more than once, and the passive-twin creation
> must refuse a parent whose destroy has begun.
> Based on UEK commits:
> e2f5005adf63 net/rds: Add krefs to struct rds_connection
> https://github.com/oracle/linux-uek/commit/e2f5005adf63
>
> 6c53ef92f46e net/rds: Merge uses of conn->c_destroy_in_prog & RDS_DESTROY_PENDING
> https://github.com/oracle/linux-uek/commit/6c53ef92f46e
>
> Patch 5 splits rds_conn_destroy() into a synchronous quiesce and a
> kref-governed free.
> Based on UEK commits:
> 2c8569e4c880 ("net/rds: Add krefs to struct rds_connection").
> https://github.com/oracle/linux-uek/commit/2c8569e4c880
>
> Patch 6 makes each transport's exit path wait for its connections to
> actually be freed before the module goes away. The wait is
> unbounded and warns every ten seconds: a connection reference can be
> held for an application-controlled time (unread data), so a timeout
> would only move the use-after-free from freed memory to unloaded
> module text. rmmod blocking while data is queued and unread is the
> historical RDS contract. For IB the wait re-sweeps the nodev list,
> since device connections migrate to it asynchronously.
> Based on UEK commits:
> ece4b4e39afa ("net/rds: wait_event_timeout until zero connections during rmmod")
> https://github.com/oracle/linux-uek/commit/ece4b4e39afa
>
> 905ec90e6166 ("net/rds: Each RDS transport should keep its own connection count")
> https://github.com/oracle/linux-uek/commit/905ec90e6166
>
> Patch 7 unlinks each transport node before its destroy, so a
> free deferred past the teardown loop cannot write into the loop's
> stack-local list head. IB marks a gathered node as claimed by the
> sweep, since its connect and shutdown paths move the node too.
>
> Patch 8 hands out real references everywhere a connection pointer
> previously escaped bare, RCU-annotates parent->c_passive, refuses to
> revive a passive connection whose destroy has begun, and serializes
> the SIOCRDSSETTOS check with the rs_conn cache.
> Based on UEK commits:
> 2c8569e4c880 ("net/rds: Add krefs to struct rds_connection")
> https://github.com/oracle/linux-uek/commit/2c8569e4c880
>
> 0e9e3a72b7f7 ("net/rds: rds_sendmsg must use rs_conn only when not being destroyed").
> https://github.com/oracle/linux-uek/commit/0e9e3a72b7f7
>
> Patch 9 has the quiesce purge cp_send_queue under cp_lock, as every
> adder to that queue holds it. No sender can be in flight at any of
> today's destroy triggers (netns teardown, module unload), so this is
> hygiene rather than a race fix, and the v5 refusals in the senders
> that guarded against that impossible state are gone.
>
> Patch 10 pins the connection across the RDMA-CM event handler
> and rejects a connect request for a connection whose destroy has
> already quiesced it.
>
> Patch 11 drops the now-unused global rds_conn_count.
> Based on UEK commits:
> 905ec90e6166 ("net/rds: Each RDS transport should keep its own connection count").
> https://github.com/oracle/linux-uek/commit/905ec90e6166
>
> Patch 12 makes struct rds_incoming hold a reference on i_conn, the
> fix for the KASAN use-after-free Chengfeng Ye reported [4].
> Based on UEK commits:
> 99b9a3715419 ("net/rds: fix crash by expanding kref coverage to rds_incoming.i_conn").
> https://github.com/oracle/linux-uek/commit/99b9a3715419
>
> Patches 5, 8, 6 and 12 are ports of the connection kref work Sharath
> Srinivasan did for Oracle UEK, adapted to the upstream code.
>
> The series has been validated with the RDS selftests over loopback-TCP
> and RXE-RDMA, plus churn tests that delete network namespaces and
> unload the modules under live rds-stress traffic.
>
> Changes since v5 [7]:
> - Rebased; the version-mismatch destroy is now an rds_conn_drop() in
> net-next (f97d8c7bab78), so patch 4's Fixes: tag is gone and its
> changelog describes what still needs the per-connection flag, and
> the CM-callback deadlock reports against earlier versions no longer
> apply.
> - Patch 2: described as the refactor it is (the passive twin is
> IB-only, single path); Fixes: tag dropped.
> - Patch 9: reduced to the cp_lock purge; the send-side refusals and
> the negative *queued signalling are dropped, since no sender can be
> in flight at a destroy trigger.
> - Patch 6: wait comment describes the initial-sweep-plus-resweep
> contract; changelog notes the guarantee is complete only once incs
> hold references.
> - Patch 8: c_destroy_in_prog read through rds_destroy_pending() in
> __rds_conn_create(); the c_passive puts documented as possibly the
> last; READ_ONCE() on the unlocked rs_tos sample; rs_conn comment
> notes the rds_release() exception.
> - Patch 10: stale forward-reference comment fixed.
>
> Changes since v4 [6]:
> - Patch 7: the IB sweep marks each gathered node with a new
> i_ib_node_detached flag instead of relying on list emptiness. A
> node parked on the sweep's stack list is not empty, so a concurrent
> rds_ib_add_conn() would have unlinked it from under the sweep's
> lockless walk; rds_ib_add_conn(), rds_ib_remove_conn() and
> rds_ib_conn_free() now leave a claimed node alone.
>
> Changes since v3 [3]:
> - Rebased; the version-mismatch destroy that patch 10 also had to
> cope with is now an rds_conn_drop() in net (f97d8c7bab78), and
> patch 10's changelog says what remains for it to cover.
> - The uninitialised transport pointer in the CM event handler that
> the second v2 review pass raised turned out to be the bug Aohan
> Mei already posted a fix for (v2 at [5], stalled after review); it
> is carried forward separately for net rather than added here.
> - Patch 7: the teardown walks take a reference on each gathered
> connection, so a connection destroyed earlier and freed by a
> pending holder cannot vanish under the iterator, and the two IB
> list movers tolerate a node the sweep already unlinked instead of
> BUG_ON()ing; the nodev sweep gathers entry by entry, so the
> resweep from the unload wait no longer relies on list_splice_init().
> - Patch 8: SIOCRDSSETTOS check-then-act closed - the install in
> rds_sendmsg() re-checks the socket's ToS under rs_lock; rs_conn
> documented as a referenced, rs_lock-serialised cache; contract
> comment restored to rds_conn_lookup().
> - Patch 9: rds_send_probe() gets the same rds_destroy_pending()
> test under cp_lock as rds_send_queue_rm() (a probe queued after the
> purge pinned the connection for good).
> - v3 patch 10 (rds_tcp_accept_one() destroy check) dropped: the check
> was not serialised against the destroy, and the race it aimed at is
> not reachable - every TCP destroy path stops the listener, flushing
> the accept work, first. A comment now records that ordering.
> - Changelog and comment corrections from the second v2 review pass
> (self-requeue exemption in patch 3, c_refcount comment in patch 5,
> the uninterruptible wait and the transport-text wake in patch 6,
> "lock-free" wording in patch 11, cross-netns comment in patch 12).
>
> Changes since v2 [2]:
> - New patches 1 and 2: pre-existing rds_ib_conn_free() interrupt
> state clobber and passive-path transport data leak, surfaced by
> review of the teardown changes.
> - Patch 6: rds_ib_destroy_nodev_conns() uses list_splice_init(), so
> the resweep from the unload wait cannot splice a stale list head.
> - Patch 8: SIOCRDSGETTOS reads rs_tos under rs_lock like SETTOS.
> - New patch 9: rds_send_queue_rm() refuses a connection whose destroy
> has begun, under cp_lock, and the quiesce purges cp_send_queue
> under cp_lock (list corruption against an in-flight sender, and a
> message that would pin the connection forever).
> - New patch 10: rds_tcp_accept_one() does not install a socket on a
> connection whose destroy has begun (socket left pointing at a freed
> path).
>
> Changes since v1 [1]:
> - New patch 1: guard the five work-arming sites that never tested
> rds_destroy_pending(); patch 2's changelog and comments narrowed
> to what it actually newly covers.
> - Patch 4 moved ahead of the reference holders, so no bisect point
> has references without the unload wait; wait made unbounded with a
> periodic warning instead of a 10 s timeout; IB exit re-sweeps the
> nodev list for connections that detach from their device late.
> - New patch 5: transport nodes unlinked before destroy (stack list
> head use-after-free from a deferred conn_free).
> - Patch 6: c_passive RCU-annotated; a destroyed passive child is
> refused by __rds_conn_create() and clears the parent's pointer
> itself; SIOCRDSSETTOS uses rs_lock; lookup comment reworded;
> explicit not-for-stable note.
> - New patch 7: reference across the CM event handler, and a
> destroy-pending re-check in rds_ib_cm_handle_connect().
> - Patch 9: changelog states what a lingering inc keeps alive and
> that it blocks module unload.
> - Changelog corrections throughout (netns teardown paths named as the
> non-rmmod destroyers, stale rds_conn_path_destroy() reference).
>
> [1] https://lore.kernel.org/netdev/20260904070248.160384-1-achender@kernel.org/
> [2] https://lore.kernel.org/netdev/20260912035027.27447-1-achender@kernel.org/
> [3] https://lore.kernel.org/netdev/20260914033719.138057-1-achender@kernel.org/
> [4] https://lore.kernel.org/netdev/20260720184955.3008978-1-nicoyip.dev@gmail.com/
> [5] https://lore.kernel.org/netdev/20260825021223.3483044-1-ljp1205831794@gmail.com/
> [6] https://lore.kernel.org/netdev/20260917073958.174056-1-achender@kernel.org/
> [7] https://lore.kernel.org/netdev/20260919061149.250658-1-achender@kernel.org/
>
> Allison
>
>
> Allison Henderson (8):
> net/rds: ib: don't enable interrupts in rds_ib_conn_free()
> net/rds: free every path's transport data on the passive create paths
> net/rds: guard every work-requeueing site with rds_destroy_pending()
> net/rds: make rds_destroy_pending() report a connection's own destroy
> net/rds: unlink transport nodes before a possibly deferred connection
> free
> net/rds: take cp_lock to purge cp_send_queue in the quiesce
> net/rds: pin the connection across RDMA-CM event handling
> net/rds: drop rds_conn_count in favor of t_conn_count
>
> Sharath Srinivasan (4):
> net/rds: split connection destroy into quiesce and kref-governed free
> net/rds: wait for connections to be freed on transport unload
> net/rds: hold connection references in lookup, sockets and c_passive
> net/rds: hold a connection reference from struct rds_incoming
>
> net/rds/af_rds.c | 24 ++-
> net/rds/connection.c | 330 +++++++++++++++++++++++++++++++++------
> net/rds/ib.c | 22 ++-
> net/rds/ib.h | 4 +
> net/rds/ib_cm.c | 29 +++-
> net/rds/ib_rdma.c | 67 ++++++--
> net/rds/ib_recv.c | 6 +-
> net/rds/ib_send.c | 18 ++-
> net/rds/loop.c | 61 ++++++--
> net/rds/message.c | 16 +-
> net/rds/rdma_transport.c | 16 +-
> net/rds/rds.h | 46 +++++-
> net/rds/recv.c | 26 ++-
> net/rds/send.c | 68 +++++++-
> net/rds/tcp.c | 53 ++++++-
> net/rds/tcp_listen.c | 21 ++-
> 16 files changed, 690 insertions(+), 117 deletions(-)
>
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-09-22 16:47 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-22 16:43 [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 01/12] net/rds: ib: don't enable interrupts in rds_ib_conn_free() Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 02/12] net/rds: undo conn_alloc() the same way on every __rds_conn_create() exit Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 03/12] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
2026-09-22 16:43 ` [PATCH net-next v6 04/12] net/rds: make rds_destroy_pending() report a connection's own destroy Allison Henderson
2026-09-22 16:47 ` [PATCH net-next v6 00/12] net/rds: make connection lifetime reference-counted Allison Henderson
-- strict thread matches above, loose matches on Subject: below --
2026-09-22 8:53 Allison Henderson
2026-09-22 8:54 ` [PATCH net-next v6 03/12] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox