From: netdev-bot+sashiko@kernel.org
To: achender@kernel.org
Cc: netdev@vger.kernel.org, linux-rdma@vger.kernel.org,
pabeni@redhat.com, edumazet@google.com, kuba@kernel.org,
horms@kernel.org, nicoyip.dev@gmail.com
Subject: Re: [PATCH net-next v2 1/9] net/rds: guard every work-requeueing site with rds_destroy_pending()
Date: Wed, 16 Sep 2026 04:36:43 +0000 [thread overview]
Message-ID: <178953340335.22033.8574444684748203869@kernel.org> (raw)
In-Reply-To: <20260912035027.27447-2-achender@kernel.org>
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider:
Critical: 0 · High: 0 · Medium: 0 · Low: 2
- [Low] At this commit the five newly added `rcu_read_lock() +
rds_destroy_pending()` guards (net/rds/ib_send.c:302-306 and :429-433,…
- [Low] The subject claims to guard "every work-requeueing site" and the
body states exactly "Five arming sites never got that guard", but four…
--
net/rds: guard every work-requeueing site with rds_destroy_pending()
This adds rcu_read_lock() + rds_destroy_pending() guards around five
queue_delayed_work() sites that re-arm cp_send_w, cp_recv_w and cp_conn_w:
two in net/rds/ib_send.c, one in net/rds/ib_recv.c, one in
net/rds/tcp_listen.c and one in net/rds/send.c. The stated goal is that no
new path work can be queued once rds_conn_path_destroy() has started
cancelling works and is about to call destroy_workqueue().
A couple of questions on the changelog and on the effect of the guards at
this commit.
> net/rds: guard every work-requeueing site with rds_destroy_pending()
[ ... ]
> Five arming sites never got that guard:
[Severity: Low]
Are those five really all of the requeueing sites? net/rds/threads.c has
four more queue_delayed_work() calls on the same works that stay unguarded.
rds_send_worker():
switch (ret) {
case -EAGAIN:
rds_stats_inc(s_send_immediate_retry);
queue_delayed_work(cp->cp_wq, &cp->cp_send_w, 0);
break;
case -ENOMEM:
rds_stats_inc(s_send_delayed_retry);
queue_delayed_work(cp->cp_wq, &cp->cp_send_w, 2);
rds_recv_worker() has the matching pair for cp_recv_w.
These four are self-requeues issued from inside the running work item, and
in this tree cancel_delayed_work_sync() -> __cancel_work_sync() disables the
work with WORK_CANCEL_DISABLE before __flush_work(), so a requeue from the
still-running callback is rejected while the cancel is in progress and no
use-after-free follows.
The subject says "every work-requeueing site" and the body gives an exact
count of five, and the following patch in the series repeats that "all of
those sites guard the queueing with rds_destroy_pending() under
rcu_read_lock() (the last stragglers were converted by the previous patch)".
Could the changelog either convert the threads.c sites too, or say that
in-callback self-requeues are exempt and why?
The Fixes: tag points at ebeeb1ad9b8a, and on trees before
cancel_delayed_work_sync() gained the disable semantics a self-requeue from
the callback was not rejected, so a backporter reading this message cannot
tell whether those sites were considered.
> diff --git a/net/rds/ib_recv.c b/net/rds/ib_recv.c
> index bd6cb3ffaa571..7d45808544a0d 100644
> --- a/net/rds/ib_recv.c
> +++ b/net/rds/ib_recv.c
> @@ -458,7 +458,11 @@ void rds_ib_recv_refill(struct rds_connection *conn, int prefill, gfp_t gfp)
> (must_wake ||
> (can_wait && rds_ib_ring_low(&ic->i_recv_ring)) ||
> rds_ib_ring_empty(&ic->i_recv_ring))) {
> - queue_delayed_work(conn->c_path->cp_wq, &conn->c_recv_w, 1);
> + rcu_read_lock();
> + if (!rds_destroy_pending(conn))
> + queue_delayed_work(conn->c_path->cp_wq,
> + &conn->c_recv_w, 1);
> + rcu_read_unlock();
> }
[Severity: Low]
Can these guards close the cancel -> destroy_workqueue() window as of this
commit? rds_destroy_pending() in net/rds/rds.h carries no per-connection
destroy state here:
static inline bool rds_destroy_pending(struct rds_connection *conn)
{
return !check_net(rds_conn_net(conn)) ||
(conn->c_trans->t_unloading && conn->c_trans->t_unloading(conn));
}
and rds_conn_destroy() publishes nothing that the guards can observe before
the grace period:
spin_lock_irq(&rds_conn_lock);
hlist_del_init_rcu(&conn->c_hash_node);
spin_unlock_irq(&rds_conn_lock);
synchronize_rcu();
For a single-connection teardown with a live netns and the module loaded,
for example the protocol version mismatch path
rds_ib_cm_connect_complete() -> rds_conn_destroy(), the predicate reads
false, so a completion landing in rds_ib_recv_refill() or
rds_ib_send_cqe_handler() can still arm cp_recv_w/cp_send_w between
cancel_delayed_work_sync() and destroy_workqueue() in
rds_conn_path_destroy().
The next patch in the series, "net/rds: make rds_destroy_pending() cover
single-connection destroy", adds conn->c_destroy_in_prog, sets it under
rds_conn_lock before hlist_del_init_rcu() and synchronize_rcu(), and tests
it first in rds_destroy_pending(), which is what gives these five guards
their effect, so at the end of the series the window is closed.
Given the Fixes: tag invites a standalone backport, would it help to state
in this changelog that the guards only take effect together with the
c_destroy_in_prog change, or to reorder the two patches?
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260912035027.27447-1-achender%40kernel.org
next prev parent reply other threads:[~2026-09-16 4:36 UTC|newest]
Thread overview: 28+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-12 3:50 [PATCH net-next v2 0/9] net/rds: make connection lifetime reference-counted Allison Henderson
2026-09-12 3:50 ` [PATCH net-next v2 1/9] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko [this message]
2026-09-12 3:50 ` [PATCH net-next v2 2/9] net/rds: make rds_destroy_pending() cover single-connection destroy Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko
2026-09-12 3:50 ` [PATCH net-next v2 3/9] net/rds: split connection destroy into quiesce and kref-governed free Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko
2026-09-12 3:50 ` [PATCH net-next v2 4/9] net/rds: wait for connections to be freed on transport unload Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko
2026-09-12 3:50 ` [PATCH net-next v2 5/9] net/rds: unlink transport nodes before a possibly deferred connection free Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko
2026-09-12 3:50 ` [PATCH net-next v2 6/9] net/rds: hold connection references in lookup, sockets and c_passive Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko
2026-09-12 3:50 ` [PATCH net-next v2 7/9] net/rds: pin the connection across RDMA-CM event handling Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko
2026-09-12 3:50 ` [PATCH net-next v2 8/9] net/rds: drop rds_conn_count in favor of t_conn_count Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko
2026-09-12 3:50 ` [PATCH net-next v2 9/9] net/rds: hold a connection reference from struct rds_incoming Allison Henderson
2026-09-13 3:50 ` sashiko-bot
2026-09-16 4:36 ` netdev-bot+sashiko
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=178953340335.22033.8574444684748203869@kernel.org \
--to=netdev-bot+sashiko@kernel.org \
--cc=achender@kernel.org \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=kuba@kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=nicoyip.dev@gmail.com \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.