Netdev List
 help / color / mirror / Atom feed
From: netdev-bot+sashiko@kernel.org
To: achender@kernel.org
Cc: netdev@vger.kernel.org, linux-rdma@vger.kernel.org,
	pabeni@redhat.com, edumazet@google.com, kuba@kernel.org,
	horms@kernel.org, nicoyip.dev@gmail.com
Subject: Re: [PATCH net-next v2 1/9] net/rds: guard every work-requeueing site with rds_destroy_pending()
Date: Wed, 16 Sep 2026 04:36:43 +0000	[thread overview]
Message-ID: <178953340335.22033.8574444684748203869@kernel.org> (raw)
In-Reply-To: <20260912035027.27447-2-achender@kernel.org>

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider:

Critical: 0 · High: 0 · Medium: 0 · Low: 2

- [Low] At this commit the five newly added `rcu_read_lock() +
  rds_destroy_pending()` guards (net/rds/ib_send.c:302-306 and :429-433,…
- [Low] The subject claims to guard "every work-requeueing site" and the
  body states exactly "Five arming sites never got that guard", but four…

--

net/rds: guard every work-requeueing site with rds_destroy_pending()

This adds rcu_read_lock() + rds_destroy_pending() guards around five
queue_delayed_work() sites that re-arm cp_send_w, cp_recv_w and cp_conn_w:
two in net/rds/ib_send.c, one in net/rds/ib_recv.c, one in
net/rds/tcp_listen.c and one in net/rds/send.c.  The stated goal is that no
new path work can be queued once rds_conn_path_destroy() has started
cancelling works and is about to call destroy_workqueue().

A couple of questions on the changelog and on the effect of the guards at
this commit.

> net/rds: guard every work-requeueing site with rds_destroy_pending()

[ ... ]

>     Five arming sites never got that guard:

[Severity: Low]
Are those five really all of the requeueing sites?  net/rds/threads.c has
four more queue_delayed_work() calls on the same works that stay unguarded.

rds_send_worker():

	switch (ret) {
	case -EAGAIN:
		rds_stats_inc(s_send_immediate_retry);
		queue_delayed_work(cp->cp_wq, &cp->cp_send_w, 0);
		break;
	case -ENOMEM:
		rds_stats_inc(s_send_delayed_retry);
		queue_delayed_work(cp->cp_wq, &cp->cp_send_w, 2);

rds_recv_worker() has the matching pair for cp_recv_w.

These four are self-requeues issued from inside the running work item, and
in this tree cancel_delayed_work_sync() -> __cancel_work_sync() disables the
work with WORK_CANCEL_DISABLE before __flush_work(), so a requeue from the
still-running callback is rejected while the cancel is in progress and no
use-after-free follows.

The subject says "every work-requeueing site" and the body gives an exact
count of five, and the following patch in the series repeats that "all of
those sites guard the queueing with rds_destroy_pending() under
rcu_read_lock() (the last stragglers were converted by the previous patch)".
Could the changelog either convert the threads.c sites too, or say that
in-callback self-requeues are exempt and why?

The Fixes: tag points at ebeeb1ad9b8a, and on trees before
cancel_delayed_work_sync() gained the disable semantics a self-requeue from
the callback was not rejected, so a backporter reading this message cannot
tell whether those sites were considered.

> diff --git a/net/rds/ib_recv.c b/net/rds/ib_recv.c
> index bd6cb3ffaa571..7d45808544a0d 100644
> --- a/net/rds/ib_recv.c
> +++ b/net/rds/ib_recv.c
> @@ -458,7 +458,11 @@ void rds_ib_recv_refill(struct rds_connection *conn, int prefill, gfp_t gfp)
>  	    (must_wake ||
>  	    (can_wait && rds_ib_ring_low(&ic->i_recv_ring)) ||
>  	    rds_ib_ring_empty(&ic->i_recv_ring))) {
> -		queue_delayed_work(conn->c_path->cp_wq, &conn->c_recv_w, 1);
> +		rcu_read_lock();
> +		if (!rds_destroy_pending(conn))
> +			queue_delayed_work(conn->c_path->cp_wq,
> +					   &conn->c_recv_w, 1);
> +		rcu_read_unlock();
>  	}

[Severity: Low]
Can these guards close the cancel -> destroy_workqueue() window as of this
commit?  rds_destroy_pending() in net/rds/rds.h carries no per-connection
destroy state here:

static inline bool rds_destroy_pending(struct rds_connection *conn)
{
	return !check_net(rds_conn_net(conn)) ||
	       (conn->c_trans->t_unloading && conn->c_trans->t_unloading(conn));
}

and rds_conn_destroy() publishes nothing that the guards can observe before
the grace period:

	spin_lock_irq(&rds_conn_lock);
	hlist_del_init_rcu(&conn->c_hash_node);
	spin_unlock_irq(&rds_conn_lock);
	synchronize_rcu();

For a single-connection teardown with a live netns and the module loaded,
for example the protocol version mismatch path
rds_ib_cm_connect_complete() -> rds_conn_destroy(), the predicate reads
false, so a completion landing in rds_ib_recv_refill() or
rds_ib_send_cqe_handler() can still arm cp_recv_w/cp_send_w between
cancel_delayed_work_sync() and destroy_workqueue() in
rds_conn_path_destroy().

The next patch in the series, "net/rds: make rds_destroy_pending() cover
single-connection destroy", adds conn->c_destroy_in_prog, sets it under
rds_conn_lock before hlist_del_init_rcu() and synchronize_rcu(), and tests
it first in rds_destroy_pending(), which is what gives these five guards
their effect, so at the end of the series the window is closed.

Given the Fixes: tag invites a standalone backport, would it help to state
in this changelog that the guards only take effect together with the
c_destroy_in_prog change, or to reorder the two patches?

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260912035027.27447-1-achender%40kernel.org

  reply	other threads:[~2026-09-16  4:36 UTC|newest]

Thread overview: 19+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-12  3:50 [PATCH net-next v2 0/9] net/rds: make connection lifetime reference-counted Allison Henderson
2026-09-12  3:50 ` [PATCH net-next v2 1/9] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko [this message]
2026-09-12  3:50 ` [PATCH net-next v2 2/9] net/rds: make rds_destroy_pending() cover single-connection destroy Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko
2026-09-12  3:50 ` [PATCH net-next v2 3/9] net/rds: split connection destroy into quiesce and kref-governed free Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko
2026-09-12  3:50 ` [PATCH net-next v2 4/9] net/rds: wait for connections to be freed on transport unload Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko
2026-09-12  3:50 ` [PATCH net-next v2 5/9] net/rds: unlink transport nodes before a possibly deferred connection free Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko
2026-09-12  3:50 ` [PATCH net-next v2 6/9] net/rds: hold connection references in lookup, sockets and c_passive Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko
2026-09-12  3:50 ` [PATCH net-next v2 7/9] net/rds: pin the connection across RDMA-CM event handling Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko
2026-09-12  3:50 ` [PATCH net-next v2 8/9] net/rds: drop rds_conn_count in favor of t_conn_count Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko
2026-09-12  3:50 ` [PATCH net-next v2 9/9] net/rds: hold a connection reference from struct rds_incoming Allison Henderson
2026-09-16  4:36   ` netdev-bot+sashiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=178953340335.22033.8574444684748203869@kernel.org \
    --to=netdev-bot+sashiko@kernel.org \
    --cc=achender@kernel.org \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=nicoyip.dev@gmail.com \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox