Linux RDMA and InfiniBand development
 help / color / mirror / Atom feed
From: netdev-bot+sashiko@kernel.org
To: achender@kernel.org
Cc: netdev@vger.kernel.org, linux-rdma@vger.kernel.org,
	pabeni@redhat.com, edumazet@google.com, kuba@kernel.org,
	horms@kernel.org, nicoyip.dev@gmail.com
Subject: Re: [PATCH net-next v3 03/13] net/rds: guard every work-requeueing site with rds_destroy_pending()
Date: Thu, 17 Sep 2026 09:38:39 +0000	[thread overview]
Message-ID: <178963791995.22033.5627133034354921592@kernel.org> (raw)
In-Reply-To: <20260914033719.138057-4-achender@kernel.org>

Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider:

Critical: 0 · High: 0 · Medium: 1 · Low: 0

- [Medium] Incomplete coverage vs. the patch's own claim: the subject
  says "guard every work-requeueing site with rds_destroy_pending()",…

--

net/rds: guard every work-requeueing site with rds_destroy_pending()

This adds the rcu_read_lock() + rds_destroy_pending() guard around five
path-work arming sites: the two cp_send_w re-arms in ib_send.c, the
cp_recv_w re-arm in rds_ib_recv_refill(), the cp_recv_w kick in
rds_tcp_accept_one(), and the cp_conn_w arm in rds_sendmsg().

> net/rds: guard every work-requeueing site with rds_destroy_pending()

[Severity: Medium]
Is "every work-requeueing site" accurate here?  After this patch four
arming sites in net/rds/threads.c still queue path works with no
rcu_read_lock() and no rds_destroy_pending() test.

rds_send_worker():

	switch (ret) {
	case -EAGAIN:
		rds_stats_inc(s_send_immediate_retry);
		queue_delayed_work(cp->cp_wq, &cp->cp_send_w, 0);
		break;
	case -ENOMEM:
		rds_stats_inc(s_send_delayed_retry);
		queue_delayed_work(cp->cp_wq, &cp->cp_send_w, 2);
		break;

rds_recv_worker():

	switch (ret) {
	case -EAGAIN:
		rds_stats_inc(s_recv_immediate_retry);
		queue_delayed_work(cp->cp_wq, &cp->cp_recv_w, 0);
		break;
	case -ENOMEM:
		rds_stats_inc(s_recv_delayed_retry);
		queue_delayed_work(cp->cp_wq, &cp->cp_recv_w, 2);
		break;

These four look functionally safe to me, since they are self-requeues
issued from inside the running work item and rds_conn_path_destroy()
disables the work before flushing it:

net/rds/connection.c:rds_conn_path_destroy() {
	/* make sure lingering queued work won't try to ref the conn */
	cancel_delayed_work_sync(&cp->cp_send_w);
	cancel_delayed_work_sync(&cp->cp_recv_w);
	...
}

__cancel_work_sync() grabs the pending bit (deleting any armed
delayed-work timer) and calls work_offqd_disable() before flushing, and
queue_delayed_work_on() then drops the queue attempt via
clear_pending_if_disabled(), so a requeue from the callback that is
being flushed is discarded and the callback cannot run again afterwards.

The following patch in the series also adds a comment on
struct rds_connection::c_destroy_in_prog that states the rule with no
exception:

	 * Set once, by rds_conn_destroy(), before it cancels the path
	 * works; read through rds_destroy_pending().  A site that arms
	 * a path work must test the predicate and queue the work inside
	 * one rcu_read_lock() section: ...

Could either the subject and changelog note the self-requeue exemption
(and the cancel_delayed_work_sync() reason it holds), or the four
threads.c sites be converted too, so the documented invariant and the
code agree?

Also, the changelog argues for the IB sites:

>     The IB completion sites are reachable from soft-irq at any point
>     before the QP is drained, so a completion landing in the window
>     between the cancel and destroy_workqueue() in rds_conn_path_destroy()
>     re-arms a work on a workqueue that is about to be destroyed: with
>     delay 0 the work is queued directly on the freed workqueue, and with
>     delay 1 the timer survives destroy_workqueue() unseen and fires
>     afterwards, queueing from a timer_list that lives in the freed c_path
>     array.

The same delay-2 timer shape appears in the threads.c -ENOMEM cases, so
a reader may conclude those are equally exposed.  Would it help to say
explicitly why the threads.c requeues are not in the same category?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260914033719.138057-1-achender%40kernel.org

  parent reply	other threads:[~2026-09-17  9:38 UTC|newest]

Thread overview: 39+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-14  3:37 [PATCH net-next v3 00/13] net/rds: make connection lifetime reference-counted Allison Henderson
2026-09-14  3:37 ` [PATCH net-next v3 01/13] net/rds: ib: don't enable interrupts in rds_ib_conn_free() Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-14  3:37 ` [PATCH net-next v3 02/13] net/rds: free every path's transport data on the passive create paths Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 03/13] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko [this message]
2026-09-14  3:37 ` [PATCH net-next v3 04/13] net/rds: make rds_destroy_pending() cover single-connection destroy Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 05/13] net/rds: split connection destroy into quiesce and kref-governed free Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 06/13] net/rds: wait for connections to be freed on transport unload Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 07/13] net/rds: unlink transport nodes before a possibly deferred connection free Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 08/13] net/rds: hold connection references in lookup, sockets and c_passive Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 09/13] net/rds: refuse to queue a message on a connection being destroyed Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 10/13] net/rds: tcp: don't attach an accepted socket to " Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 11/13] net/rds: pin the connection across RDMA-CM event handling Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 12/13] net/rds: drop rds_conn_count in favor of t_conn_count Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko
2026-09-14  3:37 ` [PATCH net-next v3 13/13] net/rds: hold a connection reference from struct rds_incoming Allison Henderson
2026-09-15  3:37   ` sashiko-bot
2026-09-17  9:38   ` netdev-bot+sashiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=178963791995.22033.5627133034354921592@kernel.org \
    --to=netdev-bot+sashiko@kernel.org \
    --cc=achender@kernel.org \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=nicoyip.dev@gmail.com \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox