Netdev List
 help / color / mirror / Atom feed
From: netdev-bot+sashiko@kernel.org
To: achender@kernel.org
Cc: netdev@vger.kernel.org, linux-rdma@vger.kernel.org,
	pabeni@redhat.com, edumazet@google.com, kuba@kernel.org,
	horms@kernel.org
Subject: Re: [PATCH net-next v8 03/13] net/rds: ib: refuse to attach a connection to a device being removed
Date: Sun, 04 Oct 2026 16:34:58 +0000	[thread overview]
Message-ID: <179113169845.434549.16777871408089243882@kernel.org> (raw)
In-Reply-To: <20261003163215.250253-4-achender@kernel.org>

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The commit message gets the trigger wrong.

--

net/rds: ib: refuse to attach a connection to a device being removed

This patch adds a shutting_down flag to struct rds_ib_device.
rds_ib_dev_shutdown() sets it under rds_ibdev->spinlock before walking
conn_list. rds_ib_add_conn() checks the flag under the same lock and
returns -ENODEV, which makes rds_ib_setup_qp() fail the connect.

> rds_ib_remove_one() drops every connection on the device's conn_list
> in rds_ib_dev_shutdown(), then clears the client data so that no new
> connect can find the device.  A connect that is already past
> rds_ib_get_client_data() when the walk runs is not covered by either:
> rds_ib_setup_qp() goes on to rds_ib_add_conn(), which moves the
> connection onto the conn_list the walk has just finished with, and
> builds a QP on a device that is on its way out.  Nothing drops that
> connection afterwards - the device's shutdown walk is over, and the
> connection never returns to ib_nodev_conns, which is the only list the
> transport exit sweeps - so it outlives the device, and the module.

[Severity: Low]
Does this description hold for a real IB device hot-removal?

disable_device() in drivers/infiniband/core/device.c removes clients in
LIFO order. That means rds_ib_remove_one() runs before rdma_cm's
cma_remove_one().

cma_process_remove() then sends DEVICE_REMOVAL to every cm_id still on the
device, including the late connection's ic->i_cm_id. It then waits for
those ids to be destroyed:

cma_remove_one()
  cma_process_remove()
    cma_send_device_removal_put(id_priv)
      ...
        rds_rdma_cm_event_handler_cmn()

net/rds/rdma_transport.c:rds_rdma_cm_event_handler_cmn() {
    ...
	case RDMA_CM_EVENT_DEVICE_REMOVAL:
	case RDMA_CM_EVENT_ADDR_CHANGE:
		if (conn)
			rds_conn_drop(conn);
		break;
    ...
}

During a plain hot-removal, rds_destroy_pending() is false, so this
rds_conn_drop() queues the shutdown. The shutdown tears down the QP and
cm_id, and rds_ib_remove_conn() puts the connection back on
ib_nodev_conns. So on hot-removal the connection does not seem to outlive
the device.

The leak described here seems to need the module unload path instead:

rds_ib_exit()
  rds_ib_set_unloading()
  rds_ib_unregister_client()
    rds_ib_remove_one()    <- device stays registered, no DEVICE_REMOVAL
  rds_ib_destroy_nodev_conns()

On this path rds_destroy_pending() is true, so every non-destroy
rds_conn_drop() returns early. rds_ib_destroy_nodev_conns() only sweeps
ib_nodev_conns. A connection attached after the walk is then leaked.

The code change looks correct for both cases. Could the commit message
name the trigger as unloading rds_rdma while a connect is in flight,
rather than device removal in general? The patch carries a Fixes: tag, so
stable backporters will use this description to judge how the bug
reproduces and what it affects.

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261003163215.250253-1-achender%40kernel.org

  reply	other threads:[~2026-10-04 16:34 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-03 16:32 [PATCH net-next v8 00/13] net/rds: make connection lifetime reference-counted Allison Henderson
2026-10-03 16:32 ` [PATCH net-next v8 01/13] net/rds: ib: don't enable interrupts in rds_ib_conn_free() Allison Henderson
2026-10-03 16:32 ` [PATCH net-next v8 02/13] net/rds: undo conn_alloc() the same way on every __rds_conn_create() exit Allison Henderson
2026-10-03 16:32 ` [PATCH net-next v8 03/13] net/rds: ib: refuse to attach a connection to a device being removed Allison Henderson
2026-10-04 16:34   ` netdev-bot+sashiko [this message]
2026-10-03 16:32 ` [PATCH net-next v8 04/13] net/rds: guard every work-requeueing site with rds_destroy_pending() Allison Henderson
2026-10-04 16:34   ` netdev-bot+sashiko
2026-10-03 16:32 ` [PATCH net-next v8 05/13] net/rds: make rds_destroy_pending() report a connection's own destroy Allison Henderson
2026-10-03 16:32 ` [PATCH net-next v8 06/13] net/rds: split connection destroy into quiesce and kref-governed free Allison Henderson
2026-10-03 16:32 ` [PATCH net-next v8 07/13] net/rds: unlink transport nodes before a possibly deferred connection free Allison Henderson
2026-10-04 16:35   ` netdev-bot+sashiko
2026-10-03 16:32 ` [PATCH net-next v8 08/13] net/rds: wait for connections to be freed on transport unload Allison Henderson
2026-10-04 16:35   ` netdev-bot+sashiko
2026-10-03 16:32 ` [PATCH net-next v8 09/13] net/rds: hold a connection reference from struct rds_incoming Allison Henderson
2026-10-04 16:35   ` netdev-bot+sashiko
2026-10-03 16:32 ` [PATCH net-next v8 10/13] net/rds: take cp_lock to purge cp_send_queue in the quiesce Allison Henderson
2026-10-03 16:32 ` [PATCH net-next v8 11/13] net/rds: hold connection references in lookup, sockets and c_passive Allison Henderson
2026-10-04 16:35   ` netdev-bot+sashiko
2026-10-03 16:32 ` [PATCH net-next v8 12/13] net/rds: pin the connection across RDMA-CM event handling Allison Henderson
2026-10-04 16:35   ` netdev-bot+sashiko
2026-10-03 16:32 ` [PATCH net-next v8 13/13] net/rds: drop rds_conn_count in favor of t_conn_count Allison Henderson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=179113169845.434549.16777871408089243882@kernel.org \
    --to=netdev-bot+sashiko@kernel.org \
    --cc=achender@kernel.org \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox