Netdev List
 help / color / mirror / Atom feed
From: Yuqi Xu <xuyuqiabc@gmail.com>
To: netdev@vger.kernel.org
Cc: Jon Maloy <jmaloy@redhat.com>,
	Tung Quang Nguyen <tung.quang.nguyen@est.tech>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>,
	Ying Xue <ying.xue@windriver.com>,
	Parthasarathy Bhuvaragan <parthasarathy.bhuvaragan@ericsson.com>,
	Kuniyuki Iwashima <kuniyu@google.com>,
	stable@vger.kernel.org, Vega <vega@nebusec.ai>,
	Ren Wei <weir@nebusec.ai>,
	xuyq21@lenovo.com
Subject: Re: [PATCH net 0/2] tipc: fix connection lifetime during netns teardown
Date: Sun, 20 Sep 2026 15:34:40 +0800	[thread overview]
Message-ID: <20260920073440.64223-1-xuyuqiabc@gmail.com> (raw)
In-Reply-To: <cover.1789722780.git.xuyuqiabc@gmail.com>

Hi all,

Thanks for the review. We verified each point against the tree with
1/2 and 2/2 applied (net/tipc/topsrv.c line numbers below are from
that tree).

=== sashiko: [patch 1/2] infinite spin / softlockup ===

> Infinite spin loop and softlockup in tipc_topsrv_stop() when waiting
> for connections with pending work items to close: loop
> `for (id = 0; srv->idr_in_use; id++) { con = idr_find(...) ... }`
> holds `spin_lock_bh(&srv->idr_lock)`; when `con == NULL` the lock is
> not dropped before continuing -> spins ~2^32 times.
> locations: net/tipc/topsrv.c:713 tipc_topsrv_stop;
> net/tipc/topsrv.c:133 tipc_conn_kref_release.

Correct, and this is precisely the failure this series fixes. In the
pre-patch code the walk is:

	for (id = 0; srv->idr_in_use; id++) {
		con = idr_find(&srv->conn_idr, id);
		if (con) {
			conn_get(con);
			spin_unlock_bh(&srv->idr_lock);
			tipc_conn_close(con);
			conn_put(con);
			spin_lock_bh(&srv->idr_lock);
		}
	}

When `idr_find()` returns NULL the loop keeps incrementing `id` while
still holding `idr_lock`. Any connection whose last reference is being
dropped in tipc_conn_kref_release() then blocks forever on
spin_lock_bh(&s->idr_lock), so `srv->idr_in_use` never reaches zero
and the CPU spins under the lock. That is the
`__radix_tree_lookup -> tipc_topsrv_exit_net` RCU stall in the crash
log of the cover letter.

2/2 rewrites exactly this walk:

	for (id = 0; srv->idr_in_use;) {
		con = idr_get_next(&srv->conn_idr, &id);
		if (!con || !kref_get_unless_zero(&con->kref)) {
			spin_unlock_bh(&srv->idr_lock);
			cond_resched();
			spin_lock_bh(&srv->idr_lock);
			id = 0;
			continue;
		}
		id++;
		spin_unlock_bh(&srv->idr_lock);
		tipc_conn_close(con);
		conn_put(con);
		spin_lock_bh(&srv->idr_lock);
	}

i.e. it uses idr_get_next() and, whenever no entry can be taken, drops
idr_lock, reschedules and retries, so a pending
tipc_conn_kref_release() can always make progress and the loop cannot
spin under the lock. So the finding is correct for 1/2 standing alone
but is fully addressed by 2/2; no further change is needed for this
point.

=== sashiko: [patch 2/2] sock_release() in atomic context ===

> `sock_release()` called in atomic/softirq context:
> `tipc_sub_timeout()` (timer softirq) -> `tipc_topsrv_queue_evt()`
> -> `conn_put()` -> `tipc_conn_kref_release()` -> `sock_release()`
> sleeps (`lock_sock()`), sleep-in-atomic.
> locations: net/tipc/subscr.c:110, net/tipc/topsrv.c:322,
> net/tipc/topsrv.c:120.

The chain is not quite as drawn. tipc_conn_kref_release() does not
call sock_release() unconditionally: it already has `if (con->sock)`
(line 134), so in-kernel connections never take that path.

For socket-backed connections the remaining sleep-in-atomic concern
is real but narrower than the chain suggests, and this series does
not change it:

- tipc_topsrv_queue_evt() only calls conn_put() synchronously on its
  error path (line 340): the connection is no longer connected after
  the lookup, kmalloc fails, or queue_work() returns false because
  ->swork is already queued. On the normal path the lookup reference
  is handed to tipc_conn_send_work(), which runs in process context
  and does the final conn_put() there.
- So the remaining sock_release() in timer softirq only happens when
  that specific conn_put() drops the last reference of a socket-backed
  connection, i.e. when the connection is being torn down concurrently.
- Neither 1/2 nor 2/2 touch subscr.c, tipc_topsrv_queue_evt() or
  tipc_conn_kref_release(). 1/2 only detaches the listener and cancels
  srv->awork; it does not touch the subscription timer path.

So this is pre-existing and orthogonal to the two patches, not for
this series. A separate follow-up could avoid dropping the last
reference for a socket-backed connection from softirq; that would be
a separate patch.

=== sashiko: [patch 2/2] NULL deref in tipc_conn_close() ===

> Unconditional `con->sock->sk` deref in `tipc_conn_close()` panics for
> in-kernel subscriptions (`tipc_topsrv_kern_subscr()` passes NULL sock
> -> `con->sock == NULL`); `tipc_conn_kref_release()` checks
> `if (con->sock)` but `tipc_conn_close()` does not.
> locations: net/tipc/topsrv.c:158 tipc_conn_close,
> net/tipc/topsrv.c:724 tipc_topsrv_stop.

This one is a valid latent bug, and we agree with the asymmetry:
tipc_conn_close() does `struct sock *sk = con->sock->sk;` (line 158)
unconditionally, while tipc_conn_kref_release() guards its
sock_release() with `if (con->sock)` (line 134).
tipc_topsrv_kern_subscr() really does allocate with sock == NULL
(line 587) and such a connection is inserted into conn_idr, so if one
is still there when tipc_topsrv_stop() walks the idr, 2/2's loop calls
tipc_conn_close() on a connection whose con->sock is NULL.

About reachability in the netns teardown path: the in-kernel
subscriber is created from tipc_group_create() (net/tipc/group.c:190),
and it is normally removed synchronously by tipc_group_delete() ->
tipc_topsrv_kern_unsubscr() when the owning socket is released
(tipc_release() -> tipc_sk_leave()). Since the socket holds a net
reference, netns teardown cannot overtake that. The NULL deref
therefore needs a kernel connection that is still in conn_idr at stop
time, e.g. when tipc_topsrv_kern_unsubscr()'s two conn_put() calls do
not drop the last reference because a pending con->swork still holds
one. That is a narrow race, not the common path.

It is also pre-existing: the pre-patch loop already called
tipc_conn_close() on every entry found in conn_idr, including
sock == NULL kernel connections, so 1/2 + 2/2 do not introduce it.
(kref_get_unless_zero() in 2/2 only narrows the
already-dropped-reference window; it does not dereference con->sock.)

Orthogonal to this series. A separate follow-up could mirror
tipc_conn_kref_release()'s check in tipc_conn_close(), i.e. skip the
con->sock->sk access when con->sock == NULL; that would be a
separate patch.

Best regards,
Yuqi Xu

      parent reply	other threads:[~2026-09-20  7:34 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-18  9:21 [PATCH net 0/2] tipc: fix connection lifetime during netns teardown Yuqi Xu
2026-09-18  9:21 ` [PATCH net 1/2] tipc: stop the listener before draining connections Yuqi Xu
2026-09-21  2:55   ` Tung Quang Nguyen
2026-09-18  9:21 ` [PATCH net 2/2] tipc: make conn_idr teardown safe Yuqi Xu
2026-09-21  3:01   ` Tung Quang Nguyen
2026-09-20  7:34 ` Yuqi Xu [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260920073440.64223-1-xuyuqiabc@gmail.com \
    --to=xuyuqiabc@gmail.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=jmaloy@redhat.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=parthasarathy.bhuvaragan@ericsson.com \
    --cc=stable@vger.kernel.org \
    --cc=tung.quang.nguyen@est.tech \
    --cc=vega@nebusec.ai \
    --cc=weir@nebusec.ai \
    --cc=xuyq21@lenovo.com \
    --cc=ying.xue@windriver.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox