From: gang.yan@linux.dev
To: "Runyu Xiao" <runyu.xiao@seu.edu.cn>,
"Matthieu Baerts" <matttbe@kernel.org>,
"Mat Martineau" <martineau@kernel.org>
Cc: "Geliang Tang" <geliang@kernel.org>,
"David S . Miller" <davem@davemloft.net>,
"Eric Dumazet" <edumazet@google.com>,
"Jakub Kicinski" <kuba@kernel.org>,
"Paolo Abeni" <pabeni@redhat.com>,
"Simon Horman" <horms@kernel.org>,
mptcp@lists.linux.dev, netdev@vger.kernel.org,
linux-kernel@vger.kernel.org, runyu.xiao@seu.edu.cn,
jianhao.xu@seu.edu.cn, stable@vger.kernel.org
Subject: Re: [PATCH net] mptcp: upgrade network refcount before socket lock
Date: Mon, 10 Aug 2026 01:59:28 +0000 [thread overview]
Message-ID: <4e57509daf45aa8968140e7a7279ad0699249bf6@linux.dev> (raw)
In-Reply-To: <20260809091949.3618191-1-runyu.xiao@seu.edu.cn>
August 9, 2026 at 5:19 PM, "Runyu Xiao" <runyu.xiao@seu.edu.cn mailto:runyu.xiao@seu.edu.cn?to=%22Runyu%20Xiao%22%20%3Crunyu.xiao%40seu.edu.cn%3E > wrote:
Hi,
Thanks for your patch.
>
> sk_net_refcnt_upgrade() calls get_net_track() with GFP_KERNEL and can enter
> direct reclaim. Calling it while holding the newly created subflow socket
> lock can create a reclaim-to-socket-lock dependency cycle.
>
> Upgrade the network reference before taking the socket lock. The socket is
> newly created and has not been exposed to other code at this point, so the
> fields changed by sk_net_refcnt_upgrade() are not accessed concurrently.
> The error path still releases the socket normally after the upgrade.
>
> The PatchProof static-analysis tool detected a GFP_KERNEL allocation while
> the socket lock is held. Manual source review of v7.1.5 and current
> mainline confirmed the lock and allocation ordering.
>
> A source-level check found `sk_net_refcnt_upgrade()` after
> `lock_sock_nested()` in the original function and before it after this
> change. A POSIX-thread lock-order model made the reclaim lock unavailable
> while the socket lock was held, observed `EBUSY` for the reclaim lock, and
> then completed with the reclaim-first order. The model checks the ordering
> invariant only; it does not execute the kernel MPTCP path. No live lockdep
> MPTCP test or reclaim fault injection was run.
>
Maybe the commit message seems too long. Could you please send a v2 with a more
concise commit message? For your reference, here is a suggested simplified version:
'''
sk_net_refcnt_upgrade() performs a GFP_KERNEL allocation (via get_net_track()
→ ref_tracker_alloc()), which can enter direct reclaim and establish a
socket_lock → fs_reclaim dependency. Move it before lock_sock_nested(), mirroring
the convention documented in net/rds/tcp.c:rds_tcp_tune().
The socket is freshly created and unpublished at this point, so sk_net_refcnt/
ns_tracker are not touched by any concurrent path; the error path still releases
via sock_release() which handles both refcounted and non-refcounted trackers.
'''
Please note that this suggested text is generated by AI, so please review it
carefully before adopting it. Also, the next patch can only be sent to mptcp@lists.linux.dev,
no need to cc to others.
> Fixes: 1d2f3d3c6268 ("mptcp: adjust to use netns refcount tracker")
> Cc: stable@vger.kernel.org
> Signed-off-by: Runyu Xiao <runyu.xiao@seu.edu.cn>
> ---
> net/mptcp/subflow.c | 11 ++++++-----
> 1 file changed, 6 insertions(+), 5 deletions(-)
>
> diff --git a/net/mptcp/subflow.c b/net/mptcp/subflow.c
> index e1f20ff8fdb4..a9f951cc6a0e 100644
> --- a/net/mptcp/subflow.c
> +++ b/net/mptcp/subflow.c
> @@ -1786,6 +1786,12 @@ int mptcp_subflow_create_socket(struct sock *sk, unsigned short family,
> if (err)
> return err;
>
> + /* kernel sockets do not by default acquire net ref, but TCP timer
> + * needs it.
> + * Update ns_tracker to current stack trace and refcounted tracker.
> + */
> + sk_net_refcnt_upgrade(sf->sk);
> +
> lock_sock_nested(sf->sk, SINGLE_DEPTH_NESTING);
>
> err = security_mptcp_add_subflow(sk, sf->sk);
> @@ -1795,11 +1801,6 @@ int mptcp_subflow_create_socket(struct sock *sk, unsigned short family,
> /* the newly created socket has to be in the same cgroup as its parent */
> mptcp_attach_cgroup(sk, sf->sk);
>
> - /* kernel sockets do not by default acquire net ref, but TCP timer
> - * needs it.
> - * Update ns_tracker to current stack trace and refcounted tracker.
> - */
> - sk_net_refcnt_upgrade(sf->sk);
> err = tcp_set_ulp(sf->sk, "mptcp");
> if (err)
> goto err_free;
LKGM!
You can add 'Acked-by: Gang Yan <gang.yan@linux.dev>' in your next patch.
Thanks
Gang
> --
> 2.34.1
>
next prev parent reply other threads:[~2026-08-10 1:59 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-09 9:19 [PATCH net] mptcp: upgrade network refcount before socket lock Runyu Xiao
2026-08-10 1:59 ` gang.yan [this message]
2026-08-10 10:01 ` Matthieu Baerts
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4e57509daf45aa8968140e7a7279ad0699249bf6@linux.dev \
--to=gang.yan@linux.dev \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=geliang@kernel.org \
--cc=horms@kernel.org \
--cc=jianhao.xu@seu.edu.cn \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=martineau@kernel.org \
--cc=matttbe@kernel.org \
--cc=mptcp@lists.linux.dev \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=runyu.xiao@seu.edu.cn \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox