From: "Asbjørn Sloth Tønnesen" <ast@fiberby.net>
To: Kuniyuki Iwashima <kuniyu@google.com>
Cc: Eric Dumazet <edumazet@google.com>,
Neal Cardwell <ncardwell@google.com>,
"David S. Miller" <davem@davemloft.net>,
Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
Simon Horman <horms@kernel.org>,
Al Viro <viro@zeniv.linux.org.uk>,
netdev@vger.kernel.org, linux-kernel@vger.kernel.org,
Kristian Nielsen <knielsen@knielsen-hq.org>,
stable@vger.kernel.org
Subject: Re: [PATCH net v2] tcp: reset late connection after listening socket close
Date: Tue, 11 Aug 2026 07:44:11 +0000 [thread overview]
Message-ID: <54866197-9523-41ca-b596-37f0c05b9e7b@fiberby.net> (raw)
In-Reply-To: <CAAVpQUDGPtKhv9XfiKgKg7NOqHGzSb4vs7ApQ=rf-RzbaCDQWA@mail.gmail.com>
On 8/11/26 6:37 AM, Kuniyuki Iwashima wrote:
> On Mon, Aug 10, 2026 at 1:57 PM Asbjørn Sloth Tønnesen <ast@fiberby.net> wrote:
>>
>> In commit c82199061009 ("task_work: remove fifo ordering guarantee")
>> Eric removed the ordering guarantee, thereby changing it from a
>> guaranteed FIFO to currently LIFO ordering, in an effort to reduce
>> jitter.
>>
>> This significantly increases the probability for a race to occur
>> between a TCP handshake and the closing of the listening socket.
>
> I looked into the repro, and the description above is too specific
> to your reproducer working on localhost.
>
> In your repro, a server creates a listening socket, accept()s
> some child connections to get multiple file descriptors, and
> exit()s the process, which close()s file descriptors from listener
> (smaller fd) to child fds.
>
> The last fput() queues each fd to the delayed work, which is
> currently executed in LIFO ordering.
>
> The child fd sends RST to the client and later the listener is
> TCP_CLOSE'd.
>
> On the client side, it tries to detect the server's exit() by checking
> RST on the established connections (child fds on listener), and then
> creates a new connection, which creates a reqsk with the alive listener
> but the last ACK will hit TCP_CLOSE'd listener.
>
> I checked the last ACK generated by the repro fails at
> __inet_inherit_port() OR inet_csk_reqsk_queue_add().
>
> That's why reverting c82199061009 prevents reproduction because
> it's too late to create a new connection when the client notices the
> listener's exit() via RST of child connections; new connection's SYN
> always gets RST with the revert.
>
> So, I would drop the description above and simply explain it like
>
> "When __inet_inherit_port() returns -ENOENT, the new connection
> is dropped silently."
>
>
>> In that case the client sees the connection as ESTABLISHED, however in
>> tcp_v{4,6}_syn_recv_sock() the call to __inet_inherit_port() returns
>> -ENOENT, and the new connection is dropped silently by put_and_exit.
>>
>> A client may therefore hang indefinitely on a blocking read() if the
>> used data communication protocol is initiated by the server, like SMTP
>> and the reporter[1]'s MariaDB protocol both are.
>>
>> Had the new connection been processed before the listening socket was
>> closed, it would have been in the accept queue,
>
> or inet_csk_reqsk_queue_add() should have sent RST.
>
>> and been notified when
>> inet_csk_listen_stop() was run, and the client would have got a reset.
>>
>> The call to __inet_inherit_port() returns -ENOENT because
>> inet_csk(sk)->icsk_bind_hash is NULL, after inet_put_port() has been
>> called by tcp_set_state(sk, TCP_CLOSE).
>>
>> This patch adds a check on the return value of the __inet_inherit_port()
>> call, and jumps to a new label, where it resets the new connection,
>> before proceeding with the put_and_exit label.
>>
>> The blamed commit was identified by testing on ancient Debian stable
>
> This commit should be more appropriate since __inet_inherit_port()
> did not fail (except for OOM) before this commit:
>
> Fixes: c2f34a65a61c ("tcp/dccp: fix potential NULL deref in
> __inet_inherit_port()")
Thank you, for you detailed analysis. I didn't find this because I didn't
notice the error condition changing over time, and this being a v4.4 commit
while I was looking for earlier commits.
>> releases to get a rough scope of where to look, then identifying in
>> which release between v3.16 and v4.19 it broke, and finally bisecting
>> v4.2..v4.3 on a fresh Debian then-stable (jessie) VM with tooling from
>> that era, and confirmed by reverting it from v4.3, v5.10.y and net.
>>
>> Additionally the race can be reproduced back to v3.6, by backporting
>> the blamed commit. Beyond v3.6 there are too many conflicts.
>>
>> Reproducer:
>> https://files.fiberby.net/ast/2026/kernel/socket_teardown_test.c
>>
>> Reported-by: Kristian Nielsen <knielsen@knielsen-hq.org>
>> Link: https://lore.kernel.org/87sf0ldk41.fsf@urd.knielsen-hq.org # [1]
>> Fixes: c82199061009 ("task_work: remove fifo ordering guarantee")
>> Cc: <stable@vger.kernel.org>
>> Signed-off-by: Asbjørn Sloth Tønnesen <ast@fiberby.net>
>> ---
>>
>> While the blamed commit properly only made the race observable, then
>> the commit that made the race possible is less realistic to track down,
>> and I don't think is worth the time to continue to the search for it
>> somewhere before v3.6.
>>
>> I'm not submitting a selftest at this time, as I have only found an
>> efficient way to often detect the issue, not disprove it, and I would
>> need more test data from different systems, before I can reliably
>> disprove it with a low runtime budget, without getting false negatives.
>>
>> Changelog:
>> v2:
>> - Use return from __inet_inherit_port() to trigger send_reply()
>> - Use req->rsk_ops->send_reset.
>> - Clarity commit message, and update to reflect the changes.
>> (Thanks Kuniyuki)
>> v1: https://lore.kernel.org/20260807194513.1263310-1-ast@fiberby.net
>>
>> net/ipv4/tcp_ipv4.c | 8 +++++++-
>> net/ipv6/tcp_ipv6.c | 8 +++++++-
>> 2 files changed, 14 insertions(+), 2 deletions(-)
>>
>> diff --git a/net/ipv4/tcp_ipv4.c b/net/ipv4/tcp_ipv4.c
>> index b8887cdd66c5..e0fba4579a57 100644
>> --- a/net/ipv4/tcp_ipv4.c
>> +++ b/net/ipv4/tcp_ipv4.c
>> @@ -1690,6 +1690,7 @@ struct sock *tcp_v4_syn_recv_sock(const struct sock *sk, struct sk_buff *skb,
>> int l3index;
>> #endif
>> struct ip_options_rcu *inet_opt;
>> + int ret;
>>
>> if (sk_acceptq_is_full(sk))
>> goto exit_overflow;
>> @@ -1756,7 +1757,10 @@ struct sock *tcp_v4_syn_recv_sock(const struct sock *sk, struct sk_buff *skb,
>> goto put_and_exit; /* OOM, release back memory */
>> #endif
>>
>> - if (__inet_inherit_port(sk, newsk) < 0)
>> + ret = __inet_inherit_port(sk, newsk);
>> + if (ret == -ENOENT)
>> + goto send_reset_and_exit;
>> + else if (ret < 0)
>> goto put_and_exit;
>
> nit: Please move the -ENOENT branch inside (ret < 0)
>
> if (unlikely(ret < 0) {
> if (ret == -ENOENT)
> goto send_reset_and_exit;
> goto put_and_exit;
> }
>
>
>> *own_req = inet_ehash_nolisten(newsk, req_to_sk(req_unhash),
>> &found_dup_sk);
>> @@ -1784,6 +1788,8 @@ struct sock *tcp_v4_syn_recv_sock(const struct sock *sk, struct sk_buff *skb,
>> exit:
>> tcp_listendrop(sk);
>> return NULL;
>> +send_reset_and_exit:
>> + req->rsk_ops->send_reset(newsk, skb, SK_RST_REASON_TCP_STATE);
>
> Sorry, I think this can be direct call, tcp_v4_send_reset().
>
> Also please pass the listener instead of newsk, it's more
> like a call in tcp_check_req().
>
> Same remarks for IPv6.
Thank you for your very detailed and insightful review.
I have v3 ready, and will post it when the 24 hours are up.
--
pw-bot: cr
prev parent reply other threads:[~2026-08-11 7:45 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-10 20:56 [PATCH net v2] tcp: reset late connection after listening socket close Asbjørn Sloth Tønnesen
2026-08-11 6:37 ` Kuniyuki Iwashima
2026-08-11 7:44 ` Asbjørn Sloth Tønnesen [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=54866197-9523-41ca-b596-37f0c05b9e7b@fiberby.net \
--to=ast@fiberby.net \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=knielsen@knielsen-hq.org \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=ncardwell@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=stable@vger.kernel.org \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).