From: Mahanta Jambigi <mjambigi@linux.ibm.com>
To: netdev-bot+sashiko@kernel.org
Cc: andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
kuba@kernel.org, pabeni@redhat.com, alibuda@linux.alibaba.com,
dust.li@linux.alibaba.com, sidraya@linux.ibm.com,
hidayath@linux.ibm.com, pasic@linux.ibm.com, horms@kernel.org,
tonylu@linux.alibaba.com, guwen@linux.alibaba.com,
stable@vger.kernel.org, netdev@vger.kernel.org,
linux-s390@vger.kernel.org, linux-rdma@vger.kernel.org
Subject: Re: [PATCH net v7] net/smc: fix lgr/lnk lifetime vs diag reader race
Date: Mon, 5 Oct 2026 18:40:16 +0530 [thread overview]
Message-ID: <46e5b6b2-a0ee-4381-816c-6435c61e320f@linux.ibm.com> (raw)
In-Reply-To: <179110012456.434549.13805687630302688897@kernel.org>
On 04/10/26 1:18 pm, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 2 potential
> issue(s) to consider.
>
> Critical: 0 · High: 1 · Medium: 1 · Low: 0
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [High] The patch does not fully close the diag-reader vs. link-lifetime
> race it says it fixes.
> - [Medium] The new safety gates in __smc_diag_dump() depend on a plain,
> lockless load of sk->sk_state.
>
> --
>
> net/smc: fix lgr/lnk lifetime vs diag reader race
>
> This patch tries to close races between the SMC sock_diag dump and
> teardown paths. The dump walks the socket hash under read_lock, while the
> teardown paths free lgr, lnk or sndbuf_desc state. The patch adds sk_state
> based guards for SMC_INIT and SMC_CLOSED sockets to __smc_diag_dump(), and
> unhashes the socket earlier in smc_conn_kill(), smc_close_active_abort()
> and smc_close_passive_work().
>
>> diff --git a/net/smc/smc_diag.c b/net/smc/smc_diag.c
>> index bf0beaa23bdb6..b29bd34acf7b5 100644
>> --- a/net/smc/smc_diag.c
>> +++ b/net/smc/smc_diag.c
>> @@ -90,7 +90,8 @@ static int __smc_diag_dump(struct sock *sk, struct sk_buff *skb,
>> r->diag_state = sk->sk_state;
>> if (smc->use_fallback)
>> r->diag_mode = SMC_DIAG_MODE_FALLBACK_TCP;
>> - else if (smc_conn_lgr_valid(&smc->conn) && smc->conn.lgr->is_smcd)
>> + else if (sk->sk_state != SMC_INIT &&
>> + smc_conn_lgr_valid(&smc->conn) && smc->conn.lgr->is_smcd)
>> r->diag_mode = SMC_DIAG_MODE_SMCD;
>> else
>> r->diag_mode = SMC_DIAG_MODE_SMCR;
>
> [Severity: Medium]
> Is a plain load of sk->sk_state enough to order the later loads of
> conn->alert_token_local, conn->lgr->is_smcd and conn->sndbuf_desc->len?
>
> This check, and the SMC_INIT/SMC_CLOSED guard added below, share no lock
> with the writer. Nothing sits between the sk_state load and the conn field
> loads: no READ_ONCE(), no smp_load_acquire(), no smp_rmb(). Only a control
> dependency links them, and that doesn't order load->load on weakly ordered
> CPUs.
>
> On the early abort path the writer runs under lock_sock():
>
> smc_conn_abort()
> smc_conn_free()
> smc_lgr_unregister_conn() /* alert_token_local = 0 */
> smc_lgr_cleanup_early()
> __smc_lgr_terminate()
> smc_lgr_free() /* lgr and buffers freed, conn->lgr and
> conn->sndbuf_desc left set */
> smc_switch_to_fallback()
> ...
> sk->sk_state = SMC_ACTIVE
>
> On arm64, could the reader see sk_state == SMC_ACTIVE together with a
> stale non-zero alert_token_local? It would then read lgr->is_smcd or
> sndbuf_desc->len from freed memory and copy the value into the netlink
> reply.
>
> This can't happen on x86 or s390, which keep load-load order, and the
> writer's window is long. The Linux kernel memory model still allows it.
>
> Would READ_ONCE() or smp_load_acquire() on sk_state (or an smp_rmb()
> before the conn field loads), paired with release-ordered state changes,
> be needed here?
Agreed. Under the LKMM, a control dependency does not order load-load on
weakly ordered architectures (such as ARM64).
In v8, I will address this with acquire-release ordering:
- In __smc_diag_dump(), I use smp_load_acquire(&sk->sk_state) and reuse
the acquired state (in r->diag_state) across the subsequent mode/state
checks.
- In af_smc.c, I paired it with smp_store_release(&sk->sk_state,
SMC_ACTIVE) when transitioning out of SMC_INIT.
>
>> @@ -103,6 +104,9 @@ static int __smc_diag_dump(struct sock *sk, struct sk_buff *skb,
>> if (nla_put(skb, SMC_DIAG_FALLBACK, sizeof(fallback), &fallback) < 0)
>> goto errout;
>>
>> + if (sk->sk_state == SMC_INIT || sk->sk_state == SMC_CLOSED)
>> + goto out;
>> +
>
> [Severity: High]
> Does this leave the conn->lnk replacement during SMC-R link failover still
> racing with the diag reader?
>
> For SMC_ACTIVE and closing-state sockets, the LGRINFO block further down
> still does this with only read_lock(&smc_hash->lock) held:
>
> net/smc/smc_diag.c:__smc_diag_dump() {
> ...
> struct smc_link *link = smc->conn.lnk;
> ...
> memcpy(linfo.lnk[0].ibname, link->smcibdev->ibdev->name,
> sizeof(link->smcibdev->ibdev->name));
> ...
> }
>
> The failover side doesn't take the hash lock and doesn't unhash the
> socket. For example, when the peer sends DELETE_LINK:
>
> smc_llc_process_cli_delete_link()
> smc_switch_conns(lgr, lnk_del, false)
> smc_switch_link_and_count(conn, to_lnk)
> smcr_link_put(conn->lnk);
> conn->lnk = to_lnk;
> smcr_link_clear(lnk_del, true)
> ...
> __smcr_link_clear()
> put_device(&lnk->smcibdev->ibdev->dev);
> smcibdev = lnk->smcibdev;
> memset(lnk, 0, sizeof(struct smc_link));
> ...
> atomic_dec_return(&smcibdev->lnk_cnt)
>
> A local RNIC or port going down reaches smc_switch_conns() the same way,
> through smcr_link_down().
>
> smc_switch_link_and_count() runs only under lgr->conns_lock and
> conn->send_lock. The diag reader takes neither lock.
>
> Suppose the reader loaded the old conn->lnk before the switch and was
> delayed before the later loads. Could it then dereference a NULL
> link->smcibdev after the memset()? Could it instead read a freed smcibdev
> or ibdev and copy those bytes to userspace?
>
> An unprivileged SOCK_DIAG dump that asks for SMC_DIAG_LGRINFO (for
> example smcss -R) can reach this.
>
> The socket stays hashed in SMC_ACTIVE or a closing state. So neither this
> guard nor the !list_empty(&smc->conn.lgr->list) check from 9dbe086c69b8
> filters it out
Good catch. When smcr_link_clear() drops the final base refcount after
smc_switch_conns(), __smcr_link_clear() can execute memset() and
put_device() while a concurrent diag reader holds a pointer to the old link.
To eliminate the pointer dereference and external device lifetime
dependency entirely, in v8 I switched from chasing
link->smcibdev->ibdev->name to copying directly from the cached
link->ibname (which is embedded in struct smc_link and initialized once
at link creation). Since struct smc_link is embedded in the link group
(pinned while the socket is hashed), this avoids the NULL dereference
and UAF on smcibdev/ibdev without introducing extra locking on the
critical TX path.
I'll post v8 shortly with these 2 changes.
pw-bot: cr
prev parent reply other threads:[~2026-10-05 13:11 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 7:30 [PATCH net v7] net/smc: fix lgr/lnk lifetime vs diag reader race Mahanta Jambigi
2026-09-30 7:33 ` netdev-bot+sinfo
2026-09-30 8:32 ` Mahanta Jambigi
2026-10-01 6:24 ` Hidayath Khan
2026-10-01 7:55 ` Sidraya Jayagond
2026-10-04 7:48 ` netdev-bot+sashiko
2026-10-05 13:10 ` Mahanta Jambigi [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=46e5b6b2-a0ee-4381-816c-6435c61e320f@linux.ibm.com \
--to=mjambigi@linux.ibm.com \
--cc=alibuda@linux.alibaba.com \
--cc=andrew+netdev@lunn.ch \
--cc=davem@davemloft.net \
--cc=dust.li@linux.alibaba.com \
--cc=edumazet@google.com \
--cc=guwen@linux.alibaba.com \
--cc=hidayath@linux.ibm.com \
--cc=horms@kernel.org \
--cc=kuba@kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=linux-s390@vger.kernel.org \
--cc=netdev-bot+sashiko@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=pasic@linux.ibm.com \
--cc=sidraya@linux.ibm.com \
--cc=stable@vger.kernel.org \
--cc=tonylu@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox