Linux s390 Architecture development
 help / color / mirror / Atom feed
From: Mahanta Jambigi <mjambigi@linux.ibm.com>
To: dust.li@linux.alibaba.com, andrew+netdev@lunn.ch,
	davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, alibuda@linux.alibaba.com,
	sidraya@linux.ibm.com, hidayath@linux.ibm.com
Cc: pasic@linux.ibm.com, horms@kernel.org, tonylu@linux.alibaba.com,
	guwen@linux.alibaba.com, stable@vger.kernel.org,
	netdev@vger.kernel.org, linux-s390@vger.kernel.org,
	linux-rdma@vger.kernel.org
Subject: Re: [PATCH net v2 0/2] net/smc: fix diag dump lifetime races
Date: Tue, 1 Sep 2026 20:43:11 +0530	[thread overview]
Message-ID: <85fa9f02-7516-4595-b5a8-4ae4ca845121@linux.ibm.com> (raw)
In-Reply-To: <apbMyLc0IhLMKwwA@linux.alibaba.com>



On 01/09/26 6:32 pm, Dust Li wrote:
> On 2026-08-31 20:27:38, Mahanta Jambigi wrote:
>>
>>
>> On 31/08/26 7:12 pm, Dust Li wrote:
>>> On 2026-08-28 08:54:37, Mahanta Jambigi wrote:
>>>> This series fixes multiple lifetime races in the SMC diag dump path.
>>>>
>>>> The first patch adds the basic infrastructure needed to synchronize diag readers
>>>> against connection-owned conn->lgr/conn->lnk updates. It introduces a
>>>> per-connection spinlock and uses it in the link switch and connection free
>>>> handoff paths. conn->lgr and conn->lnk are NULLed under the lock before the
>>>> borrowed references are released, so a non-NULL conn->lgr seen under the lock
>>>> guarantees the lgr object is alive. The diag reader can rely on this invariant
>>>> without borrowing any extra reference.
>>>>
>>>> The second patch fixes two races in smc_diag itself:
>>>>
>>>> - serialize clcsock field access against smc_clcsock_release() with
>>>>  mutex_trylock()
>>>> - take conn->lgr_lnk_lock when reading conn->lgr and conn->lnk; use
>>>>  smc_conn_lgr_valid() inside the lock to check that the connection is
>>>>  still registered, then snapshot all required fields and call nla_put()
>>>>  after releasing the lock
>>>
>>> Hi Mahanta,
>>>
>>> As discussed in the other thread, I think we should defer the release of
>>> smc->clcsock and remove clcsock_release_lock.
>>>
>>> In that case, we should no longer need these two patches. Also,
>>> introducing more locks in SMC is the last thing I want to do :)
>>
>> Thanks for the new series "[RFC net-next 0/7] net/smc: tie clcsock
>> lifetime to the smc socket and remove clcsock_release_lock" — once it
>> lands, we can drop the mutex_trylock() fix for Race 1 (clcsock).
>>
>> However, Race 2 remains open. Your series does not touch smc_core.c or
>> smc_cdc.c, so smc_conn_free(), smc_switch_link_and_count(), and
>> smc_cdc_msg_validate() still write conn->lgr/conn->lnk with no
>> synchronization against the diag reader.
> 
> Hi Mahanta,
> 
> Thanks for the detailed explanation. You are right that a per-connection
> spinlock can work here, and I agree none of the lock sites are on the
> per-message hot path. But I think we can also do the same thing we did with
> clcsock_release_lock: instead of adding a lock, tie the lifetime of
> conn->lgr/conn->lnk to the point where the connection stops being observable,
> and remove the need for synchronization altogether.
> 
> For lgr/lnk, that point is the hash table. Once the connection is unhashed, the
> diag dump (which iterates under the hash read_lock) can no longer reach it. So
> if we make sure the connection-owned references are only dropped after unhash,
> the invariant becomes: holding the hash read_lock and seeing a non-NULL
> conn->lgr implies it is safe to dereference. The diag path then reduces to
> `hold hash read_lock -> read conn->lgr -> if non-NULL, use it -> done` with no
> new lock, no extra reference, and no trylock. The invariant is carried by
> object lifetime rather than by a lock, which I find easier to keep correct over
> time.

I looked carefully at the new design and found one remaining gap.

The *unhash* invariant — "any socket in the hash has its
connection-owned lgr/lnk refs held" — protects against smc_conn_free()
dropping refs while the socket is still hashed. However it does not
protect the conn->lnk->smcibdev->ibdev->name access in the
SMC_DIAG_LGRINFO block in smc_diag.c file.

The gap is in *smc_switch_link_and_count*(). It is called under
send_lock (not under any hash-related lock) and calls
smcr_link_put(conn->lnk) on the old link before reassigning conn->lnk.
If that put drops the last reference, __smcr_link_clear() runs
immediately, doing memset(lnk, 0, sizeof(struct smc_link)) which zeroes
lnk->smcibdev. The socket remains hashed throughout — so the *unhash*
invariant is not violated — but the diag reader can hold a stale pointer
to the old link and race this memset. The timeline:

diag reader [hash read_lock held]:
    conn->lnk → old_lnk (non-NULL, socket hashed ✓)
    [about to read old_lnk->smcibdev->ibdev->name]

smc_switch_link_and_count() [send_lock held]:
    smcr_link_put(old_lnk) → last ref → __smcr_link_clear()
        memset(old_lnk, 0, ...) ← smcibdev = NULL

diag reader:
    old_lnk->smcibdev->ibdev->name ← NULL deref

The *unhash* invariant says nothing about the link a connection used to
point at before *smc_switch_link_and_count*() swapped it. Hash
membership of the socket provides no protection here because the socket
is still hashed — the link pointer simply changed underneath the diag
reader.

Any ideas on this?

      reply	other threads:[~2026-09-01 15:13 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-28  6:54 [PATCH net v2 0/2] net/smc: fix diag dump lifetime races Mahanta Jambigi
2026-08-28  6:54 ` [PATCH net v2 1/2] net/smc: add connection lifetime infrastructure for diag Mahanta Jambigi
2026-08-29  6:55   ` sashiko-bot
2026-08-28  6:54 ` [PATCH net v2 2/2] net/smc: fix races in smc_diag dump path Mahanta Jambigi
2026-08-29  6:55   ` sashiko-bot
2026-08-31 13:42 ` [PATCH net v2 0/2] net/smc: fix diag dump lifetime races Dust Li
2026-08-31 14:08   ` Mahanta Jambigi
2026-08-31 14:57   ` Mahanta Jambigi
2026-09-01 13:02     ` Dust Li
2026-09-01 15:13       ` Mahanta Jambigi [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=85fa9f02-7516-4595-b5a8-4ae4ca845121@linux.ibm.com \
    --to=mjambigi@linux.ibm.com \
    --cc=alibuda@linux.alibaba.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=davem@davemloft.net \
    --cc=dust.li@linux.alibaba.com \
    --cc=edumazet@google.com \
    --cc=guwen@linux.alibaba.com \
    --cc=hidayath@linux.ibm.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=pasic@linux.ibm.com \
    --cc=sidraya@linux.ibm.com \
    --cc=stable@vger.kernel.org \
    --cc=tonylu@linux.alibaba.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox