From: "Rao, Nikhil" <nikhirao@amd.com>
To: Paolo Abeni <pabeni@redhat.com>, netdev@vger.kernel.org
Cc: kuba@kernel.org, brett.creeley@amd.com, eric.joyner@amd.com,
andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
"Nikhil P. Rao" <nikhil.rao@amd.com>
Subject: Re: [PATCH net] pds_core: fix cmd_regs access racing BAR unmap on reset
Date: Tue, 4 Aug 2026 17:28:00 -0700 [thread overview]
Message-ID: <1c2694d9-3d4e-4dca-b8be-40cc7f59f594@amd.com> (raw)
In-Reply-To: <917affb8-f750-4e11-aade-e4ebf13e96ce@redhat.com>
On 8/4/2026 3:26 AM, Paolo Abeni wrote:
>
> On 7/29/26 7:52 AM, Nikhil P. Rao wrote:
>> pdsc_reset_prepare() and pdsc_reset_done()'s pdsc_map_bars() error path
>> clear/iounmap cmd_regs without devcmd_lock, and pdsc_firmware_update()'s
>> download loop derefs cmd_regs after dropping and retaking the lock
>> without re-checking. An FLR concurrent with a devlink flash can unmap
>> cmd_regs under an in-flight devcmd, causing a NULL deref or a write to
>> unmapped MMIO.
>>
>> Take devcmd_lock across the BAR unmap/remap. Only the PF maps cmd_regs
>> and runs devcmd, so guard the locking to the PF.
>>
>> pdsc_unmap_bars() also clears info_regs, which has its own readers under
>> config_lock (and a lockless debugfs reader); that teardown race is
>> pre-existing and handled separately.
>>
>> Fixes: e96094c1d11c ("pds_core: Clear BARs on reset")
>> Reported-by: sashiko-bot <sashiko-bot@kernel.org>
>> Closes: https://sashiko.dev/#/patchset/20260708212222.296202-1-nikhil.rao%40amd.com?part=3
>> Assisted-by: Claude:claude-opus-4.8
>> Signed-off-by: Nikhil P. Rao <nikhil.rao@amd.com>
>> ---
>> drivers/net/ethernet/amd/pds_core/fw.c | 6 ++++++
>> drivers/net/ethernet/amd/pds_core/main.c | 10 +++++++++-
>> 2 files changed, 15 insertions(+), 1 deletion(-)
>>
>> diff --git a/drivers/net/ethernet/amd/pds_core/fw.c b/drivers/net/ethernet/amd/pds_core/fw.c
>> index fa626719e68d..cd7616ed9ef3 100644
>> --- a/drivers/net/ethernet/amd/pds_core/fw.c
>> +++ b/drivers/net/ethernet/amd/pds_core/fw.c
>> @@ -134,6 +134,12 @@ int pdsc_firmware_update(struct pdsc *pdsc, const struct firmware *fw,
>>
>> copy_sz = min_t(unsigned int, buf_sz, fw->size - offset);
>> mutex_lock(&pdsc->devcmd_lock);
>> + if (!pdsc->cmd_regs) {
>
> Sashiko notes this is still racy:
>
> https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260729055258.1416225-1-nikhil.rao%40amd.com
>
> The issue is marked as a pre-existing one, but IMHO is so strictly
> related that deserve fixing in the same change.
Agreed, v2 folds in the cmd_regs part:
https://lore.kernel.org/netdev/20260804235946.177762-1-nikhil.rao@amd.com/T/#u
A note on reachability: the path the review describes
(pdsc_devcmd_locked() re-arming the health worker during reset, the
worker then running pdsc_fw_up() -> pdsc_setup() -> pdsc_identify()) is
no longer reachable since cd09971dcc1c ("pds_core: keep the health
thread stopped during reset"), which disables the work item across the
reset instead of cancelling it.
I left the download loop alone. An interrupted download is not committed
to flash: a reset clears the device's update session so a resumed
download is rejected, and the device verifies the staged image before
writing it to a flash slot, reports PDS_RC_BAD_FW rather than activating it.
On centralizing the check, pdsc_devcmd_with_data() in the
PLDM series under review for net-next does exactly that. It isn't in
net, I'll convert pdsc_identify() and pdsc_core_init() to it once that
series lands.
On intr_ctrl, intr_status and db_pages: the readers are quiesced before
the unmap -- pdsc_fw_down() runs pdsc_teardown() -> pdsc_dev_uninit(),
which frees the interrupts and the queues, so the interrupt and
start/stop paths cannot run by then. v2 says this in the commit message.
What is left is debugfs, which outlives a reset: identity_show() reads
info_regs with no lock or NULL check, and the intr_ctrl regset reads
through a base captured at file creation. I have not posted a fix for
those; v2 says they are out of scope rather than claiming they are
handled elsewhere, as v1 did.
Thanks,
Nikhil
prev parent reply other threads:[~2026-08-05 0:28 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-29 5:52 [PATCH net] pds_core: fix cmd_regs access racing BAR unmap on reset Nikhil P. Rao
2026-08-04 10:26 ` Paolo Abeni
2026-08-05 0:28 ` Rao, Nikhil [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1c2694d9-3d4e-4dca-b8be-40cc7f59f594@amd.com \
--to=nikhirao@amd.com \
--cc=andrew+netdev@lunn.ch \
--cc=brett.creeley@amd.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=eric.joyner@amd.com \
--cc=kuba@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=nikhil.rao@amd.com \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox