From: Greg KH <gregkh@linuxfoundation.org>
To: Dragos Tatulea <dtatulea@nvidia.com>
Cc: stable@vger.kernel.org, saeedm@nvidia.com, tariqt@nvidia.com,
netdev@vger.kernel.org, phaddad@nvidia.com, Paul Saab <ps@mu.org>,
Simon Horman <horms@kernel.org>, Jakub Kicinski <kuba@kernel.org>
Subject: Re: [PATCH 6.1.y 6.6.y 6.12.y 6.18.y] net/mlx5e: xsk: Fix unlocked writing to ICOSQ
Date: Wed, 9 Sep 2026 14:55:50 +0200 [thread overview]
Message-ID: <2026090941-starting-coveting-d93a@gregkh> (raw)
In-Reply-To: <6a430d55-72e8-4333-b369-af215c9f0c3a@nvidia.com>
On Wed, Sep 09, 2026 at 02:47:16PM +0200, Dragos Tatulea wrote:
>
>
> On 09.09.26 14:21, Greg KH wrote:
> > On Wed, Sep 09, 2026 at 10:03:10AM +0000, Dragos Tatulea wrote:
> >> commit c326f9c68921e2f14dfcecb2f6b4216313d50248 upstream.
> >>
> >> During napi poll, when the affinity changes and there's still XSK work
> >> to be done, we trigger an ICOSQ interrupt on the new CPU. However, this
> >> triggering on the ICOSQ is done unprotected.
> >>
> >> mlx5e_trigger_irq() is called while mlx5e_xsk_alloc_rx_mpwqe() is
> >> running from a different CPU due to affinity change. This can happen
> >> because IRQ triggering is done after napi_complete_done(). At this point
> >> the NAPI can be scheduled on a different CPU. Like this:
> >>
> >> CPU A (old affinity, NAPI tail) CPU B (new affinity, fresh NAPI)
> >> ------------------------------- --------------------------------
> >> napi_complete_done() clears SCHED
> >> mlx5e_cq_arm(...)
> >> napi_schedule_prep() sets SCHED
> >> mlx5e_napi_poll()
> >> mlx5e_xsk_alloc_rx_mpwqe()
> >> memcpy 640 B UMR body
> >> advance sq->pc by 10
> >> mlx5e_trigger_irq(&c->icosq)
> >> wqe_info[pi] = {NOP, 1}
> >> mlx5e_post_nop() advances sq->pc
> >>
> >> The obvious fix would be to lock the ICOSQ. But the ICOSQ is expected to
> >> be accessed only from the channel's NAPI and has no locking. Kick the
> >> async ICOSQ instead which is always locked.
> >>
> >> This issue was noticed in the wild with the following splat:
> >>
> >> netdevice: ge-0-0-1: Bad OP in ICOSQ CQE: 0xd
> >> WARNING: drivers/net/ethernet/mellanox/mlx5/core/en_rx.c:826 [...]
> >> [...]
> >> Call Trace:
> >> <IRQ>
> >> mlx5e_napi_poll+0x11d/0x7f0 [mlx5_core]
> >> __napi_poll+0x30/0x200
> >> ? skb_defer_free_flush+0x9c/0xc0
> >> net_rx_action+0x2fe/0x3f0
> >> handle_softirqs+0xd8/0x340
> >> __irq_exit_rcu+0xbc/0xe0
> >> common_interrupt+0x85/0xa0
> >> </IRQ>
> >> <TASK>
> >> asm_common_interrupt+0x26/0x40
> >> [...]
> >> ---[ end trace 0000000000000000 ]---
> >> mlx5_core 0000:08:00.0 ge-0-0-1: Error cqe on cqn 0x548, ci 0x2022, qn 0x8f4,
> >> opcode 0xd, syndrome 0x2, vendor syndrome 0x68
> >
> > Why does the changelog here differ from what is in Linus's tree?
> >
> > Please don't do that :(
> >
> Apologies. Will fix.
>
> >> [ Backport to 6.18.y and older: upstream commit calls
> >> mlx5e_trigger_napi_async_icosq(), which was introduced by commit
> >> 0da1dba72616 ("net/mlx5e: XSK, Fix unintended ICOSQ change") and is not
> >> present here. In these trees mlx5e_trigger_napi_icosq() is the
> >> equivalent helper: it takes c->async_icosq_lock and triggers
> >> c->async_icosq, which is unconditionally opened, activated, polled and
> >> armed for every channel. Race B of the upstream commit message does not
> >> apply, as it concerns the sync-ICOSQ variant of
> >> mlx5e_trigger_napi_icosq() that only exists upstream. ]
> >
> > That part is ok, but changing the overall changelog text for no obvious
> > reson isn't ok.
> >
> Do you prefer that I drop it in v2.
No, that's ok, it's great to have, thanks for that.
greg k-h
prev parent reply other threads:[~2026-09-09 12:57 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-09 10:03 [PATCH 6.1.y 6.6.y 6.12.y 6.18.y] net/mlx5e: xsk: Fix unlocked writing to ICOSQ Dragos Tatulea
2026-09-09 12:21 ` Greg KH
2026-09-09 12:47 ` Dragos Tatulea
2026-09-09 12:55 ` Greg KH
2026-09-09 12:55 ` Greg KH [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=2026090941-starting-coveting-d93a@gregkh \
--to=gregkh@linuxfoundation.org \
--cc=dtatulea@nvidia.com \
--cc=horms@kernel.org \
--cc=kuba@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=phaddad@nvidia.com \
--cc=ps@mu.org \
--cc=saeedm@nvidia.com \
--cc=stable@vger.kernel.org \
--cc=tariqt@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox