From: Dragos Tatulea <dtatulea@nvidia.com>
To: Greg KH <gregkh@linuxfoundation.org>
Cc: stable@vger.kernel.org, saeedm@nvidia.com, tariqt@nvidia.com,
netdev@vger.kernel.org, phaddad@nvidia.com, Paul Saab <ps@mu.org>,
Simon Horman <horms@kernel.org>, Jakub Kicinski <kuba@kernel.org>
Subject: Re: [PATCH 6.1.y 6.6.y 6.12.y 6.18.y] net/mlx5e: xsk: Fix unlocked writing to ICOSQ
Date: Wed, 9 Sep 2026 14:47:16 +0200 [thread overview]
Message-ID: <6a430d55-72e8-4333-b369-af215c9f0c3a@nvidia.com> (raw)
In-Reply-To: <2026090925-myspace-precision-c68e@gregkh>
On 09.09.26 14:21, Greg KH wrote:
> On Wed, Sep 09, 2026 at 10:03:10AM +0000, Dragos Tatulea wrote:
>> commit c326f9c68921e2f14dfcecb2f6b4216313d50248 upstream.
>>
>> During napi poll, when the affinity changes and there's still XSK work
>> to be done, we trigger an ICOSQ interrupt on the new CPU. However, this
>> triggering on the ICOSQ is done unprotected.
>>
>> mlx5e_trigger_irq() is called while mlx5e_xsk_alloc_rx_mpwqe() is
>> running from a different CPU due to affinity change. This can happen
>> because IRQ triggering is done after napi_complete_done(). At this point
>> the NAPI can be scheduled on a different CPU. Like this:
>>
>> CPU A (old affinity, NAPI tail) CPU B (new affinity, fresh NAPI)
>> ------------------------------- --------------------------------
>> napi_complete_done() clears SCHED
>> mlx5e_cq_arm(...)
>> napi_schedule_prep() sets SCHED
>> mlx5e_napi_poll()
>> mlx5e_xsk_alloc_rx_mpwqe()
>> memcpy 640 B UMR body
>> advance sq->pc by 10
>> mlx5e_trigger_irq(&c->icosq)
>> wqe_info[pi] = {NOP, 1}
>> mlx5e_post_nop() advances sq->pc
>>
>> The obvious fix would be to lock the ICOSQ. But the ICOSQ is expected to
>> be accessed only from the channel's NAPI and has no locking. Kick the
>> async ICOSQ instead which is always locked.
>>
>> This issue was noticed in the wild with the following splat:
>>
>> netdevice: ge-0-0-1: Bad OP in ICOSQ CQE: 0xd
>> WARNING: drivers/net/ethernet/mellanox/mlx5/core/en_rx.c:826 [...]
>> [...]
>> Call Trace:
>> <IRQ>
>> mlx5e_napi_poll+0x11d/0x7f0 [mlx5_core]
>> __napi_poll+0x30/0x200
>> ? skb_defer_free_flush+0x9c/0xc0
>> net_rx_action+0x2fe/0x3f0
>> handle_softirqs+0xd8/0x340
>> __irq_exit_rcu+0xbc/0xe0
>> common_interrupt+0x85/0xa0
>> </IRQ>
>> <TASK>
>> asm_common_interrupt+0x26/0x40
>> [...]
>> ---[ end trace 0000000000000000 ]---
>> mlx5_core 0000:08:00.0 ge-0-0-1: Error cqe on cqn 0x548, ci 0x2022, qn 0x8f4,
>> opcode 0xd, syndrome 0x2, vendor syndrome 0x68
>
> Why does the changelog here differ from what is in Linus's tree?
>
> Please don't do that :(
>
Apologies. Will fix.
>> [ Backport to 6.18.y and older: upstream commit calls
>> mlx5e_trigger_napi_async_icosq(), which was introduced by commit
>> 0da1dba72616 ("net/mlx5e: XSK, Fix unintended ICOSQ change") and is not
>> present here. In these trees mlx5e_trigger_napi_icosq() is the
>> equivalent helper: it takes c->async_icosq_lock and triggers
>> c->async_icosq, which is unconditionally opened, activated, polled and
>> armed for every channel. Race B of the upstream commit message does not
>> apply, as it concerns the sync-ICOSQ variant of
>> mlx5e_trigger_napi_icosq() that only exists upstream. ]
>
> That part is ok, but changing the overall changelog text for no obvious
> reson isn't ok.
>
Do you prefer that I drop it in v2.
Also, should I prefix the next patch with v2? (First time sending a stable
patch like this)
Thanks,
Dragos
next prev parent reply other threads:[~2026-09-09 12:47 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-09 10:03 [PATCH 6.1.y 6.6.y 6.12.y 6.18.y] net/mlx5e: xsk: Fix unlocked writing to ICOSQ Dragos Tatulea
2026-09-09 12:21 ` Greg KH
2026-09-09 12:47 ` Dragos Tatulea [this message]
2026-09-09 12:55 ` Greg KH
2026-09-09 12:55 ` Greg KH
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=6a430d55-72e8-4333-b369-af215c9f0c3a@nvidia.com \
--to=dtatulea@nvidia.com \
--cc=gregkh@linuxfoundation.org \
--cc=horms@kernel.org \
--cc=kuba@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=phaddad@nvidia.com \
--cc=ps@mu.org \
--cc=saeedm@nvidia.com \
--cc=stable@vger.kernel.org \
--cc=tariqt@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox