From: Linkui Xiao <xiaolinkui@126.com>
To: Przemek Kitszel <przemyslaw.kitszel@intel.com>,
Tony Nguyen <anthony.l.nguyen@intel.com>
Cc: andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
kuba@kernel.org, netdev-bot+sinfo@kernel.org, pabeni@redhat.com,
intel-wired-lan@lists.osuosl.org, netdev@vger.kernel.org,
linux-kernel@vger.kernel.org, Linkui Xiao <xiaolinkui@kylinos.cn>,
stable@vger.kernel.org
Subject: Re: [PATCH iwl-net v2 1/2] ice: detach the VF representor when ice_start_vfs() fails
Date: Thu, 8 Oct 2026 21:08:57 +0800 [thread overview]
Message-ID: <b402a837-2223-4d72-978a-b2108df78334@126.com> (raw)
In-Reply-To: <0c464819-16ee-4bae-b15c-3337516ef3c9@intel.com>
On 2026/10/7 19:11, Przemek Kitszel wrote:
> On 9/29/26 8:11 PM, Tony Nguyen wrote:
>>
>>
>> On 9/28/2026 12:14 AM, Linkui Xiao wrote:
>>> Hi,
>>>
>>> Thanks for the review. Please find the missing information below.
>>>
>>> - How the issue was discovered:
>>> Found during manual code inspection of the ice VF setup/teardown
>>> error paths. It was not reported by syzbot, a static analysis tool,
>>> or an LLM scan, and was not hit in production.
>>>
>>> - Whether the issue was actually triggered:
>>> Not actually triggered. It is a theoretical error-path cleanup
>>> issue found by inspection; no stack trace or error message was observed.
>>>
>>> - Hardware tested:
>>> Not tested on real hardware. The change was only compile-tested;
>>> no affected Intel NIC/firmware test was performed.
>>>
>>> The same applies to patch 2/2: it was also found by the same manual
>>> code inspection, was not triggered at runtime, and was only compile-
>>> tested, not tested on real hardware.
>>>
>>> If a v3 is needed for other review reasons, I will include this
>>> information in the commit messages.
>>>
>>> Thanks,
>>> Linkui Xiao
>>>
>>> On 2026/9/28 14:59, netdev-bot+sinfo@kernel.org wrote:
>>>> Hi!
>>>>
>>>> This is an automated message. This series looks like a fix, but its
>>>> commit messages seem to be missing some information:
>>>>
>>>> - How the issue was discovered, e.g. hit in production, hit during
>>>> development, syzbot report, manual code inspection, LLM or static
>>>> analysis tool scan.
>>>>
>>>> - Whether the issue was actually triggered, or is only theoretical
>>>> (e.g. found by code inspection). If it was triggered please include
>>>> the symptoms, like the stack trace or error messages.
>>>>
>>>> - What hardware the change was tested on. For driver fixes please
>>>> mention the device (and if relevant firmware version) used for
>>>> testing, or say that the change was not tested on real hardware.
>>>>
>>>> Please do not repost the series just to address the above. Instead,
>>>> reply to this email with the missing information, so that reviewers
>>>> can take it into account. If the series needs another revision for
>>>> other reasons, please include the information in the commit messages
>>>> then.
>>
>> Hi Jakub,
>>
>> For patches that go through iwl-net, how did you want this handled? It
>> says to not repost just to address the above, but since this will get
>> submitted later, did you want the commit updated to make the latter
>> check happy, did you want me to copy/paste the response to the commit,
>> or something else?
>
> In general, I would say it's up to you Tony.
> Perhaps we could decide based on the amount of additional changes.
>
> For this particular patch sashiko already posted complains:
> https://sashiko.dev/#/patchset/20260928065306.1514795-1-xiaolinkui%40126.com?part=1
>
> I've encountered the same problem (although during my OOT encounters),
> and the solution is to detach/attach VF representors outside of
> vf->cfg_lock. More rationale by AI follows:
>
> ice_reset_all_vfs() and ice_free_vfs() call ice_eswitch_detach() and
> ice_eswitch_attach() with vf->cfg_lock held. Both take the devlink
> instance lock and then register or unregister the representor netdev,
> which takes RTNL. ndo_set_vf_mac(), ndo_set_vf_vlan() and the
> representor's ethtool reset take vf->cfg_lock under RTNL, so the order
> is inverted. The ice_check_vf_ready_for_cfg() check in those callers
> runs before vf->cfg_lock is taken, so it doesn't keep them away from a
> reset that is already in progress.
>
> Detach before taking vf->cfg_lock and attach after releasing it.
> Representors are added and removed only under pf->vfs.table_lock, and
> ice_start_vfs() already attaches without vf->cfg_lock. Both functions
> set ICE_VF_DIS first, so a concurrent ice_reset_vf() bails out before it
> reaches ice_eswitch_update_repr().
Thanks. That is the AB-BA Sashiko flagged, and v3 does exactly that. The
series is three patches now, reposted as a new thread whose cover is
[PATCH iwl-net v3 0/3] ice: fix VF representor lock ordering and teardown
error paths
with the per-patch changes under "Changes in v3:" in each patch:
1/3 ice: attach and detach VF representors outside of vf->cfg_lock (new)
2/3 ice: detach the VF representor when ice_start_vfs() fails (was 1/2)
3/3 ice: free the VF MSI-X vectors when VF start fails (was 2/2)
1/3 is new and is the "same problem" you describe. ice_free_vfs() and
ice_reset_all_vfs() are where the inversion was introduced: fff292b47ac1
turned the single ice_eswitch_release() that ice_free_vfs() made before its
loop into a per-VF detach under cfg_lock, and c9663f79cd82 did the same to
the ice_eswitch_rebuild() call in ice_reset_all_vfs(). 1/3 moves the detach
ahead of cfg_lock in both loops, and in ice_reset_all_vfs() the attach back
after the unlock. It also takes cfg_lock once before the detach, so that a
reset which is already inside the lock is waited out instead of being left
with a window: ICE_VF_DIS, raised before both loops, keeps later resets out,
but not one that got in ahead of it, and that one can reach
ice_eswitch_update_repr() while ice_repr_destroy() is freeing the
representor it looks up, which is free_netdev() + kfree() with no RCU grace
period. 2/3 is the original ice_start_vfs() fix with the same reordering,
plus the disabled marking described below. 3/3 is the MSI-X leak, the same
three calls, with its Fixes: tag corrected. Tony, if you would rather keep
the -net series at two patches, say so and I will drop 1/3.
The edge that had to go is cfg_lock -> RTNL. ice_eswitch_detach_vf() takes
devl_lock() and then RTNL through ice_repr_rem_vf() -> unregister_netdev(),
while __ice_set_vf_mac(), ice_set_vf_port_vlan() and the representor's
ethtool reset take cfg_lock with RTNL already held. The representor itself
does not need cfg_lock there: it is created and destroyed under
pf->vfs.table_lock, and ice_free_vfs(), ice_reset_all_vfs() and
ice_start_vfs() all run with that lock held.
On the ice_start_vfs() path I kept cfg_lock around the rest of the teardown,
because the ICE_VF_DIS half of the rationale does not hold there.
ice_free_vfs() and ice_reset_all_vfs() raise ICE_VF_DIS before they
touch any
VF, the SR-IOV enable path never does (the clear_bit() in ice_ena_vfs() has
no matching set), and the VFs unwound by that loop already reached
ICE_VF_STATE_INIT, so "ip link set dev <pf> vf N ..." still passes
ice_check_vf_ready_for_cfg() and can run ice_reset_vf() against the VSI that
ice_vf_vsi_release() is tearing down.
That is why 2/3 marks the VF disabled under cfg_lock before the detach. As
you said, ice_check_vf_ready_for_cfg() runs before the lock is taken, so it
does not keep a reset that is already in progress away; the check that does
keep one out is the ice_is_vf_disabled() test ice_reset_vf() makes after
taking cfg_lock. Taking cfg_lock once before the detach waits that one out
while its representor is still attached, and the bit keeps later ones from
reaching ice_eswitch_update_repr(). Without it a reset already past its own
check could xa_load() the representor while ice_repr_destroy() frees it, and
that is free_netdev() + kfree(), with no RCU grace period in between.
v3 also carries the origin and testing information netdev-bot asked for, in
the commit messages, since the series is being reposted anyway. 2/3 changed
code, and in 3/3 the call in the teardown loop moved inside cfg_lock, so I
did not carry over the Reviewed-by tags from Tomasz Lichwala and Aleksandr
Loktionov. Tomasz's nit on the eswitch attach failure path is in: the IRQs
are freed before the VSI is released there, and the teardown loop keeps that
order too. The release_vsi label of ice_init_vf_vsi_res() still
releases the
VSI first, because the !vsi path has to skip that call.
The v3 series was posted here:
https://lore.kernel.org/netdev/20261008125754.3520773-2-xiaolinkui@126.com/
next prev parent reply other threads:[~2026-10-08 13:10 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-28 6:53 [PATCH iwl-net v2 1/2] ice: detach the VF representor when ice_start_vfs() fails Linkui Xiao
2026-09-28 6:53 ` [PATCH iwl-net v2 2/2] ice: release the VF MSI-X window " Linkui Xiao
2026-09-28 13:08 ` Tomasz Lichwala
2026-09-29 0:55 ` Linkui Xiao
2026-09-28 15:16 ` Loktionov, Aleksandr
2026-09-28 6:59 ` [PATCH iwl-net v2 1/2] ice: detach the VF representor " netdev-bot+sinfo
2026-09-28 7:14 ` Linkui Xiao
2026-09-29 18:11 ` Tony Nguyen
2026-10-07 11:11 ` Przemek Kitszel
2026-10-08 13:08 ` Linkui Xiao [this message]
2026-09-28 13:08 ` Tomasz Lichwala
2026-09-28 15:16 ` Loktionov, Aleksandr
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b402a837-2223-4d72-978a-b2108df78334@126.com \
--to=xiaolinkui@126.com \
--cc=andrew+netdev@lunn.ch \
--cc=anthony.l.nguyen@intel.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=intel-wired-lan@lists.osuosl.org \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=netdev-bot+sinfo@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=przemyslaw.kitszel@intel.com \
--cc=stable@vger.kernel.org \
--cc=xiaolinkui@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox