From: Linkui Xiao <xiaolinkui@126.com>
To: anthony.l.nguyen@intel.com, przemyslaw.kitszel@intel.com,
andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
kuba@kernel.org, pabeni@redhat.com
Cc: intel-wired-lan@lists.osuosl.org, netdev@vger.kernel.org,
linux-kernel@vger.kernel.org, Linkui Xiao <xiaolinkui@kylinos.cn>,
stable@vger.kernel.org
Subject: [PATCH iwl-net v3 2/3] ice: detach the VF representor when ice_start_vfs() fails
Date: Thu, 8 Oct 2026 20:57:53 +0800 [thread overview]
Message-ID: <20261008125754.3520773-3-xiaolinkui@126.com> (raw)
In-Reply-To: <20261008125754.3520773-1-xiaolinkui@126.com>
From: Linkui Xiao <xiaolinkui@kylinos.cn>
ice_start_vfs() attaches every VF it brings up to the eswitch with
ice_eswitch_attach_vf(), but the teardown path only undoes the queue
mappings and the VF VSI. Nothing calls ice_eswitch_detach_vf() for the
VFs that were attached before the failure, and the caller,
ice_ena_vfs(), goes straight to ice_free_vf_entries(), which drops the
last reference on every VF.
The port representors created for those VFs therefore outlive the
failed VF creation:
- the representor netdev stays registered and its devlink port stays
allocated, so both leak;
- repr->vf keeps pointing at the struct ice_vf that ice_put_vf() has
just freed through ice_sriov_free_vf(), and repr->src_vsi keeps
pointing at the VF VSI that ice_vf_vsi_release() tore down, so any
later use of a leftover netdev, for example
ice_eswitch_stop_all_tx_queues() walking pf->eswitch.reprs during a
PF reset, dereferences freed memory;
- the virtchnl ops of that VF, which ice_repr_add_vf() replaced with
ice_virtchnl_set_repr_ops(), are never handed back to
ice_virtchnl_set_dflt_ops();
- pf->eswitch.reprs never becomes empty, so ice_eswitch_detach()
never calls ice_eswitch_disable_switchdev().
pf->eswitch.is_running stays true with the bridge offloads and the
devlink rate topology still up, and ice_eswitch_release_env() is
skipped, leaving the uplink VSI in the switchdev configuration that
ice_eswitch_setup_env() gave it: local loopback enabled, Rx
filtering disabled and the default VSI steering removed.
Detach the representor in the teardown loop the way ice_free_vfs()
does, ahead of ice_vf_vsi_release(), because ice_repr_rem_vf() and
ice_eswitch_release_repr() both need repr->src_vsi to still be valid.
Every VF the teardown loop walks completed ice_eswitch_attach_vf()
successfully, and ice_eswitch_detach_vf() already returns early for a
VF without a representor, so no extra condition is needed.
The detach runs outside of vf->cfg_lock, the way the previous patch
leaves it in ice_free_vfs() and ice_reset_all_vfs(): taking the
devlink instance lock and then RTNL under cfg_lock is the wrong way
round against the ndo_set_vf_mac(), ndo_set_vf_vlan() and representor
ethtool reset paths. The rest of the loop body still runs under
cfg_lock, as in ice_free_vfs().
The VF is marked disabled first, because nothing else keeps a reset
away here. ICE_VF_DIS in pf->state is only set once ice_ena_vfs()
succeeds, and ice_sriov_configure() runs under the PCI device lock
rather than RTNL, so ICE_VF_STATE_DIS is what makes
ice_check_vf_ready_for_cfg() reject __ice_set_vf_mac() and
ice_set_vf_port_vlan(). Setting it under cfg_lock also waits out an
ice_reset_vf() that is already running, which would otherwise reach
ice_eswitch_update_repr() on a destroyed representor.
ice_vc_process_vf_msg() tests the same bit before it reads
vf->virtchnl_ops, which ice_repr_rem_vf() restores.
Found by code inspection of the VF setup and teardown error paths. It
was not triggered and no stack trace or error message was observed.
Compile-tested only, not run on hardware.
Fixes: fff292b47ac1 ("ice: add VF representors one by one")
Cc: stable@vger.kernel.org
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
---
Changes in v3:
- Detach the representor outside vf->cfg_lock instead of under it, as Przemek
suggested, and mark the VF disabled under cfg_lock before the detach so that
an ice_reset_vf() that is already past its own readiness check cannot reach
ice_eswitch_update_repr() while the representor goes away.
(Przemek Kitszel, Sashiko AI review)
- Include how the issue was found, that it has not been triggered, and that the
change is compile tested only, as netdev-bot asked for.
- Not carrying over the Reviewed-by tags from Tomasz Lichwala and
Aleksandr Loktionov, as the code changed after their reviews.
drivers/net/ethernet/intel/ice/ice_sriov.c | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
diff --git a/drivers/net/ethernet/intel/ice/ice_sriov.c b/drivers/net/ethernet/intel/ice/ice_sriov.c
index 471c1e29a865..470aec8849b6 100644
--- a/drivers/net/ethernet/intel/ice/ice_sriov.c
+++ b/drivers/net/ethernet/intel/ice/ice_sriov.c
@@ -520,8 +520,27 @@ static int ice_start_vfs(struct ice_pf *pf)
if (it_cnt == 0)
break;
+ /* Mark the VF disabled before its representor and its VSI go
+ * away, the way ice_free_vfs() does, and take cfg_lock
+ * around it to wait out an ice_reset_vf() already in
+ * progress. pf->state has no ICE_VF_DIS on this path and the
+ * loop leaves ICE_VF_STATE_INIT set, so without the bit a
+ * concurrent ice_reset_vf() would pass ice_is_vf_disabled()
+ * and reach ice_eswitch_update_repr() on a destroyed
+ * representor, or trip WARN_ON(!vsi) in ice_dis_vf_mappings().
+ */
+ mutex_lock(&vf->cfg_lock);
+ set_bit(ICE_VF_STATE_DIS, vf->vf_states);
+ mutex_unlock(&vf->cfg_lock);
+
+ /* detach outside of cfg_lock, see ice_free_vfs() */
+ ice_eswitch_detach_vf(pf, vf);
+
+ mutex_lock(&vf->cfg_lock);
ice_dis_vf_mappings(vf);
ice_vf_vsi_release(vf);
+ mutex_unlock(&vf->cfg_lock);
+
it_cnt--;
}
--
2.25.1
next prev parent reply other threads:[~2026-10-08 12:58 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 12:57 [PATCH iwl-net v3 0/3] ice: fix VF representor lock ordering and teardown error paths Linkui Xiao
2026-10-08 12:57 ` [PATCH iwl-net v3 1/3] ice: attach and detach VF representors outside of vf->cfg_lock Linkui Xiao
2026-10-08 12:57 ` Linkui Xiao [this message]
2026-10-08 12:57 ` [PATCH iwl-net v3 3/3] ice: free the VF MSI-X vectors when VF start fails Linkui Xiao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261008125754.3520773-3-xiaolinkui@126.com \
--to=xiaolinkui@126.com \
--cc=andrew+netdev@lunn.ch \
--cc=anthony.l.nguyen@intel.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=intel-wired-lan@lists.osuosl.org \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=przemyslaw.kitszel@intel.com \
--cc=stable@vger.kernel.org \
--cc=xiaolinkui@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox