Netdev List
 help / color / mirror / Atom feed
From: Linkui Xiao <xiaolinkui@126.com>
To: anthony.l.nguyen@intel.com, przemyslaw.kitszel@intel.com,
	andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
	kuba@kernel.org, pabeni@redhat.com
Cc: intel-wired-lan@lists.osuosl.org, netdev@vger.kernel.org,
	linux-kernel@vger.kernel.org, Linkui Xiao <xiaolinkui@kylinos.cn>,
	stable@vger.kernel.org
Subject: [PATCH iwl-net v3 2/3] ice: detach the VF representor when ice_start_vfs() fails
Date: Thu,  8 Oct 2026 20:57:53 +0800	[thread overview]
Message-ID: <20261008125754.3520773-3-xiaolinkui@126.com> (raw)
In-Reply-To: <20261008125754.3520773-1-xiaolinkui@126.com>

From: Linkui Xiao <xiaolinkui@kylinos.cn>

ice_start_vfs() attaches every VF it brings up to the eswitch with
ice_eswitch_attach_vf(), but the teardown path only undoes the queue
mappings and the VF VSI. Nothing calls ice_eswitch_detach_vf() for the
VFs that were attached before the failure, and the caller,
ice_ena_vfs(), goes straight to ice_free_vf_entries(), which drops the
last reference on every VF.

The port representors created for those VFs therefore outlive the
failed VF creation:

  - the representor netdev stays registered and its devlink port stays
    allocated, so both leak;

  - repr->vf keeps pointing at the struct ice_vf that ice_put_vf() has
    just freed through ice_sriov_free_vf(), and repr->src_vsi keeps
    pointing at the VF VSI that ice_vf_vsi_release() tore down, so any
    later use of a leftover netdev, for example
    ice_eswitch_stop_all_tx_queues() walking pf->eswitch.reprs during a
    PF reset, dereferences freed memory;

  - the virtchnl ops of that VF, which ice_repr_add_vf() replaced with
    ice_virtchnl_set_repr_ops(), are never handed back to
    ice_virtchnl_set_dflt_ops();

  - pf->eswitch.reprs never becomes empty, so ice_eswitch_detach()
    never calls ice_eswitch_disable_switchdev().
    pf->eswitch.is_running stays true with the bridge offloads and the
    devlink rate topology still up, and ice_eswitch_release_env() is
    skipped, leaving the uplink VSI in the switchdev configuration that
    ice_eswitch_setup_env() gave it: local loopback enabled, Rx
    filtering disabled and the default VSI steering removed.

Detach the representor in the teardown loop the way ice_free_vfs()
does, ahead of ice_vf_vsi_release(), because ice_repr_rem_vf() and
ice_eswitch_release_repr() both need repr->src_vsi to still be valid.
Every VF the teardown loop walks completed ice_eswitch_attach_vf()
successfully, and ice_eswitch_detach_vf() already returns early for a
VF without a representor, so no extra condition is needed.

The detach runs outside of vf->cfg_lock, the way the previous patch
leaves it in ice_free_vfs() and ice_reset_all_vfs(): taking the
devlink instance lock and then RTNL under cfg_lock is the wrong way
round against the ndo_set_vf_mac(), ndo_set_vf_vlan() and representor
ethtool reset paths. The rest of the loop body still runs under
cfg_lock, as in ice_free_vfs().

The VF is marked disabled first, because nothing else keeps a reset
away here. ICE_VF_DIS in pf->state is only set once ice_ena_vfs()
succeeds, and ice_sriov_configure() runs under the PCI device lock
rather than RTNL, so ICE_VF_STATE_DIS is what makes
ice_check_vf_ready_for_cfg() reject __ice_set_vf_mac() and
ice_set_vf_port_vlan(). Setting it under cfg_lock also waits out an
ice_reset_vf() that is already running, which would otherwise reach
ice_eswitch_update_repr() on a destroyed representor.
ice_vc_process_vf_msg() tests the same bit before it reads
vf->virtchnl_ops, which ice_repr_rem_vf() restores.

Found by code inspection of the VF setup and teardown error paths. It
was not triggered and no stack trace or error message was observed.
Compile-tested only, not run on hardware.

Fixes: fff292b47ac1 ("ice: add VF representors one by one")
Cc: stable@vger.kernel.org
Signed-off-by: Linkui Xiao <xiaolinkui@kylinos.cn>
---
Changes in v3:
- Detach the representor outside vf->cfg_lock instead of under it, as Przemek
  suggested, and mark the VF disabled under cfg_lock before the detach so that
  an ice_reset_vf() that is already past its own readiness check cannot reach
  ice_eswitch_update_repr() while the representor goes away.
  (Przemek Kitszel, Sashiko AI review)
- Include how the issue was found, that it has not been triggered, and that the
  change is compile tested only, as netdev-bot asked for.
- Not carrying over the Reviewed-by tags from Tomasz Lichwala and
  Aleksandr Loktionov, as the code changed after their reviews.
 drivers/net/ethernet/intel/ice/ice_sriov.c | 19 +++++++++++++++++++
 1 file changed, 19 insertions(+)

diff --git a/drivers/net/ethernet/intel/ice/ice_sriov.c b/drivers/net/ethernet/intel/ice/ice_sriov.c
index 471c1e29a865..470aec8849b6 100644
--- a/drivers/net/ethernet/intel/ice/ice_sriov.c
+++ b/drivers/net/ethernet/intel/ice/ice_sriov.c
@@ -520,8 +520,27 @@ static int ice_start_vfs(struct ice_pf *pf)
 		if (it_cnt == 0)
 			break;
 
+		/* Mark the VF disabled before its representor and its VSI go
+		 * away, the way ice_free_vfs() does, and take cfg_lock
+		 * around it to wait out an ice_reset_vf() already in
+		 * progress. pf->state has no ICE_VF_DIS on this path and the
+		 * loop leaves ICE_VF_STATE_INIT set, so without the bit a
+		 * concurrent ice_reset_vf() would pass ice_is_vf_disabled()
+		 * and reach ice_eswitch_update_repr() on a destroyed
+		 * representor, or trip WARN_ON(!vsi) in ice_dis_vf_mappings().
+		 */
+		mutex_lock(&vf->cfg_lock);
+		set_bit(ICE_VF_STATE_DIS, vf->vf_states);
+		mutex_unlock(&vf->cfg_lock);
+
+		/* detach outside of cfg_lock, see ice_free_vfs() */
+		ice_eswitch_detach_vf(pf, vf);
+
+		mutex_lock(&vf->cfg_lock);
 		ice_dis_vf_mappings(vf);
 		ice_vf_vsi_release(vf);
+		mutex_unlock(&vf->cfg_lock);
+
 		it_cnt--;
 	}
 
-- 
2.25.1


  parent reply	other threads:[~2026-10-08 12:58 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08 12:57 [PATCH iwl-net v3 0/3] ice: fix VF representor lock ordering and teardown error paths Linkui Xiao
2026-10-08 12:57 ` [PATCH iwl-net v3 1/3] ice: attach and detach VF representors outside of vf->cfg_lock Linkui Xiao
2026-10-08 12:57 ` Linkui Xiao [this message]
2026-10-08 12:57 ` [PATCH iwl-net v3 3/3] ice: free the VF MSI-X vectors when VF start fails Linkui Xiao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261008125754.3520773-3-xiaolinkui@126.com \
    --to=xiaolinkui@126.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=anthony.l.nguyen@intel.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=intel-wired-lan@lists.osuosl.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=przemyslaw.kitszel@intel.com \
    --cc=stable@vger.kernel.org \
    --cc=xiaolinkui@kylinos.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox