From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.16]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A7EB2419304 for ; Thu, 8 Oct 2026 21:57:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.16 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791496642; cv=none; b=FfkQtQUULBBnT5qeBYteXEGJsD5mSPTKfMiYiO+01A3YZRwlfesovYmivQWV+SE2DXGmSU3ZHWf75Q9R0lMpfJM20+dneFA1UmlVcErJ9s3kjo3vC4p6l1ebFVQ2JJT1Ee+vBXn9E6pCkccBnNLkOfDiggCVvwFhUwk2gF0/qyM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791496642; c=relaxed/simple; bh=7+9wfRtDak+xZW+K1hZGTkZE1Hg4uzA9L6EeVYX/HlQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=ORyG45epuhNYFYQqKQKWYTVC9SpyZxFwgYUppsCwm+CKkmalWx/7piseX6ISfc38EcObLF93imiBdf2yqHQDqJPHhFe+olBpxw1d64xqUOIrwUT8N7YQ/golLSJ+OKt5M7h44X/OzBnuUsvtmdfH8h4gPYGNyBGXs+wYDowl31k= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=fI/CKlWE; arc=none smtp.client-ip=192.198.163.16 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="fI/CKlWE" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1791496640; x=1823032640; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=7+9wfRtDak+xZW+K1hZGTkZE1Hg4uzA9L6EeVYX/HlQ=; b=fI/CKlWEDHRdx0eiQgthNlEGXE3gqvsfwTN7+nI7715pVjs49DAIh5+B XZeLwV0vBTArtKD0nF6LkGgRkHeh5kb+pgVI3ZtT+TQygkokZ+Xbpz8cE lkg4W2fCv9x67i4HKbTbH4uuk0RObsG/3WEAJTog5ljiXHbVv7VlTTH9Q DMHPl+2OyYH9xE8TbWrSfE3F0Lai/rWKIJqDOkfkIXKObwvlPQoMq2jQf 9hh15yA7kyf6Om2RdCOIDHnQd8yW3sIeH6ds/7WklVdtVtCtZScwQC3Gp ZuDwOuQ+xartTosNghowux0HXHOrUlMqZcmsDeHVNBo//xaKLuEarlyPr A==; X-CSE-ConnectionGUID: J6mzYaK8Tdi4zyAlk8/EsQ== X-CSE-MsgGUID: qvefD/13TG6Bf7vacxeO+A== X-IronPort-AV: E=McAfee;i="6800,10657,11929"; a="294628" X-IronPort-AV: E=Sophos;i="6.27,147,1787036400"; d="scan'208";a="294628" Received: from orviesa003.jf.intel.com ([10.64.159.143]) by fmvoesa110.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 08 Oct 2026 14:57:17 -0700 X-CSE-ConnectionGUID: 0woJ8DH5S12TbNN0oUD1Rg== X-CSE-MsgGUID: mQOZ5+A3SIaQnU1L4d1qUQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.27,147,1787036400"; d="scan'208";a="150400" Received: from anguy11-upstream.jf.intel.com ([10.166.9.133]) by orviesa003.jf.intel.com with ESMTP; 08 Oct 2026 14:57:16 -0700 From: Tony Nguyen To: davem@davemloft.net, kuba@kernel.org, pabeni@redhat.com, edumazet@kernel.org, andrew+netdev@lunn.ch, netdev@vger.kernel.org Cc: Jacob Keller , anthony.l.nguyen@intel.com, maciej.machnikowski@intel.com, przemyslaw.korba@intel.com, grzegorz.nitka@intel.com, sergey.temerkhanov@intel.com, arkadiusz.kubalewski@intel.com, poros@redhat.com, richardcochran@gmail.com, horms@kernel.org, Alexander Nowlin Subject: [PATCH net v2 02/15] ice: use reference counting and SRCU for PTP port access Date: Thu, 8 Oct 2026 14:55:59 -0700 Message-ID: <20261008215614.1987250-3-anthony.l.nguyen@intel.com> X-Mailer: git-send-email 2.47.1 In-Reply-To: <20261008215614.1987250-1-anthony.l.nguyen@intel.com> References: <20261008215614.1987250-1-anthony.l.nguyen@intel.com> Precedence: bulk X-Mailing-List: netdev@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit From: Jacob Keller The ice adapter structure maintains a list of ports associated with the adapter. This is used for supporting PTP, where the clock owner must handle many operations that require access to the PTP port structures of the associated PFs. This is implemented using a linked list and a mutex. This sort of works, but a few places within the code do not acquire the mutex when iterating the list. This includes ice_ptp_flush_all_tx_tracker(), ice_ptp_restart_all_phy(), and ice_ptp_prepare_rebuild_sec(). Fixing this is tricky, especially since it is not clear if we can simply acquire the lock around the complete iterations. The pattern of use for the port list is read-mostly with modifications only happening during PF initialization when elements are inserted. This typically only happens during early boot, though a PF could in principle be removed or loaded at arbitrary times via bind and unbind operations. The use of a mutex does mean the driver can sleep while holding it, but it still creates complicates with lock ordering and prevents iterating the list in any code path that *can't* sleep. Instead, use SRCU primitives for the port linked list, along with a reference count on the port. The kref reference counter ensures that a PF will have a valid lifetime and not be removed until the reference is released. Note that this change focuses solely on the port list and does not make an effort to resolve access to ctrl_pf, which is currently being investigated by another developer. Use of sleepable RCU is required because we often iterate the PTP port list and perform operations that might sleep. Attempts at implementing regular RCU have thus far not proven to be acceptable. The remove path first removes the port from the list, and we use kref_get_unless_zero to ensure that such ports are skipped when iterating the list. This ensures that once a port starts removing we will drop references and no longer be able to acquire new ones. This avoids loop iterations chaining together to indefinitely block removal. The ice_ptp_release_port_srcu() function is used as the release function for the kref_put() call. This uses wake_up_var() to wake the removing thread. The waiting thread will block until the final reference has been removed, then it will use synchronize_srcu() to ensure any outstanding SRCU critical sections have had the necessary grace period. This flow ensures that accesses to ports via the adapter port list will remain valid until both the SRCU critical sections have ended and all the references to the ports have been dropped. Strictly speaking, SRCU alone might be sufficient for existing code paths, but the reference count allows the option for passing a pointer to the port on to other functions if necessary in the future. Fixes: e800654e85b5 ("ice: Use ice_adapter for PTP shared data instead of auxdev") Reviewed-by: Maciek Machnikowski Signed-off-by: Jacob Keller Tested-by: Alexander Nowlin Signed-off-by: Tony Nguyen --- drivers/net/ethernet/intel/ice/ice_adapter.c | 17 ++- drivers/net/ethernet/intel/ice/ice_adapter.h | 16 ++- drivers/net/ethernet/intel/ice/ice_ptp.c | 121 ++++++++++++++----- drivers/net/ethernet/intel/ice/ice_ptp.h | 4 + 4 files changed, 116 insertions(+), 42 deletions(-) diff --git a/drivers/net/ethernet/intel/ice/ice_adapter.c b/drivers/net/ethernet/intel/ice/ice_adapter.c index 2dc3629d6d0f..536923b6ae97 100644 --- a/drivers/net/ethernet/intel/ice/ice_adapter.c +++ b/drivers/net/ethernet/intel/ice/ice_adapter.c @@ -6,6 +6,7 @@ #include #include #include +#include #include #include "ice_adapter.h" #include "ice.h" @@ -54,11 +55,18 @@ static unsigned long ice_adapter_xa_index(struct pci_dev *pdev) static struct ice_adapter *ice_adapter_new(struct pci_dev *pdev) { struct ice_adapter *adapter; + int err; adapter = kzalloc_obj(*adapter); if (!adapter) return NULL; + err = init_srcu_struct(&adapter->ports.srcu); + if (err) { + kfree(adapter); + return NULL; + } + adapter->index = ice_adapter_index(pdev); spin_lock_init(&adapter->ptp_gltsyn_time_lock); spin_lock_init(&adapter->txq_ctx_lock); @@ -66,18 +74,19 @@ static struct ice_adapter *ice_adapter_new(struct pci_dev *pdev) mutex_init(&adapter->cpi_phy_lock[i]); refcount_set(&adapter->refcount, 1); - mutex_init(&adapter->ports.lock); - INIT_LIST_HEAD(&adapter->ports.ports); + spin_lock_init(&adapter->ports.lock); + INIT_LIST_HEAD(&adapter->ports.list); return adapter; } static void ice_adapter_free(struct ice_adapter *adapter) { - WARN_ON(!list_empty(&adapter->ports.ports)); + WARN_ON(!list_empty(&adapter->ports.list)); for (int i = 0; i < ARRAY_SIZE(adapter->cpi_phy_lock); i++) mutex_destroy(&adapter->cpi_phy_lock[i]); - mutex_destroy(&adapter->ports.lock); + + cleanup_srcu_struct(&adapter->ports.srcu); kfree(adapter); } diff --git a/drivers/net/ethernet/intel/ice/ice_adapter.h b/drivers/net/ethernet/intel/ice/ice_adapter.h index 4f695f32da3d..0b01c7f5cf0d 100644 --- a/drivers/net/ethernet/intel/ice/ice_adapter.h +++ b/drivers/net/ethernet/intel/ice/ice_adapter.h @@ -17,15 +17,19 @@ struct ice_pf; /** * struct ice_port_list - data used to store the list of adapter ports * - * This structure contains data used to maintain a list of adapter ports + * This structure contains data used to maintain a list of adapter ports. + * Writers modifying the list *must* acquire the lock, and use SRCU safe list + * operations. Readers should use srcu_read_lock() on the provided domain. * - * @ports: list of ports - * @lock: protect access to the ports list + * @list: list of ports + * @lock: protect write access to the list + * @srcu: Sleepable RCU domain for this adapter */ struct ice_port_list { - struct list_head ports; - /* To synchronize the ports list operations */ - struct mutex lock; + struct list_head list; + /* To synchronize write operations on the port list */ + spinlock_t lock; + struct srcu_struct srcu; }; /** diff --git a/drivers/net/ethernet/intel/ice/ice_ptp.c b/drivers/net/ethernet/intel/ice/ice_ptp.c index fe21cee4f9de..94a66e9d8c05 100644 --- a/drivers/net/ethernet/intel/ice/ice_ptp.c +++ b/drivers/net/ethernet/intel/ice/ice_ptp.c @@ -1,6 +1,9 @@ // SPDX-License-Identifier: GPL-2.0 /* Copyright (C) 2021, Intel Corporation. */ +#include +#include +#include #include "ice.h" #include "ice_lib.h" #include "ice_trace.h" @@ -673,20 +676,33 @@ static void ice_ptp_process_tx_tstamp(struct ice_ptp_tx *tx) pf->ptp.tx_hwtstamp_good += tstamp_good; } +static void ice_ptp_release_port_srcu(struct kref *ref) +{ + wake_up_var(ref); +} + static void ice_ptp_tx_tstamp_owner(struct ice_pf *pf) { + struct ice_port_list *ports = &pf->adapter->ports; struct ice_ptp_port *port; + int srcu_idx; - mutex_lock(&pf->adapter->ports.lock); - list_for_each_entry(port, &pf->adapter->ports.ports, list_node) { + srcu_idx = srcu_read_lock(&ports->srcu); + list_for_each_entry_srcu(port, &ports->list, list_node, + srcu_read_lock_held(&ports->srcu)) { struct ice_ptp_tx *tx = &port->tx; - if (!tx || !tx->init) + if (!tx->init) + continue; + + if (!kref_get_unless_zero(&port->ref)) continue; ice_ptp_process_tx_tstamp(tx); + + kref_put(&port->ref, ice_ptp_release_port_srcu); } - mutex_unlock(&pf->adapter->ports.lock); + srcu_read_unlock(&ports->srcu, srcu_idx); } /** @@ -806,10 +822,19 @@ ice_ptp_mark_tx_tracker_stale(struct ice_ptp_tx *tx) static void ice_ptp_flush_all_tx_tracker(struct ice_pf *pf) { + struct ice_port_list *ports = &pf->adapter->ports; struct ice_ptp_port *port; + int srcu_idx; - list_for_each_entry(port, &pf->adapter->ports.ports, list_node) + srcu_idx = srcu_read_lock(&ports->srcu); + list_for_each_entry_srcu(port, &ports->list, list_node, + srcu_read_lock_held(&ports->srcu)) { + if (!kref_get_unless_zero(&port->ref)) + continue; ice_ptp_flush_tx_tracker(ptp_port_to_pf(port), &port->tx); + kref_put(&port->ref, ice_ptp_release_port_srcu); + } + srcu_read_unlock(&ports->srcu, srcu_idx); } /** @@ -1424,16 +1449,22 @@ static void ice_ptp_reset_phy_timestamping(struct ice_pf *pf) */ static void ice_ptp_restart_all_phy(struct ice_pf *pf) { - struct list_head *entry; + struct ice_port_list *ports = &pf->adapter->ports; + struct ice_ptp_port *port; + int srcu_idx; - list_for_each(entry, &pf->adapter->ports.ports) { - struct ice_ptp_port *port = list_entry(entry, - struct ice_ptp_port, - list_node); + srcu_idx = srcu_read_lock(&ports->srcu); + list_for_each_entry_srcu(port, &ports->list, list_node, + srcu_read_lock_held(&ports->srcu)) { + if (!kref_get_unless_zero(&port->ref)) + continue; if (port->link_up) ice_ptp_port_phy_restart(port); + + kref_put(&port->ref, ice_ptp_release_port_srcu); } + srcu_read_unlock(&ports->srcu, srcu_idx); } /** @@ -2694,19 +2725,28 @@ static bool ice_port_has_timestamps(struct ice_ptp_tx *tx) static bool ice_any_port_has_timestamps(struct ice_pf *pf) { + struct ice_port_list *ports = &pf->adapter->ports; + bool have_tstamps = false; struct ice_ptp_port *port; + int srcu_idx; - scoped_guard(mutex, &pf->adapter->ports.lock) { - list_for_each_entry(port, &pf->adapter->ports.ports, - list_node) { - struct ice_ptp_tx *tx = &port->tx; + srcu_idx = srcu_read_lock(&ports->srcu); + list_for_each_entry_srcu(port, &ports->list, list_node, + srcu_read_lock_held(&ports->srcu)) { + if (!kref_get_unless_zero(&port->ref)) + continue; - if (ice_port_has_timestamps(tx)) - return true; - } + if (ice_port_has_timestamps(&port->tx)) + have_tstamps = true; + + kref_put(&port->ref, ice_ptp_release_port_srcu); + + if (have_tstamps) + break; } + srcu_read_unlock(&ports->srcu, srcu_idx); - return false; + return have_tstamps; } bool ice_ptp_tx_tstamps_pending(struct ice_pf *pf) @@ -2890,14 +2930,18 @@ void ice_ptp_queue_work(struct ice_pf *pf) static void ice_ptp_prepare_rebuild_sec(struct ice_pf *pf, bool rebuild, enum ice_reset_req reset_type) { - struct list_head *entry; + struct ice_port_list *ports = &pf->adapter->ports; + struct ice_ptp_port *port; + int srcu_idx; - list_for_each(entry, &pf->adapter->ports.ports) { - struct ice_ptp_port *port = list_entry(entry, - struct ice_ptp_port, - list_node); + srcu_idx = srcu_read_lock(&ports->srcu); + list_for_each_entry_srcu(port, &ports->list, list_node, + srcu_read_lock_held(&ports->srcu)) { struct ice_pf *peer_pf = ptp_port_to_pf(port); + if (!kref_get_unless_zero(&port->ref)) + continue; + if (!ice_is_primary(&peer_pf->hw)) { if (rebuild) { /* TODO: When implementing rebuild=true: @@ -2909,7 +2953,10 @@ static void ice_ptp_prepare_rebuild_sec(struct ice_pf *pf, bool rebuild, ice_ptp_prepare_for_reset(peer_pf, reset_type); } } + + kref_put(&port->ref, ice_ptp_release_port_srcu); } + srcu_read_unlock(&ports->srcu, srcu_idx); } /** @@ -3084,11 +3131,11 @@ static int ice_ptp_setup_pf(struct ice_pf *pf) return -ENODEV; INIT_LIST_HEAD(&ptp->port.list_node); - mutex_lock(&pf->adapter->ports.lock); + kref_init(&ptp->port.ref); - list_add(&ptp->port.list_node, - &pf->adapter->ports.ports); - mutex_unlock(&pf->adapter->ports.lock); + spin_lock(&pf->adapter->ports.lock); + list_add_rcu(&ptp->port.list_node, &pf->adapter->ports.list); + spin_unlock(&pf->adapter->ports.lock); /* Seed the per-PHY Tx reference clock usage map for this port. * Only meaningful on E825 (other MAC types don't expose tx-clk @@ -3110,13 +3157,23 @@ static int ice_ptp_setup_pf(struct ice_pf *pf) static void ice_ptp_cleanup_pf(struct ice_pf *pf) { + struct ice_port_list *ports = &pf->adapter->ports; struct ice_ptp *ptp = &pf->ptp; + struct kref *ref; - if (pf->hw.mac_type != ICE_MAC_UNKNOWN) { - mutex_lock(&pf->adapter->ports.lock); - list_del(&ptp->port.list_node); - mutex_unlock(&pf->adapter->ports.lock); - } + if (pf->hw.mac_type == ICE_MAC_UNKNOWN) + return; + + spin_lock(&ports->lock); + list_del_rcu(&ptp->port.list_node); + spin_unlock(&ports->lock); + + ref = &ptp->port.ref; + kref_put(ref, ice_ptp_release_port_srcu); + + wait_var_event(ref, !kref_read(ref)); + + synchronize_srcu(&ports->srcu); } /** diff --git a/drivers/net/ethernet/intel/ice/ice_ptp.h b/drivers/net/ethernet/intel/ice/ice_ptp.h index c4b0da7ce20e..da2003ba3bb0 100644 --- a/drivers/net/ethernet/intel/ice/ice_ptp.h +++ b/drivers/net/ethernet/intel/ice/ice_ptp.h @@ -5,6 +5,8 @@ #define _ICE_PTP_H_ #include +#include +#include #include #include "ice_ptp_hw.h" @@ -138,6 +140,7 @@ struct ice_ptp_tx { * and determine when the port's PHY offset is valid. * * @list_node: list member structure + * @ref: reference counter for use with adapter ports list * @tx: Tx timestamp tracking for this port * @ov_work: delayed work task for tracking when PHY offset is valid * @ps_lock: mutex used to protect the overall PTP PHY start procedure @@ -149,6 +152,7 @@ struct ice_ptp_tx { */ struct ice_ptp_port { struct list_head list_node; + struct kref ref; struct ice_ptp_tx tx; struct kthread_delayed_work ov_work; struct mutex ps_lock; /* protects overall PTP PHY start procedure */ -- 2.47.1