* [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
@ 2026-09-28 23:04 Tony Nguyen
2026-09-28 23:04 ` [PATCH net 1/6] idpf: fix possible race on remove during a reset Tony Nguyen
` (8 more replies)
0 siblings, 9 replies; 28+ messages in thread
From: Tony Nguyen @ 2026-09-28 23:04 UTC (permalink / raw)
To: davem, kuba, pabeni, edumazet, andrew+netdev, netdev
Cc: Tony Nguyen, emil.s.tantilov, luoxuanqiang, bryan.fraschetti,
tristan, tomasz.lichwala, david.butler, horms
For idpf:
Emil fixes a remove/reset race by always stopping the vport during
idpf_stop() so NAPI is properly torn down.
For ice:
Xuanqiang Luo resolves a use-after-free issue by changing order of
operations so index is used before being freed.
Bryan Fraschetti restores ordered MMIO writes for ice Tx doorbells
by replacing writel_relaxed() call with writel().
Tristan Madani fixes representor use-after-free by releasing
metadata_dst through dst_release() to ensure it is not freed until
all references are dropped.
For iavf:
Tomasz fixes reporting of statistics by always requesting statistics
while the adapter is running as PTP commands can interfere with the
previous fallback stats update mechanism.
Dave Butler caps the advertised maximum packet size to the hardware's
single-buffer limit, preventing PF queue-configuration failures caused
by oversized multi-buffer frame limits.
The following are changes since commit a7bfaba4823e3c165bb2004c74eff7c096672bc7:
ipv6: fix prefix route expiry in modify_prefix_route()
and are available in the git repository at:
git://git.kernel.org/pub/scm/linux/kernel/git/tnguy/net-queue 200GbE
Bryan Fraschetti (1):
ice: Restore Ordered MMIO Writes for Tx Doorbells
Dave Butler (1):
iavf: cap advertised max_pkt_size at the single-buffer HW limit
Emil Tantilov (1):
idpf: fix possible race on remove during a reset
Tomasz Lichwala (1):
iavf: fix VF stats not updating due to PTP command preemption
Tristan Madani (1):
ice: fix metadata_dst refcount handling on representor teardown
Xuanqiang Luo (1):
ice: fix use-after-free in dynamic port cleanup
drivers/net/ethernet/intel/iavf/iavf_main.c | 14 ++++----------
drivers/net/ethernet/intel/iavf/iavf_virtchnl.c | 8 ++++++++
drivers/net/ethernet/intel/ice/devlink/port.c | 2 +-
drivers/net/ethernet/intel/ice/ice_eswitch.c | 2 +-
drivers/net/ethernet/intel/ice/ice_txrx.c | 4 ++--
drivers/net/ethernet/intel/idpf/idpf_lib.c | 4 ----
6 files changed, 16 insertions(+), 18 deletions(-)
--
2.47.1
^ permalink raw reply [flat|nested] 28+ messages in thread
* [PATCH net 1/6] idpf: fix possible race on remove during a reset
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
@ 2026-09-28 23:04 ` Tony Nguyen
2026-09-30 0:58 ` netdev-bot+sashiko
2026-09-28 23:04 ` [PATCH net 2/6] ice: fix use-after-free in dynamic port cleanup Tony Nguyen
` (7 subsequent siblings)
8 siblings, 1 reply; 28+ messages in thread
From: Tony Nguyen @ 2026-09-28 23:04 UTC (permalink / raw)
To: davem, kuba, pabeni, edumazet, andrew+netdev, netdev
Cc: Emil Tantilov, anthony.l.nguyen, luoxuanqiang, bryan.fraschetti,
tristan, tomasz.lichwala, david.butler, horms, Joshua Hay,
Samuel Salin
From: Emil Tantilov <emil.s.tantilov@intel.com>
Reset and remove can race, leaving NAPI registered and enabled:
modprobe idpf& sleep 1; ip link set eth0 up& rmmod idpf
[145561.805104] WARNING: net/core/dev.c:7699 at __netif_napi_del_locked+0x11a/0x130, CPU#30: rmmod/22393
...
[145561.810238] RIP: 0010:__netif_napi_del_locked+0x11a/0x130
...
[145561.817678] Call Trace:
[145561.818125] <TASK>
[145561.818653] free_netdev+0x110/0x2a0
[145561.819109] idpf_vport_dealloc+0x452/0x460 [idpf]
[145561.819668] ? enable_work+0x9f/0x100
[145561.820133] idpf_deinit_task+0x51/0x70 [idpf]
[145561.820694] idpf_vc_core_deinit+0x32/0x170 [idpf]
[145561.821193] idpf_remove+0x40/0x200 [idpf]
[145561.821667] pci_device_remove+0x40/0xa0
[145561.822132] device_release_driver_internal+0x1a9/0x210
[145561.822691] driver_detach+0x4b/0x90
[145561.823161] bus_remove_driver+0x70/0x100
[145561.823721] pci_unregister_driver+0x2e/0xb0
[145561.824207] __do_sys_delete_module.constprop.0+0x190/0x2e0
[145561.824715] ? kmem_cache_free+0x312/0x550
[145561.825213] do_syscall_64+0xc8/0x6b0
[145561.825709] ? clear_bhb_loop+0x30/0x80
[145561.826216] entry_SYSCALL_64_after_hwframe+0x76/0x7e
[145561.826602] RIP: 0033:0x7fe628130beb
Make sure to call idpf_vport_stop() in idpf_stop(), irrespective of the
IDPF_REMOVE_IN_PROG state, to allow a reset racing with remove to tear
down NAPI in idpf_detach_and_close(), which runs under RTNL lock.
Fixes: 2e281e1155fc ("idpf: detach and close netdevs while handling a reset")
Signed-off-by: Emil Tantilov <emil.s.tantilov@intel.com>
Reviewed-by: Joshua Hay <joshua.a.hay@intel.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Tested-by: Samuel Salin <Samuel.salin@intel.com>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
---
drivers/net/ethernet/intel/idpf/idpf_lib.c | 4 ----
1 file changed, 4 deletions(-)
diff --git a/drivers/net/ethernet/intel/idpf/idpf_lib.c b/drivers/net/ethernet/intel/idpf/idpf_lib.c
index 827c795afcb6..2c148377540c 100644
--- a/drivers/net/ethernet/intel/idpf/idpf_lib.c
+++ b/drivers/net/ethernet/intel/idpf/idpf_lib.c
@@ -1033,12 +1033,8 @@ static void idpf_vport_stop(struct idpf_vport *vport, bool rtnl)
*/
static int idpf_stop(struct net_device *netdev)
{
- struct idpf_netdev_priv *np = netdev_priv(netdev);
struct idpf_vport *vport;
- if (test_bit(IDPF_REMOVE_IN_PROG, np->adapter->flags))
- return 0;
-
idpf_vport_ctrl_lock(netdev);
vport = idpf_netdev_to_vport(netdev);
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH net 2/6] ice: fix use-after-free in dynamic port cleanup
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
2026-09-28 23:04 ` [PATCH net 1/6] idpf: fix possible race on remove during a reset Tony Nguyen
@ 2026-09-28 23:04 ` Tony Nguyen
2026-09-28 23:04 ` [PATCH net 3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells Tony Nguyen
` (6 subsequent siblings)
8 siblings, 0 replies; 28+ messages in thread
From: Tony Nguyen @ 2026-09-28 23:04 UTC (permalink / raw)
To: davem, kuba, pabeni, edumazet, andrew+netdev, netdev
Cc: Xuanqiang Luo, anthony.l.nguyen, emil.s.tantilov,
bryan.fraschetti, tristan, tomasz.lichwala, david.butler, horms,
kees, michal.swiatkowski, stable, Marcin Szycik,
Aleksandr Loktionov, Patryk Holda
From: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
ice_dealloc_dynamic_port() uses dyn_port->vsi->idx to erase the dynamic
port from pf->dyn_ports. However, it frees the VSI before reading the
index for the erase, resulting in a use-after-free.
Follow the reverse of the allocation order in ice_alloc_dynamic_port()
by erasing the xarray entry before freeing the VSI.
Fixes: eda69d654c7e ("ice: add basic devlink subfunctions support")
Cc: stable@vger.kernel.org
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Reviewed-by: Marcin Szycik <marcin.szycik@linux.intel.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Tested-by: Patryk Holda <patryk.holda@intel.com>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
---
drivers/net/ethernet/intel/ice/devlink/port.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/intel/ice/devlink/port.c b/drivers/net/ethernet/intel/ice/devlink/port.c
index 2a2e56777f9f..3ede24649002 100644
--- a/drivers/net/ethernet/intel/ice/devlink/port.c
+++ b/drivers/net/ethernet/intel/ice/devlink/port.c
@@ -590,8 +590,8 @@ static void ice_dealloc_dynamic_port(struct ice_dynamic_port *dyn_port)
xa_erase(&pf->sf_nums, devlink_port->attrs.pci_sf.sf);
ice_eswitch_detach_sf(pf, dyn_port);
- ice_vsi_free(dyn_port->vsi);
xa_erase(&pf->dyn_ports, dyn_port->vsi->idx);
+ ice_vsi_free(dyn_port->vsi);
kfree(dyn_port);
}
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH net 3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
2026-09-28 23:04 ` [PATCH net 1/6] idpf: fix possible race on remove during a reset Tony Nguyen
2026-09-28 23:04 ` [PATCH net 2/6] ice: fix use-after-free in dynamic port cleanup Tony Nguyen
@ 2026-09-28 23:04 ` Tony Nguyen
2026-09-30 0:58 ` netdev-bot+sashiko
2026-09-28 23:04 ` [PATCH net 4/6] ice: fix metadata_dst refcount handling on representor teardown Tony Nguyen
` (5 subsequent siblings)
8 siblings, 1 reply; 28+ messages in thread
From: Tony Nguyen @ 2026-09-28 23:04 UTC (permalink / raw)
To: davem, kuba, pabeni, edumazet, andrew+netdev, netdev
Cc: Bryan Fraschetti, anthony.l.nguyen, emil.s.tantilov, luoxuanqiang,
tristan, tomasz.lichwala, david.butler, horms, paul.greenwalt,
maciej.fijalkowski, Alexander Nowlin
From: Bryan Fraschetti <bryan.fraschetti@canonical.com>
The DQL accounting state which tracks the number of bytes queued for the
NIC to transmit, namely dql->num_queued, is updated by the ICE driver
when it invokes __netdev_tx_sent_queue(). Subsequently the driver
updates the hardware queue tail allowing the NIC to begin processing the
work queue. The queue accounting update should be ordered before the NIC
begins transmitting descriptors.
The transmit path currently updates the hardware doorbell using
writel_relaxed(), which (on arm64) does not provide the same ordering
guarantees between writes to normal memory and writes to MMIO registers
that writel() does. This introduces a potential race where the NIC
begins transmitting descriptors before the dql->num_queued update is
globally visible. If the NIC finishes before the update is observed by
dql_completed(), it detects an invalid state where more bytes have been
completed than have been queued. When this happens the following BUG_ON
is triggered.
BUG_ON(count > num_queued - dql->num_completed);
This has been observed and manifests as the following crash (note that
the trace has been trimmed) in an environment with sustained network
load that uses an Intel Corporation Ethernet Controller E810-XXV for SFP
(rev 02) on an arm64 machine. Replacing writel_relaxed() with writel()
in a test kernel eliminated the crash in the user's workload, which
previously reproduced the issue reliably.
kernel BUG at lib/dynamic_queue_limits.c:99
Internal error: Oops - BUG: 00000000f2000800 [#1] SMP
pc : dql_completed+0x268/0x2a0
lr : ice_clean_tx_irq+0x1d4/0x620 [ice]
Call trace:
dql_completed+0x268/0x2a0 (P)
ice_napi_poll+0x94/0x520 [ice]
__napi_poll+0x48/0x3f0
net_rx_action+0x194/0x420
This restores the behaviour prior to commit ccde82e90946 ("ice: add E830
Earliest TxTime First Offload support"), which changed the notification
mechanism from writel() to writel_relaxed().
Link: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2161572
Fixes: ccde82e90946 ("ice: add E830 Earliest TxTime First Offload support")
Signed-off-by: Bryan Fraschetti <bryan.fraschetti@canonical.com>
Tested-by: Alexander Nowlin <alexander.nowlin@intel.com>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
---
drivers/net/ethernet/intel/ice/ice_txrx.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/drivers/net/ethernet/intel/ice/ice_txrx.c b/drivers/net/ethernet/intel/ice/ice_txrx.c
index 31303ab5be17..a2c7c4962882 100644
--- a/drivers/net/ethernet/intel/ice/ice_txrx.c
+++ b/drivers/net/ethernet/intel/ice/ice_txrx.c
@@ -1561,10 +1561,10 @@ ice_tx_map(struct ice_tx_ring *tx_ring, struct ice_tx_buf *first,
}
}
tstamp_ring->next_to_use = j;
- writel_relaxed(j, tstamp_ring->tail);
+ writel(j, tstamp_ring->tail);
} else {
ring_kick:
- writel_relaxed(i, tx_ring->tail);
+ writel(i, tx_ring->tail);
}
return;
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH net 4/6] ice: fix metadata_dst refcount handling on representor teardown
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
` (2 preceding siblings ...)
2026-09-28 23:04 ` [PATCH net 3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells Tony Nguyen
@ 2026-09-28 23:04 ` Tony Nguyen
2026-09-28 23:04 ` [PATCH net 5/6] iavf: fix VF stats not updating due to PTP command preemption Tony Nguyen
` (4 subsequent siblings)
8 siblings, 0 replies; 28+ messages in thread
From: Tony Nguyen @ 2026-09-28 23:04 UTC (permalink / raw)
To: davem, kuba, pabeni, edumazet, andrew+netdev, netdev
Cc: Tristan Madani, anthony.l.nguyen, emil.s.tantilov, luoxuanqiang,
bryan.fraschetti, tomasz.lichwala, david.butler, horms,
grzegorz.nitka, michal.swiatkowski, stable
From: Tristan Madani <tristan@talencesecurity.com>
ice_eswitch_release_repr() uses metadata_dst_free() to release the
representor's metadata_dst. metadata_dst_free() directly frees the
underlying memory without checking the dst_entry refcount.
When ice_eswitch_port_start_xmit() processes a packet, it takes a
reference via dst_hold() and attaches the metadata_dst to the skb.
If the representor is torn down while packets are still queued on
the lower device (e.g. in a qdisc), the metadata_dst is freed while
references are still held.
Use dst_release() instead, which correctly decrements the refcount
and only frees the object when all references are dropped. The dst
subsystem already handles metadata_dst cleanup in dst_destroy() when
DST_METADATA is set.
Other drivers sharing this pattern (nfp, airoha, bnxt) already use
dst_release() for their metadata_dst lifecycle.
Fixes: f5396b8a663f7 ("ice: switchdev slow path")
Cc: stable@vger.kernel.org
Signed-off-by: Tristan Madani <tristan@talencesecurity.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
---
drivers/net/ethernet/intel/ice/ice_eswitch.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/intel/ice/ice_eswitch.c b/drivers/net/ethernet/intel/ice/ice_eswitch.c
index b069e6c514fb..6e7bba473898 100644
--- a/drivers/net/ethernet/intel/ice/ice_eswitch.c
+++ b/drivers/net/ethernet/intel/ice/ice_eswitch.c
@@ -95,7 +95,7 @@ ice_eswitch_release_repr(struct ice_pf *pf, struct ice_repr *repr)
return;
ice_vsi_update_security(vsi, ice_vsi_ctx_set_antispoof);
- metadata_dst_free(repr->dst);
+ dst_release(&repr->dst->dst);
repr->dst = NULL;
ice_fltr_add_mac_and_broadcast(vsi, repr->parent_mac,
ICE_FWD_TO_VSI);
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH net 5/6] iavf: fix VF stats not updating due to PTP command preemption
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
` (3 preceding siblings ...)
2026-09-28 23:04 ` [PATCH net 4/6] ice: fix metadata_dst refcount handling on representor teardown Tony Nguyen
@ 2026-09-28 23:04 ` Tony Nguyen
2026-09-30 0:58 ` netdev-bot+sashiko
2026-09-28 23:04 ` [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit Tony Nguyen
` (3 subsequent siblings)
8 siblings, 1 reply; 28+ messages in thread
From: Tony Nguyen @ 2026-09-28 23:04 UTC (permalink / raw)
To: davem, kuba, pabeni, edumazet, andrew+netdev, netdev
Cc: Tomasz Lichwala, anthony.l.nguyen, emil.s.tantilov, luoxuanqiang,
bryan.fraschetti, tristan, david.butler, horms, richardcochran,
jacob.e.keller, Aleksandr Loktionov, Rafal Romanowski
From: Tomasz Lichwala <tomasz.lichwala@linux.intel.com>
Since the introduction of PTP support in iavf, commit 7c01dbfc8a1c
("iavf: periodically cache PHC time"), the periodic PTP clock caching
task always provides a pending admin queue command, causing
iavf_process_aq_command() to always return success and permanently
preventing the stats fallback path from executing, which results in VF
statistics remaining at zero despite traffic flowing.
Fix this by making the stats request unconditional when the adapter is in
the running state, rather than relying on it as a fallback when no other
admin queue commands were processed.
Fixes: 7c01dbfc8a1c ("iavf: periodically cache PHC time")
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Tomasz Lichwala <tomasz.lichwala@linux.intel.com>
Tested-by: Rafal Romanowski <rafal.romanowski@intel.com>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
---
drivers/net/ethernet/intel/iavf/iavf_main.c | 14 ++++----------
1 file changed, 4 insertions(+), 10 deletions(-)
diff --git a/drivers/net/ethernet/intel/iavf/iavf_main.c b/drivers/net/ethernet/intel/iavf/iavf_main.c
index 29b8403a066b..c0686ad5c411 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_main.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_main.c
@@ -2932,18 +2932,12 @@ static int iavf_watchdog_step(struct iavf_adapter *adapter)
iavf_send_api_ver(adapter);
}
} else {
- int ret = iavf_process_aq_command(adapter);
-
- /* An error will be returned if no commands were
- * processed; use this opportunity to update stats
- * if the error isn't -ENOTSUPP
- */
- if (ret && ret != -EOPNOTSUPP &&
- adapter->state == __IAVF_RUNNING)
- iavf_request_stats(adapter);
+ iavf_process_aq_command(adapter);
}
- if (adapter->state == __IAVF_RUNNING)
+ if (adapter->state == __IAVF_RUNNING) {
+ iavf_request_stats(adapter);
iavf_detect_recover_hung(&adapter->vsi);
+ }
break;
case __IAVF_REMOVE:
default:
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
` (4 preceding siblings ...)
2026-09-28 23:04 ` [PATCH net 5/6] iavf: fix VF stats not updating due to PTP command preemption Tony Nguyen
@ 2026-09-28 23:04 ` Tony Nguyen
2026-09-29 16:18 ` Alexander Lobakin
2026-09-30 0:58 ` netdev-bot+sashiko
2026-09-28 23:10 ` [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) netdev-bot+sinfo
` (2 subsequent siblings)
8 siblings, 2 replies; 28+ messages in thread
From: Tony Nguyen @ 2026-09-28 23:04 UTC (permalink / raw)
To: davem, kuba, pabeni, edumazet, andrew+netdev, netdev
Cc: Dave Butler, anthony.l.nguyen, emil.s.tantilov, luoxuanqiang,
bryan.fraschetti, tristan, tomasz.lichwala, horms,
aleksander.lobakin, stable, Jacob Keller, Aleksandr Loktionov
From: Dave Butler <david.butler@appgate.com>
Since commit 5fa4caff59f2 ("iavf: switch to Page Pool")
iavf_configure_queues() advertises max_pkt_size to the PF as:
max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
max_frame = min_not_zero(adapter->vf_res->max_mtu, max_frame);
LIBIE_MAX_RX_FRM_LEN (16382) is the multi-descriptor scatter/gather frame
ceiling, not a single-queue value, and it exceeds the E810 MAC frame size
maximum of 9728. Per the E810 datasheet (613875-009 section 13.2.2.17.1)
the Tx frame-size register PRTDCB_TDPUC.MAX_TXFRAME has a maximum of
0x2600 (9728); larger frames are discarded. The in-tree ice driver encodes
the same value as ICE_AQ_SET_MAC_FRAME_SIZE_MAX (== LIBIE_MAX_RX_BUF_LEN ==
9728), and the VF clamped max_frame to IAVF_MAX_RXBUFFER (9728) before this
commit.
When the PF advertises vf_res->max_mtu as 0, min_not_zero() leaves
max_frame at 16382. The Linux ice PF advertises max_mtu = port MAC frame
size (<= 9728), so a VF behind ice never sends more than that. The ESXi
"icen" PF on E810 advertises max_mtu as 0, so the VF sends
max_pkt_size = 16382, which icen rejects while programming the queue
context for VIRTCHNL_OP_CONFIG_VSI_QUEUES (opcode 6):
icen_ConfigureTxQueue: VSI 8: Failed to set LAN Tx queue context for
absolute Tx queue 64, Error: ICE_ERR_PARAM
indrv_SendMsgToVf: VF 0: Failed opcode 6, Error -5
iavf 0000:03:00.0: PF returned error -5 (IAVF_ERR_PARAM) to our request 6
iavf 0000:03:00.0 ethX: NETDEV WATCHDOG: transmit queue N timed out
The VF's queues never come up; under SR-IOV passthrough the mis-programmed
queue can also trigger a fatal IOMMU fault in the guest. Forcing only
max_pkt_size back to 9728 (and leaving the Page Pool rx_buf_len/
databuffer_size untouched) makes the VF come up; databuffer_size is not
involved. This was confirmed on two E810 NVM revisions (3.00 and 4.51) and
two icen versions (1.14.2.0 and the latest 2.3.3.0): all reject the
unpatched VF and accept the patched one, so the trigger is the icen PF
behaviour, not the firmware or icen revision. Reported by several users on
E810 + ESXi icen with v6.10+ guests:
Link: https://community.intel.com/t5/Ethernet-Products/E810-C-iavf-driver-issue-on-Linux-6-12/m-p/1737490
Link: https://access.redhat.com/solutions/6973766
Link: https://knowledge.broadcom.com/external/article/404315/sriov-enabled-vms-network-adaptor-goes-d.html
Cap max_frame at the single-buffer hardware limit, restoring the
pre-Page-Pool behaviour while keeping the Page Pool rx_buf_len unchanged.
Cc: stable@vger.kernel.org # v6.10+
Fixes: 5fa4caff59f2 ("iavf: switch to Page Pool")
Signed-off-by: Dave Butler <david.butler@appgate.com>
Acked-by: Jacob Keller <jacob.e.keller@intel.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Tony Nguyen <anthony.l.nguyen@intel.com>
---
drivers/net/ethernet/intel/iavf/iavf_virtchnl.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
index ec234cc8bd9d..680a28a739bf 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
@@ -382,6 +382,14 @@ void iavf_configure_queues(struct iavf_adapter *adapter)
max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
max_frame = min_not_zero(adapter->vf_res->max_mtu, max_frame);
+ /* The PF programs max_pkt_size into the per-queue Rx context "rxmax".
+ * LIBIE_MAX_RX_FRM_LEN is the multi-descriptor (S/G) frame ceiling
+ * (16382), but that exceeds the E810 max MAC frame size (9728); some
+ * PFs reject the out-of-range value with VIRTCHNL_STATUS_ERR_PARAM.
+ * Cap it at the single-buffer HW limit (== the MAC frame max),
+ * restoring the pre-Page-Pool behaviour.
+ */
+ max_frame = min(max_frame, LIBIE_MAX_RX_BUF_LEN);
if (adapter->current_op != VIRTCHNL_OP_UNKNOWN) {
/* bail because we already have a command pending */
--
2.47.1
^ permalink raw reply related [flat|nested] 28+ messages in thread
* Re: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
` (5 preceding siblings ...)
2026-09-28 23:04 ` [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit Tony Nguyen
@ 2026-09-28 23:10 ` netdev-bot+sinfo
2026-09-29 1:41 ` Dave Butler
` (2 more replies)
2026-10-01 23:47 ` Tony Nguyen
2026-10-02 20:50 ` patchwork-bot+netdevbpf
8 siblings, 3 replies; 28+ messages in thread
From: netdev-bot+sinfo @ 2026-09-28 23:10 UTC (permalink / raw)
To: Tony Nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
emil.s.tantilov, luoxuanqiang, bryan.fraschetti, tristan,
tomasz.lichwala, david.butler, horms
Hi!
This is an automated message. This series looks like a fix, but its
commit messages seem to be missing some information:
- How the issue was discovered, e.g. hit in production, hit during
development, syzbot report, manual code inspection, LLM or static
analysis tool scan.
- Whether the issue was actually triggered, or is only theoretical
(e.g. found by code inspection). If it was triggered please include
the symptoms, like the stack trace or error messages.
- What hardware the change was tested on. For driver fixes please
mention the device (and if relevant firmware version) used for
testing, or say that the change was not tested on real hardware.
Please do not repost the series just to address the above. Instead,
reply to this email with the missing information, so that reviewers
can take it into account. If the series needs another revision for
other reasons, please include the information in the commit messages
then.
The evaluation is done by an LLM so it may be wrong, if you think
that is the case please reply and explain.
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
2026-09-28 23:10 ` [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) netdev-bot+sinfo
@ 2026-09-29 1:41 ` Dave Butler
[not found] ` <IA3PR05MB22078467309FA5DFF33EF712488B892@IA3PR05MB220784.namprd05.prod.outlook.com>
2026-09-29 17:16 ` Tantilov, Emil S
2026-09-30 15:06 ` Tomasz Lichwala
2 siblings, 1 reply; 28+ messages in thread
From: Dave Butler @ 2026-09-29 1:41 UTC (permalink / raw)
To: netdev-bot+sinfo@kernel.org, Tony Nguyen
Cc: davem@davemloft.net, kuba@kernel.org, pabeni@redhat.com,
edumazet@kernel.org, andrew+netdev@lunn.ch,
netdev@vger.kernel.org, emil.s.tantilov@intel.com,
luoxuanqiang@kylinos.cn, bryan.fraschetti@canonical.com,
tristan@talencesecurity.com, tomasz.lichwala@linux.intel.com,
horms@kernel.org
Answers inline for :
iavf: cap advertised max_pkt_size at the single-buffer HW limit
> How the issue was discovered
Found in the field, in a real production environment, and reproduced in lab.
> Whether the issue was actually triggered, or is only theoretical. If it was triggered please include the symptoms, like the stack trace or error messages.
The issue was actually triggered. VM and/or hypervisor can lose all NIC functionality, VM may stall or crash.
esxi icen logs may include the following:
icen_ConfigureTxQueue: VSI 8: Failed to set LAN Tx queue context for absolute Tx queue 64, Error: ICE_ERR_PARAM
indrv_SendMsgToVf: VF 0: Failed opcode 6, Error -5
kernel messages may include the following:
iavf 0000:03:00.0: PF returned error -5 (IAVF_ERR_PARAM) to our request 6
iavf 0000:03:00.0 ethX: NETDEV WATCHDOG: transmit queue N timed out
> What hardware the change was tested on. For driver fixes please mention the device (and if relevant firmware version) used for testing, or say that the change was not tested on real hardware.
Key hardware for reproduction is E810 NIC, running on esxi (icen) with SR-IOV enabled.
E810 NVM revisions: 3.00, 4.51
icen versions: 1.14.2.0, 2.3.3.0
>The evaluation is done by an LLM so it may be wrong, if you think that is the case please reply and explain.
This information was covered comprehensively in the commit messages, but happy to repeat here.
The information contained in this electronic mail is confidential information intended only for the use of the individual(s) or entity(s) named. If the reader of the message is not the addressee (or authorized to receive for the addressee), you are hereby notified that any dissemination, distribution or copying of this communication is strictly prohibited. If you have received this communication in error, please immediately notify the sender by reply e-mail and/or by telephone and destroy the original message.
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit
2026-09-28 23:04 ` [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit Tony Nguyen
@ 2026-09-29 16:18 ` Alexander Lobakin
2026-09-29 20:32 ` Dave Butler
2026-09-30 0:58 ` netdev-bot+sashiko
1 sibling, 1 reply; 28+ messages in thread
From: Alexander Lobakin @ 2026-09-29 16:18 UTC (permalink / raw)
To: Tony Nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev, Dave Butler,
emil.s.tantilov, luoxuanqiang, bryan.fraschetti, tristan,
tomasz.lichwala, horms, stable, Jacob Keller, Aleksandr Loktionov
From: Anthony Nguyen <anthony.l.nguyen@intel.com>
Date: Mon, 28 Sep 2026 16:04:27 -0700
> From: Dave Butler <david.butler@appgate.com>
>
> Since commit 5fa4caff59f2 ("iavf: switch to Page Pool")
> iavf_configure_queues() advertises max_pkt_size to the PF as:
>
> max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
> max_frame = min_not_zero(adapter->vf_res->max_mtu, max_frame);
>
> LIBIE_MAX_RX_FRM_LEN (16382) is the multi-descriptor scatter/gather frame
> ceiling, not a single-queue value, and it exceeds the E810 MAC frame size
> maximum of 9728. Per the E810 datasheet (613875-009 section 13.2.2.17.1)
> the Tx frame-size register PRTDCB_TDPUC.MAX_TXFRAME has a maximum of
> 0x2600 (9728); larger frames are discarded. The in-tree ice driver encodes
> the same value as ICE_AQ_SET_MAC_FRAME_SIZE_MAX (== LIBIE_MAX_RX_BUF_LEN ==
> 9728), and the VF clamped max_frame to IAVF_MAX_RXBUFFER (9728) before this
> commit.
[...]
> diff --git a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
> index ec234cc8bd9d..680a28a739bf 100644
> --- a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
> +++ b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
> @@ -382,6 +382,14 @@ void iavf_configure_queues(struct iavf_adapter *adapter)
>
> max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
> max_frame = min_not_zero(adapter->vf_res->max_mtu, max_frame);
> + /* The PF programs max_pkt_size into the per-queue Rx context "rxmax".
> + * LIBIE_MAX_RX_FRM_LEN is the multi-descriptor (S/G) frame ceiling
> + * (16382), but that exceeds the E810 max MAC frame size (9728); some
> + * PFs reject the out-of-range value with VIRTCHNL_STATUS_ERR_PARAM.
> + * Cap it at the single-buffer HW limit (== the MAC frame max),
> + * restoring the pre-Page-Pool behaviour.
> + */
> + max_frame = min(max_frame, LIBIE_MAX_RX_BUF_LEN);
Shouldn't we cap it to 9728 instead?>
> if (adapter->current_op != VIRTCHNL_OP_UNKNOWN) {
> /* bail because we already have a command pending */
Thanks,
Olek
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
2026-09-28 23:10 ` [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) netdev-bot+sinfo
2026-09-29 1:41 ` Dave Butler
@ 2026-09-29 17:16 ` Tantilov, Emil S
2026-09-30 15:06 ` Tomasz Lichwala
2 siblings, 0 replies; 28+ messages in thread
From: Tantilov, Emil S @ 2026-09-29 17:16 UTC (permalink / raw)
To: netdev-bot+sinfo, Tony Nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
luoxuanqiang, bryan.fraschetti, tristan, tomasz.lichwala,
david.butler, horms
On 9/28/2026 4:10 PM, netdev-bot+sinfo@kernel.org wrote:
> Hi!
>
> This is an automated message. This series looks like a fix, but its
> commit messages seem to be missing some information:
For [PATCH net 1/6] idpf: fix possible race on remove during a reset
https://lore.kernel.org/netdev/20260928230429.495442-2-anthony.l.nguyen@intel.com/#r>
> - How the issue was discovered, e.g. hit in production, hit during
> development, syzbot report, manual code inspection, LLM or static
> analysis tool scan.
Stress testing during development.
>
> - Whether the issue was actually triggered, or is only theoretical
> (e.g. found by code inspection). If it was triggered please include
> the symptoms, like the stack trace or error messages.
It was triggered, the commands are in the patch description:
modprobe idpf& sleep 1; ip link set eth0 up& rmmod idpf
followed by the kernel trace.
>
> - What hardware the change was tested on. For driver fixes please
> mention the device (and if relevant firmware version) used for
> testing, or say that the change was not tested on real hardware.
The bug is not HW specific, but it was tested on Infrastructure Data
Path Function (dev id 0x1452).
Thanks,
Emil
>
> Please do not repost the series just to address the above. Instead,
> reply to this email with the missing information, so that reviewers
> can take it into account. If the series needs another revision for
> other reasons, please include the information in the commit messages
> then.
>
> The evaluation is done by an LLM so it may be wrong, if you think
> that is the case please reply and explain.
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit
2026-09-29 16:18 ` Alexander Lobakin
@ 2026-09-29 20:32 ` Dave Butler
2026-09-30 11:32 ` Alexander Lobakin
0 siblings, 1 reply; 28+ messages in thread
From: Dave Butler @ 2026-09-29 20:32 UTC (permalink / raw)
To: Alexander Lobakin, Tony Nguyen
Cc: davem@davemloft.net, kuba@kernel.org, pabeni@redhat.com,
edumazet@kernel.org, andrew+netdev@lunn.ch,
netdev@vger.kernel.org, emil.s.tantilov@intel.com,
luoxuanqiang@kylinos.cn, bryan.fraschetti@canonical.com,
tristan@talencesecurity.com, tomasz.lichwala@linux.intel.com,
horms@kernel.org, stable@vger.kernel.org, Jacob Keller,
Aleksandr Loktionov
>> + max_frame = min(max_frame, LIBIE_MAX_RX_BUF_LEN);
> Shouldn't we cap it to 9728 instead?
No objection. To be clear, LIBIE_MAX_RX_BUF_LEN is 9728, so it's
the same value.
Before 5fa4caff59f2 the cap was IAVF_MAX_RXBUFFER (9728, "largest size for
single descriptor"), and that same commit replaced it with LIBIE_MAX_RX_BUF_LEN
(9728U, "The largest size for a single descriptor as per HW").
That said, I'm happy to do whatever. Just let me know your preference.
Intended recipient is public
The information contained in this electronic mail is confidential information intended only for the use of the individual(s) or entity(s) named. If the reader of the message is not the addressee (or authorized to receive for the addressee), you are hereby notified that any dissemination, distribution or copying of this communication is strictly prohibited. If you have received this communication in error, please immediately notify the sender by reply e-mail and/or by telephone and destroy the original message.
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 1/6] idpf: fix possible race on remove during a reset
2026-09-28 23:04 ` [PATCH net 1/6] idpf: fix possible race on remove during a reset Tony Nguyen
@ 2026-09-30 0:58 ` netdev-bot+sashiko
2026-10-01 23:40 ` Tantilov, Emil S
0 siblings, 1 reply; 28+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30 0:58 UTC (permalink / raw)
To: anthony.l.nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
david.butler, emil.s.tantilov, luoxuanqiang, bryan.fraschetti,
tristan, tomasz.lichwala, horms, aleksander.lobakin, stable,
jacob.e.keller, aleksandr.loktionov
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 1 (2) · Medium: 0 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] The patch removes the IDPF_REMOVE_IN_PROG early return from
idpf_stop(), which opens a double-teardown race between a userspace…
Pre-existing issues:
- [High] This problem wasn't introduced by this patch, but the patch
relies on the loop that contains it.
- [High] This problem wasn't introduced by this patch, but the patch now
sends the remove-racing-reset case through it.
--
idpf: fix possible race on remove during a reset
This patch drops the IDPF_REMOVE_IN_PROG early return from idpf_stop(),
so idpf_vport_stop() always runs on close. The goal is to let a reset
that races with remove tear down NAPI in idpf_detach_and_close(), which
runs under RTNL.
> diff --git a/drivers/net/ethernet/intel/idpf/idpf_lib.c b/drivers/net/ethernet/intel/idpf/idpf_lib.c
> index 827c795afcb62..2c148377540cd 100644
> --- a/drivers/net/ethernet/intel/idpf/idpf_lib.c
> +++ b/drivers/net/ethernet/intel/idpf/idpf_lib.c
> @@ -1033,12 +1033,8 @@ static void idpf_vport_stop(struct idpf_vport *vport, bool rtnl)
> */
> static int idpf_stop(struct net_device *netdev)
> {
> - struct idpf_netdev_priv *np = netdev_priv(netdev);
> struct idpf_vport *vport;
>
> - if (test_bit(IDPF_REMOVE_IN_PROG, np->adapter->flags))
> - return 0;
> -
[Severity: High]
Without this check, can a userspace close that runs during idpf_remove()
tear down the vport a second time?
The remove path is idpf_remove() -> idpf_vc_core_deinit() ->
idpf_deinit_task() -> idpf_vport_dealloc() -> idpf_vport_stop(vport, true).
It holds neither vport_ctrl_lock nor RTNL on entry. idpf_vport_stop() tests
the UP bit before it takes RTNL and does not test it again afterwards:
idpf_vport_stop() {
if (!test_bit(IDPF_VPORT_UP, np->state))
return;
if (rtnl)
rtnl_lock();
...
}
This is the interleaving in question:
CPU1 (rmmod) CPU2 (ip link set dev down)
__dev_close_many() holds RTNL
idpf_vport_stop(vport, true)
IDPF_VPORT_UP is set
rtnl_lock() blocks
idpf_stop()
idpf_vport_stop(vport, false)
queues/vectors released
IDPF_VPORT_UP cleared
rtnl_unlock()
teardown continues on released state
idpf_vport_intr_rel() sets q_vectors to NULL but does not change
num_q_vectors:
kfree(rsrc->q_vectors);
rsrc->q_vectors = NULL;
On the second pass, idpf_vport_intr_dis_irq_all() does:
struct idpf_q_vector *q_vector = rsrc->q_vectors;
...
for (u16 q_idx = 0; q_idx < rsrc->num_q_vectors; q_idx++)
writel(0, q_vector[q_idx].intr_reg.dyn_ctl);
Won't that dereference a NULL q_vector? The virtchnl message builders
that walk the NULLed rsrc->txq_grps/rxq_grps may also crash before that
point.
Before this patch, a close that reached idpf_stop() after remove had set
IDPF_REMOVE_IN_PROG returned 0, so only one teardown ran.
Would it help to re-check IDPF_VPORT_UP in idpf_vport_stop() after
rtnl_lock() is taken, or to hold vport_ctrl_lock around the stop in the
remove path?
> idpf_vport_ctrl_lock(netdev);
> vport = idpf_netdev_to_vport(netdev);
>
[Severity: High]
This is a pre-existing issue and was not introduced by this patch. It
matters here because this fix now depends on the loop in
idpf_detach_and_close():
for (int i = 0; i < max_vports; i++) {
struct net_device *netdev = adapter->netdevs[i];
...
if (!netif_device_present(netdev))
continue;
idpf_vc_core_init() sizes adapter->netdevs by caps.max_vports. However,
idpf_init_task() only creates netdevs up to default_num_vports. If the
device reports max_vports > default_num_vports, can the first hard reset
pass a NULL netdev to netif_device_present()?
idpf_vc_event_task() -> idpf_init_hard_reset() -> idpf_detach_and_close()
-> netif_device_present(NULL)
The other max_vports-bounded loops in the driver check for empty slots.
idpf_remove() does "if (!adapter->netdevs[i]) continue;", idpf_init_task()
checks !netdev, and idpf_decfg_netdev() stores NULL into the slot on
purpose.
It is not clear which shipping firmware reports max_vports larger than
default_num_vports. virtchnl2 defines them as separate fields, though.
This has been present since 2e281e1155fc, the commit named in Fixes:.
[Severity: High]
This is also a pre-existing issue and was not introduced by this patch.
However, the remove-racing-reset case now goes through it as well.
On a software-initiated function reset, the transaction manager is shut
down before the close path runs:
idpf_vc_event_task() {
...
func_reset:
if (adapter->xnm)
libie_ctlq_xn_shutdown(adapter->xnm);
drv_load:
set_bit(IDPF_HR_RESET_IN_PROG, adapter->flags);
idpf_init_hard_reset(adapter);
...
}
After that, libie_ctlq_xn_pop_free() refuses new transactions:
if (unlikely(xnm->shutdown))
return NULL;
The call chain is idpf_init_hard_reset() -> idpf_detach_and_close() ->
dev_close() -> idpf_stop() -> idpf_vport_stop(). The DISABLE_VPORT and
DISABLE_QUEUES messages fail silently because their return values are
ignored. Then the rings and buffers are freed:
idpf_vport_intr_deinit(vport, rsrc);
idpf_xdp_rxq_info_deinit_all(rsrc);
idpf_vport_queues_rel(vport, rsrc);
idpf_vport_intr_rel(rsrc);
reg_ops->trigger_reset() is only called later, after
idpf_vc_core_deinit(). Can the still-enabled Rx queues DMA into the freed
descriptor rings or buffers during that window?
On a VF, idpf_vf_trigger_reset() skips the reset entirely during remove:
if (trig_cause == IDPF_HR_FUNC_RESET &&
!test_bit(IDPF_REMOVE_IN_PROG, adapter->flags))
idpf_send_vf_reset_msg(adapter);
In that case, no device reset follows the free at all.
Before this patch, the remove-racing-reset case skipped idpf_vport_stop()
in idpf_stop(). The rings leaked, but they were not freed while still in
use.
Would it help to shut down the transaction manager only after
idpf_detach_and_close() has disabled the queues? Another option is to
keep the DMA memory until the device has been reset.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928230429.495442-1-anthony.l.nguyen%40intel.com
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells
2026-09-28 23:04 ` [PATCH net 3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells Tony Nguyen
@ 2026-09-30 0:58 ` netdev-bot+sashiko
2026-10-01 16:35 ` Tony Nguyen
0 siblings, 1 reply; 28+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30 0:58 UTC (permalink / raw)
To: anthony.l.nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
david.butler, emil.s.tantilov, luoxuanqiang, bryan.fraschetti,
tristan, tomasz.lichwala, horms, aleksander.lobakin, stable,
jacob.e.keller, aleksandr.loktionov
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] The patch fixes only the writer side of the ordering chain
behind the dql_completed() BUG_ON.
--
ice: Restore Ordered MMIO Writes for Tx Doorbells
This commit changes the Tx ring and tstamp ring tail doorbell writes in
ice_tx_map() from writel_relaxed() back to writel(). The goal is to make
the dql->num_queued update from __netdev_tx_sent_queue() visible before
the NIC can start on the new descriptors. That avoids the BUG_ON in
dql_completed() seen on arm64.
Link: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2161572
> diff --git a/drivers/net/ethernet/intel/ice/ice_txrx.c b/drivers/net/ethernet/intel/ice/ice_txrx.c
> index 31303ab5be175..a2c7c4962882d 100644
> --- a/drivers/net/ethernet/intel/ice/ice_txrx.c
> +++ b/drivers/net/ethernet/intel/ice/ice_txrx.c
> @@ -1561,10 +1561,10 @@ ice_tx_map(struct ice_tx_ring *tx_ring, struct ice_tx_buf *first,
> }
> }
> tstamp_ring->next_to_use = j;
> - writel_relaxed(j, tstamp_ring->tail);
> + writel(j, tstamp_ring->tail);
> } else {
> ring_kick:
> - writel_relaxed(i, tx_ring->tail);
> + writel(i, tx_ring->tail);
> }
> return;
[Severity: Medium]
This orders the xmit side. Is there a matching read barrier on the
completion side?
ice_clean_tx_irq() reads the DONE bit that the device writes, and then
completes the queue:
drivers/net/ethernet/intel/ice/ice_txrx.c:ice_clean_tx_irq() {
...
if (!(eop_desc->cmd_type_offset_bsz &
cpu_to_le64(ICE_TX_DESC_DTYPE_DESC_DONE)))
break;
...
netdev_tx_completed_queue(txring_txq(tx_ring), total_pkts, total_bytes);
...
}
netdev_tx_completed_queue() then calls dql_completed(), which uses a
relaxed load:
lib/dynamic_queue_limits.c:dql_completed() {
num_queued = READ_ONCE(dql->num_queued);
...
BUG_ON(count > num_queued - dql->num_completed);
...
}
The smp_mb() in netdev_tx_completed_queue() runs after dql_completed(), so
it does not order this load. On arm64, a control dependency on the DONE
bit does not order a later load either.
The barriers that happen to sit on this path are not always there:
- The smp_rmb() at the top of the ice_clean_tx_irq() loop runs again only
when the next tx_buf has a non-NULL next_to_watch. It is skipped when
the loop stops on a NULL next_to_watch or when the budget runs out.
- skb_unref() issues smp_rmb() only when skb->users == 1 and
CONFIG_DEBUG_NET is off. With users > 1 (for example pktgen with burst,
which does refcount_add(burst, &skb->users)), napi_consume_skb() skips
the defer path. The failing refcount_dec_and_test() then gives only
release ordering.
- u64_stats_update_begin()/end() in ice_update_tx_ring_stats() are no-ops
on 64-bit. dma_unmap_single() does nothing for coherent dma-direct
without an IOMMU.
On that path, can the num_queued load complete before the DONE bit load?
If so, it would return a value from before the xmit CPU's dql_queued()
update and hit the same BUG_ON this commit is meant to fix.
The window is probably very small in practice. The num_queued store comes
before the doorbell, the descriptor fetch, the transmit and the DONE
write-back, and testing shows the observed crash is gone. This reader-side
pattern is also older than both this patch and commit ccde82e90946.
Would a dma_rmb() or smp_rmb() in ice_clean_tx_irq() help close this
fully? It could go after the DONE check or before the
netdev_tx_completed_queue() call.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928230429.495442-1-anthony.l.nguyen%40intel.com
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 5/6] iavf: fix VF stats not updating due to PTP command preemption
2026-09-28 23:04 ` [PATCH net 5/6] iavf: fix VF stats not updating due to PTP command preemption Tony Nguyen
@ 2026-09-30 0:58 ` netdev-bot+sashiko
2026-09-30 15:07 ` Tomasz Lichwala
0 siblings, 1 reply; 28+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30 0:58 UTC (permalink / raw)
To: anthony.l.nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
david.butler, emil.s.tantilov, luoxuanqiang, bryan.fraschetti,
tristan, tomasz.lichwala, horms, aleksander.lobakin, stable,
jacob.e.keller, aleksandr.loktionov
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] iavf_watchdog_step()
(drivers/net/ethernet/intel/iavf/iavf_main.c) now calls…
--
iavf: fix VF stats not updating due to PTP command preemption
With this commit, iavf_watchdog_step() requests VF stats on every pass
while the adapter is in __IAVF_RUNNING. Before, it only did so when
iavf_process_aq_command() found nothing to do. The periodic PHC caching
task keeps a PTP command pending, so the old stats fallback never ran.
> diff --git a/drivers/net/ethernet/intel/iavf/iavf_main.c b/drivers/net/ethernet/intel/iavf/iavf_main.c
> index 29b8403a066bc..c0686ad5c4116 100644
> --- a/drivers/net/ethernet/intel/iavf/iavf_main.c
> +++ b/drivers/net/ethernet/intel/iavf/iavf_main.c
> @@ -2932,18 +2932,12 @@ static int iavf_watchdog_step(struct iavf_adapter *adapter)
> iavf_send_api_ver(adapter);
> }
> } else {
> - int ret = iavf_process_aq_command(adapter);
> -
> - /* An error will be returned if no commands were
> - * processed; use this opportunity to update stats
> - * if the error isn't -ENOTSUPP
> - */
> - if (ret && ret != -EOPNOTSUPP &&
> - adapter->state == __IAVF_RUNNING)
> - iavf_request_stats(adapter);
> + iavf_process_aq_command(adapter);
> }
> - if (adapter->state == __IAVF_RUNNING)
> + if (adapter->state == __IAVF_RUNNING) {
> + iavf_request_stats(adapter);
[Severity: Medium]
Can this put two virtchnl messages in flight during one watchdog pass?
Suppose iavf_process_aq_command() has just taken the
IAVF_FLAG_AQ_SEND_PTP_CMD branch. In that case
iavf_virtchnl_send_ptp_cmd() has sent VIRTCHNL_OP_1588_PTP_GET_TIME
without setting adapter->current_op:
iavf_virtchnl.c:iavf_virtchnl_send_ptp_cmd() {
...
err = iavf_send_pf_msg(adapter, cmd->v_opcode, cmd->msg, cmd->msglen);
if (!err) {
list_del(&cmd->list);
kfree(cmd);
...
}
The only guard in iavf_request_stats() then still sees an idle channel,
so VIRTCHNL_OP_GET_STATS goes out right behind GET_TIME:
if (adapter->current_op != VIRTCHNL_OP_UNKNOWN) {
/* no error message, this isn't crucial */
return;
}
...
adapter->current_op = VIRTCHNL_OP_GET_STATS;
When the GET_TIME reply arrives, iavf_virtchnl_completion() breaks out
of the switch and reaches the unconditional reset at the end:
} /* switch v_opcode */
adapter->current_op = VIRTCHNL_OP_UNKNOWN;
That clears the GET_STATS tracking while the GET_STATS reply is still
outstanding.
iavf_adminq_task() may drop the netdev lock before the GET_STATS reply
arrives. A watchdog pass can then send another tracked command. The
watchdog can be kicked by iavf_schedule_aq_request(), by the 20ms re-arm
while aq_required is set, or by gettimex64() queuing another PTP read
with mod_delayed_work(..., 0). The late GET_STATS reply would then clear
that command's current_op as well.
Flow Director is one place where this could do real damage.
iavf_add_fdir_filter() relies on current_op to keep only one
ADD_PENDING filter in flight. If two adds overlap, the first successful
VIRTCHNL_OP_ADD_FDIR_FILTER reply does this for every pending filter:
if (fdir->state == IAVF_FDIR_FLTR_ADD_PENDING) {
if (add_fltr->status == VIRTCHNL_FDIR_SUCCESS) {
...
fdir->state = IAVF_FDIR_FLTR_ACTIVE;
fdir->flow_id = add_fltr->flow_id;
Would the second filter then get the wrong PF flow_id? If so, a later
delete of that filter could remove the wrong PF rule and leak the real
one.
The missing current_op tracking in the PTP sender predates this patch.
Before this change, though, GET_STATS was only sent in a pass where
iavf_process_aq_command() returned an error, so it never followed a PTP
message back to back. Now every PTP send is followed by a GET_STATS in
the same pass. That covers the roughly once-per-second PHC cache update
and every gettimex64() read.
This also adds one GET_STATS mailbox message per PHC read, which the
commit message doesn't mention.
> iavf_detect_recover_hung(&adapter->vsi);
> + }
> break;
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928230429.495442-1-anthony.l.nguyen%40intel.com
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit
2026-09-28 23:04 ` [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit Tony Nguyen
2026-09-29 16:18 ` Alexander Lobakin
@ 2026-09-30 0:58 ` netdev-bot+sashiko
2026-09-30 6:02 ` Dave Butler
1 sibling, 1 reply; 28+ messages in thread
From: netdev-bot+sashiko @ 2026-09-30 0:58 UTC (permalink / raw)
To: anthony.l.nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
david.butler, emil.s.tantilov, luoxuanqiang, bryan.fraschetti,
tristan, tomasz.lichwala, horms, aleksander.lobakin, stable,
jacob.e.keller, aleksandr.loktionov
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] The patch caps the Rx max_pkt_size sent to the PF (the per-
queue Rx context 'rxmax') at LIBIE_MAX_RX_BUF_LEN (9728) in…
--
iavf: cap advertised max_pkt_size at the single-buffer HW limit
This commit caps the max_pkt_size that iavf_configure_queues() sends to
the PF in VIRTCHNL_OP_CONFIG_VSI_QUEUES at LIBIE_MAX_RX_BUF_LEN (9728).
Without the cap, a PF that advertises vf_res->max_mtu as 0 (the ESXi icen
PF on E810) receives 16382 and rejects the queue configuration.
Link: https://community.intel.com/t5/Ethernet-Products/E810-C-iavf-driver-issue-on-Linux-6-12/m-p/1737490
Link: https://access.redhat.com/solutions/6973766
Link: https://knowledge.broadcom.com/external/article/404315/sriov-enabled-vms-network-adaptor-goes-d.html
> diff --git a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
> index ec234cc8bd9db..680a28a739bf0 100644
> --- a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
> +++ b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
> @@ -382,6 +382,14 @@ void iavf_configure_queues(struct iavf_adapter *adapter)
>
> max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
> max_frame = min_not_zero(adapter->vf_res->max_mtu, max_frame);
> + /* The PF programs max_pkt_size into the per-queue Rx context "rxmax".
> + * LIBIE_MAX_RX_FRM_LEN is the multi-descriptor (S/G) frame ceiling
> + * (16382), but that exceeds the E810 max MAC frame size (9728); some
> + * PFs reject the out-of-range value with VIRTCHNL_STATUS_ERR_PARAM.
> + * Cap it at the single-buffer HW limit (== the MAC frame max),
> + * restoring the pre-Page-Pool behaviour.
> + */
> + max_frame = min(max_frame, LIBIE_MAX_RX_BUF_LEN);
[Severity: Medium]
Should netdev->max_mtu be clamped to match? Rx max_pkt_size is now capped
at 9728. However, iavf_init_config_adapter() in iavf_main.c still derives
the MTU ceiling from the 16382 S/G limit when the PF sends max_mtu == 0:
drivers/net/ethernet/intel/iavf/iavf_main.c:iavf_init_config_adapter() {
...
if (!adapter->vf_res->max_mtu) {
netdev->max_mtu = LIBIE_MAX_MTU;
} else if (adapter->vf_res->max_mtu < LIBETH_RX_LL_LEN + ETH_MIN_MTU ||
adapter->vf_res->max_mtu >
LIBETH_RX_LL_LEN + LIBIE_MAX_MTU) {
...
netdev->max_mtu = LIBIE_MAX_MTU;
} else {
netdev->max_mtu = adapter->vf_res->max_mtu - LIBETH_RX_LL_LEN;
}
...
}
LIBIE_MAX_MTU is 16382 - 26 = 16356. A PF-advertised frame size above
9728 that still passes the range check also gives a max_mtu above 9702.
In the icen case this patch targets, can a user still set an MTU the
queues can't carry? For example:
ip link set dev ethX mtu 12000
dev_validate_mtu() only checks against dev->max_mtu. iavf_change_mtu()
stores the new MTU and triggers a reset without a check of its own. The
queues are then reprogrammed with max_pkt_size = 9728.
Received frames larger than 9728 bytes would then be silently dropped by
the Rx queue context, instead of the MTU change failing. Per the commit
message, E810 also discards Tx frames above MAX_TXFRAME (9728).
Before this patch, rxmax and max_mtu both came from the 16382 ceiling and
matched each other.
The commit message and the new comment both describe this as "restoring
the pre-Page-Pool behaviour". Commit 5fa4caff59f2 changed two things. It
changed max_frame from IAVF_MAX_RXBUFFER to LIBIE_MAX_RX_FRM_LEN(). It
also changed netdev->max_mtu from IAVF_MAX_RXBUFFER - IAVF_PACKET_HDR_PAD
(9710) to LIBIE_MAX_MTU. This patch only reverts the first change.
Could netdev->max_mtu also be clamped to LIBIE_MAX_RX_BUF_LEN -
LIBETH_RX_LL_LEN, in both the zero-value branch and the invalid-value
fallback?
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260928230429.495442-1-anthony.l.nguyen%40intel.com
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit
2026-09-30 0:58 ` netdev-bot+sashiko
@ 2026-09-30 6:02 ` Dave Butler
[not found] ` <IA3PR05MB22078430E404FAD12912A706298B892@IA3PR05MB220784.namprd05.prod.outlook.com>
0 siblings, 1 reply; 28+ messages in thread
From: Dave Butler @ 2026-09-30 6:02 UTC (permalink / raw)
To: netdev-bot+sashiko@kernel.org, anthony.l.nguyen@intel.com
Cc: davem@davemloft.net, kuba@kernel.org, pabeni@redhat.com,
edumazet@kernel.org, andrew+netdev@lunn.ch,
netdev@vger.kernel.org, emil.s.tantilov@intel.com,
luoxuanqiang@kylinos.cn, bryan.fraschetti@canonical.com,
tristan@talencesecurity.com, tomasz.lichwala@linux.intel.com,
horms@kernel.org, aleksander.lobakin@intel.com,
stable@vger.kernel.org, jacob.e.keller@intel.com,
aleksandr.loktionov@intel.com
> [Severity: Medium]
> Should netdev->max_mtu be clamped to match?
Yes, the review feedback is valid. The bug outlined by the review is where the user sets an MTU in the 9728 to ~16356 range and frames above 9728 are silently dropped instead of the MTU set failing.
We should still move forward with this patch but reword the commit message. It will likely be at least several weeks before I will be able to acquire and setup a lab environment for reproduction and testing for a new patch revision. This patch as-is does not introduce any code regressions. In every case the behavior is unchanged or strictly better. Practically speaking, most users would not be impacted by the frame drop. Standard jumbo (≤9728, e.g. 9000) is entirely unaffected.
Since 5fa4caff59f2, E810 + ESXi icen + SR-IOV has been unusable: the VF never finishes CONFIG_VSI_QUEUES, regardless of MTU. The effects are not contained to the guest that triggers it. Under pass-through the mis-programmed queue can raise an IOMMU fault that wedges the PF/VF on the ESXi host. This usually takes out that interface for subsequent guests too (even if they have this patch). Recovery requires a host reboot.
A follow-up to adopt the reviewer's fix is warranted. I can provide the patch if someone else thinks they can test it sooner than I can.
A revised commit message follows:
iavf: cap advertised max_pkt_size at the single-buffer HW limit
Since commit 5fa4caff59f2 ("iavf: switch to Page Pool")
iavf_configure_queues() advertises max_pkt_size to the PF as:
max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
max_frame = min_not_zero(adapter->vf_res->max_mtu, max_frame);
LIBIE_MAX_RX_FRM_LEN (16382) is the multi-descriptor scatter/gather frame
ceiling, not a single-queue value, and it exceeds the E810 MAC frame size
maximum of 9728. Per the E810 datasheet (613875-009 section 13.2.2.17.1)
the Tx frame-size register PRTDCB_TDPUC.MAX_TXFRAME has a maximum of
0x2600 (9728); larger frames are discarded. The in-tree ice driver encodes
the same value as ICE_AQ_SET_MAC_FRAME_SIZE_MAX (== LIBIE_MAX_RX_BUF_LEN ==
9728), and the VF clamped max_frame to IAVF_MAX_RXBUFFER (9728) before this
commit.
When the PF advertises vf_res->max_mtu as 0, min_not_zero() leaves
max_frame at 16382. The Linux ice PF advertises max_mtu = port MAC frame
size (<= 9728), so a VF behind ice never sends more than that. The ESXi
"icen" PF on E810 advertises max_mtu as 0, so the VF sends
max_pkt_size = 16382, which icen rejects while programming the queue
context for VIRTCHNL_OP_CONFIG_VSI_QUEUES (opcode 6):
icen_ConfigureTxQueue: VSI 8: Failed to set LAN Tx queue context for
absolute Tx queue 64, Error: ICE_ERR_PARAM
indrv_SendMsgToVf: VF 0: Failed opcode 6, Error -5
iavf 0000:03:00.0: PF returned error -5 (IAVF_ERR_PARAM) to our request 6
iavf 0000:03:00.0 ethX: NETDEV WATCHDOG: transmit queue N timed out
The VF's queues never come up; under SR-IOV passthrough the mis-programmed
queue can also trigger a fatal IOMMU fault in the guest. Forcing only
max_pkt_size back to 9728 (and leaving the Page Pool rx_buf_len/
databuffer_size untouched) makes the VF come up; databuffer_size is not
involved. This was confirmed on two E810 NVM versions (3.00 and 4.51) under
icen, so the trigger is the icen PF, not the firmware revision. Reported by
several users on E810 + ESXi icen with v6.10+ guests:
Link: https://community.intel.com/t5/Ethernet-Products/E810-C-iavf-driver-issue-on-Linux-6-12/m-p/1737490
Link: https://access.redhat.com/solutions/6973766
Link: https://knowledge.broadcom.com/external/article/404315/sriov-enabled-vms-network-adaptor-goes-d.html
Cap max_frame at the single-buffer HW limit (LIBIE_MAX_RX_BUF_LEN, 9728) so
the VF advertises a max_pkt_size the PF can program into the Rx queue
context. This is the value the VF used before 5fa4caff59f2 (IAVF_MAX_RXBUFFER).
Note this only caps the advertised max_pkt_size. 5fa4caff59f2 also raised
netdev->max_mtu to LIBIE_MAX_MTU (derived from the same 16382 ceiling); when
the PF advertises max_mtu == 0 that leaves max_mtu above 9728, so an MTU in
the 9728 < mtu <= LIBIE_MAX_MTU range can still be set while frames above
9728 are dropped by the queue context. Clamping netdev->max_mtu to match is
left to a follow-up.
Fixes: 5fa4caff59f2 ("iavf: switch to Page Pool")
Signed-off-by: Dave Butler <david.butler@appgate.com>
Cc: stable@vger.kernel.org # v6.10+
.....
The information contained in this electronic mail is confidential information intended only for the use of the individual(s) or entity(s) named. If the reader of the message is not the addressee (or authorized to receive for the addressee), you are hereby notified that any dissemination, distribution or copying of this communication is strictly prohibited. If you have received this communication in error, please immediately notify the sender by reply e-mail and/or by telephone and destroy the original message.
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit
2026-09-29 20:32 ` Dave Butler
@ 2026-09-30 11:32 ` Alexander Lobakin
0 siblings, 0 replies; 28+ messages in thread
From: Alexander Lobakin @ 2026-09-30 11:32 UTC (permalink / raw)
To: Dave Butler
Cc: Tony Nguyen, davem@davemloft.net, kuba@kernel.org,
pabeni@redhat.com, edumazet@kernel.org, andrew+netdev@lunn.ch,
netdev@vger.kernel.org, emil.s.tantilov@intel.com,
luoxuanqiang@kylinos.cn, bryan.fraschetti@canonical.com,
tristan@talencesecurity.com, tomasz.lichwala@linux.intel.com,
horms@kernel.org, stable@vger.kernel.org, Jacob Keller,
Aleksandr Loktionov
From: Dave Butler <david.butler@appgate.com>
Date: Tue, 29 Sep 2026 20:32:53 +0000
>>> + max_frame = min(max_frame, LIBIE_MAX_RX_BUF_LEN);
>> Shouldn't we cap it to 9728 instead?
>
> No objection. To be clear, LIBIE_MAX_RX_BUF_LEN is 9728, so it's
> the same value.
Oops sorry, I did misread. I thought for some reason that you now limit
it to (4096 - hr/tr overhead) >_<
Acked-by: Alexander Lobakin <aleksander.lobakin@intel.com>
>
> Before 5fa4caff59f2 the cap was IAVF_MAX_RXBUFFER (9728, "largest size for
> single descriptor"), and that same commit replaced it with LIBIE_MAX_RX_BUF_LEN
> (9728U, "The largest size for a single descriptor as per HW").
>
> That said, I'm happy to do whatever. Just let me know your preference.
>
> Intended recipient is public
> The information contained in this electronic mail is confidential information intended only for the use of the individual(s) or entity(s) named. If the reader of the message is not the addressee (or authorized to receive for the addressee), you are hereby notified that any dissemination, distribution or copying of this communication is strictly prohibited. If you have received this communication in error, please immediately notify the sender by reply e-mail and/or by telephone and destroy the original message.
Thanks,
Olek
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
2026-09-28 23:10 ` [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) netdev-bot+sinfo
2026-09-29 1:41 ` Dave Butler
2026-09-29 17:16 ` Tantilov, Emil S
@ 2026-09-30 15:06 ` Tomasz Lichwala
2 siblings, 0 replies; 28+ messages in thread
From: Tomasz Lichwala @ 2026-09-30 15:06 UTC (permalink / raw)
To: netdev-bot+sinfo, Tony Nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
emil.s.tantilov, luoxuanqiang, bryan.fraschetti, tristan,
david.butler, horms
On 29.09.2026 01:10, netdev-bot+sinfo@kernel.org wrote:
> Hi!
>
> This is an automated message. This series looks like a fix, but its
> commit messages seem to be missing some information:
>
> - How the issue was discovered, e.g. hit in production, hit during
> development, syzbot report, manual code inspection, LLM or static
> analysis tool scan.
>
> - Whether the issue was actually triggered, or is only theoretical
> (e.g. found by code inspection). If it was triggered please include
> the symptoms, like the stack trace or error messages.
>
> - What hardware the change was tested on. For driver fixes please
> mention the device (and if relevant firmware version) used for
> testing, or say that the change was not tested on real hardware.
>
> Please do not repost the series just to address the above. Instead,
> reply to this email with the missing information, so that reviewers
> can take it into account. If the series needs another revision for
> other reasons, please include the information in the commit messages
> then.
>
> The evaluation is done by an LLM so it may be wrong, if you think
> that is the case please reply and explain.
This information will be included in patch 5/6 iavf: fix VF stats not updating due to PTP command preemption:
How discovered: found while investigating a bug report where VF statistics in sysfs stayed at zero despite working traffic.
Triggered: yes, actually reproduced — ping/iperf succeeded while rx_packets/tx_packets stayed at 0 in both sysfs and ethtool -S.
Hardware tested: Intel E810-C (ice PF) with an iavf VF passed through to a KVM guest. The fix was validated on the same setup by swapping an unpatched vs. patched iavf.ko in the same guest kernel.
Thanks,
Tomasz
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 5/6] iavf: fix VF stats not updating due to PTP command preemption
2026-09-30 0:58 ` netdev-bot+sashiko
@ 2026-09-30 15:07 ` Tomasz Lichwala
0 siblings, 0 replies; 28+ messages in thread
From: Tomasz Lichwala @ 2026-09-30 15:07 UTC (permalink / raw)
To: netdev-bot+sashiko, anthony.l.nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
david.butler, emil.s.tantilov, luoxuanqiang, bryan.fraschetti,
tristan, horms, aleksander.lobakin, stable, jacob.e.keller,
aleksandr.loktionov
On 30.09.2026 02:58, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 0
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [Medium] iavf_watchdog_step()
> (drivers/net/ethernet/intel/iavf/iavf_main.c) now calls…
>
> --
>
> iavf: fix VF stats not updating due to PTP command preemption
>
> With this commit, iavf_watchdog_step() requests VF stats on every pass
> while the adapter is in __IAVF_RUNNING. Before, it only did so when
> iavf_process_aq_command() found nothing to do. The periodic PHC caching
> task keeps a PTP command pending, so the old stats fallback never ran.
>
>> diff --git a/drivers/net/ethernet/intel/iavf/iavf_main.c b/drivers/net/ethernet/intel/iavf/iavf_main.c
>> index 29b8403a066bc..c0686ad5c4116 100644
>> --- a/drivers/net/ethernet/intel/iavf/iavf_main.c
>> +++ b/drivers/net/ethernet/intel/iavf/iavf_main.c
>> @@ -2932,18 +2932,12 @@ static int iavf_watchdog_step(struct iavf_adapter *adapter)
>> iavf_send_api_ver(adapter);
>> }
>> } else {
>> - int ret = iavf_process_aq_command(adapter);
>> -
>> - /* An error will be returned if no commands were
>> - * processed; use this opportunity to update stats
>> - * if the error isn't -ENOTSUPP
>> - */
>> - if (ret && ret != -EOPNOTSUPP &&
>> - adapter->state == __IAVF_RUNNING)
>> - iavf_request_stats(adapter);
>> + iavf_process_aq_command(adapter);
>> }
>> - if (adapter->state == __IAVF_RUNNING)
>> + if (adapter->state == __IAVF_RUNNING) {
>> + iavf_request_stats(adapter);
>
> [Severity: Medium]
> Can this put two virtchnl messages in flight during one watchdog pass?
>
> Suppose iavf_process_aq_command() has just taken the
> IAVF_FLAG_AQ_SEND_PTP_CMD branch. In that case
> iavf_virtchnl_send_ptp_cmd() has sent VIRTCHNL_OP_1588_PTP_GET_TIME
> without setting adapter->current_op:
>
> iavf_virtchnl.c:iavf_virtchnl_send_ptp_cmd() {
> ...
> err = iavf_send_pf_msg(adapter, cmd->v_opcode, cmd->msg, cmd->msglen);
> if (!err) {
> list_del(&cmd->list);
> kfree(cmd);
> ...
> }
>
> The only guard in iavf_request_stats() then still sees an idle channel,
> so VIRTCHNL_OP_GET_STATS goes out right behind GET_TIME:
>
> if (adapter->current_op != VIRTCHNL_OP_UNKNOWN) {
> /* no error message, this isn't crucial */
> return;
> }
> ...
> adapter->current_op = VIRTCHNL_OP_GET_STATS;
>
> When the GET_TIME reply arrives, iavf_virtchnl_completion() breaks out
> of the switch and reaches the unconditional reset at the end:
>
> } /* switch v_opcode */
> adapter->current_op = VIRTCHNL_OP_UNKNOWN;
>
> That clears the GET_STATS tracking while the GET_STATS reply is still
> outstanding.
>
> iavf_adminq_task() may drop the netdev lock before the GET_STATS reply
> arrives. A watchdog pass can then send another tracked command. The
> watchdog can be kicked by iavf_schedule_aq_request(), by the 20ms re-arm
> while aq_required is set, or by gettimex64() queuing another PTP read
> with mod_delayed_work(..., 0). The late GET_STATS reply would then clear
> that command's current_op as well.
>
> Flow Director is one place where this could do real damage.
> iavf_add_fdir_filter() relies on current_op to keep only one
> ADD_PENDING filter in flight. If two adds overlap, the first successful
> VIRTCHNL_OP_ADD_FDIR_FILTER reply does this for every pending filter:
>
> if (fdir->state == IAVF_FDIR_FLTR_ADD_PENDING) {
> if (add_fltr->status == VIRTCHNL_FDIR_SUCCESS) {
> ...
> fdir->state = IAVF_FDIR_FLTR_ACTIVE;
> fdir->flow_id = add_fltr->flow_id;
>
> Would the second filter then get the wrong PF flow_id? If so, a later
> delete of that filter could remove the wrong PF rule and leak the real
> one.
>
> The missing current_op tracking in the PTP sender predates this patch.
> Before this change, though, GET_STATS was only sent in a pass where
> iavf_process_aq_command() returned an error, so it never followed a PTP
> message back to back. Now every PTP send is followed by a GET_STATS in
> the same pass. That covers the roughly once-per-second PHC cache update
> and every gettimex64() read.
>
> This also adds one GET_STATS mailbox message per PHC read, which the
> commit message doesn't mention.
>
>> iavf_detect_recover_hung(&adapter->vsi);
>> + }
>> break;
>
> [ ... ]
>
Thank you, this is a valid finding. iavf_virtchnl_send_ptp_cmd() does not set adapter->current_op before sending VIRTCHNL_OP_1588_PTP_GET_TIME, unlike every other virtchnl sender, and iavf_virtchnl_completion() unconditionally clears current_op on any reply. This gap predates this patch, but making the stats request unconditional every watchdog pass makes it trivially reachable, since GET_STATS can now follow a PTP send back-to-back in the same pass - something the old fallback design structurally prevented.
This will be fixed.
Thanks,
Tomasz
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells
2026-09-30 0:58 ` netdev-bot+sashiko
@ 2026-10-01 16:35 ` Tony Nguyen
0 siblings, 0 replies; 28+ messages in thread
From: Tony Nguyen @ 2026-10-01 16:35 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
david.butler, emil.s.tantilov, luoxuanqiang, bryan.fraschetti,
tristan, tomasz.lichwala, horms, aleksander.lobakin, stable,
jacob.e.keller, aleksandr.loktionov
On 9/29/2026 5:58 PM, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 0
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [Medium] The patch fixes only the writer side of the ordering chain
> behind the dql_completed() BUG_ON.
>
> --
>
> ice: Restore Ordered MMIO Writes for Tx Doorbells
>
> This commit changes the Tx ring and tstamp ring tail doorbell writes in
> ice_tx_map() from writel_relaxed() back to writel(). The goal is to make
> the dql->num_queued update from __netdev_tx_sent_queue() visible before
> the NIC can start on the new descriptors. That avoids the BUG_ON in
> dql_completed() seen on arm64.
>
> Link: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2161572
>
>> diff --git a/drivers/net/ethernet/intel/ice/ice_txrx.c b/drivers/net/ethernet/intel/ice/ice_txrx.c
>> index 31303ab5be175..a2c7c4962882d 100644
>> --- a/drivers/net/ethernet/intel/ice/ice_txrx.c
>> +++ b/drivers/net/ethernet/intel/ice/ice_txrx.c
>> @@ -1561,10 +1561,10 @@ ice_tx_map(struct ice_tx_ring *tx_ring, struct ice_tx_buf *first,
>> }
>> }
>> tstamp_ring->next_to_use = j;
>> - writel_relaxed(j, tstamp_ring->tail);
>> + writel(j, tstamp_ring->tail);
>> } else {
>> ring_kick:
>> - writel_relaxed(i, tx_ring->tail);
>> + writel(i, tx_ring->tail);
>> }
>> return;
>
> [Severity: Medium]
>
> This orders the xmit side. Is there a matching read barrier on the
> completion side?
>
> ice_clean_tx_irq() reads the DONE bit that the device writes, and then
> completes the queue:
>
> drivers/net/ethernet/intel/ice/ice_txrx.c:ice_clean_tx_irq() {
> ...
> if (!(eop_desc->cmd_type_offset_bsz &
> cpu_to_le64(ICE_TX_DESC_DTYPE_DESC_DONE)))
> break;
> ...
> netdev_tx_completed_queue(txring_txq(tx_ring), total_pkts, total_bytes);
> ...
> }
>
> netdev_tx_completed_queue() then calls dql_completed(), which uses a
> relaxed load:
>
> lib/dynamic_queue_limits.c:dql_completed() {
> num_queued = READ_ONCE(dql->num_queued);
> ...
> BUG_ON(count > num_queued - dql->num_completed);
> ...
> }
>
> The smp_mb() in netdev_tx_completed_queue() runs after dql_completed(), so
> it does not order this load. On arm64, a control dependency on the DONE
> bit does not order a later load either.
>
> The barriers that happen to sit on this path are not always there:
>
> - The smp_rmb() at the top of the ice_clean_tx_irq() loop runs again only
> when the next tx_buf has a non-NULL next_to_watch. It is skipped when
> the loop stops on a NULL next_to_watch or when the budget runs out.
>
> - skb_unref() issues smp_rmb() only when skb->users == 1 and
> CONFIG_DEBUG_NET is off. With users > 1 (for example pktgen with burst,
> which does refcount_add(burst, &skb->users)), napi_consume_skb() skips
> the defer path. The failing refcount_dec_and_test() then gives only
> release ordering.
>
> - u64_stats_update_begin()/end() in ice_update_tx_ring_stats() are no-ops
> on 64-bit. dma_unmap_single() does nothing for coherent dma-direct
> without an IOMMU.
>
> On that path, can the num_queued load complete before the DONE bit load?
> If so, it would return a value from before the xmit CPU's dql_queued()
> update and hit the same BUG_ON this commit is meant to fix.
>
> The window is probably very small in practice. The num_queued store comes
> before the doorbell, the descriptor fetch, the transmit and the DONE
> write-back, and testing shows the observed crash is gone. This reader-side
> pattern is also older than both this patch and commit ccde82e90946.
Seems this is pre-existing as the review states this is older than the
patch and cited Fixes commit. Sashiko points out that this "window is
probably very small in practice." This patch is essentially reverting
the change from writel_relaxed() to writel() that was done in the cited
Fixes: commit and existed prior to it. Can we take this as-is as it's
resolving a real issue and this situation seems unlikely as stated by AI
and we have not had any reports of this.
Thanks,
Tony
> Would a dma_rmb() or smp_rmb() in ice_clean_tx_irq() help close this
> fully? It could go after the DONE check or before the
> netdev_tx_completed_queue() call.
>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 1/6] idpf: fix possible race on remove during a reset
2026-09-30 0:58 ` netdev-bot+sashiko
@ 2026-10-01 23:40 ` Tantilov, Emil S
0 siblings, 0 replies; 28+ messages in thread
From: Tantilov, Emil S @ 2026-10-01 23:40 UTC (permalink / raw)
To: netdev-bot+sashiko, anthony.l.nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
david.butler, luoxuanqiang, bryan.fraschetti, tristan,
tomasz.lichwala, horms, aleksander.lobakin, stable,
jacob.e.keller, aleksandr.loktionov
On 9/29/2026 5:58 PM, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
>
> Critical: 0 · High: 1 (2) · Medium: 0 · Low: 0
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [High] The patch removes the IDPF_REMOVE_IN_PROG early return from
> idpf_stop(), which opens a double-teardown race between a userspace…
>
> Pre-existing issues:
> - [High] This problem wasn't introduced by this patch, but the patch
> relies on the loop that contains it.
> - [High] This problem wasn't introduced by this patch, but the patch now
> sends the remove-racing-reset case through it.
>
> --
>
> idpf: fix possible race on remove during a reset
>
> This patch drops the IDPF_REMOVE_IN_PROG early return from idpf_stop(),
> so idpf_vport_stop() always runs on close. The goal is to let a reset
> that races with remove tear down NAPI in idpf_detach_and_close(), which
> runs under RTNL.
>
>> diff --git a/drivers/net/ethernet/intel/idpf/idpf_lib.c b/drivers/net/ethernet/intel/idpf/idpf_lib.c
>> index 827c795afcb62..2c148377540cd 100644
>> --- a/drivers/net/ethernet/intel/idpf/idpf_lib.c
>> +++ b/drivers/net/ethernet/intel/idpf/idpf_lib.c
>> @@ -1033,12 +1033,8 @@ static void idpf_vport_stop(struct idpf_vport *vport, bool rtnl)
>> */
>> static int idpf_stop(struct net_device *netdev)
>> {
>> - struct idpf_netdev_priv *np = netdev_priv(netdev);
>> struct idpf_vport *vport;
>>
>> - if (test_bit(IDPF_REMOVE_IN_PROG, np->adapter->flags))
>> - return 0;
>> -
>
> [Severity: High]
> Without this check, can a userspace close that runs during idpf_remove()
> tear down the vport a second time?
>
> The remove path is idpf_remove() -> idpf_vc_core_deinit() ->
> idpf_deinit_task() -> idpf_vport_dealloc() -> idpf_vport_stop(vport, true).
> It holds neither vport_ctrl_lock nor RTNL on entry. idpf_vport_stop() tests
> the UP bit before it takes RTNL and does not test it again afterwards:
>
> idpf_vport_stop() {
> if (!test_bit(IDPF_VPORT_UP, np->state))
> return;
>
> if (rtnl)
> rtnl_lock();
> ...
> }
>
> This is the interleaving in question:
>
> CPU1 (rmmod) CPU2 (ip link set dev down)
> __dev_close_many() holds RTNL
> idpf_vport_stop(vport, true)
> IDPF_VPORT_UP is set
> rtnl_lock() blocks
> idpf_stop()
> idpf_vport_stop(vport, false)
> queues/vectors released
> IDPF_VPORT_UP cleared
> rtnl_unlock()
> teardown continues on released state
>
> idpf_vport_intr_rel() sets q_vectors to NULL but does not change
> num_q_vectors:
>
> kfree(rsrc->q_vectors);
> rsrc->q_vectors = NULL;
>
> On the second pass, idpf_vport_intr_dis_irq_all() does:
>
> struct idpf_q_vector *q_vector = rsrc->q_vectors;
> ...
> for (u16 q_idx = 0; q_idx < rsrc->num_q_vectors; q_idx++)
> writel(0, q_vector[q_idx].intr_reg.dyn_ctl);
>
> Won't that dereference a NULL q_vector? The virtchnl message builders
> that walk the NULLed rsrc->txq_grps/rxq_grps may also crash before that
> point.
>
> Before this patch, a close that reached idpf_stop() after remove had set
> IDPF_REMOVE_IN_PROG returned 0, so only one teardown ran.
>
> Would it help to re-check IDPF_VPORT_UP in idpf_vport_stop() after
> rtnl_lock() is taken, or to hold vport_ctrl_lock around the stop in the
> remove path?
>
>> idpf_vport_ctrl_lock(netdev);
>> vport = idpf_netdev_to_vport(netdev);
>>
This is a valid hit. I will post v2 to address it. For the pre-existing
below, I will review and post fixes as needed.
Thanks,
Emil
>
> [Severity: High]
> This is a pre-existing issue and was not introduced by this patch. It
> matters here because this fix now depends on the loop in
> idpf_detach_and_close():
>
> for (int i = 0; i < max_vports; i++) {
> struct net_device *netdev = adapter->netdevs[i];
> ...
> if (!netif_device_present(netdev))
> continue;
>
> idpf_vc_core_init() sizes adapter->netdevs by caps.max_vports. However,
> idpf_init_task() only creates netdevs up to default_num_vports. If the
> device reports max_vports > default_num_vports, can the first hard reset
> pass a NULL netdev to netif_device_present()?
>
> idpf_vc_event_task() -> idpf_init_hard_reset() -> idpf_detach_and_close()
> -> netif_device_present(NULL)
>
> The other max_vports-bounded loops in the driver check for empty slots.
> idpf_remove() does "if (!adapter->netdevs[i]) continue;", idpf_init_task()
> checks !netdev, and idpf_decfg_netdev() stores NULL into the slot on
> purpose.
>
> It is not clear which shipping firmware reports max_vports larger than
> default_num_vports. virtchnl2 defines them as separate fields, though.
> This has been present since 2e281e1155fc, the commit named in Fixes:.
>
> [Severity: High]
> This is also a pre-existing issue and was not introduced by this patch.
> However, the remove-racing-reset case now goes through it as well.
>
> On a software-initiated function reset, the transaction manager is shut
> down before the close path runs:
>
> idpf_vc_event_task() {
> ...
> func_reset:
> if (adapter->xnm)
> libie_ctlq_xn_shutdown(adapter->xnm);
> drv_load:
> set_bit(IDPF_HR_RESET_IN_PROG, adapter->flags);
> idpf_init_hard_reset(adapter);
> ...
> }
>
> After that, libie_ctlq_xn_pop_free() refuses new transactions:
>
> if (unlikely(xnm->shutdown))
> return NULL;
>
> The call chain is idpf_init_hard_reset() -> idpf_detach_and_close() ->
> dev_close() -> idpf_stop() -> idpf_vport_stop(). The DISABLE_VPORT and
> DISABLE_QUEUES messages fail silently because their return values are
> ignored. Then the rings and buffers are freed:
>
> idpf_vport_intr_deinit(vport, rsrc);
> idpf_xdp_rxq_info_deinit_all(rsrc);
> idpf_vport_queues_rel(vport, rsrc);
> idpf_vport_intr_rel(rsrc);
>
> reg_ops->trigger_reset() is only called later, after
> idpf_vc_core_deinit(). Can the still-enabled Rx queues DMA into the freed
> descriptor rings or buffers during that window?
>
> On a VF, idpf_vf_trigger_reset() skips the reset entirely during remove:
>
> if (trig_cause == IDPF_HR_FUNC_RESET &&
> !test_bit(IDPF_REMOVE_IN_PROG, adapter->flags))
> idpf_send_vf_reset_msg(adapter);
>
> In that case, no device reset follows the free at all.
>
> Before this patch, the remove-racing-reset case skipped idpf_vport_stop()
> in idpf_stop(). The rings leaked, but they were not freed while still in
> use.
>
> Would it help to shut down the transaction manager only after
> idpf_detach_and_close() has disabled the queues? Another option is to
> keep the DMA memory until the device has been reset.
>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
` (6 preceding siblings ...)
2026-09-28 23:10 ` [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) netdev-bot+sinfo
@ 2026-10-01 23:47 ` Tony Nguyen
2026-10-02 20:39 ` Jakub Kicinski
2026-10-02 20:50 ` patchwork-bot+netdevbpf
8 siblings, 1 reply; 28+ messages in thread
From: Tony Nguyen @ 2026-10-01 23:47 UTC (permalink / raw)
To: davem, kuba, pabeni, edumazet, andrew+netdev, netdev
Cc: emil.s.tantilov, luoxuanqiang, bryan.fraschetti, tristan,
tomasz.lichwala, david.butler, horms
On 9/28/2026 4:04 PM, Tony Nguyen wrote:
> For idpf:
> Emil fixes a remove/reset race by always stopping the vport during
> idpf_stop() so NAPI is properly torn down.
>
> For ice:
> Xuanqiang Luo resolves a use-after-free issue by changing order of
> operations so index is used before being freed.
>
> Bryan Fraschetti restores ordered MMIO writes for ice Tx doorbells
> by replacing writel_relaxed() call with writel().
>
> Tristan Madani fixes representor use-after-free by releasing
> metadata_dst through dst_release() to ensure it is not freed until
> all references are dropped.
>
> For iavf:
> Tomasz fixes reporting of statistics by always requesting statistics
> while the adapter is running as PTP commands can interfere with the
> previous fallback stats update mechanism.
>
> Dave Butler caps the advertised maximum packet size to the hardware's
> single-buffer limit, preventing PF queue-configuration failures caused
> by oversized multi-buffer frame limits.
Patches 1 & 5 will need to be reworked and resubmitted; responses to
Sashiko are on 3 & 6. I'm hoping would be willing take the acceptable
ones and I will resubmit the reworked ones later.
Thanks,
Tony
> The following are changes since commit a7bfaba4823e3c165bb2004c74eff7c096672bc7:
> ipv6: fix prefix route expiry in modify_prefix_route()
> and are available in the git repository at:
> git://git.kernel.org/pub/scm/linux/kernel/git/tnguy/net-queue 200GbE
>
> Bryan Fraschetti (1):
> ice: Restore Ordered MMIO Writes for Tx Doorbells
>
> Dave Butler (1):
> iavf: cap advertised max_pkt_size at the single-buffer HW limit
>
> Emil Tantilov (1):
> idpf: fix possible race on remove during a reset
>
> Tomasz Lichwala (1):
> iavf: fix VF stats not updating due to PTP command preemption
>
> Tristan Madani (1):
> ice: fix metadata_dst refcount handling on representor teardown
>
> Xuanqiang Luo (1):
> ice: fix use-after-free in dynamic port cleanup
>
> drivers/net/ethernet/intel/iavf/iavf_main.c | 14 ++++----------
> drivers/net/ethernet/intel/iavf/iavf_virtchnl.c | 8 ++++++++
> drivers/net/ethernet/intel/ice/devlink/port.c | 2 +-
> drivers/net/ethernet/intel/ice/ice_eswitch.c | 2 +-
> drivers/net/ethernet/intel/ice/ice_txrx.c | 4 ++--
> drivers/net/ethernet/intel/idpf/idpf_lib.c | 4 ----
> 6 files changed, 16 insertions(+), 18 deletions(-)
>
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
2026-10-01 23:47 ` Tony Nguyen
@ 2026-10-02 20:39 ` Jakub Kicinski
[not found] ` <IA3PR05MB220784D3846E12367B0CA1D3738B892@IA3PR05MB220784.namprd05.prod.outlook.com>
0 siblings, 1 reply; 28+ messages in thread
From: Jakub Kicinski @ 2026-10-02 20:39 UTC (permalink / raw)
To: Tony Nguyen
Cc: davem, pabeni, edumazet, andrew+netdev, netdev, emil.s.tantilov,
luoxuanqiang, bryan.fraschetti, tristan, tomasz.lichwala,
david.butler, horms
On Thu, 1 Oct 2026 16:47:22 -0700 Tony Nguyen wrote:
> Patches 1 & 5 will need to be reworked and resubmitted; responses to
> Sashiko are on 3 & 6. I'm hoping would be willing take the acceptable
> ones and I will resubmit the reworked ones later.
I'm not touching 6 since Dave's response has the confidentiality clause.
So basically I'm picking ice patches as-is, and dropping the rest.
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
` (7 preceding siblings ...)
2026-10-01 23:47 ` Tony Nguyen
@ 2026-10-02 20:50 ` patchwork-bot+netdevbpf
8 siblings, 0 replies; 28+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-10-02 20:50 UTC (permalink / raw)
To: Tony Nguyen
Cc: davem, kuba, pabeni, edumazet, andrew+netdev, netdev,
emil.s.tantilov, luoxuanqiang, bryan.fraschetti, tristan,
tomasz.lichwala, david.butler, horms
Hello:
This series was applied to netdev/net.git (main)
by Jakub Kicinski <kuba@kernel.org>:
On Mon, 28 Sep 2026 16:04:21 -0700 you wrote:
> For idpf:
> Emil fixes a remove/reset race by always stopping the vport during
> idpf_stop() so NAPI is properly torn down.
>
> For ice:
> Xuanqiang Luo resolves a use-after-free issue by changing order of
> operations so index is used before being freed.
>
> [...]
Here is the summary with links:
- [net,1/6] idpf: fix possible race on remove during a reset
(no matching commit)
- [net,2/6] ice: fix use-after-free in dynamic port cleanup
https://git.kernel.org/netdev/net/c/1402fc67d2f4
- [net,3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells
https://git.kernel.org/netdev/net/c/7a579dec8423
- [net,4/6] ice: fix metadata_dst refcount handling on representor teardown
https://git.kernel.org/netdev/net/c/58858b8f1480
- [net,5/6] iavf: fix VF stats not updating due to PTP command preemption
(no matching commit)
- [net,6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit
(no matching commit)
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: Fw: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
[not found] ` <CANm61jco98RoBmBAtbnjRZCPwqS0Vt6SqXjYAEUwSD4-bWLuZA@mail.gmail.com>
@ 2026-10-02 21:06 ` David Butler
0 siblings, 0 replies; 28+ messages in thread
From: David Butler @ 2026-10-02 21:06 UTC (permalink / raw)
To: Jakub Kicinski, anthony.l.nguyen
Cc: Dave Butler, davem, pabeni, edumazet, andrew+netdev, netdev,
emil.s.tantilov, luoxuanqiang, bryan.fraschetti, tristan,
tomasz.lichwala, horms
> I'm not touching 6 since Dave's response has the confidentiality clause.
Apologies, I don't know how to disable it.
(Also, apologies again, as last attempt contained HTML)
I will resend my review response for 6 from my personal email (this)
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: Fw: [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit
[not found] ` <CANm61jc37jivo=XmRwN8ic8PNJFXZK+6Gsw5VRy=b-oZKpMsMA@mail.gmail.com>
@ 2026-10-02 21:09 ` David Butler
0 siblings, 0 replies; 28+ messages in thread
From: David Butler @ 2026-10-02 21:09 UTC (permalink / raw)
To: netdev-bot+sashiko, anthony.l.nguyen, kuba
Cc: davem, pabeni, edumazet, andrew+netdev, netdev, emil.s.tantilov,
luoxuanqiang, bryan.fraschetti, tristan, tomasz.lichwala, horms,
aleksander.lobakin, stable, jacob.e.keller, aleksandr.loktionov
> [Severity: Medium]
> Should netdev->max_mtu be clamped to match?
Yes, the review feedback is valid. The bug outlined by the review is
where the user sets an MTU
in the 9728 to ~16356 range and frames above 9728 are silently dropped
instead of the MTU set failing.
We should still move forward with this patch but reword the commit
message. It will likely be at
least several weeks before I will be able to acquire and setup a lab
environment for reproduction
and testing for a new patch revision. This patch as-is does not
introduce any code regressions.
In every case the behavior is unchanged or strictly better.
Practically speaking, most users would
not be impacted by the frame drop. Standard jumbo (≤9728, e.g. 9000)
is entirely unaffected.
Since 5fa4caff59f2, E810 + ESXi icen + SR-IOV has been unusable: the
VF never finishes
CONFIG_VSI_QUEUES, regardless of MTU. The effects are not contained
to the guest
that triggers it. Under pass-through the mis-programmed queue can
raise an IOMMU fault that
wedges the PF/VF on the ESXi host. This usually takes out that
interface for subsequent
guests too (even if they have this patch). Recovery requires a host reboot.
A follow-up to adopt the reviewer's fix is warranted. I can provide
the patch if someone else
thinks they can test it sooner than I can.
A revised commit message follows:
iavf: cap advertised max_pkt_size at the single-buffer HW limit
Since commit 5fa4caff59f2 ("iavf: switch to Page Pool")
iavf_configure_queues() advertises max_pkt_size to the PF as:
max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
max_frame = min_not_zero(adapter->vf_res->max_mtu, max_frame);
LIBIE_MAX_RX_FRM_LEN (16382) is the multi-descriptor scatter/gather frame
ceiling, not a single-queue value, and it exceeds the E810 MAC frame size
maximum of 9728. Per the E810 datasheet (613875-009 section 13.2.2.17.1)
the Tx frame-size register PRTDCB_TDPUC.MAX_TXFRAME has a maximum of
0x2600 (9728); larger frames are discarded. The in-tree ice driver encodes
the same value as ICE_AQ_SET_MAC_FRAME_SIZE_MAX (== LIBIE_MAX_RX_BUF_LEN ==
9728), and the VF clamped max_frame to IAVF_MAX_RXBUFFER (9728) before this
commit.
When the PF advertises vf_res->max_mtu as 0, min_not_zero() leaves
max_frame at 16382. The Linux ice PF advertises max_mtu = port MAC frame
size (<= 9728), so a VF behind ice never sends more than that. The ESXi
"icen" PF on E810 advertises max_mtu as 0, so the VF sends
max_pkt_size = 16382, which icen rejects while programming the queue
context for VIRTCHNL_OP_CONFIG_VSI_QUEUES (opcode 6):
icen_ConfigureTxQueue: VSI 8: Failed to set LAN Tx queue context for
absolute Tx queue 64, Error: ICE_ERR_PARAM
indrv_SendMsgToVf: VF 0: Failed opcode 6, Error -5
iavf 0000:03:00.0: PF returned error -5 (IAVF_ERR_PARAM) to
our request 6
iavf 0000:03:00.0 ethX: NETDEV WATCHDOG: transmit queue N timed out
The VF's queues never come up; under SR-IOV passthrough the mis-programmed
queue can also trigger a fatal IOMMU fault in the guest. Forcing only
max_pkt_size back to 9728 (and leaving the Page Pool rx_buf_len/
databuffer_size untouched) makes the VF come up; databuffer_size is not
involved. This was confirmed on two E810 NVM versions (3.00 and 4.51) under
icen, so the trigger is the icen PF, not the firmware revision. Reported by
several users on E810 + ESXi icen with v6.10+ guests:
Link: https://community.intel.com/t5/Ethernet-Products/E810-C-iavf-driver-issue-on-Linux-6-12/m-p/1737490
Link: https://access.redhat.com/solutions/6973766
Link: https://knowledge.broadcom.com/external/article/404315/sriov-enabled-vms-network-adaptor-goes-d.html
Cap max_frame at the single-buffer HW limit (LIBIE_MAX_RX_BUF_LEN, 9728) so
the VF advertises a max_pkt_size the PF can program into the Rx queue
context. This is the value the VF used before 5fa4caff59f2 (IAVF_MAX_RXBUFFER).
Note this only caps the advertised max_pkt_size. 5fa4caff59f2 also raised
netdev->max_mtu to LIBIE_MAX_MTU (derived from the same 16382 ceiling); when
the PF advertises max_mtu == 0 that leaves max_mtu above 9728, so an MTU in
the 9728 < mtu <= LIBIE_MAX_MTU range can still be set while frames above
9728 are dropped by the queue context. Clamping netdev->max_mtu to match is
left to a follow-up.
Fixes: 5fa4caff59f2 ("iavf: switch to Page Pool")
Signed-off-by: Dave Butler <david.butler@appgate.com>
Cc: stable@vger.kernel.org # v6.10+
^ permalink raw reply [flat|nested] 28+ messages in thread
* Re: Fw: [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf)
[not found] ` <IA3PR05MB22078467309FA5DFF33EF712488B892@IA3PR05MB220784.namprd05.prod.outlook.com>
@ 2026-10-02 21:21 ` David Butler
0 siblings, 0 replies; 28+ messages in thread
From: David Butler @ 2026-10-02 21:21 UTC (permalink / raw)
To: netdev-bot+sinfo, anthony.l.nguyen
Cc: davem, Jakub Kicinski, pabeni, edumazet, andrew+netdev, netdev,
emil.s.tantilov, luoxuanqiang, bryan.fraschetti, tristan,
tomasz.lichwala, horms
Answers inline for :
iavf: cap advertised max_pkt_size at the single-buffer HW limit
> How the issue was discovered
Found in the field, in a real production environment, and reproduced in lab.
> Whether the issue was actually triggered, or is only theoretical. If it was triggered please include the symptoms, like the stack trace or error messages.
The issue was actually triggered. VM and/or hypervisor can lose all
NIC functionality, VM may stall or crash.
esxi icen logs may include the following:
icen_ConfigureTxQueue: VSI 8: Failed to set LAN Tx queue
context for absolute Tx queue 64, Error: ICE_ERR_PARAM
indrv_SendMsgToVf: VF 0: Failed opcode 6, Error -5
kernel messages may include the following:
iavf 0000:03:00.0: PF returned error -5 (IAVF_ERR_PARAM) to
our request 6
iavf 0000:03:00.0 ethX: NETDEV WATCHDOG: transmit queue N timed out
> What hardware the change was tested on. For driver fixes please mention the device (and if relevant firmware version) used for testing, or say that the change was not tested on real hardware.
Key hardware for reproduction is E810 NIC, running on esxi (icen) with
SR-IOV enabled.
E810 NVM revisions: 3.00, 4.51
icen versions: 1.14.2.0, 2.3.3.0
>The evaluation is done by an LLM so it may be wrong, if you think that is the case please reply and explain.
This information was covered comprehensively in the commit messages,
but happy to repeat here.
^ permalink raw reply [flat|nested] 28+ messages in thread
end of thread, other threads:[~2026-10-02 21:21 UTC | newest]
Thread overview: 28+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 23:04 [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) Tony Nguyen
2026-09-28 23:04 ` [PATCH net 1/6] idpf: fix possible race on remove during a reset Tony Nguyen
2026-09-30 0:58 ` netdev-bot+sashiko
2026-10-01 23:40 ` Tantilov, Emil S
2026-09-28 23:04 ` [PATCH net 2/6] ice: fix use-after-free in dynamic port cleanup Tony Nguyen
2026-09-28 23:04 ` [PATCH net 3/6] ice: Restore Ordered MMIO Writes for Tx Doorbells Tony Nguyen
2026-09-30 0:58 ` netdev-bot+sashiko
2026-10-01 16:35 ` Tony Nguyen
2026-09-28 23:04 ` [PATCH net 4/6] ice: fix metadata_dst refcount handling on representor teardown Tony Nguyen
2026-09-28 23:04 ` [PATCH net 5/6] iavf: fix VF stats not updating due to PTP command preemption Tony Nguyen
2026-09-30 0:58 ` netdev-bot+sashiko
2026-09-30 15:07 ` Tomasz Lichwala
2026-09-28 23:04 ` [PATCH net 6/6] iavf: cap advertised max_pkt_size at the single-buffer HW limit Tony Nguyen
2026-09-29 16:18 ` Alexander Lobakin
2026-09-29 20:32 ` Dave Butler
2026-09-30 11:32 ` Alexander Lobakin
2026-09-30 0:58 ` netdev-bot+sashiko
2026-09-30 6:02 ` Dave Butler
[not found] ` <IA3PR05MB22078430E404FAD12912A706298B892@IA3PR05MB220784.namprd05.prod.outlook.com>
[not found] ` <CANm61jc37jivo=XmRwN8ic8PNJFXZK+6Gsw5VRy=b-oZKpMsMA@mail.gmail.com>
2026-10-02 21:09 ` Fw: " David Butler
2026-09-28 23:10 ` [PATCH net 0/6][pull request] Intel Wired LAN Driver Updates 2026-09-28 (idpf, ice, iavf) netdev-bot+sinfo
2026-09-29 1:41 ` Dave Butler
[not found] ` <IA3PR05MB22078467309FA5DFF33EF712488B892@IA3PR05MB220784.namprd05.prod.outlook.com>
2026-10-02 21:21 ` Fw: " David Butler
2026-09-29 17:16 ` Tantilov, Emil S
2026-09-30 15:06 ` Tomasz Lichwala
2026-10-01 23:47 ` Tony Nguyen
2026-10-02 20:39 ` Jakub Kicinski
[not found] ` <IA3PR05MB220784D3846E12367B0CA1D3738B892@IA3PR05MB220784.namprd05.prod.outlook.com>
[not found] ` <CANm61jco98RoBmBAtbnjRZCPwqS0Vt6SqXjYAEUwSD4-bWLuZA@mail.gmail.com>
2026-10-02 21:06 ` Fw: " David Butler
2026-10-02 20:50 ` patchwork-bot+netdevbpf
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).