Netdev List
 help / color / mirror / Atom feed
* [PATCH iwl-net 0/2] idpf: stop stray device writes after a warm reboot
@ 2026-09-17 10:52 Tian Xun Ng
  2026-09-17 10:52 ` [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown Tian Xun Ng
  2026-09-17 10:52 ` [PATCH iwl-net 2/2] idpf: reset the function " Tian Xun Ng
  0 siblings, 2 replies; 12+ messages in thread
From: Tian Xun Ng @ 2026-09-17 10:52 UTC (permalink / raw)
  To: intel-wired-lan
  Cc: netdev, anthony.l.nguyen, przemyslaw.kitszel, aleksander.lobakin,
	emil.s.tantilov, andrew+netdev, davem, edumazet, kuba, pabeni,
	Tian Xun Ng

From: Tian Xun Ng <tianxun.ng@bytedance.com>

After a warm reboot with idpf loaded, the device writes SW_MARKER TX
completions to DMA addresses that belong to the previous kernel, about
20 seconds into the next boot. With the IOMMU translating, this shows up
as a burst of F_TRANSLATION faults from the idpf PCI functions on every
such boot. In IOMMU pass-through mode the writes corrupt memory the new
kernel has already reused, and nodes crash in unrelated code (page table
freeing, slab allocation).

The cause is in idpf_shutdown(). Since commit 4c9106f4906a ("idpf: fix
adapter NULL pointer dereference on reboot") it no longer goes through
idpf_remove(), so idpf_vc_core_deinit() shuts the virtchnl transaction
manager down before the vports are torn down. The disable and destroy
messages then fail and the device keeps its queues enabled.

Patch 1 keeps the mailbox up for that teardown unless a hard reset is in
progress or has been detected. Patch 2 restores the function reset that
idpf_remove() performs on exit.

Testing:
- The same two changes, backported to a 6.17 kernel, were run on arm64
  (64K pages) servers with two idpf functions: no stray device writes in
  151 warm reboots (130 in IOMMU pass-through mode with page_poison=1,
  21 with the IOMMU translating), against stray writes after 20 of 20
  warm reboots on an unpatched control node.
- Not covered: restarts after a kernel crash, and a device that cannot
  answer at shutdown; the writes still occur in those cases.
- This series is those changes ported to the dev-queue branch of
  tnguy/net-queue. On this tree it is build-tested only: W=1 builds of
  drivers/net/ethernet/intel/idpf with arm64 defconfig (plus IDPF),
  allmodconfig and allyesconfig, with no warnings. It has not been run
  on hardware on this tree.

An LLM coding assistant helped analyse the crash dumps and fault logs,
locate the shutdown ordering problem, and draft both changes and their
changelogs. Both patches carry an Assisted-by tag.

Tian Xun Ng (2):
  idpf: keep the mailbox up while tearing down vports on shutdown
  idpf: reset the function on shutdown

 drivers/net/ethernet/intel/idpf/idpf_main.c    |  3 +++
 .../net/ethernet/intel/idpf/idpf_virtchnl.c    | 18 +++++++++++++-----
 2 files changed, 16 insertions(+), 5 deletions(-)


base-commit: a98bd9f12dc5ed64029d00a8192a28685a54a587
-- 
2.50.1 (Apple Git-155)


^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
  2026-09-17 10:52 [PATCH iwl-net 0/2] idpf: stop stray device writes after a warm reboot Tian Xun Ng
@ 2026-09-17 10:52 ` Tian Xun Ng
  2026-09-18 15:45   ` Loktionov, Aleksandr
  2026-09-18 17:59   ` Tantilov, Emil S
  2026-09-17 10:52 ` [PATCH iwl-net 2/2] idpf: reset the function " Tian Xun Ng
  1 sibling, 2 replies; 12+ messages in thread
From: Tian Xun Ng @ 2026-09-17 10:52 UTC (permalink / raw)
  To: intel-wired-lan
  Cc: netdev, anthony.l.nguyen, przemyslaw.kitszel, aleksander.lobakin,
	emil.s.tantilov, andrew+netdev, davem, edumazet, kuba, pabeni,
	Tian Xun Ng

From: Tian Xun Ng <tianxun.ng@bytedance.com>

Since commit 4c9106f4906a ("idpf: fix adapter NULL pointer dereference
on reboot"), idpf_shutdown() calls idpf_vc_core_deinit() directly instead
of idpf_remove(), so IDPF_REMOVE_IN_PROG is not set on shutdown.
idpf_vc_core_deinit() uses that flag to decide when to shut the virtchnl
transaction manager down. Without it, libie_ctlq_xn_shutdown() runs
before idpf_deinit_task() tears the vports down, so every message sent
during that teardown (disable vport, disable queues, destroy vport)
fails at once. The device is never told to stop its queues and keeps
them enabled, with the ring addresses of the kernel that is going away.

After a warm reboot, the first queue reconfiguration of a port that has
not been opened yet (udev setting the MTU) sends VIRTCHNL2_OP_DEL_QUEUES.
The device then drains the queues it still considers live and writes
SW_MARKER TX completions (qid_comptype_gen 0x2800 / 0xa800) to the
previous kernel's completion rings. With the IOMMU translating, those
writes fault on every such boot:

  arm-smmu-v3 arm-smmu-v3.17.auto: event: F_TRANSLATION client: 0016:01:00.0 sid: 0x30100 ssid: 0x0 iova: 0xffb70000 ipa: 0x0
  arm-smmu-v3 arm-smmu-v3.17.auto: unpriv data write s1 "Input address caused fault" stag: 0x0

In IOMMU pass-through mode they land on pages the new kernel has already
reused. With page_poison=1 that shows as "pagealloc: memory corruption"
on most boots. Without it, nodes crash in unrelated code, e.g.:

  Unable to handle kernel NULL pointer dereference at virtual address 000000000000a830
  pc : __tlb_remove_table_free+0x48/0x118
  Call trace:
   __tlb_remove_table_free
   tlb_remove_table_rcu
   rcu_do_batch

or with a slab free pointer that decodes from a slot holding 0x2800:

  Unable to handle kernel paging request at virtual address 002613b73e862c6e
  pc : kmem_cache_alloc_noprof+0xc0/0x3d0

The early shutdown exists to avoid waiting for transaction timeouts when
the mailbox is already gone. That is the hard reset case, where
idpf_vc_event_task() shuts the transaction manager down and sets
IDPF_HR_RESET_IN_PROG before idpf_init_hard_reset() calls
idpf_vc_core_deinit(). Key the early shutdown on a hard reset being in
progress or detected instead of on remove, so that both remove and
shutdown keep the mailbox up until the vports are gone.

Hard reset behaviour is unchanged. Remove is unchanged unless a hardware
reset has been detected, in which case it no longer waits for message
timeouts. Shutdown now delivers the teardown messages; if the device
stops responding without a detectable reset, shutdown can wait for the
transaction timeouts, as remove already does.

Tested on arm64 (64K pages) servers with two idpf functions, with this
change and the next patch backported to a 6.17 kernel: no stray device
writes in 151 warm reboots across IOMMU translated and pass-through
modes, against stray writes after 20 of 20 warm reboots without them.

Fixes: 4c9106f4906a ("idpf: fix adapter NULL pointer dereference on reboot")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Tian Xun Ng <tianxun.ng@bytedance.com>
---
 .../net/ethernet/intel/idpf/idpf_virtchnl.c    | 18 +++++++++++++-----
 1 file changed, 13 insertions(+), 5 deletions(-)

diff --git a/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c b/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
index 1caf52706..646b6e074 100644
--- a/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
+++ b/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
@@ -3197,14 +3197,22 @@ int idpf_vc_core_init(struct idpf_adapter *adapter)
  */
 void idpf_vc_core_deinit(struct idpf_adapter *adapter)
 {
-	bool remove_in_prog;
+	bool reset_in_prog;
 
 	if (!test_bit(IDPF_VC_CORE_INIT, adapter->flags))
 		return;
 
-	/* Avoid transaction timeouts when called during reset */
-	remove_in_prog = test_bit(IDPF_REMOVE_IN_PROG, adapter->flags);
-	if (!remove_in_prog)
+	/* Shut the transaction manager down early only when the mailbox is
+	 * already gone, i.e. a hard reset is in progress or has been detected,
+	 * to avoid waiting for transaction timeouts. On remove and on shutdown
+	 * the mailbox still works and must stay up until the vports are torn
+	 * down. Otherwise the disable and destroy messages never reach the
+	 * device, which keeps its queues enabled with ring addresses from this
+	 * kernel after a warm reboot.
+	 */
+	reset_in_prog = test_bit(IDPF_HR_RESET_IN_PROG, adapter->flags) ||
+			idpf_is_reset_detected(adapter);
+	if (reset_in_prog)
 		libie_ctlq_xn_shutdown(adapter->xnm);
 
 	idpf_ptp_release(adapter);
@@ -3213,7 +3221,7 @@ void idpf_vc_core_deinit(struct idpf_adapter *adapter)
 	idpf_rel_rx_pt_lkup(adapter);
 	idpf_intr_rel(adapter);
 
-	if (remove_in_prog)
+	if (!reset_in_prog)
 		libie_ctlq_xn_shutdown(adapter->xnm);
 
 	cancel_delayed_work_sync(&adapter->serv_task);
-- 
2.50.1 (Apple Git-155)


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH iwl-net 2/2] idpf: reset the function on shutdown
  2026-09-17 10:52 [PATCH iwl-net 0/2] idpf: stop stray device writes after a warm reboot Tian Xun Ng
  2026-09-17 10:52 ` [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown Tian Xun Ng
@ 2026-09-17 10:52 ` Tian Xun Ng
  2026-09-18 15:46   ` Loktionov, Aleksandr
  2026-09-18 17:54   ` Tantilov, Emil S
  1 sibling, 2 replies; 12+ messages in thread
From: Tian Xun Ng @ 2026-09-17 10:52 UTC (permalink / raw)
  To: intel-wired-lan
  Cc: netdev, anthony.l.nguyen, przemyslaw.kitszel, aleksander.lobakin,
	emil.s.tantilov, andrew+netdev, davem, edumazet, kuba, pabeni,
	Tian Xun Ng

From: Tian Xun Ng <tianxun.ng@bytedance.com>

idpf_remove() ends with a function reset to leave the device clean for
whoever binds it next. Commit 4c9106f4906a ("idpf: fix adapter NULL
pointer dereference on reboot") replaced the idpf_remove() call in
idpf_shutdown() with idpf_vc_core_deinit() and idpf_deinit_dflt_mbx(),
and the reset was lost along the way.

Restore it, so that the kernel started by a warm reboot or kexec finds
the device in the same state as after a module unload.

The reset alone does not stop the stray completion writes fixed by the
previous patch: the next kernel's load-time reset already performs a
reset and stale queue state was observed to survive it. It complements
the previous patch, which makes the device tear its queues down first.

Fixes: 4c9106f4906a ("idpf: fix adapter NULL pointer dereference on reboot")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Tian Xun Ng <tianxun.ng@bytedance.com>
---
 drivers/net/ethernet/intel/idpf/idpf_main.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/drivers/net/ethernet/intel/idpf/idpf_main.c b/drivers/net/ethernet/intel/idpf/idpf_main.c
index 129bccaa6..d7cd449fc 100644
--- a/drivers/net/ethernet/intel/idpf/idpf_main.c
+++ b/drivers/net/ethernet/intel/idpf/idpf_main.c
@@ -197,6 +197,9 @@ static void idpf_shutdown(struct pci_dev *pdev)
 	cancel_delayed_work_sync(&adapter->serv_task);
 	cancel_delayed_work_sync(&adapter->vc_event_task);
 	idpf_vc_core_deinit(adapter);
+
+	/* Leave the device clean for the next kernel, as idpf_remove() does */
+	adapter->dev_ops.reg_ops.trigger_reset(adapter, IDPF_HR_FUNC_RESET);
 	idpf_deinit_dflt_mbx(adapter);
 
 	if (system_state == SYSTEM_POWER_OFF)
-- 
2.50.1 (Apple Git-155)


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* RE: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
  2026-09-17 10:52 ` [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown Tian Xun Ng
@ 2026-09-18 15:45   ` Loktionov, Aleksandr
  2026-09-18 17:59   ` Tantilov, Emil S
  1 sibling, 0 replies; 12+ messages in thread
From: Loktionov, Aleksandr @ 2026-09-18 15:45 UTC (permalink / raw)
  To: Tian Xun Ng, intel-wired-lan@lists.osuosl.org
  Cc: netdev@vger.kernel.org, Nguyen, Anthony L, Kitszel, Przemyslaw,
	Lobakin, Aleksander, Tantilov, Emil S, andrew+netdev@lunn.ch,
	davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, Tian Xun Ng



> -----Original Message-----
> From: Tian Xun Ng <luckilystar08@gmail.com>
> Sent: Thursday, September 17, 2026 12:52 PM
> To: intel-wired-lan@lists.osuosl.org
> Cc: netdev@vger.kernel.org; Nguyen, Anthony L
> <anthony.l.nguyen@intel.com>; Kitszel, Przemyslaw
> <przemyslaw.kitszel@intel.com>; Lobakin, Aleksander
> <aleksander.lobakin@intel.com>; Tantilov, Emil S
> <emil.s.tantilov@intel.com>; andrew+netdev@lunn.ch;
> davem@davemloft.net; edumazet@google.com; kuba@kernel.org;
> pabeni@redhat.com; Tian Xun Ng <tianxun.ng@bytedance.com>
> Subject: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing
> down vports on shutdown
> 
> From: Tian Xun Ng <tianxun.ng@bytedance.com>
> 
> Since commit 4c9106f4906a ("idpf: fix adapter NULL pointer dereference
> on reboot"), idpf_shutdown() calls idpf_vc_core_deinit() directly
> instead of idpf_remove(), so IDPF_REMOVE_IN_PROG is not set on
> shutdown.
> idpf_vc_core_deinit() uses that flag to decide when to shut the
> virtchnl transaction manager down. Without it,
> libie_ctlq_xn_shutdown() runs before idpf_deinit_task() tears the
> vports down, so every message sent during that teardown (disable
> vport, disable queues, destroy vport) fails at once. The device is
> never told to stop its queues and keeps them enabled, with the ring
> addresses of the kernel that is going away.
> 
> After a warm reboot, the first queue reconfiguration of a port that
> has not been opened yet (udev setting the MTU) sends
> VIRTCHNL2_OP_DEL_QUEUES.
> The device then drains the queues it still considers live and writes
> SW_MARKER TX completions (qid_comptype_gen 0x2800 / 0xa800) to the
> previous kernel's completion rings. With the IOMMU translating, those
> writes fault on every such boot:
> 
>   arm-smmu-v3 arm-smmu-v3.17.auto: event: F_TRANSLATION client:
> 0016:01:00.0 sid: 0x30100 ssid: 0x0 iova: 0xffb70000 ipa: 0x0
>   arm-smmu-v3 arm-smmu-v3.17.auto: unpriv data write s1 "Input address
> caused fault" stag: 0x0
> 
> In IOMMU pass-through mode they land on pages the new kernel has
> already reused. With page_poison=1 that shows as "pagealloc: memory
> corruption"
> on most boots. Without it, nodes crash in unrelated code, e.g.:
> 
>   Unable to handle kernel NULL pointer dereference at virtual address
> 000000000000a830
>   pc : __tlb_remove_table_free+0x48/0x118
>   Call trace:
>    __tlb_remove_table_free
>    tlb_remove_table_rcu
>    rcu_do_batch
> 
> or with a slab free pointer that decodes from a slot holding 0x2800:
> 
>   Unable to handle kernel paging request at virtual address
> 002613b73e862c6e
>   pc : kmem_cache_alloc_noprof+0xc0/0x3d0
> 
> The early shutdown exists to avoid waiting for transaction timeouts
> when the mailbox is already gone. That is the hard reset case, where
> idpf_vc_event_task() shuts the transaction manager down and sets
> IDPF_HR_RESET_IN_PROG before idpf_init_hard_reset() calls
> idpf_vc_core_deinit(). Key the early shutdown on a hard reset being in
> progress or detected instead of on remove, so that both remove and
> shutdown keep the mailbox up until the vports are gone.
> 
> Hard reset behaviour is unchanged. Remove is unchanged unless a
> hardware reset has been detected, in which case it no longer waits for
> message timeouts. Shutdown now delivers the teardown messages; if the
> device stops responding without a detectable reset, shutdown can wait
> for the transaction timeouts, as remove already does.
> 
> Tested on arm64 (64K pages) servers with two idpf functions, with this
> change and the next patch backported to a 6.17 kernel: no stray device
> writes in 151 warm reboots across IOMMU translated and pass-through
> modes, against stray writes after 20 of 20 warm reboots without them.
> 
> Fixes: 4c9106f4906a ("idpf: fix adapter NULL pointer dereference on
> reboot")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Tian Xun Ng <tianxun.ng@bytedance.com>
> ---
>  .../net/ethernet/intel/idpf/idpf_virtchnl.c    | 18 +++++++++++++----
> -
>  1 file changed, 13 insertions(+), 5 deletions(-)
> 
> diff --git a/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
> b/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
> index 1caf52706..646b6e074 100644
> --- a/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
> +++ b/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
> @@ -3197,14 +3197,22 @@ int idpf_vc_core_init(struct idpf_adapter
> *adapter)
>   */
>  void idpf_vc_core_deinit(struct idpf_adapter *adapter)  {
> -	bool remove_in_prog;
> +	bool reset_in_prog;
> 
>  	if (!test_bit(IDPF_VC_CORE_INIT, adapter->flags))
>  		return;
> 
> -	/* Avoid transaction timeouts when called during reset */
> -	remove_in_prog = test_bit(IDPF_REMOVE_IN_PROG, adapter->flags);
> -	if (!remove_in_prog)
> +	/* Shut the transaction manager down early only when the
> mailbox is
> +	 * already gone, i.e. a hard reset is in progress or has been
> detected,
> +	 * to avoid waiting for transaction timeouts. On remove and on
> shutdown
> +	 * the mailbox still works and must stay up until the vports
> are torn
> +	 * down. Otherwise the disable and destroy messages never reach
> the
> +	 * device, which keeps its queues enabled with ring addresses
> from this
> +	 * kernel after a warm reboot.
> +	 */
> +	reset_in_prog = test_bit(IDPF_HR_RESET_IN_PROG, adapter->flags)
> ||
> +			idpf_is_reset_detected(adapter);
> +	if (reset_in_prog)
>  		libie_ctlq_xn_shutdown(adapter->xnm);
> 
>  	idpf_ptp_release(adapter);
> @@ -3213,7 +3221,7 @@ void idpf_vc_core_deinit(struct idpf_adapter
> *adapter)
>  	idpf_rel_rx_pt_lkup(adapter);
>  	idpf_intr_rel(adapter);
> 
> -	if (remove_in_prog)
> +	if (!reset_in_prog)
>  		libie_ctlq_xn_shutdown(adapter->xnm);
> 
>  	cancel_delayed_work_sync(&adapter->serv_task);
> --
> 2.50.1 (Apple Git-155)

Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>

^ permalink raw reply	[flat|nested] 12+ messages in thread

* RE: [PATCH iwl-net 2/2] idpf: reset the function on shutdown
  2026-09-17 10:52 ` [PATCH iwl-net 2/2] idpf: reset the function " Tian Xun Ng
@ 2026-09-18 15:46   ` Loktionov, Aleksandr
  2026-09-18 17:54   ` Tantilov, Emil S
  1 sibling, 0 replies; 12+ messages in thread
From: Loktionov, Aleksandr @ 2026-09-18 15:46 UTC (permalink / raw)
  To: Tian Xun Ng, intel-wired-lan@lists.osuosl.org
  Cc: netdev@vger.kernel.org, Nguyen, Anthony L, Kitszel, Przemyslaw,
	Lobakin, Aleksander, Tantilov, Emil S, andrew+netdev@lunn.ch,
	davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, Tian Xun Ng



> -----Original Message-----
> From: Tian Xun Ng <luckilystar08@gmail.com>
> Sent: Thursday, September 17, 2026 12:52 PM
> To: intel-wired-lan@lists.osuosl.org
> Cc: netdev@vger.kernel.org; Nguyen, Anthony L
> <anthony.l.nguyen@intel.com>; Kitszel, Przemyslaw
> <przemyslaw.kitszel@intel.com>; Lobakin, Aleksander
> <aleksander.lobakin@intel.com>; Tantilov, Emil S
> <emil.s.tantilov@intel.com>; andrew+netdev@lunn.ch;
> davem@davemloft.net; edumazet@google.com; kuba@kernel.org;
> pabeni@redhat.com; Tian Xun Ng <tianxun.ng@bytedance.com>
> Subject: [PATCH iwl-net 2/2] idpf: reset the function on shutdown
> 
> From: Tian Xun Ng <tianxun.ng@bytedance.com>
> 
> idpf_remove() ends with a function reset to leave the device clean for
> whoever binds it next. Commit 4c9106f4906a ("idpf: fix adapter NULL
> pointer dereference on reboot") replaced the idpf_remove() call in
> idpf_shutdown() with idpf_vc_core_deinit() and idpf_deinit_dflt_mbx(),
> and the reset was lost along the way.
> 
> Restore it, so that the kernel started by a warm reboot or kexec finds
> the device in the same state as after a module unload.
> 
> The reset alone does not stop the stray completion writes fixed by the
> previous patch: the next kernel's load-time reset already performs a
> reset and stale queue state was observed to survive it. It complements
> the previous patch, which makes the device tear its queues down first.
> 
> Fixes: 4c9106f4906a ("idpf: fix adapter NULL pointer dereference on
> reboot")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Tian Xun Ng <tianxun.ng@bytedance.com>
> ---
>  drivers/net/ethernet/intel/idpf/idpf_main.c | 3 +++
>  1 file changed, 3 insertions(+)
> 
> diff --git a/drivers/net/ethernet/intel/idpf/idpf_main.c
> b/drivers/net/ethernet/intel/idpf/idpf_main.c
> index 129bccaa6..d7cd449fc 100644
> --- a/drivers/net/ethernet/intel/idpf/idpf_main.c
> +++ b/drivers/net/ethernet/intel/idpf/idpf_main.c
> @@ -197,6 +197,9 @@ static void idpf_shutdown(struct pci_dev *pdev)
>  	cancel_delayed_work_sync(&adapter->serv_task);
>  	cancel_delayed_work_sync(&adapter->vc_event_task);
>  	idpf_vc_core_deinit(adapter);
> +
> +	/* Leave the device clean for the next kernel, as idpf_remove()
> does */
> +	adapter->dev_ops.reg_ops.trigger_reset(adapter,
> IDPF_HR_FUNC_RESET);
>  	idpf_deinit_dflt_mbx(adapter);
> 
>  	if (system_state == SYSTEM_POWER_OFF)
> --
> 2.50.1 (Apple Git-155)

Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH iwl-net 2/2] idpf: reset the function on shutdown
  2026-09-17 10:52 ` [PATCH iwl-net 2/2] idpf: reset the function " Tian Xun Ng
  2026-09-18 15:46   ` Loktionov, Aleksandr
@ 2026-09-18 17:54   ` Tantilov, Emil S
  2026-09-21  3:27     ` Tian Xun Ng
  1 sibling, 1 reply; 12+ messages in thread
From: Tantilov, Emil S @ 2026-09-18 17:54 UTC (permalink / raw)
  To: Tian Xun Ng, intel-wired-lan
  Cc: netdev, anthony.l.nguyen, przemyslaw.kitszel, aleksander.lobakin,
	andrew+netdev, davem, edumazet, kuba, pabeni, Tian Xun Ng



On 9/17/2026 3:52 AM, Tian Xun Ng wrote:
> From: Tian Xun Ng <tianxun.ng@bytedance.com>
> 
> idpf_remove() ends with a function reset to leave the device clean for
> whoever binds it next. Commit 4c9106f4906a ("idpf: fix adapter NULL
> pointer dereference on reboot") replaced the idpf_remove() call in
> idpf_shutdown() with idpf_vc_core_deinit() and idpf_deinit_dflt_mbx(),
> and the reset was lost along the way.
> 
> Restore it, so that the kernel started by a warm reboot or kexec finds
> the device in the same state as after a module unload.
> 
> The reset alone does not stop the stray completion writes fixed by the
> previous patch: the next kernel's load-time reset already performs a
> reset and stale queue state was observed to survive it. It complements
> the previous patch, which makes the device tear its queues down first.
> 
> Fixes: 4c9106f4906a ("idpf: fix adapter NULL pointer dereference on reboot")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Tian Xun Ng <tianxun.ng@bytedance.com>
> ---
>   drivers/net/ethernet/intel/idpf/idpf_main.c | 3 +++
>   1 file changed, 3 insertions(+)
> 
> diff --git a/drivers/net/ethernet/intel/idpf/idpf_main.c b/drivers/net/ethernet/intel/idpf/idpf_main.c
> index 129bccaa6..d7cd449fc 100644
> --- a/drivers/net/ethernet/intel/idpf/idpf_main.c
> +++ b/drivers/net/ethernet/intel/idpf/idpf_main.c
> @@ -197,6 +197,9 @@ static void idpf_shutdown(struct pci_dev *pdev)
>   	cancel_delayed_work_sync(&adapter->serv_task);
>   	cancel_delayed_work_sync(&adapter->vc_event_task);
>   	idpf_vc_core_deinit(adapter);
> +
> +	/* Leave the device clean for the next kernel, as idpf_remove() does */
> +	adapter->dev_ops.reg_ops.trigger_reset(adapter, IDPF_HR_FUNC_RESET);
>   	idpf_deinit_dflt_mbx(adapter);

This alone would not be enough, you need to make sure the reset handling 
does not trigger in the driver. Sashiko is also marking it:
https://sashiko.dev/#/patchset/20260917105205.37561-1-luckilystar08%40gmail.com

you will have to set IDPF_REMOVE_IN_PROG, though I am not sure if that 
alone would be sufficient.

Thanks,
Emil

>   
>   	if (system_state == SYSTEM_POWER_OFF)


^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
  2026-09-17 10:52 ` [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown Tian Xun Ng
  2026-09-18 15:45   ` Loktionov, Aleksandr
@ 2026-09-18 17:59   ` Tantilov, Emil S
  2026-09-21  3:27     ` Tian Xun Ng
  1 sibling, 1 reply; 12+ messages in thread
From: Tantilov, Emil S @ 2026-09-18 17:59 UTC (permalink / raw)
  To: Tian Xun Ng, intel-wired-lan
  Cc: netdev, anthony.l.nguyen, przemyslaw.kitszel, aleksander.lobakin,
	andrew+netdev, davem, edumazet, kuba, pabeni, Tian Xun Ng



On 9/17/2026 3:52 AM, Tian Xun Ng wrote:
> From: Tian Xun Ng <tianxun.ng@bytedance.com>
> 
> Since commit 4c9106f4906a ("idpf: fix adapter NULL pointer dereference
> on reboot"), idpf_shutdown() calls idpf_vc_core_deinit() directly instead
> of idpf_remove(), so IDPF_REMOVE_IN_PROG is not set on shutdown.
> idpf_vc_core_deinit() uses that flag to decide when to shut the virtchnl
> transaction manager down. Without it, libie_ctlq_xn_shutdown() runs
> before idpf_deinit_task() tears the vports down, so every message sent
> during that teardown (disable vport, disable queues, destroy vport)
> fails at once. The device is never told to stop its queues and keeps
> them enabled, with the ring addresses of the kernel that is going away.
> 
> After a warm reboot, the first queue reconfiguration of a port that has
> not been opened yet (udev setting the MTU) sends VIRTCHNL2_OP_DEL_QUEUES.
> The device then drains the queues it still considers live and writes
> SW_MARKER TX completions (qid_comptype_gen 0x2800 / 0xa800) to the
> previous kernel's completion rings. With the IOMMU translating, those
> writes fault on every such boot:
> 
>    arm-smmu-v3 arm-smmu-v3.17.auto: event: F_TRANSLATION client: 0016:01:00.0 sid: 0x30100 ssid: 0x0 iova: 0xffb70000 ipa: 0x0
>    arm-smmu-v3 arm-smmu-v3.17.auto: unpriv data write s1 "Input address caused fault" stag: 0x0
> 
> In IOMMU pass-through mode they land on pages the new kernel has already
> reused. With page_poison=1 that shows as "pagealloc: memory corruption"
> on most boots. Without it, nodes crash in unrelated code, e.g.:
> 
>    Unable to handle kernel NULL pointer dereference at virtual address 000000000000a830
>    pc : __tlb_remove_table_free+0x48/0x118
>    Call trace:
>     __tlb_remove_table_free
>     tlb_remove_table_rcu
>     rcu_do_batch
> 
> or with a slab free pointer that decodes from a slot holding 0x2800:
> 
>    Unable to handle kernel paging request at virtual address 002613b73e862c6e
>    pc : kmem_cache_alloc_noprof+0xc0/0x3d0
> 
> The early shutdown exists to avoid waiting for transaction timeouts when
> the mailbox is already gone. That is the hard reset case, where
> idpf_vc_event_task() shuts the transaction manager down and sets
> IDPF_HR_RESET_IN_PROG before idpf_init_hard_reset() calls
> idpf_vc_core_deinit(). Key the early shutdown on a hard reset being in
> progress or detected instead of on remove, so that both remove and
> shutdown keep the mailbox up until the vports are gone.
> 
> Hard reset behaviour is unchanged. Remove is unchanged unless a hardware
> reset has been detected, in which case it no longer waits for message
> timeouts. Shutdown now delivers the teardown messages; if the device
> stops responding without a detectable reset, shutdown can wait for the
> transaction timeouts, as remove already does.

Actually we can't wait on shutdown. If the MBX is defunct, like CP is 
down or unresponsive, the shutdown will hang for a very long time. This 
is the reason why we wanted to avoid communication on shutdown. Have you 
tested the shutdown after stopping the control plane?

> 
> Tested on arm64 (64K pages) servers with two idpf functions, with this
> change and the next patch backported to a 6.17 kernel: no stray device
> writes in 151 warm reboots across IOMMU translated and pass-through
> modes, against stray writes after 20 of 20 warm reboots without them.
> 
> Fixes: 4c9106f4906a ("idpf: fix adapter NULL pointer dereference on reboot")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Tian Xun Ng <tianxun.ng@bytedance.com>
> ---
>   .../net/ethernet/intel/idpf/idpf_virtchnl.c    | 18 +++++++++++++-----
>   1 file changed, 13 insertions(+), 5 deletions(-)
> 
> diff --git a/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c b/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
> index 1caf52706..646b6e074 100644
> --- a/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
> +++ b/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
> @@ -3197,14 +3197,22 @@ int idpf_vc_core_init(struct idpf_adapter *adapter)
>    */
>   void idpf_vc_core_deinit(struct idpf_adapter *adapter)
>   {
> -	bool remove_in_prog;
> +	bool reset_in_prog;
>   
>   	if (!test_bit(IDPF_VC_CORE_INIT, adapter->flags))
>   		return;
>   
> -	/* Avoid transaction timeouts when called during reset */
> -	remove_in_prog = test_bit(IDPF_REMOVE_IN_PROG, adapter->flags);
> -	if (!remove_in_prog)
> +	/* Shut the transaction manager down early only when the mailbox is
> +	 * already gone, i.e. a hard reset is in progress or has been detected,
> +	 * to avoid waiting for transaction timeouts. On remove and on shutdown
> +	 * the mailbox still works and must stay up until the vports are torn
> +	 * down. Otherwise the disable and destroy messages never reach the
> +	 * device, which keeps its queues enabled with ring addresses from this
> +	 * kernel after a warm reboot.
> +	 */
> +	reset_in_prog = test_bit(IDPF_HR_RESET_IN_PROG, adapter->flags) ||
> +			idpf_is_reset_detected(adapter);
> +	if (reset_in_prog)
>   		libie_ctlq_xn_shutdown(adapter->xnm);

This logic already exists in the reset handling, there should be no need 
to replicate it here. Do you have a trace and/or exact scenario that 
leads to remove being called while in a reset, but MBX is still alive?>
>   	idpf_ptp_release(adapter);
> @@ -3213,7 +3221,7 @@ void idpf_vc_core_deinit(struct idpf_adapter *adapter)
>   	idpf_rel_rx_pt_lkup(adapter);
>   	idpf_intr_rel(adapter);
>   
> -	if (remove_in_prog)
> +	if (!reset_in_prog)
>   		libie_ctlq_xn_shutdown(adapter->xnm);
>   
>   	cancel_delayed_work_sync(&adapter->serv_task);


^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH iwl-net 2/2] idpf: reset the function on shutdown
  2026-09-18 17:54   ` Tantilov, Emil S
@ 2026-09-21  3:27     ` Tian Xun Ng
  0 siblings, 0 replies; 12+ messages in thread
From: Tian Xun Ng @ 2026-09-21  3:27 UTC (permalink / raw)
  To: emil.s.tantilov
  Cc: Tian Xun Ng, intel-wired-lan, netdev, anthony.l.nguyen,
	przemyslaw.kitszel, aleksander.lobakin, aleksandr.loktionov,
	andrew+netdev, davem, edumazet, kuba, pabeni, Tian Xun Ng

On 9/18/2026, Tantilov, Emil S wrote:
> This alone would not be enough, you need to make sure the reset handling
> does not trigger in the driver. Sashiko is also marking it:
> you will have to set IDPF_REMOVE_IN_PROG, though I am not sure if that
> alone would be sufficient.

You are right, and so is the review bot. I missed it because our hardware
takes the PF path.

On the VF path idpf_vf_trigger_reset() sends VIRTCHNL2_OP_RESET_VF unless
IDPF_REMOVE_IN_PROG is set, and idpf_shutdown() calls
idpf_deinit_dflt_mbx() immediately after the trigger, which frees the ASQ
while that message is still outstanding. That is an enqueue-then-free of
DMA memory, and on a system with the IOMMU translating it is exactly the
class of fault this series is trying to remove. On the PF path
idpf_trigger_reset() only writes PFGEN_CTRL, which is why our testing
never hit it.

I will drop this patch from v2 rather than paper over it with a flag.
It was never load bearing: the stray writes are stopped by patch 1, and
this one only restored parity with idpf_remove(). If it turns out to be
worth having, it belongs in its own series with the VF path handled
properly.

Thanks for catching it.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
  2026-09-18 17:59   ` Tantilov, Emil S
@ 2026-09-21  3:27     ` Tian Xun Ng
  2026-09-21 21:06       ` Tantilov, Emil S
  0 siblings, 1 reply; 12+ messages in thread
From: Tian Xun Ng @ 2026-09-21  3:27 UTC (permalink / raw)
  To: emil.s.tantilov
  Cc: Tian Xun Ng, intel-wired-lan, netdev, anthony.l.nguyen,
	przemyslaw.kitszel, aleksander.lobakin, aleksandr.loktionov,
	andrew+netdev, davem, edumazet, kuba, pabeni, Tian Xun Ng

On 9/18/2026, Tantilov, Emil S wrote:
> Actually we can't wait on shutdown. If the MBX is defunct, like CP is
> down or unresponsive, the shutdown will hang for a very long time. This
> is the reason why we wanted to avoid communication on shutdown. Have you
> tested the shutdown after stopping the control plane?

No, I have not, and I cannot on this platform: the control plane sits
behind the device and I have no way to stop it from the host. So I have
to take your point as given, and as written the patch is not acceptable:
each teardown transaction uses IDPF_VC_XN_DEFAULT_TIMEOUT_MSEC (60 s),
and the teardown per vport is disable_vport, disable_queues (which also
waits for the SW marker) and destroy_vport. With a dead CP and two vports
that is minutes of hang on every reboot, to remove a fault that only
shows up on warm reboots. That trade is wrong.

What I would like to propose for v2 is to keep the teardown but bound it:
a shutdown-specific timeout, on the order of a second or two, used for
those three transactions when the driver is shutting down. If the CP
answers, the device is told to stop its queues and the stray writes go
away; if it does not, shutdown loses a bounded couple of seconds instead
of minutes. Does that direction look acceptable to you, and is there a
timeout value you would consider safe? If you would rather not have any
mailbox traffic on shutdown at all, then I think the fix has to come from
the device side instead, and I would rather know that before sending v2.

> This logic already exists in the reset handling, there should be no need
> to replicate it here. Do you have a trace and/or exact scenario that
> leads to remove being called while in a reset, but MBX is still alive?

No, I do not have such a trace. I added idpf_is_reset_detected() defensively
rather than from an observed case, and I will drop it in v2.

For the record, what we do see without any of this, on arm64 with two idpf
functions: after a warm reboot the device still has its queues enabled with
the previous kernel's ring addresses, and the first queue reconfiguration in
the next boot makes it write SW_MARKER completions into memory that kernel
has already reused. 20 of 20 warm reboots on an unpatched control node, none
in 151 with the teardown messages delivered.

Thanks for the review.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
  2026-09-21  3:27     ` Tian Xun Ng
@ 2026-09-21 21:06       ` Tantilov, Emil S
  2026-10-06 10:56         ` Tian Xun Ng
  0 siblings, 1 reply; 12+ messages in thread
From: Tantilov, Emil S @ 2026-09-21 21:06 UTC (permalink / raw)
  To: Tian Xun Ng
  Cc: intel-wired-lan, netdev, anthony.l.nguyen, przemyslaw.kitszel,
	aleksander.lobakin, aleksandr.loktionov, andrew+netdev, davem,
	edumazet, kuba, pabeni, Tian Xun Ng



On 9/20/2026 8:27 PM, Tian Xun Ng wrote:
> On 9/18/2026, Tantilov, Emil S wrote:
>> Actually we can't wait on shutdown. If the MBX is defunct, like CP is
>> down or unresponsive, the shutdown will hang for a very long time. This
>> is the reason why we wanted to avoid communication on shutdown. Have you
>> tested the shutdown after stopping the control plane?
> 
> No, I have not, and I cannot on this platform: the control plane sits
> behind the device and I have no way to stop it from the host. So I have
> to take your point as given, and as written the patch is not acceptable:
> each teardown transaction uses IDPF_VC_XN_DEFAULT_TIMEOUT_MSEC (60 s),
> and the teardown per vport is disable_vport, disable_queues (which also
> waits for the SW marker) and destroy_vport. With a dead CP and two vports
> that is minutes of hang on every reboot, to remove a fault that only
> shows up on warm reboots. That trade is wrong.
> 
> What I would like to propose for v2 is to keep the teardown but bound it:
> a shutdown-specific timeout, on the order of a second or two, used for
> those three transactions when the driver is shutting down. If the CP
> answers, the device is told to stop its queues and the stray writes go
> away; if it does not, shutdown loses a bounded couple of seconds instead
> of minutes. Does that direction look acceptable to you, and is there a
> timeout value you would consider safe? If you would rather not have any
> mailbox traffic on shutdown at all, then I think the fix has to come from
> the device side instead, and I would rather know that before sending v2.

If there is some clean way to shortcut the MBX on shutdown then I guess 
it would be acceptable, but I don't know what a "safe" timeout would be. 
As you can see the timeouts are already quite long, so you could 
potentially still bail out on a working CP that just so happens to be 
busy on the replies.

Also, consider the case where a reset on the PF will kill the MBX for 
the VFs associated with it, so a shutdown on such a VF will always end 
up timing out, since the VF reset is a message to the FW.

> 
>> This logic already exists in the reset handling, there should be no need
>> to replicate it here. Do you have a trace and/or exact scenario that
>> leads to remove being called while in a reset, but MBX is still alive?
> 
> No, I do not have such a trace. I added idpf_is_reset_detected() defensively
> rather than from an observed case, and I will drop it in v2.

OK, that makes more sense. I was curious if you are actually seeing an 
issue in your testing.

> 
> For the record, what we do see without any of this, on arm64 with two idpf
> functions: after a warm reboot the device still has its queues enabled with
> the previous kernel's ring addresses, and the first queue reconfiguration in
> the next boot makes it write SW_MARKER completions into memory that kernel
> has already reused. 20 of 20 warm reboots on an unpatched control node, none
> in 151 with the teardown messages delivered.

In that case wouldn't just the reset on shutdown be sufficient? The FW 
should clear the resources on reset, which should take care of the stale 
vports.

> 
> Thanks for the review.
> 

Thanks,
Emil

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
  2026-09-21 21:06       ` Tantilov, Emil S
@ 2026-10-06 10:56         ` Tian Xun Ng
  2026-10-06 20:03           ` Tantilov, Emil S
  0 siblings, 1 reply; 12+ messages in thread
From: Tian Xun Ng @ 2026-10-06 10:56 UTC (permalink / raw)
  To: Emil Tantilov
  Cc: intel-wired-lan, netdev, anthony.l.nguyen, przemyslaw.kitszel,
	aleksander.lobakin, aleksandr.loktionov, andrew+netdev, davem,
	edumazet, kuba, pabeni, Tian Xun Ng

On 9/21/2026 2:06 PM, Tantilov, Emil S wrote:
> In that case wouldn't just the reset on shutdown be sufficient? The FW
> should clear the resources on reset, which should take care of the stale
> vports.

Sorry for the slow reply; I wanted data before answering. On these
devices it is not sufficient.

Same arm64 hosts with two idpf PFs, IOMMU translating so every stray
write is blocked and logged as an F_TRANSLATION fault, warm reboot in a
loop, counting boots with idpf faults (always ~19 s in, at the first
queue reconfiguration):

  shutdown as in net-queue (mailbox shut first, no reset):    10 of 10
  same + PF reset after idpf_vc_core_deinit():                 2 of 2
  same + PF reset, then poll PFGEN_RSTAT for completion:       10 of 10
  teardown messages delivered, no reset (IDPF_REMOVE_IN_PROG
  set in idpf_shutdown(), the effect of this patch):           0 of 100

In the third case the serial console shows the reset happening:
PFR_STATE goes from 0x2 to 0x1 within the poll on both functions, and
reading the register mid-reset raises a TLP error on the function, so
the reset is not being lost to the reboot. The device still has the
previous kernel's queues afterwards. The reset the next kernel does at
probe (IDPF_HR_DRV_LOAD) does not clear them either; only the
disable/destroy messages do.

Is a PF reset expected to drop the vport and queue configuration on
your parts? If it is, this looks like a device firmware issue on our
side and I will raise it with the vendor. Either way the driver cannot
rely on it here.

> If there is some clean way to shortcut the MBX on shutdown then I guess
> it would be acceptable, but I don't know what a "safe" timeout would be.
> As you can see the timeouts are already quite long, so you could
> potentially still bail out on a working CP that just so happens to be
> busy on the replies.
>
> Also, consider the case where a reset on the PF will kill the MBX for
> the VFs associated with it, so a shutdown on such a VF will always end
> up timing out, since the VF reset is a message to the FW.

Understood. Given the above, some mailbox traffic on shutdown seems
unavoidable if the stale queues are to go away. What I would propose for
v2:

  - keep the vport teardown on shutdown, but give those transactions
    a short shutdown-only timeout, and stop at the first timeout
    instead of waiting on every remaining message, so a dead CP costs
    one timeout rather than several minutes;
  - skip it on a VF, where the PF/CP owns the VF's resources and the
    mailbox may already be gone;
  - drop idpf_is_reset_detected() as you suggested, and drop patch 2.

A busy CP that misses the short timeout leaves us where net-queue is
today, which is no worse than now. Would that be acceptable?

Thanks,
Tian Xun

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
  2026-10-06 10:56         ` Tian Xun Ng
@ 2026-10-06 20:03           ` Tantilov, Emil S
  0 siblings, 0 replies; 12+ messages in thread
From: Tantilov, Emil S @ 2026-10-06 20:03 UTC (permalink / raw)
  To: Tian Xun Ng
  Cc: intel-wired-lan, netdev, anthony.l.nguyen, przemyslaw.kitszel,
	aleksander.lobakin, aleksandr.loktionov, andrew+netdev, davem,
	edumazet, kuba, pabeni, Tian Xun Ng, decot@google.com



On 10/6/2026 3:56 AM, Tian Xun Ng wrote:
> On 9/21/2026 2:06 PM, Tantilov, Emil S wrote:
>> In that case wouldn't just the reset on shutdown be sufficient? The FW
>> should clear the resources on reset, which should take care of the stale
>> vports.
> 
> Sorry for the slow reply; I wanted data before answering. On these
> devices it is not sufficient.

No worries, thanks for following up.

> 
> Same arm64 hosts with two idpf PFs, IOMMU translating so every stray
> write is blocked and logged as an F_TRANSLATION fault, warm reboot in a
> loop, counting boots with idpf faults (always ~19 s in, at the first
> queue reconfiguration):
> 
>    shutdown as in net-queue (mailbox shut first, no reset):    10 of 10
>    same + PF reset after idpf_vc_core_deinit():                 2 of 2
>    same + PF reset, then poll PFGEN_RSTAT for completion:       10 of 10
>    teardown messages delivered, no reset (IDPF_REMOVE_IN_PROG
>    set in idpf_shutdown(), the effect of this patch):           0 of 100

I think this change is warranted - setting IDPF_REMOVE_IN_PROG in 
idpf_shutdown(). We just need to figure out a way to bail early on 
unresponsive FW. For the moment I think it will be OK if we prioritize 
the working case and make sure a shutdown is handled properly.

> 
> In the third case the serial console shows the reset happening:
> PFR_STATE goes from 0x2 to 0x1 within the poll on both functions, and
> reading the register mid-reset raises a TLP error on the function, so
> the reset is not being lost to the reboot. The device still has the
> previous kernel's queues afterwards. The reset the next kernel does at
> probe (IDPF_HR_DRV_LOAD) does not clear them either; only the
> disable/destroy messages do.
> 
> Is a PF reset expected to drop the vport and queue configuration on
> your parts? If it is, this looks like a device firmware issue on our
> side and I will raise it with the vendor. Either way the driver cannot
> rely on it here.

In general yes, a reset is expected to provide a clean slate for the 
driver. This is why the driver issues a reset on load. Even if we make 
the shutdown logic flawless there is no guarantee that it will actually 
run. In some cases like a system hang, we'd still end up being reliant 
on the reset clearing the state.

> 
>> If there is some clean way to shortcut the MBX on shutdown then I guess
>> it would be acceptable, but I don't know what a "safe" timeout would be.
>> As you can see the timeouts are already quite long, so you could
>> potentially still bail out on a working CP that just so happens to be
>> busy on the replies.
>>
>> Also, consider the case where a reset on the PF will kill the MBX for
>> the VFs associated with it, so a shutdown on such a VF will always end
>> up timing out, since the VF reset is a message to the FW.
> 
> Understood. Given the above, some mailbox traffic on shutdown seems
> unavoidable if the stale queues are to go away. What I would propose for
> v2:
> 
>    - keep the vport teardown on shutdown, but give those transactions
>      a short shutdown-only timeout, and stop at the first timeout
>      instead of waiting on every remaining message, so a dead CP costs
>      one timeout rather than several minutes;
>    - skip it on a VF, where the PF/CP owns the VF's resources and the
>      mailbox may already be gone;
>    - drop idpf_is_reset_detected() as you suggested, and drop patch 2.
> 
> A busy CP that misses the short timeout leaves us where net-queue is
> today, which is no worse than now. Would that be acceptable?

Agreed, this seems like a decent approach.

Thanks,
Emil
> 
> Thanks,
> Tian Xun


^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-10-06 20:04 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 10:52 [PATCH iwl-net 0/2] idpf: stop stray device writes after a warm reboot Tian Xun Ng
2026-09-17 10:52 ` [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown Tian Xun Ng
2026-09-18 15:45   ` Loktionov, Aleksandr
2026-09-18 17:59   ` Tantilov, Emil S
2026-09-21  3:27     ` Tian Xun Ng
2026-09-21 21:06       ` Tantilov, Emil S
2026-10-06 10:56         ` Tian Xun Ng
2026-10-06 20:03           ` Tantilov, Emil S
2026-09-17 10:52 ` [PATCH iwl-net 2/2] idpf: reset the function " Tian Xun Ng
2026-09-18 15:46   ` Loktionov, Aleksandr
2026-09-18 17:54   ` Tantilov, Emil S
2026-09-21  3:27     ` Tian Xun Ng

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox