From: Tian Xun Ng <luckilystar08@gmail.com>
To: intel-wired-lan@lists.osuosl.org
Cc: netdev@vger.kernel.org, anthony.l.nguyen@intel.com,
przemyslaw.kitszel@intel.com, aleksander.lobakin@intel.com,
emil.s.tantilov@intel.com, andrew+netdev@lunn.ch,
davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, Tian Xun Ng <tianxun.ng@bytedance.com>
Subject: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
Date: Thu, 17 Sep 2026 18:52:04 +0800 [thread overview]
Message-ID: <20260917105205.37561-2-luckilystar08@gmail.com> (raw)
In-Reply-To: <20260917105205.37561-1-luckilystar08@gmail.com>
From: Tian Xun Ng <tianxun.ng@bytedance.com>
Since commit 4c9106f4906a ("idpf: fix adapter NULL pointer dereference
on reboot"), idpf_shutdown() calls idpf_vc_core_deinit() directly instead
of idpf_remove(), so IDPF_REMOVE_IN_PROG is not set on shutdown.
idpf_vc_core_deinit() uses that flag to decide when to shut the virtchnl
transaction manager down. Without it, libie_ctlq_xn_shutdown() runs
before idpf_deinit_task() tears the vports down, so every message sent
during that teardown (disable vport, disable queues, destroy vport)
fails at once. The device is never told to stop its queues and keeps
them enabled, with the ring addresses of the kernel that is going away.
After a warm reboot, the first queue reconfiguration of a port that has
not been opened yet (udev setting the MTU) sends VIRTCHNL2_OP_DEL_QUEUES.
The device then drains the queues it still considers live and writes
SW_MARKER TX completions (qid_comptype_gen 0x2800 / 0xa800) to the
previous kernel's completion rings. With the IOMMU translating, those
writes fault on every such boot:
arm-smmu-v3 arm-smmu-v3.17.auto: event: F_TRANSLATION client: 0016:01:00.0 sid: 0x30100 ssid: 0x0 iova: 0xffb70000 ipa: 0x0
arm-smmu-v3 arm-smmu-v3.17.auto: unpriv data write s1 "Input address caused fault" stag: 0x0
In IOMMU pass-through mode they land on pages the new kernel has already
reused. With page_poison=1 that shows as "pagealloc: memory corruption"
on most boots. Without it, nodes crash in unrelated code, e.g.:
Unable to handle kernel NULL pointer dereference at virtual address 000000000000a830
pc : __tlb_remove_table_free+0x48/0x118
Call trace:
__tlb_remove_table_free
tlb_remove_table_rcu
rcu_do_batch
or with a slab free pointer that decodes from a slot holding 0x2800:
Unable to handle kernel paging request at virtual address 002613b73e862c6e
pc : kmem_cache_alloc_noprof+0xc0/0x3d0
The early shutdown exists to avoid waiting for transaction timeouts when
the mailbox is already gone. That is the hard reset case, where
idpf_vc_event_task() shuts the transaction manager down and sets
IDPF_HR_RESET_IN_PROG before idpf_init_hard_reset() calls
idpf_vc_core_deinit(). Key the early shutdown on a hard reset being in
progress or detected instead of on remove, so that both remove and
shutdown keep the mailbox up until the vports are gone.
Hard reset behaviour is unchanged. Remove is unchanged unless a hardware
reset has been detected, in which case it no longer waits for message
timeouts. Shutdown now delivers the teardown messages; if the device
stops responding without a detectable reset, shutdown can wait for the
transaction timeouts, as remove already does.
Tested on arm64 (64K pages) servers with two idpf functions, with this
change and the next patch backported to a 6.17 kernel: no stray device
writes in 151 warm reboots across IOMMU translated and pass-through
modes, against stray writes after 20 of 20 warm reboots without them.
Fixes: 4c9106f4906a ("idpf: fix adapter NULL pointer dereference on reboot")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Tian Xun Ng <tianxun.ng@bytedance.com>
---
.../net/ethernet/intel/idpf/idpf_virtchnl.c | 18 +++++++++++++-----
1 file changed, 13 insertions(+), 5 deletions(-)
diff --git a/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c b/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
index 1caf52706..646b6e074 100644
--- a/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
+++ b/drivers/net/ethernet/intel/idpf/idpf_virtchnl.c
@@ -3197,14 +3197,22 @@ int idpf_vc_core_init(struct idpf_adapter *adapter)
*/
void idpf_vc_core_deinit(struct idpf_adapter *adapter)
{
- bool remove_in_prog;
+ bool reset_in_prog;
if (!test_bit(IDPF_VC_CORE_INIT, adapter->flags))
return;
- /* Avoid transaction timeouts when called during reset */
- remove_in_prog = test_bit(IDPF_REMOVE_IN_PROG, adapter->flags);
- if (!remove_in_prog)
+ /* Shut the transaction manager down early only when the mailbox is
+ * already gone, i.e. a hard reset is in progress or has been detected,
+ * to avoid waiting for transaction timeouts. On remove and on shutdown
+ * the mailbox still works and must stay up until the vports are torn
+ * down. Otherwise the disable and destroy messages never reach the
+ * device, which keeps its queues enabled with ring addresses from this
+ * kernel after a warm reboot.
+ */
+ reset_in_prog = test_bit(IDPF_HR_RESET_IN_PROG, adapter->flags) ||
+ idpf_is_reset_detected(adapter);
+ if (reset_in_prog)
libie_ctlq_xn_shutdown(adapter->xnm);
idpf_ptp_release(adapter);
@@ -3213,7 +3221,7 @@ void idpf_vc_core_deinit(struct idpf_adapter *adapter)
idpf_rel_rx_pt_lkup(adapter);
idpf_intr_rel(adapter);
- if (remove_in_prog)
+ if (!reset_in_prog)
libie_ctlq_xn_shutdown(adapter->xnm);
cancel_delayed_work_sync(&adapter->serv_task);
--
2.50.1 (Apple Git-155)
next prev parent reply other threads:[~2026-09-17 16:07 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 10:52 [PATCH iwl-net 0/2] idpf: stop stray device writes after a warm reboot Tian Xun Ng
2026-09-17 10:52 ` Tian Xun Ng [this message]
2026-09-18 15:45 ` [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown Loktionov, Aleksandr
2026-09-18 17:59 ` Tantilov, Emil S
2026-09-21 3:27 ` Tian Xun Ng
2026-09-21 21:06 ` Tantilov, Emil S
2026-10-06 10:56 ` Tian Xun Ng
2026-10-06 20:03 ` Tantilov, Emil S
2026-09-17 10:52 ` [PATCH iwl-net 2/2] idpf: reset the function " Tian Xun Ng
2026-09-18 15:46 ` Loktionov, Aleksandr
2026-09-18 17:54 ` Tantilov, Emil S
2026-09-21 3:27 ` Tian Xun Ng
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260917105205.37561-2-luckilystar08@gmail.com \
--to=luckilystar08@gmail.com \
--cc=aleksander.lobakin@intel.com \
--cc=andrew+netdev@lunn.ch \
--cc=anthony.l.nguyen@intel.com \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=emil.s.tantilov@intel.com \
--cc=intel-wired-lan@lists.osuosl.org \
--cc=kuba@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=przemyslaw.kitszel@intel.com \
--cc=tianxun.ng@bytedance.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox