Netdev List
 help / color / mirror / Atom feed
From: "Tantilov, Emil S" <emil.s.tantilov@intel.com>
To: Tian Xun Ng <luckilystar08@gmail.com>
Cc: <intel-wired-lan@lists.osuosl.org>, <netdev@vger.kernel.org>,
	<anthony.l.nguyen@intel.com>, <przemyslaw.kitszel@intel.com>,
	<aleksander.lobakin@intel.com>, <aleksandr.loktionov@intel.com>,
	<andrew+netdev@lunn.ch>, <davem@davemloft.net>,
	<edumazet@google.com>, <kuba@kernel.org>, <pabeni@redhat.com>,
	Tian Xun Ng <tianxun.ng@bytedance.com>,
	"decot@google.com" <decot@google.com>
Subject: Re: [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown
Date: Tue, 6 Oct 2026 13:03:52 -0700	[thread overview]
Message-ID: <f6a32b39-6319-4389-bb7b-986e840589bc@intel.com> (raw)
In-Reply-To: <20261006105629.7481-1-luckilystar08@gmail.com>



On 10/6/2026 3:56 AM, Tian Xun Ng wrote:
> On 9/21/2026 2:06 PM, Tantilov, Emil S wrote:
>> In that case wouldn't just the reset on shutdown be sufficient? The FW
>> should clear the resources on reset, which should take care of the stale
>> vports.
> 
> Sorry for the slow reply; I wanted data before answering. On these
> devices it is not sufficient.

No worries, thanks for following up.

> 
> Same arm64 hosts with two idpf PFs, IOMMU translating so every stray
> write is blocked and logged as an F_TRANSLATION fault, warm reboot in a
> loop, counting boots with idpf faults (always ~19 s in, at the first
> queue reconfiguration):
> 
>    shutdown as in net-queue (mailbox shut first, no reset):    10 of 10
>    same + PF reset after idpf_vc_core_deinit():                 2 of 2
>    same + PF reset, then poll PFGEN_RSTAT for completion:       10 of 10
>    teardown messages delivered, no reset (IDPF_REMOVE_IN_PROG
>    set in idpf_shutdown(), the effect of this patch):           0 of 100

I think this change is warranted - setting IDPF_REMOVE_IN_PROG in 
idpf_shutdown(). We just need to figure out a way to bail early on 
unresponsive FW. For the moment I think it will be OK if we prioritize 
the working case and make sure a shutdown is handled properly.

> 
> In the third case the serial console shows the reset happening:
> PFR_STATE goes from 0x2 to 0x1 within the poll on both functions, and
> reading the register mid-reset raises a TLP error on the function, so
> the reset is not being lost to the reboot. The device still has the
> previous kernel's queues afterwards. The reset the next kernel does at
> probe (IDPF_HR_DRV_LOAD) does not clear them either; only the
> disable/destroy messages do.
> 
> Is a PF reset expected to drop the vport and queue configuration on
> your parts? If it is, this looks like a device firmware issue on our
> side and I will raise it with the vendor. Either way the driver cannot
> rely on it here.

In general yes, a reset is expected to provide a clean slate for the 
driver. This is why the driver issues a reset on load. Even if we make 
the shutdown logic flawless there is no guarantee that it will actually 
run. In some cases like a system hang, we'd still end up being reliant 
on the reset clearing the state.

> 
>> If there is some clean way to shortcut the MBX on shutdown then I guess
>> it would be acceptable, but I don't know what a "safe" timeout would be.
>> As you can see the timeouts are already quite long, so you could
>> potentially still bail out on a working CP that just so happens to be
>> busy on the replies.
>>
>> Also, consider the case where a reset on the PF will kill the MBX for
>> the VFs associated with it, so a shutdown on such a VF will always end
>> up timing out, since the VF reset is a message to the FW.
> 
> Understood. Given the above, some mailbox traffic on shutdown seems
> unavoidable if the stale queues are to go away. What I would propose for
> v2:
> 
>    - keep the vport teardown on shutdown, but give those transactions
>      a short shutdown-only timeout, and stop at the first timeout
>      instead of waiting on every remaining message, so a dead CP costs
>      one timeout rather than several minutes;
>    - skip it on a VF, where the PF/CP owns the VF's resources and the
>      mailbox may already be gone;
>    - drop idpf_is_reset_detected() as you suggested, and drop patch 2.
> 
> A busy CP that misses the short timeout leaves us where net-queue is
> today, which is no worse than now. Would that be acceptable?

Agreed, this seems like a decent approach.

Thanks,
Emil
> 
> Thanks,
> Tian Xun


  reply	other threads:[~2026-10-06 20:04 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-17 10:52 [PATCH iwl-net 0/2] idpf: stop stray device writes after a warm reboot Tian Xun Ng
2026-09-17 10:52 ` [PATCH iwl-net 1/2] idpf: keep the mailbox up while tearing down vports on shutdown Tian Xun Ng
2026-09-18 15:45   ` Loktionov, Aleksandr
2026-09-18 17:59   ` Tantilov, Emil S
2026-09-21  3:27     ` Tian Xun Ng
2026-09-21 21:06       ` Tantilov, Emil S
2026-10-06 10:56         ` Tian Xun Ng
2026-10-06 20:03           ` Tantilov, Emil S [this message]
2026-09-17 10:52 ` [PATCH iwl-net 2/2] idpf: reset the function " Tian Xun Ng
2026-09-18 15:46   ` Loktionov, Aleksandr
2026-09-18 17:54   ` Tantilov, Emil S
2026-09-21  3:27     ` Tian Xun Ng

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=f6a32b39-6319-4389-bb7b-986e840589bc@intel.com \
    --to=emil.s.tantilov@intel.com \
    --cc=aleksander.lobakin@intel.com \
    --cc=aleksandr.loktionov@intel.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=anthony.l.nguyen@intel.com \
    --cc=davem@davemloft.net \
    --cc=decot@google.com \
    --cc=edumazet@google.com \
    --cc=intel-wired-lan@lists.osuosl.org \
    --cc=kuba@kernel.org \
    --cc=luckilystar08@gmail.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=przemyslaw.kitszel@intel.com \
    --cc=tianxun.ng@bytedance.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox