Linux wireless drivers development
 help / color / mirror / Atom feed
From: Baochen Qiang <baochen.qiang@oss.qualcomm.com>
To: Michael Pfeifroth <micpf@westermo.com>,
	Jeff Johnson <jjohnson@kernel.org>
Cc: Kalle Valo <kvalo@kernel.org>,
	ath11k@lists.infradead.org, linux-wireless@vger.kernel.org,
	linux-kernel@vger.kernel.org
Subject: Re: [PATCH v3] wifi: ath11k: release peer accounting on peer delete timeout
Date: Thu, 8 Oct 2026 19:02:38 +0800	[thread overview]
Message-ID: <c3564de9-c228-429e-8e04-3ff65619d1fc@oss.qualcomm.com> (raw)
In-Reply-To: <ar0LVqHZF8UlRB0o@mpf-ESPRIMO-P9012>



On 9/30/2026 9:15 PM, Michael Pfeifroth wrote:
> ath11k_peer_delete() only decrements ar->num_peers when
> __ath11k_peer_delete() returns 0. On a peer-delete-confirmation timeout
> __ath11k_peer_delete() returns -ETIMEDOUT, so the decrement is skipped
> and one ar->num_peers slot is leaked. Once enough slots have leaked the
> ar->num_peers > (ar->max_num_peers - 1) gate in ath11k_peer_create()
> rejects every new station with "insufficient peer entry resource in
> firmware", and the AP stops accepting associations until the radio is
> restarted with a wifi down/up.
> 
> This was observed in the field on an AP after hours of uptime with
> frequent roaming and reconnects: clients could no longer associate even
> though the firmware peer table was not actually exhausted, only the
> host-side ar->num_peers accounting had leaked.
> 
> The delete-confirmation timeout itself is otherwise harmless. The host
> first waits for the peer unmap event in ath11k_wait_for_peer_deleted(),
> so by the time the delete-response completion times out the peer has
> already been removed from ab->peers and freed by
> ath11k_peer_unmap_event(). Only the ar->num_peers counter is left
> inconsistent.
> 
> The timeout is reached when the peer unmap event arrives but the peer
> delete response is missed, e.g. because the response event is dropped in
> ath11k_peer_delete_resp_event() on an unresolved vdev id (logged as
> "invalid vdev id in peer delete resp ev"). Do not treat the delete
> confirmation timeout as fatal so that ath11k_peer_delete() releases the
> ar->num_peers slot instead of leaking it.
> 
> Fixes: 690ace20ff79 ("ath11k: peer delete synchronization with firmware")
> Cc: stable@vger.kernel.org
> Signed-off-by: Michael Pfeifroth <micpf@westermo.com>
> ---
> v3:
>  - Drop the defensive peer list_del()/kfree() branch; it is dead code
>    since the peer is already freed via the unmap event or on recovery
>    (Baochen Qiang), leaving a minimal fix that just ignores the delete
>    timeout.
>  - Reframe the commit message around the field-observed num_peers leak
>    and the resulting station association failures.
> v2:
>  - Correct the root-cause description and switch to netdev comment style.
>  drivers/net/wireless/ath/ath11k/peer.c | 7 ++++---
>  1 file changed, 4 insertions(+), 3 deletions(-)
> 
> diff --git a/drivers/net/wireless/ath/ath11k/peer.c b/drivers/net/wireless/ath/ath11k/peer.c
> index b30a906..f573a2c 100644
> --- a/drivers/net/wireless/ath/ath11k/peer.c
> +++ b/drivers/net/wireless/ath/ath11k/peer.c
> @@ -340,9 +340,10 @@ static int __ath11k_peer_delete(struct ath11k *ar, u32 vdev_id, const u8 *addr)
>  		return ret;
>  	}
>  
> -	ret = ath11k_wait_for_peer_delete_done(ar, vdev_id, addr);
> -	if (ret)
> -		return ret;
> +	/* Ignore the return value: the peer is already freed, only its
> +	 * ar->num_peers slot would otherwise leak on a delete timeout.
> +	 */
> +	ath11k_wait_for_peer_delete_done(ar, vdev_id, addr);

as stated in the commit message, the purpose is to ease the case where firmware peer table
is not exhausted but host's table is, however the change here is to ignore the return
value by default, regardless of whether firmware peer table has free slots. So what about
ath11k_wait_for_peer_delete_done() timeouts due to firmware peer table exhaustion in fact?

More importantly, host rejects new peer association only after the

ar->num_peers > (ar->max_num_peers - 1)

But if we really hit it, I would suspect something wrong with firmware. So even if we give
it one more chance I don't think we can survive in the end.

>  
>  	return 0;
>  }


      reply	other threads:[~2026-10-08 11:02 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 12:25 [PATCH] wifi: ath11k: release peer accounting on peer delete timeout Michael Pfeifroth
2026-09-24  7:29 ` Baochen Qiang
2026-09-24  9:13   ` Michael Pfeifroth
2026-09-24  9:14   ` [PATCH v2] " Michael Pfeifroth
2026-09-29  6:19     ` Baochen Qiang
2026-09-30 13:15       ` [PATCH v3] " Michael Pfeifroth
2026-10-08 11:02         ` Baochen Qiang [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=c3564de9-c228-429e-8e04-3ff65619d1fc@oss.qualcomm.com \
    --to=baochen.qiang@oss.qualcomm.com \
    --cc=ath11k@lists.infradead.org \
    --cc=jjohnson@kernel.org \
    --cc=kvalo@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-wireless@vger.kernel.org \
    --cc=micpf@westermo.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox