All of lore.kernel.org
 help / color / mirror / Atom feed
From: Bobby Eshleman <bobbyeshleman@gmail.com>
To: Tariq Toukan <tariqt@nvidia.com>
Cc: Saeed Mahameed <saeedm@nvidia.com>,
	Mark Bloch <mbloch@nvidia.com>, Leon Romanovsky <leon@kernel.org>,
	Andrew Lunn <andrew+netdev@lunn.ch>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Dragos Tatulea <dtatulea@nvidia.com>,
	Mina Almasry <almasrymina@google.com>,
	netdev@vger.kernel.org, linux-rdma@vger.kernel.org,
	linux-kernel@vger.kernel.org,
	Bobby Eshleman <bobbyeshleman@meta.com>
Subject: Re: [PATCH net v3] net/mlx5e: fix BQL reset on SQ re-activation
Date: Tue, 4 Aug 2026 10:14:49 -0700	[thread overview]
Message-ID: <anIeCTh6IUJaP2mz@devvm29614.prn0.facebook.com> (raw)
In-Reply-To: <6ac643bf-c6f4-4ff2-ba57-964a645d9323@nvidia.com>

On Tue, Aug 04, 2026 at 11:23:53AM +0300, Tariq Toukan wrote:
> 
> 
> On 04/08/2026 2:47, Bobby Eshleman wrote:
> > From: Bobby Eshleman <bobbyeshleman@meta.com>
> > 
> > mlx5e_queue_start() deactivates and re-activates all channels but closes
> > only the queue being restarted. mlx5e_activate_txqsq() then
> > unconditionally calls netdev_tx_reset_queue(), zeroing the BQL counters
> > of channels that kept their in-flight TX WQEs. The next completion then
> > over-charges and trips the BUG_ON() in dql_completed():
> > 
> >    kernel BUG at lib/dynamic_queue_limits.c:99!
> >    RIP: 0010:dql_completed+0x23d/0x280
> >    Call Trace:
> >     <IRQ>
> >     mlx5e_poll_tx_cq+0x668/0xa60
> >     mlx5e_napi_poll+0x5b/0x7b0
> >     net_rx_action+0x15a/0x580
> > 
> > Reset BQL only when the SQ has no bytes in flight (sq->cc == sq->pc).
> > 
> > In the case that reset is skipped, the outstanding WQEs will eventually
> > complete and rebalance the dql. The dql->limit is carried across the
> > reset.
> > 
> > Fixes: b2588ea40ec9 ("net/mlx5e: Implement queue mgmt ops and single channel swap")
> > Signed-off-by: Bobby Eshleman <bobbyeshleman@meta.com>
> 
> Reviewed-by: Tariq Toukan <tariqt@nvidia.com>
> 
> Thanks.

Thanks Tariq, and my apologies for forgetting to carry the R-b forward.

> 
> > ---
> > Changes in v3:
> > - Sample sq->cc/sq->pc before setting MLX5E_SQ_STATE_ENABLED, so a
> >    concurrent mlx5e_poll_tx_cq() cannot un-charge BQL between the check
> >    and netdev_tx_reset_queue() (Sashiko:
> >    https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260729-mlx5-bql-v2-1-cdf285e586df@meta.com)
> > - Link to v2: https://lore.kernel.org/r/20260729-mlx5-bql-v2-1-cdf285e586df@meta.com
> > 
> > Changes in v2:
> > - remove misleading comment (Tariq)
> > - Link to v1: https://lore.kernel.org/r/20260728-mlx5-bql-v1-1-99a0ab77ede7@meta.com
> > ---
> >   drivers/net/ethernet/mellanox/mlx5/core/en_main.c | 4 +++-
> >   1 file changed, 3 insertions(+), 1 deletion(-)
> > 
> > diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
> > index 7d47a1da8b6b..f0407a850ea8 100644
> > --- a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
> > +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
> > @@ -1939,8 +1939,10 @@ int mlx5e_open_txqsq(struct mlx5e_channel *c, u32 tisn, int txq_ix,
> >   void mlx5e_activate_txqsq(struct mlx5e_txqsq *sq)
> >   {
> >   	sq->txq = netdev_get_tx_queue(sq->netdev, sq->txq_ix);
> > +	/* Reset BQL only when the SQ has no bytes in flight. */
> > +	if (sq->cc == sq->pc)
> > +		netdev_tx_reset_queue(sq->txq);
> >   	set_bit(MLX5E_SQ_STATE_ENABLED, &sq->state);
> > -	netdev_tx_reset_queue(sq->txq);
> >   	netif_tx_start_queue(sq->txq);
> >   	netif_queue_set_napi(sq->netdev, sq->txq_ix, NETDEV_QUEUE_TYPE_TX, sq->cq.napi);
> >   }
> > 
> > ---
> > base-commit: d1000fd7995e51deec872d154e0a40d82f7a539f
> > change-id: 20260728-mlx5-bql-84e02a5a46c8
> > 
> > Best regards,
> 

  reply	other threads:[~2026-08-04 17:14 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-03 23:47 [PATCH net v3] net/mlx5e: fix BQL reset on SQ re-activation Bobby Eshleman
2026-08-04  8:23 ` Tariq Toukan
2026-08-04 17:14   ` Bobby Eshleman [this message]
2026-08-05  2:30 ` patchwork-bot+netdevbpf

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=anIeCTh6IUJaP2mz@devvm29614.prn0.facebook.com \
    --to=bobbyeshleman@gmail.com \
    --cc=almasrymina@google.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=bobbyeshleman@meta.com \
    --cc=davem@davemloft.net \
    --cc=dtatulea@nvidia.com \
    --cc=edumazet@google.com \
    --cc=kuba@kernel.org \
    --cc=leon@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=mbloch@nvidia.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=saeedm@nvidia.com \
    --cc=tariqt@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.