DPDK-dev Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Stephen Hemminger <stephen@networkplumber.org>
To: Rita Ruvinsky <rita.ruvinsky@weka.io>
Cc: dev@dpdk.org, longli@microsoft.com, weh@microsoft.com, stable@dpdk.org
Subject: Re: [PATCH] net/mana: fix Tx stall from send queue free-space unit mismatch
Date: Mon, 14 Sep 2026 09:14:19 -0700	[thread overview]
Message-ID: <20260914091419.22e3ef6e@phoenix.local> (raw)
In-Reply-To: <20260914105809.919580-1-rita.ruvinsky@weka.io>

On Mon, 14 Sep 2026 13:58:08 +0300
Rita Ruvinsky <rita.ruvinsky@weka.io> wrote:

> gdma_post_work_request() subtracted a unit count from an entry count:
> 
>   queue_free_units = queue->count - (queue->head - queue->tail);
> 
> queue->count is in entries, while head and tail are in WQE alignment
> units. On a 512-entry, 128KB send queue the check saw 512 units of
> capacity instead of queue->size / GDMA_WQE_ALIGNMENT_UNIT_SIZE = 4096,
> and returned -EBUSY with the queue one eighth full. A workload that
> fills that window faster than it drains makes rte_eth_tx_burst() return
> 0 for long enough to look like a dead port.
> 
> Derive the capacity from queue->size, which is also what the ring wrap
> in gdma_get_wqe_pointer() uses. Rx is unaffected: its WQEs occupy
> exactly one unit, so entries and units coincide.
> 
> Fixes: 56dd45c0ce7b ("net/mana: implement hardware layer operations")
> Cc: stable@dpdk.org
> 
> Signed-off-by: Rita Ruvinsky <rita.ruvinsky@weka.io>
> ---

Applied to next-net

The long form AI review had some observations worth including:

On Mon, 14 Sep 2026 13:58:08 +0300
Rita Ruvinsky <rita.ruvinsky@weka.io> wrote:

> gdma_post_work_request() subtracted a unit count from an entry count:

The unit analysis is right.  head/tail are advanced in alignment
units (queue->head += wqe_size / GDMA_WQE_ALIGNMENT_UNIT_SIZE, and
gdma_get_wqe_pointer() multiplies head by the same constant), while
sq_count comes from rdma-core as attr->cap.max_send_wr and sq_size as
align_hw_size(max_send_wr * get_wqe_size(max_send_sge)).  Deriving the
capacity from size is the only self-consistent choice, and it is what
mana_gd_wq_avail_space() in the kernel driver does.

Info:

1. The debug line in the -EBUSY path still reports queue->count:

	DP_LOG(DEBUG, "WQE size %u queue count %u head %u tail %u",
	       wqe_size, queue->count, queue->head, queue->tail);

After this patch count no longer takes part in the decision for the
send or receive queue; only gdma_poll_completion_queue() still uses it,
for the CQ.  The one line printed when a post is rejected no longer
shows what it was rejected against.  Suggest:

	DP_LOG(DEBUG, "WQE size %u queue size %u free %u head %u tail %u",
	       wqe_size, queue->size, queue_free_units,
	       queue->head, queue->tail);

2. The comment describes the old bug rather than the invariant:

	/* head/tail count WQE alignment units, so the capacity they are
	 * compared against must too: queue->count is in entries and
	 * undercounts the queue, stalling Tx well below capacity.
	 */

The stall belongs in the commit message, where it already is.  In the
source the invariant is enough:

	/* head and tail are in WQE alignment units, so the capacity must
	 * come from the queue size in bytes, not the entry count.
	 */

3. Worth a sentence in the commit message that the kernel mana driver
computes the same limit in mana_gd_wq_avail_space(), in bytes:

	u32 used_space = (wq->head - wq->tail) * GDMA_WQE_BU_SIZE;
	return wq->queue_size - used_space;

It is independent confirmation of the unit convention and tells anyone
backporting this that the two drivers now agree.

  reply	other threads:[~2026-09-14 16:14 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-14 10:58 [PATCH] net/mana: fix Tx stall from send queue free-space unit mismatch Rita Ruvinsky
2026-09-14 16:14 ` Stephen Hemminger [this message]
2026-09-14 16:24   ` [EXTERNAL] " Wei Hu
2026-09-16 11:12     ` Rita Ruvinsky
2026-09-16 15:17       ` Stephen Hemminger
2026-09-17  6:23 ` [PATCH v2] " Rita Ruvinsky

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260914091419.22e3ef6e@phoenix.local \
    --to=stephen@networkplumber.org \
    --cc=dev@dpdk.org \
    --cc=longli@microsoft.com \
    --cc=rita.ruvinsky@weka.io \
    --cc=stable@dpdk.org \
    --cc=weh@microsoft.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox