Netdev List
 help / color / mirror / Atom feed
* [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
@ 2026-09-25 17:43 Joe Damato
  2026-09-25 17:43 ` [RFC net v4 1/4] bnxt_en: return the RING_FREE status to callers Joe Damato
                   ` (3 more replies)
  0 siblings, 4 replies; 12+ messages in thread
From: Joe Damato @ 2026-09-25 17:43 UTC (permalink / raw)
  To: netdev
  Cc: andrew+netdev, davem, edumazet, kuba, pabeni, horms, michael.chan,
	pavan.chebbi, linux-kernel, Joe Damato

Greetings:

This is a follow up to the previous RFC (linked below), updated based on
feedback from Michael.

On two production systems, I saw the following dmesg pattern:

  NETDEV WATCHDOG: transmit queue 0 timed out 6073 ms
  Resp cmpl intr err msg: 0x51                  x20
  hwrm_ring_free type 1 failed                  x12
  hwrm_ring_free type 2 failed                  x8
  AMD-Vi: IO_PAGE_FAULT  x3

This suggests that, for some currently unknown reason, TX completions stall
and the netdev watchdog fires. The driver asks FW to free the rings, this
times out, but the driver ignores the possible failure and frees ring memory.
Since the FW didn't respond to the ring free command, it is possible that the
FW is still DMAing to the memory which was freed.

This series tries to prevent this by:

 - Returning and checking ring free command return values
 - Examining the FW response if the ring free command times out. It is
   possible that, for some reason, the FW did complete the ring free but was
   unable to respond with an IRQ. This seems unlikely given what appears to be
   a use after free in dmesg, but worth logging just in case.
 - Stop DMA before the driver frees ring memory, which should prevent
   any possible use after free.
 - Set a bit in the state flags to signal that DMA was stopped. This prevents
   the device from being reopened without user intervention.

In the future, this could possibly be extended to make recovery automatic.

Sending this as an RFC so that the Broadcom folks have some time to take a
look and test as needed.

Thanks,
Joe

v4:
  - No changes to patch 1 or 2
  - Patch 3: reworded the message logged when DMA is stopped, a rebind is
    what recovers the device, not a firmware reset. No functional change.
  - Patch 4 added which adds a new bit (BNXT_STATE_DMA_STOPPED) that is set
    when DMA is disabled. When this bit is set, the device requires user
    intervention to bring back up.

v3: https://lore.kernel.org/netdev/20260923210744.3406861-1-joe@dama.to/
  - No changes to patch 1
  - Patch 2: Don't poll for the valid bit as Michael suggested.
  - Patch 3: bnxt_hwrm_ring_free now returns -EIO instead of stopping the
    device, so the remaining resources can be freed and the remaining commands
    can be sent before stopping the device, as Michael suggested. Note the
    switch to using pci_clear_master in this patch instead of
    pci_disable_device. This was done so that the normal shutdown paths can
    call pci_disable_device without generating a warning.

v2: https://lore.kernel.org/netdev/20260922182405.1290749-1-joe@dama.to/
  - No changes to patch 1
  - Patch 2 from v1 dropped
  - Patch 2 in the v2 now checks the response and logs state before giving up
  - Patch 3 in the v2 disables the device to stop DMA before freeing ring
    memory

RFCv1: https://lore.kernel.org/netdev/20260917233218.1160001-1-joe@dama.to/

Joe Damato (4):
  bnxt_en: return the RING_FREE status to callers
  bnxt_en: check HWRM response if completion never arrives
  bnxt_en: stop DMA before releasing rings the firmware did not free
  bnxt_en: refuse to open a device with stopped DMA

 drivers/net/ethernet/broadcom/bnxt/bnxt.c     | 122 ++++++++++++------
 drivers/net/ethernet/broadcom/bnxt/bnxt.h     |   1 +
 .../net/ethernet/broadcom/bnxt/bnxt_hwrm.c    |  33 ++++-
 3 files changed, 115 insertions(+), 41 deletions(-)


base-commit: 11536ee3d3e0b1bd35b6f3f8df55a6053eb0c71d
-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 12+ messages in thread

* [RFC net v4 1/4] bnxt_en: return the RING_FREE status to callers
  2026-09-25 17:43 [RFC net v4 0/4] bnxt_en: Make RING FREE more robust Joe Damato
@ 2026-09-25 17:43 ` Joe Damato
  2026-09-25 17:43 ` [RFC net v4 2/4] bnxt_en: check HWRM response if completion never arrives Joe Damato
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 12+ messages in thread
From: Joe Damato @ 2026-09-25 17:43 UTC (permalink / raw)
  To: netdev, Michael Chan, Pavan Chebbi, Andrew Lunn, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Prashant Sreedharan
  Cc: edumazet, horms, linux-kernel, Joe Damato

hwrm_ring_free_send_msg() reports failure to its caller, returning -EIO
when the firmware rejects HWRM_RING_FREE or never answers it. All three
ring free helpers that send the command discard the value.

Return it instead. No caller acts on it yet, so there is no functional
change.

Fixes: 74608fc98d28 ("bnxt_en: Ring free response from close path should use completion ring")
Signed-off-by: Joe Damato <joe@dama.to>
---
 drivers/net/ethernet/broadcom/bnxt/bnxt.c | 48 +++++++++++++----------
 1 file changed, 27 insertions(+), 21 deletions(-)

diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
index d7728d0c5b6e..a7f6facca7b4 100644
--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
+++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
@@ -7660,50 +7660,55 @@ static int hwrm_ring_free_send_msg(struct bnxt *bp,
 	return 0;
 }
 
-static void bnxt_hwrm_tx_ring_free(struct bnxt *bp,
-				   struct bnxt_tx_ring_info *txr,
-				   bool close_path)
+static int bnxt_hwrm_tx_ring_free(struct bnxt *bp,
+				  struct bnxt_tx_ring_info *txr,
+				  bool close_path)
 {
 	struct bnxt_ring_struct *ring = &txr->tx_ring_struct;
 	u32 cmpl_ring_id;
+	int rc;
 
 	if (ring->fw_ring_id == INVALID_HW_RING_ID)
-		return;
+		return 0;
 
 	cmpl_ring_id = close_path ? bnxt_cp_ring_for_tx(bp, txr) :
 		       INVALID_HW_RING_ID;
-	hwrm_ring_free_send_msg(bp, ring, RING_FREE_REQ_RING_TYPE_TX,
-				cmpl_ring_id);
+	rc = hwrm_ring_free_send_msg(bp, ring, RING_FREE_REQ_RING_TYPE_TX,
+				     cmpl_ring_id);
 	ring->fw_ring_id = INVALID_HW_RING_ID;
+	return rc;
 }
 
-static void bnxt_hwrm_rx_ring_free(struct bnxt *bp,
-				   struct bnxt_rx_ring_info *rxr,
-				   bool close_path)
+static int bnxt_hwrm_rx_ring_free(struct bnxt *bp,
+				  struct bnxt_rx_ring_info *rxr,
+				  bool close_path)
 {
 	struct bnxt_ring_struct *ring = &rxr->rx_ring_struct;
 	u32 grp_idx = rxr->bnapi->index;
 	u32 cmpl_ring_id;
+	int rc;
 
 	if (ring->fw_ring_id == INVALID_HW_RING_ID)
-		return;
+		return 0;
 
 	cmpl_ring_id = bnxt_cp_ring_for_rx(bp, rxr);
-	hwrm_ring_free_send_msg(bp, ring,
-				RING_FREE_REQ_RING_TYPE_RX,
-				close_path ? cmpl_ring_id :
-				INVALID_HW_RING_ID);
+	rc = hwrm_ring_free_send_msg(bp, ring,
+				     RING_FREE_REQ_RING_TYPE_RX,
+				     close_path ? cmpl_ring_id :
+				     INVALID_HW_RING_ID);
 	ring->fw_ring_id = INVALID_HW_RING_ID;
 	bp->grp_info[grp_idx].rx_fw_ring_id = INVALID_HW_RING_ID;
+	return rc;
 }
 
-static void bnxt_hwrm_rx_agg_ring_free(struct bnxt *bp,
-				       struct bnxt_rx_ring_info *rxr,
-				       bool close_path)
+static int bnxt_hwrm_rx_agg_ring_free(struct bnxt *bp,
+				      struct bnxt_rx_ring_info *rxr,
+				      bool close_path)
 {
 	struct bnxt_ring_struct *ring = &rxr->rx_agg_ring_struct;
 	u32 grp_idx = rxr->bnapi->index;
 	u32 type, cmpl_ring_id;
+	int rc;
 
 	if (bp->flags & BNXT_FLAG_CHIP_P5_PLUS)
 		type = RING_FREE_REQ_RING_TYPE_RX_AGG;
@@ -7711,14 +7716,15 @@ static void bnxt_hwrm_rx_agg_ring_free(struct bnxt *bp,
 		type = RING_FREE_REQ_RING_TYPE_RX;
 
 	if (ring->fw_ring_id == INVALID_HW_RING_ID)
-		return;
+		return 0;
 
 	cmpl_ring_id = bnxt_cp_ring_for_rx(bp, rxr);
-	hwrm_ring_free_send_msg(bp, ring, type,
-				close_path ? cmpl_ring_id :
-				INVALID_HW_RING_ID);
+	rc = hwrm_ring_free_send_msg(bp, ring, type,
+				     close_path ? cmpl_ring_id :
+				     INVALID_HW_RING_ID);
 	ring->fw_ring_id = INVALID_HW_RING_ID;
 	bp->grp_info[grp_idx].agg_fw_ring_id = INVALID_HW_RING_ID;
+	return rc;
 }
 
 static void bnxt_hwrm_cp_ring_free(struct bnxt *bp,
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [RFC net v4 2/4] bnxt_en: check HWRM response if completion never arrives
  2026-09-25 17:43 [RFC net v4 0/4] bnxt_en: Make RING FREE more robust Joe Damato
  2026-09-25 17:43 ` [RFC net v4 1/4] bnxt_en: return the RING_FREE status to callers Joe Damato
@ 2026-09-25 17:43 ` Joe Damato
  2026-09-25 17:44 ` [RFC net v4 3/4] bnxt_en: stop DMA before releasing rings the firmware did not free Joe Damato
  2026-09-25 17:44 ` [RFC net v4 4/4] bnxt_en: refuse to open a device with stopped DMA Joe Damato
  3 siblings, 0 replies; 12+ messages in thread
From: Joe Damato @ 2026-09-25 17:43 UTC (permalink / raw)
  To: netdev, Michael Chan, Pavan Chebbi, Andrew Lunn, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Prashant Sreedharan
  Cc: edumazet, horms, linux-kernel, Joe Damato

When a command is sent over a completion ring, __hwrm_send() waits for
NAPI to consume the completion and gives up if it never arrives, without
looking at the response.

If a completion is not posted within the timeout, check the response
before giving up. If resp_len is set, the sequence id matches, and the
valid byte is set then the firmware completed the command and only the
notification was lost. Fall through to the normal error_code handling in
that case.

Several seconds are spent waiting for the completion, so a response that
was written at all is complete by the time the wait gives up. There is no
need to poll for the valid byte here the way the polling path below has
to, where the poll is for a non-zero length and the valid byte at the end
of the message may still be on its way.

Log the response state on both paths so there is more data when this rare
event occurs.

Fixes: 74608fc98d28 ("bnxt_en: Ring free response from close path should use completion ring")
Signed-off-by: Joe Damato <joe@dama.to>
---
 .../net/ethernet/broadcom/bnxt/bnxt_hwrm.c    | 33 ++++++++++++++++---
 1 file changed, 29 insertions(+), 4 deletions(-)

diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt_hwrm.c b/drivers/net/ethernet/broadcom/bnxt/bnxt_hwrm.c
index 5bfabdca7d0e..4feba90f0bf6 100644
--- a/drivers/net/ethernet/broadcom/bnxt/bnxt_hwrm.c
+++ b/drivers/net/ethernet/broadcom/bnxt/bnxt_hwrm.c
@@ -582,11 +582,36 @@ static int __hwrm_send(struct bnxt *bp, struct bnxt_hwrm_ctx *ctx)
 		}
 
 		if (READ_ONCE(token->state) != BNXT_HWRM_COMPLETE) {
-			hwrm_err(bp, ctx, "Resp cmpl intr err msg: 0x%x\n",
-				 req_type);
-			goto exit;
+			__le16 resp_seq_id;
+			u8 valid_byte = 0;
+
+			/* The completion ring entry was not delivered for
+			 * some reason. It might be possible that the command
+			 * was carried out even without a completion being
+			 * posted. Check the response before giving up and log
+			 * the state.
+			 */
+			dma_rmb();
+			resp_seq_id = READ_ONCE(ctx->resp->seq_id);
+			len = le16_to_cpu(READ_ONCE(ctx->resp->resp_len));
+			if (len && resp_seq_id == ctx->req->seq_id)
+				valid_byte = *((u8 *)ctx->resp + len - 1);
+
+			if (!valid_byte) {
+				hwrm_err(bp, ctx,
+					 "Resp cmpl intr err msg: 0x%x len:%d seq:0x%x/0x%x\n",
+					 req_type, len,
+					 le16_to_cpu(resp_seq_id),
+					 le16_to_cpu(ctx->req->seq_id));
+				goto exit;
+			}
+			netdev_warn(bp->dev,
+				    "Resp cmpl intr not delivered, msg: 0x%x completed anyway (len:%d valid:0x%x err:0x%x)\n",
+				    req_type, len, valid_byte,
+				    le16_to_cpu(ctx->resp->error_code));
+		} else {
+			len = le16_to_cpu(READ_ONCE(ctx->resp->resp_len));
 		}
-		len = le16_to_cpu(READ_ONCE(ctx->resp->resp_len));
 		valid = ((u8 *)ctx->resp) + len - 1;
 	} else {
 		__le16 seen_out_of_seq = ctx->req->seq_id; /* will never see */
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [RFC net v4 3/4] bnxt_en: stop DMA before releasing rings the firmware did not free
  2026-09-25 17:43 [RFC net v4 0/4] bnxt_en: Make RING FREE more robust Joe Damato
  2026-09-25 17:43 ` [RFC net v4 1/4] bnxt_en: return the RING_FREE status to callers Joe Damato
  2026-09-25 17:43 ` [RFC net v4 2/4] bnxt_en: check HWRM response if completion never arrives Joe Damato
@ 2026-09-25 17:44 ` Joe Damato
  2026-09-25 17:44 ` [RFC net v4 4/4] bnxt_en: refuse to open a device with stopped DMA Joe Damato
  3 siblings, 0 replies; 12+ messages in thread
From: Joe Damato @ 2026-09-25 17:44 UTC (permalink / raw)
  To: netdev, Michael Chan, Pavan Chebbi, Andrew Lunn, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Prashant Sreedharan
  Cc: edumazet, horms, linux-kernel, Joe Damato

When HWRM_RING_FREE is not answered, bnxt_hwrm_ring_free() clears
fw_ring_id and __bnxt_close_nic() goes on to call bnxt_free_mem(), which
unmaps the ring memory and the RX buffers that the FW may still be
using.

This is reachable in production. On a BCM57504 the first sign is the TX
watchdog; the close that follows times out a subset of its RING_FREEs and
the driver releases those rings anyway:

  05:30:12  NETDEV WATCHDOG: transmit queue 0 timed out 6073 ms
  05:30:12  Resp cmpl intr err msg: 0x51                  x20
  05:30:12  hwrm_ring_free type 1 failed                  x12
  05:30:12  hwrm_ring_free type 2 failed                  x8
  05:30:12  AMD-Vi: IO_PAGE_FAULT  x3

Count the rings the firmware did not free and report that to the caller
so it can decide what to do. bnxt_hwrm_resource_free() still frees the
remaining firmware resources before returning the error, so the shutdown
can send every message it needs to before the device is stopped.

The device stays unusable until the driver is rebound.

Fixes: 74608fc98d28 ("bnxt_en: Ring free response from close path should use completion ring")
Signed-off-by: Joe Damato <joe@dama.to>
---
 drivers/net/ethernet/broadcom/bnxt/bnxt.c | 48 +++++++++++++++++------
 1 file changed, 35 insertions(+), 13 deletions(-)

diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
index a7f6facca7b4..33e9ce8eb449 100644
--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
+++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
@@ -7754,21 +7754,25 @@ static void bnxt_clear_one_cp_ring(struct bnxt *bp, struct bnxt_cp_ring_info *cp
 			memset(cpr->cp_desc_ring[i], 0, size);
 }
 
-static void bnxt_hwrm_ring_free(struct bnxt *bp, bool close_path)
+static int bnxt_hwrm_ring_free(struct bnxt *bp, bool close_path)
 {
+	int stuck = 0;
 	u32 type;
 	int i;
 
 	if (!bp->bnapi)
-		return;
+		return 0;
 
 	for (i = 0; i < bp->tx_nr_rings; i++)
-		bnxt_hwrm_tx_ring_free(bp, &bp->tx_ring[i], close_path);
+		if (bnxt_hwrm_tx_ring_free(bp, &bp->tx_ring[i], close_path))
+			stuck++;
 
 	bnxt_cancel_dim(bp);
 	for (i = 0; i < bp->rx_nr_rings; i++) {
-		bnxt_hwrm_rx_ring_free(bp, &bp->rx_ring[i], close_path);
-		bnxt_hwrm_rx_agg_ring_free(bp, &bp->rx_ring[i], close_path);
+		if (bnxt_hwrm_rx_ring_free(bp, &bp->rx_ring[i], close_path))
+			stuck++;
+		if (bnxt_hwrm_rx_agg_ring_free(bp, &bp->rx_ring[i], close_path))
+			stuck++;
 	}
 
 	/* The completion rings are about to be freed.  After that the
@@ -7798,6 +7802,19 @@ static void bnxt_hwrm_ring_free(struct bnxt *bp, bool close_path)
 			bp->grp_info[i].cp_fw_ring_id = INVALID_HW_RING_ID;
 		}
 	}
+
+	if (!stuck)
+		return 0;
+
+	netdev_err(bp->dev, "Firmware did not free %d ring(s)\n", stuck);
+	return -EIO;
+}
+
+static void bnxt_stop_dma(struct bnxt *bp)
+{
+	netdev_err(bp->dev,
+		   "Disabling DMA before releasing ring memory, the driver must be rebound to recover\n");
+	pci_clear_master(bp->pdev);
 }
 
 static int __bnxt_trim_rings(struct bnxt *bp, int *rx, int *tx, int max,
@@ -10832,16 +10849,19 @@ static void bnxt_clear_vnic(struct bnxt *bp)
 		bnxt_hwrm_vnic_ctx_free(bp);
 }
 
-static void bnxt_hwrm_resource_free(struct bnxt *bp, bool close_path,
-				    bool irq_re_init)
+static int bnxt_hwrm_resource_free(struct bnxt *bp, bool close_path,
+				   bool irq_re_init)
 {
+	int rc;
+
 	bnxt_clear_vnic(bp);
-	bnxt_hwrm_ring_free(bp, close_path);
+	rc = bnxt_hwrm_ring_free(bp, close_path);
 	bnxt_hwrm_ring_grp_free(bp);
 	if (irq_re_init) {
 		bnxt_hwrm_stat_ctx_free(bp);
 		bnxt_hwrm_free_tunnel_ports(bp);
 	}
+	return rc;
 }
 
 static int bnxt_hwrm_set_br_mode(struct bnxt *bp, u16 br_mode)
@@ -11363,15 +11383,15 @@ static int bnxt_init_chip(struct bnxt *bp, bool irq_re_init)
 	return 0;
 
 err_out:
-	bnxt_hwrm_resource_free(bp, 0, true);
+	if (bnxt_hwrm_resource_free(bp, 0, true))
+		bnxt_stop_dma(bp);
 
 	return rc;
 }
 
 static int bnxt_shutdown_nic(struct bnxt *bp, bool irq_re_init)
 {
-	bnxt_hwrm_resource_free(bp, 1, irq_re_init);
-	return 0;
+	return bnxt_hwrm_resource_free(bp, 1, irq_re_init);
 }
 
 static int bnxt_init_nic(struct bnxt *bp, bool irq_re_init)
@@ -13422,7 +13442,8 @@ int bnxt_half_open_nic(struct bnxt *bp)
  */
 void bnxt_half_close_nic(struct bnxt *bp)
 {
-	bnxt_hwrm_resource_free(bp, false, true);
+	if (bnxt_hwrm_resource_free(bp, false, true))
+		bnxt_stop_dma(bp);
 	bnxt_del_napi(bp);
 	bnxt_free_skbs(bp);
 	bnxt_free_mem(bp, true);
@@ -13501,7 +13522,8 @@ static void __bnxt_close_nic(struct bnxt *bp, bool irq_re_init,
 	if (BNXT_SUPPORTS_MULTI_RSS_CTX(bp))
 		bnxt_clear_rss_ctxs(bp);
 	/* Flush rings and disable interrupts */
-	bnxt_shutdown_nic(bp, irq_re_init);
+	if (bnxt_shutdown_nic(bp, irq_re_init))
+		bnxt_stop_dma(bp);
 
 	/* TODO CHIMP_FW: Link/PHY related cleanup if (link_re_init) */
 
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [RFC net v4 4/4] bnxt_en: refuse to open a device with stopped DMA
  2026-09-25 17:43 [RFC net v4 0/4] bnxt_en: Make RING FREE more robust Joe Damato
                   ` (2 preceding siblings ...)
  2026-09-25 17:44 ` [RFC net v4 3/4] bnxt_en: stop DMA before releasing rings the firmware did not free Joe Damato
@ 2026-09-25 17:44 ` Joe Damato
  2026-09-29  0:20   ` Joe Damato
  3 siblings, 1 reply; 12+ messages in thread
From: Joe Damato @ 2026-09-25 17:44 UTC (permalink / raw)
  To: netdev, Michael Chan, Pavan Chebbi, Andrew Lunn, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Prashant Sreedharan
  Cc: edumazet, horms, linux-kernel, Joe Damato

Add BNXT_STATE_DMA_STOPPED and set it in bnxt_stop_dma() to signal that
the device had DMA disabled. When opening the device later, check this
bit and exit with an error.

This is a very rare case, but when it happens the device needs manual
intervention by the user to become usable again.

Fixes: 74608fc98d28 ("bnxt_en: Ring free response from close path should use completion ring")
Signed-off-by: Joe Damato <joe@dama.to>
---
 drivers/net/ethernet/broadcom/bnxt/bnxt.c | 26 ++++++++++++++++++++---
 drivers/net/ethernet/broadcom/bnxt/bnxt.h |  1 +
 2 files changed, 24 insertions(+), 3 deletions(-)

diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
index 33e9ce8eb449..21da989e136e 100644
--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
+++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
@@ -7814,9 +7814,20 @@ static void bnxt_stop_dma(struct bnxt *bp)
 {
 	netdev_err(bp->dev,
 		   "Disabling DMA before releasing ring memory, the driver must be rebound to recover\n");
+	set_bit(BNXT_STATE_DMA_STOPPED, &bp->state);
 	pci_clear_master(bp->pdev);
 }
 
+static int bnxt_check_dma_stopped(struct bnxt *bp)
+{
+	if (!test_bit(BNXT_STATE_DMA_STOPPED, &bp->state))
+		return 0;
+
+	netdev_err(bp->dev,
+		   "DMA was disabled after the firmware failed to free rings, rebind the driver to recover\n");
+	return -ENODEV;
+}
+
 static int __bnxt_trim_rings(struct bnxt *bp, int *rx, int *tx, int max,
 			     bool shared);
 static int bnxt_trim_rings(struct bnxt *bp, int *rx, int *tx, int max,
@@ -13387,9 +13398,10 @@ static int __bnxt_open_nic(struct bnxt *bp, bool irq_re_init, bool link_re_init)
 
 int bnxt_open_nic(struct bnxt *bp, bool irq_re_init, bool link_re_init)
 {
-	int rc = 0;
+	int rc;
 
-	if (test_bit(BNXT_STATE_ABORT_ERR, &bp->state))
+	rc = bnxt_check_dma_stopped(bp);
+	if (!rc && test_bit(BNXT_STATE_ABORT_ERR, &bp->state))
 		rc = -EIO;
 	if (!rc)
 		rc = __bnxt_open_nic(bp, irq_re_init, link_re_init);
@@ -13406,7 +13418,11 @@ int bnxt_open_nic(struct bnxt *bp, bool irq_re_init, bool link_re_init)
  */
 int bnxt_half_open_nic(struct bnxt *bp)
 {
-	int rc = 0;
+	int rc;
+
+	rc = bnxt_check_dma_stopped(bp);
+	if (rc)
+		goto half_open_err;
 
 	if (test_bit(BNXT_STATE_ABORT_ERR, &bp->state)) {
 		netdev_err(bp->dev, "A previous firmware reset has not completed, aborting half open\n");
@@ -13466,6 +13482,10 @@ static int bnxt_open(struct net_device *dev)
 	struct bnxt *bp = netdev_priv(dev);
 	int rc;
 
+	rc = bnxt_check_dma_stopped(bp);
+	if (rc)
+		return rc;
+
 	if (test_bit(BNXT_STATE_ABORT_ERR, &bp->state)) {
 		rc = bnxt_reinit_after_abort(bp);
 		if (rc) {
diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.h b/drivers/net/ethernet/broadcom/bnxt/bnxt.h
index c673b2ce4a0d..061f57824c7a 100644
--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.h
+++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.h
@@ -2470,6 +2470,7 @@ struct bnxt {
 #define BNXT_STATE_DRV_REGISTERED	7
 #define BNXT_STATE_PCI_CHANNEL_IO_FROZEN	8
 #define BNXT_STATE_NAPI_DISABLED	9
+#define BNXT_STATE_DMA_STOPPED		10
 #define BNXT_STATE_FW_ACTIVATE		11
 #define BNXT_STATE_RECOVER		12
 #define BNXT_STATE_FW_NON_FATAL_COND	13
-- 
2.53.0-Meta


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* Re: [RFC net v4 4/4] bnxt_en: refuse to open a device with stopped DMA
  2026-09-25 17:44 ` [RFC net v4 4/4] bnxt_en: refuse to open a device with stopped DMA Joe Damato
@ 2026-09-29  0:20   ` Joe Damato
  0 siblings, 0 replies; 12+ messages in thread
From: Joe Damato @ 2026-09-29  0:20 UTC (permalink / raw)
  To: netdev, Michael Chan, Pavan Chebbi, Andrew Lunn, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Prashant Sreedharan
  Cc: edumazet, horms, linux-kernel

On Fri, Sep 25, 2026 at 10:44:01AM -0700, Joe Damato wrote:
> Add BNXT_STATE_DMA_STOPPED and set it in bnxt_stop_dma() to signal that
> the device had DMA disabled. When opening the device later, check this
> bit and exit with an error.
> 
> This is a very rare case, but when it happens the device needs manual
> intervention by the user to become usable again.
> 
> Fixes: 74608fc98d28 ("bnxt_en: Ring free response from close path should use completion ring")
> Signed-off-by: Joe Damato <joe@dama.to>
> ---
>  drivers/net/ethernet/broadcom/bnxt/bnxt.c | 26 ++++++++++++++++++++---
>  drivers/net/ethernet/broadcom/bnxt/bnxt.h |  1 +
>  2 files changed, 24 insertions(+), 3 deletions(-)
 
[...]

> diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.h b/drivers/net/ethernet/broadcom/bnxt/bnxt.h
> index c673b2ce4a0d..061f57824c7a 100644
> --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.h
> +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.h
> @@ -2470,6 +2470,7 @@ struct bnxt {
>  #define BNXT_STATE_DRV_REGISTERED	7
>  #define BNXT_STATE_PCI_CHANNEL_IO_FROZEN	8
>  #define BNXT_STATE_NAPI_DISABLED	9
> +#define BNXT_STATE_DMA_STOPPED		10
>  #define BNXT_STATE_FW_ACTIVATE		11

FWIW: while applying this series to an older kernel to test it on some prod
hardware, I noted that this bit was previously taken.

If this is re-submit as a real series and not an RFC, I suppose I'll tweak
this to take a different bit to make backporting easier.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
@ 2026-10-06 16:43 anmory
  2026-10-06 20:00 ` Joe Damato
  2026-10-06 20:22 ` Joe Damato
  0 siblings, 2 replies; 12+ messages in thread
From: anmory @ 2026-10-06 16:43 UTC (permalink / raw)
  To: joe; +Cc: netdev, michael.chan, pavan.chebbi, linux-kernel


[-- Attachment #1.1: Type: text/plain, Size: 2671 bytes --]

Hi Joe,

we appear to be hitting a very similar issue to the one described in
this RFC on two Broadcom BCM57504 systems.

One detail that may be particularly relevant is the ordering in our
reproductions: on both systems the AMD-Vi IO_PAGE_FAULT occurs before
the NETDEV WATCHDOG / TX timeout.

We originally encountered the problem after updating from Debian
6.12.107-1 to 6.12.111-1.

We subsequently reproduced it deliberately on two separate systems.

Test results:

affected-host-1:
BCM57504 [14e4:1751], bnxt_en
BIOS: 1.16.2
NIC FW: 36.11.55.00
FW mgmt: 236.1.153.0

6.12.107-1: stable
6.12.111-1: failure reproduced

affected-host-2:
BCM57504 [14e4:1751], bnxt_en
BIOS: 1.18.2
NIC FW: 36.11.73.00
FW mgmt: 236.1.173.0

6.12.107-1: stable
6.12.111-1: failure reproduced

The failure sequence on affected-host-1 was:

06:49:00 AMD-Vi IO_PAGE_FAULT
06:49:14 NETDEV WATCHDOG: transmit queue 0 timed out
06:49:14 TX timeout detected, starting reset task
06:49:18 HWRM/RING_FREE failures begin
06:49:25 HWRM_RING_ALLOC fails
06:49:25 bnxt_init_nic fails
06:49:25 nic open fails

On affected-host-2:

15:29:09 AMD-Vi IO_PAGE_FAULT
15:29:15 NETDEV WATCHDOG: transmit queue 4 timed out
15:29:15 TX timeout detected, starting reset task
15:29:19 HWRM/RING_FREE failures begin
15:29:37 HWRM_RING_ALLOC fails
15:29:37 bnxt_init_nic fails
15:29:37 nic open fails

So in our reproductions the IOMMU fault precedes the TX watchdog by
approximately 6-14 seconds.

We have also observed the timeout on different TX queues across
different occurrences, so it does not appear to be tied to a specific
queue.

Both interfaces were up and operating at 25 Gbit/s before the failure.
The systems use the IOMMU in translated mode.

After the reset attempt fails, the interface cannot be reopened and a
reboot is required.

The same problem therefore reproduces across:
- two physical systems
- two BIOS revisions
- two NIC firmware revisions
- different TX queues

while reverting to 6.12.107-1 has been stable with the same workload.

I've attached sanitized diagnostic output from both reproductions,
including PCI/device information, firmware versions, IOMMU
configuration and the complete failure event sequence.

We have not bisected the changes between 6.12.107 and 6.12.111 yet.

Given that the IO_PAGE_FAULT precedes the watchdog in our case, I
wonder whether this may provide some information about the initial
failure, in addition to the RING_FREE recovery problem addressed by
your RFC.

Please let me know if there are additional diagnostics that would be
useful, or if you would like us to test the RFC series on one of these
systems.

Thanks,
R.

https://proton.me/mail/home

[-- Attachment #1.2: Type: text/html, Size: 3865 bytes --]

[-- Attachment #2: affected-host-1-bnxt.txt --]
[-- Type: text/plain, Size: 4027 bytes --]

=== SYSTEM ===
hostname: affected-host-1
bios_version: 1.16.2
bios_date: 01/16/2026
eth1_pci_device: 0000:c4:00.2

=== NIC / DRIVER ===
c4:00.2 Ethernet controller [0200]: Broadcom Inc. and subsidiaries BCM57504 NetXtreme-E 10Gb/25Gb/40Gb/50Gb/100Gb Ethernet [14e4:1751] (rev 12)
	DeviceName: Integrated NIC 1 Port 3-1
	Subsystem: Broadcom Inc. and subsidiaries NetXtreme-E BCM57504 4x25G OCP3.0 [14e4:5045]
	Kernel driver in use: bnxt_en
	Kernel modules: bnxt_en

=== DEVLINK ===
pci/0000:c4:00.2:
  driver bnxt_en
  serial_number <redacted>
  versions:
      fixed:
        board.id BCM957504
        asic.id 1751
        asic.rev B2
      running:
        fw.psid 232.0.2
        fw 36.11.55.00
        fw.mgmt 236.1.153.0
        fw.mgmt.api 1.10.3
      stored:
        fw.psid 232.0.2
        fw 36.11.55.00
        fw.mgmt 236.1.153.0

=== CAPTURE BASELINE ===
2026-10-06T06:28:26+00:00
affected-host-1
6.12.111+deb13-amd64
1.16.2
01/16/2026
eth1             UP             XX:XX:XX:XX:XX:XX <BROADCAST,MULTICAST,UP,LOWER_UP> 

=== RELEVANT BOOT / HARDWARE ===
2026-10-06T06:28:25+00:00 affected-host-1 kernel: Linux version 6.12.111+deb13-amd64 (debian-kernel@lists.debian.org) (x86_64-linux-gnu-gcc-14 (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44) #1 SMP PREEMPT_DYNAMIC Debian 6.12.111-1 (2026-09-28)
2026-10-06T06:28:25+00:00 affected-host-1 kernel: DMI: Dell Inc. PowerEdge R7615, BIOS 1.16.2 01/16/2026
2026-10-06T06:28:25+00:00 affected-host-1 kernel: pci 0000:c4:00.2: [14e4:1751] type 00 class 0x020000 PCIe Endpoint
2026-10-06T06:28:25+00:00 affected-host-1 kernel: iommu: Default domain type: Translated
2026-10-06T06:28:25+00:00 affected-host-1 kernel: iommu: DMA domain TLB invalidation policy: lazy mode
2026-10-06T06:28:25+00:00 affected-host-1 kernel: pci 0000:c4:00.2: Adding to iommu group 14
2026-10-06T06:28:25+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth4: Broadcom BCM57504 NetXtreme-E 10Gb/25Gb/50Gb/100Gb/200Gb Ethernet found at mem cb010000, node addr XX:XX:XX:XX:XX:XX
2026-10-06T06:28:26+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: NIC Link is Up, 25000 Mbps (NRZ) full duplex, Flow control: ON - receive & transmit
2026-10-06T06:28:26+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: FEC autoneg off encoding: Clause 74 BaseR

=== FAILURE EVENTS ===
2026-10-06T06:49:00+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2: AMD-Vi: Event logged [IO_PAGE_FAULT domain=0x000f address=0xfb387000 flags=0x0000]
2026-10-06T06:49:14+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: NETDEV WATCHDOG: CPU: 2: transmit queue 0 timed out 5312 ms
2026-10-06T06:49:14+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: TX timeout detected, starting reset task!
2026-10-06T06:49:18+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T06:49:18+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: hwrm_ring_free type 1 failed. rc:fffffff0 err:2
2026-10-06T06:49:21+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T06:49:21+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: hwrm_ring_free type 2 failed. rc:fffffff0 err:0
2026-10-06T06:49:25+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T06:49:25+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: hwrm_ring_free type 4 failed. rc:fffffff0 err:0
2026-10-06T06:49:25+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: hwrm_ring_free type 0 failed. rc:fffffffb err:1
2026-10-06T06:49:25+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: hwrm_ring_free type 0 failed. rc:fffffffb err:1
2026-10-06T06:49:25+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: hwrm_ring_alloc type 1 failed. rc:ffffffe4 err:4
2026-10-06T06:49:25+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: bnxt_init_nic err: fffffffb
2026-10-06T06:49:25+00:00 affected-host-1 kernel: bnxt_en 0000:c4:00.2 eth1: nic open fail (rc: fffffffb)

[-- Attachment #3: affected-host-2-bnxt.txt --]
[-- Type: text/plain, Size: 5471 bytes --]

=== SYSTEM ===
hostname: affected-host-2
bios_version: 1.18.2
bios_date: 07/22/2026
eth1_pci_device: 0000:c4:00.3

=== NIC / DRIVER ===
c4:00.3 Ethernet controller [0200]: Broadcom Inc. and subsidiaries BCM57504 NetXtreme-E 10Gb/25Gb/40Gb/50Gb/100Gb Ethernet [14e4:1751] (rev 12)
	DeviceName: Integrated NIC 1 Port 4-1
	Subsystem: Broadcom Inc. and subsidiaries NetXtreme-E BCM57504 4x25G OCP3.0 [14e4:5045]
	Kernel driver in use: bnxt_en
	Kernel modules: bnxt_en

=== DEVLINK ===
pci/0000:c4:00.3:
  driver bnxt_en
  serial_number <redacted>
  versions:
      fixed:
        board.id BCM957504
        asic.id 1751
        asic.rev B2
      running:
        fw.psid 232.0.2
        fw 36.11.73.00
        fw.mgmt 236.1.173.0
        fw.mgmt.api 1.10.3
      stored:
        fw.psid 232.0.2
        fw 36.11.73.00
        fw.mgmt 236.1.173.0

=== CAPTURE BASELINE ===
2026-10-06T15:29:09+00:00
affected-host-2
6.12.111+deb13-amd64
1.18.2
07/22/2026
eth1             UP             XX:XX:XX:XX:XX:XX <BROADCAST,MULTICAST,UP,LOWER_UP> 

=== RELEVANT BOOT / HARDWARE ===
2026-10-06T15:29:07+00:00 affected-host-2 kernel: Linux version 6.12.111+deb13-amd64 (debian-kernel@lists.debian.org) (x86_64-linux-gnu-gcc-14 (Debian 14.2.0-19) 14.2.0, GNU ld (GNU Binutils for Debian) 2.44) #1 SMP PREEMPT_DYNAMIC Debian 6.12.111-1 (2026-09-28)
2026-10-06T15:29:07+00:00 affected-host-2 kernel: DMI: Dell Inc. PowerEdge R7615, BIOS 1.18.2 07/22/2026
2026-10-06T15:29:07+00:00 affected-host-2 kernel: pci 0000:c4:00.3: [14e4:1751] type 00 class 0x020000 PCIe Endpoint
2026-10-06T15:29:07+00:00 affected-host-2 kernel: iommu: Default domain type: Translated
2026-10-06T15:29:07+00:00 affected-host-2 kernel: iommu: DMA domain TLB invalidation policy: lazy mode
2026-10-06T15:29:07+00:00 affected-host-2 kernel: pci 0000:c4:00.3: Adding to iommu group 15
2026-10-06T15:29:07+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth5: Broadcom BCM57504 NetXtreme-E 10Gb/25Gb/50Gb/100Gb/200Gb Ethernet found at mem cb000000, node addr XX:XX:XX:XX:XX:XX
2026-10-06T15:29:09+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: NIC Link is Up, 25000 Mbps (NRZ) full duplex, Flow control: ON - receive & transmit
2026-10-06T15:29:09+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: FEC autoneg off encoding: Clause 74 BaseR

=== FAILURE EVENTS ===
2026-10-06T15:29:09+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3: AMD-Vi: Event logged [IO_PAGE_FAULT domain=0x0010 address=0xfb526000 flags=0x0000]
2026-10-06T15:29:15+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: NETDEV WATCHDOG: CPU: 23: transmit queue 4 timed out 5628 ms
2026-10-06T15:29:15+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: TX timeout detected, starting reset task!
2026-10-06T15:29:19+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T15:29:19+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 1 failed. rc:fffffff0 err:0
2026-10-06T15:29:19+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 1 failed. rc:ffffffea err:2
2026-10-06T15:29:19+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 1 failed. rc:ffffffea err:2
2026-10-06T15:29:22+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T15:29:22+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 1 failed. rc:fffffff0 err:2
2026-10-06T15:29:22+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 1 failed. rc:ffffffea err:2
2026-10-06T15:29:26+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T15:29:26+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 2 failed. rc:fffffff0 err:0
2026-10-06T15:29:30+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T15:29:30+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 4 failed. rc:fffffff0 err:0
2026-10-06T15:29:33+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T15:29:33+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 2 failed. rc:fffffff0 err:0
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: Resp cmpl intr err msg: 0x51
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 4 failed. rc:fffffff0 err:0
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 0 failed. rc:fffffffb err:1
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 0 failed. rc:fffffffb err:1
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 0 failed. rc:fffffffb err:1
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 0 failed. rc:fffffffb err:1
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_free type 0 failed. rc:fffffffb err:1
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: hwrm_ring_alloc type 1 failed. rc:ffffffe4 err:4
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: bnxt_init_nic err: fffffffb
2026-10-06T15:29:37+00:00 affected-host-2 kernel: bnxt_en 0000:c4:00.3 eth1: nic open fail (rc: fffffffb)

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
  2026-10-06 16:43 [RFC net v4 0/4] bnxt_en: Make RING FREE more robust anmory
@ 2026-10-06 20:00 ` Joe Damato
  2026-10-06 20:22 ` Joe Damato
  1 sibling, 0 replies; 12+ messages in thread
From: Joe Damato @ 2026-10-06 20:00 UTC (permalink / raw)
  To: anmory; +Cc: netdev, michael.chan, pavan.chebbi, linux-kernel

On Tue, Oct 06, 2026 at 04:43:42PM +0000, anmory wrote:
> Hi Joe,
> 
> we appear to be hitting a very similar issue to the one described in
> this RFC on two Broadcom BCM57504 systems.
> 
> One detail that may be particularly relevant is the ordering in our
> reproductions: on both systems the AMD-Vi IO_PAGE_FAULT occurs before
> the NETDEV WATCHDOG / TX timeout.
> 
> We originally encountered the problem after updating from Debian
> 6.12.107-1 to 6.12.111-1.
> 
> We subsequently reproduced it deliberately on two separate systems.

Can you provide a sample program or explanation on how *exactly* to reproduce
this issue?

I've hit the same one you are describing (IO_PAGE_FAULT before the NETDEV TX
timeout) but have not been able to reliably reproduce it. It happens fairly
rarely on my systems, but I am very eager to debug this and fix it.

If you have a way to reliably reproduce this, that would make it *much* easier
to debug.

Thanks,
Joe

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
  2026-10-06 16:43 [RFC net v4 0/4] bnxt_en: Make RING FREE more robust anmory
  2026-10-06 20:00 ` Joe Damato
@ 2026-10-06 20:22 ` Joe Damato
  2026-10-06 22:22   ` anmory
  1 sibling, 1 reply; 12+ messages in thread
From: Joe Damato @ 2026-10-06 20:22 UTC (permalink / raw)
  To: anmory; +Cc: netdev, michael.chan, pavan.chebbi, linux-kernel

On Tue, Oct 06, 2026 at 04:43:42PM +0000, anmory wrote:
> Hi Joe,
> 
> we appear to be hitting a very similar issue to the one described in
> this RFC on two Broadcom BCM57504 systems.
> 
> One detail that may be particularly relevant is the ordering in our
> reproductions: on both systems the AMD-Vi IO_PAGE_FAULT occurs before
> the NETDEV WATCHDOG / TX timeout.
> 
> We originally encountered the problem after updating from Debian
> 6.12.107-1 to 6.12.111-1.

BTW there was recently this patch [1] proposed by Eric which helps prevent
a bug Stefan reported [2].

I am not sure whether the bug you are hitting is Eric's bug or the bug I am
chasing.

Other than explaining how you repro'd the issue, if you want to test Eric's
patch you could apply his patch to v6.12.111 and see if the issue still
occurs.

If it stops, it's probably the bug Eric fixed and not the one I'm chasing :)
but if it continues, it could be what I've been looking for and I'd really
want to know how you reproduced it, please.

[1]: https://lore.kernel.org/netdev/20261006042153.199444-1-edumazet@kernel.org/
[2]: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
  2026-10-06 20:22 ` Joe Damato
@ 2026-10-06 22:22   ` anmory
  2026-10-06 22:37     ` Joe Damato
  0 siblings, 1 reply; 12+ messages in thread
From: anmory @ 2026-10-06 22:22 UTC (permalink / raw)
  To: Joe Damato; +Cc: netdev, michael.chan, pavan.chebbi, linux-kernel


Hi Joe,
one important detail I omitted from my previous mail: the affected eth1 interface is an 802.1Q trunk with multiple VLAN subinterfaces.
We also use Keepalived VRRP with use_vmac, so virtual MAC interfaces are layered on the VLAN interfaces. Traffic is routed/forwarded through the host and uses conntrack/SNAT before leaving through another physical interface.
We are not deliberately toggling VLAN offload during the test. The reproduction so far is simply to boot Debian 6.12.111-1 with this normal network configuration and workload. On one host it reproduced after about 20 minutes; on another it reproduced almost immediately. We have also seen a previous occurrence after several hours, so we do not yet have a deterministic packet-level trigger.
Given the VLAN configuration, Eric's patch looks especially relevant. I'll test 6.12.111 with that patch first as you suggested and report whether the failure still reproduces.
Thanks, R.

-------- Original Message --------
On Tuesday, 10/06/26 at 22:22 Joe Damato <joe@dama.to> wrote:
On Tue, Oct 06, 2026 at 04:43:42PM +0000, anmory wrote:
> Hi Joe,
>
> we appear to be hitting a very similar issue to the one described in
> this RFC on two Broadcom BCM57504 systems.
>
> One detail that may be particularly relevant is the ordering in our
> reproductions: on both systems the AMD-Vi IO_PAGE_FAULT occurs before
> the NETDEV WATCHDOG / TX timeout.
>
> We originally encountered the problem after updating from Debian
> 6.12.107-1 to 6.12.111-1.

BTW there was recently this patch [1] proposed by Eric which helps prevent
a bug Stefan reported [2].

I am not sure whether the bug you are hitting is Eric's bug or the bug I am
chasing.

Other than explaining how you repro'd the issue, if you want to test Eric's
patch you could apply his patch to v6.12.111 and see if the issue still
occurs.

If it stops, it's probably the bug Eric fixed and not the one I'm chasing :)
but if it continues, it could be what I've been looking for and I'd really
want to know how you reproduced it, please.

[1]: https://lore.kernel.org/netdev/20261006042153.199444-1-edumazet@kernel.org/
[2]: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/


^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
  2026-10-06 22:22   ` anmory
@ 2026-10-06 22:37     ` Joe Damato
  2026-10-09  4:24       ` anmory
  0 siblings, 1 reply; 12+ messages in thread
From: Joe Damato @ 2026-10-06 22:37 UTC (permalink / raw)
  To: anmory; +Cc: netdev, michael.chan, pavan.chebbi, linux-kernel

On Tue, Oct 06, 2026 at 10:22:51PM +0000, anmory wrote:
> 
> Hi Joe,
> one important detail I omitted from my previous mail: the affected eth1 interface is an 802.1Q trunk with multiple VLAN subinterfaces.
> We also use Keepalived VRRP with use_vmac, so virtual MAC interfaces are layered on the VLAN interfaces. Traffic is routed/forwarded through the host and uses conntrack/SNAT before leaving through another physical interface.
> We are not deliberately toggling VLAN offload during the test. The reproduction so far is simply to boot Debian 6.12.111-1 with this normal network configuration and workload. On one host it reproduced after about 20 minutes; on another it reproduced almost immediately. We have also seen a previous occurrence after several hours, so we do not yet have a deterministic packet-level trigger.
> Given the VLAN configuration, Eric's patch looks especially relevant. I'll test 6.12.111 with that patch first as you suggested and report whether the failure still reproduces.

Seems likely you are hitting the bug Eric just fixed, IMHO. If his patch
resolves your issue, please let me know.

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [RFC net v4 0/4] bnxt_en: Make RING FREE more robust
  2026-10-06 22:37     ` Joe Damato
@ 2026-10-09  4:24       ` anmory
  0 siblings, 0 replies; 12+ messages in thread
From: anmory @ 2026-10-09  4:24 UTC (permalink / raw)
  To: Joe Damato; +Cc: netdev, michael.chan, pavan.chebbi, linux-kernel, edumazet


Hi Joe, Eric, Michael, Pavan,
quick update from my testing.
I backported Eric's patch bnxt_en: fix DMA mapping length for padded small packets onto the exact Debian 6.12.111-1 source and built a Debian test kernel from it.
The affected host has now been running the patched kernel under the same normal workload and network configuration that reproduced the problem with stock 6.12.111-1.
The relevant setup is:
Broadcom BCM57504 / bnxt_en
affected interface is an 802.1Q trunk with multiple VLAN subinterfaces
Keepalived VRRP with use_vmac
forwarding/conntrack/SNAT traffic through the host
25 Gbit/s link
So far the patched kernel has been stable. It ran for the entire day yesterday and is still running without reproducing the failure.
With stock 6.12.111-1 we had repeatedly seen the following sequence:
AMD-Vi IO_PAGE_FAULT
NETDEV WATCHDOG / TX timeout
hwrm_ring_free / hwrm_ring_alloc failures
bnxt_init_nic failure
On the patched kernel, the capture has so far only recorded the normal bnxt_en initialization and link-up messages. No IOMMU fault, TX timeout, or HWRM recovery failure has occurred.
For comparison:
6.12.107-1                     stable
6.12.111-1 stock               reproduces the issue
6.12.111-1 + Eric's patch      stable so far
On this particular host, stock 6.12.111-1 reproduced once after roughly 20 minutes, although we have also seen occurrences take several hours. Given that the patched kernel has now survived the full day yesterday under normal traffic, this is looking increasingly consistent with Eric's DMA padding fix addressing the issue.
I will keep the patched kernel running and continue monitoring it. If the problem does reproduce, I also have the kernel capture running and will collect the NIC coredump as requested.
Thanks, R.



-------- Original Message --------
On Wednesday, 10/07/26 at 00:37 Joe Damato <joe@dama.to> wrote:
On Tue, Oct 06, 2026 at 10:22:51PM +0000, anmory wrote:
>
> Hi Joe,
> one important detail I omitted from my previous mail: the affected eth1 interface is an 802.1Q trunk with multiple VLAN subinterfaces.
> We also use Keepalived VRRP with use_vmac, so virtual MAC interfaces are layered on the VLAN interfaces. Traffic is routed/forwarded through the host and uses conntrack/SNAT before leaving through another physical interface.
> We are not deliberately toggling VLAN offload during the test. The reproduction so far is simply to boot Debian 6.12.111-1 with this normal network configuration and workload. On one host it reproduced after about 20 minutes; on another it reproduced almost immediately. We have also seen a previous occurrence after several hours, so we do not yet have a deterministic packet-level trigger.
> Given the VLAN configuration, Eric's patch looks especially relevant. I'll test 6.12.111 with that patch first as you suggested and report whether the failure still reproduces.

Seems likely you are hitting the bug Eric just fixed, IMHO. If his patch
resolves your issue, please let me know.


^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-10-09  4:25 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-25 17:43 [RFC net v4 0/4] bnxt_en: Make RING FREE more robust Joe Damato
2026-09-25 17:43 ` [RFC net v4 1/4] bnxt_en: return the RING_FREE status to callers Joe Damato
2026-09-25 17:43 ` [RFC net v4 2/4] bnxt_en: check HWRM response if completion never arrives Joe Damato
2026-09-25 17:44 ` [RFC net v4 3/4] bnxt_en: stop DMA before releasing rings the firmware did not free Joe Damato
2026-09-25 17:44 ` [RFC net v4 4/4] bnxt_en: refuse to open a device with stopped DMA Joe Damato
2026-09-29  0:20   ` Joe Damato
  -- strict thread matches above, loose matches on Subject: below --
2026-10-06 16:43 [RFC net v4 0/4] bnxt_en: Make RING FREE more robust anmory
2026-10-06 20:00 ` Joe Damato
2026-10-06 20:22 ` Joe Damato
2026-10-06 22:22   ` anmory
2026-10-06 22:37     ` Joe Damato
2026-10-09  4:24       ` anmory

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox