Netdev List
 help / color / mirror / Atom feed
From: Bernhard Schmidt <berni@debian.org>
To: Salvatore Bonaccorso <carnil@debian.org>
Cc: Eric Dumazet <edumazet@kernel.org>,
	"David S . Miller" <davem@davemloft.net>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>,
	netdev@vger.kernel.org, stable@vger.kernel.org,
	Stefan Fleischmann <sfle@kth.se>,
	Eric Dumazet <edumazet@google.com>,
	Michael Chan <michael.chan@broadcom.com>,
	Pavan Chebbi <pavan.chebbi@broadcom.com>,
	Andrew Lunn <andrew+netdev@lunn.ch>
Subject: Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
Date: Mon, 5 Oct 2026 16:26:25 +0200	[thread overview]
Message-ID: <asOzkZIrKusyF1d0@fliwatuet.svr02.mucip.net> (raw)
In-Reply-To: <asOHdc6HMSVrS9z9@eldamar.lan>

On 05/10/26 01:18 PM, Salvatore Bonaccorso wrote:
> Hi,
> 
> On Mon, Oct 05, 2026 at 04:38:12AM +0200, Eric Dumazet wrote:
> > Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412
> > NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces:
> > 
> >   DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000
> >         [fault reason 0x06] PTE Read access is not set
> >   bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
> >   ...
> >   NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out
> > 
> > The fault address (0xfc499000) is on an exact 4KB page boundary,
> > pointing to a DMA read buffer overrun.
> > 
> > In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes),
> > such as 42-byte untagged ARP frames, are padded:
> > 
> >     if (length < BNXT_MIN_PKT_SIZE) {
> >         pad = BNXT_MIN_PKT_SIZE - length;
> >         if (skb_pad(skb, pad))
> >             goto tx_kick_pending;
> >         length = BNXT_MIN_PKT_SIZE;
> >     }
> > 
> >     mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
> >     ...
> >     dma_unmap_len_set(tx_buf, len, len);
> > 
> > However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes)
> > and is left unadjusted after padding. Consequently, dma_map_single() and
> > dma_unmap_len_set() map and track only 42 bytes.
> > 
> > Later, the hardware TX buffer descriptor is programmed with the padded length:
> > 
> >     txbd->tx_bd_len_flags_type =
> >         cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
> >                     TX_BD_FLAGS_PACKET_END);
> > 
> > The NIC DMA engine is thus instructed to read 52 bytes from a region where
> > only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB
> > page (within 'pad' bytes of the next page), the hardware DMA read overruns
> > into the unmapped adjacent page, triggering an IOMMU fault.
> > 
> > This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
> > and races when toggling HW VLAN offload") because reserving extra VLAN
> > headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
> > offsets and potentially causing small frames to land right against page
> > boundaries.
> > 
> > Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx
> > path. This ensures skb->len and skb_headlen(skb) reflect the padded size so
> > that dma_map_single() maps the full buffer and the descriptor length is
> > consistent. This also removes the temporary 'pad' variable and masking logic.
> > 
> > Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
> > Cc: stable@vger.kernel.org
> > Reported-by: Stefan Fleischmann <sfle@kth.se>
> > Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
> > Signed-off-by: Eric Dumazet <edumazet@google.com>
> > ---
> > Cc: Michael Chan <michael.chan@broadcom.com>
> > Cc: Pavan Chebbi <pavan.chebbi@broadcom.com>
> > Cc: Andrew Lunn <andrew+netdev@lunn.ch>
> > ---
> >  drivers/net/ethernet/broadcom/bnxt/bnxt.c | 19 +++++++------------
> >  1 file changed, 7 insertions(+), 12 deletions(-)
> > 
> > diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> > index d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..7ea27e81e88c5ca82a453b449982b791e8acc831 100644
> > --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> > +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> > @@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> >  	struct netdev_queue *txq;
> >  	int i;
> >  	dma_addr_t mapping;
> > -	unsigned int length, pad = 0;
> > +	unsigned int length;
> >  	u32 len, free_size, vlan_tag_flags, cfa_action, flags;
> >  	struct bnxt_ptp_cfg *ptp = bp->ptp_cfg;
> >  	struct pci_dev *pdev = bp->pdev;
> > @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> >  	}
> >  
> >  normal_tx:
> > -	if (length < BNXT_MIN_PKT_SIZE) {
> > -		pad = BNXT_MIN_PKT_SIZE - length;
> > -		if (skb_pad(skb, pad))
> > -			/* SKB already freed. */
> > -			goto tx_kick_pending;
> > -		length = BNXT_MIN_PKT_SIZE;
> > +	if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
> > +		/* SKB already freed. */
> > +		goto tx_kick_pending;
> >  	}
> > -
> > +	length = skb->len;
> > +	len = skb_headlen(skb);
> >  	mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
> >  
> >  	if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
> > @@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> >  		txbd->tx_bd_len_flags_type = cpu_to_le32(flags);
> >  	}
> >  
> > -	flags &= ~TX_BD_LEN;
> > -	txbd->tx_bd_len_flags_type =
> > -		cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
> > -			    TX_BD_FLAGS_PACKET_END);
> > +	txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
> >  
> >  	netdev_tx_sent_queue(txq, skb->len);
> >  
> > -- 
> > 2.53.0
> 
> FWIW, got as well reported in Debian for an update in the 6.12.y
> series: https://bugs.debian.org/1149564 , in case you would like to
> add a further Link/Closes reference. Bernhard Schmidt is testing the
> patch as well on top of 6.12.111 (what we have right now in Debian)
> and looks promissing: https://bugs.debian.org/1149564#89 .
> 
> Berhard, want to report back a Tested-by from you?

Looks good. Machine has been stable with that patch applied to 6.12

Tested-by: Bernhard Schmidt <berni@debian.org>

Bernhard

  reply	other threads:[~2026-10-05 14:26 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05  2:38 [PATCH net] bnxt_en: fix DMA mapping length for padded small packets Eric Dumazet
2026-10-05  4:14 ` Michael Chan
2026-10-05  9:49   ` Stefan Fleischmann
2026-10-05 11:18 ` Salvatore Bonaccorso
2026-10-05 14:26   ` Bernhard Schmidt [this message]
2026-10-07  6:51     ` Fabian Grünbichler
2026-10-05 21:16 ` netdev-bot+sashiko
2026-10-06  3:56   ` Eric Dumazet
2026-10-06  3:59     ` Eric Dumazet

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=asOzkZIrKusyF1d0@fliwatuet.svr02.mucip.net \
    --to=berni@debian.org \
    --cc=andrew+netdev@lunn.ch \
    --cc=carnil@debian.org \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=edumazet@kernel.org \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=michael.chan@broadcom.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=pavan.chebbi@broadcom.com \
    --cc=sfle@kth.se \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox