* [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
@ 2026-10-05 2:38 Eric Dumazet
2026-10-05 4:14 ` Michael Chan
` (2 more replies)
0 siblings, 3 replies; 9+ messages in thread
From: Eric Dumazet @ 2026-10-05 2:38 UTC (permalink / raw)
To: David S . Miller, Jakub Kicinski, Paolo Abeni
Cc: Simon Horman, netdev, Eric Dumazet, stable, Stefan Fleischmann,
Eric Dumazet, Michael Chan, Pavan Chebbi, Andrew Lunn
Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412
NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces:
DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000
[fault reason 0x06] PTE Read access is not set
bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
...
NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out
The fault address (0xfc499000) is on an exact 4KB page boundary,
pointing to a DMA read buffer overrun.
In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes),
such as 42-byte untagged ARP frames, are padded:
if (length < BNXT_MIN_PKT_SIZE) {
pad = BNXT_MIN_PKT_SIZE - length;
if (skb_pad(skb, pad))
goto tx_kick_pending;
length = BNXT_MIN_PKT_SIZE;
}
mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
...
dma_unmap_len_set(tx_buf, len, len);
However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes)
and is left unadjusted after padding. Consequently, dma_map_single() and
dma_unmap_len_set() map and track only 42 bytes.
Later, the hardware TX buffer descriptor is programmed with the padded length:
txbd->tx_bd_len_flags_type =
cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
TX_BD_FLAGS_PACKET_END);
The NIC DMA engine is thus instructed to read 52 bytes from a region where
only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB
page (within 'pad' bytes of the next page), the hardware DMA read overruns
into the unmapped adjacent page, triggering an IOMMU fault.
This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
and races when toggling HW VLAN offload") because reserving extra VLAN
headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
offsets and potentially causing small frames to land right against page
boundaries.
Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx
path. This ensures skb->len and skb_headlen(skb) reflect the padded size so
that dma_map_single() maps the full buffer and the descriptor length is
consistent. This also removes the temporary 'pad' variable and masking logic.
Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
Cc: stable@vger.kernel.org
Reported-by: Stefan Fleischmann <sfle@kth.se>
Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
Cc: Michael Chan <michael.chan@broadcom.com>
Cc: Pavan Chebbi <pavan.chebbi@broadcom.com>
Cc: Andrew Lunn <andrew+netdev@lunn.ch>
---
drivers/net/ethernet/broadcom/bnxt/bnxt.c | 19 +++++++------------
1 file changed, 7 insertions(+), 12 deletions(-)
diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
index d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..7ea27e81e88c5ca82a453b449982b791e8acc831 100644
--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
+++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
@@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
struct netdev_queue *txq;
int i;
dma_addr_t mapping;
- unsigned int length, pad = 0;
+ unsigned int length;
u32 len, free_size, vlan_tag_flags, cfa_action, flags;
struct bnxt_ptp_cfg *ptp = bp->ptp_cfg;
struct pci_dev *pdev = bp->pdev;
@@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
}
normal_tx:
- if (length < BNXT_MIN_PKT_SIZE) {
- pad = BNXT_MIN_PKT_SIZE - length;
- if (skb_pad(skb, pad))
- /* SKB already freed. */
- goto tx_kick_pending;
- length = BNXT_MIN_PKT_SIZE;
+ if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
+ /* SKB already freed. */
+ goto tx_kick_pending;
}
-
+ length = skb->len;
+ len = skb_headlen(skb);
mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
@@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
txbd->tx_bd_len_flags_type = cpu_to_le32(flags);
}
- flags &= ~TX_BD_LEN;
- txbd->tx_bd_len_flags_type =
- cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
- TX_BD_FLAGS_PACKET_END);
+ txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
netdev_tx_sent_queue(txq, skb->len);
--
2.53.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
2026-10-05 2:38 [PATCH net] bnxt_en: fix DMA mapping length for padded small packets Eric Dumazet
@ 2026-10-05 4:14 ` Michael Chan
2026-10-05 9:49 ` Stefan Fleischmann
2026-10-05 11:18 ` Salvatore Bonaccorso
2026-10-05 21:16 ` netdev-bot+sashiko
2 siblings, 1 reply; 9+ messages in thread
From: Michael Chan @ 2026-10-05 4:14 UTC (permalink / raw)
To: Eric Dumazet
Cc: David S . Miller, Jakub Kicinski, Paolo Abeni, Simon Horman,
netdev, stable, Stefan Fleischmann, Eric Dumazet, Pavan Chebbi,
Andrew Lunn
[-- Attachment #1: Type: text/plain, Size: 1063 bytes --]
On Sun, Oct 4, 2026 at 7:38 PM Eric Dumazet <edumazet@kernel.org> wrote:
> This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
> and races when toggling HW VLAN offload") because reserving extra VLAN
> headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
> offsets and potentially causing small frames to land right against page
> boundaries.
>
> Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx
> path. This ensures skb->len and skb_headlen(skb) reflect the padded size so
> that dma_map_single() maps the full buffer and the descriptor length is
> consistent. This also removes the temporary 'pad' variable and masking logic.
>
> Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
> Cc: stable@vger.kernel.org
> Reported-by: Stefan Fleischmann <sfle@kth.se>
> Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
> Signed-off-by: Eric Dumazet <edumazet@google.com>
Thanks.
Reviewed-by: Michael Chan <michael.chan@broadcom.com>
[-- Attachment #2: S/MIME Cryptographic Signature --]
[-- Type: application/pkcs7-signature, Size: 5469 bytes --]
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
2026-10-05 4:14 ` Michael Chan
@ 2026-10-05 9:49 ` Stefan Fleischmann
0 siblings, 0 replies; 9+ messages in thread
From: Stefan Fleischmann @ 2026-10-05 9:49 UTC (permalink / raw)
To: Michael Chan, Eric Dumazet
Cc: David S . Miller, Jakub Kicinski, Paolo Abeni, Simon Horman,
netdev, stable, Eric Dumazet, Pavan Chebbi, Andrew Lunn
On Sun, 4 Oct 2026 21:14:31 -0700
Michael Chan <michael.chan@broadcom.com> wrote:
> On Sun, Oct 4, 2026 at 7:38 PM Eric Dumazet <edumazet@kernel.org>
> wrote:
>
> > This issue was exposed after commit 447cbe95ebb9 ("vlan: fix
> > skb_under_panic and races when toggling HW VLAN offload") because
> > reserving extra VLAN headroom rounded LL_RESERVED_SPACE from 48 up
> > to 64 bytes, shifting skb->data offsets and potentially causing
> > small frames to land right against page boundaries.
> >
> > Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the
> > normal_tx path. This ensures skb->len and skb_headlen(skb) reflect
> > the padded size so that dma_map_single() maps the full buffer and
> > the descriptor length is consistent. This also removes the
> > temporary 'pad' variable and masking logic.
> >
> > Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
> > Cc: stable@vger.kernel.org
> > Reported-by: Stefan Fleischmann <sfle@kth.se>
> > Closes:
> > https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
> > Signed-off-by: Eric Dumazet <edumazet@google.com>
>
> Thanks.
> Reviewed-by: Michael Chan <michael.chan@broadcom.com>
Thanks a lot! I am testing the patch right now on 6.18.55. The system
is already up for two hours without issues. Before the patch I couldn't
get past 45 minutes without hitting the IOMMU DMA read fault.
Best,
Stefan
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
2026-10-05 2:38 [PATCH net] bnxt_en: fix DMA mapping length for padded small packets Eric Dumazet
2026-10-05 4:14 ` Michael Chan
@ 2026-10-05 11:18 ` Salvatore Bonaccorso
2026-10-05 14:26 ` Bernhard Schmidt
2026-10-05 21:16 ` netdev-bot+sashiko
2 siblings, 1 reply; 9+ messages in thread
From: Salvatore Bonaccorso @ 2026-10-05 11:18 UTC (permalink / raw)
To: Eric Dumazet
Cc: David S . Miller, Jakub Kicinski, Paolo Abeni, Simon Horman,
netdev, stable, Stefan Fleischmann, Eric Dumazet, Michael Chan,
Pavan Chebbi, Andrew Lunn, Bernhard Schmidt
Hi,
On Mon, Oct 05, 2026 at 04:38:12AM +0200, Eric Dumazet wrote:
> Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412
> NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces:
>
> DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000
> [fault reason 0x06] PTE Read access is not set
> bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
> ...
> NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out
>
> The fault address (0xfc499000) is on an exact 4KB page boundary,
> pointing to a DMA read buffer overrun.
>
> In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes),
> such as 42-byte untagged ARP frames, are padded:
>
> if (length < BNXT_MIN_PKT_SIZE) {
> pad = BNXT_MIN_PKT_SIZE - length;
> if (skb_pad(skb, pad))
> goto tx_kick_pending;
> length = BNXT_MIN_PKT_SIZE;
> }
>
> mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
> ...
> dma_unmap_len_set(tx_buf, len, len);
>
> However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes)
> and is left unadjusted after padding. Consequently, dma_map_single() and
> dma_unmap_len_set() map and track only 42 bytes.
>
> Later, the hardware TX buffer descriptor is programmed with the padded length:
>
> txbd->tx_bd_len_flags_type =
> cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
> TX_BD_FLAGS_PACKET_END);
>
> The NIC DMA engine is thus instructed to read 52 bytes from a region where
> only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB
> page (within 'pad' bytes of the next page), the hardware DMA read overruns
> into the unmapped adjacent page, triggering an IOMMU fault.
>
> This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
> and races when toggling HW VLAN offload") because reserving extra VLAN
> headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
> offsets and potentially causing small frames to land right against page
> boundaries.
>
> Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx
> path. This ensures skb->len and skb_headlen(skb) reflect the padded size so
> that dma_map_single() maps the full buffer and the descriptor length is
> consistent. This also removes the temporary 'pad' variable and masking logic.
>
> Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
> Cc: stable@vger.kernel.org
> Reported-by: Stefan Fleischmann <sfle@kth.se>
> Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
> Signed-off-by: Eric Dumazet <edumazet@google.com>
> ---
> Cc: Michael Chan <michael.chan@broadcom.com>
> Cc: Pavan Chebbi <pavan.chebbi@broadcom.com>
> Cc: Andrew Lunn <andrew+netdev@lunn.ch>
> ---
> drivers/net/ethernet/broadcom/bnxt/bnxt.c | 19 +++++++------------
> 1 file changed, 7 insertions(+), 12 deletions(-)
>
> diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> index d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..7ea27e81e88c5ca82a453b449982b791e8acc831 100644
> --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> @@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> struct netdev_queue *txq;
> int i;
> dma_addr_t mapping;
> - unsigned int length, pad = 0;
> + unsigned int length;
> u32 len, free_size, vlan_tag_flags, cfa_action, flags;
> struct bnxt_ptp_cfg *ptp = bp->ptp_cfg;
> struct pci_dev *pdev = bp->pdev;
> @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> }
>
> normal_tx:
> - if (length < BNXT_MIN_PKT_SIZE) {
> - pad = BNXT_MIN_PKT_SIZE - length;
> - if (skb_pad(skb, pad))
> - /* SKB already freed. */
> - goto tx_kick_pending;
> - length = BNXT_MIN_PKT_SIZE;
> + if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
> + /* SKB already freed. */
> + goto tx_kick_pending;
> }
> -
> + length = skb->len;
> + len = skb_headlen(skb);
> mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
>
> if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
> @@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> txbd->tx_bd_len_flags_type = cpu_to_le32(flags);
> }
>
> - flags &= ~TX_BD_LEN;
> - txbd->tx_bd_len_flags_type =
> - cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
> - TX_BD_FLAGS_PACKET_END);
> + txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
>
> netdev_tx_sent_queue(txq, skb->len);
>
> --
> 2.53.0
FWIW, got as well reported in Debian for an update in the 6.12.y
series: https://bugs.debian.org/1149564 , in case you would like to
add a further Link/Closes reference. Bernhard Schmidt is testing the
patch as well on top of 6.12.111 (what we have right now in Debian)
and looks promissing: https://bugs.debian.org/1149564#89 .
Berhard, want to report back a Tested-by from you?
Regards,
Salvatore
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
2026-10-05 11:18 ` Salvatore Bonaccorso
@ 2026-10-05 14:26 ` Bernhard Schmidt
2026-10-07 6:51 ` Fabian Grünbichler
0 siblings, 1 reply; 9+ messages in thread
From: Bernhard Schmidt @ 2026-10-05 14:26 UTC (permalink / raw)
To: Salvatore Bonaccorso
Cc: Eric Dumazet, David S . Miller, Jakub Kicinski, Paolo Abeni,
Simon Horman, netdev, stable, Stefan Fleischmann, Eric Dumazet,
Michael Chan, Pavan Chebbi, Andrew Lunn
On 05/10/26 01:18 PM, Salvatore Bonaccorso wrote:
> Hi,
>
> On Mon, Oct 05, 2026 at 04:38:12AM +0200, Eric Dumazet wrote:
> > Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412
> > NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces:
> >
> > DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000
> > [fault reason 0x06] PTE Read access is not set
> > bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
> > ...
> > NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out
> >
> > The fault address (0xfc499000) is on an exact 4KB page boundary,
> > pointing to a DMA read buffer overrun.
> >
> > In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes),
> > such as 42-byte untagged ARP frames, are padded:
> >
> > if (length < BNXT_MIN_PKT_SIZE) {
> > pad = BNXT_MIN_PKT_SIZE - length;
> > if (skb_pad(skb, pad))
> > goto tx_kick_pending;
> > length = BNXT_MIN_PKT_SIZE;
> > }
> >
> > mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
> > ...
> > dma_unmap_len_set(tx_buf, len, len);
> >
> > However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes)
> > and is left unadjusted after padding. Consequently, dma_map_single() and
> > dma_unmap_len_set() map and track only 42 bytes.
> >
> > Later, the hardware TX buffer descriptor is programmed with the padded length:
> >
> > txbd->tx_bd_len_flags_type =
> > cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
> > TX_BD_FLAGS_PACKET_END);
> >
> > The NIC DMA engine is thus instructed to read 52 bytes from a region where
> > only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB
> > page (within 'pad' bytes of the next page), the hardware DMA read overruns
> > into the unmapped adjacent page, triggering an IOMMU fault.
> >
> > This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
> > and races when toggling HW VLAN offload") because reserving extra VLAN
> > headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
> > offsets and potentially causing small frames to land right against page
> > boundaries.
> >
> > Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx
> > path. This ensures skb->len and skb_headlen(skb) reflect the padded size so
> > that dma_map_single() maps the full buffer and the descriptor length is
> > consistent. This also removes the temporary 'pad' variable and masking logic.
> >
> > Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
> > Cc: stable@vger.kernel.org
> > Reported-by: Stefan Fleischmann <sfle@kth.se>
> > Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
> > Signed-off-by: Eric Dumazet <edumazet@google.com>
> > ---
> > Cc: Michael Chan <michael.chan@broadcom.com>
> > Cc: Pavan Chebbi <pavan.chebbi@broadcom.com>
> > Cc: Andrew Lunn <andrew+netdev@lunn.ch>
> > ---
> > drivers/net/ethernet/broadcom/bnxt/bnxt.c | 19 +++++++------------
> > 1 file changed, 7 insertions(+), 12 deletions(-)
> >
> > diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> > index d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..7ea27e81e88c5ca82a453b449982b791e8acc831 100644
> > --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> > +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> > @@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> > struct netdev_queue *txq;
> > int i;
> > dma_addr_t mapping;
> > - unsigned int length, pad = 0;
> > + unsigned int length;
> > u32 len, free_size, vlan_tag_flags, cfa_action, flags;
> > struct bnxt_ptp_cfg *ptp = bp->ptp_cfg;
> > struct pci_dev *pdev = bp->pdev;
> > @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> > }
> >
> > normal_tx:
> > - if (length < BNXT_MIN_PKT_SIZE) {
> > - pad = BNXT_MIN_PKT_SIZE - length;
> > - if (skb_pad(skb, pad))
> > - /* SKB already freed. */
> > - goto tx_kick_pending;
> > - length = BNXT_MIN_PKT_SIZE;
> > + if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
> > + /* SKB already freed. */
> > + goto tx_kick_pending;
> > }
> > -
> > + length = skb->len;
> > + len = skb_headlen(skb);
> > mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
> >
> > if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
> > @@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> > txbd->tx_bd_len_flags_type = cpu_to_le32(flags);
> > }
> >
> > - flags &= ~TX_BD_LEN;
> > - txbd->tx_bd_len_flags_type =
> > - cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
> > - TX_BD_FLAGS_PACKET_END);
> > + txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
> >
> > netdev_tx_sent_queue(txq, skb->len);
> >
> > --
> > 2.53.0
>
> FWIW, got as well reported in Debian for an update in the 6.12.y
> series: https://bugs.debian.org/1149564 , in case you would like to
> add a further Link/Closes reference. Bernhard Schmidt is testing the
> patch as well on top of 6.12.111 (what we have right now in Debian)
> and looks promissing: https://bugs.debian.org/1149564#89 .
>
> Berhard, want to report back a Tested-by from you?
Looks good. Machine has been stable with that patch applied to 6.12
Tested-by: Bernhard Schmidt <berni@debian.org>
Bernhard
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
2026-10-05 2:38 [PATCH net] bnxt_en: fix DMA mapping length for padded small packets Eric Dumazet
2026-10-05 4:14 ` Michael Chan
2026-10-05 11:18 ` Salvatore Bonaccorso
@ 2026-10-05 21:16 ` netdev-bot+sashiko
2026-10-06 3:56 ` Eric Dumazet
2 siblings, 1 reply; 9+ messages in thread
From: netdev-bot+sashiko @ 2026-10-05 21:16 UTC (permalink / raw)
To: edumazet
Cc: davem, kuba, pabeni, horms, netdev, stable, sfle, edumazet,
michael.chan, pavan.chebbi, andrew+netdev
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 1 · Medium: 0 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] The fix is incomplete for short non-linear skbs, and for them it
changes the wire frame for the worse.
--
bnxt_en: fix DMA mapping length for padded small packets
This patch replaces skb_pad() with skb_put_padto() in the normal_tx path
of bnxt_start_xmit(). After the change, skb->len and skb_headlen() include
the padding, so dma_map_single() maps the whole padded buffer. The pad
variable and the rewrite of the last BD length are removed.
Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
> diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> index d7728d0c5b6e6..7ea27e81e88c5 100644
> --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
[ ... ]
> @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> }
>
> normal_tx:
> - if (length < BNXT_MIN_PKT_SIZE) {
> - pad = BNXT_MIN_PKT_SIZE - length;
> - if (skb_pad(skb, pad))
> - /* SKB already freed. */
> - goto tx_kick_pending;
> - length = BNXT_MIN_PKT_SIZE;
> + if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
> + /* SKB already freed. */
> + goto tx_kick_pending;
> }
> -
> + length = skb->len;
> + len = skb_headlen(skb);
[Severity: High]
What happens here when the skb is non-linear and shorter than
BNXT_MIN_PKT_SIZE?
bnxt_start_xmit() reads last_frag and tx_buf->nr_frags earlier, before
any padding:
last_frag = skb_shinfo(skb)->nr_frags;
...
tx_buf->nr_frags = last_frag;
For a non-linear skb, skb_put_padto()->__skb_pad() skips the tailroom
fast path and calls skb_linearize(). __pskb_pull_tail() then:
- copies the frags into the head
- calls skb_frag_unref() on each one
- sets nr_frags to 0
The frags[] slots themselves are not cleared.
After that, len = skb_headlen(skb) covers all 52 padded bytes, but
last_frag still holds the old count. TX_BD_CNT(last_frag + 2) and
SET_TX_OPAQUE(bp, txr, prod, 2 + last_frag) still count the old
fragments. The frag loop also still runs over them:
for (i = 0; i < last_frag; i++) {
frag = &skb_shinfo(skb)->frags[i];
...
mapping = skb_frag_dma_map(&pdev->dev, frag, 0, len,
DMA_TO_DEVICE);
This emits a BD for each stale frag, and TX_BD_FLAGS_PACKET_END is set on
the last stale BD:
txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
As an example, take a 42-byte frame with a 14-byte linear head and one
28-byte frag. AF_PACKET PACKET_TX_RING can build one through
tpacket_fill_skb(). The head BD would describe 52 bytes and the stale frag
BD another 28. Wouldn't that put an 80-byte frame on the wire, with the
payload duplicated?
The NIC would also DMA-read pages that the skb no longer holds a
reference on. If this skb held the last reference, those pages could
already be freed or recycled.
The completion path in __bnxt_tx_int() (last = tx_buf->nr_frags) and the
tx_dma_error unwind path also trust the stale count.
bnxt sets NETIF_F_SG, so the core does not linearize these skbs first.
They reach normal_tx whenever the push path is skipped, for example when
tx_push_thresh is 0, the ring is not empty, or lflags is set.
The old code mapped the same stale frags, so that part is not new. Its
total descriptor length still came to 52, though, because the pad was
added to the last BD. The oversized frame with duplicated data seems to
come from this change. It also doesn't match this part of the commit
message:
so that dma_map_single() maps the full buffer and the descriptor
length is consistent.
Could the padding be done before length, len, last_frag and
tx_buf->nr_frags are read, and before the free_size check? Or could
last_frag and tx_buf->nr_frags be read again after skb_put_padto()
succeeds?
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261005023812.130639-1-edumazet%40kernel.org
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
2026-10-05 21:16 ` netdev-bot+sashiko
@ 2026-10-06 3:56 ` Eric Dumazet
2026-10-06 3:59 ` Eric Dumazet
0 siblings, 1 reply; 9+ messages in thread
From: Eric Dumazet @ 2026-10-06 3:56 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: davem, kuba, pabeni, horms, netdev, stable, sfle, edumazet,
michael.chan, pavan.chebbi, andrew+netdev
Le lun. 5 oct. 2026 à 23:16, <netdev-bot+sashiko@kernel.org> a écrit :
>
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
>
> Critical: 0 · High: 1 · Medium: 0 · Low: 0
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [High] The fix is incomplete for short non-linear skbs, and for them it
> changes the wire frame for the worse.
>
> --
>
> bnxt_en: fix DMA mapping length for padded small packets
>
> This patch replaces skb_pad() with skb_put_padto() in the normal_tx path
> of bnxt_start_xmit(). After the change, skb->len and skb_headlen() include
> the padding, so dma_map_single() maps the whole padded buffer. The pad
> variable and the rewrite of the last BD length are removed.
>
> Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
>
> > diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> > index d7728d0c5b6e6..7ea27e81e88c5 100644
> > --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
> > +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
>
> [ ... ]
>
> > @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
> > }
> >
> > normal_tx:
> > - if (length < BNXT_MIN_PKT_SIZE) {
> > - pad = BNXT_MIN_PKT_SIZE - length;
> > - if (skb_pad(skb, pad))
> > - /* SKB already freed. */
> > - goto tx_kick_pending;
> > - length = BNXT_MIN_PKT_SIZE;
> > + if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
> > + /* SKB already freed. */
> > + goto tx_kick_pending;
> > }
> > -
> > + length = skb->len;
> > + len = skb_headlen(skb);
>
> [Severity: High]
> What happens here when the skb is non-linear and shorter than
> BNXT_MIN_PKT_SIZE?
>
> bnxt_start_xmit() reads last_frag and tx_buf->nr_frags earlier, before
> any padding:
>
> last_frag = skb_shinfo(skb)->nr_frags;
> ...
> tx_buf->nr_frags = last_frag;
>
> For a non-linear skb, skb_put_padto()->__skb_pad() skips the tailroom
> fast path and calls skb_linearize(). __pskb_pull_tail() then:
>
> - copies the frags into the head
> - calls skb_frag_unref() on each one
> - sets nr_frags to 0
>
> The frags[] slots themselves are not cleared.
>
> After that, len = skb_headlen(skb) covers all 52 padded bytes, but
> last_frag still holds the old count. TX_BD_CNT(last_frag + 2) and
> SET_TX_OPAQUE(bp, txr, prod, 2 + last_frag) still count the old
> fragments. The frag loop also still runs over them:
>
> for (i = 0; i < last_frag; i++) {
> frag = &skb_shinfo(skb)->frags[i];
> ...
> mapping = skb_frag_dma_map(&pdev->dev, frag, 0, len,
> DMA_TO_DEVICE);
>
> This emits a BD for each stale frag, and TX_BD_FLAGS_PACKET_END is set on
> the last stale BD:
>
> txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
>
> As an example, take a 42-byte frame with a 14-byte linear head and one
> 28-byte frag. AF_PACKET PACKET_TX_RING can build one through
> tpacket_fill_skb(). The head BD would describe 52 bytes and the stale frag
> BD another 28. Wouldn't that put an 80-byte frame on the wire, with the
> payload duplicated?
>
> The NIC would also DMA-read pages that the skb no longer holds a
> reference on. If this skb held the last reference, those pages could
> already be freed or recycled.
>
> The completion path in __bnxt_tx_int() (last = tx_buf->nr_frags) and the
> tx_dma_error unwind path also trust the stale count.
>
> bnxt sets NETIF_F_SG, so the core does not linearize these skbs first.
> They reach normal_tx whenever the push path is skipped, for example when
> tx_push_thresh is 0, the ring is not empty, or lflags is set.
>
> The old code mapped the same stale frags, so that part is not new. Its
> total descriptor length still came to 52, though, because the pad was
> added to the last BD. The oversized frame with duplicated data seems to
> come from this change. It also doesn't match this part of the commit
> message:
>
> so that dma_map_single() maps the full buffer and the descriptor
> length is consistent.
>
> Could the padding be done before length, len, last_frag and
> tx_buf->nr_frags are read, and before the free_size check? Or could
> last_frag and tx_buf->nr_frags be read again after skb_put_padto()
> succeeds?
Looks legit, I am cooking a V2, thanks.
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
2026-10-06 3:56 ` Eric Dumazet
@ 2026-10-06 3:59 ` Eric Dumazet
0 siblings, 0 replies; 9+ messages in thread
From: Eric Dumazet @ 2026-10-06 3:59 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: davem, kuba, pabeni, horms, netdev, stable, sfle, edumazet,
michael.chan, pavan.chebbi, andrew+netdev
Le mar. 6 oct. 2026 à 05:56, Eric Dumazet <edumazet@kernel.org> a écrit :
>
> Looks legit, I am cooking a V2, thanks.
pw-bot: cr
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH net] bnxt_en: fix DMA mapping length for padded small packets
2026-10-05 14:26 ` Bernhard Schmidt
@ 2026-10-07 6:51 ` Fabian Grünbichler
0 siblings, 0 replies; 9+ messages in thread
From: Fabian Grünbichler @ 2026-10-07 6:51 UTC (permalink / raw)
To: Bernhard Schmidt, Salvatore Bonaccorso
Cc: Andrew Lunn, David S . Miller, Eric Dumazet, Eric Dumazet,
Simon Horman, Jakub Kicinski, Michael Chan, netdev, Paolo Abeni,
Pavan Chebbi, Stefan Fleischmann, stable
On October 5, 2026 4:26 pm, Bernhard Schmidt wrote:
> On 05/10/26 01:18 PM, Salvatore Bonaccorso wrote:
>> Hi,
>>
>> On Mon, Oct 05, 2026 at 04:38:12AM +0200, Eric Dumazet wrote:
>> > Stefan Fleischmann reported Intel IOMMU DMA Read faults on BCM57412
>> > NetXtreme-E NICs when transmitting packets on VLAN/macvlan interfaces:
>> >
>> > DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000
>> > [fault reason 0x06] PTE Read access is not set
>> > bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
>> > ...
>> > NETDEV WATCHDOG: eno1np0 (bnxt_en): transmit queue 0 timed out
>> >
>> > The fault address (0xfc499000) is on an exact 4KB page boundary,
>> > pointing to a DMA read buffer overrun.
>> >
>> > In bnxt_start_xmit(), packets smaller than BNXT_MIN_PKT_SIZE (52 bytes),
>> > such as 42-byte untagged ARP frames, are padded:
>> >
>> > if (length < BNXT_MIN_PKT_SIZE) {
>> > pad = BNXT_MIN_PKT_SIZE - length;
>> > if (skb_pad(skb, pad))
>> > goto tx_kick_pending;
>> > length = BNXT_MIN_PKT_SIZE;
>> > }
>> >
>> > mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
>> > ...
>> > dma_unmap_len_set(tx_buf, len, len);
>> >
>> > However, 'len' was initialized earlier to skb_headlen(skb) (e.g. 42 bytes)
>> > and is left unadjusted after padding. Consequently, dma_map_single() and
>> > dma_unmap_len_set() map and track only 42 bytes.
>> >
>> > Later, the hardware TX buffer descriptor is programmed with the padded length:
>> >
>> > txbd->tx_bd_len_flags_type =
>> > cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
>> > TX_BD_FLAGS_PACKET_END);
>> >
>> > The NIC DMA engine is thus instructed to read 52 bytes from a region where
>> > only 42 bytes were DMA-mapped. If skb->data ends near the boundary of a 4KB
>> > page (within 'pad' bytes of the next page), the hardware DMA read overruns
>> > into the unmapped adjacent page, triggering an IOMMU fault.
>> >
>> > This issue was exposed after commit 447cbe95ebb9 ("vlan: fix skb_under_panic
>> > and races when toggling HW VLAN offload") because reserving extra VLAN
>> > headroom rounded LL_RESERVED_SPACE from 48 up to 64 bytes, shifting skb->data
>> > offsets and potentially causing small frames to land right against page
>> > boundaries.
>> >
>> > Fix this by using skb_put_padto(skb, BNXT_MIN_PKT_SIZE) in the normal_tx
>> > path. This ensures skb->len and skb_headlen(skb) reflect the padded size so
>> > that dma_map_single() maps the full buffer and the descriptor length is
>> > consistent. This also removes the temporary 'pad' variable and masking logic.
>> >
>> > Fixes: c0c050c58d84 ("bnxt_en: New Broadcom ethernet driver.")
>> > Cc: stable@vger.kernel.org
>> > Reported-by: Stefan Fleischmann <sfle@kth.se>
>> > Closes: https://lore.kernel.org/netdev/20261004122616.56714cbd@nargothrond/
>> > Signed-off-by: Eric Dumazet <edumazet@google.com>
>> > ---
>> > Cc: Michael Chan <michael.chan@broadcom.com>
>> > Cc: Pavan Chebbi <pavan.chebbi@broadcom.com>
>> > Cc: Andrew Lunn <andrew+netdev@lunn.ch>
>> > ---
>> > drivers/net/ethernet/broadcom/bnxt/bnxt.c | 19 +++++++------------
>> > 1 file changed, 7 insertions(+), 12 deletions(-)
>> >
>> > diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
>> > index d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..7ea27e81e88c5ca82a453b449982b791e8acc831 100644
>> > --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
>> > +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
>> > @@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
>> > struct netdev_queue *txq;
>> > int i;
>> > dma_addr_t mapping;
>> > - unsigned int length, pad = 0;
>> > + unsigned int length;
>> > u32 len, free_size, vlan_tag_flags, cfa_action, flags;
>> > struct bnxt_ptp_cfg *ptp = bp->ptp_cfg;
>> > struct pci_dev *pdev = bp->pdev;
>> > @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
>> > }
>> >
>> > normal_tx:
>> > - if (length < BNXT_MIN_PKT_SIZE) {
>> > - pad = BNXT_MIN_PKT_SIZE - length;
>> > - if (skb_pad(skb, pad))
>> > - /* SKB already freed. */
>> > - goto tx_kick_pending;
>> > - length = BNXT_MIN_PKT_SIZE;
>> > + if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) {
>> > + /* SKB already freed. */
>> > + goto tx_kick_pending;
>> > }
>> > -
>> > + length = skb->len;
>> > + len = skb_headlen(skb);
>> > mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
>> >
>> > if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
>> > @@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
>> > txbd->tx_bd_len_flags_type = cpu_to_le32(flags);
>> > }
>> >
>> > - flags &= ~TX_BD_LEN;
>> > - txbd->tx_bd_len_flags_type =
>> > - cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags |
>> > - TX_BD_FLAGS_PACKET_END);
>> > + txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END);
>> >
>> > netdev_tx_sent_queue(txq, skb->len);
>> >
>> > --
>> > 2.53.0
>>
>> FWIW, got as well reported in Debian for an update in the 6.12.y
>> series: https://bugs.debian.org/1149564 , in case you would like to
>> add a further Link/Closes reference. Bernhard Schmidt is testing the
>> patch as well on top of 6.12.111 (what we have right now in Debian)
>> and looks promissing: https://bugs.debian.org/1149564#89 .
>>
>> Berhard, want to report back a Tested-by from you?
>
> Looks good. Machine has been stable with that patch applied to 6.12
>
> Tested-by: Bernhard Schmidt <berni@debian.org>
>
> Bernhard
Our users' testing (on top of 7.0.14 with 7.1.x and 7.2.x stable patches
mostly added on top) came back positive as well. Since I don't have
access to the hardware myself, and can neither reproduce nor rest, not
sending a T-b..
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-10-07 6:51 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 2:38 [PATCH net] bnxt_en: fix DMA mapping length for padded small packets Eric Dumazet
2026-10-05 4:14 ` Michael Chan
2026-10-05 9:49 ` Stefan Fleischmann
2026-10-05 11:18 ` Salvatore Bonaccorso
2026-10-05 14:26 ` Bernhard Schmidt
2026-10-07 6:51 ` Fabian Grünbichler
2026-10-05 21:16 ` netdev-bot+sashiko
2026-10-06 3:56 ` Eric Dumazet
2026-10-06 3:59 ` Eric Dumazet
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox