* [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en
@ 2026-10-04 10:26 Stefan Fleischmann
2026-10-04 11:46 ` Eric Dumazet
0 siblings, 1 reply; 15+ messages in thread
From: Stefan Fleischmann @ 2026-10-04 10:26 UTC (permalink / raw)
To: netdev, stable; +Cc: Michael Chan, Pavan Chebbi, regressions, edumazet
[-- Attachment #1: Type: text/plain, Size: 1605 bytes --]
Hi,
I have bisected a regression introduced in kernel 6.18.51 (and 7.2.3).
The issue is still present in 6.18.55 and 7.2.9. Verified by reverting
1517d1996b52 (the equivalent of 447cbe95ebb9 in the 6.18 branch) on top
of 6.18.51 and 6.18.55.
The hardware is a Dell R640 server with Intel Xeon Silver 4210 and
Broadcom BCM57412 NetXtreme-E NIC.
The network interface is configured as an untagged interface and has a
couple of tagged VLAN interfaces configured on top. All of these are
used by LXC containers with Macvlan.
The interface comes up fine after boot, but after some time (anywhere
between 2 to 40 minutes) I see a DMA error in the kernel log (attached),
and the driver ends up tearing down the interface. It can be brought up
again after that, but it will hit a DMA error again after a while and
the interface goes down again.
I'm happy to provide more debug info if needed, or test patches.
I've already updated to the latest firmware released by Dell, the problem
persist. Here is some more info on the network device:
18:00.0 Ethernet controller [0200]: Broadcom Inc. and subsidiaries BCM57412 NetXtreme-E 10Gb RDMA Ethernet Controller [14e4:16d6] (rev 01)
Subsystem: Broadcom Inc. and subsidiaries NetXtreme E-Series Advanced Dual-port 10Gb SFP+ Ethernet Network Daughter Card [14e4:4120]
# ethtool -i eno1np0
driver: bnxt_en
version: 7.2.9-1-generic
firmware-version: 238.1.168.0/pkg 38.11.70.00
expansion-rom-version:
bus-info: 0000:18:00.0
supports-statistics: yes
supports-test: yes
supports-eeprom-access: yes
supports-register-dump: yes
supports-priv-flags: no
Best,
Stefan
[-- Attachment #2: bnxt_en-kernel-log.txt --]
[-- Type: text/plain, Size: 10806 bytes --]
2026-09-28T01:41:06.242437+02:00 r640-01 kernel: DMAR: DRHD: handling fault status reg 2
2026-09-28T01:41:06.242461+02:00 r640-01 kernel: DMAR: [DMA Read NO_PASID] Request device [18:00.0] fault addr 0xfc499000 [fault reason 0x06] PTE Read access is not set
2026-09-28T01:41:06.310547+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x41a} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:07.088637+02:00 r640-01 kernel: kauditd_printk_skb: 488 callbacks suppressed
2026-09-28T01:41:07.321589+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x23 seq id 0x41b error 0xf
2026-09-28T01:41:07.330570+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0xb4 seq id 0x41c error 0xf
2026-09-28T01:41:08.348631+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x41d} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:09.372623+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x420} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:10.396561+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x423} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:12.444634+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x429} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:13.468622+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x42c} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:14.492626+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x42f} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:15.516628+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x432} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:16.540617+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x435} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:19.612571+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x43d} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:20.636560+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x440} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:21.660627+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0xb4 0x444} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:23.718601+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x447} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:26.780616+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x44f} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:27.092594+02:00 r640-01 kernel: kauditd_printk_skb: 222 callbacks suppressed
2026-09-28T01:41:27.804557+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x452} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:31.900559+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x23 0x45c} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:34.906960+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: NETDEV WATCHDOG: CPU: 9: transmit queue 1 timed out 6009 ms
2026-09-28T01:41:34.922831+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [0.0]: tx{fw_ring: 0 prod: 2c5e cons: 2c50}
2026-09-28T01:41:34.922849+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [0]: rx{fw_ring: 1 prod: 4ba7} rx_agg{fw_ring: 6 agg_prod: a6b sw_agg_prod: 26b}
2026-09-28T01:41:34.922853+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [0]: cp{fw_ring: 0 raw_cons: a7d7}
2026-09-28T01:41:34.922854+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [1.0]: tx{fw_ring: 1 prod: 58a1 cons: 56b2}
2026-09-28T01:41:34.922856+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [1]: rx{fw_ring: 2 prod: 1025} rx_agg{fw_ring: 7 agg_prod: 8d0 sw_agg_prod: d0}
2026-09-28T01:41:34.922857+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [1]: cp{fw_ring: 16 raw_cons: 42d4}
2026-09-28T01:41:34.922859+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [2.0]: tx{fw_ring: 2 prod: 300c cons: 2fe2}
2026-09-28T01:41:34.922861+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [2]: rx{fw_ring: 3 prod: 1152} rx_agg{fw_ring: 8 agg_prod: 8ce sw_agg_prod: ce}
2026-09-28T01:41:34.922862+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [2]: cp{fw_ring: 17 raw_cons: 3300}
2026-09-28T01:41:34.922864+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [3.0]: tx{fw_ring: 3 prod: 1e3c cons: 1e18}
2026-09-28T01:41:34.922865+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [3]: rx{fw_ring: 4 prod: 23aa} rx_agg{fw_ring: 9 agg_prod: 937 sw_agg_prod: 137}
2026-09-28T01:41:34.922866+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [3]: cp{fw_ring: 18 raw_cons: 4fe4}
2026-09-28T01:41:34.922867+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [4.0]: tx{fw_ring: 4 prod: 1dd4 cons: 1dad}
2026-09-28T01:41:34.922869+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [4]: rx{fw_ring: 5 prod: 1a86} rx_agg{fw_ring: 10 agg_prod: 153e sw_agg_prod: 53e}
2026-09-28T01:41:34.922870+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: [4]: cp{fw_ring: 19 raw_cons: 4a5d}
2026-09-28T01:41:34.922871+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x91 seq id 0x465 error 0x3
2026-09-28T01:41:34.931036+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x46 seq id 0x466 error 0x2
2026-09-28T01:41:34.931549+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x71 seq id 0x467 error 0x3
2026-09-28T01:41:34.939553+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x44 seq id 0x468 error 0x3
2026-09-28T01:41:34.955420+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm vnic set tpa failure rc for vnic 0: fffffff3
2026-09-28T01:41:34.963548+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x41 seq id 0x469 error 0x3
2026-09-28T01:41:34.982633+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x41 seq id 0x46b error 0x3
2026-09-28T01:41:34.982651+02:00 r640-01 kernel: workqueue: drm_fb_helper_damage_work hogged CPU for >10000us 7 times, consider switching to WQ_UNBOUND
2026-09-28T01:41:35.036321+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Resp cmpl intr abandoning msg: 0x51 due to firmware status: 0x2000001
2026-09-28T01:41:35.045234+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 1 failed. rc:fffffff0 err:3
2026-09-28T01:41:35.065242+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 1 failed. rc:fffffff0 err:3
2026-09-28T01:41:35.116722+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Resp cmpl intr abandoning msg: 0x51 due to firmware status: 0x2000001
2026-09-28T01:41:35.125713+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:3
2026-09-28T01:41:35.171632+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Resp cmpl intr abandoning msg: 0x51 due to firmware status: 0x2000001
2026-09-28T01:41:35.207067+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Resp cmpl intr abandoning msg: 0x51 due to firmware status: 0x2000001
2026-09-28T01:41:35.207079+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:3
2026-09-28T01:41:35.227564+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:3
2026-09-28T01:41:35.249621+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Resp cmpl intr abandoning msg: 0x51 due to firmware status: 0x2000001
2026-09-28T01:41:35.268237+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:3
2026-09-28T01:41:35.278721+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Resp cmpl intr abandoning msg: 0x51 due to firmware status: 0x2000001
2026-09-28T01:41:35.287205+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:3
2026-09-28T01:41:35.297692+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Resp cmpl intr abandoning msg: 0x51 due to firmware status: 0x2000001
2026-09-28T01:41:35.306162+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 2 failed. rc:fffffff0 err:3
2026-09-28T01:41:35.314227+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x51 seq id 0x47e error 0x3
2026-09-28T01:41:35.355835+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x51 seq id 0x480 error 0x3
2026-09-28T01:41:35.363894+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x51 seq id 0x481 error 0x3
2026-09-28T01:41:35.372400+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 0 failed. rc:fffffff3 err:3
2026-09-28T01:41:35.389012+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm_ring_free type 0 failed. rc:fffffff3 err:3
2026-09-28T01:41:35.397068+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x61 seq id 0x483 error 0x3
2026-09-28T01:41:35.415553+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x61 seq id 0x485 error 0x3
2026-09-28T01:41:35.423615+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x61 seq id 0x486 error 0x3
2026-09-28T01:41:35.423632+02:00 r640-01 kernel: workqueue: drm_fb_helper_damage_work hogged CPU for >10000us 35 times, consider switching to WQ_UNBOUND
2026-09-28T01:41:35.431702+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0x61 seq id 0x487 error 0x3
2026-09-28T01:41:35.439788+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0xb1 seq id 0x488 error 0x3
2026-09-28T01:41:35.447860+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0xb1 seq id 0x489 error 0x3
2026-09-28T01:41:35.455951+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0xb1 seq id 0x48a error 0x3
2026-09-28T01:41:35.464036+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0xb1 seq id 0x48b error 0x3
2026-09-28T01:41:35.472127+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0xb1 seq id 0x48c error 0x3
2026-09-28T01:41:35.502327+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm req_type 0xb0 seq id 0x48d error 0x3
2026-09-28T01:41:35.502351+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: hwrm stat ctx alloc failure rc: fffffff3
2026-09-28T01:41:35.513566+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: bnxt_init_nic err: fffffff3
2026-09-28T01:41:35.541674+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: nic open fail (rc: fffffff3)
2026-09-28T01:41:35.541691+02:00 r640-01 kernel: bnxt_en 0000:18:00.0 eno1np0: Abandoning msg {0x20 0x48e} len: 0 due to firmware status: 0x2000001
2026-09-28T01:41:37.304638+02:00 r640-01 kernel: kauditd_printk_skb: 7 callbacks suppressed
^ permalink raw reply [flat|nested] 15+ messages in thread* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-04 10:26 [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en Stefan Fleischmann @ 2026-10-04 11:46 ` Eric Dumazet 2026-10-04 14:35 ` Stefan Fleischmann 2026-10-05 17:38 ` Joe Damato 0 siblings, 2 replies; 15+ messages in thread From: Eric Dumazet @ 2026-10-04 11:46 UTC (permalink / raw) To: Stefan Fleischmann, netdev, stable Cc: Michael Chan, Pavan Chebbi, regressions, Joe Damato On 10/4/26 12:26, Stefan Fleischmann wrote: > Hi, > > I have bisected a regression introduced in kernel 6.18.51 (and 7.2.3). > The issue is still present in 6.18.55 and 7.2.9. Verified by reverting > 1517d1996b52 (the equivalent of 447cbe95ebb9 in the 6.18 branch) on top > of 6.18.51 and 6.18.55. > > The hardware is a Dell R640 server with Intel Xeon Silver 4210 and > Broadcom BCM57412 NetXtreme-E NIC. > > The network interface is configured as an untagged interface and has a > couple of tagged VLAN interfaces configured on top. All of these are > used by LXC containers with Macvlan. > > The interface comes up fine after boot, but after some time (anywhere > between 2 to 40 minutes) I see a DMA error in the kernel log (attached), > and the driver ends up tearing down the interface. It can be brought up > again after that, but it will hit a DMA error again after a while and > the interface goes down again. > > I'm happy to provide more debug info if needed, or test patches. Hi Stefan, Thanks for the bisect and detailed report. Notice that 0xfc499000 is on an exact 4KB page boundary. This points to a DMA read overrun where the Broadcom DMA engine reads past the end of a buffer mapped in page 0xfc498xxx into the adjacent unmapped page 0xfc499000. Have you tried a recent net kernel ? Could you please test whether disabling TSO or hardware VLAN offload prevents the crash? (maybe adding one option at a time) # ethtool -K eno1np0 tso off # ethtool -K eno1np0 tx-vlan-offload off Adding Joe to this thread, because some parts in tso_dma_map_init() or tso_start() might have bugs. # Disable USO on the physical interface ethtool -K eno1np0 tx-udp-segmentation off This reminds me on a prior attempt I made months ago to sanitize tso_start() Note that my email address has changed to edumazet@kernel.org Thanks. > > I've already updated to the latest firmware released by Dell, the problem > persist. Here is some more info on the network device: > > 18:00.0 Ethernet controller [0200]: Broadcom Inc. and subsidiaries BCM57412 NetXtreme-E 10Gb RDMA Ethernet Controller [14e4:16d6] (rev 01) > Subsystem: Broadcom Inc. and subsidiaries NetXtreme E-Series Advanced Dual-port 10Gb SFP+ Ethernet Network Daughter Card [14e4:4120] > > # ethtool -i eno1np0 > driver: bnxt_en > version: 7.2.9-1-generic > firmware-version: 238.1.168.0/pkg 38.11.70.00 > expansion-rom-version: > bus-info: 0000:18:00.0 > supports-statistics: yes > supports-test: yes > supports-eeprom-access: yes > supports-register-dump: yes > supports-priv-flags: no > > Best, > Stefan ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-04 11:46 ` Eric Dumazet @ 2026-10-04 14:35 ` Stefan Fleischmann 2026-10-04 17:29 ` Stefan Fleischmann 2026-10-05 17:38 ` Joe Damato 1 sibling, 1 reply; 15+ messages in thread From: Stefan Fleischmann @ 2026-10-04 14:35 UTC (permalink / raw) To: Eric Dumazet Cc: netdev, stable, Michael Chan, Pavan Chebbi, regressions, Joe Damato On Sun, 4 Oct 2026 13:46:08 +0200 Eric Dumazet <edumazet@kernel.org> wrote: > > Notice that 0xfc499000 is on an exact 4KB page boundary. This points > to a DMA read overrun where the Broadcom DMA engine reads past the > end of a buffer mapped in page 0xfc498xxx into the adjacent unmapped > page 0xfc499000. > > Have you tried a recent net kernel ? Hi Eric, that might be a bit tricky. We use ZFS on this server and the version we have installed only supports up to kernel 7.2. > Could you please test whether disabling TSO or hardware VLAN offload > prevents the crash? (maybe adding one option at a time) > > # ethtool -K eno1np0 tso off > > # ethtool -K eno1np0 tx-vlan-offload off I tested this now and none of these options made a difference. I've also tested disabling other offloads, no difference. Best, Stefan > Adding Joe to this thread, because some parts in > tso_dma_map_init() or tso_start() might have bugs. > > # Disable USO on the physical interface > ethtool -K eno1np0 tx-udp-segmentation off > > This reminds me on a prior attempt I made months ago to sanitize > tso_start() > > Note that my email address has changed to edumazet@kernel.org > > Thanks. > ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-04 14:35 ` Stefan Fleischmann @ 2026-10-04 17:29 ` Stefan Fleischmann 2026-10-04 20:52 ` Eric Dumazet 0 siblings, 1 reply; 15+ messages in thread From: Stefan Fleischmann @ 2026-10-04 17:29 UTC (permalink / raw) To: Eric Dumazet Cc: netdev, stable, Michael Chan, Pavan Chebbi, regressions, Joe Damato On Sun, 4 Oct 2026 16:35:32 +0200 Stefan Fleischmann <sfle@kth.se> wrote: > On Sun, 4 Oct 2026 13:46:08 +0200 > Eric Dumazet <edumazet@kernel.org> wrote: > > > > > > Notice that 0xfc499000 is on an exact 4KB page boundary. This points > > to a DMA read overrun where the Broadcom DMA engine reads past the > > end of a buffer mapped in page 0xfc498xxx into the adjacent unmapped > > page 0xfc499000. > > > > Have you tried a recent net kernel ? > > Hi Eric, > > that might be a bit tricky. We use ZFS on this server and the version > we have installed only supports up to kernel 7.2. Scratch that, I noticed that this even happens with none of the LXC containers running. So I tested with the main branch from https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net.git (commit 6dc989ea46b9) Same issue. Best, Stefan > > Could you please test whether disabling TSO or hardware VLAN > > offload prevents the crash? (maybe adding one option at a time) > > > > # ethtool -K eno1np0 tso off > > > > # ethtool -K eno1np0 tx-vlan-offload off > > I tested this now and none of these options made a difference. I've > also tested disabling other offloads, no difference. > > Best, > Stefan > > > Adding Joe to this thread, because some parts in > > tso_dma_map_init() or tso_start() might have bugs. > > > > # Disable USO on the physical interface > > ethtool -K eno1np0 tx-udp-segmentation off > > > > This reminds me on a prior attempt I made months ago to sanitize > > tso_start() > > > > Note that my email address has changed to edumazet@kernel.org > > > > Thanks. > > ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-04 17:29 ` Stefan Fleischmann @ 2026-10-04 20:52 ` Eric Dumazet 2026-10-04 21:11 ` Eric Dumazet 0 siblings, 1 reply; 15+ messages in thread From: Eric Dumazet @ 2026-10-04 20:52 UTC (permalink / raw) To: Stefan Fleischmann Cc: netdev, stable, Michael Chan, Pavan Chebbi, regressions, Joe Damato On 10/4/26 19:29, Stefan Fleischmann wrote: > On Sun, 4 Oct 2026 16:35:32 +0200 > Stefan Fleischmann <sfle@kth.se> wrote: > >> On Sun, 4 Oct 2026 13:46:08 +0200 >> Eric Dumazet <edumazet@kernel.org> wrote: >> >> >>> >>> Notice that 0xfc499000 is on an exact 4KB page boundary. This points >>> to a DMA read overrun where the Broadcom DMA engine reads past the >>> end of a buffer mapped in page 0xfc498xxx into the adjacent unmapped >>> page 0xfc499000. >>> >>> Have you tried a recent net kernel ? >> >> Hi Eric, >> >> that might be a bit tricky. We use ZFS on this server and the version >> we have installed only supports up to kernel 7.2. > > Scratch that, I noticed that this even happens with none of the LXC > containers running. So I tested with the main branch from > https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net.git > (commit 6dc989ea46b9) > > Same issue. Okay, this must be a bnxt issue, that has been hidden years because of some skb->data headroom/offset. LL_RESERVED_SPACE() has been increased from 48 to 64. So perhaps small packets are now crossing a page boundary (which should be fine) I see one bug in the skb_pad() vicinity. Could you try: diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c index d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..f9feeb4471a8cf310788ef8d00b68a1508a7e78e 100644 --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c @@ -678,6 +678,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) /* SKB already freed. */ goto tx_kick_pending; length = BNXT_MIN_PKT_SIZE; + len += pad; } mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE); ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-04 20:52 ` Eric Dumazet @ 2026-10-04 21:11 ` Eric Dumazet 2026-10-04 22:29 ` Michael Chan 0 siblings, 1 reply; 15+ messages in thread From: Eric Dumazet @ 2026-10-04 21:11 UTC (permalink / raw) To: Stefan Fleischmann Cc: netdev, stable, Michael Chan, Pavan Chebbi, regressions, Joe Damato On 10/4/26 22:52, Eric Dumazet wrote: > > > On 10/4/26 19:29, Stefan Fleischmann wrote: >> On Sun, 4 Oct 2026 16:35:32 +0200 >> Stefan Fleischmann <sfle@kth.se> wrote: >> >>> On Sun, 4 Oct 2026 13:46:08 +0200 >>> Eric Dumazet <edumazet@kernel.org> wrote: >>> >>> >>>> >>>> Notice that 0xfc499000 is on an exact 4KB page boundary. This points >>>> to a DMA read overrun where the Broadcom DMA engine reads past the >>>> end of a buffer mapped in page 0xfc498xxx into the adjacent unmapped >>>> page 0xfc499000. >>>> >>>> Have you tried a recent net kernel ? >>> >>> Hi Eric, >>> >>> that might be a bit tricky. We use ZFS on this server and the version >>> we have installed only supports up to kernel 7.2. >> >> Scratch that, I noticed that this even happens with none of the LXC >> containers running. So I tested with the main branch from >> https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net.git >> (commit 6dc989ea46b9) >> >> Same issue. > > Okay, this must be a bnxt issue, that has been hidden years because of > some skb->data headroom/offset. > > LL_RESERVED_SPACE() has been increased from 48 to 64. So perhaps small > packets are now crossing a page boundary (which should be fine) > > I see one bug in the skb_pad() vicinity. Could you try: > > > diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ > ethernet/broadcom/bnxt/bnxt.c > index > d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..f9feeb4471a8cf310788ef8d00b68a1508a7e78e 100644 > --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c > +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c > @@ -678,6 +678,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff > *skb, struct net_device *dev) > /* SKB already freed. */ > goto tx_kick_pending; > length = BNXT_MIN_PKT_SIZE; > + len += pad; > } > > mapping = dma_map_single(&pdev->dev, skb->data, len, > DMA_TO_DEVICE); A more polished patch would be this one. bnxt_start_xmit() needs an audit I think :/ diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c index d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..48e544544754b04cff891bda5bc86cc370e0aa0b 100644 --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c @@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) struct netdev_queue *txq; int i; dma_addr_t mapping; - unsigned int length, pad = 0; + unsigned int length; u32 len, free_size, vlan_tag_flags, cfa_action, flags; struct bnxt_ptp_cfg *ptp = bp->ptp_cfg; struct pci_dev *pdev = bp->pdev; @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) } normal_tx: - if (length < BNXT_MIN_PKT_SIZE) { - pad = BNXT_MIN_PKT_SIZE - length; - if (skb_pad(skb, pad)) - /* SKB already freed. */ - goto tx_kick_pending; - length = BNXT_MIN_PKT_SIZE; + if (skb_padto(skb, BNXT_MIN_PKT_SIZE)) { + /* SKB already freed. */ + goto tx_kick_pending; } - + length = skb->len; + len = skb_headlen(skb); mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE); if (unlikely(dma_mapping_error(&pdev->dev, mapping))) @@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) txbd->tx_bd_len_flags_type = cpu_to_le32(flags); } - flags &= ~TX_BD_LEN; - txbd->tx_bd_len_flags_type = - cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags | - TX_BD_FLAGS_PACKET_END); + txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END); netdev_tx_sent_queue(txq, skb->len); ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-04 21:11 ` Eric Dumazet @ 2026-10-04 22:29 ` Michael Chan 2026-10-05 1:59 ` Eric Dumazet 0 siblings, 1 reply; 15+ messages in thread From: Michael Chan @ 2026-10-04 22:29 UTC (permalink / raw) To: Eric Dumazet Cc: Stefan Fleischmann, netdev, stable, Pavan Chebbi, regressions, Joe Damato [-- Attachment #1: Type: text/plain, Size: 2442 bytes --] On Sun, Oct 4, 2026 at 2:11 PM Eric Dumazet <edumazet@kernel.org> wrote: > A more polished patch would be this one. > > bnxt_start_xmit() needs an audit I think :/ > > diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c > b/drivers/net/ethernet/broadcom/bnxt/bnxt.c > index > d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..48e544544754b04cff891bda5bc86cc370e0aa0b > 100644 > --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c > +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c > @@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff > *skb, struct net_device *dev) > struct netdev_queue *txq; > int i; > dma_addr_t mapping; > - unsigned int length, pad = 0; > + unsigned int length; > u32 len, free_size, vlan_tag_flags, cfa_action, flags; > struct bnxt_ptp_cfg *ptp = bp->ptp_cfg; > struct pci_dev *pdev = bp->pdev; > @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff > *skb, struct net_device *dev) > } > > normal_tx: > - if (length < BNXT_MIN_PKT_SIZE) { > - pad = BNXT_MIN_PKT_SIZE - length; > - if (skb_pad(skb, pad)) > - /* SKB already freed. */ > - goto tx_kick_pending; > - length = BNXT_MIN_PKT_SIZE; > + if (skb_padto(skb, BNXT_MIN_PKT_SIZE)) { > + /* SKB already freed. */ > + goto tx_kick_pending; > } > - > + length = skb->len; I think skb->len is not updated with the padded length here. So the HW will drop the packet seeing that the length is too short. There is another skb_put_padto() that might work better? > + len = skb_headlen(skb); > mapping = dma_map_single(&pdev->dev, skb->data, len, > DMA_TO_DEVICE); > > if (unlikely(dma_mapping_error(&pdev->dev, mapping))) > @@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff > *skb, struct net_device *dev) > txbd->tx_bd_len_flags_type = cpu_to_le32(flags); > } > > - flags &= ~TX_BD_LEN; > - txbd->tx_bd_len_flags_type = > - cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags | > - TX_BD_FLAGS_PACKET_END); > + txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END); > > netdev_tx_sent_queue(txq, skb->len); [-- Attachment #2: S/MIME Cryptographic Signature --] [-- Type: application/pkcs7-signature, Size: 5469 bytes --] ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-04 22:29 ` Michael Chan @ 2026-10-05 1:59 ` Eric Dumazet 2026-10-05 9:44 ` Fabian Grünbichler 0 siblings, 1 reply; 15+ messages in thread From: Eric Dumazet @ 2026-10-05 1:59 UTC (permalink / raw) To: Michael Chan Cc: Stefan Fleischmann, netdev, stable, Pavan Chebbi, regressions, Joe Damato On 10/5/26 00:29, Michael Chan wrote: > > I think skb->len is not updated with the padded length here. So the > HW will drop the packet seeing that the length is too short. There is > another skb_put_padto() that might work better? +1 Exactly, thanks! diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c index d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..7ea27e81e88c5ca82a453b449982b791e8acc831 100644 --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c @@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) struct netdev_queue *txq; int i; dma_addr_t mapping; - unsigned int length, pad = 0; + unsigned int length; u32 len, free_size, vlan_tag_flags, cfa_action, flags; struct bnxt_ptp_cfg *ptp = bp->ptp_cfg; struct pci_dev *pdev = bp->pdev; @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) } normal_tx: - if (length < BNXT_MIN_PKT_SIZE) { - pad = BNXT_MIN_PKT_SIZE - length; - if (skb_pad(skb, pad)) - /* SKB already freed. */ - goto tx_kick_pending; - length = BNXT_MIN_PKT_SIZE; + if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) { + /* SKB already freed. */ + goto tx_kick_pending; } - + length = skb->len; + len = skb_headlen(skb); mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE); if (unlikely(dma_mapping_error(&pdev->dev, mapping))) @@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev) txbd->tx_bd_len_flags_type = cpu_to_le32(flags); } - flags &= ~TX_BD_LEN; - txbd->tx_bd_len_flags_type = - cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags | - TX_BD_FLAGS_PACKET_END); + txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END); netdev_tx_sent_queue(txq, skb->len); ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-05 1:59 ` Eric Dumazet @ 2026-10-05 9:44 ` Fabian Grünbichler 2026-10-05 10:14 ` Eric Dumazet 0 siblings, 1 reply; 15+ messages in thread From: Fabian Grünbichler @ 2026-10-05 9:44 UTC (permalink / raw) To: Eric Dumazet, Michael Chan Cc: Joe Damato, netdev, Pavan Chebbi, regressions, Stefan Fleischmann, stable On October 5, 2026 3:59 am, Eric Dumazet wrote: > > > On 10/5/26 00:29, Michael Chan wrote: >> >> I think skb->len is not updated with the padded length here. So the >> HW will drop the packet seeing that the length is too short. There is >> another skb_put_padto() that might work better? > > +1 Exactly, thanks! FWIW, we suspect we have quite a few users running into this since importing upstream changes from 7.2.2-7.2.5. Will report back once we've provided those users with a kernel with the below diff. > diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c > b/drivers/net/ethernet/broadcom/bnxt/bnxt.c > index > d7728d0c5b6e63ee72de9dea54426bb4c8b7a9fc..7ea27e81e88c5ca82a453b449982b791e8acc831 > 100644 > --- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c > +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c > @@ -486,7 +486,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff > *skb, struct net_device *dev) > struct netdev_queue *txq; > int i; > dma_addr_t mapping; > - unsigned int length, pad = 0; > + unsigned int length; > u32 len, free_size, vlan_tag_flags, cfa_action, flags; > struct bnxt_ptp_cfg *ptp = bp->ptp_cfg; > struct pci_dev *pdev = bp->pdev; > @@ -672,14 +672,12 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff > *skb, struct net_device *dev) > } > > normal_tx: > - if (length < BNXT_MIN_PKT_SIZE) { > - pad = BNXT_MIN_PKT_SIZE - length; > - if (skb_pad(skb, pad)) > - /* SKB already freed. */ > - goto tx_kick_pending; > - length = BNXT_MIN_PKT_SIZE; > + if (skb_put_padto(skb, BNXT_MIN_PKT_SIZE)) { > + /* SKB already freed. */ > + goto tx_kick_pending; > } > - > + length = skb->len; > + len = skb_headlen(skb); > mapping = dma_map_single(&pdev->dev, skb->data, len, > DMA_TO_DEVICE); > > if (unlikely(dma_mapping_error(&pdev->dev, mapping))) > @@ -759,10 +757,7 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff > *skb, struct net_device *dev) > txbd->tx_bd_len_flags_type = cpu_to_le32(flags); > } > > - flags &= ~TX_BD_LEN; > - txbd->tx_bd_len_flags_type = > - cpu_to_le32(((len + pad) << TX_BD_LEN_SHIFT) | flags | > - TX_BD_FLAGS_PACKET_END); > + txbd->tx_bd_len_flags_type |= cpu_to_le32(TX_BD_FLAGS_PACKET_END); > > netdev_tx_sent_queue(txq, skb->len); > > > > > ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-05 9:44 ` Fabian Grünbichler @ 2026-10-05 10:14 ` Eric Dumazet 2026-10-05 12:54 ` Stefan Fleischmann 0 siblings, 1 reply; 15+ messages in thread From: Eric Dumazet @ 2026-10-05 10:14 UTC (permalink / raw) To: Fabian Grünbichler Cc: Michael Chan, Joe Damato, netdev, Pavan Chebbi, regressions, Stefan Fleischmann, stable Le lun. 5 oct. 2026 à 11:44, Fabian Grünbichler <f.gruenbichler@proxmox.com> a écrit : > > On October 5, 2026 3:59 am, Eric Dumazet wrote: > > > > > > On 10/5/26 00:29, Michael Chan wrote: > >> > >> I think skb->len is not updated with the padded length here. So the > >> HW will drop the packet seeing that the length is too short. There is > >> another skb_put_padto() that might work better? > > > > +1 Exactly, thanks! > > FWIW, we suspect we have quite a few users running into this since > importing upstream changes from 7.2.2-7.2.5. > > Will report back once we've provided those users with a kernel with the > below diff. Patch has been sent for review: https://lore.kernel.org/netdev/20261005023812.130639-1-edumazet@kernel.org/ ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-05 10:14 ` Eric Dumazet @ 2026-10-05 12:54 ` Stefan Fleischmann 2026-10-06 21:35 ` Sasha Levin 0 siblings, 1 reply; 15+ messages in thread From: Stefan Fleischmann @ 2026-10-05 12:54 UTC (permalink / raw) To: Eric Dumazet, Fabian Grünbichler Cc: Michael Chan, Joe Damato, netdev, Pavan Chebbi, regressions, stable On Mon, 5 Oct 2026 12:14:00 +0200 Eric Dumazet <edumazet@kernel.org> wrote: > Le lun. 5 oct. 2026 à 11:44, Fabian Grünbichler > <f.gruenbichler@proxmox.com> a écrit : > > > > On October 5, 2026 3:59 am, Eric Dumazet wrote: > > > > > > > > > On 10/5/26 00:29, Michael Chan wrote: > > >> > > >> I think skb->len is not updated with the padded length here. So > > >> the HW will drop the packet seeing that the length is too short. > > >> There is another skb_put_padto() that might work better? > > > > > > +1 Exactly, thanks! > > > > FWIW, we suspect we have quite a few users running into this since > > importing upstream changes from 7.2.2-7.2.5. > > > > Will report back once we've provided those users with a kernel with > > the below diff. > > Patch has been sent for review: > > https://lore.kernel.org/netdev/20261005023812.130639-1-edumazet@kernel.org/ I'm testing the patch on v6.18.55, and it looks good so far. 5 hours and counting without any issues. Best, Stefan ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-05 12:54 ` Stefan Fleischmann @ 2026-10-06 21:35 ` Sasha Levin 2026-10-07 10:33 ` Thorsten Leemhuis 0 siblings, 1 reply; 15+ messages in thread From: Sasha Levin @ 2026-10-06 21:35 UTC (permalink / raw) To: Eric Dumazet, Fabian Grünbichler Cc: Sasha Levin, Michael Chan, Joe Damato, netdev, Pavan Chebbi, regressions, stable, Stefan Fleischmann > I'm testing the patch on v6.18.55, and it looks good so far. 5 hours and > counting without any issues. Thanks for the bisect. The root cause is the bnxt_en small-packet padding and DMA length bug Eric Dumazet identified, not 447cbe95ebb9 ("vlan: fix skb_under_panic and races when toggling HW VLAN offload") itself, so we will not revert the vlan fix. Once Eric's "bnxt_en: fix DMA mapping length for padded small packets" lands in mainline, we will backport it to all affected stable trees. -- Thanks, Sasha ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-06 21:35 ` Sasha Levin @ 2026-10-07 10:33 ` Thorsten Leemhuis 2026-10-07 11:08 ` Eric Dumazet 0 siblings, 1 reply; 15+ messages in thread From: Thorsten Leemhuis @ 2026-10-07 10:33 UTC (permalink / raw) To: Sasha Levin, Eric Dumazet, Fabian Grünbichler Cc: Michael Chan, Joe Damato, netdev, Pavan Chebbi, regressions, stable, Stefan Fleischmann On 10/6/26 23:35, Sasha Levin wrote: >> I'm testing the patch on v6.18.55, and it looks good so far. 5 hours and >> counting without any issues. > > Thanks for the bisect. The root cause is the bnxt_en small-packet padding > and DMA length bug Eric Dumazet identified, not 447cbe95ebb9 ("vlan: fix > skb_under_panic and races when toggling HW VLAN offload") itself, so we > will not revert the vlan fix. Once Eric's "bnxt_en: fix DMA mapping length > for padded small packets" lands in mainline, we will backport it to all > affected stable trees. FWIW, better wait a little longer till a fix for another regression 447cbe95ebb9 exposed in another driver is fixed, too: https://lore.kernel.org/all/CAPa5EdCj3v17tB-SF2JNecq5Q8s1Pranr-XzAFWNpNTBPgPG7w@mail.gmail.com/ Ciao, Thorsten ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-07 10:33 ` Thorsten Leemhuis @ 2026-10-07 11:08 ` Eric Dumazet 0 siblings, 0 replies; 15+ messages in thread From: Eric Dumazet @ 2026-10-07 11:08 UTC (permalink / raw) To: Thorsten Leemhuis Cc: Sasha Levin, Fabian Grünbichler, Michael Chan, Joe Damato, netdev, Pavan Chebbi, regressions, stable, Stefan Fleischmann Le mer. 7 oct. 2026 à 12:35, Thorsten Leemhuis <regressions@leemhuis.info> a écrit : > > On 10/6/26 23:35, Sasha Levin wrote: > >> I'm testing the patch on v6.18.55, and it looks good so far. 5 hours and > >> counting without any issues. > > > > Thanks for the bisect. The root cause is the bnxt_en small-packet padding > > and DMA length bug Eric Dumazet identified, not 447cbe95ebb9 ("vlan: fix > > skb_under_panic and races when toggling HW VLAN offload") itself, so we > > will not revert the vlan fix. Once Eric's "bnxt_en: fix DMA mapping length > > for padded small packets" lands in mainline, we will backport it to all > > affected stable trees. > > FWIW, better wait a little longer till a fix for another regression > 447cbe95ebb9 exposed in another driver is fixed, too: > https://lore.kernel.org/all/CAPa5EdCj3v17tB-SF2JNecq5Q8s1Pranr-XzAFWNpNTBPgPG7w@mail.gmail.com/ > Different issue, and already fixed. https://patchwork.kernel.org/project/netdevbpf/patch/20261007073027.459868-1-edumazet@kernel.org/ This different bug fix will not fix the pre-existing bugs in bnxt_en. > Ciao, Thorsten > ^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en 2026-10-04 11:46 ` Eric Dumazet 2026-10-04 14:35 ` Stefan Fleischmann @ 2026-10-05 17:38 ` Joe Damato 1 sibling, 0 replies; 15+ messages in thread From: Joe Damato @ 2026-10-05 17:38 UTC (permalink / raw) To: Eric Dumazet Cc: Stefan Fleischmann, netdev, stable, Michael Chan, Pavan Chebbi, regressions On Sun, Oct 04, 2026 at 01:46:08PM +0200, Eric Dumazet wrote: > > > On 10/4/26 12:26, Stefan Fleischmann wrote: > > Hi, > > > > I have bisected a regression introduced in kernel 6.18.51 (and 7.2.3). > > The issue is still present in 6.18.55 and 7.2.9. Verified by reverting > > 1517d1996b52 (the equivalent of 447cbe95ebb9 in the 6.18 branch) on top > > of 6.18.51 and 6.18.55. > > > > The hardware is a Dell R640 server with Intel Xeon Silver 4210 and > > Broadcom BCM57412 NetXtreme-E NIC. > > > > The network interface is configured as an untagged interface and has a > > couple of tagged VLAN interfaces configured on top. All of these are > > used by LXC containers with Macvlan. > > > > The interface comes up fine after boot, but after some time (anywhere > > between 2 to 40 minutes) I see a DMA error in the kernel log (attached), > > and the driver ends up tearing down the interface. It can be brought up > > again after that, but it will hit a DMA error again after a while and > > the interface goes down again. > > > > I'm happy to provide more debug info if needed, or test patches. > > Hi Stefan, > > Thanks for the bisect and detailed report. > > Notice that 0xfc499000 is on an exact 4KB page boundary. This points to a > DMA read overrun where the Broadcom DMA engine reads past the end of a > buffer mapped in page 0xfc498xxx into the adjacent unmapped page 0xfc499000. > > Have you tried a recent net kernel ? > > Could you please test whether disabling TSO or hardware VLAN offload > prevents the crash? (maybe adding one option at a time) > > # ethtool -K eno1np0 tso off > > # ethtool -K eno1np0 tx-vlan-offload off > > Adding Joe to this thread, because some parts in > tso_dma_map_init() or tso_start() might have bugs. Thanks for adding me and debugging this. We've been hitting a similar read fault for a very long time I've had trouble tracking down (it's been reproducing since at least kernel 6.9, long before my SW USO changes). In the case I am tracking, it seems to reproduce (fairly rarely) and seems to be caused by a tcpv6 TSO packet with a read fault in the last page of a 32k frag. I have yet to find a reproducer for my issue, though :( I'll try this patch, but IIUC padding only applies to frames under 52 bytes, so I'm not sure if it's the same bug I have been searching for. ^ permalink raw reply [flat|nested] 15+ messages in thread
end of thread, other threads:[~2026-10-07 11:09 UTC | newest] Thread overview: 15+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-10-04 10:26 [REGRESSION] Commit 447cbe95ebb9 causes IOMMU DMA faults on macvlan/vlan with bnxt_en Stefan Fleischmann 2026-10-04 11:46 ` Eric Dumazet 2026-10-04 14:35 ` Stefan Fleischmann 2026-10-04 17:29 ` Stefan Fleischmann 2026-10-04 20:52 ` Eric Dumazet 2026-10-04 21:11 ` Eric Dumazet 2026-10-04 22:29 ` Michael Chan 2026-10-05 1:59 ` Eric Dumazet 2026-10-05 9:44 ` Fabian Grünbichler 2026-10-05 10:14 ` Eric Dumazet 2026-10-05 12:54 ` Stefan Fleischmann 2026-10-06 21:35 ` Sasha Levin 2026-10-07 10:33 ` Thorsten Leemhuis 2026-10-07 11:08 ` Eric Dumazet 2026-10-05 17:38 ` Joe Damato
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox