All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711
@ 2026-09-18 22:50 Md Rayhanul Islam
  2026-09-18 22:50 ` [PATCH 1/2] eal/linux: apply PCIe inbound DMA translation Md Rayhanul Islam
                   ` (3 more replies)
  0 siblings, 4 replies; 11+ messages in thread
From: Md Rayhanul Islam @ 2026-09-18 22:50 UTC (permalink / raw)
  To: dev
  Cc: bruce.richardson, Chuanyu Xue, Song Han, Md Rayhanul Islam,
	Md Rayhanul Islam

DPDK cannot transmit a single packet on a Raspberry Pi Compute Module 4
with an Intel I210.  Nothing reports an error: rte_eth_tx_burst() returns
the full count, the link is up at 1 Gbps, and testpmd in txonly mode
still shows TX-packets: 0.  Received data is DMA'd somewhere other than
the mbuf.  The same card works on x86, and the same board works through
the kernel's igb driver.

This has been reported twice and never explained: on Stack Overflow in
October 2023 with DPDK 23.07 [1], where TX-dropped equalled TX-total at
55 million with both vfio-noiommu and uio_pci_generic, and on dpdk-users
in January 2026 with DPDK 25.03 [2].  Two years apart, so this is the
platform, not a regression.

There are two independent causes, hence two patches.

First, the PCIe host bridge does not present memory to devices at CPU
physical addresses.  Its device tree says so, and DPDK does not look:

  $ hexdump -C /proc/device-tree/scb/pcie@7d500000/dma-ranges
  02000000 00000004 00000000  00000000 00000000  00000001 00000000

  PCI memory space | bus 0x4_0000_0000 | CPU 0x0 | size 4 GiB

A device reaching CPU physical address P must therefore be programmed
with P + 0x4_0000_0000.  In IOVA_PA mode every ring and mbuf address
DPDK hands the NIC falls outside the inbound window and is discarded
silently.  Patch 1 reads the translation from the host bridge and
applies it to IOVAs; that alone made the NIC transmit.

Second, the bus is not cache coherent -- no "dma-coherent" on the bridge
or any ancestor -- so the NIC read stale rings: zeroed descriptors with
DD set and null buffer addresses.  Patch 2 does the cache maintenance
the kernel DMA API would do.  Two details took the longest to find:

  - Cleaning to the point of unification (DC CVAU) was not enough.  Only
    a clean to the point of coherency made the device see CPU stores.

  - A clean writes back a whole 64-byte line, which holds four
    descriptors, so cleaning one erased DD bits the NIC had just set on
    its neighbours.  TX completion therefore comes from the hardware
    head register, and RX descriptors are refilled a whole line at a
    time, once every descriptor in that line has come back.

The cache maintenance sits in the driver.  Every driver on such a bus
needs the same thing, so an arch or EAL helper may be the better home; I
kept it local to e1000 for a first submission and am happy to move it.

Each patch documents its own half, since the failure gives nothing to
search for.

The device is bound to uio_pci_generic throughout: this SoC has no
IOMMU, so VFIO cannot be used to sidestep the address question.

Tested on a Compute Module 4 (BCM2711, 4 GB, Cortex-A72, 64-byte cache
lines) with an I210 (8086:1533 rev 03) on uio_pci_generic, Ubuntu 22.04
arm64, kernel 5.15, pcie_aspm=off.  With this series applied to current
main, testpmd txonly reaches 1.42 Mpps at 64-byte frames, which is 1 GbE
line rate, with 0 TX errors; without it, TX-packets stays 0.  The same
code based on v25.03 also passed a 300k-frame MAC loopback with no loss
and byte-exact payloads, and 390-run round-trip campaigns against a
second board with ICMP and UDP probes, 100 to 1500 byte packets, 1k
pkt/s to line rate, with no unexplained loss.

Not tested: any other board, SoC or NIC.  The offset is read from the
device tree rather than hardcoded, but I have only seen this platform.
Only e1000 was changed, so other drivers on a non-coherent bus still
read stale descriptors.

[1] https://stackoverflow.com/questions/77225289/
[2] https://mails.dpdk.org/archives/users/2026-January/008433.html

Md Rayhanul Islam (2):
  eal/linux: apply PCIe inbound DMA translation
  net/e1000: maintain caches on non-coherent DMA

 .mailmap                                  |   1 +
 doc/guides/linux_gsg/bcm2711_platform.rst |  71 ++++++
 doc/guides/linux_gsg/index.rst            |   1 +
 doc/guides/rel_notes/release_26_11.rst    |  16 ++
 drivers/net/intel/e1000/igb_rxtx.c        | 356 +++++++++++++++++++++++++++++-
 lib/eal/linux/eal_memory.c                | 265 +++++++++++++++++++++-
 6 files changed, 697 insertions(+), 13 deletions(-)
 create mode 100644 doc/guides/linux_gsg/bcm2711_platform.rst

-- 
2.34.1


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH 1/2] eal/linux: apply PCIe inbound DMA translation
  2026-09-18 22:50 [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Md Rayhanul Islam
@ 2026-09-18 22:50 ` Md Rayhanul Islam
  2026-09-18 22:50 ` [PATCH 2/2] net/e1000: maintain caches on non-coherent DMA Md Rayhanul Islam
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 11+ messages in thread
From: Md Rayhanul Islam @ 2026-09-18 22:50 UTC (permalink / raw)
  To: dev
  Cc: bruce.richardson, Chuanyu Xue, Song Han, Md Rayhanul Islam,
	Md Rayhanul Islam, Thomas Monjalon, Anatoly Burakov

On some platforms a device does not see memory at the CPU's physical
addresses: the PCIe host bridge translates inbound traffic.  An IOVA in
RTE_IOVA_PA mode is then an address the device cannot reach, so its
reads and writes fall outside the bridge's inbound window and are
dropped with no error reported.

On a Raspberry Pi Compute Module 4 (BCM2711) with an Intel I210, the
host bridge declares in its device tree "dma-ranges" that CPU physical
address 0 is reached at bus address 0x4_0000_0000.  Without that offset
the NIC reported link up but never completed a DMA: testpmd in txonly
mode left TX-packets at 0, and received data never reached the mbufs.

Read the translation from the PCI host bridge "dma-ranges" in the device
tree and add it to IOVAs in rte_mem_virt2iova() and the legacy memory
path.  Platforms that declare no translation, and those with no device
tree, keep an identity mapping.  DPDK_IOVA_PA_OFFSET overrides it.

Signed-off-by: Md Rayhanul Islam <r97yhan@gmail.com>
---
 .mailmap                                  |   1 +
 doc/guides/linux_gsg/bcm2711_platform.rst |  45 ++++
 doc/guides/linux_gsg/index.rst            |   1 +
 doc/guides/rel_notes/release_26_11.rst    |   8 +
 lib/eal/linux/eal_memory.c                | 265 +++++++++++++++++++++-
 5 files changed, 318 insertions(+), 2 deletions(-)
 create mode 100644 doc/guides/linux_gsg/bcm2711_platform.rst

diff --git a/.mailmap b/.mailmap
index 9e45cdce8f..9fd78db7e2 100644
--- a/.mailmap
+++ b/.mailmap
@@ -1032,6 +1032,7 @@ Marek Mical <marekx.mical@intel.com>
 Marek Zalfresso-jundzillo <marekx.zalfresso-jundzillo@intel.com>
 Maria Lingemark <maria.lingemark@ericsson.com>
 Mario Carrillo <mario.alfredo.c.arevalo@intel.com>
+Md Rayhanul Islam <r97yhan@gmail.com>
 Mário Kuka <kuka@cesnet.cz>
 Mariusz Drost <mariuszx.drost@intel.com>
 Mark Asselstine <mark.asselstine@windriver.com>
diff --git a/doc/guides/linux_gsg/bcm2711_platform.rst b/doc/guides/linux_gsg/bcm2711_platform.rst
new file mode 100644
index 0000000000..5a774607da
--- /dev/null
+++ b/doc/guides/linux_gsg/bcm2711_platform.rst
@@ -0,0 +1,45 @@
+.. SPDX-License-Identifier: BSD-3-Clause
+   Copyright(c) 2026 Md Rayhanul Islam
+
+Running DPDK on Broadcom BCM2711 platforms
+==========================================
+
+The Broadcom BCM2711, used on the Raspberry Pi 4 and Compute Module 4,
+has a platform property that must be taken into account before a PCIe
+device can DMA correctly.  It is handled automatically; this page
+describes what happens and how to check it.
+
+PCIe inbound address translation
+--------------------------------
+
+The PCIe host bridge does not present system memory to devices at the CPU
+physical addresses.  On a Compute Module 4 the host bridge node declares::
+
+   $ hexdump -C /proc/device-tree/scb/pcie@7d500000/dma-ranges
+   02000000 00000004 00000000  00000000 00000000  00000001 00000000
+
+which places CPU physical address 0 at PCIe bus address ``0x4_0000_0000``.
+
+In ``RTE_IOVA_PA`` mode, an IOVA must therefore be the CPU physical
+address plus that offset.  EAL reads the translation from the host
+bridge's ``dma-ranges`` property and applies it.  Platforms that declare
+no translation, and platforms without a device tree, are unaffected.
+
+A DPDK application logs the offset it found at startup::
+
+   EAL: PCIe bus addresses are offset by 0x400000000 from CPU physical
+   addresses (/proc/device-tree/scb/pcie@7d500000/dma-ranges); applying
+   it to IOVAs
+
+The value can be overridden with the ``DPDK_IOVA_PA_OFFSET`` environment
+variable, given in hexadecimal.  Without the translation, a device
+reports link up and counts packets in its own registers while never
+completing a DMA.
+
+Recommended settings
+--------------------
+
+* Add ``pcie_aspm=off`` to the kernel command line, so the PCIe link does
+  not enter a low-power state during a run.
+* Bind the device to ``uio_pci_generic``.  There is no IOMMU on this
+  SoC, so ``vfio-pci`` can only be used in unsafe no-IOMMU mode.
diff --git a/doc/guides/linux_gsg/index.rst b/doc/guides/linux_gsg/index.rst
index f739edd6ca..480909ec2c 100644
--- a/doc/guides/linux_gsg/index.rst
+++ b/doc/guides/linux_gsg/index.rst
@@ -20,3 +20,4 @@ Getting Started Guide for Linux
     enable_func
     nic_perf_intel_platform
     amd_platform
+    bcm2711_platform
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 4b3e5d995c..37452b2a3f 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,14 @@ New Features
      Also, make sure to start the actual text at the margin.
      =======================================================
 
+* **Added PCIe inbound DMA address translation on Linux.**
+
+  EAL now reads the PCIe host bridge "dma-ranges" property from the device
+  tree and applies the translation it declares to IOVAs, so devices on a
+  platform whose inbound window is not identity-mapped, such as the
+  Broadcom BCM2711 on Raspberry Pi 4 and Compute Module 4, can reach
+  system memory.
+
 
 Removed Items
 -------------
diff --git a/lib/eal/linux/eal_memory.c b/lib/eal/linux/eal_memory.c
index d9d505d865..75c6ba07c8 100644
--- a/lib/eal/linux/eal_memory.c
+++ b/lib/eal/linux/eal_memory.c
@@ -3,6 +3,7 @@
  * Copyright(c) 2013 6WIND S.A.
  */
 
+#include <dirent.h>
 #include <errno.h>
 #include <fcntl.h>
 #include <stdbool.h>
@@ -25,6 +26,7 @@
 #include <numaif.h>
 #endif
 
+#include <rte_byteorder.h>
 #include <rte_errno.h>
 #include <rte_log.h>
 #include <rte_memory.h>
@@ -83,6 +85,262 @@ uint64_t eal_get_baseaddr(void)
 #endif
 }
 
+/*
+ * On some platforms the address space seen by PCIe devices is not the CPU
+ * physical address space: the host bridge applies a fixed translation to
+ * inbound (device-to-memory) traffic.  There, an IOVA in RTE_IOVA_PA mode is
+ * *not* a CPU physical address -- a device DMAing to CPU physical address P
+ * must be programmed with P + offset, or its reads and writes fall outside the
+ * bridge's inbound window and are silently discarded.
+ *
+ * This was found on a Broadcom BCM2711 (Raspberry Pi Compute Module 4, 4GB),
+ * whose host bridge node pcie@7d500000 declares in "dma-ranges" that CPU
+ * physical 0 is reached by devices at PCIe address 0x4_0000_0000 (its outbound
+ * window, in "ranges", starts at PCIe address 0xc000_0000).  Without the
+ * translation every DPDK descriptor ring and mbuf address handed to the NIC is
+ * unreachable, and the observed symptom was a device that reported link up and
+ * counted packets in its own registers while never completing a single DMA.
+ *
+ * The translation is read from the "dma-ranges" property of the PCI host
+ * bridge in the device tree, so each platform gets the value it declares
+ * rather than one assumed here.  It is zero on identity-mapped platforms and
+ * absent where there is no device tree (x86), and everything below is then a
+ * no-op.
+ */
+#define DT_ROOT_PATH		"/proc/device-tree"
+#define IOVA_PA_OFFSET_ENV	"DPDK_IOVA_PA_OFFSET"
+#define IOVA_PA_OFFSET_UNSET	UINT64_MAX
+#define DT_MAX_SCAN_DEPTH	4
+
+/*
+ * Cached PCIe-bus-address-minus-CPU-physical-address translation.  Written at
+ * most once with a value that does not depend on who computes it, and it is a
+ * single aligned 64-bit store, so the benign race between threads racing to
+ * fill it in cannot produce a torn or inconsistent result.
+ */
+static uint64_t iova_pa_offset = IOVA_PA_OFFSET_UNSET;
+
+/* Read a device-tree property into buf, storing its length in *outlen. */
+static int
+dt_read_prop(const char *dir, const char *prop, void *buf, size_t buflen,
+		size_t *outlen)
+{
+	char path[PATH_MAX];
+	size_t n;
+	FILE *f;
+
+	if (snprintf(path, sizeof(path), "%s/%s", dir, prop) >= (int)sizeof(path))
+		return -1;
+	f = fopen(path, "rb");
+	if (f == NULL)
+		return -1;
+	n = fread(buf, 1, buflen, f);
+	if (ferror(f)) {
+		fclose(f);
+		return -1;
+	}
+	fclose(f);
+	*outlen = n;
+	return 0;
+}
+
+/* Read a single-cell device-tree property, e.g. #address-cells. */
+static int
+dt_read_u32(const char *dir, const char *prop, uint32_t *out)
+{
+	uint32_t val;
+	size_t len;
+
+	if (dt_read_prop(dir, prop, &val, sizeof(val), &len) < 0 ||
+			len != sizeof(val))
+		return -1;
+	*out = rte_be_to_cpu_32(val);
+	return 0;
+}
+
+/* Device-tree addresses are big-endian sequences of 32-bit cells. */
+static uint64_t
+dt_read_cells(const uint32_t *cells, uint32_t n)
+{
+	uint64_t val = 0;
+	uint32_t i;
+
+	for (i = 0; i < n; i++)
+		val = (val << 32) | rte_be_to_cpu_32(cells[i]);
+	return val;
+}
+
+/*
+ * Derive the inbound translation from the "dma-ranges" of one PCI host bridge.
+ * Each entry is <pci-address> <parent-address> <size>, where a PCI address is
+ * always 3 cells (phys.hi holds flags, phys.mid/phys.lo hold the address), the
+ * parent address is the parent bus' #address-cells, and the size is this node's
+ * #size-cells.
+ */
+static int
+dt_pci_dma_offset(const char *node, const char *parent, uint64_t *offset)
+{
+	uint32_t cells[256];
+	uint32_t parent_ac, size_c;
+	size_t len, ncells, per_entry, i;
+	uint64_t off = 0;
+	bool first = true;
+
+	if (dt_read_prop(node, "dma-ranges", cells, sizeof(cells), &len) < 0)
+		return -1;
+
+	/* An empty "dma-ranges" declares the bus to be identity mapped. */
+	if (len == 0) {
+		*offset = 0;
+		return 0;
+	}
+
+	if (dt_read_u32(parent, "#address-cells", &parent_ac) < 0 ||
+			dt_read_u32(node, "#size-cells", &size_c) < 0)
+		return -1;
+	/* More than 2 cells cannot be held in a uint64_t. */
+	if (parent_ac == 0 || parent_ac > 2 || size_c > 2)
+		return -1;
+
+	per_entry = 3 + parent_ac + size_c;
+	ncells = len / sizeof(uint32_t);
+	if (ncells == 0 || ncells % per_entry != 0)
+		return -1;
+
+	for (i = 0; i < ncells; i += per_entry) {
+		uint64_t pci_addr = dt_read_cells(&cells[i + 1], 2);
+		uint64_t cpu_addr = dt_read_cells(&cells[i + 3], parent_ac);
+		uint64_t entry_off = pci_addr - cpu_addr;
+
+		if (first) {
+			off = entry_off;
+			first = false;
+		} else if (entry_off != off) {
+			/*
+			 * The bridge translates different regions differently,
+			 * which a single offset cannot express.  Refuse to
+			 * guess rather than corrupt every IOVA.
+			 */
+			EAL_LOG(WARNING,
+				"%s: non-uniform dma-ranges, cannot derive an IOVA offset",
+				node);
+			return -1;
+		}
+	}
+
+	*offset = off;
+	return 0;
+}
+
+/*
+ * Walk the device tree looking for a PCI host bridge that describes a non-zero
+ * inbound translation, and report the first one found.  Systems with several
+ * host bridges translating differently would need per-device IOVAs, which the
+ * IOVA-as-PA model cannot express; we log the node we used so a mismatch is
+ * visible rather than silent.
+ */
+static int
+dt_scan_pci_dma_offset(const char *dir, const char *parent, int depth,
+		uint64_t *offset, char *node, size_t node_len)
+{
+	struct dirent *ent;
+	int ret = -1;
+	DIR *d;
+
+	if (depth > DT_MAX_SCAN_DEPTH)
+		return -1;
+
+	if (parent != NULL) {
+		char type[16];
+		size_t len;
+
+		if (dt_read_prop(dir, "device_type", type, sizeof(type) - 1,
+				&len) == 0) {
+			type[len] = '\0';
+			if (strcmp(type, "pci") == 0 &&
+					dt_pci_dma_offset(dir, parent,
+						offset) == 0 && *offset != 0) {
+				strlcpy(node, dir, node_len);
+				return 0;
+			}
+		}
+	}
+
+	d = opendir(dir);
+	if (d == NULL)
+		return -1;
+	while ((ent = readdir(d)) != NULL) {
+		char child[PATH_MAX];
+
+		/* Skip ".", "..", and the device tree's own dot-properties. */
+		if (ent->d_name[0] == '.')
+			continue;
+		/* procfs may not fill in d_type; opendir() filters non-dirs. */
+		if (ent->d_type != DT_DIR && ent->d_type != DT_UNKNOWN)
+			continue;
+		if (snprintf(child, sizeof(child), "%s/%s", dir,
+				ent->d_name) >= (int)sizeof(child))
+			continue;
+		if (dt_scan_pci_dma_offset(child, dir, depth + 1, offset,
+				node, node_len) == 0) {
+			ret = 0;
+			break;
+		}
+	}
+	closedir(d);
+	return ret;
+}
+
+/* Offset to add to a CPU physical address to obtain the address a PCIe device
+ * must use to reach it.  Zero on identity-mapped platforms.
+ */
+static uint64_t
+eal_iova_pa_offset(void)
+{
+	char node[PATH_MAX] = "";
+	uint64_t offset = 0;
+	const char *env;
+	char *end;
+
+	if (iova_pa_offset != IOVA_PA_OFFSET_UNSET)
+		return iova_pa_offset;
+
+	env = getenv(IOVA_PA_OFFSET_ENV);
+	if (env != NULL && env[0] != '\0') {
+		errno = 0;
+		offset = strtoull(env, &end, 0);
+		if (errno != 0 || *end != '\0') {
+			EAL_LOG(ERR, "Invalid %s value '%s', assuming no offset",
+				IOVA_PA_OFFSET_ENV, env);
+			offset = 0;
+		} else {
+			EAL_LOG(NOTICE,
+				"Using IOVA offset 0x%" PRIx64 " from %s",
+				offset, IOVA_PA_OFFSET_ENV);
+		}
+	} else if (dt_scan_pci_dma_offset(DT_ROOT_PATH, NULL, 0, &offset,
+			node, sizeof(node)) == 0) {
+		EAL_LOG(NOTICE,
+			"PCIe bus addresses are offset by 0x%" PRIx64
+			" from CPU physical addresses (%s/dma-ranges); applying it to IOVAs",
+			offset, node);
+	} else {
+		offset = 0;
+	}
+
+	iova_pa_offset = offset;
+	return offset;
+}
+
+/* Convert a CPU physical address into the IOVA a device must be given. */
+static rte_iova_t
+eal_pa_to_iova(phys_addr_t pa)
+{
+	if (pa == RTE_BAD_IOVA)
+		return RTE_BAD_IOVA;
+	return pa + eal_iova_pa_offset();
+}
+
 /*
  * Get physical address of any mapped virtual address in the current process.
  */
@@ -150,7 +408,7 @@ rte_mem_virt2iova(const void *virtaddr)
 {
 	if (rte_eal_iova_mode() == RTE_IOVA_VA)
 		return (uintptr_t)virtaddr;
-	return rte_mem_virt2phy(virtaddr);
+	return eal_pa_to_iova(rte_mem_virt2phy(virtaddr));
 }
 
 /*
@@ -805,7 +1063,10 @@ remap_segment(struct hugepage_file *hugepages, int seg_start, int seg_end)
 		ms->addr = addr;
 		ms->hugepage_sz = page_sz;
 		ms->len = memseg_len;
-		ms->iova = hfile->physaddr;
+		/* In IOVA-as-VA mode physaddr has been rewritten to the VA. */
+		ms->iova = rte_eal_iova_mode() == RTE_IOVA_VA ?
+				hfile->physaddr :
+				eal_pa_to_iova(hfile->physaddr);
 		ms->socket_id = hfile->socket_id;
 		ms->nchannel = rte_memory_get_nchannel();
 		ms->nrank = rte_memory_get_nrank();
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [PATCH 2/2] net/e1000: maintain caches on non-coherent DMA
  2026-09-18 22:50 [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Md Rayhanul Islam
  2026-09-18 22:50 ` [PATCH 1/2] eal/linux: apply PCIe inbound DMA translation Md Rayhanul Islam
@ 2026-09-18 22:50 ` Md Rayhanul Islam
  2026-09-19 21:05 ` [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Stephen Hemminger
  2026-10-02 20:09 ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Md Rayhanul Islam
  3 siblings, 0 replies; 11+ messages in thread
From: Md Rayhanul Islam @ 2026-09-18 22:50 UTC (permalink / raw)
  To: dev
  Cc: bruce.richardson, Chuanyu Xue, Song Han, Md Rayhanul Islam,
	Md Rayhanul Islam

When the PCIe bus is not cache coherent, the NIC can read stale memory.
The kernel's DMA API cleans the cache before handing a buffer to a
device; DPDK writes descriptors straight into hugepage memory and rings
the doorbell.

On a BCM2711 (Raspberry Pi Compute Module 4, Cortex-A72, 64-byte cache
lines) with an Intel I210, the NIC fetched the previous contents of the
rings -- zeroed descriptors with DD set and null buffer addresses -- so
transmits never completed and received data was DMA'd to address zero.
Cleaning to the point of unification was not enough; the device saw the
CPU's stores only after a clean to the point of coherency.

Cleaning writes back a whole 64-byte line, which holds four descriptors,
so cleaning one just written also restored stale status for its
neighbours and erased DD bits the NIC had set.

When the device is described by a device tree and no ancestor declares
"dma-coherent", on arm64:

 - clean the D-cache (DC CVAC) over descriptors and packet data before
   the tail register lets the NIC read them;
 - invalidate (DC CIVAC) descriptor status and received data before
   reading what the NIC wrote;
 - take TX completion from the hardware head register instead of the DD
   bits;
 - write RX buffer addresses back one whole cache line at a time, once
   every descriptor in that line has been returned and consumed.

This is a no-op on cache-coherent platforms and non-arm64 builds, and
DPDK_DMA_NONCOHERENT forces it either way.

Signed-off-by: Md Rayhanul Islam <r97yhan@gmail.com>
---
 doc/guides/linux_gsg/bcm2711_platform.rst |  30 +-
 doc/guides/rel_notes/release_26_11.rst    |   8 +
 drivers/net/intel/e1000/igb_rxtx.c        | 356 +++++++++++++++++++++-
 3 files changed, 381 insertions(+), 13 deletions(-)

diff --git a/doc/guides/linux_gsg/bcm2711_platform.rst b/doc/guides/linux_gsg/bcm2711_platform.rst
index 5a774607da..3be149e2f7 100644
--- a/doc/guides/linux_gsg/bcm2711_platform.rst
+++ b/doc/guides/linux_gsg/bcm2711_platform.rst
@@ -5,8 +5,8 @@ Running DPDK on Broadcom BCM2711 platforms
 ==========================================
 
 The Broadcom BCM2711, used on the Raspberry Pi 4 and Compute Module 4,
-has a platform property that must be taken into account before a PCIe
-device can DMA correctly.  It is handled automatically; this page
+needs two platform properties to be taken into account before a PCIe
+device can DMA correctly.  Both are handled automatically; this page
 describes what happens and how to check it.
 
 PCIe inbound address translation
@@ -36,6 +36,32 @@ variable, given in hexadecimal.  Without the translation, a device
 reports link up and counts packets in its own registers while never
 completing a DMA.
 
+Non-cache-coherent DMA
+----------------------
+
+The PCIe bus is not cache coherent: neither ``pcie@7d500000`` nor any of
+its ancestors declares ``dma-coherent``.  A device therefore does not see
+data the CPU has only written to its caches, and the CPU does not see
+data the device has written to memory.
+
+The ``e1000`` (igb) driver handles this when it detects a
+device-tree-described device with no ``dma-coherent`` ancestor, on arm64.
+It cleans the data cache over descriptors and packet data before the tail
+register is written, and invalidates it before reading what the device
+wrote.  Because a cache line covers four descriptors, TX completion is
+taken from the hardware head register rather than the DD bits, and RX
+descriptors are refilled one whole cache line at a time.
+
+This is logged at device probe::
+
+   E1000_INIT: 0000:01:00.0: PCIe DMA is not cache coherent,
+   cleaning 64-byte D-cache lines before each DMA
+
+It can be forced on or off with ``DPDK_DMA_NONCOHERENT=1`` or ``0``.
+
+Other drivers running on this platform need equivalent handling; only
+``e1000`` implements it today.
+
 Recommended settings
 --------------------
 
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 37452b2a3f..e292cc3714 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -63,6 +63,14 @@ New Features
   Broadcom BCM2711 on Raspberry Pi 4 and Compute Module 4, can reach
   system memory.
 
+* **Added non-coherent DMA support to the e1000 (igb) driver.**
+
+  The driver now performs the cache maintenance required on a PCIe bus
+  that is not cache coherent, taking TX completion from the hardware head
+  register and refilling RX descriptors a cache line at a time.  It is
+  enabled when the device is described by a device tree whose bus does not
+  declare "dma-coherent", on arm64.
+
 
 Removed Items
 -------------
diff --git a/drivers/net/intel/e1000/igb_rxtx.c b/drivers/net/intel/e1000/igb_rxtx.c
index 4fda5d57a0..ee81772a0e 100644
--- a/drivers/net/intel/e1000/igb_rxtx.c
+++ b/drivers/net/intel/e1000/igb_rxtx.c
@@ -4,6 +4,10 @@
 
 #include <sys/queue.h>
 
+#include <limits.h>
+#include <stdbool.h>
+#include <unistd.h>
+
 #include <stdio.h>
 #include <stdlib.h>
 #include <string.h>
@@ -42,6 +46,211 @@
 #include "base/e1000_api.h"
 #include "e1000_ethdev.h"
 
+/*
+ * Non-cache-coherent DMA support.
+ *
+ * When the PCIe host bridge is not cache coherent -- no "dma-coherent" on the
+ * device tree bus -- a bus-master read by the NIC can observe stale memory:
+ * data the CPU has just stored is not yet visible to the device.  The Linux
+ * DMA API hides this by cleaning the CPU data cache over a buffer before
+ * handing it to the device (arch_sync_dma_for_device), but DPDK writes
+ * descriptors and packet data straight into hugepage memory and rings the
+ * doorbell immediately.
+ *
+ * This was developed and tested on one such platform: a BCM2711 (Raspberry Pi
+ * Compute Module 4, 4GB, Cortex-A72, 64-byte D-cache lines) with an Intel I210
+ * on uio_pci_generic.  There the NIC fetched the previous contents of the
+ * rings -- zeroed descriptors with DD set and null buffer addresses -- so
+ * transmits never completed and received data was DMA'd to address zero.
+ * Cleaning only to the point of unification (DC CVAU) did not help; the NIC
+ * saw the CPU's stores only after a clean to the point of coherency.
+ *
+ * Mirror the kernel: clean the D-cache to the point of coherency (DC CVAC, as
+ * arm64 __dma_clean_area does) over every descriptor and packet buffer before
+ * the tail register tells the NIC it may read them, and discard cached copies
+ * (DC CIVAC, as __dma_inv_area does) of descriptor status and received data
+ * before reading what the NIC wrote.
+ *
+ * Cleaning a line writes the CPU's copy of the whole line back to RAM, so it
+ * must never cover memory the NIC may still be writing: with 4 descriptors
+ * per 64-byte line, flushing a freshly written descriptor also restored stale
+ * status for its neighbours and erased the DD bits the NIC had just set (the
+ * kernel avoids this by mapping rings non-cacheable).  Therefore, when DMA is
+ * non-coherent:
+ *  - TX completion is taken from the hardware head (TDH), not from DD bits;
+ *  - RX buffer addresses are written back one whole cache line at a time,
+ *    and only once every descriptor in that line has been handed back by the
+ *    NIC and consumed.
+ * All of this is a no-op on cache-coherent platforms and non-arm64 builds.
+ */
+#if defined(RTE_ARCH_ARM64)
+/* void igb_arm64_dcache_clean(uintptr_t start, uintptr_t end, uint64_t line) */
+void igb_arm64_dcache_clean(uintptr_t start, uintptr_t end, uint64_t line);
+/* void igb_arm64_dcache_inval(uintptr_t start, uintptr_t end, uint64_t line) */
+void igb_arm64_dcache_inval(uintptr_t start, uintptr_t end, uint64_t line);
+/* uint64_t igb_arm64_read_ctr(void) */
+uint64_t igb_arm64_read_ctr(void);
+__asm__(
+	"	.pushsection .text\n"
+	"	.p2align 2\n"
+	"	.globl igb_arm64_dcache_clean\n"
+	"	.hidden igb_arm64_dcache_clean\n"
+	"	.type igb_arm64_dcache_clean, %function\n"
+	"igb_arm64_dcache_clean:\n"
+	"	sub	x3, x2, #1\n"
+	"	bic	x0, x0, x3\n"
+	"1:	dc	cvac, x0\n"
+	"	add	x0, x0, x2\n"
+	"	cmp	x0, x1\n"
+	"	b.lo	1b\n"
+	"	ret\n"
+	"	.size igb_arm64_dcache_clean, .-igb_arm64_dcache_clean\n"
+	"	.globl igb_arm64_dcache_inval\n"
+	"	.hidden igb_arm64_dcache_inval\n"
+	"	.type igb_arm64_dcache_inval, %function\n"
+	"igb_arm64_dcache_inval:\n"
+	"	sub	x3, x2, #1\n"
+	"	bic	x0, x0, x3\n"
+	"1:	dc	civac, x0\n"
+	"	add	x0, x0, x2\n"
+	"	cmp	x0, x1\n"
+	"	b.lo	1b\n"
+	"	ret\n"
+	"	.size igb_arm64_dcache_inval, .-igb_arm64_dcache_inval\n"
+	"	.globl igb_arm64_read_ctr\n"
+	"	.hidden igb_arm64_read_ctr\n"
+	"	.type igb_arm64_read_ctr, %function\n"
+	"igb_arm64_read_ctr:\n"
+	"	mrs	x0, ctr_el0\n"
+	"	ret\n"
+	"	.size igb_arm64_read_ctr, .-igb_arm64_read_ctr\n"
+	"	.popsection\n");
+#endif
+
+#define IGB_DMA_NONCOHERENT_ENV "DPDK_DMA_NONCOHERENT"
+
+/* -1: not yet probed, 0: cache coherent, 1: needs explicit cache cleaning */
+static int igb_dma_noncoherent = -1;
+static uint64_t igb_dcache_line;
+
+/*
+ * Linux only instantiates "of_node" links for devices described by a device
+ * tree.  Walk the PCI device's sysfs ancestry (device -> bridges -> host
+ * controller -> SoC bus) the way of_dma_is_coherent() walks DT parents, and
+ * treat the device as non-coherent if it is DT-described and no ancestor
+ * declares "dma-coherent".  ACPI and x86 systems have no of_node and are left
+ * alone.
+ */
+static void
+igb_probe_dma_coherence(const char *pci_name)
+{
+	const char *env = getenv(IGB_DMA_NONCOHERENT_ENV);
+	int noncoherent = 0;
+
+	if (igb_dma_noncoherent >= 0)
+		return;
+
+	if (env != NULL && env[0] != '\0') {
+		noncoherent = atoi(env) != 0;
+	} else {
+#if defined(RTE_ARCH_ARM64)
+		char path[PATH_MAX], real[PATH_MAX], probe[PATH_MAX];
+		bool dt_described = false, coherent = false;
+		char *slash;
+
+		if (snprintf(path, sizeof(path), "/sys/bus/pci/devices/%s",
+				pci_name) < (int)sizeof(path) &&
+				realpath(path, real) != NULL) {
+			while ((slash = strrchr(real, '/')) != NULL &&
+					slash != real) {
+				if (snprintf(probe, sizeof(probe), "%s/of_node",
+						real) < (int)sizeof(probe) &&
+						access(probe, F_OK) == 0) {
+					dt_described = true;
+					if (snprintf(probe, sizeof(probe),
+							"%s/of_node/dma-coherent",
+							real) < (int)sizeof(probe) &&
+							access(probe, F_OK) == 0) {
+						coherent = true;
+						break;
+					}
+				}
+				*slash = '\0';
+			}
+		}
+		noncoherent = dt_described && !coherent;
+#else
+		RTE_SET_USED(pci_name);
+#endif
+	}
+
+#if defined(RTE_ARCH_ARM64)
+	if (noncoherent) {
+		/* CTR_EL0.DCacheLine is log2 of the line size in words. */
+		igb_dcache_line = 4ULL << ((igb_arm64_read_ctr() >> 16) & 0xf);
+		PMD_INIT_LOG(NOTICE,
+			"%s: PCIe DMA is not cache coherent, cleaning %" PRIu64
+			"-byte D-cache lines before each DMA",
+			pci_name, igb_dcache_line);
+	}
+#else
+	if (noncoherent)
+		PMD_INIT_LOG(WARNING,
+			"%s: %s set, but cache cleaning is only implemented for arm64",
+			pci_name, IGB_DMA_NONCOHERENT_ENV);
+	noncoherent = 0;
+#endif
+	igb_dma_noncoherent = noncoherent;
+}
+
+/* Make [addr, addr + len) visible to a bus-master read by the NIC. */
+static inline void
+igb_dma_sync_for_device(const volatile void *addr, size_t len)
+{
+#if defined(RTE_ARCH_ARM64)
+	if (unlikely(igb_dma_noncoherent > 0) && len != 0)
+		igb_arm64_dcache_clean((uintptr_t)addr, (uintptr_t)addr + len,
+				igb_dcache_line);
+#else
+	RTE_SET_USED(addr);
+	RTE_SET_USED(len);
+#endif
+}
+
+/*
+ * Make memory the NIC has written visible to the CPU: discard any cached copy
+ * (DC CIVAC, as arm64 __dma_inv_area does) so the next read comes from RAM.
+ */
+static inline void
+igb_dma_sync_for_cpu(const volatile void *addr, size_t len)
+{
+#if defined(RTE_ARCH_ARM64)
+	if (unlikely(igb_dma_noncoherent > 0) && len != 0)
+		igb_arm64_dcache_inval((uintptr_t)addr, (uintptr_t)addr + len,
+				igb_dcache_line);
+#else
+	RTE_SET_USED(addr);
+	RTE_SET_USED(len);
+#endif
+}
+
+/* Sync descriptors [from, to) of a ring of nb_desc entries, handling wrap. */
+static inline void
+igb_dma_sync_ring(const volatile void *ring, size_t desc_size, uint16_t nb_desc,
+		uint16_t from, uint16_t to)
+{
+	if (likely(igb_dma_noncoherent <= 0) || from == to)
+		return;
+	if (from < to) {
+		igb_dma_sync_for_device(RTE_PTR_ADD(ring, from * desc_size),
+				(to - from) * desc_size);
+	} else {
+		igb_dma_sync_for_device(RTE_PTR_ADD(ring, from * desc_size),
+				(nb_desc - from) * desc_size);
+		igb_dma_sync_for_device(ring, to * desc_size);
+	}
+}
+
 #ifdef RTE_LIBRTE_IEEE1588
 #define IGB_TX_IEEE1588_TMST RTE_MBUF_F_TX_IEEE1588_TMST
 #else
@@ -98,6 +307,11 @@ struct igb_rx_queue {
 	struct rte_mbuf *pkt_last_seg;  /**< Last segment of current packet. */
 	uint16_t            nb_rx_desc; /**< number of RX descriptors. */
 	uint16_t            rx_tail;    /**< current value of RDT register. */
+	/**
+	 * Non-coherent DMA only: first consumed descriptor whose refilled
+	 * buffer address has not been written back to the ring yet.
+	 */
+	uint16_t            rx_unwritten;
 	uint16_t            nb_rx_hold; /**< number of held free RX desc. */
 	uint16_t            rx_free_thresh; /**< max free RX desc to hold. */
 	uint16_t            queue_id;   /**< RX queue index. */
@@ -167,6 +381,7 @@ struct igb_tx_queue {
 	uint64_t               tx_ring_phys_addr; /**< TX ring DMA address. */
 	struct igb_tx_entry    *sw_ring; /**< virtual address of SW ring. */
 	volatile uint32_t      *tdt_reg_addr; /**< Address of TDT register. */
+	volatile uint32_t      *tdh_reg_addr; /**< Address of TDH register. */
 	uint32_t               txd_type;      /**< Device-specific TXD type */
 	uint16_t               nb_tx_desc;    /**< number of TX descriptors. */
 	uint16_t               tx_tail; /**< Current value of TDT register. */
@@ -384,6 +599,90 @@ tx_desc_vlan_flags_to_cmdtype(uint64_t ol_flags)
 	return cmdtype;
 }
 
+/*
+ * Non-coherent DMA: has the NIC completed descriptor desc?  The NIC owns the
+ * circular range [TDH, tx_tail); everything else has been processed.  *head
+ * caches one TDH read per burst (-1 = not read yet) and is refreshed once
+ * before reporting a descriptor as still busy.
+ */
+static inline bool
+igb_tx_desc_done(struct igb_tx_queue *txq, uint16_t desc, int32_t *head)
+{
+	uint16_t n = txq->nb_tx_desc;
+	int pass;
+
+	if (likely(igb_dma_noncoherent <= 0))
+		return (txq->tx_ring[desc].wb.status &
+			rte_cpu_to_le_32(E1000_TXD_STAT_DD)) != 0;
+
+	for (pass = 0; pass < 2; pass++) {
+		uint16_t owned, dist;
+
+		if (*head < 0 || pass > 0)
+			*head = rte_le_to_cpu_32(E1000_PCI_REG(txq->tdh_reg_addr)) % n;
+		owned = (uint16_t)((txq->tx_tail + n - *head) % n);
+		dist = (uint16_t)((desc + n - *head) % n);
+		if (dist >= owned)
+			return true;
+	}
+	return false;
+}
+
+/* Descriptors covered by one D-cache line, or 0 if lines don't tile rings. */
+static inline uint16_t
+igb_rx_desc_per_line(struct igb_rx_queue *rxq)
+{
+	uint16_t per_line = igb_dcache_line / sizeof(*rxq->rx_ring);
+
+	if (per_line <= 1 || rxq->nb_rx_desc % per_line != 0 ||
+			((uintptr_t)rxq->rx_ring & (igb_dcache_line - 1)) != 0)
+		return 1;
+	return per_line;
+}
+
+/* Write refilled buffer addresses for descriptors [from, to) and flush them. */
+static inline void
+igb_rx_write_range(struct igb_rx_queue *rxq, uint16_t from, uint16_t to)
+{
+	uint16_t i;
+
+	for (i = from; i < to; i++) {
+		volatile union e1000_adv_rx_desc *rxd = &rxq->rx_ring[i];
+
+		rxd->read.hdr_addr = 0;
+		rxd->read.pkt_addr = rte_cpu_to_le_64(
+			rte_mbuf_data_iova_default(rxq->sw_ring[i].mbuf));
+	}
+	igb_dma_sync_for_device(&rxq->rx_ring[from],
+			(to - from) * sizeof(*rxq->rx_ring));
+}
+
+/*
+ * Non-coherent DMA: write back every whole cache line of descriptors that
+ * lies entirely before the software head rx_id, i.e. that the NIC has
+ * finished with.  Returns the first descriptor still unwritten, which bounds
+ * how far RDT may advance.
+ */
+static inline uint16_t
+igb_rx_flush_refills(struct igb_rx_queue *rxq, uint16_t rx_id)
+{
+	uint16_t per_line = igb_rx_desc_per_line(rxq);
+	uint16_t start = rxq->rx_unwritten;
+	uint16_t end = rx_id - rx_id % per_line;
+
+	if (end == start)
+		return start;
+	if (end < start) {
+		/* Wrapped: finish the ring, then the aligned part from 0. */
+		igb_rx_write_range(rxq, start, rxq->nb_rx_desc);
+		start = 0;
+	}
+	if (end > start)
+		igb_rx_write_range(rxq, start, end);
+	rxq->rx_unwritten = end;
+	return end;
+}
+
 uint16_t
 eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 	       uint16_t nb_pkts)
@@ -402,6 +701,7 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 	uint16_t slen;
 	uint64_t ol_flags;
 	uint16_t tx_end;
+	int32_t hw_head = -1;
 	uint16_t tx_id;
 	uint16_t tx_last;
 	uint16_t nb_tx;
@@ -508,7 +808,7 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 		/*
 		 * Check that this descriptor is free.
 		 */
-		if (! (txr[tx_end].wb.status & E1000_TXD_STAT_DD)) {
+		if (!igb_tx_desc_done(txq, tx_end, &hw_head)) {
 			if (nb_tx == 0)
 				return 0;
 			goto end_of_tx;
@@ -595,6 +895,8 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 			 */
 			slen = (uint16_t) m_seg->data_len;
 			buf_dma_addr = rte_mbuf_data_iova(m_seg);
+			igb_dma_sync_for_device(rte_pktmbuf_mtod(m_seg, void *),
+					slen);
 			txd->read.buffer_addr =
 				rte_cpu_to_le_64(buf_dma_addr);
 			txd->read.cmd_type_len =
@@ -617,6 +919,9 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
  end_of_tx:
 	rte_wmb();
 
+	igb_dma_sync_ring(txq->tx_ring, sizeof(*txq->tx_ring), txq->nb_tx_desc,
+			txq->tx_tail, tx_id);
+
 	/*
 	 * Set the Transmit Descriptor Tail (TDT).
 	 */
@@ -854,6 +1159,7 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		 * using invalid descriptor fields when read from rxd.
 		 */
 		rxdp = &rx_ring[rx_id];
+		igb_dma_sync_for_cpu(rxdp, sizeof(*rxdp));
 		staterr = rxdp->wb.upper.status_error;
 		if (! (staterr & rte_cpu_to_le_32(E1000_RXD_STAT_DD)))
 			break;
@@ -923,8 +1229,11 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		rxe->mbuf = nmb;
 		dma_addr =
 			rte_cpu_to_le_64(rte_mbuf_data_iova_default(nmb));
-		rxdp->read.hdr_addr = 0;
-		rxdp->read.pkt_addr = dma_addr;
+		/* Non-coherent DMA defers this to igb_rx_flush_refills(). */
+		if (likely(igb_dma_noncoherent <= 0)) {
+			rxdp->read.hdr_addr = 0;
+			rxdp->read.pkt_addr = dma_addr;
+		}
 
 		/*
 		 * Initialize the returned mbuf.
@@ -942,6 +1251,8 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		pkt_len = (uint16_t) (rte_le_to_cpu_16(rxd.wb.upper.length) -
 				      rxq->crc_len);
 		rxm->data_off = RTE_PKTMBUF_HEADROOM;
+		igb_dma_sync_for_cpu((char *)rxm->buf_addr + rxm->data_off,
+				rte_le_to_cpu_16(rxd.wb.upper.length));
 		rte_packet_prefetch((char *)rxm->buf_addr + rxm->data_off);
 		rxm->nb_segs = 1;
 		rxm->next = NULL;
@@ -993,6 +1304,8 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 			   (unsigned) rxq->port_id, (unsigned) rxq->queue_id,
 			   (unsigned) rx_id, (unsigned) nb_hold,
 			   (unsigned) nb_rx);
+		if (unlikely(igb_dma_noncoherent > 0))
+			rx_id = igb_rx_flush_refills(rxq, rx_id);
 		rx_id = (uint16_t) ((rx_id == 0) ?
 				     (rxq->nb_rx_desc - 1) : (rx_id - 1));
 		E1000_PCI_REG_WRITE(rxq->rdt_reg_addr, rx_id);
@@ -1049,6 +1362,7 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		 * using invalid descriptor fields when read from rxd.
 		 */
 		rxdp = &rx_ring[rx_id];
+		igb_dma_sync_for_cpu(rxdp, sizeof(*rxdp));
 		staterr = rxdp->wb.upper.status_error;
 		if (! (staterr & rte_cpu_to_le_32(E1000_RXD_STAT_DD)))
 			break;
@@ -1117,8 +1431,11 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		rxm = rxe->mbuf;
 		rxe->mbuf = nmb;
 		dma = rte_cpu_to_le_64(rte_mbuf_data_iova_default(nmb));
-		rxdp->read.pkt_addr = dma;
-		rxdp->read.hdr_addr = 0;
+		/* Non-coherent DMA defers this to igb_rx_flush_refills(). */
+		if (likely(igb_dma_noncoherent <= 0)) {
+			rxdp->read.pkt_addr = dma;
+			rxdp->read.hdr_addr = 0;
+		}
 
 		/*
 		 * Set data length & data buffer address of mbuf.
@@ -1126,6 +1443,8 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		data_len = rte_le_to_cpu_16(rxd.wb.upper.length);
 		rxm->data_len = data_len;
 		rxm->data_off = RTE_PKTMBUF_HEADROOM;
+		igb_dma_sync_for_cpu((char *)rxm->buf_addr + rxm->data_off,
+				data_len);
 
 		/*
 		 * If this is the first buffer of the received packet,
@@ -1255,6 +1574,8 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 			   (unsigned) rxq->port_id, (unsigned) rxq->queue_id,
 			   (unsigned) rx_id, (unsigned) nb_hold,
 			   (unsigned) nb_rx);
+		if (unlikely(igb_dma_noncoherent > 0))
+			rx_id = igb_rx_flush_refills(rxq, rx_id);
 		rx_id = (uint16_t) ((rx_id == 0) ?
 				     (rxq->nb_rx_desc - 1) : (rx_id - 1));
 		E1000_PCI_REG_WRITE(rxq->rdt_reg_addr, rx_id);
@@ -1308,18 +1629,17 @@ static int
 igb_tx_done_cleanup(struct igb_tx_queue *txq, uint32_t free_cnt)
 {
 	struct igb_tx_entry *sw_ring;
-	volatile union e1000_adv_tx_desc *txr;
 	uint16_t tx_first; /* First segment analyzed. */
 	uint16_t tx_id;    /* Current segment being processed. */
 	uint16_t tx_last;  /* Last segment in the current packet. */
 	uint16_t tx_next;  /* First segment of the next packet. */
+	int32_t hw_head = -1;
 	int count = 0;
 
 	if (!txq)
 		return -ENODEV;
 
 	sw_ring = txq->sw_ring;
-	txr = txq->tx_ring;
 
 	/* tx_tail is the last sent packet on the sw_ring. Goto the end
 	 * of that packet (the last segment in the packet chain) and
@@ -1345,8 +1665,7 @@ igb_tx_done_cleanup(struct igb_tx_queue *txq, uint32_t free_cnt)
 		tx_last = sw_ring[tx_id].last_id;
 
 		if (sw_ring[tx_last].mbuf) {
-			if (txr[tx_last].wb.status &
-			    E1000_TXD_STAT_DD) {
+			if (igb_tx_desc_done(txq, tx_last, &hw_head)) {
 				/* Increment the number of packets
 				 * freed.
 				 */
@@ -1510,6 +1829,7 @@ eth_igb_tx_queue_setup(struct rte_eth_dev *dev,
 	offloads = tx_conf->offloads | dev->data->dev_conf.txmode.offloads;
 
 	hw = E1000_DEV_PRIVATE_TO_HW(dev->data->dev_private);
+	igb_probe_dma_coherence(dev->device->name);
 
 	/*
 	 * Validate number of transmit descriptors.
@@ -1575,6 +1895,7 @@ eth_igb_tx_queue_setup(struct rte_eth_dev *dev,
 	txq->port_id = dev->data->port_id;
 
 	txq->tdt_reg_addr = E1000_PCI_REG_ADDR(hw, E1000_TDT(txq->reg_idx));
+	txq->tdh_reg_addr = E1000_PCI_REG_ADDR(hw, E1000_TDH(txq->reg_idx));
 	txq->tx_ring_phys_addr = tz->iova;
 
 	txq->tx_ring = (union e1000_adv_tx_desc *) tz->addr;
@@ -1642,6 +1963,7 @@ igb_reset_rx_queue(struct igb_rx_queue *rxq)
 	}
 
 	rxq->rx_tail = 0;
+	rxq->rx_unwritten = 0;
 	rxq->pkt_first_seg = NULL;
 	rxq->pkt_last_seg = NULL;
 }
@@ -1710,6 +2032,7 @@ eth_igb_rx_queue_setup(struct rte_eth_dev *dev,
 	offloads = rx_conf->offloads | dev->data->dev_conf.rxmode.offloads;
 
 	hw = E1000_DEV_PRIVATE_TO_HW(dev->data->dev_private);
+	igb_probe_dma_coherence(dev->device->name);
 
 	/*
 	 * Validate number of receive descriptors.
@@ -1801,7 +2124,8 @@ eth_igb_rx_queue_count(void *rx_queue)
 	rxdp = &(rxq->rx_ring[rxq->rx_tail]);
 
 	while ((desc < rxq->nb_rx_desc) &&
-		(rxdp->wb.upper.status_error & E1000_RXD_STAT_DD)) {
+		(igb_dma_sync_for_cpu(rxdp, sizeof(*rxdp)),
+		 rxdp->wb.upper.status_error & E1000_RXD_STAT_DD)) {
 		desc += IGB_RXQ_SCAN_INTERVAL;
 		rxdp += IGB_RXQ_SCAN_INTERVAL;
 		if (rxq->rx_tail + desc >= rxq->nb_rx_desc)
@@ -1830,6 +2154,7 @@ eth_igb_rx_descriptor_status(void *rx_queue, uint16_t offset)
 		desc -= rxq->nb_rx_desc;
 
 	status = &rxq->rx_ring[desc].wb.upper.status_error;
+	igb_dma_sync_for_cpu(status, sizeof(*status));
 	if (*status & rte_cpu_to_le_32(E1000_RXD_STAT_DD))
 		return RTE_ETH_RX_DESC_DONE;
 
@@ -1851,8 +2176,15 @@ eth_igb_tx_descriptor_status(void *tx_queue, uint16_t offset)
 		desc -= txq->nb_tx_desc;
 
 	status = &txq->tx_ring[desc].wb.status;
-	if (*status & rte_cpu_to_le_32(E1000_TXD_STAT_DD))
+	if (unlikely(igb_dma_noncoherent > 0)) {
+		int32_t hw_head = -1;
+
+		RTE_SET_USED(status);
+		if (igb_tx_desc_done(txq, desc, &hw_head))
+			return RTE_ETH_TX_DESC_DONE;
+	} else if (*status & rte_cpu_to_le_32(E1000_TXD_STAT_DD)) {
 		return RTE_ETH_TX_DESC_DONE;
+	}
 
 	return RTE_ETH_TX_DESC_FULL;
 }
@@ -2276,6 +2608,8 @@ igb_alloc_rx_queue_mbufs(struct igb_rx_queue *rxq)
 		rxd->read.pkt_addr = dma_addr;
 		rxe[i].mbuf = mbuf;
 	}
+	igb_dma_sync_for_device(rxq->rx_ring,
+			rxq->nb_rx_desc * sizeof(*rxq->rx_ring));
 
 	return 0;
 }
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* Re: [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711
  2026-09-18 22:50 [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Md Rayhanul Islam
  2026-09-18 22:50 ` [PATCH 1/2] eal/linux: apply PCIe inbound DMA translation Md Rayhanul Islam
  2026-09-18 22:50 ` [PATCH 2/2] net/e1000: maintain caches on non-coherent DMA Md Rayhanul Islam
@ 2026-09-19 21:05 ` Stephen Hemminger
  2026-09-21  8:56   ` Bruce Richardson
  2026-10-02 20:09 ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Md Rayhanul Islam
  3 siblings, 1 reply; 11+ messages in thread
From: Stephen Hemminger @ 2026-09-19 21:05 UTC (permalink / raw)
  To: Md Rayhanul Islam
  Cc: dev, bruce.richardson, Chuanyu Xue, Song Han, Md Rayhanul Islam

On Fri, 18 Sep 2026 18:50:42 -0400
Md Rayhanul Islam <r97yhan@gmail.com> wrote:

> DPDK cannot transmit a single packet on a Raspberry Pi Compute Module 4
> with an Intel I210.  Nothing reports an error: rte_eth_tx_burst() returns
> the full count, the link is up at 1 Gbps, and testpmd in txonly mode
> still shows TX-packets: 0.  Received data is DMA'd somewhere other than
> the mbuf.  The same card works on x86, and the same board works through
> the kernel's igb driver.
> 
> This has been reported twice and never explained: on Stack Overflow in
> October 2023 with DPDK 23.07 [1], where TX-dropped equalled TX-total at
> 55 million with both vfio-noiommu and uio_pci_generic, and on dpdk-users
> in January 2026 with DPDK 25.03 [2].  Two years apart, so this is the
> platform, not a regression.
> 
> There are two independent causes, hence two patches.
> 
> First, the PCIe host bridge does not present memory to devices at CPU
> physical addresses.  Its device tree says so, and DPDK does not look:
> 
>   $ hexdump -C /proc/device-tree/scb/pcie@7d500000/dma-ranges
>   02000000 00000004 00000000  00000000 00000000  00000001 00000000
> 
>   PCI memory space | bus 0x4_0000_0000 | CPU 0x0 | size 4 GiB
> 
> A device reaching CPU physical address P must therefore be programmed
> with P + 0x4_0000_0000.  In IOVA_PA mode every ring and mbuf address
> DPDK hands the NIC falls outside the inbound window and is discarded
> silently.  Patch 1 reads the translation from the host bridge and
> applies it to IOVAs; that alone made the NIC transmit.
> 
> Second, the bus is not cache coherent -- no "dma-coherent" on the bridge
> or any ancestor -- so the NIC read stale rings: zeroed descriptors with
> DD set and null buffer addresses.  Patch 2 does the cache maintenance
> the kernel DMA API would do.  Two details took the longest to find:
> 
>   - Cleaning to the point of unification (DC CVAU) was not enough.  Only
>     a clean to the point of coherency made the device see CPU stores.
> 
>   - A clean writes back a whole 64-byte line, which holds four
>     descriptors, so cleaning one erased DD bits the NIC had just set on
>     its neighbours.  TX completion therefore comes from the hardware
>     head register, and RX descriptors are refilled a whole line at a
>     time, once every descriptor in that line has come back.
> 
> The cache maintenance sits in the driver.  Every driver on such a bus
> needs the same thing, so an arch or EAL helper may be the better home; I
> kept it local to e1000 for a first submission and am happy to move it.
> 
> Each patch documents its own half, since the failure gives nothing to
> search for.
> 
> The device is bound to uio_pci_generic throughout: this SoC has no
> IOMMU, so VFIO cannot be used to sidestep the address question.
> 
> Tested on a Compute Module 4 (BCM2711, 4 GB, Cortex-A72, 64-byte cache
> lines) with an I210 (8086:1533 rev 03) on uio_pci_generic, Ubuntu 22.04
> arm64, kernel 5.15, pcie_aspm=off.  With this series applied to current
> main, testpmd txonly reaches 1.42 Mpps at 64-byte frames, which is 1 GbE
> line rate, with 0 TX errors; without it, TX-packets stays 0.  The same
> code based on v25.03 also passed a 300k-frame MAC loopback with no loss
> and byte-exact payloads, and 390-run round-trip campaigns against a
> second board with ICMP and UDP probes, 100 to 1500 byte packets, 1k
> pkt/s to line rate, with no unexplained loss.
> 
> Not tested: any other board, SoC or NIC.  The offset is read from the
> device tree rather than hardcoded, but I have only seen this platform.
> Only e1000 was changed, so other drivers on a non-coherent bus still
> read stale descriptors.
> 
> [1] https://stackoverflow.com/questions/77225289/
> [2] https://mails.dpdk.org/archives/users/2026-January/008433.html
> 
> Md Rayhanul Islam (2):
>   eal/linux: apply PCIe inbound DMA translation
>   net/e1000: maintain caches on non-coherent DMA

Short observations:
  1. Too much AI generated slop, extra docs, comments on everything.
     New code should look like the surrounding code.
     This looks like AI wasn't quite sure and left lots of docs for future self.
  2. Handling DMA coherence should not be in the driver.
     It should be done in EAL.
  3. EAL helpers should be similar to helpers used in other OS (Linux and FreeBSD)
     when handling DMA issues.

AI observations:


Patch 1

- The offset is a property of one host bridge but is applied to every
  IOVA in the process, and it comes from whichever PCI node readdir()
  finds first under /proc/device-tree. BCM2712 already has several
  PCIe controllers with different dma-ranges. The offset has to come
  from the bridge the probed device sits behind. That is bus driver
  work: walk /sys/bus/pci/devices/<bdf> up to the host bridge of_node
  and read dma-ranges there, the same walk patch 2 does for
  dma-coherent. Bus scan runs before memory init and already feeds
  IOVA mode selection, so the offset can be delivered the same way.

- Lazy scan inside rte_mem_virt2iova() with a racy static is the
  wrong lifecycle. Compute once at init and keep it in shared
  mem_config so secondaries see the same value.

- No environment variables. If an override is needed it is an EAL
  option.

Patch 2

- The asm has no dsb after the dc loops, and DPDK's arm64 barriers
  are all dmb, which do not order cache maintenance. The clean can
  still be in flight when the tail register is written.

- EL0 cannot execute dc ivac, only civac, so every buffer the device
  may write must be cleaned before it is handed over. TX data and
  descriptors are cleaned, RX buffers on refill are not. Dirty lines
  from the application's previous use of the mbuf get written back
  over received data. txonly/rxonly/loopback never modify packets in
  place so the tests would not catch it; macswap would.

- TX completion from TDH: check the I210 datasheet. TDH advances when
  the descriptor is fetched, not when the data DMA completes, and the
  kernel igb driver never uses it for cleanup. Freeing the mbuf on TDH
  can corrupt frames still in flight.

- Coherence state is a process-global static set in queue_setup. Two
  ports on different buses get one answer, and a secondary process
  never runs queue_setup so it treats the device as coherent. Needs to
  be per-device, set by the bus, in shared device data.

- Every descriptor on every platform now tests igb_dma_noncoherent.
  Use separate burst functions selected at setup and leave the
  coherent path alone.

- em_rxtx.c in the same PMD is untouched.

The right next step is an RFC for the EAL sync API and the bus hook,
with igb as the first user, rather than a v2 of these patches.




^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711
  2026-09-19 21:05 ` [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Stephen Hemminger
@ 2026-09-21  8:56   ` Bruce Richardson
  2026-09-21 15:42     ` Stephen Hemminger
  0 siblings, 1 reply; 11+ messages in thread
From: Bruce Richardson @ 2026-09-21  8:56 UTC (permalink / raw)
  To: Stephen Hemminger
  Cc: Md Rayhanul Islam, dev, Chuanyu Xue, Song Han, Md Rayhanul Islam

On Sat, Sep 19, 2026 at 02:05:49PM -0700, Stephen Hemminger wrote:
> On Fri, 18 Sep 2026 18:50:42 -0400
> Md Rayhanul Islam <r97yhan@gmail.com> wrote:
> 
> > DPDK cannot transmit a single packet on a Raspberry Pi Compute Module 4
> > with an Intel I210.  Nothing reports an error: rte_eth_tx_burst() returns
> > the full count, the link is up at 1 Gbps, and testpmd in txonly mode
> > still shows TX-packets: 0.  Received data is DMA'd somewhere other than
> > the mbuf.  The same card works on x86, and the same board works through
> > the kernel's igb driver.
> > 
> > This has been reported twice and never explained: on Stack Overflow in
> > October 2023 with DPDK 23.07 [1], where TX-dropped equalled TX-total at
> > 55 million with both vfio-noiommu and uio_pci_generic, and on dpdk-users
> > in January 2026 with DPDK 25.03 [2].  Two years apart, so this is the
> > platform, not a regression.
> > 
> > There are two independent causes, hence two patches.
> > 
> > First, the PCIe host bridge does not present memory to devices at CPU
> > physical addresses.  Its device tree says so, and DPDK does not look:
> > 
> >   $ hexdump -C /proc/device-tree/scb/pcie@7d500000/dma-ranges
> >   02000000 00000004 00000000  00000000 00000000  00000001 00000000
> > 
> >   PCI memory space | bus 0x4_0000_0000 | CPU 0x0 | size 4 GiB
> > 
> > A device reaching CPU physical address P must therefore be programmed
> > with P + 0x4_0000_0000.  In IOVA_PA mode every ring and mbuf address
> > DPDK hands the NIC falls outside the inbound window and is discarded
> > silently.  Patch 1 reads the translation from the host bridge and
> > applies it to IOVAs; that alone made the NIC transmit.
> > 
> > Second, the bus is not cache coherent -- no "dma-coherent" on the bridge
> > or any ancestor -- so the NIC read stale rings: zeroed descriptors with
> > DD set and null buffer addresses.  Patch 2 does the cache maintenance
> > the kernel DMA API would do.  Two details took the longest to find:
> > 
> >   - Cleaning to the point of unification (DC CVAU) was not enough.  Only
> >     a clean to the point of coherency made the device see CPU stores.
> > 
> >   - A clean writes back a whole 64-byte line, which holds four
> >     descriptors, so cleaning one erased DD bits the NIC had just set on
> >     its neighbours.  TX completion therefore comes from the hardware
> >     head register, and RX descriptors are refilled a whole line at a
> >     time, once every descriptor in that line has come back.
> > 
> > The cache maintenance sits in the driver.  Every driver on such a bus
> > needs the same thing, so an arch or EAL helper may be the better home; I
> > kept it local to e1000 for a first submission and am happy to move it.
> > 
> > Each patch documents its own half, since the failure gives nothing to
> > search for.
> > 
> > The device is bound to uio_pci_generic throughout: this SoC has no
> > IOMMU, so VFIO cannot be used to sidestep the address question.
> > 
> > Tested on a Compute Module 4 (BCM2711, 4 GB, Cortex-A72, 64-byte cache
> > lines) with an I210 (8086:1533 rev 03) on uio_pci_generic, Ubuntu 22.04
> > arm64, kernel 5.15, pcie_aspm=off.  With this series applied to current
> > main, testpmd txonly reaches 1.42 Mpps at 64-byte frames, which is 1 GbE
> > line rate, with 0 TX errors; without it, TX-packets stays 0.  The same
> > code based on v25.03 also passed a 300k-frame MAC loopback with no loss
> > and byte-exact payloads, and 390-run round-trip campaigns against a
> > second board with ICMP and UDP probes, 100 to 1500 byte packets, 1k
> > pkt/s to line rate, with no unexplained loss.
> > 
> > Not tested: any other board, SoC or NIC.  The offset is read from the
> > device tree rather than hardcoded, but I have only seen this platform.
> > Only e1000 was changed, so other drivers on a non-coherent bus still
> > read stale descriptors.
> > 
> > [1] https://stackoverflow.com/questions/77225289/
> > [2] https://mails.dpdk.org/archives/users/2026-January/008433.html
> > 
> > Md Rayhanul Islam (2):
> >   eal/linux: apply PCIe inbound DMA translation
> >   net/e1000: maintain caches on non-coherent DMA
> 
> Short observations:
>   1. Too much AI generated slop, extra docs, comments on everything.
>      New code should look like the surrounding code.
>      This looks like AI wasn't quite sure and left lots of docs for future self.

I really don't view this as a bad thing. A little too much detail in the
comments is better for the future than too little. Given these patches
touch areas outside the usual behaviour we expect from DPDK platforms,
having extra comments is definitely useful.


/Bruce

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711
  2026-09-21  8:56   ` Bruce Richardson
@ 2026-09-21 15:42     ` Stephen Hemminger
  2026-09-24 19:22       ` Md Rayhanul Islam
  0 siblings, 1 reply; 11+ messages in thread
From: Stephen Hemminger @ 2026-09-21 15:42 UTC (permalink / raw)
  To: Bruce Richardson
  Cc: Md Rayhanul Islam, dev, Chuanyu Xue, Song Han, Md Rayhanul Islam

On Mon, 21 Sep 2026 09:56:17 +0100
Bruce Richardson <bruce.richardson@intel.com> wrote:

> > 
> > Short observations:
> >   1. Too much AI generated slop, extra docs, comments on everything.
> >      New code should look like the surrounding code.
> >      This looks like AI wasn't quite sure and left lots of docs for future self.  
> 
> I really don't view this as a bad thing. A little too much detail in the
> comments is better for the future than too little. Given these patches
> touch areas outside the usual behaviour we expect from DPDK platforms,
> having extra comments is definitely useful.
> 
> 
> /Bruce

There is a fine line between a couple of sentences and three paragraphs

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711
  2026-09-21 15:42     ` Stephen Hemminger
@ 2026-09-24 19:22       ` Md Rayhanul Islam
  0 siblings, 0 replies; 11+ messages in thread
From: Md Rayhanul Islam @ 2026-09-24 19:22 UTC (permalink / raw)
  To: Stephen Hemminger
  Cc: Bruce Richardson, dev, Chuanyu Xue, Song Han, Md Rayhanul Islam

[-- Attachment #1: Type: text/plain, Size: 1304 bytes --]

Hello,

Thank you for the detailed review. I am reworking the series as an RFC for
the EAL DMA API and PCI bus hook, with igb as the first user.

DMA properties now come from each device’s host bridge, and cache handling
is in EAL. I will also revise RX buffer reuse, TX completion, and
descriptor ownership, and add separate burst functions for non-coherent DMA.

I will also remove the extra guide and shorten the comments.

Best regards,
Rayhan

On Mon, Sep 21, 2026 at 11:42 AM Stephen Hemminger <
stephen@networkplumber.org> wrote:

> On Mon, 21 Sep 2026 09:56:17 +0100
> Bruce Richardson <bruce.richardson@intel.com> wrote:
>
> > >
> > > Short observations:
> > >   1. Too much AI generated slop, extra docs, comments on everything.
> > >      New code should look like the surrounding code.
> > >      This looks like AI wasn't quite sure and left lots of docs for
> future self.
> >
> > I really don't view this as a bad thing. A little too much detail in the
> > comments is better for the future than too little. Given these patches
> > touch areas outside the usual behaviour we expect from DPDK platforms,
> > having extra comments is definitely useful.
> >
> >
> > /Bruce
>
> There is a fine line between a couple of sentences and three paragraphs
>

[-- Attachment #2: Type: text/html, Size: 1915 bytes --]

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [RFC PATCH 0/2] per-device DMA properties, with igb as the first user
  2026-09-18 22:50 [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Md Rayhanul Islam
                   ` (2 preceding siblings ...)
  2026-09-19 21:05 ` [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Stephen Hemminger
@ 2026-10-02 20:09 ` Md Rayhanul Islam
  2026-10-02 20:09   ` [RFC PATCH 1/2] eal/pci: add per-device DMA translation and synchronization Md Rayhanul Islam
                     ` (2 more replies)
  3 siblings, 3 replies; 11+ messages in thread
From: Md Rayhanul Islam @ 2026-10-02 20:09 UTC (permalink / raw)
  To: dev

This RFC reworks the BCM2711 DMA fix around a per-device DMA context and
an EAL sync API, after the feedback on the first posting [1].

On Raspberry Pi 4 / Compute Module 4 the PCIe host bridge translates DMA
addresses and is not cache coherent, so a device such as the Intel I210
cannot DMA at all without both being handled.

What changed since the first posting:

* The translation and the coherency are read per device from its own
  bridge, not kept as one process-wide offset.  EAL no longer rewrites
  IOVAs; the driver translates where it programs the hardware.
* A driver opts in with RTE_PCI_DRV_DMA_NONCOHERENT, and the bus refuses
  such a device to any driver that has not.  That replaces the refusal
  written into em by hand.
* The cache maintenance is an experimental EAL interface with explicit
  directions, and the loops end in DSB SY.
* Completion comes from the Done bits, not the head register.
* Both environment variables are gone.

Tested between two Compute Module 4 boards with Intel I210s, both running
this series: testpmd txonly holds 1.42 Mpps with 64-byte frames, and a
512 MB UDP transfer with DPDK at both ends arrives byte-identical.  Built
with GCC and with clang 14 -Werror.

[1] https://inbox.dpdk.org/dev/20260918225045.177125-1-r97yhan@gmail.com/

Md Rayhanul Islam (2):
  eal/pci: add per-device DMA translation and synchronization
  net/e1000: support non-coherent DMA

 .mailmap                               |   1 +
 doc/guides/rel_notes/release_26_11.rst |  20 ++
 drivers/bus/pci/bus_pci_driver.h       |   2 +
 drivers/bus/pci/linux/pci.c            |   6 +
 drivers/bus/pci/linux/pci_dma_ranges.c | 168 ++++++++++
 drivers/bus/pci/linux/pci_init.h       |   2 +
 drivers/bus/pci/meson.build            |   1 +
 drivers/bus/pci/pci_common.c           |  17 +
 drivers/bus/pci/private.h              |   2 +
 drivers/bus/pci/rte_bus_pci.h          |  52 ++++
 drivers/net/intel/e1000/e1000_ethdev.h |  19 ++
 drivers/net/intel/e1000/igb_ethdev.c   |   9 +-
 drivers/net/intel/e1000/igb_rxtx.c     | 409 +++++++++++++++++++++++--
 lib/eal/include/meson.build            |   1 +
 lib/eal/include/rte_mem_sync.h         | 133 ++++++++
 15 files changed, 815 insertions(+), 27 deletions(-)
 create mode 100644 drivers/bus/pci/linux/pci_dma_ranges.c
 create mode 100644 lib/eal/include/rte_mem_sync.h

-- 
2.34.1


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [RFC PATCH 1/2] eal/pci: add per-device DMA translation and synchronization
  2026-10-02 20:09 ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Md Rayhanul Islam
@ 2026-10-02 20:09   ` Md Rayhanul Islam
  2026-10-02 20:09   ` [RFC PATCH 2/2] net/e1000: support non-coherent DMA Md Rayhanul Islam
  2026-10-03 16:51   ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Stephen Hemminger
  2 siblings, 0 replies; 11+ messages in thread
From: Md Rayhanul Islam @ 2026-10-02 20:09 UTC (permalink / raw)
  To: dev

Some PCIe platforms, such as BCM2711, translate DMA addresses and are not
cache coherent.  Both belong to the bridge above a device, so one offset
applied to the whole process is wrong on a board with two controllers.

Read them per device.  The PCI bus walks the device's sysfs ancestry to
its bridge, and a driver copies the result with rte_pci_get_dma_info().
Addresses are translated where they are programmed into the hardware,
with rte_pci_dma_iova(), so EAL IOVAs are untouched.  Properties this
parser cannot express mark that one device unusable rather than failing
the scan.

A driver must set RTE_PCI_DRV_DMA_NONCOHERENT before the bus hands it
such a device.  Only bridges in the compatible table are read, so other
platforms keep the behavior they have today.

Add rte_mem_sync_for_device() and rte_mem_sync_for_cpu(), named after
dma_sync_single_for_{device,cpu}(), with explicit transfer directions,
and rte_mem_dcache_line_size().  On arm64 they are DC CVAC and DC CIVAC
loops closed with DSB SY, since the DPDK barriers are DMB and do not
order cache maintenance.  EL0 cannot execute DC IVAC, so the CPU-side
helper cleans as well as invalidates.

Signed-off-by: Md Rayhanul Islam <r97yhan@gmail.com>
---
 .mailmap                               |   1 +
 doc/guides/rel_notes/release_26_11.rst |  15 +++
 drivers/bus/pci/bus_pci_driver.h       |   2 +
 drivers/bus/pci/linux/pci.c            |   6 +
 drivers/bus/pci/linux/pci_dma_ranges.c | 168 +++++++++++++++++++++++++
 drivers/bus/pci/linux/pci_init.h       |   2 +
 drivers/bus/pci/meson.build            |   1 +
 drivers/bus/pci/pci_common.c           |  17 +++
 drivers/bus/pci/private.h              |   2 +
 drivers/bus/pci/rte_bus_pci.h          |  52 ++++++++
 lib/eal/include/meson.build            |   1 +
 lib/eal/include/rte_mem_sync.h         | 133 ++++++++++++++++++++
 12 files changed, 400 insertions(+)
 create mode 100644 drivers/bus/pci/linux/pci_dma_ranges.c
 create mode 100644 lib/eal/include/rte_mem_sync.h

diff --git a/.mailmap b/.mailmap
index 9e45cdce8f..9fd78db7e2 100644
--- a/.mailmap
+++ b/.mailmap
@@ -1032,6 +1032,7 @@ Marek Mical <marekx.mical@intel.com>
 Marek Zalfresso-jundzillo <marekx.zalfresso-jundzillo@intel.com>
 Maria Lingemark <maria.lingemark@ericsson.com>
 Mario Carrillo <mario.alfredo.c.arevalo@intel.com>
+Md Rayhanul Islam <r97yhan@gmail.com>
 Mário Kuka <kuka@cesnet.cz>
 Mariusz Drost <mariuszx.drost@intel.com>
 Mark Asselstine <mark.asselstine@windriver.com>
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 4b3e5d995c..a752c94bbf 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,21 @@ New Features
      Also, make sure to start the actual text at the margin.
      =======================================================
 
+* **Added per-device DMA properties to the PCI bus.**
+
+  The Linux PCI bus reads the address translation and the cache coherency
+  of the host bridge above a device from the device tree, and exposes them
+  through ``rte_pci_get_dma_info()`` and ``rte_pci_dma_iova()``.  A driver
+  must set ``RTE_PCI_DRV_DMA_NONCOHERENT`` before the bus gives it a device
+  that needs either.  Devices behind a translating, non-coherent bridge,
+  such as those on the Broadcom BCM2711, can now be used.
+
+* **Added cache maintenance helpers for non-coherent DMA.**
+
+  ``rte_mem_sync_for_device()`` and ``rte_mem_sync_for_cpu()`` hand a buffer
+  between the CPU and a device whose DMA is not coherent with the caches,
+  with an explicit transfer direction.
+
 
 Removed Items
 -------------
diff --git a/drivers/bus/pci/bus_pci_driver.h b/drivers/bus/pci/bus_pci_driver.h
index c04ebddf59..216e4beb20 100644
--- a/drivers/bus/pci/bus_pci_driver.h
+++ b/drivers/bus/pci/bus_pci_driver.h
@@ -140,6 +140,8 @@ struct rte_pci_driver {
 #define RTE_PCI_DRV_KEEP_MAPPED_RES 0x0020
 /** Device driver needs IOVA as VA and cannot work with IOVA as PA */
 #define RTE_PCI_DRV_NEED_IOVA_AS_VA 0x0040
+/** Driver translates DMA addresses and maintains caches itself. */
+#define RTE_PCI_DRV_DMA_NONCOHERENT 0x0080
 
 /**
  * Register a PCI driver.
diff --git a/drivers/bus/pci/linux/pci.c b/drivers/bus/pci/linux/pci.c
index 9aae0a5d14..55c989c55d 100644
--- a/drivers/bus/pci/linux/pci.c
+++ b/drivers/bus/pci/linux/pci.c
@@ -220,6 +220,8 @@ pci_scan_one(const char *dirname, const struct rte_pci_addr *addr)
 	dev = &pdev->device;
 	dev->addr = *addr;
 
+	pci_dt_read_dma_info(dev, dirname);
+
 	/* get vendor id */
 	snprintf(filename, sizeof(filename), "%s/vendor", dirname);
 	if (eal_parse_sysfs_value(filename, &tmp) < 0) {
@@ -335,6 +337,7 @@ pci_scan_one(const char *dirname, const struct rte_pci_addr *addr)
 				rte_bus_insert_device(&rte_pci_bus, &dev2->device, &dev->device);
 			} else { /* already registered */
 				if (!rte_dev_is_probed(&dev2->device)) {
+					RTE_PCI_DEVICE_INTERNAL(dev2)->dma = pdev->dma;
 					dev2->kdrv = dev->kdrv;
 					dev2->max_vfs = dev->max_vfs;
 					dev2->id = dev->id;
@@ -591,6 +594,9 @@ pci_device_iova_mode(const struct rte_pci_driver *pdrv,
 {
 	enum rte_iova_mode iova_mode = RTE_IOVA_DC;
 
+	if (RTE_PCI_DEVICE_INTERNAL_CONST(pdev)->dma.size != 0)
+		return RTE_IOVA_PA;
+
 	switch (pdev->kdrv) {
 	case RTE_PCI_KDRV_VFIO: {
 		static int is_vfio_noiommu_enabled = -1;
diff --git a/drivers/bus/pci/linux/pci_dma_ranges.c b/drivers/bus/pci/linux/pci_dma_ranges.c
new file mode 100644
index 0000000000..da5edbd5cc
--- /dev/null
+++ b/drivers/bus/pci/linux/pci_dma_ranges.c
@@ -0,0 +1,168 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Md Rayhanul Islam
+ */
+
+#include <errno.h>
+#include <limits.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+
+#include <rte_byteorder.h>
+#include <rte_common.h>
+#include <rte_string_fns.h>
+
+#include "pci_init.h"
+#include "private.h"
+
+static int
+dt_read(const char *node, const char *name, void *buf, size_t size)
+{
+	char path[PATH_MAX];
+	size_t len;
+	int ret;
+	FILE *f;
+
+	if (snprintf(path, sizeof(path), "%s/%s", node, name) >= (int)sizeof(path))
+		return -ENAMETOOLONG;
+	f = fopen(path, "rb");
+	if (f == NULL)
+		return -errno;
+	len = fread(buf, 1, size, f);
+	ret = ferror(f) ? -EIO : (int)len;
+	if (ret >= 0 && fgetc(f) != EOF)
+		ret = -E2BIG;
+	fclose(f);
+	return ret;
+}
+
+static bool
+dt_is_known_bridge(const char *node)
+{
+	static const char * const compatible[] = { "brcm,bcm2711-pcie" };
+	char buf[256];
+	size_t pos, n, i;
+	int len;
+
+	len = dt_read(node, "compatible", buf, sizeof(buf));
+	if (len <= 0 || buf[len - 1] != '\0')
+		return false;
+	for (pos = 0; pos < (size_t)len; pos += n + 1) {
+		n = strlen(buf + pos);
+		for (i = 0; i < RTE_DIM(compatible); i++)
+			if (strcmp(buf + pos, compatible[i]) == 0)
+				return true;
+	}
+	return false;
+}
+
+static int
+dt_cells(const char *node, const char *name)
+{
+	rte_be32_t value;
+
+	if (dt_read(node, name, &value, sizeof(value)) != sizeof(value))
+		return -EINVAL;
+	return rte_be_to_cpu_32(value);
+}
+
+static uint64_t
+dt_address(const rte_be32_t *cells, int n)
+{
+	uint64_t value = 0;
+
+	while (n-- > 0)
+		value = (value << 32) | rte_be_to_cpu_32(*cells++);
+	return value;
+}
+
+/* One <pci-address> <parent-address> <size> entry in "dma-ranges". */
+static int
+dt_dma_window(const char *node, struct rte_pci_dma_info *info)
+{
+	char parent[PATH_MAX], *slash;
+	rte_be32_t cells[7];
+	int ac, len;
+
+	strlcpy(parent, node, sizeof(parent));
+	slash = strrchr(parent, '/');
+	if (slash == NULL || slash == parent)
+		return -EINVAL;
+	*slash = '\0';
+
+	ac = dt_cells(parent, "#address-cells");
+	if ((ac != 1 && ac != 2) || dt_cells(node, "#address-cells") != 3 ||
+			dt_cells(node, "#size-cells") != 2)
+		return -ENOTSUP;
+
+	len = dt_read(node, "dma-ranges", cells, sizeof(cells));
+	if (len < 0)
+		return len;
+	if (len != (3 + ac + 2) * (int)sizeof(cells[0]) ||
+			(rte_be_to_cpu_32(cells[0]) & 0x03000000) != 0x02000000)
+		return -ENOTSUP;
+
+	info->bus_base = dt_address(cells + 1, 2);
+	info->cpu_base = dt_address(cells + 3, ac);
+	info->size = dt_address(cells + 3 + ac, 2);
+	if (info->size == 0 || info->size > UINT64_MAX - info->cpu_base ||
+			info->size > UINT64_MAX - info->bus_base)
+		return -EINVAL;
+	return 0;
+}
+
+/* The nearest ancestor that declares coherency decides it. */
+static bool
+dt_is_coherent(const char *node)
+{
+	char path[PATH_MAX], value;
+	char *slash;
+
+	strlcpy(path, node, sizeof(path));
+	for (;;) {
+		if (dt_read(path, "dma-coherent", &value, sizeof(value)) >= 0)
+			return true;
+		if (dt_read(path, "dma-noncoherent", &value, sizeof(value)) >= 0)
+			return false;
+		slash = strrchr(path, '/');
+		if (slash == NULL || slash == path)
+			return false;
+		*slash = '\0';
+	}
+}
+
+/*
+ * Unparseable properties mark only this device unusable, so the scan of the
+ * other devices is not affected.
+ */
+void
+pci_dt_read_dma_info(struct rte_pci_device *dev, const char *dirname)
+{
+	struct rte_pci_dma_info *info = &RTE_PCI_DEVICE_INTERNAL(dev)->dma;
+	char node[PATH_MAX], path[PATH_MAX], of_node[PATH_MAX];
+	char *slash;
+	int ret;
+
+	memset(info, 0, sizeof(*info));
+	if (realpath(dirname, node) == NULL)
+		return;
+
+	while ((slash = strrchr(node, '/')) != NULL && slash != node) {
+		if (snprintf(path, sizeof(path), "%s/of_node", node) <
+				(int)sizeof(path) &&
+				realpath(path, of_node) != NULL &&
+				dt_is_known_bridge(of_node)) {
+			ret = dt_dma_window(of_node, info);
+			if (ret < 0) {
+				PCI_LOG(ERR, "%s: cannot use the DMA properties of %s",
+					dev->name, of_node);
+				info->unusable = true;
+				return;
+			}
+			info->noncoherent = !dt_is_coherent(of_node);
+			return;
+		}
+		*slash = '\0';
+	}
+}
diff --git a/drivers/bus/pci/linux/pci_init.h b/drivers/bus/pci/linux/pci_init.h
index 6949dd57d9..fcbf5cdfa2 100644
--- a/drivers/bus/pci/linux/pci_init.h
+++ b/drivers/bus/pci/linux/pci_init.h
@@ -73,4 +73,6 @@ int pci_vfio_unmap_resource(struct rte_pci_device *dev);
 
 int pci_vfio_is_enabled(void);
 
+void pci_dt_read_dma_info(struct rte_pci_device *dev, const char *dirname);
+
 #endif /* EAL_PCI_INIT_H_ */
diff --git a/drivers/bus/pci/meson.build b/drivers/bus/pci/meson.build
index fede114dc7..9e9b27ef4d 100644
--- a/drivers/bus/pci/meson.build
+++ b/drivers/bus/pci/meson.build
@@ -10,6 +10,7 @@ if is_linux
     sources += files(
             'pci_common_uio.c',
             'linux/pci.c',
+            'linux/pci_dma_ranges.c',
             'linux/pci_uio.c',
             'linux/pci_vfio.c',
     )
diff --git a/drivers/bus/pci/pci_common.c b/drivers/bus/pci/pci_common.c
index dc8db80d3b..14efb46d23 100644
--- a/drivers/bus/pci/pci_common.c
+++ b/drivers/bus/pci/pci_common.c
@@ -142,6 +142,14 @@ pci_unmap_resource(void *requested_addr, size_t size)
 		PCI_LOG(DEBUG, "  PCI memory unmapped at %p", requested_addr);
 }
 
+RTE_EXPORT_INTERNAL_SYMBOL(rte_pci_get_dma_info)
+void
+rte_pci_get_dma_info(const struct rte_pci_device *dev,
+		struct rte_pci_dma_info *info)
+{
+	*info = RTE_PCI_DEVICE_INTERNAL_CONST(dev)->dma;
+}
+
 static bool
 pci_bus_match(const struct rte_driver *drv, const struct rte_device *dev)
 {
@@ -186,6 +194,7 @@ pci_probe_device(struct rte_driver *drv, struct rte_device *dev)
 	struct rte_pci_device *pci_dev = RTE_BUS_DEVICE(dev, *pci_dev);
 	struct rte_pci_driver *pci_drv = RTE_BUS_DRIVER(drv, *pci_drv);
 	struct rte_pci_addr *loc = &pci_dev->addr;
+	const struct rte_pci_dma_info *dma;
 	bool already_probed;
 	int ret;
 
@@ -196,6 +205,14 @@ pci_probe_device(struct rte_driver *drv, struct rte_device *dev)
 	if (pci_dev->device.numa_node < 0 && rte_socket_count() > 1)
 		PCI_LOG(INFO, "Device %s is not NUMA-aware", pci_dev->name);
 
+	dma = &RTE_PCI_DEVICE_INTERNAL(pci_dev)->dma;
+	if (dma->unusable || ((dma->noncoherent || dma->size != 0) &&
+			!(pci_drv->drv_flags & RTE_PCI_DRV_DMA_NONCOHERENT))) {
+		PCI_LOG(ERR, "%s: %s does not handle this device's DMA properties",
+			pci_dev->name, pci_drv->driver.name);
+		return -ENOTSUP;
+	}
+
 	already_probed = (pci_dev->intr_handle != NULL);
 	if (already_probed && !(pci_drv->drv_flags & RTE_PCI_DRV_PROBE_AGAIN)) {
 		PCI_LOG(DEBUG, "Device %s is already probed", pci_dev->device.name);
diff --git a/drivers/bus/pci/private.h b/drivers/bus/pci/private.h
index 8103c32881..5a0286dbc0 100644
--- a/drivers/bus/pci/private.h
+++ b/drivers/bus/pci/private.h
@@ -13,6 +13,7 @@
 #include <rte_log.h>
 #include <rte_os_shim.h>
 #include <rte_pci.h>
+#include <rte_bus_pci.h>
 
 extern int pci_bus_logtype;
 #define RTE_LOGTYPE_PCI_BUS pci_bus_logtype
@@ -41,6 +42,7 @@ struct rte_pci_region {
 
 struct rte_pci_device_internal {
 	struct rte_pci_device device;
+	struct rte_pci_dma_info dma;
 	/* PCI regions provided by e.g. VFIO. */
 	struct rte_pci_region region[RTE_MAX_PCI_REGIONS];
 };
diff --git a/drivers/bus/pci/rte_bus_pci.h b/drivers/bus/pci/rte_bus_pci.h
index 19a7b15b99..1db4fdbdd8 100644
--- a/drivers/bus/pci/rte_bus_pci.h
+++ b/drivers/bus/pci/rte_bus_pci.h
@@ -172,6 +172,58 @@ __rte_internal
 int rte_pci_pasid_set_state(const struct rte_pci_device *dev,
 		off_t offset, bool enable);
 
+/**
+ * @internal
+ * DMA properties of the host bridge above a device.  A zero size means
+ * addresses are not translated.
+ */
+struct rte_pci_dma_info {
+	uint64_t cpu_base;
+	uint64_t bus_base;
+	uint64_t size;
+	bool noncoherent;
+	bool unusable;		/**< properties the bus could not parse. */
+};
+
+/**
+ * @internal
+ * Copy a device's DMA properties.  Drivers supporting secondary processes
+ * keep the copy in shared data.
+ *
+ * @param dev
+ *   The PCI device.
+ * @param info
+ *   Filled in with the bridge's DMA properties.
+ */
+__rte_internal
+void rte_pci_get_dma_info(const struct rte_pci_device *dev,
+		struct rte_pci_dma_info *info);
+
+/**
+ * @internal
+ * Translate an IOVA range into a device address.
+ *
+ * @return
+ *   The device address, or RTE_BAD_IOVA if the range is outside the window.
+ */
+static inline rte_iova_t
+rte_pci_dma_iova(const struct rte_pci_dma_info *info, rte_iova_t iova,
+		size_t len)
+{
+	uint64_t offset;
+
+	if (iova == RTE_BAD_IOVA || len == 0)
+		return RTE_BAD_IOVA;
+	if (info->size == 0)
+		return iova;
+	if (iova < info->cpu_base)
+		return RTE_BAD_IOVA;
+	offset = iova - info->cpu_base;
+	if (offset >= info->size || len > info->size - offset)
+		return RTE_BAD_IOVA;
+	return info->bus_base + offset;
+}
+
 /**
  * Read PCI config space.
  *
diff --git a/lib/eal/include/meson.build b/lib/eal/include/meson.build
index aef5824e5f..9eabd91773 100644
--- a/lib/eal/include/meson.build
+++ b/lib/eal/include/meson.build
@@ -32,6 +32,7 @@ headers += files(
         'rte_lock_annotations.h',
         'rte_malloc.h',
         'rte_mcslock.h',
+        'rte_mem_sync.h',
         'rte_memory.h',
         'rte_memzone.h',
         'rte_pci_dev_feature_defs.h',
diff --git a/lib/eal/include/rte_mem_sync.h b/lib/eal/include/rte_mem_sync.h
new file mode 100644
index 0000000000..6ac517f243
--- /dev/null
+++ b/lib/eal/include/rte_mem_sync.h
@@ -0,0 +1,133 @@
+/* SPDX-License-Identifier: BSD-3-Clause
+ * Copyright(c) 2026 Md Rayhanul Islam
+ */
+
+#ifndef RTE_MEM_SYNC_H
+#define RTE_MEM_SYNC_H
+
+/**
+ * @file
+ * @warning
+ * @b EXPERIMENTAL: this API may change without prior notice.
+ *
+ * Cache maintenance for devices whose DMA is not coherent with the CPU
+ * caches, after dma_sync_single_for_{device,cpu}().  Ownership is per cache
+ * line, and a sync does not wait for the transfer to complete.  Empty on
+ * coherent platforms.
+ */
+
+#include <stdbool.h>
+#include <stddef.h>
+#include <stdint.h>
+
+#include <rte_common.h>
+#include <rte_compat.h>
+
+#ifdef __cplusplus
+extern "C" {
+#endif
+
+/** Transfer direction, as the device sees it. */
+enum rte_mem_sync_direction {
+	RTE_MEM_SYNC_TO_DEVICE,
+	RTE_MEM_SYNC_FROM_DEVICE,
+	RTE_MEM_SYNC_BIDIRECTIONAL,
+};
+
+/** Size of a data cache line on this CPU. */
+__rte_experimental
+static inline size_t
+rte_mem_dcache_line_size(void)
+{
+#ifdef RTE_ARCH_ARM64
+	uint64_t ctr;
+
+	/* CTR_EL0.DminLine is log2 of the line size in words. */
+	asm volatile("mrs %0, ctr_el0" : "=r" (ctr));
+	return 4U << ((ctr >> 16) & 0xf);
+#else
+	return RTE_CACHE_LINE_SIZE;
+#endif
+}
+
+/** Whether this build implements the maintenance below. */
+__rte_experimental
+static inline bool
+rte_mem_sync_supported(void)
+{
+#if defined(RTE_ARCH_ARM64) && defined(RTE_EXEC_ENV_LINUX)
+	return true;
+#else
+	return false;
+#endif
+}
+
+#if defined(RTE_ARCH_ARM64) && defined(RTE_EXEC_ENV_LINUX)
+
+/**
+ * Hand [addr, addr + len) to the device.  EL0 cannot execute DC IVAC, so a
+ * buffer the device writes is cleaned as well as invalidated.
+ */
+__rte_experimental
+static inline void
+rte_mem_sync_for_device(const void *addr, size_t len,
+		enum rte_mem_sync_direction direction)
+{
+	uintptr_t p = (uintptr_t)addr & ~(uintptr_t)(RTE_CACHE_LINE_MIN_SIZE - 1);
+	uintptr_t last = ((uintptr_t)addr + len - 1) &
+			~(uintptr_t)(RTE_CACHE_LINE_MIN_SIZE - 1);
+
+	if (len == 0)
+		return;
+	for (;;) {
+		if (direction == RTE_MEM_SYNC_TO_DEVICE)
+			asm volatile("dc cvac, %0" : : "r" (p) : "memory");
+		else
+			asm volatile("dc civac, %0" : : "r" (p) : "memory");
+		if (p == last)
+			break;
+		p += RTE_CACHE_LINE_MIN_SIZE;
+	}
+	/* DMB does not order cache maintenance. */
+	asm volatile("dsb sy" : : : "memory");
+}
+
+/** Take [addr, addr + len) back from the device once the transfer is done. */
+__rte_experimental
+static inline void
+rte_mem_sync_for_cpu(const void *addr, size_t len,
+		enum rte_mem_sync_direction direction)
+{
+	if (direction != RTE_MEM_SYNC_TO_DEVICE)
+		rte_mem_sync_for_device(addr, len, RTE_MEM_SYNC_FROM_DEVICE);
+}
+
+#else
+
+__rte_experimental
+static inline void
+rte_mem_sync_for_device(const void *addr, size_t len,
+		enum rte_mem_sync_direction direction)
+{
+	RTE_SET_USED(addr);
+	RTE_SET_USED(len);
+	RTE_SET_USED(direction);
+}
+
+__rte_experimental
+static inline void
+rte_mem_sync_for_cpu(const void *addr, size_t len,
+		enum rte_mem_sync_direction direction)
+{
+	RTE_SET_USED(addr);
+	RTE_SET_USED(len);
+	RTE_SET_USED(direction);
+}
+
+#endif /* RTE_ARCH_ARM64 && RTE_EXEC_ENV_LINUX */
+
+#ifdef __cplusplus
+}
+#endif
+
+#endif /* RTE_MEM_SYNC_H */
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* [RFC PATCH 2/2] net/e1000: support non-coherent DMA
  2026-10-02 20:09 ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Md Rayhanul Islam
  2026-10-02 20:09   ` [RFC PATCH 1/2] eal/pci: add per-device DMA translation and synchronization Md Rayhanul Islam
@ 2026-10-02 20:09   ` Md Rayhanul Islam
  2026-10-03 16:51   ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Stephen Hemminger
  2 siblings, 0 replies; 11+ messages in thread
From: Md Rayhanul Islam @ 2026-10-02 20:09 UTC (permalink / raw)
  To: dev

On non-coherent systems the NIC reads stale descriptors and packet data.
Seen on a BCM2711 Compute Module 4 with an Intel I210: transmits never
completed and received data went to address zero.

Take the DMA context from the bus at probe, keep it in the shared adapter
data so secondary processes see it, and install separate TX and RX bursts
for such devices.  They translate every address against the device window
and synchronize buffers and descriptors.  Both come from one inlined
template with a constant flag, so the coherent datapath is unchanged.

Four descriptors share a cache line.  Invalidate a line before writing
into it, so cleaning it afterwards keeps the Done bits the NIC wrote on
the neighbours; without that a burst loses about three packets of every
sixteen.  Completion is the Done bit, never the head register, which
advances when a descriptor is fetched rather than when its data has been
read.  Receive descriptors are refilled a whole line at a time, and
receive buffers are cleaned before handover and rejected unless they are
direct, cache-line aligned and inside the window.

The em driver, which shares net/e1000 with igb, does not set
RTE_PCI_DRV_DMA_NONCOHERENT, so on non-coherent systems the PCI
bus does not probe em devices; they fail cleanly instead of
running with stale DMA.

Tested between two Compute Module 4 boards with Intel I210s: testpmd
txonly holds 1.42 Mpps at 1 GbE line rate, and a 512 MB transfer with
DPDK at both ends arrives byte-identical.

Signed-off-by: Md Rayhanul Islam <r97yhan@gmail.com>
---
 doc/guides/rel_notes/release_26_11.rst |   5 +
 drivers/net/intel/e1000/e1000_ethdev.h |  19 ++
 drivers/net/intel/e1000/igb_ethdev.c   |   9 +-
 drivers/net/intel/e1000/igb_rxtx.c     | 409 +++++++++++++++++++++++--
 4 files changed, 415 insertions(+), 27 deletions(-)

diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index a752c94bbf..5883f0aeb8 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -70,6 +70,11 @@ New Features
   between the CPU and a device whose DMA is not coherent with the caches,
   with an explicit transfer direction.
 
+* **Updated e1000 driver.**
+
+  * Added non-coherent DMA support to ``igb``, through burst functions
+    selected at probe so that coherent platforms are unaffected.
+
 
 Removed Items
 -------------
diff --git a/drivers/net/intel/e1000/e1000_ethdev.h b/drivers/net/intel/e1000/e1000_ethdev.h
index 0907c7c259..979c661f09 100644
--- a/drivers/net/intel/e1000/e1000_ethdev.h
+++ b/drivers/net/intel/e1000/e1000_ethdev.h
@@ -11,6 +11,8 @@
 #include <rte_flow.h>
 #include <rte_time.h>
 #include <rte_pci.h>
+#include <rte_bus_pci.h>
+#include <rte_mem_sync.h>
 
 #define E1000_INTEL_VENDOR_ID 0x8086
 
@@ -283,6 +285,8 @@ struct e1000_adapter {
 	struct e1000_vf_info    *vfdata;
 	struct e1000_filter_info filter;
 	bool stopped;
+	struct rte_pci_dma_info dma;	/**< DMA window of the bridge above. */
+	bool dma_active;		/**< translate and maintain caches. */
 	struct rte_timecounter  systime_tc;
 	struct rte_timecounter  rx_tstamp_tc;
 	struct rte_timecounter  tx_tstamp_tc;
@@ -428,6 +432,21 @@ int eth_igb_tx_queue_setup(struct rte_eth_dev *dev, uint16_t tx_queue_id,
 
 int eth_igb_tx_done_cleanup(void *txq, uint32_t free_cnt);
 
+int igb_dma_probe(struct rte_eth_dev *dev);
+
+void igb_dma_set_burst(struct rte_eth_dev *dev);
+
+#if defined(RTE_ARCH_ARM64)
+uint16_t eth_igb_recv_pkts_nc(void *rxq, struct rte_mbuf **rx_pkts,
+		uint16_t nb_pkts);
+
+uint16_t eth_igb_recv_scattered_pkts_nc(void *rxq, struct rte_mbuf **rx_pkts,
+		uint16_t nb_pkts);
+
+uint16_t eth_igb_xmit_pkts_nc(void *txq, struct rte_mbuf **tx_pkts,
+		uint16_t nb_pkts);
+#endif
+
 int eth_igb_rx_init(struct rte_eth_dev *dev);
 
 void eth_igb_tx_init(struct rte_eth_dev *dev);
diff --git a/drivers/net/intel/e1000/igb_ethdev.c b/drivers/net/intel/e1000/igb_ethdev.c
index 524c030be6..287f2a0105 100644
--- a/drivers/net/intel/e1000/igb_ethdev.c
+++ b/drivers/net/intel/e1000/igb_ethdev.c
@@ -809,9 +809,15 @@ eth_igb_dev_init(struct rte_eth_dev *eth_dev)
 	if (rte_eal_process_type() != RTE_PROC_PRIMARY){
 		if (eth_dev->data->scattered_rx)
 			eth_dev->rx_pkt_burst = &eth_igb_recv_scattered_pkts;
+		igb_dma_set_burst(eth_dev);
 		return 0;
 	}
 
+	error = igb_dma_probe(eth_dev);
+	if (error != 0)
+		return error;
+	igb_dma_set_burst(eth_dev);
+
 	rte_eth_copy_pci_info(eth_dev, pci_dev);
 
 	hw->hw_addr= (void *)pci_dev->mem_resource[0].addr;
@@ -1097,7 +1103,8 @@ static int eth_igb_pci_remove(struct rte_pci_device *pci_dev)
 
 static struct rte_pci_driver rte_igb_pmd = {
 	.id_table = pci_id_igb_map,
-	.drv_flags = RTE_PCI_DRV_NEED_MAPPING | RTE_PCI_DRV_INTR_LSC,
+	.drv_flags = RTE_PCI_DRV_NEED_MAPPING | RTE_PCI_DRV_INTR_LSC |
+		RTE_PCI_DRV_DMA_NONCOHERENT,
 	.probe = eth_igb_pci_probe,
 	.remove = eth_igb_pci_remove,
 };
diff --git a/drivers/net/intel/e1000/igb_rxtx.c b/drivers/net/intel/e1000/igb_rxtx.c
index 4fda5d57a0..3996f23b03 100644
--- a/drivers/net/intel/e1000/igb_rxtx.c
+++ b/drivers/net/intel/e1000/igb_rxtx.c
@@ -4,6 +4,10 @@
 
 #include <sys/queue.h>
 
+#include <limits.h>
+#include <stdbool.h>
+#include <unistd.h>
+
 #include <stdio.h>
 #include <stdlib.h>
 #include <string.h>
@@ -31,6 +35,9 @@
 #include <rte_mbuf.h>
 #include <rte_ether.h>
 #include <ethdev_driver.h>
+#include <bus_pci_driver.h>
+#include <rte_bus_pci.h>
+#include <rte_mem_sync.h>
 #include <rte_prefetch.h>
 #include <rte_udp.h>
 #include <rte_tcp.h>
@@ -42,6 +49,107 @@
 #include "base/e1000_api.h"
 #include "e1000_ethdev.h"
 
+/*
+ * Non-coherent DMA: addresses go through the device's DMA window, and
+ * descriptors and buffers are synced around every handover.
+ */
+void
+igb_dma_set_burst(struct rte_eth_dev *dev)
+{
+#if defined(RTE_ARCH_ARM64)
+	struct e1000_adapter *adapter = E1000_DEV_PRIVATE(dev->data->dev_private);
+
+	if (!adapter->dma_active)
+		return;
+	if (dev->rx_pkt_burst == eth_igb_recv_pkts)
+		dev->rx_pkt_burst = eth_igb_recv_pkts_nc;
+	else if (dev->rx_pkt_burst == eth_igb_recv_scattered_pkts)
+		dev->rx_pkt_burst = eth_igb_recv_scattered_pkts_nc;
+	if (dev->tx_pkt_burst == eth_igb_xmit_pkts)
+		dev->tx_pkt_burst = eth_igb_xmit_pkts_nc;
+#else
+	RTE_SET_USED(dev);
+#endif
+}
+
+int
+igb_dma_probe(struct rte_eth_dev *dev)
+{
+	struct e1000_adapter *adapter = E1000_DEV_PRIVATE(dev->data->dev_private);
+	struct rte_pci_device *pci_dev = RTE_CLASS_TO_BUS_DEVICE(dev, *pci_dev);
+
+	rte_pci_get_dma_info(pci_dev, &adapter->dma);
+	adapter->dma_active = adapter->dma.noncoherent || adapter->dma.size != 0;
+	if (!adapter->dma_active)
+		return 0;
+	if (!rte_mem_sync_supported() || rte_mem_dcache_line_size() !=
+			4 * sizeof(union e1000_adv_rx_desc)) {
+		PMD_INIT_LOG(ERR, "%s: this DMA handling needs 64-byte cache lines",
+			dev->device->name);
+		return -ENOTSUP;
+	}
+	PMD_INIT_LOG(NOTICE, "%s: DMA is not cache coherent, maintaining %zu"
+		"-byte D-cache lines", dev->device->name,
+		rte_mem_dcache_line_size());
+	return 0;
+}
+
+static inline void
+igb_dma_sync_for_device(const volatile void *addr, size_t len,
+		enum rte_mem_sync_direction dir)
+{
+	if (len != 0)
+		rte_mem_sync_for_device((const void *)(uintptr_t)addr, len, dir);
+}
+
+static inline void
+igb_dma_sync_for_cpu(const volatile void *addr, size_t len)
+{
+	if (len != 0)
+		rte_mem_sync_for_cpu((const void *)(uintptr_t)addr, len,
+				RTE_MEM_SYNC_FROM_DEVICE);
+}
+
+static inline void
+igb_rx_buf_sync_for_device(struct rte_mbuf *mb)
+{
+	igb_dma_sync_for_device((char *)mb->buf_addr + RTE_PKTMBUF_HEADROOM,
+			mb->buf_len - RTE_PKTMBUF_HEADROOM,
+			RTE_MEM_SYNC_FROM_DEVICE);
+}
+
+static inline bool
+igb_rx_buf_usable(const struct rte_pci_dma_info *dma, struct rte_mbuf *mb)
+{
+	size_t line = rte_mem_dcache_line_size();
+	uintptr_t addr = (uintptr_t)mb->buf_addr + RTE_PKTMBUF_HEADROOM;
+	size_t len;
+
+	if (!RTE_MBUF_DIRECT(mb) || RTE_MBUF_HAS_EXTBUF(mb) ||
+			mb->buf_len <= RTE_PKTMBUF_HEADROOM)
+		return false;
+	len = mb->buf_len - RTE_PKTMBUF_HEADROOM;
+	return (addr % line) == 0 && (len % line) == 0 &&
+		rte_pci_dma_iova(dma, rte_mbuf_data_iova_default(mb), len) !=
+			RTE_BAD_IOVA;
+}
+
+static inline void
+igb_dma_sync_ring(const volatile void *ring, size_t desc_size, uint16_t nb_desc,
+		uint16_t from, uint16_t to)
+{
+	if (from == to)
+		return;
+	if (from < to) {
+		igb_dma_sync_for_device(RTE_PTR_ADD(ring, from * desc_size),
+				(to - from) * desc_size, RTE_MEM_SYNC_TO_DEVICE);
+	} else {
+		igb_dma_sync_for_device(RTE_PTR_ADD(ring, from * desc_size),
+				(nb_desc - from) * desc_size, RTE_MEM_SYNC_TO_DEVICE);
+		igb_dma_sync_for_device(ring, to * desc_size, RTE_MEM_SYNC_TO_DEVICE);
+	}
+}
+
 #ifdef RTE_LIBRTE_IEEE1588
 #define IGB_TX_IEEE1588_TMST RTE_MBUF_F_TX_IEEE1588_TMST
 #else
@@ -88,6 +196,9 @@ enum igb_rxq_flags {
  * Structure associated with each RX queue.
  */
 struct igb_rx_queue {
+	struct rte_pci_dma_info dma;    /**< translation window of the bridge. */
+	bool                   dma_active; /**< translate and maintain caches. */
+	uint16_t               dma_per_line; /**< descriptors in one cache line. */
 	struct rte_mempool  *mb_pool;   /**< mbuf pool to populate RX ring. */
 	volatile union e1000_adv_rx_desc *rx_ring; /**< RX ring virtual address. */
 	uint64_t            rx_ring_phys_addr; /**< RX ring DMA address. */
@@ -98,6 +209,7 @@ struct igb_rx_queue {
 	struct rte_mbuf *pkt_last_seg;  /**< Last segment of current packet. */
 	uint16_t            nb_rx_desc; /**< number of RX descriptors. */
 	uint16_t            rx_tail;    /**< current value of RDT register. */
+	uint16_t            rx_unwritten; /**< first refill not written back. */
 	uint16_t            nb_rx_hold; /**< number of held free RX desc. */
 	uint16_t            rx_free_thresh; /**< max free RX desc to hold. */
 	uint16_t            queue_id;   /**< RX queue index. */
@@ -163,6 +275,10 @@ struct igb_advctx_info {
  * Structure associated with each TX queue.
  */
 struct igb_tx_queue {
+	struct rte_pci_dma_info dma;    /**< translation window of the bridge. */
+	bool                   dma_active; /**< translate and maintain caches. */
+	uint16_t               dma_per_line; /**< descriptors in one cache line. */
+	uint16_t               last_rs; /**< newest descriptor asking for status. */
 	volatile union e1000_adv_tx_desc *tx_ring; /**< TX ring address */
 	uint64_t               tx_ring_phys_addr; /**< TX ring DMA address. */
 	struct igb_tx_entry    *sw_ring; /**< virtual address of SW ring. */
@@ -384,9 +500,97 @@ tx_desc_vlan_flags_to_cmdtype(uint64_t ol_flags)
 	return cmdtype;
 }
 
-uint16_t
-eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
-	       uint16_t nb_pkts)
+/*
+ * Refresh a descriptor line before writing it, so the clean that publishes it
+ * does not erase Done bits the NIC set on its neighbours.
+ */
+static inline void
+igb_tx_line_acquire(struct igb_tx_queue *txq, uint16_t desc)
+{
+	uint16_t line = desc & ~(uint16_t)(txq->dma_per_line - 1);
+
+	igb_dma_sync_for_cpu(&txq->tx_ring[line],
+			txq->dma_per_line * sizeof(*txq->tx_ring));
+}
+
+/* Completion is in order, so Done on the last RS descriptor covers this one. */
+static inline bool
+igb_tx_desc_done(struct igb_tx_queue *txq, uint16_t desc, const bool nc)
+{
+	volatile uint32_t *status = &txq->tx_ring[desc].wb.status;
+
+	if (!nc)
+		return (*status & rte_cpu_to_le_32(E1000_TXD_STAT_DD)) != 0;
+
+	igb_dma_sync_for_cpu(status, sizeof(*status));
+	if (*status & rte_cpu_to_le_32(E1000_TXD_STAT_DD))
+		return true;
+
+	status = &txq->tx_ring[txq->last_rs].wb.status;
+	igb_dma_sync_for_cpu(status, sizeof(*status));
+	return (*status & rte_cpu_to_le_32(E1000_TXD_STAT_DD)) != 0;
+}
+
+static bool
+igb_tx_pkt_reachable(struct igb_tx_queue *txq, struct rte_mbuf *mb)
+{
+	uint16_t n = 0, expected = mb->nb_segs;
+
+	for (; mb != NULL && n < txq->nb_tx_desc; mb = mb->next, n++)
+		if (rte_pci_dma_iova(&txq->dma, rte_mbuf_data_iova(mb),
+				mb->data_len) == RTE_BAD_IOVA)
+			return false;
+	return mb == NULL && n == expected;
+}
+
+static inline uint16_t
+igb_rx_desc_per_line(struct igb_rx_queue *rxq)
+{
+	return rxq->dma_per_line;
+}
+
+static inline void
+igb_rx_write_range(struct igb_rx_queue *rxq, uint16_t from, uint16_t to)
+{
+	uint16_t i;
+
+	for (i = from; i < to; i++) {
+		volatile union e1000_adv_rx_desc *rxd = &rxq->rx_ring[i];
+		struct rte_mbuf *mb = rxq->sw_ring[i].mbuf;
+
+		igb_rx_buf_sync_for_device(mb);
+		rxd->read.hdr_addr = 0;
+		rxd->read.pkt_addr = rte_cpu_to_le_64(rte_pci_dma_iova(&rxq->dma,
+			rte_mbuf_data_iova_default(mb),
+			mb->buf_len - RTE_PKTMBUF_HEADROOM));
+	}
+	igb_dma_sync_for_device(&rxq->rx_ring[from],
+			(to - from) * sizeof(*rxq->rx_ring), RTE_MEM_SYNC_TO_DEVICE);
+}
+
+/* Post whole descriptor lines; returns the first one not yet posted. */
+static inline uint16_t
+igb_rx_flush_refills(struct igb_rx_queue *rxq, uint16_t rx_id)
+{
+	uint16_t mask = igb_rx_desc_per_line(rxq) - 1;
+	uint16_t start = rxq->rx_unwritten;
+	uint16_t end = rx_id & ~mask;
+
+	if (end == start)
+		return start;
+	if (end < start) {
+		igb_rx_write_range(rxq, start, rxq->nb_rx_desc);
+		start = 0;
+	}
+	if (end > start)
+		igb_rx_write_range(rxq, start, end);
+	rxq->rx_unwritten = end;
+	return end;
+}
+
+static __rte_always_inline uint16_t
+igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts, uint16_t nb_pkts,
+	      const bool nc)
 {
 	struct igb_tx_queue *txq;
 	struct igb_tx_entry *sw_ring;
@@ -419,6 +623,10 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 
 	for (nb_tx = 0; nb_tx < nb_pkts; nb_tx++) {
 		tx_pkt = *tx_pkts++;
+		if (nc && (tx_pkt->nb_segs == 0 ||
+				tx_pkt->nb_segs >= txq->nb_tx_desc ||
+				!igb_tx_pkt_reachable(txq, tx_pkt)))
+			break;
 		pkt_len = tx_pkt->pkt_len;
 
 		RTE_MBUF_PREFETCH_TO_FREE(txe->mbuf);
@@ -508,7 +716,7 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 		/*
 		 * Check that this descriptor is free.
 		 */
-		if (! (txr[tx_end].wb.status & E1000_TXD_STAT_DD)) {
+		if (!igb_tx_desc_done(txq, tx_end, nc)) {
 			if (nb_tx == 0)
 				return 0;
 			goto end_of_tx;
@@ -595,6 +803,15 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 			 */
 			slen = (uint16_t) m_seg->data_len;
 			buf_dma_addr = rte_mbuf_data_iova(m_seg);
+			if (nc) {
+				if ((tx_id & (txq->dma_per_line - 1)) == 0 ||
+						tx_id == txq->tx_tail)
+					igb_tx_line_acquire(txq, tx_id);
+				buf_dma_addr = rte_pci_dma_iova(&txq->dma,
+						buf_dma_addr, slen);
+				igb_dma_sync_for_device(rte_pktmbuf_mtod(m_seg, void *),
+						slen, RTE_MEM_SYNC_TO_DEVICE);
+			}
 			txd->read.buffer_addr =
 				rte_cpu_to_le_64(buf_dma_addr);
 			txd->read.cmd_type_len =
@@ -613,10 +830,15 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 		 */
 		txd->read.cmd_type_len |=
 			rte_cpu_to_le_32(E1000_TXD_CMD_EOP | E1000_TXD_CMD_RS);
+		txq->last_rs = tx_last;
 	}
  end_of_tx:
 	rte_wmb();
 
+	if (nc)
+		igb_dma_sync_ring(txq->tx_ring, sizeof(*txq->tx_ring),
+				txq->nb_tx_desc, txq->tx_tail, tx_id);
+
 	/*
 	 * Set the Transmit Descriptor Tail (TDT).
 	 */
@@ -629,6 +851,20 @@ eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts,
 	return nb_tx;
 }
 
+uint16_t
+eth_igb_xmit_pkts(void *tx_queue, struct rte_mbuf **tx_pkts, uint16_t nb_pkts)
+{
+	return igb_xmit_pkts(tx_queue, tx_pkts, nb_pkts, false);
+}
+
+#if defined(RTE_ARCH_ARM64)
+uint16_t
+eth_igb_xmit_pkts_nc(void *tx_queue, struct rte_mbuf **tx_pkts, uint16_t nb_pkts)
+{
+	return igb_xmit_pkts(tx_queue, tx_pkts, nb_pkts, true);
+}
+#endif
+
 /*********************************************************************
  *
  *  TX prep functions
@@ -817,9 +1053,9 @@ rx_desc_error_to_pkt_flags(uint32_t rx_status)
 		E1000_RXD_ERR_CKSUM_BIT) & E1000_RXD_ERR_CKSUM_MSK];
 }
 
-uint16_t
-eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
-	       uint16_t nb_pkts)
+static __rte_always_inline uint16_t
+igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts, uint16_t nb_pkts,
+	      const bool nc)
 {
 	struct igb_rx_queue *rxq;
 	volatile union e1000_adv_rx_desc *rx_ring;
@@ -854,6 +1090,8 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		 * using invalid descriptor fields when read from rxd.
 		 */
 		rxdp = &rx_ring[rx_id];
+		if (nc)
+			igb_dma_sync_for_cpu(rxdp, sizeof(*rxdp));
 		staterr = rxdp->wb.upper.status_error;
 		if (! (staterr & rte_cpu_to_le_32(E1000_RXD_STAT_DD)))
 			break;
@@ -900,6 +1138,11 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 			break;
 		}
 
+		if (nc && !igb_rx_buf_usable(&rxq->dma, nmb)) {
+			rte_pktmbuf_free(nmb);
+			rte_eth_devices[rxq->port_id].data->rx_mbuf_alloc_failed++;
+			break;
+		}
 		nb_hold++;
 		rxe = &sw_ring[rx_id];
 		rx_id++;
@@ -923,8 +1166,10 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		rxe->mbuf = nmb;
 		dma_addr =
 			rte_cpu_to_le_64(rte_mbuf_data_iova_default(nmb));
-		rxdp->read.hdr_addr = 0;
-		rxdp->read.pkt_addr = dma_addr;
+		if (!nc) {	/* deferred to igb_rx_flush_refills() */
+			rxdp->read.hdr_addr = 0;
+			rxdp->read.pkt_addr = dma_addr;
+		}
 
 		/*
 		 * Initialize the returned mbuf.
@@ -942,6 +1187,9 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		pkt_len = (uint16_t) (rte_le_to_cpu_16(rxd.wb.upper.length) -
 				      rxq->crc_len);
 		rxm->data_off = RTE_PKTMBUF_HEADROOM;
+		if (nc)
+			igb_dma_sync_for_cpu((char *)rxm->buf_addr + rxm->data_off,
+					rte_le_to_cpu_16(rxd.wb.upper.length));
 		rte_packet_prefetch((char *)rxm->buf_addr + rxm->data_off);
 		rxm->nb_segs = 1;
 		rxm->next = NULL;
@@ -993,18 +1241,39 @@ eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 			   (unsigned) rxq->port_id, (unsigned) rxq->queue_id,
 			   (unsigned) rx_id, (unsigned) nb_hold,
 			   (unsigned) nb_rx);
+		if (nc) {
+			uint16_t first = rxq->rx_unwritten;
+
+			rx_id = igb_rx_flush_refills(rxq, rx_id);
+			nb_hold -= (rx_id + rxq->nb_rx_desc - first) % rxq->nb_rx_desc;
+		} else {
+			nb_hold = 0;
+		}
 		rx_id = (uint16_t) ((rx_id == 0) ?
 				     (rxq->nb_rx_desc - 1) : (rx_id - 1));
 		E1000_PCI_REG_WRITE(rxq->rdt_reg_addr, rx_id);
-		nb_hold = 0;
 	}
 	rxq->nb_rx_hold = nb_hold;
 	return nb_rx;
 }
 
 uint16_t
-eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
-			 uint16_t nb_pkts)
+eth_igb_recv_pkts(void *rx_queue, struct rte_mbuf **rx_pkts, uint16_t nb_pkts)
+{
+	return igb_recv_pkts(rx_queue, rx_pkts, nb_pkts, false);
+}
+
+#if defined(RTE_ARCH_ARM64)
+uint16_t
+eth_igb_recv_pkts_nc(void *rx_queue, struct rte_mbuf **rx_pkts, uint16_t nb_pkts)
+{
+	return igb_recv_pkts(rx_queue, rx_pkts, nb_pkts, true);
+}
+#endif
+
+static __rte_always_inline uint16_t
+igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
+			uint16_t nb_pkts, const bool nc)
 {
 	struct igb_rx_queue *rxq;
 	volatile union e1000_adv_rx_desc *rx_ring;
@@ -1049,6 +1318,8 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		 * using invalid descriptor fields when read from rxd.
 		 */
 		rxdp = &rx_ring[rx_id];
+		if (nc)
+			igb_dma_sync_for_cpu(rxdp, sizeof(*rxdp));
 		staterr = rxdp->wb.upper.status_error;
 		if (! (staterr & rte_cpu_to_le_32(E1000_RXD_STAT_DD)))
 			break;
@@ -1091,6 +1362,11 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 			break;
 		}
 
+		if (nc && !igb_rx_buf_usable(&rxq->dma, nmb)) {
+			rte_pktmbuf_free(nmb);
+			rte_eth_devices[rxq->port_id].data->rx_mbuf_alloc_failed++;
+			break;
+		}
 		nb_hold++;
 		rxe = &sw_ring[rx_id];
 		rx_id++;
@@ -1117,8 +1393,10 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		rxm = rxe->mbuf;
 		rxe->mbuf = nmb;
 		dma = rte_cpu_to_le_64(rte_mbuf_data_iova_default(nmb));
-		rxdp->read.pkt_addr = dma;
-		rxdp->read.hdr_addr = 0;
+		if (!nc) {	/* deferred to igb_rx_flush_refills() */
+			rxdp->read.pkt_addr = dma;
+			rxdp->read.hdr_addr = 0;
+		}
 
 		/*
 		 * Set data length & data buffer address of mbuf.
@@ -1126,6 +1404,9 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 		data_len = rte_le_to_cpu_16(rxd.wb.upper.length);
 		rxm->data_len = data_len;
 		rxm->data_off = RTE_PKTMBUF_HEADROOM;
+		if (nc)
+			igb_dma_sync_for_cpu((char *)rxm->buf_addr + rxm->data_off,
+					data_len);
 
 		/*
 		 * If this is the first buffer of the received packet,
@@ -1255,15 +1536,38 @@ eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
 			   (unsigned) rxq->port_id, (unsigned) rxq->queue_id,
 			   (unsigned) rx_id, (unsigned) nb_hold,
 			   (unsigned) nb_rx);
+		if (nc) {
+			uint16_t first = rxq->rx_unwritten;
+
+			rx_id = igb_rx_flush_refills(rxq, rx_id);
+			nb_hold -= (rx_id + rxq->nb_rx_desc - first) % rxq->nb_rx_desc;
+		} else {
+			nb_hold = 0;
+		}
 		rx_id = (uint16_t) ((rx_id == 0) ?
 				     (rxq->nb_rx_desc - 1) : (rx_id - 1));
 		E1000_PCI_REG_WRITE(rxq->rdt_reg_addr, rx_id);
-		nb_hold = 0;
 	}
 	rxq->nb_rx_hold = nb_hold;
 	return nb_rx;
 }
 
+uint16_t
+eth_igb_recv_scattered_pkts(void *rx_queue, struct rte_mbuf **rx_pkts,
+			    uint16_t nb_pkts)
+{
+	return igb_recv_scattered_pkts(rx_queue, rx_pkts, nb_pkts, false);
+}
+
+#if defined(RTE_ARCH_ARM64)
+uint16_t
+eth_igb_recv_scattered_pkts_nc(void *rx_queue, struct rte_mbuf **rx_pkts,
+			       uint16_t nb_pkts)
+{
+	return igb_recv_scattered_pkts(rx_queue, rx_pkts, nb_pkts, true);
+}
+#endif
+
 /*
  * Maximum number of Ring Descriptors.
  *
@@ -1308,7 +1612,6 @@ static int
 igb_tx_done_cleanup(struct igb_tx_queue *txq, uint32_t free_cnt)
 {
 	struct igb_tx_entry *sw_ring;
-	volatile union e1000_adv_tx_desc *txr;
 	uint16_t tx_first; /* First segment analyzed. */
 	uint16_t tx_id;    /* Current segment being processed. */
 	uint16_t tx_last;  /* Last segment in the current packet. */
@@ -1319,7 +1622,6 @@ igb_tx_done_cleanup(struct igb_tx_queue *txq, uint32_t free_cnt)
 		return -ENODEV;
 
 	sw_ring = txq->sw_ring;
-	txr = txq->tx_ring;
 
 	/* tx_tail is the last sent packet on the sw_ring. Goto the end
 	 * of that packet (the last segment in the packet chain) and
@@ -1345,8 +1647,7 @@ igb_tx_done_cleanup(struct igb_tx_queue *txq, uint32_t free_cnt)
 		tx_last = sw_ring[tx_id].last_id;
 
 		if (sw_ring[tx_last].mbuf) {
-			if (txr[tx_last].wb.status &
-			    E1000_TXD_STAT_DD) {
+			if (igb_tx_desc_done(txq, tx_last, txq->dma_active)) {
 				/* Increment the number of packets
 				 * freed.
 				 */
@@ -1466,6 +1767,10 @@ igb_reset_tx_queue(struct igb_tx_queue *txq, struct rte_eth_dev *dev)
 		txq->ctx_start = txq->queue_id * IGB_CTX_NUM;
 
 	igb_reset_tx_queue_stat(txq);
+	if (txq->dma_active)
+		igb_dma_sync_for_device(txq->tx_ring,
+				txq->nb_tx_desc * sizeof(*txq->tx_ring),
+				RTE_MEM_SYNC_TO_DEVICE);
 }
 
 uint64_t
@@ -1573,9 +1878,21 @@ eth_igb_tx_queue_setup(struct rte_eth_dev *dev,
 	txq->reg_idx = (uint16_t)((RTE_ETH_DEV_SRIOV(dev).active == 0) ?
 		queue_idx : RTE_ETH_DEV_SRIOV(dev).def_pool_q_idx + queue_idx);
 	txq->port_id = dev->data->port_id;
+	txq->dma = E1000_DEV_PRIVATE(dev->data->dev_private)->dma;
+	txq->dma_active = E1000_DEV_PRIVATE(dev->data->dev_private)->dma_active;
+	txq->dma_per_line = 1;
+	if (txq->dma_active) {
+		txq->dma_per_line = rte_mem_dcache_line_size() /
+				sizeof(*txq->tx_ring);
+		txq->wthresh = 0;
+	}
 
 	txq->tdt_reg_addr = E1000_PCI_REG_ADDR(hw, E1000_TDT(txq->reg_idx));
-	txq->tx_ring_phys_addr = tz->iova;
+	txq->tx_ring_phys_addr = rte_pci_dma_iova(&txq->dma, tz->iova, size);
+	if (txq->tx_ring_phys_addr == RTE_BAD_IOVA) {
+		igb_tx_queue_release(txq);
+		return -EINVAL;
+	}
 
 	txq->tx_ring = (union e1000_adv_tx_desc *) tz->addr;
 	/* Allocate software ring */
@@ -1591,6 +1908,7 @@ eth_igb_tx_queue_setup(struct rte_eth_dev *dev,
 
 	igb_reset_tx_queue(txq, dev);
 	dev->tx_pkt_burst = eth_igb_xmit_pkts;
+	igb_dma_set_burst(dev);
 	dev->tx_pkt_prepare = &eth_igb_prep_pkts;
 	dev->data->tx_queues[queue_idx] = txq;
 	txq->offloads = offloads;
@@ -1642,6 +1960,9 @@ igb_reset_rx_queue(struct igb_rx_queue *rxq)
 	}
 
 	rxq->rx_tail = 0;
+	rxq->rx_unwritten = 0;
+	if (rxq->dma_active)
+		rxq->nb_rx_hold = 0;
 	rxq->pkt_first_seg = NULL;
 	rxq->pkt_last_seg = NULL;
 }
@@ -1748,6 +2069,9 @@ eth_igb_rx_queue_setup(struct rte_eth_dev *dev,
 	rxq->reg_idx = (uint16_t)((RTE_ETH_DEV_SRIOV(dev).active == 0) ?
 		queue_idx : RTE_ETH_DEV_SRIOV(dev).def_pool_q_idx + queue_idx);
 	rxq->port_id = dev->data->port_id;
+	rxq->dma = E1000_DEV_PRIVATE(dev->data->dev_private)->dma;
+	rxq->dma_active = E1000_DEV_PRIVATE(dev->data->dev_private)->dma_active;
+	rxq->dma_per_line = 1;
 	if (dev->data->dev_conf.rxmode.offloads & RTE_ETH_RX_OFFLOAD_KEEP_CRC)
 		rxq->crc_len = RTE_ETHER_CRC_LEN;
 	else
@@ -1769,8 +2093,20 @@ eth_igb_rx_queue_setup(struct rte_eth_dev *dev,
 	rxq->mz = rz;
 	rxq->rdt_reg_addr = E1000_PCI_REG_ADDR(hw, E1000_RDT(rxq->reg_idx));
 	rxq->rdh_reg_addr = E1000_PCI_REG_ADDR(hw, E1000_RDH(rxq->reg_idx));
-	rxq->rx_ring_phys_addr = rz->iova;
+	rxq->rx_ring_phys_addr = rte_pci_dma_iova(&rxq->dma, rz->iova, size);
+	if (rxq->rx_ring_phys_addr == RTE_BAD_IOVA) {
+		igb_rx_queue_release(rxq);
+		return -EINVAL;
+	}
 	rxq->rx_ring = (union e1000_adv_rx_desc *) rz->addr;
+	if (rxq->dma_active) {
+		size_t line = rte_mem_dcache_line_size();
+		uint16_t per_line = line / sizeof(*rxq->rx_ring);
+
+		if (per_line > 1 && rxq->nb_rx_desc % per_line == 0 &&
+				((uintptr_t)rxq->rx_ring & (line - 1)) == 0)
+			rxq->dma_per_line = per_line;
+	}
 
 	/* Allocate software ring. */
 	rxq->sw_ring = rte_zmalloc("rxq->sw_ring",
@@ -1800,8 +2136,11 @@ eth_igb_rx_queue_count(void *rx_queue)
 	rxq = rx_queue;
 	rxdp = &(rxq->rx_ring[rxq->rx_tail]);
 
-	while ((desc < rxq->nb_rx_desc) &&
-		(rxdp->wb.upper.status_error & E1000_RXD_STAT_DD)) {
+	while (desc < rxq->nb_rx_desc) {
+		if (rxq->dma_active)
+			igb_dma_sync_for_cpu(rxdp, sizeof(*rxdp));
+		if ((rxdp->wb.upper.status_error & E1000_RXD_STAT_DD) == 0)
+			break;
 		desc += IGB_RXQ_SCAN_INTERVAL;
 		rxdp += IGB_RXQ_SCAN_INTERVAL;
 		if (rxq->rx_tail + desc >= rxq->nb_rx_desc)
@@ -1830,6 +2169,8 @@ eth_igb_rx_descriptor_status(void *rx_queue, uint16_t offset)
 		desc -= rxq->nb_rx_desc;
 
 	status = &rxq->rx_ring[desc].wb.upper.status_error;
+	if (rxq->dma_active)
+		igb_dma_sync_for_cpu(status, sizeof(*status));
 	if (*status & rte_cpu_to_le_32(E1000_RXD_STAT_DD))
 		return RTE_ETH_RX_DESC_DONE;
 
@@ -1840,7 +2181,6 @@ int
 eth_igb_tx_descriptor_status(void *tx_queue, uint16_t offset)
 {
 	struct igb_tx_queue *txq = tx_queue;
-	volatile uint32_t *status;
 	uint32_t desc;
 
 	if (unlikely(offset >= txq->nb_tx_desc))
@@ -1850,8 +2190,7 @@ eth_igb_tx_descriptor_status(void *tx_queue, uint16_t offset)
 	if (desc >= txq->nb_tx_desc)
 		desc -= txq->nb_tx_desc;
 
-	status = &txq->tx_ring[desc].wb.status;
-	if (*status & rte_cpu_to_le_32(E1000_TXD_STAT_DD))
+	if (igb_tx_desc_done(txq, desc, txq->dma_active))
 		return RTE_ETH_TX_DESC_DONE;
 
 	return RTE_ETH_TX_DESC_FULL;
@@ -2271,11 +2610,25 @@ igb_alloc_rx_queue_mbufs(struct igb_rx_queue *rxq)
 		}
 		dma_addr =
 			rte_cpu_to_le_64(rte_mbuf_data_iova_default(mbuf));
+		if (rxq->dma_active) {
+			if (!igb_rx_buf_usable(&rxq->dma, mbuf)) {
+				rte_pktmbuf_free(mbuf);
+				return -EINVAL;
+			}
+			dma_addr = rte_cpu_to_le_64(rte_pci_dma_iova(&rxq->dma,
+					rte_mbuf_data_iova_default(mbuf),
+					mbuf->buf_len - RTE_PKTMBUF_HEADROOM));
+			igb_rx_buf_sync_for_device(mbuf);
+		}
 		rxd = &rxq->rx_ring[i];
 		rxd->read.hdr_addr = 0;
 		rxd->read.pkt_addr = dma_addr;
 		rxe[i].mbuf = mbuf;
 	}
+	if (rxq->dma_active)
+		igb_dma_sync_for_device(rxq->rx_ring,
+				rxq->nb_rx_desc * sizeof(*rxq->rx_ring),
+				RTE_MEM_SYNC_TO_DEVICE);
 
 	return 0;
 }
@@ -2583,6 +2936,7 @@ eth_igb_rx_init(struct rte_eth_dev *dev)
 		E1000_WRITE_REG(hw, E1000_RDH(rxq->reg_idx), 0);
 		E1000_WRITE_REG(hw, E1000_RDT(rxq->reg_idx), rxq->nb_rx_desc - 1);
 	}
+	igb_dma_set_burst(dev);
 
 	return 0;
 }
@@ -2624,6 +2978,8 @@ eth_igb_tx_init(struct rte_eth_dev *dev)
 
 		/* Setup Transmit threshold registers. */
 		txdctl = E1000_READ_REG(hw, E1000_TXDCTL(txq->reg_idx));
+		if (txq->dma_active)
+			txdctl &= ~(0x1FU << 16);	/* the field is OR-ed below */
 		txdctl |= txq->pthresh & 0x1F;
 		txdctl |= ((txq->hthresh & 0x1F) << 8);
 		txdctl |= ((txq->wthresh & 0x1F) << 16);
@@ -2793,6 +3149,7 @@ eth_igbvf_rx_init(struct rte_eth_dev *dev)
 		E1000_WRITE_REG(hw, E1000_RDH(i), 0);
 		E1000_WRITE_REG(hw, E1000_RDT(i), rxq->nb_rx_desc - 1);
 	}
+	igb_dma_set_burst(dev);
 
 	return 0;
 }
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 11+ messages in thread

* Re: [RFC PATCH 0/2] per-device DMA properties, with igb as the first user
  2026-10-02 20:09 ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Md Rayhanul Islam
  2026-10-02 20:09   ` [RFC PATCH 1/2] eal/pci: add per-device DMA translation and synchronization Md Rayhanul Islam
  2026-10-02 20:09   ` [RFC PATCH 2/2] net/e1000: support non-coherent DMA Md Rayhanul Islam
@ 2026-10-03 16:51   ` Stephen Hemminger
  2 siblings, 0 replies; 11+ messages in thread
From: Stephen Hemminger @ 2026-10-03 16:51 UTC (permalink / raw)
  To: Md Rayhanul Islam; +Cc: dev

On Fri,  2 Oct 2026 16:09:25 -0400
Md Rayhanul Islam <r97yhan@gmail.com> wrote:

> This RFC reworks the BCM2711 DMA fix around a per-device DMA context and
> an EAL sync API, after the feedback on the first posting [1].
> 
> On Raspberry Pi 4 / Compute Module 4 the PCIe host bridge translates DMA
> addresses and is not cache coherent, so a device such as the Intel I210
> cannot DMA at all without both being handled.
> 
> What changed since the first posting:
> 
> * The translation and the coherency are read per device from its own
>   bridge, not kept as one process-wide offset.  EAL no longer rewrites
>   IOVAs; the driver translates where it programs the hardware.
> * A driver opts in with RTE_PCI_DRV_DMA_NONCOHERENT, and the bus refuses
>   such a device to any driver that has not.  That replaces the refusal
>   written into em by hand.
> * The cache maintenance is an experimental EAL interface with explicit
>   directions, and the loops end in DSB SY.
> * Completion comes from the Done bits, not the head register.
> * Both environment variables are gone.
> 
> Tested between two Compute Module 4 boards with Intel I210s, both running
> this series: testpmd txonly holds 1.42 Mpps with 64-byte frames, and a
> 512 MB UDP transfer with DPDK at both ends arrives byte-identical.  Built
> with GCC and with clang 14 -Werror.

More detailed review with AI, I had to direct it to ignore using AF_XDP
in this case.

Native PMDs on Raspberry Pi class boards are a reasonable goal, and the
analysis of descriptor line sharing is good. Needs rework before a
non-RFC version.

Direction
- Split the series: (1) detection and refusing to probe, which is a fix
  on its own since igb on a CM4 today DMAs to bus address 0; (2) the EAL
  cache maintenance API; (3) driver support.
- Coherency detection should not depend on the bridge allowlist, only
  window parsing does. On arm64 DT, no dma-coherent means non-coherent
  (same as of_dma_is_coherent()). At least warn for any DT host bridge
  without it.
- igb will not be the last PMD (igc is next). Line ownership rules for
  rings, mempool validation against the window and burst variant
  selection belong in common code, not copied per driver.
- An address limit is a property of memory. Enforce it at hugepage
  allocation (extend rte_mem_set_dma_mask(); a bit mask can't express
  3 GB) and keep only the offset add in the datapath.
- Document that this needs uio_pci_generic or vfio-noiommu. VFIO with an
  IOMMU refuses non-coherent devices.

Translation
- BCM2711 does not translate: pcie0 dma-ranges is 1:1 with a 3 GB limit.
- BCM2712 (Pi 5, CM5) does (PCIe 0x10_0000_0000 maps to CPU 0), but the
  parser rejects it: 64-bit memory space code, more than one dma-ranges
  entry (MSI window, plus a 32-bit window on pcie2), and
  brcm,bcm2712-pcie is not in the table. The offset path has never run.

Bus and EAL
- PCI_LOG uses dev->name before pci_common_set() sets it.
- --iova-mode=va overrides the forced PA. Probe must check
  rte_eal_iova_mode() == RTE_IOVA_PA.
- The DT walk runs for every PCI device on every platform. Skip it when
  /sys/firmware/devicetree does not exist.
- rte_pci_dma_info is driver-only API; it goes in bus_pci_driver.h.
- rte_mem_sync.h: arch code goes in lib/eal/arm/include behind a
  generic/ header. "Empty on coherent platforms" is wrong (it is empty
  on non-arm64). Read CTR_EL0 once at init, since it can trap. Stride by
  the CTR_EL0 line size, not RTE_CACHE_LINE_MIN_SIZE. for_cpu is
  clean+invalidate, so RTE_ASSERT alignment. Add a test in app/test.

igb
- Per-packet validation (igb_tx_pkt_reachable, igb_rx_buf_usable)
  belongs at queue setup or in tx_prepare. TX silently stops the burst,
  so retrying applications spin. RX stalls and counts as mbuf allocation
  failure.
- Context descriptors are written before the segment loop, so their
  lines are never acquired.
- The acquire narrows the DD loss window but does not close it.
  Correctness rests on the last_rs fallback; document that.
- last_rs is updated per packet, so the in-burst fallback checks a
  descriptor not yet posted. Snapshot it at burst entry and reset it in
  igb_reset_tx_queue().
- The dma_per_line fallback to 1 is dead code, and if ever reached it
  would reintroduce the bug. Make it an error.
- The new fields at the head of the queue structs push the coherent
  path's hot fields down. Move them to the end and show pahole.
- TXDCTL: if OR-ing into WTHRESH is wrong, it is wrong on every
  platform. Fix it separately with a Fixes: tag, and explain forcing
  wthresh to 0.
- Nits: limits.h and unistd.h are unused, igb_dma_set_burst() in the
  igbvf path is dead, and the error on 32-bit Arm is misleading.

Testing
- 1.42 Mpps is not 64 byte line rate (1.488). What frame size?
- Need RX and io/mac fwd numbers, runs with checksum, VLAN insert and
  TSO offloads, and a Pi 5.
- Board RAM size? On a 4 or 8 GB CM4, hugepages can land above 3 GB.

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-10-03 16:51 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-18 22:50 [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Md Rayhanul Islam
2026-09-18 22:50 ` [PATCH 1/2] eal/linux: apply PCIe inbound DMA translation Md Rayhanul Islam
2026-09-18 22:50 ` [PATCH 2/2] net/e1000: maintain caches on non-coherent DMA Md Rayhanul Islam
2026-09-19 21:05 ` [PATCH 0/2] make PCIe DMA work on Broadcom BCM2711 Stephen Hemminger
2026-09-21  8:56   ` Bruce Richardson
2026-09-21 15:42     ` Stephen Hemminger
2026-09-24 19:22       ` Md Rayhanul Islam
2026-10-02 20:09 ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Md Rayhanul Islam
2026-10-02 20:09   ` [RFC PATCH 1/2] eal/pci: add per-device DMA translation and synchronization Md Rayhanul Islam
2026-10-02 20:09   ` [RFC PATCH 2/2] net/e1000: support non-coherent DMA Md Rayhanul Islam
2026-10-03 16:51   ` [RFC PATCH 0/2] per-device DMA properties, with igb as the first user Stephen Hemminger

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.