* [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools
@ 2026-10-03 21:22 Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table Luigi Rizzo
` (21 more replies)
0 siblings, 22 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Here is a subsystem called DMA_PMD on which I would like feedback on
architecture, possible enhancements, or kernel components that could be
reused to avoid duplication.
I think it can be extremely useful for those affected by the HW/SW overhead
of IOMMU (especially IOTLB thrashing, see [1]), or confidential computing,
or preemptable VMs.
The (not too exaggerated) pitch line is
DMA_PMD has the performance of identity, guarantees that only IO buffers
can ever have active IOMMU or IOTLB mappings, removes the bounce buffer
overhead in confidential computing and preemptable VMs, and integrates
smoothly with existing kernel APIs.
=== ARCHITECTURE
DMA_PMD was initially designed to address IOTLB thrashing (details in [1]),
but turned out to also resolve nicely the strict IOMMU overhead, and avoid
the bounce buffer overhead in preemptible VMs and Confidential Computing.
It works by combining several known techniques:
- transparently feed allocators of device memory (skb_page_frag_refill(),
dma_alloc_attrs(), pagepool, ...) with DMA_PMD pages, i.e. physically
contiguous PMD_SIZE (2MB) pages mapped via PDE_SIZE IOMMU entries
- heavy recycling of DMA_PMD pages (like pagepool) and lazy IOMMU unmapping
(like DMA-FQ), BUT:
- like strict IOMMU (DMA), safely release memory back to the kernel only
after destroying all of its IOMMU mappings and a synchronous IOTLB flush
- on allocation, DMA_PMD pages can be configured to be pinned in the host
(hence suitable for preemptible VMs) and/or unencrypted (hence suitable
for Confidential Computing), removing the need for bounce buffers
- there are global and per-device optin /sys/device/.../dma_pmd_*
and /proc/sys/net/core/tx_enable_dma_pmd
NIC drivers can typically use dma_pmd with no modifications for rings,
tx buffers, rx buffers (if they use pagepool), and rx headers. For tx
headers, almost all drivers need changes (see later patches in the series)
to implement cheap bounce buffers and avoid individual 4K mappings.
The series has the following main components:
- a sparse array (similar to pageblock_flags) to quickly attach metadata
to a 2MB page without fiddling with the struct page. Cost is 128B per 2MB
page used as an IO buffer, totally negligible.
- dma_pmd_pool, a replacement for alloc_pages() that can be instantiated
per-CPU or per receive queue. It handles the split of PMD_SIZE pages
into order-N blocks, handles dma_map and unmap, and aggressively recycles
entries. It is used to feed skb_page_frag_refill(), pagepool, and
receive buffers for drivers that do not use pagepool.
See [2] for "WHY NOT PAGEPOOL FOR TX AND EVERYTHING"
- dma_pmd_arena, is another allocator backed by DMA_PMD pages and
is used exclusively as the backend for dma_alloc_attrs().
Used for longer-lived allocations (descriptor/completion rings,
rx and tx header buffers)
- glue code to hook DMA_PMD into pagepool [2], skb_page_frag_refill(),
dma_alloc_attrs(), dma_map/unmap...
- per-driver patches, where necessary (e.g. tx header buffers [3])
=== PERFORMANCE BENEFITS
Your mileage may vary. Enabling the IOMMU may have no throughput
impact, until it does when some system components (bus, IOTLB, CPU)
become overloaded. Aside from throughput reduction, one interesting
parameter is the effectiveness of the IOTLB. Here is a sample of SMMU
performance counters for a large ARM system with 2x200G NICs doing
bidirectional traffic:
=== DMA-FQ MODE (total throughput ~440Gbps)
25,757,488 smmuv3_pmcg_*/event=0x80/ IOTLB lookups
16,634,941 smmuv3_pmcg_*/event=0x81/ IOTLB misses
=== DMA_PMD on top of strict DMA (total throughput ~745Gbps)
22,693,218 smmuv3_pmcg_*/event=0x80/ IOTLB lookups
30,536 smmuv3_pmcg_*/event=0x81/ IOTLB misses
(not a mistake, also lookups went down despite the higher rate
because the tx side can use larger segments)
The 2MB mappings made IOTLB misses almost non existent, because
the working set is reduced by a factor of ~512.
Note, the code has more verbose comments than I would like.
Several of them are there to avoid Sashiko getting confused and
flagging false positives.
=== NOTES
[1] IOMMU IMPACT AND IOTLB THRASHING
This is documented in more detail in Documentation/core-api/dma-pmd.rst
but the compact version is below.
There are three main costs involved with using the IOMMU:
- CPU cost for dma map/unmap
- CPU cost and latency for synchronous IOTLB flush (required for strong security)
- IOTLB thrashing, shows up dramatically when the IO access pattern
exceeds the IOTLB size, and IOMMU page walks slow down bus activity
up to a point where we see over 30..60% throughput reduction just for
this reason. Some relevant references
https://lore.kernel.org/all/4b42f2eb-dc29-153e-ace9-5584ea2e5070@redhat.com/
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10280580/
Specifically, NICs access 4 or more pages per packet (descriptor queue,
completion queue, rx or tx header, one or more buffers), with very little
locality especially for data buffers and tx headers. Each queue has 1-4K
entries, a fast NIC uses 16 or more queues per direction, so just the
data buffers cover 32..128K pages. Using 4KB mappings makes the IOTLB
ineffective, as shown by the performance counters shown earlier.
[2] WHY NOT PAGEPOOL FOR TX AND EVERYTHING
There is some overlap between pagepool and dma_pmd_pool in that both
implement a cache, but pagepool is missing some important features:
- pagepool does not handle splitting 2M pages in smaller chunks, or
the mappings. That would need to be implemented (and is what dma_pmd_pool does)
- pagepool's provided pages cannot be cleanly managed by
get_page()/put_page(), which is what all the consumers of
skb_page_frag_refill(). Fixing that would require dozen of changes
with high risk of missing some paths.
- even for the receive path, a provider for pagepool has more restrictions
than alloc_pages-supplied memory, so if I wanted to wrap dma_pmd_pool
as a provider I would have needed changes to the provider interface.
- pagepool requires a device at allocation time, which is not known when
skb_page_frag_refill() runs, hence the IOMMU mappings cannot be handled
at allocation time.
- page_pool_alloc_pages assumes a single NAPI consumer with BH disabled,
whereas skb_page_frag_refill() runs from a preemptive context.
- and last not least, we need to handle the dma_alloc_attrs() allocations
which don't match pagepool API
Of course all of the above could be modified, but the complexity
would be similar to that of implementing dma_pmd_pool, and with a
huge risk of breaking existing functionality or missing some path.
[3] DRIVER CHANGES
Most drivers need no change for rings, transmit buffers, or receive buffers
(if they use pagepool). Tx headers are generally mapped on the fly and
that is enough to trigger IOTLB thrashing, so most drivers need a small
change to implement very cheap tx bounce buffers backed by DMA_PMD pages.
We could avoid the bounce buffers using DMA_PMD for the skb->head region,
but that memory is contiguous to skb_shinfo so it would be exposed to
IO device access, which may be undesirable for security.
Luigi Rizzo (22):
iommu/dma: introduce CONFIG_DMA_PMD and metadata table
iommu/dma: add DMA_PMD pool lifecycle and page recycle hook
mm: Add split_page_compound()
iommu/dma: add DMA_PMD pool block allocation
iommu/dma: Global cap and shrinker for DMA_PMD pool memory
iommu/dma: reserve a per-domain IOVA window for DMA_PMD pages
iommu/dma: release DMA_PMD domain mappings on domain teardown
iommu/dma: use per-domain IOVA window to map DMA_PMD memory
iommu/dma: Add DMA_PMD arena allocator
driver core: Add per-device dma_pmd_* sysfs attributes
dma-mapping: Use DMA_PMD arena for dma_alloc_attrs()
net/core: Use per-CPU DMA_PMD pools for skb_page_frag_refill()
net/core: Use DMA_PMD for page_pool memory
iommu/dma: Support decrypted and pinned DMA_PMD pages
iommu/dma: Add background page scrubber for DMA_PMD pools
iommu/dma: Add per-NUMA-node PMD page reservoir
net/gve: Use DMA_PMD memory for RX buffers
net/gve: Use DMA_PMD memory for tx header bounce buffers
net/mlx5e: Use DMA_PMD memory for tx header bounce buffers
net/idpf: Use DMA_PMD memory for tx header bounce buffers
net/bnxt: Use DMA_PMD memory for tx header bounce buffers
iommu/dma: Add DMA_PMD statistics and debugfs
Documentation/core-api/dma-pmd.rst | 320 ++++
Documentation/core-api/index.rst | 1 +
drivers/base/core.c | 51 +
drivers/iommu/Kconfig | 33 +
drivers/iommu/Makefile | 3 +
drivers/iommu/dma-iommu.c | 72 +-
drivers/iommu/dma-iommu.h | 8 +
drivers/iommu/dma-pmd-arena.c | 412 +++++
drivers/iommu/dma-pmd-kunit.c | 225 +++
drivers/iommu/dma-pmd-map.c | 531 ++++++
drivers/iommu/dma-pmd-meta.c | 403 +++++
drivers/iommu/dma-pmd-pool.c | 1538 +++++++++++++++++
drivers/iommu/dma-pmd-priv.h | 385 +++++
drivers/iommu/iommu.c | 2 +-
drivers/net/ethernet/broadcom/bnxt/bnxt.c | 45 +-
drivers/net/ethernet/broadcom/bnxt/bnxt.h | 2 +
drivers/net/ethernet/google/gve/gve.h | 5 +
drivers/net/ethernet/google/gve/gve_main.c | 5 +
drivers/net/ethernet/google/gve/gve_rx.c | 61 +-
drivers/net/ethernet/google/gve/gve_tx.c | 38 +-
drivers/net/ethernet/google/gve/gve_tx_dqo.c | 45 +-
.../ethernet/intel/idpf/idpf_singleq_txrx.c | 6 +-
drivers/net/ethernet/intel/idpf/idpf_txrx.c | 21 +-
drivers/net/ethernet/intel/idpf/idpf_txrx.h | 27 +-
drivers/net/ethernet/mellanox/mlx5/core/en.h | 2 +
.../net/ethernet/mellanox/mlx5/core/en/txrx.h | 4 +-
.../net/ethernet/mellanox/mlx5/core/en_main.c | 14 +
.../net/ethernet/mellanox/mlx5/core/en_tx.c | 28 +-
include/linux/device.h | 5 +
include/linux/dma-pmd.h | 198 +++
include/linux/mm.h | 2 +
include/net/libeth/tx.h | 5 +-
include/net/page_pool/types.h | 3 +
include/net/sock.h | 3 +
kernel/dma/direct.c | 6 +-
kernel/dma/mapping.c | 39 +-
mm/Kconfig.debug | 11 +
mm/Makefile | 1 +
mm/page_alloc.c | 89 +
mm/split_page_compound_kunit.c | 113 ++
net/core/page_pool.c | 61 +-
net/core/sock.c | 86 +-
net/core/sysctl_net_core.c | 7 +
43 files changed, 4835 insertions(+), 81 deletions(-)
create mode 100644 Documentation/core-api/dma-pmd.rst
create mode 100644 drivers/iommu/dma-pmd-arena.c
create mode 100644 drivers/iommu/dma-pmd-kunit.c
create mode 100644 drivers/iommu/dma-pmd-map.c
create mode 100644 drivers/iommu/dma-pmd-meta.c
create mode 100644 drivers/iommu/dma-pmd-pool.c
create mode 100644 drivers/iommu/dma-pmd-priv.h
create mode 100644 include/linux/dma-pmd.h
create mode 100644 mm/split_page_compound_kunit.c
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:44 ` Randy Dunlap
2026-10-03 21:22 ` [RFC: DMA_PMD 02/22] iommu/dma: add DMA_PMD pool lifecycle and page recycle hook Luigi Rizzo
` (20 subsequent siblings)
21 siblings, 1 reply; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
DMA_PMD is a new subsystem to back IO buffers with physically contiguous
PMD_SIZE pages, allowing the use of larger leaves in the IOMMU. This
significantly reduces IOTLB pressure, with great performance benefits. It
support very common cases (2MB PMD, !PREEMPT_RT). The 2MB backing pages
are dynamically allocated and released. DMA_PMD overhead is 128B of
metadata per physical page (64KB per 1GB, or 0.01% of physical RAM).
Add a CONFIG_DMA_PMD option to enable it only when applicable, and
introduce a sparse per-PMD metadata table to store the necessary side
information efficiently (similar in principle to pageblock_flags).
The table is initialized on first use, reserving KVA across the physical
address space up to iomem_resource.end (capped at MAX_PHYSMEM_BITS,
64KB of KVA per 1GB of address space) along with a chunk-populated
bitmap (1 bit per 64MB chunk) for fast lockless lookups. Only chunks
covering online RAM are backed by physical pages at init; memory
hotplugged later within the reserved range is populated on demand when
allocated by a DMA_PMD pool, and safely declined otherwise.
Also add a built-in KUnit test suite (CONFIG_DMA_PMD_META_KUNIT_TEST)
exercising initialization idempotency, PFN/physical round-trip lookups,
out-of-bounds physical address handling, and DMA_PMD membership toggling.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
Documentation/core-api/dma-pmd.rst | 320 +++++++++++++++++++++++++
Documentation/core-api/index.rst | 1 +
drivers/iommu/Kconfig | 32 +++
drivers/iommu/Makefile | 2 +
drivers/iommu/dma-pmd-kunit.c | 102 ++++++++
drivers/iommu/dma-pmd-meta.c | 365 +++++++++++++++++++++++++++++
drivers/iommu/dma-pmd-priv.h | 56 +++++
include/linux/dma-pmd.h | 36 +++
8 files changed, 914 insertions(+)
create mode 100644 Documentation/core-api/dma-pmd.rst
create mode 100644 drivers/iommu/dma-pmd-kunit.c
create mode 100644 drivers/iommu/dma-pmd-meta.c
create mode 100644 drivers/iommu/dma-pmd-priv.h
create mode 100644 include/linux/dma-pmd.h
diff --git a/Documentation/core-api/dma-pmd.rst b/Documentation/core-api/dma-pmd.rst
new file mode 100644
index 0000000000000..1d113b8c15876
--- /dev/null
+++ b/Documentation/core-api/dma-pmd.rst
@@ -0,0 +1,320 @@
+.. SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause
+
+=========================================================
+DMA_PMD: PMD_SIZE IOMMU mappings for DMA-coherent devices
+=========================================================
+
+Overview
+========
+
+DMA_PMD is a transparent enhancement of the current IOMMU modes, that
+gives the performance of identity mode, and like strict IOMMU mode
+(DMA), guarantees that memory will go back to the buddy allocator only
+when all IOMMU mappings have been removed and IOTLB flushed.
+
+The original design aimed at removing IOTLB thrashing, but its mechanism
+also give security guarantees comparable to strict IOMMU, and remove the
+bounce buffer overhead from Confidential Computing and preemptible VMs.
+
+Background and motivations
+==========================
+
+I/O devices typically access at least 3-4 different memory regions on
+each transaction (network packet or disk request):
+
+- command, completion and buffer queues
+- optionally, header or metadata buffers (e.g., NIC packet headers, NVMe SGL...)
+- data buffers (TX socket buffers, NIC RX buffers, disk I/O buffers)
+
+Enabling the IOMMU impacts performance for the following reasons:
+
+- CPU cost to install/remove IOMMU Page Table Entries (PTEs)
+ (``dma_map_*()``, ``dma_unmap_*()``)
+- CPU/system overhead to flush the IOTLB, ie remove stale PTEs that could give
+ access to sensitive data or code when memory is recycled for other purposes.
+- Higher DMA read/write latency from IOMMU page walks when the working set
+ exceeds the IOTLB size, backpressuring the bus via flow control.
+ This is one of the biggest bottleneck for high speed devices:
+ typical network devices have very little IOVA locality, and that makes the IOTLB
+ almost completely ineffective.
+
+There are known methods to mitigate these problems:
+
+1. Recycling I/O buffers and their IOMMU mappings removes most of the DMA map/unmap
+ costs. This is common practice for command, completion and buffer
+ queues, as well as disk I/O and network receive buffers. Network
+ TX buffers are a noticeable exception. They normally come from
+ ``alloc_pages()`` calls unaware of the leaf device, and require
+ mapping/unmapping on each use.
+
+2. Lazily flushing the IOTLB (as implemented by DMA-FQ) amortizes a
+ very expensive operation (up to several microseconds to wait for
+ completion, necessary for security), but the way it is implemented
+ in DMA-FQ compromises on security: IOTLB flush requests are run
+ periodically in batches, but the underlying memory is recycled without
+ waiting. During that window, the physical memory is still accessible,
+ opening the door to exfiltration or corruption of sensitive data/code.
+
+3. The set of PTEs in active use can be significantly reduced by mapping larger
+ blocks, e.g. 2MB (called PMD) instead of the usual 4KB (PTE). This is not
+ currently pursued by the kernel.
+
+DMA_PMD addresses all the above with a combination of simple, known concepts,
+implemented in a way that integrates smoothly with the kernel and drivers:
+
+- intercept calls to the small set of functions used to allocate, map, unmap, and free
+ device-accessible memory (rings and buffers)
+- return memory backed by physically contiguous 2MB pages (PMD) and 2MB IOMMU
+ mappings
+- heavily recycle memory and mappings
+- wait until the IOTLB has been flushed before returning unused DMA_PMD memory
+ to the system pool
+
+Execution is what makes DMA_PMD practical: the implementation intercepts
+a small set of kernel APIs so that device will inherit the benefits with
+little if any changes, and the design is such that these intercepts have
+minimal impact on the original functions.
+
+Furthermore, since DMA_PMD intercepts allocations of device-accessible
+memory, it gives two important benefits:
+- transparent pinning/unpining, required for preemptible VMs
+- transparent configuration of unencrypted mode, required for confidential computing
+
+Limitations
+===========
+
+DMA_PMD is currently limited to systems with ``PMD_SIZE`` equal to 2MB, which
+is the vast majority of existing platforms. It could be extended to larger
+pages without much effort. This is enforced at compile time.
+
+The implementation assumes non-preemptible spinlocks, so DMA_PMD is not
+compatible with ``PREEMPT_RT``. This is enforced at compile time.
+
+DMA_PMD can only be used by DMA-coherent devices. This is enforced at runtime.
+
+IOMMU mappings for data buffers are always ``DMA_BIDIRECTIONAL``. This
+simplifies the handling of e.g. forwarding between different network
+interfaces.
+
+IOMMU mappings for queues are also ``DMA_BIDIRECTIONAL``, though stricter modes
+can be implemented trivially.
+
+Implementation details
+======================
+
+Most of the code size and design complexity relates to the handling
+of exceptional events (device teardown, memory hotplug) and to
+the safe release of pages through deferred tasks, RCU and grace
+periods.
+
+The implementation lives in ``drivers/iommu/dma-pmd-*.c`` and
+``include/linux/dma-pmd.h``, and relies on the components described below.
+We use the name "DMA_PMD" to refer to the second-level physically contiguous
+pages (typically 2MB) used for I/O memory.
+
+The set of functions that need to be intercepted is relatively small:
+
+- ``page_pool_dev_alloc*()``, ``skb_page_frag_refill()``, ``__free_pages()``
+- ``dma_alloc_attrs()``, ``dma_map_*()``, ``dma_unmap_*()``
+
+and most devices will automatically inherit the benefits.
+
+Requirements (for performance and functionality)
+================================================
+
+- efficient identification of DMA_PMD pages, based on either their Physical
+ Address (PA) or their I/O Virtual Address (IOVA)
+- fast access to DMA_PMD page metadata
+- fast allocation under highly concurrent usage
+- efficient DMA map/unmap
+- comprehensive lifetime management
+- support for confidential computing (decrypted DMA buffers)
+- support for preemptible VMs (pinned DMA buffers)
+
+Internal mechanisms
+===================
+
+We use the following components:
+
+- **DMA_PMD metadata** (``struct dma_pmd_meta``):
+
+ Identification of DMA_PMD pages could be done in principle with a flag in
+ ``struct page``, but that would not solve the management of per-page
+ metadata.
+
+ DMA_PMD addresses both requirements without depending on ``struct page``
+ using an idea similar to ``pageblock_flags``. Each 2MB page used as a
+ buffer requires only 256B of metadata, or 0.01% overhead. For metadata,
+ we reserve a sparse virtual memory region covering of approximate
+ size ``max_pfn * 256 / PMD_SIZE``, so the metadata can be accessed
+ directly using the 2MB page index. Metadata pages are populated on
+ demand, with a compact bitmap (1 bit per 32MB chunk, or 4KB per 1TB
+ of address space) tracking which metadata pages are backed, allowing
+ fast lockless validation and direct array indexing. Memory hotplugged
+ later is also supported.
+
+ ``struct dma_pmd_meta`` has flags to mark the page status (used as
+ DMA_PMD buffer, decrypted for confidential computing, or pinned against
+ memory compaction), and tracks which subpages are free, which domains
+ have an active IOMMU mapping for the page, and holds list/RCU linkage
+ to manage its lifetime.
+
+- **Per-domain reserved DMA_PMD IOVA ranges**:
+
+ Another key requirement is to quickly resolve the PA-IOVA mapping for a given domain. Since
+ a DMA_PMD page can be mapped in multiple domains (e.g. a network buffer used
+ to route packets between different NICs), it would be too expensive to manage random
+ PA-IOVA mappings created on the fly.
+ DMA_PMD reserves on each domain an IOVA range
+ covering the physical memory range plus an extra region, allowing a
+ fixed PA-IOVA offset, possibly different for each domain.
+ The IOVA allocation happens the first time a domain is used, hence a
+ IOVA within the reserved range will uniquely identify a DMA_PMD page.
+
+- **alloc_pages() compatible dma_pmd_pool allocator for I/O buffers**:
+
+ DMA_PMD implements a dma_pmd_pool object for fixed-order page allocations
+ that is a natural fit for the allocation of NIC receive buffers and tx
+ socket buffers. The API is similar to ``alloc_pages()`` (in fact, it is an
+ almost direct replacement) and dma_pmd_pools are instantiated as follow:
+ - per NIC-receive-queue (or page_pool) to provide order-0 pages to each queue
+ - per-CPU to provide order-0 and order-3 pages to refill tx socket buffers.
+
+- **dma_pmd_arena allocator to back dma_alloc_attrs()**
+
+ Devices use ``dma_alloc_attrs()`` or variants for long-lived, variable size
+ allocations to back device queues and header buffers. dma_pmd_arena
+ has the same functions: it both allocates and maps memory in the
+ reserved IOVA, and is used exclusively within dma_alloc_attrs() to
+ handle suitable requests (4K or larger, GFP_KERNEL, for devices that
+ specifically enable DMA_PMD)
+
+
+- **Fine grained control**
+
+ Both for experimentation and production use, it is useful to have
+ some form of control on which devices want to use DMA_PMD and for
+ what. DMA_PMD exposes the following controls. Keep in mind that full
+ performance can only be achieved if all regions are mapped using DMA_PMD.
+
+ - /sys/devices/*/*/dma_pmd_{rings,rxbuf,tx_hdrs} per-device entries
+ that can be used within device drivers to decide whether to use DMA_PMD for each of its regions
+
+ - /proc/sys/net/core/tx_enable_dma_pmd controls whether tx socket buffers should use DMA_PMD
+
+- **Safe release of memory to the system**
+
+ A key feature of DMA_PMD is that memory is not released to the system
+ while there are active IOMMU mappings. Release of buffers is managed
+ as follows
+
+ 1. on ``__free_pages()`` or ``dma_free_attrs()``, blocks are returned to DMA_PMD,
+ IOMMU mappings are still active, and they can be recycled
+
+ 2. when all blocks in a DMA_PMD page are unused, the page is moved to an
+ "idle" list, still with active mappings and available for reuse
+
+ 3. upon memory pressure or when above some global threshold, a shrinker
+ thread collects idle pages from all pools and starts removing the iommu
+ mappings and issues a synchronous IOTLB flush. The pages are not usable for allocations
+ but not returned to the pool yet
+
+ 4. once the previous step is complete, the system schedules ``call_rcu()``
+ and waits for an RCU grace period (and re-encrypts the page if it was
+ decrypted) before returning the 2MB page to the buddy allocator.
+
+- **Handling domain destruction**
+
+ When a device/domain is destroyed, all I/O issued by the device must have
+ been quiesced and memory released via ``__free_pages()`` or ``dma_free*()``.
+ DMA_PMD hooks into the domain destructor and removes all existing mappings
+ for the disappearing domain and synchronously flushes the IOTLB, similarly
+ to steps #3 and #4 shown before.
+
+- **Allocation and compound splitting**:
+
+ A pool that needs memory calls ``dma_pmd_add_page()`` to allocate a
+ physically contiguous 2MB page on the calling CPU's NUMA node, and
+ ``split_page_compound()`` to split it into independent blocks of
+ ``pool->order``. These blocks can be independently passed through the
+ networking stack.
+
+Modified kernel APIs and runtime controls
+=========================================
+
+DMA_PMD integrates transparently into existing kernel memory, DMA, and
+networking APIs. Allocation of DMA_PMD memory is opt-in via per-device sysfs
+attributes or sysctls, while mapping, unmapping, and freeing automatically
+detect whether a buffer or IOVA belongs to DMA_PMD:
+
+1. **Streaming DMA map and unmap** (``dma_map_page_attrs()``,
+ ``dma_map_single_attrs()``, ``dma_map_sg_attrs()`` and
+ ``dma_unmap_page_attrs()``, ``dma_unmap_single_attrs()``,
+ ``dma_unmap_sg_attrs()`` via ``drivers/iommu/dma-iommu.c`` and
+ ``kernel/dma/direct.c``):
+
+ - On map, any buffer belonging to a DMA_PMD page (``dma_is_pmd_phys(phys)``)
+ is mapped via the domain's reserved ``DMA_PMD`` IOVA window with a simple
+ addition, typically reusing existing mappings.
+ Under ``dma-direct``, decrypted or pinned ``DMA_PMD`` pages
+ (``dma_is_pmd_direct(phys)``) can be mapped directly without ``swiotlb`` bounce.
+
+ On unmap, IOVAs can be identified as DMA_PMD buffers with a simple range check
+ and that results in a no-operation.
+
+2. **Coherent DMA allocation and free** (``dma_alloc_attrs()`` /
+ ``dma_alloc_coherent()`` and ``dma_free_attrs()`` / ``dma_free_coherent()``
+ in ``kernel/dma/mapping.c``):
+
+ - DMA_PMD is enabled per device via ``/sys/devices/.../dma_pmd_rings``
+ (``dev->dma_pmd_rings``). When set, sleepable ``GFP_KERNEL`` coherent
+ allocations are rounded up to 4KB and carved out of DMA_PMD buffers
+ ``dma_pmd_arena`` and mapped accordingly.
+ ``dma_free_attrs()`` automatically detects arena IOVAs via
+ ``dma_pmd_free()`` and returns the sub-blocks to the arena.
+
+3. **Page allocator release** (``__free_pages()`` / ``put_page()`` via
+ ``__free_pages_prepare()`` in ``mm/page_alloc.c``):
+
+ - ``dma_pmd_free_page()`` checks
+ ``dma_is_pmd_page(page_to_pfn(page))`` and recycles DMA_PMD subpages back to
+ their owning ``dma_pmd_pool`` instead of releasing them to the buddy
+ allocator while their 2MB IOMMU mappings remain active.
+
+4. **IOMMU domain teardown** (``iommu_put_dma_cookie()`` in
+ ``drivers/iommu/dma-iommu.c``):
+
+ - A call to ``dma_pmd_domain_release()`` before freeing a
+ domain's IOVA cookie and page tables clears cached domain state across
+ all pools. Existing IOMMU mappings and IOTLB flush has already happened.
+
+5. **Networking ``page_pool``**
+
+ - DMA_PMD is enabled per device via ``/sys/devices/.../dma_pmd_rxbuf``
+ (``dev->dma_pmd_rxbuf``). Any NIC driver using ``page_pool`` with this
+ flag set will back its RX buffers with a per-``page_pool``
+ ``dma_pmd_pool``, and opportunistically replaces plain 4KB pages when pooled 2MB
+ blocks are available.
+
+6. **Socket TX page fragments** (``skb_page_frag_refill()`` in
+ ``net/core/sock.c``):
+
+ - When enabled globally via the ``net.core.tx_enable_dma_pmd`` sysctl
+ socket TX fragments are allocated from per-CPU ``dma_pmd_pool`` instances.
+
+7. **NIC driver RX pools and TX header bounce buffers**
+
+ - **RX buffers:** Controlled per device by ``/sys/devices/.../dma_pmd_rxbuf``
+ (``dev->dma_pmd_rxbuf``) for queue modes managing their own page rings
+ (``idpf``, ``gve`` GQI-RDA, ``gq``).
+ - **TX headers:** Controlled per device by
+ ``/sys/devices/.../dma_pmd_tx_hdrs`` (``dev->dma_pmd_tx_hdrs``),
+ copying linear packet headers into pre-allocated coherent bounce buffers
+ (backed by ``dma_pmd_arena`` when ``dma_pmd_rings`` is also enabled).
+
+8. **Global pool limit and observability:**
+
+ - ``/sys/module/kernel/parameters/dma_pmd_max_pages``: Global cap on 2MB
+ pages held across all pools (defaults to 1/8th of physical RAM).
+ - ``/sys/kernel/debug/dma_pmd/pools``: Debugfs summary of global and per-pool
+ allocation, mapping, and fallback counters.
diff --git a/Documentation/core-api/index.rst b/Documentation/core-api/index.rst
index 92f91c6a0d79d..49608d2b7509b 100644
--- a/Documentation/core-api/index.rst
+++ b/Documentation/core-api/index.rst
@@ -113,6 +113,7 @@ more memory-management documentation in Documentation/mm/index.rst.
dma-attributes
dma-isa-lpc
swiotlb
+ dma-pmd
mm-api
cgroup
genalloc
diff --git a/drivers/iommu/Kconfig b/drivers/iommu/Kconfig
index 6e07bd69467a3..1bb347fd9da4a 100644
--- a/drivers/iommu/Kconfig
+++ b/drivers/iommu/Kconfig
@@ -158,6 +158,38 @@ config IOMMU_DMA
select NEED_SG_DMA_LENGTH
select NEED_SG_DMA_FLAGS if SWIOTLB
+# PMD_SIZE IOMMU backing for DMA-coherent devices
+config DMA_PMD
+ bool "PMD_SIZE IOMMU backing for DMA-coherent devices"
+ depends on IOMMU_DMA
+ depends on X86_64 || (ARM64 && ARM64_4K_PAGES)
+ depends on !PREEMPT_RT
+ default y
+ help
+ Hand out DMA buffers carved out of PMD_SIZE physically contiguous
+ blocks that the IOMMU maps with a single leaf PTE. Mapping,
+ unmapping and IOTLB flushing are done lazily to reduce CPU overhead.
+ The large mappings greatly reduce IOTLB usage. As in strict iommu
+ mode, pages are returned to the buddy allocator only when any existing
+ mappings have been removed and IOTLB flushed.
+ Costs 128KB of metadata per GB of DMA_PMD buffers, allocated on demand.
+
+ The config option enforces requirements (2MB PMD and !PREEMPT_RT) at compile
+ time. Other constraints (e.g. dma-coherent devices) are verified at runtime.
+
+ If unsure, say Y.
+
+config DMA_PMD_META_KUNIT_TEST
+ bool "KUnit test for DMA_PMD metadata table (built-in)" if !KUNIT_ALL_TESTS
+ depends on DMA_PMD && KUNIT=y
+ default KUNIT_ALL_TESTS
+ help
+ Builds KUnit unit tests for the sparse, per-PMD-frame metadata table,
+ PMD page pools, and coherent DMA arenas in drivers/iommu/dma-pmd*.c,
+ verifying metadata lookup, pool recycling, and arena block allocation.
+
+ If unsure, say N.
+
# Shared Virtual Addressing
config IOMMU_SVA
select IOMMU_MM_DATA
diff --git a/drivers/iommu/Makefile b/drivers/iommu/Makefile
index 2f05725eaab18..2ad9b2eefd741 100644
--- a/drivers/iommu/Makefile
+++ b/drivers/iommu/Makefile
@@ -11,6 +11,8 @@ obj-$(CONFIG_IOMMU_API) += iommu-traces.o
obj-$(CONFIG_IOMMU_API) += iommu-sysfs.o
obj-$(CONFIG_IOMMU_DEBUGFS) += iommu-debugfs.o
obj-$(CONFIG_IOMMU_DMA) += dma-iommu.o
+obj-$(CONFIG_DMA_PMD) += dma-pmd-meta.o
+obj-$(CONFIG_DMA_PMD_META_KUNIT_TEST) += dma-pmd-kunit.o
obj-$(CONFIG_IOMMU_IO_PGTABLE) += io-pgtable.o
obj-$(CONFIG_IOMMU_IO_PGTABLE_ARMV7S) += io-pgtable-arm-v7s.o
obj-$(CONFIG_IOMMU_IO_PGTABLE_LPAE) += io-pgtable-arm.o
diff --git a/drivers/iommu/dma-pmd-kunit.c b/drivers/iommu/dma-pmd-kunit.c
new file mode 100644
index 0000000000000..f80527eaf2e84
--- /dev/null
+++ b/drivers/iommu/dma-pmd-kunit.c
@@ -0,0 +1,102 @@
+// SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause
+/*
+ * KUnit tests for the opaque struct dma_pmd_meta table API.
+ */
+#include <kunit/test.h>
+#include <linux/gfp.h>
+#include <linux/dma-pmd.h>
+#include <linux/mm.h>
+
+#include "dma-pmd-priv.h"
+
+static void test_meta_init_and_roundtrip(struct kunit *test)
+{
+ struct dma_pmd_meta *m_pfn, *m_phys;
+ unsigned long pfn, base_pfn;
+ phys_addr_t pa, base_pa;
+ struct page *page;
+
+ KUNIT_ASSERT_EQ(test, dma_pmd_meta_init(), 0);
+ /* Second call must be idempotent. */
+ KUNIT_ASSERT_EQ(test, dma_pmd_meta_init(), 0);
+
+ page = alloc_page(GFP_KERNEL);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+
+ pfn = page_to_pfn(page);
+ base_pfn = ALIGN_DOWN(pfn, 1UL << PMD_ORDER);
+ pa = page_to_phys(page);
+ base_pa = ALIGN_DOWN(pa, PMD_SIZE);
+
+ m_pfn = dma_pmd_meta_of_pfn(pfn);
+ m_phys = dma_pmd_meta_from_phys(pa);
+
+ KUNIT_EXPECT_PTR_EQ(test, m_pfn, m_phys);
+ KUNIT_EXPECT_EQ(test, dma_pmd_meta_to_pfn(m_pfn), base_pfn);
+ KUNIT_EXPECT_EQ(test, dma_pmd_meta_to_phys(m_phys), base_pa);
+
+ __free_page(page);
+}
+
+static void test_meta_invalid_phys(struct kunit *test)
+{
+ struct dma_pmd_meta *m;
+
+ KUNIT_ASSERT_EQ(test, dma_pmd_meta_init(), 0);
+
+ m = dma_pmd_meta_from_phys(PHYS_ADDR_MAX);
+ KUNIT_ASSERT_NOT_NULL(test, m);
+ KUNIT_EXPECT_EQ(test, dma_pmd_meta_to_phys(m), PHYS_ADDR_MAX);
+ KUNIT_EXPECT_FALSE(test, dma_is_pmd_page(ULONG_MAX >> PAGE_SHIFT));
+}
+
+static void dma_pmd_meta_set_pooled(struct dma_pmd_meta *m, bool pooled)
+{
+ WRITE_ONCE(m->pooled, pooled);
+}
+
+static void test_meta_pooled_toggle(struct kunit *test)
+{
+ unsigned long pfn, base_pfn;
+ struct dma_pmd_meta *m;
+ struct page *page;
+
+ KUNIT_ASSERT_EQ(test, dma_pmd_meta_init(), 0);
+
+ page = alloc_pages(GFP_KERNEL, PMD_ORDER);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+
+ pfn = page_to_pfn(page);
+ base_pfn = ALIGN_DOWN(pfn, 1UL << PMD_ORDER);
+ /* The buddy allocator returns naturally aligned blocks. */
+ KUNIT_EXPECT_EQ(test, pfn, base_pfn);
+ m = dma_pmd_meta_of_pfn(pfn);
+
+ KUNIT_EXPECT_FALSE(test, dma_is_pmd_page(pfn));
+
+ dma_pmd_meta_set_pooled(m, true);
+ KUNIT_EXPECT_TRUE(test, dma_is_pmd_page(base_pfn));
+ KUNIT_EXPECT_TRUE(test, dma_is_pmd_page(pfn));
+ KUNIT_EXPECT_TRUE(test, dma_is_pmd_page(base_pfn + (1UL << PMD_ORDER) - 1));
+
+ dma_pmd_meta_set_pooled(m, false);
+ KUNIT_EXPECT_FALSE(test, dma_is_pmd_page(pfn));
+
+ __free_pages(page, PMD_ORDER);
+}
+
+static struct kunit_case dma_pmd_meta_test_cases[] = {
+ KUNIT_CASE(test_meta_init_and_roundtrip),
+ KUNIT_CASE(test_meta_invalid_phys),
+ KUNIT_CASE(test_meta_pooled_toggle),
+ {}
+};
+
+static struct kunit_suite dma_pmd_meta_test_suite = {
+ .name = "dma_pmd_meta",
+ .test_cases = dma_pmd_meta_test_cases,
+};
+
+kunit_test_suite(dma_pmd_meta_test_suite);
+MODULE_DESCRIPTION("KUnit tests for struct dma_pmd_meta table");
+MODULE_LICENSE("Dual BSD/GPL");
diff --git a/drivers/iommu/dma-pmd-meta.c b/drivers/iommu/dma-pmd-meta.c
new file mode 100644
index 0000000000000..fcaf5b121e66a
--- /dev/null
+++ b/drivers/iommu/dma-pmd-meta.c
@@ -0,0 +1,365 @@
+// SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause
+/*
+ * DMA_PMD sparse per-PMD-frame metadata table.
+ *
+ * See Documentation/core-api/dma-pmd.rst for the architecture overview.
+ */
+
+#include <linux/bitmap.h>
+#include <linux/cache.h>
+#include <linux/cacheflush.h>
+#include <linux/dma-pmd.h>
+#include <linux/export.h>
+#include <linux/gfp.h>
+#include <linux/ioport.h>
+#include <linux/list.h>
+#include <linux/memblock.h>
+#include <linux/mm.h>
+#include <linux/mutex.h>
+#include <linux/pgtable.h>
+#include <linux/spinlock.h>
+#include <linux/vmalloc.h>
+
+#include "dma-pmd-priv.h"
+
+/*
+ * Each PAGE_SIZE (4KB) metadata page holds (PAGE_SIZE >> DMA_PMD_META_SHIFT)
+ * struct dma_pmd_meta entries (32 entries of 128B, or 16 entries of 256B with
+ * spinlock debugging), each covering one PMD_SIZE (2MB) physical frame.
+ * One metadata page therefore covers a chunk of DMA_PMD_CHUNK_PAGES 4KB pages
+ * (64MB normally, or 32MB with spinlock debugging), i.e. order
+ * DMA_PMD_CHUNK_ORDER.
+ */
+#define DMA_PMD_CHUNK_ORDER (PMD_ORDER + PAGE_SHIFT - DMA_PMD_META_SHIFT)
+#define DMA_PMD_CHUNK_PAGES BIT(DMA_PMD_CHUNK_ORDER)
+
+static_assert(PAGE_SIZE >= DMA_PMD_META_SIZE);
+static_assert(PMD_SIZE == SZ_2M);
+static_assert(PMD_ORDER <= MAX_PAGE_ORDER);
+static_assert(PMD_SHIFT <= SUBSECTION_SHIFT);
+
+void *dma_pmd_meta_array __read_mostly;
+EXPORT_SYMBOL(dma_pmd_meta_array);
+unsigned long dma_pmd_meta_nframes __read_mostly;
+static unsigned long *dma_pmd_chunk_bitmap __read_mostly;
+
+static struct dma_pmd_meta dma_pmd_meta_nil;
+
+static unsigned long dma_pmd_meta_pages;
+static DEFINE_MUTEX(dma_pmd_meta_mutex);
+
+static __always_inline struct dma_pmd_meta *__dma_pmd_meta_of_pfn(unsigned long pfn)
+{
+ struct dma_pmd_meta *array = dma_pmd_meta_base();
+ unsigned long frame = pfn >> PMD_ORDER;
+
+ if (unlikely(!array || frame >= dma_pmd_meta_nframes ||
+ !test_bit_acquire(pfn >> DMA_PMD_CHUNK_ORDER, dma_pmd_chunk_bitmap)))
+ return &dma_pmd_meta_nil;
+
+ return &array[frame];
+}
+
+bool __dma_is_pmd_page(unsigned long pfn)
+{
+ return READ_ONCE(__dma_pmd_meta_of_pfn(pfn)->pooled);
+}
+EXPORT_SYMBOL(__dma_is_pmd_page);
+
+/**
+ * dma_pmd_meta_of_pfn - Metadata for a PFN known to be used by DMA_PMD.
+ * @pfn: PFN the caller already holds a struct page for
+ */
+struct dma_pmd_meta *dma_pmd_meta_of_pfn(unsigned long pfn)
+{
+ return __dma_pmd_meta_of_pfn(pfn);
+}
+EXPORT_SYMBOL(dma_pmd_meta_of_pfn);
+
+/**
+ * dma_pmd_meta_from_phys - Metadata for the PMD frame containing @pa
+ * @pa: Any physical address
+ *
+ * Safe against MMIO above max_pfn and PFNs inside physical holes: those
+ * resolve to the shared sink entry.
+ *
+ * Return: Pointer to metadata entry (never NULL).
+ */
+struct dma_pmd_meta *dma_pmd_meta_from_phys(phys_addr_t pa)
+{
+ return __dma_pmd_meta_of_pfn(PHYS_PFN(pa));
+}
+EXPORT_SYMBOL(dma_pmd_meta_from_phys);
+
+/**
+ * dma_pmd_meta_to_pfn - Base PFN of @m's PMD frame
+ * @m: Entry in @dma_pmd_meta_array
+ *
+ * Return: PMD-aligned base PFN, or ULONG_MAX for the sink.
+ */
+unsigned long dma_pmd_meta_to_pfn(const struct dma_pmd_meta *m)
+{
+ if (unlikely(m == &dma_pmd_meta_nil))
+ return ULONG_MAX;
+
+ return (unsigned long)(m - dma_pmd_meta_base()) << PMD_ORDER;
+}
+EXPORT_SYMBOL(dma_pmd_meta_to_pfn);
+
+/**
+ * dma_pmd_meta_to_phys - Base physical address of @m's PMD frame
+ * @m: Entry returned by dma_pmd_meta_of_pfn() or dma_pmd_meta_from_phys()
+ *
+ * Return: PMD-aligned physical address, or PHYS_ADDR_MAX for the sink.
+ */
+phys_addr_t dma_pmd_meta_to_phys(const struct dma_pmd_meta *m)
+{
+ if (unlikely(m == &dma_pmd_meta_nil))
+ return PHYS_ADDR_MAX;
+
+ return (phys_addr_t)(m - dma_pmd_meta_base()) << PMD_SHIFT;
+}
+EXPORT_SYMBOL(dma_pmd_meta_to_phys);
+
+/*
+ * Install one preallocated page. The page is allocated by the caller rather
+ * than here because apply_to_page_range() runs this callback under lazy-MMU
+ * mode with the pte level pinned, which is not a context to allocate from.
+ *
+ * @data points at the caller's page pointer and is cleared once the page has
+ * been consumed, so the caller can free it if it was not needed.
+ */
+static int dma_pmd_meta_set_pte(pte_t *ptep, unsigned long addr, void *data)
+{
+ struct page **pagep = data;
+ pte_t pte;
+
+ if (!pte_none(ptep_get(ptep)))
+ return 0;
+
+ pte = pfn_pte(page_to_pfn(*pagep), PAGE_KERNEL);
+
+ spin_lock(&init_mm.page_table_lock);
+ if (likely(pte_none(ptep_get(ptep)))) {
+ set_pte_at(&init_mm, addr, ptep, pte);
+ *pagep = NULL;
+ }
+ spin_unlock(&init_mm.page_table_lock);
+
+ return 0;
+}
+
+/* True iff @pfn is backed by online system RAM (excludes holes and ZONE_DEVICE). */
+static inline bool dma_pmd_pfn_online(unsigned long pfn)
+{
+ struct mem_section *ms;
+
+ if (unlikely(!pfn_valid(pfn)))
+ return false;
+
+ ms = __pfn_to_section(pfn);
+ if (unlikely(!online_section(ms)))
+ return false;
+
+ if (unlikely(online_device_section(ms) && is_zone_device_page(pfn_to_page(pfn))))
+ return false;
+
+ return true;
+}
+
+/**
+ * dma_pmd_meta_populate - back the array for [@start_pfn, @end_pfn)
+ * @array: base virtual address of the reservation
+ * @nframes: number of valid PMD frames in the reservation
+ * @start_pfn: first PFN of the range
+ * @end_pfn: one past the last PFN of the range
+ * @force: if true, populate every chunk unconditionally (hotplug / arena)
+ *
+ * Maps and zeroes one page of the array per chunk of the range that contains
+ * at least one online RAM PFN (or unconditionally when @force is set), and
+ * skips chunks that are already mapped. A chunk
+ * is DMA_PMD_CHUNK_PAGES of physical address space: 64MB normally, 32MB when
+ * spinlock debugging doubles the entry size.
+ * Locking: caller must hold @dma_pmd_meta_mutex.
+ *
+ * Return: 0, or -ENOMEM with the range partially backed.
+ */
+static int dma_pmd_meta_populate(struct dma_pmd_meta *array, unsigned long nframes,
+ unsigned long start_pfn, unsigned long end_pfn, bool force)
+{
+ unsigned long chunk, last, pfn;
+
+ lockdep_assert_held(&dma_pmd_meta_mutex);
+
+ end_pfn = min(end_pfn, nframes << PMD_ORDER);
+ if (start_pfn >= end_pfn)
+ return 0;
+
+ chunk = start_pfn >> DMA_PMD_CHUNK_ORDER;
+ last = (end_pfn - 1) >> DMA_PMD_CHUNK_ORDER;
+
+ for (; chunk <= last; chunk++) {
+ unsigned long addr, base = chunk << DMA_PMD_CHUNK_ORDER;
+ int ret, nid = NUMA_NO_NODE;
+ bool has_valid = false;
+ struct page *page;
+
+ addr = (unsigned long)array + (chunk << PAGE_SHIFT);
+ if (vmalloc_to_page((void *)addr)) {
+ set_bit(chunk, dma_pmd_chunk_bitmap);
+ continue; /* already backed */
+ }
+
+ if (force) {
+ has_valid = true;
+ } else {
+ for (pfn = base; pfn < base + DMA_PMD_CHUNK_PAGES;
+ pfn += 1UL << PMD_ORDER) {
+ if (dma_pmd_pfn_online(pfn)) {
+ has_valid = true;
+ nid = page_to_nid(pfn_to_page(pfn));
+ if (nid != NUMA_NO_NODE && node_online(nid))
+ break;
+ nid = NUMA_NO_NODE;
+ }
+ }
+ }
+ if (!has_valid)
+ continue; /* pure hole, leave it unmapped */
+
+ page = alloc_pages_node(nid, GFP_KERNEL | __GFP_ZERO, 0);
+ if (!page)
+ return -ENOMEM;
+
+ ret = apply_to_page_range(&init_mm, addr, PAGE_SIZE,
+ dma_pmd_meta_set_pte, &page);
+ if (ret) {
+ __free_page(page);
+ return ret;
+ }
+ if (page) {
+ /*
+ * Unreachable while dma_pmd_meta_mutex serialises
+ * every install, but kept so that the defensive
+ * pte_none() re-test in dma_pmd_meta_set_pte() can
+ * never leak the page it declined to consume.
+ */
+ __free_page(page);
+ set_bit(chunk, dma_pmd_chunk_bitmap);
+ continue;
+ }
+
+ flush_cache_vmap(addr, addr + PAGE_SIZE);
+ /* Pair with test_bit_acquire() in readers. */
+ smp_mb__before_atomic();
+ set_bit(chunk, dma_pmd_chunk_bitmap);
+ dma_pmd_meta_pages++;
+ }
+
+ return 0;
+}
+
+bool dma_pmd_meta_ensure_pfn(unsigned long pfn, bool can_block)
+{
+ int ret;
+
+ if (likely(__dma_pmd_meta_of_pfn(pfn) != &dma_pmd_meta_nil))
+ return true;
+
+ if (!can_block || !dma_pmd_meta_base() ||
+ (pfn >> PMD_ORDER) >= dma_pmd_meta_nframes)
+ return false;
+
+ mutex_lock(&dma_pmd_meta_mutex);
+ ret = dma_pmd_meta_populate(dma_pmd_meta_base(), dma_pmd_meta_nframes,
+ pfn, pfn + (1UL << PMD_ORDER), true);
+ mutex_unlock(&dma_pmd_meta_mutex);
+
+ return !ret;
+}
+EXPORT_SYMBOL(dma_pmd_meta_ensure_pfn);
+
+/**
+ * dma_pmd_meta_init - Reserve and populate the sparse per-PMD metadata array
+ *
+ * Populates all present RAM pages before publishing @dma_pmd_meta_array so
+ * no concurrent reader ever sees an unmapped entry.
+ *
+ * Return: 0 on success, or negative errno on failure.
+ */
+int dma_pmd_meta_init(void)
+{
+ unsigned long nframes, nchunks, size;
+ struct dma_pmd_meta *array;
+ struct vm_struct *vm;
+ int ret = 0;
+
+ if (likely(dma_pmd_meta_base()))
+ return 0;
+
+ mutex_lock(&dma_pmd_meta_mutex);
+ if (dma_pmd_meta_array)
+ goto out_unlock;
+
+ nframes = DIV_ROUND_UP(max3((unsigned long)max_pfn,
+ (unsigned long)max_possible_pfn,
+ (unsigned long)min_t(u64, iomem_resource.end >> PAGE_SHIFT,
+ 1ULL << (MAX_PHYSMEM_BITS - PAGE_SHIFT))),
+ 1UL << PMD_ORDER);
+ size = PAGE_ALIGN(nframes << DMA_PMD_META_SHIFT);
+ nchunks = size >> PAGE_SHIFT;
+
+ dma_pmd_chunk_bitmap = bitmap_zalloc(nchunks, GFP_KERNEL);
+ if (!dma_pmd_chunk_bitmap) {
+ ret = -ENOMEM;
+ goto out_unlock;
+ }
+
+ vm = get_vm_area(size, VM_MAP);
+ if (!vm) {
+ bitmap_free(dma_pmd_chunk_bitmap);
+ dma_pmd_chunk_bitmap = NULL;
+ ret = -ENOMEM;
+ goto out_unlock;
+ }
+ array = vm->addr;
+
+ ret = dma_pmd_meta_populate(array, nframes, 0, max_pfn, false);
+ if (ret) {
+ struct page *p, *next;
+ unsigned long chunk;
+ LIST_HEAD(pages);
+
+ for_each_set_bit(chunk, dma_pmd_chunk_bitmap, nchunks) {
+ p = vmalloc_to_page((void *)array + (chunk << PAGE_SHIFT));
+ if (p)
+ list_add(&p->lru, &pages);
+ }
+ dma_pmd_meta_pages = 0;
+ free_vm_area(vm);
+ bitmap_free(dma_pmd_chunk_bitmap);
+ dma_pmd_chunk_bitmap = NULL;
+ list_for_each_entry_safe(p, next, &pages, lru) {
+ list_del_init(&p->lru);
+ __free_page(p);
+ }
+ goto out_unlock;
+ }
+
+ /*
+ * Publish @nframes and the populated pages before the array pointer:
+ * readers pair with smp_load_acquire(&dma_pmd_meta_array).
+ */
+ dma_pmd_meta_nframes = nframes;
+ /* Pairs with smp_load_acquire() in dma_is_pmd_page(). */
+ smp_store_release(&dma_pmd_meta_array, array);
+
+ pr_info("dma_pmd: %lu frames, %lu KB of KVA, %lu pages backed (%lu KB)\n",
+ nframes, size / 1024, dma_pmd_meta_pages,
+ dma_pmd_meta_pages * (PAGE_SIZE / 1024));
+
+out_unlock:
+ mutex_unlock(&dma_pmd_meta_mutex);
+ return ret;
+}
+EXPORT_SYMBOL(dma_pmd_meta_init);
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
new file mode 100644
index 0000000000000..a5ebede11dac2
--- /dev/null
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -0,0 +1,56 @@
+/* SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause */
+#ifndef _DRIVERS_IOMMU_DMA_PMD_PRIV_H
+#define _DRIVERS_IOMMU_DMA_PMD_PRIV_H
+
+#include <linux/dma-pmd.h>
+#include <linux/spinlock.h>
+#include <linux/types.h>
+
+#ifdef CONFIG_DMA_PMD
+
+/*
+ * Normally 128 B per entry (64 KB per GB of RAM), or 256 B when spinlock
+ * debugging enlarges struct dma_pmd_meta. Verified by static_assert().
+ */
+#if defined(CONFIG_DEBUG_SPINLOCK) || defined(CONFIG_DEBUG_LOCK_ALLOC)
+#define DMA_PMD_META_SHIFT 8
+#else
+#define DMA_PMD_META_SHIFT 7
+#endif
+#define DMA_PMD_META_SIZE BIT(DMA_PMD_META_SHIFT)
+
+/**
+ * struct dma_pmd_meta - Metadata for a single DMA_PMD page
+ * @pooled: True while this page is owned by an dma_pmd_pool (at offset 0)
+ *
+ * Lives in the sparse per-PMD-frame array @dma_pmd_meta_array indexed by
+ * (pfn >> PMD_ORDER). Only chunks covering valid RAM are backed by physical
+ * pages (64 KB per GB of RAM).
+ */
+struct dma_pmd_meta {
+ bool pooled;
+} __aligned(DMA_PMD_META_SIZE);
+
+static_assert(offsetof(struct dma_pmd_meta, pooled) == 0);
+static_assert(sizeof(struct dma_pmd_meta) == DMA_PMD_META_SIZE);
+
+extern unsigned long dma_pmd_meta_nframes;
+
+static inline struct dma_pmd_meta *dma_pmd_meta_base(void)
+{
+ /*
+ * Acquire pairs with smp_store_release() in dma_pmd_meta_init():
+ * orders both @dma_pmd_meta_nframes and the populated backing pages.
+ */
+ return smp_load_acquire(&dma_pmd_meta_array);
+}
+
+int dma_pmd_meta_init(void);
+bool dma_pmd_meta_ensure_pfn(unsigned long pfn, bool can_block);
+struct dma_pmd_meta *dma_pmd_meta_of_pfn(unsigned long pfn);
+struct dma_pmd_meta *dma_pmd_meta_from_phys(phys_addr_t pa);
+unsigned long dma_pmd_meta_to_pfn(const struct dma_pmd_meta *m);
+phys_addr_t dma_pmd_meta_to_phys(const struct dma_pmd_meta *m);
+
+#endif /* CONFIG_DMA_PMD */
+#endif /* _DRIVERS_IOMMU_DMA_PMD_PRIV_H */
diff --git a/include/linux/dma-pmd.h b/include/linux/dma-pmd.h
new file mode 100644
index 0000000000000..07df815945257
--- /dev/null
+++ b/include/linux/dma-pmd.h
@@ -0,0 +1,36 @@
+/* SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause */
+/* See Documentation/core-api/dma-pmd.rst for the architecture overview. */
+#ifndef _LINUX_DMA_PMD_H
+#define _LINUX_DMA_PMD_H
+
+#include <linux/compiler.h>
+#include <linux/mm_types.h>
+#include <linux/types.h>
+
+#ifdef CONFIG_DMA_PMD
+
+extern void *dma_pmd_meta_array;
+
+bool __dma_is_pmd_page(unsigned long pfn);
+
+static inline bool dma_is_pmd_page(unsigned long pfn)
+{
+ /*
+ * Acquire pairs with smp_store_release() in dma_pmd_meta_init():
+ * orders both @dma_pmd_meta_nframes and the populated backing pages.
+ */
+ if (likely(!smp_load_acquire(&dma_pmd_meta_array)))
+ return false;
+
+ return __dma_is_pmd_page(pfn);
+}
+
+#else /* !CONFIG_DMA_PMD */
+
+static inline bool dma_is_pmd_page(unsigned long pfn)
+{
+ return false;
+}
+
+#endif /* CONFIG_DMA_PMD */
+#endif /* _LINUX_DMA_PMD_H */
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 02/22] iommu/dma: add DMA_PMD pool lifecycle and page recycle hook
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 03/22] mm: Add split_page_compound() Luigi Rizzo
` (19 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
DMA_PMD manages pools of PMD_SIZE physically contiguous memory from which
IO buffers can be allocated. Recycling memory to the pool amortizes the
cost of mapping, unmapping and IOTLB flush, but must be handled so that
memory with active IOMMU mappings cannot be used for other data or code.
Add the pool lifecycle (struct dma_pmd_pool, dma_pmd_pool_create(),
dma_pmd_pool_destroy()) and the page release side. In particular,
__dma_pmd_free_page() is hooked into __free_pages_prepare() so that
DMA_PMD pages are returned to the pool.
Also implement asynchronous DMA_PMD page retirement and release to
the buddy allocator via a WQ_MEM_RECLAIM worker. Retirement removes
DMA_PMD pages from pools and destroys the mappings flushing the IOTLB,
but defers the buddy release by an RCU grace period, so a reader that
already passed dma_is_pmd_page() can safely finish reading p2m.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/Makefile | 2 +-
drivers/iommu/dma-pmd-meta.c | 4 +-
drivers/iommu/dma-pmd-pool.c | 379 +++++++++++++++++++++++++++++++++++
drivers/iommu/dma-pmd-priv.h | 129 +++++++++++-
include/linux/dma-pmd.h | 73 ++++++-
mm/page_alloc.c | 4 +
6 files changed, 579 insertions(+), 12 deletions(-)
create mode 100644 drivers/iommu/dma-pmd-pool.c
diff --git a/drivers/iommu/Makefile b/drivers/iommu/Makefile
index 2ad9b2eefd741..de2c3bad9aa3e 100644
--- a/drivers/iommu/Makefile
+++ b/drivers/iommu/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_IOMMU_API) += iommu-traces.o
obj-$(CONFIG_IOMMU_API) += iommu-sysfs.o
obj-$(CONFIG_IOMMU_DEBUGFS) += iommu-debugfs.o
obj-$(CONFIG_IOMMU_DMA) += dma-iommu.o
-obj-$(CONFIG_DMA_PMD) += dma-pmd-meta.o
+obj-$(CONFIG_DMA_PMD) += dma-pmd-meta.o dma-pmd-pool.o
obj-$(CONFIG_DMA_PMD_META_KUNIT_TEST) += dma-pmd-kunit.o
obj-$(CONFIG_IOMMU_IO_PGTABLE) += io-pgtable.o
obj-$(CONFIG_IOMMU_IO_PGTABLE_ARMV7S) += io-pgtable-arm-v7s.o
diff --git a/drivers/iommu/dma-pmd-meta.c b/drivers/iommu/dma-pmd-meta.c
index fcaf5b121e66a..27b8c2568c749 100644
--- a/drivers/iommu/dma-pmd-meta.c
+++ b/drivers/iommu/dma-pmd-meta.c
@@ -43,7 +43,9 @@ EXPORT_SYMBOL(dma_pmd_meta_array);
unsigned long dma_pmd_meta_nframes __read_mostly;
static unsigned long *dma_pmd_chunk_bitmap __read_mostly;
-static struct dma_pmd_meta dma_pmd_meta_nil;
+static struct dma_pmd_meta dma_pmd_meta_nil = {
+ .map_lock = __SPIN_LOCK_UNLOCKED(dma_pmd_meta_nil.map_lock),
+};
static unsigned long dma_pmd_meta_pages;
static DEFINE_MUTEX(dma_pmd_meta_mutex);
diff --git a/drivers/iommu/dma-pmd-pool.c b/drivers/iommu/dma-pmd-pool.c
new file mode 100644
index 0000000000000..5020265d52cd5
--- /dev/null
+++ b/drivers/iommu/dma-pmd-pool.c
@@ -0,0 +1,379 @@
+// SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause
+/*
+ * DMA_PMD streaming page pool, async page reclaim, shrinker, and debugfs.
+ *
+ * See Documentation/core-api/dma-pmd.rst for the architecture overview.
+ */
+
+#include <linux/atomic.h>
+#include <linux/bitmap.h>
+#include <linux/bitops.h>
+#include <linux/cache.h>
+#include <linux/debugfs.h>
+#include <linux/dma-mapping.h>
+#include <linux/dma-pmd.h>
+#include <linux/export.h>
+#include <linux/gfp.h>
+#include <linux/kref.h>
+#include <linux/list.h>
+#include <linux/llist.h>
+#include <linux/mm.h>
+#include <linux/moduleparam.h>
+#include <linux/mutex.h>
+#include <linux/rcupdate.h>
+#include <linux/seq_file.h>
+#include <linux/shrinker.h>
+#include <linux/slab.h>
+#include <linux/spinlock.h>
+#include <linux/srcu.h>
+#include <linux/workqueue.h>
+
+#include "dma-pmd-priv.h"
+
+/*
+ * DMA_PMD pages whose last block has just been freed and which are over the
+ * pool's idle watermark are pushed here locklessly for later release.
+ */
+static LLIST_HEAD(dma_pmd_free_list);
+
+/* Dedicated WQ_MEM_RECLAIM workqueue, can run without allocations. */
+static struct workqueue_struct *dma_pmd_wq __ro_after_init;
+
+/* All live pools. */
+static LIST_HEAD(dma_pmd_pools);
+static DEFINE_MUTEX(dma_pmd_pools_lock);
+static LLIST_HEAD(dma_pmd_dead_pools);
+
+static void dma_pmd_schedule_reclaim(void);
+
+static void dma_pmd_pool_free_kref(struct kref *kref)
+{
+ struct dma_pmd_pool *pool = container_of(kref, struct dma_pmd_pool, refcount);
+
+ llist_add(&pool->dead_node, &dma_pmd_dead_pools);
+ dma_pmd_schedule_reclaim();
+}
+
+/*
+ * Releasing an idle PMD page proceeds in three stages:
+ *
+ * 1. Detach from pool->idle (or when the last block of a destroyed/over-watermark
+ * PMD page frees in __dma_pmd_free_page()). Because __free_pages_prepare()
+ * may run in hardirq/softirq or with allocator locks held, the PMD page is
+ * pushed locklessly onto dma_pmd_free_list and dma_pmd_reclaim_work is
+ * scheduled on dma_pmd_wq.
+ * 2. In process context, dma_pmd_release_page() unmaps all IOMMU domains
+ * (dma_pmd_unmap_all(), added in a later commit), clears meta->pooled so
+ * new lockless readers stop entering @meta, and queues
+ * dma_pmd_release_page_rcu() via call_rcu().
+ * 3. After an RCU grace period (once no concurrent dma_pmd_free_page() reader
+ * can still be dereferencing @meta), dma_pmd_release_page_rcu() unfreezes
+ * all order-pool->order blocks (already split by split_page_compound()),
+ * returns them to the buddy allocator, and drops pool->refcount.
+ */
+
+/**
+ * dma_pmd_release_page_rcu - Return a retired PMD page to the buddy allocator
+ * @head: @rcu member of the PMD page being released
+ *
+ * Runs once no reader can still be resolving a PFN in this PMD page. All blocks
+ * are unfrozen and handed back to the buddy allocator, and the pool reference
+ * is dropped last.
+ */
+static void dma_pmd_release_page_rcu(struct rcu_head *head)
+{
+ /*
+ * Read @meta->pool before returning the blocks to buddy: once the last
+ * block is freed, the PMD frame can be reallocated and @meta overwritten.
+ */
+ struct dma_pmd_meta *meta = container_of(head, struct dma_pmd_meta, rcu);
+ struct page *page = pfn_to_page(dma_pmd_meta_to_pfn(meta));
+ struct dma_pmd_pool *pool = READ_ONCE(meta->pool);
+ unsigned int nr = DMA_PMD_BLOCKS(pool->order);
+ unsigned int order = pool->order;
+ unsigned long flags;
+ unsigned int i;
+
+ for (i = 0; i < nr; i++) {
+ struct page *block = page + (i << order);
+
+ page_ref_unfreeze(block, 1);
+ __free_pages(block, order);
+ }
+
+ spin_lock_irqsave(&pool->lock, flags);
+ pool->pmd_free_cnt++;
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+ kref_put(&pool->refcount, dma_pmd_pool_free_kref);
+}
+
+/**
+ * dma_pmd_release_page - Retire a PMD page and schedule its buddy release
+ * @meta: PMD page metadata structure (must have all usable blocks idle)
+ *
+ * Cost: Slow path; clears the membership bit. The blocks themselves are
+ * returned after an RCU grace period.
+ * Locking: Must not hold @pool->lock.
+ * Frequency: Rare (background reclaim workqueue, pool destruction).
+ */
+static void dma_pmd_release_page(struct dma_pmd_meta *meta)
+{
+ /* dma_pmd_unmap_all(meta) will be called here once IOMMU mappings are added. */
+
+ /*
+ * Clear membership before the grace period. A reader that passed
+ * dma_is_pmd_page() just before this may still be reading @meta, so
+ * the PMD page and its metadata entry must outlive the grace period.
+ */
+ WRITE_ONCE(meta->pooled, false);
+ call_rcu(&meta->rcu, dma_pmd_release_page_rcu);
+}
+
+static void dma_pmd_reclaim_work_fn(struct work_struct *work)
+{
+ struct llist_node *node = llist_del_all(&dma_pmd_free_list);
+ struct dma_pmd_pool *pool, *ptmp;
+ struct dma_pmd_meta *meta, *tmp;
+
+ llist_for_each_entry_safe(meta, tmp, node, llnode)
+ dma_pmd_release_page(meta);
+
+ node = llist_del_all(&dma_pmd_dead_pools);
+ if (node) {
+ mutex_lock(&dma_pmd_pools_lock);
+ llist_for_each_entry_safe(pool, ptmp, node, dead_node)
+ list_del(&pool->node);
+ mutex_unlock(&dma_pmd_pools_lock);
+
+ llist_for_each_entry_safe(pool, ptmp, node, dead_node)
+ kfree(pool);
+ }
+}
+
+static DECLARE_WORK(dma_pmd_reclaim_work, dma_pmd_reclaim_work_fn);
+
+static void dma_pmd_schedule_reclaim(void)
+{
+ struct workqueue_struct *wq = READ_ONCE(dma_pmd_wq);
+
+ if (likely(wq))
+ queue_work(wq, &dma_pmd_reclaim_work);
+ else
+ schedule_work(&dma_pmd_reclaim_work);
+}
+
+/**
+ * dma_pmd_pool_create - Create a DMA_PMD page pool
+ * @order: Block order to dispense, must be <= PMD_ORDER.
+ * @max_idle_pages: Maximum number of completely idle PMD pages to keep cached in
+ * the pool before asynchronously releasing excess PMD pages to buddy
+ * (0 uses default of 16 = 32 MB).
+ *
+ * New PMD physical pages are allocated on demand on the NUMA node of the
+ * calling CPU (the NAPI CPU for RX queue refill, or the application CPU for TX).
+ *
+ * Cost: Process-context control plane; allocates pool metadata structure, and
+ * on first call the global membership bitmap.
+ * Locking: May sleep (GFP_KERNEL).
+ * Frequency: Rare (driver queue initialization or subsystem init).
+ *
+ * Return: Pointer to the created pool, or NULL on failure.
+ */
+struct dma_pmd_pool *dma_pmd_pool_create(unsigned int order, unsigned int max_idle_pages)
+{
+ struct dma_pmd_pool *pool;
+
+ if (order > PMD_ORDER)
+ return NULL;
+
+ pool = kzalloc(sizeof(*pool), GFP_KERNEL);
+ if (!pool)
+ return NULL;
+
+ kref_init(&pool->refcount);
+ spin_lock_init(&pool->lock);
+ pool->order = order;
+ pool->max_idle_pages = max_idle_pages ? : 16;
+ pool->next_alloc_attempt = jiffies;
+ INIT_LIST_HEAD(&pool->partial);
+ INIT_LIST_HEAD(&pool->idle);
+ INIT_LIST_HEAD(&pool->full);
+
+ mutex_lock(&dma_pmd_pools_lock);
+ if (dma_pmd_meta_init()) {
+ mutex_unlock(&dma_pmd_pools_lock);
+ kfree(pool);
+ return NULL;
+ }
+ list_add(&pool->node, &dma_pmd_pools);
+ mutex_unlock(&dma_pmd_pools_lock);
+
+ return pool;
+}
+EXPORT_SYMBOL(dma_pmd_pool_create);
+
+/**
+ * dma_pmd_pool_destroy - Destroy a DMA_PMD page pool and unmap cached PMD pages
+ * @pool: Pool to destroy
+ *
+ * Frees completely idle 2M pages immediately. Device DMA must be quiesced by
+ * the caller first, but blocks already handed to the networking stack (e.g.,
+ * RX skbs waiting in socket queues or TX skbs in TCP retransmit queues) may
+ * still be in flight; those 2M pages stay on @pool->partial / @pool->full and
+ * keep @pool alive on @dma_pmd_pools via @pool->refcount until their last
+ * blocks free.
+ *
+ * Cost: Control plane teardown; flushes async reclaim workqueue.
+ * Locking: Process context (may sleep in flush_work). Acquires @pool->lock.
+ * Frequency: Rare (driver queue teardown or module unload).
+ *
+ * Return: Always NULL, so callers can write `pool = dma_pmd_pool_destroy(pool)`.
+ */
+struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *pool)
+{
+ struct dma_pmd_meta *meta, *tmp;
+ LIST_HEAD(release_list);
+ unsigned long flags;
+
+ if (!pool)
+ return NULL;
+
+ spin_lock_irqsave(&pool->lock, flags);
+ pool->destroyed = true;
+
+ /* Take out the idle PMD pages; the unmap happens outside the lock. */
+ list_splice_init(&pool->idle, &release_list);
+ pool->num_idle_pages = 0;
+
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+ list_for_each_entry_safe(meta, tmp, &release_list, list) {
+ list_del(&meta->list);
+ dma_pmd_release_page(meta);
+ }
+
+ kref_put(&pool->refcount, dma_pmd_pool_free_kref);
+ flush_work(&dma_pmd_reclaim_work);
+ return NULL;
+}
+EXPORT_SYMBOL(dma_pmd_pool_destroy);
+
+/**
+ * __dma_pmd_free_page - Recycle a freed block back into its DMA_PMD pool
+ * @page: Block being freed
+ *
+ * Called from __free_pages_prepare() when a block's refcount reaches 0.
+ * Marks the block index available in @meta->dirty_bitmap[] (or @meta->free_bitmap[]
+ * when init_on_free zeroes it). If the PMD page becomes completely idle and the
+ * pool exceeds @max_idle_pages, queues the PMD page to dma_pmd_free_list for
+ * asynchronous unmapping and buddy release.
+ *
+ * Cost: O(1) bit set under spinlock (~15-20 cycles).
+ * Locking: Runs inside the rcu_read_lock() section opened by
+ * dma_pmd_free_page(), and acquires @pool->lock (irqsave).
+ * Safe from any context (IRQ, softirq, process). Never sleeps.
+ * Frequency: Very high (called on every kfree_skb / napi_consume_skb / put_page
+ * for buffers allocated from an dma_pmd_pool).
+ *
+ * Return: true if the page belonged to a DMA_PMD pool and was recycled, false
+ * if it is a normal system page that should be freed by the buddy
+ * allocator.
+ */
+bool __dma_pmd_free_page(struct page *page)
+{
+ struct dma_pmd_meta *meta = dma_pmd_meta_of_pfn(page_to_pfn(page));
+ unsigned int nr = DMA_PMD_BLOCKS(meta->pool->order), idx;
+ struct dma_pmd_pool *pool = meta->pool;
+ unsigned long pfn = page_to_pfn(page);
+ bool should_release = false;
+ bool zeroed_on_free = false;
+ unsigned int avail;
+ unsigned long flags;
+
+ idx = (pfn - dma_pmd_meta_to_pfn(meta)) >> pool->order;
+
+ /*
+ * Never reuse or return a hwpoisoned block to buddy: leave its bit
+ * clear in free_bitmap/dirty_bitmap so the PMD page stays pinned and isolated.
+ */
+ if (unlikely(folio_contain_hwpoisoned_page(page_folio(page)))) {
+ pr_err_once("dma_pmd: poisoned block at pfn %lu, leaking it to keep pool %p intact\n",
+ pfn, pool);
+ return true;
+ }
+
+ /*
+ * Reject double frees (bit already set in free_bitmap or dirty_bitmap),
+ * out-of-range indices, or mismatched orders: leak the block rather than
+ * corrupt the pool or hand a live PMD page's block to the buddy allocator.
+ */
+ if (WARN_ON_ONCE(idx >= nr || compound_order(page) != pool->order ||
+ test_bit(idx, meta->free_bitmap) ||
+ test_bit(idx, meta->dirty_bitmap)))
+ return true;
+
+ /*
+ * Pooled blocks stay allocated from buddy until dma_pmd_release_page_rcu()
+ * (which runs KASAN/KMSAN/pgalloc_tag_sub via __free_pages()), and never
+ * come from HIGHMEM (CONFIG_DMA_PMD is 64-bit only).
+ */
+ if (want_init_on_free()) {
+ memset(page_address(page), 0, PAGE_SIZE << pool->order);
+ zeroed_on_free = true;
+ }
+
+ spin_lock_irqsave(&pool->lock, flags);
+
+ if (WARN_ON_ONCE(test_bit(idx, meta->free_bitmap) ||
+ test_bit(idx, meta->dirty_bitmap))) {
+ spin_unlock_irqrestore(&pool->lock, flags);
+ return true;
+ }
+
+ if (zeroed_on_free) {
+ __set_bit(idx, meta->free_bitmap);
+ meta->nr_free++;
+ } else {
+ __set_bit(idx, meta->dirty_bitmap);
+ meta->nr_dirty++;
+ }
+ avail = dma_pmd_meta_avail(meta);
+ pool->block_free_cnt++;
+
+ if (avail == 1 && !pool->destroyed)
+ list_move(&meta->list, &pool->partial);
+
+ if (avail == nr) {
+ if (pool->destroyed || pool->num_idle_pages >= pool->max_idle_pages) {
+ list_del_init(&meta->list);
+ llist_add(&meta->llnode, &dma_pmd_free_list);
+ should_release = true;
+ } else {
+ list_move(&meta->list, &pool->idle);
+ pool->num_idle_pages++;
+ }
+ }
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+ if (should_release)
+ dma_pmd_schedule_reclaim();
+
+ return true;
+}
+EXPORT_SYMBOL(__dma_pmd_free_page);
+
+static int __init dma_pmd_init(void)
+{
+ /*
+ * Reclaim frees memory, so it must not queue behind arbitrary work on
+ * system_wq when the machine is already short of it. Failure is not
+ * fatal: dma_pmd_schedule_reclaim() falls back to system_wq.
+ */
+ dma_pmd_wq = alloc_workqueue("dma_pmd", WQ_MEM_RECLAIM, 0);
+ if (!dma_pmd_wq)
+ pr_warn("dma_pmd: no reclaim workqueue, falling back to system_wq\n");
+
+ return 0;
+}
+subsys_initcall(dma_pmd_init);
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index a5ebede11dac2..99bd7d77fd4ef 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -3,37 +3,150 @@
#define _DRIVERS_IOMMU_DMA_PMD_PRIV_H
#include <linux/dma-pmd.h>
+#include <linux/kref.h>
+#include <linux/list.h>
+#include <linux/llist.h>
#include <linux/spinlock.h>
#include <linux/types.h>
#ifdef CONFIG_DMA_PMD
+#define DMA_PMD_BLOCKS(order) (1U << (PMD_ORDER - (order)))
+
+/*
+ * Locking
+ * -------
+ * dma_pmd_meta_mutex Serialises population of dma_pmd_meta_array. Taken
+ * from init and on-demand from dma_pmd_meta_ensure_pfn()
+ * when can_block is true. It is not nested with pool->lock
+ * or meta->map_lock, and sits inside dma_pmd_pools_lock
+ * when dma_pmd_pool_create() calls dma_pmd_meta_init().
+ *
+ * dma_pmd_pools_lock Mutex over the list of live pools. Outermost of the
+ * three, and it must not be taken from a context that
+ * cannot sleep.
+ *
+ * pool->lock IRQ-safe spinlock over one pool's partial/idle/full
+ * lists and its counters. The alloc and free fast paths
+ * take it on its own.
+ *
+ * meta->map_lock IRQ-safe spinlock over one PMD frame's domains_mapped
+ * bitmap. Innermost, and never held across anything that
+ * sleeps.
+ */
+
/*
- * Normally 128 B per entry (64 KB per GB of RAM), or 256 B when spinlock
- * debugging enlarges struct dma_pmd_meta. Verified by static_assert().
+ * 256 B per entry (128 KB per GB of RAM). Verified by static_assert().
*/
-#if defined(CONFIG_DEBUG_SPINLOCK) || defined(CONFIG_DEBUG_LOCK_ALLOC)
#define DMA_PMD_META_SHIFT 8
-#else
-#define DMA_PMD_META_SHIFT 7
-#endif
#define DMA_PMD_META_SIZE BIT(DMA_PMD_META_SHIFT)
/**
* struct dma_pmd_meta - Metadata for a single DMA_PMD page
- * @pooled: True while this page is owned by an dma_pmd_pool (at offset 0)
+ * @pooled: True while this page is owned by an dma_pmd_pool (offset 0)
+ * read locklessly by dma_is_pmd_page().
+ * @nr_free: Number of zeroed available blocks in @free_bitmap
+ * @nr_dirty: Number of dirty available blocks in @dirty_bitmap
+ * @map_lock: Spinlock protecting slow-path updates to @domains_mapped
+ *
+ * Map fast path:
+ * @domains_mapped: One bit per domain index: set once this PMD page's leaf
+ * PTE is installed in that domain.
+ *
+ * Alloc/free fast path, all under pool->lock:
+ * @pool: Owning dma_pmd_pool (holds a kref on the pool while pooled)
+ * @list: Node in pool->partial, pool->idle or pool->full
+ *
+ * Slow path only (disjoint lifetimes):
+ * @llnode: Node in lockless dma_pmd_free_list for async buddy release
+ * @rcu: RCU head used to defer buddy release past lockless readers
+ *
+ * @free_bitmap: Bitmap of zeroed available block indices (up to 1 << PMD_ORDER)
+ * @dirty_bitmap: Bitmap of dirty available block indices
*
* Lives in the sparse per-PMD-frame array @dma_pmd_meta_array indexed by
* (pfn >> PMD_ORDER). Only chunks covering valid RAM are backed by physical
- * pages (64 KB per GB of RAM).
+ * pages (128 KB per GB of RAM).
*/
struct dma_pmd_meta {
bool pooled;
+ u16 nr_free;
+ u16 nr_dirty;
+ spinlock_t map_lock;
+
+ unsigned long domains_mapped;
+
+ struct dma_pmd_pool *pool;
+ struct list_head list;
+
+ union {
+ struct llist_node llnode;
+ struct rcu_head rcu;
+ };
+
+ DECLARE_BITMAP(free_bitmap, 1U << PMD_ORDER) ____cacheline_aligned;
+ DECLARE_BITMAP(dirty_bitmap, 1U << PMD_ORDER) ____cacheline_aligned;
} __aligned(DMA_PMD_META_SIZE);
static_assert(offsetof(struct dma_pmd_meta, pooled) == 0);
static_assert(sizeof(struct dma_pmd_meta) == DMA_PMD_META_SIZE);
+static inline unsigned int dma_pmd_meta_avail(const struct dma_pmd_meta *m)
+{
+ return m->nr_free + m->nr_dirty;
+}
+
+/**
+ * struct dma_pmd_pool - Pool managing PMD-backed blocks of a fixed order
+ *
+ * @lock: Spinlock protecting @partial, @idle, @full, @num_idle_pages,
+ * block bitmaps and the statistics counters
+ * @order: Block order managed by this pool (<= PMD_ORDER)
+ * @destroyed: Set when dma_pmd_pool_destroy() has been called
+ * @partial: List of PMD pages with at least 1 free block, but not all
+ * of them (MRU ordered). Candidates for allocation.
+ * @idle: List of PMD pages whose every usable block is free. Held
+ * separately because they can be released under pressure.
+ * @full: List of PMD pages with 0 free blocks
+ * @num_idle_pages: Length of @idle
+ * @max_idle_pages: High watermark of completely idle PMD pages retained in
+ * @idle before triggering asynchronous release to buddy
+ * @pmd_alloc_cnt: Statistics counter of PMD pages allocated from buddy
+ * @pmd_free_cnt: Statistics counter of PMD pages released back to buddy
+ * @block_alloc_cnt: Statistics counter of block allocations satisfied
+ * @block_free_cnt: Statistics counter of block frees recycled into pool
+ * @next_alloc_attempt: Do not attempt a new order-9 allocation before this time.
+ * Damps repeated high-order GFP_ATOMIC failures under
+ * fragmentation, which would otherwise be retried on every
+ * pool miss from softirq context.
+ * @refcount: Reference count; held by pool creator and each active PMD page
+ * @node: Node in the global dma_pmd_pools list
+ * @dead_node: Node in dma_pmd_dead_pools once @refcount reaches 0
+ */
+struct dma_pmd_pool {
+ /* First cacheline: hot fields touched on block alloc and free. */
+ spinlock_t lock;
+ u8 order;
+ bool destroyed;
+ struct list_head partial;
+ struct list_head idle;
+ struct list_head full;
+ unsigned int num_idle_pages;
+ unsigned int max_idle_pages;
+
+ /* Statistics, all updated under @lock. */
+ u64 pmd_alloc_cnt;
+ u64 pmd_free_cnt;
+ u64 block_alloc_cnt;
+ u64 block_free_cnt;
+
+ /* Cold. */
+ unsigned long next_alloc_attempt;
+ struct kref refcount;
+ struct list_head node;
+ struct llist_node dead_node;
+} ____cacheline_aligned;
+
extern unsigned long dma_pmd_meta_nframes;
static inline struct dma_pmd_meta *dma_pmd_meta_base(void)
diff --git a/include/linux/dma-pmd.h b/include/linux/dma-pmd.h
index 07df815945257..236be50b7349f 100644
--- a/include/linux/dma-pmd.h
+++ b/include/linux/dma-pmd.h
@@ -3,9 +3,12 @@
#ifndef _LINUX_DMA_PMD_H
#define _LINUX_DMA_PMD_H
-#include <linux/compiler.h>
-#include <linux/mm_types.h>
#include <linux/types.h>
+#include <linux/mm.h>
+#include <linux/rcupdate.h>
+
+struct device;
+struct dma_pmd_pool;
#ifdef CONFIG_DMA_PMD
@@ -25,6 +28,56 @@ static inline bool dma_is_pmd_page(unsigned long pfn)
return __dma_is_pmd_page(pfn);
}
+bool __dma_pmd_free_page(struct page *page);
+
+static inline bool dma_pmd_free_page(struct page *page)
+{
+ bool ret;
+
+ rcu_read_lock();
+ ret = dma_is_pmd_page(page_to_pfn(page)) && __dma_pmd_free_page(page);
+ rcu_read_unlock();
+
+ return ret;
+}
+
+/*
+ * External API: dma_pmd_pool (streaming DMA page pool)
+ * -----------------------------------------------------
+ * Opportunistic drop-in for alloc_pages_node() that carves fixed-order
+ * (order <= PMD_ORDER) pages out of 2MB physically contiguous buddy pages.
+ *
+ * - Creation (dma_pmd_pool_create):
+ * Records @order and @max_idle_pages; no 2MB pages or IOMMU mappings
+ * are allocated yet.
+ *
+ * - Allocation (dma_pmd_pool_alloc_node / dma_pmd_pool_alloc):
+ * Dispenses an order-@order compound page with refcount 1 (or NULL on
+ * unsupported GFP flags / OOM, so callers fall back to alloc_pages_node()).
+ * On a pool miss, allocates a 2MB page on @nid and splits it into subpages.
+ *
+ * - Mapping (dma_map_page / dma_map_single / dma_map_phys / dma_map_sg):
+ * Intercepts single-buffer and scatterlist maps. On first use of a 2MB page
+ * in an IOMMU domain, lazily installs a 2MB bidirectional leaf PTE in the
+ * domain's DMA_PMD IOVA window; subsequent maps are O(1) arithmetic.
+ *
+ * - Unmapping & Recycling (dma_unmap_* / put_page):
+ * dma_unmap_page/single/phys/sg() is a no-op for IOVAs in the DMA_PMD window,
+ * keeping the 2MB PTE resident across buffer reuse. When the page's last
+ * reference drops (put_page() / __free_pages()), __free_pages_prepare()
+ * intercepts it via dma_pmd_free_page() and recycles the subpage back into
+ * @pool.
+ *
+ * - Unmap & Buddy Release (reclaim / shrinker / dma_pmd_pool_destroy):
+ * A 2MB page is released back to the buddy allocator only when all of its
+ * subpages are free in @pool and either the pool exceeds @max_idle_pages, the
+ * system shrinker reclaims idle pages, or dma_pmd_pool_destroy() is called.
+ * Before __free_pages(PMD_ORDER) is called, the 2MB leaf PTE is unmapped from
+ * every domain that mapped it and the IOTLB is synchronously flushed.
+ */
+struct dma_pmd_pool *dma_pmd_pool_create(unsigned int order, unsigned int max_idle_pages);
+struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *pool);
+
#else /* !CONFIG_DMA_PMD */
static inline bool dma_is_pmd_page(unsigned long pfn)
@@ -32,5 +85,21 @@ static inline bool dma_is_pmd_page(unsigned long pfn)
return false;
}
+static inline bool dma_pmd_free_page(struct page *page)
+{
+ return false;
+}
+
+static inline struct dma_pmd_pool *
+dma_pmd_pool_create(unsigned int order, unsigned int max_idle_pages)
+{
+ return NULL;
+}
+
+static inline struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *pool)
+{
+ return NULL;
+}
+
#endif /* CONFIG_DMA_PMD */
#endif /* _LINUX_DMA_PMD_H */
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 12fac9084c483..e8e9905cdea5b 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -17,6 +17,7 @@
#include <linux/stddef.h>
#include <linux/mm.h>
#include <linux/highmem.h>
+#include <linux/dma-pmd.h>
#include <linux/interrupt.h>
#include <linux/jiffies.h>
#include <linux/compiler.h>
@@ -1327,6 +1328,9 @@ static __always_inline bool __free_pages_prepare(struct page *page,
VM_BUG_ON_PAGE(PageTail(page), page);
+ if (unlikely(dma_pmd_free_page(page)))
+ return false;
+
trace_mm_page_free(page, order);
kmsan_free_page(page, order);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 03/22] mm: Add split_page_compound()
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 02/22] iommu/dma: add DMA_PMD pool lifecycle and page recycle hook Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 04/22] iommu/dma: add DMA_PMD pool block allocation Luigi Rizzo
` (18 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Add split_page_compound(), which splits an order-@old_order page into
independently refcounted order-@new_order compound pages (or order-0
pages when @new_order == 0). This is the high-order counterpart of
split_page(), used by DMA_PMD pools to carve PMD pages into smaller
compound blocks.
All resulting pieces are returned frozen (refcount 0) so the caller can
park them in a pool and publish each piece with page_ref_unfreeze(piece, 1)
when handed out.
Also add a KUnit test suite (CONFIG_SPLIT_PAGE_COMPOUND_KUNIT_TEST) in
mm/split_page_compound_kunit.c.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
include/linux/mm.h | 2 +
mm/Kconfig.debug | 11 ++++
mm/Makefile | 1 +
mm/page_alloc.c | 85 +++++++++++++++++++++++++
mm/split_page_compound_kunit.c | 113 +++++++++++++++++++++++++++++++++
5 files changed, 212 insertions(+)
create mode 100644 mm/split_page_compound_kunit.c
diff --git a/include/linux/mm.h b/include/linux/mm.h
index dd09c438fa23e..799a42005627f 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -1985,6 +1985,8 @@ static inline struct folio *virt_to_folio(const void *x)
void __folio_put(struct folio *folio);
void split_page(struct page *page, unsigned int order);
+int split_page_compound(struct page *page, unsigned int old_order,
+ unsigned int new_order);
void folio_copy(struct folio *dst, struct folio *src);
int folio_mc_copy(struct folio *dst, struct folio *src);
diff --git a/mm/Kconfig.debug b/mm/Kconfig.debug
index 15dca19dd07da..e1a04671c3b0d 100644
--- a/mm/Kconfig.debug
+++ b/mm/Kconfig.debug
@@ -347,3 +347,14 @@ config MEM_ALLOC_PROFILING_DEBUG
help
Adds warnings with helpful error messages for memory allocation
profiling.
+
+config SPLIT_PAGE_COMPOUND_KUNIT_TEST
+ bool "KUnit test for split_page_compound() (built-in)" if !KUNIT_ALL_TESTS
+ depends on KUNIT=y
+ default KUNIT_ALL_TESTS
+ help
+ Builds KUnit unit tests for split_page_compound() in mm/page_alloc.c,
+ verifying compound splitting, sub-block freezing, and speculative
+ reference handling.
+
+ If unsure, say N.
diff --git a/mm/Makefile b/mm/Makefile
index e7245cb88c665..448c0f5f783f3 100644
--- a/mm/Makefile
+++ b/mm/Makefile
@@ -148,3 +148,4 @@ obj-$(CONFIG_EXECMEM) += execmem.o
obj-$(CONFIG_TMPFS_QUOTA) += shmem_quota.o
obj-$(CONFIG_LAZY_MMU_MODE_KUNIT_TEST) += tests/lazy_mmu_mode_kunit.o
obj-$(CONFIG_MEM_ALLOC_PROFILING) += alloc_tag.o
+obj-$(CONFIG_SPLIT_PAGE_COMPOUND_KUNIT_TEST) += split_page_compound_kunit.o
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index e8e9905cdea5b..a663ec81a588b 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -3133,6 +3133,91 @@ void split_page(struct page *page, unsigned int order)
}
EXPORT_SYMBOL_GPL(split_page);
+/*
+ * split_page_compound() - split an order-@old_order page into compound pages
+ * of order @new_order
+ * @page: the page to split
+ * @old_order: the current order of @page
+ * @new_order: the order of the resulting pages
+ *
+ * Unlike split_page(), which produces order-0 pages, this produces
+ * 1 << (@old_order - @new_order) compound pages, each with its own head.
+ * @page must not be compound, and the caller must hold the only reference to
+ * it: the block is frozen while the new heads are built.
+ *
+ * All pieces (including @page at index 0) are returned frozen, with a zero
+ * refcount. Publish each piece with
+ * page_ref_unfreeze(piece, 1) before handing it out: its release store orders
+ * the compound layout built here against anyone who then observes the
+ * refcount. A relaxed set_page_count() would not, and on a weakly ordered
+ * machine a PFN walker could see a live refcount on a head whose layout is
+ * not visible yet.
+ *
+ * Memcg-charged pages are rejected: the accounting helpers assume a split
+ * produces order-0 pages and would mis-attribute the new compound heads.
+ *
+ * Return: 0 on success, -EINVAL if @page cannot be split or if @new_order is
+ * larger than @old_order, -EBUSY if anyone but the caller holds a reference
+ * to it.
+ */
+int split_page_compound(struct page *page, unsigned int old_order,
+ unsigned int new_order)
+{
+ unsigned int i, step = 1U << new_order, nr = 1U << old_order;
+
+ if (WARN_ON_ONCE(PageCompound(page) || !page_count(page)))
+ return -EINVAL;
+
+ if (WARN_ON_ONCE(new_order > old_order))
+ return -EINVAL;
+
+ if (WARN_ON_ONCE(memcg_kmem_online() && PageMemcgKmem(page)))
+ return -EINVAL;
+
+ /*
+ * Shape the new compound pages while the block is frozen, which is
+ * the order __alloc_pages_noprof() itself uses: prep_new_page()
+ * builds the page with a zero refcount and the set_page_refcounted()
+ * that follows is what makes it visible.
+ *
+ * Freezing stops a speculative PFN walker from resolving a compound
+ * page that is only half built: get_page_unless_zero() fails for as
+ * long as the layout below is being written. It does not order
+ * against memory_failure(), which sets PG_hwpoison before taking any
+ * reference, so the non-atomic __SetPageHead() below can still lose a
+ * concurrent poison bit - exactly as it can for every other
+ * prep_compound_page() caller, the page allocator included.
+ *
+ * A failed freeze is not a bug, it means someone holds a speculative
+ * reference. Every PFN walker in the tree gates
+ * get_page_unless_zero() on PageLRU, which a freshly allocated block
+ * is not, so in practice only the hwpoison machinery gets here. Let
+ * the caller fall back rather than warn: a machine running
+ * panic_on_warn should not die of a condition the caller handles.
+ */
+ if (!page_ref_freeze(page, 1))
+ return -EBUSY;
+
+ if (!new_order) {
+ /*
+ * The subpages of a non-compound block are already
+ * independent frozen pages, so only the bookkeeping applies.
+ */
+ split_page_owner(page, old_order, 0);
+ pgalloc_tag_split(page_folio(page), old_order, 0);
+ split_page_memcg(page, old_order);
+ return 0;
+ }
+
+ split_page_owner(page, old_order, new_order);
+ pgalloc_tag_split(page_folio(page), old_order, new_order);
+
+ for (i = 0; i < nr; i += step)
+ prep_compound_page(page + i, new_order);
+
+ return 0;
+}
+
int __isolate_free_page(struct page *page, unsigned int order)
{
struct zone *zone = page_zone(page);
diff --git a/mm/split_page_compound_kunit.c b/mm/split_page_compound_kunit.c
new file mode 100644
index 0000000000000..bb0ac908bbc1e
--- /dev/null
+++ b/mm/split_page_compound_kunit.c
@@ -0,0 +1,113 @@
+// SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause
+/*
+ * KUnit tests for split_page_compound().
+ */
+#include <kunit/test.h>
+#include <linux/gfp.h>
+#include <linux/mm.h>
+
+static void test_split_compound_pieces(struct kunit *test)
+{
+ const unsigned int old_order = 5, new_order = 2;
+ const unsigned int step = 1U << new_order;
+ const unsigned int nr = 1U << old_order;
+ struct page *page;
+ unsigned int i, j;
+ int ret;
+
+ page = alloc_pages(GFP_KERNEL, old_order);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+ KUNIT_EXPECT_FALSE(test, PageCompound(page));
+
+ ret = split_page_compound(page, old_order, new_order);
+ if (ret)
+ __free_pages(page, old_order);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+
+ for (i = 0; i < nr; i += step) {
+ struct page *sub = page + i;
+
+ /* All heads 0..N-1 are returned frozen (refcount 0). */
+ KUNIT_EXPECT_EQ(test, page_count(sub), 0);
+ KUNIT_EXPECT_TRUE(test, PageHead(sub));
+ KUNIT_EXPECT_EQ(test, compound_order(sub), new_order);
+
+ for (j = 1; j < step; j++) {
+ KUNIT_EXPECT_TRUE(test, PageTail(sub + j));
+ KUNIT_EXPECT_PTR_EQ(test, compound_head(sub + j), sub);
+ }
+
+ page_ref_unfreeze(sub, 1);
+
+ /* Tail get_page()/put_page() must operate on this piece's head. */
+ get_page(sub + 1);
+ KUNIT_EXPECT_EQ(test, page_count(sub), 2);
+ put_page(sub + 1);
+ KUNIT_EXPECT_EQ(test, page_count(sub), 1);
+ }
+
+ /* Each compound piece frees independently at new_order. */
+ for (i = 0; i < nr; i += step)
+ __free_pages(page + i, new_order);
+}
+
+static void test_split_order0_pieces(struct kunit *test)
+{
+ const unsigned int old_order = 3, nr = 1U << old_order;
+ struct page *page;
+ unsigned int i;
+ int ret;
+
+ page = alloc_pages(GFP_KERNEL, old_order);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+
+ ret = split_page_compound(page, old_order, 0);
+ if (ret)
+ __free_pages(page, old_order);
+ KUNIT_ASSERT_EQ(test, ret, 0);
+
+ for (i = 0; i < nr; i++) {
+ struct page *sub = page + i;
+
+ KUNIT_EXPECT_FALSE(test, PageCompound(sub));
+ KUNIT_EXPECT_EQ(test, page_count(sub), 0);
+ page_ref_unfreeze(sub, 1);
+ __free_pages(sub, 0);
+ }
+}
+
+static void test_split_ebusy_extra_ref(struct kunit *test)
+{
+ const unsigned int old_order = 4;
+ const unsigned int new_order = 2;
+ struct page *page;
+ int ret;
+
+ page = alloc_pages(GFP_KERNEL, old_order);
+ KUNIT_ASSERT_NOT_NULL(test, page);
+
+ /* Simulate a concurrent speculative reference. */
+ get_page(page);
+ ret = split_page_compound(page, old_order, new_order);
+ KUNIT_EXPECT_EQ(test, ret, -EBUSY);
+ KUNIT_EXPECT_FALSE(test, PageCompound(page));
+
+ put_page(page);
+ __free_pages(page, old_order);
+}
+
+static struct kunit_case split_page_compound_test_cases[] = {
+ KUNIT_CASE(test_split_compound_pieces),
+ KUNIT_CASE(test_split_order0_pieces),
+ KUNIT_CASE(test_split_ebusy_extra_ref),
+ {}
+};
+
+static struct kunit_suite split_page_compound_test_suite = {
+ .name = "split_page_compound",
+ .test_cases = split_page_compound_test_cases,
+};
+
+kunit_test_suite(split_page_compound_test_suite);
+MODULE_DESCRIPTION("KUnit tests for split_page_compound()");
+MODULE_LICENSE("Dual BSD/GPL");
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 04/22] iommu/dma: add DMA_PMD pool block allocation
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (2 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 03/22] mm: Add split_page_compound() Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 05/22] iommu/dma: Global cap and shrinker for DMA_PMD pool memory Luigi Rizzo
` (17 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Add dma_pmd_add_page(), dma_pmd_pool_alloc_bulk_node() (with single-page
inline wrappers dma_pmd_pool_alloc_node() and dma_pmd_pool_alloc()), and
dma_pmd_pool_has_free() to carve PMD_SIZE physical blocks into subpages
of a pool's order with split_page_compound() and hand them out to callers.
A parked subpage is kept frozen (refcount 0) and published with
page_ref_unfreeze() when it is handed out, mirroring the way the page
allocator shapes a page before publishing it, so a stray put_page() on a
parked subpage trips a refcount check instead of freeing it twice.
GFP flags asking for a placement a pool cannot choose - __GFP_THISNODE,
__GFP_DMA, __GFP_DMA32 - are ignored, as documented at
dma_pmd_pool_alloc_bulk_node(); __GFP_ZERO is honoured by zeroing at
dispense.
Also add a KUnit test case (test_pool_alloc_and_recycle) exercising pool
allocation, recycling via __free_page(), and deferred teardown with
in-flight blocks.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/dma-pmd-kunit.c | 41 ++++-
drivers/iommu/dma-pmd-pool.c | 279 ++++++++++++++++++++++++++++++++++
include/linux/dma-pmd.h | 40 +++++
3 files changed, 359 insertions(+), 1 deletion(-)
diff --git a/drivers/iommu/dma-pmd-kunit.c b/drivers/iommu/dma-pmd-kunit.c
index f80527eaf2e84..24978a5123f54 100644
--- a/drivers/iommu/dma-pmd-kunit.c
+++ b/drivers/iommu/dma-pmd-kunit.c
@@ -70,6 +70,7 @@ static void test_meta_pooled_toggle(struct kunit *test)
base_pfn = ALIGN_DOWN(pfn, 1UL << PMD_ORDER);
/* The buddy allocator returns naturally aligned blocks. */
KUNIT_EXPECT_EQ(test, pfn, base_pfn);
+ KUNIT_ASSERT_TRUE(test, dma_pmd_meta_ensure_pfn(pfn, true));
m = dma_pmd_meta_of_pfn(pfn);
KUNIT_EXPECT_FALSE(test, dma_is_pmd_page(pfn));
@@ -85,10 +86,48 @@ static void test_meta_pooled_toggle(struct kunit *test)
__free_pages(page, PMD_ORDER);
}
+static void test_pool_alloc_and_recycle(struct kunit *test)
+{
+ struct dma_pmd_pool *pool;
+ struct page *p1, *p2;
+ unsigned int order;
+
+ KUNIT_EXPECT_NULL(test, dma_pmd_pool_destroy(NULL));
+
+ for (order = 0; order <= PMD_ORDER + 1; order++) {
+ pool = dma_pmd_pool_create(order, 2);
+ if (order > PMD_ORDER) {
+ KUNIT_EXPECT_NULL(test, pool);
+ continue;
+ }
+ KUNIT_ASSERT_NOT_NULL(test, pool);
+ KUNIT_EXPECT_FALSE(test, dma_pmd_pool_has_free(pool));
+
+ p1 = dma_pmd_pool_alloc(pool, GFP_KERNEL | __GFP_ZERO);
+ if (!p1)
+ dma_pmd_pool_destroy(pool);
+ KUNIT_ASSERT_NOT_NULL(test, p1);
+ KUNIT_EXPECT_TRUE(test, dma_is_pmd_page(page_to_pfn(p1)));
+ KUNIT_EXPECT_EQ(test, dma_pmd_pool_has_free(pool), order < PMD_ORDER);
+
+ /* Freeing p1 returns it to the pool; next allocation reuses index 0. */
+ __free_pages(p1, order);
+ KUNIT_EXPECT_TRUE(test, dma_pmd_pool_has_free(pool));
+ p2 = dma_pmd_pool_alloc(pool, GFP_KERNEL);
+ KUNIT_EXPECT_PTR_EQ(test, p1, p2);
+
+ /* Destroy with p2 still in flight; freeing p2 releases the 2M page. */
+ dma_pmd_pool_destroy(pool);
+ if (p2)
+ __free_pages(p2, order);
+ }
+}
+
static struct kunit_case dma_pmd_meta_test_cases[] = {
KUNIT_CASE(test_meta_init_and_roundtrip),
KUNIT_CASE(test_meta_invalid_phys),
KUNIT_CASE(test_meta_pooled_toggle),
+ KUNIT_CASE(test_pool_alloc_and_recycle),
{}
};
@@ -98,5 +137,5 @@ static struct kunit_suite dma_pmd_meta_test_suite = {
};
kunit_test_suite(dma_pmd_meta_test_suite);
-MODULE_DESCRIPTION("KUnit tests for struct dma_pmd_meta table");
+MODULE_DESCRIPTION("KUnit tests for struct dma_pmd_meta table and DMA_PMD pool");
MODULE_LICENSE("Dual BSD/GPL");
diff --git a/drivers/iommu/dma-pmd-pool.c b/drivers/iommu/dma-pmd-pool.c
index 5020265d52cd5..e6e5ba7c65f27 100644
--- a/drivers/iommu/dma-pmd-pool.c
+++ b/drivers/iommu/dma-pmd-pool.c
@@ -259,6 +259,285 @@ struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *pool)
}
EXPORT_SYMBOL(dma_pmd_pool_destroy);
+/**
+ * dma_pmd_add_page - Allocate, split, and register a new PMD page
+ * @pool: Owning pool
+ * @gfp: GFP allocation flags
+ * @nid: Target NUMA node (or NUMA_NO_NODE for local node)
+ *
+ * Cost: Slow path (pool miss); allocates PMD page from buddy on @nid,
+ * splits into compound blocks, and sets the PMD page's membership bit.
+ * Locking: Lockless with respect to @pool->lock (caller inserts into list).
+ * Frequency: Rare after warmup (only when the pool holds no PMD page with a
+ * free block).
+ *
+ * Return: Pointer to new dma_pmd_meta with all usable blocks free, or NULL
+ * on failure.
+ */
+static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
+ gfp_t gfp, int nid)
+{
+ /*
+ * Require full GFP_KERNEL (not just gfpflags_allow_blocking()):
+ * dma_pmd_meta_ensure_pfn() and dma_pmd_page_prepare() use GFP_KERNEL.
+ */
+ bool can_block = (gfp & GFP_KERNEL) == GFP_KERNEL;
+ unsigned int nr = DMA_PMD_BLOCKS(pool->order);
+ struct dma_pmd_meta *meta;
+ gfp_t alloc_gfp, base_gfp;
+ struct page *page;
+
+ /*
+ * A high-order allocation is expensive when it fails, and under fragmentation
+ * it fails on every pool miss. For GFP_ATOMIC, back off briefly.
+ */
+ if (!can_block && time_before(jiffies, READ_ONCE(pool->next_alloc_attempt)))
+ return NULL;
+
+ /*
+ * The following flags are discarded:
+ *
+ * __GFP_MOVABLE would let the PMD page come from ZONE_MOVABLE or a CMA
+ * pageblock. A pooled PMD page is pinned for the life of the pool and
+ * has no movable_operations, so memory hot-remove and CMA allocation
+ * would fail on it forever.
+ *
+ * __GFP_HIGHMEM goes because pooled blocks are handed out as
+ * directly addressable DMA buffers, and __GFP_COMP because the PMD page
+ * is shaped into compound pieces by hand. What is left is also free
+ * of everything slab treats as a bug in GFP_SLAB_BUG_MASK, so the
+ * metadata allocation below can use it as it stands.
+ * __GFP_ZERO is handled per-block in dma_pmd_pool_alloc_bulk_node().
+ */
+ base_gfp = gfp & ~(__GFP_COMP | __GFP_HIGHMEM | __GFP_MOVABLE | __GFP_ZERO);
+
+ /*
+ * A caller that can block gets __GFP_RETRY_MAYFAIL so the allocation
+ * is allowed to compact for an order-9 page, which is the difference
+ * between a pool that is populated before traffic starts and one that
+ * never populates at all on a fragmented machine. It still fails
+ * rather than triggering the OOM killer: this is an optimisation, and
+ * every caller has a working fallback.
+ */
+ alloc_gfp = base_gfp | __GFP_NOWARN;
+ alloc_gfp |= can_block ? __GFP_RETRY_MAYFAIL : __GFP_NORETRY;
+
+ page = alloc_pages_node(nid, alloc_gfp, PMD_ORDER);
+ if (!page) {
+ /* Retry backoff; could be made shorter or tunable in future. */
+ if (!can_block)
+ WRITE_ONCE(pool->next_alloc_attempt,
+ jiffies + DIV_ROUND_UP(HZ, 100));
+ return NULL;
+ }
+
+ /*
+ * Ensure the metadata chunk for this PFN is populated. Fails if the PFN
+ * is past the reservation, or if it belongs to hotplugged memory and the
+ * caller cannot block to populate its metadata page.
+ */
+ if (unlikely(!dma_pmd_meta_ensure_pfn(page_to_pfn(page), can_block))) {
+ __free_pages(page, PMD_ORDER);
+ return NULL;
+ }
+
+ /* Fails with -EBUSY if a concurrent PFN walker holds a speculative ref. */
+ if (split_page_compound(page, PMD_ORDER, pool->order)) {
+ __free_pages(page, PMD_ORDER);
+ return NULL;
+ }
+
+ meta = dma_pmd_meta_of_pfn(page_to_pfn(page));
+ memset(meta, 0, sizeof(*meta));
+ meta->pool = pool;
+ spin_lock_init(&meta->map_lock);
+
+ bitmap_set(meta->dirty_bitmap, 0, nr);
+ meta->nr_dirty = nr;
+
+ kref_get(&pool->refcount);
+
+ /*
+ * Publish last. Once the bit is set, __free_pages_prepare() on any CPU
+ * will route this PMD page's blocks back into the pool, so every @meta
+ * field a reader consults has to be in place first.
+ *
+ * No explicit barrier: a block cannot reach any reader until this
+ * function returns and the caller publishes the PMD page under
+ * @pool->lock, whose release orders both stores against every consumer.
+ */
+ WRITE_ONCE(meta->pooled, true);
+
+ return meta;
+}
+
+static unsigned int dma_pmd_scan_bitmap(struct dma_pmd_pool *pool,
+ unsigned long *bitmap, u16 *countp,
+ unsigned long base_pfn, unsigned int nr,
+ unsigned long want, struct page **out,
+ bool unfreeze)
+{
+ unsigned int idx, got = 0, used = 0;
+
+ if (!*countp || !want)
+ return 0;
+
+ for_each_set_bit(idx, bitmap, nr) {
+ struct page *block = pfn_to_page(base_pfn + (idx << pool->order));
+
+ __clear_bit(idx, bitmap);
+ used++;
+ if (unlikely(folio_contain_hwpoisoned_page(page_folio(block)))) {
+ pr_err_once("dma_pmd: poisoned block at pfn %lu, leaking it to keep pool %p intact\n",
+ page_to_pfn(block), pool);
+ } else {
+ if (unfreeze)
+ page_ref_unfreeze(block, 1);
+ out[got++] = block;
+ }
+
+ if (used == *countp || got == want)
+ break;
+ }
+ *countp -= used;
+ pool->block_alloc_cnt += used;
+ return got;
+}
+
+/* Extract up to @want available blocks from @meta under @pool->lock. */
+static unsigned long dma_pmd_take_blocks(struct dma_pmd_pool *pool,
+ struct dma_pmd_meta *meta,
+ unsigned long want,
+ struct page **out, bool zero)
+{
+ unsigned long base_pfn = dma_pmd_meta_to_pfn(meta);
+ unsigned int avail_before = dma_pmd_meta_avail(meta);
+ unsigned int nr = DMA_PMD_BLOCKS(pool->order);
+ unsigned int got = 0, used;
+
+ if (zero) {
+ got += dma_pmd_scan_bitmap(pool, meta->free_bitmap, &meta->nr_free,
+ base_pfn, nr, want, out, true);
+ got += dma_pmd_scan_bitmap(pool, meta->dirty_bitmap, &meta->nr_dirty,
+ base_pfn, nr, want - got, out + got, false);
+ } else {
+ got += dma_pmd_scan_bitmap(pool, meta->dirty_bitmap, &meta->nr_dirty,
+ base_pfn, nr, want, out, false);
+ got += dma_pmd_scan_bitmap(pool, meta->free_bitmap, &meta->nr_free,
+ base_pfn, nr, want - got, out + got, true);
+ }
+ used = avail_before - dma_pmd_meta_avail(meta);
+ if (!dma_pmd_meta_avail(meta) || WARN_ON_ONCE(!used)) {
+ meta->nr_free = 0;
+ meta->nr_dirty = 0;
+ list_move(&meta->list, &pool->full);
+ }
+
+ return got;
+}
+
+/**
+ * dma_pmd_pool_alloc_bulk_node - Allocate up to @nr_pages blocks from @pool
+ * @pool: Pool to allocate from
+ * @gfp: GFP flags used if a new PMD page must be allocated from buddy
+ * @nid: Target NUMA node if a new PMD page is allocated (or NUMA_NO_NODE)
+ * @nr_pages: Number of blocks requested
+ * @page_array: Output array of at least @nr_pages entries
+ *
+ * Cost: Fast path is O(nr_pages) bitmap scan from @pool->partial or @pool->idle
+ * under a single lock acquisition. Slow path (both empty) calls
+ * dma_pmd_add_page().
+ * Locking: Acquires @pool->lock (irqsave) for the bitmap scan. Drops lock
+ * during slow-path PMD buddy allocation and per-block initialization.
+ * Frequency: High (called per RX buffer refill or per SKB TX page frag refill).
+ *
+ * A pool hands out a block of a PMD page it already owns, on the node that
+ * PMD page came from, so it cannot satisfy a hard placement constraint:
+ * requests with __GFP_THISNODE, __GFP_DMA or __GFP_DMA32 are declined, as is
+ * __GFP_ACCOUNT (pooled blocks are shared across callers and cannot be charged
+ * to a single memcg), so the caller falls back to the page allocator. Flags
+ * that only widen what is acceptable, such as __GFP_HIGHMEM, are ignored
+ * harmlessly.
+ *
+ * Return: Number of blocks placed in @page_array (0 on failure).
+ */
+unsigned long dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp,
+ int nid, unsigned long nr_pages,
+ struct page **page_array)
+{
+ unsigned long allocated = 0, flags, i;
+ bool zero;
+
+ if (unlikely(!pool || pool->destroyed || !nr_pages ||
+ (gfp & (__GFP_DMA | __GFP_DMA32 | __GFP_THISNODE |
+ __GFP_ACCOUNT))))
+ return 0;
+
+ zero = want_init_on_alloc(gfp);
+
+ spin_lock_irqsave(&pool->lock, flags);
+ while (allocated < nr_pages) {
+ struct dma_pmd_meta *meta;
+
+ /* Partially used pages first, idle ones next. */
+ meta = list_first_entry_or_null(&pool->partial, struct dma_pmd_meta, list);
+ if (!meta && !list_empty(&pool->idle)) {
+ meta = list_first_entry(&pool->idle, struct dma_pmd_meta, list);
+ list_move(&meta->list, &pool->partial);
+ pool->num_idle_pages--;
+ }
+ if (!meta) {
+ /* Fallback to a new allocation */
+ spin_unlock_irqrestore(&pool->lock, flags);
+ meta = dma_pmd_add_page(pool, gfp, nid);
+ if (!meta)
+ goto out_prep;
+ spin_lock_irqsave(&pool->lock, flags);
+ pool->pmd_alloc_cnt++;
+ list_add(&meta->list, &pool->partial);
+ }
+
+ allocated += dma_pmd_take_blocks(pool, meta,
+ nr_pages - allocated,
+ &page_array[allocated], zero);
+ }
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+out_prep:
+ for (i = 0; i < allocated; i++) {
+ if (!page_count(page_array[i])) {
+ page_ref_unfreeze(page_array[i], 1);
+ if (zero)
+ memset(page_address(page_array[i]), 0, PAGE_SIZE << pool->order);
+ }
+ }
+ return allocated;
+}
+EXPORT_SYMBOL(dma_pmd_pool_alloc_bulk_node);
+
+/**
+ * dma_pmd_pool_has_free - Can @pool hand out a block without a new PMD page?
+ * @pool: Pool to query
+ *
+ * Cost: O(1) lockless list check.
+ * Locking: None. The answer is advisory and can be stale the moment it is
+ * returned.
+ * Frequency: High; intended for per-buffer decisions on the RX path.
+ *
+ * For callers choosing between an existing non-pooled buffer and a pooled
+ * replacement. Answering "no" when the pool is empty keeps them from trading a
+ * perfectly reusable page for a failed allocation and another singly-mapped
+ * page, which is strictly worse than leaving it alone.
+ *
+ * Return: true if a PMD page with a free block is present.
+ */
+bool dma_pmd_pool_has_free(struct dma_pmd_pool *pool)
+{
+ return pool && !READ_ONCE(pool->destroyed) &&
+ (!list_empty_careful(&pool->partial) || !list_empty_careful(&pool->idle));
+}
+EXPORT_SYMBOL(dma_pmd_pool_has_free);
+
/**
* __dma_pmd_free_page - Recycle a freed block back into its DMA_PMD pool
* @page: Block being freed
diff --git a/include/linux/dma-pmd.h b/include/linux/dma-pmd.h
index 236be50b7349f..085d9a730c278 100644
--- a/include/linux/dma-pmd.h
+++ b/include/linux/dma-pmd.h
@@ -5,6 +5,7 @@
#include <linux/types.h>
#include <linux/mm.h>
+#include <linux/numa.h>
#include <linux/rcupdate.h>
struct device;
@@ -77,6 +78,22 @@ static inline bool dma_pmd_free_page(struct page *page)
*/
struct dma_pmd_pool *dma_pmd_pool_create(unsigned int order, unsigned int max_idle_pages);
struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *pool);
+unsigned long dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp,
+ int nid, unsigned long nr_pages,
+ struct page **page_array);
+static inline struct page *dma_pmd_pool_alloc_node(struct dma_pmd_pool *pool,
+ gfp_t gfp, int nid)
+{
+ struct page *page = NULL;
+
+ dma_pmd_pool_alloc_bulk_node(pool, gfp, nid, 1, &page);
+ return page;
+}
+static inline struct page *dma_pmd_pool_alloc(struct dma_pmd_pool *pool, gfp_t gfp)
+{
+ return dma_pmd_pool_alloc_node(pool, gfp, NUMA_NO_NODE);
+}
+bool dma_pmd_pool_has_free(struct dma_pmd_pool *pool);
#else /* !CONFIG_DMA_PMD */
@@ -101,5 +118,28 @@ static inline struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *poo
return NULL;
}
+static inline unsigned long
+dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp, int nid,
+ unsigned long nr_pages, struct page **page_array)
+{
+ return 0;
+}
+
+static inline struct page *
+dma_pmd_pool_alloc_node(struct dma_pmd_pool *pool, gfp_t gfp, int nid)
+{
+ return NULL;
+}
+
+static inline struct page *dma_pmd_pool_alloc(struct dma_pmd_pool *pool, gfp_t gfp)
+{
+ return NULL;
+}
+
+static inline bool dma_pmd_pool_has_free(struct dma_pmd_pool *pool)
+{
+ return false;
+}
+
#endif /* CONFIG_DMA_PMD */
#endif /* _LINUX_DMA_PMD_H */
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 05/22] iommu/dma: Global cap and shrinker for DMA_PMD pool memory
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (3 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 04/22] iommu/dma: add DMA_PMD pool block allocation Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 06/22] iommu/dma: reserve a per-domain IOVA window for DMA_PMD pages Luigi Rizzo
` (16 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Each pool holds some idle pages, for faster future allocations, but on
large systems there is one pool per CPU and the number may become large.
Add two mechanism to bound the number of unused DMA_PMD pages:
- a global ceiling on the number of DMA_PMD pages pinned across all
pools, defaulting to one eighth of RAM with a floor for small machines
if not configured via /sys/module/kernel/parameters/dma_pmd_max_pages
It is checked before each DMA_PMD page allocation and is meant to bound
a pathological configuration, not to be reached in normal operation.
- a shrinker that, under reclaim, offers every completely idle page in
every pool and retires as many as the reclaim path asks for, stopping
there. Idle pages are detached under the pool lock and retired outside
it, since retiring one unmaps it from the IOMMU.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/dma-pmd-kunit.c | 21 ++++-
drivers/iommu/dma-pmd-pool.c | 146 +++++++++++++++++++++++++++++++++-
drivers/iommu/dma-pmd-priv.h | 2 +
3 files changed, 165 insertions(+), 4 deletions(-)
diff --git a/drivers/iommu/dma-pmd-kunit.c b/drivers/iommu/dma-pmd-kunit.c
index 24978a5123f54..4e38cc5970433 100644
--- a/drivers/iommu/dma-pmd-kunit.c
+++ b/drivers/iommu/dma-pmd-kunit.c
@@ -88,8 +88,8 @@ static void test_meta_pooled_toggle(struct kunit *test)
static void test_pool_alloc_and_recycle(struct kunit *test)
{
- struct dma_pmd_pool *pool;
- struct page *p1, *p2;
+ struct dma_pmd_pool *pool, *pool_wm;
+ struct page *p1, *p2, *b1, *b2, *b3;
unsigned int order;
KUNIT_EXPECT_NULL(test, dma_pmd_pool_destroy(NULL));
@@ -121,6 +121,23 @@ static void test_pool_alloc_and_recycle(struct kunit *test)
if (p2)
__free_pages(p2, order);
}
+
+ /*
+ * Exercise max_idle_2m = 1 watermark release: an order-(PMD_ORDER - 1)
+ * pool has 2 blocks per 2M page, so 3 allocations span two 2M pages.
+ */
+ pool_wm = dma_pmd_pool_create(PMD_ORDER - 1, 1);
+ KUNIT_ASSERT_NOT_NULL(test, pool_wm);
+ b1 = dma_pmd_pool_alloc(pool_wm, GFP_KERNEL);
+ b2 = dma_pmd_pool_alloc(pool_wm, GFP_KERNEL);
+ b3 = dma_pmd_pool_alloc(pool_wm, GFP_KERNEL);
+ if (b1)
+ __free_pages(b1, PMD_ORDER - 1);
+ if (b2)
+ __free_pages(b2, PMD_ORDER - 1);
+ if (b3)
+ __free_pages(b3, PMD_ORDER - 1);
+ dma_pmd_pool_destroy(pool_wm);
}
static struct kunit_case dma_pmd_meta_test_cases[] = {
diff --git a/drivers/iommu/dma-pmd-pool.c b/drivers/iommu/dma-pmd-pool.c
index e6e5ba7c65f27..78f2537772d1d 100644
--- a/drivers/iommu/dma-pmd-pool.c
+++ b/drivers/iommu/dma-pmd-pool.c
@@ -39,6 +39,17 @@ static LLIST_HEAD(dma_pmd_free_list);
/* Dedicated WQ_MEM_RECLAIM workqueue, can run without allocations. */
static struct workqueue_struct *dma_pmd_wq __ro_after_init;
+/*
+ * Running total and global ceiling on DMA_PMD pages pinned across all pools.
+ * If zero, it is set at boot at 1/8 of total RAM.
+ * Without this the bound is per-pool, so the footprint scales with the number
+ * of pools (one per RX queue, two per CPU for sockets) with nothing watching
+ * the aggregate.
+ */
+static atomic_long_t dma_pmd_nr_pages __cacheline_aligned_in_smp;
+static unsigned long dma_pmd_max_pages __read_mostly;
+core_param(dma_pmd_max_pages, dma_pmd_max_pages, ulong, 0644);
+
/* All live pools. */
static LIST_HEAD(dma_pmd_pools);
static DEFINE_MUTEX(dma_pmd_pools_lock);
@@ -101,6 +112,8 @@ static void dma_pmd_release_page_rcu(struct rcu_head *head)
__free_pages(block, order);
}
+ atomic_long_dec(&dma_pmd_nr_pages);
+
spin_lock_irqsave(&pool->lock, flags);
pool->pmd_free_cnt++;
spin_unlock_irqrestore(&pool->lock, flags);
@@ -115,7 +128,7 @@ static void dma_pmd_release_page_rcu(struct rcu_head *head)
* Cost: Slow path; clears the membership bit. The blocks themselves are
* returned after an RCU grace period.
* Locking: Must not hold @pool->lock.
- * Frequency: Rare (background reclaim workqueue, pool destruction).
+ * Frequency: Rare (background reclaim workqueue, shrinker, pool destruction).
*/
static void dma_pmd_release_page(struct dma_pmd_meta *meta)
{
@@ -168,7 +181,9 @@ static void dma_pmd_schedule_reclaim(void)
* @order: Block order to dispense, must be <= PMD_ORDER.
* @max_idle_pages: Maximum number of completely idle PMD pages to keep cached in
* the pool before asynchronously releasing excess PMD pages to buddy
- * (0 uses default of 16 = 32 MB).
+ * (0 uses default of 16 = 32 MB). This is a per-pool bound; the
+ * aggregate across all pools is additionally capped by
+ * dma_pmd_max_pages and trimmed by the shrinker.
*
* New PMD physical pages are allocated on demand on the NUMA node of the
* calling CPU (the NAPI CPU for RX queue refill, or the application CPU for TX).
@@ -294,6 +309,15 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
if (!can_block && time_before(jiffies, READ_ONCE(pool->next_alloc_attempt)))
return NULL;
+ /* Aggregate ceiling across every pool, independent of max_idle_pages. */
+ if (atomic_long_inc_return(&dma_pmd_nr_pages) > READ_ONCE(dma_pmd_max_pages)) {
+ atomic_long_dec(&dma_pmd_nr_pages);
+ if (!can_block)
+ WRITE_ONCE(pool->next_alloc_attempt,
+ jiffies + DIV_ROUND_UP(HZ, 100));
+ return NULL;
+ }
+
/*
* The following flags are discarded:
*
@@ -324,6 +348,7 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
page = alloc_pages_node(nid, alloc_gfp, PMD_ORDER);
if (!page) {
+ atomic_long_dec(&dma_pmd_nr_pages);
/* Retry backoff; could be made shorter or tunable in future. */
if (!can_block)
WRITE_ONCE(pool->next_alloc_attempt,
@@ -338,12 +363,14 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
*/
if (unlikely(!dma_pmd_meta_ensure_pfn(page_to_pfn(page), can_block))) {
__free_pages(page, PMD_ORDER);
+ atomic_long_dec(&dma_pmd_nr_pages);
return NULL;
}
/* Fails with -EBUSY if a concurrent PFN walker holds a speculative ref. */
if (split_page_compound(page, PMD_ORDER, pool->order)) {
__free_pages(page, PMD_ORDER);
+ atomic_long_dec(&dma_pmd_nr_pages);
return NULL;
}
@@ -642,8 +669,112 @@ bool __dma_pmd_free_page(struct page *page)
}
EXPORT_SYMBOL(__dma_pmd_free_page);
+/*
+ * Idle PMD pages are retained per pool up to max_idle_pages, which is a throughput
+ * knob and deliberately does not react to memory pressure. The shrinker is
+ * what makes that safe: under reclaim every completely idle PMD page in every
+ * pool becomes available, so the pools give the memory back instead of
+ * pinning their high-water mark for the lifetime of the machine.
+ *
+ * Both callbacks use mutex_trylock() rather than mutex_lock(): a thread that
+ * already holds dma_pmd_pools_lock can enter direct reclaim from an
+ * allocation it makes under that lock, and reclaim calls straight back in
+ * here. Giving up is always correct, since there is nothing this shrinker
+ * could free that the holder is not already about to account for.
+ */
+static unsigned long dma_pmd_shrink_count(struct shrinker *shrink,
+ struct shrink_control *sc)
+{
+ struct dma_pmd_pool *pool;
+ unsigned long nr = 0;
+
+ if (!mutex_trylock(&dma_pmd_pools_lock))
+ return 0;
+
+ list_for_each_entry(pool, &dma_pmd_pools, node)
+ nr += READ_ONCE(pool->num_idle_pages);
+
+ mutex_unlock(&dma_pmd_pools_lock);
+
+ return nr << PMD_ORDER;
+}
+
+static unsigned long dma_pmd_shrink_scan(struct shrinker *shrink,
+ struct shrink_control *sc)
+{
+ struct dma_pmd_meta *meta, *tmp;
+ struct dma_pmd_pool *pool;
+ unsigned long freed = 0;
+ LIST_HEAD(release_list);
+ unsigned long flags;
+
+ if (!mutex_trylock(&dma_pmd_pools_lock))
+ return SHRINK_STOP;
+
+ list_for_each_entry(pool, &dma_pmd_pools, node) {
+ /*
+ * Every PMD page on @idle is a candidate, so this detaches one
+ * per iteration and stops as soon as the reclaim path has
+ * what it asked for: the IRQs-off section is bounded by the
+ * work done, not by how many PMD pages the pool holds.
+ */
+ spin_lock_irqsave(&pool->lock, flags);
+ list_for_each_entry_safe(meta, tmp, &pool->idle, list) {
+ if (freed >= sc->nr_to_scan)
+ break;
+ list_move(&meta->list, &release_list);
+ pool->num_idle_pages--;
+ freed += 1UL << PMD_ORDER;
+ }
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+ if (freed >= sc->nr_to_scan)
+ break;
+ }
+
+ mutex_unlock(&dma_pmd_pools_lock);
+
+ /* Retiring must not run under @pool->lock. */
+ list_for_each_entry_safe(meta, tmp, &release_list, list) {
+ list_del(&meta->list);
+ dma_pmd_release_page(meta);
+ }
+
+ /*
+ * Both counts are in pages, matching dma_pmd_shrink_count(). A
+ * 2M page is 512 of them and cannot be freed in parts, so a
+ * SHRINK_BATCH sized request always overshoots. Report what was
+ * really reclaimed so that do_shrink_slab() charges its budget for
+ * the whole PMD page instead of the 128 pages it asked for, and does
+ * not come back four times over for the same pressure.
+ *
+ * Zero scanned pages would leave that budget untouched and spin,
+ * so a scan that reclaimed nothing has to say SHRINK_STOP.
+ */
+ sc->nr_scanned = freed;
+
+ return freed ?: SHRINK_STOP;
+}
+
static int __init dma_pmd_init(void)
{
+ struct shrinker *shrink;
+
+ /*
+ * Hard ceiling on memory diverted from the buddy allocator into 2MB
+ * pools, in PMD pages. One eighth of RAM is far above what the per-pool
+ * watermarks should ever reach; it exists to bound a pathological
+ * configuration (many queues, many CPUs, all pools at their high
+ * water mark) rather than to be hit in normal operation. The floor
+ * keeps small machines usable. Zero here means the dma_pmd_max_pages=
+ * boot parameter did not override it.
+ *
+ * Lowering the cap later only stops pools from growing; the PMD pages
+ * already held are returned by the shrinker as they fall idle.
+ */
+ if (!dma_pmd_max_pages)
+ dma_pmd_max_pages = max(totalram_pages() >> (PMD_ORDER + 3), 16UL);
+
/*
* Reclaim frees memory, so it must not queue behind arbitrary work on
* system_wq when the machine is already short of it. Failure is not
@@ -653,6 +784,17 @@ static int __init dma_pmd_init(void)
if (!dma_pmd_wq)
pr_warn("dma_pmd: no reclaim workqueue, falling back to system_wq\n");
+ shrink = shrinker_alloc(0, "dma_pmd");
+ if (!shrink) {
+ pr_warn("dma_pmd: shrinker registration failed\n");
+ return 0;
+ }
+
+ shrink->count_objects = dma_pmd_shrink_count;
+ shrink->scan_objects = dma_pmd_shrink_scan;
+ shrink->seeks = DEFAULT_SEEKS;
+ shrinker_register(shrink);
+
return 0;
}
subsys_initcall(dma_pmd_init);
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index 99bd7d77fd4ef..39295d0add51d 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -21,6 +21,8 @@
* when can_block is true. It is not nested with pool->lock
* or meta->map_lock, and sits inside dma_pmd_pools_lock
* when dma_pmd_pool_create() calls dma_pmd_meta_init().
+ * The only reverse edge on dma_pmd_pools_lock is the
+ * shrinker's mutex_trylock(), which cannot block.
*
* dma_pmd_pools_lock Mutex over the list of live pools. Outermost of the
* three, and it must not be taken from a context that
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 06/22] iommu/dma: reserve a per-domain IOVA window for DMA_PMD pages
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (4 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 05/22] iommu/dma: Global cap and shrinker for DMA_PMD pool memory Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 07/22] iommu/dma: release DMA_PMD domain mappings on domain teardown Luigi Rizzo
` (15 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Embed struct dma_pmd_window in struct iommu_dma_cookie to record the
per-domain IOVA range reserved for DMA_PMD mappings, and expose the
cookie accessors (dma_pmd_dma_window(), dma_pmd_dma_iovad(),
dma_pmd_window_owns()) and dma_info_to_prot() needed by the DMA_PMD
mapping layer in subsequent patches.
No functional change.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/dma-iommu.c | 39 ++++++++++++++++++-
drivers/iommu/dma-iommu.h | 8 ++++
drivers/iommu/dma-pmd-kunit.c | 19 +++++++++
drivers/iommu/dma-pmd-priv.h | 72 +++++++++++++++++++++++++++++++++++
4 files changed, 137 insertions(+), 1 deletion(-)
diff --git a/drivers/iommu/dma-iommu.c b/drivers/iommu/dma-iommu.c
index 58c624513cd43..70aee7a3020c8 100644
--- a/drivers/iommu/dma-iommu.c
+++ b/drivers/iommu/dma-iommu.c
@@ -18,6 +18,7 @@
#include <linux/gfp.h>
#include <linux/huge_mm.h>
#include <linux/iommu.h>
+#include <linux/dma-pmd.h>
#include <linux/iommu-dma.h>
#include <linux/iova.h>
#include <linux/irq.h>
@@ -36,6 +37,7 @@
#include <trace/events/swiotlb.h>
#include "dma-iommu.h"
+#include "dma-pmd-priv.h"
#include "iommu-pages.h"
struct iommu_dma_msi_page {
@@ -75,6 +77,8 @@ struct iommu_dma_cookie {
struct iommu_domain *fq_domain;
/* Options for dma-iommu use */
struct iommu_dma_options options;
+ /* IOVA window reserved for DMA_PMD pages. Nothing else can go here. */
+ struct dma_pmd_window win_dma_pmd;
};
struct iommu_dma_msi_cookie {
@@ -732,7 +736,7 @@ static int iommu_dma_init_domain(struct iommu_domain *domain, struct device *dev
*
* Return: corresponding IOMMU API page protection flags
*/
-static int dma_info_to_prot(enum dma_data_direction dir, bool coherent,
+int dma_info_to_prot(enum dma_data_direction dir, bool coherent,
unsigned long attrs)
{
int prot;
@@ -1214,6 +1218,39 @@ static inline size_t iova_unaligned(struct iova_domain *iovad, phys_addr_t phys,
return iova_offset(iovad, phys | size);
}
+/**
+ * dma_pmd_dma_window - This domain's IOVA window for DMA_PMD
+ * @domain: Domain to look at
+ *
+ * For the DMA_PMD code's slow paths, which need a window but do not have the
+ * cookie layout. The DMA map path is handed the pointer by its caller instead,
+ * so this is never called at map frequency.
+ *
+ * Return: the window, or NULL if @domain is NULL or does not carry a DMA-IOVA
+ * cookie. The cookie shares a union with the MSI, iommufd and fault-handler
+ * ones, so the type has to be tested rather than the pointer. An identity,
+ * passthrough, or MSI-only domain has no window, and a device can be moved to
+ * one after its pool was created (e.g. via sysfs domain type changes or VFIO
+ * attachment).
+ */
+struct dma_pmd_window *dma_pmd_dma_window(struct iommu_domain *domain)
+{
+ /*
+ * Also decline in kdump kernels (iommu_deferred_attach_enabled);
+ * dma_pmd_dma_iovad() uses this check so win->size stays 0.
+ */
+ if (static_branch_unlikely(&iommu_deferred_attach_enabled) ||
+ !domain || domain->cookie_type != IOMMU_COOKIE_DMA_IOVA)
+ return NULL;
+
+ return &domain->iova_cookie->win_dma_pmd;
+}
+
+struct iova_domain *dma_pmd_dma_iovad(struct iommu_domain *domain)
+{
+ return dma_pmd_dma_window(domain) ? &domain->iova_cookie->iovad : NULL;
+}
+
dma_addr_t iommu_dma_map_phys(struct device *dev, phys_addr_t phys, size_t size,
enum dma_data_direction dir, unsigned long attrs)
{
diff --git a/drivers/iommu/dma-iommu.h b/drivers/iommu/dma-iommu.h
index 040d002525632..241f8de002e50 100644
--- a/drivers/iommu/dma-iommu.h
+++ b/drivers/iommu/dma-iommu.h
@@ -7,6 +7,9 @@
#include <linux/iommu.h>
+struct dma_pmd_window;
+struct iova_domain;
+
#ifdef CONFIG_IOMMU_DMA
void iommu_setup_dma_ops(struct device *dev, struct iommu_domain *domain);
@@ -22,6 +25,11 @@ void iommu_dma_get_resv_regions(struct device *dev, struct list_head *list);
int iommu_dma_sw_msi(struct iommu_domain *domain, struct msi_desc *desc,
phys_addr_t msi_addr);
+int dma_info_to_prot(enum dma_data_direction dir, bool coherent, unsigned long attrs);
+
+struct iova_domain *dma_pmd_dma_iovad(struct iommu_domain *domain);
+struct dma_pmd_window *dma_pmd_dma_window(struct iommu_domain *domain);
+
extern bool iommu_dma_forcedac;
#else /* CONFIG_IOMMU_DMA */
diff --git a/drivers/iommu/dma-pmd-kunit.c b/drivers/iommu/dma-pmd-kunit.c
index 4e38cc5970433..16312014b8976 100644
--- a/drivers/iommu/dma-pmd-kunit.c
+++ b/drivers/iommu/dma-pmd-kunit.c
@@ -140,11 +140,30 @@ static void test_pool_alloc_and_recycle(struct kunit *test)
dma_pmd_pool_destroy(pool_wm);
}
+static void test_window_helpers(struct kunit *test)
+{
+ struct dma_pmd_window win = {};
+
+ KUNIT_ASSERT_EQ(test, dma_pmd_meta_init(), 0);
+
+ /* Empty window (size == 0) never owns any IOVA. */
+ KUNIT_EXPECT_FALSE(test, dma_pmd_window_owns(&win, 0));
+ KUNIT_EXPECT_FALSE(test, dma_pmd_window_owns(&win, SZ_4G));
+
+ win.base = SZ_4G;
+ win.size = SZ_2G;
+ KUNIT_EXPECT_FALSE(test, dma_pmd_window_owns(&win, SZ_4G - 1));
+ KUNIT_EXPECT_TRUE(test, dma_pmd_window_owns(&win, SZ_4G));
+ KUNIT_EXPECT_TRUE(test, dma_pmd_window_owns(&win, SZ_4G + SZ_2G - 1));
+ KUNIT_EXPECT_FALSE(test, dma_pmd_window_owns(&win, SZ_4G + SZ_2G));
+}
+
static struct kunit_case dma_pmd_meta_test_cases[] = {
KUNIT_CASE(test_meta_init_and_roundtrip),
KUNIT_CASE(test_meta_invalid_phys),
KUNIT_CASE(test_meta_pooled_toggle),
KUNIT_CASE(test_pool_alloc_and_recycle),
+ KUNIT_CASE(test_window_helpers),
{}
};
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index 39295d0add51d..8104897572921 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -9,10 +9,46 @@
#include <linux/spinlock.h>
#include <linux/types.h>
+/**
+ * struct dma_pmd_window - A domain's IOVA window reserved for DMA_PMD pages
+ * @base: First IOVA of the window. A page at @phys is mapped, in every domain
+ * that has a window, at @base + @phys.
+ * @size: Window size in bytes, or 0 if this domain has no window, either
+ * because nothing has pooled through it yet, or because it could not
+ * find a free range that large. 0 must make both the map and the unmap
+ * side decline, see dma_pmd_window_owns().
+ * @domain_idx: Dense domain index + DMA_PMD_IDX_FIRST, used for per-PMD-page
+ * presence (bit @domain_idx - DMA_PMD_IDX_FIRST), or one of
+ * DMA_PMD_IDX_NONE, DMA_PMD_IDX_NOHUGE or DMA_PMD_IDX_NOSPACE while
+ * @size is 0. Of the three sentinels, only DMA_PMD_IDX_NONE is retried.
+ *
+ * Lives in the domain's iommu_dma_cookie. The window is a real allocation out
+ * of the domain's iova_domain, so no ordinary IOVA can ever fall inside it,
+ * which is what lets the unmap path recognise a pooled IOVA by range alone.
+ */
+struct dma_pmd_window {
+ dma_addr_t base;
+ u64 size;
+ u16 domain_idx;
+};
+
#ifdef CONFIG_DMA_PMD
#define DMA_PMD_BLOCKS(order) (1U << (PMD_ORDER - (order)))
+/*
+ * Sentinel values of dma_pmd_window.domain_idx below DMA_PMD_IDX_FIRST:
+ * never attempted (0, matching kzalloc), or permanently declined. A real
+ * domain index is >= DMA_PMD_IDX_FIRST (subtract DMA_PMD_IDX_FIRST for
+ * the bitmap bit position).
+ */
+enum {
+ DMA_PMD_IDX_NONE, /* not attempted yet */
+ DMA_PMD_IDX_NOHUGE, /* @domain can never pool */
+ DMA_PMD_IDX_NOSPACE, /* no index or no IOVA range */
+ DMA_PMD_IDX_FIRST, /* first valid domain index */
+};
+
/*
* Locking
* -------
@@ -167,5 +203,41 @@ struct dma_pmd_meta *dma_pmd_meta_from_phys(phys_addr_t pa);
unsigned long dma_pmd_meta_to_pfn(const struct dma_pmd_meta *m);
phys_addr_t dma_pmd_meta_to_phys(const struct dma_pmd_meta *m);
+static inline bool dma_is_pmd_phys(phys_addr_t phys)
+{
+ return dma_is_pmd_page(phys >> PAGE_SHIFT);
+}
+
+/**
+ * dma_pmd_window_owns - Was @dma handed out by a DMA_PMD mapping?
+ * @win: The unmapping domain's window
+ * @dma: IOVA being unmapped
+ *
+ * Only DMA_PMD pages can be mapped within the reserved IOVA window
+ * so a range check suffices.
+ * A domain that has no window has @size 0, so the unsigned subtraction
+ * underflows to a huge value and the test is always false. That is the same
+ * check that stops the map path from using a window the domain does not own.
+ */
+static inline bool dma_pmd_window_owns(const struct dma_pmd_window *win, dma_addr_t dma)
+{
+ /* Acquire pairs with the release of @size in dma_pmd_window_assign(). */
+ u64 size = smp_load_acquire(&win->size);
+
+ return size && dma - win->base < size;
+}
+
+#else /* !CONFIG_DMA_PMD */
+
+static inline bool dma_is_pmd_phys(phys_addr_t phys)
+{
+ return false;
+}
+
+static inline bool dma_pmd_window_owns(const struct dma_pmd_window *win, dma_addr_t dma)
+{
+ return false;
+}
+
#endif /* CONFIG_DMA_PMD */
#endif /* _DRIVERS_IOMMU_DMA_PMD_PRIV_H */
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 07/22] iommu/dma: release DMA_PMD domain mappings on domain teardown
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (5 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 06/22] iommu/dma: reserve a per-domain IOVA window for DMA_PMD pages Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 08/22] iommu/dma: use per-domain IOVA window to map DMA_PMD memory Luigi Rizzo
` (14 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Track per-domain PMD_SIZE leaf PTE presence in a 64-entry domain registry
(dma_pmd_domains[]) and a per-PMD-page bitmask (p2m->domains_mapped),
unmapping active domains in dma_pmd_unmap_all() when a DMA_PMD page is
retired.
When an IOMMU DMA domain is torn down at runtime (for example across
VFIO bind/unbind or Dynamic Memory Security domain switches), its page
tables and IOVA domain are freed by the caller, but DMA_PMD pages that were
mapped into it may outlive the domain in per-queue or per-CPU pools.
Add dma_pmd_domain_release() and call it from iommu_put_dma_cookie()
before freeing the cookie:
- Clear the domain's slot in dma_pmd_domains[] and mark it draining so
a racing dma_pmd_unmap_all() skips the dying domain instead of
calling iommu_unmap() on freed page tables.
- Walk all live pools (including destroyed pools that still hold
in-flight DMA_PMD pages via pool->refcount) and clear the domain's bit in
every p2m->domains_mapped bitmap via dma_pmd_forget_domain().
- Flush dma_pmd_reclaim_work and wait on dma_pmd_srcu so any in-flight
unmapper finishes before the domain index is released for reuse.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/Makefile | 2 +-
drivers/iommu/dma-iommu.c | 6 +
drivers/iommu/dma-pmd-kunit.c | 8 ++
drivers/iommu/dma-pmd-map.c | 220 ++++++++++++++++++++++++++++++++++
drivers/iommu/dma-pmd-pool.c | 59 ++++++---
drivers/iommu/dma-pmd-priv.h | 34 ++++++
6 files changed, 314 insertions(+), 15 deletions(-)
create mode 100644 drivers/iommu/dma-pmd-map.c
diff --git a/drivers/iommu/Makefile b/drivers/iommu/Makefile
index de2c3bad9aa3e..e701b36b9b605 100644
--- a/drivers/iommu/Makefile
+++ b/drivers/iommu/Makefile
@@ -11,7 +11,7 @@ obj-$(CONFIG_IOMMU_API) += iommu-traces.o
obj-$(CONFIG_IOMMU_API) += iommu-sysfs.o
obj-$(CONFIG_IOMMU_DEBUGFS) += iommu-debugfs.o
obj-$(CONFIG_IOMMU_DMA) += dma-iommu.o
-obj-$(CONFIG_DMA_PMD) += dma-pmd-meta.o dma-pmd-pool.o
+obj-$(CONFIG_DMA_PMD) += dma-pmd-meta.o dma-pmd-pool.o dma-pmd-map.o
obj-$(CONFIG_DMA_PMD_META_KUNIT_TEST) += dma-pmd-kunit.o
obj-$(CONFIG_IOMMU_IO_PGTABLE) += io-pgtable.o
obj-$(CONFIG_IOMMU_IO_PGTABLE_ARMV7S) += io-pgtable-arm-v7s.o
diff --git a/drivers/iommu/dma-iommu.c b/drivers/iommu/dma-iommu.c
index 70aee7a3020c8..ee14175b79bc2 100644
--- a/drivers/iommu/dma-iommu.c
+++ b/drivers/iommu/dma-iommu.c
@@ -430,6 +430,12 @@ void iommu_put_dma_cookie(struct iommu_domain *domain)
struct iommu_dma_cookie *cookie = domain->iova_cookie;
struct iommu_dma_msi_page *msi, *tmp;
+ /*
+ * Drop any DMA_PMD mappings cached against this domain before its
+ * page tables and IOVA domain go away.
+ */
+ dma_pmd_domain_release(domain);
+
if (cookie->iovad.granule) {
iommu_dma_free_fq(cookie);
put_iova_domain(&cookie->iovad);
diff --git a/drivers/iommu/dma-pmd-kunit.c b/drivers/iommu/dma-pmd-kunit.c
index 16312014b8976..7f4251845ba49 100644
--- a/drivers/iommu/dma-pmd-kunit.c
+++ b/drivers/iommu/dma-pmd-kunit.c
@@ -5,6 +5,7 @@
#include <kunit/test.h>
#include <linux/gfp.h>
#include <linux/dma-pmd.h>
+#include <linux/iommu.h>
#include <linux/mm.h>
#include "dma-pmd-priv.h"
@@ -142,7 +143,9 @@ static void test_pool_alloc_and_recycle(struct kunit *test)
static void test_window_helpers(struct kunit *test)
{
+ struct iommu_domain dummy_domain = {};
struct dma_pmd_window win = {};
+ struct dma_pmd_pool *pool;
KUNIT_ASSERT_EQ(test, dma_pmd_meta_init(), 0);
@@ -156,6 +159,11 @@ static void test_window_helpers(struct kunit *test)
KUNIT_EXPECT_TRUE(test, dma_pmd_window_owns(&win, SZ_4G));
KUNIT_EXPECT_TRUE(test, dma_pmd_window_owns(&win, SZ_4G + SZ_2G - 1));
KUNIT_EXPECT_FALSE(test, dma_pmd_window_owns(&win, SZ_4G + SZ_2G));
+ /* Releasing an unregistered domain with an active pool is a safe no-op. */
+ pool = dma_pmd_pool_create(0, 1);
+ KUNIT_ASSERT_NOT_NULL(test, pool);
+ dma_pmd_domain_release(&dummy_domain);
+ dma_pmd_pool_destroy(pool);
}
static struct kunit_case dma_pmd_meta_test_cases[] = {
diff --git a/drivers/iommu/dma-pmd-map.c b/drivers/iommu/dma-pmd-map.c
new file mode 100644
index 0000000000000..cffa94a7c7c84
--- /dev/null
+++ b/drivers/iommu/dma-pmd-map.c
@@ -0,0 +1,220 @@
+// SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause
+/*
+ * DMA_PMD per-domain IOVA window management and IOMMU mapping.
+ *
+ * See Documentation/core-api/dma-pmd.rst for the architecture overview.
+ */
+
+#include <linux/atomic.h>
+#include <linux/bitops.h>
+#include <linux/cache.h>
+#include <linux/dma-map-ops.h>
+#include <linux/dma-mapping.h>
+#include <linux/dma-pmd.h>
+#include <linux/export.h>
+#include <linux/iommu.h>
+#include <linux/mm.h>
+#include <linux/spinlock.h>
+#include <linux/srcu.h>
+
+#include "dma-iommu.h"
+#include "dma-pmd-priv.h"
+
+/*
+ * Index -> domain, for PMD pages that have to unmap themselves from every domain
+ * that installed them. A slot is cleared before the matching bits are cleared
+ * in dma_pmd_domain_release(), so a PMD page racing with a dying domain sees
+ * NULL and correctly leaves the doomed page tables alone.
+ *
+ * @base is cached here rather than read back from the domain's DMA cookie, so
+ * that teardown never has to dereference a cookie that may be about to be
+ * freed.
+ *
+ * The lock also serialises index allocation, which happens from the DMA map
+ * path and so cannot sleep.
+ */
+static struct {
+ struct iommu_domain *domain;
+ dma_addr_t base;
+ bool draining;
+} dma_pmd_domains[DMA_PMD_MAX_DOMAINS] __read_mostly;
+static DEFINE_SPINLOCK(dma_pmd_domains_lock);
+
+/*
+ * Held by dma_pmd_unmap_all() across the registry read and the iommu_unmap()
+ * that follows, and waited on by dma_pmd_domain_release() once it has cleared
+ * the slot. Without it an unmapper that sampled a domain a moment before the
+ * slot was cleared would walk into page tables their owner is already freeing.
+ * SRCU rather than RCU to avoid holding RCU across IOTLB flushes.
+ */
+DEFINE_SRCU(dma_pmd_srcu);
+
+/**
+ * dma_pmd_unmap_all - Remove a PMD page's PTE from every domain holding it
+ * @meta: PMD page metadata structure
+ *
+ * The window itself is a permanent allocation out of each domain's
+ * iova_domain and is deliberately left alone; only the leaf PTE goes away.
+ *
+ * The bits are cleared before the unmaps, not after. A mapper that observes a
+ * cleared bit takes the slow path and blocks on @meta->map_lock, so it either
+ * re-installs the PTE or waits; the reverse order would leave a window in
+ * which the bit still claims a translation that has already been torn down.
+ *
+ * A domain that has since been released has a NULL registry slot and is
+ * skipped: its page tables are being freed by their owner and must not be
+ * touched here.
+ *
+ * Cost: Slow path (teardown / shrinker only); one iommu_unmap() per domain.
+ * Locking: Acquires @meta->map_lock and @dma_pmd_domains_lock (irqsave), and
+ * drops both before unmapping. Holds @dma_pmd_srcu across the
+ * registry read and the unmap, which is what keeps the domain alive
+ * in between.
+ * Frequency: Rare (only when a PMD page is released back to the buddy
+ * allocator or when a pool is destroyed).
+ */
+void dma_pmd_unmap_all(struct dma_pmd_meta *meta)
+{
+ phys_addr_t phys = dma_pmd_meta_to_phys(meta);
+ unsigned long mapped, flags;
+ int idx, srcu_idx;
+
+ srcu_idx = srcu_read_lock(&dma_pmd_srcu);
+ spin_lock_irqsave(&meta->map_lock, flags);
+ mapped = meta->domains_mapped;
+ meta->domains_mapped = 0;
+ spin_unlock_irqrestore(&meta->map_lock, flags);
+
+ if (!mapped) {
+ srcu_read_unlock(&dma_pmd_srcu, srcu_idx);
+ return;
+ }
+
+ for_each_set_bit(idx, &mapped, DMA_PMD_MAX_DOMAINS) {
+ struct iommu_domain *domain;
+ dma_addr_t base;
+
+ spin_lock_irqsave(&dma_pmd_domains_lock, flags);
+ domain = dma_pmd_domains[idx].domain;
+ base = dma_pmd_domains[idx].base;
+ spin_unlock_irqrestore(&dma_pmd_domains_lock, flags);
+
+ if (!domain)
+ continue;
+
+ iommu_unmap(domain, base + phys, PMD_SIZE);
+ }
+ srcu_read_unlock(&dma_pmd_srcu, srcu_idx);
+}
+
+/**
+ * dma_pmd_forget_domain - Drop a PMD page's PTE record for a dying domain
+ * @meta: PMD page metadata structure
+ * @idx: Index of the domain being torn down
+ *
+ * Deliberately does not unmap: @domain is being destroyed by its owner, which
+ * frees the page tables and the IOVA domain wholesale. Touching it here is
+ * exactly the use-after-free this is meant to prevent, so the bit is only
+ * forgotten.
+ *
+ * Locking: Acquires @meta->map_lock (irqsave).
+ *
+ * Return: 1 if this PMD page had a mapping in @idx, 0 otherwise.
+ */
+unsigned int dma_pmd_forget_domain(struct dma_pmd_meta *meta, int idx)
+{
+ unsigned int dropped = 0;
+ unsigned long flags;
+
+ if (!(READ_ONCE(meta->domains_mapped) & BIT(idx)))
+ return 0;
+
+ spin_lock_irqsave(&meta->map_lock, flags);
+ if (meta->domains_mapped & BIT(idx)) {
+ meta->domains_mapped &= ~BIT(idx);
+ dropped = 1;
+ }
+ spin_unlock_irqrestore(&meta->map_lock, flags);
+
+ return dropped;
+}
+
+/**
+ * dma_pmd_domain_release - Forget every PMD page mapping for @domain
+ * @domain: DMA domain being destroyed
+ *
+ * A PMD page records, as a bit per domain index, which domains hold its PMD leaf
+ * PTE. Those bits are otherwise only cleared when the PMD page itself is
+ * released, so a domain that comes and goes - which this tree does at runtime,
+ * see the DMS comments in __iommu_dma_map_phys() - would permanently consume
+ * one index per generation, and every PMD page that ever mapped through it would
+ * keep claiming a mapping in page tables that no longer exist.
+ *
+ * The registry slot is cleared first and the per-PMD-page bits afterwards.
+ * dma_pmd_unmap_all() reads the bit and then the slot, so that order means it
+ * either sees no bit, or sees a NULL slot and skips: it can never unmap
+ * through a domain whose page tables its owner is freeing. An unmapper that
+ * sampled the slot just before it was cleared is waited out on @dma_pmd_srcu.
+ *
+ * The slot stays reserved until the walk and drain are done, so the index
+ * cannot be handed to a new domain while PMD pages still carry its bit.
+ *
+ * Cost: O(PMD pages) across all pools, but only on domain teardown.
+ * Locking: Takes @dma_pmd_domains_lock, then @dma_pmd_pools_lock, then each
+ * @pool->lock, then each @meta->map_lock. Must be called from process
+ * context.
+ */
+void dma_pmd_domain_release(struct iommu_domain *domain)
+{
+ unsigned int dropped = 0;
+ unsigned long flags;
+ int idx;
+
+ /* Nothing has ever been pooled, so nothing can reference @domain. */
+ if (likely(!dma_pmd_meta_base()))
+ return;
+
+ spin_lock_irqsave(&dma_pmd_domains_lock, flags);
+ for (idx = 0; idx < DMA_PMD_MAX_DOMAINS; idx++) {
+ if (dma_pmd_domains[idx].domain == domain) {
+ dma_pmd_domains[idx].domain = NULL;
+ /*
+ * Keep the slot reserved until every PMD page below has
+ * forgotten it. A new domain handed this index now
+ * would inherit the stale bits, and its first map of
+ * such a PMD page would skip iommu_map() and return an
+ * IOVA with no PTE behind it.
+ */
+ dma_pmd_domains[idx].draining = true;
+ break;
+ }
+ }
+ spin_unlock_irqrestore(&dma_pmd_domains_lock, flags);
+
+ /* @domain never pooled, so no PMD page can be holding a bit for it. */
+ if (idx == DMA_PMD_MAX_DOMAINS)
+ return;
+
+ dropped += dma_pmd_pools_forget_domain(idx);
+
+ /*
+ * Any PMD page removed from a pool list before the walk above is either
+ * already on @dma_pmd_free_list (placed there under @pool->lock in
+ * __dma_pmd_free_page()) or being unmapped under @dma_pmd_srcu (in
+ * the shrinker or dma_pmd_pool_destroy()). Drain the reclaim list and
+ * wait out @dma_pmd_srcu before allowing @idx to be reused.
+ *
+ * synchronize_srcu() also waits out any dma_pmd_unmap_all() that
+ * sampled @domain before the slot was cleared above.
+ */
+ synchronize_srcu(&dma_pmd_srcu);
+
+ /* No PMD page claims @idx any more, so it can be reused. */
+ spin_lock_irqsave(&dma_pmd_domains_lock, flags);
+ dma_pmd_domains[idx].draining = false;
+ spin_unlock_irqrestore(&dma_pmd_domains_lock, flags);
+
+ if (dropped)
+ pr_debug("dma_pmd: released %u PMD page mappings for domain %p (index %d)\n",
+ dropped, domain, idx);
+}
diff --git a/drivers/iommu/dma-pmd-pool.c b/drivers/iommu/dma-pmd-pool.c
index 78f2537772d1d..1d4244ba422a3 100644
--- a/drivers/iommu/dma-pmd-pool.c
+++ b/drivers/iommu/dma-pmd-pool.c
@@ -74,9 +74,8 @@ static void dma_pmd_pool_free_kref(struct kref *kref)
* pushed locklessly onto dma_pmd_free_list and dma_pmd_reclaim_work is
* scheduled on dma_pmd_wq.
* 2. In process context, dma_pmd_release_page() unmaps all IOMMU domains
- * (dma_pmd_unmap_all(), added in a later commit), clears meta->pooled so
- * new lockless readers stop entering @meta, and queues
- * dma_pmd_release_page_rcu() via call_rcu().
+ * (dma_pmd_unmap_all()), clears meta->pooled so new lockless readers stop
+ * entering @meta, and queues dma_pmd_release_page_rcu() via call_rcu().
* 3. After an RCU grace period (once no concurrent dma_pmd_free_page() reader
* can still be dereferencing @meta), dma_pmd_release_page_rcu() unfreezes
* all order-pool->order blocks (already split by split_page_compound()),
@@ -125,14 +124,14 @@ static void dma_pmd_release_page_rcu(struct rcu_head *head)
* dma_pmd_release_page - Retire a PMD page and schedule its buddy release
* @meta: PMD page metadata structure (must have all usable blocks idle)
*
- * Cost: Slow path; clears the membership bit. The blocks themselves are
- * returned after an RCU grace period.
- * Locking: Must not hold @pool->lock.
+ * Cost: Slow path; unmaps the IOMMU domains and clears the membership bit.
+ * The blocks themselves are returned after an RCU grace period.
+ * Locking: Process context (calls iommu_unmap). Must not hold @pool->lock.
* Frequency: Rare (background reclaim workqueue, shrinker, pool destruction).
*/
static void dma_pmd_release_page(struct dma_pmd_meta *meta)
{
- /* dma_pmd_unmap_all(meta) will be called here once IOMMU mappings are added. */
+ dma_pmd_unmap_all(meta);
/*
* Clear membership before the grace period. A reader that passed
@@ -176,6 +175,31 @@ static void dma_pmd_schedule_reclaim(void)
schedule_work(&dma_pmd_reclaim_work);
}
+unsigned int dma_pmd_pools_forget_domain(int idx)
+{
+ struct dma_pmd_pool *pool;
+ struct dma_pmd_meta *meta;
+ unsigned int dropped = 0;
+ unsigned long flags;
+
+ mutex_lock(&dma_pmd_pools_lock);
+ list_for_each_entry(pool, &dma_pmd_pools, node) {
+ spin_lock_irqsave(&pool->lock, flags);
+ list_for_each_entry(meta, &pool->partial, list)
+ dropped += dma_pmd_forget_domain(meta, idx);
+ list_for_each_entry(meta, &pool->idle, list)
+ dropped += dma_pmd_forget_domain(meta, idx);
+ list_for_each_entry(meta, &pool->full, list)
+ dropped += dma_pmd_forget_domain(meta, idx);
+ spin_unlock_irqrestore(&pool->lock, flags);
+ }
+ mutex_unlock(&dma_pmd_pools_lock);
+
+ dma_pmd_schedule_reclaim();
+ flush_work(&dma_pmd_reclaim_work);
+ return dropped;
+}
+
/**
* dma_pmd_pool_create - Create a DMA_PMD page pool
* @order: Block order to dispense, must be <= PMD_ORDER.
@@ -232,12 +256,10 @@ EXPORT_SYMBOL(dma_pmd_pool_create);
* dma_pmd_pool_destroy - Destroy a DMA_PMD page pool and unmap cached PMD pages
* @pool: Pool to destroy
*
- * Frees completely idle 2M pages immediately. Device DMA must be quiesced by
- * the caller first, but blocks already handed to the networking stack (e.g.,
- * RX skbs waiting in socket queues or TX skbs in TCP retransmit queues) may
- * still be in flight; those 2M pages stay on @pool->partial / @pool->full and
- * keep @pool alive on @dma_pmd_pools via @pool->refcount until their last
- * blocks free.
+ * Unmaps and frees completely idle PMD pages. If any blocks are still in flight
+ * (e.g., held by socket queues or SKBs), the pool structure and active pages
+ * remain on @dma_pmd_pools via @pool->refcount (visible to dma_pmd_domain_release())
+ * and are automatically unmapped and reclaimed as their last blocks free.
*
* Cost: Control plane teardown; flushes async reclaim workqueue.
* Locking: Process context (may sleep in flush_work). Acquires @pool->lock.
@@ -250,10 +272,12 @@ struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *pool)
struct dma_pmd_meta *meta, *tmp;
LIST_HEAD(release_list);
unsigned long flags;
+ int srcu_idx;
if (!pool)
return NULL;
+ srcu_idx = srcu_read_lock(&dma_pmd_srcu);
spin_lock_irqsave(&pool->lock, flags);
pool->destroyed = true;
@@ -267,6 +291,7 @@ struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *pool)
list_del(&meta->list);
dma_pmd_release_page(meta);
}
+ srcu_read_unlock(&dma_pmd_srcu, srcu_idx);
kref_put(&pool->refcount, dma_pmd_pool_free_kref);
flush_work(&dma_pmd_reclaim_work);
@@ -707,10 +732,12 @@ static unsigned long dma_pmd_shrink_scan(struct shrinker *shrink,
unsigned long freed = 0;
LIST_HEAD(release_list);
unsigned long flags;
+ int srcu_idx;
if (!mutex_trylock(&dma_pmd_pools_lock))
return SHRINK_STOP;
+ srcu_idx = srcu_read_lock(&dma_pmd_srcu);
list_for_each_entry(pool, &dma_pmd_pools, node) {
/*
* Every PMD page on @idle is a candidate, so this detaches one
@@ -734,11 +761,15 @@ static unsigned long dma_pmd_shrink_scan(struct shrinker *shrink,
mutex_unlock(&dma_pmd_pools_lock);
- /* Retiring must not run under @pool->lock. */
+ /*
+ * Retiring is done outside both locks: dma_pmd_release_page() unmaps
+ * from the IOMMU, which is far too long to hold pool->lock for.
+ */
list_for_each_entry_safe(meta, tmp, &release_list, list) {
list_del(&meta->list);
dma_pmd_release_page(meta);
}
+ srcu_read_unlock(&dma_pmd_srcu, srcu_idx);
/*
* Both counts are in pages, matching dma_pmd_shrink_count(). A
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index 8104897572921..8477a8edfac63 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -7,8 +7,11 @@
#include <linux/list.h>
#include <linux/llist.h>
#include <linux/spinlock.h>
+#include <linux/srcu.h>
#include <linux/types.h>
+struct iommu_domain;
+
/**
* struct dma_pmd_window - A domain's IOVA window reserved for DMA_PMD pages
* @base: First IOVA of the window. A page at @phys is mapped, in every domain
@@ -36,6 +39,15 @@ struct dma_pmd_window {
#define DMA_PMD_BLOCKS(order) (1U << (PMD_ORDER - (order)))
+/*
+ * Number of IOMMU domains that may use DMA_PMD at once.
+ *
+ * Each such domain is given a dense index, which is the bit position it uses
+ * in dma_pmd_meta.domains_mapped. One word per PMD page records every domain
+ * that has mapped this PMD page. Used to speed up map and unmap.
+ */
+#define DMA_PMD_MAX_DOMAINS BITS_PER_LONG
+
/*
* Sentinel values of dma_pmd_window.domain_idx below DMA_PMD_IDX_FIRST:
* never attempted (0, matching kzalloc), or permanently declined. A real
@@ -71,6 +83,17 @@ enum {
* meta->map_lock IRQ-safe spinlock over one PMD frame's domains_mapped
* bitmap. Innermost, and never held across anything that
* sleeps.
+ *
+ * The full order is dma_pmd_pools_lock -> pool->lock -> meta->map_lock, as
+ * taken by dma_pmd_domain_release().
+ *
+ * dma_pmd_domains_lock IRQ-safe spinlock over the domain index table and the
+ * per-domain windows. Taken on its own; it is never
+ * nested with dma_pmd_pools_lock in either direction.
+ *
+ * dma_pmd_srcu Read side lets the unmap path dereference a domain out
+ * of dma_pmd_domains[] without blocking a concurrent
+ * dma_pmd_domain_release().
*/
/*
@@ -186,6 +209,7 @@ struct dma_pmd_pool {
} ____cacheline_aligned;
extern unsigned long dma_pmd_meta_nframes;
+extern struct srcu_struct dma_pmd_srcu;
static inline struct dma_pmd_meta *dma_pmd_meta_base(void)
{
@@ -203,6 +227,10 @@ struct dma_pmd_meta *dma_pmd_meta_from_phys(phys_addr_t pa);
unsigned long dma_pmd_meta_to_pfn(const struct dma_pmd_meta *m);
phys_addr_t dma_pmd_meta_to_phys(const struct dma_pmd_meta *m);
+unsigned int dma_pmd_pools_forget_domain(int idx);
+void dma_pmd_unmap_all(struct dma_pmd_meta *meta);
+unsigned int dma_pmd_forget_domain(struct dma_pmd_meta *meta, int idx);
+
static inline bool dma_is_pmd_phys(phys_addr_t phys)
{
return dma_is_pmd_page(phys >> PAGE_SHIFT);
@@ -227,6 +255,8 @@ static inline bool dma_pmd_window_owns(const struct dma_pmd_window *win, dma_add
return size && dma - win->base < size;
}
+void dma_pmd_domain_release(struct iommu_domain *domain);
+
#else /* !CONFIG_DMA_PMD */
static inline bool dma_is_pmd_phys(phys_addr_t phys)
@@ -239,5 +269,9 @@ static inline bool dma_pmd_window_owns(const struct dma_pmd_window *win, dma_add
return false;
}
+static inline void dma_pmd_domain_release(struct iommu_domain *domain)
+{
+}
+
#endif /* CONFIG_DMA_PMD */
#endif /* _DRIVERS_IOMMU_DMA_PMD_PRIV_H */
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 08/22] iommu/dma: use per-domain IOVA window to map DMA_PMD memory
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (6 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 07/22] iommu/dma: release DMA_PMD domain mappings on domain teardown Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 09/22] iommu/dma: Add DMA_PMD arena allocator Luigi Rizzo
` (13 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Map each DMA_PMD page at win->base + phys inside the domain's reserved
IOVA window and track per-domain leaf PDE residency in
p2m->domains_mapped:
- dma_pmd_window_assign() assigns a dense domain index in
dma_pmd_domains[] and allocates the per-domain IOVA window via
dma_pmd_dma_window_alloc() on first use.
- dma_pmd_dma_window_alloc() reserves a PMD_SIZE-aligned range via
alloc_iova() from the domain's iova_domain so no ordinary IOVA can
land inside it, allowing unmap to identify pooled IOVAs by range check.
- dma_pmd_dma_map_phys() returns win->base + phys in O(1) when the
domain's bit is set in p2m->domains_mapped, installing the PDE_ORDER leaf
PTE on the first map in a domain.
- iommu_dma_map_phys() and iommu_dma_unmap_phys() route DMA_PMD
buffers through the per-domain window, making dma_unmap_*() an O(1)
range check with no page-table walk or IOTLB flush.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/Kconfig | 1 +
drivers/iommu/dma-iommu.c | 27 +++-
drivers/iommu/dma-pmd-map.c | 275 +++++++++++++++++++++++++++++++++++
drivers/iommu/dma-pmd-priv.h | 33 +++++
drivers/iommu/iommu.c | 2 +-
5 files changed, 336 insertions(+), 2 deletions(-)
diff --git a/drivers/iommu/Kconfig b/drivers/iommu/Kconfig
index 1bb347fd9da4a..7cb5b036e0f3b 100644
--- a/drivers/iommu/Kconfig
+++ b/drivers/iommu/Kconfig
@@ -164,6 +164,7 @@ config DMA_PMD
depends on IOMMU_DMA
depends on X86_64 || (ARM64 && ARM64_4K_PAGES)
depends on !PREEMPT_RT
+ select NEED_SG_DMA_FLAGS
default y
help
Hand out DMA buffers carved out of PMD_SIZE physically contiguous
diff --git a/drivers/iommu/dma-iommu.c b/drivers/iommu/dma-iommu.c
index ee14175b79bc2..6d33e245530ec 100644
--- a/drivers/iommu/dma-iommu.c
+++ b/drivers/iommu/dma-iommu.c
@@ -1286,6 +1286,14 @@ dma_addr_t iommu_dma_map_phys(struct device *dev, phys_addr_t phys, size_t size,
arch_sync_dma_flush();
}
+ if (dma_is_pmd_phys(phys)) {
+ iova = dma_pmd_dma_map_phys(dev, domain,
+ &domain->iova_cookie->win_dma_pmd,
+ phys, size, prot, dma_mask);
+ if (likely(iova != DMA_MAPPING_ERROR))
+ return iova;
+ }
+
iova = __iommu_dma_map(dev, phys, size, prot, dma_mask);
if (iova == DMA_MAPPING_ERROR &&
!(attrs & (DMA_ATTR_MMIO | DMA_ATTR_REQUIRE_COHERENT)))
@@ -1296,14 +1304,19 @@ dma_addr_t iommu_dma_map_phys(struct device *dev, phys_addr_t phys, size_t size,
void iommu_dma_unmap_phys(struct device *dev, dma_addr_t dma_handle,
size_t size, enum dma_data_direction dir, unsigned long attrs)
{
+ struct iommu_domain *domain = iommu_get_dma_domain(dev);
phys_addr_t phys;
+ /* DMA_PMD mapped buffer: nothing to unmap or release. */
+ if (dma_pmd_window_owns(&domain->iova_cookie->win_dma_pmd, dma_handle))
+ return;
+
if (attrs & (DMA_ATTR_MMIO | DMA_ATTR_REQUIRE_COHERENT)) {
__iommu_dma_unmap(dev, dma_handle, size);
return;
}
- phys = iommu_iova_to_phys(iommu_get_dma_domain(dev), dma_handle);
+ phys = iommu_iova_to_phys(domain, dma_handle);
if (WARN_ON(!phys))
return;
@@ -1499,6 +1512,18 @@ int iommu_dma_map_sg(struct device *dev, struct scatterlist *sg, int nents,
*/
break;
case PCI_P2PDMA_MAP_NONE:
+ if (dma_is_pmd_phys(sg_phys(s))) {
+ iova = dma_pmd_dma_map_phys(dev, domain,
+ &cookie->win_dma_pmd,
+ sg_phys(s), s_length,
+ prot, dma_get_mask(dev));
+ if (likely(iova != DMA_MAPPING_ERROR)) {
+ s->dma_address = iova;
+ sg_dma_len(s) = s_length;
+ sg_dma_mark_bus_address(s);
+ continue;
+ }
+ }
break;
case PCI_P2PDMA_MAP_BUS_ADDR:
/*
diff --git a/drivers/iommu/dma-pmd-map.c b/drivers/iommu/dma-pmd-map.c
index cffa94a7c7c84..565e10dd1a2e8 100644
--- a/drivers/iommu/dma-pmd-map.c
+++ b/drivers/iommu/dma-pmd-map.c
@@ -13,6 +13,8 @@
#include <linux/dma-pmd.h>
#include <linux/export.h>
#include <linux/iommu.h>
+#include <linux/iova.h>
+#include <linux/memblock.h>
#include <linux/mm.h>
#include <linux/spinlock.h>
#include <linux/srcu.h>
@@ -218,3 +220,276 @@ void dma_pmd_domain_release(struct iommu_domain *domain)
pr_debug("dma_pmd: released %u PMD page mappings for domain %p (index %d)\n",
dropped, domain, idx);
}
+
+/**
+ * dma_pmd_dma_window_alloc - Allocate an IOVA window for DMA_PMD mappings
+ * @dev: Device whose addressing limits the window must respect
+ * @domain: Domain to take the window from
+ * @size: Window size in bytes
+ *
+ * Allocates a contiguous IOVA range out of @domain's iova_domain on first use,
+ * aligned to PMD_SIZE matching the mapping. Callers on the map path still
+ * check each device's own dma_mask/bus_dma_limit against the resulting IOVA and
+ * fall back to per-buffer mapping if a narrower device shares @domain.
+ *
+ * Return: base IOVA, or 0 if no range that large is available.
+ */
+static dma_addr_t dma_pmd_dma_window_alloc(struct device *dev,
+ struct iommu_domain *domain, u64 *sizep)
+{
+ u64 limit = min_not_zero(dma_get_mask(dev), dev->bus_dma_limit);
+ struct iova_domain *iovad = dma_pmd_dma_iovad(domain);
+ u64 min_size = ALIGN(PFN_PHYS(max_pfn), PMD_SIZE);
+ unsigned long shift, iova_len;
+ struct iova *new_iova;
+
+ if (!iovad)
+ return 0;
+
+ if (domain->geometry.force_aperture)
+ limit = min_t(u64, limit, domain->geometry.aperture_end);
+
+ if (*sizep > limit / 2 && limit / 2 >= min_size)
+ *sizep = ALIGN_DOWN(limit / 2, PMD_SIZE);
+
+ /* Ignore devices with too many constraints. */
+ if (limit <= DMA_BIT_MASK(32) || *sizep > limit - PMD_SIZE)
+ return 0;
+
+ shift = iova_shift(iovad);
+ iova_len = (*sizep + PMD_SIZE) >> shift;
+
+ /*
+ * Note: dev->iommu->pci_32bit_workaround is intentionally not consulted
+ * here. It is an advisory preference (enabled by default on all PCI
+ * devices).
+ */
+ new_iova = alloc_iova(iovad, iova_len, limit >> shift, false);
+ if (!new_iova)
+ return 0;
+
+ return ALIGN((dma_addr_t)new_iova->pfn_lo << shift, PMD_SIZE);
+}
+
+/**
+ * dma_pmd_window_assign - Give @domain an IOVA window and a domain index
+ * @dev: Device whose addressing limits the window must respect
+ * @domain: Domain to set up
+ * @win: @domain's window, from its DMA cookie
+ *
+ * Called the first time anything tries to pool through @domain. Failure is
+ * recorded in @win->domain_idx rather than retried: the allocation is a
+ * multi-terabyte search of the IOVA tree, and a domain that has no room for it
+ * once will not have room on the next buffer either. The caller tests that
+ * record without the lock, so the reason has to survive in the sentinel.
+ *
+ * A domain that cannot install a PMD leaf is declined. The whole benefit
+ * is the single large PTE; without it a 2M page would cost 512 4KB PTEs, which
+ * is worse than not pooling.
+ *
+ * Locking: Acquires @dma_pmd_domains_lock (irqsave). Callable from the DMA
+ * map path, so it never sleeps.
+ *
+ * Return: 0 if @win is usable on return, negative otherwise.
+ */
+int dma_pmd_window_assign(struct device *dev, struct iommu_domain *domain,
+ struct dma_pmd_window *win)
+{
+ u64 size = ALIGN(PFN_PHYS(dma_pmd_top_pfn()), PMD_SIZE);
+ unsigned long flags;
+ int idx, ret = 0;
+ dma_addr_t base;
+
+ if (!dev_is_dma_coherent(dev))
+ return -EOPNOTSUPP;
+
+ if (min_not_zero(dma_get_mask(dev), dev->bus_dma_limit) <= DMA_BIT_MASK(32))
+ return -ENOSPC;
+
+ spin_lock_irqsave(&dma_pmd_domains_lock, flags);
+
+ if (win->size) /* another CPU got here first */
+ goto out;
+
+ /* Declined earlier; only DMA_PMD_IDX_NONE means "not tried yet". */
+ if (win->domain_idx != DMA_PMD_IDX_NONE) {
+ ret = win->domain_idx == DMA_PMD_IDX_NOHUGE ? -EOPNOTSUPP : -ENOSPC;
+ goto out;
+ }
+
+ if (!(domain->pgsize_bitmap & PMD_SIZE)) {
+ dev_warn_ratelimited(dev,
+ "dma_pmd: domain lacks a %luMB page size, not pooling\n",
+ (unsigned long)(PMD_SIZE >> 20));
+ ret = -EOPNOTSUPP;
+ goto decline_nohuge;
+ }
+
+ for (idx = 0; idx < DMA_PMD_MAX_DOMAINS; idx++) {
+ if (!dma_pmd_domains[idx].domain && !dma_pmd_domains[idx].draining)
+ break;
+ }
+ if (idx == DMA_PMD_MAX_DOMAINS) {
+ pr_warn_ratelimited("dma_pmd: all %d domain indices in use, not pooling for domain %p\n",
+ DMA_PMD_MAX_DOMAINS, domain);
+ ret = -ENOSPC;
+ goto decline;
+ }
+
+ base = dma_pmd_dma_window_alloc(dev, domain, &size);
+ if (!base) {
+ dev_warn_ratelimited(dev,
+ "dma_pmd: no free %llu GB IOVA range for a DMA_PMD window\n",
+ size >> 30);
+ ret = -ENOSPC;
+ goto decline;
+ }
+
+ dma_pmd_domains[idx].domain = domain;
+ dma_pmd_domains[idx].base = base;
+ win->base = base;
+ win->domain_idx = idx + DMA_PMD_IDX_FIRST;
+ /*
+ * Publish .base and .domain_idx before .size. A non-zero .size is what
+ * tells both the map and the unmap side that the other two are valid;
+ * see dma_pmd_window_owns().
+ */
+ smp_store_release(&win->size, size);
+
+ dev_info(dev, "dma_pmd: domain %p index %d window %pad + %llu GB\n",
+ domain, idx, &base, size >> 30);
+ goto out;
+
+decline_nohuge:
+ win->domain_idx = DMA_PMD_IDX_NOHUGE;
+ goto out;
+
+decline:
+ win->domain_idx = DMA_PMD_IDX_NOSPACE;
+out:
+ spin_unlock_irqrestore(&dma_pmd_domains_lock, flags);
+ return ret;
+}
+
+/**
+ * dma_pmd_dma_map_phys - Derive the IOVA of a pool address, mapping if needed
+ * @dev: Device performing DMA
+ * @domain: @dev's DMA domain
+ * @win: @domain's DMA_PMD window, from its DMA cookie
+ * @phys: Physical address within a DMA_PMD page
+ * @size: Mapping size requested
+ * @prot: IOMMU protection flags requested by the caller
+ * @dma_mask: Device DMA mask
+ *
+ * The IOVA is not allocated or cached: it is @win->base plus the PMD page's
+ * physical address, so every block of every PMD page has a fixed
+ * address in this domain that both this function and the unmap path can
+ * compute. All the PMD page has to remember is whether the PMD leaf PTE has
+ * actually been installed here, which is one bit in @meta->domains_mapped.
+ *
+ * The PMD page is mapped with IOMMU_CACHE | IOMMU_READ | IOMMU_WRITE so that
+ * blocks can be used in any DMA direction. Callers requesting additional
+ * protection flags in @prot fall back to per-buffer mapping.
+ *
+ * Cost:
+ * - Fast path: O(1) lockless - one bit test and an add. No IOVA tree, no
+ * page table, and no metadata load beyond the bitmask word.
+ * - First use of a PMD page in a domain: one iommu_map() of the whole PMD
+ * page, once, under @meta->map_lock.
+ * Locking: Fast path is lockless. Slow path acquires @meta->map_lock
+ * (irqsave), and @dma_pmd_domains_lock on the very first PMD page
+ * mapped through @domain.
+ * Frequency: Very high (called on every dma_map_page / dma_map_single for
+ * buffers allocated from an dma_pmd_pool).
+ *
+ * Return: Mapped IOVA, or DMA_MAPPING_ERROR on failure.
+ */
+dma_addr_t dma_pmd_dma_map_phys(struct device *dev, struct iommu_domain *domain,
+ struct dma_pmd_window *win, phys_addr_t phys,
+ size_t size, int prot, u64 dma_mask)
+{
+ int block_prot = IOMMU_CACHE | IOMMU_READ | IOMMU_WRITE;
+ unsigned long pfn = phys >> PAGE_SHIFT;
+ struct dma_pmd_meta *meta;
+ unsigned int domain_idx;
+ unsigned long flags;
+ dma_addr_t iova;
+ u64 win_size;
+ int ret;
+
+ if (unlikely((prot & ~block_prot) || !dev_is_dma_coherent(dev) ||
+ !dma_is_pmd_page(pfn) || (phys & (PMD_SIZE - 1)) + size > PMD_SIZE))
+ return DMA_MAPPING_ERROR;
+
+ meta = dma_pmd_meta_of_pfn(pfn);
+
+ /* Acquire pairs with the .size release in dma_pmd_window_assign(). */
+ win_size = smp_load_acquire(&win->size);
+
+ if (unlikely(!win_size)) {
+ u16 win_idx = READ_ONCE(win->domain_idx);
+
+ /*
+ * A declined domain stays declined, so the answer is cached in
+ * @domain_idx and read without @dma_pmd_domains_lock. Otherwise
+ * every map through such a domain would queue on a global
+ * spinlock only to be told no.
+ */
+ if (unlikely(win_idx == DMA_PMD_IDX_NOHUGE))
+ ret = -EOPNOTSUPP;
+ else if (unlikely(win_idx == DMA_PMD_IDX_NOSPACE))
+ ret = -ENOSPC;
+ else
+ ret = dma_pmd_window_assign(dev, domain, win);
+
+ if (ret)
+ return DMA_MAPPING_ERROR;
+ win_size = READ_ONCE(win->size);
+ }
+
+ iova = dma_pmd_window_iova(win, phys);
+
+ /*
+ * The window was sized against the first device to pool through this
+ * domain (and may have been clamped to the domain aperture). Mirror
+ * __iommu_dma_map() and bound by the window size, bus limit, and
+ * device mask.
+ */
+ if (unlikely(iova - win->base + size > win_size ||
+ iova + size - 1 > min_not_zero(dma_mask, dev->bus_dma_limit)))
+ return DMA_MAPPING_ERROR;
+
+ domain_idx = win->domain_idx - DMA_PMD_IDX_FIRST;
+ if (likely(test_bit(domain_idx, &meta->domains_mapped)))
+ return iova;
+
+ /* First use of this PMD page in this domain: install the leaf PTE. */
+ spin_lock_irqsave(&meta->map_lock, flags);
+ if (test_bit(domain_idx, &meta->domains_mapped)) {
+ /* raced, someone else mapped */
+ ret = 0;
+ } else {
+ ret = iommu_map(domain,
+ dma_pmd_window_iova(win, dma_pmd_meta_to_phys(meta)),
+ dma_pmd_meta_to_phys(meta), PMD_SIZE,
+ block_prot, GFP_ATOMIC);
+ if (!ret) {
+ /*
+ * A lockless reader that sees this bit skips the
+ * iommu_map() above and hands the IOVA straight to the
+ * device, so the leaf PTE has to be visible first.
+ * set_bit() carries no ordering of its own on a weakly
+ * ordered machine; the barrier is a no-op on x86,
+ * where the LOCK prefix already provides it.
+ */
+ smp_mb__before_atomic();
+ set_bit(domain_idx, &meta->domains_mapped);
+ }
+ }
+ spin_unlock_irqrestore(&meta->map_lock, flags);
+
+ if (unlikely(ret))
+ return DMA_MAPPING_ERROR;
+
+ return iova;
+}
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index 8477a8edfac63..355f865f44057 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -2,6 +2,7 @@
#ifndef _DRIVERS_IOMMU_DMA_PMD_PRIV_H
#define _DRIVERS_IOMMU_DMA_PMD_PRIV_H
+#include <linux/dma-mapping.h>
#include <linux/dma-pmd.h>
#include <linux/kref.h>
#include <linux/list.h>
@@ -10,6 +11,7 @@
#include <linux/srcu.h>
#include <linux/types.h>
+struct device;
struct iommu_domain;
/**
@@ -220,6 +222,22 @@ static inline struct dma_pmd_meta *dma_pmd_meta_base(void)
return smp_load_acquire(&dma_pmd_meta_array);
}
+/*
+ * Highest PFN covered by the allocated @dma_pmd_meta_array reservation.
+ * Every domain's IOVA window is sized to match this bound.
+ */
+static inline unsigned long dma_pmd_top_pfn(void)
+{
+ return dma_pmd_meta_nframes << PMD_ORDER;
+}
+
+/* IOVA of @phys in @win. Only valid once the PMD page's PTE is installed. */
+static inline dma_addr_t dma_pmd_window_iova(const struct dma_pmd_window *win,
+ phys_addr_t phys)
+{
+ return win->base + phys;
+}
+
int dma_pmd_meta_init(void);
bool dma_pmd_meta_ensure_pfn(unsigned long pfn, bool can_block);
struct dma_pmd_meta *dma_pmd_meta_of_pfn(unsigned long pfn);
@@ -230,6 +248,8 @@ phys_addr_t dma_pmd_meta_to_phys(const struct dma_pmd_meta *m);
unsigned int dma_pmd_pools_forget_domain(int idx);
void dma_pmd_unmap_all(struct dma_pmd_meta *meta);
unsigned int dma_pmd_forget_domain(struct dma_pmd_meta *meta, int idx);
+int dma_pmd_window_assign(struct device *dev, struct iommu_domain *domain,
+ struct dma_pmd_window *win);
static inline bool dma_is_pmd_phys(phys_addr_t phys)
{
@@ -255,6 +275,10 @@ static inline bool dma_pmd_window_owns(const struct dma_pmd_window *win, dma_add
return size && dma - win->base < size;
}
+dma_addr_t dma_pmd_dma_map_phys(struct device *dev, struct iommu_domain *domain,
+ struct dma_pmd_window *win, phys_addr_t phys,
+ size_t size, int prot, u64 dma_mask);
+
void dma_pmd_domain_release(struct iommu_domain *domain);
#else /* !CONFIG_DMA_PMD */
@@ -269,6 +293,15 @@ static inline bool dma_pmd_window_owns(const struct dma_pmd_window *win, dma_add
return false;
}
+static inline dma_addr_t dma_pmd_dma_map_phys(struct device *dev,
+ struct iommu_domain *domain,
+ struct dma_pmd_window *win,
+ phys_addr_t phys, size_t size,
+ int prot, u64 dma_mask)
+{
+ return DMA_MAPPING_ERROR;
+}
+
static inline void dma_pmd_domain_release(struct iommu_domain *domain)
{
}
diff --git a/drivers/iommu/iommu.c b/drivers/iommu/iommu.c
index cd1bca7ede9af..d918d1294eb0d 100644
--- a/drivers/iommu/iommu.c
+++ b/drivers/iommu/iommu.c
@@ -2867,7 +2867,7 @@ ssize_t iommu_map_sg(struct iommu_domain *domain, unsigned long iova,
while (i <= nents) {
phys_addr_t s_phys = sg_phys(sg);
- if (len && s_phys != start + len) {
+ if (len && (sg_dma_is_bus_address(sg) || s_phys != start + len)) {
ret = iommu_map_nosync(domain, iova + mapped, start,
len, prot, gfp);
if (ret)
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 09/22] iommu/dma: Add DMA_PMD arena allocator
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (7 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 08/22] iommu/dma: use per-domain IOVA window to map DMA_PMD memory Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 10/22] driver core: Add per-device dma_pmd_* sysfs attributes Luigi Rizzo
` (12 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Device driver use dma_alloc_attrs() for several long lived
device-accessible regions, such as queues and header buffers.
Add a compatible dma_pmd_arena allocator for DMA_PMD pages, which will be
later used by dma_alloc_attrs() for eligible allocations. Reserve a IOVA
region right after the one used to map physical ram, used specifically for
dma_pmd_arena allocations so they can be easily identifyed from their IOVA.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/Makefile | 3 +-
drivers/iommu/dma-pmd-arena.c | 385 ++++++++++++++++++++++++++++++++++
drivers/iommu/dma-pmd-kunit.c | 35 ++++
drivers/iommu/dma-pmd-map.c | 19 +-
drivers/iommu/dma-pmd-meta.c | 61 ++++--
drivers/iommu/dma-pmd-priv.h | 41 +++-
include/linux/dma-pmd.h | 9 +
7 files changed, 519 insertions(+), 34 deletions(-)
create mode 100644 drivers/iommu/dma-pmd-arena.c
diff --git a/drivers/iommu/Makefile b/drivers/iommu/Makefile
index e701b36b9b605..e96fcf309a0b2 100644
--- a/drivers/iommu/Makefile
+++ b/drivers/iommu/Makefile
@@ -11,7 +11,8 @@ obj-$(CONFIG_IOMMU_API) += iommu-traces.o
obj-$(CONFIG_IOMMU_API) += iommu-sysfs.o
obj-$(CONFIG_IOMMU_DEBUGFS) += iommu-debugfs.o
obj-$(CONFIG_IOMMU_DMA) += dma-iommu.o
-obj-$(CONFIG_DMA_PMD) += dma-pmd-meta.o dma-pmd-pool.o dma-pmd-map.o
+obj-$(CONFIG_DMA_PMD) += dma-pmd-meta.o dma-pmd-pool.o dma-pmd-map.o \
+ dma-pmd-arena.o
obj-$(CONFIG_DMA_PMD_META_KUNIT_TEST) += dma-pmd-kunit.o
obj-$(CONFIG_IOMMU_IO_PGTABLE) += io-pgtable.o
obj-$(CONFIG_IOMMU_IO_PGTABLE_ARMV7S) += io-pgtable-arm-v7s.o
diff --git a/drivers/iommu/dma-pmd-arena.c b/drivers/iommu/dma-pmd-arena.c
new file mode 100644
index 0000000000000..578df5be86a86
--- /dev/null
+++ b/drivers/iommu/dma-pmd-arena.c
@@ -0,0 +1,385 @@
+// SPDX-License-Identifier: GPL-2.0 OR BSD-3-Clause
+/*
+ * DMA_PMD coherent arena allocator (ARENA_REGION).
+ *
+ * See Documentation/core-api/dma-pmd.rst for the architecture overview.
+ */
+
+#include <linux/bitmap.h>
+#include <linux/bitops.h>
+#include <linux/cache.h>
+#include <linux/dma-map-ops.h>
+#include <linux/dma-mapping.h>
+#include <linux/dma-pmd.h>
+#include <linux/export.h>
+#include <linux/gfp.h>
+#include <linux/iommu.h>
+#include <linux/mm.h>
+#include <linux/sched/mm.h>
+#include <linux/slab.h>
+#include <linux/spinlock.h>
+#include <linux/vmalloc.h>
+
+#include "dma-iommu.h"
+#include "dma-pmd-priv.h"
+/*
+ * dma_pmd_arena (ARENA_REGION allocator for coherent buffers)
+ * -----------------------------------------------------------
+ * Packs long-lived coherent DMA allocations into PMD physical pages mapped
+ * via PMD IOMMU leaf PTEs in the 16GB ARENA_REGION at the bottom of each
+ * domain's DMA_PMD IOVA window.
+ *
+ * - Allocation & Mapping (dma_pmd_arena_alloc / dma_pmd_dma_alloc):
+ * Allocates a PAGE_SIZE-aligned, zeroed region of @size bytes in
+ * ARENA_REGION, returning the CPU virtual address and storing the IOVA in
+ * *@dma. Each PMD arena page is mapped into a single IOMMU domain and is
+ * never shared across domains. Allocations <= 1MB first try partially used
+ * PMD arena pages already mapped in the caller's domain before allocating a
+ * fresh PMD page. Allocations > 1MB allocate contiguous empty PMD slots in
+ * ARENA_REGION and vmap() them when spanning multiple PMD pages; if the
+ * trailing PMD page has leftover 4KB blocks, it becomes available for
+ * subsequent <= 1MB allocations in the same domain.
+ *
+ * - Unmapping & Release (dma_free_coherent / dma_pmd_free):
+ * dma_pmd_free() checks whether @dma_handle falls in @dev's ARENA_REGION and,
+ * if so, unreserves the 4KB blocks via dma_pmd_arena_free(). Once all 512 4KB
+ * blocks of a PMD slot are free, its PMD PTE is unmapped from its domain and
+ * the backing PMD page is returned to the buddy allocator.
+ */
+
+struct dma_pmd_arena {
+ spinlock_t lock; /* protects bitmaps and arena metadata */
+ DECLARE_BITMAP(fully_empty, DMA_PMD_ARENA_PAGES);
+ DECLARE_BITMAP(partial_empty, DMA_PMD_ARENA_PAGES);
+};
+
+/* Currently one instance for all allocations */
+static struct dma_pmd_arena dma_pmd_arena __cacheline_aligned_in_smp = {
+ .lock = __SPIN_LOCK_UNLOCKED(dma_pmd_arena.lock),
+ .fully_empty = { [0 ... BITS_TO_LONGS(DMA_PMD_ARENA_PAGES) - 1] = ~0UL },
+};
+
+/* Provide a contiguous VA range for multi-page allocations */
+static void *dma_pmd_arena_vmap(unsigned int slot, unsigned int npages)
+{
+ unsigned int nr_subpages = npages * DMA_PMD_BLOCKS(0);
+ unsigned int subpages_per_2m = DMA_PMD_BLOCKS(0);
+ struct page *page, **pages;
+ unsigned int i, j;
+ void *va;
+
+ pages = kvmalloc_array(nr_subpages, sizeof(*pages), GFP_KERNEL);
+ if (!pages)
+ return NULL;
+
+ for (i = 0; i < npages; i++) {
+ page = dma_pmd_arena_meta(slot + i)->arena_page;
+ for (j = 0; j < subpages_per_2m; j++)
+ pages[i * subpages_per_2m + j] = page + j;
+ }
+
+ va = dma_common_pages_remap(pages, nr_subpages << PAGE_SHIFT,
+ PAGE_KERNEL, __builtin_return_address(0));
+ if (!va)
+ kvfree(pages);
+ return va;
+}
+
+/* Provide a contiguous IOVA range for allocations */
+static int dma_pmd_arena_map_domain(struct device *dev, struct iommu_domain *domain,
+ struct dma_pmd_window *win, struct dma_pmd_meta *m)
+{
+ unsigned int domain_idx;
+ unsigned long flags;
+ dma_addr_t iova;
+ int prot, ret;
+ u64 limit;
+
+ if (!domain)
+ return 0;
+
+ iova = win->base + dma_pmd_meta_win_offset(m);
+ limit = min_not_zero(dma_get_mask(dev), dev->bus_dma_limit);
+ if (iova + PMD_SIZE - 1 > limit)
+ return -ENOSPC;
+
+ domain_idx = win->domain_idx - DMA_PMD_IDX_FIRST;
+ if (test_bit(domain_idx, &m->domains_mapped))
+ return 0;
+
+ prot = dma_info_to_prot(DMA_BIDIRECTIONAL, true, 0);
+ spin_lock_irqsave(&m->map_lock, flags);
+ if (test_bit(domain_idx, &m->domains_mapped)) {
+ ret = 0;
+ } else if (unlikely(!dma_pmd_domain_active(domain_idx, domain))) {
+ ret = -ENODEV;
+ } else {
+ ret = iommu_map(domain, iova, page_to_phys(m->arena_page),
+ PMD_SIZE, prot, GFP_ATOMIC);
+ if (!ret) {
+ /* Ensure the leaf PTE is visible before setting the bit. */
+ smp_mb__before_atomic();
+ set_bit(domain_idx, &m->domains_mapped);
+ }
+ }
+ spin_unlock_irqrestore(&m->map_lock, flags);
+ return ret;
+}
+
+/**
+ * dma_pmd_arena_free - Unreserve blocks in an ARENA_REGION allocation
+ * @dev: Device passed to dma_free_coherent()
+ * @size: Size of the allocation being freed
+ * @cpu_addr: CPU virtual address passed to dma_free_coherent()
+ * @dma: DMA address passed to dma_free_coherent()
+ *
+ * Return: true if @dma belonged to ARENA_REGION and was freed.
+ */
+bool dma_pmd_arena_free(struct device *dev, size_t size, void *cpu_addr, dma_addr_t dma)
+{
+ unsigned int slot, off, nblocks, bpp = DMA_PMD_BLOCKS(0);
+ struct iommu_domain *domain;
+ struct dma_pmd_window *win;
+ dma_addr_t arena_base;
+ unsigned long flags;
+
+ if (!dev || !cpu_addr)
+ return false;
+
+ if (!dma_pmd_meta_base())
+ return false;
+
+ if (dev->iommu_group) {
+ domain = iommu_get_dma_domain(dev);
+ win = domain ? dma_pmd_dma_window(domain) : NULL;
+ /* Pairs with smp_store_release() in dma_pmd_window_assign(). */
+ if (!win || !smp_load_acquire(&win->size))
+ return false;
+ arena_base = win->base;
+ } else if (IS_ENABLED(CONFIG_DMA_PMD_META_KUNIT_TEST)) {
+ /* Reachable only from KUnit tests with a dummy device. */
+ arena_base = 0;
+ } else {
+ return false;
+ }
+
+ if (dma < arena_base || dma - arena_base >= ARENA_REGION_SIZE)
+ return false;
+
+ if (is_vmalloc_addr(cpu_addr)) {
+ kvfree(dma_common_find_pages(cpu_addr));
+ dma_common_free_remap(cpu_addr, size);
+ }
+
+ slot = (dma - arena_base) >> PMD_SHIFT;
+ off = ((dma - arena_base) & (PMD_SIZE - 1)) >> PAGE_SHIFT;
+ nblocks = ALIGN(size, PAGE_SIZE) >> PAGE_SHIFT;
+
+ while (nblocks > 0 && slot < DMA_PMD_ARENA_PAGES) {
+ struct dma_pmd_meta *m = dma_pmd_arena_meta(slot);
+ unsigned int n = min(nblocks, bpp - off);
+ struct page *free_page = NULL;
+
+ spin_lock_irqsave(&dma_pmd_arena.lock, flags);
+ if (!WARN_ON_ONCE(!m->arena_page)) {
+ bitmap_clear(m->free_bitmap, off, n);
+ m->nr_free += n;
+ if (m->nr_free == bpp) {
+ free_page = m->arena_page;
+ m->arena_page = NULL;
+ m->nr_free = 0;
+ __clear_bit(slot, dma_pmd_arena.partial_empty);
+ } else {
+ __set_bit(slot, dma_pmd_arena.partial_empty);
+ }
+ }
+ spin_unlock_irqrestore(&dma_pmd_arena.lock, flags);
+
+ if (free_page) {
+ dma_pmd_unmap_all(m);
+ __free_pages(free_page, PMD_ORDER);
+ spin_lock_irqsave(&dma_pmd_arena.lock, flags);
+ __set_bit(slot, dma_pmd_arena.fully_empty);
+ spin_unlock_irqrestore(&dma_pmd_arena.lock, flags);
+ }
+
+ nblocks -= n;
+ off = 0;
+ slot++;
+ }
+
+ return true;
+}
+EXPORT_SYMBOL(dma_pmd_arena_free);
+
+/**
+ * dma_pmd_arena_alloc - Carve a DMA-mapped buffer out of ARENA_REGION
+ * @dev: Device the buffer is allocated and mapped for
+ * @size: Requested size, rounded up to 4KB
+ * @dma: Out parameter, DMA address of the returned buffer in ARENA_REGION
+ * @node: Preferred NUMA node, or NUMA_NO_NODE for the device's own node
+ *
+ * Locking: Acquires @dma_pmd_arena.lock (irqsave) when searching or updating
+ * slot bitmaps.
+ *
+ * Return: Kernel virtual address of the zeroed buffer, or NULL if the caller
+ * should fall back to its own allocation (e.g. dma_alloc_coherent()).
+ */
+void *dma_pmd_arena_alloc(struct device *dev, size_t size, dma_addr_t *dma, int node)
+{
+ unsigned int npages, nblocks, slot, i, bpp = DMA_PMD_BLOCKS(0);
+ unsigned long *bitmap = dma_pmd_arena.fully_empty;
+ struct iommu_domain *domain;
+ struct dma_pmd_window *win;
+ dma_addr_t arena_base;
+ unsigned long flags;
+ void *va;
+
+ if (!dev || (!dev->iommu_group && !IS_ENABLED(CONFIG_DMA_PMD_META_KUNIT_TEST)))
+ return NULL;
+
+ if (node == NUMA_NO_NODE)
+ node = dev_to_node(dev);
+
+ size = ALIGN(size, PAGE_SIZE);
+ if (!size || size > ARENA_REGION_SIZE)
+ return NULL;
+
+ domain = dev->iommu_group ? iommu_get_dma_domain(dev) : NULL;
+ win = domain ? dma_pmd_dma_window(domain) : NULL;
+ if (dev->iommu_group && (!win || !(domain->pgsize_bitmap & PMD_SIZE)))
+ goto err_fallback;
+
+ if (unlikely(dma_pmd_meta_init()))
+ goto err_fallback;
+
+ /* Pairs with smp_store_release() in dma_pmd_window_assign(). */
+ if (win && !smp_load_acquire(&win->size) &&
+ (READ_ONCE(win->domain_idx) != DMA_PMD_IDX_NONE ||
+ dma_pmd_window_assign(dev, domain, win)))
+ goto err_fallback;
+
+ arena_base = win ? win->base : 0;
+ nblocks = size >> PAGE_SHIFT;
+
+ /*
+ * Allocations <= 1MB first try partially empty pages already mapped in
+ * @domain (arena PMD pages are single-domain and never shared across
+ * IOMMU domains).
+ */
+ if (size <= SZ_1M) {
+ unsigned long domain_mask = win ? BIT(win->domain_idx - DMA_PMD_IDX_FIRST) : 0;
+ u64 limit = min_not_zero(dma_get_mask(dev), dev->bus_dma_limit);
+
+ spin_lock_irqsave(&dma_pmd_arena.lock, flags);
+ for_each_set_bit(slot, dma_pmd_arena.partial_empty, DMA_PMD_ARENA_PAGES) {
+ struct dma_pmd_meta *m = dma_pmd_arena_meta(slot);
+ unsigned long off;
+
+ if (READ_ONCE(m->domains_mapped) != domain_mask ||
+ (domain &&
+ arena_base + ((dma_addr_t)(slot + 1) << PMD_SHIFT) - 1 > limit))
+ continue;
+ if (m->nr_free < nblocks)
+ continue;
+
+ off = bitmap_find_next_zero_area(m->free_bitmap, bpp, 0, nblocks,
+ (1UL << get_order(size)) - 1);
+ if (off >= bpp)
+ continue;
+
+ bitmap_set(m->free_bitmap, off, nblocks);
+ m->nr_free -= nblocks;
+ if (m->nr_free == 0)
+ __clear_bit(slot, dma_pmd_arena.partial_empty);
+ va = page_address(m->arena_page) + (off << PAGE_SHIFT);
+ *dma = arena_base + ((dma_addr_t)slot << PMD_SHIFT) +
+ (off << PAGE_SHIFT);
+ spin_unlock_irqrestore(&dma_pmd_arena.lock, flags);
+
+ memset(va, 0, size);
+ return va;
+ }
+ spin_unlock_irqrestore(&dma_pmd_arena.lock, flags);
+ }
+
+ /*
+ * Allocations > 1MB (or <= 1MB when no partial page fits) look for
+ * @npages contiguous fully empty pages.
+ */
+ npages = DIV_ROUND_UP(size, PMD_SIZE);
+ spin_lock_irqsave(&dma_pmd_arena.lock, flags);
+ slot = 0;
+ while (slot + npages <= DMA_PMD_ARENA_PAGES) {
+ unsigned int next_zero;
+
+ slot = find_next_bit(bitmap, DMA_PMD_ARENA_PAGES, slot);
+ if (slot + npages > DMA_PMD_ARENA_PAGES)
+ break;
+ next_zero = find_next_zero_bit(bitmap, slot + npages, slot);
+ if (next_zero >= slot + npages) {
+ bitmap_clear(bitmap, slot, npages);
+ break;
+ }
+ slot = next_zero + 1;
+ }
+ spin_unlock_irqrestore(&dma_pmd_arena.lock, flags);
+
+ if (slot + npages > DMA_PMD_ARENA_PAGES)
+ goto err_fallback;
+
+ for (i = 0; i < npages; i++) {
+ struct dma_pmd_meta *m = dma_pmd_arena_meta(slot + i);
+ gfp_t gfp = GFP_KERNEL | __GFP_ZERO | __GFP_NOWARN;
+ struct page *page;
+
+ page = (node == NUMA_NO_NODE) ? alloc_pages(gfp, PMD_ORDER) :
+ alloc_pages_node(node, gfp, PMD_ORDER);
+ if (!page)
+ goto err_free_pages;
+
+ m->arena_page = page;
+ if (dma_pmd_arena_map_domain(dev, domain, win, m)) {
+ m->arena_page = NULL;
+ __free_pages(page, PMD_ORDER);
+ goto err_free_pages;
+ }
+ }
+
+ va = (npages == 1) ? page_address(dma_pmd_arena_meta(slot)->arena_page) :
+ dma_pmd_arena_vmap(slot, npages);
+ if (!va)
+ goto err_free_pages;
+
+ spin_lock_irqsave(&dma_pmd_arena.lock, flags);
+ for (i = 0; i < npages; i++) {
+ struct dma_pmd_meta *m = dma_pmd_arena_meta(slot + i);
+ unsigned int n = min(nblocks - i * bpp, bpp);
+
+ bitmap_zero(m->free_bitmap, bpp);
+ bitmap_set(m->free_bitmap, 0, n);
+ m->nr_free = bpp - n;
+ if (m->nr_free > 0)
+ __set_bit(slot + i, dma_pmd_arena.partial_empty);
+ }
+ spin_unlock_irqrestore(&dma_pmd_arena.lock, flags);
+
+ *dma = arena_base + ((dma_addr_t)slot << PMD_SHIFT);
+ return va;
+
+err_free_pages:
+ while (i--) {
+ struct dma_pmd_meta *m = dma_pmd_arena_meta(slot + i);
+ struct page *page = m->arena_page;
+
+ m->arena_page = NULL;
+ dma_pmd_unmap_all(m);
+ __free_pages(page, PMD_ORDER);
+ }
+ spin_lock_irqsave(&dma_pmd_arena.lock, flags);
+ bitmap_set(bitmap, slot, npages);
+ spin_unlock_irqrestore(&dma_pmd_arena.lock, flags);
+err_fallback:
+ return NULL;
+}
+EXPORT_SYMBOL(dma_pmd_arena_alloc);
diff --git a/drivers/iommu/dma-pmd-kunit.c b/drivers/iommu/dma-pmd-kunit.c
index 7f4251845ba49..b91dbcb468879 100644
--- a/drivers/iommu/dma-pmd-kunit.c
+++ b/drivers/iommu/dma-pmd-kunit.c
@@ -166,12 +166,47 @@ static void test_window_helpers(struct kunit *test)
dma_pmd_pool_destroy(pool);
}
+static void test_arena_alloc_and_free(struct kunit *test)
+{
+ struct device dev = { .numa_node = NUMA_NO_NODE };
+ void *v1, *v2, *v3, *vl1;
+ dma_addr_t d1, d2, d3, dl1;
+
+ KUNIT_EXPECT_NULL(test,
+ dma_pmd_arena_alloc(NULL, SZ_4K, &d1, NUMA_NO_NODE));
+
+ v1 = dma_pmd_arena_alloc(&dev, SZ_64K, &d1, NUMA_NO_NODE);
+ KUNIT_ASSERT_NOT_NULL(test, v1);
+ v2 = dma_pmd_arena_alloc(&dev, SZ_64K, &d2, NUMA_NO_NODE);
+ KUNIT_ASSERT_NOT_NULL(test, v2);
+ KUNIT_EXPECT_PTR_EQ(test, v2, v1 + SZ_64K);
+ KUNIT_EXPECT_EQ(test, d2, d1 + SZ_64K);
+
+ /* Freeing v1 unreserves [0, 64K); next 64K alloc reuses it. */
+ KUNIT_EXPECT_TRUE(test, dma_pmd_arena_free(&dev, SZ_64K, v1, d1));
+ v3 = dma_pmd_arena_alloc(&dev, SZ_64K, &d3, NUMA_NO_NODE);
+ KUNIT_EXPECT_PTR_EQ(test, v3, v1);
+ KUNIT_EXPECT_EQ(test, d3, d1);
+
+ /*
+ * Allocate 3MB (spans 2 contiguous fully_empty slots in ARENA_REGION);
+ * the trailing 1MB in the second slot is marked partial_empty and can
+ * satisfy a <= 1MB allocation before both are freed.
+ */
+ vl1 = dma_pmd_arena_alloc(&dev, SZ_2M + SZ_1M, &dl1, NUMA_NO_NODE);
+ KUNIT_ASSERT_NOT_NULL(test, vl1);
+ KUNIT_EXPECT_TRUE(test, dma_pmd_arena_free(&dev, SZ_2M + SZ_1M, vl1, dl1));
+ KUNIT_EXPECT_TRUE(test, dma_pmd_arena_free(&dev, SZ_64K, v2, d2));
+ KUNIT_EXPECT_TRUE(test, dma_pmd_arena_free(&dev, SZ_64K, v3, d3));
+}
+
static struct kunit_case dma_pmd_meta_test_cases[] = {
KUNIT_CASE(test_meta_init_and_roundtrip),
KUNIT_CASE(test_meta_invalid_phys),
KUNIT_CASE(test_meta_pooled_toggle),
KUNIT_CASE(test_pool_alloc_and_recycle),
KUNIT_CASE(test_window_helpers),
+ KUNIT_CASE(test_arena_alloc_and_free),
{}
};
diff --git a/drivers/iommu/dma-pmd-map.c b/drivers/iommu/dma-pmd-map.c
index 565e10dd1a2e8..1147874ac8068 100644
--- a/drivers/iommu/dma-pmd-map.c
+++ b/drivers/iommu/dma-pmd-map.c
@@ -51,6 +51,11 @@ static DEFINE_SPINLOCK(dma_pmd_domains_lock);
*/
DEFINE_SRCU(dma_pmd_srcu);
+bool dma_pmd_domain_active(unsigned int idx, const struct iommu_domain *domain)
+{
+ return READ_ONCE(dma_pmd_domains[idx].domain) == domain;
+}
+
/**
* dma_pmd_unmap_all - Remove a PMD page's PTE from every domain holding it
* @meta: PMD page metadata structure
@@ -77,7 +82,7 @@ DEFINE_SRCU(dma_pmd_srcu);
*/
void dma_pmd_unmap_all(struct dma_pmd_meta *meta)
{
- phys_addr_t phys = dma_pmd_meta_to_phys(meta);
+ dma_addr_t offset = dma_pmd_meta_win_offset(meta);
unsigned long mapped, flags;
int idx, srcu_idx;
@@ -104,7 +109,7 @@ void dma_pmd_unmap_all(struct dma_pmd_meta *meta)
if (!domain)
continue;
- iommu_unmap(domain, base + phys, PMD_SIZE);
+ iommu_unmap(domain, base + offset, PMD_SIZE);
}
srcu_read_unlock(&dma_pmd_srcu, srcu_idx);
}
@@ -170,6 +175,7 @@ void dma_pmd_domain_release(struct iommu_domain *domain)
{
unsigned int dropped = 0;
unsigned long flags;
+ unsigned int i;
int idx;
/* Nothing has ever been pooled, so nothing can reference @domain. */
@@ -199,6 +205,9 @@ void dma_pmd_domain_release(struct iommu_domain *domain)
dropped += dma_pmd_pools_forget_domain(idx);
+ for (i = 0; i < DMA_PMD_ARENA_PAGES; i++)
+ dropped += dma_pmd_forget_domain(dma_pmd_arena_meta(i), idx);
+
/*
* Any PMD page removed from a pool list before the walk above is either
* already on @dma_pmd_free_list (placed there under @pool->lock in
@@ -225,7 +234,7 @@ void dma_pmd_domain_release(struct iommu_domain *domain)
* dma_pmd_dma_window_alloc - Allocate an IOVA window for DMA_PMD mappings
* @dev: Device whose addressing limits the window must respect
* @domain: Domain to take the window from
- * @size: Window size in bytes
+ * @sizep: In/out window size in bytes
*
* Allocates a contiguous IOVA range out of @domain's iova_domain on first use,
* aligned to PMD_SIZE matching the mapping. Callers on the map path still
@@ -237,9 +246,9 @@ void dma_pmd_domain_release(struct iommu_domain *domain)
static dma_addr_t dma_pmd_dma_window_alloc(struct device *dev,
struct iommu_domain *domain, u64 *sizep)
{
+ u64 min_size = ARENA_REGION_SIZE + ALIGN(PFN_PHYS(max_pfn), PMD_SIZE);
u64 limit = min_not_zero(dma_get_mask(dev), dev->bus_dma_limit);
struct iova_domain *iovad = dma_pmd_dma_iovad(domain);
- u64 min_size = ALIGN(PFN_PHYS(max_pfn), PMD_SIZE);
unsigned long shift, iova_len;
struct iova *new_iova;
@@ -295,7 +304,7 @@ static dma_addr_t dma_pmd_dma_window_alloc(struct device *dev,
int dma_pmd_window_assign(struct device *dev, struct iommu_domain *domain,
struct dma_pmd_window *win)
{
- u64 size = ALIGN(PFN_PHYS(dma_pmd_top_pfn()), PMD_SIZE);
+ u64 size = ALIGN(PFN_PHYS(dma_pmd_top_pfn()), PMD_SIZE) + ARENA_REGION_SIZE;
unsigned long flags;
int idx, ret = 0;
dma_addr_t base;
diff --git a/drivers/iommu/dma-pmd-meta.c b/drivers/iommu/dma-pmd-meta.c
index 27b8c2568c749..8f2bc7d77e5bd 100644
--- a/drivers/iommu/dma-pmd-meta.c
+++ b/drivers/iommu/dma-pmd-meta.c
@@ -186,8 +186,9 @@ static inline bool dma_pmd_pfn_online(unsigned long pfn)
*
* Return: 0, or -ENOMEM with the range partially backed.
*/
-static int dma_pmd_meta_populate(struct dma_pmd_meta *array, unsigned long nframes,
- unsigned long start_pfn, unsigned long end_pfn, bool force)
+static int dma_pmd_meta_populate(struct dma_pmd_meta *array, unsigned long *bitmap,
+ unsigned long nframes, unsigned long start_pfn,
+ unsigned long end_pfn, bool force)
{
unsigned long chunk, last, pfn;
@@ -208,7 +209,8 @@ static int dma_pmd_meta_populate(struct dma_pmd_meta *array, unsigned long nfram
addr = (unsigned long)array + (chunk << PAGE_SHIFT);
if (vmalloc_to_page((void *)addr)) {
- set_bit(chunk, dma_pmd_chunk_bitmap);
+ if (bitmap)
+ set_bit(chunk, bitmap);
continue; /* already backed */
}
@@ -247,14 +249,17 @@ static int dma_pmd_meta_populate(struct dma_pmd_meta *array, unsigned long nfram
* never leak the page it declined to consume.
*/
__free_page(page);
- set_bit(chunk, dma_pmd_chunk_bitmap);
+ if (bitmap)
+ set_bit(chunk, bitmap);
continue;
}
flush_cache_vmap(addr, addr + PAGE_SIZE);
- /* Pair with test_bit_acquire() in readers. */
- smp_mb__before_atomic();
- set_bit(chunk, dma_pmd_chunk_bitmap);
+ if (bitmap) {
+ /* Pair with test_bit_acquire() in readers. */
+ smp_mb__before_atomic();
+ set_bit(chunk, bitmap);
+ }
dma_pmd_meta_pages++;
}
@@ -273,8 +278,9 @@ bool dma_pmd_meta_ensure_pfn(unsigned long pfn, bool can_block)
return false;
mutex_lock(&dma_pmd_meta_mutex);
- ret = dma_pmd_meta_populate(dma_pmd_meta_base(), dma_pmd_meta_nframes,
- pfn, pfn + (1UL << PMD_ORDER), true);
+ ret = dma_pmd_meta_populate(dma_pmd_meta_base(), dma_pmd_chunk_bitmap,
+ dma_pmd_meta_nframes, pfn,
+ pfn + (1UL << PMD_ORDER), true);
mutex_unlock(&dma_pmd_meta_mutex);
return !ret;
@@ -284,18 +290,20 @@ EXPORT_SYMBOL(dma_pmd_meta_ensure_pfn);
/**
* dma_pmd_meta_init - Reserve and populate the sparse per-PMD metadata array
*
- * Populates all present RAM pages before publishing @dma_pmd_meta_array so
- * no concurrent reader ever sees an unmapped entry.
+ * Populates all present RAM pages and ARENA_REGION entries before publishing
+ * @dma_pmd_meta_array so no concurrent reader ever sees an unmapped entry.
*
* Return: 0 on success, or negative errno on failure.
*/
int dma_pmd_meta_init(void)
{
- unsigned long nframes, nchunks, size;
- struct dma_pmd_meta *array;
+ unsigned long nframes, total_frames, nchunks, size, i;
+ struct dma_pmd_meta *raw, *array;
struct vm_struct *vm;
int ret = 0;
+ static_assert(IS_ALIGNED(DMA_PMD_ARENA_PAGES << DMA_PMD_META_SHIFT, PAGE_SIZE));
+
if (likely(dma_pmd_meta_base()))
return 0;
@@ -308,8 +316,9 @@ int dma_pmd_meta_init(void)
(unsigned long)min_t(u64, iomem_resource.end >> PAGE_SHIFT,
1ULL << (MAX_PHYSMEM_BITS - PAGE_SHIFT))),
1UL << PMD_ORDER);
- size = PAGE_ALIGN(nframes << DMA_PMD_META_SHIFT);
- nchunks = size >> PAGE_SHIFT;
+ total_frames = DMA_PMD_ARENA_PAGES + nframes;
+ size = PAGE_ALIGN(total_frames << DMA_PMD_META_SHIFT);
+ nchunks = PAGE_ALIGN(nframes << DMA_PMD_META_SHIFT) >> PAGE_SHIFT;
dma_pmd_chunk_bitmap = bitmap_zalloc(nchunks, GFP_KERNEL);
if (!dma_pmd_chunk_bitmap) {
@@ -324,14 +333,25 @@ int dma_pmd_meta_init(void)
ret = -ENOMEM;
goto out_unlock;
}
- array = vm->addr;
-
- ret = dma_pmd_meta_populate(array, nframes, 0, max_pfn, false);
+ raw = vm->addr;
+ array = raw + DMA_PMD_ARENA_PAGES;
+
+ ret = dma_pmd_meta_populate(raw, NULL, DMA_PMD_ARENA_PAGES,
+ 0, DMA_PMD_ARENA_PAGES << PMD_ORDER, true);
+ if (!ret)
+ ret = dma_pmd_meta_populate(array, dma_pmd_chunk_bitmap,
+ nframes, 0, max_pfn, false);
if (ret) {
+ unsigned long off, chunk;
struct page *p, *next;
- unsigned long chunk;
LIST_HEAD(pages);
+ for (off = 0; off < (DMA_PMD_ARENA_PAGES << DMA_PMD_META_SHIFT);
+ off += PAGE_SIZE) {
+ p = vmalloc_to_page((void *)raw + off);
+ if (p)
+ list_add(&p->lru, &pages);
+ }
for_each_set_bit(chunk, dma_pmd_chunk_bitmap, nchunks) {
p = vmalloc_to_page((void *)array + (chunk << PAGE_SHIFT));
if (p)
@@ -348,6 +368,9 @@ int dma_pmd_meta_init(void)
goto out_unlock;
}
+ for (i = 0; i < DMA_PMD_ARENA_PAGES; i++)
+ spin_lock_init(&raw[i].map_lock);
+
/*
* Publish @nframes and the populated pages before the array pointer:
* readers pair with smp_load_acquire(&dma_pmd_meta_array).
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index 355f865f44057..b411ecb1fdafc 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -16,8 +16,9 @@ struct iommu_domain;
/**
* struct dma_pmd_window - A domain's IOVA window reserved for DMA_PMD pages
- * @base: First IOVA of the window. A page at @phys is mapped, in every domain
- * that has a window, at @base + @phys.
+ * @base: First IOVA of the window. The leading ARENA_REGION_SIZE bytes
+ * [base, base + ARENA_REGION_SIZE) form ARENA_REGION; a pooled RAM
+ * page at @phys is mapped at @base + ARENA_REGION_SIZE + @phys.
* @size: Window size in bytes, or 0 if this domain has no window, either
* because nothing has pooled through it yet, or because it could not
* find a free range that large. 0 must make both the map and the unmap
@@ -40,6 +41,8 @@ struct dma_pmd_window {
#ifdef CONFIG_DMA_PMD
#define DMA_PMD_BLOCKS(order) (1U << (PMD_ORDER - (order)))
+#define DMA_PMD_ARENA_PAGES 8192U
+#define ARENA_REGION_SIZE ((u64)DMA_PMD_ARENA_PAGES * PMD_SIZE)
/*
* Number of IOMMU domains that may use DMA_PMD at once.
@@ -116,8 +119,9 @@ enum {
* @domains_mapped: One bit per domain index: set once this PMD page's leaf
* PTE is installed in that domain.
*
- * Alloc/free fast path, all under pool->lock:
+ * Alloc/free fast path, all under pool->lock (or dma_pmd_arena.lock):
* @pool: Owning dma_pmd_pool (holds a kref on the pool while pooled)
+ * @arena_page: Allocated 2MB struct page for an ARENA_REGION entry
* @list: Node in pool->partial, pool->idle or pool->full
*
* Slow path only (disjoint lifetimes):
@@ -125,11 +129,13 @@ enum {
* @rcu: RCU head used to defer buddy release past lockless readers
*
* @free_bitmap: Bitmap of zeroed available block indices (up to 1 << PMD_ORDER)
- * @dirty_bitmap: Bitmap of dirty available block indices
+ * for pools, or allocated 4KB blocks for ARENA_REGION entries
+ * @dirty_bitmap: Bitmap of dirty available block indices for pools
*
* Lives in the sparse per-PMD-frame array @dma_pmd_meta_array indexed by
- * (pfn >> PMD_ORDER). Only chunks covering valid RAM are backed by physical
- * pages (128 KB per GB of RAM).
+ * (pfn >> PMD_ORDER), preceded by DMA_PMD_ARENA_PAGES entries for
+ * ARENA_REGION. Only chunks covering valid RAM and ARENA_REGION are backed by
+ * physical pages (128 KB per GB of RAM).
*/
struct dma_pmd_meta {
bool pooled;
@@ -139,7 +145,10 @@ struct dma_pmd_meta {
unsigned long domains_mapped;
- struct dma_pmd_pool *pool;
+ union {
+ struct dma_pmd_pool *pool;
+ struct page *arena_page;
+ };
struct list_head list;
union {
@@ -224,18 +233,29 @@ static inline struct dma_pmd_meta *dma_pmd_meta_base(void)
/*
* Highest PFN covered by the allocated @dma_pmd_meta_array reservation.
- * Every domain's IOVA window is sized to match this bound.
+ * Every domain's IOVA window is sized up to this bound.
*/
static inline unsigned long dma_pmd_top_pfn(void)
{
return dma_pmd_meta_nframes << PMD_ORDER;
}
+static inline struct dma_pmd_meta *dma_pmd_arena_meta(unsigned int slot)
+{
+ return (dma_pmd_meta_base() - DMA_PMD_ARENA_PAGES) + slot;
+}
+
/* IOVA of @phys in @win. Only valid once the PMD page's PTE is installed. */
static inline dma_addr_t dma_pmd_window_iova(const struct dma_pmd_window *win,
phys_addr_t phys)
{
- return win->base + phys;
+ return win->base + ARENA_REGION_SIZE + phys;
+}
+
+/* Offset of @m (either arena or RAM) from the start of a domain's window. */
+static inline dma_addr_t dma_pmd_meta_win_offset(const struct dma_pmd_meta *m)
+{
+ return (dma_addr_t)(m - dma_pmd_arena_meta(0)) << PMD_SHIFT;
}
int dma_pmd_meta_init(void);
@@ -250,6 +270,7 @@ void dma_pmd_unmap_all(struct dma_pmd_meta *meta);
unsigned int dma_pmd_forget_domain(struct dma_pmd_meta *meta, int idx);
int dma_pmd_window_assign(struct device *dev, struct iommu_domain *domain,
struct dma_pmd_window *win);
+bool dma_pmd_domain_active(unsigned int idx, const struct iommu_domain *domain);
static inline bool dma_is_pmd_phys(phys_addr_t phys)
{
@@ -281,6 +302,8 @@ dma_addr_t dma_pmd_dma_map_phys(struct device *dev, struct iommu_domain *domain,
void dma_pmd_domain_release(struct iommu_domain *domain);
+void *dma_pmd_arena_alloc(struct device *dev, size_t size, dma_addr_t *dma, int node);
+
#else /* !CONFIG_DMA_PMD */
static inline bool dma_is_pmd_phys(phys_addr_t phys)
diff --git a/include/linux/dma-pmd.h b/include/linux/dma-pmd.h
index 085d9a730c278..55c770fe232b0 100644
--- a/include/linux/dma-pmd.h
+++ b/include/linux/dma-pmd.h
@@ -95,6 +95,9 @@ static inline struct page *dma_pmd_pool_alloc(struct dma_pmd_pool *pool, gfp_t g
}
bool dma_pmd_pool_has_free(struct dma_pmd_pool *pool);
+/* Hooks for kernel/dma/mapping.c */
+bool dma_pmd_arena_free(struct device *dev, size_t size, void *cpu_addr, dma_addr_t dma);
+
#else /* !CONFIG_DMA_PMD */
static inline bool dma_is_pmd_page(unsigned long pfn)
@@ -141,5 +144,11 @@ static inline bool dma_pmd_pool_has_free(struct dma_pmd_pool *pool)
return false;
}
+static inline bool dma_pmd_arena_free(struct device *dev, size_t size,
+ void *cpu_addr, dma_addr_t dma)
+{
+ return false;
+}
+
#endif /* CONFIG_DMA_PMD */
#endif /* _LINUX_DMA_PMD_H */
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 10/22] driver core: Add per-device dma_pmd_* sysfs attributes
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (8 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 09/22] iommu/dma: Add DMA_PMD arena allocator Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 11/22] dma-mapping: Use DMA_PMD arena for dma_alloc_attrs() Luigi Rizzo
` (11 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Add four boolean attributes to struct device exposed under
/sys/devices/*/*/dma_pmd_*:
- dma_pmd_rxbuf: enable DMA_PMD pools for RX buffers / page_pool
- dma_pmd_tx_hdrs: enable DMA_PMD pools for TX header bounce buffers
- dma_pmd_rings: enable DMA_PMD arena for coherent rings
- dma_pmd_debug: log coherent DMA alloc timing and arena usage
These can be used by device drivers to selectively enable DMA_PMD for
their allocations.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/base/core.c | 51 ++++++++++++++++++++++++++++++++++++++++++
include/linux/device.h | 5 +++++
2 files changed, 56 insertions(+)
diff --git a/drivers/base/core.c b/drivers/base/core.c
index 4c0c373998a19..f8c3c42d84b8d 100644
--- a/drivers/base/core.c
+++ b/drivers/base/core.c
@@ -2905,6 +2905,43 @@ static ssize_t removable_show(struct device *dev, const struct device_attribute
}
static const DEVICE_ATTR_RO(removable);
+/* Per-device configuration to selectively enable DMA_PMD */
+#define DMA_PMD_BOOL_ATTR(name) \
+static ssize_t name##_show(struct device *dev, \
+ struct device_attribute *attr, char *buf) \
+{ \
+ return sysfs_emit(buf, "%u\n", READ_ONCE(dev->name)); \
+} \
+static ssize_t name##_store(struct device *dev, \
+ struct device_attribute *attr, \
+ const char *buf, size_t count) \
+{ \
+ bool val; \
+ int ret = kstrtobool(buf, &val); \
+ if (ret < 0) \
+ return ret; \
+ WRITE_ONCE(dev->name, val); \
+ return count; \
+} \
+static DEVICE_ATTR_RW(name)
+
+DMA_PMD_BOOL_ATTR(dma_pmd_rxbuf);
+DMA_PMD_BOOL_ATTR(dma_pmd_tx_hdrs);
+DMA_PMD_BOOL_ATTR(dma_pmd_rings);
+DMA_PMD_BOOL_ATTR(dma_pmd_debug);
+
+static struct attribute *dev_attr_dma_pmd[] = {
+ &dev_attr_dma_pmd_rxbuf.attr,
+ &dev_attr_dma_pmd_tx_hdrs.attr,
+ &dev_attr_dma_pmd_rings.attr,
+ &dev_attr_dma_pmd_debug.attr,
+ NULL,
+};
+
+static const struct attribute_group dev_attr_dma_pmd_group = {
+ .attrs = dev_attr_dma_pmd,
+};
+
int device_add_groups(struct device *dev,
const struct attribute_group *const *groups)
{
@@ -3012,8 +3049,20 @@ static int device_add_attrs(struct device *dev)
goto err_remove_dev_removable;
}
+ if (dev->dma_mask) {
+ error = device_add_group(dev, &dev_attr_dma_pmd_group);
+ if (error)
+ goto err_remove_dev_physical_location;
+ }
+
return 0;
+ err_remove_dev_physical_location:
+ if (dev->physical_location) {
+ device_remove_group(dev, &dev_attr_physical_location_group);
+ kfree(dev->physical_location);
+ dev->physical_location = NULL;
+ }
err_remove_dev_removable:
device_remove_file(dev, &dev_attr_removable);
err_remove_dev_waiting_for_supplier:
@@ -3037,6 +3086,8 @@ static void device_remove_attrs(struct device *dev)
const struct class *class = dev->class;
const struct device_type *type = dev->type;
+ device_remove_group(dev, &dev_attr_dma_pmd_group);
+
if (dev->physical_location) {
device_remove_group(dev, &dev_attr_physical_location_group);
kfree(dev->physical_location);
diff --git a/include/linux/device.h b/include/linux/device.h
index aee79fd6b32b4..52aa4a7e1cc45 100644
--- a/include/linux/device.h
+++ b/include/linux/device.h
@@ -794,6 +794,11 @@ struct device {
enum device_removable removable;
+ bool dma_pmd_rxbuf;
+ bool dma_pmd_tx_hdrs;
+ bool dma_pmd_rings;
+ bool dma_pmd_debug;
+
DECLARE_BITMAP(flags, DEV_FLAG_COUNT);
};
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 11/22] dma-mapping: Use DMA_PMD arena for dma_alloc_attrs()
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (9 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 10/22] driver core: Add per-device dma_pmd_* sysfs attributes Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 12/22] net/core: Use per-CPU DMA_PMD pools for skb_page_frag_refill() Luigi Rizzo
` (10 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Wire dma_alloc_attrs() and dma_free_attrs() into the DMA_PMD arena
when dev->dma_pmd_rings is enabled on a device:
- Add dma_pmd_dma_alloc() and dma_is_pmd_dma() backed by
dma_pmd_arena.
- Route GFP_KERNEL dma_alloc_attrs() requests through dma_pmd_dma_alloc()
when dev->dma_pmd_rings is set, logging latency when
dev->dma_pmd_debug is enabled.
- Reclaim sub-allocations back to the arena in dma_free_attrs() via
dma_pmd_arena_free() while keeping the PMD_SIZE IOMMU mapping active.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/dma-pmd-arena.c | 27 ++++++++++++++++++++++++
drivers/iommu/dma-pmd-kunit.c | 11 +++++++---
drivers/iommu/dma-pmd-map.c | 17 +++++++++++++++
include/linux/dma-pmd.h | 29 ++++++++++++++++++++++++++
kernel/dma/mapping.c | 39 +++++++++++++++++++++++++++++------
5 files changed, 114 insertions(+), 9 deletions(-)
diff --git a/drivers/iommu/dma-pmd-arena.c b/drivers/iommu/dma-pmd-arena.c
index 578df5be86a86..e1a7d9d6f0f60 100644
--- a/drivers/iommu/dma-pmd-arena.c
+++ b/drivers/iommu/dma-pmd-arena.c
@@ -383,3 +383,30 @@ void *dma_pmd_arena_alloc(struct device *dev, size_t size, dma_addr_t *dma, int
return NULL;
}
EXPORT_SYMBOL(dma_pmd_arena_alloc);
+
+void *dma_pmd_dma_alloc(struct device *dev, size_t size, dma_addr_t *dma,
+ gfp_t gfp, unsigned long attrs)
+{
+ void *va;
+
+ if (!dev || !dev->iommu_group || !dev_is_dma_coherent(dev))
+ return NULL;
+ /*
+ * The arena allocates PMD pages, IOMMU page tables, and vmap() areas
+ * with GFP_KERNEL, and shares PMD pages across allocations on the same
+ * device. Decline non-GFP_KERNEL requests (including GFP_NOFS/GFP_NOIO
+ * or scoped memalloc_nofs/noio contexts) and per-allocation modifiers
+ * such as __GFP_ACCOUNT or __GFP_NORETRY so the caller falls back to
+ * the standard DMA allocator.
+ */
+ if ((current_gfp_context(gfp) & ~(__GFP_ZERO | __GFP_NOWARN)) != GFP_KERNEL ||
+ (attrs & ~DMA_ATTR_NO_WARN))
+ return NULL;
+
+ va = dma_pmd_arena_alloc(dev, size, dma, dev_to_node(dev));
+ if (!va && !(gfp & __GFP_NOWARN) && !(attrs & DMA_ATTR_NO_WARN))
+ dev_warn_ratelimited(dev, "DMA_PMD arena alloc failed (size %zu)\n",
+ size);
+ return va;
+}
+EXPORT_SYMBOL(dma_pmd_dma_alloc);
diff --git a/drivers/iommu/dma-pmd-kunit.c b/drivers/iommu/dma-pmd-kunit.c
index b91dbcb468879..c2363277e2ecd 100644
--- a/drivers/iommu/dma-pmd-kunit.c
+++ b/drivers/iommu/dma-pmd-kunit.c
@@ -169,11 +169,16 @@ static void test_window_helpers(struct kunit *test)
static void test_arena_alloc_and_free(struct kunit *test)
{
struct device dev = { .numa_node = NUMA_NO_NODE };
- void *v1, *v2, *v3, *vl1;
dma_addr_t d1, d2, d3, dl1;
+ void *v1, *v2, *v3, *vl1;
- KUNIT_EXPECT_NULL(test,
- dma_pmd_arena_alloc(NULL, SZ_4K, &d1, NUMA_NO_NODE));
+ KUNIT_EXPECT_NULL(test, dma_pmd_arena_alloc(NULL, SZ_4K, &d1, NUMA_NO_NODE));
+ KUNIT_EXPECT_NULL(test, dma_pmd_dma_alloc(NULL, SZ_4K, &d1, GFP_KERNEL, 0));
+ KUNIT_EXPECT_NULL(test, dma_pmd_dma_alloc(&dev, SZ_4K, &d1, GFP_KERNEL, 0));
+ KUNIT_EXPECT_NULL(test, dma_pmd_dma_alloc(&dev, SZ_4K, &d1,
+ GFP_KERNEL | __GFP_ACCOUNT, 0));
+ KUNIT_EXPECT_FALSE(test, dma_pmd_free(&dev, SZ_4K, NULL, 0));
+ KUNIT_EXPECT_FALSE(test, dma_is_pmd_dma(&dev, 0));
v1 = dma_pmd_arena_alloc(&dev, SZ_64K, &d1, NUMA_NO_NODE);
KUNIT_ASSERT_NOT_NULL(test, v1);
diff --git a/drivers/iommu/dma-pmd-map.c b/drivers/iommu/dma-pmd-map.c
index 1147874ac8068..76cb5d01cb1e8 100644
--- a/drivers/iommu/dma-pmd-map.c
+++ b/drivers/iommu/dma-pmd-map.c
@@ -380,6 +380,23 @@ int dma_pmd_window_assign(struct device *dev, struct iommu_domain *domain,
return ret;
}
+/**
+ * dma_is_pmd_dma - Test whether @dma lies in @dev's DMA_PMD IOVA window
+ * @dev: Device performing DMA
+ * @dma: DMA address to test
+ */
+bool dma_is_pmd_dma(struct device *dev, dma_addr_t dma)
+{
+ struct dma_pmd_window *win;
+
+ if (!dev || !dev->iommu_group)
+ return false;
+
+ win = dma_pmd_dma_window(iommu_get_dma_domain(dev));
+ return win && dma_pmd_window_owns(win, dma);
+}
+EXPORT_SYMBOL(dma_is_pmd_dma);
+
/**
* dma_pmd_dma_map_phys - Derive the IOVA of a pool address, mapping if needed
* @dev: Device performing DMA
diff --git a/include/linux/dma-pmd.h b/include/linux/dma-pmd.h
index 55c770fe232b0..7b7e9477d44e1 100644
--- a/include/linux/dma-pmd.h
+++ b/include/linux/dma-pmd.h
@@ -96,7 +96,19 @@ static inline struct page *dma_pmd_pool_alloc(struct dma_pmd_pool *pool, gfp_t g
bool dma_pmd_pool_has_free(struct dma_pmd_pool *pool);
/* Hooks for kernel/dma/mapping.c */
+void *dma_pmd_dma_alloc(struct device *dev, size_t size, dma_addr_t *dma,
+ gfp_t gfp, unsigned long attrs);
bool dma_pmd_arena_free(struct device *dev, size_t size, void *cpu_addr, dma_addr_t dma);
+bool dma_is_pmd_dma(struct device *dev, dma_addr_t dma);
+
+static inline bool dma_pmd_free(struct device *dev, size_t size,
+ void *cpu_addr, dma_addr_t dma_handle)
+{
+ /* Pairs with smp_store_release() in dma_pmd_meta_init(). */
+ if (likely(!smp_load_acquire(&dma_pmd_meta_array)))
+ return false;
+ return dma_pmd_arena_free(dev, size, cpu_addr, dma_handle);
+}
#else /* !CONFIG_DMA_PMD */
@@ -144,11 +156,28 @@ static inline bool dma_pmd_pool_has_free(struct dma_pmd_pool *pool)
return false;
}
+static inline void *dma_pmd_dma_alloc(struct device *dev, size_t size, dma_addr_t *dma,
+ gfp_t gfp, unsigned long attrs)
+{
+ return NULL;
+}
+
static inline bool dma_pmd_arena_free(struct device *dev, size_t size,
void *cpu_addr, dma_addr_t dma)
{
return false;
}
+static inline bool dma_is_pmd_dma(struct device *dev, dma_addr_t dma)
+{
+ return false;
+}
+
+static inline bool dma_pmd_free(struct device *dev, size_t size,
+ void *cpu_addr, dma_addr_t dma_handle)
+{
+ return false;
+}
+
#endif /* CONFIG_DMA_PMD */
#endif /* _LINUX_DMA_PMD_H */
diff --git a/kernel/dma/mapping.c b/kernel/dma/mapping.c
index bf2651a70b7c2..cc18896017e35 100644
--- a/kernel/dma/mapping.c
+++ b/kernel/dma/mapping.c
@@ -10,6 +10,8 @@
#include <linux/dma-map-ops.h>
#include <linux/export.h>
#include <linux/gfp.h>
+#include <linux/ktime.h>
+#include <linux/dma-pmd.h>
#include <linux/iommu-dma.h>
#include <linux/kmsan.h>
#include <linux/of_device.h>
@@ -180,7 +182,8 @@ dma_addr_t dma_map_phys(struct device *dev, phys_addr_t phys, size_t size,
if (!is_mmio)
kmsan_handle_dma(phys, size, dir);
trace_dma_map_phys(dev, phys, addr, size, dir, attrs);
- debug_dma_map_phys(dev, phys, size, dir, addr, attrs);
+ if (IS_ENABLED(CONFIG_DMA_API_DEBUG) && !dma_is_pmd_dma(dev, addr))
+ debug_dma_map_phys(dev, phys, size, dir, addr, attrs);
return addr;
}
@@ -223,7 +226,8 @@ void dma_unmap_phys(struct device *dev, dma_addr_t addr, size_t size,
else if (ops->unmap_phys)
ops->unmap_phys(dev, addr, size, dir, attrs);
trace_dma_unmap_phys(dev, addr, size, dir, attrs);
- debug_dma_unmap_phys(dev, addr, size, dir, attrs);
+ if (IS_ENABLED(CONFIG_DMA_API_DEBUG) && !dma_is_pmd_dma(dev, addr))
+ debug_dma_unmap_phys(dev, addr, size, dir, attrs);
}
EXPORT_SYMBOL_GPL(dma_unmap_phys);
@@ -633,7 +637,11 @@ EXPORT_SYMBOL_GPL(dma_get_required_mask);
void *dma_alloc_attrs(struct device *dev, size_t size, dma_addr_t *dma_handle,
gfp_t flag, unsigned long attrs)
{
+ bool debug = unlikely(READ_ONCE(dev->dma_pmd_debug)) && gfpflags_allow_blocking(flag);
const struct dma_map_ops *ops = get_dma_ops(dev);
+ bool can_block = gfpflags_allow_blocking(flag);
+ bool from_arena = false;
+ u64 start_ns = 0;
void *cpu_addr;
WARN_ON_ONCE(!dev->coherent_dma_mask);
@@ -655,10 +663,14 @@ void *dma_alloc_attrs(struct device *dev, size_t size, dma_addr_t *dma_handle,
if (force_dma_unencrypted(dev))
attrs |= __DMA_ATTR_ALLOC_CC_SHARED;
+ if (debug) {
+ start_ns = ktime_get_ns();
+ *dma_handle = 0;
+ }
if (dma_alloc_from_dev_coherent(dev, size, dma_handle, &cpu_addr)) {
trace_dma_alloc(dev, cpu_addr, *dma_handle, size,
DMA_BIDIRECTIONAL, flag, attrs);
- return cpu_addr;
+ goto out_debug;
}
/* let the implementation decide on the zone to allocate from: */
@@ -667,18 +679,31 @@ void *dma_alloc_attrs(struct device *dev, size_t size, dma_addr_t *dma_handle,
if (dma_alloc_direct(dev, ops) || arch_dma_alloc_direct(dev)) {
cpu_addr = dma_direct_alloc(dev, size, dma_handle, flag, attrs);
} else if (use_dma_iommu(dev)) {
- cpu_addr = iommu_dma_alloc(dev, size, dma_handle, flag, attrs);
+ if (READ_ONCE(dev->dma_pmd_rings) && can_block) {
+ cpu_addr = dma_pmd_dma_alloc(dev, size, dma_handle, flag, attrs);
+ from_arena = !!cpu_addr;
+ }
+ if (!from_arena)
+ cpu_addr = iommu_dma_alloc(dev, size, dma_handle, flag, attrs);
} else if (ops->alloc) {
cpu_addr = ops->alloc(dev, size, dma_handle, flag, attrs);
} else {
+ cpu_addr = NULL;
trace_dma_alloc(dev, NULL, 0, size, DMA_BIDIRECTIONAL, flag,
attrs);
- return NULL;
+ goto out_debug;
}
trace_dma_alloc(dev, cpu_addr, *dma_handle, size, DMA_BIDIRECTIONAL,
flag, attrs);
- debug_dma_alloc_coherent(dev, size, *dma_handle, cpu_addr, attrs);
+ if (!from_arena)
+ debug_dma_alloc_coherent(dev, size, *dma_handle, cpu_addr, attrs);
+out_debug:
+ if (debug)
+ dev_info(dev,
+ "dma_alloc_coherent(size=%zu, gfp=%pGg, attrs=%#lx) -> va=%p, dma=%pad, arena=%d in %llu ns\n",
+ size, &flag, attrs, cpu_addr, dma_handle, from_arena,
+ ktime_get_ns() - start_ns);
return cpu_addr;
}
EXPORT_SYMBOL(dma_alloc_attrs);
@@ -703,6 +728,8 @@ void dma_free_attrs(struct device *dev, size_t size, void *cpu_addr,
attrs);
if (!cpu_addr)
return;
+ if (use_dma_iommu(dev) && dma_pmd_free(dev, size, cpu_addr, dma_handle))
+ return;
debug_dma_free_coherent(dev, size, cpu_addr, dma_handle, attrs);
if (dma_alloc_direct(dev, ops) || arch_dma_free_direct(dev, dma_handle))
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 12/22] net/core: Use per-CPU DMA_PMD pools for skb_page_frag_refill()
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (10 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 11/22] dma-mapping: Use DMA_PMD arena for dma_alloc_attrs() Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 13/22] net/core: Use DMA_PMD for page_pool memory Luigi Rizzo
` (9 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Add sysctl net.core.tx_enable_dma_pmd to back skb_page_frag_refill() with
per-CPU order-3 (32KB block) and order-0 (4KB block) DMA_PMD page pools.
When enabled, TX socket buffer page fragment allocations draw fragments
from DMA_PMD physical pages registered with per-CPU dma_pmd_pool. The
first allocation maps the entire DMA_PMD page, subsequent allocations
reuse the cached PMD_SIZE IOVA mapping locklessly without per-packet unmap
or IOTLB flush.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
include/net/sock.h | 3 ++
net/core/sock.c | 86 +++++++++++++++++++++++++++++++++++++-
net/core/sysctl_net_core.c | 7 ++++
3 files changed, 94 insertions(+), 2 deletions(-)
diff --git a/include/net/sock.h b/include/net/sock.h
index 60ea55dc18854..a06b53e9ff225 100644
--- a/include/net/sock.h
+++ b/include/net/sock.h
@@ -3092,6 +3092,9 @@ extern __u32 sysctl_rmem_default;
#define SKB_FRAG_PAGE_ORDER get_order(32768)
DECLARE_STATIC_KEY_FALSE(net_high_order_alloc_disable_key);
+DECLARE_STATIC_KEY_FALSE(net_tx_enable_dma_pmd_key);
+int net_tx_dma_pmd_sysctl(const struct ctl_table *table, int write,
+ void *buffer, size_t *lenp, loff_t *ppos);
static inline int sk_get_wmem0(const struct sock *sk, const struct proto *proto)
{
diff --git a/net/core/sock.c b/net/core/sock.c
index e8551df8330ff..1fd338a547ce1 100644
--- a/net/core/sock.c
+++ b/net/core/sock.c
@@ -115,6 +115,7 @@
#include <linux/memcontrol.h>
#include <linux/prefetch.h>
#include <linux/compat.h>
+#include <linux/dma-pmd.h>
#include <linux/mroute.h>
#include <linux/mroute6.h>
#include <linux/icmpv6.h>
@@ -3173,6 +3174,87 @@ static void sk_leave_memory_pressure(struct sock *sk)
}
DEFINE_STATIC_KEY_FALSE(net_high_order_alloc_disable_key);
+DEFINE_STATIC_KEY_FALSE(net_tx_enable_dma_pmd_key);
+
+static DEFINE_PER_CPU(struct dma_pmd_pool *, tx_pmd_pool_high);
+static DEFINE_PER_CPU(struct dma_pmd_pool *, tx_pmd_pool_order0);
+
+/*
+ * Lazily initialize the per-CPU DMA_PMD page pools on the first write to sysctl
+ * net.core.tx_enable_dma_pmd.
+ *
+ * Once initialized, the struct dma_pmd_pool descriptors remain allocated for
+ * the lifetime of the kernel so that lockless raw_cpu_read() in
+ * __alloc_pmd() is always safe against concurrent sysctl
+ * toggles. If pool creation fails partway through, all pools created so far are
+ * destroyed and per-CPU pointers are reset to NULL.
+ */
+static int net_tx_dma_pmd_init(void)
+{
+ static DEFINE_MUTEX(mutex);
+ struct dma_pmd_pool *pool;
+ static bool initialized;
+ int cpu;
+
+ if (!IS_ENABLED(CONFIG_DMA_PMD))
+ return -EOPNOTSUPP;
+
+ mutex_lock(&mutex);
+ if (initialized) {
+ mutex_unlock(&mutex);
+ return 0;
+ }
+
+ for_each_possible_cpu(cpu) {
+ if (SKB_FRAG_PAGE_ORDER) {
+ pool = dma_pmd_pool_create(SKB_FRAG_PAGE_ORDER, 16);
+ if (!pool)
+ goto err_cleanup;
+ per_cpu(tx_pmd_pool_high, cpu) = pool;
+ }
+
+ pool = dma_pmd_pool_create(0, 16);
+ if (!pool)
+ goto err_cleanup;
+ per_cpu(tx_pmd_pool_order0, cpu) = pool;
+ }
+
+ initialized = true;
+ mutex_unlock(&mutex);
+ return 0;
+
+err_cleanup:
+ for_each_possible_cpu(cpu) {
+ per_cpu(tx_pmd_pool_high, cpu) =
+ dma_pmd_pool_destroy(per_cpu(tx_pmd_pool_high, cpu));
+ per_cpu(tx_pmd_pool_order0, cpu) =
+ dma_pmd_pool_destroy(per_cpu(tx_pmd_pool_order0, cpu));
+ }
+ mutex_unlock(&mutex);
+ return -ENOMEM;
+}
+
+int net_tx_dma_pmd_sysctl(const struct ctl_table *table, int write,
+ void *buffer, size_t *lenp, loff_t *ppos)
+{
+ if (write) {
+ int ret = net_tx_dma_pmd_init();
+
+ if (ret)
+ return ret;
+ }
+
+ return proc_do_static_key(table, write, buffer, lenp, ppos);
+}
+
+static struct page *__alloc_pmd(gfp_t gfp, unsigned int order)
+{
+ if (!static_branch_unlikely(&net_tx_enable_dma_pmd_key))
+ return alloc_pages(gfp, order);
+ if (order == SKB_FRAG_PAGE_ORDER)
+ return dma_pmd_pool_alloc(raw_cpu_read(tx_pmd_pool_high), gfp);
+ return dma_pmd_pool_alloc(raw_cpu_read(tx_pmd_pool_order0), gfp) ?: alloc_page(gfp);
+}
/**
* skb_page_frag_refill - check that a page_frag contains enough room
@@ -3200,7 +3282,7 @@ bool skb_page_frag_refill(unsigned int sz, struct page_frag *pfrag, gfp_t gfp)
if (SKB_FRAG_PAGE_ORDER &&
!static_branch_unlikely(&net_high_order_alloc_disable_key)) {
/* Avoid direct reclaim but allow kswapd to wake */
- pfrag->page = alloc_pages((gfp & ~__GFP_DIRECT_RECLAIM) |
+ pfrag->page = __alloc_pmd((gfp & ~__GFP_DIRECT_RECLAIM) |
__GFP_COMP | __GFP_NOWARN |
__GFP_NORETRY,
SKB_FRAG_PAGE_ORDER);
@@ -3209,7 +3291,7 @@ bool skb_page_frag_refill(unsigned int sz, struct page_frag *pfrag, gfp_t gfp)
return true;
}
}
- pfrag->page = alloc_page(gfp);
+ pfrag->page = __alloc_pmd(gfp, 0);
if (likely(pfrag->page)) {
pfrag->size = PAGE_SIZE;
return true;
diff --git a/net/core/sysctl_net_core.c b/net/core/sysctl_net_core.c
index eb35da3556f4a..d9a7d9cb349e0 100644
--- a/net/core/sysctl_net_core.c
+++ b/net/core/sysctl_net_core.c
@@ -651,6 +651,13 @@ static struct ctl_table net_core_table[] = {
.mode = 0644,
.proc_handler = proc_do_static_key,
},
+ {
+ .procname = "tx_enable_dma_pmd",
+ .data = &net_tx_enable_dma_pmd_key.key,
+ .maxlen = sizeof(net_tx_enable_dma_pmd_key),
+ .mode = 0644,
+ .proc_handler = net_tx_dma_pmd_sysctl,
+ },
{
.procname = "gro_normal_batch",
.data = &net_hotdata.gro_normal_batch,
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 13/22] net/core: Use DMA_PMD for page_pool memory
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (11 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 12/22] net/core: Use per-CPU DMA_PMD pools for skb_page_frag_refill() Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 14/22] iommu/dma: Support decrypted and pinned DMA_PMD pages Luigi Rizzo
` (8 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
When dev->dma_pmd_rxbuf is enabled on the pool device,
use a per-page_pool DMA_PMD pool for page_pool allocations
(__page_pool_alloc_netmems_slow() and __page_pool_alloc_page_order()).
When enabled, pages leaving the page_pool cache (e.g. held by socket
buffers beyond NAPI) return to the DMA_PMD pool on final put_page(),
retaining their cached IOVA mapping, avoiding per-page 4KB IOMMU
PTE updates and IOTLB flushes on replenishment. Plain 4KB pages are
opportunistically replaced in __page_pool_page_can_be_recycled() when
free pooled blocks are available.
Also skip inserting and erasing DMA_PMD pages in the pool->dma_mapped
XArray (used for DMA_ATTR_SKIP_CPU_SYNC tracking), as dma_unmap_page_attrs()
is a no-op for DMA_PMD pages and skipping the XArray avoids xa_lock
contention on refill and release.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
include/net/page_pool/types.h | 3 ++
net/core/page_pool.c | 61 ++++++++++++++++++++++++++++-------
2 files changed, 52 insertions(+), 12 deletions(-)
diff --git a/include/net/page_pool/types.h b/include/net/page_pool/types.h
index 03da138722f58..62a5c50dda43f 100644
--- a/include/net/page_pool/types.h
+++ b/include/net/page_pool/types.h
@@ -34,6 +34,8 @@
#define PP_FLAG_ALL (PP_FLAG_DMA_MAP | PP_FLAG_DMA_SYNC_DEV | \
PP_FLAG_SYSTEM_POOL | PP_FLAG_ALLOW_UNREADABLE_NETMEM)
+struct dma_pmd_pool;
+
/* Index limit to stay within PP_DMA_INDEX_BITS for DMA indices */
#define PP_DMA_INDEX_LIMIT XA_LIMIT(1, BIT(PP_DMA_INDEX_BITS) - 1)
@@ -234,6 +236,7 @@ struct page_pool {
void *mp_priv;
const struct memory_provider_ops *mp_ops;
+ struct dma_pmd_pool *dma_pmd_pool;
struct xarray dma_mapped;
diff --git a/net/core/page_pool.c b/net/core/page_pool.c
index d7c88c0b67e53..5774247a06291 100644
--- a/net/core/page_pool.c
+++ b/net/core/page_pool.c
@@ -22,6 +22,7 @@
#include <linux/page-flags.h>
#include <linux/mm.h> /* for put_page() */
#include <linux/poison.h>
+#include <linux/dma-pmd.h>
#include <linux/ethtool.h>
#include <linux/netdevice.h>
@@ -304,6 +305,10 @@ static int page_pool_init(struct page_pool *pool,
goto free_ptr_ring;
}
+ if (pool->dma_map && pool->p.dev && !pool->mp_ops &&
+ READ_ONCE(pool->p.dev->dma_pmd_rxbuf))
+ pool->dma_pmd_pool = dma_pmd_pool_create(pool->p.order, 8);
+
return 0;
free_ptr_ring:
@@ -318,6 +323,7 @@ static int page_pool_init(struct page_pool *pool,
static void page_pool_uninit(struct page_pool *pool)
{
+ pool->dma_pmd_pool = dma_pmd_pool_destroy(pool->dma_pmd_pool);
ptr_ring_cleanup(&pool->ring, NULL);
xa_destroy(&pool->dma_mapped);
@@ -406,7 +412,9 @@ static noinline netmem_ref page_pool_refill_alloc_cache(struct page_pool *pool)
if (unlikely(!netmem))
break;
- if (likely(netmem_is_pref_nid(netmem, pref_nid))) {
+ if (likely(netmem_is_pref_nid(netmem, pref_nid)) ||
+ (pool->dma_pmd_pool &&
+ dma_is_pmd_page(page_to_pfn(netmem_to_page(netmem))))) {
pool->alloc.cache[pool->alloc.count++] = netmem;
} else {
/* NUMA mismatch;
@@ -565,9 +573,20 @@ static bool page_pool_dma_map(struct page_pool *pool, netmem_ref netmem, gfp_t g
goto unmap_failed;
}
- err = page_pool_register_dma_index(pool, netmem, gfp);
- if (err)
- goto unset_failed;
+ /*
+ * DMA_PMD mappings are owned by dma_pmd_pool (and torn down on pool or
+ * IOMMU domain release without pool->p.dev), not unmapped per-page in
+ * page_pool_scrub(). Leave dma_index == 0 so page_pool_release_dma_index()
+ * skips dma_unmap_page_attrs() without dereferencing pool->p.dev on
+ * late returns after device removal.
+ */
+ if (pool->dma_pmd_pool && dma_is_pmd_dma(pool->p.dev, dma)) {
+ netmem_set_dma_index(netmem, 0);
+ } else {
+ err = page_pool_register_dma_index(pool, netmem, gfp);
+ if (err)
+ goto unset_failed;
+ }
page_pool_dma_sync_for_device(pool, netmem, pool->p.max_len);
@@ -585,10 +604,13 @@ static bool page_pool_dma_map(struct page_pool *pool, netmem_ref netmem, gfp_t g
static struct page *__page_pool_alloc_page_order(struct page_pool *pool,
gfp_t gfp)
{
- struct page *page;
+ struct page *page = NULL;
gfp |= __GFP_COMP;
- page = alloc_pages_node(pool->p.nid, gfp, pool->p.order);
+ if (pool->dma_pmd_pool)
+ page = dma_pmd_pool_alloc_node(pool->dma_pmd_pool, gfp, pool->p.nid);
+ if (!page)
+ page = alloc_pages_node(pool->p.nid, gfp, pool->p.order);
if (unlikely(!page))
return NULL;
@@ -634,8 +656,14 @@ static noinline netmem_ref __page_pool_alloc_netmems_slow(struct page_pool *pool
/* Mark empty alloc.cache slots "empty" for alloc_pages_bulk */
memset(&pool->alloc.cache, 0, sizeof(void *) * bulk);
- nr_pages = alloc_pages_bulk_node(gfp, pool->p.nid, bulk,
- (struct page **)pool->alloc.cache);
+ nr_pages = 0;
+ if (pool->dma_pmd_pool)
+ nr_pages = dma_pmd_pool_alloc_bulk_node(pool->dma_pmd_pool, gfp,
+ pool->p.nid, bulk,
+ (struct page **)pool->alloc.cache);
+ if (!nr_pages)
+ nr_pages = alloc_pages_bulk_node(gfp, pool->p.nid, bulk,
+ (struct page **)pool->alloc.cache);
if (unlikely(!nr_pages))
return 0;
@@ -825,9 +853,18 @@ static bool page_pool_recycle_in_cache(netmem_ref netmem,
static bool __page_pool_page_can_be_recycled(netmem_ref netmem)
{
- return netmem_is_net_iov(netmem) ||
- (page_ref_count(netmem_to_page(netmem)) == 1 &&
- !page_is_pfmemalloc(netmem_to_page(netmem)));
+ struct page *page;
+
+ if (netmem_is_net_iov(netmem))
+ return true;
+
+ page = netmem_to_page(netmem);
+ if (page_ref_count(page) != 1 || page_is_pfmemalloc(page))
+ return false;
+ if (page->pp->dma_pmd_pool && !dma_is_pmd_page(page_to_pfn(page)) &&
+ dma_pmd_pool_has_free(page->pp->dma_pmd_pool))
+ return false;
+ return true;
}
/* If the page refcnt == 1, this will try to recycle the page.
@@ -1178,7 +1215,7 @@ static void page_pool_scrub(struct page_pool *pool)
* if there are no outstanding mapped pages.
*/
if (dma_dev_need_sync(pool->p.dev) &&
- !xa_empty(&pool->dma_mapped))
+ (pool->dma_pmd_pool || !xa_empty(&pool->dma_mapped)))
synchronize_net();
}
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 14/22] iommu/dma: Support decrypted and pinned DMA_PMD pages
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (12 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 13/22] net/core: Use DMA_PMD for page_pool memory Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 15/22] iommu/dma: Add background page scrubber for DMA_PMD pools Luigi Rizzo
` (7 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
In Confidential Computing (CoCo) guests (AMD SEV-SNP, Intel TDX, ARM CCA),
guest memory is private/encrypted by default and DMA buffers must be
converted to shared/decrypted in the host page tables before a device
can access them. Similarly, in preemptible VMs, IO-accessible memory may
need to be pinned in the host on allocation and unpinned on release.
Doing these conversions on a per-packet 4K basis is prohibitive. With
DMA_PMD pools, a 2MB page is allocated once from the buddy allocator,
handed out in sub-blocks, and only returned when completely idle.
Add two module parameters (dma_pmd.decrypt and dma_pmd.pin, with
dma_pmd.decrypt auto-enabled when CC_ATTR_GUEST_MEM_ENCRYPT is active)
and per-PMD-page state flags in meta->flags (DMA_PMD_DECRYPTED,
DMA_PMD_PINNED):
- On 2MB page allocation (dma_pmd_add_page), call dma_pmd_page_prepare()
before splitting the compound page to invoke set_memory_decrypted()
and/or the host pin hook. Skip conversion if the caller cannot sleep
(!gfpflags_allow_blocking), since set_memory_decrypted() acquires
cpa_lock and allocates PTE pages.
- On 2MB page release (dma_pmd_release_page / dma_pmd_release_page_rcu),
call dma_pmd_page_unprepare() before RCU free to restore
set_memory_encrypted() and/or unpin before returning the 2MB page to
the buddy allocator. If restoration fails, leak the 2MB page rather
than returning a page with inconsistent encryption/pin state to buddy.
- In dma_pmd_shrink_scan(), defer decrypted/pinned pages to the async
reclaim workqueue (dma_pmd_free_list) instead of calling
dma_pmd_release_page() inline under fs_reclaim.
- Add dma_is_pmd_direct(phys) and use it in dma_direct_map_phys() to
bypass swiotlb_map() bounce buffering for decrypted/pinned DMA_PMD
pages, using phys_to_dma_direct() so force_dma_unencrypted() devices
receive unencrypted DMA addresses.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/dma-pmd-kunit.c | 2 +-
drivers/iommu/dma-pmd-meta.c | 15 +++-
drivers/iommu/dma-pmd-pool.c | 147 ++++++++++++++++++++++++++++------
drivers/iommu/dma-pmd-priv.h | 14 +++-
include/linux/dma-pmd.h | 15 ++++
kernel/dma/direct.c | 6 +-
6 files changed, 168 insertions(+), 31 deletions(-)
diff --git a/drivers/iommu/dma-pmd-kunit.c b/drivers/iommu/dma-pmd-kunit.c
index c2363277e2ecd..7fcc2bcb315bc 100644
--- a/drivers/iommu/dma-pmd-kunit.c
+++ b/drivers/iommu/dma-pmd-kunit.c
@@ -53,7 +53,7 @@ static void test_meta_invalid_phys(struct kunit *test)
static void dma_pmd_meta_set_pooled(struct dma_pmd_meta *m, bool pooled)
{
- WRITE_ONCE(m->pooled, pooled);
+ WRITE_ONCE(m->flags, pooled ? DMA_PMD_POOLED : 0);
}
static void test_meta_pooled_toggle(struct kunit *test)
diff --git a/drivers/iommu/dma-pmd-meta.c b/drivers/iommu/dma-pmd-meta.c
index 8f2bc7d77e5bd..744964d2c7974 100644
--- a/drivers/iommu/dma-pmd-meta.c
+++ b/drivers/iommu/dma-pmd-meta.c
@@ -8,6 +8,7 @@
#include <linux/bitmap.h>
#include <linux/cache.h>
#include <linux/cacheflush.h>
+#include <linux/cc_platform.h>
#include <linux/dma-pmd.h>
#include <linux/export.h>
#include <linux/gfp.h>
@@ -64,10 +65,22 @@ static __always_inline struct dma_pmd_meta *__dma_pmd_meta_of_pfn(unsigned long
bool __dma_is_pmd_page(unsigned long pfn)
{
- return READ_ONCE(__dma_pmd_meta_of_pfn(pfn)->pooled);
+ return READ_ONCE(__dma_pmd_meta_of_pfn(pfn)->flags) & DMA_PMD_POOLED;
}
EXPORT_SYMBOL(__dma_is_pmd_page);
+bool __dma_is_pmd_direct(unsigned long pfn)
+{
+ u8 flags = READ_ONCE(__dma_pmd_meta_of_pfn(pfn)->flags);
+
+ if (cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT) &&
+ !(flags & DMA_PMD_DECRYPTED))
+ return false;
+ return (flags & DMA_PMD_POOLED) &&
+ (flags & (DMA_PMD_DECRYPTED | DMA_PMD_PINNED));
+}
+EXPORT_SYMBOL(__dma_is_pmd_direct);
+
/**
* dma_pmd_meta_of_pfn - Metadata for a PFN known to be used by DMA_PMD.
* @pfn: PFN the caller already holds a struct page for
diff --git a/drivers/iommu/dma-pmd-pool.c b/drivers/iommu/dma-pmd-pool.c
index 1d4244ba422a3..deacb6f7d3610 100644
--- a/drivers/iommu/dma-pmd-pool.c
+++ b/drivers/iommu/dma-pmd-pool.c
@@ -9,6 +9,7 @@
#include <linux/bitmap.h>
#include <linux/bitops.h>
#include <linux/cache.h>
+#include <linux/cc_platform.h>
#include <linux/debugfs.h>
#include <linux/dma-mapping.h>
#include <linux/dma-pmd.h>
@@ -22,6 +23,7 @@
#include <linux/mutex.h>
#include <linux/rcupdate.h>
#include <linux/seq_file.h>
+#include <linux/set_memory.h>
#include <linux/shrinker.h>
#include <linux/slab.h>
#include <linux/spinlock.h>
@@ -39,6 +41,14 @@ static LLIST_HEAD(dma_pmd_free_list);
/* Dedicated WQ_MEM_RECLAIM workqueue, can run without allocations. */
static struct workqueue_struct *dma_pmd_wq __ro_after_init;
+/* Mark pooled PMD pages decrypted for DMA (auto-enabled on CoCo guests). */
+static bool dma_pmd_decrypt __read_mostly;
+module_param_named(decrypt, dma_pmd_decrypt, bool, 0644);
+
+/* Pin pooled PMD pages in the host hypervisor on allocation and unpin on release. */
+static bool dma_pmd_pin __read_mostly;
+module_param_named(pin, dma_pmd_pin, bool, 0644);
+
/*
* Running total and global ceiling on DMA_PMD pages pinned across all pools.
* If zero, it is set at boot at 1/8 of total RAM.
@@ -65,6 +75,70 @@ static void dma_pmd_pool_free_kref(struct kref *kref)
dma_pmd_schedule_reclaim();
}
+static int dma_pmd_page_prepare(struct page *page, struct dma_pmd_meta *meta,
+ bool can_block)
+{
+ bool want_decrypt = READ_ONCE(dma_pmd_decrypt);
+ bool want_pin = READ_ONCE(dma_pmd_pin);
+ void *vaddr = page_address(page);
+ int ret;
+
+ if (!want_decrypt && !want_pin)
+ return 0;
+
+ /*
+ * Decryption and host-pinning hypercalls may sleep. In atomic context,
+ * leave DMA_PMD_DECRYPTED/PINNED unset so the page is still pooled but
+ * uses normal swiotlb bounce buffering when required.
+ */
+ if (!can_block) {
+ pr_warn_ratelimited("%s: cannot block but want %s %s\n", __func__,
+ want_decrypt ? "decrypt" : "", want_pin ? "pin" : "");
+ return 0;
+ }
+
+ /*
+ * Explicit page decryption is for CoCo guests (where force_dma_unencrypted()
+ * is true), not bare-metal SME hosts where 64-bit devices and the IOMMU
+ * set the encryption bit in DMA addresses and IOMMU PTEs.
+ */
+ if (want_decrypt && !WARN_ON_ONCE(cc_platform_has(CC_ATTR_HOST_MEM_ENCRYPT))) {
+ ret = set_memory_decrypted((unsigned long)vaddr, 1 << PMD_ORDER);
+ if (ret)
+ return ret;
+ /* Initialize cleartext cachelines after flipping encryption state. */
+ memset(vaddr, 0, PMD_SIZE);
+ meta->flags |= DMA_PMD_DECRYPTED;
+ }
+
+ if (want_pin) {
+ /* Hook point for hypervisor GPA range pinning. */
+ meta->flags |= DMA_PMD_PINNED;
+ }
+
+ return 0;
+}
+
+static int dma_pmd_page_unprepare(struct dma_pmd_meta *meta)
+{
+ struct page *page = pfn_to_page(dma_pmd_meta_to_pfn(meta));
+
+ if (meta->flags & DMA_PMD_PINNED) {
+ /* Hook point for hypervisor GPA range unpinning. */
+ WRITE_ONCE(meta->flags, meta->flags & ~DMA_PMD_PINNED);
+ }
+
+ if (meta->flags & DMA_PMD_DECRYPTED) {
+ if (set_memory_encrypted((unsigned long)page_address(page),
+ 1 << PMD_ORDER)) {
+ pr_warn_ratelimited("dma_pmd: leaking PMD page that could not be re-encrypted\n");
+ return -EIO;
+ }
+ WRITE_ONCE(meta->flags, meta->flags & ~DMA_PMD_DECRYPTED);
+ }
+ return 0;
+}
+
/*
* Releasing an idle PMD page proceeds in three stages:
*
@@ -74,7 +148,7 @@ static void dma_pmd_pool_free_kref(struct kref *kref)
* pushed locklessly onto dma_pmd_free_list and dma_pmd_reclaim_work is
* scheduled on dma_pmd_wq.
* 2. In process context, dma_pmd_release_page() unmaps all IOMMU domains
- * (dma_pmd_unmap_all()), clears meta->pooled so new lockless readers stop
+ * (dma_pmd_unmap_all()), clears DMA_PMD_POOLED so new lockless readers stop
* entering @meta, and queues dma_pmd_release_page_rcu() via call_rcu().
* 3. After an RCU grace period (once no concurrent dma_pmd_free_page() reader
* can still be dereferencing @meta), dma_pmd_release_page_rcu() unfreezes
@@ -97,22 +171,21 @@ static void dma_pmd_release_page_rcu(struct rcu_head *head)
* block is freed, the PMD frame can be reallocated and @meta overwritten.
*/
struct dma_pmd_meta *meta = container_of(head, struct dma_pmd_meta, rcu);
- struct page *page = pfn_to_page(dma_pmd_meta_to_pfn(meta));
+ struct page *block, *page = pfn_to_page(dma_pmd_meta_to_pfn(meta));
struct dma_pmd_pool *pool = READ_ONCE(meta->pool);
- unsigned int nr = DMA_PMD_BLOCKS(pool->order);
- unsigned int order = pool->order;
+ unsigned int order = pool->order, i, nr = DMA_PMD_BLOCKS(order);
+ bool leak = READ_ONCE(meta->flags) & DMA_PMD_DECRYPTED;
unsigned long flags;
- unsigned int i;
-
- for (i = 0; i < nr; i++) {
- struct page *block = page + (i << order);
- page_ref_unfreeze(block, 1);
- __free_pages(block, order);
+ if (!leak) {
+ for (i = 0; i < nr; i++) {
+ block = page + (i << order);
+ page_ref_unfreeze(block, 1);
+ __free_pages(block, order);
+ }
+ atomic_long_dec(&dma_pmd_nr_pages);
}
- atomic_long_dec(&dma_pmd_nr_pages);
-
spin_lock_irqsave(&pool->lock, flags);
pool->pmd_free_cnt++;
spin_unlock_irqrestore(&pool->lock, flags);
@@ -138,7 +211,9 @@ static void dma_pmd_release_page(struct dma_pmd_meta *meta)
* dma_is_pmd_page() just before this may still be reading @meta, so
* the PMD page and its metadata entry must outlive the grace period.
*/
- WRITE_ONCE(meta->pooled, false);
+ WRITE_ONCE(meta->flags, meta->flags & ~DMA_PMD_POOLED);
+ dma_pmd_page_unprepare(meta);
+
call_rcu(&meta->rcu, dma_pmd_release_page_rcu);
}
@@ -392,20 +467,33 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
return NULL;
}
- /* Fails with -EBUSY if a concurrent PFN walker holds a speculative ref. */
- if (split_page_compound(page, PMD_ORDER, pool->order)) {
+ meta = dma_pmd_meta_of_pfn(page_to_pfn(page));
+ memset(meta, 0, sizeof(*meta));
+ meta->pool = pool;
+ spin_lock_init(&meta->map_lock);
+
+ if (dma_pmd_page_prepare(page, meta, can_block)) {
__free_pages(page, PMD_ORDER);
atomic_long_dec(&dma_pmd_nr_pages);
return NULL;
}
- meta = dma_pmd_meta_of_pfn(page_to_pfn(page));
- memset(meta, 0, sizeof(*meta));
- meta->pool = pool;
- spin_lock_init(&meta->map_lock);
+ /* Fails with -EBUSY if a concurrent PFN walker holds a speculative ref. */
+ if (split_page_compound(page, PMD_ORDER, pool->order)) {
+ if (!dma_pmd_page_unprepare(meta)) {
+ __free_pages(page, PMD_ORDER);
+ atomic_long_dec(&dma_pmd_nr_pages);
+ }
+ return NULL;
+ }
- bitmap_set(meta->dirty_bitmap, 0, nr);
- meta->nr_dirty = nr;
+ if (meta->flags & DMA_PMD_DECRYPTED) {
+ bitmap_set(meta->free_bitmap, 0, nr);
+ meta->nr_free = nr;
+ } else {
+ bitmap_set(meta->dirty_bitmap, 0, nr);
+ meta->nr_dirty = nr;
+ }
kref_get(&pool->refcount);
@@ -418,7 +506,7 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
* function returns and the caller publishes the PMD page under
* @pool->lock, whose release orders both stores against every consumer.
*/
- WRITE_ONCE(meta->pooled, true);
+ WRITE_ONCE(meta->flags, meta->flags | DMA_PMD_POOLED);
return meta;
}
@@ -764,10 +852,18 @@ static unsigned long dma_pmd_shrink_scan(struct shrinker *shrink,
/*
* Retiring is done outside both locks: dma_pmd_release_page() unmaps
* from the IOMMU, which is far too long to hold pool->lock for.
+ * Pages that need set_memory_encrypted() cannot be released under
+ * fs_reclaim (CPA allocates PTE pages on x86); hand them to the
+ * reclaim workqueue instead.
*/
list_for_each_entry_safe(meta, tmp, &release_list, list) {
- list_del(&meta->list);
- dma_pmd_release_page(meta);
+ list_del_init(&meta->list);
+ if (meta->flags & (DMA_PMD_DECRYPTED | DMA_PMD_PINNED)) {
+ llist_add(&meta->llnode, &dma_pmd_free_list);
+ dma_pmd_schedule_reclaim();
+ } else {
+ dma_pmd_release_page(meta);
+ }
}
srcu_read_unlock(&dma_pmd_srcu, srcu_idx);
@@ -806,6 +902,9 @@ static int __init dma_pmd_init(void)
if (!dma_pmd_max_pages)
dma_pmd_max_pages = max(totalram_pages() >> (PMD_ORDER + 3), 16UL);
+ if (cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT))
+ dma_pmd_decrypt = true;
+
/*
* Reclaim frees memory, so it must not queue behind arbitrary work on
* system_wq when the machine is already short of it. Failure is not
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index b411ecb1fdafc..8c7a50d5b7220 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -107,10 +107,16 @@ enum {
#define DMA_PMD_META_SHIFT 8
#define DMA_PMD_META_SIZE BIT(DMA_PMD_META_SHIFT)
+enum {
+ DMA_PMD_POOLED = BIT(0),
+ DMA_PMD_DECRYPTED = BIT(1),
+ DMA_PMD_PINNED = BIT(2),
+};
+
/**
* struct dma_pmd_meta - Metadata for a single DMA_PMD page
- * @pooled: True while this page is owned by an dma_pmd_pool (offset 0)
- * read locklessly by dma_is_pmd_page().
+ * @flags: Bitmask of DMA_PMD_POOLED, DMA_PMD_DECRYPTED, DMA_PMD_PINNED
+ * (offset 0), read locklessly by dma_is_pmd_page() and dma_is_pmd_direct().
* @nr_free: Number of zeroed available blocks in @free_bitmap
* @nr_dirty: Number of dirty available blocks in @dirty_bitmap
* @map_lock: Spinlock protecting slow-path updates to @domains_mapped
@@ -138,7 +144,7 @@ enum {
* physical pages (128 KB per GB of RAM).
*/
struct dma_pmd_meta {
- bool pooled;
+ u8 flags;
u16 nr_free;
u16 nr_dirty;
spinlock_t map_lock;
@@ -160,7 +166,7 @@ struct dma_pmd_meta {
DECLARE_BITMAP(dirty_bitmap, 1U << PMD_ORDER) ____cacheline_aligned;
} __aligned(DMA_PMD_META_SIZE);
-static_assert(offsetof(struct dma_pmd_meta, pooled) == 0);
+static_assert(offsetof(struct dma_pmd_meta, flags) == 0);
static_assert(sizeof(struct dma_pmd_meta) == DMA_PMD_META_SIZE);
static inline unsigned int dma_pmd_meta_avail(const struct dma_pmd_meta *m)
diff --git a/include/linux/dma-pmd.h b/include/linux/dma-pmd.h
index 7b7e9477d44e1..c9775e86979f9 100644
--- a/include/linux/dma-pmd.h
+++ b/include/linux/dma-pmd.h
@@ -16,6 +16,7 @@ struct dma_pmd_pool;
extern void *dma_pmd_meta_array;
bool __dma_is_pmd_page(unsigned long pfn);
+bool __dma_is_pmd_direct(unsigned long pfn);
static inline bool dma_is_pmd_page(unsigned long pfn)
{
@@ -29,6 +30,15 @@ static inline bool dma_is_pmd_page(unsigned long pfn)
return __dma_is_pmd_page(pfn);
}
+static inline bool dma_is_pmd_direct(phys_addr_t phys)
+{
+ /* Pairs with smp_store_release() in dma_pmd_meta_init(). */
+ if (likely(!smp_load_acquire(&dma_pmd_meta_array)))
+ return false;
+
+ return __dma_is_pmd_direct(phys >> PAGE_SHIFT);
+}
+
bool __dma_pmd_free_page(struct page *page);
static inline bool dma_pmd_free_page(struct page *page)
@@ -117,6 +127,11 @@ static inline bool dma_is_pmd_page(unsigned long pfn)
return false;
}
+static inline bool dma_is_pmd_direct(phys_addr_t phys)
+{
+ return false;
+}
+
static inline bool dma_pmd_free_page(struct page *page)
{
return false;
diff --git a/kernel/dma/direct.c b/kernel/dma/direct.c
index da665ca22d5c0..a786c2869b869 100644
--- a/kernel/dma/direct.c
+++ b/kernel/dma/direct.c
@@ -7,6 +7,7 @@
#include <linux/memblock.h> /* for max_pfn */
#include <linux/export.h>
#include <linux/mm.h>
+#include <linux/dma-pmd.h>
#include <linux/dma-map-ops.h>
#include <linux/scatterlist.h>
#include <linux/pfn.h>
@@ -679,7 +680,10 @@ dma_addr_t dma_direct_map_phys(struct device *dev, phys_addr_t phys,
attrs |= DMA_ATTR_CC_SHARED;
}
- if (is_swiotlb_force_bounce(dev)) {
+ if (dma_is_pmd_direct(phys)) {
+ if (force_dma_unencrypted(dev))
+ attrs |= DMA_ATTR_CC_SHARED;
+ } else if (is_swiotlb_force_bounce(dev)) {
if (attrs & (DMA_ATTR_MMIO | DMA_ATTR_REQUIRE_COHERENT))
return DMA_MAPPING_ERROR;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 15/22] iommu/dma: Add background page scrubber for DMA_PMD pools
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (13 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 14/22] iommu/dma: Support decrypted and pinned DMA_PMD pages Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 16/22] iommu/dma: Add per-NUMA-node PMD page reservoir Luigi Rizzo
` (6 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Add a background scrubber worker (dma_pmd_scrub_work) that zeroes dirty
blocks in meta->dirty_bitmap[] and moves them to meta->free_bitmap[],
and change __GFP_ZERO (want_init_on_alloc) allocations to take blocks
exclusively from free_bitmap[]:
- Remove the per-block memset() and out_prep: loop from
dma_pmd_pool_alloc_bulk_node(). When dma_pmd_add_page() allocates a
new PMD page for a zeroed request, zero the 2MB page once and place
all blocks in free_bitmap[].
- Maintain pool->partial (nr_free > 0, with clean-only pages at head and
mixed pages at tail), pool->partial_dirty (nr_free == 0, nr_dirty > 0),
and pool->idle (clean pages at head, dirty pages at tail) so both
allocation and scrubbing find candidate PMD pages in O(1) time.
- Trigger dma_pmd_scrub_work only when meta->nr_dirty transitions from 0
to 1 on block free (or when a non-zeroed PMD page is added).
- Detach up to DMA_PMD_SCRUB_BATCH (64) dirty blocks under pool->lock,
zero them outside the lock, and re-insert them into free_bitmap[].
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/dma-pmd-pool.c | 336 ++++++++++++++++++++++++++++++-----
drivers/iommu/dma-pmd-priv.h | 9 +-
2 files changed, 296 insertions(+), 49 deletions(-)
diff --git a/drivers/iommu/dma-pmd-pool.c b/drivers/iommu/dma-pmd-pool.c
index deacb6f7d3610..d774cd05347b8 100644
--- a/drivers/iommu/dma-pmd-pool.c
+++ b/drivers/iommu/dma-pmd-pool.c
@@ -32,6 +32,8 @@
#include "dma-pmd-priv.h"
+#define DMA_PMD_SCRUB_BATCH 64U
+
/*
* DMA_PMD pages whose last block has just been freed and which are over the
* pool's idle watermark are pushed here locklessly for later release.
@@ -66,6 +68,7 @@ static DEFINE_MUTEX(dma_pmd_pools_lock);
static LLIST_HEAD(dma_pmd_dead_pools);
static void dma_pmd_schedule_reclaim(void);
+static void dma_pmd_schedule_scrub(void);
static void dma_pmd_pool_free_kref(struct kref *kref)
{
@@ -250,6 +253,216 @@ static void dma_pmd_schedule_reclaim(void)
schedule_work(&dma_pmd_reclaim_work);
}
+/*
+ * Background page scrubber
+ * ------------------------
+ * Freed blocks are returned to @meta->dirty_bitmap so __dma_pmd_free_page()
+ * never pays a memset() on the free hot path. Callers that do not require
+ * zeroed memory take dirty blocks first, while callers with __GFP_ZERO (or
+ * init_on_alloc) take pre-zeroed blocks from @meta->free_bitmap.
+ *
+ * The scrubber runs on system_wq to asynchronously convert dirty blocks into
+ * zeroed blocks in @meta->free_bitmap. To avoid holding @pool->lock across
+ * memset(), it detaches up to DMA_PMD_SCRUB_BATCH (64) dirty bits under
+ * @pool->lock, zeroes the pages unlocked (while their bits are clear in both
+ * bitmaps so no allocator or double-free check can race on them), and
+ * republishes them in @meta->free_bitmap under @pool->lock.
+ *
+ * Triggered lazily (without touching any list on every __dma_pmd_free_page()):
+ * when a PMD page transitions from 0 to 1 dirty block, when a page with dirty
+ * blocks becomes completely idle, or when a __GFP_ZERO allocation misses on
+ * clean pages while dirty pages are present.
+ */
+
+/*
+ * Place @meta on the appropriate pool list according to its available clean
+ * (@nr_free) and dirty (@nr_dirty) block counts. Caller must hold @pool->lock
+ * and update @pool->num_idle_pages when moving into or out of @pool->idle.
+ *
+ * List invariants:
+ * - @pool->full: avail == 0
+ * - @pool->partial: 0 < avail < nr, nr_free > 0
+ * (nr_dirty == 0 at head, nr_dirty > 0 at tail)
+ * - @pool->partial_dirty: 0 < avail < nr, nr_free == 0, nr_dirty > 0
+ * - @pool->idle: avail == nr
+ * (nr_dirty == 0 at head, nr_dirty > 0 at tail)
+ */
+static void dma_pmd_place_meta(struct dma_pmd_pool *pool,
+ struct dma_pmd_meta *meta)
+{
+ unsigned int avail = dma_pmd_meta_avail(meta);
+ unsigned int nr = DMA_PMD_BLOCKS(pool->order);
+
+ if (!avail) {
+ list_move(&meta->list, &pool->full);
+ } else if (avail == nr) {
+ if (!meta->nr_dirty)
+ list_move(&meta->list, &pool->idle);
+ else
+ list_move_tail(&meta->list, &pool->idle);
+ } else if (!meta->nr_free) {
+ list_move_tail(&meta->list, &pool->partial_dirty);
+ } else if (!meta->nr_dirty) {
+ list_move(&meta->list, &pool->partial);
+ } else {
+ list_move_tail(&meta->list, &pool->partial);
+ }
+}
+
+static struct dma_pmd_meta *dma_pmd_pick_dirty_meta(struct dma_pmd_pool *pool)
+{
+ struct dma_pmd_meta *meta;
+
+ meta = list_first_entry_or_null(&pool->partial_dirty,
+ struct dma_pmd_meta, list);
+ if (meta)
+ return meta;
+
+ if (!list_empty(&pool->partial)) {
+ meta = list_last_entry(&pool->partial, struct dma_pmd_meta, list);
+ if (meta->nr_dirty)
+ return meta;
+ }
+
+ if (!list_empty(&pool->idle)) {
+ meta = list_last_entry(&pool->idle, struct dma_pmd_meta, list);
+ if (meta->nr_dirty) {
+ pool->num_idle_pages--;
+ return meta;
+ }
+ }
+
+ return NULL;
+}
+
+/**
+ * dma_pmd_scrub_pool - Scrub one batch of dirty blocks in @pool
+ * @pool: Pool to scrub
+ *
+ * Cost: Zeroes up to DMA_PMD_SCRUB_BATCH blocks (e.g. up to 256 KB for
+ * order-0) outside @pool->lock; two brief O(batch) bitmap passes under
+ * @pool->lock.
+ * Locking: Process context. Acquires @pool->lock (irqsave) to detach and
+ * republish block bits; drops lock during memset().
+ * Frequency: Background worker only (on system_wq).
+ *
+ * Return: true if a batch was scrubbed (more dirty blocks may remain),
+ * false if @pool has no dirty blocks left or is destroyed.
+ */
+static bool dma_pmd_scrub_pool(struct dma_pmd_pool *pool)
+{
+ unsigned int nr, order, idx, count = 0, used = 0;
+ u16 idxs[DMA_PMD_SCRUB_BATCH];
+ struct dma_pmd_meta *meta;
+ bool should_release = false;
+ unsigned long base_pfn;
+ unsigned long flags;
+
+ spin_lock_irqsave(&pool->lock, flags);
+ if (pool->destroyed) {
+ spin_unlock_irqrestore(&pool->lock, flags);
+ return false;
+ }
+
+ meta = dma_pmd_pick_dirty_meta(pool);
+ if (!meta) {
+ spin_unlock_irqrestore(&pool->lock, flags);
+ return false;
+ }
+
+ order = pool->order;
+ nr = DMA_PMD_BLOCKS(order);
+ base_pfn = dma_pmd_meta_to_pfn(meta);
+
+ for_each_set_bit(idx, meta->dirty_bitmap, nr) {
+ struct page *block = pfn_to_page(base_pfn + (idx << order));
+
+ __clear_bit(idx, meta->dirty_bitmap);
+ used++;
+ if (unlikely(folio_contain_hwpoisoned_page(page_folio(block)))) {
+ pr_err_once("dma_pmd: poisoned block at pfn %lu, leaking it to keep pool %p intact\n",
+ page_to_pfn(block), pool);
+ } else {
+ idxs[count++] = idx;
+ }
+ if (used == meta->nr_dirty || count == DMA_PMD_SCRUB_BATCH)
+ break;
+ }
+ meta->nr_dirty -= used;
+ dma_pmd_place_meta(pool, meta);
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+ if (unlikely(!count))
+ return used > 0;
+
+ for (idx = 0; idx < count; idx++) {
+ struct page *block = pfn_to_page(base_pfn + (idxs[idx] << order));
+
+ memset(page_address(block), 0, PAGE_SIZE << order);
+ }
+
+ spin_lock_irqsave(&pool->lock, flags);
+ for (idx = 0; idx < count; idx++)
+ __set_bit(idxs[idx], meta->free_bitmap);
+ meta->nr_free += count;
+ pool->block_scrub_cnt += count;
+
+ if (dma_pmd_meta_avail(meta) == nr) {
+ if (pool->destroyed || pool->num_idle_pages >= pool->max_idle_pages) {
+ list_del_init(&meta->list);
+ llist_add(&meta->llnode, &dma_pmd_free_list);
+ should_release = true;
+ } else {
+ pool->num_idle_pages++;
+ dma_pmd_place_meta(pool, meta);
+ }
+ } else if (!pool->destroyed) {
+ dma_pmd_place_meta(pool, meta);
+ }
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+ if (should_release)
+ dma_pmd_schedule_reclaim();
+
+ return true;
+}
+
+/*
+ * Walk all live pools and scrub dirty blocks until every pool is clean.
+ *
+ * Cost: Proportional to total dirty blocks across all pools; yields via
+ * cond_resched() between batches of DMA_PMD_SCRUB_BATCH blocks.
+ * Locking: Process context (system_wq). Holds @dma_pmd_pools_lock across the
+ * pool walk.
+ * Frequency: Scheduled on demand via dma_pmd_schedule_scrub(); coalesced by
+ * work_pending().
+ */
+static void dma_pmd_scrub_work_fn(struct work_struct *work)
+{
+ struct dma_pmd_pool *pool;
+ bool progress;
+
+ do {
+ progress = false;
+ mutex_lock(&dma_pmd_pools_lock);
+ list_for_each_entry(pool, &dma_pmd_pools, node) {
+ while (dma_pmd_scrub_pool(pool)) {
+ progress = true;
+ cond_resched();
+ }
+ }
+ mutex_unlock(&dma_pmd_pools_lock);
+ } while (progress);
+}
+
+static DECLARE_WORK(dma_pmd_scrub_work, dma_pmd_scrub_work_fn);
+
+static void dma_pmd_schedule_scrub(void)
+{
+ if (!work_pending(&dma_pmd_scrub_work))
+ schedule_work(&dma_pmd_scrub_work);
+}
+
unsigned int dma_pmd_pools_forget_domain(int idx)
{
struct dma_pmd_pool *pool;
@@ -262,6 +475,8 @@ unsigned int dma_pmd_pools_forget_domain(int idx)
spin_lock_irqsave(&pool->lock, flags);
list_for_each_entry(meta, &pool->partial, list)
dropped += dma_pmd_forget_domain(meta, idx);
+ list_for_each_entry(meta, &pool->partial_dirty, list)
+ dropped += dma_pmd_forget_domain(meta, idx);
list_for_each_entry(meta, &pool->idle, list)
dropped += dma_pmd_forget_domain(meta, idx);
list_for_each_entry(meta, &pool->full, list)
@@ -311,6 +526,7 @@ struct dma_pmd_pool *dma_pmd_pool_create(unsigned int order, unsigned int max_id
pool->max_idle_pages = max_idle_pages ? : 16;
pool->next_alloc_attempt = jiffies;
INIT_LIST_HEAD(&pool->partial);
+ INIT_LIST_HEAD(&pool->partial_dirty);
INIT_LIST_HEAD(&pool->idle);
INIT_LIST_HEAD(&pool->full);
@@ -368,6 +584,7 @@ struct dma_pmd_pool *dma_pmd_pool_destroy(struct dma_pmd_pool *pool)
}
srcu_read_unlock(&dma_pmd_srcu, srcu_idx);
+ flush_work(&dma_pmd_scrub_work);
kref_put(&pool->refcount, dma_pmd_pool_free_kref);
flush_work(&dma_pmd_reclaim_work);
return NULL;
@@ -379,6 +596,7 @@ EXPORT_SYMBOL(dma_pmd_pool_destroy);
* @pool: Owning pool
* @gfp: GFP allocation flags
* @nid: Target NUMA node (or NUMA_NO_NODE for local node)
+ * @zero: True if caller requires zeroed blocks in @meta->free_bitmap
*
* Cost: Slow path (pool miss); allocates PMD page from buddy on @nid,
* splits into compound blocks, and sets the PMD page's membership bit.
@@ -390,7 +608,7 @@ EXPORT_SYMBOL(dma_pmd_pool_destroy);
* on failure.
*/
static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
- gfp_t gfp, int nid)
+ gfp_t gfp, int nid, bool zero)
{
/*
* Require full GFP_KERNEL (not just gfpflags_allow_blocking()):
@@ -431,7 +649,7 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
* is shaped into compound pieces by hand. What is left is also free
* of everything slab treats as a bug in GFP_SLAB_BUG_MASK, so the
* metadata allocation below can use it as it stands.
- * __GFP_ZERO is handled per-block in dma_pmd_pool_alloc_bulk_node().
+ * __GFP_ZERO is handled explicitly below after dma_pmd_page_prepare().
*/
base_gfp = gfp & ~(__GFP_COMP | __GFP_HIGHMEM | __GFP_MOVABLE | __GFP_ZERO);
@@ -487,7 +705,9 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
return NULL;
}
- if (meta->flags & DMA_PMD_DECRYPTED) {
+ if (zero || (meta->flags & DMA_PMD_DECRYPTED)) {
+ if (!(meta->flags & DMA_PMD_DECRYPTED))
+ memset(page_address(page), 0, PMD_SIZE);
bitmap_set(meta->free_bitmap, 0, nr);
meta->nr_free = nr;
} else {
@@ -514,8 +734,7 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
static unsigned int dma_pmd_scan_bitmap(struct dma_pmd_pool *pool,
unsigned long *bitmap, u16 *countp,
unsigned long base_pfn, unsigned int nr,
- unsigned long want, struct page **out,
- bool unfreeze)
+ unsigned long want, struct page **out)
{
unsigned int idx, got = 0, used = 0;
@@ -531,8 +750,7 @@ static unsigned int dma_pmd_scan_bitmap(struct dma_pmd_pool *pool,
pr_err_once("dma_pmd: poisoned block at pfn %lu, leaking it to keep pool %p intact\n",
page_to_pfn(block), pool);
} else {
- if (unfreeze)
- page_ref_unfreeze(block, 1);
+ page_ref_unfreeze(block, 1);
out[got++] = block;
}
@@ -544,6 +762,41 @@ static unsigned int dma_pmd_scan_bitmap(struct dma_pmd_pool *pool,
return got;
}
+static struct dma_pmd_meta *dma_pmd_pick_meta(struct dma_pmd_pool *pool, bool zero)
+{
+ struct dma_pmd_meta *meta;
+
+ if (!zero) {
+ meta = list_first_entry_or_null(&pool->partial_dirty,
+ struct dma_pmd_meta, list);
+ if (meta)
+ return meta;
+ if (!list_empty(&pool->partial))
+ return list_last_entry(&pool->partial,
+ struct dma_pmd_meta, list);
+ if (!list_empty(&pool->idle)) {
+ meta = list_last_entry(&pool->idle,
+ struct dma_pmd_meta, list);
+ pool->num_idle_pages--;
+ return meta;
+ }
+ return NULL;
+ }
+
+ meta = list_first_entry_or_null(&pool->partial, struct dma_pmd_meta, list);
+ if (meta)
+ return meta;
+
+ list_for_each_entry(meta, &pool->idle, list) {
+ if (meta->nr_free) {
+ pool->num_idle_pages--;
+ return meta;
+ }
+ }
+
+ return NULL;
+}
+
/* Extract up to @want available blocks from @meta under @pool->lock. */
static unsigned long dma_pmd_take_blocks(struct dma_pmd_pool *pool,
struct dma_pmd_meta *meta,
@@ -555,23 +808,18 @@ static unsigned long dma_pmd_take_blocks(struct dma_pmd_pool *pool,
unsigned int nr = DMA_PMD_BLOCKS(pool->order);
unsigned int got = 0, used;
- if (zero) {
- got += dma_pmd_scan_bitmap(pool, meta->free_bitmap, &meta->nr_free,
- base_pfn, nr, want, out, true);
+ if (!zero)
got += dma_pmd_scan_bitmap(pool, meta->dirty_bitmap, &meta->nr_dirty,
- base_pfn, nr, want - got, out + got, false);
- } else {
- got += dma_pmd_scan_bitmap(pool, meta->dirty_bitmap, &meta->nr_dirty,
- base_pfn, nr, want, out, false);
- got += dma_pmd_scan_bitmap(pool, meta->free_bitmap, &meta->nr_free,
- base_pfn, nr, want - got, out + got, true);
- }
+ base_pfn, nr, want, out);
+ got += dma_pmd_scan_bitmap(pool, meta->free_bitmap, &meta->nr_free,
+ base_pfn, nr, want - got, out + got);
+
used = avail_before - dma_pmd_meta_avail(meta);
- if (!dma_pmd_meta_avail(meta) || WARN_ON_ONCE(!used)) {
+ if (WARN_ON_ONCE(!used)) {
meta->nr_free = 0;
meta->nr_dirty = 0;
- list_move(&meta->list, &pool->full);
}
+ dma_pmd_place_meta(pool, meta);
return got;
}
@@ -588,7 +836,7 @@ static unsigned long dma_pmd_take_blocks(struct dma_pmd_pool *pool,
* under a single lock acquisition. Slow path (both empty) calls
* dma_pmd_add_page().
* Locking: Acquires @pool->lock (irqsave) for the bitmap scan. Drops lock
- * during slow-path PMD buddy allocation and per-block initialization.
+ * during slow-path PMD buddy allocation.
* Frequency: High (called per RX buffer refill or per SKB TX page frag refill).
*
* A pool hands out a block of a PMD page it already owns, on the node that
@@ -605,7 +853,8 @@ unsigned long dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp,
int nid, unsigned long nr_pages,
struct page **page_array)
{
- unsigned long allocated = 0, flags, i;
+ unsigned long allocated = 0, flags;
+ bool should_scrub = false;
bool zero;
if (unlikely(!pool || pool->destroyed || !nr_pages ||
@@ -619,22 +868,18 @@ unsigned long dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp,
while (allocated < nr_pages) {
struct dma_pmd_meta *meta;
- /* Partially used pages first, idle ones next. */
- meta = list_first_entry_or_null(&pool->partial, struct dma_pmd_meta, list);
- if (!meta && !list_empty(&pool->idle)) {
- meta = list_first_entry(&pool->idle, struct dma_pmd_meta, list);
- list_move(&meta->list, &pool->partial);
- pool->num_idle_pages--;
- }
+ meta = dma_pmd_pick_meta(pool, zero);
if (!meta) {
/* Fallback to a new allocation */
spin_unlock_irqrestore(&pool->lock, flags);
- meta = dma_pmd_add_page(pool, gfp, nid);
+ meta = dma_pmd_add_page(pool, gfp, nid, zero);
if (!meta)
- goto out_prep;
+ goto out;
spin_lock_irqsave(&pool->lock, flags);
pool->pmd_alloc_cnt++;
list_add(&meta->list, &pool->partial);
+ if (meta->nr_dirty)
+ should_scrub = true;
}
allocated += dma_pmd_take_blocks(pool, meta,
@@ -643,14 +888,9 @@ unsigned long dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp,
}
spin_unlock_irqrestore(&pool->lock, flags);
-out_prep:
- for (i = 0; i < allocated; i++) {
- if (!page_count(page_array[i])) {
- page_ref_unfreeze(page_array[i], 1);
- if (zero)
- memset(page_address(page_array[i]), 0, PAGE_SIZE << pool->order);
- }
- }
+out:
+ if (should_scrub)
+ dma_pmd_schedule_scrub();
return allocated;
}
EXPORT_SYMBOL(dma_pmd_pool_alloc_bulk_node);
@@ -674,7 +914,9 @@ EXPORT_SYMBOL(dma_pmd_pool_alloc_bulk_node);
bool dma_pmd_pool_has_free(struct dma_pmd_pool *pool)
{
return pool && !READ_ONCE(pool->destroyed) &&
- (!list_empty_careful(&pool->partial) || !list_empty_careful(&pool->idle));
+ (!list_empty_careful(&pool->partial) ||
+ !list_empty_careful(&pool->partial_dirty) ||
+ !list_empty_careful(&pool->idle));
}
EXPORT_SYMBOL(dma_pmd_pool_has_free);
@@ -706,8 +948,8 @@ bool __dma_pmd_free_page(struct page *page)
struct dma_pmd_pool *pool = meta->pool;
unsigned long pfn = page_to_pfn(page);
bool should_release = false;
+ bool should_scrub = false;
bool zeroed_on_free = false;
- unsigned int avail;
unsigned long flags;
idx = (pfn - dma_pmd_meta_to_pfn(meta)) >> pool->order;
@@ -755,28 +997,28 @@ bool __dma_pmd_free_page(struct page *page)
meta->nr_free++;
} else {
__set_bit(idx, meta->dirty_bitmap);
- meta->nr_dirty++;
+ should_scrub = !meta->nr_dirty++;
}
- avail = dma_pmd_meta_avail(meta);
pool->block_free_cnt++;
- if (avail == 1 && !pool->destroyed)
- list_move(&meta->list, &pool->partial);
-
- if (avail == nr) {
+ if (dma_pmd_meta_avail(meta) == nr) {
if (pool->destroyed || pool->num_idle_pages >= pool->max_idle_pages) {
list_del_init(&meta->list);
llist_add(&meta->llnode, &dma_pmd_free_list);
should_release = true;
} else {
- list_move(&meta->list, &pool->idle);
pool->num_idle_pages++;
+ dma_pmd_place_meta(pool, meta);
}
+ } else if (!pool->destroyed) {
+ dma_pmd_place_meta(pool, meta);
}
spin_unlock_irqrestore(&pool->lock, flags);
if (should_release)
dma_pmd_schedule_reclaim();
+ else if (should_scrub)
+ dma_pmd_schedule_scrub();
return true;
}
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index 8c7a50d5b7220..82a5c869aa216 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -181,8 +181,10 @@ static inline unsigned int dma_pmd_meta_avail(const struct dma_pmd_meta *m)
* block bitmaps and the statistics counters
* @order: Block order managed by this pool (<= PMD_ORDER)
* @destroyed: Set when dma_pmd_pool_destroy() has been called
- * @partial: List of PMD pages with at least 1 free block, but not all
- * of them (MRU ordered). Candidates for allocation.
+ * @partial: List of partially used PMD pages with @nr_free > 0 (clean-only
+ * at head, mixed clean/dirty at tail)
+ * @partial_dirty: List of partially used PMD pages with @nr_free == 0 and
+ * @nr_dirty > 0
* @idle: List of PMD pages whose every usable block is free. Held
* separately because they can be released under pressure.
* @full: List of PMD pages with 0 free blocks
@@ -193,6 +195,7 @@ static inline unsigned int dma_pmd_meta_avail(const struct dma_pmd_meta *m)
* @pmd_free_cnt: Statistics counter of PMD pages released back to buddy
* @block_alloc_cnt: Statistics counter of block allocations satisfied
* @block_free_cnt: Statistics counter of block frees recycled into pool
+ * @block_scrub_cnt: Statistics counter of dirty blocks zeroed by scrubber
* @next_alloc_attempt: Do not attempt a new order-9 allocation before this time.
* Damps repeated high-order GFP_ATOMIC failures under
* fragmentation, which would otherwise be retried on every
@@ -207,6 +210,7 @@ struct dma_pmd_pool {
u8 order;
bool destroyed;
struct list_head partial;
+ struct list_head partial_dirty;
struct list_head idle;
struct list_head full;
unsigned int num_idle_pages;
@@ -217,6 +221,7 @@ struct dma_pmd_pool {
u64 pmd_free_cnt;
u64 block_alloc_cnt;
u64 block_free_cnt;
+ u64 block_scrub_cnt;
/* Cold. */
unsigned long next_alloc_attempt;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 16/22] iommu/dma: Add per-NUMA-node PMD page reservoir
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (14 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 15/22] iommu/dma: Add background page scrubber for DMA_PMD pools Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 17/22] net/gve: Use DMA_PMD memory for RX buffers Luigi Rizzo
` (5 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
Add a per-NUMA-node reservoir of pre-allocated, pre-zeroed, and
pre-prepared 2MB pages (dma_pmd_reservoirs[MAX_NUMNODES]) refilled in
process context by a background worker (dma_pmd_reservoir_work):
- Controlled by dma_pmd.reservoir_max (default 16 PMD pages = 32 MB per
N_MEMORY node, subject to the global dma_pmd_max_pages ceiling).
- dma_pmd_add_page() first tries dma_pmd_reservoir_get(nid) (preferring
the target NUMA node and falling back to other N_MEMORY nodes),
avoiding blocking or failing order-9 buddy allocations on pool misses.
- When a node's reservoir drops to <= reservoir_max / 2 (or on pool
creation), dma_pmd_reservoir_work is scheduled to refill in the
background.
- Integrate the reservoir with dma_pmd_shrink_count() and
dma_pmd_shrink_scan() so reservoir pages are reclaimed first under
memory pressure.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/dma-pmd-pool.c | 307 +++++++++++++++++++++++++++++++----
1 file changed, 277 insertions(+), 30 deletions(-)
diff --git a/drivers/iommu/dma-pmd-pool.c b/drivers/iommu/dma-pmd-pool.c
index d774cd05347b8..d48c8e19741ab 100644
--- a/drivers/iommu/dma-pmd-pool.c
+++ b/drivers/iommu/dma-pmd-pool.c
@@ -21,6 +21,7 @@
#include <linux/mm.h>
#include <linux/moduleparam.h>
#include <linux/mutex.h>
+#include <linux/nodemask.h>
#include <linux/rcupdate.h>
#include <linux/seq_file.h>
#include <linux/set_memory.h>
@@ -62,6 +63,18 @@ static atomic_long_t dma_pmd_nr_pages __cacheline_aligned_in_smp;
static unsigned long dma_pmd_max_pages __read_mostly;
core_param(dma_pmd_max_pages, dma_pmd_max_pages, ulong, 0644);
+/* Target number of pre-allocated PMD pages per NUMA node in reservoir. */
+static unsigned int dma_pmd_reservoir_max __read_mostly = 16;
+module_param_named(reservoir_max, dma_pmd_reservoir_max, uint, 0644);
+
+struct dma_pmd_reservoir {
+ spinlock_t lock; /* protects @pages and @nr_pages */
+ struct list_head pages;
+ unsigned int nr_pages;
+};
+
+static struct dma_pmd_reservoir dma_pmd_reservoirs[MAX_NUMNODES];
+
/* All live pools. */
static LIST_HEAD(dma_pmd_pools);
static DEFINE_MUTEX(dma_pmd_pools_lock);
@@ -69,6 +82,7 @@ static LLIST_HEAD(dma_pmd_dead_pools);
static void dma_pmd_schedule_reclaim(void);
static void dma_pmd_schedule_scrub(void);
+static void dma_pmd_schedule_reservoir(void);
static void dma_pmd_pool_free_kref(struct kref *kref)
{
@@ -81,8 +95,10 @@ static void dma_pmd_pool_free_kref(struct kref *kref)
static int dma_pmd_page_prepare(struct page *page, struct dma_pmd_meta *meta,
bool can_block)
{
- bool want_decrypt = READ_ONCE(dma_pmd_decrypt);
- bool want_pin = READ_ONCE(dma_pmd_pin);
+ bool want_decrypt = READ_ONCE(dma_pmd_decrypt) &&
+ !(meta->flags & DMA_PMD_DECRYPTED);
+ bool want_pin = READ_ONCE(dma_pmd_pin) &&
+ !(meta->flags & DMA_PMD_PINNED);
void *vaddr = page_address(page);
int ret;
@@ -142,6 +158,16 @@ static int dma_pmd_page_unprepare(struct dma_pmd_meta *meta)
return 0;
}
+static void dma_pmd_free_reservoir_page(struct dma_pmd_meta *meta)
+{
+ struct page *page = pfn_to_page(dma_pmd_meta_to_pfn(meta));
+
+ if (!dma_pmd_page_unprepare(meta)) {
+ __free_pages(page, PMD_ORDER);
+ atomic_long_dec(&dma_pmd_nr_pages);
+ }
+}
+
/*
* Releasing an idle PMD page proceeds in three stages:
*
@@ -226,8 +252,12 @@ static void dma_pmd_reclaim_work_fn(struct work_struct *work)
struct dma_pmd_pool *pool, *ptmp;
struct dma_pmd_meta *meta, *tmp;
- llist_for_each_entry_safe(meta, tmp, node, llnode)
- dma_pmd_release_page(meta);
+ llist_for_each_entry_safe(meta, tmp, node, llnode) {
+ if (meta->pool)
+ dma_pmd_release_page(meta);
+ else
+ dma_pmd_free_reservoir_page(meta);
+ }
node = llist_del_all(&dma_pmd_dead_pools);
if (node) {
@@ -463,6 +493,186 @@ static void dma_pmd_schedule_scrub(void)
schedule_work(&dma_pmd_scrub_work);
}
+/*
+ * Per-NUMA-node PMD page reservoir
+ * --------------------------------
+ * Order-9 (2MB) buddy allocations can compact/reclaim and fail or stall in
+ * atomic context (such as NAPI RX refill or softirq TX), and CoCo decryption /
+ * hypervisor pinning cannot run in atomic context at all.
+ *
+ * Each NUMA node maintains a small reservoir of up to @dma_pmd_reservoir_max
+ * (default 8 = 16 MB/node) pre-allocated, pre-zeroed, and pre-decrypted/pinned
+ * PMD pages. On a pool miss, dma_pmd_add_page() pops a ready PMD page from the
+ * target node's reservoir in O(1) under a spinlock (falling back to other
+ * nodes before calling alloc_pages_node() directly).
+ *
+ * A background worker on system_wq refills nodes whenever a node's count drops
+ * to or below half of @dma_pmd_reservoir_max, on a reservoir miss, or when a
+ * new pool is created. Pages held in the reservoir count toward the global
+ * @dma_pmd_nr_pages budget and are reclaimable by the shrinker under memory
+ * pressure.
+ */
+
+/**
+ * dma_pmd_reservoir_refill_node - Refill node @nid's reservoir up to @max
+ * @nid: NUMA node ID
+ * @max: Target number of PMD pages to hold in @nid's reservoir
+ *
+ * Cost: Slow path; allocates order-9 pages with GFP_KERNEL | __GFP_RETRY_MAYFAIL,
+ * populates metadata, decrypts/pins if configured, and zeroes 2MB per page.
+ * Locking: Process context (system_wq); sleeps during buddy allocation, CPA
+ * decryption, and cond_resched(). Takes @res->lock briefly to enqueue.
+ * Frequency: Rare (when reservoir drops to <= @max / 2 or on pool creation).
+ */
+static void dma_pmd_reservoir_refill_node(int nid, unsigned int max)
+{
+ gfp_t gfp = GFP_KERNEL | __GFP_THISNODE | __GFP_NOWARN | __GFP_RETRY_MAYFAIL;
+ struct dma_pmd_reservoir *res = &dma_pmd_reservoirs[nid];
+
+ while (READ_ONCE(res->nr_pages) < max) {
+ struct dma_pmd_meta *meta;
+ unsigned long flags;
+ struct page *page;
+
+ if (atomic_long_inc_return(&dma_pmd_nr_pages) >
+ READ_ONCE(dma_pmd_max_pages)) {
+ atomic_long_dec(&dma_pmd_nr_pages);
+ break;
+ }
+
+ page = alloc_pages_node(nid, gfp, PMD_ORDER);
+ if (!page) {
+ atomic_long_dec(&dma_pmd_nr_pages);
+ break;
+ }
+
+ if (unlikely(!dma_pmd_meta_ensure_pfn(page_to_pfn(page), true))) {
+ __free_pages(page, PMD_ORDER);
+ atomic_long_dec(&dma_pmd_nr_pages);
+ break;
+ }
+
+ meta = dma_pmd_meta_of_pfn(page_to_pfn(page));
+ memset(meta, 0, sizeof(*meta));
+ spin_lock_init(&meta->map_lock);
+
+ if (dma_pmd_page_prepare(page, meta, true)) {
+ __free_pages(page, PMD_ORDER);
+ atomic_long_dec(&dma_pmd_nr_pages);
+ break;
+ }
+
+ if (!(meta->flags & DMA_PMD_DECRYPTED))
+ memset(page_address(page), 0, PMD_SIZE);
+
+ spin_lock_irqsave(&res->lock, flags);
+ list_add(&meta->list, &res->pages);
+ res->nr_pages++;
+ spin_unlock_irqrestore(&res->lock, flags);
+
+ cond_resched();
+ }
+}
+
+static void dma_pmd_reservoir_work_fn(struct work_struct *work)
+{
+ unsigned int max = READ_ONCE(dma_pmd_reservoir_max);
+ int nid;
+
+ if (!max || !dma_pmd_meta_base())
+ return;
+
+ for_each_node_state(nid, N_MEMORY)
+ dma_pmd_reservoir_refill_node(nid, max);
+}
+
+static DECLARE_WORK(dma_pmd_reservoir_work, dma_pmd_reservoir_work_fn);
+
+static void dma_pmd_schedule_reservoir(void)
+{
+ if (READ_ONCE(dma_pmd_reservoir_max) &&
+ !work_pending(&dma_pmd_reservoir_work))
+ schedule_work(&dma_pmd_reservoir_work);
+}
+
+/*
+ * Pop one pre-prepared PMD page from node @nid's reservoir.
+ *
+ * Cost: O(1) list pop under @res->lock.
+ * Locking: Acquires @res->lock (irqsave). Safe from any context (IRQ, softirq,
+ * process). Never sleeps.
+ * Frequency: Only on pool miss (dma_pmd_add_page()).
+ */
+static struct dma_pmd_meta *dma_pmd_reservoir_pop_node(int nid, unsigned int max)
+{
+ struct dma_pmd_reservoir *res = &dma_pmd_reservoirs[nid];
+ struct dma_pmd_meta *meta;
+ unsigned long flags;
+ bool refill = false;
+
+ if (!READ_ONCE(res->nr_pages))
+ return NULL;
+
+ spin_lock_irqsave(&res->lock, flags);
+ meta = list_first_entry_or_null(&res->pages, struct dma_pmd_meta, list);
+ if (meta) {
+ list_del_init(&meta->list);
+ res->nr_pages--;
+ refill = res->nr_pages <= max / 2;
+ }
+ spin_unlock_irqrestore(&res->lock, flags);
+
+ if (refill)
+ dma_pmd_schedule_reservoir();
+
+ return meta;
+}
+
+/**
+ * dma_pmd_reservoir_get - Obtain a pre-prepared PMD page, preferring @nid
+ * @nid: Preferred NUMA node (or NUMA_NO_NODE for local node)
+ *
+ * Tries @nid's reservoir first, then falls back to other memory nodes and
+ * schedules an async refill if a fallback or empty reservoir is encountered.
+ *
+ * Cost: O(1) on local node hit; O(MAX_NUMNODES) lockless checks on miss.
+ * Locking: Acquires @res->lock (irqsave) on non-empty node. Safe from any
+ * context. Never sleeps.
+ * Frequency: Only on pool miss (dma_pmd_add_page()).
+ *
+ * Return: Metadata pointer of a ready PMD page (not yet compound-split or
+ * bound to a pool), or NULL if all reservoirs are empty.
+ */
+static struct dma_pmd_meta *dma_pmd_reservoir_get(int nid)
+{
+ unsigned int max = READ_ONCE(dma_pmd_reservoir_max);
+ struct dma_pmd_meta *meta;
+ int target_nid, n;
+
+ if (!max)
+ return NULL;
+
+ target_nid = (nid == NUMA_NO_NODE) ? numa_mem_id() : nid;
+ if (target_nid >= 0 && target_nid < MAX_NUMNODES) {
+ meta = dma_pmd_reservoir_pop_node(target_nid, max);
+ if (meta)
+ return meta;
+ }
+
+ for_each_node_state(n, N_MEMORY) {
+ if (n == target_nid)
+ continue;
+ meta = dma_pmd_reservoir_pop_node(n, max);
+ if (meta) {
+ dma_pmd_schedule_reservoir();
+ return meta;
+ }
+ }
+
+ dma_pmd_schedule_reservoir();
+ return NULL;
+}
+
unsigned int dma_pmd_pools_forget_domain(int idx)
{
struct dma_pmd_pool *pool;
@@ -538,6 +748,7 @@ struct dma_pmd_pool *dma_pmd_pool_create(unsigned int order, unsigned int max_id
}
list_add(&pool->node, &dma_pmd_pools);
mutex_unlock(&dma_pmd_pools_lock);
+ dma_pmd_schedule_reservoir();
return pool;
}
@@ -616,10 +827,18 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
*/
bool can_block = (gfp & GFP_KERNEL) == GFP_KERNEL;
unsigned int nr = DMA_PMD_BLOCKS(pool->order);
+ bool from_reservoir = false;
struct dma_pmd_meta *meta;
gfp_t alloc_gfp, base_gfp;
struct page *page;
+ meta = dma_pmd_reservoir_get(nid);
+ if (meta) {
+ page = pfn_to_page(dma_pmd_meta_to_pfn(meta));
+ from_reservoir = true;
+ goto prepare;
+ }
+
/*
* A high-order allocation is expensive when it fails, and under fragmentation
* it fails on every pool miss. For GFP_ATOMIC, back off briefly.
@@ -687,9 +906,10 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
meta = dma_pmd_meta_of_pfn(page_to_pfn(page));
memset(meta, 0, sizeof(*meta));
- meta->pool = pool;
spin_lock_init(&meta->map_lock);
+prepare:
+ meta->pool = pool;
if (dma_pmd_page_prepare(page, meta, can_block)) {
__free_pages(page, PMD_ORDER);
atomic_long_dec(&dma_pmd_nr_pages);
@@ -705,8 +925,8 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
return NULL;
}
- if (zero || (meta->flags & DMA_PMD_DECRYPTED)) {
- if (!(meta->flags & DMA_PMD_DECRYPTED))
+ if (from_reservoir || zero || (meta->flags & DMA_PMD_DECRYPTED)) {
+ if (!from_reservoir && !(meta->flags & DMA_PMD_DECRYPTED))
memset(page_address(page), 0, PMD_SIZE);
bitmap_set(meta->free_bitmap, 0, nr);
meta->nr_free = nr;
@@ -1042,14 +1262,17 @@ static unsigned long dma_pmd_shrink_count(struct shrinker *shrink,
{
struct dma_pmd_pool *pool;
unsigned long nr = 0;
+ int nid;
- if (!mutex_trylock(&dma_pmd_pools_lock))
- return 0;
+ for_each_node_state(nid, N_MEMORY)
+ nr += READ_ONCE(dma_pmd_reservoirs[nid].nr_pages);
- list_for_each_entry(pool, &dma_pmd_pools, node)
- nr += READ_ONCE(pool->num_idle_pages);
+ if (mutex_trylock(&dma_pmd_pools_lock)) {
+ list_for_each_entry(pool, &dma_pmd_pools, node)
+ nr += READ_ONCE(pool->num_idle_pages);
- mutex_unlock(&dma_pmd_pools_lock);
+ mutex_unlock(&dma_pmd_pools_lock);
+ }
return nr << PMD_ORDER;
}
@@ -1062,34 +1285,50 @@ static unsigned long dma_pmd_shrink_scan(struct shrinker *shrink,
unsigned long freed = 0;
LIST_HEAD(release_list);
unsigned long flags;
- int srcu_idx;
+ int srcu_idx, nid;
- if (!mutex_trylock(&dma_pmd_pools_lock))
- return SHRINK_STOP;
+ for_each_node_state(nid, N_MEMORY) {
+ struct dma_pmd_reservoir *res = &dma_pmd_reservoirs[nid];
- srcu_idx = srcu_read_lock(&dma_pmd_srcu);
- list_for_each_entry(pool, &dma_pmd_pools, node) {
- /*
- * Every PMD page on @idle is a candidate, so this detaches one
- * per iteration and stops as soon as the reclaim path has
- * what it asked for: the IRQs-off section is bounded by the
- * work done, not by how many PMD pages the pool holds.
- */
- spin_lock_irqsave(&pool->lock, flags);
- list_for_each_entry_safe(meta, tmp, &pool->idle, list) {
+ if (!READ_ONCE(res->nr_pages))
+ continue;
+ spin_lock_irqsave(&res->lock, flags);
+ list_for_each_entry_safe(meta, tmp, &res->pages, list) {
if (freed >= sc->nr_to_scan)
break;
list_move(&meta->list, &release_list);
- pool->num_idle_pages--;
+ res->nr_pages--;
freed += 1UL << PMD_ORDER;
}
- spin_unlock_irqrestore(&pool->lock, flags);
-
+ spin_unlock_irqrestore(&res->lock, flags);
if (freed >= sc->nr_to_scan)
break;
}
- mutex_unlock(&dma_pmd_pools_lock);
+ srcu_idx = srcu_read_lock(&dma_pmd_srcu);
+ if (freed < sc->nr_to_scan && mutex_trylock(&dma_pmd_pools_lock)) {
+ list_for_each_entry(pool, &dma_pmd_pools, node) {
+ /*
+ * Every PMD page on @idle is a candidate, so this detaches one
+ * per iteration and stops as soon as the reclaim path has
+ * what it asked for: the IRQs-off section is bounded by the
+ * work done, not by how many PMD pages the pool holds.
+ */
+ spin_lock_irqsave(&pool->lock, flags);
+ list_for_each_entry_safe(meta, tmp, &pool->idle, list) {
+ if (freed >= sc->nr_to_scan)
+ break;
+ list_move(&meta->list, &release_list);
+ pool->num_idle_pages--;
+ freed += 1UL << PMD_ORDER;
+ }
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+ if (freed >= sc->nr_to_scan)
+ break;
+ }
+ mutex_unlock(&dma_pmd_pools_lock);
+ }
/*
* Retiring is done outside both locks: dma_pmd_release_page() unmaps
@@ -1103,8 +1342,10 @@ static unsigned long dma_pmd_shrink_scan(struct shrinker *shrink,
if (meta->flags & (DMA_PMD_DECRYPTED | DMA_PMD_PINNED)) {
llist_add(&meta->llnode, &dma_pmd_free_list);
dma_pmd_schedule_reclaim();
- } else {
+ } else if (meta->pool) {
dma_pmd_release_page(meta);
+ } else {
+ dma_pmd_free_reservoir_page(meta);
}
}
srcu_read_unlock(&dma_pmd_srcu, srcu_idx);
@@ -1128,6 +1369,12 @@ static unsigned long dma_pmd_shrink_scan(struct shrinker *shrink,
static int __init dma_pmd_init(void)
{
struct shrinker *shrink;
+ int nid;
+
+ for (nid = 0; nid < MAX_NUMNODES; nid++) {
+ spin_lock_init(&dma_pmd_reservoirs[nid].lock);
+ INIT_LIST_HEAD(&dma_pmd_reservoirs[nid].pages);
+ }
/*
* Hard ceiling on memory diverted from the buddy allocator into 2MB
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 17/22] net/gve: Use DMA_PMD memory for RX buffers
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (15 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 16/22] iommu/dma: Add per-NUMA-node PMD page reservoir Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 18/22] net/gve: Use DMA_PMD memory for tx header bounce buffers Luigi Rizzo
` (4 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
When dev->dma_pmd_rxbuf is enabled, back GQI-RDA RX packet buffers with
a per-queue DMA_PMD page pool (rx->dma_pmd_pool); DQO-RDA uses page_pool,
which checks dev->dma_pmd_rxbuf directly.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/net/ethernet/google/gve/gve.h | 2 +
drivers/net/ethernet/google/gve/gve_main.c | 5 ++
drivers/net/ethernet/google/gve/gve_rx.c | 61 ++++++++++++++++------
3 files changed, 53 insertions(+), 15 deletions(-)
diff --git a/drivers/net/ethernet/google/gve/gve.h b/drivers/net/ethernet/google/gve/gve.h
index c280ff35ee771..11330404f681b 100644
--- a/drivers/net/ethernet/google/gve/gve.h
+++ b/drivers/net/ethernet/google/gve/gve.h
@@ -10,6 +10,7 @@
#include <linux/dma-mapping.h>
#include <linux/dmapool.h>
#include <linux/ethtool_netlink.h>
+#include <linux/dma-pmd.h>
#include <linux/netdevice.h>
#include <linux/net_tstamp.h>
#include <linux/pci.h>
@@ -258,6 +259,7 @@ struct gve_rx_ring {
u32 qpl_copy_pool_mask;
u32 qpl_copy_pool_head;
struct gve_rx_slot_page_info *qpl_copy_pool;
+ struct dma_pmd_pool *dma_pmd_pool;
};
/* DQO fields. */
diff --git a/drivers/net/ethernet/google/gve/gve_main.c b/drivers/net/ethernet/google/gve/gve_main.c
index 9cc343a162712..e5631987b050f 100644
--- a/drivers/net/ethernet/google/gve/gve_main.c
+++ b/drivers/net/ethernet/google/gve/gve_main.c
@@ -1055,6 +1055,11 @@ static int gve_queues_mem_alloc(struct gve_priv *priv,
if (err)
goto free_tx;
+ if (rx_alloc_cfg->raw_addressing && priv->pdev->dev.dma_pmd_rxbuf)
+ dev_info(&priv->pdev->dev,
+ "DMA_PMD rx_bufs pool: enabled on %u RX queue(s)\n",
+ rx_alloc_cfg->qcfg_rx->num_queues);
+
return 0;
free_tx:
diff --git a/drivers/net/ethernet/google/gve/gve_rx.c b/drivers/net/ethernet/google/gve/gve_rx.c
index 81ea800e66e94..45b5e841efcd3 100644
--- a/drivers/net/ethernet/google/gve/gve_rx.c
+++ b/drivers/net/ethernet/google/gve/gve_rx.c
@@ -121,6 +121,7 @@ void gve_rx_free_ring_gqi(struct gve_priv *priv, struct gve_rx_ring *rx,
}
gve_rx_unfill_pages(priv, rx, cfg);
+ rx->dma_pmd_pool = dma_pmd_pool_destroy(rx->dma_pmd_pool);
if (rx->data.data_ring) {
bytes = sizeof(*rx->data.data_ring) * slots;
@@ -159,14 +160,29 @@ static void gve_setup_rx_buffer(struct gve_rx_ring *rx,
static int gve_rx_alloc_buffer(struct gve_priv *priv, struct device *dev,
struct gve_rx_slot_page_info *page_info,
union gve_rx_data_slot *data_slot,
- struct gve_rx_ring *rx)
+ struct gve_rx_ring *rx, gfp_t gfp)
{
- struct page *page;
+ struct page *page = NULL;
dma_addr_t dma;
- int err;
+ int err = 0;
- err = gve_alloc_page(priv, dev, &page, &dma, DMA_FROM_DEVICE,
- GFP_ATOMIC);
+ if (rx->dma_pmd_pool) {
+ page = dma_pmd_pool_alloc_node(rx->dma_pmd_pool,
+ gfp | __GFP_NOWARN, priv->numa_node);
+ if (page) {
+ dma = dma_map_page(dev, page, 0, PAGE_SIZE,
+ DMA_FROM_DEVICE);
+ if (dma_mapping_error(dev, dma)) {
+ priv->dma_mapping_error++;
+ put_page(page);
+ page = NULL;
+ err = -ENOMEM;
+ }
+ }
+ }
+ if (!page && !err)
+ err = gve_alloc_page(priv, dev, &page, &dma, DMA_FROM_DEVICE,
+ gfp);
if (err) {
u64_stats_update_begin(&rx->statss);
rx->rx_buf_alloc_fail++;
@@ -209,7 +225,8 @@ static int gve_rx_prefill_pages(struct gve_rx_ring *rx,
}
err = gve_rx_alloc_buffer(priv, &priv->pdev->dev,
&rx->data.page_info[i],
- &rx->data.data_ring[i], rx);
+ &rx->data.data_ring[i], rx,
+ GFP_KERNEL);
if (err)
goto alloc_err_rda;
}
@@ -319,6 +336,8 @@ int gve_rx_alloc_ring_gqi(struct gve_priv *priv,
err = -ENOMEM;
goto abort_with_copy_pool;
}
+ } else if (hdev->dma_pmd_rxbuf) {
+ rx->dma_pmd_pool = dma_pmd_pool_create(0, 8);
}
filled_pages = gve_rx_prefill_pages(rx, cfg);
@@ -365,6 +384,7 @@ int gve_rx_alloc_ring_gqi(struct gve_priv *priv,
abort_filled:
gve_rx_unfill_pages(priv, rx, cfg);
abort_with_qpl:
+ rx->dma_pmd_pool = dma_pmd_pool_destroy(rx->dma_pmd_pool);
if (!rx->data.raw_addressing) {
gve_free_queue_page_list(priv, rx->data.qpl, qpl_id);
rx->data.qpl = NULL;
@@ -495,13 +515,19 @@ static void gve_rx_flip_buff(struct gve_rx_slot_page_info *page_info, __be64 *sl
*(slot_addr) ^= offset;
}
-static int gve_rx_can_recycle_buffer(struct gve_rx_slot_page_info *page_info)
+static int gve_rx_can_recycle_buffer(struct gve_rx_ring *rx,
+ struct gve_rx_slot_page_info *page_info)
{
int pagecount = page_count(page_info->page);
/* This page is not being used by any SKBs - reuse */
- if (pagecount == page_info->pagecnt_bias)
+ if (pagecount == page_info->pagecnt_bias) {
+ if (rx && rx->dma_pmd_pool &&
+ !dma_is_pmd_page(page_to_pfn(page_info->page)) &&
+ dma_pmd_pool_has_free(rx->dma_pmd_pool))
+ return 0;
return 1;
+ }
/* This page is still being used by an SKB - we can't reuse */
else if (pagecount > page_info->pagecnt_bias)
return 0;
@@ -545,7 +571,7 @@ static struct sk_buff *gve_rx_copy_to_pool(struct gve_rx_ring *rx,
copy_page_info = &rx->qpl_copy_pool[pool_idx];
if (!copy_page_info->can_flip) {
- int recycle = gve_rx_can_recycle_buffer(copy_page_info);
+ int recycle = gve_rx_can_recycle_buffer(NULL, copy_page_info);
if (unlikely(recycle < 0)) {
gve_schedule_reset(rx->gve);
@@ -665,7 +691,7 @@ static struct sk_buff *gve_rx_skb(struct gve_priv *priv, struct gve_rx_ring *rx,
u64_stats_update_end(&rx->statss);
}
} else {
- int recycle = gve_rx_can_recycle_buffer(page_info);
+ int recycle = gve_rx_can_recycle_buffer(rx, page_info);
if (unlikely(recycle < 0)) {
gve_schedule_reset(priv);
@@ -971,7 +997,7 @@ static bool gve_rx_refill_buffers(struct gve_priv *priv, struct gve_rx_ring *rx)
* owns half the page it is impossible to tell which half. Either
* the whole page is free or it needs to be replaced.
*/
- int recycle = gve_rx_can_recycle_buffer(page_info);
+ int recycle = gve_rx_can_recycle_buffer(rx, page_info);
if (recycle < 0) {
if (!rx->data.raw_addressing)
@@ -983,11 +1009,16 @@ static bool gve_rx_refill_buffers(struct gve_priv *priv, struct gve_rx_ring *rx)
union gve_rx_data_slot *data_slot =
&rx->data.data_ring[idx];
struct device *dev = &priv->pdev->dev;
- gve_rx_free_buffer(dev, page_info, data_slot);
- page_info->page = NULL;
+ struct gve_rx_slot_page_info old_info = *page_info;
+ union gve_rx_data_slot old_slot = *data_slot;
+
if (gve_rx_alloc_buffer(priv, dev, page_info,
- data_slot, rx)) {
- break;
+ data_slot, rx, GFP_ATOMIC)) {
+ if (page_count(page_info->page) >
+ page_info->pagecnt_bias)
+ break;
+ } else {
+ gve_rx_free_buffer(dev, &old_info, &old_slot);
}
}
}
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 18/22] net/gve: Use DMA_PMD memory for tx header bounce buffers
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (16 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 17/22] net/gve: Use DMA_PMD memory for RX buffers Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 19/22] net/mlx5e: " Luigi Rizzo
` (3 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
When dev->dma_pmd_tx_hdrs is enabled, allocate a MAX_TCP_HEADER
bounce buffer per TX slot/packet out of the DMA_PMD arena in raw-addressing
mode (both GQI and DQO). When skb_headlen(skb) <= MAX_TCP_HEADER, the
linear portion is copied into the pre-mapped bounce buffer instead of
calling dma_map_single()/dma_unmap_single() on every packet.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/net/ethernet/google/gve/gve.h | 3 ++
drivers/net/ethernet/google/gve/gve_tx.c | 38 ++++++++++++++---
drivers/net/ethernet/google/gve/gve_tx_dqo.c | 45 +++++++++++++++-----
3 files changed, 69 insertions(+), 17 deletions(-)
diff --git a/drivers/net/ethernet/google/gve/gve.h b/drivers/net/ethernet/google/gve/gve.h
index 11330404f681b..00e438fe05a50 100644
--- a/drivers/net/ethernet/google/gve/gve.h
+++ b/drivers/net/ethernet/google/gve/gve.h
@@ -18,6 +18,7 @@
#include <linux/ptp_clock_kernel.h>
#include <linux/u64_stats_sync.h>
#include <net/page_pool/helpers.h>
+#include <net/tcp.h>
#include <net/xdp.h>
#include "gve_desc.h"
@@ -642,6 +643,8 @@ struct gve_tx_ring {
};
} dqo;
} ____cacheline_aligned;
+ void *tx_hdr_bufs;
+ dma_addr_t tx_hdr_bufs_dma;
struct netdev_queue *netdev_txq;
struct gve_queue_resources *q_resources; /* head and tail pointer idx */
struct device *dev;
diff --git a/drivers/net/ethernet/google/gve/gve_tx.c b/drivers/net/ethernet/google/gve/gve_tx.c
index 79338fd9bf66e..9610b54ef4792 100644
--- a/drivers/net/ethernet/google/gve/gve_tx.c
+++ b/drivers/net/ethernet/google/gve/gve_tx.c
@@ -224,6 +224,11 @@ static void gve_tx_free_ring_gqi(struct gve_priv *priv, struct gve_tx_ring *tx,
u32 slots;
slots = tx->mask + 1;
+ if (tx->tx_hdr_bufs) {
+ dma_free_coherent(hdev, slots * MAX_TCP_HEADER,
+ tx->tx_hdr_bufs, tx->tx_hdr_bufs_dma);
+ tx->tx_hdr_bufs = NULL;
+ }
dma_free_coherent(hdev, sizeof(*tx->q_resources),
tx->q_resources, tx->q_resources_bus);
tx->q_resources = NULL;
@@ -298,6 +303,11 @@ static int gve_tx_alloc_ring_gqi(struct gve_priv *priv,
/* map Tx FIFO */
if (gve_tx_fifo_init(priv, &tx->tx_fifo))
goto abort_with_qpl;
+ } else if (hdev->dma_pmd_tx_hdrs) {
+ tx->tx_hdr_bufs = dma_alloc_coherent(hdev,
+ cfg->ring_size * MAX_TCP_HEADER,
+ &tx->tx_hdr_bufs_dma,
+ GFP_KERNEL);
}
tx->q_resources =
@@ -311,6 +321,12 @@ static int gve_tx_alloc_ring_gqi(struct gve_priv *priv,
return 0;
abort_with_fifo:
+ if (tx->tx_hdr_bufs) {
+ dma_free_coherent(hdev,
+ cfg->ring_size * MAX_TCP_HEADER,
+ tx->tx_hdr_bufs, tx->tx_hdr_bufs_dma);
+ tx->tx_hdr_bufs = NULL;
+ }
if (!tx->raw_addressing)
gve_tx_fifo_release(priv, &tx->tx_fifo);
abort_with_qpl:
@@ -423,6 +439,8 @@ static inline int gve_skb_fifo_bytes_required(struct gve_tx_ring *tx,
#define MAX_TX_DESC_NEEDED (MAX_SKB_FRAGS + 4)
static void gve_tx_unmap_buf(struct device *dev, struct gve_tx_buffer_state *info)
{
+ if (!dma_unmap_len(info, len))
+ return;
if (info->skb) {
dma_unmap_single(dev, dma_unmap_addr(info, dma),
dma_unmap_len(info, len),
@@ -657,13 +675,21 @@ static int gve_tx_add_skb_no_copy(struct gve_priv *priv, struct gve_tx_ring *tx,
info->skb = skb;
- addr = dma_map_single(tx->dev, skb->data, len, DMA_TO_DEVICE);
- if (unlikely(dma_mapping_error(tx->dev, addr))) {
- tx->dma_mapping_error++;
- goto drop;
+ if (tx->tx_hdr_bufs && len <= MAX_TCP_HEADER) {
+ u32 hdr_off = idx * MAX_TCP_HEADER;
+
+ memcpy(tx->tx_hdr_bufs + hdr_off, skb->data, len);
+ addr = tx->tx_hdr_bufs_dma + hdr_off;
+ dma_unmap_len_set(info, len, 0);
+ } else {
+ addr = dma_map_single(tx->dev, skb->data, len, DMA_TO_DEVICE);
+ if (unlikely(dma_mapping_error(tx->dev, addr))) {
+ tx->dma_mapping_error++;
+ goto drop;
+ }
+ dma_unmap_len_set(info, len, len);
+ dma_unmap_addr_set(info, dma, addr);
}
- dma_unmap_len_set(info, len, len);
- dma_unmap_addr_set(info, dma, addr);
num_descriptors = 1 + shinfo->nr_frags;
if (hlen < len)
diff --git a/drivers/net/ethernet/google/gve/gve_tx_dqo.c b/drivers/net/ethernet/google/gve/gve_tx_dqo.c
index ad99cbb77e291..ac69793492c9f 100644
--- a/drivers/net/ethernet/google/gve/gve_tx_dqo.c
+++ b/drivers/net/ethernet/google/gve/gve_tx_dqo.c
@@ -176,8 +176,9 @@ static void gve_unmap_packet(struct device *dev,
return;
/* SKB linear portion is guaranteed to be mapped */
- dma_unmap_single(dev, dma_unmap_addr(pkt, dma[0]),
- dma_unmap_len(pkt, len[0]), DMA_TO_DEVICE);
+ if (dma_unmap_len(pkt, len[0]))
+ dma_unmap_single(dev, dma_unmap_addr(pkt, dma[0]),
+ dma_unmap_len(pkt, len[0]), DMA_TO_DEVICE);
for (i = 1; i < pkt->num_bufs; i++) {
netmem_dma_unmap_page_attrs(dev, dma_unmap_addr(pkt, dma[i]),
dma_unmap_len(pkt, len[i]),
@@ -232,6 +233,13 @@ static void gve_tx_free_ring_dqo(struct gve_priv *priv, struct gve_tx_ring *tx,
size_t bytes;
u32 qpl_id;
+ if (tx->tx_hdr_bufs) {
+ dma_free_coherent(hdev,
+ tx->dqo.num_pending_packets * MAX_TCP_HEADER,
+ tx->tx_hdr_bufs, tx->tx_hdr_bufs_dma);
+ tx->tx_hdr_bufs = NULL;
+ }
+
if (tx->q_resources) {
dma_free_coherent(hdev, sizeof(*tx->q_resources),
tx->q_resources, tx->q_resources_bus);
@@ -399,6 +407,12 @@ static int gve_tx_alloc_ring_dqo(struct gve_priv *priv,
if (gve_tx_qpl_buf_init(tx))
goto err;
+ } else if (hdev->dma_pmd_tx_hdrs) {
+ tx->tx_hdr_bufs =
+ dma_alloc_coherent(hdev,
+ tx->dqo.num_pending_packets * MAX_TCP_HEADER,
+ &tx->tx_hdr_bufs_dma,
+ GFP_KERNEL);
}
return 0;
@@ -712,12 +726,20 @@ static int gve_tx_add_skb_no_copy_dqo(struct gve_tx_ring *tx,
u32 len = skb_headlen(skb);
dma_addr_t addr;
- addr = dma_map_single(tx->dev, skb->data, len, DMA_TO_DEVICE);
- if (unlikely(dma_mapping_error(tx->dev, addr)))
- goto err;
+ if (tx->tx_hdr_bufs && len <= MAX_TCP_HEADER) {
+ u32 hdr_off = (u32)completion_tag * MAX_TCP_HEADER;
- dma_unmap_len_set(pkt, len[pkt->num_bufs], len);
- dma_unmap_addr_set(pkt, dma[pkt->num_bufs], addr);
+ memcpy(tx->tx_hdr_bufs + hdr_off, skb->data, len);
+ addr = tx->tx_hdr_bufs_dma + hdr_off;
+ dma_unmap_len_set(pkt, len[pkt->num_bufs], 0);
+ } else {
+ addr = dma_map_single(tx->dev, skb->data, len,
+ DMA_TO_DEVICE);
+ if (unlikely(dma_mapping_error(tx->dev, addr)))
+ goto err;
+ dma_unmap_len_set(pkt, len[pkt->num_bufs], len);
+ dma_unmap_addr_set(pkt, dma[pkt->num_bufs], addr);
+ }
++pkt->num_bufs;
gve_tx_fill_pkt_desc_dqo(tx, desc_idx, enable_csum, len, addr,
@@ -748,10 +770,11 @@ static int gve_tx_add_skb_no_copy_dqo(struct gve_tx_ring *tx,
err:
for (i = 0; i < pkt->num_bufs; i++) {
if (i == 0) {
- dma_unmap_single(tx->dev,
- dma_unmap_addr(pkt, dma[i]),
- dma_unmap_len(pkt, len[i]),
- DMA_TO_DEVICE);
+ if (dma_unmap_len(pkt, len[i]))
+ dma_unmap_single(tx->dev,
+ dma_unmap_addr(pkt, dma[i]),
+ dma_unmap_len(pkt, len[i]),
+ DMA_TO_DEVICE);
} else {
dma_unmap_page(tx->dev,
dma_unmap_addr(pkt, dma[i]),
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 19/22] net/mlx5e: Use DMA_PMD memory for tx header bounce buffers
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (17 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 18/22] net/gve: Use DMA_PMD memory for tx header bounce buffers Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 20/22] net/idpf: " Luigi Rizzo
` (2 subsequent siblings)
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
When dev->dma_pmd_tx_hdrs is enabled, allocate a MAX_TCP_HEADER
bounce buffer per TX DMA FIFO slot out of the DMA_PMD arena. When the
linear portion of a transmitted packet is at most MAX_TCP_HEADER bytes,
copy it into the pre-mapped bounce buffer instead of calling
dma_map_single()/dma_unmap_single() on every packet.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/net/ethernet/mellanox/mlx5/core/en.h | 2 ++
.../net/ethernet/mellanox/mlx5/core/en/txrx.h | 4 ++-
.../net/ethernet/mellanox/mlx5/core/en_main.c | 14 ++++++++++
.../net/ethernet/mellanox/mlx5/core/en_tx.c | 28 +++++++++++++++----
4 files changed, 42 insertions(+), 6 deletions(-)
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en.h b/drivers/net/ethernet/mellanox/mlx5/core/en.h
index a3945195f454e..d762ad1c1ffaf 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en.h
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en.h
@@ -442,6 +442,8 @@ struct mlx5e_txqsq {
struct mlx5e_sq_dma *dma_fifo;
struct mlx5e_skb_fifo skb_fifo;
struct mlx5e_tx_wqe_info *wqe_info;
+ void *tx_hdr_bufs;
+ dma_addr_t tx_hdr_bufs_dma;
} db;
void __iomem *uar_map;
struct netdev_queue *txq;
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en/txrx.h b/drivers/net/ethernet/mellanox/mlx5/core/en/txrx.h
index f2a8453d8dce0..6de7a2a430d72 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en/txrx.h
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en/txrx.h
@@ -365,7 +365,9 @@ mlx5e_tx_dma_unmap(struct device *pdev, struct mlx5e_sq_dma *dma)
{
switch (dma->type) {
case MLX5E_DMA_MAP_SINGLE:
- dma_unmap_single(pdev, dma->addr, dma->size, DMA_TO_DEVICE);
+ if (dma->size)
+ dma_unmap_single(pdev, dma->addr, dma->size,
+ DMA_TO_DEVICE);
break;
case MLX5E_DMA_MAP_PAGE:
netmem_dma_unmap_page_attrs(pdev, dma->addr, dma->size,
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
index 4c6060d54fbed..efd766aa43a02 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c
@@ -1666,6 +1666,13 @@ static void mlx5e_free_icosq(struct mlx5e_icosq *sq)
void mlx5e_free_txqsq_db(struct mlx5e_txqsq *sq)
{
+ if (sq->db.tx_hdr_bufs) {
+ int wq_sz = mlx5_wq_cyc_get_size(&sq->wq);
+ int df_sz = wq_sz * MLX5_SEND_WQEBB_NUM_DS;
+
+ dma_free_coherent(sq->pdev, df_sz * MAX_TCP_HEADER,
+ sq->db.tx_hdr_bufs, sq->db.tx_hdr_bufs_dma);
+ }
kvfree(sq->db.wqe_info);
kvfree(sq->db.skb_fifo.fifo);
kvfree(sq->db.dma_fifo);
@@ -1690,6 +1697,13 @@ int mlx5e_alloc_txqsq_db(struct mlx5e_txqsq *sq, int numa)
return -ENOMEM;
}
+ if (sq->pdev->dma_pmd_tx_hdrs)
+ sq->db.tx_hdr_bufs =
+ dma_alloc_coherent(sq->pdev,
+ df_sz * MAX_TCP_HEADER,
+ &sq->db.tx_hdr_bufs_dma,
+ GFP_KERNEL | __GFP_NOWARN);
+
sq->dma_fifo_mask = df_sz - 1;
sq->db.skb_fifo.pc = &sq->skb_fifo_pc;
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_tx.c b/drivers/net/ethernet/mellanox/mlx5/core/en_tx.c
index 0b5e600e4a6a9..3b6a7a8b0128e 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en_tx.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en_tx.c
@@ -177,6 +177,27 @@ mlx5e_tx_get_gso_ihs(struct mlx5e_txqsq *sq, struct sk_buff *skb)
return ihs;
}
+static inline dma_addr_t
+mlx5e_tx_map_hdr(struct mlx5e_txqsq *sq, unsigned char *data, u32 len)
+{
+ dma_addr_t dma_addr;
+
+ if (sq->db.tx_hdr_bufs && len <= MAX_TCP_HEADER) {
+ u32 off = (sq->dma_fifo_pc & sq->dma_fifo_mask) *
+ MAX_TCP_HEADER;
+
+ memcpy(sq->db.tx_hdr_bufs + off, data, len);
+ dma_addr = sq->db.tx_hdr_bufs_dma + off;
+ mlx5e_dma_push_single(sq, dma_addr, 0);
+ return dma_addr;
+ }
+
+ dma_addr = dma_map_single(sq->pdev, data, len, DMA_TO_DEVICE);
+ if (likely(!dma_mapping_error(sq->pdev, dma_addr)))
+ mlx5e_dma_push_single(sq, dma_addr, len);
+ return dma_addr;
+}
+
static inline int
mlx5e_txwqe_build_dsegs(struct mlx5e_txqsq *sq, struct sk_buff *skb,
unsigned char *skb_data, u16 headlen,
@@ -187,8 +208,7 @@ mlx5e_txwqe_build_dsegs(struct mlx5e_txqsq *sq, struct sk_buff *skb,
int i;
if (headlen) {
- dma_addr = dma_map_single(sq->pdev, skb_data, headlen,
- DMA_TO_DEVICE);
+ dma_addr = mlx5e_tx_map_hdr(sq, skb_data, headlen);
if (unlikely(dma_mapping_error(sq->pdev, dma_addr)))
goto dma_unmap_wqe_err;
@@ -196,7 +216,6 @@ mlx5e_txwqe_build_dsegs(struct mlx5e_txqsq *sq, struct sk_buff *skb,
dseg->lkey = sq->mkey_be;
dseg->byte_count = cpu_to_be32(headlen);
- mlx5e_dma_push_single(sq, dma_addr, headlen);
num_dma++;
dseg++;
}
@@ -578,7 +597,7 @@ mlx5e_sq_xmit_mpwqe(struct mlx5e_txqsq *sq, struct sk_buff *skb,
txd.data = skb->data;
txd.len = skb->len;
- txd.dma_addr = dma_map_single(sq->pdev, txd.data, txd.len, DMA_TO_DEVICE);
+ txd.dma_addr = mlx5e_tx_map_hdr(sq, txd.data, txd.len);
if (unlikely(dma_mapping_error(sq->pdev, txd.dma_addr)))
goto err_unmap;
@@ -591,7 +610,6 @@ mlx5e_sq_xmit_mpwqe(struct mlx5e_txqsq *sq, struct sk_buff *skb,
sq->stats->xmit_more += xmit_more;
- mlx5e_dma_push_single(sq, txd.dma_addr, txd.len);
mlx5e_skb_fifo_push(&sq->db.skb_fifo, skb);
mlx5e_tx_mpwqe_add_dseg(sq, &txd);
mlx5e_tx_skb_update_ts_flags(skb);
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 20/22] net/idpf: Use DMA_PMD memory for tx header bounce buffers
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (18 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 19/22] net/mlx5e: " Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 21/22] net/bnxt: " Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 22/22] iommu/dma: Add DMA_PMD statistics and debugfs Luigi Rizzo
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
When dev->dma_pmd_tx_hdrs is enabled, allocate a MAX_TCP_HEADER
bounce buffer per TX buffer slot out of the DMA_PMD arena. When the
linear portion of a transmitted packet is at most MAX_TCP_HEADER bytes,
copy it into the pre-mapped bounce buffer and set the SQE unmap length
to 0 instead of calling dma_map_single()/dma_unmap_page() on every
packet.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
.../ethernet/intel/idpf/idpf_singleq_txrx.c | 6 ++---
drivers/net/ethernet/intel/idpf/idpf_txrx.c | 21 ++++++++++++---
drivers/net/ethernet/intel/idpf/idpf_txrx.h | 27 ++++++++++++++++++-
include/net/libeth/tx.h | 5 ++--
4 files changed, 50 insertions(+), 9 deletions(-)
diff --git a/drivers/net/ethernet/intel/idpf/idpf_singleq_txrx.c b/drivers/net/ethernet/intel/idpf/idpf_singleq_txrx.c
index e3ddf18dcbf51..24a2dfec6212d 100644
--- a/drivers/net/ethernet/intel/idpf/idpf_singleq_txrx.c
+++ b/drivers/net/ethernet/intel/idpf/idpf_singleq_txrx.c
@@ -261,7 +261,7 @@ static void idpf_tx_singleq_map(struct idpf_tx_queue *tx_q,
tx_desc = &tx_q->base_tx[i];
- dma = dma_map_single(tx_q->dev, skb->data, size, DMA_TO_DEVICE);
+ dma = idpf_tx_map_hdr(tx_q, skb, first, size);
/* write each descriptor with CRC bit */
if (idpf_queue_has(CRC_EN, tx_q))
@@ -274,8 +274,7 @@ static void idpf_tx_singleq_map(struct idpf_tx_queue *tx_q,
return idpf_tx_singleq_dma_map_error(tx_q, skb,
first, i);
- /* record length, and DMA address */
- dma_unmap_len_set(tx_buf, len, size);
+ /* record DMA address */
dma_unmap_addr_set(tx_buf, dma, dma);
tx_buf->type = LIBETH_SQE_FRAG;
@@ -326,6 +325,7 @@ static void idpf_tx_singleq_map(struct idpf_tx_queue *tx_q,
size = skb_frag_size(frag);
data_len -= size;
+ dma_unmap_len_set(tx_buf, len, size);
dma = skb_frag_dma_map(tx_q->dev, frag, 0, size,
DMA_TO_DEVICE);
diff --git a/drivers/net/ethernet/intel/idpf/idpf_txrx.c b/drivers/net/ethernet/intel/idpf/idpf_txrx.c
index 4311ffa30bb18..4bdec468ea2ce 100644
--- a/drivers/net/ethernet/intel/idpf/idpf_txrx.c
+++ b/drivers/net/ethernet/intel/idpf/idpf_txrx.c
@@ -90,6 +90,14 @@ static void idpf_tx_buf_rel_all(struct idpf_tx_queue *txq)
else
idpf_tx_buf_clean(txq);
+ if (txq->hdr_buf_va) {
+ dma_free_coherent(txq->dev,
+ txq->buf_pool_size * MAX_TCP_HEADER,
+ txq->hdr_buf_va, txq->hdr_buf_dma);
+ txq->hdr_buf_va = NULL;
+ txq->hdr_buf_dma = 0;
+ }
+
kfree(txq->tx_buf);
txq->tx_buf = NULL;
}
@@ -187,6 +195,13 @@ static int idpf_tx_buf_alloc_all(struct idpf_tx_queue *tx_q)
if (!tx_q->tx_buf)
return -ENOMEM;
+ if (tx_q->dev->dma_pmd_tx_hdrs)
+ tx_q->hdr_buf_va =
+ dma_alloc_coherent(tx_q->dev,
+ tx_q->buf_pool_size * MAX_TCP_HEADER,
+ &tx_q->hdr_buf_dma,
+ GFP_KERNEL | __GFP_NOWARN);
+
return 0;
}
@@ -2663,7 +2678,7 @@ static void idpf_tx_splitq_map(struct idpf_tx_queue *tx_q,
tx_desc = &tx_q->flex_tx[i];
- dma = dma_map_single(tx_q->dev, skb->data, size, DMA_TO_DEVICE);
+ dma = idpf_tx_map_hdr(tx_q, skb, first, size);
tx_buf = first;
first->nr_frags = 0;
@@ -2680,8 +2695,7 @@ static void idpf_tx_splitq_map(struct idpf_tx_queue *tx_q,
first->nr_frags++;
tx_buf->type = LIBETH_SQE_FRAG;
- /* record length, and DMA address */
- dma_unmap_len_set(tx_buf, len, size);
+ /* record DMA address */
dma_unmap_addr_set(tx_buf, dma, dma);
/* buf_addr is in same location for both desc types */
@@ -2784,6 +2798,7 @@ static void idpf_tx_splitq_map(struct idpf_tx_queue *tx_q,
size = skb_frag_size(frag);
data_len -= size;
+ dma_unmap_len_set(tx_buf, len, size);
dma = skb_frag_dma_map(tx_q->dev, frag, 0, size,
DMA_TO_DEVICE);
diff --git a/drivers/net/ethernet/intel/idpf/idpf_txrx.h b/drivers/net/ethernet/intel/idpf/idpf_txrx.h
index 93547597efd2a..39458a8e5baa0 100644
--- a/drivers/net/ethernet/intel/idpf/idpf_txrx.h
+++ b/drivers/net/ethernet/intel/idpf/idpf_txrx.h
@@ -8,6 +8,7 @@
#include <linux/net/intel/virtchnl2_lan_desc.h>
#include <net/libeth/cache.h>
+#include <net/libeth/tx.h>
#include <net/libeth/types.h>
#include <net/netdev_queues.h>
#include <net/tcp.h>
@@ -639,6 +640,8 @@ libeth_cacheline_set_assert(struct idpf_rx_queue,
* @q_vector: Backreference to associated vector
* @buf_pool_size: Total number of idpf_tx_buf
* @rel_q_id: relative virtchnl queue index
+ * @hdr_buf_va: Base of the TX header bounce buffer array, or NULL
+ * @hdr_buf_dma: DMA address matching @hdr_buf_va
*/
struct idpf_tx_queue {
__cacheline_group_begin_aligned(read_mostly);
@@ -715,6 +718,8 @@ struct idpf_tx_queue {
u32 buf_pool_size;
u32 rel_q_id;
+ void *hdr_buf_va;
+ dma_addr_t hdr_buf_dma;
__cacheline_group_end_aligned(cold);
};
libeth_cacheline_set_assert(struct idpf_tx_queue, 64,
@@ -723,7 +728,7 @@ libeth_cacheline_set_assert(struct idpf_tx_queue, 64,
offsetofend(struct idpf_tx_queue, timer) +
offsetof(struct idpf_tx_queue, q_stats) -
offsetofend(struct idpf_tx_queue, tstamp_task),
- 32);
+ 48);
/**
* struct idpf_buf_queue - software structure representing a buffer queue
@@ -1131,4 +1136,24 @@ int idpf_tso(struct sk_buff *skb, struct idpf_tx_offload_params *off);
void idpf_wait_for_sw_marker_completion(const struct idpf_tx_queue *txq);
+static inline dma_addr_t idpf_tx_map_hdr(struct idpf_tx_queue *tx_q,
+ const struct sk_buff *skb,
+ struct libeth_sqe *first,
+ unsigned int size)
+{
+ u32 buf_idx = first - tx_q->tx_buf;
+
+ if (tx_q->hdr_buf_va && buf_idx < tx_q->buf_pool_size &&
+ size <= MAX_TCP_HEADER) {
+ u32 offset = buf_idx * MAX_TCP_HEADER;
+
+ memcpy(tx_q->hdr_buf_va + offset, skb->data, size);
+ dma_unmap_len_set(first, len, 0);
+ return tx_q->hdr_buf_dma + offset;
+ }
+
+ dma_unmap_len_set(first, len, size);
+ return dma_map_single(tx_q->dev, skb->data, size, DMA_TO_DEVICE);
+}
+
#endif /* !_IDPF_TXRX_H_ */
diff --git a/include/net/libeth/tx.h b/include/net/libeth/tx.h
index c3db5c6f16410..2720ab9393e90 100644
--- a/include/net/libeth/tx.h
+++ b/include/net/libeth/tx.h
@@ -130,8 +130,9 @@ static inline void libeth_tx_complete(struct libeth_sqe *sqe,
case LIBETH_SQE_SKB:
case LIBETH_SQE_FRAG:
case LIBETH_SQE_SLAB:
- dma_unmap_page(cp->dev, dma_unmap_addr(sqe, dma),
- dma_unmap_len(sqe, len), DMA_TO_DEVICE);
+ if (dma_unmap_len(sqe, len))
+ dma_unmap_page(cp->dev, dma_unmap_addr(sqe, dma),
+ dma_unmap_len(sqe, len), DMA_TO_DEVICE);
break;
default:
break;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 21/22] net/bnxt: Use DMA_PMD memory for tx header bounce buffers
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (19 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 20/22] net/idpf: " Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 22/22] iommu/dma: Add DMA_PMD statistics and debugfs Luigi Rizzo
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
When dev->dma_pmd_tx_hdrs is enabled, allocate a MAX_TCP_HEADER
bounce buffer per TX ring slot out of the DMA_PMD arena. When the
linear portion of a transmitted packet is at most MAX_TCP_HEADER bytes,
copy it into the pre-mapped bounce buffer instead of calling
dma_map_single()/dma_unmap_single() on every packet.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/net/ethernet/broadcom/bnxt/bnxt.c | 45 ++++++++++++++++++-----
drivers/net/ethernet/broadcom/bnxt/bnxt.h | 2 +
2 files changed, 37 insertions(+), 10 deletions(-)
diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.c b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
index d7728d0c5b6e6..b9c91eb3aaddd 100644
--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.c
+++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.c
@@ -680,13 +680,22 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
length = BNXT_MIN_PKT_SIZE;
}
- mapping = dma_map_single(&pdev->dev, skb->data, len, DMA_TO_DEVICE);
-
- if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
- goto tx_free;
-
- dma_unmap_addr_set(tx_buf, mapping, mapping);
- dma_unmap_len_set(tx_buf, len, len);
+ if (txr->tx_hdr_bufs && len + pad <= MAX_TCP_HEADER) {
+ u32 off = RING_TX(bp, prod) * MAX_TCP_HEADER;
+
+ memcpy(txr->tx_hdr_bufs + off, skb->data, len);
+ if (pad && !last_frag)
+ memset(txr->tx_hdr_bufs + off + len, 0, pad);
+ mapping = txr->tx_hdr_bufs_dma + off;
+ dma_unmap_len_set(tx_buf, len, 0);
+ } else {
+ mapping = dma_map_single(&pdev->dev, skb->data, len,
+ DMA_TO_DEVICE);
+ if (unlikely(dma_mapping_error(&pdev->dev, mapping)))
+ goto tx_free;
+ dma_unmap_addr_set(tx_buf, mapping, mapping);
+ dma_unmap_len_set(tx_buf, len, len);
+ }
flags = (len << TX_BD_LEN_SHIFT) | TX_BD_TYPE_LONG_TX_BD |
TX_BD_CNT(last_frag + 2);
@@ -800,8 +809,9 @@ static netdev_tx_t bnxt_start_xmit(struct sk_buff *skb, struct net_device *dev)
/* start back at beginning and unmap skb */
prod = txr->tx_prod;
tx_buf = &txr->tx_buf_ring[RING_TX(bp, prod)];
- dma_unmap_single(&pdev->dev, dma_unmap_addr(tx_buf, mapping),
- skb_headlen(skb), DMA_TO_DEVICE);
+ if (dma_unmap_len(tx_buf, len))
+ dma_unmap_single(&pdev->dev, dma_unmap_addr(tx_buf, mapping),
+ skb_headlen(skb), DMA_TO_DEVICE);
prod = NEXT_TX(prod);
/* unmap remaining mapped pages */
@@ -885,7 +895,6 @@ static bool __bnxt_tx_int(struct bnxt *bp, struct bnxt_tx_ring_info *txr,
dma_unmap_single(&pdev->dev, dma_addr, dma_len,
DMA_TO_DEVICE);
}
-
last = tx_buf->nr_frags;
for (j = 0; j < last; j++) {
@@ -4111,6 +4120,15 @@ static void bnxt_free_tx_rings(struct bnxt *bp)
ring = &txr->tx_ring_struct;
+ if (txr->tx_hdr_bufs) {
+ dma_free_coherent(&pdev->dev,
+ ring->ring_mem.nr_pages *
+ TX_DESC_CNT * MAX_TCP_HEADER,
+ txr->tx_hdr_bufs,
+ txr->tx_hdr_bufs_dma);
+ txr->tx_hdr_bufs = NULL;
+ }
+
bnxt_free_ring(bp, &ring->ring_mem);
}
}
@@ -4181,6 +4199,13 @@ static int bnxt_alloc_tx_rings(struct bnxt *bp)
if (rc)
return rc;
}
+ if (pdev->dev.dma_pmd_tx_hdrs && i >= bp->tx_nr_rings_xdp)
+ txr->tx_hdr_bufs =
+ dma_alloc_coherent(&pdev->dev,
+ ring->ring_mem.nr_pages *
+ TX_DESC_CNT * MAX_TCP_HEADER,
+ &txr->tx_hdr_bufs_dma,
+ GFP_KERNEL | __GFP_NOWARN);
qidx = bp->tc_to_qidx[j];
ring->queue_id = bp->q_info[qidx].queue_id;
spin_lock_init(&txr->xdp_tx_lock);
diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt.h b/drivers/net/ethernet/broadcom/bnxt/bnxt.h
index c673b2ce4a0d2..b42bbabc0eb98 100644
--- a/drivers/net/ethernet/broadcom/bnxt/bnxt.h
+++ b/drivers/net/ethernet/broadcom/bnxt/bnxt.h
@@ -1003,6 +1003,8 @@ struct bnxt_tx_ring_info {
struct tx_push_buffer *tx_push;
dma_addr_t tx_push_mapping;
__le64 data_mapping;
+ void *tx_hdr_bufs;
+ dma_addr_t tx_hdr_bufs_dma;
void *tx_inline_buf;
dma_addr_t tx_inline_dma;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* [RFC: DMA_PMD 22/22] iommu/dma: Add DMA_PMD statistics and debugfs
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
` (20 preceding siblings ...)
2026-10-03 21:22 ` [RFC: DMA_PMD 21/22] net/bnxt: " Luigi Rizzo
@ 2026-10-03 21:22 ` Luigi Rizzo
21 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
Luigi Rizzo
DMA_PMD pools fail silently. dma_pmd_pool_alloc() returns NULL for four
distinct reasons - allocation backoff, the global DMA_PMD page ceiling, an
PMD_ORDER allocation failure, and a split_page_compound() rejection - and
every caller quietly falls back to a plain page, so a pool that has
stopped pooling looks exactly like one that is working. The map path is no
better: a buffer can be mapped individually because the domain has no
usable IOVA window, because the domain has no 2MB page size, or because
the PMD_ORDER page mapping itself failed, and none of that was observable
either. The only symptom is a throughput loss with no counter to point at.
Add a counter for each of those cases plus the number of 2M page mappings
established and dropped, and export them, along with the existing per-pool
statistics, through /sys/kernel/debug/dma_pmd/pools.
The new counters are atomic64_t rather than being folded into the existing
lock-protected statistics because their update sites cannot take
pool->lock: dma_pmd_add_page() runs outside it, and the map slow path
holds p2m->map_lock, so taking pool->lock there would invert the order
against dma_pmd_domain_release().
Also emit a one-shot warning on the first PMD_ORDER allocation failure.
That request carries __GFP_NOWARN, so the page allocator says nothing, and
this is the one failure mode that is worth a line in dmesg rather than a
counter somebody has to know to look at.
Signed-off-by: Luigi Rizzo <lrizzo@google.com>
---
drivers/iommu/dma-pmd-map.c | 16 +++-
drivers/iommu/dma-pmd-pool.c | 143 ++++++++++++++++++++++++++++++++---
drivers/iommu/dma-pmd-priv.h | 43 ++++++++++-
3 files changed, 186 insertions(+), 16 deletions(-)
diff --git a/drivers/iommu/dma-pmd-map.c b/drivers/iommu/dma-pmd-map.c
index 76cb5d01cb1e8..46c6ab5c1b7de 100644
--- a/drivers/iommu/dma-pmd-map.c
+++ b/drivers/iommu/dma-pmd-map.c
@@ -468,8 +468,13 @@ dma_addr_t dma_pmd_dma_map_phys(struct device *dev, struct iommu_domain *domain,
else
ret = dma_pmd_window_assign(dev, domain, win);
- if (ret)
+ if (ret) {
+ if (ret == -EOPNOTSUPP)
+ atomic64_inc(&meta->pool->fallback_nopool);
+ else
+ atomic64_inc(&meta->pool->fallback_nowindow);
return DMA_MAPPING_ERROR;
+ }
win_size = READ_ONCE(win->size);
}
@@ -482,8 +487,10 @@ dma_addr_t dma_pmd_dma_map_phys(struct device *dev, struct iommu_domain *domain,
* device mask.
*/
if (unlikely(iova - win->base + size > win_size ||
- iova + size - 1 > min_not_zero(dma_mask, dev->bus_dma_limit)))
+ iova + size - 1 > min_not_zero(dma_mask, dev->bus_dma_limit))) {
+ atomic64_inc(&meta->pool->fallback_nowindow);
return DMA_MAPPING_ERROR;
+ }
domain_idx = win->domain_idx - DMA_PMD_IDX_FIRST;
if (likely(test_bit(domain_idx, &meta->domains_mapped)))
@@ -510,12 +517,15 @@ dma_addr_t dma_pmd_dma_map_phys(struct device *dev, struct iommu_domain *domain,
*/
smp_mb__before_atomic();
set_bit(domain_idx, &meta->domains_mapped);
+ atomic64_inc(&meta->pool->pmd_map_cnt);
}
}
spin_unlock_irqrestore(&meta->map_lock, flags);
- if (unlikely(ret))
+ if (unlikely(ret)) {
+ atomic64_inc(&meta->pool->fallback_maperr);
return DMA_MAPPING_ERROR;
+ }
return iova;
}
diff --git a/drivers/iommu/dma-pmd-pool.c b/drivers/iommu/dma-pmd-pool.c
index d48c8e19741ab..a85aa550b500c 100644
--- a/drivers/iommu/dma-pmd-pool.c
+++ b/drivers/iommu/dma-pmd-pool.c
@@ -682,16 +682,21 @@ unsigned int dma_pmd_pools_forget_domain(int idx)
mutex_lock(&dma_pmd_pools_lock);
list_for_each_entry(pool, &dma_pmd_pools, node) {
+ unsigned int n = 0;
+
spin_lock_irqsave(&pool->lock, flags);
list_for_each_entry(meta, &pool->partial, list)
- dropped += dma_pmd_forget_domain(meta, idx);
+ n += dma_pmd_forget_domain(meta, idx);
list_for_each_entry(meta, &pool->partial_dirty, list)
- dropped += dma_pmd_forget_domain(meta, idx);
+ n += dma_pmd_forget_domain(meta, idx);
list_for_each_entry(meta, &pool->idle, list)
- dropped += dma_pmd_forget_domain(meta, idx);
+ n += dma_pmd_forget_domain(meta, idx);
list_for_each_entry(meta, &pool->full, list)
- dropped += dma_pmd_forget_domain(meta, idx);
+ n += dma_pmd_forget_domain(meta, idx);
spin_unlock_irqrestore(&pool->lock, flags);
+
+ atomic64_add(n, &pool->domain_forget_cnt);
+ dropped += n;
}
mutex_unlock(&dma_pmd_pools_lock);
@@ -843,8 +848,10 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
* A high-order allocation is expensive when it fails, and under fragmentation
* it fails on every pool miss. For GFP_ATOMIC, back off briefly.
*/
- if (!can_block && time_before(jiffies, READ_ONCE(pool->next_alloc_attempt)))
+ if (!can_block && time_before(jiffies, READ_ONCE(pool->next_alloc_attempt))) {
+ atomic64_inc(&pool->fail_backoff);
return NULL;
+ }
/* Aggregate ceiling across every pool, independent of max_idle_pages. */
if (atomic_long_inc_return(&dma_pmd_nr_pages) > READ_ONCE(dma_pmd_max_pages)) {
@@ -852,6 +859,7 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
if (!can_block)
WRITE_ONCE(pool->next_alloc_attempt,
jiffies + DIV_ROUND_UP(HZ, 100));
+ atomic64_inc(&pool->fail_budget);
return NULL;
}
@@ -890,6 +898,14 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
if (!can_block)
WRITE_ONCE(pool->next_alloc_attempt,
jiffies + DIV_ROUND_UP(HZ, 100));
+ atomic64_inc(&pool->fail_nomem);
+ /*
+ * __GFP_NOWARN suppresses the page allocator's own report, so
+ * without this the pool can stop pooling entirely and look
+ * exactly like one that is working.
+ */
+ pr_warn_once("dma_pmd: order-%u allocation failed for pool %p; further failures are counted in debugfs only\n",
+ PMD_ORDER, pool);
return NULL;
}
@@ -922,6 +938,7 @@ static struct dma_pmd_meta *dma_pmd_add_page(struct dma_pmd_pool *pool,
__free_pages(page, PMD_ORDER);
atomic_long_dec(&dma_pmd_nr_pages);
}
+ atomic64_inc(&pool->fail_split);
return NULL;
}
@@ -1020,7 +1037,7 @@ static struct dma_pmd_meta *dma_pmd_pick_meta(struct dma_pmd_pool *pool, bool ze
/* Extract up to @want available blocks from @meta under @pool->lock. */
static unsigned long dma_pmd_take_blocks(struct dma_pmd_pool *pool,
struct dma_pmd_meta *meta,
- unsigned long want,
+ int target_nid, unsigned long want,
struct page **out, bool zero)
{
unsigned long base_pfn = dma_pmd_meta_to_pfn(meta);
@@ -1040,6 +1057,8 @@ static unsigned long dma_pmd_take_blocks(struct dma_pmd_pool *pool,
meta->nr_dirty = 0;
}
dma_pmd_place_meta(pool, meta);
+ if (page_to_nid(pfn_to_page(base_pfn)) != target_nid)
+ pool->numa_mismatch_cnt += used;
return got;
}
@@ -1075,13 +1094,18 @@ unsigned long dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp,
{
unsigned long allocated = 0, flags;
bool should_scrub = false;
+ int target_nid;
bool zero;
- if (unlikely(!pool || pool->destroyed || !nr_pages ||
- (gfp & (__GFP_DMA | __GFP_DMA32 | __GFP_THISNODE |
- __GFP_ACCOUNT))))
+ if (unlikely(!pool || pool->destroyed || !nr_pages))
+ return 0;
+
+ if (unlikely(gfp & (__GFP_DMA | __GFP_DMA32 | __GFP_THISNODE | __GFP_ACCOUNT))) {
+ atomic64_inc(&pool->fail_gfp);
return 0;
+ }
+ target_nid = (nid == NUMA_NO_NODE) ? numa_mem_id() : nid;
zero = want_init_on_alloc(gfp);
spin_lock_irqsave(&pool->lock, flags);
@@ -1093,8 +1117,11 @@ unsigned long dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp,
/* Fallback to a new allocation */
spin_unlock_irqrestore(&pool->lock, flags);
meta = dma_pmd_add_page(pool, gfp, nid, zero);
- if (!meta)
+ if (!meta) {
+ if (!allocated)
+ atomic64_inc(&pool->block_alloc_fail);
goto out;
+ }
spin_lock_irqsave(&pool->lock, flags);
pool->pmd_alloc_cnt++;
list_add(&meta->list, &pool->partial);
@@ -1102,7 +1129,7 @@ unsigned long dma_pmd_pool_alloc_bulk_node(struct dma_pmd_pool *pool, gfp_t gfp,
should_scrub = true;
}
- allocated += dma_pmd_take_blocks(pool, meta,
+ allocated += dma_pmd_take_blocks(pool, meta, target_nid,
nr_pages - allocated,
&page_array[allocated], zero);
}
@@ -1366,6 +1393,91 @@ static unsigned long dma_pmd_shrink_scan(struct shrinker *shrink,
return freed ?: SHRINK_STOP;
}
+/*
+ * One block of lines per pool, in /sys/kernel/debug/dma_pmd/pools.
+ *
+ * Pools that have seen no activity at all are skipped: two are created per
+ * possible CPU and one per RX queue, so on a large machine the overwhelming
+ * majority are idle and would bury the few that carry traffic. A pool that
+ * tried and failed is never skipped - that is the case this file exists for.
+ *
+ * The u64 statistics are snapshotted under @pool->lock and printed after it is
+ * dropped; the atomic64 ones are read locklessly, so a line is not a
+ * consistent instant. That is fine for what this is used for - watching which
+ * way a counter moves - and keeps the file off the locking hot path.
+ *
+ * The partial/full split is not reported: counting it means walking both lists
+ * with @pool->lock held and interrupts off, which is unbounded in the number of
+ * PMD pages a pool holds. "held" is derived from the got/put counters instead.
+ */
+static int dma_pmd_pools_show(struct seq_file *s, void *unused)
+{
+ unsigned int reservoir_nr = 0;
+ struct dma_pmd_pool *pool;
+ int nid;
+
+ for_each_node_state(nid, N_MEMORY)
+ reservoir_nr += READ_ONCE(dma_pmd_reservoirs[nid].nr_pages);
+
+ seq_printf(s, "2m_pages %ld/%lu reservoir %u; pools with no activity are omitted\n",
+ atomic_long_read(&dma_pmd_nr_pages), dma_pmd_max_pages, reservoir_nr);
+
+ mutex_lock(&dma_pmd_pools_lock);
+ list_for_each_entry(pool, &dma_pmd_pools, node) {
+ u64 alloc_2m, free_2m, block_alloc, block_free, block_scrub;
+ u64 block_fail, fail_gfp, numa_mismatch;
+ unsigned int num_idle, max_idle;
+ unsigned long flags;
+
+ block_fail = atomic64_read(&pool->block_alloc_fail);
+ fail_gfp = atomic64_read(&pool->fail_gfp);
+
+ spin_lock_irqsave(&pool->lock, flags);
+ alloc_2m = pool->pmd_alloc_cnt;
+ free_2m = pool->pmd_free_cnt;
+ block_alloc = pool->block_alloc_cnt;
+ block_free = pool->block_free_cnt;
+ block_scrub = pool->block_scrub_cnt;
+ numa_mismatch = pool->numa_mismatch_cnt;
+ num_idle = pool->num_idle_pages;
+ max_idle = pool->max_idle_pages;
+ spin_unlock_irqrestore(&pool->lock, flags);
+
+ /*
+ * A pool that acquired no PMD page is the one worth looking at,
+ * not the one to hide: it is indistinguishable from a healthy
+ * pool everywhere else, because its callers just fall back.
+ * Omit a pool only when nothing has ever happened to it, which
+ * is the idle-per-CPU case this filter exists for.
+ */
+ if (!alloc_2m && !block_alloc && !block_fail && !fail_gfp)
+ continue;
+
+ seq_printf(s, "pool %p order %u\n", pool, pool->order);
+ seq_printf(s, " 2m_pages got %llu put %llu held %llu idle %u/%u\n",
+ alloc_2m, free_2m, alloc_2m - free_2m, num_idle, max_idle);
+ seq_printf(s, " blocks alloc %llu free %llu scrub %llu numa_mismatch %llu failed %llu gfp %llu\n",
+ block_alloc, block_free, block_scrub, numa_mismatch,
+ block_fail, fail_gfp);
+ seq_printf(s, " no_2m backoff %lld budget %lld nomem %lld split %lld\n",
+ atomic64_read(&pool->fail_backoff),
+ atomic64_read(&pool->fail_budget),
+ atomic64_read(&pool->fail_nomem),
+ atomic64_read(&pool->fail_split));
+ seq_printf(s, " domains mapped %lld forgotten %lld\n",
+ atomic64_read(&pool->pmd_map_cnt),
+ atomic64_read(&pool->domain_forget_cnt));
+ seq_printf(s, " per_buf no_window %lld no_pool %lld map_err %lld\n",
+ atomic64_read(&pool->fallback_nowindow),
+ atomic64_read(&pool->fallback_nopool),
+ atomic64_read(&pool->fallback_maperr));
+ }
+ mutex_unlock(&dma_pmd_pools_lock);
+
+ return 0;
+}
+DEFINE_SHOW_ATTRIBUTE(dma_pmd_pools);
+
static int __init dma_pmd_init(void)
{
struct shrinker *shrink;
@@ -1377,7 +1489,7 @@ static int __init dma_pmd_init(void)
}
/*
- * Hard ceiling on memory diverted from the buddy allocator into 2MB
+ * Hard ceiling on memory diverted from the buddy allocator into DMA_PMD
* pools, in PMD pages. One eighth of RAM is far above what the per-pool
* watermarks should ever reach; it exists to bound a pathological
* configuration (many queues, many CPUs, all pools at their high
@@ -1403,6 +1515,13 @@ static int __init dma_pmd_init(void)
if (!dma_pmd_wq)
pr_warn("dma_pmd: no reclaim workqueue, falling back to system_wq\n");
+ /*
+ * No error handling and no dentry kept: this is built in and never
+ * unloaded, and the stubs are no-ops when debugfs is disabled.
+ */
+ debugfs_create_file("pools", 0444, debugfs_create_dir("dma_pmd", NULL),
+ NULL, &dma_pmd_pools_fops);
+
shrink = shrinker_alloc(0, "dma_pmd");
if (!shrink) {
pr_warn("dma_pmd: shrinker registration failed\n");
diff --git a/drivers/iommu/dma-pmd-priv.h b/drivers/iommu/dma-pmd-priv.h
index 82a5c869aa216..b9bf48eb7f0f7 100644
--- a/drivers/iommu/dma-pmd-priv.h
+++ b/drivers/iommu/dma-pmd-priv.h
@@ -178,7 +178,8 @@ static inline unsigned int dma_pmd_meta_avail(const struct dma_pmd_meta *m)
* struct dma_pmd_pool - Pool managing PMD-backed blocks of a fixed order
*
* @lock: Spinlock protecting @partial, @idle, @full, @num_idle_pages,
- * block bitmaps and the statistics counters
+ * the block bitmaps and the u64 statistics counters. The
+ * atomic64_t counters below are deliberately outside it.
* @order: Block order managed by this pool (<= PMD_ORDER)
* @destroyed: Set when dma_pmd_pool_destroy() has been called
* @partial: List of partially used PMD pages with @nr_free > 0 (clean-only
@@ -196,6 +197,7 @@ static inline unsigned int dma_pmd_meta_avail(const struct dma_pmd_meta *m)
* @block_alloc_cnt: Statistics counter of block allocations satisfied
* @block_free_cnt: Statistics counter of block frees recycled into pool
* @block_scrub_cnt: Statistics counter of dirty blocks zeroed by scrubber
+ * @numa_mismatch_cnt: Block allocations satisfied from a non-target NUMA node
* @next_alloc_attempt: Do not attempt a new order-9 allocation before this time.
* Damps repeated high-order GFP_ATOMIC failures under
* fragmentation, which would otherwise be retried on every
@@ -203,6 +205,26 @@ static inline unsigned int dma_pmd_meta_avail(const struct dma_pmd_meta *m)
* @refcount: Reference count; held by pool creator and each active PMD page
* @node: Node in the global dma_pmd_pools list
* @dead_node: Node in dma_pmd_dead_pools once @refcount reaches 0
+ * @fail_backoff: PMD page acquisitions skipped because @next_alloc_attempt was set
+ * @fail_budget: PMD page acquisitions refused by global dma_pmd_max_pages
+ * @fail_nomem: PMD page acquisitions that ran out of memory for the order-9
+ * allocation. Silent otherwise: the request carries __GFP_NOWARN
+ * so the page allocator says nothing.
+ * @fail_split: split_page_compound() rejections
+ * @block_alloc_fail: PMD page acquisitions that failed, making
+ * dma_pmd_pool_alloc() return NULL so the caller had to fall
+ * back to a plain page
+ * @fail_gfp: dma_pmd_pool_alloc() calls declined due to unsupported GFP
+ * flags (__GFP_DMA, __GFP_DMA32, __GFP_THISNODE, __GFP_ACCOUNT)
+ * @pmd_map_cnt: PMD page mappings established across all domains
+ * @fallback_nowindow: Buffers mapped individually because the domain has no
+ * usable IOVA window - no free range large enough, no free
+ * domain index, or a window the device cannot address
+ * @fallback_nopool: Buffers mapped individually because the domain can never
+ * pool: direct isolation, or no PMD page size
+ * @fallback_maperr: Buffers mapped individually because the PMD page mapping
+ * itself failed
+ * @domain_forget_cnt: Cached mappings dropped by dma_pmd_domain_release()
*/
struct dma_pmd_pool {
/* First cacheline: hot fields touched on block alloc and free. */
@@ -222,12 +244,31 @@ struct dma_pmd_pool {
u64 block_alloc_cnt;
u64 block_free_cnt;
u64 block_scrub_cnt;
+ u64 numa_mismatch_cnt;
/* Cold. */
unsigned long next_alloc_attempt;
struct kref refcount;
struct list_head node;
struct llist_node dead_node;
+
+ /*
+ * Statistics updated without @lock. dma_pmd_add_page() runs outside
+ * it, and the map slow path holds @meta->map_lock - taking @lock there
+ * would invert the order against dma_pmd_domain_release(), which
+ * walks @lock then map_lock.
+ */
+ atomic64_t fail_backoff;
+ atomic64_t fail_budget;
+ atomic64_t fail_nomem;
+ atomic64_t fail_split;
+ atomic64_t block_alloc_fail;
+ atomic64_t fail_gfp;
+ atomic64_t pmd_map_cnt;
+ atomic64_t fallback_nowindow;
+ atomic64_t fallback_nopool;
+ atomic64_t fallback_maperr;
+ atomic64_t domain_forget_cnt;
} ____cacheline_aligned;
extern unsigned long dma_pmd_meta_nframes;
--
2.56.0.rc1.315.gc6ed9934b7-goog
^ permalink raw reply related [flat|nested] 25+ messages in thread
* Re: [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table
2026-10-03 21:22 ` [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table Luigi Rizzo
@ 2026-10-03 21:44 ` Randy Dunlap
2026-10-04 9:14 ` Luigi Rizzo
0 siblings, 1 reply; 25+ messages in thread
From: Randy Dunlap @ 2026-10-03 21:44 UTC (permalink / raw)
To: Luigi Rizzo, Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel
On 10/3/26 2:22 PM, Luigi Rizzo wrote:
> diff --git a/drivers/iommu/Kconfig b/drivers/iommu/Kconfig
> index 6e07bd69467a3..1bb347fd9da4a 100644
> --- a/drivers/iommu/Kconfig
> +++ b/drivers/iommu/Kconfig
> @@ -158,6 +158,38 @@ config IOMMU_DMA
> select NEED_SG_DMA_LENGTH
> select NEED_SG_DMA_FLAGS if SWIOTLB
>
> +# PMD_SIZE IOMMU backing for DMA-coherent devices
> +config DMA_PMD
> + bool "PMD_SIZE IOMMU backing for DMA-coherent devices"
> + depends on IOMMU_DMA
> + depends on X86_64 || (ARM64 && ARM64_4K_PAGES)
> + depends on !PREEMPT_RT
> + default y
> + help
> + Hand out DMA buffers carved out of PMD_SIZE physically contiguous
> + blocks that the IOMMU maps with a single leaf PTE. Mapping,
> + unmapping and IOTLB flushing are done lazily to reduce CPU overhead.
> + The large mappings greatly reduce IOTLB usage. As in strict iommu
> + mode, pages are returned to the buddy allocator only when any existing
> + mappings have been removed and IOTLB flushed.
> + Costs 128KB of metadata per GB of DMA_PMD buffers, allocated on demand.
Paragraph 1 of the patch description says 64KB per 1GB, not 128KB.
Or am I mixing things up?
--
~Randy
^ permalink raw reply [flat|nested] 25+ messages in thread
* Re: [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table
2026-10-03 21:44 ` Randy Dunlap
@ 2026-10-04 9:14 ` Luigi Rizzo
0 siblings, 0 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-04 9:14 UTC (permalink / raw)
To: Randy Dunlap
Cc: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
Christoph Hellwig, Marek Szyprowski, Andrew Morton,
Vlastimil Babka, David Hildenbrand, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Greg Kroah-Hartman,
Rafael J . Wysocki, Danilo Krummrich, Jonathan Corbet,
Jesper Dangaard Brouer, Ilias Apalodimas, Willem de Bruijn,
Kuniyuki Iwashima, Joshua Washington, Harshitha Ramamurthy,
Saeed Mahameed, Tariq Toukan, Tony Nguyen, Przemek Kitszel,
Alexander Lobakin, Michael Chan, Pavan Chebbi, iommu, netdev,
linux-mm, driver-core, linux-doc, linux-kernel
On Sat, Oct 3, 2026 at 11:44 PM Randy Dunlap <rdunlap@infradead.org> wrote:
>
>
>
> On 10/3/26 2:22 PM, Luigi Rizzo wrote:
> > diff --git a/drivers/iommu/Kconfig b/drivers/iommu/Kconfig
> > index 6e07bd69467a3..1bb347fd9da4a 100644
> > --- a/drivers/iommu/Kconfig
> > +++ b/drivers/iommu/Kconfig
> > @@ -158,6 +158,38 @@ config IOMMU_DMA
> > select NEED_SG_DMA_LENGTH
> > select NEED_SG_DMA_FLAGS if SWIOTLB
> >
> > +# PMD_SIZE IOMMU backing for DMA-coherent devices
> > +config DMA_PMD
...
> > + Costs 128KB of metadata per GB of DMA_PMD buffers, allocated on demand.
>
> Paragraph 1 of the patch description says 64KB per 1GB, not 128KB.
> Or am I mixing things up?
256B per 2MB page, 128KB per 1GB is the correct version.
The number in the patch description is a leftover from an earlier
version, I will update.
cheers
luigi
^ permalink raw reply [flat|nested] 25+ messages in thread
end of thread, other threads:[~2026-10-04 9:15 UTC | newest]
Thread overview: 25+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table Luigi Rizzo
2026-10-03 21:44 ` Randy Dunlap
2026-10-04 9:14 ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 02/22] iommu/dma: add DMA_PMD pool lifecycle and page recycle hook Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 03/22] mm: Add split_page_compound() Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 04/22] iommu/dma: add DMA_PMD pool block allocation Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 05/22] iommu/dma: Global cap and shrinker for DMA_PMD pool memory Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 06/22] iommu/dma: reserve a per-domain IOVA window for DMA_PMD pages Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 07/22] iommu/dma: release DMA_PMD domain mappings on domain teardown Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 08/22] iommu/dma: use per-domain IOVA window to map DMA_PMD memory Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 09/22] iommu/dma: Add DMA_PMD arena allocator Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 10/22] driver core: Add per-device dma_pmd_* sysfs attributes Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 11/22] dma-mapping: Use DMA_PMD arena for dma_alloc_attrs() Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 12/22] net/core: Use per-CPU DMA_PMD pools for skb_page_frag_refill() Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 13/22] net/core: Use DMA_PMD for page_pool memory Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 14/22] iommu/dma: Support decrypted and pinned DMA_PMD pages Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 15/22] iommu/dma: Add background page scrubber for DMA_PMD pools Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 16/22] iommu/dma: Add per-NUMA-node PMD page reservoir Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 17/22] net/gve: Use DMA_PMD memory for RX buffers Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 18/22] net/gve: Use DMA_PMD memory for tx header bounce buffers Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 19/22] net/mlx5e: " Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 20/22] net/idpf: " Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 21/22] net/bnxt: " Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 22/22] iommu/dma: Add DMA_PMD statistics and debugfs Luigi Rizzo
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox