Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools
@ 2026-10-03 21:22 Luigi Rizzo
  2026-10-03 21:22 ` [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table Luigi Rizzo
                   ` (21 more replies)
  0 siblings, 22 replies; 25+ messages in thread
From: Luigi Rizzo @ 2026-10-03 21:22 UTC (permalink / raw)
  To: Luigi Rizzo, Joerg Roedel, Will Deacon, Robin Murphy,
	Christoph Hellwig, Marek Szyprowski, Andrew Morton,
	Vlastimil Babka, David Hildenbrand, David S . Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Greg Kroah-Hartman, Rafael J . Wysocki, Danilo Krummrich,
	Jonathan Corbet, Jesper Dangaard Brouer, Ilias Apalodimas,
	Willem de Bruijn, Kuniyuki Iwashima, Joshua Washington,
	Harshitha Ramamurthy, Saeed Mahameed, Tariq Toukan, Tony Nguyen,
	Przemek Kitszel, Alexander Lobakin, Michael Chan, Pavan Chebbi,
	iommu, netdev, linux-mm, driver-core, linux-doc, linux-kernel,
	Luigi Rizzo

Here is a subsystem called DMA_PMD on which I would like feedback on
architecture, possible enhancements, or kernel components that could be
reused to avoid duplication.

I think it can be extremely useful for those affected by the HW/SW overhead
of IOMMU (especially IOTLB thrashing, see [1]), or confidential computing,
or preemptable VMs.

The (not too exaggerated) pitch line is

DMA_PMD has the performance of identity, guarantees that only IO buffers
can ever have active IOMMU or IOTLB mappings, removes the bounce buffer
overhead in confidential computing and preemptable VMs, and integrates
smoothly with existing kernel APIs.

=== ARCHITECTURE

DMA_PMD was initially designed to address IOTLB thrashing (details in [1]),
but turned out to also resolve nicely the strict IOMMU overhead, and avoid
the bounce buffer overhead in preemptible VMs and Confidential Computing.

It works by combining several known techniques:

- transparently feed allocators of device memory (skb_page_frag_refill(),
  dma_alloc_attrs(), pagepool, ...) with DMA_PMD pages, i.e. physically
  contiguous PMD_SIZE (2MB) pages mapped via PDE_SIZE IOMMU entries

- heavy recycling of DMA_PMD pages (like pagepool) and lazy IOMMU unmapping
  (like DMA-FQ), BUT:

- like strict IOMMU (DMA), safely release memory back to the kernel only
  after destroying all of its IOMMU mappings and a synchronous IOTLB flush

- on allocation, DMA_PMD pages can be configured to be pinned in the host
  (hence suitable for preemptible VMs) and/or unencrypted (hence suitable
  for Confidential Computing), removing the need for bounce buffers

- there are global and per-device optin /sys/device/.../dma_pmd_*
  and /proc/sys/net/core/tx_enable_dma_pmd

NIC drivers can typically use dma_pmd with no modifications for rings,
tx buffers, rx buffers (if they use pagepool), and rx headers. For tx
headers, almost all drivers need changes (see later patches in the series)
to implement cheap bounce buffers and avoid individual 4K mappings.

The series has the following main components:

- a sparse array (similar to pageblock_flags) to quickly attach metadata
  to a 2MB page without fiddling with the struct page. Cost is 128B per 2MB
  page used as an IO buffer, totally negligible.

- dma_pmd_pool, a replacement for alloc_pages() that can be instantiated
  per-CPU or per receive queue. It handles the split of PMD_SIZE pages
  into order-N blocks, handles dma_map and unmap, and aggressively recycles
  entries. It is used to feed skb_page_frag_refill(), pagepool, and
  receive buffers for drivers that do not use pagepool.
  See [2] for "WHY NOT PAGEPOOL FOR TX AND EVERYTHING"

- dma_pmd_arena, is another allocator backed by DMA_PMD pages and
  is used exclusively as the backend for dma_alloc_attrs().
  Used for longer-lived allocations (descriptor/completion rings,
  rx and tx header buffers)

- glue code to hook DMA_PMD into pagepool [2], skb_page_frag_refill(),
  dma_alloc_attrs(), dma_map/unmap...

- per-driver patches, where necessary (e.g. tx header buffers [3])

=== PERFORMANCE BENEFITS

Your mileage may vary. Enabling the IOMMU may have no throughput
impact, until it does when some system components (bus, IOTLB, CPU)
become overloaded.  Aside from throughput reduction, one interesting
parameter is the effectiveness of the IOTLB. Here is a sample of SMMU
performance counters for a large ARM system with 2x200G NICs doing
bidirectional traffic:

    === DMA-FQ MODE (total throughput ~440Gbps)
    25,757,488      smmuv3_pmcg_*/event=0x80/   IOTLB lookups
    16,634,941      smmuv3_pmcg_*/event=0x81/   IOTLB  misses

    === DMA_PMD on top of strict DMA (total throughput ~745Gbps)
    22,693,218      smmuv3_pmcg_*/event=0x80/   IOTLB lookups
        30,536      smmuv3_pmcg_*/event=0x81/   IOTLB  misses
    (not a mistake, also lookups went down despite the higher rate
    because the tx side can use larger segments)

The 2MB mappings made IOTLB misses almost non existent, because
the working set is reduced by a factor of ~512.
       
Note, the code has more verbose comments than I would like.
Several of them are there to avoid Sashiko getting confused and
flagging false positives.

=== NOTES

[1] IOMMU IMPACT AND IOTLB THRASHING
This is documented in more detail in Documentation/core-api/dma-pmd.rst
but the compact version is below.

There are three main costs involved with using the IOMMU:
- CPU cost for dma map/unmap
- CPU cost and latency for synchronous IOTLB flush (required for strong security)
- IOTLB thrashing, shows up dramatically when the IO access pattern
  exceeds the IOTLB size, and IOMMU page walks slow down bus activity
  up to a point where we see over 30..60% throughput reduction just for
  this reason.  Some relevant references

  https://lore.kernel.org/all/4b42f2eb-dc29-153e-ace9-5584ea2e5070@redhat.com/
  https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10280580/

Specifically, NICs access 4 or more pages per packet (descriptor queue,
completion queue, rx or tx header, one or more buffers), with very little
locality especially for data buffers and tx headers. Each queue has 1-4K
entries, a fast NIC uses 16 or more queues per direction, so just the
data buffers cover 32..128K pages. Using 4KB mappings makes the IOTLB
ineffective, as shown by the performance counters shown earlier.

[2] WHY NOT PAGEPOOL FOR TX AND EVERYTHING
There is some overlap between pagepool and dma_pmd_pool in that both
implement a cache, but pagepool is missing some important features:

- pagepool does not handle splitting 2M pages in smaller chunks, or
  the mappings. That would need to be implemented (and is what dma_pmd_pool does)

- pagepool's provided pages cannot be cleanly managed by
  get_page()/put_page(), which is what all the consumers of
  skb_page_frag_refill(). Fixing that would require dozen of changes
  with high risk of missing some paths.

- even for the receive path, a provider for pagepool has more restrictions
  than alloc_pages-supplied memory, so if I wanted to wrap dma_pmd_pool
  as a provider I would have needed changes to the provider interface.

- pagepool requires a device at allocation time, which is not known when
  skb_page_frag_refill() runs, hence the IOMMU mappings cannot be handled
  at allocation time.

- page_pool_alloc_pages assumes a single NAPI consumer with BH disabled,
  whereas skb_page_frag_refill() runs from a preemptive context.

- and last not least, we need to handle the dma_alloc_attrs() allocations
  which don't match pagepool API

Of course all of the above could be modified, but the complexity
would be similar to that of implementing dma_pmd_pool, and with a
huge risk of breaking existing functionality or missing some path.

[3] DRIVER CHANGES

Most drivers need no change for rings, transmit buffers, or receive buffers
(if they use pagepool). Tx headers are generally mapped on the fly and
that is enough to trigger IOTLB thrashing, so most drivers need a small
change to implement very cheap tx bounce buffers backed by DMA_PMD pages.

We could avoid the bounce buffers using DMA_PMD for the skb->head region,
but that memory is contiguous to skb_shinfo so it would be exposed to
IO device access, which may be undesirable for security.

Luigi Rizzo (22):
  iommu/dma: introduce CONFIG_DMA_PMD and metadata table
  iommu/dma: add DMA_PMD pool lifecycle and page recycle hook
  mm: Add split_page_compound()
  iommu/dma: add DMA_PMD pool block allocation
  iommu/dma: Global cap and shrinker for DMA_PMD pool memory
  iommu/dma: reserve a per-domain IOVA window for DMA_PMD pages
  iommu/dma: release DMA_PMD domain mappings on domain teardown
  iommu/dma: use per-domain IOVA window to map DMA_PMD memory
  iommu/dma: Add DMA_PMD arena allocator
  driver core: Add per-device dma_pmd_* sysfs attributes
  dma-mapping: Use DMA_PMD arena for dma_alloc_attrs()
  net/core: Use per-CPU DMA_PMD pools for skb_page_frag_refill()
  net/core: Use DMA_PMD for page_pool memory
  iommu/dma: Support decrypted and pinned DMA_PMD pages
  iommu/dma: Add background page scrubber for DMA_PMD pools
  iommu/dma: Add per-NUMA-node PMD page reservoir
  net/gve: Use DMA_PMD memory for RX buffers
  net/gve: Use DMA_PMD memory for tx header bounce buffers
  net/mlx5e: Use DMA_PMD memory for tx header bounce buffers
  net/idpf: Use DMA_PMD memory for tx header bounce buffers
  net/bnxt: Use DMA_PMD memory for tx header bounce buffers
  iommu/dma: Add DMA_PMD statistics and debugfs

 Documentation/core-api/dma-pmd.rst            |  320 ++++
 Documentation/core-api/index.rst              |    1 +
 drivers/base/core.c                           |   51 +
 drivers/iommu/Kconfig                         |   33 +
 drivers/iommu/Makefile                        |    3 +
 drivers/iommu/dma-iommu.c                     |   72 +-
 drivers/iommu/dma-iommu.h                     |    8 +
 drivers/iommu/dma-pmd-arena.c                 |  412 +++++
 drivers/iommu/dma-pmd-kunit.c                 |  225 +++
 drivers/iommu/dma-pmd-map.c                   |  531 ++++++
 drivers/iommu/dma-pmd-meta.c                  |  403 +++++
 drivers/iommu/dma-pmd-pool.c                  | 1538 +++++++++++++++++
 drivers/iommu/dma-pmd-priv.h                  |  385 +++++
 drivers/iommu/iommu.c                         |    2 +-
 drivers/net/ethernet/broadcom/bnxt/bnxt.c     |   45 +-
 drivers/net/ethernet/broadcom/bnxt/bnxt.h     |    2 +
 drivers/net/ethernet/google/gve/gve.h         |    5 +
 drivers/net/ethernet/google/gve/gve_main.c    |    5 +
 drivers/net/ethernet/google/gve/gve_rx.c      |   61 +-
 drivers/net/ethernet/google/gve/gve_tx.c      |   38 +-
 drivers/net/ethernet/google/gve/gve_tx_dqo.c  |   45 +-
 .../ethernet/intel/idpf/idpf_singleq_txrx.c   |    6 +-
 drivers/net/ethernet/intel/idpf/idpf_txrx.c   |   21 +-
 drivers/net/ethernet/intel/idpf/idpf_txrx.h   |   27 +-
 drivers/net/ethernet/mellanox/mlx5/core/en.h  |    2 +
 .../net/ethernet/mellanox/mlx5/core/en/txrx.h |    4 +-
 .../net/ethernet/mellanox/mlx5/core/en_main.c |   14 +
 .../net/ethernet/mellanox/mlx5/core/en_tx.c   |   28 +-
 include/linux/device.h                        |    5 +
 include/linux/dma-pmd.h                       |  198 +++
 include/linux/mm.h                            |    2 +
 include/net/libeth/tx.h                       |    5 +-
 include/net/page_pool/types.h                 |    3 +
 include/net/sock.h                            |    3 +
 kernel/dma/direct.c                           |    6 +-
 kernel/dma/mapping.c                          |   39 +-
 mm/Kconfig.debug                              |   11 +
 mm/Makefile                                   |    1 +
 mm/page_alloc.c                               |   89 +
 mm/split_page_compound_kunit.c                |  113 ++
 net/core/page_pool.c                          |   61 +-
 net/core/sock.c                               |   86 +-
 net/core/sysctl_net_core.c                    |    7 +
 43 files changed, 4835 insertions(+), 81 deletions(-)
 create mode 100644 Documentation/core-api/dma-pmd.rst
 create mode 100644 drivers/iommu/dma-pmd-arena.c
 create mode 100644 drivers/iommu/dma-pmd-kunit.c
 create mode 100644 drivers/iommu/dma-pmd-map.c
 create mode 100644 drivers/iommu/dma-pmd-meta.c
 create mode 100644 drivers/iommu/dma-pmd-pool.c
 create mode 100644 drivers/iommu/dma-pmd-priv.h
 create mode 100644 include/linux/dma-pmd.h
 create mode 100644 mm/split_page_compound_kunit.c

-- 
2.56.0.rc1.315.gc6ed9934b7-goog



^ permalink raw reply	[flat|nested] 25+ messages in thread

end of thread, other threads:[~2026-10-04  9:15 UTC | newest]

Thread overview: 25+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-03 21:22 [RFC: DMA_PMD 00/22] DMA_PMD: PMD_SIZE-backed IO buffer pools Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 01/22] iommu/dma: introduce CONFIG_DMA_PMD and metadata table Luigi Rizzo
2026-10-03 21:44   ` Randy Dunlap
2026-10-04  9:14     ` Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 02/22] iommu/dma: add DMA_PMD pool lifecycle and page recycle hook Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 03/22] mm: Add split_page_compound() Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 04/22] iommu/dma: add DMA_PMD pool block allocation Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 05/22] iommu/dma: Global cap and shrinker for DMA_PMD pool memory Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 06/22] iommu/dma: reserve a per-domain IOVA window for DMA_PMD pages Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 07/22] iommu/dma: release DMA_PMD domain mappings on domain teardown Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 08/22] iommu/dma: use per-domain IOVA window to map DMA_PMD memory Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 09/22] iommu/dma: Add DMA_PMD arena allocator Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 10/22] driver core: Add per-device dma_pmd_* sysfs attributes Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 11/22] dma-mapping: Use DMA_PMD arena for dma_alloc_attrs() Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 12/22] net/core: Use per-CPU DMA_PMD pools for skb_page_frag_refill() Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 13/22] net/core: Use DMA_PMD for page_pool memory Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 14/22] iommu/dma: Support decrypted and pinned DMA_PMD pages Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 15/22] iommu/dma: Add background page scrubber for DMA_PMD pools Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 16/22] iommu/dma: Add per-NUMA-node PMD page reservoir Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 17/22] net/gve: Use DMA_PMD memory for RX buffers Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 18/22] net/gve: Use DMA_PMD memory for tx header bounce buffers Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 19/22] net/mlx5e: " Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 20/22] net/idpf: " Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 21/22] net/bnxt: " Luigi Rizzo
2026-10-03 21:22 ` [RFC: DMA_PMD 22/22] iommu/dma: Add DMA_PMD statistics and debugfs Luigi Rizzo

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox