From: Pranjal Shrivastava <praan@google.com>
To: Joerg Roedel <joro@8bytes.org>, Will Deacon <will@kernel.org>,
Robin Murphy <robin.murphy@arm.com>,
Jason Gunthorpe <jgg@ziepe.ca>, Kevin Tian <kevin.tian@intel.com>,
Alex Williamson <alex@shazbot.org>,
David Matlack <dmatlack@google.com>,
Jonathan Corbet <corbet@lwn.net>,
Shuah Khan <skhan@linuxfoundation.org>,
Randy Dunlap <rdunlap@infradead.org>
Cc: Mostafa Saleh <smostafa@google.com>,
Daniel Mentz <danielmentz@google.com>,
Samiullah Khawaja <skhawaja@google.com>,
iommu@lists.linux.dev, kvm@vger.kernel.org,
linux-kselftest@vger.kernel.org, linux-doc@vger.kernel.org,
Logan Odell <loganodell@google.com>,
Pranjal Shrivastava <praan@google.com>
Subject: [RFC PATCH 0/7] iommu: Introduce Page Table Observability Framework
Date: Thu, 1 Oct 2026 22:45:24 +0000 [thread overview]
Message-ID: <20261001224531.765278-1-praan@google.com> (raw)
Introduce an observability framework to track and expose IOMMU page table
memory usage per domain. As VMMs and userspace drivers map and unmap large,
sparse IOVA regions through VFIO and iommufd, page table directories are
often left allocated but completely empty. generic_pt frees a table when a
single unmap covers it entirely, but tables that empty through a series of
partial unmaps stay allocated until the domain is destroyed.
Currently, this *stranded* memory is only visible in aggregate
(nr_iommu_pages in /proc/vmstat, and sec_pagetables in the memcg
memory.stat). There is no way to tell which domain, and hence which VFIO
container or iommufd context, owns it. On a host, it shows up as a drop in
available memory with no corresponding owner and has resulted in OOMs
without any diagnostic information.
The series implements the tracking infrastructure to attribute this
memory to its domain and exposes it per fd, so that userspace / telemetry
can see which VFIO container or iommufd context holds it.
Design
======
The tracking is implemented once, within the generic_pt library:
- struct iommu_domain gains an atomic_long_t nr_pages counter.
- To avoid expanding existing metadata structures, struct ioptdesc
repurposes the unused page->private slot to alias the owning
struct iommu_domain pointer. The ioptdesc layout stays identical to
struct page.
- The core page allocator (iommu_alloc_pages_node_sz()) is wrapped with
an attributed variant, iommu_alloc_pages_node_sz_attributed(), which
charges the domain on allocation. The common free path uncharges it
using the stored domain pointer, so no free call site changes.
- generic_pt's table allocator uses the attributed variant, so every
driver built on generic_pt gains accounting without driver changes.
The cost is one atomic operation per page table allocation and free. Leaf
map and unmap are unchanged.
Userspace interface
===================
The per-domain counts are aggregated when fdinfo is read, and exposed
through /proc/<pid>/fdinfo/<fd>:
- VFIO container fd (type1): summed over all attached domains.
- iommufd fd: summed over all paging HWPTs in the context. Nested HWPTs
are skipped as their stage-1 page tables are owned by userspace.
Both backends report the same key, so tooling does not need to know
whether a VMM uses VFIO type1 or iommufd:
iommu-nr-pages: <pages>
Only domains whose page table is implemented by generic_pt account their
memory. If any relevant domain does not, the key is omitted rather than
reporting a misleading partial count. The field is documented in
Documentation/filesystems/proc.rst.
Built on generic_pt
=====================
All users of generic_pt in this tree: Intel VT-d (first and second stage),
AMD IOMMU (v1 and v2 page tables) and RISC-V. ARM SMMUv3 will gain
accounting once its conversion to generic_pt lands. [1]
Open questions
==============
- Is a single, backend-agnostic fdinfo key preferred over per-backend
keys (e.g. vfio-nr-pages / iommufd-nr-pages)?
- Is omitting the key preferable to reporting a partial count when a
domain without accounting is attached?
- The value is in pages, consistent with nr_iommu_pages in /proc/vmstat.
Would bytes be preferred?
- Is aliasing page->private through the ioptdesc overlay acceptable?
Upcoming Work / Roadmap
=======================
An IO page table shrinker, which reclaims empty leaf directories under
memory pressure, is posted separately as an RFC. The two series are
independent. Per-domain accounting is useful on its own, and also
quantifies the memory such reclaim could recover.
There's an alignment session at Linux Plumbers Conference 2026 for these [2]
[1] https://lore.kernel.org/all/0-v2-563ee63886f0+1209-iommupt_armv8_jgg@nvidia.com/
[2] https://lpc.events/event/20/contributions/2624/
Logan Odell (1):
vfio/type1: Expose IO page table usage via fdinfo
Pranjal Shrivastava (6):
iommu: Add infrastructure for per-domain IOPT accounting
iommu: Implement domain-attributed page allocation
iommupt: Enable per-domain IOPT attribution
iommufd: Expose IO page table usage via fdinfo
iommufd/selftest: Add observability test for iommu-nr-pages
vfio/selftests: Add observability test for IO page table usage
Documentation/filesystems/proc.rst | 34 +++++
drivers/iommu/generic_pt/iommu_pt.h | 7 +-
drivers/iommu/iommu-pages.c | 35 +++++-
drivers/iommu/iommu-pages.h | 25 +++-
drivers/iommu/iommufd/main.c | 38 ++++++
drivers/vfio/container.c | 21 ++++
drivers/vfio/vfio.h | 2 +
drivers/vfio/vfio_iommu_type1.c | 21 ++++
include/linux/iommu.h | 2 +
tools/testing/selftests/iommu/iommufd.c | 48 +++++++
tools/testing/selftests/vfio/Makefile | 1 +
.../selftests/vfio/vfio_nr_pages_test.c | 119 ++++++++++++++++++
12 files changed, 347 insertions(+), 6 deletions(-)
create mode 100644 tools/testing/selftests/vfio/vfio_nr_pages_test.c
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.56.0.rc1.315.gc6ed9934b7-goog
next reply other threads:[~2026-10-01 22:45 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-01 22:45 Pranjal Shrivastava [this message]
2026-10-01 22:45 ` [RFC PATCH 1/7] iommu: Add infrastructure for per-domain IOPT accounting Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 2/7] iommu: Implement domain-attributed page allocation Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 3/7] iommupt: Enable per-domain IOPT attribution Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 4/7] vfio/type1: Expose IO page table usage via fdinfo Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 5/7] iommufd: " Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 6/7] iommufd/selftest: Add observability test for iommu-nr-pages Pranjal Shrivastava
2026-10-01 22:45 ` [RFC PATCH 7/7] vfio/selftests: Add observability test for IO page table usage Pranjal Shrivastava
2026-10-02 9:14 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261001224531.765278-1-praan@google.com \
--to=praan@google.com \
--cc=alex@shazbot.org \
--cc=corbet@lwn.net \
--cc=danielmentz@google.com \
--cc=dmatlack@google.com \
--cc=iommu@lists.linux.dev \
--cc=jgg@ziepe.ca \
--cc=joro@8bytes.org \
--cc=kevin.tian@intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-doc@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=loganodell@google.com \
--cc=rdunlap@infradead.org \
--cc=robin.murphy@arm.com \
--cc=skhan@linuxfoundation.org \
--cc=skhawaja@google.com \
--cc=smostafa@google.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox