From: Vincent Donnefort <vdonnefort@google.com>
To: Thierry Reding <thierry.reding@kernel.org>
Cc: "Rob Herring" <robh@kernel.org>,
"Krzysztof Kozlowski" <krzk+dt@kernel.org>,
"Conor Dooley" <conor+dt@kernel.org>,
"Jonathan Hunter" <jonathanh@nvidia.com>,
"David Airlie" <airlied@gmail.com>,
"Simona Vetter" <simona@ffwll.ch>,
"Maarten Lankhorst" <maarten.lankhorst@linux.intel.com>,
"Maxime Ripard" <mripard@kernel.org>,
"Thomas Zimmermann" <tzimmermann@suse.de>,
"Sowjanya Komatineni" <skomatineni@nvidia.com>,
"Luca Ceresoli" <luca.ceresoli@bootlin.com>,
"Mikko Perttunen" <mperttunen@nvidia.com>,
"Yury Norov" <yury.norov@gmail.com>,
"Rasmus Villemoes" <linux@rasmusvillemoes.dk>,
"Russell King" <linux@armlinux.org.uk>,
"Alexander Gordeev" <agordeev@linux.ibm.com>,
"Gerald Schaefer" <gerald.schaefer@linux.ibm.com>,
"Heiko Carstens" <hca@linux.ibm.com>,
"Vasily Gorbik" <gor@linux.ibm.com>,
"Christian Borntraeger" <borntraeger@linux.ibm.com>,
"Sven Schnelle" <svens@linux.ibm.com>,
"Andrew Morton" <akpm@linux-foundation.org>,
"David Hildenbrand" <david@kernel.org>,
"Lorenzo Stoakes" <ljs@kernel.org>,
"Liam R. Howlett" <liam@infradead.org>,
"Vlastimil Babka" <vbabka@kernel.org>,
"Mike Rapoport" <rppt@kernel.org>,
"Suren Baghdasaryan" <surenb@google.com>,
"Michal Hocko" <mhocko@suse.com>,
"Marek Szyprowski" <m.szyprowski@samsung.com>,
"Robin Murphy" <robin.murphy@arm.com>,
"Sumit Semwal" <sumit.semwal@linaro.org>,
"Benjamin Gaignard" <benjamin.gaignard@collabora.com>,
"Brian Starkey" <Brian.Starkey@arm.com>,
"John Stultz" <jstultz@google.com>,
"T.J. Mercier" <tjmercier@google.com>,
"Christian König" <christian.koenig@amd.com>,
"Steven Rostedt" <rostedt@goodmis.org>,
"Masami Hiramatsu" <mhiramat@kernel.org>,
"Mathieu Desnoyers" <mathieu.desnoyers@efficios.com>,
"Catalin Marinas" <catalin.marinas@arm.com>,
"Will Deacon" <will@kernel.org>, "Chun Ng" <chunn@nvidia.com>,
"Thierry Reding" <thierry.reding@gmail.com>,
devicetree@vger.kernel.org, linux-tegra@vger.kernel.org,
linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org,
linux-media@vger.kernel.org,
linux-arm-kernel@lists.infradead.org, linux-s390@vger.kernel.org,
linux-mm@kvack.org, iommu@lists.linux.dev,
linaro-mm-sig@lists.linaro.org,
linux-trace-kernel@vger.kernel.org,
"Thierry Reding" <treding@nvidia.com>,
pavan.kondeti@oss.qualcomm.com
Subject: Re: [PATCH v4 07/10] dma-buf: heaps: Add support for Tegra VPR
Date: Mon, 17 Aug 2026 18:22:13 +0100 [thread overview]
Message-ID: <aoNDRTKCUkzd2n5G@google.com> (raw)
In-Reply-To: <an8K2gb1pIqbDj6-@orome>
On Fri, Aug 14, 2026 at 02:56:05PM +0200, Thierry Reding wrote:
> On Thu, Aug 13, 2026 at 07:25:20PM +0100, Vincent Donnefort wrote:
> > On Fri, Aug 07, 2026 at 05:54:27PM +0200, Thierry Reding wrote:
> > > From: Thierry Reding <treding@nvidia.com>
> > >
> > > NVIDIA Tegra SoCs commonly define a Video-Protection-Region, which is a
> > > region of memory dedicated to content-protected video decode and
> > > playback. This memory cannot be accessed by the CPU and only certain
> > > hardware devices have access to it.
> > >
> > > Expose the VPR as a DMA heap so that applications and drivers can
> > > allocate buffers from this region for use-cases that require this kind
> > > of protected memory.
> > >
> > > VPR has a few very critical peculiarities. First, it must be a single
> > > contiguous region of memory (there is a single pair of registers that
> > > set the base address and size of the region), which is configured by
> > > calling back into the secure monitor. The memory region also needs to
> > > quite large for some use-cases because it needs to fit multiple video
> > > frames (8K video should be supported), so VPR sizes of ~2 GiB are
> > > expected. However, some devices cannot afford to reserve this amount
> > > of memory for a particular use-case, and therefore the VPR must be
> > > resizable.
> > >
> > > Unfortunately, resizing the VPR is slightly tricky because the GPU found
> > > on Tegra SoCs must be in reset during the VPR resize operation. This is
> > > currently implemented by freezing all userspace processes and calling
> > > invoking the GPU's freeze() implementation, resizing and the thawing the
> > > GPU and userspace processes. This is quite heavy-handed, so eventually
> > > it might be better to implement thawing/freezing in the GPU driver in
> > > such a way that they block accesses to the GPU so that the VPR resize
> > > operation can happen without suspending all userspace.
> > >
> > > In order to balance the memory usage versus the amount of resizing that
> > > needs to happen, the VPR is divided into multiple chunks. Each chunk is
> > > implemented as a CMA area that is completely allocated on first use to
> > > guarantee the contiguity of the VPR. Once all buffers from a chunk have
> > > been freed, the CMA area is deallocated and the memory returned to the
> > > system.
> >
> > Hi,
> >
> > I believe we (the Android team) are trying to solve similar issue to yours: Arm
> > CPUs can still speculatively read memory after it has been transitioned to the
> > Secure state, as long as they retain a cacheable mapping to it.
> >
> > As modifying the direct mapping is difficult, the workaround ended up in the
> > hypervisor which unmaps the pages from the host stage-2 on intercepted FF-A Lend
> > invocations.
> >
> > If convenient, this is nonetheless the wrong place to do it. Scattering the host
> > stage-2 is really terrible for performance and we would like to move it where it
> > should be, directly into the kernel...
> >
> > This is what I thought was the attempt in v3, but I am now a bit confused
> > because I see you are using set_direct_map_invalid_noflush()
> > set_direct_map_default_noflush(), but I am not sure that works if rodata=full is
> > not set? So does the memory for the NVIDIA IP still need to be unmapped?
>
> Yeah, this currently relies on the circumstances being such that
> can_set_direct_map() returns true, and in the case where we want to use
> the resizable VPR functionality, we're going to have page-granular
> mappings anyway.
>
> > If so, how about having an option where the CMA allocation is backed by a direct
> > map with the same page granularity, or at least a smaller and aligned granule?
> >
> > With that, we know that for whatever CMA allocation we do, we can safely unmap
> > from the direct map without risking splitting blocks. On CMA free, we can safely
> > remap into the direct map as I do not believe we coalesce yet.
>
> This sounds intriguing. For VPR we could possibly make the size a
> multiple of the memblock size (or whatever might be appropriate). If we
> can create a page-granular linear mapping specifically for that region,
> that'd be ideal. I don't know if the linear mapping can be subdivided in
> this fashion, though.
Actually I don't think modifying the memblock is necessary at all!
Here's what I have so far:
diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c
index d4de88770ecf..a9f81413b70d 100644
--- a/arch/arm64/mm/mmu.c
+++ b/arch/arm64/mm/mmu.c
@@ -22,6 +22,7 @@
#include <linux/fs.h>
#include <linux/io.h>
#include <linux/mm.h>
+#include <linux/of_fdt.h>
#include <linux/vmalloc.h>
#include <linux/set_memory.h>
#include <linux/kfence.h>
@@ -1138,6 +1139,27 @@ static inline void arm64_kfence_map_pool(void) { }
#endif /* CONFIG_KFENCE */
+#define MAX_FORCE_PTE_REGIONS 64 /* Greater or equal to MAX_RESERVED_REGIONS */
+
+static struct {
+ phys_addr_t start;
+ phys_addr_t end;
+} force_pte_regions[MAX_FORCE_PTE_REGIONS] __initdata;
+
+void __init early_init_dt_force_pte_arch(phys_addr_t start, phys_addr_t end)
+{
+ static int count;
+
+ if (force_pte_mapping())
+ return;
+
+ if (count >= MAX_FORCE_PTE_REGIONS)
+ return;
+
+ force_pte_regions[count].start = start;
+ force_pte_regions[count++].end = end;
+}
+
static void __init map_mem(void)
{
static const u64 direct_map_end = _PAGE_END(VA_BITS_MIN);
@@ -1186,6 +1208,15 @@ static void __init map_mem(void)
__map_memblock(init_end, kernel_end, pgprot_tagged(PAGE_KERNEL),
flags);
+ for (i = 0; i < ARRAY_SIZE(force_pte_regions); i++) {
+ if (!force_pte_regions[i].end)
+ break;
+
+ __map_memblock(force_pte_regions[i].start, force_pte_regions[i].end,
+ pgprot_tagged(PAGE_KERNEL),
+ flags | NO_BLOCK_MAPPINGS | NO_CONT_MAPPINGS);
+ }
+
/* map all the memory banks */
for_each_mem_range(i, &start, &end) {
/*
diff --git a/drivers/of/of_reserved_mem.c b/drivers/of/of_reserved_mem.c
index 42e3e2d8a2b8..445f204f01e4 100644
--- a/drivers/of/of_reserved_mem.c
+++ b/drivers/of/of_reserved_mem.c
@@ -136,6 +136,8 @@ static int __init early_init_dt_reserve_memory(phys_addr_t base,
return memblock_reserve(base, size);
}
+void __weak early_init_dt_force_pte_arch(phys_addr_t start, phys_addr_t end) { }
+
/*
* __reserved_mem_reserve_reg() - reserve memory described in the
* first entry in 'reg' property
@@ -168,6 +170,9 @@ static int __init __reserved_mem_reserve_reg(unsigned long node,
size = s;
if (size && early_init_dt_reserve_memory(base, size, nomap) == 0) {
+ if (of_get_flat_dt_prop(node, "force-pte", NULL))
+ early_init_dt_force_pte_arch(base, base + size);
+
fdt_fixup_reserved_mem_node(node, base, size);
pr_debug("Reserved memory: reserved region for node '%s': base %pa, size %lu MiB\n",
uname, &base, (unsigned long)(size / SZ_1M));
@@ -517,6 +522,9 @@ static int __init __reserved_mem_alloc_size(unsigned long node, const char *unam
return -ENOMEM;
}
+ if (of_get_flat_dt_prop(node, "force-pte", NULL))
+ early_init_dt_force_pte_arch(base, base + size);
+
fdt_fixup_reserved_mem_node(node, base, size);
fdt_init_reserved_mem_node(node, uname, base, size);
diff --git a/include/linux/of_fdt.h b/include/linux/of_fdt.h
index 51dadbaa3d63..6c73c76dee1b 100644
--- a/include/linux/of_fdt.h
+++ b/include/linux/of_fdt.h
@@ -74,6 +74,7 @@ extern void early_init_dt_check_for_usable_mem_range(void);
extern int early_init_dt_scan_chosen_stdout(void);
extern void early_init_fdt_scan_reserved_mem(void);
extern void early_init_fdt_reserve_self(void);
+extern void early_init_dt_force_pte_arch(phys_addr_t start, phys_addr_t end);
extern void early_init_dt_add_memory_arch(u64 base, u64 size);
extern u64 dt_mem_next_cell(int s, const __be32 **cellp);
>
> > (This alignment is not necessary for upcoming systems with BBML3)
> >
> > However, to implement this, we'd need to split memblocks and I believe the only
> > way at the moment is temporarily mark it as "nomap" which didn't seem very
> > popular in the comments on v3.
>
> Could this be simplified if this type of allocation is always memblock
> aligned?
>
> Looking at map_mem(), it seems like we could add some sort of special-
> casing in the for_each_mem_range() block to check if the memory is VPR
> (or generic, page-granular carveout, or whatever we want to call it) and
> set NO_BLOCK_MAPPINGS | NO_CONT_MAPPINGS in that case.
>
> If that works, set_direct_map_*() could be enhanced to detect such cases
> and always work. Or perhaps a more specific API could be introduced.
>
> Thierry
for set_direct_map() I think we could extend it to check if it is mapped at the
PTE-level and if it is we can proceed?
I am currently looking at extending CMA with an option "unmap-on-alloc;" that
would only be available if CONFIG_ARCH_HAS_SET_DIRECT_MAP, or
cma_set_unmap_on_alloc, or if "force-pte;" is set.
Hopefully I can share something this week.
--
Vincent
next prev parent reply other threads:[~2026-08-17 17:22 UTC|newest]
Thread overview: 30+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-07 15:54 [PATCH v4 00/10] dma-buf: heaps: Add support for Tegra VPR Thierry Reding
2026-08-07 15:54 ` [PATCH v4 01/10] dt-bindings: reserved-memory: Document " Thierry Reding
2026-08-07 16:08 ` sashiko-bot
2026-08-07 15:54 ` [PATCH v4 02/10] dt-bindings: display: tegra: Document memory regions Thierry Reding
2026-08-07 16:05 ` sashiko-bot
2026-08-12 22:33 ` Rob Herring (Arm)
2026-08-07 15:54 ` [PATCH v4 03/10] dt-bindings: gpu: host1x: Document memory-regions for NVDEC Thierry Reding
2026-08-07 16:05 ` sashiko-bot
2026-08-12 22:34 ` Rob Herring (Arm)
2026-08-07 15:54 ` [PATCH v4 04/10] bitmap: Add bitmap_allocate() function Thierry Reding
2026-08-07 16:11 ` sashiko-bot
2026-08-07 15:54 ` [PATCH v4 05/10] mm/cma: Allow dynamically creating CMA areas Thierry Reding
2026-08-07 16:15 ` sashiko-bot
2026-08-07 16:16 ` David Hildenbrand (Arm)
2026-08-10 14:14 ` Thierry Reding
2026-08-10 19:06 ` David Hildenbrand (Arm)
2026-08-07 15:54 ` [PATCH v4 06/10] dma-buf: heaps: Add debugfs support Thierry Reding
2026-08-07 16:19 ` sashiko-bot
2026-08-10 12:25 ` Christian König
2026-08-07 15:54 ` [PATCH v4 07/10] dma-buf: heaps: Add support for Tegra VPR Thierry Reding
2026-08-07 16:12 ` sashiko-bot
2026-08-13 18:25 ` Vincent Donnefort
2026-08-14 12:56 ` Thierry Reding
2026-08-17 17:22 ` Vincent Donnefort [this message]
2026-08-07 15:54 ` [PATCH v4 08/10] arm64: tegra: Add VPR placeholder node on Tegra234 Thierry Reding
2026-08-07 16:06 ` sashiko-bot
2026-08-07 15:54 ` [PATCH v4 09/10] arm64: tegra: Hook up VPR to host1x Thierry Reding
2026-08-07 16:16 ` sashiko-bot
2026-08-07 15:54 ` [PATCH v4 10/10] arm64: tegra: Add VPR placeholder node on Tegra264 Thierry Reding
2026-08-07 16:09 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aoNDRTKCUkzd2n5G@google.com \
--to=vdonnefort@google.com \
--cc=Brian.Starkey@arm.com \
--cc=agordeev@linux.ibm.com \
--cc=airlied@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=benjamin.gaignard@collabora.com \
--cc=borntraeger@linux.ibm.com \
--cc=catalin.marinas@arm.com \
--cc=christian.koenig@amd.com \
--cc=chunn@nvidia.com \
--cc=conor+dt@kernel.org \
--cc=david@kernel.org \
--cc=devicetree@vger.kernel.org \
--cc=dri-devel@lists.freedesktop.org \
--cc=gerald.schaefer@linux.ibm.com \
--cc=gor@linux.ibm.com \
--cc=hca@linux.ibm.com \
--cc=iommu@lists.linux.dev \
--cc=jonathanh@nvidia.com \
--cc=jstultz@google.com \
--cc=krzk+dt@kernel.org \
--cc=liam@infradead.org \
--cc=linaro-mm-sig@lists.linaro.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-media@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=linux-s390@vger.kernel.org \
--cc=linux-tegra@vger.kernel.org \
--cc=linux-trace-kernel@vger.kernel.org \
--cc=linux@armlinux.org.uk \
--cc=linux@rasmusvillemoes.dk \
--cc=ljs@kernel.org \
--cc=luca.ceresoli@bootlin.com \
--cc=m.szyprowski@samsung.com \
--cc=maarten.lankhorst@linux.intel.com \
--cc=mathieu.desnoyers@efficios.com \
--cc=mhiramat@kernel.org \
--cc=mhocko@suse.com \
--cc=mperttunen@nvidia.com \
--cc=mripard@kernel.org \
--cc=pavan.kondeti@oss.qualcomm.com \
--cc=robh@kernel.org \
--cc=robin.murphy@arm.com \
--cc=rostedt@goodmis.org \
--cc=rppt@kernel.org \
--cc=simona@ffwll.ch \
--cc=skomatineni@nvidia.com \
--cc=sumit.semwal@linaro.org \
--cc=surenb@google.com \
--cc=svens@linux.ibm.com \
--cc=thierry.reding@gmail.com \
--cc=thierry.reding@kernel.org \
--cc=tjmercier@google.com \
--cc=treding@nvidia.com \
--cc=tzimmermann@suse.de \
--cc=vbabka@kernel.org \
--cc=will@kernel.org \
--cc=yury.norov@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.