On Wed, Aug 19, 2026 at 10:10:09AM +0100, Vincent Donnefort wrote: > On Tue, Aug 18, 2026 at 12:48:09PM +0200, Thierry Reding wrote: > > On Mon, Aug 17, 2026 at 06:22:13PM +0100, Vincent Donnefort wrote: > > > On Fri, Aug 14, 2026 at 02:56:05PM +0200, Thierry Reding wrote: > > > > On Thu, Aug 13, 2026 at 07:25:20PM +0100, Vincent Donnefort wrote: > > > > > On Fri, Aug 07, 2026 at 05:54:27PM +0200, Thierry Reding wrote: > > > > > > From: Thierry Reding > > > > > > > > > > > > NVIDIA Tegra SoCs commonly define a Video-Protection-Region, which is a > > > > > > region of memory dedicated to content-protected video decode and > > > > > > playback. This memory cannot be accessed by the CPU and only certain > > > > > > hardware devices have access to it. > > > > > > > > > > > > Expose the VPR as a DMA heap so that applications and drivers can > > > > > > allocate buffers from this region for use-cases that require this kind > > > > > > of protected memory. > > > > > > > > > > > > VPR has a few very critical peculiarities. First, it must be a single > > > > > > contiguous region of memory (there is a single pair of registers that > > > > > > set the base address and size of the region), which is configured by > > > > > > calling back into the secure monitor. The memory region also needs to > > > > > > quite large for some use-cases because it needs to fit multiple video > > > > > > frames (8K video should be supported), so VPR sizes of ~2 GiB are > > > > > > expected. However, some devices cannot afford to reserve this amount > > > > > > of memory for a particular use-case, and therefore the VPR must be > > > > > > resizable. > > > > > > > > > > > > Unfortunately, resizing the VPR is slightly tricky because the GPU found > > > > > > on Tegra SoCs must be in reset during the VPR resize operation. This is > > > > > > currently implemented by freezing all userspace processes and calling > > > > > > invoking the GPU's freeze() implementation, resizing and the thawing the > > > > > > GPU and userspace processes. This is quite heavy-handed, so eventually > > > > > > it might be better to implement thawing/freezing in the GPU driver in > > > > > > such a way that they block accesses to the GPU so that the VPR resize > > > > > > operation can happen without suspending all userspace. > > > > > > > > > > > > In order to balance the memory usage versus the amount of resizing that > > > > > > needs to happen, the VPR is divided into multiple chunks. Each chunk is > > > > > > implemented as a CMA area that is completely allocated on first use to > > > > > > guarantee the contiguity of the VPR. Once all buffers from a chunk have > > > > > > been freed, the CMA area is deallocated and the memory returned to the > > > > > > system. > > > > > > > > > > Hi, > > > > > > > > > > I believe we (the Android team) are trying to solve similar issue to yours: Arm > > > > > CPUs can still speculatively read memory after it has been transitioned to the > > > > > Secure state, as long as they retain a cacheable mapping to it. > > > > > > > > > > As modifying the direct mapping is difficult, the workaround ended up in the > > > > > hypervisor which unmaps the pages from the host stage-2 on intercepted FF-A Lend > > > > > invocations. > > > > > > > > > > If convenient, this is nonetheless the wrong place to do it. Scattering the host > > > > > stage-2 is really terrible for performance and we would like to move it where it > > > > > should be, directly into the kernel... > > > > > > > > > > This is what I thought was the attempt in v3, but I am now a bit confused > > > > > because I see you are using set_direct_map_invalid_noflush() > > > > > set_direct_map_default_noflush(), but I am not sure that works if rodata=full is > > > > > not set? So does the memory for the NVIDIA IP still need to be unmapped? > > > > > > > > Yeah, this currently relies on the circumstances being such that > > > > can_set_direct_map() returns true, and in the case where we want to use > > > > the resizable VPR functionality, we're going to have page-granular > > > > mappings anyway. > > > > > > > > > If so, how about having an option where the CMA allocation is backed by a direct > > > > > map with the same page granularity, or at least a smaller and aligned granule? > > > > > > > > > > With that, we know that for whatever CMA allocation we do, we can safely unmap > > > > > from the direct map without risking splitting blocks. On CMA free, we can safely > > > > > remap into the direct map as I do not believe we coalesce yet. > > > > > > > > This sounds intriguing. For VPR we could possibly make the size a > > > > multiple of the memblock size (or whatever might be appropriate). If we > > > > can create a page-granular linear mapping specifically for that region, > > > > that'd be ideal. I don't know if the linear mapping can be subdivided in > > > > this fashion, though. > > > > > > Actually I don't think modifying the memblock is necessary at all! > > > > > > Here's what I have so far: > > > > > > diff --git a/arch/arm64/mm/mmu.c b/arch/arm64/mm/mmu.c > > > index d4de88770ecf..a9f81413b70d 100644 > > > --- a/arch/arm64/mm/mmu.c > > > +++ b/arch/arm64/mm/mmu.c > > > @@ -22,6 +22,7 @@ > > > #include > > > #include > > > #include > > > +#include > > > #include > > > #include > > > #include > > > @@ -1138,6 +1139,27 @@ static inline void arm64_kfence_map_pool(void) { } > > > > > > #endif /* CONFIG_KFENCE */ > > > > > > +#define MAX_FORCE_PTE_REGIONS 64 /* Greater or equal to MAX_RESERVED_REGIONS */ > > > + > > > +static struct { > > > + phys_addr_t start; > > > + phys_addr_t end; > > > +} force_pte_regions[MAX_FORCE_PTE_REGIONS] __initdata; > > > + > > > +void __init early_init_dt_force_pte_arch(phys_addr_t start, phys_addr_t end) > > > +{ > > > + static int count; > > > + > > > + if (force_pte_mapping()) > > > + return; > > > + > > > + if (count >= MAX_FORCE_PTE_REGIONS) > > > + return; > > > + > > > + force_pte_regions[count].start = start; > > > + force_pte_regions[count++].end = end; > > > +} > > > + > > > static void __init map_mem(void) > > > { > > > static const u64 direct_map_end = _PAGE_END(VA_BITS_MIN); > > > @@ -1186,6 +1208,15 @@ static void __init map_mem(void) > > > __map_memblock(init_end, kernel_end, pgprot_tagged(PAGE_KERNEL), > > > flags); > > > > > > + for (i = 0; i < ARRAY_SIZE(force_pte_regions); i++) { > > > + if (!force_pte_regions[i].end) > > > + break; > > > + > > > + __map_memblock(force_pte_regions[i].start, force_pte_regions[i].end, > > > + pgprot_tagged(PAGE_KERNEL), > > > + flags | NO_BLOCK_MAPPINGS | NO_CONT_MAPPINGS); > > > + } > > > > Interesting. I had been thinking more along the lines of adding this > > directly into memblock using a new entry in enum memblock_flags. None of > > the existing ones seem to do what we need here, though some are closely > > related. MEMBLOCK_SECURE or MEMBLOCK_PROTECTED are quite specific and > > don't necessarily mean that we need the mapping to be page-granular, so > > MEMBLOCK_PAGE_GRANULAR is perhaps clearer. > > I thought it would be interesting to not have to modify memblock since it is > used for all architectures, while the DT is at least slightly less widespread. This isn't something that is DT specific. On the Tegra side, I suspect we might need this on ACPI platforms eventually, too. And it's not an ARM specific feature either. Other platforms could use shared carveouts, too. > Also, cma_init_reserved_mem() could use an order 9. In that case PTE-level isn't > necessary and we could just use PMD-level. Extending that support would be more > cumbersome with a memblock flag than with a callback. > > rmem_cma_setup() hardcoding an order 0, perhaps this isn't really a problem at > the moment? How so? memblock is really just used as a way of backing the CMA. How CMA subdivides it doesn't really matter, right? Or are you suggesting that if we use an entire PMD as granularity, we could equally well remove the entire PMD from the linear mapping instead of doing it page-by-page? I think that'd work really well for at least VPR, since it is 1 MiB aligned anyway. Using a granularity of 2 MiB is easily doable. For anything that doesn't require a single contiguous area to be protected this might be a bit more challenging since it potentially wastes a lot of memory. On the other hand, a lot of this is heavily custom code anyway, so the entire stack could be modified to make efficient use of this (i.e. userspace could allocate a larger chunk for a pool of buffers, etc.). > But yeah, the alternative is to create a "memblock_mark_forcepte" (it seems > memblock_setclr_flag does split memblocks) and let of_reserved_mem call that > function. Finally the arm64 mmu code can simply check for the flag before > calling __map_memblock. Perhaps it isn't that bad in the end? It sounds like the right level of abstraction to me. But I'm not too familiar with this code, so it'd be good to hear from the MM and/or ARM maintainers what they think about this. [...] > I had in mind to extend "shared-dma-pool" to handle > set_direct_map_invalid_noflush()/set_direct_map_default_noflush() based on an > option. But perhaps it is better to create another separate driver. And VPR > needing a specific dma-heap driver anyway, it could call the direct-map > functions there too without relying on CMA to do anything? I think it'd be nice to have separate APIs for this case where we know the memory region is already page-granular (or, I suppose, PMD granular) and removing from (or adding back to) the linear mapping is safe. That way we could avoid the checks for can_set_direct_map() for each page. It'd also be nice to have a version that can update the protection bits for a range of pages (would map directly to update_range_prot()) instead of having to manually iterate over each page in a range. A good middle-ground might be to have helpers that do the grunt work and they can then be called from VPR and the shared-dma-pool drivers to have pages removed from the linear mapping. VPR and similar can then perform the hardware protection bits on top of that. Thierry