* Re: [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values [not found] <20260727164715.2866609-1-usama.anjum@arm.com> @ 2026-07-28 19:26 ` David Hildenbrand (Arm) 2026-07-29 10:13 ` Alexander Gordeev [not found] ` <20260727164715.2866609-10-usama.anjum@arm.com> 1 sibling, 1 reply; 9+ messages in thread From: David Hildenbrand (Arm) @ 2026-07-28 19:26 UTC (permalink / raw) To: Muhammad Usama Anjum, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, pfalcato, agordeev, ryan.roberts Cc: linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On 7/27/26 18:46, Muhammad Usama Anjum wrote: > Hi, > > pte_t currently describes both a logical PTE value and an element stored in > a PTE table. Consequently, pte_t * can point either to a standalone value, > often a stack copy, or to a PTE-table slot. The compiler cannot distinguish > these cases. A value pointer can therefore be passed to an interface that > expects table storage, while table storage can be read by direct > dereference instead of the architecture accessor. > > This series begins a staged conversion at the PTE level. It introduces > hw_pte_t as the element type for PTE-table storage and converts generic MM > to use hw_pte_t *. Logical PTE values remain pte_t. Interfaces that > intentionally return a value through pte_t *, such as install_pte, remain > value interfaces; the relevant parameters are named ptentp to make that > distinction explicit. > > The generic definition aliases hw_pte_t to pte_t, so this series preserves > the representation and behaviour of every architecture. ptep_get() keeps > its existing READ_ONCE() semantics and converts the stored element through > __pte_from_hw(). An architecture can later define a distinct hw_pte_t and > convert its PTE interfaces to make the distinction compiler-enforced. > Architecture PTE implementations and most architecture code are > deliberately left for those later opt-in conversions. > > Here, hw_pte_t identifies PTE-table storage rather than table lifetime: > complete PTE tables use hw_pte_t whether or not they are currently linked > into a page-table hierarchy, while standalone copied values use pte_t. The > distinction between complete but unlinked tables and hardware-reachable > tables was raised during discussion and remains an important point for > review. > > PMD, PUD, P4D and PGD storage are deliberately out of scope. They can be > converted in later series after the PTE boundary is agreed, avoiding the > PMD-specific cases that made an all-level conversion difficult to review. > > Most mechanical pointer conversions were generated with the Coccinelle > script included below, then audited and fixed by hand. > > This series does not add a second ptep_get_once() accessor and does not > remove or replace STRICT_MM_TYPECHECKS. Do you have a pointer at the arm64 part, so people can get a feeling for how an actual hw_pte_t implementation can look like. -- Cheers, David ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values 2026-07-28 19:26 ` [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values David Hildenbrand (Arm) @ 2026-07-29 10:13 ` Alexander Gordeev 2026-07-29 11:05 ` David Hildenbrand (Arm) 0 siblings, 1 reply; 9+ messages in thread From: Alexander Gordeev @ 2026-07-29 10:13 UTC (permalink / raw) To: David Hildenbrand (Arm) Cc: Muhammad Usama Anjum, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, pfalcato, ryan.roberts, linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On Tue, Jul 28, 2026 at 09:26:06PM +0200, David Hildenbrand (Arm) wrote: > > The generic definition aliases hw_pte_t to pte_t, so this series preserves > > the representation and behaviour of every architecture. ptep_get() keeps > > its existing READ_ONCE() semantics and converts the stored element through > > __pte_from_hw(). An architecture can later define a distinct hw_pte_t and > > convert its PTE interfaces to make the distinction compiler-enforced. > > Architecture PTE implementations and most architecture code are > > deliberately left for those later opt-in conversions. ... > Do you have a pointer at the arm64 part, so people can get a feeling for how an > actual hw_pte_t implementation can look like. May be we need the generic hw_pte_t implementation as { pte_t pte; } right away? Also, can you envision a case when sizeof(pte_t) != sizeof(hw_pte_t)? If not, the compile-time check is worth adding. > -- > Cheers, > > David Thanks! ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values 2026-07-29 10:13 ` Alexander Gordeev @ 2026-07-29 11:05 ` David Hildenbrand (Arm) 2026-07-29 11:44 ` Alexander Gordeev [not found] ` <46319967-0338-4532-b953-dbea77084fe3@arm.com> 0 siblings, 2 replies; 9+ messages in thread From: David Hildenbrand (Arm) @ 2026-07-29 11:05 UTC (permalink / raw) To: Alexander Gordeev Cc: Muhammad Usama Anjum, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, pfalcato, ryan.roberts, linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On 7/29/26 12:13, Alexander Gordeev wrote: > On Tue, Jul 28, 2026 at 09:26:06PM +0200, David Hildenbrand (Arm) wrote: >>> The generic definition aliases hw_pte_t to pte_t, so this series preserves >>> the representation and behaviour of every architecture. ptep_get() keeps >>> its existing READ_ONCE() semantics and converts the stored element through >>> __pte_from_hw(). An architecture can later define a distinct hw_pte_t and >>> convert its PTE interfaces to make the distinction compiler-enforced. >>> Architecture PTE implementations and most architecture code are >>> deliberately left for those later opt-in conversions. > ... >> Do you have a pointer at the arm64 part, so people can get a feeling for how an >> actual hw_pte_t implementation can look like. > > May be we need the generic hw_pte_t implementation as { pte_t pte; } > right away? Indeed, that makes sense. We just need a way for the architecture to opt-in that it did the conversion. > > Also, can you envision a case when sizeof(pte_t) != sizeof(hw_pte_t)? > If not, the compile-time check is worth adding. That wouldn't work as is. We'd have to intercept ptep++ and instead have a helper to advance the ptep pointer. One idea for that would be to let the compiler catch that by marking hw_pte_t an undefined struct (unknown size), such that the compiler would indicate any usage of ptep++ properly. That can be done when it would actually required, so as a first step having a generic hw_pte_t would just work. -- Cheers, David ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values 2026-07-29 11:05 ` David Hildenbrand (Arm) @ 2026-07-29 11:44 ` Alexander Gordeev 2026-07-29 11:51 ` David Hildenbrand (Arm) [not found] ` <46319967-0338-4532-b953-dbea77084fe3@arm.com> 1 sibling, 1 reply; 9+ messages in thread From: Alexander Gordeev @ 2026-07-29 11:44 UTC (permalink / raw) To: David Hildenbrand (Arm) Cc: Muhammad Usama Anjum, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, pfalcato, ryan.roberts, linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On Wed, Jul 29, 2026 at 01:05:35PM +0200, David Hildenbrand (Arm) wrote: > On 7/29/26 12:13, Alexander Gordeev wrote: > > On Tue, Jul 28, 2026 at 09:26:06PM +0200, David Hildenbrand (Arm) wrote: > >>> The generic definition aliases hw_pte_t to pte_t, so this series preserves > >>> the representation and behaviour of every architecture. ptep_get() keeps > >>> its existing READ_ONCE() semantics and converts the stored element through > >>> __pte_from_hw(). An architecture can later define a distinct hw_pte_t and > >>> convert its PTE interfaces to make the distinction compiler-enforced. > >>> Architecture PTE implementations and most architecture code are > >>> deliberately left for those later opt-in conversions. > > ... > >> Do you have a pointer at the arm64 part, so people can get a feeling for how an > >> actual hw_pte_t implementation can look like. > > > > May be we need the generic hw_pte_t implementation as { pte_t pte; } > > right away? > > Indeed, that makes sense. We just need a way for the architecture to opt-in that > it did the conversion. > > > > > Also, can you envision a case when sizeof(pte_t) != sizeof(hw_pte_t)? > > If not, the compile-time check is worth adding. > > That wouldn't work as is. We'd have to intercept ptep++ and instead have a > helper to advance the ptep pointer. I think I am missing the point. That is my takeaway from your previous mail: https://lore.kernel.org/lkml/31d36023-d728-4eee-90f8-158c7066f565@kernel.org/ <quote> > { > page_table_check_ptes_set(mm, addr, ptep, pte, nr); > for (;;) { > set_pte(ptep, pte); > if (--nr == 0) > break; > - ptep++; > + ptep = hw_pte_next(ptep); We should really just let ptep++ work as before. > pte = pte_next_pfn(pte); > } > } </quote> So are we going after hw_pte_next(ptep) or ptep++? > One idea for that would be to let the compiler catch that by marking hw_pte_t an > undefined struct (unknown size), such that the compiler would indicate any usage > of ptep++ properly. What is the purpose of that? I mean when pte_t vs hw_pte_t uses are sorted out what is the benefit of preventing hw_ptep++? > That can be done when it would actually required, so as a first step having a > generic hw_pte_t would just work. > > -- > Cheers, > > David Thanks! ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values 2026-07-29 11:44 ` Alexander Gordeev @ 2026-07-29 11:51 ` David Hildenbrand (Arm) 0 siblings, 0 replies; 9+ messages in thread From: David Hildenbrand (Arm) @ 2026-07-29 11:51 UTC (permalink / raw) To: Alexander Gordeev Cc: Muhammad Usama Anjum, Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, pfalcato, ryan.roberts, linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On 7/29/26 13:44, Alexander Gordeev wrote: > On Wed, Jul 29, 2026 at 01:05:35PM +0200, David Hildenbrand (Arm) wrote: >> On 7/29/26 12:13, Alexander Gordeev wrote: >>> ... >>> >>> May be we need the generic hw_pte_t implementation as { pte_t pte; } >>> right away? >> >> Indeed, that makes sense. We just need a way for the architecture to opt-in that >> it did the conversion. >> >>> >>> Also, can you envision a case when sizeof(pte_t) != sizeof(hw_pte_t)? >>> If not, the compile-time check is worth adding. >> >> That wouldn't work as is. We'd have to intercept ptep++ and instead have a >> helper to advance the ptep pointer. > > I think I am missing the point. That is my takeaway from your previous mail: > https://lore.kernel.org/lkml/31d36023-d728-4eee-90f8-158c7066f565@kernel.org/ > > <quote> >> { >> page_table_check_ptes_set(mm, addr, ptep, pte, nr); >> for (;;) { >> set_pte(ptep, pte); >> if (--nr == 0) >> break; >> - ptep++; >> + ptep = hw_pte_next(ptep); > > We should really just let ptep++ work as before. > >> pte = pte_next_pfn(pte); >> } >> } > </quote> > > So are we going after hw_pte_next(ptep) or ptep++? As I said "That wouldn't work as is. We'd have ". So in this series here we are clearly going for ptep++ and sizeof(pte_t) == sizeof(hw_pte_t). Because otherwise it wouldn't work. > >> One idea for that would be to let the compiler catch that by marking hw_pte_t an >> undefined struct (unknown size), such that the compiler would indicate any usage >> of ptep++ properly. > > What is the purpose of that? I mean when pte_t vs hw_pte_t uses are sorted > out what is the benefit of preventing hw_ptep++? One thing I could pull out of my magic hat is that you might be able to decide at runtime the size of your underlying page table entries. E.g., have 128bit pteval, but allow running on HW with either 64bit ptes or 128bit ptes. I'm sure there are more challenges to that, but that's an easy thing to imagine. But again, the focus of this patch set here is ptep++ to just keep working. -- Cheers, David ^ permalink raw reply [flat|nested] 9+ messages in thread
[parent not found: <46319967-0338-4532-b953-dbea77084fe3@arm.com>]
* Re: [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values [not found] ` <46319967-0338-4532-b953-dbea77084fe3@arm.com> @ 2026-07-29 11:52 ` David Hildenbrand (Arm) [not found] ` <38e14b3f-52b3-45d8-b151-776be75d44d8@arm.com> 0 siblings, 1 reply; 9+ messages in thread From: David Hildenbrand (Arm) @ 2026-07-29 11:52 UTC (permalink / raw) To: Muhammad Usama Anjum, Alexander Gordeev Cc: Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, pfalcato, ryan.roberts, linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On 7/29/26 13:33, Muhammad Usama Anjum wrote: > On 29/07/2026 12:05 pm, David Hildenbrand (Arm) wrote: >> On 7/29/26 12:13, Alexander Gordeev wrote: >>> ... >>> >>> May be we need the generic hw_pte_t implementation as { pte_t pte; } >>> right away? >> >> Indeed, that makes sense. We just need a way for the architecture to opt-in that >> it did the conversion. > Yeah and architecture cannot opt-in until its converted. So we should leave > it to architecture to define hw_pte_t. In the context of this series, we should have something like an CONFIG_ARCH_HAS_XXX and select the definition based on that. So we'd have a generic variant. -- Cheers, David ^ permalink raw reply [flat|nested] 9+ messages in thread
[parent not found: <38e14b3f-52b3-45d8-b151-776be75d44d8@arm.com>]
* Re: [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values [not found] ` <38e14b3f-52b3-45d8-b151-776be75d44d8@arm.com> @ 2026-07-29 12:28 ` David Hildenbrand (Arm) 0 siblings, 0 replies; 9+ messages in thread From: David Hildenbrand (Arm) @ 2026-07-29 12:28 UTC (permalink / raw) To: Muhammad Usama Anjum, Alexander Gordeev Cc: Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, pfalcato, ryan.roberts, linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On 7/29/26 14:21, Muhammad Usama Anjum wrote: > On 29/07/2026 12:52 pm, David Hildenbrand (Arm) wrote: >> On 7/29/26 13:33, Muhammad Usama Anjum wrote: >>> Yeah and architecture cannot opt-in until its converted. So we should leave >>> it to architecture to define hw_pte_t. >> >> In the context of this series, we should have something like an >> CONFIG_ARCH_HAS_XXX and select the definition based on that. >> >> So we'd have a generic variant. > I'll add CONFIG_ARCH_HAS_XXX config in generic version. > > I don't understand why generic code would define hw_pte_t if arch optionally > opts-in. > > Are you saying pte_t may have different definition in different arches. But > the hw_pte_t would always be structure of pte_t. Hence definition > > typedef struct { pte_t __pte; }; hw_pte_t; > > can be moved to the generic code? > > Moving definition to generic side is fine at this point in time. But if > a arch wants to hw_pte_t opaque and don't want the generic code to perform > arithmatic (ptep++), it'll not work. Yes, but we can tackle this once some arch actually needs that. For now, it makes this patch set easier to digest. -- Cheers, David ^ permalink raw reply [flat|nested] 9+ messages in thread
[parent not found: <20260727164715.2866609-10-usama.anjum@arm.com>]
* Re: [PATCH RFC 09/11] misc/sgi-gru: use ptep_get() for page-table reads [not found] ` <20260727164715.2866609-10-usama.anjum@arm.com> @ 2026-07-29 12:36 ` Pedro Falcato 2026-07-29 12:42 ` David Hildenbrand (Arm) 0 siblings, 1 reply; 9+ messages in thread From: Pedro Falcato @ 2026-07-29 12:36 UTC (permalink / raw) To: Muhammad Usama Anjum Cc: Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, David Hildenbrand, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, agordeev, ryan.roberts, linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On Mon, Jul 27, 2026 at 05:47:00PM +0100, Muhammad Usama Anjum wrote: > A leaf PMD is being read through ptep_get() by treating the PMD address > as PTE-sized table storage. ptep_get() now accepts hw_pte_t *, so update > the cast accordingly. > > pte_offset_kernel() also returns hw_pte_t *. Get pte_t value by calling > ptep_get(). > > Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com> > --- > drivers/misc/sgi-gru/grufault.c | 4 ++-- > 1 file changed, 2 insertions(+), 2 deletions(-) > > diff --git a/drivers/misc/sgi-gru/grufault.c b/drivers/misc/sgi-gru/grufault.c > index 3557d78ee47a2..ff89d34ad2aa4 100644 > --- a/drivers/misc/sgi-gru/grufault.c > +++ b/drivers/misc/sgi-gru/grufault.c > @@ -228,10 +228,10 @@ static int atomic_pte_lookup(struct vm_area_struct *vma, unsigned long vaddr, > goto err; > #ifdef CONFIG_X86_64 > if (unlikely(pmd_leaf(*pmdp))) > - pte = ptep_get((pte_t *)pmdp); > + pte = ptep_get((hw_pte_t *)pmdp); > else > #endif > - pte = *pte_offset_kernel(pmdp, vaddr); > + pte = ptep_get(pte_offset_kernel(pmdp, vaddr)); > > if (unlikely(!pte_present(pte) || > (write && (!pte_write(pte) || !pte_dirty(pte))))) This code is super, super broken. Can we remove this ASAP? For starters, we're using is_vm_hugetlb_page() to detect page shift, the code does not grab refs on the pages, uses pte_offset_kernel() on user page tables, does not handle PUD-level hugepages, does not handle PMD-level hugepages on !x86_64, does not hold page table locks nor check for pte table retraction, etc I don't think I need to go on. -- Pedro ^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH RFC 09/11] misc/sgi-gru: use ptep_get() for page-table reads 2026-07-29 12:36 ` [PATCH RFC 09/11] misc/sgi-gru: use ptep_get() for page-table reads Pedro Falcato @ 2026-07-29 12:42 ` David Hildenbrand (Arm) 0 siblings, 0 replies; 9+ messages in thread From: David Hildenbrand (Arm) @ 2026-07-29 12:42 UTC (permalink / raw) To: Pedro Falcato, Muhammad Usama Anjum Cc: Jani Nikula, Joonas Lahtinen, Rodrigo Vivi, Tvrtko Ursulin, David Airlie, Simona Vetter, Dimitri Sivanich, Arnd Bergmann, Greg Kroah-Hartman, James E.J. Bottomley, Helge Deller, Juergen Gross, Stefano Stabellini, Muchun Song, Oscar Salvador, Andrew Morton, Liam R. Howlett, Lorenzo Stoakes, Will Deacon, Aneesh Kumar K.V, Nick Piggin, Peter Zijlstra, Andrey Ryabinin, Pasha Tatashin, Chris Li, Kairui Song, Uladzislau Rezki, Steven Rostedt, Masami Hiramatsu, Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko, Eduard Zingerman, Kumar Kartikeya Dwivedi, Ingo Molnar, Arnaldo Carvalho de Melo, Namhyung Kim, SJ Park, Matthew Wilcox (Oracle), Jan Kara, Jason Gunthorpe, Leon Romanovsky, Miaohe Lin, Dennis Zhou, Tejun Heo, Christoph Lameter, Mike Rapoport, Johannes Weiner, ziy, agordeev, ryan.roberts, linux-kernel, intel-gfx, dri-devel, linux-parisc, xen-devel, linux-mm, linux-fsdevel, linux-arch, kasan-dev, linux-trace-kernel, bpf, linux-perf-users, damon On 7/29/26 14:36, Pedro Falcato wrote: > On Mon, Jul 27, 2026 at 05:47:00PM +0100, Muhammad Usama Anjum wrote: >> A leaf PMD is being read through ptep_get() by treating the PMD address >> as PTE-sized table storage. ptep_get() now accepts hw_pte_t *, so update >> the cast accordingly. >> >> pte_offset_kernel() also returns hw_pte_t *. Get pte_t value by calling >> ptep_get(). >> >> Signed-off-by: Muhammad Usama Anjum <usama.anjum@arm.com> >> --- >> drivers/misc/sgi-gru/grufault.c | 4 ++-- >> 1 file changed, 2 insertions(+), 2 deletions(-) >> >> diff --git a/drivers/misc/sgi-gru/grufault.c b/drivers/misc/sgi-gru/grufault.c >> index 3557d78ee47a2..ff89d34ad2aa4 100644 >> --- a/drivers/misc/sgi-gru/grufault.c >> +++ b/drivers/misc/sgi-gru/grufault.c >> @@ -228,10 +228,10 @@ static int atomic_pte_lookup(struct vm_area_struct *vma, unsigned long vaddr, >> goto err; >> #ifdef CONFIG_X86_64 >> if (unlikely(pmd_leaf(*pmdp))) >> - pte = ptep_get((pte_t *)pmdp); >> + pte = ptep_get((hw_pte_t *)pmdp); >> else >> #endif >> - pte = *pte_offset_kernel(pmdp, vaddr); >> + pte = ptep_get(pte_offset_kernel(pmdp, vaddr)); >> >> if (unlikely(!pte_present(pte) || >> (write && (!pte_write(pte) || !pte_dirty(pte))))) > > This code is super, super broken. Can we remove this ASAP? For starters, > we're using is_vm_hugetlb_page() to detect page shift, the code does not > grab refs on the pages, uses pte_offset_kernel() on user page tables, > does not handle PUD-level hugepages, does not handle PMD-level hugepages on !x86_64, > does not hold page table locks nor check for pte table retraction, etc > > I don't think I need to go on. > Heh, when Usama first showed me an early diffstat I was like "what the hell are drivers doing with page tables here, they shouldn't be doing that". atomic_pte_lookup documents: "Only supports Intel large pages (2MB only) on x86_64. ZZZ - hugepage support is incomplete" What is this crap? :) -- Cheers, David ^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-07-29 12:42 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <20260727164715.2866609-1-usama.anjum@arm.com>
2026-07-28 19:26 ` [PATCH RFC 00/11] mm: distinguish PTE table storage from PTE values David Hildenbrand (Arm)
2026-07-29 10:13 ` Alexander Gordeev
2026-07-29 11:05 ` David Hildenbrand (Arm)
2026-07-29 11:44 ` Alexander Gordeev
2026-07-29 11:51 ` David Hildenbrand (Arm)
[not found] ` <46319967-0338-4532-b953-dbea77084fe3@arm.com>
2026-07-29 11:52 ` David Hildenbrand (Arm)
[not found] ` <38e14b3f-52b3-45d8-b151-776be75d44d8@arm.com>
2026-07-29 12:28 ` David Hildenbrand (Arm)
[not found] ` <20260727164715.2866609-10-usama.anjum@arm.com>
2026-07-29 12:36 ` [PATCH RFC 09/11] misc/sgi-gru: use ptep_get() for page-table reads Pedro Falcato
2026-07-29 12:42 ` David Hildenbrand (Arm)
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox