All of lore.kernel.org
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Lance Yang <lance.yang@linux.dev>,
	pjw@kernel.org, palmer@dabbelt.com, aou@eecs.berkeley.edu
Cc: alex@ghiti.fr, akpm@linux-foundation.org,
	zhangchunyan@iscas.ac.cn, rppt@kernel.org, kas@kernel.org,
	andrew+kernel@donnellan.id.au, rmclure@linux.ibm.com,
	debug@rivosinc.com, baolin.wang@linux.alibaba.com,
	usama.arif@linux.dev, wangruikang@iscas.ac.cn,
	namcao@linutronix.de, linux-riscv@lists.infradead.org,
	linux-kernel@vger.kernel.org, alexghiti@rivosinc.com,
	viro@zeniv.linux.org.uk, ajones@ventanamicro.com, arnd@arndb.de,
	axelrasmussen@google.com, brauner@kernel.org,
	conor.dooley@microchip.com, conor@kernel.org, jack@suse.cz,
	liam@infradead.org, ljs@kernel.org, mhocko@suse.com,
	paul.walmsley@sifive.com, peterx@redhat.com, robh@kernel.org,
	surenb@google.com, vbabka@kernel.org, yuanchu@google.com,
	stable@vger.kernel.org, pasha.tatashin@soleen.com,
	linux-mm@kvack.org, me@ziyao.cc
Subject: Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
Date: Mon, 5 Oct 2026 12:37:20 +0200	[thread overview]
Message-ID: <fabd52b6-36a7-442e-9cfb-47a39e4ae3d6@kernel.org> (raw)
In-Reply-To: <20261004030312.90163-1-lance.yang@linux.dev>

On 10/4/26 05:03, Lance Yang wrote:
> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
> a problem for PMD migration entries, since pmd_present() also checks
> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
> cleared during splitting.

I'm curious: why do we have to set leaf indications for non-present things? The
HW sure will ignore it, right?

Is this a sw problem? Who needs that?

> 
> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
> migration PMD as a present THP! The fault handler skips
> pmd_migration_entry_wait(), and a write fault can end up in
> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
> as a mapped PFN.

That sounds bad.

> 
> Move the swap soft-dirty bit to bit 12 and start the swap offset at
> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
> in migration PMDs and lets us keep the existing pmd_present() check
> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
> tracking is disabled.

That reduces the effective swap size (and PFN we can store). Could that be a
problem?

> 
> Fixes: 2a3ebad4db63 ("riscv: mm: add soft-dirty page tracking support")
> Cc: stable@vger.kernel.org
> Signed-off-by: Lance Yang <lance.yang@linux.dev>
> ---
> I found this while reviewing Usama's PMD-level swap entries series [1]
> and comparing the swap flag bits with the PMD presence checks across
> architectures.
> 
> [1] https://lore.kernel.org/all/20261002095503.3585565-1-usama.arif@linux.dev/
> 
>  arch/riscv/include/asm/pgtable-bits.h |  6 +++---
>  arch/riscv/include/asm/pgtable.h      | 10 ++++++----
>  2 files changed, 9 insertions(+), 7 deletions(-)
> 
> diff --git a/arch/riscv/include/asm/pgtable-bits.h b/arch/riscv/include/asm/pgtable-bits.h
> index d5a86b4df3ce6..d84b459e010c1 100644
> --- a/arch/riscv/include/asm/pgtable-bits.h
> +++ b/arch/riscv/include/asm/pgtable-bits.h
> @@ -27,12 +27,12 @@
>  	((riscv_has_extension_unlikely(RISCV_ISA_EXT_SVRSW60T59B)) ?	\
>  	 (1UL << 59) : 0)
>  /*
> - * Bit 3 is always zero for swap entry computation, so we
> - * can borrow it for swap page soft-dirty tracking.
> + * Bit 12 is reserved for swap soft-dirty tracking. The swap offset
> + * starts at bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled.
>   */
>  #define _PAGE_SWP_SOFT_DIRTY						\
>  	((riscv_has_extension_unlikely(RISCV_ISA_EXT_SVRSW60T59B)) ?	\
> -	 _PAGE_EXEC : 0)
> +	 (1UL << 12) : 0)
>  #else
>  #define _PAGE_SOFT_DIRTY	0
>  #define _PAGE_SWP_SOFT_DIRTY	0
> diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
> index b644db16bda94..acaf6d3b54d04 100644
> --- a/arch/riscv/include/asm/pgtable.h
> +++ b/arch/riscv/include/asm/pgtable.h
> @@ -1178,18 +1178,20 @@ static inline pud_t pud_modify(pud_t pud, pgprot_t newprot)
>   *
>   * Format of swap PTE:
>   *	bit            0:	_PAGE_PRESENT (zero)
> - *	bit       1 to 2:	(zero)
> - *	bit            3:	_PAGE_SWP_SOFT_DIRTY
> + *	bit       1 to 3:	_PAGE_LEAF (zero)
>   *	bit            4:	_PAGE_SWP_UFFD
>   *	bit            5:	_PAGE_PROT_NONE (zero)
>   *	bit            6:	exclusive marker
>   *	bits      7 to 11:	swap type
> - *	bits 12 to XLEN-1:	swap offset
> + *	bit           12:	_PAGE_SWP_SOFT_DIRTY (CONFIG_MEM_SOFT_DIRTY)
> + *	bits 12/13 to XLEN-1:	swap offset (without/with CONFIG_MEM_SOFT_DIRTY)
>   */
>  #define __SWP_TYPE_SHIFT	7
>  #define __SWP_TYPE_BITS		5
>  #define __SWP_TYPE_MASK		((1UL << __SWP_TYPE_BITS) - 1)
> -#define __SWP_OFFSET_SHIFT	(__SWP_TYPE_BITS + __SWP_TYPE_SHIFT)
> +#define __SWP_OFFSET_SHIFT \
> +	(__SWP_TYPE_BITS + __SWP_TYPE_SHIFT + \
> +	 IS_ENABLED(CONFIG_MEM_SOFT_DIRTY))
>  
>  #define MAX_SWAPFILES_CHECK()	\
>  	BUILD_BUG_ON(MAX_SWAPFILES_SHIFT > __SWP_TYPE_BITS)


-- 
Cheers,

David

_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv

WARNING: multiple messages have this Message-ID (diff)
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: Lance Yang <lance.yang@linux.dev>,
	pjw@kernel.org, palmer@dabbelt.com, aou@eecs.berkeley.edu
Cc: alex@ghiti.fr, akpm@linux-foundation.org,
	zhangchunyan@iscas.ac.cn, rppt@kernel.org, kas@kernel.org,
	andrew+kernel@donnellan.id.au, rmclure@linux.ibm.com,
	debug@rivosinc.com, baolin.wang@linux.alibaba.com,
	usama.arif@linux.dev, wangruikang@iscas.ac.cn,
	namcao@linutronix.de, linux-riscv@lists.infradead.org,
	linux-kernel@vger.kernel.org, alexghiti@rivosinc.com,
	viro@zeniv.linux.org.uk, ajones@ventanamicro.com, arnd@arndb.de,
	axelrasmussen@google.com, brauner@kernel.org,
	conor.dooley@microchip.com, conor@kernel.org, jack@suse.cz,
	liam@infradead.org, ljs@kernel.org, mhocko@suse.com,
	paul.walmsley@sifive.com, peterx@redhat.com, robh@kernel.org,
	surenb@google.com, vbabka@kernel.org, yuanchu@google.com,
	stable@vger.kernel.org, pasha.tatashin@soleen.com,
	linux-mm@kvack.org, me@ziyao.cc
Subject: Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
Date: Mon, 5 Oct 2026 12:37:20 +0200	[thread overview]
Message-ID: <fabd52b6-36a7-442e-9cfb-47a39e4ae3d6@kernel.org> (raw)
In-Reply-To: <20261004030312.90163-1-lance.yang@linux.dev>

On 10/4/26 05:03, Lance Yang wrote:
> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
> a problem for PMD migration entries, since pmd_present() also checks
> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
> cleared during splitting.

I'm curious: why do we have to set leaf indications for non-present things? The
HW sure will ignore it, right?

Is this a sw problem? Who needs that?

> 
> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
> migration PMD as a present THP! The fault handler skips
> pmd_migration_entry_wait(), and a write fault can end up in
> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
> as a mapped PFN.

That sounds bad.

> 
> Move the swap soft-dirty bit to bit 12 and start the swap offset at
> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
> in migration PMDs and lets us keep the existing pmd_present() check
> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
> tracking is disabled.

That reduces the effective swap size (and PFN we can store). Could that be a
problem?

> 
> Fixes: 2a3ebad4db63 ("riscv: mm: add soft-dirty page tracking support")
> Cc: stable@vger.kernel.org
> Signed-off-by: Lance Yang <lance.yang@linux.dev>
> ---
> I found this while reviewing Usama's PMD-level swap entries series [1]
> and comparing the swap flag bits with the PMD presence checks across
> architectures.
> 
> [1] https://lore.kernel.org/all/20261002095503.3585565-1-usama.arif@linux.dev/
> 
>  arch/riscv/include/asm/pgtable-bits.h |  6 +++---
>  arch/riscv/include/asm/pgtable.h      | 10 ++++++----
>  2 files changed, 9 insertions(+), 7 deletions(-)
> 
> diff --git a/arch/riscv/include/asm/pgtable-bits.h b/arch/riscv/include/asm/pgtable-bits.h
> index d5a86b4df3ce6..d84b459e010c1 100644
> --- a/arch/riscv/include/asm/pgtable-bits.h
> +++ b/arch/riscv/include/asm/pgtable-bits.h
> @@ -27,12 +27,12 @@
>  	((riscv_has_extension_unlikely(RISCV_ISA_EXT_SVRSW60T59B)) ?	\
>  	 (1UL << 59) : 0)
>  /*
> - * Bit 3 is always zero for swap entry computation, so we
> - * can borrow it for swap page soft-dirty tracking.
> + * Bit 12 is reserved for swap soft-dirty tracking. The swap offset
> + * starts at bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled.
>   */
>  #define _PAGE_SWP_SOFT_DIRTY						\
>  	((riscv_has_extension_unlikely(RISCV_ISA_EXT_SVRSW60T59B)) ?	\
> -	 _PAGE_EXEC : 0)
> +	 (1UL << 12) : 0)
>  #else
>  #define _PAGE_SOFT_DIRTY	0
>  #define _PAGE_SWP_SOFT_DIRTY	0
> diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
> index b644db16bda94..acaf6d3b54d04 100644
> --- a/arch/riscv/include/asm/pgtable.h
> +++ b/arch/riscv/include/asm/pgtable.h
> @@ -1178,18 +1178,20 @@ static inline pud_t pud_modify(pud_t pud, pgprot_t newprot)
>   *
>   * Format of swap PTE:
>   *	bit            0:	_PAGE_PRESENT (zero)
> - *	bit       1 to 2:	(zero)
> - *	bit            3:	_PAGE_SWP_SOFT_DIRTY
> + *	bit       1 to 3:	_PAGE_LEAF (zero)
>   *	bit            4:	_PAGE_SWP_UFFD
>   *	bit            5:	_PAGE_PROT_NONE (zero)
>   *	bit            6:	exclusive marker
>   *	bits      7 to 11:	swap type
> - *	bits 12 to XLEN-1:	swap offset
> + *	bit           12:	_PAGE_SWP_SOFT_DIRTY (CONFIG_MEM_SOFT_DIRTY)
> + *	bits 12/13 to XLEN-1:	swap offset (without/with CONFIG_MEM_SOFT_DIRTY)
>   */
>  #define __SWP_TYPE_SHIFT	7
>  #define __SWP_TYPE_BITS		5
>  #define __SWP_TYPE_MASK		((1UL << __SWP_TYPE_BITS) - 1)
> -#define __SWP_OFFSET_SHIFT	(__SWP_TYPE_BITS + __SWP_TYPE_SHIFT)
> +#define __SWP_OFFSET_SHIFT \
> +	(__SWP_TYPE_BITS + __SWP_TYPE_SHIFT + \
> +	 IS_ENABLED(CONFIG_MEM_SOFT_DIRTY))
>  
>  #define MAX_SWAPFILES_CHECK()	\
>  	BUILD_BUG_ON(MAX_SWAPFILES_SHIFT > __SWP_TYPE_BITS)


-- 
Cheers,

David


  reply	other threads:[~2026-10-05 10:37 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-04  3:03 [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present Lance Yang
2026-10-04  3:03 ` Lance Yang
2026-10-05 10:37 ` David Hildenbrand (Arm) [this message]
2026-10-05 10:37   ` David Hildenbrand (Arm)
2026-10-05 13:46   ` Lance Yang
2026-10-05 13:46     ` Lance Yang
2026-10-05 14:23     ` Lance Yang
2026-10-05 14:23       ` Lance Yang
2026-10-07 11:01     ` David Hildenbrand (Arm)
2026-10-07 11:01       ` David Hildenbrand (Arm)
2026-10-07 11:33       ` Lance Yang
2026-10-07 11:33         ` Lance Yang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=fabd52b6-36a7-442e-9cfb-47a39e4ae3d6@kernel.org \
    --to=david@kernel.org \
    --cc=ajones@ventanamicro.com \
    --cc=akpm@linux-foundation.org \
    --cc=alex@ghiti.fr \
    --cc=alexghiti@rivosinc.com \
    --cc=andrew+kernel@donnellan.id.au \
    --cc=aou@eecs.berkeley.edu \
    --cc=arnd@arndb.de \
    --cc=axelrasmussen@google.com \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=brauner@kernel.org \
    --cc=conor.dooley@microchip.com \
    --cc=conor@kernel.org \
    --cc=debug@rivosinc.com \
    --cc=jack@suse.cz \
    --cc=kas@kernel.org \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-riscv@lists.infradead.org \
    --cc=ljs@kernel.org \
    --cc=me@ziyao.cc \
    --cc=mhocko@suse.com \
    --cc=namcao@linutronix.de \
    --cc=palmer@dabbelt.com \
    --cc=pasha.tatashin@soleen.com \
    --cc=paul.walmsley@sifive.com \
    --cc=peterx@redhat.com \
    --cc=pjw@kernel.org \
    --cc=rmclure@linux.ibm.com \
    --cc=robh@kernel.org \
    --cc=rppt@kernel.org \
    --cc=stable@vger.kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    --cc=wangruikang@iscas.ac.cn \
    --cc=yuanchu@google.com \
    --cc=zhangchunyan@iscas.ac.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.