All of lore.kernel.org
 help / color / mirror / Atom feed
From: Lance Yang <lance.yang@linux.dev>
To: david@kernel.org
Cc: pjw@kernel.org, palmer@dabbelt.com, aou@eecs.berkeley.edu,
	alex@ghiti.fr, akpm@linux-foundation.org,
	zhangchunyan@iscas.ac.cn, rppt@kernel.org, kas@kernel.org,
	andrew+kernel@donnellan.id.au, rmclure@linux.ibm.com,
	debug@rivosinc.com, baolin.wang@linux.alibaba.com,
	usama.arif@linux.dev, wangruikang@iscas.ac.cn,
	namcao@linutronix.de, linux-riscv@lists.infradead.org,
	linux-kernel@vger.kernel.org, alexghiti@rivosinc.com,
	viro@zeniv.linux.org.uk, ajones@ventanamicro.com, arnd@arndb.de,
	axelrasmussen@google.com, brauner@kernel.org,
	conor.dooley@microchip.com, conor@kernel.org, jack@suse.cz,
	liam@infradead.org, ljs@kernel.org, mhocko@suse.com,
	paul.walmsley@sifive.com, peterx@redhat.com, robh@kernel.org,
	surenb@google.com, vbabka@kernel.org, yuanchu@google.com,
	stable@vger.kernel.org, pasha.tatashin@soleen.com,
	linux-mm@kvack.org, me@ziyao.cc
Subject: Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
Date: Mon, 5 Oct 2026 22:23:57 +0800	[thread overview]
Message-ID: <b928b8a4-5ec5-4642-ac02-9ecbbbfd6218@linux.dev> (raw)
In-Reply-To: <20261005134641.5801-1-lance.yang@linux.dev>



On 2026/10/5 21:46, Lance Yang wrote:
> 
> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/4/26 05:03, Lance Yang wrote:
>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>> a problem for PMD migration entries, since pmd_present() also checks
>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>> cleared during splitting.
>>
>> I'm curious: why do we have to set leaf indications for non-present things? The
>> HW sure will ignore it, right?
> 
> Yeah, that surprised me too :) Still wrapping my head around the details
> ...
> 
>> Is this a sw problem? Who needs that?
> 
> IIUC, it's for software during a PMD split.
> 
> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
> before installing the PTE table. Software still needs pmd_present() and
> pmd_trans_huge() to recognize the THP in between.

Also, x86 keeps _PAGE_PSE to identify the huge PMD. RISC-V doesn't
have a separate leaf bit, so it keeps the R/W/X bits to identify
the PMD as a leaf entry :)

> 
> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
> to hardware but still identifiable as a THP by software.
> 
> Hopefully I didn't miss something.
> 
>>>
>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>> migration PMD as a present THP! The fault handler skips
>>> pmd_migration_entry_wait(), and a write fault can end up in
>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>> as a mapped PFN.
>>
>> That sounds bad.
> 
> YES, looks a bit off ...
> 
>>>
>>> Move the swap soft-dirty bit to bit 12 and start the swap offset at
>>> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
>>> in migration PMDs and lets us keep the existing pmd_present() check
>>> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
>>> tracking is disabled.
>>
>> That reduces the effective swap size (and PFN we can store). Could that be a
>> problem?
> 
> We don't need all 52 bits of the swap offset.
> 
> RV64 PFNs only need 44 bits for migration entries, and actual swap is
> already limited to about 16 TiB per area with 4 KiB pages by
> last_page (__u32) and swap_info_struct.max (unsigned int).
> 
> So there's room to reserve a bit without reducing the supported swap
> size or PFN range.
> 
> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
> offset.
> 
> [...]
> 
> Cheers, Lance



WARNING: multiple messages have this Message-ID (diff)
From: Lance Yang <lance.yang@linux.dev>
To: david@kernel.org
Cc: pjw@kernel.org, palmer@dabbelt.com, aou@eecs.berkeley.edu,
	alex@ghiti.fr, akpm@linux-foundation.org,
	zhangchunyan@iscas.ac.cn, rppt@kernel.org, kas@kernel.org,
	andrew+kernel@donnellan.id.au, rmclure@linux.ibm.com,
	debug@rivosinc.com, baolin.wang@linux.alibaba.com,
	usama.arif@linux.dev, wangruikang@iscas.ac.cn,
	namcao@linutronix.de, linux-riscv@lists.infradead.org,
	linux-kernel@vger.kernel.org, alexghiti@rivosinc.com,
	viro@zeniv.linux.org.uk, ajones@ventanamicro.com, arnd@arndb.de,
	axelrasmussen@google.com, brauner@kernel.org,
	conor.dooley@microchip.com, conor@kernel.org, jack@suse.cz,
	liam@infradead.org, ljs@kernel.org, mhocko@suse.com,
	paul.walmsley@sifive.com, peterx@redhat.com, robh@kernel.org,
	surenb@google.com, vbabka@kernel.org, yuanchu@google.com,
	stable@vger.kernel.org, pasha.tatashin@soleen.com,
	linux-mm@kvack.org, me@ziyao.cc
Subject: Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
Date: Mon, 5 Oct 2026 22:23:57 +0800	[thread overview]
Message-ID: <b928b8a4-5ec5-4642-ac02-9ecbbbfd6218@linux.dev> (raw)
In-Reply-To: <20261005134641.5801-1-lance.yang@linux.dev>



On 2026/10/5 21:46, Lance Yang wrote:
> 
> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/4/26 05:03, Lance Yang wrote:
>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>> a problem for PMD migration entries, since pmd_present() also checks
>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>> cleared during splitting.
>>
>> I'm curious: why do we have to set leaf indications for non-present things? The
>> HW sure will ignore it, right?
> 
> Yeah, that surprised me too :) Still wrapping my head around the details
> ...
> 
>> Is this a sw problem? Who needs that?
> 
> IIUC, it's for software during a PMD split.
> 
> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
> before installing the PTE table. Software still needs pmd_present() and
> pmd_trans_huge() to recognize the THP in between.

Also, x86 keeps _PAGE_PSE to identify the huge PMD. RISC-V doesn't
have a separate leaf bit, so it keeps the R/W/X bits to identify
the PMD as a leaf entry :)

> 
> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
> to hardware but still identifiable as a THP by software.
> 
> Hopefully I didn't miss something.
> 
>>>
>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>> migration PMD as a present THP! The fault handler skips
>>> pmd_migration_entry_wait(), and a write fault can end up in
>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>> as a mapped PFN.
>>
>> That sounds bad.
> 
> YES, looks a bit off ...
> 
>>>
>>> Move the swap soft-dirty bit to bit 12 and start the swap offset at
>>> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
>>> in migration PMDs and lets us keep the existing pmd_present() check
>>> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
>>> tracking is disabled.
>>
>> That reduces the effective swap size (and PFN we can store). Could that be a
>> problem?
> 
> We don't need all 52 bits of the swap offset.
> 
> RV64 PFNs only need 44 bits for migration entries, and actual swap is
> already limited to about 16 TiB per area with 4 KiB pages by
> last_page (__u32) and swap_info_struct.max (unsigned int).
> 
> So there's room to reserve a bit without reducing the supported swap
> size or PFN range.
> 
> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
> offset.
> 
> [...]
> 
> Cheers, Lance


_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv

  reply	other threads:[~2026-10-05 14:24 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-04  3:03 [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present Lance Yang
2026-10-04  3:03 ` Lance Yang
2026-10-05 10:37 ` David Hildenbrand (Arm)
2026-10-05 10:37   ` David Hildenbrand (Arm)
2026-10-05 13:46   ` Lance Yang
2026-10-05 13:46     ` Lance Yang
2026-10-05 14:23     ` Lance Yang [this message]
2026-10-05 14:23       ` Lance Yang
2026-10-07 11:01     ` David Hildenbrand (Arm)
2026-10-07 11:01       ` David Hildenbrand (Arm)
2026-10-07 11:33       ` Lance Yang
2026-10-07 11:33         ` Lance Yang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=b928b8a4-5ec5-4642-ac02-9ecbbbfd6218@linux.dev \
    --to=lance.yang@linux.dev \
    --cc=ajones@ventanamicro.com \
    --cc=akpm@linux-foundation.org \
    --cc=alex@ghiti.fr \
    --cc=alexghiti@rivosinc.com \
    --cc=andrew+kernel@donnellan.id.au \
    --cc=aou@eecs.berkeley.edu \
    --cc=arnd@arndb.de \
    --cc=axelrasmussen@google.com \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=brauner@kernel.org \
    --cc=conor.dooley@microchip.com \
    --cc=conor@kernel.org \
    --cc=david@kernel.org \
    --cc=debug@rivosinc.com \
    --cc=jack@suse.cz \
    --cc=kas@kernel.org \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=linux-riscv@lists.infradead.org \
    --cc=ljs@kernel.org \
    --cc=me@ziyao.cc \
    --cc=mhocko@suse.com \
    --cc=namcao@linutronix.de \
    --cc=palmer@dabbelt.com \
    --cc=pasha.tatashin@soleen.com \
    --cc=paul.walmsley@sifive.com \
    --cc=peterx@redhat.com \
    --cc=pjw@kernel.org \
    --cc=rmclure@linux.ibm.com \
    --cc=robh@kernel.org \
    --cc=rppt@kernel.org \
    --cc=stable@vger.kernel.org \
    --cc=surenb@google.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    --cc=wangruikang@iscas.ac.cn \
    --cc=yuanchu@google.com \
    --cc=zhangchunyan@iscas.ac.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.