* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
@ 2026-10-05 13:46 ` Lance Yang
0 siblings, 0 replies; 12+ messages in thread
From: Lance Yang @ 2026-10-05 13:46 UTC (permalink / raw)
To: david
Cc: lance.yang, pjw, palmer, aou, alex, akpm, zhangchunyan, rppt, kas,
andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
yuanchu, stable, pasha.tatashin, linux-mm, me
On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>On 10/4/26 05:03, Lance Yang wrote:
>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>> a problem for PMD migration entries, since pmd_present() also checks
>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>> cleared during splitting.
>
>I'm curious: why do we have to set leaf indications for non-present things? The
>HW sure will ignore it, right?
Yeah, that surprised me too :) Still wrapping my head around the details
...
>Is this a sw problem? Who needs that?
IIUC, it's for software during a PMD split.
__split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
before installing the PTE table. Software still needs pmd_present() and
pmd_trans_huge() to recognize the THP in between.
RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
to hardware but still identifiable as a THP by software.
Hopefully I didn't miss something.
>>
>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>> migration PMD as a present THP! The fault handler skips
>> pmd_migration_entry_wait(), and a write fault can end up in
>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>> as a mapped PFN.
>
>That sounds bad.
YES, looks a bit off ...
>>
>> Move the swap soft-dirty bit to bit 12 and start the swap offset at
>> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
>> in migration PMDs and lets us keep the existing pmd_present() check
>> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
>> tracking is disabled.
>
>That reduces the effective swap size (and PFN we can store). Could that be a
>problem?
We don't need all 52 bits of the swap offset.
RV64 PFNs only need 44 bits for migration entries, and actual swap is
already limited to about 16 TiB per area with 4 KiB pages by
last_page (__u32) and swap_info_struct.max (unsigned int).
So there's room to reserve a bit without reducing the supported swap
size or PFN range.
CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
offset.
[...]
Cheers, Lance
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
2026-10-05 13:46 ` Lance Yang
@ 2026-10-05 14:23 ` Lance Yang
-1 siblings, 0 replies; 12+ messages in thread
From: Lance Yang @ 2026-10-05 14:23 UTC (permalink / raw)
To: david
Cc: pjw, palmer, aou, alex, akpm, zhangchunyan, rppt, kas,
andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
yuanchu, stable, pasha.tatashin, linux-mm, me
On 2026/10/5 21:46, Lance Yang wrote:
>
> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/4/26 05:03, Lance Yang wrote:
>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>> a problem for PMD migration entries, since pmd_present() also checks
>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>> cleared during splitting.
>>
>> I'm curious: why do we have to set leaf indications for non-present things? The
>> HW sure will ignore it, right?
>
> Yeah, that surprised me too :) Still wrapping my head around the details
> ...
>
>> Is this a sw problem? Who needs that?
>
> IIUC, it's for software during a PMD split.
>
> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
> before installing the PTE table. Software still needs pmd_present() and
> pmd_trans_huge() to recognize the THP in between.
Also, x86 keeps _PAGE_PSE to identify the huge PMD. RISC-V doesn't
have a separate leaf bit, so it keeps the R/W/X bits to identify
the PMD as a leaf entry :)
>
> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
> to hardware but still identifiable as a THP by software.
>
> Hopefully I didn't miss something.
>
>>>
>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>> migration PMD as a present THP! The fault handler skips
>>> pmd_migration_entry_wait(), and a write fault can end up in
>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>> as a mapped PFN.
>>
>> That sounds bad.
>
> YES, looks a bit off ...
>
>>>
>>> Move the swap soft-dirty bit to bit 12 and start the swap offset at
>>> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
>>> in migration PMDs and lets us keep the existing pmd_present() check
>>> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
>>> tracking is disabled.
>>
>> That reduces the effective swap size (and PFN we can store). Could that be a
>> problem?
>
> We don't need all 52 bits of the swap offset.
>
> RV64 PFNs only need 44 bits for migration entries, and actual swap is
> already limited to about 16 TiB per area with 4 KiB pages by
> last_page (__u32) and swap_info_struct.max (unsigned int).
>
> So there's room to reserve a bit without reducing the supported swap
> size or PFN range.
>
> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
> offset.
>
> [...]
>
> Cheers, Lance
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
@ 2026-10-05 14:23 ` Lance Yang
0 siblings, 0 replies; 12+ messages in thread
From: Lance Yang @ 2026-10-05 14:23 UTC (permalink / raw)
To: david
Cc: pjw, palmer, aou, alex, akpm, zhangchunyan, rppt, kas,
andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
yuanchu, stable, pasha.tatashin, linux-mm, me
On 2026/10/5 21:46, Lance Yang wrote:
>
> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/4/26 05:03, Lance Yang wrote:
>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>> a problem for PMD migration entries, since pmd_present() also checks
>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>> cleared during splitting.
>>
>> I'm curious: why do we have to set leaf indications for non-present things? The
>> HW sure will ignore it, right?
>
> Yeah, that surprised me too :) Still wrapping my head around the details
> ...
>
>> Is this a sw problem? Who needs that?
>
> IIUC, it's for software during a PMD split.
>
> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
> before installing the PTE table. Software still needs pmd_present() and
> pmd_trans_huge() to recognize the THP in between.
Also, x86 keeps _PAGE_PSE to identify the huge PMD. RISC-V doesn't
have a separate leaf bit, so it keeps the R/W/X bits to identify
the PMD as a leaf entry :)
>
> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
> to hardware but still identifiable as a THP by software.
>
> Hopefully I didn't miss something.
>
>>>
>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>> migration PMD as a present THP! The fault handler skips
>>> pmd_migration_entry_wait(), and a write fault can end up in
>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>> as a mapped PFN.
>>
>> That sounds bad.
>
> YES, looks a bit off ...
>
>>>
>>> Move the swap soft-dirty bit to bit 12 and start the swap offset at
>>> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
>>> in migration PMDs and lets us keep the existing pmd_present() check
>>> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
>>> tracking is disabled.
>>
>> That reduces the effective swap size (and PFN we can store). Could that be a
>> problem?
>
> We don't need all 52 bits of the swap offset.
>
> RV64 PFNs only need 44 bits for migration entries, and actual swap is
> already limited to about 16 TiB per area with 4 KiB pages by
> last_page (__u32) and swap_info_struct.max (unsigned int).
>
> So there's room to reserve a bit without reducing the supported swap
> size or PFN range.
>
> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
> offset.
>
> [...]
>
> Cheers, Lance
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
2026-10-05 13:46 ` Lance Yang
@ 2026-10-07 11:01 ` David Hildenbrand (Arm)
-1 siblings, 0 replies; 12+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-07 11:01 UTC (permalink / raw)
To: Lance Yang
Cc: pjw, palmer, aou, alex, akpm, zhangchunyan, rppt, kas,
andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
yuanchu, stable, pasha.tatashin, linux-mm, me
On 10/5/26 15:46, Lance Yang wrote:
>
> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/4/26 05:03, Lance Yang wrote:
>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>> a problem for PMD migration entries, since pmd_present() also checks
>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>> cleared during splitting.
>>
>> I'm curious: why do we have to set leaf indications for non-present things? The
>> HW sure will ignore it, right?
>
> Yeah, that surprised me too :) Still wrapping my head around the details
> ...
>
>> Is this a sw problem? Who needs that?
>
> IIUC, it's for software during a PMD split.
>
> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
> before installing the PTE table. Software still needs pmd_present() and
> pmd_trans_huge() to recognize the THP in between.
>
> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
> to hardware but still identifiable as a THP by software.
>
> Hopefully I didn't miss something.
I thought some change we performed to GUP-fast wouldn't require that anymore.
But my memory is a bit vague ... :)
>
>>>
>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>> migration PMD as a present THP! The fault handler skips
>>> pmd_migration_entry_wait(), and a write fault can end up in
>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>> as a mapped PFN.
>>
>> That sounds bad.
>
[...]
> YES, looks a bit off ...
>> That reduces the effective swap size (and PFN we can store). Could that be a
>> problem?
>
> We don't need all 52 bits of the swap offset.
>
> RV64 PFNs only need 44 bits for migration entries, and actual swap is
> already limited to about 16 TiB per area with 4 KiB pages by
> last_page (__u32) and swap_info_struct.max (unsigned int).
>
> So there's room to reserve a bit without reducing the supported swap
> size or PFN range.
>
> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
> offset.
Good, can you spell that out in the patch description?
--
Cheers,
David
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
@ 2026-10-07 11:01 ` David Hildenbrand (Arm)
0 siblings, 0 replies; 12+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-07 11:01 UTC (permalink / raw)
To: Lance Yang
Cc: pjw, palmer, aou, alex, akpm, zhangchunyan, rppt, kas,
andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
yuanchu, stable, pasha.tatashin, linux-mm, me
On 10/5/26 15:46, Lance Yang wrote:
>
> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/4/26 05:03, Lance Yang wrote:
>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>> a problem for PMD migration entries, since pmd_present() also checks
>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>> cleared during splitting.
>>
>> I'm curious: why do we have to set leaf indications for non-present things? The
>> HW sure will ignore it, right?
>
> Yeah, that surprised me too :) Still wrapping my head around the details
> ...
>
>> Is this a sw problem? Who needs that?
>
> IIUC, it's for software during a PMD split.
>
> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
> before installing the PTE table. Software still needs pmd_present() and
> pmd_trans_huge() to recognize the THP in between.
>
> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
> to hardware but still identifiable as a THP by software.
>
> Hopefully I didn't miss something.
I thought some change we performed to GUP-fast wouldn't require that anymore.
But my memory is a bit vague ... :)
>
>>>
>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>> migration PMD as a present THP! The fault handler skips
>>> pmd_migration_entry_wait(), and a write fault can end up in
>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>> as a mapped PFN.
>>
>> That sounds bad.
>
[...]
> YES, looks a bit off ...
>> That reduces the effective swap size (and PFN we can store). Could that be a
>> problem?
>
> We don't need all 52 bits of the swap offset.
>
> RV64 PFNs only need 44 bits for migration entries, and actual swap is
> already limited to about 16 TiB per area with 4 KiB pages by
> last_page (__u32) and swap_info_struct.max (unsigned int).
>
> So there's room to reserve a bit without reducing the supported swap
> size or PFN range.
>
> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
> offset.
Good, can you spell that out in the patch description?
--
Cheers,
David
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
2026-10-07 11:01 ` David Hildenbrand (Arm)
@ 2026-10-07 11:33 ` Lance Yang
-1 siblings, 0 replies; 12+ messages in thread
From: Lance Yang @ 2026-10-07 11:33 UTC (permalink / raw)
To: david
Cc: lance.yang, pjw, palmer, aou, alex, akpm, zhangchunyan, rppt, kas,
andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
yuanchu, stable, pasha.tatashin, linux-mm, me
On Wed, Oct 07, 2026 at 01:01:14PM +0200, David Hildenbrand (Arm) wrote:
>On 10/5/26 15:46, Lance Yang wrote:
>>
>> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>>> On 10/4/26 05:03, Lance Yang wrote:
>>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>>> a problem for PMD migration entries, since pmd_present() also checks
>>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>>> cleared during splitting.
>>>
>>> I'm curious: why do we have to set leaf indications for non-present things? The
>>> HW sure will ignore it, right?
>>
>> Yeah, that surprised me too :) Still wrapping my head around the details
>> ...
>>
>>> Is this a sw problem? Who needs that?
>>
>> IIUC, it's for software during a PMD split.
>>
>> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
>> before installing the PTE table. Software still needs pmd_present() and
>> pmd_trans_huge() to recognize the THP in between.
>>
>> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
>> to hardware but still identifiable as a THP by software.
>>
>> Hopefully I didn't miss something.
>
>I thought some change we performed to GUP-fast wouldn't require that anymore.
>But my memory is a bit vague ... :)
Yeah, I'll have another look at that :)
>>
>>>>
>>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>>> migration PMD as a present THP! The fault handler skips
>>>> pmd_migration_entry_wait(), and a write fault can end up in
>>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>>> as a mapped PFN.
>>>
>>> That sounds bad.
>>
>
>[...]
>
>> YES, looks a bit off ...
>>> That reduces the effective swap size (and PFN we can store). Could that be a
>>> problem?
>>
>> We don't need all 52 bits of the swap offset.
>>
>> RV64 PFNs only need 44 bits for migration entries, and actual swap is
>> already limited to about 16 TiB per area with 4 KiB pages by
>> last_page (__u32) and swap_info_struct.max (unsigned int).
>>
>> So there's room to reserve a bit without reducing the supported swap
>> size or PFN range.
>>
>> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
>> offset.
>
>Good, can you spell that out in the patch description?
Sure, will add that to the commit message in v2 once I've heard back
from maintainers :)
Cheers, Lance
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
@ 2026-10-07 11:33 ` Lance Yang
0 siblings, 0 replies; 12+ messages in thread
From: Lance Yang @ 2026-10-07 11:33 UTC (permalink / raw)
To: david
Cc: lance.yang, pjw, palmer, aou, alex, akpm, zhangchunyan, rppt, kas,
andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
yuanchu, stable, pasha.tatashin, linux-mm, me
On Wed, Oct 07, 2026 at 01:01:14PM +0200, David Hildenbrand (Arm) wrote:
>On 10/5/26 15:46, Lance Yang wrote:
>>
>> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>>> On 10/4/26 05:03, Lance Yang wrote:
>>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>>> a problem for PMD migration entries, since pmd_present() also checks
>>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>>> cleared during splitting.
>>>
>>> I'm curious: why do we have to set leaf indications for non-present things? The
>>> HW sure will ignore it, right?
>>
>> Yeah, that surprised me too :) Still wrapping my head around the details
>> ...
>>
>>> Is this a sw problem? Who needs that?
>>
>> IIUC, it's for software during a PMD split.
>>
>> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
>> before installing the PTE table. Software still needs pmd_present() and
>> pmd_trans_huge() to recognize the THP in between.
>>
>> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
>> to hardware but still identifiable as a THP by software.
>>
>> Hopefully I didn't miss something.
>
>I thought some change we performed to GUP-fast wouldn't require that anymore.
>But my memory is a bit vague ... :)
Yeah, I'll have another look at that :)
>>
>>>>
>>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>>> migration PMD as a present THP! The fault handler skips
>>>> pmd_migration_entry_wait(), and a write fault can end up in
>>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>>> as a mapped PFN.
>>>
>>> That sounds bad.
>>
>
>[...]
>
>> YES, looks a bit off ...
>>> That reduces the effective swap size (and PFN we can store). Could that be a
>>> problem?
>>
>> We don't need all 52 bits of the swap offset.
>>
>> RV64 PFNs only need 44 bits for migration entries, and actual swap is
>> already limited to about 16 TiB per area with 4 KiB pages by
>> last_page (__u32) and swap_info_struct.max (unsigned int).
>>
>> So there's room to reserve a bit without reducing the supported swap
>> size or PFN range.
>>
>> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
>> offset.
>
>Good, can you spell that out in the patch description?
Sure, will add that to the commit message in v2 once I've heard back
from maintainers :)
Cheers, Lance
_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv
^ permalink raw reply [flat|nested] 12+ messages in thread