Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Balbir Singh <balbirs@nvidia.com>
To: Lance Yang <lance.yang@linux.dev>, akpm@linux-foundation.org
Cc: gourry@gourry.net, linux-mm@kvack.org,
	linux-kernel@vger.kernel.org, kernel-team@meta.com,
	david@kernel.org, ljs@kernel.org, ziy@nvidia.com,
	baolin.wang@linux.alibaba.com, liam@infradead.org,
	nico.pache@linux.dev, ryan.roberts@arm.com, dev.jain@arm.com,
	baohua@kernel.org, usama.arif@linux.dev, vbabka@kernel.org,
	jannh@google.com, matthew.brost@intel.com,
	joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com,
	ying.huang@linux.alibaba.com, apopple@nvidia.com
Subject: Re: [PATCH v2 0/3] mm: reject zone device folios in more folio walkers
Date: Tue, 1 Sep 2026 07:57:43 +1000	[thread overview]
Message-ID: <03fc6a4f-8393-4785-bf73-d1c9f2c202be@nvidia.com> (raw)
In-Reply-To: <20260830074751.25370-1-lance.yang@linux.dev>

On 8/30/26 5:47 PM, Lance Yang wrote:
> 
> On Sun, Aug 30, 2026 at 01:18:24PM +0800, Lance Yang wrote:
>>
>>
>> On 2026/8/30 08:18, Andrew Morton wrote:
>>> On Tue, 18 Aug 2026 12:31:43 +0800 Lance Yang <lance.yang@linux.dev> wrote:
>>>
>>>>
>>>> On Mon, Aug 17, 2026 at 06:08:07PM -0400, Gregory Price wrote:
>>>>> Several LRU-oriented mm walkers resolve the folio backing a PMD entry
>>>>> (or a physical pfn) and then reclaim, age, migrate, or lazyfree it
>>>>> without ever checking for ZONE_DEVICE memory.
>>>>>
>>>>> This series adds missing folio_is_zone_device() rejections, matching
>>>>> the checks that comparable walkers already perform.
>>>>>
>>>>> - mm/huge_memory, mm/madvise: the !pmd_present branch above these sites
>>>>>   only filters device-private entries (which are non-present).
>>>>>
>>>>>   A present zone device PMD (e.g. device-coherent) would still reach the
>>>>>   folio and be lazyfreed / aged / paged out. Add an explicit check.
>>>>>
>>>>> - mm/mempolicy: queue_folios_pmd() can see a present zone device PMD
>>>>>   (e.g. device-coherent) and queue it for migration.
>>>>>
>>>>> No crash reproducer - this is a correctness/hardening cleanup found by
>>>>> inspection. All checks are placed after the folio is resolved and before
>>>>> it is acted upon, on paths that already hold the relevant page-table lock,
>>>>> so no locking or refcount changes are involved.
>>>>
>>>> Cool!
>>>>
>>>> Gave the whole series a spin on x86_64 QEMU with a PMD-mapped
>>>> device-coherent THP. Without these patches, partial MADV_FREE and
>>>> MADV_COLD reliably hit a kernel panic in remove_migration_pte(), while
>>>> mbind(MPOL_MF_MOVE | MPOL_MF_STRICT) returned -EIO.
>>>>
>>>> With v2, all three worked fine, PMD mapping stayed intact, and data
>>>> checked out :)
>>>>
>>>> Note that both kernels used the same small change to the in-kernel HMM
>>>> test driver, allowing its coherent device memory to be allocated as 2 MB
>>>> folios so the PMD-mapped test case could be exercised.
>>>>
>>>> Tested-by: Lance Yang <lance.yang@linux.dev>
>>>
>>> Thanks Lance, you're so diligent.
>>>
>>> I'm wondering what to do here.  Gregory told us
>>>
>>> : No crash reproducer - this is a correctness/hardening cleanup found by
>>> : inspection. All checks are placed after the folio is resolved and before
>>> : it is acted upon, on paths that already hold the relevant page-table lock,
>>> : so no locking or refcount changes are involved.
>>>
>>> And you had to tweak the hmm-test driver to reproduce the bug(s).
>>>
>>> So when do we push this series out to -stable?  As a hair-on-fire
>>> hotfix, or as a leisurely next-merge-window thing?
>>
>> Thanks, Andrew :) Yeah, I'd say next merge window should be fine :)
>>
>> The crash is real once the mapping exists, but I had to tweak test_hmm
>> to create that PMD-mapped device-coherent folio, and I couldn't find
>> any in-tree production driver doing that today.
>>
>> So no need to rush this one, I guess.
> 
> BTW, noticed that the ZONE_DEVICE split handling only covers
> device-private folios, so device-coherent folios aren't supported ...
> 
> The call chains are:
> 
> split_folio()
>   -> __folio_split()
>     -> folio_check_splittable()
>     -> __folio_freeze_and_split_unmapped()
> 
> migrate_vma_pages()
>   -> __migrate_device_pages()
>     -> migrate_vma_split_unmapped_folio()
>       -> folio_split_unmapped()
>         -> __folio_freeze_and_split_unmapped()
> 
> And I added the device-coherent check to folio_check_splittable() and
> folio_split_unmapped(). See below. They can go away once device-coherent
> folio splitting is supported :)
> 
> If folks think it's worth having, I can send it as a follow-up :)
> 
> ---8<---
> Subject: [PATCH] mm/huge_memory: don't split device-coherent folios
> 
> From: Lance Yang <lance.yang@linux.dev>
> 
> The ZONE_DEVICE split handling only covers device-private folios.
> Device-coherent folios are not supported.
> 
> The call chains are:
> 
> split_folio()
>   -> __folio_split()
>     -> folio_check_splittable()
>     -> __folio_freeze_and_split_unmapped()
> 
> migrate_vma_pages()
>   -> __migrate_device_pages()
>     -> migrate_vma_split_unmapped_folio()
>       -> folio_split_unmapped()
>         -> __folio_freeze_and_split_unmapped()
> 
> Reject device-coherent folios in folio_check_splittable() and
> folio_split_unmapped().
> 
> Fixes: a30b48bf1b24 ("mm/migrate_device: implement THP migration of zone device pages")
> Signed-off-by: Lance Yang <lance.yang@linux.dev>
> ---
>  mm/huge_memory.c | 14 +++++++++++---
>  1 file changed, 11 insertions(+), 3 deletions(-)
> 
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index 54494c3fa983..a5dd38e9a8de 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -3937,6 +3937,10 @@ int folio_check_splittable(struct folio *folio, unsigned int new_order,
>  	if (!folio->mapping && !folio_test_anon(folio))
>  		return -EBUSY;
> 
> +	/* TODO: Support splitting device-coherent folios. */
> +	if (folio_is_device_coherent(folio))
> +		return -EOPNOTSUPP;
> +
>  	/* order-1 is not supported for anonymous THP. */
>  	if (folio_test_anon(folio) && new_order == 1)
>  		return -EINVAL;
> @@ -4354,15 +4358,16 @@ static int __folio_split(struct folio *folio, unsigned int new_order,
>   *
>   * anon_vma_lock is not required to be held, mmap_read_lock() or
>   * mmap_write_lock() should be held. @folio is expected to be locked by the
> - * caller. device-private and non device-private folios are supported along
> + * caller. device-private and non-ZONE_DEVICE folios are supported along
>   * with folios that are in the swapcache. @folio should also be unmapped and
>   * isolated from LRU (if applicable)
>   *
>   * Upon return, the folio is not remapped, split folios are not added to LRU,
>   * free_folio_and_swap_cache() is not called, and new folios remain locked.
>   *
> - * Return: 0 on success, -EAGAIN if the folio cannot be split (e.g., due to
> - *         insufficient reference count or extra pins).
> + * Return: 0 on success, -EOPNOTSUPP for device-coherent folios, or -EAGAIN if
> + *         the folio cannot be split (e.g., due to insufficient reference
> + *         count or extra pins).
>   */
>  int folio_split_unmapped(struct folio *folio, unsigned int new_order)
>  {
> @@ -4373,6 +4378,9 @@ int folio_split_unmapped(struct folio *folio, unsigned int new_order)
>  	VM_WARN_ON_ONCE_FOLIO(!folio_test_large(folio), folio);
>  	VM_WARN_ON_ONCE_FOLIO(!folio_test_anon(folio), folio);
> 
> +	if (folio_is_device_coherent(folio))
> +		return -EOPNOTSUPP;
> +
>  	if (folio_expected_ref_count(folio) != folio_ref_count(folio) - 1)
>  		return -EAGAIN;
> 
> --
> 

FYI: Device Coherent THP is not yet supported, support should be easy to add. it
is definitely desirable, we should get it working along with mTHP support as
well (TODO).

Balbir





      parent reply	other threads:[~2026-08-31 21:58 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-17 22:08 [PATCH v2 0/3] mm: reject zone device folios in more folio walkers Gregory Price
2026-08-17 22:08 ` [PATCH v2 1/3] mm/huge_memory: skip zone device folios in madvise_free_huge_pmd() Gregory Price
2026-08-17 22:08 ` [PATCH v2 2/3] mm/madvise: skip zone device folios in cold/pageout PMD range Gregory Price
2026-08-18  8:27   ` Balbir Singh
2026-08-17 22:08 ` [PATCH v2 3/3] mm/mempolicy: skip zone device folios when queueing folios Gregory Price
2026-08-17 23:34   ` Balbir Singh
2026-08-18 12:51     ` Gregory Price
2026-08-18  7:20   ` Lorenzo Stoakes (ARM)
2026-08-18 17:19   ` David Hildenbrand (Arm)
2026-08-18  4:31 ` [PATCH v2 0/3] mm: reject zone device folios in more folio walkers Lance Yang
2026-08-30  0:18   ` Andrew Morton
2026-08-30  5:18     ` Lance Yang
2026-08-30  7:47       ` Lance Yang
2026-08-30 16:50         ` Gregory Price
2026-08-31 21:57         ` Balbir Singh [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=03fc6a4f-8393-4785-bf73-d1c9f2c202be@nvidia.com \
    --to=balbirs@nvidia.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=byungchul@sk.com \
    --cc=david@kernel.org \
    --cc=dev.jain@arm.com \
    --cc=gourry@gourry.net \
    --cc=jannh@google.com \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kernel-team@meta.com \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=matthew.brost@intel.com \
    --cc=nico.pache@linux.dev \
    --cc=rakie.kim@sk.com \
    --cc=ryan.roberts@arm.com \
    --cc=usama.arif@linux.dev \
    --cc=vbabka@kernel.org \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox