Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: "David Hildenbrand (Arm)" <david@kernel.org>
To: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>,
	Andrew Morton <akpm@linux-foundation.org>,
	Zi Yan <ziy@nvidia.com>,
	Baolin Wang <baolin.wang@linux.alibaba.com>,
	"Liam R. Howlett" <liam@infradead.org>,
	Nico Pache <npache@redhat.com>,
	Ryan Roberts <ryan.roberts@arm.com>, Dev Jain <dev.jain@arm.com>,
	Barry Song <baohua@kernel.org>, Lance Yang <lance.yang@linux.dev>,
	Usama Arif <usama.arif@linux.dev>,
	Pankaj Raghav <p.raghav@samsung.com>,
	Hannes Reinecke <hare@suse.de>, Hugh Dickins <hughd@google.com>,
	Yang Shi <shy828301@gmail.com>, Kiryl Shutsemau <kas@kernel.org>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
	Hengbin Zhang <uqbarz@gmail.com>,
	stable@vger.kernel.org
Subject: Re: [PATCH mm-hotfixes v2 1/2] mm/huge_memory: fix huge_zero_pfn race
Date: Thu, 30 Jul 2026 16:18:38 +0200	[thread overview]
Message-ID: <7c19c4f3-2d89-48ff-bdf1-5bad418e719b@kernel.org> (raw)
In-Reply-To: <20260730-fix-refcounted-huge-zero-v2-1-c5d8a41b317f@kernel.org>

On 7/30/26 12:55, Lorenzo Stoakes (ARM) wrote:
> If !CONFIG_PERSISTENT_HUGE_ZERO_FOLIO, the huge_zero_folio is refcounted by
> huge_zero_refcount and returned by mm_get_huge_zero_folio().
> 
> When the caller is done with the huge zero page, its reference count is
> decremented. Only a shrinker can set the reference count to zero.
> 
> A race can unfortunately occur between a shrinker decrementing the
> reference count to zero and a concurrent page fault.
> 
> This is because shrink_huge_zero_folio_scan() might, if very
> unlucky, be preempted between setting huge_zero_refcount to zero
> and writing an invalid value.
> 
> During this time get_huge_zero_folio() could write to huge_zero_pfn
> before shrink_huge_zero_folio_scan() resumes.
> 
> In this event the huge zero folio will be persistently
> misidentified causing the THP code path to be entered
> inappropriately for the huge zero folio:
> 
>                 CPU 0                                   CPU 1
> =======================================|=================================
> shrink_huge_zero_folio_scan()          |
>    atomic_cmpxchg() sets refcount to 0 |
>    xchg() sets huge_zero_folio to NULL | get_huge_zero_folio()
>                  |                     |    atomic_inc_not_zero() -> zero
>       preempted for a long time        |    Allocate new huge zero folio
>                  |                     |    Write valid huge_zero_folio
>                  v                     |    Write valid huge_zero_pfn
>   Overwrite huge_zero_pfn with ~0UL   <--- Invalid overwrite!
> 
> This results in is_huge_zero_pfn() and is_huge_zero_pmd() incorrectly
> returning false for a huge zero page which could result in issues like the
> huge zero folio being incorrectly split.
> 
> Note that the issue is with huge_zero_pfn not huge_zero_folio, as
> get_huge_zero_folio() uses cmpxchg() gated on huge_zero_folio being NULL
> with a retry loop and shrink_huge_zero_folio_scan() uses xchg() to set
> huge_zero_folio.
> 
> Fix the issue by introducing a spinlock, huge_zero_lock, to prevent
> concurrent write of huge_zero_folio, huge_zero_pfn and huge_zero_refcount.
> 
> There needs to be significant care taken here to ensure correctness:
> 
> The fast path in get_huge_zero_folio() uses atomic_inc_not_zero(), which is
> outside of the critical section, and means huge zero allocation is gated on
> zero huge_zero_refcount.
> 
> The fast path doesn't use huge_zero_lock, so the critical section is
> irrelevant to it.
> 
> So invariants are required - huge_zero_refcount MUST:
> 
> * Only be set in the huge_zero_lock critical section to ensure
>   serialisation of huge_zero_pfn, huge_zero_folio and huge_zero_refcount
>   writes.
> 
> * Be set non-zero only AFTER huge_zero_[pfn, folio] are set to valid values
>   so installation of the huge zero folio on read page fault ensures
>   concurrent is_huge_zero_*() calls correctly identify the huge zero folio.
> 
> * Be set zero only BEFORE huge_zero_[pfn, folio] are set to NULL and ~0UL
>   respectively, and atomically.
> 
> Establish these by:
> 
> * Only setting huge_zero_refcount to zero or an absolute value in the
>   huge_zero_lock critical section in get_huge_zero_folio() and
>   shrink_huge_zero_folio_scan(), and always updating atomically there
>   and elsewhere.
> 
> * Using atomic_set_release(&huge_zero_refcount) in get_huge_zero_folio()
>   after huge_zero_[pfn, folio] are set. This is paired with
>   atomic_inc_not_zero() to ensure atomic_inc_not_zero() only observes a
>   non-zero value if huge_zero_[pfn, folio] are set.
> 
> * Using atomic_cmpxchg() in shrink_huge_zero_folio_scan() (as before) to
>   ensure that it is set zero only when equal to 1 and set atomically.
> 
> * atomic_cmpxchg() being fully ordered ensures this is done prior to
>   huge_zero_[folio, pfn] being set to NULL and ~0UL respectively.
> 
> Eliminate the retry loop in get_huge_zero_folio() as the atomic_cmpxchg()
> in shrink_huge_zero_folio_scan() is now performed under the lock, and
> replace with an equally locked atomic_inc() to set the reference count
> should the caller be raced on huge zero folio installation.
> 
> folio_put() naturally implies a full memory barrier so its ordering is
> maintained correctly.
> 
> The huge zero folio also cannot be released except when the shrinker does
> so as it is non-LRU and non-rmappable.
> 
> Note that only the huge zero shrinker (via shrink_huge_zero_folio_scan())
> can actually set huge_zero_refcount to zero, which is the count of mm's
> which have at least one huge zero folio installed plus one shrinker pin.
> 
> Additionally convert a BUG_ON() to a VM_WARN_ON_ONCE().
> 
> Suggested-by: David Hildenbrand (Arm) <david@kernel.org>
> Reported-by: Hengbin Zhang <uqbarz@gmail.com>
> Closes: https://lore.kernel.org/linux-mm/20260727154001.4102341-1-uqbarz@gmail.com/
> Fixes: 3b77e8c8cde5 ("mm/thp: make is_huge_zero_pmd() safe and quicker")
> Cc: stable@vger.kernel.org
> Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
> ---

Thanks!

Acked-by: David Hildenbrand (Arm) <david@kernel.org>

-- 
Cheers,

David


  reply	other threads:[~2026-07-30 14:18 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-30 10:55 [PATCH mm-hotfixes v2 0/2] mm/huge_memory: fix huge_zero_pfn race Lorenzo Stoakes (ARM)
2026-07-30 10:55 ` [PATCH mm-hotfixes v2 1/2] " Lorenzo Stoakes (ARM)
2026-07-30 14:18   ` David Hildenbrand (Arm) [this message]
2026-07-30 10:55 ` [PATCH mm-hotfixes v2 2/2] mm/huge_memory: separate out CONFIG_PERSISTENT_HUGE_ZERO_FOLIO logic Lorenzo Stoakes (ARM)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=7c19c4f3-2d89-48ff-bdf1-5bad418e719b@kernel.org \
    --to=david@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=baohua@kernel.org \
    --cc=baolin.wang@linux.alibaba.com \
    --cc=dev.jain@arm.com \
    --cc=hare@suse.de \
    --cc=hughd@google.com \
    --cc=kas@kernel.org \
    --cc=lance.yang@linux.dev \
    --cc=liam@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=npache@redhat.com \
    --cc=p.raghav@samsung.com \
    --cc=ryan.roberts@arm.com \
    --cc=shy828301@gmail.com \
    --cc=stable@vger.kernel.org \
    --cc=uqbarz@gmail.com \
    --cc=usama.arif@linux.dev \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox