From: Pedro Falcato <pfalcato@suse.de>
To: Vernon Yang <vernon2gm@gmail.com>
Cc: tglx@kernel.org, mingo@redhat.com, bp@alien8.de,
dave.hansen@linux.intel.com, akpm@linux-foundation.org,
david@kernel.org, hpa@zytor.com, rmclure@linux.ibm.com,
andrew+kernel@donnellan.id.au, pasha.tatashin@soleen.com,
kas@kernel.org, tj@kernel.org, rppt@kernel.org,
rick.p.edgecombe@intel.com, yu-cheng.yu@intel.com,
orsonpeters@gmail.com, linux-kernel@vger.kernel.org,
x86@kernel.org, linux-mm@kvack.org,
Vernon Yang <yanglincheng@kylinos.cn>,
stable@vger.kernel.org
Subject: Re: [PATCH] x86/mm: Fix pmd_modify() dropping the dirty bit
Date: Mon, 7 Sep 2026 10:51:50 +0100 [thread overview]
Message-ID: <ap6IYeTM8PMinxo-@pedro-suse.tail5790ac.ts.net> (raw)
In-Reply-To: <20260903031608.1194238-1-vernon2gm@gmail.com>
On Thu, Sep 03, 2026 at 11:16:08AM +0800, Vernon Yang wrote:
> From: Vernon Yang <yanglincheng@kylinos.cn>
>
> pmd_modify() masks the old value with (_HPAGE_CHG_MASK & ~_PAGE_DIRTY),
> silently discarding the hardware dirty bit. The subsequent
> pmd_mksaveddirty() call is supposed to transfer _PAGE_DIRTY into
> _PAGE_SAVED_DIRTY when write-protecting, but the dirty bit was already
> stripped from the value, so there is nothing left to transfer.
>
> Contrast with pte_modify(), which keeps _PAGE_DIRTY_BITS in its mask,
> and pud_modify(), which keeps _HPAGE_CHG_MASK untouched: pmd_modify()
> is the odd one out. Any pmd_modify() on a writable, dirty PMD loses
> the dirty state.
>
> One visible consequence is data loss with MADV_FREE on PMD-mapped THP:
>
> memset(buf, 0x5A, size); // PMD-mapped THP, PMD dirty
> madvise(buf, size, MADV_FREE); // PMD cleaned but left writable,
> // folio marked lazyfree
> memset(buf, 0x5A, size); // hardware sets _PAGE_DIRTY again
> mprotect(buf, size, PROT_READ); // pmd_modify() drops the dirty bit
> mprotect(buf, size, PROT_READ|PROT_WRITE);
> // ... memory pressure ...
>
> Reclaim (e.g. under memcg pressure) then finds the lazyfree folio with
> no dirty bit set anywhere and frees it in
> __discard_anon_folio_pmd_locked(), even though the data was rewritten
> after MADV_FREE; subsequent reads fault in fresh zero pages. NUMA
> hinting alone can trigger the same loss, as do_huge_pmd_numa_page()
> restores the PMD through pmd_modify() as well.
>
> PMD-mapped file THPs are affected too: mprotect()/NUMA hinting dropping
> the dirty bit means rewritten data is never written back.
>
> Fix it by keeping _PAGE_DIRTY in the preserved mask, exactly like
> pte_modify() and pud_modify() do. The existing
> pmd_mksaveddirty()/pmd_clear_saveddirty() pair then performs the
> hardware-dirty <-> saved-dirty transition based on the write bit,
> preserving the shadow-stack encoding rules.
>
> Closes: https://lore.kernel.org/r/CAJxLxMUGu1-L+O_nAONOwOXnS=cNbNApCWqdthRjd76LThtSPg@mail.gmail.com/
> Fixes: bb3aadf7d446 ("x86/mm: Start actually marking _PAGE_SAVED_DIRTY")
> Cc: stable@vger.kernel.org
> Signed-off-by: Vernon Yang <yanglincheng@kylinos.cn>
FYI we just had a customer case for our downstream kernel for this exact same
issue. This fixed it beautifully (I'm wondering if polars started doing
something interesting lately, that nobody else does, and that's why...).
Reviewed-by: Pedro Falcato <pfalcato@suse.de>
Tested-by: Pedro Falcato <pfalcato@suse.de>
--
Pedro
prev parent reply other threads:[~2026-09-07 9:52 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 3:16 [PATCH] x86/mm: Fix pmd_modify() dropping the dirty bit Vernon Yang
2026-09-03 4:10 ` Andrew Morton
2026-09-03 5:24 ` Vernon Yang
2026-09-03 7:34 ` Orson Peters
2026-09-03 15:40 ` Dave Hansen
2026-09-03 11:54 ` Kiryl Shutsemau
2026-09-03 14:18 ` Dave Hansen
2026-09-03 17:21 ` Andrew Morton
2026-09-03 17:58 ` Edgecombe, Rick P
2026-09-03 18:01 ` Dave Hansen
2026-09-03 18:33 ` Edgecombe, Rick P
2026-09-07 9:51 ` Pedro Falcato [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ap6IYeTM8PMinxo-@pedro-suse.tail5790ac.ts.net \
--to=pfalcato@suse.de \
--cc=akpm@linux-foundation.org \
--cc=andrew+kernel@donnellan.id.au \
--cc=bp@alien8.de \
--cc=dave.hansen@linux.intel.com \
--cc=david@kernel.org \
--cc=hpa@zytor.com \
--cc=kas@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mingo@redhat.com \
--cc=orsonpeters@gmail.com \
--cc=pasha.tatashin@soleen.com \
--cc=rick.p.edgecombe@intel.com \
--cc=rmclure@linux.ibm.com \
--cc=rppt@kernel.org \
--cc=stable@vger.kernel.org \
--cc=tglx@kernel.org \
--cc=tj@kernel.org \
--cc=vernon2gm@gmail.com \
--cc=x86@kernel.org \
--cc=yanglincheng@kylinos.cn \
--cc=yu-cheng.yu@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox