From: "Huang, Ying" <ying.huang@linux.alibaba.com>
To: Gregory Price <gourry@gourry.net>
Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org,
kernel-team@meta.com, akpm@linux-foundation.org,
david@kernel.org, ziy@nvidia.com, matthew.brost@intel.com,
joshua.hahnjy@gmail.com, rakie.kim@sk.com, byungchul@sk.com,
apopple@nvidia.com, urezki@gmail.com, chenwandun@huawei.com,
Chelsy Ratnawat <chelsyratnawat2001@gmail.com>,
stable@vger.kernel.org
Subject: Re: [PATCH] mm/mempolicy: take a cpuset cookie for the interleave node count
Date: Tue, 01 Sep 2026 19:24:21 +0800 [thread overview]
Message-ID: <877bl5qknu.fsf@DESKTOP-5N7EMDA> (raw)
In-Reply-To: <20260828193111.1023497-1-gourry@gourry.net> (Gregory Price's message of "Fri, 28 Aug 2026 15:31:11 -0400")
Gregory Price <gourry@gourry.net> writes:
> alloc_pages_bulk_interleave() counts pol->nodes without a cpuset cookie:
>
> nodes = nodes_weight(pol->nodes);
> nr_pages_per_node = nr_pages / nodes;
>
> nodemask_t spans several words once MAX_NUMNODES exceeds BITS_PER_LONG, so
> a concurrent cpuset rebind can tear that read and yield an empty mask even
> though neither version of it was empty. The call then allocates nothing
> and returns 0.
>
> Some compilers will hoist the loop entry test above the division,
> because nr_pages_per_node is dead when the loop does not run.
>
> 682e: call ... <- nodes_weight()
> 6838: test %eax,%eax
> 683a: jle 692d <- nodes <= 0 skips the loop
> 684a: div %rcx
>
> So in most deployments, this div/0 is unreachable - but nothing in the
> source guarantees that, it's just not easily exercised.
>
> Take the cookie around the count and bail if the mask really is empty.
> Only the count needs it, interleave_nodes() takes the cookie itself so
> so a torn read there is already retried.
>
> A rebind landing mid-loop can still leave the count disagreeing with the
> mask, so the loop may revisit a node or skip one - but a rebind where
> nodes change causes migration, so a handful of misplaced pages isn't
> catastrophic in any sense.
>
> Measured on a 72 node VM (NODES_SHIFT=10) with a cgroup v2 cpuset flipping
> cpuset.mems between a word 0 and a word 1 node set, and the two word read
> artificially widened: 330 zero counts in 130414 calls without the cookie,
> and 401 retries with it.
>
> Reported-by: Chelsy Ratnawat <chelsyratnawat2001@gmail.com>
> Link: https://lore.kernel.org/all/20250907160829.91628-1-chelsyratnawat2001@gmail.com/
> Fixes: c00b6b961099 ("mm/vmalloc: introduce alloc_pages_bulk_array_mempolicy to accelerate memory allocation")
> Cc: stable@vger.kernel.org
> Assisted-by: Claude:claude-opus-5
> Signed-off-by: Gregory Price (Meta) <gourry@gourry.net>
> ---
> mm/mempolicy.c | 12 +++++++++++-
> 1 file changed, 11 insertions(+), 1 deletion(-)
>
> diff --git a/mm/mempolicy.c b/mm/mempolicy.c
> index 79053ece02cd..060a0eb26917 100644
> --- a/mm/mempolicy.c
> +++ b/mm/mempolicy.c
> @@ -2592,6 +2592,7 @@ static unsigned long alloc_pages_bulk_interleave(gfp_t gfp,
> struct mempolicy *pol, unsigned long nr_pages,
> struct page **page_array)
> {
> + unsigned int cpuset_mems_cookie;
> int nodes;
> unsigned long nr_pages_per_node;
> int delta;
> @@ -2599,7 +2600,16 @@ static unsigned long alloc_pages_bulk_interleave(gfp_t gfp,
> unsigned long nr_allocated;
> unsigned long total_allocated = 0;
>
> - nodes = nodes_weight(pol->nodes);
> + /* count the nodes, retry if a rebind happened during the read */
> + do {
> + cpuset_mems_cookie = read_mems_allowed_begin();
> + nodes = nodes_weight(pol->nodes);
> + } while (read_mems_allowed_retry(cpuset_mems_cookie));
> +
> + /* if the nodemask has become invalid, we cannot do anything */
> + if (!nodes)
> + return 0;
> +
> nr_pages_per_node = nr_pages / nodes;
> delta = nr_pages - nodes * nr_pages_per_node;
LGTM, feel free to add my
Reviewed-by: Huang Ying <ying.huang@linux.alibaba.com>
in the future version.
---
Best Regards,
Huang, Ying
next prev parent reply other threads:[~2026-09-01 11:24 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-28 19:31 [PATCH] mm/mempolicy: take a cpuset cookie for the interleave node count Gregory Price
2026-09-01 11:24 ` Huang, Ying [this message]
2026-09-01 15:54 ` Gregory Price
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=877bl5qknu.fsf@DESKTOP-5N7EMDA \
--to=ying.huang@linux.alibaba.com \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=byungchul@sk.com \
--cc=chelsyratnawat2001@gmail.com \
--cc=chenwandun@huawei.com \
--cc=david@kernel.org \
--cc=gourry@gourry.net \
--cc=joshua.hahnjy@gmail.com \
--cc=kernel-team@meta.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=matthew.brost@intel.com \
--cc=rakie.kim@sk.com \
--cc=stable@vger.kernel.org \
--cc=urezki@gmail.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox