From: Rakie Kim <rakie.kim@sk.com>
To: Gregory Price <gourry@gourry.net>
Cc: linux-kernel@vger.kernel.org, kernel-team@meta.com,
akpm@linux-foundation.org, david@kernel.org, ziy@nvidia.com,
matthew.brost@intel.com, joshua.hahnjy@gmail.com,
byungchul@sk.com, ying.huang@linux.alibaba.com,
apopple@nvidia.com, urezki@gmail.com, chenwandun@huawei.com,
linux-mm@kvack.org, kernel_team@skhynix.com,
Rakie Kim <rakie.kim@sk.com>
Subject: Re: [PATCH 2/2] mm/mempolicy: stop copying the nodemask in the interleave paths
Date: Fri, 4 Sep 2026 17:01:37 +0900 [thread overview]
Message-ID: <20260904080140.1992-1-rakie.kim@sk.com> (raw)
In-Reply-To: <apgzMZQPTxS9QtDy@gourry-fedora-PF4VCD3F>
On Wed, 2 Sep 2026 10:52:45 -0400 Gregory Price <gourry@gourry.net> wrote:
> I don't think this actually fixes anything?
>
> But basically the proposal is to widen the SRCU() window further to
> include the cpuset cookie entirely.
>
> e.g.
>
> SRCU() {
> cpuset_cookie() {
> for_each_node_mask(node, pol->nodes)
> weight_total += ...
> nnodes++;
> }
>
> /* ... snip - single node quick-exit ... */
>
> /* ... actual multi-node bulk allocation ... */
> for_each_node_mask(node, pol->nodes) {
> nr_allocated = __alloc_pages_bulk(gfp, node, ...);
>
> /*
> * At this point, due to a torn read from pol->nodes
> * we can visit a node that wasn't present previously
> * or we can skip a node that was present previously.
> *
> * In either case, weight_total is the wrong value for
> * the set of nodes being walked anyway - we are going
> * to skew in the distribution no matter what.
> */
> }
> }
>
> I'm not sure widening the SRCU window is worth it here, it doesn't
> actually buy us anything.
>
> Also we'd be calculating the weight total every time even when there's a
> scenario where we quick-exit because the entire allocation fits in the
> first node in the mask.
Sorry, I did not explain that well. I was not suggesting moving the
sum into the cookie loop or widening the SRCU section - only adding
the counter to the sum loop where it already is:
/* calculate total, detect system default usage */
nnodes = 0;
for_each_node_mask(node, pol->nodes) {
weight_total += table ? table[node] : 1;
nnodes++;
}
The order of the function does not change, so the quick-exit path
still returns before this loop runs - the sum is not computed in
that case either way.
What I had in mind is the gap between the two reads. Say the policy
starts with two nodes, every weight is 10, and 100 pages are
requested:
/* the mask is {0,1} here */
do {
cpuset_mems_cookie = read_mems_allowed_begin();
nnodes = nodes_weight(pol->nodes); /* nnodes = 2 */
} while (read_mems_allowed_retry(cpuset_mems_cookie));
/* a rebind grows the mask to {0,1,2,3} at this point */
/* calculate total, detect system default usage */
for_each_node_mask(node, pol->nodes)
weight_total += ...; /* 10 * 4 = 40 */
rounds = rem_pages / weight_total; /* 100 / 40 = 2 */
for (i = 0; i < nnodes; i++) /* bounded by 2 */
...
The distribution is planned from a total of 40, which spreads the
100 pages over four nodes, but the walk is still bounded by the two
nodes counted earlier. It places 60 pages and returns, and the
caller allocates the remaining 40 one page at a time. The mask does
not have to change again during the walk for this to happen.
You are right that a torn read during the walk still skews the
distribution, and this does not change that. It only removes the
case where the bound and the total start out inconsistent. If that
is not worth the extra line, I am fine either way.
Thanks for looking at it.
Rakie Kim
next prev parent reply other threads:[~2026-09-04 8:01 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-29 1:59 [PATCH 0/2] mm/mempolicy: stop copying state in the interleave paths Gregory Price
2026-08-29 1:59 ` [PATCH 1/2] mm/mempolicy: use SRCU for the weighted interleave state Gregory Price
2026-08-29 1:59 ` [PATCH 2/2] mm/mempolicy: stop copying the nodemask in the interleave paths Gregory Price
2026-09-02 9:00 ` Rakie Kim
2026-09-02 14:52 ` Gregory Price
2026-09-04 8:01 ` Rakie Kim [this message]
2026-08-29 23:18 ` [PATCH 0/2] mm/mempolicy: stop copying state " Andrew Morton
2026-08-30 16:39 ` Gregory Price
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260904080140.1992-1-rakie.kim@sk.com \
--to=rakie.kim@sk.com \
--cc=akpm@linux-foundation.org \
--cc=apopple@nvidia.com \
--cc=byungchul@sk.com \
--cc=chenwandun@huawei.com \
--cc=david@kernel.org \
--cc=gourry@gourry.net \
--cc=joshua.hahnjy@gmail.com \
--cc=kernel-team@meta.com \
--cc=kernel_team@skhynix.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=matthew.brost@intel.com \
--cc=urezki@gmail.com \
--cc=ying.huang@linux.alibaba.com \
--cc=ziy@nvidia.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox