All of lore.kernel.org
 help / color / mirror / Atom feed
From: Sourabh Jain <sourabhjain@linux.ibm.com>
To: George Guo <dongtai.guo@linux.dev>,
	rppt@kernel.org, pasha.tatashin@soleen.com, pratyush@kernel.org
Cc: graf@amazon.com, changyuanl@google.com,
	akpm@linux-foundation.org, chenhuacai@kernel.org,
	liukexin@kylinos.cn, guodongtai@kylinos.cn,
	kexec@lists.infradead.org, linux-mm@kvack.org,
	loongarch@lists.linux.dev, linux-kernel@vger.kernel.org
Subject: Re: [PATCH 1/1] liveupdate: kho: calculate per-node scratch sizes before allocation
Date: Sat, 12 Sep 2026 12:03:25 +0530	[thread overview]
Message-ID: <616daf17-4598-4a30-8574-16480dfc23cb@linux.ibm.com> (raw)
In-Reply-To: <20260904025101.9959-1-dongtai.guo@linux.dev>

Hello George,

On 04/09/26 08:21, George Guo wrote:
> From: George Guo <guodongtai@kylinos.cn>
>
> The default percentage-based policy sizes scratch areas from the current
> kernel's MEMBLOCK_RSRV_KERN footprint. This is a reasonable heuristic for
> predicting the early memory demand of the next kernel.
>
> scratch_size_update() calculates the lowmem and global sizes before either
> area is allocated. However, kho_reserve_scratch() calculates each per-node
> size only after allocating the lowmem and global areas. Since memblock
> allocations are marked MEMBLOCK_RSRV_KERN, the per-node calculation
> includes those newly allocated scratch areas and scales them again.

I may be missing something here, but doesn't memblock_reserved_kern_size()
check the node ID of the region before checking the region flag 
(MEMBLOCK_RSRV_KERN)?

Code snippet from memblock_reserved_kern_size()
```
if (nid == memblock_get_region_node(r) || !numa_valid_node(nid))
     if (r->flags & MEMBLOCK_RSRV_KERN)
         total += size;
```

For a valid nid, my understanding is that the global and lowmem scratch
areas should not be counted because they are allocated with NUMA_NO_NODE
(-1). So, ideally, these regions should be excluded when calculating the
reserved memory for a specific node ID.

Based on this, I am not sure that marking the lowmem and global areas as
MEMBLOCK_RSRV_KERN is what causes the per-node size calculation to be
inflated. I am looking into the code further to better understand the
actual cause of the issue that this patch is trying to address.

With that said, I wonder if this fix might be more of a stop-gap solution.
As mentioned above, since NUMA_NO_NODE (-1) is used for the lowmem and
global allocations, my understanding is that these areas ideally should
not be included when calculating the size for a specific node ID.

I could be missing something in my understanding, so I would appreciate
your thoughts on these observations.

- Sourabh Jain

> Fixes: 3dc92c311498 ("kexec: add Kexec HandOver (KHO) generation helpers")
> Reported-by: Kexin Liu <liukexin@kylinos.cn>
> Co-developed-by: Kexin Liu <liukexin@kylinos.cn>
> Signed-off-by: Kexin Liu <liukexin@kylinos.cn>
> Signed-off-by: George Guo <guodongtai@kylinos.cn>
> ---
>   kernel/liveupdate/kexec_handover.c | 13 ++++++++++++-
>   1 file changed, 12 insertions(+), 1 deletion(-)
>
> diff --git a/kernel/liveupdate/kexec_handover.c b/kernel/liveupdate/kexec_handover.c
> index 7c4d86daf86d..39f489a258d9 100644
> --- a/kernel/liveupdate/kexec_handover.c
> +++ b/kernel/liveupdate/kexec_handover.c
> @@ -847,6 +847,17 @@ static void __init kho_reserve_scratch(void)
>   		goto err_disable_kho;
>   	}
>   
> +	/*
> +	 * Calculate the per-node sizes before reserving any scratch areas.
> +	 * memblock allocations are marked MEMBLOCK_RSRV_KERN, so calculating
> +	 * them later would count the lowmem and global scratch areas as kernel
> +	 * allocations and scale them again.
> +	 */
> +	i = 2;
> +	for_each_node_state(nid, N_MEMORY)
> +		kho_scratch[i++].size = scratch_size_node(nid);
> +	i = 0;
> +
>   	/*
>   	 * reserve scratch area in low memory for lowmem allocations in the
>   	 * next kernel
> @@ -880,7 +891,7 @@ static void __init kho_reserve_scratch(void)
>   	 * memoryless nodes, as we can not allocate scratch areas there.
>   	 */
>   	for_each_node_state(nid, N_MEMORY) {
> -		size = scratch_size_node(nid);
> +		size = kho_scratch[i].size;
>   		addr = memblock_alloc_range_nid(size, SCRATCH_ALIGNMENT_BYTES,
>   						0, MEMBLOCK_ALLOC_ACCESSIBLE,
>   						nid, true);



  parent reply	other threads:[~2026-09-12  6:33 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04  2:51 [PATCH 1/1] liveupdate: kho: calculate per-node scratch sizes before allocation George Guo
2026-09-06 20:07 ` Mike Rapoport
2026-09-07 10:24   ` George Guo
2026-09-08  3:05     ` Sourabh Jain
2026-09-12  6:33 ` Sourabh Jain [this message]
2026-09-17 10:30   ` Sourabh Jain
2026-09-17 22:53     ` Pratyush Yadav
2026-09-18  9:33       ` George Guo
2026-09-22 13:19         ` Sourabh Jain
2026-09-21  5:31       ` Sourabh Jain

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=616daf17-4598-4a30-8574-16480dfc23cb@linux.ibm.com \
    --to=sourabhjain@linux.ibm.com \
    --cc=akpm@linux-foundation.org \
    --cc=changyuanl@google.com \
    --cc=chenhuacai@kernel.org \
    --cc=dongtai.guo@linux.dev \
    --cc=graf@amazon.com \
    --cc=guodongtai@kylinos.cn \
    --cc=kexec@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=liukexin@kylinos.cn \
    --cc=loongarch@lists.linux.dev \
    --cc=pasha.tatashin@soleen.com \
    --cc=pratyush@kernel.org \
    --cc=rppt@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.