Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Pratyush Yadav <pratyush@kernel.org>
To: Mike Rapoport <rppt@kernel.org>
Cc: Pratyush Yadav <pratyush@kernel.org>,
	 Pasha Tatashin <pasha.tatashin@soleen.com>,
	 Alexander Graf <graf@amazon.com>,
	 Muchun Song <muchun.song@linux.dev>,
	 Oscar Salvador <osalvador@suse.de>,
	 David Hildenbrand <david@kernel.org>,
	 Andrew Morton <akpm@linux-foundation.org>,
	 Jason Miu <jasonmiu@google.com>,
	 Jork Loeser <jloeser@linux.microsoft.com>,
	 kexec@lists.infradead.org, linux-mm@kvack.org,
	 linux-kernel@vger.kernel.org
Subject: Re: [PATCH v3 17/21] mm/mm_init: don't rely on memblock to get KHO scratch migratetype
Date: Sat, 25 Jul 2026 17:56:08 +0200	[thread overview]
Message-ID: <2vxzfr17gj1j.fsf@kernel.org> (raw)
In-Reply-To: <178410811939.820181.7932476003120213021.b4-review@b4> (Mike Rapoport's message of "Wed, 15 Jul 2026 12:35:19 +0300")

On Wed, Jul 15 2026, Mike Rapoport wrote:

>> Currently struct page init via memmap_init() or deferred_init_memmap()
>> only queries the migrate type from KHO for each discrete memory range.
>> That works currently since KHO scratch memory has a different memory
>> type so it is always it its own region.
>> 
>> An upcoming patch will add support for discovering blocks of memory with
>> no preservations and it will mark it as MEMBLOCK_KHO_SCRATCH to allow
>> allocations from them. This can lead to the bootmem KHO scratch areas to
>> be merged into larger free ranges. This merging breaks the selection of
>> migrate type.
>> 
>> Get rid of kho_scratch_migratetype(). Instead, let the struct page init
>> logic set the migratetype to MIGRATE_MOVABLE for all pages, and then
>> give KHO a chance to update the migrate type of its bootmem scratch
>> areas to MIGRATE_CMA. Since scratch areas should be a few order of
>> magnitues smaller than the total system memory, doing it this way is
>> more efficient than searching the kho_scratch array for every pageblock.
>
> It's true for non-deferred case, but with deferred initialization of
> struct pages the static part is small and deferred chunks are small, so
> we are going to update the migratetype nearly as much as we would check
> it with kho_scratch_migratetype().
>
> For simplicity of this patchset I'd keep kho_scratch_migratetype(), just
> move it over to kexec_handover.h and make it rely on kho_scratch_overlap().
>
> The bulk optimization can be a separate patch on top.

Sure, will do.

>
>> Since kho_scratch_migratetype() no longer exists, drop the migratetype
>> argument to memmap_init_zone_range() and deferred_init_pages().
>> 
>> Signed-off-by: Pratyush Yadav (Google) <pratyush@kernel.org>
>>
[...]
>> +/*
>> + * Mark all KHO scratch pageblocks between start_pfn and end_pfn as
>> + * MIGRATE_CMA.
>> + */
>> +void kho_init_scratch_migratetype(unsigned long start_pfn, unsigned long end_pfn)
>> +{
>> +	for (unsigned int i = 0; i < kho_scratch_cnt; i++) {
>> +		unsigned long scratch_start, scratch_end, pfn;
>> +
>> +		scratch_start = PHYS_PFN(kho_scratch[i].addr);
>> +		scratch_end = PHYS_PFN(kho_scratch[i].addr + kho_scratch[i].size);
>> +
>> +		/* Target range doesn't overlap with scratch. */
>> +		if (start_pfn >= scratch_end || end_pfn <= scratch_start)
>> +			continue;
>> +
>> +		pfn = max_t(unsigned long, start_pfn, scratch_start);
>> +		pfn = pageblock_start_pfn(pfn);
>> +
>> +		while (pfn < end_pfn && pfn < scratch_end) {
>> +			init_pageblock_migratetype(pfn_to_page(pfn),
>> +						   MIGRATE_CMA, false);
>> +			pfn += pageblock_nr_pages;
>> +		}
>> +	}
>> +}
>
> static inline enum migratetype kho_scratch_migratetype(unsigned long pfn,
> 						       enum migratetype mt)
> {
> 	if (kho_scratch_overlap(PFN_PHYS(pfn)), PAGE_SIZE)
> 		return MIGRATE_CMA;
> 	return mt;
> }
>
> Is way easier to parse and reason about ;-)
>
>> +
>>  /**
>>   * kho_reserve_scratch - Reserve a contiguous chunk of memory for kexec
>>   *
[...]
>> @@ -2012,8 +2021,7 @@ static inline void __init pgdat_init_report_one_done(void)
>>   * Return number of pages initialized.
>>   */
>>  static unsigned long __init deferred_init_pages(struct zone *zone,
>> -		unsigned long start_pfn, unsigned long end_pfn,
>> -		enum migratetype mt)
>> +		unsigned long start_pfn, unsigned long end_pfn)
>>  {
>>  	int nid = zone_to_nid(zone);
>>  	unsigned long nr_pages = end_pfn - start_pfn, pfn = start_pfn;
>> @@ -2027,7 +2035,14 @@ static unsigned long __init deferred_init_pages(struct zone *zone,
>>  	pfn = pageblock_align(start_pfn);
>>  	page = pfn_to_page(pfn);
>>  	for (; pfn < end_pfn; pfn += pageblock_nr_pages, page += pageblock_nr_pages)
>> -		init_pageblock_migratetype(page, mt, false);
>> +		init_pageblock_migratetype(page, MIGRATE_MOVABLE, false);
>> +
>> +	/*
>> +	 * Update the migrate type for any pageblocks that fall in a KHO scratch
>> +	 * area. Since most pageblocks will not be KHO scratch, do it outside
>> +	 * the loop to only do the search once.
>> +	 */
>> +	kho_init_scratch_migratetype(start_pfn, end_pfn);
>
> And here the chunks are at most MAX_ORDER_NR_PAGES, so bulk init also
> wouldn't save much.

Hmm, right. I didn't think of that. On my system pageblock order is 9
and max page order is 10, so it does halve the amount of work done, but
still, I see your point. It took me a couple of tries to get
kho_init_scratch_migratetype() right, and kho_scratch_overlap() is
certainly easier to reason about.

Oh well, I suppose this gets added to the list of performance
improvements we do at some point down the line

-- 
Regards,
Pratyush Yadav


  parent reply	other threads:[~2026-07-25 15:56 UTC|newest]

Thread overview: 26+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-09 17:37 [PATCH v3 00/21] kho: make boot time huge page allocation work nicely with KHO Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 01/21] kho: generalize radix tree APIs Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 02/21] kho: make radix max key width more obvious Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 03/21] kho: disallow wide keys in radix tree Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 04/21] kho: return virtual address of mem_map Pratyush Yadav
     [not found]   ` <178410811938.820181.15709073379301681705.b4-review@b4>
2026-07-24 16:45     ` Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 05/21] kho: store incoming radix tree in kho_in Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 06/21] kho: move all memory retrieval logic to kho_mem_retrieve() Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 07/21] kho: add a struct for radix callbacks Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 08/21] kho: add callback for table pages Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 09/21] kho: add data argument to radix walk callback Pratyush Yadav
2026-07-09 17:37 ` [PATCH v3 10/21] kho: allow early-boot usage of the KHO radix tree Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 11/21] kho: allow destroying " Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 12/21] kho: add kho_radix_init_tree() Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 13/21] kho: expose kho_scratch_overlap() to kexec_handover.h Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 14/21] kho: initialize kho_scratch pointer earlier in boot Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 15/21] kho: initialize preserved memory map radix tree earlier Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 16/21] mm/mm_init: init deferred page migratetype in deferred_init_pages() Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 17/21] mm/mm_init: don't rely on memblock to get KHO scratch migratetype Pratyush Yadav
     [not found]   ` <178410811939.820181.7932476003120213021.b4-review@b4>
2026-07-25 15:56     ` Pratyush Yadav [this message]
2026-07-09 17:38 ` [PATCH v3 18/21] kho: extend scratch Pratyush Yadav
     [not found]   ` <178410811939.820181.9175877917181976639.b4-review@b4>
2026-07-24 16:54     ` Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 19/21] memblock: make HugeTLB bootmem allocation work with KHO Pratyush Yadav
     [not found]   ` <178410811939.820181.7248948033992831535.b4-review@b4>
2026-07-24 16:51     ` Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 20/21] memblock: add memblock_reserved_hugetlb_size() Pratyush Yadav
2026-07-09 17:38 ` [PATCH v3 21/21] kho: exclude hugetlb memory from scratch size calculation Pratyush Yadav

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=2vxzfr17gj1j.fsf@kernel.org \
    --to=pratyush@kernel.org \
    --cc=akpm@linux-foundation.org \
    --cc=david@kernel.org \
    --cc=graf@amazon.com \
    --cc=jasonmiu@google.com \
    --cc=jloeser@linux.microsoft.com \
    --cc=kexec@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=muchun.song@linux.dev \
    --cc=osalvador@suse.de \
    --cc=pasha.tatashin@soleen.com \
    --cc=rppt@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox