From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9F9D13D9058 for ; Fri, 24 Jul 2026 16:54:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784912048; cv=none; b=aPkK0zzx7hMzTQIBWACt+gBmzNCv+uGJQWQNL/rmfldZ07A4FiUf484EHyS5uTeyopQnWYzHQ3si8x6xwsAy3u4xWc4MsuHiNQP+k8N2bNZjjJRcg8yK3hPWoP21b7ZypVx4Jitaior7AWVHtWomppUOjxaUYw68myzDPVmD2yY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784912048; c=relaxed/simple; bh=0L24YJpXuFl+ogFcPH3o7FIX7j0vtDJCx7HUJRT8n8o=; h=From:To:Cc:Subject:In-Reply-To:References:Date:Message-ID: MIME-Version:Content-Type; b=DcgNBG/E4RA1mEYkc5BTqKtuu1zLNi/wNHI4G5M1Da0LltIeDwKu7J2LRFRseMfOc6bgci5ByLTMDoVMafjCcnBAb9PD5SVpw+0kXsOoDPQfk4vNX4Aw0FEY71U048ZH4f4FdzUoGpeKDol2DMsql5Nm0mp5NI0zadBhsMogp0A= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Yxe39jIv; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Yxe39jIv" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 0B3291F000E9; Fri, 24 Jul 2026 16:54:04 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1784912047; bh=uNcmPxfRJQMtu5X0pj3wcAJCBsBKUUg7nsN4VhDio8k=; h=From:To:Cc:Subject:In-Reply-To:References:Date; b=Yxe39jIvTCcqCfDHGXA9bJRbgJwrAzkgmnmNq/txHjKzKdJSjcqFSGp338WI2WcOh JNM+sR68lVt7AEID/otkdwv/7fK5enEazDCwQqgMvQu0QeqMQiZJtcvoR2/st3Edco 78otWKQOtrReQtPWBgwpIcXClrZcLDljk/tj+K87zZ5NRnnmfEvPQ5LIQfYHrQ7ROo 64pkPRyj6JlTsIhc3wZ5VF8fEXVcGLPfmOf2OcVEfIq3mwYE8/xQG5k6oQg7dKcN5G xDc1jiT20zPWL7YRk2QWS7ZDVheuKx9TJek4/5v7B/sZHnsigg74iCzNYEJ6ypoETW gCPLpM5dpaxlw== From: Pratyush Yadav To: Mike Rapoport Cc: Pratyush Yadav , Pasha Tatashin , Alexander Graf , Muchun Song , Oscar Salvador , David Hildenbrand , Andrew Morton , Jason Miu , Jork Loeser , kexec@lists.infradead.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH v3 18/21] kho: extend scratch In-Reply-To: <178410811939.820181.9175877917181976639.b4-review@b4> (Mike Rapoport's message of "Wed, 15 Jul 2026 12:35:19 +0300") References: <20260709173821.429921-1-pratyush@kernel.org> <20260709173821.429921-19-pratyush@kernel.org> <178410811939.820181.9175877917181976639.b4-review@b4> Date: Fri, 24 Jul 2026 18:54:03 +0200 Message-ID: <2vxzjyqkgwgk.fsf@kernel.org> User-Agent: Gnus/5.13 (Gnus v5.13) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain On Wed, Jul 15 2026, Mike Rapoport wrote: >> Motivation >> ========== >> >> The scratch space is allocated by the first kernel in the KHO chain, and >> is reused by all subsequent kernels. The size of the space is either set >> via the commandline by the system administrator or by calculating the >> amount of memory used by the kernel and adding a multiplier. In either >> case, the scratch space is a heuristic and is liable to fill up and fail >> allocation if a kernel uses more memory than expected. >> >> In addition, gigantic huge pages (usually 1 GiB) are allocated via >> memblock, and in a KHO boot that memory comes from the scratch space. In >> hypervisors it is common to dedicate a major part of the system's memory >> to gigantic hugepages for VM memory. >> >> If this memory needs to come from scratch space, then scratch needs to >> be greater than the memory needed for huge pages, which is impractical. >> In addition, hugepages can be preserved memory. Allocating them from >> scratch violates the assumption that scratch contains no preserved >> memory. >> >> Methodology >> =========== >> >> Discover areas that don't contain any preserved memory at boot by >> walking the preserved memory radix tree. Mark them as scratch to allow >> allocations from them. This makes KHO more resilient to memory pressure >> and allows supporting huge page preservation. >> >> Since the preserved memory radix tree mixes both physical address and >> order into a single key, and does not track table pages, it is difficult >> to identify free areas from it directly. Walk the tree and digest it >> down into another radix tree. The latter tracks blocks of >> KHO_EXT_SHIFT (1 GiB as of now) granularity. Then walk the digested tree >> and mark the areas between the present keys as scratch. >> >> Performance >> =========== >> >> The discovery algorithm traverses the preserved memory radix tree >> exactly once. While it does use memory for the digested radix tree, >> since the blocks are split by 1 GiB, a single bitmap with 4k pages can >> track up to 32 TiB of memory. So there are likely to be very few radix >> tree pages used in this tracking. For systems with all physical memory >> below 32 TiB, this should result in a total of 6 pages being >> used (KHO_TREE_MAX_DEPTH == 6). >> >> An alternate way of achieving this would be to call kho_mem_retrieve() >> earlier in boot and mark all the KHO preservations as reserved. But that >> can blow up memblock.reserved with a bunch of 4K pages scattered >> everywhere, which will reduce performance of subsequent allocations. >> Since the free blocks are tracked in chunks of 1 GiB, this won't blow up >> memblock.memory as much. >> >> There is no inherent reason for using 1 GiB as the discovered block >> size. This can be changed later if needed. Currently, KHO is mainly >> targeted for server grade systems with hundreds of gigabytes to >> terabytes of memory. So 1 GiB is a reasonable granularity for those >> systems. For smaller systems this doesn't work as well, but we can >> arrive at a better heuristic when we have concrete use cases. >> >> Practical evaluation >> ==================== >> >> The testing is done on a x86_64 qemu VM running under KVM with 64G >> memory and 12 CPUs. The machine pre-allocates 50 1G pages. >> >> Since the performance scales with how busy the radix tree is, tests are >> done with 2 preservation patterns: first with two 1M memfds, second with >> two 1G memfds, both using 4k pages. >> >> Test case 1 - 1M memfd >> ~~~~~~~~~~~~~~~~~~~~~~ >> >> This test case has two memfds with 1M memory each in 4k pages, plus >> other preservations from LUO core and other KHO users. >> >> This is how the radix tree stats look like: >> >> radix_nodes: 0x13 >> nr_preservations: 0x214 >> mem_preserved: 0x227000 >> >> per order preservations: >> order 0: 0x20f >> order 1: 0x4 >> order 4: 0x1 >> >> and this is how long it takes to extend the scratch after KHO boot: >> >> KHO: KHO extend time: 47 us >> KHO: KHO extend total mem: 0xe6c17b000 (~57G) >> >> Test case 2 - 1G memfd >> ~~~~~~~~~~~~~~~~~~~~~~ >> >> This test case has two memfds with 1G memory each in 4k pages, plus >> other preservations from LUO core and other KHO users. >> >> This is how the radix tree stats look like: >> >> radix_nodes: 0x28 >> nr_preservations: 0x80816 >> mem_preserved: 0x80829000 >> >> per order preservations: >> order 0: 0x80811 >> order 1: 0x4 >> order 4: 0x1 >> >> and this is how long it takes to extend the scratch after KHO boot: >> >> KHO: KHO extend time: 22514 us >> KHO: KHO extend total mem: 0xd3f200000 (~52G) >> >> Signed-off-by: Pratyush Yadav (Google) >> >> diff --git a/kernel/liveupdate/kexec_handover.c b/kernel/liveupdate/kexec_handover.c >> index b400976851e76..2c48488746290 100644 >> --- a/kernel/liveupdate/kexec_handover.c >> +++ b/kernel/liveupdate/kexec_handover.c >> @@ -84,6 +84,23 @@ static struct kho_out kho_out = { >> }, >> }; >> >> +struct kho_in { >> + phys_addr_t fdt_phys; >> + phys_addr_t scratch_phys; >> + char previous_release[__NEW_UTS_LEN + 1]; >> + u32 kexec_count; >> + struct kho_debugfs dbg; >> + struct kho_radix_tree radix_tree; >> +}; >> + >> +static struct kho_in kho_in = { >> +}; >> + >> +static const void *kho_get_fdt(void) >> +{ >> + return kho_in.fdt_phys ? phys_to_virt(kho_in.fdt_phys) : NULL; >> +} >> + >> /** >> * kho_encode_radix_key - Encodes a physical address and order into a radix key. >> * @phys: The physical address of the page. >> @@ -895,6 +912,128 @@ static void __init kho_reserve_scratch(void) >> kho_enable = false; >> } >> >> +/* >> + * Look for free blocks of 1G. This is a heuristic chosen to work efficiently >> + * with large systems with hundreds of gigabytes of memory. It will work poorly >> + * on smaller systems. The algorithm itself doesn't depend on the actual value, >> + * so it can be changed to a different heuristic later if needed. >> + */ >> +#define KHO_EXT_BLKSIZE SZ_1G >> +#define KHO_EXT_SHIFT const_ilog2(KHO_EXT_BLKSIZE) > > Nit: EXT_ what? ;-) > >> + >> +/* Called for the KHO preserved memory radix tree. */ >> +static int __init kho_ext_walk_leaf(unsigned long key, void *data) > > I have to say this function still confuses me :) Honestly, I think this whole tree walking API makes things really hard to reason about. I think we should replace it with something like kho_radix_for_each(). It would improve code readability a lot. I plan to do it at some point, but not as a part of this series. > >> +{ >> + /* Radix tree tracking free blocks. */ >> + struct kho_radix_tree *tree = data; >> + phys_addr_t start, end; >> + unsigned int order; >> + int err; >> + >> + start = kho_decode_radix_key(key, &order); > > I would add a comment here, that the key is from the memory tracker and > it is decoded to an address. ACK, will add. -- Regards, Pratyush Yadav