From: Xu Yilun <yilun.xu@linux.intel.com>
To: Kiryl Shutsemau <kas@kernel.org>
Cc: x86@kernel.org, linux-coco@lists.linux.dev,
linux-kernel@vger.kernel.org, rick.p.edgecombe@intel.com,
yilun.xu@intel.com, xiaoyao.li@intel.com, sohil.mehta@intel.com,
adrian.hunter@intel.com, kishen.maloor@intel.com,
tony.lindgren@linux.intel.com, peter.fang@intel.com,
baolu.lu@linux.intel.com, zhenzhong.duan@intel.com,
chao.gao@intel.com, artem.bityutskiy@linux.intel.com,
kvm@vger.kernel.org
Subject: Re: [PATCH 4/6] x86/virt/tdx: Add extra memory to TDX module for the extensions
Date: Mon, 24 Aug 2026 17:18:25 +0800 [thread overview]
Message-ID: <aowMYbHKK4b0O5P6@yilunxu-OptiPlex-7050> (raw)
In-Reply-To: <aohvlMO7ehqcVW96@thinkstation>
> > The TDX module accepts the memory in the form of a PFN array. This array
> > is passed via a single 64-bit SEAMCALL leaf parameter, which encodes two
> > values: the PFN of the container page holding the array, and the number
> > of entries in the array. Create a helper to encode this format and name
> > it after the TDX module term: HPA_LIST_INFO.
>
> The array entries are physical addresses, not PFNs. HPA_LIST_INFO encodes
> a PFN, the array does not.
OK. I'll change PFN array => HPA array
[...]
> > +#define TDX_HPA_LIST_MAX_NR_PAGES (PAGE_SIZE / sizeof(u64))
> > +
> > +struct tdx_hpa_list {
> > + u64 phys[TDX_HPA_LIST_MAX_NR_PAGES];
> > +};
> > +
> > +static_assert(sizeof(struct tdx_hpa_list) == PAGE_SIZE);
[...]
> > + hpa_list = kzalloc_obj(*hpa_list);
> > + if (!hpa_list)
> > + return -ENOMEM;
>
> to_hpa_list_info() expects hpa_list to be page-aligned. It happens to
> work with kmalloc for PAGE_SIZE allocation.
The struct tdx_hpa_list definition follows the TDX ABI and is guarenteed
to be PAGE_SIZE by:
static_assert(sizeof(struct tdx_hpa_list) == PAGE_SIZE);
and kmalloc guarentees the page alignment.
59bb47985c1d ("mm, sl[aou]b: guarantee natural alignment for kmalloc(power-of-two)")
So I think it's OK, not "happen to work".
>
> Maybe it is better to allocate it with buddy allocator instead?
It can be, but then we need an extra variable to record the
struct page *, which seems redundant?
>
> > +
> > + page = alloc_contig_pages(required_pages, GFP_KERNEL, numa_mem_id(),
> > + &node_online_map);
>
> Why contiguous? TDH.EXT.MEM.ADD takes a list of page addresses and the loop
> below writes every one of them out separately.
>
> alloc_pages_bulk() fits the chunking that is already here, and a short
> return can be handled per chunk. alloc_contig_pages() isolates and migrates
> to get its range and fails TDX init outright when it cannot find one. PAMT
Yeah, this is not the ABI requirement, but the kernel's consideration. A
brief reasoning in the commit log: avoiding permanent memory fragmentation
and buddy allocator efficiency loss.
Also there is some discussion:
https://lore.kernel.org/all/167d9540-2d9a-4367-bc68-b96494bc4044@intel.com/
TL;DR
- The memory will never return to the kernel.
- There is chance that this tens of megabytes will fragment tens of
gigabytes of memory forever.
- The chance of fragmentation is actually low since at boot up, but
let the buddy allocator take care of these never-returned memory
is not necessary and lowers its efficiency.
> needs it because the TDMR ABI describes each PAMT as base+size. This does
> not.
>
> > + if (!page) {
> > + ret = -ENOMEM;
> > + goto out_free_hpa_list;
> > + }
> > +
> > + added_pages = 0;
> > + while (added_pages < required_pages) {
> > + unsigned int chunk_pages = min(required_pages - added_pages,
> > + TDX_HPA_LIST_MAX_NR_PAGES);
> > + struct page *chunk = page + added_pages;
> > + unsigned int i;
> > +
> > + for (i = 0; i < chunk_pages; i++)
> > + hpa_list->phys[i] = page_to_phys(chunk + i);
> > +
> > + ret = tdx_ext_mem_add(hpa_list, chunk_pages);
> > + if (ret) {
> > + /*
> > + * This SEAMCALL leaf shouldn't fail, and if it does,
> > + * things are broken enough that complex error handling
> > + * isn't worth it. Intentionally leak all pages,
> > + * including un-added pages.
> > + */
> > + WARN(1, "Fatal: TDX module rejected memory for extensions, stranded all pages\n");
> > + break;
>
> It supposed to be
> goto out_free_hpa_list;
>
> No?
The difference is to print the memory amount or not. For simple error
handling, we stranded all pages, we let users know the cost even if the
initialization fails.
>
>
> > + }
> > +
> > + added_pages += chunk_pages;
> > + }
> > +
> > + /*
> > + * Memory for TDX module extensions is never reclaimed and can be tens
> > + * of megabytes. Print the amount so users know the cost.
> > + */
> > + pr_info("%lu KB consumed for TDX module extensions\n",
> > + required_pages * PAGE_SIZE / 1024);
> > +
> > +out_free_hpa_list:
> > + kfree(hpa_list);
> > +
> > + return ret;
> > +}
> > +
[...]
> > --- a/arch/x86/virt/vmx/tdx/tdx_global_metadata.c
> > +++ b/arch/x86/virt/vmx/tdx/tdx_global_metadata.c
> > @@ -137,6 +137,12 @@ static __init int get_tdx_sys_info_ext(struct tdx_sys_info_ext *sysinfo_ext)
> > int ret;
> > u64 val;
> >
> > + ret = read_sys_metadata_field(0x3100000200000000, &val);
> > + if (ret)
> > + return ret;
> > +
> > + sysinfo_ext->memory_pool_required_pages = val;
> > +
>
> Why above ext_required read?
I want to sort them in ascending order of the field ID, so reviewers can
seach them more easily.
> Is it even valid to read it in such case?
It is OK. The two ext metadata are both valid after TDX feature
configurations.
The ext_required == 0 && memory_pool_required_pages > 0 is highly
suspicious based on our current understanding, but that's more of a
module BUG, not caused by metadata reading order.
>
> > ret = read_sys_metadata_field(0x3100000000000001, &val);
> > if (ret)
> > return ret;
> > --
> > 2.25.1
> >
>
> --
> Kiryl Shutsemau / Kirill A. Shutemov
next prev parent reply other threads:[~2026-08-24 9:18 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-21 3:29 [PATCH 0/6] Enable TDX module extensions Xu Yilun
2026-08-21 3:29 ` [PATCH 1/6] x86/virt/tdx: Wrap TDH.SYS.CONFIG/UPDATE operations in helpers Xu Yilun
2026-08-21 20:53 ` Edgecombe, Rick P
2026-08-24 4:52 ` Xu Yilun
2026-08-21 3:29 ` [PATCH 2/6] x86/virt/tdx: Configure add-on features on TDX module init and update Xu Yilun
2026-08-21 14:38 ` Dave Hansen
2026-08-21 21:18 ` Edgecombe, Rick P
2026-08-24 6:37 ` Xu Yilun
2026-08-21 22:01 ` Edgecombe, Rick P
2026-08-21 3:29 ` [PATCH 3/6] x86/virt/tdx: Detect if the extensions initialization is required Xu Yilun
2026-08-21 15:22 ` Kiryl Shutsemau
2026-08-21 22:22 ` Edgecombe, Rick P
2026-08-24 12:16 ` Kiryl Shutsemau
2026-08-21 3:29 ` [PATCH 4/6] x86/virt/tdx: Add extra memory to TDX module for the extensions Xu Yilun
2026-08-21 15:44 ` Kiryl Shutsemau
2026-08-24 9:18 ` Xu Yilun [this message]
2026-08-24 12:22 ` Kiryl Shutsemau
2026-08-21 3:29 ` [PATCH 5/6] x86/virt/tdx: Make TDX module initialize " Xu Yilun
2026-08-21 23:55 ` Edgecombe, Rick P
2026-08-21 3:29 ` [PATCH 6/6] x86/virt/tdx: Re-initialize the extensions on runtime TDX module update Xu Yilun
2026-08-22 0:01 ` Edgecombe, Rick P
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aowMYbHKK4b0O5P6@yilunxu-OptiPlex-7050 \
--to=yilun.xu@linux.intel.com \
--cc=adrian.hunter@intel.com \
--cc=artem.bityutskiy@linux.intel.com \
--cc=baolu.lu@linux.intel.com \
--cc=chao.gao@intel.com \
--cc=kas@kernel.org \
--cc=kishen.maloor@intel.com \
--cc=kvm@vger.kernel.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=peter.fang@intel.com \
--cc=rick.p.edgecombe@intel.com \
--cc=sohil.mehta@intel.com \
--cc=tony.lindgren@linux.intel.com \
--cc=x86@kernel.org \
--cc=xiaoyao.li@intel.com \
--cc=yilun.xu@intel.com \
--cc=zhenzhong.duan@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).