linux-coco.lists.linux.dev archive mirror
 help / color / mirror / Atom feed
From: Xu Yilun <yilun.xu@linux.intel.com>
To: Kiryl Shutsemau <kas@kernel.org>
Cc: x86@kernel.org, linux-coco@lists.linux.dev,
	linux-kernel@vger.kernel.org, rick.p.edgecombe@intel.com,
	yilun.xu@intel.com, xiaoyao.li@intel.com, sohil.mehta@intel.com,
	adrian.hunter@intel.com, kishen.maloor@intel.com,
	tony.lindgren@linux.intel.com, peter.fang@intel.com,
	baolu.lu@linux.intel.com, zhenzhong.duan@intel.com,
	chao.gao@intel.com, artem.bityutskiy@linux.intel.com,
	kvm@vger.kernel.org
Subject: Re: [PATCH 4/6] x86/virt/tdx: Add extra memory to TDX module for the extensions
Date: Mon, 24 Aug 2026 17:18:25 +0800	[thread overview]
Message-ID: <aowMYbHKK4b0O5P6@yilunxu-OptiPlex-7050> (raw)
In-Reply-To: <aohvlMO7ehqcVW96@thinkstation>

> > The TDX module accepts the memory in the form of a PFN array. This array
> > is passed via a single 64-bit SEAMCALL leaf parameter, which encodes two
> > values: the PFN of the container page holding the array, and the number
> > of entries in the array. Create a helper to encode this format and name
> > it after the TDX module term: HPA_LIST_INFO.
> 
> The array entries are physical addresses, not PFNs. HPA_LIST_INFO encodes
> a PFN, the array does not.

OK. I'll change PFN array => HPA array

[...]

> > +#define TDX_HPA_LIST_MAX_NR_PAGES	(PAGE_SIZE / sizeof(u64))
> > +
> > +struct tdx_hpa_list {
> > +	u64 phys[TDX_HPA_LIST_MAX_NR_PAGES];
> > +};
> > +
> > +static_assert(sizeof(struct tdx_hpa_list) == PAGE_SIZE);

[...]

> > +	hpa_list = kzalloc_obj(*hpa_list);
> > +	if (!hpa_list)
> > +		return -ENOMEM;
> 
> to_hpa_list_info() expects hpa_list to be page-aligned. It happens to
> work with kmalloc for PAGE_SIZE allocation.

The struct tdx_hpa_list definition follows the TDX ABI and is guarenteed
to be PAGE_SIZE by:

	static_assert(sizeof(struct tdx_hpa_list) == PAGE_SIZE);

and kmalloc guarentees the page alignment.

	59bb47985c1d ("mm, sl[aou]b: guarantee natural alignment for kmalloc(power-of-two)")

So I think it's OK, not "happen to work".

> 
> Maybe it is better to allocate it with buddy allocator instead?

It can be, but then we need an extra variable to record the
struct page *, which seems redundant?

> 
> > +
> > +	page = alloc_contig_pages(required_pages, GFP_KERNEL, numa_mem_id(),
> > +				  &node_online_map);
> 
> Why contiguous? TDH.EXT.MEM.ADD takes a list of page addresses and the loop
> below writes every one of them out separately.
> 
> alloc_pages_bulk() fits the chunking that is already here, and a short
> return can be handled per chunk. alloc_contig_pages() isolates and migrates
> to get its range and fails TDX init outright when it cannot find one. PAMT

Yeah, this is not the ABI requirement, but the kernel's consideration. A
brief reasoning in the commit log: avoiding permanent memory fragmentation
and buddy allocator efficiency loss.

Also there is some discussion:

https://lore.kernel.org/all/167d9540-2d9a-4367-bc68-b96494bc4044@intel.com/

TL;DR
  - The memory will never return to the kernel.
  - There is chance that this tens of megabytes will fragment tens of
    gigabytes of memory forever.
  - The chance of fragmentation is actually low since at boot up, but
    let the buddy allocator take care of these never-returned memory
    is not necessary and lowers its efficiency.

> needs it because the TDMR ABI describes each PAMT as base+size. This does
> not.
> 
> > +	if (!page) {
> > +		ret = -ENOMEM;
> > +		goto out_free_hpa_list;
> > +	}
> > +
> > +	added_pages = 0;
> > +	while (added_pages < required_pages) {
> > +		unsigned int chunk_pages = min(required_pages - added_pages,
> > +					       TDX_HPA_LIST_MAX_NR_PAGES);
> > +		struct page *chunk = page + added_pages;
> > +		unsigned int i;
> > +
> > +		for (i = 0; i < chunk_pages; i++)
> > +			hpa_list->phys[i] = page_to_phys(chunk + i);
> > +
> > +		ret = tdx_ext_mem_add(hpa_list, chunk_pages);
> > +		if (ret) {
> > +			/*
> > +			 * This SEAMCALL leaf shouldn't fail, and if it does,
> > +			 * things are broken enough that complex error handling
> > +			 * isn't worth it. Intentionally leak all pages,
> > +			 * including un-added pages.
> > +			 */
> > +			WARN(1, "Fatal: TDX module rejected memory for extensions, stranded all pages\n");
> > +			break;
> 
> It supposed to be
> 			goto out_free_hpa_list;
> 
> No?

The difference is to print the memory amount or not. For simple error
handling, we stranded all pages, we let users know the cost even if the
initialization fails.

> 
> 
> > +		}
> > +
> > +		added_pages += chunk_pages;
> > +	}
> > +
> > +	/*
> > +	 * Memory for TDX module extensions is never reclaimed and can be tens
> > +	 * of megabytes. Print the amount so users know the cost.
> > +	 */
> > +	pr_info("%lu KB consumed for TDX module extensions\n",
> > +		required_pages * PAGE_SIZE / 1024);
> > +
> > +out_free_hpa_list:
> > +	kfree(hpa_list);
> > +
> > +	return ret;
> > +}
> > +

[...]

> > --- a/arch/x86/virt/vmx/tdx/tdx_global_metadata.c
> > +++ b/arch/x86/virt/vmx/tdx/tdx_global_metadata.c
> > @@ -137,6 +137,12 @@ static __init int get_tdx_sys_info_ext(struct tdx_sys_info_ext *sysinfo_ext)
> >  	int ret;
> >  	u64 val;
> >  
> > +	ret = read_sys_metadata_field(0x3100000200000000, &val);
> > +	if (ret)
> > +		return ret;
> > +
> > +	sysinfo_ext->memory_pool_required_pages = val;
> > +
> 
> Why above ext_required read?

I want to sort them in ascending order of the field ID, so reviewers can
seach them more easily.

> Is it even valid to read it in such case?

It is OK. The two ext metadata are both valid after TDX feature
configurations.

The ext_required == 0 && memory_pool_required_pages > 0 is highly
suspicious based on our current understanding, but that's more of a
module BUG, not caused by metadata reading order.

> 
> >  	ret = read_sys_metadata_field(0x3100000000000001, &val);
> >  	if (ret)
> >  		return ret;
> > -- 
> > 2.25.1
> > 
> 
> -- 
>   Kiryl Shutsemau / Kirill A. Shutemov

  reply	other threads:[~2026-08-24  9:18 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-21  3:29 [PATCH 0/6] Enable TDX module extensions Xu Yilun
2026-08-21  3:29 ` [PATCH 1/6] x86/virt/tdx: Wrap TDH.SYS.CONFIG/UPDATE operations in helpers Xu Yilun
2026-08-21 20:53   ` Edgecombe, Rick P
2026-08-24  4:52     ` Xu Yilun
2026-08-21  3:29 ` [PATCH 2/6] x86/virt/tdx: Configure add-on features on TDX module init and update Xu Yilun
2026-08-21 14:38   ` Dave Hansen
2026-08-21 21:18     ` Edgecombe, Rick P
2026-08-24  6:37     ` Xu Yilun
2026-08-21 22:01   ` Edgecombe, Rick P
2026-08-21  3:29 ` [PATCH 3/6] x86/virt/tdx: Detect if the extensions initialization is required Xu Yilun
2026-08-21 15:22   ` Kiryl Shutsemau
2026-08-21 22:22     ` Edgecombe, Rick P
2026-08-24 12:16       ` Kiryl Shutsemau
2026-08-21  3:29 ` [PATCH 4/6] x86/virt/tdx: Add extra memory to TDX module for the extensions Xu Yilun
2026-08-21 15:44   ` Kiryl Shutsemau
2026-08-24  9:18     ` Xu Yilun [this message]
2026-08-24 12:22       ` Kiryl Shutsemau
2026-08-21  3:29 ` [PATCH 5/6] x86/virt/tdx: Make TDX module initialize " Xu Yilun
2026-08-21 23:55   ` Edgecombe, Rick P
2026-08-21  3:29 ` [PATCH 6/6] x86/virt/tdx: Re-initialize the extensions on runtime TDX module update Xu Yilun
2026-08-22  0:01   ` Edgecombe, Rick P

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aowMYbHKK4b0O5P6@yilunxu-OptiPlex-7050 \
    --to=yilun.xu@linux.intel.com \
    --cc=adrian.hunter@intel.com \
    --cc=artem.bityutskiy@linux.intel.com \
    --cc=baolu.lu@linux.intel.com \
    --cc=chao.gao@intel.com \
    --cc=kas@kernel.org \
    --cc=kishen.maloor@intel.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-coco@lists.linux.dev \
    --cc=linux-kernel@vger.kernel.org \
    --cc=peter.fang@intel.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=sohil.mehta@intel.com \
    --cc=tony.lindgren@linux.intel.com \
    --cc=x86@kernel.org \
    --cc=xiaoyao.li@intel.com \
    --cc=yilun.xu@intel.com \
    --cc=zhenzhong.duan@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).