From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.11]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ED80C305680; Thu, 3 Sep 2026 01:51:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.11 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788400294; cv=none; b=NcKEDyWtsHvPS3llVKjEJAwQTR6K5WCoXKJJltC0BJI1hHoXTTJ1FzUllsYdu0bzlagRwKpyrUI6BeMLUtaKpTOvQ2P8T5d/GaQPLIqd1J18d6mokIm1O3W38DnJeEjojDLcouwbrhebyhQ0FGSBjXC+Pax3yUewe5Osxdi8gJg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1788400294; c=relaxed/simple; bh=La5TSnve7LEhFz6cfbRU27H/X/gt9pcVEVuJ+3YN1X4=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=j6yEcJFizkondx7mmbvHHHmidJ6MAdjbJsh1s3pkMEUYH1ffiyH4ZpFkm/7ehZabs/1iaVQcufdlLJs0u1H7uOd//qzne4nYj55CESNCnpDYuVWfSswqTyscG5KzCvtWwenHAdsMUm+dksODVA+Jb6A5lXHrGgp1u/qb7xR2Noo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=F4Gl76hD; arc=none smtp.client-ip=192.198.163.11 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="F4Gl76hD" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1788400290; x=1819936290; h=from:to:cc:subject:date:message-id:in-reply-to: references:mime-version:content-transfer-encoding; bh=La5TSnve7LEhFz6cfbRU27H/X/gt9pcVEVuJ+3YN1X4=; b=F4Gl76hDYjRXSRs3wkg5jkrcvTOXpluChd/xrRUyLXx4eRh5+N+4h8ZM ICpxmWWnDc8kGUFxNIvS4wGcw74Cfd8V0akjCnuE+t2abNfmj0LK3MnTo 0j3EeBMVqxSHSX5MLG35Go9GwhjcejRriYit1z3VfzS6hsgFnOA3QmTLB 1C3JheTpX0xvzEIuE7AfvcK4JydLCeiXoNQHtS8+2IIBvyBppuHpHkhr0 2UHUFtqUC+VDeLkVmTuDPuek5B5ktm7ENwGBKfOX1TIQtDozUqK3tM+S+ g01e7J1Wp6mQ4CpUv/GDqIZQxL7VKA8GNpCxCcW9YZ7hefWhRolYTrK7W g==; X-CSE-ConnectionGUID: O7Gy742iRB2vEMjluZ0JbA== X-CSE-MsgGUID: hGam1aYpQ+evBiqJsn0RPQ== X-IronPort-AV: E=McAfee;i="6800,10657,11894"; a="99469165" X-IronPort-AV: E=Sophos;i="6.25,258,1779174000"; d="scan'208";a="99469165" Received: from orviesa001.jf.intel.com ([10.64.159.141]) by fmvoesa105.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 02 Sep 2026 18:51:23 -0700 X-CSE-ConnectionGUID: 4Ku2yq3FRcKvvw9mp6XvMQ== X-CSE-MsgGUID: n0XP+eJ0RU6qnuYqoQr5vQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,258,1779174000"; d="scan'208";a="307770229" Received: from rpedgeco-desk.jf.intel.com ([10.88.27.135]) by smtpauth.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 02 Sep 2026 18:51:22 -0700 From: Rick Edgecombe To: bp@alien8.de, dave.hansen@intel.com, hpa@zytor.com, kas@kernel.org, kvm@vger.kernel.org, linux-coco@lists.linux.dev, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, mingo@redhat.com, nik.borisov@suse.com, pbonzini@redhat.com, seanjc@google.com, tglx@kernel.org, vannapurve@google.com, x86@kernel.org, chao.gao@intel.com, yan.y.zhao@intel.com, kai.huang@intel.com, tony.lindgren@linux.intel.com, binbin.wu@intel.com, sohil.mehta@intel.com Cc: rick.p.edgecombe@intel.com, Hongyu Ning , Binbin Wu , Dave Hansen Subject: [PATCH v10 06/11] KVM: TDX: Allocate PAMT memory for TD and vCPU control structures Date: Wed, 2 Sep 2026 18:51:08 -0700 Message-ID: <20260903015113.93343-7-rick.p.edgecombe@intel.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260903015113.93343-1-rick.p.edgecombe@intel.com> References: <20260903015113.93343-1-rick.p.edgecombe@intel.com> Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit From: "Kirill A. Shutemov" Use control page helpers for allocating and freeing TD control structures, such that these operations can work for Dynamic PAMT. The TDX module tracks some state for each page of physical memory that it might use. It calls this state the PAMT. It includes separate state for each page size a physical page could be utilized at within the TDX module (1GB, 2MB, 4KB). In Dynamic PAMT, only the 4KB page size state is allocated dynamically. So the kernel must ensure PAMT backing is installed for any 4KB page being gifted to the TDX module, and must tear down the backing when all associated gifted pages are reclaimed. TD scoped control pages (TDR, TDCS) and vCPU scoped control pages (TDVPR, TDCX) are all handed to the TDX module at 4KB page size and are therefore subject to this requirement. Replace the raw alloc_page()/__free_page() calls for these pages with tdx_alloc/free_control_page(). Switching between special Dynamic PAMT operations or normal page alloc/free operations is handled internally in tdx_alloc/free_control_page(). So don't check for Dynamic PAMT around these calls. Just call them unconditionally. Similarly, drop the NULL checks before freeing, as tdx_free_control_page() handles NULL internally. No functional change intended when DPAMT is not in use. Signed-off-by: Kirill A. Shutemov [sean: handle alloc+free+reclaim in one patch] Signed-off-by: Sean Christopherson [rick: enhance log, reviewing, rebase, with help from AI tooling] Signed-off-by: Rick Edgecombe Tested-by: Hongyu Ning Reviewed-by: Binbin Wu Reviewed-by: Chao Gao Reviewed-by: Yan Zhao Reviewed-by: Tony Lindgren Reviewed-by: Nikolay Borisov Reviewed-by: Dave Hansen Reviewed-by: Vishal Annapurve Acked-by: Sohil Mehta Acked-by: Sean Christopherson --- v7: - Fixup tags (Sean) - Missing word in log (Binbin) - Log smoothness (Yan) --- arch/x86/kvm/vmx/tdx.c | 35 ++++++++++++++--------------------- 1 file changed, 14 insertions(+), 21 deletions(-) diff --git a/arch/x86/kvm/vmx/tdx.c b/arch/x86/kvm/vmx/tdx.c index b272c20586a74..3592596b5649a 100644 --- a/arch/x86/kvm/vmx/tdx.c +++ b/arch/x86/kvm/vmx/tdx.c @@ -362,7 +362,7 @@ static void tdx_reclaim_control_page(struct page *ctrl_page) if (tdx_reclaim_page(ctrl_page)) return; - __free_page(ctrl_page); + tdx_free_control_page(ctrl_page); } struct tdx_flush_vp_arg { @@ -589,7 +589,7 @@ static void tdx_reclaim_td_control_pages(struct kvm *kvm) tdx_quirk_reset_paddr(page_to_phys(kvm_tdx->td.tdr_page), PAGE_SIZE); - __free_page(kvm_tdx->td.tdr_page); + tdx_free_control_page(kvm_tdx->td.tdr_page); kvm_tdx->td.tdr_page = NULL; } @@ -2456,7 +2456,7 @@ static int __tdx_td_init(struct kvm *kvm, struct td_params *td_params, ret = -ENOMEM; - tdr_page = alloc_page(GFP_KERNEL_ACCOUNT); + tdr_page = tdx_alloc_control_page(); if (!tdr_page) goto free_hkid; @@ -2469,7 +2469,7 @@ static int __tdx_td_init(struct kvm *kvm, struct td_params *td_params, goto free_tdr; for (i = 0; i < kvm_tdx->td.tdcs_nr_pages; i++) { - tdcs_pages[i] = alloc_page(GFP_KERNEL_ACCOUNT); + tdcs_pages[i] = tdx_alloc_control_page(); if (!tdcs_pages[i]) goto free_tdcs; } @@ -2587,10 +2587,8 @@ static int __tdx_td_init(struct kvm *kvm, struct td_params *td_params, teardown: /* Only free pages not yet added, so start at 'i' */ for (; i < kvm_tdx->td.tdcs_nr_pages; i++) { - if (tdcs_pages[i]) { - __free_page(tdcs_pages[i]); - tdcs_pages[i] = NULL; - } + tdx_free_control_page(tdcs_pages[i]); + tdcs_pages[i] = NULL; } if (!kvm_tdx->td.tdcs_pages) kfree(tdcs_pages); @@ -2605,16 +2603,13 @@ static int __tdx_td_init(struct kvm *kvm, struct td_params *td_params, free_cpumask_var(packages); free_tdcs: - for (i = 0; i < kvm_tdx->td.tdcs_nr_pages; i++) { - if (tdcs_pages[i]) - __free_page(tdcs_pages[i]); - } + for (i = 0; i < kvm_tdx->td.tdcs_nr_pages; i++) + tdx_free_control_page(tdcs_pages[i]); kfree(tdcs_pages); kvm_tdx->td.tdcs_pages = NULL; free_tdr: - if (tdr_page) - __free_page(tdr_page); + tdx_free_control_page(tdr_page); kvm_tdx->td.tdr_page = NULL; free_hkid: @@ -2948,7 +2943,7 @@ static int tdx_td_vcpu_init(struct kvm_vcpu *vcpu, u64 vcpu_rcx) int ret, i; u64 err; - page = alloc_page(GFP_KERNEL_ACCOUNT); + page = tdx_alloc_control_page(); if (!page) return -ENOMEM; tdx->vp.tdvpr_page = page; @@ -2968,7 +2963,7 @@ static int tdx_td_vcpu_init(struct kvm_vcpu *vcpu, u64 vcpu_rcx) } for (i = 0; i < kvm_tdx->td.tdcx_nr_pages; i++) { - page = alloc_page(GFP_KERNEL_ACCOUNT); + page = tdx_alloc_control_page(); if (!page) { ret = -ENOMEM; goto free_tdcx; @@ -2990,7 +2985,7 @@ static int tdx_td_vcpu_init(struct kvm_vcpu *vcpu, u64 vcpu_rcx) * method, but the rest are freed here. */ for (; i < kvm_tdx->td.tdcx_nr_pages; i++) { - __free_page(tdx->vp.tdcx_pages[i]); + tdx_free_control_page(tdx->vp.tdcx_pages[i]); tdx->vp.tdcx_pages[i] = NULL; } return -EIO; @@ -3018,16 +3013,14 @@ static int tdx_td_vcpu_init(struct kvm_vcpu *vcpu, u64 vcpu_rcx) free_tdcx: for (i = 0; i < kvm_tdx->td.tdcx_nr_pages; i++) { - if (tdx->vp.tdcx_pages[i]) - __free_page(tdx->vp.tdcx_pages[i]); + tdx_free_control_page(tdx->vp.tdcx_pages[i]); tdx->vp.tdcx_pages[i] = NULL; } kfree(tdx->vp.tdcx_pages); tdx->vp.tdcx_pages = NULL; free_tdvpr: - if (tdx->vp.tdvpr_page) - __free_page(tdx->vp.tdvpr_page); + tdx_free_control_page(tdx->vp.tdvpr_page); tdx->vp.tdvpr_page = NULL; tdx->vp.tdvpr_pa = 0; -- 2.55.0