From: Zack Rusin <zack.rusin@broadcom.com>
To: Borislav Petkov <bp@alien8.de>,
Ajay Kaher <ajay.kaher@broadcom.com>,
Alexey Makhalov <alexey.makhalov@broadcom.com>,
x86@kernel.org, Dennis Zhou <dennis@kernel.org>,
Tejun Heo <tj@kernel.org>, Arnd Bergmann <arnd@arndb.de>,
Kiryl Shutsemau <kas@kernel.org>,
Rick Edgecombe <rick.p.edgecombe@intel.com>
Cc: Thomas Gleixner <tglx@kernel.org>, Ingo Molnar <mingo@redhat.com>,
Dave Hansen <dave.hansen@linux.intel.com>,
"H . Peter Anvin" <hpa@zytor.com>,
virtualization@lists.linux.dev,
bcm-kernel-feedback-list@broadcom.com,
linux-kernel@vger.kernel.org, Christoph Lameter <cl@gentwo.org>,
Andrew Morton <akpm@linux-foundation.org>,
Tom Lendacky <thomas.lendacky@amd.com>,
Bo Gan <bo.gan@broadcom.com>,
linux-mm@kvack.org, linux-arch@vger.kernel.org,
linux-coco@lists.linux.dev, kvm@vger.kernel.org,
Zack Rusin <zack.rusin@broadcom.com>
Subject: [PATCH v1 2/2] x86/vmware: Decrypt steal-time storage before sharing it
Date: Wed, 16 Sep 2026 01:05:40 -0400 [thread overview]
Message-ID: <bbdb89067cd07d07fe4963ad229932faf87b0ef9.1789488039.git.zack.rusin@broadcom.com> (raw)
In-Reply-To: <cover.1789488039.git.zack.rusin@broadcom.com>
VMware's steal-time counter must be in shared memory so the host can
update it. Encrypted guests currently register its address without
decrypting the storage. Move their setup to an early initcall, when the
memory allocator is available for page-table splitting, but before
secondary CPUs start. Leave ordinary guests' setup unchanged.
Convert every possible CPU's storage before registering any address.
Zero each object after conversion, since its contents may not survive
conversion and the host initializes only the counter. Register the boot
CPU with preemption disabled; the existing hotplug callbacks handle the
others. If conversion or boot-CPU registration fails, attempt to re-encrypt
all affected ranges, including a partially converted failing range, and
disable steal time.
TDX's conversion path requires directly mapped memory. Check all per-CPU
objects before converting anything, and disable steal time if a TDX guest
uses the page per-CPU allocator, which supplies vmalloc mappings. This
also covers automatic fallback from the embedded allocator. AMD guests
are unaffected by this restriction because their conversion path supports
those mappings.
Link: https://lore.kernel.org/r/20260309235250.2611115-5-alexey.makhalov@broadcom.com
Co-developed-by: Bo Gan <bo.gan@broadcom.com>
Signed-off-by: Bo Gan <bo.gan@broadcom.com>
Co-developed-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
Signed-off-by: Alexey Makhalov <alexey.makhalov@broadcom.com>
Signed-off-by: Zack Rusin <zack.rusin@broadcom.com>
---
arch/x86/kernel/cpu/vmware.c | 96 +++++++++++++++++++++++++++++++++++-
1 file changed, 95 insertions(+), 1 deletion(-)
diff --git a/arch/x86/kernel/cpu/vmware.c b/arch/x86/kernel/cpu/vmware.c
index 34b73573b108..f49898275d87 100644
--- a/arch/x86/kernel/cpu/vmware.c
+++ b/arch/x86/kernel/cpu/vmware.c
@@ -25,14 +25,19 @@
#include <linux/init.h>
#include <linux/export.h>
#include <linux/clocksource.h>
+#include <linux/cc_platform.h>
#include <linux/cpu.h>
#include <linux/efi.h>
+#include <linux/mm.h>
+#include <linux/preempt.h>
#include <linux/reboot.h>
+#include <linux/set_memory.h>
#include <linux/static_call.h>
#include <linux/sched/cputime.h>
#include <asm/div64.h>
#include <asm/x86_init.h>
#include <asm/hypervisor.h>
+#include <asm/cpufeature.h>
#include <asm/cpuid/api.h>
#include <asm/timer.h>
#include <asm/apic.h>
@@ -147,6 +152,7 @@ static struct cyc2ns_data vmware_cyc2ns __ro_after_init;
static bool vmw_sched_clock __initdata = true;
static DEFINE_PER_CPU_DECRYPTED(struct vmware_steal_time, vmw_steal_time) __aligned(64);
static bool has_steal_clock;
+static bool vmw_steal_time_ready;
static bool steal_acc __initdata = true; /* steal time accounting */
static __init int setup_vmw_sched_clock(char *s)
@@ -280,7 +286,9 @@ static void vmware_disable_steal_time(void)
static void vmware_guest_cpu_init(void)
{
- if (has_steal_clock)
+ if (has_steal_clock &&
+ (!cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT) ||
+ vmw_steal_time_ready))
vmware_register_steal_time();
}
@@ -325,6 +333,92 @@ static int vmware_cpu_down_prepare(unsigned int cpu)
}
#endif
+static void __init vmware_steal_time_range(int cpu, unsigned long *start,
+ int *numpages)
+{
+ unsigned long addr = (unsigned long)per_cpu_ptr(&vmw_steal_time, cpu);
+
+ *start = addr & PAGE_MASK;
+ *numpages = DIV_ROUND_UP(offset_in_page(addr) +
+ sizeof(struct vmware_steal_time), PAGE_SIZE);
+}
+
+static int __init vmware_decrypt_steal_time(void)
+{
+ int cpu, failed_cpu, numpages, ret;
+ unsigned long start;
+
+ if (!has_steal_clock ||
+ !cc_platform_has(CC_ATTR_GUEST_MEM_ENCRYPT))
+ return 0;
+
+ /*
+ * TDX's conversion callback derives the physical range with __pa(),
+ * so the storage must live in the direct map. The page per-CPU
+ * allocator, which is also the automatic fallback when the embedding
+ * allocator fails, hands out vmalloc addresses instead. Reject those
+ * before converting anything, so no page is left converted.
+ */
+ if (cpu_feature_enabled(X86_FEATURE_TDX_GUEST)) {
+ for_each_possible_cpu(cpu) {
+ if (!is_vmalloc_addr(per_cpu_ptr(&vmw_steal_time, cpu)))
+ continue;
+
+ pr_warn("steal time disabled: TDX requires a direct mapping\n");
+ has_steal_clock = false;
+ return 0;
+ }
+ }
+
+ for_each_possible_cpu(cpu) {
+ vmware_steal_time_range(cpu, &start, &numpages);
+ ret = set_memory_decrypted(start, numpages);
+ if (ret) {
+ failed_cpu = cpu;
+ goto rollback;
+ }
+
+ /*
+ * Conversion need not preserve the zeroes. The host
+ * initializes only the counter on enable, so the guest
+ * must initialize the reserved words itself.
+ */
+ memset(per_cpu_ptr(&vmw_steal_time, cpu), 0,
+ sizeof(struct vmware_steal_time));
+ }
+
+ vmw_steal_time_ready = true;
+ preempt_disable();
+ vmware_guest_cpu_init();
+ preempt_enable();
+ if (!has_steal_clock) {
+ pr_warn("failed to register boot CPU steal-time memory\n");
+ failed_cpu = nr_cpu_ids;
+ goto rollback_pages;
+ }
+ return 0;
+
+rollback:
+ pr_warn("failed to decrypt steal-time memory for CPU %d: %d\n",
+ failed_cpu, ret);
+
+rollback_pages:
+ for_each_possible_cpu(cpu) {
+ vmware_steal_time_range(cpu, &start, &numpages);
+ ret = set_memory_encrypted(start, numpages);
+ if (ret)
+ pr_warn("failed to re-encrypt steal-time memory for CPU %d: %d\n",
+ cpu, ret);
+ if (cpu == failed_cpu)
+ break;
+ }
+
+ vmw_steal_time_ready = false;
+ has_steal_clock = false;
+ return 0;
+}
+early_initcall(vmware_decrypt_steal_time);
+
static __init int activate_jump_labels(void)
{
if (has_steal_clock) {
--
2.53.0
next prev parent reply other threads:[~2026-09-16 5:06 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-16 5:05 [PATCH v1 0/2] x86/vmware: Share steal-time storage in encrypted guests Zack Rusin
2026-09-16 5:05 ` [PATCH v1 1/2] percpu: Use X86_MEM_ENCRYPT for decrypted per-CPU data Zack Rusin
2026-09-16 11:14 ` Kiryl Shutsemau
2026-09-16 5:05 ` Zack Rusin [this message]
2026-09-16 12:18 ` [PATCH v1 0/2] x86/vmware: Share steal-time storage in encrypted guests Kiryl Shutsemau
2026-09-16 15:41 ` Zack Rusin
2026-09-17 14:11 ` Kiryl Shutsemau
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=bbdb89067cd07d07fe4963ad229932faf87b0ef9.1789488039.git.zack.rusin@broadcom.com \
--to=zack.rusin@broadcom.com \
--cc=ajay.kaher@broadcom.com \
--cc=akpm@linux-foundation.org \
--cc=alexey.makhalov@broadcom.com \
--cc=arnd@arndb.de \
--cc=bcm-kernel-feedback-list@broadcom.com \
--cc=bo.gan@broadcom.com \
--cc=bp@alien8.de \
--cc=cl@gentwo.org \
--cc=dave.hansen@linux.intel.com \
--cc=dennis@kernel.org \
--cc=hpa@zytor.com \
--cc=kas@kernel.org \
--cc=kvm@vger.kernel.org \
--cc=linux-arch@vger.kernel.org \
--cc=linux-coco@lists.linux.dev \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mingo@redhat.com \
--cc=rick.p.edgecombe@intel.com \
--cc=tglx@kernel.org \
--cc=thomas.lendacky@amd.com \
--cc=tj@kernel.org \
--cc=virtualization@lists.linux.dev \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox