Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Tom Lendacky <thomas.lendacky@amd.com>
To: Ashish Kalra <Ashish.Kalra@amd.com>,
	tglx@kernel.org, mingo@redhat.com, bp@alien8.de,
	dave.hansen@linux.intel.com, x86@kernel.org, hpa@zytor.com,
	seanjc@google.com, peterz@infradead.org,
	herbert@gondor.apana.org.au, davem@davemloft.net,
	ardb@kernel.org
Cc: pbonzini@redhat.com, aik@amd.com, Michael.Roth@amd.com,
	KPrateek.Nayak@amd.com, Tycho.Andersen@amd.com,
	Nathan.Fontenot@amd.com, ackerleytng@google.com,
	jackyli@google.com, pgonda@google.com, rientjes@google.com,
	jacobhxu@google.com, xin@zytor.com,
	pawan.kumar.gupta@linux.intel.com, babu.moger@amd.com,
	dyoung@redhat.com, nikunj@amd.com, john.allen@amd.com,
	darwi@linutronix.de, linux-kernel@vger.kernel.org,
	linux-crypto@vger.kernel.org, kvm@vger.kernel.org,
	linux-coco@lists.linux.dev
Subject: Re: [PATCH v12 0/5] Add RMPOPT support.
Date: Thu, 27 Aug 2026 11:29:25 -0500	[thread overview]
Message-ID: <c5ee7f05-a85b-46d9-88ff-199961c04bfc@amd.com> (raw)
In-Reply-To: <cover.1786389115.git.ashish.kalra@amd.com>

On 8/10/26 14:20, Ashish Kalra wrote:
> From: Ashish Kalra <ashish.kalra@amd.com>
> 
> In the SEV-SNP architecture, hypervisor and non-SNP guests are subject
> to RMP checks on writes to provide integrity of SEV-SNP guest memory.
> 
> The RMPOPT architecture enables optimizations whereby the RMP checks
> can be skipped if 1GB regions of memory are known to not contain any
> SNP guest memory.
> 
> RMPOPT is a new instruction designed to minimize the performance
> overhead of RMP checks for the hypervisor and non-SNP guests.
> 
> RMPOPT instruction currently supports two functions. In case of the
> verify and report status function the CPU will read the RMP contents,
> verify the entire 1GB region starting at the provided SPA is HV-owned.
> For the entire 1GB region it checks that all RMP entries in this region
> are HV-owned (i.e, not in assigned state) and then accordingly updates
> the RMPOPT table to indicate if optimization has been enabled and
> provide indication to software if the optimization was successful.
> 
> In case of report status function, the CPU returns the optimization
> status for the 1GB region.
> 
> The RMPOPT table is managed by a combination of software and hardware.
> Software uses the RMPOPT instruction to set bits in the table,
> indicating that regions of memory are entirely HV-owned.  Hardware
> automatically clears bits in the RMPOPT table when RMP contents are
> changed during RMPUPDATE instruction.
> 
> For more information on the RMPOPT instruction, see the AMD64 RMPOPT
> technical documentation.
> 
> As SNP is enabled by default the hypervisor and non-SNP guests are
> subject to RMP write checks to provide integrity of SNP guest memory.
> 
> This patch-series adds support to enable RMP optimizations for up to
> 2TB of system RAM across the system and allow RMPUPDATE to disable
> those optimizations as SNP guests are launched.
> 
> Support for RAM larger than 2 TB will be added in follow-on series.
> 
> This series also adds support to disable CPU hotplug while SNP is
> active, as the SEV firmware enumerates CPUs at SNP initialization and is
> not aware of the OS bringing CPUs online or offline afterwards.  This
> also keeps the set of CPUs stable for the asynchronous RMPOPT scan, so
> the per-core RMPOPT_BASE MSRs programmed during setup remain valid.
> 
> This series also introduces support to re-enable RMP optimizations
> during SNP guest termination, after guest pages have been converted
> back to shared.
> 
> RMP optimizations are performed asynchronously by queuing work on a
> dedicated workqueue after a 10 second delay.
> 
> Delaying work allows batching of multiple SNP guest terminations.
> 
> Once 1GB hugetlb guest_memfd support is merged, support for
> re-enabling RMPOPT optimizations during 1GB page cleanup will be added
> in follow-on series.

For the series:

Reviewed-by: Tom Lendacky <thomas.lendacky@amd.com>

> 
> v12:
> - Merged the former "Add interface to re-enable RMP optimizations" and
>   "KVM: SEV: Perform RMP optimizations on SNP guest shutdown" patches into a
>   single patch (now 5/5), since the interface has no user without the KVM
>   caller.  The series is now 5 patches.
> - 2/5 (Disable CPU hotplug): drop a stray reference to later patches from the
>   commit message and trim the snp_prepare() comment.  Move cpu_hotplug_enable()
>   to the end of snp_shutdown(), after clear_rmp() and the mfd_reconfigure() IPI,
>   so no all-CPU operation runs after hotplug is re-enabled.
> - 3/5 (Initialize RMPOPT MSRs): restructure snp_probe_rmptable_info() so the
>   contiguous probe (which clears X86_FEATURE_RMPOPT) is the common fall-through,
>   also covering X86_FEATURE_SEGMENTED_RMP advertised but disabled by firmware.
>   Move snp_setup_rmpopt() to the end of __sev_snp_init_locked().  Give
>   snp_cleanup_rmpopt() its own short comment in snp_shutdown().
> - 4/5 (async RMPOPT): document the optimization trigger points in the commit
>   message.  Rename enum rmpopt_function to rmpopt_op_type.  Merge __rmpopt() into
>   rmpopt() and replace rmpopt_smp() with a new rmpopt_scan_range() that loops as
>   the on_each_cpu() callback, so the follower scan issues one IPI per core rather
>   than one per 1GB.  Use u64 for physical addresses.  Allocate the follower
>   cpumask once at setup.  Document rmpopt_wq as the setup sentinel.  Move the
>   RMPOPT_WORK_TIMEOUT define to the guest-shutdown patch (its only user) as
>   (10 * MSEC_PER_SEC).  Remove a spammy pr_info() and reword comments.
> - 5/5 (Re-enable RMP optimizations on SNP guest shutdown): document why
>   re-optimization is driven by guest teardown (the only event that returns
>   guest memory to hypervisor ownership); reword the commit message
>   (s/clear/disable/ for optimizations, drop "Conversely").
> 
>   Review feedback from Borislav Petkov.
> 
> v11:
> - Reordered so "Disable CPU hotplug while SNP is active" (2/6) precedes
>   "Initialize RMPOPT configuration MSRs" (3/6): the RMPOPT setup/cleanup code is
>   then introduced with CPU hotplug already disabled and never takes
>   cpus_read_lock().
> - 1/6 (cpufeatures): drop the tools/arch/x86/include/asm/cpufeatures.h change and
>   adopt the commit message as applied by Boris.
> - 2/6 (Disable CPU hotplug): reword the commit message.  Drop the redundant
>   cpus_read_lock()/cpus_read_unlock() in snp_prepare() in this same patch -- with
>   hotplug disabled the online CPU mask is stable, so the read lock is not needed
>   and hotplug handling has no hole when bisecting.
> - 3/6 (Initialize RMPOPT MSRs): replace the rmpopt_capable bool with a small
>   helper local to arch/x86/virt/svm/sev.c --
>   cpu_feature_enabled(X86_FEATURE_RMPOPT) && cc_platform_has(CC_ATTR_HOST_SEV_SNP)
>   -- clearing X86_FEATURE_RMPOPT for a contiguous (non-segmented) RMP in
>   snp_probe_rmptable_info() (runs at BSP init, before alternatives).  Rename
>   rmpopt_cleanup() to snp_cleanup_rmpopt() to match snp_setup_rmpopt().  Both are
>   introduced without cpus_read_lock() (hotplug is already disabled by 2/6).
>   Simplify the RMPOPT_BASE comments.  Drop the now-unused <asm/sev.h> include from
>   core.c.
> - 4/6 (async RMPOPT): the follower scan is likewise introduced without
>   cpus_read_lock().  Drop the cond_resched() calls (nops on x86).  On
>   re-initialization after a legacy SNP shutdown, re-queue the optimization pass
>   instead of skipping it.  pr_warn() on cpumask allocation failure.
> 
>   Review feedback from Borislav Petkov and K Prateek Nayak.
> 
> v10:
> - Rework the CPU-hotplug patch (3/6): disable CPU hotplug in
>   snp_prepare(), before SnpEn is set, instead of late in
>   __sev_snp_init_locked(), so no CPU can come online without SnpEn during
>   SNP initialization (per upstream review).  Tie hotplug to SnpEn: it
>   stays disabled while SnpEn is set -- including across a failed SNP_INIT
>   and across the legacy SNP_SHUTDOWN_EX path -- and is re-enabled only
>   once the firmware clears SnpEn on the x86_snp_shutdown path.  Drop the
>   separate idempotent flag: snp_prepare() re-enables hotplug on its own
>   early failure, and a kexec target that boots with SnpEn already set
>   disables hotplug once in snp_rmptable_init().  Reword the commit log and
>   comments accordingly.
> - Emit a pr_warn() in rmpopt_work_handler() (4/6) when the follower
>   cpumask allocation fails, instead of silently skipping the optimization
>   pass.
> 
>   Sashiko AI upstream review identified several of the above issues.
> 
> v9:
> - Rename rmpopt_configured to rmpopt_capable.
> - Make rmpopt_cpumask a cpumask_var_t (allocated/freed at setup/cleanup)
>   instead of a static cpumask_t.
> - Drop the v8 WARN_ON_ONCE() on the RMPOPT_BASE writes; use a plain
>   wrmsrq_on_cpu(), matching the SNP MSR-write convention in this file.
> - Disable CPU hotplug with cpu_hotplug_disable()/cpu_hotplug_enable()
>   (per tglx); re-enable only on the full x86_snp_shutdown path.
> - Simplify rmpopt_work_handler() to a single leader-then-followers path:
>   with CPU hotplug disabled while SNP is active and snp_prepare()
>   requiring all CPUs online when RMPOPT_BASE is programmed, every core is
>   always programmed, so the explicit-leader fallback is now unreachable.
>   Drop it along with the v8 work_on_cpu()/rmpopt_leader_fn() helper.
> - Drop the debugfs interface (was patch 7/7) and its report-only
>   plumbing; observability will be revisited after this series is merged.
> - Restrict snp_rmpopt_all_physmem()'s export to the kvm-amd module.
> - Use scoped_guard(cpus_read_lock) for the per-CPU MSR and follower
>   loops.
> 
>   Sashiko AI upstream review identified several of the above issues.
> 
> v8:
> - Add a new patch to disable CPU hotplug while SNP is active, keeping
>   the CPU set stable for the RMPOPT work handler.
> - Drop the setup_clear_cpu_cap(X86_FEATURE_RMPOPT) calls; the
>   rmpopt_configured bool is the runtime guard.
> - WARN_ON_ONCE() on the RMPOPT_BASE MSR writes that previously ignored
>   their return value.
> - Simplify rmpopt_work_handler() by removing the explicit-leader
>   fallback: with CPU hotplug disabled while SNP is active and
>   snp_prepare() requiring all CPUs online when RMPOPT_BASE is programmed,
>   every core is always programmed, so the running CPU can always be the
>   leader.  This drops the smp_call_function_single() fallback (and with
>   it the AB-BA deadlock and IRQ-latency concerns) and collapses the
>   leader selection into a single leader-then-followers path.
> - Use mod_delayed_work() in snp_rmpopt_all_physmem() so the batching
>   delay tracks the last SNP guest termination.
> 
>   Sashiko AI code review identified several of the above issues.
> 
> v7:
> - Sync tools/arch/x86/include/asm/cpufeatures.h to mirror the kernel
>   header for X86_FEATURE_RMPOPT.
> - Fix commit title to use X86_FEATURE_RMPOPT to match the code
>   (was X86_FEATURE_AMD_RMPOPT).
> - Add static bool rmpopt_configured, set only when segmented RMP setup
>   succeeds in setup_rmptable().  Check rmpopt_configured alongside
>   cpu_feature_enabled(X86_FEATURE_RMPOPT) in snp_setup_rmpopt() and
>   snp_rmpopt_all_physmem(), because setup_clear_cpu_cap() is unreliable
>   after alternatives are patched.  Add snp_clear_rmpopt_configured()
>   called from amd_cc_platform_clear() when CC_ATTR_HOST_SEV_SNP is
>   cleared.  Do not use __ro_after_init on rmpopt_configured since the
>   writer snp_clear_rmpopt_configured() is not __init.
> - Add cond_resched() to all three leader loops in rmpopt_work_handler()
>   to prevent soft lockups on systems with up to 2TB of RAM.
> - Add comment above __rmpopt() documenting the RMPOPT instruction
>   encoding (F2 0F 01 FC) and register interface (RAX = system physical
>   address input, RCX = operation type input, RFLAGS.CF = output).
>   Note: RMPOPT does not modify RAX unlike PVALIDATE/RMPUPDATE, so
>   the existing "a" (input-only) constraint is correct.
> 
>   Sashiko AI code review identified several of the above issues.
> 
> v6:
> - Drop wrmsrq_on_cpus() helper; use for_each_cpu() with wrmsrq_on_cpu()
>   instead, as RMPOPT_BASE MSR programming is not performance-critical.
> - Rewrite rmpopt_work_handler() leader selection to use a local
>   follower_mask copy instead of modifying the global rmpopt_cpumask.
>   This eliminates the current_cpu_cleared tracking and the restore at
>   the end, and removes the need for synchronization comments about
>   transient cpumask inconsistency.
> - Add three-way leader selection in rmpopt_work_handler():
>   1. Current CPU is a primary thread in cpumask: run leader locally.
>   2. Current CPU is a sibling thread whose primary is in cpumask:
>      run leader locally (RMPOPT_BASE MSR is per-core), remove the
>      primary from followers via cpumask_andnot(topology_sibling_cpumask).
>   3. Current CPU's core has no RMPOPT_BASE MSR programmed: pick an
>      explicit leader via cpumask_first() + smp_call_function_single()
>      to avoid #UD, with cpus_read_lock() around the IPI loop.
> - Add WARN_ON_ONCE guard for empty cpumask in the explicit leader
>   fallback path, with migrate_enable() before goto out.
> - Add .llseek = seq_lseek to rmpopt_table_fops for consistency with
>   other seq_file-based debugfs files and to support tools like "less".
> - Change debugfs file permissions from 0444 to 0400 to restrict access
>   to root only.
> - Add comment in rmpopt_table_seq_show() explaining why cpu_online_mask
>   is safe: RMPOPT_BASE MSR is per-core and snp_prepare() ensures all
>   CPUs are online when the MSR is programmed.
> 
>   Sashiko AI code review identified several of the above issues.
> 
> v5:
> - Introduce rmpopt_cleanup() to tear down workqueue, debugfs, cpumask,
>   and MSR state, called from snp_shutdown().
> - Introduce rmpopt_wq_mutex to serialize snp_setup_rmpopt(),
>   snp_rmpopt_all_physmem(), and rmpopt_cleanup().
> - Introduce rmpopt_show_mutex to serialize debugfs reporting of
>   rmpopt_report_cpumask.
> - Move snp_rmpopt_all_physmem() call after SNP DECOMMISSION during
>   guest shutdown.
> - Use migrate_disable()/migrate_enable() for CPU pinning in the
>   rmpopt_work_handler() leader loop to maintain CPU affinity without
>   disabling preemption for the entire RMPOPT scan.
> - Add cpus_read_lock()/cpus_read_unlock() around the follower
>   on_each_cpu_mask() loop in rmpopt_work_handler().
> - Guard snp_setup_rmpopt() against re-initialization when
>   SNP_SHUTDOWN_EX with x86_snp_shutdown=0 skips rmpopt_cleanup()
>   but clears snp_initialized, preventing workqueue and resource
>   leaks on repeated init/shutdown cycles.
> - Replace setup_clear_cpu_cap() with pr_err() on alloc_workqueue()
>   failure in snp_setup_rmpopt(), as setup_clear_cpu_cap() cannot be
>   used after alternatives are patched; callers check rmpopt_wq != NULL
>   as the runtime guard instead.
> - Add pr_info() when RMPOPT coverage is capped at 2TB.
> - Add comments noting CPU hotplug is not supported with SNP enabled
>   and only online primary threads are covered by rmpopt_cpumask.
> - Add comment in setup_rmptable() noting Segmented RMP must be
>   enabled to enable RMPOPT.
> - Simplify cpumask setup loop to set if primary thread rather than
>   skip if not primary.
> - Improve grammar and clarity in snp_setup_rmpopt() comments.
> - Added Reviewed-by's.
> 
>   Sashiko AI code review identified several of the above issues.
> 
> v4:
> - Add new wrmsrq_on_cpus() helper to write same u64 value to a
>   per-CPU MSR across a cpumask without per-cpu struct allocation
>   overhead.
> - Rename configure_and_enable_rmpopt() to snp_setup_rmpopt().
> - Use wrmsrq_on_cpus() instead of wrmsrq_on_cpu() loop for
>   programming RMPOPT_BASE MSRs.
> - Add setup_clear_cpu_cap(X86_FEATURE_RMPOPT) if segmented RMP
>   setup fails or workqueue allocation fails.
> - Add X86_FEATURE_RMPOPT feature clear logic in amd_cc_platform_clear()
>   for CC_ATTR_HOST_SEV_SNP.
> - All of the above allow checking for only X86_FEATURE_RMPOPT for both
>   RMPOPT setup/enable and RMP re-optimizations.
> - Rename snp_perform_rmp_optimization() to snp_rmpopt_all_physmem().
> - Split rmpopt() into rmpopt() and rmpopt_smp() for SMP callback use.
> - Introduce separate rmpopt_report_cpumask for debugfs reporting,
>   distinct from rmpopt_cpumask used for primary thread tracking.
> - Remove snp_perform_rmp_optimization() call from __sev_snp_init_locked()
>   and instead setup and enable RMPOPT after SNP is enabled and
>   initialized.
> 
> v3:
> - Drop all RMPOPT kthread support and introduce adding custom and
>   dedicated workqueue to schedule delayed and asynchronous RMPOPT work.
> - Drop the guest_memfd inode cleanup interface and add support to
>   re-enable RMP optimizations during guest shutdown using the
>   asynchronous and delayed workqueue interface.
> - Introduce new __rmpopt() helper and rmpopt() and
>   rmpopt_report_status() wrappers on top which use rax and rcx
>   parameters to closely match RMPOPT specs.
> - Use new optimized RMPOPT loop to issue RMPOPT instructions on all
>   system RAM upto 2TB and all CPUs, by optimizing each range on one CPU
>   first, then let other CPUs execute RMPOPT in parallel so they can skip
>   most work as the range has already been optimized.
> - Also add support for running the optimized RMPOPT loop only on
>   one thread per core.
> - Replace all PUD_SIZE references with SZ_1G to conform to 1GB regions
>   as specified by RMPOPT specifications and not be dependent on PUD_SIZE
>   which makes the RMPOPT patch-set independent of x86 page table sizes.
> - Use wrmsrq_on_cpu() to program the RMPOPT_BASE MSR registers on
>   all CPUs that removes all ugly casting to use on_each_cpu_mask().
> - Fix inline commits and patch commit messages
> 
> 
> v2:
> - Drop all NUMA and Socket configuration and enablement support and
>   enable RMPOPT support for up to 2TB of system RAM.
> - Drop get_cpumask_of_primary_threads() and enable per-core RMPOPT
>   base MSRs and issue RMPOPT instruction on all CPUs.
> - Drop the configfs interface to manually re-enable RMP optimizations.
> - Add new guest_memfd cleanup interface to automatically re-enable
>   RMP optimizations during guest shutdown.
> - Include references to the public RMPOPT documentation.
> - Move debugfs directory for RMPOPT under architecuture specific
>   parent directory.
> 
> Ashish Kalra (5):
>   x86/cpufeatures: Add X86_FEATURE_RMPOPT feature flag
>   x86/sev: Disable CPU hotplug while SNP is active
>   x86/sev: Initialize RMPOPT configuration MSRs
>   x86/sev: Add support to perform RMP optimizations asynchronously
>   x86/sev: Re-enable RMP optimizations on SNP guest shutdown
> 
>  arch/x86/include/asm/cpufeatures.h |   2 +-
>  arch/x86/include/asm/msr-index.h   |   3 +
>  arch/x86/include/asm/sev.h         |   4 +
>  arch/x86/kernel/cpu/scattered.c    |   1 +
>  arch/x86/kvm/svm/sev.c             |  10 +
>  arch/x86/virt/svm/sev.c            | 289 +++++++++++++++++++++++++++--
>  drivers/crypto/ccp/sev-dev.c       |   2 +
>  7 files changed, 296 insertions(+), 15 deletions(-)
> 


      parent reply	other threads:[~2026-08-27 16:29 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-10 19:20 [PATCH v12 0/5] Add RMPOPT support Ashish Kalra
2026-08-10 19:21 ` [PATCH v12 1/5] x86/cpufeatures: Add X86_FEATURE_RMPOPT feature flag Ashish Kalra
2026-08-10 19:21 ` [PATCH v12 2/5] x86/sev: Disable CPU hotplug while SNP is active Ashish Kalra
2026-08-10 19:21 ` [PATCH v12 3/5] x86/sev: Initialize RMPOPT configuration MSRs Ashish Kalra
2026-08-27  2:06   ` Borislav Petkov
2026-08-10 19:22 ` [PATCH v12 4/5] x86/sev: Add support to perform RMP optimizations asynchronously Ashish Kalra
2026-08-29 20:52   ` Borislav Petkov
2026-08-31 20:00     ` Kalra, Ashish
2026-09-01 19:20       ` Borislav Petkov
2026-08-10 19:22 ` [PATCH v12 5/5] x86/sev: Re-enable RMP optimizations on SNP guest shutdown Ashish Kalra
2026-08-27 16:29 ` Tom Lendacky [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=c5ee7f05-a85b-46d9-88ff-199961c04bfc@amd.com \
    --to=thomas.lendacky@amd.com \
    --cc=Ashish.Kalra@amd.com \
    --cc=KPrateek.Nayak@amd.com \
    --cc=Michael.Roth@amd.com \
    --cc=Nathan.Fontenot@amd.com \
    --cc=Tycho.Andersen@amd.com \
    --cc=ackerleytng@google.com \
    --cc=aik@amd.com \
    --cc=ardb@kernel.org \
    --cc=babu.moger@amd.com \
    --cc=bp@alien8.de \
    --cc=darwi@linutronix.de \
    --cc=dave.hansen@linux.intel.com \
    --cc=davem@davemloft.net \
    --cc=dyoung@redhat.com \
    --cc=herbert@gondor.apana.org.au \
    --cc=hpa@zytor.com \
    --cc=jackyli@google.com \
    --cc=jacobhxu@google.com \
    --cc=john.allen@amd.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-coco@lists.linux.dev \
    --cc=linux-crypto@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mingo@redhat.com \
    --cc=nikunj@amd.com \
    --cc=pawan.kumar.gupta@linux.intel.com \
    --cc=pbonzini@redhat.com \
    --cc=peterz@infradead.org \
    --cc=pgonda@google.com \
    --cc=rientjes@google.com \
    --cc=seanjc@google.com \
    --cc=tglx@kernel.org \
    --cc=x86@kernel.org \
    --cc=xin@zytor.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox