Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Christoph Schlameuss" <schlameuss@linux.ibm.com>
Cc: "Heiko Carstens" <hca@linux.ibm.com>,
	"Alexander Gordeev" <agordeev@linux.ibm.com>,
	"Christian Borntraeger" <borntraeger@linux.ibm.vom>,
	kvm@vger.kernel.org, "Vasily Gorbik" <gor@linux.ibm.com>,
	linux-s390@vger.kernel.org
Subject: Re: [PATCH v4 14/20] KVM: s390: vsie: Shadow VSIE SCA in guest-1
Date: Wed, 12 Aug 2026 16:10:01 +0000	[thread overview]
Message-ID: <20260812161002.332CB1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260812-vsie-sigpi-v4-14-f19e49666315@linux.ibm.com>

Thank you for your contribution! Sashiko AI review found 10 potential issue(s) to consider:
- [Critical] Reusing a `vsie_sca` unconditionally zeroes and unpins `vsie_page` structures that may be actively executing in hardware via a concurrent non-SCA VSIE run.
- [Critical] Reusing a `vsie_page` for a different SCB without resetting the shadow control block leads to cross-VM state leakage.
- [Critical] Unvalidated guest-controlled `icpua` field leads to out-of-bounds array indexing, enabling massive out-of-bounds host memory accesses.
- [High] VSIE error paths improperly convert guest-level validity intercepts and addressing exceptions into fatal host-level negative errors, allowing a nested guest to abort the host VM.
- [High] The shadow SCA is never linked to the shadow SIE control block, completely breaking the VSIE-Interpretation-Extension feature.
- [High] A Time-Of-Check to Time-Of-Use (TOCTOU) vulnerability exists when iterating over the guest-modifiable `mcn` mask, allowing malicious guests to evade error cleanup and leave dangling host pointers in the shadow SCA.
- [High] An unbounded memory leak of XArray nodes occurs because stale mappings are never erased from `addr_to_page`.
- [High] Failing to pin a reused `vsie_sca` leaves a dangling `sca_gpa` that causes erroneous XArray erasure upon the next reuse attempt, corrupting the SCA lookup cache.
- [Medium] Overwriting `vsie_page_old->scb_gpa` with `ULONG_MAX` causes KVM to pass a bogus GPA to `mark_page_dirty()`, silently losing dirty tracking for the guest SCB.
- [Medium] `shadow_sca()` fails to detect guest modifications to the SCA CPU mask (`mcn`) after the first shadow, breaking dynamic CPU hotplug in the nested guest.
--

--- Patch [14]: [PATCH v4 14/20] KVM: s390: vsie: Shadow VSIE SCA in guest-1 ---
commit 4635296fb400229acd33f3f98111bf5ac95e07d2
Author: Christoph Schlameuss <schlameuss@linux.ibm.com>

KVM: s390: vsie: Shadow VSIE SCA in guest-1

Restructure kvm_s390_handle_vsie() to create a guest-1 shadow of the SCA
if guest-2 attempts to enter SIE with an SCA. When active, shadow SCAs and
CPUs are created to allow direct machine monitoring of intervention bits,
reducing VM exits.

> diff --git a/arch/s390/kvm/vsie.c b/arch/s390/kvm/vsie.c
> index e632311222002..2fd349e52bd6d 100644
> --- a/arch/s390/kvm/vsie.c
> +++ b/arch/s390/kvm/vsie.c

[ ... ]

> +static struct vsie_sca *get_vsie_sca(struct kvm_vcpu *vcpu, struct kvm_s390_sie_block *scb_o)
> +{
> +	struct vsie_sca *vsie_sca, *vsie_sca_new = NULL;

[ ... ]

> +	if (vsie_sca_new && kvm->arch.vsie.sca_count < max_vsie_sca) {
> +		/* make use of vsie_sca just created */
> +		vsie_sca = vsie_sca_new;
> +		vsie_sca_new = NULL;
> +
> +		kvm->arch.vsie.scas[kvm->arch.vsie.sca_count] = vsie_sca;
> +		kvm->arch.vsie.sca_count++;
> +		atomic_set(&vsie_sca->ref_count, 1);
> +	} else {
> +		/* reuse previously created vsie_sca allocation for different osca */
> +		vsie_sca = get_reuseable_vsie_sca(kvm);
> +		/* with nr_vcpus scas one must be reusable */
> +		if (IS_ERR(vsie_sca))
> +			goto out;
> +		WARN_ON_ONCE(atomic_read(&vsie_sca->ref_count) != 1);
> +
> +		xa_erase(&kvm->arch.vsie.osca_to_sca, vsie_sca->sca_gpa >> SCA_ALIGNMENT_SHIFT);
> +		for (cpu_nr = 0; cpu_nr < KVM_S390_MAX_VSIE_VCPUS; cpu_nr++) {
> +			vsie_page_n = vsie_sca->pages[cpu_nr];
> +			if (!vsie_page_n)
> +				continue;
> +
> +			/* unpin but keep the vsie_page for reuse */
> +			unpin_scb(kvm, vsie_page_n);
> +			release_gmap_shadow_safe(kvm, vsie_page_n);
> +			memset(vsie_page_n, 0, sizeof(struct vsie_page));
> +			vsie_page_n->scb_gpa = ULONG_MAX;

[Severity: Critical]
Does this unconditionally zero and unpin a vsie_page that might be actively 
executing in hardware via a concurrent non-SCA VSIE run?

If a guest uses the same SCB address in two different virtual CPUs, one vCPU 
might enter VSIE without an SCA and claim the vsie_page without incrementing 
vsie_sca->ref_count. Could this reuse path then destroy the memory while the 
first vCPU is actively running it?

> +		}
> +		unpin_sca(kvm, vsie_sca);
> +		clear_vsie_sca(vsie_sca);
> +	}
> +
> +	if (sie_uses_esca(scb_o))
> +		__set_bit(VSIE_SCA_ESCA, &vsie_sca->flags);
> +	vsie_sca->sca_gpa = sca_gpa;
> +
> +	/*
> +	 * The pinned original sca will only be unpinned lazily to limit the
> +	 * required amount of pins/unpins on each vsie entry/exit.
> +	 * The unpin is done in the reuse vsie_sca allocation path above and
> +	 * kvm_s390_vsie_destroy().
> +	 */
> +	rc = pin_sca(kvm, vsie_sca);
> +	if (rc) {
> +		put_vsie_sca(vsie_sca);
> +		vsie_sca = ERR_PTR(rc);
> +		goto out;

[Severity: High]
Does failing to pin a reused vsie_sca leave a dangling sca_gpa that causes an 
erroneous XArray erasure on the next reuse? 

The sca_gpa retains the failing address when it is released to the pool. When 
get_reuseable_vsie_sca() picks it up again, it appears it will call xa_erase() 
on the old sca_gpa, potentially deleting a valid entry stored by another 
thread.

> +	}
> +
> +	WARN_ON_ONCE(xa_store(&kvm->arch.vsie.osca_to_sca,
> +			      vsie_sca->sca_gpa >> SCA_ALIGNMENT_SHIFT, vsie_sca, GFP_KERNEL));
> +
> +out:
> +	up_write(&kvm->arch.vsie.vsie_sca_lock);
> +	if (vsie_sca_new)
> +		free_vsie_sca(kvm, vsie_sca_new);
> +	return vsie_sca;
> +}

[ ... ]

>  static int vsie_page_init(struct kvm_vcpu *vcpu, struct vsie_page *vsie_page, unsigned long scb_gpa)
>  {
> +	struct vsie_page *vsie_page_old;
>  	struct kvm *kvm = vcpu->kvm;
>  	int rc;
>  
> -	if (vsie_page->scb_gpa != ULONG_MAX)
> -		xa_erase(&kvm->arch.vsie.addr_to_page, vsie_page->scb_gpa >> SCB_ALIGNMENT_SHIFT);

[Severity: High]
Are old scb_gpa mappings permanently leaking in the XArray here? 

By removing the xa_erase() call, a guest executing VSIE in a tight loop with 
varying scb_addrs might force KVM to recycle vsie_pages while continuously 
accumulating stale entries in addr_to_page, eventually exhausting host memory.

> -
> +	vsie_page->scb_gpa = scb_gpa;
> +	rc = pin_scb(vcpu, vsie_page);
> +	if (rc) {
> +		vsie_page->scb_gpa = ULONG_MAX;
> +		return -ENOMEM;
> +	}

[Severity: High]
Can this return a fatal host-level error for a guest-level addressing 
exception? 

If a nested guest enters VSIE with an invalid scb_o address, pin_scb() returns 
1. By converting this and returning -ENOMEM here, does this allow a nested 
guest to abort the host VM instead of correctly injecting a program interrupt 
into the guest?

>  
>  	vsie_page->sca_gpa = read_scao(kvm, vsie_page->scb_o);
> -	WARN_ON_ONCE(xa_insert(&kvm->arch.vsie.addr_to_page, scb_gpa >> SCB_ALIGNMENT_SHIFT,
> -			       vsie_page, GFP_KERNEL_ACCOUNT));
> +
> +	/*
> +	 * store the vsie_page in addr_to_page
> +	 * mind that g2 may have reused the sca - make sure we do not remove the sca from
> +	 * the new config when reusing the vsie_page_old
> +	 */
> +	vsie_page_old = xa_store(&kvm->arch.vsie.addr_to_page, scb_gpa >> SCB_ALIGNMENT_SHIFT,
> +				 vsie_page, GFP_KERNEL_ACCOUNT);
> +	if (WARN_ON_ONCE(xa_err(vsie_page_old)))
> +		return 0;
> +	if (vsie_page_old && vsie_page_old != vsie_page)
> +		WRITE_ONCE(vsie_page_old->scb_gpa, ULONG_MAX);

[Severity: Medium]
Does overwriting the old vsie_page's scb_gpa with ULONG_MAX silently lose 
dirty tracking? 

When this evicted page is later unpinned via unpin_scb(), it will call 
mark_page_dirty() using ULONG_MAX instead of the actual guest physical 
address of the SCB, losing the migration state.

>  
>  	return 0;
>  }

[ ... ]

> +static struct vsie_page *get_vsie_page_cpu_nr(struct kvm_vcpu *vcpu, struct vsie_sca *vsie_sca,
> +					      gpa_t scb_gpa, u16 cpu_nr)
> +{
> +	struct vsie_page *vsie_page, *vsie_page_new = NULL;
> +	int rc;
> +
> +	vsie_page = vsie_sca->pages[cpu_nr];

[Severity: Critical]
Does this use an unvalidated guest-controlled icpua field directly as an array 
index? 

Since cpu_nr comes from scb_o->icpua without prior bounds checking against 
KVM_S390_MAX_VSIE_VCPUS in kvm_s390_handle_vsie(), a malicious guest could 
provide an out-of-bounds value (up to 65535), leading to massive out-of-bounds 
read/write on the kernel heap.

> +	if (!vsie_page) {
> +		vsie_page_new = alloc_vsie_page(vcpu->kvm);

[ ... ]

> +	if (vsie_page->scb_gpa != scb_gpa || vsie_page->sca_gpa != vsie_sca->sca_gpa) {
> +		scoped_guard(mutex, &vcpu->kvm->arch.vsie.mutex) {
> +			unpin_scb(vcpu->kvm, vsie_page);
> +			rc = vsie_page_init(vcpu, vsie_page, scb_gpa);
> +		}
> +		if (WARN_ON_ONCE(rc)) {
> +			put_vsie_page(vsie_page);
> +			return ERR_PTR(rc);
> +		}
> +	}

[Severity: Critical]
When reusing a vsie_page for a different SCB, does this lead to cross-VM state 
leakage? 

By updating the pinning via vsie_page_init() but omitting critical state reset 
steps like memset(&vsie_page->scb_s, 0, ...), release_gmap_shadow_safe(), and 
prefix_unmapped(), this might leave old shadow control block state intact and 
stale GMAP translations active from a different nested VM.

> +
> +	return vsie_page;
> +}

[ ... ]

> +static int _shadow_sca(struct kvm_vcpu *vcpu, struct vsie_page *vsie_page,
> +		       struct vsie_sca *vsie_sca)
> +{
> +	bool is_esca = sie_uses_esca(vsie_page->scb_o);
> +	unsigned int cpu_nr, cpu_slots;
> +	struct vsie_page *vsie_page_n;
> +	hpa_t sca_o_entry_hpa;
> +	hva_t sca_o_entry_hva;
> +	unsigned long *mcn;
> +	gpa_t scb_o_gpa;
> +	int rc;
> +
> +	if (is_esca)
> +		mcn = phys_to_virt(sca_o_hpa(vsie_sca)) + offsetof(struct esca_block, mcn);
> +	else
> +		mcn = phys_to_virt(sca_o_hpa(vsie_sca)) + offsetof(struct bsca_block, mcn);
> +
> +	/* pin and make shadow for ALL scb in the sca */
> +	cpu_slots = is_esca ? KVM_S390_MAX_VSIE_VCPUS : KVM_S390_BSCA_CPU_SLOTS;
> +	for_each_set_bit_inv(cpu_nr, mcn, cpu_slots) {
> +		rc = get_sca_entry_addr(vcpu->kvm, vsie_sca, cpu_nr, NULL, &sca_o_entry_hpa);
> +		if (rc)
> +			goto err;

[ ... ]

> +	vsie_sca->ssca.osca = sca_o_hpa(vsie_sca);
> +
> +	return 0;
> +
> +err:
> +	for_each_set_bit_inv(cpu_nr, mcn, cpu_slots) {
> +		vsie_sca->ssca.cpu[cpu_nr].ssda = 0;
> +		vsie_sca->ssca.cpu[cpu_nr].ossea = 0;
> +	}

[Severity: High]
Is there a TOCTOU race when iterating over the guest-modifiable mcn mask here? 

Since mcn is read directly from guest memory twice, could a malicious guest 
toggle bits between the initialization loop and this error cleanup loop? By 
unsetting a bit before the error loop, the guest could force KVM to skip 
clearing the ssda for that CPU, leaving a valid physical pointer to a shadow 
SCB inside the active shadow SCA.

> +	return rc;
> +}
> +
> +/* Shadow or reshadow the SCA on VSIE enter. */
> +static int shadow_sca(struct kvm_vcpu *vcpu, struct vsie_page *vsie_page, struct vsie_sca *vsie_sca)
> +{
> +	int rc = 0;
> +
> +	guard(rwsem_write)(&vcpu->kvm->arch.vsie.vsie_sca_lock);
> +	if (!vsie_sca->ssca.osca)
> +		rc = _shadow_sca(vcpu, vsie_page, vsie_sca);

[Severity: Medium]
Does this fail to detect guest modifications to the SCA CPU mask after the 
first shadow? 

If the guest dynamically updates the CPU mask (e.g., setting a new bit in mcn 
to bring up a new guest-3 vCPU) and re-enters VSIE, KVM will immediately 
return without updating the shadow SCA since ssca.osca is already populated. 
Could this break dynamic CPU hotplug for nested guests?

> +
> +	return rc;
> +}
> +
>  int kvm_s390_handle_vsie(struct kvm_vcpu *vcpu)
>  {

[ ... ]

>  
> -	rc = pin_scb(vcpu, vsie_page);
> -	if (rc)
> -		goto out_put;
>  	rc = shadow_scb(vcpu, vsie_page);
>  	if (rc)
> -		goto out_unpin_scb;
> +		goto out_put;

[Severity: High]
Does the shadow SCA need to be linked to the shadow SIE control block during 
shadow_scb() setup? 

If write_scao() is not called for the active VCPU within shadow_scb() (and 
pin_blocks() also skips setting scaol when vsie_page->vsie_sca is present), it 
appears scb_s->scaol will remain 0, which would break the newly added 
VSIE-Interpretation-Extension feature.

> +	if (vsie_sca) {
> +		/* pin and shadow the sca including all scb_o in the g3 conf */
> +		rc = shadow_sca(vcpu, vsie_page, vsie_sca);
> +		if (rc)
> +			goto out_put;
> +	}
> +
>  	rc = pin_blocks(vcpu, vsie_page);
>  	if (rc)
>  		goto out_unshadow;
>  	register_shadow_scb(vcpu, vsie_page);
> +
>  	rc = vsie_run(vcpu, vsie_page);
> +
>  	unregister_shadow_scb(vcpu);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260812-vsie-sigpi-v4-0-f19e49666315@linux.ibm.com?part=14

  reply	other threads:[~2026-08-12 16:10 UTC|newest]

Thread overview: 41+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-12 13:52 [PATCH v4 00/20] KVM: s390: vsie: Add VSIE SIGP ECALL and SRSI Interpretation Christoph Schlameuss
2026-08-12 13:52 ` [PATCH v4 01/20] KVM: s390: vsie: Add SCAO read and write helpers Christoph Schlameuss
2026-08-12 13:58   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 02/20] KVM: s390: vsie: Move SCAO validation into a function Christoph Schlameuss
2026-08-12 15:13   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 03/20] KVM: s390: vsie: Add vsie_interp_extf detection Christoph Schlameuss
2026-08-12 13:58   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 04/20] KVM: s390: vsie: Add ssca_block and ssca_entry structs Christoph Schlameuss
2026-08-12 13:57   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 05/20] KVM: s390: vsie: Move pin/unpin_scb methods Christoph Schlameuss
2026-08-12 14:03   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 06/20] KVM: s390: vsie: Move pin/unpin guest page Christoph Schlameuss
2026-08-12 14:08   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 07/20] KVM: s390: vsie: Move release/acquire gmap shadow Christoph Schlameuss
2026-08-12 14:02   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 08/20] KVM: s390: vsie: Create helpers to alloc and free vsie_pages Christoph Schlameuss
2026-08-12 14:04   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 09/20] KVM: s390: vsie: Replace radix_tree with xarray addr_to_page Christoph Schlameuss
2026-08-12 14:03   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 10/20] KVM: s390: vsie: Add helper to release gmap shadow Christoph Schlameuss
2026-08-12 14:04   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 11/20] KVM: s390: vsie: Lazily keep original scb pinned after vsie exit Christoph Schlameuss
2026-08-12 15:28   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 12/20] KVM: s390: vsie: Add helper to pin and unpin multiple guest pages Christoph Schlameuss
2026-08-12 14:11   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 13/20] KVM: s390: vsie: Add struct vsie_sca with pin and unpin methods Christoph Schlameuss
2026-08-12 15:43   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 14/20] KVM: s390: vsie: Shadow VSIE SCA in guest-1 Christoph Schlameuss
2026-08-12 16:10   ` sashiko-bot [this message]
2026-08-12 13:52 ` [PATCH v4 15/20] KVM: s390: vsie: Guard against invalid CPU address Christoph Schlameuss
2026-08-12 16:23   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 16/20] KVM: s390: vsie: Allow guest-3 cpu add and remove with ssca Christoph Schlameuss
2026-08-12 16:40   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 17/20] KVM: s390: vsie: Add VSIE max shadow configuration Christoph Schlameuss
2026-08-12 16:53   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 18/20] KVM: s390: vsie: Add VSIE shadow stat counters Christoph Schlameuss
2026-08-12 17:02   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 19/20] KVM: s390: vsie: Create minimal scb shadows for not running g3 blocks Christoph Schlameuss
2026-08-12 17:26   ` sashiko-bot
2026-08-12 13:52 ` [PATCH v4 20/20] KVM: s390: vsie: Enable use of VSIE SSCA Christoph Schlameuss
2026-08-12 17:45   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260812161002.332CB1F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=agordeev@linux.ibm.com \
    --cc=borntraeger@linux.ibm.vom \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=schlameuss@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox