* [PATCH 1/3] KVM: PPC: Book3S HV: Ensure that calls to idr_alloc are synchronized
2026-10-01 10:02 [PATCH 0/3] Fixes for KVM HV around synchronization and initializations Gautam Menghani
@ 2026-10-01 10:02 ` Gautam Menghani
2026-10-01 10:02 ` [PATCH 2/3] KVM: PPC: Book3S HV: XICS: Update irq_type for cleanup to work Gautam Menghani
2026-10-01 10:02 ` [PATCH 3/3] KVM: PPC: Book3s HV: Take SRCU read lock around kvmppc_h_page_init() call Gautam Menghani
2 siblings, 0 replies; 6+ messages in thread
From: Gautam Menghani @ 2026-10-01 10:02 UTC (permalink / raw)
To: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, amachhiw,
harshpb
Cc: linuxppc-dev, kvm, linux-kernel, stable
Calls to idr_alloc() in __prealloc_nested() are not currently synchronized.
According to the documentation [1], the caller should provide their own
locking to prevent concurrent modifications to the idr to prevent bugs.
Wrap the call to __prealloc_nested() in a spinlock to prevent concurrent
modifications to idr.
[1]: lib/idr.c
Fixes: c0f00a18e2a8 ("KVM: PPC: Book3S HV Nested: Change nested guest lookup to use idr")
Cc: stable@vger.kernel.org # 5.19+
Co-developed-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Signed-off-by: Ritesh Harjani (IBM) <ritesh.list@gmail.com>
Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
---
arch/powerpc/kvm/book3s_hv_nested.c | 34 +++++++++--------------------
1 file changed, 10 insertions(+), 24 deletions(-)
diff --git a/arch/powerpc/kvm/book3s_hv_nested.c b/arch/powerpc/kvm/book3s_hv_nested.c
index a6ff42d7666c..e893fdd24f0a 100644
--- a/arch/powerpc/kvm/book3s_hv_nested.c
+++ b/arch/powerpc/kvm/book3s_hv_nested.c
@@ -702,20 +702,6 @@ static struct kvm_nested_guest *__find_nested(struct kvm *kvm, int lpid)
return idr_find(&kvm->arch.kvm_nested_guest_idr, lpid);
}
-static bool __prealloc_nested(struct kvm *kvm, int lpid)
-{
- if (idr_alloc(&kvm->arch.kvm_nested_guest_idr,
- NULL, lpid, lpid + 1, GFP_KERNEL) != lpid)
- return false;
- return true;
-}
-
-static void __add_nested(struct kvm *kvm, int lpid, struct kvm_nested_guest *gp)
-{
- if (idr_replace(&kvm->arch.kvm_nested_guest_idr, gp, lpid))
- WARN_ON(1);
-}
-
static void __remove_nested(struct kvm *kvm, int lpid)
{
idr_remove(&kvm->arch.kvm_nested_guest_idr, lpid);
@@ -862,21 +848,21 @@ struct kvm_nested_guest *kvmhv_get_nested(struct kvm *kvm, int l1_lpid,
if (!newgp)
return NULL;
- if (!__prealloc_nested(kvm, l1_lpid)) {
- kvmhv_release_nested(newgp);
- return NULL;
- }
-
+ idr_preload(GFP_KERNEL);
spin_lock(&kvm->mmu_lock);
gp = __find_nested(kvm, l1_lpid);
if (!gp) {
- __add_nested(kvm, l1_lpid, newgp);
- ++newgp->refcnt;
- gp = newgp;
- newgp = NULL;
+ if (idr_alloc(&kvm->arch.kvm_nested_guest_idr, newgp,
+ l1_lpid, l1_lpid + 1, GFP_NOWAIT) >= 0) {
+ ++newgp->refcnt;
+ gp = newgp;
+ newgp = NULL;
+ }
}
- ++gp->refcnt;
+ if (gp)
+ ++gp->refcnt;
spin_unlock(&kvm->mmu_lock);
+ idr_preload_end();
if (newgp)
kvmhv_release_nested(newgp);
--
2.52.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* [PATCH 2/3] KVM: PPC: Book3S HV: XICS: Update irq_type for cleanup to work
2026-10-01 10:02 [PATCH 0/3] Fixes for KVM HV around synchronization and initializations Gautam Menghani
2026-10-01 10:02 ` [PATCH 1/3] KVM: PPC: Book3S HV: Ensure that calls to idr_alloc are synchronized Gautam Menghani
@ 2026-10-01 10:02 ` Gautam Menghani
2026-10-01 10:10 ` sashiko-bot
2026-10-01 10:02 ` [PATCH 3/3] KVM: PPC: Book3s HV: Take SRCU read lock around kvmppc_h_page_init() call Gautam Menghani
2 siblings, 1 reply; 6+ messages in thread
From: Gautam Menghani @ 2026-10-01 10:02 UTC (permalink / raw)
To: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, amachhiw,
harshpb
Cc: linuxppc-dev, kvm, linux-kernel, stable
In kvmppc_xive_connect_vcpu(), if any failures are encountered, the
control flow jumps to the 'bail' label where kvmppc_xive_cleanup_vcpu()
is called. This cleanup function exits prematurely since
vcpu->arch.irq_type is set to KVMPPC_IRQ_DEFAULT at this point. This ends
up resulting in wasted memory.
Fix this by setting irq_type to KVMPPC_IRQ_XICS before vcpu configuration
starts. This makes XICS initialization symmetrical to XIVE
initialization done in kvmppc_xive_native_connect_vcpu().
Fixes: 5af50993850a ("KVM: PPC: Book3S HV: Native usage of the XIVE interrupt controller")
Cc: stable@vger.kernel.org # 4.12+
Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
---
arch/powerpc/kvm/book3s_xive.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/arch/powerpc/kvm/book3s_xive.c b/arch/powerpc/kvm/book3s_xive.c
index 1d67237783b7..67219044fb8c 100644
--- a/arch/powerpc/kvm/book3s_xive.c
+++ b/arch/powerpc/kvm/book3s_xive.c
@@ -1937,6 +1937,7 @@ int kvmppc_xive_connect_vcpu(struct kvm_device *dev,
xc->vp_id = vp_id;
xc->mfrr = 0xff;
xc->valid = true;
+ vcpu->arch.irq_type = KVMPPC_IRQ_XICS;
r = xive_native_get_vp_info(xc->vp_id, &xc->vp_cam, &xc->vp_chip_id);
if (r)
@@ -2025,7 +2026,6 @@ int kvmppc_xive_connect_vcpu(struct kvm_device *dev,
return r;
}
- vcpu->arch.irq_type = KVMPPC_IRQ_XICS;
return 0;
}
--
2.52.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* Re: [PATCH 2/3] KVM: PPC: Book3S HV: XICS: Update irq_type for cleanup to work
2026-10-01 10:02 ` [PATCH 2/3] KVM: PPC: Book3S HV: XICS: Update irq_type for cleanup to work Gautam Menghani
@ 2026-10-01 10:10 ` sashiko-bot
2026-10-09 8:21 ` Gautam Menghani
0 siblings, 1 reply; 6+ messages in thread
From: sashiko-bot @ 2026-10-01 10:10 UTC (permalink / raw)
To: Gautam Menghani; +Cc: kvm
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Critical] xive: host kernel panic in kvmppc_xive_cleanup_vcpu() due to uninitialized eoi_mmio
--
commit 2d2dc3956268ca839bb2cf6d3521ccdcf2a1831b
Author: Gautam Menghani <gautam@linux.ibm.com>
KVM: PPC: Book3S HV: XICS: Update irq_type for cleanup to work
This commit updates the assignment of vcpu->arch.irq_type to occur before
vcpu configuration starts. This is intended to ensure that
kvmppc_xive_cleanup_vcpu() can properly execute if early setup failures are
encountered during kvmppc_xive_connect_vcpu().
> diff --git a/arch/powerpc/kvm/book3s_xive.c b/arch/powerpc/kvm/book3s_xive.c
> index 1d67237783b71..67219044fb8c7 100644
> --- a/arch/powerpc/kvm/book3s_xive.c
> +++ b/arch/powerpc/kvm/book3s_xive.c
> @@ -1937,6 +1937,7 @@ int kvmppc_xive_connect_vcpu(struct kvm_device *dev,
> xc->vp_id = vp_id;
> xc->mfrr = 0xff;
> xc->valid = true;
> + vcpu->arch.irq_type = KVMPPC_IRQ_XICS;
[Severity: Critical]
Could this cause a regression where an early failure leads to a host
kernel panic?
If xive_native_get_vp_info() or xive_native_alloc_irq() fails shortly after
this assignment, the code jumps to the bail label and calls
kvmppc_xive_cleanup_vcpu().
Because irq_type is now KVMPPC_IRQ_XICS, kvmppc_xive_cleanup_vcpu() no
longer bails out early, and unconditionally attempts to mask the VP IPI:
kvmppc_xive_cleanup_vcpu()
/* Mask the VP IPI */
xive_vm_esb_load(&xc->vp_ipi_data, XIVE_ESB_SET_PQ_01);
Since xive_native_populate_irq_data() has not been successfully executed yet,
xc->vp_ipi_data.eoi_mmio is still uninitialized. When xive_vm_esb_load()
is called:
xive_vm_esb_load()
val = __raw_readq(__x_eoi_page(xd) + offset);
Would this dereference the uninitialized eoi_mmio pointer and cause a
fatal page fault?
Could an unprivileged host userspace process intentionally trigger this
early failure path (e.g., by exhausting IRQ resources) during VM creation,
resulting in a host denial of service?
>
> r = xive_native_get_vp_info(xc->vp_id, &xc->vp_cam, &xc->vp_chip_id);
> if (r)
[ ... ]
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261001100208.1271628-1-gautam@linux.ibm.com?part=2
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH 2/3] KVM: PPC: Book3S HV: XICS: Update irq_type for cleanup to work
2026-10-01 10:10 ` sashiko-bot
@ 2026-10-09 8:21 ` Gautam Menghani
0 siblings, 0 replies; 6+ messages in thread
From: Gautam Menghani @ 2026-10-09 8:21 UTC (permalink / raw)
To: sashiko-reviews; +Cc: kvm, linuxppc-dev
On Thu, Oct 01, 2026 at 10:10:51AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [Critical] xive: host kernel panic in kvmppc_xive_cleanup_vcpu() due to uninitialized eoi_mmio
> --
>
> commit 2d2dc3956268ca839bb2cf6d3521ccdcf2a1831b
> Author: Gautam Menghani <gautam@linux.ibm.com>
>
> KVM: PPC: Book3S HV: XICS: Update irq_type for cleanup to work
>
> This commit updates the assignment of vcpu->arch.irq_type to occur before
> vcpu configuration starts. This is intended to ensure that
> kvmppc_xive_cleanup_vcpu() can properly execute if early setup failures are
> encountered during kvmppc_xive_connect_vcpu().
>
> > diff --git a/arch/powerpc/kvm/book3s_xive.c b/arch/powerpc/kvm/book3s_xive.c
> > index 1d67237783b71..67219044fb8c7 100644
> > --- a/arch/powerpc/kvm/book3s_xive.c
> > +++ b/arch/powerpc/kvm/book3s_xive.c
> > @@ -1937,6 +1937,7 @@ int kvmppc_xive_connect_vcpu(struct kvm_device *dev,
> > xc->vp_id = vp_id;
> > xc->mfrr = 0xff;
> > xc->valid = true;
> > + vcpu->arch.irq_type = KVMPPC_IRQ_XICS;
>
> [Severity: Critical]
> Could this cause a regression where an early failure leads to a host
> kernel panic?
>
> If xive_native_get_vp_info() or xive_native_alloc_irq() fails shortly after
> this assignment, the code jumps to the bail label and calls
> kvmppc_xive_cleanup_vcpu().
>
> Because irq_type is now KVMPPC_IRQ_XICS, kvmppc_xive_cleanup_vcpu() no
> longer bails out early, and unconditionally attempts to mask the VP IPI:
>
> kvmppc_xive_cleanup_vcpu()
> /* Mask the VP IPI */
> xive_vm_esb_load(&xc->vp_ipi_data, XIVE_ESB_SET_PQ_01);
>
> Since xive_native_populate_irq_data() has not been successfully executed yet,
> xc->vp_ipi_data.eoi_mmio is still uninitialized. When xive_vm_esb_load()
> is called:
>
> xive_vm_esb_load()
> val = __raw_readq(__x_eoi_page(xd) + offset);
>
> Would this dereference the uninitialized eoi_mmio pointer and cause a
> fatal page fault?
Yes, this is a valid bug. I'll fix it in v2.
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH 3/3] KVM: PPC: Book3s HV: Take SRCU read lock around kvmppc_h_page_init() call
2026-10-01 10:02 [PATCH 0/3] Fixes for KVM HV around synchronization and initializations Gautam Menghani
2026-10-01 10:02 ` [PATCH 1/3] KVM: PPC: Book3S HV: Ensure that calls to idr_alloc are synchronized Gautam Menghani
2026-10-01 10:02 ` [PATCH 2/3] KVM: PPC: Book3S HV: XICS: Update irq_type for cleanup to work Gautam Menghani
@ 2026-10-01 10:02 ` Gautam Menghani
2 siblings, 0 replies; 6+ messages in thread
From: Gautam Menghani @ 2026-10-01 10:02 UTC (permalink / raw)
To: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, amachhiw,
harshpb
Cc: linuxppc-dev, kvm, linux-kernel, stable
Since there is no srcu read lock around kvmppc_h_page_init(), a
"suspicious rcu_dereference_check() usage" can be triggered by running
host linux with CONFIG_PROVE_RCU=y and the following code in the guest:
plpar_hcall_norets(H_PAGE_INIT, H_ZERO_PAGE, phys_addr, 0);
=============================
WARNING: suspicious RCU usage
7.0.0-dirty #227 Not tainted
-----------------------------
./include/linux/kvm_host.h:1074 suspicious rcu_dereference_check() usage!
other info that might help us debug this:
rcu_scheduler_active = 2, debug_locks = 1
1 lock held by qemu-system-ppc/2093:
#0: c000000354197ab0 (&vcpu->mutex){+.+.}-{3:3}, at: kvm_vcpu_ioctl+0x10c/0xaf0
stack backtrace:
CPU: 30 UID: 0 PID: 2093 Comm: qemu-system-ppc Not tainted 7.0.0-dirty #227 PREEMPT(full)
Hardware name: IBM,9080-HEX POWER10 (architected) 0x800200 0xf000006 of:IBM,FW1060.60 (NH1060_158) hv:phyp pSeries
Call Trace:
[c00000033e1bf430] [c000000001b3dbd4] dump_stack_lvl+0xc8/0x130 (unreliable)
[c00000033e1bf470] [c0000000003da964] lockdep_rcu_suspicious+0x1f4/0x290
[c00000033e1bf520] [c000000000231424] gfn_to_memslot+0x1b4/0x1c0
[c00000033e1bf560] [c00000000023637c] kvm_clear_guest+0x9c/0x110
[c00000033e1bf5c0] [c0000000002714a4] kvmppc_h_page_init+0x160/0x1a0
[c00000033e1bf610] [c00000000026c040] kvmppc_pseries_do_hcall+0x1930/0x1940
[c00000033e1bf6d0] [c00000000026fb18] kvmppc_vcpu_run_hv+0x268/0x800
[c00000033e1bf7a0] [c000000000246b30] kvmppc_vcpu_run+0x30/0x50
[c00000033e1bf7c0] [c000000000241ed4] kvm_arch_vcpu_ioctl_run+0x364/0x530
[c00000033e1bf860] [c00000000022b870] kvm_vcpu_ioctl+0x1b0/0xaf0
[c00000033e1bfa50] [c0000000009b4514] sys_ioctl+0x144/0x190
[c00000033e1bfab0] [c000000000034330] system_call_exception+0x170/0x370
[c00000033e1bfe50] [c00000000000d05c] system_call_vectored_common+0x15c/0x2ec
Fix this by taking a srcu read lock around kvmppc_h_page_init().
Fixes: 2d34d1c3bbfd ("KVM: PPC: Book3S HV: Implement virtual mode H_PAGE_INIT handler")
Cc: stable@vger.kernel.org # 5.2+
Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
---
arch/powerpc/kvm/book3s_hv.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
index dbac3573b2c8..8678a93bac51 100644
--- a/arch/powerpc/kvm/book3s_hv.c
+++ b/arch/powerpc/kvm/book3s_hv.c
@@ -1386,9 +1386,11 @@ int kvmppc_pseries_do_hcall(struct kvm_vcpu *vcpu)
ret = kvmhv_copy_tofrom_guest_nested(vcpu);
break;
case H_PAGE_INIT:
+ kvm_vcpu_srcu_read_lock(vcpu);
ret = kvmppc_h_page_init(vcpu, kvmppc_get_gpr(vcpu, 4),
kvmppc_get_gpr(vcpu, 5),
kvmppc_get_gpr(vcpu, 6));
+ kvm_vcpu_srcu_read_unlock(vcpu);
break;
case H_SVM_PAGE_IN:
ret = H_UNSUPPORTED;
--
2.52.0
^ permalink raw reply related [flat|nested] 6+ messages in thread