* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
2026-09-21 11:10 [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit Gautam Menghani
@ 2026-09-21 11:25 ` sashiko-bot
2026-09-22 5:59 ` Gautam Menghani
2026-09-23 3:38 ` Narayana Murty N
` (2 subsequent siblings)
3 siblings, 1 reply; 9+ messages in thread
From: sashiko-bot @ 2026-09-21 11:25 UTC (permalink / raw)
To: Gautam Menghani; +Cc: kvm
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] kvm: ppc: lost LPCR_MER interrupts for XICS due to XIVE state leak
--
commit 215dda3f95780a965ce67c9c452b6f7d1948e486
Author: Gautam Menghani <gautam@linux.ibm.com>
KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
This commit avoids spurious interrupts for KVM on PowerNV guests by not
setting the LPCR_MER bit when native XIVE is available. It introduces a new
helper, kvmppc_xive_native_enabled(), to check for native XIVE and skips
setting LPCR_MER if it evaluates to true.
Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> index 169ea6a7fbad4..580ad2548c2bf 100644
> --- a/arch/powerpc/include/asm/kvm_ppc.h
> +++ b/arch/powerpc/include/asm/kvm_ppc.h
> @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> }
>
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> +{
> + return kvm->arch.xive_devices.native;
> +}
> +
[ ... ]
> diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> index dbac3573b2c89..cb2bb29451a7f 100644
> --- a/arch/powerpc/kvm/book3s_hv.c
> +++ b/arch/powerpc/kvm/book3s_hv.c
> @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> kvmppc_inject_interrupt_hv(vcpu,
> BOOK3S_INTERRUPT_EXTERNAL, 0);
> - else
> + else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> lpcr |= LPCR_MER;
> } else {
> /*
[Severity: High]
Does kvmppc_xive_native_enabled() accurately reflect whether XIVE is currently
in use?
If a VM switches from XIVE to XICS (such as during a kexec or CAS
renegotiation), the XIVE native device is destroyed. However, the native
pointer is intentionally not cleared:
arch/powerpc/kvm/book3s_xive_native.c:kvmppc_xive_native_release() {
...
/*
* A reference of the kvmppc_xive pointer is now kept under
* the xive_devices struct of the machine for reuse. It is
* freed when the VM is destroyed for now until we fix all the
* execution paths.
*/
kfree(dev);
}
When running in XICS mode after this switch, kvmppc_xive_native_enabled()
will continue to evaluate to true.
Will this cause KVM to bypass setting LPCR_MER, and subsequently fail to
inject external interrupts when MSR_EE=0, resulting in a guest hang?
Could this check instead evaluate the active vCPU IRQ routing, for example by
using !kvmppc_xive_enabled(vcpu)?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260921111043.49825-1-gautam@linux.ibm.com?part=1
^ permalink raw reply [flat|nested] 9+ messages in thread* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
2026-09-21 11:25 ` sashiko-bot
@ 2026-09-22 5:59 ` Gautam Menghani
0 siblings, 0 replies; 9+ messages in thread
From: Gautam Menghani @ 2026-09-22 5:59 UTC (permalink / raw)
To: sashiko-reviews; +Cc: kvm, linuxppc-dev
On Mon, Sep 21, 2026 at 11:25:38AM +0000, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] kvm: ppc: lost LPCR_MER interrupts for XICS due to XIVE state leak
> --
>
> commit 215dda3f95780a965ce67c9c452b6f7d1948e486
> Author: Gautam Menghani <gautam@linux.ibm.com>
>
> KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
>
> This commit avoids spurious interrupts for KVM on PowerNV guests by not
> setting the LPCR_MER bit when native XIVE is available. It introduces a new
> helper, kvmppc_xive_native_enabled(), to check for native XIVE and skips
> setting LPCR_MER if it evaluates to true.
>
> Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
>
> > diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> > index 169ea6a7fbad4..580ad2548c2bf 100644
> > --- a/arch/powerpc/include/asm/kvm_ppc.h
> > +++ b/arch/powerpc/include/asm/kvm_ppc.h
> > @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> > return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> > }
> >
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> > +{
> > + return kvm->arch.xive_devices.native;
> > +}
> > +
>
> [ ... ]
>
> > diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> > index dbac3573b2c89..cb2bb29451a7f 100644
> > --- a/arch/powerpc/kvm/book3s_hv.c
> > +++ b/arch/powerpc/kvm/book3s_hv.c
> > @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> > if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> > kvmppc_inject_interrupt_hv(vcpu,
> > BOOK3S_INTERRUPT_EXTERNAL, 0);
> > - else
> > + else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> > lpcr |= LPCR_MER;
> > } else {
> > /*
>
> [Severity: High]
> Does kvmppc_xive_native_enabled() accurately reflect whether XIVE is currently
> in use?
>
> If a VM switches from XIVE to XICS (such as during a kexec or CAS
> renegotiation), the XIVE native device is destroyed. However, the native
> pointer is intentionally not cleared:
The kexec case is not supported for guests - booting a KVM guest in XIVE mode and
kexec to XICS mode.
>
> arch/powerpc/kvm/book3s_xive_native.c:kvmppc_xive_native_release() {
> ...
> /*
> * A reference of the kvmppc_xive pointer is now kept under
> * the xive_devices struct of the machine for reuse. It is
> * freed when the VM is destroyed for now until we fix all the
> * execution paths.
> */
>
> kfree(dev);
> }
>
> When running in XICS mode after this switch, kvmppc_xive_native_enabled()
> will continue to evaluate to true.
>
> Will this cause KVM to bypass setting LPCR_MER, and subsequently fail to
> inject external interrupts when MSR_EE=0, resulting in a guest hang?
No. If LPCR_MER is not set, the only effect of that is the vCPU will have to
exit to the host for the host to inject interrupts into the vCPU. One
extra exit isn't a big deal.
>
> Could this check instead evaluate the active vCPU IRQ routing, for example by
> using !kvmppc_xive_enabled(vcpu)?
kvmppc_xive_enabled() does not check for native/irqchip=on case. It
returns true even if emulation is used.
>
> --
> Sashiko AI review · https://sashiko.dev/#/patchset/20260921111043.49825-1-gautam@linux.ibm.com?part=1
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
2026-09-21 11:10 [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit Gautam Menghani
2026-09-21 11:25 ` sashiko-bot
@ 2026-09-23 3:38 ` Narayana Murty N
2026-10-01 3:58 ` Gautam Menghani
2026-09-23 7:01 ` Mukesh Kumar Chaurasiya
2026-10-05 8:00 ` Amit Machhiwal
3 siblings, 1 reply; 9+ messages in thread
From: Narayana Murty N @ 2026-09-23 3:38 UTC (permalink / raw)
To: Gautam Menghani, maddy, npiggin, mpe, chleroy, ritesh.list,
sshegde, amachhiw
Cc: linuxppc-dev, kvm, linux-kernel, stable, Timothy Pearson
Hi Gautam,
On 21/09/26 4:40 PM, Gautam Menghani wrote:
> A huge number of spurious interrupts can be seen immediately after a KVM
> on PowerNV guest boots up in XIVE mode.
>
> $ cat /proc/interrupts | grep SPU
> SPU: 223705 192439 273526 147623 Spurious interrupts
>
> This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> Handle pending exceptions on guest entry with MSR_EE"). The root cause
> is that once LPCR_MER bit is set, it is supposed to be reset by
> software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> not exit back to the host until the decrementer expires or there is an
> hcall, etc. This is because KVM on PowerNV guests have support for
> native XIVE, so they are not dependent on host for interrupt emulation.
> Due to this behaviour, a huge number of spurious interrupts are seen
> since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> vCPU exits to the host.
>
> Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> available (currently in case of KVM on PowerNV only), as the XIVE hardware
> can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> functionality is not required. This reduces the number of spurious
> interrupts drastically.
>
> Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> Cc: stable@vger.kernel.org # 6.8+
> Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
> Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
> ---
> v3:
> 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
>
> v2:
> 1. Handle the case where xive_interrupt_pending() is true and also the
> external exception bit is set. (Narayana)
>
> arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
> arch/powerpc/kvm/book3s_hv.c | 2 +-
> 2 files changed, 8 insertions(+), 1 deletion(-)
>
> diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> index 169ea6a7fbad..580ad2548c2b 100644
> --- a/arch/powerpc/include/asm/kvm_ppc.h
> +++ b/arch/powerpc/include/asm/kvm_ppc.h
> @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> }
>
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> +{
> + return kvm->arch.xive_devices.native;
> +}
Also, should 'kvmppc_xive_native_enabled()' check the active XIVE
device rather than just whether a native device has been created?
'kvmppc_xive_native_release()' clears 'kvm->arch.xive', but explicitly
keeps the 'kvmppc_xive' pointer under 'xive_devices' for reuse. The
comment in 'book3s_xive.c' also says that when switching between XICS
and native XIVE, the previous KVM device is released before the new one
is created.
So 'xive_devices.native != NULL' seems to mean that a native XIVE
device has been allocated/cached, rather than that it is currently
active.
Would this be more appropriate?
return kvm->arch.xive_devices.native &&
kvm->arch.xive == kvm->arch.xive_devices.native;
> +
> extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> struct kvm_vcpu *vcpu, u32 cpu);
> extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
>
> static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> { return 0; }
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> +
> static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
> static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> index dbac3573b2c8..cb2bb29451a7 100644
> --- a/arch/powerpc/kvm/book3s_hv.c
> +++ b/arch/powerpc/kvm/book3s_hv.c
> @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> kvmppc_inject_interrupt_hv(vcpu,
> BOOK3S_INTERRUPT_EXTERNAL, 0);
> - else
> + else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> lpcr |= LPCR_MER;
Is the intent here to suppress LPCR_MER only for interrupts that
native XIVE can deliver directly?
BOOK3S_IRQPRIO_EXTERNAL is a separate software-pending exception. If
MSR_EE is clear it remains pending, so it seems this path should still
set MER even when native XIVE is active.
IOW, should the behavior be:
else if (!kvmppc_xive_native_enabled(vcpu->kvm) ||
test_bit(BOOK3S_IRQPRIO_EXTERNAL,
&vcpu->arch.pending_exceptions))
lpcr |= LPCR_MER;
so that only the native-XIVE-pending case suppresses MER?
Thanks,
narayana Murty N
> } else {
> /*
^ permalink raw reply [flat|nested] 9+ messages in thread* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
2026-09-23 3:38 ` Narayana Murty N
@ 2026-10-01 3:58 ` Gautam Menghani
0 siblings, 0 replies; 9+ messages in thread
From: Gautam Menghani @ 2026-10-01 3:58 UTC (permalink / raw)
To: Narayana Murty N
Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, amachhiw,
linuxppc-dev, kvm, linux-kernel, stable, Timothy Pearson
On Wed, Sep 23, 2026 at 09:08:57AM +0530, Narayana Murty N wrote:
> Hi Gautam,
>
> On 21/09/26 4:40 PM, Gautam Menghani wrote:
> > A huge number of spurious interrupts can be seen immediately after a KVM
> > on PowerNV guest boots up in XIVE mode.
> >
> > $ cat /proc/interrupts | grep SPU
> > SPU: 223705 192439 273526 147623 Spurious interrupts
> >
> > This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> > Handle pending exceptions on guest entry with MSR_EE"). The root cause
> > is that once LPCR_MER bit is set, it is supposed to be reset by
> > software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> > not exit back to the host until the decrementer expires or there is an
> > hcall, etc. This is because KVM on PowerNV guests have support for
> > native XIVE, so they are not dependent on host for interrupt emulation.
> > Due to this behaviour, a huge number of spurious interrupts are seen
> > since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> > vCPU exits to the host.
> >
> > Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> > available (currently in case of KVM on PowerNV only), as the XIVE hardware
> > can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> > functionality is not required. This reduces the number of spurious
> > interrupts drastically.
> >
> > Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> > Cc: stable@vger.kernel.org # 6.8+
> > Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
> > Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> > Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
> > ---
> > v3:
> > 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
> >
> > v2:
> > 1. Handle the case where xive_interrupt_pending() is true and also the
> > external exception bit is set. (Narayana)
> >
> > arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
> > arch/powerpc/kvm/book3s_hv.c | 2 +-
> > 2 files changed, 8 insertions(+), 1 deletion(-)
> >
> > diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> > index 169ea6a7fbad..580ad2548c2b 100644
> > --- a/arch/powerpc/include/asm/kvm_ppc.h
> > +++ b/arch/powerpc/include/asm/kvm_ppc.h
> > @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> > return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> > }
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> > +{
> > + return kvm->arch.xive_devices.native;
> > +}
> Also, should 'kvmppc_xive_native_enabled()' check the active XIVE
> device rather than just whether a native device has been created?
>
> 'kvmppc_xive_native_release()' clears 'kvm->arch.xive', but explicitly
> keeps the 'kvmppc_xive' pointer under 'xive_devices' for reuse. The
> comment in 'book3s_xive.c' also says that when switching between XICS
> and native XIVE, the previous KVM device is released before the new one
> is created.
>
> So 'xive_devices.native != NULL' seems to mean that a native XIVE
> device has been allocated/cached, rather than that it is currently
> active.
In a KVM guest on PowerNV, whatever interrupt mode the guest is booted
in (XIVE/ XICS), you cannot switch it with kexec. So if
'xive_devices.native' exists, we can be confident that native XIVE is
being used. So the code in the patch should be fine.
>
> Would this be more appropriate?
>
> return kvm->arch.xive_devices.native &&
> kvm->arch.xive == kvm->arch.xive_devices.native;
>
> > +
> > extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> > struct kvm_vcpu *vcpu, u32 cpu);
> > extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> > @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
> > static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> > { return 0; }
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> > +
> > static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> > struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
> > static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> > diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> > index dbac3573b2c8..cb2bb29451a7 100644
> > --- a/arch/powerpc/kvm/book3s_hv.c
> > +++ b/arch/powerpc/kvm/book3s_hv.c
> > @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> > if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> > kvmppc_inject_interrupt_hv(vcpu,
> > BOOK3S_INTERRUPT_EXTERNAL, 0);
> > - else
> > + else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> > lpcr |= LPCR_MER;
> Is the intent here to suppress LPCR_MER only for interrupts that
> native XIVE can deliver directly?
>
> BOOK3S_IRQPRIO_EXTERNAL is a separate software-pending exception. If
> MSR_EE is clear it remains pending, so it seems this path should still
> set MER even when native XIVE is active.
>
> IOW, should the behavior be:
>
> else if (!kvmppc_xive_native_enabled(vcpu->kvm) ||
> test_bit(BOOK3S_IRQPRIO_EXTERNAL,
> &vcpu->arch.pending_exceptions))
> lpcr |= LPCR_MER;
>
> so that only the native-XIVE-pending case suppresses MER?
No this would still result in spurious flood. If native XIVE is being
used, LPCR_MER cannot be set in any case. Software interrupts will
continue to work in the same way they work in an LPAR.
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
2026-09-21 11:10 [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit Gautam Menghani
2026-09-21 11:25 ` sashiko-bot
2026-09-23 3:38 ` Narayana Murty N
@ 2026-09-23 7:01 ` Mukesh Kumar Chaurasiya
2026-09-23 12:50 ` Gautam Menghani
2026-10-05 8:00 ` Amit Machhiwal
3 siblings, 1 reply; 9+ messages in thread
From: Mukesh Kumar Chaurasiya @ 2026-09-23 7:01 UTC (permalink / raw)
To: Gautam Menghani
Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux,
amachhiw, linuxppc-dev, kvm, linux-kernel, stable,
Timothy Pearson
On Mon, Sep 21, 2026 at 04:40:42PM +0530, Gautam Menghani wrote:
> A huge number of spurious interrupts can be seen immediately after a KVM
> on PowerNV guest boots up in XIVE mode.
>
> $ cat /proc/interrupts | grep SPU
> SPU: 223705 192439 273526 147623 Spurious interrupts
>
> This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> Handle pending exceptions on guest entry with MSR_EE"). The root cause
> is that once LPCR_MER bit is set, it is supposed to be reset by
> software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> not exit back to the host until the decrementer expires or there is an
> hcall, etc. This is because KVM on PowerNV guests have support for
> native XIVE, so they are not dependent on host for interrupt emulation.
> Due to this behaviour, a huge number of spurious interrupts are seen
> since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> vCPU exits to the host.
>
> Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> available (currently in case of KVM on PowerNV only), as the XIVE hardware
> can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> functionality is not required. This reduces the number of spurious
> interrupts drastically.
Hey Gautam,
Thanks for this. Can you also share the numbers after the change.
Regards,
Mukesh
[...]
^ permalink raw reply [flat|nested] 9+ messages in thread* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
2026-09-23 7:01 ` Mukesh Kumar Chaurasiya
@ 2026-09-23 12:50 ` Gautam Menghani
0 siblings, 0 replies; 9+ messages in thread
From: Gautam Menghani @ 2026-09-23 12:50 UTC (permalink / raw)
To: Mukesh Kumar Chaurasiya
Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux,
amachhiw, linuxppc-dev, kvm, linux-kernel, stable,
Timothy Pearson
On Wed, Sep 23, 2026 at 12:31:05PM +0530, Mukesh Kumar Chaurasiya wrote:
> On Mon, Sep 21, 2026 at 04:40:42PM +0530, Gautam Menghani wrote:
> > A huge number of spurious interrupts can be seen immediately after a KVM
> > on PowerNV guest boots up in XIVE mode.
> >
> > $ cat /proc/interrupts | grep SPU
> > SPU: 223705 192439 273526 147623 Spurious interrupts
> >
> > This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> > Handle pending exceptions on guest entry with MSR_EE"). The root cause
> > is that once LPCR_MER bit is set, it is supposed to be reset by
> > software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> > not exit back to the host until the decrementer expires or there is an
> > hcall, etc. This is because KVM on PowerNV guests have support for
> > native XIVE, so they are not dependent on host for interrupt emulation.
> > Due to this behaviour, a huge number of spurious interrupts are seen
> > since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> > vCPU exits to the host.
> >
> > Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> > available (currently in case of KVM on PowerNV only), as the XIVE hardware
> > can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> > functionality is not required. This reduces the number of spurious
> > interrupts drastically.
> Hey Gautam,
>
> Thanks for this. Can you also share the numbers after the change.
>
Sure, here are the numbers seen immediately on boot with this patch
applied:
$ cat /proc/interrupts | grep SPU
SPU: 123 139 99 144 Spurious interrupts
> Regards,
> Mukesh
> [...]
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
2026-09-21 11:10 [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit Gautam Menghani
` (2 preceding siblings ...)
2026-09-23 7:01 ` Mukesh Kumar Chaurasiya
@ 2026-10-05 8:00 ` Amit Machhiwal
2026-10-05 8:31 ` Gautam Menghani
3 siblings, 1 reply; 9+ messages in thread
From: Amit Machhiwal @ 2026-10-05 8:00 UTC (permalink / raw)
To: Gautam Menghani
Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux,
amachhiw, linuxppc-dev, kvm, linux-kernel, stable,
Timothy Pearson
Hi Gautam,
Thanks for the patch and working on this. I have a couple of questions though.
On 2026/09/21 04:40 PM, Gautam Menghani wrote:
> A huge number of spurious interrupts can be seen immediately after a KVM
> on PowerNV guest boots up in XIVE mode.
>
> $ cat /proc/interrupts | grep SPU
> SPU: 223705 192439 273526 147623 Spurious interrupts
>
> This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> Handle pending exceptions on guest entry with MSR_EE"). The root cause
> is that once LPCR_MER bit is set, it is supposed to be reset by
> software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> not exit back to the host until the decrementer expires or there is an
> hcall, etc. This is because KVM on PowerNV guests have support for
> native XIVE, so they are not dependent on host for interrupt emulation.
> Due to this behaviour, a huge number of spurious interrupts are seen
> since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> vCPU exits to the host.
>
> Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> available (currently in case of KVM on PowerNV only), as the XIVE hardware
> can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> functionality is not required. This reduces the number of spurious
> interrupts drastically.
>
> Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> Cc: stable@vger.kernel.org # 6.8+
> Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
> Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
> ---
> v3:
> 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
>
> v2:
> 1. Handle the case where xive_interrupt_pending() is true and also the
> external exception bit is set. (Narayana)
>
> arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
> arch/powerpc/kvm/book3s_hv.c | 2 +-
> 2 files changed, 8 insertions(+), 1 deletion(-)
>
> diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> index 169ea6a7fbad..580ad2548c2b 100644
> --- a/arch/powerpc/include/asm/kvm_ppc.h
> +++ b/arch/powerpc/include/asm/kvm_ppc.h
> @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> }
>
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> +{
> + return kvm->arch.xive_devices.native;
> +}
> +
> extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> struct kvm_vcpu *vcpu, u32 cpu);
> extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
>
> static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> { return 0; }
> +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> +
> static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
> static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> index dbac3573b2c8..cb2bb29451a7 100644
> --- a/arch/powerpc/kvm/book3s_hv.c
> +++ b/arch/powerpc/kvm/book3s_hv.c
> @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> kvmppc_inject_interrupt_hv(vcpu,
> BOOK3S_INTERRUPT_EXTERNAL, 0);
> - else
> + else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> lpcr |= LPCR_MER;
1. What happens when an L1 KVM guest is booted with (native) XIVE and then the
guest is rebooted with `xive=off` i.e., with XICS? Will we stop setting
LPCR_MER as kvm->arch.xive_devices.native would still be set?
2. What happends when an L1 KVM guest is booted with XIVE and then we kexec into
a new kernel with `xive=off`? You did mention in your other reply that with
kexec, `xive=off` is ignored currently but IMO, we would want to understand
where this limiation lies and fix that if need be. I do understand we are
trying to fix spurious interrupts problem with XIVE in this patch and this
particular problem can be taken separately but it worth investigating.
Thanks,
Amit
^ permalink raw reply [flat|nested] 9+ messages in thread* Re: [PATCH v3] KVM: PPC: Book3S HV: Avoid spurious interrupts caused by LPCR_MER bit
2026-10-05 8:00 ` Amit Machhiwal
@ 2026-10-05 8:31 ` Gautam Menghani
0 siblings, 0 replies; 9+ messages in thread
From: Gautam Menghani @ 2026-10-05 8:31 UTC (permalink / raw)
To: Amit Machhiwal
Cc: maddy, npiggin, mpe, chleroy, ritesh.list, sshegde, nnmlinux,
linuxppc-dev, kvm, linux-kernel, stable, Timothy Pearson
On Mon, Oct 05, 2026 at 01:30:59PM +0530, Amit Machhiwal wrote:
> Hi Gautam,
>
> Thanks for the patch and working on this. I have a couple of questions though.
>
> On 2026/09/21 04:40 PM, Gautam Menghani wrote:
> > A huge number of spurious interrupts can be seen immediately after a KVM
> > on PowerNV guest boots up in XIVE mode.
> >
> > $ cat /proc/interrupts | grep SPU
> > SPU: 223705 192439 273526 147623 Spurious interrupts
> >
> > This bug was introduced by commit ecd10702baae5 ("KVM: PPC: Book3S HV:
> > Handle pending exceptions on guest entry with MSR_EE"). The root cause
> > is that once LPCR_MER bit is set, it is supposed to be reset by
> > software. But when a vCPU starts running with LPCR_MER set, the vCPU does
> > not exit back to the host until the decrementer expires or there is an
> > hcall, etc. This is because KVM on PowerNV guests have support for
> > native XIVE, so they are not dependent on host for interrupt emulation.
> > Due to this behaviour, a huge number of spurious interrupts are seen
> > since LPCR_MER continues to be set and LPCR_MER cannot be reset until the
> > vCPU exits to the host.
> >
> > Fix this behaviour by not using the LPCR_MER bit whenever native XIVE is
> > available (currently in case of KVM on PowerNV only), as the XIVE hardware
> > can present interrupts to the KVM guest vCPU directly. So the LPCR_MER
> > functionality is not required. This reduces the number of spurious
> > interrupts drastically.
> >
> > Fixes: ecd10702baae5 ("KVM: PPC: Book3S HV: Handle pending exceptions on guest entry with MSR_EE")
> > Cc: stable@vger.kernel.org # 6.8+
> > Reported-by: Timothy Pearson <tpearson@raptorengineering.com>
> > Closes: https://lore.kernel.org/linuxppc-dev/582904882.11159.1786719390349.JavaMail.zimbra@raptorengineeringinc.com
> > Signed-off-by: Gautam Menghani <gautam@linux.ibm.com>
> > ---
> > v3:
> > 1. Continue the use of LPCR_MER when kernel-irqchip=off (Sashiko)
> >
> > v2:
> > 1. Handle the case where xive_interrupt_pending() is true and also the
> > external exception bit is set. (Narayana)
> >
> > arch/powerpc/include/asm/kvm_ppc.h | 7 +++++++
> > arch/powerpc/kvm/book3s_hv.c | 2 +-
> > 2 files changed, 8 insertions(+), 1 deletion(-)
> >
> > diff --git a/arch/powerpc/include/asm/kvm_ppc.h b/arch/powerpc/include/asm/kvm_ppc.h
> > index 169ea6a7fbad..580ad2548c2b 100644
> > --- a/arch/powerpc/include/asm/kvm_ppc.h
> > +++ b/arch/powerpc/include/asm/kvm_ppc.h
> > @@ -747,6 +747,11 @@ static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> > return vcpu->arch.irq_type == KVMPPC_IRQ_XIVE;
> > }
> >
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm)
> > +{
> > + return kvm->arch.xive_devices.native;
> > +}
> > +
> > extern int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> > struct kvm_vcpu *vcpu, u32 cpu);
> > extern void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu);
> > @@ -782,6 +787,8 @@ static inline bool kvmppc_xive_rearm_escalation(struct kvm_vcpu *vcpu) { return
> >
> > static inline int kvmppc_xive_enabled(struct kvm_vcpu *vcpu)
> > { return 0; }
> > +static inline bool kvmppc_xive_native_enabled(struct kvm *kvm) { return false; }
> > +
> > static inline int kvmppc_xive_native_connect_vcpu(struct kvm_device *dev,
> > struct kvm_vcpu *vcpu, u32 cpu) { return -EBUSY; }
> > static inline void kvmppc_xive_native_cleanup_vcpu(struct kvm_vcpu *vcpu) { }
> > diff --git a/arch/powerpc/kvm/book3s_hv.c b/arch/powerpc/kvm/book3s_hv.c
> > index dbac3573b2c8..cb2bb29451a7 100644
> > --- a/arch/powerpc/kvm/book3s_hv.c
> > +++ b/arch/powerpc/kvm/book3s_hv.c
> > @@ -4980,7 +4980,7 @@ int kvmhv_run_single_vcpu(struct kvm_vcpu *vcpu, u64 time_limit,
> > if (!kvmhv_on_pseries() && (__kvmppc_get_msr_hv(vcpu) & MSR_EE))
> > kvmppc_inject_interrupt_hv(vcpu,
> > BOOK3S_INTERRUPT_EXTERNAL, 0);
> > - else
> > + else if (!kvmppc_xive_native_enabled(vcpu->kvm))
> > lpcr |= LPCR_MER;
>
> 1. What happens when an L1 KVM guest is booted with (native) XIVE and then the
> guest is rebooted with `xive=off` i.e., with XICS? Will we stop setting
> LPCR_MER as kvm->arch.xive_devices.native would still be set?
Yes, that's a good catch. kvm->arch.xive_devices.native is cleared only
when vm is destroyed. So this case has to be handled.
> 2. What happends when an L1 KVM guest is booted with XIVE and then we kexec into
> a new kernel with `xive=off`? You did mention in your other reply that with
> kexec, `xive=off` is ignored currently but IMO, we would want to understand
> where this limiation lies and fix that if need be. I do understand we are
> trying to fix spurious interrupts problem with XIVE in this patch and this
> particular problem can be taken separately but it worth investigating.
Yes right, the kexec case is to be understood, but I'll take it up separately.
Meanwhile, I'll send a v4 to fix the reboot case.
>
> Thanks,
> Amit
^ permalink raw reply [flat|nested] 9+ messages in thread