* [PATCH v3 1/2] KVM: arm64: Fix spurious warning for benign stage 2 teardown race
2026-09-01 17:28 [PATCH v3 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race Lorenzo Stoakes (ARM)
@ 2026-09-01 17:28 ` Lorenzo Stoakes (ARM)
2026-09-01 17:46 ` sashiko-bot
2026-09-01 17:29 ` [PATCH v3 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, " Lorenzo Stoakes (ARM)
2026-09-15 22:25 ` [PATCH v3 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 " Oliver Upton
2 siblings, 1 reply; 8+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-01 17:28 UTC (permalink / raw)
To: Marc Zyngier, Oliver Upton, Joey Gouly, Steffen Eiden,
Suzuki K Poulose, Zenghui Yu, Catalin Marinas, Will Deacon,
Christoffer Dall, Fuad Tabba
Cc: Wei-Lin Chang, Yao Yuan, linux-arm-kernel, kvmarm, linux-kernel,
Lorenzo Stoakes (ARM), stable
kvmtool was used to establish an L1 guest with 8 CPUs and 8 GiB of RAM, an
L2 guest with 4 CPUs and 4 GiB of RAM and an L3 guest with 2 CPUs and 2 GiB
of RAM, all of which was then exited.
Under memory pressure in the L0 host warnings were observed due to
migration triggered by compaction:
WARNING: arch/arm64/kvm/mmu.c:336 at __unmap_stage2_range+0x64/0x80,
CPU#5: kcompactd0/66
Which was, in turn, triggered by an MMU notifier for the host invalidation:
mmu_notifier_invalidate_range_start()
-> ... -> kvm_mmu_notifier_invalidate_range_start()
-> kvm_mmu_unmap_gfn_range()
-> kvm_unmap_gfn_range()
-> kvm_nested_s2_unmap()
-> kvm_stage2_unmap_range()
-> __unmap_stage2_range()
-> stage2_apply_range()
<- -EINVAL, triggering a WARN_ON()
Racing with L0's teardown of stage 2 page tables:
exit_mm()
-> mmput()
-> __mmput()
-> exit_mmap()
-> mmu_notifier_release()
-> ... -> kvm_mmu_notifier_release()
-> kvm_flush_shadow_all()
-> kvm_arch_flush_shadow_all()
-> kvm_free_stage2_pgd()
-> [ acquire kvm->mmu_lock for write ]
-> mmu->pgt = NULL [ among other tasks ]
-> [ release kvm->mmu_lock for write ]
It turns out there is a benign race resulting in a spurious warning:
Thread A - notify: migration | Thread B - notify: release
-------------------------------|---------------------------------
< kvm->mmu_lock held > |
stage2_apply_range() |
get mmu->pgt, check !NULL |
... | kvm_arch_flush_shadow_all()
cond_resched_rwlock_write(); | < contend, sleep kvm->mmu_lock >
< drop kvm->mmu_lock > | < acquire kvm->mmu_lock>
| ...
| kvm_free_stage2_pgd()
| mmu->pgt = NULL
| < invalidate MMU >
| ...
| < release kvm->mmu_lock >
[ scheduled ] |
stage2_apply_range() |
< loop to next > |
get, mmu->pgt, check !NULL |
is NULL, return -EINVAL |
__unmap_stage2_range() |
WARN_ON(-EINVAL) <--- entirely spurious - the race was handled
correctly.
Fix the spurious warning by updating stage2_apply_range() to no longer
treat concurrent PGT teardown on lock release as an error - whether the
walker is tearing down page tables or doing something else this is a
legitimate reason to abort the operation without error.
This keeps the warning in place for all other circumstances.
In practice only __unmap_stage2_range() actually does anything with the
error so this only impacts that.
Fixes: ec14c272408a ("KVM: arm64: nv: Unmap/flush shadow stage 2 page tables")
Cc: stable@vger.kernel.org
Reviewed-by: Yuan Yao <yaoyuan@linux.alibaba.com>
Reviewed-by: Marc Zyngier <maz@kernel.org>
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
---
arch/arm64/kvm/mmu.c | 15 ++++++++++++---
1 file changed, 12 insertions(+), 3 deletions(-)
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index 9ba86450fe4a..2d44cd6a5aed 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -59,27 +59,36 @@ static phys_addr_t stage2_range_addr_end(phys_addr_t addr, phys_addr_t end)
* long will also starve other vCPUs. We have to also make sure that the page
* tables are not freed while we released the lock.
*/
-static int stage2_apply_range(struct kvm_s2_mmu *mmu, phys_addr_t addr,
+static int stage2_apply_range(struct kvm_s2_mmu *mmu, phys_addr_t start,
phys_addr_t end,
int (*fn)(struct kvm_pgtable *, u64, u64),
bool resched)
{
struct kvm *kvm = kvm_s2_mmu_to_kvm(mmu);
+ bool lock_dropped = false;
+ phys_addr_t addr = start;
int ret;
u64 next;
do {
struct kvm_pgtable *pgt = mmu->pgt;
+ /*
+ * We may be raced on PGT teardown when we release the
+ * kvm->mmu_lock. That's fine as the PGT is legitimately no
+ * longer present.
+ */
if (!pgt)
- return -EINVAL;
+ return lock_dropped ? 0 : -EINVAL;
next = stage2_range_addr_end(addr, end);
ret = fn(pgt, addr, next - addr);
if (ret)
break;
- if (resched && next != end)
+ if (resched && next != end) {
cond_resched_rwlock_write(&kvm->mmu_lock);
+ lock_dropped = true;
+ }
} while (addr = next, addr != end);
return ret;
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* Re: [PATCH v3 1/2] KVM: arm64: Fix spurious warning for benign stage 2 teardown race
2026-09-01 17:28 ` [PATCH v3 1/2] KVM: arm64: Fix spurious warning for benign stage 2 " Lorenzo Stoakes (ARM)
@ 2026-09-01 17:46 ` sashiko-bot
0 siblings, 0 replies; 8+ messages in thread
From: sashiko-bot @ 2026-09-01 17:46 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM); +Cc: kvmarm, Oliver Upton, Marc Zyngier
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
Pre-existing issues:
- [Critical] Use-After-Free of nested_mmus array in stage2_apply_range due to concurrent kvfree
--
commit 2cabe3fbde6089df082f059c9d4b617266797a8f
Author: Lorenzo Stoakes (ARM) <ljs@kernel.org>
KVM: arm64: Fix spurious warning for benign stage 2 teardown race
This commit addresses a spurious warning observed during stage 2 page table
teardown by preventing concurrent PGT teardown from being treated as an error
when the mmu_lock has been dropped and re-acquired.
> diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
> index 9ba86450fe4af..2d44cd6a5aed9 100644
> --- a/arch/arm64/kvm/mmu.c
> +++ b/arch/arm64/kvm/mmu.c
[ ... ]
> @@ -59,27 +59,36 @@ static phys_addr_t stage2_range_addr_end(phys_addr_t addr, phys_addr_t end)
> -static int stage2_apply_range(struct kvm_s2_mmu *mmu, phys_addr_t addr,
> +static int stage2_apply_range(struct kvm_s2_mmu *mmu, phys_addr_t start,
> phys_addr_t end,
> int (*fn)(struct kvm_pgtable *, u64, u64),
> bool resched)
> {
> struct kvm *kvm = kvm_s2_mmu_to_kvm(mmu);
> + bool lock_dropped = false;
> + phys_addr_t addr = start;
> int ret;
> u64 next;
>
> do {
> struct kvm_pgtable *pgt = mmu->pgt;
[Severity: Critical]
This isn't a bug introduced by this patch, but while this change correctly
avoids the warning when pgt is NULL, does the lock-dropping behavior here
expose a Use-After-Free on the mmu pointer itself?
Looking at the teardown paths, it appears stage2_apply_range() can race with
concurrent freeing of the kvm->arch.nested_mmus array.
If Thread A (e.g., kcompactd) is iterating over the array in
kvm_nested_s2_unmap() and yields kvm->mmu_lock inside this loop via
cond_resched_rwlock_write(), Thread B (such as exit_mmap ->
mmu_notifier_release) could concurrently call kvm_arch_flush_shadow_all().
Thread B acquires mmu_lock to clear the PGDs, and then locklessly calls
kvfree(kvm->arch.nested_mmus) after dropping the lock.
When Thread A wakes up and re-acquires mmu_lock to resume the loop, its
local mmu pointer would point into the freed nested_mmus array, leading to a
Use-After-Free right here when it accesses mmu->pgt.
A similar concurrent kvfree() appears to exist in kvm_vcpu_init_nested(),
which reallocates and frees the array (kvfree(tmp)) while readers might be
yielding here.
Is there a need to defer these kvfree() calls or better protect the yielding
readers in stage2_apply_range() to prevent this race?
> + /*
> + * We may be raced on PGT teardown when we release the
> + * kvm->mmu_lock. That's fine as the PGT is legitimately no
> + * longer present.
> + */
> if (!pgt)
> - return -EINVAL;
> + return lock_dropped ? 0 : -EINVAL;
>
> next = stage2_range_addr_end(addr, end);
> ret = fn(pgt, addr, next - addr);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260901-kvm-arm-nested-virt-fix-v3-0-b154676f7e4c@kernel.org?part=1
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH v3 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race
2026-09-01 17:28 [PATCH v3 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race Lorenzo Stoakes (ARM)
2026-09-01 17:28 ` [PATCH v3 1/2] KVM: arm64: Fix spurious warning for benign stage 2 " Lorenzo Stoakes (ARM)
@ 2026-09-01 17:29 ` Lorenzo Stoakes (ARM)
2026-09-01 17:46 ` sashiko-bot
2026-09-11 8:12 ` Jonathan Davies
2026-09-15 22:25 ` [PATCH v3 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 " Oliver Upton
2 siblings, 2 replies; 8+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-01 17:29 UTC (permalink / raw)
To: Marc Zyngier, Oliver Upton, Joey Gouly, Steffen Eiden,
Suzuki K Poulose, Zenghui Yu, Catalin Marinas, Will Deacon,
Christoffer Dall, Fuad Tabba
Cc: Wei-Lin Chang, Yao Yuan, linux-arm-kernel, kvmarm, linux-kernel,
Lorenzo Stoakes (ARM), stable
Commit 7270cc9157f4 ("KVM: arm64: nv: Handle VNCR_EL2 invalidation from MMU
notifiers") introduced VNCR_EL2 invalidation in both kvm_nested_s2_unmap()
and kvm_nested_s2_wp().
However at the point of this being performed concurrent stage 2 teardown of
a nested guest can cause kvm->arch.mmu.pgt to be set to NULL.
This happens in kvm_flush_shadow_all() -> kvm_arch_flush_shadow_all() ->
kvm_free_stage2_pgd() and is performed under the kvm->mmu_lock.
Commit ec14c272408a ("KVM: arm64: nv: Unmap/flush shadow stage 2 page
tables") introduced the teardown of the entire nested MMU range, which then
invokes stage2_apply_range() with resched=true:
mmu_notifier_invalidate_range_start()
-> ... -> kvm_mmu_notifier_invalidate_range_start()
-> kvm_mmu_unmap_gfn_range()
-> kvm_unmap_gfn_range()
-> kvm_nested_s2_unmap()
-> kvm_stage2_unmap_range()
-> __unmap_stage2_range()
-> stage2_apply_range()
This means that stage2_apply_range() can drop the kvm->mmu_lock and thus
concurrent progress can be made in lockstep with
kvm_arch_flush_shadow_all().
If kvm_arch_flush_shadow_all() advances ahead of stage2_apply_range() and
completes its operation it guarantees a NULL pointer deref.
Since kvm_free_stage2_pgd() is performed under the kvm->mmu_lock this will
either be observed NULL or not and serialised against
kvm_free_stage2_pgd().
Resolve the issue by abstracting the invalidation to a new function,
kvm_invalidate_vncr_ipa_all(), and check that the pgt is non-NULL before
dereferencing it.
Fixes: 7270cc9157f4 ("KVM: arm64: nv: Handle VNCR_EL2 invalidation from MMU notifiers")
Cc: stable@vger.kernel.org
Reviewed-by: Marc Zyngier <maz@kernel.org>
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
---
arch/arm64/kvm/nested.c | 15 +++++++++++++--
1 file changed, 13 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c
index 17123f0b6dab..f69722e1592a 100644
--- a/arch/arm64/kvm/nested.c
+++ b/arch/arm64/kvm/nested.c
@@ -1260,6 +1260,17 @@ void kvm_handle_s1e2_tlbi(struct kvm_vcpu *vcpu, u32 inst, u64 val)
invalidate_vncr_va(vcpu->kvm, &scope);
}
+static void kvm_invalidate_vncr_ipa_all(struct kvm *kvm)
+{
+ struct kvm_pgtable *pgt = kvm->arch.mmu.pgt;
+
+ lockdep_assert_held_write(&kvm->mmu_lock);
+
+ /* if the mmu lock was dropped, pgt teardown may have raced. */
+ if (pgt)
+ kvm_invalidate_vncr_ipa(kvm, 0, BIT(pgt->ia_bits));
+}
+
void kvm_nested_s2_wp(struct kvm *kvm)
{
int i;
@@ -1276,7 +1287,7 @@ void kvm_nested_s2_wp(struct kvm *kvm)
kvm_stage2_wp_range(mmu, 0, kvm_phys_size(mmu));
}
- kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits));
+ kvm_invalidate_vncr_ipa_all(kvm);
}
void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block)
@@ -1295,7 +1306,7 @@ void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block)
kvm_stage2_unmap_range(mmu, 0, kvm_phys_size(mmu), may_block);
}
- kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits));
+ kvm_invalidate_vncr_ipa_all(kvm);
}
void kvm_nested_s2_flush(struct kvm *kvm)
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* Re: [PATCH v3 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race
2026-09-01 17:29 ` [PATCH v3 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, " Lorenzo Stoakes (ARM)
@ 2026-09-01 17:46 ` sashiko-bot
2026-09-11 8:12 ` Jonathan Davies
1 sibling, 0 replies; 8+ messages in thread
From: sashiko-bot @ 2026-09-01 17:46 UTC (permalink / raw)
To: Lorenzo Stoakes (ARM); +Cc: Marc Zyngier, kvmarm, Oliver Upton
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
Pre-existing issues:
- [High] Concurrent VM teardown frees the `kvm->arch.nested_mmus` array without holding `kvm->mmu_lock`, causing a Use-After-Free when `kvm_nested_s2_unmap`, `kvm_nested_s2_wp`, or `kvm_nested_s2_flush` iterate over the array.
--
commit 6d817ae312afc7a3c0a4cae55f93ea78c14d26e1
Author: Lorenzo Stoakes (ARM) <ljs@kernel.org>
KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race
This commit fixes a NULL pointer dereference caused by concurrent stage 2
teardown racing with MMU notifiers by wrapping VNCR IPA invalidation in a
new helper that checks for a NULL page table pointer.
> diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c
> index 17123f0b6daba..f69722e1592a5 100644
> --- a/arch/arm64/kvm/nested.c
> +++ b/arch/arm64/kvm/nested.c
[ ... ]
> @@ -1295,7 +1306,7 @@ void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block)
> kvm_stage2_unmap_range(mmu, 0, kvm_phys_size(mmu), may_block);
[Severity: High]
This is a pre-existing issue, but since this patch explicitly deals with the
race between kvm_arch_flush_shadow_all() and MMU notifiers dropping the
mmu_lock, is there a Use-After-Free risk with kvm->arch.nested_mmus here?
If kvm_stage2_unmap_range() ends up calling stage2_apply_range() which
drops kvm->mmu_lock, kvm_arch_flush_shadow_all() can proceed to tear
down the nested MMUs.
In arch/arm64/kvm/nested.c:kvm_arch_flush_shadow_all():
kvfree(kvm->arch.nested_mmus);
kvm->arch.nested_mmus = NULL;
kvm->arch.nested_mmus_size = 0;
This kvfree() happens without holding kvm->mmu_lock.
When the MMU notifier path in kvm_nested_s2_unmap() reacquires the lock
and continues, the mmu pointer (which points into the now-freed
kvm->arch.nested_mmus array) is completely dangling.
For example, in arch/arm64/kvm/mmu.c:stage2_apply_range():
do {
struct kvm_pgtable *pgt = mmu->pgt;
...
Can this dereference of mmu->pgt cause memory corruption or crashes if the
underlying array has been freed by a concurrent VM teardown?
> }
>
> - kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits));
> + kvm_invalidate_vncr_ipa_all(kvm);
> }
>
> void kvm_nested_s2_flush(struct kvm *kvm)
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260901-kvm-arm-nested-virt-fix-v3-0-b154676f7e4c@kernel.org?part=2
^ permalink raw reply [flat|nested] 8+ messages in thread* Re: [PATCH v3 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race
2026-09-01 17:29 ` [PATCH v3 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, " Lorenzo Stoakes (ARM)
2026-09-01 17:46 ` sashiko-bot
@ 2026-09-11 8:12 ` Jonathan Davies
2026-09-11 8:44 ` Lorenzo Stoakes (ARM)
1 sibling, 1 reply; 8+ messages in thread
From: Jonathan Davies @ 2026-09-11 8:12 UTC (permalink / raw)
To: ljs
Cc: catalin.marinas, christoffer.dall, fuad.tabba, joey.gouly, kvmarm,
linux-arm-kernel, linux-kernel, maz, oupton, seiden, stable,
suzuki.poulose, weilin.chang, will, yaoyuan, yuzenghui
On Tue, 01 Sep 2026, "Lorenzo Stoakes (ARM)" <ljs@kernel.org> wrote:
> This means that stage2_apply_range() can drop the kvm->mmu_lock and
> thus concurrent progress can be made in lockstep with
> kvm_arch_flush_shadow_all().
>
> If kvm_arch_flush_shadow_all() advances ahead of stage2_apply_range()
> and completes its operation it guarantees a NULL pointer deref.
I've reproducibly hit exactly this problem when running nested guests
with concurrent memory compaction on the L0 host. This patch fixes it
for me perfectly (applied to a 6.18 kernel).
So, for what it's worth:
Tested-by: Jonathan Davies <jonathan.davies@nutanix.com>
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v3 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race
2026-09-11 8:12 ` Jonathan Davies
@ 2026-09-11 8:44 ` Lorenzo Stoakes (ARM)
0 siblings, 0 replies; 8+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-11 8:44 UTC (permalink / raw)
To: Jonathan Davies
Cc: catalin.marinas, christoffer.dall, fuad.tabba, joey.gouly, kvmarm,
linux-arm-kernel, linux-kernel, maz, oupton, seiden, stable,
suzuki.poulose, weilin.chang, will, yaoyuan, yuzenghui
On Fri, Sep 11, 2026 at 09:12:20AM +0100, Jonathan Davies wrote:
> On Tue, 01 Sep 2026, "Lorenzo Stoakes (ARM)" <ljs@kernel.org> wrote:
> > This means that stage2_apply_range() can drop the kvm->mmu_lock and
> > thus concurrent progress can be made in lockstep with
> > kvm_arch_flush_shadow_all().
> >
> > If kvm_arch_flush_shadow_all() advances ahead of stage2_apply_range()
> > and completes its operation it guarantees a NULL pointer deref.
>
> I've reproducibly hit exactly this problem when running nested guests
> with concurrent memory compaction on the L0 host. This patch fixes it
> for me perfectly (applied to a 6.18 kernel).
>
> So, for what it's worth:
>
> Tested-by: Jonathan Davies <jonathan.davies@nutanix.com>
Amazing, thanks :)
--
Cheers, Lorenzo
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v3 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race
2026-09-01 17:28 [PATCH v3 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race Lorenzo Stoakes (ARM)
2026-09-01 17:28 ` [PATCH v3 1/2] KVM: arm64: Fix spurious warning for benign stage 2 " Lorenzo Stoakes (ARM)
2026-09-01 17:29 ` [PATCH v3 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, " Lorenzo Stoakes (ARM)
@ 2026-09-15 22:25 ` Oliver Upton
2 siblings, 0 replies; 8+ messages in thread
From: Oliver Upton @ 2026-09-15 22:25 UTC (permalink / raw)
To: Marc Zyngier, Joey Gouly, Steffen Eiden, Suzuki K Poulose,
Zenghui Yu, Catalin Marinas, Will Deacon, Christoffer Dall,
Fuad Tabba, Lorenzo Stoakes (ARM)
Cc: Oliver Upton, Wei-Lin Chang, Yao Yuan, linux-arm-kernel, kvmarm,
linux-kernel, stable
On Tue, 01 Sep 2026 18:28:58 +0100, Lorenzo Stoakes (ARM) wrote:
> When GFNs are invalidated in L0 an MMU notifier triggers
> kvm_unmap_gfn_range() which tears down all of the stage 2 shadow page
> tables for nested guests via kvm_nested_s2_unmap().
>
> To avoid lockup, the kvm->mmu_lock is dropped while doing this and the task
> rescheduled once for each block of physical address space (32 MiB for 16
> KiB page size), with the lock being reacquired once the task is scheduled
> again.
>
> [...]
Applied to fixes, thanks!
[1/2] KVM: arm64: Fix spurious warning for benign stage 2 teardown race
https://git.kernel.org/kvmarm/kvmarm/c/38b70fc453c3
[2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race
https://git.kernel.org/kvmarm/kvmarm/c/4c74e233cded
--
Best,
Oliver
^ permalink raw reply [flat|nested] 8+ messages in thread