From: "Lorenzo Stoakes (ARM)" <ljs@kernel.org>
To: Marc Zyngier <maz@kernel.org>, Oliver Upton <oupton@kernel.org>,
Fuad Tabba <tabba@google.com>, Joey Gouly <joey.gouly@arm.com>,
Steffen Eiden <seiden@linux.ibm.com>,
Suzuki K Poulose <suzuki.poulose@arm.com>,
Zenghui Yu <yuzenghui@huawei.com>,
Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>,
Christoffer Dall <christoffer.dall@arm.com>
Cc: Wei-Lin Chang <weilin.chang@arm.com>,
Yao Yuan <yaoyuan@linux.alibaba.com>,
linux-arm-kernel@lists.infradead.org, kvmarm@lists.linux.dev,
linux-kernel@vger.kernel.org,
"Lorenzo Stoakes (ARM)" <ljs@kernel.org>,
stable@vger.kernel.org
Subject: [PATCH v2 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race
Date: Sat, 22 Aug 2026 18:46:54 +0100 [thread overview]
Message-ID: <20260822-kvm-arm-nested-virt-fix-v2-2-ac4059a0eaa6@kernel.org> (raw)
In-Reply-To: <20260822-kvm-arm-nested-virt-fix-v2-0-ac4059a0eaa6@kernel.org>
Commit 7270cc9157f4 ("KVM: arm64: nv: Handle VNCR_EL2 invalidation from MMU
notifiers") introduced VNCR_EL2 invalidation in both kvm_nested_s2_unmap()
and kvm_nested_s2_wp().
However at the point of this being performed concurrent stage 2 teardown of
a nested guest can cause kvm->arch.mmu.pgt to be set to NULL.
This happens in kvm_flush_shadow_all() -> kvm_arch_flush_shadow_all() ->
kvm_free_stage2_pgd() and is performed under the kvm->mmu_lock.
Commit ec14c272408a ("KVM: arm64: nv: Unmap/flush shadow stage 2 page
tables") introduced the teardown of the entire nested MMU range, which then
invokes stage2_apply_range() with resched=true:
mmu_notifier_invalidate_range_start()
-> ... -> kvm_mmu_notifier_invalidate_range_start()
-> kvm_mmu_unmap_gfn_range()
-> kvm_unmap_gfn_range()
-> kvm_nested_s2_unmap()
-> kvm_stage2_unmap_range()
-> __unmap_stage2_range()
-> stage2_apply_range()
This means that stage2_apply_range() can drop the kvm->mmu_lock and thus
concurrent progress can be made in lockstep with
kvm_arch_flush_shadow_all().
If kvm_arch_flush_shadow_all() advances ahead of stage2_apply_range() and
completes its operation it guarantees a NULL pointer deref.
Since kvm_free_stage2_pgd() is performed under the kvm->mmu_lock this will
either be observed NULL or not and serialised against
kvm_free_stage2_pgd().
Resolve the issue by abstracting the invalidation to a new function,
kvm_invalidate_vncr_ipa_all(), and check that the pgt is non-NULL before
dereferencing it.
Fixes: 7270cc9157f4 ("KVM: arm64: nv: Handle VNCR_EL2 invalidation from MMU notifiers")
Cc: stable@vger.kernel.org
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
---
arch/arm64/kvm/nested.c | 15 +++++++++++++--
1 file changed, 13 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c
index 17123f0b6dab..f69722e1592a 100644
--- a/arch/arm64/kvm/nested.c
+++ b/arch/arm64/kvm/nested.c
@@ -1260,6 +1260,17 @@ void kvm_handle_s1e2_tlbi(struct kvm_vcpu *vcpu, u32 inst, u64 val)
invalidate_vncr_va(vcpu->kvm, &scope);
}
+static void kvm_invalidate_vncr_ipa_all(struct kvm *kvm)
+{
+ struct kvm_pgtable *pgt = kvm->arch.mmu.pgt;
+
+ lockdep_assert_held_write(&kvm->mmu_lock);
+
+ /* if the mmu lock was dropped, pgt teardown may have raced. */
+ if (pgt)
+ kvm_invalidate_vncr_ipa(kvm, 0, BIT(pgt->ia_bits));
+}
+
void kvm_nested_s2_wp(struct kvm *kvm)
{
int i;
@@ -1276,7 +1287,7 @@ void kvm_nested_s2_wp(struct kvm *kvm)
kvm_stage2_wp_range(mmu, 0, kvm_phys_size(mmu));
}
- kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits));
+ kvm_invalidate_vncr_ipa_all(kvm);
}
void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block)
@@ -1295,7 +1306,7 @@ void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block)
kvm_stage2_unmap_range(mmu, 0, kvm_phys_size(mmu), may_block);
}
- kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits));
+ kvm_invalidate_vncr_ipa_all(kvm);
}
void kvm_nested_s2_flush(struct kvm *kvm)
--
2.55.0
next prev parent reply other threads:[~2026-08-22 17:47 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-22 17:46 [PATCH v2 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 teardown race Lorenzo Stoakes (ARM)
2026-08-22 17:46 ` [PATCH v2 1/2] KVM: arm64: Fix spurious warning for benign stage 2 " Lorenzo Stoakes (ARM)
2026-08-22 18:02 ` sashiko-bot
2026-08-23 5:38 ` Yao Yuan
2026-08-24 13:29 ` Lorenzo Stoakes (ARM)
2026-08-22 17:46 ` Lorenzo Stoakes (ARM) [this message]
2026-08-22 18:05 ` [PATCH v2 2/2] KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, " sashiko-bot
2026-08-23 7:53 ` [PATCH v2 0/2] KVM: arm64: Fix spurious warn, null ptr deref on S2 " Marc Zyngier
2026-08-24 13:15 ` Lorenzo Stoakes (ARM)
2026-08-24 15:56 ` Marc Zyngier
2026-08-24 17:34 ` Lorenzo Stoakes (ARM)
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260822-kvm-arm-nested-virt-fix-v2-2-ac4059a0eaa6@kernel.org \
--to=ljs@kernel.org \
--cc=catalin.marinas@arm.com \
--cc=christoffer.dall@arm.com \
--cc=joey.gouly@arm.com \
--cc=kvmarm@lists.linux.dev \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=maz@kernel.org \
--cc=oupton@kernel.org \
--cc=seiden@linux.ibm.com \
--cc=stable@vger.kernel.org \
--cc=suzuki.poulose@arm.com \
--cc=tabba@google.com \
--cc=weilin.chang@arm.com \
--cc=will@kernel.org \
--cc=yaoyuan@linux.alibaba.com \
--cc=yuzenghui@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.