All of lore.kernel.org
 help / color / mirror / Atom feed
From: Mushahid Hussain <hmushi@amazon.co.uk>
To: Sean Christopherson <seanjc@google.com>,
	Paolo Bonzini <pbonzini@redhat.com>
Cc: David Hildenbrand <david@kernel.org>, <kvm@vger.kernel.org>,
	<linux-kernel@vger.kernel.org>, <nh-open-source@amazon.com>
Subject: [PATCH] KVM: Use kvcalloc() to allocate lpage_info arrays and dirty bitmaps
Date: Sat, 15 Aug 2026 14:22:18 +0000	[thread overview]
Message-ID: <20260815142218.85067-1-hmushi@amazon.co.uk> (raw)

Use kvcalloc() instead of __vcalloc() for the lpage_info arrays and
memslot dirty bitmaps, and switch their frees from vfree() back to
kvfree(). kvcalloc() serves sub-page requests from the slab and
falls back to vmalloc for anything larger; kvfree() handles either
allocator. The vfree() pairing came from commit a952d608f0be ("KVM:
Use vfree for memory allocated by vcalloc()/__vcalloc()").

Commit 37b2a6510a48 ("KVM: use __vcalloc for very large allocations")
moved these sites onto __vcalloc() to escape kvmalloc()'s INT_MAX
size cap, since memslot metadata sizes scale with the memslot size
and can exceed 2 GiB. Neither site can reach the cap: lpage_info
holds one 4 byte entry per hugepage granule (per 2 MiB at level 2,
per 1 GiB at level 3), so reaching INT_MAX would take a 1 PiB
memslot, while userspace slots are capped at 8 TiB by
KVM_MEM_MAX_NR_PAGES and x86's internal slots at 4 GiB by
__x86_set_memory_region()'s u32 size. The dirty bitmap is two
bitmaps at one bit per page, npages/4 bytes, at most 512 MiB, and
internal slots never have one because kvm_set_internal_memslot()
rejects any flags.

__vcalloc() makes every allocation at least a page, so a single page
memslot consumes 8 KiB of vmalloc for 8 bytes of lpage_info and
another 4 KiB for a 16 byte dirty bitmap when dirty logging is
enabled. This overhead scales with the number of slots and VMs on a
host, adding up to memory pressure when guest address spaces are
fragmented into small slots.

The rmap and gfn_write_track arrays keep __vcalloc() and vfree():
the 4K rmap and gfn_write_track are per-page arrays, 8 and 2 bytes
per 4 KiB page, which legitimately cross INT_MAX below the 8 TiB
slot ceiling; the smaller higher-level rmaps share the 4K rmap's
allocation loop; and none of them allocate under the TDP MMU, where
the waste above was observed.

Fixes: 37b2a6510a48 ("KVM: use __vcalloc for very large allocations")
Assisted-by: Kiro:claude-fable-5
Signed-off-by: Mushahid Hussain <hmushi@amazon.co.uk>
---
 arch/x86/kvm/x86.c  | 6 +++---
 virt/kvm/kvm_main.c | 4 ++--
 2 files changed, 5 insertions(+), 5 deletions(-)

diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c
index afcac1042947..77567aa20d83 100644
--- a/arch/x86/kvm/x86.c
+++ b/arch/x86/kvm/x86.c
@@ -13488,7 +13488,7 @@ void kvm_arch_free_memslot(struct kvm *kvm, struct kvm_memory_slot *slot)
 	memslot_rmap_free(slot);
 
 	for (i = 1; i < KVM_NR_PAGE_SIZES; ++i) {
-		vfree(slot->arch.lpage_info[i - 1]);
+		kvfree(slot->arch.lpage_info[i - 1]);
 		slot->arch.lpage_info[i - 1] = NULL;
 	}
 
@@ -13544,7 +13544,7 @@ static int kvm_alloc_memslot_metadata(struct kvm *kvm,
 
 		lpages = __kvm_mmu_slot_lpages(slot, npages, level);
 
-		linfo = __vcalloc(lpages, sizeof(*linfo), GFP_KERNEL_ACCOUNT);
+		linfo = kvcalloc(lpages, sizeof(*linfo), GFP_KERNEL_ACCOUNT);
 		if (!linfo)
 			goto out_free;
 
@@ -13580,7 +13580,7 @@ static int kvm_alloc_memslot_metadata(struct kvm *kvm,
 	memslot_rmap_free(slot);
 
 	for (i = 1; i < KVM_NR_PAGE_SIZES; ++i) {
-		vfree(slot->arch.lpage_info[i - 1]);
+		kvfree(slot->arch.lpage_info[i - 1]);
 		slot->arch.lpage_info[i - 1] = NULL;
 	}
 	return -ENOMEM;
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index e44c20c04961..52af1ebdd14c 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -927,7 +927,7 @@ static void kvm_destroy_dirty_bitmap(struct kvm_memory_slot *memslot)
 	if (!memslot->dirty_bitmap)
 		return;
 
-	vfree(memslot->dirty_bitmap);
+	kvfree(memslot->dirty_bitmap);
 	memslot->dirty_bitmap = NULL;
 }
 
@@ -1422,7 +1422,7 @@ static int kvm_alloc_dirty_bitmap(struct kvm_memory_slot *memslot)
 {
 	unsigned long dirty_bytes = kvm_dirty_bitmap_bytes(memslot);
 
-	memslot->dirty_bitmap = __vcalloc(2, dirty_bytes, GFP_KERNEL_ACCOUNT);
+	memslot->dirty_bitmap = kvcalloc(2, dirty_bytes, GFP_KERNEL_ACCOUNT);
 	if (!memslot->dirty_bitmap)
 		return -ENOMEM;
 

base-commit: 8cdeaa50eae8dad34885515f62559ee83e7e8dda
-- 
2.47.3


             reply	other threads:[~2026-08-15 14:22 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-15 14:22 Mushahid Hussain [this message]
2026-09-28 23:16 ` [PATCH] KVM: Use kvcalloc() to allocate lpage_info arrays and dirty bitmaps Sean Christopherson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260815142218.85067-1-hmushi@amazon.co.uk \
    --to=hmushi@amazon.co.uk \
    --cc=david@kernel.org \
    --cc=kvm@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=nh-open-source@amazon.com \
    --cc=pbonzini@redhat.com \
    --cc=seanjc@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.