Kernel KVM virtualization development
 help / color / mirror / Atom feed
From: Sean Christopherson <seanjc@google.com>
To: Yan Zhao <yan.y.zhao@intel.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>,
	pbonzini@redhat.com,  Andrew Morton <akpm@linux-foundation.org>,
	David Hildenbrand <david@kernel.org>, Zi Yan <ziy@nvidia.com>,
	 Matthew Brost <matthew.brost@intel.com>,
	Joshua Hahn <joshua.hahnjy@gmail.com>,
	 Rakie Kim <rakie.kim@sk.com>, Byungchul Park <byungchul@sk.com>,
	Gregory Price <gourry@gourry.net>,
	 Ying Huang <ying.huang@linux.alibaba.com>,
	Alistair Popple <apopple@nvidia.com>,
	linux-mm@kvack.org,  linux-kernel@vger.kernel.org,
	Neha Gholkar <nehagholkar@gmail.com>,
	kvm@vger.kernel.org,  rick.p.edgecombe@intel.com,
	vishal.l.verma@intel.com
Subject: Re: [PATCH] mm: mempolicy: fix automatic numa balancing for shmem
Date: Thu, 23 Jul 2026 06:51:26 -0700	[thread overview]
Message-ID: <amIcXrS7nGX-adpG@google.com> (raw)
In-Reply-To: <amGpEEqxoLq/Y7ZO@yzhao56-desk.sh.intel.com>

On Thu, Jul 23, 2026, Yan Zhao wrote:
> On Mon, Jun 29, 2026 at 12:33:37PM -0400, Johannes Weiner wrote:
> > Neha reports that mapped shmem aren't considered for NUMA balancing,
> > noting convergence problems and bandwidth bottlenecking for cachelib
> > based workloads on tiered memory systems.
> > 
> > Looking at the code and going through the git history, this doesn't
> > actually seem intentional:
> > 
> > Commit fc3147245d19 ("mm: numa: Limit NUMA scanning to migrate-on-fault
> > VMAs") added a vma_policy_mof() gate to task_numa_work() so VMAs whose
> > policy lacks MPOL_F_MOF are skipped from NUMA balancing scans. The
> > motivation was a real usecase: Oracle was pinning shared segments with
> > mbind(MPOL_BIND) so trapping faults was both expensive and pointless.
> > 
> > The handling of NULL from vm_ops->get_policy, however, treated "user
> > explicitly opted out" the same as "user never specified anything." For
> > VMAs whose shared policy is absent - the common case for shmem - the
> > scan was disabled too.
> > 
> > This issue is old. It probably hurts less in conventional NUMA. But it's
> > very noticable on tiered systems, where entire tmpfs workingsets can get
> > stuck on lower-bandwidth memory.
> > 
> > Fix this by having vma_policy_mof() use __get_vma_policy() directly, and
> > thereby handle the fallback to task policy (-> preferred_node_policy()
> > has MPOL_F_MOF per default). Every other consumer of vm_ops->get_policy
> > already handles it this way, the scan-eligibility check was the outlier.
> > 
> > This preserves Mel's intended fix: don't scan stuff the user explicitly
> > pinned. But allow default policy vmas to participate in balancing.
> Hi,
> 
> This patch introduces a performance regression of a KVM stress test, which I
> addressed in the KVM selftest itself (see the analysis in the patch log).
> Could you share your thoughts on whether the userspace fix is the appropriate
> approach?

Yikes.  This could have meaningful "real world" impact on VMs backed with shmem,
not just on KVM's convoluted stress test.  NUMA balancing generally performs
poorly for VMs due to the higher costs of VM-Exits versus page faults, and due
to inefficiencies in the mmu_notifier interface (KVM does a full TLB shootdown
of the affected VM on every MMU_NOTIFY_PROTECTION_VMA event).

My stance is that using NUMA balancing with KVM guests is a terrible idea, and
that anyone that insists on using such a setup gets to suffer the consequences.
But in this case, IIUC, this change will "silently" enable NUMA balancing for
shmem-based KVM setups where it was previously disabled (albeit unintentionally).

I'm not fundamentally opposed to the change, but I do worry that downstream KVM
users could be in for a nasty surprise.

       reply	other threads:[~2026-07-23 13:51 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <20260629163337.1264881-1-hannes@cmpxchg.org>
     [not found] ` <amGpEEqxoLq/Y7ZO@yzhao56-desk.sh.intel.com>
2026-07-23 13:51   ` Sean Christopherson [this message]
2026-07-23 14:54     ` [PATCH] mm: mempolicy: fix automatic numa balancing for shmem Johannes Weiner
2026-07-23 15:58       ` Gregory Price

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amIcXrS7nGX-adpG@google.com \
    --to=seanjc@google.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=byungchul@sk.com \
    --cc=david@kernel.org \
    --cc=gourry@gourry.net \
    --cc=hannes@cmpxchg.org \
    --cc=joshua.hahnjy@gmail.com \
    --cc=kvm@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=matthew.brost@intel.com \
    --cc=nehagholkar@gmail.com \
    --cc=pbonzini@redhat.com \
    --cc=rakie.kim@sk.com \
    --cc=rick.p.edgecombe@intel.com \
    --cc=vishal.l.verma@intel.com \
    --cc=yan.y.zhao@intel.com \
    --cc=ying.huang@linux.alibaba.com \
    --cc=ziy@nvidia.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox