Linux KVM/arm64 development list
 help / color / mirror / Atom feed
From: Gavin Shan <gshan@redhat.com>
To: Ricardo Koller <ricarkol@google.com>,
	pbonzini@redhat.com, maz@kernel.org, oupton@google.com,
	yuzenghui@huawei.com, dmatlack@google.com
Cc: kvm@vger.kernel.org, kvmarm@lists.linux.dev, qperret@google.com,
	catalin.marinas@arm.com, andrew.jones@linux.dev,
	seanjc@google.com, alexandru.elisei@arm.com,
	suzuki.poulose@arm.com, eric.auger@redhat.com, reijiw@google.com,
	rananta@google.com, bgardon@google.com, ricarkol@gmail.com
Subject: Re: [PATCH v2 00/12] Implement Eager Page Splitting for ARM.
Date: Tue, 14 Feb 2023 16:57:59 +1100	[thread overview]
Message-ID: <1a3afa6d-3478-31dd-6f34-52075875c2fa@redhat.com> (raw)
In-Reply-To: <20230206165851.3106338-1-ricarkol@google.com>

Hi Ricardo,

On 2/7/23 3:58 AM, Ricardo Koller wrote:
> Eager Page Splitting improves the performance of dirty-logging (used
> in live migrations) when guest memory is backed by huge-pages.  It's
> an optimization used in Google Cloud since 2016 on x86, and for the
> last couple of months on ARM.
> 
> Background and motivation
> =========================
> Dirty logging is typically used for live-migration iterative copying.
> KVM implements dirty-logging at the PAGE_SIZE granularity (will refer
> to 4K pages from now on).  It does it by faulting on write-protected
> 4K pages.  Therefore, enabling dirty-logging on a huge-page requires
> breaking it into 4K pages in the first place.  KVM does this breaking
> on fault, and because it's in the critical path it only maps the 4K
> page that faulted; every other 4K page is left unmapped.  This is not
> great for performance on ARM for a couple of reasons:
> 
> - Splitting on fault can halt vcpus for milliseconds in some
>    implementations. Splitting a block PTE requires using a broadcasted
>    TLB invalidation (TLBI) for every huge-page (due to the
>    break-before-make requirement). Note that x86 doesn't need this. We
>    observed some implementations that take millliseconds to complete
>    broadcasted TLBIs when done in parallel from multiple vcpus.  And
>    that's exactly what happens when doing it on fault: multiple vcpus
>    fault at the same time triggering TLBIs in parallel.
> 
> - Read intensive guest workloads end up paying for dirty-logging.
>    Only mapping the faulting 4K page means that all the other pages
>    that were part of the huge-page will now be unmapped. The effect is
>    that any access, including reads, now has to fault.
> 
> Eager Page Splitting (on ARM)
> =============================
> Eager Page Splitting fixes the above two issues by eagerly splitting
> huge-pages when enabling dirty logging. The goal is to avoid doing it
> while faulting on write-protected pages. This is what the TDP MMU does
> for x86 [0], except that x86 does it for different reasons: to avoid
> grabbing the MMU lock on fault. Note that taking care of
> write-protection faults still requires grabbing the MMU lock on ARM,
> but not on x86 (with the fast_page_fault path).
> 
> An additional benefit of eagerly splitting huge-pages is that it can
> be done in a controlled way (e.g., via an IOCTL). This series provides
> two knobs for doing it, just like its x86 counterpart: when enabling
> dirty logging, and when using the KVM_CLEAR_DIRTY_LOG ioctl. The
> benefit of doing it on KVM_CLEAR_DIRTY_LOG is that this ioctl takes
> ranges, and not complete memslots like when enabling dirty logging.
> This means that the cost of splitting (mainly broadcasted TLBIs) can
> be throttled: split a range, wait for a bit, split another range, etc.
> The benefits of this approach were presented by Oliver Upton at KVM
> Forum 2022 [1].
> 

[...]

Sorry for raising questions about the design lately. There are two operations
regarding the existing huge page mapping. Here, lets take PMD and PTE mapping
as an example for discussion: (a) The existing PMD mapping is split to contiguous
512 PTE mappings when all sub-pages are written in sequence and dirty logging has
been enabled (b) The contiguous 512 PTE mappings are combined to one PMD mapping
when dirty logging is disabled.

Before this series is applied, both (a) and (b) are handled by the page fault handler.
After this series is applied, (a) is handled in the ioctl handler while (b) is still
handled in the page fault handler. I'm not sure why we can't eagerly split the PMD
mapping into 512 PTE mapping in the page fault handler? In this way, the implementation
may be simplified by extending kvm_pgtable_stage2_map(). In the implementation, the
newly introduced API kvm_pgtable_stage2_split() calls to kvm_pgtable_stage2_create_unlinked()
and then stage2_map_walker(), which is part of kvm_pgtable_stage2_map(), to create the
unlinked page tables. It's why I have the question.

Thanks,
Gavin



  parent reply	other threads:[~2023-02-14  5:58 UTC|newest]

Thread overview: 30+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2023-02-06 16:58 [PATCH v2 00/12] Implement Eager Page Splitting for ARM Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 01/12] KVM: arm64: Add KVM_PGTABLE_WALK ctx->flags for skipping BBM and CMO Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 02/12] KVM: arm64: Rename free_unlinked to free_removed Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 03/12] KVM: arm64: Add helper for creating unlinked stage2 subtrees Ricardo Koller
2023-02-07 14:51   ` Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 04/12] KVM: arm64: Add kvm_pgtable_stage2_split() Ricardo Koller
2023-02-09  5:58   ` Gavin Shan
2023-02-09 12:40     ` Ricardo Koller
2023-02-09 16:17       ` Ricardo Koller
2023-02-09 22:48         ` Gavin Shan
2023-02-15 17:43           ` Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 05/12] KVM: arm64: Refactor kvm_arch_commit_memory_region() Ricardo Koller
2023-02-09  6:02   ` Gavin Shan
2023-02-15 17:47     ` Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 06/12] KVM: arm64: Add kvm_uninit_stage2_mmu() Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 07/12] KVM: arm64: Export kvm_are_all_memslots_empty() Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 08/12] KVM: arm64: Add KVM_CAP_ARM_EAGER_SPLIT_CHUNK_SIZE Ricardo Koller
2023-02-08 11:05   ` Gavin Shan
2023-02-06 16:58 ` [PATCH v2 09/12] KVM: arm64: Split huge pages when dirty logging is enabled Ricardo Koller
2023-02-09  6:26   ` Gavin Shan
2023-02-09 12:50     ` Ricardo Koller
2023-02-09 23:09       ` Gavin Shan
2023-02-15 16:25       ` Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 10/12] KVM: arm64: Open-code kvm_mmu_write_protect_pt_masked() Ricardo Koller
2023-02-06 16:58 ` [PATCH v2 11/12] KVM: arm64: Split huge pages during KVM_CLEAR_DIRTY_LOG Ricardo Koller
2023-02-09  6:29   ` Gavin Shan
2023-02-06 16:58 ` [PATCH v2 12/12] KVM: arm64: Use local TLBI on permission relaxation Ricardo Koller
2023-02-14  5:57 ` Gavin Shan [this message]
2023-02-14  7:33   ` [PATCH v2 00/12] Implement Eager Page Splitting for ARM Oliver Upton
2023-02-14 21:59     ` Ricardo Koller

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=1a3afa6d-3478-31dd-6f34-52075875c2fa@redhat.com \
    --to=gshan@redhat.com \
    --cc=alexandru.elisei@arm.com \
    --cc=andrew.jones@linux.dev \
    --cc=bgardon@google.com \
    --cc=catalin.marinas@arm.com \
    --cc=dmatlack@google.com \
    --cc=eric.auger@redhat.com \
    --cc=kvm@vger.kernel.org \
    --cc=kvmarm@lists.linux.dev \
    --cc=maz@kernel.org \
    --cc=oupton@google.com \
    --cc=pbonzini@redhat.com \
    --cc=qperret@google.com \
    --cc=rananta@google.com \
    --cc=reijiw@google.com \
    --cc=ricarkol@gmail.com \
    --cc=ricarkol@google.com \
    --cc=seanjc@google.com \
    --cc=suzuki.poulose@arm.com \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox