Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Yize Wang <wangyize7@huawei.com>
To: Marc Zyngier <maz@kernel.org>
Cc: <kvmarm@lists.linux.dev>, <linux-arm-kernel@lists.infradead.org>,
	<linux-kernel@vger.kernel.org>, <catalin.marinas@arm.com>,
	<will@kernel.org>, <fuad.tabba@linux.dev>, <joey.gouly@arm.com>,
	<seiden@linux.ibm.com>, <suzuki.poulose@arm.com>,
	<yuzenghui@huawei.com>, <mark.rutland@arm.com>,
	<zhengchuan@huawei.com>, <jiangjiacheng@huawei.com>,
	<yize_w1110@163.com>
Subject: Re: [RESEND][PATCH 0/2] Batch register access for live migration optimization
Date: Thu, 8 Oct 2026 20:50:57 +0800	[thread overview]
Message-ID: <6890def8-168d-4951-b6bd-bcbbcbf3595a@huawei.com> (raw)
In-Reply-To: <86ece63kes.wl-maz@kernel.org>


在 2026/10/3 22:53, Marc Zyngier 写道:
> On Sun, 20 Sep 2026 10:15:25 +0100,
> Yize Wang <wangyize7@huawei.com> wrote:
>> 在 2026/9/18 20:08, Marc Zyngier 写道:
>>> On Fri, 18 Sep 2026 09:18:13 +0100,
>>> Yize Wang<wangyize7@huawei.com> wrote:
>>>> This series adds batch register access support to KVM/arm64 to reduce
>>>> syscall overhead during VM live migration.
>>>>
>>>> Currently, QEMU issues one ioctl per register when saving/restoring VGIC
>>>> state. On large VM configurations this means tens of thousands of syscalls,
>>>> where lock acquisition and context switch overhead dominates migration
>>>> downtime. Thus, we provide a batch register method to allow userspace
>>>> read/write multiple distributor and redistributor registers in a single call.
>>>> In this way, we can significantly reduce syscalls and migration downtime.
>>>>
>>>> Test the VM migration time under pressure conditions.
>>>> The VM specifications for migration are as follows:
>>>> - VM use 4-K page;
>>>> - the number of VCPU is 160;
>>>> - the total memory is 320Gigabit;
>>>> - use 'Redis SET-benchmark' to pressurize VM;
>>>>
>>>> Performance results (3-run average, ms):
>>>>       | Metric              | Without patch | With patch | Improvement |
>>>>       |---------------------|---------------|------------|-------------|
>>>>       | Migration downtime  |        536    |     321    |     40%     |
>>>>       | Source (total)      |        344    |     230    |     33%     |
>>>>       |   - VGIC put        |        158    |      40    |     75%     |
>>>>       |   - VGIC get        |        120    |      19    |     84%     |
>>>>       | Destination (total) |        192    |      91    |     53%     |
>>>>       |   - VGIC put        |        132    |      27    |     80%     |
>>>>
>>>> Yize Wang (2):
>>>>     KVM: arm64: Add batch group constant and data structure to UAPI header
>>>>     KVM: arm64: Add VGIC v3 batch register access implementation
>>> Questions:
>>>
>>> - Why only the MMIO registers?
>>>
>>> - Why not the sysregs?
>>>
>>> - Why only the GIC?
>>>
>>> - Why not all of the state?
>>>
>>> - Where is the corresponding userspace code?
>>>
>>> More importantly, since this is about batching system calls:
>>>
>>> - Why can't this be done with io_uring instead?
>>>
>>> 	M.
>>
>> Hi, Marc! Thank you for the review.
>>
>>    These patches focus on optimizing GICv3 register access during live
>> migration. We found that there are a large number of locks (kvm->lock,
>> vcpus, config_lock) in the GIC, these lock operations wil cost large
>> time waste. The batches of sysreg for vcpu optimization will come in
>> follow as a separate series. And let me address these questions one by
>> one.
> All these locks should be non-blocking by the time you save anything
> related to the GIC, because:
>
> - none of the vcpu can be running
>
> - this must be a single threaded operation
>
> So if you are seeing anything contended, this is either the sign of a
> bad KVM bug, or an indication that you are violating the above
> requirements.
>
> 	M.


Hi, Marc! Thanks for your review.

You are right that when saving the GIC state during live migration process, all vCPUs are not running and it is a single threaded operation. Our patches process the packaged register states from the userspace under the same lock. These operations only aim to reduce the lock opreation downtime and do not involve lock contention.

In addition, these patches focus on reducing the syscalls of ioctl. Now, every register requires an ioctl, resulting in (num_cpu * registers_per_cpu) system calls. This will cost lots of time during the large-scale VM  migration. Our patches batch registers by category in qemu and then send them to kernel in batches. This operations significantly reduce the number of ioctl syscalls and save a great amount of downtime.

The comparison is as follow:
BEFORE:

        QEMU (Userspace)                  Kernel (Kernelspace)
   +------------------------+        +----------------------------+
   |                        |        |                            |
   |  kvm_arm_gicv3_get()   |        |  vgic_v3_attr_regs_access  |
   |                        |        |                            |
   |  for ncpu in num_cpu:  |        |                            |
   |    for each register:  |        |                            |
   |                        |        |                            |
   |  GICR_CTLR             | ioctl->|  mutex_lock(&kvm->lock)    |
   |                        |        |  kvm_trylock_all_vcpus     |
   |                        |        |  mutex_lock(&config_lock)  |
   |                        | <-val  |  read/write register       |
   |                        |        |  mutex_unlock              |
   |                        |        |  kvm_unlock_all_vcpus      |
   |                        |        |  mutex_unlock(&kvm->lock)  |
   |                        |        |                            |
   |  GICR_STATUSR          | ioctl->|  (repeat lock/read         |
   |                        | <-val  |   -> unlock)               |
   |  ...                   |        |                            |
   |                        |        |                            |
   |  GICR_WAKER            | ioctl->|  (repeat lock/read         |
   |                        | <-val  |   -> unlock)               |
   |  ...                   |        |                            |
   |                        |        |                            |
   |  ICC_SRE_EL1           | ioctl->|  (repeat lock/read         |
   |                        | <-val  |   -> unlock)               |
   |  ...                   |        |                            |
   |                        |        |                            |
   +------------------------+        +----------------------------+


AFTER:

        QEMU (Userspace)                  Kernel (Kernelspace)
   +------------------------+        +----------------------------+
   |                        |        |                            |
   |  kvm_arm_gicv3_get()   |        |  vgic_v3_batch_access      |
   |                        |        |                            |
   |  Collect Phase:        |        |                            |
   |    add(GICR_CTLR)      |        |                            |
   |    add(GICR_STATUSR)   |        |                            |
   |    add(GICR_WAKER)     |        |                            |
   |    add(GICR_IGROUPR0)  |        |                            |
   |    ...                 |        |                            |
   |    add(ICC_SRE_EL1)    |        |                            |
   |    ...                 |        |                            |
   |                        |        |                            |
   |  Commit Phase:         | BATCH->|  copy_from_user            |
   |    batch_commit()      |        |  mutex_lock(&kvm->lock)    |
   |                        |        |  kvm_trylock_all_vcpus     |
   |                        |        |  mutex_lock(&config_lock)  |
   |                        |        |                            |
   |                        |        |  Process Each Entry:       |
   |                        |        |    entry[0]: GICR_CTLR     |
   |                        |        |    entry[1]: GICR_STATUSR  |
   |                        |        |    entry[2]: GICR_WAKER    |
   |                        |        |    ...                     |
   |                        |        |    entry[N]: ICC_SRE       |
   |                        |        |    FLAG_LOCKED: skip       |
   |                        |        |                            |
   |  Unpack Phase:         |<-result|  mutex_unlock              |
   |    Write back          |        |  kvm_unlock_all_vcpus      |
   |    Post-process:       |        |  mutex_unlock(&kvm->lock)  |
   |      unshuffle,split   |        |  copy_to_user              |
   |                        |        |                            |
   +------------------------+        +----------------------------+





      parent reply	other threads:[~2026-10-08 12:52 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-18  8:18 [PATCH 0/2] Batch register access for live migration optimization Yize Wang
2026-09-18  8:18 ` [PATCH 1/2] KVM: arm64: Add batch group constant and data structure to UAPI header Yize Wang
2026-09-18  8:18 ` [PATCH 2/2] KVM: arm64: Add VGIC v3 batch register access implementation Yize Wang
2026-09-18 12:08 ` [PATCH 0/2] Batch register access for live migration optimization Marc Zyngier
2026-09-20 12:12   ` [RESEND][PATCH " Yize Wang
2026-09-23  7:57     ` Yize Wang
     [not found]   ` <f70c6fe3-e3c5-4912-b4bb-8a418fec4346@huawei.com>
2026-10-03 14:53     ` [PATCH " Marc Zyngier
2026-10-08  8:32       ` Yize Wang
2026-10-08 12:50       ` Yize Wang [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=6890def8-168d-4951-b6bd-bcbbcbf3595a@huawei.com \
    --to=wangyize7@huawei.com \
    --cc=catalin.marinas@arm.com \
    --cc=fuad.tabba@linux.dev \
    --cc=jiangjiacheng@huawei.com \
    --cc=joey.gouly@arm.com \
    --cc=kvmarm@lists.linux.dev \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=maz@kernel.org \
    --cc=seiden@linux.ibm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=will@kernel.org \
    --cc=yize_w1110@163.com \
    --cc=yuzenghui@huawei.com \
    --cc=zhengchuan@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox