From: Yize Wang <wangyize7@huawei.com>
To: Marc Zyngier <maz@kernel.org>
Cc: <kvmarm@lists.linux.dev>, <linux-arm-kernel@lists.infradead.org>,
<linux-kernel@vger.kernel.org>, <catalin.marinas@arm.com>,
<will@kernel.org>, <fuad.tabba@linux.dev>, <joey.gouly@arm.com>,
<seiden@linux.ibm.com>, <suzuki.poulose@arm.com>,
<yuzenghui@huawei.com>, <mark.rutland@arm.com>,
<zhengchuan@huawei.com>, <jiangjiacheng@huawei.com>,
<yize_w1110@163.com>
Subject: Re: [PATCH 0/2] Batch register access for live migration optimization
Date: Thu, 8 Oct 2026 16:32:51 +0800 [thread overview]
Message-ID: <77a9ab65-2997-47f8-8e99-04c4467df7b8@huawei.com> (raw)
In-Reply-To: <86ece63kes.wl-maz@kernel.org>
在 2026/10/3 22:53, Marc Zyngier 写道:
> On Sun, 20 Sep 2026 10:15:25 +0100,
> Yize Wang <wangyize7@huawei.com> wrote:
>> 在 2026/9/18 20:08, Marc Zyngier 写道:
>>> On Fri, 18 Sep 2026 09:18:13 +0100,
>>> Yize Wang<wangyize7@huawei.com> wrote:
>>>> This series adds batch register access support to KVM/arm64 to reduce
>>>> syscall overhead during VM live migration.
>>>>
>>>> Currently, QEMU issues one ioctl per register when saving/restoring VGIC
>>>> state. On large VM configurations this means tens of thousands of syscalls,
>>>> where lock acquisition and context switch overhead dominates migration
>>>> downtime. Thus, we provide a batch register method to allow userspace
>>>> read/write multiple distributor and redistributor registers in a single call.
>>>> In this way, we can significantly reduce syscalls and migration downtime.
>>>>
>>>> Test the VM migration time under pressure conditions.
>>>> The VM specifications for migration are as follows:
>>>> - VM use 4-K page;
>>>> - the number of VCPU is 160;
>>>> - the total memory is 320Gigabit;
>>>> - use 'Redis SET-benchmark' to pressurize VM;
>>>>
>>>> Performance results (3-run average, ms):
>>>> | Metric | Without patch | With patch | Improvement |
>>>> |---------------------|---------------|------------|-------------|
>>>> | Migration downtime | 536 | 321 | 40% |
>>>> | Source (total) | 344 | 230 | 33% |
>>>> | - VGIC put | 158 | 40 | 75% |
>>>> | - VGIC get | 120 | 19 | 84% |
>>>> | Destination (total) | 192 | 91 | 53% |
>>>> | - VGIC put | 132 | 27 | 80% |
>>>>
>>>> Yize Wang (2):
>>>> KVM: arm64: Add batch group constant and data structure to UAPI header
>>>> KVM: arm64: Add VGIC v3 batch register access implementation
>>> Questions:
>>>
>>> - Why only the MMIO registers?
>>>
>>> - Why not the sysregs?
>>>
>>> - Why only the GIC?
>>>
>>> - Why not all of the state?
>>>
>>> - Where is the corresponding userspace code?
>>>
>>> More importantly, since this is about batching system calls:
>>>
>>> - Why can't this be done with io_uring instead?
>>>
>>> M.
>>
>> Hi, Marc! Thank you for the review.
>>
>> These patches focus on optimizing GICv3 register access during live
>> migration. We found that there are a large number of locks (kvm->lock,
>> vcpus, config_lock) in the GIC, these lock operations wil cost large
>> time waste. The batches of sysreg for vcpu optimization will come in
>> follow as a separate series. And let me address these questions one by
>> one.
> All these locks should be non-blocking by the time you save anything
> related to the GIC, because:
>
> - none of the vcpu can be running
>
> - this must be a single threaded operation
>
> So if you are seeing anything contended, this is either the sign of a
> bad KVM bug, or an indication that you are violating the above
> requirements.
>
> M.
Hi, Marc! Thanks for your review.
You are right that when saving the GIC state during live migration
process, all vCPUs are not running and it is a single threaded
operation. Our patches process the packaged register states from the
userspace under the same lock. These operations only aim to reduce the
lock opreation downtime and do not involve lock contention.
In addition, these patches focus on reducing the syscalls of ioctl. Now,
every register requires an ioctl, resulting in (num_cpu *
registers_per_cpu) system calls. This will cost lots of time during the
large-scale VM migration. Our patches batch registers by category in
qemu and then send them to kernel in batches. This operations
significantly reduce the number of ioctl syscalls and save a great
amount of downtime.
The comparison is as follow:
BEFORE:
QEMU (Userspace)Kernel (Kernelspace)
+--------------------------++------------------------------+
||||
|kvm_arm_gicv3_get()||vgic_v3_attr_regs_access()|
||||
|for ncpu in num_cpu:|||
|+--------------+|||
|| GICR_CTLR|----|--- ioctl ------> |mutex_lock(&kvm->lock)|
|+--------------+ <--|--- val --------- |kvm_trylock_all_vcpus()|
|+--------------+||mutex_lock(&config_lock)|
|| GICR_STATUSR |----|--- ioctl ------> ||
|+--------------+ <--|--- val --------- |read/write single register|
|+--------------+|||
|| GICR_WAKER|----|--- ioctl ------> |mutex_unlock(&config_lock)|
|+--------------+ <--|--- val --------- |kvm_unlock_all_vcpus()|
|...||mutex_unlock(&kvm->lock)|
|+--------------+|||
|| ICC_SRE_EL1|----|--- ioctl ------> |(repeat lock -> read -> unlock
|+--------------+ <--|--- val --------- |for each register)|
|...|||
||||
+--------------------------++------------------------------+
AFTER:
QEMU (Userspace)Kernel (Kernelspace)
+----------------------------++--------------------------------+
||||
|kvm_arm_gicv3_get()||vgic_v3_batch_access()|
||| |
|+- Collect Phase ----+|||
||| |||
|| batch_add(GICR_CTLR)| |||
|| batch_add(GICR_STATUSR) |||
|| batch_add(GICR_WAKER)||||
|| batch_add(GICR_IGROUPR0)| ||
|| ...||||
|| batch_add(ICC_SRE_EL1) |||
|| ...||||
|+---------------------+ |||
||1 ioctl|copy_from_user(entries)|
|+- Commit Phase -----+|||
||| --- ioctl(BATCH)---> |mutex_lock(&kvm->lock)|
||batch_commit()| ||kvm_trylock_all_vcpus()|
|||||mutex_lock(&config_lock)|
||||||
|+---------------------+ ||+-- Process Each Entry --+|
||||||
|+- Unpack Phase -----+|||entry[0]: read GICR_CTLR| |
||| <---batch result ---||entry[1]: read GICR_STATUSR|
||Write back to QEMU | | ||entry[2]: read GICR_WAKER| |
||state||||...||
||(post-process:| | ||entry[N]: read ICC_SRE | |
||unshuffle, split| | ||| |
||bytes, etc.)| |||* FLAG_LOCKED: skip||
||| | ||locking (already held)| |
|+---------------------+ ||+-------------------------+ |
||| |
+----------------------------++--------------------------------+
next prev parent reply other threads:[~2026-10-08 8:33 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 8:18 [PATCH 0/2] Batch register access for live migration optimization Yize Wang
2026-09-18 8:18 ` [PATCH 1/2] KVM: arm64: Add batch group constant and data structure to UAPI header Yize Wang
2026-09-18 8:18 ` [PATCH 2/2] KVM: arm64: Add VGIC v3 batch register access implementation Yize Wang
2026-09-18 12:08 ` [PATCH 0/2] Batch register access for live migration optimization Marc Zyngier
2026-09-20 12:12 ` [RESEND][PATCH " Yize Wang
2026-09-23 7:57 ` Yize Wang
[not found] ` <f70c6fe3-e3c5-4912-b4bb-8a418fec4346@huawei.com>
2026-10-03 14:53 ` [PATCH " Marc Zyngier
2026-10-08 8:32 ` Yize Wang [this message]
2026-10-08 12:50 ` [RESEND][PATCH " Yize Wang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=77a9ab65-2997-47f8-8e99-04c4467df7b8@huawei.com \
--to=wangyize7@huawei.com \
--cc=catalin.marinas@arm.com \
--cc=fuad.tabba@linux.dev \
--cc=jiangjiacheng@huawei.com \
--cc=joey.gouly@arm.com \
--cc=kvmarm@lists.linux.dev \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mark.rutland@arm.com \
--cc=maz@kernel.org \
--cc=seiden@linux.ibm.com \
--cc=suzuki.poulose@arm.com \
--cc=will@kernel.org \
--cc=yize_w1110@163.com \
--cc=yuzenghui@huawei.com \
--cc=zhengchuan@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox