From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id D33DBCA6007 for ; Thu, 8 Oct 2026 12:52:02 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:Content-Transfer-Encoding: Content-Type:In-Reply-To:From:References:CC:To:Subject:MIME-Version:Date: Message-ID:Reply-To:Content-ID:Content-Description:Resent-Date:Resent-From: Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=azu39hy4sP9l9bRMcKId7i+KbDup1QkTOPoJ+2kNYCw=; b=spkgvMn6MnIpNGZs7VIhnXBlLS 34Af4AMkUPdgykp77e4R7zLeNoQJ6nTc/6uT5gKKnetJbP6/UkHxsSPszOzt9MExpYTHKdc5OjyEY cnBzLofjuY+nmSEa1YYQRZjkiWRtUswLSe7k2PGkKZK/WWeCEc8rwSF+sNt8lHkqmrnnjO7yRpn66 Hk77cQvDSdHob8nyme5bE2P/WDUk5Qv6GAEbUIm2I5Vt5imqbmqtOKoNXWeYyN7E0wuYZ1uZy5PEv c+YonNFArucS2OJwpyvO2Qxu0TKs/XHm2h0KQfuzGCDaiZAE5ys33fxsHCSM5mdNvqffQNj0hxKBT zDdFyS1A==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1xEnb5-00000004OgQ-3ic6; Thu, 08 Oct 2026 12:51:43 +0000 Received: from desiato.infradead.org ([2001:8b0:10b:1:d65d:64ff:fe57:4e05]) by bombadil.infradead.org with esmtps (Exim 4.99.1 #2 (Red Hat Linux)) id 1xEnb4-00000004OgA-0x3I for linux-arm-kernel@bombadil.infradead.org; Thu, 08 Oct 2026 12:51:42 +0000 DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=infradead.org; s=desiato.20200630; h=Content-Transfer-Encoding:Content-Type :In-Reply-To:From:References:CC:To:Subject:MIME-Version:Date:Message-ID: Sender:Reply-To:Content-ID:Content-Description; bh=azu39hy4sP9l9bRMcKId7i+KbDup1QkTOPoJ+2kNYCw=; b=Wm0mec4PE2hnToMjQyrpBsFYaU kqhaiX3/m1AmRjpJ5DIdf+msbBF+xSqvqLewI37ogMnZ6dBJFQQRA1Rl557Mq7wSSPLkvM7Vyov0b M7StxmEcOiavTQdm0lX4HdlZWMQ9D8/SnYbEgf7JVSjasIKwNKUUO3ORi3TuQCZetGRpL/ZZKZbfr Mw1hpp7SYycBkHG6IWl8JTOX+b7Mb4/zpgDJrkUVVfSF42n5+C7TT3YNmMIWEJ5xTr+OoglgaQZwp xzjE9jpJimmYv12jSdFVYHeVNFihaMIvybFnqXDzgUJ9UVeARbQnmI9K2SLZZUfGSu9zJXwtBdyzR 2L4ln/wg==; Received: from canpmsgout01.his.huawei.com ([113.46.200.216]) by desiato.infradead.org with esmtps (Exim 4.99.2 #2 (Red Hat Linux)) id 1xEnb0-0000000ANfi-3864 for linux-arm-kernel@lists.infradead.org; Thu, 08 Oct 2026 12:51:41 +0000 dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=azu39hy4sP9l9bRMcKId7i+KbDup1QkTOPoJ+2kNYCw=; b=cdcy/2djDSiUkkPO0QEaErPKpvMVBP3ghfWe4K54nsje1g/e2wppO6EK36bXKK/gC6SPjDq4g MZl/77my+iIk6VNg+ZZFh79SQu/VJt3dHA4IRb8SwoYyBKsiP07Lho2A92wcK2rcpEL/skRCsex VlU74hvEOtmS6lwk0qd+HOg= Received: from mail.maildlp.com (unknown [172.19.163.104]) by canpmsgout01.his.huawei.com (SkyGuard) with ESMTPS id 4j0qKy1bMLz1T4JM; Thu, 8 Oct 2026 20:38:46 +0800 (CST) Received: from dggpemr200003.china.huawei.com (unknown [7.185.36.25]) by mail.maildlp.com (Postfix) with ESMTPS id 4B1004057F; Thu, 8 Oct 2026 20:50:59 +0800 (CST) Received: from [10.174.185.234] (10.174.185.234) by dggpemr200003.china.huawei.com (7.185.36.25) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.46; Thu, 8 Oct 2026 20:50:58 +0800 Message-ID: <6890def8-168d-4951-b6bd-bcbbcbf3595a@huawei.com> Date: Thu, 8 Oct 2026 20:50:57 +0800 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [RESEND][PATCH 0/2] Batch register access for live migration optimization To: Marc Zyngier CC: , , , , , , , , , , , , , References: <20260918081930.4014735-1-wangyize7@huawei.com> <861paq69te.wl-maz@kernel.org> <86ece63kes.wl-maz@kernel.org> From: Yize Wang In-Reply-To: <86ece63kes.wl-maz@kernel.org> Content-Type: text/plain; charset="UTF-8"; format=flowed Content-Transfer-Encoding: 8bit X-Originating-IP: [10.174.185.234] X-ClientProxiedBy: kwepems500001.china.huawei.com (7.221.188.70) To dggpemr200003.china.huawei.com (7.185.36.25) X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.9.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20261008_135139_505469_BFADB08C X-CRM114-Status: GOOD ( 18.07 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org 在 2026/10/3 22:53, Marc Zyngier 写道: > On Sun, 20 Sep 2026 10:15:25 +0100, > Yize Wang wrote: >> 在 2026/9/18 20:08, Marc Zyngier 写道: >>> On Fri, 18 Sep 2026 09:18:13 +0100, >>> Yize Wang wrote: >>>> This series adds batch register access support to KVM/arm64 to reduce >>>> syscall overhead during VM live migration. >>>> >>>> Currently, QEMU issues one ioctl per register when saving/restoring VGIC >>>> state. On large VM configurations this means tens of thousands of syscalls, >>>> where lock acquisition and context switch overhead dominates migration >>>> downtime. Thus, we provide a batch register method to allow userspace >>>> read/write multiple distributor and redistributor registers in a single call. >>>> In this way, we can significantly reduce syscalls and migration downtime. >>>> >>>> Test the VM migration time under pressure conditions. >>>> The VM specifications for migration are as follows: >>>> - VM use 4-K page; >>>> - the number of VCPU is 160; >>>> - the total memory is 320Gigabit; >>>> - use 'Redis SET-benchmark' to pressurize VM; >>>> >>>> Performance results (3-run average, ms): >>>> | Metric | Without patch | With patch | Improvement | >>>> |---------------------|---------------|------------|-------------| >>>> | Migration downtime | 536 | 321 | 40% | >>>> | Source (total) | 344 | 230 | 33% | >>>> | - VGIC put | 158 | 40 | 75% | >>>> | - VGIC get | 120 | 19 | 84% | >>>> | Destination (total) | 192 | 91 | 53% | >>>> | - VGIC put | 132 | 27 | 80% | >>>> >>>> Yize Wang (2): >>>> KVM: arm64: Add batch group constant and data structure to UAPI header >>>> KVM: arm64: Add VGIC v3 batch register access implementation >>> Questions: >>> >>> - Why only the MMIO registers? >>> >>> - Why not the sysregs? >>> >>> - Why only the GIC? >>> >>> - Why not all of the state? >>> >>> - Where is the corresponding userspace code? >>> >>> More importantly, since this is about batching system calls: >>> >>> - Why can't this be done with io_uring instead? >>> >>> M. >> >> Hi, Marc! Thank you for the review. >> >>   These patches focus on optimizing GICv3 register access during live >> migration. We found that there are a large number of locks (kvm->lock, >> vcpus, config_lock) in the GIC, these lock operations wil cost large >> time waste. The batches of sysreg for vcpu optimization will come in >> follow as a separate series. And let me address these questions one by >> one. > All these locks should be non-blocking by the time you save anything > related to the GIC, because: > > - none of the vcpu can be running > > - this must be a single threaded operation > > So if you are seeing anything contended, this is either the sign of a > bad KVM bug, or an indication that you are violating the above > requirements. > > M. Hi, Marc! Thanks for your review. You are right that when saving the GIC state during live migration process, all vCPUs are not running and it is a single threaded operation. Our patches process the packaged register states from the userspace under the same lock. These operations only aim to reduce the lock opreation downtime and do not involve lock contention. In addition, these patches focus on reducing the syscalls of ioctl. Now, every register requires an ioctl, resulting in (num_cpu * registers_per_cpu) system calls. This will cost lots of time during the large-scale VM migration. Our patches batch registers by category in qemu and then send them to kernel in batches. This operations significantly reduce the number of ioctl syscalls and save a great amount of downtime. The comparison is as follow: BEFORE: QEMU (Userspace) Kernel (Kernelspace) +------------------------+ +----------------------------+ | | | | | kvm_arm_gicv3_get() | | vgic_v3_attr_regs_access | | | | | | for ncpu in num_cpu: | | | | for each register: | | | | | | | | GICR_CTLR | ioctl->| mutex_lock(&kvm->lock) | | | | kvm_trylock_all_vcpus | | | | mutex_lock(&config_lock) | | | <-val | read/write register | | | | mutex_unlock | | | | kvm_unlock_all_vcpus | | | | mutex_unlock(&kvm->lock) | | | | | | GICR_STATUSR | ioctl->| (repeat lock/read | | | <-val | -> unlock) | | ... | | | | | | | | GICR_WAKER | ioctl->| (repeat lock/read | | | <-val | -> unlock) | | ... | | | | | | | | ICC_SRE_EL1 | ioctl->| (repeat lock/read | | | <-val | -> unlock) | | ... | | | | | | | +------------------------+ +----------------------------+ AFTER: QEMU (Userspace) Kernel (Kernelspace) +------------------------+ +----------------------------+ | | | | | kvm_arm_gicv3_get() | | vgic_v3_batch_access | | | | | | Collect Phase: | | | | add(GICR_CTLR) | | | | add(GICR_STATUSR) | | | | add(GICR_WAKER) | | | | add(GICR_IGROUPR0) | | | | ... | | | | add(ICC_SRE_EL1) | | | | ... | | | | | | | | Commit Phase: | BATCH->| copy_from_user | | batch_commit() | | mutex_lock(&kvm->lock) | | | | kvm_trylock_all_vcpus | | | | mutex_lock(&config_lock) | | | | | | | | Process Each Entry: | | | | entry[0]: GICR_CTLR | | | | entry[1]: GICR_STATUSR | | | | entry[2]: GICR_WAKER | | | | ... | | | | entry[N]: ICC_SRE | | | | FLAG_LOCKED: skip | | | | | | Unpack Phase: |<-result| mutex_unlock | | Write back | | kvm_unlock_all_vcpus | | Post-process: | | mutex_unlock(&kvm->lock) | | unshuffle,split | | copy_to_user | | | | | +------------------------+ +----------------------------+