From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id CAD74C55182 for ; Wed, 5 Aug 2026 03:43:53 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:List-Subscribe:List-Help :List-Post:List-Archive:List-Unsubscribe:List-Id:Content-Transfer-Encoding: Content-Type:In-Reply-To:From:References:CC:To:Subject:MIME-Version:Date: Message-ID:Reply-To:Content-ID:Content-Description:Resent-Date:Resent-From: Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:List-Owner; bh=B3lSrMpocDV/Hvep88z36Dgthq0eziOupD1/vUR+w/A=; b=LoebcCzb9BGADtBAMSkK2m+Aam i3xxa3u+Re0Uz5vsEwB7wzNbx4MxMfdA6C3kgpqtDPP7YCoWKjtYFERflIv2Id6znCgVBTkw0P1jl J3muw7vMCdybvy2+OxCJTkniQd8kIdr6KzN9LZEQwW8It56VB5PBC1qWdEmC66I1utksM88VIC4Nv 9cwQXphjo7kGIK5JvTKPZ3LqNQxf/+OEBCExde3lhmkSZj5nK4r+MdYyMMnnmqWZeBs9jHkGUUgFF C8MtXsyRfDlj+1/8GMbRcisHFMML2oQoWi2CttKlIyS+8h+zkhO2va26mHcZkwmycIOA/ey+5WPXn 6pexYBzQ==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1wrSXj-00000003BWd-0v9z; Wed, 05 Aug 2026 03:43:47 +0000 Received: from canpmsgout07.his.huawei.com ([113.46.200.222]) by bombadil.infradead.org with esmtps (Exim 4.99.1 #2 (Red Hat Linux)) id 1wrSXg-00000003BWJ-2oeV for linux-arm-kernel@lists.infradead.org; Wed, 05 Aug 2026 03:43:46 +0000 dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=B3lSrMpocDV/Hvep88z36Dgthq0eziOupD1/vUR+w/A=; b=m/eLKphe0U0vdf134ra5FOXe/Jdx1SdB1VrNfeMXmz90/Cr9j09gSN/YbpqKLP3HFt1XUe99f pbkE9eQb3ekisjALdFrXvyV6mRMMdOeegaLcL4JGFDKbG1CoScMBS4Qmt5wbplmmD0QEK6zlySO gVDyzmzcDl9p6kDJ/wTQ8UI= Received: from mail.maildlp.com (unknown [172.19.163.127]) by canpmsgout07.his.huawei.com (SkyGuard) with ESMTPS id 4hFGGw6rGDzLlTg; Wed, 5 Aug 2026 11:34:00 +0800 (CST) Received: from kwepemr100010.china.huawei.com (unknown [7.202.195.125]) by mail.maildlp.com (Postfix) with ESMTPS id 2BCA840572; Wed, 5 Aug 2026 11:43:32 +0800 (CST) Received: from [10.67.120.103] (10.67.120.103) by kwepemr100010.china.huawei.com (7.202.195.125) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.36; Wed, 5 Aug 2026 11:43:31 +0800 Message-ID: Date: Wed, 5 Aug 2026 11:43:30 +0800 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v4 3/6] KVM: arm64: Add auto DBM support for hardware dirty tracking To: Leonardo Bras CC: Oliver Upton , , , , , , , , , , , , , , , , , , References: <96516762-f004-4c2a-a9c6-6fbad49ad6ae@huawei.com> <8e36e2c8-587f-4228-ab13-d6783927a281@huawei.com> <45d76d0d-48e9-46a6-b1f9-691f840eba46@huawei.com> <95f1ebd4-be04-4684-be68-388f752a1379@huawei.com> <2070002d-94aa-4a5b-8df1-e8e0f0fd265c@huawei.com> From: Tian Zheng In-Reply-To: Content-Type: text/plain; charset="UTF-8"; format=flowed Content-Transfer-Encoding: 8bit X-Originating-IP: [10.67.120.103] X-ClientProxiedBy: kwepems500002.china.huawei.com (7.221.188.17) To kwepemr100010.china.huawei.com (7.202.195.125) X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.9.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20260804_204345_046226_17C8CB20 X-CRM114-Status: GOOD ( 45.14 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On 8/4/2026 7:10 PM, Leonardo Bras wrote: > On Tue, Aug 04, 2026 at 12:54:16PM +0800, Tian Zheng wrote: >> >> >> On 8/4/2026 12:32 AM, Leonardo Bras wrote: >>> On Mon, Aug 03, 2026 at 09:57:46PM +0800, Tian Zheng wrote: >>>> >>>> >>>> On 8/3/2026 6:21 PM, Leonardo Bras wrote: >>>>> On Mon, Aug 03, 2026 at 12:04:24PM +0800, Tian Zheng wrote: >>>>>> >>>>>> >>>>>> On 8/3/2026 9:33 AM, Tian Zheng wrote: >>>>>>>>>>>>> 09, 2026 at 06:40:23PM +0800, Tian Zheng wrote: >>>>>>>>>>>>>> -    if (prot & KVM_PGTABLE_PROT_W) >>>>>>>>>>>>>> +    if (prot & KVM_PGTABLE_PROT_W) { >>>>>>>>>>>>>>             set |= KVM_PTE_LEAF_ATTR_LO_S2_S2AP_W; >>>>>>>>>>>>>> >>>>>>>>>>>>>> +        /* >>>>>>>>>>>>>> +         * No DEVICE filter needed here: >>>>>>>>>>>>>> relax_perms is only called >>>>>>>>>>>>>> +         * on FSC_PERM faults. Device pages >>>>>>>>>>>>>> always get full RW from >>>>>>>>>>>>>> +         * initial mapping and are never write-protected during >>>>>>>>>>>>>> +         * migration, so they never trigger a permission fault. >>>>>>>>>>>>>> +         */ >>>>>>>>>>>>>> +        if (pgt->flags & KVM_PGTABLE_S2_DBM) >>>>>>>>>>>>>> +            set |= KVM_PTE_LEAF_ATTR_HI_S2_DBM; >>>>>>>>>>>>>> +    } else { >>>>>>>>>>>>>> +        /* >>>>>>>>>>>>>> +         * Clear DBM on W→RO downgrade to prevent hardware from >>>>>>>>>>>>>> +         * silently upgrading RO+DBM back to W+dirty, which would >>>>>>>>>>>>>> +         * bypass KVM's write tracking and cause data corruption. >>>>>>>>>>>>>> +         */ >>>>>>>>>>>>>> +        clr |= KVM_PTE_LEAF_ATTR_HI_S2_DBM; >>>>>>>>>>>>>> +    } >>>>>>>>>>>>>> + >>>>>>>>>>>>> This block makes it pretty evident that the DBM bit really *is* the >>>>>>>>>>>>> write permission bit. I'd much rather we >>>>>>>>>>>>> introduce the concept of dirty >>>>>>>>>>>>> state to the page table library and migrate the abstract write >>>>>>>>>>>>> permission to the DBM field, even if we don't have FEAT_HAFDBS. >>>>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> Ohh, that's an amazing idea! >>>>>>>>>> >>>>>>>>>> Thinking about that again... >>>>>>>>>> If we adopt the encoding with DBM being the write-permission >>>>>>>>>> bit, and all >>>>>>>>>> PTEs have it since the start, how can we have lazy-splitting happening? >>>>>>>>>> >>>>>>>>>> Only way I think of is removing both DBM and S2_S2AP_W bit >>>>>>>>>> from writable >>>>>>>>>> PTEs during dirty-track enable, and re-adding them during >>>>>>>>>> the first write >>>>>>>>>> fault. If we don't remove the DBM bit, systems with HDBSS >>>>>>>>>> would just dirty >>>>>>>>>> it by hardware, without causing a fault. >>>>>>>>>> >>>>>>>>>> DBM=0 would need to happen only in the first write-protect (only on >>>>>>>>>> lazy-splitting). All other write-protecting would just clean >>>>>>>>>> the S2_S2AP_W >>>>>>>>>> bit, as everything is already split. >>>>>>>>>> >>>>>>>>>> Is that what was intended? >>>>>>>>>> >>>>>>>>>> Thanks! >>>>>>>>>> Leo >>>>>>>>>> >>>>>>>>> Hi Leo, >>>>>>>>> >>>>>>>>> I think the cleanest way to handle this is to simply avoid setting DBM >>>>>>>>> on block mappings. If we only set DBM on page-level PTEs, then block >>>>>>>>> mappings will naturally stay DBM=0 and trigger a write fault on first >>>>>>>>> access — exactly what we need for lazy splitting. >>>>>>>>> >>>>>>>>> When the fault occurs, the block gets split into page-level PTEs, and at >>>>>>>>> that point we can set DBM=1 on the resulting leaf entries. This way: >>>>>>>>> >>>>>>>>> 1. Lazy split works naturally (fault -> split -> set DBM=1) >>>>>>>>> >>>>>>>>> 2. No need to clear DBM globally at dirty-track enable >>>>>>>>> >>>>>>>>> 3. No special handling for block mappings >>>>>>>>> >>>>>>>>> So I think global DBM is still viable — we just need to filter out block >>>>>>>>> mappings when setting the DBM bit. That way the lazy split path >>>>>>>>> is preserved >>>>>>>>> without extra complexity. >>>>>>>> >>>>>>>> Hi Tian, >>>>>>>> >>>>>>>> Humm, but would not that be contrary to what Oliver suggested: >>>>>>>> changing the >>>>>>>> encoding from the PTE for all entries? >>>>>>>> >>>>>>>> (Like, if the PTE is writable, it has to have DBM set) >>>>>>>> >>>>>>>> IIUC what you said, on first faulting of the page in the VM: >>>>>>>> - If the entry is a page (level-3 leaf) and writable, add DBM >>>>>>>> - If it's a block entry (leaf but not a level-3), don't add DBM >>>>>>>> >>>>>>>> So after we enable dirty-logging: >>>>>>>> - a level-3 entry would not fault, using HDBSS, and >>>>>>>> - a block entry would fault, do the splitting, and add DBM to level-3 >>>>>>>>    entries during the split. >>>>>>>> >>>>>>>> If I got that correct, that would be clean indeed. >>>>>>>> >>>>>>>> But then we would have a different encoding for block entries and page >>>>>>>> entries. In page entries, DBM could be used to say if the page is >>>>>>>> writable, >>>>>>>> but on block entries one would have to look at the 'dirty-bit'. >>>>>>>> >>>>>>>> Would that be ok? >>>>>>>> >>>>>>>> Thanks! >>>>>>>> Leo >>>>>>>> >>>>>>> Hi Leo, >>>>>>> >>>>>>> My initial concern was that clearing all DBM bits at the start of >>>>>>> migration would be too expensive, so I thought distinguishing between >>>>>>> level-3 entries and block entries would be better. >>>>>>> >>>>> >>>>> I think we expect it to be expensive, but since we already clean the >>>>> dirty-bit (ro/rw) bit, we can have both happening in the same write :) >>>>> >>>>> (since we only mark the DBM bit when we fault the memory on lazy-splitting, >>>>> we are expecting to have the same amount of writes to pagetable as we have >>>>> before HDBSS, both on faulting and 1st iteration cleaning) >>>>> >>>> Hi, Leo >>>> >>>> Actually, I have thought about this approach too, but if we clear DBM in >>>> kvm_pgtable_stage2_wrprotect(), then during the first round of >>>> migration, we will fault and release RO -> W, and then add DBM. >>> >>> Yeah, that's only for lazy-splitting, though. >>> >>>> >>>> But next time, when we migrate the dirty pages in round two, we will run >>>> kvm_pgtable_stage2_wrprotect() again, which will clear DBM again. And >>>> finally, HDBSS will be useless during migration. >>> >>> Right, on lazy splitting, we have to clean the DBM bit on the >>> write-protect only if it's a block entry (hugepage). >>> >>> Once it faults for the first time, it will lazy-split, and we don't need to >>> clean the DBM bit. >>> >>>> >>>>> >>>>>>> However, I ran a quick test on a 400GB VM (4 vCPUs), and the overhead >>>>>>> turned out to be around 30ns — which I think is acceptable. >>>>>> >>>>>> Just a quick correction — I misstated the unit in my previous email. The >>>>>> overhead for clearing DBM on the 400GB VM (4 vCPUs) was around 32 µs, not 30 >>>>>> ns. >>>>>> >>>>> >>>>> Oh, that seems more likely :) >>>>> >>>>> Question: is tha above amount of memory initially in Level-1 blocks, >>>>> level-2 blocks or level-3 pages? (aka: were you using explicit/transparent >>>>> hugepages?) >>>>> >>>> >>>> I'm using transparent hugepages. However, if we were to use level-3 stage-2 >>>> pages with -mem-prealloc enabled in QEMU, I believe the time cost would be >>>> extremely high — potentially out of our control. >>>> >>> >>> Yeah, that's the issue. >>> For this not to explode like this, we need to mark as RO only when the >>> entries are blocks AND we are doing lazy splitting. >>> >>> We have: >>> Mode DBM Dirty bit >>> RO 0 X >>> WC 1 0 >>> WD 1 1 >>> >>> On write-protect: >>> - Lazy splitting + block entry (hugepage, level 2-) -> RO >>> - Otherwise -> WC >>> >>> On first fault, the block entry will be lazy-splitten, and we can set DBM=1 >>> in every new page. >>> >>> That way we guarantee that we are not faulting level-3 pages unecessarily, >>> nor need to go through the whole tree setting DBM=1 or DBM=0 on level-3 >>> pages. >>> >>> How does that sound? >>> >>> Thanks! >>> Leo >>> > >> >> Hi Leo, >> >> I've also been thinking about this approach: clear >> KVM_PTE_LEAF_ATTR_HI_S2_DBM when kvm_pgtable_stage2_wrprotect() calls >> stage2_update_leaf_attrs(). And we can check whether a page is a block >> page during the page walk, right? > > Hi Tian, > That was what I was thinking :) > >> >> So we can check the page level in the walker callback >> stage2_attr_walker(), filter there, clear DBM for block pages and >> preserve DBM on level-3 pages. Something like this: >> >> ``` >> pte &= ~data->attr_clr; // wrprotect: clears S2AP_W only >> pte |= data->attr_set; >> if (ctx->level < KVM_PGTABLE_LAST_LEVEL) > > Only on lazy splitting, right? > > Or maybe we get the DBM bit on during eager splitting... > >> pte &= ~KVM_PTE_LEAF_ATTR_HI_S2_DBM; // strip DBM from blocks >> ``` >> >> My only concern is whether this breaks Oliver's model of treating DBM as >> the write permission bit > > That was my point in my previous message, we are not breaking Oliver's > model. IIUC in the new model/encoding, we have: > > Mode DBM Dirty bit > RO 0 X > WC 1 0 > WD 1 1 > > Let's not think about individual bits for now: > Blocks (hugepages) are marked RO, pages are marked WC > > We do that so the blocks can be faulted, split and the resulting pages can > be marked WC. (except the page written to, which gets WD) > > That means we are following the new encoding, but deciding to WC/RO > depending on the context. For dirty-logging both can be used to mark a page > that is not dirty, so we should be fine. > >> — though DBM will be set back to 1 once the >> block is split into level-3 pages. So the inconsistency is temporary. > > No inconsistency for now :) > > What do you think? > > Thanks! > Leo > Right, totally agree. Thanks! Tian