Linux-ARM-Kernel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Tian Zheng <zhengtian10@huawei.com>
To: Leonardo Bras <leo.bras@arm.com>
Cc: Oliver Upton <oupton@kernel.org>, <maz@kernel.org>,
	<catalin.marinas@arm.com>, <will@kernel.org>,
	<yuzenghui@huawei.com>, <wangzhou1@hisilicon.com>,
	<yangjinqian1@huawei.com>, <caijian11@h-partners.com>,
	<liuyonglong@huawei.com>, <yezhenyu2@huawei.com>,
	<yubihong@huawei.com>, <linuxarm@huawei.com>,
	<joey.gouly@arm.com>, <kvmarm@lists.linux.dev>,
	<kvm@vger.kernel.org>, <linux-arm-kernel@lists.infradead.org>,
	<linux-kernel@vger.kernel.org>, <seiden@linux.ibm.com>,
	<suzuki.poulose@arm.com>
Subject: Re: [PATCH v4 3/6] KVM: arm64: Add auto DBM support for hardware dirty tracking
Date: Mon, 3 Aug 2026 09:33:02 +0800	[thread overview]
Message-ID: <8e36e2c8-587f-4228-ab13-d6783927a281@huawei.com> (raw)
In-Reply-To: <amoZY8uGLf1PEpjK@LeoBrasDK>



On 7/29/2026 11:16 PM, Leonardo Bras wrote:
> On Wed, Jul 29, 2026 at 04:51:59PM +0800, Tian Zheng wrote:
>>
>>
>> On 7/20/2026 8:58 PM, Leonardo Bras wrote:
>>> On Fri, Jul 17, 2026 at 04:21:32PM +0100, Leonardo Bras wrote:
>>>> On Fri, Jul 17, 2026 at 11:58:06AM +0800, Tian Zheng wrote:
>>>>>
>>>>> On 7/16/2026 3:39 PM, Oliver Upton wrote:
>>>>>> Hi Tian,
>>>>>>
>>>>>> On Thu, Jul 09, 2026 at 06:40:23PM +0800, Tian Zheng wrote:
>>>>>>> -	if (prot & KVM_PGTABLE_PROT_W)
>>>>>>> +	if (prot & KVM_PGTABLE_PROT_W) {
>>>>>>>     		set |= KVM_PTE_LEAF_ATTR_LO_S2_S2AP_W;
>>>>>>>
>>>>>>> +		/*
>>>>>>> +		 * No DEVICE filter needed here: relax_perms is only called
>>>>>>> +		 * on FSC_PERM faults. Device pages always get full RW from
>>>>>>> +		 * initial mapping and are never write-protected during
>>>>>>> +		 * migration, so they never trigger a permission fault.
>>>>>>> +		 */
>>>>>>> +		if (pgt->flags & KVM_PGTABLE_S2_DBM)
>>>>>>> +			set |= KVM_PTE_LEAF_ATTR_HI_S2_DBM;
>>>>>>> +	} else {
>>>>>>> +		/*
>>>>>>> +		 * Clear DBM on W→RO downgrade to prevent hardware from
>>>>>>> +		 * silently upgrading RO+DBM back to W+dirty, which would
>>>>>>> +		 * bypass KVM's write tracking and cause data corruption.
>>>>>>> +		 */
>>>>>>> +		clr |= KVM_PTE_LEAF_ATTR_HI_S2_DBM;
>>>>>>> +	}
>>>>>>> +
>>>>>> This block makes it pretty evident that the DBM bit really *is* the
>>>>>> write permission bit. I'd much rather we introduce the concept of dirty
>>>>>> state to the page table library and migrate the abstract write
>>>>>> permission to the DBM field, even if we don't have FEAT_HAFDBS.
>>>>>>
>>>>
>>>> Ohh, that's an amazing idea!
>>>
>>> Thinking about that again...
>>> If we adopt the encoding with DBM being the write-permission bit, and all
>>> PTEs have it since the start, how can we have lazy-splitting happening?
>>>
>>> Only way I think of is removing both DBM and S2_S2AP_W bit from writable
>>> PTEs during dirty-track enable, and re-adding them during the first write
>>> fault. If we don't remove the DBM bit, systems with HDBSS would just dirty
>>> it by hardware, without causing a fault.
>>>
>>> DBM=0 would need to happen only in the first write-protect (only on
>>> lazy-splitting). All other write-protecting would just clean the S2_S2AP_W
>>> bit, as everything is already split.
>>>
>>> Is that what was intended?
>>>
>>> Thanks!
>>> Leo
>>>
>> Hi Leo,
>>
>> I think the cleanest way to handle this is to simply avoid setting DBM
>> on block mappings. If we only set DBM on page-level PTEs, then block
>> mappings will naturally stay DBM=0 and trigger a write fault on first
>> access — exactly what we need for lazy splitting.
>>
>> When the fault occurs, the block gets split into page-level PTEs, and at
>> that point we can set DBM=1 on the resulting leaf entries. This way:
>>
>> 1. Lazy split works naturally (fault -> split -> set DBM=1)
>>
>> 2. No need to clear DBM globally at dirty-track enable
>>
>> 3. No special handling for block mappings
>>
>> So I think global DBM is still viable — we just need to filter out block
>> mappings when setting the DBM bit. That way the lazy split path is preserved
>> without extra complexity.
> 
> Hi Tian,
> 
> Humm, but would not that be contrary to what Oliver suggested: changing the
> encoding from the PTE for all entries?
> 
> (Like, if the PTE is writable, it has to have DBM set)
> 
> IIUC what you said, on first faulting of the page in the VM:
> - If the entry is a page (level-3 leaf) and writable, add DBM
> - If it's a block entry (leaf but not a level-3), don't add DBM
> 
> So after we enable dirty-logging:
> - a level-3 entry would not fault, using HDBSS, and
> - a block entry would fault, do the splitting, and add DBM to level-3
>    entries during the split.
> 
> If I got that correct, that would be clean indeed.
> 
> But then we would have a different encoding for block entries and page
> entries. In page entries, DBM could be used to say if the page is writable,
> but on block entries one would have to look at the 'dirty-bit'.
> 
> Would that be ok?
> 
> Thanks!
> Leo
> 
Hi Leo,

My initial concern was that clearing all DBM bits at the start of
migration would be too expensive, so I thought distinguishing between
level-3 entries and block entries would be better.

However, I ran a quick test on a 400GB VM (4 vCPUs), and the overhead
turned out to be around 30ns — which I think is acceptable.

So I think we can go with your approach: simply clear DBM globally in
kvm_arch_commit_memory_region() when dirty logging starts, before write-
protecting the memslot.

```
void kvm_arch_commit_memory_region(...)
{
     // ...
     if (log_dirty_pages) {
         if (change == KVM_MR_DELETE)
             return;

         kvm_mmu_clear_dbm_memory_region(kvm, new->id);
         // ...
     }
     // ...
}
```

Thanks!
Tian


  reply	other threads:[~2026-08-03  1:33 UTC|newest]

Thread overview: 71+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-09 10:40 [PATCH v4 0/6] Support the FEAT_HDBSS introduced in Armv9.5 Tian Zheng
2026-07-09 10:40 ` [PATCH v4 1/6] KVM: arm64: Enable eager hugepage splitting if HDBSS is available Tian Zheng
2026-07-09 10:40 ` [PATCH v4 2/6] KVM: arm64: Add support for FEAT_HDBSS Tian Zheng
2026-07-09 10:40 ` [PATCH v4 3/6] KVM: arm64: Add auto DBM support for hardware dirty tracking Tian Zheng
2026-07-13 11:17   ` Leonardo Bras
2026-07-14  1:14     ` Tian Zheng
2026-07-14  7:23       ` Marc Zyngier
2026-07-14  7:44         ` Tian Zheng
2026-07-14 10:20           ` Leonardo Bras
2026-07-16  7:39   ` Oliver Upton
2026-07-17  3:58     ` Tian Zheng
2026-07-17 15:21       ` Leonardo Bras
2026-07-20 12:58         ` Leonardo Bras
2026-07-29  8:51           ` Tian Zheng
2026-07-29 15:16             ` Leonardo Bras
2026-08-03  1:33               ` Tian Zheng [this message]
2026-08-03  4:04                 ` Tian Zheng
2026-08-03 10:21                   ` Leonardo Bras
2026-08-03 13:57                     ` Tian Zheng
2026-08-03 16:32                       ` Leonardo Bras
2026-08-03 16:34                         ` Leonardo Bras
2026-08-04  4:54                         ` Tian Zheng
2026-08-04 11:10                           ` Leonardo Bras
2026-08-05  3:41                             ` Tian Zheng
2026-08-05  3:43                             ` Tian Zheng
2026-07-28  8:49         ` Tian Zheng
2026-07-09 10:40 ` [PATCH v4 4/6] KVM: arm64: Add HDBSS per-vCPU buffer management Tian Zheng
2026-07-13 13:39   ` Leonardo Bras
2026-07-14  7:15     ` Tian Zheng
2026-07-14 10:47       ` Leonardo Bras
2026-07-15  9:16         ` Tian Zheng
2026-07-15 14:28           ` Leonardo Bras
2026-07-17  4:06             ` Tian Zheng
2026-07-09 10:40 ` [PATCH v4 5/6] KVM: arm64: Add HDBSS fault handling and buffer flush Tian Zheng
2026-07-13 14:06   ` Leonardo Bras
2026-07-14  7:38     ` Tian Zheng
2026-07-14 10:50       ` Leonardo Bras
2026-07-14 13:27         ` Tian Zheng
2026-07-14 14:19           ` Leonardo Bras
2026-07-17  6:51             ` Tian Zheng
2026-07-17 15:44               ` Leonardo Bras
2026-07-21  8:53                 ` Inochi Amaoto
2026-07-21 14:18                   ` Leonardo Bras
2026-07-22  5:14                     ` Inochi Amaoto
2026-07-22 11:04                       ` Leonardo Bras
2026-07-23  1:17                         ` Inochi Amaoto
2026-07-28  7:52                     ` Tian Zheng
2026-07-29 15:30                       ` Leonardo Bras
2026-08-03  3:15                         ` Tian Zheng
2026-08-03 10:43                           ` Leonardo Bras
2026-08-04  8:26                             ` Tian Zheng
2026-08-03  3:22                 ` Tian Zheng
2026-08-03 10:24                   ` Leonardo Bras
2026-07-09 10:40 ` [PATCH v4 6/6] KVM: arm64: Add auto HDBSS enable/disable on dirty logging change Tian Zheng
2026-07-13 14:50   ` Leonardo Bras
2026-07-14  8:58     ` Tian Zheng
2026-07-14 11:16       ` Leonardo Bras
2026-07-14 14:33         ` Leonardo Bras
2026-07-16  8:37           ` Tian Zheng
2026-07-17  7:23         ` Tian Zheng
2026-07-17 15:50           ` Leonardo Bras
2026-07-16  7:15     ` Tian Zheng
2026-07-17 15:53       ` Leonardo Bras
2026-07-13 10:31 ` [PATCH v4 0/6] Support the FEAT_HDBSS introduced in Armv9.5 Leonardo Bras
2026-07-13 16:27   ` Leonardo Bras
2026-07-14 10:39     ` Tian Zheng
2026-07-14 11:20       ` Leonardo Bras
2026-07-14 13:29         ` Tian Zheng
2026-07-14  9:37   ` Tian Zheng
2026-07-14 10:19     ` Leonardo Bras
2026-07-14 13:34       ` Tian Zheng

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=8e36e2c8-587f-4228-ab13-d6783927a281@huawei.com \
    --to=zhengtian10@huawei.com \
    --cc=caijian11@h-partners.com \
    --cc=catalin.marinas@arm.com \
    --cc=joey.gouly@arm.com \
    --cc=kvm@vger.kernel.org \
    --cc=kvmarm@lists.linux.dev \
    --cc=leo.bras@arm.com \
    --cc=linux-arm-kernel@lists.infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linuxarm@huawei.com \
    --cc=liuyonglong@huawei.com \
    --cc=maz@kernel.org \
    --cc=oupton@kernel.org \
    --cc=seiden@linux.ibm.com \
    --cc=suzuki.poulose@arm.com \
    --cc=wangzhou1@hisilicon.com \
    --cc=will@kernel.org \
    --cc=yangjinqian1@huawei.com \
    --cc=yezhenyu2@huawei.com \
    --cc=yubihong@huawei.com \
    --cc=yuzenghui@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox