From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mta0.migadu.com (out-16.mta0.migadu.com [91.218.175.16]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 36A0D368D7E for ; Wed, 30 Sep 2026 03:36:39 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.16 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790739402; cv=none; b=VFeotowBdJMJhY8q85r0baGm2sNjCG56NL54Vnxb+A9PLn5S1tYSQW9KjNjvGDBt9lnY9K6crE7XFMCJsMGmurOZW+6byzSONDDyS7zDzxaddqj3dVdGDi679S/fzNigW1fYm3kd4/nOxBjbBhL/QTZzwEK6Kt4FJuHDQxGzXYA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790739402; c=relaxed/simple; bh=UGu6hnp7zOzIji6/SbzCysLcwMM7yV0BdKnkgtzfNIQ=; h=Content-Type:Mime-Version:Subject:From:In-Reply-To:Date:Cc: Message-Id:References:To; b=m5qhkW6FohNykscqmsVWe4MV/Y+q09nzDH70fkU2C8Tuf/tbFIXyN0WwnUSQqMfZU5mcwJsx1hwx5dsVd6hEa908awYR1PHSZi/PREVIVpiTdJXG1kKNA9MrlCwxG5LcYSpPDRw8gzykY59RNslyj7U/L/yS3F4jojTBG6oaEko= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=VvU+1/Vn; arc=none smtp.client-ip=91.218.175.16 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="VvU+1/Vn" X-Envelope-To: damon@lists.linux.dev DKIM-Signature: a=rsa-sha256; bh=UGu6hnp7zOzIji6/SbzCysLcwMM7yV0BdKnkgtzfNIQ=; c=simple/simple; d=linux.dev; h=from:to:subject:date:message-id:mime-version:content-type; s=key1; t=1790739398; v=1; x=1791344198; b=VvU+1/VnbPIh/soOlHQ3yAjfFdjnE05G4a6JM3kzNVkS3xXrPEIeaBD61771dm0Zmp234NjV 9qdtKPRd8uIc3Xu+mbZMgaDYyVXOWS3g0DEZ8f4ZnI50BKpzZ1EhAqxgbGEL5e7pMroJuex7rH/ nnA9j4+F/DeUPIBYsHmqYrBM= X-Envelope-To: damon@lists.linux.dev Received: by mta10.migadu.com with ESMTPS id 9bc00279546608ff; Wed, 30 Sep 2026 03:36:37 +0000 X-Mizu-Trace-ID: 9bc00279546608ff X-Migadu-Flow: FLOW_OUT Content-Type: text/plain; charset=us-ascii Precedence: bulk X-Mailing-List: damon@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 (Mac OS X Mail 16.0 \(3901.100.1.1.11\)) Subject: Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper From: Muchun Song In-Reply-To: Date: Wed, 30 Sep 2026 11:36:23 +0800 Cc: Jialiang Huang , lance.yang@linux.dev, baohua@kernel.org, damon@lists.linux.dev, david@kernel.org, kunwu.chan@gmail.com, lianux.mm@gmail.com, linux-mm@kvack.org, mst@redhat.com, ryncsn@gmail.com, sj@kernel.org, virtualization@lists.linux.dev, xueyuan.chen21@gmail.com Content-Transfer-Encoding: 7bit Message-Id: References: <20260925054408.10431-1-lance.yang@linux.dev> <20260929123241.1408414-1-huang-jl@deepseek.com> To: Gao Xiang X-Mailer: Apple Mail (2.3901.100.1.1.11) > On Sep 29, 2026, at 20:41, Gao Xiang wrote: > > On Tue, Sep 29, 2026 at 08:32:41PM +0800, Jialiang Huang wrote: >> Hi all, >> >> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks >> to everyone working on DAMON and virtio-balloon free-page reporting. >> They have been very useful for our workloads. >> >> Gao Xiang wrote: >>> It can cause sync 4K faults on the host in the worst case >> >> This is one of our concerns with virtio-pmem as well: moving I/O onto >> the page-fault path can introduce performance trade-offs. The other >> concern is the substantial struct page overhead for large images. >> >> For now, we enable virtio-pmem only for moderately sized, frequently >> used read-only images, where there is more opportunity to share the >> same host page cache across sandboxes, as Gao pointed out. >> >> Muchun's vmemmap work is also interesting to us. My understanding is >> that it allocates private backing for struct page metadata on demand, >> which could help reduce the upfront memory overhead for large images. > > Although I haven't had a chance and time to look into that, the main > concern from me is that mmap() access will call > "dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox > workloads (or not malicious, just valid mmap workloads) can cause guest > memory OOMs due to "struct page balloon" for large rootfs in the worst > cases and cause the follow-up mmap access failure, because the guest > memory size may not even fulfill "struct page" for large rootfs. I agree that this is a real issue with v1. One detail is that merely establishing the mapping does not materialize the metadata. Materialization happens when a fault resolves to an allocated DAX extent, before its PFN is inserted into a userspace mapping. However, that distinction does not remove the problem. In the worst case, a workload can fault enough of the pmem range to restore the full vmemmap cost, about 1.56% of the pmem size. Since v1 does not dematerialize private vmemmap pages, even a one-time scan can retain that cost until the device is removed. The follow-up mentioned in the cover letter is intended to make the optimization reversible. One possible direction would be to invalidate clean DAX entries under memory pressure and zap their userspace mappings. Once a DAX entry has been removed and the corresponding PFNs have no remaining mappings, references, or pins that require private metadata, the associated vmemmap backing could be remapped to the shared read-only page. A later access would fault and materialize it again. This would address the persistent ballooning caused by a one-time scan, but it would not help while a large DAX working set remains actively used or pinned. As a separate policy mechanism for that case, one possible direction would be to account the faulted pmem range to the faulting mm's memory cgroup: PAGE_SIZE for a PTE mapping and PMD_SIZE for a PMD mapping. This would intentionally account the mapped DAX capacity. Such accounting could be explicitly enabled through a cgroup v2 mount option, following the opt-in model of memory_hugetlb_accounting, so the existing FS-DAX accounting semantics would remain unchanged by default. With such accounting, an active or pinned DAX working set could still reach its cgroup limit, but it could not grow private vmemmap without corresponding cgroup usage. This is only a possible direction for bounding the peak working set, though, and I have not worked through all of its implementation details yet. Thanks, Muchun > > Thanks, > Gao Xiang > >> >> For disks without virtio-pmem, DAMON with virtio-balloon free-page >> reporting lets us reclaim cold guest page-cache pages and return >> the memory to the host, without those virtio-pmem-specific issues. >> The two approaches complement each other in our setup. >> >> We have not yet fully explored how best to tune the DAMON and free-page >> reporting parameters for our workloads. For example, the kernel's default >> free-page reporting granularity is 2 MiB, which is fairly coarse: >> reclaiming cold pages does not necessarily produce free blocks of that >> size. We still need to evaluate how finer reporting granularity and >> different DAMON settings affect memory savings and workload performance. >> >> Best, >> Jialiang Huang > >