Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Muchun Song <muchun.song@linux.dev>
To: Gao Xiang <xiang@kernel.org>
Cc: Jialiang Huang <huang-jl@deepseek.com>,
	lance.yang@linux.dev, baohua@kernel.org, damon@lists.linux.dev,
	david@kernel.org, kunwu.chan@gmail.com, lianux.mm@gmail.com,
	linux-mm@kvack.org, mst@redhat.com, ryncsn@gmail.com,
	sj@kernel.org, virtualization@lists.linux.dev,
	xueyuan.chen21@gmail.com
Subject: Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
Date: Wed, 30 Sep 2026 11:36:23 +0800	[thread overview]
Message-ID: <AA71E56E-F29A-44B2-B5A0-7806CD684B21@linux.dev> (raw)
In-Reply-To: <aruyFPzEe3pS_Qzx@XiangdeMacBook-Pro.local>



> On Sep 29, 2026, at 20:41, Gao Xiang <xiang@kernel.org> wrote:
> 
> On Tue, Sep 29, 2026 at 08:32:41PM +0800, Jialiang Huang wrote:
>> Hi all,
>> 
>> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
>> to everyone working on DAMON and virtio-balloon free-page reporting.
>> They have been very useful for our workloads.
>> 
>> Gao Xiang wrote:
>>> It can cause sync 4K faults on the host in the worst case
>> 
>> This is one of our concerns with virtio-pmem as well: moving I/O onto
>> the page-fault path can introduce performance trade-offs. The other
>> concern is the substantial struct page overhead for large images.
>> 
>> For now, we enable virtio-pmem only for moderately sized, frequently
>> used read-only images, where there is more opportunity to share the
>> same host page cache across sandboxes, as Gao pointed out.
>> 
>> Muchun's vmemmap work is also interesting to us. My understanding is
>> that it allocates private backing for struct page metadata on demand,
>> which could help reduce the upfront memory overhead for large images.
> 
> Although I haven't had a chance and time to look into that, the main
> concern from me is that mmap() access will call
> "dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox
> workloads (or not malicious, just valid mmap workloads) can cause guest
> memory OOMs due to "struct page balloon" for large rootfs in the worst
> cases and cause the follow-up mmap access failure, because the guest
> memory size may not even fulfill "struct page" for large rootfs.

I agree that this is a real issue with v1.

One detail is that merely establishing the mapping does not materialize
the metadata. Materialization happens when a fault resolves to an
allocated DAX extent, before its PFN is inserted into a userspace
mapping. However, that distinction does not remove the problem.

In the worst case, a workload can fault enough of the pmem range to
restore the full vmemmap cost, about 1.56% of the pmem size. Since v1
does not dematerialize private vmemmap pages, even a one-time scan can
retain that cost until the device is removed.

The follow-up mentioned in the cover letter is intended to make the
optimization reversible. One possible direction would be to invalidate
clean DAX entries under memory pressure and zap their userspace
mappings. Once a DAX entry has been removed and the corresponding PFNs
have no remaining mappings, references, or pins that require private
metadata, the associated vmemmap backing could be remapped to the
shared read-only page. A later access would fault and materialize it
again.

This would address the persistent ballooning caused by a one-time scan,
but it would not help while a large DAX working set remains actively
used or pinned.

As a separate policy mechanism for that case, one possible direction
would be to account the faulted pmem range to the faulting mm's memory
cgroup: PAGE_SIZE for a PTE mapping and PMD_SIZE for a PMD mapping.
This would intentionally account the mapped DAX capacity.

Such accounting could be explicitly enabled through a cgroup v2 mount
option, following the opt-in model of memory_hugetlb_accounting, so the
existing FS-DAX accounting semantics would remain unchanged by default.

With such accounting, an active or pinned DAX working set could still
reach its cgroup limit, but it could not grow private vmemmap without
corresponding cgroup usage. This is only a possible direction for
bounding the peak working set, though, and I have not worked through
all of its implementation details yet.

Thanks,
Muchun

> 
> Thanks,
> Gao Xiang
> 
>> 
>> For disks without virtio-pmem, DAMON with virtio-balloon free-page
>> reporting lets us reclaim cold guest page-cache pages and return
>> the memory to the host, without those virtio-pmem-specific issues.
>> The two approaches complement each other in our setup.
>> 
>> We have not yet fully explored how best to tune the DAMON and free-page
>> reporting parameters for our workloads. For example, the kernel's default
>> free-page reporting granularity is 2 MiB, which is fairly coarse:
>> reclaiming cold pages does not necessarily produce free blocks of that
>> size. We still need to evaluate how finer reporting granularity and
>> different DAMON settings affect memory savings and workload performance.
>> 
>> Best,
>> Jialiang Huang
> 
> 



  parent reply	other threads:[~2026-09-30  3:36 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25  5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
2026-09-25  5:57 ` Lian Wang
2026-09-25  6:46   ` KunWu Chan
2026-09-25  7:04 ` David Hildenbrand (Arm)
2026-09-25  7:28   ` Lian Wang
2026-09-25  8:15   ` Lance Yang
2026-09-25 10:03     ` David Hildenbrand (Arm)
2026-09-25 10:16     ` Gao Xiang
2026-09-25 10:27       ` Gao Xiang
2026-09-25 10:13 ` SJ Park
2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41   ` Gao Xiang
2026-09-29 12:50     ` Jialiang Huang
2026-09-30  3:36     ` Muchun Song [this message]
2026-09-30  7:24       ` Gao Xiang
2026-09-30  9:37         ` Muchun Song
2026-09-30 10:33           ` Gao Xiang
2026-09-29 16:56   ` SJ Park
2026-09-30  3:07     ` Lian Wang
2026-09-30  8:08       ` SJ Park
2026-09-29 18:14   ` Pratyush Mallick

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=AA71E56E-F29A-44B2-B5A0-7806CD684B21@linux.dev \
    --to=muchun.song@linux.dev \
    --cc=baohua@kernel.org \
    --cc=damon@lists.linux.dev \
    --cc=david@kernel.org \
    --cc=huang-jl@deepseek.com \
    --cc=kunwu.chan@gmail.com \
    --cc=lance.yang@linux.dev \
    --cc=lianux.mm@gmail.com \
    --cc=linux-mm@kvack.org \
    --cc=mst@redhat.com \
    --cc=ryncsn@gmail.com \
    --cc=sj@kernel.org \
    --cc=virtualization@lists.linux.dev \
    --cc=xiang@kernel.org \
    --cc=xueyuan.chen21@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox