Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Muchun Song <muchun.song@linux.dev>
To: Gao Xiang <xiang@kernel.org>
Cc: Jialiang Huang <huang-jl@deepseek.com>,
	lance.yang@linux.dev, baohua@kernel.org, damon@lists.linux.dev,
	david@kernel.org, kunwu.chan@gmail.com, lianux.mm@gmail.com,
	linux-mm@kvack.org, mst@redhat.com, ryncsn@gmail.com,
	sj@kernel.org, virtualization@lists.linux.dev,
	xueyuan.chen21@gmail.com
Subject: Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
Date: Wed, 30 Sep 2026 17:37:11 +0800	[thread overview]
Message-ID: <CEE47F27-874B-486A-995F-E223249659FE@linux.dev> (raw)
In-Reply-To: <ary5Nn9le4xG7_-N@MacBookPro>



> On Sep 30, 2026, at 15:24, Gao Xiang <xiang@kernel.org> wrote:
> 
> On Wed, Sep 30, 2026 at 11:36:23AM +0800, Muchun Song wrote:
>> 
>> 
>>> On Sep 29, 2026, at 20:41, Gao Xiang <xiang@kernel.org> wrote:
>>> 
>>> On Tue, Sep 29, 2026 at 08:32:41PM +0800, Jialiang Huang wrote:
>>>> Hi all,
>>>> 
>>>> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
>>>> to everyone working on DAMON and virtio-balloon free-page reporting.
>>>> They have been very useful for our workloads.
>>>> 
>>>> Gao Xiang wrote:
>>>>> It can cause sync 4K faults on the host in the worst case
>>>> 
>>>> This is one of our concerns with virtio-pmem as well: moving I/O onto
>>>> the page-fault path can introduce performance trade-offs. The other
>>>> concern is the substantial struct page overhead for large images.
>>>> 
>>>> For now, we enable virtio-pmem only for moderately sized, frequently
>>>> used read-only images, where there is more opportunity to share the
>>>> same host page cache across sandboxes, as Gao pointed out.
>>>> 
>>>> Muchun's vmemmap work is also interesting to us. My understanding is
>>>> that it allocates private backing for struct page metadata on demand,
>>>> which could help reduce the upfront memory overhead for large images.
>>> 
>>> Although I haven't had a chance and time to look into that, the main
>>> concern from me is that mmap() access will call
>>> "dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox
>>> workloads (or not malicious, just valid mmap workloads) can cause guest
>>> memory OOMs due to "struct page balloon" for large rootfs in the worst
>>> cases and cause the follow-up mmap access failure, because the guest
>>> memory size may not even fulfill "struct page" for large rootfs.
>> 
>> I agree that this is a real issue with v1.
>> 
>> One detail is that merely establishing the mapping does not materialize
>> the metadata. Materialization happens when a fault resolves to an
>> allocated DAX extent, before its PFN is inserted into a userspace
>> mapping. However, that distinction does not remove the problem.
>> 
>> In the worst case, a workload can fault enough of the pmem range to
>> restore the full vmemmap cost, about 1.56% of the pmem size. Since v1
>> does not dematerialize private vmemmap pages, even a one-time scan can
>> retain that cost until the device is removed.
>> 
>> The follow-up mentioned in the cover letter is intended to make the
>> optimization reversible. One possible direction would be to invalidate
>> clean DAX entries under memory pressure and zap their userspace
>> mappings. Once a DAX entry has been removed and the corresponding PFNs
>> have no remaining mappings, references, or pins that require private
>> metadata, the associated vmemmap backing could be remapped to the
>> shared read-only page. A later access would fault and materialize it
>> again.
> 
> BTW, it's impossible for shared DAX entries (like the current XFS DAX
> with reflink and EROFS will support this feature later too for chunk
> memory sharing), since you cannot just use mapping and index to get
> the VMA like page cache does unless you invent another new mechanism
> for this.

You're right. For the reflink scenario, reclamation is currently
difficult. If we want to reclaim, it would also be in three stages:
1) reclaim the struct page corresponding to PFNs that no longer have
any mapping; 2) reclaim the cases where mappings exist but are not
reflink; 3) reclaim the reflink scenario.

These three stages go from simple to difficult. Of course, I hadn't
thought this far ahead before. So in the v1 version, not even the
first stage was implemented. At the very least, before I act, I need
enough planning and thought.

> 
> Reclaiming page entry mechanism seems it can be used for or overlapped
> to another types of memory (in order to save struct page memory in
> general): I'm not sure if it needs wider discussion on reclaiming
> "struct page" in general first.

Of course, I don't think struct page saving is the focus of the discussion
here, so we can stop discussing it.

> 
> Anyway, it'd be better to get some numbers with RL or agent workloads
> (especially the host memory is under reasonable pressure) before
> landing all these new infras upstream if proceeding in this way.

At least for now, as far as I'm concerned, I don't intend to land all the
features mentioned here. The first thing I want to address is on-demand
allocation of struct page.

Thanks,
Muchun

> 
> Thanks,
> Gao Xiang




  reply	other threads:[~2026-09-30  9:37 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25  5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
2026-09-25  5:57 ` Lian Wang
2026-09-25  6:46   ` KunWu Chan
2026-09-25  7:04 ` David Hildenbrand (Arm)
2026-09-25  7:28   ` Lian Wang
2026-09-25  8:15   ` Lance Yang
2026-09-25 10:03     ` David Hildenbrand (Arm)
2026-09-25 10:16     ` Gao Xiang
2026-09-25 10:27       ` Gao Xiang
2026-09-25 10:13 ` SJ Park
2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41   ` Gao Xiang
2026-09-29 12:50     ` Jialiang Huang
2026-09-30  3:36     ` Muchun Song
2026-09-30  7:24       ` Gao Xiang
2026-09-30  9:37         ` Muchun Song [this message]
2026-09-30 10:33           ` Gao Xiang
2026-09-29 16:56   ` SJ Park
2026-09-30  3:07     ` Lian Wang
2026-09-30  8:08       ` SJ Park
2026-09-29 18:14   ` Pratyush Mallick

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=CEE47F27-874B-486A-995F-E223249659FE@linux.dev \
    --to=muchun.song@linux.dev \
    --cc=baohua@kernel.org \
    --cc=damon@lists.linux.dev \
    --cc=david@kernel.org \
    --cc=huang-jl@deepseek.com \
    --cc=kunwu.chan@gmail.com \
    --cc=lance.yang@linux.dev \
    --cc=lianux.mm@gmail.com \
    --cc=linux-mm@kvack.org \
    --cc=mst@redhat.com \
    --cc=ryncsn@gmail.com \
    --cc=sj@kernel.org \
    --cc=virtualization@lists.linux.dev \
    --cc=xiang@kernel.org \
    --cc=xueyuan.chen21@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox