From: Muchun Song <muchun.song@linux.dev>
To: Gao Xiang <xiang@kernel.org>
Cc: Jialiang Huang <huang-jl@deepseek.com>,
lance.yang@linux.dev, baohua@kernel.org, damon@lists.linux.dev,
david@kernel.org, kunwu.chan@gmail.com, lianux.mm@gmail.com,
linux-mm@kvack.org, mst@redhat.com, ryncsn@gmail.com,
sj@kernel.org, virtualization@lists.linux.dev,
xueyuan.chen21@gmail.com
Subject: Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
Date: Wed, 30 Sep 2026 17:37:11 +0800 [thread overview]
Message-ID: <CEE47F27-874B-486A-995F-E223249659FE@linux.dev> (raw)
In-Reply-To: <ary5Nn9le4xG7_-N@MacBookPro>
> On Sep 30, 2026, at 15:24, Gao Xiang <xiang@kernel.org> wrote:
>
> On Wed, Sep 30, 2026 at 11:36:23AM +0800, Muchun Song wrote:
>>
>>
>>> On Sep 29, 2026, at 20:41, Gao Xiang <xiang@kernel.org> wrote:
>>>
>>> On Tue, Sep 29, 2026 at 08:32:41PM +0800, Jialiang Huang wrote:
>>>> Hi all,
>>>>
>>>> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
>>>> to everyone working on DAMON and virtio-balloon free-page reporting.
>>>> They have been very useful for our workloads.
>>>>
>>>> Gao Xiang wrote:
>>>>> It can cause sync 4K faults on the host in the worst case
>>>>
>>>> This is one of our concerns with virtio-pmem as well: moving I/O onto
>>>> the page-fault path can introduce performance trade-offs. The other
>>>> concern is the substantial struct page overhead for large images.
>>>>
>>>> For now, we enable virtio-pmem only for moderately sized, frequently
>>>> used read-only images, where there is more opportunity to share the
>>>> same host page cache across sandboxes, as Gao pointed out.
>>>>
>>>> Muchun's vmemmap work is also interesting to us. My understanding is
>>>> that it allocates private backing for struct page metadata on demand,
>>>> which could help reduce the upfront memory overhead for large images.
>>>
>>> Although I haven't had a chance and time to look into that, the main
>>> concern from me is that mmap() access will call
>>> "dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox
>>> workloads (or not malicious, just valid mmap workloads) can cause guest
>>> memory OOMs due to "struct page balloon" for large rootfs in the worst
>>> cases and cause the follow-up mmap access failure, because the guest
>>> memory size may not even fulfill "struct page" for large rootfs.
>>
>> I agree that this is a real issue with v1.
>>
>> One detail is that merely establishing the mapping does not materialize
>> the metadata. Materialization happens when a fault resolves to an
>> allocated DAX extent, before its PFN is inserted into a userspace
>> mapping. However, that distinction does not remove the problem.
>>
>> In the worst case, a workload can fault enough of the pmem range to
>> restore the full vmemmap cost, about 1.56% of the pmem size. Since v1
>> does not dematerialize private vmemmap pages, even a one-time scan can
>> retain that cost until the device is removed.
>>
>> The follow-up mentioned in the cover letter is intended to make the
>> optimization reversible. One possible direction would be to invalidate
>> clean DAX entries under memory pressure and zap their userspace
>> mappings. Once a DAX entry has been removed and the corresponding PFNs
>> have no remaining mappings, references, or pins that require private
>> metadata, the associated vmemmap backing could be remapped to the
>> shared read-only page. A later access would fault and materialize it
>> again.
>
> BTW, it's impossible for shared DAX entries (like the current XFS DAX
> with reflink and EROFS will support this feature later too for chunk
> memory sharing), since you cannot just use mapping and index to get
> the VMA like page cache does unless you invent another new mechanism
> for this.
You're right. For the reflink scenario, reclamation is currently
difficult. If we want to reclaim, it would also be in three stages:
1) reclaim the struct page corresponding to PFNs that no longer have
any mapping; 2) reclaim the cases where mappings exist but are not
reflink; 3) reclaim the reflink scenario.
These three stages go from simple to difficult. Of course, I hadn't
thought this far ahead before. So in the v1 version, not even the
first stage was implemented. At the very least, before I act, I need
enough planning and thought.
>
> Reclaiming page entry mechanism seems it can be used for or overlapped
> to another types of memory (in order to save struct page memory in
> general): I'm not sure if it needs wider discussion on reclaiming
> "struct page" in general first.
Of course, I don't think struct page saving is the focus of the discussion
here, so we can stop discussing it.
>
> Anyway, it'd be better to get some numbers with RL or agent workloads
> (especially the host memory is under reasonable pressure) before
> landing all these new infras upstream if proceeding in this way.
At least for now, as far as I'm concerned, I don't intend to land all the
features mentioned here. The first thing I want to address is on-demand
allocation of struct page.
Thanks,
Muchun
>
> Thanks,
> Gao Xiang
next prev parent reply other threads:[~2026-09-30 9:37 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-25 5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
2026-09-25 5:57 ` Lian Wang
2026-09-25 6:46 ` KunWu Chan
2026-09-25 7:04 ` David Hildenbrand (Arm)
2026-09-25 7:28 ` Lian Wang
2026-09-25 8:15 ` Lance Yang
2026-09-25 10:03 ` David Hildenbrand (Arm)
2026-09-25 10:16 ` Gao Xiang
2026-09-25 10:27 ` Gao Xiang
2026-09-25 10:13 ` SJ Park
2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41 ` Gao Xiang
2026-09-29 12:50 ` Jialiang Huang
2026-09-30 3:36 ` Muchun Song
2026-09-30 7:24 ` Gao Xiang
2026-09-30 9:37 ` Muchun Song [this message]
2026-09-30 10:33 ` Gao Xiang
2026-09-29 16:56 ` SJ Park
2026-09-30 3:07 ` Lian Wang
2026-09-30 8:08 ` SJ Park
2026-09-29 18:14 ` Pratyush Mallick
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=CEE47F27-874B-486A-995F-E223249659FE@linux.dev \
--to=muchun.song@linux.dev \
--cc=baohua@kernel.org \
--cc=damon@lists.linux.dev \
--cc=david@kernel.org \
--cc=huang-jl@deepseek.com \
--cc=kunwu.chan@gmail.com \
--cc=lance.yang@linux.dev \
--cc=lianux.mm@gmail.com \
--cc=linux-mm@kvack.org \
--cc=mst@redhat.com \
--cc=ryncsn@gmail.com \
--cc=sj@kernel.org \
--cc=virtualization@lists.linux.dev \
--cc=xiang@kernel.org \
--cc=xueyuan.chen21@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox