* [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
@ 2026-09-25 5:44 Lance Yang
2026-09-25 5:57 ` Lian Wang
` (3 more replies)
0 siblings, 4 replies; 21+ messages in thread
From: Lance Yang @ 2026-09-25 5:44 UTC (permalink / raw)
To: sj, mst, david
Cc: damon, linux-mm, virtualization, ryncsn, kunwu.chan, lianux.mm,
baohua, xueyuan.chen21, Lance Yang
Hi all,
I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
and virtio-balloon:
With this kind of workload, an agent may read a file once and never touch
it again, while those pages remain in the guest page cache. Without memory
pressure in the guest, they can stay cached even though the host would
like that memory back ...
The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
cold file pages from the guest page cache; buddy gets a chance to coalesce
them into reportable blocks, and virtio-balloon passes those blocks to
Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
And it worked: time-integrated host memory use fell by 21.2%, while peak
usage stayed about the same, with no significant CPU overhead. Pretty cool
to see these pieces show up in another real, large-scale system :)
Kernel work still changes the world. Cheers to that!
Anyhow, I still don't know DAMON and virtio-balloon as well as I should
(mostly I just know the people who built them) :P
[1] https://arxiv.org/html/2609.22978
Cheers, Lance
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
@ 2026-09-25 5:57 ` Lian Wang
2026-09-25 6:46 ` KunWu Chan
2026-09-25 7:04 ` David Hildenbrand (Arm)
` (2 subsequent siblings)
3 siblings, 1 reply; 21+ messages in thread
From: Lian Wang @ 2026-09-25 5:57 UTC (permalink / raw)
To: Lance Yang
Cc: sj, mst, david, damon, linux-mm, virtualization, Kairui Song,
kunwu.chan, lianux.mm, baohua, xueyuan.chen21
Hi Lance,
Thank you very much for sharing this and for your interest in the related
work. It is great to see DAMON and virtio-balloon being used together in
another real, large-scale system.
One scenario we have recently been studying came from Sangfor. In a KVM/QEMU
deployment with guest memory backed by shared tmpfs and host THP enabled,
host-side DAMON can report a much larger hot-memory proportion than the
guest's fine-grained activity suggests. The related discussion is here:
https://lore.kernel.org/all/20260915025514.2434-1-lianux.mm@gmail.com/
We are preparing more complete experiments and improvements around this
scenario, and we plan to discuss it in more depth at LPC.
I am still actively running experiments based in part on SJ's earlier
introductions and our discussions. After LPC, we hope to share results on a
more regular basis, together with discussion outcomes and reports from more
scenarios. I believe DAMON will continue to become even better as these users
and perspectives come together.
I am also contacting the authors of the DeepSeek paper to learn more about
their setup and experience. We hope they can join the discussion so that we
can work together to improve this area further.
Kairui is also very interested in this DAMON use case and may already have
discussed related ideas with SJ, so I have kept him in Cc.
Thank you again, Lance, and thanks everyone!
Best,
Lian
On Fri, 25 Sep 2026 13:44:08 +0800 Lance Yang <lance.yang@linux.dev> wrote:
> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
> and virtio-balloon:
>
> And it worked: time-integrated host memory use fell by 21.2%, while peak
> usage stayed about the same, with no significant CPU overhead.
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 5:57 ` Lian Wang
@ 2026-09-25 6:46 ` KunWu Chan
0 siblings, 0 replies; 21+ messages in thread
From: KunWu Chan @ 2026-09-25 6:46 UTC (permalink / raw)
To: Lian Wang
Cc: Lance Yang, sj, mst, david, damon, linux-mm, virtualization,
Kairui Song, baohua, xueyuan.chen21
Hi Lance, Lian,
Thanks for sharing this, Lance. I also saw SJ's post yesterday and
was very happy to see DAMON being used in a real AI production
environment and delivering measurable benefits.
Hopefully, DAMON and other MM features can continue to evolve in the
AI era, better supporting diverse workloads and real-world use cases.
Looking forward to learning more from the DSec experience and
discussing this with everyone at LPC.
Thanks,
Kunwu
On Fri, Sep 25, 2026 at 1:58 PM Lian Wang <lianux.mm@gmail.com> wrote:
>
> Hi Lance,
>
> Thank you very much for sharing this and for your interest in the related
> work. It is great to see DAMON and virtio-balloon being used together in
> another real, large-scale system.
>
> One scenario we have recently been studying came from Sangfor. In a KVM/QEMU
> deployment with guest memory backed by shared tmpfs and host THP enabled,
> host-side DAMON can report a much larger hot-memory proportion than the
> guest's fine-grained activity suggests. The related discussion is here:
>
> https://lore.kernel.org/all/20260915025514.2434-1-lianux.mm@gmail.com/
>
> We are preparing more complete experiments and improvements around this
> scenario, and we plan to discuss it in more depth at LPC.
>
> I am still actively running experiments based in part on SJ's earlier
> introductions and our discussions. After LPC, we hope to share results on a
> more regular basis, together with discussion outcomes and reports from more
> scenarios. I believe DAMON will continue to become even better as these users
> and perspectives come together.
>
> I am also contacting the authors of the DeepSeek paper to learn more about
> their setup and experience. We hope they can join the discussion so that we
> can work together to improve this area further.
>
> Kairui is also very interested in this DAMON use case and may already have
> discussed related ideas with SJ, so I have kept him in Cc.
>
> Thank you again, Lance, and thanks everyone!
>
> Best,
> Lian
>
> On Fri, 25 Sep 2026 13:44:08 +0800 Lance Yang <lance.yang@linux.dev> wrote:
>
> > I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
> > and virtio-balloon:
> >
> > And it worked: time-integrated host memory use fell by 21.2%, while peak
> > usage stayed about the same, with no significant CPU overhead.
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
2026-09-25 5:57 ` Lian Wang
@ 2026-09-25 7:04 ` David Hildenbrand (Arm)
2026-09-25 7:28 ` Lian Wang
2026-09-25 8:15 ` Lance Yang
2026-09-25 10:13 ` SJ Park
2026-09-29 12:32 ` Jialiang Huang
3 siblings, 2 replies; 21+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-25 7:04 UTC (permalink / raw)
To: Lance Yang, sj, mst
Cc: damon, linux-mm, virtualization, ryncsn, kunwu.chan, lianux.mm,
baohua, xueyuan.chen21
On 9/25/26 07:44, Lance Yang wrote:
> Hi all,
>
> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
> and virtio-balloon:
>
> With this kind of workload, an agent may read a file once and never touch
> it again, while those pages remain in the guest page cache. Without memory
> pressure in the guest, they can stay cached even though the host would
> like that memory back ...
>
> The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
Heh, I read "virtio-balloon" and thought "balloon inflation/deflation, what year
is it?!". Free-page reporting makes much more sense.
> cold file pages from the guest page cache; buddy gets a chance to coalesce
> them into reportable blocks, and virtio-balloon passes those blocks to
> Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
When I was at RH we were looking at this issue as well. virtio-pmem was one way
of avoiding the page cache in VM entirely. But it has its own limitations.
Thanks for sharing!
--
Cheers,
David
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 7:04 ` David Hildenbrand (Arm)
@ 2026-09-25 7:28 ` Lian Wang
2026-09-25 8:15 ` Lance Yang
1 sibling, 0 replies; 21+ messages in thread
From: Lian Wang @ 2026-09-25 7:28 UTC (permalink / raw)
To: David Hildenbrand (Arm)
Cc: Lance Yang, sj, mst, damon, linux-mm, virtualization, Kairui Song,
kunwu.chan, lianux.mm, baohua, xueyuan.chen21
Hi David,
Indeed, both directions are very exciting. virtio-pmem avoids duplicated
page-cache residency, while DAMON with free-page reporting provides a
practical path to reclaim cold guest memory and return it to the host.
This connects closely with the work we are already doing. We are continuing
to improve our DAMON-based experiments, and we would also like to explore how
DAMON-based memory reclamation can be coordinated with broader CPU and
resource monitoring.
At LPC, I plan to bring detailed data and discuss the follow-up design with
everyone. I hope we can turn this into an executable milestone, bring the
related ideas together, and try more scenarios as a community. Everything
looks very encouraging and gives us much to look forward to. With SJ helping
coordinate the broader roadmap and everyone contributing, I am confident that
we will achieve many more exciting results. Kernel work like this can continue
making the world better.
Thanks,
Lian
On Fri, 25 Sep 2026 09:04:41 +0200 "David Hildenbrand (Arm)" <david@kernel.org> wrote:
> Free-page reporting makes much more sense.
>
> When I was at RH we were looking at this issue as well. virtio-pmem was one
> way of avoiding the page cache in VM entirely. But it has its own limitations.
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 7:04 ` David Hildenbrand (Arm)
2026-09-25 7:28 ` Lian Wang
@ 2026-09-25 8:15 ` Lance Yang
2026-09-25 10:03 ` David Hildenbrand (Arm)
2026-09-25 10:16 ` Gao Xiang
1 sibling, 2 replies; 21+ messages in thread
From: Lance Yang @ 2026-09-25 8:15 UTC (permalink / raw)
To: david, muchun.song, research
Cc: lance.yang, sj, mst, damon, linux-mm, virtualization, ryncsn,
kunwu.chan, lianux.mm, baohua, xueyuan.chen21
+Cc DeepSeek and Muchun
On Fri, Sep 25, 2026 at 09:04:41AM +0200, David Hildenbrand (Arm) wrote:
>On 9/25/26 07:44, Lance Yang wrote:
>> Hi all,
>>
>> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
>> and virtio-balloon:
>>
>> With this kind of workload, an agent may read a file once and never touch
>> it again, while those pages remain in the guest page cache. Without memory
>> pressure in the guest, they can stay cached even though the host would
>> like that memory back ...
>>
>> The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
>
>Heh, I read "virtio-balloon" and thought "balloon inflation/deflation, what year
>is it?!". Free-page reporting makes much more sense.
>
>> cold file pages from the guest page cache; buddy gets a chance to coalesce
>> them into reportable blocks, and virtio-balloon passes those blocks to
>> Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
>
>When I was at RH we were looking at this issue as well. virtio-pmem was one way
>of avoiding the page cache in VM entirely. But it has its own limitations.
YES, they enable virtio-pmem with DAX for the read-only EROFS base-image
and toolkit layers, while using DAMON with balloon free-page reporting to
reclaim cold file pages from the guest page cache on larger writable disks.
They also point out that the guest must allocate struct page metadata for
the entire pmem-backed address range. So virtio-pmem is not free either :)
BTW, Muchun recently posted a pretty cool series for exactly that:
https://lore.kernel.org/linux-mm/20260903122128.12264-1-songmuchun@bytedance.com/
(It shares vmemmap backing until a DAX fault needs private metadata,
avoiding the full per-PFN cost up front.)
I have a feeling the DeepSeek team will be watching this one closely :P
>Thanks for sharing!
Cheers!
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 8:15 ` Lance Yang
@ 2026-09-25 10:03 ` David Hildenbrand (Arm)
2026-09-25 10:16 ` Gao Xiang
1 sibling, 0 replies; 21+ messages in thread
From: David Hildenbrand (Arm) @ 2026-09-25 10:03 UTC (permalink / raw)
To: Lance Yang, muchun.song, research
Cc: sj, mst, damon, linux-mm, virtualization, ryncsn, kunwu.chan,
lianux.mm, baohua, xueyuan.chen21
On 9/25/26 10:15, Lance Yang wrote:
> +Cc DeepSeek and Muchun
>
> On Fri, Sep 25, 2026 at 09:04:41AM +0200, David Hildenbrand (Arm) wrote:
>> On 9/25/26 07:44, Lance Yang wrote:
>>> Hi all,
>>>
>>> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
>>> and virtio-balloon:
>>>
>>> With this kind of workload, an agent may read a file once and never touch
>>> it again, while those pages remain in the guest page cache. Without memory
>>> pressure in the guest, they can stay cached even though the host would
>>> like that memory back ...
>>>
>>> The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
>>
>> Heh, I read "virtio-balloon" and thought "balloon inflation/deflation, what year
>> is it?!". Free-page reporting makes much more sense.
>>
>>> cold file pages from the guest page cache; buddy gets a chance to coalesce
>>> them into reportable blocks, and virtio-balloon passes those blocks to
>>> Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
>>
>> When I was at RH we were looking at this issue as well. virtio-pmem was one way
>> of avoiding the page cache in VM entirely. But it has its own limitations.
>
> YES, they enable virtio-pmem with DAX for the read-only EROFS base-image
> and toolkit layers, while using DAMON with balloon free-page reporting to
> reclaim cold file pages from the guest page cache on larger writable disks.
>
> They also point out that the guest must allocate struct page metadata for
> the entire pmem-backed address range. So virtio-pmem is not free either :)
>
> BTW, Muchun recently posted a pretty cool series for exactly that:
>
> https://lore.kernel.org/linux-mm/20260903122128.12264-1-songmuchun@bytedance.com/
Yes, that will be helpful in that regard.
>
> (It shares vmemmap backing until a DAX fault needs private metadata,
> avoiding the full per-PFN cost up front.)
>
> I have a feeling the DeepSeek team will be watching this one closely :P
;)
--
Cheers,
David
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
2026-09-25 5:57 ` Lian Wang
2026-09-25 7:04 ` David Hildenbrand (Arm)
@ 2026-09-25 10:13 ` SJ Park
2026-09-29 12:32 ` Jialiang Huang
3 siblings, 0 replies; 21+ messages in thread
From: SJ Park @ 2026-09-25 10:13 UTC (permalink / raw)
To: Lance Yang
Cc: SJ Park, mst, david, damon, linux-mm, virtualization, ryncsn,
kunwu.chan, lianux.mm, baohua, xueyuan.chen21
Hi Lance,
On Fri, 25 Sep 2026 13:44:08 +0800 Lance Yang <lance.yang@linux.dev> wrote:
> Hi all,
>
> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
> and virtio-balloon:
>
> With this kind of workload, an agent may read a file once and never touch
> it again, while those pages remain in the guest page cache. Without memory
> pressure in the guest, they can stay cached even though the host would
> like that memory back ...
>
> The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
> cold file pages from the guest page cache; buddy gets a chance to coalesce
> them into reportable blocks, and virtio-balloon passes those blocks to
> Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
>
> And it worked: time-integrated host memory use fell by 21.2%, while peak
> usage stayed about the same, with no significant CPU overhead. Pretty cool
> to see these pieces show up in another real, large-scale system :)
Thank you for sharing this, Lance!
>
> Kernel work still changes the world. Cheers to that!
Indeed, the community is making the work a great place.
>
> Anyhow, I still don't know DAMON and virtio-balloon as well as I should
> (mostly I just know the people who built them) :P
Please feel free to ask any question to DAMON community and me, whenever you
get :)
>
> [1] https://arxiv.org/html/2609.22978
I only briefly read the paper. But it looks like DeepSeek's DAMON usage is
similar to AWS' DAMON usage [1] for their serverless service.
While I was in AWS, I proposed [2] a more advanced version called
Access/Contiguity-aware Memory Autoscaling (ACMA). After the RFC idea v2 [2],
no progress is made so far, though. I shortly covered it in LSFMMBPF'25 [3],
but all I mentioned was that the project is suspended. Recently a useer has
personally reached out to me asking the status of ACMA. It is too early stage
of the discussion to say or promise something in my humble perspective, though.
[1] https://cdn.amazon.science/ee/a4/41ff11374f2f865e5e24de11bd17/resource-management-in-aurora-serverless.pdf
[2] https://lore.kernel.org/all/20240512193657.79298-1-sj@kernel.org/
[3] https://lore.kernel.org/all/20250121183103.42877-1-sj@kernel.org/
Thanks,
SJ
[...]
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 8:15 ` Lance Yang
2026-09-25 10:03 ` David Hildenbrand (Arm)
@ 2026-09-25 10:16 ` Gao Xiang
2026-09-25 10:27 ` Gao Xiang
1 sibling, 1 reply; 21+ messages in thread
From: Gao Xiang @ 2026-09-25 10:16 UTC (permalink / raw)
To: Lance Yang
Cc: david, muchun.song, research, sj, mst, damon, linux-mm,
virtualization, ryncsn, kunwu.chan, lianux.mm, baohua,
xueyuan.chen21
Hi,
On Fri, Sep 25, 2026 at 04:15:04PM +0800, Lance Yang wrote:
> +Cc DeepSeek and Muchun
>
> On Fri, Sep 25, 2026 at 09:04:41AM +0200, David Hildenbrand (Arm) wrote:
> >On 9/25/26 07:44, Lance Yang wrote:
> >> Hi all,
> >>
> >> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
> >> and virtio-balloon:
> >>
> >> With this kind of workload, an agent may read a file once and never touch
> >> it again, while those pages remain in the guest page cache. Without memory
> >> pressure in the guest, they can stay cached even though the host would
> >> like that memory back ...
> >>
> >> The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
> >
> >Heh, I read "virtio-balloon" and thought "balloon inflation/deflation, what year
> >is it?!". Free-page reporting makes much more sense.
> >
> >> cold file pages from the guest page cache; buddy gets a chance to coalesce
> >> them into reportable blocks, and virtio-balloon passes those blocks to
> >> Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
> >
> >When I was at RH we were looking at this issue as well. virtio-pmem was one way
> >of avoiding the page cache in VM entirely. But it has its own limitations.
>
> YES, they enable virtio-pmem with DAX for the read-only EROFS base-image
> and toolkit layers, while using DAMON with balloon free-page reporting to
> reclaim cold file pages from the guest page cache on larger writable disks.
>
> They also point out that the guest must allocate struct page metadata for
> the entire pmem-backed address range. So virtio-pmem is not free either :)
>
> BTW, Muchun recently posted a pretty cool series for exactly that:
>
> https://lore.kernel.org/linux-mm/20260903122128.12264-1-songmuchun@bytedance.com/
>
> (It shares vmemmap backing until a DAX fault needs private metadata,
> avoiding the full per-PFN cost up front.)
>
> I have a feeling the DeepSeek team will be watching this one closely :P
There are several points virtio-pmem RW from my own viewpoints:
- It makes the write async I/O synchronously, note that write I/Os are
not quite the same as read I/Os (read I/Os are mostly sync). Storage
also support multi-queues which can better leverage that, and that is
why sometimes brd block device is not good at high-performance nvme
for example.
- Note that sandbox usually has a memory limit, but dax RW makes the
whole rootfs addressable, e.g. if you have 128GiB rootfs, which means
you could fault 128GiB on the host, instead of the sandbox memory
size, so it might cause some security concern (as long as users
shouldn't expect the host memory can be used up to 128GiB + memsize).
- It can cause sync 4K faults on the host in the worst case (maybe
large folios on the host can improve a bit yet not quite), in
constant to the guest memory + THP usage. I don't know how the
reclaim overhead is measured currently, but it seems the dsec paper
also mentioned in this case.
- The guest workload will still use mmap() for many sandbox apps, so
`struct page` optimization is just for the optimized case, but not
for the worst cases, the malicious VM sandboxes can still take
much more `struct page` in the guest.
There would be better to have some benchmark here for typical RL
training RW virtio-pmem: but block storage semantics cannot already
be replaced with the memory semantics.
I think virtio-pmem RO is useful simply because it can reuse the same
page cache among multiple sandboxes on the host, which can even
warm-up other sandbox workloads, although it still has some security
concern but I guess for RL training it doesn't matter and read is
almost synchronous unlike writes.
Thanks,
Gao Xiang
>
> >Thanks for sharing!
>
> Cheers!
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 10:16 ` Gao Xiang
@ 2026-09-25 10:27 ` Gao Xiang
0 siblings, 0 replies; 21+ messages in thread
From: Gao Xiang @ 2026-09-25 10:27 UTC (permalink / raw)
To: Lance Yang
Cc: Gao Xiang, david, muchun.song, research, sj, mst, damon, linux-mm,
virtualization, ryncsn, kunwu.chan, lianux.mm, baohua,
xueyuan.chen21
On Fri, Sep 25, 2026 at 12:16:55PM +0200, Gao Xiang wrote:
> Hi,
>
> On Fri, Sep 25, 2026 at 04:15:04PM +0800, Lance Yang wrote:
> > +Cc DeepSeek and Muchun
> >
> > On Fri, Sep 25, 2026 at 09:04:41AM +0200, David Hildenbrand (Arm) wrote:
> > >On 9/25/26 07:44, Lance Yang wrote:
> > >> Hi all,
> > >>
> > >> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
> > >> and virtio-balloon:
> > >>
> > >> With this kind of workload, an agent may read a file once and never touch
> > >> it again, while those pages remain in the guest page cache. Without memory
> > >> pressure in the guest, they can stay cached even though the host would
> > >> like that memory back ...
> > >>
> > >> The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
> > >
> > >Heh, I read "virtio-balloon" and thought "balloon inflation/deflation, what year
> > >is it?!". Free-page reporting makes much more sense.
> > >
> > >> cold file pages from the guest page cache; buddy gets a chance to coalesce
> > >> them into reportable blocks, and virtio-balloon passes those blocks to
> > >> Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
> > >
> > >When I was at RH we were looking at this issue as well. virtio-pmem was one way
> > >of avoiding the page cache in VM entirely. But it has its own limitations.
> >
> > YES, they enable virtio-pmem with DAX for the read-only EROFS base-image
> > and toolkit layers, while using DAMON with balloon free-page reporting to
> > reclaim cold file pages from the guest page cache on larger writable disks.
> >
> > They also point out that the guest must allocate struct page metadata for
> > the entire pmem-backed address range. So virtio-pmem is not free either :)
> >
> > BTW, Muchun recently posted a pretty cool series for exactly that:
> >
> > https://lore.kernel.org/linux-mm/20260903122128.12264-1-songmuchun@bytedance.com/
> >
> > (It shares vmemmap backing until a DAX fault needs private metadata,
> > avoiding the full per-PFN cost up front.)
> >
> > I have a feeling the DeepSeek team will be watching this one closely :P
>
> There are several points virtio-pmem RW from my own viewpoints:
>
> - It makes the write async I/O synchronously, note that write I/Os are
> not quite the same as read I/Os (read I/Os are mostly sync). Storage
> also support multi-queues which can better leverage that, and that is
> why sometimes brd block device is not good at high-performance nvme
> for example.
>
> - Note that sandbox usually has a memory limit, but dax RW makes the
> whole rootfs addressable, e.g. if you have 128GiB rootfs, which means
> you could fault 128GiB on the host, instead of the sandbox memory
> size, so it might cause some security concern (as long as users
> shouldn't expect the host memory can be used up to 128GiB + memsize).
>
> - It can cause sync 4K faults on the host in the worst case (maybe
> large folios on the host can improve a bit yet not quite), in
> constant to the guest memory + THP usage. I don't know how the
> reclaim overhead is measured currently, but it seems the dsec paper
> also mentioned in this case.
>
> - The guest workload will still use mmap() for many sandbox apps, so
> `struct page` optimization is just for the optimized case, but not
> for the worst cases, the malicious VM sandboxes can still take
> much more `struct page` in the guest.
>
> There would be better to have some benchmark here for typical RL
> training RW virtio-pmem: but block storage semantics cannot already
> be replaced with the memory semantics.
>
> I think virtio-pmem RO is useful simply because it can reuse the same
> page cache among multiple sandboxes on the host, which can even
> warm-up other sandbox workloads, although it still has some security
> concern but I guess for RL training it doesn't matter and read is
> almost synchronous unlike writes.
BTW, I've thought about the sandbox writable layers for a while, maybe
EROFS could have its own dedicated efficient writable layers as a
optional feature at some time, but I need to think carefully first
and look forward to get more numbers before landing a premature
implementation to the upstream (memory semantics, storage semantics
or just a overlay + hybrid approaches); also there are some
non-technical points to move forward in this direction.
There are some other important features which are more like low-hanging
fruits, so I'm more in a wait-and-see mode until I get a sensible
direction on this sandboxing scenario.
Thanks,
Gao Xiang
>
> Thanks,
> Gao Xiang
>
> >
> > >Thanks for sharing!
> >
> > Cheers!
>
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-25 5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
` (2 preceding siblings ...)
2026-09-25 10:13 ` SJ Park
@ 2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41 ` Gao Xiang
` (2 more replies)
3 siblings, 3 replies; 21+ messages in thread
From: Jialiang Huang @ 2026-09-29 12:32 UTC (permalink / raw)
To: lance.yang
Cc: baohua, damon, david, kunwu.chan, lianux.mm, linux-mm, mst,
ryncsn, sj, virtualization, xueyuan.chen21, huang-jl, xiang
Hi all,
I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
to everyone working on DAMON and virtio-balloon free-page reporting.
They have been very useful for our workloads.
Gao Xiang wrote:
> It can cause sync 4K faults on the host in the worst case
This is one of our concerns with virtio-pmem as well: moving I/O onto
the page-fault path can introduce performance trade-offs. The other
concern is the substantial struct page overhead for large images.
For now, we enable virtio-pmem only for moderately sized, frequently
used read-only images, where there is more opportunity to share the
same host page cache across sandboxes, as Gao pointed out.
Muchun's vmemmap work is also interesting to us. My understanding is
that it allocates private backing for struct page metadata on demand,
which could help reduce the upfront memory overhead for large images.
For disks without virtio-pmem, DAMON with virtio-balloon free-page
reporting lets us reclaim cold guest page-cache pages and return
the memory to the host, without those virtio-pmem-specific issues.
The two approaches complement each other in our setup.
We have not yet fully explored how best to tune the DAMON and free-page
reporting parameters for our workloads. For example, the kernel's default
free-page reporting granularity is 2 MiB, which is fairly coarse:
reclaiming cold pages does not necessarily produce free blocks of that
size. We still need to evaluate how finer reporting granularity and
different DAMON settings affect memory savings and workload performance.
Best,
Jialiang Huang
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-29 12:32 ` Jialiang Huang
@ 2026-09-29 12:41 ` Gao Xiang
2026-09-29 12:50 ` Jialiang Huang
2026-09-30 3:36 ` Muchun Song
2026-09-29 16:56 ` SJ Park
2026-09-29 18:14 ` Pratyush Mallick
2 siblings, 2 replies; 21+ messages in thread
From: Gao Xiang @ 2026-09-29 12:41 UTC (permalink / raw)
To: Jialiang Huang
Cc: lance.yang, baohua, damon, david, kunwu.chan, lianux.mm, linux-mm,
mst, ryncsn, sj, virtualization, xueyuan.chen21, xiang
On Tue, Sep 29, 2026 at 08:32:41PM +0800, Jialiang Huang wrote:
> Hi all,
>
> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
> to everyone working on DAMON and virtio-balloon free-page reporting.
> They have been very useful for our workloads.
>
> Gao Xiang wrote:
> > It can cause sync 4K faults on the host in the worst case
>
> This is one of our concerns with virtio-pmem as well: moving I/O onto
> the page-fault path can introduce performance trade-offs. The other
> concern is the substantial struct page overhead for large images.
>
> For now, we enable virtio-pmem only for moderately sized, frequently
> used read-only images, where there is more opportunity to share the
> same host page cache across sandboxes, as Gao pointed out.
>
> Muchun's vmemmap work is also interesting to us. My understanding is
> that it allocates private backing for struct page metadata on demand,
> which could help reduce the upfront memory overhead for large images.
Although I haven't had a chance and time to look into that, the main
concern from me is that mmap() access will call
"dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox
workloads (or not malicious, just valid mmap workloads) can cause guest
memory OOMs due to "struct page balloon" for large rootfs in the worst
cases and cause the follow-up mmap access failure, because the guest
memory size may not even fulfill "struct page" for large rootfs.
Thanks,
Gao Xiang
>
> For disks without virtio-pmem, DAMON with virtio-balloon free-page
> reporting lets us reclaim cold guest page-cache pages and return
> the memory to the host, without those virtio-pmem-specific issues.
> The two approaches complement each other in our setup.
>
> We have not yet fully explored how best to tune the DAMON and free-page
> reporting parameters for our workloads. For example, the kernel's default
> free-page reporting granularity is 2 MiB, which is fairly coarse:
> reclaiming cold pages does not necessarily produce free blocks of that
> size. We still need to evaluate how finer reporting granularity and
> different DAMON settings affect memory savings and workload performance.
>
> Best,
> Jialiang Huang
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-29 12:41 ` Gao Xiang
@ 2026-09-29 12:50 ` Jialiang Huang
2026-09-30 3:36 ` Muchun Song
1 sibling, 0 replies; 21+ messages in thread
From: Jialiang Huang @ 2026-09-29 12:50 UTC (permalink / raw)
To: xiang
Cc: baohua, damon, david, huang-jl, kunwu.chan, lance.yang, lianux.mm,
linux-mm, mst, ryncsn, sj, virtualization, xueyuan.chen21
>> Muchun's vmemmap work is also interesting to us. My understanding is
>> that it allocates private backing for struct page metadata on demand,
>> which could help reduce the upfront memory overhead for large images.
>
> Although I haven't had a chance and time to look into that, the main
> concern from me is that mmap() access will call
> "dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox
> workloads (or not malicious, just valid mmap workloads) can cause guest
> memory OOMs due to "struct page balloon" for large rootfs in the worst
> cases and cause the follow-up mmap access failure, because the guest
> memory size may not even fulfill "struct page" for large rootfs.
>
> Thanks,
> Gao Xiang
Agreed, if a workload actually accesses a sufficiently large range of the
image, allocating struct page metadata on demand could indeed exhaust
guest memory.
Best,
Jialiang Huang
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41 ` Gao Xiang
@ 2026-09-29 16:56 ` SJ Park
2026-09-30 3:07 ` Lian Wang
2026-09-29 18:14 ` Pratyush Mallick
2 siblings, 1 reply; 21+ messages in thread
From: SJ Park @ 2026-09-29 16:56 UTC (permalink / raw)
To: Jialiang Huang
Cc: SJ Park, lance.yang, baohua, damon, david, kunwu.chan, lianux.mm,
linux-mm, mst, ryncsn, virtualization, xueyuan.chen21, xiang
On Tue, 29 Sep 2026 20:32:41 +0800 Jialiang Huang <huang-jl@deepseek.com> wrote:
> Hi all,
>
> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
> to everyone working on DAMON and virtio-balloon free-page reporting.
> They have been very useful for our workloads.
So glad to hear this. And thank you for sharing your use case!
[...]
> We have not yet fully explored how best to tune the DAMON and free-page
> reporting parameters for our workloads. For example, the kernel's default
> free-page reporting granularity is 2 MiB, which is fairly coarse:
> reclaiming cold pages does not necessarily produce free blocks of that
> size. We still need to evaluate how finer reporting granularity and
> different DAMON settings affect memory savings and workload performance.
I agree the reclaim-reporting granularity mismatch could be a room to improve.
As I also mentioned on my previous reply, this was one of the main motivations
of Access/Contiguity-aware Memory Autoscaling [1] proposal. As also mentioned
on the previous reply, the project made no much progress yet, though.
[1] https://lore.kernel.org/all/20240512193657.79298-1-sj@kernel.org/
Thanks,
SJ
[...]
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41 ` Gao Xiang
2026-09-29 16:56 ` SJ Park
@ 2026-09-29 18:14 ` Pratyush Mallick
2 siblings, 0 replies; 21+ messages in thread
From: Pratyush Mallick @ 2026-09-29 18:14 UTC (permalink / raw)
To: huang-jl
Cc: baohua, damon, david, kunwu.chan, lance.yang, lianux.mm, linux-mm,
mst, ryncsn, sj, virtualization, xiang, xueyuan.chen21
> We have not yet fully explored how best to tune the DAMON and free-page
> reporting parameters for our workloads.
You can also try tuning page_reporting_delay_ms, which can help you with
saving time-integrated host memory.
[1] https://lore.kernel.org/20260731193705.2902728-1-pratmal@google.com
Thanks,
Pratyush
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-29 16:56 ` SJ Park
@ 2026-09-30 3:07 ` Lian Wang
2026-09-30 8:08 ` SJ Park
0 siblings, 1 reply; 21+ messages in thread
From: Lian Wang @ 2026-09-30 3:07 UTC (permalink / raw)
To: SJ Park, Jialiang Huang
Cc: lance.yang, KunWu Chan, baohua, damon, david, linux-mm, mst,
ryncsn, virtualization, xueyuan.chen21, xiang
Hi SJ,
On Tue, 29 Sep 2026 09:56:03 -0700 SJ Park <sj@kernel.org> wrote:
> I agree the reclaim-reporting granularity mismatch could be a room to improve.
Thanks for pointing us back to the Access/Contiguity-aware Memory
Autoscaling proposal. Jialiang's description gives us a useful deployment
context and a concrete problem to investigate.
The guest reclaim/reporting case and our host-side THP/tiering case are
different, but both raise questions about how workload behaviour and
monitoring granularity should guide kernel actions.
KunWu and I would like to help move this work forward together, alongside
the existing DAMON roadmap. For now, I would like to share a few possible
areas for discussion:
1. Workloads, DAMON policies, and tiering. My current work includes
improving masim-based experiments and studying TPP-inspired memory
tiering. We would like to make the workloads more representative and
better understand how DAMON's observations and policy settings affect
placement decisions. The aim is to understand where configuration
changes are sufficient and where code changes, if any, would help.
2. Testing and contributing to the existing PMU work. Our initial focus
would be on helping with the Arm SPE integration and testing the IBS
and PEBS work on x86, in coordination with the people already working
on these. The experiments could evaluate whether finer-grained
sampling provides useful additional information, and at what CPU
cost. We could leave huge-page-specific mechanisms for a later
discussion, guided by the results.
3. Following up on the DeepSeek workload. If the team is interested, we
could learn more about the workloads and tuning questions they can
share, and study the path from workload behaviour and monitoring
parameters through reclaim and free-page reporting. Pratyush's
reporting-delay suggestion adds another useful dimension. This may
provide a practical starting point for revisiting parts of the
access/contiguity-aware autoscaling work and evaluating their
relevance to the reported workload.
Could we discuss whether some of these could become follow-up milestones
in the DAMON discussions at LPC, or here on the mailing list? We would be
happy to help organise the discussion and take on agreed pieces of
testing or development. We would also welcome closer collaboration with
DeepSeek and other interested MM developers, including Muchun, where our
work overlaps.
This is only a brief outline for now. I plan to bring the experimental
results and remaining questions to LPC. After discussing and aligning on
the details there, we plan to share a more detailed proposal with the
community and write up the planned work and experimental findings in
blog posts.
I appreciate how much is already on your schedule. We hope to support
the existing plans and help with agreed tasks as the scope becomes
clearer. There is no urgency to reply; we can discuss these ideas
whenever it fits your schedule.
Thanks,
Lian
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-29 12:41 ` Gao Xiang
2026-09-29 12:50 ` Jialiang Huang
@ 2026-09-30 3:36 ` Muchun Song
2026-09-30 7:24 ` Gao Xiang
1 sibling, 1 reply; 21+ messages in thread
From: Muchun Song @ 2026-09-30 3:36 UTC (permalink / raw)
To: Gao Xiang
Cc: Jialiang Huang, lance.yang, baohua, damon, david, kunwu.chan,
lianux.mm, linux-mm, mst, ryncsn, sj, virtualization,
xueyuan.chen21
> On Sep 29, 2026, at 20:41, Gao Xiang <xiang@kernel.org> wrote:
>
> On Tue, Sep 29, 2026 at 08:32:41PM +0800, Jialiang Huang wrote:
>> Hi all,
>>
>> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
>> to everyone working on DAMON and virtio-balloon free-page reporting.
>> They have been very useful for our workloads.
>>
>> Gao Xiang wrote:
>>> It can cause sync 4K faults on the host in the worst case
>>
>> This is one of our concerns with virtio-pmem as well: moving I/O onto
>> the page-fault path can introduce performance trade-offs. The other
>> concern is the substantial struct page overhead for large images.
>>
>> For now, we enable virtio-pmem only for moderately sized, frequently
>> used read-only images, where there is more opportunity to share the
>> same host page cache across sandboxes, as Gao pointed out.
>>
>> Muchun's vmemmap work is also interesting to us. My understanding is
>> that it allocates private backing for struct page metadata on demand,
>> which could help reduce the upfront memory overhead for large images.
>
> Although I haven't had a chance and time to look into that, the main
> concern from me is that mmap() access will call
> "dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox
> workloads (or not malicious, just valid mmap workloads) can cause guest
> memory OOMs due to "struct page balloon" for large rootfs in the worst
> cases and cause the follow-up mmap access failure, because the guest
> memory size may not even fulfill "struct page" for large rootfs.
I agree that this is a real issue with v1.
One detail is that merely establishing the mapping does not materialize
the metadata. Materialization happens when a fault resolves to an
allocated DAX extent, before its PFN is inserted into a userspace
mapping. However, that distinction does not remove the problem.
In the worst case, a workload can fault enough of the pmem range to
restore the full vmemmap cost, about 1.56% of the pmem size. Since v1
does not dematerialize private vmemmap pages, even a one-time scan can
retain that cost until the device is removed.
The follow-up mentioned in the cover letter is intended to make the
optimization reversible. One possible direction would be to invalidate
clean DAX entries under memory pressure and zap their userspace
mappings. Once a DAX entry has been removed and the corresponding PFNs
have no remaining mappings, references, or pins that require private
metadata, the associated vmemmap backing could be remapped to the
shared read-only page. A later access would fault and materialize it
again.
This would address the persistent ballooning caused by a one-time scan,
but it would not help while a large DAX working set remains actively
used or pinned.
As a separate policy mechanism for that case, one possible direction
would be to account the faulted pmem range to the faulting mm's memory
cgroup: PAGE_SIZE for a PTE mapping and PMD_SIZE for a PMD mapping.
This would intentionally account the mapped DAX capacity.
Such accounting could be explicitly enabled through a cgroup v2 mount
option, following the opt-in model of memory_hugetlb_accounting, so the
existing FS-DAX accounting semantics would remain unchanged by default.
With such accounting, an active or pinned DAX working set could still
reach its cgroup limit, but it could not grow private vmemmap without
corresponding cgroup usage. This is only a possible direction for
bounding the peak working set, though, and I have not worked through
all of its implementation details yet.
Thanks,
Muchun
>
> Thanks,
> Gao Xiang
>
>>
>> For disks without virtio-pmem, DAMON with virtio-balloon free-page
>> reporting lets us reclaim cold guest page-cache pages and return
>> the memory to the host, without those virtio-pmem-specific issues.
>> The two approaches complement each other in our setup.
>>
>> We have not yet fully explored how best to tune the DAMON and free-page
>> reporting parameters for our workloads. For example, the kernel's default
>> free-page reporting granularity is 2 MiB, which is fairly coarse:
>> reclaiming cold pages does not necessarily produce free blocks of that
>> size. We still need to evaluate how finer reporting granularity and
>> different DAMON settings affect memory savings and workload performance.
>>
>> Best,
>> Jialiang Huang
>
>
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-30 3:36 ` Muchun Song
@ 2026-09-30 7:24 ` Gao Xiang
2026-09-30 9:37 ` Muchun Song
0 siblings, 1 reply; 21+ messages in thread
From: Gao Xiang @ 2026-09-30 7:24 UTC (permalink / raw)
To: Muchun Song
Cc: Gao Xiang, Jialiang Huang, lance.yang, baohua, damon, david,
kunwu.chan, lianux.mm, linux-mm, mst, ryncsn, sj, virtualization,
xueyuan.chen21
On Wed, Sep 30, 2026 at 11:36:23AM +0800, Muchun Song wrote:
>
>
> > On Sep 29, 2026, at 20:41, Gao Xiang <xiang@kernel.org> wrote:
> >
> > On Tue, Sep 29, 2026 at 08:32:41PM +0800, Jialiang Huang wrote:
> >> Hi all,
> >>
> >> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
> >> to everyone working on DAMON and virtio-balloon free-page reporting.
> >> They have been very useful for our workloads.
> >>
> >> Gao Xiang wrote:
> >>> It can cause sync 4K faults on the host in the worst case
> >>
> >> This is one of our concerns with virtio-pmem as well: moving I/O onto
> >> the page-fault path can introduce performance trade-offs. The other
> >> concern is the substantial struct page overhead for large images.
> >>
> >> For now, we enable virtio-pmem only for moderately sized, frequently
> >> used read-only images, where there is more opportunity to share the
> >> same host page cache across sandboxes, as Gao pointed out.
> >>
> >> Muchun's vmemmap work is also interesting to us. My understanding is
> >> that it allocates private backing for struct page metadata on demand,
> >> which could help reduce the upfront memory overhead for large images.
> >
> > Although I haven't had a chance and time to look into that, the main
> > concern from me is that mmap() access will call
> > "dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox
> > workloads (or not malicious, just valid mmap workloads) can cause guest
> > memory OOMs due to "struct page balloon" for large rootfs in the worst
> > cases and cause the follow-up mmap access failure, because the guest
> > memory size may not even fulfill "struct page" for large rootfs.
>
> I agree that this is a real issue with v1.
>
> One detail is that merely establishing the mapping does not materialize
> the metadata. Materialization happens when a fault resolves to an
> allocated DAX extent, before its PFN is inserted into a userspace
> mapping. However, that distinction does not remove the problem.
>
> In the worst case, a workload can fault enough of the pmem range to
> restore the full vmemmap cost, about 1.56% of the pmem size. Since v1
> does not dematerialize private vmemmap pages, even a one-time scan can
> retain that cost until the device is removed.
>
> The follow-up mentioned in the cover letter is intended to make the
> optimization reversible. One possible direction would be to invalidate
> clean DAX entries under memory pressure and zap their userspace
> mappings. Once a DAX entry has been removed and the corresponding PFNs
> have no remaining mappings, references, or pins that require private
> metadata, the associated vmemmap backing could be remapped to the
> shared read-only page. A later access would fault and materialize it
> again.
BTW, it's impossible for shared DAX entries (like the current XFS DAX
with reflink and EROFS will support this feature later too for chunk
memory sharing), since you cannot just use mapping and index to get
the VMA like page cache does unless you invent another new mechanism
for this.
Reclaiming page entry mechanism seems it can be used for or overlapped
to another types of memory (in order to save struct page memory in
general): I'm not sure if it needs wider discussion on reclaiming
"struct page" in general first.
Anyway, it'd be better to get some numbers with RL or agent workloads
(especially the host memory is under reasonable pressure) before
landing all these new infras upstream if proceeding in this way.
Thanks,
Gao Xiang
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-30 3:07 ` Lian Wang
@ 2026-09-30 8:08 ` SJ Park
0 siblings, 0 replies; 21+ messages in thread
From: SJ Park @ 2026-09-30 8:08 UTC (permalink / raw)
To: Lian Wang
Cc: SJ Park, Jialiang Huang, lance.yang, KunWu Chan, baohua, damon,
david, linux-mm, mst, ryncsn, virtualization, xueyuan.chen21,
xiang
Hi Lian,
On Wed, 30 Sep 2026 11:07:29 +0800 Lian Wang <lianux.mm@gmail.com> wrote:
> Hi SJ,
>
> On Tue, 29 Sep 2026 09:56:03 -0700 SJ Park <sj@kernel.org> wrote:
>
> > I agree the reclaim-reporting granularity mismatch could be a room to improve.
>
> Thanks for pointing us back to the Access/Contiguity-aware Memory
> Autoscaling proposal. Jialiang's description gives us a useful deployment
> context and a concrete problem to investigate.
>
> The guest reclaim/reporting case and our host-side THP/tiering case are
> different, but both raise questions about how workload behaviour and
> monitoring granularity should guide kernel actions.
>
> KunWu and I would like to help move this work forward together, alongside
> the existing DAMON roadmap. For now, I would like to share a few possible
> areas for discussion:
>
> 1. Workloads, DAMON policies, and tiering. My current work includes
> improving masim-based experiments and studying TPP-inspired memory
> tiering. We would like to make the workloads more representative and
> better understand how DAMON's observations and policy settings affect
> placement decisions. The aim is to understand where configuration
> changes are sufficient and where code changes, if any, would help.
>
> 2. Testing and contributing to the existing PMU work. Our initial focus
> would be on helping with the Arm SPE integration and testing the IBS
> and PEBS work on x86, in coordination with the people already working
> on these. The experiments could evaluate whether finer-grained
> sampling provides useful additional information, and at what CPU
> cost. We could leave huge-page-specific mechanisms for a later
> discussion, guided by the results.
>
> 3. Following up on the DeepSeek workload. If the team is interested, we
> could learn more about the workloads and tuning questions they can
> share, and study the path from workload behaviour and monitoring
> parameters through reclaim and free-page reporting. Pratyush's
> reporting-delay suggestion adds another useful dimension. This may
> provide a practical starting point for revisiting parts of the
> access/contiguity-aware autoscaling work and evaluating their
> relevance to the reported workload.
>
> Could we discuss whether some of these could become follow-up milestones
> in the DAMON discussions at LPC, or here on the mailing list?
Thank you for adding topics to discuss.
Of course we could disucss anywhere anytime. I think Access/Contituity-aware
Memory Autoscaling (ACMA) could be just a parallel project. It doens't need to
be an extension of, or blocked by "beyond-page table access bit" project, in my
humble opinion.
ACMA is still on my TODO list. It is mainly a matter of prioritization. As it
becomes clear it requires more attention, I could allocate more time for the
project. There was a user showing interest recently. I'm planning to
prioritize it more. If DeepSeek could confirm their interest, it could help me
prioritizing it further.
As always, coding may not be that difficult and take that long time. Testing
might be the biggest challenge. I don't have good test infrastructure now.
Actually the lack of good test infrastructure is one of my biggest DAMON
maintenance concerns nowadays. Others' help on testing or test infra would be
super helpful.
If someone wants to implement and contribute it based on the proposed idea, it
would also be nice.
> We would be
> happy to help organise the discussion and take on agreed pieces of
> testing or development. We would also welcome closer collaboration with
> DeepSeek and other interested MM developers, including Muchun, where our
> work overlaps.
Sounds great. That would be super productive collaborations.
>
> This is only a brief outline for now. I plan to bring the experimental
> results and remaining questions to LPC. After discussing and aligning on
> the details there, we plan to share a more detailed proposal with the
> community and write up the planned work and experimental findings in
> blog posts.
We could keep discussing in mailing list, but of course in person discussions
are always helpful. I'm eagerly looking forward to the next week :)
>
> I appreciate how much is already on your schedule. We hope to support
> the existing plans and help with agreed tasks as the scope becomes
> clearer. There is no urgency to reply; we can discuss these ideas
> whenever it fits your schedule.
Appreciate your continued and grateful contributions, Lian.
Thanks,
SJ
[...]
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-30 7:24 ` Gao Xiang
@ 2026-09-30 9:37 ` Muchun Song
2026-09-30 10:33 ` Gao Xiang
0 siblings, 1 reply; 21+ messages in thread
From: Muchun Song @ 2026-09-30 9:37 UTC (permalink / raw)
To: Gao Xiang
Cc: Jialiang Huang, lance.yang, baohua, damon, david, kunwu.chan,
lianux.mm, linux-mm, mst, ryncsn, sj, virtualization,
xueyuan.chen21
> On Sep 30, 2026, at 15:24, Gao Xiang <xiang@kernel.org> wrote:
>
> On Wed, Sep 30, 2026 at 11:36:23AM +0800, Muchun Song wrote:
>>
>>
>>> On Sep 29, 2026, at 20:41, Gao Xiang <xiang@kernel.org> wrote:
>>>
>>> On Tue, Sep 29, 2026 at 08:32:41PM +0800, Jialiang Huang wrote:
>>>> Hi all,
>>>>
>>>> I'm an engineer at DeepSeek. Thanks for the discussion, and thanks
>>>> to everyone working on DAMON and virtio-balloon free-page reporting.
>>>> They have been very useful for our workloads.
>>>>
>>>> Gao Xiang wrote:
>>>>> It can cause sync 4K faults on the host in the worst case
>>>>
>>>> This is one of our concerns with virtio-pmem as well: moving I/O onto
>>>> the page-fault path can introduce performance trade-offs. The other
>>>> concern is the substantial struct page overhead for large images.
>>>>
>>>> For now, we enable virtio-pmem only for moderately sized, frequently
>>>> used read-only images, where there is more opportunity to share the
>>>> same host page cache across sandboxes, as Gao pointed out.
>>>>
>>>> Muchun's vmemmap work is also interesting to us. My understanding is
>>>> that it allocates private backing for struct page metadata on demand,
>>>> which could help reduce the upfront memory overhead for large images.
>>>
>>> Although I haven't had a chance and time to look into that, the main
>>> concern from me is that mmap() access will call
>>> "dax_fault_iter->vmemmap_materialize_page()", and malicious sandbox
>>> workloads (or not malicious, just valid mmap workloads) can cause guest
>>> memory OOMs due to "struct page balloon" for large rootfs in the worst
>>> cases and cause the follow-up mmap access failure, because the guest
>>> memory size may not even fulfill "struct page" for large rootfs.
>>
>> I agree that this is a real issue with v1.
>>
>> One detail is that merely establishing the mapping does not materialize
>> the metadata. Materialization happens when a fault resolves to an
>> allocated DAX extent, before its PFN is inserted into a userspace
>> mapping. However, that distinction does not remove the problem.
>>
>> In the worst case, a workload can fault enough of the pmem range to
>> restore the full vmemmap cost, about 1.56% of the pmem size. Since v1
>> does not dematerialize private vmemmap pages, even a one-time scan can
>> retain that cost until the device is removed.
>>
>> The follow-up mentioned in the cover letter is intended to make the
>> optimization reversible. One possible direction would be to invalidate
>> clean DAX entries under memory pressure and zap their userspace
>> mappings. Once a DAX entry has been removed and the corresponding PFNs
>> have no remaining mappings, references, or pins that require private
>> metadata, the associated vmemmap backing could be remapped to the
>> shared read-only page. A later access would fault and materialize it
>> again.
>
> BTW, it's impossible for shared DAX entries (like the current XFS DAX
> with reflink and EROFS will support this feature later too for chunk
> memory sharing), since you cannot just use mapping and index to get
> the VMA like page cache does unless you invent another new mechanism
> for this.
You're right. For the reflink scenario, reclamation is currently
difficult. If we want to reclaim, it would also be in three stages:
1) reclaim the struct page corresponding to PFNs that no longer have
any mapping; 2) reclaim the cases where mappings exist but are not
reflink; 3) reclaim the reflink scenario.
These three stages go from simple to difficult. Of course, I hadn't
thought this far ahead before. So in the v1 version, not even the
first stage was implemented. At the very least, before I act, I need
enough planning and thought.
>
> Reclaiming page entry mechanism seems it can be used for or overlapped
> to another types of memory (in order to save struct page memory in
> general): I'm not sure if it needs wider discussion on reclaiming
> "struct page" in general first.
Of course, I don't think struct page saving is the focus of the discussion
here, so we can stop discussing it.
>
> Anyway, it'd be better to get some numbers with RL or agent workloads
> (especially the host memory is under reasonable pressure) before
> landing all these new infras upstream if proceeding in this way.
At least for now, as far as I'm concerned, I don't intend to land all the
features mentioned here. The first thing I want to address is on-demand
allocation of struct page.
Thanks,
Muchun
>
> Thanks,
> Gao Xiang
^ permalink raw reply [flat|nested] 21+ messages in thread
* Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
2026-09-30 9:37 ` Muchun Song
@ 2026-09-30 10:33 ` Gao Xiang
0 siblings, 0 replies; 21+ messages in thread
From: Gao Xiang @ 2026-09-30 10:33 UTC (permalink / raw)
To: Muchun Song
Cc: Gao Xiang, Jialiang Huang, lance.yang, baohua, damon, david,
kunwu.chan, lianux.mm, linux-mm, mst, ryncsn, sj, virtualization,
xueyuan.chen21, Matthew Wilcox
On Wed, Sep 30, 2026 at 05:37:11PM +0800, Muchun Song wrote:
>
...
>
> >
> > Anyway, it'd be better to get some numbers with RL or agent workloads
> > (especially the host memory is under reasonable pressure) before
> > landing all these new infras upstream if proceeding in this way.
>
> At least for now, as far as I'm concerned, I don't intend to land all the
> features mentioned here. The first thing I want to address is on-demand
> allocation of struct page.
From my point of view, on-demend allocation of `struct page` for FSDAX
is indeed useful and can be landed upstream, mainly because `struct page`
initialization takes much long time for large FSDAX devices even a small
range is used, that impacts the guest startup time a lot.
But in order to make it safe for production, I suggest reserve the
guest memory in advance for the worst cases of `struct page` at the first
stage, otherwise it could cause unexpected issues for many guest workloads:
it's hard to perdict if the guest memory is still enough, it's unfriendly
to system designers as well (that would completely based on their
experience and hard to ensure the service quality beforehand.)
Since the guest memory for `struct page` is touched on-demand, the memory
over-committing on the host is still workable, the memory reservation
can be removed after the reclaim path is done.
Thanks,
Gao Xiang
>
> Thanks,
> Muchun
>
> >
> > Thanks,
> > Gao Xiang
>
>
^ permalink raw reply [flat|nested] 21+ messages in thread
end of thread, other threads:[~2026-09-30 10:33 UTC | newest]
Thread overview: 21+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-25 5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
2026-09-25 5:57 ` Lian Wang
2026-09-25 6:46 ` KunWu Chan
2026-09-25 7:04 ` David Hildenbrand (Arm)
2026-09-25 7:28 ` Lian Wang
2026-09-25 8:15 ` Lance Yang
2026-09-25 10:03 ` David Hildenbrand (Arm)
2026-09-25 10:16 ` Gao Xiang
2026-09-25 10:27 ` Gao Xiang
2026-09-25 10:13 ` SJ Park
2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41 ` Gao Xiang
2026-09-29 12:50 ` Jialiang Huang
2026-09-30 3:36 ` Muchun Song
2026-09-30 7:24 ` Gao Xiang
2026-09-30 9:37 ` Muchun Song
2026-09-30 10:33 ` Gao Xiang
2026-09-29 16:56 ` SJ Park
2026-09-30 3:07 ` Lian Wang
2026-09-30 8:08 ` SJ Park
2026-09-29 18:14 ` Pratyush Mallick
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox