From: Gao Xiang <xiang@kernel.org>
To: Lance Yang <lance.yang@linux.dev>
Cc: Gao Xiang <xiang@kernel.org>,
david@kernel.org, muchun.song@linux.dev, research@deepseek.com,
sj@kernel.org, mst@redhat.com, damon@lists.linux.dev,
linux-mm@kvack.org, virtualization@lists.linux.dev,
ryncsn@gmail.com, kunwu.chan@gmail.com, lianux.mm@gmail.com,
baohua@kernel.org, xueyuan.chen21@gmail.com
Subject: Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
Date: Fri, 25 Sep 2026 12:27:09 +0200 [thread overview]
Message-ID: <arZMfTVu4t4hdHKD@MacBookPro.fritz.box> (raw)
In-Reply-To: <arZKFznn4Z2LvoJN@MacBookPro.fritz.box>
On Fri, Sep 25, 2026 at 12:16:55PM +0200, Gao Xiang wrote:
> Hi,
>
> On Fri, Sep 25, 2026 at 04:15:04PM +0800, Lance Yang wrote:
> > +Cc DeepSeek and Muchun
> >
> > On Fri, Sep 25, 2026 at 09:04:41AM +0200, David Hildenbrand (Arm) wrote:
> > >On 9/25/26 07:44, Lance Yang wrote:
> > >> Hi all,
> > >>
> > >> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
> > >> and virtio-balloon:
> > >>
> > >> With this kind of workload, an agent may read a file once and never touch
> > >> it again, while those pages remain in the guest page cache. Without memory
> > >> pressure in the guest, they can stay cached even though the host would
> > >> like that memory back ...
> > >>
> > >> The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
> > >
> > >Heh, I read "virtio-balloon" and thought "balloon inflation/deflation, what year
> > >is it?!". Free-page reporting makes much more sense.
> > >
> > >> cold file pages from the guest page cache; buddy gets a chance to coalesce
> > >> them into reportable blocks, and virtio-balloon passes those blocks to
> > >> Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
> > >
> > >When I was at RH we were looking at this issue as well. virtio-pmem was one way
> > >of avoiding the page cache in VM entirely. But it has its own limitations.
> >
> > YES, they enable virtio-pmem with DAX for the read-only EROFS base-image
> > and toolkit layers, while using DAMON with balloon free-page reporting to
> > reclaim cold file pages from the guest page cache on larger writable disks.
> >
> > They also point out that the guest must allocate struct page metadata for
> > the entire pmem-backed address range. So virtio-pmem is not free either :)
> >
> > BTW, Muchun recently posted a pretty cool series for exactly that:
> >
> > https://lore.kernel.org/linux-mm/20260903122128.12264-1-songmuchun@bytedance.com/
> >
> > (It shares vmemmap backing until a DAX fault needs private metadata,
> > avoiding the full per-PFN cost up front.)
> >
> > I have a feeling the DeepSeek team will be watching this one closely :P
>
> There are several points virtio-pmem RW from my own viewpoints:
>
> - It makes the write async I/O synchronously, note that write I/Os are
> not quite the same as read I/Os (read I/Os are mostly sync). Storage
> also support multi-queues which can better leverage that, and that is
> why sometimes brd block device is not good at high-performance nvme
> for example.
>
> - Note that sandbox usually has a memory limit, but dax RW makes the
> whole rootfs addressable, e.g. if you have 128GiB rootfs, which means
> you could fault 128GiB on the host, instead of the sandbox memory
> size, so it might cause some security concern (as long as users
> shouldn't expect the host memory can be used up to 128GiB + memsize).
>
> - It can cause sync 4K faults on the host in the worst case (maybe
> large folios on the host can improve a bit yet not quite), in
> constant to the guest memory + THP usage. I don't know how the
> reclaim overhead is measured currently, but it seems the dsec paper
> also mentioned in this case.
>
> - The guest workload will still use mmap() for many sandbox apps, so
> `struct page` optimization is just for the optimized case, but not
> for the worst cases, the malicious VM sandboxes can still take
> much more `struct page` in the guest.
>
> There would be better to have some benchmark here for typical RL
> training RW virtio-pmem: but block storage semantics cannot already
> be replaced with the memory semantics.
>
> I think virtio-pmem RO is useful simply because it can reuse the same
> page cache among multiple sandboxes on the host, which can even
> warm-up other sandbox workloads, although it still has some security
> concern but I guess for RL training it doesn't matter and read is
> almost synchronous unlike writes.
BTW, I've thought about the sandbox writable layers for a while, maybe
EROFS could have its own dedicated efficient writable layers as a
optional feature at some time, but I need to think carefully first
and look forward to get more numbers before landing a premature
implementation to the upstream (memory semantics, storage semantics
or just a overlay + hybrid approaches); also there are some
non-technical points to move forward in this direction.
There are some other important features which are more like low-hanging
fruits, so I'm more in a wait-and-see mode until I get a sensible
direction on this sandboxing scenario.
Thanks,
Gao Xiang
>
> Thanks,
> Gao Xiang
>
> >
> > >Thanks for sharing!
> >
> > Cheers!
>
next prev parent reply other threads:[~2026-09-25 10:27 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-25 5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
2026-09-25 5:57 ` Lian Wang
2026-09-25 6:46 ` KunWu Chan
2026-09-25 7:04 ` David Hildenbrand (Arm)
2026-09-25 7:28 ` Lian Wang
2026-09-25 8:15 ` Lance Yang
2026-09-25 10:03 ` David Hildenbrand (Arm)
2026-09-25 10:16 ` Gao Xiang
2026-09-25 10:27 ` Gao Xiang [this message]
2026-09-25 10:13 ` SJ Park
2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41 ` Gao Xiang
2026-09-29 12:50 ` Jialiang Huang
2026-09-30 3:36 ` Muchun Song
2026-09-30 7:24 ` Gao Xiang
2026-09-30 9:37 ` Muchun Song
2026-09-30 10:33 ` Gao Xiang
2026-09-29 16:56 ` SJ Park
2026-09-30 3:07 ` Lian Wang
2026-09-30 8:08 ` SJ Park
2026-09-29 18:14 ` Pratyush Mallick
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arZMfTVu4t4hdHKD@MacBookPro.fritz.box \
--to=xiang@kernel.org \
--cc=baohua@kernel.org \
--cc=damon@lists.linux.dev \
--cc=david@kernel.org \
--cc=kunwu.chan@gmail.com \
--cc=lance.yang@linux.dev \
--cc=lianux.mm@gmail.com \
--cc=linux-mm@kvack.org \
--cc=mst@redhat.com \
--cc=muchun.song@linux.dev \
--cc=research@deepseek.com \
--cc=ryncsn@gmail.com \
--cc=sj@kernel.org \
--cc=virtualization@lists.linux.dev \
--cc=xueyuan.chen21@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox