Linux-mm Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Gao Xiang <xiang@kernel.org>
To: Lance Yang <lance.yang@linux.dev>
Cc: Gao Xiang <xiang@kernel.org>,
	david@kernel.org, muchun.song@linux.dev, research@deepseek.com,
	sj@kernel.org, mst@redhat.com, damon@lists.linux.dev,
	linux-mm@kvack.org, virtualization@lists.linux.dev,
	ryncsn@gmail.com, kunwu.chan@gmail.com, lianux.mm@gmail.com,
	baohua@kernel.org, xueyuan.chen21@gmail.com
Subject: Re: [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper
Date: Fri, 25 Sep 2026 12:27:09 +0200	[thread overview]
Message-ID: <arZMfTVu4t4hdHKD@MacBookPro.fritz.box> (raw)
In-Reply-To: <arZKFznn4Z2LvoJN@MacBookPro.fritz.box>

On Fri, Sep 25, 2026 at 12:16:55PM +0200, Gao Xiang wrote:
> Hi,
> 
> On Fri, Sep 25, 2026 at 04:15:04PM +0800, Lance Yang wrote:
> > +Cc DeepSeek and Muchun
> > 
> > On Fri, Sep 25, 2026 at 09:04:41AM +0200, David Hildenbrand (Arm) wrote:
> > >On 9/25/26 07:44, Lance Yang wrote:
> > >> Hi all,
> > >> 
> > >> I was reading DeepSeek's new DSec paper[1] and found a nice use of DAMON
> > >> and virtio-balloon:
> > >> 
> > >> With this kind of workload, an agent may read a file once and never touch
> > >> it again, while those pages remain in the guest page cache. Without memory
> > >> pressure in the guest, they can stay cached even though the host would
> > >> like that memory back ...
> > >> 
> > >> The trick is DAMON + virtio-balloon free-page reporting :) DAMON reclaims
> > >
> > >Heh, I read "virtio-balloon" and thought "balloon inflation/deflation, what year
> > >is it?!". Free-page reporting makes much more sense.
> > >
> > >> cold file pages from the guest page cache; buddy gets a chance to coalesce
> > >> them into reportable blocks, and virtio-balloon passes those blocks to
> > >> Firecracker. Firecracker can then drop the host backing with MADV_DONTNEED.
> > >
> > >When I was at RH we were looking at this issue as well. virtio-pmem was one way
> > >of avoiding the page cache in VM entirely. But it has its own limitations.
> > 
> > YES, they enable virtio-pmem with DAX for the read-only EROFS base-image
> > and toolkit layers, while using DAMON with balloon free-page reporting to
> > reclaim cold file pages from the guest page cache on larger writable disks.
> > 
> > They also point out that the guest must allocate struct page metadata for
> > the entire pmem-backed address range. So virtio-pmem is not free either :)
> > 
> > BTW, Muchun recently posted a pretty cool series for exactly that:
> > 
> > https://lore.kernel.org/linux-mm/20260903122128.12264-1-songmuchun@bytedance.com/
> > 
> > (It shares vmemmap backing until a DAX fault needs private metadata,
> > avoiding the full per-PFN cost up front.)
> > 
> > I have a feeling the DeepSeek team will be watching this one closely :P
> 
> There are several points virtio-pmem RW from my own viewpoints:
> 
>  - It makes the write async I/O synchronously, note that write I/Os are
>    not quite the same as read I/Os (read I/Os are mostly sync). Storage
>    also support multi-queues which can better leverage that, and that is
>    why sometimes brd block device is not good at high-performance nvme
>    for example.
> 
>  - Note that sandbox usually has a memory limit, but dax RW makes the
>    whole rootfs addressable, e.g. if you have 128GiB rootfs, which means
>    you could fault 128GiB on the host, instead of the sandbox memory
>    size, so it might cause some security concern (as long as users
>    shouldn't expect the host memory can be used up to 128GiB + memsize).
> 
>  - It can cause sync 4K faults on the host in the worst case (maybe
>    large folios on the host can improve a bit yet not quite), in
>    constant to the guest memory + THP usage. I don't know how the
>    reclaim overhead is measured currently, but it seems the dsec paper
>    also mentioned in this case.
> 
>  - The guest workload will still use mmap() for many sandbox apps, so
>    `struct page` optimization is just for the optimized case, but not
>    for the worst cases, the malicious VM sandboxes can still take
>    much more `struct page` in the guest.
> 
> There would be better to have some benchmark here for typical RL
> training RW virtio-pmem: but block storage semantics cannot already
> be replaced with the memory semantics.
> 
> I think virtio-pmem RO is useful simply because it can reuse the same
> page cache among multiple sandboxes on the host, which can even
> warm-up other sandbox workloads, although it still has some security
> concern but I guess for RL training it doesn't matter and read is
> almost synchronous unlike writes.

BTW, I've thought about the sandbox writable layers for a while, maybe
EROFS could have its own dedicated efficient writable layers as a
optional feature at some time, but I need to think carefully first
and look forward to get more numbers before landing a premature
implementation to the upstream (memory semantics, storage semantics
or just a overlay + hybrid approaches); also there are some
non-technical points to move forward in this direction.

There are some other important features which are more like low-hanging
fruits, so I'm more in a wait-and-see mode until I get a sensible
direction on this sandboxing scenario.

Thanks,
Gao Xiang

> 
> Thanks,
> Gao Xiang 
> 
> > 
> > >Thanks for sharing!
> > 
> > Cheers!
> 


  reply	other threads:[~2026-09-25 10:27 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25  5:44 [FYI] DAMON and virtio-balloon in DeepSeek's DSec paper Lance Yang
2026-09-25  5:57 ` Lian Wang
2026-09-25  6:46   ` KunWu Chan
2026-09-25  7:04 ` David Hildenbrand (Arm)
2026-09-25  7:28   ` Lian Wang
2026-09-25  8:15   ` Lance Yang
2026-09-25 10:03     ` David Hildenbrand (Arm)
2026-09-25 10:16     ` Gao Xiang
2026-09-25 10:27       ` Gao Xiang [this message]
2026-09-25 10:13 ` SJ Park
2026-09-29 12:32 ` Jialiang Huang
2026-09-29 12:41   ` Gao Xiang
2026-09-29 12:50     ` Jialiang Huang
2026-09-30  3:36     ` Muchun Song
2026-09-30  7:24       ` Gao Xiang
2026-09-30  9:37         ` Muchun Song
2026-09-30 10:33           ` Gao Xiang
2026-09-29 16:56   ` SJ Park
2026-09-30  3:07     ` Lian Wang
2026-09-30  8:08       ` SJ Park
2026-09-29 18:14   ` Pratyush Mallick

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arZMfTVu4t4hdHKD@MacBookPro.fritz.box \
    --to=xiang@kernel.org \
    --cc=baohua@kernel.org \
    --cc=damon@lists.linux.dev \
    --cc=david@kernel.org \
    --cc=kunwu.chan@gmail.com \
    --cc=lance.yang@linux.dev \
    --cc=lianux.mm@gmail.com \
    --cc=linux-mm@kvack.org \
    --cc=mst@redhat.com \
    --cc=muchun.song@linux.dev \
    --cc=research@deepseek.com \
    --cc=ryncsn@gmail.com \
    --cc=sj@kernel.org \
    --cc=virtualization@lists.linux.dev \
    --cc=xueyuan.chen21@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox