All of lore.kernel.org
 help / color / mirror / Atom feed
From: 高翔 <gaoxiang17@xiaomi.com>
To: Andrii Nakryiko <andrii.nakryiko@gmail.com>,
	"David Hildenbrand (Arm)" <david@kernel.org>
Cc: "Xiang Gao" <gxxa03070307@gmail.com>,
	"Andrii Nakryiko" <andrii@kernel.org>,
	"Alexei Starovoitov" <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	"Andrew Morton" <akpm@linux-foundation.org>,
	印闯 <yinchuang1@xiaomi.com>,
	"bpf@vger.kernel.org" <bpf@vger.kernel.org>,
	"linux-mm@kvack.org" <linux-mm@kvack.org>,
	"linux-fsdevel@vger.kernel.org" <linux-fsdevel@vger.kernel.org>,
	"linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
	"Lorenzo Stoakes (Arm)" <ljs@kernel.org>,
	"Steven Rostedt" <rostedt@goodmis.org>
Subject: 答复: [External Mail]Re: [RFC] bpf: account ring buffer backing pages separately from Lost RAM
Date: Tue, 18 Aug 2026 13:46:09 +0000	[thread overview]
Message-ID: <8d2e20842c24460296e4c83e6dc0dde3@xiaomi.com> (raw)
In-Reply-To: <CAEf4BzbwZkoacX-iP1j4+p8YNbWsq0pc3eTermh2uT5ir66zbw@mail.gmail.com>

[-- Attachment #1: Type: text/plain, Size: 5304 bytes --]

Thanks for the pointer. Understood — no new NR_* counter or
/proc/meminfo entry.


The remaining question is on the consumer side: Android's Lost RAM
accounting would need to enumerate all live BPF ringbuf maps
(BPF_MAP_GET_NEXT_ID) and read each map's fdinfo memlock to sum them.


Is per-map fdinfo enumeration the intended way for userspace to get the
aggregate, or is there a more efficient BPF-specific aggregate interface
that I'm missing?



________________________________
发件人: Andrii Nakryiko <andrii.nakryiko@gmail.com>
发送时间: 2026年8月18日 2:22:07
收件人: David Hildenbrand (Arm)
抄送: Xiang Gao; Andrii Nakryiko; Alexei Starovoitov; Daniel Borkmann; Andrew Morton; 印闯; bpf@vger.kernel.org; linux-mm@kvack.org; linux-fsdevel@vger.kernel.org; linux-kernel@vger.kernel.org; 高翔; Lorenzo Stoakes (Arm); Steven Rostedt
主题: [External Mail]Re: [RFC] bpf: account ring buffer backing pages separately from Lost RAM

[外部邮件] 此邮件来源于小米公司外部,请谨慎处理。若对邮件安全性存疑,请将邮件转发给misec@xiaomi.com进行反馈

On Mon, Aug 17, 2026 at 11:10 AM David Hildenbrand (Arm)
<david@kernel.org> wrote:
>
> On 8/15/26 11:18, Xiang Gao wrote:
> > Hi,
>
> Hi,
>
> >
> > I would like to discuss accounting BPF ring buffer backing pages in
> > system-wide memory reports.
> >
> > BPF ring buffers allocate their data and metadata as order-0 pages directly
> > from the buddy allocator, and then map those pages with vmap().
>
> I assume there is a reason the slab isn't used, right? Are these pages mapped
> into user space such that page->mapcount would get used?
>
> Can you point me at relevant code?

See code in [0]. And yes, these pages are meant to be mapped into user space.

  [0] https://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next.git/tree/kernel/bpf/ringbuf.c#n93

>
> >
> > Because vmap() maps caller-owned pages, these backing pages are not counted
> > by VmallocUsed. They are also not slab pages. As a result, most BPF ring
> > buffer memory is not represented by an existing named /proc/meminfo category
> > and appears as Lost RAM in Android memory reports.
> >
> > We measured this on an Android 6.18 kernel.
> >
> > Test case:
> >
> >   32 BPF ring buffer maps
> >   16 MiB data area per map
> >   512 MiB total data area
> >
> > Observed changes:
> >
> >   Lost RAM:       approximately +529 MiB
> >   VmallocUsed:      approximately +2 MiB
> >   Slab:        approximately unchanged
> >
> > After destroying all maps, the values returned close to baseline.
> >
> > The question is whether the kernel should expose the unique physical backing
> > pages of live BPF ring buffers through a dedicated global counter and a
> > /proc/meminfo entry, for example:
> >
> >   BpfRingbuf: <value in kB>
>
> This looks a bit too specific for my taste. And I think we should try to no
> inflate these statistics here too much.
>

+1, way too specific

> >
> > The proposed counter would include:
> >
> >   * ring buffer data pages;
> >   * metadata pages;
> >   * consumer and producer position pages.
> >
> > It would exclude:
> >
> >   * the second virtual mapping of data pages;
> >   * the pages[] pointer array;
> >   * map metadata allocations;
> >   * vmap page tables.
> >
> > The goal is to account for the currently unclassified direct backing pages.
> > Slab- and vmalloc-backed auxiliary allocations are already represented by
> > existing memory categories and should not be counted again.
> >
> > A possible implementation is an NR_BPF_RINGBUF vmstat counter maintained by
> > the ring buffer allocation and free paths, with the aggregate exposed through
> > /proc/meminfo.
> >
> > Questions:
> >
> > 1. Is a dedicated BPF ring buffer counter appropriate?
>
> I don't think so.
>
> See [1] where we just had the same discussion for tracing buffers. For them,
> Steve [2] had an idea on how to expose them more fine-grained and tracing specific.
>
> [1] https://lore.kernel.org/r/20260810094025.136705-1-gaoxiang17@xiaomi.com
> [2] https://lore.kernel.org/r/20260810105710.6ee5e493@gandalf.local.home
>

We already report per-BPF ringbuf memory usage either through bpf()
syscall or map's fdinfo. E.g., with `sudo bpftool map show` you'll
see"

1455733: ringbuf  name event_ringbuf  flags 0x0
        key 0B  value 0B  max_entries 262144  memlock 275776B
        btf_id 2193434
        pids tcpeventd(2549812)

where memlock is how much memory is allocated for the ringbuf data area.

> > 2. Should this be represented as an NR_* vmstat counter?
>
> I don't think so.
>
> > 3. Is /proc/meminfo an acceptable interface for this information?
>
> Again, I don't think so. "Lost RAM" really is just "excessive memory allocated
> by some other subsystem".
>
> I agree that some users might want to figure out what is consuming that much
> memory, but I don't think growing /proc/meminfo in that way is really what we want.
>
> > 4. Is counting only unique physical backing pages the correct accounting unit?
>
> I'd assume the "It would exclude" part above should not be accounted there, if
> that's what you mean.
>
> --
> Cheers,
>
> David

[-- Attachment #2: Type: text/html, Size: 8041 bytes --]

  reply	other threads:[~2026-08-18 13:46 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-15  9:18 [RFC] bpf: account ring buffer backing pages separately from Lost RAM Xiang Gao
2026-08-17 18:10 ` David Hildenbrand (Arm)
2026-08-17 18:22   ` Andrii Nakryiko
2026-08-18 13:46     ` 高翔 [this message]
2026-08-19 17:24       ` [External Mail]Re: " Andrii Nakryiko

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=8d2e20842c24460296e4c83e6dc0dde3@xiaomi.com \
    --to=gaoxiang17@xiaomi.com \
    --cc=akpm@linux-foundation.org \
    --cc=andrii.nakryiko@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=david@kernel.org \
    --cc=gxxa03070307@gmail.com \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=ljs@kernel.org \
    --cc=rostedt@goodmis.org \
    --cc=yinchuang1@xiaomi.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.