BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next v3 0/5] File descriptor interface for BPF streams
@ 2026-09-25  4:55 Kumar Kartikeya Dwivedi
  2026-09-25  4:55 ` [PATCH bpf-next v3 1/5] bpf: Skip zero-length stream writes Kumar Kartikeya Dwivedi
                   ` (6 more replies)
  0 siblings, 7 replies; 11+ messages in thread
From: Kumar Kartikeya Dwivedi @ 2026-09-25  4:55 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Andrii Nakryiko, Daniel Borkmann,
	Eduard Zingerman, Emil Tsalapatis, Tejun Heo, kkd, kernel-team

BPF program stdout and stderr streams can currently be consumed only
through BPF_PROG_STREAM_READ_BY_FD. This requires repeated bpf() calls and
provides no way to block for output or integrate with poll-based event
loops.

Add BPF_PROG_STREAM_OPEN to return a read-only, close-on-exec descriptor
for a program stream. Reads block by default, with an option for
non-blocking mode. poll and epoll report readable data and report hangup
when the program is freed, while allowing buffered data to be drained
before EOF. Stream descriptors retain the stream storage without retaining
the program itself. Waiters are notified through irq_work only when a
publication turns an empty stream readable, as bpf_ringbuf does, so a
program whose stream nobody drains pays for a single notification.

Expose the command through libbpf and add a -w/--wait option to bpftool
prog tracelog that blocks for new output until the program is unloaded or
bpftool is interrupted. The default dump keeps using
BPF_PROG_STREAM_READ_BY_FD and exits once buffered output is drained, so
existing scripts behave the same on old and new kernels. The legacy
command remains supported.

The first patch fixes a pre-existing issue where empty stream writes
allocate elements that escape capacity accounting; the readiness tracking
added later relies on every element carrying data.

See commit logs for details.

Changelog:
----------
v2 -> v3
v2: https://lore.kernel.org/bpf/20260924162641.1922423-1-memxor@gmail.com

 * Queue the wakeup irq_work only when a publication turns an empty stream
   readable instead of on every write, as bpf_ringbuf does. (Alexei)
 * Make waiting for stream output opt-in through a new bpftool -w/--wait
   option; the default dump keeps using BPF_PROG_STREAM_READ_BY_FD and
   exits as before. (Alexei)
 * Exit from the bpftool signal handlers like the trace pipe command instead
   of checking a stop flag that a signal arriving before read() blocks would
   miss, and drop the disposition save and restore. (Alexei)
 * Fail bpftool --wait with a clear error on kernels without
   BPF_PROG_STREAM_OPEN instead of degrading into a dump.
 * Test that output published after a full drain notifies an edge-triggered
   epoll waiter.
 * Rebase on bpf-next.

v1 -> v2
v1: https://lore.kernel.org/bpf/20260830093514.4105972-1-memxor@gmail.com

 * Fold the irq_work and readable-counter patches into the interface patch
   so no intermediate state wakes waiters directly or derives readiness
   from the capacity counter. (BPF CI)
 * Accept only BPF_F_STREAM_NONBLOCK; BPF_F_RDONLY is no longer accepted
   since the descriptor is always read-only.
 * Allocate streams only for programs loaded through BPF_PROG_LOAD, not for
   classic BPF filters, JIT subprograms or shim programs.
 * Add a patch that skips zero-length stream writes instead of allocating
   elements that bypass capacity accounting. (Emil, Sashiko)
 * Report EOF only when the stream was already dead before it was found
   empty, so data published right before program teardown is not lost.
   (Sashiko)
 * Document that POLLHUP follows program destruction, not the caller's own
   release, and that it may lag the final reference drop.
 * Drop the llseek operation so lseek fails with ESPIPE like other stream
   descriptors. (Emil)
 * Drop the unreachable length check in the file read path; the VFS caps
   read sizes below INT_MAX.
 * Let __bpf_prog_free() clean up after a failed stream allocation instead
   of freeing the streams twice, and drop redundant zeroing of the stream
   counters after kzalloc().
 * Explain why wakeups are always deferred through irq_work and drop the
   batching rationale. (Emil)
 * Skip irq_work_sync() at teardown for streams that never queued a
   notification; on PREEMPT_RT it waits for an RCU grace period that every
   program free, including cBPF filters, would otherwise pay. (Sashiko)
 * Qualify the POLLIN followed by EAGAIN claim to a single reader. (BPF CI)
 * Handle SIGINT, SIGHUP and SIGTERM in bpftool so a followed stream exits
   cleanly, and document the follow behavior. (Emil, BPF CI)
 * Drop bpftool's program reference once the stream is open so following
   ends with EOF when the program is unloaded, and restore the previous
   signal dispositions afterwards for batch mode. (Sashiko)
 * Report bpftool read errors on the fallback path as well.
 * Test lseek rejection and empty stream writes.
 * Retry the NMI test write on a later sample if the first attempt fails.
 * Drop Emil's Reviewed-by from the patches that changed.
 * Rebase on bpf-next.

Kumar Kartikeya Dwivedi (5):
  bpf: Skip zero-length stream writes
  bpf: Add file descriptor interface for program streams
  libbpf: Add bpf_prog_stream_open()
  bpftool: Add option to wait for program stream output
  selftests/bpf: Test program stream file descriptors

 include/linux/bpf.h                           |  14 +-
 include/uapi/linux/bpf.h                      |  39 ++
 kernel/bpf/core.c                             |   8 +-
 kernel/bpf/stream.c                           | 220 +++++++-
 kernel/bpf/syscall.c                          |  29 ++
 .../bpftool/Documentation/bpftool-prog.rst    |  12 +-
 tools/bpf/bpftool/bash-completion/bpftool     |   2 +-
 tools/bpf/bpftool/main.c                      |   7 +-
 tools/bpf/bpftool/main.h                      |   1 +
 tools/bpf/bpftool/prog.c                      |  62 ++-
 tools/include/uapi/linux/bpf.h                |  39 ++
 tools/lib/bpf/bpf.c                           |  19 +
 tools/lib/bpf/bpf.h                           |  24 +
 tools/lib/bpf/libbpf.map                      |   1 +
 .../testing/selftests/bpf/prog_tests/stream.c | 486 ++++++++++++++++++
 tools/testing/selftests/bpf/progs/stream.c    |  20 +
 16 files changed, 946 insertions(+), 37 deletions(-)


base-commit: 51455305b0d6ab5a3432e0d7858f4e73f397cc95
-- 
2.53.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-09-25 21:11 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-25  4:55 [PATCH bpf-next v3 0/5] File descriptor interface for BPF streams Kumar Kartikeya Dwivedi
2026-09-25  4:55 ` [PATCH bpf-next v3 1/5] bpf: Skip zero-length stream writes Kumar Kartikeya Dwivedi
2026-09-25  4:55 ` [PATCH bpf-next v3 2/5] bpf: Add file descriptor interface for program streams Kumar Kartikeya Dwivedi
2026-09-25  4:55 ` [PATCH bpf-next v3 3/5] libbpf: Add bpf_prog_stream_open() Kumar Kartikeya Dwivedi
2026-09-25  4:55 ` [PATCH bpf-next v3 4/5] bpftool: Add option to wait for program stream output Kumar Kartikeya Dwivedi
2026-09-25  5:15   ` sashiko-bot
2026-09-25  6:13     ` Kumar Kartikeya Dwivedi
2026-09-25 14:22   ` Quentin Monnet
2026-09-25  4:55 ` [PATCH bpf-next v3 5/5] selftests/bpf: Test program stream file descriptors Kumar Kartikeya Dwivedi
2026-09-25  6:14 ` [PATCH bpf-next v3 0/5] File descriptor interface for BPF streams Kumar Kartikeya Dwivedi
2026-09-25 21:10 ` patchwork-bot+netdevbpf

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox