All of lore.kernel.org
 help / color / mirror / Atom feed
From: Jay Wang <wanjay@amazon.com>
To: <bpf@vger.kernel.org>, Alexei Starovoitov <ast@kernel.org>,
	"Daniel Borkmann" <daniel@iogearbox.net>,
	Andrii Nakryiko <andrii@kernel.org>,
	"Eduard Zingerman" <eddyz87@gmail.com>,
	Kumar Kartikeya Dwivedi <memxor@gmail.com>
Cc: "Alan Maguire" <alan.maguire@oracle.com>,
	"Martin KaFai Lau" <martin.lau@linux.dev>,
	"Yonghong Song" <yonghong.song@linux.dev>,
	"Jiri Olsa" <jolsa@kernel.org>,
	"Ihor Solodrai" <ihor.solodrai@linux.dev>,
	"Quentin Monnet" <qmo@kernel.org>,
	"Nathan Chancellor" <nathan@kernel.org>,
	"Nicolas Schier" <nsc@kernel.org>,
	linux-kbuild@vger.kernel.org,
	"Thomas Weißschuh" <linux@weissschuh.net>,
	"Christian Heusel" <christian@heusel.eu>,
	"Luis Chamberlain" <mcgrof@kernel.org>,
	"Petr Pavlu" <petr.pavlu@suse.com>,
	"Sami Tolvanen" <samitolvanen@google.com>,
	linux-modules@vger.kernel.org,
	"Steven Rostedt" <rostedt@goodmis.org>,
	"Masami Hiramatsu" <mhiramat@kernel.org>,
	"Mathieu Desnoyers" <mathieu.desnoyers@efficios.com>,
	linux-trace-kernel@vger.kernel.org,
	"Arnaldo Carvalho de Melo" <acme@kernel.org>,
	"Namhyung Kim" <namhyung@kernel.org>,
	"Ian Rogers" <irogers@google.com>,
	linux-perf-users@vger.kernel.org,
	"Jiri Kosina" <jikos@kernel.org>,
	"Benjamin Tissoires" <bentiss@kernel.org>,
	linux-input@vger.kernel.org, "Tejun Heo" <tj@kernel.org>,
	"David Vernet" <void@manifault.com>,
	"Andrea Righi" <arighi@nvidia.com>,
	"Changwoo Min" <changwoo@igalia.com>,
	sched-ext@lists.linux.dev, "Shuah Khan" <shuah@kernel.org>,
	linux-kselftest@vger.kernel.org,
	"Miguel Ojeda" <ojeda@kernel.org>,
	rust-for-linux@vger.kernel.org, "Arnd Bergmann" <arnd@arndb.de>,
	linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org,
	"Hazem Mohamed Abuelfotoh" <abuehaze@amazon.com>,
	"Bjoern Doebel" <doebel@amazon.de>,
	"Martin Pohlack" <mpohlack@amazon.de>,
	jay.wang.upstream@gmail.com
Subject: [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start
Date: Thu, 1 Oct 2026 22:52:07 +0000	[thread overview]
Message-ID: <20261001225214.12351-6-wanjay@amazon.com> (raw)
In-Reply-To: <20261001225214.12351-1-wanjay@amazon.com>

With CONFIG_DEBUG_INFO_BTF=m, bpf_get_btf_vmlinux() and
bpf_find_btf_id() do not load the vmlinux BTF: loading waits for user
space, and their callers were not written for that.  Besides the bpf()
system call and /sys/kernel/btf/vmlinux, some tracefs and bpffs requests
need the BTF.  Load it at the start of those, with
bpf_load_btf_vmlinux(), and have the code that only uses the BTF if it
happens to be there peek:

 - Reading a tracepoint's btf_ids file (events/*/btf_ids): built-in
   events use the vmlinux BTF, and module BTF is only registered once
   that is loaded.  event_btf_ids_read() looks the BTF up under
   event_mutex, which the trace notifier of the btf_vmlinux module
   takes, so load it before taking the mutex, on the first read() of
   the file (not again for the one that returns EOF).
 - Probe events with BTF arguments ($argN, argument names, $retval,
   $current, typecasts): the parser loads the BTF before looking a
   function or struct up.  It holds dyn_event_ops_mutex, which loading
   a module never takes: besides event creation, only
   dyn_event_register() takes it, from built-in init code.  Loading
   here also keeps a $retval from silently losing its type.
 - The ftrace function argument printer (func-args, funcgraph-args)
   runs in the trace output path, which includes ftrace_dump() with
   interrupts disabled.  It only prints the arguments if the BTF is
   already loaded, and never loads it: that also avoids one modprobe
   per trace line when the module is not installed.
 - bpffs: parsing delegate_* mount options (fs_context) loads the BTF
   when a value names commands or types, which are looked up in it;
   "any" and numeric masks need no BTF and do not load it.  Showing the
   options in /proc/*/mountinfo runs under namespace_sem; it only uses
   the names if the BTF is already there and falls back to hex, as it
   already does without BTF.

With CONFIG_DEBUG_INFO_BTF=y bpf_load_btf_vmlinux() is
bpf_get_btf_vmlinux() and the BTF is parsed at boot, so nothing changes.

Signed-off-by: Jay Wang <wanjay@amazon.com>
---
 kernel/bpf/inode.c          | 44 ++++++++++++++++++++++++-------------
 kernel/trace/trace_events.c | 10 +++++++++
 kernel/trace/trace_output.c |  7 ++++++
 kernel/trace/trace_probe.c  | 16 ++++++++++++++
 4 files changed, 62 insertions(+), 15 deletions(-)

diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c
index 7837968c0842..d05bbb61a593 100644
--- a/kernel/bpf/inode.c
+++ b/kernel/bpf/inode.c
@@ -658,7 +658,11 @@ struct bpffs_btf_enums {
 	const struct btf_type *attach_t;
 };
 
-static int find_bpffs_btf_enums(struct bpffs_btf_enums *info)
+/*
+ * @load: load the vmlinux BTF if necessary (CONFIG_DEBUG_INFO_BTF=m), see
+ * bpf_load_btf_vmlinux(); otherwise only use it if it is already parsed.
+ */
+static int find_bpffs_btf_enums(struct bpffs_btf_enums *info, bool load)
 {
 	struct {
 		const struct btf_type **type;
@@ -674,7 +678,7 @@ static int find_bpffs_btf_enums(struct bpffs_btf_enums *info)
 
 	memset(info, 0, sizeof(*info));
 
-	btf = bpf_get_btf_vmlinux();
+	btf = load ? bpf_load_btf_vmlinux() : bpf_peek_btf_vmlinux();
 	if (IS_ERR(btf))
 		return PTR_ERR(btf);
 	if (!btf)
@@ -795,8 +799,11 @@ static int bpf_show_options(struct seq_file *m, struct dentry *root)
 	    opts->delegate_progs || opts->delegate_attachs) {
 		struct bpffs_btf_enums info;
 
-		/* ignore errors, fallback to hex */
-		(void)find_bpffs_btf_enums(&info);
+		/*
+		 * ignore errors, fallback to hex; this runs under
+		 * namespace_sem, so do not load the BTF from here
+		 */
+		(void)find_bpffs_btf_enums(&info, false);
 
 		mask = (1ULL << __MAX_BPF_CMD) - 1;
 		seq_print_delegate_opts(m, "delegate_cmds",
@@ -1052,35 +1059,33 @@ static int bpf_parse_param(struct fs_context *fc, struct fs_parameter *param)
 	case OPT_DELEGATE_MAPS:
 	case OPT_DELEGATE_PROGS:
 	case OPT_DELEGATE_ATTACHS: {
-		struct bpffs_btf_enums info;
-		const struct btf_type *enum_t;
+		struct bpffs_btf_enums info = {};
+		const struct btf_type **enum_t;
+		bool enums_tried = false;
 		const char *enum_pfx;
-		u64 *delegate_msk, msk = 0;
+		u64 *delegate_msk, msk = 0, num;
 		char *p, *str;
 		int val;
 
-		/* ignore errors, fallback to hex */
-		(void)find_bpffs_btf_enums(&info);
-
 		switch (opt) {
 		case OPT_DELEGATE_CMDS:
 			delegate_msk = &opts->delegate_cmds;
-			enum_t = info.cmd_t;
+			enum_t = &info.cmd_t;
 			enum_pfx = "BPF_";
 			break;
 		case OPT_DELEGATE_MAPS:
 			delegate_msk = &opts->delegate_maps;
-			enum_t = info.map_t;
+			enum_t = &info.map_t;
 			enum_pfx = "BPF_MAP_TYPE_";
 			break;
 		case OPT_DELEGATE_PROGS:
 			delegate_msk = &opts->delegate_progs;
-			enum_t = info.prog_t;
+			enum_t = &info.prog_t;
 			enum_pfx = "BPF_PROG_TYPE_";
 			break;
 		case OPT_DELEGATE_ATTACHS:
 			delegate_msk = &opts->delegate_attachs;
-			enum_t = info.attach_t;
+			enum_t = &info.attach_t;
 			enum_pfx = "BPF_";
 			break;
 		default:
@@ -1089,9 +1094,18 @@ static int bpf_parse_param(struct fs_context *fc, struct fs_parameter *param)
 
 		str = param->string;
 		while ((p = strsep(&str, ":"))) {
+			/*
+			 * Only names need the vmlinux BTF: "any" and numbers do
+			 * not load it.  Ignore errors, fallback to hex.
+			 */
+			if (strcmp(p, "any") && kstrtou64(p, 0, &num) && !enums_tried) {
+				(void)find_bpffs_btf_enums(&info, true);
+				enums_tried = true;
+			}
+
 			if (strcmp(p, "any") == 0) {
 				msk |= ~0ULL;
-			} else if (find_btf_enum_const(info.btf, enum_t, enum_pfx, p, &val)) {
+			} else if (find_btf_enum_const(info.btf, *enum_t, enum_pfx, p, &val)) {
 				msk |= 1ULL << val;
 			} else {
 				err = kstrtou64(p, 0, &msk);
diff --git a/kernel/trace/trace_events.c b/kernel/trace/trace_events.c
index 30c0ddf90887..c887ac6a4857 100644
--- a/kernel/trace/trace_events.c
+++ b/kernel/trace/trace_events.c
@@ -23,6 +23,7 @@
 #include <linux/sort.h>
 #include <linux/slab.h>
 #include <linux/delay.h>
+#include <linux/bpf.h>
 #include <linux/btf.h>
 
 #include <trace/events/sched.h>
@@ -2245,6 +2246,15 @@ event_btf_ids_read(struct file *filp, char __user *ubuf, size_t cnt, loff_t *ppo
 	char buf[128];
 	int len;
 
+	/*
+	 * Built-in events use the vmlinux BTF, and with CONFIG_DEBUG_INFO_BTF=m
+	 * module BTF is only registered once that is loaded.  Loading it loads
+	 * a module, whose trace notifier takes event_mutex: load it before
+	 * taking that, and only for the first read, not again for the EOF one.
+	 */
+	if (!*ppos)
+		bpf_load_btf_vmlinux();
+
 	/* Module unload could free call->class and ids[] mid-read. */
 	scoped_guard(mutex, &event_mutex) {
 		file = event_file_file(filp);
diff --git a/kernel/trace/trace_output.c b/kernel/trace/trace_output.c
index a5ad76175d10..1f346e524ee6 100644
--- a/kernel/trace/trace_output.c
+++ b/kernel/trace/trace_output.c
@@ -739,6 +739,13 @@ void print_function_args(struct trace_seq *s, unsigned long *args,
 	if (lookup_symbol_name(func, name))
 		goto out;
 
+	/*
+	 * This can run with interrupts disabled (ftrace_dump()): only use
+	 * the vmlinux BTF if it is parsed, never load it from here.
+	 */
+	if (IS_ERR_OR_NULL(bpf_peek_btf_vmlinux()))
+		goto out;
+
 	/* TODO: Pass module name here too */
 	t = btf_find_func_proto(name, &btf);
 	if (IS_ERR_OR_NULL(t))
diff --git a/kernel/trace/trace_probe.c b/kernel/trace/trace_probe.c
index 804442b2f7d2..d53ee1ef820c 100644
--- a/kernel/trace/trace_probe.c
+++ b/kernel/trace/trace_probe.c
@@ -531,6 +531,19 @@ static const char *fetch_type_from_btf_type(struct btf *btf,
 	return NULL;
 }
 
+/*
+ * Arguments described by BTF need the vmlinux BTF.  With
+ * CONFIG_DEBUG_INFO_BTF=m it may not be loaded yet, so load it before looking
+ * anything up.  Parsing holds dyn_event_ops_mutex, which loading a module
+ * never takes: besides event creation, only dyn_event_register() takes it,
+ * from built-in init code.
+ */
+static void trace_probe_load_btf(void)
+{
+	lockdep_assert_held(&dyn_event_ops_mutex);
+	bpf_load_btf_vmlinux();
+}
+
 static int query_btf_context(struct traceprobe_parse_context *ctx)
 {
 	const struct btf_param *param;
@@ -544,6 +557,7 @@ static int query_btf_context(struct traceprobe_parse_context *ctx)
 	if (!ctx->funcname)
 		return -EINVAL;
 
+	trace_probe_load_btf();
 	type = btf_find_func_proto(ctx->funcname, &btf);
 	if (!type)
 		return -ENOENT;
@@ -762,6 +776,7 @@ static int parse_btf_arg(char *varname,
 	if (!strcmp(varname, "$current")) {
 		code->op = FETCH_OP_CURRENT;
 		/* If no typecast is specified for $current, use task_struct by default */
+		trace_probe_load_btf();
 		ret = bpf_find_btf_id("task_struct", BTF_KIND_STRUCT, &ctx->struct_btf);
 		if (ret < 0) {
 			trace_probe_log_err(ctx->offset, NO_BTF_ENTRY);
@@ -890,6 +905,7 @@ static int query_btf_struct(const char *sname, struct traceprobe_parse_context *
 		ctx->struct_btf = NULL;
 	}
 
+	trace_probe_load_btf();
 	id = bpf_find_btf_id(sname, BTF_KIND_STRUCT, &btf);
 	if (id < 0)
 		return id;
-- 
2.47.3


  parent reply	other threads:[~2026-10-01 22:53 UTC|newest]

Thread overview: 41+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 22:52 [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Jay Wang
2026-10-01 22:52 ` [PATCH bpf-next v4 01/12] bpf: pass the vmlinux BTF to btf_parse_module() and let it adopt the data Jay Wang
2026-10-02  9:14   ` sashiko-bot
2026-10-01 22:52 ` [PATCH bpf-next v4 02/12] bpf: split the kfunc, dtor kfunc and struct_ops registration bodies Jay Wang
2026-10-02  9:14   ` sashiko-bot
2026-10-01 22:52 ` [PATCH bpf-next v4 03/12] bpf: fetch the vmlinux BTF where kernel types enter a program Jay Wang
2026-10-02  9:14   ` sashiko-bot
2026-10-01 22:52 ` [PATCH bpf-next v4 04/12] bpf: take the vmlinux BTF from the btf_vmlinux module Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-02  9:14   ` sashiko-bot
2026-10-02 11:48   ` Alexei Starovoitov
2026-10-01 22:52 ` Jay Wang [this message]
2026-10-02  9:14   ` [PATCH bpf-next v4 05/12] bpf, tracing: load the vmlinux BTF where tracefs and bpffs requests start sashiko-bot
2026-10-01 22:52 ` [PATCH bpf-next v4 06/12] bpf: defer vmlinux kfunc and struct_ops registrations Jay Wang
2026-10-02  9:14   ` sashiko-bot
2026-10-01 22:52 ` [PATCH bpf-next v4 07/12] bpf: keep module BTF until the vmlinux BTF is available Jay Wang
2026-10-01 23:45   ` bot+bpf-ci
2026-10-02  9:14   ` sashiko-bot
2026-10-01 22:52 ` [PATCH bpf-next v4 08/12] bpf: expose deferred .BTF.base module BTF in sysfs from module load Jay Wang
2026-10-02  9:14   ` sashiko-bot
2026-10-01 22:52 ` [PATCH bpf-next v4 09/12] bpf, trace, net: prepare CONFIG_DEBUG_INFO_BTF checks for a tristate Jay Wang
2026-10-02  9:14   ` sashiko-bot
2026-10-05 11:32   ` Nicolas Schier
2026-10-01 22:52 ` [PATCH bpf-next v4 10/12] resolve_btfids: add --btf_link to fill in .BTF.link records Jay Wang
2026-10-01 23:29   ` bot+bpf-ci
2026-10-02  9:14   ` sashiko-bot
2026-10-01 22:52 ` [PATCH bpf-next v4 11/12] tools, samples: take the vmlinux BTF from vmlinux.unstripped first Jay Wang
2026-10-02  9:14   ` sashiko-bot
2026-10-05 11:17   ` Nicolas Schier
2026-10-01 22:52 ` [PATCH bpf-next v4 12/12] kbuild, bpf: allow building the vmlinux BTF as a module Jay Wang
2026-10-02  9:14   ` sashiko-bot
2026-10-02  9:47   ` Alan Maguire
2026-10-02  4:36 ` [PATCH bpf-next v4 00/12] bpf: make the vmlinux BTF an on-demand loadable module (CONFIG_DEBUG_INFO_BTF=m) to save ~5.4 MB memory Ihor Solodrai
2026-10-02  7:34   ` Jay Wang
2026-10-02 10:05     ` Alan Maguire
2026-10-02 20:58     ` Ihor Solodrai
2026-10-03  6:38       ` Alexei Starovoitov
2026-10-03 11:45         ` Alan Maguire
2026-10-03 12:19           ` Alexei Starovoitov
2026-10-04 22:21       ` Jay Wang
2026-10-05 19:01         ` Ihor Solodrai

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001225214.12351-6-wanjay@amazon.com \
    --to=wanjay@amazon.com \
    --cc=abuehaze@amazon.com \
    --cc=acme@kernel.org \
    --cc=alan.maguire@oracle.com \
    --cc=andrii@kernel.org \
    --cc=arighi@nvidia.com \
    --cc=arnd@arndb.de \
    --cc=ast@kernel.org \
    --cc=bentiss@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=changwoo@igalia.com \
    --cc=christian@heusel.eu \
    --cc=daniel@iogearbox.net \
    --cc=doebel@amazon.de \
    --cc=eddyz87@gmail.com \
    --cc=ihor.solodrai@linux.dev \
    --cc=irogers@google.com \
    --cc=jay.wang.upstream@gmail.com \
    --cc=jikos@kernel.org \
    --cc=jolsa@kernel.org \
    --cc=linux-doc@vger.kernel.org \
    --cc=linux-input@vger.kernel.org \
    --cc=linux-kbuild@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=linux-modules@vger.kernel.org \
    --cc=linux-perf-users@vger.kernel.org \
    --cc=linux-trace-kernel@vger.kernel.org \
    --cc=linux@weissschuh.net \
    --cc=martin.lau@linux.dev \
    --cc=mathieu.desnoyers@efficios.com \
    --cc=mcgrof@kernel.org \
    --cc=memxor@gmail.com \
    --cc=mhiramat@kernel.org \
    --cc=mpohlack@amazon.de \
    --cc=namhyung@kernel.org \
    --cc=nathan@kernel.org \
    --cc=nsc@kernel.org \
    --cc=ojeda@kernel.org \
    --cc=petr.pavlu@suse.com \
    --cc=qmo@kernel.org \
    --cc=rostedt@goodmis.org \
    --cc=rust-for-linux@vger.kernel.org \
    --cc=samitolvanen@google.com \
    --cc=sched-ext@lists.linux.dev \
    --cc=shuah@kernel.org \
    --cc=tj@kernel.org \
    --cc=void@manifault.com \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.