BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next v5 0/2] bpftool: Batch bounded hash map dumps
@ 2026-09-24 16:34 Tianyi Chen
  2026-09-24 16:34 ` [PATCH bpf-next v5 1/2] bpftool: Use batch lookups for " Tianyi Chen
  2026-09-24 16:34 ` [PATCH bpf-next v5 2/2] selftests/bpf: Check bpftool batch map dump contents Tianyi Chen
  0 siblings, 2 replies; 5+ messages in thread
From: Tianyi Chen @ 2026-09-24 16:34 UTC (permalink / raw)
  To: bpf; +Cc: qmo, andrii, eddyz87, ihor.solodrai, linux-kselftest

Use lookup batches for ordinary hash maps whose maximum key/value
storage fits within 4 MiB, preserving output and conservative fallback.

Changes in v5, addressing Quentin's review:
- Rebase onto current bpf-next and preserve the updated per-CPU printing
  interfaces. Eligibility, batching and fallback logic are unchanged.
- Replace the old performance paragraph in patch 1 with fresh matched-base
  measurements redirecting stdout to /dev/null. Median elapsed time falls
  by 37.7% for plain output and 41.2% for JSON in this fixture; separately
  measured BPF syscall counts remain 200,004 versus 395.
- The selftest patch is unchanged.

Validation:
- Built baseline and patched bpftool from the same base with GCC 16.2.1.
- All 11 bpftool_map_batch subtests pass in an x86-64 KVM guest.
- The 100,000-entry fixture produces byte-identical baseline/patched JSON.
  Plain and JSON per-CPU fallback output also matches the baseline.
- bpftool synchronization and diff checks pass.

The performance test uses Linux 7.2.5, 4-byte keys/values, one pinned vCPU,
two warm-ups and 15 alternating untraced runs per binary and output mode.
Patch 1 includes timings and methodology. This rerun uses a different
base and guest environment from the old measurement, so the change in
speedup is not attributed solely to output redirection.

v4: https://lore.kernel.org/r/20260911034732.219752-1-diannaaav@gmail.com
Review: https://lore.kernel.org/r/141264c9-2d05-4c05-a76a-306818e854ba@kernel.org

Tianyi Chen (2):
  bpftool: Use batch lookups for bounded hash map dumps
  selftests/bpf: Check bpftool batch map dump contents

 tools/bpf/bpftool/map.c                       | 117 ++++++++++-
 .../bpf/prog_tests/bpftool_map_batch.c        | 187 ++++++++++++++++++
 2 files changed, 296 insertions(+), 8 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/bpftool_map_batch.c


base-commit: 0e4cf80d0d4893d8227ba816d0559ab778af8125
-- 
2.55.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH bpf-next v5 1/2] bpftool: Use batch lookups for bounded hash map dumps
  2026-09-24 16:34 [PATCH bpf-next v5 0/2] bpftool: Batch bounded hash map dumps Tianyi Chen
@ 2026-09-24 16:34 ` Tianyi Chen
  2026-09-25 17:05   ` Quentin Monnet
  2026-09-24 16:34 ` [PATCH bpf-next v5 2/2] selftests/bpf: Check bpftool batch map dump contents Tianyi Chen
  1 sibling, 1 reply; 5+ messages in thread
From: Tianyi Chen @ 2026-09-24 16:34 UTC (permalink / raw)
  To: bpf; +Cc: qmo, andrii, eddyz87, ihor.solodrai, linux-kselftest

Use BPF_MAP_LOOKUP_BATCH when dumping hash maps to reduce the number
of BPF syscalls while preserving plain, JSON and BTF formatting.

Compare against the unpatched base using a populated 100,000-entry hash
map with 4-byte keys and values, with stdout redirected to /dev/null.
In an x86-64 KVM guest running Linux 7.2.5, both builds use GCC 16.2.1
and run on one vCPU. After two warm-ups per binary and output mode,
alternate baseline and patched runs for 15 untraced samples each:

                 baseline       batch       elapsed-time reduction
  plain          88.165 ms      54.951 ms    37.7%
  JSON           81.973 ms      48.216 ms    41.2%

These are median wall times, including process startup and formatting.
Separately, strace counts 200,004 versus 395 BPF syscalls for the plain
dump. Complete JSON output is identical for this map. These results are
specific to this fixture and guest, not a general speedup guarantee.

Hash batch lookup must fit an entire bucket. Restrict eligibility to
maps whose maximum key/value storage fits in 4 MiB, so even a worst-case
bucket can fit without restarting a partially printed dump. Start with
up to 256 entries and grow on ENOSPC using the same input cursor.

Fall back to individual lookups only when the initial batch operation
is unsupported. Restarting after output has begun would duplicate
entries. Process the final partial batch on ENOENT, but do not trust
count or output buffers after other errors. Keep fatal diagnostics on
stderr so JSON element arrays contain only map entries.

Link: https://github.com/libbpf/bpftool/issues/63

Assisted-by: LLM
Signed-off-by: Tianyi Chen <hi@tychen.cc>
---
 tools/bpf/bpftool/map.c | 117 +++++++++++++++++++++++++++++++++++++---
 1 file changed, 109 insertions(+), 8 deletions(-)

diff --git a/tools/bpf/bpftool/map.c b/tools/bpf/bpftool/map.c
index 20d59eab09a1..b12c06a6f755 100644
--- a/tools/bpf/bpftool/map.c
+++ b/tools/bpf/bpftool/map.c
@@ -741,15 +741,10 @@ static int do_show(int argc, char **argv)
 	return errno == ENOENT ? 0 : -1;
 }
 
-static int dump_map_elem(int fd, void *key, void *value,
-			 struct bpf_map_info *map_info, struct btf *btf,
-			 json_writer_t *btf_wtr, const int *cpu_ids, int cpu_cnt)
+static void print_map_elem(void *key, void *value,
+			   struct bpf_map_info *map_info, struct btf *btf,
+			   json_writer_t *btf_wtr, const int *cpu_ids, int cpu_cnt)
 {
-	if (bpf_map_lookup_elem(fd, key, value)) {
-		print_entry_error(map_info, key, errno);
-		return -1;
-	}
-
 	if (json_output) {
 		print_entry_json(map_info, key, value, btf, cpu_ids, cpu_cnt);
 	} else if (btf) {
@@ -763,10 +758,112 @@ static int dump_map_elem(int fd, void *key, void *value,
 	} else {
 		print_entry_plain(map_info, key, value, cpu_ids, cpu_cnt);
 	}
+}
+
+static int dump_map_elem(int fd, void *key, void *value,
+			 struct bpf_map_info *map_info, struct btf *btf,
+			 json_writer_t *btf_wtr, const int *cpu_ids, int cpu_cnt)
+{
+	if (bpf_map_lookup_elem(fd, key, value)) {
+		print_entry_error(map_info, key, errno);
+		return -1;
+	}
 
+	print_map_elem(key, value, map_info, btf, btf_wtr, cpu_ids, cpu_cnt);
 	return 0;
 }
 
+#define MAP_DUMP_BATCH_FALLBACK 1
+#define MAP_DUMP_BATCH_SIZE 256U
+#define MAP_DUMP_BATCH_MAX_BYTES (4 * 1024 * 1024)
+
+/* Return MAP_DUMP_BATCH_FALLBACK only before batch traversal starts. */
+static int dump_map_batch(int fd, void *key, void *value,
+			  struct bpf_map_info *info, struct btf *btf,
+			  json_writer_t *wtr, unsigned int *num_elems)
+{
+	__u32 capacity, count, batch = 0, next_batch = 0, i;
+	void *keys = NULL, *values = NULL, *buf;
+	bool first = true, can_fallback = true;
+	int err;
+
+	/*
+	 * Hash lookup batches must accommodate a whole bucket. Restrict the
+	 * optimization to maps whose worst-case bucket fits the memory budget,
+	 * so a later ENOSPC never forces a restart after printing some entries.
+	 * Division also bounds the allocation multiplications on 32-bit hosts.
+	 */
+	if (info->type != BPF_MAP_TYPE_HASH || !info->max_entries ||
+	    (__u64)info->key_size + info->value_size >
+	    MAP_DUMP_BATCH_MAX_BYTES / info->max_entries)
+		return MAP_DUMP_BATCH_FALLBACK;
+
+	capacity = min(info->max_entries, MAP_DUMP_BATCH_SIZE);
+resize:
+	buf = realloc(keys, (size_t)capacity * info->key_size);
+	if (!buf) {
+		err = ENOMEM;
+		goto error;
+	}
+	keys = buf;
+	buf = realloc(values, (size_t)capacity * info->value_size);
+	if (!buf) {
+		err = ENOMEM;
+		goto error;
+	}
+	values = buf;
+
+	while (true) {
+		count = capacity;
+		err = bpf_map_lookup_batch(fd, first ? NULL : &batch,
+					   &next_batch, keys, values, &count, NULL);
+		err = err ? errno : 0;
+		/*
+		 * Older kernels reject the command before updating count. Do not
+		 * inspect the buffers on these errors, or fall back after progress.
+		 */
+		if (can_fallback && (err == EINVAL || err == EOPNOTSUPP ||
+				     err == 524 /* ENOTSUPP */)) {
+			err = MAP_DUMP_BATCH_FALLBACK;
+			goto out;
+		}
+		can_fallback = false;
+		if (err == ENOSPC) {
+			if (capacity == info->max_entries)
+				goto error;
+			capacity += min(capacity, info->max_entries - capacity);
+			/* Preserve the input cursor: the oversized bucket was not read. */
+			goto resize;
+		}
+		/* In particular, EFAULT can leave count and the buffers invalid. */
+		if (err && err != ENOENT)
+			goto error;
+		for (i = 0; i < count; i++) {
+			/*
+			 * Keep the alignment provided by individual lookups, including
+			 * for BTF types whose map key/value size is not aligned.
+			 */
+			memcpy(key, keys + (size_t)i * info->key_size, info->key_size);
+			memcpy(value, values + (size_t)i * info->value_size, info->value_size);
+			print_map_elem(key, value, info, btf, wtr, NULL, 0);
+			(*num_elems)++;
+		}
+		if (err == ENOENT) {
+			err = 0;
+			goto out;
+		}
+		first = false;
+		batch = next_batch;
+	}
+error:
+	fprintf(stderr, "Error: can't lookup map batch: %s\n", strerror(err));
+	err = -1;
+out:
+	free(keys);
+	free(values);
+	return err;
+}
+
 static int maps_have_btf(int *fds, int nb_fds)
 {
 	struct bpf_map_info info = {};
@@ -880,6 +977,9 @@ map_dump(int fd, struct bpf_map_info *info, json_writer_t *wtr,
 		p_info("Warning: cannot read values from %s map with value_size != 8",
 		       map_type_str);
 	}
+	err = dump_map_batch(fd, key, value, info, btf, wtr, &num_elems);
+	if (err != MAP_DUMP_BATCH_FALLBACK)
+		goto end_dump;
 	while (true) {
 		err = bpf_map_get_next_key(fd, prev_key, key);
 		if (err) {
@@ -893,6 +993,7 @@ map_dump(int fd, struct bpf_map_info *info, json_writer_t *wtr,
 		prev_key = key;
 	}
 
+end_dump:
 	if (wtr) {
 		jsonw_end_array(wtr);	/* elements */
 		if (show_header)
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH bpf-next v5 2/2] selftests/bpf: Check bpftool batch map dump contents
  2026-09-24 16:34 [PATCH bpf-next v5 0/2] bpftool: Batch bounded hash map dumps Tianyi Chen
  2026-09-24 16:34 ` [PATCH bpf-next v5 1/2] bpftool: Use batch lookups for " Tianyi Chen
@ 2026-09-24 16:34 ` Tianyi Chen
  1 sibling, 0 replies; 5+ messages in thread
From: Tianyi Chen @ 2026-09-24 16:34 UTC (permalink / raw)
  To: bpf; +Cc: qmo, andrii, eddyz87, ihor.solodrai, linux-kselftest

Exercise hash map dumps around the initial batch size and across
multiple batches, including empty and single-entry maps. Compare each
complete unordered key/value set with the input data in plain, JSON
and pretty JSON output.

Cover one-byte keys, three-byte values and BTF-formatted maps to catch
cursor sizing, buffer alignment and formatting regressions.

Assisted-by: LLM
Signed-off-by: Tianyi Chen <hi@tychen.cc>
---
 .../bpf/prog_tests/bpftool_map_batch.c        | 187 ++++++++++++++++++
 1 file changed, 187 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/bpftool_map_batch.c

diff --git a/tools/testing/selftests/bpf/prog_tests/bpftool_map_batch.c b/tools/testing/selftests/bpf/prog_tests/bpftool_map_batch.c
new file mode 100644
index 000000000000..b4216ed778ef
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/bpftool_map_batch.c
@@ -0,0 +1,187 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <test_progs.h>
+#include <bpftool_helpers.h>
+#include <bpf/btf.h>
+#include <ctype.h>
+
+#define MAX_ENTRIES 1025
+#define RECORD_SIZE 256
+#define OUTPUT_SIZE (MAX_ENTRIES * RECORD_SIZE + 1024)
+
+struct dump_case {
+	const char *name;
+	unsigned int count;
+	unsigned int key_size;
+	unsigned int value_size;
+	bool btf;
+};
+
+static void hex_bytes(char *out, const void *data, unsigned int size, bool json)
+{
+	const unsigned char *bytes = data;
+	unsigned int i;
+
+	if (json)
+		*out++ = '[';
+	for (i = 0; i < size; i++) {
+		if (json && i)
+			*out++ = ',';
+		out += sprintf(out, json ? "\"0x%02x\"" : "%02x", bytes[i]);
+	}
+	if (json)
+		*out++ = ']';
+	*out = '\0';
+}
+
+static void expected_record(char *record, const struct dump_case *test,
+			    unsigned int index, bool json)
+{
+	__u32 key = index, value = index * 37 + 11;
+	unsigned char short_key = index;
+	char key_hex[64], value_hex[64], formatted[96];
+
+	hex_bytes(key_hex, test->key_size == 1 ? (void *)&short_key : &key,
+		  test->key_size, json);
+	hex_bytes(value_hex, &value, test->value_size, json);
+	snprintf(formatted, sizeof(formatted), "{\"key\":%u,\"value\":%u}",
+		 key, value);
+	if (json && test->btf)
+		snprintf(record, RECORD_SIZE,
+			 "{\"key\":%s,\"value\":%s,\"formatted\":%s}",
+			 key_hex, value_hex, formatted);
+	else if (json)
+		snprintf(record, RECORD_SIZE, "{\"key\":%s,\"value\":%s}",
+			 key_hex, value_hex);
+	else if (test->btf)
+		snprintf(record, RECORD_SIZE, "%s", formatted);
+	else
+		snprintf(record, RECORD_SIZE, "key:%svalue:%s", key_hex, value_hex);
+}
+
+static void check_dump(const struct dump_case *test, __u32 id, bool json, bool pretty)
+{
+	bool array = json || test->btf;
+	char command[MAX_BPFTOOL_CMD_LEN], expected[RECORD_SIZE], footer[64];
+	bool seen[MAX_ENTRIES] = {};
+	char *output, *src, *dst, *cursor;
+	unsigned int i, n;
+	int err;
+
+	output = calloc(1, OUTPUT_SIZE);
+	if (!ASSERT_OK_PTR(output, "alloc_output"))
+		return;
+	snprintf(command, sizeof(command), "%smap dump id %u",
+		 pretty ? "-p " : json ? "-j " : "", id);
+	err = get_bpftool_command_output(command, output, OUTPUT_SIZE);
+	if (!ASSERT_OK(err, "map_dump"))
+		goto out;
+	/*
+	 * Ignore presentation whitespace, but compare complete records and all
+	 * punctuation. Expected contents come only from the input data, never
+	 * from another map walk or bpftool invocation.
+	 */
+	for (src = output, dst = output; *src; src++)
+		if (!isspace((unsigned char)*src))
+			*dst++ = *src;
+	*dst = '\0';
+	cursor = output;
+	if (array) {
+		if (!ASSERT_EQ(*cursor, '[', "array_start"))
+			goto out;
+		cursor++;
+	}
+	for (n = 0; n < test->count; n++) {
+		if (array && n) {
+			if (!ASSERT_EQ(*cursor, ',', "record_separator"))
+				goto out;
+			cursor++;
+		}
+		for (i = 0; i < test->count; i++) {
+			if (seen[i])
+				continue;
+			expected_record(expected, test, i, json);
+			if (!strncmp(cursor, expected, strlen(expected)))
+				break;
+		}
+		if (!ASSERT_LT(i, test->count, "unique_expected_record"))
+			goto out;
+		seen[i] = true;
+		cursor += strlen(expected);
+	}
+	if (array) {
+		ASSERT_STREQ(cursor, "]", "array_end_and_count");
+	} else {
+		snprintf(footer, sizeof(footer), "Found%uelement%s", test->count,
+			 test->count == 1 ? "" : "s");
+		ASSERT_STREQ(cursor, footer, "plain_count");
+	}
+out:
+	free(output);
+}
+
+static void run_dump_case(const struct dump_case *test)
+{
+	LIBBPF_OPTS(bpf_map_create_opts, opts);
+	struct bpf_map_info info = {};
+	__u32 info_len = sizeof(info);
+	struct btf *btf = NULL;
+	unsigned int i;
+	int fd = -1;
+
+	if (test->btf) {
+		btf = btf__new_empty();
+		if (!ASSERT_OK_PTR(btf, "btf_new"))
+			return;
+		if (!ASSERT_EQ(btf__add_int(btf, "unsigned int", 4, 0), 1,
+			       "btf_int") ||
+		    !ASSERT_OK(btf__load_into_kernel(btf), "btf_load"))
+			goto out;
+		opts.btf_fd = btf__fd(btf);
+		opts.btf_key_type_id = 1;
+		opts.btf_value_type_id = 1;
+	}
+	fd = bpf_map_create(BPF_MAP_TYPE_HASH, "dump_batch", test->key_size,
+			    test->value_size, test->count ?: 1, &opts);
+	if (!ASSERT_OK_FD(fd, "map_create"))
+		goto out;
+	for (i = 0; i < test->count; i++) {
+		__u32 key = i, value = i * 37 + 11;
+		unsigned char short_key = i;
+		void *key_ptr = test->key_size == 1 ? (void *)&short_key : &key;
+
+		if (!ASSERT_OK(bpf_map_update_elem(fd, key_ptr, &value, BPF_ANY),
+			       "map_update"))
+			goto out;
+	}
+	if (!ASSERT_OK(bpf_map_get_info_by_fd(fd, &info, &info_len), "map_info"))
+		goto out;
+	check_dump(test, info.id, false, false);
+	check_dump(test, info.id, true, false);
+	check_dump(test, info.id, true, true);
+out:
+	if (fd >= 0)
+		close(fd);
+	btf__free(btf);
+}
+
+void test_bpftool_map_batch(void)
+{
+	static const struct dump_case cases[] = {
+		{ "empty", 0, 4, 4 },
+		{ "single", 1, 4, 4 },
+		{ "below_batch", 255, 4, 4 },
+		{ "exact_batch", 256, 4, 4 },
+		{ "above_batch", 257, 4, 4 },
+		{ "multiple_batches", 1025, 4, 4 },
+		{ "one_byte_key", 256, 1, 4 },
+		{ "odd_value_size", 257, 4, 3 },
+		{ "btf_empty", 0, 4, 4, true },
+		{ "btf_single", 1, 4, 4, true },
+		{ "btf_multiple_batches", 1025, 4, 4, true },
+	};
+	unsigned int i;
+
+	for (i = 0; i < ARRAY_SIZE(cases); i++)
+		if (test__start_subtest(cases[i].name))
+			run_dump_case(&cases[i]);
+}
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH bpf-next v5 1/2] bpftool: Use batch lookups for bounded hash map dumps
  2026-09-24 16:34 ` [PATCH bpf-next v5 1/2] bpftool: Use batch lookups for " Tianyi Chen
@ 2026-09-25 17:05   ` Quentin Monnet
  2026-09-25 17:46     ` Tianyi Chen
  0 siblings, 1 reply; 5+ messages in thread
From: Quentin Monnet @ 2026-09-25 17:05 UTC (permalink / raw)
  To: Tianyi Chen, bpf; +Cc: andrii, eddyz87, ihor.solodrai, linux-kselftest

2026-09-25 01:34 UTC+0900 ~ Tianyi Chen <hi@tychen.cc>
> Use BPF_MAP_LOOKUP_BATCH when dumping hash maps to reduce the number


Thanks for this!

Why hash maps only? I think there are some other types of map that could
do this too, no? LRU_HASH has a good chance to be highly populated and
sounds like a good candidate. (It's OK to start just with hash maps, I'm
only asking to understand if I missed a particular reason to not include
other types).


> of BPF syscalls while preserving plain, JSON and BTF formatting.
> 
> Compare against the unpatched base using a populated 100,000-entry hash
> map with 4-byte keys and values, with stdout redirected to /dev/null.
> In an x86-64 KVM guest running Linux 7.2.5, both builds use GCC 16.2.1
> and run on one vCPU. After two warm-ups per binary and output mode,
> alternate baseline and patched runs for 15 untraced samples each:
> 
>                  baseline       batch       elapsed-time reduction
>   plain          88.165 ms      54.951 ms    37.7%
>   JSON           81.973 ms      48.216 ms    41.2%


That's nicer than the initial numbers, thanks for re-running your tests.


> 
> These are median wall times, including process startup and formatting.
> Separately, strace counts 200,004 versus 395 BPF syscalls for the plain
> dump. Complete JSON output is identical for this map. These results are
> specific to this fixture and guest, not a general speedup guarantee.
> 
> Hash batch lookup must fit an entire bucket. Restrict eligibility to
> maps whose maximum key/value storage fits in 4 MiB, so even a worst-case
> bucket can fit without restarting a partially printed dump. Start with
> up to 256 entries and grow on ENOSPC using the same input cursor.
> 
> Fall back to individual lookups only when the initial batch operation
> is unsupported. Restarting after output has begun would duplicate
> entries. Process the final partial batch on ENOENT, but do not trust
> count or output buffers after other errors. Keep fatal diagnostics on
> stderr so JSON element arrays contain only map entries.
> 
> Link: https://github.com/libbpf/bpftool/issues/63
> 
> Assisted-by: LLM
> Signed-off-by: Tianyi Chen <hi@tychen.cc>
> ---
>  tools/bpf/bpftool/map.c | 117 +++++++++++++++++++++++++++++++++++++---
>  1 file changed, 109 insertions(+), 8 deletions(-)
> 
> diff --git a/tools/bpf/bpftool/map.c b/tools/bpf/bpftool/map.c
> index 20d59eab09a1..b12c06a6f755 100644
> --- a/tools/bpf/bpftool/map.c
> +++ b/tools/bpf/bpftool/map.c
> @@ -741,15 +741,10 @@ static int do_show(int argc, char **argv)
>  	return errno == ENOENT ? 0 : -1;
>  }
>  
> -static int dump_map_elem(int fd, void *key, void *value,
> -			 struct bpf_map_info *map_info, struct btf *btf,
> -			 json_writer_t *btf_wtr, const int *cpu_ids, int cpu_cnt)
> +static void print_map_elem(void *key, void *value,
> +			   struct bpf_map_info *map_info, struct btf *btf,
> +			   json_writer_t *btf_wtr, const int *cpu_ids, int cpu_cnt)
>  {
> -	if (bpf_map_lookup_elem(fd, key, value)) {
> -		print_entry_error(map_info, key, errno);
> -		return -1;
> -	}
> -
>  	if (json_output) {
>  		print_entry_json(map_info, key, value, btf, cpu_ids, cpu_cnt);
>  	} else if (btf) {
> @@ -763,10 +758,112 @@ static int dump_map_elem(int fd, void *key, void *value,
>  	} else {
>  		print_entry_plain(map_info, key, value, cpu_ids, cpu_cnt);
>  	}
> +}
> +
> +static int dump_map_elem(int fd, void *key, void *value,
> +			 struct bpf_map_info *map_info, struct btf *btf,
> +			 json_writer_t *btf_wtr, const int *cpu_ids, int cpu_cnt)
> +{
> +	if (bpf_map_lookup_elem(fd, key, value)) {
> +		print_entry_error(map_info, key, errno);
> +		return -1;
> +	}
>  
> +	print_map_elem(key, value, map_info, btf, btf_wtr, cpu_ids, cpu_cnt);
>  	return 0;
>  }
>  
> +#define MAP_DUMP_BATCH_FALLBACK 1
> +#define MAP_DUMP_BATCH_SIZE 256U
> +#define MAP_DUMP_BATCH_MAX_BYTES (4 * 1024 * 1024)
> +
> +/* Return MAP_DUMP_BATCH_FALLBACK only before batch traversal starts. */
> +static int dump_map_batch(int fd, void *key, void *value,
> +			  struct bpf_map_info *info, struct btf *btf,
> +			  json_writer_t *wtr, unsigned int *num_elems)
> +{
> +	__u32 capacity, count, batch = 0, next_batch = 0, i;


__u32 batch: That's the correct size for hash maps, but some other map
types have different key size so we may have to allocate something of
size map->key_size instead in the future. Can you add a comment about
the size being tied to supported map types, please?


> +	void *keys = NULL, *values = NULL, *buf;
> +	bool first = true, can_fallback = true;
> +	int err;
> +
> +	/*
> +	 * Hash lookup batches must accommodate a whole bucket. Restrict the
> +	 * optimization to maps whose worst-case bucket fits the memory budget,
> +	 * so a later ENOSPC never forces a restart after printing some entries.
> +	 * Division also bounds the allocation multiplications on 32-bit hosts.
> +	 */
> +	if (info->type != BPF_MAP_TYPE_HASH || !info->max_entries ||
> +	    (__u64)info->key_size + info->value_size >
> +	    MAP_DUMP_BATCH_MAX_BYTES / info->max_entries)
> +		return MAP_DUMP_BATCH_FALLBACK;


(I think eligibility would be worth its own function, don't step into
dump_map_batch() if we're not eligible. It would also make it clearer
what are the conditions to be eligible, rather that conditions to _not_
be eligible in the current form.)

So looking at this, this means we must have:

    info->key_size + info->value_size
        <= MAP_DUMP_BATCH_MAX_BYTES / info->max_entries

... in order to make the map eligible to batch dump, in other words: the
bigger the map, the less likely it is to go through batch dump (although
batch dump precisely becomes interesting for larger maps).

I understand that MAP_DUMP_BATCH_MAX_BYTES mostly bounds the allocation
we are ready to do; could we make it a limit for the "buffer" resizing,
rather than eligibility? Or am I missing something?

Quentin

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH bpf-next v5 1/2] bpftool: Use batch lookups for bounded hash map dumps
  2026-09-25 17:05   ` Quentin Monnet
@ 2026-09-25 17:46     ` Tianyi Chen
  0 siblings, 0 replies; 5+ messages in thread
From: Tianyi Chen @ 2026-09-25 17:46 UTC (permalink / raw)
  To: Quentin Monnet, bpf; +Cc: andrii, eddyz87, ihor.solodrai, linux-kselftest

Hi Quentin,

Thanks for the review.

> Why hash maps only?

This was an initial scope choice, not a lack of lookup-batch support for
LRU_HASH. It uses the same u32 bucket cursor and is a good candidate for
follow-up coverage, including eviction and concurrent updates. For now,
other map types retain the existing individual-lookup path.

> Can you add a comment about

Added a comment tying the u32 cursor to HASH bucket indices, independently
of the map's key size.

> eligibility would be worth its own function

Done: map_dump_can_batch() is a positive predicate checked by the caller.

> limit for the "buffer" resizing

Yes, the whole-map bound was stricter than necessary. In v6, 4 MiB caps
the combined key/value batch buffers, not the map's maximum capacity.
If ENOSPC persists at that cap, individual lookups continue from the last
emitted key rather than restarting from NULL. If no entries were emitted,
traversal starts at the beginning. This keeps the existing non-atomic
iteration semantics.

All 12 subtests pass, including a large sparse map. I also injected
ENOSPC before and after output, forcing growth to the 4 MiB cap: plain
and JSON output matched the baseline without missing or repeated entries.
Initial unsupported-operation fallback and fatal-error handling were
checked too. Fresh /dev/null timings and methodology are in patch 1.

v6 is a new thread:
https://lore.kernel.org/r/20260925174412.2028749-1-hi@tychen.cc

The v5 CI failure is the same s390x test_verifier callx diagnostic mismatch
seen on for-next_test, not a batch-dump failure. All 11 v5 batch subtests
passed on s390x:
https://github.com/kernel-patches/bpf/actions/runs/36028683139/job/107736923087
https://github.com/kernel-patches/bpf/actions/runs/36040586092/job/107777296289

Thanks,
Tianyi

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-25 17:46 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 16:34 [PATCH bpf-next v5 0/2] bpftool: Batch bounded hash map dumps Tianyi Chen
2026-09-24 16:34 ` [PATCH bpf-next v5 1/2] bpftool: Use batch lookups for " Tianyi Chen
2026-09-25 17:05   ` Quentin Monnet
2026-09-25 17:46     ` Tianyi Chen
2026-09-24 16:34 ` [PATCH bpf-next v5 2/2] selftests/bpf: Check bpftool batch map dump contents Tianyi Chen

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox