All of lore.kernel.org
 help / color / mirror / Atom feed
From: Tao Cui <cui.tao@linux.dev>
To: tj@kernel.org, josef@toxicopanda.com, axboe@kernel.dk,
	ameryhung@gmail.com
Cc: cgroups@vger.kernel.org, linux-block@vger.kernel.org,
	linux-kernel@vger.kernel.org, bpf@vger.kernel.org,
	andrii@kernel.org, eddyz87@gmail.com, ast@kernel.org,
	daniel@iogearbox.net, linux-kselftest@vger.kernel.org,
	cui.tao@linux.dev, cuitao@kylinos.cn
Subject: [RFC PATCH v5 4/5] selftests/bpf: add multi-stream sequentiality example model
Date: Fri, 18 Sep 2026 11:17:50 +0800	[thread overview]
Message-ID: <20260918031751.1255420-5-cui.tao@linux.dev> (raw)
In-Reply-To: <20260918031751.1255420-1-cui.tao@linux.dev>

From: Tao Cui <cuitao@kylinos.cn>

Add a second example cost model which replaces the builtin
single-cursor sequentiality heuristic with a per-cgroup table of
stream slots: an IO is sequential iff its sector matches the expected
next sector of any tracked stream.  Interleaved sequential readers in
one cgroup keep their own slots instead of ping-ponging a single
cursor, and random IO inside a hot window rarely matches a moving
expectation.  Merged bios skip the base cost but still advance the
matched stream position, so a merge at the expected sector does not
make the following new IO look random.

Stream state lives in a CGRP_STORAGE map keyed by the cgroup of the
blkcg argument, following the cgroup lifetime; there is no
fixed-size registry to exhaust.

Measured (QEMU, virtio-blk with the HDD profile, 4k IOs, w=1000):
two sequential readers in one cgroup are priced 1961us/op by the
builtin model (judged random) and 23us/op by this model (judged
sequential), the completed IO count rises from 4495 to 207505;
random IO inside an 8M window is priced 24us/op by builtin
(undercharge) and 2607us/op by this model; single-stream sequential
and whole-disk random pricing are unchanged.

Stream updates are lockless like the builtin cursor; a lost update
misclassifies a single IO.  Slots are only advanced for bios the
builtin prices.  The streams test lives here with the model it
loads.  The page count truncates like the builtin.

Signed-off-by: Tao Cui <cuitao@kylinos.cn>
---
 .../selftests/bpf/prog_tests/iocost_model.c   |  34 ++++
 tools/testing/selftests/bpf/progs/iocost_ms.c | 156 ++++++++++++++++++
 2 files changed, 190 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/progs/iocost_ms.c

diff --git a/tools/testing/selftests/bpf/prog_tests/iocost_model.c b/tools/testing/selftests/bpf/prog_tests/iocost_model.c
index 85b5ea7482eed..ac734c0e09cb0 100644
--- a/tools/testing/selftests/bpf/prog_tests/iocost_model.c
+++ b/tools/testing/selftests/bpf/prog_tests/iocost_model.c
@@ -4,6 +4,7 @@
 #include <fcntl.h>
 #include <unistd.h>
 #include "iocost_model.skel.h"
+#include "iocost_ms.skel.h"
 
 /*
  * Write a line to io.cost.model with write(2) and return the errno of
@@ -164,3 +165,36 @@ void serial_test_iocost_model(void)
 
 	iocost_model__destroy(skel);
 }
+
+/*
+ * Same check for the multi-stream example model.  Only one model can
+ * be bound to a device at a time; both tests bind and restore, so
+ * they are serial and independent.
+ */
+void serial_test_iocost_model_streams(void)
+{
+	struct iocost_ms *skel;
+	char *dev;
+	int err;
+
+	dev = getenv("IOCOST_TEST_DEV");
+	if (!dev || geteuid() != 0) {
+		test__skip();
+		return;
+	}
+	if (!dev_has_iocost(dev)) {
+		printf("skip: %s has no iocost enabled\n", dev);
+		test__skip();
+		return;
+	}
+
+	skel = iocost_ms__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "skel_open_load"))
+		return;
+
+	err = iocost_ms__attach(skel);
+	if (ASSERT_OK(err, "attach"))
+		ASSERT_OK(bind_model(dev, "iocost_ms"), "bind_and_readback");
+
+	iocost_ms__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/iocost_ms.c b/tools/testing/selftests/bpf/progs/iocost_ms.c
new file mode 100644
index 0000000000000..70019db0fd790
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/iocost_ms.c
@@ -0,0 +1,156 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Example multi-stream sequentiality detection cost model.
+ *
+ * The builtin model keeps a single cursor per cgroup, so two
+ * interleaved sequential readers in one cgroup are all priced random
+ * (measured 89x overcharge, 12.9x throughput collapse), while random
+ * IO inside a hot window smaller than the 16MB seek threshold is
+ * priced sequential (measured 107x undercharge).  This model replaces
+ * the single cursor with a per-cgroup table of stream slots: an IO is
+ * sequential iff its sector matches the expected next sector of any
+ * tracked stream.  Interleaved streams keep their own slots, and
+ * windowed random IO rarely matches a moving expectation.
+ *
+ * Stream state lives in a CGRP_STORAGE map, so it is created and
+ * freed with the cgroup.  The model implements the full builtin
+ * linear formula itself, including flush pricing.
+ */
+#include "vmlinux.h"
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+
+/*
+ * VTIME_PER_SEC, IOC_PAGE_SIZE/SHIFT, IOC_SECT_TO_PAGE_SHIFT and
+ * IOCOST_COST_F_MERGE come from vmlinux.h (BTF enum constants)
+ */
+#define IOCOST_REQ_OP_MASK	0xff		/* REQ_OP_MASK, not in BTF */
+
+/*
+ * DIV64_U64_ROUND_UP / DIV_ROUND_UP_ULL equivalents, folded at
+ * compile time
+ */
+#define RU(x, y)		((x) / (y) + (((x) % (y)) ? 1 : 0))
+
+#define RBPS	174019176ULL
+#define RSEQIOPS	41708ULL
+#define RRANDIOPS	370ULL
+#define WBPS	178075866ULL
+#define WSEQIOPS	42705ULL
+#define WRANDIOPS	378ULL
+
+#define RPAGE	(RU(VTIME_PER_SEC, RU(RBPS, IOC_PAGE_SIZE)))
+#define RSEQIO	(RU(VTIME_PER_SEC, RSEQIOPS) - RPAGE)
+#define RRANDIO	(RU(VTIME_PER_SEC, RRANDIOPS) - RPAGE)
+#define WPAGE	(RU(VTIME_PER_SEC, RU(WBPS, IOC_PAGE_SIZE)))
+#define WSEQIO	(RU(VTIME_PER_SEC, WSEQIOPS) - WPAGE)
+#define WRANDIO	(RU(VTIME_PER_SEC, WRANDIOPS) - WPAGE)
+
+#define NSLOTS	4
+
+struct streams {
+	__u64 expected[NSLOTS];	/* next expected sector, per stream */
+	__u64 stamp[NSLOTS];	/* LRU stamp, 0 = empty */
+};
+
+/*
+ * per-cgroup stream table: keyed by the cgroup, freed with it
+ */
+struct {
+	__uint(type, BPF_MAP_TYPE_CGRP_STORAGE);
+	__uint(map_flags, BPF_F_NO_PREALLOC);
+	__type(key, int);
+	__type(value, struct streams);
+} stream_tab SEC(".maps");
+
+SEC("struct_ops")
+u64 BPF_PROG(iocost_ms_calc_cost, u64 opf, u64 nbytes, u64 sector,
+	     struct blkcg *blkcg, u64 model_flags)
+{
+	struct streams *s;
+	u64 pages, base, coef_page, randio, advance, now;
+	u32 i, victim = 0, found = 0xFFFFFFFF;
+
+	if ((opf & IOCOST_REQ_OP_MASK) == REQ_OP_READ) {
+		base = RSEQIO; coef_page = RPAGE; randio = RRANDIO;
+	} else if ((opf & IOCOST_REQ_OP_MASK) == REQ_OP_WRITE) {
+		base = WSEQIO; coef_page = WPAGE; randio = WRANDIO;
+	} else {
+		/*
+		 * a fully owning model must price every op; unknown
+		 * ops are priced as per-page writes
+		 */
+		base = 0; coef_page = WPAGE; randio = 0;
+	}
+	advance = RU(nbytes, 512);	/* sectors */
+
+	/* only bios the builtin prices participate in stream tracking */
+	if (!(((opf & IOCOST_REQ_OP_MASK) == REQ_OP_READ ||
+	       (opf & IOCOST_REQ_OP_MASK) == REQ_OP_WRITE) && nbytes)) {
+		pages = nbytes >> IOC_PAGE_SHIFT;
+		if (!pages)
+			pages = 1;
+		return base + pages * coef_page;
+	}
+
+	s = bpf_cgrp_storage_get(&stream_tab, blkcg->css.cgroup, NULL,
+				 BPF_LOCAL_STORAGE_GET_F_CREATE);
+	if (!s) {
+		/* no storage: price per page, truncating like the builtin */
+		pages = nbytes >> IOC_PAGE_SHIFT;
+		if (!pages)
+			pages = 1;
+		return base + pages * coef_page;
+	}
+
+	/*
+	 * Slot access is lockless, mirroring the builtin single-cursor
+	 * update in ioc_rqos_throttle(): concurrent CPUs submitting for
+	 * the same cgroup can race on slot updates; mispricing is
+	 * bounded and acceptable for an example model.
+	 */
+	now = bpf_ktime_get_ns();
+	for (i = 0; i < NSLOTS; i++) {
+		if (s->expected[i] == sector && s->stamp[i]) {
+			found = i;
+			break;
+		}
+	}
+	if (found != 0xFFFFFFFF) {
+		/* sequential: keep the seq base from the op branch */
+		s->expected[found] = sector + advance;
+		s->stamp[found] = now;
+	} else {
+		base = randio;
+		for (i = 1; i < NSLOTS; i++) {
+			if (s->stamp[i] < s->stamp[victim])
+				victim = i;
+		}
+		s->expected[victim] = sector + advance;
+		s->stamp[victim] = now;
+	}
+
+	/* builtin truncates: max(sectors >> IOC_SECT_TO_PAGE_SHIFT, 1) */
+	pages = nbytes >> IOC_PAGE_SHIFT;
+	if (!pages)
+		pages = 1;
+	if (model_flags & IOCOST_COST_F_MERGE) {
+		/*
+		 * merged bios skip the base cost but still advance
+		 * the stream position above, so a merge at the
+		 * expected sector does not make the following new IO
+		 * look random
+		 */
+		base = 0;
+	}
+
+	return base + pages * coef_page;
+}
+
+SEC(".struct_ops")
+struct iocost_model_ops iocost_ms = {
+	.calc_cost = (void *)iocost_ms_calc_cost,
+	.name = "iocost_ms",
+};
+
+char LICENSE[] SEC("license") = "GPL";
-- 
2.43.0


  parent reply	other threads:[~2026-09-18  3:18 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-18  3:17 [RFC PATCH v5 0/5] blk-iocost: BPF struct_ops cost model Tao Cui
2026-09-18  3:17 ` [RFC PATCH v5 1/5] blk-iocost: add BPF struct_ops cost model support Tao Cui
2026-09-18  3:29   ` sashiko-bot
2026-09-18 15:43   ` Alexei Starovoitov
2026-09-19  7:34     ` Tao Cui
2026-09-18  3:17 ` [RFC PATCH v5 2/5] selftests/bpf: add iocost cost model test Tao Cui
2026-09-18  3:17 ` [RFC PATCH v5 3/5] blk-iocost: add iocost_ioc_tick tracepoint for per-period device summary Tao Cui
2026-09-18  3:17 ` Tao Cui [this message]
2026-09-18  3:17 ` [RFC PATCH v5 5/5] docs: cgroup-v2: document io.cost model=<name> binding Tao Cui
2026-09-18  5:48 ` [RFC PATCH v5 0/5] blk-iocost: BPF struct_ops cost model Tao Cui

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260918031751.1255420-5-cui.tao@linux.dev \
    --to=cui.tao@linux.dev \
    --cc=ameryhung@gmail.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=axboe@kernel.dk \
    --cc=bpf@vger.kernel.org \
    --cc=cgroups@vger.kernel.org \
    --cc=cuitao@kylinos.cn \
    --cc=daniel@iogearbox.net \
    --cc=eddyz87@gmail.com \
    --cc=josef@toxicopanda.com \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=tj@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.