From: Jeff Layton <jlayton@kernel.org>
To: Benjamin Coddington <ben.coddington@hammerspace.com>,
Chuck Lever <cel@kernel.org>, NeilBrown <neil@brown.name>
Cc: linux-nfs@vger.kernel.org, Daire Byrne <daire@dneg.com>,
bpf@vger.kernel.org, Martin KaFai Lau <martin.lau@linux.dev>
Subject: Re: [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF
Date: Thu, 08 Oct 2026 20:48:10 +0200 [thread overview]
Message-ID: <a0f7d257143bee30db0cda0ca186b06443710bf7.camel@kernel.org> (raw)
In-Reply-To: <cover.1791402701.git.bcodding@hammerspace.com>
On Wed, 2026-10-07 at 15:59 -0400, Benjamin Coddington wrote:
> From: Benjamin Coddington <bcodding@hammerspace.com>
>
> This is v2 of [3], with the change I floated in that thread: the kernel
> no longer decides which transports belong together. By default nothing
> changes. An administrator who wants peers grouped loads a small BPF
> program that returns a class for each accepted connection, and the pool
> takes turns across classes.
>
> What's different from v1
> ------------------------
>
> The table I described on [3] became a hook. Instead of a kernel table
> of address prefixes, a netlink interface for it and an nfs-utils side,
> sunrpc registers a struct_ops with one callback: given the accepted
> svc_xprt, return a class. 0 means the anonymous client (shared with
> everything unclassified, today's order among them); any other number is
> a class every transport returning it shares; a flag bit on top means
> "and one client per peer address within that class". So "all movers are
> one class, everybody else per host" is two entries in a prefix map, and
> the posted v1 behaviour is one entry.
>
> The hook runs once per accepted connection, in process context, after
> TCP and RDMA have both set the peer address. Nothing runs per request.
> No program loaded means every transport is on the
> anonymous client, and with patch 6 the pool uses its old flat queue
> until a classifier attaches, so the dispatch path is today's code for
> anyone who hasn't asked for anything. That's the answer to Chuck's cost
> question from [3]: measured on bare metal the two-level queue cost about
> half a microsecond a dispatch (1.7% at 484k IOPS on one client's eight
> connections); with no program it now costs a static branch on enqueue
> and an empty-queue test on dequeue.
>
> Attaching is a struct_ops link, bound to the attaching task's network
> namespace, one per namespace. A second attach gets -EBUSY, a link
> update replaces the program, closing the link removes it, and a
> namespace that exits leaves its link with nothing behind. Transports
> keep the class they got at accept; a new program only affects new
> connections. The reference program and the attach tests are in the BPF
> selftests (patch 7). svc-classify (patch 8) attaches the prefix
> classifier and manages its map by address prefix, so an administrator
> never touches bpftool. The page in patch 9,
> Documentation/filesystems/nfs/rpc-server-clients.rst, has the class
> word, the install procedure, and six class maps with the share each one
> gives its clients.
>
> Still not in here: moving a transport between clients, so a v4.1
> clientid key, which Chuck asked about on [3], is not in this version.
> A multi-homed host or an IPv6 host with a prefix's worth of addresses
> is handled by the prefix map instead.
>
> Also from the [3] review: control events (accept, close, handshake) go
> ahead of data within a client (patch 5), and the dequeue tracepoint
> prints the class (patch 4).
>
> Numbers
> -------
>
> The dispatch numbers are the v1 series' on a two-socket box (2 x Xeon
> 6542Y, 96 threads, no KASAN), from [3]. The dispatch code is the same
> here: v2 puts the classifier in front of it, and the flat queue behind
> a static branch when nothing is attached. Loopback, NFSv3, 16 threads,
> 10 ms injected service time with 50% jitter. A is v7.2, B the series.
> v4.1 is within 2% in every cell.
>
> One interactive client (bursts of 32) against one host with K
> backlogged connections, burst completion p50 in ms, floor 37:
>
> K 4 8 16 32
> A 70 195 351 671
> B 56 62 61 63
>
> N (K=16) 1 8 32 64 96 128
> A 21 98 350 690 1035 1365
> B 11 31 62 103 144 179
>
> M hosts x 4 1 2 4 8
> A 70 192 351 670
> B 57 81 121 192
>
> S: 2x8 vs K, share 4 8 16 32
> A 41.0 20.0 11.1 5.9
> B 50.3 50.0 50.0 50.0
>
> A real NFSv3 client walking 500 files against a 16-connection
> aggressor: 37.8 s on A, 2.5 s on B, 0.12 s alone on both.
>
> Classes. The v1 kernel keyed on source address, so on [3] I emulated
> "all movers are one class" by giving six movers one address. With this
> series that's one map entry (the movers' prefix -> 1, everything else
> per address), and the six-address column is what "everything per
> address" gives. Six movers at 4 or 8 connections each against
> customers, customer share of dispatches:
>
> A per address movers one class
> one customer, 1 conn x 16 4.0 / 2.0 14.3 / 14.3 49.9 / 49.9
> one customer, 4 x 4 14.3 / 7.7 14.3 / 14.3 49.8 / 49.7
> four customers, 4 x 4 each 40 / 25 40 / 40 80 / 80
>
> Burst of 32 against M busy hosts, p50 ms: hosts on their own addresses
> 57 81 121 192, as one class 56 62 62 62. A: 71 193 348 673 either way.
>
> Cost. With a classifier attached, dispatch takes two or three more
> lock-free queue operations per request. On the same box that's about half a
> microsecond when the enqueue and the dequeue run on different cores,
> fio 4k O_DIRECT randread over loopback, 5 x 20s:
>
> A B
> nconnect=1, 16 jobs 152.8k 152.2k -0.4%
> nconnect=8, 16 jobs 483.8k 475.5k -1.7%
> nconnect=8, 64 jobs 584.2k 581.0k -0.5%
>
> The serial walk and 4 jobs on one connection don't move. With no
> classifier attached the pool is on its old flat queue (patch 6). I only
> have my VM for that path so far (KASAN, 10 vCPUs, loopback, the same
> fio), and there the no-program kernel measures the same as v7.2 within
> the run-to-run spread: 16 jobs 94.9k vs 95.5k, 64 jobs 138.6k vs 139.0k.
> So the static branch adds nothing I can see, but that's a VM number.
>
> Not tested: RDMA. The hook sits in the accept path TCP and RDMA share,
> but I could not bring a soft-RoCE listener up on the VM.
>
> Daire - the same nine patches on v7.2 (plus the wake fix) are at
> https://github.com/bcodding/linux branch nfsd-clientq-v2-7.2 if you
> want to run it; the loader builds from tools/net/sunrpc/svc-classify in
> that tree.
>
> The series is on nfsd-testing at 56589cdb5881, which already has the
> svc_clean_up_xprts() wake fix from [3].
>
> patch 1 track clients by class, no change to dispatch
> patch 2 dispatch round-robin across clients
> patch 3 the struct_ops hook
> patch 4 class in the svc_xprt_dequeue tracepoint
> patch 5 control events ahead of data within a client
> patch 6 flat queue while no classifier is attached
> patch 7 selftests: the prefix program and the attach tests
> patch 8 tools/net/sunrpc/svc-classify: the prefix classifier and its
> loader (load/replace/unload, add/del/list by prefix, status)
> patch 9 Documentation: the model, the class word, installing a
> classifier with svc-classify, six class maps and the share
> each gives, how it works underneath
>
> [1] https://lore.kernel.org/linux-nfs/cover.1780498019.git.bcodding@hammerspace.com
> [2] https://lore.kernel.org/linux-nfs/cover.1782314746.git.bcodding@hammerspace.com
> [3] https://lore.kernel.org/linux-nfs/cover.1790953694.git.bcodding@hammerspace.com
>
> Benjamin Coddington (9):
> SUNRPC: track service clients by class
> SUNRPC: dispatch ready transports round-robin across clients
> SUNRPC: add a BPF struct_ops hook to classify accepted transports
> SUNRPC: report the transport's class in the svc_xprt_dequeue
> tracepoint
> SUNRPC: dispatch control events ahead of data within a client
> SUNRPC: keep the single transport queue while no classifier is
> attached
> selftests/bpf: add svc_classifier tests
> tools/net/sunrpc: add svc-classify, a prefix classifier and its loader
> Documentation: describe RPC server transport classes and the BPF
> classifier
>
> Documentation/filesystems/nfs/index.rst | 1 +
> .../filesystems/nfs/rpc-server-clients.rst | 394 ++++++++++++++++++
> include/linux/sunrpc/svc.h | 72 +++-
> include/linux/sunrpc/svc_xprt.h | 30 ++
> include/trace/events/sunrpc.h | 11 +-
> net/sunrpc/Kconfig | 12 +
> net/sunrpc/Makefile | 1 +
> net/sunrpc/netns.h | 4 +
> net/sunrpc/sunrpc_syms.c | 4 +
> net/sunrpc/svc.c | 36 ++
> net/sunrpc/svc_classify.c | 210 ++++++++++
> net/sunrpc/svc_xprt.c | 281 ++++++++++++-
> tools/net/sunrpc/svc-classify/.gitignore | 1 +
> tools/net/sunrpc/svc-classify/Makefile | 65 +++
> tools/net/sunrpc/svc-classify/svc-classify.c | 357 ++++++++++++++++
> .../sunrpc/svc-classify/svc_classify.bpf.c | 83 ++++
> tools/testing/selftests/bpf/config | 2 +
> .../selftests/bpf/prog_tests/svc_classifier.c | 104 +++++
> .../selftests/bpf/progs/bpf_svc_classifier.c | 80 ++++
> 19 files changed, 1724 insertions(+), 24 deletions(-)
> create mode 100644 Documentation/filesystems/nfs/rpc-server-clients.rst
> create mode 100644 net/sunrpc/svc_classify.c
> create mode 100644 tools/net/sunrpc/svc-classify/.gitignore
> create mode 100644 tools/net/sunrpc/svc-classify/Makefile
> create mode 100644 tools/net/sunrpc/svc-classify/svc-classify.c
> create mode 100644 tools/net/sunrpc/svc-classify/svc_classify.bpf.c
> create mode 100644 tools/testing/selftests/bpf/prog_tests/svc_classifier.c
> create mode 100644 tools/testing/selftests/bpf/progs/bpf_svc_classifier.c
My Comments:
Conceptually, I like this idea, and I personally like that I now have a
clear (to me) example of how to plug eBPF programs into kernel code, so
thanks for that!
I think the userland bits should be made part of nfs-utils, and the
LLM-slop docs need to be rewritten for humans.
--
Jeff Layton <jlayton@kernel.org>
prev parent reply other threads:[~2026-10-08 18:48 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 1/9] SUNRPC: track service clients by class Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 2/9] SUNRPC: dispatch ready transports round-robin across clients Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 3/9] SUNRPC: add a BPF struct_ops hook to classify accepted transports Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 4/9] SUNRPC: report the transport's class in the svc_xprt_dequeue tracepoint Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 5/9] SUNRPC: dispatch control events ahead of data within a client Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 6/9] SUNRPC: keep the single transport queue while no classifier is attached Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 7/9] selftests/bpf: add svc_classifier tests Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader Benjamin Coddington
2026-10-08 18:36 ` Jeff Layton
2026-10-08 19:38 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier Benjamin Coddington
2026-10-08 18:42 ` Jeff Layton
2026-10-08 19:40 ` Benjamin Coddington
2026-10-08 15:19 ` [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Chuck Lever
2026-10-08 15:28 ` Benjamin Coddington
2026-10-08 18:48 ` Jeff Layton [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=a0f7d257143bee30db0cda0ca186b06443710bf7.camel@kernel.org \
--to=jlayton@kernel.org \
--cc=ben.coddington@hammerspace.com \
--cc=bpf@vger.kernel.org \
--cc=cel@kernel.org \
--cc=daire@dneg.com \
--cc=linux-nfs@vger.kernel.org \
--cc=martin.lau@linux.dev \
--cc=neil@brown.name \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox