BPF List
 help / color / mirror / Atom feed
From: Jeff Layton <jlayton@kernel.org>
To: Benjamin Coddington <ben.coddington@hammerspace.com>,
	Chuck Lever <cel@kernel.org>, NeilBrown <neil@brown.name>
Cc: linux-nfs@vger.kernel.org, Daire Byrne <daire@dneg.com>,
	 bpf@vger.kernel.org, Martin KaFai Lau <martin.lau@linux.dev>
Subject: Re: [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF
Date: Thu, 08 Oct 2026 20:48:10 +0200	[thread overview]
Message-ID: <a0f7d257143bee30db0cda0ca186b06443710bf7.camel@kernel.org> (raw)
In-Reply-To: <cover.1791402701.git.bcodding@hammerspace.com>

On Wed, 2026-10-07 at 15:59 -0400, Benjamin Coddington wrote:
> From: Benjamin Coddington <bcodding@hammerspace.com>
> 
> This is v2 of [3], with the change I floated in that thread: the kernel
> no longer decides which transports belong together.  By default nothing
> changes.  An administrator who wants peers grouped loads a small BPF
> program that returns a class for each accepted connection, and the pool
> takes turns across classes.
> 
> What's different from v1
> ------------------------
> 
> The table I described on [3] became a hook.  Instead of a kernel table
> of address prefixes, a netlink interface for it and an nfs-utils side,
> sunrpc registers a struct_ops with one callback: given the accepted
> svc_xprt, return a class.  0 means the anonymous client (shared with
> everything unclassified, today's order among them); any other number is
> a class every transport returning it shares; a flag bit on top means
> "and one client per peer address within that class".  So "all movers are
> one class, everybody else per host" is two entries in a prefix map, and
> the posted v1 behaviour is one entry.
> 
> The hook runs once per accepted connection, in process context, after
> TCP and RDMA have both set the peer address.  Nothing runs per request.
> No program loaded means every transport is on the
> anonymous client, and with patch 6 the pool uses its old flat queue
> until a classifier attaches, so the dispatch path is today's code for
> anyone who hasn't asked for anything.  That's the answer to Chuck's cost
> question from [3]: measured on bare metal the two-level queue cost about
> half a microsecond a dispatch (1.7% at 484k IOPS on one client's eight
> connections); with no program it now costs a static branch on enqueue
> and an empty-queue test on dequeue.
> 
> Attaching is a struct_ops link, bound to the attaching task's network
> namespace, one per namespace.  A second attach gets -EBUSY, a link
> update replaces the program, closing the link removes it, and a
> namespace that exits leaves its link with nothing behind.  Transports
> keep the class they got at accept; a new program only affects new
> connections.  The reference program and the attach tests are in the BPF
> selftests (patch 7).  svc-classify (patch 8) attaches the prefix
> classifier and manages its map by address prefix, so an administrator
> never touches bpftool.  The page in patch 9,
> Documentation/filesystems/nfs/rpc-server-clients.rst, has the class
> word, the install procedure, and six class maps with the share each one
> gives its clients.
> 
> Still not in here: moving a transport between clients, so a v4.1
> clientid key, which Chuck asked about on [3], is not in this version.
> A multi-homed host or an IPv6 host with a prefix's worth of addresses
> is handled by the prefix map instead.
> 
> Also from the [3] review: control events (accept, close, handshake) go
> ahead of data within a client (patch 5), and the dequeue tracepoint
> prints the class (patch 4).
> 
> Numbers
> -------
> 
> The dispatch numbers are the v1 series' on a two-socket box (2 x Xeon
> 6542Y, 96 threads, no KASAN), from [3].  The dispatch code is the same
> here: v2 puts the classifier in front of it, and the flat queue behind
> a static branch when nothing is attached.  Loopback, NFSv3, 16 threads,
> 10 ms injected service time with 50% jitter.  A is v7.2, B the series.
> v4.1 is within 2% in every cell.
> 
> One interactive client (bursts of 32) against one host with K
> backlogged connections, burst completion p50 in ms, floor 37:
> 
>     K             4      8     16     32
>     A            70    195    351    671
>     B            56     62     61     63
> 
>     N (K=16)      1      8     32     64     96    128
>     A            21     98    350    690   1035   1365
>     B            11     31     62    103    144    179
> 
>     M hosts x 4   1      2      4      8
>     A            70    192    351    670
>     B            57     81    121    192
> 
>     S: 2x8 vs K, share   4      8     16     32
>     A                 41.0   20.0   11.1    5.9
>     B                 50.3   50.0   50.0   50.0
> 
> A real NFSv3 client walking 500 files against a 16-connection
> aggressor: 37.8 s on A, 2.5 s on B, 0.12 s alone on both.
> 
> Classes.  The v1 kernel keyed on source address, so on [3] I emulated
> "all movers are one class" by giving six movers one address.  With this
> series that's one map entry (the movers' prefix -> 1, everything else
> per address), and the six-address column is what "everything per
> address" gives.  Six movers at 4 or 8 connections each against
> customers, customer share of dispatches:
> 
>                                  A           per address   movers one class
>     one customer, 1 conn x 16    4.0 / 2.0   14.3 / 14.3   49.9 / 49.9
>     one customer, 4 x 4         14.3 / 7.7   14.3 / 14.3   49.8 / 49.7
>     four customers, 4 x 4 each  40 / 25      40 / 40       80 / 80
> 
> Burst of 32 against M busy hosts, p50 ms: hosts on their own addresses
> 57 81 121 192, as one class 56 62 62 62.  A: 71 193 348 673 either way.
> 
> Cost.  With a classifier attached, dispatch takes two or three more
> lock-free queue operations per request.  On the same box that's about half a
> microsecond when the enqueue and the dequeue run on different cores,
> fio 4k O_DIRECT randread over loopback, 5 x 20s:
> 
>                                    A        B
>     nconnect=1, 16 jobs        152.8k   152.2k   -0.4%
>     nconnect=8, 16 jobs        483.8k   475.5k   -1.7%
>     nconnect=8, 64 jobs        584.2k   581.0k   -0.5%
> 
> The serial walk and 4 jobs on one connection don't move.  With no
> classifier attached the pool is on its old flat queue (patch 6).  I only
> have my VM for that path so far (KASAN, 10 vCPUs, loopback, the same
> fio), and there the no-program kernel measures the same as v7.2 within
> the run-to-run spread: 16 jobs 94.9k vs 95.5k, 64 jobs 138.6k vs 139.0k.
> So the static branch adds nothing I can see, but that's a VM number.
> 
> Not tested: RDMA.  The hook sits in the accept path TCP and RDMA share,
> but I could not bring a soft-RoCE listener up on the VM.
> 
> Daire - the same nine patches on v7.2 (plus the wake fix) are at
> https://github.com/bcodding/linux branch nfsd-clientq-v2-7.2 if you
> want to run it; the loader builds from tools/net/sunrpc/svc-classify in
> that tree.
> 
> The series is on nfsd-testing at 56589cdb5881, which already has the
> svc_clean_up_xprts() wake fix from [3].
> 
>   patch 1  track clients by class, no change to dispatch
>   patch 2  dispatch round-robin across clients
>   patch 3  the struct_ops hook
>   patch 4  class in the svc_xprt_dequeue tracepoint
>   patch 5  control events ahead of data within a client
>   patch 6  flat queue while no classifier is attached
>   patch 7  selftests: the prefix program and the attach tests
>   patch 8  tools/net/sunrpc/svc-classify: the prefix classifier and its
>            loader (load/replace/unload, add/del/list by prefix, status)
>   patch 9  Documentation: the model, the class word, installing a
>            classifier with svc-classify, six class maps and the share
>            each gives, how it works underneath
> 
> [1] https://lore.kernel.org/linux-nfs/cover.1780498019.git.bcodding@hammerspace.com
> [2] https://lore.kernel.org/linux-nfs/cover.1782314746.git.bcodding@hammerspace.com
> [3] https://lore.kernel.org/linux-nfs/cover.1790953694.git.bcodding@hammerspace.com
> 
> Benjamin Coddington (9):
>   SUNRPC: track service clients by class
>   SUNRPC: dispatch ready transports round-robin across clients
>   SUNRPC: add a BPF struct_ops hook to classify accepted transports
>   SUNRPC: report the transport's class in the svc_xprt_dequeue
>     tracepoint
>   SUNRPC: dispatch control events ahead of data within a client
>   SUNRPC: keep the single transport queue while no classifier is
>     attached
>   selftests/bpf: add svc_classifier tests
>   tools/net/sunrpc: add svc-classify, a prefix classifier and its loader
>   Documentation: describe RPC server transport classes and the BPF
>     classifier
> 
>  Documentation/filesystems/nfs/index.rst       |   1 +
>  .../filesystems/nfs/rpc-server-clients.rst    | 394 ++++++++++++++++++
>  include/linux/sunrpc/svc.h                    |  72 +++-
>  include/linux/sunrpc/svc_xprt.h               |  30 ++
>  include/trace/events/sunrpc.h                 |  11 +-
>  net/sunrpc/Kconfig                            |  12 +
>  net/sunrpc/Makefile                           |   1 +
>  net/sunrpc/netns.h                            |   4 +
>  net/sunrpc/sunrpc_syms.c                      |   4 +
>  net/sunrpc/svc.c                              |  36 ++
>  net/sunrpc/svc_classify.c                     | 210 ++++++++++
>  net/sunrpc/svc_xprt.c                         | 281 ++++++++++++-
>  tools/net/sunrpc/svc-classify/.gitignore      |   1 +
>  tools/net/sunrpc/svc-classify/Makefile        |  65 +++
>  tools/net/sunrpc/svc-classify/svc-classify.c  | 357 ++++++++++++++++
>  .../sunrpc/svc-classify/svc_classify.bpf.c    |  83 ++++
>  tools/testing/selftests/bpf/config            |   2 +
>  .../selftests/bpf/prog_tests/svc_classifier.c | 104 +++++
>  .../selftests/bpf/progs/bpf_svc_classifier.c  |  80 ++++
>  19 files changed, 1724 insertions(+), 24 deletions(-)
>  create mode 100644 Documentation/filesystems/nfs/rpc-server-clients.rst
>  create mode 100644 net/sunrpc/svc_classify.c
>  create mode 100644 tools/net/sunrpc/svc-classify/.gitignore
>  create mode 100644 tools/net/sunrpc/svc-classify/Makefile
>  create mode 100644 tools/net/sunrpc/svc-classify/svc-classify.c
>  create mode 100644 tools/net/sunrpc/svc-classify/svc_classify.bpf.c
>  create mode 100644 tools/testing/selftests/bpf/prog_tests/svc_classifier.c
>  create mode 100644 tools/testing/selftests/bpf/progs/bpf_svc_classifier.c

My Comments:

Conceptually, I like this idea, and I personally like that I now have a
clear (to me) example of how to plug eBPF programs into kernel code, so
thanks for that!

I think the userland bits should be made part of nfs-utils, and the
LLM-slop docs need to be rewritten for humans.

-- 
Jeff Layton <jlayton@kernel.org>

      parent reply	other threads:[~2026-10-08 18:48 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 1/9] SUNRPC: track service clients by class Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 2/9] SUNRPC: dispatch ready transports round-robin across clients Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 3/9] SUNRPC: add a BPF struct_ops hook to classify accepted transports Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 4/9] SUNRPC: report the transport's class in the svc_xprt_dequeue tracepoint Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 5/9] SUNRPC: dispatch control events ahead of data within a client Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 6/9] SUNRPC: keep the single transport queue while no classifier is attached Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 7/9] selftests/bpf: add svc_classifier tests Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader Benjamin Coddington
2026-10-08 18:36   ` Jeff Layton
2026-10-08 19:38     ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier Benjamin Coddington
2026-10-08 18:42   ` Jeff Layton
2026-10-08 19:40     ` Benjamin Coddington
2026-10-08 15:19 ` [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Chuck Lever
2026-10-08 15:28   ` Benjamin Coddington
2026-10-08 18:48 ` Jeff Layton [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=a0f7d257143bee30db0cda0ca186b06443710bf7.camel@kernel.org \
    --to=jlayton@kernel.org \
    --cc=ben.coddington@hammerspace.com \
    --cc=bpf@vger.kernel.org \
    --cc=cel@kernel.org \
    --cc=daire@dneg.com \
    --cc=linux-nfs@vger.kernel.org \
    --cc=martin.lau@linux.dev \
    --cc=neil@brown.name \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox