* [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF
@ 2026-10-07 19:59 Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 1/9] SUNRPC: track service clients by class Benjamin Coddington
` (10 more replies)
0 siblings, 11 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
This is v2 of [3], with the change I floated in that thread: the kernel
no longer decides which transports belong together. By default nothing
changes. An administrator who wants peers grouped loads a small BPF
program that returns a class for each accepted connection, and the pool
takes turns across classes.
What's different from v1
------------------------
The table I described on [3] became a hook. Instead of a kernel table
of address prefixes, a netlink interface for it and an nfs-utils side,
sunrpc registers a struct_ops with one callback: given the accepted
svc_xprt, return a class. 0 means the anonymous client (shared with
everything unclassified, today's order among them); any other number is
a class every transport returning it shares; a flag bit on top means
"and one client per peer address within that class". So "all movers are
one class, everybody else per host" is two entries in a prefix map, and
the posted v1 behaviour is one entry.
The hook runs once per accepted connection, in process context, after
TCP and RDMA have both set the peer address. Nothing runs per request.
No program loaded means every transport is on the
anonymous client, and with patch 6 the pool uses its old flat queue
until a classifier attaches, so the dispatch path is today's code for
anyone who hasn't asked for anything. That's the answer to Chuck's cost
question from [3]: measured on bare metal the two-level queue cost about
half a microsecond a dispatch (1.7% at 484k IOPS on one client's eight
connections); with no program it now costs a static branch on enqueue
and an empty-queue test on dequeue.
Attaching is a struct_ops link, bound to the attaching task's network
namespace, one per namespace. A second attach gets -EBUSY, a link
update replaces the program, closing the link removes it, and a
namespace that exits leaves its link with nothing behind. Transports
keep the class they got at accept; a new program only affects new
connections. The reference program and the attach tests are in the BPF
selftests (patch 7). svc-classify (patch 8) attaches the prefix
classifier and manages its map by address prefix, so an administrator
never touches bpftool. The page in patch 9,
Documentation/filesystems/nfs/rpc-server-clients.rst, has the class
word, the install procedure, and six class maps with the share each one
gives its clients.
Still not in here: moving a transport between clients, so a v4.1
clientid key, which Chuck asked about on [3], is not in this version.
A multi-homed host or an IPv6 host with a prefix's worth of addresses
is handled by the prefix map instead.
Also from the [3] review: control events (accept, close, handshake) go
ahead of data within a client (patch 5), and the dequeue tracepoint
prints the class (patch 4).
Numbers
-------
The dispatch numbers are the v1 series' on a two-socket box (2 x Xeon
6542Y, 96 threads, no KASAN), from [3]. The dispatch code is the same
here: v2 puts the classifier in front of it, and the flat queue behind
a static branch when nothing is attached. Loopback, NFSv3, 16 threads,
10 ms injected service time with 50% jitter. A is v7.2, B the series.
v4.1 is within 2% in every cell.
One interactive client (bursts of 32) against one host with K
backlogged connections, burst completion p50 in ms, floor 37:
K 4 8 16 32
A 70 195 351 671
B 56 62 61 63
N (K=16) 1 8 32 64 96 128
A 21 98 350 690 1035 1365
B 11 31 62 103 144 179
M hosts x 4 1 2 4 8
A 70 192 351 670
B 57 81 121 192
S: 2x8 vs K, share 4 8 16 32
A 41.0 20.0 11.1 5.9
B 50.3 50.0 50.0 50.0
A real NFSv3 client walking 500 files against a 16-connection
aggressor: 37.8 s on A, 2.5 s on B, 0.12 s alone on both.
Classes. The v1 kernel keyed on source address, so on [3] I emulated
"all movers are one class" by giving six movers one address. With this
series that's one map entry (the movers' prefix -> 1, everything else
per address), and the six-address column is what "everything per
address" gives. Six movers at 4 or 8 connections each against
customers, customer share of dispatches:
A per address movers one class
one customer, 1 conn x 16 4.0 / 2.0 14.3 / 14.3 49.9 / 49.9
one customer, 4 x 4 14.3 / 7.7 14.3 / 14.3 49.8 / 49.7
four customers, 4 x 4 each 40 / 25 40 / 40 80 / 80
Burst of 32 against M busy hosts, p50 ms: hosts on their own addresses
57 81 121 192, as one class 56 62 62 62. A: 71 193 348 673 either way.
Cost. With a classifier attached, dispatch takes two or three more
lock-free queue operations per request. On the same box that's about half a
microsecond when the enqueue and the dequeue run on different cores,
fio 4k O_DIRECT randread over loopback, 5 x 20s:
A B
nconnect=1, 16 jobs 152.8k 152.2k -0.4%
nconnect=8, 16 jobs 483.8k 475.5k -1.7%
nconnect=8, 64 jobs 584.2k 581.0k -0.5%
The serial walk and 4 jobs on one connection don't move. With no
classifier attached the pool is on its old flat queue (patch 6). I only
have my VM for that path so far (KASAN, 10 vCPUs, loopback, the same
fio), and there the no-program kernel measures the same as v7.2 within
the run-to-run spread: 16 jobs 94.9k vs 95.5k, 64 jobs 138.6k vs 139.0k.
So the static branch adds nothing I can see, but that's a VM number.
Not tested: RDMA. The hook sits in the accept path TCP and RDMA share,
but I could not bring a soft-RoCE listener up on the VM.
Daire - the same nine patches on v7.2 (plus the wake fix) are at
https://github.com/bcodding/linux branch nfsd-clientq-v2-7.2 if you
want to run it; the loader builds from tools/net/sunrpc/svc-classify in
that tree.
The series is on nfsd-testing at 56589cdb5881, which already has the
svc_clean_up_xprts() wake fix from [3].
patch 1 track clients by class, no change to dispatch
patch 2 dispatch round-robin across clients
patch 3 the struct_ops hook
patch 4 class in the svc_xprt_dequeue tracepoint
patch 5 control events ahead of data within a client
patch 6 flat queue while no classifier is attached
patch 7 selftests: the prefix program and the attach tests
patch 8 tools/net/sunrpc/svc-classify: the prefix classifier and its
loader (load/replace/unload, add/del/list by prefix, status)
patch 9 Documentation: the model, the class word, installing a
classifier with svc-classify, six class maps and the share
each gives, how it works underneath
[1] https://lore.kernel.org/linux-nfs/cover.1780498019.git.bcodding@hammerspace.com
[2] https://lore.kernel.org/linux-nfs/cover.1782314746.git.bcodding@hammerspace.com
[3] https://lore.kernel.org/linux-nfs/cover.1790953694.git.bcodding@hammerspace.com
Benjamin Coddington (9):
SUNRPC: track service clients by class
SUNRPC: dispatch ready transports round-robin across clients
SUNRPC: add a BPF struct_ops hook to classify accepted transports
SUNRPC: report the transport's class in the svc_xprt_dequeue
tracepoint
SUNRPC: dispatch control events ahead of data within a client
SUNRPC: keep the single transport queue while no classifier is
attached
selftests/bpf: add svc_classifier tests
tools/net/sunrpc: add svc-classify, a prefix classifier and its loader
Documentation: describe RPC server transport classes and the BPF
classifier
Documentation/filesystems/nfs/index.rst | 1 +
.../filesystems/nfs/rpc-server-clients.rst | 394 ++++++++++++++++++
include/linux/sunrpc/svc.h | 72 +++-
include/linux/sunrpc/svc_xprt.h | 30 ++
include/trace/events/sunrpc.h | 11 +-
net/sunrpc/Kconfig | 12 +
net/sunrpc/Makefile | 1 +
net/sunrpc/netns.h | 4 +
net/sunrpc/sunrpc_syms.c | 4 +
net/sunrpc/svc.c | 36 ++
net/sunrpc/svc_classify.c | 210 ++++++++++
net/sunrpc/svc_xprt.c | 281 ++++++++++++-
tools/net/sunrpc/svc-classify/.gitignore | 1 +
tools/net/sunrpc/svc-classify/Makefile | 65 +++
tools/net/sunrpc/svc-classify/svc-classify.c | 357 ++++++++++++++++
.../sunrpc/svc-classify/svc_classify.bpf.c | 83 ++++
tools/testing/selftests/bpf/config | 2 +
.../selftests/bpf/prog_tests/svc_classifier.c | 104 +++++
.../selftests/bpf/progs/bpf_svc_classifier.c | 80 ++++
19 files changed, 1724 insertions(+), 24 deletions(-)
create mode 100644 Documentation/filesystems/nfs/rpc-server-clients.rst
create mode 100644 net/sunrpc/svc_classify.c
create mode 100644 tools/net/sunrpc/svc-classify/.gitignore
create mode 100644 tools/net/sunrpc/svc-classify/Makefile
create mode 100644 tools/net/sunrpc/svc-classify/svc-classify.c
create mode 100644 tools/net/sunrpc/svc-classify/svc_classify.bpf.c
create mode 100644 tools/testing/selftests/bpf/prog_tests/svc_classifier.c
create mode 100644 tools/testing/selftests/bpf/progs/bpf_svc_classifier.c
--
2.53.0
^ permalink raw reply [flat|nested] 17+ messages in thread
* [PATCH RFC v2 1/9] SUNRPC: track service clients by class
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 2/9] SUNRPC: dispatch ready transports round-robin across clients Benjamin Coddington
` (9 subsequent siblings)
10 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
A pool dispatches its ready transports in FIFO order, one RPC per turn,
so a peer's share of the service grows with the number of connections
it holds: a client with K connections is served K times as often as a
client with one. Dispatching per client instead needs the service to
know which transports belong together, and which transports belong
together is a question for the administrator, not the kernel: one site
wants every host served alike, another wants a set of bulk movers to
share one turn among them.
Add struct svc_client, one per class of transports in a network
namespace of a service, or one per peer address within a class, found
through a small hash table under a spinlock. The class of an accepted
transport comes from svc_classify(), which for now returns
SVC_CLASS_NONE, so every transport stays on the service's anonymous
client; a later patch lets an administrator's BPF program supply it.
The table is touched only when a connection is accepted and when a
client's last transport is freed; dispatch will follow the transport's
xpt_client pointer, which is set before the transport can be enqueued
(it is still XPT_BUSY from svc_xprt_init()) and is kept alive by the
transport's reference. Each client carries a per-pool lwq of its ready
transports and a node for the pool's queue of clients, unused as yet.
Listeners, UDP sockets, transports of no class, and accepted transports
for which no client can be allocated are bound to the anonymous client.
Dispatch is unchanged by this patch.
Suggested-by: NeilBrown <neil@brown.name>
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
include/linux/sunrpc/svc.h | 43 +++++++++++++
include/linux/sunrpc/svc_xprt.h | 8 +++
net/sunrpc/svc.c | 34 ++++++++++
net/sunrpc/svc_xprt.c | 108 ++++++++++++++++++++++++++++++++
4 files changed, 193 insertions(+)
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index dfadea50e0c6..e45f4fdea20b 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -58,6 +58,44 @@ enum {
SP_TASK_STARTING, /* Task has started but not added to idle yet */
};
+/*
+ * A group of transports that share turns at the service: those of one
+ * class in one network namespace, or of one peer address within a class.
+ * Ready transports are queued per client and per pool, and a pool
+ * dispatches round-robin across its queued clients, so a client's share
+ * of the service does not grow with its connection count.
+ */
+struct svc_client_pool {
+ struct lwq cp_xprts; /* ready transports */
+ struct lwq_node cp_ready; /* link in svc_pool.sp_clients */
+ unsigned long cp_flags;
+ struct svc_client *cp_client;
+};
+
+enum {
+ SVC_CP_QUEUED, /* on sp_clients, or held by the thread that took it */
+};
+
+struct svc_client {
+ struct hlist_node cl_hash;
+ refcount_t cl_ref;
+ struct net *cl_net;
+ u32 cl_class;
+ struct sockaddr_storage cl_addr; /* with SVC_CLASS_PER_ADDR; port ignored */
+ struct svc_client_pool cl_pool[]; /* one per pool */
+};
+
+#define SVC_CLIENT_HASH_BITS 8
+
+/*
+ * Classes returned by svc_classify(). SVC_CLASS_NONE leaves the transport
+ * on the service's anonymous client; any other value names a client shared
+ * by every transport of that class, or, with SVC_CLASS_PER_ADDR set, one
+ * client per peer address within the class.
+ */
+#define SVC_CLASS_NONE 0
+#define SVC_CLASS_PER_ADDR BIT_U32(31)
+
struct svc_rqst;
@@ -96,6 +134,10 @@ struct svc_serv {
struct svc_pool * sv_pools; /* array of thread pools */
int (*sv_threadfn)(void *data);
+ spinlock_t sv_client_lock; /* protects sv_client_hash */
+ struct hlist_head sv_client_hash[1 << SVC_CLIENT_HASH_BITS];
+ struct svc_client *sv_anon_client; /* transports without a peer */
+
#if defined(CONFIG_SUNRPC_BACKCHANNEL)
struct lwq sv_cb_list; /* queue for callback requests
* that arrive over the same
@@ -503,6 +545,7 @@ void svc_reserve(struct svc_rqst *rqstp, int space);
void svc_pool_wake_idle_thread(struct svc_pool *pool);
struct svc_pool *svc_pool_for_cpu(struct svc_serv *serv);
unsigned int svc_serv_nrpools(const struct svc_serv *serv);
+struct svc_client *svc_client_alloc(unsigned int nrpools, gfp_t gfp);
char * svc_print_addr(struct svc_rqst *, char *, size_t);
const char * svc_proc_name(const struct svc_rqst *rqstp);
int svc_encode_result_payload(struct svc_rqst *rqstp,
diff --git a/include/linux/sunrpc/svc_xprt.h b/include/linux/sunrpc/svc_xprt.h
index 7176c42f19d7..d4f6dd5c8db3 100644
--- a/include/linux/sunrpc/svc_xprt.h
+++ b/include/linux/sunrpc/svc_xprt.h
@@ -63,6 +63,7 @@ struct svc_xprt {
unsigned long xpt_flags;
struct svc_serv *xpt_server; /* service for transport */
+ struct svc_client *xpt_client; /* peer this transport belongs to */
atomic_t xpt_reserved; /* outq space rsvd, UDP only */
atomic_t xpt_nr_rqsts; /* Number of requests */
struct mutex xpt_mutex; /* to serialize sending data */
@@ -177,6 +178,13 @@ int svc_xprt_create(struct svc_serv *serv, const char *xprt_name,
void svc_xprt_destroy_all(struct svc_serv *serv, struct net *net,
bool unregister);
void svc_xprt_received(struct svc_xprt *xprt);
+
+/* The class of an accepted transport; SVC_CLASS_* in svc.h. */
+static inline u32 svc_classify(const struct svc_xprt *xprt)
+{
+ return SVC_CLASS_NONE;
+}
+
void svc_xprt_enqueue(struct svc_xprt *xprt);
void svc_xprt_put(struct svc_xprt *xprt);
void svc_xprt_copy_addrs(struct svc_rqst *rqstp, struct svc_xprt *xprt);
diff --git a/net/sunrpc/svc.c b/net/sunrpc/svc.c
index 0a2c040ad096..cb2b7caaa627 100644
--- a/net/sunrpc/svc.c
+++ b/net/sunrpc/svc.c
@@ -387,6 +387,30 @@ static void svc_pool_destroy_counters(struct svc_pool *pool)
percpu_counter_destroy(&pool->sp_threads_woken);
}
+/**
+ * svc_client_alloc - allocate a service client
+ * @nrpools: number of thread pools in the service
+ * @gfp: allocation flags
+ *
+ * Return: the client, holding one reference, or %NULL.
+ */
+struct svc_client *svc_client_alloc(unsigned int nrpools, gfp_t gfp)
+{
+ struct svc_client *cl;
+ unsigned int i;
+
+ cl = kzalloc(struct_size(cl, cl_pool, nrpools), gfp);
+ if (!cl)
+ return NULL;
+ refcount_set(&cl->cl_ref, 1);
+ INIT_HLIST_NODE(&cl->cl_hash);
+ for (i = 0; i < nrpools; i++) {
+ lwq_init(&cl->cl_pool[i].cp_xprts);
+ cl->cl_pool[i].cp_client = cl;
+ }
+ return cl;
+}
+
/*
* Create an RPC service
*/
@@ -432,8 +456,16 @@ __svc_create(struct svc_program *prog, int nprogs, struct svc_stat *stats,
__svc_init_bc(serv);
+ spin_lock_init(&serv->sv_client_lock);
+ serv->sv_anon_client = svc_client_alloc(npools, GFP_KERNEL);
+ if (!serv->sv_anon_client) {
+ kfree(serv);
+ return NULL;
+ }
+
serv->sv_pools = kzalloc_objs(struct svc_pool, npools);
if (!serv->sv_pools) {
+ kfree(serv->sv_anon_client);
kfree(serv);
return NULL;
}
@@ -459,6 +491,7 @@ __svc_create(struct svc_program *prog, int nprogs, struct svc_stat *stats,
while (i--)
svc_pool_destroy_counters(&serv->sv_pools[i]);
kfree(serv->sv_pools);
+ kfree(serv->sv_anon_client);
kfree(serv);
return NULL;
}
@@ -545,6 +578,7 @@ svc_destroy(struct svc_serv **servp)
if (serv->sv_is_pooled)
svc_pool_map_put();
+ kfree(serv->sv_anon_client);
kfree(serv->sv_pools);
kfree(serv);
}
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 9858dfcb846a..c7da55335506 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -9,8 +9,11 @@
#include <linux/sched/mm.h>
#include <linux/errno.h>
#include <linux/freezer.h>
+#include <linux/hash.h>
#include <linux/slab.h>
#include <net/sock.h>
+#include <net/ip.h>
+#include <net/ipv6.h>
#include <linux/sunrpc/addr.h>
#include <linux/sunrpc/stats.h>
#include <linux/sunrpc/svc_xprt.h>
@@ -167,6 +170,107 @@ void svc_xprt_deferred_close(struct svc_xprt *xprt)
}
EXPORT_SYMBOL_GPL(svc_xprt_deferred_close);
+static unsigned int svc_client_hash(const struct net *net, u32 class,
+ const struct sockaddr *sa)
+{
+ u32 h = hash_ptr(net, 32) ^ class;
+
+ if (class & SVC_CLASS_PER_ADDR) {
+ switch (sa->sa_family) {
+ case AF_INET:
+ h = __ipv4_addr_hash(((const struct sockaddr_in *)sa)->sin_addr.s_addr, h);
+ break;
+ case AF_INET6:
+ h = __ipv6_addr_jhash(&((const struct sockaddr_in6 *)sa)->sin6_addr, h);
+ break;
+ }
+ }
+ return hash_32(h, SVC_CLIENT_HASH_BITS);
+}
+
+static struct svc_client *svc_client_get(struct svc_client *cl)
+{
+ refcount_inc(&cl->cl_ref);
+ return cl;
+}
+
+static void svc_client_put(struct svc_serv *serv, struct svc_client *cl)
+{
+ if (!refcount_dec_and_test(&cl->cl_ref))
+ return;
+ spin_lock_bh(&serv->sv_client_lock);
+ hlist_del(&cl->cl_hash);
+ spin_unlock_bh(&serv->sv_client_lock);
+ kfree(cl);
+}
+
+/* Caller holds sv_client_lock. */
+static struct svc_client *svc_client_find(struct hlist_head *head,
+ const struct net *net, u32 class,
+ const struct sockaddr *sa)
+{
+ struct svc_client *cl;
+
+ hlist_for_each_entry(cl, head, cl_hash)
+ if (cl->cl_net == net && cl->cl_class == class &&
+ (!(class & SVC_CLASS_PER_ADDR) ||
+ rpc_cmp_addr((struct sockaddr *)&cl->cl_addr, sa)) &&
+ refcount_inc_not_zero(&cl->cl_ref))
+ return cl;
+ return NULL;
+}
+
+/*
+ * Bind @xprt to the client for its class. Runs once per accepted
+ * transport, in process context, while the transport is still XPT_BUSY
+ * from svc_xprt_init(), so no enqueue can see the pointer change. A
+ * transport of no class, one keyed by an address that is not IP, or one
+ * for which no client can be allocated, stays on the service's anonymous
+ * client.
+ */
+static void svc_client_bind(struct svc_serv *serv, struct svc_xprt *xprt)
+{
+ struct sockaddr *sa = (struct sockaddr *)&xprt->xpt_remote;
+ struct svc_client *cl, *new = NULL;
+ struct net *net = xprt->xpt_net;
+ struct hlist_head *head;
+ u32 class;
+
+ class = svc_classify(xprt);
+ if (class == SVC_CLASS_NONE)
+ return;
+ if ((class & SVC_CLASS_PER_ADDR) &&
+ sa->sa_family != AF_INET && sa->sa_family != AF_INET6)
+ return;
+ head = &serv->sv_client_hash[svc_client_hash(net, class, sa)];
+
+ spin_lock_bh(&serv->sv_client_lock);
+ cl = svc_client_find(head, net, class, sa);
+ spin_unlock_bh(&serv->sv_client_lock);
+ if (!cl) {
+ new = svc_client_alloc(svc_serv_nrpools(serv), GFP_KERNEL);
+ if (!new)
+ return;
+ new->cl_net = net;
+ new->cl_class = class;
+ if (class & SVC_CLASS_PER_ADDR)
+ memcpy(&new->cl_addr, sa, xprt->xpt_remotelen);
+
+ spin_lock_bh(&serv->sv_client_lock);
+ cl = svc_client_find(head, net, class, sa);
+ if (!cl) {
+ hlist_add_head(&new->cl_hash, head);
+ cl = new;
+ new = NULL;
+ }
+ spin_unlock_bh(&serv->sv_client_lock);
+ kfree(new);
+ }
+
+ svc_client_put(serv, xprt->xpt_client);
+ xprt->xpt_client = cl;
+}
+
static void svc_xprt_free(struct kref *kref)
{
struct svc_xprt *xprt =
@@ -176,6 +280,7 @@ static void svc_xprt_free(struct kref *kref)
trace_svc_xprt_free(xprt);
if (test_bit(XPT_CACHE_AUTH, &xprt->xpt_flags))
svcauth_unix_info_release(xprt);
+ svc_client_put(xprt->xpt_server, xprt->xpt_client);
put_cred(xprt->xpt_cred);
put_net_track(xprt->xpt_net, &xprt->ns_tracker);
/* See comment on corresponding get in xs_setup_bc_tcp(): */
@@ -213,6 +318,7 @@ void svc_xprt_init(struct net *net, struct svc_xprt_class *xcl,
xprt->xpt_ops = xcl->xcl_ops;
kref_init(&xprt->xpt_ref);
xprt->xpt_server = serv;
+ xprt->xpt_client = svc_client_get(serv->sv_anon_client);
INIT_LIST_HEAD(&xprt->xpt_list);
INIT_LIST_HEAD(&xprt->xpt_deferred);
INIT_LIST_HEAD(&xprt->xpt_users);
@@ -836,6 +942,8 @@ static bool svc_thread_wait_for_work(struct svc_rqst *rqstp, long timeo)
static void svc_add_new_temp_xprt(struct svc_serv *serv, struct svc_xprt *newxpt)
{
+ svc_client_bind(serv, newxpt);
+
spin_lock_bh(&serv->sv_lock);
set_bit(XPT_TEMP, &newxpt->xpt_flags);
list_add(&newxpt->xpt_list, &serv->sv_tempsocks);
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH RFC v2 2/9] SUNRPC: dispatch ready transports round-robin across clients
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 1/9] SUNRPC: track service clients by class Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 3/9] SUNRPC: add a BPF struct_ops hook to classify accepted transports Benjamin Coddington
` (8 subsequent siblings)
10 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
A pool keeps its ready transports in one FIFO and serves one RPC per
turn, so a peer is served in proportion to how many of its connections
are backlogged. A client with K connections, or a deep NFSv4.1 slot
table spread over nconnect transports, takes K turns for every one a
single-connection client gets, and the single-connection client's
latency grows with everyone else's connection count.
Queue ready transports on their client instead, and have the pool
dispatch round-robin across clients: take the client at the head of
the pool's client queue, dispatch one of its transports, and put the
client back at the tail if it has more. Every client with work queued
gets one dispatch per round however many connections it holds, and a
client's wait is bounded by the number of busy clients rather than by
the number of busy connections. Within a client, transports are still
served in FIFO order.
Enqueue stays lockless: the transport goes on the client's per-pool
lwq, and the client goes on the pool's lwq the first time one of its
transports becomes ready, guarded by a per-pool SVC_CP_QUEUED bit that
the dequeuing thread clears when it finds the client empty. The bit is
cleared before the client's queue is re-tested, so an enqueue that saw
it set and did not queue the client is caught by the re-test. Dequeue
takes one more lwq spinlock than before. Control events (XPT_CONN,
XPT_CLOSE, XPT_HANDSHAKE) take turns like data: a close queued behind
k of its own client's transports is dispatched k rounds later, where
it used to wait behind every queued transport in the pool.
Network namespace teardown walks the pool's clients: those of the
namespace being destroyed are drained and their transports deleted; the
anonymous client, which may hold listeners from several namespaces, is
filtered per transport; the rest are re-queued and a thread is woken.
While nothing returns a class from svc_classify(), every transport is
on the anonymous client, a pool's client queue holds one entry, and
dispatch is in today's FIFO order.
Suggested-by: NeilBrown <neil@brown.name>
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
include/linux/sunrpc/svc.h | 2 +-
net/sunrpc/svc.c | 2 +-
net/sunrpc/svc_xprt.c | 97 +++++++++++++++++++++++++++++---------
3 files changed, 77 insertions(+), 24 deletions(-)
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index e45f4fdea20b..a27dc2942167 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -38,7 +38,7 @@ struct svc_pool {
unsigned int sp_nrthreads; /* # of threads currently running in pool */
unsigned int sp_nrthrmin; /* Min number of threads to run per pool */
unsigned int sp_nrthrmax; /* Max requested number of threads in pool */
- struct lwq sp_xprts; /* pending transports */
+ struct lwq sp_clients; /* clients with ready transports */
struct list_head sp_all_threads; /* all server threads */
struct llist_head sp_idle_threads; /* idle server threads */
diff --git a/net/sunrpc/svc.c b/net/sunrpc/svc.c
index cb2b7caaa627..4adf113d067b 100644
--- a/net/sunrpc/svc.c
+++ b/net/sunrpc/svc.c
@@ -477,7 +477,7 @@ __svc_create(struct svc_program *prog, int nprogs, struct svc_stat *stats,
i, serv->sv_name);
pool->sp_id = i;
- lwq_init(&pool->sp_xprts);
+ lwq_init(&pool->sp_clients);
INIT_LIST_HEAD(&pool->sp_all_threads);
init_llist_head(&pool->sp_idle_threads);
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index c7da55335506..735f4d92bf5d 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -627,6 +627,7 @@ static bool svc_xprt_ready(struct svc_xprt *xprt)
*/
void svc_xprt_enqueue(struct svc_xprt *xprt)
{
+ struct svc_client_pool *cp;
struct svc_pool *pool;
if (!svc_xprt_ready(xprt))
@@ -641,25 +642,56 @@ void svc_xprt_enqueue(struct svc_xprt *xprt)
return;
pool = svc_pool_for_cpu(xprt->xpt_server);
+ cp = &xprt->xpt_client->cl_pool[pool->sp_id];
percpu_counter_inc(&pool->sp_sockets_queued);
xprt->xpt_qtime = ktime_get();
- lwq_enqueue(&xprt->xpt_ready, &pool->sp_xprts);
+ lwq_enqueue(&xprt->xpt_ready, &cp->cp_xprts);
+ if (!test_and_set_bit(SVC_CP_QUEUED, &cp->cp_flags))
+ lwq_enqueue(&cp->cp_ready, &pool->sp_clients);
svc_pool_wake_idle_thread(pool);
}
EXPORT_SYMBOL_GPL(svc_xprt_enqueue);
/*
- * Dequeue the first transport, if there is one.
+ * Put @cp back on the pool's client queue if it still has ready
+ * transports. Otherwise clear SVC_CP_QUEUED and look once more: an
+ * enqueue that found the bit set has left the client for us to queue.
+ */
+static void svc_client_pool_requeue(struct svc_pool *pool,
+ struct svc_client_pool *cp)
+{
+ if (!lwq_empty(&cp->cp_xprts)) {
+ lwq_enqueue(&cp->cp_ready, &pool->sp_clients);
+ return;
+ }
+ clear_bit(SVC_CP_QUEUED, &cp->cp_flags);
+ /* the clear is visible before the re-test; see svc_xprt_enqueue() */
+ smp_mb__after_atomic();
+ if (!lwq_empty(&cp->cp_xprts) &&
+ !test_and_set_bit(SVC_CP_QUEUED, &cp->cp_flags))
+ lwq_enqueue(&cp->cp_ready, &pool->sp_clients);
+}
+
+/*
+ * Dequeue the next transport: one from the client at the head of the
+ * pool's client queue, which then goes to the tail if it has more.
*/
static struct svc_xprt *svc_xprt_dequeue(struct svc_pool *pool)
{
- struct svc_xprt *xprt = NULL;
+ struct svc_client_pool *cp;
+ struct svc_xprt *xprt;
- xprt = lwq_dequeue(&pool->sp_xprts, struct svc_xprt, xpt_ready);
- if (xprt)
- svc_xprt_get(xprt);
+ do {
+ cp = lwq_dequeue(&pool->sp_clients, struct svc_client_pool,
+ cp_ready);
+ if (!cp)
+ return NULL;
+ xprt = lwq_dequeue(&cp->cp_xprts, struct svc_xprt, xpt_ready);
+ svc_client_pool_requeue(pool, cp);
+ } while (!xprt);
+ svc_xprt_get(xprt);
return xprt;
}
@@ -886,7 +918,7 @@ svc_thread_should_sleep(struct svc_rqst *rqstp)
return false;
/* was a socket queued? */
- if (!lwq_empty(&pool->sp_xprts))
+ if (!lwq_empty(&pool->sp_clients))
return false;
/* are we shutting down? */
@@ -1338,28 +1370,49 @@ static int svc_close_list(struct svc_serv *serv, struct list_head *xprt_list, st
return ret;
}
-static void svc_clean_up_xprts(struct svc_serv *serv, struct net *net)
+/*
+ * Delete the ready transports of @net that are queued on @cp; the rest go
+ * back in their order.
+ */
+static void svc_client_pool_clean(struct svc_client_pool *cp, struct net *net)
{
+ struct llist_node *q, **t1, *t2;
struct svc_xprt *xprt;
+
+ q = lwq_dequeue_all(&cp->cp_xprts);
+ lwq_for_each_safe(xprt, t1, t2, &q, xpt_ready) {
+ if (xprt->xpt_net == net) {
+ set_bit(XPT_CLOSE, &xprt->xpt_flags);
+ svc_delete_xprt(xprt);
+ xprt = NULL;
+ }
+ }
+ if (q)
+ lwq_enqueue_batch(q, &cp->cp_xprts);
+}
+
+static void svc_clean_up_xprts(struct svc_serv *serv, struct net *net)
+{
int i;
for (i = 0; i < svc_serv_nrpools(serv); i++) {
struct svc_pool *pool = &serv->sv_pools[i];
- struct llist_node *q, **t1, *t2;
-
- q = lwq_dequeue_all(&pool->sp_xprts);
- lwq_for_each_safe(xprt, t1, t2, &q, xpt_ready) {
- if (xprt->xpt_net == net) {
- set_bit(XPT_CLOSE, &xprt->xpt_flags);
- svc_delete_xprt(xprt);
- xprt = NULL;
- }
- }
-
- if (q) {
- lwq_enqueue_batch(q, &pool->sp_xprts);
- svc_pool_wake_idle_thread(pool);
+ struct svc_client_pool *cp;
+ struct svc_client *cl;
+ struct llist_node *q;
+
+ q = lwq_dequeue_all(&pool->sp_clients);
+ while (q) {
+ cp = container_of(q, struct svc_client_pool, cp_ready.node);
+ q = q->next;
+ /* Deleting the last transport frees the client. */
+ cl = svc_client_get(cp->cp_client);
+ if (cl->cl_net == net || cl == serv->sv_anon_client)
+ svc_client_pool_clean(cp, net);
+ svc_client_pool_requeue(pool, cp);
+ svc_client_put(serv, cl);
}
+ svc_pool_wake_idle_thread(pool);
}
}
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH RFC v2 3/9] SUNRPC: add a BPF struct_ops hook to classify accepted transports
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 1/9] SUNRPC: track service clients by class Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 2/9] SUNRPC: dispatch ready transports round-robin across clients Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 4/9] SUNRPC: report the transport's class in the svc_xprt_dequeue tracepoint Benjamin Coddington
` (7 subsequent siblings)
10 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
Which transports should share turns at a service is site policy: one
server wants every host served alike, another wants a set of bulk
movers to count as one peer while its interactive clients are served
per host, and a server behind a NAT or a re-export gateway wants the
gateway's address split some other way. The kernel has no good basis
for choosing, and the address table this started as would need its own
configuration interface and a userspace tool to fill it.
Let the administrator supply the classification instead. Add
struct svc_classifier, a BPF struct_ops with one callback that is
handed each accepted svc_xprt, in process context with its peer and
local addresses set, and returns an SVC_CLASS_* value: SVC_CLASS_NONE
to leave the transport on the anonymous client, a class number shared
by every transport in that class, or a class number with
SVC_CLASS_PER_ADDR set for one client per peer address within the
class. A program keyed on an address-prefix map does the whole job in
a few lines. The callback runs under rcu_read_lock(), so sleepable
programs are refused.
Attaching a classifier (a struct_ops link) binds it to the attaching
task's network namespace, one per namespace and one namespace per
instance; a second attach fails with -EBUSY, a link update replaces
it, detaching or closing the link removes it, and a namespace that
exits leaves its link with nothing behind it. The classifier applies
to every RPC service in the namespace; a program that cares which one
reads xpt_server->sv_name. Transports keep the class they were given
at accept, so a change of classifier affects new connections only.
With no classifier attached, or with CONFIG_SUNRPC_BPF_CLASSIFY off,
svc_classify() returns SVC_CLASS_NONE and nothing changes. The
struct_ops registers from sunrpc's module init; a module without BTF
registers nothing, and the BPF core says so when the configuration
promised module BTF.
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
include/linux/sunrpc/svc.h | 24 ++++
include/linux/sunrpc/svc_xprt.h | 15 +++
net/sunrpc/Kconfig | 12 ++
net/sunrpc/Makefile | 1 +
net/sunrpc/netns.h | 4 +
net/sunrpc/sunrpc_syms.c | 4 +
net/sunrpc/svc_classify.c | 200 ++++++++++++++++++++++++++++++++
7 files changed, 260 insertions(+)
create mode 100644 net/sunrpc/svc_classify.c
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index a27dc2942167..5d8e35c5379e 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -96,6 +96,30 @@ struct svc_client {
#define SVC_CLASS_NONE 0
#define SVC_CLASS_PER_ADDR BIT_U32(31)
+#define SVC_CLASSIFIER_NAME_MAX 16
+
+struct svc_xprt;
+
+/**
+ * struct svc_classifier - BPF struct_ops that classifies accepted transports
+ * @name: a label for the instance, set by the program and not interpreted
+ * @classify: called once per accepted transport, in process context, under
+ * rcu_read_lock(), with the peer and local addresses set; returns an
+ * SVC_CLASS_* value
+ *
+ * Attaching one binds it to the attaching task's network namespace, where
+ * it classifies every service's transports until it is detached; one per
+ * namespace. Transports keep the class they were given when accepted.
+ */
+struct svc_classifier {
+ /* private: */
+ struct net *net;
+
+ /* public: */
+ char name[SVC_CLASSIFIER_NAME_MAX];
+ u32 (*classify)(const struct svc_xprt *xprt);
+};
+
struct svc_rqst;
diff --git a/include/linux/sunrpc/svc_xprt.h b/include/linux/sunrpc/svc_xprt.h
index d4f6dd5c8db3..10c18e948b42 100644
--- a/include/linux/sunrpc/svc_xprt.h
+++ b/include/linux/sunrpc/svc_xprt.h
@@ -180,11 +180,26 @@ void svc_xprt_destroy_all(struct svc_serv *serv, struct net *net,
void svc_xprt_received(struct svc_xprt *xprt);
/* The class of an accepted transport; SVC_CLASS_* in svc.h. */
+#if IS_ENABLED(CONFIG_SUNRPC_BPF_CLASSIFY)
+u32 svc_classify(const struct svc_xprt *xprt);
+void svc_classifier_exit_net(struct net *net);
+int svc_classifier_init_module(void);
+#else
static inline u32 svc_classify(const struct svc_xprt *xprt)
{
return SVC_CLASS_NONE;
}
+static inline void svc_classifier_exit_net(struct net *net)
+{
+}
+
+static inline int svc_classifier_init_module(void)
+{
+ return 0;
+}
+#endif
+
void svc_xprt_enqueue(struct svc_xprt *xprt);
void svc_xprt_put(struct svc_xprt *xprt);
void svc_xprt_copy_addrs(struct svc_rqst *rqstp, struct svc_xprt *xprt);
diff --git a/net/sunrpc/Kconfig b/net/sunrpc/Kconfig
index e7808e5714dc..2b96ac8d1ba2 100644
--- a/net/sunrpc/Kconfig
+++ b/net/sunrpc/Kconfig
@@ -16,6 +16,18 @@ config SUNRPC_SWAP
bool
depends on SUNRPC
+config SUNRPC_BPF_CLASSIFY
+ bool "BPF classification of RPC service transports"
+ depends on SUNRPC && BPF_SYSCALL && BPF_JIT
+ default y
+ help
+ Let a BPF struct_ops program attached in a network namespace say
+ which accepted transports of an RPC service (NFS, NLM, the NFSv4
+ callback service) share turns at the service's threads. With no
+ program attached, dispatch is unchanged.
+
+ If unsure, say Y.
+
config RPCSEC_GSS_KRB5
tristate "Secure RPC: Kerberos V mechanism"
depends on SUNRPC && CRYPTO
diff --git a/net/sunrpc/Makefile b/net/sunrpc/Makefile
index 96727df3aa85..fdcb860b616e 100644
--- a/net/sunrpc/Makefile
+++ b/net/sunrpc/Makefile
@@ -19,3 +19,4 @@ sunrpc-$(CONFIG_SUNRPC_DEBUG) += debugfs.o
sunrpc-$(CONFIG_SUNRPC_BACKCHANNEL) += backchannel_rqst.o
sunrpc-$(CONFIG_PROC_FS) += stats.o
sunrpc-$(CONFIG_SYSCTL) += sysctl.o
+sunrpc-$(CONFIG_SUNRPC_BPF_CLASSIFY) += svc_classify.o
diff --git a/net/sunrpc/netns.h b/net/sunrpc/netns.h
index 4efb5f28d881..ed01bc5296dd 100644
--- a/net/sunrpc/netns.h
+++ b/net/sunrpc/netns.h
@@ -34,6 +34,10 @@ struct sunrpc_net {
atomic_t pipe_users;
struct proc_dir_entry *use_gssp_proc;
struct proc_dir_entry *gss_krb5_enctypes;
+
+#if IS_ENABLED(CONFIG_SUNRPC_BPF_CLASSIFY)
+ struct svc_classifier __rcu *svc_classifier;
+#endif
};
extern unsigned int sunrpc_net_id;
diff --git a/net/sunrpc/sunrpc_syms.c b/net/sunrpc/sunrpc_syms.c
index 1a3884a0376a..4fdb676baa30 100644
--- a/net/sunrpc/sunrpc_syms.c
+++ b/net/sunrpc/sunrpc_syms.c
@@ -74,6 +74,7 @@ static __net_exit void sunrpc_exit_net(struct net *net)
{
struct sunrpc_net *sn = net_generic(net, sunrpc_net_id);
+ svc_classifier_exit_net(net);
rpc_pipefs_exit_net(net);
unix_gid_cache_destroy(net);
ip_map_cache_destroy(net);
@@ -122,6 +123,9 @@ init_sunrpc(void)
#endif
svc_init_xprt_sock(); /* svc sock transport */
init_socket_xprt(); /* clnt sock transport */
+ err = svc_classifier_init_module();
+ if (err)
+ pr_warn("RPC: no BPF transport classifier (%d)\n", err);
return 0;
out6:
diff --git a/net/sunrpc/svc_classify.c b/net/sunrpc/svc_classify.c
new file mode 100644
index 000000000000..aa63e0a3cb95
--- /dev/null
+++ b/net/sunrpc/svc_classify.c
@@ -0,0 +1,200 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * BPF struct_ops classification of accepted RPC service transports.
+ */
+
+#include <linux/bpf.h>
+#include <linux/bpf_verifier.h>
+#include <linux/btf.h>
+#include <linux/nsproxy.h>
+#include <linux/sunrpc/svc.h>
+#include <linux/sunrpc/svc_xprt.h>
+
+#include "netns.h"
+
+/* Orders attach, detach and replacement against namespace exit. */
+static DEFINE_SPINLOCK(svc_classifier_lock);
+
+/**
+ * svc_classify - ask the namespace's classifier for a transport's class
+ * @xprt: an accepted transport, peer and local addresses set
+ *
+ * Return: an SVC_CLASS_* value; %SVC_CLASS_NONE when no classifier is
+ * attached in the transport's network namespace.
+ */
+u32 svc_classify(const struct svc_xprt *xprt)
+{
+ struct sunrpc_net *sn = net_generic(xprt->xpt_net, sunrpc_net_id);
+ struct svc_classifier *c;
+ u32 class = SVC_CLASS_NONE;
+
+ rcu_read_lock();
+ c = rcu_dereference(sn->svc_classifier);
+ if (c)
+ class = c->classify(xprt);
+ rcu_read_unlock();
+ return class;
+}
+
+/* Caller holds svc_classifier_lock. */
+static void svc_classifier_detach(struct svc_classifier *c)
+{
+ struct sunrpc_net *sn;
+
+ if (!c->net)
+ return;
+ sn = net_generic(c->net, sunrpc_net_id);
+ RCU_INIT_POINTER(sn->svc_classifier, NULL);
+ c->net = NULL;
+}
+
+/**
+ * svc_classifier_exit_net - forget the classifier of a dying namespace
+ * @net: the network namespace being torn down
+ *
+ * The attached link, if any, survives with nothing behind it.
+ */
+void svc_classifier_exit_net(struct net *net)
+{
+ struct sunrpc_net *sn = net_generic(net, sunrpc_net_id);
+ struct svc_classifier *c;
+
+ spin_lock(&svc_classifier_lock);
+ c = rcu_dereference_protected(sn->svc_classifier,
+ lockdep_is_held(&svc_classifier_lock));
+ if (c)
+ svc_classifier_detach(c);
+ spin_unlock(&svc_classifier_lock);
+}
+
+static int svc_classifier_reg(void *kdata, struct bpf_link *link)
+{
+ struct svc_classifier *c = kdata;
+ struct net *net = current->nsproxy->net_ns;
+ struct sunrpc_net *sn = net_generic(net, sunrpc_net_id);
+ int err = 0;
+
+ if (!link)
+ return -EOPNOTSUPP;
+
+ spin_lock(&svc_classifier_lock);
+ /* one namespace per classifier, one classifier per namespace */
+ if (c->net || rcu_access_pointer(sn->svc_classifier)) {
+ err = -EBUSY;
+ } else {
+ c->net = net;
+ rcu_assign_pointer(sn->svc_classifier, c);
+ }
+ spin_unlock(&svc_classifier_lock);
+ return err;
+}
+
+static void svc_classifier_unreg(void *kdata, struct bpf_link *link)
+{
+ spin_lock(&svc_classifier_lock);
+ svc_classifier_detach(kdata);
+ spin_unlock(&svc_classifier_lock);
+ synchronize_rcu();
+}
+
+static int svc_classifier_update(void *kdata, void *old_kdata,
+ struct bpf_link *link)
+{
+ struct svc_classifier *c = kdata, *old = old_kdata;
+ struct sunrpc_net *sn;
+ int err = 0;
+
+ if (c == old)
+ return 0;
+
+ spin_lock(&svc_classifier_lock);
+ if (c->net) {
+ err = -EBUSY;
+ } else if (!old->net) {
+ err = -ENOENT;
+ } else {
+ c->net = old->net;
+ old->net = NULL;
+ sn = net_generic(c->net, sunrpc_net_id);
+ rcu_assign_pointer(sn->svc_classifier, c);
+ }
+ spin_unlock(&svc_classifier_lock);
+ if (!err)
+ synchronize_rcu();
+ return err;
+}
+
+static int svc_classifier_validate(void *kdata)
+{
+ struct svc_classifier *c = kdata;
+
+ return c->classify ? 0 : -EINVAL;
+}
+
+/* classify() runs under rcu_read_lock() */
+static int svc_classifier_check_member(const struct btf_type *t,
+ const struct btf_member *member,
+ const struct bpf_prog *prog)
+{
+ return prog->sleepable ? -EINVAL : 0;
+}
+
+static int svc_classifier_init_member(const struct btf_type *t,
+ const struct btf_member *member,
+ void *kdata, const void *udata)
+{
+ const struct svc_classifier *uc = udata;
+ struct svc_classifier *kc = kdata;
+ u32 moff = __btf_member_bit_offset(t, member) / 8;
+
+ if (moff == offsetof(struct svc_classifier, name)) {
+ if (bpf_obj_name_cpy(kc->name, uc->name, sizeof(uc->name)) <= 0)
+ return -EINVAL;
+ return 1;
+ }
+ return 0;
+}
+
+static int svc_classifier_init(struct btf *btf)
+{
+ return 0;
+}
+
+static u32 __svc_classifier_stub_classify(const struct svc_xprt *xprt)
+{
+ return SVC_CLASS_NONE;
+}
+
+static struct svc_classifier __svc_classifier_stubs = {
+ .classify = __svc_classifier_stub_classify,
+};
+
+static const struct bpf_func_proto *
+svc_classifier_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog)
+{
+ return bpf_base_func_proto(func_id, prog);
+}
+
+static const struct bpf_verifier_ops svc_classifier_verifier_ops = {
+ .get_func_proto = svc_classifier_func_proto,
+ .is_valid_access = bpf_tracing_btf_ctx_access,
+};
+
+static struct bpf_struct_ops bpf_svc_classifier_ops = {
+ .verifier_ops = &svc_classifier_verifier_ops,
+ .init = svc_classifier_init,
+ .check_member = svc_classifier_check_member,
+ .init_member = svc_classifier_init_member,
+ .reg = svc_classifier_reg,
+ .unreg = svc_classifier_unreg,
+ .update = svc_classifier_update,
+ .validate = svc_classifier_validate,
+ .cfi_stubs = &__svc_classifier_stubs,
+ .name = "svc_classifier",
+ .owner = THIS_MODULE,
+};
+
+int __init svc_classifier_init_module(void)
+{
+ return register_bpf_struct_ops(&bpf_svc_classifier_ops, svc_classifier);
+}
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH RFC v2 4/9] SUNRPC: report the transport's class in the svc_xprt_dequeue tracepoint
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
` (2 preceding siblings ...)
2026-10-07 19:59 ` [PATCH RFC v2 3/9] SUNRPC: add a BPF struct_ops hook to classify accepted transports Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 5/9] SUNRPC: dispatch control events ahead of data within a client Benjamin Coddington
` (6 subsequent siblings)
10 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
With dispatch round-robin across clients, which client a transport
was served as matters as much as its address. Print the class the
transport was bound to at accept: 0 for the anonymous client, the
class number otherwise, with "+addr" when the client is one address
within the class. The peer address is already in the event.
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
include/trace/events/sunrpc.h | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
diff --git a/include/trace/events/sunrpc.h b/include/trace/events/sunrpc.h
index 77efe63d2c14..107fbed5d556 100644
--- a/include/trace/events/sunrpc.h
+++ b/include/trace/events/sunrpc.h
@@ -2042,20 +2042,27 @@ TRACE_EVENT(svc_xprt_dequeue,
TP_STRUCT__entry(
SVC_XPRT_ENDPOINT_FIELDS(rqst->rq_xprt)
+ __field(u32, class)
+ __field(bool, per_addr)
__field(unsigned long, wakeup)
__field(unsigned long, qtime)
),
TP_fast_assign(
ktime_t ktime = ktime_get();
+ u32 class = rqst->rq_xprt->xpt_client->cl_class;
SVC_XPRT_ENDPOINT_ASSIGNMENTS(rqst->rq_xprt);
+ __entry->class = class & ~SVC_CLASS_PER_ADDR;
+ __entry->per_addr = class & SVC_CLASS_PER_ADDR;
__entry->wakeup = ktime_to_us(ktime_sub(ktime, rqst->rq_qtime));
__entry->qtime = ktime_to_us(ktime_sub(ktime, rqst->rq_xprt->xpt_qtime));
),
- TP_printk(SVC_XPRT_ENDPOINT_FORMAT " wakeup-us=%lu qtime-us=%lu",
- SVC_XPRT_ENDPOINT_VARARGS, __entry->wakeup, __entry->qtime)
+ TP_printk(SVC_XPRT_ENDPOINT_FORMAT " class=%u%s wakeup-us=%lu qtime-us=%lu",
+ SVC_XPRT_ENDPOINT_VARARGS,
+ __entry->class, __entry->per_addr ? "+addr" : "",
+ __entry->wakeup, __entry->qtime)
);
DECLARE_EVENT_CLASS(svc_xprt_event,
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH RFC v2 5/9] SUNRPC: dispatch control events ahead of data within a client
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
` (3 preceding siblings ...)
2026-10-07 19:59 ` [PATCH RFC v2 4/9] SUNRPC: report the transport's class in the svc_xprt_dequeue tracepoint Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 6/9] SUNRPC: keep the single transport queue while no classifier is attached Benjamin Coddington
` (5 subsequent siblings)
10 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
A transport that needs a connection accepted, closed, or a handshake
run waits its turn like a transport with data, behind every transport
of its client queued before it, and a client with many connections can
hold a close back for k rounds of all the busy clients. None of those
events costs the service a request's worth of work, and a close frees
resources the sooner it is served.
Give each client a second per-pool queue for transports with XPT_CONN,
XPT_CLOSE or XPT_HANDSHAKE set when they are enqueued, and have the
dispatcher serve that queue before the client's data queue. The
client's turn in the round-robin is unchanged; only which of its
transports goes first. A transport that is already queued with data
when it is closed is served in its original order, as before.
Suggested-by: Chuck Lever <chuck.lever@oracle.com>
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
include/linux/sunrpc/svc.h | 2 ++
include/linux/sunrpc/svc_xprt.h | 3 ++
net/sunrpc/svc.c | 1 +
net/sunrpc/svc_xprt.c | 60 +++++++++++++++++++++++++++------
4 files changed, 56 insertions(+), 10 deletions(-)
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index 5d8e35c5379e..02a8eb5d3357 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -66,6 +66,7 @@ enum {
* of the service does not grow with its connection count.
*/
struct svc_client_pool {
+ struct lwq cp_ctrl; /* ready transports with a control event */
struct lwq cp_xprts; /* ready transports */
struct lwq_node cp_ready; /* link in svc_pool.sp_clients */
unsigned long cp_flags;
@@ -74,6 +75,7 @@ struct svc_client_pool {
enum {
SVC_CP_QUEUED, /* on sp_clients, or held by the thread that took it */
+ SVC_CP_DATA_NEXT, /* last turn served a control event; data is owed */
};
struct svc_client {
diff --git a/include/linux/sunrpc/svc_xprt.h b/include/linux/sunrpc/svc_xprt.h
index 10c18e948b42..e05d7c8610ff 100644
--- a/include/linux/sunrpc/svc_xprt.h
+++ b/include/linux/sunrpc/svc_xprt.h
@@ -116,6 +116,9 @@ enum {
*/
};
+/* events a client's queue serves ahead of its data */
+#define SVC_XPT_CTRL_EVENTS (BIT(XPT_CONN) | BIT(XPT_CLOSE) | BIT(XPT_HANDSHAKE))
+
/*
* Maximum number of "tmp" connections - those without XPT_PEER_VALID -
* permitted on any service.
diff --git a/net/sunrpc/svc.c b/net/sunrpc/svc.c
index 4adf113d067b..92c6737e5905 100644
--- a/net/sunrpc/svc.c
+++ b/net/sunrpc/svc.c
@@ -405,6 +405,7 @@ struct svc_client *svc_client_alloc(unsigned int nrpools, gfp_t gfp)
refcount_set(&cl->cl_ref, 1);
INIT_HLIST_NODE(&cl->cl_hash);
for (i = 0; i < nrpools; i++) {
+ lwq_init(&cl->cl_pool[i].cp_ctrl);
lwq_init(&cl->cl_pool[i].cp_xprts);
cl->cl_pool[i].cp_client = cl;
}
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 735f4d92bf5d..14a7483886fc 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -646,7 +646,10 @@ void svc_xprt_enqueue(struct svc_xprt *xprt)
percpu_counter_inc(&pool->sp_sockets_queued);
xprt->xpt_qtime = ktime_get();
- lwq_enqueue(&xprt->xpt_ready, &cp->cp_xprts);
+ if (READ_ONCE(xprt->xpt_flags) & SVC_XPT_CTRL_EVENTS)
+ lwq_enqueue(&xprt->xpt_ready, &cp->cp_ctrl);
+ else
+ lwq_enqueue(&xprt->xpt_ready, &cp->cp_xprts);
if (!test_and_set_bit(SVC_CP_QUEUED, &cp->cp_flags))
lwq_enqueue(&cp->cp_ready, &pool->sp_clients);
@@ -654,6 +657,11 @@ void svc_xprt_enqueue(struct svc_xprt *xprt)
}
EXPORT_SYMBOL_GPL(svc_xprt_enqueue);
+static bool svc_client_pool_empty(struct svc_client_pool *cp)
+{
+ return lwq_empty(&cp->cp_ctrl) && lwq_empty(&cp->cp_xprts);
+}
+
/*
* Put @cp back on the pool's client queue if it still has ready
* transports. Otherwise clear SVC_CP_QUEUED and look once more: an
@@ -662,21 +670,47 @@ EXPORT_SYMBOL_GPL(svc_xprt_enqueue);
static void svc_client_pool_requeue(struct svc_pool *pool,
struct svc_client_pool *cp)
{
- if (!lwq_empty(&cp->cp_xprts)) {
+ if (!svc_client_pool_empty(cp)) {
lwq_enqueue(&cp->cp_ready, &pool->sp_clients);
return;
}
clear_bit(SVC_CP_QUEUED, &cp->cp_flags);
/* the clear is visible before the re-test; see svc_xprt_enqueue() */
smp_mb__after_atomic();
- if (!lwq_empty(&cp->cp_xprts) &&
+ if (!svc_client_pool_empty(cp) &&
!test_and_set_bit(SVC_CP_QUEUED, &cp->cp_flags))
lwq_enqueue(&cp->cp_ready, &pool->sp_clients);
}
/*
- * Dequeue the next transport: one from the client at the head of the
- * pool's client queue, which then goes to the tail if it has more.
+ * A client's control events go before its data, but not two turns in a
+ * row while data waits: a listener re-arms XPT_CONN after every accept
+ * and would otherwise hold the anonymous client's turns for as long as
+ * connections keep arriving. Only the thread holding @cp touches
+ * SVC_CP_DATA_NEXT.
+ */
+static struct svc_xprt *svc_client_pool_dequeue(struct svc_client_pool *cp)
+{
+ struct svc_xprt *xprt = NULL;
+ bool data_owed;
+
+ data_owed = test_and_clear_bit(SVC_CP_DATA_NEXT, &cp->cp_flags);
+ if (!data_owed)
+ xprt = lwq_dequeue(&cp->cp_ctrl, struct svc_xprt, xpt_ready);
+ if (!xprt) {
+ xprt = lwq_dequeue(&cp->cp_xprts, struct svc_xprt, xpt_ready);
+ if (xprt || !data_owed)
+ return xprt;
+ xprt = lwq_dequeue(&cp->cp_ctrl, struct svc_xprt, xpt_ready);
+ }
+ if (xprt && !lwq_empty(&cp->cp_xprts))
+ set_bit(SVC_CP_DATA_NEXT, &cp->cp_flags);
+ return xprt;
+}
+
+/*
+ * Dequeue the next transport from the pool's client queue: one from the
+ * client at the head, which then goes to the tail if it has more.
*/
static struct svc_xprt *svc_xprt_dequeue(struct svc_pool *pool)
{
@@ -688,7 +722,7 @@ static struct svc_xprt *svc_xprt_dequeue(struct svc_pool *pool)
cp_ready);
if (!cp)
return NULL;
- xprt = lwq_dequeue(&cp->cp_xprts, struct svc_xprt, xpt_ready);
+ xprt = svc_client_pool_dequeue(cp);
svc_client_pool_requeue(pool, cp);
} while (!xprt);
svc_xprt_get(xprt);
@@ -1371,15 +1405,15 @@ static int svc_close_list(struct svc_serv *serv, struct list_head *xprt_list, st
}
/*
- * Delete the ready transports of @net that are queued on @cp; the rest go
+ * Delete the ready transports of @net that are queued on @lq; the rest go
* back in their order.
*/
-static void svc_client_pool_clean(struct svc_client_pool *cp, struct net *net)
+static void svc_xprt_lwq_clean(struct lwq *lq, struct net *net)
{
struct llist_node *q, **t1, *t2;
struct svc_xprt *xprt;
- q = lwq_dequeue_all(&cp->cp_xprts);
+ q = lwq_dequeue_all(lq);
lwq_for_each_safe(xprt, t1, t2, &q, xpt_ready) {
if (xprt->xpt_net == net) {
set_bit(XPT_CLOSE, &xprt->xpt_flags);
@@ -1388,7 +1422,13 @@ static void svc_client_pool_clean(struct svc_client_pool *cp, struct net *net)
}
}
if (q)
- lwq_enqueue_batch(q, &cp->cp_xprts);
+ lwq_enqueue_batch(q, lq);
+}
+
+static void svc_client_pool_clean(struct svc_client_pool *cp, struct net *net)
+{
+ svc_xprt_lwq_clean(&cp->cp_ctrl, net);
+ svc_xprt_lwq_clean(&cp->cp_xprts, net);
}
static void svc_clean_up_xprts(struct svc_serv *serv, struct net *net)
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH RFC v2 6/9] SUNRPC: keep the single transport queue while no classifier is attached
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
` (4 preceding siblings ...)
2026-10-07 19:59 ` [PATCH RFC v2 5/9] SUNRPC: dispatch control events ahead of data within a client Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 7/9] selftests/bpf: add svc_classifier tests Benjamin Coddington
` (4 subsequent siblings)
10 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
With no classifier attached every transport is on the anonymous client
and a pool's client queue holds one entry, so dispatch is in today's
order but takes one more lwq dequeue and a re-queue per request than
today's code. That is about half a microsecond a dispatch on current
hardware, paid by every server whether or not it has asked for
anything.
Keep the pool's flat transport queue, sp_xprts, and use it while no
classifier is attached anywhere, switched by a static key that
attaching a classifier raises and detaching it drops. The dispatcher
looks first at whichever queue is not in use, so transports queued
before a switch are served promptly after it and nothing is lost in
either direction; the thread's sleep test covers both queues, and
network namespace teardown cleans both. With the key off the enqueue
and dequeue paths are today's, plus one empty test of the client queue
at dequeue.
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
include/linux/sunrpc/svc.h | 1 +
include/linux/sunrpc/svc_xprt.h | 4 +++
net/sunrpc/svc.c | 1 +
net/sunrpc/svc_classify.c | 20 ++++++++++----
net/sunrpc/svc_xprt.c | 46 ++++++++++++++++++++++++++++++---
5 files changed, 63 insertions(+), 9 deletions(-)
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index 02a8eb5d3357..49c1cebe5702 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -38,6 +38,7 @@ struct svc_pool {
unsigned int sp_nrthreads; /* # of threads currently running in pool */
unsigned int sp_nrthrmin; /* Min number of threads to run per pool */
unsigned int sp_nrthrmax; /* Max requested number of threads in pool */
+ struct lwq sp_xprts; /* ready transports, no classifier */
struct lwq sp_clients; /* clients with ready transports */
struct list_head sp_all_threads; /* all server threads */
struct llist_head sp_idle_threads; /* idle server threads */
diff --git a/include/linux/sunrpc/svc_xprt.h b/include/linux/sunrpc/svc_xprt.h
index e05d7c8610ff..ef21b2e57b98 100644
--- a/include/linux/sunrpc/svc_xprt.h
+++ b/include/linux/sunrpc/svc_xprt.h
@@ -8,6 +8,7 @@
#ifndef SUNRPC_SVC_XPRT_H
#define SUNRPC_SVC_XPRT_H
+#include <linux/jump_label.h>
#include <linux/sunrpc/svc.h>
struct module;
@@ -182,6 +183,9 @@ void svc_xprt_destroy_all(struct svc_serv *serv, struct net *net,
bool unregister);
void svc_xprt_received(struct svc_xprt *xprt);
+/* Set while a classifier is attached anywhere: dispatch by client. */
+DECLARE_STATIC_KEY_FALSE(svc_classify_key);
+
/* The class of an accepted transport; SVC_CLASS_* in svc.h. */
#if IS_ENABLED(CONFIG_SUNRPC_BPF_CLASSIFY)
u32 svc_classify(const struct svc_xprt *xprt);
diff --git a/net/sunrpc/svc.c b/net/sunrpc/svc.c
index 92c6737e5905..5f6a0fc06e66 100644
--- a/net/sunrpc/svc.c
+++ b/net/sunrpc/svc.c
@@ -478,6 +478,7 @@ __svc_create(struct svc_program *prog, int nprogs, struct svc_stat *stats,
i, serv->sv_name);
pool->sp_id = i;
+ lwq_init(&pool->sp_xprts);
lwq_init(&pool->sp_clients);
INIT_LIST_HEAD(&pool->sp_all_threads);
init_llist_head(&pool->sp_idle_threads);
diff --git a/net/sunrpc/svc_classify.c b/net/sunrpc/svc_classify.c
index aa63e0a3cb95..8d93d63157f5 100644
--- a/net/sunrpc/svc_classify.c
+++ b/net/sunrpc/svc_classify.c
@@ -36,16 +36,17 @@ u32 svc_classify(const struct svc_xprt *xprt)
return class;
}
-/* Caller holds svc_classifier_lock. */
-static void svc_classifier_detach(struct svc_classifier *c)
+/* Caller holds svc_classifier_lock; the static key is dropped after. */
+static bool svc_classifier_detach(struct svc_classifier *c)
{
struct sunrpc_net *sn;
if (!c->net)
- return;
+ return false;
sn = net_generic(c->net, sunrpc_net_id);
RCU_INIT_POINTER(sn->svc_classifier, NULL);
c->net = NULL;
+ return true;
}
/**
@@ -58,13 +59,16 @@ void svc_classifier_exit_net(struct net *net)
{
struct sunrpc_net *sn = net_generic(net, sunrpc_net_id);
struct svc_classifier *c;
+ bool detached = false;
spin_lock(&svc_classifier_lock);
c = rcu_dereference_protected(sn->svc_classifier,
lockdep_is_held(&svc_classifier_lock));
if (c)
- svc_classifier_detach(c);
+ detached = svc_classifier_detach(c);
spin_unlock(&svc_classifier_lock);
+ if (detached)
+ static_branch_dec(&svc_classify_key);
}
static int svc_classifier_reg(void *kdata, struct bpf_link *link)
@@ -86,14 +90,20 @@ static int svc_classifier_reg(void *kdata, struct bpf_link *link)
rcu_assign_pointer(sn->svc_classifier, c);
}
spin_unlock(&svc_classifier_lock);
+ if (!err)
+ static_branch_inc(&svc_classify_key);
return err;
}
static void svc_classifier_unreg(void *kdata, struct bpf_link *link)
{
+ bool detached;
+
spin_lock(&svc_classifier_lock);
- svc_classifier_detach(kdata);
+ detached = svc_classifier_detach(kdata);
spin_unlock(&svc_classifier_lock);
+ if (detached)
+ static_branch_dec(&svc_classify_key);
synchronize_rcu();
}
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 14a7483886fc..e18fd20923cd 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -170,6 +170,8 @@ void svc_xprt_deferred_close(struct svc_xprt *xprt)
}
EXPORT_SYMBOL_GPL(svc_xprt_deferred_close);
+DEFINE_STATIC_KEY_FALSE(svc_classify_key);
+
static unsigned int svc_client_hash(const struct net *net, u32 class,
const struct sockaddr *sa)
{
@@ -642,17 +644,22 @@ void svc_xprt_enqueue(struct svc_xprt *xprt)
return;
pool = svc_pool_for_cpu(xprt->xpt_server);
- cp = &xprt->xpt_client->cl_pool[pool->sp_id];
percpu_counter_inc(&pool->sp_sockets_queued);
xprt->xpt_qtime = ktime_get();
+ if (!static_branch_unlikely(&svc_classify_key)) {
+ lwq_enqueue(&xprt->xpt_ready, &pool->sp_xprts);
+ goto wake;
+ }
+
+ cp = &xprt->xpt_client->cl_pool[pool->sp_id];
if (READ_ONCE(xprt->xpt_flags) & SVC_XPT_CTRL_EVENTS)
lwq_enqueue(&xprt->xpt_ready, &cp->cp_ctrl);
else
lwq_enqueue(&xprt->xpt_ready, &cp->cp_xprts);
if (!test_and_set_bit(SVC_CP_QUEUED, &cp->cp_flags))
lwq_enqueue(&cp->cp_ready, &pool->sp_clients);
-
+wake:
svc_pool_wake_idle_thread(pool);
}
EXPORT_SYMBOL_GPL(svc_xprt_enqueue);
@@ -712,7 +719,7 @@ static struct svc_xprt *svc_client_pool_dequeue(struct svc_client_pool *cp)
* Dequeue the next transport from the pool's client queue: one from the
* client at the head, which then goes to the tail if it has more.
*/
-static struct svc_xprt *svc_xprt_dequeue(struct svc_pool *pool)
+static struct svc_xprt *svc_client_dequeue(struct svc_pool *pool)
{
struct svc_client_pool *cp;
struct svc_xprt *xprt;
@@ -729,6 +736,36 @@ static struct svc_xprt *svc_xprt_dequeue(struct svc_pool *pool)
return xprt;
}
+/*
+ * Dequeue the next transport. Transports are queued flat on sp_xprts
+ * while no classifier is attached and by client otherwise; whichever
+ * queue is not in use is looked at first, so that what was queued
+ * before a switch is served promptly after it.
+ */
+static struct svc_xprt *svc_xprt_dequeue(struct svc_pool *pool)
+{
+ struct svc_xprt *xprt = NULL;
+
+ if (static_branch_unlikely(&svc_classify_key)) {
+ if (!lwq_empty(&pool->sp_xprts))
+ xprt = lwq_dequeue(&pool->sp_xprts, struct svc_xprt,
+ xpt_ready);
+ if (xprt)
+ svc_xprt_get(xprt);
+ else
+ xprt = svc_client_dequeue(pool);
+ return xprt;
+ }
+ if (!lwq_empty(&pool->sp_clients))
+ xprt = svc_client_dequeue(pool);
+ if (!xprt) {
+ xprt = lwq_dequeue(&pool->sp_xprts, struct svc_xprt, xpt_ready);
+ if (xprt)
+ svc_xprt_get(xprt);
+ }
+ return xprt;
+}
+
/**
* svc_reserve - change the space reserved for the reply to a request.
* @rqstp: The request in question
@@ -952,7 +989,7 @@ svc_thread_should_sleep(struct svc_rqst *rqstp)
return false;
/* was a socket queued? */
- if (!lwq_empty(&pool->sp_clients))
+ if (!lwq_empty(&pool->sp_xprts) || !lwq_empty(&pool->sp_clients))
return false;
/* are we shutting down? */
@@ -1441,6 +1478,7 @@ static void svc_clean_up_xprts(struct svc_serv *serv, struct net *net)
struct svc_client *cl;
struct llist_node *q;
+ svc_xprt_lwq_clean(&pool->sp_xprts, net);
q = lwq_dequeue_all(&pool->sp_clients);
while (q) {
cp = container_of(q, struct svc_client_pool, cp_ready.node);
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH RFC v2 7/9] selftests/bpf: add svc_classifier tests
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
` (5 preceding siblings ...)
2026-10-07 19:59 ` [PATCH RFC v2 6/9] SUNRPC: keep the single transport queue while no classifier is attached Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader Benjamin Coddington
` (3 subsequent siblings)
10 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
A classifier for RPC service transports that looks the peer address
up in an LPM trie keyed by {family, address}, the reference for the
sunrpc svc_classifier struct_ops, and tests of the attach model: a
second attach in the same network namespace fails with -EBUSY, a link
update replaces the classifier, a detached instance can be attached
again, a classifier attached inside a namespace does not occupy the
initial one, and a namespace that exits leaves its link to detach into
nothing. Skipped where the struct_ops is not registered.
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
tools/testing/selftests/bpf/config | 2 +
.../selftests/bpf/prog_tests/svc_classifier.c | 104 ++++++++++++++++++
.../selftests/bpf/progs/bpf_svc_classifier.c | 80 ++++++++++++++
3 files changed, 186 insertions(+)
create mode 100644 tools/testing/selftests/bpf/prog_tests/svc_classifier.c
create mode 100644 tools/testing/selftests/bpf/progs/bpf_svc_classifier.c
diff --git a/tools/testing/selftests/bpf/config b/tools/testing/selftests/bpf/config
index ea7044f30adc..2cb7080d7da2 100644
--- a/tools/testing/selftests/bpf/config
+++ b/tools/testing/selftests/bpf/config
@@ -114,6 +114,7 @@ CONFIG_IP_NF_IPTABLES=y
CONFIG_IP6_NF_IPTABLES=y
CONFIG_IP6_NF_FILTER=y
CONFIG_NF_NAT=y
+CONFIG_NFS_FS=y
CONFIG_PACKET=y
CONFIG_RC_CORE=y
CONFIG_SAMPLES=y
@@ -133,6 +134,7 @@ CONFIG_TCP_CONG_BBR=y
CONFIG_INFINIBAND=y
CONFIG_SMC=y
CONFIG_SMC_HS_CTRL_BPF=y
+CONFIG_SUNRPC_BPF_CLASSIFY=y
CONFIG_DIBS=y
CONFIG_DIBS_LO=y
CONFIG_PM_WAKELOCKS=y
diff --git a/tools/testing/selftests/bpf/prog_tests/svc_classifier.c b/tools/testing/selftests/bpf/prog_tests/svc_classifier.c
new file mode 100644
index 000000000000..887af29b21c7
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/svc_classifier.c
@@ -0,0 +1,104 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * The svc_classifier struct_ops: one per network namespace, bound to
+ * the attaching task's namespace; replace by link update; detach by
+ * link close; a namespace that exits leaves its link empty.
+ */
+#include <test_progs.h>
+#include "bpf_svc_classifier.skel.h"
+
+static struct bpf_svc_classifier *load(void)
+{
+ struct bpf_svc_classifier *skel;
+
+ skel = bpf_svc_classifier__open_and_load();
+ if (!skel && (errno == ENOENT || errno == EINVAL || errno == EOPNOTSUPP)) {
+ test__skip();
+ return NULL;
+ }
+ ASSERT_OK_PTR(skel, "bpf_svc_classifier__open_and_load");
+ return skel;
+}
+
+static void test_attach(void)
+{
+ struct bpf_svc_classifier *one, *two;
+ struct bpf_link *link, *busy;
+ int err;
+
+ one = load();
+ if (!one)
+ return;
+ two = load();
+ if (!two)
+ goto out_one;
+
+ link = bpf_map__attach_struct_ops(one->maps.prefix);
+ if (!ASSERT_OK_PTR(link, "attach"))
+ goto out_two;
+
+ busy = bpf_map__attach_struct_ops(two->maps.prefix);
+ ASSERT_ERR_PTR(busy, "second attach");
+ ASSERT_EQ(errno, EBUSY, "second attach errno");
+
+ err = bpf_link__update_map(link, two->maps.prefix);
+ ASSERT_OK(err, "replace");
+
+ /* the replaced instance can be attached again once released */
+ bpf_link__destroy(link);
+ link = bpf_map__attach_struct_ops(one->maps.prefix);
+ ASSERT_OK_PTR(link, "attach after detach");
+ bpf_link__destroy(link);
+out_two:
+ bpf_svc_classifier__destroy(two);
+out_one:
+ bpf_svc_classifier__destroy(one);
+}
+
+static void test_netns(void)
+{
+ struct bpf_svc_classifier *inner, *outer;
+ struct netns_obj *ns;
+ struct bpf_link *link, *link2;
+
+ inner = load();
+ if (!inner)
+ return;
+ outer = load();
+ if (!outer)
+ goto out_inner;
+
+ ns = netns_new("svc_classifier_ns", true);
+ if (!ASSERT_OK_PTR(ns, "netns_new"))
+ goto out_outer;
+
+ /* attached inside the namespace */
+ link = bpf_map__attach_struct_ops(inner->maps.prefix);
+ if (!ASSERT_OK_PTR(link, "attach in netns"))
+ goto out_ns;
+
+ /* the initial namespace is free */
+ netns_free(ns);
+ ns = NULL;
+ link2 = bpf_map__attach_struct_ops(outer->maps.prefix);
+ ASSERT_OK_PTR(link2, "attach in init_net");
+ bpf_link__destroy(link2);
+
+ /* the link of the dead namespace detaches into nothing */
+ bpf_link__destroy(link);
+out_ns:
+ if (ns)
+ netns_free(ns);
+out_outer:
+ bpf_svc_classifier__destroy(outer);
+out_inner:
+ bpf_svc_classifier__destroy(inner);
+}
+
+void test_svc_classifier(void)
+{
+ if (test__start_subtest("attach"))
+ test_attach();
+ if (test__start_subtest("netns"))
+ test_netns();
+}
diff --git a/tools/testing/selftests/bpf/progs/bpf_svc_classifier.c b/tools/testing/selftests/bpf/progs/bpf_svc_classifier.c
new file mode 100644
index 000000000000..bae9d115ee74
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/bpf_svc_classifier.c
@@ -0,0 +1,80 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * A svc_classifier for RPC service transports: the peer address is looked
+ * up in an LPM trie keyed by {family, address}; a hit returns the class
+ * word stored there, a miss returns 0, the anonymous client. A key of
+ * prefixlen 0 is the default for every family, 8 the default for one.
+ */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+
+char _license[] SEC("license") = "GPL";
+
+#define AF_INET 2
+#define AF_INET6 10
+
+#define SVC_CLASSIFIER_NAME_MAX 16
+
+struct svc_xprt___local {
+ struct __kernel_sockaddr_storage xpt_remote;
+} __attribute__((preserve_access_index));
+
+struct svc_classifier___local {
+ char name[SVC_CLASSIFIER_NAME_MAX];
+ __u32 (*classify)(const struct svc_xprt___local *xprt);
+};
+
+struct prefix_key {
+ __u32 prefixlen; /* 8 (the family byte) + address bits */
+ __u8 family;
+ __u8 addr[16];
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_LPM_TRIE);
+ __type(key, struct prefix_key);
+ __type(value, __u32);
+ __uint(max_entries, 1024);
+ __uint(map_flags, BPF_F_NO_PREALLOC);
+} prefixes SEC(".maps");
+
+__u64 classified;
+
+SEC("struct_ops/classify")
+__u32 BPF_PROG(prefix_classify, const struct svc_xprt___local *xprt)
+{
+ struct __kernel_sockaddr_storage ss;
+ struct prefix_key key = {};
+ __u32 *class;
+
+ if (bpf_core_read(&ss, sizeof(ss), &xprt->xpt_remote))
+ return 0;
+
+ key.family = ss.ss_family;
+ switch (ss.ss_family) {
+ case AF_INET:
+ /* struct sockaddr_in: port at 2, address at 4 */
+ __builtin_memcpy(key.addr, &ss.__data[2], 4);
+ key.prefixlen = 8 + 32;
+ break;
+ case AF_INET6:
+ /* struct sockaddr_in6: port at 2, flowinfo at 4, address at 8 */
+ __builtin_memcpy(key.addr, &ss.__data[6], 16);
+ key.prefixlen = 8 + 128;
+ break;
+ default:
+ return 0;
+ }
+
+ __sync_fetch_and_add(&classified, 1);
+ class = bpf_map_lookup_elem(&prefixes, &key);
+ return class ? *class : 0;
+}
+
+SEC(".struct_ops.link")
+struct svc_classifier___local prefix = {
+ .name = "prefix",
+ .classify = (void *)prefix_classify,
+};
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
` (6 preceding siblings ...)
2026-10-07 19:59 ` [PATCH RFC v2 7/9] selftests/bpf: add svc_classifier tests Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-08 18:36 ` Jeff Layton
2026-10-07 19:59 ` [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier Benjamin Coddington
` (2 subsequent siblings)
10 siblings, 1 reply; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
An administrator who wants RPC service transports dispatched per client
needs a classifier attached in the service's network namespace and a
way to fill its map. Add svc-classify: the prefix classifier from the
BPF selftests, with its own type definitions in place of vmlinux.h,
embedded in a loader as a skeleton so the binary stands alone. load
attaches it and pins the link and the map, replace swaps the program
and keeps the map, unload detaches; add, del and list manage the map
with address prefixes and class words written as N, N+addr or addr;
status shows the link. The Makefile builds the in-tree libbpf and
bpftool it needs.
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
tools/net/sunrpc/svc-classify/.gitignore | 1 +
tools/net/sunrpc/svc-classify/Makefile | 65 ++++
tools/net/sunrpc/svc-classify/svc-classify.c | 357 ++++++++++++++++++
.../sunrpc/svc-classify/svc_classify.bpf.c | 83 ++++
4 files changed, 506 insertions(+)
create mode 100644 tools/net/sunrpc/svc-classify/.gitignore
create mode 100644 tools/net/sunrpc/svc-classify/Makefile
create mode 100644 tools/net/sunrpc/svc-classify/svc-classify.c
create mode 100644 tools/net/sunrpc/svc-classify/svc_classify.bpf.c
diff --git a/tools/net/sunrpc/svc-classify/.gitignore b/tools/net/sunrpc/svc-classify/.gitignore
new file mode 100644
index 000000000000..567609b1234a
--- /dev/null
+++ b/tools/net/sunrpc/svc-classify/.gitignore
@@ -0,0 +1 @@
+build/
diff --git a/tools/net/sunrpc/svc-classify/Makefile b/tools/net/sunrpc/svc-classify/Makefile
new file mode 100644
index 000000000000..9d86bc88417b
--- /dev/null
+++ b/tools/net/sunrpc/svc-classify/Makefile
@@ -0,0 +1,65 @@
+# SPDX-License-Identifier: GPL-2.0
+OUTPUT ?= $(CURDIR)/build/
+override OUTPUT := $(abspath $(OUTPUT))/
+$(shell mkdir -p $(OUTPUT))
+
+include ../../../build/Build.include
+include ../../../scripts/Makefile.arch
+include ../../../scripts/Makefile.include
+
+TOOLSDIR := $(abspath ../../..)
+BPFDIR := $(TOOLSDIR)/lib/bpf
+BPFTOOLDIR := $(TOOLSDIR)/bpf/bpftool
+APIDIR := $(TOOLSDIR)/include/uapi
+INCLUDE_DIR := $(OUTPUT)include
+BPFOBJ := $(OUTPUT)libbpf/libbpf.a
+DEFAULT_BPFTOOL := $(OUTPUT)sbin/bpftool
+BPFTOOL ?= $(DEFAULT_BPFTOOL)
+CLANG ?= clang
+msg = $(if $(Q),@printf ' %-8s %s\n' "$(1)" "$(3)";)
+
+prefix ?= /usr/local
+sbindir ?= $(prefix)/sbin
+INSTALL ?= install
+
+CFLAGS += -g -O2 -Wall -I$(INCLUDE_DIR) -I$(OUTPUT)
+LDLIBS += -lelf -lz
+BPF_CFLAGS := -g -O2 -target bpf -D__TARGET_ARCH_$(SRCARCH) \
+ -I$(INCLUDE_DIR) -I$(APIDIR) -Wall
+
+all: $(OUTPUT)svc-classify
+
+$(OUTPUT)libbpf $(OUTPUT)bpftool:
+ $(Q)mkdir -p $@
+
+$(BPFOBJ): $(wildcard $(BPFDIR)/*.[ch] $(BPFDIR)/Makefile) | $(OUTPUT)libbpf
+ $(Q)$(MAKE) $(submake_extras) -C $(BPFDIR) OUTPUT=$(OUTPUT)libbpf/ \
+ DESTDIR=$(OUTPUT) prefix= all install_headers
+
+$(DEFAULT_BPFTOOL): $(wildcard $(BPFTOOLDIR)/*.[ch] $(BPFTOOLDIR)/Makefile) \
+ $(BPFOBJ) | $(OUTPUT)bpftool
+ $(Q)$(MAKE) $(submake_extras) -C $(BPFTOOLDIR) OUTPUT=$(OUTPUT)bpftool/ \
+ LIBBPF_OUTPUT=$(OUTPUT)libbpf/ LIBBPF_DESTDIR=$(OUTPUT) \
+ prefix= DESTDIR=$(OUTPUT) install-bin
+
+$(OUTPUT)svc_classify.bpf.o: svc_classify.bpf.c $(BPFOBJ)
+ $(call msg,CLNG-BPF,,$(notdir $@))
+ $(Q)$(CLANG) $(BPF_CFLAGS) -c $< -o $@
+
+$(OUTPUT)svc_classify.skel.h: $(OUTPUT)svc_classify.bpf.o $(BPFTOOL)
+ $(call msg,GEN-SKEL,,$(notdir $@))
+ $(Q)$(BPFTOOL) gen skeleton $< name svc_classify > $@
+
+$(OUTPUT)svc-classify: svc-classify.c $(OUTPUT)svc_classify.skel.h $(BPFOBJ)
+ $(call msg,CC,,$(notdir $@))
+ $(Q)$(CC) $(CFLAGS) -o $@ $< $(BPFOBJ) $(LDLIBS)
+
+install: $(OUTPUT)svc-classify
+ $(Q)$(INSTALL) -d $(DESTDIR)$(sbindir)
+ $(Q)$(INSTALL) -m 755 $(OUTPUT)svc-classify $(DESTDIR)$(sbindir)/
+
+clean:
+ $(Q)rm -rf $(OUTPUT)
+
+.PHONY: all install clean
+.DELETE_ON_ERROR:
diff --git a/tools/net/sunrpc/svc-classify/svc-classify.c b/tools/net/sunrpc/svc-classify/svc-classify.c
new file mode 100644
index 000000000000..f722b891be2c
--- /dev/null
+++ b/tools/net/sunrpc/svc-classify/svc-classify.c
@@ -0,0 +1,357 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * svc-classify - attach the prefix classifier to the RPC services of this
+ * network namespace and manage its prefix map.
+ *
+ * svc-classify [-p DIR] load attach; pin the link and the map
+ * svc-classify [-p DIR] replace replace the program, keep the map
+ * svc-classify [-p DIR] unload detach
+ * svc-classify [-p DIR] add PREFIX CLASS
+ * svc-classify [-p DIR] del PREFIX
+ * svc-classify [-p DIR] list
+ * svc-classify [-p DIR] status
+ *
+ * PREFIX is a.b.c.d[/len], x:y::z[/len], any, any4 or any6. CLASS is N
+ * (one client for every peer that matches), N+addr (one client per peer
+ * address within class N), addr (the same as 0+addr) or 0 (the anonymous
+ * client). DIR, where the link and the map are pinned, defaults to
+ * /sys/fs/bpf/svc_classify.
+ */
+#include <arpa/inet.h>
+#include <errno.h>
+#include <limits.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/stat.h>
+#include <unistd.h>
+
+#include <bpf/bpf.h>
+#include <bpf/libbpf.h>
+
+#include "svc_classify.skel.h"
+
+#define PIN_DIR_DEFAULT "/sys/fs/bpf/svc_classify"
+#define SVC_CLASS_PER_ADDR (1U << 31)
+
+struct prefix_key {
+ __u32 prefixlen;
+ __u8 family;
+ __u8 addr[16];
+};
+
+static const char *progname;
+static char pin_link[PATH_MAX], pin_map[PATH_MAX];
+static const char *pin_dir = PIN_DIR_DEFAULT;
+
+static void usage(void)
+{
+ fprintf(stderr,
+ "usage: %s [-p DIR] load | replace | unload | list | status |\n"
+ " add PREFIX CLASS | del PREFIX\n"
+ "PREFIX: a.b.c.d[/len], x:y::z[/len], any, any4, any6\n"
+ "CLASS: N, N+addr, addr, 0\n", progname);
+ exit(2);
+}
+
+static int do_load(void)
+{
+ struct svc_classify *skel;
+ struct bpf_link *link;
+ int err;
+
+ if (access(pin_link, F_OK) == 0) {
+ fprintf(stderr, "%s: already loaded (%s exists)\n", progname,
+ pin_link);
+ return 1;
+ }
+ mkdir(pin_dir, 0700);
+
+ skel = svc_classify__open_and_load();
+ if (!skel) {
+ fprintf(stderr, "%s: load: %s\n", progname, strerror(errno));
+ rmdir(pin_dir);
+ return 1;
+ }
+ link = bpf_map__attach_struct_ops(skel->maps.prefix);
+ if (!link) {
+ fprintf(stderr, "%s: attach: %s\n", progname, strerror(errno));
+ rmdir(pin_dir);
+ return 1;
+ }
+ err = bpf_link__pin(link, pin_link);
+ if (!err)
+ err = bpf_map__pin(skel->maps.prefixes, pin_map);
+ if (err) {
+ fprintf(stderr, "%s: pin: %s\n", progname, strerror(-err));
+ bpf_link__destroy(link);
+ unlink(pin_link);
+ rmdir(pin_dir);
+ return 1;
+ }
+ return 0;
+}
+
+static int do_replace(void)
+{
+ struct svc_classify *skel;
+ struct bpf_link *link;
+ int link_fd, err;
+
+ link_fd = bpf_obj_get(pin_link);
+ if (link_fd < 0) {
+ fprintf(stderr, "%s: not loaded (%s)\n", progname,
+ strerror(errno));
+ return 1;
+ }
+ skel = svc_classify__open();
+ if (!skel) {
+ fprintf(stderr, "%s: open: %s\n", progname, strerror(errno));
+ return 1;
+ }
+ /*
+ * Reuse the pinned map so the entries survive. A build whose map
+ * definition differs fails here; unload and load then.
+ */
+ err = bpf_map__set_pin_path(skel->maps.prefixes, pin_map);
+ if (!err)
+ err = svc_classify__load(skel);
+ if (err) {
+ fprintf(stderr, "%s: load: %s\n", progname, strerror(-err));
+ return 1;
+ }
+ /*
+ * bpf_link__update_map() wants the link that attached the map, not
+ * one reopened from a pin. bpf_map__attach_struct_ops() writes the
+ * program into the new map's value before it tries to create a link,
+ * and the link is refused with EBUSY while the old one is attached;
+ * the written map is what the link update needs.
+ */
+ link = bpf_map__attach_struct_ops(skel->maps.prefix);
+ if (link) {
+ bpf_link__destroy(link);
+ fprintf(stderr, "%s: nothing attached in this namespace; unload and load\n",
+ progname);
+ return 1;
+ }
+ if (errno != EBUSY) {
+ fprintf(stderr, "%s: attach: %s\n", progname, strerror(errno));
+ return 1;
+ }
+ err = bpf_link_update(link_fd, bpf_map__fd(skel->maps.prefix), NULL);
+ if (err) {
+ fprintf(stderr, "%s: replace: %s\n", progname, strerror(errno));
+ return 1;
+ }
+ return 0;
+}
+
+static int do_unload(void)
+{
+ int ret = 0;
+
+ if (unlink(pin_link) && errno != ENOENT) {
+ perror(pin_link);
+ ret = 1;
+ }
+ if (unlink(pin_map) && errno != ENOENT) {
+ perror(pin_map);
+ ret = 1;
+ }
+ rmdir(pin_dir);
+ return ret;
+}
+
+static int parse_prefix(const char *s, struct prefix_key *key)
+{
+ char buf[64], *slash;
+ int bits, max;
+
+ memset(key, 0, sizeof(*key));
+ if (!strcmp(s, "any"))
+ return 0;
+ if (!strcmp(s, "any4")) {
+ key->family = AF_INET;
+ key->prefixlen = 8;
+ return 0;
+ }
+ if (!strcmp(s, "any6")) {
+ key->family = AF_INET6;
+ key->prefixlen = 8;
+ return 0;
+ }
+ if (strlen(s) >= sizeof(buf))
+ return -1;
+ strcpy(buf, s);
+ slash = strchr(buf, '/');
+ if (slash)
+ *slash++ = '\0';
+ if (inet_pton(AF_INET, buf, key->addr) == 1) {
+ key->family = AF_INET;
+ max = 32;
+ } else if (inet_pton(AF_INET6, buf, key->addr) == 1) {
+ key->family = AF_INET6;
+ max = 128;
+ } else {
+ return -1;
+ }
+ bits = max;
+ if (slash) {
+ char *end;
+
+ if (*slash < '0' || *slash > '9')
+ return -1;
+ bits = strtoul(slash, &end, 10);
+ if (*end || bits > max)
+ return -1;
+ }
+ key->prefixlen = 8 + bits;
+ return 0;
+}
+
+static int parse_class(const char *s, __u32 *class)
+{
+ unsigned long n;
+ char *end;
+
+ if (!strcmp(s, "addr")) {
+ *class = SVC_CLASS_PER_ADDR;
+ return 0;
+ }
+ n = strtoul(s, &end, 0);
+ if (end == s || n >= SVC_CLASS_PER_ADDR)
+ return -1;
+ *class = n;
+ if (!strcmp(end, "+addr"))
+ *class |= SVC_CLASS_PER_ADDR;
+ else if (*end)
+ return -1;
+ return 0;
+}
+
+static void print_entry(const struct prefix_key *key, __u32 class)
+{
+ char addr[INET6_ADDRSTRLEN];
+
+ if (key->prefixlen == 0)
+ printf("any");
+ else if (key->prefixlen == 8)
+ printf("any%d", key->family == AF_INET ? 4 : 6);
+ else
+ printf("%s/%u", inet_ntop(key->family, key->addr, addr,
+ sizeof(addr)), key->prefixlen - 8);
+ printf(" -> %u%s\n", class & ~SVC_CLASS_PER_ADDR,
+ class & SVC_CLASS_PER_ADDR ? "+addr" : "");
+}
+
+static int open_map(void)
+{
+ int fd = bpf_obj_get(pin_map);
+
+ if (fd < 0) {
+ fprintf(stderr, "%s: not loaded (%s: %s)\n", progname, pin_map,
+ strerror(errno));
+ exit(1);
+ }
+ return fd;
+}
+
+static int do_add(const char *prefix, const char *cls)
+{
+ struct prefix_key key;
+ __u32 class;
+
+ if (parse_prefix(prefix, &key) || parse_class(cls, &class))
+ usage();
+ if (bpf_map_update_elem(open_map(), &key, &class, BPF_ANY)) {
+ fprintf(stderr, "%s: add: %s\n", progname, strerror(errno));
+ return 1;
+ }
+ return 0;
+}
+
+static int do_del(const char *prefix)
+{
+ struct prefix_key key;
+
+ if (parse_prefix(prefix, &key))
+ usage();
+ if (bpf_map_delete_elem(open_map(), &key)) {
+ fprintf(stderr, "%s: del: %s\n", progname, strerror(errno));
+ return 1;
+ }
+ return 0;
+}
+
+static int do_list(void)
+{
+ struct prefix_key key, next;
+ __u32 class;
+ int fd;
+
+ fd = open_map();
+ if (bpf_map_get_next_key(fd, NULL, &next))
+ return 0;
+ do {
+ key = next;
+ if (!bpf_map_lookup_elem(fd, &key, &class))
+ print_entry(&key, class);
+ } while (!bpf_map_get_next_key(fd, &key, &next));
+ return 0;
+}
+
+static int do_status(void)
+{
+ struct bpf_link_info info = {};
+ __u32 len = sizeof(info);
+ int fd;
+
+ fd = bpf_obj_get(pin_link);
+ if (fd < 0) {
+ printf("not loaded\n");
+ return 1;
+ }
+ if (bpf_link_get_info_by_fd(fd, &info, &len))
+ printf("loaded\n");
+ else
+ printf("loaded, link id %u, struct_ops map id %u\n", info.id,
+ info.struct_ops.map_id);
+ return 0;
+}
+
+int main(int argc, char **argv)
+{
+ const char *cmd;
+ int opt;
+
+ progname = argv[0];
+ while ((opt = getopt(argc, argv, "p:")) != -1) {
+ if (opt != 'p')
+ usage();
+ pin_dir = optarg;
+ }
+ argc -= optind;
+ argv += optind;
+ if (argc < 1)
+ usage();
+ cmd = argv[0];
+ snprintf(pin_link, sizeof(pin_link), "%s/link", pin_dir);
+ snprintf(pin_map, sizeof(pin_map), "%s/prefixes", pin_dir);
+
+ if (!strcmp(cmd, "load") && argc == 1)
+ return do_load();
+ if (!strcmp(cmd, "replace") && argc == 1)
+ return do_replace();
+ if (!strcmp(cmd, "unload") && argc == 1)
+ return do_unload();
+ if (!strcmp(cmd, "add") && argc == 3)
+ return do_add(argv[1], argv[2]);
+ if (!strcmp(cmd, "del") && argc == 2)
+ return do_del(argv[1]);
+ if (!strcmp(cmd, "list") && argc == 1)
+ return do_list();
+ if (!strcmp(cmd, "status") && argc == 1)
+ return do_status();
+ usage();
+ return 2;
+}
diff --git a/tools/net/sunrpc/svc-classify/svc_classify.bpf.c b/tools/net/sunrpc/svc-classify/svc_classify.bpf.c
new file mode 100644
index 000000000000..3d012f102ec3
--- /dev/null
+++ b/tools/net/sunrpc/svc-classify/svc_classify.bpf.c
@@ -0,0 +1,83 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * A svc_classifier for RPC service transports: the peer address is looked
+ * up in an LPM trie keyed by {family, address}; a hit returns the class
+ * word stored there, a miss returns 0, the anonymous client. A key of
+ * prefixlen 0 is the default for every family, 8 the default for one.
+ */
+#include <linux/types.h>
+#include <linux/bpf.h>
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_tracing.h>
+#include <bpf/bpf_core_read.h>
+
+char _license[] SEC("license") = "GPL";
+
+#define AF_INET 2
+#define AF_INET6 10
+
+#define SVC_CLASSIFIER_NAME_MAX 16
+
+struct sockaddr_storage___local {
+ unsigned short ss_family;
+ char __data[126];
+};
+
+struct svc_xprt___local {
+ struct sockaddr_storage___local xpt_remote;
+} __attribute__((preserve_access_index));
+
+struct svc_classifier___local {
+ char name[SVC_CLASSIFIER_NAME_MAX];
+ __u32 (*classify)(const struct svc_xprt___local *xprt);
+};
+
+struct prefix_key {
+ __u32 prefixlen; /* 8 (the family byte) + address bits */
+ __u8 family;
+ __u8 addr[16];
+};
+
+struct {
+ __uint(type, BPF_MAP_TYPE_LPM_TRIE);
+ __type(key, struct prefix_key);
+ __type(value, __u32);
+ __uint(max_entries, 1024);
+ __uint(map_flags, BPF_F_NO_PREALLOC);
+} prefixes SEC(".maps");
+
+SEC("struct_ops/classify")
+__u32 BPF_PROG(prefix_classify, const struct svc_xprt___local *xprt)
+{
+ struct sockaddr_storage___local ss;
+ struct prefix_key key = {};
+ __u32 *class;
+
+ if (bpf_core_read(&ss, sizeof(ss), &xprt->xpt_remote))
+ return 0;
+
+ key.family = ss.ss_family;
+ switch (ss.ss_family) {
+ case AF_INET:
+ /* struct sockaddr_in: port at 2, address at 4 */
+ __builtin_memcpy(key.addr, &ss.__data[2], 4);
+ key.prefixlen = 8 + 32;
+ break;
+ case AF_INET6:
+ /* struct sockaddr_in6: port at 2, flowinfo at 4, address at 8 */
+ __builtin_memcpy(key.addr, &ss.__data[6], 16);
+ key.prefixlen = 8 + 128;
+ break;
+ default:
+ return 0;
+ }
+
+ class = bpf_map_lookup_elem(&prefixes, &key);
+ return class ? *class : 0;
+}
+
+SEC(".struct_ops.link")
+struct svc_classifier___local prefix = {
+ .name = "prefix",
+ .classify = (void *)prefix_classify,
+};
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
` (7 preceding siblings ...)
2026-10-07 19:59 ` [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader Benjamin Coddington
@ 2026-10-07 19:59 ` Benjamin Coddington
2026-10-08 18:42 ` Jeff Layton
2026-10-08 15:19 ` [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Chuck Lever
2026-10-08 18:48 ` Jeff Layton
10 siblings, 1 reply; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-07 19:59 UTC (permalink / raw)
To: Chuck Lever, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
From: Benjamin Coddington <bcodding@hammerspace.com>
Dispatch round-robin across clients changes how a service shares its
threads, and the classes come from a BPF program the administrator
loads, so the administrator needs the model, the class word, and the
procedure. Add a page with those: install a classifier with
svc-classify (build, attach, fill the map, check, change, remove, keep
across reboots), six class maps and the share each gives its clients,
the tracepoint that shows the class, and how the hook, the program and
the attach model work underneath.
Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
---
Documentation/filesystems/nfs/index.rst | 1 +
.../filesystems/nfs/rpc-server-clients.rst | 394 ++++++++++++++++++
2 files changed, 395 insertions(+)
create mode 100644 Documentation/filesystems/nfs/rpc-server-clients.rst
diff --git a/Documentation/filesystems/nfs/index.rst b/Documentation/filesystems/nfs/index.rst
index a29a212b5b4d..57a61ce0533d 100644
--- a/Documentation/filesystems/nfs/index.rst
+++ b/Documentation/filesystems/nfs/index.rst
@@ -16,3 +16,4 @@ NFS
nfsd-io-modes
knfsd-stats
reexport
+ rpc-server-clients
diff --git a/Documentation/filesystems/nfs/rpc-server-clients.rst b/Documentation/filesystems/nfs/rpc-server-clients.rst
new file mode 100644
index 000000000000..b08351a88314
--- /dev/null
+++ b/Documentation/filesystems/nfs/rpc-server-clients.rst
@@ -0,0 +1,394 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+======================================================
+RPC server: per-client dispatch and transport classes
+======================================================
+
+An RPC service thread pool serves its ready transports in FIFO order,
+one request per turn. A peer with many connections (``nconnect``, a
+deep NFSv4.1 slot table, a data mover with a thread per file) takes a
+turn per connection, and a peer with one connection waits behind all
+of them.
+
+With a classifier attached, the pool instead serves its *clients* round
+robin: each client with work queued gets one request per round, however
+many connections it holds. Which transports form a client is decided
+by a BPF program the administrator loads, once per accepted connection.
+With no program loaded nothing changes.
+
+This document describes the model, how to install a classifier with
+the ``svc-classify`` tool, what some class maps do, and how it works
+underneath.
+
+Clients, classes and turns
+==========================
+
+A *client* is a set of transports that share one turn. Every service
+(nfsd, lockd, the NFSv4 callback service) keeps its own clients, found
+by *class* and network namespace, or by class, namespace and peer
+address.
+
+The classifier returns a 32-bit class word for each accepted transport:
+
+``0``
+ ``SVC_CLASS_NONE``: the transport stays on the service's *anonymous
+ client*.
+
+``N``, 1 to 2^31 - 1
+ one client for every transport of class ``N``.
+
+``N | SVC_CLASS_PER_ADDR``
+ one client per peer address within class ``N``. ``N`` may be 0, so
+ ``SVC_CLASS_PER_ADDR`` alone means one client per peer address.
+
+``SVC_CLASS_PER_ADDR`` is bit 31. ``svc-classify`` and this document
+write a class word with the bit set as ``N+addr``; ``svc-classify``
+also accepts ``addr`` for ``0+addr``.
+
+Dispatch rules:
+
+- A pool serves the clients that have transports queued round robin,
+ one request per client per round.
+- Within a client, transports are served in FIFO order, except that a
+ transport needing a connection accepted, closed, or a TLS handshake
+ run goes ahead of transports with data, though not two such turns in
+ a row while data waits.
+- A client's share of the pool does not grow with its connection count.
+ A client with one connection and a client with thirty get the same
+ number of turns while both have work queued.
+
+Things that follow from the rules and are easy to get wrong:
+
+- The anonymous client is one client. Once any classifier is attached,
+ every transport whose class is ``0`` shares a single turn with every
+ other such transport of that service. Within that one client the old
+ FIFO order applies, so those peers share the turn in proportion to
+ their connection counts. A map that classifies only the hosts it
+ cares about and leaves the rest at ``0`` gives "the rest" one turn
+ in total (see example 5 below).
+- Per-address clients ignore the port. Only IPv4 and IPv6 peers can be
+ keyed by address; a transport whose peer is anything else is left on
+ the anonymous client.
+- Listeners and UDP sockets are never classified; UDP traffic is served
+ from the anonymous client.
+- A transport keeps the class it was given when accepted. Changing the
+ map, or replacing or removing the classifier, affects connections
+ accepted afterwards. A connection that must be reclassified has to
+ reconnect.
+- nfsd has one service per network namespace, so its anonymous client
+ is per namespace. lockd and the NFSv4 callback service are one
+ service for all namespaces; their anonymous clients span namespaces.
+- With no classifier attached anywhere, the pool uses a single FIFO;
+ the per-request cost of the feature is then a static branch on
+ enqueue and one empty-queue test on dequeue.
+ ``CONFIG_SUNRPC_BPF_CLASSIFY`` (default y) builds the hook; without
+ it there is no classification.
+
+Installing a classifier
+=======================
+
+There is one tool to do it with: ``svc-classify``, in
+``tools/net/sunrpc/svc-classify`` of the kernel source. It carries the
+classifier program inside the binary, attaches it, and manages the
+class map by address prefix. ``bpftool`` is not needed, and its
+``struct_ops`` subcommands cannot be used in its place (see "How it
+works").
+
+What you need
+-------------
+
+- A kernel with ``CONFIG_SUNRPC_BPF_CLASSIFY`` (default y) and BTF:
+ ``CONFIG_DEBUG_INFO_BTF``, and ``CONFIG_DEBUG_INFO_BTF_MODULES`` when
+ ``sunrpc`` is a module.
+- ``/sys/fs/bpf`` mounted (systemd mounts it).
+- To build the tool: ``clang``, and the ``libelf`` and ``zlib``
+ development files. The tool builds the libbpf and bpftool it needs
+ from the kernel source tree.
+
+Build and install the tool
+--------------------------
+
+In the kernel source tree::
+
+ $ make -C tools/net/sunrpc/svc-classify
+ # make -C tools/net/sunrpc/svc-classify install
+
+The binary goes to ``/usr/local/sbin/svc-classify`` (``prefix=`` and
+``sbindir=`` override that). At run time it needs only libelf and
+zlib.
+
+Attach it
+---------
+
+As root, in the network namespace the service runs in (on a host that
+is the initial namespace; in a container, ``nsenter`` or
+``ip netns exec`` into it first)::
+
+ # svc-classify load
+ # svc-classify status
+ loaded, link id 46, struct_ops map id 140
+
+The classifier is now attached to every RPC service in the namespace,
+with an empty map: every connection accepted from now on returns class
+``0`` and stays on the anonymous client, so nothing has changed yet.
+One classifier per namespace: a second ``load`` is refused.
+
+Install the class map
+---------------------
+
+Add one line per address prefix. The longest matching prefix wins; a
+peer that matches nothing gets class ``0``::
+
+ # svc-classify add 192.0.2.0/24 1
+ # svc-classify add any addr
+ # svc-classify list
+ 192.0.2.0/24 -> 1
+ any -> 0+addr
+
+``PREFIX`` is ``a.b.c.d[/len]``, ``x:y::z[/len]``, ``any`` (every
+address of either family), ``any4`` or ``any6``. ``CLASS`` is ``N``
+(one client for every peer that matches), ``N+addr`` (one client per
+peer address, within class ``N``), ``addr`` (one client per peer
+address, the same as ``0+addr``), or ``0`` (the anonymous client).
+
+Entries apply to connections accepted after they are added; to
+reclassify existing connections, have the clients reconnect (restarting
+the service does that for all of them). ``svc-classify del PREFIX``
+removes an entry.
+
+Check it
+--------
+
+``svc-classify status`` and ``svc-classify list`` show the link and the
+map. The ``sunrpc:svc_xprt_dequeue`` tracepoint shows the class every
+transport is dispatched as (see "Observing"), which is the check that
+the map does what was meant.
+
+Change or remove it
+-------------------
+
+- ``svc-classify add`` and ``del`` change the map at any time.
+- ``svc-classify replace`` installs a new build of the tool's program
+ on the attached link and keeps the map, provided the new build's map
+ definition is unchanged; otherwise unload and load.
+- ``svc-classify unload`` detaches the classifier. Connections accepted
+ afterwards are anonymous; when no classifier is attached in any
+ namespace, the service is back on its single FIFO.
+
+Across reboots
+--------------
+
+Nothing persists: the link and the map live in ``/sys/fs/bpf`` and are
+gone at boot. Run the ``load`` and ``add`` commands before the service
+starts, for example from a unit ordered before ``nfs-server.service``::
+
+ [Unit]
+ Description=RPC service transport classifier
+ Before=nfs-server.service
+
+ [Service]
+ Type=oneshot
+ RemainAfterExit=yes
+ ExecStart=/usr/local/sbin/svc-classify load
+ ExecStart=/usr/local/sbin/svc-classify add 192.0.2.0/24 1
+ ExecStart=/usr/local/sbin/svc-classify add any addr
+ ExecStop=/usr/local/sbin/svc-classify unload
+
+ [Install]
+ WantedBy=nfs-server.service
+
+Loading after the service is up works too; it only misses the
+connections already accepted.
+
+Examples
+========
+
+Each example gives the ``svc-classify add`` lines, the hosts they are
+applied to, and the share of the pool each client gets while all of
+them have work queued.
+
+1. Every host its own client
+----------------------------
+
+::
+
+ # svc-classify add any addr
+
+Three hosts: one with eight connections, one with one, one with two.
+Each host is a client, and each gets a third of the turns. Without a
+classifier they would be served in proportion to their connections,
+8:1:2.
+
+2. A set of movers as one client, everyone else per host
+--------------------------------------------------------
+
+::
+
+ # svc-classify add 192.0.2.0/24 1
+ # svc-classify add any addr
+
+Three movers in ``192.0.2.0/24`` with four connections each, and two
+other hosts, one with one connection and one with two. The movers are
+one client; each of the other hosts is a client of its own. Each of
+the three clients gets a third of the pool; the movers share theirs by
+connection count, a ninth each. Without a classifier the movers'
+twelve connections would take twelve fifteenths of the pool and the
+one-connection host one fifteenth.
+
+3. Two mover groups, one turn each
+----------------------------------
+
+::
+
+ # svc-classify add 192.0.2.0/25 1
+ # svc-classify add 192.0.2.128/25 2
+ # svc-classify add any addr
+
+Two hosts in ``192.0.2.0/25`` and two in ``192.0.2.128/25``, four
+connections each, and one other host with one connection. Each group
+is a client and the other host is a client: three clients, a third
+each. Within a group the two hosts split the group's third by
+connection count, a sixth each.
+
+4. Per host inside a campus, the rest of the world as one client
+----------------------------------------------------------------
+
+::
+
+ # svc-classify add 198.51.100.0/24 1+addr
+ # svc-classify add any 2
+
+Two hosts in ``198.51.100.0/24`` and two hosts outside it. Each campus
+host is a client of its own (class 1, one client per address); the
+outside hosts together are one client (class 2). Three clients, a
+third each; the outside hosts split their third by connection count.
+
+5. Leaving hosts unclassified
+-----------------------------
+
+::
+
+ # svc-classify add 203.0.113.0/24 0
+ # svc-classify add any addr
+
+Two hosts in ``203.0.113.0/24`` and two hosts elsewhere. The two in
+``203.0.113.0/24`` return ``0`` and so share the anonymous client, one
+turn between them; the two elsewhere are a client each. Three
+clients, a third each; the anonymous pair split their third by
+connection count.
+
+The common mistake is the map with only the movers in it::
+
+ # svc-classify add 192.0.2.0/24 1
+
+Three movers in ``192.0.2.0/24`` and two other hosts, one with one
+connection and one with two. There are exactly two clients: the
+movers, and everyone else on the anonymous client. The movers get
+half the pool. The other half goes to the two other hosts in
+proportion to their connections, a sixth and a third of the pool. Add
+``any addr`` to serve them per host.
+
+6. IPv6
+-------
+
+::
+
+ # svc-classify add 2001:db8:1::/48 1
+ # svc-classify add any addr
+
+Entries are per family; ``any`` covers both families, ``any6`` IPv6
+only. Two hosts in ``2001:db8:1::/48`` and one host elsewhere: the two
+are one client, the other is a client of its own, half the pool each.
+
+Observing
+=========
+
+The ``sunrpc:svc_xprt_dequeue`` tracepoint reports the class of every
+transport as it is dispatched::
+
+ svc_xprt_dequeue: server=127.0.0.1:3049 client=127.0.1.1:38209 xpt_id=1164 flags=BUSY|DATA|TEMP|CACHE_AUTH|LOCAL|CONG_CTRL class=1 wakeup-us=30 qtime-us=4
+ svc_xprt_dequeue: server=127.0.0.1:3049 client=127.0.2.1:45965 xpt_id=1162 flags=BUSY|DATA|TEMP|CACHE_AUTH|LOCAL|CONG_CTRL class=0+addr wakeup-us=60 qtime-us=10
+ svc_xprt_dequeue: server=[::1]:3049 client=[2001:db8:1::1]:44401 xpt_id=1152 flags=BUSY|DATA|TEMP|CACHE_AUTH|LOCAL|CONG_CTRL class=1 wakeup-us=62 qtime-us=7
+ svc_xprt_dequeue: server=[::]:3049 client=(einval) xpt_id=2 flags=BUSY|CONN|CHNGBUF|LISTENER|CACHE_AUTH|CONG_CTRL|RPCB_UNREG class=0 wakeup-us=25 qtime-us=25
+
+``class=0`` is the anonymous client; ``+addr`` marks a per-address
+client; listeners show ``class=0`` and ``LISTENER`` in their flags.
+``bpftool link show`` lists the attached classifier's link, and
+``bpftool map dump pinned /sys/fs/bpf/svc_classify/prefixes`` the map
+with its raw keys.
+
+How it works
+============
+
+The classifier
+--------------
+
+.. kernel-doc:: include/linux/sunrpc/svc.h
+ :identifiers: svc_classifier
+
+The callback runs once per accepted transport, in process context,
+under ``rcu_read_lock()``; sleepable programs are refused at load.
+Only the base BPF helpers are available (map lookups, the usual). The
+program is handed the ``struct svc_xprt`` and may read its fields
+directly:
+
+- ``xpt_remote`` and ``xpt_remotelen``: the peer address;
+- ``xpt_local``: the address the connection arrived on;
+- ``xpt_net``: the network namespace;
+- ``xpt_server->sv_name``: which service, ``"nfsd"``, ``"lockd"``,
+ ``"NFSv4 callback"``;
+- ``xpt_class->xcl_name``: the transport class, ``"tcp"``, ``"rdma"``.
+
+A classifier applies to every service in its namespace. A program that
+wants to treat services differently reads ``xpt_server->sv_name``.
+
+The program ``svc-classify`` carries,
+``tools/net/sunrpc/svc-classify/svc_classify.bpf.c``, is the prefix
+classifier from the BPF selftests
+(``tools/testing/selftests/bpf/progs/bpf_svc_classifier.c``). Its one
+map is an LPM trie keyed by ``{prefixlen, family, addr[16]}`` with
+``prefixlen`` counting the family byte plus the address bits, and the
+class word as the value; a miss returns ``0``. Nothing else in the
+program is specific to this policy. A classifier keyed on the local
+address, the service name, or anything else the ``svc_xprt`` shows is
+the same program with a different lookup, and ``svc-classify replace``
+installs a rebuilt one as long as its map definition is unchanged.
+
+Attaching
+---------
+
+The classifier is a ``struct_ops`` map. Creating its link attaches the
+classifier to the network namespace of the task that creates the link:
+
+- one classifier per namespace; a second attach fails with ``-EBUSY``;
+- ``BPF_LINK_UPDATE`` on the link replaces the program;
+- closing or detaching the link, or unpinning its last reference,
+ removes the classifier;
+- a namespace that exits leaves its link attached to nothing;
+ closing it is harmless.
+
+``svc-classify load`` creates the link with
+``bpf_map__attach_struct_ops()`` and pins it, with the map, under
+``/sys/fs/bpf/svc_classify`` (``-p DIR`` chooses another directory,
+for a second namespace); ``replace`` is ``bpf_link__update_map()`` on
+the pinned link; ``unload`` unpins both.
+
+``bpftool struct_ops`` (``register``, ``dump``, ``unregister``) resolves
+the struct_ops type in the kernel's own BTF only, so when ``sunrpc`` is
+a module those subcommands do not find ``svc_classifier`` maps, and
+``register`` fails after creating the link. Use ``svc-classify``.
+
+Limits
+======
+
+- A transport cannot move between clients; a client identity that is
+ only known after the connection is accepted (an NFSv4.1 client id,
+ say) cannot be used. NFSv4.1 sessions over ``nconnect`` are one
+ client by peer address.
+- RDMA transports are classified like TCP, by the peer address of the
+ connection; UDP is not classified.
+- The cost with no classifier attached is a static branch on enqueue
+ and one empty-queue test on dequeue. With one attached, dispatch
+ takes two or three more lock-free queue operations per request than
+ before: the client's queue, the pool's queue of clients, and a
+ requeue when the client has more.
--
2.53.0
^ permalink raw reply related [flat|nested] 17+ messages in thread
* Re: [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
` (8 preceding siblings ...)
2026-10-07 19:59 ` [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier Benjamin Coddington
@ 2026-10-08 15:19 ` Chuck Lever
2026-10-08 15:28 ` Benjamin Coddington
2026-10-08 18:48 ` Jeff Layton
10 siblings, 1 reply; 17+ messages in thread
From: Chuck Lever @ 2026-10-08 15:19 UTC (permalink / raw)
To: Benjamin Coddington, Jeff Layton, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
On Wed, Oct 7, 2026, at 3:59 PM, Benjamin Coddington wrote:
> This is v2 of [3], with the change I floated in that thread: the kernel
> no longer decides which transports belong together. By default nothing
> changes. An administrator who wants peers grouped loads a small BPF
> program that returns a class for each accepted connection, and the pool
> takes turns across classes.
As a quick response to this posting: I'd like to hear opinions from others
on this approach rather than just fetching it straight into nfsd-testing.
Having patches to mull over is great. Thanks Ben!
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF
2026-10-08 15:19 ` [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Chuck Lever
@ 2026-10-08 15:28 ` Benjamin Coddington
0 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-08 15:28 UTC (permalink / raw)
To: Chuck Lever
Cc: Benjamin Coddington, Jeff Layton, NeilBrown, linux-nfs,
Daire Byrne, bpf, Martin KaFai Lau
On 8 Oct 2026, at 11:19, Chuck Lever wrote:
> On Wed, Oct 7, 2026, at 3:59 PM, Benjamin Coddington wrote:
>> This is v2 of [3], with the change I floated in that thread: the kernel
>> no longer decides which transports belong together. By default nothing
>> changes. An administrator who wants peers grouped loads a small BPF
>> program that returns a class for each accepted connection, and the pool
>> takes turns across classes.
>
> As a quick response to this posting: I'd like to hear opinions from others
> on this approach rather than just fetching it straight into nfsd-testing.
>
> Having patches to mull over is great. Thanks Ben!
No problem, thanks for all the feedback and discussion so far. One point of
trouble on the bpf-list side is their CI bot can't apply these patches
because there are conflicts from the nfsd-testing base. It might be that
the bpf folks won't even bother to review if their CI can't build it..
If this goes further, I'll rebase a next version onto something the bpf-list
tooling can handle, then back to nfsd-testing if it looks like its mergeable.
Ben
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader
2026-10-07 19:59 ` [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader Benjamin Coddington
@ 2026-10-08 18:36 ` Jeff Layton
2026-10-08 19:38 ` Benjamin Coddington
0 siblings, 1 reply; 17+ messages in thread
From: Jeff Layton @ 2026-10-08 18:36 UTC (permalink / raw)
To: Benjamin Coddington, Chuck Lever, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
On Wed, 2026-10-07 at 15:59 -0400, Benjamin Coddington wrote:
> From: Benjamin Coddington <bcodding@hammerspace.com>
>
> An administrator who wants RPC service transports dispatched per client
> needs a classifier attached in the service's network namespace and a
> way to fill its map. Add svc-classify: the prefix classifier from the
> BPF selftests, with its own type definitions in place of vmlinux.h,
> embedded in a loader as a skeleton so the binary stands alone. load
> attaches it and pins the link and the map, replace swaps the program
> and keeps the map, unload detaches; add, del and list manage the map
> with address prefixes and class words written as N, N+addr or addr;
> status shows the link. The Makefile builds the in-tree libbpf and
> bpftool it needs.
>
> Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
> ---
> tools/net/sunrpc/svc-classify/.gitignore | 1 +
> tools/net/sunrpc/svc-classify/Makefile | 65 ++++
> tools/net/sunrpc/svc-classify/svc-classify.c | 357 ++++++++++++++++++
> .../sunrpc/svc-classify/svc_classify.bpf.c | 83 ++++
> 4 files changed, 506 insertions(+)
> create mode 100644 tools/net/sunrpc/svc-classify/.gitignore
> create mode 100644 tools/net/sunrpc/svc-classify/Makefile
> create mode 100644 tools/net/sunrpc/svc-classify/svc-classify.c
> create mode 100644 tools/net/sunrpc/svc-classify/svc_classify.bpf.c
>
> diff --git a/tools/net/sunrpc/svc-classify/.gitignore b/tools/net/sunrpc/svc-classify/.gitignore
> new file mode 100644
> index 000000000000..567609b1234a
> --- /dev/null
> +++ b/tools/net/sunrpc/svc-classify/.gitignore
> @@ -0,0 +1 @@
> +build/
> diff --git a/tools/net/sunrpc/svc-classify/Makefile b/tools/net/sunrpc/svc-classify/Makefile
> new file mode 100644
> index 000000000000..9d86bc88417b
> --- /dev/null
> +++ b/tools/net/sunrpc/svc-classify/Makefile
> @@ -0,0 +1,65 @@
> +# SPDX-License-Identifier: GPL-2.0
> +OUTPUT ?= $(CURDIR)/build/
> +override OUTPUT := $(abspath $(OUTPUT))/
> +$(shell mkdir -p $(OUTPUT))
> +
> +include ../../../build/Build.include
> +include ../../../scripts/Makefile.arch
> +include ../../../scripts/Makefile.include
> +
> +TOOLSDIR := $(abspath ../../..)
> +BPFDIR := $(TOOLSDIR)/lib/bpf
> +BPFTOOLDIR := $(TOOLSDIR)/bpf/bpftool
> +APIDIR := $(TOOLSDIR)/include/uapi
> +INCLUDE_DIR := $(OUTPUT)include
> +BPFOBJ := $(OUTPUT)libbpf/libbpf.a
> +DEFAULT_BPFTOOL := $(OUTPUT)sbin/bpftool
> +BPFTOOL ?= $(DEFAULT_BPFTOOL)
> +CLANG ?= clang
> +msg = $(if $(Q),@printf ' %-8s %s\n' "$(1)" "$(3)";)
> +
> +prefix ?= /usr/local
> +sbindir ?= $(prefix)/sbin
> +INSTALL ?= install
> +
> +CFLAGS += -g -O2 -Wall -I$(INCLUDE_DIR) -I$(OUTPUT)
> +LDLIBS += -lelf -lz
> +BPF_CFLAGS := -g -O2 -target bpf -D__TARGET_ARCH_$(SRCARCH) \
> + -I$(INCLUDE_DIR) -I$(APIDIR) -Wall
> +
> +all: $(OUTPUT)svc-classify
> +
> +$(OUTPUT)libbpf $(OUTPUT)bpftool:
> + $(Q)mkdir -p $@
> +
> +$(BPFOBJ): $(wildcard $(BPFDIR)/*.[ch] $(BPFDIR)/Makefile) | $(OUTPUT)libbpf
> + $(Q)$(MAKE) $(submake_extras) -C $(BPFDIR) OUTPUT=$(OUTPUT)libbpf/ \
> + DESTDIR=$(OUTPUT) prefix= all install_headers
> +
> +$(DEFAULT_BPFTOOL): $(wildcard $(BPFTOOLDIR)/*.[ch] $(BPFTOOLDIR)/Makefile) \
> + $(BPFOBJ) | $(OUTPUT)bpftool
> + $(Q)$(MAKE) $(submake_extras) -C $(BPFTOOLDIR) OUTPUT=$(OUTPUT)bpftool/ \
> + LIBBPF_OUTPUT=$(OUTPUT)libbpf/ LIBBPF_DESTDIR=$(OUTPUT) \
> + prefix= DESTDIR=$(OUTPUT) install-bin
> +
> +$(OUTPUT)svc_classify.bpf.o: svc_classify.bpf.c $(BPFOBJ)
> + $(call msg,CLNG-BPF,,$(notdir $@))
> + $(Q)$(CLANG) $(BPF_CFLAGS) -c $< -o $@
> +
> +$(OUTPUT)svc_classify.skel.h: $(OUTPUT)svc_classify.bpf.o $(BPFTOOL)
> + $(call msg,GEN-SKEL,,$(notdir $@))
> + $(Q)$(BPFTOOL) gen skeleton $< name svc_classify > $@
> +
> +$(OUTPUT)svc-classify: svc-classify.c $(OUTPUT)svc_classify.skel.h $(BPFOBJ)
> + $(call msg,CC,,$(notdir $@))
> + $(Q)$(CC) $(CFLAGS) -o $@ $< $(BPFOBJ) $(LDLIBS)
> +
> +install: $(OUTPUT)svc-classify
> + $(Q)$(INSTALL) -d $(DESTDIR)$(sbindir)
> + $(Q)$(INSTALL) -m 755 $(OUTPUT)svc-classify $(DESTDIR)$(sbindir)/
> +
> +clean:
> + $(Q)rm -rf $(OUTPUT)
> +
> +.PHONY: all install clean
> +.DELETE_ON_ERROR:
> diff --git a/tools/net/sunrpc/svc-classify/svc-classify.c b/tools/net/sunrpc/svc-classify/svc-classify.c
> new file mode 100644
> index 000000000000..f722b891be2c
> --- /dev/null
> +++ b/tools/net/sunrpc/svc-classify/svc-classify.c
> @@ -0,0 +1,357 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * svc-classify - attach the prefix classifier to the RPC services of this
> + * network namespace and manage its prefix map.
> + *
> + * svc-classify [-p DIR] load attach; pin the link and the map
> + * svc-classify [-p DIR] replace replace the program, keep the map
> + * svc-classify [-p DIR] unload detach
> + * svc-classify [-p DIR] add PREFIX CLASS
> + * svc-classify [-p DIR] del PREFIX
> + * svc-classify [-p DIR] list
> + * svc-classify [-p DIR] status
> + *
> + * PREFIX is a.b.c.d[/len], x:y::z[/len], any, any4 or any6. CLASS is N
> + * (one client for every peer that matches), N+addr (one client per peer
> + * address within class N), addr (the same as 0+addr) or 0 (the anonymous
> + * client). DIR, where the link and the map are pinned, defaults to
> + * /sys/fs/bpf/svc_classify.
> + */
> +#include <arpa/inet.h>
> +#include <errno.h>
> +#include <limits.h>
> +#include <stdio.h>
> +#include <stdlib.h>
> +#include <string.h>
> +#include <sys/stat.h>
> +#include <unistd.h>
> +
> +#include <bpf/bpf.h>
> +#include <bpf/libbpf.h>
> +
> +#include "svc_classify.skel.h"
> +
> +#define PIN_DIR_DEFAULT "/sys/fs/bpf/svc_classify"
> +#define SVC_CLASS_PER_ADDR (1U << 31)
> +
> +struct prefix_key {
> + __u32 prefixlen;
> + __u8 family;
> + __u8 addr[16];
> +};
> +
> +static const char *progname;
> +static char pin_link[PATH_MAX], pin_map[PATH_MAX];
> +static const char *pin_dir = PIN_DIR_DEFAULT;
> +
> +static void usage(void)
> +{
> + fprintf(stderr,
> + "usage: %s [-p DIR] load | replace | unload | list | status |\n"
> + " add PREFIX CLASS | del PREFIX\n"
> + "PREFIX: a.b.c.d[/len], x:y::z[/len], any, any4, any6\n"
> + "CLASS: N, N+addr, addr, 0\n", progname);
> + exit(2);
> +}
> +
> +static int do_load(void)
> +{
> + struct svc_classify *skel;
> + struct bpf_link *link;
> + int err;
> +
> + if (access(pin_link, F_OK) == 0) {
> + fprintf(stderr, "%s: already loaded (%s exists)\n", progname,
> + pin_link);
> + return 1;
> + }
> + mkdir(pin_dir, 0700);
> +
> + skel = svc_classify__open_and_load();
> + if (!skel) {
> + fprintf(stderr, "%s: load: %s\n", progname, strerror(errno));
> + rmdir(pin_dir);
> + return 1;
> + }
> + link = bpf_map__attach_struct_ops(skel->maps.prefix);
> + if (!link) {
> + fprintf(stderr, "%s: attach: %s\n", progname, strerror(errno));
> + rmdir(pin_dir);
> + return 1;
> + }
> + err = bpf_link__pin(link, pin_link);
> + if (!err)
> + err = bpf_map__pin(skel->maps.prefixes, pin_map);
> + if (err) {
> + fprintf(stderr, "%s: pin: %s\n", progname, strerror(-err));
> + bpf_link__destroy(link);
> + unlink(pin_link);
> + rmdir(pin_dir);
> + return 1;
> + }
> + return 0;
> +}
> +
> +static int do_replace(void)
> +{
> + struct svc_classify *skel;
> + struct bpf_link *link;
> + int link_fd, err;
> +
> + link_fd = bpf_obj_get(pin_link);
> + if (link_fd < 0) {
> + fprintf(stderr, "%s: not loaded (%s)\n", progname,
> + strerror(errno));
> + return 1;
> + }
> + skel = svc_classify__open();
> + if (!skel) {
> + fprintf(stderr, "%s: open: %s\n", progname, strerror(errno));
> + return 1;
> + }
> + /*
> + * Reuse the pinned map so the entries survive. A build whose map
> + * definition differs fails here; unload and load then.
> + */
> + err = bpf_map__set_pin_path(skel->maps.prefixes, pin_map);
> + if (!err)
> + err = svc_classify__load(skel);
> + if (err) {
> + fprintf(stderr, "%s: load: %s\n", progname, strerror(-err));
> + return 1;
> + }
> + /*
> + * bpf_link__update_map() wants the link that attached the map, not
> + * one reopened from a pin. bpf_map__attach_struct_ops() writes the
> + * program into the new map's value before it tries to create a link,
> + * and the link is refused with EBUSY while the old one is attached;
> + * the written map is what the link update needs.
> + */
> + link = bpf_map__attach_struct_ops(skel->maps.prefix);
> + if (link) {
> + bpf_link__destroy(link);
> + fprintf(stderr, "%s: nothing attached in this namespace; unload and load\n",
> + progname);
> + return 1;
> + }
> + if (errno != EBUSY) {
> + fprintf(stderr, "%s: attach: %s\n", progname, strerror(errno));
> + return 1;
> + }
> + err = bpf_link_update(link_fd, bpf_map__fd(skel->maps.prefix), NULL);
> + if (err) {
> + fprintf(stderr, "%s: replace: %s\n", progname, strerror(errno));
> + return 1;
> + }
> + return 0;
> +}
> +
> +static int do_unload(void)
> +{
> + int ret = 0;
> +
> + if (unlink(pin_link) && errno != ENOENT) {
> + perror(pin_link);
> + ret = 1;
> + }
> + if (unlink(pin_map) && errno != ENOENT) {
> + perror(pin_map);
> + ret = 1;
> + }
> + rmdir(pin_dir);
> + return ret;
> +}
> +
> +static int parse_prefix(const char *s, struct prefix_key *key)
> +{
> + char buf[64], *slash;
> + int bits, max;
> +
> + memset(key, 0, sizeof(*key));
> + if (!strcmp(s, "any"))
> + return 0;
> + if (!strcmp(s, "any4")) {
> + key->family = AF_INET;
> + key->prefixlen = 8;
> + return 0;
> + }
> + if (!strcmp(s, "any6")) {
> + key->family = AF_INET6;
> + key->prefixlen = 8;
> + return 0;
> + }
> + if (strlen(s) >= sizeof(buf))
> + return -1;
> + strcpy(buf, s);
> + slash = strchr(buf, '/');
> + if (slash)
> + *slash++ = '\0';
> + if (inet_pton(AF_INET, buf, key->addr) == 1) {
> + key->family = AF_INET;
> + max = 32;
> + } else if (inet_pton(AF_INET6, buf, key->addr) == 1) {
> + key->family = AF_INET6;
> + max = 128;
> + } else {
> + return -1;
> + }
> + bits = max;
> + if (slash) {
> + char *end;
> +
> + if (*slash < '0' || *slash > '9')
> + return -1;
> + bits = strtoul(slash, &end, 10);
> + if (*end || bits > max)
> + return -1;
> + }
> + key->prefixlen = 8 + bits;
> + return 0;
> +}
> +
> +static int parse_class(const char *s, __u32 *class)
> +{
> + unsigned long n;
> + char *end;
> +
> + if (!strcmp(s, "addr")) {
> + *class = SVC_CLASS_PER_ADDR;
> + return 0;
> + }
> + n = strtoul(s, &end, 0);
> + if (end == s || n >= SVC_CLASS_PER_ADDR)
> + return -1;
> + *class = n;
> + if (!strcmp(end, "+addr"))
> + *class |= SVC_CLASS_PER_ADDR;
> + else if (*end)
> + return -1;
> + return 0;
> +}
> +
> +static void print_entry(const struct prefix_key *key, __u32 class)
> +{
> + char addr[INET6_ADDRSTRLEN];
> +
> + if (key->prefixlen == 0)
> + printf("any");
> + else if (key->prefixlen == 8)
> + printf("any%d", key->family == AF_INET ? 4 : 6);
> + else
> + printf("%s/%u", inet_ntop(key->family, key->addr, addr,
> + sizeof(addr)), key->prefixlen - 8);
> + printf(" -> %u%s\n", class & ~SVC_CLASS_PER_ADDR,
> + class & SVC_CLASS_PER_ADDR ? "+addr" : "");
> +}
> +
> +static int open_map(void)
> +{
> + int fd = bpf_obj_get(pin_map);
> +
> + if (fd < 0) {
> + fprintf(stderr, "%s: not loaded (%s: %s)\n", progname, pin_map,
> + strerror(errno));
> + exit(1);
> + }
> + return fd;
> +}
> +
> +static int do_add(const char *prefix, const char *cls)
> +{
> + struct prefix_key key;
> + __u32 class;
> +
> + if (parse_prefix(prefix, &key) || parse_class(cls, &class))
> + usage();
> + if (bpf_map_update_elem(open_map(), &key, &class, BPF_ANY)) {
> + fprintf(stderr, "%s: add: %s\n", progname, strerror(errno));
> + return 1;
> + }
> + return 0;
> +}
> +
> +static int do_del(const char *prefix)
> +{
> + struct prefix_key key;
> +
> + if (parse_prefix(prefix, &key))
> + usage();
> + if (bpf_map_delete_elem(open_map(), &key)) {
> + fprintf(stderr, "%s: del: %s\n", progname, strerror(errno));
> + return 1;
> + }
> + return 0;
> +}
> +
> +static int do_list(void)
> +{
> + struct prefix_key key, next;
> + __u32 class;
> + int fd;
> +
> + fd = open_map();
> + if (bpf_map_get_next_key(fd, NULL, &next))
> + return 0;
> + do {
> + key = next;
> + if (!bpf_map_lookup_elem(fd, &key, &class))
> + print_entry(&key, class);
> + } while (!bpf_map_get_next_key(fd, &key, &next));
> + return 0;
> +}
> +
> +static int do_status(void)
> +{
> + struct bpf_link_info info = {};
> + __u32 len = sizeof(info);
> + int fd;
> +
> + fd = bpf_obj_get(pin_link);
> + if (fd < 0) {
> + printf("not loaded\n");
> + return 1;
> + }
> + if (bpf_link_get_info_by_fd(fd, &info, &len))
> + printf("loaded\n");
> + else
> + printf("loaded, link id %u, struct_ops map id %u\n", info.id,
> + info.struct_ops.map_id);
> + return 0;
> +}
> +
> +int main(int argc, char **argv)
> +{
> + const char *cmd;
> + int opt;
> +
> + progname = argv[0];
> + while ((opt = getopt(argc, argv, "p:")) != -1) {
> + if (opt != 'p')
> + usage();
> + pin_dir = optarg;
> + }
> + argc -= optind;
> + argv += optind;
> + if (argc < 1)
> + usage();
> + cmd = argv[0];
> + snprintf(pin_link, sizeof(pin_link), "%s/link", pin_dir);
> + snprintf(pin_map, sizeof(pin_map), "%s/prefixes", pin_dir);
> +
> + if (!strcmp(cmd, "load") && argc == 1)
> + return do_load();
> + if (!strcmp(cmd, "replace") && argc == 1)
> + return do_replace();
> + if (!strcmp(cmd, "unload") && argc == 1)
> + return do_unload();
> + if (!strcmp(cmd, "add") && argc == 3)
> + return do_add(argv[1], argv[2]);
> + if (!strcmp(cmd, "del") && argc == 2)
> + return do_del(argv[1]);
> + if (!strcmp(cmd, "list") && argc == 1)
> + return do_list();
> + if (!strcmp(cmd, "status") && argc == 1)
> + return do_status();
> + usage();
> + return 2;
> +}
> diff --git a/tools/net/sunrpc/svc-classify/svc_classify.bpf.c b/tools/net/sunrpc/svc-classify/svc_classify.bpf.c
> new file mode 100644
> index 000000000000..3d012f102ec3
> --- /dev/null
> +++ b/tools/net/sunrpc/svc-classify/svc_classify.bpf.c
> @@ -0,0 +1,83 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * A svc_classifier for RPC service transports: the peer address is looked
> + * up in an LPM trie keyed by {family, address}; a hit returns the class
> + * word stored there, a miss returns 0, the anonymous client. A key of
> + * prefixlen 0 is the default for every family, 8 the default for one.
> + */
> +#include <linux/types.h>
> +#include <linux/bpf.h>
> +#include <bpf/bpf_helpers.h>
> +#include <bpf/bpf_tracing.h>
> +#include <bpf/bpf_core_read.h>
> +
> +char _license[] SEC("license") = "GPL";
> +
> +#define AF_INET 2
> +#define AF_INET6 10
> +
> +#define SVC_CLASSIFIER_NAME_MAX 16
> +
> +struct sockaddr_storage___local {
> + unsigned short ss_family;
> + char __data[126];
> +};
> +
> +struct svc_xprt___local {
> + struct sockaddr_storage___local xpt_remote;
> +} __attribute__((preserve_access_index));
> +
> +struct svc_classifier___local {
> + char name[SVC_CLASSIFIER_NAME_MAX];
> + __u32 (*classify)(const struct svc_xprt___local *xprt);
> +};
> +
> +struct prefix_key {
> + __u32 prefixlen; /* 8 (the family byte) + address bits */
> + __u8 family;
> + __u8 addr[16];
> +};
> +
> +struct {
> + __uint(type, BPF_MAP_TYPE_LPM_TRIE);
> + __type(key, struct prefix_key);
> + __type(value, __u32);
> + __uint(max_entries, 1024);
> + __uint(map_flags, BPF_F_NO_PREALLOC);
> +} prefixes SEC(".maps");
> +
> +SEC("struct_ops/classify")
> +__u32 BPF_PROG(prefix_classify, const struct svc_xprt___local *xprt)
> +{
> + struct sockaddr_storage___local ss;
> + struct prefix_key key = {};
> + __u32 *class;
> +
> + if (bpf_core_read(&ss, sizeof(ss), &xprt->xpt_remote))
> + return 0;
> +
> + key.family = ss.ss_family;
> + switch (ss.ss_family) {
> + case AF_INET:
> + /* struct sockaddr_in: port at 2, address at 4 */
> + __builtin_memcpy(key.addr, &ss.__data[2], 4);
> + key.prefixlen = 8 + 32;
> + break;
> + case AF_INET6:
> + /* struct sockaddr_in6: port at 2, flowinfo at 4, address at 8 */
> + __builtin_memcpy(key.addr, &ss.__data[6], 16);
> + key.prefixlen = 8 + 128;
> + break;
> + default:
> + return 0;
> + }
> +
> + class = bpf_map_lookup_elem(&prefixes, &key);
> + return class ? *class : 0;
> +}
> +
> +SEC(".struct_ops.link")
> +struct svc_classifier___local prefix = {
> + .name = "prefix",
> + .classify = (void *)prefix_classify,
> +};
The userland programs in the kernel are often not packaged, so it seems
to me that svc-classify would be better integrated into nfs-utils. That
would allow you to more easily bundle it with the systemd service file.
It might even make sense to integrate this functionality into nfsdctl,
but I guess the classifier affects all rpc_serv objects as well (e.g.
NLM), so it may be best to keep it as its own thing.
--
Jeff Layton <jlayton@kernel.org>
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier
2026-10-07 19:59 ` [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier Benjamin Coddington
@ 2026-10-08 18:42 ` Jeff Layton
2026-10-08 19:40 ` Benjamin Coddington
0 siblings, 1 reply; 17+ messages in thread
From: Jeff Layton @ 2026-10-08 18:42 UTC (permalink / raw)
To: Benjamin Coddington, Chuck Lever, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
On Wed, 2026-10-07 at 15:59 -0400, Benjamin Coddington wrote:
> From: Benjamin Coddington <bcodding@hammerspace.com>
>
> Dispatch round-robin across clients changes how a service shares its
> threads, and the classes come from a BPF program the administrator
> loads, so the administrator needs the model, the class word, and the
> procedure. Add a page with those: install a classifier with
> svc-classify (build, attach, fill the map, check, change, remove, keep
> across reboots), six class maps and the share each gives its clients,
> the tracepoint that shows the class, and how the hook, the program and
> the attach model work underneath.
>
> Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
> ---
> Documentation/filesystems/nfs/index.rst | 1 +
> .../filesystems/nfs/rpc-server-clients.rst | 394 ++++++++++++++++++
> 2 files changed, 395 insertions(+)
> create mode 100644 Documentation/filesystems/nfs/rpc-server-clients.rst
>
> diff --git a/Documentation/filesystems/nfs/index.rst b/Documentation/filesystems/nfs/index.rst
> index a29a212b5b4d..57a61ce0533d 100644
> --- a/Documentation/filesystems/nfs/index.rst
> +++ b/Documentation/filesystems/nfs/index.rst
> @@ -16,3 +16,4 @@ NFS
> nfsd-io-modes
> knfsd-stats
> reexport
> + rpc-server-clients
> diff --git a/Documentation/filesystems/nfs/rpc-server-clients.rst b/Documentation/filesystems/nfs/rpc-server-clients.rst
> new file mode 100644
> index 000000000000..b08351a88314
> --- /dev/null
> +++ b/Documentation/filesystems/nfs/rpc-server-clients.rst
> @@ -0,0 +1,394 @@
> +.. SPDX-License-Identifier: GPL-2.0
> +
> +======================================================
> +RPC server: per-client dispatch and transport classes
> +======================================================
> +
> +An RPC service thread pool serves its ready transports in FIFO order,
> +one request per turn. A peer with many connections (``nconnect``, a
> +deep NFSv4.1 slot table, a data mover with a thread per file) takes a
> +turn per connection, and a peer with one connection waits behind all
> +of them.
> +
> +With a classifier attached, the pool instead serves its *clients* round
> +robin: each client with work queued gets one request per round, however
> +many connections it holds. Which transports form a client is decided
> +by a BPF program the administrator loads, once per accepted connection.
> +With no program loaded nothing changes.
> +
> +This document describes the model, how to install a classifier with
> +the ``svc-classify`` tool, what some class maps do, and how it works
> +underneath.
> +
> +Clients, classes and turns
> +==========================
> +
> +A *client* is a set of transports that share one turn. Every service
> +(nfsd, lockd, the NFSv4 callback service) keeps its own clients, found
> +by *class* and network namespace, or by class, namespace and peer
> +address.
> +
> +The classifier returns a 32-bit class word for each accepted transport:
> +
> +``0``
> + ``SVC_CLASS_NONE``: the transport stays on the service's *anonymous
> + client*.
> +
> +``N``, 1 to 2^31 - 1
> + one client for every transport of class ``N``.
> +
> +``N | SVC_CLASS_PER_ADDR``
> + one client per peer address within class ``N``. ``N`` may be 0, so
> + ``SVC_CLASS_PER_ADDR`` alone means one client per peer address.
> +
> +``SVC_CLASS_PER_ADDR`` is bit 31. ``svc-classify`` and this document
> +write a class word with the bit set as ``N+addr``; ``svc-classify``
> +also accepts ``addr`` for ``0+addr``.
> +
> +Dispatch rules:
> +
> +- A pool serves the clients that have transports queued round robin,
> + one request per client per round.
> +- Within a client, transports are served in FIFO order, except that a
> + transport needing a connection accepted, closed, or a TLS handshake
> + run goes ahead of transports with data, though not two such turns in
> + a row while data waits.
> +- A client's share of the pool does not grow with its connection count.
> + A client with one connection and a client with thirty get the same
> + number of turns while both have work queued.
> +
> +Things that follow from the rules and are easy to get wrong:
> +
> +- The anonymous client is one client. Once any classifier is attached,
> + every transport whose class is ``0`` shares a single turn with every
> + other such transport of that service. Within that one client the old
> + FIFO order applies, so those peers share the turn in proportion to
> + their connection counts. A map that classifies only the hosts it
> + cares about and leaves the rest at ``0`` gives "the rest" one turn
> + in total (see example 5 below).
> +- Per-address clients ignore the port. Only IPv4 and IPv6 peers can be
> + keyed by address; a transport whose peer is anything else is left on
> + the anonymous client.
> +- Listeners and UDP sockets are never classified; UDP traffic is served
> + from the anonymous client.
> +- A transport keeps the class it was given when accepted. Changing the
> + map, or replacing or removing the classifier, affects connections
> + accepted afterwards. A connection that must be reclassified has to
> + reconnect.
> +- nfsd has one service per network namespace, so its anonymous client
> + is per namespace. lockd and the NFSv4 callback service are one
> + service for all namespaces; their anonymous clients span namespaces.
> +- With no classifier attached anywhere, the pool uses a single FIFO;
> + the per-request cost of the feature is then a static branch on
> + enqueue and one empty-queue test on dequeue.
> + ``CONFIG_SUNRPC_BPF_CLASSIFY`` (default y) builds the hook; without
> + it there is no classification.
> +
> +Installing a classifier
> +=======================
> +
> +There is one tool to do it with: ``svc-classify``, in
> +``tools/net/sunrpc/svc-classify`` of the kernel source. It carries the
> +classifier program inside the binary, attaches it, and manages the
> +class map by address prefix. ``bpftool`` is not needed, and its
> +``struct_ops`` subcommands cannot be used in its place (see "How it
> +works").
> +
"There is one tool to do it with: ..."
The obviously LLM-written documentation gives me the Ick and I stopped
reading. There is also a lot of info in this doc that is not
particularly helpful, like "the things that follow from the rules and
are easy to get wrong".
If you want humans to actually read this, then it probably needs to be
written by a human. If you think the LLM slop is actually useful for
other LLMs, then maybe put those bits at the end in a clearly-
delineated section.
> +What you need
> +-------------
> +
> +- A kernel with ``CONFIG_SUNRPC_BPF_CLASSIFY`` (default y) and BTF:
> + ``CONFIG_DEBUG_INFO_BTF``, and ``CONFIG_DEBUG_INFO_BTF_MODULES`` when
> + ``sunrpc`` is a module.
> +- ``/sys/fs/bpf`` mounted (systemd mounts it).
> +- To build the tool: ``clang``, and the ``libelf`` and ``zlib``
> + development files. The tool builds the libbpf and bpftool it needs
> + from the kernel source tree.
> +
> +Build and install the tool
> +--------------------------
> +
> +In the kernel source tree::
> +
> + $ make -C tools/net/sunrpc/svc-classify
> + # make -C tools/net/sunrpc/svc-classify install
> +
> +The binary goes to ``/usr/local/sbin/svc-classify`` (``prefix=`` and
> +``sbindir=`` override that). At run time it needs only libelf and
> +zlib.
> +
> +Attach it
> +---------
> +
> +As root, in the network namespace the service runs in (on a host that
> +is the initial namespace; in a container, ``nsenter`` or
> +``ip netns exec`` into it first)::
> +
> + # svc-classify load
> + # svc-classify status
> + loaded, link id 46, struct_ops map id 140
> +
> +The classifier is now attached to every RPC service in the namespace,
> +with an empty map: every connection accepted from now on returns class
> +``0`` and stays on the anonymous client, so nothing has changed yet.
> +One classifier per namespace: a second ``load`` is refused.
> +
> +Install the class map
> +---------------------
> +
> +Add one line per address prefix. The longest matching prefix wins; a
> +peer that matches nothing gets class ``0``::
> +
> + # svc-classify add 192.0.2.0/24 1
> + # svc-classify add any addr
> + # svc-classify list
> + 192.0.2.0/24 -> 1
> + any -> 0+addr
> +
> +``PREFIX`` is ``a.b.c.d[/len]``, ``x:y::z[/len]``, ``any`` (every
> +address of either family), ``any4`` or ``any6``. ``CLASS`` is ``N``
> +(one client for every peer that matches), ``N+addr`` (one client per
> +peer address, within class ``N``), ``addr`` (one client per peer
> +address, the same as ``0+addr``), or ``0`` (the anonymous client).
> +
> +Entries apply to connections accepted after they are added; to
> +reclassify existing connections, have the clients reconnect (restarting
> +the service does that for all of them). ``svc-classify del PREFIX``
> +removes an entry.
> +
> +Check it
> +--------
> +
> +``svc-classify status`` and ``svc-classify list`` show the link and the
> +map. The ``sunrpc:svc_xprt_dequeue`` tracepoint shows the class every
> +transport is dispatched as (see "Observing"), which is the check that
> +the map does what was meant.
> +
> +Change or remove it
> +-------------------
> +
> +- ``svc-classify add`` and ``del`` change the map at any time.
> +- ``svc-classify replace`` installs a new build of the tool's program
> + on the attached link and keeps the map, provided the new build's map
> + definition is unchanged; otherwise unload and load.
> +- ``svc-classify unload`` detaches the classifier. Connections accepted
> + afterwards are anonymous; when no classifier is attached in any
> + namespace, the service is back on its single FIFO.
> +
> +Across reboots
> +--------------
> +
> +Nothing persists: the link and the map live in ``/sys/fs/bpf`` and are
> +gone at boot. Run the ``load`` and ``add`` commands before the service
> +starts, for example from a unit ordered before ``nfs-server.service``::
> +
> + [Unit]
> + Description=RPC service transport classifier
> + Before=nfs-server.service
> +
> + [Service]
> + Type=oneshot
> + RemainAfterExit=yes
> + ExecStart=/usr/local/sbin/svc-classify load
> + ExecStart=/usr/local/sbin/svc-classify add 192.0.2.0/24 1
> + ExecStart=/usr/local/sbin/svc-classify add any addr
> + ExecStop=/usr/local/sbin/svc-classify unload
> +
> + [Install]
> + WantedBy=nfs-server.service
> +
> +Loading after the service is up works too; it only misses the
> +connections already accepted.
> +
> +Examples
> +========
> +
> +Each example gives the ``svc-classify add`` lines, the hosts they are
> +applied to, and the share of the pool each client gets while all of
> +them have work queued.
> +
> +1. Every host its own client
> +----------------------------
> +
> +::
> +
> + # svc-classify add any addr
> +
> +Three hosts: one with eight connections, one with one, one with two.
> +Each host is a client, and each gets a third of the turns. Without a
> +classifier they would be served in proportion to their connections,
> +8:1:2.
> +
> +2. A set of movers as one client, everyone else per host
> +--------------------------------------------------------
> +
> +::
> +
> + # svc-classify add 192.0.2.0/24 1
> + # svc-classify add any addr
> +
> +Three movers in ``192.0.2.0/24`` with four connections each, and two
> +other hosts, one with one connection and one with two. The movers are
> +one client; each of the other hosts is a client of its own. Each of
> +the three clients gets a third of the pool; the movers share theirs by
> +connection count, a ninth each. Without a classifier the movers'
> +twelve connections would take twelve fifteenths of the pool and the
> +one-connection host one fifteenth.
> +
> +3. Two mover groups, one turn each
> +----------------------------------
> +
> +::
> +
> + # svc-classify add 192.0.2.0/25 1
> + # svc-classify add 192.0.2.128/25 2
> + # svc-classify add any addr
> +
> +Two hosts in ``192.0.2.0/25`` and two in ``192.0.2.128/25``, four
> +connections each, and one other host with one connection. Each group
> +is a client and the other host is a client: three clients, a third
> +each. Within a group the two hosts split the group's third by
> +connection count, a sixth each.
> +
> +4. Per host inside a campus, the rest of the world as one client
> +----------------------------------------------------------------
> +
> +::
> +
> + # svc-classify add 198.51.100.0/24 1+addr
> + # svc-classify add any 2
> +
> +Two hosts in ``198.51.100.0/24`` and two hosts outside it. Each campus
> +host is a client of its own (class 1, one client per address); the
> +outside hosts together are one client (class 2). Three clients, a
> +third each; the outside hosts split their third by connection count.
> +
> +5. Leaving hosts unclassified
> +-----------------------------
> +
> +::
> +
> + # svc-classify add 203.0.113.0/24 0
> + # svc-classify add any addr
> +
> +Two hosts in ``203.0.113.0/24`` and two hosts elsewhere. The two in
> +``203.0.113.0/24`` return ``0`` and so share the anonymous client, one
> +turn between them; the two elsewhere are a client each. Three
> +clients, a third each; the anonymous pair split their third by
> +connection count.
> +
> +The common mistake is the map with only the movers in it::
> +
> + # svc-classify add 192.0.2.0/24 1
> +
> +Three movers in ``192.0.2.0/24`` and two other hosts, one with one
> +connection and one with two. There are exactly two clients: the
> +movers, and everyone else on the anonymous client. The movers get
> +half the pool. The other half goes to the two other hosts in
> +proportion to their connections, a sixth and a third of the pool. Add
> +``any addr`` to serve them per host.
> +
> +6. IPv6
> +-------
> +
> +::
> +
> + # svc-classify add 2001:db8:1::/48 1
> + # svc-classify add any addr
> +
> +Entries are per family; ``any`` covers both families, ``any6`` IPv6
> +only. Two hosts in ``2001:db8:1::/48`` and one host elsewhere: the two
> +are one client, the other is a client of its own, half the pool each.
> +
> +Observing
> +=========
> +
> +The ``sunrpc:svc_xprt_dequeue`` tracepoint reports the class of every
> +transport as it is dispatched::
> +
> + svc_xprt_dequeue: server=127.0.0.1:3049 client=127.0.1.1:38209 xpt_id=1164 flags=BUSY|DATA|TEMP|CACHE_AUTH|LOCAL|CONG_CTRL class=1 wakeup-us=30 qtime-us=4
> + svc_xprt_dequeue: server=127.0.0.1:3049 client=127.0.2.1:45965 xpt_id=1162 flags=BUSY|DATA|TEMP|CACHE_AUTH|LOCAL|CONG_CTRL class=0+addr wakeup-us=60 qtime-us=10
> + svc_xprt_dequeue: server=[::1]:3049 client=[2001:db8:1::1]:44401 xpt_id=1152 flags=BUSY|DATA|TEMP|CACHE_AUTH|LOCAL|CONG_CTRL class=1 wakeup-us=62 qtime-us=7
> + svc_xprt_dequeue: server=[::]:3049 client=(einval) xpt_id=2 flags=BUSY|CONN|CHNGBUF|LISTENER|CACHE_AUTH|CONG_CTRL|RPCB_UNREG class=0 wakeup-us=25 qtime-us=25
> +
> +``class=0`` is the anonymous client; ``+addr`` marks a per-address
> +client; listeners show ``class=0`` and ``LISTENER`` in their flags.
> +``bpftool link show`` lists the attached classifier's link, and
> +``bpftool map dump pinned /sys/fs/bpf/svc_classify/prefixes`` the map
> +with its raw keys.
> +
> +How it works
> +============
> +
> +The classifier
> +--------------
> +
> +.. kernel-doc:: include/linux/sunrpc/svc.h
> + :identifiers: svc_classifier
> +
> +The callback runs once per accepted transport, in process context,
> +under ``rcu_read_lock()``; sleepable programs are refused at load.
> +Only the base BPF helpers are available (map lookups, the usual). The
> +program is handed the ``struct svc_xprt`` and may read its fields
> +directly:
> +
> +- ``xpt_remote`` and ``xpt_remotelen``: the peer address;
> +- ``xpt_local``: the address the connection arrived on;
> +- ``xpt_net``: the network namespace;
> +- ``xpt_server->sv_name``: which service, ``"nfsd"``, ``"lockd"``,
> + ``"NFSv4 callback"``;
> +- ``xpt_class->xcl_name``: the transport class, ``"tcp"``, ``"rdma"``.
> +
> +A classifier applies to every service in its namespace. A program that
> +wants to treat services differently reads ``xpt_server->sv_name``.
> +
> +The program ``svc-classify`` carries,
> +``tools/net/sunrpc/svc-classify/svc_classify.bpf.c``, is the prefix
> +classifier from the BPF selftests
> +(``tools/testing/selftests/bpf/progs/bpf_svc_classifier.c``). Its one
> +map is an LPM trie keyed by ``{prefixlen, family, addr[16]}`` with
> +``prefixlen`` counting the family byte plus the address bits, and the
> +class word as the value; a miss returns ``0``. Nothing else in the
> +program is specific to this policy. A classifier keyed on the local
> +address, the service name, or anything else the ``svc_xprt`` shows is
> +the same program with a different lookup, and ``svc-classify replace``
> +installs a rebuilt one as long as its map definition is unchanged.
> +
> +Attaching
> +---------
> +
> +The classifier is a ``struct_ops`` map. Creating its link attaches the
> +classifier to the network namespace of the task that creates the link:
> +
> +- one classifier per namespace; a second attach fails with ``-EBUSY``;
> +- ``BPF_LINK_UPDATE`` on the link replaces the program;
> +- closing or detaching the link, or unpinning its last reference,
> + removes the classifier;
> +- a namespace that exits leaves its link attached to nothing;
> + closing it is harmless.
> +
> +``svc-classify load`` creates the link with
> +``bpf_map__attach_struct_ops()`` and pins it, with the map, under
> +``/sys/fs/bpf/svc_classify`` (``-p DIR`` chooses another directory,
> +for a second namespace); ``replace`` is ``bpf_link__update_map()`` on
> +the pinned link; ``unload`` unpins both.
> +
> +``bpftool struct_ops`` (``register``, ``dump``, ``unregister``) resolves
> +the struct_ops type in the kernel's own BTF only, so when ``sunrpc`` is
> +a module those subcommands do not find ``svc_classifier`` maps, and
> +``register`` fails after creating the link. Use ``svc-classify``.
> +
> +Limits
> +======
> +
> +- A transport cannot move between clients; a client identity that is
> + only known after the connection is accepted (an NFSv4.1 client id,
> + say) cannot be used. NFSv4.1 sessions over ``nconnect`` are one
> + client by peer address.
> +- RDMA transports are classified like TCP, by the peer address of the
> + connection; UDP is not classified.
> +- The cost with no classifier attached is a static branch on enqueue
> + and one empty-queue test on dequeue. With one attached, dispatch
> + takes two or three more lock-free queue operations per request than
> + before: the client's queue, the pool's queue of clients, and a
> + requeue when the client has more.
--
Jeff Layton <jlayton@kernel.org>
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
` (9 preceding siblings ...)
2026-10-08 15:19 ` [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Chuck Lever
@ 2026-10-08 18:48 ` Jeff Layton
10 siblings, 0 replies; 17+ messages in thread
From: Jeff Layton @ 2026-10-08 18:48 UTC (permalink / raw)
To: Benjamin Coddington, Chuck Lever, NeilBrown
Cc: linux-nfs, Daire Byrne, bpf, Martin KaFai Lau
On Wed, 2026-10-07 at 15:59 -0400, Benjamin Coddington wrote:
> From: Benjamin Coddington <bcodding@hammerspace.com>
>
> This is v2 of [3], with the change I floated in that thread: the kernel
> no longer decides which transports belong together. By default nothing
> changes. An administrator who wants peers grouped loads a small BPF
> program that returns a class for each accepted connection, and the pool
> takes turns across classes.
>
> What's different from v1
> ------------------------
>
> The table I described on [3] became a hook. Instead of a kernel table
> of address prefixes, a netlink interface for it and an nfs-utils side,
> sunrpc registers a struct_ops with one callback: given the accepted
> svc_xprt, return a class. 0 means the anonymous client (shared with
> everything unclassified, today's order among them); any other number is
> a class every transport returning it shares; a flag bit on top means
> "and one client per peer address within that class". So "all movers are
> one class, everybody else per host" is two entries in a prefix map, and
> the posted v1 behaviour is one entry.
>
> The hook runs once per accepted connection, in process context, after
> TCP and RDMA have both set the peer address. Nothing runs per request.
> No program loaded means every transport is on the
> anonymous client, and with patch 6 the pool uses its old flat queue
> until a classifier attaches, so the dispatch path is today's code for
> anyone who hasn't asked for anything. That's the answer to Chuck's cost
> question from [3]: measured on bare metal the two-level queue cost about
> half a microsecond a dispatch (1.7% at 484k IOPS on one client's eight
> connections); with no program it now costs a static branch on enqueue
> and an empty-queue test on dequeue.
>
> Attaching is a struct_ops link, bound to the attaching task's network
> namespace, one per namespace. A second attach gets -EBUSY, a link
> update replaces the program, closing the link removes it, and a
> namespace that exits leaves its link with nothing behind. Transports
> keep the class they got at accept; a new program only affects new
> connections. The reference program and the attach tests are in the BPF
> selftests (patch 7). svc-classify (patch 8) attaches the prefix
> classifier and manages its map by address prefix, so an administrator
> never touches bpftool. The page in patch 9,
> Documentation/filesystems/nfs/rpc-server-clients.rst, has the class
> word, the install procedure, and six class maps with the share each one
> gives its clients.
>
> Still not in here: moving a transport between clients, so a v4.1
> clientid key, which Chuck asked about on [3], is not in this version.
> A multi-homed host or an IPv6 host with a prefix's worth of addresses
> is handled by the prefix map instead.
>
> Also from the [3] review: control events (accept, close, handshake) go
> ahead of data within a client (patch 5), and the dequeue tracepoint
> prints the class (patch 4).
>
> Numbers
> -------
>
> The dispatch numbers are the v1 series' on a two-socket box (2 x Xeon
> 6542Y, 96 threads, no KASAN), from [3]. The dispatch code is the same
> here: v2 puts the classifier in front of it, and the flat queue behind
> a static branch when nothing is attached. Loopback, NFSv3, 16 threads,
> 10 ms injected service time with 50% jitter. A is v7.2, B the series.
> v4.1 is within 2% in every cell.
>
> One interactive client (bursts of 32) against one host with K
> backlogged connections, burst completion p50 in ms, floor 37:
>
> K 4 8 16 32
> A 70 195 351 671
> B 56 62 61 63
>
> N (K=16) 1 8 32 64 96 128
> A 21 98 350 690 1035 1365
> B 11 31 62 103 144 179
>
> M hosts x 4 1 2 4 8
> A 70 192 351 670
> B 57 81 121 192
>
> S: 2x8 vs K, share 4 8 16 32
> A 41.0 20.0 11.1 5.9
> B 50.3 50.0 50.0 50.0
>
> A real NFSv3 client walking 500 files against a 16-connection
> aggressor: 37.8 s on A, 2.5 s on B, 0.12 s alone on both.
>
> Classes. The v1 kernel keyed on source address, so on [3] I emulated
> "all movers are one class" by giving six movers one address. With this
> series that's one map entry (the movers' prefix -> 1, everything else
> per address), and the six-address column is what "everything per
> address" gives. Six movers at 4 or 8 connections each against
> customers, customer share of dispatches:
>
> A per address movers one class
> one customer, 1 conn x 16 4.0 / 2.0 14.3 / 14.3 49.9 / 49.9
> one customer, 4 x 4 14.3 / 7.7 14.3 / 14.3 49.8 / 49.7
> four customers, 4 x 4 each 40 / 25 40 / 40 80 / 80
>
> Burst of 32 against M busy hosts, p50 ms: hosts on their own addresses
> 57 81 121 192, as one class 56 62 62 62. A: 71 193 348 673 either way.
>
> Cost. With a classifier attached, dispatch takes two or three more
> lock-free queue operations per request. On the same box that's about half a
> microsecond when the enqueue and the dequeue run on different cores,
> fio 4k O_DIRECT randread over loopback, 5 x 20s:
>
> A B
> nconnect=1, 16 jobs 152.8k 152.2k -0.4%
> nconnect=8, 16 jobs 483.8k 475.5k -1.7%
> nconnect=8, 64 jobs 584.2k 581.0k -0.5%
>
> The serial walk and 4 jobs on one connection don't move. With no
> classifier attached the pool is on its old flat queue (patch 6). I only
> have my VM for that path so far (KASAN, 10 vCPUs, loopback, the same
> fio), and there the no-program kernel measures the same as v7.2 within
> the run-to-run spread: 16 jobs 94.9k vs 95.5k, 64 jobs 138.6k vs 139.0k.
> So the static branch adds nothing I can see, but that's a VM number.
>
> Not tested: RDMA. The hook sits in the accept path TCP and RDMA share,
> but I could not bring a soft-RoCE listener up on the VM.
>
> Daire - the same nine patches on v7.2 (plus the wake fix) are at
> https://github.com/bcodding/linux branch nfsd-clientq-v2-7.2 if you
> want to run it; the loader builds from tools/net/sunrpc/svc-classify in
> that tree.
>
> The series is on nfsd-testing at 56589cdb5881, which already has the
> svc_clean_up_xprts() wake fix from [3].
>
> patch 1 track clients by class, no change to dispatch
> patch 2 dispatch round-robin across clients
> patch 3 the struct_ops hook
> patch 4 class in the svc_xprt_dequeue tracepoint
> patch 5 control events ahead of data within a client
> patch 6 flat queue while no classifier is attached
> patch 7 selftests: the prefix program and the attach tests
> patch 8 tools/net/sunrpc/svc-classify: the prefix classifier and its
> loader (load/replace/unload, add/del/list by prefix, status)
> patch 9 Documentation: the model, the class word, installing a
> classifier with svc-classify, six class maps and the share
> each gives, how it works underneath
>
> [1] https://lore.kernel.org/linux-nfs/cover.1780498019.git.bcodding@hammerspace.com
> [2] https://lore.kernel.org/linux-nfs/cover.1782314746.git.bcodding@hammerspace.com
> [3] https://lore.kernel.org/linux-nfs/cover.1790953694.git.bcodding@hammerspace.com
>
> Benjamin Coddington (9):
> SUNRPC: track service clients by class
> SUNRPC: dispatch ready transports round-robin across clients
> SUNRPC: add a BPF struct_ops hook to classify accepted transports
> SUNRPC: report the transport's class in the svc_xprt_dequeue
> tracepoint
> SUNRPC: dispatch control events ahead of data within a client
> SUNRPC: keep the single transport queue while no classifier is
> attached
> selftests/bpf: add svc_classifier tests
> tools/net/sunrpc: add svc-classify, a prefix classifier and its loader
> Documentation: describe RPC server transport classes and the BPF
> classifier
>
> Documentation/filesystems/nfs/index.rst | 1 +
> .../filesystems/nfs/rpc-server-clients.rst | 394 ++++++++++++++++++
> include/linux/sunrpc/svc.h | 72 +++-
> include/linux/sunrpc/svc_xprt.h | 30 ++
> include/trace/events/sunrpc.h | 11 +-
> net/sunrpc/Kconfig | 12 +
> net/sunrpc/Makefile | 1 +
> net/sunrpc/netns.h | 4 +
> net/sunrpc/sunrpc_syms.c | 4 +
> net/sunrpc/svc.c | 36 ++
> net/sunrpc/svc_classify.c | 210 ++++++++++
> net/sunrpc/svc_xprt.c | 281 ++++++++++++-
> tools/net/sunrpc/svc-classify/.gitignore | 1 +
> tools/net/sunrpc/svc-classify/Makefile | 65 +++
> tools/net/sunrpc/svc-classify/svc-classify.c | 357 ++++++++++++++++
> .../sunrpc/svc-classify/svc_classify.bpf.c | 83 ++++
> tools/testing/selftests/bpf/config | 2 +
> .../selftests/bpf/prog_tests/svc_classifier.c | 104 +++++
> .../selftests/bpf/progs/bpf_svc_classifier.c | 80 ++++
> 19 files changed, 1724 insertions(+), 24 deletions(-)
> create mode 100644 Documentation/filesystems/nfs/rpc-server-clients.rst
> create mode 100644 net/sunrpc/svc_classify.c
> create mode 100644 tools/net/sunrpc/svc-classify/.gitignore
> create mode 100644 tools/net/sunrpc/svc-classify/Makefile
> create mode 100644 tools/net/sunrpc/svc-classify/svc-classify.c
> create mode 100644 tools/net/sunrpc/svc-classify/svc_classify.bpf.c
> create mode 100644 tools/testing/selftests/bpf/prog_tests/svc_classifier.c
> create mode 100644 tools/testing/selftests/bpf/progs/bpf_svc_classifier.c
My Comments:
Conceptually, I like this idea, and I personally like that I now have a
clear (to me) example of how to plug eBPF programs into kernel code, so
thanks for that!
I think the userland bits should be made part of nfs-utils, and the
LLM-slop docs need to be rewritten for humans.
--
Jeff Layton <jlayton@kernel.org>
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader
2026-10-08 18:36 ` Jeff Layton
@ 2026-10-08 19:38 ` Benjamin Coddington
0 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-08 19:38 UTC (permalink / raw)
To: Jeff Layton
Cc: Benjamin Coddington, Chuck Lever, NeilBrown, linux-nfs,
Daire Byrne, bpf, Martin KaFai Lau
On 8 Oct 2026, at 14:36, Jeff Layton wrote:
> The userland programs in the kernel are often not packaged, so it seems
> to me that svc-classify would be better integrated into nfs-utils. That
> would allow you to more easily bundle it with the systemd service file.
>
> It might even make sense to integrate this functionality into nfsdctl,
> but I guess the classifier affects all rpc_serv objects as well (e.g.
> NLM), so it may be best to keep it as its own thing.
Technically we don't need this classifer, we could just expect that if you
want to use this feature you must write your own. I added it to make things
easier for potential testers - but it is limited. For example, it doesn't
classify on anything other than src address.
Ben
^ permalink raw reply [flat|nested] 17+ messages in thread
* Re: [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier
2026-10-08 18:42 ` Jeff Layton
@ 2026-10-08 19:40 ` Benjamin Coddington
0 siblings, 0 replies; 17+ messages in thread
From: Benjamin Coddington @ 2026-10-08 19:40 UTC (permalink / raw)
To: Jeff Layton
Cc: Benjamin Coddington, Chuck Lever, NeilBrown, linux-nfs,
Daire Byrne, bpf, Martin KaFai Lau
On 8 Oct 2026, at 14:42, Jeff Layton wrote:
> On Wed, 2026-10-07 at 15:59 -0400, Benjamin Coddington wrote:
>> From: Benjamin Coddington <bcodding@hammerspace.com>
>>
>> Dispatch round-robin across clients changes how a service shares its
>> threads, and the classes come from a BPF program the administrator
>> loads, so the administrator needs the model, the class word, and the
>> procedure. Add a page with those: install a classifier with
>> svc-classify (build, attach, fill the map, check, change, remove, keep
>> across reboots), six class maps and the share each gives its clients,
>> the tracepoint that shows the class, and how the hook, the program and
>> the attach model work underneath.
>>
>> Signed-off-by: Benjamin Coddington <bcodding@hammerspace.com>
>> ---
>> Documentation/filesystems/nfs/index.rst | 1 +
>> .../filesystems/nfs/rpc-server-clients.rst | 394 ++++++++++++++++++
>> 2 files changed, 395 insertions(+)
>> create mode 100644 Documentation/filesystems/nfs/rpc-server-clients.rst
>>
>> diff --git a/Documentation/filesystems/nfs/index.rst b/Documentation/filesystems/nfs/index.rst
>> index a29a212b5b4d..57a61ce0533d 100644
>> --- a/Documentation/filesystems/nfs/index.rst
>> +++ b/Documentation/filesystems/nfs/index.rst
>> @@ -16,3 +16,4 @@ NFS
>> nfsd-io-modes
>> knfsd-stats
>> reexport
>> + rpc-server-clients
>> diff --git a/Documentation/filesystems/nfs/rpc-server-clients.rst b/Documentation/filesystems/nfs/rpc-server-clients.rst
>> new file mode 100644
>> index 000000000000..b08351a88314
>> --- /dev/null
>> +++ b/Documentation/filesystems/nfs/rpc-server-clients.rst
>> @@ -0,0 +1,394 @@
>> +.. SPDX-License-Identifier: GPL-2.0
>> +
>> +======================================================
>> +RPC server: per-client dispatch and transport classes
>> +======================================================
>> +
>> +An RPC service thread pool serves its ready transports in FIFO order,
>> +one request per turn. A peer with many connections (``nconnect``, a
>> +deep NFSv4.1 slot table, a data mover with a thread per file) takes a
>> +turn per connection, and a peer with one connection waits behind all
>> +of them.
>> +
>> +With a classifier attached, the pool instead serves its *clients* round
>> +robin: each client with work queued gets one request per round, however
>> +many connections it holds. Which transports form a client is decided
>> +by a BPF program the administrator loads, once per accepted connection.
>> +With no program loaded nothing changes.
>> +
>> +This document describes the model, how to install a classifier with
>> +the ``svc-classify`` tool, what some class maps do, and how it works
>> +underneath.
>> +
>> +Clients, classes and turns
>> +==========================
>> +
>> +A *client* is a set of transports that share one turn. Every service
>> +(nfsd, lockd, the NFSv4 callback service) keeps its own clients, found
>> +by *class* and network namespace, or by class, namespace and peer
>> +address.
>> +
>> +The classifier returns a 32-bit class word for each accepted transport:
>> +
>> +``0``
>> + ``SVC_CLASS_NONE``: the transport stays on the service's *anonymous
>> + client*.
>> +
>> +``N``, 1 to 2^31 - 1
>> + one client for every transport of class ``N``.
>> +
>> +``N | SVC_CLASS_PER_ADDR``
>> + one client per peer address within class ``N``. ``N`` may be 0, so
>> + ``SVC_CLASS_PER_ADDR`` alone means one client per peer address.
>> +
>> +``SVC_CLASS_PER_ADDR`` is bit 31. ``svc-classify`` and this document
>> +write a class word with the bit set as ``N+addr``; ``svc-classify``
>> +also accepts ``addr`` for ``0+addr``.
>> +
>> +Dispatch rules:
>> +
>> +- A pool serves the clients that have transports queued round robin,
>> + one request per client per round.
>> +- Within a client, transports are served in FIFO order, except that a
>> + transport needing a connection accepted, closed, or a TLS handshake
>> + run goes ahead of transports with data, though not two such turns in
>> + a row while data waits.
>> +- A client's share of the pool does not grow with its connection count.
>> + A client with one connection and a client with thirty get the same
>> + number of turns while both have work queued.
>> +
>> +Things that follow from the rules and are easy to get wrong:
>> +
>> +- The anonymous client is one client. Once any classifier is attached,
>> + every transport whose class is ``0`` shares a single turn with every
>> + other such transport of that service. Within that one client the old
>> + FIFO order applies, so those peers share the turn in proportion to
>> + their connection counts. A map that classifies only the hosts it
>> + cares about and leaves the rest at ``0`` gives "the rest" one turn
>> + in total (see example 5 below).
>> +- Per-address clients ignore the port. Only IPv4 and IPv6 peers can be
>> + keyed by address; a transport whose peer is anything else is left on
>> + the anonymous client.
>> +- Listeners and UDP sockets are never classified; UDP traffic is served
>> + from the anonymous client.
>> +- A transport keeps the class it was given when accepted. Changing the
>> + map, or replacing or removing the classifier, affects connections
>> + accepted afterwards. A connection that must be reclassified has to
>> + reconnect.
>> +- nfsd has one service per network namespace, so its anonymous client
>> + is per namespace. lockd and the NFSv4 callback service are one
>> + service for all namespaces; their anonymous clients span namespaces.
>> +- With no classifier attached anywhere, the pool uses a single FIFO;
>> + the per-request cost of the feature is then a static branch on
>> + enqueue and one empty-queue test on dequeue.
>> + ``CONFIG_SUNRPC_BPF_CLASSIFY`` (default y) builds the hook; without
>> + it there is no classification.
>> +
>> +Installing a classifier
>> +=======================
>> +
>> +There is one tool to do it with: ``svc-classify``, in
>> +``tools/net/sunrpc/svc-classify`` of the kernel source. It carries the
>> +classifier program inside the binary, attaches it, and manages the
>> +class map by address prefix. ``bpftool`` is not needed, and its
>> +``struct_ops`` subcommands cannot be used in its place (see "How it
>> +works").
>> +
>
> "There is one tool to do it with: ..."
barf.
> The obviously LLM-written documentation gives me the Ick and I stopped
> reading. There is also a lot of info in this doc that is not
> particularly helpful, like "the things that follow from the rules and
> are easy to get wrong".
>
> If you want humans to actually read this, then it probably needs to be
> written by a human. If you think the LLM slop is actually useful for
> other LLMs, then maybe put those bits at the end in a clearly-
> delineated section.
Yeah, you're right. This really got spit out of the AI tubes without a lot
of refinement by me - thanks for reading even this far.
Ben
^ permalink raw reply [flat|nested] 17+ messages in thread
end of thread, other threads:[~2026-10-08 19:40 UTC | newest]
Thread overview: 17+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07 19:59 [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 1/9] SUNRPC: track service clients by class Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 2/9] SUNRPC: dispatch ready transports round-robin across clients Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 3/9] SUNRPC: add a BPF struct_ops hook to classify accepted transports Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 4/9] SUNRPC: report the transport's class in the svc_xprt_dequeue tracepoint Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 5/9] SUNRPC: dispatch control events ahead of data within a client Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 6/9] SUNRPC: keep the single transport queue while no classifier is attached Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 7/9] selftests/bpf: add svc_classifier tests Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 8/9] tools/net/sunrpc: add svc-classify, a prefix classifier and its loader Benjamin Coddington
2026-10-08 18:36 ` Jeff Layton
2026-10-08 19:38 ` Benjamin Coddington
2026-10-07 19:59 ` [PATCH RFC v2 9/9] Documentation: describe RPC server transport classes and the BPF classifier Benjamin Coddington
2026-10-08 18:42 ` Jeff Layton
2026-10-08 19:40 ` Benjamin Coddington
2026-10-08 15:19 ` [PATCH RFC v2 0/9] SUNRPC: dispatch ready transports round-robin across clients, classes from BPF Chuck Lever
2026-10-08 15:28 ` Benjamin Coddington
2026-10-08 18:48 ` Jeff Layton
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox