* [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 23:00 ` NeilBrown
2026-09-18 17:21 ` [PATCH v5 02/11] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
` (9 subsequent siblings)
10 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
nfsd_cache_size_limit() scales the DRC entry limit with the square
root of low memory, then clamps the result at 256k entries. The
clamp binds on any host with more than 64 GB of low memory. Once
the cache reaches max_drc_entries, nfsd_prune_bucket_locked()
evicts oldest-first regardless of age, so on such a host the cap
governs retention rather than RC_EXPIRE. A server that completes
more than about 2200 calls per second fills 256k entries inside
the 120 second RC_EXPIRE window, and from then on entries are
evicted before they expire. Those are the entries a retransmit
could still hit.
Commit 0338dd157282 ("nfsd: dynamically allocate DRC entries")
added the clamp in 2013, when the formula's worst case of 1 KB
per entry made 256 MB a reasonable ceiling for the hosts of the
day. The formula was already sized to memory; the clamp froze it
at that generation of hardware.
Remove the cap and let the square-root formula govern sizing. It
scales sub-linearly with memory, and the hash table already uses
kvzalloc, so larger sizes need no physical contiguity. The limit is
unchanged at 64 GB and below. At 1 TB the formula yields 1048576
entries, four times the old cap. Even at 1 KB per entry, the worst
case the old comment assumed, a full cache is 1 GB, 0.1% of that
host's memory. That memory is not reclaimable: the shrinker frees
only entries older than RC_EXPIRE, and each net namespace sizes its
own cache.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfscache.c | 33 ++++++++++++++++++---------------
1 file changed, 18 insertions(+), 15 deletions(-)
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index 80364b91331a..d0f65cc9c07a 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -47,34 +47,37 @@ static unsigned long nfsd_reply_cache_scan(struct shrinker *shrink,
struct shrink_control *sc);
/*
- * Put a cap on the size of the DRC based on the amount of available
- * low memory in the machine.
+ * Size the DRC by the amount of low memory in the machine. The
+ * limit scales with the square root of available pages; with 4 KB
+ * pages:
*
* 64MB: 8192
- * 128MB: 11585
+ * 128MB: 11584
* 256MB: 16384
- * 512MB: 23170
+ * 512MB: 23168
* 1GB: 32768
- * 2GB: 46340
+ * 2GB: 46336
* 4GB: 65536
- * 8GB: 92681
+ * 8GB: 92672
* 16GB: 131072
+ * 32GB: 185344
+ * 64GB: 262144
+ * 128GB: 370688
+ * 256GB: 524288
+ * 512GB: 741440
+ * 1TB: 1048576
*
- * ...with a hard cap of 256k entries. In the worst case, each entry will be
- * ~1k, so the above numbers should give a rough max of the amount of memory
- * used in k.
- *
- * XXX: these limits are per-container, so memory used will increase
- * linearly with number of containers. Maybe that's OK.
+ * These limits are per-net-namespace, so memory used increases
+ * linearly with the number of namespaces. The shrinker frees only
+ * entries older than RC_EXPIRE, so a cache below its limit is not
+ * reclaimable under memory pressure.
*/
static unsigned int
nfsd_cache_size_limit(void)
{
- unsigned int limit;
unsigned long low_pages = totalram_pages() - totalhigh_pages();
- limit = (16 * int_sqrt(low_pages)) << (PAGE_SHIFT-10);
- return min_t(unsigned int, limit, 256*1024);
+ return (16 * int_sqrt(low_pages)) << (PAGE_SHIFT - 10);
}
/*
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* Re: [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size
2026-09-18 17:21 ` [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
@ 2026-09-18 23:00 ` NeilBrown
2026-09-20 17:47 ` Chuck Lever
0 siblings, 1 reply; 19+ messages in thread
From: NeilBrown @ 2026-09-18 23:00 UTC (permalink / raw)
To: Chuck Lever
Cc: Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey, Rick Macklem,
linux-nfs, Chuck Lever
On Sat, 19 Sep 2026, Chuck Lever wrote:
> nfsd_cache_size_limit() scales the DRC entry limit with the square
> root of low memory, then clamps the result at 256k entries. The
> clamp binds on any host with more than 64 GB of low memory. Once
> the cache reaches max_drc_entries, nfsd_prune_bucket_locked()
> evicts oldest-first regardless of age, so on such a host the cap
> governs retention rather than RC_EXPIRE. A server that completes
> more than about 2200 calls per second fills 256k entries inside
> the 120 second RC_EXPIRE window, and from then on entries are
> evicted before they expire. Those are the entries a retransmit
> could still hit.
>
> Commit 0338dd157282 ("nfsd: dynamically allocate DRC entries")
> added the clamp in 2013, when the formula's worst case of 1 KB
> per entry made 256 MB a reasonable ceiling for the hosts of the
> day. The formula was already sized to memory; the clamp froze it
> at that generation of hardware.
>
> Remove the cap and let the square-root formula govern sizing. It
> scales sub-linearly with memory, and the hash table already uses
> kvzalloc, so larger sizes need no physical contiguity. The limit is
> unchanged at 64 GB and below. At 1 TB the formula yields 1048576
> entries, four times the old cap. Even at 1 KB per entry, the worst
> case the old comment assumed, a full cache is 1 GB, 0.1% of that
> host's memory. That memory is not reclaimable: the shrinker frees
> only entries older than RC_EXPIRE, and each net namespace sizes its
> own cache.
>
> Assisted-by: LLM
> Signed-off-by: Chuck Lever <cel@kernel.org>
> ---
> fs/nfsd/nfscache.c | 33 ++++++++++++++++++---------------
> 1 file changed, 18 insertions(+), 15 deletions(-)
>
> diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
> index 80364b91331a..d0f65cc9c07a 100644
> --- a/fs/nfsd/nfscache.c
> +++ b/fs/nfsd/nfscache.c
> @@ -47,34 +47,37 @@ static unsigned long nfsd_reply_cache_scan(struct shrinker *shrink,
> struct shrink_control *sc);
>
> /*
> - * Put a cap on the size of the DRC based on the amount of available
> - * low memory in the machine.
> + * Size the DRC by the amount of low memory in the machine. The
> + * limit scales with the square root of available pages; with 4 KB
> + * pages:
> *
> * 64MB: 8192
> - * 128MB: 11585
> + * 128MB: 11584
> * 256MB: 16384
> - * 512MB: 23170
> + * 512MB: 23168
> * 1GB: 32768
> - * 2GB: 46340
> + * 2GB: 46336
> * 4GB: 65536
> - * 8GB: 92681
> + * 8GB: 92672
> * 16GB: 131072
> + * 32GB: 185344
> + * 64GB: 262144
> + * 128GB: 370688
> + * 256GB: 524288
> + * 512GB: 741440
> + * 1TB: 1048576
> *
> - * ...with a hard cap of 256k entries. In the worst case, each entry will be
> - * ~1k, so the above numbers should give a rough max of the amount of memory
> - * used in k.
> - *
> - * XXX: these limits are per-container, so memory used will increase
> - * linearly with number of containers. Maybe that's OK.
> + * These limits are per-net-namespace, so memory used increases
> + * linearly with the number of namespaces. The shrinker frees only
> + * entries older than RC_EXPIRE, so a cache below its limit is not
> + * reclaimable under memory pressure.
> */
> static unsigned int
> nfsd_cache_size_limit(void)
> {
> - unsigned int limit;
> unsigned long low_pages = totalram_pages() - totalhigh_pages();
>
> - limit = (16 * int_sqrt(low_pages)) << (PAGE_SHIFT-10);
> - return min_t(unsigned int, limit, 256*1024);
> + return (16 * int_sqrt(low_pages)) << (PAGE_SHIFT - 10);
> }
This patch looks good and does what it says on the tin....
But that formula at the end (which the patch doesn't change) looks weird.
As you note in the table above, 8GB on 4K pages results in 92681 target
cache entries.
8GB with 64K pages comes to 370727 cache entries - 4 times as many.
That doesn't seem justified.
If we made it
32 * int_sqrt(low_pages << (PAGE_SHIFT-10))
or even
int_sqrt(low_pages << PAGE_SHIFT)
it would be 92681 entries no matter the page size which is what the
original table says.
Does this need fixing? Should we fix it now?
NeilBrown
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size
2026-09-18 23:00 ` NeilBrown
@ 2026-09-20 17:47 ` Chuck Lever
2026-09-20 22:10 ` NeilBrown
0 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2026-09-20 17:47 UTC (permalink / raw)
To: NeilBrown
Cc: Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey, Rick Macklem,
linux-nfs
On Fri, Sep 18, 2026, at 7:00 PM, NeilBrown wrote:
> This patch looks good and does what it says on the tin....
> But that formula at the end (which the patch doesn't change) looks weird.
>
> As you note in the table above, 8GB on 4K pages results in 92681 target
> cache entries.
> 8GB with 64K pages comes to 370727 cache entries - 4 times as many.
> That doesn't seem justified.
>
> If we made it
>
> 32 * int_sqrt(low_pages << (PAGE_SHIFT-10))
> or even
> int_sqrt(low_pages << PAGE_SHIFT)
>
> it would be 92681 entries no matter the page size which is what the
> original table says.
>
> Does this need fixing? Should we fix it now?
I'll add it as a separate patch ahead of the cap removal, using
int_sqrt(low_pages << PAGE_SHIFT)
The table in the comment then matches on every page size, so the
"with 4 KB pages" qualifier goes away.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size
2026-09-20 17:47 ` Chuck Lever
@ 2026-09-20 22:10 ` NeilBrown
0 siblings, 0 replies; 19+ messages in thread
From: NeilBrown @ 2026-09-20 22:10 UTC (permalink / raw)
To: Chuck Lever
Cc: Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey, Rick Macklem,
linux-nfs
On Mon, 21 Sep 2026, Chuck Lever wrote:
> On Fri, Sep 18, 2026, at 7:00 PM, NeilBrown wrote:
> > This patch looks good and does what it says on the tin....
> > But that formula at the end (which the patch doesn't change) looks weird.
> >
> > As you note in the table above, 8GB on 4K pages results in 92681 target
> > cache entries.
> > 8GB with 64K pages comes to 370727 cache entries - 4 times as many.
> > That doesn't seem justified.
> >
> > If we made it
> >
> > 32 * int_sqrt(low_pages << (PAGE_SHIFT-10))
> > or even
> > int_sqrt(low_pages << PAGE_SHIFT)
> >
> > it would be 92681 entries no matter the page size which is what the
> > original table says.
> >
> > Does this need fixing? Should we fix it now?
>
> I'll add it as a separate patch ahead of the cap removal, using
>
> int_sqrt(low_pages << PAGE_SHIFT)
I worry a bit about overflow. sqrt() accepts an unsigned long
and low_memory is by definition addresses with a pointer the same size
of an unsigned long. So the only possible problem is that if there
were 4GB of low pages on a 32bit system. Then "low_pages << PAGE_SHIFT"
would be zero.
But that isn't possible - for exactly the reason that we can put error
numbers in the two 4096 values uses for addresses. low memory never
quite reaches the top.
So it can never overflow. (good)
Thanks,
NeilBrown
>
> The table in the comment then matches on every page size, so the
> "with 4 KB pages" qualifier goes away.
>
>
> --
> Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
>
^ permalink raw reply [flat|nested] 19+ messages in thread
* [PATCH v5 02/11] SUNRPC: Assign a unique identifier to each svc_xprt
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-18 17:21 ` [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 03/11] NFSD: Track transport in DRC entries Chuck Lever
` (8 subsequent siblings)
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
A consumer that associates state with a transport without holding a
reference cannot detect ABA collisions: once a transport is freed,
SLUB may hand out a new svc_xprt at the same address.
Assign each transport a 64-bit identifier from a global sequential
counter in svc_xprt_init(). The counter cannot wrap within the
lifetime of a transport, so two live transports never share an
identifier and a consumer needs no release when the transport is
freed. The NFSD duplicate reply cache is the first consumer.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
include/linux/sunrpc/svc_xprt.h | 1 +
include/trace/events/sunrpc.h | 8 ++++++--
net/sunrpc/svc_xprt.c | 17 ++++++++++++++---
3 files changed, 21 insertions(+), 5 deletions(-)
diff --git a/include/linux/sunrpc/svc_xprt.h b/include/linux/sunrpc/svc_xprt.h
index 2af222f3ea2c..52082d8ea6dd 100644
--- a/include/linux/sunrpc/svc_xprt.h
+++ b/include/linux/sunrpc/svc_xprt.h
@@ -56,6 +56,7 @@ struct svc_xprt {
struct svc_xprt_class *xpt_class;
const struct svc_xprt_ops *xpt_ops;
struct kref xpt_ref;
+ u64 xpt_id;
ktime_t xpt_qtime;
struct list_head xpt_list;
struct lwq_node xpt_ready;
diff --git a/include/trace/events/sunrpc.h b/include/trace/events/sunrpc.h
index ff855197880d..77efe63d2c14 100644
--- a/include/trace/events/sunrpc.h
+++ b/include/trace/events/sunrpc.h
@@ -1986,7 +1986,8 @@ TRACE_EVENT(svc_xprt_create_err,
__sockaddr(server, (x)->xpt_locallen) \
__sockaddr(client, (x)->xpt_remotelen) \
__field(unsigned long, flags) \
- __field(unsigned int, netns_ino)
+ __field(unsigned int, netns_ino) \
+ __field(u64, xpt_id)
#define SVC_XPRT_ENDPOINT_ASSIGNMENTS(x) \
do { \
@@ -1996,13 +1997,15 @@ TRACE_EVENT(svc_xprt_create_err,
(x)->xpt_remotelen); \
__entry->flags = (x)->xpt_flags; \
__entry->netns_ino = (x)->xpt_net->ns.inum; \
+ __entry->xpt_id = (x)->xpt_id; \
} while (0)
#define SVC_XPRT_ENDPOINT_FORMAT \
- "server=%pISpc client=%pISpc flags=%s"
+ "server=%pISpc client=%pISpc xpt_id=%llu flags=%s"
#define SVC_XPRT_ENDPOINT_VARARGS \
__get_sockaddr(server), __get_sockaddr(client), \
+ __entry->xpt_id, \
show_svc_xprt_flags(__entry->flags)
TRACE_EVENT(svc_xprt_enqueue,
@@ -2024,6 +2027,7 @@ TRACE_EVENT(svc_xprt_enqueue,
xprt->xpt_remotelen);
__entry->flags = flags;
__entry->netns_ino = xprt->xpt_net->ns.inum;
+ __entry->xpt_id = xprt->xpt_id;
),
TP_printk(SVC_XPRT_ENDPOINT_FORMAT, SVC_XPRT_ENDPOINT_VARARGS)
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index d5634dd6d6cc..463ba8d77f0c 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -26,6 +26,8 @@
static unsigned int svc_rpc_per_connection_limit __read_mostly;
module_param(svc_rpc_per_connection_limit, uint, 0644);
+static atomic64_t svc_xprt_id_seq;
+
static struct svc_deferred_req *svc_deferred_dequeue(struct svc_xprt *xprt);
static int svc_deferred_recv(struct svc_rqst *rqstp);
@@ -190,9 +192,17 @@ void svc_xprt_put(struct svc_xprt *xprt)
}
EXPORT_SYMBOL_GPL(svc_xprt_put);
-/*
- * Called by transport drivers to initialize the transport independent
- * portion of the transport instance.
+/**
+ * svc_xprt_init - initialize transport-independent portion of a transport
+ * @net: network namespace in which the transport operates
+ * @xcl: transport class providing operations and metadata
+ * @xprt: svc_xprt to initialize
+ * @serv: RPC service that owns this transport
+ *
+ * @xprt->xpt_id is never reused: a 64-bit counter does not wrap
+ * within the lifetime of a transport.
+ *
+ * Context: Process context. May sleep.
*/
void svc_xprt_init(struct net *net, struct svc_xprt_class *xcl,
struct svc_xprt *xprt, struct svc_serv *serv)
@@ -210,6 +220,7 @@ void svc_xprt_init(struct net *net, struct svc_xprt_class *xcl,
set_bit(XPT_BUSY, &xprt->xpt_flags);
xprt->xpt_net = get_net_track(net, &xprt->ns_tracker, GFP_ATOMIC);
strcpy(xprt->xpt_remotebuf, "uninitialized");
+ xprt->xpt_id = atomic64_inc_return(&svc_xprt_id_seq);
}
EXPORT_SYMBOL_GPL(svc_xprt_init);
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 03/11] NFSD: Track transport in DRC entries
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-18 17:21 ` [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
2026-09-18 17:21 ` [PATCH v5 02/11] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 04/11] NFSD: Prepare bucket pruning for additional eviction reasons Chuck Lever
` (7 subsequent siblings)
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
A cached reply carries no record of the transport its request arrived
on. Transport-aware eviction needs that association.
Record the xprt id in each entry. struct nfsd_cacherep grows from 128
to 136 bytes. No behavior change.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/cache.h | 1 +
fs/nfsd/nfscache.c | 1 +
2 files changed, 2 insertions(+)
diff --git a/fs/nfsd/cache.h b/fs/nfsd/cache.h
index 3bc4856e34b8..589ffe4762d5 100644
--- a/fs/nfsd/cache.h
+++ b/fs/nfsd/cache.h
@@ -37,6 +37,7 @@ struct nfsd_cacherep {
unsigned char c_state, /* unused, inprog, done */
c_type, /* status, buffer */
c_secure : 1; /* req came from port < 1024 */
+ u64 c_xprt; /* svc_xprt that carried req */
unsigned long c_timestamp;
union {
struct kvec u_vec;
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index d0f65cc9c07a..474bfeecde9e 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -112,6 +112,7 @@ nfsd_cacherep_alloc(struct svc_rqst *rqstp, __wsum csum,
rp->c_key.k_vers = rqstp->rq_vers;
rp->c_key.k_len = rqstp->rq_arg.len;
rp->c_key.k_csum = csum;
+ rp->c_xprt = rqstp->rq_xprt->xpt_id;
}
return rp;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 04/11] NFSD: Prepare bucket pruning for additional eviction reasons
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
` (2 preceding siblings ...)
2026-09-18 17:21 ` [PATCH v5 03/11] NFSD: Track transport in DRC entries Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 05/11] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
` (6 subsequent siblings)
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
nfsd_prune_bucket_locked() combines its two eviction reasons,
memory pressure and RC_EXPIRE, in a single test and stops the walk
at the first entry the test does not evict. A reason that needs
its own handling, a tracepoint per reason or a test that must run
ahead of the stop, has nowhere to go.
Test each reason on its own and jump to a shared eviction label.
The walk still stops at the first entry no reason applies to.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfscache.c | 9 ++++++---
1 file changed, 6 insertions(+), 3 deletions(-)
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index 474bfeecde9e..be0b27af8ff0 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -275,10 +275,13 @@ nfsd_prune_bucket_locked(struct nfsd_net *nn, struct nfsd_drc_bucket *b,
/* The bucket LRU is ordered oldest-first. */
list_for_each_entry_safe(rp, tmp, &b->lru_head, c_lru) {
- if (atomic_read(&nn->num_drc_entries) <= nn->max_drc_entries &&
- time_before(expiry, rp->c_timestamp))
- break;
+ if (atomic_read(&nn->num_drc_entries) > nn->max_drc_entries)
+ goto evict;
+ if (time_before_eq(rp->c_timestamp, expiry))
+ goto evict;
+ break;
+evict:
nfsd_cacherep_unlink_locked(nn, b, rp);
list_add(&rp->c_lru, dispose);
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 05/11] NFSD: Add tracepoints for DRC entry eviction
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
` (3 preceding siblings ...)
2026-09-18 17:21 ` [PATCH v5 04/11] NFSD: Prepare bucket pruning for additional eviction reasons Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 06/11] NFSD: Record DRC population in lookup tracepoints Chuck Lever
` (5 subsequent siblings)
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
The DRC pruning path evicts entries for memory pressure and expiry,
and neither emits a trace event, so there is no way to see which
reason is retiring entries or how old they are when it happens. A
further eviction reason needs the same visibility.
Add an event class that records the cache population, the XID, the
transport, and the entry's age, and define an event for each
eviction reason.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfscache.c | 8 ++++++--
fs/nfsd/trace.h | 37 +++++++++++++++++++++++++++++++++++++
2 files changed, 43 insertions(+), 2 deletions(-)
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index be0b27af8ff0..4d0101e93add 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -275,10 +275,14 @@ nfsd_prune_bucket_locked(struct nfsd_net *nn, struct nfsd_drc_bucket *b,
/* The bucket LRU is ordered oldest-first. */
list_for_each_entry_safe(rp, tmp, &b->lru_head, c_lru) {
- if (atomic_read(&nn->num_drc_entries) > nn->max_drc_entries)
+ if (atomic_read(&nn->num_drc_entries) > nn->max_drc_entries) {
+ trace_nfsd_drc_evict_pressure(nn, rp);
goto evict;
- if (time_before_eq(rp->c_timestamp, expiry))
+ }
+ if (time_before_eq(rp->c_timestamp, expiry)) {
+ trace_nfsd_drc_evict_expired(nn, rp);
goto evict;
+ }
break;
evict:
diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h
index 2ae7f150a72c..ed4efcfa3961 100644
--- a/fs/nfsd/trace.h
+++ b/fs/nfsd/trace.h
@@ -1557,6 +1557,43 @@ TRACE_EVENT(nfsd_drc_mismatch,
__entry->ingress)
);
+DECLARE_EVENT_CLASS(nfsd_drc_entry_class,
+ TP_PROTO(
+ const struct nfsd_net *nn,
+ const struct nfsd_cacherep *rp
+ ),
+ TP_ARGS(nn, rp),
+ TP_STRUCT__entry(
+ __field(unsigned long long, boot_time)
+ __field(unsigned int, num_drc_entries)
+ __field(u32, xid)
+ __field(u64, xprt)
+ __field(unsigned long, age)
+ ),
+ TP_fast_assign(
+ __entry->boot_time = nn->boot_time;
+ __entry->num_drc_entries = atomic_read(&nn->num_drc_entries);
+ __entry->xid = be32_to_cpu(rp->c_key.k_xid);
+ __entry->xprt = rp->c_xprt;
+ __entry->age = time_is_after_jiffies(rp->c_timestamp) ?
+ 1 : jiffies - rp->c_timestamp;
+ ),
+ TP_printk("boot_time=%16llx entries=%u xid=0x%08x xprt=%llu age=%lu",
+ __entry->boot_time, __entry->num_drc_entries,
+ __entry->xid, __entry->xprt, __entry->age)
+);
+
+#define DEFINE_NFSD_DRC_ENTRY_EVENT(name) \
+DEFINE_EVENT(nfsd_drc_entry_class, nfsd_drc_##name, \
+ TP_PROTO( \
+ const struct nfsd_net *nn, \
+ const struct nfsd_cacherep *rp \
+ ), \
+ TP_ARGS(nn, rp))
+
+DEFINE_NFSD_DRC_ENTRY_EVENT(evict_pressure);
+DEFINE_NFSD_DRC_ENTRY_EVENT(evict_expired);
+
TRACE_EVENT(nfsd_cb_args,
TP_PROTO(
const struct nfs4_client *clp,
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 06/11] NFSD: Record DRC population in lookup tracepoints
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
` (4 preceding siblings ...)
2026-09-18 17:21 ` [PATCH v5 05/11] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 07/11] SUNRPC: Publish reply positions for upper-layer consumers Chuck Lever
` (4 subsequent siblings)
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
nfsd_drc_found reports the outcome of a lookup but not how full the
cache was at the time, so a trace cannot show whether retransmits are
being caught while the cache runs near its cap.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfscache.c | 4 +++-
fs/nfsd/trace.h | 11 +++++++----
2 files changed, 10 insertions(+), 5 deletions(-)
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index 4d0101e93add..cd1011bd148d 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -487,6 +487,7 @@ int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
struct nfsd_drc_bucket *b;
int type = ntli->ntli_cachetype;
LIST_HEAD(dispose);
+ unsigned int entries;
int rtn = RC_DOIT;
if (type == RC_NOCACHE) {
@@ -556,7 +557,8 @@ int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
}
out_trace:
- trace_nfsd_drc_found(nn, rqstp, rtn);
+ entries = atomic_read(&nn->num_drc_entries);
+ trace_nfsd_drc_found(nn, entries, rqstp, rtn);
out_unlock:
spin_unlock(&b->cache_lock);
out:
diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h
index ed4efcfa3961..631682a76f9c 100644
--- a/fs/nfsd/trace.h
+++ b/fs/nfsd/trace.h
@@ -1513,23 +1513,26 @@ TRACE_DEFINE_ENUM(RC_DOIT);
TRACE_EVENT(nfsd_drc_found,
TP_PROTO(
const struct nfsd_net *nn,
+ unsigned int num_drc_entries,
const struct svc_rqst *rqstp,
int result
),
- TP_ARGS(nn, rqstp, result),
+ TP_ARGS(nn, num_drc_entries, rqstp, result),
TP_STRUCT__entry(
__field(unsigned long long, boot_time)
+ __field(unsigned int, num_drc_entries)
__field(unsigned long, result)
__field(u32, xid)
),
TP_fast_assign(
__entry->boot_time = nn->boot_time;
+ __entry->num_drc_entries = num_drc_entries;
__entry->result = result;
__entry->xid = be32_to_cpu(rqstp->rq_xid);
),
- TP_printk("boot_time=%16llx xid=0x%08x result=%s",
- __entry->boot_time, __entry->xid,
- show_drc_retval(__entry->result))
+ TP_printk("boot_time=%16llx entries=%u xid=0x%08x result=%s",
+ __entry->boot_time, __entry->num_drc_entries,
+ __entry->xid, show_drc_retval(__entry->result))
);
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 07/11] SUNRPC: Publish reply positions for upper-layer consumers
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
` (5 preceding siblings ...)
2026-09-18 17:21 ` [PATCH v5 06/11] NFSD: Record DRC population in lookup tracepoints Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 08/11] SUNRPC: Publish TCP reply positions Chuck Lever
` (3 subsequent siblings)
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
An RPC server's duplicate reply cache has to hold a cached reply
until the client has received it. SUNRPC gives an upper layer no way
to learn that, so a cache entry lives for the full retention
interval even when the reply was delivered at once.
Add two transport-neutral values that an upper layer can compare:
the position of each reply on its transport, recorded in the svc_rqst
by xpo_sendto(), and the position up to which the peer has
acknowledged, recorded on the svc_xprt. The units are private to the
transport (bytes for a stream, Send sequence for RDMA); a consumer
learns only whether one is less than or equal to the other. A reply
that xpo_sendto() does not hand to the wire keeps position zero and
is never reported as acknowledged.
The acknowledged position is an atomic64_t. A transport publishes it
outside any lock the consumer holds, and a torn 64-bit read on a
32-bit host could run ahead of the peer.
Add a per-service hook that runs once per request when its reply
phase ends, whether svc_send() sent the reply or svc_process()
dropped it, so an upper layer can record the reply's position
against its own state either way. On the drop path the position is
zero.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
include/linux/sunrpc/svc.h | 16 ++++++++++++++++
include/linux/sunrpc/svc_xprt.h | 2 ++
net/sunrpc/svc.c | 3 +++
net/sunrpc/svc_xprt.c | 2 ++
4 files changed, 23 insertions(+)
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index 24698856eb40..01036e093012 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -59,6 +59,8 @@ enum {
};
+struct svc_rqst;
+
/*
* RPC service.
*
@@ -96,6 +98,13 @@ struct svc_serv {
* connection */
bool sv_bc_enabled; /* service uses backchannel */
#endif /* CONFIG_SUNRPC_BACKCHANNEL */
+
+ /*
+ * Called once per request after its reply phase, whether
+ * xpo_sendto() succeeded, failed, or was never reached because
+ * svc_process() dropped the reply. See rq_reply_pos.
+ */
+ void (*sv_reply_sent)(struct svc_rqst *rqstp);
};
/* This is used by pool_stats to find and lock an svc */
@@ -267,6 +276,13 @@ struct svc_rqst {
unsigned int bc_to_retries;
unsigned int rq_status_counter; /* RPC processing counter */
void *rq_private; /* For use by the service thread */
+
+ /*
+ * Transport position of this request's reply, set by
+ * xpo_sendto(); zero when the reply did not reach the wire.
+ * Comparable only with the xpt_acked_pos of the same svc_xprt.
+ */
+ u64 rq_reply_pos;
};
/* bits for rq_flags */
diff --git a/include/linux/sunrpc/svc_xprt.h b/include/linux/sunrpc/svc_xprt.h
index 52082d8ea6dd..7176c42f19d7 100644
--- a/include/linux/sunrpc/svc_xprt.h
+++ b/include/linux/sunrpc/svc_xprt.h
@@ -66,6 +66,8 @@ struct svc_xprt {
atomic_t xpt_reserved; /* outq space rsvd, UDP only */
atomic_t xpt_nr_rqsts; /* Number of requests */
struct mutex xpt_mutex; /* to serialize sending data */
+ atomic64_t xpt_acked_pos; /* replies up to this position
+ * are acknowledged; 0 = none */
spinlock_t xpt_lock; /* protects sk_deferred
* and xpt_auth_cache */
void *xpt_auth_cache;/* auth cache */
diff --git a/net/sunrpc/svc.c b/net/sunrpc/svc.c
index f73412e123a1..9ef0661bb422 100644
--- a/net/sunrpc/svc.c
+++ b/net/sunrpc/svc.c
@@ -1488,6 +1488,7 @@ svc_process_common(struct svc_rqst *rqstp)
/* Reset the accept_stat for the RPC */
rqstp->rq_accept_statp = NULL;
+ rqstp->rq_reply_pos = 0;
/* Will be turned off only when NFSv4 Sessions are used */
set_bit(RQ_USEDEFERRAL, &rqstp->rq_flags);
@@ -1733,6 +1734,8 @@ void svc_process(struct svc_rqst *rqstp)
goto out_baddir;
if (!svc_process_common(rqstp)) {
+ if (rqstp->rq_server->sv_reply_sent)
+ rqstp->rq_server->sv_reply_sent(rqstp);
svc_release_rqst(rqstp);
goto out_drop;
}
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 463ba8d77f0c..1668d96ac9ff 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -1028,6 +1028,8 @@ void svc_send(struct svc_rqst *rqstp)
trace_svc_stats_latency(rqstp);
status = xprt->xpt_ops->xpo_sendto(rqstp);
+ if (xprt->xpt_server->sv_reply_sent)
+ xprt->xpt_server->sv_reply_sent(rqstp);
trace_svc_send(rqstp, status);
}
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 08/11] SUNRPC: Publish TCP reply positions
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
` (6 preceding siblings ...)
2026-09-18 17:21 ` [PATCH v5 07/11] SUNRPC: Publish reply positions for upper-layer consumers Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 09/11] svcrdma: Publish RDMA " Chuck Lever
` (2 subsequent siblings)
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
Add the svcsock side of reply-position reporting. A reply's position
is the running count of bytes handed to the socket, kept per socket
under xpt_mutex. A TCP sequence number cannot serve as the position:
the 32-bit space wraps in a few hundred milliseconds at 40 GbE, and
a duplicate reply cache entry lives for two minutes.
The acknowledged position is derived from the socket's unacknowledged
byte count, write_seq minus snd_una, after each successful send. The
two are read without the socket lock, as tcp_ioctl() reads them for
SIOCOUTQ: write_seq advances only under xpt_mutex, and a stale
snd_una only lowers the result. Refreshing the position only in the
send path means a consumer sees the acknowledgment state as of the
previous reply on that connection, stale by one round trip but never
ahead of the peer. A failed or refused send leaves rq_reply_pos at
zero, so the reply is never reported as acknowledged.
Under kTLS the socket counts ciphertext while svcsock counts
plaintext, and tls_sw_sendmsg() can return the full plaintext count
with part of a record still waiting for socket write space. The
unacknowledged count then understates what the peer has yet to
receive. Whether a record is pending is state private to net/tls, so
a TLS session publishes no acknowledged position and its replies
are retained for the full interval.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
include/linux/sunrpc/svcsock.h | 3 +++
net/sunrpc/svcsock.c | 29 +++++++++++++++++++++++++++++
2 files changed, 32 insertions(+)
diff --git a/include/linux/sunrpc/svcsock.h b/include/linux/sunrpc/svcsock.h
index 372a00882ca6..b6e0767de35a 100644
--- a/include/linux/sunrpc/svcsock.h
+++ b/include/linux/sunrpc/svcsock.h
@@ -41,6 +41,9 @@ struct svc_sock {
struct page_frag_cache sk_frag_cache;
+ /* reply bytes handed to the socket; protected by xpt_mutex */
+ u64 sk_send_pos;
+
struct completion sk_handshake_done;
/* received data */
diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c
index e5459d504b6a..b9bab3751e76 100644
--- a/net/sunrpc/svcsock.c
+++ b/net/sunrpc/svcsock.c
@@ -1390,6 +1390,32 @@ static int svc_tcp_sendmsg(struct svc_sock *svsk, struct svc_rqst *rqstp,
return ret;
}
+/*
+ * Bytes the socket has not seen acknowledged all belong to the most
+ * recent replies, so sk_send_pos less that count is the acknowledged
+ * position. write_seq advances only under xpt_mutex, which the caller
+ * holds, and a stale snd_una only lowers the result.
+ *
+ * Under kTLS, tls_sw_sendmsg() can return the full plaintext count
+ * with part of a record still waiting for socket write space, leaving
+ * write_seq short of the reply. Whether a record is pending is private
+ * to net/tls, so a TLS session publishes no acknowledged position.
+ */
+static void svc_tcp_update_acked(struct svc_sock *svsk)
+{
+ struct tcp_sock *tp = tcp_sk(svsk->sk_sk);
+ u64 unacked, acked;
+
+ if (test_bit(XPT_TLS_SESSION, &svsk->sk_xprt.xpt_flags))
+ return;
+ unacked = READ_ONCE(tp->write_seq) - READ_ONCE(tp->snd_una);
+ if (unacked >= svsk->sk_send_pos)
+ return;
+ acked = svsk->sk_send_pos - unacked;
+ if (acked > atomic64_read(&svsk->sk_xprt.xpt_acked_pos))
+ atomic64_set(&svsk->sk_xprt.xpt_acked_pos, acked);
+}
+
/**
* svc_tcp_sendto - Send out a reply on a TCP socket
* @rqstp: completed svc_rqst
@@ -1418,6 +1444,9 @@ static int svc_tcp_sendto(struct svc_rqst *rqstp)
trace_svcsock_tcp_send(xprt, sent);
if (sent < 0 || sent != (xdr->len + sizeof(marker)))
goto out_close;
+ svsk->sk_send_pos += sent;
+ rqstp->rq_reply_pos = svsk->sk_send_pos;
+ svc_tcp_update_acked(svsk);
mutex_unlock(&xprt->xpt_mutex);
return sent;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 09/11] svcrdma: Publish RDMA reply positions
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
` (7 preceding siblings ...)
2026-09-18 17:21 ` [PATCH v5 08/11] SUNRPC: Publish TCP reply positions Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 10/11] NFSD: Evict acknowledged DRC entries Chuck Lever
2026-09-18 17:21 ` [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
Add the svcrdma side of reply-position reporting. A reply's position
is the sequence number of its Send on the transport, and the Send
completion for that reply publishes the same number as the
transport's acknowledged position. A successful Send completion means
the peer's HCA has acknowledged the message, so a reply whose Send
has completed is in the client's receive buffer.
Completions on one queue pair arrive in posting order. Several nfsd
threads post on one queue pair with no serialization above the
provider, so the order in which Sends are queued is decided inside
the provider. Assign the sequence number and post the Send under one
lock so positions match queuing order. ib_post_send() does not
sleep; the Send Queue wait stays outside the lock.
Only a posted RPC Reply reports a position. The RDMA_ERROR message
and backchannel calls post with no position pointer, and a Send that
fails to post leaves rq_reply_pos at zero.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
include/linux/sunrpc/svc_rdma.h | 5 ++++-
net/sunrpc/xprtrdma/svc_rdma_backchannel.c | 2 +-
net/sunrpc/xprtrdma/svc_rdma_sendto.c | 27 ++++++++++++++++++++++++---
net/sunrpc/xprtrdma/svc_rdma_transport.c | 1 +
4 files changed, 30 insertions(+), 5 deletions(-)
diff --git a/include/linux/sunrpc/svc_rdma.h b/include/linux/sunrpc/svc_rdma.h
index 76aa5ec4ab40..7627afed16bb 100644
--- a/include/linux/sunrpc/svc_rdma.h
+++ b/include/linux/sunrpc/svc_rdma.h
@@ -96,6 +96,8 @@ struct svcxprt_rdma {
spinlock_t sc_send_lock;
struct llist_head sc_send_ctxts;
+ spinlock_t sc_post_lock; /* orders Send posting */
+ u64 sc_post_seq; /* Sends posted so far */
spinlock_t sc_rw_ctxt_lock;
struct llist_head sc_rw_ctxts;
@@ -242,6 +244,7 @@ struct svc_rdma_send_ctxt {
struct ib_send_wr *sc_wr_chain;
int sc_sqecount;
struct ib_cqe sc_cqe;
+ u64 sc_pos;
struct xdr_buf sc_hdrbuf;
struct xdr_stream sc_stream;
@@ -305,7 +308,7 @@ extern struct svc_rdma_send_ctxt *
extern void svc_rdma_send_ctxt_put(struct svcxprt_rdma *rdma,
struct svc_rdma_send_ctxt *ctxt);
extern int svc_rdma_post_send(struct svcxprt_rdma *rdma,
- struct svc_rdma_send_ctxt *ctxt);
+ struct svc_rdma_send_ctxt *ctxt, u64 *pos);
extern int svc_rdma_map_reply_msg(struct svcxprt_rdma *rdma,
struct svc_rdma_send_ctxt *sctxt,
const struct svc_rdma_pcl *write_pcl,
diff --git a/net/sunrpc/xprtrdma/svc_rdma_backchannel.c b/net/sunrpc/xprtrdma/svc_rdma_backchannel.c
index e5a78b761012..549c0c39a97a 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_backchannel.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_backchannel.c
@@ -90,7 +90,7 @@ static int svc_rdma_bc_sendto(struct svcxprt_rdma *rdma,
*/
get_page(virt_to_page(rqst->rq_buffer));
sctxt->sc_send_wr.opcode = IB_WR_SEND;
- return svc_rdma_post_send(rdma, sctxt);
+ return svc_rdma_post_send(rdma, sctxt, NULL);
}
/* Server-side transport endpoint wants a whole page for its send
diff --git a/net/sunrpc/xprtrdma/svc_rdma_sendto.c b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
index c09659b17351..47391cc9d750 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_sendto.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_sendto.c
@@ -222,6 +222,7 @@ struct svc_rdma_send_ctxt *svc_rdma_send_ctxt_get(struct svcxprt_rdma *rdma)
ctxt->sc_page_count = 0;
ctxt->sc_wr_chain = &ctxt->sc_send_wr;
ctxt->sc_sqecount = 1;
+ ctxt->sc_pos = 0;
return ctxt;
@@ -471,6 +472,8 @@ static void svc_rdma_wc_send(struct ib_cq *cq, struct ib_wc *wc)
goto flushed;
trace_svcrdma_wc_send(&ctxt->sc_cid);
+ if (ctxt->sc_pos > atomic64_read(&rdma->sc_xprt.xpt_acked_pos))
+ atomic64_set(&rdma->sc_xprt.xpt_acked_pos, ctxt->sc_pos);
svc_rdma_send_ctxt_put(rdma, ctxt);
return;
@@ -487,23 +490,29 @@ static void svc_rdma_wc_send(struct ib_cq *cq, struct ib_wc *wc)
* svc_rdma_post_send - Post a WR chain to the Send Queue
* @rdma: transport context
* @ctxt: WR chain to post
+ * @pos: OUT: position of this Send on @rdma, or NULL
*
* Copy fields in @ctxt to stack variables in order to guarantee
* that these values remain available after the ib_post_send() call.
* In some error flow cases, svc_rdma_wc_send() releases @ctxt.
*
+ * @pos is written only when the chain was posted. Positions on one
+ * transport increase in posting order, and the Send completion of
+ * @ctxt publishes its position in xpt_acked_pos.
+ *
* Return values:
* %0: @ctxt's WR chain was posted successfully
* %-ENOTCONN: The connection was lost
*/
int svc_rdma_post_send(struct svcxprt_rdma *rdma,
- struct svc_rdma_send_ctxt *ctxt)
+ struct svc_rdma_send_ctxt *ctxt, u64 *pos)
{
struct ib_send_wr *first_wr = ctxt->sc_wr_chain;
struct ib_send_wr *send_wr = &ctxt->sc_send_wr;
const struct ib_send_wr *bad_wr = first_wr;
struct rpc_rdma_cid cid = ctxt->sc_cid;
int ret, sqecount = ctxt->sc_sqecount;
+ u64 seq;
might_sleep();
@@ -518,10 +527,22 @@ int svc_rdma_post_send(struct svcxprt_rdma *rdma,
return ret;
trace_svcrdma_post_send(ctxt);
+
+ /*
+ * Assign the position and post under one lock so positions
+ * match the order the provider queues the Sends, and thus
+ * completion order.
+ */
+ spin_lock(&rdma->sc_post_lock);
+ seq = ++rdma->sc_post_seq;
+ ctxt->sc_pos = seq;
ret = ib_post_send(rdma->sc_qp, first_wr, &bad_wr);
+ spin_unlock(&rdma->sc_post_lock);
if (ret)
return svc_rdma_post_send_err(rdma, &cid, bad_wr,
first_wr, sqecount, ret);
+ if (pos)
+ *pos = seq;
return 0;
}
@@ -1052,7 +1073,7 @@ static int svc_rdma_send_reply_msg(struct svcxprt_rdma *rdma,
send_wr->opcode = IB_WR_SEND;
}
- return svc_rdma_post_send(rdma, sctxt);
+ return svc_rdma_post_send(rdma, sctxt, &rqstp->rq_reply_pos);
}
/**
@@ -1122,7 +1143,7 @@ void svc_rdma_send_error_msg(struct svcxprt_rdma *rdma,
*/
sctxt->sc_wr_chain = &sctxt->sc_send_wr;
sctxt->sc_sqecount = 1;
- if (svc_rdma_post_send(rdma, sctxt))
+ if (svc_rdma_post_send(rdma, sctxt, NULL))
goto put_ctxt;
return;
diff --git a/net/sunrpc/xprtrdma/svc_rdma_transport.c b/net/sunrpc/xprtrdma/svc_rdma_transport.c
index f949601b2144..84042688bb3d 100644
--- a/net/sunrpc/xprtrdma/svc_rdma_transport.c
+++ b/net/sunrpc/xprtrdma/svc_rdma_transport.c
@@ -207,6 +207,7 @@ static struct svcxprt_rdma *svc_rdma_create_xprt(struct svc_serv *serv,
lockdep_set_class(&cma_xprt->sc_send_lock, &svcrdma_sctx_lock);
spin_lock_init(&cma_xprt->sc_rw_ctxt_lock);
lockdep_set_class(&cma_xprt->sc_rw_ctxt_lock, &svcrdma_rwctx_lock);
+ spin_lock_init(&cma_xprt->sc_post_lock);
/*
* Note that this implies that the underlying transport support
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 10/11] NFSD: Evict acknowledged DRC entries
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
` (8 preceding siblings ...)
2026-09-18 17:21 ` [PATCH v5 09/11] svcrdma: Publish RDMA " Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 17:21 ` [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
10 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
A duplicate reply cache entry exists to answer a retransmit. Once
the client has received the reply, no retransmit of that XID can
follow, yet the entry stays in the cache until RC_EXPIRE elapses or
memory pressure forces it out. On a busy server the cache is mostly
replies the client already has.
Record each reply's transport position in its cache entry from the
sv_reply_sent hook, and compare it with the transport's acknowledged
position when a later request prunes the same bucket. An entry whose
position the peer has acknowledged is evicted before the age and
pressure checks run. Only entries that arrived on the current
request's transport are compared: the prune holds a reference on
that transport alone, so its acknowledged position is the only one
safe to read. The shrinker passes no transport and never evicts on
this basis. The walk still stops at the first entry that has neither
expired nor been acknowledged, so an acknowledged entry behind one
from another transport, or behind a UDP or TLS entry, stays until
that entry expires. Only the entries the existing walk reaches are
evicted here; a per-transport list of unacknowledged replies would
reach the rest.
A transport acknowledgment means the peer's transport holds the
reply, not that the RPC layer above it has read the reply. Two
retransmits can therefore find the entry already gone. A client that
discards unread socket data when it resets a connection retransmits
the call over the new connection. A client whose RPC timer fires
after the reply was acknowledged but before its RPC layer read it
retransmits the call on the same connection. In the second case the
call is executed again, but TCP delivers the first reply ahead of
the second, so the client completes the call with the first reply
and discards the second as an unknown XID. The re-execution is
visible only to a third party that reused the name in between. A
call still in progress when its retransmit arrives has sent no
reply, so its entry is not acknowledged and the retransmit is
dropped as before. Retention therefore ends at delivery rather than
at consumption; only NFSv4.1 sessions close that gap.
Between nfsd_cache_update() and the hook the entry is pinned against
eviction, and the nfsd thread keeps its pointer in thread-local
state. Every nfsd_dispatch() path that leaves an entry in RC_DONE
returns 1, so svc_process() runs the hook whether it sends the reply
or drops it after svc_authorise() fails, and the pin is always
cleared. The drop and encode-error paths update with RC_NOCACHE,
which frees the entry without pinning it.
A reply that never reached the wire, or that went out over UDP,
keeps position zero and is retained for the full RC_EXPIRE interval.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/cache.h | 6 ++++-
fs/nfsd/nfscache.c | 68 +++++++++++++++++++++++++++++++++++++++++++++++++++---
fs/nfsd/nfsd.h | 3 +++
fs/nfsd/nfssvc.c | 1 +
fs/nfsd/trace.h | 1 +
5 files changed, 75 insertions(+), 4 deletions(-)
diff --git a/fs/nfsd/cache.h b/fs/nfsd/cache.h
index 589ffe4762d5..894e0b61cbfc 100644
--- a/fs/nfsd/cache.h
+++ b/fs/nfsd/cache.h
@@ -36,8 +36,11 @@ struct nfsd_cacherep {
struct list_head c_lru;
unsigned char c_state, /* unused, inprog, done */
c_type, /* status, buffer */
- c_secure : 1; /* req came from port < 1024 */
+ c_secure : 1, /* req came from port < 1024 */
+ c_inflight : 1; /* reply cached but not yet sent */
u64 c_xprt; /* svc_xprt that carried req */
+ u64 c_pos; /* reply's position on c_xprt;
+ * 0 = never acknowledged */
unsigned long c_timestamp;
union {
struct kvec u_vec;
@@ -88,6 +91,7 @@ int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
unsigned int len, struct nfsd_cacherep **cacherep);
void nfsd_cache_update(struct svc_rqst *rqstp, struct nfsd_cacherep *rp,
int cachetype, __be32 *statp);
+void nfsd_cache_reply_sent(struct svc_rqst *rqstp);
int nfsd_reply_cache_stats_show(struct seq_file *m, void *v);
#endif /* NFSCACHE_H */
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index cd1011bd148d..41b8098d3528 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -113,6 +113,8 @@ nfsd_cacherep_alloc(struct svc_rqst *rqstp, __wsum csum,
rp->c_key.k_len = rqstp->rq_arg.len;
rp->c_key.k_csum = csum;
rp->c_xprt = rqstp->rq_xprt->xpt_id;
+ rp->c_inflight = 0;
+ rp->c_pos = 0;
}
return rp;
}
@@ -259,13 +261,34 @@ nfsd_cache_bucket_find(__be32 xid, struct nfsd_net *nn)
return &nn->drc_hashtbl[hash];
}
+static bool
+nfsd_cacherep_acked(const struct nfsd_cacherep *rp,
+ const struct svc_xprt *xprt)
+{
+ return rp->c_state == RC_DONE && rp->c_pos && xprt &&
+ rp->c_xprt == xprt->xpt_id &&
+ rp->c_pos <= atomic64_read(&xprt->xpt_acked_pos);
+}
+
/*
* Remove and return no more than @max expired entries in bucket @b.
* If @max is zero, do not limit the number of removed entries.
+ *
+ * @xprt is the transport of the current request, or NULL when the
+ * caller holds no transport. Only entries that arrived on @xprt can
+ * be evicted as acknowledged: the caller holds a reference on that
+ * transport alone, so it is the only xpt_acked_pos safe to read.
+ * The walk stops at the first entry that has neither expired nor
+ * been acknowledged, so an acknowledged entry behind it stays until
+ * that entry expires.
+ *
+ * An in-flight entry is never evicted: its svc_rqst still points to
+ * it and nfsd_cache_reply_sent() will write to it.
*/
static void
nfsd_prune_bucket_locked(struct nfsd_net *nn, struct nfsd_drc_bucket *b,
- unsigned int max, struct list_head *dispose)
+ unsigned int max, struct list_head *dispose,
+ struct svc_xprt *xprt)
{
unsigned long expiry = jiffies - RC_EXPIRE;
struct nfsd_cacherep *rp, *tmp;
@@ -275,6 +298,12 @@ nfsd_prune_bucket_locked(struct nfsd_net *nn, struct nfsd_drc_bucket *b,
/* The bucket LRU is ordered oldest-first. */
list_for_each_entry_safe(rp, tmp, &b->lru_head, c_lru) {
+ if (rp->c_inflight)
+ continue;
+ if (nfsd_cacherep_acked(rp, xprt)) {
+ trace_nfsd_drc_evict_acked(nn, rp);
+ goto evict;
+ }
if (atomic_read(&nn->num_drc_entries) > nn->max_drc_entries) {
trace_nfsd_drc_evict_pressure(nn, rp);
goto evict;
@@ -338,7 +367,7 @@ nfsd_reply_cache_scan(struct shrinker *shrink, struct shrink_control *sc)
continue;
spin_lock(&b->cache_lock);
- nfsd_prune_bucket_locked(nn, b, 0, &dispose);
+ nfsd_prune_bucket_locked(nn, b, 0, &dispose, NULL);
spin_unlock(&b->cache_lock);
freed += nfsd_cacherep_dispose(&dispose);
@@ -512,7 +541,7 @@ int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
goto found_entry;
*cacherep = rp;
rp->c_state = RC_INPROG;
- nfsd_prune_bucket_locked(nn, b, 3, &dispose);
+ nfsd_prune_bucket_locked(nn, b, 3, &dispose, rqstp->rq_xprt);
spin_unlock(&b->cache_lock);
nfsd_cacherep_dispose(&dispose);
@@ -590,6 +619,7 @@ void nfsd_cache_update(struct svc_rqst *rqstp, struct nfsd_cacherep *rp,
int cachetype, __be32 *statp)
{
struct nfsd_net *nn = net_generic(SVC_NET(rqstp), nfsd_net_id);
+ struct nfsd_thread_local_info *ntli = rqstp->rq_private;
struct kvec *resv = &rqstp->rq_res.head[0], *cachv;
struct nfsd_drc_bucket *b;
int len;
@@ -636,10 +666,42 @@ void nfsd_cache_update(struct svc_rqst *rqstp, struct nfsd_cacherep *rp,
rp->c_secure = test_bit(RQ_SECURE, &rqstp->rq_flags);
rp->c_type = cachetype;
rp->c_state = RC_DONE;
+ rp->c_inflight = 1;
+ ntli->ntli_cacherep = rp;
spin_unlock(&b->cache_lock);
return;
}
+/**
+ * nfsd_cache_reply_sent - record a reply's transport position
+ * @rqstp: RPC transaction whose reply phase has just ended
+ *
+ * Installed as the svc_serv's sv_reply_sent hook. A transport that
+ * does not publish positions (UDP), or a reply that svc_process()
+ * dropped after dispatch, leaves rq_reply_pos at zero, and the entry
+ * is then never evicted as acknowledged.
+ *
+ * Context: nfsd thread context. Takes and releases cache_lock.
+ */
+void nfsd_cache_reply_sent(struct svc_rqst *rqstp)
+{
+ struct nfsd_thread_local_info *ntli = rqstp->rq_private;
+ struct nfsd_cacherep *rp = ntli->ntli_cacherep;
+ struct nfsd_drc_bucket *b;
+ struct nfsd_net *nn;
+
+ if (!rp)
+ return;
+ ntli->ntli_cacherep = NULL;
+
+ nn = net_generic(SVC_NET(rqstp), nfsd_net_id);
+ b = nfsd_cache_bucket_find(rp->c_key.k_xid, nn);
+ spin_lock(&b->cache_lock);
+ rp->c_pos = rqstp->rq_reply_pos;
+ rp->c_inflight = 0;
+ spin_unlock(&b->cache_lock);
+}
+
static int
nfsd_cache_append(struct svc_rqst *rqstp, struct kvec *data)
{
diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h
index a145294c59c8..0ae78ed00da4 100644
--- a/fs/nfsd/nfsd.h
+++ b/fs/nfsd/nfsd.h
@@ -53,9 +53,12 @@ extern atomic_t nfsd_th_cnt; /* number of available threads */
extern const struct seq_operations nfs_exports_op;
+struct nfsd_cacherep;
+
struct nfsd_thread_local_info {
struct nfs4_client **ntli_lease_breaker;
int ntli_cachetype;
+ struct nfsd_cacherep *ntli_cacherep;
};
/*
diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c
index c04ef9d180ce..6fc54399cb90 100644
--- a/fs/nfsd/nfssvc.c
+++ b/fs/nfsd/nfssvc.c
@@ -634,6 +634,7 @@ int nfsd_create_serv(struct net *net)
percpu_ref_exit(&nn->nfsd_net_ref);
return -ENOMEM;
}
+ serv->sv_reply_sent = nfsd_cache_reply_sent;
error = svc_bind(serv, net);
if (error < 0) {
diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h
index 631682a76f9c..7abd46e8a752 100644
--- a/fs/nfsd/trace.h
+++ b/fs/nfsd/trace.h
@@ -1596,6 +1596,7 @@ DEFINE_EVENT(nfsd_drc_entry_class, nfsd_drc_##name, \
DEFINE_NFSD_DRC_ENTRY_EVENT(evict_pressure);
DEFINE_NFSD_DRC_ENTRY_EVENT(evict_expired);
+DEFINE_NFSD_DRC_ENTRY_EVENT(evict_acked);
TRACE_EVENT(nfsd_cb_args,
TP_PROTO(
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
` (9 preceding siblings ...)
2026-09-18 17:21 ` [PATCH v5 10/11] NFSD: Evict acknowledged DRC entries Chuck Lever
@ 2026-09-18 17:21 ` Chuck Lever
2026-09-18 23:39 ` NeilBrown
10 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
Cc: Rick Macklem, linux-nfs, Chuck Lever
nfsd_cache_csum() costs CPU on every non-idempotent request. It also
prevents the use of zero-copy RDMA receives for WRITE and SYMLINK
because the payload must be in the server's memory before the DRC
lookup can run.
Commit 01a7decf7593 ("nfsd: keep a checksum of the first 256 bytes
of request") added the checksum when a growing cache made XID
collisions easier to hit. It is the only guard against a reused
XID: two calls with the same procedure and equal-length arguments
match on every other key field.
Acknowledgment-based eviction retires a TCP or RDMA entry once the
transport confirms delivery and a later request from the same
connection visits its bucket. Three kinds of entry outlast that: the
entries a connection has issued since it last visited each bucket,
entries from a connection the client has since closed, which no
later request can evict because their transport is gone, and UDP
entries. Those wait out RC_EXPIRE. A collision needs the same XID
from the same address and port within that window. The Linux NFS
client's XIDs are sequential from a seed drawn with
get_random_u32(), so an XID recurs only after 2^32 calls, and a
rebooted client does not replay its previous sequence. A client that
restarts its sequence from a fixed value and rebinds its previous
source port within RC_EXPIRE is unguarded, but a client that repeats
a live XID cannot reliably match replies to its own calls anyway.
An acknowledged entry stays in its bucket until the next prune
visits it, and a lookup matches it in the meantime. The client
already holds that reply, so treat a call carrying its XID as a
miss: evict the entry and insert the new one in its place.
Remove the checksum, the nfsd_drc_mismatch tracepoint, and the
payload_misses stat that counted checksum-detected collisions.
nfsd_dispatch() no longer snapshots the argument stream before
decoding, and with k_csum gone, struct nfsd_cacherep shrinks to 136
bytes.
The "payload misses" line disappears from
/proc/fs/nfsd/reply_cache_stats. No known userspace tool parses it.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
.../ABI/testing/procfs-nfsd-reply_cache_stats | 11 +--
fs/nfsd/cache.h | 9 +-
fs/nfsd/netns.h | 2 -
fs/nfsd/nfscache.c | 103 +++++----------------
fs/nfsd/nfssvc.c | 10 +-
fs/nfsd/stats.h | 5 -
fs/nfsd/trace.h | 24 -----
7 files changed, 28 insertions(+), 136 deletions(-)
diff --git a/Documentation/ABI/testing/procfs-nfsd-reply_cache_stats b/Documentation/ABI/testing/procfs-nfsd-reply_cache_stats
index 57ed5f8e6597..7a22ac7c1ccd 100644
--- a/Documentation/ABI/testing/procfs-nfsd-reply_cache_stats
+++ b/Documentation/ABI/testing/procfs-nfsd-reply_cache_stats
@@ -19,19 +19,16 @@ Description:
cache misses s64 Requests not found in cache
not cached s64 Idempotent requests that
bypass the cache
- payload misses s64 XID matched but request
- checksum did not
longest chain len u32 Longest hash chain observed
cachesize at longest u32 Cache size when longest
chain was recorded
======================= ====== ==========================
Counter fields (cache hits, cache misses, not cached,
- payload misses, mem usage) are maintained with per-cpu
- counters and may briefly show stale values under
- concurrent load. There is no way to reset these
- counters; consumers should compute rates by sampling
- over time.
+ mem usage) are maintained with per-cpu counters and
+ may briefly show stale values under concurrent load.
+ There is no way to reset these counters; consumers
+ should compute rates by sampling over time.
New fields may be appended in future kernels. Parsers
should match on field name, not line position.
diff --git a/fs/nfsd/cache.h b/fs/nfsd/cache.h
index 894e0b61cbfc..3491a10524d7 100644
--- a/fs/nfsd/cache.h
+++ b/fs/nfsd/cache.h
@@ -22,9 +22,7 @@ struct nfsd_net;
*/
struct nfsd_cacherep {
struct {
- /* Keep often-read xid, csum in the same cache line: */
__be32 k_xid;
- __wsum k_csum;
u32 k_proc;
u32 k_prot;
u32 k_vers;
@@ -80,15 +78,12 @@ enum {
/* Cache entries expire after this time period */
#define RC_EXPIRE (120 * HZ)
-/* Checksum this amount of the request */
-#define RC_CSUMLEN (256U)
-
int nfsd_drc_slab_create(void);
void nfsd_drc_slab_free(void);
int nfsd_reply_cache_init(struct nfsd_net *);
void nfsd_reply_cache_shutdown(struct nfsd_net *);
-int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
- unsigned int len, struct nfsd_cacherep **cacherep);
+int nfsd_cache_lookup(struct svc_rqst *rqstp,
+ struct nfsd_cacherep **cacherep);
void nfsd_cache_update(struct svc_rqst *rqstp, struct nfsd_cacherep *rp,
int cachetype, __be32 *statp);
void nfsd_cache_reply_sent(struct svc_rqst *rqstp);
diff --git a/fs/nfsd/netns.h b/fs/nfsd/netns.h
index 0ce7da20aba3..30231b9027dd 100644
--- a/fs/nfsd/netns.h
+++ b/fs/nfsd/netns.h
@@ -39,8 +39,6 @@ enum nfsd_net_flag {
};
enum {
- /* cache misses due only to checksum comparison failures */
- NFSD_STATS_PAYLOAD_MISSES,
/* amount of memory (in bytes) currently consumed by the DRC */
NFSD_STATS_DRC_MEM_USAGE,
NFSD_STATS_RC_HITS, /* repcache hits */
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index 41b8098d3528..f5bc29cd8eee 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -13,10 +13,8 @@
#include <linux/slab.h>
#include <linux/vmalloc.h>
#include <linux/sunrpc/addr.h>
-#include <linux/highmem.h>
#include <linux/log2.h>
#include <linux/hash.h>
-#include <net/checksum.h>
#include "nfsd.h"
#include "nfserr.h"
@@ -91,8 +89,7 @@ nfsd_hashsize(unsigned int limit)
}
static struct nfsd_cacherep *
-nfsd_cacherep_alloc(struct svc_rqst *rqstp, __wsum csum,
- struct nfsd_net *nn)
+nfsd_cacherep_alloc(struct svc_rqst *rqstp, struct nfsd_net *nn)
{
struct nfsd_cacherep *rp;
@@ -111,7 +108,6 @@ nfsd_cacherep_alloc(struct svc_rqst *rqstp, __wsum csum,
rp->c_key.k_prot = rqstp->rq_prot;
rp->c_key.k_vers = rqstp->rq_vers;
rp->c_key.k_len = rqstp->rq_arg.len;
- rp->c_key.k_csum = csum;
rp->c_xprt = rqstp->rq_xprt->xpt_id;
rp->c_inflight = 0;
rp->c_pos = 0;
@@ -377,68 +373,10 @@ nfsd_reply_cache_scan(struct shrinker *shrink, struct shrink_control *sc)
return freed;
}
-/**
- * nfsd_cache_csum - Checksum incoming NFS Call arguments
- * @buf: buffer containing a whole RPC Call message
- * @start: starting byte of the NFS Call header
- * @remaining: size of the NFS Call header, in bytes
- *
- * Compute a weak checksum of the leading bytes of an NFS procedure
- * call header to help verify that a retransmitted Call matches an
- * entry in the duplicate reply cache.
- *
- * To avoid assumptions about how the RPC message is laid out in
- * @buf and what else it might contain (eg, a GSS MIC suffix), the
- * caller passes us the exact location and length of the NFS Call
- * header.
- *
- * Returns a 32-bit checksum value, as defined in RFC 793.
- */
-static __wsum nfsd_cache_csum(struct xdr_buf *buf, unsigned int start,
- unsigned int remaining)
-{
- unsigned int base, len;
- struct xdr_buf subbuf;
- __wsum csum = 0;
- void *p;
- int idx;
-
- if (remaining > RC_CSUMLEN)
- remaining = RC_CSUMLEN;
- if (xdr_buf_subsegment(buf, &subbuf, start, remaining))
- return csum;
-
- /* rq_arg.head first */
- if (subbuf.head[0].iov_len) {
- len = min_t(unsigned int, subbuf.head[0].iov_len, remaining);
- csum = csum_partial(subbuf.head[0].iov_base, len, csum);
- remaining -= len;
- }
-
- /* Continue into page array */
- idx = subbuf.page_base / PAGE_SIZE;
- base = subbuf.page_base & ~PAGE_MASK;
- while (remaining) {
- p = page_address(subbuf.pages[idx]) + base;
- len = min_t(unsigned int, PAGE_SIZE - base, remaining);
- csum = csum_partial(p, len, csum);
- remaining -= len;
- base = 0;
- ++idx;
- }
- return csum;
-}
-
static int
nfsd_cache_key_cmp(const struct nfsd_cacherep *key,
- const struct nfsd_cacherep *rp, struct nfsd_net *nn)
+ const struct nfsd_cacherep *rp)
{
- if (key->c_key.k_xid == rp->c_key.k_xid &&
- key->c_key.k_csum != rp->c_key.k_csum) {
- nfsd_stats_payload_misses_inc(nn);
- trace_nfsd_drc_mismatch(nn, key, rp);
- }
-
return memcmp(&key->c_key, &rp->c_key, sizeof(key->c_key));
}
@@ -462,7 +400,7 @@ nfsd_cache_insert(struct nfsd_drc_bucket *b, struct nfsd_cacherep *key,
parent = *p;
rp = rb_entry(parent, struct nfsd_cacherep, c_node);
- cmp = nfsd_cache_key_cmp(key, rp, nn);
+ cmp = nfsd_cache_key_cmp(key, rp);
if (cmp < 0)
p = &parent->rb_left;
else if (cmp > 0)
@@ -491,28 +429,23 @@ nfsd_cache_insert(struct nfsd_drc_bucket *b, struct nfsd_cacherep *key,
/**
* nfsd_cache_lookup - Find an entry in the duplicate reply cache
* @rqstp: Incoming Call to find
- * @start: starting byte in @rqstp->rq_arg of the NFS Call header
- * @len: size of the NFS Call header, in bytes
* @cacherep: OUT: DRC entry for this request
*
- * Try to find an entry matching the current call in the cache. When none
- * is found, we try to grab the oldest expired entry off the LRU list. If
- * a suitable one isn't there, then drop the cache_lock and allocate a
- * new one, then search again in case one got inserted while this thread
- * didn't hold the lock.
+ * On a miss, the entry created for this call is returned in
+ * @cacherep. On a hit, the cached reply is encoded into @rqstp's
+ * response.
*
* Return values:
* %RC_DOIT: Process the request normally
* %RC_REPLY: Reply from cache
* %RC_DROPIT: Do not process the request further
*/
-int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
- unsigned int len, struct nfsd_cacherep **cacherep)
+int nfsd_cache_lookup(struct svc_rqst *rqstp,
+ struct nfsd_cacherep **cacherep)
{
struct nfsd_net *nn = net_generic(SVC_NET(rqstp), nfsd_net_id);
struct nfsd_thread_local_info *ntli = rqstp->rq_private;
struct nfsd_cacherep *rp, *found;
- __wsum csum;
struct nfsd_drc_bucket *b;
int type = ntli->ntli_cachetype;
LIST_HEAD(dispose);
@@ -524,21 +457,29 @@ int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
goto out;
}
- csum = nfsd_cache_csum(&rqstp->rq_arg, start, len);
-
/*
* Since the common case is a cache miss followed by an insert,
* preallocate an entry.
*/
- rp = nfsd_cacherep_alloc(rqstp, csum, nn);
+ rp = nfsd_cacherep_alloc(rqstp, nn);
if (!rp)
goto out;
b = nfsd_cache_bucket_find(rqstp->rq_xid, nn);
spin_lock(&b->cache_lock);
found = nfsd_cache_insert(b, rp, nn);
- if (found != rp)
- goto found_entry;
+ if (found != rp) {
+ /*
+ * The client already holds the reply for an acknowledged
+ * entry, so a call carrying its XID is a new call.
+ */
+ if (!nfsd_cacherep_acked(found, rqstp->rq_xprt))
+ goto found_entry;
+ trace_nfsd_drc_evict_acked(nn, found);
+ nfsd_cacherep_unlink_locked(nn, b, found);
+ list_add(&found->c_lru, &dispose);
+ nfsd_cache_insert(b, rp, nn);
+ }
*cacherep = rp;
rp->c_state = RC_INPROG;
nfsd_prune_bucket_locked(nn, b, 3, &dispose, rqstp->rq_xprt);
@@ -737,8 +678,6 @@ int nfsd_reply_cache_stats_show(struct seq_file *m, void *v)
percpu_counter_sum_positive(&nn->counter[NFSD_STATS_RC_MISSES]));
seq_printf(m, "not cached: %lld\n",
percpu_counter_sum_positive(&nn->counter[NFSD_STATS_RC_NOCACHE]));
- seq_printf(m, "payload misses: %lld\n",
- percpu_counter_sum_positive(&nn->counter[NFSD_STATS_PAYLOAD_MISSES]));
seq_printf(m, "longest chain len: %u\n", nn->longest_chain);
seq_printf(m, "cachesize at longest: %u\n", nn->longest_chain_cachesize);
return 0;
diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c
index 6fc54399cb90..07e5e8549f6d 100644
--- a/fs/nfsd/nfssvc.c
+++ b/fs/nfsd/nfssvc.c
@@ -1005,7 +1005,6 @@ int nfsd_dispatch(struct svc_rqst *rqstp)
const struct svc_procedure *proc = rqstp->rq_procinfo;
__be32 *statp = rqstp->rq_accept_statp;
struct nfsd_cacherep *rp;
- unsigned int start, len;
__be32 *nfs_reply;
/*
@@ -1014,13 +1013,6 @@ int nfsd_dispatch(struct svc_rqst *rqstp)
*/
ntli->ntli_cachetype = proc->pc_cachetype;
- /*
- * ->pc_decode advances the argument stream past the NFS
- * Call header, so grab the header's starting location and
- * size now for the call to nfsd_cache_lookup().
- */
- start = xdr_stream_pos(&rqstp->rq_arg_stream);
- len = xdr_stream_remaining(&rqstp->rq_arg_stream);
if (!proc->pc_decode(rqstp, &rqstp->rq_arg_stream))
goto out_decode_err;
@@ -1034,7 +1026,7 @@ int nfsd_dispatch(struct svc_rqst *rqstp)
smp_store_release(&rqstp->rq_status_counter, rqstp->rq_status_counter | 1);
rp = NULL;
- switch (nfsd_cache_lookup(rqstp, start, len, &rp)) {
+ switch (nfsd_cache_lookup(rqstp, &rp)) {
case RC_DOIT:
break;
case RC_REPLY:
diff --git a/fs/nfsd/stats.h b/fs/nfsd/stats.h
index aabfbb1a9c71..4a556dfbf64a 100644
--- a/fs/nfsd/stats.h
+++ b/fs/nfsd/stats.h
@@ -97,11 +97,6 @@ static inline void nfsd_stats_io_write_add(struct nfsd_net *nn,
amount);
}
-static inline void nfsd_stats_payload_misses_inc(struct nfsd_net *nn)
-{
- percpu_counter_inc(&nn->counter[NFSD_STATS_PAYLOAD_MISSES]);
-}
-
/**
* nfsd_stats_drc_mem_usage_add - Add memory used by a cache item
* @nn: target network namespace
diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h
index 7abd46e8a752..1febb42a008f 100644
--- a/fs/nfsd/trace.h
+++ b/fs/nfsd/trace.h
@@ -1536,30 +1536,6 @@ TRACE_EVENT(nfsd_drc_found,
);
-TRACE_EVENT(nfsd_drc_mismatch,
- TP_PROTO(
- const struct nfsd_net *nn,
- const struct nfsd_cacherep *key,
- const struct nfsd_cacherep *rp
- ),
- TP_ARGS(nn, key, rp),
- TP_STRUCT__entry(
- __field(unsigned long long, boot_time)
- __field(u32, xid)
- __field(u32, cached)
- __field(u32, ingress)
- ),
- TP_fast_assign(
- __entry->boot_time = nn->boot_time;
- __entry->xid = be32_to_cpu(key->c_key.k_xid);
- __entry->cached = (__force u32)key->c_key.k_csum;
- __entry->ingress = (__force u32)rp->c_key.k_csum;
- ),
- TP_printk("boot_time=%16llx xid=0x%08x cached-csum=0x%08x ingress-csum=0x%08x",
- __entry->boot_time, __entry->xid, __entry->cached,
- __entry->ingress)
-);
-
DECLARE_EVENT_CLASS(nfsd_drc_entry_class,
TP_PROTO(
const struct nfsd_net *nn,
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* Re: [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat
2026-09-18 17:21 ` [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
@ 2026-09-18 23:39 ` NeilBrown
2026-09-19 16:11 ` Chuck Lever
0 siblings, 1 reply; 19+ messages in thread
From: NeilBrown @ 2026-09-18 23:39 UTC (permalink / raw)
To: Chuck Lever
Cc: Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey, Rick Macklem,
linux-nfs, Chuck Lever
On Sat, 19 Sep 2026, Chuck Lever wrote:
> nfsd_cache_csum() costs CPU on every non-idempotent request. It also
> prevents the use of zero-copy RDMA receives for WRITE and SYMLINK
> because the payload must be in the server's memory before the DRC
> lookup can run.
>
> Commit 01a7decf7593 ("nfsd: keep a checksum of the first 256 bytes
> of request") added the checksum when a growing cache made XID
> collisions easier to hit. It is the only guard against a reused
> XID: two calls with the same procedure and equal-length arguments
> match on every other key field.
>
> Acknowledgment-based eviction retires a TCP or RDMA entry once the
> transport confirms delivery and a later request from the same
> connection visits its bucket. Three kinds of entry outlast that: the
> entries a connection has issued since it last visited each bucket,
> entries from a connection the client has since closed, which no
> later request can evict because their transport is gone, and UDP
> entries. Those wait out RC_EXPIRE. A collision needs the same XID
> from the same address and port within that window. The Linux NFS
> client's XIDs are sequential from a seed drawn with
> get_random_u32(), so an XID recurs only after 2^32 calls, and a
> rebooted client does not replay its previous sequence. A client that
> restarts its sequence from a fixed value and rebinds its previous
> source port within RC_EXPIRE is unguarded, but a client that repeats
> a live XID cannot reliably match replies to its own calls anyway.
>
> An acknowledged entry stays in its bucket until the next prune
> visits it, and a lookup matches it in the meantime. The client
> already holds that reply, so treat a call carrying its XID as a
> miss: evict the entry and insert the new one in its place.
I don't follow the logic here.
If we get a request on a particular transport and it is a good match
for a request that we recently replied to on the same transport, then
the chance of it being a timed-out retransmit is 100%. A client that
re-uses xids that quickly is buggy, isn't it?
So if nfsd_cacherep_acked() then I think we can ignore the request.
In fact if the "found" reply was sent on the same xprt that we received
a new request on, I think we can ignore the new request even if the
reply hasn't been acked yet.
Everything else in the series is good and I really appreciate how
thoroughly you have explained the logic in a lot of places.
Thanks,
NeilBrown
>
> Remove the checksum, the nfsd_drc_mismatch tracepoint, and the
> payload_misses stat that counted checksum-detected collisions.
> nfsd_dispatch() no longer snapshots the argument stream before
> decoding, and with k_csum gone, struct nfsd_cacherep shrinks to 136
> bytes.
>
> The "payload misses" line disappears from
> /proc/fs/nfsd/reply_cache_stats. No known userspace tool parses it.
>
> Assisted-by: LLM
> Signed-off-by: Chuck Lever <cel@kernel.org>
> ---
> .../ABI/testing/procfs-nfsd-reply_cache_stats | 11 +--
> fs/nfsd/cache.h | 9 +-
> fs/nfsd/netns.h | 2 -
> fs/nfsd/nfscache.c | 103 +++++----------------
> fs/nfsd/nfssvc.c | 10 +-
> fs/nfsd/stats.h | 5 -
> fs/nfsd/trace.h | 24 -----
> 7 files changed, 28 insertions(+), 136 deletions(-)
>
> diff --git a/Documentation/ABI/testing/procfs-nfsd-reply_cache_stats b/Documentation/ABI/testing/procfs-nfsd-reply_cache_stats
> index 57ed5f8e6597..7a22ac7c1ccd 100644
> --- a/Documentation/ABI/testing/procfs-nfsd-reply_cache_stats
> +++ b/Documentation/ABI/testing/procfs-nfsd-reply_cache_stats
> @@ -19,19 +19,16 @@ Description:
> cache misses s64 Requests not found in cache
> not cached s64 Idempotent requests that
> bypass the cache
> - payload misses s64 XID matched but request
> - checksum did not
> longest chain len u32 Longest hash chain observed
> cachesize at longest u32 Cache size when longest
> chain was recorded
> ======================= ====== ==========================
>
> Counter fields (cache hits, cache misses, not cached,
> - payload misses, mem usage) are maintained with per-cpu
> - counters and may briefly show stale values under
> - concurrent load. There is no way to reset these
> - counters; consumers should compute rates by sampling
> - over time.
> + mem usage) are maintained with per-cpu counters and
> + may briefly show stale values under concurrent load.
> + There is no way to reset these counters; consumers
> + should compute rates by sampling over time.
>
> New fields may be appended in future kernels. Parsers
> should match on field name, not line position.
> diff --git a/fs/nfsd/cache.h b/fs/nfsd/cache.h
> index 894e0b61cbfc..3491a10524d7 100644
> --- a/fs/nfsd/cache.h
> +++ b/fs/nfsd/cache.h
> @@ -22,9 +22,7 @@ struct nfsd_net;
> */
> struct nfsd_cacherep {
> struct {
> - /* Keep often-read xid, csum in the same cache line: */
> __be32 k_xid;
> - __wsum k_csum;
> u32 k_proc;
> u32 k_prot;
> u32 k_vers;
> @@ -80,15 +78,12 @@ enum {
> /* Cache entries expire after this time period */
> #define RC_EXPIRE (120 * HZ)
>
> -/* Checksum this amount of the request */
> -#define RC_CSUMLEN (256U)
> -
> int nfsd_drc_slab_create(void);
> void nfsd_drc_slab_free(void);
> int nfsd_reply_cache_init(struct nfsd_net *);
> void nfsd_reply_cache_shutdown(struct nfsd_net *);
> -int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
> - unsigned int len, struct nfsd_cacherep **cacherep);
> +int nfsd_cache_lookup(struct svc_rqst *rqstp,
> + struct nfsd_cacherep **cacherep);
> void nfsd_cache_update(struct svc_rqst *rqstp, struct nfsd_cacherep *rp,
> int cachetype, __be32 *statp);
> void nfsd_cache_reply_sent(struct svc_rqst *rqstp);
> diff --git a/fs/nfsd/netns.h b/fs/nfsd/netns.h
> index 0ce7da20aba3..30231b9027dd 100644
> --- a/fs/nfsd/netns.h
> +++ b/fs/nfsd/netns.h
> @@ -39,8 +39,6 @@ enum nfsd_net_flag {
> };
>
> enum {
> - /* cache misses due only to checksum comparison failures */
> - NFSD_STATS_PAYLOAD_MISSES,
> /* amount of memory (in bytes) currently consumed by the DRC */
> NFSD_STATS_DRC_MEM_USAGE,
> NFSD_STATS_RC_HITS, /* repcache hits */
> diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
> index 41b8098d3528..f5bc29cd8eee 100644
> --- a/fs/nfsd/nfscache.c
> +++ b/fs/nfsd/nfscache.c
> @@ -13,10 +13,8 @@
> #include <linux/slab.h>
> #include <linux/vmalloc.h>
> #include <linux/sunrpc/addr.h>
> -#include <linux/highmem.h>
> #include <linux/log2.h>
> #include <linux/hash.h>
> -#include <net/checksum.h>
>
> #include "nfsd.h"
> #include "nfserr.h"
> @@ -91,8 +89,7 @@ nfsd_hashsize(unsigned int limit)
> }
>
> static struct nfsd_cacherep *
> -nfsd_cacherep_alloc(struct svc_rqst *rqstp, __wsum csum,
> - struct nfsd_net *nn)
> +nfsd_cacherep_alloc(struct svc_rqst *rqstp, struct nfsd_net *nn)
> {
> struct nfsd_cacherep *rp;
>
> @@ -111,7 +108,6 @@ nfsd_cacherep_alloc(struct svc_rqst *rqstp, __wsum csum,
> rp->c_key.k_prot = rqstp->rq_prot;
> rp->c_key.k_vers = rqstp->rq_vers;
> rp->c_key.k_len = rqstp->rq_arg.len;
> - rp->c_key.k_csum = csum;
> rp->c_xprt = rqstp->rq_xprt->xpt_id;
> rp->c_inflight = 0;
> rp->c_pos = 0;
> @@ -377,68 +373,10 @@ nfsd_reply_cache_scan(struct shrinker *shrink, struct shrink_control *sc)
> return freed;
> }
>
> -/**
> - * nfsd_cache_csum - Checksum incoming NFS Call arguments
> - * @buf: buffer containing a whole RPC Call message
> - * @start: starting byte of the NFS Call header
> - * @remaining: size of the NFS Call header, in bytes
> - *
> - * Compute a weak checksum of the leading bytes of an NFS procedure
> - * call header to help verify that a retransmitted Call matches an
> - * entry in the duplicate reply cache.
> - *
> - * To avoid assumptions about how the RPC message is laid out in
> - * @buf and what else it might contain (eg, a GSS MIC suffix), the
> - * caller passes us the exact location and length of the NFS Call
> - * header.
> - *
> - * Returns a 32-bit checksum value, as defined in RFC 793.
> - */
> -static __wsum nfsd_cache_csum(struct xdr_buf *buf, unsigned int start,
> - unsigned int remaining)
> -{
> - unsigned int base, len;
> - struct xdr_buf subbuf;
> - __wsum csum = 0;
> - void *p;
> - int idx;
> -
> - if (remaining > RC_CSUMLEN)
> - remaining = RC_CSUMLEN;
> - if (xdr_buf_subsegment(buf, &subbuf, start, remaining))
> - return csum;
> -
> - /* rq_arg.head first */
> - if (subbuf.head[0].iov_len) {
> - len = min_t(unsigned int, subbuf.head[0].iov_len, remaining);
> - csum = csum_partial(subbuf.head[0].iov_base, len, csum);
> - remaining -= len;
> - }
> -
> - /* Continue into page array */
> - idx = subbuf.page_base / PAGE_SIZE;
> - base = subbuf.page_base & ~PAGE_MASK;
> - while (remaining) {
> - p = page_address(subbuf.pages[idx]) + base;
> - len = min_t(unsigned int, PAGE_SIZE - base, remaining);
> - csum = csum_partial(p, len, csum);
> - remaining -= len;
> - base = 0;
> - ++idx;
> - }
> - return csum;
> -}
> -
> static int
> nfsd_cache_key_cmp(const struct nfsd_cacherep *key,
> - const struct nfsd_cacherep *rp, struct nfsd_net *nn)
> + const struct nfsd_cacherep *rp)
> {
> - if (key->c_key.k_xid == rp->c_key.k_xid &&
> - key->c_key.k_csum != rp->c_key.k_csum) {
> - nfsd_stats_payload_misses_inc(nn);
> - trace_nfsd_drc_mismatch(nn, key, rp);
> - }
> -
> return memcmp(&key->c_key, &rp->c_key, sizeof(key->c_key));
> }
>
> @@ -462,7 +400,7 @@ nfsd_cache_insert(struct nfsd_drc_bucket *b, struct nfsd_cacherep *key,
> parent = *p;
> rp = rb_entry(parent, struct nfsd_cacherep, c_node);
>
> - cmp = nfsd_cache_key_cmp(key, rp, nn);
> + cmp = nfsd_cache_key_cmp(key, rp);
> if (cmp < 0)
> p = &parent->rb_left;
> else if (cmp > 0)
> @@ -491,28 +429,23 @@ nfsd_cache_insert(struct nfsd_drc_bucket *b, struct nfsd_cacherep *key,
> /**
> * nfsd_cache_lookup - Find an entry in the duplicate reply cache
> * @rqstp: Incoming Call to find
> - * @start: starting byte in @rqstp->rq_arg of the NFS Call header
> - * @len: size of the NFS Call header, in bytes
> * @cacherep: OUT: DRC entry for this request
> *
> - * Try to find an entry matching the current call in the cache. When none
> - * is found, we try to grab the oldest expired entry off the LRU list. If
> - * a suitable one isn't there, then drop the cache_lock and allocate a
> - * new one, then search again in case one got inserted while this thread
> - * didn't hold the lock.
> + * On a miss, the entry created for this call is returned in
> + * @cacherep. On a hit, the cached reply is encoded into @rqstp's
> + * response.
> *
> * Return values:
> * %RC_DOIT: Process the request normally
> * %RC_REPLY: Reply from cache
> * %RC_DROPIT: Do not process the request further
> */
> -int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
> - unsigned int len, struct nfsd_cacherep **cacherep)
> +int nfsd_cache_lookup(struct svc_rqst *rqstp,
> + struct nfsd_cacherep **cacherep)
> {
> struct nfsd_net *nn = net_generic(SVC_NET(rqstp), nfsd_net_id);
> struct nfsd_thread_local_info *ntli = rqstp->rq_private;
> struct nfsd_cacherep *rp, *found;
> - __wsum csum;
> struct nfsd_drc_bucket *b;
> int type = ntli->ntli_cachetype;
> LIST_HEAD(dispose);
> @@ -524,21 +457,29 @@ int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
> goto out;
> }
>
> - csum = nfsd_cache_csum(&rqstp->rq_arg, start, len);
> -
> /*
> * Since the common case is a cache miss followed by an insert,
> * preallocate an entry.
> */
> - rp = nfsd_cacherep_alloc(rqstp, csum, nn);
> + rp = nfsd_cacherep_alloc(rqstp, nn);
> if (!rp)
> goto out;
>
> b = nfsd_cache_bucket_find(rqstp->rq_xid, nn);
> spin_lock(&b->cache_lock);
> found = nfsd_cache_insert(b, rp, nn);
> - if (found != rp)
> - goto found_entry;
> + if (found != rp) {
> + /*
> + * The client already holds the reply for an acknowledged
> + * entry, so a call carrying its XID is a new call.
> + */
> + if (!nfsd_cacherep_acked(found, rqstp->rq_xprt))
> + goto found_entry;
> + trace_nfsd_drc_evict_acked(nn, found);
> + nfsd_cacherep_unlink_locked(nn, b, found);
> + list_add(&found->c_lru, &dispose);
> + nfsd_cache_insert(b, rp, nn);
> + }
> *cacherep = rp;
> rp->c_state = RC_INPROG;
> nfsd_prune_bucket_locked(nn, b, 3, &dispose, rqstp->rq_xprt);
> @@ -737,8 +678,6 @@ int nfsd_reply_cache_stats_show(struct seq_file *m, void *v)
> percpu_counter_sum_positive(&nn->counter[NFSD_STATS_RC_MISSES]));
> seq_printf(m, "not cached: %lld\n",
> percpu_counter_sum_positive(&nn->counter[NFSD_STATS_RC_NOCACHE]));
> - seq_printf(m, "payload misses: %lld\n",
> - percpu_counter_sum_positive(&nn->counter[NFSD_STATS_PAYLOAD_MISSES]));
> seq_printf(m, "longest chain len: %u\n", nn->longest_chain);
> seq_printf(m, "cachesize at longest: %u\n", nn->longest_chain_cachesize);
> return 0;
> diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c
> index 6fc54399cb90..07e5e8549f6d 100644
> --- a/fs/nfsd/nfssvc.c
> +++ b/fs/nfsd/nfssvc.c
> @@ -1005,7 +1005,6 @@ int nfsd_dispatch(struct svc_rqst *rqstp)
> const struct svc_procedure *proc = rqstp->rq_procinfo;
> __be32 *statp = rqstp->rq_accept_statp;
> struct nfsd_cacherep *rp;
> - unsigned int start, len;
> __be32 *nfs_reply;
>
> /*
> @@ -1014,13 +1013,6 @@ int nfsd_dispatch(struct svc_rqst *rqstp)
> */
> ntli->ntli_cachetype = proc->pc_cachetype;
>
> - /*
> - * ->pc_decode advances the argument stream past the NFS
> - * Call header, so grab the header's starting location and
> - * size now for the call to nfsd_cache_lookup().
> - */
> - start = xdr_stream_pos(&rqstp->rq_arg_stream);
> - len = xdr_stream_remaining(&rqstp->rq_arg_stream);
> if (!proc->pc_decode(rqstp, &rqstp->rq_arg_stream))
> goto out_decode_err;
>
> @@ -1034,7 +1026,7 @@ int nfsd_dispatch(struct svc_rqst *rqstp)
> smp_store_release(&rqstp->rq_status_counter, rqstp->rq_status_counter | 1);
>
> rp = NULL;
> - switch (nfsd_cache_lookup(rqstp, start, len, &rp)) {
> + switch (nfsd_cache_lookup(rqstp, &rp)) {
> case RC_DOIT:
> break;
> case RC_REPLY:
> diff --git a/fs/nfsd/stats.h b/fs/nfsd/stats.h
> index aabfbb1a9c71..4a556dfbf64a 100644
> --- a/fs/nfsd/stats.h
> +++ b/fs/nfsd/stats.h
> @@ -97,11 +97,6 @@ static inline void nfsd_stats_io_write_add(struct nfsd_net *nn,
> amount);
> }
>
> -static inline void nfsd_stats_payload_misses_inc(struct nfsd_net *nn)
> -{
> - percpu_counter_inc(&nn->counter[NFSD_STATS_PAYLOAD_MISSES]);
> -}
> -
> /**
> * nfsd_stats_drc_mem_usage_add - Add memory used by a cache item
> * @nn: target network namespace
> diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h
> index 7abd46e8a752..1febb42a008f 100644
> --- a/fs/nfsd/trace.h
> +++ b/fs/nfsd/trace.h
> @@ -1536,30 +1536,6 @@ TRACE_EVENT(nfsd_drc_found,
>
> );
>
> -TRACE_EVENT(nfsd_drc_mismatch,
> - TP_PROTO(
> - const struct nfsd_net *nn,
> - const struct nfsd_cacherep *key,
> - const struct nfsd_cacherep *rp
> - ),
> - TP_ARGS(nn, key, rp),
> - TP_STRUCT__entry(
> - __field(unsigned long long, boot_time)
> - __field(u32, xid)
> - __field(u32, cached)
> - __field(u32, ingress)
> - ),
> - TP_fast_assign(
> - __entry->boot_time = nn->boot_time;
> - __entry->xid = be32_to_cpu(key->c_key.k_xid);
> - __entry->cached = (__force u32)key->c_key.k_csum;
> - __entry->ingress = (__force u32)rp->c_key.k_csum;
> - ),
> - TP_printk("boot_time=%16llx xid=0x%08x cached-csum=0x%08x ingress-csum=0x%08x",
> - __entry->boot_time, __entry->xid, __entry->cached,
> - __entry->ingress)
> -);
> -
> DECLARE_EVENT_CLASS(nfsd_drc_entry_class,
> TP_PROTO(
> const struct nfsd_net *nn,
>
> --
> 2.55.0
>
>
>
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat
2026-09-18 23:39 ` NeilBrown
@ 2026-09-19 16:11 ` Chuck Lever
2026-09-20 10:37 ` NeilBrown
0 siblings, 1 reply; 19+ messages in thread
From: Chuck Lever @ 2026-09-19 16:11 UTC (permalink / raw)
To: NeilBrown
Cc: Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey, Rick Macklem,
linux-nfs
On Fri, Sep 18, 2026, at 7:39 PM, NeilBrown wrote:
> If we get a request on a particular transport and it is a good match
> for a request that we recently replied to on the same transport, then
> the chance of it being a timed-out retransmit is 100%. A client that
> re-uses xids that quickly is buggy, isn't it?
> So if nfsd_cacherep_acked() then I think we can ignore the request.
> In fact if the "found" reply was sent on the same xprt that we received
> a new request on, I think we can ignore the new request even if the
> reply hasn't been acked yet.
Just to be clear, by "ignore" do you mean "drop the request" or "reply
with the cached reply" ?
> Everything else in the series is good and I really appreciate how
> thoroughly you have explained the logic in a lot of places.
Thanks, your comments have made this a much better implementation.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat
2026-09-19 16:11 ` Chuck Lever
@ 2026-09-20 10:37 ` NeilBrown
2026-09-20 17:49 ` Chuck Lever
0 siblings, 1 reply; 19+ messages in thread
From: NeilBrown @ 2026-09-20 10:37 UTC (permalink / raw)
To: Chuck Lever
Cc: Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey, Rick Macklem,
linux-nfs
On Sun, 20 Sep 2026, Chuck Lever wrote:
>
> On Fri, Sep 18, 2026, at 7:39 PM, NeilBrown wrote:
>
> > If we get a request on a particular transport and it is a good match
> > for a request that we recently replied to on the same transport, then
> > the chance of it being a timed-out retransmit is 100%. A client that
> > re-uses xids that quickly is buggy, isn't it?
> > So if nfsd_cacherep_acked() then I think we can ignore the request.
> > In fact if the "found" reply was sent on the same xprt that we received
> > a new request on, I think we can ignore the new request even if the
> > reply hasn't been acked yet.
>
> Just to be clear, by "ignore" do you mean "drop the request" or "reply
> with the cached reply" ?
I mean "RC_DROPIT". If the connection we sent the reply on is still
active - active enough to get requests - and we haven't seen the ack
yet, then in a sense the request is in "RC_INPROG" and so RC_DROPIT
is appropriate.
Even if we *have* seen the ack, I think RC_DROPIT is appropriate.
NeilBrown
>
>
> > Everything else in the series is good and I really appreciate how
> > thoroughly you have explained the logic in a lot of places.
>
> Thanks, your comments have made this a much better implementation.
>
> --
> Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
>
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat
2026-09-20 10:37 ` NeilBrown
@ 2026-09-20 17:49 ` Chuck Lever
0 siblings, 0 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-20 17:49 UTC (permalink / raw)
To: NeilBrown
Cc: Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey, Rick Macklem,
linux-nfs
On Sun, Sep 20, 2026, at 6:37 AM, NeilBrown wrote:
> On Sun, 20 Sep 2026, Chuck Lever wrote:
>> Just to be clear, by "ignore" do you mean "drop the request" or "reply
>> with the cached reply" ?
>
> I mean "RC_DROPIT". If the connection we sent the reply on is still
> active - active enough to get requests - and we haven't seen the ack
> yet, then in a sense the request is in "RC_INPROG" and so RC_DROPIT
> is appropriate.
>
> Even if we *have* seen the ack, I think RC_DROPIT is appropriate.
I'll remove the "treat an acknowledged match as a new call" block,
then. A match whose entry was sent on the same connection now
returns RC_DROPIT whether or not the reply has been acknowledged.
The test is a recorded reply position plus a matching xpt_id, so
UDP entries, which carry no position, are still replayed, and so is
an entry whose reply svc_process() dropped.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 19+ messages in thread