From: Chuck Lever <cel@kernel.org>
To: Jeff Layton <jlayton@kernel.org>, NeilBrown <neil@brown.name>,
Olga Kornievskaia <okorniev@redhat.com>,
Dai Ngo <Dai.Ngo@oracle.com>, Tom Talpey <tom@talpey.com>
Cc: Rick Macklem <rmacklem@uoguelph.ca>,
linux-nfs@vger.kernel.org, Chuck Lever <cel@kernel.org>
Subject: [PATCH v6 08/12] SUNRPC: Publish reply positions for upper-layer consumers
Date: Mon, 21 Sep 2026 09:22:34 -0400 [thread overview]
Message-ID: <20260921-duplicate-reply-cache-v6-8-db5e13fd9944@kernel.org> (raw)
In-Reply-To: <20260921-duplicate-reply-cache-v6-0-db5e13fd9944@kernel.org>
An RPC server's duplicate reply cache has to hold a cached reply
until the client has received it. SUNRPC gives an upper layer no way
to learn that, so a cache entry lives for the full retention
interval even when the reply was delivered at once.
Add two transport-neutral values that an upper layer can compare:
the position of each reply on its transport, recorded in the svc_rqst
by xpo_sendto(), and the position up to which the peer has
acknowledged, recorded on the svc_xprt. The units are private to the
transport (bytes for a stream, Send sequence for RDMA); a consumer
learns only whether one is less than or equal to the other. A reply
that xpo_sendto() does not hand to the wire keeps position zero and
is never reported as acknowledged.
The acknowledged position is an atomic64_t. A transport publishes it
outside any lock the consumer holds, and a torn 64-bit read on a
32-bit host could run ahead of the peer.
Add a per-service hook that runs once per request when its reply
phase ends, whether svc_send() sent the reply or svc_process()
dropped it, so an upper layer can record the reply's position
against its own state either way. On the drop path the position is
zero.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
include/linux/sunrpc/svc.h | 16 ++++++++++++++++
include/linux/sunrpc/svc_xprt.h | 2 ++
net/sunrpc/svc.c | 3 +++
net/sunrpc/svc_xprt.c | 2 ++
4 files changed, 23 insertions(+)
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index 24698856eb40..01036e093012 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -59,6 +59,8 @@ enum {
};
+struct svc_rqst;
+
/*
* RPC service.
*
@@ -96,6 +98,13 @@ struct svc_serv {
* connection */
bool sv_bc_enabled; /* service uses backchannel */
#endif /* CONFIG_SUNRPC_BACKCHANNEL */
+
+ /*
+ * Called once per request after its reply phase, whether
+ * xpo_sendto() succeeded, failed, or was never reached because
+ * svc_process() dropped the reply. See rq_reply_pos.
+ */
+ void (*sv_reply_sent)(struct svc_rqst *rqstp);
};
/* This is used by pool_stats to find and lock an svc */
@@ -267,6 +276,13 @@ struct svc_rqst {
unsigned int bc_to_retries;
unsigned int rq_status_counter; /* RPC processing counter */
void *rq_private; /* For use by the service thread */
+
+ /*
+ * Transport position of this request's reply, set by
+ * xpo_sendto(); zero when the reply did not reach the wire.
+ * Comparable only with the xpt_acked_pos of the same svc_xprt.
+ */
+ u64 rq_reply_pos;
};
/* bits for rq_flags */
diff --git a/include/linux/sunrpc/svc_xprt.h b/include/linux/sunrpc/svc_xprt.h
index 52082d8ea6dd..7176c42f19d7 100644
--- a/include/linux/sunrpc/svc_xprt.h
+++ b/include/linux/sunrpc/svc_xprt.h
@@ -66,6 +66,8 @@ struct svc_xprt {
atomic_t xpt_reserved; /* outq space rsvd, UDP only */
atomic_t xpt_nr_rqsts; /* Number of requests */
struct mutex xpt_mutex; /* to serialize sending data */
+ atomic64_t xpt_acked_pos; /* replies up to this position
+ * are acknowledged; 0 = none */
spinlock_t xpt_lock; /* protects sk_deferred
* and xpt_auth_cache */
void *xpt_auth_cache;/* auth cache */
diff --git a/net/sunrpc/svc.c b/net/sunrpc/svc.c
index f73412e123a1..9ef0661bb422 100644
--- a/net/sunrpc/svc.c
+++ b/net/sunrpc/svc.c
@@ -1488,6 +1488,7 @@ svc_process_common(struct svc_rqst *rqstp)
/* Reset the accept_stat for the RPC */
rqstp->rq_accept_statp = NULL;
+ rqstp->rq_reply_pos = 0;
/* Will be turned off only when NFSv4 Sessions are used */
set_bit(RQ_USEDEFERRAL, &rqstp->rq_flags);
@@ -1733,6 +1734,8 @@ void svc_process(struct svc_rqst *rqstp)
goto out_baddir;
if (!svc_process_common(rqstp)) {
+ if (rqstp->rq_server->sv_reply_sent)
+ rqstp->rq_server->sv_reply_sent(rqstp);
svc_release_rqst(rqstp);
goto out_drop;
}
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 463ba8d77f0c..1668d96ac9ff 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -1028,6 +1028,8 @@ void svc_send(struct svc_rqst *rqstp)
trace_svc_stats_latency(rqstp);
status = xprt->xpt_ops->xpo_sendto(rqstp);
+ if (xprt->xpt_server->sv_reply_sent)
+ xprt->xpt_server->sv_reply_sent(rqstp);
trace_svc_send(rqstp, status);
}
--
2.55.0
next prev parent reply other threads:[~2026-09-21 13:22 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 13:22 [PATCH v6 00/12] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-21 13:22 ` [PATCH v6 01/12] NFSD: Make the DRC size limit independent of page size Chuck Lever
2026-09-21 13:22 ` [PATCH v6 02/12] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
2026-09-21 13:22 ` [PATCH v6 03/12] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
2026-09-21 13:22 ` [PATCH v6 04/12] NFSD: Track transport in DRC entries Chuck Lever
2026-09-21 13:22 ` [PATCH v6 05/12] NFSD: Prepare bucket pruning for additional eviction reasons Chuck Lever
2026-09-21 13:22 ` [PATCH v6 06/12] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
2026-09-21 13:22 ` [PATCH v6 07/12] NFSD: Record DRC population in lookup tracepoints Chuck Lever
2026-09-21 13:22 ` Chuck Lever [this message]
2026-09-21 13:22 ` [PATCH v6 09/12] SUNRPC: Publish TCP reply positions Chuck Lever
2026-09-21 13:22 ` [PATCH v6 10/12] svcrdma: Publish RDMA " Chuck Lever
2026-09-21 13:22 ` [PATCH v6 11/12] NFSD: Evict acknowledged DRC entries Chuck Lever
2026-09-21 13:22 ` [PATCH v6 12/12] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
2026-09-23 14:56 ` Jeff Layton
2026-09-23 16:36 ` Chuck Lever
2026-09-22 6:55 ` [PATCH v6 00/12] Improve the scalability of NFSD's classic DRC NeilBrown
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260921-duplicate-reply-cache-v6-8-db5e13fd9944@kernel.org \
--to=cel@kernel.org \
--cc=Dai.Ngo@oracle.com \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=neil@brown.name \
--cc=okorniev@redhat.com \
--cc=rmacklem@uoguelph.ca \
--cc=tom@talpey.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox