From: Chuck Lever <cel@kernel.org>
To: Jeff Layton <jlayton@kernel.org>, NeilBrown <neil@brown.name>,
Olga Kornievskaia <okorniev@redhat.com>,
Dai Ngo <Dai.Ngo@oracle.com>, Tom Talpey <tom@talpey.com>
Cc: Rick Macklem <rmacklem@uoguelph.ca>,
linux-nfs@vger.kernel.org, Chuck Lever <cel@kernel.org>
Subject: [PATCH v5 08/11] SUNRPC: Publish TCP reply positions
Date: Fri, 18 Sep 2026 13:21:20 -0400 [thread overview]
Message-ID: <20260918-duplicate-reply-cache-v5-8-b6aba9ebf2f4@kernel.org> (raw)
In-Reply-To: <20260918-duplicate-reply-cache-v5-0-b6aba9ebf2f4@kernel.org>
Add the svcsock side of reply-position reporting. A reply's position
is the running count of bytes handed to the socket, kept per socket
under xpt_mutex. A TCP sequence number cannot serve as the position:
the 32-bit space wraps in a few hundred milliseconds at 40 GbE, and
a duplicate reply cache entry lives for two minutes.
The acknowledged position is derived from the socket's unacknowledged
byte count, write_seq minus snd_una, after each successful send. The
two are read without the socket lock, as tcp_ioctl() reads them for
SIOCOUTQ: write_seq advances only under xpt_mutex, and a stale
snd_una only lowers the result. Refreshing the position only in the
send path means a consumer sees the acknowledgment state as of the
previous reply on that connection, stale by one round trip but never
ahead of the peer. A failed or refused send leaves rq_reply_pos at
zero, so the reply is never reported as acknowledged.
Under kTLS the socket counts ciphertext while svcsock counts
plaintext, and tls_sw_sendmsg() can return the full plaintext count
with part of a record still waiting for socket write space. The
unacknowledged count then understates what the peer has yet to
receive. Whether a record is pending is state private to net/tls, so
a TLS session publishes no acknowledged position and its replies
are retained for the full interval.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
include/linux/sunrpc/svcsock.h | 3 +++
net/sunrpc/svcsock.c | 29 +++++++++++++++++++++++++++++
2 files changed, 32 insertions(+)
diff --git a/include/linux/sunrpc/svcsock.h b/include/linux/sunrpc/svcsock.h
index 372a00882ca6..b6e0767de35a 100644
--- a/include/linux/sunrpc/svcsock.h
+++ b/include/linux/sunrpc/svcsock.h
@@ -41,6 +41,9 @@ struct svc_sock {
struct page_frag_cache sk_frag_cache;
+ /* reply bytes handed to the socket; protected by xpt_mutex */
+ u64 sk_send_pos;
+
struct completion sk_handshake_done;
/* received data */
diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c
index e5459d504b6a..b9bab3751e76 100644
--- a/net/sunrpc/svcsock.c
+++ b/net/sunrpc/svcsock.c
@@ -1390,6 +1390,32 @@ static int svc_tcp_sendmsg(struct svc_sock *svsk, struct svc_rqst *rqstp,
return ret;
}
+/*
+ * Bytes the socket has not seen acknowledged all belong to the most
+ * recent replies, so sk_send_pos less that count is the acknowledged
+ * position. write_seq advances only under xpt_mutex, which the caller
+ * holds, and a stale snd_una only lowers the result.
+ *
+ * Under kTLS, tls_sw_sendmsg() can return the full plaintext count
+ * with part of a record still waiting for socket write space, leaving
+ * write_seq short of the reply. Whether a record is pending is private
+ * to net/tls, so a TLS session publishes no acknowledged position.
+ */
+static void svc_tcp_update_acked(struct svc_sock *svsk)
+{
+ struct tcp_sock *tp = tcp_sk(svsk->sk_sk);
+ u64 unacked, acked;
+
+ if (test_bit(XPT_TLS_SESSION, &svsk->sk_xprt.xpt_flags))
+ return;
+ unacked = READ_ONCE(tp->write_seq) - READ_ONCE(tp->snd_una);
+ if (unacked >= svsk->sk_send_pos)
+ return;
+ acked = svsk->sk_send_pos - unacked;
+ if (acked > atomic64_read(&svsk->sk_xprt.xpt_acked_pos))
+ atomic64_set(&svsk->sk_xprt.xpt_acked_pos, acked);
+}
+
/**
* svc_tcp_sendto - Send out a reply on a TCP socket
* @rqstp: completed svc_rqst
@@ -1418,6 +1444,9 @@ static int svc_tcp_sendto(struct svc_rqst *rqstp)
trace_svcsock_tcp_send(xprt, sent);
if (sent < 0 || sent != (xdr->len + sizeof(marker)))
goto out_close;
+ svsk->sk_send_pos += sent;
+ rqstp->rq_reply_pos = svsk->sk_send_pos;
+ svc_tcp_update_acked(svsk);
mutex_unlock(&xprt->xpt_mutex);
return sent;
--
2.55.0
next prev parent reply other threads:[~2026-09-18 17:21 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-18 17:21 ` [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
2026-09-18 23:00 ` NeilBrown
2026-09-20 17:47 ` Chuck Lever
2026-09-20 22:10 ` NeilBrown
2026-09-18 17:21 ` [PATCH v5 02/11] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
2026-09-18 17:21 ` [PATCH v5 03/11] NFSD: Track transport in DRC entries Chuck Lever
2026-09-18 17:21 ` [PATCH v5 04/11] NFSD: Prepare bucket pruning for additional eviction reasons Chuck Lever
2026-09-18 17:21 ` [PATCH v5 05/11] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
2026-09-18 17:21 ` [PATCH v5 06/11] NFSD: Record DRC population in lookup tracepoints Chuck Lever
2026-09-18 17:21 ` [PATCH v5 07/11] SUNRPC: Publish reply positions for upper-layer consumers Chuck Lever
2026-09-18 17:21 ` Chuck Lever [this message]
2026-09-18 17:21 ` [PATCH v5 09/11] svcrdma: Publish RDMA reply positions Chuck Lever
2026-09-18 17:21 ` [PATCH v5 10/11] NFSD: Evict acknowledged DRC entries Chuck Lever
2026-09-18 17:21 ` [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
2026-09-18 23:39 ` NeilBrown
2026-09-19 16:11 ` Chuck Lever
2026-09-20 10:37 ` NeilBrown
2026-09-20 17:49 ` Chuck Lever
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260918-duplicate-reply-cache-v5-8-b6aba9ebf2f4@kernel.org \
--to=cel@kernel.org \
--cc=Dai.Ngo@oracle.com \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=neil@brown.name \
--cc=okorniev@redhat.com \
--cc=rmacklem@uoguelph.ca \
--cc=tom@talpey.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox