From: Chuck Lever <cel@kernel.org>
To: Jeff Layton <jlayton@kernel.org>, NeilBrown <neil@brown.name>,
Olga Kornievskaia <okorniev@redhat.com>,
Dai Ngo <Dai.Ngo@oracle.com>, Tom Talpey <tom@talpey.com>
Cc: Rick Macklem <rmacklem@uoguelph.ca>,
linux-nfs@vger.kernel.org, Chuck Lever <cel@kernel.org>
Subject: [PATCH v3 07/12] SUNRPC: Add TCP sequence-number ACK tracking for reply delivery
Date: Thu, 10 Sep 2026 09:54:47 -0400 [thread overview]
Message-ID: <20260910-duplicate-reply-cache-v3-7-31532a4c7449@kernel.org> (raw)
In-Reply-To: <20260910-duplicate-reply-cache-v3-0-31532a4c7449@kernel.org>
A DRC entry for a reply sent over TCP lingers until RC_EXPIRE even
though the peer's TCP acknowledgment has already confirmed that the
reply arrived.
The acknowledgment confirms only that the reply reached the peer's
TCP stack. If the connection drops before the NFS client reads it,
the client retransmits the call and the evicted entry cannot answer
it. Memory pressure and RC_EXPIRE already open the same window.
When the upper layer caches a reply it stores a cookie on the
svc_rqst. After svc_tcp_sendto() transmits the reply, record the
cookie with the current write_seq in a fixed-size per-socket ring.
Before appending, drain entries whose sequence number is at or before
snd_una and report each as delivered. The send path is
single-threaded, so the ring needs no locking. sk_wmem_queued cannot
be used for this because it counts per-skb truesize, not payload
bytes.
On a kTLS session the snapshot needs one more check. A short push
inside the TLS layer leaves the rest of the record as
partially_sent_record, and tls_sw_sendmsg() still returns the full
plaintext count, so write_seq can fall short of the reply. Record the
reply only when no partially_sent_record remains after the send.
Some replies get no ring entry: the ring is full, the send fails, the
transport is already dead, or a kTLS record is still partly unsent.
Report each of these as untracked at once, so the upper layer knows no
delivery report will follow. When the socket is freed, report entries
still in the ring as undelivered. Set XPT_REPLY_ACK on TCP transports
so the upper layer expects these reports.
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfscache.c | 15 +++++---
include/linux/sunrpc/svc.h | 3 ++
include/linux/sunrpc/svcsock.h | 10 +++++
net/sunrpc/svc.c | 1 +
net/sunrpc/svcsock.c | 85 ++++++++++++++++++++++++++++++++++++++++++
5 files changed, 108 insertions(+), 6 deletions(-)
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index b0231c659237..473bf825c37a 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -289,12 +289,14 @@ nfsd_cache_bucket_find(__be32 xid, struct nfsd_net *nn)
}
/*
- * The generation match keeps a stale cookie from acknowledging a
- * later entry that reuses the same XID and transport. The walk
- * starts at the MRU end of the bucket LRU, where the entry for a
- * just-sent reply sits. The visit cap bounds the time spent under
- * cache_lock when a client packs the bucket; an entry missed under
- * the cap is left to the other eviction reasons.
+ * The generation match keeps a stale cookie from acknowledging a later
+ * entry that reuses the same XID and transport.
+ *
+ * Average bucket occupancy is small (TARGET_BUCKET_SIZE), so a linear walk
+ * of the bucket LRU beats an rb-tree lookup. The walk starts at the MRU
+ * end, where the entry for a just-sent reply sits. The visit cap bounds
+ * the time spent under cache_lock when a client packs the bucket; an entry
+ * missed under the cap is left to the other eviction reasons.
*/
static void nfsd_reply_ack(void *data, const svc_ack_cookie_t *cookie,
bool delivered)
@@ -723,6 +725,7 @@ void nfsd_cache_update(struct svc_rqst *rqstp, struct nfsd_cacherep *rp,
rp->c_ack_gen = nfsd_cache_next_ack_gen();
rp->c_type = cachetype;
rp->c_state = RC_DONE;
+ rqstp->rq_ack_cookie = nfsd_cache_ack_cookie(rp);
spin_unlock(&b->cache_lock);
return;
}
diff --git a/include/linux/sunrpc/svc.h b/include/linux/sunrpc/svc.h
index 8c9e27752698..da3822c54f3e 100644
--- a/include/linux/sunrpc/svc.h
+++ b/include/linux/sunrpc/svc.h
@@ -315,6 +315,9 @@ struct svc_rqst {
unsigned int bc_to_retries;
unsigned int rq_status_counter; /* RPC processing counter */
void *rq_private; /* For use by the service thread */
+
+ svc_ack_cookie_t rq_ack_cookie; /* zero when no reply-ack
+ * tracking is requested */
};
/* bits for rq_flags */
diff --git a/include/linux/sunrpc/svcsock.h b/include/linux/sunrpc/svcsock.h
index 372a00882ca6..595be631b68b 100644
--- a/include/linux/sunrpc/svcsock.h
+++ b/include/linux/sunrpc/svcsock.h
@@ -41,6 +41,16 @@ struct svc_sock {
struct page_frag_cache sk_frag_cache;
+ /* TCP reply-ACK tracking (single-threaded send path) */
+#define SVC_ACK_RING_BITS 6
+#define SVC_ACK_RING_SIZE (1 << SVC_ACK_RING_BITS)
+ struct {
+ u32 ae_pos; /* write_seq after send */
+ svc_ack_cookie_t ae_cookie; /* opaque DRC cookie */
+ } sk_ack_ring[SVC_ACK_RING_SIZE];
+ unsigned int sk_ack_head;
+ unsigned int sk_ack_tail;
+
struct completion sk_handshake_done;
/* received data */
diff --git a/net/sunrpc/svc.c b/net/sunrpc/svc.c
index f73412e123a1..33dd9cbaff42 100644
--- a/net/sunrpc/svc.c
+++ b/net/sunrpc/svc.c
@@ -1488,6 +1488,7 @@ svc_process_common(struct svc_rqst *rqstp)
/* Reset the accept_stat for the RPC */
rqstp->rq_accept_statp = NULL;
+ rqstp->rq_ack_cookie = (svc_ack_cookie_t){};
/* Will be turned off only when NFSv4 Sessions are used */
set_bit(RQ_USEDEFERRAL, &rqstp->rq_flags);
diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c
index 625aebbbc6b3..a28905181589 100644
--- a/net/sunrpc/svcsock.c
+++ b/net/sunrpc/svcsock.c
@@ -28,6 +28,7 @@
#include <linux/file.h>
#include <linux/freezer.h>
#include <linux/bvec.h>
+#include <linux/circ_buf.h>
#include <net/sock.h>
#include <net/checksum.h>
@@ -36,6 +37,7 @@
#include <net/udp.h>
#include <net/tcp.h>
#include <net/tcp_states.h>
+#include <net/tls.h>
#include <net/tls_prot.h>
#include <net/handshake.h>
#include <linux/uaccess.h>
@@ -88,6 +90,7 @@ static void svc_sock_free(struct svc_xprt *);
static struct svc_xprt *svc_create_socket(struct svc_serv *, int,
struct net *, struct sockaddr *,
int, int);
+
#ifdef CONFIG_DEBUG_LOCK_ALLOC
static struct lock_class_key svc_key[2];
static struct lock_class_key svc_slock_key[2];
@@ -379,6 +382,45 @@ static void svc_data_ready(struct sock *sk)
}
}
+/*
+ * Report as delivered each recorded reply that the peer's cumulative ACK
+ * now covers.
+ */
+static void svc_tcp_ack_drain(struct svc_sock *svsk)
+{
+ struct svc_serv *serv = svsk->sk_xprt.xpt_server;
+ u32 snd_una = READ_ONCE(tcp_sk(svsk->sk_sk)->snd_una);
+
+ while (svsk->sk_ack_head != svsk->sk_ack_tail) {
+ unsigned int idx = svsk->sk_ack_tail &
+ (SVC_ACK_RING_SIZE - 1);
+
+ if (after(svsk->sk_ack_ring[idx].ae_pos, snd_una))
+ break;
+ svc_reply_acked(serv,
+ &svsk->sk_ack_ring[idx].ae_cookie, true);
+ svsk->sk_ack_tail++;
+ }
+}
+
+/*
+ * Once the socket is freed, acknowledgments for replies still in the ring
+ * can no longer be observed.
+ */
+static void svc_tcp_ack_purge(struct svc_sock *svsk)
+{
+ struct svc_serv *serv = svsk->sk_xprt.xpt_server;
+
+ while (svsk->sk_ack_head != svsk->sk_ack_tail) {
+ unsigned int idx = svsk->sk_ack_tail &
+ (SVC_ACK_RING_SIZE - 1);
+
+ svc_reply_acked(serv,
+ &svsk->sk_ack_ring[idx].ae_cookie, false);
+ svsk->sk_ack_tail++;
+ }
+}
+
/*
* INET callback when space is newly available on the socket.
*/
@@ -1392,6 +1434,19 @@ static int svc_tcp_sendmsg(struct svc_sock *svsk, struct svc_rqst *rqstp,
return ret;
}
+/*
+ * tls_sw_sendmsg() can return the full plaintext count with part of a
+ * record still waiting for socket write space, leaving write_seq short
+ * of the reply. The TLS layer keeps that record as partially_sent_record
+ * until it is pushed.
+ */
+static bool svc_tcp_reply_queued(struct svc_sock *svsk)
+{
+ if (!test_bit(XPT_TLS_SESSION, &svsk->sk_xprt.xpt_flags))
+ return true;
+ return !READ_ONCE(tls_get_ctx(svsk->sk_sk)->partially_sent_record);
+}
+
/**
* svc_tcp_sendto - Send out a reply on a TCP socket
* @rqstp: completed svc_rqst
@@ -1420,10 +1475,32 @@ static int svc_tcp_sendto(struct svc_rqst *rqstp)
trace_svcsock_tcp_send(xprt, sent);
if (sent < 0 || sent != (xdr->len + sizeof(marker)))
goto out_close;
+
+ svc_tcp_ack_drain(svsk);
+ if (svc_ack_cookie_present(&rqstp->rq_ack_cookie)) {
+ if (svc_tcp_reply_queued(svsk) &&
+ CIRC_SPACE(svsk->sk_ack_head, svsk->sk_ack_tail,
+ SVC_ACK_RING_SIZE) > 0) {
+ unsigned int idx = svsk->sk_ack_head &
+ (SVC_ACK_RING_SIZE - 1);
+
+ svsk->sk_ack_ring[idx].ae_pos =
+ tcp_sk(svsk->sk_sk)->write_seq;
+ svsk->sk_ack_ring[idx].ae_cookie =
+ rqstp->rq_ack_cookie;
+ svsk->sk_ack_head++;
+ } else {
+ svc_reply_acked(xprt->xpt_server, &rqstp->rq_ack_cookie,
+ false);
+ }
+ }
+
mutex_unlock(&xprt->xpt_mutex);
return sent;
out_notconn:
+ if (svc_ack_cookie_present(&rqstp->rq_ack_cookie))
+ svc_reply_acked(xprt->xpt_server, &rqstp->rq_ack_cookie, false);
mutex_unlock(&xprt->xpt_mutex);
return -ENOTCONN;
out_close:
@@ -1431,6 +1508,8 @@ static int svc_tcp_sendto(struct svc_rqst *rqstp)
xprt->xpt_server->sv_name,
(sent < 0) ? "got error" : "sent",
sent, xdr->len + sizeof(marker));
+ if (svc_ack_cookie_present(&rqstp->rq_ack_cookie))
+ svc_reply_acked(xprt->xpt_server, &rqstp->rq_ack_cookie, false);
svc_xprt_deferred_close(xprt);
mutex_unlock(&xprt->xpt_mutex);
return -EAGAIN;
@@ -1487,6 +1566,7 @@ static bool svc_tcp_init(struct svc_sock *svsk, struct svc_serv *serv)
return false;
set_bit(XPT_CACHE_AUTH, &svsk->sk_xprt.xpt_flags);
set_bit(XPT_CONG_CTRL, &svsk->sk_xprt.xpt_flags);
+ set_bit(XPT_REPLY_ACK, &svsk->sk_xprt.xpt_flags);
if (sk->sk_state == TCP_LISTEN) {
strcpy(svsk->sk_xprt.xpt_remotebuf, "listener");
set_bit(XPT_LISTENER, &svsk->sk_xprt.xpt_flags);
@@ -1593,6 +1673,9 @@ static struct svc_sock *svc_setup_socket(struct svc_serv *serv,
}
}
+ svsk->sk_ack_head = 0;
+ svsk->sk_ack_tail = 0;
+
svsk->sk_sock = sock;
svsk->sk_sk = inet;
svsk->sk_ostate = inet->sk_state_change;
@@ -1810,6 +1893,8 @@ static void svc_sock_free(struct svc_xprt *xprt)
trace_svcsock_free(svsk, sock);
+ svc_tcp_ack_purge(svsk);
+
tls_handshake_cancel(sock->sk);
if (sock->file)
sockfd_put(sock);
--
2.55.0
next prev parent reply other threads:[~2026-09-10 13:55 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 13:54 [PATCH v3 00/12] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-10 13:54 ` [PATCH v3 01/12] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
2026-09-10 13:54 ` [PATCH v3 02/12] NFSD: Track transport in DRC entries Chuck Lever
2026-09-10 13:54 ` [PATCH v3 03/12] NFSD: Prepare bucket pruning for out-of-order eviction Chuck Lever
2026-09-10 13:54 ` [PATCH v3 04/12] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
2026-09-10 13:54 ` [PATCH v3 05/12] NFSD: Record DRC population in lookup tracepoints Chuck Lever
2026-09-10 13:54 ` [PATCH v3 06/12] NFSD: Add reply-acknowledged callback infrastructure Chuck Lever
2026-09-10 13:54 ` Chuck Lever [this message]
2026-09-10 13:54 ` [PATCH v3 08/12] svcrdma: Fire reply-acknowledged callback on Send completion Chuck Lever
2026-09-10 13:54 ` [PATCH v3 09/12] SUNRPC: Record last-request timestamp on svc_xprt Chuck Lever
2026-09-10 13:54 ` [PATCH v3 10/12] NFSD: Evict unacknowledged DRC entries via implied ACK Chuck Lever
2026-09-10 13:54 ` [PATCH v3 11/12] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
2026-09-10 13:54 ` [PATCH v3 12/12] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
2026-09-10 17:25 ` [PATCH v3 00/12] Improve the scalability of NFSD's classic DRC Jeff Layton
2026-09-10 23:02 ` NeilBrown
2026-09-11 14:42 ` Chuck Lever
2026-09-11 23:20 ` NeilBrown
2026-09-12 16:22 ` Chuck Lever
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260910-duplicate-reply-cache-v3-7-31532a4c7449@kernel.org \
--to=cel@kernel.org \
--cc=Dai.Ngo@oracle.com \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=neil@brown.name \
--cc=okorniev@redhat.com \
--cc=rmacklem@uoguelph.ca \
--cc=tom@talpey.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.