From: "Chuck Lever" <cel@kernel.org>
To: NeilBrown <neil@brown.name>
Cc: "Jeff Layton" <jlayton@kernel.org>,
"Olga Kornievskaia" <okorniev@redhat.com>,
"Dai Ngo" <Dai.Ngo@oracle.com>, "Tom Talpey" <tom@talpey.com>,
"Rick Macklem" <rmacklem@uoguelph.ca>,
linux-nfs@vger.kernel.org
Subject: Re: [PATCH v3 00/12] Improve the scalability of NFSD's classic DRC
Date: Sat, 12 Sep 2026 12:22:08 -0400 [thread overview]
Message-ID: <6293796e-8a68-4032-90f3-2adfc1386e55@app.fastmail.com> (raw)
In-Reply-To: <178916883696.207413.18145968545647470391@noble.neil.brown.name>
On Fri, Sep 11, 2026, at 7:20 PM, NeilBrown wrote:
> On Sat, 12 Sep 2026, Chuck Lever wrote:
>> On Fri, 11 Sep 2026, NeilBrown wrote:
>> > On Thu, 10 Sep 2026, Chuck Lever wrote:
>> > > 1. The transport reports delivery outright where it can. TCP now
>> > > reports when snd_una passes the reply's sequence number, and RDMA
>> > > reports on each Send completion.
>> >
>> > As you note in the relevant patch, the TCP ack can arrive before the NFS
>> > client has seen the reply. But I don't think you drop the cached reply
>> > immediately, so that probably don't matter.
>>
>> Right. The ACK marks the entry, and the next prune of that bucket
>> evicts it.
>>
>>
>> > > 2. Where no report will come, a later request on the same connection
>> > > stands in as an implied ACK, since a client on a live TCP or RDMA
>> > > connection does not retransmit within it (RFC 1813 Section 4.5).
>> >
>> > I don't follow this. Why would no report come? TCP will always send
>> > and ACK.
>>
>> The cover letter was unclear. TCP always ACKs, but svc_tcp_sendto()
>> does not always track the reply. The per-socket ring is fixed-size,
>> so a burst overflows it. A kTLS send that leaves part of the record
>> unsent is not recorded, and neither is a failed send. Each of those
>> is reported as untracked right away, so the DRC entry has no report
>> to wait for.
>
> I don't see the need for the per-socket ring.
> When a reply is sent - record the TCP sequence number in the DRC.
> When pruning the DRC, get the TCP ack number first and compare it
> against sequence numbers.
> Do this often enough that that entries are pruned before the sequence
> wraps past them.
>
> kTLS is clearly more subtle but there must be some way record an
> approximate seq number for an incomplete send .. maybe current seq +
> header-size + queued size ??
>
> Unsent replies will never get an ack, but presumably the client will
> resend, be answered using the cached reply, and we can then get an ack
> on the new connection. I think it would be wrong to drop an unsent
> reply before the timeout, unless it gets re-sent and transport-acked
> before then.
>
> I think the opaque cookie aspect of the design is a mistake. Everything
> is fully sequenced over a network connection so a sequence number
> (clearly paired with xpt_id) makes more sense...
> Or does RDMA not sequence replies? If so then it can only ack a single
> reply at a time, not a range of replies. But that needn't be a barrier.
>
> When a message is transmitted the transport hands a seq-number (and
> xpt_id) to the cache. When it gets an ack it hands a range of
> seq-numbers to the cache. For a transport that had out-of-order acks,
> this range would always be of length one (that transport might need a
> private atomic_t to generate the sequence numbers). For a transport
> that batches acks (like TCP) the range could be much longer.
>
> Also, I wonder if the xpt_id should be 64bit assigned sequentially with no
> xa keeping track of used one. 64bits never wraps.
>
> I think that with this design the implied ACK would add no value.
Thanks, all makes sense to me and reduces code complexity.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
prev parent reply other threads:[~2026-09-12 16:22 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 13:54 [PATCH v3 00/12] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-10 13:54 ` [PATCH v3 01/12] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
2026-09-10 13:54 ` [PATCH v3 02/12] NFSD: Track transport in DRC entries Chuck Lever
2026-09-10 13:54 ` [PATCH v3 03/12] NFSD: Prepare bucket pruning for out-of-order eviction Chuck Lever
2026-09-10 13:54 ` [PATCH v3 04/12] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
2026-09-10 13:54 ` [PATCH v3 05/12] NFSD: Record DRC population in lookup tracepoints Chuck Lever
2026-09-10 13:54 ` [PATCH v3 06/12] NFSD: Add reply-acknowledged callback infrastructure Chuck Lever
2026-09-10 13:54 ` [PATCH v3 07/12] SUNRPC: Add TCP sequence-number ACK tracking for reply delivery Chuck Lever
2026-09-10 13:54 ` [PATCH v3 08/12] svcrdma: Fire reply-acknowledged callback on Send completion Chuck Lever
2026-09-10 13:54 ` [PATCH v3 09/12] SUNRPC: Record last-request timestamp on svc_xprt Chuck Lever
2026-09-10 13:54 ` [PATCH v3 10/12] NFSD: Evict unacknowledged DRC entries via implied ACK Chuck Lever
2026-09-10 13:54 ` [PATCH v3 11/12] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
2026-09-10 13:54 ` [PATCH v3 12/12] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
2026-09-10 17:25 ` [PATCH v3 00/12] Improve the scalability of NFSD's classic DRC Jeff Layton
2026-09-10 23:02 ` NeilBrown
2026-09-11 14:42 ` Chuck Lever
2026-09-11 23:20 ` NeilBrown
2026-09-12 16:22 ` Chuck Lever [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=6293796e-8a68-4032-90f3-2adfc1386e55@app.fastmail.com \
--to=cel@kernel.org \
--cc=Dai.Ngo@oracle.com \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=neil@brown.name \
--cc=okorniev@redhat.com \
--cc=rmacklem@uoguelph.ca \
--cc=tom@talpey.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.