From: "Chuck Lever" <cel@kernel.org>
To: "Jeff Layton" <jlayton@kernel.org>, NeilBrown <neil@brown.name>
Cc: "Rick Macklem" <rmacklem@uoguelph.ca>,
linux-nfs@vger.kernel.org,
"Olga Kornievskaia" <okorniev@redhat.com>,
"Dai Ngo" <Dai.Ngo@oracle.com>, "Tom Talpey" <tom@talpey.com>
Subject: Re: [PATCH v6 12/12] NFSD: Remove DRC checksum and payload_misses stat
Date: Wed, 23 Sep 2026 12:36:15 -0400 [thread overview]
Message-ID: <9592daa8-4fe4-4e42-8baf-18899480da3a@app.fastmail.com> (raw)
In-Reply-To: <4fe822079c7154ba9f712ca9f7432d345dbc0783.camel@kernel.org>
On Wed, Sep 23, 2026, at 10:56 AM, Jeff Layton wrote:
> On Mon, 2026-09-21 at 09:22 -0400, Chuck Lever wrote:
>> nfsd_cache_csum() costs CPU on every non-idempotent request. It also
>> prevents the use of zero-copy RDMA receives for WRITE and SYMLINK
>> because the payload must be in the server's memory before the DRC
>> lookup can run.
>>
>> Commit 01a7decf7593 ("nfsd: keep a checksum of the first 256 bytes
>> of request") added the checksum when a growing cache made XID
>> collisions easier to hit. It is the only guard against a reused
>> XID: two calls with the same procedure and equal-length arguments
>> match on every other key field.
>>
>> Acknowledgment-based eviction retires a TCP or RDMA entry once the
>> transport confirms delivery and a later request from the same
>> connection visits its bucket. Three kinds of entry outlast that: the
>> entries a connection has issued since it last visited each bucket,
>> entries from a connection the client has since closed, which no
>> later request can evict because their transport is gone, and UDP
>> entries. Those wait out RC_EXPIRE. A collision needs the same XID
>> from the same address and port within that window. The Linux NFS
>> client's XIDs are sequential from a seed drawn with
>> get_random_u32(), so an XID recurs only after 2^32 calls, and a
>> rebooted client does not replay its previous sequence. A client that
>> restarts its sequence from a fixed value and rebinds its previous
>> source port within RC_EXPIRE is unguarded, but a client that repeats
>> a live XID cannot reliably match replies to its own calls anyway.
>>
>> A call whose XID matches an entry sent on the same connection is
>> therefore a retransmit. TCP and RDMA deliver the queued reply or
>> drop the connection, so drop the call as though the entry were
>> still in progress rather than replaying it. A match from another
>> connection, or from UDP, is replayed as before.
>>
>> Remove the checksum, the nfsd_drc_mismatch tracepoint, and the
>> payload_misses stat that counted checksum-detected collisions.
>> nfsd_dispatch() no longer snapshots the argument stream before
>> decoding, and with k_csum gone, struct nfsd_cacherep shrinks to 136
>> bytes.
>>
>> The "payload misses" line disappears from
>> /proc/fs/nfsd/reply_cache_stats. No known userspace tool parses it.
>>
>> Assisted-by: LLM
>> Signed-off-by: Chuck Lever <cel@kernel.org>
>> @@ -561,6 +492,13 @@ int nfsd_cache_lookup(struct svc_rqst *rqstp, unsigned int start,
>> if (rp->c_state == RC_INPROG)
>> goto out_trace;
>>
>> + /*
>> + * The connection that carried the reply delivers it, so a
>> + * retransmit on that connection needs no replay.
>> + */
>> + if (rp->c_pos && rp->c_xprt == rqstp->rq_xprt->xpt_id)
>> + goto out_trace;
>> +
>
> This patch seems to break pynfs RPLY14, which does a mkdir and then a
> replay of it later.
Note: The replay is on the same connection.
> The test is arguably broken though, since it does the replay over the
> same connection as the original request. I'm looking into fixing that
> (and will check for others that might be similarly problematic).
The new NFSD behavior is by design, suggested by Neil. Any client
that reuses an XID on the same connection is either broken or has
sent 4 billion requests in 120 seconds. Clients are supposed to
send retransmits on a different connection.
So, it's possible the test is wrong, but we need to think hard
about whether this new behavior will be a problem for existing
NFSv4.0 clients.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
next prev parent reply other threads:[~2026-09-23 16:36 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 13:22 [PATCH v6 00/12] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-21 13:22 ` [PATCH v6 01/12] NFSD: Make the DRC size limit independent of page size Chuck Lever
2026-09-21 13:22 ` [PATCH v6 02/12] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
2026-09-21 13:22 ` [PATCH v6 03/12] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
2026-09-21 13:22 ` [PATCH v6 04/12] NFSD: Track transport in DRC entries Chuck Lever
2026-09-21 13:22 ` [PATCH v6 05/12] NFSD: Prepare bucket pruning for additional eviction reasons Chuck Lever
2026-09-21 13:22 ` [PATCH v6 06/12] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
2026-09-21 13:22 ` [PATCH v6 07/12] NFSD: Record DRC population in lookup tracepoints Chuck Lever
2026-09-21 13:22 ` [PATCH v6 08/12] SUNRPC: Publish reply positions for upper-layer consumers Chuck Lever
2026-09-21 13:22 ` [PATCH v6 09/12] SUNRPC: Publish TCP reply positions Chuck Lever
2026-09-21 13:22 ` [PATCH v6 10/12] svcrdma: Publish RDMA " Chuck Lever
2026-09-21 13:22 ` [PATCH v6 11/12] NFSD: Evict acknowledged DRC entries Chuck Lever
2026-09-21 13:22 ` [PATCH v6 12/12] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
2026-09-23 14:56 ` Jeff Layton
2026-09-23 16:36 ` Chuck Lever [this message]
2026-09-22 6:55 ` [PATCH v6 00/12] Improve the scalability of NFSD's classic DRC NeilBrown
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9592daa8-4fe4-4e42-8baf-18899480da3a@app.fastmail.com \
--to=cel@kernel.org \
--cc=Dai.Ngo@oracle.com \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=neil@brown.name \
--cc=okorniev@redhat.com \
--cc=rmacklem@uoguelph.ca \
--cc=tom@talpey.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox