Linux NFS development
 help / color / mirror / Atom feed
From: Jeff Layton <jlayton@kernel.org>
To: Chuck Lever <cel@kernel.org>, NeilBrown <neil@brown.name>,
	Olga Kornievskaia <okorniev@redhat.com>,
	Dai Ngo <Dai.Ngo@oracle.com>, Tom Talpey <tom@talpey.com>
Cc: Rick Macklem <rmacklem@uoguelph.ca>, linux-nfs@vger.kernel.org
Subject: Re: [PATCH v3 00/12] Improve the scalability of NFSD's classic DRC
Date: Thu, 10 Sep 2026 13:25:36 -0400	[thread overview]
Message-ID: <bf2256fabd3190f08232033561f2ec374440a1bd.camel@kernel.org> (raw)
In-Reply-To: <20260910-duplicate-reply-cache-v3-0-31532a4c7449@kernel.org>

On Thu, 2026-09-10 at 09:54 -0400, Chuck Lever wrote:
> A completed DRC entry currently stays in its bucket for 120 seconds
> whether or not the client already holds the reply. On a busy server
> those entries lengthen every bucket walk, and under pressure they
> crowd out entries a retransmit could still hit.
> 
> This series modifies the DRC to retire an entry once there is reason
> to believe its reply was delivered.
> 
> 1. The transport reports delivery outright where it can. TCP now
>    reports when snd_una passes the reply's sequence number, and RDMA
>    reports on each Send completion.
> 2. Where no report will come, a later request on the same connection
>    stands in as an implied ACK, since a client on a live TCP or RDMA
>    connection does not retransmit within it (RFC 1813 Section 4.5).
> 
> UDP continues to use a traditional time-based eviction mechanism.
> 
> The first mechanism is the more reliable of the two, and the backup
> mechanism is more of an informed guess. A pipelined client sends its
> next request before the previous reply lands, and an entry evicted
> in that window is lost if the connection then drops and the client
> retransmits. In the current NFSD code, memory pressure and RC_EXPIRE
> already open this same window. For instance, on a server with fast
> networking and storage, pressure eviction already evicts entries far
> younger than 120 seconds.
> 
> Two DRC capacity guards that were sized for a cache full of stale
> replies can be removed, now that reply delivery reports keep the DRC
> small. The commit messages for those patches explain the rationale
> in detail.
> 
> v2 of this series (implied ACK only) was profiled with "perf record
> -e cycles -e cpu-clock -e LLC-load-misses -e branch-misses" during
> an NFSv3/RDMA 4KB random-write workload. v3 has not been re-profiled.
> nfsd_cache_lookup overhead dropped from 1.53% to 0.76% of CPU
> cycles during this test. The rb-tree operations (rb_erase,
> rb_insert_color) that dominated LLC cache misses fell from a
> combined 13.2% to 1.1% of all LLC-load-misses, because shorter-lived
> entries keep the per-bucket trees small.
> 
> ---
> Changes in v3:
> - Add the reply-acknowledged callback and its TCP and RDMA reporters
>   ahead of implied ACK, which now defers to a pending report.
> - Split the prune-loop restructure into its own patch.
> - Move the eviction tracepoints ahead of the ack patches; the
>   implied-ACK event now lands with implied ACK.
> - Keep err local to the pmap_register block in svc_setup_socket()
>   (per Jeff's review).
> - Never move xpt_last_recv backwards when nfsd threads race.
> - Fold the XPT_ORDERED patch into its consumer (per Jeff's review).
> - State the re-execution exposure of implied-ACK eviction instead of
>   calling the DRC advisory (per Jeff's review).
> - New patch: remove the DRC request checksum and payload_misses stat.
> - New patch: remove the 256k-entry cap on the DRC size.
> - Link to v2: https://patch.msgid.link/20260828-duplicate-reply-cache-v2-0-25069e660a7b@kernel.org
> 
> Changes in v2:
> - Print the DRC eviction tracepoints' age field as unsigned.
> - Link to v1: https://patch.msgid.link/20260826-duplicate-reply-cache-v1-0-b1d51e1af5c7@kernel.org
> 
> ---
> Chuck Lever (12):
>       SUNRPC: Assign a unique identifier to each svc_xprt
>       NFSD: Track transport in DRC entries
>       NFSD: Prepare bucket pruning for out-of-order eviction
>       NFSD: Add tracepoints for DRC entry eviction
>       NFSD: Record DRC population in lookup tracepoints
>       NFSD: Add reply-acknowledged callback infrastructure
>       SUNRPC: Add TCP sequence-number ACK tracking for reply delivery
>       svcrdma: Fire reply-acknowledged callback on Send completion
>       SUNRPC: Record last-request timestamp on svc_xprt
>       NFSD: Evict unacknowledged DRC entries via implied ACK
>       NFSD: Remove DRC checksum and payload_misses stat
>       NFSD: Remove hard cap on duplicate reply cache size
> 
>  .../ABI/testing/procfs-nfsd-reply_cache_stats      |  11 +-
>  fs/nfsd/cache.h                                    |  20 +-
>  fs/nfsd/netns.h                                    |   2 -
>  fs/nfsd/nfscache.c                                 | 310 +++++++++++++--------
>  fs/nfsd/nfssvc.c                                   |  20 +-
>  fs/nfsd/stats.h                                    |   5 -
>  fs/nfsd/trace.h                                    |  81 +++++-
>  include/linux/sunrpc/svc.h                         |  51 ++++
>  include/linux/sunrpc/svc_rdma.h                    |   1 +
>  include/linux/sunrpc/svc_xprt.h                    |   6 +-
>  include/linux/sunrpc/svcsock.h                     |  10 +
>  include/trace/events/sunrpc.h                      |  10 +-
>  net/sunrpc/netns.h                                 |   4 +
>  net/sunrpc/sunrpc_syms.c                           |   2 +
>  net/sunrpc/svc.c                                   |   1 +
>  net/sunrpc/svc_xprt.c                              |  53 +++-
>  net/sunrpc/svcsock.c                               | 119 +++++++-
>  net/sunrpc/xprtrdma/svc_rdma_sendto.c              |  14 +
>  net/sunrpc/xprtrdma/svc_rdma_transport.c           |   7 +-
>  19 files changed, 541 insertions(+), 186 deletions(-)
> ---
> base-commit: abf2077ee32058e60d3f49ecab265d3c4bd953d9
> change-id: 20260325-duplicate-reply-cache-f0fe7c7b740c
> 
> Best regards,
> --  
> Chuck Lever <cel@kernel.org>

I like the idea overall. I'm still a bit nervous that might see a non-
idempotent operations fail occasionally, since this is aggressively
expiring them, but let's see how it does.

Reviewed-by: Jeff Layton <jlayton@kernel.org>

  parent reply	other threads:[~2026-09-10 17:25 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-10 13:54 [PATCH v3 00/12] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-10 13:54 ` [PATCH v3 01/12] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
2026-09-10 13:54 ` [PATCH v3 02/12] NFSD: Track transport in DRC entries Chuck Lever
2026-09-10 13:54 ` [PATCH v3 03/12] NFSD: Prepare bucket pruning for out-of-order eviction Chuck Lever
2026-09-10 13:54 ` [PATCH v3 04/12] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
2026-09-10 13:54 ` [PATCH v3 05/12] NFSD: Record DRC population in lookup tracepoints Chuck Lever
2026-09-10 13:54 ` [PATCH v3 06/12] NFSD: Add reply-acknowledged callback infrastructure Chuck Lever
2026-09-10 13:54 ` [PATCH v3 07/12] SUNRPC: Add TCP sequence-number ACK tracking for reply delivery Chuck Lever
2026-09-10 13:54 ` [PATCH v3 08/12] svcrdma: Fire reply-acknowledged callback on Send completion Chuck Lever
2026-09-10 13:54 ` [PATCH v3 09/12] SUNRPC: Record last-request timestamp on svc_xprt Chuck Lever
2026-09-10 13:54 ` [PATCH v3 10/12] NFSD: Evict unacknowledged DRC entries via implied ACK Chuck Lever
2026-09-10 13:54 ` [PATCH v3 11/12] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
2026-09-10 13:54 ` [PATCH v3 12/12] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
2026-09-10 17:25 ` Jeff Layton [this message]
2026-09-10 23:02 ` [PATCH v3 00/12] Improve the scalability of NFSD's classic DRC NeilBrown
2026-09-11 14:42   ` Chuck Lever
2026-09-11 23:20     ` NeilBrown
2026-09-12 16:22       ` Chuck Lever

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=bf2256fabd3190f08232033561f2ec374440a1bd.camel@kernel.org \
    --to=jlayton@kernel.org \
    --cc=Dai.Ngo@oracle.com \
    --cc=cel@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=neil@brown.name \
    --cc=okorniev@redhat.com \
    --cc=rmacklem@uoguelph.ca \
    --cc=tom@talpey.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox