All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC
@ 2026-09-18 17:21 Chuck Lever
  2026-09-18 17:21 ` [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
                   ` (10 more replies)
  0 siblings, 11 replies; 19+ messages in thread
From: Chuck Lever @ 2026-09-18 17:21 UTC (permalink / raw)
  To: Jeff Layton, NeilBrown, Olga Kornievskaia, Dai Ngo, Tom Talpey
  Cc: Rick Macklem, linux-nfs, Chuck Lever

A completed DRC entry currently stays in its bucket for at most 120
seconds, whether or not the client already holds the reply. On a
busy server those entries lengthen every bucket walk, and under
pressure they crowd out entries a retransmit could still hit.

On a host with more than 64 GB of low memory the cap, not RC_EXPIRE,
governs retention, so entries are evicted before a retransmit can
hit them. Since 120 seconds is shorter than the default retransmit
timeout for connected transport, entries are frequently dropped from
the DRC before a client might retransmit.

After this series, the DRC retires an entry once the server transport
has seen the client peer acknowledge its reply. Each transport
publishes two positions, the reply it has just sent and the point up
to which the peer has acknowledged, and the DRC compares only the
two. Nothing else about transport state is exposed to fs/nfsd.

UDP entries and replies sent over a TLS session keep the current
time-based eviction. RC_EXPIRE itself is unchanged for every
transport. A longer lifetime for connected transports is follow-on
work. The bucket walk still stops at the first entry that has
neither expired nor been acknowledged, so an acknowledged entry
behind one from another transport stays until it expires. A later
series moves the cache to the connection and removes that limit.

Retention now ends at delivery rather than at consumption. A client
whose RPC timer fires after the reply was acknowledged but before
its RPC layer read it retransmits, and NFSD executes the call again.
The client completes with the first reply and discards the second.
Patch 10 works through the cases.

The XID checksum was sized for a cache full of stale replies. With
acknowledged entries retired, it guards only entries that wait out
RC_EXPIRE (patch 11).

---
Changes in v5:
- Actually make xpt_id 64 bits wide (per Neil's review).
- Move the DRC cap removal first (per Neil's review).
- Rewrite the cap removal's description to stand on its own.
- Drop the 4x visit bound from the prune walk (per Neil's review).
- Evict acknowledged entries only where the prune walk reaches them.
- Restore the shrinker's unlimited per-bucket scan.
- Link to v4: https://patch.msgid.link/20260914-duplicate-reply-cache-v4-0-bcf9d2628a94@kernel.org

Changes in v4:
- Track reply positions instead of ack cookies (per Neil's review).
- Make xpt_id a 64-bit counter and drop the xarray (per Neil's review).
- Drop implied ACK and the last-request timestamp (per Neil's review).
- Keep an unsent reply until RC_EXPIRE (per Neil's review).
- Pin a DRC entry against eviction until its reply has been sent.
- svcrdma: Serialize Send posting so positions match completion order.
- Drop the nfsd_drc_reply_acked tracepoint.
- Link to v3: https://patch.msgid.link/20260910-duplicate-reply-cache-v3-0-31532a4c7449@kernel.org

Changes in v3:
- Add the reply-acknowledged callback and its TCP and RDMA reporters
  ahead of implied ACK, which now defers to a pending report.
- Split the prune-loop restructure into its own patch.
- Move the eviction tracepoints ahead of the ack patches; the
  implied-ACK event now lands with implied ACK.
- Keep err local to the pmap_register block in svc_setup_socket()
  (per Jeff's review).
- Never move xpt_last_recv backwards when nfsd threads race.
- Fold the XPT_ORDERED patch into its consumer (per Jeff's review).
- State the re-execution exposure of implied-ACK eviction instead of
  calling the DRC advisory (per Jeff's review).
- New patch: remove the DRC request checksum and payload_misses stat.
- New patch: remove the 256k-entry cap on the DRC size.
- Link to v2: https://patch.msgid.link/20260828-duplicate-reply-cache-v2-0-25069e660a7b@kernel.org

Changes in v2:
- Print the DRC eviction tracepoints' age field as unsigned.
- Link to v1: https://patch.msgid.link/20260826-duplicate-reply-cache-v1-0-b1d51e1af5c7@kernel.org

---
Chuck Lever (11):
      NFSD: Remove hard cap on duplicate reply cache size
      SUNRPC: Assign a unique identifier to each svc_xprt
      NFSD: Track transport in DRC entries
      NFSD: Prepare bucket pruning for additional eviction reasons
      NFSD: Add tracepoints for DRC entry eviction
      NFSD: Record DRC population in lookup tracepoints
      SUNRPC: Publish reply positions for upper-layer consumers
      SUNRPC: Publish TCP reply positions
      svcrdma: Publish RDMA reply positions
      NFSD: Evict acknowledged DRC entries
      NFSD: Remove DRC checksum and payload_misses stat

 .../ABI/testing/procfs-nfsd-reply_cache_stats      |  11 +-
 fs/nfsd/cache.h                                    |  16 +-
 fs/nfsd/netns.h                                    |   2 -
 fs/nfsd/nfscache.c                                 | 222 +++++++++++----------
 fs/nfsd/nfsd.h                                     |   3 +
 fs/nfsd/nfssvc.c                                   |  11 +-
 fs/nfsd/stats.h                                    |   5 -
 fs/nfsd/trace.h                                    |  47 +++--
 include/linux/sunrpc/svc.h                         |  16 ++
 include/linux/sunrpc/svc_rdma.h                    |   5 +-
 include/linux/sunrpc/svc_xprt.h                    |   3 +
 include/linux/sunrpc/svcsock.h                     |   3 +
 include/trace/events/sunrpc.h                      |   8 +-
 net/sunrpc/svc.c                                   |   3 +
 net/sunrpc/svc_xprt.c                              |  19 +-
 net/sunrpc/svcsock.c                               |  29 +++
 net/sunrpc/xprtrdma/svc_rdma_backchannel.c         |   2 +-
 net/sunrpc/xprtrdma/svc_rdma_sendto.c              |  27 ++-
 net/sunrpc/xprtrdma/svc_rdma_transport.c           |   1 +
 19 files changed, 273 insertions(+), 160 deletions(-)
---
base-commit: abf2077ee32058e60d3f49ecab265d3c4bd953d9
change-id: 20260325-duplicate-reply-cache-f0fe7c7b740c

Best regards,
--  
Chuck Lever <cel@kernel.org>


^ permalink raw reply	[flat|nested] 19+ messages in thread

end of thread, other threads:[~2026-09-20 22:10 UTC | newest]

Thread overview: 19+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-18 17:21 ` [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size Chuck Lever
2026-09-18 23:00   ` NeilBrown
2026-09-20 17:47     ` Chuck Lever
2026-09-20 22:10       ` NeilBrown
2026-09-18 17:21 ` [PATCH v5 02/11] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
2026-09-18 17:21 ` [PATCH v5 03/11] NFSD: Track transport in DRC entries Chuck Lever
2026-09-18 17:21 ` [PATCH v5 04/11] NFSD: Prepare bucket pruning for additional eviction reasons Chuck Lever
2026-09-18 17:21 ` [PATCH v5 05/11] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
2026-09-18 17:21 ` [PATCH v5 06/11] NFSD: Record DRC population in lookup tracepoints Chuck Lever
2026-09-18 17:21 ` [PATCH v5 07/11] SUNRPC: Publish reply positions for upper-layer consumers Chuck Lever
2026-09-18 17:21 ` [PATCH v5 08/11] SUNRPC: Publish TCP reply positions Chuck Lever
2026-09-18 17:21 ` [PATCH v5 09/11] svcrdma: Publish RDMA " Chuck Lever
2026-09-18 17:21 ` [PATCH v5 10/11] NFSD: Evict acknowledged DRC entries Chuck Lever
2026-09-18 17:21 ` [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
2026-09-18 23:39   ` NeilBrown
2026-09-19 16:11     ` Chuck Lever
2026-09-20 10:37       ` NeilBrown
2026-09-20 17:49         ` Chuck Lever

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.