From: Chuck Lever <cel@kernel.org>
To: Jeff Layton <jlayton@kernel.org>, NeilBrown <neil@brown.name>,
Olga Kornievskaia <okorniev@redhat.com>,
Dai Ngo <Dai.Ngo@oracle.com>, Tom Talpey <tom@talpey.com>
Cc: Rick Macklem <rmacklem@uoguelph.ca>,
linux-nfs@vger.kernel.org, Chuck Lever <cel@kernel.org>
Subject: [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size
Date: Fri, 18 Sep 2026 13:21:13 -0400 [thread overview]
Message-ID: <20260918-duplicate-reply-cache-v5-1-b6aba9ebf2f4@kernel.org> (raw)
In-Reply-To: <20260918-duplicate-reply-cache-v5-0-b6aba9ebf2f4@kernel.org>
nfsd_cache_size_limit() scales the DRC entry limit with the square
root of low memory, then clamps the result at 256k entries. The
clamp binds on any host with more than 64 GB of low memory. Once
the cache reaches max_drc_entries, nfsd_prune_bucket_locked()
evicts oldest-first regardless of age, so on such a host the cap
governs retention rather than RC_EXPIRE. A server that completes
more than about 2200 calls per second fills 256k entries inside
the 120 second RC_EXPIRE window, and from then on entries are
evicted before they expire. Those are the entries a retransmit
could still hit.
Commit 0338dd157282 ("nfsd: dynamically allocate DRC entries")
added the clamp in 2013, when the formula's worst case of 1 KB
per entry made 256 MB a reasonable ceiling for the hosts of the
day. The formula was already sized to memory; the clamp froze it
at that generation of hardware.
Remove the cap and let the square-root formula govern sizing. It
scales sub-linearly with memory, and the hash table already uses
kvzalloc, so larger sizes need no physical contiguity. The limit is
unchanged at 64 GB and below. At 1 TB the formula yields 1048576
entries, four times the old cap. Even at 1 KB per entry, the worst
case the old comment assumed, a full cache is 1 GB, 0.1% of that
host's memory. That memory is not reclaimable: the shrinker frees
only entries older than RC_EXPIRE, and each net namespace sizes its
own cache.
Assisted-by: LLM
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfscache.c | 33 ++++++++++++++++++---------------
1 file changed, 18 insertions(+), 15 deletions(-)
diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c
index 80364b91331a..d0f65cc9c07a 100644
--- a/fs/nfsd/nfscache.c
+++ b/fs/nfsd/nfscache.c
@@ -47,34 +47,37 @@ static unsigned long nfsd_reply_cache_scan(struct shrinker *shrink,
struct shrink_control *sc);
/*
- * Put a cap on the size of the DRC based on the amount of available
- * low memory in the machine.
+ * Size the DRC by the amount of low memory in the machine. The
+ * limit scales with the square root of available pages; with 4 KB
+ * pages:
*
* 64MB: 8192
- * 128MB: 11585
+ * 128MB: 11584
* 256MB: 16384
- * 512MB: 23170
+ * 512MB: 23168
* 1GB: 32768
- * 2GB: 46340
+ * 2GB: 46336
* 4GB: 65536
- * 8GB: 92681
+ * 8GB: 92672
* 16GB: 131072
+ * 32GB: 185344
+ * 64GB: 262144
+ * 128GB: 370688
+ * 256GB: 524288
+ * 512GB: 741440
+ * 1TB: 1048576
*
- * ...with a hard cap of 256k entries. In the worst case, each entry will be
- * ~1k, so the above numbers should give a rough max of the amount of memory
- * used in k.
- *
- * XXX: these limits are per-container, so memory used will increase
- * linearly with number of containers. Maybe that's OK.
+ * These limits are per-net-namespace, so memory used increases
+ * linearly with the number of namespaces. The shrinker frees only
+ * entries older than RC_EXPIRE, so a cache below its limit is not
+ * reclaimable under memory pressure.
*/
static unsigned int
nfsd_cache_size_limit(void)
{
- unsigned int limit;
unsigned long low_pages = totalram_pages() - totalhigh_pages();
- limit = (16 * int_sqrt(low_pages)) << (PAGE_SHIFT-10);
- return min_t(unsigned int, limit, 256*1024);
+ return (16 * int_sqrt(low_pages)) << (PAGE_SHIFT - 10);
}
/*
--
2.55.0
next prev parent reply other threads:[~2026-09-18 17:21 UTC|newest]
Thread overview: 19+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 17:21 [PATCH v5 00/11] Improve the scalability of NFSD's classic DRC Chuck Lever
2026-09-18 17:21 ` Chuck Lever [this message]
2026-09-18 23:00 ` [PATCH v5 01/11] NFSD: Remove hard cap on duplicate reply cache size NeilBrown
2026-09-20 17:47 ` Chuck Lever
2026-09-20 22:10 ` NeilBrown
2026-09-18 17:21 ` [PATCH v5 02/11] SUNRPC: Assign a unique identifier to each svc_xprt Chuck Lever
2026-09-18 17:21 ` [PATCH v5 03/11] NFSD: Track transport in DRC entries Chuck Lever
2026-09-18 17:21 ` [PATCH v5 04/11] NFSD: Prepare bucket pruning for additional eviction reasons Chuck Lever
2026-09-18 17:21 ` [PATCH v5 05/11] NFSD: Add tracepoints for DRC entry eviction Chuck Lever
2026-09-18 17:21 ` [PATCH v5 06/11] NFSD: Record DRC population in lookup tracepoints Chuck Lever
2026-09-18 17:21 ` [PATCH v5 07/11] SUNRPC: Publish reply positions for upper-layer consumers Chuck Lever
2026-09-18 17:21 ` [PATCH v5 08/11] SUNRPC: Publish TCP reply positions Chuck Lever
2026-09-18 17:21 ` [PATCH v5 09/11] svcrdma: Publish RDMA " Chuck Lever
2026-09-18 17:21 ` [PATCH v5 10/11] NFSD: Evict acknowledged DRC entries Chuck Lever
2026-09-18 17:21 ` [PATCH v5 11/11] NFSD: Remove DRC checksum and payload_misses stat Chuck Lever
2026-09-18 23:39 ` NeilBrown
2026-09-19 16:11 ` Chuck Lever
2026-09-20 10:37 ` NeilBrown
2026-09-20 17:49 ` Chuck Lever
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260918-duplicate-reply-cache-v5-1-b6aba9ebf2f4@kernel.org \
--to=cel@kernel.org \
--cc=Dai.Ngo@oracle.com \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=neil@brown.name \
--cc=okorniev@redhat.com \
--cc=rmacklem@uoguelph.ca \
--cc=tom@talpey.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.