From: Piyush Sachdeva <s.piyush1024@gmail.com>
To: Anna Schumaker <anna@nowheycreamery.com>,
Tigran Mkrtchyan <tigran.mkrtchyan@desy.de>,
Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs <linux-nfs@vger.kernel.org>,
Chuck Lever <cel@kernel.org>,
Trond Myklebust <trondmy@kernel.org>,
sfrench@samba.org, sprasad@microsoft.com,
vaibsharma@microsoft.com
Subject: Re: NFS delegations behavior analysis
Date: Wed, 09 Sep 2026 15:40:43 +0530 [thread overview]
Message-ID: <m2a4pqwx98.fsf@gmail.com> (raw)
In-Reply-To: <a9885b28-a3fb-4ec6-b6fb-1be78ef16b2e@app.fastmail.com>
"Anna Schumaker" <anna@nowheycreamery.com> writes:
> On Tue, Jun 23, 2026, at 7:04 AM, Mkrtchyan, Tigran wrote:
>> ----- Original Message -----
>>> From: "Jeff Layton" <jlayton@kernel.org>
>>> To: "Piyush Sachdeva" <s.piyush1024@gmail.com>, "linux-nfs" <linux-nfs@vger.kernel.org>, "Chuck Lever" <cel@kernel.org>,
>>> "trondmy" <trondmy@kernel.org>, sfrench@samba.org, sprasad@microsoft.com
>>> Cc: vaibsharma@microsoft.com
>>> Sent: Tuesday, 23 June, 2026 12:50:16
>>> Subject: Re: NFS delegations behavior analysis
>>
>>> On Tue, 2026-06-23 at 15:31 +0530, Piyush Sachdeva wrote:
>>>> Hi,
>>>> Lately I have been running micro benchmarks around the `ls` command and
>>>> reading through the code documentation of the NFS client to better
>>>> understand the client side caching behavior with and without
>>>> delegations.
>>>>
>>>> Understanding so far:
>>>> Delegations (both file and directory) are granted by the server to the
>>>> client, indefinitely (until revoked or under the watermark) to cache
>>>> attributes. The caching of data is a result of the attribute
>>>> cache. Hence forth, a directory delegation will cache the directory
>>>> attributes and the names of the files in the directory, and a file
>>>> delegation will cache the attributes of the file and the file data.
>>>>
>>>> Workload run:
>>>> I focused on the 2 workloads below, doing 2 passes of a large flat
>>>> directory (with close to 100K files) -
>>>> a cold pass, and warm pass using the cache from the cold pass:
>>>> - lslr - ls -lR on both runs
>>>> - lsmix - ls -R (cold) and then ls -lR (warm)
>>>>
>>>> I also played with the rdirplus behavior using both the default
>>>> heuristic behavior and the `rdirplus=force` set at mount time.
>>>>
>>>> Numbers:
>>>> actimeo=5s, rdirplus=force, ACLs off, flat_dir
>>>> ==================================================================
>>>>
>>>> | LSLR | LSMIX
>>>> | (ls -lR cold / warm) | (p1 ls -R / p2 ls -lR)
>>>> Operation | flat cold | flat warm | flat p1 | flat p2
>>>> -----------------+-------------+-----------+-------------+---------
>>>> READDIR calls | 27 | 0 | 27 | 0
>>>> READDIR recv B | 23,603,024 | 0 | 23,603,024 | 0
>>>> call type | readdirplus | -- | readdirplus | --
>>>> LOOKUP | 1 | 0 | 1 | 0
>>>> GETATTR | 3 | 100,000 | 2 | 100,001
>>>> ACCESS | 2 | 0 | 2 | 0
>>>> -----------------+-------------+-----------+-------------+---------
>>>> Elapsed (age) | ~14 s | ~62 s | ~16 s | ~63 s
>>>>
>>>>
>>>> Observations:
>>>> When doing `ls` or `ls -l` on a directory, due to the open(2) on the
>>>> directory, the client gets a directory delegation - caching the
>>>> directory attributes and file names. However, as we don't have file
>>>> delegations due to no open(2) calls to any of the files. Henceforth,
>>>> the cache of file attributes is governed by `actimeo`.
>>>> Now here is the interesting bit, if the next `ls -l` is issued after
>>>> the `actimeo`, a massive GETATTR storm hits the server, doing stat()
>>>> calls for every file in the directory. As a result, the performance of
>>>> this warm `ls -l` run ends up being worse than the cold pass. I am
>>>> guessing this is most likely due to the compounded "rdirplus" being more
>>>> efficient than stat() calls.
>>>>
>>>>
>>>> Proposal:
>>>> For large directories, this ends up being a massive problem, taking 1-2
>>>> minutes when enumerating a directory on the warm passes.
>>>> - An easier way to tackle this could be to do a rdirplus=[auto | forced]
>>>> instead of issuing the stat(2) storm to the server: When the client
>>>> notices that there are cache misses, which would be the case of file
>>>> attributes, instead of fetching file names from the directory-delegation
>>>> cache and attributes from GETATTR, the client does a READDIRPLUS to
>>>> the server, nonetheless.
>>>> - A more tedious would be the to cache file attributes as well, as a part
>>>> of the directory delegation. This would end up requiring a change in the
>>>> NFS protocol spec though.
>>>> - Bulk GETATTR calls: I am uncertain of the feasibility of this, but
>>>> what if, the client could do 1 GETATTR call for getting attributes
>>>> for multiple files.
>>>
>>>
>>> ls is such a hard workload to get right, because we don't really get an
>>
>> 100% agree. And there were a couple of attempts to address this issue
>> (second ls that is slow).
>>
>>> indication in the kernel of what userland's intentions are. It's
>>> basically a readdir() call followed by a bunch of stat()'s, but at the
>>> point where we're getting the readdir() call, we don't know if userland
>>> intends to stat() those files or not. We have to make a guess about
>>> that intention.
>>>
>>> In this case, it sounds like the directory cache was valid, so the
>>> client decided it didn't need to do a READDIR at all, but the
>>> individual files had caches that timed out.
>>>
>>> So imagine you're the kernel client and have been given that second
>>> readdir() call: Why should you decide to do a READDIRPLUS at that point
>>> instead of a regular READDIR?
>>
>> May we need some kind of client-side heuristics, like on the server side
>> for open-delegations, where after seeing some `stats` for files in the
>> In the same directory, the client will decide to switch to READDIR (v4)
>> to get all attributes in one go.
>
> We do something like that already. I don't think we'll ever have a readdir
> plus heuristic that makes everybody happy without userspace somehow telling
> us their intentions. I know a statx-based readdir plus system call probably
> sounds crazy, but something like that would go a long way to take all the
> guesswork out of things on our end.
>
> Anna
>
>>
>> Best regards,
>> Tigran.
>>
>>> --
>>> Jeff Layton <jlayton@kernel.org>
>>
>> Attachments:
>> * smime.p7s
Hi,
I believe the problem from this discussion has been the cost associated
with GETATTRs during directory enumerations, compounded by the fact that
the kernel can't know whether more GETATTRs on the same directory are
coming.
Since the readdirplus heuristic rework in v5.18, plaing READDIRs
fetching the file type in v6.8, and introduction of rdirplus=force in
v6.14, the client has become quite performant especially for the
ls-style workloads. find and du, still suffer from GETATTR storms for
large-enough directories, as they enumerate the directory once then
stat() every child. Note that nfs_getattr() already feeds the
cache_hit/cache_miss counters (since v4.10), but that signal lives on
the open-directory context and is consumed only inside nfs_readdir().
This is where the performance hurts the most.
As a benchmark, I prototyped a solution for older kernels (v4.18, v5.14,
v6.4-based distros). An eBPF-based detector recognizes GETATTR/LOOKUP
storms on large flat directories and triggers a userspace helper to
enumerate the directory (READDIR(PLUS), start -> EOF). The returned
entries refresh the attribute/dcache so subsequent GETATTR/LOOKUP calls
are served from cache. We measured directory-enumeration time and
GETATTR/LOOKUP counts drop by ~50–80x on large directories, for ls,
find, and du.
With that as evidence, I'd like to propose exploring two in-kernel
directions and get your opinion on feasibility and desirability:
1. Extend the existing getattr miss-tracking: when cache_misses on a
directory cross a threshold, trigger a deferred, background
READDIRPLUS to refresh that directory's child attributes. This gives
the find / du case a consumer for the miss signal it already
generates.
2. A batched-GETATTR path so tools like du /find can request many
children's attributes in fewer round-trips. I recognize there's no
batched-stat() interface today and that this spans userspace and the
kernel, so I'm raising it mainly to ask whether a metadata-batching
interface (or wiring stat readahead to an NFSv4 multi-GETATTR
compound) is something the community would consider.
I'd appreciate your thoughts on both, particularly whether (1) is best
framed as an extension of the v5.18 heuristic.
--
Best,
Piyush Sachdeva
next prev parent reply other threads:[~2026-09-09 10:10 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-06-23 10:01 NFS delegations behavior analysis Piyush Sachdeva
2026-06-23 10:50 ` Jeff Layton
2026-06-23 11:04 ` Mkrtchyan, Tigran
2026-06-23 11:10 ` Jeff Layton
2026-06-23 13:11 ` Benjamin Coddington
2026-06-23 13:31 ` Daire Byrne
2026-06-23 13:32 ` Benjamin Coddington
2026-06-23 13:40 ` Jeff Layton
2026-06-23 13:59 ` Benjamin Coddington
2026-06-23 16:29 ` Trond Myklebust
2026-06-23 13:33 ` Jeff Layton
2026-06-23 13:11 ` Anna Schumaker
2026-09-09 10:10 ` Piyush Sachdeva [this message]
2026-09-09 12:07 ` Piyush Sachdeva
2026-09-09 13:13 ` Jeff Layton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=m2a4pqwx98.fsf@gmail.com \
--to=s.piyush1024@gmail.com \
--cc=anna@nowheycreamery.com \
--cc=cel@kernel.org \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=sfrench@samba.org \
--cc=sprasad@microsoft.com \
--cc=tigran.mkrtchyan@desy.de \
--cc=trondmy@kernel.org \
--cc=vaibsharma@microsoft.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox