From: Piyush Sachdeva <s.piyush1024@gmail.com>
To: Anna Schumaker <anna@nowheycreamery.com>,
Tigran Mkrtchyan <tigran.mkrtchyan@desy.de>,
Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs <linux-nfs@vger.kernel.org>,
Chuck Lever <cel@kernel.org>,
Trond Myklebust <trondmy@kernel.org>,
sfrench@samba.org, sprasad@microsoft.com,
vaibsharma@microsoft.com
Subject: Re: NFS delegations behavior analysis
Date: Wed, 09 Sep 2026 17:37:08 +0530 [thread overview]
Message-ID: <m25x0ewrv7.fsf@gmail.com> (raw)
In-Reply-To: <m2a4pqwx98.fsf@gmail.com>
Piyush Sachdeva <s.piyush1024@gmail.com> writes:
> "Anna Schumaker" <anna@nowheycreamery.com> writes:
>
>> On Tue, Jun 23, 2026, at 7:04 AM, Mkrtchyan, Tigran wrote:
>>> ----- Original Message -----
>>>> From: "Jeff Layton" <jlayton@kernel.org>
>>>> To: "Piyush Sachdeva" <s.piyush1024@gmail.com>, "linux-nfs" <linux-nfs@vger.kernel.org>, "Chuck Lever" <cel@kernel.org>,
>>>> "trondmy" <trondmy@kernel.org>, sfrench@samba.org, sprasad@microsoft.com
>>>> Cc: vaibsharma@microsoft.com
>>>> Sent: Tuesday, 23 June, 2026 12:50:16
>>>> Subject: Re: NFS delegations behavior analysis
>>>
>>>> On Tue, 2026-06-23 at 15:31 +0530, Piyush Sachdeva wrote:
>>>>> Hi,
>>>>> Lately I have been running micro benchmarks around the `ls` command and
>>>>> reading through the code documentation of the NFS client to better
>>>>> understand the client side caching behavior with and without
>>>>> delegations.
>>>>>
>>>>> Understanding so far:
>>>>> Delegations (both file and directory) are granted by the server to the
>>>>> client, indefinitely (until revoked or under the watermark) to cache
>>>>> attributes. The caching of data is a result of the attribute
>>>>> cache. Hence forth, a directory delegation will cache the directory
>>>>> attributes and the names of the files in the directory, and a file
>>>>> delegation will cache the attributes of the file and the file data.
>>>>>
>>>>> Workload run:
>>>>> I focused on the 2 workloads below, doing 2 passes of a large flat
>>>>> directory (with close to 100K files) -
>>>>> a cold pass, and warm pass using the cache from the cold pass:
>>>>> - lslr - ls -lR on both runs
>>>>> - lsmix - ls -R (cold) and then ls -lR (warm)
>>>>>
>>>>> I also played with the rdirplus behavior using both the default
>>>>> heuristic behavior and the `rdirplus=force` set at mount time.
>>>>>
>>>>> Numbers:
>>>>> actimeo=5s, rdirplus=force, ACLs off, flat_dir
>>>>> ==================================================================
>>>>>
>>>>> | LSLR | LSMIX
>>>>> | (ls -lR cold / warm) | (p1 ls -R / p2 ls -lR)
>>>>> Operation | flat cold | flat warm | flat p1 | flat p2
>>>>> -----------------+-------------+-----------+-------------+---------
>>>>> READDIR calls | 27 | 0 | 27 | 0
>>>>> READDIR recv B | 23,603,024 | 0 | 23,603,024 | 0
>>>>> call type | readdirplus | -- | readdirplus | --
>>>>> LOOKUP | 1 | 0 | 1 | 0
>>>>> GETATTR | 3 | 100,000 | 2 | 100,001
>>>>> ACCESS | 2 | 0 | 2 | 0
>>>>> -----------------+-------------+-----------+-------------+---------
>>>>> Elapsed (age) | ~14 s | ~62 s | ~16 s | ~63 s
>>>>>
>>>>>
>>>>> Observations:
>>>>> When doing `ls` or `ls -l` on a directory, due to the open(2) on the
>>>>> directory, the client gets a directory delegation - caching the
>>>>> directory attributes and file names. However, as we don't have file
>>>>> delegations due to no open(2) calls to any of the files. Henceforth,
>>>>> the cache of file attributes is governed by `actimeo`.
>>>>> Now here is the interesting bit, if the next `ls -l` is issued after
>>>>> the `actimeo`, a massive GETATTR storm hits the server, doing stat()
>>>>> calls for every file in the directory. As a result, the performance of
>>>>> this warm `ls -l` run ends up being worse than the cold pass. I am
>>>>> guessing this is most likely due to the compounded "rdirplus" being more
>>>>> efficient than stat() calls.
>>>>>
>>>>>
>>>>> Proposal:
>>>>> For large directories, this ends up being a massive problem, taking 1-2
>>>>> minutes when enumerating a directory on the warm passes.
>>>>> - An easier way to tackle this could be to do a rdirplus=[auto | forced]
>>>>> instead of issuing the stat(2) storm to the server: When the client
>>>>> notices that there are cache misses, which would be the case of file
>>>>> attributes, instead of fetching file names from the directory-delegation
>>>>> cache and attributes from GETATTR, the client does a READDIRPLUS to
>>>>> the server, nonetheless.
>>>>> - A more tedious would be the to cache file attributes as well, as a part
>>>>> of the directory delegation. This would end up requiring a change in the
>>>>> NFS protocol spec though.
>>>>> - Bulk GETATTR calls: I am uncertain of the feasibility of this, but
>>>>> what if, the client could do 1 GETATTR call for getting attributes
>>>>> for multiple files.
>>>>
>>>>
>>>> ls is such a hard workload to get right, because we don't really get an
>>>
>>> 100% agree. And there were a couple of attempts to address this issue
>>> (second ls that is slow).
>>>
>>>> indication in the kernel of what userland's intentions are. It's
>>>> basically a readdir() call followed by a bunch of stat()'s, but at the
>>>> point where we're getting the readdir() call, we don't know if userland
>>>> intends to stat() those files or not. We have to make a guess about
>>>> that intention.
>>>>
>>>> In this case, it sounds like the directory cache was valid, so the
>>>> client decided it didn't need to do a READDIR at all, but the
>>>> individual files had caches that timed out.
>>>>
>>>> So imagine you're the kernel client and have been given that second
>>>> readdir() call: Why should you decide to do a READDIRPLUS at that point
>>>> instead of a regular READDIR?
>>>
>>> May we need some kind of client-side heuristics, like on the server side
>>> for open-delegations, where after seeing some `stats` for files in the
>>> In the same directory, the client will decide to switch to READDIR (v4)
>>> to get all attributes in one go.
>>
>> We do something like that already. I don't think we'll ever have a readdir
>> plus heuristic that makes everybody happy without userspace somehow telling
>> us their intentions. I know a statx-based readdir plus system call probably
>> sounds crazy, but something like that would go a long way to take all the
>> guesswork out of things on our end.
>>
>> Anna
>>
>>>
>>> Best regards,
>>> Tigran.
>>>
>>>> --
>>>> Jeff Layton <jlayton@kernel.org>
>>>
>>> Attachments:
>>> * smime.p7s
>
> Hi,
> I believe the problem from this discussion has been the cost associated
> with GETATTRs during directory enumerations, compounded by the fact that
> the kernel can't know whether more GETATTRs on the same directory are
> coming.
>
> Since the readdirplus heuristic rework in v5.18, plaing READDIRs
> fetching the file type in v6.8, and introduction of rdirplus=force in
> v6.14, the client has become quite performant especially for the
> ls-style workloads. find and du, still suffer from GETATTR storms for
> large-enough directories, as they enumerate the directory once then
> stat() every child. Note that nfs_getattr() already feeds the
> cache_hit/cache_miss counters (since v4.10), but that signal lives on
> the open-directory context and is consumed only inside nfs_readdir().
> This is where the performance hurts the most.
>
> As a benchmark, I prototyped a solution for older kernels (v4.18, v5.14,
> v6.4-based distros). An eBPF-based detector recognizes GETATTR/LOOKUP
> storms on large flat directories and triggers a userspace helper to
> enumerate the directory (READDIR(PLUS), start -> EOF). The returned
> entries refresh the attribute/dcache so subsequent GETATTR/LOOKUP calls
> are served from cache. We measured directory-enumeration time and
> GETATTR/LOOKUP counts drop by ~50–80x on large directories, for ls,
> find, and du.
>
> With that as evidence, I'd like to propose exploring two in-kernel
> directions and get your opinion on feasibility and desirability:
> 1. Extend the existing getattr miss-tracking: when cache_misses on a
> directory cross a threshold, trigger a deferred, background
> READDIRPLUS to refresh that directory's child attributes. This gives
> the find / du case a consumer for the miss signal it already
> generates.
> 2. A batched-GETATTR path so tools like du /find can request many
> children's attributes in fewer round-trips. I recognize there's no
> batched-stat() interface today and that this spans userspace and the
> kernel, so I'm raising it mainly to ask whether a metadata-batching
> interface (or wiring stat readahead to an NFSv4 multi-GETATTR
> compound) is something the community would consider.
>
> I'd appreciate your thoughts on both, particularly whether (1) is best
> framed as an extension of the v5.18 heuristic.
>
> --
> Best,
> Piyush Sachdeva
I am posting the benchmark numbers of the directory enumeration workload
both with and without the eBPF-based tool for reference.
Different workloads:
- ls_r: 2 iterations of - ls -R "$TARGET_DIR"
- ls_lr: 2 iterations of - ls -lR "$TARGET_DIR"
- lsmix: ls -R "$TARGET_DIR" followed by ls -lR "$TARGET_DIR"
- find: 2 iterations of - find "$TARGET_DIR" -type f -size +0c
- du: 2 iterations of - du "$TARGET_DIR"
- The share gets remounted in between each workload runs.
- Each workload has 2 successive iterations (no gap) - a cold run that
populates the cache, and after, a warm run that re-uses the cache.
- Distro(kernel): SLES 15 SP5 (v5.14)
- $TARGET_DIR - Flat large directory; file count - 1 Million files
- actimeo=30
Numbers:
- Iteration 1 (cold) - without eBPF / with eBPF
op ls_r ls_lr lsmix du find
--------------- -------------------- -------------------- -------------------- ---------------------- ----------------------
READDIR calls 217 / 455 263 / 263 217 / 454 127 / 422 127 / 421
READDIR recv(B) 210,687,496 / 248,029,456 / 210,687,496 / 115,395,048 / 115,395,048 /
440,573,116 248,029,456 439,524,444 405,689,200 404,640,792
LOOKUP 30,178 / 6,161 1 / 1 30,178 / 6,691 733,573 / 19,465 733,573 / 5,837
GETATTR 61 / 5 1 / 1 63 / 5 177,623 / 5 177,622 / 4
ACCESS 2 / 2 2 / 2 2 / 2 2 / 2 2 / 2
elapsed (ms) 919,684 / 40,657 23,938 / 24,984 947,506 / 41,133 2,643,906 / 78,684 2,627,536 / 38,337
- Iteration 2 (warm) - without eBPF / with eBPF
op ls_r ls_lr lsmix du find
--------------- -------------------- -------------------- -------------------- ---------------------- ----------------------
READDIR calls 217 / 0 0 / 0 258 / 256 124 / 317 123 / 417
READDIR recv(B) 208,601,368 / 0 0 / 0 248,909,816 / 109,134,232 / 109,134,232 /
248,909,816 398,379,976 398,379,976
LOOKUP 0 / 0 0 / 0 0 / 0 0 / 0 0 / 0
GETATTR 301,214 / 1 1 / 1 127 / 127 822,266 / 26,329 822,267 / 25,611
ACCESS 0 / 0 0 / 0 0 / 0 0 / 0 0 / 0
elapsed (ms) 550,011 / 2,648 5,231 / 6,419 23,264 / 24,090 1,494,821 / 65,103 1,527,773 / 64,990
You can see the reduction in total time, GETATTRs and LOOKUPs.
--
Best,
Piyush Sachdeva <s.piyush1024@gmail.com>
next prev parent reply other threads:[~2026-09-09 12:07 UTC|newest]
Thread overview: 15+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-06-23 10:01 NFS delegations behavior analysis Piyush Sachdeva
2026-06-23 10:50 ` Jeff Layton
2026-06-23 11:04 ` Mkrtchyan, Tigran
2026-06-23 11:10 ` Jeff Layton
2026-06-23 13:11 ` Benjamin Coddington
2026-06-23 13:31 ` Daire Byrne
2026-06-23 13:32 ` Benjamin Coddington
2026-06-23 13:40 ` Jeff Layton
2026-06-23 13:59 ` Benjamin Coddington
2026-06-23 16:29 ` Trond Myklebust
2026-06-23 13:33 ` Jeff Layton
2026-06-23 13:11 ` Anna Schumaker
2026-09-09 10:10 ` Piyush Sachdeva
2026-09-09 12:07 ` Piyush Sachdeva [this message]
2026-09-09 13:13 ` Jeff Layton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=m25x0ewrv7.fsf@gmail.com \
--to=s.piyush1024@gmail.com \
--cc=anna@nowheycreamery.com \
--cc=cel@kernel.org \
--cc=jlayton@kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=sfrench@samba.org \
--cc=sprasad@microsoft.com \
--cc=tigran.mkrtchyan@desy.de \
--cc=trondmy@kernel.org \
--cc=vaibsharma@microsoft.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox