Linux NFS development
 help / color / mirror / Atom feed
From: Piyush Sachdeva <s.piyush1024@gmail.com>
To: Anna Schumaker <anna@nowheycreamery.com>,
	Tigran Mkrtchyan <tigran.mkrtchyan@desy.de>,
	Jeff Layton <jlayton@kernel.org>
Cc: linux-nfs <linux-nfs@vger.kernel.org>,
	Chuck Lever <cel@kernel.org>,
	Trond Myklebust <trondmy@kernel.org>,
	sfrench@samba.org, sprasad@microsoft.com,
	vaibsharma@microsoft.com
Subject: Re: NFS delegations behavior analysis
Date: Wed, 09 Sep 2026 17:37:08 +0530	[thread overview]
Message-ID: <m25x0ewrv7.fsf@gmail.com> (raw)
In-Reply-To: <m2a4pqwx98.fsf@gmail.com>

Piyush Sachdeva <s.piyush1024@gmail.com> writes:

> "Anna Schumaker" <anna@nowheycreamery.com> writes:
>
>> On Tue, Jun 23, 2026, at 7:04 AM, Mkrtchyan, Tigran wrote:
>>> ----- Original Message -----
>>>> From: "Jeff Layton" <jlayton@kernel.org>
>>>> To: "Piyush Sachdeva" <s.piyush1024@gmail.com>, "linux-nfs" <linux-nfs@vger.kernel.org>, "Chuck Lever" <cel@kernel.org>,
>>>> "trondmy" <trondmy@kernel.org>, sfrench@samba.org, sprasad@microsoft.com
>>>> Cc: vaibsharma@microsoft.com
>>>> Sent: Tuesday, 23 June, 2026 12:50:16
>>>> Subject: Re: NFS delegations behavior analysis
>>>
>>>> On Tue, 2026-06-23 at 15:31 +0530, Piyush Sachdeva wrote:
>>>>> Hi,
>>>>> Lately I have been running micro benchmarks around the `ls` command and
>>>>> reading through the code documentation of the NFS client to better
>>>>> understand the client side caching behavior with and without
>>>>> delegations.
>>>>> 
>>>>> Understanding so far:
>>>>> Delegations (both file and directory) are granted by the server to the
>>>>> client, indefinitely (until revoked or under the watermark) to cache
>>>>> attributes. The caching of data is a result of the attribute
>>>>> cache. Hence forth, a directory delegation will cache the directory
>>>>> attributes and the names of the files in the directory, and a file
>>>>> delegation will cache the attributes of the file and the file data.
>>>>> 
>>>>> Workload run:
>>>>> I focused on the 2 workloads below, doing 2 passes of a large flat
>>>>> directory (with close to 100K files) -
>>>>> a cold pass, and warm pass using the cache from the cold pass:
>>>>> - lslr - ls -lR on both runs
>>>>> - lsmix - ls -R (cold) and then ls -lR (warm)
>>>>> 
>>>>> I also played with the rdirplus behavior using both the default
>>>>> heuristic behavior and the `rdirplus=force` set at mount time.
>>>>> 
>>>>> Numbers:
>>>>> actimeo=5s, rdirplus=force, ACLs off, flat_dir
>>>>> ==================================================================
>>>>> 
>>>>>                  |         LSLR          |         LSMIX
>>>>>                  |  (ls -lR cold / warm) |  (p1 ls -R / p2 ls -lR)
>>>>> Operation        |  flat cold  | flat warm |   flat p1   | flat p2
>>>>> -----------------+-------------+-----------+-------------+---------
>>>>> READDIR calls    |    27       |     0     |   27        |    0
>>>>> READDIR recv B   | 23,603,024  |     0     | 23,603,024  |    0
>>>>>    call type     | readdirplus |    --     | readdirplus |    --
>>>>> LOOKUP           |     1       |     0     |    1        |    0
>>>>> GETATTR          |     3       |  100,000  |    2        | 100,001
>>>>> ACCESS           |     2       |     0     |    2        |    0
>>>>> -----------------+-------------+-----------+-------------+---------
>>>>> Elapsed (age)    |  ~14 s      |  ~62 s    |   ~16 s     |  ~63 s
>>>>> 
>>>>> 
>>>>> Observations:
>>>>> When doing `ls` or `ls -l` on a directory, due to the open(2) on the
>>>>> directory, the client gets a directory delegation - caching the
>>>>> directory attributes and file names. However, as we don't have file
>>>>> delegations due to no open(2) calls to any of the files. Henceforth,
>>>>> the cache of file attributes is governed by `actimeo`.
>>>>> Now here is the interesting bit, if the next `ls -l` is issued after
>>>>> the `actimeo`, a massive GETATTR storm hits the server, doing stat()
>>>>> calls for every file in the directory. As a result, the performance of
>>>>> this warm `ls -l` run ends up being worse than the cold pass. I am
>>>>> guessing this is most likely due to the compounded "rdirplus" being more
>>>>> efficient than stat() calls.
>>>>> 
>>>>> 
>>>>> Proposal:
>>>>> For large directories, this ends up being a massive problem, taking 1-2
>>>>> minutes when enumerating a directory on the warm passes.
>>>>> - An easier way to tackle this could be to do a rdirplus=[auto | forced]
>>>>>   instead of issuing the stat(2) storm to the server: When the client
>>>>>   notices that there are cache misses, which would be the case of file
>>>>>   attributes, instead of fetching file names from the directory-delegation
>>>>>   cache and attributes from GETATTR, the client does a READDIRPLUS to
>>>>>   the server, nonetheless.
>>>>> - A more tedious would be the to cache file attributes as well, as a part
>>>>>   of the directory delegation. This would end up requiring a change in the
>>>>>   NFS protocol spec though.
>>>>> - Bulk GETATTR calls: I am uncertain of the feasibility of this, but
>>>>>   what if, the client could do 1 GETATTR call for getting attributes
>>>>>   for multiple files.
>>>> 
>>>> 
>>>> ls is such a hard workload to get right, because we don't really get an
>>>
>>> 100% agree. And there were a couple of attempts to address this issue
>>> (second ls that is slow).
>>>
>>>> indication in the kernel of what userland's intentions are. It's
>>>> basically a readdir() call followed by a bunch of stat()'s, but at the
>>>> point where we're getting the readdir() call, we don't know if userland
>>>> intends to stat() those files or not. We have to make a guess about
>>>> that intention.
>>>> 
>>>> In this case, it sounds like the directory cache was valid, so the
>>>> client decided it didn't need to do a READDIR at all, but the
>>>> individual files had caches that timed out.
>>>> 
>>>> So imagine you're the kernel client and have been given that second
>>>> readdir() call: Why should you decide to do a READDIRPLUS at that point
>>>> instead of a regular READDIR?
>>>
>>> May we need some kind of client-side heuristics, like on the server side
>>> for open-delegations, where after seeing some `stats` for files in the
>>> In the same directory, the client will decide to switch to READDIR (v4)
>>> to get all attributes in one go.
>>
>> We do something like that already. I don't think we'll ever have a readdir
>> plus heuristic that makes everybody happy without userspace somehow telling
>> us their intentions. I know a statx-based readdir plus system call probably
>> sounds crazy, but something like that would go a long way to take all the
>> guesswork out of things on our end.
>>
>> Anna
>>
>>>
>>> Best regards,
>>>    Tigran.
>>>
>>>> --
>>>> Jeff Layton <jlayton@kernel.org>
>>>
>>> Attachments:
>>> * smime.p7s
>
> Hi,
> I believe the problem from this discussion has been the cost associated
> with GETATTRs during directory enumerations, compounded by the fact that
> the kernel can't know whether more GETATTRs on the same directory are
> coming.
>
> Since the readdirplus heuristic rework in v5.18, plaing READDIRs
> fetching the file type in v6.8, and introduction of rdirplus=force in
> v6.14, the client has become quite performant especially for the
> ls-style workloads. find and du, still suffer from GETATTR storms for
> large-enough directories, as they enumerate the directory once then
> stat() every child. Note that nfs_getattr() already feeds the
> cache_hit/cache_miss counters (since v4.10), but that signal lives on
> the open-directory context and is consumed only inside nfs_readdir().
> This is where the performance hurts the most.
>
> As a benchmark, I prototyped a solution for older kernels (v4.18, v5.14,
> v6.4-based distros). An eBPF-based detector recognizes GETATTR/LOOKUP
> storms on large flat directories and triggers a userspace helper to
> enumerate the directory (READDIR(PLUS), start -> EOF). The returned
> entries refresh the attribute/dcache so subsequent GETATTR/LOOKUP calls
> are served from cache.  We measured directory-enumeration time and
> GETATTR/LOOKUP counts drop by ~50–80x on large directories, for ls,
> find, and du.
>
> With that as evidence, I'd like to propose exploring two in-kernel
> directions and get your opinion on feasibility and desirability:
> 1. Extend the existing getattr miss-tracking: when cache_misses on a
>    directory cross a threshold, trigger a deferred, background
>    READDIRPLUS to refresh that directory's child attributes. This gives
>    the find / du case a consumer for the miss signal it already
>    generates.
> 2. A batched-GETATTR path so tools like du /find can request many
>    children's attributes in fewer round-trips. I recognize there's no
>    batched-stat() interface today and that this spans userspace and the
>    kernel, so I'm raising it mainly to ask whether a metadata-batching
>    interface (or wiring stat readahead to an NFSv4 multi-GETATTR
>    compound) is something the community would consider.
>
> I'd appreciate your thoughts on both, particularly whether (1) is best
> framed as an extension of the v5.18 heuristic.
>
> -- 
> Best,
> Piyush Sachdeva

I am posting the benchmark numbers of the directory enumeration workload
both with and without the eBPF-based tool for reference.

Different workloads:
- ls_r: 2 iterations of - ls -R "$TARGET_DIR"
- ls_lr: 2 iterations of - ls -lR "$TARGET_DIR"
- lsmix: ls -R "$TARGET_DIR" followed by ls -lR "$TARGET_DIR"
- find: 2 iterations of - find "$TARGET_DIR" -type f -size +0c
- du: 2 iterations of - du "$TARGET_DIR"

- The share gets remounted in between each workload runs.
- Each workload has 2 successive iterations (no gap) - a cold run that
  populates the cache, and after, a warm run that re-uses the cache.
- Distro(kernel): SLES 15 SP5 (v5.14)
- $TARGET_DIR - Flat large directory; file count - 1 Million files
- actimeo=30

Numbers:
- Iteration 1 (cold) - without eBPF / with eBPF

op               ls_r                  ls_lr                 lsmix                 du                      find
---------------  --------------------  --------------------  --------------------  ----------------------  ----------------------
READDIR calls    217 / 455             263 / 263             217 / 454             127 / 422               127 / 421
READDIR recv(B)  210,687,496 /         248,029,456 /         210,687,496 /         115,395,048 /           115,395,048 /
                 440,573,116           248,029,456           439,524,444           405,689,200             404,640,792
LOOKUP           30,178 / 6,161        1 / 1                 30,178 / 6,691        733,573 / 19,465        733,573 / 5,837
GETATTR          61 / 5                1 / 1                 63 / 5                177,623 / 5             177,622 / 4
ACCESS           2 / 2                 2 / 2                 2 / 2                 2 / 2                   2 / 2
elapsed (ms)     919,684 / 40,657      23,938 / 24,984       947,506 / 41,133      2,643,906 / 78,684      2,627,536 / 38,337


- Iteration 2 (warm) - without eBPF / with eBPF

op               ls_r                  ls_lr                 lsmix                 du                      find
---------------  --------------------  --------------------  --------------------  ----------------------  ----------------------
READDIR calls    217 / 0               0 / 0                 258 / 256             124 / 317               123 / 417
READDIR recv(B)  208,601,368 / 0       0 / 0                 248,909,816 /         109,134,232 /           109,134,232 /
                                                             248,909,816           398,379,976             398,379,976
LOOKUP           0 / 0                 0 / 0                 0 / 0                 0 / 0                   0 / 0
GETATTR          301,214 / 1           1 / 1                 127 / 127             822,266 / 26,329        822,267 / 25,611
ACCESS           0 / 0                 0 / 0                 0 / 0                 0 / 0                   0 / 0
elapsed (ms)     550,011 / 2,648       5,231 / 6,419         23,264 / 24,090       1,494,821 / 65,103      1,527,773 / 64,990


You can see the reduction in total time, GETATTRs and LOOKUPs. 

-- 
Best,
Piyush Sachdeva <s.piyush1024@gmail.com>

  reply	other threads:[~2026-09-09 12:07 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-23 10:01 NFS delegations behavior analysis Piyush Sachdeva
2026-06-23 10:50 ` Jeff Layton
2026-06-23 11:04   ` Mkrtchyan, Tigran
2026-06-23 11:10     ` Jeff Layton
2026-06-23 13:11       ` Benjamin Coddington
2026-06-23 13:31         ` Daire Byrne
2026-06-23 13:32         ` Benjamin Coddington
2026-06-23 13:40           ` Jeff Layton
2026-06-23 13:59             ` Benjamin Coddington
2026-06-23 16:29           ` Trond Myklebust
2026-06-23 13:33         ` Jeff Layton
2026-06-23 13:11     ` Anna Schumaker
2026-09-09 10:10       ` Piyush Sachdeva
2026-09-09 12:07         ` Piyush Sachdeva [this message]
2026-09-09 13:13         ` Jeff Layton

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=m25x0ewrv7.fsf@gmail.com \
    --to=s.piyush1024@gmail.com \
    --cc=anna@nowheycreamery.com \
    --cc=cel@kernel.org \
    --cc=jlayton@kernel.org \
    --cc=linux-nfs@vger.kernel.org \
    --cc=sfrench@samba.org \
    --cc=sprasad@microsoft.com \
    --cc=tigran.mkrtchyan@desy.de \
    --cc=trondmy@kernel.org \
    --cc=vaibsharma@microsoft.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox