* [PATCH] exportfs: Release the get_name() directory file synchronously
@ 2026-10-07 16:12 Chuck Lever
0 siblings, 0 replies; 5+ messages in thread
From: Chuck Lever @ 2026-10-07 16:12 UTC (permalink / raw)
To: danny_pidutti; +Cc: Christian Brauner, Al Viro, Amir Goldstein, linux-fsdevel
get_name() opens the parent directory with dentry_open() and
releases it with fput(). From a kernel thread, fput() defers the
final __fput() to the delayed_fput work item. NFSD threads are
kernel threads, and NFSD reaches get_name() through
exportfs_decode_fh() whenever a file handle's dentry is not
connected to the root.
On a filesystem with more directories than the dcache retains,
nearly every READDIRPLUS needs a reconnect. The deferred releases
arrive faster than the work item retires them, and each pending
file pins the directory's readdir state. On ext4 that state is the
htree fname cache for the last leaf block read. A 64-thread NFSv3
server under a parallel tree walk showed unreclaimable slab growing
by about 90 MB per second, and the same workload on a 5.15 kernel
ended in a global OOM.
Commit 5ff318f645eb ("nfsd: use __fput_sync() to avoid delayed
closing of files.") stopped NFSD's own closes from feeding the
delayed_fput queue but did not cover the exportfs reconnect path.
Release the file with __fput_sync(). get_name() opened the file
itself, read-only, on a directory, so __fput() involves nothing
that could wait on the caller.
Fixes: 4a9d4b024a31 ("switch fput to task_work_add")
Reported-by: Danny Pidutti <danny_pidutti@hotmail.com>
Closes: https://lore.kernel.org/linux-fsdevel/DS4PR04MB9846BA3C16BC5ACED8B589228D952@DS4PR04MB9846.namprd04.prod.outlook.com/
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/exportfs/expfs.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
diff --git a/fs/exportfs/expfs.c b/fs/exportfs/expfs.c
index eafd99507afe..b5768f94fbb5 100644
--- a/fs/exportfs/expfs.c
+++ b/fs/exportfs/expfs.c
@@ -336,7 +336,11 @@ static int get_name(const struct path *path, char *name, struct dentry *child)
}
out_close:
- fput(file);
+ /*
+ * @file is the read-only directory opened above, so __fput()
+ * cannot block on anything the caller holds.
+ */
+ __fput_sync(file);
out:
return error;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread* [PATCH] exportfs: Release the get_name() directory file synchronously
@ 2026-10-07 16:46 Chuck Lever
[not found] ` <LV2PR15MB538107FDEB65AB1A4CEA8D5EC6942@LV2PR15MB5381.namprd15.prod.outlook.com>
2026-10-08 13:01 ` Benjamin Coddington
0 siblings, 2 replies; 5+ messages in thread
From: Chuck Lever @ 2026-10-07 16:46 UTC (permalink / raw)
To: Danny, Christian Brauner, Al Viro, Amir Goldstein
Cc: linux-fsdevel, linux-nfs, Danny Pidutti
get_name() opens the parent directory with dentry_open() and
releases it with fput(). From a kernel thread, fput() defers the
final __fput() to the delayed_fput work item. NFSD threads are
kernel threads, and NFSD reaches get_name() through
exportfs_decode_fh() whenever a file handle's dentry is not
connected to the root.
On a filesystem with more directories than the dcache retains,
nearly every READDIRPLUS needs a reconnect. The deferred releases
arrive faster than the work item retires them, and each pending
file pins the directory's readdir state. On ext4 that state is the
htree fname cache for the last leaf block read. A 64-thread NFSv3
server under a parallel tree walk showed unreclaimable slab growing
by about 90 MB per second, and the same workload on a 5.15 kernel
ended in a global OOM.
Commit 5ff318f645eb ("nfsd: use __fput_sync() to avoid delayed
closing of files.") stopped NFSD's own closes from feeding the
delayed_fput queue but did not cover the exportfs reconnect path.
Release the file with __fput_sync(). get_name() opened the file
itself, read-only, on a directory, so __fput() involves nothing
that could wait on the caller.
Fixes: 4a9d4b024a31 ("switch fput to task_work_add")
Reported-by: Danny Pidutti <danny_pidutti@hotmail.com>
Closes: https://lore.kernel.org/linux-fsdevel/DS4PR04MB9846BA3C16BC5ACED8B589228D952@DS4PR04MB9846.namprd04.prod.outlook.com/
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/exportfs/expfs.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
Mail from kernel.org to hotmail.com was rejected. kernel.org appears
to be on some deny-list. Reposting.
diff --git a/fs/exportfs/expfs.c b/fs/exportfs/expfs.c
index eafd99507afe..b5768f94fbb5 100644
--- a/fs/exportfs/expfs.c
+++ b/fs/exportfs/expfs.c
@@ -336,7 +336,11 @@ static int get_name(const struct path *path, char *name, struct dentry *child)
}
out_close:
- fput(file);
+ /*
+ * @file is the read-only directory opened above, so __fput()
+ * cannot block on anything the caller holds.
+ */
+ __fput_sync(file);
out:
return error;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 5+ messages in thread[parent not found: <LV2PR15MB538107FDEB65AB1A4CEA8D5EC6942@LV2PR15MB5381.namprd15.prod.outlook.com>]
* Re: [PATCH] exportfs: Release the get_name() directory file synchronously
[not found] ` <LV2PR15MB538107FDEB65AB1A4CEA8D5EC6942@LV2PR15MB5381.namprd15.prod.outlook.com>
@ 2026-10-07 19:03 ` Danny Pidutti
2026-10-07 21:43 ` Chuck Lever
0 siblings, 1 reply; 5+ messages in thread
From: Danny Pidutti @ 2026-10-07 19:03 UTC (permalink / raw)
To: Daniel Pidutti, Chuck Lever, Christian Brauner, Al Viro,
Amir Goldstein
Cc: linux-fsdevel@vger.kernel.org, linux-nfs@vger.kernel.org
Thanks for the quick patch. The description matches what I saw on
our server.
I tested it on the kernel where I found the problem: Ubuntu
7.0.0-38 (7.0.14) built with this patch only, 64 nfsd threads,
vfs_cache_pressure at the default of 100, and the same parallel tree
walk over about 43 million files and 5.5 million directories.
Before the patch, unreclaimable slab grew by about 90 MB per second
and reached the 8 GB limit of my test script. With the patch it
peaked at 1.5 GB in the first two minutes, came down to 0.6 GB while
all 64 threads were still busy, and stayed flat until the walk ended.
Open files went back to the idle count once the walk finished.
One question about the stable backport. The Fixes: tag goes back to
3.10, so stable may pick this up for 5.15 and 6.1. As far as I
remember, __fput_sync() on those kernels had a BUG_ON for callers
that are not kernel threads, and that was removed around 6.6.
get_name() can also be reached from open_by_handle_at() in process
context. Could you check whether the older stable kernels need a
different fix, or a guard on current->flags & PF_KTHREAD?
Danny Pidutti
________________________________________
From: Chuck Lever <cel@kernel.org>
Sent: Wednesday, October 7, 2026 12:46 PM
To: Daniel Pidutti <Danny@lightningiq.io>; Christian Brauner <brauner@kernel.org>; Al Viro <viro@zeniv.linux.org.uk>; Amir Goldstein <amir73il@gmail.com>
Cc: linux-fsdevel@vger.kernel.org <linux-fsdevel@vger.kernel.org>; linux-nfs@vger.kernel.org <linux-nfs@vger.kernel.org>; Danny Pidutti <danny_pidutti@hotmail.com>
Subject: [PATCH] exportfs: Release the get_name() directory file synchronously
get_name() opens the parent directory with dentry_open() and
releases it with fput(). From a kernel thread, fput() defers the
final __fput() to the delayed_fput work item. NFSD threads are
kernel threads, and NFSD reaches get_name() through
exportfs_decode_fh() whenever a file handle's dentry is not
connected to the root.
On a filesystem with more directories than the dcache retains,
nearly every READDIRPLUS needs a reconnect. The deferred releases
arrive faster than the work item retires them, and each pending
file pins the directory's readdir state. On ext4 that state is the
htree fname cache for the last leaf block read. A 64-thread NFSv3
server under a parallel tree walk showed unreclaimable slab growing
by about 90 MB per second, and the same workload on a 5.15 kernel
ended in a global OOM.
Commit 5ff318f645eb ("nfsd: use __fput_sync() to avoid delayed
closing of files.") stopped NFSD's own closes from feeding the
delayed_fput queue but did not cover the exportfs reconnect path.
Release the file with __fput_sync(). get_name() opened the file
itself, read-only, on a directory, so __fput() involves nothing
that could wait on the caller.
Fixes: 4a9d4b024a31 ("switch fput to task_work_add")
Reported-by: Danny Pidutti <danny_pidutti@hotmail.com>
Closes: https://lore.kernel.org/linux-fsdevel/DS4PR04MB9846BA3C16BC5ACED8B589228D952@DS4PR04MB9846.namprd04.prod.outlook.com/
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/exportfs/expfs.c | 6 +++++-
1 file changed, 5 insertions(+), 1 deletion(-)
Mail from kernel.org to hotmail.com was rejected. kernel.org appears
to be on some deny-list. Reposting.
diff --git a/fs/exportfs/expfs.c b/fs/exportfs/expfs.c
index eafd99507afe..b5768f94fbb5 100644
--- a/fs/exportfs/expfs.c
+++ b/fs/exportfs/expfs.c
@@ -336,7 +336,11 @@ static int get_name(const struct path *path, char *name, struct dentry *child)
}
out_close:
- fput(file);
+ /*
+ * @file is the read-only directory opened above, so __fput()
+ * cannot block on anything the caller holds.
+ */
+ __fput_sync(file);
out:
return error;
}
--
2.55.0
Disclaimer
The information contained in this communication from the sender is confidential. It is intended solely for use by the recipient and others authorized to receive it. If you are not the recipient, you are hereby notified that any disclosure, copying, distribution or taking action in relation of the contents of this information is strictly prohibited and may be unlawful.
This email has been scanned for viruses and malware, and may have been automatically archived by Mimecast, a leader in email security and cyber resilience. Mimecast integrates email defenses with brand protection, security awareness training, web security, compliance and other essential capabilities. Mimecast helps protect large and small organizations from malicious activity, human error and technology failure; and to lead the movement toward building a more resilient world. To find out more, visit our website.
^ permalink raw reply related [flat|nested] 5+ messages in thread* Re: [PATCH] exportfs: Release the get_name() directory file synchronously
2026-10-07 19:03 ` Danny Pidutti
@ 2026-10-07 21:43 ` Chuck Lever
0 siblings, 0 replies; 5+ messages in thread
From: Chuck Lever @ 2026-10-07 21:43 UTC (permalink / raw)
To: Daniel Pidutti
Cc: Christian Brauner, Al Viro, Amir Goldstein,
linux-fsdevel@vger.kernel.org, linux-nfs@vger.kernel.org
On 10/7/26 3:03 PM, Danny Pidutti wrote:
> Thanks for the quick patch. The description matches what I saw on
> our server.
>
> I tested it on the kernel where I found the problem: Ubuntu
> 7.0.0-38 (7.0.14) built with this patch only, 64 nfsd threads,
> vfs_cache_pressure at the default of 100, and the same parallel tree
> walk over about 43 million files and 5.5 million directories.
>
> Before the patch, unreclaimable slab grew by about 90 MB per second
> and reached the 8 GB limit of my test script. With the patch it
> peaked at 1.5 GB in the first two minutes, came down to 0.6 GB while
> all 64 threads were still busy, and stayed flat until the walk ended.
> Open files went back to the idle count once the walk finished.
Thank you for testing. May I add your Tested-by: to the patch?
> One question about the stable backport. The Fixes: tag goes back to
> 3.10, so stable may pick this up for 5.15 and 6.1. As far as I
> remember, __fput_sync() on those kernels had a BUG_ON for callers
> that are not kernel threads, and that was removed around 6.6.
> get_name() can also be reached from open_by_handle_at() in process
> context. Could you check whether the older stable kernels need a
> different fix, or a guard on current->flags & PF_KTHREAD?
Confirmed: the tip of linux-5.15.y and linux-6.1.y still have
BUG_ON(!(task->flags & PF_KTHREAD)) in __fput_sync(); it was
removed by commit 021a160abf62 ("fs: use __fput_sync in close(2)"),
which went into v6.6.
The Fixes: tag has to remain unchanged, since it names the commit
that introduced the behavior. Since this fix will go through one of
Christian's VFS trees, I can post a v2 with a recommended Cc: stable
tag, such as:
Cc: <stable@vger.kernel.org> # see patch description, needs adjustments for <= 6.1
and note in the patch description that kernels before v6.6 need the
__fput_sync() call guarded by current->flags & PF_KTHREAD, with a
plain fput() otherwise. Among the choices sanctioned in Documentation/
this one comes closest.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH] exportfs: Release the get_name() directory file synchronously
2026-10-07 16:46 Chuck Lever
[not found] ` <LV2PR15MB538107FDEB65AB1A4CEA8D5EC6942@LV2PR15MB5381.namprd15.prod.outlook.com>
@ 2026-10-08 13:01 ` Benjamin Coddington
1 sibling, 0 replies; 5+ messages in thread
From: Benjamin Coddington @ 2026-10-08 13:01 UTC (permalink / raw)
To: Chuck Lever
Cc: Danny, Christian Brauner, Al Viro, Amir Goldstein, linux-fsdevel,
linux-nfs, Danny Pidutti
On 7 Oct 2026, at 12:46, Chuck Lever wrote:
> get_name() opens the parent directory with dentry_open() and
> releases it with fput(). From a kernel thread, fput() defers the
> final __fput() to the delayed_fput work item. NFSD threads are
> kernel threads, and NFSD reaches get_name() through
> exportfs_decode_fh() whenever a file handle's dentry is not
> connected to the root.
>
> On a filesystem with more directories than the dcache retains,
> nearly every READDIRPLUS needs a reconnect. The deferred releases
> arrive faster than the work item retires them, and each pending
> file pins the directory's readdir state. On ext4 that state is the
> htree fname cache for the last leaf block read. A 64-thread NFSv3
> server under a parallel tree walk showed unreclaimable slab growing
> by about 90 MB per second, and the same workload on a 5.15 kernel
> ended in a global OOM.
>
> Commit 5ff318f645eb ("nfsd: use __fput_sync() to avoid delayed
> closing of files.") stopped NFSD's own closes from feeding the
> delayed_fput queue but did not cover the exportfs reconnect path.
>
> Release the file with __fput_sync(). get_name() opened the file
> itself, read-only, on a directory, so __fput() involves nothing
> that could wait on the caller.
>
> Fixes: 4a9d4b024a31 ("switch fput to task_work_add")
> Reported-by: Danny Pidutti <danny_pidutti@hotmail.com>
> Closes: https://lore.kernel.org/linux-fsdevel/DS4PR04MB9846BA3C16BC5ACED8B589228D952@DS4PR04MB9846.namprd04.prod.outlook.com/
> Signed-off-by: Chuck Lever <cel@kernel.org>
Reviewed-by: Benjamin Coddington <bcodding@hammerspace.com>
Ben
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-08 13:01 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07 16:12 [PATCH] exportfs: Release the get_name() directory file synchronously Chuck Lever
-- strict thread matches above, loose matches on Subject: below --
2026-10-07 16:46 Chuck Lever
[not found] ` <LV2PR15MB538107FDEB65AB1A4CEA8D5EC6942@LV2PR15MB5381.namprd15.prod.outlook.com>
2026-10-07 19:03 ` Danny Pidutti
2026-10-07 21:43 ` Chuck Lever
2026-10-08 13:01 ` Benjamin Coddington
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox