From: NeilBrown <neilb@ownmail.net>
To: Alexander Viro <viro@zeniv.linux.org.uk>,
Christian Brauner <brauner@kernel.org>,
Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>,
Jori Koolstra <jkoolstra@xs4all.nl>,
Mateusz Guzik <mjguzik@gmail.com>,
Dorjoy Chowdhury <dorjoychy111@gmail.com>
Cc: Trond Myklebust <trondmy@kernel.org>,
Anna Schumaker <anna@kernel.org>,
Andreas Gruenbacher <agruenba@redhat.com>,
gfs2@lists.linux.dev, Ilya Dryomov <idryomov@gmail.com>,
Alex Markuze <amarkuze@redhat.com>,
Viacheslav Dubeyko <slava@dubeyko.com>,
ceph-devel@vger.kernel.org, Paulo Alcantara <pc@manguebit.org>,
Namjae Jeon <linkinjeon@kernel.org>,
linux-cifs@vger.kernel.org, linux-fsdevel@vger.kernel.org,
linux-nfs@vger.kernel.org
Subject: [PATCH v2 01/14] VFS: revise and expand documentation for atomic_open.
Date: Sat, 19 Sep 2026 12:06:05 +1000 [thread overview]
Message-ID: <20260919022441.3305170-2-neilb@ownmail.net> (raw)
In-Reply-To: <20260919022441.3305170-1-neilb@ownmail.net>
From: NeilBrown <neil@brown.name>
atomic_open is a complex operation which different filesystems implement
quite differently. The available documentation doesn't give clear
guidance on how it should be implemented.
nfsd has a particular need to open only regular files, but to get
precise information about what was found if it wasn't a regular file.
This is slightly different to the syscall calling needs. In particular
it suggests that __O_REGULAR shouldn't always result in -EFTYPE.
In any case that does involve creating open state, using
finish_no_open() is simplest as it reduces the need to check
__O_REGULAR, O_DIRECTORY, O_NOFOLLOW.
So refresh the documentation to give guidance on the choice between
finish_no_open, finish_open, and an error. Efficiency always wins, but
when that isn't an issue, prefer finish_no_open().
Also clarify the required behaviour when __O_REGULAR is given. This
should return -EISDIR if a directory is found as nfsd needs this. If a
symlink is found then __O_REGULAR does NOT apply: O_NOFOLLOW must be
used to decided if it is safe to not return the looked-up dentry.
Signed-off-by: NeilBrown <neil@brown.name>
---
Documentation/filesystems/vfs.rst | 67 ++++++++++++++++++++++++++-----
fs/namei.c | 3 ++
2 files changed, 59 insertions(+), 11 deletions(-)
diff --git a/Documentation/filesystems/vfs.rst b/Documentation/filesystems/vfs.rst
index d3a93eec3945..00ada8cc85ae 100644
--- a/Documentation/filesystems/vfs.rst
+++ b/Documentation/filesystems/vfs.rst
@@ -599,17 +599,62 @@ otherwise noted.
``atomic_open``
called on the last component of an open. Using this optional
- method the filesystem can look up, possibly create and open the
- file in one atomic operation. If it wants to leave actual
- opening to the caller (e.g. if the file turned out to be a
- symlink, device, or just something filesystem won't do atomic
- open for), it may signal this by returning finish_no_open(file,
- dentry). This method is only called if the last component is
- negative or needs lookup. Cached positive dentries are still
- handled by f_op->open(). If the file was created, FMODE_CREATED
- flag should be set in file->f_mode. In case of O_EXCL the
- method must only succeed if the file didn't exist and hence
- FMODE_CREATED shall always be set on success.
+ method the filesystem can look up, create, truncate, and open
+ the file in one atomic operation. This is needed if the
+ filesystem content can be changed asynchronously and
+ specifically if a negative dentry is not a guarantee that the
+ object doesn't exist. It is also useful if it is possible to
+ perform combinations of revalidate, lookup, create, open, and
+ truncate more efficiently what with a sequence of individual
+ operations.
+
+ If the object found is not a file or directory, or if
+ lookup/create succeeded without establishing any "open" state,
+ then finish_no_open() should be called to confirm that the
+ dentry is ready to be handled by normal VFS processing.
+ FMODE_CREATED should be set in the "file" if the object was
+ created, and this will prevent further access permission checks,
+ or handling of O_TRUNC and O_EXCL.
+
+ If the lookup/create operation established some open state for a
+ file or directory, the open should be completed by calling
+ finish_open(). Passing NULL as the "open" function to
+ finish_open() is unlikely to be useful as that assumes that no
+ open state has been established.
+
+ atomic_open() may generate errors related to O_DIRECTORY,
+ __O_REGULAR, O_EXCL, O_NOFOLLOW but is not required to as the
+ caller will check those against the resulting dentry and
+ generate any error needed, possibly closing the file if it was
+ opened by finish_open(). atomic_open() is encouraged to handle
+ these flags only when doing so is more efficient than not.
+
+ If __O_REGULAR is handled, it should generate -EISDIR if the
+ name is known to be a directory or -EFTYPE if it is some other
+ non-regular file other than a symbolic link. Handling of a
+ symbolic link should be guided by O_NOFOLLOW, not __O_REGULAR:
+ -ELOOP can be return if O_NOFOLLOW is set, otherwise the symlink
+ should be returned through finish_no_open().
+
+ The focus for atomic_open() is to provide the correct dentry and
+ to set FMODE_CREATED as accurately as possible. If O_EXCL was
+ set, FMODE_CREATED should only be set if this operation
+ certainly created the object. If O_EXCL was not set,
+ FMODE_CREATE should be set if it is possible that this operation
+ created the object.
+
+ This method is only called if the last component is negative or
+ needs lookup. Cached positive dentries are still handled by
+ f_op->open().
+
+ If the dentry provided is negative (not in-lookup) and O_CREAT
+ isn't set, then there is no guarantee of exclusive access to the
+ dentry - another thread might call ->atomic_open() on the same
+ dentry at the same time. If needed a filesystem can ensure this
+ doesn't happen by returning 0 from ->d_revalidate when that is
+ called with LOOKUP_OPEN on a negative dentry. This will ensure
+ that ->atomic_open() only receives an in-lookup dentry, which
+ always ensures exclusive access.
``tmpfile``
called in the end of O_TMPFILE open(). Optional, equivalent to
diff --git a/fs/namei.c b/fs/namei.c
index d95249dd527c..0f69abb3743b 100644
--- a/fs/namei.c
+++ b/fs/namei.c
@@ -5007,6 +5007,9 @@ static struct file *path_openat(struct nameidata *nd,
error = -EINVAL;
}
fput_close(file);
+ if (error == -EISDIR &&
+ (op->open_flag & __O_REGULAR))
+ error = -EFTYPE;
if (error == -EOPENSTALE) {
if (flags & LOOKUP_RCU)
error = -ECHILD;
base-commit: 9189e6a6f89e32d3a604b221ea64e67e1a35957c
--
2.50.0.107.gf914562f5916.dirty
next prev parent reply other threads:[~2026-09-19 2:25 UTC|newest]
Thread overview: 47+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-19 2:06 [PATCH v2 01/14] fixes for vfs_lookup_open, and integration with nfsd NeilBrown
2026-09-19 2:06 ` NeilBrown [this message]
2026-09-24 12:25 ` [PATCH v2 01/14] VFS: revise and expand documentation for atomic_open Jeff Layton
2026-09-19 2:06 ` [PATCH v2 02/14] nfs: correctly handle NFS4ERR_WRONG_TYPE from v4 OPEN request NeilBrown
2026-09-24 12:55 ` Jeff Layton
2026-09-24 13:01 ` Jeff Layton
2026-09-19 2:06 ` [PATCH v2 03/14] gfs2: simplify atomic_open handling NeilBrown
2026-09-19 16:49 ` Andreas Gruenbacher
2026-09-19 22:22 ` NeilBrown
2026-09-20 16:31 ` Andreas Gruenbacher
2026-09-22 21:16 ` NeilBrown
2026-09-19 2:06 ` [PATCH v2 04/14] ceph: simplify atomic_open to use finish_no_open() NeilBrown
2026-09-24 13:02 ` Jeff Layton
2026-09-25 21:49 ` NeilBrown
2026-09-26 11:48 ` Jeff Layton
2026-09-19 2:06 ` [PATCH v2 05/14] cifs: allow -EISDIR precedence over -EFTYPE in __cifs_do_create() NeilBrown
2026-09-23 5:54 ` Namjae Jeon
2026-09-23 6:50 ` NeilBrown
2026-09-23 7:52 ` Namjae Jeon
2026-09-23 8:13 ` Namjae Jeon
2026-09-19 2:06 ` [PATCH v2 06/14] vfs: add some allowed open flags to vfs_lookup_open() NeilBrown
2026-09-24 12:54 ` Jeff Layton
2026-09-29 15:18 ` Jori Koolstra
2026-09-29 22:04 ` NeilBrown
2026-09-19 2:06 ` [PATCH v2 07/14] vfs: O_NONBLOCK|O_CREAT open shouldn't wait for directory delegation NeilBrown
2026-09-24 12:51 ` Jeff Layton
2026-09-19 2:06 ` [PATCH v2 08/14] vfs: don't return -ENODEV from vfs_lookup_open() NeilBrown
2026-09-24 13:07 ` Jeff Layton
2026-09-19 2:06 ` [PATCH v2 09/14] vfs: change vfs_lookup_open() to use do_open(), not vfs_open() NeilBrown
2026-09-24 13:13 ` Jeff Layton
2026-09-19 2:06 ` [PATCH v2 10/14] nfsd: make EEXIST checks in nfsd4_create_file() more consistent NeilBrown
2026-09-24 13:23 ` Jeff Layton
2026-09-25 21:57 ` NeilBrown
2026-09-19 2:06 ` [PATCH v2 11/14] nfsd: check for mountpoints after non-creating open NeilBrown
2026-09-24 13:03 ` Jeff Layton
2026-09-25 22:14 ` NeilBrown
2026-09-19 2:06 ` [PATCH v2 12/14] nfsd: switch NFS4 OPEN to use vfs_lookup_open() NeilBrown
2026-09-20 17:10 ` Chuck Lever
2026-09-22 21:55 ` NeilBrown
2026-09-23 4:07 ` NeilBrown
2026-09-19 2:06 ` [PATCH v2 13/14] nfsd: change nfsd_check_obj_isreg() to use nfs error codes NeilBrown
2026-09-24 13:25 ` Jeff Layton
2026-09-19 2:06 ` [PATCH v2 14/14] nfsd: use vfs_lookup_open() for non-creating open requests too NeilBrown
2026-09-20 17:13 ` Chuck Lever
2026-09-25 22:26 ` NeilBrown
2026-09-25 16:04 ` [PATCH v2 01/14] fixes for vfs_lookup_open, and integration with nfsd Christian Brauner
2026-09-25 16:56 ` Chuck Lever
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260919022441.3305170-2-neilb@ownmail.net \
--to=neilb@ownmail.net \
--cc=agruenba@redhat.com \
--cc=amarkuze@redhat.com \
--cc=anna@kernel.org \
--cc=brauner@kernel.org \
--cc=cel@kernel.org \
--cc=ceph-devel@vger.kernel.org \
--cc=dorjoychy111@gmail.com \
--cc=gfs2@lists.linux.dev \
--cc=idryomov@gmail.com \
--cc=jkoolstra@xs4all.nl \
--cc=jlayton@kernel.org \
--cc=linkinjeon@kernel.org \
--cc=linux-cifs@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=mjguzik@gmail.com \
--cc=neil@brown.name \
--cc=pc@manguebit.org \
--cc=slava@dubeyko.com \
--cc=trondmy@kernel.org \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox