* [PATCH v3 1/9] dax: replace exported dax_dev_get() with non-allocating dax_dev_find()
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
2026-10-06 18:10 ` sashiko-bot
2026-10-06 18:01 ` [PATCH v3 2/9] dax: use READ_ONCE() in dax_holder() Miklos Szeredi
` (7 subsequent siblings)
8 siblings, 1 reply; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, John Groves, Amir Goldstein, Darrick J . Wong,
Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl
From: John Groves <John@Groves.net>
This fix is in response to a Sashiko review, and some subsequent
analysis.
dax_dev_get() uses iget5_locked() which creates a new inode if no
matching one exists. This is correct for the internal caller
(alloc_dax), but dangerous for external callers that look up devices
from user-supplied or metadata-supplied dev_t values:
1. A new inode is created with DAXDEV_ALIVE set but no backing driver,
no ops, and no IDA-allocated minor number.
2. On teardown, dax_destroy_inode() warns because kill_dax() was never
called, and dax_free_inode() calls ida_free() for a minor that was
never ida_alloc'd -- potentially freeing the minor of a real device.
Add dax_dev_find() which uses ilookup5() for lookup-only semantics:
it returns an existing dax_device with an elevated inode reference, or
NULL if no device with the given dev_t exists. It never creates inodes.
A dax_alive() check under dax_read_lock() guards against returning a
device that is concurrently being torn down by kill_dax().
Make dax_dev_get() static again (internal to super.c for alloc_dax),
export dax_dev_find() instead. Also add the missing CONFIG_DAX=n stub.
About the 'fixes' tag: this removes the export of dax_dev_get(),
which was flawed, and replaces is with dax_dev_find(). It feels like
the fixes tag makes sense for correcting an ABI error.
Fixes: 2ae624d5a555d ("dax: export dax_dev_get()")
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
Signed-off-by: John Groves <john@groves.net>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
drivers/dax/super.c | 38 ++++++++++++++++++++++++++++++++++++--
include/linux/dax.h | 6 +++++-
2 files changed, 41 insertions(+), 3 deletions(-)
diff --git a/drivers/dax/super.c b/drivers/dax/super.c
index 45f84b0eb909..824e1f6df378 100644
--- a/drivers/dax/super.c
+++ b/drivers/dax/super.c
@@ -565,7 +565,7 @@ static int dax_set(struct inode *inode, void *data)
return 0;
}
-struct dax_device *dax_dev_get(dev_t devt)
+static struct dax_device *dax_dev_get(dev_t devt)
{
struct dax_device *dax_dev;
struct inode *inode;
@@ -588,7 +588,41 @@ struct dax_device *dax_dev_get(dev_t devt)
return dax_dev;
}
-EXPORT_SYMBOL_GPL(dax_dev_get);
+
+/**
+ * dax_dev_find - look up an existing dax_device by dev_t
+ * @devt: the device number to find
+ *
+ * Returns a dax_device with an elevated inode reference, or NULL if no
+ * device with the given dev_t exists. Unlike dax_dev_get(), this never
+ * allocates a new inode -- it is safe for external callers that are looking
+ * up devices from user-supplied or metadata-supplied dev_t values.
+ *
+ * Caller must put_dax() the returned device when done.
+ */
+struct dax_device *dax_dev_find(dev_t devt)
+{
+ struct dax_device *dax_dev;
+ struct inode *inode;
+ int id;
+
+ inode = ilookup5(dax_superblock, hash_32(devt + DAXFS_MAGIC, 31),
+ dax_test, &devt);
+ if (!inode)
+ return NULL;
+
+ dax_dev = to_dax_dev(inode);
+ id = dax_read_lock();
+ if (!dax_alive(dax_dev)) {
+ dax_read_unlock(id);
+ iput(inode);
+ return NULL;
+ }
+ dax_read_unlock(id);
+
+ return dax_dev;
+}
+EXPORT_SYMBOL_GPL(dax_dev_find);
struct dax_device *alloc_dax(void *private, const struct dax_operations *ops)
{
diff --git a/include/linux/dax.h b/include/linux/dax.h
index fe6c3ded1b50..29113eb95e72 100644
--- a/include/linux/dax.h
+++ b/include/linux/dax.h
@@ -54,7 +54,7 @@ struct dax_device *alloc_dax(void *private, const struct dax_operations *ops);
void *dax_holder(struct dax_device *dax_dev);
void put_dax(struct dax_device *dax_dev);
void kill_dax(struct dax_device *dax_dev);
-struct dax_device *dax_dev_get(dev_t devt);
+struct dax_device *dax_dev_find(dev_t devt);
void dax_write_cache(struct dax_device *dax_dev, bool wc);
bool dax_write_cache_enabled(struct dax_device *dax_dev);
bool dax_synchronous(struct dax_device *dax_dev);
@@ -92,6 +92,10 @@ static inline void put_dax(struct dax_device *dax_dev)
static inline void kill_dax(struct dax_device *dax_dev)
{
}
+static inline struct dax_device *dax_dev_find(dev_t devt)
+{
+ return NULL;
+}
static inline void dax_write_cache(struct dax_device *dax_dev, bool wc)
{
}
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v3 1/9] dax: replace exported dax_dev_get() with non-allocating dax_dev_find()
2026-10-06 18:01 ` [PATCH v3 1/9] dax: replace exported dax_dev_get() with non-allocating dax_dev_find() Miklos Szeredi
@ 2026-10-06 18:10 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-10-06 18:10 UTC (permalink / raw)
To: Miklos Szeredi; +Cc: nvdimm, linux-cxl
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] dax: partially initialized dax_device returned by dax_dev_find() due to premature unlock_new_inode()
--
commit 9b2ac5a45a1d5eca0b23e84e793ce64303f44b4e
Author: John Groves <John@Groves.net>
dax: replace exported dax_dev_get() with non-allocating dax_dev_find()
dax_dev_get() uses iget5_locked() which creates a new inode if no
matching one exists, which is dangerous for external callers. This patch
introduces dax_dev_find() to provide lookup-only semantics without allocating
new inodes, and restricts dax_dev_get() to internal allocation use.
> diff --git a/drivers/dax/super.c b/drivers/dax/super.c
> index 45f84b0eb909a..824e1f6df3780 100644
> --- a/drivers/dax/super.c
> +++ b/drivers/dax/super.c
[ ... ]
> @@ -588,7 +588,41 @@ struct dax_device *dax_dev_get(dev_t devt)
>
> return dax_dev;
> }
> -EXPORT_SYMBOL_GPL(dax_dev_get);
> +
> +/**
> + * dax_dev_find - look up an existing dax_device by dev_t
> + * @devt: the device number to find
> + *
> + * Returns a dax_device with an elevated inode reference, or NULL if no
> + * device with the given dev_t exists. Unlike dax_dev_get(), this never
> + * allocates a new inode -- it is safe for external callers that are looking
> + * up devices from user-supplied or metadata-supplied dev_t values.
> + *
> + * Caller must put_dax() the returned device when done.
> + */
> +struct dax_device *dax_dev_find(dev_t devt)
> +{
> + struct dax_device *dax_dev;
> + struct inode *inode;
> + int id;
> +
> + inode = ilookup5(dax_superblock, hash_32(devt + DAXFS_MAGIC, 31),
> + dax_test, &devt);
> + if (!inode)
> + return NULL;
> +
> + dax_dev = to_dax_dev(inode);
> + id = dax_read_lock();
> + if (!dax_alive(dax_dev)) {
^^^^^^^^^^^^^^^^^^^^^^^
[Severity: High]
Does this introduce a regression where a concurrent external caller of
dax_dev_find() can obtain a partially initialized device?
When a thread calls alloc_dax(), it calls dax_dev_get() to allocate the inode.
Inside dax_dev_get(), the new inode is instantiated, DAXDEV_ALIVE is set, and
unlock_new_inode() is called:
drivers/dax/super.c:dax_dev_get() {
...
if (inode_state_read_once(inode) & I_NEW) {
set_bit(DAXDEV_ALIVE, &dax_dev->flags);
...
unlock_new_inode(inode);
}
return dax_dev;
}
This clears the I_NEW flag, making the inode fully accessible to concurrent
VFS lookups before dax_dev_get() returns.
If a concurrent thread calls dax_dev_find() at this moment, ilookup5() will
find the unlocked inode and return it. Since DAXDEV_ALIVE is already set,
dax_dev_find() will successfully return the device.
However, the allocating thread has not yet initialized the operations and
private data in alloc_dax():
drivers/dax/super.c:alloc_dax() {
...
dax_dev = dax_dev_get(devt);
if (!dax_dev)
goto err_dev;
dax_dev->ops = ops;
dax_dev->private = private;
return dax_dev;
}
Will this sequence allow an external caller to access a dax_device with
uninitialized ops and private pointers?
> + dax_read_unlock(id);
> + iput(inode);
> + return NULL;
> + }
> + dax_read_unlock(id);
> +
> + return dax_dev;
> +}
> +EXPORT_SYMBOL_GPL(dax_dev_find);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006180115.1425232-1-mszeredi@redhat.com?part=1
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v3 2/9] dax: use READ_ONCE() in dax_holder()
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
2026-10-06 18:01 ` [PATCH v3 1/9] dax: replace exported dax_dev_get() with non-allocating dax_dev_find() Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
2026-10-06 18:13 ` sashiko-bot
2026-10-08 1:02 ` Alison Schofield
2026-10-06 18:01 ` [PATCH v3 3/9] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
` (6 subsequent siblings)
8 siblings, 2 replies; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
Make sure that a non-NULL check in callers remains valid even when racing
with fs_put_dax().
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
drivers/dax/super.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/dax/super.c b/drivers/dax/super.c
index 824e1f6df378..0a440f8b42e4 100644
--- a/drivers/dax/super.c
+++ b/drivers/dax/super.c
@@ -677,7 +677,7 @@ EXPORT_SYMBOL_GPL(put_dax);
*/
void *dax_holder(struct dax_device *dax_dev)
{
- return dax_dev->holder_data;
+ return READ_ONCE(dax_dev->holder_data);
}
EXPORT_SYMBOL_GPL(dax_holder);
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v3 2/9] dax: use READ_ONCE() in dax_holder()
2026-10-06 18:01 ` [PATCH v3 2/9] dax: use READ_ONCE() in dax_holder() Miklos Szeredi
@ 2026-10-06 18:13 ` sashiko-bot
2026-10-08 1:02 ` Alison Schofield
1 sibling, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-10-06 18:13 UTC (permalink / raw)
To: Miklos Szeredi; +Cc: nvdimm, linux-cxl
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] xfs: missing NULL check in xfs_dax_notify_failure() causes NULL pointer dereference
- [High] dax: missing synchronize_srcu() in fs_put_dax() leads to Use-After-Free of xfs_mount
--
commit b62dcd3b2d7873db8fcf1fc4a0c2b9c52417ad0a
Author: Miklos Szeredi <mszeredi@redhat.com>
dax: use READ_ONCE() in dax_holder()
Make sure that a non-NULL check in callers remains valid even when racing
with fs_put_dax().
> diff --git a/drivers/dax/super.c b/drivers/dax/super.c
> index 824e1f6df3780..0a440f8b42e4f 100644
> --- a/drivers/dax/super.c
> +++ b/drivers/dax/super.c
> @@ -677,7 +677,7 @@ EXPORT_SYMBOL_GPL(put_dax);
> */
> void *dax_holder(struct dax_device *dax_dev)
> {
> - return dax_dev->holder_data;
> + return READ_ONCE(dax_dev->holder_data);
[Severity: High]
Does this code still leave a window for a NULL pointer dereference?
The commit message mentions protecting a non-NULL check in callers, but the
caller in xfs_dax_notify_failure() doesn't appear to have one:
fs/xfs/xfs_notify_failure.c:xfs_dax_notify_failure() {
struct xfs_mount *mp = dax_holder(dax_dev);
if (!(mp->m_super->s_flags & SB_BORN)) {
...
}
If a concurrent unmount clears dax_dev->holder_data to NULL via
fs_put_dax(), won't this result in assigning NULL to mp and immediately
dereferencing mp->m_super?
[Severity: High]
Can a use-after-free occur on the returned xfs_mount pointer even if a NULL
check is added?
When a memory failure triggers dax_holder_notify_failure(), it acquires
dax_read_lock() (which is SRCU) and calls xfs_dax_notify_failure().
Concurrently, if an unmount calls fs_put_dax(), it clears the holder data:
drivers/dax/super.c:fs_put_dax() {
...
WRITE_ONCE(dax_dev->holder_ops, NULL);
prev = cmpxchg(&dax_dev->holder_data, holder, NULL);
...
}
Since fs_put_dax() returns without calling synchronize_srcu(&dax_srcu) to
wait for active readers, the unmount process can proceed and destroy the
xfs_mount structure.
Won't the still-executing xfs_dax_notify_failure() then access the freed
xfs_mount, resulting in a use-after-free?
> }
> EXPORT_SYMBOL_GPL(dax_holder);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006180115.1425232-1-mszeredi@redhat.com?part=2
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v3 2/9] dax: use READ_ONCE() in dax_holder()
2026-10-06 18:01 ` [PATCH v3 2/9] dax: use READ_ONCE() in dax_holder() Miklos Szeredi
2026-10-06 18:13 ` sashiko-bot
@ 2026-10-08 1:02 ` Alison Schofield
2026-10-08 6:58 ` Miklos Szeredi
1 sibling, 1 reply; 23+ messages in thread
From: Alison Schofield @ 2026-10-08 1:02 UTC (permalink / raw)
To: Miklos Szeredi
Cc: fuse-devel, John Groves, Amir Goldstein, Darrick J . Wong,
Vishal Verma, Dave Jiang, nvdimm, linux-cxl
On Tue, Oct 06, 2026 at 08:01:07PM +0200, Miklos Szeredi wrote:
> Make sure that a non-NULL check in callers remains valid even when racing
> with fs_put_dax().
Hi Miklos,
I understand this is in preparation for the new FUSE DAX backing support, yet it
does need a changelog that stands on its own.
As written, it suggests that READ_ONCE() makes a non-NULL result safe against
concurrent fs_put_dax(). However, READ_ONCE() only controls the pointer load, but
IIUC it does not protect the lifetime of the returned holder.
Could the changelog explain what READ_ONCE() provides and describe the intended
RCU-protected access? Include what guarantees that the holder remains valid
while it is being used.
Also, is this RCU-based lifetime handling specific to the new FUSE implementation,
or is it intended to establish a general requirement for DAX holders and callers
of dax_holder()?
-- Alison
>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
> ---
> drivers/dax/super.c | 2 +-
> 1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/drivers/dax/super.c b/drivers/dax/super.c
> index 824e1f6df378..0a440f8b42e4 100644
> --- a/drivers/dax/super.c
> +++ b/drivers/dax/super.c
> @@ -677,7 +677,7 @@ EXPORT_SYMBOL_GPL(put_dax);
> */
> void *dax_holder(struct dax_device *dax_dev)
> {
> - return dax_dev->holder_data;
> + return READ_ONCE(dax_dev->holder_data);
> }
> EXPORT_SYMBOL_GPL(dax_holder);
>
> --
> 2.54.0
>
>
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v3 2/9] dax: use READ_ONCE() in dax_holder()
2026-10-08 1:02 ` Alison Schofield
@ 2026-10-08 6:58 ` Miklos Szeredi
0 siblings, 0 replies; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-08 6:58 UTC (permalink / raw)
To: Alison Schofield
Cc: Miklos Szeredi, fuse-devel, John Groves, Amir Goldstein,
Darrick J . Wong, Vishal Verma, Dave Jiang, nvdimm, linux-cxl
On Thu, 8 Oct 2026 at 03:03, Alison Schofield
<alison.schofield@intel.com> wrote:
> As written, it suggests that READ_ONCE() makes a non-NULL result safe against
> concurrent fs_put_dax(). However, READ_ONCE() only controls the pointer load, but
> IIUC it does not protect the lifetime of the returned holder.
Right, the changelog was a bit hurried. Will fix.
> Also, is this RCU-based lifetime handling specific to the new FUSE implementation,
> or is it intended to establish a general requirement for DAX holders and callers
> of dax_holder()?
I haven't looked. The RCU based fix only works as long as the memory
failure handler doesn't need to sleep. I cannot make a comment on
the xfs code.
Thanks,
Miklos
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v3 3/9] fuse: add helpers for EIO return value with kernel message
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
2026-10-06 18:01 ` [PATCH v3 1/9] dax: replace exported dax_dev_get() with non-allocating dax_dev_find() Miklos Szeredi
2026-10-06 18:01 ` [PATCH v3 2/9] dax: use READ_ONCE() in dax_holder() Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
2026-10-06 21:03 ` Amir Goldstein
2026-10-06 18:01 ` [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
` (5 subsequent siblings)
8 siblings, 1 reply; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
Conditions are triggered by a buggy fuse server will result in -EIO which
is hard to interpret without any context.
Add helpers that print a short message to the kernel log as well as
returning the error value. This serves a dual purpose:
- documents the error condition inside the code
- allows the implementor of the fuse server to get the context for the
error
Some debug messages are removed in favor of this.
This also changes the mmap return value from -ETXTBSY to -ENODEV if a
passthrough open raced with the prior check. This makes both cases return
-ENODEV if the file is in passthrough mode.
Additionally change the return value of fuse_file_cached_io_open() from an
error value to a bool (true on success) as the error value is translated
anyway.
Currently these use the pr_notice_once() variant. Possibly should be
changed to a ratelimit of e.g. once per minute.
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
fs/fuse/file.c | 6 ++--
fs/fuse/fuse_i.h | 4 ++-
fs/fuse/iomode.c | 75 ++++++++++++++++++-------------------------
fs/fuse/passthrough.c | 19 ++++-------
4 files changed, 43 insertions(+), 61 deletions(-)
diff --git a/fs/fuse/file.c b/fs/fuse/file.c
index 6d707f2b3bff..bde07948d1f0 100644
--- a/fs/fuse/file.c
+++ b/fs/fuse/file.c
@@ -2415,7 +2415,6 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma)
struct fuse_file *ff = file->private_data;
struct fuse_conn *fc = ff->fm->fc;
struct inode *inode = file_inode(file);
- int rc;
/* DAX mmap is superior to direct_io mmap */
if (FUSE_IS_VDAX(inode))
@@ -2457,9 +2456,8 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma)
* After first mmap, the inode stays in caching io mode until
* the direct_io file release.
*/
- rc = fuse_file_cached_io_open(inode, ff);
- if (rc)
- return rc;
+ if (!fuse_file_cached_io_open(inode, ff))
+ return -ENODEV;
}
if ((vma->vm_flags & VM_SHARED) && (vma->vm_flags & VM_MAYWRITE))
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 8546855386b5..d9ec88028d7b 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -11,6 +11,8 @@
# define pr_fmt(fmt) "fuse: " fmt
#endif
+#define fuse_EIO(fmt, ...) (pr_notice_once("%s: " fmt "\n", __func__, ##__VA_ARGS__), -EIO)
+
#include "args.h"
#include <linux/fuse.h>
#include <linux/fs.h>
@@ -1251,7 +1253,7 @@ int fuse_fileattr_set(struct mnt_idmap *idmap,
struct dentry *dentry, struct file_kattr *fa);
/* iomode.c */
-int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff);
+bool fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff);
int fuse_inode_uncached_io_start(struct fuse_inode *fi,
struct fuse_backing *fb);
void fuse_inode_uncached_io_end(struct fuse_inode *fi);
diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
index 79637c09e883..1b8b267d0c0d 100644
--- a/fs/fuse/iomode.c
+++ b/fs/fuse/iomode.c
@@ -27,15 +27,15 @@ static inline bool fuse_is_io_cache_wait(struct fuse_inode *fi)
* Blocks new parallel dio writes and waits for the in-progress parallel dio
* writes to complete.
*/
-int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
+bool fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
{
struct fuse_inode *fi = get_fuse_inode(inode);
/* There are no io modes if server does not implement open */
if (!ff->args)
- return 0;
+ return true;
- spin_lock(&fi->lock);
+ guard(spinlock)(&fi->lock);
/*
* Setting the bit advises new direct-io writes to use an exclusive
* lock - without it the wait below might be forever.
@@ -53,8 +53,7 @@ int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
*/
if (fuse_inode_backing(fi)) {
clear_bit(FUSE_I_CACHE_IO_MODE, &fi->state);
- spin_unlock(&fi->lock);
- return -ETXTBSY;
+ return false;
}
WARN_ON(ff->iomode == IOM_UNCACHED);
@@ -64,8 +63,7 @@ int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
set_bit(FUSE_I_CACHE_IO_MODE, &fi->state);
fi->iocachectr++;
}
- spin_unlock(&fi->lock);
- return 0;
+ return true;
}
static void fuse_file_cached_io_release(struct fuse_file *ff,
@@ -85,19 +83,16 @@ static void fuse_file_cached_io_release(struct fuse_file *ff,
int fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
{
struct fuse_backing *oldfb;
- int err = 0;
- spin_lock(&fi->lock);
+ guard(spinlock)(&fi->lock);
/* deny conflicting backing files on same fuse inode */
oldfb = fuse_inode_backing(fi);
- if (fb && oldfb && oldfb != fb) {
- err = -EBUSY;
- goto unlock;
- }
- if (fi->iocachectr > 0) {
- err = -ETXTBSY;
- goto unlock;
- }
+ if (fb && oldfb && oldfb != fb)
+ return -EBUSY;
+
+ if (fi->iocachectr > 0)
+ return -ETXTBSY;
+
fi->iocachectr--;
/* fuse inode holds a single refcount of backing file */
@@ -107,9 +102,7 @@ int fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
} else {
fuse_backing_put(fb);
}
-unlock:
- spin_unlock(&fi->lock);
- return err;
+ return 0;
}
/* Takes uncached_io inode mode reference to be dropped on file release */
@@ -121,8 +114,11 @@ static int fuse_file_uncached_io_open(struct inode *inode,
int err;
err = fuse_inode_uncached_io_start(fi, fb);
- if (err)
- return err;
+ if (err) {
+ if (err == -EBUSY)
+ return fuse_EIO("mismatched backing");
+ return fuse_EIO("conflicting caching mode");
+ }
WARN_ON(ff->iomode != IOM_NONE);
ff->iomode = IOM_UNCACHED;
@@ -173,9 +169,11 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
int err;
/* Check allowed conditions for file open in passthrough mode */
- if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) || !fc->passthrough ||
- (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK))
- return -EINVAL;
+ if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) || !fc->passthrough)
+ return fuse_EIO("passthrough not enabled");
+
+ if (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK)
+ return fuse_EIO("conflicting open flags");
fb = fuse_passthrough_open(file, ff->args->open_outarg.backing_id);
if (IS_ERR(fb))
@@ -197,7 +195,6 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
{
struct fuse_file *ff = file->private_data;
struct fuse_inode *fi = get_fuse_inode(inode);
- int err;
/*
* io modes are not relevant with virtiofs DAX and with server that does not
@@ -208,11 +205,12 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
/*
* Server is expected to use FOPEN_PASSTHROUGH for all opens of an inode
- * which is already open for passthrough.
+ * which is already open for passthrough. Using incorrect open mode is
+ * a server mistake, which results in user visible failure of open()
+ * with EIO error.
*/
- err = -EINVAL;
if (fuse_inode_backing(fi) && !(ff->open_flags & FOPEN_PASSTHROUGH))
- goto fail;
+ return fuse_EIO("FOPEN_PASSTHROUGH expected");
/*
* FOPEN_PARALLEL_DIRECT_WRITES requires FOPEN_DIRECT_IO.
@@ -233,23 +231,12 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
return 0;
if (ff->open_flags & FOPEN_PASSTHROUGH)
- err = fuse_file_passthrough_open(inode, file);
- else
- err = fuse_file_cached_io_open(inode, ff);
- if (err)
- goto fail;
+ return fuse_file_passthrough_open(inode, file);
- return 0;
+ if (!fuse_file_cached_io_open(inode, ff))
+ return fuse_EIO("conflicting passthrough open");
-fail:
- pr_debug("failed to open file in requested io mode (open_flags=0x%x, err=%i).\n",
- ff->open_flags, err);
- /*
- * The file open mode determines the inode io mode.
- * Using incorrect open mode is a server mistake, which results in
- * user visible failure of open() with EIO error.
- */
- return -EIO;
+ return 0;
}
/* No more pending io and no new io possible to inode via open/mmapped file */
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index b43d3e0f7081..489838462d1e 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -159,34 +159,29 @@ struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
struct fuse_conn *fc = ff->fm->fc;
struct fuse_backing *fb = NULL;
struct file *backing_file;
- int err;
- err = -EINVAL;
if (backing_id <= 0)
- goto out;
+ return ERR_PTR(fuse_EIO("invalid backing_id"));
- err = -ENOENT;
fb = fuse_backing_lookup(fc, backing_id);
if (!fb)
- goto out;
+ return ERR_PTR(fuse_EIO("backing not found"));
/* Allocate backing file per fuse file to store fuse path */
backing_file = backing_file_open(file, file->f_flags,
&fb->file->f_path, fb->cred);
- err = PTR_ERR(backing_file);
if (IS_ERR(backing_file)) {
fuse_backing_put(fb);
- goto out;
+ return ERR_PTR(fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file)));
}
- err = 0;
ff->passthrough = backing_file;
ff->cred = get_cred(fb->cred);
-out:
- pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p, err=%i\n", __func__,
- backing_id, fb, ff->passthrough, err);
- return err ? ERR_PTR(err) : fb;
+ pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p\n", __func__,
+ backing_id, fb, ff->passthrough);
+
+ return fb;
}
void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb)
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v3 3/9] fuse: add helpers for EIO return value with kernel message
2026-10-06 18:01 ` [PATCH v3 3/9] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
@ 2026-10-06 21:03 ` Amir Goldstein
0 siblings, 0 replies; 23+ messages in thread
From: Amir Goldstein @ 2026-10-06 21:03 UTC (permalink / raw)
To: Miklos Szeredi
Cc: fuse-devel, John Groves, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
On Tue, Oct 6, 2026 at 8:01 PM Miklos Szeredi <mszeredi@redhat.com> wrote:
>
> Conditions are triggered by a buggy fuse server will result in -EIO which
> is hard to interpret without any context.
>
> Add helpers that print a short message to the kernel log as well as
> returning the error value. This serves a dual purpose:
>
> - documents the error condition inside the code
>
> - allows the implementor of the fuse server to get the context for the
> error
>
> Some debug messages are removed in favor of this.
>
> This also changes the mmap return value from -ETXTBSY to -ENODEV if a
> passthrough open raced with the prior check. This makes both cases return
> -ENODEV if the file is in passthrough mode.
>
> Additionally change the return value of fuse_file_cached_io_open() from an
> error value to a bool (true on success) as the error value is translated
> anyway.
>
> Currently these use the pr_notice_once() variant. Possibly should be
> changed to a ratelimit of e.g. once per minute.
>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
Reviewed-by: Amir Goldstein <amir73il@gmail.com>
> ---
> fs/fuse/file.c | 6 ++--
> fs/fuse/fuse_i.h | 4 ++-
> fs/fuse/iomode.c | 75 ++++++++++++++++++-------------------------
> fs/fuse/passthrough.c | 19 ++++-------
> 4 files changed, 43 insertions(+), 61 deletions(-)
>
> diff --git a/fs/fuse/file.c b/fs/fuse/file.c
> index 6d707f2b3bff..bde07948d1f0 100644
> --- a/fs/fuse/file.c
> +++ b/fs/fuse/file.c
> @@ -2415,7 +2415,6 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma)
> struct fuse_file *ff = file->private_data;
> struct fuse_conn *fc = ff->fm->fc;
> struct inode *inode = file_inode(file);
> - int rc;
>
> /* DAX mmap is superior to direct_io mmap */
> if (FUSE_IS_VDAX(inode))
> @@ -2457,9 +2456,8 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma)
> * After first mmap, the inode stays in caching io mode until
> * the direct_io file release.
> */
> - rc = fuse_file_cached_io_open(inode, ff);
> - if (rc)
> - return rc;
> + if (!fuse_file_cached_io_open(inode, ff))
> + return -ENODEV;
> }
>
> if ((vma->vm_flags & VM_SHARED) && (vma->vm_flags & VM_MAYWRITE))
> diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> index 8546855386b5..d9ec88028d7b 100644
> --- a/fs/fuse/fuse_i.h
> +++ b/fs/fuse/fuse_i.h
> @@ -11,6 +11,8 @@
> # define pr_fmt(fmt) "fuse: " fmt
> #endif
>
> +#define fuse_EIO(fmt, ...) (pr_notice_once("%s: " fmt "\n", __func__, ##__VA_ARGS__), -EIO)
> +
> #include "args.h"
> #include <linux/fuse.h>
> #include <linux/fs.h>
> @@ -1251,7 +1253,7 @@ int fuse_fileattr_set(struct mnt_idmap *idmap,
> struct dentry *dentry, struct file_kattr *fa);
>
> /* iomode.c */
> -int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff);
> +bool fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff);
> int fuse_inode_uncached_io_start(struct fuse_inode *fi,
> struct fuse_backing *fb);
> void fuse_inode_uncached_io_end(struct fuse_inode *fi);
> diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
> index 79637c09e883..1b8b267d0c0d 100644
> --- a/fs/fuse/iomode.c
> +++ b/fs/fuse/iomode.c
> @@ -27,15 +27,15 @@ static inline bool fuse_is_io_cache_wait(struct fuse_inode *fi)
> * Blocks new parallel dio writes and waits for the in-progress parallel dio
> * writes to complete.
> */
> -int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
> +bool fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
> {
> struct fuse_inode *fi = get_fuse_inode(inode);
>
> /* There are no io modes if server does not implement open */
> if (!ff->args)
> - return 0;
> + return true;
>
> - spin_lock(&fi->lock);
> + guard(spinlock)(&fi->lock);
> /*
> * Setting the bit advises new direct-io writes to use an exclusive
> * lock - without it the wait below might be forever.
> @@ -53,8 +53,7 @@ int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
> */
> if (fuse_inode_backing(fi)) {
> clear_bit(FUSE_I_CACHE_IO_MODE, &fi->state);
> - spin_unlock(&fi->lock);
> - return -ETXTBSY;
> + return false;
> }
>
> WARN_ON(ff->iomode == IOM_UNCACHED);
> @@ -64,8 +63,7 @@ int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
> set_bit(FUSE_I_CACHE_IO_MODE, &fi->state);
> fi->iocachectr++;
> }
> - spin_unlock(&fi->lock);
> - return 0;
> + return true;
> }
>
> static void fuse_file_cached_io_release(struct fuse_file *ff,
> @@ -85,19 +83,16 @@ static void fuse_file_cached_io_release(struct fuse_file *ff,
> int fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
> {
> struct fuse_backing *oldfb;
> - int err = 0;
>
> - spin_lock(&fi->lock);
> + guard(spinlock)(&fi->lock);
> /* deny conflicting backing files on same fuse inode */
> oldfb = fuse_inode_backing(fi);
> - if (fb && oldfb && oldfb != fb) {
> - err = -EBUSY;
> - goto unlock;
> - }
> - if (fi->iocachectr > 0) {
> - err = -ETXTBSY;
> - goto unlock;
> - }
> + if (fb && oldfb && oldfb != fb)
> + return -EBUSY;
> +
> + if (fi->iocachectr > 0)
> + return -ETXTBSY;
> +
> fi->iocachectr--;
>
> /* fuse inode holds a single refcount of backing file */
> @@ -107,9 +102,7 @@ int fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
> } else {
> fuse_backing_put(fb);
> }
> -unlock:
> - spin_unlock(&fi->lock);
> - return err;
> + return 0;
> }
>
> /* Takes uncached_io inode mode reference to be dropped on file release */
> @@ -121,8 +114,11 @@ static int fuse_file_uncached_io_open(struct inode *inode,
> int err;
>
> err = fuse_inode_uncached_io_start(fi, fb);
> - if (err)
> - return err;
> + if (err) {
> + if (err == -EBUSY)
> + return fuse_EIO("mismatched backing");
> + return fuse_EIO("conflicting caching mode");
> + }
>
> WARN_ON(ff->iomode != IOM_NONE);
> ff->iomode = IOM_UNCACHED;
> @@ -173,9 +169,11 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
> int err;
>
> /* Check allowed conditions for file open in passthrough mode */
> - if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) || !fc->passthrough ||
> - (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK))
> - return -EINVAL;
> + if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) || !fc->passthrough)
> + return fuse_EIO("passthrough not enabled");
> +
> + if (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK)
> + return fuse_EIO("conflicting open flags");
>
> fb = fuse_passthrough_open(file, ff->args->open_outarg.backing_id);
> if (IS_ERR(fb))
> @@ -197,7 +195,6 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
> {
> struct fuse_file *ff = file->private_data;
> struct fuse_inode *fi = get_fuse_inode(inode);
> - int err;
>
> /*
> * io modes are not relevant with virtiofs DAX and with server that does not
> @@ -208,11 +205,12 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
>
> /*
> * Server is expected to use FOPEN_PASSTHROUGH for all opens of an inode
> - * which is already open for passthrough.
> + * which is already open for passthrough. Using incorrect open mode is
> + * a server mistake, which results in user visible failure of open()
> + * with EIO error.
> */
> - err = -EINVAL;
> if (fuse_inode_backing(fi) && !(ff->open_flags & FOPEN_PASSTHROUGH))
> - goto fail;
> + return fuse_EIO("FOPEN_PASSTHROUGH expected");
>
> /*
> * FOPEN_PARALLEL_DIRECT_WRITES requires FOPEN_DIRECT_IO.
> @@ -233,23 +231,12 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
> return 0;
>
> if (ff->open_flags & FOPEN_PASSTHROUGH)
> - err = fuse_file_passthrough_open(inode, file);
> - else
> - err = fuse_file_cached_io_open(inode, ff);
> - if (err)
> - goto fail;
> + return fuse_file_passthrough_open(inode, file);
>
> - return 0;
> + if (!fuse_file_cached_io_open(inode, ff))
> + return fuse_EIO("conflicting passthrough open");
>
> -fail:
> - pr_debug("failed to open file in requested io mode (open_flags=0x%x, err=%i).\n",
> - ff->open_flags, err);
> - /*
> - * The file open mode determines the inode io mode.
> - * Using incorrect open mode is a server mistake, which results in
> - * user visible failure of open() with EIO error.
> - */
> - return -EIO;
> + return 0;
> }
>
> /* No more pending io and no new io possible to inode via open/mmapped file */
> diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
> index b43d3e0f7081..489838462d1e 100644
> --- a/fs/fuse/passthrough.c
> +++ b/fs/fuse/passthrough.c
> @@ -159,34 +159,29 @@ struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
> struct fuse_conn *fc = ff->fm->fc;
> struct fuse_backing *fb = NULL;
> struct file *backing_file;
> - int err;
>
> - err = -EINVAL;
> if (backing_id <= 0)
> - goto out;
> + return ERR_PTR(fuse_EIO("invalid backing_id"));
>
> - err = -ENOENT;
> fb = fuse_backing_lookup(fc, backing_id);
> if (!fb)
> - goto out;
> + return ERR_PTR(fuse_EIO("backing not found"));
>
> /* Allocate backing file per fuse file to store fuse path */
> backing_file = backing_file_open(file, file->f_flags,
> &fb->file->f_path, fb->cred);
> - err = PTR_ERR(backing_file);
> if (IS_ERR(backing_file)) {
> fuse_backing_put(fb);
> - goto out;
> + return ERR_PTR(fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file)));
> }
>
> - err = 0;
> ff->passthrough = backing_file;
> ff->cred = get_cred(fb->cred);
> -out:
> - pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p, err=%i\n", __func__,
> - backing_id, fb, ff->passthrough, err);
>
> - return err ? ERR_PTR(err) : fb;
> + pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p\n", __func__,
> + backing_id, fb, ff->passthrough);
> +
> + return fb;
> }
>
> void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb)
> --
> 2.54.0
>
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
` (2 preceding siblings ...)
2026-10-06 18:01 ` [PATCH v3 3/9] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
2026-10-06 18:20 ` sashiko-bot
` (2 more replies)
2026-10-06 18:01 ` [PATCH v3 5/9] fuse: support opening 64 bit " Miklos Szeredi
` (4 subsequent siblings)
8 siblings, 3 replies; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
Add support for server allocated 64-bit backing IDs alongside the
existing kernel allocated 32-bit IDs.
Opened with FUSE_DEV_IOC_BACKING_CREATE, the backing ID sent via
fuse_backing_create_in.backing_id.
Closed with FUSE_NOTIFY_BACKING_REMOVE. Since close only provides the
backing ID, not the file descriptor, it doesn't have to be done with an
ioctl.
Backing ID can take any 64 bit value other than zero.
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
fs/fuse/backing.c | 179 ++++++++++++++++++++++++++++----------
fs/fuse/dev.c | 24 +++++
fs/fuse/dev.h | 3 +
fs/fuse/fuse_i.h | 24 ++++-
fs/fuse/inode.c | 15 +++-
fs/fuse/notify.c | 28 ++++++
fs/fuse/passthrough.c | 3 +
include/uapi/linux/fuse.h | 22 ++++-
8 files changed, 246 insertions(+), 52 deletions(-)
diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index 433fa3098d71..bc40818778df 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -9,6 +9,7 @@
#include "fuse_i.h"
#include <linux/file.h>
+#include <linux/rhashtable.h>
static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb)
{
@@ -48,6 +49,8 @@ static int fuse_backing_id_alloc(struct fuse_conn *fc, struct fuse_backing *fb)
id = idr_alloc_cyclic(&fc->backing_files_map, fb, 1, 0, GFP_ATOMIC);
spin_unlock(&fc->lock);
idr_preload_end();
+ if (id > 0)
+ fb->backing_id = id;
WARN_ON_ONCE(id == 0);
return id;
@@ -61,81 +64,133 @@ static struct fuse_backing *fuse_backing_id_remove(struct fuse_conn *fc,
spin_lock(&fc->lock);
fb = idr_remove(&fc->backing_files_map, id);
spin_unlock(&fc->lock);
+ if (fb)
+ fb->backing_id = 0;
return fb;
}
-static int fuse_backing_id_free(int id, void *p, void *data)
-{
- struct fuse_backing *fb = p;
+static const struct rhashtable_params fuse_backing_prm = {
+ .head_offset = offsetof(struct fuse_backing, hash_node),
+ .key_offset = offsetof(struct fuse_backing, backing_id),
+ .key_len = sizeof_field(struct fuse_backing, backing_id),
+};
- WARN_ON_ONCE(refcount_read(&fb->count) != 1);
- fuse_backing_free(fb);
- return 0;
+static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
+{
+ return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
}
-void fuse_backing_files_free(struct fuse_conn *fc)
+int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)
{
- idr_for_each(&fc->backing_files_map, fuse_backing_id_free, NULL);
- idr_destroy(&fc->backing_files_map);
+ struct fuse_backing *fb;
+ int err;
+
+ if (!fc->backing_id_64)
+ return -EINVAL;
+
+ scoped_guard(spinlock, &fc->lock) {
+ fb = rhashtable_lookup_fast(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
+ if (!fb)
+ return -ENOENT;
+
+ err = rhashtable_remove_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
+ WARN_ON(err);
+ }
+ fb->backing_id = 0;
+ fuse_backing_put(fb);
+
+ return 0;
}
-int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
+static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
{
- struct file *file;
+ struct fuse_backing *fb;
struct super_block *backing_sb;
- struct fuse_backing *fb = NULL;
- int res;
-
- pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
+ struct file *file;
/* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
- res = -EPERM;
if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
- goto out;
+ return ERR_PTR(-EPERM);
- res = -EINVAL;
- if (map->flags || map->padding)
- goto out;
+ CLASS(fd_raw, f)(fd);
+ if (fd_empty(f))
+ return ERR_PTR(-EBADF);
- file = fget_raw(map->fd);
- res = -EBADF;
- if (!file)
- goto out;
+ file = fd_file(f);
/* read/write/splice/mmap passthrough only relevant for regular files */
- res = d_is_dir(file->f_path.dentry) ? -EISDIR : -EINVAL;
if (!d_is_reg(file->f_path.dentry))
- goto out_fput;
+ return d_is_dir(file->f_path.dentry) ? ERR_PTR(-EISDIR) : ERR_PTR(-EINVAL);
backing_sb = file_inode(file)->i_sb;
- res = -ELOOP;
if (backing_sb->s_stack_depth >= fc->max_stack_depth)
- goto out_fput;
+ return ERR_PTR(-ELOOP);
fb = kmalloc_obj(struct fuse_backing);
- res = -ENOMEM;
if (!fb)
- goto out_fput;
+ return ERR_PTR(-ENOMEM);
- fb->file = file;
+ fb->file = get_file(file);
fb->cred = get_current_cred();
refcount_set(&fb->count, 1);
- res = fuse_backing_id_alloc(fc, fb);
- if (res < 0) {
+ return fb;
+}
+
+int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map)
+{
+ struct fuse_backing *fb;
+ int res;
+
+ if (map->padding || map->spare[0] || map->spare[1] || !map->backing_id)
+ return -EINVAL;
+
+ if (!fc->backing_id_64)
+ return -EINVAL;
+
+ fb = fuse_backing_new(fc, map->fd);
+ if (IS_ERR(fb))
+ return PTR_ERR(fb);
+
+ fb->backing_id = map->backing_id;
+ res = fuse_backing_add_64(fc, fb);
+ if (res < 0)
fuse_backing_free(fb);
- fb = NULL;
+
+ return res;
+}
+
+int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
+{
+ struct fuse_backing *fb = NULL;
+ int res;
+
+ pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
+
+ res = -EINVAL;
+ if (map->flags || map->padding)
+ goto out;
+
+ scoped_guard(spinlock, &fc->lock) {
+ if (fc->backing_id_64)
+ goto out;
+ fc->backing_id_32 = true;
}
+ fb = fuse_backing_new(fc, map->fd);
+ res = PTR_ERR(fb);
+ if (!IS_ERR(fb)) {
+ res = fuse_backing_id_alloc(fc, fb);
+ if (res < 0) {
+ fuse_backing_free(fb);
+ fb = NULL;
+ }
+ }
out:
pr_debug("%s: fb=0x%p, ret=%i\n", __func__, fb, res);
return res;
-
-out_fput:
- fput(file);
- goto out;
}
int fuse_backing_close(struct fuse_conn *fc, int backing_id)
@@ -145,6 +200,9 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
pr_debug("%s: backing_id=%d\n", __func__, backing_id);
+ if (fc->backing_id_64)
+ return -EINVAL;
+
/* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
err = -EPERM;
if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
@@ -167,14 +225,47 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
return err;
}
-struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id)
+struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id)
{
struct fuse_backing *fb;
- rcu_read_lock();
- fb = idr_find(&fc->backing_files_map, backing_id);
- fb = fuse_backing_get(fb);
- rcu_read_unlock();
+ guard(rcu)();
+ if (!fc->backing_id_64)
+ fb = idr_find(&fc->backing_files_map, backing_id);
+ else
+ fb = rhashtable_lookup(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
- return fb;
+ return fuse_backing_get(fb);
+}
+
+static void fuse_backing_check_free(struct fuse_backing *fb)
+{
+ WARN_ON_ONCE(refcount_read(&fb->count) != 1);
+ fuse_backing_free(fb);
+}
+
+static int fuse_backing_idr_free(int id, void *p, void *data)
+{
+ fuse_backing_check_free(p);
+ return 0;
+}
+
+static void fuse_backing_rht_free(void *p, void *data)
+{
+ fuse_backing_check_free(p);
+}
+
+void fuse_backing_files_free(struct fuse_conn *fc)
+{
+ if (fc->backing_id_64) {
+ rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
+ } else {
+ idr_for_each(&fc->backing_files_map, fuse_backing_idr_free, NULL);
+ idr_destroy(&fc->backing_files_map);
+ }
+}
+
+void fuse_backing_files_init_64(struct fuse_conn *fc)
+{
+ rhashtable_init(&fc->backing_64_ht, &fuse_backing_prm);
}
diff --git a/fs/fuse/dev.c b/fs/fuse/dev.c
index 2d7ee5498f1c..15876a534ad0 100644
--- a/fs/fuse/dev.c
+++ b/fs/fuse/dev.c
@@ -2339,6 +2339,27 @@ static long fuse_dev_ioctl_backing_open(struct file *file,
return fuse_backing_open(fud->chan->conn, &map);
}
+static long fuse_dev_ioctl_backing_create(struct file *file,
+ struct fuse_backing_create_in __user *argp)
+{
+ struct fuse_dev *fud = fuse_get_dev(file);
+ struct fuse_backing_create_in map;
+
+ if (IS_ERR(fud))
+ return PTR_ERR(fud);
+
+ if (!smp_load_acquire(&fud->chan->initialized))
+ return -ENOTCONN;
+
+ if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
+ return -EOPNOTSUPP;
+
+ if (copy_from_user(&map, argp, sizeof(map)))
+ return -EFAULT;
+
+ return fuse_backing_open_64(fud->chan->conn, &map);
+}
+
static long fuse_dev_ioctl_backing_close(struct file *file, __u32 __user *argp)
{
struct fuse_dev *fud = fuse_get_dev(file);
@@ -2379,6 +2400,9 @@ static long fuse_dev_ioctl(struct file *file, unsigned int cmd,
case FUSE_DEV_IOC_BACKING_OPEN:
return fuse_dev_ioctl_backing_open(file, argp);
+ case FUSE_DEV_IOC_BACKING_CREATE:
+ return fuse_dev_ioctl_backing_create(file, argp);
+
case FUSE_DEV_IOC_BACKING_CLOSE:
return fuse_dev_ioctl_backing_close(file, argp);
diff --git a/fs/fuse/dev.h b/fs/fuse/dev.h
index f6c47ae0395b..dbffd5bed2f5 100644
--- a/fs/fuse/dev.h
+++ b/fs/fuse/dev.h
@@ -14,6 +14,7 @@ struct fuse_dev;
struct fuse_args;
struct fuse_copy_state;
struct fuse_backing_map;
+struct fuse_backing_create_in;
struct file;
struct folio;
enum fuse_notify_code;
@@ -87,6 +88,8 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map);
int fuse_backing_close(struct fuse_conn *fc, int backing_id);
+int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map);
+int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id);
int fuse_copy_one(struct fuse_copy_state *cs, void *val, unsigned size);
int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop,
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index d9ec88028d7b..6def89667409 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -32,6 +32,7 @@
#include <linux/pid_namespace.h>
#include <linux/refcount.h>
#include <linux/user_namespace.h>
+#include <linux/rhashtable-types.h>
/** Default max number of pages that can be used in a single read request */
#define FUSE_DEFAULT_MAX_PAGES_PER_REQ 32
@@ -92,7 +93,8 @@ struct fuse_submount_lookup {
struct fuse_backing {
struct file *file;
const struct cred *cred;
-
+ u64 backing_id;
+ struct rhash_head hash_node;
/* refcount */
refcount_t count;
struct rcu_head rcu;
@@ -689,6 +691,12 @@ struct fuse_conn {
/** @init_security: Initialize security xattrs when creating a new inode */
unsigned int init_security:1;
+ /** Legacy 32 bit backing ID */
+ bool backing_id_32:1;
+
+ /** Backing ID is 64 bit and allocated by the server */
+ bool backing_id_64:1;
+
/**
* @create_supp_group: Add supplementary group info when creating
* a new inode
@@ -770,8 +778,14 @@ struct fuse_conn {
struct fuse_sync_bucket __rcu *curr_bucket;
#ifdef CONFIG_FUSE_PASSTHROUGH
- /** @backing_files_map: IDR for backing files ids */
- struct idr backing_files_map;
+ /* Selected by backing_id_64 */
+ union {
+ /** @backing_files_map: IDR for backing files ids */
+ struct idr backing_files_map;
+
+ /** @backing_64_ht: 64 bit ID lookup hash table */
+ struct rhashtable backing_64_ht;
+ };
#endif
};
@@ -1270,7 +1284,7 @@ void fuse_file_release(struct inode *inode, struct fuse_file *ff,
/* backing.c */
#ifdef CONFIG_FUSE_PASSTHROUGH
void fuse_backing_put(struct fuse_backing *fb);
-struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id);
+
#else
static inline void fuse_backing_put(struct fuse_backing *fb)
@@ -1278,7 +1292,9 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
}
#endif
+struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
void fuse_backing_files_init(struct fuse_conn *fc);
+void fuse_backing_files_init_64(struct fuse_conn *fc);
void fuse_backing_files_free(struct fuse_conn *fc);
/* passthrough.c */
diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
index cbb10e19e7e8..b5b51865d59f 100644
--- a/fs/fuse/inode.c
+++ b/fs/fuse/inode.c
@@ -1408,13 +1408,24 @@ static void process_init_reply(struct fuse_args *args, int error)
* them together.
*/
if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) &&
- (flags & FUSE_PASSTHROUGH) &&
+ (flags & (FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2)) &&
arg->max_stack_depth > 0 &&
arg->max_stack_depth <= FILESYSTEM_MAX_STACK_DEPTH &&
!(flags & FUSE_WRITEBACK_CACHE)) {
fc->passthrough = 1;
fc->max_stack_depth = arg->max_stack_depth;
fm->sb->s_stack_depth = arg->max_stack_depth;
+ if (flags & FUSE_PASSTHROUGH_V2) {
+ /* Prevent race with fuse_backing_open() */
+ scoped_guard(spinlock, &fc->lock) {
+ if (fc->backing_id_32)
+ ok = false;
+ else
+ fc->backing_id_64 = true;
+ }
+ if (ok)
+ fuse_backing_files_init_64(fc);
+ }
}
if (flags & FUSE_NO_EXPORT_SUPPORT)
fm->sb->s_export_op = &fuse_export_fid_operations;
@@ -1500,7 +1511,7 @@ static struct fuse_init_args *fuse_new_init(struct fuse_mount *fm)
if (fm->fc->auto_submounts)
flags |= FUSE_SUBMOUNTS;
if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
- flags |= FUSE_PASSTHROUGH;
+ flags |= FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2;
/* Only offered to sufficiently privileged servers; see
* fuse_syncfs_enable().
*/
diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
index c0428e03d138..93e916a16ac9 100644
--- a/fs/fuse/notify.c
+++ b/fs/fuse/notify.c
@@ -409,6 +409,31 @@ static int fuse_notify_prune(struct fuse_conn *fc, unsigned int size,
return 0;
}
+static int fuse_notify_backing_remove(struct fuse_conn *fc, unsigned int size,
+ struct fuse_copy_state *cs)
+{
+ struct fuse_notify_backing_remove_out outarg;
+ int err;
+
+ if (size != sizeof(outarg))
+ return -EINVAL;
+
+ err = fuse_copy_one(cs, &outarg, sizeof(outarg));
+ if (err)
+ return err;
+
+ if (outarg.reserved)
+ return -EINVAL;
+
+ if (!fc->backing_id_64)
+ return -EINVAL;
+
+ if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
+ return -EOPNOTSUPP;
+
+ return fuse_backing_close_64(fc, outarg.backing_id);
+}
+
int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
unsigned int size, struct fuse_copy_state *cs)
{
@@ -440,6 +465,9 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
case FUSE_NOTIFY_PRUNE:
return fuse_notify_prune(fc, size, cs);
+ case FUSE_NOTIFY_BACKING_REMOVE:
+ return fuse_notify_backing_remove(fc, size, cs);
+
default:
return -EINVAL;
}
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index 489838462d1e..313c8d7ffc09 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -160,6 +160,9 @@ struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
struct fuse_backing *fb = NULL;
struct file *backing_file;
+ if (fc->backing_id_64)
+ return ERR_PTR(fuse_EIO("incompatible backing version"));
+
if (backing_id <= 0)
return ERR_PTR(fuse_EIO("invalid backing_id"));
diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
index 10a7f31c4bdf..37e9559783c3 100644
--- a/include/uapi/linux/fuse.h
+++ b/include/uapi/linux/fuse.h
@@ -251,6 +251,9 @@
*
* 7.47
* - add FUSE_HAS_SYNCFS opt-in flag for privileged userspace servers
+ * - add FUSE_PASSTHROUGH_V2
+ * - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
+ * - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
*/
#ifndef _LINUX_FUSE_H
@@ -473,6 +476,7 @@ struct fuse_file_lock {
* with CAP_SYS_ADMIN in the initial user namespace (the same
* privilege that mounting virtiofs or fuseblk requires).
* Insufficiently privileged servers ignore it.
+ * FUSE_PASSTHROUGH_V2: use 64 bit server allocated backing ID
*/
#define FUSE_ASYNC_READ (1 << 0)
#define FUSE_POSIX_LOCKS (1 << 1)
@@ -522,6 +526,7 @@ struct fuse_file_lock {
#define FUSE_REQUEST_TIMEOUT (1ULL << 42)
#define FUSE_HAS_IO_URING_BUFPOOL (1ULL << 43)
#define FUSE_HAS_SYNCFS (1ULL << 44)
+#define FUSE_PASSTHROUGH_V2 (1ULL << 45)
/**
* CUSE INIT request/reply flags
@@ -709,6 +714,7 @@ enum fuse_notify_code {
FUSE_NOTIFY_RESEND = 7,
FUSE_NOTIFY_INC_EPOCH = 8,
FUSE_NOTIFY_PRUNE = 9,
+ FUSE_NOTIFY_BACKING_REMOVE = 10,
};
/* The read buffer is required to be at least 8k, but may be much larger */
@@ -1159,13 +1165,20 @@ struct fuse_backing_map {
uint64_t padding;
};
+struct fuse_backing_create_in {
+ int32_t fd;
+ uint32_t padding;
+ uint64_t backing_id; /* Zero value is reserved */
+ uint64_t spare[2];
+};
+
/* Device ioctls: */
#define FUSE_DEV_IOC_MAGIC 229
#define FUSE_DEV_IOC_CLONE _IOR(FUSE_DEV_IOC_MAGIC, 0, uint32_t)
-#define FUSE_DEV_IOC_BACKING_OPEN _IOW(FUSE_DEV_IOC_MAGIC, 1, \
- struct fuse_backing_map)
+#define FUSE_DEV_IOC_BACKING_OPEN _IOW(FUSE_DEV_IOC_MAGIC, 1, struct fuse_backing_map)
#define FUSE_DEV_IOC_BACKING_CLOSE _IOW(FUSE_DEV_IOC_MAGIC, 2, uint32_t)
#define FUSE_DEV_IOC_SYNC_INIT _IO(FUSE_DEV_IOC_MAGIC, 3)
+#define FUSE_DEV_IOC_BACKING_CREATE _IOW(FUSE_DEV_IOC_MAGIC, 4, struct fuse_backing_create_in)
struct fuse_lseek_in {
uint64_t fh;
@@ -1193,6 +1206,11 @@ struct fuse_copy_file_range_out {
uint64_t bytes_copied;
};
+struct fuse_notify_backing_remove_out {
+ uint64_t backing_id;
+ uint64_t reserved;
+};
+
#define FUSE_SETUPMAPPING_FLAG_WRITE (1ull << 0)
#define FUSE_SETUPMAPPING_FLAG_READ (1ull << 1)
struct fuse_setupmapping_in {
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID
2026-10-06 18:01 ` [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
@ 2026-10-06 18:20 ` sashiko-bot
2026-10-06 20:10 ` John Groves
2026-10-06 22:08 ` Amir Goldstein
2 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-10-06 18:20 UTC (permalink / raw)
To: Miklos Szeredi; +Cc: linux-cxl, nvdimm
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] fuse: Uninitialized rhashtable destruction on failed init
- [High] fuse: TOCTOU race between fuse_backing_close and process_init_reply
- [Medium] fuse: rhashtable_insert_fast allows duplicate backing_id
--
commit 90eba0277f715071875b550790d8bf829971f8b3
Author: Miklos Szeredi <mszeredi@redhat.com>
fuse: support 64 bit, server allocated backing ID
This commit adds support for server-allocated 64-bit backing IDs alongside
the existing kernel-allocated 32-bit IDs. It introduces a new ioctl to
create these backing IDs and handles their removal.
> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index 433fa3098d71f..bc40818778dff 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
[ ... ]
> +static const struct rhashtable_params fuse_backing_prm = {
> + .head_offset = offsetof(struct fuse_backing, hash_node),
> + .key_offset = offsetof(struct fuse_backing, backing_id),
> + .key_len = sizeof_field(struct fuse_backing, backing_id),
> +};
>
> +static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
> +{
> + return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
> }
[Severity: Medium]
Does this code allow duplicate backing_id values?
If a FUSE daemon provides an already-existing backing_id,
rhashtable_insert_fast() will insert the new object instead of rejecting
it with -EEXIST.
Should this use rhashtable_lookup_insert_fast() instead to safely enforce
uniqueness?
[ ... ]
> int fuse_backing_close(struct fuse_conn *fc, int backing_id)
> {
> struct fuse_backing *fb = NULL;
> int err;
>
> pr_debug("%s: backing_id=%d\n", __func__, backing_id);
>
> + if (fc->backing_id_64)
> + return -EINVAL;
> +
[Severity: High]
Can this TOCTOU race lead to memory corruption? Since fc->backing_id_64 is
read here outside of fc->lock, it can be observed as false concurrently with
process_init_reply() setting it to true:
CPU1 (fuse_dev_ioctl_backing_close):
if (fc->backing_id_64) // reads false
...
fuse_backing_id_remove()
spin_lock(&fc->lock);
idr_remove(&fc->backing_files_map, id);
CPU2 (process_init_reply):
scoped_guard(spinlock, &fc->lock) {
fc->backing_id_64 = true;
}
if (ok)
fuse_backing_files_init_64(fc); // calls rhashtable_init()
Because fuse_backing_close() proceeds, it calls idr_remove() on the union
concurrently with rhashtable_init() in CPU2, corrupting the shared memory.
> /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> err = -EPERM;
> if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
[ ... ]
> diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
> index cbb10e19e7e86..b5b51865d59f8 100644
> --- a/fs/fuse/inode.c
> +++ b/fs/fuse/inode.c
[ ... ]
> if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) &&
> - (flags & FUSE_PASSTHROUGH) &&
> + (flags & (FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2)) &&
> arg->max_stack_depth > 0 &&
> arg->max_stack_depth <= FILESYSTEM_MAX_STACK_DEPTH &&
> !(flags & FUSE_WRITEBACK_CACHE)) {
> fc->passthrough = 1;
> fc->max_stack_depth = arg->max_stack_depth;
> fm->sb->s_stack_depth = arg->max_stack_depth;
> + if (flags & FUSE_PASSTHROUGH_V2) {
> + /* Prevent race with fuse_backing_open() */
> + scoped_guard(spinlock, &fc->lock) {
> + if (fc->backing_id_32)
> + ok = false;
> + else
> + fc->backing_id_64 = true;
> + }
> + if (ok)
> + fuse_backing_files_init_64(fc);
> + }
[Severity: High]
Does this unconditionally set fc->backing_id_64 = true but skip initialization
if ok is false?
If an earlier validation check sets ok = false, this code still sets
fc->backing_id_64 = true but skips calling fuse_backing_files_init_64(fc).
When the FUSE connection is later torn down, fuse_backing_files_free() will
see fc->backing_id_64 == true and attempt to call rhashtable_free_and_destroy()
on an uninitialized rhashtable, causing a kernel crash.
> }
> if (flags & FUSE_NO_EXPORT_SUPPORT)
> fm->sb->s_export_op = &fuse_export_fid_operations;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006180115.1425232-1-mszeredi@redhat.com?part=4
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID
2026-10-06 18:01 ` [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
2026-10-06 18:20 ` sashiko-bot
@ 2026-10-06 20:10 ` John Groves
2026-10-07 12:49 ` Miklos Szeredi
2026-10-06 22:08 ` Amir Goldstein
2 siblings, 1 reply; 23+ messages in thread
From: John Groves @ 2026-10-06 20:10 UTC (permalink / raw)
To: Miklos Szeredi, fuse-devel
Cc: Amir Goldstein, Darrick J . Wong, Vishal Verma, Dave Jiang,
Alison Schofield, nvdimm, linux-cxl
On Tue, Oct 6, 2026, at 1:01 PM, Miklos Szeredi wrote:
> Add support for server allocated 64-bit backing IDs alongside the
> existing kernel allocated 32-bit IDs.
>
> Opened with FUSE_DEV_IOC_BACKING_CREATE, the backing ID sent via
> fuse_backing_create_in.backing_id.
>
> Closed with FUSE_NOTIFY_BACKING_REMOVE. Since close only provides the
> backing ID, not the file descriptor, it doesn't have to be done with an
> ioctl.
>
> Backing ID can take any 64 bit value other than zero.
>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
> ---
> fs/fuse/backing.c | 179 ++++++++++++++++++++++++++++----------
> fs/fuse/dev.c | 24 +++++
> fs/fuse/dev.h | 3 +
> fs/fuse/fuse_i.h | 24 ++++-
> fs/fuse/inode.c | 15 +++-
> fs/fuse/notify.c | 28 ++++++
> fs/fuse/passthrough.c | 3 +
> include/uapi/linux/fuse.h | 22 ++++-
> 8 files changed, 246 insertions(+), 52 deletions(-)
>
> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index 433fa3098d71..bc40818778df 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
> @@ -9,6 +9,7 @@
> #include "fuse_i.h"
>
> #include <linux/file.h>
> +#include <linux/rhashtable.h>
>
> static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb)
> {
> @@ -48,6 +49,8 @@ static int fuse_backing_id_alloc(struct fuse_conn *fc, struct fuse_backing *fb)
> id = idr_alloc_cyclic(&fc->backing_files_map, fb, 1, 0, GFP_ATOMIC);
> spin_unlock(&fc->lock);
> idr_preload_end();
> + if (id > 0)
> + fb->backing_id = id;
>
> WARN_ON_ONCE(id == 0);
> return id;
> @@ -61,81 +64,133 @@ static struct fuse_backing *fuse_backing_id_remove(struct fuse_conn *fc,
> spin_lock(&fc->lock);
> fb = idr_remove(&fc->backing_files_map, id);
> spin_unlock(&fc->lock);
> + if (fb)
> + fb->backing_id = 0;
>
> return fb;
> }
>
> -static int fuse_backing_id_free(int id, void *p, void *data)
> -{
> - struct fuse_backing *fb = p;
> +static const struct rhashtable_params fuse_backing_prm = {
> + .head_offset = offsetof(struct fuse_backing, hash_node),
> + .key_offset = offsetof(struct fuse_backing, backing_id),
> + .key_len = sizeof_field(struct fuse_backing, backing_id),
> +};
>
> - WARN_ON_ONCE(refcount_read(&fb->count) != 1);
> - fuse_backing_free(fb);
> - return 0;
> +static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
> +{
> + return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
> }
>
> -void fuse_backing_files_free(struct fuse_conn *fc)
> +int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)
> {
> - idr_for_each(&fc->backing_files_map, fuse_backing_id_free, NULL);
> - idr_destroy(&fc->backing_files_map);
> + struct fuse_backing *fb;
> + int err;
> +
> + if (!fc->backing_id_64)
> + return -EINVAL;
> +
> + scoped_guard(spinlock, &fc->lock) {
> + fb = rhashtable_lookup_fast(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
> + if (!fb)
> + return -ENOENT;
> +
> + err = rhashtable_remove_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
> + WARN_ON(err);
> + }
> + fb->backing_id = 0;
> + fuse_backing_put(fb);
> +
> + return 0;
> }
>
> -int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
> +static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
> {
> - struct file *file;
> + struct fuse_backing *fb;
> struct super_block *backing_sb;
> - struct fuse_backing *fb = NULL;
> - int res;
> -
> - pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
> + struct file *file;
>
> /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> - res = -EPERM;
> if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> - goto out;
> + return ERR_PTR(-EPERM);
>
> - res = -EINVAL;
> - if (map->flags || map->padding)
> - goto out;
> + CLASS(fd_raw, f)(fd);
> + if (fd_empty(f))
> + return ERR_PTR(-EBADF);
>
> - file = fget_raw(map->fd);
> - res = -EBADF;
> - if (!file)
> - goto out;
> + file = fd_file(f);
>
> /* read/write/splice/mmap passthrough only relevant for regular files */
> - res = d_is_dir(file->f_path.dentry) ? -EISDIR : -EINVAL;
> if (!d_is_reg(file->f_path.dentry))
> - goto out_fput;
> + return d_is_dir(file->f_path.dentry) ? ERR_PTR(-EISDIR) : ERR_PTR(-EINVAL);
>
> backing_sb = file_inode(file)->i_sb;
> - res = -ELOOP;
> if (backing_sb->s_stack_depth >= fc->max_stack_depth)
> - goto out_fput;
> + return ERR_PTR(-ELOOP);
>
> fb = kmalloc_obj(struct fuse_backing);
> - res = -ENOMEM;
> if (!fb)
> - goto out_fput;
> + return ERR_PTR(-ENOMEM);
>
> - fb->file = file;
> + fb->file = get_file(file);
> fb->cred = get_current_cred();
> refcount_set(&fb->count, 1);
>
> - res = fuse_backing_id_alloc(fc, fb);
> - if (res < 0) {
> + return fb;
> +}
> +
> +int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map)
> +{
> + struct fuse_backing *fb;
> + int res;
> +
> + if (map->padding || map->spare[0] || map->spare[1] || !map->backing_id)
> + return -EINVAL;
> +
> + if (!fc->backing_id_64)
> + return -EINVAL;
> +
> + fb = fuse_backing_new(fc, map->fd);
> + if (IS_ERR(fb))
> + return PTR_ERR(fb);
> +
> + fb->backing_id = map->backing_id;
> + res = fuse_backing_add_64(fc, fb);
> + if (res < 0)
> fuse_backing_free(fb);
> - fb = NULL;
> +
> + return res;
> +}
> +
> +int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
> +{
> + struct fuse_backing *fb = NULL;
> + int res;
> +
> + pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
> +
> + res = -EINVAL;
> + if (map->flags || map->padding)
> + goto out;
> +
> + scoped_guard(spinlock, &fc->lock) {
> + if (fc->backing_id_64)
> + goto out;
> + fc->backing_id_32 = true;
> }
>
> + fb = fuse_backing_new(fc, map->fd);
> + res = PTR_ERR(fb);
> + if (!IS_ERR(fb)) {
> + res = fuse_backing_id_alloc(fc, fb);
> + if (res < 0) {
> + fuse_backing_free(fb);
> + fb = NULL;
> + }
> + }
> out:
> pr_debug("%s: fb=0x%p, ret=%i\n", __func__, fb, res);
>
> return res;
> -
> -out_fput:
> - fput(file);
> - goto out;
> }
>
> int fuse_backing_close(struct fuse_conn *fc, int backing_id)
> @@ -145,6 +200,9 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
>
> pr_debug("%s: backing_id=%d\n", __func__, backing_id);
>
> + if (fc->backing_id_64)
> + return -EINVAL;
> +
> /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> err = -EPERM;
> if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> @@ -167,14 +225,47 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
> return err;
> }
>
> -struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id)
> +struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id)
> {
> struct fuse_backing *fb;
>
> - rcu_read_lock();
> - fb = idr_find(&fc->backing_files_map, backing_id);
> - fb = fuse_backing_get(fb);
> - rcu_read_unlock();
> + guard(rcu)();
> + if (!fc->backing_id_64)
> + fb = idr_find(&fc->backing_files_map, backing_id);
> + else
> + fb = rhashtable_lookup(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
>
> - return fb;
> + return fuse_backing_get(fb);
> +}
> +
> +static void fuse_backing_check_free(struct fuse_backing *fb)
> +{
> + WARN_ON_ONCE(refcount_read(&fb->count) != 1);
> + fuse_backing_free(fb);
> +}
> +
> +static int fuse_backing_idr_free(int id, void *p, void *data)
> +{
> + fuse_backing_check_free(p);
> + return 0;
> +}
> +
> +static void fuse_backing_rht_free(void *p, void *data)
> +{
> + fuse_backing_check_free(p);
> +}
> +
> +void fuse_backing_files_free(struct fuse_conn *fc)
> +{
> + if (fc->backing_id_64) {
> + rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
> + } else {
> + idr_for_each(&fc->backing_files_map, fuse_backing_idr_free, NULL);
> + idr_destroy(&fc->backing_files_map);
> + }
> +}
> +
> +void fuse_backing_files_init_64(struct fuse_conn *fc)
> +{
> + rhashtable_init(&fc->backing_64_ht, &fuse_backing_prm);
> }
> diff --git a/fs/fuse/dev.c b/fs/fuse/dev.c
> index 2d7ee5498f1c..15876a534ad0 100644
> --- a/fs/fuse/dev.c
> +++ b/fs/fuse/dev.c
> @@ -2339,6 +2339,27 @@ static long fuse_dev_ioctl_backing_open(struct file *file,
> return fuse_backing_open(fud->chan->conn, &map);
> }
>
> +static long fuse_dev_ioctl_backing_create(struct file *file,
> + struct fuse_backing_create_in __user *argp)
> +{
> + struct fuse_dev *fud = fuse_get_dev(file);
> + struct fuse_backing_create_in map;
> +
> + if (IS_ERR(fud))
> + return PTR_ERR(fud);
> +
> + if (!smp_load_acquire(&fud->chan->initialized))
> + return -ENOTCONN;
> +
> + if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
> + return -EOPNOTSUPP;
> +
> + if (copy_from_user(&map, argp, sizeof(map)))
> + return -EFAULT;
> +
> + return fuse_backing_open_64(fud->chan->conn, &map);
> +}
> +
> static long fuse_dev_ioctl_backing_close(struct file *file, __u32 __user *argp)
> {
> struct fuse_dev *fud = fuse_get_dev(file);
> @@ -2379,6 +2400,9 @@ static long fuse_dev_ioctl(struct file *file, unsigned int cmd,
> case FUSE_DEV_IOC_BACKING_OPEN:
> return fuse_dev_ioctl_backing_open(file, argp);
>
> + case FUSE_DEV_IOC_BACKING_CREATE:
> + return fuse_dev_ioctl_backing_create(file, argp);
> +
> case FUSE_DEV_IOC_BACKING_CLOSE:
> return fuse_dev_ioctl_backing_close(file, argp);
>
> diff --git a/fs/fuse/dev.h b/fs/fuse/dev.h
> index f6c47ae0395b..dbffd5bed2f5 100644
> --- a/fs/fuse/dev.h
> +++ b/fs/fuse/dev.h
> @@ -14,6 +14,7 @@ struct fuse_dev;
> struct fuse_args;
> struct fuse_copy_state;
> struct fuse_backing_map;
> +struct fuse_backing_create_in;
> struct file;
> struct folio;
> enum fuse_notify_code;
> @@ -87,6 +88,8 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
>
> int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map);
> int fuse_backing_close(struct fuse_conn *fc, int backing_id);
> +int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map);
> +int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id);
>
> int fuse_copy_one(struct fuse_copy_state *cs, void *val, unsigned size);
> int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop,
> diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> index d9ec88028d7b..6def89667409 100644
> --- a/fs/fuse/fuse_i.h
> +++ b/fs/fuse/fuse_i.h
> @@ -32,6 +32,7 @@
> #include <linux/pid_namespace.h>
> #include <linux/refcount.h>
> #include <linux/user_namespace.h>
> +#include <linux/rhashtable-types.h>
>
> /** Default max number of pages that can be used in a single read request */
> #define FUSE_DEFAULT_MAX_PAGES_PER_REQ 32
> @@ -92,7 +93,8 @@ struct fuse_submount_lookup {
> struct fuse_backing {
> struct file *file;
> const struct cred *cred;
> -
> + u64 backing_id;
> + struct rhash_head hash_node;
> /* refcount */
> refcount_t count;
> struct rcu_head rcu;
> @@ -689,6 +691,12 @@ struct fuse_conn {
> /** @init_security: Initialize security xattrs when creating a new inode */
> unsigned int init_security:1;
>
> + /** Legacy 32 bit backing ID */
> + bool backing_id_32:1;
> +
> + /** Backing ID is 64 bit and allocated by the server */
> + bool backing_id_64:1;
> +
> /**
> * @create_supp_group: Add supplementary group info when creating
> * a new inode
> @@ -770,8 +778,14 @@ struct fuse_conn {
> struct fuse_sync_bucket __rcu *curr_bucket;
>
> #ifdef CONFIG_FUSE_PASSTHROUGH
> - /** @backing_files_map: IDR for backing files ids */
> - struct idr backing_files_map;
> + /* Selected by backing_id_64 */
> + union {
> + /** @backing_files_map: IDR for backing files ids */
> + struct idr backing_files_map;
> +
> + /** @backing_64_ht: 64 bit ID lookup hash table */
> + struct rhashtable backing_64_ht;
> + };
> #endif
> };
>
> @@ -1270,7 +1284,7 @@ void fuse_file_release(struct inode *inode, struct fuse_file *ff,
> /* backing.c */
> #ifdef CONFIG_FUSE_PASSTHROUGH
> void fuse_backing_put(struct fuse_backing *fb);
> -struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id);
> +
> #else
>
> static inline void fuse_backing_put(struct fuse_backing *fb)
> @@ -1278,7 +1292,9 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
> }
> #endif
>
> +struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
> void fuse_backing_files_init(struct fuse_conn *fc);
> +void fuse_backing_files_init_64(struct fuse_conn *fc);
> void fuse_backing_files_free(struct fuse_conn *fc);
>
> /* passthrough.c */
> diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
> index cbb10e19e7e8..b5b51865d59f 100644
> --- a/fs/fuse/inode.c
> +++ b/fs/fuse/inode.c
> @@ -1408,13 +1408,24 @@ static void process_init_reply(struct fuse_args *args, int error)
> * them together.
> */
> if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) &&
> - (flags & FUSE_PASSTHROUGH) &&
> + (flags & (FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2)) &&
> arg->max_stack_depth > 0 &&
> arg->max_stack_depth <= FILESYSTEM_MAX_STACK_DEPTH &&
> !(flags & FUSE_WRITEBACK_CACHE)) {
> fc->passthrough = 1;
> fc->max_stack_depth = arg->max_stack_depth;
> fm->sb->s_stack_depth = arg->max_stack_depth;
> + if (flags & FUSE_PASSTHROUGH_V2) {
> + /* Prevent race with fuse_backing_open() */
> + scoped_guard(spinlock, &fc->lock) {
> + if (fc->backing_id_32)
> + ok = false;
> + else
> + fc->backing_id_64 = true;
> + }
> + if (ok)
> + fuse_backing_files_init_64(fc);
> + }
> }
> if (flags & FUSE_NO_EXPORT_SUPPORT)
> fm->sb->s_export_op = &fuse_export_fid_operations;
> @@ -1500,7 +1511,7 @@ static struct fuse_init_args *fuse_new_init(struct fuse_mount *fm)
> if (fm->fc->auto_submounts)
> flags |= FUSE_SUBMOUNTS;
> if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
> - flags |= FUSE_PASSTHROUGH;
> + flags |= FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2;
> /* Only offered to sufficiently privileged servers; see
> * fuse_syncfs_enable().
> */
> diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
> index c0428e03d138..93e916a16ac9 100644
> --- a/fs/fuse/notify.c
> +++ b/fs/fuse/notify.c
> @@ -409,6 +409,31 @@ static int fuse_notify_prune(struct fuse_conn *fc, unsigned int size,
> return 0;
> }
>
> +static int fuse_notify_backing_remove(struct fuse_conn *fc, unsigned int size,
> + struct fuse_copy_state *cs)
> +{
> + struct fuse_notify_backing_remove_out outarg;
> + int err;
> +
> + if (size != sizeof(outarg))
> + return -EINVAL;
> +
> + err = fuse_copy_one(cs, &outarg, sizeof(outarg));
> + if (err)
> + return err;
> +
> + if (outarg.reserved)
> + return -EINVAL;
> +
> + if (!fc->backing_id_64)
> + return -EINVAL;
> +
> + if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
> + return -EOPNOTSUPP;
> +
> + return fuse_backing_close_64(fc, outarg.backing_id);
> +}
> +
> int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
> unsigned int size, struct fuse_copy_state *cs)
> {
> @@ -440,6 +465,9 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
> case FUSE_NOTIFY_PRUNE:
> return fuse_notify_prune(fc, size, cs);
>
> + case FUSE_NOTIFY_BACKING_REMOVE:
> + return fuse_notify_backing_remove(fc, size, cs);
> +
> default:
> return -EINVAL;
> }
> diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
> index 489838462d1e..313c8d7ffc09 100644
> --- a/fs/fuse/passthrough.c
> +++ b/fs/fuse/passthrough.c
> @@ -160,6 +160,9 @@ struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
> struct fuse_backing *fb = NULL;
> struct file *backing_file;
>
> + if (fc->backing_id_64)
> + return ERR_PTR(fuse_EIO("incompatible backing version"));
> +
> if (backing_id <= 0)
> return ERR_PTR(fuse_EIO("invalid backing_id"));
>
> diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
> index 10a7f31c4bdf..37e9559783c3 100644
> --- a/include/uapi/linux/fuse.h
> +++ b/include/uapi/linux/fuse.h
> @@ -251,6 +251,9 @@
> *
> * 7.47
> * - add FUSE_HAS_SYNCFS opt-in flag for privileged userspace servers
> + * - add FUSE_PASSTHROUGH_V2
> + * - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
> + * - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
> */
>
> #ifndef _LINUX_FUSE_H
> @@ -473,6 +476,7 @@ struct fuse_file_lock {
> * with CAP_SYS_ADMIN in the initial user namespace (the same
> * privilege that mounting virtiofs or fuseblk requires).
> * Insufficiently privileged servers ignore it.
> + * FUSE_PASSTHROUGH_V2: use 64 bit server allocated backing ID
> */
> #define FUSE_ASYNC_READ (1 << 0)
> #define FUSE_POSIX_LOCKS (1 << 1)
> @@ -522,6 +526,7 @@ struct fuse_file_lock {
> #define FUSE_REQUEST_TIMEOUT (1ULL << 42)
> #define FUSE_HAS_IO_URING_BUFPOOL (1ULL << 43)
> #define FUSE_HAS_SYNCFS (1ULL << 44)
> +#define FUSE_PASSTHROUGH_V2 (1ULL << 45)
>
> /**
> * CUSE INIT request/reply flags
> @@ -709,6 +714,7 @@ enum fuse_notify_code {
> FUSE_NOTIFY_RESEND = 7,
> FUSE_NOTIFY_INC_EPOCH = 8,
> FUSE_NOTIFY_PRUNE = 9,
> + FUSE_NOTIFY_BACKING_REMOVE = 10,
> };
>
> /* The read buffer is required to be at least 8k, but may be much larger */
> @@ -1159,13 +1165,20 @@ struct fuse_backing_map {
> uint64_t padding;
> };
>
> +struct fuse_backing_create_in {
> + int32_t fd;
> + uint32_t padding;
> + uint64_t backing_id; /* Zero value is reserved */
> + uint64_t spare[2];
> +};
> +
> /* Device ioctls: */
> #define FUSE_DEV_IOC_MAGIC 229
> #define FUSE_DEV_IOC_CLONE _IOR(FUSE_DEV_IOC_MAGIC, 0, uint32_t)
> -#define FUSE_DEV_IOC_BACKING_OPEN _IOW(FUSE_DEV_IOC_MAGIC, 1, \
> - struct fuse_backing_map)
> +#define FUSE_DEV_IOC_BACKING_OPEN _IOW(FUSE_DEV_IOC_MAGIC, 1, struct fuse_backing_map)
> #define FUSE_DEV_IOC_BACKING_CLOSE _IOW(FUSE_DEV_IOC_MAGIC, 2, uint32_t)
> #define FUSE_DEV_IOC_SYNC_INIT _IO(FUSE_DEV_IOC_MAGIC, 3)
> +#define FUSE_DEV_IOC_BACKING_CREATE _IOW(FUSE_DEV_IOC_MAGIC, 4, struct fuse_backing_create_in)
>
> struct fuse_lseek_in {
> uint64_t fh;
> @@ -1193,6 +1206,11 @@ struct fuse_copy_file_range_out {
> uint64_t bytes_copied;
> };
>
> +struct fuse_notify_backing_remove_out {
> + uint64_t backing_id;
> + uint64_t reserved;
> +};
> +
> #define FUSE_SETUPMAPPING_FLAG_WRITE (1ull << 0)
> #define FUSE_SETUPMAPPING_FLAG_READ (1ull << 1)
> struct fuse_setupmapping_in {
> --
> 2.54.0
>
>
This doesn't build with CONFIG_FUSE=m. Maybe do something like this?
CLASS(fd_raw, ...) expands to fdget_raw(), which is not exported, so
fuse fails to link as a module:
ERROR: modpost: fs/fuse/fuse.ko: symbol 'fdget_raw' undefined!
Go back to fget_raw()/fput(), which the code used before the rewrite
and which is exported. Keeps O_PATH file descriptors working, unlike
switching to CLASS(fd, ...).
diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index ea6f5d7b2cf5..ec7e1ebb2635 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -183,16 +183,18 @@ static int fuse_backing_open_file(struct fuse_conn *fc, struct fuse_backing *fb,
static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
{
struct fuse_backing *fb __free(kfree) = kzalloc_obj(*fb);
+ struct file *file;
int err;
if (!fb)
return ERR_PTR(-ENOMEM);
- CLASS(fd_raw, f)(fd);
- if (fd_empty(f))
+ file = fget_raw(fd);
+ if (!file)
return ERR_PTR(-EBADF);
- err = fuse_backing_open_file(fc, fb, fd_file(f));
+ err = fuse_backing_open_file(fc, fb, file);
+ fput(file);
if (err)
return ERR_PTR(err);
-John
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID
2026-10-06 20:10 ` John Groves
@ 2026-10-07 12:49 ` Miklos Szeredi
0 siblings, 0 replies; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-07 12:49 UTC (permalink / raw)
To: John Groves
Cc: Miklos Szeredi, fuse-devel, Amir Goldstein, Darrick J . Wong,
Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl
On Tue, 6 Oct 2026 at 22:12, John Groves <john@groves.net> wrote:
> This doesn't build with CONFIG_FUSE=m. Maybe do something like this?
I'll just add EXPORT_SYMBOL(fdget_raw). There's no reason to not
export this if fget_raw() is exported.
Thanks,
Miklos
^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID
2026-10-06 18:01 ` [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
2026-10-06 18:20 ` sashiko-bot
2026-10-06 20:10 ` John Groves
@ 2026-10-06 22:08 ` Amir Goldstein
2026-10-07 12:54 ` Miklos Szeredi
2 siblings, 1 reply; 23+ messages in thread
From: Amir Goldstein @ 2026-10-06 22:08 UTC (permalink / raw)
To: Miklos Szeredi
Cc: fuse-devel, John Groves, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
On Tue, Oct 6, 2026 at 8:01 PM Miklos Szeredi <mszeredi@redhat.com> wrote:
>
> Add support for server allocated 64-bit backing IDs alongside the
> existing kernel allocated 32-bit IDs.
>
> Opened with FUSE_DEV_IOC_BACKING_CREATE, the backing ID sent via
> fuse_backing_create_in.backing_id.
>
> Closed with FUSE_NOTIFY_BACKING_REMOVE. Since close only provides the
> backing ID, not the file descriptor, it doesn't have to be done with an
> ioctl.
>
> Backing ID can take any 64 bit value other than zero.
>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
> ---
> fs/fuse/backing.c | 179 ++++++++++++++++++++++++++++----------
> fs/fuse/dev.c | 24 +++++
> fs/fuse/dev.h | 3 +
> fs/fuse/fuse_i.h | 24 ++++-
> fs/fuse/inode.c | 15 +++-
> fs/fuse/notify.c | 28 ++++++
> fs/fuse/passthrough.c | 3 +
> include/uapi/linux/fuse.h | 22 ++++-
[...]
> diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
> index cbb10e19e7e8..b5b51865d59f 100644
> --- a/fs/fuse/inode.c
> +++ b/fs/fuse/inode.c
> @@ -1408,13 +1408,24 @@ static void process_init_reply(struct fuse_args *args, int error)
> * them together.
> */
> if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) &&
> - (flags & FUSE_PASSTHROUGH) &&
> + (flags & (FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2)) &&
> arg->max_stack_depth > 0 &&
> arg->max_stack_depth <= FILESYSTEM_MAX_STACK_DEPTH &&
> !(flags & FUSE_WRITEBACK_CACHE)) {
> fc->passthrough = 1;
> fc->max_stack_depth = arg->max_stack_depth;
> fm->sb->s_stack_depth = arg->max_stack_depth;
> + if (flags & FUSE_PASSTHROUGH_V2) {
Since I was contemplating myself on whether I should also add an init flag
FUSE_PASSTHROUGH_INO, it occured to me that arg->max_stack_depth
has plenty of spare bits to host FUSE_PASSTHROUGH_* feature flags.
So if we rebrand it as arg->passthrough_params
#define FUSE_PASSTROUGH_PARAMS(max_backing_stack_depth, flags) \
((max_backing_stack_depth & 0xff) + 1) | ((flags & 0xff) << 8)
We could use a feature flag FUSE_PASSTROUGH_SERVER_ID_64
instead of wasting another init flag
It also avoid the unneeded questions/documentation/awkwardness
regarding FUSE_PASSTHROUGH_V2 with/out FUSE_PASSTHROUGH
But mostly, it leaves room for more extensions.
> + /* Prevent race with fuse_backing_open() */
> + scoped_guard(spinlock, &fc->lock) {
> + if (fc->backing_id_32)
> + ok = false;
> + else
> + fc->backing_id_64 = true;
> + }
> + if (ok)
> + fuse_backing_files_init_64(fc);
> + }
See my comments on backing_id_32 gating on the v2 patch.
Thanks,
Amir.
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID
2026-10-06 22:08 ` Amir Goldstein
@ 2026-10-07 12:54 ` Miklos Szeredi
0 siblings, 0 replies; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-07 12:54 UTC (permalink / raw)
To: Amir Goldstein
Cc: Miklos Szeredi, fuse-devel, John Groves, Darrick J . Wong,
Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl
On Wed, 7 Oct 2026 at 00:12, Amir Goldstein <amir73il@gmail.com> wrote:
> #define FUSE_PASSTROUGH_PARAMS(max_backing_stack_depth, flags) \
> ((max_backing_stack_depth & 0xff) + 1) | ((flags & 0xff) << 8)
I don't think that beast belongs in a public header. That alone is a
reason not to go this way.
> It also avoid the unneeded questions/documentation/awkwardness
> regarding FUSE_PASSTHROUGH_V2 with/out FUSE_PASSTHROUGH
"If FUSE_PASSTHROUGH_V2 is set, FUSE_PASSTHROUGH is ignored"
Thanks,
Miklos
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v3 5/9] fuse: support opening 64 bit backing ID
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
` (3 preceding siblings ...)
2026-10-06 18:01 ` [PATCH v3 4/9] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
2026-10-06 18:01 ` [PATCH v3 6/9] fuse: add support for opening dax device as backing Miklos Szeredi
` (3 subsequent siblings)
8 siblings, 0 replies; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
Add backing_id_64 field to fuse_open_out to allow opening files with
64-bit server-allocated backing IDs.
If the connection was initialized with FUSE_PASSTHROUGH_V2, the kernel
reads the backing ID from open_out.backing_id_64 instead of
open_out.backing_id.
The open reply is now variable-length for backward compatibility with
servers that don't send the new backing_id_64 field.
Reviewed-by: Amir Goldstein <amir73il@gmail.com>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
fs/fuse/dir.c | 3 ++-
fs/fuse/file.c | 6 +++++-
fs/fuse/fuse_i.h | 2 +-
fs/fuse/iomode.c | 35 +++++++++++++++++++++++++++++------
fs/fuse/passthrough.c | 28 ++++++----------------------
include/uapi/linux/fuse.h | 2 ++
6 files changed, 45 insertions(+), 31 deletions(-)
diff --git a/fs/fuse/dir.c b/fs/fuse/dir.c
index f48fafccce4b..874ec7cfffb1 100644
--- a/fs/fuse/dir.c
+++ b/fs/fuse/dir.c
@@ -876,6 +876,7 @@ static int fuse_create_open(struct mnt_idmap *idmap, struct inode *dir,
args.out_args[0].value = &outentry;
/* Store outarg for fuse_finish_open() */
outopenp = &ff->args->open_outarg;
+ args.out_argvar = true; /* compat */
args.out_args[1].size = sizeof(*outopenp);
args.out_args[1].value = outopenp;
@@ -885,7 +886,7 @@ static int fuse_create_open(struct mnt_idmap *idmap, struct inode *dir,
err = fuse_simple_idmap_request(idmap, fm, &args);
free_ext_value(&args);
- if (err)
+ if (err < 0)
goto out_free_ff;
err = -EIO;
diff --git a/fs/fuse/file.c b/fs/fuse/file.c
index bde07948d1f0..5273957f0399 100644
--- a/fs/fuse/file.c
+++ b/fs/fuse/file.c
@@ -28,6 +28,7 @@ static int fuse_send_open(struct fuse_mount *fm, u64 nodeid,
{
struct fuse_open_in inarg;
FUSE_ARGS(args);
+ int res;
memset(&inarg, 0, sizeof(inarg));
inarg.flags = open_flags & ~(O_CREAT | O_EXCL | O_NOCTTY);
@@ -45,10 +46,13 @@ static int fuse_send_open(struct fuse_mount *fm, u64 nodeid,
args.in_args[0].size = sizeof(inarg);
args.in_args[0].value = &inarg;
args.out_numargs = 1;
+ args.out_argvar = true; /* compat */
args.out_args[0].size = sizeof(*outargp);
args.out_args[0].value = outargp;
- return fuse_simple_request(fm, &args);
+ res = fuse_simple_request(fm, &args);
+
+ return res < 0 ? res : 0;
}
struct fuse_file *fuse_file_alloc(struct fuse_mount *fm, bool release)
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 6def89667409..f34607eda2c0 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -1317,7 +1317,7 @@ static inline struct fuse_backing *fuse_inode_backing_set(struct fuse_inode *fi,
#endif
}
-struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id);
+int fuse_passthrough_open(struct file *file, struct fuse_backing *fb);
void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb);
static inline bool fuse_is_passthrough(struct fuse_file *ff)
diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
index 1b8b267d0c0d..8b4774c80b11 100644
--- a/fs/fuse/iomode.c
+++ b/fs/fuse/iomode.c
@@ -165,7 +165,9 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
{
struct fuse_file *ff = file->private_data;
struct fuse_conn *fc = get_fuse_conn(inode);
+ struct fuse_open_out *outarg = &ff->args->open_outarg;
struct fuse_backing *fb;
+ u64 backing_id;
int err;
/* Check allowed conditions for file open in passthrough mode */
@@ -175,18 +177,39 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
if (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK)
return fuse_EIO("conflicting open flags");
- fb = fuse_passthrough_open(file, ff->args->open_outarg.backing_id);
- if (IS_ERR(fb))
- return PTR_ERR(fb);
+ if (!fc->backing_id_64) {
+ if (outarg->backing_id_64 != 0)
+ return fuse_EIO("64 bit backing ID set");
+
+ if (outarg->backing_id <= 0)
+ return fuse_EIO("invalid backing ID");
+
+ backing_id = outarg->backing_id;
+ } else {
+ if (outarg->backing_id != 0)
+ return fuse_EIO("32 bit backing ID set");
+
+ backing_id = outarg->backing_id_64;
+ }
+ fb = fuse_backing_lookup(fc, backing_id);
+ if (!fb)
+ return fuse_EIO("backing not found");
+
+ err = fuse_passthrough_open(file, fb);
+ if (err)
+ goto backing_put;
/* First passthrough file open denies caching inode io mode */
err = fuse_file_uncached_io_open(inode, ff, fb);
- if (!err)
- return 0;
+ if (err)
+ goto passthrough_release;
+ return 0;
+
+passthrough_release:
fuse_passthrough_release(ff, fb);
+backing_put:
fuse_backing_put(fb);
-
return err;
}
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index 313c8d7ffc09..4894842ad6d0 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -150,41 +150,25 @@ ssize_t fuse_passthrough_mmap(struct file *file, struct vm_area_struct *vma)
/*
* Setup passthrough to a backing file.
- *
- * Returns an fb object with elevated refcount to be stored in fuse inode.
*/
-struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
+int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
{
struct fuse_file *ff = file->private_data;
- struct fuse_conn *fc = ff->fm->fc;
- struct fuse_backing *fb = NULL;
struct file *backing_file;
- if (fc->backing_id_64)
- return ERR_PTR(fuse_EIO("incompatible backing version"));
-
- if (backing_id <= 0)
- return ERR_PTR(fuse_EIO("invalid backing_id"));
-
- fb = fuse_backing_lookup(fc, backing_id);
- if (!fb)
- return ERR_PTR(fuse_EIO("backing not found"));
-
/* Allocate backing file per fuse file to store fuse path */
backing_file = backing_file_open(file, file->f_flags,
&fb->file->f_path, fb->cred);
- if (IS_ERR(backing_file)) {
- fuse_backing_put(fb);
- return ERR_PTR(fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file)));
- }
+ if (IS_ERR(backing_file))
+ return fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file));
ff->passthrough = backing_file;
ff->cred = get_cred(fb->cred);
- pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p\n", __func__,
- backing_id, fb, ff->passthrough);
+ pr_debug("%s: backing_id=%llu, fb=0x%p, backing_file=0x%p\n", __func__,
+ fb->backing_id, fb, ff->passthrough);
- return fb;
+ return 0;
}
void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb)
diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
index 37e9559783c3..4b3904b3acb7 100644
--- a/include/uapi/linux/fuse.h
+++ b/include/uapi/linux/fuse.h
@@ -254,6 +254,7 @@
* - add FUSE_PASSTHROUGH_V2
* - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
* - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
+ * - add backing_id_64 to fuse_open_out
*/
#ifndef _LINUX_FUSE_H
@@ -841,6 +842,7 @@ struct fuse_open_out {
uint64_t fh;
uint32_t open_flags;
int32_t backing_id;
+ uint64_t backing_id_64;
};
struct fuse_release_in {
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* [PATCH v3 6/9] fuse: add support for opening dax device as backing
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
` (4 preceding siblings ...)
2026-10-06 18:01 ` [PATCH v3 5/9] fuse: support opening 64 bit " Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
2026-10-06 18:18 ` sashiko-bot
2026-10-06 18:01 ` [PATCH v3 7/9] fuse: add extent map data structure Miklos Szeredi
` (2 subsequent siblings)
8 siblings, 1 reply; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
This is only possible with FUSE_PASSTHROUGH_V2 enabled.
Mark the inode with S_DAX if FUSE_LOOKUP returns with FUSE_ATTR_DAX set.
This patch does not yet provide a way actually use the dax dev backing:
when such a backing ID is provided in reply to FUSE_OPEN with
FOPEN_PASSTHROUGH flag set, an error will be returned.
Originally-by: John Groves <john@groves.net>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
fs/fuse/backing.c | 105 +++++++++++++++++++++++++++++++++---------
fs/fuse/file.c | 2 +-
fs/fuse/fuse_i.h | 33 +++++++++++--
fs/fuse/inode.c | 11 ++++-
fs/fuse/iomode.c | 10 ++--
fs/fuse/passthrough.c | 6 ++-
6 files changed, 134 insertions(+), 33 deletions(-)
diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index bc40818778df..0ed850ddf8cd 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -9,6 +9,7 @@
#include "fuse_i.h"
#include <linux/file.h>
+#include <linux/dax.h>
#include <linux/rhashtable.h>
static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb)
@@ -22,9 +23,16 @@ static void fuse_backing_free(struct fuse_backing *fb)
{
pr_debug("%s: fb=0x%p\n", __func__, fb);
- if (fb->file)
- fput(fb->file);
- put_cred(fb->cred);
+ switch (fb->type) {
+ case FUSE_BACKING_PATH:
+ path_put(&fb->path);
+ put_cred(fb->cred);
+ break;
+
+ case FUSE_BACKING_DAXDEV:
+ fs_put_dax(fb->dax_dev, fb);
+ break;
+ }
kfree_rcu(fb, rcu);
}
@@ -103,39 +111,90 @@ int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)
return 0;
}
-static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
+static int fuse_dax_notify_failure(struct dax_device *daxdev, u64 offset, u64 len, int mf_flags)
{
struct fuse_backing *fb;
- struct super_block *backing_sb;
- struct file *file;
- /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
- if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
- return ERR_PTR(-EPERM);
+ guard(rcu)();
- CLASS(fd_raw, f)(fd);
- if (fd_empty(f))
- return ERR_PTR(-EBADF);
+ fb = dax_holder(daxdev);
+ if (fb)
+ fb->dax_error = true;
+
+ return 0;
+}
+
+static const struct dax_holder_operations fuse_dax_holder_ops = {
+ .notify_failure = fuse_dax_notify_failure,
+};
+
+static int fuse_backing_open_file(struct fuse_conn *fc, struct fuse_backing *fb, struct file *file)
+{
+ struct inode *inode = file_inode(file);
+ struct dax_device *daxdev;
+ int err;
+
+ switch (inode->i_mode & S_IFMT) {
+ case S_IFREG:
+ /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
+ if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
+ return -EPERM;
+
+ if (inode->i_sb->s_stack_depth >= fc->max_stack_depth)
+ return -ELOOP;
+
+ fb->type = FUSE_BACKING_PATH;
+ fb->path = file->f_path;
+ path_get(&fb->path);
+ fb->cred = get_current_cred();
+ return 0;
+
+ case S_IFCHR:
+ if (!fc->passthrough || !fc->backing_id_64)
+ return -EINVAL;
+
+ daxdev = dax_dev_find(inode->i_rdev);
+ if (!daxdev)
+ return -EINVAL;
+
+ err = -EPERM;
+ if (capable(CAP_SYS_RAWIO)) {
+ err = fs_dax_get(daxdev, fb, &fuse_dax_holder_ops);
+ if (!err) {
+ fb->type = FUSE_BACKING_DAXDEV;
+ fb->dax_dev = daxdev;
+ }
+ }
+ put_dax(daxdev);
+ return err;
- file = fd_file(f);
+ case S_IFDIR:
+ return -EISDIR;
- /* read/write/splice/mmap passthrough only relevant for regular files */
- if (!d_is_reg(file->f_path.dentry))
- return d_is_dir(file->f_path.dentry) ? ERR_PTR(-EISDIR) : ERR_PTR(-EINVAL);
+ default:
+ return -EINVAL;
+ }
+}
- backing_sb = file_inode(file)->i_sb;
- if (backing_sb->s_stack_depth >= fc->max_stack_depth)
- return ERR_PTR(-ELOOP);
+static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
+{
+ struct fuse_backing *fb __free(kfree) = kzalloc_obj(*fb);
+ int err;
- fb = kmalloc_obj(struct fuse_backing);
if (!fb)
return ERR_PTR(-ENOMEM);
- fb->file = get_file(file);
- fb->cred = get_current_cred();
+ CLASS(fd_raw, f)(fd);
+ if (fd_empty(f))
+ return ERR_PTR(-EBADF);
+
+ err = fuse_backing_open_file(fc, fb, fd_file(f));
+ if (err)
+ return ERR_PTR(err);
+
refcount_set(&fb->count, 1);
- return fb;
+ return_ptr(fb);
}
int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map)
diff --git a/fs/fuse/file.c b/fs/fuse/file.c
index 5273957f0399..e0d72b9d3c4a 100644
--- a/fs/fuse/file.c
+++ b/fs/fuse/file.c
@@ -297,7 +297,7 @@ static int fuse_open(struct inode *inode, struct file *file)
if (!err) {
if (is_truncate)
truncate_pagecache(inode, 0);
- else if (!(ff->open_flags & FOPEN_KEEP_CACHE))
+ else if (!(ff->open_flags & FOPEN_KEEP_CACHE) && !IS_DAX(inode))
invalidate_inode_pages2(inode->i_mapping);
}
out_unlock:
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index f34607eda2c0..7acf860865de 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -89,10 +89,25 @@ struct fuse_submount_lookup {
struct fuse_forget_link *forget;
};
+enum fuse_backing_type {
+ FUSE_BACKING_PATH,
+ FUSE_BACKING_DAXDEV,
+};
+
/* Container for data related to mapping to backing file */
struct fuse_backing {
- struct file *file;
- const struct cred *cred;
+ enum fuse_backing_type type;
+
+ union {
+ struct {
+ struct path path;
+ const struct cred *cred;
+ };
+ struct {
+ struct dax_device *dax_dev;
+ bool dax_error;
+ };
+ };
u64 backing_id;
struct rhash_head hash_node;
/* refcount */
@@ -1242,7 +1257,19 @@ void fuse_free_conn(struct fuse_conn *fc);
/* dax.c */
-#define FUSE_IS_VDAX(inode) (IS_ENABLED(CONFIG_FUSE_VDAX) && IS_DAX(inode))
+static inline bool fuse_inode_vdax(struct inode *inode)
+{
+#ifdef CONFIG_FUSE_VDAX
+ return get_fuse_inode(inode)->vdax;
+#else
+ return false;
+#endif
+}
+
+static inline bool FUSE_IS_VDAX(struct inode *inode)
+{
+ return fuse_inode_vdax(inode) && IS_DAX(inode);
+}
ssize_t fuse_vdax_read_iter(struct kiocb *iocb, struct iov_iter *to);
ssize_t fuse_vdax_write_iter(struct kiocb *iocb, struct iov_iter *from);
diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
index b5b51865d59f..750e092c3971 100644
--- a/fs/fuse/inode.c
+++ b/fs/fuse/inode.c
@@ -148,7 +148,7 @@ static void fuse_evict_inode(struct inode *inode)
/* Will write inode on close/munmap and in all other dirtiers */
WARN_ON(inode_state_read_once(inode) & I_DIRTY_INODE);
- if (FUSE_IS_VDAX(inode))
+ if (IS_DAX(inode))
dax_break_layout_final(inode);
truncate_inode_pages_final(&inode->i_data);
@@ -403,6 +403,10 @@ static void fuse_init_submount_lookup(struct fuse_submount_lookup *sl,
refcount_set(&sl->count, 1);
}
+static const struct address_space_operations fuse_dax_aops = {
+ .dirty_folio = noop_dirty_folio,
+};
+
static void fuse_init_inode(struct inode *inode, struct fuse_attr *attr,
struct fuse_conn *fc)
{
@@ -430,6 +434,11 @@ static void fuse_init_inode(struct inode *inode, struct fuse_attr *attr,
*/
if (!fc->posix_acl)
inode->i_acl = inode->i_default_acl = ACL_DONT_CACHE;
+
+ if ((attr->flags & FUSE_ATTR_DAX) && !fuse_inode_vdax(inode)) {
+ inode->i_flags |= S_DAX;
+ inode->i_data.a_ops = &fuse_dax_aops;
+ }
}
static int fuse_inode_eq(struct inode *inode, void *_nodeidp)
diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
index 8b4774c80b11..38afe1f238ef 100644
--- a/fs/fuse/iomode.c
+++ b/fs/fuse/iomode.c
@@ -230,10 +230,14 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
* Server is expected to use FOPEN_PASSTHROUGH for all opens of an inode
* which is already open for passthrough. Using incorrect open mode is
* a server mistake, which results in user visible failure of open()
- * with EIO error.
+ * with EIO error. Same with DAX inodes.
*/
- if (fuse_inode_backing(fi) && !(ff->open_flags & FOPEN_PASSTHROUGH))
- return fuse_EIO("FOPEN_PASSTHROUGH expected");
+ if (!(ff->open_flags & FOPEN_PASSTHROUGH)) {
+ if (fuse_inode_backing(fi))
+ return fuse_EIO("FOPEN_PASSTHROUGH expected");
+ if (IS_DAX(inode))
+ return fuse_EIO("DAX inode without FOPEN_PASSTHROUGH");
+ }
/*
* FOPEN_PARALLEL_DIRECT_WRITES requires FOPEN_DIRECT_IO.
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index 4894842ad6d0..e9ab1aea34e2 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -156,9 +156,11 @@ int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
struct fuse_file *ff = file->private_data;
struct file *backing_file;
+ if (fb->type != FUSE_BACKING_PATH)
+ return fuse_EIO("invalid backing type");
+
/* Allocate backing file per fuse file to store fuse path */
- backing_file = backing_file_open(file, file->f_flags,
- &fb->file->f_path, fb->cred);
+ backing_file = backing_file_open(file, file->f_flags, &fb->path, fb->cred);
if (IS_ERR(backing_file))
return fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file));
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v3 6/9] fuse: add support for opening dax device as backing
2026-10-06 18:01 ` [PATCH v3 6/9] fuse: add support for opening dax device as backing Miklos Szeredi
@ 2026-10-06 18:18 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-10-06 18:18 UTC (permalink / raw)
To: Miklos Szeredi; +Cc: linux-cxl, nvdimm
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Critical] fuse: FUSE_ATTR_DAX applied to non-regular files causes symlink DoS
- [High] fuse: sleep-in-atomic bug in fuse_backing_free() during RCU inode teardown
--
commit 122a624f241dd19a7e4f459c49f52f0a5d738881
Author: Miklos Szeredi <mszeredi@redhat.com>
Subject: fuse: add support for opening dax device as backing
This commit adds support for using a DAX device as a backing file in FUSE
passthrough mode. It introduces a new DAX backing type and marks the inode
with S_DAX if FUSE_ATTR_DAX is set during lookup.
> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index bc40818778dff..0ed850ddf8cde 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
[ ... ]
> @@ -22,9 +23,16 @@ static void fuse_backing_free(struct fuse_backing *fb)
> {
> pr_debug("%s: fb=0x%p\n", __func__, fb);
>
> - if (fb->file)
> - fput(fb->file);
> - put_cred(fb->cred);
> + switch (fb->type) {
> + case FUSE_BACKING_PATH:
> + path_put(&fb->path);
> + put_cred(fb->cred);
> + break;
> +
> + case FUSE_BACKING_DAXDEV:
> + fs_put_dax(fb->dax_dev, fb);
> + break;
> + }
[Severity: High]
Can path_put() or fs_put_dax() be called from atomic context here?
When the final reference is dropped during inode teardown in
fuse_free_inode(), it executes from the VFS RCU softirq callback
(i_callback).
If fuse_backing_put() is called from that RCU softirq context and drops
the final reference, invoking fuse_backing_free() synchronously, could
the sleeping functions path_put() and fs_put_dax() (which calls iput())
trigger a sleep-in-atomic bug?
> kfree_rcu(fb, rcu);
> }
>
[ ... ]
> diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
> index b5b51865d59f8..750e092c3971d 100644
> --- a/fs/fuse/inode.c
> +++ b/fs/fuse/inode.c
[ ... ]
> @@ -430,6 +434,11 @@ static void fuse_init_inode(struct inode *inode, struct fuse_attr *attr,
> */
> if (!fc->posix_acl)
> inode->i_acl = inode->i_default_acl = ACL_DONT_CACHE;
> +
> + if ((attr->flags & FUSE_ATTR_DAX) && !fuse_inode_vdax(inode)) {
> + inode->i_flags |= S_DAX;
> + inode->i_data.a_ops = &fuse_dax_aops;
> + }
[Severity: Critical]
Does this overwrite the address space operations for non-regular files
like symlinks?
If a FUSE server returns FUSE_ATTR_DAX on a symlink lookup, this code
unconditionally sets inode->i_data.a_ops = &fuse_dax_aops.
Since fuse_dax_aops does not implement a .read_folio callback, won't
this cause a NULL pointer dereference when a readlink() syscall attempts
to read the symlink (e.g., via fuse_get_link() -> page_get_link_raw() ->
read_mapping_folio())?
Should there be a check for S_ISREG(inode->i_mode) before setting the
DAX address space operations?
> }
>
> static int fuse_inode_eq(struct inode *inode, void *_nodeidp)
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006180115.1425232-1-mszeredi@redhat.com?part=6
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v3 7/9] fuse: add extent map data structure
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
` (5 preceding siblings ...)
2026-10-06 18:01 ` [PATCH v3 6/9] fuse: add support for opening dax device as backing Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
2026-10-06 18:17 ` sashiko-bot
2026-10-06 18:01 ` [PATCH v3 8/9] fuse: add extent map I/O support Miklos Szeredi
2026-10-06 18:01 ` [PATCH v3 9/9] fuse: add support for striped backing Miklos Szeredi
8 siblings, 1 reply; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
Add support for creating and managing extent maps that map file regions to
dax device regions.
Introduce FUSE_NOTIFY_BACKING_MAP for populating extent maps from
userspace.
Extent maps are handled as a new type of backing and are assigned a 64 bit
backing ID by the server.
One extent consists of
- offset within the containing backing
- length of extent
- target backing ID
- offset into target backing
Extents must be non-overlapping, and can only refer to dax device backings
for now.
At this point only support creating and populating the extent map in one
operation, though the interface is not limited by this and later may be
changed, so that partial population of an extent map would be possible.
Originally-by: John Groves <john@groves.net>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
fs/fuse/Makefile | 2 +-
fs/fuse/backing.c | 38 +++++++++--
fs/fuse/ext_map.c | 131 ++++++++++++++++++++++++++++++++++++++
fs/fuse/fuse_i.h | 14 +++-
fs/fuse/notify.c | 44 +++++++++++++
include/uapi/linux/fuse.h | 26 ++++++++
6 files changed, 248 insertions(+), 7 deletions(-)
create mode 100644 fs/fuse/ext_map.c
diff --git a/fs/fuse/Makefile b/fs/fuse/Makefile
index 5858feafa916..da9e802d7ae1 100644
--- a/fs/fuse/Makefile
+++ b/fs/fuse/Makefile
@@ -15,7 +15,7 @@ fuse-y += dev.o dir.o file.o inode.o control.o xattr.o acl.o readdir.o ioctl.o r
fuse-y += poll.o notify.o
fuse-y += iomode.o
fuse-$(CONFIG_FUSE_VDAX) += dax.o
-fuse-$(CONFIG_FUSE_PASSTHROUGH) += passthrough.o backing.o
+fuse-$(CONFIG_FUSE_PASSTHROUGH) += passthrough.o backing.o ext_map.o
fuse-$(CONFIG_SYSCTL) += sysctl.o
fuse-$(CONFIG_FUSE_IO_URING) += dev_uring.o
diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index 0ed850ddf8cd..b7ebc951dc31 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -32,6 +32,10 @@ static void fuse_backing_free(struct fuse_backing *fb)
case FUSE_BACKING_DAXDEV:
fs_put_dax(fb->dax_dev, fb);
break;
+
+ case FUSE_BACKING_EXTMAP:
+ fuse_ext_map_destroy(&fb->extents);
+ break;
}
kfree_rcu(fb, rcu);
}
@@ -84,7 +88,7 @@ static const struct rhashtable_params fuse_backing_prm = {
.key_len = sizeof_field(struct fuse_backing, backing_id),
};
-static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
+int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
{
return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
}
@@ -316,12 +320,38 @@ static void fuse_backing_rht_free(void *p, void *data)
void fuse_backing_files_free(struct fuse_conn *fc)
{
- if (fc->backing_id_64) {
- rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
- } else {
+ struct rhashtable_iter iter;
+ struct fuse_backing *fb;
+
+ if (!fc->backing_id_64) {
idr_for_each(&fc->backing_files_map, fuse_backing_idr_free, NULL);
idr_destroy(&fc->backing_files_map);
+ return;
}
+
+ /*
+ * extents are referencing other backings, put these refs before
+ * destroying the backings themselves
+ */
+ rhashtable_walk_enter(&fc->backing_64_ht, &iter);
+ rhashtable_walk_start(&iter);
+ while ((fb = rhashtable_walk_next(&iter))) {
+ if (IS_ERR(fb)) {
+ if (PTR_ERR(fb) == -EAGAIN)
+ continue;
+ break;
+ }
+ if (fb->type == FUSE_BACKING_EXTMAP && !RB_EMPTY_ROOT(&fb->extents)) {
+ rhashtable_walk_stop(&iter);
+ fuse_ext_map_destroy(&fb->extents);
+ fb->extents.rb_node = NULL;
+ rhashtable_walk_start(&iter);
+ }
+ }
+ rhashtable_walk_stop(&iter);
+ rhashtable_walk_exit(&iter);
+
+ rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
}
void fuse_backing_files_init_64(struct fuse_conn *fc)
diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
new file mode 100644
index 000000000000..39d7bb521873
--- /dev/null
+++ b/fs/fuse/ext_map.c
@@ -0,0 +1,131 @@
+// SPDX-License-Identifier: GPL-2.0-only
+
+#include "fuse_i.h"
+#include <linux/rbtree.h>
+#include <linux/pagemap.h>
+
+struct fuse_iext {
+ struct rb_node rb;
+ u64 start;
+ u64 end; /* exclusive */
+ struct fuse_backing *backing;
+ u64 backing_offset;
+};
+
+void fuse_ext_map_destroy(struct rb_root *extents)
+{
+ struct fuse_iext *fie, *tmp;
+
+ rbtree_postorder_for_each_entry_safe(fie, tmp, extents, rb) {
+ fuse_backing_put(fie->backing);
+ kfree(fie);
+ }
+}
+
+static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
+ struct fuse_extent *ext)
+{
+ struct fuse_iext *new_fie __free(kfree) = kzalloc_obj(*new_fie);
+ struct rb_node *parent = NULL, **link = &extents->rb_node;
+ loff_t end;
+ u64 tmp;
+
+ if (!new_fie)
+ return -ENOMEM;
+
+ if (ext->reserved[0] || ext->reserved[1])
+ return fuse_EIO("reserved fields set");
+
+ if (!PAGE_ALIGNED(ext->offset) || !PAGE_ALIGNED(ext->length) || !PAGE_ALIGNED(ext->addr))
+ return fuse_EIO("not page aligned");
+
+ if (!ext->length)
+ return fuse_EIO("zero sized extent");
+
+ if (overflows_type(ext->offset, loff_t) ||
+ check_add_overflow(ext->offset, ext->length, &end) ||
+ check_add_overflow(ext->addr, ext->length, &tmp))
+ return fuse_EIO("offset overflow");
+
+ new_fie->start = ext->offset;
+ new_fie->end = end;
+ new_fie->backing_offset = ext->addr;
+
+ while (*link) {
+ struct fuse_iext *fie = rb_entry(*link, typeof(*fie), rb);
+
+ parent = *link;
+ if (new_fie->end <= fie->start)
+ link = &parent->rb_left;
+ else if (new_fie->start >= fie->end)
+ link = &parent->rb_right;
+ else
+ return fuse_EIO("overlap");
+ }
+
+ new_fie->backing = fuse_backing_lookup(fc, ext->backing_id);
+ if (!new_fie->backing)
+ return fuse_EIO("backing not found");
+
+ if (new_fie->backing->type != FUSE_BACKING_DAXDEV) {
+ fuse_backing_put(new_fie->backing);
+ return fuse_EIO("backing is not dax device");
+ }
+
+ rb_link_node(&new_fie->rb, parent, link);
+ rb_insert_color(&no_free_ptr(new_fie)->rb, extents);
+
+ return 0;
+}
+
+bool fuse_ext_map_is_dax(struct fuse_backing *fb)
+{
+ struct fuse_iext *fie;
+
+ if (WARN_ON(RB_EMPTY_ROOT(&fb->extents)))
+ return false;
+
+ fie = rb_entry(fb->extents.rb_node, typeof(*fie), rb);
+ return fie->backing->type == FUSE_BACKING_DAXDEV;
+}
+
+int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_out *arg,
+ struct fuse_extent *ext)
+{
+ struct fuse_backing *fb;
+ unsigned int i;
+ int err;
+
+ if (!arg->num_extents)
+ return fuse_EIO("no extents");
+
+ if (arg->flags & FUSE_BACKING_MAP_CREATE) {
+ fb = kzalloc_obj(*fb);
+ if (!fb)
+ return -ENOMEM;
+
+ refcount_set(&fb->count, 1);
+ fb->type = FUSE_BACKING_EXTMAP;
+ fb->backing_id = arg->backing_id;
+ } else {
+ /* Adding extents to an existing extmap backing is not yet supported */
+ return -EINVAL;
+ }
+
+ for (i = 0; i < arg->num_extents; i++) {
+ err = fuse_add_extent(fc, &fb->extents, &ext[i]);
+ if (err)
+ goto err_put;
+ }
+
+ err = 0;
+ if (arg->flags & FUSE_BACKING_MAP_CREATE) {
+ err = fuse_backing_add_64(fc, fb);
+ if (!err)
+ return 0;
+ }
+
+err_put:
+ fuse_backing_put(fb);
+ return err;
+}
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 7acf860865de..75eec5c17985 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -24,7 +24,7 @@
#include <linux/backing-dev.h>
#include <linux/mutex.h>
#include <linux/rwsem.h>
-#include <linux/rbtree.h>
+#include <linux/rbtree_types.h>
#include <linux/poll.h>
#include <linux/workqueue.h>
#include <linux/kref.h>
@@ -92,6 +92,7 @@ struct fuse_submount_lookup {
enum fuse_backing_type {
FUSE_BACKING_PATH,
FUSE_BACKING_DAXDEV,
+ FUSE_BACKING_EXTMAP,
};
/* Container for data related to mapping to backing file */
@@ -107,6 +108,9 @@ struct fuse_backing {
struct dax_device *dax_dev;
bool dax_error;
};
+ struct {
+ struct rb_root extents;
+ };
};
u64 backing_id;
struct rhash_head hash_node;
@@ -1311,7 +1315,6 @@ void fuse_file_release(struct inode *inode, struct fuse_file *ff,
/* backing.c */
#ifdef CONFIG_FUSE_PASSTHROUGH
void fuse_backing_put(struct fuse_backing *fb);
-
#else
static inline void fuse_backing_put(struct fuse_backing *fb)
@@ -1320,6 +1323,7 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
#endif
struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
+int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb);
void fuse_backing_files_init(struct fuse_conn *fc);
void fuse_backing_files_init_64(struct fuse_conn *fc);
void fuse_backing_files_free(struct fuse_conn *fc);
@@ -1370,4 +1374,10 @@ extern void fuse_sysctl_unregister(void);
#define fuse_sysctl_unregister() do { } while (0)
#endif /* CONFIG_SYSCTL */
+/* ext_map.c */
+
+void fuse_ext_map_destroy(struct rb_root *extents);
+bool fuse_ext_map_is_dax(struct fuse_backing *fb);
+int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_out *arg,
+ struct fuse_extent *ext);
#endif /* _FS_FUSE_I_H */
diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
index 93e916a16ac9..7c427f88bbc5 100644
--- a/fs/fuse/notify.c
+++ b/fs/fuse/notify.c
@@ -434,6 +434,47 @@ static int fuse_notify_backing_remove(struct fuse_conn *fc, unsigned int size,
return fuse_backing_close_64(fc, outarg.backing_id);
}
+static int fuse_notify_map(struct fuse_conn *fc, unsigned int size,
+ struct fuse_copy_state *cs)
+{
+ struct fuse_notify_backing_map_out outarg;
+ struct fuse_extent *ext __free(kvfree) = NULL;
+ int err;
+
+ if (size < sizeof(outarg))
+ return -EINVAL;
+
+ err = fuse_copy_one(cs, &outarg, sizeof(outarg));
+ if (err)
+ return err;
+
+ if (outarg.num_extents > FUSE_MAX_EXTENTS)
+ return -EINVAL;
+
+ size -= sizeof(outarg);
+ if (size != outarg.num_extents * sizeof(*ext))
+ return -EINVAL;
+
+ if (outarg.reserved[0] != 0 || outarg.reserved[1] != 0)
+ return -EINVAL;
+
+ if (outarg.flags & ~FUSE_BACKING_MAP_CREATE)
+ return -EINVAL;
+
+ if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
+ return -EOPNOTSUPP;
+
+ ext = kvmalloc_objs(*ext, outarg.num_extents);
+ if (!ext)
+ return -ENOMEM;
+
+ err = fuse_copy_one(cs, ext, size);
+ if (err)
+ return err;
+
+ return fuse_ext_map_populate(fc, &outarg, ext);
+}
+
int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
unsigned int size, struct fuse_copy_state *cs)
{
@@ -468,6 +509,9 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
case FUSE_NOTIFY_BACKING_REMOVE:
return fuse_notify_backing_remove(fc, size, cs);
+ case FUSE_NOTIFY_BACKING_MAP:
+ return fuse_notify_map(fc, size, cs);
+
default:
return -EINVAL;
}
diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
index 4b3904b3acb7..60060da98e8d 100644
--- a/include/uapi/linux/fuse.h
+++ b/include/uapi/linux/fuse.h
@@ -255,6 +255,7 @@
* - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
* - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
* - add backing_id_64 to fuse_open_out
+ * - add FUSE_NOTIFY_BACKING_MAP, fuse_notify_backing_map_out, fuse_extent, FUSE_BACKING_MAP_CREATE
*/
#ifndef _LINUX_FUSE_H
@@ -716,6 +717,7 @@ enum fuse_notify_code {
FUSE_NOTIFY_INC_EPOCH = 8,
FUSE_NOTIFY_PRUNE = 9,
FUSE_NOTIFY_BACKING_REMOVE = 10,
+ FUSE_NOTIFY_BACKING_MAP = 11,
};
/* The read buffer is required to be at least 8k, but may be much larger */
@@ -1213,6 +1215,30 @@ struct fuse_notify_backing_remove_out {
uint64_t reserved;
};
+/**
+ * notify_map flags
+ *
+ * FUSE_BACKING_MAP_CREATE: create backing with the supplied ID
+ */
+#define FUSE_BACKING_MAP_CREATE (1 << 0)
+
+struct fuse_notify_backing_map_out {
+ uint64_t backing_id;
+ uint32_t num_extents;
+ uint32_t flags;
+ uint64_t reserved[2];
+};
+
+#define FUSE_MAX_EXTENTS 1365 /* (1 << 16) / sizeof(struct fuse_extent) */
+
+struct fuse_extent {
+ uint64_t offset; /* offset of extent into parent backing */
+ uint64_t length; /* extent length */
+ uint64_t backing_id; /* target backing */
+ uint64_t addr; /* target offset within backing file/device */
+ uint64_t reserved[2];
+};
+
#define FUSE_SETUPMAPPING_FLAG_WRITE (1ull << 0)
#define FUSE_SETUPMAPPING_FLAG_READ (1ull << 1)
struct fuse_setupmapping_in {
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v3 7/9] fuse: add extent map data structure
2026-10-06 18:01 ` [PATCH v3 7/9] fuse: add extent map data structure Miklos Szeredi
@ 2026-10-06 18:17 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-10-06 18:17 UTC (permalink / raw)
To: Miklos Szeredi; +Cc: linux-cxl, nvdimm
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] fuse: NULL pointer dereference in fuse_notify_map() due to missing backing_id_64 check
- [High] fuse: type confusion and wild pointer dereference in fuse_ext_map_is_dax()
- [Medium] fuse: duplicate backing_id insertion due to rhashtable_insert_fast usage
--
commit 7c30614b0de0feab07ac16cfbaac90b3eedab05c
Author: Miklos Szeredi <mszeredi@redhat.com>
fuse: add extent map data structure
Add support for creating and managing extent maps that map file regions to
dax device regions. Introduce FUSE_NOTIFY_BACKING_MAP for populating extent
maps from userspace.
> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index 0ed850ddf8cde..b7ebc951dc313 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
[ ... ]
> @@ -84,7 +88,7 @@ static const struct rhashtable_params fuse_backing_prm = {
> .key_len = sizeof_field(struct fuse_backing, backing_id),
> };
>
> -static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
> +int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
> {
> return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
> }
[Severity: Medium]
Will this blindly insert duplicate backing IDs into the hash table?
Calling rhashtable_insert_fast() passes NULL as the key to the internal
__rhashtable_insert_fast(), which short-circuits the duplicate comparison
logic. Since the newly added FUSE_NOTIFY_BACKING_MAP accepts user-provided
backing IDs and inserts them here, could an untrusted FUSE server send
multiple requests with the same backing ID and cause inconsistent state or
unbounded memory accumulation?
Should this use rhashtable_lookup_insert_fast() instead to properly check
for uniqueness?
> diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
> new file mode 100644
> index 0000000000000..39d7bb521873c
> --- /dev/null
> +++ b/fs/fuse/ext_map.c
[ ... ]
> +bool fuse_ext_map_is_dax(struct fuse_backing *fb)
> +{
> + struct fuse_iext *fie;
> +
> + if (WARN_ON(RB_EMPTY_ROOT(&fb->extents)))
> + return false;
> +
> + fie = rb_entry(fb->extents.rb_node, typeof(*fie), rb);
> + return fie->backing->type == FUSE_BACKING_DAXDEV;
> +}
[Severity: High]
Does this safely handle other backing types passed in via the fb pointer?
The extents field is part of an anonymous union overlaying dax_dev. If this
function is called on a FUSE_BACKING_DAXDEV backing, it will treat the
dax_dev pointer as an rb_root. If dax_dev is non-NULL, RB_EMPTY_ROOT()
returns false, and the code then casts the pointer via rb_entry() and
dereferences it.
Can this cause an out-of-bounds read or a wild pointer dereference? Should
there be a check to ensure fb->type == FUSE_BACKING_EXTMAP before attempting
to access fb->extents?
> diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
> index 93e916a16ac93..7c427f88bbc52 100644
> --- a/fs/fuse/notify.c
> +++ b/fs/fuse/notify.c
[ ... ]
> +static int fuse_notify_map(struct fuse_conn *fc, unsigned int size,
> + struct fuse_copy_state *cs)
> +{
> + struct fuse_notify_backing_map_out outarg;
> + struct fuse_extent *ext __free(kvfree) = NULL;
> + int err;
[ ... ]
> + if (outarg.flags & ~FUSE_BACKING_MAP_CREATE)
> + return -EINVAL;
> +
> + if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
> + return -EOPNOTSUPP;
> +
> + ext = kvmalloc_objs(*ext, outarg.num_extents);
> + if (!ext)
> + return -ENOMEM;
> +
> + err = fuse_copy_one(cs, ext, size);
> + if (err)
> + return err;
> +
> + return fuse_ext_map_populate(fc, &outarg, ext);
> +}
[Severity: High]
Is a check for fc->backing_id_64 missing here before attempting to populate
the extent map?
If the backing_id_64 feature was not negotiated with the FUSE server during
connection initialization, fc->backing_64_ht will remain uninitialized
(zero-filled) because fuse_backing_files_init_64() is never called.
When fuse_ext_map_populate() subsequently calls fuse_backing_add_64(), it
will use rhashtable_insert_fast() on an uninitialized hash table. Could this
result in a guaranteed NULL pointer dereference on ht->tbl, allowing a buggy
or malicious FUSE server to crash the kernel?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006180115.1425232-1-mszeredi@redhat.com?part=7
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v3 8/9] fuse: add extent map I/O support
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
` (6 preceding siblings ...)
2026-10-06 18:01 ` [PATCH v3 7/9] fuse: add extent map data structure Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
2026-10-06 18:22 ` sashiko-bot
2026-10-06 18:01 ` [PATCH v3 9/9] fuse: add support for striped backing Miklos Szeredi
8 siblings, 1 reply; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
Wire up read, write, splice and mmap operations for extent-mapped
files through the iomap/DAX infrastructure.
When a passthrough file is opened with an EXTMAP backing, I/O is
dispatched to the dax devices referenced by the extent map.
If an address is not mapped, EIO is returned on the I/O operation.
Originally-by: John Groves <john@groves.net>
Reviewed-by: Amir Goldstein <amir73il@gmail.com>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
fs/fuse/backing.c | 15 ++++
fs/fuse/ext_map.c | 161 ++++++++++++++++++++++++++++++++++++++++++
fs/fuse/file.c | 4 ++
fs/fuse/fuse_i.h | 4 ++
fs/fuse/notify.c | 3 +
fs/fuse/passthrough.c | 29 ++++++--
6 files changed, 210 insertions(+), 6 deletions(-)
diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index b7ebc951dc31..c53538448dfb 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -288,6 +288,21 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
return err;
}
+bool fuse_backing_is_dax(struct fuse_backing *fb)
+{
+ switch (fb->type) {
+ case FUSE_BACKING_PATH:
+ return false;
+ case FUSE_BACKING_DAXDEV:
+ return true;
+ case FUSE_BACKING_EXTMAP:
+ return fuse_ext_map_is_dax(fb);
+ default:
+ WARN_ON(1);
+ return false;
+ }
+}
+
struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id)
{
struct fuse_backing *fb;
diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
index 39d7bb521873..4f418185c0b7 100644
--- a/fs/fuse/ext_map.c
+++ b/fs/fuse/ext_map.c
@@ -3,6 +3,14 @@
#include "fuse_i.h"
#include <linux/rbtree.h>
#include <linux/pagemap.h>
+#include <linux/iomap.h>
+#include <linux/dax.h>
+
+#ifndef CONFIG_FS_DAX
+#define dax_iomap_rw(...) (-EIO)
+#define dax_iomap_fault(...) (-EIO)
+#define dax_finish_sync_fault(...) (-EIO)
+#endif
struct fuse_iext {
struct rb_node rb;
@@ -22,6 +30,24 @@ void fuse_ext_map_destroy(struct rb_root *extents)
}
}
+static struct fuse_iext *fuse_find_extent(struct rb_root *extents, u64 offset)
+{
+ struct rb_node *node = extents->rb_node;
+
+ while (node) {
+ struct fuse_iext *fie = rb_entry(node, typeof(*fie), rb);
+
+ if (offset < fie->start)
+ node = node->rb_left;
+ else if (offset >= fie->end)
+ node = node->rb_right;
+ else
+ return fie;
+ }
+
+ return NULL;
+}
+
static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
struct fuse_extent *ext)
{
@@ -78,6 +104,141 @@ static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
return 0;
}
+static int fuse_ext_map_iomap_begin(struct inode *inode, loff_t offset, loff_t length,
+ unsigned int flags, struct iomap *iomap, struct iomap *srcmap)
+{
+ struct fuse_backing *fb = fuse_inode_backing(get_fuse_inode(inode));
+ struct fuse_iext *fie;
+ loff_t ext_len;
+
+ if (!fb || fb->type != FUSE_BACKING_EXTMAP)
+ return fuse_EIO("missing or wrong type backing");
+
+ fie = fuse_find_extent(&fb->extents, offset);
+ if (!fie)
+ return fuse_EIO("missing mapping");
+
+ if (WARN_ON(fie->backing->type != FUSE_BACKING_DAXDEV))
+ return fuse_EIO("wrong type backing for extent");
+
+ if (fie->backing->dax_error) {
+ fuse_make_bad(inode);
+ return fuse_EIO("dax error");
+ }
+
+ ext_len = fie->end - fie->start;
+
+ iomap->offset = fie->start;
+ iomap->addr = fie->backing_offset;
+ iomap->length = ext_len;
+ iomap->dax_dev = fie->backing->dax_dev;
+ iomap->type = IOMAP_MAPPED;
+ iomap->flags = 0;
+
+ return 0;
+}
+
+static const struct iomap_ops fuse_ext_map_iomap_ops = {
+ .iomap_begin = fuse_ext_map_iomap_begin,
+};
+
+static vm_fault_t fuse_ext_map_huge_fault(struct vm_fault *vmf, unsigned int order)
+{
+ struct inode *inode = file_inode(vmf->vma->vm_file);
+ bool write_fault = (vmf->flags & FAULT_FLAG_WRITE) && (vmf->vma->vm_flags & VM_SHARED);
+ vm_fault_t ret;
+ unsigned long pfn;
+
+ if (WARN_ON_ONCE(!IS_DAX(inode)))
+ return VM_FAULT_SIGBUS;
+
+ if (write_fault) {
+ sb_start_pagefault(inode->i_sb);
+ file_update_time(vmf->vma->vm_file);
+ }
+
+ filemap_invalidate_lock_shared(inode->i_mapping);
+
+ ret = dax_iomap_fault(vmf, order, &pfn, NULL, &fuse_ext_map_iomap_ops);
+ if (ret & VM_FAULT_NEEDDSYNC)
+ ret = dax_finish_sync_fault(vmf, order, pfn);
+
+ filemap_invalidate_unlock_shared(inode->i_mapping);
+
+ if (write_fault)
+ sb_end_pagefault(inode->i_sb);
+
+ return ret;
+}
+
+static vm_fault_t fuse_ext_map_fault(struct vm_fault *vmf)
+{
+ return fuse_ext_map_huge_fault(vmf, 0);
+}
+
+static const struct vm_operations_struct fuse_ext_map_vm_ops = {
+ .fault = fuse_ext_map_fault,
+ .huge_fault = fuse_ext_map_huge_fault,
+ .page_mkwrite = fuse_ext_map_fault,
+ .pfn_mkwrite = fuse_ext_map_fault,
+};
+
+static void fuse_rw_clamp(struct kiocb *iocb, struct iov_iter *ubuf)
+{
+ struct inode *inode = iocb->ki_filp->f_mapping->host;
+ loff_t i_size = i_size_read(inode);
+ loff_t max_count = iocb->ki_pos >= i_size ? 0 : i_size - iocb->ki_pos;
+
+ if (iov_iter_count(ubuf) > max_count)
+ iov_iter_truncate(ubuf, max_count);
+}
+
+ssize_t fuse_ext_map_read_iter(struct kiocb *iocb, struct iov_iter *to)
+{
+ struct inode *inode = iocb->ki_filp->f_mapping->host;
+ ssize_t res;
+
+ guard(rwsem_read)(&inode->i_rwsem);
+
+ fuse_rw_clamp(iocb, to);
+
+ if (!iov_iter_count(to))
+ return 0;
+
+ res = dax_iomap_rw(iocb, to, &fuse_ext_map_iomap_ops);
+
+ file_accessed(iocb->ki_filp);
+ return res;
+}
+
+ssize_t fuse_ext_map_write_iter(struct kiocb *iocb, struct iov_iter *from)
+{
+ ssize_t res;
+
+ res = generic_write_checks(iocb, from);
+ if (res <= 0)
+ return res;
+
+ fuse_rw_clamp(iocb, from);
+
+ if (!iov_iter_count(from))
+ return -EFBIG;
+
+ res = kiocb_modified(iocb);
+ if (res)
+ return res;
+
+ return dax_iomap_rw(iocb, from, &fuse_ext_map_iomap_ops);
+}
+
+int fuse_ext_map_mmap(struct file *file, struct vm_area_struct *vma)
+{
+ file_accessed(file);
+ vma->vm_ops = &fuse_ext_map_vm_ops;
+ vm_flags_set(vma, VM_HUGEPAGE);
+ return 0;
+}
+
bool fuse_ext_map_is_dax(struct fuse_backing *fb)
{
struct fuse_iext *fie;
diff --git a/fs/fuse/file.c b/fs/fuse/file.c
index e0d72b9d3c4a..8d12d6162424 100644
--- a/fs/fuse/file.c
+++ b/fs/fuse/file.c
@@ -551,6 +551,10 @@ static int fuse_fsync(struct file *file, loff_t start, loff_t end,
if (fuse_is_bad(inode))
return -EIO;
+ /* Skip sync for DAX inodes */
+ if (IS_DAX(inode) && !fuse_inode_vdax(inode))
+ return 0;
+
inode_lock(inode);
/*
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 75eec5c17985..5d38dc1b6f0c 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -1324,6 +1324,7 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb);
+bool fuse_backing_is_dax(struct fuse_backing *fb);
void fuse_backing_files_init(struct fuse_conn *fc);
void fuse_backing_files_init_64(struct fuse_conn *fc);
void fuse_backing_files_free(struct fuse_conn *fc);
@@ -1377,6 +1378,9 @@ extern void fuse_sysctl_unregister(void);
/* ext_map.c */
void fuse_ext_map_destroy(struct rb_root *extents);
+ssize_t fuse_ext_map_write_iter(struct kiocb *iocb, struct iov_iter *from);
+ssize_t fuse_ext_map_read_iter(struct kiocb *iocb, struct iov_iter *to);
+int fuse_ext_map_mmap(struct file *file, struct vm_area_struct *vma);
bool fuse_ext_map_is_dax(struct fuse_backing *fb);
int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_out *arg,
struct fuse_extent *ext);
diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
index 7c427f88bbc5..3bfedd0cdbdd 100644
--- a/fs/fuse/notify.c
+++ b/fs/fuse/notify.c
@@ -448,6 +448,9 @@ static int fuse_notify_map(struct fuse_conn *fc, unsigned int size,
if (err)
return err;
+ if (!fc->backing_id_64)
+ return -EINVAL;
+
if (outarg.num_extents > FUSE_MAX_EXTENTS)
return -EINVAL;
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index e9ab1aea34e2..040817ad76e9 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -35,7 +35,6 @@ ssize_t fuse_passthrough_read_iter(struct kiocb *iocb, struct iov_iter *iter)
struct fuse_file *ff = file->private_data;
struct file *backing_file = fuse_file_passthrough(ff);
size_t count = iov_iter_count(iter);
- ssize_t ret;
struct backing_file_ctx ctx = {
.cred = ff->cred,
.accessed = fuse_file_accessed,
@@ -48,10 +47,10 @@ ssize_t fuse_passthrough_read_iter(struct kiocb *iocb, struct iov_iter *iter)
if (!count)
return 0;
- ret = backing_file_read_iter(backing_file, iter, iocb, iocb->ki_flags,
- &ctx);
+ if (!backing_file)
+ return fuse_ext_map_read_iter(iocb, iter);
- return ret;
+ return backing_file_read_iter(backing_file, iter, iocb, iocb->ki_flags, &ctx);
}
ssize_t fuse_passthrough_write_iter(struct kiocb *iocb,
@@ -74,10 +73,13 @@ ssize_t fuse_passthrough_write_iter(struct kiocb *iocb,
if (!count)
return 0;
- inode_lock(inode);
+ guard(rwsem_write)(&inode->i_rwsem);
+
+ if (!backing_file)
+ return fuse_ext_map_write_iter(iocb, iter);
+
ret = backing_file_write_iter(backing_file, iter, iocb, iocb->ki_flags,
&ctx);
- inode_unlock(inode);
return ret;
}
@@ -98,6 +100,9 @@ ssize_t fuse_passthrough_splice_read(struct file *in, loff_t *ppos,
pr_debug("%s: backing_file=0x%p, pos=%lld, len=%zu, flags=0x%x\n", __func__,
backing_file, *ppos, len, flags);
+ if (!backing_file)
+ return copy_splice_read(in, ppos, pipe, len, flags);
+
init_sync_kiocb(&iocb, in);
iocb.ki_pos = *ppos;
ret = backing_file_splice_read(backing_file, &iocb, pipe, len, flags, &ctx);
@@ -123,6 +128,9 @@ ssize_t fuse_passthrough_splice_write(struct pipe_inode_info *pipe,
pr_debug("%s: backing_file=0x%p, pos=%lld, len=%zu, flags=0x%x\n", __func__,
backing_file, *ppos, len, flags);
+ if (!backing_file)
+ return iter_file_splice_write(pipe, out, ppos, len, flags);
+
inode_lock(inode);
init_sync_kiocb(&iocb, out);
iocb.ki_pos = *ppos;
@@ -145,6 +153,9 @@ ssize_t fuse_passthrough_mmap(struct file *file, struct vm_area_struct *vma)
pr_debug("%s: backing_file=0x%p, start=%lu, end=%lu\n", __func__,
backing_file, vma->vm_start, vma->vm_end);
+ if (!backing_file)
+ return fuse_ext_map_mmap(file, vma);
+
return backing_file_mmap(backing_file, vma, &ctx);
}
@@ -156,6 +167,12 @@ int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
struct fuse_file *ff = file->private_data;
struct file *backing_file;
+ if (fb->type == FUSE_BACKING_EXTMAP) {
+ if (fuse_backing_is_dax(fb) != !!IS_DAX(file_inode(file)))
+ return fuse_EIO("dax mode mismatch");
+ return 0;
+ }
+
if (fb->type != FUSE_BACKING_PATH)
return fuse_EIO("invalid backing type");
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread* Re: [PATCH v3 8/9] fuse: add extent map I/O support
2026-10-06 18:01 ` [PATCH v3 8/9] fuse: add extent map I/O support Miklos Szeredi
@ 2026-10-06 18:22 ` sashiko-bot
0 siblings, 0 replies; 23+ messages in thread
From: sashiko-bot @ 2026-10-06 18:22 UTC (permalink / raw)
To: Miklos Szeredi; +Cc: linux-cxl, nvdimm
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] fuse: fsync skips CPU cache flushing and metadata sync for EXTMAP DAX
- [High] fuse: dax_iomap_fault stubs return -EIO instead of valid vm_fault_t
- [Medium] fuse: incorrect -EFBIG error on write past EOF for EXTMAP DAX
--
commit 6c74cf2f1c1750d6fe1df2fa7ef25325ee0d117e
Author: Miklos Szeredi <mszeredi@redhat.com>
fuse: add extent map I/O support
This commit wires up read, write, splice, and mmap operations for
extent-mapped files using the iomap and DAX infrastructure. When a
passthrough file is opened with an EXTMAP backing, I/O is dispatched
to the associated DAX devices.
> diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
> index 39d7bb521873c..4f418185c0b79 100644
> --- a/fs/fuse/ext_map.c
> +++ b/fs/fuse/ext_map.c
> @@ -3,6 +3,14 @@
> #include "fuse_i.h"
> #include <linux/rbtree.h>
> #include <linux/pagemap.h>
> +#include <linux/iomap.h>
> +#include <linux/dax.h>
> +
> +#ifndef CONFIG_FS_DAX
> +#define dax_iomap_rw(...) (-EIO)
> +#define dax_iomap_fault(...) (-EIO)
> +#define dax_finish_sync_fault(...) (-EIO)
> +#endif
[Severity: High]
Can returning -EIO here cause memory management fault handler corruption?
When CONFIG_FS_DAX is disabled, these macros stub out the DAX fault
handlers to return -EIO. However, the return type of the fault handlers
is vm_fault_t, which expects a bitmask of VM_FAULT_* flags (such as
VM_FAULT_SIGBUS).
Returning a negative integer like -EIO is interpreted as setting almost
all VM_FAULT_x bits simultaneously. This might lead to unpredictable
kernel behavior or infinite fault loops when handled by the core memory
management fault logic.
[ ... ]
> +static vm_fault_t fuse_ext_map_huge_fault(struct vm_fault *vmf, unsigned int order)
> +{
> + struct inode *inode = file_inode(vmf->vma->vm_file);
> + bool write_fault = (vmf->flags & FAULT_FLAG_WRITE) && (vmf->vma->vm_flags & VM_SHARED);
> + vm_fault_t ret;
> + unsigned long pfn;
> +
> + if (WARN_ON_ONCE(!IS_DAX(inode)))
> + return VM_FAULT_SIGBUS;
> +
[ ... ]
> +
> + filemap_invalidate_lock_shared(inode->i_mapping);
> +
> + ret = dax_iomap_fault(vmf, order, &pfn, NULL, &fuse_ext_map_iomap_ops);
> + if (ret & VM_FAULT_NEEDDSYNC)
> + ret = dax_finish_sync_fault(vmf, order, pfn);
> +
> + filemap_invalidate_unlock_shared(inode->i_mapping);
> +
> + if (write_fault)
> + sb_end_pagefault(inode->i_sb);
> +
> + return ret;
> +}
[ ... ]
> +ssize_t fuse_ext_map_write_iter(struct kiocb *iocb, struct iov_iter *from)
> +{
> + ssize_t res;
> +
> + res = generic_write_checks(iocb, from);
> + if (res <= 0)
> + return res;
> +
> + fuse_rw_clamp(iocb, from);
> +
> + if (!iov_iter_count(from))
> + return -EFBIG;
[Severity: Medium]
Is -EFBIG the correct error code to return for writes past the end of an
extent-mapped file?
The fuse_rw_clamp() call truncates the iterator if the write starts at or
beyond the current file size. Returning -EFBIG in this condition maps a
standard EOF or out-of-space scenario to a file too large error.
POSIX semantics generally expect -ENOSPC or 0 (a short write) when writing
past a non-extensible block device or fixed-size mapping. Returning -EFBIG
might cause applications like cp or dd to fail incorrectly or inadvertently
trigger a SIGXFSZ signal.
> +
> + res = kiocb_modified(iocb);
> + if (res)
> + return res;
> +
> + return dax_iomap_rw(iocb, from, &fuse_ext_map_iomap_ops);
> +}
[ ... ]
> diff --git a/fs/fuse/file.c b/fs/fuse/file.c
> index b5bef7b9e466e..b5c66f63acd02 100644
> --- a/fs/fuse/file.c
> +++ b/fs/fuse/file.c
> @@ -551,6 +551,10 @@ static int fuse_fsync(struct file *file, loff_t start, loff_t end,
> if (fuse_is_bad(inode))
> return -EIO;
>
> + /* Skip sync for DAX inodes */
> + if (IS_DAX(inode) && !fuse_inode_vdax(inode))
> + return 0;
[Severity: High]
Does this early return bypass CPU cache flushing and metadata sync for
EXTMAP DAX files?
For FUSE passthrough EXTMAP files, IS_DAX(inode) will be true, but
fuse_inode_vdax(inode) evaluates to false since vdax is specific to the
traditional virtio-fs DAX window.
Returning 0 here completely skips file_write_and_wait_range(), which is
needed to flush CPU caches for DAX memory via dax_writeback_mapping_range().
It also skips sync_inode_metadata(), which syncs metadata updates to the
FUSE server. This bypass could cause CPU-cached DAX writes to be lost on a
crash, violating DAX durability guarantees.
> +
> inode_lock(inode);
>
> /*
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006180115.1425232-1-mszeredi@redhat.com?part=8
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v3 9/9] fuse: add support for striped backing
2026-10-06 18:01 [PATCH v3 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
` (7 preceding siblings ...)
2026-10-06 18:01 ` [PATCH v3 8/9] fuse: add extent map I/O support Miklos Szeredi
@ 2026-10-06 18:01 ` Miklos Szeredi
8 siblings, 0 replies; 23+ messages in thread
From: Miklos Szeredi @ 2026-10-06 18:01 UTC (permalink / raw)
To: fuse-devel
Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
Dave Jiang, Alison Schofield, nvdimm, linux-cxl
Add the FUSE_BACKING_MAP_CYCLIC flag for FUSE_NOTIFY_BACKING_MAP
notifications to enable creating striped maps.
This is just a set of regular extents that repeats after the end of the
last extent using a striping pattern (i.e. on the target backings the
stripes have a contiguous footprint).
Two additional limits apply compared to non-cyclic extent map population:
- the size of the extents (chunk size) must be equal
- the extents must start from zero offset and follow each other without
gaps.
Originally-by: John Groves <john@groves.net>
Reviewed-by: Amir Goldstein <amir73il@gmail.com>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
fs/fuse/ext_map.c | 30 +++++++++++++++++++++++++++---
fs/fuse/fuse_i.h | 1 +
fs/fuse/notify.c | 2 +-
include/uapi/linux/fuse.h | 5 ++++-
4 files changed, 33 insertions(+), 5 deletions(-)
diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
index 4f418185c0b7..049e0b4514bc 100644
--- a/fs/fuse/ext_map.c
+++ b/fs/fuse/ext_map.c
@@ -109,12 +109,18 @@ static int fuse_ext_map_iomap_begin(struct inode *inode, loff_t offset, loff_t l
{
struct fuse_backing *fb = fuse_inode_backing(get_fuse_inode(inode));
struct fuse_iext *fie;
+ u64 ncycle = 0, seq_off = 0, addr, tmp;
loff_t ext_len;
if (!fb || fb->type != FUSE_BACKING_EXTMAP)
return fuse_EIO("missing or wrong type backing");
- fie = fuse_find_extent(&fb->extents, offset);
+ if (fb->cycle_length) {
+ ncycle = div64_u64(offset, fb->cycle_length);
+ seq_off = ncycle * fb->cycle_length;
+ }
+
+ fie = fuse_find_extent(&fb->extents, offset - seq_off);
if (!fie)
return fuse_EIO("missing mapping");
@@ -127,9 +133,15 @@ static int fuse_ext_map_iomap_begin(struct inode *inode, loff_t offset, loff_t l
}
ext_len = fie->end - fie->start;
+ addr = fie->backing_offset;
+ if (check_mul_overflow(ncycle, ext_len, &tmp) || check_add_overflow(addr, tmp, &addr))
+ return fuse_EIO("extent address overflow");
- iomap->offset = fie->start;
- iomap->addr = fie->backing_offset;
+ if (check_add_overflow(addr, ext_len, &tmp))
+ return fuse_EIO("extent end overflow");
+
+ iomap->offset = fie->start + seq_off;
+ iomap->addr = addr;
iomap->length = ext_len;
iomap->dax_dev = fie->backing->dax_dev;
iomap->type = IOMAP_MAPPED;
@@ -256,6 +268,7 @@ int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_o
struct fuse_backing *fb;
unsigned int i;
int err;
+ u64 chunk_size = 0;
if (!arg->num_extents)
return fuse_EIO("no extents");
@@ -273,7 +286,18 @@ int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_o
return -EINVAL;
}
+ if (arg->flags & FUSE_BACKING_MAP_CYCLIC) {
+ chunk_size = ext[0].length;
+ err = -EINVAL;
+ if (check_mul_overflow(chunk_size, arg->num_extents, &fb->cycle_length))
+ goto err_put;
+ }
+
for (i = 0; i < arg->num_extents; i++) {
+ err = -EINVAL;
+ if (chunk_size && (ext[i].offset != chunk_size * i || ext[i].length != chunk_size))
+ goto err_put;
+
err = fuse_add_extent(fc, &fb->extents, &ext[i]);
if (err)
goto err_put;
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 5d38dc1b6f0c..cd10156e09d4 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -110,6 +110,7 @@ struct fuse_backing {
};
struct {
struct rb_root extents;
+ u64 cycle_length;
};
};
u64 backing_id;
diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
index 3bfedd0cdbdd..31172de72477 100644
--- a/fs/fuse/notify.c
+++ b/fs/fuse/notify.c
@@ -461,7 +461,7 @@ static int fuse_notify_map(struct fuse_conn *fc, unsigned int size,
if (outarg.reserved[0] != 0 || outarg.reserved[1] != 0)
return -EINVAL;
- if (outarg.flags & ~FUSE_BACKING_MAP_CREATE)
+ if (outarg.flags & ~(FUSE_BACKING_MAP_CREATE | FUSE_BACKING_MAP_CYCLIC))
return -EINVAL;
if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
index 60060da98e8d..6b2fe68b4ab5 100644
--- a/include/uapi/linux/fuse.h
+++ b/include/uapi/linux/fuse.h
@@ -255,7 +255,8 @@
* - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
* - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
* - add backing_id_64 to fuse_open_out
- * - add FUSE_NOTIFY_BACKING_MAP, fuse_notify_backing_map_out, fuse_extent, FUSE_BACKING_MAP_CREATE
+ * - add FUSE_NOTIFY_BACKING_MAP, fuse_notify_backing_map_out, fuse_extent
+ * - add FUSE_BACKING_MAP_CREATE, FUSE_BACKING_MAP_CYCLIC
*/
#ifndef _LINUX_FUSE_H
@@ -1219,8 +1220,10 @@ struct fuse_notify_backing_remove_out {
* notify_map flags
*
* FUSE_BACKING_MAP_CREATE: create backing with the supplied ID
+ * FUSE_BACKING_MAP_CYCLIC: map repeats after last extent
*/
#define FUSE_BACKING_MAP_CREATE (1 << 0)
+#define FUSE_BACKING_MAP_CYCLIC (1 << 1)
struct fuse_notify_backing_map_out {
uint64_t backing_id;
--
2.54.0
^ permalink raw reply related [flat|nested] 23+ messages in thread