NVDIMM Device and Persistent Memory development
 help / color / mirror / Atom feed
* [PATCH v2 0/8] fuse: DAX device based extent maps (famfs)
@ 2026-10-01 15:07 Miklos Szeredi
  2026-10-01 15:07 ` [PATCH v2 1/8] dax: replace exported dax_dev_get() with non-allocating dax_dev_find() Miklos Szeredi
                   ` (8 more replies)
  0 siblings, 9 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

This is the FUSIfication of the famfs-fuse patchset.

 - extent maps are assigned to a backing ID

 - 64 bit backing ID's managed by server

 - maps can be linear or cyclic (striped)

 - API supports nesting, but not implemented yet (i.e. multiple striped
   extents per inode, a-la famfs)

libfuse/famfs trees for testing:

 https://github.com/szmi/libfuse.git#extent-map-v2
 https://github.com/szmi/famfs.git#extent-map-v2

This patchset:

 git://git.kernel.org/pub/scm/linux/kernel/git/mszeredi/fuse.git#extent-map-v2

Passes famfs smoke test on emulated dax device.

Comments are welcome.

Thanks,
Miklos

---
Changes since v1:

 - moved 3 prep patches to fuse.git#for-next
 - introduce FUSE_PASSTHROUGH_V2 (Amir)
 - lots of small fixes and cleanups (Amir, Sashiko)

---
John Groves (1):
  dax: replace exported dax_dev_get() with non-allocating dax_dev_find()

Miklos Szeredi (7):
  fuse: add helpers for EIO return value with kernel message
  fuse: support 64 bit, server allocated backing ID
  fuse: support opening 64 bit backing ID
  fuse: add support for opening dax device as backing
  fuse: add extent map data structure
  fuse: add extent map I/O support
  fuse: add support for striped backing

 drivers/dax/super.c       |  38 ++++-
 fs/fuse/Makefile          |   2 +-
 fs/fuse/backing.c         | 292 ++++++++++++++++++++++++++++++-------
 fs/fuse/dev.c             |  21 +++
 fs/fuse/dev.h             |   3 +
 fs/fuse/dir.c             |   3 +-
 fs/fuse/ext_map.c         | 299 ++++++++++++++++++++++++++++++++++++++
 fs/fuse/file.c            |  16 +-
 fs/fuse/fuse_i.h          |  72 +++++++--
 fs/fuse/inode.c           |  17 ++-
 fs/fuse/iomode.c          | 108 +++++++-------
 fs/fuse/notify.c          |  75 ++++++++++
 fs/fuse/passthrough.c     |  63 ++++----
 include/linux/dax.h       |   6 +-
 include/uapi/linux/fuse.h |  53 ++++++-
 15 files changed, 904 insertions(+), 164 deletions(-)
 create mode 100644 fs/fuse/ext_map.c

-- 
2.54.0


^ permalink raw reply	[flat|nested] 29+ messages in thread

* [PATCH v2 1/8] dax: replace exported dax_dev_get() with non-allocating dax_dev_find()
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
@ 2026-10-01 15:07 ` Miklos Szeredi
  2026-10-01 15:07 ` [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
                   ` (7 subsequent siblings)
  8 siblings, 0 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, John Groves, Amir Goldstein, Darrick J . Wong,
	Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl

From: John Groves <John@Groves.net>

This fix is in response to a Sashiko review, and some subsequent
analysis.

dax_dev_get() uses iget5_locked() which creates a new inode if no
matching one exists. This is correct for the internal caller
(alloc_dax), but dangerous for external callers that look up devices
from user-supplied or metadata-supplied dev_t values:

1. A new inode is created with DAXDEV_ALIVE set but no backing driver,
   no ops, and no IDA-allocated minor number.

2. On teardown, dax_destroy_inode() warns because kill_dax() was never
   called, and dax_free_inode() calls ida_free() for a minor that was
   never ida_alloc'd -- potentially freeing the minor of a real device.

Add dax_dev_find() which uses ilookup5() for lookup-only semantics:
it returns an existing dax_device with an elevated inode reference, or
NULL if no device with the given dev_t exists. It never creates inodes.
A dax_alive() check under dax_read_lock() guards against returning a
device that is concurrently being torn down by kill_dax().

Make dax_dev_get() static again (internal to super.c for alloc_dax),
export dax_dev_find() instead.  Also add the missing CONFIG_DAX=n stub.

About the 'fixes' tag: this removes the export of dax_dev_get(),
which was flawed, and replaces is with dax_dev_find(). It feels like
the fixes tag makes sense for correcting an ABI error.

Fixes: 2ae624d5a555d ("dax: export dax_dev_get()")
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Reviewed-by: Alison Schofield <alison.schofield@intel.com>
Reviewed-by: "Darrick J. Wong" <djwong@kernel.org>
Signed-off-by: John Groves <john@groves.net>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
 drivers/dax/super.c | 38 ++++++++++++++++++++++++++++++++++++--
 include/linux/dax.h |  6 +++++-
 2 files changed, 41 insertions(+), 3 deletions(-)

diff --git a/drivers/dax/super.c b/drivers/dax/super.c
index 45f84b0eb909..824e1f6df378 100644
--- a/drivers/dax/super.c
+++ b/drivers/dax/super.c
@@ -565,7 +565,7 @@ static int dax_set(struct inode *inode, void *data)
 	return 0;
 }
 
-struct dax_device *dax_dev_get(dev_t devt)
+static struct dax_device *dax_dev_get(dev_t devt)
 {
 	struct dax_device *dax_dev;
 	struct inode *inode;
@@ -588,7 +588,41 @@ struct dax_device *dax_dev_get(dev_t devt)
 
 	return dax_dev;
 }
-EXPORT_SYMBOL_GPL(dax_dev_get);
+
+/**
+ * dax_dev_find - look up an existing dax_device by dev_t
+ * @devt: the device number to find
+ *
+ * Returns a dax_device with an elevated inode reference, or NULL if no
+ * device with the given dev_t exists. Unlike dax_dev_get(), this never
+ * allocates a new inode -- it is safe for external callers that are looking
+ * up devices from user-supplied or metadata-supplied dev_t values.
+ *
+ * Caller must put_dax() the returned device when done.
+ */
+struct dax_device *dax_dev_find(dev_t devt)
+{
+	struct dax_device *dax_dev;
+	struct inode *inode;
+	int id;
+
+	inode = ilookup5(dax_superblock, hash_32(devt + DAXFS_MAGIC, 31),
+			 dax_test, &devt);
+	if (!inode)
+		return NULL;
+
+	dax_dev = to_dax_dev(inode);
+	id = dax_read_lock();
+	if (!dax_alive(dax_dev)) {
+		dax_read_unlock(id);
+		iput(inode);
+		return NULL;
+	}
+	dax_read_unlock(id);
+
+	return dax_dev;
+}
+EXPORT_SYMBOL_GPL(dax_dev_find);
 
 struct dax_device *alloc_dax(void *private, const struct dax_operations *ops)
 {
diff --git a/include/linux/dax.h b/include/linux/dax.h
index fe6c3ded1b50..29113eb95e72 100644
--- a/include/linux/dax.h
+++ b/include/linux/dax.h
@@ -54,7 +54,7 @@ struct dax_device *alloc_dax(void *private, const struct dax_operations *ops);
 void *dax_holder(struct dax_device *dax_dev);
 void put_dax(struct dax_device *dax_dev);
 void kill_dax(struct dax_device *dax_dev);
-struct dax_device *dax_dev_get(dev_t devt);
+struct dax_device *dax_dev_find(dev_t devt);
 void dax_write_cache(struct dax_device *dax_dev, bool wc);
 bool dax_write_cache_enabled(struct dax_device *dax_dev);
 bool dax_synchronous(struct dax_device *dax_dev);
@@ -92,6 +92,10 @@ static inline void put_dax(struct dax_device *dax_dev)
 static inline void kill_dax(struct dax_device *dax_dev)
 {
 }
+static inline struct dax_device *dax_dev_find(dev_t devt)
+{
+	return NULL;
+}
 static inline void dax_write_cache(struct dax_device *dax_dev, bool wc)
 {
 }
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
  2026-10-01 15:07 ` [PATCH v2 1/8] dax: replace exported dax_dev_get() with non-allocating dax_dev_find() Miklos Szeredi
@ 2026-10-01 15:07 ` Miklos Szeredi
  2026-10-01 15:18   ` sashiko-bot
  2026-10-01 16:32   ` Amir Goldstein
  2026-10-01 15:07 ` [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
                   ` (6 subsequent siblings)
  8 siblings, 2 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

Conditions are triggered by a buggy fuse server will result in -EIO which
is hard to interpret without any context.

Add helpers that print a short message to the kernel log as well as
returning the error value.  This serves a dual purpose:

 - documents the error condition inside the code

 - allows the implementor of the fuse server to get the context for the
   error

Some debug messages are removed in favor of this.

This also changes the mmap return value from -ETXTBSY to -ENODEV if a
passthrough open raced with the prior check.  This makes both cases return
-ENODEV if the file is in passthrough mode.

Additionally change the return value of fuse_inode_uncached_io_start() and
fuse_file_cached_io_open() from an error value to a bool (true on success)
as the error value is translated anyway.

Currently these use the pr_notice_once() variant.  Possibly should be
changed to a ratelimit of e.g. once per minute.

Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
 fs/fuse/file.c        |  8 ++---
 fs/fuse/fuse_i.h      |  6 ++--
 fs/fuse/iomode.c      | 74 ++++++++++++++++---------------------------
 fs/fuse/passthrough.c | 19 ++++-------
 4 files changed, 42 insertions(+), 65 deletions(-)

diff --git a/fs/fuse/file.c b/fs/fuse/file.c
index 6d707f2b3bff..92976906ab05 100644
--- a/fs/fuse/file.c
+++ b/fs/fuse/file.c
@@ -1478,7 +1478,7 @@ static void fuse_dio_lock(struct kiocb *iocb, struct iov_iter *from,
 		 * have raced, so check it again.
 		 */
 		if (fuse_io_past_eof(iocb, from) ||
-		    fuse_inode_uncached_io_start(fi, NULL) != 0) {
+		    !fuse_inode_uncached_io_start(fi, NULL)) {
 			inode_unlock_shared(inode);
 			inode_lock(inode);
 			*exclusive = true;
@@ -2415,7 +2415,6 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma)
 	struct fuse_file *ff = file->private_data;
 	struct fuse_conn *fc = ff->fm->fc;
 	struct inode *inode = file_inode(file);
-	int rc;
 
 	/* DAX mmap is superior to direct_io mmap */
 	if (FUSE_IS_VDAX(inode))
@@ -2457,9 +2456,8 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma)
 		 * After first mmap, the inode stays in caching io mode until
 		 * the direct_io file release.
 		 */
-		rc = fuse_file_cached_io_open(inode, ff);
-		if (rc)
-			return rc;
+		if (!fuse_file_cached_io_open(inode, ff))
+			return -ENODEV;
 	}
 
 	if ((vma->vm_flags & VM_SHARED) && (vma->vm_flags & VM_MAYWRITE))
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 8546855386b5..3dd4dff24c50 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -11,6 +11,8 @@
 # define pr_fmt(fmt) "fuse: " fmt
 #endif
 
+#define fuse_EIO(fmt, ...) (pr_notice_once("%s: " fmt "\n", __func__, ##__VA_ARGS__), -EIO)
+
 #include "args.h"
 #include <linux/fuse.h>
 #include <linux/fs.h>
@@ -1251,8 +1253,8 @@ int fuse_fileattr_set(struct mnt_idmap *idmap,
 		      struct dentry *dentry, struct file_kattr *fa);
 
 /* iomode.c */
-int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff);
-int fuse_inode_uncached_io_start(struct fuse_inode *fi,
+bool fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff);
+bool fuse_inode_uncached_io_start(struct fuse_inode *fi,
 				 struct fuse_backing *fb);
 void fuse_inode_uncached_io_end(struct fuse_inode *fi);
 
diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
index 79637c09e883..1a10bc65fb31 100644
--- a/fs/fuse/iomode.c
+++ b/fs/fuse/iomode.c
@@ -27,15 +27,15 @@ static inline bool fuse_is_io_cache_wait(struct fuse_inode *fi)
  * Blocks new parallel dio writes and waits for the in-progress parallel dio
  * writes to complete.
  */
-int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
+bool fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
 {
 	struct fuse_inode *fi = get_fuse_inode(inode);
 
 	/* There are no io modes if server does not implement open */
 	if (!ff->args)
-		return 0;
+		return true;
 
-	spin_lock(&fi->lock);
+	guard(spinlock)(&fi->lock);
 	/*
 	 * Setting the bit advises new direct-io writes to use an exclusive
 	 * lock - without it the wait below might be forever.
@@ -53,8 +53,7 @@ int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
 	 */
 	if (fuse_inode_backing(fi)) {
 		clear_bit(FUSE_I_CACHE_IO_MODE, &fi->state);
-		spin_unlock(&fi->lock);
-		return -ETXTBSY;
+		return false;
 	}
 
 	WARN_ON(ff->iomode == IOM_UNCACHED);
@@ -64,8 +63,7 @@ int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
 			set_bit(FUSE_I_CACHE_IO_MODE, &fi->state);
 		fi->iocachectr++;
 	}
-	spin_unlock(&fi->lock);
-	return 0;
+	return true;
 }
 
 static void fuse_file_cached_io_release(struct fuse_file *ff,
@@ -82,22 +80,19 @@ static void fuse_file_cached_io_release(struct fuse_file *ff,
 }
 
 /* Start strictly uncached io mode where cache access is not allowed */
-int fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
+bool fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
 {
 	struct fuse_backing *oldfb;
-	int err = 0;
 
-	spin_lock(&fi->lock);
+	guard(spinlock)(&fi->lock);
 	/* deny conflicting backing files on same fuse inode */
 	oldfb = fuse_inode_backing(fi);
-	if (fb && oldfb && oldfb != fb) {
-		err = -EBUSY;
-		goto unlock;
-	}
-	if (fi->iocachectr > 0) {
-		err = -ETXTBSY;
-		goto unlock;
-	}
+	if (fb && oldfb && oldfb != fb)
+		return false;
+
+	if (fi->iocachectr > 0)
+		return false;
+
 	fi->iocachectr--;
 
 	/* fuse inode holds a single refcount of backing file */
@@ -107,9 +102,7 @@ int fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
 	} else {
 		fuse_backing_put(fb);
 	}
-unlock:
-	spin_unlock(&fi->lock);
-	return err;
+	return true;
 }
 
 /* Takes uncached_io inode mode reference to be dropped on file release */
@@ -118,11 +111,9 @@ static int fuse_file_uncached_io_open(struct inode *inode,
 				      struct fuse_backing *fb)
 {
 	struct fuse_inode *fi = get_fuse_inode(inode);
-	int err;
 
-	err = fuse_inode_uncached_io_start(fi, fb);
-	if (err)
-		return err;
+	if (!fuse_inode_uncached_io_start(fi, fb))
+		return fuse_EIO("failed to start uncached I/O");
 
 	WARN_ON(ff->iomode != IOM_NONE);
 	ff->iomode = IOM_UNCACHED;
@@ -173,9 +164,11 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
 	int err;
 
 	/* Check allowed conditions for file open in passthrough mode */
-	if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) || !fc->passthrough ||
-	    (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK))
-		return -EINVAL;
+	if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) || !fc->passthrough)
+		return fuse_EIO("passthrough not enabled");
+
+	if (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK)
+		return fuse_EIO("conflicting open flags");
 
 	fb = fuse_passthrough_open(file, ff->args->open_outarg.backing_id);
 	if (IS_ERR(fb))
@@ -208,11 +201,12 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
 
 	/*
 	 * Server is expected to use FOPEN_PASSTHROUGH for all opens of an inode
-	 * which is already open for passthrough.
+	 * which is already open for passthrough.  Using incorrect open mode is
+	 * a server mistake, which results in user visible failure of open()
+	 * with EIO error.
 	 */
-	err = -EINVAL;
 	if (fuse_inode_backing(fi) && !(ff->open_flags & FOPEN_PASSTHROUGH))
-		goto fail;
+		return fuse_EIO("FOPEN_PASSTHROUGH expected");
 
 	/*
 	 * FOPEN_PARALLEL_DIRECT_WRITES requires FOPEN_DIRECT_IO.
@@ -234,22 +228,10 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
 
 	if (ff->open_flags & FOPEN_PASSTHROUGH)
 		err = fuse_file_passthrough_open(inode, file);
-	else
-		err = fuse_file_cached_io_open(inode, ff);
-	if (err)
-		goto fail;
+	else if (!fuse_file_cached_io_open(inode, ff))
+		err = fuse_EIO("conflicting passthrough open");
 
-	return 0;
-
-fail:
-	pr_debug("failed to open file in requested io mode (open_flags=0x%x, err=%i).\n",
-		 ff->open_flags, err);
-	/*
-	 * The file open mode determines the inode io mode.
-	 * Using incorrect open mode is a server mistake, which results in
-	 * user visible failure of open() with EIO error.
-	 */
-	return -EIO;
+	return err;
 }
 
 /* No more pending io and no new io possible to inode via open/mmapped file */
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index b43d3e0f7081..489838462d1e 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -159,34 +159,29 @@ struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
 	struct fuse_conn *fc = ff->fm->fc;
 	struct fuse_backing *fb = NULL;
 	struct file *backing_file;
-	int err;
 
-	err = -EINVAL;
 	if (backing_id <= 0)
-		goto out;
+		return ERR_PTR(fuse_EIO("invalid backing_id"));
 
-	err = -ENOENT;
 	fb = fuse_backing_lookup(fc, backing_id);
 	if (!fb)
-		goto out;
+		return ERR_PTR(fuse_EIO("backing not found"));
 
 	/* Allocate backing file per fuse file to store fuse path */
 	backing_file = backing_file_open(file, file->f_flags,
 					 &fb->file->f_path, fb->cred);
-	err = PTR_ERR(backing_file);
 	if (IS_ERR(backing_file)) {
 		fuse_backing_put(fb);
-		goto out;
+		return ERR_PTR(fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file)));
 	}
 
-	err = 0;
 	ff->passthrough = backing_file;
 	ff->cred = get_cred(fb->cred);
-out:
-	pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p, err=%i\n", __func__,
-		 backing_id, fb, ff->passthrough, err);
 
-	return err ? ERR_PTR(err) : fb;
+	pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p\n", __func__,
+		 backing_id, fb, ff->passthrough);
+
+	return fb;
 }
 
 void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb)
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
  2026-10-01 15:07 ` [PATCH v2 1/8] dax: replace exported dax_dev_get() with non-allocating dax_dev_find() Miklos Szeredi
  2026-10-01 15:07 ` [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
@ 2026-10-01 15:07 ` Miklos Szeredi
  2026-10-01 15:26   ` sashiko-bot
  2026-10-01 17:07   ` Amir Goldstein
  2026-10-01 15:07 ` [PATCH v2 4/8] fuse: support opening 64 bit " Miklos Szeredi
                   ` (5 subsequent siblings)
  8 siblings, 2 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

Add support for server allocated 64-bit backing IDs alongside the
existing kernel allocated 32-bit IDs.

Opened with FUSE_DEV_IOC_BACKING_OPEN, but the backing ID is sent via
fuse_backing_map.backing_id, with FUSE_BACKING_ID_64 being set in .flags.

Closed with FUSE_NOTIFY_BACKING_REMOVE.  Since close only provides the
backing ID, not the file descriptor, it doesn't have to be done with an
ioctl.

All 64 bit values are valid.

Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
 fs/fuse/backing.c         | 180 ++++++++++++++++++++++++++++----------
 fs/fuse/dev.c             |  21 +++++
 fs/fuse/dev.h             |   3 +
 fs/fuse/fuse_i.h          |  21 ++++-
 fs/fuse/inode.c           |   6 +-
 fs/fuse/notify.c          |  28 ++++++
 fs/fuse/passthrough.c     |   3 +
 include/uapi/linux/fuse.h |  22 ++++-
 8 files changed, 231 insertions(+), 53 deletions(-)

diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index 433fa3098d71..3c879df7989c 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -9,6 +9,7 @@
 #include "fuse_i.h"
 
 #include <linux/file.h>
+#include <linux/rhashtable.h>
 
 static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb)
 {
@@ -48,6 +49,8 @@ static int fuse_backing_id_alloc(struct fuse_conn *fc, struct fuse_backing *fb)
 	id = idr_alloc_cyclic(&fc->backing_files_map, fb, 1, 0, GFP_ATOMIC);
 	spin_unlock(&fc->lock);
 	idr_preload_end();
+	if (id > 0)
+		fb->backing_id = id;
 
 	WARN_ON_ONCE(id == 0);
 	return id;
@@ -61,81 +64,131 @@ static struct fuse_backing *fuse_backing_id_remove(struct fuse_conn *fc,
 	spin_lock(&fc->lock);
 	fb = idr_remove(&fc->backing_files_map, id);
 	spin_unlock(&fc->lock);
+	if (fb)
+		fb->backing_id = 0;
 
 	return fb;
 }
 
-static int fuse_backing_id_free(int id, void *p, void *data)
-{
-	struct fuse_backing *fb = p;
+static const struct rhashtable_params fuse_backing_prm = {
+	.head_offset = offsetof(struct fuse_backing, hash_node),
+	.key_offset = offsetof(struct fuse_backing, backing_id),
+	.key_len = sizeof_field(struct fuse_backing, backing_id),
+};
 
-	WARN_ON_ONCE(refcount_read(&fb->count) != 1);
-	fuse_backing_free(fb);
-	return 0;
+static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
+{
+	return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
 }
 
-void fuse_backing_files_free(struct fuse_conn *fc)
+int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)
 {
-	idr_for_each(&fc->backing_files_map, fuse_backing_id_free, NULL);
-	idr_destroy(&fc->backing_files_map);
+	struct fuse_backing *fb;
+	int err;
+
+	if (!fc->backing_id_64)
+		return -EINVAL;
+
+	scoped_guard(spinlock, &fc->lock) {
+		fb = rhashtable_lookup_fast(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
+		if (!fb)
+			return -ENOENT;
+
+		err = rhashtable_remove_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
+		WARN_ON(err);
+	}
+	fb->backing_id = 0;
+	fuse_backing_put(fb);
+
+	return 0;
 }
 
-int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
+static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
 {
-	struct file *file;
+	struct fuse_backing *fb;
 	struct super_block *backing_sb;
-	struct fuse_backing *fb = NULL;
-	int res;
-
-	pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
+	struct file *file;
 
 	/* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
-	res = -EPERM;
 	if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
-		goto out;
+		return ERR_PTR(-EPERM);
 
-	res = -EINVAL;
-	if (map->flags || map->padding)
-		goto out;
+	CLASS(fd_raw, f)(fd);
+	if (fd_empty(f))
+		return ERR_PTR(-EBADF);
 
-	file = fget_raw(map->fd);
-	res = -EBADF;
-	if (!file)
-		goto out;
+	file = fd_file(f);
 
 	/* read/write/splice/mmap passthrough only relevant for regular files */
-	res = d_is_dir(file->f_path.dentry) ? -EISDIR : -EINVAL;
 	if (!d_is_reg(file->f_path.dentry))
-		goto out_fput;
+		return d_is_dir(file->f_path.dentry) ? ERR_PTR(-EISDIR) : ERR_PTR(-EINVAL);
 
 	backing_sb = file_inode(file)->i_sb;
-	res = -ELOOP;
 	if (backing_sb->s_stack_depth >= fc->max_stack_depth)
-		goto out_fput;
+		return ERR_PTR(-ELOOP);
 
 	fb = kmalloc_obj(struct fuse_backing);
-	res = -ENOMEM;
 	if (!fb)
-		goto out_fput;
+		return ERR_PTR(-ENOMEM);
 
-	fb->file = file;
+	fb->file = get_file(file);
 	fb->cred = get_current_cred();
 	refcount_set(&fb->count, 1);
 
-	res = fuse_backing_id_alloc(fc, fb);
-	if (res < 0) {
-		fuse_backing_free(fb);
-		fb = NULL;
+	return fb;
+}
+
+int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map)
+{
+	struct fuse_backing *fb;
+	int res;
+
+	res = -EINVAL;
+	if (map->padding || map->spare[0] || map->spare[1])
+		goto out;
+
+	if (!fc->backing_id_64)
+		goto out;
+
+	fb = fuse_backing_new(fc, map->fd);
+	res = PTR_ERR(fb);
+	if (!IS_ERR(fb)) {
+		fb->backing_id = map->backing_id;
+		res = fuse_backing_add_64(fc, fb);
+		if (res < 0)
+			fuse_backing_free(fb);
 	}
+out:
+	return res;
+}
+
+int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
+{
+	struct fuse_backing *fb = NULL;
+	int res;
+
+	pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
 
+	res = -EINVAL;
+	if (map->flags || map->padding)
+		goto out;
+
+	if (fc->backing_id_64)
+		goto out;
+
+	fb = fuse_backing_new(fc, map->fd);
+	res = PTR_ERR(fb);
+	if (!IS_ERR(fb)) {
+		res = fuse_backing_id_alloc(fc, fb);
+		if (res < 0) {
+			fuse_backing_free(fb);
+			fb = NULL;
+		}
+	}
 out:
 	pr_debug("%s: fb=0x%p, ret=%i\n", __func__, fb, res);
 
 	return res;
-
-out_fput:
-	fput(file);
-	goto out;
 }
 
 int fuse_backing_close(struct fuse_conn *fc, int backing_id)
@@ -145,6 +198,9 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
 
 	pr_debug("%s: backing_id=%d\n", __func__, backing_id);
 
+	if (fc->backing_id_64)
+		return -EINVAL;
+
 	/* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
 	err = -EPERM;
 	if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
@@ -167,14 +223,48 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
 	return err;
 }
 
-struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id)
+struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id)
 {
 	struct fuse_backing *fb;
 
-	rcu_read_lock();
-	fb = idr_find(&fc->backing_files_map, backing_id);
-	fb = fuse_backing_get(fb);
-	rcu_read_unlock();
+	guard(rcu)();
+	if (!fc->backing_id_64)
+		fb = idr_find(&fc->backing_files_map, backing_id);
+	else
+		fb = rhashtable_lookup(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
 
-	return fb;
+	return fuse_backing_get(fb);
+}
+
+static void fuse_backing_check_free(struct fuse_backing *fb)
+{
+	WARN_ON_ONCE(refcount_read(&fb->count) != 1);
+	fuse_backing_free(fb);
+}
+
+static int fuse_backing_idr_free(int id, void *p, void *data)
+{
+	fuse_backing_check_free(p);
+	return 0;
+}
+
+static void fuse_backing_rht_free(void *p, void *data)
+{
+	fuse_backing_check_free(p);
+}
+
+void fuse_backing_files_free(struct fuse_conn *fc)
+{
+	if (fc->backing_id_64) {
+		rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
+	} else {
+		idr_for_each(&fc->backing_files_map, fuse_backing_idr_free, NULL);
+		idr_destroy(&fc->backing_files_map);
+	}
+}
+
+void fuse_backing_files_init_64(struct fuse_conn *fc)
+{
+	rhashtable_init(&fc->backing_64_ht, &fuse_backing_prm);
+	fc->backing_id_64 = true;
 }
diff --git a/fs/fuse/dev.c b/fs/fuse/dev.c
index 2d7ee5498f1c..8b63a08a26b6 100644
--- a/fs/fuse/dev.c
+++ b/fs/fuse/dev.c
@@ -2339,6 +2339,24 @@ static long fuse_dev_ioctl_backing_open(struct file *file,
 	return fuse_backing_open(fud->chan->conn, &map);
 }
 
+static long fuse_dev_ioctl_backing_create(struct file *file,
+					  struct fuse_backing_create_in __user *argp)
+{
+	struct fuse_dev *fud = fuse_get_dev(file);
+	struct fuse_backing_create_in map;
+
+	if (IS_ERR(fud))
+		return PTR_ERR(fud);
+
+	if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
+		return -EOPNOTSUPP;
+
+	if (copy_from_user(&map, argp, sizeof(map)))
+		return -EFAULT;
+
+	return fuse_backing_open_64(fud->chan->conn, &map);
+}
+
 static long fuse_dev_ioctl_backing_close(struct file *file, __u32 __user *argp)
 {
 	struct fuse_dev *fud = fuse_get_dev(file);
@@ -2379,6 +2397,9 @@ static long fuse_dev_ioctl(struct file *file, unsigned int cmd,
 	case FUSE_DEV_IOC_BACKING_OPEN:
 		return fuse_dev_ioctl_backing_open(file, argp);
 
+	case FUSE_DEV_IOC_BACKING_CREATE:
+		return fuse_dev_ioctl_backing_create(file, argp);
+
 	case FUSE_DEV_IOC_BACKING_CLOSE:
 		return fuse_dev_ioctl_backing_close(file, argp);
 
diff --git a/fs/fuse/dev.h b/fs/fuse/dev.h
index f6c47ae0395b..dbffd5bed2f5 100644
--- a/fs/fuse/dev.h
+++ b/fs/fuse/dev.h
@@ -14,6 +14,7 @@ struct fuse_dev;
 struct fuse_args;
 struct fuse_copy_state;
 struct fuse_backing_map;
+struct fuse_backing_create_in;
 struct file;
 struct folio;
 enum fuse_notify_code;
@@ -87,6 +88,8 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
 
 int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map);
 int fuse_backing_close(struct fuse_conn *fc, int backing_id);
+int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map);
+int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id);
 
 int fuse_copy_one(struct fuse_copy_state *cs, void *val, unsigned size);
 int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop,
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 3dd4dff24c50..003ef3c35d7a 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -32,6 +32,7 @@
 #include <linux/pid_namespace.h>
 #include <linux/refcount.h>
 #include <linux/user_namespace.h>
+#include <linux/rhashtable-types.h>
 
 /** Default max number of pages that can be used in a single read request */
 #define FUSE_DEFAULT_MAX_PAGES_PER_REQ 32
@@ -92,7 +93,8 @@ struct fuse_submount_lookup {
 struct fuse_backing {
 	struct file *file;
 	const struct cred *cred;
-
+	u64 backing_id;
+	struct rhash_head hash_node;
 	/* refcount */
 	refcount_t count;
 	struct rcu_head rcu;
@@ -689,6 +691,9 @@ struct fuse_conn {
 	/** @init_security: Initialize security xattrs when creating a new inode */
 	unsigned int init_security:1;
 
+	/** Backing ID is 64 bit and allocated by the server */
+	bool backing_id_64:1;
+
 	/**
 	 * @create_supp_group: Add supplementary group info when creating
 	 * a new inode
@@ -770,8 +775,14 @@ struct fuse_conn {
 	struct fuse_sync_bucket __rcu *curr_bucket;
 
 #ifdef CONFIG_FUSE_PASSTHROUGH
-	/** @backing_files_map: IDR for backing files ids */
-	struct idr backing_files_map;
+	/* Selected by backing_id_64 */
+	union {
+		/** @backing_files_map: IDR for backing files ids */
+		struct idr backing_files_map;
+
+		/** @backing_64_ht: 64 bit ID lookup hash table */
+		struct rhashtable backing_64_ht;
+	};
 #endif
 };
 
@@ -1270,7 +1281,7 @@ void fuse_file_release(struct inode *inode, struct fuse_file *ff,
 /* backing.c */
 #ifdef CONFIG_FUSE_PASSTHROUGH
 void fuse_backing_put(struct fuse_backing *fb);
-struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id);
+
 #else
 
 static inline void fuse_backing_put(struct fuse_backing *fb)
@@ -1278,7 +1289,9 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
 }
 #endif
 
+struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
 void fuse_backing_files_init(struct fuse_conn *fc);
+void fuse_backing_files_init_64(struct fuse_conn *fc);
 void fuse_backing_files_free(struct fuse_conn *fc);
 
 /* passthrough.c */
diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
index cbb10e19e7e8..bb76bddfc241 100644
--- a/fs/fuse/inode.c
+++ b/fs/fuse/inode.c
@@ -1408,13 +1408,15 @@ static void process_init_reply(struct fuse_args *args, int error)
 			 * them together.
 			 */
 			if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) &&
-			    (flags & FUSE_PASSTHROUGH) &&
+			    (flags & (FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2)) &&
 			    arg->max_stack_depth > 0 &&
 			    arg->max_stack_depth <= FILESYSTEM_MAX_STACK_DEPTH &&
 			    !(flags & FUSE_WRITEBACK_CACHE))  {
 				fc->passthrough = 1;
 				fc->max_stack_depth = arg->max_stack_depth;
 				fm->sb->s_stack_depth = arg->max_stack_depth;
+				if (flags & FUSE_PASSTHROUGH_V2)
+					fuse_backing_files_init_64(fc);
 			}
 			if (flags & FUSE_NO_EXPORT_SUPPORT)
 				fm->sb->s_export_op = &fuse_export_fid_operations;
@@ -1500,7 +1502,7 @@ static struct fuse_init_args *fuse_new_init(struct fuse_mount *fm)
 	if (fm->fc->auto_submounts)
 		flags |= FUSE_SUBMOUNTS;
 	if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
-		flags |= FUSE_PASSTHROUGH;
+		flags |= FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2;
 	/* Only offered to sufficiently privileged servers; see
 	 * fuse_syncfs_enable().
 	 */
diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
index c0428e03d138..93e916a16ac9 100644
--- a/fs/fuse/notify.c
+++ b/fs/fuse/notify.c
@@ -409,6 +409,31 @@ static int fuse_notify_prune(struct fuse_conn *fc, unsigned int size,
 	return 0;
 }
 
+static int fuse_notify_backing_remove(struct fuse_conn *fc, unsigned int size,
+				     struct fuse_copy_state *cs)
+{
+	struct fuse_notify_backing_remove_out outarg;
+	int err;
+
+	if (size != sizeof(outarg))
+		return -EINVAL;
+
+	err = fuse_copy_one(cs, &outarg, sizeof(outarg));
+	if (err)
+		return err;
+
+	if (outarg.reserved)
+		return -EINVAL;
+
+	if (!fc->backing_id_64)
+		return -EINVAL;
+
+	if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
+		return -EOPNOTSUPP;
+
+	return fuse_backing_close_64(fc, outarg.backing_id);
+}
+
 int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
 		unsigned int size, struct fuse_copy_state *cs)
 {
@@ -440,6 +465,9 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
 	case FUSE_NOTIFY_PRUNE:
 		return fuse_notify_prune(fc, size, cs);
 
+	case FUSE_NOTIFY_BACKING_REMOVE:
+		return fuse_notify_backing_remove(fc, size, cs);
+
 	default:
 		return -EINVAL;
 	}
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index 489838462d1e..313c8d7ffc09 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -160,6 +160,9 @@ struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
 	struct fuse_backing *fb = NULL;
 	struct file *backing_file;
 
+	if (fc->backing_id_64)
+		return ERR_PTR(fuse_EIO("incompatible backing version"));
+
 	if (backing_id <= 0)
 		return ERR_PTR(fuse_EIO("invalid backing_id"));
 
diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
index 10a7f31c4bdf..c88d7094af44 100644
--- a/include/uapi/linux/fuse.h
+++ b/include/uapi/linux/fuse.h
@@ -251,6 +251,9 @@
  *
  *  7.47
  *  - add FUSE_HAS_SYNCFS opt-in flag for privileged userspace servers
+ *  - add FUSE_PASSTHROUGH_V2
+ *  - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
+ *  - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
  */
 
 #ifndef _LINUX_FUSE_H
@@ -473,6 +476,7 @@ struct fuse_file_lock {
  *		with CAP_SYS_ADMIN in the initial user namespace (the same
  *		privilege that mounting virtiofs or fuseblk requires).
  *		Insufficiently privileged servers ignore it.
+ * FUSE_PASSTHROUGH_V2: use 64 bit server allocated backing ID
  */
 #define FUSE_ASYNC_READ		(1 << 0)
 #define FUSE_POSIX_LOCKS	(1 << 1)
@@ -522,6 +526,7 @@ struct fuse_file_lock {
 #define FUSE_REQUEST_TIMEOUT	(1ULL << 42)
 #define FUSE_HAS_IO_URING_BUFPOOL (1ULL << 43)
 #define FUSE_HAS_SYNCFS		(1ULL << 44)
+#define FUSE_PASSTHROUGH_V2	(1ULL << 45)
 
 /**
  * CUSE INIT request/reply flags
@@ -709,6 +714,7 @@ enum fuse_notify_code {
 	FUSE_NOTIFY_RESEND = 7,
 	FUSE_NOTIFY_INC_EPOCH = 8,
 	FUSE_NOTIFY_PRUNE = 9,
+	FUSE_NOTIFY_BACKING_REMOVE = 10,
 };
 
 /* The read buffer is required to be at least 8k, but may be much larger */
@@ -1159,13 +1165,20 @@ struct fuse_backing_map {
 	uint64_t	padding;
 };
 
+struct fuse_backing_create_in {
+	int32_t		fd;
+	uint32_t	padding;
+	uint64_t	backing_id;
+	uint64_t	spare[2];
+};
+
 /* Device ioctls: */
 #define FUSE_DEV_IOC_MAGIC		229
 #define FUSE_DEV_IOC_CLONE		_IOR(FUSE_DEV_IOC_MAGIC, 0, uint32_t)
-#define FUSE_DEV_IOC_BACKING_OPEN	_IOW(FUSE_DEV_IOC_MAGIC, 1, \
-					     struct fuse_backing_map)
+#define FUSE_DEV_IOC_BACKING_OPEN	_IOW(FUSE_DEV_IOC_MAGIC, 1, struct fuse_backing_map)
 #define FUSE_DEV_IOC_BACKING_CLOSE	_IOW(FUSE_DEV_IOC_MAGIC, 2, uint32_t)
 #define FUSE_DEV_IOC_SYNC_INIT		_IO(FUSE_DEV_IOC_MAGIC, 3)
+#define FUSE_DEV_IOC_BACKING_CREATE	_IOW(FUSE_DEV_IOC_MAGIC, 4, struct fuse_backing_create_in)
 
 struct fuse_lseek_in {
 	uint64_t	fh;
@@ -1193,6 +1206,11 @@ struct fuse_copy_file_range_out {
 	uint64_t	bytes_copied;
 };
 
+struct fuse_notify_backing_remove_out {
+	uint64_t	backing_id;
+	uint64_t	reserved;
+};
+
 #define FUSE_SETUPMAPPING_FLAG_WRITE (1ull << 0)
 #define FUSE_SETUPMAPPING_FLAG_READ (1ull << 1)
 struct fuse_setupmapping_in {
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 4/8] fuse: support opening 64 bit backing ID
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
                   ` (2 preceding siblings ...)
  2026-10-01 15:07 ` [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
@ 2026-10-01 15:07 ` Miklos Szeredi
  2026-10-01 15:22   ` sashiko-bot
  2026-10-01 17:09   ` Amir Goldstein
  2026-10-01 15:07 ` [PATCH v2 5/8] fuse: add support for opening dax device as backing Miklos Szeredi
                   ` (4 subsequent siblings)
  8 siblings, 2 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

Add backing_id_64 field to fuse_open_out to allow opening files with
64-bit server-allocated backing IDs.

When the server sets FUSE_BACKING_ID_64 in open_out.open_flags, the
kernel reads the backing ID from open_out.backing_id_64 instead of
open_out.backing_id.

The open reply is now variable-length for backward compatibility with
servers that don't send the new backing_id_64 field.

Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
 fs/fuse/dir.c             |  3 ++-
 fs/fuse/file.c            |  6 +++++-
 fs/fuse/fuse_i.h          |  2 +-
 fs/fuse/iomode.c          | 34 ++++++++++++++++++++++++++++------
 fs/fuse/passthrough.c     | 28 ++++++----------------------
 include/uapi/linux/fuse.h |  2 ++
 6 files changed, 44 insertions(+), 31 deletions(-)

diff --git a/fs/fuse/dir.c b/fs/fuse/dir.c
index f48fafccce4b..874ec7cfffb1 100644
--- a/fs/fuse/dir.c
+++ b/fs/fuse/dir.c
@@ -876,6 +876,7 @@ static int fuse_create_open(struct mnt_idmap *idmap, struct inode *dir,
 	args.out_args[0].value = &outentry;
 	/* Store outarg for fuse_finish_open() */
 	outopenp = &ff->args->open_outarg;
+	args.out_argvar = true; /* compat */
 	args.out_args[1].size = sizeof(*outopenp);
 	args.out_args[1].value = outopenp;
 
@@ -885,7 +886,7 @@ static int fuse_create_open(struct mnt_idmap *idmap, struct inode *dir,
 
 	err = fuse_simple_idmap_request(idmap, fm, &args);
 	free_ext_value(&args);
-	if (err)
+	if (err < 0)
 		goto out_free_ff;
 
 	err = -EIO;
diff --git a/fs/fuse/file.c b/fs/fuse/file.c
index 92976906ab05..1036c06fe35a 100644
--- a/fs/fuse/file.c
+++ b/fs/fuse/file.c
@@ -28,6 +28,7 @@ static int fuse_send_open(struct fuse_mount *fm, u64 nodeid,
 {
 	struct fuse_open_in inarg;
 	FUSE_ARGS(args);
+	int res;
 
 	memset(&inarg, 0, sizeof(inarg));
 	inarg.flags = open_flags & ~(O_CREAT | O_EXCL | O_NOCTTY);
@@ -45,10 +46,13 @@ static int fuse_send_open(struct fuse_mount *fm, u64 nodeid,
 	args.in_args[0].size = sizeof(inarg);
 	args.in_args[0].value = &inarg;
 	args.out_numargs = 1;
+	args.out_argvar = true; /* compat */
 	args.out_args[0].size = sizeof(*outargp);
 	args.out_args[0].value = outargp;
 
-	return fuse_simple_request(fm, &args);
+	res = fuse_simple_request(fm, &args);
+
+	return res < 0 ? res : 0;
 }
 
 struct fuse_file *fuse_file_alloc(struct fuse_mount *fm, bool release)
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 003ef3c35d7a..59afebc3dd2b 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -1314,7 +1314,7 @@ static inline struct fuse_backing *fuse_inode_backing_set(struct fuse_inode *fi,
 #endif
 }
 
-struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id);
+int fuse_passthrough_open(struct file *file, struct fuse_backing *fb);
 void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb);
 
 static inline bool fuse_is_passthrough(struct fuse_file *ff)
diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
index 1a10bc65fb31..811ea603b776 100644
--- a/fs/fuse/iomode.c
+++ b/fs/fuse/iomode.c
@@ -160,7 +160,9 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
 {
 	struct fuse_file *ff = file->private_data;
 	struct fuse_conn *fc = get_fuse_conn(inode);
+	struct fuse_open_out *outarg = &ff->args->open_outarg;
 	struct fuse_backing *fb;
+	u64 backing_id;
 	int err;
 
 	/* Check allowed conditions for file open in passthrough mode */
@@ -170,18 +172,38 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
 	if (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK)
 		return fuse_EIO("conflicting open flags");
 
-	fb = fuse_passthrough_open(file, ff->args->open_outarg.backing_id);
-	if (IS_ERR(fb))
-		return PTR_ERR(fb);
+	if (!fc->backing_id_64) {
+		if (outarg->backing_id_64 != 0)
+			return fuse_EIO("64 bit backing ID set");
+
+		backing_id = outarg->backing_id;
+		if (backing_id <= 0)
+			return fuse_EIO("invalid backing ID");
+	} else {
+		if (outarg->backing_id != 0)
+			return fuse_EIO("32 bit backing ID set");
+
+		backing_id = outarg->backing_id_64;
+	}
+	fb = fuse_backing_lookup(fc, backing_id);
+	if (!fb)
+		return fuse_EIO("backing not found");
+
+	err = fuse_passthrough_open(file, fb);
+	if (err)
+		goto backing_put;
 
 	/* First passthrough file open denies caching inode io mode */
 	err = fuse_file_uncached_io_open(inode, ff, fb);
-	if (!err)
-		return 0;
+	if (err)
+		goto passthrough_release;
+
+	return 0;
 
+passthrough_release:
 	fuse_passthrough_release(ff, fb);
+backing_put:
 	fuse_backing_put(fb);
-
 	return err;
 }
 
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index 313c8d7ffc09..4894842ad6d0 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -150,41 +150,25 @@ ssize_t fuse_passthrough_mmap(struct file *file, struct vm_area_struct *vma)
 
 /*
  * Setup passthrough to a backing file.
- *
- * Returns an fb object with elevated refcount to be stored in fuse inode.
  */
-struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
+int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
 {
 	struct fuse_file *ff = file->private_data;
-	struct fuse_conn *fc = ff->fm->fc;
-	struct fuse_backing *fb = NULL;
 	struct file *backing_file;
 
-	if (fc->backing_id_64)
-		return ERR_PTR(fuse_EIO("incompatible backing version"));
-
-	if (backing_id <= 0)
-		return ERR_PTR(fuse_EIO("invalid backing_id"));
-
-	fb = fuse_backing_lookup(fc, backing_id);
-	if (!fb)
-		return ERR_PTR(fuse_EIO("backing not found"));
-
 	/* Allocate backing file per fuse file to store fuse path */
 	backing_file = backing_file_open(file, file->f_flags,
 					 &fb->file->f_path, fb->cred);
-	if (IS_ERR(backing_file)) {
-		fuse_backing_put(fb);
-		return ERR_PTR(fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file)));
-	}
+	if (IS_ERR(backing_file))
+		return fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file));
 
 	ff->passthrough = backing_file;
 	ff->cred = get_cred(fb->cred);
 
-	pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p\n", __func__,
-		 backing_id, fb, ff->passthrough);
+	pr_debug("%s: backing_id=%llu, fb=0x%p, backing_file=0x%p\n", __func__,
+		 fb->backing_id, fb, ff->passthrough);
 
-	return fb;
+	return 0;
 }
 
 void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb)
diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
index c88d7094af44..60fb2add5e12 100644
--- a/include/uapi/linux/fuse.h
+++ b/include/uapi/linux/fuse.h
@@ -254,6 +254,7 @@
  *  - add FUSE_PASSTHROUGH_V2
  *  - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
  *  - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
+ *  - add backing_id_64 to fuse_open_out
  */
 
 #ifndef _LINUX_FUSE_H
@@ -841,6 +842,7 @@ struct fuse_open_out {
 	uint64_t	fh;
 	uint32_t	open_flags;
 	int32_t		backing_id;
+	uint64_t	backing_id_64;
 };
 
 struct fuse_release_in {
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 5/8] fuse: add support for opening dax device as backing
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
                   ` (3 preceding siblings ...)
  2026-10-01 15:07 ` [PATCH v2 4/8] fuse: support opening 64 bit " Miklos Szeredi
@ 2026-10-01 15:07 ` Miklos Szeredi
  2026-10-01 15:30   ` sashiko-bot
  2026-10-01 16:07   ` Amir Goldstein
  2026-10-01 15:07 ` [PATCH v2 6/8] fuse: add extent map data structure Miklos Szeredi
                   ` (3 subsequent siblings)
  8 siblings, 2 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

This is only possible with FUSE_PASSTHROUGH_V2 enabled.

Mark the inode with S_DAX if FUSE_LOOKUP returns with FUSE_ATTR_DAX set.

This patch does not yet provide a way actually use the dax dev backing:
when such a backing ID is provided in reply to FUSE_OPEN with
FOPEN_PASSTHROUGH flag set, an error will be returned.

Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
 fs/fuse/backing.c     | 121 ++++++++++++++++++++++++++++++------------
 fs/fuse/file.c        |   2 +-
 fs/fuse/fuse_i.h      |  28 ++++++++--
 fs/fuse/inode.c       |  11 +++-
 fs/fuse/passthrough.c |   6 ++-
 5 files changed, 126 insertions(+), 42 deletions(-)

diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index 3c879df7989c..c852f0498961 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -9,6 +9,7 @@
 #include "fuse_i.h"
 
 #include <linux/file.h>
+#include <linux/dax.h>
 #include <linux/rhashtable.h>
 
 static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb)
@@ -22,9 +23,16 @@ static void fuse_backing_free(struct fuse_backing *fb)
 {
 	pr_debug("%s: fb=0x%p\n", __func__, fb);
 
-	if (fb->file)
-		fput(fb->file);
-	put_cred(fb->cred);
+	switch (fb->type) {
+	case FUSE_BACKING_PATH:
+		path_put(&fb->path);
+		put_cred(fb->cred);
+		break;
+
+	case FUSE_BACKING_DAXDEV:
+		fs_put_dax(fb->dax_dev, fb);
+		break;
+	}
 	kfree_rcu(fb, rcu);
 }
 
@@ -103,39 +111,83 @@ int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)
 	return 0;
 }
 
-static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
+static int fuse_dax_notify_failure(struct dax_device *daxdev, u64 offset, u64 len, int mf_flags)
 {
-	struct fuse_backing *fb;
-	struct super_block *backing_sb;
-	struct file *file;
+	struct fuse_backing *fb = dax_holder(daxdev);
 
-	/* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
-	if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
-		return ERR_PTR(-EPERM);
+	fb->dax_error = true;
 
-	CLASS(fd_raw, f)(fd);
-	if (fd_empty(f))
-		return ERR_PTR(-EBADF);
+	return 0;
+}
 
-	file = fd_file(f);
+static const struct dax_holder_operations fuse_dax_holder_ops = {
+	.notify_failure		= fuse_dax_notify_failure,
+};
+
+static int fuse_backing_open_file(struct fuse_conn *fc, struct fuse_backing *fb, struct file *file)
+{
+	struct inode *inode = file_inode(file);
+	struct dax_device *daxdev;
+	int err;
 
-	/* read/write/splice/mmap passthrough only relevant for regular files */
-	if (!d_is_reg(file->f_path.dentry))
-		return d_is_dir(file->f_path.dentry) ? ERR_PTR(-EISDIR) : ERR_PTR(-EINVAL);
+	switch (inode->i_mode & S_IFMT) {
+	case S_IFREG:
+		/* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
+		if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
+			return -EPERM;
+
+		if (inode->i_sb->s_stack_depth >= fc->max_stack_depth)
+			return -ELOOP;
+
+		fb->type = FUSE_BACKING_PATH;
+		fb->path = file->f_path;
+		path_get(&fb->path);
+		fb->cred = get_current_cred();
+		return 0;
+
+	case S_IFCHR:
+		daxdev = dax_dev_find(inode->i_rdev);
+		if (!daxdev)
+			return -EINVAL;
+
+		err = -EPERM;
+		if (capable(CAP_SYS_RAWIO)) {
+			err = fs_dax_get(daxdev, fb, &fuse_dax_holder_ops);
+			if (!err) {
+				fb->type = FUSE_BACKING_DAXDEV;
+				fb->dax_dev = daxdev;
+			}
+		}
+		put_dax(daxdev);
+		return err;
 
-	backing_sb = file_inode(file)->i_sb;
-	if (backing_sb->s_stack_depth >= fc->max_stack_depth)
-		return ERR_PTR(-ELOOP);
+	case S_IFDIR:
+		return -EISDIR;
+
+	default:
+		return -EINVAL;
+	}
+}
+
+static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
+{
+	struct fuse_backing *fb __free(kfree) = kmalloc_obj(*fb);
+	int err;
 
-	fb = kmalloc_obj(struct fuse_backing);
 	if (!fb)
 		return ERR_PTR(-ENOMEM);
 
-	fb->file = get_file(file);
-	fb->cred = get_current_cred();
+	CLASS(fd_raw, f)(fd);
+	if (fd_empty(f))
+		return ERR_PTR(-EBADF);
+
+	err = fuse_backing_open_file(fc, fb, fd_file(f));
+	if (err)
+		return ERR_PTR(err);
+
 	refcount_set(&fb->count, 1);
 
-	return fb;
+	return_ptr(fb);
 }
 
 int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map)
@@ -143,22 +195,21 @@ int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *ma
 	struct fuse_backing *fb;
 	int res;
 
-	res = -EINVAL;
 	if (map->padding || map->spare[0] || map->spare[1])
-		goto out;
+		return -EINVAL;
 
 	if (!fc->backing_id_64)
-		goto out;
+		return -EINVAL;
 
 	fb = fuse_backing_new(fc, map->fd);
-	res = PTR_ERR(fb);
-	if (!IS_ERR(fb)) {
-		fb->backing_id = map->backing_id;
-		res = fuse_backing_add_64(fc, fb);
-		if (res < 0)
-			fuse_backing_free(fb);
-	}
-out:
+	if (IS_ERR(fb))
+		return PTR_ERR(fb);
+
+	fb->backing_id = map->backing_id;
+	res = fuse_backing_add_64(fc, fb);
+	if (res < 0)
+		fuse_backing_free(fb);
+
 	return res;
 }
 
diff --git a/fs/fuse/file.c b/fs/fuse/file.c
index 1036c06fe35a..fb8e0a9fa4f6 100644
--- a/fs/fuse/file.c
+++ b/fs/fuse/file.c
@@ -297,7 +297,7 @@ static int fuse_open(struct inode *inode, struct file *file)
 	if (!err) {
 		if (is_truncate)
 			truncate_pagecache(inode, 0);
-		else if (!(ff->open_flags & FOPEN_KEEP_CACHE))
+		else if (!(ff->open_flags & FOPEN_KEEP_CACHE) && !IS_DAX(inode))
 			invalidate_inode_pages2(inode->i_mapping);
 	}
 out_unlock:
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 59afebc3dd2b..3e7ffb9bb2c3 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -89,10 +89,25 @@ struct fuse_submount_lookup {
 	struct fuse_forget_link *forget;
 };
 
+enum fuse_backing_type {
+	FUSE_BACKING_PATH,
+	FUSE_BACKING_DAXDEV,
+};
+
 /* Container for data related to mapping to backing file */
 struct fuse_backing {
-	struct file *file;
-	const struct cred *cred;
+	enum fuse_backing_type type;
+
+	union {
+		struct {
+			struct path path;
+			const struct cred *cred;
+		};
+		struct {
+			struct dax_device *dax_dev;
+			bool dax_error;
+		};
+	};
 	u64 backing_id;
 	struct rhash_head hash_node;
 	/* refcount */
@@ -1239,7 +1254,14 @@ void fuse_free_conn(struct fuse_conn *fc);
 
 /* dax.c */
 
-#define FUSE_IS_VDAX(inode) (IS_ENABLED(CONFIG_FUSE_VDAX) && IS_DAX(inode))
+static inline bool FUSE_IS_VDAX(struct inode *inode)
+{
+#ifdef CONFIG_FUSE_VDAX
+	return get_fuse_inode(inode)->vdax && IS_DAX(inode);
+#else
+	return false;
+#endif
+}
 
 ssize_t fuse_vdax_read_iter(struct kiocb *iocb, struct iov_iter *to);
 ssize_t fuse_vdax_write_iter(struct kiocb *iocb, struct iov_iter *from);
diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
index bb76bddfc241..d42c824fd6af 100644
--- a/fs/fuse/inode.c
+++ b/fs/fuse/inode.c
@@ -148,7 +148,7 @@ static void fuse_evict_inode(struct inode *inode)
 	/* Will write inode on close/munmap and in all other dirtiers */
 	WARN_ON(inode_state_read_once(inode) & I_DIRTY_INODE);
 
-	if (FUSE_IS_VDAX(inode))
+	if (IS_DAX(inode))
 		dax_break_layout_final(inode);
 
 	truncate_inode_pages_final(&inode->i_data);
@@ -403,6 +403,10 @@ static void fuse_init_submount_lookup(struct fuse_submount_lookup *sl,
 	refcount_set(&sl->count, 1);
 }
 
+static const struct address_space_operations fuse_dax_aops = {
+	.dirty_folio	= noop_dirty_folio,
+};
+
 static void fuse_init_inode(struct inode *inode, struct fuse_attr *attr,
 			    struct fuse_conn *fc)
 {
@@ -430,6 +434,11 @@ static void fuse_init_inode(struct inode *inode, struct fuse_attr *attr,
 	 */
 	if (!fc->posix_acl)
 		inode->i_acl = inode->i_default_acl = ACL_DONT_CACHE;
+
+	if ((attr->flags & FUSE_ATTR_DAX) && !fc->vdax) {
+		inode->i_flags |= S_DAX;
+		inode->i_data.a_ops = &fuse_dax_aops;
+	}
 }
 
 static int fuse_inode_eq(struct inode *inode, void *_nodeidp)
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index 4894842ad6d0..e9ab1aea34e2 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -156,9 +156,11 @@ int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
 	struct fuse_file *ff = file->private_data;
 	struct file *backing_file;
 
+	if (fb->type != FUSE_BACKING_PATH)
+		return fuse_EIO("invalid backing type");
+
 	/* Allocate backing file per fuse file to store fuse path */
-	backing_file = backing_file_open(file, file->f_flags,
-					 &fb->file->f_path, fb->cred);
+	backing_file = backing_file_open(file, file->f_flags, &fb->path, fb->cred);
 	if (IS_ERR(backing_file))
 		return fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file));
 
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 6/8] fuse: add extent map data structure
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
                   ` (4 preceding siblings ...)
  2026-10-01 15:07 ` [PATCH v2 5/8] fuse: add support for opening dax device as backing Miklos Szeredi
@ 2026-10-01 15:07 ` Miklos Szeredi
  2026-10-01 15:24   ` sashiko-bot
  2026-10-01 15:07 ` [PATCH v2 7/8] fuse: add extent map I/O support Miklos Szeredi
                   ` (2 subsequent siblings)
  8 siblings, 1 reply; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

Add support for creating and managing extent maps that map file regions to
dax device regions.

Introduce FUSE_NOTIFY_BACKING_MAP for populating extent maps from
userspace.

Extent maps are handled as a new type of backing and are assigned a 64 bit
backing ID by the server.

One extent consists of

 - offset within the containing backing
 - length of extent
 - target backing ID
 - offset into target backing

Extents must be non-overlapping, and can only refer to dax device backings
for now.

At this point only support creating and populating the extent map in one
operation, though the interface is not limited by this and later may be
changed, so that partial population of an extent map would be possible.

Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
 fs/fuse/Makefile          |   2 +-
 fs/fuse/backing.c         |  36 +++++++++--
 fs/fuse/ext_map.c         | 129 ++++++++++++++++++++++++++++++++++++++
 fs/fuse/fuse_i.h          |  14 ++++-
 fs/fuse/notify.c          |  44 +++++++++++++
 include/uapi/linux/fuse.h |  26 ++++++++
 6 files changed, 244 insertions(+), 7 deletions(-)
 create mode 100644 fs/fuse/ext_map.c

diff --git a/fs/fuse/Makefile b/fs/fuse/Makefile
index 5858feafa916..da9e802d7ae1 100644
--- a/fs/fuse/Makefile
+++ b/fs/fuse/Makefile
@@ -15,7 +15,7 @@ fuse-y += dev.o dir.o file.o inode.o control.o xattr.o acl.o readdir.o ioctl.o r
 fuse-y += poll.o notify.o
 fuse-y += iomode.o
 fuse-$(CONFIG_FUSE_VDAX) += dax.o
-fuse-$(CONFIG_FUSE_PASSTHROUGH) += passthrough.o backing.o
+fuse-$(CONFIG_FUSE_PASSTHROUGH) += passthrough.o backing.o ext_map.o
 fuse-$(CONFIG_SYSCTL) += sysctl.o
 fuse-$(CONFIG_FUSE_IO_URING) += dev_uring.o
 
diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index c852f0498961..e954391d696d 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -32,6 +32,10 @@ static void fuse_backing_free(struct fuse_backing *fb)
 	case FUSE_BACKING_DAXDEV:
 		fs_put_dax(fb->dax_dev, fb);
 		break;
+
+	case FUSE_BACKING_EXTMAP:
+		fuse_ext_map_destroy(&fb->extents);
+		break;
 	}
 	kfree_rcu(fb, rcu);
 }
@@ -84,7 +88,7 @@ static const struct rhashtable_params fuse_backing_prm = {
 	.key_len = sizeof_field(struct fuse_backing, backing_id),
 };
 
-static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
+int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
 {
 	return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
 }
@@ -306,12 +310,36 @@ static void fuse_backing_rht_free(void *p, void *data)
 
 void fuse_backing_files_free(struct fuse_conn *fc)
 {
-	if (fc->backing_id_64) {
-		rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
-	} else {
+	struct rhashtable_iter iter;
+	struct fuse_backing *fb;
+
+	if (!fc->backing_id_64) {
 		idr_for_each(&fc->backing_files_map, fuse_backing_idr_free, NULL);
 		idr_destroy(&fc->backing_files_map);
+		return;
 	}
+
+	/*
+	 * extents are referencing other backings, put these refs before
+	 * destroying the backings themselves
+	 */
+	rhashtable_walk_enter(&fc->backing_64_ht, &iter);
+	rhashtable_walk_start(&iter);
+	while ((fb = rhashtable_walk_next(&iter))) {
+		if (IS_ERR(fb)) {
+			if (PTR_ERR(fb) == -EAGAIN)
+				continue;
+			break;
+		}
+		if (fb->type == FUSE_BACKING_EXTMAP) {
+			fuse_ext_map_destroy(&fb->extents);
+			fb->extents.rb_node = NULL;
+		}
+	}
+	rhashtable_walk_stop(&iter);
+	rhashtable_walk_exit(&iter);
+
+	rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
 }
 
 void fuse_backing_files_init_64(struct fuse_conn *fc)
diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
new file mode 100644
index 000000000000..ab01d678d08e
--- /dev/null
+++ b/fs/fuse/ext_map.c
@@ -0,0 +1,129 @@
+// SPDX-License-Identifier: GPL-2.0-only
+
+#include "fuse_i.h"
+#include <linux/rbtree.h>
+#include <linux/pagemap.h>
+
+struct fuse_iext {
+	struct rb_node rb;
+	u64 start;
+	u64 end; /* exclusive */
+	struct fuse_backing *backing;
+	u64 backing_offset;
+};
+
+void fuse_ext_map_destroy(struct rb_root *extents)
+{
+	struct fuse_iext *fie, *tmp;
+
+	rbtree_postorder_for_each_entry_safe(fie, tmp, extents, rb) {
+		fuse_backing_put(fie->backing);
+		kfree(fie);
+	}
+}
+
+static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
+			   struct fuse_extent *ext)
+{
+	struct fuse_iext *new_fie __free(kfree) = kzalloc_obj(*new_fie);
+	struct rb_node *parent = NULL, **link = &extents->rb_node;
+	loff_t end;
+
+	if (!new_fie)
+		return -ENOMEM;
+
+	if (ext->reserved[0] || ext->reserved[1])
+		return fuse_EIO("reserved fields set");
+
+	if (!PAGE_ALIGNED(ext->offset) || !PAGE_ALIGNED(ext->length) || !PAGE_ALIGNED(ext->addr))
+		return fuse_EIO("not page aligned");
+
+	if (!ext->length)
+		return fuse_EIO("zero sized extent");
+
+	if (overflows_type(ext->offset, loff_t) ||
+	    check_add_overflow(ext->offset, ext->length, &end))
+		return fuse_EIO("offset overflow");
+
+	new_fie->start = ext->offset;
+	new_fie->end = end;
+	new_fie->backing_offset = ext->addr;
+
+	while (*link) {
+		struct fuse_iext *fie = rb_entry(*link, typeof(*fie), rb);
+
+		parent = *link;
+		if (new_fie->end <= fie->start)
+			link = &parent->rb_left;
+		else if (new_fie->start >= fie->end)
+			link = &parent->rb_right;
+		else
+			return fuse_EIO("overlap");
+	}
+
+	new_fie->backing = fuse_backing_lookup(fc, ext->backing_id);
+	if (!new_fie->backing)
+		return fuse_EIO("backing not found");
+
+	if (new_fie->backing->type != FUSE_BACKING_DAXDEV) {
+		fuse_backing_put(new_fie->backing);
+		return fuse_EIO("backing is not dax device");
+	}
+
+	rb_link_node(&new_fie->rb, parent, link);
+	rb_insert_color(&no_free_ptr(new_fie)->rb, extents);
+
+	return 0;
+}
+
+bool fuse_ext_map_is_dax(struct fuse_backing *fb)
+{
+	struct fuse_iext *fie;
+
+	if (WARN_ON(RB_EMPTY_ROOT(&fb->extents)))
+		return false;
+
+	fie = rb_entry(fb->extents.rb_node, typeof(*fie), rb);
+	return fie->backing->type == FUSE_BACKING_DAXDEV;
+}
+
+int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_out *arg,
+			  struct fuse_extent *ext)
+{
+	struct fuse_backing *fb;
+	unsigned int i;
+	int err;
+
+	if (!arg->num_extents)
+		return fuse_EIO("no extents");
+
+	if (arg->flags & FUSE_BACKING_MAP_CREATE) {
+		fb = kzalloc_obj(*fb);
+		if (!fb)
+			return -ENOMEM;
+
+		refcount_set(&fb->count, 1);
+		fb->type = FUSE_BACKING_EXTMAP;
+		fb->backing_id = arg->backing_id;
+	} else {
+		/* Adding extents to an existing extmap backing is not yet supported */
+		return -EINVAL;
+	}
+
+	for (i = 0; i < arg->num_extents; i++) {
+		err = fuse_add_extent(fc, &fb->extents, &ext[i]);
+		if (err)
+			goto err_put;
+	}
+
+	err = 0;
+	if (arg->flags & FUSE_BACKING_MAP_CREATE) {
+		err = fuse_backing_add_64(fc, fb);
+		if (!err)
+			return 0;
+	}
+
+err_put:
+	fuse_backing_put(fb);
+	return err;
+}
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 3e7ffb9bb2c3..5a20f3863ee2 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -24,7 +24,7 @@
 #include <linux/backing-dev.h>
 #include <linux/mutex.h>
 #include <linux/rwsem.h>
-#include <linux/rbtree.h>
+#include <linux/rbtree_types.h>
 #include <linux/poll.h>
 #include <linux/workqueue.h>
 #include <linux/kref.h>
@@ -92,6 +92,7 @@ struct fuse_submount_lookup {
 enum fuse_backing_type {
 	FUSE_BACKING_PATH,
 	FUSE_BACKING_DAXDEV,
+	FUSE_BACKING_EXTMAP,
 };
 
 /* Container for data related to mapping to backing file */
@@ -107,6 +108,9 @@ struct fuse_backing {
 			struct dax_device *dax_dev;
 			bool dax_error;
 		};
+		struct {
+			struct rb_root extents;
+		};
 	};
 	u64 backing_id;
 	struct rhash_head hash_node;
@@ -1303,7 +1307,6 @@ void fuse_file_release(struct inode *inode, struct fuse_file *ff,
 /* backing.c */
 #ifdef CONFIG_FUSE_PASSTHROUGH
 void fuse_backing_put(struct fuse_backing *fb);
-
 #else
 
 static inline void fuse_backing_put(struct fuse_backing *fb)
@@ -1312,6 +1315,7 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
 #endif
 
 struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
+int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb);
 void fuse_backing_files_init(struct fuse_conn *fc);
 void fuse_backing_files_init_64(struct fuse_conn *fc);
 void fuse_backing_files_free(struct fuse_conn *fc);
@@ -1362,4 +1366,10 @@ extern void fuse_sysctl_unregister(void);
 #define fuse_sysctl_unregister()	do { } while (0)
 #endif /* CONFIG_SYSCTL */
 
+/* ext_map.c */
+
+void fuse_ext_map_destroy(struct rb_root *extents);
+bool fuse_ext_map_is_dax(struct fuse_backing *fb);
+int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_out *arg,
+			  struct fuse_extent *ext);
 #endif /* _FS_FUSE_I_H */
diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
index 93e916a16ac9..7c427f88bbc5 100644
--- a/fs/fuse/notify.c
+++ b/fs/fuse/notify.c
@@ -434,6 +434,47 @@ static int fuse_notify_backing_remove(struct fuse_conn *fc, unsigned int size,
 	return fuse_backing_close_64(fc, outarg.backing_id);
 }
 
+static int fuse_notify_map(struct fuse_conn *fc, unsigned int size,
+			   struct fuse_copy_state *cs)
+{
+	struct fuse_notify_backing_map_out outarg;
+	struct fuse_extent *ext __free(kvfree) = NULL;
+	int err;
+
+	if (size < sizeof(outarg))
+		return -EINVAL;
+
+	err = fuse_copy_one(cs, &outarg, sizeof(outarg));
+	if (err)
+		return err;
+
+	if (outarg.num_extents > FUSE_MAX_EXTENTS)
+		return -EINVAL;
+
+	size -= sizeof(outarg);
+	if (size != outarg.num_extents * sizeof(*ext))
+		return -EINVAL;
+
+	if (outarg.reserved[0] != 0 || outarg.reserved[1] != 0)
+		return -EINVAL;
+
+	if (outarg.flags & ~FUSE_BACKING_MAP_CREATE)
+		return -EINVAL;
+
+	if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
+		return -EOPNOTSUPP;
+
+	ext = kvmalloc_objs(*ext, outarg.num_extents);
+	if (!ext)
+		return -ENOMEM;
+
+	err = fuse_copy_one(cs, ext, size);
+	if (err)
+		return err;
+
+	return fuse_ext_map_populate(fc, &outarg, ext);
+}
+
 int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
 		unsigned int size, struct fuse_copy_state *cs)
 {
@@ -468,6 +509,9 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
 	case FUSE_NOTIFY_BACKING_REMOVE:
 		return fuse_notify_backing_remove(fc, size, cs);
 
+	case FUSE_NOTIFY_BACKING_MAP:
+		return fuse_notify_map(fc, size, cs);
+
 	default:
 		return -EINVAL;
 	}
diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
index 60fb2add5e12..b17786492f0b 100644
--- a/include/uapi/linux/fuse.h
+++ b/include/uapi/linux/fuse.h
@@ -255,6 +255,7 @@
  *  - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
  *  - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
  *  - add backing_id_64 to fuse_open_out
+ *  - add FUSE_NOTIFY_BACKING_MAP, fuse_notify_backing_map_out, fuse_extent, FUSE_BACKING_MAP_CREATE
  */
 
 #ifndef _LINUX_FUSE_H
@@ -716,6 +717,7 @@ enum fuse_notify_code {
 	FUSE_NOTIFY_INC_EPOCH = 8,
 	FUSE_NOTIFY_PRUNE = 9,
 	FUSE_NOTIFY_BACKING_REMOVE = 10,
+	FUSE_NOTIFY_BACKING_MAP = 11,
 };
 
 /* The read buffer is required to be at least 8k, but may be much larger */
@@ -1213,6 +1215,30 @@ struct fuse_notify_backing_remove_out {
 	uint64_t	reserved;
 };
 
+/**
+ * notify_map flags
+ *
+ * FUSE_BACKING_MAP_CREATE:	create backing with the supplied ID
+ */
+#define FUSE_BACKING_MAP_CREATE	(1 << 0)
+
+struct fuse_notify_backing_map_out {
+	uint64_t	backing_id;
+	uint32_t	num_extents;
+	uint32_t	flags;
+	uint64_t	reserved[2];
+};
+
+#define FUSE_MAX_EXTENTS 1365		/*  (1 << 16) / sizeof(struct fuse_extent) */
+
+struct fuse_extent {
+	uint64_t	offset;		/* offset of extent into parent backing */
+	uint64_t	length;		/* extent length */
+	uint64_t	backing_id;	/* target backing */
+	uint64_t	addr;		/* target offset within backing file/device */
+	uint64_t	reserved[2];
+};
+
 #define FUSE_SETUPMAPPING_FLAG_WRITE (1ull << 0)
 #define FUSE_SETUPMAPPING_FLAG_READ (1ull << 1)
 struct fuse_setupmapping_in {
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 7/8] fuse: add extent map I/O support
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
                   ` (5 preceding siblings ...)
  2026-10-01 15:07 ` [PATCH v2 6/8] fuse: add extent map data structure Miklos Szeredi
@ 2026-10-01 15:07 ` Miklos Szeredi
  2026-10-01 15:27   ` sashiko-bot
  2026-10-01 16:11   ` Amir Goldstein
  2026-10-01 15:07 ` [PATCH v2 8/8] fuse: add support for striped backing Miklos Szeredi
  2026-10-05 23:27 ` [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) John Groves
  8 siblings, 2 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

Wire up read, write, splice and mmap operations for extent-mapped
files through the iomap/DAX infrastructure.

When a passthrough file is opened with an EXTMAP backing, I/O is
dispatched to the dax devices referenced by the extent map.

If an address is not mapped, EIO is returned on the I/O operation.

Reviewed-by: Amir Goldstein <amir73il@gmail.com>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
 fs/fuse/backing.c     |  15 +++++
 fs/fuse/ext_map.c     | 149 ++++++++++++++++++++++++++++++++++++++++++
 fs/fuse/fuse_i.h      |   4 ++
 fs/fuse/notify.c      |   3 +
 fs/fuse/passthrough.c |  29 ++++++--
 5 files changed, 194 insertions(+), 6 deletions(-)

diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
index e954391d696d..2708f140b94f 100644
--- a/fs/fuse/backing.c
+++ b/fs/fuse/backing.c
@@ -278,6 +278,21 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
 	return err;
 }
 
+bool fuse_backing_is_dax(struct fuse_backing *fb)
+{
+	switch (fb->type) {
+	case FUSE_BACKING_PATH:
+		return false;
+	case FUSE_BACKING_DAXDEV:
+		return true;
+	case FUSE_BACKING_EXTMAP:
+		return fuse_ext_map_is_dax(fb);
+	default:
+		WARN_ON(1);
+		return false;
+	}
+}
+
 struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id)
 {
 	struct fuse_backing *fb;
diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
index ab01d678d08e..61cb22e821b6 100644
--- a/fs/fuse/ext_map.c
+++ b/fs/fuse/ext_map.c
@@ -3,6 +3,8 @@
 #include "fuse_i.h"
 #include <linux/rbtree.h>
 #include <linux/pagemap.h>
+#include <linux/iomap.h>
+#include <linux/dax.h>
 
 struct fuse_iext {
 	struct rb_node rb;
@@ -22,6 +24,24 @@ void fuse_ext_map_destroy(struct rb_root *extents)
 	}
 }
 
+static struct fuse_iext *fuse_find_extent(struct rb_root *extents, u64 offset)
+{
+	struct rb_node *node = extents->rb_node;
+
+	while (node) {
+		struct fuse_iext *fie = rb_entry(node, typeof(*fie), rb);
+
+		if (offset < fie->start)
+			node = node->rb_left;
+		else if (offset >= fie->end)
+			node = node->rb_right;
+		else
+			return fie;
+	}
+
+	return NULL;
+}
+
 static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
 			   struct fuse_extent *ext)
 {
@@ -76,6 +96,135 @@ static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
 	return 0;
 }
 
+static int fuse_ext_map_iomap_begin(struct inode *inode, loff_t offset, loff_t length,
+				    unsigned int flags, struct iomap *iomap, struct iomap *srcmap)
+{
+	struct fuse_backing *fb = fuse_inode_backing(get_fuse_inode(inode));
+	struct fuse_iext *fie;
+	loff_t ext_len;
+
+	if (!fb || fb->type != FUSE_BACKING_EXTMAP)
+		return fuse_EIO("missing or wrong type backing");
+
+	fie = fuse_find_extent(&fb->extents, offset);
+	if (!fie)
+		return fuse_EIO("missing mapping");
+
+	if (WARN_ON(fie->backing->type != FUSE_BACKING_DAXDEV))
+		return fuse_EIO("wrong type backing for extent");
+
+	if (fie->backing->dax_error) {
+		fuse_make_bad(inode);
+		return fuse_EIO("dax error");
+	}
+
+	ext_len = fie->end - fie->start;
+
+	iomap->offset = fie->start;
+	iomap->addr = fie->backing_offset;
+	iomap->length = ext_len;
+	iomap->dax_dev = fie->backing->dax_dev;
+	iomap->type = IOMAP_MAPPED;
+	iomap->flags = 0;
+
+	return 0;
+}
+
+static const struct iomap_ops fuse_ext_map_iomap_ops = {
+	.iomap_begin		= fuse_ext_map_iomap_begin,
+};
+
+static vm_fault_t fuse_ext_map_huge_fault(struct vm_fault *vmf, unsigned int order)
+{
+	struct inode *inode = file_inode(vmf->vma->vm_file);
+	bool write_fault = (vmf->flags & FAULT_FLAG_WRITE) && (vmf->vma->vm_flags & VM_SHARED);
+	vm_fault_t ret;
+	unsigned long pfn;
+
+	if (WARN_ON_ONCE(!IS_DAX(inode)))
+		return VM_FAULT_SIGBUS;
+
+	if (write_fault) {
+		sb_start_pagefault(inode->i_sb);
+		file_update_time(vmf->vma->vm_file);
+	}
+
+	filemap_invalidate_lock_shared(inode->i_mapping);
+
+	ret = dax_iomap_fault(vmf, order, &pfn, NULL, &fuse_ext_map_iomap_ops);
+	if (ret & VM_FAULT_NEEDDSYNC)
+		ret = dax_finish_sync_fault(vmf, order, pfn);
+
+	filemap_invalidate_unlock_shared(inode->i_mapping);
+
+	if (write_fault)
+		sb_end_pagefault(inode->i_sb);
+
+	return ret;
+}
+
+static vm_fault_t fuse_ext_map_fault(struct vm_fault *vmf)
+{
+	return fuse_ext_map_huge_fault(vmf, 0);
+}
+
+static const struct vm_operations_struct fuse_ext_map_vm_ops = {
+	.fault		= fuse_ext_map_fault,
+	.huge_fault	= fuse_ext_map_huge_fault,
+	.page_mkwrite	= fuse_ext_map_fault,
+	.pfn_mkwrite	= fuse_ext_map_fault,
+};
+
+static void fuse_rw_clamp(struct kiocb *iocb, struct iov_iter *ubuf)
+{
+	struct inode *inode = iocb->ki_filp->f_mapping->host;
+	loff_t i_size = i_size_read(inode);
+	loff_t max_count = iocb->ki_pos >= i_size ? 0 : i_size - iocb->ki_pos;
+
+	if (iov_iter_count(ubuf) > max_count)
+		iov_iter_truncate(ubuf, max_count);
+}
+
+ssize_t fuse_ext_map_read_iter(struct kiocb *iocb, struct iov_iter *to)
+{
+	ssize_t res;
+
+	fuse_rw_clamp(iocb, to);
+
+	if (!iov_iter_count(to))
+		return 0;
+
+	res = dax_iomap_rw(iocb, to, &fuse_ext_map_iomap_ops);
+
+	file_accessed(iocb->ki_filp);
+	return res;
+}
+
+ssize_t fuse_ext_map_write_iter(struct kiocb *iocb, struct iov_iter *from)
+{
+	ssize_t res;
+
+	fuse_rw_clamp(iocb, from);
+
+	res = generic_write_checks(iocb, from);
+	if (res <= 0)
+		return res;
+
+	res = kiocb_modified(iocb);
+	if (res)
+		return res;
+
+	return dax_iomap_rw(iocb, from, &fuse_ext_map_iomap_ops);
+}
+
+int fuse_ext_map_mmap(struct file *file, struct vm_area_struct *vma)
+{
+	file_accessed(file);
+	vma->vm_ops = &fuse_ext_map_vm_ops;
+	vm_flags_set(vma, VM_HUGEPAGE);
+	return 0;
+}
+
 bool fuse_ext_map_is_dax(struct fuse_backing *fb)
 {
 	struct fuse_iext *fie;
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 5a20f3863ee2..7248588b45c4 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -1316,6 +1316,7 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
 
 struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
 int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb);
+bool fuse_backing_is_dax(struct fuse_backing *fb);
 void fuse_backing_files_init(struct fuse_conn *fc);
 void fuse_backing_files_init_64(struct fuse_conn *fc);
 void fuse_backing_files_free(struct fuse_conn *fc);
@@ -1369,6 +1370,9 @@ extern void fuse_sysctl_unregister(void);
 /* ext_map.c */
 
 void fuse_ext_map_destroy(struct rb_root *extents);
+ssize_t fuse_ext_map_write_iter(struct kiocb *iocb, struct iov_iter *from);
+ssize_t fuse_ext_map_read_iter(struct kiocb *iocb, struct iov_iter *to);
+int fuse_ext_map_mmap(struct file *file, struct vm_area_struct *vma);
 bool fuse_ext_map_is_dax(struct fuse_backing *fb);
 int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_out *arg,
 			  struct fuse_extent *ext);
diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
index 7c427f88bbc5..3bfedd0cdbdd 100644
--- a/fs/fuse/notify.c
+++ b/fs/fuse/notify.c
@@ -448,6 +448,9 @@ static int fuse_notify_map(struct fuse_conn *fc, unsigned int size,
 	if (err)
 		return err;
 
+	if (!fc->backing_id_64)
+		return -EINVAL;
+
 	if (outarg.num_extents > FUSE_MAX_EXTENTS)
 		return -EINVAL;
 
diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
index e9ab1aea34e2..040817ad76e9 100644
--- a/fs/fuse/passthrough.c
+++ b/fs/fuse/passthrough.c
@@ -35,7 +35,6 @@ ssize_t fuse_passthrough_read_iter(struct kiocb *iocb, struct iov_iter *iter)
 	struct fuse_file *ff = file->private_data;
 	struct file *backing_file = fuse_file_passthrough(ff);
 	size_t count = iov_iter_count(iter);
-	ssize_t ret;
 	struct backing_file_ctx ctx = {
 		.cred = ff->cred,
 		.accessed = fuse_file_accessed,
@@ -48,10 +47,10 @@ ssize_t fuse_passthrough_read_iter(struct kiocb *iocb, struct iov_iter *iter)
 	if (!count)
 		return 0;
 
-	ret = backing_file_read_iter(backing_file, iter, iocb, iocb->ki_flags,
-				     &ctx);
+	if (!backing_file)
+		return fuse_ext_map_read_iter(iocb, iter);
 
-	return ret;
+	return backing_file_read_iter(backing_file, iter, iocb, iocb->ki_flags, &ctx);
 }
 
 ssize_t fuse_passthrough_write_iter(struct kiocb *iocb,
@@ -74,10 +73,13 @@ ssize_t fuse_passthrough_write_iter(struct kiocb *iocb,
 	if (!count)
 		return 0;
 
-	inode_lock(inode);
+	guard(rwsem_write)(&inode->i_rwsem);
+
+	if (!backing_file)
+		return fuse_ext_map_write_iter(iocb, iter);
+
 	ret = backing_file_write_iter(backing_file, iter, iocb, iocb->ki_flags,
 				      &ctx);
-	inode_unlock(inode);
 
 	return ret;
 }
@@ -98,6 +100,9 @@ ssize_t fuse_passthrough_splice_read(struct file *in, loff_t *ppos,
 	pr_debug("%s: backing_file=0x%p, pos=%lld, len=%zu, flags=0x%x\n", __func__,
 		 backing_file, *ppos, len, flags);
 
+	if (!backing_file)
+		return copy_splice_read(in, ppos, pipe, len, flags);
+
 	init_sync_kiocb(&iocb, in);
 	iocb.ki_pos = *ppos;
 	ret = backing_file_splice_read(backing_file, &iocb, pipe, len, flags, &ctx);
@@ -123,6 +128,9 @@ ssize_t fuse_passthrough_splice_write(struct pipe_inode_info *pipe,
 	pr_debug("%s: backing_file=0x%p, pos=%lld, len=%zu, flags=0x%x\n", __func__,
 		 backing_file, *ppos, len, flags);
 
+	if (!backing_file)
+		return iter_file_splice_write(pipe, out, ppos, len, flags);
+
 	inode_lock(inode);
 	init_sync_kiocb(&iocb, out);
 	iocb.ki_pos = *ppos;
@@ -145,6 +153,9 @@ ssize_t fuse_passthrough_mmap(struct file *file, struct vm_area_struct *vma)
 	pr_debug("%s: backing_file=0x%p, start=%lu, end=%lu\n", __func__,
 		 backing_file, vma->vm_start, vma->vm_end);
 
+	if (!backing_file)
+		return fuse_ext_map_mmap(file, vma);
+
 	return backing_file_mmap(backing_file, vma, &ctx);
 }
 
@@ -156,6 +167,12 @@ int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
 	struct fuse_file *ff = file->private_data;
 	struct file *backing_file;
 
+	if (fb->type == FUSE_BACKING_EXTMAP) {
+		if (fuse_backing_is_dax(fb) != !!IS_DAX(file_inode(file)))
+			return fuse_EIO("dax mode mismatch");
+		return 0;
+	}
+
 	if (fb->type != FUSE_BACKING_PATH)
 		return fuse_EIO("invalid backing type");
 
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* [PATCH v2 8/8] fuse: add support for striped backing
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
                   ` (6 preceding siblings ...)
  2026-10-01 15:07 ` [PATCH v2 7/8] fuse: add extent map I/O support Miklos Szeredi
@ 2026-10-01 15:07 ` Miklos Szeredi
  2026-10-05 23:27 ` [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) John Groves
  8 siblings, 0 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-01 15:07 UTC (permalink / raw)
  To: fuse-devel
  Cc: John Groves, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

Add the FUSE_BACKING_MAP_CYCLIC flag for FUSE_NOTIFY_BACKING_MAP
notifications to enable creating striped maps.

This is just a set of regular extents that repeats after the end of the
last extent using a striping pattern (i.e. on the target backings the
stripes have a contiguous footprint).

Two additional limits apply compared to non-cyclic extent map population:

 - the size of the extents (chunk size) must be equal

 - the extents must start from zero offset and follow each other without
   gaps.

Reviewed-by: Amir Goldstein <amir73il@gmail.com>
Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
---
 fs/fuse/ext_map.c         | 27 ++++++++++++++++++++++++---
 fs/fuse/fuse_i.h          |  1 +
 fs/fuse/notify.c          |  2 +-
 include/uapi/linux/fuse.h |  5 ++++-
 4 files changed, 30 insertions(+), 5 deletions(-)

diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
index 61cb22e821b6..356a5ec4f5e1 100644
--- a/fs/fuse/ext_map.c
+++ b/fs/fuse/ext_map.c
@@ -101,12 +101,18 @@ static int fuse_ext_map_iomap_begin(struct inode *inode, loff_t offset, loff_t l
 {
 	struct fuse_backing *fb = fuse_inode_backing(get_fuse_inode(inode));
 	struct fuse_iext *fie;
+	u64 ncycle = 0, seq_off = 0, addr, tmp;
 	loff_t ext_len;
 
 	if (!fb || fb->type != FUSE_BACKING_EXTMAP)
 		return fuse_EIO("missing or wrong type backing");
 
-	fie = fuse_find_extent(&fb->extents, offset);
+	if (fb->cycle_length) {
+		ncycle = div64_u64(offset, fb->cycle_length);
+		seq_off = ncycle * fb->cycle_length;
+	}
+
+	fie = fuse_find_extent(&fb->extents, offset - seq_off);
 	if (!fie)
 		return fuse_EIO("missing mapping");
 
@@ -119,9 +125,12 @@ static int fuse_ext_map_iomap_begin(struct inode *inode, loff_t offset, loff_t l
 	}
 
 	ext_len = fie->end - fie->start;
+	addr = fie->backing_offset;
+	if (check_mul_overflow(ncycle, ext_len, &tmp) || check_add_overflow(addr, tmp, &addr))
+		return fuse_EIO("extent address overflow");
 
-	iomap->offset = fie->start;
-	iomap->addr = fie->backing_offset;
+	iomap->offset = fie->start + seq_off;
+	iomap->addr = addr;
 	iomap->length = ext_len;
 	iomap->dax_dev = fie->backing->dax_dev;
 	iomap->type = IOMAP_MAPPED;
@@ -242,6 +251,7 @@ int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_o
 	struct fuse_backing *fb;
 	unsigned int i;
 	int err;
+	u64 chunk_size = 0;
 
 	if (!arg->num_extents)
 		return fuse_EIO("no extents");
@@ -259,7 +269,18 @@ int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_o
 		return -EINVAL;
 	}
 
+	if (arg->flags & FUSE_BACKING_MAP_CYCLIC) {
+		chunk_size = ext[0].length;
+		err = -EINVAL;
+		if (check_mul_overflow(chunk_size, arg->num_extents, &fb->cycle_length))
+			goto err_put;
+	}
+
 	for (i = 0; i < arg->num_extents; i++) {
+		err = -EINVAL;
+		if (chunk_size && (ext[i].offset != chunk_size * i || ext[i].length != chunk_size))
+			goto err_put;
+
 		err = fuse_add_extent(fc, &fb->extents, &ext[i]);
 		if (err)
 			goto err_put;
diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
index 7248588b45c4..83cc1dd67cef 100644
--- a/fs/fuse/fuse_i.h
+++ b/fs/fuse/fuse_i.h
@@ -110,6 +110,7 @@ struct fuse_backing {
 		};
 		struct {
 			struct rb_root extents;
+			u64 cycle_length;
 		};
 	};
 	u64 backing_id;
diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
index 3bfedd0cdbdd..31172de72477 100644
--- a/fs/fuse/notify.c
+++ b/fs/fuse/notify.c
@@ -461,7 +461,7 @@ static int fuse_notify_map(struct fuse_conn *fc, unsigned int size,
 	if (outarg.reserved[0] != 0 || outarg.reserved[1] != 0)
 		return -EINVAL;
 
-	if (outarg.flags & ~FUSE_BACKING_MAP_CREATE)
+	if (outarg.flags & ~(FUSE_BACKING_MAP_CREATE | FUSE_BACKING_MAP_CYCLIC))
 		return -EINVAL;
 
 	if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
index b17786492f0b..f371acb73cda 100644
--- a/include/uapi/linux/fuse.h
+++ b/include/uapi/linux/fuse.h
@@ -255,7 +255,8 @@
  *  - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
  *  - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
  *  - add backing_id_64 to fuse_open_out
- *  - add FUSE_NOTIFY_BACKING_MAP, fuse_notify_backing_map_out, fuse_extent, FUSE_BACKING_MAP_CREATE
+ *  - add FUSE_NOTIFY_BACKING_MAP, fuse_notify_backing_map_out, fuse_extent
+ *  - add FUSE_BACKING_MAP_CREATE, FUSE_BACKING_MAP_CYCLIC
  */
 
 #ifndef _LINUX_FUSE_H
@@ -1219,8 +1220,10 @@ struct fuse_notify_backing_remove_out {
  * notify_map flags
  *
  * FUSE_BACKING_MAP_CREATE:	create backing with the supplied ID
+ * FUSE_BACKING_MAP_CYCLIC:	map repeats after last extent
  */
 #define FUSE_BACKING_MAP_CREATE	(1 << 0)
+#define FUSE_BACKING_MAP_CYCLIC	(1 << 1)
 
 struct fuse_notify_backing_map_out {
 	uint64_t	backing_id;
-- 
2.54.0


^ permalink raw reply related	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message
  2026-10-01 15:07 ` [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
@ 2026-10-01 15:18   ` sashiko-bot
  2026-10-01 16:32   ` Amir Goldstein
  1 sibling, 0 replies; 29+ messages in thread
From: sashiko-bot @ 2026-10-01 15:18 UTC (permalink / raw)
  To: Miklos Szeredi; +Cc: linux-cxl, nvdimm

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] fs/fuse: uninitialized err returned in fuse_file_io_open()
--

commit bfd7cbb3d1064d14b9422fa08119f1cc83817677
Author: Miklos Szeredi <mszeredi@redhat.com>

fuse: add helpers for EIO return value with kernel message

This commit introduces helpers to print a short message to the kernel log
and return an error value when a buggy fuse server triggers an EIO
condition. It also updates various I/O and passthrough open handlers to
utilize these helpers, removing older debug messages and changing some
functions to return boolean values on success.

> diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
> index 79637c09e8839..1a10bc65fb310 100644
> --- a/fs/fuse/iomode.c
> +++ b/fs/fuse/iomode.c
[ ... ]
> @@ -208,11 +201,12 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
>  
>  	/*
>  	 * Server is expected to use FOPEN_PASSTHROUGH for all opens of an inode
> -	 * which is already open for passthrough.
> +	 * which is already open for passthrough.  Using incorrect open mode is
> +	 * a server mistake, which results in user visible failure of open()
> +	 * with EIO error.
>  	 */
> -	err = -EINVAL;
>  	if (fuse_inode_backing(fi) && !(ff->open_flags & FOPEN_PASSTHROUGH))
> -		goto fail;
> +		return fuse_EIO("FOPEN_PASSTHROUGH expected");
[ ... ]
>  	if (ff->open_flags & FOPEN_PASSTHROUGH)
>  		err = fuse_file_passthrough_open(inode, file);
> -	else
> -		err = fuse_file_cached_io_open(inode, ff);
> -	if (err)
> -		goto fail;
> +	else if (!fuse_file_cached_io_open(inode, ff))
> +		err = fuse_EIO("conflicting passthrough open");
>  
> -	return 0;
> -
> -fail:
> -	pr_debug("failed to open file in requested io mode (open_flags=0x%x, err=%i).\n",
> -		 ff->open_flags, err);
> -	/*
> -	 * The file open mode determines the inode io mode.
> -	 * Using incorrect open mode is a server mistake, which results in
> -	 * user visible failure of open() with EIO error.
> -	 */
> -	return -EIO;
> +	return err;
>  }

[Severity: High]
Does this code leak an uninitialized stack value on success?

On the standard cached I/O open path, if the FOPEN_PASSTHROUGH flag is
not set and fuse_file_cached_io_open() returns true for success,
execution will bypass both branches of the if/else-if statement.

Since the initial err = -EINVAL assignment was removed earlier in the
function, the err variable remains unassigned in this scenario. This means
a successful file open will return an uninitialized stack variable,
potentially leading to undefined VFS behavior or random system call
failures depending on the stack contents.

Could we initialize err to 0 at declaration, or explicitly return 0 on
the success path?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001150935.655979-1-mszeredi@redhat.com?part=2

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 4/8] fuse: support opening 64 bit backing ID
  2026-10-01 15:07 ` [PATCH v2 4/8] fuse: support opening 64 bit " Miklos Szeredi
@ 2026-10-01 15:22   ` sashiko-bot
  2026-10-01 17:09   ` Amir Goldstein
  1 sibling, 0 replies; 29+ messages in thread
From: sashiko-bot @ 2026-10-01 15:22 UTC (permalink / raw)
  To: Miklos Szeredi; +Cc: linux-cxl, nvdimm

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Low] fuse: negative backing ID validation bypass due to unsigned conversion
--

commit 38737a0048b9ea224c699e27954c4cb00bdfc596
Author: Miklos Szeredi <mszeredi@redhat.com>

fuse: support opening 64 bit backing ID

This commit adds a backing_id_64 field to fuse_open_out to allow opening
files with 64-bit server-allocated backing IDs. When the server sets
FUSE_BACKING_ID_64, the kernel reads the backing ID from backing_id_64
instead of backing_id.

> diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
> index 1a10bc65fb310..811ea603b7767 100644
> --- a/fs/fuse/iomode.c
> +++ b/fs/fuse/iomode.c
> @@ -160,7 +160,9 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
>  {
>  	struct fuse_file *ff = file->private_data;
>  	struct fuse_conn *fc = get_fuse_conn(inode);
> +	struct fuse_open_out *outarg = &ff->args->open_outarg;
>  	struct fuse_backing *fb;
> +	u64 backing_id;
>  	int err;
>  
>  	/* Check allowed conditions for file open in passthrough mode */
> @@ -170,18 +172,38 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
>  	if (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK)
>  		return fuse_EIO("conflicting open flags");
>  
> -	fb = fuse_passthrough_open(file, ff->args->open_outarg.backing_id);
> -	if (IS_ERR(fb))
> -		return PTR_ERR(fb);
> +	if (!fc->backing_id_64) {
> +		if (outarg->backing_id_64 != 0)
> +			return fuse_EIO("64 bit backing ID set");
> +
> +		backing_id = outarg->backing_id;
> +		if (backing_id <= 0)

[Severity: Low]
Since outarg->backing_id is a signed 32-bit integer and backing_id is an
unsigned 64-bit integer, will negative IDs bypass this explicit error check
due to sign-extension during the assignment?

If outarg->backing_id is negative, it gets sign-extended into a large
positive value when assigned to the u64 backing_id. This causes the <= 0
check to evaluate to false, skipping the intended "invalid backing ID"
validation.

The code will then proceed and fail in fuse_backing_lookup(), logging
"backing not found" instead of the expected "invalid backing ID" error.

> +			return fuse_EIO("invalid backing ID");
> +	} else {
> +		if (outarg->backing_id != 0)
> +			return fuse_EIO("32 bit backing ID set");
> +
> +		backing_id = outarg->backing_id_64;
> +	}
> +	fb = fuse_backing_lookup(fc, backing_id);
> +	if (!fb)
> +		return fuse_EIO("backing not found");

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001150935.655979-1-mszeredi@redhat.com?part=4

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 6/8] fuse: add extent map data structure
  2026-10-01 15:07 ` [PATCH v2 6/8] fuse: add extent map data structure Miklos Szeredi
@ 2026-10-01 15:24   ` sashiko-bot
  0 siblings, 0 replies; 29+ messages in thread
From: sashiko-bot @ 2026-10-01 15:24 UTC (permalink / raw)
  To: Miklos Szeredi; +Cc: nvdimm, linux-cxl

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] fuse: sleep in RCU read-side critical section during backing files cleanup
- [Medium] fuse: missing overflow validation on ext->addr and ext->length
--

commit 791e33c26af3a517c133a125042b17d8f5ff57ec
Author: Miklos Szeredi <mszeredi@redhat.com>

fuse: add extent map data structure

Add support for creating and managing extent maps that map file regions to
dax device regions. Introduce FUSE_NOTIFY_BACKING_MAP for populating extent
maps from userspace.

> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index c852f04989619..e954391d696d0 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
[ ... ]
> @@ -306,12 +310,36 @@ static void fuse_backing_rht_free(void *p, void *data)
>  
>  void fuse_backing_files_free(struct fuse_conn *fc)
>  {
[ ... ]
> +	/*
> +	 * extents are referencing other backings, put these refs before
> +	 * destroying the backings themselves
> +	 */
> +	rhashtable_walk_enter(&fc->backing_64_ht, &iter);
> +	rhashtable_walk_start(&iter);
> +	while ((fb = rhashtable_walk_next(&iter))) {
> +		if (IS_ERR(fb)) {
> +			if (PTR_ERR(fb) == -EAGAIN)
> +				continue;
> +			break;
> +		}
> +		if (fb->type == FUSE_BACKING_EXTMAP) {
> +			fuse_ext_map_destroy(&fb->extents);

[Severity: High]
Does this code sleep in an RCU read-side critical section?

rhashtable_walk_start() explicitly acquires the RCU read lock.
When this code drops the final reference to the underlying DAXDEV backing,
the call chain is:

fuse_ext_map_destroy()
  fuse_backing_put()
    fuse_backing_free()
      fs_put_dax()
        iput()

Since iput() can sleep, it appears this could trigger 'scheduling while
atomic' panics or lockdep warnings when unmounting a FUSE passthrough
connection.

> +			fb->extents.rb_node = NULL;
> +		}
> +	}
> +	rhashtable_walk_stop(&iter);
> +	rhashtable_walk_exit(&iter);
> +
> +	rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
>  }

> diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
> new file mode 100644
> index 0000000000000..ab01d678d08e5
> --- /dev/null
> +++ b/fs/fuse/ext_map.c
[ ... ]
> +static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
> +			   struct fuse_extent *ext)
> +{
> +	struct fuse_iext *new_fie __free(kfree) = kzalloc_obj(*new_fie);
> +	struct rb_node *parent = NULL, **link = &extents->rb_node;
> +	loff_t end;
[ ... ]
> +	if (overflows_type(ext->offset, loff_t) ||
> +	    check_add_overflow(ext->offset, ext->length, &end))
> +		return fuse_EIO("offset overflow");
> +
> +	new_fie->start = ext->offset;
> +	new_fie->end = end;
> +	new_fie->backing_offset = ext->addr;

[Severity: Medium]
Is it possible for the physical address calculation to wrap around?

The bounds check above validates ext->offset, but appears to omit checking
ext->addr + ext->length for overflow.

Could a user-space process supplying an ext->addr near U64_MAX cause an
integer overflow on the physical address range calculation, resulting in a
non-contiguous backing mapping?

> +
> +	while (*link) {

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001150935.655979-1-mszeredi@redhat.com?part=6

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID
  2026-10-01 15:07 ` [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
@ 2026-10-01 15:26   ` sashiko-bot
  2026-10-01 17:07   ` Amir Goldstein
  1 sibling, 0 replies; 29+ messages in thread
From: sashiko-bot @ 2026-10-01 15:26 UTC (permalink / raw)
  To: Miklos Szeredi; +Cc: linux-cxl, nvdimm

Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] fuse: unchecked rhashtable_init() return leads to NULL dereference
- [High] fuse: race between process_init_reply() and ioctl corrupts IDR/rhashtable union
- [Medium] fuse: duplicate 64-bit backing IDs lead to memory leaks
--

commit 4b1027b3a4203f8caa807855de45ddb14b02bbb9
Author: Miklos Szeredi <mszeredi@redhat.com>

fuse: support 64 bit, server allocated backing ID

This commit adds support for server allocated 64-bit backing IDs
alongside the existing kernel allocated 32-bit IDs. It introduces
FUSE_DEV_IOC_BACKING_CREATE and FUSE_NOTIFY_BACKING_REMOVE for
managing the 64-bit backing IDs.

> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index 433fa3098d71f..3c879df7989c9 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
[ ... ]
> +static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
> +{
> +	return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
>  }

[Severity: Medium]
Does this allow duplicate backing IDs, leading to a memory leak?

rhashtable_insert_fast() does not check for duplicate keys. If a user
invokes FUSE_DEV_IOC_BACKING_CREATE multiple times with the exact same
64-bit map->backing_id, they will all be inserted.

When removing via fuse_backing_close_64(), rhashtable_remove_fast() will
only retrieve and delete one instance, leaving the duplicates orphaned and
leaking the struct fuse_backing objects.

[ ... ]
> +void fuse_backing_files_init_64(struct fuse_conn *fc)
> +{
> +	rhashtable_init(&fc->backing_64_ht, &fuse_backing_prm);
> +	fc->backing_id_64 = true;
>  }

[Severity: High]
What happens if rhashtable_init() fails here due to memory pressure?

Since the return value is ignored, fc->backing_id_64 = true is set
unconditionally. If rhashtable_init() returns an error, the hash table
remains uninitialized.

A subsequent ioctl triggering fuse_backing_open_64() would call
rhashtable_insert_fast(), which dereferences the NULL ht->tbl, resulting
in a kernel panic.

> diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> index 2a1aadc08e2a8..9b65543c6cc79 100644
> --- a/fs/fuse/fuse_i.h
> +++ b/fs/fuse/fuse_i.h
[ ... ]
> @@ -770,8 +775,14 @@ struct fuse_conn {
>  	struct fuse_sync_bucket __rcu *curr_bucket;
>  
>  #ifdef CONFIG_FUSE_PASSTHROUGH
> -	/** @backing_files_map: IDR for backing files ids */
> -	struct idr backing_files_map;
> +	/* Selected by backing_id_64 */
> +	union {
> +		/** @backing_files_map: IDR for backing files ids */
> +		struct idr backing_files_map;
> +
> +		/** @backing_64_ht: 64 bit ID lookup hash table */
> +		struct rhashtable backing_64_ht;
> +	};
>  #endif
>  };
>  
> diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
> index cbb10e19e7e86..bb76bddfc2417 100644
> --- a/fs/fuse/inode.c
> +++ b/fs/fuse/inode.c
[ ... ]
> @@ -1408,13 +1408,15 @@ static void process_init_reply(struct fuse_args *args, int error)
>  			 * them together.
>  			 */
>  			if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) &&
> -			    (flags & FUSE_PASSTHROUGH) &&
> +			    (flags & (FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2)) &&
>  			    arg->max_stack_depth > 0 &&
>  			    arg->max_stack_depth <= FILESYSTEM_MAX_STACK_DEPTH &&
>  			    !(flags & FUSE_WRITEBACK_CACHE))  {
>  				fc->passthrough = 1;
>  				fc->max_stack_depth = arg->max_stack_depth;
>  				fm->sb->s_stack_depth = arg->max_stack_depth;
> +				if (flags & FUSE_PASSTHROUGH_V2)
> +					fuse_backing_files_init_64(fc);

[Severity: High]
Can this sequence cause memory corruption in the union?

If fc->passthrough = 1 is set locklessly before fuse_backing_files_init_64(fc)
completes, a concurrent FUSE_DEV_IOC_BACKING_OPEN ioctl in fuse_backing_open()
could see fc->passthrough == 1 and fc->backing_id_64 == 0.

It would proceed to call fuse_backing_id_alloc(), modifying the 32-bit
backing_files_map IDR. Meanwhile, fuse_backing_files_init_64() executes
rhashtable_init(&fc->backing_64_ht), which would blindly overwrite the very
same memory union, corrupting the IDR state.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001150935.655979-1-mszeredi@redhat.com?part=3

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 7/8] fuse: add extent map I/O support
  2026-10-01 15:07 ` [PATCH v2 7/8] fuse: add extent map I/O support Miklos Szeredi
@ 2026-10-01 15:27   ` sashiko-bot
  2026-10-01 16:11   ` Amir Goldstein
  1 sibling, 0 replies; 29+ messages in thread
From: sashiko-bot @ 2026-10-01 15:27 UTC (permalink / raw)
  To: Miklos Szeredi; +Cc: linux-cxl, nvdimm

Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] fuse: missing inode_lock_shared() in DAX read path causes lockdep splats and races
- [High] fuse: fuse_rw_clamp() truncates writes at EOF to 0 bytes, causing infinite loops
- [High] fuse: O_APPEND writes bypass fuse_rw_clamp limits due to ordering
- [High] fuse: ABBA deadlock between mmap_lock and i_rwsem in DAX page fault
--

commit 900b63ca4cd5e837e1d2beda0ce7ca37d122ab2d
Author: Miklos Szeredi <mszeredi@redhat.com>
Subject: fuse: add extent map I/O support

Wire up read, write, splice and mmap operations for extent-mapped
files through the iomap/DAX infrastructure.

> diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
> index ab01d678d08e5..61cb22e821b60 100644
> --- a/fs/fuse/ext_map.c
> +++ b/fs/fuse/ext_map.c
[ ... ]
> +static vm_fault_t fuse_ext_map_huge_fault(struct vm_fault *vmf, unsigned int order)
> +{
> +	struct inode *inode = file_inode(vmf->vma->vm_file);
> +	bool write_fault = (vmf->flags & FAULT_FLAG_WRITE) && (vmf->vma->vm_flags & VM_SHARED);
> +	vm_fault_t ret;
> +	unsigned long pfn;
> +
> +	if (WARN_ON_ONCE(!IS_DAX(inode)))
> +		return VM_FAULT_SIGBUS;
> +
> +	if (write_fault) {
> +		sb_start_pagefault(inode->i_sb);
> +		file_update_time(vmf->vma->vm_file);
> +	}
> +
> +	filemap_invalidate_lock_shared(inode->i_mapping);
> +
> +	ret = dax_iomap_fault(vmf, order, &pfn, NULL, &fuse_ext_map_iomap_ops);
> +	if (ret & VM_FAULT_NEEDDSYNC)
> +		ret = dax_finish_sync_fault(vmf, order, pfn);
        ^^^^

[Severity: High]
Can this create an ABBA deadlock between mmap_lock and i_rwsem?

This page fault handler is entered holding the VFS mmap_lock, and then
takes filemap_invalidate_lock_shared(). If the fault requires a data
sync (O_DSYNC), it calls dax_finish_sync_fault() which synchronously
invokes vfs_fsync_range(). 

The filesystem's fsync handler, fuse_fsync(), unconditionally acquires
inode_lock(inode). This creates an mmap_lock -> inode_lock dependency,
which violates the standard kernel lock hierarchy and can deadlock if
another thread holds the inode_lock while taking a page fault.

> +
> +	filemap_invalidate_unlock_shared(inode->i_mapping);
> +
> +	if (write_fault)
> +		sb_end_pagefault(inode->i_sb);
> +
> +	return ret;
> +}
[ ... ]
> +ssize_t fuse_ext_map_write_iter(struct kiocb *iocb, struct iov_iter *from)
> +{
> +	ssize_t res;
> +
> +	fuse_rw_clamp(iocb, from);
> +
> +	res = generic_write_checks(iocb, from);
        ^^^^

[Severity: High]
Can O_APPEND writes bypass the file size limits checked in
fuse_rw_clamp() due to this ordering?

Because fuse_rw_clamp() runs before generic_write_checks(),
kiocb->ki_pos is not yet updated to i_size for O_APPEND writes. Clamping
will evaluate based on the old, smaller position, and
generic_write_checks() will then advance ki_pos to i_size, allowing an
out-of-bounds write via dax_iomap_rw().

[Severity: High]
Also, does fuse_rw_clamp() cause writes at or past EOF to return 0
instead of a standard POSIX error?

When a write starts at or past i_size, fuse_rw_clamp() truncates the
iterator to 0. Then, generic_write_checks() returns 0, which propagates
up to the VFS. This can cause standard utilities (like cp and dd) to
retry indefinitely in an infinite loop instead of receiving -ENOSPC or
-EFBIG.

> +	if (res <= 0)
> +		return res;
> +
> +	res = kiocb_modified(iocb);
> +	if (res)
> +		return res;
> +
> +	return dax_iomap_rw(iocb, from, &fuse_ext_map_iomap_ops);
> +}

> diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
> index e9ab1aea34e27..040817ad76e94 100644
> --- a/fs/fuse/passthrough.c
> +++ b/fs/fuse/passthrough.c
> @@ -48,10 +47,10 @@ ssize_t fuse_passthrough_read_iter(struct kiocb *iocb, struct iov_iter *iter)
>  	if (!count)
>  		return 0;
>  
> -	ret = backing_file_read_iter(backing_file, iter, iocb, iocb->ki_flags,
> -				     &ctx);
> +	if (!backing_file)
> +		return fuse_ext_map_read_iter(iocb, iter);
        ^^^^

[Severity: High]
Does this path miss acquiring the inode read lock before calling into
the DAX read operations?

When fuse_ext_map_read_iter() processes the read, it delegates to
dax_iomap_rw() which expects the VFS inode lock to be held for
synchronization against truncates. Skipping
inode_lock_shared(&inode->i_rwsem) here could lead to lockdep splats or
out-of-bounds reads during concurrent mapping changes.

>  
> -	return ret;
> +	return backing_file_read_iter(backing_file, iter, iocb, iocb->ki_flags, &ctx);
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001150935.655979-1-mszeredi@redhat.com?part=7

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 5/8] fuse: add support for opening dax device as backing
  2026-10-01 15:07 ` [PATCH v2 5/8] fuse: add support for opening dax device as backing Miklos Szeredi
@ 2026-10-01 15:30   ` sashiko-bot
  2026-10-01 16:07   ` Amir Goldstein
  1 sibling, 0 replies; 29+ messages in thread
From: sashiko-bot @ 2026-10-01 15:30 UTC (permalink / raw)
  To: Miklos Szeredi; +Cc: linux-cxl, nvdimm

Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [High] mm: missing a_ops methods on passthrough DAX inodes
- [High] mm: uninitialized memory in fuse_backing_new for dax_error
- [High] mm: NULL pointer dereference in fuse_dax_notify_failure race
- [High] mm: virtiofs DAX cache consistency broken by IS_DAX check
- [Medium] mm: RCU data race on backing_id in fuse_backing_close_64
--

commit db1b69a7f08e8b3b4063c0a40f3bd08ae1c2dbc6
Author: Miklos Szeredi <mszeredi@redhat.com>
Subject: fuse: add support for opening dax device as backing

This commit adds support for using a DAX device as a backing file for FUSE.
It marks the inode with S_DAX when FUSE_ATTR_DAX is set during lookup, and
prepares the backing object structures and ops for DAX devices.

> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
[ ... ]
> @@ -103,39 +111,83 @@ int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)

[Severity: Medium]
Could modifying fb->backing_id to 0 outside of the rhashtable lock while
lockless readers might still be inspecting it cause an RCU data race?
RCU readers could be concurrently comparing the key in rhashtable_lookup_fast,
potentially leading to KCSAN warnings.

>  	return 0;
>  }
>  
> -static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
> +static int fuse_dax_notify_failure(struct dax_device *daxdev, u64 offset, u64 len, int mf_flags)
>  {
> -	struct fuse_backing *fb;
> -	struct super_block *backing_sb;
> -	struct file *file;
> +	struct fuse_backing *fb = dax_holder(daxdev);
>  
> -	/* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> -	if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> -		return ERR_PTR(-EPERM);
> +	fb->dax_error = true;

[Severity: High]
Does this code risk a NULL pointer dereference if it races with fs_put_dax?
If a memory failure occurs concurrently with fuse_backing_free clearing
the DAX holder data, dax_holder could return NULL. Should we verify fb
is not NULL before accessing dax_error?

> +
> +	return 0;
> +}
[ ... ]
> +static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
> +{
> +	struct fuse_backing *fb __free(kfree) = kmalloc_obj(*fb);

[Severity: High]
Does this allocation leave fb->dax_error uninitialized? Since kmalloc_obj
does not zero memory, the adjacent union member dax_error might contain
garbage when the dax_dev is explicitly initialized for S_IFCHR later. If
subsequent patches in this series read dax_error (such as in
fuse_ext_map_iomap_begin), it could falsely trigger an error path and cause
erratic failures.

> +	int err;
>  
> -	fb = kmalloc_obj(struct fuse_backing);
>  	if (!fb)
>  		return ERR_PTR(-ENOMEM);
[ ... ]
> --- a/fs/fuse/file.c
> +++ b/fs/fuse/file.c
> @@ -297,7 +297,7 @@ static int fuse_open(struct inode *inode, struct file *file)
>  	if (!err) {
>  		if (is_truncate)
>  			truncate_pagecache(inode, 0);
> -		else if (!(ff->open_flags & FOPEN_KEEP_CACHE))
> +		else if (!(ff->open_flags & FOPEN_KEEP_CACHE) && !IS_DAX(inode))
>  			invalidate_inode_pages2(inode->i_mapping);

[Severity: High]
Could skipping invalidate_inode_pages2 for DAX inodes break virtiofs DAX
cache consistency? Virtiofs DAX relies on invalidate_inode_pages2 to clear
stale DAX exceptional entries from the page cache XArray. Without this
invalidation, stale mappings might remain resident and cause data corruption.

>  	}
>  out_unlock:
> --- a/fs/fuse/inode.c
> +++ b/fs/fuse/inode.c
[ ... ]
> @@ -403,6 +403,10 @@ static void fuse_init_submount_lookup(struct fuse_submount_lookup *sl,
>  	refcount_set(&sl->count, 1);
>  }
>  
> +static const struct address_space_operations fuse_dax_aops = {
> +	.dirty_folio	= noop_dirty_folio,
> +};
> +

[Severity: High]
Is it intentional that this struct lacks read_folio and write_begin? If a
DAX-enabled inode is opened without passthrough (e.g., FOPEN_PASSTHROUGH is
missing) or direct I/O, normal non-passthrough reads might be routed to the
page cache. The VFS could then dereference mapping->a_ops->read_folio which
is NULL, potentially causing a kernel panic.

>  static void fuse_init_inode(struct inode *inode, struct fuse_attr *attr,
>  			    struct fuse_conn *fc)
>  {

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20261001150935.655979-1-mszeredi@redhat.com?part=5

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 5/8] fuse: add support for opening dax device as backing
  2026-10-01 15:07 ` [PATCH v2 5/8] fuse: add support for opening dax device as backing Miklos Szeredi
  2026-10-01 15:30   ` sashiko-bot
@ 2026-10-01 16:07   ` Amir Goldstein
  1 sibling, 0 replies; 29+ messages in thread
From: Amir Goldstein @ 2026-10-01 16:07 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: fuse-devel, John Groves, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Thu, Oct 1, 2026 at 5:09 PM Miklos Szeredi <mszeredi@redhat.com> wrote:
>
> This is only possible with FUSE_PASSTHROUGH_V2 enabled.
>
> Mark the inode with S_DAX if FUSE_LOOKUP returns with FUSE_ATTR_DAX set.
>
> This patch does not yet provide a way actually use the dax dev backing:
> when such a backing ID is provided in reply to FUSE_OPEN with
> FOPEN_PASSTHROUGH flag set, an error will be returned.
>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
> ---
>  fs/fuse/backing.c     | 121 ++++++++++++++++++++++++++++++------------
>  fs/fuse/file.c        |   2 +-
>  fs/fuse/fuse_i.h      |  28 ++++++++--
>  fs/fuse/inode.c       |  11 +++-
>  fs/fuse/passthrough.c |   6 ++-
>  5 files changed, 126 insertions(+), 42 deletions(-)
>
> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index 3c879df7989c..c852f0498961 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
> @@ -9,6 +9,7 @@
>  #include "fuse_i.h"
>
>  #include <linux/file.h>
> +#include <linux/dax.h>
>  #include <linux/rhashtable.h>
>
>  static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb)
> @@ -22,9 +23,16 @@ static void fuse_backing_free(struct fuse_backing *fb)
>  {
>         pr_debug("%s: fb=0x%p\n", __func__, fb);
>
> -       if (fb->file)
> -               fput(fb->file);
> -       put_cred(fb->cred);
> +       switch (fb->type) {
> +       case FUSE_BACKING_PATH:
> +               path_put(&fb->path);
> +               put_cred(fb->cred);
> +               break;
> +
> +       case FUSE_BACKING_DAXDEV:
> +               fs_put_dax(fb->dax_dev, fb);
> +               break;
> +       }
>         kfree_rcu(fb, rcu);
>  }
>
> @@ -103,39 +111,83 @@ int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)
>         return 0;
>  }
>
> -static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
> +static int fuse_dax_notify_failure(struct dax_device *daxdev, u64 offset, u64 len, int mf_flags)
>  {
> -       struct fuse_backing *fb;
> -       struct super_block *backing_sb;
> -       struct file *file;
> +       struct fuse_backing *fb = dax_holder(daxdev);
>
> -       /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> -       if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> -               return ERR_PTR(-EPERM);
> +       fb->dax_error = true;
>
> -       CLASS(fd_raw, f)(fd);
> -       if (fd_empty(f))
> -               return ERR_PTR(-EBADF);
> +       return 0;
> +}
>
> -       file = fd_file(f);
> +static const struct dax_holder_operations fuse_dax_holder_ops = {
> +       .notify_failure         = fuse_dax_notify_failure,
> +};
> +
> +static int fuse_backing_open_file(struct fuse_conn *fc, struct fuse_backing *fb, struct file *file)
> +{
> +       struct inode *inode = file_inode(file);
> +       struct dax_device *daxdev;
> +       int err;
>
> -       /* read/write/splice/mmap passthrough only relevant for regular files */
> -       if (!d_is_reg(file->f_path.dentry))
> -               return d_is_dir(file->f_path.dentry) ? ERR_PTR(-EISDIR) : ERR_PTR(-EINVAL);
> +       switch (inode->i_mode & S_IFMT) {
> +       case S_IFREG:
> +               /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> +               if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> +                       return -EPERM;
> +
> +               if (inode->i_sb->s_stack_depth >= fc->max_stack_depth)
> +                       return -ELOOP;
> +
> +               fb->type = FUSE_BACKING_PATH;
> +               fb->path = file->f_path;
> +               path_get(&fb->path);
> +               fb->cred = get_current_cred();
> +               return 0;
> +
> +       case S_IFCHR:
> +               daxdev = dax_dev_find(inode->i_rdev);
> +               if (!daxdev)
> +                       return -EINVAL;
> +
> +               err = -EPERM;
> +               if (capable(CAP_SYS_RAWIO)) {
> +                       err = fs_dax_get(daxdev, fb, &fuse_dax_holder_ops);
> +                       if (!err) {
> +                               fb->type = FUSE_BACKING_DAXDEV;
> +                               fb->dax_dev = daxdev;
> +                       }
> +               }
> +               put_dax(daxdev);
> +               return err;
>
> -       backing_sb = file_inode(file)->i_sb;
> -       if (backing_sb->s_stack_depth >= fc->max_stack_depth)
> -               return ERR_PTR(-ELOOP);
> +       case S_IFDIR:
> +               return -EISDIR;
> +
> +       default:
> +               return -EINVAL;
> +       }
> +}
> +
> +static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
> +{
> +       struct fuse_backing *fb __free(kfree) = kmalloc_obj(*fb);
> +       int err;
>
> -       fb = kmalloc_obj(struct fuse_backing);
>         if (!fb)
>                 return ERR_PTR(-ENOMEM);
>
> -       fb->file = get_file(file);
> -       fb->cred = get_current_cred();
> +       CLASS(fd_raw, f)(fd);
> +       if (fd_empty(f))
> +               return ERR_PTR(-EBADF);
> +
> +       err = fuse_backing_open_file(fc, fb, fd_file(f));
> +       if (err)
> +               return ERR_PTR(err);
> +
>         refcount_set(&fb->count, 1);
>
> -       return fb;
> +       return_ptr(fb);
>  }
>
>  int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map)
> @@ -143,22 +195,21 @@ int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *ma
>         struct fuse_backing *fb;
>         int res;
>
> -       res = -EINVAL;
>         if (map->padding || map->spare[0] || map->spare[1])
> -               goto out;
> +               return -EINVAL;
>
>         if (!fc->backing_id_64)
> -               goto out;
> +               return -EINVAL;
>
>         fb = fuse_backing_new(fc, map->fd);
> -       res = PTR_ERR(fb);
> -       if (!IS_ERR(fb)) {
> -               fb->backing_id = map->backing_id;
> -               res = fuse_backing_add_64(fc, fb);
> -               if (res < 0)
> -                       fuse_backing_free(fb);
> -       }
> -out:
> +       if (IS_ERR(fb))
> +               return PTR_ERR(fb);
> +
> +       fb->backing_id = map->backing_id;
> +       res = fuse_backing_add_64(fc, fb);
> +       if (res < 0)
> +               fuse_backing_free(fb);
> +
>         return res;
>  }
>
> diff --git a/fs/fuse/file.c b/fs/fuse/file.c
> index 1036c06fe35a..fb8e0a9fa4f6 100644
> --- a/fs/fuse/file.c
> +++ b/fs/fuse/file.c
> @@ -297,7 +297,7 @@ static int fuse_open(struct inode *inode, struct file *file)
>         if (!err) {
>                 if (is_truncate)
>                         truncate_pagecache(inode, 0);
> -               else if (!(ff->open_flags & FOPEN_KEEP_CACHE))
> +               else if (!(ff->open_flags & FOPEN_KEEP_CACHE) && !IS_DAX(inode))
>                         invalidate_inode_pages2(inode->i_mapping);
>         }
>  out_unlock:
> diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> index 59afebc3dd2b..3e7ffb9bb2c3 100644
> --- a/fs/fuse/fuse_i.h
> +++ b/fs/fuse/fuse_i.h
> @@ -89,10 +89,25 @@ struct fuse_submount_lookup {
>         struct fuse_forget_link *forget;
>  };
>
> +enum fuse_backing_type {
> +       FUSE_BACKING_PATH,
> +       FUSE_BACKING_DAXDEV,
> +};
> +
>  /* Container for data related to mapping to backing file */
>  struct fuse_backing {
> -       struct file *file;
> -       const struct cred *cred;
> +       enum fuse_backing_type type;
> +
> +       union {
> +               struct {
> +                       struct path path;
> +                       const struct cred *cred;
> +               };
> +               struct {
> +                       struct dax_device *dax_dev;
> +                       bool dax_error;
> +               };
> +       };
>         u64 backing_id;
>         struct rhash_head hash_node;
>         /* refcount */
> @@ -1239,7 +1254,14 @@ void fuse_free_conn(struct fuse_conn *fc);
>
>  /* dax.c */
>
> -#define FUSE_IS_VDAX(inode) (IS_ENABLED(CONFIG_FUSE_VDAX) && IS_DAX(inode))
> +static inline bool FUSE_IS_VDAX(struct inode *inode)
> +{
> +#ifdef CONFIG_FUSE_VDAX
> +       return get_fuse_inode(inode)->vdax && IS_DAX(inode);
> +#else
> +       return false;
> +#endif
> +}
>
>  ssize_t fuse_vdax_read_iter(struct kiocb *iocb, struct iov_iter *to);
>  ssize_t fuse_vdax_write_iter(struct kiocb *iocb, struct iov_iter *from);
> diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
> index bb76bddfc241..d42c824fd6af 100644
> --- a/fs/fuse/inode.c
> +++ b/fs/fuse/inode.c
> @@ -148,7 +148,7 @@ static void fuse_evict_inode(struct inode *inode)
>         /* Will write inode on close/munmap and in all other dirtiers */
>         WARN_ON(inode_state_read_once(inode) & I_DIRTY_INODE);
>
> -       if (FUSE_IS_VDAX(inode))
> +       if (IS_DAX(inode))
>                 dax_break_layout_final(inode);
>
>         truncate_inode_pages_final(&inode->i_data);
> @@ -403,6 +403,10 @@ static void fuse_init_submount_lookup(struct fuse_submount_lookup *sl,
>         refcount_set(&sl->count, 1);
>  }
>
> +static const struct address_space_operations fuse_dax_aops = {
> +       .dirty_folio    = noop_dirty_folio,
> +};
> +
>  static void fuse_init_inode(struct inode *inode, struct fuse_attr *attr,
>                             struct fuse_conn *fc)
>  {
> @@ -430,6 +434,11 @@ static void fuse_init_inode(struct inode *inode, struct fuse_attr *attr,
>          */
>         if (!fc->posix_acl)
>                 inode->i_acl = inode->i_default_acl = ACL_DONT_CACHE;
> +
> +       if ((attr->flags & FUSE_ATTR_DAX) && !fc->vdax) {

This (fc->vdax) fails build without CONFIG_FUSE_DAX

> +               inode->i_flags |= S_DAX;
> +               inode->i_data.a_ops = &fuse_dax_aops;
> +       }
>  }
>
>  static int fuse_inode_eq(struct inode *inode, void *_nodeidp)
> diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
> index 4894842ad6d0..e9ab1aea34e2 100644
> --- a/fs/fuse/passthrough.c
> +++ b/fs/fuse/passthrough.c
> @@ -156,9 +156,11 @@ int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
>         struct fuse_file *ff = file->private_data;
>         struct file *backing_file;
>
> +       if (fb->type != FUSE_BACKING_PATH)
> +               return fuse_EIO("invalid backing type");
> +
>         /* Allocate backing file per fuse file to store fuse path */
> -       backing_file = backing_file_open(file, file->f_flags,
> -                                        &fb->file->f_path, fb->cred);
> +       backing_file = backing_file_open(file, file->f_flags, &fb->path, fb->cred);
>         if (IS_ERR(backing_file))
>                 return fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file));
>
> --
> 2.54.0
>

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 7/8] fuse: add extent map I/O support
  2026-10-01 15:07 ` [PATCH v2 7/8] fuse: add extent map I/O support Miklos Szeredi
  2026-10-01 15:27   ` sashiko-bot
@ 2026-10-01 16:11   ` Amir Goldstein
  1 sibling, 0 replies; 29+ messages in thread
From: Amir Goldstein @ 2026-10-01 16:11 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: fuse-devel, John Groves, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Thu, Oct 1, 2026 at 5:09 PM Miklos Szeredi <mszeredi@redhat.com> wrote:
>
> Wire up read, write, splice and mmap operations for extent-mapped
> files through the iomap/DAX infrastructure.
>
> When a passthrough file is opened with an EXTMAP backing, I/O is
> dispatched to the dax devices referenced by the extent map.
>
> If an address is not mapped, EIO is returned on the I/O operation.
>
> Reviewed-by: Amir Goldstein <amir73il@gmail.com>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
> ---
>  fs/fuse/backing.c     |  15 +++++
>  fs/fuse/ext_map.c     | 149 ++++++++++++++++++++++++++++++++++++++++++
>  fs/fuse/fuse_i.h      |   4 ++
>  fs/fuse/notify.c      |   3 +
>  fs/fuse/passthrough.c |  29 ++++++--
>  5 files changed, 194 insertions(+), 6 deletions(-)
>
> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index e954391d696d..2708f140b94f 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
> @@ -278,6 +278,21 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
>         return err;
>  }
>
> +bool fuse_backing_is_dax(struct fuse_backing *fb)
> +{
> +       switch (fb->type) {
> +       case FUSE_BACKING_PATH:
> +               return false;
> +       case FUSE_BACKING_DAXDEV:
> +               return true;
> +       case FUSE_BACKING_EXTMAP:
> +               return fuse_ext_map_is_dax(fb);
> +       default:
> +               WARN_ON(1);
> +               return false;
> +       }
> +}
> +
>  struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id)
>  {
>         struct fuse_backing *fb;
> diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
> index ab01d678d08e..61cb22e821b6 100644
> --- a/fs/fuse/ext_map.c
> +++ b/fs/fuse/ext_map.c
> @@ -3,6 +3,8 @@
>  #include "fuse_i.h"
>  #include <linux/rbtree.h>
>  #include <linux/pagemap.h>
> +#include <linux/iomap.h>
> +#include <linux/dax.h>
>
>  struct fuse_iext {
>         struct rb_node rb;
> @@ -22,6 +24,24 @@ void fuse_ext_map_destroy(struct rb_root *extents)
>         }
>  }
>
> +static struct fuse_iext *fuse_find_extent(struct rb_root *extents, u64 offset)
> +{
> +       struct rb_node *node = extents->rb_node;
> +
> +       while (node) {
> +               struct fuse_iext *fie = rb_entry(node, typeof(*fie), rb);
> +
> +               if (offset < fie->start)
> +                       node = node->rb_left;
> +               else if (offset >= fie->end)
> +                       node = node->rb_right;
> +               else
> +                       return fie;
> +       }
> +
> +       return NULL;
> +}
> +
>  static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
>                            struct fuse_extent *ext)
>  {
> @@ -76,6 +96,135 @@ static int fuse_add_extent(struct fuse_conn *fc, struct rb_root *extents,
>         return 0;
>  }
>
> +static int fuse_ext_map_iomap_begin(struct inode *inode, loff_t offset, loff_t length,
> +                                   unsigned int flags, struct iomap *iomap, struct iomap *srcmap)
> +{
> +       struct fuse_backing *fb = fuse_inode_backing(get_fuse_inode(inode));
> +       struct fuse_iext *fie;
> +       loff_t ext_len;
> +
> +       if (!fb || fb->type != FUSE_BACKING_EXTMAP)
> +               return fuse_EIO("missing or wrong type backing");
> +
> +       fie = fuse_find_extent(&fb->extents, offset);
> +       if (!fie)
> +               return fuse_EIO("missing mapping");
> +
> +       if (WARN_ON(fie->backing->type != FUSE_BACKING_DAXDEV))
> +               return fuse_EIO("wrong type backing for extent");
> +
> +       if (fie->backing->dax_error) {
> +               fuse_make_bad(inode);
> +               return fuse_EIO("dax error");
> +       }
> +
> +       ext_len = fie->end - fie->start;
> +
> +       iomap->offset = fie->start;
> +       iomap->addr = fie->backing_offset;
> +       iomap->length = ext_len;
> +       iomap->dax_dev = fie->backing->dax_dev;
> +       iomap->type = IOMAP_MAPPED;
> +       iomap->flags = 0;
> +
> +       return 0;
> +}
> +
> +static const struct iomap_ops fuse_ext_map_iomap_ops = {
> +       .iomap_begin            = fuse_ext_map_iomap_begin,
> +};
> +
> +static vm_fault_t fuse_ext_map_huge_fault(struct vm_fault *vmf, unsigned int order)
> +{
> +       struct inode *inode = file_inode(vmf->vma->vm_file);
> +       bool write_fault = (vmf->flags & FAULT_FLAG_WRITE) && (vmf->vma->vm_flags & VM_SHARED);
> +       vm_fault_t ret;
> +       unsigned long pfn;
> +
> +       if (WARN_ON_ONCE(!IS_DAX(inode)))
> +               return VM_FAULT_SIGBUS;
> +
> +       if (write_fault) {
> +               sb_start_pagefault(inode->i_sb);
> +               file_update_time(vmf->vma->vm_file);
> +       }
> +
> +       filemap_invalidate_lock_shared(inode->i_mapping);
> +
> +       ret = dax_iomap_fault(vmf, order, &pfn, NULL, &fuse_ext_map_iomap_ops);
> +       if (ret & VM_FAULT_NEEDDSYNC)
> +               ret = dax_finish_sync_fault(vmf, order, pfn);
> +
> +       filemap_invalidate_unlock_shared(inode->i_mapping);
> +
> +       if (write_fault)
> +               sb_end_pagefault(inode->i_sb);
> +
> +       return ret;
> +}
> +
> +static vm_fault_t fuse_ext_map_fault(struct vm_fault *vmf)
> +{
> +       return fuse_ext_map_huge_fault(vmf, 0);
> +}
> +
> +static const struct vm_operations_struct fuse_ext_map_vm_ops = {
> +       .fault          = fuse_ext_map_fault,
> +       .huge_fault     = fuse_ext_map_huge_fault,
> +       .page_mkwrite   = fuse_ext_map_fault,
> +       .pfn_mkwrite    = fuse_ext_map_fault,
> +};
> +
> +static void fuse_rw_clamp(struct kiocb *iocb, struct iov_iter *ubuf)
> +{
> +       struct inode *inode = iocb->ki_filp->f_mapping->host;
> +       loff_t i_size = i_size_read(inode);
> +       loff_t max_count = iocb->ki_pos >= i_size ? 0 : i_size - iocb->ki_pos;
> +
> +       if (iov_iter_count(ubuf) > max_count)
> +               iov_iter_truncate(ubuf, max_count);
> +}
> +
> +ssize_t fuse_ext_map_read_iter(struct kiocb *iocb, struct iov_iter *to)
> +{
> +       ssize_t res;
> +
> +       fuse_rw_clamp(iocb, to);
> +
> +       if (!iov_iter_count(to))
> +               return 0;
> +
> +       res = dax_iomap_rw(iocb, to, &fuse_ext_map_iomap_ops);

These fail link without CONFIG_FS_DAX
Because this file build does not depend on CONFIG_FUSE_DAX
Probably simplest to #define dax_iomap_rw() (-EIO)
locally in this case

> +
> +       file_accessed(iocb->ki_filp);
> +       return res;
> +}
> +
> +ssize_t fuse_ext_map_write_iter(struct kiocb *iocb, struct iov_iter *from)
> +{
> +       ssize_t res;
> +
> +       fuse_rw_clamp(iocb, from);
> +
> +       res = generic_write_checks(iocb, from);
> +       if (res <= 0)
> +               return res;
> +
> +       res = kiocb_modified(iocb);
> +       if (res)
> +               return res;
> +
> +       return dax_iomap_rw(iocb, from, &fuse_ext_map_iomap_ops);
> +}
> +
> +int fuse_ext_map_mmap(struct file *file, struct vm_area_struct *vma)
> +{
> +       file_accessed(file);
> +       vma->vm_ops = &fuse_ext_map_vm_ops;
> +       vm_flags_set(vma, VM_HUGEPAGE);
> +       return 0;
> +}
> +
>  bool fuse_ext_map_is_dax(struct fuse_backing *fb)
>  {
>         struct fuse_iext *fie;
> diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> index 5a20f3863ee2..7248588b45c4 100644
> --- a/fs/fuse/fuse_i.h
> +++ b/fs/fuse/fuse_i.h
> @@ -1316,6 +1316,7 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
>
>  struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
>  int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb);
> +bool fuse_backing_is_dax(struct fuse_backing *fb);
>  void fuse_backing_files_init(struct fuse_conn *fc);
>  void fuse_backing_files_init_64(struct fuse_conn *fc);
>  void fuse_backing_files_free(struct fuse_conn *fc);
> @@ -1369,6 +1370,9 @@ extern void fuse_sysctl_unregister(void);
>  /* ext_map.c */
>
>  void fuse_ext_map_destroy(struct rb_root *extents);
> +ssize_t fuse_ext_map_write_iter(struct kiocb *iocb, struct iov_iter *from);
> +ssize_t fuse_ext_map_read_iter(struct kiocb *iocb, struct iov_iter *to);
> +int fuse_ext_map_mmap(struct file *file, struct vm_area_struct *vma);
>  bool fuse_ext_map_is_dax(struct fuse_backing *fb);
>  int fuse_ext_map_populate(struct fuse_conn *fc, struct fuse_notify_backing_map_out *arg,
>                           struct fuse_extent *ext);
> diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c
> index 7c427f88bbc5..3bfedd0cdbdd 100644
> --- a/fs/fuse/notify.c
> +++ b/fs/fuse/notify.c
> @@ -448,6 +448,9 @@ static int fuse_notify_map(struct fuse_conn *fc, unsigned int size,
>         if (err)
>                 return err;
>
> +       if (!fc->backing_id_64)
> +               return -EINVAL;
> +
>         if (outarg.num_extents > FUSE_MAX_EXTENTS)
>                 return -EINVAL;
>
> diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
> index e9ab1aea34e2..040817ad76e9 100644
> --- a/fs/fuse/passthrough.c
> +++ b/fs/fuse/passthrough.c
> @@ -35,7 +35,6 @@ ssize_t fuse_passthrough_read_iter(struct kiocb *iocb, struct iov_iter *iter)
>         struct fuse_file *ff = file->private_data;
>         struct file *backing_file = fuse_file_passthrough(ff);
>         size_t count = iov_iter_count(iter);
> -       ssize_t ret;
>         struct backing_file_ctx ctx = {
>                 .cred = ff->cred,
>                 .accessed = fuse_file_accessed,
> @@ -48,10 +47,10 @@ ssize_t fuse_passthrough_read_iter(struct kiocb *iocb, struct iov_iter *iter)
>         if (!count)
>                 return 0;
>
> -       ret = backing_file_read_iter(backing_file, iter, iocb, iocb->ki_flags,
> -                                    &ctx);
> +       if (!backing_file)
> +               return fuse_ext_map_read_iter(iocb, iter);
>
> -       return ret;
> +       return backing_file_read_iter(backing_file, iter, iocb, iocb->ki_flags, &ctx);
>  }
>
>  ssize_t fuse_passthrough_write_iter(struct kiocb *iocb,
> @@ -74,10 +73,13 @@ ssize_t fuse_passthrough_write_iter(struct kiocb *iocb,
>         if (!count)
>                 return 0;
>
> -       inode_lock(inode);
> +       guard(rwsem_write)(&inode->i_rwsem);
> +
> +       if (!backing_file)
> +               return fuse_ext_map_write_iter(iocb, iter);
> +
>         ret = backing_file_write_iter(backing_file, iter, iocb, iocb->ki_flags,
>                                       &ctx);
> -       inode_unlock(inode);
>
>         return ret;
>  }
> @@ -98,6 +100,9 @@ ssize_t fuse_passthrough_splice_read(struct file *in, loff_t *ppos,
>         pr_debug("%s: backing_file=0x%p, pos=%lld, len=%zu, flags=0x%x\n", __func__,
>                  backing_file, *ppos, len, flags);
>
> +       if (!backing_file)
> +               return copy_splice_read(in, ppos, pipe, len, flags);
> +
>         init_sync_kiocb(&iocb, in);
>         iocb.ki_pos = *ppos;
>         ret = backing_file_splice_read(backing_file, &iocb, pipe, len, flags, &ctx);
> @@ -123,6 +128,9 @@ ssize_t fuse_passthrough_splice_write(struct pipe_inode_info *pipe,
>         pr_debug("%s: backing_file=0x%p, pos=%lld, len=%zu, flags=0x%x\n", __func__,
>                  backing_file, *ppos, len, flags);
>
> +       if (!backing_file)
> +               return iter_file_splice_write(pipe, out, ppos, len, flags);
> +
>         inode_lock(inode);
>         init_sync_kiocb(&iocb, out);
>         iocb.ki_pos = *ppos;
> @@ -145,6 +153,9 @@ ssize_t fuse_passthrough_mmap(struct file *file, struct vm_area_struct *vma)
>         pr_debug("%s: backing_file=0x%p, start=%lu, end=%lu\n", __func__,
>                  backing_file, vma->vm_start, vma->vm_end);
>
> +       if (!backing_file)
> +               return fuse_ext_map_mmap(file, vma);
> +
>         return backing_file_mmap(backing_file, vma, &ctx);
>  }
>
> @@ -156,6 +167,12 @@ int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
>         struct fuse_file *ff = file->private_data;
>         struct file *backing_file;
>
> +       if (fb->type == FUSE_BACKING_EXTMAP) {
> +               if (fuse_backing_is_dax(fb) != !!IS_DAX(file_inode(file)))
> +                       return fuse_EIO("dax mode mismatch");
> +               return 0;
> +       }
> +
>         if (fb->type != FUSE_BACKING_PATH)
>                 return fuse_EIO("invalid backing type");
>
> --
> 2.54.0
>

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message
  2026-10-01 15:07 ` [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
  2026-10-01 15:18   ` sashiko-bot
@ 2026-10-01 16:32   ` Amir Goldstein
  2026-10-05  9:46     ` Miklos Szeredi
  1 sibling, 1 reply; 29+ messages in thread
From: Amir Goldstein @ 2026-10-01 16:32 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: fuse-devel, John Groves, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Thu, Oct 1, 2026 at 5:09 PM Miklos Szeredi <mszeredi@redhat.com> wrote:
>
> Conditions are triggered by a buggy fuse server will result in -EIO which
> is hard to interpret without any context.
>
> Add helpers that print a short message to the kernel log as well as
> returning the error value.  This serves a dual purpose:
>
>  - documents the error condition inside the code
>
>  - allows the implementor of the fuse server to get the context for the
>    error
>
> Some debug messages are removed in favor of this.
>
> This also changes the mmap return value from -ETXTBSY to -ENODEV if a
> passthrough open raced with the prior check.  This makes both cases return
> -ENODEV if the file is in passthrough mode.
>
> Additionally change the return value of fuse_inode_uncached_io_start() and
> fuse_file_cached_io_open() from an error value to a bool (true on success)
> as the error value is translated anyway.
>
> Currently these use the pr_notice_once() variant.  Possibly should be
> changed to a ratelimit of e.g. once per minute.
>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
> ---
>  fs/fuse/file.c        |  8 ++---
>  fs/fuse/fuse_i.h      |  6 ++--
>  fs/fuse/iomode.c      | 74 ++++++++++++++++---------------------------
>  fs/fuse/passthrough.c | 19 ++++-------
>  4 files changed, 42 insertions(+), 65 deletions(-)
>
> diff --git a/fs/fuse/file.c b/fs/fuse/file.c
> index 6d707f2b3bff..92976906ab05 100644
> --- a/fs/fuse/file.c
> +++ b/fs/fuse/file.c
> @@ -1478,7 +1478,7 @@ static void fuse_dio_lock(struct kiocb *iocb, struct iov_iter *from,
>                  * have raced, so check it again.
>                  */
>                 if (fuse_io_past_eof(iocb, from) ||
> -                   fuse_inode_uncached_io_start(fi, NULL) != 0) {
> +                   !fuse_inode_uncached_io_start(fi, NULL)) {
>                         inode_unlock_shared(inode);
>                         inode_lock(inode);
>                         *exclusive = true;
> @@ -2415,7 +2415,6 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma)
>         struct fuse_file *ff = file->private_data;
>         struct fuse_conn *fc = ff->fm->fc;
>         struct inode *inode = file_inode(file);
> -       int rc;
>
>         /* DAX mmap is superior to direct_io mmap */
>         if (FUSE_IS_VDAX(inode))
> @@ -2457,9 +2456,8 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma)
>                  * After first mmap, the inode stays in caching io mode until
>                  * the direct_io file release.
>                  */
> -               rc = fuse_file_cached_io_open(inode, ff);
> -               if (rc)
> -                       return rc;
> +               if (!fuse_file_cached_io_open(inode, ff))
> +                       return -ENODEV;
>         }
>
>         if ((vma->vm_flags & VM_SHARED) && (vma->vm_flags & VM_MAYWRITE))
> diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> index 8546855386b5..3dd4dff24c50 100644
> --- a/fs/fuse/fuse_i.h
> +++ b/fs/fuse/fuse_i.h
> @@ -11,6 +11,8 @@
>  # define pr_fmt(fmt) "fuse: " fmt
>  #endif
>
> +#define fuse_EIO(fmt, ...) (pr_notice_once("%s: " fmt "\n", __func__, ##__VA_ARGS__), -EIO)
> +
>  #include "args.h"
>  #include <linux/fuse.h>
>  #include <linux/fs.h>
> @@ -1251,8 +1253,8 @@ int fuse_fileattr_set(struct mnt_idmap *idmap,
>                       struct dentry *dentry, struct file_kattr *fa);
>
>  /* iomode.c */
> -int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff);
> -int fuse_inode_uncached_io_start(struct fuse_inode *fi,
> +bool fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff);
> +bool fuse_inode_uncached_io_start(struct fuse_inode *fi,
>                                  struct fuse_backing *fb);
>  void fuse_inode_uncached_io_end(struct fuse_inode *fi);
>
> diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
> index 79637c09e883..1a10bc65fb31 100644
> --- a/fs/fuse/iomode.c
> +++ b/fs/fuse/iomode.c
> @@ -27,15 +27,15 @@ static inline bool fuse_is_io_cache_wait(struct fuse_inode *fi)
>   * Blocks new parallel dio writes and waits for the in-progress parallel dio
>   * writes to complete.
>   */
> -int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
> +bool fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
>  {
>         struct fuse_inode *fi = get_fuse_inode(inode);
>
>         /* There are no io modes if server does not implement open */
>         if (!ff->args)
> -               return 0;
> +               return true;
>
> -       spin_lock(&fi->lock);
> +       guard(spinlock)(&fi->lock);
>         /*
>          * Setting the bit advises new direct-io writes to use an exclusive
>          * lock - without it the wait below might be forever.
> @@ -53,8 +53,7 @@ int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
>          */
>         if (fuse_inode_backing(fi)) {
>                 clear_bit(FUSE_I_CACHE_IO_MODE, &fi->state);
> -               spin_unlock(&fi->lock);
> -               return -ETXTBSY;
> +               return false;
>         }
>
>         WARN_ON(ff->iomode == IOM_UNCACHED);
> @@ -64,8 +63,7 @@ int fuse_file_cached_io_open(struct inode *inode, struct fuse_file *ff)
>                         set_bit(FUSE_I_CACHE_IO_MODE, &fi->state);
>                 fi->iocachectr++;
>         }
> -       spin_unlock(&fi->lock);
> -       return 0;
> +       return true;
>  }
>
>  static void fuse_file_cached_io_release(struct fuse_file *ff,
> @@ -82,22 +80,19 @@ static void fuse_file_cached_io_release(struct fuse_file *ff,
>  }
>
>  /* Start strictly uncached io mode where cache access is not allowed */
> -int fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
> +bool fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
>  {
>         struct fuse_backing *oldfb;
> -       int err = 0;
>
> -       spin_lock(&fi->lock);
> +       guard(spinlock)(&fi->lock);
>         /* deny conflicting backing files on same fuse inode */
>         oldfb = fuse_inode_backing(fi);
> -       if (fb && oldfb && oldfb != fb) {
> -               err = -EBUSY;
> -               goto unlock;
> -       }
> -       if (fi->iocachectr > 0) {
> -               err = -ETXTBSY;
> -               goto unlock;
> -       }
> +       if (fb && oldfb && oldfb != fb)
> +               return false;
> +
> +       if (fi->iocachectr > 0)
> +               return false;
> +
>         fi->iocachectr--;
>
>         /* fuse inode holds a single refcount of backing file */
> @@ -107,9 +102,7 @@ int fuse_inode_uncached_io_start(struct fuse_inode *fi, struct fuse_backing *fb)
>         } else {
>                 fuse_backing_put(fb);
>         }
> -unlock:
> -       spin_unlock(&fi->lock);
> -       return err;
> +       return true;
>  }
>
>  /* Takes uncached_io inode mode reference to be dropped on file release */
> @@ -118,11 +111,9 @@ static int fuse_file_uncached_io_open(struct inode *inode,
>                                       struct fuse_backing *fb)
>  {
>         struct fuse_inode *fi = get_fuse_inode(inode);
> -       int err;
>
> -       err = fuse_inode_uncached_io_start(fi, fb);
> -       if (err)
> -               return err;
> +       if (!fuse_inode_uncached_io_start(fi, fb))
> +               return fuse_EIO("failed to start uncached I/O");

Why lose the internal error code?
If server implementation is wrong that is a good hint of what went wrong.
ETXTBSY is trying to mix cache mode and passthrough mode.
EBUSY is trying to passthrough to two different backing inodes.
I am not claiming this is a good API but at least we (or LLM) could do
something useful with this lost information.

>
>         WARN_ON(ff->iomode != IOM_NONE);
>         ff->iomode = IOM_UNCACHED;
> @@ -173,9 +164,11 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
>         int err;
>
>         /* Check allowed conditions for file open in passthrough mode */
> -       if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) || !fc->passthrough ||
> -           (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK))
> -               return -EINVAL;
> +       if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) || !fc->passthrough)
> +               return fuse_EIO("passthrough not enabled");
> +
> +       if (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK)
> +               return fuse_EIO("conflicting open flags");
>
>         fb = fuse_passthrough_open(file, ff->args->open_outarg.backing_id);
>         if (IS_ERR(fb))
> @@ -208,11 +201,12 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
>
>         /*
>          * Server is expected to use FOPEN_PASSTHROUGH for all opens of an inode
> -        * which is already open for passthrough.
> +        * which is already open for passthrough.  Using incorrect open mode is
> +        * a server mistake, which results in user visible failure of open()
> +        * with EIO error.
>          */
> -       err = -EINVAL;
>         if (fuse_inode_backing(fi) && !(ff->open_flags & FOPEN_PASSTHROUGH))
> -               goto fail;
> +               return fuse_EIO("FOPEN_PASSTHROUGH expected");
>
>         /*
>          * FOPEN_PARALLEL_DIRECT_WRITES requires FOPEN_DIRECT_IO.
> @@ -234,22 +228,10 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
>
>         if (ff->open_flags & FOPEN_PASSTHROUGH)
>                 err = fuse_file_passthrough_open(inode, file);
> -       else
> -               err = fuse_file_cached_io_open(inode, ff);
> -       if (err)
> -               goto fail;
> +       else if (!fuse_file_cached_io_open(inode, ff))
> +               err = fuse_EIO("conflicting passthrough open");

Same here.
If we can keep this bit of information and print it - it is helpful IMO,
even if we don't go all the way to annotate every error case with
a pretty log.

Thanks,
Amir.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID
  2026-10-01 15:07 ` [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
  2026-10-01 15:26   ` sashiko-bot
@ 2026-10-01 17:07   ` Amir Goldstein
  2026-10-01 18:58     ` Amir Goldstein
  1 sibling, 1 reply; 29+ messages in thread
From: Amir Goldstein @ 2026-10-01 17:07 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: fuse-devel, John Groves, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Thu, Oct 1, 2026 at 5:09 PM Miklos Szeredi <mszeredi@redhat.com> wrote:
>
> Add support for server allocated 64-bit backing IDs alongside the
> existing kernel allocated 32-bit IDs.
>
> Opened with FUSE_DEV_IOC_BACKING_OPEN, but the backing ID is sent via
> fuse_backing_map.backing_id, with FUSE_BACKING_ID_64 being set in .flags.

Description outdated.

>
> Closed with FUSE_NOTIFY_BACKING_REMOVE.  Since close only provides the
> backing ID, not the file descriptor, it doesn't have to be done with an
> ioctl.
>
> All 64 bit values are valid.

Can you reserve id 0? or have a reason to need it?

>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
> ---
>  fs/fuse/backing.c         | 180 ++++++++++++++++++++++++++++----------
>  fs/fuse/dev.c             |  21 +++++
>  fs/fuse/dev.h             |   3 +
>  fs/fuse/fuse_i.h          |  21 ++++-
>  fs/fuse/inode.c           |   6 +-
>  fs/fuse/notify.c          |  28 ++++++
>  fs/fuse/passthrough.c     |   3 +
>  include/uapi/linux/fuse.h |  22 ++++-
>  8 files changed, 231 insertions(+), 53 deletions(-)
>
> diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> index 433fa3098d71..3c879df7989c 100644
> --- a/fs/fuse/backing.c
> +++ b/fs/fuse/backing.c
> @@ -9,6 +9,7 @@
>  #include "fuse_i.h"
>
>  #include <linux/file.h>
> +#include <linux/rhashtable.h>
>
>  static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb)
>  {
> @@ -48,6 +49,8 @@ static int fuse_backing_id_alloc(struct fuse_conn *fc, struct fuse_backing *fb)
>         id = idr_alloc_cyclic(&fc->backing_files_map, fb, 1, 0, GFP_ATOMIC);
>         spin_unlock(&fc->lock);
>         idr_preload_end();
> +       if (id > 0)
> +               fb->backing_id = id;
>
>         WARN_ON_ONCE(id == 0);
>         return id;
> @@ -61,81 +64,131 @@ static struct fuse_backing *fuse_backing_id_remove(struct fuse_conn *fc,
>         spin_lock(&fc->lock);
>         fb = idr_remove(&fc->backing_files_map, id);
>         spin_unlock(&fc->lock);
> +       if (fb)
> +               fb->backing_id = 0;
>
>         return fb;
>  }
>
> -static int fuse_backing_id_free(int id, void *p, void *data)
> -{
> -       struct fuse_backing *fb = p;
> +static const struct rhashtable_params fuse_backing_prm = {
> +       .head_offset = offsetof(struct fuse_backing, hash_node),
> +       .key_offset = offsetof(struct fuse_backing, backing_id),
> +       .key_len = sizeof_field(struct fuse_backing, backing_id),
> +};
>
> -       WARN_ON_ONCE(refcount_read(&fb->count) != 1);
> -       fuse_backing_free(fb);
> -       return 0;
> +static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
> +{
> +       return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
>  }
>
> -void fuse_backing_files_free(struct fuse_conn *fc)
> +int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)
>  {
> -       idr_for_each(&fc->backing_files_map, fuse_backing_id_free, NULL);
> -       idr_destroy(&fc->backing_files_map);
> +       struct fuse_backing *fb;
> +       int err;
> +
> +       if (!fc->backing_id_64)
> +               return -EINVAL;
> +
> +       scoped_guard(spinlock, &fc->lock) {
> +               fb = rhashtable_lookup_fast(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
> +               if (!fb)
> +                       return -ENOENT;
> +
> +               err = rhashtable_remove_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
> +               WARN_ON(err);
> +       }
> +       fb->backing_id = 0;
> +       fuse_backing_put(fb);
> +
> +       return 0;
>  }
>
> -int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
> +static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
>  {
> -       struct file *file;
> +       struct fuse_backing *fb;
>         struct super_block *backing_sb;
> -       struct fuse_backing *fb = NULL;
> -       int res;
> -
> -       pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
> +       struct file *file;
>
>         /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> -       res = -EPERM;
>         if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> -               goto out;
> +               return ERR_PTR(-EPERM);
>
> -       res = -EINVAL;
> -       if (map->flags || map->padding)
> -               goto out;
> +       CLASS(fd_raw, f)(fd);
> +       if (fd_empty(f))
> +               return ERR_PTR(-EBADF);
>
> -       file = fget_raw(map->fd);
> -       res = -EBADF;
> -       if (!file)
> -               goto out;
> +       file = fd_file(f);
>
>         /* read/write/splice/mmap passthrough only relevant for regular files */
> -       res = d_is_dir(file->f_path.dentry) ? -EISDIR : -EINVAL;
>         if (!d_is_reg(file->f_path.dentry))
> -               goto out_fput;
> +               return d_is_dir(file->f_path.dentry) ? ERR_PTR(-EISDIR) : ERR_PTR(-EINVAL);
>
>         backing_sb = file_inode(file)->i_sb;
> -       res = -ELOOP;
>         if (backing_sb->s_stack_depth >= fc->max_stack_depth)
> -               goto out_fput;
> +               return ERR_PTR(-ELOOP);
>
>         fb = kmalloc_obj(struct fuse_backing);
> -       res = -ENOMEM;
>         if (!fb)
> -               goto out_fput;
> +               return ERR_PTR(-ENOMEM);
>
> -       fb->file = file;
> +       fb->file = get_file(file);
>         fb->cred = get_current_cred();
>         refcount_set(&fb->count, 1);
>
> -       res = fuse_backing_id_alloc(fc, fb);
> -       if (res < 0) {
> -               fuse_backing_free(fb);
> -               fb = NULL;
> +       return fb;
> +}
> +
> +int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map)
> +{
> +       struct fuse_backing *fb;
> +       int res;
> +
> +       res = -EINVAL;
> +       if (map->padding || map->spare[0] || map->spare[1])

Please reserve backing_id 0

> +               goto out;
> +
> +       if (!fc->backing_id_64)
> +               goto out;
> +
> +       fb = fuse_backing_new(fc, map->fd);
> +       res = PTR_ERR(fb);
> +       if (!IS_ERR(fb)) {
> +               fb->backing_id = map->backing_id;
> +               res = fuse_backing_add_64(fc, fb);
> +               if (res < 0)
> +                       fuse_backing_free(fb);
>         }
> +out:
> +       return res;
> +}
> +
> +int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
> +{
> +       struct fuse_backing *fb = NULL;
> +       int res;
> +
> +       pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
>
> +       res = -EINVAL;
> +       if (map->flags || map->padding)
> +               goto out;
> +
> +       if (fc->backing_id_64)
> +               goto out;
> +
> +       fb = fuse_backing_new(fc, map->fd);
> +       res = PTR_ERR(fb);
> +       if (!IS_ERR(fb)) {
> +               res = fuse_backing_id_alloc(fc, fb);
> +               if (res < 0) {
> +                       fuse_backing_free(fb);
> +                       fb = NULL;
> +               }
> +       }
>  out:
>         pr_debug("%s: fb=0x%p, ret=%i\n", __func__, fb, res);
>
>         return res;
> -
> -out_fput:
> -       fput(file);
> -       goto out;
>  }
>
>  int fuse_backing_close(struct fuse_conn *fc, int backing_id)
> @@ -145,6 +198,9 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
>
>         pr_debug("%s: backing_id=%d\n", __func__, backing_id);
>
> +       if (fc->backing_id_64)
> +               return -EINVAL;
> +
>         /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
>         err = -EPERM;
>         if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> @@ -167,14 +223,48 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
>         return err;
>  }
>
> -struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id)
> +struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id)
>  {
>         struct fuse_backing *fb;
>
> -       rcu_read_lock();
> -       fb = idr_find(&fc->backing_files_map, backing_id);
> -       fb = fuse_backing_get(fb);
> -       rcu_read_unlock();
> +       guard(rcu)();
> +       if (!fc->backing_id_64)
> +               fb = idr_find(&fc->backing_files_map, backing_id);
> +       else
> +               fb = rhashtable_lookup(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
>
> -       return fb;
> +       return fuse_backing_get(fb);
> +}
> +
> +static void fuse_backing_check_free(struct fuse_backing *fb)
> +{
> +       WARN_ON_ONCE(refcount_read(&fb->count) != 1);
> +       fuse_backing_free(fb);
> +}
> +
> +static int fuse_backing_idr_free(int id, void *p, void *data)
> +{
> +       fuse_backing_check_free(p);
> +       return 0;
> +}
> +
> +static void fuse_backing_rht_free(void *p, void *data)
> +{
> +       fuse_backing_check_free(p);
> +}
> +
> +void fuse_backing_files_free(struct fuse_conn *fc)
> +{
> +       if (fc->backing_id_64) {
> +               rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
> +       } else {
> +               idr_for_each(&fc->backing_files_map, fuse_backing_idr_free, NULL);
> +               idr_destroy(&fc->backing_files_map);
> +       }
> +}
> +
> +void fuse_backing_files_init_64(struct fuse_conn *fc)
> +{
> +       rhashtable_init(&fc->backing_64_ht, &fuse_backing_prm);
> +       fc->backing_id_64 = true;
>  }
> diff --git a/fs/fuse/dev.c b/fs/fuse/dev.c
> index 2d7ee5498f1c..8b63a08a26b6 100644
> --- a/fs/fuse/dev.c
> +++ b/fs/fuse/dev.c
> @@ -2339,6 +2339,24 @@ static long fuse_dev_ioctl_backing_open(struct file *file,
>         return fuse_backing_open(fud->chan->conn, &map);
>  }
>
> +static long fuse_dev_ioctl_backing_create(struct file *file,
> +                                         struct fuse_backing_create_in __user *argp)
> +{
> +       struct fuse_dev *fud = fuse_get_dev(file);
> +       struct fuse_backing_create_in map;
> +
> +       if (IS_ERR(fud))
> +               return PTR_ERR(fud);
> +
> +       if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
> +               return -EOPNOTSUPP;
> +
> +       if (copy_from_user(&map, argp, sizeof(map)))
> +               return -EFAULT;
> +
> +       return fuse_backing_open_64(fud->chan->conn, &map);
> +}
> +
>  static long fuse_dev_ioctl_backing_close(struct file *file, __u32 __user *argp)
>  {
>         struct fuse_dev *fud = fuse_get_dev(file);
> @@ -2379,6 +2397,9 @@ static long fuse_dev_ioctl(struct file *file, unsigned int cmd,
>         case FUSE_DEV_IOC_BACKING_OPEN:
>                 return fuse_dev_ioctl_backing_open(file, argp);
>
> +       case FUSE_DEV_IOC_BACKING_CREATE:
> +               return fuse_dev_ioctl_backing_create(file, argp);
> +
>         case FUSE_DEV_IOC_BACKING_CLOSE:
>                 return fuse_dev_ioctl_backing_close(file, argp);
>
> diff --git a/fs/fuse/dev.h b/fs/fuse/dev.h
> index f6c47ae0395b..dbffd5bed2f5 100644
> --- a/fs/fuse/dev.h
> +++ b/fs/fuse/dev.h
> @@ -14,6 +14,7 @@ struct fuse_dev;
>  struct fuse_args;
>  struct fuse_copy_state;
>  struct fuse_backing_map;
> +struct fuse_backing_create_in;
>  struct file;
>  struct folio;
>  enum fuse_notify_code;
> @@ -87,6 +88,8 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
>
>  int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map);
>  int fuse_backing_close(struct fuse_conn *fc, int backing_id);
> +int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map);
> +int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id);
>
>  int fuse_copy_one(struct fuse_copy_state *cs, void *val, unsigned size);
>  int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop,
> diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> index 3dd4dff24c50..003ef3c35d7a 100644
> --- a/fs/fuse/fuse_i.h
> +++ b/fs/fuse/fuse_i.h
> @@ -32,6 +32,7 @@
>  #include <linux/pid_namespace.h>
>  #include <linux/refcount.h>
>  #include <linux/user_namespace.h>
> +#include <linux/rhashtable-types.h>
>
>  /** Default max number of pages that can be used in a single read request */
>  #define FUSE_DEFAULT_MAX_PAGES_PER_REQ 32
> @@ -92,7 +93,8 @@ struct fuse_submount_lookup {
>  struct fuse_backing {
>         struct file *file;
>         const struct cred *cred;
> -
> +       u64 backing_id;
> +       struct rhash_head hash_node;
>         /* refcount */
>         refcount_t count;
>         struct rcu_head rcu;
> @@ -689,6 +691,9 @@ struct fuse_conn {
>         /** @init_security: Initialize security xattrs when creating a new inode */
>         unsigned int init_security:1;
>
> +       /** Backing ID is 64 bit and allocated by the server */
> +       bool backing_id_64:1;
> +
>         /**
>          * @create_supp_group: Add supplementary group info when creating
>          * a new inode
> @@ -770,8 +775,14 @@ struct fuse_conn {
>         struct fuse_sync_bucket __rcu *curr_bucket;
>
>  #ifdef CONFIG_FUSE_PASSTHROUGH
> -       /** @backing_files_map: IDR for backing files ids */
> -       struct idr backing_files_map;
> +       /* Selected by backing_id_64 */
> +       union {
> +               /** @backing_files_map: IDR for backing files ids */
> +               struct idr backing_files_map;
> +
> +               /** @backing_64_ht: 64 bit ID lookup hash table */
> +               struct rhashtable backing_64_ht;
> +       };
>  #endif
>  };
>
> @@ -1270,7 +1281,7 @@ void fuse_file_release(struct inode *inode, struct fuse_file *ff,
>  /* backing.c */
>  #ifdef CONFIG_FUSE_PASSTHROUGH
>  void fuse_backing_put(struct fuse_backing *fb);
> -struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id);
> +
>  #else
>
>  static inline void fuse_backing_put(struct fuse_backing *fb)
> @@ -1278,7 +1289,9 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
>  }
>  #endif
>
> +struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
>  void fuse_backing_files_init(struct fuse_conn *fc);
> +void fuse_backing_files_init_64(struct fuse_conn *fc);
>  void fuse_backing_files_free(struct fuse_conn *fc);
>
>  /* passthrough.c */
> diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
> index cbb10e19e7e8..bb76bddfc241 100644
> --- a/fs/fuse/inode.c
> +++ b/fs/fuse/inode.c
> @@ -1408,13 +1408,15 @@ static void process_init_reply(struct fuse_args *args, int error)
>                          * them together.
>                          */
>                         if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) &&
> -                           (flags & FUSE_PASSTHROUGH) &&
> +                           (flags & (FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2)) &&
>                             arg->max_stack_depth > 0 &&
>                             arg->max_stack_depth <= FILESYSTEM_MAX_STACK_DEPTH &&
>                             !(flags & FUSE_WRITEBACK_CACHE))  {
>                                 fc->passthrough = 1;
>                                 fc->max_stack_depth = arg->max_stack_depth;
>                                 fm->sb->s_stack_depth = arg->max_stack_depth;
> +                               if (flags & FUSE_PASSTHROUGH_V2)
> +                                       fuse_backing_files_init_64(fc);

It would feel a little less hacky if we do:
    else
          fuse_backing_files_init_32(fc);

and set
         fc->passthrough = 1;

inside both of them to gate against fuse_backing_files_free()
called without processing init response,
or without either FUSE_PASSTHROUGH in the init response.

Thanks,
Amir.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 4/8] fuse: support opening 64 bit backing ID
  2026-10-01 15:07 ` [PATCH v2 4/8] fuse: support opening 64 bit " Miklos Szeredi
  2026-10-01 15:22   ` sashiko-bot
@ 2026-10-01 17:09   ` Amir Goldstein
  1 sibling, 0 replies; 29+ messages in thread
From: Amir Goldstein @ 2026-10-01 17:09 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: fuse-devel, John Groves, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Thu, Oct 1, 2026 at 5:09 PM Miklos Szeredi <mszeredi@redhat.com> wrote:
>
> Add backing_id_64 field to fuse_open_out to allow opening files with
> 64-bit server-allocated backing IDs.
>
> When the server sets FUSE_BACKING_ID_64 in open_out.open_flags, the
> kernel reads the backing ID from open_out.backing_id_64 instead of
> open_out.backing_id.
>
> The open reply is now variable-length for backward compatibility with
> servers that don't send the new backing_id_64 field.
>
> Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>

Reviewed-by: Amir Goldstein <amir73il@gmail.com>

> ---
>  fs/fuse/dir.c             |  3 ++-
>  fs/fuse/file.c            |  6 +++++-
>  fs/fuse/fuse_i.h          |  2 +-
>  fs/fuse/iomode.c          | 34 ++++++++++++++++++++++++++++------
>  fs/fuse/passthrough.c     | 28 ++++++----------------------
>  include/uapi/linux/fuse.h |  2 ++
>  6 files changed, 44 insertions(+), 31 deletions(-)
>
> diff --git a/fs/fuse/dir.c b/fs/fuse/dir.c
> index f48fafccce4b..874ec7cfffb1 100644
> --- a/fs/fuse/dir.c
> +++ b/fs/fuse/dir.c
> @@ -876,6 +876,7 @@ static int fuse_create_open(struct mnt_idmap *idmap, struct inode *dir,
>         args.out_args[0].value = &outentry;
>         /* Store outarg for fuse_finish_open() */
>         outopenp = &ff->args->open_outarg;
> +       args.out_argvar = true; /* compat */
>         args.out_args[1].size = sizeof(*outopenp);
>         args.out_args[1].value = outopenp;
>
> @@ -885,7 +886,7 @@ static int fuse_create_open(struct mnt_idmap *idmap, struct inode *dir,
>
>         err = fuse_simple_idmap_request(idmap, fm, &args);
>         free_ext_value(&args);
> -       if (err)
> +       if (err < 0)
>                 goto out_free_ff;
>
>         err = -EIO;
> diff --git a/fs/fuse/file.c b/fs/fuse/file.c
> index 92976906ab05..1036c06fe35a 100644
> --- a/fs/fuse/file.c
> +++ b/fs/fuse/file.c
> @@ -28,6 +28,7 @@ static int fuse_send_open(struct fuse_mount *fm, u64 nodeid,
>  {
>         struct fuse_open_in inarg;
>         FUSE_ARGS(args);
> +       int res;
>
>         memset(&inarg, 0, sizeof(inarg));
>         inarg.flags = open_flags & ~(O_CREAT | O_EXCL | O_NOCTTY);
> @@ -45,10 +46,13 @@ static int fuse_send_open(struct fuse_mount *fm, u64 nodeid,
>         args.in_args[0].size = sizeof(inarg);
>         args.in_args[0].value = &inarg;
>         args.out_numargs = 1;
> +       args.out_argvar = true; /* compat */
>         args.out_args[0].size = sizeof(*outargp);
>         args.out_args[0].value = outargp;
>
> -       return fuse_simple_request(fm, &args);
> +       res = fuse_simple_request(fm, &args);
> +
> +       return res < 0 ? res : 0;
>  }
>
>  struct fuse_file *fuse_file_alloc(struct fuse_mount *fm, bool release)
> diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> index 003ef3c35d7a..59afebc3dd2b 100644
> --- a/fs/fuse/fuse_i.h
> +++ b/fs/fuse/fuse_i.h
> @@ -1314,7 +1314,7 @@ static inline struct fuse_backing *fuse_inode_backing_set(struct fuse_inode *fi,
>  #endif
>  }
>
> -struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id);
> +int fuse_passthrough_open(struct file *file, struct fuse_backing *fb);
>  void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb);
>
>  static inline bool fuse_is_passthrough(struct fuse_file *ff)
> diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c
> index 1a10bc65fb31..811ea603b776 100644
> --- a/fs/fuse/iomode.c
> +++ b/fs/fuse/iomode.c
> @@ -160,7 +160,9 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
>  {
>         struct fuse_file *ff = file->private_data;
>         struct fuse_conn *fc = get_fuse_conn(inode);
> +       struct fuse_open_out *outarg = &ff->args->open_outarg;
>         struct fuse_backing *fb;
> +       u64 backing_id;
>         int err;
>
>         /* Check allowed conditions for file open in passthrough mode */
> @@ -170,18 +172,38 @@ static int fuse_file_passthrough_open(struct inode *inode, struct file *file)
>         if (ff->open_flags & ~FOPEN_PASSTHROUGH_MASK)
>                 return fuse_EIO("conflicting open flags");
>
> -       fb = fuse_passthrough_open(file, ff->args->open_outarg.backing_id);
> -       if (IS_ERR(fb))
> -               return PTR_ERR(fb);
> +       if (!fc->backing_id_64) {
> +               if (outarg->backing_id_64 != 0)
> +                       return fuse_EIO("64 bit backing ID set");
> +
> +               backing_id = outarg->backing_id;
> +               if (backing_id <= 0)
> +                       return fuse_EIO("invalid backing ID");
> +       } else {
> +               if (outarg->backing_id != 0)
> +                       return fuse_EIO("32 bit backing ID set");
> +
> +               backing_id = outarg->backing_id_64;
> +       }
> +       fb = fuse_backing_lookup(fc, backing_id);
> +       if (!fb)
> +               return fuse_EIO("backing not found");
> +
> +       err = fuse_passthrough_open(file, fb);
> +       if (err)
> +               goto backing_put;
>
>         /* First passthrough file open denies caching inode io mode */
>         err = fuse_file_uncached_io_open(inode, ff, fb);
> -       if (!err)
> -               return 0;
> +       if (err)
> +               goto passthrough_release;
> +
> +       return 0;
>
> +passthrough_release:
>         fuse_passthrough_release(ff, fb);
> +backing_put:
>         fuse_backing_put(fb);
> -
>         return err;
>  }
>
> diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
> index 313c8d7ffc09..4894842ad6d0 100644
> --- a/fs/fuse/passthrough.c
> +++ b/fs/fuse/passthrough.c
> @@ -150,41 +150,25 @@ ssize_t fuse_passthrough_mmap(struct file *file, struct vm_area_struct *vma)
>
>  /*
>   * Setup passthrough to a backing file.
> - *
> - * Returns an fb object with elevated refcount to be stored in fuse inode.
>   */
> -struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id)
> +int fuse_passthrough_open(struct file *file, struct fuse_backing *fb)
>  {
>         struct fuse_file *ff = file->private_data;
> -       struct fuse_conn *fc = ff->fm->fc;
> -       struct fuse_backing *fb = NULL;
>         struct file *backing_file;
>
> -       if (fc->backing_id_64)
> -               return ERR_PTR(fuse_EIO("incompatible backing version"));
> -
> -       if (backing_id <= 0)
> -               return ERR_PTR(fuse_EIO("invalid backing_id"));
> -
> -       fb = fuse_backing_lookup(fc, backing_id);
> -       if (!fb)
> -               return ERR_PTR(fuse_EIO("backing not found"));
> -
>         /* Allocate backing file per fuse file to store fuse path */
>         backing_file = backing_file_open(file, file->f_flags,
>                                          &fb->file->f_path, fb->cred);
> -       if (IS_ERR(backing_file)) {
> -               fuse_backing_put(fb);
> -               return ERR_PTR(fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file)));
> -       }
> +       if (IS_ERR(backing_file))
> +               return fuse_EIO("failed to open backing file (%ld)", PTR_ERR(backing_file));
>
>         ff->passthrough = backing_file;
>         ff->cred = get_cred(fb->cred);
>
> -       pr_debug("%s: backing_id=%d, fb=0x%p, backing_file=0x%p\n", __func__,
> -                backing_id, fb, ff->passthrough);
> +       pr_debug("%s: backing_id=%llu, fb=0x%p, backing_file=0x%p\n", __func__,
> +                fb->backing_id, fb, ff->passthrough);
>
> -       return fb;
> +       return 0;
>  }
>
>  void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb)
> diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h
> index c88d7094af44..60fb2add5e12 100644
> --- a/include/uapi/linux/fuse.h
> +++ b/include/uapi/linux/fuse.h
> @@ -254,6 +254,7 @@
>   *  - add FUSE_PASSTHROUGH_V2
>   *  - add FUSE_DEV_IOC_BACKING_CREATE, struct fuse_backing_create_in
>   *  - add FUSE_NOTIFY_BACKING_REMOVE, struct fuse_notify_backing_remove_out
> + *  - add backing_id_64 to fuse_open_out
>   */
>
>  #ifndef _LINUX_FUSE_H
> @@ -841,6 +842,7 @@ struct fuse_open_out {
>         uint64_t        fh;
>         uint32_t        open_flags;
>         int32_t         backing_id;
> +       uint64_t        backing_id_64;
>  };
>
>  struct fuse_release_in {
> --
> 2.54.0
>

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID
  2026-10-01 17:07   ` Amir Goldstein
@ 2026-10-01 18:58     ` Amir Goldstein
  2026-10-05 13:33       ` Miklos Szeredi
  0 siblings, 1 reply; 29+ messages in thread
From: Amir Goldstein @ 2026-10-01 18:58 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: fuse-devel, John Groves, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Thu, Oct 1, 2026 at 7:07 PM Amir Goldstein <amir73il@gmail.com> wrote:
>
> On Thu, Oct 1, 2026 at 5:09 PM Miklos Szeredi <mszeredi@redhat.com> wrote:
> >
> > Add support for server allocated 64-bit backing IDs alongside the
> > existing kernel allocated 32-bit IDs.
> >
> > Opened with FUSE_DEV_IOC_BACKING_OPEN, but the backing ID is sent via
> > fuse_backing_map.backing_id, with FUSE_BACKING_ID_64 being set in .flags.
>
> Description outdated.
>
> >
> > Closed with FUSE_NOTIFY_BACKING_REMOVE.  Since close only provides the
> > backing ID, not the file descriptor, it doesn't have to be done with an
> > ioctl.
> >
> > All 64 bit values are valid.
>
> Can you reserve id 0? or have a reason to need it?
>
> >
> > Signed-off-by: Miklos Szeredi <mszeredi@redhat.com>
> > ---
> >  fs/fuse/backing.c         | 180 ++++++++++++++++++++++++++++----------
> >  fs/fuse/dev.c             |  21 +++++
> >  fs/fuse/dev.h             |   3 +
> >  fs/fuse/fuse_i.h          |  21 ++++-
> >  fs/fuse/inode.c           |   6 +-
> >  fs/fuse/notify.c          |  28 ++++++
> >  fs/fuse/passthrough.c     |   3 +
> >  include/uapi/linux/fuse.h |  22 ++++-
> >  8 files changed, 231 insertions(+), 53 deletions(-)
> >
> > diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c
> > index 433fa3098d71..3c879df7989c 100644
> > --- a/fs/fuse/backing.c
> > +++ b/fs/fuse/backing.c
> > @@ -9,6 +9,7 @@
> >  #include "fuse_i.h"
> >
> >  #include <linux/file.h>
> > +#include <linux/rhashtable.h>
> >
> >  static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb)
> >  {
> > @@ -48,6 +49,8 @@ static int fuse_backing_id_alloc(struct fuse_conn *fc, struct fuse_backing *fb)
> >         id = idr_alloc_cyclic(&fc->backing_files_map, fb, 1, 0, GFP_ATOMIC);
> >         spin_unlock(&fc->lock);
> >         idr_preload_end();
> > +       if (id > 0)
> > +               fb->backing_id = id;
> >
> >         WARN_ON_ONCE(id == 0);
> >         return id;
> > @@ -61,81 +64,131 @@ static struct fuse_backing *fuse_backing_id_remove(struct fuse_conn *fc,
> >         spin_lock(&fc->lock);
> >         fb = idr_remove(&fc->backing_files_map, id);
> >         spin_unlock(&fc->lock);
> > +       if (fb)
> > +               fb->backing_id = 0;
> >
> >         return fb;
> >  }
> >
> > -static int fuse_backing_id_free(int id, void *p, void *data)
> > -{
> > -       struct fuse_backing *fb = p;
> > +static const struct rhashtable_params fuse_backing_prm = {
> > +       .head_offset = offsetof(struct fuse_backing, hash_node),
> > +       .key_offset = offsetof(struct fuse_backing, backing_id),
> > +       .key_len = sizeof_field(struct fuse_backing, backing_id),
> > +};
> >
> > -       WARN_ON_ONCE(refcount_read(&fb->count) != 1);
> > -       fuse_backing_free(fb);
> > -       return 0;
> > +static int fuse_backing_add_64(struct fuse_conn *fc, struct fuse_backing *fb)
> > +{
> > +       return rhashtable_insert_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
> >  }
> >
> > -void fuse_backing_files_free(struct fuse_conn *fc)
> > +int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id)
> >  {
> > -       idr_for_each(&fc->backing_files_map, fuse_backing_id_free, NULL);
> > -       idr_destroy(&fc->backing_files_map);
> > +       struct fuse_backing *fb;
> > +       int err;
> > +
> > +       if (!fc->backing_id_64)
> > +               return -EINVAL;
> > +
> > +       scoped_guard(spinlock, &fc->lock) {
> > +               fb = rhashtable_lookup_fast(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
> > +               if (!fb)
> > +                       return -ENOENT;
> > +
> > +               err = rhashtable_remove_fast(&fc->backing_64_ht, &fb->hash_node, fuse_backing_prm);
> > +               WARN_ON(err);
> > +       }
> > +       fb->backing_id = 0;
> > +       fuse_backing_put(fb);
> > +
> > +       return 0;
> >  }
> >
> > -int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
> > +static struct fuse_backing *fuse_backing_new(struct fuse_conn *fc, int fd)
> >  {
> > -       struct file *file;
> > +       struct fuse_backing *fb;
> >         struct super_block *backing_sb;
> > -       struct fuse_backing *fb = NULL;
> > -       int res;
> > -
> > -       pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
> > +       struct file *file;
> >
> >         /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> > -       res = -EPERM;
> >         if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> > -               goto out;
> > +               return ERR_PTR(-EPERM);
> >
> > -       res = -EINVAL;
> > -       if (map->flags || map->padding)
> > -               goto out;
> > +       CLASS(fd_raw, f)(fd);
> > +       if (fd_empty(f))
> > +               return ERR_PTR(-EBADF);
> >
> > -       file = fget_raw(map->fd);
> > -       res = -EBADF;
> > -       if (!file)
> > -               goto out;
> > +       file = fd_file(f);
> >
> >         /* read/write/splice/mmap passthrough only relevant for regular files */
> > -       res = d_is_dir(file->f_path.dentry) ? -EISDIR : -EINVAL;
> >         if (!d_is_reg(file->f_path.dentry))
> > -               goto out_fput;
> > +               return d_is_dir(file->f_path.dentry) ? ERR_PTR(-EISDIR) : ERR_PTR(-EINVAL);
> >
> >         backing_sb = file_inode(file)->i_sb;
> > -       res = -ELOOP;
> >         if (backing_sb->s_stack_depth >= fc->max_stack_depth)
> > -               goto out_fput;
> > +               return ERR_PTR(-ELOOP);
> >
> >         fb = kmalloc_obj(struct fuse_backing);
> > -       res = -ENOMEM;
> >         if (!fb)
> > -               goto out_fput;
> > +               return ERR_PTR(-ENOMEM);
> >
> > -       fb->file = file;
> > +       fb->file = get_file(file);
> >         fb->cred = get_current_cred();
> >         refcount_set(&fb->count, 1);
> >
> > -       res = fuse_backing_id_alloc(fc, fb);
> > -       if (res < 0) {
> > -               fuse_backing_free(fb);
> > -               fb = NULL;
> > +       return fb;
> > +}
> > +
> > +int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map)
> > +{
> > +       struct fuse_backing *fb;
> > +       int res;
> > +
> > +       res = -EINVAL;
> > +       if (map->padding || map->spare[0] || map->spare[1])
>
> Please reserve backing_id 0
>
> > +               goto out;
> > +
> > +       if (!fc->backing_id_64)
> > +               goto out;
> > +
> > +       fb = fuse_backing_new(fc, map->fd);
> > +       res = PTR_ERR(fb);
> > +       if (!IS_ERR(fb)) {
> > +               fb->backing_id = map->backing_id;
> > +               res = fuse_backing_add_64(fc, fb);
> > +               if (res < 0)
> > +                       fuse_backing_free(fb);
> >         }
> > +out:
> > +       return res;
> > +}
> > +
> > +int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map)
> > +{
> > +       struct fuse_backing *fb = NULL;
> > +       int res;
> > +
> > +       pr_debug("%s: fd=%d flags=0x%x\n", __func__, map->fd, map->flags);
> >
> > +       res = -EINVAL;
> > +       if (map->flags || map->padding)
> > +               goto out;
> > +
> > +       if (fc->backing_id_64)
> > +               goto out;
> > +
> > +       fb = fuse_backing_new(fc, map->fd);
> > +       res = PTR_ERR(fb);
> > +       if (!IS_ERR(fb)) {
> > +               res = fuse_backing_id_alloc(fc, fb);
> > +               if (res < 0) {
> > +                       fuse_backing_free(fb);
> > +                       fb = NULL;
> > +               }
> > +       }
> >  out:
> >         pr_debug("%s: fb=0x%p, ret=%i\n", __func__, fb, res);
> >
> >         return res;
> > -
> > -out_fput:
> > -       fput(file);
> > -       goto out;
> >  }
> >
> >  int fuse_backing_close(struct fuse_conn *fc, int backing_id)
> > @@ -145,6 +198,9 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
> >
> >         pr_debug("%s: backing_id=%d\n", __func__, backing_id);
> >
> > +       if (fc->backing_id_64)
> > +               return -EINVAL;
> > +
> >         /* TODO: relax CAP_SYS_ADMIN once backing files are visible to lsof */
> >         err = -EPERM;
> >         if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
> > @@ -167,14 +223,48 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id)
> >         return err;
> >  }
> >
> > -struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id)
> > +struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id)
> >  {
> >         struct fuse_backing *fb;
> >
> > -       rcu_read_lock();
> > -       fb = idr_find(&fc->backing_files_map, backing_id);
> > -       fb = fuse_backing_get(fb);
> > -       rcu_read_unlock();
> > +       guard(rcu)();
> > +       if (!fc->backing_id_64)
> > +               fb = idr_find(&fc->backing_files_map, backing_id);
> > +       else
> > +               fb = rhashtable_lookup(&fc->backing_64_ht, &backing_id, fuse_backing_prm);
> >
> > -       return fb;
> > +       return fuse_backing_get(fb);
> > +}
> > +
> > +static void fuse_backing_check_free(struct fuse_backing *fb)
> > +{
> > +       WARN_ON_ONCE(refcount_read(&fb->count) != 1);
> > +       fuse_backing_free(fb);
> > +}
> > +
> > +static int fuse_backing_idr_free(int id, void *p, void *data)
> > +{
> > +       fuse_backing_check_free(p);
> > +       return 0;
> > +}
> > +
> > +static void fuse_backing_rht_free(void *p, void *data)
> > +{
> > +       fuse_backing_check_free(p);
> > +}
> > +
> > +void fuse_backing_files_free(struct fuse_conn *fc)
> > +{
> > +       if (fc->backing_id_64) {
> > +               rhashtable_free_and_destroy(&fc->backing_64_ht, fuse_backing_rht_free, NULL);
> > +       } else {
> > +               idr_for_each(&fc->backing_files_map, fuse_backing_idr_free, NULL);
> > +               idr_destroy(&fc->backing_files_map);
> > +       }
> > +}
> > +
> > +void fuse_backing_files_init_64(struct fuse_conn *fc)
> > +{
> > +       rhashtable_init(&fc->backing_64_ht, &fuse_backing_prm);
> > +       fc->backing_id_64 = true;
> >  }
> > diff --git a/fs/fuse/dev.c b/fs/fuse/dev.c
> > index 2d7ee5498f1c..8b63a08a26b6 100644
> > --- a/fs/fuse/dev.c
> > +++ b/fs/fuse/dev.c
> > @@ -2339,6 +2339,24 @@ static long fuse_dev_ioctl_backing_open(struct file *file,
> >         return fuse_backing_open(fud->chan->conn, &map);
> >  }
> >
> > +static long fuse_dev_ioctl_backing_create(struct file *file,
> > +                                         struct fuse_backing_create_in __user *argp)
> > +{
> > +       struct fuse_dev *fud = fuse_get_dev(file);
> > +       struct fuse_backing_create_in map;
> > +
> > +       if (IS_ERR(fud))
> > +               return PTR_ERR(fud);
> > +
> > +       if (!IS_ENABLED(CONFIG_FUSE_PASSTHROUGH))
> > +               return -EOPNOTSUPP;
> > +
> > +       if (copy_from_user(&map, argp, sizeof(map)))
> > +               return -EFAULT;
> > +
> > +       return fuse_backing_open_64(fud->chan->conn, &map);
> > +}
> > +
> >  static long fuse_dev_ioctl_backing_close(struct file *file, __u32 __user *argp)
> >  {
> >         struct fuse_dev *fud = fuse_get_dev(file);
> > @@ -2379,6 +2397,9 @@ static long fuse_dev_ioctl(struct file *file, unsigned int cmd,
> >         case FUSE_DEV_IOC_BACKING_OPEN:
> >                 return fuse_dev_ioctl_backing_open(file, argp);
> >
> > +       case FUSE_DEV_IOC_BACKING_CREATE:
> > +               return fuse_dev_ioctl_backing_create(file, argp);
> > +
> >         case FUSE_DEV_IOC_BACKING_CLOSE:
> >                 return fuse_dev_ioctl_backing_close(file, argp);
> >
> > diff --git a/fs/fuse/dev.h b/fs/fuse/dev.h
> > index f6c47ae0395b..dbffd5bed2f5 100644
> > --- a/fs/fuse/dev.h
> > +++ b/fs/fuse/dev.h
> > @@ -14,6 +14,7 @@ struct fuse_dev;
> >  struct fuse_args;
> >  struct fuse_copy_state;
> >  struct fuse_backing_map;
> > +struct fuse_backing_create_in;
> >  struct file;
> >  struct folio;
> >  enum fuse_notify_code;
> > @@ -87,6 +88,8 @@ int fuse_notify(struct fuse_conn *fc, enum fuse_notify_code code,
> >
> >  int fuse_backing_open(struct fuse_conn *fc, struct fuse_backing_map *map);
> >  int fuse_backing_close(struct fuse_conn *fc, int backing_id);
> > +int fuse_backing_open_64(struct fuse_conn *fc, struct fuse_backing_create_in *map);
> > +int fuse_backing_close_64(struct fuse_conn *fc, u64 backing_id);
> >
> >  int fuse_copy_one(struct fuse_copy_state *cs, void *val, unsigned size);
> >  int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop,
> > diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h
> > index 3dd4dff24c50..003ef3c35d7a 100644
> > --- a/fs/fuse/fuse_i.h
> > +++ b/fs/fuse/fuse_i.h
> > @@ -32,6 +32,7 @@
> >  #include <linux/pid_namespace.h>
> >  #include <linux/refcount.h>
> >  #include <linux/user_namespace.h>
> > +#include <linux/rhashtable-types.h>
> >
> >  /** Default max number of pages that can be used in a single read request */
> >  #define FUSE_DEFAULT_MAX_PAGES_PER_REQ 32
> > @@ -92,7 +93,8 @@ struct fuse_submount_lookup {
> >  struct fuse_backing {
> >         struct file *file;
> >         const struct cred *cred;
> > -
> > +       u64 backing_id;
> > +       struct rhash_head hash_node;
> >         /* refcount */
> >         refcount_t count;
> >         struct rcu_head rcu;
> > @@ -689,6 +691,9 @@ struct fuse_conn {
> >         /** @init_security: Initialize security xattrs when creating a new inode */
> >         unsigned int init_security:1;
> >
> > +       /** Backing ID is 64 bit and allocated by the server */
> > +       bool backing_id_64:1;
> > +
> >         /**
> >          * @create_supp_group: Add supplementary group info when creating
> >          * a new inode
> > @@ -770,8 +775,14 @@ struct fuse_conn {
> >         struct fuse_sync_bucket __rcu *curr_bucket;
> >
> >  #ifdef CONFIG_FUSE_PASSTHROUGH
> > -       /** @backing_files_map: IDR for backing files ids */
> > -       struct idr backing_files_map;
> > +       /* Selected by backing_id_64 */
> > +       union {
> > +               /** @backing_files_map: IDR for backing files ids */
> > +               struct idr backing_files_map;
> > +
> > +               /** @backing_64_ht: 64 bit ID lookup hash table */
> > +               struct rhashtable backing_64_ht;
> > +       };
> >  #endif
> >  };
> >
> > @@ -1270,7 +1281,7 @@ void fuse_file_release(struct inode *inode, struct fuse_file *ff,
> >  /* backing.c */
> >  #ifdef CONFIG_FUSE_PASSTHROUGH
> >  void fuse_backing_put(struct fuse_backing *fb);
> > -struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id);
> > +
> >  #else
> >
> >  static inline void fuse_backing_put(struct fuse_backing *fb)
> > @@ -1278,7 +1289,9 @@ static inline void fuse_backing_put(struct fuse_backing *fb)
> >  }
> >  #endif
> >
> > +struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, u64 backing_id);
> >  void fuse_backing_files_init(struct fuse_conn *fc);
> > +void fuse_backing_files_init_64(struct fuse_conn *fc);
> >  void fuse_backing_files_free(struct fuse_conn *fc);
> >
> >  /* passthrough.c */
> > diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c
> > index cbb10e19e7e8..bb76bddfc241 100644
> > --- a/fs/fuse/inode.c
> > +++ b/fs/fuse/inode.c
> > @@ -1408,13 +1408,15 @@ static void process_init_reply(struct fuse_args *args, int error)
> >                          * them together.
> >                          */
> >                         if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) &&
> > -                           (flags & FUSE_PASSTHROUGH) &&
> > +                           (flags & (FUSE_PASSTHROUGH | FUSE_PASSTHROUGH_V2)) &&
> >                             arg->max_stack_depth > 0 &&
> >                             arg->max_stack_depth <= FILESYSTEM_MAX_STACK_DEPTH &&
> >                             !(flags & FUSE_WRITEBACK_CACHE))  {
> >                                 fc->passthrough = 1;
> >                                 fc->max_stack_depth = arg->max_stack_depth;
> >                                 fm->sb->s_stack_depth = arg->max_stack_depth;
> > +                               if (flags & FUSE_PASSTHROUGH_V2)
> > +                                       fuse_backing_files_init_64(fc);
>
> It would feel a little less hacky if we do:
>     else
>           fuse_backing_files_init_32(fc);
>
> and set
>          fc->passthrough = 1;
>
> inside both of them to gate against fuse_backing_files_free()
> called without processing init response,
> or without either FUSE_PASSTHROUGH in the init response.

Also we should not enable passthrough if rhashtable_init has failed.

Thanks,
Amir.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message
  2026-10-01 16:32   ` Amir Goldstein
@ 2026-10-05  9:46     ` Miklos Szeredi
  0 siblings, 0 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-05  9:46 UTC (permalink / raw)
  To: Amir Goldstein
  Cc: Miklos Szeredi, fuse-devel, John Groves, Darrick J . Wong,
	Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Thu, 1 Oct 2026 at 18:39, Amir Goldstein <amir73il@gmail.com> wrote:
>
> On Thu, Oct 1, 2026 at 5:09 PM Miklos Szeredi <mszeredi@redhat.com> wrote:

> > @@ -118,11 +111,9 @@ static int fuse_file_uncached_io_open(struct inode *inode,
> >                                       struct fuse_backing *fb)
> >  {
> >         struct fuse_inode *fi = get_fuse_inode(inode);
> > -       int err;
> >
> > -       err = fuse_inode_uncached_io_start(fi, fb);
> > -       if (err)
> > -               return err;
> > +       if (!fuse_inode_uncached_io_start(fi, fb))
> > +               return fuse_EIO("failed to start uncached I/O");
>
> Why lose the internal error code?
> If server implementation is wrong that is a good hint of what went wrong.
> ETXTBSY is trying to mix cache mode and passthrough mode.
> EBUSY is trying to passthrough to two different backing inodes.

I'll move the error codes back, and turn them into fuse_EIO with an
appropriate message in the caller.

> > @@ -234,22 +228,10 @@ int fuse_file_io_open(struct file *file, struct inode *inode)
> >
> >         if (ff->open_flags & FOPEN_PASSTHROUGH)
> >                 err = fuse_file_passthrough_open(inode, file);
> > -       else
> > -               err = fuse_file_cached_io_open(inode, ff);
> > -       if (err)
> > -               goto fail;
> > +       else if (!fuse_file_cached_io_open(inode, ff))
> > +               err = fuse_EIO("conflicting passthrough open");
>
> Same here.
> If we can keep this bit of information and print it - it is helpful IMO,
> even if we don't go all the way to annotate every error case with
> a pretty log.

There is just one way to fail fuse_file_cached_io_open() so I think
this is okay.

Thanks,
Miklos

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID
  2026-10-01 18:58     ` Amir Goldstein
@ 2026-10-05 13:33       ` Miklos Szeredi
  2026-10-06 21:24         ` Amir Goldstein
  0 siblings, 1 reply; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-05 13:33 UTC (permalink / raw)
  To: Amir Goldstein
  Cc: Miklos Szeredi, fuse-devel, John Groves, Darrick J . Wong,
	Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Thu, 1 Oct 2026 at 20:59, Amir Goldstein <amir73il@gmail.com> wrote:

> > > All 64 bit values are valid.
> >
> > Can you reserve id 0? or have a reason to need it?

I thought that removing all restrictions on the value would be nice.
But okay, zero value we can live without.

> > It would feel a little less hacky if we do:
> >     else
> >           fuse_backing_files_init_32(fc);
> >
> > and set
> >          fc->passthrough = 1;
> >
> > inside both of them to gate against fuse_backing_files_free()
> > called without processing init response,
> > or without either FUSE_PASSTHROUGH in the init response.

Bigger problem is that IOC_BACKING_OPEN can be done before or during
INIT.  So we need to take care of that case and the race as well.
Fixed by:

 - leaving the idr initialization fuse_conn_init()

 - adding fc->backing_id_32 flag which is set by fuse_backing_open()
after chekcking that fc->backing_id_64 is not set.  All under fc->lock

 - in process_init_reply() checking fc->backing_id_32 before setting
fc->backing_id_64 and failing the connection if so.  Also under
fc->lock

 - in fuse_dev_ioctl_backing_create() if fch->initialized is not set
return ENOTCONN

> Also we should not enable passthrough if rhashtable_init has failed.

rhashtable_init() can't fail if parameters are correct since it falls
back to __GFP_NOFAIL() for allocation.

Thanks,
Miklos

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 0/8] fuse: DAX device based extent maps (famfs)
  2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
                   ` (7 preceding siblings ...)
  2026-10-01 15:07 ` [PATCH v2 8/8] fuse: add support for striped backing Miklos Szeredi
@ 2026-10-05 23:27 ` John Groves
  2026-10-06  9:48   ` Miklos Szeredi
  8 siblings, 1 reply; 29+ messages in thread
From: John Groves @ 2026-10-05 23:27 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: fuse-devel, Amir Goldstein, Darrick J . Wong, Vishal Verma,
	Dave Jiang, Alison Schofield, nvdimm, linux-cxl, linux-fsdevel,
	david, willy, brauner

On 26/10/01 05:07PM, Miklos Szeredi wrote:
> This is the FUSIfication of the famfs-fuse patchset.

This series (v1) was sent on the day I left on a vacation from which I
returned at the end of last week.

I feel I should acknowledge this and assure everybody that I'm working
through it. This really runs my famfs/fuse patches through a 
chipper/shredder, and it is gonna take me some time to review it, 
understand it and test it.

A few comments below, and specific patch comments should start trickling out.

One very high level quip though: If you're comfortable with this, go ahead 
and merge it. We can fix bugs and performance problems, if any, in due 
course. I'm serious about that.

> 
>  - extent maps are assigned to a backing ID
> 
>  - 64 bit backing ID's managed by server

I thought we agreed you weren't gonna do these two, but if it does no harm,
whatever.

> 
>  - maps can be linear or cyclic (striped)

Please correct me if any of my analysis below is wrong...

Here what you've done is remove my explicitly interleaved extent lists,
which have a header containing a chunk size and a set of strip extents (which
are basically fully self-describing), and replace is with just simple extent 
lists that have an optional CYCLIC flag in the header. If the flag is set on 
an extent list, (and let's say the strip extents need to be 1GiB and the 
chunk size is 2MiB), each CYCLIC extent gets its size set to the chunk size 
and the kernel trusts that the actual size allocated per stirp is actually 
sufficient that i_size will not cycle past the end of the *actually* 
allocated size of each strip extent.

This resembles an idea I considered, to try to make the metadata *look* 
simpler while still carrying the info that needs to be conveyed. But this 
implementation feels pretty janky to me, and it should at least be 
meticulously documented. Better might be to put the chunk_size in the header,
and then you arguably don't need the CYCLIC flag, since non-zero chunk_size
means the same thing. But actually, you can accept multiple CYCLIC strip
lists if you set chunk_size and then set the CYCLIC flag on the first of
each strip set.

I see the kernel-side rejects these s/size/chunk_size/ strip extents if they
are not all the same size. Why not put the chunk size in the header and let
the extents honestly describe the offset and range that they cover? 
The current design can't test for sufficient strip extent sizes, although 
it replaces code that did verify this.

> 
>  - API supports nesting, but not implemented yet (i.e. multiple striped
>    extents per inode, a-la famfs)
> 
> libfuse/famfs trees for testing:
> 
>  https://github.com/szmi/libfuse.git#extent-map-v2
>  https://github.com/szmi/famfs.git#extent-map-v2
> 
> This patchset:
> 
>  git://git.kernel.org/pub/scm/linux/kernel/git/mszeredi/fuse.git#extent-map-v2
> 
> Passes famfs smoke test on emulated dax device.

Thank you for doing that. Prior to this there have been many strong opinions
but none of the holders of said opinions built or tested famfs.

However, the smoke tests passing isn't fully legit, because the patch to the
famfs fuse server passes the entire strip size on CYCLIC mappings, meaning
it just produces a concatenation of ranges - which does not achieve the 
interleaving goal at all. That case degenerates to simple extents.

If this is the new design for passing cycling maps into the kernel, a test
will be added to catch this - but it's brand new and didn't get a test added,
so it wasn't caught automatically.

> 
> Comments are welcome.

One more... This patch inserts an rbtree lookup into the fault path via the
very-performance-critical fuse_ext_map_iomap_begin() path. This is 
performance-critical because it handles TLB/PTE/PMD faults when the memory
is never sparse (no I/O to amortize over). The code you replaced did several
things to avoid this sort of chasing:

- With interleaved extent lists, strip and strip-offset can be calculated
  in order 1 (and then strip overflow can be checked and avoided)
- Simple extent lists were checked for constant extent sizes; if constant,
  we can just index into the list (also order 1)

Why do I want to avoid an rbtree lookup or the bpf nonsense in the fault path?
"Because this is memory and it must run at memory speeds" --yours truly

To be specific: each chase/dereference that results in an L3 cache miss adds
a stall of (1 * MEMORY_LATENCY) to the fault time. This is the good reason 
why I didn't do any chasing in the fault path of the famfs code that you're
replacing.

OK, one more. You're using the rbtree for indexing by file offset, which is
a zero-based non-sparse range in this case. If this could not be avoided 
(which I think it can) you should use an xarray or radix tree for that...
right?

> 
> Thanks,
> Miklos
> 
> ---
> Changes since v1:
> 
>  - moved 3 prep patches to fuse.git#for-next
>  - introduce FUSE_PASSTHROUGH_V2 (Amir)
>  - lots of small fixes and cleanups (Amir, Sashiko)
> 
> ---
> John Groves (1):
>   dax: replace exported dax_dev_get() with non-allocating dax_dev_find()

Golly, it seems like there's more than 1 patch worth of content from me here,
but the chipper/shredder might meet the letter of copyright law in some 
jurisdictions...

> 
> Miklos Szeredi (7):
>   fuse: add helpers for EIO return value with kernel message
>   fuse: support 64 bit, server allocated backing ID
>   fuse: support opening 64 bit backing ID
>   fuse: add support for opening dax device as backing
>   fuse: add extent map data structure
>   fuse: add extent map I/O support
>   fuse: add support for striped backing
> 
>  drivers/dax/super.c       |  38 ++++-
>  fs/fuse/Makefile          |   2 +-
>  fs/fuse/backing.c         | 292 ++++++++++++++++++++++++++++++-------
>  fs/fuse/dev.c             |  21 +++
>  fs/fuse/dev.h             |   3 +
>  fs/fuse/dir.c             |   3 +-
>  fs/fuse/ext_map.c         | 299 ++++++++++++++++++++++++++++++++++++++
>  fs/fuse/file.c            |  16 +-
>  fs/fuse/fuse_i.h          |  72 +++++++--
>  fs/fuse/inode.c           |  17 ++-
>  fs/fuse/iomode.c          | 108 +++++++-------
>  fs/fuse/notify.c          |  75 ++++++++++
>  fs/fuse/passthrough.c     |  63 ++++----
>  include/linux/dax.h       |   6 +-
>  include/uapi/linux/fuse.h |  53 ++++++-
>  15 files changed, 904 insertions(+), 164 deletions(-)
>  create mode 100644 fs/fuse/ext_map.c
> 
> -- 
> 2.54.0
> 

John


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 0/8] fuse: DAX device based extent maps (famfs)
  2026-10-05 23:27 ` [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) John Groves
@ 2026-10-06  9:48   ` Miklos Szeredi
  2026-10-08 23:00     ` John Groves
  0 siblings, 1 reply; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-06  9:48 UTC (permalink / raw)
  To: John Groves
  Cc: Miklos Szeredi, fuse-devel, Amir Goldstein, Darrick J . Wong,
	Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl,
	linux-fsdevel, david, willy, brauner

On Tue, 6 Oct 2026 at 01:37, John Groves <john@groves.net> wrote:

> I see the kernel-side rejects these s/size/chunk_size/ strip extents if they
> are not all the same size. Why not put the chunk size in the header and let
> the extents honestly describe the offset and range that they cover?
> The current design can't test for sufficient strip extent sizes, although
> it replaces code that did verify this.

Not adding chunk size to the header is intentional to keep the API as
simple as possible.

> However, the smoke tests passing isn't fully legit, because the patch to the
> famfs fuse server passes the entire strip size on CYCLIC mappings, meaning
> it just produces a concatenation of ranges - which does not achieve the
> interleaving goal at all. That case degenerates to simple extents.

Are you talking about the case where there are multiple simple extents
each with a different striped mapping?

That case is supported by the API, but not the implementation (since
there was not test case).  If you do a test case, I'll add the
implementation.

> To be specific: each chase/dereference that results in an L3 cache miss adds
> a stall of (1 * MEMORY_LATENCY) to the fault time. This is the good reason
> why I didn't do any chasing in the fault path of the famfs code that you're
> replacing.

Please benchmark against the in-kernel implementation.  If there's
significant (meaning it may have a real effect on a real-life
workload) performance hit from using the rb-tree, then we can add an
optimization.  I don't want to add complexity without actually being
able to measure the gain.

> OK, one more. You're using the rbtree for indexing by file offset, which is
> a zero-based non-sparse range in this case. If this could not be avoided
> (which I think it can) you should use an xarray or radix tree for that...
> right?

Unlike xarray, the rb_tree API is one I'm familiar with. But nothing
prevents switching to xarray if that's a better choice.

Thanks,
Miklos


>
> >
> > Thanks,
> > Miklos
> >
> > ---
> > Changes since v1:
> >
> >  - moved 3 prep patches to fuse.git#for-next
> >  - introduce FUSE_PASSTHROUGH_V2 (Amir)
> >  - lots of small fixes and cleanups (Amir, Sashiko)
> >
> > ---
> > John Groves (1):
> >   dax: replace exported dax_dev_get() with non-allocating dax_dev_find()
>
> Golly, it seems like there's more than 1 patch worth of content from me here,
> but the chipper/shredder might meet the letter of copyright law in some
> jurisdictions...
>
> >
> > Miklos Szeredi (7):
> >   fuse: add helpers for EIO return value with kernel message
> >   fuse: support 64 bit, server allocated backing ID
> >   fuse: support opening 64 bit backing ID
> >   fuse: add support for opening dax device as backing
> >   fuse: add extent map data structure
> >   fuse: add extent map I/O support
> >   fuse: add support for striped backing
> >
> >  drivers/dax/super.c       |  38 ++++-
> >  fs/fuse/Makefile          |   2 +-
> >  fs/fuse/backing.c         | 292 ++++++++++++++++++++++++++++++-------
> >  fs/fuse/dev.c             |  21 +++
> >  fs/fuse/dev.h             |   3 +
> >  fs/fuse/dir.c             |   3 +-
> >  fs/fuse/ext_map.c         | 299 ++++++++++++++++++++++++++++++++++++++
> >  fs/fuse/file.c            |  16 +-
> >  fs/fuse/fuse_i.h          |  72 +++++++--
> >  fs/fuse/inode.c           |  17 ++-
> >  fs/fuse/iomode.c          | 108 +++++++-------
> >  fs/fuse/notify.c          |  75 ++++++++++
> >  fs/fuse/passthrough.c     |  63 ++++----
> >  include/linux/dax.h       |   6 +-
> >  include/uapi/linux/fuse.h |  53 ++++++-
> >  15 files changed, 904 insertions(+), 164 deletions(-)
> >  create mode 100644 fs/fuse/ext_map.c
> >
> > --
> > 2.54.0
> >
>
> John
>
>

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID
  2026-10-05 13:33       ` Miklos Szeredi
@ 2026-10-06 21:24         ` Amir Goldstein
  2026-10-07 12:46           ` Miklos Szeredi
  0 siblings, 1 reply; 29+ messages in thread
From: Amir Goldstein @ 2026-10-06 21:24 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: Miklos Szeredi, fuse-devel, John Groves, Darrick J . Wong,
	Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Mon, Oct 5, 2026 at 3:33 PM Miklos Szeredi <miklos@szeredi.hu> wrote:

> > > It would feel a little less hacky if we do:
> > >     else
> > >           fuse_backing_files_init_32(fc);
> > >
> > > and set
> > >          fc->passthrough = 1;
> > >
> > > inside both of them to gate against fuse_backing_files_free()
> > > called without processing init response,
> > > or without either FUSE_PASSTHROUGH in the init response.
>
> Bigger problem is that IOC_BACKING_OPEN can be done before or during
> INIT.  So we need to take care of that case and the race as well.
> Fixed by:
>
>  - leaving the idr initialization fuse_conn_init()
>
>  - adding fc->backing_id_32 flag which is set by fuse_backing_open()
> after chekcking that fc->backing_id_64 is not set.  All under fc->lock
>
>  - in process_init_reply() checking fc->backing_id_32 before setting
> fc->backing_id_64 and failing the connection if so.  Also under
> fc->lock
>
>  - in fuse_dev_ioctl_backing_create() if fch->initialized is not set
> return ENOTCONN
>

I guess that works, but why the asymmetry between the two modes?

We could check fch->initialized in both _open() and _create()
and then no need for the backing_id_32 flag?

Unless we gate _open() with fch->initialized, then it could fail on
either one of these cases:

                if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
                        return -EPERM;

                if (inode->i_sb->s_stack_depth >= fc->max_stack_depth)
                        return -ELOOP;

Quite randomly.

Maybe I am missing something.

Thanks,
Amir.

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID
  2026-10-06 21:24         ` Amir Goldstein
@ 2026-10-07 12:46           ` Miklos Szeredi
  0 siblings, 0 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-07 12:46 UTC (permalink / raw)
  To: Amir Goldstein
  Cc: Miklos Szeredi, fuse-devel, John Groves, Darrick J . Wong,
	Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl

On Tue, 6 Oct 2026 at 23:24, Amir Goldstein <amir73il@gmail.com> wrote:

> Unless we gate _open() with fch->initialized, then it could fail on
> either one of these cases:
>
>                 if (!fc->passthrough || !capable(CAP_SYS_ADMIN))
>                         return -EPERM;
>
>                 if (inode->i_sb->s_stack_depth >= fc->max_stack_depth)
>                         return -ELOOP;
>
> Quite randomly.
>
> Maybe I am missing something.

No, you are right.

Thanks,
Miklos

^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 0/8] fuse: DAX device based extent maps (famfs)
  2026-10-06  9:48   ` Miklos Szeredi
@ 2026-10-08 23:00     ` John Groves
  2026-10-09 10:39       ` Miklos Szeredi
  0 siblings, 1 reply; 29+ messages in thread
From: John Groves @ 2026-10-08 23:00 UTC (permalink / raw)
  To: Miklos Szeredi
  Cc: Miklos Szeredi, fuse-devel, Amir Goldstein, Darrick J . Wong,
	Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl,
	linux-fsdevel, david, willy, brauner

On 26/10/06 11:48AM, Miklos Szeredi wrote:
> On Tue, 6 Oct 2026 at 01:37, John Groves <john@groves.net> wrote:
> 
> > I see the kernel-side rejects these s/size/chunk_size/ strip extents if they
> > are not all the same size. Why not put the chunk size in the header and let
> > the extents honestly describe the offset and range that they cover?
> > The current design can't test for sufficient strip extent sizes, although
> > it replaces code that did verify this.
> 
> Not adding chunk size to the header is intentional to keep the API as
> simple as possible.

Fewer fields is nice, but it's on the honor system that the server allocated
large enough strips - since strips carry the chunk_size rather than their
actual size when CYCLIC is set. IMO it would be an improvement to add the
chunk_size to the header; then the strip sizes could be validated.

But it's your call, and I won't argue further about this unless something
changes.

> 
> > However, the smoke tests passing isn't fully legit, because the patch to the
> > famfs fuse server passes the entire strip size on CYCLIC mappings, meaning
> > it just produces a concatenation of ranges - which does not achieve the
> > interleaving goal at all. That case degenerates to simple extents.
> 
> Are you talking about the case where there are multiple simple extents
> each with a different striped mapping?

No

I'm saying your patch to the famfs user space, when famfs sends interleaved
file maps, sets CYCLIC but sets strip extent sizes to the actual strip size,
not the chunk_size. So chunking is at the size of the entire strip, which is
to say: the net is a multi-extent file that is not interleaved.

So that's a bug. I've fixed it locally and will share that in the next day
or two. I also needed to refactor your patch the famfs user space because
it disabled ABI compatibility with some older versions of the famfs kernel
that I'm still supporting users with. If you can wait to patch that further
until I share an update, it will be cleaner on my end.

> 
> That case is supported by the API, but not the implementation (since
> there was not test case).  If you do a test case, I'll add the
> implementation.
> 
> > To be specific: each chase/dereference that results in an L3 cache miss adds
> > a stall of (1 * MEMORY_LATENCY) to the fault time. This is the good reason
> > why I didn't do any chasing in the fault path of the famfs code that you're
> > replacing.
> 
> Please benchmark against the in-kernel implementation.  If there's
> significant (meaning it may have a real effect on a real-life
> workload) performance hit from using the rb-tree, then we can add an
> optimization.  I don't want to add complexity without actually being
> able to measure the gain.

I will do that, but it's not a small undertaking. Back to that in a moment.

I have an idea for how to make fault handling order-1 in the normal famfs
cases but fall back to order log-n if necessary (e.g. if not all extents are 
the same size). I'm testing that now and will share it soon if I don't run 
into any significant problems.

I think we can probably agree than things otherwise being equal, order-1
fault handling (or order-1 falling back to order-n if necessary) is 
preferable to order-log-n. If you or anybody doesn't agree with that, I 
hope we can discuss it.

As for benchmarking to "prove" that the rbtree is a problem, I'm sure there
are use cases where it is not a problem (like small virtual daxdevs on
workstations or laptops). What Micron worries about is huge working sets on 
many terabyte disaggregated memory. If the memory is 45TB (the biggest 
system I have "access" to), the data sets exceed the processor cache by a 
substantially higher factor than even in a maxed out server. That means
more L3 cache misses. I'm trying to avoid an approach with built-in chasing
on performance paths, because Micron's experience is that pointer chasing 
becomes L3-cache-miss bound.

Anyway, one workload (and working-set size) that doesn't exhibit significant 
L3 cache miss performance hit does not prove that others won't be negatively
impacted. Proving a negative, etc...

Having said all that, this topic seems to recur and be contentious - we are
working on tests that can demonstrate bottlenecks - but my current ask is
that if I can offer an alternative approach to rbtree that is better on 
paper, can we please use it?

> 
> > OK, one more. You're using the rbtree for indexing by file offset, which is
> > a zero-based non-sparse range in this case. If this could not be avoided
> > (which I think it can) you should use an xarray or radix tree for that...
> > right?
> 
> Unlike xarray, the rb_tree API is one I'm familiar with. But nothing
> prevents switching to xarray if that's a better choice.

That's fair. And I withdraw the xarray suggestion because I have a better
idea - will share that in a day or two, before Monday almost certainly.
Need to run it through all of my CI...

Thanks,
John

<snip>


^ permalink raw reply	[flat|nested] 29+ messages in thread

* Re: [PATCH v2 0/8] fuse: DAX device based extent maps (famfs)
  2026-10-08 23:00     ` John Groves
@ 2026-10-09 10:39       ` Miklos Szeredi
  0 siblings, 0 replies; 29+ messages in thread
From: Miklos Szeredi @ 2026-10-09 10:39 UTC (permalink / raw)
  To: John Groves
  Cc: Miklos Szeredi, fuse-devel, Amir Goldstein, Darrick J . Wong,
	Vishal Verma, Dave Jiang, Alison Schofield, nvdimm, linux-cxl,
	linux-fsdevel, david, willy, brauner

On Fri, 9 Oct 2026 at 01:00, John Groves <john@groves.net> wrote:

> Fewer fields is nice, but it's on the honor system that the server allocated
> large enough strips - since strips carry the chunk_size rather than their
> actual size when CYCLIC is set. IMO it would be an improvement to add the
> chunk_size to the header; then the strip sizes could be validated.

Cyclic maps could easily be generalized to not require a fixed chunk
size.  Giving a different meaning to extent length (strip length)
would break that, it would no longer be a CYCLIC map.  It's an
un-generalization that has dubious value, since the strip length is
implicit in either the file size or the parent extent length.

In other words: make the server validate the strip size, and all will
be good.  The kernel really doesn't need this information.

> So that's a bug. I've fixed it locally and will share that in the next day
> or two. I also needed to refactor your patch the famfs user space because
> it disabled ABI compatibility with some older versions of the famfs kernel
> that I'm still supporting users with. If you can wait to patch that further
> until I share an update, it will be cleaner on my end.

Yeah, thanks.

> Having said all that, this topic seems to recur and be contentious - we are
> working on tests that can demonstrate bottlenecks - but my current ask is
> that if I can offer an alternative approach to rbtree that is better on
> paper, can we please use it?

If it's not adding complexity, then sure.

Thanks,
Miklos

^ permalink raw reply	[flat|nested] 29+ messages in thread

end of thread, other threads:[~2026-10-09 10:40 UTC | newest]

Thread overview: 29+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-01 15:07 [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) Miklos Szeredi
2026-10-01 15:07 ` [PATCH v2 1/8] dax: replace exported dax_dev_get() with non-allocating dax_dev_find() Miklos Szeredi
2026-10-01 15:07 ` [PATCH v2 2/8] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
2026-10-01 15:18   ` sashiko-bot
2026-10-01 16:32   ` Amir Goldstein
2026-10-05  9:46     ` Miklos Szeredi
2026-10-01 15:07 ` [PATCH v2 3/8] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
2026-10-01 15:26   ` sashiko-bot
2026-10-01 17:07   ` Amir Goldstein
2026-10-01 18:58     ` Amir Goldstein
2026-10-05 13:33       ` Miklos Szeredi
2026-10-06 21:24         ` Amir Goldstein
2026-10-07 12:46           ` Miklos Szeredi
2026-10-01 15:07 ` [PATCH v2 4/8] fuse: support opening 64 bit " Miklos Szeredi
2026-10-01 15:22   ` sashiko-bot
2026-10-01 17:09   ` Amir Goldstein
2026-10-01 15:07 ` [PATCH v2 5/8] fuse: add support for opening dax device as backing Miklos Szeredi
2026-10-01 15:30   ` sashiko-bot
2026-10-01 16:07   ` Amir Goldstein
2026-10-01 15:07 ` [PATCH v2 6/8] fuse: add extent map data structure Miklos Szeredi
2026-10-01 15:24   ` sashiko-bot
2026-10-01 15:07 ` [PATCH v2 7/8] fuse: add extent map I/O support Miklos Szeredi
2026-10-01 15:27   ` sashiko-bot
2026-10-01 16:11   ` Amir Goldstein
2026-10-01 15:07 ` [PATCH v2 8/8] fuse: add support for striped backing Miklos Szeredi
2026-10-05 23:27 ` [PATCH v2 0/8] fuse: DAX device based extent maps (famfs) John Groves
2026-10-06  9:48   ` Miklos Szeredi
2026-10-08 23:00     ` John Groves
2026-10-09 10:39       ` Miklos Szeredi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox