From: sashiko-bot@kernel.org
To: "Miklos Szeredi" <mszeredi@redhat.com>
Cc: linux-cxl@vger.kernel.org, nvdimm@lists.linux.dev
Subject: Re: [PATCH v4 8/9] fuse: add extent map I/O support
Date: Thu, 08 Oct 2026 11:36:46 +0000 [thread overview]
Message-ID: <sashiko-outbox-164097@kernel.org> (raw)
In-Reply-To: <20261008112004.1899560-9-mszeredi@redhat.com>
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] fuse: silent data loss on fsync for DAX extent-mapped files
- [High] fuse: ABBA deadlock between pipe_lock and inode lock in splice write
- [High] fuse: concurrent EXTMAP DAX faults and truncate/punch_hole race due to missing lock
--
commit 5f3e8f02d15211d77c4f619bc19ce42ddb86c547
Author: Miklos Szeredi <mszeredi@redhat.com>
fuse: add extent map I/O support
Wire up read, write, splice and mmap operations for extent-mapped files
through the iomap/DAX infrastructure. When a passthrough file is opened
with an EXTMAP backing, I/O is dispatched to the dax devices referenced
by the extent map.
> diff --git a/fs/fuse/ext_map.c b/fs/fuse/ext_map.c
> index afac5de51c802..744940cbf36e1 100644
> --- a/fs/fuse/ext_map.c
> +++ b/fs/fuse/ext_map.c
[ ... ]
> +static vm_fault_t fuse_ext_map_huge_fault(struct vm_fault *vmf, unsigned int order)
> +{
> + struct inode *inode = file_inode(vmf->vma->vm_file);
> + bool write_fault = (vmf->flags & FAULT_FLAG_WRITE) && (vmf->vma->vm_flags & VM_SHARED);
> + vm_fault_t ret;
> + unsigned long pfn;
> +
> + if (!IS_ENABLED(CONFIG_FS_DAX))
> + return VM_FAULT_SIGBUS;
> +
> + if (WARN_ON_ONCE(!IS_DAX(inode)))
> + return VM_FAULT_SIGBUS;
> +
> + if (write_fault) {
> + sb_start_pagefault(inode->i_sb);
> + file_update_time(vmf->vma->vm_file);
> + }
> +
> + filemap_invalidate_lock_shared(inode->i_mapping);
[Severity: High]
Does this allow concurrent truncate or punch_hole operations to race with page
faults?
The page fault handler acquires filemap_invalidate_lock_shared() here,
expecting writers to hold it exclusively during invalidation. However, writer
paths like fuse_do_setattr() and fuse_file_fallocate() check FUSE_IS_VDAX()
to decide whether to acquire the exclusive filemap_invalidate_lock. For EXTMAP
DAX files, FUSE_IS_VDAX() evaluates to false, causing the exclusive lock
acquisition to be skipped. Consequently, page faults are not blocked during
truncate_pagecache() or truncate_pagecache_range().
Can this leave stale DAX entries in the mapping tree, leading to
use-after-free of storage blocks when memory-mapped I/O accesses blocks that
have been freed or punched out?
> +
> + ret = dax_iomap_fault(vmf, order, &pfn, NULL, &fuse_ext_map_iomap_ops);
> + if (ret & VM_FAULT_NEEDDSYNC)
> + ret = dax_finish_sync_fault(vmf, order, pfn);
> +
> + filemap_invalidate_unlock_shared(inode->i_mapping);
> +
> + if (write_fault)
> + sb_end_pagefault(inode->i_sb);
> +
> + return ret;
> +}
[ ... ]
> diff --git a/fs/fuse/file.c b/fs/fuse/file.c
> index b5bef7b9e466e..77d7a54066e47 100644
> --- a/fs/fuse/file.c
> +++ b/fs/fuse/file.c
> @@ -551,6 +551,17 @@ static int fuse_fsync(struct file *file, loff_t start, loff_t end,
> if (fuse_is_bad(inode))
> return -EIO;
>
> + if (IS_DAX(inode) && !fuse_inode_vdax(inode)) {
> + /*
> + * Note: the DAX fault path calls this via dax_finish_sync_fault()
> + * and taking inode lock in that context is prohibited.
> + *
> + * FIXME: need to flush CPU caches for the DAX memory range
> + * FIXME: need to sync metadata to the server
> + */
> + return 0;
[Severity: High]
Can this cause silent data loss on fsync() for DAX extent-mapped files?
When an application calls fsync() on a FUSE DAX extent-mapped file, this
bypasses flushing the CPU caches for the DAX memory range or syncing metadata
to the server, while still returning 0 (success).
Because the operation falsely reports success, applications relying on fsync()
or MAP_SYNC (via dax_finish_sync_fault()) to ensure data durability will not
take fallback actions. Could this lead to silent data loss on a power failure
or crash?
> + }
> +
> inode_lock(inode);
>
> /*
> diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c
> index e9ab1aea34e27..040817ad76e94 100644
> --- a/fs/fuse/passthrough.c
> +++ b/fs/fuse/passthrough.c
[ ... ]
> @@ -74,10 +73,13 @@ ssize_t fuse_passthrough_write_iter(struct kiocb *iocb,
> if (!count)
> return 0;
>
> - inode_lock(inode);
> + guard(rwsem_write)(&inode->i_rwsem);
> +
> + if (!backing_file)
> + return fuse_ext_map_write_iter(iocb, iter);
[Severity: High]
Does this introduce an ABBA deadlock between pipe_lock and the inode lock?
For normal passthrough, fuse_passthrough_splice_write() acquires the inode
lock first, and then calls backing_file_splice_write(), which ultimately
acquires pipe_lock() via the backing file's splice_write callback.
However, the EXTMAP DAX path added below for splice_write() directly calls
iter_file_splice_write(), which acquires pipe_lock() first. That function then
invokes fuse_passthrough_write_iter() where the inode lock is acquired here
via guard(rwsem_write).
If a malicious FUSE server provides a normal passthrough backing for one open
and an EXTMAP backing for another open of the same inode, and two threads
perform splice writes using the same pipe and the two different file
descriptors, it results in a literal ABBA deadlock. Even operating on
different inodes would trigger a lockdep lock inversion splat.
> +
> ret = backing_file_write_iter(backing_file, iter, iocb, iocb->ki_flags,
> &ctx);
> - inode_unlock(inode);
[ ... ]
> @@ -123,6 +128,9 @@ ssize_t fuse_passthrough_splice_write(struct pipe_inode_info *pipe,
> pr_debug("%s: backing_file=0x%p, pos=%lld, len=%zu, flags=0x%x\n", __func__,
> backing_file, *ppos, len, flags);
>
> + if (!backing_file)
> + return iter_file_splice_write(pipe, out, ppos, len, flags);
> +
> inode_lock(inode);
> init_sync_kiocb(&iocb, out);
> iocb.ki_pos = *ppos;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261008112004.1899560-1-mszeredi@redhat.com?part=8
next prev parent reply other threads:[~2026-10-08 11:36 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-08 11:19 [PATCH v4 0/9] fuse: DAX device based extent maps (famfs) Miklos Szeredi
2026-10-08 11:19 ` [PATCH v4 1/9] dax: add fsdev_dax_from_file() Miklos Szeredi
2026-10-08 11:34 ` sashiko-bot
2026-10-08 11:19 ` [PATCH v4 2/9] dax: use READ_ONCE() in dax_holder() Miklos Szeredi
2026-10-08 11:32 ` sashiko-bot
2026-10-08 13:20 ` Gregory Price
2026-10-08 11:19 ` [PATCH v4 3/9] fuse: add helpers for EIO return value with kernel message Miklos Szeredi
2026-10-08 11:19 ` [PATCH v4 4/9] fuse: support 64 bit, server allocated backing ID Miklos Szeredi
2026-10-09 7:36 ` Amir Goldstein
2026-10-08 11:19 ` [PATCH v4 5/9] fuse: support opening 64 bit " Miklos Szeredi
2026-10-08 11:19 ` [PATCH v4 6/9] fuse: add support for opening dax device as backing Miklos Szeredi
2026-10-08 11:34 ` sashiko-bot
2026-10-08 11:19 ` [PATCH v4 7/9] fuse: add extent map data structure Miklos Szeredi
2026-10-08 11:19 ` [PATCH v4 8/9] fuse: add extent map I/O support Miklos Szeredi
2026-10-08 11:36 ` sashiko-bot [this message]
2026-10-08 11:20 ` [PATCH v4 9/9] fuse: add support for striped backing Miklos Szeredi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=sashiko-outbox-164097@kernel.org \
--to=sashiko-bot@kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=mszeredi@redhat.com \
--cc=nvdimm@lists.linux.dev \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox