* [PATCH v2 0/2] nilfs2: eliminate blockdev_direct_IO() call
@ 2026-08-08 0:35 Viacheslav Dubeyko
2026-08-08 0:35 ` [PATCH v2 1/2] nilfs2: add iomap operations for direct I/O Viacheslav Dubeyko
2026-08-08 0:35 ` [PATCH v2 2/2] nilfs2: switch O_DIRECT to iomap based operations Viacheslav Dubeyko
0 siblings, 2 replies; 5+ messages in thread
From: Viacheslav Dubeyko @ 2026-08-08 0:35 UTC (permalink / raw)
To: konishi.ryusuke, hch
Cc: linux-nilfs, linux-fsdevel, lilinmao, Viacheslav Dubeyko
The patchset eliminates blockdev_direct_IO() from
nilfs2 entirely. It adds iomap.h with declaration of iomaps
operations and iomap.c with implementation of
nilfs_iomap_begin(). The nilfs_direct_IO() and the .direct_IO
callback on nilfs_aops are removed. The nilfs_file_read_iter()
dispatches O_DIRECT reads to iomap_dio_rw() using nilfs_iomap_ops.
Everything else still goes through generic_file_read_iter().
The unnecessary "select LEGACY_DIRECT_IO" has been exchanged on
"select FS_IOMAP" in Kconfig.
v2
The bug has been discovered in nilfs_iomap_begin() of patchset's
first version thanks to Linmao Li.
nilfs_iomap_begin() looks up the filesystem block containing the requested
offset and returns the physical address of that block. However, it sets
iomap->offset to the original, possibly sub-block, offset while leaving
iomap->addr at the start of the physical block.
iomap_sector() adds the difference between the I/O position and
iomap->offset to iomap->addr. For an O_DIRECT read at offset 512 on a
filesystem with 4 KiB blocks, the two file offsets are equal and the I/O is
therefore submitted at the start of the physical block instead of 512 bytes
into it. iomap direct I/O permits this alignment when the device logical
block size is 512 bytes.
Describe mapped and hole extents from the filesystem-block-aligned file
offset. This makes the physical and file offsets refer to the same byte.
It also prevents a one-block hole mapping from extending into the following
block.
Viacheslav Dubeyko (2):
nilfs2: add iomap operations for direct I/O
nilfs2: switch O_DIRECT to iomap based operations
fs/nilfs2/Kconfig | 2 +-
fs/nilfs2/Makefile | 2 +-
fs/nilfs2/file.c | 40 ++++++++++++++++++++++++--
fs/nilfs2/inode.c | 13 ---------
fs/nilfs2/iomap.c | 72 ++++++++++++++++++++++++++++++++++++++++++++++
fs/nilfs2/iomap.h | 13 +++++++++
6 files changed, 124 insertions(+), 18 deletions(-)
create mode 100644 fs/nilfs2/iomap.c
create mode 100644 fs/nilfs2/iomap.h
--
2.43.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH v2 1/2] nilfs2: add iomap operations for direct I/O
2026-08-08 0:35 [PATCH v2 0/2] nilfs2: eliminate blockdev_direct_IO() call Viacheslav Dubeyko
@ 2026-08-08 0:35 ` Viacheslav Dubeyko
2026-08-08 0:35 ` [PATCH v2 2/2] nilfs2: switch O_DIRECT to iomap based operations Viacheslav Dubeyko
1 sibling, 0 replies; 5+ messages in thread
From: Viacheslav Dubeyko @ 2026-08-08 0:35 UTC (permalink / raw)
To: konishi.ryusuke, hch
Cc: linux-nilfs, linux-fsdevel, lilinmao, Viacheslav Dubeyko
The patch adds iomap.h with declaration of iomaps
operations and iomap.c with implementation of
nilfs_iomap_begin(). This method is based on calling
the nilfs_bmap_lookup_contig() method. It resolves
the file's block-mapping to find the sequence of
physically-contiguous blocks starting at starting
block. Offsets at or beyond i_size are reported
as a hole. For offsets within i_size, the requested
length is first clamped to the block-aligned end of the file.
Both the hole and mapped cases then report iomap->offset
aligned to the containing block (blkoff << i_blkbits)
rather than the raw request offset.
Signed-off-by: Viacheslav Dubeyko <slava@dubeyko.com>
Co-developed-by: Linmao Li <lilinmao@kylinos.cn>
Signed-off-by: Linmao Li <lilinmao@kylinos.cn>
cc: Christoph Hellwig <hch@lst.de>
cc: Ryusuke Konishi <konishi.ryusuke@gmail.com>
cc: linux-nilfs@vger.kernel.org
cc: linux-fsdevel@vger.kernel.org
---
fs/nilfs2/Makefile | 2 +-
fs/nilfs2/iomap.c | 72 ++++++++++++++++++++++++++++++++++++++++++++++
fs/nilfs2/iomap.h | 13 +++++++++
3 files changed, 86 insertions(+), 1 deletion(-)
create mode 100644 fs/nilfs2/iomap.c
create mode 100644 fs/nilfs2/iomap.h
diff --git a/fs/nilfs2/Makefile b/fs/nilfs2/Makefile
index 43b60b8a4d07..516e6b85a03c 100644
--- a/fs/nilfs2/Makefile
+++ b/fs/nilfs2/Makefile
@@ -3,4 +3,4 @@ obj-$(CONFIG_NILFS2_FS) += nilfs2.o
nilfs2-y := inode.o file.o dir.o super.o namei.o page.o mdt.o \
btnode.o bmap.o btree.o direct.o dat.o recovery.o \
the_nilfs.o segbuf.o segment.o cpfile.o sufile.o \
- ifile.o alloc.o gcinode.o ioctl.o sysfs.o
+ ifile.o alloc.o gcinode.o ioctl.o sysfs.o iomap.o
diff --git a/fs/nilfs2/iomap.c b/fs/nilfs2/iomap.c
new file mode 100644
index 000000000000..e130ed63abd9
--- /dev/null
+++ b/fs/nilfs2/iomap.c
@@ -0,0 +1,72 @@
+// SPDX-License-Identifier: GPL-2.0+
+/*
+ * NILFS iomap support implementation.
+ *
+ * Written by Viacheslav Dubeyko.
+ */
+
+#include <linux/iomap.h>
+#include <linux/pagemap.h>
+#include "nilfs.h"
+#include "mdt.h"
+#include "iomap.h"
+
+static int nilfs_iomap_begin(struct inode *inode, loff_t offset,
+ loff_t length, unsigned int flags,
+ struct iomap *iomap, struct iomap *srcmap)
+{
+ struct the_nilfs *nilfs = inode->i_sb->s_fs_info;
+ struct nilfs_inode_info *ii = NILFS_I(inode);
+ sector_t blkoff = offset >> inode->i_blkbits;
+ loff_t iomap_offset = (loff_t)blkoff << inode->i_blkbits;
+ unsigned int maxblocks;
+ __u64 blknum = 0;
+ int ret;
+
+ /* Completely beyond EOF. Treat as hole */
+ if (i_size_read(inode) <= offset) {
+ iomap->type = IOMAP_HOLE;
+ iomap->addr = IOMAP_NULL_ADDR;
+ iomap->offset = offset;
+ iomap->length = length;
+ return 0;
+ }
+
+ /* Clamp length if the requested range goes beyond i_size */
+ if (offset + length > i_size_read(inode)) {
+ loff_t i_size = i_size_read(inode);
+ unsigned int blocksize = i_blocksize(inode);
+
+ length = round_up(i_size, blocksize) - offset;
+ }
+
+ maxblocks = min_t(loff_t, length >> inode->i_blkbits, INT_MAX);
+ if (maxblocks == 0)
+ maxblocks = 1;
+
+ down_read(&NILFS_MDT(nilfs->ns_dat)->mi_sem);
+ ret = nilfs_bmap_lookup_contig(ii->i_bmap, blkoff, &blknum, maxblocks);
+ up_read(&NILFS_MDT(nilfs->ns_dat)->mi_sem);
+
+ if (ret == -ENOENT) {
+ iomap->type = IOMAP_HOLE;
+ iomap->addr = IOMAP_NULL_ADDR;
+ iomap->offset = iomap_offset;
+ iomap->length = i_blocksize(inode);
+ return 0;
+ } else if (ret < 0)
+ return ret;
+
+ iomap->bdev = inode->i_sb->s_bdev;
+ iomap->offset = iomap_offset;
+ iomap->length = (loff_t)ret << inode->i_blkbits;
+ iomap->addr = (loff_t)blknum << inode->i_blkbits;
+ iomap->type = IOMAP_MAPPED;
+ iomap->flags = IOMAP_F_MERGED;
+
+ return 0;
+}
+
+const struct iomap_ops nilfs_iomap_ops = {
+ .iomap_begin = nilfs_iomap_begin,
+};
diff --git a/fs/nilfs2/iomap.h b/fs/nilfs2/iomap.h
new file mode 100644
index 000000000000..adef3e22346d
--- /dev/null
+++ b/fs/nilfs2/iomap.h
@@ -0,0 +1,13 @@
+/* SPDX-License-Identifier: GPL-2.0+ */
+/*
+ * NILFS iomap support declarations.
+ *
+ * Written by Viacheslav Dubeyko.
+ */
+
+#ifndef _NILFS_IOMAP_H
+#define _NILFS_IOMAP_H
+
+extern const struct iomap_ops nilfs_iomap_ops;
+
+#endif /* _NILFS_IOMAP_H */
--
2.43.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH v2 2/2] nilfs2: switch O_DIRECT to iomap based operations
2026-08-08 0:35 [PATCH v2 0/2] nilfs2: eliminate blockdev_direct_IO() call Viacheslav Dubeyko
2026-08-08 0:35 ` [PATCH v2 1/2] nilfs2: add iomap operations for direct I/O Viacheslav Dubeyko
@ 2026-08-08 0:35 ` Viacheslav Dubeyko
2026-08-08 23:51 ` Ryusuke Konishi
1 sibling, 1 reply; 5+ messages in thread
From: Viacheslav Dubeyko @ 2026-08-08 0:35 UTC (permalink / raw)
To: konishi.ryusuke, hch
Cc: linux-nilfs, linux-fsdevel, lilinmao, Viacheslav Dubeyko
The patch eliminates blockdev_direct_IO() from
nilfs2 entirely:
- nilfs_file_open() now sets FMODE_CAN_ODIRECT explicitly, since
permission to open the file O_DIRECT was previously implied by
aops->direct_IO being non-NULL.
- nilfs_file_read_iter() dispatches O_DIRECT reads to iomap_dio_rw()
using nilfs_iomap_ops; everything else still goes through
generic_file_read_iter() as before.
- nilfs_file_write_iter() strips IOCB_DIRECT and falls through to
generic_file_write_iter()'s ordinary buffered path. NILFS2 cannot
perform true direct I/O writes: new blocks are delay-allocated and
only given a real disk address by the segment constructor, which
works on buffer_head lists, not iomap. This reproduces today's
actual behavior: the old nilfs_direct_IO() already just returned 0
for WRITE.
- nilfs_direct_IO() and the .direct_IO callback on nilfs_aops are
removed.
- drop the unnecessary "select LEGACY_DIRECT_IO" from Kconfig
in favor of "select FS_IOMAP".
Signed-off-by: Viacheslav Dubeyko <slava@dubeyko.com>
cc: Christoph Hellwig <hch@lst.de>
cc: Ryusuke Konishi <konishi.ryusuke@gmail.com>
cc: linux-nilfs@vger.kernel.org
cc: linux-fsdevel@vger.kernel.org
---
fs/nilfs2/Kconfig | 2 +-
fs/nilfs2/file.c | 40 +++++++++++++++++++++++++++++++++++++---
fs/nilfs2/inode.c | 13 -------------
3 files changed, 38 insertions(+), 17 deletions(-)
diff --git a/fs/nilfs2/Kconfig b/fs/nilfs2/Kconfig
index 7dae168e346e..0a5ace60e6ab 100644
--- a/fs/nilfs2/Kconfig
+++ b/fs/nilfs2/Kconfig
@@ -3,7 +3,7 @@ config NILFS2_FS
tristate "NILFS2 file system support"
select BUFFER_HEAD
select CRC32
- select LEGACY_DIRECT_IO
+ select FS_IOMAP
help
NILFS2 is a log-structured file system (LFS) supporting continuous
snapshotting. In addition to versioning capability of the entire
diff --git a/fs/nilfs2/file.c b/fs/nilfs2/file.c
index f93b68c4877c..ad2e87c049c9 100644
--- a/fs/nilfs2/file.c
+++ b/fs/nilfs2/file.c
@@ -10,9 +10,12 @@
#include <linux/fs.h>
#include <linux/filelock.h>
#include <linux/mm.h>
+#include <linux/uio.h>
+#include <linux/iomap.h>
#include <linux/writeback.h>
#include "nilfs.h"
#include "segment.h"
+#include "iomap.h"
int nilfs_sync_file(struct file *file, loff_t start, loff_t end, int datasync)
{
@@ -133,20 +136,51 @@ static int nilfs_file_mmap_prepare(struct vm_area_desc *desc)
return 0;
}
+static int nilfs_file_open(struct inode *inode, struct file *file)
+{
+ file->f_mode |= FMODE_CAN_ODIRECT;
+ return generic_file_open(inode, file);
+}
+
+static ssize_t nilfs_file_read_iter(struct kiocb *iocb, struct iov_iter *to)
+{
+ if (iocb->ki_flags & IOCB_DIRECT) {
+ return iomap_dio_rw(iocb, to, &nilfs_iomap_ops,
+ NULL, 0, NULL, 0);
+ } else
+ return generic_file_read_iter(iocb, to);
+}
+
+static ssize_t nilfs_file_write_iter(struct kiocb *iocb, struct iov_iter *from)
+{
+ /*
+ * NILFS2 cannot perform true direct I/O writes: new blocks are
+ * delay-allocated and are only given a real disk address when
+ * the segment constructor writes them out as part of a log,
+ * which works directly on buffer_head lists rather than
+ * through iomap. Fall back to the ordinary buffered write path
+ * for O_DIRECT writes.
+ */
+ if (iocb->ki_flags & IOCB_DIRECT)
+ iocb->ki_flags &= ~IOCB_DIRECT;
+
+ return generic_file_write_iter(iocb, from);
+}
+
/*
* We have mostly NULL's here: the current defaults are ok for
* the nilfs filesystem.
*/
const struct file_operations nilfs_file_operations = {
.llseek = generic_file_llseek,
- .read_iter = generic_file_read_iter,
- .write_iter = generic_file_write_iter,
+ .read_iter = nilfs_file_read_iter,
+ .write_iter = nilfs_file_write_iter,
.unlocked_ioctl = nilfs_ioctl,
#ifdef CONFIG_COMPAT
.compat_ioctl = nilfs_compat_ioctl,
#endif /* CONFIG_COMPAT */
.mmap_prepare = nilfs_file_mmap_prepare,
- .open = generic_file_open,
+ .open = nilfs_file_open,
/* .release = nilfs_release_file, */
.fsync = nilfs_sync_file,
.splice_read = filemap_splice_read,
diff --git a/fs/nilfs2/inode.c b/fs/nilfs2/inode.c
index 51f7e125a311..f4a9d9ea9c3f 100644
--- a/fs/nilfs2/inode.c
+++ b/fs/nilfs2/inode.c
@@ -257,18 +257,6 @@ static int nilfs_write_end(const struct kiocb *iocb,
return err ? : copied;
}
-static ssize_t
-nilfs_direct_IO(struct kiocb *iocb, struct iov_iter *iter)
-{
- struct inode *inode = file_inode(iocb->ki_filp);
-
- if (iov_iter_rw(iter) == WRITE)
- return 0;
-
- /* Needs synchronization with the cleaner */
- return blockdev_direct_IO(iocb, inode, iter, nilfs_get_block);
-}
-
const struct address_space_operations nilfs_aops = {
.read_folio = nilfs_read_folio,
.writepages = nilfs_writepages,
@@ -277,7 +265,6 @@ const struct address_space_operations nilfs_aops = {
.write_begin = nilfs_write_begin,
.write_end = nilfs_write_end,
.invalidate_folio = block_invalidate_folio,
- .direct_IO = nilfs_direct_IO,
.migrate_folio = buffer_migrate_folio_norefs,
.is_partially_uptodate = block_is_partially_uptodate,
};
--
2.43.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH v2 2/2] nilfs2: switch O_DIRECT to iomap based operations
2026-08-08 0:35 ` [PATCH v2 2/2] nilfs2: switch O_DIRECT to iomap based operations Viacheslav Dubeyko
@ 2026-08-08 23:51 ` Ryusuke Konishi
2026-08-10 0:15 ` Viacheslav Dubeyko
0 siblings, 1 reply; 5+ messages in thread
From: Ryusuke Konishi @ 2026-08-08 23:51 UTC (permalink / raw)
To: Viacheslav Dubeyko; +Cc: hch, linux-nilfs, linux-fsdevel, lilinmao
Hi Viacheslav,
On Sat, Aug 8, 2026 at 9:35 AM Viacheslav Dubeyko wrote:
...
> +static ssize_t nilfs_file_read_iter(struct kiocb *iocb, struct iov_iter *to)
> +{
> + if (iocb->ki_flags & IOCB_DIRECT) {
> + return iomap_dio_rw(iocb, to, &nilfs_iomap_ops,
> + NULL, 0, NULL, 0);
> + } else
> + return generic_file_read_iter(iocb, to);
> +}
Among the tests that used to pass before the iomap conversion, I found that
running xfstests generic/418 and generic/465 now fails with a "broken bmap"
error, causing the filesystem to remount read-only.
This happens because iomap_dio_rw() is called without holding the shared
inode lock. To prevent this race, I think iomap_dio_rw() should be
protected as follows:
struct inode *inode = file_inode(iocb->ki_filp);
ssize_t ret;
if (iocb->ki_flags & IOCB_DIRECT) {
inode_lock_shared(inode);
ret = iomap_dio_rw(iocb, to, &nilfs_iomap_ops,
NULL, 0, NULL, 0);
inode_unlock_shared(inode);
} else {
ret = generic_file_read_iter(iocb, to);
}
return ret;
Adding this locking resolved the failures in both generic/418 and
generic/465 in my tests.
> +
> +static ssize_t nilfs_file_write_iter(struct kiocb *iocb, struct iov_iter *from)
> +{
> + /*
> + * NILFS2 cannot perform true direct I/O writes: new blocks are
> + * delay-allocated and are only given a real disk address when
> + * the segment constructor writes them out as part of a log,
> + * which works directly on buffer_head lists rather than
> + * through iomap. Fall back to the ordinary buffered write path
> + * for O_DIRECT writes.
> + */
Also, regarding this comment, following Christoph's suggestion, how
about simplifying it as follows?
/*
* NILFS2 lacks direct I/O write support; fall back to buffered writes.
*/
Thanks,
Ryusuke Konishi
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v2 2/2] nilfs2: switch O_DIRECT to iomap based operations
2026-08-08 23:51 ` Ryusuke Konishi
@ 2026-08-10 0:15 ` Viacheslav Dubeyko
0 siblings, 0 replies; 5+ messages in thread
From: Viacheslav Dubeyko @ 2026-08-10 0:15 UTC (permalink / raw)
To: Ryusuke Konishi; +Cc: hch, linux-nilfs, linux-fsdevel, lilinmao
Hi Ryusuke,
On Sun, 2026-08-09 at 08:51 +0900, Ryusuke Konishi wrote:
> Hi Viacheslav,
>
> On Sat, Aug 8, 2026 at 9:35 AM Viacheslav Dubeyko wrote:
> ...
> > +static ssize_t nilfs_file_read_iter(struct kiocb *iocb, struct
> > iov_iter *to)
> > +{
> > + if (iocb->ki_flags & IOCB_DIRECT) {
> > + return iomap_dio_rw(iocb, to, &nilfs_iomap_ops,
> > + NULL, 0, NULL, 0);
> > + } else
> > + return generic_file_read_iter(iocb, to);
> > +}
>
> Among the tests that used to pass before the iomap conversion, I
> found that
> running xfstests generic/418 and generic/465 now fails with a "broken
> bmap"
> error, causing the filesystem to remount read-only.
>
> This happens because iomap_dio_rw() is called without holding the
> shared
> inode lock. To prevent this race, I think iomap_dio_rw() should be
> protected as follows:
>
> struct inode *inode = file_inode(iocb->ki_filp);
> ssize_t ret;
>
> if (iocb->ki_flags & IOCB_DIRECT) {
> inode_lock_shared(inode);
> ret = iomap_dio_rw(iocb, to, &nilfs_iomap_ops,
> NULL, 0, NULL, 0);
> inode_unlock_shared(inode);
> } else {
> ret = generic_file_read_iter(iocb, to);
> }
>
> return ret;
>
> Adding this locking resolved the failures in both generic/418 and
> generic/465 in my tests.
Yeah, makes sense. Let double check this.
>
> > +
> > +static ssize_t nilfs_file_write_iter(struct kiocb *iocb, struct
> > iov_iter *from)
> > +{
> > + /*
> > + * NILFS2 cannot perform true direct I/O writes: new blocks
> > are
> > + * delay-allocated and are only given a real disk address
> > when
> > + * the segment constructor writes them out as part of a
> > log,
> > + * which works directly on buffer_head lists rather than
> > + * through iomap. Fall back to the ordinary buffered write
> > path
> > + * for O_DIRECT writes.
> > + */
>
> Also, regarding this comment, following Christoph's suggestion, how
> about simplifying it as follows?
>
> /*
> * NILFS2 lacks direct I/O write support; fall back to
> buffered writes.
> */
Sorry, my bad. I missed this. Let me rework the comment.
Thanks,
Slava.
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-08-10 0:15 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-08 0:35 [PATCH v2 0/2] nilfs2: eliminate blockdev_direct_IO() call Viacheslav Dubeyko
2026-08-08 0:35 ` [PATCH v2 1/2] nilfs2: add iomap operations for direct I/O Viacheslav Dubeyko
2026-08-08 0:35 ` [PATCH v2 2/2] nilfs2: switch O_DIRECT to iomap based operations Viacheslav Dubeyko
2026-08-08 23:51 ` Ryusuke Konishi
2026-08-10 0:15 ` Viacheslav Dubeyko
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox