* [PATCH v5 1/8] fs: add write-stream management ioctls
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
@ 2026-09-21 9:31 ` Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 2/8] fs: add generic write-stream management Kanchan Joshi
` (8 subsequent siblings)
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:31 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev, Anuj Gupta, Kanchan Joshi
From: Anuj Gupta <anuj20.g@samsung.com>
Add three ioctls for write stream management:
FS_IOC_WRITE_STREAM_GET_MAX number of streams the filesystem offers
FS_IOC_WRITE_STREAM_ALLOC reserve a stream, returns an fd
FS_IOC_WRITE_STREAM_SET set or clear a stream on a file
Write streams are a resource the kernel manages and applications ask
for.
ALLOC reserves a stream and returns a stream fd; the stream stays
reserved as long as that fd is open. ALLOC fails once all streams are
reserved.
SET assigns the stream named by the stream_fd in struct
fs_write_stream_set to an open file; the same stream fd can be used to
set the stream on any number of files. SET with FS_WRITE_STREAM_SET_CLEAR
and a stream_fd of -1 clears whatever stream the file carries.
Suggested-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
---
include/uapi/linux/fs.h | 15 +++++++++++++++
1 file changed, 15 insertions(+)
diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index 34c6f219462a..1974052aa476 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -345,6 +345,21 @@ struct file_attr {
/* Get logical block metadata capability details */
#define FS_IOC_GETLBMD_CAP _IOWR(0x15, 2, struct logical_block_metadata_cap)
+struct fs_write_stream_set {
+ __s32 stream_fd; /* IN: stream to attach, -1 to clear */
+ __u32 flags; /* IN: FS_WRITE_STREAM_SET_* */
+};
+
+/* Clear the file's write stream. Requires stream_fd to be -1. */
+#define FS_WRITE_STREAM_SET_CLEAR (1U << 0)
+
+/* GET_MAX returns the count through its argument; ALLOC returns a stream fd
+ * as the ioctl return value and takes none.
+ */
+#define FS_IOC_WRITE_STREAM_GET_MAX _IOR(0x15, 3, __u32)
+#define FS_IOC_WRITE_STREAM_ALLOC _IO(0x15, 4)
+#define FS_IOC_WRITE_STREAM_SET _IOW(0x15, 5, struct fs_write_stream_set)
+
/*
* Inode flags (FS_IOC_GETFLAGS / FS_IOC_SETFLAGS)
*
--
2.25.1
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH v5 2/8] fs: add generic write-stream management
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 1/8] fs: add write-stream management ioctls Kanchan Joshi
@ 2026-09-21 9:31 ` Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint Kanchan Joshi
` (7 subsequent siblings)
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:31 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev, Anuj Gupta, Kanchan Joshi
From: Anuj Gupta <anuj20.g@samsung.com>
Add a bitmap-based stream pool and anonymous fds for reserved streams.
Helpers initialize and destroy a pool, reserve a stream as an fd, and
resolve a stream fd back to its id.
Suggested-by: Christoph Hellwig <hch@lst.de>
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
---
fs/Makefile | 2 +-
fs/write_streams.c | 138 ++++++++++++++++++++++++++++++++++
include/linux/write_streams.h | 33 ++++++++
3 files changed, 172 insertions(+), 1 deletion(-)
create mode 100644 fs/write_streams.c
create mode 100644 include/linux/write_streams.h
diff --git a/fs/Makefile b/fs/Makefile
index 055dfc23d82b..494da034cef2 100644
--- a/fs/Makefile
+++ b/fs/Makefile
@@ -23,7 +23,7 @@ obj-$(CONFIG_PROC_FS) += proc_namespace.o
obj-$(CONFIG_LEGACY_DIRECT_IO) += direct-io.o
obj-y += notify/
obj-$(CONFIG_EPOLL) += eventpoll.o
-obj-y += anon_inodes.o
+obj-y += anon_inodes.o write_streams.o
obj-$(CONFIG_SIGNALFD) += signalfd.o
obj-$(CONFIG_TIMERFD) += timerfd.o
obj-$(CONFIG_EVENTFD) += eventfd.o
diff --git a/fs/write_streams.c b/fs/write_streams.c
new file mode 100644
index 000000000000..b7f6ca6cf725
--- /dev/null
+++ b/fs/write_streams.c
@@ -0,0 +1,138 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Generic write-stream fd management. */
+#include <linux/anon_inodes.h>
+#include <linux/bitmap.h>
+#include <linux/err.h>
+#include <linux/file.h>
+#include <linux/fs.h>
+#include <linux/mount.h>
+#include <linux/slab.h>
+#include <linux/write_streams.h>
+
+struct write_stream {
+ struct write_stream_pool *pool;
+ struct vfsmount *mnt; /* pins the mount */
+ u16 id; /* 1-based slot number */
+};
+
+static int write_stream_release(struct inode *inode, struct file *file)
+{
+ struct write_stream *ws = file->private_data;
+
+ spin_lock(&ws->pool->lock);
+ __clear_bit(ws->id - 1, ws->pool->streams_in_use);
+ spin_unlock(&ws->pool->lock);
+ mntput(ws->mnt);
+ kfree(ws);
+ return 0;
+}
+
+static const struct file_operations write_stream_fops = {
+ .release = write_stream_release,
+ .llseek = noop_llseek,
+};
+
+/**
+ * write_stream_pool_init - size a write stream pool
+ * @pool: pool to initialise
+ * @nr: number of slots, may be zero
+ *
+ * Returns 0, or -ENOMEM if the bitmap cannot be allocated.
+ */
+int write_stream_pool_init(struct write_stream_pool *pool, unsigned int nr)
+{
+ spin_lock_init(&pool->lock);
+ pool->streams_in_use = NULL;
+ pool->nr_streams = 0;
+
+ if (!nr)
+ return 0;
+
+ pool->streams_in_use = bitmap_zalloc(nr, GFP_KERNEL);
+ if (!pool->streams_in_use)
+ return -ENOMEM;
+ pool->nr_streams = nr;
+ return 0;
+}
+EXPORT_SYMBOL_GPL(write_stream_pool_init);
+
+/**
+ * write_stream_pool_destroy - free a write stream pool
+ * @pool: pool to release
+ */
+void write_stream_pool_destroy(struct write_stream_pool *pool)
+{
+ bitmap_free(pool->streams_in_use);
+ pool->streams_in_use = NULL;
+ pool->nr_streams = 0;
+}
+EXPORT_SYMBOL_GPL(write_stream_pool_destroy);
+
+/**
+ * write_stream_alloc_fd - reserve a slot and return an fd naming it
+ * @pool: pool to reserve from
+ * @file: file whose mount is pinned for the life of the stream
+ *
+ * Returns an O_RDONLY | O_CLOEXEC fd, -EOPNOTSUPP if the pool has no slots,
+ * or -EBUSY if they are all taken.
+ */
+int write_stream_alloc_fd(struct write_stream_pool *pool, struct file *file)
+{
+ struct write_stream *ws;
+ unsigned int slot;
+ int fd;
+
+ if (!pool->nr_streams)
+ return -EOPNOTSUPP;
+
+ ws = kzalloc_obj(*ws, GFP_KERNEL);
+ if (!ws)
+ return -ENOMEM;
+
+ spin_lock(&pool->lock);
+ slot = find_first_zero_bit(pool->streams_in_use, pool->nr_streams);
+ if (slot >= pool->nr_streams) {
+ spin_unlock(&pool->lock);
+ kfree(ws);
+ return -EBUSY;
+ }
+ __set_bit(slot, pool->streams_in_use);
+ spin_unlock(&pool->lock);
+
+ ws->pool = pool;
+ ws->mnt = mntget(file->f_path.mnt);
+ ws->id = slot + 1;
+
+ fd = anon_inode_getfd("[write_stream]", &write_stream_fops, ws,
+ O_RDONLY | O_CLOEXEC);
+ if (fd < 0) {
+ spin_lock(&pool->lock);
+ __clear_bit(slot, pool->streams_in_use);
+ spin_unlock(&pool->lock);
+ mntput(ws->mnt);
+ kfree(ws);
+ }
+ return fd;
+}
+EXPORT_SYMBOL_GPL(write_stream_alloc_fd);
+
+/**
+ * write_stream_get_id - the id a stream file names
+ * @file: candidate stream file
+ * @pool: pool the stream must belong to, or NULL to accept any
+ *
+ * Returns the 1-based id, or -EINVAL if @file is not a stream file, or is
+ * a stream of a different pool.
+ */
+int write_stream_get_id(struct file *file, const struct write_stream_pool *pool)
+{
+ struct write_stream *ws;
+
+ if (file->f_op != &write_stream_fops)
+ return -EINVAL;
+ ws = file->private_data;
+ if (pool && ws->pool != pool)
+ return -EINVAL;
+ return ws->id;
+}
+EXPORT_SYMBOL_GPL(write_stream_get_id);
diff --git a/include/linux/write_streams.h b/include/linux/write_streams.h
new file mode 100644
index 000000000000..d27f42df5984
--- /dev/null
+++ b/include/linux/write_streams.h
@@ -0,0 +1,33 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LINUX_WRITE_STREAMS_H
+#define _LINUX_WRITE_STREAMS_H
+
+#include <linux/spinlock.h>
+#include <linux/types.h>
+
+struct file;
+
+/*
+ * Write stream slot pool, embedded in whatever filesystem object scopes
+ * an id space.
+ */
+struct write_stream_pool {
+ unsigned int nr_streams;
+ unsigned long *streams_in_use;
+ spinlock_t lock;
+};
+
+static inline unsigned int
+write_stream_pool_count(const struct write_stream_pool *pool)
+{
+ return pool->nr_streams;
+}
+
+int write_stream_pool_init(struct write_stream_pool *pool, unsigned int nr);
+void write_stream_pool_destroy(struct write_stream_pool *pool);
+
+int write_stream_alloc_fd(struct write_stream_pool *pool, struct file *file);
+int write_stream_get_id(struct file *file,
+ const struct write_stream_pool *pool);
+
+#endif /* _LINUX_WRITE_STREAMS_H */
--
2.25.1
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 1/8] fs: add write-stream management ioctls Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 2/8] fs: add generic write-stream management Kanchan Joshi
@ 2026-09-21 9:31 ` Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 4/8] xfs: implement software write-stream management Kanchan Joshi
` (6 subsequent siblings)
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:31 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi, Anuj Gupta
Keep the attached stream id in the VFS inode rather than the
filesystem's own, so that fcntl_set_rw_hint() can see it. It is a u16 in
the 32-bit hole after i_state, so struct inode does not grow.
Suggested-by: Christoph Hellwig <hch@lst.de>
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
---
fs/fcntl.c | 7 +++++++
fs/inode.c | 1 +
include/linux/fs.h | 2 +-
3 files changed, 9 insertions(+), 1 deletion(-)
diff --git a/fs/fcntl.c b/fs/fcntl.c
index c158f082f1da..a5439e076197 100644
--- a/fs/fcntl.c
+++ b/fs/fcntl.c
@@ -380,7 +380,14 @@ static long fcntl_set_rw_hint(struct file *file, unsigned long arg)
if (!rw_hint_valid(hint))
return -EINVAL;
+ /* keep write stream and write life time hint mutually exclusive */
+ spin_lock(&inode->i_lock);
+ if (hint != WRITE_LIFE_NOT_SET && READ_ONCE(inode->i_write_stream)) {
+ spin_unlock(&inode->i_lock);
+ return -EBUSY;
+ }
WRITE_ONCE(inode->i_write_hint, hint);
+ spin_unlock(&inode->i_lock);
/*
* file->f_mapping->host may differ from inode. As an example,
diff --git a/fs/inode.c b/fs/inode.c
index ba7da39be4a3..50b91d449a38 100644
--- a/fs/inode.c
+++ b/fs/inode.c
@@ -246,6 +246,7 @@ int inode_init_always_gfp(struct super_block *sb, struct inode *inode, gfp_t gfp
atomic_set(&inode->i_writecount, 0);
inode->i_size = 0;
inode->i_write_hint = WRITE_LIFE_NOT_SET;
+ inode->i_write_stream = 0;
inode->i_blocks = 0;
inode->i_bytes = 0;
inode->i_generation = 0;
diff --git a/include/linux/fs.h b/include/linux/fs.h
index f9d1e05e8ae6..3d9ce4d30cd5 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -812,7 +812,7 @@ struct inode {
/* Misc */
struct inode_state_flags i_state;
- /* 32-bit hole */
+ u16 i_write_stream;
struct rw_semaphore i_rwsem;
unsigned long dirtied_when; /* jiffies of first dirtying */
--
2.25.1
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH v5 4/8] xfs: implement software write-stream management
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
` (2 preceding siblings ...)
2026-09-21 9:31 ` [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint Kanchan Joshi
@ 2026-09-21 9:31 ` Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 5/8] xfs: write stream based AG placement Kanchan Joshi
` (5 subsequent siblings)
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:31 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev, Anuj Gupta, Kanchan Joshi
From: Anuj Gupta <anuj20.g@samsung.com>
Implement the FS_IOC_WRITE_STREAM_* handlers.
Keep a stream pool in the data device buftarg, sized from the AG count.
The id is stored in inode->i_write_stream and is not written to disk.
Write streams, filestreams and write life time hints are mutually
exclusive. GET_MAX reports zero for filestream and realtime inodes, SET
rejects those and any inode carrying a hint, and setattr refuses to turn
on filestream or realtime placement while a stream is attached.
ALLOC is restricted to CAP_SYS_ADMIN.
Suggested-by: Christoph Hellwig <hch@lst.de>
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
---
fs/xfs/xfs_buf.c | 30 ++++++++++++++++++
fs/xfs/xfs_buf.h | 6 ++++
fs/xfs/xfs_inode.c | 69 ++++++++++++++++++++++++++++++++++++++++
fs/xfs/xfs_inode.h | 4 +++
fs/xfs/xfs_ioctl.c | 79 ++++++++++++++++++++++++++++++++++++++++++++++
fs/xfs/xfs_super.c | 5 +++
6 files changed, 193 insertions(+)
diff --git a/fs/xfs/xfs_buf.c b/fs/xfs/xfs_buf.c
index 8256c1d13ce2..4e84528b05cb 100644
--- a/fs/xfs/xfs_buf.c
+++ b/fs/xfs/xfs_buf.c
@@ -1647,6 +1647,7 @@ void
xfs_free_buftarg(
struct xfs_buftarg *btp)
{
+ write_stream_pool_destroy(&btp->bt_stream_pool);
xfs_destroy_buftarg(btp);
fs_put_dax(btp->bt_daxdev, btp->bt_mount);
/* the main block device is closed by kill_block_super */
@@ -1684,6 +1685,35 @@ xfs_configure_buftarg_atomic_writes(
btp->bt_awu_max = max_bytes;
}
+#define XFS_MAX_SW_WRITE_STREAMS U8_MAX
+
+/* Heuristic to derive software write stream count from a group topology */
+static unsigned int
+xfs_sw_write_stream_count(
+ unsigned int nr_groups)
+{
+ unsigned int group_set_size;
+
+ if (nr_groups >= 16)
+ group_set_size = 4;
+ else if (nr_groups >= 8)
+ group_set_size = 2;
+ else
+ group_set_size = 1;
+ return min(nr_groups / group_set_size, XFS_MAX_SW_WRITE_STREAMS);
+}
+
+int
+xfs_buftarg_init_streams(
+ struct xfs_buftarg *btp,
+ unsigned int nr_groups)
+{
+ unsigned int nr_streams;
+
+ nr_streams = xfs_sw_write_stream_count(nr_groups);
+ return write_stream_pool_init(&btp->bt_stream_pool, nr_streams);
+}
+
/* Configure a buffer target that abstracts a block device. */
int
xfs_configure_buftarg(
diff --git a/fs/xfs/xfs_buf.h b/fs/xfs/xfs_buf.h
index a4729253b56f..b5292e5731e3 100644
--- a/fs/xfs/xfs_buf.h
+++ b/fs/xfs/xfs_buf.h
@@ -15,6 +15,7 @@
#include <linux/uio.h>
#include <linux/list_lru.h>
#include <linux/lockref.h>
+#include <linux/write_streams.h>
extern struct kmem_cache *xfs_buf_cache;
@@ -102,6 +103,9 @@ struct xfs_buftarg {
unsigned int bt_awu_min;
unsigned int bt_awu_max;
+ /* slot pool for stream fds on this device */
+ struct write_stream_pool bt_stream_pool;
+
struct rhashtable bt_hash;
};
@@ -360,6 +364,8 @@ extern void xfs_buftarg_wait(struct xfs_buftarg *);
extern void xfs_buftarg_drain(struct xfs_buftarg *);
int xfs_configure_buftarg(struct xfs_buftarg *btp, unsigned int sectorsize,
xfs_fsblock_t nr_blocks);
+int xfs_buftarg_init_streams(struct xfs_buftarg *btp,
+ unsigned int nr_groups);
#define xfs_readonly_buftarg(buftarg) bdev_read_only((buftarg)->bt_bdev)
diff --git a/fs/xfs/xfs_inode.c b/fs/xfs/xfs_inode.c
index 030a7c8f2c12..df0a37c4d68a 100644
--- a/fs/xfs/xfs_inode.c
+++ b/fs/xfs/xfs_inode.c
@@ -4,6 +4,7 @@
* All Rights Reserved.
*/
#include <linux/iversion.h>
+#include <linux/write_streams.h>
#include "xfs_platform.h"
#include "xfs_fs.h"
@@ -47,6 +48,74 @@
struct kmem_cache *xfs_inode_cache;
+/* Number of write streams available to this inode. */
+int
+xfs_inode_max_write_streams(
+ struct xfs_inode *ip)
+{
+ xfs_assert_ilocked(ip, XFS_ILOCK_SHARED | XFS_ILOCK_EXCL);
+
+ if (xfs_inode_is_filestream(ip))
+ return 0;
+ if (XFS_IS_REALTIME_INODE(ip))
+ return 0;
+ return write_stream_pool_count(&xfs_inode_buftarg(ip)->bt_stream_pool);
+}
+
+/* Bind the write stream named by @stream_fd to @ip */
+int
+xfs_inode_set_write_stream(
+ struct xfs_inode *ip,
+ int stream_fd)
+{
+ CLASS(fd, f)(stream_fd);
+ struct xfs_buftarg *target;
+ int id, error = 0;
+
+ if (!fd_file(f))
+ return -EBADF;
+ xfs_ilock(ip, XFS_ILOCK_EXCL);
+ if (XFS_IS_REALTIME_INODE(ip)) {
+ error = -EINVAL;
+ goto out_unlock;
+ }
+
+ target = xfs_inode_buftarg(ip);
+ id = write_stream_get_id(fd_file(f), &target->bt_stream_pool);
+ if (id < 0) {
+ error = id;
+ goto out_unlock;
+ }
+
+ /* Filestream and write-stream are mutually exclusive */
+ if (xfs_inode_is_filestream(ip)) {
+ error = -EINVAL;
+ goto out_unlock;
+ }
+
+ /* keep write stream and write life time hint mutually exclusive */
+ spin_lock(&VFS_I(ip)->i_lock);
+ if (VFS_I(ip)->i_write_hint != WRITE_LIFE_NOT_SET) {
+ spin_unlock(&VFS_I(ip)->i_lock);
+ error = -EBUSY;
+ goto out_unlock;
+ }
+ WRITE_ONCE(VFS_I(ip)->i_write_stream, id);
+ spin_unlock(&VFS_I(ip)->i_lock);
+out_unlock:
+ xfs_iunlock(ip, XFS_ILOCK_EXCL);
+ return error;
+}
+
+void
+xfs_inode_clear_write_stream(
+ struct xfs_inode *ip)
+{
+ xfs_ilock(ip, XFS_ILOCK_EXCL);
+ WRITE_ONCE(VFS_I(ip)->i_write_stream, 0);
+ xfs_iunlock(ip, XFS_ILOCK_EXCL);
+}
+
/*
* These two are wrapper routines around the xfs_ilock() routine used to
* centralize some grungy code. They are used in places that wish to lock the
diff --git a/fs/xfs/xfs_inode.h b/fs/xfs/xfs_inode.h
index 1602027cd0aa..c053f105d2fa 100644
--- a/fs/xfs/xfs_inode.h
+++ b/fs/xfs/xfs_inode.h
@@ -672,4 +672,8 @@ int xfs_icreate_dqalloc(const struct xfs_icreate_args *args,
struct xfs_dquot **udqpp, struct xfs_dquot **gdqpp,
struct xfs_dquot **pdqpp);
+int xfs_inode_max_write_streams(struct xfs_inode *ip);
+int xfs_inode_set_write_stream(struct xfs_inode *ip, int stream_fd);
+void xfs_inode_clear_write_stream(struct xfs_inode *ip);
+
#endif /* __XFS_INODE_H__ */
diff --git a/fs/xfs/xfs_ioctl.c b/fs/xfs/xfs_ioctl.c
index 96ca3e480cb9..bc6957ad06cb 100644
--- a/fs/xfs/xfs_ioctl.c
+++ b/fs/xfs/xfs_ioctl.c
@@ -557,6 +557,11 @@ xfs_ioctl_setattr_xflags(
bool rtflag = (fa->fsx_xflags & FS_XFLAG_REALTIME);
uint64_t i_flags2;
+ /* refuse a filestream/realtime flag change while a stream is attached */
+ if (READ_ONCE(VFS_I(ip)->i_write_stream) &&
+ ((fa->fsx_xflags & FS_XFLAG_FILESTREAM) || rtflag))
+ return -EINVAL;
+
if (rtflag != XFS_IS_REALTIME_INODE(ip)) {
/* Can't change realtime flag if any extents are allocated. */
if (xfs_inode_has_filedata(ip))
@@ -1200,6 +1205,73 @@ xfs_ioctl_fs_counts(
return 0;
}
+static int
+xfs_ioc_write_stream_get_max(
+ struct file *filp,
+ void __user *arg)
+{
+ struct xfs_inode *ip = XFS_I(file_inode(filp));
+ __u32 max;
+
+ xfs_ilock(ip, XFS_ILOCK_SHARED);
+ max = xfs_inode_max_write_streams(ip);
+ xfs_iunlock(ip, XFS_ILOCK_SHARED);
+
+ return put_user(max, (__u32 __user *)arg);
+}
+
+static int
+xfs_ioc_write_stream_alloc(
+ struct file *filp)
+{
+ struct xfs_inode *ip = XFS_I(file_inode(filp));
+ struct xfs_buftarg *target;
+ int max;
+
+ if (!capable(CAP_SYS_ADMIN))
+ return -EPERM;
+
+ xfs_ilock(ip, XFS_ILOCK_SHARED);
+ max = xfs_inode_max_write_streams(ip);
+ target = xfs_inode_buftarg(ip);
+ xfs_iunlock(ip, XFS_ILOCK_SHARED);
+
+ if (!max)
+ return -EOPNOTSUPP;
+ return write_stream_alloc_fd(&target->bt_stream_pool, filp);
+}
+
+static int
+xfs_ioc_write_stream_set(
+ struct file *filp,
+ void __user *arg)
+{
+ struct xfs_inode *ip = XFS_I(file_inode(filp));
+ struct fs_write_stream_set set;
+
+ if (!(filp->f_mode & FMODE_WRITE))
+ return -EBADF;
+ if (copy_from_user(&set, arg, sizeof(set)))
+ return -EFAULT;
+ if (set.flags & ~FS_WRITE_STREAM_SET_CLEAR)
+ return -EINVAL;
+
+ /* Regular files only. */
+ if (!S_ISREG(VFS_I(ip)->i_mode))
+ return -EINVAL;
+
+ if (!inode_owner_or_capable(file_mnt_idmap(filp), VFS_I(ip)))
+ return -EPERM;
+
+ if (set.flags & FS_WRITE_STREAM_SET_CLEAR) {
+ if (set.stream_fd != -1)
+ return -EINVAL;
+ xfs_inode_clear_write_stream(ip);
+ return 0;
+ }
+ return xfs_inode_set_write_stream(ip, set.stream_fd);
+}
+
/*
* These long-unused ioctls were removed from the official ioctl API in 5.17,
* but retain these definitions so that we can log warnings about them.
@@ -1466,6 +1538,13 @@ xfs_file_ioctl(
case XFS_IOC_VERIFY_MEDIA:
return xfs_ioc_verify_media(filp, arg);
+ case FS_IOC_WRITE_STREAM_GET_MAX:
+ return xfs_ioc_write_stream_get_max(filp, arg);
+ case FS_IOC_WRITE_STREAM_ALLOC:
+ return xfs_ioc_write_stream_alloc(filp);
+ case FS_IOC_WRITE_STREAM_SET:
+ return xfs_ioc_write_stream_set(filp, (void __user *)arg);
+
default:
return -ENOTTY;
}
diff --git a/fs/xfs/xfs_super.c b/fs/xfs/xfs_super.c
index b24db75eaedc..443b3d847150 100644
--- a/fs/xfs/xfs_super.c
+++ b/fs/xfs/xfs_super.c
@@ -608,6 +608,11 @@ xfs_setup_devices(
if (error)
return error;
+ error = xfs_buftarg_init_streams(mp->m_ddev_targp,
+ mp->m_sb.sb_agcount);
+ if (error)
+ return error;
+
if (mp->m_logdev_targp && mp->m_logdev_targp != mp->m_ddev_targp) {
unsigned int log_sector_size = BBSIZE;
--
2.25.1
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH v5 5/8] xfs: write stream based AG placement
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
` (3 preceding siblings ...)
2026-09-21 9:31 ` [PATCH v5 4/8] xfs: implement software write-stream management Kanchan Joshi
@ 2026-09-21 9:31 ` Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 6/8] xfs: support write streams on realtime volumes Kanchan Joshi
` (4 subsequent siblings)
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:31 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi, Anuj Gupta
Choose the starting AG from the write stream set on the file.
Isolating streams into separate allocation groups reduces block
interleaving between concurrent writers, AGF lock contention and logical
file fragmentation.
AGs are partitioned among the streams. The stream value selects the AG
set and the inode number selects the AG within it, so intra-stream
concurrency comes from the AG set size.
Example: 8 Allocation Groups, 4 write streams
AG set size = 2 AGs per write stream
Stream 1 (ID: 1) Stream 2 (ID: 2) Streams 3 & 4
+---------+---------+ +---------+---------+ +-------------
| AG0 | AG1 | | AG2 | AG3 | | AG4...AG7
+---------+---------+ +---------+---------+ +-------------
^ ^ ^ ^
| | | |
| File B (ino: 101) | File D (ino: 201)
| 101 % 2 = 1 -> AG 1 | 201 % 2 = 1 -> AG 3
| |
File A (ino: 100) File C (ino: 200)
100 % 2 = 0 -> AG 0 200 % 2 = 0 -> AG 2
If the AGs do not divide evenly, the last stream absorbs the remainder.
Set boundaries are a hint, not a partition: file contiguity is preserved
and the full space stays usable with a single stream.
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
---
fs/xfs/libxfs/xfs_bmap.c | 58 ++++++++++++++++++++++++++++++++++++++--
1 file changed, 56 insertions(+), 2 deletions(-)
diff --git a/fs/xfs/libxfs/xfs_bmap.c b/fs/xfs/libxfs/xfs_bmap.c
index d64defeda645..2e83811cefcb 100644
--- a/fs/xfs/libxfs/xfs_bmap.c
+++ b/fs/xfs/libxfs/xfs_bmap.c
@@ -3579,8 +3579,31 @@ xfs_bmap_btalloc_filestreams(
return xfs_bmap_btalloc_low_space(ap, args);
}
+static xfs_agnumber_t
+xfs_bmap_write_stream_agno(
+ struct xfs_inode *ip,
+ unsigned int stream_id)
+{
+ struct xfs_mount *mp = ip->i_mount;
+ xfs_agnumber_t nr_ags = mp->m_sb.sb_agcount;
+ unsigned int nr_streams =
+ write_stream_pool_count(&mp->m_ddev_targp->bt_stream_pool);
+ xfs_agnumber_t ag_set_size, start_agno;
+
+ stream_id -= 1; /* convert from 1-based to 0-based */
+ ag_set_size = nr_ags / nr_streams;
+ start_agno = stream_id * ag_set_size;
+
+ /* last stream absorbs any uneven remainder */
+ if (stream_id == nr_streams - 1)
+ ag_set_size = nr_ags - start_agno;
+
+ return start_agno + I_INO(ip) % ag_set_size;
+}
+
+/* Core AG allocator. The caller sets ap->blkno to the target AG start. */
static int
-xfs_bmap_btalloc_best_length(
+xfs_bmap_btalloc_from_blkno(
struct xfs_bmalloca *ap,
struct xfs_alloc_arg *args,
int stripe_align)
@@ -3588,7 +3611,6 @@ xfs_bmap_btalloc_best_length(
xfs_extlen_t blen = 0;
int error;
- ap->blkno = XFS_INODE_TO_FSB(ap->ip);
if (!xfs_bmap_adjacent(ap))
ap->eof = false;
@@ -3621,6 +3643,32 @@ xfs_bmap_btalloc_best_length(
return xfs_bmap_btalloc_low_space(ap, args);
}
+/* Start a write-stream file in the AG set that backs its stream. */
+static int
+xfs_bmap_btalloc_write_stream(
+ struct xfs_bmalloca *ap,
+ struct xfs_alloc_arg *args,
+ int stripe_align,
+ unsigned int stream_id)
+{
+ struct xfs_mount *mp = ap->ip->i_mount;
+ xfs_agnumber_t agno;
+
+ agno = xfs_bmap_write_stream_agno(ap->ip, stream_id);
+ ap->blkno = XFS_AGB_TO_FSB(mp, agno, 0);
+ return xfs_bmap_btalloc_from_blkno(ap, args, stripe_align);
+}
+
+static int
+xfs_bmap_btalloc_best_length(
+ struct xfs_bmalloca *ap,
+ struct xfs_alloc_arg *args,
+ int stripe_align)
+{
+ ap->blkno = XFS_INODE_TO_FSB(ap->ip);
+ return xfs_bmap_btalloc_from_blkno(ap, args, stripe_align);
+}
+
static int
xfs_bmap_btalloc(
struct xfs_bmalloca *ap)
@@ -3640,6 +3688,7 @@ xfs_bmap_btalloc(
};
xfs_fileoff_t orig_offset;
xfs_extlen_t orig_length;
+ unsigned int stream_id;
int error;
int stripe_align;
@@ -3652,11 +3701,16 @@ xfs_bmap_btalloc(
/* Trim the allocation back to the maximum an AG can fit. */
args.maxlen = min(ap->length, mp->m_ag_max_usable);
+ stream_id = READ_ONCE(VFS_I(ap->ip)->i_write_stream);
+
if (unlikely(XFS_TEST_ERROR(mp, XFS_ERRTAG_BMAP_ALLOC_MINLEN_EXTENT)))
error = xfs_bmap_exact_minlen_extent_alloc(ap, &args);
else if ((ap->datatype & XFS_ALLOC_USERDATA) &&
xfs_inode_is_filestream(ap->ip))
error = xfs_bmap_btalloc_filestreams(ap, &args, stripe_align);
+ else if ((ap->datatype & XFS_ALLOC_USERDATA) && stream_id)
+ error = xfs_bmap_btalloc_write_stream(ap, &args, stripe_align,
+ stream_id);
else
error = xfs_bmap_btalloc_best_length(ap, &args, stripe_align);
if (error)
--
2.25.1
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH v5 6/8] xfs: support write streams on realtime volumes
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
` (4 preceding siblings ...)
2026-09-21 9:31 ` [PATCH v5 5/8] xfs: write stream based AG placement Kanchan Joshi
@ 2026-09-21 9:31 ` Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 7/8] iomap: introduce and propagate write_stream Kanchan Joshi
` (3 subsequent siblings)
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:31 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev, Anuj Gupta, Kanchan Joshi
From: Anuj Gupta <anuj20.g@samsung.com>
Enable write streams on non-zoned realtime volumes that have realtime
groups, partitioning RTGs among the streams as the AG side does.
Stream based RTG selection applies to the first allocation of a file;
later ones keep using the adjacent-allocation hint so the file stays
contiguous.
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
---
fs/xfs/xfs_inode.c | 4 ++--
fs/xfs/xfs_ioctl.c | 3 ++-
fs/xfs/xfs_rtalloc.c | 27 +++++++++++++++++++++++++++
fs/xfs/xfs_super.c | 7 +++++++
4 files changed, 38 insertions(+), 3 deletions(-)
diff --git a/fs/xfs/xfs_inode.c b/fs/xfs/xfs_inode.c
index df0a37c4d68a..c3503fe870ba 100644
--- a/fs/xfs/xfs_inode.c
+++ b/fs/xfs/xfs_inode.c
@@ -57,7 +57,7 @@ xfs_inode_max_write_streams(
if (xfs_inode_is_filestream(ip))
return 0;
- if (XFS_IS_REALTIME_INODE(ip))
+ if (XFS_IS_REALTIME_INODE(ip) && xfs_has_zoned(ip->i_mount))
return 0;
return write_stream_pool_count(&xfs_inode_buftarg(ip)->bt_stream_pool);
}
@@ -75,7 +75,7 @@ xfs_inode_set_write_stream(
if (!fd_file(f))
return -EBADF;
xfs_ilock(ip, XFS_ILOCK_EXCL);
- if (XFS_IS_REALTIME_INODE(ip)) {
+ if (XFS_IS_REALTIME_INODE(ip) && xfs_has_zoned(ip->i_mount)) {
error = -EINVAL;
goto out_unlock;
}
diff --git a/fs/xfs/xfs_ioctl.c b/fs/xfs/xfs_ioctl.c
index bc6957ad06cb..4c0aff34e3c3 100644
--- a/fs/xfs/xfs_ioctl.c
+++ b/fs/xfs/xfs_ioctl.c
@@ -559,7 +559,8 @@ xfs_ioctl_setattr_xflags(
/* refuse a filestream/realtime flag change while a stream is attached */
if (READ_ONCE(VFS_I(ip)->i_write_stream) &&
- ((fa->fsx_xflags & FS_XFLAG_FILESTREAM) || rtflag))
+ ((fa->fsx_xflags & FS_XFLAG_FILESTREAM) ||
+ rtflag != XFS_IS_REALTIME_INODE(ip)))
return -EINVAL;
if (rtflag != XFS_IS_REALTIME_INODE(ip)) {
diff --git a/fs/xfs/xfs_rtalloc.c b/fs/xfs/xfs_rtalloc.c
index 84efe5a8fb11..a291f1aca6eb 100644
--- a/fs/xfs/xfs_rtalloc.c
+++ b/fs/xfs/xfs_rtalloc.c
@@ -2160,6 +2160,28 @@ xfs_rtallocate_align(
return 0;
}
+static xfs_rtblock_t
+xfs_bmap_write_stream_rtbno(
+ struct xfs_inode *ip,
+ unsigned int stream_id)
+{
+ struct xfs_mount *mp = ip->i_mount;
+ unsigned int nr_rtgs = mp->m_sb.sb_rgcount;
+ unsigned int nr_streams =
+ write_stream_pool_count(&mp->m_rtdev_targp->bt_stream_pool);
+ unsigned int set_size = nr_rtgs / nr_streams;
+ xfs_rgnumber_t start_rgno, target_rgno;
+
+ stream_id -= 1; /* convert from 1-based to 0-based */
+ start_rgno = stream_id * set_size;
+ if (stream_id == nr_streams - 1)
+ set_size = nr_rtgs - start_rgno;
+ target_rgno = start_rgno + I_INO(ip) % set_size;
+
+ return (xfs_rtblock_t)target_rgno <<
+ mp->m_groups[XG_TYPE_RTG].blklog;
+}
+
int
xfs_bmap_rtalloc(
struct xfs_bmalloca *ap)
@@ -2174,10 +2196,13 @@ xfs_bmap_rtalloc(
bool noalign = false;
bool initial_user_data =
ap->datatype & XFS_ALLOC_INITIAL_USER_DATA;
+ unsigned int stream_id;
int error;
ASSERT(!xfs_has_zoned(ap->tp->t_mountp));
+ stream_id = READ_ONCE(VFS_I(ap->ip)->i_write_stream);
+
retry:
error = xfs_rtallocate_align(ap, &ralen, &raminlen, &prod, &noalign);
if (error)
@@ -2185,6 +2210,8 @@ xfs_bmap_rtalloc(
if (xfs_bmap_adjacent(ap))
bno_hint = ap->blkno;
+ else if (stream_id && xfs_has_rtgroups(ap->ip->i_mount))
+ bno_hint = xfs_bmap_write_stream_rtbno(ap->ip, stream_id);
if (xfs_has_rtgroups(ap->ip->i_mount)) {
error = xfs_rtallocate_rtgs(ap->tp, bno_hint, raminlen, ralen,
diff --git a/fs/xfs/xfs_super.c b/fs/xfs/xfs_super.c
index 443b3d847150..5c92f8ebd8a7 100644
--- a/fs/xfs/xfs_super.c
+++ b/fs/xfs/xfs_super.c
@@ -637,6 +637,13 @@ xfs_setup_devices(
if (error)
return error;
xfs_update_bdi_rahead(mp);
+
+ if (xfs_has_rtgroups(mp) && !xfs_has_zoned(mp)) {
+ error = xfs_buftarg_init_streams(mp->m_rtdev_targp,
+ mp->m_sb.sb_rgcount);
+ if (error)
+ return error;
+ }
}
return 0;
--
2.25.1
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH v5 7/8] iomap: introduce and propagate write_stream
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
` (5 preceding siblings ...)
2026-09-21 9:31 ` [PATCH v5 6/8] xfs: support write streams on realtime volumes Kanchan Joshi
@ 2026-09-21 9:31 ` Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 8/8] xfs: support hardware write streams Kanchan Joshi
` (2 subsequent siblings)
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:31 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi, Anuj Gupta
Add a new write_stream field to struct iomap. Propagate write_stream from
iomap to bio in both direct I/O and buffered writeback paths.
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
---
fs/iomap/direct-io.c | 1 +
fs/iomap/ioend.c | 3 +++
include/linux/iomap.h | 1 +
3 files changed, 5 insertions(+)
diff --git a/fs/iomap/direct-io.c b/fs/iomap/direct-io.c
index 8b4039d16ce8..f406a7bc8804 100644
--- a/fs/iomap/direct-io.c
+++ b/fs/iomap/direct-io.c
@@ -349,6 +349,7 @@ static ssize_t iomap_dio_bio_iter_one(struct iomap_iter *iter,
fscrypt_set_bio_crypt_ctx(bio, iter->inode, pos, GFP_KERNEL);
bio->bi_iter.bi_sector = iomap_sector(&iter->iomap, pos);
bio->bi_write_hint = iter->inode->i_write_hint;
+ bio->bi_write_stream = iter->iomap.write_stream;
bio->bi_ioprio = dio->iocb->ki_ioprio;
bio->bi_private = dio;
bio->bi_end_io = iomap_dio_bio_end_io;
diff --git a/fs/iomap/ioend.c b/fs/iomap/ioend.c
index 7bbbb417f915..31a229c12b13 100644
--- a/fs/iomap/ioend.c
+++ b/fs/iomap/ioend.c
@@ -166,6 +166,7 @@ static struct iomap_ioend *iomap_alloc_ioend(struct iomap_writepage_ctx *wpc,
GFP_NOFS, &iomap_ioend_bioset);
bio->bi_iter.bi_sector = iomap_sector(&wpc->iomap, pos);
bio->bi_write_hint = wpc->inode->i_write_hint;
+ bio->bi_write_stream = wpc->iomap.write_stream;
wbc_init_bio(wpc->wbc, bio);
wpc->nr_folios = 0;
return iomap_init_ioend(wpc->inode, bio, pos, ioend_flags);
@@ -189,6 +190,8 @@ static bool iomap_can_add_to_ioend(struct iomap_writepage_ctx *wpc, loff_t pos,
if (!(wpc->iomap.flags & IOMAP_F_ANON_WRITE) &&
iomap_sector(&wpc->iomap, pos) != bio_end_sector(&ioend->io_bio))
return false;
+ if (wpc->iomap.write_stream != ioend->io_bio.bi_write_stream)
+ return false;
/*
* Limit ioend bio chain lengths to minimise IO completion latency. This
* also prevents long tight loops ending page writeback on all the
diff --git a/include/linux/iomap.h b/include/linux/iomap.h
index bc7ae6327dbf..f923778f08bd 100644
--- a/include/linux/iomap.h
+++ b/include/linux/iomap.h
@@ -133,6 +133,7 @@ struct iomap {
u64 length; /* length of mapping, bytes */
u16 type; /* type of mapping */
u16 flags; /* flags for mapping */
+ u16 write_stream; /* write stream for I/O */
struct block_device *bdev; /* block device for I/O */
struct dax_device *dax_dev; /* dax_dev for dax operations */
void *inline_data;
--
2.25.1
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH v5 8/8] xfs: support hardware write streams
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
` (6 preceding siblings ...)
2026-09-21 9:31 ` [PATCH v5 7/8] iomap: introduce and propagate write_stream Kanchan Joshi
@ 2026-09-21 9:31 ` Kanchan Joshi
2026-09-21 9:43 ` [PATCH v5 0/8] xfs " Kanchan Joshi
2026-09-21 23:23 ` Dave Chinner
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:31 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi, Anuj Gupta
Use the stream count reported by the data and realtime block devices
where they have one, falling back to the geometry-derived count. The
count is capped at the group count so that every stream backs onto at
least one group; a device with more streams than the filesystem has
groups loses the excess, so mkfs should create at least as many groups
as the device has streams.
Set iomap->write_stream from the inode so that the id reaches the
device.
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
---
fs/xfs/xfs_buf.c | 7 ++++++-
fs/xfs/xfs_iomap.c | 1 +
2 files changed, 7 insertions(+), 1 deletion(-)
diff --git a/fs/xfs/xfs_buf.c b/fs/xfs/xfs_buf.c
index 4e84528b05cb..1748be9aa7d6 100644
--- a/fs/xfs/xfs_buf.c
+++ b/fs/xfs/xfs_buf.c
@@ -1709,8 +1709,13 @@ xfs_buftarg_init_streams(
unsigned int nr_groups)
{
unsigned int nr_streams;
+ unsigned int hw_streams;
- nr_streams = xfs_sw_write_stream_count(nr_groups);
+ hw_streams = bdev_max_write_streams(btp->bt_bdev);
+ if (hw_streams)
+ nr_streams = min(hw_streams, nr_groups);
+ else
+ nr_streams = xfs_sw_write_stream_count(nr_groups);
return write_stream_pool_init(&btp->bt_stream_pool, nr_streams);
}
diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
index 7c6238fed61e..3e9df790a544 100644
--- a/fs/xfs/xfs_iomap.c
+++ b/fs/xfs/xfs_iomap.c
@@ -144,6 +144,7 @@ xfs_bmbt_to_iomap(
}
iomap->offset = XFS_FSB_TO_B(mp, imap->br_startoff);
iomap->length = XFS_FSB_TO_B(mp, imap->br_blockcount);
+ iomap->write_stream = READ_ONCE(VFS_I(ip)->i_write_stream);
if (mapping_flags & IOMAP_DAX) {
iomap->dax_dev = target->bt_daxdev;
} else {
--
2.25.1
^ permalink raw reply related [flat|nested] 12+ messages in thread* Re: [PATCH v5 0/8] xfs write streams
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
` (7 preceding siblings ...)
2026-09-21 9:31 ` [PATCH v5 8/8] xfs: support hardware write streams Kanchan Joshi
@ 2026-09-21 9:43 ` Kanchan Joshi
2026-09-21 23:23 ` Dave Chinner
9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21 9:43 UTC (permalink / raw)
To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
Cc: linux-xfs, linux-fsdevel, gost.dev
Patches are on top of 7.3-rc3.
Branch: https://github.com/SamsungDS/linux/tree/feat/xfs_fdp_v5
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH v5 0/8] xfs write streams
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
` (8 preceding siblings ...)
2026-09-21 9:43 ` [PATCH v5 0/8] xfs " Kanchan Joshi
@ 2026-09-21 23:23 ` Dave Chinner
2026-09-25 15:13 ` Kanchan Joshi
9 siblings, 1 reply; 12+ messages in thread
From: Dave Chinner @ 2026-09-21 23:23 UTC (permalink / raw)
To: Kanchan Joshi
Cc: brauner, hch, djwong, cem, jack, axboe, kbusch, linux-xfs,
linux-fsdevel, gost.dev
On Mon, Sep 21, 2026 at 03:01:39PM +0530, Kanchan Joshi wrote:
> This series introduces a generic interface [2,3] for write stream management on
> files.
> It enables spatial isolation (at sw and hw level) and concurrency improvments [1]
> in xfs by
> (a) steering each write stream to its own set of allocation groups (patch #5)
> and realtime groups (patch #6).
> (b) connecting xfs write streams to block write streams (FDP capable NVMe).
>
> Write streams allow the abstraction provider (fs, block, raid etc.) to
> leverage application's intent (file relationships/lifecycle).
> - application: reserves a stream and sets it on the files it wants
> placed together.
> - xfs: maps streams to AGs/RGs; allocates without interleaving; gains
> higher concurrency due to reduced lock contention.
> - hardware: maps streams to underlying allocation unit; reduces device
> internal write amplification, improved life, predictable QoS.
>
> Also
> - A stream is handed out using an fd. The filesystem is free to map streams
> onto its own geometry, and the reservation/fd machinery is generic
> (fs/write_streams.c) so other filesystems can reuse it.
>
> - Since high-level write stream (in xfs) and logical placement can work
> without the low-level write streams (in block device), series has a general
> value beyond the hardware that provides spatial isolation. Patches 1-7 are
> software only and work on any block device; patch 8 aligns the stream
> count to the hardware streams when the device has them.
So, how would I set up a stream that directs all writes to AG 1, and
returns ENOSPC to write operations if that AG is full?
What about having several streams, each pointing at a different,
known AG that the application directly controls (i.e. a known 1:1
mapping between stream_fd and agno)?
This is functionality that we could use in xfs_fsr to get rid of all
the historic tmpdir/tmpfile heuristics that allow it to "control"
locality of the data placement of files that it defragments. We also
need such control of data placement to empty AGs for shrink
operations. Yes, this only occurred to me a couple of days ago, but
now that I've made the connection between write streams and fs
allocation policy direction, it seems like a natural fit.
FWIW, that also means that, for XFS, write streams need to be
applicable to directories, so that directory block allocation also
gets placed according to the stream ID, and that new inodes in that
directory are created in the same AG as the stream ID points to.
This directly allows us to rebuild directories using FICLONE/UNSHARE
tricks and have all the indoes, data and metadata placed in the AG
we desire (i.e. necessary functionality for online shrink).
So:
> [2]
> ### Application interface
>
> Three new ioctls:
> FS_IOC_WRITE_STREAM_GET_MAX number of streams the filesystem offers
> FS_IOC_WRITE_STREAM_ALLOC reserve a stream, returns an fd
> FS_IOC_WRITE_STREAM_SET set or clear a stream on a file
How does this interface enable such usage of write streams to direct
purely filesystem level allocation policy requirements? I don't see
how I can use this API to direct where in the filesystem to map the
stream to. I can see that there is a 'set' command, but all it has
is a flags field. There's nothing passed to the ALLOC command to
allow the application to indicate to the filesystem where it wants
the new stream to point to.
IOWs, it appears taht there is no way to provide the FS with any
sort of direction as to how the stream should be set up. If we want
a FS specific stream (e.g. an AG) then we don't want it mapped to a
hardware stream, and we don't want it mapped to some random set of
AGs (like this patchset implements). The only real reason for
filesystem level write streams is to expose allocation locality
control to applications, so it seems kinda silly to have an API that
prevents any real application level control...
Indeed, how do we ask for a fs-level write stream instead of a
hardware-level write stream? They are different things, and have
different use cases (obviously!) so there definitely needs to be
some level of though put into this. And, FWIW, the write stream id
will probably need to be a u32 if we are going to support FS level
write streams, as we can have more than 65536 AGs in a
filesystem....
> - Usage model: application needs to get a handle (fd) for a write stream
> before being able to use it. This avoids multi-application conflicts.
Why can't multiple applications use the same write stream locality
mapping independently? If the application wants an exclusive write
stream (i.e. exclusive access to a set of AGs in the filesystem),
then surely that's a flag for the ALLOC API, right?
Cheers,
Dave.
--
Dave Chinner
dgc@kernel.org
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH v5 0/8] xfs write streams
2026-09-21 23:23 ` Dave Chinner
@ 2026-09-25 15:13 ` Kanchan Joshi
0 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-25 15:13 UTC (permalink / raw)
To: Dave Chinner
Cc: brauner, hch, djwong, cem, jack, axboe, kbusch, linux-xfs,
linux-fsdevel, gost.dev
On 9/22/2026 4:53 AM, Dave Chinner wrote:
> On Mon, Sep 21, 2026 at 03:01:39PM +0530, Kanchan Joshi wrote:
>> This series introduces a generic interface [2,3] for write stream management on
>> files.
>> It enables spatial isolation (at sw and hw level) and concurrency improvments [1]
>> in xfs by
>> (a) steering each write stream to its own set of allocation groups (patch #5)
>> and realtime groups (patch #6).
>> (b) connecting xfs write streams to block write streams (FDP capable NVMe).
>>
>> Write streams allow the abstraction provider (fs, block, raid etc.) to
>> leverage application's intent (file relationships/lifecycle).
>> - application: reserves a stream and sets it on the files it wants
>> placed together.
>> - xfs: maps streams to AGs/RGs; allocates without interleaving; gains
>> higher concurrency due to reduced lock contention.
>> - hardware: maps streams to underlying allocation unit; reduces device
>> internal write amplification, improved life, predictable QoS.
>>
>> Also
>> - A stream is handed out using an fd. The filesystem is free to map streams
>> onto its own geometry, and the reservation/fd machinery is generic
>> (fs/write_streams.c) so other filesystems can reuse it.
>>
>> - Since high-level write stream (in xfs) and logical placement can work
>> without the low-level write streams (in block device), series has a general
>> value beyond the hardware that provides spatial isolation. Patches 1-7 are
>> software only and work on any block device; patch 8 aligns the stream
>> count to the hardware streams when the device has them.
>
> So, how would I set up a stream that directs all writes to AG 1, and
> returns ENOSPC to write operations if that AG is full?
You can't, and that's a deliberate and necessary design choice for the
usecase (and I mentioned soft boundary in patch #5). While AG/RG are
fixed-sized XFS buckets, stream is a high-level (vfs) bucket with its
opaque handle (fd) that does not come with fixed size constraint.
Applications may tag many files (more data) with one stream and tag few
(less data) with another stream.
> What about having several streams, each pointing at a different,
> known AG that the application directly controls (i.e. a known 1:1
> mapping between stream_fd and agno)?
We can't use agno here.
Stream lets the application state intent: these files belong together,
those stay apart. It does not dictate layout. AG/RG is xfs-only
vocabulary, and that's fine for xfs-only interface. But we are doing a
generic interface that needs to remain portable across filesystems.
> This is functionality that we could use in xfs_fsr to get rid of all
> the historic tmpdir/tmpfile heuristics that allow it to "control"
> locality of the data placement of files that it defragments. We also
> need such control of data placement to empty AGs for shrink
> operations. Yes, this only occurred to me a couple of days ago, but
> now that I've made the connection between write streams and fs
> allocation policy direction, it seems like a natural fit.
>
> FWIW, that also means that, for XFS, write streams need to be
> applicable to directories, so that directory block allocation also
> gets placed according to the stream ID, and that new inodes in that
> directory are created in the same AG as the stream ID points to.
> This directly allows us to rebuild directories using FICLONE/UNSHARE
> tricks and have all the indoes, data and metadata placed in the AG
> we desire (i.e. necessary functionality for online shrink).
With this, we are discussing xfs tools (and not fs-agnostic
applications) that speak nitty-gritty of xfs layout. This usecase does
not require stream and stream-fd based workflow of this series; it can
be served more cleanly with XFS only ioctl. Something like:
XFS_IOC_SET_GROUP(AG/RG no, flags) on the file
To set the 'hard' allocation-directive so that all allocations happens
from that AG/RG no and ENOSPC if not.
Also it seems shrink usecase will also require that target AG remains
exclusive (not available for allocations) while it is getting emptied.
That, exclusive capacity locking, also we don't do with write-stream. We
can't return ENOSPC to other allocators just because somebody decided to
use/abuse stream to create artificial lack of free space.
> So:
>
>> [2]
>> ### Application interface
>>
>> Three new ioctls:
>> FS_IOC_WRITE_STREAM_GET_MAX number of streams the filesystem offers
>> FS_IOC_WRITE_STREAM_ALLOC reserve a stream, returns an fd
>> FS_IOC_WRITE_STREAM_SET set or clear a stream on a file
>
> How does this interface enable such usage of write streams to direct
> purely filesystem level allocation policy requirements?
As mentioned above, you might agree that interface should remain
portable across filesystems.
> I don't see
> how I can use this API to direct where in the filesystem to map the
> stream to. I can see that there is a 'set' command, but all it has
> is a flags field. There's nothing passed to the ALLOC command to
> allow the application to indicate to the filesystem where it wants
> the new stream to point to.
FWIW, we had a 'stream-id' as part of ALLOC in the previous version.
Christoph suggested to remove that from UAPI.
https://lore.kernel.org/linux-xfs/20260825065533.GA24808@lst.de/
The direction has been to be more opaque and less explicit.
> IOWs, it appears taht there is no way to provide the FS with any
> sort of direction as to how the stream should be set up. If we want
> a FS specific stream (e.g. an AG) then we don't want it mapped to a
> hardware stream, and we don't want it mapped to some random set of
> AGs (like this patchset implements). The only real reason for
> filesystem level write streams is to expose allocation locality
> control to applications, so it seems kinda silly to have an API that
> prevents any real application level control...
> Indeed, how do we ask for a fs-level write stream instead of a
> hardware-level write stream? They are different things, and have
> different use cases (obviously!) so there definitely needs to be
> some level of though put into this. And, FWIW, the write stream id
> will probably need to be a u32 if we are going to support FS level
> write streams, as we can have more than 65536 AGs in a
> filesystem....
All this gets handled cleanly with the above xfs-ioctl based approach.
In that, one can put the actual AG/RG no into xfs-inode itself so that
future allocation can happen from that.
>
>> - Usage model: application needs to get a handle (fd) for a write stream
>> before being able to use it. This avoids multi-application conflicts.
>
> Why can't multiple applications use the same write stream locality
> mapping independently? If the application wants an exclusive write
> stream (i.e. exclusive access to a set of AGs in the filesystem),
> then surely that's a flag for the ALLOC API, right?
Fair; But I should clarity that exclusive access is only for stream
resource/handle (fd) so that another unrelated application does not pick
the same stream and start colliding. This matters more for hardware
stream. But exclusivity is not about AGs, we don't lock AGs on
per-stream basis.
^ permalink raw reply [flat|nested] 12+ messages in thread