Linux filesystem development
 help / color / mirror / Atom feed
* [PATCH v5 0/8] xfs write streams
       [not found] <CGME20260921093229epcas5p387ee10f88335ddc5fc930ca919769a60@epcas5p3.samsung.com>
@ 2026-09-21  9:31 ` Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 1/8] fs: add write-stream management ioctls Kanchan Joshi
                     ` (9 more replies)
  0 siblings, 10 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi

This series introduces a generic interface [2,3] for write stream management on
files.
It enables spatial isolation (at sw and hw level) and concurrency improvments [1]
in xfs by
(a) steering each write stream to its own set of allocation groups (patch #5)
    and realtime groups (patch #6).
(b) connecting xfs write streams to block write streams (FDP capable NVMe).

Write streams allow the abstraction provider (fs, block, raid etc.) to
leverage application's intent (file relationships/lifecycle).
- application: reserves a stream and sets it on the files it wants
  placed together.
- xfs: maps streams to AGs/RGs; allocates without interleaving; gains
  higher concurrency due to reduced lock contention.
- hardware: maps streams to underlying allocation unit; reduces device
  internal write amplification, improved life, predictable QoS.

Also
- A stream is handed out using an fd. The filesystem is free to map streams
  onto its own geometry, and the reservation/fd machinery is generic
 (fs/write_streams.c) so other filesystems can reuse it.

- Since high-level write stream (in xfs) and logical placement can work
  without the low-level write streams (in block device), series has a general
  value beyond the hardware that provides spatial isolation. Patches 1-7 are
  software only and work on any block device; patch 8 aligns the stream
  count to the hardware streams when the device has them.

- write-stream is different from existing 'filestream' allocator which
  maintains directory-to-AG associations in a global MRU cache. That
  requires state managment and memory (and its reclaim). Proposed AG-set
  based steering relies on simple, statless/lockless airthmatic that aligns
  more with the default allocator heuristics.

[1]
### Performance

1. On regular NVMe

a. Inter-stream concurrency
---------------------------
fio: 4k write, direct IO, 16 jobs, 1 directory, 16 files * 8GiB, iodepth 32
xfs: 16 AGs, 4 write-streams

base: 41 KIOPS
write-stream AG-set: 227 KIOPS (+453%)
here, 16 files are assigned 4 distinct write-streams (4 files/stream)

b. Intra-stream concurrency
----------------------------
fio: 4k write, direct IO, 4 jobs, 1 directory, 4 files * 8GiB, iodepth 32
xfs: 16 AGs, 4 write-streams, write-stream AG-set size = 4

base: 59 KIOPS
write-stream AG-set: 112 KIOPS (+89%)
here, 4 files are assigned single write-stream

2. On FDP-capable NVMe:
RocksDB YCSB
WAF (base vs write-stream): 35% Reduction


[2]
### Application interface

Three new ioctls:
      FS_IOC_WRITE_STREAM_GET_MAX  number of streams the filesystem offers
      FS_IOC_WRITE_STREAM_ALLOC    reserve a stream, returns an fd
      FS_IOC_WRITE_STREAM_SET      set or clear a stream on a file

### Comparison with Write Hints (RWH_WRITE_LIFE_*)

- Semantics: Write Hints describe 'data temperature' (e.g.,short/long/extreme),
  implying a lifetime. Write Streams describe 'data placement'
  (e.g., Bin 1/Bin 2), implying only separation.

- Scalability: Write Hints are limited to a small, fixed enum (6 values).
  Write streams are dynamic, provider-dependent values that can scale much
  higher.

- Discovery: The existing write-hint interface is advisory and decoupled
  from underlying capabilties; application has no way to probe support
  and cannot deterministically know which hints are valid. OTOH, write-streams
  provide explicit discovery.

- Usage model: application needs to get a handle (fd) for a write stream
  before being able to use it. This avoids multi-application conflicts.

Note: within the kernel, the separation between two constructs
(write-hint and write-stream) had started from 6.16 itself.

### Changelog

since v4:
https://lore.kernel.org/linux-xfs/20260717125538.508925-1-joshi.k@samsung.com/

- stream reservation and fd handling moved to generic code, and the
  stream id moved to struct inode from xfs inode (Darrick, Christoph)
- stream id no longer shows in uapi: OPEN becomes ALLOC and returns
  fd for any (unreserved) stream; OPEN_EXACT and GET are gone (Christoph)
- ALLOC requires CAP_SYS_ADMIN (Christoph)
- write stream and write life time hint are mutually exclusive at both
  setters
- stream pool lives on the buftarg (Christoph), which gives the realtime
  device streams of its own: new patch to enable RT support (Christoph, Darrick)
- software streams come first and work on any device; hardware streams
  are enabled in the last (Christoph)

since v3:
https://lore.kernel.org/linux-block/20260616180555.33338-1-joshi.k@samsung.com/
- add fd-based interface to open/set the write stream (Christoph)
- move from single multiplexed ioctl to 4 distinct ioctls (Christoph)
- add mutual exclusion checks against existing write-hint, filestream (Christoph)
- uint16_t for write-stream within iomap and other streamlining (Darrick)

since v2:
https://lore.kernel.org/linux-fsdevel/20260309052944.156054-1-joshi.k@samsung.com/
- xfs default allocator optimization using fixed-size generic AG set (Dave)
- reuse the above to simplify the write-stream AG set handling
- streamline the uapi; Use union for GET_MAX and GET/SET (Darrick)
- uint16_t for write-stream within xfs inode and other cleanups (Darrick)

since v1:
https://lore.kernel.org/linux-fsdevel/20260216052540.217920-1-joshi.k@samsung.com/
- swich from fcntl based to ioctl-based interface (Christian)
- new patch (#4) that makes xfs allocator use the write streams for AG
  selection
- new patch (#5) that introduces software write streams in xfs.

[3]
### Interface example
/*
 * Minimal write-stream example.
 *
 *      cc -Wall -o ws-example ws-example.c
 *      sudo ./ws-example /mnt/xfs        # data device
 *      sudo ./ws-example /mnt/xfs rt     # realtime volume
 *
 * Reserves two streams.  a.dat gets the first; b.dat and c.dat share the
 * second.  Each file then takes one 4K write, so a.dat is placed apart from
 * b.dat and c.dat, which are placed together.
 *
 * With "rt" the files are made realtime first.
 */

#include <sys/ioctl.h>
#include <err.h>
#include <fcntl.h>
#include <stdio.h>
#include <string.h>
#include <unistd.h>
#include <linux/fs.h>

struct fs_write_stream_set {
        __s32   stream_fd;      /* IN: stream to set, -1 to clear */
        __u32   flags;          /* IN: FS_WRITE_STREAM_SET_* */
};

#define FS_WRITE_STREAM_SET_CLEAR       (1U << 0)

#define FS_IOC_WRITE_STREAM_GET_MAX     _IOR(0x15, 3, __u32)
#define FS_IOC_WRITE_STREAM_ALLOC       _IO(0x15, 4)    /* returns an fd */
#define FS_IOC_WRITE_STREAM_SET         _IOW(0x15, 5, struct fs_write_stream_set)

static char buf[4096];

/* create @name; with @rt, put it on the realtime volume */
static int create_file(const char *dir, const char *name, int rt)
{
        struct fsxattr fa;
        char path[512];
        int fd;

        snprintf(path, sizeof(path), "%s/%s", dir, name);

        fd = open(path, O_CREAT | O_TRUNC | O_RDWR, 0644);
        if (fd < 0)
                err(1, "%s", path);
        if (!rt)
                return fd;

        if (ioctl(fd, FS_IOC_FSGETXATTR, &fa))
                err(1, "FSGETXATTR %s", name);
        fa.fsx_xflags |= FS_XFLAG_REALTIME;
        if (ioctl(fd, FS_IOC_FSSETXATTR, &fa))
                err(1, "FSSETXATTR %s (no realtime volume?)", name);
        return fd;
}

/* set @stream_fd on @fd, write 4K */
static void write_on_stream(int fd, const char *name, int stream_fd)
{
        struct fs_write_stream_set set = { .stream_fd = stream_fd, .flags = 0 };

        if (ioctl(fd, FS_IOC_WRITE_STREAM_SET, &set))
                err(1, "SET %s", name);
        if (write(fd, buf, sizeof(buf)) != sizeof(buf))
                err(1, "write %s", name);
        fsync(fd);
        close(fd);
}

int main(int argc, char **argv)
{
        __u32 max = 0;
        int rt, fa, fb, fc, s1, s2;

        if (argc < 2 || argc > 3 || (argc == 3 && strcmp(argv[2], "rt")))
                errx(2, "usage: %s <dir-on-xfs> [rt]", argv[0]);
        rt = argc == 3;

        fa = create_file(argv[1], "a.dat", rt);
        fb = create_file(argv[1], "b.dat", rt);
        fc = create_file(argv[1], "c.dat", rt);

        /* query and reserve through a file on the device being placed on */
        if (ioctl(fa, FS_IOC_WRITE_STREAM_GET_MAX, &max))
                err(1, "GET_MAX");
        printf("max streams: %u\n", max);
        if (max < 2)
                errx(1, "need at least 2 streams");

        /* ALLOC returns the stream fd as the ioctl return value */
        s1 = ioctl(fa, FS_IOC_WRITE_STREAM_ALLOC);
        if (s1 < 0)
                err(1, "ALLOC");
        s2 = ioctl(fa, FS_IOC_WRITE_STREAM_ALLOC);
        if (s2 < 0)
                err(1, "ALLOC");

        write_on_stream(fa, "a.dat", s1);
        write_on_stream(fb, "b.dat", s2);
        write_on_stream(fc, "c.dat", s2);

        /* closing a stream fd releases the stream */
        close(s1);
        close(s2);

        puts("a.dat on stream 1; b.dat and c.dat on stream 2");
        puts("confirm AG/RG steering using xfs_bmap");
        return 0;
}

Anuj Gupta (4):
  fs: add write-stream management ioctls
  fs: add generic write-stream management
  xfs: implement software write-stream management
  xfs: support write streams on realtime volumes

Kanchan Joshi (4):
  fs: add i_write_stream, exclusive with the write life time hint
  xfs: write stream based AG placement
  iomap: introduce and propagate write_stream
  xfs: support hardware write streams

 fs/Makefile                   |   2 +-
 fs/fcntl.c                    |   7 ++
 fs/inode.c                    |   1 +
 fs/iomap/direct-io.c          |   1 +
 fs/iomap/ioend.c              |   3 +
 fs/write_streams.c            | 138 ++++++++++++++++++++++++++++++++++
 fs/xfs/libxfs/xfs_bmap.c      |  58 +++++++++++++-
 fs/xfs/xfs_buf.c              |  35 +++++++++
 fs/xfs/xfs_buf.h              |   6 ++
 fs/xfs/xfs_inode.c            |  69 +++++++++++++++++
 fs/xfs/xfs_inode.h            |   4 +
 fs/xfs/xfs_ioctl.c            |  80 ++++++++++++++++++++
 fs/xfs/xfs_iomap.c            |   1 +
 fs/xfs/xfs_rtalloc.c          |  27 +++++++
 fs/xfs/xfs_super.c            |  12 +++
 include/linux/fs.h            |   2 +-
 include/linux/iomap.h         |   1 +
 include/linux/write_streams.h |  33 ++++++++
 include/uapi/linux/fs.h       |  15 ++++
 19 files changed, 491 insertions(+), 4 deletions(-)
 create mode 100644 fs/write_streams.c
 create mode 100644 include/linux/write_streams.h

-- 
2.25.1


^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH v5 1/8] fs: add write-stream management ioctls
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
@ 2026-09-21  9:31   ` Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 2/8] fs: add generic write-stream management Kanchan Joshi
                     ` (8 subsequent siblings)
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Anuj Gupta, Kanchan Joshi

From: Anuj Gupta <anuj20.g@samsung.com>

Add three ioctls for write stream management:

  FS_IOC_WRITE_STREAM_GET_MAX  number of streams the filesystem offers
  FS_IOC_WRITE_STREAM_ALLOC    reserve a stream, returns an fd
  FS_IOC_WRITE_STREAM_SET      set or clear a stream on a file

Write streams are a resource the kernel manages and applications ask
for.

ALLOC reserves a stream and returns a stream fd; the stream stays
reserved as long as that fd is open. ALLOC fails once all streams are
reserved.

SET assigns the stream named by the stream_fd in struct
fs_write_stream_set to an open file; the same stream fd can be used to
set the stream on any number of files. SET with FS_WRITE_STREAM_SET_CLEAR
and a stream_fd of -1 clears whatever stream the file carries.

Suggested-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
---
 include/uapi/linux/fs.h | 15 +++++++++++++++
 1 file changed, 15 insertions(+)

diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index 34c6f219462a..1974052aa476 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -345,6 +345,21 @@ struct file_attr {
 /* Get logical block metadata capability details */
 #define FS_IOC_GETLBMD_CAP		_IOWR(0x15, 2, struct logical_block_metadata_cap)
 
+struct fs_write_stream_set {
+	__s32		stream_fd;	/* IN: stream to attach, -1 to clear */
+	__u32		flags;		/* IN: FS_WRITE_STREAM_SET_* */
+};
+
+/* Clear the file's write stream. Requires stream_fd to be -1. */
+#define FS_WRITE_STREAM_SET_CLEAR	(1U << 0)
+
+/* GET_MAX returns the count through its argument; ALLOC returns a stream fd
+ * as the ioctl return value and takes none.
+ */
+#define FS_IOC_WRITE_STREAM_GET_MAX	_IOR(0x15, 3, __u32)
+#define FS_IOC_WRITE_STREAM_ALLOC	_IO(0x15, 4)
+#define FS_IOC_WRITE_STREAM_SET		_IOW(0x15, 5, struct fs_write_stream_set)
+
 /*
  * Inode flags (FS_IOC_GETFLAGS / FS_IOC_SETFLAGS)
  *
-- 
2.25.1


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 2/8] fs: add generic write-stream management
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 1/8] fs: add write-stream management ioctls Kanchan Joshi
@ 2026-09-21  9:31   ` Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint Kanchan Joshi
                     ` (7 subsequent siblings)
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Anuj Gupta, Kanchan Joshi

From: Anuj Gupta <anuj20.g@samsung.com>

Add a bitmap-based stream pool and anonymous fds for reserved streams.
Helpers initialize and destroy a pool, reserve a stream as an fd, and
resolve a stream fd back to its id.

Suggested-by: Christoph Hellwig <hch@lst.de>
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
---
 fs/Makefile                   |   2 +-
 fs/write_streams.c            | 138 ++++++++++++++++++++++++++++++++++
 include/linux/write_streams.h |  33 ++++++++
 3 files changed, 172 insertions(+), 1 deletion(-)
 create mode 100644 fs/write_streams.c
 create mode 100644 include/linux/write_streams.h

diff --git a/fs/Makefile b/fs/Makefile
index 055dfc23d82b..494da034cef2 100644
--- a/fs/Makefile
+++ b/fs/Makefile
@@ -23,7 +23,7 @@ obj-$(CONFIG_PROC_FS)		+= proc_namespace.o
 obj-$(CONFIG_LEGACY_DIRECT_IO)	+= direct-io.o
 obj-y				+= notify/
 obj-$(CONFIG_EPOLL)		+= eventpoll.o
-obj-y				+= anon_inodes.o
+obj-y				+= anon_inodes.o write_streams.o
 obj-$(CONFIG_SIGNALFD)		+= signalfd.o
 obj-$(CONFIG_TIMERFD)		+= timerfd.o
 obj-$(CONFIG_EVENTFD)		+= eventfd.o
diff --git a/fs/write_streams.c b/fs/write_streams.c
new file mode 100644
index 000000000000..b7f6ca6cf725
--- /dev/null
+++ b/fs/write_streams.c
@@ -0,0 +1,138 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Generic write-stream fd management. */
+#include <linux/anon_inodes.h>
+#include <linux/bitmap.h>
+#include <linux/err.h>
+#include <linux/file.h>
+#include <linux/fs.h>
+#include <linux/mount.h>
+#include <linux/slab.h>
+#include <linux/write_streams.h>
+
+struct write_stream {
+	struct write_stream_pool	*pool;
+	struct vfsmount			*mnt;	/* pins the mount */
+	u16				id;	/* 1-based slot number */
+};
+
+static int write_stream_release(struct inode *inode, struct file *file)
+{
+	struct write_stream *ws = file->private_data;
+
+	spin_lock(&ws->pool->lock);
+	__clear_bit(ws->id - 1, ws->pool->streams_in_use);
+	spin_unlock(&ws->pool->lock);
+	mntput(ws->mnt);
+	kfree(ws);
+	return 0;
+}
+
+static const struct file_operations write_stream_fops = {
+	.release	= write_stream_release,
+	.llseek		= noop_llseek,
+};
+
+/**
+ * write_stream_pool_init - size a write stream pool
+ * @pool: pool to initialise
+ * @nr: number of slots, may be zero
+ *
+ * Returns 0, or -ENOMEM if the bitmap cannot be allocated.
+ */
+int write_stream_pool_init(struct write_stream_pool *pool, unsigned int nr)
+{
+	spin_lock_init(&pool->lock);
+	pool->streams_in_use = NULL;
+	pool->nr_streams = 0;
+
+	if (!nr)
+		return 0;
+
+	pool->streams_in_use = bitmap_zalloc(nr, GFP_KERNEL);
+	if (!pool->streams_in_use)
+		return -ENOMEM;
+	pool->nr_streams = nr;
+	return 0;
+}
+EXPORT_SYMBOL_GPL(write_stream_pool_init);
+
+/**
+ * write_stream_pool_destroy - free a write stream pool
+ * @pool: pool to release
+ */
+void write_stream_pool_destroy(struct write_stream_pool *pool)
+{
+	bitmap_free(pool->streams_in_use);
+	pool->streams_in_use = NULL;
+	pool->nr_streams = 0;
+}
+EXPORT_SYMBOL_GPL(write_stream_pool_destroy);
+
+/**
+ * write_stream_alloc_fd - reserve a slot and return an fd naming it
+ * @pool: pool to reserve from
+ * @file: file whose mount is pinned for the life of the stream
+ *
+ * Returns an O_RDONLY | O_CLOEXEC fd, -EOPNOTSUPP if the pool has no slots,
+ * or -EBUSY if they are all taken.
+ */
+int write_stream_alloc_fd(struct write_stream_pool *pool, struct file *file)
+{
+	struct write_stream *ws;
+	unsigned int slot;
+	int fd;
+
+	if (!pool->nr_streams)
+		return -EOPNOTSUPP;
+
+	ws = kzalloc_obj(*ws, GFP_KERNEL);
+	if (!ws)
+		return -ENOMEM;
+
+	spin_lock(&pool->lock);
+	slot = find_first_zero_bit(pool->streams_in_use, pool->nr_streams);
+	if (slot >= pool->nr_streams) {
+		spin_unlock(&pool->lock);
+		kfree(ws);
+		return -EBUSY;
+	}
+	__set_bit(slot, pool->streams_in_use);
+	spin_unlock(&pool->lock);
+
+	ws->pool = pool;
+	ws->mnt = mntget(file->f_path.mnt);
+	ws->id = slot + 1;
+
+	fd = anon_inode_getfd("[write_stream]", &write_stream_fops, ws,
+			      O_RDONLY | O_CLOEXEC);
+	if (fd < 0) {
+		spin_lock(&pool->lock);
+		__clear_bit(slot, pool->streams_in_use);
+		spin_unlock(&pool->lock);
+		mntput(ws->mnt);
+		kfree(ws);
+	}
+	return fd;
+}
+EXPORT_SYMBOL_GPL(write_stream_alloc_fd);
+
+/**
+ * write_stream_get_id - the id a stream file names
+ * @file: candidate stream file
+ * @pool: pool the stream must belong to, or NULL to accept any
+ *
+ * Returns the 1-based id, or -EINVAL if @file is not a stream file, or is
+ * a stream of a different pool.
+ */
+int write_stream_get_id(struct file *file, const struct write_stream_pool *pool)
+{
+	struct write_stream *ws;
+
+	if (file->f_op != &write_stream_fops)
+		return -EINVAL;
+	ws = file->private_data;
+	if (pool && ws->pool != pool)
+		return -EINVAL;
+	return ws->id;
+}
+EXPORT_SYMBOL_GPL(write_stream_get_id);
diff --git a/include/linux/write_streams.h b/include/linux/write_streams.h
new file mode 100644
index 000000000000..d27f42df5984
--- /dev/null
+++ b/include/linux/write_streams.h
@@ -0,0 +1,33 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef _LINUX_WRITE_STREAMS_H
+#define _LINUX_WRITE_STREAMS_H
+
+#include <linux/spinlock.h>
+#include <linux/types.h>
+
+struct file;
+
+/*
+ * Write stream slot pool, embedded in whatever filesystem object scopes
+ * an id space.
+ */
+struct write_stream_pool {
+	unsigned int		nr_streams;
+	unsigned long		*streams_in_use;
+	spinlock_t		lock;
+};
+
+static inline unsigned int
+write_stream_pool_count(const struct write_stream_pool *pool)
+{
+	return pool->nr_streams;
+}
+
+int  write_stream_pool_init(struct write_stream_pool *pool, unsigned int nr);
+void write_stream_pool_destroy(struct write_stream_pool *pool);
+
+int  write_stream_alloc_fd(struct write_stream_pool *pool, struct file *file);
+int  write_stream_get_id(struct file *file,
+			 const struct write_stream_pool *pool);
+
+#endif /* _LINUX_WRITE_STREAMS_H */
-- 
2.25.1


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 1/8] fs: add write-stream management ioctls Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 2/8] fs: add generic write-stream management Kanchan Joshi
@ 2026-09-21  9:31   ` Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 4/8] xfs: implement software write-stream management Kanchan Joshi
                     ` (6 subsequent siblings)
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi, Anuj Gupta

Keep the attached stream id in the VFS inode rather than the
filesystem's own, so that fcntl_set_rw_hint() can see it. It is a u16 in
the 32-bit hole after i_state, so struct inode does not grow.

Suggested-by: Christoph Hellwig <hch@lst.de>
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
---
 fs/fcntl.c         | 7 +++++++
 fs/inode.c         | 1 +
 include/linux/fs.h | 2 +-
 3 files changed, 9 insertions(+), 1 deletion(-)

diff --git a/fs/fcntl.c b/fs/fcntl.c
index c158f082f1da..a5439e076197 100644
--- a/fs/fcntl.c
+++ b/fs/fcntl.c
@@ -380,7 +380,14 @@ static long fcntl_set_rw_hint(struct file *file, unsigned long arg)
 	if (!rw_hint_valid(hint))
 		return -EINVAL;
 
+	/* keep write stream and write life time hint mutually exclusive */
+	spin_lock(&inode->i_lock);
+	if (hint != WRITE_LIFE_NOT_SET && READ_ONCE(inode->i_write_stream)) {
+		spin_unlock(&inode->i_lock);
+		return -EBUSY;
+	}
 	WRITE_ONCE(inode->i_write_hint, hint);
+	spin_unlock(&inode->i_lock);
 
 	/*
 	 * file->f_mapping->host may differ from inode. As an example,
diff --git a/fs/inode.c b/fs/inode.c
index ba7da39be4a3..50b91d449a38 100644
--- a/fs/inode.c
+++ b/fs/inode.c
@@ -246,6 +246,7 @@ int inode_init_always_gfp(struct super_block *sb, struct inode *inode, gfp_t gfp
 	atomic_set(&inode->i_writecount, 0);
 	inode->i_size = 0;
 	inode->i_write_hint = WRITE_LIFE_NOT_SET;
+	inode->i_write_stream = 0;
 	inode->i_blocks = 0;
 	inode->i_bytes = 0;
 	inode->i_generation = 0;
diff --git a/include/linux/fs.h b/include/linux/fs.h
index f9d1e05e8ae6..3d9ce4d30cd5 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -812,7 +812,7 @@ struct inode {
 
 	/* Misc */
 	struct inode_state_flags i_state;
-	/* 32-bit hole */
+	u16			i_write_stream;
 	struct rw_semaphore	i_rwsem;
 
 	unsigned long		dirtied_when;	/* jiffies of first dirtying */
-- 
2.25.1


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 4/8] xfs: implement software write-stream management
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
                     ` (2 preceding siblings ...)
  2026-09-21  9:31   ` [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint Kanchan Joshi
@ 2026-09-21  9:31   ` Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 5/8] xfs: write stream based AG placement Kanchan Joshi
                     ` (5 subsequent siblings)
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Anuj Gupta, Kanchan Joshi

From: Anuj Gupta <anuj20.g@samsung.com>

Implement the FS_IOC_WRITE_STREAM_* handlers.

Keep a stream pool in the data device buftarg, sized from the AG count.
The id is stored in inode->i_write_stream and is not written to disk.

Write streams, filestreams and write life time hints are mutually
exclusive. GET_MAX reports zero for filestream and realtime inodes, SET
rejects those and any inode carrying a hint, and setattr refuses to turn
on filestream or realtime placement while a stream is attached.

ALLOC is restricted to CAP_SYS_ADMIN.

Suggested-by: Christoph Hellwig <hch@lst.de>
Suggested-by: Darrick J. Wong <djwong@kernel.org>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
---
 fs/xfs/xfs_buf.c   | 30 ++++++++++++++++++
 fs/xfs/xfs_buf.h   |  6 ++++
 fs/xfs/xfs_inode.c | 69 ++++++++++++++++++++++++++++++++++++++++
 fs/xfs/xfs_inode.h |  4 +++
 fs/xfs/xfs_ioctl.c | 79 ++++++++++++++++++++++++++++++++++++++++++++++
 fs/xfs/xfs_super.c |  5 +++
 6 files changed, 193 insertions(+)

diff --git a/fs/xfs/xfs_buf.c b/fs/xfs/xfs_buf.c
index 8256c1d13ce2..4e84528b05cb 100644
--- a/fs/xfs/xfs_buf.c
+++ b/fs/xfs/xfs_buf.c
@@ -1647,6 +1647,7 @@ void
 xfs_free_buftarg(
 	struct xfs_buftarg	*btp)
 {
+	write_stream_pool_destroy(&btp->bt_stream_pool);
 	xfs_destroy_buftarg(btp);
 	fs_put_dax(btp->bt_daxdev, btp->bt_mount);
 	/* the main block device is closed by kill_block_super */
@@ -1684,6 +1685,35 @@ xfs_configure_buftarg_atomic_writes(
 	btp->bt_awu_max = max_bytes;
 }
 
+#define XFS_MAX_SW_WRITE_STREAMS	U8_MAX
+
+/* Heuristic to derive software write stream count from a group topology */
+static unsigned int
+xfs_sw_write_stream_count(
+	unsigned int		nr_groups)
+{
+	unsigned int		group_set_size;
+
+	if (nr_groups >= 16)
+		group_set_size = 4;
+	else if (nr_groups >= 8)
+		group_set_size = 2;
+	else
+		group_set_size = 1;
+	return min(nr_groups / group_set_size, XFS_MAX_SW_WRITE_STREAMS);
+}
+
+int
+xfs_buftarg_init_streams(
+	struct xfs_buftarg	*btp,
+	unsigned int		nr_groups)
+{
+	unsigned int		nr_streams;
+
+	nr_streams = xfs_sw_write_stream_count(nr_groups);
+	return write_stream_pool_init(&btp->bt_stream_pool, nr_streams);
+}
+
 /* Configure a buffer target that abstracts a block device. */
 int
 xfs_configure_buftarg(
diff --git a/fs/xfs/xfs_buf.h b/fs/xfs/xfs_buf.h
index a4729253b56f..b5292e5731e3 100644
--- a/fs/xfs/xfs_buf.h
+++ b/fs/xfs/xfs_buf.h
@@ -15,6 +15,7 @@
 #include <linux/uio.h>
 #include <linux/list_lru.h>
 #include <linux/lockref.h>
+#include <linux/write_streams.h>
 
 extern struct kmem_cache *xfs_buf_cache;
 
@@ -102,6 +103,9 @@ struct xfs_buftarg {
 	unsigned int		bt_awu_min;
 	unsigned int		bt_awu_max;
 
+	/* slot pool for stream fds on this device */
+	struct write_stream_pool bt_stream_pool;
+
 	struct rhashtable	bt_hash;
 };
 
@@ -360,6 +364,8 @@ extern void xfs_buftarg_wait(struct xfs_buftarg *);
 extern void xfs_buftarg_drain(struct xfs_buftarg *);
 int xfs_configure_buftarg(struct xfs_buftarg *btp, unsigned int sectorsize,
 		xfs_fsblock_t nr_blocks);
+int xfs_buftarg_init_streams(struct xfs_buftarg *btp,
+		unsigned int nr_groups);
 
 #define xfs_readonly_buftarg(buftarg)	bdev_read_only((buftarg)->bt_bdev)
 
diff --git a/fs/xfs/xfs_inode.c b/fs/xfs/xfs_inode.c
index 030a7c8f2c12..df0a37c4d68a 100644
--- a/fs/xfs/xfs_inode.c
+++ b/fs/xfs/xfs_inode.c
@@ -4,6 +4,7 @@
  * All Rights Reserved.
  */
 #include <linux/iversion.h>
+#include <linux/write_streams.h>
 
 #include "xfs_platform.h"
 #include "xfs_fs.h"
@@ -47,6 +48,74 @@
 
 struct kmem_cache *xfs_inode_cache;
 
+/* Number of write streams available to this inode. */
+int
+xfs_inode_max_write_streams(
+	struct xfs_inode	*ip)
+{
+	xfs_assert_ilocked(ip, XFS_ILOCK_SHARED | XFS_ILOCK_EXCL);
+
+	if (xfs_inode_is_filestream(ip))
+		return 0;
+	if (XFS_IS_REALTIME_INODE(ip))
+		return 0;
+	return write_stream_pool_count(&xfs_inode_buftarg(ip)->bt_stream_pool);
+}
+
+/* Bind the write stream named by @stream_fd to @ip */
+int
+xfs_inode_set_write_stream(
+	struct xfs_inode	*ip,
+	int			stream_fd)
+{
+	CLASS(fd, f)(stream_fd);
+	struct xfs_buftarg	*target;
+	int			id, error = 0;
+
+	if (!fd_file(f))
+		return -EBADF;
+	xfs_ilock(ip, XFS_ILOCK_EXCL);
+	if (XFS_IS_REALTIME_INODE(ip)) {
+		error = -EINVAL;
+		goto out_unlock;
+	}
+
+	target = xfs_inode_buftarg(ip);
+	id = write_stream_get_id(fd_file(f), &target->bt_stream_pool);
+	if (id < 0) {
+		error = id;
+		goto out_unlock;
+	}
+
+	/* Filestream and write-stream are mutually exclusive */
+	if (xfs_inode_is_filestream(ip)) {
+		error = -EINVAL;
+		goto out_unlock;
+	}
+
+	/* keep write stream and write life time hint mutually exclusive */
+	spin_lock(&VFS_I(ip)->i_lock);
+	if (VFS_I(ip)->i_write_hint != WRITE_LIFE_NOT_SET) {
+		spin_unlock(&VFS_I(ip)->i_lock);
+		error = -EBUSY;
+		goto out_unlock;
+	}
+	WRITE_ONCE(VFS_I(ip)->i_write_stream, id);
+	spin_unlock(&VFS_I(ip)->i_lock);
+out_unlock:
+	xfs_iunlock(ip, XFS_ILOCK_EXCL);
+	return error;
+}
+
+void
+xfs_inode_clear_write_stream(
+	struct xfs_inode	*ip)
+{
+	xfs_ilock(ip, XFS_ILOCK_EXCL);
+	WRITE_ONCE(VFS_I(ip)->i_write_stream, 0);
+	xfs_iunlock(ip, XFS_ILOCK_EXCL);
+}
+
 /*
  * These two are wrapper routines around the xfs_ilock() routine used to
  * centralize some grungy code.  They are used in places that wish to lock the
diff --git a/fs/xfs/xfs_inode.h b/fs/xfs/xfs_inode.h
index 1602027cd0aa..c053f105d2fa 100644
--- a/fs/xfs/xfs_inode.h
+++ b/fs/xfs/xfs_inode.h
@@ -672,4 +672,8 @@ int xfs_icreate_dqalloc(const struct xfs_icreate_args *args,
 		struct xfs_dquot **udqpp, struct xfs_dquot **gdqpp,
 		struct xfs_dquot **pdqpp);
 
+int xfs_inode_max_write_streams(struct xfs_inode *ip);
+int xfs_inode_set_write_stream(struct xfs_inode *ip, int stream_fd);
+void xfs_inode_clear_write_stream(struct xfs_inode *ip);
+
 #endif	/* __XFS_INODE_H__ */
diff --git a/fs/xfs/xfs_ioctl.c b/fs/xfs/xfs_ioctl.c
index 96ca3e480cb9..bc6957ad06cb 100644
--- a/fs/xfs/xfs_ioctl.c
+++ b/fs/xfs/xfs_ioctl.c
@@ -557,6 +557,11 @@ xfs_ioctl_setattr_xflags(
 	bool			rtflag = (fa->fsx_xflags & FS_XFLAG_REALTIME);
 	uint64_t		i_flags2;
 
+	/* refuse a filestream/realtime flag change while a stream is attached */
+	if (READ_ONCE(VFS_I(ip)->i_write_stream) &&
+	    ((fa->fsx_xflags & FS_XFLAG_FILESTREAM) || rtflag))
+		return -EINVAL;
+
 	if (rtflag != XFS_IS_REALTIME_INODE(ip)) {
 		/* Can't change realtime flag if any extents are allocated. */
 		if (xfs_inode_has_filedata(ip))
@@ -1200,6 +1205,73 @@ xfs_ioctl_fs_counts(
 	return 0;
 }
 
+static int
+xfs_ioc_write_stream_get_max(
+	struct file		*filp,
+	void __user		*arg)
+{
+	struct xfs_inode	*ip = XFS_I(file_inode(filp));
+	__u32			max;
+
+	xfs_ilock(ip, XFS_ILOCK_SHARED);
+	max = xfs_inode_max_write_streams(ip);
+	xfs_iunlock(ip, XFS_ILOCK_SHARED);
+
+	return put_user(max, (__u32 __user *)arg);
+}
+
+static int
+xfs_ioc_write_stream_alloc(
+	struct file		*filp)
+{
+	struct xfs_inode	*ip = XFS_I(file_inode(filp));
+	struct xfs_buftarg	*target;
+	int			max;
+
+	if (!capable(CAP_SYS_ADMIN))
+		return -EPERM;
+
+	xfs_ilock(ip, XFS_ILOCK_SHARED);
+	max = xfs_inode_max_write_streams(ip);
+	target = xfs_inode_buftarg(ip);
+	xfs_iunlock(ip, XFS_ILOCK_SHARED);
+
+	if (!max)
+		return -EOPNOTSUPP;
+	return write_stream_alloc_fd(&target->bt_stream_pool, filp);
+}
+
+static int
+xfs_ioc_write_stream_set(
+	struct file		*filp,
+	void __user		*arg)
+{
+	struct xfs_inode	*ip = XFS_I(file_inode(filp));
+	struct fs_write_stream_set set;
+
+	if (!(filp->f_mode & FMODE_WRITE))
+		return -EBADF;
+	if (copy_from_user(&set, arg, sizeof(set)))
+		return -EFAULT;
+	if (set.flags & ~FS_WRITE_STREAM_SET_CLEAR)
+		return -EINVAL;
+
+	/* Regular files only. */
+	if (!S_ISREG(VFS_I(ip)->i_mode))
+		return -EINVAL;
+
+	if (!inode_owner_or_capable(file_mnt_idmap(filp), VFS_I(ip)))
+		return -EPERM;
+
+	if (set.flags & FS_WRITE_STREAM_SET_CLEAR) {
+		if (set.stream_fd != -1)
+			return -EINVAL;
+		xfs_inode_clear_write_stream(ip);
+		return 0;
+	}
+	return xfs_inode_set_write_stream(ip, set.stream_fd);
+}
+
 /*
  * These long-unused ioctls were removed from the official ioctl API in 5.17,
  * but retain these definitions so that we can log warnings about them.
@@ -1466,6 +1538,13 @@ xfs_file_ioctl(
 	case XFS_IOC_VERIFY_MEDIA:
 		return xfs_ioc_verify_media(filp, arg);
 
+	case FS_IOC_WRITE_STREAM_GET_MAX:
+		return xfs_ioc_write_stream_get_max(filp, arg);
+	case FS_IOC_WRITE_STREAM_ALLOC:
+		return xfs_ioc_write_stream_alloc(filp);
+	case FS_IOC_WRITE_STREAM_SET:
+		return xfs_ioc_write_stream_set(filp, (void __user *)arg);
+
 	default:
 		return -ENOTTY;
 	}
diff --git a/fs/xfs/xfs_super.c b/fs/xfs/xfs_super.c
index b24db75eaedc..443b3d847150 100644
--- a/fs/xfs/xfs_super.c
+++ b/fs/xfs/xfs_super.c
@@ -608,6 +608,11 @@ xfs_setup_devices(
 	if (error)
 		return error;
 
+	error = xfs_buftarg_init_streams(mp->m_ddev_targp,
+			mp->m_sb.sb_agcount);
+	if (error)
+		return error;
+
 	if (mp->m_logdev_targp && mp->m_logdev_targp != mp->m_ddev_targp) {
 		unsigned int	log_sector_size = BBSIZE;
 
-- 
2.25.1


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 5/8] xfs: write stream based AG placement
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
                     ` (3 preceding siblings ...)
  2026-09-21  9:31   ` [PATCH v5 4/8] xfs: implement software write-stream management Kanchan Joshi
@ 2026-09-21  9:31   ` Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 6/8] xfs: support write streams on realtime volumes Kanchan Joshi
                     ` (4 subsequent siblings)
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi, Anuj Gupta

Choose the starting AG from the write stream set on the file.

Isolating streams into separate allocation groups reduces block
interleaving between concurrent writers, AGF lock contention and logical
file fragmentation.

AGs are partitioned among the streams. The stream value selects the AG
set and the inode number selects the AG within it, so intra-stream
concurrency comes from the AG set size.

Example: 8 Allocation Groups, 4 write streams
AG set size = 2 AGs per write stream

   Stream 1 (ID: 1)         Stream 2 (ID: 2)         Streams 3 & 4
 +---------+---------+    +---------+---------+    +-------------
 |   AG0   |   AG1   |    |   AG2   |   AG3   |    |  AG4...AG7
 +---------+---------+    +---------+---------+    +-------------
      ^         ^              ^         ^
      |         |              |         |
      | File B (ino: 101)      | File D (ino: 201)
      | 101 % 2 = 1 -> AG 1    | 201 % 2 = 1 -> AG 3
      |                        |
 File A (ino: 100)        File C (ino: 200)
 100 % 2 = 0 -> AG 0      200 % 2 = 0 -> AG 2

If the AGs do not divide evenly, the last stream absorbs the remainder.

Set boundaries are a hint, not a partition: file contiguity is preserved
and the full space stays usable with a single stream.

Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
---
 fs/xfs/libxfs/xfs_bmap.c | 58 ++++++++++++++++++++++++++++++++++++++--
 1 file changed, 56 insertions(+), 2 deletions(-)

diff --git a/fs/xfs/libxfs/xfs_bmap.c b/fs/xfs/libxfs/xfs_bmap.c
index d64defeda645..2e83811cefcb 100644
--- a/fs/xfs/libxfs/xfs_bmap.c
+++ b/fs/xfs/libxfs/xfs_bmap.c
@@ -3579,8 +3579,31 @@ xfs_bmap_btalloc_filestreams(
 	return xfs_bmap_btalloc_low_space(ap, args);
 }
 
+static xfs_agnumber_t
+xfs_bmap_write_stream_agno(
+	struct xfs_inode	*ip,
+	unsigned int		stream_id)
+{
+	struct xfs_mount	*mp = ip->i_mount;
+	xfs_agnumber_t		nr_ags = mp->m_sb.sb_agcount;
+	unsigned int		nr_streams =
+		write_stream_pool_count(&mp->m_ddev_targp->bt_stream_pool);
+	xfs_agnumber_t		ag_set_size, start_agno;
+
+	stream_id -= 1;	/* convert from 1-based to 0-based */
+	ag_set_size = nr_ags / nr_streams;
+	start_agno = stream_id * ag_set_size;
+
+	/* last stream absorbs any uneven remainder */
+	if (stream_id == nr_streams - 1)
+		ag_set_size = nr_ags - start_agno;
+
+	return start_agno + I_INO(ip) % ag_set_size;
+}
+
+/* Core AG allocator.  The caller sets ap->blkno to the target AG start. */
 static int
-xfs_bmap_btalloc_best_length(
+xfs_bmap_btalloc_from_blkno(
 	struct xfs_bmalloca	*ap,
 	struct xfs_alloc_arg	*args,
 	int			stripe_align)
@@ -3588,7 +3611,6 @@ xfs_bmap_btalloc_best_length(
 	xfs_extlen_t		blen = 0;
 	int			error;
 
-	ap->blkno = XFS_INODE_TO_FSB(ap->ip);
 	if (!xfs_bmap_adjacent(ap))
 		ap->eof = false;
 
@@ -3621,6 +3643,32 @@ xfs_bmap_btalloc_best_length(
 	return xfs_bmap_btalloc_low_space(ap, args);
 }
 
+/* Start a write-stream file in the AG set that backs its stream. */
+static int
+xfs_bmap_btalloc_write_stream(
+	struct xfs_bmalloca	*ap,
+	struct xfs_alloc_arg	*args,
+	int			stripe_align,
+	unsigned int		stream_id)
+{
+	struct xfs_mount	*mp = ap->ip->i_mount;
+	xfs_agnumber_t		agno;
+
+	agno = xfs_bmap_write_stream_agno(ap->ip, stream_id);
+	ap->blkno = XFS_AGB_TO_FSB(mp, agno, 0);
+	return xfs_bmap_btalloc_from_blkno(ap, args, stripe_align);
+}
+
+static int
+xfs_bmap_btalloc_best_length(
+	struct xfs_bmalloca	*ap,
+	struct xfs_alloc_arg	*args,
+	int			stripe_align)
+{
+	ap->blkno = XFS_INODE_TO_FSB(ap->ip);
+	return xfs_bmap_btalloc_from_blkno(ap, args, stripe_align);
+}
+
 static int
 xfs_bmap_btalloc(
 	struct xfs_bmalloca	*ap)
@@ -3640,6 +3688,7 @@ xfs_bmap_btalloc(
 	};
 	xfs_fileoff_t		orig_offset;
 	xfs_extlen_t		orig_length;
+	unsigned int		stream_id;
 	int			error;
 	int			stripe_align;
 
@@ -3652,11 +3701,16 @@ xfs_bmap_btalloc(
 	/* Trim the allocation back to the maximum an AG can fit. */
 	args.maxlen = min(ap->length, mp->m_ag_max_usable);
 
+	stream_id = READ_ONCE(VFS_I(ap->ip)->i_write_stream);
+
 	if (unlikely(XFS_TEST_ERROR(mp, XFS_ERRTAG_BMAP_ALLOC_MINLEN_EXTENT)))
 		error = xfs_bmap_exact_minlen_extent_alloc(ap, &args);
 	else if ((ap->datatype & XFS_ALLOC_USERDATA) &&
 			xfs_inode_is_filestream(ap->ip))
 		error = xfs_bmap_btalloc_filestreams(ap, &args, stripe_align);
+	else if ((ap->datatype & XFS_ALLOC_USERDATA) && stream_id)
+		error = xfs_bmap_btalloc_write_stream(ap, &args, stripe_align,
+				stream_id);
 	else
 		error = xfs_bmap_btalloc_best_length(ap, &args, stripe_align);
 	if (error)
-- 
2.25.1


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 6/8] xfs: support write streams on realtime volumes
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
                     ` (4 preceding siblings ...)
  2026-09-21  9:31   ` [PATCH v5 5/8] xfs: write stream based AG placement Kanchan Joshi
@ 2026-09-21  9:31   ` Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 7/8] iomap: introduce and propagate write_stream Kanchan Joshi
                     ` (3 subsequent siblings)
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Anuj Gupta, Kanchan Joshi

From: Anuj Gupta <anuj20.g@samsung.com>

Enable write streams on non-zoned realtime volumes that have realtime
groups, partitioning RTGs among the streams as the AG side does.

Stream based RTG selection applies to the first allocation of a file;
later ones keep using the adjacent-allocation hint so the file stays
contiguous.

Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
---
 fs/xfs/xfs_inode.c   |  4 ++--
 fs/xfs/xfs_ioctl.c   |  3 ++-
 fs/xfs/xfs_rtalloc.c | 27 +++++++++++++++++++++++++++
 fs/xfs/xfs_super.c   |  7 +++++++
 4 files changed, 38 insertions(+), 3 deletions(-)

diff --git a/fs/xfs/xfs_inode.c b/fs/xfs/xfs_inode.c
index df0a37c4d68a..c3503fe870ba 100644
--- a/fs/xfs/xfs_inode.c
+++ b/fs/xfs/xfs_inode.c
@@ -57,7 +57,7 @@ xfs_inode_max_write_streams(
 
 	if (xfs_inode_is_filestream(ip))
 		return 0;
-	if (XFS_IS_REALTIME_INODE(ip))
+	if (XFS_IS_REALTIME_INODE(ip) && xfs_has_zoned(ip->i_mount))
 		return 0;
 	return write_stream_pool_count(&xfs_inode_buftarg(ip)->bt_stream_pool);
 }
@@ -75,7 +75,7 @@ xfs_inode_set_write_stream(
 	if (!fd_file(f))
 		return -EBADF;
 	xfs_ilock(ip, XFS_ILOCK_EXCL);
-	if (XFS_IS_REALTIME_INODE(ip)) {
+	if (XFS_IS_REALTIME_INODE(ip) && xfs_has_zoned(ip->i_mount)) {
 		error = -EINVAL;
 		goto out_unlock;
 	}
diff --git a/fs/xfs/xfs_ioctl.c b/fs/xfs/xfs_ioctl.c
index bc6957ad06cb..4c0aff34e3c3 100644
--- a/fs/xfs/xfs_ioctl.c
+++ b/fs/xfs/xfs_ioctl.c
@@ -559,7 +559,8 @@ xfs_ioctl_setattr_xflags(
 
 	/* refuse a filestream/realtime flag change while a stream is attached */
 	if (READ_ONCE(VFS_I(ip)->i_write_stream) &&
-	    ((fa->fsx_xflags & FS_XFLAG_FILESTREAM) || rtflag))
+	    ((fa->fsx_xflags & FS_XFLAG_FILESTREAM) ||
+	     rtflag != XFS_IS_REALTIME_INODE(ip)))
 		return -EINVAL;
 
 	if (rtflag != XFS_IS_REALTIME_INODE(ip)) {
diff --git a/fs/xfs/xfs_rtalloc.c b/fs/xfs/xfs_rtalloc.c
index 84efe5a8fb11..a291f1aca6eb 100644
--- a/fs/xfs/xfs_rtalloc.c
+++ b/fs/xfs/xfs_rtalloc.c
@@ -2160,6 +2160,28 @@ xfs_rtallocate_align(
 	return 0;
 }
 
+static xfs_rtblock_t
+xfs_bmap_write_stream_rtbno(
+	struct xfs_inode	*ip,
+	unsigned int		stream_id)
+{
+	struct xfs_mount	*mp = ip->i_mount;
+	unsigned int		nr_rtgs = mp->m_sb.sb_rgcount;
+	unsigned int		nr_streams =
+		write_stream_pool_count(&mp->m_rtdev_targp->bt_stream_pool);
+	unsigned int		set_size = nr_rtgs / nr_streams;
+	xfs_rgnumber_t		start_rgno, target_rgno;
+
+	stream_id -= 1;	/* convert from 1-based to 0-based */
+	start_rgno = stream_id * set_size;
+	if (stream_id == nr_streams - 1)
+		set_size = nr_rtgs - start_rgno;
+	target_rgno = start_rgno + I_INO(ip) % set_size;
+
+	return (xfs_rtblock_t)target_rgno <<
+		mp->m_groups[XG_TYPE_RTG].blklog;
+}
+
 int
 xfs_bmap_rtalloc(
 	struct xfs_bmalloca	*ap)
@@ -2174,10 +2196,13 @@ xfs_bmap_rtalloc(
 	bool			noalign = false;
 	bool			initial_user_data =
 		ap->datatype & XFS_ALLOC_INITIAL_USER_DATA;
+	unsigned int		stream_id;
 	int			error;
 
 	ASSERT(!xfs_has_zoned(ap->tp->t_mountp));
 
+	stream_id = READ_ONCE(VFS_I(ap->ip)->i_write_stream);
+
 retry:
 	error = xfs_rtallocate_align(ap, &ralen, &raminlen, &prod, &noalign);
 	if (error)
@@ -2185,6 +2210,8 @@ xfs_bmap_rtalloc(
 
 	if (xfs_bmap_adjacent(ap))
 		bno_hint = ap->blkno;
+	else if (stream_id && xfs_has_rtgroups(ap->ip->i_mount))
+		bno_hint = xfs_bmap_write_stream_rtbno(ap->ip, stream_id);
 
 	if (xfs_has_rtgroups(ap->ip->i_mount)) {
 		error = xfs_rtallocate_rtgs(ap->tp, bno_hint, raminlen, ralen,
diff --git a/fs/xfs/xfs_super.c b/fs/xfs/xfs_super.c
index 443b3d847150..5c92f8ebd8a7 100644
--- a/fs/xfs/xfs_super.c
+++ b/fs/xfs/xfs_super.c
@@ -637,6 +637,13 @@ xfs_setup_devices(
 		if (error)
 			return error;
 		xfs_update_bdi_rahead(mp);
+
+		if (xfs_has_rtgroups(mp) && !xfs_has_zoned(mp)) {
+			error = xfs_buftarg_init_streams(mp->m_rtdev_targp,
+					mp->m_sb.sb_rgcount);
+			if (error)
+				return error;
+		}
 	}
 
 	return 0;
-- 
2.25.1


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 7/8] iomap: introduce and propagate write_stream
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
                     ` (5 preceding siblings ...)
  2026-09-21  9:31   ` [PATCH v5 6/8] xfs: support write streams on realtime volumes Kanchan Joshi
@ 2026-09-21  9:31   ` Kanchan Joshi
  2026-09-21  9:31   ` [PATCH v5 8/8] xfs: support hardware write streams Kanchan Joshi
                     ` (2 subsequent siblings)
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi, Anuj Gupta

Add a new write_stream field to struct iomap. Propagate write_stream from
iomap to bio in both direct I/O and buffered writeback paths.

Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
---
 fs/iomap/direct-io.c  | 1 +
 fs/iomap/ioend.c      | 3 +++
 include/linux/iomap.h | 1 +
 3 files changed, 5 insertions(+)

diff --git a/fs/iomap/direct-io.c b/fs/iomap/direct-io.c
index 8b4039d16ce8..f406a7bc8804 100644
--- a/fs/iomap/direct-io.c
+++ b/fs/iomap/direct-io.c
@@ -349,6 +349,7 @@ static ssize_t iomap_dio_bio_iter_one(struct iomap_iter *iter,
 	fscrypt_set_bio_crypt_ctx(bio, iter->inode, pos, GFP_KERNEL);
 	bio->bi_iter.bi_sector = iomap_sector(&iter->iomap, pos);
 	bio->bi_write_hint = iter->inode->i_write_hint;
+	bio->bi_write_stream = iter->iomap.write_stream;
 	bio->bi_ioprio = dio->iocb->ki_ioprio;
 	bio->bi_private = dio;
 	bio->bi_end_io = iomap_dio_bio_end_io;
diff --git a/fs/iomap/ioend.c b/fs/iomap/ioend.c
index 7bbbb417f915..31a229c12b13 100644
--- a/fs/iomap/ioend.c
+++ b/fs/iomap/ioend.c
@@ -166,6 +166,7 @@ static struct iomap_ioend *iomap_alloc_ioend(struct iomap_writepage_ctx *wpc,
 			       GFP_NOFS, &iomap_ioend_bioset);
 	bio->bi_iter.bi_sector = iomap_sector(&wpc->iomap, pos);
 	bio->bi_write_hint = wpc->inode->i_write_hint;
+	bio->bi_write_stream = wpc->iomap.write_stream;
 	wbc_init_bio(wpc->wbc, bio);
 	wpc->nr_folios = 0;
 	return iomap_init_ioend(wpc->inode, bio, pos, ioend_flags);
@@ -189,6 +190,8 @@ static bool iomap_can_add_to_ioend(struct iomap_writepage_ctx *wpc, loff_t pos,
 	if (!(wpc->iomap.flags & IOMAP_F_ANON_WRITE) &&
 	    iomap_sector(&wpc->iomap, pos) != bio_end_sector(&ioend->io_bio))
 		return false;
+	if (wpc->iomap.write_stream != ioend->io_bio.bi_write_stream)
+		return false;
 	/*
 	 * Limit ioend bio chain lengths to minimise IO completion latency. This
 	 * also prevents long tight loops ending page writeback on all the
diff --git a/include/linux/iomap.h b/include/linux/iomap.h
index bc7ae6327dbf..f923778f08bd 100644
--- a/include/linux/iomap.h
+++ b/include/linux/iomap.h
@@ -133,6 +133,7 @@ struct iomap {
 	u64			length;	/* length of mapping, bytes */
 	u16			type;	/* type of mapping */
 	u16			flags;	/* flags for mapping */
+	u16			write_stream; /* write stream for I/O */
 	struct block_device	*bdev;	/* block device for I/O */
 	struct dax_device	*dax_dev; /* dax_dev for dax operations */
 	void			*inline_data;
-- 
2.25.1


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 8/8] xfs: support hardware write streams
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
                     ` (6 preceding siblings ...)
  2026-09-21  9:31   ` [PATCH v5 7/8] iomap: introduce and propagate write_stream Kanchan Joshi
@ 2026-09-21  9:31   ` Kanchan Joshi
  2026-09-21  9:43   ` [PATCH v5 0/8] xfs " Kanchan Joshi
  2026-09-21 23:23   ` Dave Chinner
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:31 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev, Kanchan Joshi, Anuj Gupta

Use the stream count reported by the data and realtime block devices
where they have one, falling back to the geometry-derived count. The
count is capped at the group count so that every stream backs onto at
least one group; a device with more streams than the filesystem has
groups loses the excess, so mkfs should create at least as many groups
as the device has streams.

Set iomap->write_stream from the inode so that the id reaches the
device.

Signed-off-by: Kanchan Joshi <joshi.k@samsung.com>
Signed-off-by: Anuj Gupta <anuj20.g@samsung.com>
---
 fs/xfs/xfs_buf.c   | 7 ++++++-
 fs/xfs/xfs_iomap.c | 1 +
 2 files changed, 7 insertions(+), 1 deletion(-)

diff --git a/fs/xfs/xfs_buf.c b/fs/xfs/xfs_buf.c
index 4e84528b05cb..1748be9aa7d6 100644
--- a/fs/xfs/xfs_buf.c
+++ b/fs/xfs/xfs_buf.c
@@ -1709,8 +1709,13 @@ xfs_buftarg_init_streams(
 	unsigned int		nr_groups)
 {
 	unsigned int		nr_streams;
+	unsigned int		hw_streams;
 
-	nr_streams = xfs_sw_write_stream_count(nr_groups);
+	hw_streams = bdev_max_write_streams(btp->bt_bdev);
+	if (hw_streams)
+		nr_streams = min(hw_streams, nr_groups);
+	else
+		nr_streams = xfs_sw_write_stream_count(nr_groups);
 	return write_stream_pool_init(&btp->bt_stream_pool, nr_streams);
 }
 
diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
index 7c6238fed61e..3e9df790a544 100644
--- a/fs/xfs/xfs_iomap.c
+++ b/fs/xfs/xfs_iomap.c
@@ -144,6 +144,7 @@ xfs_bmbt_to_iomap(
 	}
 	iomap->offset = XFS_FSB_TO_B(mp, imap->br_startoff);
 	iomap->length = XFS_FSB_TO_B(mp, imap->br_blockcount);
+	iomap->write_stream = READ_ONCE(VFS_I(ip)->i_write_stream);
 	if (mapping_flags & IOMAP_DAX) {
 		iomap->dax_dev = target->bt_daxdev;
 	} else {
-- 
2.25.1


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* Re: [PATCH v5 0/8] xfs write streams
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
                     ` (7 preceding siblings ...)
  2026-09-21  9:31   ` [PATCH v5 8/8] xfs: support hardware write streams Kanchan Joshi
@ 2026-09-21  9:43   ` Kanchan Joshi
  2026-09-21 23:23   ` Dave Chinner
  9 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-21  9:43 UTC (permalink / raw)
  To: brauner, hch, djwong, dgc, cem, jack, axboe, kbusch
  Cc: linux-xfs, linux-fsdevel, gost.dev

Patches are on top of 7.3-rc3.
Branch: https://github.com/SamsungDS/linux/tree/feat/xfs_fdp_v5

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v5 0/8] xfs write streams
  2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
                     ` (8 preceding siblings ...)
  2026-09-21  9:43   ` [PATCH v5 0/8] xfs " Kanchan Joshi
@ 2026-09-21 23:23   ` Dave Chinner
  2026-09-25 15:13     ` Kanchan Joshi
  9 siblings, 1 reply; 12+ messages in thread
From: Dave Chinner @ 2026-09-21 23:23 UTC (permalink / raw)
  To: Kanchan Joshi
  Cc: brauner, hch, djwong, cem, jack, axboe, kbusch, linux-xfs,
	linux-fsdevel, gost.dev

On Mon, Sep 21, 2026 at 03:01:39PM +0530, Kanchan Joshi wrote:
> This series introduces a generic interface [2,3] for write stream management on
> files.
> It enables spatial isolation (at sw and hw level) and concurrency improvments [1]
> in xfs by
> (a) steering each write stream to its own set of allocation groups (patch #5)
>     and realtime groups (patch #6).
> (b) connecting xfs write streams to block write streams (FDP capable NVMe).
> 
> Write streams allow the abstraction provider (fs, block, raid etc.) to
> leverage application's intent (file relationships/lifecycle).
> - application: reserves a stream and sets it on the files it wants
>   placed together.
> - xfs: maps streams to AGs/RGs; allocates without interleaving; gains
>   higher concurrency due to reduced lock contention.
> - hardware: maps streams to underlying allocation unit; reduces device
>   internal write amplification, improved life, predictable QoS.
> 
> Also
> - A stream is handed out using an fd. The filesystem is free to map streams
>   onto its own geometry, and the reservation/fd machinery is generic
>  (fs/write_streams.c) so other filesystems can reuse it.
> 
> - Since high-level write stream (in xfs) and logical placement can work
>   without the low-level write streams (in block device), series has a general
>   value beyond the hardware that provides spatial isolation. Patches 1-7 are
>   software only and work on any block device; patch 8 aligns the stream
>   count to the hardware streams when the device has them.

So, how would I set up a stream that directs all writes to AG 1, and
returns ENOSPC to write operations if that AG is full?

What about having several streams, each pointing at a different,
known AG that the application directly controls (i.e. a known 1:1
mapping between stream_fd and agno)?

This is functionality that we could use in xfs_fsr to get rid of all
the historic tmpdir/tmpfile heuristics that allow it to "control"
locality of the data placement of files that it defragments. We also
need such control of data placement to empty AGs for shrink
operations. Yes, this only occurred to me a couple of days ago, but
now that I've made the connection between write streams and fs
allocation policy direction, it seems like a natural fit.

FWIW, that also means that, for XFS, write streams need to be
applicable to directories, so that directory block allocation also
gets placed according to the stream ID, and that new inodes in that
directory are created in the same AG as the stream ID points to.
This directly allows us to rebuild directories using FICLONE/UNSHARE
tricks and have all the indoes, data and metadata placed in the AG
we desire (i.e. necessary functionality for online shrink).

So:

> [2]
> ### Application interface
> 
> Three new ioctls:
>       FS_IOC_WRITE_STREAM_GET_MAX  number of streams the filesystem offers
>       FS_IOC_WRITE_STREAM_ALLOC    reserve a stream, returns an fd
>       FS_IOC_WRITE_STREAM_SET      set or clear a stream on a file

How does this interface enable such usage of write streams to direct
purely filesystem level allocation policy requirements? I don't see
how I can use this API to direct where in the filesystem to map the
stream to. I can see that there is a 'set' command, but all it has
is a flags field. There's nothing passed to the ALLOC command to
allow the application to indicate to the filesystem where it wants
the new stream to point to.

IOWs, it appears taht there is no way to provide the FS with any
sort of direction as to how the stream should be set up. If we want
a FS specific stream (e.g. an AG) then we don't want it mapped to a
hardware stream, and we don't want it mapped to some random set of
AGs (like this patchset implements). The only real reason for
filesystem level write streams is to expose allocation locality
control to applications, so it seems kinda silly to have an API that
prevents any real application level control...

Indeed, how do we ask for a fs-level write stream instead of a
hardware-level write stream? They are different things, and have
different use cases (obviously!) so there definitely needs to be
some level of though put into this. And, FWIW, the write stream id
will probably need to be a u32 if we are going to support FS level
write streams, as we can have more than 65536 AGs in a
filesystem....

> - Usage model: application needs to get a handle (fd) for a write stream
>   before being able to use it. This avoids multi-application conflicts.

Why can't multiple applications use the same write stream locality
mapping independently? If the application wants an exclusive write
stream (i.e. exclusive access to a set of AGs in the filesystem),
then surely that's a flag for the ALLOC API, right?

Cheers,

Dave.
-- 
Dave Chinner
dgc@kernel.org

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v5 0/8] xfs write streams
  2026-09-21 23:23   ` Dave Chinner
@ 2026-09-25 15:13     ` Kanchan Joshi
  0 siblings, 0 replies; 12+ messages in thread
From: Kanchan Joshi @ 2026-09-25 15:13 UTC (permalink / raw)
  To: Dave Chinner
  Cc: brauner, hch, djwong, cem, jack, axboe, kbusch, linux-xfs,
	linux-fsdevel, gost.dev

On 9/22/2026 4:53 AM, Dave Chinner wrote:
> On Mon, Sep 21, 2026 at 03:01:39PM +0530, Kanchan Joshi wrote:
>> This series introduces a generic interface [2,3] for write stream management on
>> files.
>> It enables spatial isolation (at sw and hw level) and concurrency improvments [1]
>> in xfs by
>> (a) steering each write stream to its own set of allocation groups (patch #5)
>>      and realtime groups (patch #6).
>> (b) connecting xfs write streams to block write streams (FDP capable NVMe).
>>
>> Write streams allow the abstraction provider (fs, block, raid etc.) to
>> leverage application's intent (file relationships/lifecycle).
>> - application: reserves a stream and sets it on the files it wants
>>    placed together.
>> - xfs: maps streams to AGs/RGs; allocates without interleaving; gains
>>    higher concurrency due to reduced lock contention.
>> - hardware: maps streams to underlying allocation unit; reduces device
>>    internal write amplification, improved life, predictable QoS.
>>
>> Also
>> - A stream is handed out using an fd. The filesystem is free to map streams
>>    onto its own geometry, and the reservation/fd machinery is generic
>>   (fs/write_streams.c) so other filesystems can reuse it.
>>
>> - Since high-level write stream (in xfs) and logical placement can work
>>    without the low-level write streams (in block device), series has a general
>>    value beyond the hardware that provides spatial isolation. Patches 1-7 are
>>    software only and work on any block device; patch 8 aligns the stream
>>    count to the hardware streams when the device has them.
> 
> So, how would I set up a stream that directs all writes to AG 1, and
> returns ENOSPC to write operations if that AG is full?

You can't, and that's a deliberate and necessary design choice for the 
usecase (and I mentioned soft boundary in patch #5). While AG/RG are 
fixed-sized XFS buckets, stream is a high-level (vfs) bucket with its 
opaque handle (fd) that does not come with fixed size constraint. 
Applications may tag many files (more data) with one stream and tag few 
(less data) with another stream.


> What about having several streams, each pointing at a different,
> known AG that the application directly controls (i.e. a known 1:1
> mapping between stream_fd and agno)?

We can't use agno here.
Stream lets the application state intent: these files belong together, 
those stay apart. It does not dictate layout. AG/RG is xfs-only 
vocabulary, and that's fine for xfs-only interface. But we are doing a 
generic interface that needs to remain portable across filesystems.


> This is functionality that we could use in xfs_fsr to get rid of all
> the historic tmpdir/tmpfile heuristics that allow it to "control"
> locality of the data placement of files that it defragments. We also
> need such control of data placement to empty AGs for shrink
> operations. Yes, this only occurred to me a couple of days ago, but
> now that I've made the connection between write streams and fs
> allocation policy direction, it seems like a natural fit.
> 
> FWIW, that also means that, for XFS, write streams need to be
> applicable to directories, so that directory block allocation also
> gets placed according to the stream ID, and that new inodes in that
> directory are created in the same AG as the stream ID points to.
> This directly allows us to rebuild directories using FICLONE/UNSHARE
> tricks and have all the indoes, data and metadata placed in the AG
> we desire (i.e. necessary functionality for online shrink).


With this, we are discussing xfs tools (and not fs-agnostic 
applications) that speak nitty-gritty of xfs layout. This usecase does 
not require stream and stream-fd based workflow of this series; it can 
be served more cleanly with XFS only ioctl. Something like:

  XFS_IOC_SET_GROUP(AG/RG no, flags) on the file

To set the 'hard' allocation-directive so that all allocations happens 
from that AG/RG no and ENOSPC if not.

Also it seems shrink usecase will also require that target AG remains 
exclusive (not available for allocations) while it is getting emptied. 
That, exclusive capacity locking, also we don't do with write-stream. We 
can't return ENOSPC to other allocators just because somebody decided to 
use/abuse stream to create artificial lack of free space.

> So:
> 
>> [2]
>> ### Application interface
>>
>> Three new ioctls:
>>        FS_IOC_WRITE_STREAM_GET_MAX  number of streams the filesystem offers
>>        FS_IOC_WRITE_STREAM_ALLOC    reserve a stream, returns an fd
>>        FS_IOC_WRITE_STREAM_SET      set or clear a stream on a file
> 
> How does this interface enable such usage of write streams to direct
> purely filesystem level allocation policy requirements?

As mentioned above, you might agree that interface should remain 
portable across filesystems.

> I don't see
> how I can use this API to direct where in the filesystem to map the
> stream to. I can see that there is a 'set' command, but all it has
> is a flags field. There's nothing passed to the ALLOC command to
> allow the application to indicate to the filesystem where it wants
> the new stream to point to.

FWIW, we had a 'stream-id' as part of ALLOC in the previous version.
Christoph suggested to remove that from UAPI.
https://lore.kernel.org/linux-xfs/20260825065533.GA24808@lst.de/

The direction has been to be more opaque and less explicit.

> IOWs, it appears taht there is no way to provide the FS with any
> sort of direction as to how the stream should be set up. If we want
> a FS specific stream (e.g. an AG) then we don't want it mapped to a
> hardware stream, and we don't want it mapped to some random set of
> AGs (like this patchset implements). The only real reason for
> filesystem level write streams is to expose allocation locality
> control to applications, so it seems kinda silly to have an API that
> prevents any real application level control...
> Indeed, how do we ask for a fs-level write stream instead of a
> hardware-level write stream? They are different things, and have
> different use cases (obviously!) so there definitely needs to be
> some level of though put into this. And, FWIW, the write stream id
> will probably need to be a u32 if we are going to support FS level
> write streams, as we can have more than 65536 AGs in a
> filesystem....

All this gets handled cleanly with the above xfs-ioctl based approach.
In that, one can put the actual AG/RG no into xfs-inode itself so that 
future allocation can happen from that.


> 
>> - Usage model: application needs to get a handle (fd) for a write stream
>>    before being able to use it. This avoids multi-application conflicts.
> 
> Why can't multiple applications use the same write stream locality
> mapping independently? If the application wants an exclusive write
> stream (i.e. exclusive access to a set of AGs in the filesystem),
> then surely that's a flag for the ALLOC API, right?

Fair; But I should clarity that exclusive access is only for stream 
resource/handle (fd) so that another unrelated application does not pick 
the same stream and start colliding. This matters more for hardware 
stream. But exclusivity is not about AGs, we don't lock AGs on 
per-stream basis.





^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-09-25 15:21 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <CGME20260921093229epcas5p387ee10f88335ddc5fc930ca919769a60@epcas5p3.samsung.com>
2026-09-21  9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
2026-09-21  9:31   ` [PATCH v5 1/8] fs: add write-stream management ioctls Kanchan Joshi
2026-09-21  9:31   ` [PATCH v5 2/8] fs: add generic write-stream management Kanchan Joshi
2026-09-21  9:31   ` [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint Kanchan Joshi
2026-09-21  9:31   ` [PATCH v5 4/8] xfs: implement software write-stream management Kanchan Joshi
2026-09-21  9:31   ` [PATCH v5 5/8] xfs: write stream based AG placement Kanchan Joshi
2026-09-21  9:31   ` [PATCH v5 6/8] xfs: support write streams on realtime volumes Kanchan Joshi
2026-09-21  9:31   ` [PATCH v5 7/8] iomap: introduce and propagate write_stream Kanchan Joshi
2026-09-21  9:31   ` [PATCH v5 8/8] xfs: support hardware write streams Kanchan Joshi
2026-09-21  9:43   ` [PATCH v5 0/8] xfs " Kanchan Joshi
2026-09-21 23:23   ` Dave Chinner
2026-09-25 15:13     ` Kanchan Joshi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox