From: Dave Chinner <dgc@kernel.org>
To: Kanchan Joshi <joshi.k@samsung.com>
Cc: brauner@kernel.org, hch@lst.de, djwong@kernel.org,
cem@kernel.org, jack@suse.cz, axboe@kernel.dk, kbusch@kernel.org,
linux-xfs@vger.kernel.org, linux-fsdevel@vger.kernel.org,
gost.dev@samsung.com
Subject: Re: [PATCH v5 0/8] xfs write streams
Date: Tue, 22 Sep 2026 09:23:23 +1000 [thread overview]
Message-ID: <arG8a7Vcy6_qZi0e@dread> (raw)
In-Reply-To: <20260921093147.59935-1-joshi.k@samsung.com>
On Mon, Sep 21, 2026 at 03:01:39PM +0530, Kanchan Joshi wrote:
> This series introduces a generic interface [2,3] for write stream management on
> files.
> It enables spatial isolation (at sw and hw level) and concurrency improvments [1]
> in xfs by
> (a) steering each write stream to its own set of allocation groups (patch #5)
> and realtime groups (patch #6).
> (b) connecting xfs write streams to block write streams (FDP capable NVMe).
>
> Write streams allow the abstraction provider (fs, block, raid etc.) to
> leverage application's intent (file relationships/lifecycle).
> - application: reserves a stream and sets it on the files it wants
> placed together.
> - xfs: maps streams to AGs/RGs; allocates without interleaving; gains
> higher concurrency due to reduced lock contention.
> - hardware: maps streams to underlying allocation unit; reduces device
> internal write amplification, improved life, predictable QoS.
>
> Also
> - A stream is handed out using an fd. The filesystem is free to map streams
> onto its own geometry, and the reservation/fd machinery is generic
> (fs/write_streams.c) so other filesystems can reuse it.
>
> - Since high-level write stream (in xfs) and logical placement can work
> without the low-level write streams (in block device), series has a general
> value beyond the hardware that provides spatial isolation. Patches 1-7 are
> software only and work on any block device; patch 8 aligns the stream
> count to the hardware streams when the device has them.
So, how would I set up a stream that directs all writes to AG 1, and
returns ENOSPC to write operations if that AG is full?
What about having several streams, each pointing at a different,
known AG that the application directly controls (i.e. a known 1:1
mapping between stream_fd and agno)?
This is functionality that we could use in xfs_fsr to get rid of all
the historic tmpdir/tmpfile heuristics that allow it to "control"
locality of the data placement of files that it defragments. We also
need such control of data placement to empty AGs for shrink
operations. Yes, this only occurred to me a couple of days ago, but
now that I've made the connection between write streams and fs
allocation policy direction, it seems like a natural fit.
FWIW, that also means that, for XFS, write streams need to be
applicable to directories, so that directory block allocation also
gets placed according to the stream ID, and that new inodes in that
directory are created in the same AG as the stream ID points to.
This directly allows us to rebuild directories using FICLONE/UNSHARE
tricks and have all the indoes, data and metadata placed in the AG
we desire (i.e. necessary functionality for online shrink).
So:
> [2]
> ### Application interface
>
> Three new ioctls:
> FS_IOC_WRITE_STREAM_GET_MAX number of streams the filesystem offers
> FS_IOC_WRITE_STREAM_ALLOC reserve a stream, returns an fd
> FS_IOC_WRITE_STREAM_SET set or clear a stream on a file
How does this interface enable such usage of write streams to direct
purely filesystem level allocation policy requirements? I don't see
how I can use this API to direct where in the filesystem to map the
stream to. I can see that there is a 'set' command, but all it has
is a flags field. There's nothing passed to the ALLOC command to
allow the application to indicate to the filesystem where it wants
the new stream to point to.
IOWs, it appears taht there is no way to provide the FS with any
sort of direction as to how the stream should be set up. If we want
a FS specific stream (e.g. an AG) then we don't want it mapped to a
hardware stream, and we don't want it mapped to some random set of
AGs (like this patchset implements). The only real reason for
filesystem level write streams is to expose allocation locality
control to applications, so it seems kinda silly to have an API that
prevents any real application level control...
Indeed, how do we ask for a fs-level write stream instead of a
hardware-level write stream? They are different things, and have
different use cases (obviously!) so there definitely needs to be
some level of though put into this. And, FWIW, the write stream id
will probably need to be a u32 if we are going to support FS level
write streams, as we can have more than 65536 AGs in a
filesystem....
> - Usage model: application needs to get a handle (fd) for a write stream
> before being able to use it. This avoids multi-application conflicts.
Why can't multiple applications use the same write stream locality
mapping independently? If the application wants an exclusive write
stream (i.e. exclusive access to a set of AGs in the filesystem),
then surely that's a flag for the ALLOC API, right?
Cheers,
Dave.
--
Dave Chinner
dgc@kernel.org
next prev parent reply other threads:[~2026-09-21 23:23 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <CGME20260921093229epcas5p387ee10f88335ddc5fc930ca919769a60@epcas5p3.samsung.com>
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 1/8] fs: add write-stream management ioctls Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 2/8] fs: add generic write-stream management Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 4/8] xfs: implement software write-stream management Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 5/8] xfs: write stream based AG placement Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 6/8] xfs: support write streams on realtime volumes Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 7/8] iomap: introduce and propagate write_stream Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 8/8] xfs: support hardware write streams Kanchan Joshi
2026-09-21 9:43 ` [PATCH v5 0/8] xfs " Kanchan Joshi
2026-09-21 23:23 ` Dave Chinner [this message]
2026-09-25 15:13 ` Kanchan Joshi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arG8a7Vcy6_qZi0e@dread \
--to=dgc@kernel.org \
--cc=axboe@kernel.dk \
--cc=brauner@kernel.org \
--cc=cem@kernel.org \
--cc=djwong@kernel.org \
--cc=gost.dev@samsung.com \
--cc=hch@lst.de \
--cc=jack@suse.cz \
--cc=joshi.k@samsung.com \
--cc=kbusch@kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-xfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox