From: Kanchan Joshi <joshi.k@samsung.com>
To: Dave Chinner <dgc@kernel.org>
Cc: brauner@kernel.org, hch@lst.de, djwong@kernel.org,
cem@kernel.org, jack@suse.cz, axboe@kernel.dk, kbusch@kernel.org,
linux-xfs@vger.kernel.org, linux-fsdevel@vger.kernel.org,
gost.dev@samsung.com
Subject: Re: [PATCH v5 0/8] xfs write streams
Date: Fri, 25 Sep 2026 20:43:56 +0530 [thread overview]
Message-ID: <ef9844a0-3a9f-48e4-8092-2de2e03db1c8@samsung.com> (raw)
In-Reply-To: <arG8a7Vcy6_qZi0e@dread>
On 9/22/2026 4:53 AM, Dave Chinner wrote:
> On Mon, Sep 21, 2026 at 03:01:39PM +0530, Kanchan Joshi wrote:
>> This series introduces a generic interface [2,3] for write stream management on
>> files.
>> It enables spatial isolation (at sw and hw level) and concurrency improvments [1]
>> in xfs by
>> (a) steering each write stream to its own set of allocation groups (patch #5)
>> and realtime groups (patch #6).
>> (b) connecting xfs write streams to block write streams (FDP capable NVMe).
>>
>> Write streams allow the abstraction provider (fs, block, raid etc.) to
>> leverage application's intent (file relationships/lifecycle).
>> - application: reserves a stream and sets it on the files it wants
>> placed together.
>> - xfs: maps streams to AGs/RGs; allocates without interleaving; gains
>> higher concurrency due to reduced lock contention.
>> - hardware: maps streams to underlying allocation unit; reduces device
>> internal write amplification, improved life, predictable QoS.
>>
>> Also
>> - A stream is handed out using an fd. The filesystem is free to map streams
>> onto its own geometry, and the reservation/fd machinery is generic
>> (fs/write_streams.c) so other filesystems can reuse it.
>>
>> - Since high-level write stream (in xfs) and logical placement can work
>> without the low-level write streams (in block device), series has a general
>> value beyond the hardware that provides spatial isolation. Patches 1-7 are
>> software only and work on any block device; patch 8 aligns the stream
>> count to the hardware streams when the device has them.
>
> So, how would I set up a stream that directs all writes to AG 1, and
> returns ENOSPC to write operations if that AG is full?
You can't, and that's a deliberate and necessary design choice for the
usecase (and I mentioned soft boundary in patch #5). While AG/RG are
fixed-sized XFS buckets, stream is a high-level (vfs) bucket with its
opaque handle (fd) that does not come with fixed size constraint.
Applications may tag many files (more data) with one stream and tag few
(less data) with another stream.
> What about having several streams, each pointing at a different,
> known AG that the application directly controls (i.e. a known 1:1
> mapping between stream_fd and agno)?
We can't use agno here.
Stream lets the application state intent: these files belong together,
those stay apart. It does not dictate layout. AG/RG is xfs-only
vocabulary, and that's fine for xfs-only interface. But we are doing a
generic interface that needs to remain portable across filesystems.
> This is functionality that we could use in xfs_fsr to get rid of all
> the historic tmpdir/tmpfile heuristics that allow it to "control"
> locality of the data placement of files that it defragments. We also
> need such control of data placement to empty AGs for shrink
> operations. Yes, this only occurred to me a couple of days ago, but
> now that I've made the connection between write streams and fs
> allocation policy direction, it seems like a natural fit.
>
> FWIW, that also means that, for XFS, write streams need to be
> applicable to directories, so that directory block allocation also
> gets placed according to the stream ID, and that new inodes in that
> directory are created in the same AG as the stream ID points to.
> This directly allows us to rebuild directories using FICLONE/UNSHARE
> tricks and have all the indoes, data and metadata placed in the AG
> we desire (i.e. necessary functionality for online shrink).
With this, we are discussing xfs tools (and not fs-agnostic
applications) that speak nitty-gritty of xfs layout. This usecase does
not require stream and stream-fd based workflow of this series; it can
be served more cleanly with XFS only ioctl. Something like:
XFS_IOC_SET_GROUP(AG/RG no, flags) on the file
To set the 'hard' allocation-directive so that all allocations happens
from that AG/RG no and ENOSPC if not.
Also it seems shrink usecase will also require that target AG remains
exclusive (not available for allocations) while it is getting emptied.
That, exclusive capacity locking, also we don't do with write-stream. We
can't return ENOSPC to other allocators just because somebody decided to
use/abuse stream to create artificial lack of free space.
> So:
>
>> [2]
>> ### Application interface
>>
>> Three new ioctls:
>> FS_IOC_WRITE_STREAM_GET_MAX number of streams the filesystem offers
>> FS_IOC_WRITE_STREAM_ALLOC reserve a stream, returns an fd
>> FS_IOC_WRITE_STREAM_SET set or clear a stream on a file
>
> How does this interface enable such usage of write streams to direct
> purely filesystem level allocation policy requirements?
As mentioned above, you might agree that interface should remain
portable across filesystems.
> I don't see
> how I can use this API to direct where in the filesystem to map the
> stream to. I can see that there is a 'set' command, but all it has
> is a flags field. There's nothing passed to the ALLOC command to
> allow the application to indicate to the filesystem where it wants
> the new stream to point to.
FWIW, we had a 'stream-id' as part of ALLOC in the previous version.
Christoph suggested to remove that from UAPI.
https://lore.kernel.org/linux-xfs/20260825065533.GA24808@lst.de/
The direction has been to be more opaque and less explicit.
> IOWs, it appears taht there is no way to provide the FS with any
> sort of direction as to how the stream should be set up. If we want
> a FS specific stream (e.g. an AG) then we don't want it mapped to a
> hardware stream, and we don't want it mapped to some random set of
> AGs (like this patchset implements). The only real reason for
> filesystem level write streams is to expose allocation locality
> control to applications, so it seems kinda silly to have an API that
> prevents any real application level control...
> Indeed, how do we ask for a fs-level write stream instead of a
> hardware-level write stream? They are different things, and have
> different use cases (obviously!) so there definitely needs to be
> some level of though put into this. And, FWIW, the write stream id
> will probably need to be a u32 if we are going to support FS level
> write streams, as we can have more than 65536 AGs in a
> filesystem....
All this gets handled cleanly with the above xfs-ioctl based approach.
In that, one can put the actual AG/RG no into xfs-inode itself so that
future allocation can happen from that.
>
>> - Usage model: application needs to get a handle (fd) for a write stream
>> before being able to use it. This avoids multi-application conflicts.
>
> Why can't multiple applications use the same write stream locality
> mapping independently? If the application wants an exclusive write
> stream (i.e. exclusive access to a set of AGs in the filesystem),
> then surely that's a flag for the ALLOC API, right?
Fair; But I should clarity that exclusive access is only for stream
resource/handle (fd) so that another unrelated application does not pick
the same stream and start colliding. This matters more for hardware
stream. But exclusivity is not about AGs, we don't lock AGs on
per-stream basis.
prev parent reply other threads:[~2026-09-25 15:21 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <CGME20260921093229epcas5p387ee10f88335ddc5fc930ca919769a60@epcas5p3.samsung.com>
2026-09-21 9:31 ` [PATCH v5 0/8] xfs write streams Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 1/8] fs: add write-stream management ioctls Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 2/8] fs: add generic write-stream management Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 3/8] fs: add i_write_stream, exclusive with the write life time hint Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 4/8] xfs: implement software write-stream management Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 5/8] xfs: write stream based AG placement Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 6/8] xfs: support write streams on realtime volumes Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 7/8] iomap: introduce and propagate write_stream Kanchan Joshi
2026-09-21 9:31 ` [PATCH v5 8/8] xfs: support hardware write streams Kanchan Joshi
2026-09-21 9:43 ` [PATCH v5 0/8] xfs " Kanchan Joshi
2026-09-21 23:23 ` Dave Chinner
2026-09-25 15:13 ` Kanchan Joshi [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=ef9844a0-3a9f-48e4-8092-2de2e03db1c8@samsung.com \
--to=joshi.k@samsung.com \
--cc=axboe@kernel.dk \
--cc=brauner@kernel.org \
--cc=cem@kernel.org \
--cc=dgc@kernel.org \
--cc=djwong@kernel.org \
--cc=gost.dev@samsung.com \
--cc=hch@lst.de \
--cc=jack@suse.cz \
--cc=kbusch@kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-xfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox