* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-25 10:36 [PATCH] xfs: don't limit software atomic writes by the group alignment Pankaj Raghav
@ 2026-09-25 11:44 ` John Garry
2026-09-29 7:17 ` Pankaj Raghav
2026-10-05 11:02 ` John Garry
2026-10-05 11:02 ` John Garry
2 siblings, 1 reply; 19+ messages in thread
From: John Garry @ 2026-09-25 11:44 UTC (permalink / raw)
To: Pankaj Raghav, cem, linux-xfs
Cc: John Garry, Darrick J . Wong, gost.dev, pankaj.raghav
On 9/25/26 11:36, Pankaj Raghav wrote:
> xfs_calc_group_awu_max() clamps the atomic write unit maximum to the
> greatest power-of-two factor of the group size when the device advertises
> atomic writes, so that allocations can be made naturally aligned for
> REQ_ATOMIC. But that value is the software limit: it is only used when
> reflink is enabled, and out of place writes through the COW fork have no
> alignment requirement.
We have mkfs atomic write options for selecting atomic write limits -
you have tried that, right?
>
> A filesystem on a > 4TB device that advertises 16k atomic writes then reports
> an atomic write unit maximum of a single fsblock and an optimal maximum of 0,
> and rejects any larger RWF_ATOMIC write. That is worse than the same filesystem
> on a device with no atomic write support at all.
>
> Compute the software limit from the group size alone, and apply the
> alignment constraint in xfs_get_atomic_write_max_opt() instead, which is
> what reports the size that can be offloaded to the hardware.
>
> Results on a 8TB device with 16k hardware atomic support with 4k
> blocksize:
>
> Before patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 4096
> stx_atomic_write_unit_max: 4096
> stx_atomic_write_unit_max_opt: 0
> stx_atomic_write_segments_max: 1
>
> After patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 4096
> stx_atomic_write_unit_max: 2097152
Here the AG size may not be a multiple of 2097152B, right? If so, when
we try to naturally align the allocation for a CoW write, even if the AG
blocks are naturally aligned (per AG), the disk blocks may not be
naturnally aligned - is this correct? For HW-based atomic writes, we
want naturally aligned disk blocks.
> stx_atomic_write_unit_max_opt: 4096
> stx_atomic_write_segments_max: 1
>
> Results on a 8TB device with 16k hardware atomic support with 16k
> blocksize:
>
> Before patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 16384
> stx_atomic_write_unit_max: 16384
> stx_atomic_write_unit_max_opt: 0
> stx_atomic_write_segments_max: 1
>
> After patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 16384
> stx_atomic_write_unit_max: 33554432
> stx_atomic_write_unit_max_opt: 16384
> stx_atomic_write_segments_max: 1
>
> Fixes: 0c438dcc3150 ("xfs: add xfs_calc_atomic_write_unit_max()")
> Assisted-by: LLM
> Signed-off-by: Pankaj Raghav <p.raghav@samsung.com>
> ---
> fs/xfs/xfs_iops.c | 27 ++++++++++++++++++++++-----
> fs/xfs/xfs_mount.c | 35 ++++++++++++++++++++++++-----------
> fs/xfs/xfs_mount.h | 2 ++
> 3 files changed, 48 insertions(+), 16 deletions(-)
>
> diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c
> index d1306e723899..8c8b14f94ede 100644
> --- a/fs/xfs/xfs_iops.c
> +++ b/fs/xfs/xfs_iops.c
> @@ -620,6 +620,12 @@ xfs_get_atomic_write_min(
> return 0;
> }
>
> +static inline enum xfs_group_type
> +xfs_inode_group_type(struct xfs_inode *ip)
> +{
> + return XFS_IS_REALTIME_INODE(ip) ? XG_TYPE_RTG : XG_TYPE_AG;
> +}
> +
> unsigned int
> xfs_get_atomic_write_max(
> struct xfs_inode *ip)
> @@ -642,19 +648,20 @@ xfs_get_atomic_write_max(
> * then advertise a maximum size of whatever we can complete through
> * that means. Hardware support is reported via max_opt, not here.
> */
> - if (XFS_IS_REALTIME_INODE(ip))
> - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_RTG].awu_max);
> - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_AG].awu_max);
> + return XFS_FSB_TO_B(mp, mp->m_groups[xfs_inode_group_type(ip)].awu_max);
> }
>
> unsigned int
> xfs_get_atomic_write_max_opt(
> struct xfs_inode *ip)
> {
> + struct xfs_mount *mp = ip->i_mount;
> unsigned int awu_max = xfs_get_atomic_write_max(ip);
> + xfs_extlen_t align_max_fsb;
> + unsigned int opt;
>
> /* if the max is 1x block, then just keep behaviour that opt is 0 */
> - if (awu_max <= ip->i_mount->m_sb.sb_blocksize)
> + if (awu_max <= mp->m_sb.sb_blocksize)
> return 0;
>
> /*
> @@ -663,7 +670,17 @@ xfs_get_atomic_write_max_opt(
> * less than our out of place write limit, but we don't want to exceed
> * the awu_max.
> */
> - return min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
> + opt = min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
> +
> + /*
> + * REQ_ATOMIC writes also have to be naturally aligned on disk, so we
> + * cannot promise more than the largest extent that the allocator is
> + * able to align within a group.
> + */
> + align_max_fsb = xfs_calc_group_awu_align_max(mp,
> + xfs_inode_group_type(ip));
> +
> + return min_t(xfs_fsize_t, opt, XFS_FSB_TO_B(mp, align_max_fsb));
> }
>
> static void
> diff --git a/fs/xfs/xfs_mount.c b/fs/xfs/xfs_mount.c
> index be90c7b03994..b7110b3d023b 100644
> --- a/fs/xfs/xfs_mount.c
> +++ b/fs/xfs/xfs_mount.c
> @@ -674,15 +674,13 @@ static inline xfs_extlen_t xfs_calc_atomic_write_max(struct xfs_mount *mp)
> }
>
> /*
> - * If the underlying device advertises atomic write support, limit the size of
> - * atomic writes to the greatest power-of-two factor of the group size so
> - * that every atomic write unit aligns with the start of every group. This is
> - * required so that the allocations for an atomic write will always be
> - * aligned compatibly with the alignment requirements of the storage.
> + * The largest out of place write is the number of blocks that user files can
> + * allocate from any group.
> *
> - * If the device doesn't advertise atomic writes, then there are no alignment
> - * restrictions and the largest out-of-place write we can do ourselves is the
> - * number of blocks that user files can allocate from any group.
> + * This is a software limit, so it deliberately does not take the alignment
> + * requirements of the storage into account. An atomic write that cannot be
> + * handed to the device as a single naturally aligned REQ_ATOMIC bio is
> + * completed through the COW fork instead, which has no such requirement.
> */
> static xfs_extlen_t
> xfs_calc_group_awu_max(
> @@ -690,15 +688,30 @@ xfs_calc_group_awu_max(
> enum xfs_group_type type)
> {
> struct xfs_groups *g = &mp->m_groups[type];
> - struct xfs_buftarg *btp = xfs_group_type_buftarg(mp, type);
>
> if (g->blocks == 0)
> return 0;
> - if (btp && btp->bt_awu_min > 0)
> - return max_pow_of_two_factor(g->blocks);
> return rounddown_pow_of_two(g->blocks);
> }
>
> +/*
> + * Compute the largest atomic write unit for which the allocator can guarantee
> + * naturally aligned extents in this group type.
> + *
> + * Hardware atomic writes have to be naturally aligned on disk.
> + */
> +xfs_extlen_t
> +xfs_calc_group_awu_align_max(
> + struct xfs_mount *mp,
> + enum xfs_group_type type)
> +{
> + struct xfs_groups *g = &mp->m_groups[type];
> +
> + if (g->blocks == 0)
> + return 0;
> + return max_pow_of_two_factor(g->blocks);
> +}
> +
> /* Compute the maximum atomic write unit size for each section. */
> static inline void
> xfs_calc_atomic_write_unit_max(
> diff --git a/fs/xfs/xfs_mount.h b/fs/xfs/xfs_mount.h
> index 216a38a354e7..a2894c18e3c1 100644
> --- a/fs/xfs/xfs_mount.h
> +++ b/fs/xfs/xfs_mount.h
> @@ -812,6 +812,8 @@ static inline void xfs_mod_sb_delalloc(struct xfs_mount *mp, int64_t delta)
>
> int xfs_set_max_atomic_write_opt(struct xfs_mount *mp,
> unsigned long long new_max_bytes);
> +xfs_extlen_t xfs_calc_group_awu_align_max(struct xfs_mount *mp,
> + enum xfs_group_type type);
>
> static inline struct xfs_buftarg *
> xfs_group_type_buftarg(
>
> base-commit: 1282269a5ddb044481e7f8fd43b3195e211d5475
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-25 11:44 ` John Garry
@ 2026-09-29 7:17 ` Pankaj Raghav
2026-09-29 14:47 ` John Garry
2026-09-29 16:40 ` Darrick J. Wong
0 siblings, 2 replies; 19+ messages in thread
From: Pankaj Raghav @ 2026-09-29 7:17 UTC (permalink / raw)
To: John Garry, Pankaj Raghav, cem, linux-xfs
Cc: John Garry, Darrick J . Wong, gost.dev
On 9/25/2026 1:44 PM, John Garry wrote:
> On 9/25/26 11:36, Pankaj Raghav wrote:
> > xfs_calc_group_awu_max() clamps the atomic write unit maximum to the
> > greatest power-of-two factor of the group size when the device advertises
> > atomic writes, so that allocations can be made naturally aligned for
> > REQ_ATOMIC. But that value is the software limit: it is only used when
> > reflink is enabled, and out of place writes through the COW fork have no
> > alignment requirement.
>
> We have mkfs atomic write options for selecting atomic write limits - you have
> tried that, right?
>
Hmm, I did not try that but I restricted the agsize to a power of 2 value to
workaround the limitation.
Are you talking about this option: max_atomic_write?
In any case, with default options the atomic values we are exposing is not
correct at the moment.
>>
>> A filesystem on a > 4TB device that advertises 16k atomic writes then reports
>> an atomic write unit maximum of a single fsblock and an optimal maximum of 0,
>> and rejects any larger RWF_ATOMIC write. That is worse than the same filesystem
>> on a device with no atomic write support at all.
>>
>> Compute the software limit from the group size alone, and apply the
>> alignment constraint in xfs_get_atomic_write_max_opt() instead, which is
>> what reports the size that can be offloaded to the hardware.
>>
>> Results on a 8TB device with 16k hardware atomic support with 4k
>> blocksize:
>>
>> Before patches:
>> /media/test/hello.txt:
>> stx_atomic_write_unit_min: 4096
>> stx_atomic_write_unit_max: 4096
>> stx_atomic_write_unit_max_opt: 0
>> stx_atomic_write_segments_max: 1
>>
>> After patches:
>> /media/test/hello.txt:
>> stx_atomic_write_unit_min: 4096
>> stx_atomic_write_unit_max: 2097152
>
> Here the AG size may not be a multiple of 2097152B, right? If so, when we try to
> naturally align the allocation for a CoW write, even if the AG blocks are
> naturally aligned (per AG), the disk blocks may not be naturnally aligned - is
> this correct? For HW-based atomic writes, we want naturally aligned disk blocks.
>
Correct.
Just to give more context:
Trace of calc_atomic_write_unit_max:
mount-696 [010] 449.230543: xfs_calc_atomic_write_unit_max: dev 259:0 ag
max_write 262144 max_ioend 512 max_gsize 134217728 awu_max 512
2097152 comes from awumax being 512 fsblocks.
root@debian:~# xfs_info /dev/nvme0n1
meta-data=/dev/nvme0n1 isize=512 agcount=8, agsize=268435455 blks
= sectsz=4096 attr=2, projid32bit=1
= crc=1 finobt=1, sparse=1, rmapbt=0
= reflink=1 bigtime=1 inobtcount=1 nrext64=0
data = bsize=4096 blocks=2147483640, imaxpct=5
= sunit=0 swidth=0 blks
naming =version 2 bsize=4096 ascii-ci=0, ftype=1
log =internal log bsize=4096 blocks=521728, version=2
= sectsz=4096 sunit=1 blks, lazy-count=1
realtime =none extsz=4096 blocks=0, rtextents=0
>> stx_atomic_write_unit_max_opt: 4096
>> stx_atomic_write_segments_max: 1
>>
>> Results on a 8TB device with 16k hardware atomic support with 16k
>> blocksize:
>>
>> Before patches:
>> /media/test/hello.txt:
>> stx_atomic_write_unit_min: 16384
>> stx_atomic_write_unit_max: 16384
>> stx_atomic_write_unit_max_opt: 0
>> stx_atomic_write_segments_max: 1
>>
>> After patches:
>> /media/test/hello.txt:
>> stx_atomic_write_unit_min: 16384
>> stx_atomic_write_unit_max: 33554432
>> stx_atomic_write_unit_max_opt: 16384
>> stx_atomic_write_segments_max: 1
>>
--
Pankaj
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-29 7:17 ` Pankaj Raghav
@ 2026-09-29 14:47 ` John Garry
2026-09-30 5:10 ` Pankaj Raghav (Samsung)
2026-09-29 16:40 ` Darrick J. Wong
1 sibling, 1 reply; 19+ messages in thread
From: John Garry @ 2026-09-29 14:47 UTC (permalink / raw)
To: Pankaj Raghav, Pankaj Raghav, cem, linux-xfs
Cc: John Garry, Darrick J . Wong, gost.dev
On 9/29/26 08:17, Pankaj Raghav wrote:
>
>
> On 9/25/2026 1:44 PM, John Garry wrote:
>> On 9/25/26 11:36, Pankaj Raghav wrote:
>> > xfs_calc_group_awu_max() clamps the atomic write unit maximum to the
>> > greatest power-of-two factor of the group size when the device
>> advertises
>> > atomic writes, so that allocations can be made naturally aligned for
>> > REQ_ATOMIC. But that value is the software limit: it is only used when
>> > reflink is enabled, and out of place writes through the COW fork
>> have no
>> > alignment requirement.
>>
>> We have mkfs atomic write options for selecting atomic write limits -
>> you have tried that, right?
>>
>
> Hmm, I did not try that but I restricted the agsize to a power of 2
> value to
> workaround the limitation.
>
> Are you talking about this option: max_atomic_write?
Yeah, so there is a mkfs and also a mount option.
For the mount option you again provide a desired awu max. From the
desired awu max the FS checks whether that is possible (from AG count,
etc) and then it also sizes relevant "transaction memories" for
CoW-based atomic writes accordingly (to satisfy the awu max).
>
> In any case, with default options the atomic values we are exposing is
> not correct at the moment.
>
>>>
>>> A filesystem on a > 4TB device that advertises 16k atomic writes then
>>> reports
>>> an atomic write unit maximum of a single fsblock and an optimal
>>> maximum of 0,
>>> and rejects any larger RWF_ATOMIC write. That is worse than the same
>>> filesystem
>>> on a device with no atomic write support at all.
>>>
>>> Compute the software limit from the group size alone, and apply the
>>> alignment constraint in xfs_get_atomic_write_max_opt() instead, which is
>>> what reports the size that can be offloaded to the hardware.
>>>
>>> Results on a 8TB device with 16k hardware atomic support with 4k
>>> blocksize:
>>>
>>> Before patches:
>>> /media/test/hello.txt:
>>> stx_atomic_write_unit_min: 4096
>>> stx_atomic_write_unit_max: 4096
>>> stx_atomic_write_unit_max_opt: 0
>>> stx_atomic_write_segments_max: 1
>>>
>>> After patches:
>>> /media/test/hello.txt:
>>> stx_atomic_write_unit_min: 4096
>>> stx_atomic_write_unit_max: 2097152
>>
>> Here the AG size may not be a multiple of 2097152B, right? If so, when
>> we try to naturally align the allocation for a CoW write, even if the
>> AG blocks are naturally aligned (per AG), the disk blocks may not be
>> naturnally aligned - is this correct? For HW-based atomic writes, we
>> want naturally aligned disk blocks.
>>
>
> Correct.
>
> Just to give more context:
>
> Trace of calc_atomic_write_unit_max:
>
> mount-696 [010] 449.230543: xfs_calc_atomic_write_unit_max: dev
> 259:0 ag max_write 262144 max_ioend 512 max_gsize 134217728 awu_max 512
>
> 2097152 comes from awumax being 512 fsblocks.
>
> root@debian:~# xfs_info /dev/nvme0n1
> meta-data=/dev/nvme0n1 isize=512 agcount=8,
> agsize=268435455 blks
Well max power-of-2 factor of this is going to be 1.
> = sectsz=4096 attr=2, projid32bit=1
> = crc=1 finobt=1, sparse=1, rmapbt=0
> = reflink=1 bigtime=1 inobtcount=1
> nrext64=0
> data = bsize=4096 blocks=2147483640, imaxpct=5
> = sunit=0 swidth=0 blks
> naming =version 2 bsize=4096 ascii-ci=0, ftype=1
> log =internal log bsize=4096 blocks=521728, version=2
> = sectsz=4096 sunit=1 blks, lazy-count=1
> realtime =none extsz=4096 blocks=0, rtextents=0
>
>
>>> stx_atomic_write_unit_max_opt: 4096
>>> stx_atomic_write_segments_max: 1
>>>
>>> Results on a 8TB device with 16k hardware atomic support with 16k
>>> blocksize:
>>>
>>> Before patches:
>>> /media/test/hello.txt:
>>> stx_atomic_write_unit_min: 16384
>>> stx_atomic_write_unit_max: 16384
>>> stx_atomic_write_unit_max_opt: 0
>>> stx_atomic_write_segments_max: 1
>>>
>>> After patches:
>>> /media/test/hello.txt:
>>> stx_atomic_write_unit_min: 16384
>>> stx_atomic_write_unit_max: 33554432
>>> stx_atomic_write_unit_max_opt: 16384
>>> stx_atomic_write_segments_max: 1
>>>
> --
> Pankaj
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-29 14:47 ` John Garry
@ 2026-09-30 5:10 ` Pankaj Raghav (Samsung)
2026-09-30 8:42 ` John Garry
0 siblings, 1 reply; 19+ messages in thread
From: Pankaj Raghav (Samsung) @ 2026-09-30 5:10 UTC (permalink / raw)
To: John Garry
Cc: Pankaj Raghav, cem, linux-xfs, John Garry, Darrick J . Wong,
gost.dev
> > Just to give more context:
> >
> > Trace of calc_atomic_write_unit_max:
> >
> > mount-696 [010] 449.230543: xfs_calc_atomic_write_unit_max: dev
> > 259:0 ag max_write 262144 max_ioend 512 max_gsize 134217728 awu_max 512
> >
> > 2097152 comes from awumax being 512 fsblocks.
> >
> > root@debian:~# xfs_info /dev/nvme0n1
> > meta-data=/dev/nvme0n1 isize=512 agcount=8,
> > agsize=268435455 blks
>
> Well max power-of-2 factor of this is going to be 1.
>
Exactly. I get that we should restrict max_opt to be 1 because the AG's
will not be aligned and the best we can do is one fsblock for HW based
atomics. Why should we restrict the max (SW atomics limit) to 1 fsblock?
--
Pankaj
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-30 5:10 ` Pankaj Raghav (Samsung)
@ 2026-09-30 8:42 ` John Garry
2026-09-30 12:27 ` Pankaj Raghav
0 siblings, 1 reply; 19+ messages in thread
From: John Garry @ 2026-09-30 8:42 UTC (permalink / raw)
To: Pankaj Raghav (Samsung)
Cc: Pankaj Raghav, cem, linux-xfs, John Garry, Darrick J . Wong,
gost.dev
On 9/30/26 06:10, Pankaj Raghav (Samsung) wrote:
>>> Just to give more context:
>>>
>>> Trace of calc_atomic_write_unit_max:
>>>
>>> mount-696 [010] 449.230543: xfs_calc_atomic_write_unit_max: dev
>>> 259:0 ag max_write 262144 max_ioend 512 max_gsize 134217728 awu_max 512
>>>
>>> 2097152 comes from awumax being 512 fsblocks.
>>>
>>> root@debian:~# xfs_info /dev/nvme0n1
>>> meta-data=/dev/nvme0n1 isize=512 agcount=8,
>>> agsize=268435455 blks
>>
>> Well max power-of-2 factor of this is going to be 1.
>>
>
> Exactly. I get that we should restrict max_opt to be 1 because the AG's
> will not be aligned and the best we can do is one fsblock for HW based
> atomics. Why should we restrict the max (SW atomics limit) to 1 fsblock?
>
Maybe in this case we don't need to restrict CoW-based atomics to 1
fsblock. But why care? I mean, you have HW support, which is so much
better to use than CoW-based atomics, so better to config your FS to
avail of them.
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-30 8:42 ` John Garry
@ 2026-09-30 12:27 ` Pankaj Raghav
2026-09-30 12:48 ` John Garry
0 siblings, 1 reply; 19+ messages in thread
From: Pankaj Raghav @ 2026-09-30 12:27 UTC (permalink / raw)
To: John Garry, Darrick J . Wong
Cc: Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
On 9/30/2026 10:42 AM, John Garry wrote:
> On 9/30/26 06:10, Pankaj Raghav (Samsung) wrote:
>>>> Just to give more context:
>>>>
>>>> Trace of calc_atomic_write_unit_max:
>>>>
>>>> mount-696 [010] 449.230543: xfs_calc_atomic_write_unit_max: dev
>>>> 259:0 ag max_write 262144 max_ioend 512 max_gsize 134217728 awu_max 512
>>>>
>>>> 2097152 comes from awumax being 512 fsblocks.
>>>>
>>>> root@debian:~# xfs_info /dev/nvme0n1
>>>> meta-data=/dev/nvme0n1 isize=512 agcount=8,
>>>> agsize=268435455 blks
>>>
>>> Well max power-of-2 factor of this is going to be 1.
>>>
>>
>> Exactly. I get that we should restrict max_opt to be 1 because the AG's
>> will not be aligned and the best we can do is one fsblock for HW based
>> atomics. Why should we restrict the max (SW atomics limit) to 1 fsblock?
>>
>
> Maybe in this case we don't need to restrict CoW-based atomics to 1 fsblock. But
> why care? I mean, you have HW support, which is so much better to use than CoW-
> based atomics, so better to config your FS to avail of them.
That is true but it is a bit strange we expose a smaller CoW based atomics when
we have HW based atomic support.
For me the main issue is this function:
static xfs_extlen_t
xfs_calc_group_awu_max(
struct xfs_mount *mp,
enum xfs_group_type type)
{
struct xfs_groups *g = &mp->m_groups[type];
struct xfs_buftarg *btp = xfs_group_type_buftarg(mp, type);
if (g->blocks == 0)
return 0;
if (btp && btp->bt_awu_min > 0)
return max_pow_of_two_factor(g->blocks);
return rounddown_pow_of_two(g->blocks);
}
Why do we take into account hardware limits when we calculate SW atomic limits?
For more context:
In my drive, I am getting odd nr_blks agsize for the drive I have when I format
it with 16k block size. This is fine because block size is aligned
with HW atomic write size. The issue comes in xfs_calc_group_awu_max where
we do max_pow_of_two_factor on blocks if HW based atomics are enabled, thereby,
awu_max is reported as 1.
So this results in the following
panky@blixen:~$ mkfs.xfs -V
mkfs.xfs version 7.1.1
panky@blixen:~$ xfs_info /mnt/atomssd/
meta-data=/dev/nvme1n1 isize=512 agcount=14, agsize=67108863 blks
= sectsz=4096 attr=2, projid32bit=1
= crc=1 finobt=1, sparse=1, rmapbt=1
= reflink=1 bigtime=1 inobtcount=1 nrext64=1
= exchange=1 metadir=0
data = bsize=16384 blocks=937558016, imaxpct=5
= sunit=0 swidth=0 blks
naming =version 2 bsize=16384 ascii-ci=0, ftype=1, parent=1
log =internal log bsize=16384 blocks=130432, version=2
= sectsz=4096 sunit=1 blks, lazy-count=1
realtime =none extsz=16384 blocks=0, rtextents=0
= rgcount=0 rgsize=0 extents
= zoned=0 start=0 reserved=0
root@blixen:~# trace-cmd report
cpus=8
mount-4432 [005] ..... 24183.039820:
xfs_calc_atomic_write_unit_max: dev 259:2 ag max_write 131072 max_ioend 4096
max_gsize 1 awu_max 1
mount-4432 [005] ..... 24183.039820:
xfs_calc_atomic_write_unit_max: dev 259:2 rtg max_write 131072 max_ioend 4096
max_gsize 0 awu_max 0
panky@blixen:~/tools$ sudo ./statx /mnt/atomssd/hello.txt
/mnt/atomssd/hello.txt:
stx_atomic_write_unit_min: 16384
stx_atomic_write_unit_max: 16384
stx_atomic_write_unit_max_opt: 0
stx_atomic_write_segments_max: 1
I hope I have explained the issue clearly now.
--
Pankaj
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-30 12:27 ` Pankaj Raghav
@ 2026-09-30 12:48 ` John Garry
2026-10-01 16:30 ` Pankaj Raghav
0 siblings, 1 reply; 19+ messages in thread
From: John Garry @ 2026-09-30 12:48 UTC (permalink / raw)
To: Pankaj Raghav, Darrick J . Wong
Cc: Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
On 9/30/26 13:27, Pankaj Raghav wrote:
>>
>> Maybe in this case we don't need to restrict CoW-based atomics to 1
>> fsblock. But why care? I mean, you have HW support, which is so much
>> better to use than CoW- based atomics, so better to config your FS to
>> avail of them.
>
> That is true but it is a bit strange we expose a smaller CoW based
> atomics when we have HW based atomic support.
>
> For me the main issue is this function:
> static xfs_extlen_t
> xfs_calc_group_awu_max(
> struct xfs_mount *mp,
> enum xfs_group_type type)
> {
> struct xfs_groups *g = &mp->m_groups[type];
> struct xfs_buftarg *btp = xfs_group_type_buftarg(mp, type);
>
> if (g->blocks == 0)
> return 0;
> if (btp && btp->bt_awu_min > 0)
> return max_pow_of_two_factor(g->blocks);
> return rounddown_pow_of_two(g->blocks);
> }
----->8-----
--- a/fs/xfs/xfs_mount.c
+++ b/fs/xfs/xfs_mount.c
@@ -694,7 +694,7 @@ xfs_calc_group_awu_max(
if (g->blocks == 0)
return 0;
- if (btp && btp->bt_awu_min > 0)
+ if (btp && btp->bt_awu_max > mp->m_sb.sb_blocksize)
return max_pow_of_two_factor(g->blocks);
return rounddown_pow_of_two(g->blocks);
}
-----8<-----
You are advocating something like this (to solve the issue below), right?
note: that we still need to ensure that bt_awu_min <= sb_blocksize to
get HW atomics at all (so should keep a check for bt_awu_min in
xfs_calc_group_awu_max() or similar)
>
> Why do we take into account hardware limits when we calculate SW atomic
> limits?
>
> For more context:
>
> In my drive, I am getting odd nr_blks agsize for the drive I have when I
> format
> it with 16k block size. This is fine because block size is aligned
> with HW atomic write size. The issue comes in xfs_calc_group_awu_max where
> we do max_pow_of_two_factor on blocks if HW based atomics are enabled,
> thereby,
> awu_max is reported as 1.
>
>
> So this results in the following
>
> panky@blixen:~$ mkfs.xfs -V
> mkfs.xfs version 7.1.1
>
> panky@blixen:~$ xfs_info /mnt/atomssd/
> meta-data=/dev/nvme1n1 isize=512 agcount=14,
> agsize=67108863 blks
> = sectsz=4096 attr=2, projid32bit=1
> = crc=1 finobt=1, sparse=1, rmapbt=1
> = reflink=1 bigtime=1 inobtcount=1
> nrext64=1
> = exchange=1 metadir=0
> data = bsize=16384 blocks=937558016, imaxpct=5
> = sunit=0 swidth=0 blks
> naming =version 2 bsize=16384 ascii-ci=0, ftype=1, parent=1
> log =internal log bsize=16384 blocks=130432, version=2
> = sectsz=4096 sunit=1 blks, lazy-count=1
> realtime =none extsz=16384 blocks=0, rtextents=0
> = rgcount=0 rgsize=0 extents
> = zoned=0 start=0 reserved=0
>
>
> root@blixen:~# trace-cmd report
> cpus=8
> mount-4432 [005] ..... 24183.039820:
> xfs_calc_atomic_write_unit_max: dev 259:2 ag max_write 131072 max_ioend
> 4096 max_gsize 1 awu_max 1
> mount-4432 [005] ..... 24183.039820:
> xfs_calc_atomic_write_unit_max: dev 259:2 rtg max_write 131072 max_ioend
> 4096 max_gsize 0 awu_max 0
>
> panky@blixen:~/tools$ sudo ./statx /mnt/atomssd/hello.txt
> /mnt/atomssd/hello.txt:
> stx_atomic_write_unit_min: 16384
> stx_atomic_write_unit_max: 16384
> stx_atomic_write_unit_max_opt: 0
> stx_atomic_write_segments_max: 1
>
>
> I hope I have explained the issue clearly now.
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-30 12:48 ` John Garry
@ 2026-10-01 16:30 ` Pankaj Raghav
2026-10-02 8:45 ` John Garry
0 siblings, 1 reply; 19+ messages in thread
From: Pankaj Raghav @ 2026-10-01 16:30 UTC (permalink / raw)
To: John Garry, Darrick J . Wong
Cc: Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
On 9/30/2026 2:48 PM, John Garry wrote:
> On 9/30/26 13:27, Pankaj Raghav wrote:
>>>
>>> Maybe in this case we don't need to restrict CoW-based atomics to 1 fsblock.
>>> But why care? I mean, you have HW support, which is so much better to use
>>> than CoW- based atomics, so better to config your FS to avail of them.
>>
>> That is true but it is a bit strange we expose a smaller CoW based atomics
>> when we have HW based atomic support.
>>
>> For me the main issue is this function:
>> static xfs_extlen_t
>> xfs_calc_group_awu_max(
>> struct xfs_mount *mp,
>> enum xfs_group_type type)
>> {
>> struct xfs_groups *g = &mp->m_groups[type];
>> struct xfs_buftarg *btp = xfs_group_type_buftarg(mp, type);
>>
>> if (g->blocks == 0)
>> return 0;
>> if (btp && btp->bt_awu_min > 0)
>> return max_pow_of_two_factor(g->blocks);
>> return rounddown_pow_of_two(g->blocks);
>> }
>
>
> ----->8-----
>
> --- a/fs/xfs/xfs_mount.c
> +++ b/fs/xfs/xfs_mount.c
> @@ -694,7 +694,7 @@ xfs_calc_group_awu_max(
>
> if (g->blocks == 0)
> return 0;
> - if (btp && btp->bt_awu_min > 0)
> + if (btp && btp->bt_awu_max > mp->m_sb.sb_blocksize)
> return max_pow_of_two_factor(g->blocks);
> return rounddown_pow_of_two(g->blocks);
> }
>
> -----8<-----
>
>
> You are advocating something like this (to solve the issue below), right?
>
Hmm, this might fix the issue but I still don't understand why we have hardware
atomics check while determining SW atomics limit. Am I missing something?
> note: that we still need to ensure that bt_awu_min <= sb_blocksize to get HW
> atomics at all (so should keep a check for bt_awu_min in
> xfs_calc_group_awu_max() or similar)
>
>>
--
Pankaj
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-10-01 16:30 ` Pankaj Raghav
@ 2026-10-02 8:45 ` John Garry
2026-10-02 11:24 ` Pankaj Raghav
0 siblings, 1 reply; 19+ messages in thread
From: John Garry @ 2026-10-02 8:45 UTC (permalink / raw)
To: Pankaj Raghav, Darrick J . Wong
Cc: Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
On 10/1/26 17:30, Pankaj Raghav wrote:
>>
>> ----->8-----
>>
>> --- a/fs/xfs/xfs_mount.c
>> +++ b/fs/xfs/xfs_mount.c
>> @@ -694,7 +694,7 @@ xfs_calc_group_awu_max(
>>
>> if (g->blocks == 0)
>> return 0;
>> - if (btp && btp->bt_awu_min > 0)
>> + if (btp && btp->bt_awu_max > mp->m_sb.sb_blocksize)
>> return max_pow_of_two_factor(g->blocks);
>> return rounddown_pow_of_two(g->blocks);
>> }
>>
>> -----8<-----
>>
>>
>> You are advocating something like this (to solve the issue below), right?
>>
>
> Hmm, this might fix the issue but I still don't understand why we have
> hardware atomics check while determining SW atomics limit. Am I missing
> something?
ok, I suppose that it (i.e. whether bt_awu_max > sb_blocksize) should
not determine CoW-based atomics limits, but it should determine HW
limits (in opt max).
>
>> note: that we still need to ensure that bt_awu_min <= sb_blocksize to
>> get HW atomics at all (so should keep a check for bt_awu_min in
>> xfs_calc_group_awu_max() or similar)
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-10-02 8:45 ` John Garry
@ 2026-10-02 11:24 ` Pankaj Raghav
0 siblings, 0 replies; 19+ messages in thread
From: Pankaj Raghav @ 2026-10-02 11:24 UTC (permalink / raw)
To: John Garry, Darrick J . Wong
Cc: Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
On 10/2/2026 10:45 AM, John Garry wrote:
> On 10/1/26 17:30, Pankaj Raghav wrote:
>>>
>>> ----->8-----
>>>
>>> --- a/fs/xfs/xfs_mount.c
>>> +++ b/fs/xfs/xfs_mount.c
>>> @@ -694,7 +694,7 @@ xfs_calc_group_awu_max(
>>>
>>> if (g->blocks == 0)
>>> return 0;
>>> - if (btp && btp->bt_awu_min > 0)
>>> + if (btp && btp->bt_awu_max > mp->m_sb.sb_blocksize)
>>> return max_pow_of_two_factor(g->blocks);
>>> return rounddown_pow_of_two(g->blocks);
>>> }
>>>
>>> -----8<-----
>>>
>>>
>>> You are advocating something like this (to solve the issue below), right?
>>>
>>
>> Hmm, this might fix the issue but I still don't understand why we have
>> hardware atomics check while determining SW atomics limit. Am I missing
>> something?
>
> ok, I suppose that it (i.e. whether bt_awu_max > sb_blocksize) should not
> determine CoW-based atomics limits, but it should determine HW limits (in opt max).
>
Could you take a look the patches again and see if it makes sense? We already
take into account the HW limits while calculating opt. So the new function
xfs_calc_group_awu_align_max() checks for alignment constraints because of agsize.
Without patches:
[nix-shell:~]# trace-cmd report
cpus=16
mount-617 [009] 1184.719159: xfs_calc_atomic_write_unit_max: dev 259:2 ag
max_write 131072 max_ioend 4096 max_gsize 2 awu_max 2
[nix-shell:~]# ./statx /media/test/hello.txt
/media/test/hello.txt:
stx_atomic_write_unit_min: 16384
stx_atomic_write_unit_max: 32768
stx_atomic_write_unit_max_opt: 16384
stx_atomic_write_segments_max: 1
With patches:
[nix-shell:~]# trace-cmd report
cpus=16
mount-598 [005] 319.692000: xfs_calc_atomic_write_unit_max: dev 259:2 ag
max_write 131072 max_ioend 4096 max_gsize 33554432 awu_max 4096
[nix-shell:~]# ./statx /media/test/hello.txt
/media/test/hello.txt:
stx_atomic_write_unit_min: 16384
stx_atomic_write_unit_max: 67108864
stx_atomic_write_unit_max_opt: 16384
stx_atomic_write_segments_max: 1
>>
>>> note: that we still need to ensure that bt_awu_min <= sb_blocksize to get HW
>>> atomics at all (so should keep a check for bt_awu_min in
>>> xfs_calc_group_awu_max() or similar)
>
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-29 7:17 ` Pankaj Raghav
2026-09-29 14:47 ` John Garry
@ 2026-09-29 16:40 ` Darrick J. Wong
2026-09-29 20:42 ` Darrick J. Wong
1 sibling, 1 reply; 19+ messages in thread
From: Darrick J. Wong @ 2026-09-29 16:40 UTC (permalink / raw)
To: Pankaj Raghav
Cc: John Garry, Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
On Tue, Sep 29, 2026 at 09:17:32AM +0200, Pankaj Raghav wrote:
>
>
> On 9/25/2026 1:44 PM, John Garry wrote:
> > On 9/25/26 11:36, Pankaj Raghav wrote:
> > > xfs_calc_group_awu_max() clamps the atomic write unit maximum to the
> > > greatest power-of-two factor of the group size when the device advertises
> > > atomic writes, so that allocations can be made naturally aligned for
> > > REQ_ATOMIC. But that value is the software limit: it is only used when
> > > reflink is enabled, and out of place writes through the COW fork have no
> > > alignment requirement.
> >
> > We have mkfs atomic write options for selecting atomic write limits -
> > you have tried that, right?
> >
>
> Hmm, I did not try that but I restricted the agsize to a power of 2 value to
> workaround the limitation.
>
> Are you talking about this option: max_atomic_write?
>
> In any case, with default options the atomic values we are exposing is not
> correct at the moment.
>
> > >
> > > A filesystem on a > 4TB device that advertises 16k atomic writes then reports
> > > an atomic write unit maximum of a single fsblock and an optimal maximum of 0,
> > > and rejects any larger RWF_ATOMIC write. That is worse than the same filesystem
> > > on a device with no atomic write support at all.
> > >
> > > Compute the software limit from the group size alone, and apply the
> > > alignment constraint in xfs_get_atomic_write_max_opt() instead, which is
> > > what reports the size that can be offloaded to the hardware.
> > >
> > > Results on a 8TB device with 16k hardware atomic support with 4k
> > > blocksize:
> > >
> > > Before patches:
> > > /media/test/hello.txt:
> > > stx_atomic_write_unit_min: 4096
> > > stx_atomic_write_unit_max: 4096
> > > stx_atomic_write_unit_max_opt: 0
> > > stx_atomic_write_segments_max: 1
> > >
> > > After patches:
> > > /media/test/hello.txt:
> > > stx_atomic_write_unit_min: 4096
> > > stx_atomic_write_unit_max: 2097152
> >
> > Here the AG size may not be a multiple of 2097152B, right? If so, when
> > we try to naturally align the allocation for a CoW write, even if the AG
> > blocks are naturally aligned (per AG), the disk blocks may not be
> > naturnally aligned - is this correct? For HW-based atomic writes, we
> > want naturally aligned disk blocks.
> >
>
> Correct.
>
> Just to give more context:
>
> Trace of calc_atomic_write_unit_max:
>
> mount-696 [010] 449.230543: xfs_calc_atomic_write_unit_max: dev 259:0
> ag max_write 262144 max_ioend 512 max_gsize 134217728 awu_max 512
>
> 2097152 comes from awumax being 512 fsblocks.
>
> root@debian:~# xfs_info /dev/nvme0n1
> meta-data=/dev/nvme0n1 isize=512 agcount=8, agsize=268435455 blks
268435455? That's not aligned to the atomic write size, but that is the
max AG size. I thought mkfs would round it down for us automatically:
/*
* We've already validated (or discarded) the hardware atomic write
* geometry. Try to align the agsize to the maximum atomic write unit
* to give users maximum flexibility in choosing atomic write sizes.
*/
if (ft->data.awu_max > 0)
dsunit = max(DTOBT(ft->data.awu_max, cfg->blocklog),
dsunit);
But I guess I'll play around with it for a while and see if I get
anywhere. Assuming you're running a recent mkfs that knows about
the atomic write options and whatnot?
--D
> = sectsz=4096 attr=2, projid32bit=1
> = crc=1 finobt=1, sparse=1, rmapbt=0
> = reflink=1 bigtime=1 inobtcount=1 nrext64=0
> data = bsize=4096 blocks=2147483640, imaxpct=5
> = sunit=0 swidth=0 blks
> naming =version 2 bsize=4096 ascii-ci=0, ftype=1
> log =internal log bsize=4096 blocks=521728, version=2
> = sectsz=4096 sunit=1 blks, lazy-count=1
> realtime =none extsz=4096 blocks=0, rtextents=0
>
>
> > > stx_atomic_write_unit_max_opt: 4096
> > > stx_atomic_write_segments_max: 1
> > >
> > > Results on a 8TB device with 16k hardware atomic support with 16k
> > > blocksize:
> > >
> > > Before patches:
> > > /media/test/hello.txt:
> > > stx_atomic_write_unit_min: 16384
> > > stx_atomic_write_unit_max: 16384
> > > stx_atomic_write_unit_max_opt: 0
> > > stx_atomic_write_segments_max: 1
> > >
> > > After patches:
> > > /media/test/hello.txt:
> > > stx_atomic_write_unit_min: 16384
> > > stx_atomic_write_unit_max: 33554432
> > > stx_atomic_write_unit_max_opt: 16384
> > > stx_atomic_write_segments_max: 1
> > >
> --
> Pankaj
>
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-29 16:40 ` Darrick J. Wong
@ 2026-09-29 20:42 ` Darrick J. Wong
2026-09-30 8:48 ` Pankaj Raghav (Samsung)
0 siblings, 1 reply; 19+ messages in thread
From: Darrick J. Wong @ 2026-09-29 20:42 UTC (permalink / raw)
To: Pankaj Raghav
Cc: John Garry, Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
On Tue, Sep 29, 2026 at 09:40:10AM -0700, Darrick J. Wong wrote:
> On Tue, Sep 29, 2026 at 09:17:32AM +0200, Pankaj Raghav wrote:
> >
> >
> > On 9/25/2026 1:44 PM, John Garry wrote:
> > > On 9/25/26 11:36, Pankaj Raghav wrote:
> > > > xfs_calc_group_awu_max() clamps the atomic write unit maximum to the
> > > > greatest power-of-two factor of the group size when the device advertises
> > > > atomic writes, so that allocations can be made naturally aligned for
> > > > REQ_ATOMIC. But that value is the software limit: it is only used when
> > > > reflink is enabled, and out of place writes through the COW fork have no
> > > > alignment requirement.
> > >
> > > We have mkfs atomic write options for selecting atomic write limits -
> > > you have tried that, right?
> > >
> >
> > Hmm, I did not try that but I restricted the agsize to a power of 2 value to
> > workaround the limitation.
> >
> > Are you talking about this option: max_atomic_write?
> >
> > In any case, with default options the atomic values we are exposing is not
> > correct at the moment.
> >
> > > >
> > > > A filesystem on a > 4TB device that advertises 16k atomic writes then reports
> > > > an atomic write unit maximum of a single fsblock and an optimal maximum of 0,
> > > > and rejects any larger RWF_ATOMIC write. That is worse than the same filesystem
> > > > on a device with no atomic write support at all.
> > > >
> > > > Compute the software limit from the group size alone, and apply the
> > > > alignment constraint in xfs_get_atomic_write_max_opt() instead, which is
> > > > what reports the size that can be offloaded to the hardware.
> > > >
> > > > Results on a 8TB device with 16k hardware atomic support with 4k
> > > > blocksize:
> > > >
> > > > Before patches:
> > > > /media/test/hello.txt:
> > > > stx_atomic_write_unit_min: 4096
> > > > stx_atomic_write_unit_max: 4096
> > > > stx_atomic_write_unit_max_opt: 0
> > > > stx_atomic_write_segments_max: 1
> > > >
> > > > After patches:
> > > > /media/test/hello.txt:
> > > > stx_atomic_write_unit_min: 4096
> > > > stx_atomic_write_unit_max: 2097152
> > >
> > > Here the AG size may not be a multiple of 2097152B, right? If so, when
> > > we try to naturally align the allocation for a CoW write, even if the AG
> > > blocks are naturally aligned (per AG), the disk blocks may not be
> > > naturnally aligned - is this correct? For HW-based atomic writes, we
> > > want naturally aligned disk blocks.
> > >
> >
> > Correct.
> >
> > Just to give more context:
> >
> > Trace of calc_atomic_write_unit_max:
> >
> > mount-696 [010] 449.230543: xfs_calc_atomic_write_unit_max: dev 259:0
> > ag max_write 262144 max_ioend 512 max_gsize 134217728 awu_max 512
> >
> > 2097152 comes from awumax being 512 fsblocks.
> >
> > root@debian:~# xfs_info /dev/nvme0n1
> > meta-data=/dev/nvme0n1 isize=512 agcount=8, agsize=268435455 blks
>
> 268435455? That's not aligned to the atomic write size, but that is the
> max AG size. I thought mkfs would round it down for us automatically:
>
> /*
> * We've already validated (or discarded) the hardware atomic write
> * geometry. Try to align the agsize to the maximum atomic write unit
> * to give users maximum flexibility in choosing atomic write sizes.
> */
> if (ft->data.awu_max > 0)
> dsunit = max(DTOBT(ft->data.awu_max, cfg->blocklog),
> dsunit);
>
> But I guess I'll play around with it for a while and see if I get
> anywhere. Assuming you're running a recent mkfs that knows about
> the atomic write options and whatnot?
FWIW I tried this with scsi-debug and got the following:
# umount /dev/sdg
# rmmod scsi-debug
# modprobe scsi-debug dev_size_mb=1000 virtual_gb=4096 atomic_wr=1 atomic_wr_max_length=32
# xfs_io -c 'statx -r -m all ' /dev/sdg | grep atomic
stat.atomic_write_unit_min = 1024
stat.atomic_write_unit_max = 16384
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 0
# mkfs.xfs -f /dev/sdg
meta-data=/dev/sdg isize=512 agcount=4, agsize=268435452 blks
= sectsz=512 attr=2, projid32bit=1
= crc=1 finobt=1, sparse=1, rmapbt=1
= reflink=1 bigtime=1 inobtcount=1 nrext64=1
= exchange=1 metadir=0
data = bsize=4096 blocks=1073741808, imaxpct=5
= sunit=0 swidth=0 blks
naming =version 2 bsize=4096 ascii-ci=0, ftype=1, parent=1
log =internal log bsize=4096 blocks=521728, version=2
= sectsz=512 sunit=0 blks, lazy-count=1
realtime =none extsz=4096 blocks=0, rtextents=0
= rgcount=0 rgsize=0 extents
= zoned=0 start=0 reserved=0
# mount /dev/sdg /mnt
# touch /mnt/fubar
# xfs_io -c 'statx -r -m all' /mnt/fubar | grep atomic
stat.atomic_write_unit_min = 4096
stat.atomic_write_unit_max = 16384
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 16384
Note the agsize=268435452, which means that it rounded the AG size down
to something congruent with the 16k hardware atomic write max.
(Are you sure you're using mkfs.xfs from xfsprogs 6.16 or newer?)
> --D
>
> > = sectsz=4096 attr=2, projid32bit=1
> > = crc=1 finobt=1, sparse=1, rmapbt=0
> > = reflink=1 bigtime=1 inobtcount=1 nrext64=0
> > data = bsize=4096 blocks=2147483640, imaxpct=5
> > = sunit=0 swidth=0 blks
> > naming =version 2 bsize=4096 ascii-ci=0, ftype=1
> > log =internal log bsize=4096 blocks=521728, version=2
> > = sectsz=4096 sunit=1 blks, lazy-count=1
> > realtime =none extsz=4096 blocks=0, rtextents=0
> >
> >
> > > > stx_atomic_write_unit_max_opt: 4096
> > > > stx_atomic_write_segments_max: 1
> > > >
> > > > Results on a 8TB device with 16k hardware atomic support with 16k
> > > > blocksize:
> > > >
> > > > Before patches:
> > > > /media/test/hello.txt:
> > > > stx_atomic_write_unit_min: 16384
> > > > stx_atomic_write_unit_max: 16384
> > > > stx_atomic_write_unit_max_opt: 0
> > > > stx_atomic_write_segments_max: 1
> > > >
> > > > After patches:
> > > > /media/test/hello.txt:
> > > > stx_atomic_write_unit_min: 16384
> > > > stx_atomic_write_unit_max: 33554432
> > > > stx_atomic_write_unit_max_opt: 16384
> > > > stx_atomic_write_segments_max: 1
> > > >
> > --
> > Pankaj
> >
>
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-29 20:42 ` Darrick J. Wong
@ 2026-09-30 8:48 ` Pankaj Raghav (Samsung)
2026-10-07 17:43 ` Darrick J. Wong
0 siblings, 1 reply; 19+ messages in thread
From: Pankaj Raghav (Samsung) @ 2026-09-30 8:48 UTC (permalink / raw)
To: Darrick J. Wong, john.garry
Cc: Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
> > But I guess I'll play around with it for a while and see if I get
> > anywhere. Assuming you're running a recent mkfs that knows about
> > the atomic write options and whatnot?
>
> FWIW I tried this with scsi-debug and got the following:
>
> # umount /dev/sdg
> # rmmod scsi-debug
> # modprobe scsi-debug dev_size_mb=1000 virtual_gb=4096 atomic_wr=1 atomic_wr_max_length=32
> # xfs_io -c 'statx -r -m all ' /dev/sdg | grep atomic
> stat.atomic_write_unit_min = 1024
> stat.atomic_write_unit_max = 16384
> stat.atomic_write_segments_max = 1
> stat.atomic_write_unit_max_opt = 0
> # mkfs.xfs -f /dev/sdg
> meta-data=/dev/sdg isize=512 agcount=4, agsize=268435452 blks
> = sectsz=512 attr=2, projid32bit=1
> = crc=1 finobt=1, sparse=1, rmapbt=1
> = reflink=1 bigtime=1 inobtcount=1 nrext64=1
> = exchange=1 metadir=0
> data = bsize=4096 blocks=1073741808, imaxpct=5
> = sunit=0 swidth=0 blks
> naming =version 2 bsize=4096 ascii-ci=0, ftype=1, parent=1
> log =internal log bsize=4096 blocks=521728, version=2
> = sectsz=512 sunit=0 blks, lazy-count=1
> realtime =none extsz=4096 blocks=0, rtextents=0
> = rgcount=0 rgsize=0 extents
> = zoned=0 start=0 reserved=0
> # mount /dev/sdg /mnt
> # touch /mnt/fubar
> # xfs_io -c 'statx -r -m all' /mnt/fubar | grep atomic
> stat.atomic_write_unit_min = 4096
> stat.atomic_write_unit_max = 16384
> stat.atomic_write_segments_max = 1
> stat.atomic_write_unit_max_opt = 16384
>
> Note the agsize=268435452, which means that it rounded the AG size down
> to something congruent with the 16k hardware atomic write max.
>
> (Are you sure you're using mkfs.xfs from xfsprogs 6.16 or newer?)
I noticed the issue when I was running a test on a bare metal server
with xfsprog 7.1.1. But I tried to recreate the issue in a VM where I
was running an old xfsprog. So the values I have posted here are wrong.
Sorry for that.
> # xfs_io -c 'statx -r -m all' /mnt/fubar | grep atomic
> stat.atomic_write_unit_min = 4096
> stat.atomic_write_unit_max = 16384
> stat.atomic_write_segments_max = 1
> stat.atomic_write_unit_max_opt = 16384
However, I still do not understand why do we have to restrict
atomic_write_unit_max (SW based atomics) to 16k here?
Let me take a step back and give a history on why I started to look into
this.
I am running 7.2-rc6 with 7.1.1 xfsprogs on a bare metal server with a
NVMe device that supports atomics.
I am formatting XFS with 16k blocksize.
panky@blixen:~$ df -h /dev/nvme1n1
Filesystem Size Used Avail Use% Mounted on
/dev/nvme1n1 14T 179G 14T 2% /mnt/atomssd
panky@blixen:~$ cat /sys/block/nvme1n1/queue/atomic_write_unit_max_bytes
16384
panky@blixen:~$ xfs_info /mnt/atomssd/
meta-data=/dev/nvme1n1 isize=512 agcount=14, agsize=67108863 blks
= sectsz=4096 attr=2, projid32bit=1
= crc=1 finobt=1, sparse=1, rmapbt=1
= reflink=1 bigtime=1 inobtcount=1 nrext64=1
= exchange=1 metadir=0
data = bsize=16384 blocks=937558016, imaxpct=5
= sunit=0 swidth=0 blks
naming =version 2 bsize=16384 ascii-ci=0, ftype=1, parent=1
log =internal log bsize=16384 blocks=130432, version=2
= sectsz=4096 sunit=1 blks, lazy-count=1
realtime =none extsz=16384 blocks=0, rtextents=0
= rgcount=0 rgsize=0 extents
= zoned=0 start=0 reserved=0
panky@blixen:~/tools$ sudo ./statx /mnt/atomssd/hello.txt
[sudo: authenticate] Password:
/mnt/atomssd/hello.txt:
stx_atomic_write_unit_min: 16384
stx_atomic_write_unit_max: 16384
stx_atomic_write_unit_max_opt: 0
stx_atomic_write_segments_max: 1
We expose max_opt to be 0, which means the application can assume
max_opt to be the same as atomic_write_unit_max. But why are we not
exposing bigger atomic_write_unit_max (SW atomics) here and make max_opt
to HW based atomics value (16k) in this case?
Could you format your scsi-debug device with 16k fsblock size and see
what values you are getting here?
--
Pankaj
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-30 8:48 ` Pankaj Raghav (Samsung)
@ 2026-10-07 17:43 ` Darrick J. Wong
0 siblings, 0 replies; 19+ messages in thread
From: Darrick J. Wong @ 2026-10-07 17:43 UTC (permalink / raw)
To: Pankaj Raghav (Samsung)
Cc: john.garry, Pankaj Raghav, cem, linux-xfs, John Garry, gost.dev
On Wed, Sep 30, 2026 at 08:48:56AM +0000, Pankaj Raghav (Samsung) wrote:
> > > But I guess I'll play around with it for a while and see if I get
> > > anywhere. Assuming you're running a recent mkfs that knows about
> > > the atomic write options and whatnot?
> >
> > FWIW I tried this with scsi-debug and got the following:
> >
> > # umount /dev/sdg
> > # rmmod scsi-debug
> > # modprobe scsi-debug dev_size_mb=1000 virtual_gb=4096 atomic_wr=1 atomic_wr_max_length=32
> > # xfs_io -c 'statx -r -m all ' /dev/sdg | grep atomic
> > stat.atomic_write_unit_min = 1024
> > stat.atomic_write_unit_max = 16384
> > stat.atomic_write_segments_max = 1
> > stat.atomic_write_unit_max_opt = 0
> > # mkfs.xfs -f /dev/sdg
> > meta-data=/dev/sdg isize=512 agcount=4, agsize=268435452 blks
> > = sectsz=512 attr=2, projid32bit=1
> > = crc=1 finobt=1, sparse=1, rmapbt=1
> > = reflink=1 bigtime=1 inobtcount=1 nrext64=1
> > = exchange=1 metadir=0
> > data = bsize=4096 blocks=1073741808, imaxpct=5
> > = sunit=0 swidth=0 blks
> > naming =version 2 bsize=4096 ascii-ci=0, ftype=1, parent=1
> > log =internal log bsize=4096 blocks=521728, version=2
> > = sectsz=512 sunit=0 blks, lazy-count=1
> > realtime =none extsz=4096 blocks=0, rtextents=0
> > = rgcount=0 rgsize=0 extents
> > = zoned=0 start=0 reserved=0
> > # mount /dev/sdg /mnt
> > # touch /mnt/fubar
> > # xfs_io -c 'statx -r -m all' /mnt/fubar | grep atomic
> > stat.atomic_write_unit_min = 4096
> > stat.atomic_write_unit_max = 16384
> > stat.atomic_write_segments_max = 1
> > stat.atomic_write_unit_max_opt = 16384
> >
> > Note the agsize=268435452, which means that it rounded the AG size down
> > to something congruent with the 16k hardware atomic write max.
> >
> > (Are you sure you're using mkfs.xfs from xfsprogs 6.16 or newer?)
>
> I noticed the issue when I was running a test on a bare metal server
> with xfsprog 7.1.1. But I tried to recreate the issue in a VM where I
> was running an old xfsprog. So the values I have posted here are wrong.
> Sorry for that.
>
> > # xfs_io -c 'statx -r -m all' /mnt/fubar | grep atomic
> > stat.atomic_write_unit_min = 4096
> > stat.atomic_write_unit_max = 16384
> > stat.atomic_write_segments_max = 1
> > stat.atomic_write_unit_max_opt = 16384
>
> However, I still do not understand why do we have to restrict
> atomic_write_unit_max (SW based atomics) to 16k here?
>
> Let me take a step back and give a history on why I started to look into
> this.
>
> I am running 7.2-rc6 with 7.1.1 xfsprogs on a bare metal server with a
> NVMe device that supports atomics.
>
> I am formatting XFS with 16k blocksize.
>
> panky@blixen:~$ df -h /dev/nvme1n1
> Filesystem Size Used Avail Use% Mounted on
> /dev/nvme1n1 14T 179G 14T 2% /mnt/atomssd
>
> panky@blixen:~$ cat /sys/block/nvme1n1/queue/atomic_write_unit_max_bytes
> 16384
>
> panky@blixen:~$ xfs_info /mnt/atomssd/
> meta-data=/dev/nvme1n1 isize=512 agcount=14, agsize=67108863 blks
> = sectsz=4096 attr=2, projid32bit=1
> = crc=1 finobt=1, sparse=1, rmapbt=1
> = reflink=1 bigtime=1 inobtcount=1 nrext64=1
> = exchange=1 metadir=0
> data = bsize=16384 blocks=937558016, imaxpct=5
> = sunit=0 swidth=0 blks
> naming =version 2 bsize=16384 ascii-ci=0, ftype=1, parent=1
> log =internal log bsize=16384 blocks=130432, version=2
> = sectsz=4096 sunit=1 blks, lazy-count=1
> realtime =none extsz=16384 blocks=0, rtextents=0
> = rgcount=0 rgsize=0 extents
> = zoned=0 start=0 reserved=0
>
> panky@blixen:~/tools$ sudo ./statx /mnt/atomssd/hello.txt
> [sudo: authenticate] Password:
> /mnt/atomssd/hello.txt:
> stx_atomic_write_unit_min: 16384
> stx_atomic_write_unit_max: 16384
> stx_atomic_write_unit_max_opt: 0
> stx_atomic_write_segments_max: 1
>
> We expose max_opt to be 0, which means the application can assume
> max_opt to be the same as atomic_write_unit_max. But why are we not
> exposing bigger atomic_write_unit_max (SW atomics) here and make max_opt
> to HW based atomics value (16k) in this case?
>
> Could you format your scsi-debug device with 16k fsblock size and see
> what values you are getting here?
(Apologies, this one slipped into the cracks)
I can't get the scsi-debug device to format with 16k lbas for whatever
reason, but I can simulate the xfs parts with my dumb script:
#!/bin/bash -x
# fiddle with scsi-debug and xfs atomics
dev="$(lsscsi | grep scsi_debug | awk '{print $6}')"
mnt="${1:-/mnt/s}"
hw_awu="${2:-16834}"
sectsz="${3:-512}"
align="$4"
umount "${mnt}" ${dev}
rmmod scsi-debug
scsi_debug_args=(dev_size_mb=1000 virtual_gb=4096)
if [ "${hw_awu}" -ge 512 ]; then
scsi_debug_args+=(atomic_wr=1 atomic_wr_max_length=$((hw_awu / 512)))
fi
modprobe scsi-debug "${scsi_debug_args[@]}"
dev="$(lsscsi | grep scsi_debug | awk '{print $6}')"
xfs_io -c 'statx -r -m all' $dev | grep atomic
if [ "${sectsz}" -le 4096 ]; then
blksz=4096
else
blksz="${sectsz}"
fi
mkfs_args=(-f "${dev}" -s size="${sectsz}" -b size="${blksz}")
if [ -n "${align}" ]; then
max_agsize=$((1 << 40)) # bytes
agsize=$(((max_agsize / blksz) - 1)) # blocks
align=$((align / blksz)) # convert to blocks
mkfs_args+=(-d agsize="$(( agsize & ~(align - 1) ))b")
fi
mkfs.xfs "${mkfs_args[@]}"
mount "${dev}" "${mnt}"
touch "${mnt}/fubar"
xfs_io -c 'statx -r -m all' "${mnt}/fubar" | grep atomic
So, with an XFS sector size of 16k and an atomic write max of 32k, we
get:
# /code/t/atomicswap/align.sh '' 32768 16384
umount: /dev/sdg: not mounted.
stat.atomic_write_unit_min = 1024
stat.atomic_write_unit_max = 32768
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 0
meta-data=/dev/sdg isize=512 agcount=4, agsize=67108862 blks
= sectsz=16384 attr=2, projid32bit=1
= crc=1 finobt=1, sparse=1, rmapbt=1
= reflink=1 bigtime=1 inobtcount=1 nrext64=1
= exchange=1 metadir=1
data = bsize=16384 blocks=268435448, imaxpct=5
= sunit=0 swidth=0 blks
naming =version 2 bsize=16384 ascii-ci=0, ftype=1, parent=1
log =internal log bsize=16384 blocks=130432, version=2
= sectsz=16384 sunit=1 blks, lazy-count=1
realtime =none extsz=16384 blocks=0, rtextents=0
= rgcount=0 rgsize=67108864 extents
= zoned=0 start=0 reserved=0
stat.atomic_write_unit_min = 16384
stat.atomic_write_unit_max = 32768
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 32768
Note that the AG size is aligned to the hardware max (32k) and both the
hardware and software awu are set to 32k. Now let's try it again but
forcing the AG size to be aligned to a megabyte:
# /code/t/atomicswap/align.sh '' 32768 16384 1048576
umount: /dev/sdg: not mounted.
stat.atomic_write_unit_min = 1024
stat.atomic_write_unit_max = 32768
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 0
meta-data=/dev/sdg isize=512 agcount=4, agsize=67108800 blks
= sectsz=16384 attr=2, projid32bit=1
= crc=1 finobt=1, sparse=1, rmapbt=1
= reflink=1 bigtime=1 inobtcount=1 nrext64=1
= exchange=1 metadir=1
data = bsize=16384 blocks=268435200, imaxpct=5
= sunit=0 swidth=0 blks
naming =version 2 bsize=16384 ascii-ci=0, ftype=1, parent=1
log =internal log bsize=16384 blocks=130432, version=2
= sectsz=16384 sunit=1 blks, lazy-count=1
realtime =none extsz=16384 blocks=0, rtextents=0
= rgcount=0 rgsize=67108864 extents
= zoned=0 start=0 reserved=0
stat.atomic_write_unit_min = 16384
stat.atomic_write_unit_max = 1048576
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 32768
Now the software atomic write max jumps to 1MB.
So I think I finally (re)understand what's going on here. In the first
case, mkfs detects the 32k hw awu_max and automatically aligns the AGs
to 32k. This is done so that when we feed "align=32k" into the per-AG
allocator, it will (maybe) give us new space that's aligned to 32k
within the AG, which when combined with the AG alignment of 32k means
that the allocation is aligned to 32k on the device. Excellent!
But what you're saying is that you want the *software* awu_max to be
whatever the log will support:
# /code/t/atomicswap/align.sh '' 0 16384
umount: /dev/sdg: not mounted.
stat.atomic_write_unit_min = 0
stat.atomic_write_unit_max = 0
stat.atomic_write_segments_max = 0
stat.atomic_write_unit_max_opt = 0
meta-data=/dev/sdg isize=512 agcount=4, agsize=67108863 blks
= sectsz=16384 attr=2, projid32bit=1
= crc=1 finobt=1, sparse=1, rmapbt=1
= reflink=1 bigtime=1 inobtcount=1 nrext64=1
= exchange=1 metadir=1
data = bsize=16384 blocks=268435452, imaxpct=5
= sunit=0 swidth=0 blks
naming =version 2 bsize=16384 ascii-ci=0, ftype=1, parent=1
log =internal log bsize=16384 blocks=130432, version=2
= sectsz=16384 sunit=1 blks, lazy-count=1
realtime =none extsz=16384 blocks=0, rtextents=0
= rgcount=0 rgsize=67108864 extents
= zoned=0 start=0 reserved=0
stat.atomic_write_unit_min = 16384
stat.atomic_write_unit_max = 67108864
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 0
Which is 64MB. The software fallback doesn't care about allocation
alignment, and advertising the smaller atomic_write_unit_max is a quirk
from the days when we didn't have atomic_write_unit_max_opt, so
constraining it was the only way to prohibit programs from sending
atomic writes that most likely wouldn't turn into hardware atomic
writes.
Ok now that I grasp what you're getting at, I'll go read your v2 patch.
--D
^ permalink raw reply [flat|nested] 19+ messages in thread
* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-25 10:36 [PATCH] xfs: don't limit software atomic writes by the group alignment Pankaj Raghav
2026-09-25 11:44 ` John Garry
@ 2026-10-05 11:02 ` John Garry
2026-10-05 11:02 ` John Garry
2 siblings, 0 replies; 19+ messages in thread
From: John Garry @ 2026-10-05 11:02 UTC (permalink / raw)
To: Pankaj Raghav, cem, linux-xfs
Cc: John Garry, Darrick J . Wong, gost.dev, pankaj.raghav
On 9/25/26 11:36, Pankaj Raghav wrote:
> xfs_calc_group_awu_max() clamps the atomic write unit maximum to the
> greatest power-of-two factor of the group size when the device advertises
> atomic writes, so that allocations can be made naturally aligned for
> REQ_ATOMIC. But that value is the software limit: it is only used when
> reflink is enabled, and out of place writes through the COW fork have no
> alignment requirement.
>
> mkfs caps agsize one block below 1T, so any filesystem with maximum sized
> AGs has an odd agsize. max_pow_of_two_factor() will return 1 for odd
> agsize.
>
> A filesystem on a > 4TB device that advertises 16k atomic writes then reports
> an atomic write unit maximum of a single fsblock and an optimal maximum of 0,
> and rejects any larger RWF_ATOMIC write. That is worse than the same filesystem
> on a device with no atomic write support at all.
>
> Compute the software limit from the group size alone, and apply the
> alignment constraint in xfs_get_atomic_write_max_opt() instead, which is
> what reports the size that can be offloaded to the hardware.
>
> Results on a 8TB device with 16k hardware atomic support with 4k
> blocksize:
>
> Before patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 4096
> stx_atomic_write_unit_max: 4096
> stx_atomic_write_unit_max_opt: 0
> stx_atomic_write_segments_max: 1
>
> After patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 4096
> stx_atomic_write_unit_max: 2097152
> stx_atomic_write_unit_max_opt: 4096
> stx_atomic_write_segments_max: 1
>
> Results on a 8TB device with 16k hardware atomic support with 16k
> blocksize:
>
> Before patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 16384
> stx_atomic_write_unit_max: 16384
> stx_atomic_write_unit_max_opt: 0
> stx_atomic_write_segments_max: 1
>
> After patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 16384
> stx_atomic_write_unit_max: 33554432
> stx_atomic_write_unit_max_opt: 16384
> stx_atomic_write_segments_max: 1
>
> Fixes: 0c438dcc3150 ("xfs: add xfs_calc_atomic_write_unit_max()")
> Assisted-by: LLM
> Signed-off-by: Pankaj Raghav <p.raghav@samsung.com>
> ---
> fs/xfs/xfs_iops.c | 27 ++++++++++++++++++++++-----
> fs/xfs/xfs_mount.c | 35 ++++++++++++++++++++++++-----------
> fs/xfs/xfs_mount.h | 2 ++
> 3 files changed, 48 insertions(+), 16 deletions(-)
>
> diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c
> index d1306e723899..8c8b14f94ede 100644
> --- a/fs/xfs/xfs_iops.c
> +++ b/fs/xfs/xfs_iops.c
> @@ -620,6 +620,12 @@ xfs_get_atomic_write_min(
> return 0;
> }
>
> +static inline enum xfs_group_type
> +xfs_inode_group_type(struct xfs_inode *ip)
> +{
> + return XFS_IS_REALTIME_INODE(ip) ? XG_TYPE_RTG : XG_TYPE_AG;
> +}
would this be better is a common location (so that it could be reused)?
> +
> unsigned int
> xfs_get_atomic_write_max(
> struct xfs_inode *ip)
> @@ -642,19 +648,20 @@ xfs_get_atomic_write_max(
> * then advertise a maximum size of whatever we can complete through
> * that means. Hardware support is reported via max_opt, not here.
> */
> - if (XFS_IS_REALTIME_INODE(ip))
> - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_RTG].awu_max);
> - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_AG].awu_max);
> + return XFS_FSB_TO_B(mp, mp->m_groups[xfs_inode_group_type(ip)].awu_max);
> }
>
> unsigned int
> xfs_get_atomic_write_max_opt(
> struct xfs_inode *ip)
> {
> + struct xfs_mount *mp = ip->i_mount;
> unsigned int awu_max = xfs_get_atomic_write_max(ip);
xfs_get_atomic_write_max() value is calculated based on HW atomic
support. I am wondering if we should add a function to just give the max
CoW-based atomic, and have it called here and from
xfs_get_atomic_write_max(). Not a big deal, though.
> + xfs_extlen_t align_max_fsb;
> + unsigned int opt;
>
> /* if the max is 1x block, then just keep behaviour that opt is 0 */
> - if (awu_max <= ip->i_mount->m_sb.sb_blocksize)
> + if (awu_max <= mp->m_sb.sb_blocksize)
> return 0;
>
> /*
> @@ -663,7 +670,17 @@ xfs_get_atomic_write_max_opt(
> * less than our out of place write limit, but we don't want to exceed
> * the awu_max.
> */
> - return min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
> + opt = min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
> +
> + /*
> + * REQ_ATOMIC writes also have to be naturally aligned on disk, so we
> + * cannot promise more than the largest extent that the allocator is
> + * able to align within a group.
> + */
> + align_max_fsb = xfs_calc_group_awu_align_max(mp,
> + xfs_inode_group_type(ip));
> +
> + return min_t(xfs_fsize_t, opt, XFS_FSB_TO_B(mp, align_max_fsb));
unsigned int? But is there a possibility that the value in
XFS_FSB_TO_B(mp, align_max_fsb) can exceed an unsigned int?
> }
>
> static void
> diff --git a/fs/xfs/xfs_mount.c b/fs/xfs/xfs_mount.c
> index be90c7b03994..b7110b3d023b 100644
> --- a/fs/xfs/xfs_mount.c
> +++ b/fs/xfs/xfs_mount.c
> @@ -674,15 +674,13 @@ static inline xfs_extlen_t xfs_calc_atomic_write_max(struct xfs_mount *mp)
> }
>
> /*
> - * If the underlying device advertises atomic write support, limit the size of
> - * atomic writes to the greatest power-of-two factor of the group size so
> - * that every atomic write unit aligns with the start of every group. This is
> - * required so that the allocations for an atomic write will always be
> - * aligned compatibly with the alignment requirements of the storage.
> + * The largest out of place write is the number of blocks that user files can
> + * allocate from any group.
> *
> - * If the device doesn't advertise atomic writes, then there are no alignment
> - * restrictions and the largest out-of-place write we can do ourselves is the
> - * number of blocks that user files can allocate from any group.
> + * This is a software limit, so it deliberately does not take the alignment
> + * requirements of the storage into account. An atomic write that cannot be
> + * handed to the device as a single naturally aligned REQ_ATOMIC bio is
> + * completed through the COW fork instead, which has no such requirement.
> */
> static xfs_extlen_t
> xfs_calc_group_awu_max(
> @@ -690,15 +688,30 @@ xfs_calc_group_awu_max(
> enum xfs_group_type type)
> {
> struct xfs_groups *g = &mp->m_groups[type];
> - struct xfs_buftarg *btp = xfs_group_type_buftarg(mp, type);
>
> if (g->blocks == 0)
> return 0;
> - if (btp && btp->bt_awu_min > 0)
> - return max_pow_of_two_factor(g->blocks);
> return rounddown_pow_of_two(g->blocks);
> }
>
> +/*
> + * Compute the largest atomic write unit for which the allocator can guarantee
> + * naturally aligned extents in this group type.
> + *
> + * Hardware atomic writes have to be naturally aligned on disk.
> + */
> +xfs_extlen_t
> +xfs_calc_group_awu_align_max(
> + struct xfs_mount *mp,
> + enum xfs_group_type type)
> +{
> + struct xfs_groups *g = &mp->m_groups[type];
> +
> + if (g->blocks == 0)
> + return 0;
> + return max_pow_of_two_factor(g->blocks);
> +}
> +
> /* Compute the maximum atomic write unit size for each section. */
> static inline void
> xfs_calc_atomic_write_unit_max(
> diff --git a/fs/xfs/xfs_mount.h b/fs/xfs/xfs_mount.h
> index 216a38a354e7..a2894c18e3c1 100644
> --- a/fs/xfs/xfs_mount.h
> +++ b/fs/xfs/xfs_mount.h
> @@ -812,6 +812,8 @@ static inline void xfs_mod_sb_delalloc(struct xfs_mount *mp, int64_t delta)
>
> int xfs_set_max_atomic_write_opt(struct xfs_mount *mp,
> unsigned long long new_max_bytes);
> +xfs_extlen_t xfs_calc_group_awu_align_max(struct xfs_mount *mp,
> + enum xfs_group_type type);
>
> static inline struct xfs_buftarg *
> xfs_group_type_buftarg(
>
> base-commit: 1282269a5ddb044481e7f8fd43b3195e211d5475
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-09-25 10:36 [PATCH] xfs: don't limit software atomic writes by the group alignment Pankaj Raghav
2026-09-25 11:44 ` John Garry
2026-10-05 11:02 ` John Garry
@ 2026-10-05 11:02 ` John Garry
2026-10-05 16:24 ` Pankaj Raghav (Samsung)
2 siblings, 1 reply; 19+ messages in thread
From: John Garry @ 2026-10-05 11:02 UTC (permalink / raw)
To: Pankaj Raghav, cem, linux-xfs; +Cc: Darrick J . Wong, gost.dev, pankaj.raghav
On 9/25/26 11:36, Pankaj Raghav wrote:
> xfs_calc_group_awu_max() clamps the atomic write unit maximum to the
> greatest power-of-two factor of the group size when the device advertises
> atomic writes, so that allocations can be made naturally aligned for
> REQ_ATOMIC. But that value is the software limit: it is only used when
> reflink is enabled, and out of place writes through the COW fork have no
> alignment requirement.
>
> mkfs caps agsize one block below 1T, so any filesystem with maximum sized
> AGs has an odd agsize. max_pow_of_two_factor() will return 1 for odd
> agsize.
>
> A filesystem on a > 4TB device that advertises 16k atomic writes then reports
> an atomic write unit maximum of a single fsblock and an optimal maximum of 0,
> and rejects any larger RWF_ATOMIC write. That is worse than the same filesystem
> on a device with no atomic write support at all.
>
> Compute the software limit from the group size alone, and apply the
> alignment constraint in xfs_get_atomic_write_max_opt() instead, which is
> what reports the size that can be offloaded to the hardware.
>
> Results on a 8TB device with 16k hardware atomic support with 4k
> blocksize:
>
> Before patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 4096
> stx_atomic_write_unit_max: 4096
> stx_atomic_write_unit_max_opt: 0
> stx_atomic_write_segments_max: 1
>
> After patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 4096
> stx_atomic_write_unit_max: 2097152
> stx_atomic_write_unit_max_opt: 4096
> stx_atomic_write_segments_max: 1
>
> Results on a 8TB device with 16k hardware atomic support with 16k
> blocksize:
>
> Before patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 16384
> stx_atomic_write_unit_max: 16384
> stx_atomic_write_unit_max_opt: 0
> stx_atomic_write_segments_max: 1
>
> After patches:
> /media/test/hello.txt:
> stx_atomic_write_unit_min: 16384
> stx_atomic_write_unit_max: 33554432
> stx_atomic_write_unit_max_opt: 16384
> stx_atomic_write_segments_max: 1
>
> Fixes: 0c438dcc3150 ("xfs: add xfs_calc_atomic_write_unit_max()")
> Assisted-by: LLM
> Signed-off-by: Pankaj Raghav <p.raghav@samsung.com>
> ---
> fs/xfs/xfs_iops.c | 27 ++++++++++++++++++++++-----
> fs/xfs/xfs_mount.c | 35 ++++++++++++++++++++++++-----------
> fs/xfs/xfs_mount.h | 2 ++
> 3 files changed, 48 insertions(+), 16 deletions(-)
>
> diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c
> index d1306e723899..8c8b14f94ede 100644
> --- a/fs/xfs/xfs_iops.c
> +++ b/fs/xfs/xfs_iops.c
> @@ -620,6 +620,12 @@ xfs_get_atomic_write_min(
> return 0;
> }
>
> +static inline enum xfs_group_type
> +xfs_inode_group_type(struct xfs_inode *ip)
> +{
> + return XFS_IS_REALTIME_INODE(ip) ? XG_TYPE_RTG : XG_TYPE_AG;
> +}
would this be better is a common location (so that it could be reused)?
> +
> unsigned int
> xfs_get_atomic_write_max(
> struct xfs_inode *ip)
> @@ -642,19 +648,20 @@ xfs_get_atomic_write_max(
> * then advertise a maximum size of whatever we can complete through
> * that means. Hardware support is reported via max_opt, not here.
> */
> - if (XFS_IS_REALTIME_INODE(ip))
> - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_RTG].awu_max);
> - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_AG].awu_max);
> + return XFS_FSB_TO_B(mp, mp->m_groups[xfs_inode_group_type(ip)].awu_max);
> }
>
> unsigned int
> xfs_get_atomic_write_max_opt(
> struct xfs_inode *ip)
> {
> + struct xfs_mount *mp = ip->i_mount;
> unsigned int awu_max = xfs_get_atomic_write_max(ip);
xfs_get_atomic_write_max() value is calculated based on HW atomic
support. I am wondering if we should add a function to just give the max
CoW-based atomic, and have it called here and from
xfs_get_atomic_write_max(). Not a big deal, though.
> + xfs_extlen_t align_max_fsb;
> + unsigned int opt;
>
> /* if the max is 1x block, then just keep behaviour that opt is 0 */
> - if (awu_max <= ip->i_mount->m_sb.sb_blocksize)
> + if (awu_max <= mp->m_sb.sb_blocksize)
> return 0;
>
> /*
> @@ -663,7 +670,17 @@ xfs_get_atomic_write_max_opt(
> * less than our out of place write limit, but we don't want to exceed
> * the awu_max.
> */
> - return min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
> + opt = min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
> +
> + /*
> + * REQ_ATOMIC writes also have to be naturally aligned on disk, so we
> + * cannot promise more than the largest extent that the allocator is
> + * able to align within a group.
> + */
> + align_max_fsb = xfs_calc_group_awu_align_max(mp,
> + xfs_inode_group_type(ip));
> +
> + return min_t(xfs_fsize_t, opt, XFS_FSB_TO_B(mp, align_max_fsb));
unsigned int? But is there a possibility that the value in
XFS_FSB_TO_B(mp, align_max_fsb) can exceed an unsigned int?
> }
>
> static void
> diff --git a/fs/xfs/xfs_mount.c b/fs/xfs/xfs_mount.c
> index be90c7b03994..b7110b3d023b 100644
> --- a/fs/xfs/xfs_mount.c
> +++ b/fs/xfs/xfs_mount.c
> @@ -674,15 +674,13 @@ static inline xfs_extlen_t xfs_calc_atomic_write_max(struct xfs_mount *mp)
> }
>
> /*
> - * If the underlying device advertises atomic write support, limit the size of
> - * atomic writes to the greatest power-of-two factor of the group size so
> - * that every atomic write unit aligns with the start of every group. This is
> - * required so that the allocations for an atomic write will always be
> - * aligned compatibly with the alignment requirements of the storage.
> + * The largest out of place write is the number of blocks that user files can
> + * allocate from any group.
> *
> - * If the device doesn't advertise atomic writes, then there are no alignment
> - * restrictions and the largest out-of-place write we can do ourselves is the
> - * number of blocks that user files can allocate from any group.
> + * This is a software limit, so it deliberately does not take the alignment
> + * requirements of the storage into account. An atomic write that cannot be
> + * handed to the device as a single naturally aligned REQ_ATOMIC bio is
> + * completed through the COW fork instead, which has no such requirement.
> */
> static xfs_extlen_t
> xfs_calc_group_awu_max(
> @@ -690,15 +688,30 @@ xfs_calc_group_awu_max(
> enum xfs_group_type type)
> {
> struct xfs_groups *g = &mp->m_groups[type];
> - struct xfs_buftarg *btp = xfs_group_type_buftarg(mp, type);
>
> if (g->blocks == 0)
> return 0;
> - if (btp && btp->bt_awu_min > 0)
> - return max_pow_of_two_factor(g->blocks);
> return rounddown_pow_of_two(g->blocks);
> }
>
> +/*
> + * Compute the largest atomic write unit for which the allocator can guarantee
> + * naturally aligned extents in this group type.
> + *
> + * Hardware atomic writes have to be naturally aligned on disk.
> + */
> +xfs_extlen_t
> +xfs_calc_group_awu_align_max(
> + struct xfs_mount *mp,
> + enum xfs_group_type type)
> +{
> + struct xfs_groups *g = &mp->m_groups[type];
> +
> + if (g->blocks == 0)
> + return 0;
> + return max_pow_of_two_factor(g->blocks);
> +}
> +
> /* Compute the maximum atomic write unit size for each section. */
> static inline void
> xfs_calc_atomic_write_unit_max(
> diff --git a/fs/xfs/xfs_mount.h b/fs/xfs/xfs_mount.h
> index 216a38a354e7..a2894c18e3c1 100644
> --- a/fs/xfs/xfs_mount.h
> +++ b/fs/xfs/xfs_mount.h
> @@ -812,6 +812,8 @@ static inline void xfs_mod_sb_delalloc(struct xfs_mount *mp, int64_t delta)
>
> int xfs_set_max_atomic_write_opt(struct xfs_mount *mp,
> unsigned long long new_max_bytes);
> +xfs_extlen_t xfs_calc_group_awu_align_max(struct xfs_mount *mp,
> + enum xfs_group_type type);
>
> static inline struct xfs_buftarg *
> xfs_group_type_buftarg(
>
> base-commit: 1282269a5ddb044481e7f8fd43b3195e211d5475
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-10-05 11:02 ` John Garry
@ 2026-10-05 16:24 ` Pankaj Raghav (Samsung)
2026-10-06 7:06 ` John Garry
0 siblings, 1 reply; 19+ messages in thread
From: Pankaj Raghav (Samsung) @ 2026-10-05 16:24 UTC (permalink / raw)
To: John Garry; +Cc: Pankaj Raghav, cem, linux-xfs, Darrick J . Wong, gost.dev
> > diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c
> > index d1306e723899..8c8b14f94ede 100644
> > --- a/fs/xfs/xfs_iops.c
> > +++ b/fs/xfs/xfs_iops.c
> > @@ -620,6 +620,12 @@ xfs_get_atomic_write_min(
> > return 0;
> > }
> > +static inline enum xfs_group_type
> > +xfs_inode_group_type(struct xfs_inode *ip)
> > +{
> > + return XFS_IS_REALTIME_INODE(ip) ? XG_TYPE_RTG : XG_TYPE_AG;
> > +}
>
> would this be better is a common location (so that it could be reused)?
>
Probably to xfs_mount.h?
> > +
> > unsigned int
> > xfs_get_atomic_write_max(
> > struct xfs_inode *ip)
> > @@ -642,19 +648,20 @@ xfs_get_atomic_write_max(
> > * then advertise a maximum size of whatever we can complete through
> > * that means. Hardware support is reported via max_opt, not here.
> > */
> > - if (XFS_IS_REALTIME_INODE(ip))
> > - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_RTG].awu_max);
> > - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_AG].awu_max);
> > + return XFS_FSB_TO_B(mp, mp->m_groups[xfs_inode_group_type(ip)].awu_max);
> > }
> > unsigned int
> > xfs_get_atomic_write_max_opt(
> > struct xfs_inode *ip)
> > {
> > + struct xfs_mount *mp = ip->i_mount;
> > unsigned int awu_max = xfs_get_atomic_write_max(ip);
>
> xfs_get_atomic_write_max() value is calculated based on HW atomic support. I
> am wondering if we should add a function to just give the max CoW-based
> atomic, and have it called here and from xfs_get_atomic_write_max(). Not a
> big deal, though.
Could you elaborate this comment?
I do remove any HW dependency in xfs_get_atomic_write_max() calculation
as a part of this patch.
>
> > + xfs_extlen_t align_max_fsb;
> > + unsigned int opt;
> > /* if the max is 1x block, then just keep behaviour that opt is 0 */
> > - if (awu_max <= ip->i_mount->m_sb.sb_blocksize)
> > + if (awu_max <= mp->m_sb.sb_blocksize)
> > return 0;
> > /*
> > @@ -663,7 +670,17 @@ xfs_get_atomic_write_max_opt(
> > * less than our out of place write limit, but we don't want to exceed
> > * the awu_max.
> > */
> > - return min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
> > + opt = min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
> > +
> > + /*
> > + * REQ_ATOMIC writes also have to be naturally aligned on disk, so we
> > + * cannot promise more than the largest extent that the allocator is
> > + * able to align within a group.
> > + */
> > + align_max_fsb = xfs_calc_group_awu_align_max(mp,
> > + xfs_inode_group_type(ip));
> > +
> > + return min_t(xfs_fsize_t, opt, XFS_FSB_TO_B(mp, align_max_fsb));
>
> unsigned int? But is there a possibility that the value in XFS_FSB_TO_B(mp,
> align_max_fsb) can exceed an unsigned int?
We are limited by `opt` length which is an unsigned int. I do use
min_t(xfs_fsize_t, ..) for calculation to avoid any truncation error.
--
Pankaj
^ permalink raw reply [flat|nested] 19+ messages in thread* Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
2026-10-05 16:24 ` Pankaj Raghav (Samsung)
@ 2026-10-06 7:06 ` John Garry
0 siblings, 0 replies; 19+ messages in thread
From: John Garry @ 2026-10-06 7:06 UTC (permalink / raw)
To: Pankaj Raghav (Samsung)
Cc: Pankaj Raghav, cem, linux-xfs, Darrick J . Wong, gost.dev
On 10/5/26 17:24, Pankaj Raghav (Samsung) wrote:
>>> diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c
>>> index d1306e723899..8c8b14f94ede 100644
>>> --- a/fs/xfs/xfs_iops.c
>>> +++ b/fs/xfs/xfs_iops.c
>>> @@ -620,6 +620,12 @@ xfs_get_atomic_write_min(
>>> return 0;
>>> }
>>> +static inline enum xfs_group_type
>>> +xfs_inode_group_type(struct xfs_inode *ip)
>>> +{
>>> + return XFS_IS_REALTIME_INODE(ip) ? XG_TYPE_RTG : XG_TYPE_AG;
>>> +}
>>
>> would this be better is a common location (so that it could be reused)?
>>
>
> Probably to xfs_mount.h?
>
I'm not sure. Darrick may be able to give a good suggestion.
>>> +
>>> unsigned int
>>> xfs_get_atomic_write_max(
>>> struct xfs_inode *ip)
>>> @@ -642,19 +648,20 @@ xfs_get_atomic_write_max(
>>> * then advertise a maximum size of whatever we can complete through
>>> * that means. Hardware support is reported via max_opt, not here.
>>> */
>>> - if (XFS_IS_REALTIME_INODE(ip))
>>> - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_RTG].awu_max);
>>> - return XFS_FSB_TO_B(mp, mp->m_groups[XG_TYPE_AG].awu_max);
>>> + return XFS_FSB_TO_B(mp, mp->m_groups[xfs_inode_group_type(ip)].awu_max);
>>> }
>>> unsigned int
>>> xfs_get_atomic_write_max_opt(
>>> struct xfs_inode *ip)
>>> {
>>> + struct xfs_mount *mp = ip->i_mount;
>>> unsigned int awu_max = xfs_get_atomic_write_max(ip);
>>
>> xfs_get_atomic_write_max() value is calculated based on HW atomic support. I
>> am wondering if we should add a function to just give the max CoW-based
>> atomic, and have it called here and from xfs_get_atomic_write_max(). Not a
>> big deal, though.
>
> Could you elaborate this comment?
>
> I do remove any HW dependency in xfs_get_atomic_write_max() calculation
> as a part of this patch.
xfs_get_atomic_write_max() does still have a
xfs_inode_can_hw_atomic_write() call.
It just seems a bit awkward that xfs_get_atomic_write_max_opt() calls
xfs_get_atomic_write_max(), when it should be able to do the full
calculation itself.
This is not a deal deal which I am mentioning.
>
>>
>>> + xfs_extlen_t align_max_fsb;
>>> + unsigned int opt;
>>> /* if the max is 1x block, then just keep behaviour that opt is 0 */
>>> - if (awu_max <= ip->i_mount->m_sb.sb_blocksize)
>>> + if (awu_max <= mp->m_sb.sb_blocksize)
>>> return 0;
>>> /*
>>> @@ -663,7 +670,17 @@ xfs_get_atomic_write_max_opt(
>>> * less than our out of place write limit, but we don't want to exceed
>>> * the awu_max.
>>> */
>>> - return min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
>>> + opt = min(awu_max, xfs_inode_buftarg(ip)->bt_awu_max);
>>> +
>>> + /*
>>> + * REQ_ATOMIC writes also have to be naturally aligned on disk, so we
>>> + * cannot promise more than the largest extent that the allocator is
>>> + * able to align within a group.
>>> + */
>>> + align_max_fsb = xfs_calc_group_awu_align_max(mp,
>>> + xfs_inode_group_type(ip));
>>> +
>>> + return min_t(xfs_fsize_t, opt, XFS_FSB_TO_B(mp, align_max_fsb));
>>
>> unsigned int? But is there a possibility that the value in XFS_FSB_TO_B(mp,
>> align_max_fsb) can exceed an unsigned int?
>
> We are limited by `opt` length which is an unsigned int. I do use
> min_t(xfs_fsize_t, ..) for calculation to avoid any truncation error.
>
ok, fine
BTW, maybe call the variable max_opt, and not just opt.
^ permalink raw reply [flat|nested] 19+ messages in thread