Linux XFS filesystem development
 help / color / mirror / Atom feed
From: "Darrick J. Wong" <djwong@kernel.org>
To: "Pankaj Raghav (Samsung)" <pankaj.raghav@linux.dev>
Cc: john.garry@linux.dev, Pankaj Raghav <p.raghav@samsung.com>,
	cem@kernel.org, linux-xfs@vger.kernel.org,
	John Garry <john.g.garry@oracle.com>,
	gost.dev@samsung.com
Subject: Re: [PATCH] xfs: don't limit software atomic writes by the group alignment
Date: Wed, 7 Oct 2026 10:43:30 -0700	[thread overview]
Message-ID: <20261007174330.GK2705364@frogsfrogsfrogs> (raw)
In-Reply-To: <m72iphraf7d3hoziw55inrymk6ynsgl3zio7nyiysm4w2fkac4@jtcaqpvg7g4a>

On Wed, Sep 30, 2026 at 08:48:56AM +0000, Pankaj Raghav (Samsung) wrote:
> > > But I guess I'll play around with it for a while and see if I get
> > > anywhere.  Assuming you're running a recent mkfs that knows about
> > > the atomic write options and whatnot?
> > 
> > FWIW I tried this with scsi-debug and got the following:
> > 
> > # umount /dev/sdg
> > # rmmod scsi-debug
> > # modprobe scsi-debug dev_size_mb=1000 virtual_gb=4096 atomic_wr=1 atomic_wr_max_length=32
> > # xfs_io -c 'statx -r -m all ' /dev/sdg | grep atomic
> > stat.atomic_write_unit_min = 1024
> > stat.atomic_write_unit_max = 16384
> > stat.atomic_write_segments_max = 1
> > stat.atomic_write_unit_max_opt = 0
> > # mkfs.xfs -f /dev/sdg
> > meta-data=/dev/sdg               isize=512    agcount=4, agsize=268435452 blks
> >          =                       sectsz=512   attr=2, projid32bit=1
> >          =                       crc=1        finobt=1, sparse=1, rmapbt=1
> >          =                       reflink=1    bigtime=1 inobtcount=1 nrext64=1
> >          =                       exchange=1   metadir=0
> > data     =                       bsize=4096   blocks=1073741808, imaxpct=5
> >          =                       sunit=0      swidth=0 blks
> > naming   =version 2              bsize=4096   ascii-ci=0, ftype=1, parent=1
> > log      =internal log           bsize=4096   blocks=521728, version=2
> >          =                       sectsz=512   sunit=0 blks, lazy-count=1
> > realtime =none                   extsz=4096   blocks=0, rtextents=0
> >          =                       rgcount=0    rgsize=0 extents
> >          =                       zoned=0      start=0 reserved=0
> > # mount /dev/sdg /mnt
> > # touch /mnt/fubar
> > # xfs_io -c 'statx -r -m all' /mnt/fubar | grep atomic
> > stat.atomic_write_unit_min = 4096
> > stat.atomic_write_unit_max = 16384
> > stat.atomic_write_segments_max = 1
> > stat.atomic_write_unit_max_opt = 16384
> > 
> > Note the agsize=268435452, which means that it rounded the AG size down
> > to something congruent with the 16k hardware atomic write max.
> > 
> > (Are you sure you're using mkfs.xfs from xfsprogs 6.16 or newer?)
> 
> I noticed the issue when I was running a test on a bare metal server
> with xfsprog 7.1.1. But I tried to recreate the issue in a VM where I
> was running an old xfsprog. So the values I have posted here are wrong.
> Sorry for that.
> 
> > # xfs_io -c 'statx -r -m all' /mnt/fubar | grep atomic
> > stat.atomic_write_unit_min = 4096
> > stat.atomic_write_unit_max = 16384
> > stat.atomic_write_segments_max = 1
> > stat.atomic_write_unit_max_opt = 16384
> 
> However, I still do not understand why do we have to restrict
> atomic_write_unit_max (SW based atomics) to 16k here?
> 
> Let me take a step back and give a history on why I started to look into
> this.
> 
> I am running 7.2-rc6 with 7.1.1 xfsprogs on a bare metal server with a
> NVMe device that supports atomics.
> 
> I am formatting XFS with 16k blocksize.
> 
> panky@blixen:~$ df -h /dev/nvme1n1
> Filesystem      Size  Used Avail Use% Mounted on
> /dev/nvme1n1     14T  179G   14T   2% /mnt/atomssd
> 
> panky@blixen:~$ cat /sys/block/nvme1n1/queue/atomic_write_unit_max_bytes
> 16384
> 
> panky@blixen:~$ xfs_info /mnt/atomssd/
> meta-data=/dev/nvme1n1           isize=512    agcount=14, agsize=67108863 blks
>          =                       sectsz=4096  attr=2, projid32bit=1
>          =                       crc=1        finobt=1, sparse=1, rmapbt=1
>          =                       reflink=1    bigtime=1 inobtcount=1 nrext64=1
>          =                       exchange=1   metadir=0
> data     =                       bsize=16384  blocks=937558016, imaxpct=5
>          =                       sunit=0      swidth=0 blks
> naming   =version 2              bsize=16384  ascii-ci=0, ftype=1, parent=1
> log      =internal log           bsize=16384  blocks=130432, version=2
>          =                       sectsz=4096  sunit=1 blks, lazy-count=1
> realtime =none                   extsz=16384  blocks=0, rtextents=0
>          =                       rgcount=0    rgsize=0 extents
>          =                       zoned=0      start=0 reserved=0
> 
> panky@blixen:~/tools$ sudo ./statx /mnt/atomssd/hello.txt
> [sudo: authenticate] Password:
> /mnt/atomssd/hello.txt:
>   stx_atomic_write_unit_min:            16384
>   stx_atomic_write_unit_max:            16384
>   stx_atomic_write_unit_max_opt:        0
>   stx_atomic_write_segments_max:        1
> 
> We expose max_opt to be 0, which means the application can assume
> max_opt to be the same as atomic_write_unit_max. But why are we not
> exposing bigger atomic_write_unit_max (SW atomics) here and make max_opt
> to HW based atomics value (16k) in this case?
> 
> Could you format your scsi-debug device with 16k fsblock size and see
> what values you are getting here?

(Apologies, this one slipped into the cracks)

I can't get the scsi-debug device to format with 16k lbas for whatever
reason, but I can simulate the xfs parts with my dumb script:

#!/bin/bash -x

# fiddle with scsi-debug and xfs atomics

dev="$(lsscsi | grep scsi_debug | awk '{print $6}')"
mnt="${1:-/mnt/s}"
hw_awu="${2:-16834}"
sectsz="${3:-512}"
align="$4"

umount "${mnt}" ${dev}
rmmod scsi-debug

scsi_debug_args=(dev_size_mb=1000 virtual_gb=4096)
if [ "${hw_awu}" -ge 512 ]; then
	scsi_debug_args+=(atomic_wr=1 atomic_wr_max_length=$((hw_awu / 512)))
fi
modprobe scsi-debug "${scsi_debug_args[@]}"
dev="$(lsscsi | grep scsi_debug | awk '{print $6}')"
xfs_io -c 'statx -r -m all' $dev | grep atomic

if [ "${sectsz}" -le 4096 ]; then
	blksz=4096
else
	blksz="${sectsz}"
fi

mkfs_args=(-f "${dev}" -s size="${sectsz}" -b size="${blksz}")
if [ -n "${align}" ]; then
	max_agsize=$((1 << 40))		# bytes
	agsize=$(((max_agsize / blksz) - 1))	# blocks
	align=$((align / blksz))	# convert to blocks

	mkfs_args+=(-d agsize="$(( agsize & ~(align - 1) ))b")
fi
mkfs.xfs "${mkfs_args[@]}"
mount "${dev}" "${mnt}"

touch "${mnt}/fubar"
xfs_io -c 'statx -r -m all' "${mnt}/fubar" | grep atomic

So, with an XFS sector size of 16k and an atomic write max of 32k, we
get:

# /code/t/atomicswap/align.sh '' 32768 16384
umount: /dev/sdg: not mounted.
stat.atomic_write_unit_min = 1024
stat.atomic_write_unit_max = 32768
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 0
meta-data=/dev/sdg               isize=512    agcount=4, agsize=67108862 blks
         =                       sectsz=16384 attr=2, projid32bit=1
         =                       crc=1        finobt=1, sparse=1, rmapbt=1
         =                       reflink=1    bigtime=1 inobtcount=1 nrext64=1
         =                       exchange=1   metadir=1
data     =                       bsize=16384  blocks=268435448, imaxpct=5
         =                       sunit=0      swidth=0 blks
naming   =version 2              bsize=16384  ascii-ci=0, ftype=1, parent=1
log      =internal log           bsize=16384  blocks=130432, version=2
         =                       sectsz=16384 sunit=1 blks, lazy-count=1
realtime =none                   extsz=16384  blocks=0, rtextents=0
         =                       rgcount=0    rgsize=67108864 extents
         =                       zoned=0      start=0 reserved=0
stat.atomic_write_unit_min = 16384
stat.atomic_write_unit_max = 32768
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 32768

Note that the AG size is aligned to the hardware max (32k) and both the
hardware and software awu are set to 32k.  Now let's try it again but
forcing the AG size to be aligned to a megabyte:

# /code/t/atomicswap/align.sh '' 32768 16384 1048576
umount: /dev/sdg: not mounted.
stat.atomic_write_unit_min = 1024
stat.atomic_write_unit_max = 32768
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 0
meta-data=/dev/sdg               isize=512    agcount=4, agsize=67108800 blks
         =                       sectsz=16384 attr=2, projid32bit=1
         =                       crc=1        finobt=1, sparse=1, rmapbt=1
         =                       reflink=1    bigtime=1 inobtcount=1 nrext64=1
         =                       exchange=1   metadir=1
data     =                       bsize=16384  blocks=268435200, imaxpct=5
         =                       sunit=0      swidth=0 blks
naming   =version 2              bsize=16384  ascii-ci=0, ftype=1, parent=1
log      =internal log           bsize=16384  blocks=130432, version=2
         =                       sectsz=16384 sunit=1 blks, lazy-count=1
realtime =none                   extsz=16384  blocks=0, rtextents=0
         =                       rgcount=0    rgsize=67108864 extents
         =                       zoned=0      start=0 reserved=0
stat.atomic_write_unit_min = 16384
stat.atomic_write_unit_max = 1048576
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 32768

Now the software atomic write max jumps to 1MB.

So I think I finally (re)understand what's going on here.  In the first
case, mkfs detects the 32k hw awu_max and automatically aligns the AGs
to 32k.  This is done so that when we feed "align=32k" into the per-AG
allocator, it will (maybe) give us new space that's aligned to 32k
within the AG, which when combined with the AG alignment of 32k means
that the allocation is aligned to 32k on the device.  Excellent!

But what you're saying is that you want the *software* awu_max to be
whatever the log will support:

# /code/t/atomicswap/align.sh '' 0 16384
umount: /dev/sdg: not mounted.
stat.atomic_write_unit_min = 0
stat.atomic_write_unit_max = 0
stat.atomic_write_segments_max = 0
stat.atomic_write_unit_max_opt = 0
meta-data=/dev/sdg               isize=512    agcount=4, agsize=67108863 blks
         =                       sectsz=16384 attr=2, projid32bit=1
         =                       crc=1        finobt=1, sparse=1, rmapbt=1
         =                       reflink=1    bigtime=1 inobtcount=1 nrext64=1
         =                       exchange=1   metadir=1
data     =                       bsize=16384  blocks=268435452, imaxpct=5
         =                       sunit=0      swidth=0 blks
naming   =version 2              bsize=16384  ascii-ci=0, ftype=1, parent=1
log      =internal log           bsize=16384  blocks=130432, version=2
         =                       sectsz=16384 sunit=1 blks, lazy-count=1
realtime =none                   extsz=16384  blocks=0, rtextents=0
         =                       rgcount=0    rgsize=67108864 extents
         =                       zoned=0      start=0 reserved=0
stat.atomic_write_unit_min = 16384
stat.atomic_write_unit_max = 67108864
stat.atomic_write_segments_max = 1
stat.atomic_write_unit_max_opt = 0

Which is 64MB.  The software fallback doesn't care about allocation
alignment, and advertising the smaller atomic_write_unit_max is a quirk
from the days when we didn't have atomic_write_unit_max_opt, so
constraining it was the only way to prohibit programs from sending
atomic writes that most likely wouldn't turn into hardware atomic
writes.

Ok now that I grasp what you're getting at, I'll go read your v2 patch.

--D

  reply	other threads:[~2026-10-07 17:43 UTC|newest]

Thread overview: 19+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25 10:36 [PATCH] xfs: don't limit software atomic writes by the group alignment Pankaj Raghav
2026-09-25 11:44 ` John Garry
2026-09-29  7:17   ` Pankaj Raghav
2026-09-29 14:47     ` John Garry
2026-09-30  5:10       ` Pankaj Raghav (Samsung)
2026-09-30  8:42         ` John Garry
2026-09-30 12:27           ` Pankaj Raghav
2026-09-30 12:48             ` John Garry
2026-10-01 16:30               ` Pankaj Raghav
2026-10-02  8:45                 ` John Garry
2026-10-02 11:24                   ` Pankaj Raghav
2026-09-29 16:40     ` Darrick J. Wong
2026-09-29 20:42       ` Darrick J. Wong
2026-09-30  8:48         ` Pankaj Raghav (Samsung)
2026-10-07 17:43           ` Darrick J. Wong [this message]
2026-10-05 11:02 ` John Garry
2026-10-05 11:02 ` John Garry
2026-10-05 16:24   ` Pankaj Raghav (Samsung)
2026-10-06  7:06     ` John Garry

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261007174330.GK2705364@frogsfrogsfrogs \
    --to=djwong@kernel.org \
    --cc=cem@kernel.org \
    --cc=gost.dev@samsung.com \
    --cc=john.g.garry@oracle.com \
    --cc=john.garry@linux.dev \
    --cc=linux-xfs@vger.kernel.org \
    --cc=p.raghav@samsung.com \
    --cc=pankaj.raghav@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox