From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0E24C347FC0 for ; Tue, 29 Sep 2026 20:42:29 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790714551; cv=none; b=bi8+B4MLdsKW8Bbz01ifenR6ulJGXcRt2XizuXLHCFl8ASrmHQV+tVlYGU3BjyEHNkEA+BJwSyeh0RSt8qFUmz4HLHm5TpzEby3DhPkK/1d/9fyyi76NgoybFHh7oTX7w2DDA5hT2yj4MiWDY+19PDSCWm8/6sqtYtzGyYZ/OLI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790714551; c=relaxed/simple; bh=W+BOURyeOdUWXOcPIkdiEwofu/5PKwricI50+8vovPM=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=bsmITYj1CwLo6mSTo2uKu/KO51B9ntZMG5xsEFm606KRW77gNSkW6i/ycgkMLwB0xcjgRPa81CFwL3CA80ka1n3jZOGDHKj0q1MdkWS4eLdpnmhhDo/xfftxP0fBBu6OC6DwkZlDDSpHzssVx4m5JQgZXbDCtdbfr5yN6sFu8Uk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Mq3hrIQz; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Mq3hrIQz" Received: by smtp.kernel.org (Postfix) with UTF8SMTPSA id 793241F000FF; Tue, 29 Sep 2026 20:42:29 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1790714549; bh=W1JL/YdiU+W1ibu3y4DywbuNoGsPlxhUG6OKDeYAwj4=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=Mq3hrIQz09pHBnywnQFY5bJ2tbbKAIk6nYZCmEgbGoRpmRMWnFrLKmJ83fqbg1JGv uKARO0HWnD1snAAbeUjsEfNeHJuwbMHdJSuCVWF6bKGBRhybzMi0lO796USP5TtUTi xcd3NGimrUCmElJgpYuWtNiteA40XtVTIW9ZpdpSROQCeSCTgDnfoaJncLGk5FGJj0 tdLRB8NS98GY4CCTyFqv6MRJTScncoBO6XlhikJTXflC0L/FecxnF9hEA247C9OXEx rui6mtMLb8DdRX3Hh5Yf+uW5CsS4Ce08GQCL9NQiUTeHZyKrF4Sv5/AYjK3nU3cRU6 St3MPccZYPNUw== Date: Tue, 29 Sep 2026 13:42:28 -0700 From: "Darrick J. Wong" To: Pankaj Raghav Cc: John Garry , Pankaj Raghav , cem@kernel.org, linux-xfs@vger.kernel.org, John Garry , gost.dev@samsung.com Subject: Re: [PATCH] xfs: don't limit software atomic writes by the group alignment Message-ID: <20260929204228.GT2705364@frogsfrogsfrogs> References: <20260925103640.932735-1-p.raghav@samsung.com> <41749fa7-dcfa-46f0-af0f-ed3fa5f39f15@linux.dev> <20260929164010.GS2705364@frogsfrogsfrogs> Precedence: bulk X-Mailing-List: linux-xfs@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=iso-8859-1 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <20260929164010.GS2705364@frogsfrogsfrogs> On Tue, Sep 29, 2026 at 09:40:10AM -0700, Darrick J. Wong wrote: > On Tue, Sep 29, 2026 at 09:17:32AM +0200, Pankaj Raghav wrote: > > > > > > On 9/25/2026 1:44 PM, John Garry wrote: > > > On 9/25/26 11:36, Pankaj Raghav wrote: > > > > xfs_calc_group_awu_max() clamps the atomic write unit maximum to the > > > > greatest power-of-two factor of the group size when the device advertises > > > > atomic writes, so that allocations can be made naturally aligned for > > > > REQ_ATOMIC. But that value is the software limit: it is only used when > > > > reflink is enabled, and out of place writes through the COW fork have no > > > > alignment requirement. > > > > > > We have mkfs atomic write options for selecting atomic write limits - > > > you have tried that, right? > > > > > > > Hmm, I did not try that but I restricted the agsize to a power of 2 value to > > workaround the limitation. > > > > Are you talking about this option: max_atomic_write? > > > > In any case, with default options the atomic values we are exposing is not > > correct at the moment. > > > > > > > > > > A filesystem on a > 4TB device that advertises 16k atomic writes then reports > > > > an atomic write unit maximum of a single fsblock and an optimal maximum of 0, > > > > and rejects any larger RWF_ATOMIC write. That is worse than the same filesystem > > > > on a device with no atomic write support at all. > > > > > > > > Compute the software limit from the group size alone, and apply the > > > > alignment constraint in xfs_get_atomic_write_max_opt() instead, which is > > > > what reports the size that can be offloaded to the hardware. > > > > > > > > Results on a 8TB device with 16k hardware atomic support with 4k > > > > blocksize: > > > > > > > > Before patches: > > > > /media/test/hello.txt: > > > >    stx_atomic_write_unit_min:            4096 > > > >    stx_atomic_write_unit_max:            4096 > > > >    stx_atomic_write_unit_max_opt:        0 > > > >    stx_atomic_write_segments_max:        1 > > > > > > > > After patches: > > > > /media/test/hello.txt: > > > >    stx_atomic_write_unit_min:            4096 > > > >    stx_atomic_write_unit_max:            2097152 > > > > > > Here the AG size may not be a multiple of 2097152B, right? If so, when > > > we try to naturally align the allocation for a CoW write, even if the AG > > > blocks are naturally aligned (per AG), the disk blocks may not be > > > naturnally aligned - is this correct? For HW-based atomic writes, we > > > want naturally aligned disk blocks. > > > > > > > Correct. > > > > Just to give more context: > > > > Trace of calc_atomic_write_unit_max: > > > > mount-696 [010] 449.230543: xfs_calc_atomic_write_unit_max: dev 259:0 > > ag max_write 262144 max_ioend 512 max_gsize 134217728 awu_max 512 > > > > 2097152 comes from awumax being 512 fsblocks. > > > > root@debian:~# xfs_info /dev/nvme0n1 > > meta-data=/dev/nvme0n1 isize=512 agcount=8, agsize=268435455 blks > > 268435455? That's not aligned to the atomic write size, but that is the > max AG size. I thought mkfs would round it down for us automatically: > > /* > * We've already validated (or discarded) the hardware atomic write > * geometry. Try to align the agsize to the maximum atomic write unit > * to give users maximum flexibility in choosing atomic write sizes. > */ > if (ft->data.awu_max > 0) > dsunit = max(DTOBT(ft->data.awu_max, cfg->blocklog), > dsunit); > > But I guess I'll play around with it for a while and see if I get > anywhere. Assuming you're running a recent mkfs that knows about > the atomic write options and whatnot? FWIW I tried this with scsi-debug and got the following: # umount /dev/sdg # rmmod scsi-debug # modprobe scsi-debug dev_size_mb=1000 virtual_gb=4096 atomic_wr=1 atomic_wr_max_length=32 # xfs_io -c 'statx -r -m all ' /dev/sdg | grep atomic stat.atomic_write_unit_min = 1024 stat.atomic_write_unit_max = 16384 stat.atomic_write_segments_max = 1 stat.atomic_write_unit_max_opt = 0 # mkfs.xfs -f /dev/sdg meta-data=/dev/sdg isize=512 agcount=4, agsize=268435452 blks = sectsz=512 attr=2, projid32bit=1 = crc=1 finobt=1, sparse=1, rmapbt=1 = reflink=1 bigtime=1 inobtcount=1 nrext64=1 = exchange=1 metadir=0 data = bsize=4096 blocks=1073741808, imaxpct=5 = sunit=0 swidth=0 blks naming =version 2 bsize=4096 ascii-ci=0, ftype=1, parent=1 log =internal log bsize=4096 blocks=521728, version=2 = sectsz=512 sunit=0 blks, lazy-count=1 realtime =none extsz=4096 blocks=0, rtextents=0 = rgcount=0 rgsize=0 extents = zoned=0 start=0 reserved=0 # mount /dev/sdg /mnt # touch /mnt/fubar # xfs_io -c 'statx -r -m all' /mnt/fubar | grep atomic stat.atomic_write_unit_min = 4096 stat.atomic_write_unit_max = 16384 stat.atomic_write_segments_max = 1 stat.atomic_write_unit_max_opt = 16384 Note the agsize=268435452, which means that it rounded the AG size down to something congruent with the 16k hardware atomic write max. (Are you sure you're using mkfs.xfs from xfsprogs 6.16 or newer?) > --D > > > = sectsz=4096 attr=2, projid32bit=1 > > = crc=1 finobt=1, sparse=1, rmapbt=0 > > = reflink=1 bigtime=1 inobtcount=1 nrext64=0 > > data = bsize=4096 blocks=2147483640, imaxpct=5 > > = sunit=0 swidth=0 blks > > naming =version 2 bsize=4096 ascii-ci=0, ftype=1 > > log =internal log bsize=4096 blocks=521728, version=2 > > = sectsz=4096 sunit=1 blks, lazy-count=1 > > realtime =none extsz=4096 blocks=0, rtextents=0 > > > > > > > >    stx_atomic_write_unit_max_opt:        4096 > > > >    stx_atomic_write_segments_max:        1 > > > > > > > > Results on a 8TB device with 16k hardware atomic support with 16k > > > > blocksize: > > > > > > > > Before patches: > > > > /media/test/hello.txt: > > > >    stx_atomic_write_unit_min:            16384 > > > >    stx_atomic_write_unit_max:            16384 > > > >    stx_atomic_write_unit_max_opt:        0 > > > >    stx_atomic_write_segments_max:        1 > > > > > > > > After patches: > > > > /media/test/hello.txt: > > > >    stx_atomic_write_unit_min:            16384 > > > >    stx_atomic_write_unit_max:            33554432 > > > >    stx_atomic_write_unit_max_opt:        16384 > > > >    stx_atomic_write_segments_max:        1 > > > > > > -- > > Pankaj > > >