Linux Btrfs filesystem development
 help / color / mirror / Atom feed
From: Qu Wenruo <quwenruo.btrfs@gmx.com>
To: Boris Burkov <boris@bur.io>, Qu Wenruo <wqu@suse.com>
Cc: linux-btrfs@vger.kernel.org
Subject: Re: [PATCH v4 4/6] btrfs: introduce compression for delayed bbio
Date: Fri, 25 Sep 2026 07:44:43 +0930	[thread overview]
Message-ID: <7aeb8a6a-471f-4051-a1ab-b5741ad27b0f@gmx.com> (raw)
In-Reply-To: <20260924173029.GB2146908@zen.localdomain>



在 2026/9/25 03:00, Boris Burkov 写道:
> On Fri, Sep 18, 2026 at 09:00:29AM +0930, Qu Wenruo wrote:
>> The compressed write path inside a delayed bbio is mostly the same as
>> regular compression, but with some differences:
>>
>> - The error handling should not touch folio flags
>>    It will be handled by the parent delayed bbio.
>>    And those folios already have WRITEBACK flag set, not the LOCKED flag
>>    of the async submission path.
>>
>> - A successful compression will lead to a child compressed bio
>>    That compressed bio will be properly submitted, and if there are no
>>    more pending IOs of the delayed bbio, end the delayed bbio.
>>
>>    There is a minor note, since we're going through the regular
>>    extent_writepage_io() path, we can have multiple bbios for the same
>>    delayed ordered extent.
>>
>>    This means we may have a slightly lower compression ratio if for
>>    whatever reason the writeback path chooses to submit a smaller bio.
>>
>> - No sequential execution of data extent reservation
>>    The existing async thread has one quirk related to the ordered
>>    function execution, which is not suitable for this call site.
>>
>>    After the compressed bio is submitted, we can no longer touch the
>>    child compressed bio (it could finish immediately and also finish the
>>    parent delayed bbio).
>>    Meanwhile the async ordered function needs different entries to handle
>>    the workload and free involved structures.
>>
>>    These will be the major changes compared to the existing compressed
>>    write.
>>
>> - No extra blkcg/write flag inheritance
>>    This is a missing feature compared to the older async compression.
>>    This needs to be implemented in the future, but priority shouldn't be
>>    that high.
> 
> I believe there is reservation double counting, detailed inline.
> 
>>
>> Signed-off-by: Qu Wenruo <wqu@suse.com>
>> ---
>>   fs/btrfs/inode.c | 114 ++++++++++++++++++++++++++++++++++++++++++++++-
>>   1 file changed, 113 insertions(+), 1 deletion(-)
>>
>> diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
>> index 2d6a581d977b..fbc0e5426d45 100644
>> --- a/fs/btrfs/inode.c
>> +++ b/fs/btrfs/inode.c
>> @@ -7696,14 +7696,126 @@ struct extent_map *btrfs_create_delayed_em(struct btrfs_inode *inode,
>>   	return em;
>>   }
>>   
>> +static void end_bbio_delayed_compressed(struct btrfs_bio *bbio)
>> +{
>> +	struct delayed_bio_private *dbp = bbio->private;
>> +	struct btrfs_bio *parent = dbp->delayed_bbio;
>> +	struct folio_iter fi;
>> +
>> +	bio_for_each_folio_all(fi, &bbio->bio)
>> +		btrfs_free_compr_folio(fi.folio);
>> +	btrfs_bio_end_io(parent, bbio->bio.bi_status);
>> +	bio_put(&bbio->bio);
>> +}
>> +
>> +static bool try_submit_compressed(struct btrfs_bio *parent)
>> +{
>> +	struct delayed_bio_private *dbp = parent->private;
>> +	struct btrfs_inode *inode = parent->inode;
>> +	struct btrfs_fs_info *fs_info = inode->root->fs_info;
>> +	struct btrfs_key ins;
>> +	struct compressed_bio *cb;
>> +	struct extent_state *cached = NULL;
>> +	struct extent_map *em;
>> +	struct btrfs_ordered_extent *ordered;
>> +	struct btrfs_file_extent file_extent;
>> +	u64 alloc_hint;
>> +	const u32 len = bio_get_size(&parent->bio);
>> +	const u64 fileoff = parent->file_offset;
>> +	const u64 end = fileoff + len - 1;
>> +	u32 compressed_size;
>> +	int compress_type = fs_info->compress_type;
>> +	int compress_level = fs_info->compress_level;
>> +	int ret;
>> +
>> +	if (!btrfs_inode_can_compress(inode) ||
>> +	    !inode_need_compress(inode, fileoff, end, false))
>> +		return false;
>> +
>> +	if (inode->defrag_compress > 0 &&
>> +	    inode->defrag_compress < BTRFS_NR_COMPRESS_TYPES) {
>> +		compress_type = inode->defrag_compress;
>> +		compress_level = inode->defrag_compress_level;
>> +	} else if (inode->prop_compress) {
>> +		compress_type = inode->prop_compress;
>> +	}
>> +	cb = btrfs_compress_bio(inode, fileoff, len, compress_type,
>> +				compress_level, 0);
>> +	if (IS_ERR(cb))
>> +		return false;
>> +
>> +	round_up_last_block(cb, fs_info->sectorsize);
>> +	compressed_size = cb->bbio.bio.bi_iter.bi_size;
>> +	/* If no space is saved, abort and mark the inode NOCOMPRESS. */
>> +	if (compressed_size >= len) {
>> +		if (!btrfs_test_opt(fs_info, FORCE_COMPRESS) &&
>> +		    !inode->prop_compress)
>> +			inode->flags |= BTRFS_INODE_NOCOMPRESS;
>> +		cleanup_compressed_bio(cb);
>> +		return false;
>> +	}
>> +
>> +	alloc_hint = btrfs_get_extent_allocation_hint(inode, fileoff, len);
>> +	ret = btrfs_reserve_extent(inode->root, len,
>> +				   compressed_size, compressed_size,
>> +				   0, alloc_hint, &ins, true, true);
> 
> This consumes bytes_may_use and moves it to bytes_reserved
> 
>> +	if (ret < 0) {
>> +		cleanup_compressed_bio(cb);
>> +		return false;
>> +	}
>> +	btrfs_lock_extent(&inode->io_tree, fileoff, end, &cached);
>> +	file_extent.disk_bytenr = ins.objectid;
>> +	file_extent.disk_num_bytes = ins.offset;
>> +	file_extent.ram_bytes = len;
>> +	file_extent.num_bytes = len;
>> +	file_extent.offset = 0;
>> +	file_extent.compression = cb->compress_type;
>> +
>> +	cb->bbio.bio.bi_iter.bi_sector = ins.objectid >> SECTOR_SHIFT;
>> +	em = btrfs_create_io_em(inode, fileoff, &file_extent, BTRFS_ORDERED_COMPRESSED);
>> +	if (IS_ERR(em)) {
>> +		ret = PTR_ERR(em);
>> +		goto out_free_reserve;
>> +	}
>> +	btrfs_free_extent_map(em);
>> +
>> +	ordered = btrfs_alloc_ordered_extent(inode, fileoff, &file_extent,
>> +					     1U << BTRFS_ORDERED_COMPRESSED);
>> +	if (IS_ERR(ordered)) {
>> +		btrfs_drop_extent_map_range(inode, fileoff, end, false);
>> +		ret = PTR_ERR(ordered);
> 
> Now this fails with ENOMEM
> 
>> +		goto out_free_reserve;
>> +	}
>> +	cb->bbio.ordered = ordered;
>> +	btrfs_dec_block_group_reservations(fs_info, ins.objectid);
>> +	btrfs_unlock_extent(&inode->io_tree, fileoff, end, &cached);
>> +
>> +	cb->bbio.end_io = end_bbio_delayed_compressed;
>> +	cb->bbio.private = dbp;
>> +	atomic_inc(&parent->pending_ios);
>> +	btrfs_submit_bbio(&cb->bbio, 0);
>> +	return true;
>> +
>> +out_free_reserve:
>> +	btrfs_dec_block_group_reservations(fs_info, ins.objectid);
> 
> We free bytes_reserved, so bytes_may_use and bytes_reserved both no
> longer reflect the reservation.
> 
>> +	btrfs_free_reserved_extent(fs_info, ins.objectid, ins.offset, true);
>> +	btrfs_unlock_extent(&inode->io_tree, fileoff, end, &cached);
>> +	cleanup_compressed_bio(cb);
>> +	return false;
> 
> And fallback to uncompressed submission in run_delayed_bbio() which
> calls btrfs_reserve_extent() again.

Right, and that's also why the old async submission do not fallback to 
uncompressed submission.

The old code only fallback after a btrfs_reserve_extent() failure, but 
treat later failures as regular errors.

So here we should also follow that behavior.

Thanks a lot for exposing this problem.
Qu

> 
>> +}
>> +
>>   static void run_delayed_bbio(struct work_struct *work)
>>   {
>>   	struct delayed_bio_private *dbp = container_of(work, struct delayed_bio_private, work);
>>   	struct btrfs_bio *parent = dbp->delayed_bbio;
>>   
>> -	/* Compressed and uncompressed fallback is not yet implemented. */
>> +	if (try_submit_compressed(parent))
>> +		goto finish;
>> +
>> +	/* Uncompressed fallback is not yet implemented. */
>>   	ASSERT(0);
> 
> Sorry hard to reply across patches, but clearly it is still OK in this
> version :)
> 
>>   
>> +finish:
>>   	/*
>>   	 * Any real compressed/uncompressed bios have increased
>>   	 * parent->pending_ios, the last caller of btrfs_bio_end_io()
>> -- 
>> 2.55.0
>>
> 


  reply	other threads:[~2026-09-24 22:14 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-17 23:30 [PATCH v4 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
2026-09-17 23:30 ` [PATCH v4 1/6] btrfs: add delayed ordered extent support Qu Wenruo
2026-09-24 17:21   ` Boris Burkov
2026-09-17 23:30 ` [PATCH v4 2/6] btrfs: add skeleton for delayed btrfs bio Qu Wenruo
2026-09-17 23:30 ` [PATCH v4 3/6] btrfs: introduce the skeleton of delayed bbio endio function Qu Wenruo
2026-09-17 23:30 ` [PATCH v4 4/6] btrfs: introduce compression for delayed bbio Qu Wenruo
2026-09-24 17:30   ` Boris Burkov
2026-09-24 22:14     ` Qu Wenruo [this message]
2026-09-17 23:30 ` [PATCH v4 5/6] btrfs: implement uncompressed fallback " Qu Wenruo
2026-09-24 17:36   ` Boris Burkov
2026-09-24 22:17     ` Qu Wenruo
2026-09-17 23:30 ` [PATCH v4 6/6] btrfs: enable experimental delayed compression support Qu Wenruo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=7aeb8a6a-471f-4051-a1ab-b5741ad27b0f@gmx.com \
    --to=quwenruo.btrfs@gmx.com \
    --cc=boris@bur.io \
    --cc=linux-btrfs@vger.kernel.org \
    --cc=wqu@suse.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox