From: Qu Wenruo <quwenruo.btrfs@gmx.com>
To: Boris Burkov <boris@bur.io>, Qu Wenruo <wqu@suse.com>
Cc: linux-btrfs@vger.kernel.org
Subject: Re: [PATCH v4 4/6] btrfs: introduce compression for delayed bbio
Date: Fri, 25 Sep 2026 07:44:43 +0930 [thread overview]
Message-ID: <7aeb8a6a-471f-4051-a1ab-b5741ad27b0f@gmx.com> (raw)
In-Reply-To: <20260924173029.GB2146908@zen.localdomain>
在 2026/9/25 03:00, Boris Burkov 写道:
> On Fri, Sep 18, 2026 at 09:00:29AM +0930, Qu Wenruo wrote:
>> The compressed write path inside a delayed bbio is mostly the same as
>> regular compression, but with some differences:
>>
>> - The error handling should not touch folio flags
>> It will be handled by the parent delayed bbio.
>> And those folios already have WRITEBACK flag set, not the LOCKED flag
>> of the async submission path.
>>
>> - A successful compression will lead to a child compressed bio
>> That compressed bio will be properly submitted, and if there are no
>> more pending IOs of the delayed bbio, end the delayed bbio.
>>
>> There is a minor note, since we're going through the regular
>> extent_writepage_io() path, we can have multiple bbios for the same
>> delayed ordered extent.
>>
>> This means we may have a slightly lower compression ratio if for
>> whatever reason the writeback path chooses to submit a smaller bio.
>>
>> - No sequential execution of data extent reservation
>> The existing async thread has one quirk related to the ordered
>> function execution, which is not suitable for this call site.
>>
>> After the compressed bio is submitted, we can no longer touch the
>> child compressed bio (it could finish immediately and also finish the
>> parent delayed bbio).
>> Meanwhile the async ordered function needs different entries to handle
>> the workload and free involved structures.
>>
>> These will be the major changes compared to the existing compressed
>> write.
>>
>> - No extra blkcg/write flag inheritance
>> This is a missing feature compared to the older async compression.
>> This needs to be implemented in the future, but priority shouldn't be
>> that high.
>
> I believe there is reservation double counting, detailed inline.
>
>>
>> Signed-off-by: Qu Wenruo <wqu@suse.com>
>> ---
>> fs/btrfs/inode.c | 114 ++++++++++++++++++++++++++++++++++++++++++++++-
>> 1 file changed, 113 insertions(+), 1 deletion(-)
>>
>> diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
>> index 2d6a581d977b..fbc0e5426d45 100644
>> --- a/fs/btrfs/inode.c
>> +++ b/fs/btrfs/inode.c
>> @@ -7696,14 +7696,126 @@ struct extent_map *btrfs_create_delayed_em(struct btrfs_inode *inode,
>> return em;
>> }
>>
>> +static void end_bbio_delayed_compressed(struct btrfs_bio *bbio)
>> +{
>> + struct delayed_bio_private *dbp = bbio->private;
>> + struct btrfs_bio *parent = dbp->delayed_bbio;
>> + struct folio_iter fi;
>> +
>> + bio_for_each_folio_all(fi, &bbio->bio)
>> + btrfs_free_compr_folio(fi.folio);
>> + btrfs_bio_end_io(parent, bbio->bio.bi_status);
>> + bio_put(&bbio->bio);
>> +}
>> +
>> +static bool try_submit_compressed(struct btrfs_bio *parent)
>> +{
>> + struct delayed_bio_private *dbp = parent->private;
>> + struct btrfs_inode *inode = parent->inode;
>> + struct btrfs_fs_info *fs_info = inode->root->fs_info;
>> + struct btrfs_key ins;
>> + struct compressed_bio *cb;
>> + struct extent_state *cached = NULL;
>> + struct extent_map *em;
>> + struct btrfs_ordered_extent *ordered;
>> + struct btrfs_file_extent file_extent;
>> + u64 alloc_hint;
>> + const u32 len = bio_get_size(&parent->bio);
>> + const u64 fileoff = parent->file_offset;
>> + const u64 end = fileoff + len - 1;
>> + u32 compressed_size;
>> + int compress_type = fs_info->compress_type;
>> + int compress_level = fs_info->compress_level;
>> + int ret;
>> +
>> + if (!btrfs_inode_can_compress(inode) ||
>> + !inode_need_compress(inode, fileoff, end, false))
>> + return false;
>> +
>> + if (inode->defrag_compress > 0 &&
>> + inode->defrag_compress < BTRFS_NR_COMPRESS_TYPES) {
>> + compress_type = inode->defrag_compress;
>> + compress_level = inode->defrag_compress_level;
>> + } else if (inode->prop_compress) {
>> + compress_type = inode->prop_compress;
>> + }
>> + cb = btrfs_compress_bio(inode, fileoff, len, compress_type,
>> + compress_level, 0);
>> + if (IS_ERR(cb))
>> + return false;
>> +
>> + round_up_last_block(cb, fs_info->sectorsize);
>> + compressed_size = cb->bbio.bio.bi_iter.bi_size;
>> + /* If no space is saved, abort and mark the inode NOCOMPRESS. */
>> + if (compressed_size >= len) {
>> + if (!btrfs_test_opt(fs_info, FORCE_COMPRESS) &&
>> + !inode->prop_compress)
>> + inode->flags |= BTRFS_INODE_NOCOMPRESS;
>> + cleanup_compressed_bio(cb);
>> + return false;
>> + }
>> +
>> + alloc_hint = btrfs_get_extent_allocation_hint(inode, fileoff, len);
>> + ret = btrfs_reserve_extent(inode->root, len,
>> + compressed_size, compressed_size,
>> + 0, alloc_hint, &ins, true, true);
>
> This consumes bytes_may_use and moves it to bytes_reserved
>
>> + if (ret < 0) {
>> + cleanup_compressed_bio(cb);
>> + return false;
>> + }
>> + btrfs_lock_extent(&inode->io_tree, fileoff, end, &cached);
>> + file_extent.disk_bytenr = ins.objectid;
>> + file_extent.disk_num_bytes = ins.offset;
>> + file_extent.ram_bytes = len;
>> + file_extent.num_bytes = len;
>> + file_extent.offset = 0;
>> + file_extent.compression = cb->compress_type;
>> +
>> + cb->bbio.bio.bi_iter.bi_sector = ins.objectid >> SECTOR_SHIFT;
>> + em = btrfs_create_io_em(inode, fileoff, &file_extent, BTRFS_ORDERED_COMPRESSED);
>> + if (IS_ERR(em)) {
>> + ret = PTR_ERR(em);
>> + goto out_free_reserve;
>> + }
>> + btrfs_free_extent_map(em);
>> +
>> + ordered = btrfs_alloc_ordered_extent(inode, fileoff, &file_extent,
>> + 1U << BTRFS_ORDERED_COMPRESSED);
>> + if (IS_ERR(ordered)) {
>> + btrfs_drop_extent_map_range(inode, fileoff, end, false);
>> + ret = PTR_ERR(ordered);
>
> Now this fails with ENOMEM
>
>> + goto out_free_reserve;
>> + }
>> + cb->bbio.ordered = ordered;
>> + btrfs_dec_block_group_reservations(fs_info, ins.objectid);
>> + btrfs_unlock_extent(&inode->io_tree, fileoff, end, &cached);
>> +
>> + cb->bbio.end_io = end_bbio_delayed_compressed;
>> + cb->bbio.private = dbp;
>> + atomic_inc(&parent->pending_ios);
>> + btrfs_submit_bbio(&cb->bbio, 0);
>> + return true;
>> +
>> +out_free_reserve:
>> + btrfs_dec_block_group_reservations(fs_info, ins.objectid);
>
> We free bytes_reserved, so bytes_may_use and bytes_reserved both no
> longer reflect the reservation.
>
>> + btrfs_free_reserved_extent(fs_info, ins.objectid, ins.offset, true);
>> + btrfs_unlock_extent(&inode->io_tree, fileoff, end, &cached);
>> + cleanup_compressed_bio(cb);
>> + return false;
>
> And fallback to uncompressed submission in run_delayed_bbio() which
> calls btrfs_reserve_extent() again.
Right, and that's also why the old async submission do not fallback to
uncompressed submission.
The old code only fallback after a btrfs_reserve_extent() failure, but
treat later failures as regular errors.
So here we should also follow that behavior.
Thanks a lot for exposing this problem.
Qu
>
>> +}
>> +
>> static void run_delayed_bbio(struct work_struct *work)
>> {
>> struct delayed_bio_private *dbp = container_of(work, struct delayed_bio_private, work);
>> struct btrfs_bio *parent = dbp->delayed_bbio;
>>
>> - /* Compressed and uncompressed fallback is not yet implemented. */
>> + if (try_submit_compressed(parent))
>> + goto finish;
>> +
>> + /* Uncompressed fallback is not yet implemented. */
>> ASSERT(0);
>
> Sorry hard to reply across patches, but clearly it is still OK in this
> version :)
>
>>
>> +finish:
>> /*
>> * Any real compressed/uncompressed bios have increased
>> * parent->pending_ios, the last caller of btrfs_bio_end_io()
>> --
>> 2.55.0
>>
>
next prev parent reply other threads:[~2026-09-24 22:14 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-17 23:30 [PATCH v4 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
2026-09-17 23:30 ` [PATCH v4 1/6] btrfs: add delayed ordered extent support Qu Wenruo
2026-09-24 17:21 ` Boris Burkov
2026-09-17 23:30 ` [PATCH v4 2/6] btrfs: add skeleton for delayed btrfs bio Qu Wenruo
2026-09-17 23:30 ` [PATCH v4 3/6] btrfs: introduce the skeleton of delayed bbio endio function Qu Wenruo
2026-09-17 23:30 ` [PATCH v4 4/6] btrfs: introduce compression for delayed bbio Qu Wenruo
2026-09-24 17:30 ` Boris Burkov
2026-09-24 22:14 ` Qu Wenruo [this message]
2026-09-17 23:30 ` [PATCH v4 5/6] btrfs: implement uncompressed fallback " Qu Wenruo
2026-09-24 17:36 ` Boris Burkov
2026-09-24 22:17 ` Qu Wenruo
2026-09-17 23:30 ` [PATCH v4 6/6] btrfs: enable experimental delayed compression support Qu Wenruo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7aeb8a6a-471f-4051-a1ab-b5741ad27b0f@gmx.com \
--to=quwenruo.btrfs@gmx.com \
--cc=boris@bur.io \
--cc=linux-btrfs@vger.kernel.org \
--cc=wqu@suse.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox