From: Qu Wenruo <wqu@suse.com>
To: linux-btrfs@vger.kernel.org
Subject: [PATCH v7 0/6] btrfs: delay compression to bbio submission time
Date: Fri, 25 Sep 2026 18:37:33 +0930 [thread overview]
Message-ID: <cover.1790327143.git.wqu@suse.com> (raw)
[CHANGELOG]
v7:
- Fix the error handling of run_delalloc_delayed()
Again exposed by Sashiko and looks correct.
Where other delalloc functions all lock the whole range, then alloc
EM/OEs with extent locked.
And the error handling is already assuming the full range has extent
lock held.
Follow the common behavior by locking the whole range first.
v6:
- Fix a bitmap race which is not protected
This is reported by Sashiko and looks correct.
There can be multiple bios submitted for the same delayed OE.
Thus the child_cleanup_bitmap needs to be protected from races.
Re-use btrfs_inode::ordered_tree_lock to prevent race.
- Release qgroup reserved space on error path
- Release EXTENT_DELALLOC_NEW and EXTENT_DEFRAG flags on error path
- Call btrfs_drop_extent_map_range() during error handling
The point here is even if we failed to create the real child EM, there
is still a delayed EM created at run_delalloc_delayed().
Since the parent OE cleanup will not cover the error range, we also
need to drop the delayed EM.
- Drop a duplicated xmpchg() call
btrfs_bio_end_io() is already doing it for us.
The remaining comments from Sashiko are incorrect:
- Double metadata delalloc releaseing
The out_free_reserve tag of try_submit_compressed() is only reached
when an error happens before an OE created.
So there is no child OE to do the cleanup. Furthermore
btrfs_removed_ordered_extent() won't release metadata on delayed
parent OE.
Thus the situation mentioned in the review is incorrect.
- Data space leak
The data space is already converted from bytes_may_use to
bytes_reserved at a succesful btrfs_reserve_extent().
And later on error path btrfs_free_reserved_extent() will decrease
bytes_reserved.
Furthermore no parent OE will cleanup the range, so there is no
leak on data reserved space.
Although the io tree bits (EXTENT_DELALLOC_NEW and EXTENT_DEFRAG) is
still leaking.
v5:
- Fix an error handling path when compression writes failed after reserved
data extents
The current fallback behavior is different than the old async path,
and can lead to reserved space inbalance.
Follow the old async behavior that if an error is hit after reserved
a data extent, error out without falling back to uncompressed writes.
- Fix an error handling path that both parent and failed child OEs are
doing the same cleanup twice
If a delayed child OE range hits some critical error without creating
an OE, e.g. failed to allocate the child OE itself, the child OE range
will be cleaned up.
But later parent OE will also cleanup any range that has no child OE,
it will clean the range again.
Fix it by changing the btrfs_ordered_extent::child_bitmap to
child_cleanup bitmap, which is initialized to all 1 when allocating
the OE, then cleared when:
* A new child OE is added
* A child OE range is already cleaned up
In either try_submit_compressed() or
submit_one_uncompressed_range().
So that the parent OE cleanup will skip any range that is covered by
an child OE or already cleaned up.
- Fix reserved metadata space leak in child OE error path
If the child OE range failed to allocate an OE, the reserved metadata
space is not released, thus will require manual cleanup.
- Fix a lockdep warning where the aquire and release is not paired
- Mark the mapping as error when end_bbio_delayed() hit an error
To follow end_bbio_data_write().
v4 RESEND:
- Fix a conflict due to some rewording in for-next
Which prevents Sashiko to review the series.
v4:
- Fix a possible write path deadlock during reclaim
The delayed_bio_private structure is allocated during writeback, it
needs GFP_NOFS to avoid deadlock.
- Use a dedicated workqueue for WQ_MEM_RECLAIM ability
v3:
- Various fixes to error handling
Mostly reported by LLM.
- Use delayed bbios' pending_io to tracking child bbios
Which reduces the structure size of delayed_bio_private.
- Extra full sync requirement for delayed writes
To avoid fast sync path to grab a delayed OE, which has no reliable
on-disk bytenr.
v2:
- Rebased to the latest for-next branch
Several minor conflicts:
* The removal of folio ordered flag
* The refactor of btrfs_mod_oustanding_extents()
- Fix a random failure in btrfs/260
It turns out that the original filemap_flush() only triggers writeback
of dirty pages, but since our new compression happens after the bios
are submitted, there can be a race between inode->defrag_compression
clearing and compression path reading inode->defrag_compression.
This can cause btrfs to use the mount option other than the specified
defrag compression algo to do the compression.
Fix it by using filemap_write_and_wait_range(), which also avoids the
quirky double flush behavior.
- Fix a use-after-free bug where bio->bi_status is accessed after
bio_put()
- Remove a mapping_set_error() call when try_submit_compressed() failed
As we still have uncompressed fallback, we should not set the mapping
as error.
- Drop all allocated OEs along with the extent maps when
run_delalloc_delayed() failed
- Slightly reword the cover letter
PoC->v1:
- Fix the ordered extent leak caused by incorrect ref count of child OEs
- Fix the reserved space leakage in ranges without a real OE
- Fix the hang caused by incorrect extent lock/unlock pair
All exposed by fsstress runs
- Fix the OE range check in btrfs_wait_ordered_extents() that affects
snapshot creation
All exposed by fstests runs
Qu Wenruo (6):
btrfs: add delayed ordered extent support
btrfs: add skeleton for delayed btrfs bio
btrfs: introduce the skeleton of delayed bbio endio function
btrfs: introduce compression for delayed bbio
btrfs: implement uncompressed fallback for delayed bbio
btrfs: enable experimental delayed compression support
fs/btrfs/bio.c | 1 +
fs/btrfs/bio.h | 7 +
fs/btrfs/btrfs_inode.h | 3 +
fs/btrfs/defrag.c | 27 +-
fs/btrfs/disk-io.c | 8 +-
fs/btrfs/extent_io.c | 36 ++-
fs/btrfs/extent_map.h | 9 +-
fs/btrfs/fs.h | 5 +
fs/btrfs/inode.c | 595 +++++++++++++++++++++++++++++++++++++++-
fs/btrfs/ordered-data.c | 185 ++++++++++---
fs/btrfs/ordered-data.h | 20 ++
11 files changed, 830 insertions(+), 66 deletions(-)
--
2.55.0
next reply other threads:[~2026-09-25 9:08 UTC|newest]
Thread overview: 9+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-25 9:07 Qu Wenruo [this message]
2026-09-25 9:07 ` [PATCH v7 1/6] btrfs: add delayed ordered extent support Qu Wenruo
2026-09-25 9:32 ` Miquel Sabaté Solà
[not found] ` <6ab63fad.d0a3cccc.111ed5.df78SMTPIN_ADDED_BROKEN@mx.google.com>
2026-09-25 9:36 ` Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 2/6] btrfs: add skeleton for delayed btrfs bio Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 3/6] btrfs: introduce the skeleton of delayed bbio endio function Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 4/6] btrfs: introduce compression for delayed bbio Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 5/6] btrfs: implement uncompressed fallback " Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 6/6] btrfs: enable experimental delayed compression support Qu Wenruo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=cover.1790327143.git.wqu@suse.com \
--to=wqu@suse.com \
--cc=linux-btrfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox