Linux Btrfs filesystem development
 help / color / mirror / Atom feed
From: Qu Wenruo <wqu@suse.com>
To: linux-btrfs@vger.kernel.org
Subject: [PATCH v7 0/6] btrfs: delay compression to bbio submission time
Date: Fri, 25 Sep 2026 18:37:33 +0930	[thread overview]
Message-ID: <cover.1790327143.git.wqu@suse.com> (raw)

[CHANGELOG]
v7:
- Fix the error handling of run_delalloc_delayed()
  Again exposed by Sashiko and looks correct.

  Where other delalloc functions all lock the whole range, then alloc
  EM/OEs with extent locked.
  And the error handling is already assuming the full range has extent
  lock held.

  Follow the common behavior by locking the whole range first.

v6:
- Fix a bitmap race which is not protected
  This is reported by Sashiko and looks correct.

  There can be multiple bios submitted for the same delayed OE.
  Thus the child_cleanup_bitmap needs to be protected from races.

  Re-use btrfs_inode::ordered_tree_lock to prevent race.

- Release qgroup reserved space on error path

- Release EXTENT_DELALLOC_NEW and EXTENT_DEFRAG flags on error path

- Call btrfs_drop_extent_map_range() during error handling
  The point here is even if we failed to create the real child EM, there
  is still a delayed EM created at run_delalloc_delayed().

  Since the parent OE cleanup will not cover the error range, we also
  need to drop the delayed EM.

- Drop a duplicated xmpchg() call
  btrfs_bio_end_io() is already doing it for us.

The remaining comments from Sashiko are incorrect:

- Double metadata delalloc releaseing
  The out_free_reserve tag of try_submit_compressed() is only reached
  when an error happens before an OE created.
  So there is no child OE to do the cleanup. Furthermore
  btrfs_removed_ordered_extent() won't release metadata on delayed
  parent OE.

  Thus the situation mentioned in the review is incorrect.

- Data space leak
  The data space is already converted from bytes_may_use to
  bytes_reserved at a succesful btrfs_reserve_extent().

  And later on error path btrfs_free_reserved_extent() will decrease
  bytes_reserved.
  Furthermore no parent OE will cleanup the range, so there is no
  leak on data reserved space.

  Although the io tree bits (EXTENT_DELALLOC_NEW and EXTENT_DEFRAG) is
  still leaking.

v5:
- Fix an error handling path when compression writes failed after reserved
  data extents
  The current fallback behavior is different than the old async path,
  and can lead to reserved space inbalance.

  Follow the old async behavior that if an error is hit after reserved
  a data extent, error out without falling back to uncompressed writes.

- Fix an error handling path that both parent and failed child OEs are
  doing the same cleanup twice
  If a delayed child OE range hits some critical error without creating
  an OE, e.g. failed to allocate the child OE itself, the child OE range
  will be cleaned up.

  But later parent OE will also cleanup any range that has no child OE,
  it will clean the range again.

  Fix it by changing the btrfs_ordered_extent::child_bitmap to
  child_cleanup bitmap, which is initialized to all 1 when allocating
  the OE, then cleared when:
  * A new child OE is added
  * A child OE range is already cleaned up
    In either try_submit_compressed() or
    submit_one_uncompressed_range().

  So that the parent OE cleanup will skip any range that is covered by
  an child OE or already cleaned up.

- Fix reserved metadata space leak in child OE error path
  If the child OE range failed to allocate an OE, the reserved metadata
  space is not released, thus will require manual cleanup.

- Fix a lockdep warning where the aquire and release is not paired

- Mark the mapping as error when end_bbio_delayed() hit an error
  To follow end_bbio_data_write().

v4 RESEND:
- Fix a conflict due to some rewording in for-next
  Which prevents Sashiko to review the series.

v4:
- Fix a possible write path deadlock during reclaim
  The delayed_bio_private structure is allocated during writeback, it
  needs GFP_NOFS to avoid deadlock.

- Use a dedicated workqueue for WQ_MEM_RECLAIM ability

v3:
- Various fixes to error handling
  Mostly reported by LLM.

- Use delayed bbios' pending_io to tracking child bbios
  Which reduces the structure size of delayed_bio_private.

- Extra full sync requirement for delayed writes
  To avoid fast sync path to grab a delayed OE, which has no reliable
  on-disk bytenr.

v2:
- Rebased to the latest for-next branch
  Several minor conflicts:
  * The removal of folio ordered flag
  * The refactor of btrfs_mod_oustanding_extents()

- Fix a random failure in btrfs/260
  It turns out that the original filemap_flush() only triggers writeback
  of dirty pages, but since our new compression happens after the bios
  are submitted, there can be a race between inode->defrag_compression
  clearing and compression path reading inode->defrag_compression.

  This can cause btrfs to use the mount option other than the specified
  defrag compression algo to do the compression.

  Fix it by using filemap_write_and_wait_range(), which also avoids the
  quirky double flush behavior.
  
- Fix a use-after-free bug where bio->bi_status is accessed after
  bio_put()

- Remove a mapping_set_error() call when try_submit_compressed() failed
  As we still have uncompressed fallback, we should not set the mapping
  as error.

- Drop all allocated OEs along with the extent maps when
  run_delalloc_delayed() failed
  
- Slightly reword the cover letter

PoC->v1:
- Fix the ordered extent leak caused by incorrect ref count of child OEs
- Fix the reserved space leakage in ranges without a real OE
- Fix the hang caused by incorrect extent lock/unlock pair
  All exposed by fsstress runs

- Fix the OE range check in btrfs_wait_ordered_extents() that affects
  snapshot creation
  All exposed by fstests runs



Qu Wenruo (6):
  btrfs: add delayed ordered extent support
  btrfs: add skeleton for delayed btrfs bio
  btrfs: introduce the skeleton of delayed bbio endio function
  btrfs: introduce compression for delayed bbio
  btrfs: implement uncompressed fallback for delayed bbio
  btrfs: enable experimental delayed compression support

 fs/btrfs/bio.c          |   1 +
 fs/btrfs/bio.h          |   7 +
 fs/btrfs/btrfs_inode.h  |   3 +
 fs/btrfs/defrag.c       |  27 +-
 fs/btrfs/disk-io.c      |   8 +-
 fs/btrfs/extent_io.c    |  36 ++-
 fs/btrfs/extent_map.h   |   9 +-
 fs/btrfs/fs.h           |   5 +
 fs/btrfs/inode.c        | 595 +++++++++++++++++++++++++++++++++++++++-
 fs/btrfs/ordered-data.c | 185 ++++++++++---
 fs/btrfs/ordered-data.h |  20 ++
 11 files changed, 830 insertions(+), 66 deletions(-)

-- 
2.55.0


             reply	other threads:[~2026-09-25  9:08 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25  9:07 Qu Wenruo [this message]
2026-09-25  9:07 ` [PATCH v7 1/6] btrfs: add delayed ordered extent support Qu Wenruo
2026-09-25  9:32   ` Miquel Sabaté Solà
     [not found]   ` <6ab63fad.d0a3cccc.111ed5.df78SMTPIN_ADDED_BROKEN@mx.google.com>
2026-09-25  9:36     ` Qu Wenruo
2026-09-25  9:07 ` [PATCH v7 2/6] btrfs: add skeleton for delayed btrfs bio Qu Wenruo
2026-09-25  9:07 ` [PATCH v7 3/6] btrfs: introduce the skeleton of delayed bbio endio function Qu Wenruo
2026-09-25  9:07 ` [PATCH v7 4/6] btrfs: introduce compression for delayed bbio Qu Wenruo
2026-09-25  9:07 ` [PATCH v7 5/6] btrfs: implement uncompressed fallback " Qu Wenruo
2026-09-25  9:07 ` [PATCH v7 6/6] btrfs: enable experimental delayed compression support Qu Wenruo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=cover.1790327143.git.wqu@suse.com \
    --to=wqu@suse.com \
    --cc=linux-btrfs@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox