* [PATCH v7 0/6] btrfs: delay compression to bbio submission time
@ 2026-09-25 9:07 Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 1/6] btrfs: add delayed ordered extent support Qu Wenruo
` (5 more replies)
0 siblings, 6 replies; 9+ messages in thread
From: Qu Wenruo @ 2026-09-25 9:07 UTC (permalink / raw)
To: linux-btrfs
[CHANGELOG]
v7:
- Fix the error handling of run_delalloc_delayed()
Again exposed by Sashiko and looks correct.
Where other delalloc functions all lock the whole range, then alloc
EM/OEs with extent locked.
And the error handling is already assuming the full range has extent
lock held.
Follow the common behavior by locking the whole range first.
v6:
- Fix a bitmap race which is not protected
This is reported by Sashiko and looks correct.
There can be multiple bios submitted for the same delayed OE.
Thus the child_cleanup_bitmap needs to be protected from races.
Re-use btrfs_inode::ordered_tree_lock to prevent race.
- Release qgroup reserved space on error path
- Release EXTENT_DELALLOC_NEW and EXTENT_DEFRAG flags on error path
- Call btrfs_drop_extent_map_range() during error handling
The point here is even if we failed to create the real child EM, there
is still a delayed EM created at run_delalloc_delayed().
Since the parent OE cleanup will not cover the error range, we also
need to drop the delayed EM.
- Drop a duplicated xmpchg() call
btrfs_bio_end_io() is already doing it for us.
The remaining comments from Sashiko are incorrect:
- Double metadata delalloc releaseing
The out_free_reserve tag of try_submit_compressed() is only reached
when an error happens before an OE created.
So there is no child OE to do the cleanup. Furthermore
btrfs_removed_ordered_extent() won't release metadata on delayed
parent OE.
Thus the situation mentioned in the review is incorrect.
- Data space leak
The data space is already converted from bytes_may_use to
bytes_reserved at a succesful btrfs_reserve_extent().
And later on error path btrfs_free_reserved_extent() will decrease
bytes_reserved.
Furthermore no parent OE will cleanup the range, so there is no
leak on data reserved space.
Although the io tree bits (EXTENT_DELALLOC_NEW and EXTENT_DEFRAG) is
still leaking.
v5:
- Fix an error handling path when compression writes failed after reserved
data extents
The current fallback behavior is different than the old async path,
and can lead to reserved space inbalance.
Follow the old async behavior that if an error is hit after reserved
a data extent, error out without falling back to uncompressed writes.
- Fix an error handling path that both parent and failed child OEs are
doing the same cleanup twice
If a delayed child OE range hits some critical error without creating
an OE, e.g. failed to allocate the child OE itself, the child OE range
will be cleaned up.
But later parent OE will also cleanup any range that has no child OE,
it will clean the range again.
Fix it by changing the btrfs_ordered_extent::child_bitmap to
child_cleanup bitmap, which is initialized to all 1 when allocating
the OE, then cleared when:
* A new child OE is added
* A child OE range is already cleaned up
In either try_submit_compressed() or
submit_one_uncompressed_range().
So that the parent OE cleanup will skip any range that is covered by
an child OE or already cleaned up.
- Fix reserved metadata space leak in child OE error path
If the child OE range failed to allocate an OE, the reserved metadata
space is not released, thus will require manual cleanup.
- Fix a lockdep warning where the aquire and release is not paired
- Mark the mapping as error when end_bbio_delayed() hit an error
To follow end_bbio_data_write().
v4 RESEND:
- Fix a conflict due to some rewording in for-next
Which prevents Sashiko to review the series.
v4:
- Fix a possible write path deadlock during reclaim
The delayed_bio_private structure is allocated during writeback, it
needs GFP_NOFS to avoid deadlock.
- Use a dedicated workqueue for WQ_MEM_RECLAIM ability
v3:
- Various fixes to error handling
Mostly reported by LLM.
- Use delayed bbios' pending_io to tracking child bbios
Which reduces the structure size of delayed_bio_private.
- Extra full sync requirement for delayed writes
To avoid fast sync path to grab a delayed OE, which has no reliable
on-disk bytenr.
v2:
- Rebased to the latest for-next branch
Several minor conflicts:
* The removal of folio ordered flag
* The refactor of btrfs_mod_oustanding_extents()
- Fix a random failure in btrfs/260
It turns out that the original filemap_flush() only triggers writeback
of dirty pages, but since our new compression happens after the bios
are submitted, there can be a race between inode->defrag_compression
clearing and compression path reading inode->defrag_compression.
This can cause btrfs to use the mount option other than the specified
defrag compression algo to do the compression.
Fix it by using filemap_write_and_wait_range(), which also avoids the
quirky double flush behavior.
- Fix a use-after-free bug where bio->bi_status is accessed after
bio_put()
- Remove a mapping_set_error() call when try_submit_compressed() failed
As we still have uncompressed fallback, we should not set the mapping
as error.
- Drop all allocated OEs along with the extent maps when
run_delalloc_delayed() failed
- Slightly reword the cover letter
PoC->v1:
- Fix the ordered extent leak caused by incorrect ref count of child OEs
- Fix the reserved space leakage in ranges without a real OE
- Fix the hang caused by incorrect extent lock/unlock pair
All exposed by fsstress runs
- Fix the OE range check in btrfs_wait_ordered_extents() that affects
snapshot creation
All exposed by fstests runs
Qu Wenruo (6):
btrfs: add delayed ordered extent support
btrfs: add skeleton for delayed btrfs bio
btrfs: introduce the skeleton of delayed bbio endio function
btrfs: introduce compression for delayed bbio
btrfs: implement uncompressed fallback for delayed bbio
btrfs: enable experimental delayed compression support
fs/btrfs/bio.c | 1 +
fs/btrfs/bio.h | 7 +
fs/btrfs/btrfs_inode.h | 3 +
fs/btrfs/defrag.c | 27 +-
fs/btrfs/disk-io.c | 8 +-
fs/btrfs/extent_io.c | 36 ++-
fs/btrfs/extent_map.h | 9 +-
fs/btrfs/fs.h | 5 +
fs/btrfs/inode.c | 595 +++++++++++++++++++++++++++++++++++++++-
fs/btrfs/ordered-data.c | 185 ++++++++++---
fs/btrfs/ordered-data.h | 20 ++
11 files changed, 830 insertions(+), 66 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v7 1/6] btrfs: add delayed ordered extent support
2026-09-25 9:07 [PATCH v7 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
@ 2026-09-25 9:07 ` Qu Wenruo
2026-09-25 9:32 ` Miquel Sabaté Solà
[not found] ` <6ab63fad.d0a3cccc.111ed5.df78SMTPIN_ADDED_BROKEN@mx.google.com>
2026-09-25 9:07 ` [PATCH v7 2/6] btrfs: add skeleton for delayed btrfs bio Qu Wenruo
` (4 subsequent siblings)
5 siblings, 2 replies; 9+ messages in thread
From: Qu Wenruo @ 2026-09-25 9:07 UTC (permalink / raw)
To: linux-btrfs
A delayed ordered extent has the following features:
- A new BTRFS_ORDERED_DELAYED flag
And this new flag must be set along with the BTRFS_ORDERED_REGULAR flag.
- No allocation of any on-disk space
As a delayed ordered extent doesn't take any on-disk space yet, it
won't release any reserved data/meta space either.
- Zero or more real OEs can be added to the parent
If a real OE is allocated, it must be inside the parent OE.
And such real OE will go through the regular data/meta space
reservation path.
- Child OEs will not be added to the per-inode OE rb-tree nor
per-root list
Only the parent OE is added to the per-inode rb-tree and per-root
list.
So anything waiting for ordered extents should only work on the parent
one.
There is a special corner case for btrfs_wait_ordered_extents(), as
delayed parent OEs have 0 disk_bytenr and disk_num_bytes, they will
be considered out of the [0, U64_MAX] range.
Thus we have to always wait for any delayed OEs of a root, no matter
if a block group range is given or not.
- When the parent OE finishes, all child OEs will also be finished
And reserved space is all handled by the child OEs.
- Any range not covered by a child OE will be manually cleaned up
When adding a child OE to the parent one, the range in
child_cleanup_bitmap will be cleared.
If a range is already cleaned up but without a child OE (happens when
OE allocation failed), whoever cleans up the range should clear the bits
in the child_cleanup_bitmap.
And when the parent OE finishes, any range in child_cleanup_bitmap
will be properly cleaned up.
The above features allow us to use the existing ordered extent interfaces
to allocate new real OEs, and wait for them properly.
Signed-off-by: Qu Wenruo <wqu@suse.com>
---
fs/btrfs/inode.c | 86 ++++++++++++++++++-
fs/btrfs/ordered-data.c | 185 ++++++++++++++++++++++++++++++----------
fs/btrfs/ordered-data.h | 20 +++++
3 files changed, 243 insertions(+), 48 deletions(-)
diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
index 2a32072849cb..e9ed3e31fa88 100644
--- a/fs/btrfs/inode.c
+++ b/fs/btrfs/inode.c
@@ -3204,6 +3204,81 @@ static int insert_ordered_extent_file_extent(struct btrfs_trans_handle *trans,
update_inode_bytes, oe->qgroup_rsv);
}
+static int finish_delayed_ordered(struct btrfs_ordered_extent *oe)
+{
+ struct btrfs_inode *inode = oe->inode;
+ struct btrfs_fs_info *fs_info = inode->root->fs_info;
+ struct btrfs_ordered_extent *child;
+ struct btrfs_ordered_extent *tmp;
+ struct extent_state *cached = NULL;
+ const u32 nr_bits = oe->num_bytes >> fs_info->sectorsize_bits;
+ bool io_error = test_bit(BTRFS_ORDERED_IOERR, &oe->flags);
+ u32 cur_bit = 0;
+ int ret = 0;
+ int saved_ret = 0;
+
+ /* Finish each child OE. */
+ list_for_each_entry_safe(child, tmp, &oe->child_list, child_list) {
+ const u32 child_bit = (child->file_offset - oe->file_offset) >>
+ fs_info->sectorsize_bits;
+ const u32 child_nr_bits = child->num_bytes >> fs_info->sectorsize_bits;
+
+ list_del_init(&child->child_list);
+ refcount_inc(&child->refs);
+
+ /* The range should have been cleared in the bitmap. */
+ ASSERT(bitmap_test_range_all_zero(oe->child_cleanup_bitmap,
+ child_bit, child_nr_bits));
+
+ if (io_error)
+ set_bit(BTRFS_ORDERED_IOERR, &child->flags);
+
+ ret = btrfs_finish_one_ordered(child);
+ if (ret && !saved_ret)
+ saved_ret = ret;
+ }
+
+ while (cur_bit < nr_bits) {
+ u64 range_start;
+ u64 range_end;
+ u32 range_len;
+ unsigned int first_zero;
+
+ cur_bit = find_next_bit(oe->child_cleanup_bitmap, nr_bits, cur_bit);
+
+ if (cur_bit >= nr_bits)
+ break;
+
+ first_zero = find_next_zero_bit(oe->child_cleanup_bitmap, nr_bits,
+ cur_bit);
+ range_start = oe->file_offset + (cur_bit << fs_info->sectorsize_bits);
+ range_len = (first_zero - cur_bit) << fs_info->sectorsize_bits;
+ range_end = range_start + range_len - 1;
+ cur_bit = first_zero;
+
+ btrfs_lock_extent(&inode->io_tree, range_start, range_end, &cached);
+ /*
+ * The range has reserved data/metadata but no real OE, thus we have
+ * to manually release them.
+ */
+ btrfs_delalloc_release_space(inode, NULL, range_start, range_len, true);
+ /*
+ * Also need to remove/drop the pinned extent map range.
+ * Here we do not want the extent map to stay, as they do not represent
+ * any real extent on-disk.
+ */
+ btrfs_drop_extent_map_range(inode, range_start, range_end, false);
+ btrfs_clear_extent_bit(&inode->io_tree, range_start, range_end,
+ EXTENT_LOCKED | EXTENT_DELALLOC_NEW | EXTENT_DEFRAG |
+ EXTENT_DO_ACCOUNTING, &cached);
+ }
+
+ btrfs_remove_ordered_extent(oe);
+ btrfs_put_ordered_extent(oe);
+ btrfs_put_ordered_extent(oe);
+ return saved_ret;
+}
+
/*
* As ordered data IO finishes, this gets called so we can finish
* an ordered extent if the range of bytes in the file it covers are
@@ -3226,6 +3301,13 @@ int btrfs_finish_one_ordered(struct btrfs_ordered_extent *ordered_extent)
bool clear_reserved_extent = true;
unsigned int clear_bits = 0;
+ freespace_inode = btrfs_is_free_space_inode(inode);
+ if (!freespace_inode)
+ btrfs_lockdep_acquire(fs_info, btrfs_ordered_extent);
+
+ if (test_bit(BTRFS_ORDERED_DELAYED, &ordered_extent->flags))
+ return finish_delayed_ordered(ordered_extent);
+
start = ordered_extent->file_offset;
end = start + ordered_extent->num_bytes - 1;
@@ -3238,10 +3320,6 @@ int btrfs_finish_one_ordered(struct btrfs_ordered_extent *ordered_extent)
if (!test_bit(BTRFS_ORDERED_NOCOW, &ordered_extent->flags))
clear_bits |= EXTENT_DEFRAG;
- freespace_inode = btrfs_is_free_space_inode(inode);
- if (!freespace_inode)
- btrfs_lockdep_acquire(fs_info, btrfs_ordered_extent);
-
if (unlikely(test_bit(BTRFS_ORDERED_IOERR, &ordered_extent->flags))) {
ret = -EIO;
goto out;
diff --git a/fs/btrfs/ordered-data.c b/fs/btrfs/ordered-data.c
index b32d4eabe0ab..aad26972f8b4 100644
--- a/fs/btrfs/ordered-data.c
+++ b/fs/btrfs/ordered-data.c
@@ -155,6 +155,7 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
u64 qgroup_rsv = 0;
const bool is_nocow = (flags &
((1U << BTRFS_ORDERED_NOCOW) | (1U << BTRFS_ORDERED_PREALLOC)));
+ const bool is_delayed = test_bit(BTRFS_ORDERED_DELAYED, &flags);
/* Only one type flag can be set. */
ASSERT(has_single_bit_set(flags & BTRFS_ORDERED_EXCLUSIVE_FLAGS),
@@ -170,6 +171,17 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
if (test_bit(BTRFS_ORDERED_ENCODED, &flags))
ASSERT(test_bit(BTRFS_ORDERED_COMPRESSED, &flags));
+ /*
+ * DELAYED can only be set with REGULAR, no DIRECT/ENCODED, and should
+ * not exceed BTRFS_MAX_COMPRESSED size.
+ */
+ if (test_bit(BTRFS_ORDERED_DELAYED, &flags)) {
+ ASSERT(test_bit(BTRFS_ORDERED_REGULAR, &flags));
+ ASSERT(!test_bit(BTRFS_ORDERED_DIRECT, &flags));
+ ASSERT(!test_bit(BTRFS_ORDERED_ENCODED, &flags));
+ ASSERT(num_bytes <= BTRFS_MAX_COMPRESSED);
+ }
+
/*
* For a NOCOW write we can free the qgroup reserve right now. For a COW
* one we transfer the reserved space from the inode's iotree into the
@@ -178,13 +190,16 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
* completing the ordered extent, when running the data delayed ref it
* creates, we free the reserved data with btrfs_qgroup_free_refroot().
*/
- if (is_nocow)
- ret = btrfs_qgroup_free_data(inode, NULL, file_offset, num_bytes, &qgroup_rsv);
- else
- ret = btrfs_qgroup_release_data(inode, file_offset, num_bytes, &qgroup_rsv);
-
- if (ret < 0)
- return ERR_PTR(ret);
+ if (!is_delayed) {
+ if (is_nocow)
+ ret = btrfs_qgroup_free_data(inode, NULL, file_offset,
+ num_bytes, &qgroup_rsv);
+ else
+ ret = btrfs_qgroup_release_data(inode, file_offset,
+ num_bytes, &qgroup_rsv);
+ if (ret < 0)
+ return ERR_PTR(ret);
+ }
entry = kmem_cache_zalloc(btrfs_ordered_extent_cache, GFP_NOFS);
if (!entry) {
@@ -216,19 +231,26 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
INIT_LIST_HEAD(&entry->root_extent_list);
INIT_LIST_HEAD(&entry->work_list);
INIT_LIST_HEAD(&entry->bioc_list);
+ INIT_LIST_HEAD(&entry->child_list);
init_completion(&entry->completion);
+ RB_CLEAR_NODE(&entry->rb_node);
/*
* We don't need the count_max_extents here, we can assume that all of
* that work has been done at higher layers, so this is truly the
* smallest the extent is going to get.
*/
- spin_lock(&inode->lock);
- btrfs_mod_outstanding_extents(inode, 1);
- spin_unlock(&inode->lock);
+ if (!is_delayed) {
+ spin_lock(&inode->lock);
+ btrfs_mod_outstanding_extents(inode, 1);
+ spin_unlock(&inode->lock);
+ } else {
+ bitmap_set(entry->child_cleanup_bitmap, 0,
+ num_bytes >> inode->root->fs_info->sectorsize_bits);
+ }
out:
- if (IS_ERR(entry) && !is_nocow)
+ if (IS_ERR(entry) && !is_nocow && !is_delayed)
btrfs_qgroup_free_refroot(inode->root->fs_info,
btrfs_root_id(inode->root),
qgroup_rsv, BTRFS_QGROUP_RSV_DATA);
@@ -236,12 +258,47 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
return entry;
}
+static void add_child_oe(struct btrfs_ordered_extent *parent,
+ struct btrfs_ordered_extent *child)
+{
+ struct btrfs_inode *inode = parent->inode;
+ struct btrfs_fs_info *fs_info = inode->root->fs_info;
+ const u32 start_bit = (child->file_offset - parent->file_offset) >>
+ fs_info->sectorsize_bits;
+ const u32 nr_bits = child->num_bytes >> fs_info->sectorsize_bits;
+
+ lockdep_assert_held(&inode->ordered_tree_lock);
+ /* Basic flags check for parent and child. */
+ ASSERT(test_bit(BTRFS_ORDERED_DELAYED, &parent->flags));
+ ASSERT(!test_bit(BTRFS_ORDERED_DELAYED, &child->flags));
+
+ /* Child should not belong to any parent yet. */
+ ASSERT(list_empty(&child->child_list));
+
+ /* Child should be fully inside parent's range. */
+ ASSERT(child->file_offset >= parent->file_offset);
+ ASSERT(child->file_offset + child->num_bytes <=
+ parent->file_offset + parent->num_bytes);
+
+ /*
+ * There should be no existing child in the range, thus
+ * all cleanup bits should be set.
+ */
+ ASSERT(bitmap_test_range_all_set(parent->child_cleanup_bitmap,
+ start_bit, nr_bits));
+
+ list_add_tail(&child->child_list, &parent->child_list);
+
+ bitmap_clear(parent->child_cleanup_bitmap, start_bit, nr_bits);
+}
+
static void insert_ordered_extent(struct btrfs_ordered_extent *entry)
{
struct btrfs_inode *inode = entry->inode;
struct btrfs_root *root = inode->root;
struct btrfs_fs_info *fs_info = root->fs_info;
struct rb_node *node;
+ bool is_child = false;
trace_btrfs_ordered_extent_add(inode, entry);
@@ -254,17 +311,25 @@ static void insert_ordered_extent(struct btrfs_ordered_extent *entry)
spin_lock(&inode->ordered_tree_lock);
node = tree_insert(&inode->ordered_tree, entry->file_offset,
&entry->rb_node);
- if (unlikely(node)) {
+ if (node) {
struct btrfs_ordered_extent *exist =
rb_entry(node, struct btrfs_ordered_extent, rb_node);
- btrfs_panic(fs_info, -EEXIST,
+ if (test_bit(BTRFS_ORDERED_DELAYED, &exist->flags)) {
+ add_child_oe(exist, entry);
+ is_child = true;
+ } else {
+ btrfs_panic(fs_info, -EEXIST,
"overlapping ordered extents, existing oe file_offset %llu num_bytes %llu flags 0x%lx, new oe file_offset %llu num_bytes %llu flags 0x%lx",
exist->file_offset, exist->num_bytes, exist->flags,
entry->file_offset, entry->num_bytes, entry->flags);
+ }
}
spin_unlock(&inode->ordered_tree_lock);
+ /* Child OE shouldn't be added to per-root oe list. */
+ if (is_child)
+ return;
spin_lock(&root->ordered_extent_lock);
list_add_tail(&entry->root_extent_list,
&root->ordered_extents);
@@ -337,6 +402,20 @@ struct btrfs_ordered_extent *btrfs_alloc_ordered_extent(
return entry;
}
+struct btrfs_ordered_extent *btrfs_alloc_delayed_ordered_extent(
+ struct btrfs_inode *inode, u64 file_offset, u32 length)
+{
+ struct btrfs_ordered_extent *entry;
+
+ entry = alloc_ordered_extent(inode, file_offset, length, length, 0, 0, 0,
+ (1UL << BTRFS_ORDERED_REGULAR) |
+ (1UL << BTRFS_ORDERED_DELAYED),
+ BTRFS_COMPRESS_NONE);
+ if (!IS_ERR(entry))
+ insert_ordered_extent(entry);
+ return entry;
+}
+
/*
* Add a struct btrfs_ordered_sum into the list of checksums to be inserted
* when an ordered extent is finished. If the list covers more than one
@@ -656,8 +735,9 @@ void btrfs_remove_ordered_extent(struct btrfs_ordered_extent *entry)
struct btrfs_root *root = btrfs_inode->root;
struct btrfs_fs_info *fs_info = root->fs_info;
struct rb_node *node;
- bool pending;
+ bool pending = false;
bool freespace_inode;
+ const bool is_delayed = test_bit(BTRFS_ORDERED_DELAYED, &entry->flags);
/*
* If this is a free space inode the thread has not acquired the ordered
@@ -666,33 +746,37 @@ void btrfs_remove_ordered_extent(struct btrfs_ordered_extent *entry)
freespace_inode = btrfs_is_free_space_inode(btrfs_inode);
btrfs_lockdep_acquire(fs_info, btrfs_trans_pending_ordered);
- /* This is paired with alloc_ordered_extent(). */
- spin_lock(&btrfs_inode->lock);
- btrfs_mod_outstanding_extents(btrfs_inode, -1);
- spin_unlock(&btrfs_inode->lock);
- if (root != fs_info->tree_root) {
- u64 release;
+ if (!is_delayed) {
+ /* This is paired with alloc_ordered_extent(). */
+ spin_lock(&btrfs_inode->lock);
+ btrfs_mod_outstanding_extents(btrfs_inode, -1);
+ spin_unlock(&btrfs_inode->lock);
- if (test_bit(BTRFS_ORDERED_ENCODED, &entry->flags))
- release = entry->disk_num_bytes;
- else
- release = entry->num_bytes;
- btrfs_delalloc_release_metadata(btrfs_inode, release,
+ if (root != fs_info->tree_root) {
+ u64 release;
+
+ if (test_bit(BTRFS_ORDERED_ENCODED, &entry->flags))
+ release = entry->disk_num_bytes;
+ else
+ release = entry->num_bytes;
+ btrfs_delalloc_release_metadata(btrfs_inode, release,
test_bit(BTRFS_ORDERED_IOERR,
&entry->flags));
+ }
}
-
percpu_counter_add_batch(&fs_info->ordered_bytes, -entry->num_bytes,
fs_info->delalloc_batch);
spin_lock(&btrfs_inode->ordered_tree_lock);
- node = &entry->rb_node;
- rb_erase(node, &btrfs_inode->ordered_tree);
- RB_CLEAR_NODE(node);
- if (btrfs_inode->ordered_tree_last == node)
- btrfs_inode->ordered_tree_last = NULL;
- set_bit(BTRFS_ORDERED_COMPLETE, &entry->flags);
- pending = test_and_clear_bit(BTRFS_ORDERED_PENDING, &entry->flags);
+ if (!RB_EMPTY_NODE(&entry->rb_node)) {
+ node = &entry->rb_node;
+ rb_erase(node, &btrfs_inode->ordered_tree);
+ RB_CLEAR_NODE(node);
+ if (btrfs_inode->ordered_tree_last == node)
+ btrfs_inode->ordered_tree_last = NULL;
+ set_bit(BTRFS_ORDERED_COMPLETE, &entry->flags);
+ pending = test_and_clear_bit(BTRFS_ORDERED_PENDING, &entry->flags);
+ }
spin_unlock(&btrfs_inode->ordered_tree_lock);
/*
@@ -724,17 +808,20 @@ void btrfs_remove_ordered_extent(struct btrfs_ordered_extent *entry)
btrfs_lockdep_release(fs_info, btrfs_trans_pending_ordered);
- spin_lock(&root->ordered_extent_lock);
- list_del_init(&entry->root_extent_list);
- root->nr_ordered_extents--;
-
trace_btrfs_ordered_extent_remove(btrfs_inode, entry);
- if (!root->nr_ordered_extents) {
- spin_lock(&fs_info->ordered_root_lock);
- BUG_ON(list_empty(&root->ordered_root));
- list_del_init(&root->ordered_root);
- spin_unlock(&fs_info->ordered_root_lock);
+ spin_lock(&root->ordered_extent_lock);
+ /* For child OEs, they are not added to per-root OEs. */
+ if (!list_empty(&entry->root_extent_list)) {
+ list_del_init(&entry->root_extent_list);
+ root->nr_ordered_extents--;
+
+ if (!root->nr_ordered_extents) {
+ spin_lock(&fs_info->ordered_root_lock);
+ BUG_ON(list_empty(&root->ordered_root));
+ list_del_init(&root->ordered_root);
+ spin_unlock(&fs_info->ordered_root_lock);
+ }
}
spin_unlock(&root->ordered_extent_lock);
wake_up(&entry->wait);
@@ -783,8 +870,18 @@ u64 btrfs_wait_ordered_extents(struct btrfs_root *root, u64 nr,
ordered = list_first_entry(&splice, struct btrfs_ordered_extent,
root_extent_list);
- if (range_end <= ordered->disk_bytenr ||
- ordered->disk_bytenr + ordered->disk_num_bytes <= range_start) {
+ /*
+ * Delayed OEs have 0 disk_bytenr and 0 disk_num_bytes, thus
+ * they will be considered out of the [0, U64_MAX) range.
+ * And we do not know where they will really land until the
+ * writeback has finished.
+ *
+ * So here we must exclude delayed OEs from the block group
+ * range check, and always wait for them.
+ */
+ if (!test_bit(BTRFS_ORDERED_DELAYED, &ordered->flags) &&
+ (range_end <= ordered->disk_bytenr ||
+ ordered->disk_bytenr + ordered->disk_num_bytes <= range_start)) {
list_move_tail(&ordered->root_extent_list, &skipped);
cond_resched_lock(&root->ordered_extent_lock);
continue;
diff --git a/fs/btrfs/ordered-data.h b/fs/btrfs/ordered-data.h
index 8d5d5ba1e02f..f72588c65e62 100644
--- a/fs/btrfs/ordered-data.h
+++ b/fs/btrfs/ordered-data.h
@@ -13,6 +13,7 @@
#include <linux/rbtree.h>
#include <linux/wait.h>
#include "async-thread.h"
+#include "compression.h"
struct inode;
struct page;
@@ -87,6 +88,12 @@ enum {
*/
BTRFS_ORDERED_DIRECT,
+ /*
+ * Extra bit for delayed OE, can only be set for REGULAR.
+ * Cannot be set with COMPRESSED/ENCODED/DIRECT.
+ */
+ BTRFS_ORDERED_DELAYED,
+
BTRFS_ORDERED_NR_FLAGS,
};
static_assert(BTRFS_ORDERED_NR_FLAGS <= BITS_PER_LONG);
@@ -155,6 +162,17 @@ struct btrfs_ordered_extent {
/* a per root list of all the pending ordered extents */
struct list_head root_extent_list;
+ /* Child ordered extent list for delayed OE. */
+ struct list_head child_list;
+
+ /*
+ * Only utilized by delayed parent OE.
+ *
+ * Indicate the range that needs cleanup.
+ */
+ unsigned long child_cleanup_bitmap[BITS_TO_LONGS(
+ BTRFS_MAX_COMPRESSED / BTRFS_MIN_BLOCKSIZE)];
+
struct btrfs_work work;
struct completion completion;
@@ -192,6 +210,8 @@ struct btrfs_file_extent {
struct btrfs_ordered_extent *btrfs_alloc_ordered_extent(
struct btrfs_inode *inode, u64 file_offset,
const struct btrfs_file_extent *file_extent, unsigned long flags);
+struct btrfs_ordered_extent *btrfs_alloc_delayed_ordered_extent(
+ struct btrfs_inode *inode, u64 file_offset, u32 length);
void btrfs_add_ordered_sum(struct btrfs_ordered_extent *entry,
struct btrfs_ordered_sum *sum);
struct btrfs_ordered_extent *btrfs_lookup_ordered_extent(struct btrfs_inode *inode,
--
2.55.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v7 2/6] btrfs: add skeleton for delayed btrfs bio
2026-09-25 9:07 [PATCH v7 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 1/6] btrfs: add delayed ordered extent support Qu Wenruo
@ 2026-09-25 9:07 ` Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 3/6] btrfs: introduce the skeleton of delayed bbio endio function Qu Wenruo
` (3 subsequent siblings)
5 siblings, 0 replies; 9+ messages in thread
From: Qu Wenruo @ 2026-09-25 9:07 UTC (permalink / raw)
To: linux-btrfs
The objective of such new delayed btrfs bio infrastructure is to allow
compressed write to go through the regular extent_writepage_io() path,
instead of the async submission path.
This will make it easier to align our write path to iomap.
The core ideas of delayed btrfs bio are:
- A placeholder ordered extent created at delalloc time
No space is reserved at that time.
- A delayed extent map created at delalloc time
It will have a special disk_bytenr (-4) to indicate the range is
delayed.
And a new EXTENT_FLAG_DELAYED flag.
- Delayed btrfs bios will be limited to BTRFS_MAX_COMPRESSED size
As only compression will go through delayed btrfs bio.
- Delayed btrfs bios will have @is_delayed flag set
And such bio will have 0 as bi_sector, but will never be submitted
directly through btrfs_submit_bbio().
Currently the submission of a delayed btrfs bio is not here yet, and
will be implemented by later patches.
- Btrfs bio assembly mostly follows the regular path
There are several small exceptions:
* btrfs_bio_is_contig() needs to handle delayed disk_bytenr/bbio
* New bbio needs to have its is_delayed flag set if disk_bytenr
is EXTENT_MAP_DELAYED
- Real ordered extents will be created at bbio submission time
This part is not implemented in this patch.
Signed-off-by: Qu Wenruo <wqu@suse.com>
---
fs/btrfs/bio.c | 1 +
fs/btrfs/bio.h | 7 +++++++
fs/btrfs/btrfs_inode.h | 3 +++
fs/btrfs/extent_io.c | 31 ++++++++++++++++++++++++------
fs/btrfs/extent_map.h | 9 ++++++++-
fs/btrfs/inode.c | 43 +++++++++++++++++++++++++++++++++++++++++-
6 files changed, 86 insertions(+), 8 deletions(-)
diff --git a/fs/btrfs/bio.c b/fs/btrfs/bio.c
index 2c6234f6182f..b17b6f13e067 100644
--- a/fs/btrfs/bio.c
+++ b/fs/btrfs/bio.c
@@ -872,6 +872,7 @@ void btrfs_submit_bbio(struct btrfs_bio *bbio, int mirror_num)
{
/* If bbio->inode is not populated, its file_offset must be 0. */
ASSERT(bbio->inode || bbio->file_offset == 0);
+ ASSERT(!bbio->is_delayed);
assert_bbio_alignment(bbio);
diff --git a/fs/btrfs/bio.h b/fs/btrfs/bio.h
index bbf362b8668b..54c2ca46b6e3 100644
--- a/fs/btrfs/bio.h
+++ b/fs/btrfs/bio.h
@@ -95,6 +95,13 @@ struct btrfs_bio {
/* Whether the bio is written using zone append. */
bool can_use_append:1;
+ /*
+ * If the bio is delayed.
+ *
+ * A delayed bbio will not be directly submitted.
+ */
+ bool is_delayed:1;
+
/*
* This member must come last, bio_alloc_bioset will allocate enough
* bytes for entire btrfs_bio but relies on bio being last.
diff --git a/fs/btrfs/btrfs_inode.h b/fs/btrfs/btrfs_inode.h
index 03c00c99099d..33c8eb5c69e8 100644
--- a/fs/btrfs/btrfs_inode.h
+++ b/fs/btrfs/btrfs_inode.h
@@ -643,5 +643,8 @@ u64 btrfs_get_extent_allocation_hint(struct btrfs_inode *inode, u64 start,
struct extent_map *btrfs_create_io_em(struct btrfs_inode *inode, u64 start,
const struct btrfs_file_extent *file_extent,
int type);
+struct extent_map *btrfs_create_delayed_em(struct btrfs_inode *inode,
+ u64 start, u32 length);
+void btrfs_submit_delayed_write(struct btrfs_bio *bbio);
#endif
diff --git a/fs/btrfs/extent_io.c b/fs/btrfs/extent_io.c
index 55e9144d4759..a822f6635eba 100644
--- a/fs/btrfs/extent_io.c
+++ b/fs/btrfs/extent_io.c
@@ -183,12 +183,16 @@ static void submit_one_bio(struct btrfs_bio_ctrl *bio_ctrl)
/* Caller should ensure the bio has at least some range added */
ASSERT(bbio->bio.bi_iter.bi_size);
-
+ /* Delayed bbio is only for write. */
+ if (bbio->is_delayed)
+ ASSERT(btrfs_op(&bbio->bio) == BTRFS_MAP_WRITE);
bio_set_csum_search_commit_root(bio_ctrl);
if (btrfs_op(&bbio->bio) == BTRFS_MAP_READ &&
bio_ctrl->compress_type != BTRFS_COMPRESS_NONE)
btrfs_submit_compressed_read(bbio);
+ else if (bbio->is_delayed)
+ btrfs_submit_delayed_write(bbio);
else
btrfs_submit_bbio(bbio, 0);
@@ -729,6 +733,14 @@ static bool btrfs_bio_is_contig(struct btrfs_bio_ctrl *bio_ctrl,
struct bio *bio = &bio_ctrl->bbio->bio;
const sector_t sector = disk_bytenr >> SECTOR_SHIFT;
+ /* One is delayed bbio and one is not, definitely not contig. */
+ if (bio_ctrl->bbio->is_delayed != (disk_bytenr == EXTENT_MAP_DELAYED))
+ return false;
+
+ /* For delayed bbio, only need to check if the file range is contig. */
+ if (bio_ctrl->bbio->is_delayed)
+ return bio_ctrl->next_file_offset == file_offset;
+
if (bio_ctrl->compress_type != BTRFS_COMPRESS_NONE) {
/*
* For compression, all IO should have its logical bytenr set
@@ -754,7 +766,13 @@ static int alloc_new_bio(struct btrfs_inode *inode,
bbio = btrfs_bio_alloc(BIO_MAX_VECS, bio_ctrl->opf, inode,
file_offset, bio_ctrl->end_io_func, NULL);
- bbio->bio.bi_iter.bi_sector = disk_bytenr >> SECTOR_SHIFT;
+ if (disk_bytenr == EXTENT_MAP_DELAYED) {
+ bbio->is_delayed = true;
+ bbio->bio.bi_iter.bi_sector = 0;
+ } else {
+ bbio->is_delayed = false;
+ bbio->bio.bi_iter.bi_sector = disk_bytenr >> SECTOR_SHIFT;
+ }
bbio->bio.bi_write_hint = inode->vfs_inode.i_write_hint;
bio_ctrl->bbio = bbio;
bio_ctrl->len_to_oe_boundary = U32_MAX;
@@ -781,7 +799,7 @@ static int alloc_new_bio(struct btrfs_inode *inode,
}
bio_ctrl->len_to_oe_boundary = min_t(u32, U32_MAX,
ordered->file_offset +
- ordered->disk_num_bytes - file_offset);
+ ordered->num_bytes - file_offset);
bbio->ordered = ordered;
/*
@@ -1812,7 +1830,6 @@ static int submit_write_sector(struct btrfs_inode *inode,
{
struct btrfs_fs_info *fs_info = inode->root->fs_info;
struct btrfs_ordered_extent *oe;
- u64 block_start;
u64 disk_bytenr;
u64 extent_offset;
const u32 sectorsize = fs_info->sectorsize;
@@ -1864,8 +1881,10 @@ static int submit_write_sector(struct btrfs_inode *inode,
ASSERT(oe->compress_type == BTRFS_COMPRESS_NONE);
ASSERT(!test_bit(BTRFS_ORDERED_COMPRESSED, &oe->flags));
- block_start = oe->disk_bytenr + oe->offset;
- disk_bytenr = block_start + extent_offset;
+ if (test_bit(BTRFS_ORDERED_DELAYED, &oe->flags))
+ disk_bytenr = EXTENT_MAP_DELAYED;
+ else
+ disk_bytenr = oe->disk_bytenr + oe->offset + extent_offset;
btrfs_put_ordered_extent(oe);
diff --git a/fs/btrfs/extent_map.h b/fs/btrfs/extent_map.h
index 6f685f3c9327..aac77762948a 100644
--- a/fs/btrfs/extent_map.h
+++ b/fs/btrfs/extent_map.h
@@ -13,7 +13,8 @@
struct btrfs_inode;
struct btrfs_fs_info;
-#define EXTENT_MAP_LAST_BYTE ((u64)-4)
+#define EXTENT_MAP_LAST_BYTE ((u64)-5)
+#define EXTENT_MAP_DELAYED ((u64)-4)
#define EXTENT_MAP_HOLE ((u64)-3)
#define EXTENT_MAP_INLINE ((u64)-2)
@@ -30,6 +31,12 @@ enum {
ENUM_BIT(EXTENT_FLAG_LOGGING),
/* This em is merged from two or more physically adjacent ems */
ENUM_BIT(EXTENT_FLAG_MERGED),
+ /*
+ * The real on-disk extent allocation is delayed until bio submission.
+ * For now it's only a placeholder with EXTENT_MAP_DELAYED as
+ * its disk_bytenr.
+ */
+ ENUM_BIT(EXTENT_FLAG_DELAYED),
};
/*
diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
index e9ed3e31fa88..814bc456176e 100644
--- a/fs/btrfs/inode.c
+++ b/fs/btrfs/inode.c
@@ -7676,10 +7676,51 @@ struct extent_map *btrfs_create_io_em(struct btrfs_inode *inode, u64 start,
return ERR_PTR(ret);
}
- /* em got 2 refs now, callers needs to do btrfs_free_extent_map once. */
+ /* em got 2 refs now, callers need to do btrfs_free_extent_map once. */
return em;
}
+struct extent_map *btrfs_create_delayed_em(struct btrfs_inode *inode,
+ u64 start, u32 length)
+{
+ struct extent_map *em;
+ int ret;
+
+ em = btrfs_alloc_extent_map();
+ if (!em)
+ return ERR_PTR(-ENOMEM);
+
+ em->start = start;
+ em->len = length;
+ em->disk_bytenr = EXTENT_MAP_DELAYED;
+ em->disk_num_bytes = 0;
+ em->ram_bytes = 0;
+ em->generation = -1;
+ em->offset = 0;
+ em->flags = EXTENT_FLAG_DELAYED | EXTENT_FLAG_PINNED;
+
+ ret = btrfs_replace_extent_map_range(inode, em, true);
+ if (ret) {
+ btrfs_free_extent_map(em);
+ return ERR_PTR(ret);
+ }
+
+ /* em got 2 refs now, callers need to do btrfs_free_extent_map once. */
+ return em;
+}
+
+void btrfs_submit_delayed_write(struct btrfs_bio *bbio)
+{
+ ASSERT(bbio->is_delayed);
+
+ /*
+ * Not yet implemented, and should not hit this path as we have no
+ * caller to create delayed extent map.
+ */
+ ASSERT(0);
+ bio_put(&bbio->bio);
+}
+
/*
* For release_folio() and invalidate_folio() we have a race window where
* folio_end_writeback() is called but the subpage spinlock is not yet released.
--
2.55.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v7 3/6] btrfs: introduce the skeleton of delayed bbio endio function
2026-09-25 9:07 [PATCH v7 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 1/6] btrfs: add delayed ordered extent support Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 2/6] btrfs: add skeleton for delayed btrfs bio Qu Wenruo
@ 2026-09-25 9:07 ` Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 4/6] btrfs: introduce compression for delayed bbio Qu Wenruo
` (2 subsequent siblings)
5 siblings, 0 replies; 9+ messages in thread
From: Qu Wenruo @ 2026-09-25 9:07 UTC (permalink / raw)
To: linux-btrfs
A delayed bbio will not be directly submitted, but queued into a
workqueue, to perform compression there.
There is also a new workqueue, delayed_write_workers, for this workload.
It will eventually replace the existing delalloc_workers.
The compression and uncompressed fallback are not implemented in this
patch.
Only the main endio function and a helper to queue workload into a
workqueue are implemented.
The endio function is mostly the same as end_bbio_data_write(), except
for the extra memory allocation/freeing for the bbio->private.
Another point to note is that, we reuse btrfs_bio::pending_ios to track
the lifespan of the parent delayed bbio.
We can reuse it because the delayed bbio is never submitted for IO, thus
it will never be split by btrfs, so we can reuse that atomic to avoid
extra space inside delayed_bio_private.
Signed-off-by: Qu Wenruo <wqu@suse.com>
---
fs/btrfs/disk-io.c | 8 +++--
fs/btrfs/fs.h | 5 +++
fs/btrfs/inode.c | 84 +++++++++++++++++++++++++++++++++++++++++++---
3 files changed, 90 insertions(+), 7 deletions(-)
diff --git a/fs/btrfs/disk-io.c b/fs/btrfs/disk-io.c
index 90cf0647c47e..ea85cfc25774 100644
--- a/fs/btrfs/disk-io.c
+++ b/fs/btrfs/disk-io.c
@@ -1774,6 +1774,8 @@ static void btrfs_stop_all_workers(struct btrfs_fs_info *fs_info)
if (fs_info->fixup_workers)
destroy_workqueue(fs_info->fixup_workers);
btrfs_destroy_workqueue(fs_info->delalloc_workers);
+ if (fs_info->delayed_write_workers)
+ destroy_workqueue(fs_info->delayed_write_workers);
btrfs_destroy_workqueue(fs_info->workers);
if (fs_info->endio_workers)
destroy_workqueue(fs_info->endio_workers);
@@ -1972,7 +1974,8 @@ static int btrfs_init_workqueues(struct btrfs_fs_info *fs_info)
fs_info->delalloc_workers =
btrfs_alloc_workqueue(fs_info, "delalloc",
flags, max_active, 2);
-
+ fs_info->delayed_write_workers =
+ alloc_workqueue("btrfs-delayed-write", flags, max_active);
fs_info->flush_workers =
btrfs_alloc_workqueue(fs_info, "flush_delalloc",
flags, max_active, 0);
@@ -2003,7 +2006,7 @@ static int btrfs_init_workqueues(struct btrfs_fs_info *fs_info)
fs_info->discard_ctl.discard_workers =
alloc_ordered_workqueue("btrfs-discard", WQ_FREEZABLE);
- if (!(fs_info->workers &&
+ if (!(fs_info->workers && fs_info->delayed_write_workers &&
fs_info->delalloc_workers && fs_info->flush_workers &&
fs_info->endio_workers && fs_info->endio_meta_workers &&
fs_info->endio_write_workers &&
@@ -4434,6 +4437,7 @@ void __cold close_ctree(struct btrfs_fs_info *fs_info)
* when we call kthread_stop().
*/
btrfs_flush_workqueue(fs_info->delalloc_workers);
+ flush_workqueue(fs_info->delayed_write_workers);
/*
* We can have ordered extents getting their last reference dropped from
diff --git a/fs/btrfs/fs.h b/fs/btrfs/fs.h
index 3eba8438593c..72e3ef1de423 100644
--- a/fs/btrfs/fs.h
+++ b/fs/btrfs/fs.h
@@ -707,6 +707,11 @@ struct btrfs_fs_info {
*/
struct btrfs_workqueue *workers;
struct btrfs_workqueue *delalloc_workers;
+ /*
+ * This is for delayed writes, for now only compressed writes
+ * on non-zoned experimental builds utilize this.
+ */
+ struct workqueue_struct *delayed_write_workers;
struct btrfs_workqueue *flush_workers;
struct workqueue_struct *endio_workers;
struct workqueue_struct *endio_meta_workers;
diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
index 814bc456176e..0de121f50562 100644
--- a/fs/btrfs/inode.c
+++ b/fs/btrfs/inode.c
@@ -97,6 +97,11 @@ struct data_reloc_warn {
int mirror_num;
};
+struct delayed_bio_private {
+ struct work_struct work;
+ struct btrfs_bio *delayed_bbio;
+};
+
/*
* For the file_extent_tree, we want to hold the inode lock when we lookup and
* update the disk_i_size, but lockdep will complain because our io_tree we hold
@@ -7709,18 +7714,87 @@ struct extent_map *btrfs_create_delayed_em(struct btrfs_inode *inode,
return em;
}
-void btrfs_submit_delayed_write(struct btrfs_bio *bbio)
+static void run_delayed_bbio(struct work_struct *work)
{
- ASSERT(bbio->is_delayed);
+ struct delayed_bio_private *dbp = container_of(work, struct delayed_bio_private, work);
+ struct btrfs_bio *parent = dbp->delayed_bbio;
+
+ /* Parent OE must be set. */
+ ASSERT(parent->ordered && test_bit(BTRFS_ORDERED_DELAYED, &parent->ordered->flags));
+
+ /* Compressed and uncompressed fallback is not yet implemented. */
+ ASSERT(0);
/*
- * Not yet implemented, and should not hit this path as we have no
- * caller to create delayed extent map.
+ * Any real compressed/uncompressed bios have increased
+ * parent->pending_ios, the last caller of btrfs_bio_end_io()
+ * will execute the real endio function.
*/
- ASSERT(0);
+ btrfs_bio_end_io(parent, BLK_STS_OK);
+}
+
+static void end_bbio_delayed(struct btrfs_bio *bbio)
+{
+ struct delayed_bio_private *dbp = bbio->private;
+ struct btrfs_inode *inode = bbio->inode;
+ struct btrfs_fs_info *fs_info = inode->root->fs_info;
+ struct folio_iter fi;
+ const u32 bio_size = bio_get_size(&bbio->bio);
+ const bool uptodate = bbio->status == BLK_STS_OK;
+
+ ASSERT(bbio->is_delayed);
+
+ bio_for_each_folio_all(fi, &bbio->bio) {
+ u64 start = folio_pos(fi.folio) + fi.offset;
+ u32 len = fi.length;
+
+ btrfs_folio_clear_writeback(fs_info, fi.folio, start, len);
+ }
+ if (!uptodate)
+ mapping_set_error(inode->vfs_inode.i_mapping,
+ blk_status_to_errno(bbio->status));
+ btrfs_mark_ordered_io_finished(inode, bbio->file_offset, bio_size, uptodate);
+ kfree(dbp);
bio_put(&bbio->bio);
}
+void btrfs_submit_delayed_write(struct btrfs_bio *bbio)
+{
+ const struct btrfs_fs_info *fs_info = bbio->inode->root->fs_info;
+ struct delayed_bio_private *dbp;
+
+ ASSERT(bbio->is_delayed);
+
+ bbio->end_io = end_bbio_delayed;
+ dbp = kzalloc_obj(*dbp, GFP_NOFS);
+ if (!dbp) {
+ btrfs_bio_end_io(bbio, errno_to_blk_status(-ENOMEM));
+ return;
+ }
+ dbp->delayed_bbio = bbio;
+ bbio->private = dbp;
+ /*
+ * TODO: find a way to properly allow sequential extent allocation.
+ *
+ * Currently compression and submission happen in the same workload.
+ * We want to run compression in parallel but run allocation sequentially.
+ *
+ * The existing btrfs workqueue will execute the sequential workload
+ * twice, the second one to free the structure, thus the btrfs_work
+ * structure can have a lifespan longer than the delayed bbio.
+ *
+ * But the current structure's lifespan is bounded to the original
+ * bbio, if using btrfs workqueue, we need extra lifespan management
+ * to ensure the work structure doesn't disappear when the original
+ * bbio is finished.
+ *
+ * This means we have to either delay the original bbio's finish,
+ * or have a very complex synchronization.
+ */
+ INIT_WORK(&dbp->work, run_delayed_bbio);
+ queue_work(fs_info->delayed_write_workers, &dbp->work);
+}
+
/*
* For release_folio() and invalidate_folio() we have a race window where
* folio_end_writeback() is called but the subpage spinlock is not yet released.
--
2.55.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v7 4/6] btrfs: introduce compression for delayed bbio
2026-09-25 9:07 [PATCH v7 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
` (2 preceding siblings ...)
2026-09-25 9:07 ` [PATCH v7 3/6] btrfs: introduce the skeleton of delayed bbio endio function Qu Wenruo
@ 2026-09-25 9:07 ` Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 5/6] btrfs: implement uncompressed fallback " Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 6/6] btrfs: enable experimental delayed compression support Qu Wenruo
5 siblings, 0 replies; 9+ messages in thread
From: Qu Wenruo @ 2026-09-25 9:07 UTC (permalink / raw)
To: linux-btrfs
The compressed write path inside a delayed bbio is mostly the same as
regular compression, but with some differences:
- The error handling should not touch folio flags
It will be handled by the parent delayed bbio.
And those folios already have WRITEBACK flag set, not the LOCKED flag
of the async submission path.
- The error handling should only be called once
If we hit a critical error, e.g. unable to create an OE, the range
will be cleaned up by try_submit_compressed(), thus the parent OE
should not clean up that range again.
- A successful compression will lead to a child compressed bio
That compressed bio will be properly submitted, and if there are no
more pending IOs of the delayed bbio, end the delayed bbio.
There is a minor note, since we're going through the regular
extent_writepage_io() path, we can have multiple bbios for the same
delayed ordered extent.
This means we may have a slightly lower compression ratio if for
whatever reason the writeback path chooses to submit a smaller bio.
- No sequential execution of data extent reservation
The existing async thread has one quirk related to the ordered
function execution, which is not suitable for this call site.
After the compressed bio is submitted, we can no longer touch the
child compressed bio (it could finish immediately and also finish the
parent delayed bbio).
Meanwhile the async ordered function needs different entries to handle
the workload and free involved structures.
These will be the major changes compared to the existing compressed
write.
- No extra blkcg/write flag inheritance
This is a missing feature compared to the older async compression.
This needs to be implemented in the future, but priority shouldn't be
that high.
Signed-off-by: Qu Wenruo <wqu@suse.com>
---
fs/btrfs/inode.c | 165 ++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 163 insertions(+), 2 deletions(-)
diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
index 0de121f50562..2a131e9283dd 100644
--- a/fs/btrfs/inode.c
+++ b/fs/btrfs/inode.c
@@ -7714,23 +7714,184 @@ struct extent_map *btrfs_create_delayed_em(struct btrfs_inode *inode,
return em;
}
+static void end_bbio_delayed_compressed(struct btrfs_bio *bbio)
+{
+ struct delayed_bio_private *dbp = bbio->private;
+ struct btrfs_bio *parent = dbp->delayed_bbio;
+ struct folio_iter fi;
+
+ bio_for_each_folio_all(fi, &bbio->bio)
+ btrfs_free_compr_folio(fi.folio);
+ btrfs_bio_end_io(parent, bbio->bio.bi_status);
+ bio_put(&bbio->bio);
+}
+
+/*
+ * If we failed to allocate a child OE for a range, that range will be cleaned
+ * up by the caller, and the parent OE should not clean it up again.
+ * So here we need to clear the bits in the parent OE.
+ */
+static void clear_cleanup_bitmap(struct btrfs_ordered_extent *parent,
+ u64 file_offset, u32 len)
+{
+ const u32 sectorbits = parent->inode->root->fs_info->sectorsize_bits;
+
+ /*
+ * There can be multiple child ranges for the parent delayed OE.
+ *
+ * Re-use inode->ordered_tree_lock to prevent race.
+ */
+ spin_lock(&parent->inode->ordered_tree_lock);
+ bitmap_clear(parent->child_cleanup_bitmap,
+ (file_offset - parent->file_offset) >> sectorbits,
+ len >> sectorbits);
+ spin_unlock(&parent->inode->ordered_tree_lock);
+}
+
+/*
+ * Return 0 if a compressed write is submitted.
+ * Return >0 if a compressed write is not submitted and we need to fall back to
+ * uncompressed write.
+ * Return <0 if we hit a critical error, and the reserved space is already cleaned
+ * up.
+ */
+static int try_submit_compressed(struct btrfs_bio *parent)
+{
+ struct delayed_bio_private *dbp = parent->private;
+ struct btrfs_inode *inode = parent->inode;
+ struct btrfs_fs_info *fs_info = inode->root->fs_info;
+ struct btrfs_key ins;
+ struct compressed_bio *cb;
+ struct extent_state *cached = NULL;
+ struct extent_map *em;
+ struct btrfs_ordered_extent *ordered;
+ struct btrfs_file_extent file_extent;
+ u64 alloc_hint;
+ const u32 len = bio_get_size(&parent->bio);
+ const u64 fileoff = parent->file_offset;
+ const u64 end = fileoff + len - 1;
+ u32 compressed_size;
+ int compress_type = fs_info->compress_type;
+ int compress_level = fs_info->compress_level;
+ int ret;
+
+ if (!btrfs_inode_can_compress(inode) ||
+ !inode_need_compress(inode, fileoff, end, false))
+ return 1;
+
+ if (inode->defrag_compress > 0 &&
+ inode->defrag_compress < BTRFS_NR_COMPRESS_TYPES) {
+ compress_type = inode->defrag_compress;
+ compress_level = inode->defrag_compress_level;
+ } else if (inode->prop_compress) {
+ compress_type = inode->prop_compress;
+ }
+ cb = btrfs_compress_bio(inode, fileoff, len, compress_type,
+ compress_level, 0);
+ if (IS_ERR(cb))
+ return 1;
+
+ round_up_last_block(cb, fs_info->sectorsize);
+ compressed_size = cb->bbio.bio.bi_iter.bi_size;
+ /* If no space is saved, abort and mark the inode NOCOMPRESS. */
+ if (compressed_size >= len) {
+ if (!btrfs_test_opt(fs_info, FORCE_COMPRESS) &&
+ !inode->prop_compress)
+ inode->flags |= BTRFS_INODE_NOCOMPRESS;
+ cleanup_compressed_bio(cb);
+ return 1;
+ }
+
+ alloc_hint = btrfs_get_extent_allocation_hint(inode, fileoff, len);
+ ret = btrfs_reserve_extent(inode->root, len,
+ compressed_size, compressed_size,
+ 0, alloc_hint, &ins, true, true);
+ if (ret < 0) {
+ cleanup_compressed_bio(cb);
+ return 1;
+ }
+ btrfs_lock_extent(&inode->io_tree, fileoff, end, &cached);
+ file_extent.disk_bytenr = ins.objectid;
+ file_extent.disk_num_bytes = ins.offset;
+ file_extent.ram_bytes = len;
+ file_extent.num_bytes = len;
+ file_extent.offset = 0;
+ file_extent.compression = cb->compress_type;
+
+ cb->bbio.bio.bi_iter.bi_sector = ins.objectid >> SECTOR_SHIFT;
+ em = btrfs_create_io_em(inode, fileoff, &file_extent, BTRFS_ORDERED_COMPRESSED);
+ if (IS_ERR(em)) {
+ ret = PTR_ERR(em);
+ goto out_free_reserve;
+ }
+ btrfs_free_extent_map(em);
+
+ ordered = btrfs_alloc_ordered_extent(inode, fileoff, &file_extent,
+ 1U << BTRFS_ORDERED_COMPRESSED);
+ if (IS_ERR(ordered)) {
+ ret = PTR_ERR(ordered);
+ goto out_free_reserve;
+ }
+ cb->bbio.ordered = ordered;
+ btrfs_dec_block_group_reservations(fs_info, ins.objectid);
+ btrfs_unlock_extent(&inode->io_tree, fileoff, end, &cached);
+
+ cb->bbio.end_io = end_bbio_delayed_compressed;
+ cb->bbio.private = dbp;
+ atomic_inc(&parent->pending_ios);
+ btrfs_submit_bbio(&cb->bbio, 0);
+ return 0;
+
+out_free_reserve:
+ /*
+ * We may have not created the real EM yet, but the range will
+ * not be cleaned up by the delayed parent OE.
+ * So we still need to drop the parent delayed em for the failed range.
+ */
+ btrfs_drop_extent_map_range(inode, fileoff, end, false);
+ btrfs_dec_block_group_reservations(fs_info, ins.objectid);
+ btrfs_free_reserved_extent(fs_info, ins.objectid, ins.offset, true);
+ /*
+ * The range has EXTENT_DELALLOC cleared already, so
+ * EXTENT_DO_ACCOUNTING is not working, and bytes_reserved
+ * already decreased by the above btrfs_free_reserved_extent().
+ *
+ * Here we only need to release metadata and qgroup reserved space.
+ */
+ btrfs_delalloc_release_metadata(inode, len, true);
+ btrfs_qgroup_free_data(inode, NULL, fileoff, len, NULL);
+ btrfs_clear_extent_bit(&inode->io_tree, fileoff, end,
+ EXTENT_LOCKED | EXTENT_DELALLOC_NEW | EXTENT_DEFRAG,
+ &cached);
+ cleanup_compressed_bio(cb);
+ clear_cleanup_bitmap(parent->ordered, parent->file_offset, len);
+ ASSERT(ret < 0);
+ return ret;
+}
+
static void run_delayed_bbio(struct work_struct *work)
{
struct delayed_bio_private *dbp = container_of(work, struct delayed_bio_private, work);
struct btrfs_bio *parent = dbp->delayed_bbio;
+ int ret;
/* Parent OE must be set. */
ASSERT(parent->ordered && test_bit(BTRFS_ORDERED_DELAYED, &parent->ordered->flags));
- /* Compressed and uncompressed fallback is not yet implemented. */
+ ret = try_submit_compressed(parent);
+ if (ret <= 0)
+ goto finish;
+
+ /* Uncompressed fallback is not yet implemented. */
ASSERT(0);
+finish:
/*
* Any real compressed/uncompressed bios have increased
* parent->pending_ios, the last caller of btrfs_bio_end_io()
* will execute the real endio function.
*/
- btrfs_bio_end_io(parent, BLK_STS_OK);
+ btrfs_bio_end_io(parent, errno_to_blk_status(ret));
}
static void end_bbio_delayed(struct btrfs_bio *bbio)
--
2.55.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v7 5/6] btrfs: implement uncompressed fallback for delayed bbio
2026-09-25 9:07 [PATCH v7 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
` (3 preceding siblings ...)
2026-09-25 9:07 ` [PATCH v7 4/6] btrfs: introduce compression for delayed bbio Qu Wenruo
@ 2026-09-25 9:07 ` Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 6/6] btrfs: enable experimental delayed compression support Qu Wenruo
5 siblings, 0 replies; 9+ messages in thread
From: Qu Wenruo @ 2026-09-25 9:07 UTC (permalink / raw)
To: linux-btrfs
When compression fails (either bad ratio, fragmented free space) we have
to fall back to uncompressed writes.
The uncompressed fallback is mostly the same as cow_file_range() but
with some changes:
- Endio function is slightly different from the compressed path
Only in the folio freeing handling.
- Uncompressed fallback error handling
Since at this stage, the folios already have WRITEBACK flag set, we do
not need to do the usual page unlock/end writeback, but just free the
reserved space and call it a day.
Signed-off-by: Qu Wenruo <wqu@suse.com>
---
fs/btrfs/inode.c | 164 ++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 162 insertions(+), 2 deletions(-)
diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
index 2a131e9283dd..fcf67a2e118d 100644
--- a/fs/btrfs/inode.c
+++ b/fs/btrfs/inode.c
@@ -7869,10 +7869,159 @@ static int try_submit_compressed(struct btrfs_bio *parent)
return ret;
}
+static void end_bbio_delayed_uncompressed(struct btrfs_bio *bbio)
+{
+ struct delayed_bio_private *dbp = bbio->private;
+ struct btrfs_bio *parent = dbp->delayed_bbio;
+ struct folio_iter fi;
+
+ bio_for_each_folio_all(fi, &bbio->bio)
+ folio_put(fi.folio);
+ btrfs_bio_end_io(parent, bbio->bio.bi_status);
+ bio_put(&bbio->bio);
+}
+
+static struct btrfs_bio *child_bbio_from_page_cache(struct btrfs_bio *parent,
+ u64 fileoff, u32 len)
+{
+ struct btrfs_inode *inode = parent->inode;
+ struct btrfs_fs_info *fs_info = inode->root->fs_info;
+ struct address_space *mapping = inode->vfs_inode.i_mapping;
+ struct btrfs_bio *bbio;
+ struct folio_iter fi;
+ u64 cur = fileoff;
+ int ret;
+
+ bbio = btrfs_bio_alloc(len >> fs_info->sectorsize_bits, REQ_OP_WRITE,
+ inode, fileoff, end_bbio_delayed_uncompressed,
+ parent->private);
+
+ while (cur < fileoff + len) {
+ struct folio *folio;
+ u32 cur_len;
+ bool queued;
+
+ folio = filemap_get_folio(mapping, cur >> PAGE_SHIFT);
+ if (IS_ERR(folio)) {
+ ret = PTR_ERR(folio);
+ goto error;
+ }
+ cur_len = min_t(u64, folio_next_pos(folio), fileoff + len) - cur;
+ queued = bio_add_folio(&bbio->bio, folio, cur_len,
+ offset_in_folio(folio, cur));
+ /* There should be enough slots for the bio. */
+ if (WARN_ON(!queued)) {
+ ret = -EIO;
+ folio_put(folio);
+ goto error;
+ }
+ cur += cur_len;
+ }
+
+ return bbio;
+error:
+ bio_for_each_folio_all(fi, &bbio->bio)
+ folio_put(fi.folio);
+ bio_put(&bbio->bio);
+ return ERR_PTR(ret);
+}
+
+static int submit_one_uncompressed_range(struct btrfs_bio *parent, struct btrfs_key *ins,
+ struct extent_state **cached, u64 file_offset,
+ u32 num_bytes, u64 alloc_hint, u32 *ret_alloc_size)
+{
+ struct btrfs_inode *inode = parent->inode;
+ struct btrfs_root *root = inode->root;
+ struct btrfs_fs_info *fs_info = root->fs_info;
+ struct btrfs_ordered_extent *ordered;
+ struct btrfs_file_extent file_extent;
+ struct btrfs_bio *child = NULL;
+ struct extent_map *em;
+ u64 cur_end;
+ u32 cur_len = 0;
+ int ret;
+
+ ret = btrfs_reserve_extent(root, num_bytes, num_bytes, fs_info->sectorsize,
+ 0, alloc_hint, ins, true, true);
+ if (ret < 0)
+ return ret;
+
+ cur_len = ins->offset;
+ cur_end = file_offset + cur_len - 1;
+
+ file_extent.disk_bytenr = ins->objectid;
+ file_extent.disk_num_bytes = ins->offset;
+ file_extent.num_bytes = ins->offset;
+ file_extent.ram_bytes = ins->offset;
+ file_extent.offset = 0;
+ file_extent.compression = BTRFS_COMPRESS_NONE;
+
+ btrfs_lock_extent(&inode->io_tree, file_offset, cur_end, cached);
+
+ child = child_bbio_from_page_cache(parent, file_offset, cur_len);
+ if (IS_ERR(child)) {
+ ret = PTR_ERR(child);
+ child = NULL;
+ goto free_reserved;
+ }
+
+ em = btrfs_create_io_em(inode, file_offset, &file_extent, BTRFS_ORDERED_REGULAR);
+ if (IS_ERR(em)) {
+ ret = PTR_ERR(em);
+ goto free_reserved;
+ }
+ btrfs_free_extent_map(em);
+ ordered = btrfs_alloc_ordered_extent(inode, file_offset, &file_extent,
+ 1U << BTRFS_ORDERED_REGULAR);
+ if (IS_ERR(ordered)) {
+ ret = PTR_ERR(ordered);
+ goto free_reserved;
+ }
+ btrfs_dec_block_group_reservations(fs_info, ins->objectid);
+ btrfs_unlock_extent(&inode->io_tree, file_offset, cur_end, cached);
+
+ child->ordered = ordered;
+ child->private = parent->private;
+ child->end_io = end_bbio_delayed_uncompressed;
+ child->bio.bi_iter.bi_sector = ins->objectid >> SECTOR_SHIFT;
+ atomic_inc(&parent->pending_ios);
+ btrfs_submit_bbio(child, 0);
+ *ret_alloc_size = cur_len;
+ return 0;
+
+free_reserved:
+ if (child) {
+ struct folio_iter fi;
+
+ bio_for_each_folio_all(fi, &child->bio)
+ folio_put(fi.folio);
+ bio_put(&child->bio);
+ }
+ /* Check the comments at end of try_submit_compressed(). */
+ btrfs_drop_extent_map_range(inode, file_offset, cur_end, false);
+ btrfs_qgroup_free_data(inode, NULL, file_offset, cur_len, NULL);
+ btrfs_dec_block_group_reservations(fs_info, ins->objectid);
+ btrfs_free_reserved_extent(fs_info, ins->objectid, ins->offset, true);
+ btrfs_delalloc_release_metadata(inode, cur_len, true);
+ btrfs_clear_extent_bit(&inode->io_tree, file_offset, cur_end,
+ EXTENT_LOCKED | EXTENT_DELALLOC_NEW | EXTENT_DEFRAG,
+ cached);
+ clear_cleanup_bitmap(parent->ordered, file_offset, cur_len);
+ ASSERT(ret != -EAGAIN);
+ return ret;
+}
+
static void run_delayed_bbio(struct work_struct *work)
{
struct delayed_bio_private *dbp = container_of(work, struct delayed_bio_private, work);
struct btrfs_bio *parent = dbp->delayed_bbio;
+ struct btrfs_key ins;
+ struct extent_state *cached = NULL;
+ const u32 uncompressed_size = bio_get_size(&parent->bio);
+ const u64 start = parent->file_offset;
+ const u64 end = start + uncompressed_size - 1;
+ u64 cur = start;
+ u64 alloc_hint;
int ret;
/* Parent OE must be set. */
@@ -7882,8 +8031,19 @@ static void run_delayed_bbio(struct work_struct *work)
if (ret <= 0)
goto finish;
- /* Uncompressed fallback is not yet implemented. */
- ASSERT(0);
+ alloc_hint = btrfs_get_extent_allocation_hint(parent->inode, start,
+ uncompressed_size);
+ while (cur < end) {
+ u32 cur_len;
+
+ ret = submit_one_uncompressed_range(parent, &ins, &cached,
+ cur, end + 1 - cur,
+ alloc_hint, &cur_len);
+ if (ret < 0)
+ goto finish;
+ cur += cur_len;
+ alloc_hint = ins.objectid + ins.offset;
+ }
finish:
/*
--
2.55.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v7 6/6] btrfs: enable experimental delayed compression support
2026-09-25 9:07 [PATCH v7 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
` (4 preceding siblings ...)
2026-09-25 9:07 ` [PATCH v7 5/6] btrfs: implement uncompressed fallback " Qu Wenruo
@ 2026-09-25 9:07 ` Qu Wenruo
5 siblings, 0 replies; 9+ messages in thread
From: Qu Wenruo @ 2026-09-25 9:07 UTC (permalink / raw)
To: linux-btrfs
Instead of the existing async submission path, the new delayed bbio will
handle compressed writes by:
- Allocating delayed em/oe at run_delalloc_*() time
Thus no data extent is reserved at that time.
- Delayed bbio will be assembled at extent_writepage_io() time
- Delayed bbio will be intercepted just before submission
Which will run compression (or fall back to uncompressed writes) in
workqueue.
Data extents will only be reserved at that time, and the delayed em
will be replaced by real ones.
Meanwhile the real OE will be added as a child of the parent delayed
OE, and when the parent OE finishes, the child OEs will be finished
with their file extents inserted.
This has some benefits:
- Higher concurrency
The old async submission held the folio and io tree range locked.
This meant we cannot even read the uptodate folio.
Furthermore although the compressed write is queued into a workqueue
for submission and extent_writepage_io() will skip the compressed
range, when we need to write the next folio of the compressed range,
we will need to wait for the folio to be unlocked.
This makes async submission less async.
- Future DONTCACHE writes support
We do not support DONTCACHE because that feature requires writeback path
to clear the folio dirty and submit them sequentially.
Meanwhile async submission makes the writeback async, breaking the
sequential submission requirement.
This is also why we need complex per-block tracking for writeback
flags, while iomap only requires a counter tracking.
With the new delayed compression, the lifespan of a folio aligns with
DONTCACHE and iomap.
However there is also a change to the fsync() behavior:
- Now an inode with compressed write will require full sync
We cannot do any fast sync, as that will grab the info from OE
directly, but since delayed OE is just a placeholder, all the OE info
cannot be trusted until the parent OE is finished.
There is an extra handling for defrag, where we have to write and wait
for the defrag range.
This is to avoid clearing inode->defrag_compress before the compression
starts.
Finally this new experimental feature is not yet compatible with zoned,
so for zoned btrfs it will go through the old compression path.
Signed-off-by: Qu Wenruo <wqu@suse.com>
---
fs/btrfs/defrag.c | 27 ++++++++++++++---
fs/btrfs/extent_io.c | 5 +++-
fs/btrfs/inode.c | 71 ++++++++++++++++++++++++++++++++++++++++++--
3 files changed, 95 insertions(+), 8 deletions(-)
diff --git a/fs/btrfs/defrag.c b/fs/btrfs/defrag.c
index 6ec5dd760d42..2ee447b3cfe3 100644
--- a/fs/btrfs/defrag.c
+++ b/fs/btrfs/defrag.c
@@ -1349,6 +1349,7 @@ int btrfs_defrag_file(struct btrfs_inode *inode, struct file_ra_state *ra,
struct btrfs_fs_info *fs_info = inode->root->fs_info;
unsigned long sectors_defragged = 0;
u64 isize = i_size_read(&inode->vfs_inode);
+ const u64 start = round_down(range->start, fs_info->sectorsize);
u64 cur;
u64 last_byte;
bool do_compress = (range->flags & BTRFS_DEFRAG_RANGE_COMPRESS);
@@ -1400,7 +1401,7 @@ int btrfs_defrag_file(struct btrfs_inode *inode, struct file_ra_state *ra,
}
/* Align the range */
- cur = round_down(range->start, fs_info->sectorsize);
+ cur = start;
last_byte = round_up(last_byte, fs_info->sectorsize) - 1;
/*
@@ -1471,10 +1472,28 @@ int btrfs_defrag_file(struct btrfs_inode *inode, struct file_ra_state *ra,
* need to be written back immediately.
*/
if (range->flags & BTRFS_DEFRAG_RANGE_START_IO) {
- filemap_flush(inode->vfs_inode.i_mapping);
- if (test_bit(BTRFS_INODE_HAS_ASYNC_EXTENT,
- &inode->runtime_flags))
+ /*
+ * For experimental delayed writeback, we must wait
+ * for the range to be fully written back before
+ * clearing inode->defrag_compress.
+ *
+ * Regular filemap_flush() will only start writeback,
+ * which will only create delayed OEs. But the real
+ * compression is happening later.
+ * This means if we just flush but not wait for writeback,
+ * the inode->defrag_compress clearing can race with
+ * compression, so the requested compression algorithm
+ * is not applied.
+ */
+ if (IS_ENABLED(CONFIG_BTRFS_EXPERIMENTAL) && !btrfs_is_zoned(fs_info)) {
+ filemap_write_and_wait_range(inode->vfs_inode.i_mapping,
+ start, last_byte);
+ } else {
filemap_flush(inode->vfs_inode.i_mapping);
+ if (test_bit(BTRFS_INODE_HAS_ASYNC_EXTENT,
+ &inode->runtime_flags))
+ filemap_flush(inode->vfs_inode.i_mapping);
+ }
}
if (range->compress_type == BTRFS_COMPRESS_LZO)
btrfs_set_fs_incompat(fs_info, COMPRESS_LZO);
diff --git a/fs/btrfs/extent_io.c b/fs/btrfs/extent_io.c
index a822f6635eba..970548df2dbc 100644
--- a/fs/btrfs/extent_io.c
+++ b/fs/btrfs/extent_io.c
@@ -899,8 +899,11 @@ static int submit_one_block(struct btrfs_bio_ctrl *bio_ctrl,
* If we have accumulated decent amount of IO, send it to the block
* layer so that IO can run while we are accumulating more folios to
* write.
+ *
+ * This doesn't apply to delayed bbio which is going to be
+ * compressed.
*/
- else if (bio_ctrl->wbc &&
+ else if (bio_ctrl->wbc && !bio_ctrl->bbio->is_delayed &&
bio_ctrl->bbio->bio.bi_iter.bi_size >= fs_info->writeback_bio_size)
submit_one_bio(bio_ctrl);
return 0;
diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
index fcf67a2e118d..b3e4c6426609 100644
--- a/fs/btrfs/inode.c
+++ b/fs/btrfs/inode.c
@@ -1659,6 +1659,68 @@ static bool run_delalloc_compressed(struct btrfs_inode *inode,
return true;
}
+static int run_delalloc_delayed(struct btrfs_inode *inode, struct folio *locked_folio,
+ u64 start, u64 end)
+{
+ struct btrfs_root *root = inode->root;
+ struct btrfs_fs_info *fs_info = root->fs_info;
+ struct extent_state *cached = NULL;
+ u64 cur = start;
+ int ret;
+
+ if (btrfs_is_shutdown(fs_info)) {
+ ret = -EIO;
+ goto error;
+ }
+ btrfs_lock_extent(&inode->io_tree, start, end, &cached);
+ while (cur < end) {
+ struct extent_map *em;
+ struct btrfs_ordered_extent *oe;
+ u32 cur_len = min_t(u64, end + 1 - cur, BTRFS_MAX_COMPRESSED);
+
+ em = btrfs_create_delayed_em(inode, cur, cur_len);
+ if (IS_ERR(em)) {
+ ret = PTR_ERR(em);
+ goto error;
+ }
+ btrfs_free_extent_map(em);
+ oe = btrfs_alloc_delayed_ordered_extent(inode, cur, cur_len);
+ if (IS_ERR(oe)) {
+ btrfs_drop_extent_map_range(inode, cur, cur + cur_len - 1, false);
+ ret = PTR_ERR(oe);
+ goto error;
+ }
+ btrfs_put_ordered_extent(oe);
+
+ cur += cur_len;
+ }
+ /*
+ * Force the next fsync to wait for any delayed OE to finish,
+ * as the fast path cannot handle delayed OE correctly.
+ */
+ btrfs_set_inode_full_sync(inode);
+ extent_clear_unlock_delalloc(inode, start, end, locked_folio, &cached,
+ EXTENT_LOCKED | EXTENT_DELALLOC,
+ PAGE_UNLOCK);
+ return 0;
+error:
+ if (start < cur) {
+ btrfs_drop_extent_map_range(inode, start, cur - 1, false);
+ btrfs_cleanup_ordered_extents(inode, start, cur - start);
+ extent_clear_unlock_delalloc(inode, start, cur - 1, locked_folio, &cached,
+ EXTENT_LOCKED | EXTENT_DELALLOC,
+ PAGE_UNLOCK | PAGE_START_WRITEBACK | PAGE_END_WRITEBACK);
+ }
+ if (cur < end) {
+ extent_clear_unlock_delalloc(inode, cur, end, locked_folio, &cached,
+ EXTENT_LOCKED | EXTENT_DELALLOC | EXTENT_DELALLOC_NEW |
+ EXTENT_DEFRAG | EXTENT_DO_ACCOUNTING,
+ PAGE_UNLOCK | PAGE_START_WRITEBACK | PAGE_END_WRITEBACK);
+ btrfs_qgroup_free_data(inode, NULL, cur, end + 1 - cur, NULL);
+ }
+ return ret;
+}
+
/*
* Run the delalloc range from start to end, and write back any dirty pages
* covered by the range.
@@ -2454,9 +2516,12 @@ int btrfs_run_delalloc_range(struct btrfs_inode *inode, struct folio *locked_fol
return run_delalloc_nocow(inode, locked_folio, start, end);
if (btrfs_inode_can_compress(inode) &&
- inode_need_compress(inode, start, end, false) &&
- run_delalloc_compressed(inode, locked_folio, start, end, wbc))
- return 1;
+ inode_need_compress(inode, start, end, false)) {
+ if (IS_ENABLED(CONFIG_BTRFS_EXPERIMENTAL) && !zoned)
+ return run_delalloc_delayed(inode, locked_folio, start, end);
+ else if (run_delalloc_compressed(inode, locked_folio, start, end, wbc))
+ return 1;
+ }
if (zoned)
return run_delalloc_cow(inode, locked_folio, start, end, wbc, true);
--
2.55.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* Re: [PATCH v7 1/6] btrfs: add delayed ordered extent support
2026-09-25 9:07 ` [PATCH v7 1/6] btrfs: add delayed ordered extent support Qu Wenruo
@ 2026-09-25 9:32 ` Miquel Sabaté Solà
[not found] ` <6ab63fad.d0a3cccc.111ed5.df78SMTPIN_ADDED_BROKEN@mx.google.com>
1 sibling, 0 replies; 9+ messages in thread
From: Miquel Sabaté Solà @ 2026-09-25 9:32 UTC (permalink / raw)
To: Qu Wenruo; +Cc: linux-btrfs
[-- Attachment #1: Type: text/plain, Size: 19543 bytes --]
Qu Wenruo @ 2026-09-25 18:37 +0930:
> A delayed ordered extent has the following features:
>
> - A new BTRFS_ORDERED_DELAYED flag
> And this new flag must be set along with the BTRFS_ORDERED_REGULAR flag.
>
> - No allocation of any on-disk space
> As a delayed ordered extent doesn't take any on-disk space yet, it
> won't release any reserved data/meta space either.
>
> - Zero or more real OEs can be added to the parent
> If a real OE is allocated, it must be inside the parent OE.
> And such real OE will go through the regular data/meta space
> reservation path.
>
> - Child OEs will not be added to the per-inode OE rb-tree nor
> per-root list
> Only the parent OE is added to the per-inode rb-tree and per-root
> list.
> So anything waiting for ordered extents should only work on the parent
> one.
>
> There is a special corner case for btrfs_wait_ordered_extents(), as
> delayed parent OEs have 0 disk_bytenr and disk_num_bytes, they will
> be considered out of the [0, U64_MAX] range.
> Thus we have to always wait for any delayed OEs of a root, no matter
> if a block group range is given or not.
>
> - When the parent OE finishes, all child OEs will also be finished
> And reserved space is all handled by the child OEs.
>
> - Any range not covered by a child OE will be manually cleaned up
> When adding a child OE to the parent one, the range in
> child_cleanup_bitmap will be cleared.
>
> If a range is already cleaned up but without a child OE (happens when
> OE allocation failed), whoever cleans up the range should clear the bits
> in the child_cleanup_bitmap.
>
> And when the parent OE finishes, any range in child_cleanup_bitmap
> will be properly cleaned up.
>
> The above features allow us to use the existing ordered extent interfaces
> to allocate new real OEs, and wait for them properly.
>
> Signed-off-by: Qu Wenruo <wqu@suse.com>
> ---
> fs/btrfs/inode.c | 86 ++++++++++++++++++-
> fs/btrfs/ordered-data.c | 185 ++++++++++++++++++++++++++++++----------
> fs/btrfs/ordered-data.h | 20 +++++
> 3 files changed, 243 insertions(+), 48 deletions(-)
>
> diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
> index 2a32072849cb..e9ed3e31fa88 100644
> --- a/fs/btrfs/inode.c
> +++ b/fs/btrfs/inode.c
> @@ -3204,6 +3204,81 @@ static int insert_ordered_extent_file_extent(struct btrfs_trans_handle *trans,
> update_inode_bytes, oe->qgroup_rsv);
> }
>
> +static int finish_delayed_ordered(struct btrfs_ordered_extent *oe)
> +{
> + struct btrfs_inode *inode = oe->inode;
> + struct btrfs_fs_info *fs_info = inode->root->fs_info;
> + struct btrfs_ordered_extent *child;
> + struct btrfs_ordered_extent *tmp;
> + struct extent_state *cached = NULL;
> + const u32 nr_bits = oe->num_bytes >> fs_info->sectorsize_bits;
> + bool io_error = test_bit(BTRFS_ORDERED_IOERR, &oe->flags);
I believe this one can be made "const", right?
> + u32 cur_bit = 0;
> + int ret = 0;
> + int saved_ret = 0;
> +
> + /* Finish each child OE. */
> + list_for_each_entry_safe(child, tmp, &oe->child_list, child_list) {
> + const u32 child_bit = (child->file_offset - oe->file_offset) >>
> + fs_info->sectorsize_bits;
> + const u32 child_nr_bits = child->num_bytes >> fs_info->sectorsize_bits;
> +
> + list_del_init(&child->child_list);
> + refcount_inc(&child->refs);
> +
> + /* The range should have been cleared in the bitmap. */
> + ASSERT(bitmap_test_range_all_zero(oe->child_cleanup_bitmap,
> + child_bit, child_nr_bits));
> +
> + if (io_error)
> + set_bit(BTRFS_ORDERED_IOERR, &child->flags);
> +
> + ret = btrfs_finish_one_ordered(child);
> + if (ret && !saved_ret)
> + saved_ret = ret;
> + }
> +
> + while (cur_bit < nr_bits) {
> + u64 range_start;
> + u64 range_end;
> + u32 range_len;
> + unsigned int first_zero;
> +
> + cur_bit = find_next_bit(oe->child_cleanup_bitmap, nr_bits, cur_bit);
> +
> + if (cur_bit >= nr_bits)
> + break;
> +
> + first_zero = find_next_zero_bit(oe->child_cleanup_bitmap, nr_bits,
> + cur_bit);
> + range_start = oe->file_offset + (cur_bit << fs_info->sectorsize_bits);
> + range_len = (first_zero - cur_bit) << fs_info->sectorsize_bits;
> + range_end = range_start + range_len - 1;
> + cur_bit = first_zero;
> +
> + btrfs_lock_extent(&inode->io_tree, range_start, range_end, &cached);
> + /*
> + * The range has reserved data/metadata but no real OE, thus we have
> + * to manually release them.
> + */
> + btrfs_delalloc_release_space(inode, NULL, range_start, range_len, true);
> + /*
> + * Also need to remove/drop the pinned extent map range.
> + * Here we do not want the extent map to stay, as they do not represent
> + * any real extent on-disk.
> + */
> + btrfs_drop_extent_map_range(inode, range_start, range_end, false);
> + btrfs_clear_extent_bit(&inode->io_tree, range_start, range_end,
> + EXTENT_LOCKED | EXTENT_DELALLOC_NEW | EXTENT_DEFRAG |
> + EXTENT_DO_ACCOUNTING, &cached);
> + }
> +
> + btrfs_remove_ordered_extent(oe);
> + btrfs_put_ordered_extent(oe);
> + btrfs_put_ordered_extent(oe);
> + return saved_ret;
> +}
> +
> /*
> * As ordered data IO finishes, this gets called so we can finish
> * an ordered extent if the range of bytes in the file it covers are
> @@ -3226,6 +3301,13 @@ int btrfs_finish_one_ordered(struct btrfs_ordered_extent *ordered_extent)
> bool clear_reserved_extent = true;
> unsigned int clear_bits = 0;
>
> + freespace_inode = btrfs_is_free_space_inode(inode);
> + if (!freespace_inode)
> + btrfs_lockdep_acquire(fs_info, btrfs_ordered_extent);
> +
> + if (test_bit(BTRFS_ORDERED_DELAYED, &ordered_extent->flags))
> + return finish_delayed_ordered(ordered_extent);
> +
> start = ordered_extent->file_offset;
> end = start + ordered_extent->num_bytes - 1;
>
> @@ -3238,10 +3320,6 @@ int btrfs_finish_one_ordered(struct btrfs_ordered_extent *ordered_extent)
> if (!test_bit(BTRFS_ORDERED_NOCOW, &ordered_extent->flags))
> clear_bits |= EXTENT_DEFRAG;
>
> - freespace_inode = btrfs_is_free_space_inode(inode);
> - if (!freespace_inode)
> - btrfs_lockdep_acquire(fs_info, btrfs_ordered_extent);
> -
> if (unlikely(test_bit(BTRFS_ORDERED_IOERR, &ordered_extent->flags))) {
> ret = -EIO;
> goto out;
> diff --git a/fs/btrfs/ordered-data.c b/fs/btrfs/ordered-data.c
> index b32d4eabe0ab..aad26972f8b4 100644
> --- a/fs/btrfs/ordered-data.c
> +++ b/fs/btrfs/ordered-data.c
> @@ -155,6 +155,7 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
> u64 qgroup_rsv = 0;
> const bool is_nocow = (flags &
> ((1U << BTRFS_ORDERED_NOCOW) | (1U << BTRFS_ORDERED_PREALLOC)));
> + const bool is_delayed = test_bit(BTRFS_ORDERED_DELAYED, &flags);
>
> /* Only one type flag can be set. */
> ASSERT(has_single_bit_set(flags & BTRFS_ORDERED_EXCLUSIVE_FLAGS),
> @@ -170,6 +171,17 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
> if (test_bit(BTRFS_ORDERED_ENCODED, &flags))
> ASSERT(test_bit(BTRFS_ORDERED_COMPRESSED, &flags));
>
> + /*
> + * DELAYED can only be set with REGULAR, no DIRECT/ENCODED, and should
> + * not exceed BTRFS_MAX_COMPRESSED size.
> + */
> + if (test_bit(BTRFS_ORDERED_DELAYED, &flags)) {
> + ASSERT(test_bit(BTRFS_ORDERED_REGULAR, &flags));
> + ASSERT(!test_bit(BTRFS_ORDERED_DIRECT, &flags));
> + ASSERT(!test_bit(BTRFS_ORDERED_ENCODED, &flags));
> + ASSERT(num_bytes <= BTRFS_MAX_COMPRESSED);
> + }
> +
> /*
> * For a NOCOW write we can free the qgroup reserve right now. For a COW
> * one we transfer the reserved space from the inode's iotree into the
> @@ -178,13 +190,16 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
> * completing the ordered extent, when running the data delayed ref it
> * creates, we free the reserved data with btrfs_qgroup_free_refroot().
> */
> - if (is_nocow)
> - ret = btrfs_qgroup_free_data(inode, NULL, file_offset, num_bytes, &qgroup_rsv);
> - else
> - ret = btrfs_qgroup_release_data(inode, file_offset, num_bytes, &qgroup_rsv);
> -
> - if (ret < 0)
> - return ERR_PTR(ret);
> + if (!is_delayed) {
> + if (is_nocow)
> + ret = btrfs_qgroup_free_data(inode, NULL, file_offset,
> + num_bytes, &qgroup_rsv);
> + else
> + ret = btrfs_qgroup_release_data(inode, file_offset,
> + num_bytes, &qgroup_rsv);
> + if (ret < 0)
> + return ERR_PTR(ret);
> + }
>
> entry = kmem_cache_zalloc(btrfs_ordered_extent_cache, GFP_NOFS);
> if (!entry) {
> @@ -216,19 +231,26 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
> INIT_LIST_HEAD(&entry->root_extent_list);
> INIT_LIST_HEAD(&entry->work_list);
> INIT_LIST_HEAD(&entry->bioc_list);
> + INIT_LIST_HEAD(&entry->child_list);
> init_completion(&entry->completion);
> + RB_CLEAR_NODE(&entry->rb_node);
>
> /*
> * We don't need the count_max_extents here, we can assume that all of
> * that work has been done at higher layers, so this is truly the
> * smallest the extent is going to get.
> */
> - spin_lock(&inode->lock);
> - btrfs_mod_outstanding_extents(inode, 1);
> - spin_unlock(&inode->lock);
> + if (!is_delayed) {
> + spin_lock(&inode->lock);
> + btrfs_mod_outstanding_extents(inode, 1);
> + spin_unlock(&inode->lock);
> + } else {
> + bitmap_set(entry->child_cleanup_bitmap, 0,
> + num_bytes >> inode->root->fs_info->sectorsize_bits);
> + }
>
> out:
> - if (IS_ERR(entry) && !is_nocow)
> + if (IS_ERR(entry) && !is_nocow && !is_delayed)
> btrfs_qgroup_free_refroot(inode->root->fs_info,
> btrfs_root_id(inode->root),
> qgroup_rsv, BTRFS_QGROUP_RSV_DATA);
> @@ -236,12 +258,47 @@ static struct btrfs_ordered_extent *alloc_ordered_extent(
> return entry;
> }
>
> +static void add_child_oe(struct btrfs_ordered_extent *parent,
> + struct btrfs_ordered_extent *child)
> +{
> + struct btrfs_inode *inode = parent->inode;
> + struct btrfs_fs_info *fs_info = inode->root->fs_info;
> + const u32 start_bit = (child->file_offset - parent->file_offset) >>
> + fs_info->sectorsize_bits;
> + const u32 nr_bits = child->num_bytes >> fs_info->sectorsize_bits;
> +
> + lockdep_assert_held(&inode->ordered_tree_lock);
> + /* Basic flags check for parent and child. */
> + ASSERT(test_bit(BTRFS_ORDERED_DELAYED, &parent->flags));
> + ASSERT(!test_bit(BTRFS_ORDERED_DELAYED, &child->flags));
> +
> + /* Child should not belong to any parent yet. */
> + ASSERT(list_empty(&child->child_list));
> +
> + /* Child should be fully inside parent's range. */
> + ASSERT(child->file_offset >= parent->file_offset);
> + ASSERT(child->file_offset + child->num_bytes <=
> + parent->file_offset + parent->num_bytes);
> +
> + /*
> + * There should be no existing child in the range, thus
> + * all cleanup bits should be set.
> + */
> + ASSERT(bitmap_test_range_all_set(parent->child_cleanup_bitmap,
> + start_bit, nr_bits));
> +
> + list_add_tail(&child->child_list, &parent->child_list);
> +
> + bitmap_clear(parent->child_cleanup_bitmap, start_bit, nr_bits);
> +}
> +
> static void insert_ordered_extent(struct btrfs_ordered_extent *entry)
> {
> struct btrfs_inode *inode = entry->inode;
> struct btrfs_root *root = inode->root;
> struct btrfs_fs_info *fs_info = root->fs_info;
> struct rb_node *node;
> + bool is_child = false;
>
> trace_btrfs_ordered_extent_add(inode, entry);
>
> @@ -254,17 +311,25 @@ static void insert_ordered_extent(struct btrfs_ordered_extent *entry)
> spin_lock(&inode->ordered_tree_lock);
> node = tree_insert(&inode->ordered_tree, entry->file_offset,
> &entry->rb_node);
> - if (unlikely(node)) {
> + if (node) {
> struct btrfs_ordered_extent *exist =
> rb_entry(node, struct btrfs_ordered_extent, rb_node);
>
> - btrfs_panic(fs_info, -EEXIST,
> + if (test_bit(BTRFS_ORDERED_DELAYED, &exist->flags)) {
> + add_child_oe(exist, entry);
> + is_child = true;
> + } else {
> + btrfs_panic(fs_info, -EEXIST,
> "overlapping ordered extents, existing oe file_offset %llu num_bytes %llu flags 0x%lx, new oe file_offset %llu num_bytes %llu flags 0x%lx",
> exist->file_offset, exist->num_bytes, exist->flags,
> entry->file_offset, entry->num_bytes, entry->flags);
> + }
> }
> spin_unlock(&inode->ordered_tree_lock);
>
> + /* Child OE shouldn't be added to per-root oe list. */
> + if (is_child)
> + return;
> spin_lock(&root->ordered_extent_lock);
> list_add_tail(&entry->root_extent_list,
> &root->ordered_extents);
> @@ -337,6 +402,20 @@ struct btrfs_ordered_extent *btrfs_alloc_ordered_extent(
> return entry;
> }
>
> +struct btrfs_ordered_extent *btrfs_alloc_delayed_ordered_extent(
> + struct btrfs_inode *inode, u64 file_offset, u32 length)
> +{
> + struct btrfs_ordered_extent *entry;
> +
> + entry = alloc_ordered_extent(inode, file_offset, length, length, 0, 0, 0,
> + (1UL << BTRFS_ORDERED_REGULAR) |
> + (1UL << BTRFS_ORDERED_DELAYED),
> + BTRFS_COMPRESS_NONE);
> + if (!IS_ERR(entry))
> + insert_ordered_extent(entry);
> + return entry;
> +}
> +
> /*
> * Add a struct btrfs_ordered_sum into the list of checksums to be inserted
> * when an ordered extent is finished. If the list covers more than one
> @@ -656,8 +735,9 @@ void btrfs_remove_ordered_extent(struct btrfs_ordered_extent *entry)
> struct btrfs_root *root = btrfs_inode->root;
> struct btrfs_fs_info *fs_info = root->fs_info;
> struct rb_node *node;
> - bool pending;
> + bool pending = false;
> bool freespace_inode;
> + const bool is_delayed = test_bit(BTRFS_ORDERED_DELAYED, &entry->flags);
>
> /*
> * If this is a free space inode the thread has not acquired the ordered
> @@ -666,33 +746,37 @@ void btrfs_remove_ordered_extent(struct btrfs_ordered_extent *entry)
> freespace_inode = btrfs_is_free_space_inode(btrfs_inode);
>
> btrfs_lockdep_acquire(fs_info, btrfs_trans_pending_ordered);
> - /* This is paired with alloc_ordered_extent(). */
> - spin_lock(&btrfs_inode->lock);
> - btrfs_mod_outstanding_extents(btrfs_inode, -1);
> - spin_unlock(&btrfs_inode->lock);
> - if (root != fs_info->tree_root) {
> - u64 release;
> + if (!is_delayed) {
> + /* This is paired with alloc_ordered_extent(). */
> + spin_lock(&btrfs_inode->lock);
> + btrfs_mod_outstanding_extents(btrfs_inode, -1);
> + spin_unlock(&btrfs_inode->lock);
>
> - if (test_bit(BTRFS_ORDERED_ENCODED, &entry->flags))
> - release = entry->disk_num_bytes;
> - else
> - release = entry->num_bytes;
> - btrfs_delalloc_release_metadata(btrfs_inode, release,
> + if (root != fs_info->tree_root) {
> + u64 release;
> +
> + if (test_bit(BTRFS_ORDERED_ENCODED, &entry->flags))
> + release = entry->disk_num_bytes;
> + else
> + release = entry->num_bytes;
> + btrfs_delalloc_release_metadata(btrfs_inode, release,
> test_bit(BTRFS_ORDERED_IOERR,
> &entry->flags));
> + }
> }
> -
> percpu_counter_add_batch(&fs_info->ordered_bytes, -entry->num_bytes,
> fs_info->delalloc_batch);
>
> spin_lock(&btrfs_inode->ordered_tree_lock);
> - node = &entry->rb_node;
> - rb_erase(node, &btrfs_inode->ordered_tree);
> - RB_CLEAR_NODE(node);
> - if (btrfs_inode->ordered_tree_last == node)
> - btrfs_inode->ordered_tree_last = NULL;
> - set_bit(BTRFS_ORDERED_COMPLETE, &entry->flags);
> - pending = test_and_clear_bit(BTRFS_ORDERED_PENDING, &entry->flags);
> + if (!RB_EMPTY_NODE(&entry->rb_node)) {
> + node = &entry->rb_node;
> + rb_erase(node, &btrfs_inode->ordered_tree);
> + RB_CLEAR_NODE(node);
> + if (btrfs_inode->ordered_tree_last == node)
> + btrfs_inode->ordered_tree_last = NULL;
> + set_bit(BTRFS_ORDERED_COMPLETE, &entry->flags);
> + pending = test_and_clear_bit(BTRFS_ORDERED_PENDING, &entry->flags);
> + }
> spin_unlock(&btrfs_inode->ordered_tree_lock);
>
> /*
> @@ -724,17 +808,20 @@ void btrfs_remove_ordered_extent(struct btrfs_ordered_extent *entry)
>
> btrfs_lockdep_release(fs_info, btrfs_trans_pending_ordered);
>
> - spin_lock(&root->ordered_extent_lock);
> - list_del_init(&entry->root_extent_list);
> - root->nr_ordered_extents--;
> -
> trace_btrfs_ordered_extent_remove(btrfs_inode, entry);
>
> - if (!root->nr_ordered_extents) {
> - spin_lock(&fs_info->ordered_root_lock);
> - BUG_ON(list_empty(&root->ordered_root));
> - list_del_init(&root->ordered_root);
> - spin_unlock(&fs_info->ordered_root_lock);
> + spin_lock(&root->ordered_extent_lock);
> + /* For child OEs, they are not added to per-root OEs. */
> + if (!list_empty(&entry->root_extent_list)) {
> + list_del_init(&entry->root_extent_list);
> + root->nr_ordered_extents--;
> +
> + if (!root->nr_ordered_extents) {
> + spin_lock(&fs_info->ordered_root_lock);
> + BUG_ON(list_empty(&root->ordered_root));
> + list_del_init(&root->ordered_root);
> + spin_unlock(&fs_info->ordered_root_lock);
> + }
> }
> spin_unlock(&root->ordered_extent_lock);
> wake_up(&entry->wait);
> @@ -783,8 +870,18 @@ u64 btrfs_wait_ordered_extents(struct btrfs_root *root, u64 nr,
> ordered = list_first_entry(&splice, struct btrfs_ordered_extent,
> root_extent_list);
>
> - if (range_end <= ordered->disk_bytenr ||
> - ordered->disk_bytenr + ordered->disk_num_bytes <= range_start) {
> + /*
> + * Delayed OEs have 0 disk_bytenr and 0 disk_num_bytes, thus
> + * they will be considered out of the [0, U64_MAX) range.
> + * And we do not know where they will really land until the
> + * writeback has finished.
> + *
> + * So here we must exclude delayed OEs from the block group
> + * range check, and always wait for them.
> + */
> + if (!test_bit(BTRFS_ORDERED_DELAYED, &ordered->flags) &&
> + (range_end <= ordered->disk_bytenr ||
> + ordered->disk_bytenr + ordered->disk_num_bytes <= range_start)) {
> list_move_tail(&ordered->root_extent_list, &skipped);
> cond_resched_lock(&root->ordered_extent_lock);
> continue;
> diff --git a/fs/btrfs/ordered-data.h b/fs/btrfs/ordered-data.h
> index 8d5d5ba1e02f..f72588c65e62 100644
> --- a/fs/btrfs/ordered-data.h
> +++ b/fs/btrfs/ordered-data.h
> @@ -13,6 +13,7 @@
> #include <linux/rbtree.h>
> #include <linux/wait.h>
> #include "async-thread.h"
> +#include "compression.h"
>
> struct inode;
> struct page;
> @@ -87,6 +88,12 @@ enum {
> */
> BTRFS_ORDERED_DIRECT,
>
> + /*
> + * Extra bit for delayed OE, can only be set for REGULAR.
> + * Cannot be set with COMPRESSED/ENCODED/DIRECT.
> + */
> + BTRFS_ORDERED_DELAYED,
> +
> BTRFS_ORDERED_NR_FLAGS,
> };
> static_assert(BTRFS_ORDERED_NR_FLAGS <= BITS_PER_LONG);
> @@ -155,6 +162,17 @@ struct btrfs_ordered_extent {
> /* a per root list of all the pending ordered extents */
> struct list_head root_extent_list;
>
> + /* Child ordered extent list for delayed OE. */
> + struct list_head child_list;
> +
> + /*
> + * Only utilized by delayed parent OE.
> + *
> + * Indicate the range that needs cleanup.
> + */
> + unsigned long child_cleanup_bitmap[BITS_TO_LONGS(
> + BTRFS_MAX_COMPRESSED / BTRFS_MIN_BLOCKSIZE)];
> +
> struct btrfs_work work;
>
> struct completion completion;
> @@ -192,6 +210,8 @@ struct btrfs_file_extent {
> struct btrfs_ordered_extent *btrfs_alloc_ordered_extent(
> struct btrfs_inode *inode, u64 file_offset,
> const struct btrfs_file_extent *file_extent, unsigned long flags);
> +struct btrfs_ordered_extent *btrfs_alloc_delayed_ordered_extent(
> + struct btrfs_inode *inode, u64 file_offset, u32 length);
> void btrfs_add_ordered_sum(struct btrfs_ordered_extent *entry,
> struct btrfs_ordered_sum *sum);
> struct btrfs_ordered_extent *btrfs_lookup_ordered_extent(struct btrfs_inode *inode,
[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 897 bytes --]
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v7 1/6] btrfs: add delayed ordered extent support
[not found] ` <6ab63fad.d0a3cccc.111ed5.df78SMTPIN_ADDED_BROKEN@mx.google.com>
@ 2026-09-25 9:36 ` Qu Wenruo
0 siblings, 0 replies; 9+ messages in thread
From: Qu Wenruo @ 2026-09-25 9:36 UTC (permalink / raw)
To: Miquel Sabaté Solà; +Cc: linux-btrfs
在 2026/9/25 19:02, Miquel Sabaté Solà 写道:
> Qu Wenruo @ 2026-09-25 18:37 +0930:
>
>> A delayed ordered extent has the following features:
>>
>> - A new BTRFS_ORDERED_DELAYED flag
>> And this new flag must be set along with the BTRFS_ORDERED_REGULAR flag.
>>
>> - No allocation of any on-disk space
>> As a delayed ordered extent doesn't take any on-disk space yet, it
>> won't release any reserved data/meta space either.
>>
>> - Zero or more real OEs can be added to the parent
>> If a real OE is allocated, it must be inside the parent OE.
>> And such real OE will go through the regular data/meta space
>> reservation path.
>>
>> - Child OEs will not be added to the per-inode OE rb-tree nor
>> per-root list
>> Only the parent OE is added to the per-inode rb-tree and per-root
>> list.
>> So anything waiting for ordered extents should only work on the parent
>> one.
>>
>> There is a special corner case for btrfs_wait_ordered_extents(), as
>> delayed parent OEs have 0 disk_bytenr and disk_num_bytes, they will
>> be considered out of the [0, U64_MAX] range.
>> Thus we have to always wait for any delayed OEs of a root, no matter
>> if a block group range is given or not.
>>
>> - When the parent OE finishes, all child OEs will also be finished
>> And reserved space is all handled by the child OEs.
>>
>> - Any range not covered by a child OE will be manually cleaned up
>> When adding a child OE to the parent one, the range in
>> child_cleanup_bitmap will be cleared.
>>
>> If a range is already cleaned up but without a child OE (happens when
>> OE allocation failed), whoever cleans up the range should clear the bits
>> in the child_cleanup_bitmap.
>>
>> And when the parent OE finishes, any range in child_cleanup_bitmap
>> will be properly cleaned up.
>>
>> The above features allow us to use the existing ordered extent interfaces
>> to allocate new real OEs, and wait for them properly.
>>
>> Signed-off-by: Qu Wenruo <wqu@suse.com>
>> ---
>> fs/btrfs/inode.c | 86 ++++++++++++++++++-
>> fs/btrfs/ordered-data.c | 185 ++++++++++++++++++++++++++++++----------
>> fs/btrfs/ordered-data.h | 20 +++++
>> 3 files changed, 243 insertions(+), 48 deletions(-)
>>
>> diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
>> index 2a32072849cb..e9ed3e31fa88 100644
>> --- a/fs/btrfs/inode.c
>> +++ b/fs/btrfs/inode.c
>> @@ -3204,6 +3204,81 @@ static int insert_ordered_extent_file_extent(struct btrfs_trans_handle *trans,
>> update_inode_bytes, oe->qgroup_rsv);
>> }
>>
>> +static int finish_delayed_ordered(struct btrfs_ordered_extent *oe)
>> +{
>> + struct btrfs_inode *inode = oe->inode;
>> + struct btrfs_fs_info *fs_info = inode->root->fs_info;
>> + struct btrfs_ordered_extent *child;
>> + struct btrfs_ordered_extent *tmp;
>> + struct extent_state *cached = NULL;
>> + const u32 nr_bits = oe->num_bytes >> fs_info->sectorsize_bits;
>> + bool io_error = test_bit(BTRFS_ORDERED_IOERR, &oe->flags);
>
> I believe this one can be made "const", right?
This is a very minor one, I can do it during merge.
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-09-25 9:36 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-25 9:07 [PATCH v7 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 1/6] btrfs: add delayed ordered extent support Qu Wenruo
2026-09-25 9:32 ` Miquel Sabaté Solà
[not found] ` <6ab63fad.d0a3cccc.111ed5.df78SMTPIN_ADDED_BROKEN@mx.google.com>
2026-09-25 9:36 ` Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 2/6] btrfs: add skeleton for delayed btrfs bio Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 3/6] btrfs: introduce the skeleton of delayed bbio endio function Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 4/6] btrfs: introduce compression for delayed bbio Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 5/6] btrfs: implement uncompressed fallback " Qu Wenruo
2026-09-25 9:07 ` [PATCH v7 6/6] btrfs: enable experimental delayed compression support Qu Wenruo
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox