From: Qu Wenruo <wqu@suse.com>
To: linux-btrfs@vger.kernel.org
Subject: [PATCH v4 RESEND 6/6] btrfs: enable experimental delayed compression support
Date: Fri, 18 Sep 2026 10:14:05 +0930 [thread overview]
Message-ID: <6fa82fa88c238dcbe58e1f54e43ce8fbe7a7fdd2.1789692183.git.wqu@suse.com> (raw)
In-Reply-To: <cover.1789692182.git.wqu@suse.com>
Instead of the existing async submission path, the new delayed bbio will
handle compressed writes by:
- Allocating delayed em/oe at run_delalloc_*() time
Thus no data extent is reserved at that time.
- Delayed bbio will be assembled at extent_writepage_io() time
- Delayed bbio will be intercepted just before submission
Which will run compression (or fallback to uncompressed writes) in
workqueue.
Data extents will only be reserved at that time, and the delayed em
will be replaced by real ones.
Meanwhile the real OE will be added as a child of the parent delayed
OE, and when the parent OE finishes, the child OE will be finished
with their file extents inserted.
This has some benefits:
- Higher concurrency
Previously async submission will hold the folio and io tree range
locked, this means we cannot even read the uptodate folio.
Furthermore although the compressed write is queued into a workqueue
for submission and extent_writepage_io() will skip the compressed
range, when we need to write the next folio of the compressed range,
we will need to wait for the folio to be unlocked.
This makes async submission less async.
- Future DONTCACHE writes support
We do not support DONTCACHE because that feature requires writeback path
to clear the folio dirty and submit them sequentially.
Meanwhile async submission makes the writeback async, breaking the
sequential submission requirement.
This is also why we need complex per-block tracking for writeback
flags, while iomap only requires a counter tracking.
With the new delayed compression, the lifespan of a folio aligns with
DONTCACHE and iomap.
However there is also a change to the fsync() behavior:
- Now an inode with compressed write will require full sync
We can not do any fast sync, as that will grab the info from OE
directly, but since delayed OE is just a place holder, all the OE info
can not be trusted until the parent OE is finished.
There is an extra handling for defrag, where we have to write and wait
for the defrag range.
This is to avoid clearing inode->defrag_compress before the compression
starts.
Finally this new experimental feature is not yet compatible with zoned,
so for zoned btrfs it will go through the old compression path.
Signed-off-by: Qu Wenruo <wqu@suse.com>
---
fs/btrfs/defrag.c | 27 ++++++++++++++---
fs/btrfs/extent_io.c | 5 +++-
fs/btrfs/inode.c | 71 ++++++++++++++++++++++++++++++++++++++++++--
3 files changed, 95 insertions(+), 8 deletions(-)
diff --git a/fs/btrfs/defrag.c b/fs/btrfs/defrag.c
index 6ec5dd760d42..32801ae4da7a 100644
--- a/fs/btrfs/defrag.c
+++ b/fs/btrfs/defrag.c
@@ -1349,6 +1349,7 @@ int btrfs_defrag_file(struct btrfs_inode *inode, struct file_ra_state *ra,
struct btrfs_fs_info *fs_info = inode->root->fs_info;
unsigned long sectors_defragged = 0;
u64 isize = i_size_read(&inode->vfs_inode);
+ const u64 start = round_down(range->start, fs_info->sectorsize);
u64 cur;
u64 last_byte;
bool do_compress = (range->flags & BTRFS_DEFRAG_RANGE_COMPRESS);
@@ -1400,7 +1401,7 @@ int btrfs_defrag_file(struct btrfs_inode *inode, struct file_ra_state *ra,
}
/* Align the range */
- cur = round_down(range->start, fs_info->sectorsize);
+ cur = start;
last_byte = round_up(last_byte, fs_info->sectorsize) - 1;
/*
@@ -1471,10 +1472,28 @@ int btrfs_defrag_file(struct btrfs_inode *inode, struct file_ra_state *ra,
* need to be written back immediately.
*/
if (range->flags & BTRFS_DEFRAG_RANGE_START_IO) {
- filemap_flush(inode->vfs_inode.i_mapping);
- if (test_bit(BTRFS_INODE_HAS_ASYNC_EXTENT,
- &inode->runtime_flags))
+ /*
+ * For experimental delayed writeback, we must wait
+ * for the range to be fully written back before
+ * clearing inode->defrag_compress.
+ *
+ * Regular filemap_flush() will only start writeback,
+ * which will only create delayed OEs. But the real
+ * compression is happening later.
+ * This means if we just flush but not wait for writeback,
+ * the inode->defrag_compress clearing can race with
+ * compression, so the requested compression algorithm
+ * is not applied.
+ */
+ if (IS_ENABLED(CONFIG_BTRFS_EXPERIMENTAL)) {
+ filemap_write_and_wait_range(inode->vfs_inode.i_mapping,
+ start, last_byte);
+ } else {
filemap_flush(inode->vfs_inode.i_mapping);
+ if (test_bit(BTRFS_INODE_HAS_ASYNC_EXTENT,
+ &inode->runtime_flags))
+ filemap_flush(inode->vfs_inode.i_mapping);
+ }
}
if (range->compress_type == BTRFS_COMPRESS_LZO)
btrfs_set_fs_incompat(fs_info, COMPRESS_LZO);
diff --git a/fs/btrfs/extent_io.c b/fs/btrfs/extent_io.c
index a822f6635eba..970548df2dbc 100644
--- a/fs/btrfs/extent_io.c
+++ b/fs/btrfs/extent_io.c
@@ -899,8 +899,11 @@ static int submit_one_block(struct btrfs_bio_ctrl *bio_ctrl,
* If we have accumulated decent amount of IO, send it to the block
* layer so that IO can run while we are accumulating more folios to
* write.
+ *
+ * This doesn't apply to delayed bbio which is going to be
+ * compressed.
*/
- else if (bio_ctrl->wbc &&
+ else if (bio_ctrl->wbc && !bio_ctrl->bbio->is_delayed &&
bio_ctrl->bbio->bio.bi_iter.bi_size >= fs_info->writeback_bio_size)
submit_one_bio(bio_ctrl);
return 0;
diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c
index 656616465336..78ebb3f81d01 100644
--- a/fs/btrfs/inode.c
+++ b/fs/btrfs/inode.c
@@ -1659,6 +1659,68 @@ static bool run_delalloc_compressed(struct btrfs_inode *inode,
return true;
}
+static int run_delalloc_delayed(struct btrfs_inode *inode, struct folio *locked_folio,
+ u64 start, u64 end)
+{
+ struct btrfs_root *root = inode->root;
+ struct btrfs_fs_info *fs_info = root->fs_info;
+ struct extent_state *cached = NULL;
+ u64 cur = start;
+ int ret;
+
+ if (btrfs_is_shutdown(fs_info)) {
+ ret = -EIO;
+ goto error;
+ }
+ while (cur < end) {
+ struct extent_map *em;
+ struct btrfs_ordered_extent *oe;
+ u32 cur_len = min_t(u64, end + 1 - cur, BTRFS_MAX_COMPRESSED);
+
+ btrfs_lock_extent(&inode->io_tree, cur, cur + cur_len - 1, &cached);
+ em = btrfs_create_delayed_em(inode, cur, cur_len);
+ if (IS_ERR(em)) {
+ ret = PTR_ERR(em);
+ goto error;
+ }
+ btrfs_free_extent_map(em);
+ oe = btrfs_alloc_delayed_ordered_extent(inode, cur, cur_len);
+ if (IS_ERR(oe)) {
+ btrfs_drop_extent_map_range(inode, cur, cur + cur_len - 1, false);
+ ret = PTR_ERR(oe);
+ goto error;
+ }
+ btrfs_put_ordered_extent(oe);
+
+ cur += cur_len;
+ }
+ /*
+ * Force the next fsync to wait for any delayed OE to finish,
+ * as the fast path cannot handle delayed OE correctly.
+ */
+ btrfs_set_inode_full_sync(inode);
+ extent_clear_unlock_delalloc(inode, start, end, locked_folio, &cached,
+ EXTENT_LOCKED | EXTENT_DELALLOC,
+ PAGE_UNLOCK);
+ return 0;
+error:
+ if (start < cur) {
+ btrfs_drop_extent_map_range(inode, start, cur - 1, false);
+ btrfs_cleanup_ordered_extents(inode, start, cur - start);
+ extent_clear_unlock_delalloc(inode, start, cur - 1, locked_folio, &cached,
+ EXTENT_LOCKED | EXTENT_DELALLOC,
+ PAGE_UNLOCK | PAGE_START_WRITEBACK | PAGE_END_WRITEBACK);
+ }
+ if (cur < end) {
+ extent_clear_unlock_delalloc(inode, cur, end, locked_folio, &cached,
+ EXTENT_LOCKED | EXTENT_DELALLOC | EXTENT_DELALLOC_NEW |
+ EXTENT_DEFRAG | EXTENT_DO_ACCOUNTING,
+ PAGE_UNLOCK | PAGE_START_WRITEBACK | PAGE_END_WRITEBACK);
+ btrfs_qgroup_free_data(inode, NULL, cur, end + 1 - cur, NULL);
+ }
+ return ret;
+}
+
/*
* Run the delalloc range from start to end, and write back any dirty pages
* covered by the range.
@@ -2454,9 +2516,12 @@ int btrfs_run_delalloc_range(struct btrfs_inode *inode, struct folio *locked_fol
return run_delalloc_nocow(inode, locked_folio, start, end);
if (btrfs_inode_can_compress(inode) &&
- inode_need_compress(inode, start, end, false) &&
- run_delalloc_compressed(inode, locked_folio, start, end, wbc))
- return 1;
+ inode_need_compress(inode, start, end, false)) {
+ if (IS_ENABLED(CONFIG_BTRFS_EXPERIMENTAL) && !zoned)
+ return run_delalloc_delayed(inode, locked_folio, start, end);
+ else if (run_delalloc_compressed(inode, locked_folio, start, end, wbc))
+ return 1;
+ }
if (zoned)
return run_delalloc_cow(inode, locked_folio, start, end, wbc, true);
--
2.55.0
prev parent reply other threads:[~2026-09-18 0:44 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 0:43 [PATCH v4 RESEND 0/6] btrfs: delay compression to bbio submission time Qu Wenruo
2026-09-18 0:44 ` [PATCH v4 RESEND 1/6] btrfs: add delayed ordered extent support Qu Wenruo
2026-09-18 0:44 ` [PATCH v4 RESEND 2/6] btrfs: add skeleton for delayed btrfs bio Qu Wenruo
2026-09-18 0:44 ` [PATCH v4 RESEND 3/6] btrfs: introduce the skeleton of delayed bbio endio function Qu Wenruo
2026-09-18 0:44 ` [PATCH v4 RESEND 4/6] btrfs: introduce compression for delayed bbio Qu Wenruo
2026-09-18 0:44 ` [PATCH v4 RESEND 5/6] btrfs: implement uncompressed fallback " Qu Wenruo
2026-09-18 0:44 ` Qu Wenruo [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=6fa82fa88c238dcbe58e1f54e43ce8fbe7a7fdd2.1789692183.git.wqu@suse.com \
--to=wqu@suse.com \
--cc=linux-btrfs@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox