Linux RAID subsystem development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Yu Kuai" <yukuai@kernel.org>
Cc: linux-raid@vger.kernel.org, yukuai@fygo.io
Subject: Re: [PATCH v5 20/29] md/md-llbitmap: don't skip reshape ranges from bitmap state
Date: Sun, 02 Aug 2026 20:31:01 +0000	[thread overview]
Message-ID: <20260802203101.92D5A1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260802195038.164272-21-yukuai@kernel.org>

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:

Pre-existing issues:
- [Critical] This problem wasn't introduced by this patch, but a lockless race condition in `llbitmap_set_page_dirty` causes dirty bitmap blocks to be lost in memory and never flushed to disk.
- [High] This problem wasn't introduced by this patch, but integer truncation of `blocks` can cause an infinite loop in the `md_do_sync` thread.
--

--- Patch [20]: [PATCH v5 20/29] md/md-llbitmap: don't skip reshape ranges from bitmap state ---
commit c6db76f6bc970107f68d37fc602584cc2c0e0a1b
Author: Yu Kuai <yukuai@fygo.io>

md/md-llbitmap: don't skip reshape ranges from bitmap state

This commit modifies md-llbitmap to stop using stale bitmap state for
skipping reshape ranges. It ensures that reshape progress is tracked purely
by array metadata. This prevents data loss after a restart before checkpoints
are persisted to disk.

Tested-by: Mykola Marzhan <mykola@meshstor.io>

> diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c
> index fa16a4224c45a..a20e55fdf82b0 100644
> --- a/drivers/md/md-llbitmap.c
> +++ b/drivers/md/md-llbitmap.c
> @@ -1707,6 +1707,14 @@ static sector_t llbitmap_skip_sync_blocks(struct mddev *mddev, sector_t offset)

[Severity: High]
This isn't a bug introduced by this patch, but can integer truncation of
blocks cause an infinite loop in the md_do_sync() thread?

In llbitmap_skip_sync_blocks(), blocks is declared as an int:

    int blocks = llbitmap->chunksize - (offset & (llbitmap->chunksize - 1));

If the array size and bitmap space require chunksize to exceed INT_MAX (e.g.,
a >8TB array with a small bitmap space), chunksize will be >= 2^31. Can the
subtraction yield a value >= 2^31, which truncates into a negative 32-bit int
(e.g., -1024)?

When this negative int is returned as a 64-bit sector_t, does it sign-extend
to a massive positive value? If md_do_sync() adds this value to the current
sync sector index, could it cause the index to overflow and decrease, looping
endlessly back and forth?

>  		return 0;
>  	c = llbitmap_read(llbitmap, p);
>  
> +	/*
> +	 * Reshape progress is tracked by array metadata rather than llbitmap.
> +	 * Skipping reshape ranges from stale bitmap state can lose data after a
> +	 * restart before the corresponding bits are checkpointed to disk.
> +	 */
> +	if (test_bit(MD_RECOVERY_RESHAPE, &mddev->recovery))
> +		return 0;
> +
>  	/* always skip unwritten blocks */
>  	if (c == BitUnwritten)
>  		return blocks;

[Severity: Critical]
This is a pre-existing issue, but a lockless race condition in
llbitmap_set_page_dirty() can cause dirty bitmap blocks to be lost in memory
and never flushed to disk.

Can this sequence actually occur? If llbitmap_set_page_dirty() is called
concurrently with the llbitmap_unplug_fn() flusher thread, the writer sets
the LLPageDirty flag before setting the specific block bit:

drivers/md/md-llbitmap.c:llbitmap_set_page_dirty() {
    ...
    if (!test_bit(LLPageDirty, &pctl->flags))
        set_bit(LLPageDirty, &pctl->flags);
    
    // race window here
    
    set_bit(block, pctl->dirty);
    ...
}

If the flusher thread clears LLPageDirty and checks pctl->dirty between these
two operations:

drivers/md/md-llbitmap.c:llbitmap_unplug_fn() {
    ...
    if (!test_bit(LLPageDirty, &llbitmap->pctl[i]->flags) ||
        !test_and_clear_bit(LLPageDirty, &llbitmap->pctl[i]->flags))
        continue;
    ...
}

Will it miss the dirty block? The writer then sets the block bit, but
because LLPageDirty is now 0, the flusher will ignore this page indefinitely
until another block happens to be dirtied.

If a dirty block is not flushed to the disk bitmap, will the array fail to
resync that block after a crash or unclean shutdown, leaving stale data on
one of the drives?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260802195038.164272-1-yukuai@kernel.org?part=20

  reply	other threads:[~2026-08-02 20:31 UTC|newest]

Thread overview: 54+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-02 19:50 [PATCH v5 00/29] md: improve lockless bitmap reshape support Yu Kuai
2026-08-02 19:50 ` [PATCH v5 01/29] md/md-llbitmap: clear flush state after daemon flush Yu Kuai
2026-08-02 20:28   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 02/29] md/md-llbitmap: use GFP_NOIO for cache allocations Yu Kuai
2026-08-02 20:44   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 03/29] md/md-llbitmap: only end fully synced chunks Yu Kuai
2026-08-02 19:50 ` [PATCH v5 04/29] md/raid5: reject zero-sector reshape chunks Yu Kuai
2026-08-02 20:31   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 05/29] md/raid5: round bitmap stripes with sector division Yu Kuai
2026-08-02 20:19   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 06/29] md: wait for behind writes before destroying bitmap Yu Kuai
2026-08-02 20:40   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 07/29] md: avoid stale clone I/O accounting timestamps Yu Kuai
2026-08-02 20:45   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 08/29] md/md-llbitmap: prevent create failure bitmap UAF Yu Kuai
2026-08-02 20:39   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 09/29] md/md-llbitmap: stop daemon timer rearm on destroy Yu Kuai
2026-08-02 20:19   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 10/29] md: skip bitmap accounting for empty write ranges Yu Kuai
2026-08-02 19:50 ` [PATCH v5 11/29] md: add helper to split bios at reshape offset Yu Kuai
2026-08-02 20:19   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 12/29] md: add exact bitmap mapping and reshape hooks Yu Kuai
2026-08-02 19:50 ` [PATCH v5 13/29] md/md-llbitmap: track bitmap sync_size explicitly Yu Kuai
2026-08-02 20:24   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 14/29] md/md-llbitmap: allocate page controls independently Yu Kuai
2026-08-02 20:27   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 15/29] md/md-llbitmap: grow the page cache in place for reshape Yu Kuai
2026-08-02 20:37   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 16/29] md/md-llbitmap: track target reshape geometry fields Yu Kuai
2026-08-02 20:25   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 17/29] md/md-llbitmap: finish reshape geometry Yu Kuai
2026-08-02 20:39   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 18/29] md/md-llbitmap: refuse reshape while llbitmap still needs sync Yu Kuai
2026-08-02 20:44   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 19/29] md/md-llbitmap: add reshape range mapping helpers Yu Kuai
2026-08-02 20:31   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 20/29] md/md-llbitmap: don't skip reshape ranges from bitmap state Yu Kuai
2026-08-02 20:31   ` sashiko-bot [this message]
2026-08-02 19:50 ` [PATCH v5 21/29] md/md-llbitmap: remap checkpointed bits as reshape progresses Yu Kuai
2026-08-02 20:43   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 22/29] md/md-llbitmap: clamp state-machine walks to tracked bits Yu Kuai
2026-08-02 19:50 ` [PATCH v5 23/29] md/raid10: reject llbitmap reshape when md chunk shrinks Yu Kuai
2026-08-02 20:40   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 24/29] md/raid10: wire llbitmap reshape lifecycle Yu Kuai
2026-08-02 20:49   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 25/29] md/raid10: split reshape bios before bitmap accounting Yu Kuai
2026-08-02 20:46   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 26/29] md/raid5: add exact old and new llbitmap mapping helpers Yu Kuai
2026-08-02 19:50 ` [PATCH v5 27/29] md/raid5: reject llbitmap reshape when md chunk shrinks Yu Kuai
2026-08-02 20:42   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 28/29] md/raid5: wire llbitmap reshape lifecycle Yu Kuai
2026-08-02 20:46   ` sashiko-bot
2026-08-02 19:50 ` [PATCH v5 29/29] md/raid5: split reshape bios before bitmap accounting Yu Kuai
2026-08-03 12:11 ` [PATCH v5 00/29] md: improve lockless bitmap reshape support Mykola Marzhan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260802203101.92D5A1F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=linux-raid@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=yukuai@fygo.io \
    --cc=yukuai@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox