Linux RAID subsystem development
 help / color / mirror / Atom feed
* Re: [QUESTION] Debugging some file data corruption
From: Kenta Akagi @ 2026-02-11 14:07 UTC (permalink / raw)
  To: Calvin Owens
  Cc: quwenruo.btrfs, linux-block, linux-raid, linux-btrfs,
	linux-crypto, k
In-Reply-To: <aWU5VnDW8cHnnauK@mozart.vkv.me>



On 2026/01/13 3:11, Calvin Owens wrote:
> On Wednesday 11/12 at 07:32 +1030, Qu Wenruo wrote:
>> 在 2025/11/12 03:31, Calvin Owens 写道:
>>> Hello all,
>>>
>>> I'm looking for help debugging some corruption I recently encountered.
>>> It happened on 6.17.0, and I'm trying to reproduce it on 6.18-rc. This
>>> is not really actionable yet, I'm just looking for advice.
>>>
>>> After copying about 10TB of data to a btrfs+luks+mdraid1 across two 18TB
>>> drives,
>>
>> With LUKS in the middle, it makes any corruption pattern very human
>> unreadable.
>>
>> I guess it's not really feasible to try to reproduce the problem again since
>> it has 10TiB data involved?
>>
>> But if you can spend a lot of time waiting for data copy, mind to try
>> the following combination(s)?
>>
>> - btrfs on mdraid1
>> - btrfs RAID1 on raw two HDDs
> 
> I re-ran the copies several times and never reproduced it. But I finally
> hit it again after I gave up, while making a backup on 6.18.0 with
> btrfs-raid1+luks:
> 
>     Opening filesystem to check...
>     Checking filesystem on /dev/mapper/sdc_crypt
>     UUID: f8223856-32cc-4dcf-8cee-9312e032c005
>     [1/8] checking log skipped (none written)
>     [1/7] checking root items                      (0:00:59 elapsed, 4875792 items checked)
>     [2/7] checking extents                         (0:28:03 elapsed, 4583919 items checked)
>     [3/7] checking free space tree                 (0:01:01 elapsed, 8253 items checked)
>     [4/7] checking fs roots                        (0:20:30 elapsed, 10978 items checked)
>     mirror 2 bytenr 870744010752 csum 0xdb1b27f0ca1a0139f7c65a0c0698a9a3f9e6ca6d624da7f70eecb3fc0f14ffc7 expected csum 0x2bce1cca32d98c3f83087f09980770a101b0560b1ddde7919fbbcd58a75f7d6b
>     mirror 1 bytenr 2004063948800 csum 0xe3f0a16cc8f03705a89f81178d4617c2847d660b7171abe29b65b5b394a9aace expected csum 0xc85b7f37fb620e1a68754692fd7ca43846ca316de6485fc8a4b447bfeab78d78
>     mirror 2 bytenr 3575828525056 csum 0xb9d9ee193d29b59b2015efed6151029c340a051c0338ec7ebca200363d304be8 expected csum 0xd645b7e5add1aa540a3440b78b7d31daa9545c925d32d2650d6bc61b7fdf4813
>     mirror 1 bytenr 3714124914688 csum 0xcb2ff575e84a8e965b94bc8faf9e76d8f645742ee9cd503609efd78ac22623e4 expected csum 0x966b1d021450ffb6c47759161533e185f06a14a50c30ff881097a43b7ad6d6cc
>     mirror 2 bytenr 4211891310592 csum 0xd1c11dabfa4bf3acea463479ab444bbc4c66dc9ba3257f09f4e1815bb46afac2 expected csum 0x17fff4d14269c69d84d267466c577ced3787fd9a9a445e36642067c47f129601
>     mirror 2 bytenr 4328552914944 csum 0x02f75c04f1d921ce34f5b6b9bf40c3b0056971fc89989dad227fc45723938472 expected csum 0xdd6416d001ac47f16e68a24d1438244152355a0774142d4885e5a031c6938d93
>     mirror 2 bytenr 4681011163136 csum 0x44c0b9d90fb258659e7377e93a96039afcadb501559a0cd831bf8c36f8fa1b2b expected csum 0x36c59223538807fc604537be5d686900031866626f3a1dc788755e92a74869cb
>     mirror 1 bytenr 4808263344128 csum 0xfd6278ce98b1f15c8aefea981e5ba7a521fa1a08d0f642185abf72e215288618 expected csum 0x1108b8b604c5b21f7d7d80d32a1fd9c0c0e753cad8fb97967dfa3525105bf808
>     mirror 1 bytenr 4993017057280 csum 0x033440e421102ec0c7b3057eae89dd80f300aecba70d3f1fcf5fe81c2cd6faba expected csum 0xb0e94e343b7291df9aae42a079bd0bba307a1c2ab81315361a49b0d8b6d53f49
>     mirror 2 bytenr 5037246173184 csum 0x00b807ff8d0a31a79bb947e774c077916b0bd162dbf24fe43e4fca179e364214 expected csum 0x984635996e4c3fed6c701519e9f654b1b0adb06f3921522f6b687bbccd730271
>     mirror 1 bytenr 5316786614272 csum 0x2805b36b130667ffaf187f1cc6bdab0802e5eb26332f3bae572202c9c34585dd expected csum 0xd7a2bab5bfddacddb610c6044504474c88962ed22d837a881faba8e4e60bee40
>     mirror 2 bytenr 7293991006208 csum 0x68f917b93233bb6c7e49775c0c0deed66c47fe0027f3d4287a3b2c736fb25a1b expected csum 0x3662cf0b4a35cd0a0b71a475f35a92a6fefdcf21b7ee59a97b261bb6fbda1c8a
>     [5/7] checking csums against data              (33:30:05 elapsed, 6082759 items checked)
>     ERROR: errors found in csum tree
>     [6/7] checking root refs                       (0:00:00 elapsed, 3 items checked)
>     [8/8] checking quota groups skipped (not enabled on this FS)
>     found 8848080683008 bytes used, error(s) found
>     total csum bytes: 68538889856
>     total tree bytes: 75102781440
>     total fs tree bytes: 180338688
>     total extent tree bytes: 340852736
>     btree space waste bytes: 5344101189
>     file data blocks allocated: 8776199127040
>      referenced 8772977901568
> 
> The corruptions looked similar to the first time, couldn't get any new
> clues out of them. Unfortunately it was on LUKS again because I'd given
> up on reproducing this, I'll try without LUKS going forward.
> 
> I realized the common factor betwen the original repro and this one was
> that I was additionally running an sftp copy of the same files over the
> network while the files were being copied between the two local volumes.
> 
>     mkfs.btrfs -m raid1 -d raid1 --csum blake2 /dev/sda /dev/sdb
>     mount /dev/sda -o compress=zstd:1 /mnt
>     rsync -Pav /data/ /mnt
> 
> ...and then, concurrently from another machine:
> 
>     sftp -ar nas:/data .
> 
> Can anybody else reproduce this corruption with that combo?
> 
> Otherwise, I'll keep working on narrowing this down: the tests take ages
> to run but require very little actual human time from me, so I'm happy
> to keep trying. Any other suggestions are welcome!

Hi Calvin,

How about replacing one of your hard drives with a product from a completely 
different manufacturer for isolate the issue?
I've seen this type of problem occur twice due to hardware, 
and neither of them were using btrfs, md or luks.

The first case was about 12 years ago. I had a RAID-Z2 (RAID6) configured 
with ZFS on Linux. There was a problem only with the HDDs connected to the 
Marvell 88SE9128, and the problem was resolved when I changed to a different 
SATA controller.
My memory is a bit hazy, but I believe ZFS detected a checksum inconsistency 
during scrubbing only on the disk connected with the Marvell.
- However, since you say your USB enclosure has been in use for a long time, 
it seems the controller is not the issue.

The other case was relatively recent, and I can't go into details for some reasons, 
but there was a problem with the disk's firmware which resulted in data corruption.

Thanks,
Akagi

> 
> Thanks,
> Calvin
> 


^ permalink raw reply

* [PATCH V4 0/3] block: ignore discard return value
From: Chaitanya Kulkarni @ 2026-02-11 20:44 UTC (permalink / raw)
  To: linan122, song, yukuai, hch, sagi, axboe
  Cc: linux-raid, linux-nvme, linux-block, Chaitanya Kulkarni

Hi,

This version has remaining patches from the original series to change
__blkdev_issue_discard() return type from int to void.

First two patches make md and nvmet drivers ignore the return value of
__blkdev_issue_discard() call and last patch changes the actual function
return type from int to void.

-ck

V3->V4:-

* Rebase md and nvmet patch on latest linux-block/for-next
* Add a last patch to change the return type of
  __blkdev_issue_discard() from int to void now that all call sites
  are fixed.

Chaitanya Kulkarni (3):
  md: ignore discard return value
  nvmet: ignore discard return value
  block: change return type to void

 block/blk-lib.c                   |  3 +--
 drivers/md/md.c                   |  4 ++--
 drivers/nvme/target/io-cmd-bdev.c | 28 +++++++---------------------
 include/linux/blkdev.h            |  2 +-
 4 files changed, 11 insertions(+), 26 deletions(-)

-- 
2.39.5


^ permalink raw reply

* [PATCH V4 1/3] md: ignore discard return value
From: Chaitanya Kulkarni @ 2026-02-11 20:44 UTC (permalink / raw)
  To: linan122, song, yukuai, hch, sagi, axboe
  Cc: linux-raid, linux-nvme, linux-block, Chaitanya Kulkarni,
	Martin K . Petersen, Johannes Thumshirn
In-Reply-To: <20260211204438.22568-1-kch@nvidia.com>

__blkdev_issue_discard() always returns 0, making all error checking at
call sites dead code.

Simplify md to only check !discard_bio by ignoring the
__blkdev_issue_discard() value.

Reviewed-by: Martin K. Petersen <martin.petersen@oracle.com>
Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Chaitanya Kulkarni <kch@nvidia.com>
---
 drivers/md/md.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/drivers/md/md.c b/drivers/md/md.c
index 59cd303548de..89c9e63a9139 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -9179,8 +9179,8 @@ void md_submit_discard_bio(struct mddev *mddev, struct md_rdev *rdev,
 {
 	struct bio *discard_bio = NULL;
 
-	if (__blkdev_issue_discard(rdev->bdev, start, size, GFP_NOIO,
-			&discard_bio) || !discard_bio)
+	__blkdev_issue_discard(rdev->bdev, start, size, GFP_NOIO, &discard_bio);
+	if (!discard_bio)
 		return;
 
 	bio_chain(discard_bio, bio);
-- 
2.39.5


^ permalink raw reply related

* [PATCH V4 2/3] nvmet: ignore discard return value
From: Chaitanya Kulkarni @ 2026-02-11 20:44 UTC (permalink / raw)
  To: linan122, song, yukuai, hch, sagi, axboe
  Cc: linux-raid, linux-nvme, linux-block, Chaitanya Kulkarni,
	Martin K . Petersen, Johannes Thumshirn
In-Reply-To: <20260211204438.22568-1-kch@nvidia.com>

__blkdev_issue_discard() always returns 0, making the error checking
in nvmet_bdev_discard_range() dead code.

Kill the function nvmet_bdev_discard_range() and call
__blkdev_issue_discard() directly from nvmet_bdev_execute_discard(),
since no error handling is needed anymore for __blkdev_issue_discard()
call.

Reviewed-by: Martin K. Petersen <martin.petersen@oracle.com>
Reviewed-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Chaitanya Kulkarni <kch@nvidia.com>
---
 drivers/nvme/target/io-cmd-bdev.c | 28 +++++++---------------------
 1 file changed, 7 insertions(+), 21 deletions(-)

diff --git a/drivers/nvme/target/io-cmd-bdev.c b/drivers/nvme/target/io-cmd-bdev.c
index 0103815542d4..f15d1c213bc6 100644
--- a/drivers/nvme/target/io-cmd-bdev.c
+++ b/drivers/nvme/target/io-cmd-bdev.c
@@ -363,29 +363,14 @@ u16 nvmet_bdev_flush(struct nvmet_req *req)
 	return 0;
 }
 
-static u16 nvmet_bdev_discard_range(struct nvmet_req *req,
-		struct nvme_dsm_range *range, struct bio **bio)
-{
-	struct nvmet_ns *ns = req->ns;
-	int ret;
-
-	ret = __blkdev_issue_discard(ns->bdev,
-			nvmet_lba_to_sect(ns, range->slba),
-			le32_to_cpu(range->nlb) << (ns->blksize_shift - 9),
-			GFP_KERNEL, bio);
-	if (ret && ret != -EOPNOTSUPP) {
-		req->error_slba = le64_to_cpu(range->slba);
-		return errno_to_nvme_status(req, ret);
-	}
-	return NVME_SC_SUCCESS;
-}
-
 static void nvmet_bdev_execute_discard(struct nvmet_req *req)
 {
+	struct nvmet_ns *ns = req->ns;
 	struct nvme_dsm_range range;
 	struct bio *bio = NULL;
+	sector_t nr_sects;
 	int i;
-	u16 status;
+	u16 status = NVME_SC_SUCCESS;
 
 	for (i = 0; i <= le32_to_cpu(req->cmd->dsm.nr); i++) {
 		status = nvmet_copy_from_sgl(req, i * sizeof(range), &range,
@@ -393,9 +378,10 @@ static void nvmet_bdev_execute_discard(struct nvmet_req *req)
 		if (status)
 			break;
 
-		status = nvmet_bdev_discard_range(req, &range, &bio);
-		if (status)
-			break;
+		nr_sects = le32_to_cpu(range.nlb) << (ns->blksize_shift - 9);
+		__blkdev_issue_discard(ns->bdev,
+				nvmet_lba_to_sect(ns, range.slba), nr_sects,
+				GFP_KERNEL, &bio);
 	}
 
 	if (bio) {
-- 
2.39.5


^ permalink raw reply related

* [PATCH V4 3/3] block: change return type to void
From: Chaitanya Kulkarni @ 2026-02-11 20:44 UTC (permalink / raw)
  To: linan122, song, yukuai, hch, sagi, axboe
  Cc: linux-raid, linux-nvme, linux-block, Chaitanya Kulkarni
In-Reply-To: <20260211204438.22568-1-kch@nvidia.com>

Now that all the callers of __blkdev_issue_discard() have been changed
to ignore its return value, change its return type from int to void.

Signed-off-by: Chaitanya Kulkarni <kch@nvidia.com>
---
 block/blk-lib.c        | 3 +--
 include/linux/blkdev.h | 2 +-
 2 files changed, 2 insertions(+), 3 deletions(-)

diff --git a/block/blk-lib.c b/block/blk-lib.c
index 0be3acdc3eb5..3213afc7f0d5 100644
--- a/block/blk-lib.c
+++ b/block/blk-lib.c
@@ -60,7 +60,7 @@ struct bio *blk_alloc_discard_bio(struct block_device *bdev,
 	return bio;
 }
 
-int __blkdev_issue_discard(struct block_device *bdev, sector_t sector,
+void __blkdev_issue_discard(struct block_device *bdev, sector_t sector,
 		sector_t nr_sects, gfp_t gfp_mask, struct bio **biop)
 {
 	struct bio *bio;
@@ -68,7 +68,6 @@ int __blkdev_issue_discard(struct block_device *bdev, sector_t sector,
 	while ((bio = blk_alloc_discard_bio(bdev, &sector, &nr_sects,
 			gfp_mask)))
 		*biop = bio_chain_and_submit(*biop, bio);
-	return 0;
 }
 EXPORT_SYMBOL(__blkdev_issue_discard);
 
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index 2ae4c45e4959..5ac4c7f2c2c0 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1258,7 +1258,7 @@ extern void blk_io_schedule(void);
 
 int blkdev_issue_discard(struct block_device *bdev, sector_t sector,
 		sector_t nr_sects, gfp_t gfp_mask);
-int __blkdev_issue_discard(struct block_device *bdev, sector_t sector,
+void __blkdev_issue_discard(struct block_device *bdev, sector_t sector,
 		sector_t nr_sects, gfp_t gfp_mask, struct bio **biop);
 int blkdev_issue_secure_erase(struct block_device *bdev, sector_t sector,
 		sector_t nr_sects, gfp_t gfp);
-- 
2.39.5


^ permalink raw reply related

* Re: [PATCH V4 0/3] block: ignore discard return value
From: Christoph Hellwig @ 2026-02-12 10:15 UTC (permalink / raw)
  To: Chaitanya Kulkarni
  Cc: linan122, song, yukuai, hch, sagi, axboe, linux-raid, linux-nvme,
	linux-block
In-Reply-To: <20260211204438.22568-1-kch@nvidia.com>

Looks good:

Reviewed-by: Christoph Hellwig <hch@lst.de>

And maybe Jens can pick this up for 7.0 so that we have a clean
slate for the new merge window?


^ permalink raw reply

* Re: [PATCH V4 0/3] block: ignore discard return value
From: Jens Axboe @ 2026-02-12 11:27 UTC (permalink / raw)
  To: linan122, song, yukuai, hch, sagi, Chaitanya Kulkarni
  Cc: linux-raid, linux-nvme, linux-block
In-Reply-To: <20260211204438.22568-1-kch@nvidia.com>


On Wed, 11 Feb 2026 12:44:35 -0800, Chaitanya Kulkarni wrote:
> This version has remaining patches from the original series to change
> __blkdev_issue_discard() return type from int to void.
> 
> First two patches make md and nvmet drivers ignore the return value of
> __blkdev_issue_discard() call and last patch changes the actual function
> return type from int to void.
> 
> [...]

Applied, thanks!

[1/3] md: ignore discard return value
      commit: 699fcfb6cb80a9df67fd2086a1c930d196d709f2
[2/3] nvmet: ignore discard return value
      commit: 38d12f15c4772b5383b1249b2afb0d206a430f0f
[3/3] block: change return type to void
      commit: 453daece381e60df20da16c49ccc6a9bc5c6515a

Best regards,
-- 
Jens Axboe




^ permalink raw reply

* Re: [PATCH v2 2/3] lib/raid6: Optimizing the raid6_select_algo time through asynchronous processing
From: kernel test robot @ 2026-02-13  2:34 UTC (permalink / raw)
  To: sunliming
  Cc: oe-lkp, lkp, linux-raid, song, yukuai, linux-kernel, linan666,
	sunliming, oliver.sang
In-Reply-To: <20260206054341.106878-3-sunliming@linux.dev>



Hello,

kernel test robot noticed "Oops:int3:#[##]SMP_KASAN_NOPTI" on:

commit: 9931f8d4a042e5e2cbf74d9ee99ece44da404284 ("[PATCH v2 2/3] lib/raid6: Optimizing the raid6_select_algo time through asynchronous processing")
url: https://github.com/intel-lab-lkp/linux/commits/sunliming-linux-dev/lib-raid6-Divide-the-raid6-algorithm-selection-process-into-two-parts/20260206-134850
base: https://git.kernel.org/cgit/linux/kernel/git/akpm/mm.git mm-nonmm-unstable
patch link: https://lore.kernel.org/all/20260206054341.106878-3-sunliming@linux.dev/
patch subject: [PATCH v2 2/3] lib/raid6: Optimizing the raid6_select_algo time through asynchronous processing

in testcase: perf-sanity-tests
version: 
with following parameters:

	perf_compiler: gcc
	group: group-00



config: x86_64-rhel-9.4-bpf
compiler: gcc-14
test machine: 16 threads 1 sockets Intel(R) Xeon(R) E-2278G CPU @ 3.40GHz (Coffee Lake-E) with 32G memory

(please refer to attached dmesg/kmsg for entire log/backtrace)



If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <oliver.sang@intel.com>
| Closes: https://lore.kernel.org/oe-lkp/202602131016.5ce5a38e-lkp@intel.com




[   81.049609][   C15] Oops: int3: 0000 [#1] SMP KASAN NOPTI
[   81.049615][   C15] CPU: 15 UID: 0 PID: 109 Comm: kworker/15:0 Tainted: G S                  6.19.0-rc6-00175-g9931f8d4a042 #1 PREEMPT(full)
[   81.049621][   C15] Tainted: [S]=CPU_OUT_OF_SPEC
[   81.049622][   C15] Hardware name: Intel Corporation Mehlow UP Server Platform/Moss Beach Server, BIOS CNLSE2R1.R00.X188.B13.1903250419 03/25/2019
[   81.049625][   C15] Workqueue: events benchmark_work_func [raid6_pq]
[   81.049631][   C15] RIP: 0010:benchmark_work_func (kbuild/src/consumer/lib/raid6/recov.c:65) raid6_pq
[   81.049650][   C15] Code: cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc <cc> cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc cc
All code
========
   0:	cc                   	int3
   1:	cc                   	int3
   2:	cc                   	int3
   3:	cc                   	int3
   4:	cc                   	int3
   5:	cc                   	int3
   6:	cc                   	int3
   7:	cc                   	int3
   8:	cc                   	int3
   9:	cc                   	int3
   a:	cc                   	int3
   b:	cc                   	int3
   c:	cc                   	int3
   d:	cc                   	int3
   e:	cc                   	int3
   f:	cc                   	int3
  10:	cc                   	int3
  11:	cc                   	int3
  12:	cc                   	int3
  13:	cc                   	int3
  14:	cc                   	int3
  15:	cc                   	int3
  16:	cc                   	int3
  17:	cc                   	int3
  18:	cc                   	int3
  19:	cc                   	int3
  1a:	cc                   	int3
  1b:	cc                   	int3
  1c:	cc                   	int3
  1d:	cc                   	int3
  1e:	cc                   	int3
  1f:	cc                   	int3
  20:	cc                   	int3
  21:	cc                   	int3
  22:	cc                   	int3
  23:	cc                   	int3
  24:	cc                   	int3
  25:	cc                   	int3
  26:	cc                   	int3
  27:	cc                   	int3
  28:	cc                   	int3
  29:	cc                   	int3
  2a:*	cc                   	int3		<-- trapping instruction
  2b:	cc                   	int3
  2c:	cc                   	int3
  2d:	cc                   	int3
  2e:	cc                   	int3
  2f:	cc                   	int3
  30:	cc                   	int3
  31:	cc                   	int3
  32:	cc                   	int3
  33:	cc                   	int3
  34:	cc                   	int3
  35:	cc                   	int3
  36:	cc                   	int3
  37:	cc                   	int3
  38:	cc                   	int3
  39:	cc                   	int3
  3a:	cc                   	int3
  3b:	cc                   	int3
  3c:	cc                   	int3
  3d:	cc                   	int3
  3e:	cc                   	int3
  3f:	cc                   	int3

Code starting with the faulting instruction
===========================================
   0:	cc                   	int3
   1:	cc                   	int3
   2:	cc                   	int3
   3:	cc                   	int3
   4:	cc                   	int3
   5:	cc                   	int3
   6:	cc                   	int3
   7:	cc                   	int3
   8:	cc                   	int3
   9:	cc                   	int3
   a:	cc                   	int3
   b:	cc                   	int3
   c:	cc                   	int3
   d:	cc                   	int3
   e:	cc                   	int3
   f:	cc                   	int3
  10:	cc                   	int3
  11:	cc                   	int3
  12:	cc                   	int3
  13:	cc                   	int3
  14:	cc                   	int3
  15:	cc                   	int3
[   81.049653][   C15] RSP: 0018:ffff888101557b58 EFLAGS: 00000286
[   81.049656][   C15] RAX: ffffffffc115b5a0 RBX: 1ffff110202aaf6b RCX: 0000000000000000
[   81.049658][   C15] RDX: ffffffffc115b6e0 RSI: ffffffff81513d9c RDI: ffffed10202aaf56
[   81.049660][   C15] RBP: ffff88886bdf0000 R08: 0000000000000000 R09: fffffbfff0b21e14
[   81.049662][   C15] R10: ffffffff8590f0a7 R11: 0000000000000001 R12: ffff888101557bb8
[   81.049664][   C15] R13: ffff88886bdf8000 R14: ffff888101557bb8 R15: ffff8887887c33c0
[   81.049667][   C15] FS:  0000000000000000(0000) GS:ffff888801e6a000(0000) knlGS:0000000000000000
[   81.049669][   C15] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[   81.049671][   C15] CR2: 00007f33621da000 CR3: 00000002f20ed006 CR4: 00000000003726f0
[   81.049674][   C15] Call Trace:
[   81.049676][   C15]  <TASK>
[   81.049680][   C15]  ? rcu_is_watching (kbuild/src/consumer/arch/x86/include/asm/atomic.h:23 kbuild/src/consumer/include/linux/atomic/atomic-arch-fallback.h:457 kbuild/src/consumer/include/linux/context_tracking.h:128 kbuild/src/consumer/kernel/rcu/tree.c:751)
[   81.049685][   C15]  ? rcu_is_watching (kbuild/src/consumer/arch/x86/include/asm/atomic.h:23 kbuild/src/consumer/include/linux/atomic/atomic-arch-fallback.h:457 kbuild/src/consumer/include/linux/context_tracking.h:128 kbuild/src/consumer/kernel/rcu/tree.c:751)
[   81.049689][   C15]  ? lock_acquire (kbuild/src/consumer/include/trace/events/lock.h:24 (discriminator 2) kbuild/src/consumer/kernel/locking/lockdep.c:5831 (discriminator 2))
[   81.049694][   C15]  ? process_one_work (kbuild/src/consumer/arch/x86/include/asm/jump_label.h:37 kbuild/src/consumer/include/trace/events/workqueue.h:110 kbuild/src/consumer/kernel/workqueue.c:3262)


The kernel config and materials to reproduce are available at:
https://download.01.org/0day-ci/archive/20260213/202602131016.5ce5a38e-lkp@intel.com



-- 
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki


^ permalink raw reply

* [PATCH v3 1/3] lib/raid6: Divide the raid6 algorithm selection process into two parts
From: sunliming @ 2026-02-13  3:14 UTC (permalink / raw)
  To: song, yukuai; +Cc: linux-raid, linux-kernel, linan666, sunliming
In-Reply-To: <20260213031419.125069-1-sunliming@linux.dev>

From: sunliming <sunliming@kylinos.cn>

Divide the RAID6 algorithm selection process into two parts: fast selection
and benchmark selection. To prepare for the asynchronous processing of
the benchmark phase.

Signed-off-by: sunliming <sunliming@kylinos.cn>
---
 lib/raid6/algos.c | 76 +++++++++++++++++++++++++++++++++--------------
 1 file changed, 54 insertions(+), 22 deletions(-)

diff --git a/lib/raid6/algos.c b/lib/raid6/algos.c
index 799e0e5eac26..c21e3ad99d97 100644
--- a/lib/raid6/algos.c
+++ b/lib/raid6/algos.c
@@ -152,8 +152,32 @@ static inline const struct raid6_recov_calls *raid6_choose_recov(void)
 	return best;
 }
 
-static inline const struct raid6_calls *raid6_choose_gen(
-	void *(*const dptrs)[RAID6_TEST_DISKS], const int disks)
+/* Quick selection: select the highest priority valid algorithm. */
+static inline int raid6_choose_gen_fast(void)
+{
+	int ret = 0;
+	const struct raid6_calls *const *algo;
+	const struct raid6_calls *best = NULL;
+
+	for (best = NULL, algo = raid6_algos; *algo; algo++)
+		if (!best || (*algo)->priority > best->priority)
+			if (!(*algo)->valid || (*algo)->valid())
+				best = *algo;
+
+	if (best) {
+		raid6_call = *best;
+		pr_info("raid6: skipped pq benchmark and selected %s\n",
+				best->name);
+	} else {
+		pr_err("raid6: No valid algorithm found even for fast selection!\n");
+		ret = -EINVAL;
+	}
+
+	return ret;
+}
+
+static inline const struct raid6_calls *raid6_gen_benchmark(
+		void *(*const dptrs)[RAID6_TEST_DISKS], const int disks)
 {
 	unsigned long perf, bestgenperf, j0, j1;
 	int start = (disks>>1)-1, stop = disks-3;	/* work on the second half of the disks */
@@ -165,11 +189,6 @@ static inline const struct raid6_calls *raid6_choose_gen(
 			if ((*algo)->valid && !(*algo)->valid())
 				continue;
 
-			if (!IS_ENABLED(CONFIG_RAID6_PQ_BENCHMARK)) {
-				best = *algo;
-				break;
-			}
-
 			perf = 0;
 
 			preempt_disable();
@@ -200,12 +219,6 @@ static inline const struct raid6_calls *raid6_choose_gen(
 
 	raid6_call = *best;
 
-	if (!IS_ENABLED(CONFIG_RAID6_PQ_BENCHMARK)) {
-		pr_info("raid6: skipped pq benchmark and selected %s\n",
-			best->name);
-		goto out;
-	}
-
 	pr_info("raid6: using algorithm %s gen() %ld MB/s\n",
 		best->name,
 		(bestgenperf * HZ * (disks - 2)) >>
@@ -239,15 +252,13 @@ static inline const struct raid6_calls *raid6_choose_gen(
 /* Try to pick the best algorithm */
 /* This code uses the gfmul table as convenient data set to abuse */
 
-int __init raid6_select_algo(void)
+static int raid6_choose_gen_benmark(void)
 {
 	const int disks = RAID6_TEST_DISKS;
-
 	const struct raid6_calls *gen_best;
-	const struct raid6_recov_calls *rec_best;
 	char *disk_ptr, *p;
 	void *dptrs[RAID6_TEST_DISKS];
-	int i, cycle;
+	int i, cycle, ret = 0;
 
 	/* prepare the buffer and fill it circularly with gfmul table */
 	disk_ptr = (char *)__get_free_pages(GFP_KERNEL, RAID6_TEST_DISKS_ORDER);
@@ -269,15 +280,36 @@ int __init raid6_select_algo(void)
 	if ((disks - 2) * PAGE_SIZE % 65536)
 		memcpy(p, raid6_gfmul, (disks - 2) * PAGE_SIZE % 65536);
 
-	/* select raid gen_syndrome function */
-	gen_best = raid6_choose_gen(&dptrs, disks);
+	gen_best = raid6_gen_benchmark(&dptrs, disks);
+	if (!gen_best)
+		ret = -EINVAL;
+
+	free_pages((unsigned long)disk_ptr, RAID6_TEST_DISKS_ORDER);
+
+	return ret;
+}
+
+int __init raid6_select_algo(void)
+{
+	int ret = 0;
+	const struct raid6_recov_calls *rec_best = NULL;
+
+	/* select raid gen_syndrome functions */
+	if (!IS_ENABLED(CONFIG_RAID6_PQ_BENCHMARK))
+		ret = raid6_choose_gen_fast();
+	else
+		ret = raid6_choose_gen_benmark();
+
+	if (ret < 0)
+		goto out;
 
 	/* select raid recover functions */
 	rec_best = raid6_choose_recov();
+	if (!rec_best)
+		ret = -EINVAL;
 
-	free_pages((unsigned long)disk_ptr, RAID6_TEST_DISKS_ORDER);
-
-	return gen_best && rec_best ? 0 : -EINVAL;
+out:
+	return ret;
 }
 
 static void raid6_exit(void)
-- 
2.25.1


^ permalink raw reply related

* [PATCH v3 0/3] lib/raid6: Optimize raid6_select_algo to ensure
From: sunliming @ 2026-02-13  3:14 UTC (permalink / raw)
  To: song, yukuai; +Cc: linux-raid, linux-kernel, linan666, sunliming

From: sunliming <sunliming@kylinos.cn>

The selection of RAID6 PQ functions involves a dual-strategy approach:
prioritizing startup speed leads to quickly choosing a usable algorithm,
while prioritizing performance leads to selecting an optimal algorithm
via benchmarking. This choice is determined by the RAID6_PQ_BENCHMARK
configuration. This patch series achieves both fast startup and optimal
algorithm selection by initially choosing an algorithm quickly at startup,
then asynchronously determining the optimal algorithm through benchmarking.
Since all RAID6 PQ function algorithms are functionally equivalent despite
performance differences, this approach should be effective.

---
Changes in v3:
  Remove the __init annotation from raid6_benchmark_work and benchmark_work_func,
as they are still referenced during initialization.
Changes in v2:
  Select the highest-priority algorithm instead of the first one. 
  Add the cancel_work_sync function in the exit function to handle the 
work queue cleanup.
--- 
sunliming (3):
  ib/raid6: Divide the raid6 algorithm selection process into two parts
  lib/raid6: Optimizing the raid6_select_algo time through asynchronous
    processing
  lib/raid6: Delete the RAID6_PQ_BENCHMARK config

 include/linux/raid/pq.h |  3 --
 lib/Kconfig             |  8 ----
 lib/raid6/algos.c       | 88 ++++++++++++++++++++++++++++++-----------
 3 files changed, 65 insertions(+), 34 deletions(-)

-- 
2.25.1


^ permalink raw reply

* [PATCH v3 2/3] lib/raid6: Optimizing the raid6_select_algo time through asynchronous processing
From: sunliming @ 2026-02-13  3:14 UTC (permalink / raw)
  To: song, yukuai; +Cc: linux-raid, linux-kernel, linan666, sunliming
In-Reply-To: <20260213031419.125069-1-sunliming@linux.dev>

From: sunliming <sunliming@kylinos.cn>

Optimizing the raid6_select_algo time. In raid6_select_algo(), an raid6 gen
algorithm is first selected quickly through synchronous processing, while
the time-consuming process of selecting the optimal algorithm via benchmarking
is handled asynchronously. This approach speeds up the overall startup time
and ultimately ensures the selection of an optimal algorithm.

Signed-off-by: sunliming <sunliming@kylinos.cn>
---
 lib/raid6/algos.c | 30 ++++++++++++++++++++----------
 1 file changed, 20 insertions(+), 10 deletions(-)

diff --git a/lib/raid6/algos.c b/lib/raid6/algos.c
index c21e3ad99d97..b8b5515ac7a6 100644
--- a/lib/raid6/algos.c
+++ b/lib/raid6/algos.c
@@ -12,6 +12,7 @@
  */
 
 #include <linux/raid/pq.h>
+#include <linux/workqueue.h>
 #ifndef __KERNEL__
 #include <sys/mman.h>
 #include <stdio.h>
@@ -166,7 +167,7 @@ static inline int raid6_choose_gen_fast(void)
 
 	if (best) {
 		raid6_call = *best;
-		pr_info("raid6: skipped pq benchmark and selected %s\n",
+		pr_info("raid6: raid6: fast selected %s, async benchmark pending\n",
 				best->name);
 	} else {
 		pr_err("raid6: No valid algorithm found even for fast selection!\n");
@@ -213,7 +214,7 @@ static inline const struct raid6_calls *raid6_gen_benchmark(
 	}
 
 	if (!best) {
-		pr_err("raid6: Yikes! No algorithm found!\n");
+		pr_warn("raid6: async benchmark failed to find any algorithm\n");
 		goto out;
 	}
 
@@ -289,24 +290,33 @@ static int raid6_choose_gen_benmark(void)
 	return ret;
 }
 
+static struct work_struct raid6_benchmark_work;
+
+static void benchmark_work_func(struct work_struct *work)
+{
+	raid6_choose_gen_benmark();
+}
+
 int __init raid6_select_algo(void)
 {
 	int ret = 0;
 	const struct raid6_recov_calls *rec_best = NULL;
 
-	/* select raid gen_syndrome functions */
-	if (!IS_ENABLED(CONFIG_RAID6_PQ_BENCHMARK))
-		ret = raid6_choose_gen_fast();
-	else
-		ret = raid6_choose_gen_benmark();
-
+	/* phase 1: synchronous fast selection generation algorithm */
+	ret = raid6_choose_gen_fast();
 	if (ret < 0)
 		goto out;
 
 	/* select raid recover functions */
 	rec_best = raid6_choose_recov();
-	if (!rec_best)
+	if (!rec_best) {
 		ret = -EINVAL;
+		goto out;
+	}
+
+	/* phase 2: asynchronous performance benchmarking */
+	INIT_WORK(&raid6_benchmark_work, benchmark_work_func);
+	schedule_work(&raid6_benchmark_work);
 
 out:
 	return ret;
@@ -314,7 +324,7 @@ int __init raid6_select_algo(void)
 
 static void raid6_exit(void)
 {
-	do { } while (0);
+	cancel_work_sync(&raid6_benchmark_work);
 }
 
 subsys_initcall(raid6_select_algo);
-- 
2.25.1


^ permalink raw reply related

* [PATCH v3 3/3] lib/raid6: Delete the RAID6_PQ_BENCHMARK config
From: sunliming @ 2026-02-13  3:14 UTC (permalink / raw)
  To: song, yukuai; +Cc: linux-raid, linux-kernel, linan666, sunliming
In-Reply-To: <20260213031419.125069-1-sunliming@linux.dev>

From: sunliming <sunliming@kylinos.cn>

Now RAID6 PQ functions is automatically choosed and the
RAID6_PQ_BENCHMARK is not needed.

Signed-off-by: sunliming <sunliming@kylinos.cn>
---
 include/linux/raid/pq.h | 3 ---
 lib/Kconfig             | 8 --------
 2 files changed, 11 deletions(-)

diff --git a/include/linux/raid/pq.h b/include/linux/raid/pq.h
index 2467b3be15c9..6378ec4ae4ba 100644
--- a/include/linux/raid/pq.h
+++ b/include/linux/raid/pq.h
@@ -67,9 +67,6 @@ extern const char raid6_empty_zero_page[PAGE_SIZE];
 #define MODULE_DESCRIPTION(desc)
 #define subsys_initcall(x)
 #define module_exit(x)
-
-#define IS_ENABLED(x) (x)
-#define CONFIG_RAID6_PQ_BENCHMARK 1
 #endif /* __KERNEL__ */
 
 /* Routine choices */
diff --git a/lib/Kconfig b/lib/Kconfig
index 2923924bea78..841a0245a2c4 100644
--- a/lib/Kconfig
+++ b/lib/Kconfig
@@ -11,14 +11,6 @@ menu "Library routines"
 config RAID6_PQ
 	tristate
 
-config RAID6_PQ_BENCHMARK
-	bool "Automatically choose fastest RAID6 PQ functions"
-	depends on RAID6_PQ
-	default y
-	help
-	  Benchmark all available RAID6 PQ functions on init and choose the
-	  fastest one.
-
 config LINEAR_RANGES
 	tristate
 
-- 
2.25.1


^ permalink raw reply related

* Re: [PATCH v3 0/3] lib/raid6: Optimize raid6_select_algo to ensure
From: Christoph Hellwig @ 2026-02-13  8:15 UTC (permalink / raw)
  To: sunliming; +Cc: song, yukuai, linux-raid, linux-kernel, linan666, sunliming
In-Reply-To: <20260213031419.125069-1-sunliming@linux.dev>

This is actually a very bad idea.

The current indirect calls for the raid parity generation are very
expensive and need to be replaced by static calls.  Doing runtime
switching defeats this.

I had started on doing the static calls a while ago, let me try to
finish that off.


^ permalink raw reply

* [PATCH 0/5] md/md-llbitmap: fixes and proactive parity building support
From: Yu Kuai @ 2026-02-14  6:10 UTC (permalink / raw)
  To: song; +Cc: linan122, xni, colyli, linux-raid, linux-kernel

This series contains fixes and enhancements for the md-llbitmap (lockless
bitmap) implementation.

Patches 1-2 are bug fixes:
- Patch 1 fixes bitmap data being read from spare disks that are not yet
  in sync, which could lead to incorrect dirty bit tracking.
- Patch 2 fixes a race condition where the state machine could transition
  before the barrier is properly raised.

Patch 3 improves compatibility with older mdadm versions by detecting
on-disk bitmap version and falling back to the correct bitmap_ops when
there's a version mismatch.

Patch 4 adds support for proactive XOR parity building in RAID-456 arrays.
This allows users to pre-build parity for unwritten regions via sysfs
before any user data is written, which can improve write performance for
workloads that will eventually use all storage. New states (CleanUnwritten,
NeedSyncUnwritten, SyncingUnwritten) are added to track these regions
separately from normal dirty/syncing states.

Patch 5 optimizes initial array sync for RAID-456 arrays on devices that
support write_zeroes with unmap. By zeroing all disks upfront, parity is
automatically consistent (0 XOR 0 = 0), allowing the bitmap to be
initialized to BitCleanUnwritten and skipping the initial sync entirely.
This significantly reduces array initialization time on modern NVMe SSDs.

Yu Kuai (5):
  md/md-llbitmap: skip reading rdevs that are not in_sync
  md/md-llbitmap: raise barrier before state machine transition
  md: add fallback to correct bitmap_ops on version mismatch
  md/md-llbitmap: add CleanUnwritten state for RAID-5 proactive parity
    building
  md/md-llbitmap: optimize initial sync with write_zeroes_unmap support

 drivers/md/md-llbitmap.c | 213 +++++++++++++++++++++++++++++++++++----
 drivers/md/md.c          | 109 +++++++++++++++++++-
 2 files changed, 301 insertions(+), 21 deletions(-)

-- 
2.51.0


^ permalink raw reply

* [PATCH 1/5] md/md-llbitmap: skip reading rdevs that are not in_sync
From: Yu Kuai @ 2026-02-14  6:10 UTC (permalink / raw)
  To: song; +Cc: linan122, xni, colyli, linux-raid, linux-kernel
In-Reply-To: <20260214061013.2335604-1-yukuai@fnnas.com>

When reading bitmap pages from member disks, the code iterates through
all rdevs and attempts to read from the first available one. However,
it only checks for raid_disk assignment and Faulty flag, missing the
In_sync flag check.

This can cause bitmap data to be read from spare disks that are still
being rebuilt and don't have valid bitmap information yet. Reading
stale or uninitialized bitmap data from such disks can lead to
incorrect dirty bit tracking, potentially causing data corruption
during recovery or normal operation.

Add the In_sync flag check to ensure bitmap pages are only read from
fully synchronized member disks that have valid bitmap data.

Cc: stable@vger.kernel.org
Fixes: 5ab829f1971d ("md/md-llbitmap: introduce new lockless bitmap")
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
 drivers/md/md-llbitmap.c | 3 ++-
 1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c
index cd713a7dc270..30d7e36b22c4 100644
--- a/drivers/md/md-llbitmap.c
+++ b/drivers/md/md-llbitmap.c
@@ -459,7 +459,8 @@ static struct page *llbitmap_read_page(struct llbitmap *llbitmap, int idx)
 	rdev_for_each(rdev, mddev) {
 		sector_t sector;
 
-		if (rdev->raid_disk < 0 || test_bit(Faulty, &rdev->flags))
+		if (rdev->raid_disk < 0 || test_bit(Faulty, &rdev->flags) ||
+		    !test_bit(In_sync, &rdev->flags))
 			continue;
 
 		sector = mddev->bitmap_info.offset +
-- 
2.51.0


^ permalink raw reply related

* [PATCH 2/5] md/md-llbitmap: raise barrier before state machine transition
From: Yu Kuai @ 2026-02-14  6:10 UTC (permalink / raw)
  To: song; +Cc: linan122, xni, colyli, linux-raid, linux-kernel
In-Reply-To: <20260214061013.2335604-1-yukuai@fnnas.com>

Move the barrier raise operation before calling llbitmap_state_machine()
in both llbitmap_start_write() and llbitmap_start_discard(). This
ensures the barrier is in place before any state transitions occur,
preventing potential race conditions where the state machine could
complete before the barrier is properly raised.

Cc: stable@vger.kernel.org
Fixes: 5ab829f1971d ("md/md-llbitmap: introduce new lockless bitmap")
Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
 drivers/md/md-llbitmap.c | 8 ++++----
 1 file changed, 4 insertions(+), 4 deletions(-)

diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c
index 30d7e36b22c4..5f9e7004e3e3 100644
--- a/drivers/md/md-llbitmap.c
+++ b/drivers/md/md-llbitmap.c
@@ -1070,12 +1070,12 @@ static void llbitmap_start_write(struct mddev *mddev, sector_t offset,
 	int page_start = (start + BITMAP_DATA_OFFSET) >> PAGE_SHIFT;
 	int page_end = (end + BITMAP_DATA_OFFSET) >> PAGE_SHIFT;
 
-	llbitmap_state_machine(llbitmap, start, end, BitmapActionStartwrite);
-
 	while (page_start <= page_end) {
 		llbitmap_raise_barrier(llbitmap, page_start);
 		page_start++;
 	}
+
+	llbitmap_state_machine(llbitmap, start, end, BitmapActionStartwrite);
 }
 
 static void llbitmap_end_write(struct mddev *mddev, sector_t offset,
@@ -1102,12 +1102,12 @@ static void llbitmap_start_discard(struct mddev *mddev, sector_t offset,
 	int page_start = (start + BITMAP_DATA_OFFSET) >> PAGE_SHIFT;
 	int page_end = (end + BITMAP_DATA_OFFSET) >> PAGE_SHIFT;
 
-	llbitmap_state_machine(llbitmap, start, end, BitmapActionDiscard);
-
 	while (page_start <= page_end) {
 		llbitmap_raise_barrier(llbitmap, page_start);
 		page_start++;
 	}
+
+	llbitmap_state_machine(llbitmap, start, end, BitmapActionDiscard);
 }
 
 static void llbitmap_end_discard(struct mddev *mddev, sector_t offset,
-- 
2.51.0


^ permalink raw reply related

* [PATCH 3/5] md: add fallback to correct bitmap_ops on version mismatch
From: Yu Kuai @ 2026-02-14  6:10 UTC (permalink / raw)
  To: song; +Cc: linan122, xni, colyli, linux-raid, linux-kernel
In-Reply-To: <20260214061013.2335604-1-yukuai@fnnas.com>

If default bitmap version and on-disk version doesn't match, and mdadm
is not the latest version to set bitmap_type, set bitmap_ops based on
the disk version.

Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
 drivers/md/md.c | 103 +++++++++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 102 insertions(+), 1 deletion(-)

diff --git a/drivers/md/md.c b/drivers/md/md.c
index 59cd303548de..d2607ed5c2e9 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -6447,15 +6447,116 @@ static void md_safemode_timeout(struct timer_list *t)
 
 static int start_dirty_degraded;
 
+/*
+ * Read bitmap superblock and return the bitmap_id based on disk version.
+ * This is used as fallback when default bitmap version and on-disk version
+ * doesn't match, and mdadm is not the latest version to set bitmap_type.
+ */
+static enum md_submodule_id md_bitmap_get_id_from_sb(struct mddev *mddev)
+{
+	struct md_rdev *rdev;
+	struct page *sb_page;
+	bitmap_super_t *sb;
+	enum md_submodule_id id = ID_BITMAP_NONE;
+	sector_t sector;
+	u32 version;
+
+	if (!mddev->bitmap_info.offset)
+		return ID_BITMAP_NONE;
+
+	sb_page = alloc_page(GFP_KERNEL);
+	if (!sb_page)
+		return ID_BITMAP_NONE;
+
+	sector = mddev->bitmap_info.offset;
+
+	rdev_for_each(rdev, mddev) {
+		u32 iosize;
+
+		if (!test_bit(In_sync, &rdev->flags) ||
+		    test_bit(Faulty, &rdev->flags) ||
+		    test_bit(Bitmap_sync, &rdev->flags))
+			continue;
+
+		iosize = roundup(sizeof(bitmap_super_t),
+				 bdev_logical_block_size(rdev->bdev));
+		if (sync_page_io(rdev, sector, iosize, sb_page, REQ_OP_READ,
+				 true))
+			goto read_ok;
+	}
+	goto out;
+
+read_ok:
+	sb = kmap_local_page(sb_page);
+	if (sb->magic != cpu_to_le32(BITMAP_MAGIC))
+		goto out_unmap;
+
+	version = le32_to_cpu(sb->version);
+	switch (version) {
+	case BITMAP_MAJOR_LO:
+	case BITMAP_MAJOR_HI:
+	case BITMAP_MAJOR_CLUSTERED:
+		id = ID_BITMAP;
+		break;
+	case BITMAP_MAJOR_LOCKLESS:
+		id = ID_LLBITMAP;
+		break;
+	default:
+		pr_warn("md: %s: unknown bitmap version %u\n",
+			mdname(mddev), version);
+		break;
+	}
+
+out_unmap:
+	kunmap_local(sb);
+out:
+	__free_page(sb_page);
+	return id;
+}
+
 static int md_bitmap_create(struct mddev *mddev)
 {
+	enum md_submodule_id orig_id = mddev->bitmap_id;
+	enum md_submodule_id sb_id;
+	int err;
+
 	if (mddev->bitmap_id == ID_BITMAP_NONE)
 		return -EINVAL;
 
 	if (!mddev_set_bitmap_ops(mddev))
 		return -ENOENT;
 
-	return mddev->bitmap_ops->create(mddev);
+	err = mddev->bitmap_ops->create(mddev);
+	if (!err)
+		return 0;
+
+	/*
+	 * Create failed, if default bitmap version and on-disk version
+	 * doesn't match, and mdadm is not the latest version to set
+	 * bitmap_type, set bitmap_ops based on the disk version.
+	 */
+	mddev_clear_bitmap_ops(mddev);
+
+	sb_id = md_bitmap_get_id_from_sb(mddev);
+	if (sb_id == ID_BITMAP_NONE || sb_id == orig_id)
+		return err;
+
+	pr_info("md: %s: bitmap version mismatch, switching from %d to %d\n",
+		mdname(mddev), orig_id, sb_id);
+
+	mddev->bitmap_id = sb_id;
+	if (!mddev_set_bitmap_ops(mddev)) {
+		mddev->bitmap_id = orig_id;
+		return -ENOENT;
+	}
+
+	err = mddev->bitmap_ops->create(mddev);
+	if (err) {
+		mddev_clear_bitmap_ops(mddev);
+		mddev->bitmap_id = orig_id;
+	}
+
+	return err;
 }
 
 static void md_bitmap_destroy(struct mddev *mddev)
-- 
2.51.0


^ permalink raw reply related

* [PATCH 4/5] md/md-llbitmap: add CleanUnwritten state for RAID-5 proactive parity building
From: Yu Kuai @ 2026-02-14  6:10 UTC (permalink / raw)
  To: song; +Cc: linan122, xni, colyli, linux-raid, linux-kernel
In-Reply-To: <20260214061013.2335604-1-yukuai@fnnas.com>

Add new states to the llbitmap state machine to support proactive XOR
parity building for RAID-5 arrays. This allows users to pre-build parity
data for unwritten regions before any user data is written.

New states added:
- BitNeedSyncUnwritten: Transitional state when proactive sync is triggered
  via sysfs on Unwritten regions.
- BitSyncingUnwritten: Proactive sync in progress for unwritten region.
- BitCleanUnwritten: XOR parity has been pre-built, but no user data
  written yet. When user writes to this region, it transitions to BitDirty.

New actions added:
- BitmapActionProactiveSync: Trigger for proactive XOR parity building.
- BitmapActionClearUnwritten: Convert CleanUnwritten/NeedSyncUnwritten/
  SyncingUnwritten states back to Unwritten before recovery starts.

State flows:
- Current (lazy): Unwritten -> (write) -> NeedSync -> (sync) -> Dirty -> Clean
- New (proactive): Unwritten -> (sysfs) -> NeedSyncUnwritten -> (sync) -> CleanUnwritten
- On write to CleanUnwritten: CleanUnwritten -> (write) -> Dirty -> Clean
- On disk replacement: CleanUnwritten regions are converted to Unwritten
  before recovery starts, so recovery only rebuilds regions with user data

A new sysfs interface is added at /sys/block/mdX/md/llbitmap/proactive_sync
(write-only) to trigger proactive sync. This only works for RAID-456 arrays.

Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
 drivers/md/md-llbitmap.c | 140 +++++++++++++++++++++++++++++++++++----
 drivers/md/md.c          |   6 +-
 2 files changed, 132 insertions(+), 14 deletions(-)

diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c
index 5f9e7004e3e3..461050b2771b 100644
--- a/drivers/md/md-llbitmap.c
+++ b/drivers/md/md-llbitmap.c
@@ -208,6 +208,20 @@ enum llbitmap_state {
 	BitNeedSync,
 	/* data is synchronizing */
 	BitSyncing,
+	/*
+	 * Proactive sync requested for unwritten region (raid456 only).
+	 * Triggered via sysfs when user wants to pre-build XOR parity
+	 * for regions that have never been written.
+	 */
+	BitNeedSyncUnwritten,
+	/* Proactive sync in progress for unwritten region */
+	BitSyncingUnwritten,
+	/*
+	 * XOR parity has been pre-built for a region that has never had
+	 * user data written. When user writes to this region, it transitions
+	 * to BitDirty.
+	 */
+	BitCleanUnwritten,
 	BitStateCount,
 	BitNone = 0xff,
 };
@@ -232,6 +246,12 @@ enum llbitmap_action {
 	 * BitNeedSync.
 	 */
 	BitmapActionStale,
+	/*
+	 * Proactive sync trigger for raid456 - builds XOR parity for
+	 * Unwritten regions without requiring user data write first.
+	 */
+	BitmapActionProactiveSync,
+	BitmapActionClearUnwritten,
 	BitmapActionCount,
 	/* Init state is BitUnwritten */
 	BitmapActionInit,
@@ -304,6 +324,8 @@ static char state_machine[BitStateCount][BitmapActionCount] = {
 		[BitmapActionDaemon]		= BitNone,
 		[BitmapActionDiscard]		= BitNone,
 		[BitmapActionStale]		= BitNone,
+		[BitmapActionProactiveSync]	= BitNeedSyncUnwritten,
+		[BitmapActionClearUnwritten]	= BitNone,
 	},
 	[BitClean] = {
 		[BitmapActionStartwrite]	= BitDirty,
@@ -314,6 +336,8 @@ static char state_machine[BitStateCount][BitmapActionCount] = {
 		[BitmapActionDaemon]		= BitNone,
 		[BitmapActionDiscard]		= BitUnwritten,
 		[BitmapActionStale]		= BitNeedSync,
+		[BitmapActionProactiveSync]	= BitNone,
+		[BitmapActionClearUnwritten]	= BitNone,
 	},
 	[BitDirty] = {
 		[BitmapActionStartwrite]	= BitNone,
@@ -324,6 +348,8 @@ static char state_machine[BitStateCount][BitmapActionCount] = {
 		[BitmapActionDaemon]		= BitClean,
 		[BitmapActionDiscard]		= BitUnwritten,
 		[BitmapActionStale]		= BitNeedSync,
+		[BitmapActionProactiveSync]	= BitNone,
+		[BitmapActionClearUnwritten]	= BitNone,
 	},
 	[BitNeedSync] = {
 		[BitmapActionStartwrite]	= BitNone,
@@ -334,6 +360,8 @@ static char state_machine[BitStateCount][BitmapActionCount] = {
 		[BitmapActionDaemon]		= BitNone,
 		[BitmapActionDiscard]		= BitUnwritten,
 		[BitmapActionStale]		= BitNone,
+		[BitmapActionProactiveSync]	= BitNone,
+		[BitmapActionClearUnwritten]	= BitNone,
 	},
 	[BitSyncing] = {
 		[BitmapActionStartwrite]	= BitNone,
@@ -344,6 +372,44 @@ static char state_machine[BitStateCount][BitmapActionCount] = {
 		[BitmapActionDaemon]		= BitNone,
 		[BitmapActionDiscard]		= BitUnwritten,
 		[BitmapActionStale]		= BitNeedSync,
+		[BitmapActionProactiveSync]	= BitNone,
+		[BitmapActionClearUnwritten]	= BitNone,
+	},
+	[BitNeedSyncUnwritten] = {
+		[BitmapActionStartwrite]	= BitNeedSync,
+		[BitmapActionStartsync]		= BitSyncingUnwritten,
+		[BitmapActionEndsync]		= BitNone,
+		[BitmapActionAbortsync]		= BitUnwritten,
+		[BitmapActionReload]		= BitUnwritten,
+		[BitmapActionDaemon]		= BitNone,
+		[BitmapActionDiscard]		= BitUnwritten,
+		[BitmapActionStale]		= BitUnwritten,
+		[BitmapActionProactiveSync]	= BitNone,
+		[BitmapActionClearUnwritten]	= BitUnwritten,
+	},
+	[BitSyncingUnwritten] = {
+		[BitmapActionStartwrite]	= BitSyncing,
+		[BitmapActionStartsync]		= BitSyncingUnwritten,
+		[BitmapActionEndsync]		= BitCleanUnwritten,
+		[BitmapActionAbortsync]		= BitUnwritten,
+		[BitmapActionReload]		= BitUnwritten,
+		[BitmapActionDaemon]		= BitNone,
+		[BitmapActionDiscard]		= BitUnwritten,
+		[BitmapActionStale]		= BitUnwritten,
+		[BitmapActionProactiveSync]	= BitNone,
+		[BitmapActionClearUnwritten]	= BitUnwritten,
+	},
+	[BitCleanUnwritten] = {
+		[BitmapActionStartwrite]	= BitDirty,
+		[BitmapActionStartsync]		= BitNone,
+		[BitmapActionEndsync]		= BitNone,
+		[BitmapActionAbortsync]		= BitNone,
+		[BitmapActionReload]		= BitNone,
+		[BitmapActionDaemon]		= BitNone,
+		[BitmapActionDiscard]		= BitUnwritten,
+		[BitmapActionStale]		= BitUnwritten,
+		[BitmapActionProactiveSync]	= BitNone,
+		[BitmapActionClearUnwritten]	= BitUnwritten,
 	},
 };
 
@@ -376,6 +442,7 @@ static void llbitmap_infect_dirty_bits(struct llbitmap *llbitmap,
 			pctl->state[pos] = level_456 ? BitNeedSync : BitDirty;
 			break;
 		case BitClean:
+		case BitCleanUnwritten:
 			pctl->state[pos] = BitDirty;
 			break;
 		}
@@ -383,7 +450,7 @@ static void llbitmap_infect_dirty_bits(struct llbitmap *llbitmap,
 }
 
 static void llbitmap_set_page_dirty(struct llbitmap *llbitmap, int idx,
-				    int offset)
+				    int offset, bool infect)
 {
 	struct llbitmap_page_ctl *pctl = llbitmap->pctl[idx];
 	unsigned int io_size = llbitmap->io_size;
@@ -398,7 +465,7 @@ static void llbitmap_set_page_dirty(struct llbitmap *llbitmap, int idx,
 	 * resync all the dirty bits, hence skip infect new dirty bits to
 	 * prevent resync unnecessary data.
 	 */
-	if (llbitmap->mddev->degraded) {
+	if (llbitmap->mddev->degraded || !infect) {
 		set_bit(block, pctl->dirty);
 		return;
 	}
@@ -438,7 +505,9 @@ static void llbitmap_write(struct llbitmap *llbitmap, enum llbitmap_state state,
 
 	llbitmap->pctl[idx]->state[bit] = state;
 	if (state == BitDirty || state == BitNeedSync)
-		llbitmap_set_page_dirty(llbitmap, idx, bit);
+		llbitmap_set_page_dirty(llbitmap, idx, bit, true);
+	else if (state == BitNeedSyncUnwritten)
+		llbitmap_set_page_dirty(llbitmap, idx, bit, false);
 }
 
 static struct page *llbitmap_read_page(struct llbitmap *llbitmap, int idx)
@@ -627,11 +696,10 @@ static enum llbitmap_state llbitmap_state_machine(struct llbitmap *llbitmap,
 			goto write_bitmap;
 		}
 
-		if (c == BitNeedSync)
+		if (c == BitNeedSync || c == BitNeedSyncUnwritten)
 			need_resync = !mddev->degraded;
 
 		state = state_machine[c][action];
-
 write_bitmap:
 		if (unlikely(mddev->degraded)) {
 			/* For degraded array, mark new data as need sync. */
@@ -658,8 +726,7 @@ static enum llbitmap_state llbitmap_state_machine(struct llbitmap *llbitmap,
 		}
 
 		llbitmap_write(llbitmap, state, start);
-
-		if (state == BitNeedSync)
+		if (state == BitNeedSync || state == BitNeedSyncUnwritten)
 			need_resync = !mddev->degraded;
 		else if (state == BitDirty &&
 			 !timer_pending(&llbitmap->pending_timer))
@@ -1229,7 +1296,7 @@ static bool llbitmap_blocks_synced(struct mddev *mddev, sector_t offset)
 	unsigned long p = offset >> llbitmap->chunkshift;
 	enum llbitmap_state c = llbitmap_read(llbitmap, p);
 
-	return c == BitClean || c == BitDirty;
+	return c == BitClean || c == BitDirty || c == BitCleanUnwritten;
 }
 
 static sector_t llbitmap_skip_sync_blocks(struct mddev *mddev, sector_t offset)
@@ -1243,6 +1310,10 @@ static sector_t llbitmap_skip_sync_blocks(struct mddev *mddev, sector_t offset)
 	if (c == BitUnwritten)
 		return blocks;
 
+	/* Skip CleanUnwritten - no user data, will be reset after recovery */
+	if (c == BitCleanUnwritten)
+		return blocks;
+
 	/* For degraded array, don't skip */
 	if (mddev->degraded)
 		return 0;
@@ -1261,14 +1332,25 @@ static bool llbitmap_start_sync(struct mddev *mddev, sector_t offset,
 {
 	struct llbitmap *llbitmap = mddev->bitmap;
 	unsigned long p = offset >> llbitmap->chunkshift;
+	enum llbitmap_state state;
+
+	/*
+	 * Before recovery starts, convert CleanUnwritten to Unwritten.
+	 * This ensures the new disk won't have stale parity data.
+	 */
+	if (offset == 0 && test_bit(MD_RECOVERY_RECOVER, &mddev->recovery) &&
+	    !test_bit(MD_RECOVERY_LAZY_RECOVER, &mddev->recovery))
+		llbitmap_state_machine(llbitmap, 0, llbitmap->chunks - 1,
+				       BitmapActionClearUnwritten);
+
 
 	/*
 	 * Handle one bit at a time, this is much simpler. And it doesn't matter
 	 * if md_do_sync() loop more times.
 	 */
 	*blocks = llbitmap->chunksize - (offset & (llbitmap->chunksize - 1));
-	return llbitmap_state_machine(llbitmap, p, p,
-				      BitmapActionStartsync) == BitSyncing;
+	state = llbitmap_state_machine(llbitmap, p, p, BitmapActionStartsync);
+	return state == BitSyncing || state == BitSyncingUnwritten;
 }
 
 /* Something is wrong, sync_thread stop at @offset */
@@ -1474,9 +1556,15 @@ static ssize_t bits_show(struct mddev *mddev, char *page)
 	}
 
 	mutex_unlock(&mddev->bitmap_info.mutex);
-	return sprintf(page, "unwritten %d\nclean %d\ndirty %d\nneed sync %d\nsyncing %d\n",
+	return sprintf(page,
+		       "unwritten %d\nclean %d\ndirty %d\n"
+		       "need sync %d\nsyncing %d\n"
+		       "need sync unwritten %d\nsyncing unwritten %d\n"
+		       "clean unwritten %d\n",
 		       bits[BitUnwritten], bits[BitClean], bits[BitDirty],
-		       bits[BitNeedSync], bits[BitSyncing]);
+		       bits[BitNeedSync], bits[BitSyncing],
+		       bits[BitNeedSyncUnwritten], bits[BitSyncingUnwritten],
+		       bits[BitCleanUnwritten]);
 }
 
 static struct md_sysfs_entry llbitmap_bits = __ATTR_RO(bits);
@@ -1549,11 +1637,39 @@ barrier_idle_store(struct mddev *mddev, const char *buf, size_t len)
 
 static struct md_sysfs_entry llbitmap_barrier_idle = __ATTR_RW(barrier_idle);
 
+static ssize_t
+proactive_sync_store(struct mddev *mddev, const char *buf, size_t len)
+{
+	struct llbitmap *llbitmap;
+
+	/* Only for RAID-456 */
+	if (!raid_is_456(mddev))
+		return -EINVAL;
+
+	mutex_lock(&mddev->bitmap_info.mutex);
+	llbitmap = mddev->bitmap;
+	if (!llbitmap || !llbitmap->pctl) {
+		mutex_unlock(&mddev->bitmap_info.mutex);
+		return -ENODEV;
+	}
+
+	/* Trigger proactive sync on all Unwritten regions */
+	llbitmap_state_machine(llbitmap, 0, llbitmap->chunks - 1,
+			       BitmapActionProactiveSync);
+
+	mutex_unlock(&mddev->bitmap_info.mutex);
+	return len;
+}
+
+static struct md_sysfs_entry llbitmap_proactive_sync =
+	__ATTR(proactive_sync, 0200, NULL, proactive_sync_store);
+
 static struct attribute *md_llbitmap_attrs[] = {
 	&llbitmap_bits.attr,
 	&llbitmap_metadata.attr,
 	&llbitmap_daemon_sleep.attr,
 	&llbitmap_barrier_idle.attr,
+	&llbitmap_proactive_sync.attr,
 	NULL
 };
 
diff --git a/drivers/md/md.c b/drivers/md/md.c
index d2607ed5c2e9..270802b8a4fc 100644
--- a/drivers/md/md.c
+++ b/drivers/md/md.c
@@ -9870,8 +9870,10 @@ void md_do_sync(struct md_thread *thread)
 				 * Give other IO more of a chance.
 				 * The faster the devices, the less we wait.
 				 */
-				wait_event(mddev->recovery_wait,
-					   !atomic_read(&mddev->recovery_active));
+				wait_event_timeout(
+					mddev->recovery_wait,
+					!atomic_read(&mddev->recovery_active),
+					HZ);
 			}
 		}
 	}
-- 
2.51.0


^ permalink raw reply related

* [PATCH 5/5] md/md-llbitmap: optimize initial sync with write_zeroes_unmap support
From: Yu Kuai @ 2026-02-14  6:10 UTC (permalink / raw)
  To: song; +Cc: linan122, xni, colyli, linux-raid, linux-kernel
In-Reply-To: <20260214061013.2335604-1-yukuai@fnnas.com>

For RAID-456 arrays with llbitmap, if all underlying disks support
write_zeroes with unmap, issue write_zeroes to zero all disk data
regions and initialize the bitmap to BitCleanUnwritten instead of
BitUnwritten.

This optimization skips the initial XOR parity building because:
1. write_zeroes with unmap guarantees zeroed reads after the operation
2. For RAID-456, when all data is zero, parity is automatically
   consistent (0 XOR 0 XOR ... = 0)
3. BitCleanUnwritten indicates parity is valid but no user data
   has been written

The implementation adds two helper functions:
- llbitmap_all_disks_support_wzeroes_unmap(): Checks if all active
  disks support write_zeroes with unmap
- llbitmap_zero_all_disks(): Issues blkdev_issue_zeroout() to each
  rdev's data region to zero all disks

The zeroing and bitmap state setting happens in llbitmap_init_state()
during bitmap initialization. If any disk fails to zero, we fall back
to BitUnwritten and normal lazy recovery.

This significantly reduces array initialization time for RAID-456
arrays built on modern NVMe SSDs or other devices that support
write_zeroes with unmap.

Signed-off-by: Yu Kuai <yukuai@fnnas.com>
---
 drivers/md/md-llbitmap.c | 62 +++++++++++++++++++++++++++++++++++++++-
 1 file changed, 61 insertions(+), 1 deletion(-)

diff --git a/drivers/md/md-llbitmap.c b/drivers/md/md-llbitmap.c
index 461050b2771b..48bc6a639edd 100644
--- a/drivers/md/md-llbitmap.c
+++ b/drivers/md/md-llbitmap.c
@@ -654,13 +654,73 @@ static int llbitmap_cache_pages(struct llbitmap *llbitmap)
 	return 0;
 }
 
+/*
+ * Check if all underlying disks support write_zeroes with unmap.
+ */
+static bool llbitmap_all_disks_support_wzeroes_unmap(struct llbitmap *llbitmap)
+{
+	struct mddev *mddev = llbitmap->mddev;
+	struct md_rdev *rdev;
+
+	rdev_for_each(rdev, mddev) {
+		if (rdev->raid_disk < 0 || test_bit(Faulty, &rdev->flags))
+			continue;
+
+		if (bdev_write_zeroes_unmap_sectors(rdev->bdev) == 0)
+			return false;
+	}
+
+	return true;
+}
+
+/*
+ * Issue write_zeroes to all underlying disks to zero their data regions.
+ * This ensures parity consistency for RAID-456 (0 XOR 0 = 0).
+ * Returns true if all disks were successfully zeroed.
+ */
+static bool llbitmap_zero_all_disks(struct llbitmap *llbitmap)
+{
+	struct mddev *mddev = llbitmap->mddev;
+	struct md_rdev *rdev;
+	sector_t dev_sectors = mddev->dev_sectors;
+	int ret;
+
+	rdev_for_each(rdev, mddev) {
+		if (rdev->raid_disk < 0 || test_bit(Faulty, &rdev->flags))
+			continue;
+
+		ret = blkdev_issue_zeroout(rdev->bdev,
+					   rdev->data_offset,
+					   dev_sectors,
+					   GFP_KERNEL, 0);
+		if (ret) {
+			pr_warn("md/llbitmap: failed to zero disk %pg: %d\n",
+				rdev->bdev, ret);
+			return false;
+		}
+	}
+
+	return true;
+}
+
 static void llbitmap_init_state(struct llbitmap *llbitmap)
 {
+	struct mddev *mddev = llbitmap->mddev;
 	enum llbitmap_state state = BitUnwritten;
 	unsigned long i;
 
-	if (test_and_clear_bit(BITMAP_CLEAN, &llbitmap->flags))
+	if (test_and_clear_bit(BITMAP_CLEAN, &llbitmap->flags)) {
 		state = BitClean;
+	} else if (raid_is_456(mddev) &&
+		   llbitmap_all_disks_support_wzeroes_unmap(llbitmap)) {
+		/*
+		 * All disks support write_zeroes with unmap. Zero all disks
+		 * to ensure parity consistency, then set BitCleanUnwritten
+		 * to skip initial sync.
+		 */
+		if (llbitmap_zero_all_disks(llbitmap))
+			state = BitCleanUnwritten;
+	}
 
 	for (i = 0; i < llbitmap->chunks; i++)
 		llbitmap_write(llbitmap, state, i);
-- 
2.51.0


^ permalink raw reply related

* Re: Cannot change RAID array check speed
From: David Niklas @ 2026-02-16  3:10 UTC (permalink / raw)
  To: linux-kernel, Linux RAID
In-Reply-To: <20260215220843.6b28e632@Core-Ultra-2-x20>

I forgot to mention that I also tested with Kernel version 6.14. It has
the same problem.

Thanks again,
David

On Sun, 15 Feb 2026 22:08:43 -0500
David Niklas <simd@vfemail.net> wrote:
> Hello,
> I upgraded my kernel from 6.9 to 6.17. They're both the "same" custom
> config. I tried:
> 
> # echo 200000 > /sys/devices/virtual/block/md7/md/sync_speed
> bash: /sys/devices/virtual/block/md7/md/sync_speed: Permission denied
> # id
> uid=0(root) gid=0(root) groups=0(root)
> 
> This used to work. Any ideas as to what I could have set wrong? I
> haven't a clue!
> 
> Is there a security debugging tool or a log-file I could use to trace
> this down?
> 
> Thanks,
> David


^ permalink raw reply

* Cannot change RAID array check speed
From: David Niklas @ 2026-02-16  3:08 UTC (permalink / raw)
  To: linux-kernel, Linux RAID

Hello,
I upgraded my kernel from 6.9 to 6.17. They're both the "same" custom
config. I tried:

# echo 200000 > /sys/devices/virtual/block/md7/md/sync_speed
bash: /sys/devices/virtual/block/md7/md/sync_speed: Permission denied
# id
uid=0(root) gid=0(root) groups=0(root)

This used to work. Any ideas as to what I could have set wrong? I haven't
a clue!

Is there a security debugging tool or a log-file I could use to trace this
down?

Thanks,
David

^ permalink raw reply

* Re: Cannot change RAID array check speed
From: simd @ 2026-02-16 15:14 UTC (permalink / raw)
  To: linux-kernel, Linux RAID
In-Reply-To: <CALtW_ahB7MUVgRuf=3itf-XbYW_nuEyj_p2-BGEk71mkrVg6FA@mail.gmail.com>

It worked with LK 6.9 on Devuan (Debian Stretch) ASCII.

I upgraded the system, both HW and SW, to Daedalus and I had to go with a
newer kernel to use the newer Intel iGPU. I'm now on Devuan (Debian
Bookworm) Daedalus.

I used to change this value all the time to get the check done faster or
to slow it down because I needed to access the array.

Maybe it's a distro security policy? IDK. But as I said, I used to change
it all the time. That's why there's a sync_speed_min/max, or so I thought.

Thanks,
David

PS: Accidentally sent to the user instead of to the list. Sorry.


On Mon, 16 Feb 2026 06:20:58 +0100
Dragan Milivojević <galileo@pkm-inc.com> wrote:
> When did that work?
> /sys/devices/virtual/ etc
> is readonly on 4.18,  6.12 and  6.18 (a few boxes that I have).
> 
> 
> 
> On Mon, 16 Feb 2026 at 04:15, David Niklas <simd@vfemail.net> wrote:
> >
> > Hello,
> > I upgraded my kernel from 6.9 to 6.17. They're both the "same" custom
> > config. I tried:
> >
> > # echo 200000 > /sys/devices/virtual/block/md7/md/sync_speed
> > bash: /sys/devices/virtual/block/md7/md/sync_speed: Permission denied
> > # id
> > uid=0(root) gid=0(root) groups=0(root)
> >
> > This used to work. Any ideas as to what I could have set wrong? I
> > haven't a clue!
> >
> > Is there a security debugging tool or a log-file I could use to trace
> > this down?
> >
> > Thanks,
> > David
> >  


^ permalink raw reply

* Re: Cannot change RAID array check speed
From: David Niklas @ 2026-02-16 20:01 UTC (permalink / raw)
  To: Linux RAID; +Cc: linux-kernel
In-Reply-To: <20260216101452.2f28a76b@Core-Ultra-2-x20>

Hey,
I think I figured this out. I need to change it via the
sync_speed_min/max. I can't used sync_speed anymore.

Thanks guys!


On Mon, 16 Feb 2026 10:14:52 -0500
simd@vfemail.net wrote:
> It worked with LK 6.9 on Devuan (Debian Stretch) ASCII.
> 
> I upgraded the system, both HW and SW, to Daedalus and I had to go with
> a newer kernel to use the newer Intel iGPU. I'm now on Devuan (Debian
> Bookworm) Daedalus.
> 
> I used to change this value all the time to get the check done faster or
> to slow it down because I needed to access the array.
> 
> Maybe it's a distro security policy? IDK. But as I said, I used to
> change it all the time. That's why there's a sync_speed_min/max, or so
> I thought.
> 
> Thanks,
> David
> 
> PS: Accidentally sent to the user instead of to the list. Sorry.
> 
> 
> On Mon, 16 Feb 2026 06:20:58 +0100
> Dragan Milivojević <galileo@pkm-inc.com> wrote:
> > When did that work?
> > /sys/devices/virtual/ etc
> > is readonly on 4.18,  6.12 and  6.18 (a few boxes that I have).
> > 
> > 
> > 
> > On Mon, 16 Feb 2026 at 04:15, David Niklas <simd@vfemail.net> wrote:  
> > >
> > > Hello,
> > > I upgraded my kernel from 6.9 to 6.17. They're both the "same"
> > > custom config. I tried:
> > >
> > > # echo 200000 > /sys/devices/virtual/block/md7/md/sync_speed
> > > bash: /sys/devices/virtual/block/md7/md/sync_speed: Permission
> > > denied # id
> > > uid=0(root) gid=0(root) groups=0(root)
> > >
> > > This used to work. Any ideas as to what I could have set wrong? I
> > > haven't a clue!
> > >
> > > Is there a security debugging tool or a log-file I could use to
> > > trace this down?
> > >
> > > Thanks,
> > > David
> > >    
> 
> 


^ permalink raw reply

* Re: [PATCH 3/5] md: add fallback to correct bitmap_ops on version mismatch
From: Su Yue @ 2026-02-17  8:54 UTC (permalink / raw)
  To: Yu Kuai; +Cc: song, linan122, xni, colyli, linux-raid, linux-kernel
In-Reply-To: <20260214061013.2335604-4-yukuai@fnnas.com>

On Sat 14 Feb 2026 at 14:10, Yu Kuai <yukuai@fnnas.com> wrote:

> If default bitmap version and on-disk version doesn't match, and 
> mdadm
> is not the latest version to set bitmap_type, set bitmap_ops 
> based on
> the disk version.
>
Why not just let old version mdadm fails  since llbitmap is a new 
feature.

> Signed-off-by: Yu Kuai <yukuai@fnnas.com>
> ---
>  drivers/md/md.c | 103 
>  +++++++++++++++++++++++++++++++++++++++++++++++-
>  1 file changed, 102 insertions(+), 1 deletion(-)
>
> diff --git a/drivers/md/md.c b/drivers/md/md.c
> index 59cd303548de..d2607ed5c2e9 100644
> --- a/drivers/md/md.c
> +++ b/drivers/md/md.c
> @@ -6447,15 +6447,116 @@ static void md_safemode_timeout(struct 
> timer_list *t)
>
>  static int start_dirty_degraded;
>
> +/*
> + * Read bitmap superblock and return the bitmap_id based on 
> disk version.
> + * This is used as fallback when default bitmap version and 
> on-disk version
> + * doesn't match, and mdadm is not the latest version to set 
> bitmap_type.
> + */
> +static enum md_submodule_id md_bitmap_get_id_from_sb(struct 
> mddev *mddev)
> +{
> +	struct md_rdev *rdev;
> +	struct page *sb_page;
> +	bitmap_super_t *sb;
> +	enum md_submodule_id id = ID_BITMAP_NONE;
> +	sector_t sector;
> +	u32 version;
> +
> +	if (!mddev->bitmap_info.offset)
> +		return ID_BITMAP_NONE;
> +
> +	sb_page = alloc_page(GFP_KERNEL);
> +	if (!sb_page)
> +		return ID_BITMAP_NONE;
> +
>
Personally I don't like the way treating error as ID_BITMAP_NONE.
When wrong things happen everything looks fine, no error code, no 
error message.

> +	sector = mddev->bitmap_info.offset;
> +
> +	rdev_for_each(rdev, mddev) {
> +		u32 iosize;
> +
> +		if (!test_bit(In_sync, &rdev->flags) ||
> +		    test_bit(Faulty, &rdev->flags) ||
> +		    test_bit(Bitmap_sync, &rdev->flags))
> +			continue;
> +
> +		iosize = roundup(sizeof(bitmap_super_t),
> +				 bdev_logical_block_size(rdev->bdev));
> +		if (sync_page_io(rdev, sector, iosize, sb_page, 
> REQ_OP_READ,
> +				 true))
> +			goto read_ok;
> +	}
>
And here.

> +	goto out;
> +
> +read_ok:
> +	sb = kmap_local_page(sb_page);
> +	if (sb->magic != cpu_to_le32(BITMAP_MAGIC))
> +		goto out_unmap;
> +
> +	version = le32_to_cpu(sb->version);
> +	switch (version) {
> +	case BITMAP_MAJOR_LO:
> +	case BITMAP_MAJOR_HI:
> +	case BITMAP_MAJOR_CLUSTERED:
>
For BITMAP_MAJOR_CLUSTERED, why not ID_CLUSTER ?

--
Su
> +		id = ID_BITMAP;
> +		break;
> +	case BITMAP_MAJOR_LOCKLESS:
> +		id = ID_LLBITMAP;
> +		break;
> +	default:
> +		pr_warn("md: %s: unknown bitmap version %u\n",
> +			mdname(mddev), version);
> +		break;
> +	}
> +
> +out_unmap:
> +	kunmap_local(sb);
> +out:
> +	__free_page(sb_page);
> +	return id;
> +}
> +
>  static int md_bitmap_create(struct mddev *mddev)
>  {
> +	enum md_submodule_id orig_id = mddev->bitmap_id;
> +	enum md_submodule_id sb_id;
> +	int err;
> +
>  	if (mddev->bitmap_id == ID_BITMAP_NONE)
>  		return -EINVAL;
>
>  	if (!mddev_set_bitmap_ops(mddev))
>  		return -ENOENT;
>
> -	return mddev->bitmap_ops->create(mddev);
> +	err = mddev->bitmap_ops->create(mddev);
> +	if (!err)
> +		return 0;
>
> +
> +	/*
> +	 * Create failed, if default bitmap version and on-disk 
> version
> +	 * doesn't match, and mdadm is not the latest version to set
> +	 * bitmap_type, set bitmap_ops based on the disk version.
> +	 */
> +	mddev_clear_bitmap_ops(mddev);
> +
> +	sb_id = md_bitmap_get_id_from_sb(mddev);
> +	if (sb_id == ID_BITMAP_NONE || sb_id == orig_id)
> +		return err;
> +
> +	pr_info("md: %s: bitmap version mismatch, switching from %d to 
> %d\n",
> +		mdname(mddev), orig_id, sb_id);
> +
> +	mddev->bitmap_id = sb_id;
> +	if (!mddev_set_bitmap_ops(mddev)) {
> +		mddev->bitmap_id = orig_id;
> +		return -ENOENT;
> +	}
> +
> +	err = mddev->bitmap_ops->create(mddev);
> +	if (err) {
> +		mddev_clear_bitmap_ops(mddev);
> +		mddev->bitmap_id = orig_id;
> +	}
> +
> +	return err;
>  }
>
>  static void md_bitmap_destroy(struct mddev *mddev)

^ permalink raw reply

* Re: [f2fs-dev] [PATCH V3 0/6] block: ignore __blkdev_issue_discard() ret value
From: patchwork-bot+f2fs @ 2026-02-17 21:14 UTC (permalink / raw)
  To: Chaitanya Kulkarni
  Cc: axboe, agk, snitzer, mpatocka, song, yukuai, hch, sagi, kch,
	jaegeuk, chao, cem, dm-devel, linux-raid, linux-kernel,
	linux-nvme, linux-f2fs-devel, linux-block, bpf, linux-xfs
In-Reply-To: <20251124234806.75216-1-ckulkarnilinux@gmail.com>

Hello:

This series was applied to jaegeuk/f2fs.git (dev)
by Carlos Maiolino <cem@kernel.org>:

On Mon, 24 Nov 2025 15:48:00 -0800 you wrote:
> Hi,
> 
> __blkdev_issue_discard() only returns value 0, that makes post call
> error checking code dead. This patch series revmoes this dead code at
> all the call sites and adjust the callers.
> 
> Please note that it doesn't change the return type of the function from
> int to void in this series, it will be done once this series gets merged
> smoothly.
> 
> [...]

Here is the summary with links:
  - [f2fs-dev,V3,1/6] block: ignore discard return value
    (no matching commit)
  - [f2fs-dev,V3,2/6] md: ignore discard return value
    https://git.kernel.org/jaegeuk/f2fs/c/699fcfb6cb80
  - [f2fs-dev,V3,3/6] dm: ignore discard return value
    (no matching commit)
  - [f2fs-dev,V3,4/6] nvmet: ignore discard return value
    https://git.kernel.org/jaegeuk/f2fs/c/38d12f15c477
  - [f2fs-dev,V3,5/6] f2fs: ignore discard return value
    (no matching commit)
  - [f2fs-dev,V3,6/6] xfs: ignore discard return value
    https://git.kernel.org/jaegeuk/f2fs/c/2145f447b79a

You are awesome, thank you!
-- 
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html



^ permalink raw reply


This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox