Linux-Next discussions
 help / color / mirror / Atom feed
* next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
@ 2026-09-20 12:58 Tao Cui
  2026-09-20 13:12 ` Stephen Rothwell
  0 siblings, 1 reply; 9+ messages in thread
From: Tao Cui @ 2026-09-20 12:58 UTC (permalink / raw)
  To: Stephen Rothwell
  Cc: Tao Cui, linux-next, Jens Axboe, Damien Le Moal, linux-block,
	cuitao

Hi Stephen,

I ran into this while testing a blk-iocost series of mine (charging
zone appends as writes) against a zoned null_blk: the device simply
does not show up, whether configured through module parameters or
configfs.  null_blk prints "using native zone append" and then
add_disk() fails silently with -ENODEV.

My series only touches blk-iocost.c, and the failure reproduces on
plain linux-next, so it is not mine.  Bisecting over the tags (qemu,
"null_blk.zoned=1 null_blk.gb=1 null_blk.zone_size=64" on the command
line) narrows it down to one day:

  next-20260916: nullb0 created, /sys/block/nullb0/queue/zoned = host-managed
  next-20260917: no device, add_disk() fails
  next-20260918: same failure

The only blk-zoned.c changes between next-20260916 and next-20260917
come in through the merge of the block tree, and next-20260918 repeats
the same resolution in the same place:

  66c8b6b56c09 ("Merge branch 'for-next' of .../axboe/linux.git", 2026-09-17)
  a9ce2250a714 ("Merge branch 'for-next' of .../axboe/linux.git", 2026-09-18)

As far as I can tell the resolved disk_revalidate_capacity() matches
neither parent: it gains a

  if (args->capacity >= args->nr_zones)

check that compares a sector count against a zone count, with
args->nr_zones still 0 on the first call, so it always trips and
blk_revalidate_disk_zones() always returns -ENODEV.  The zone_sectors
power-of-two check that both parents have is gone as well.  This hits
every zoned device, not just null_blk.

Taking the axboe side of the function makes nullb0 come back on
next-20260918 here (host-managed, chunk_sectors 131072), so it looks
like the merge should just have taken that side.

Same conflict, same bad outcome twice in a row now, so tomorrow's
merge may hit it again unless the block tree has moved on in this
area.

Regards,
Tao

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
  2026-09-20 12:58 next-20260917 and next-20260918: merge of the block tree breaks all zoned devices Tao Cui
@ 2026-09-20 13:12 ` Stephen Rothwell
  2026-09-20 15:18   ` Mark Brown
  0 siblings, 1 reply; 9+ messages in thread
From: Stephen Rothwell @ 2026-09-20 13:12 UTC (permalink / raw)
  To: Tao Cui
  Cc: linux-next, Jens Axboe, Damien Le Moal, linux-block, cuitao,
	Mark Brown

[-- Attachment #1: Type: text/plain, Size: 2188 bytes --]

Hi Tao,

On Sun, 20 Sep 2026 20:58:34 +0800 Tao Cui <cui.tao@linux.dev> wrote:
>
> Hi Stephen,

I no longer run linux-next, please email Mark Brown (cc'd).

[The rest kept for Mark's benefit]

> I ran into this while testing a blk-iocost series of mine (charging
> zone appends as writes) against a zoned null_blk: the device simply
> does not show up, whether configured through module parameters or
> configfs.  null_blk prints "using native zone append" and then
> add_disk() fails silently with -ENODEV.
> 
> My series only touches blk-iocost.c, and the failure reproduces on
> plain linux-next, so it is not mine.  Bisecting over the tags (qemu,
> "null_blk.zoned=1 null_blk.gb=1 null_blk.zone_size=64" on the command
> line) narrows it down to one day:
> 
>   next-20260916: nullb0 created, /sys/block/nullb0/queue/zoned = host-managed
>   next-20260917: no device, add_disk() fails
>   next-20260918: same failure
> 
> The only blk-zoned.c changes between next-20260916 and next-20260917
> come in through the merge of the block tree, and next-20260918 repeats
> the same resolution in the same place:
> 
>   66c8b6b56c09 ("Merge branch 'for-next' of .../axboe/linux.git", 2026-09-17)
>   a9ce2250a714 ("Merge branch 'for-next' of .../axboe/linux.git", 2026-09-18)
> 
> As far as I can tell the resolved disk_revalidate_capacity() matches
> neither parent: it gains a
> 
>   if (args->capacity >= args->nr_zones)
> 
> check that compares a sector count against a zone count, with
> args->nr_zones still 0 on the first call, so it always trips and
> blk_revalidate_disk_zones() always returns -ENODEV.  The zone_sectors
> power-of-two check that both parents have is gone as well.  This hits
> every zoned device, not just null_blk.
> 
> Taking the axboe side of the function makes nullb0 come back on
> next-20260918 here (host-managed, chunk_sectors 131072), so it looks
> like the merge should just have taken that side.
> 
> Same conflict, same bad outcome twice in a row now, so tomorrow's
> merge may hit it again unless the block tree has moved on in this
> area.
> 
> Regards,
> Tao
> 

-- 
Cheers,
Stephen Rothwell

[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 488 bytes --]

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
  2026-09-20 13:12 ` Stephen Rothwell
@ 2026-09-20 15:18   ` Mark Brown
  2026-09-20 22:37     ` Damien Le Moal
  0 siblings, 1 reply; 9+ messages in thread
From: Mark Brown @ 2026-09-20 15:18 UTC (permalink / raw)
  To: Stephen Rothwell
  Cc: Tao Cui, linux-next, Jens Axboe, Damien Le Moal, linux-block,
	cuitao

[-- Attachment #1: Type: text/plain, Size: 1215 bytes --]

On Sun, Sep 20, 2026 at 11:12:26PM +1000, Stephen Rothwell wrote:
> On Sun, 20 Sep 2026 20:58:34 +0800 Tao Cui <cui.tao@linux.dev> wrote:

> I no longer run linux-next, please email Mark Brown (cc'd).

> [The rest kept for Mark's benefit]

Thanks.

> > I ran into this while testing a blk-iocost series of mine (charging
> > zone appends as writes) against a zoned null_blk: the device simply
> > does not show up, whether configured through module parameters or
> > configfs.  null_blk prints "using native zone append" and then
> > add_disk() fails silently with -ENODEV.

> > Taking the axboe side of the function makes nullb0 come back on
> > next-20260918 here (host-managed, chunk_sectors 131072), so it looks
> > like the merge should just have taken that side.

> > Same conflict, same bad outcome twice in a row now, so tomorrow's
> > merge may hit it again unless the block tree has moved on in this
> > area.

I did say when I did that merge that I had no confidence in it but
nobody responded.  Can you please send me a commit on top of current
-next which fixes up the resolution appropriately?  I can drop extra
fixups in relatively easily, it's probably safer to take something that
you have tested.

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 488 bytes --]

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
  2026-09-20 15:18   ` Mark Brown
@ 2026-09-20 22:37     ` Damien Le Moal
  2026-09-21  0:49       ` Tao Cui
  2026-09-21  8:24       ` Mark Brown
  0 siblings, 2 replies; 9+ messages in thread
From: Damien Le Moal @ 2026-09-20 22:37 UTC (permalink / raw)
  To: Mark Brown; +Cc: Tao Cui, linux-next, Jens Axboe, linux-block, cuitao

[-- Attachment #1: Type: text/plain, Size: 630 bytes --]

On 9/21/26 00:18, Mark Brown wrote:
> I did say when I did that merge that I had no confidence in it but
> nobody responded.  Can you please send me a commit on top of current
> -next which fixes up the resolution appropriately?  I can drop extra
> fixups in relatively easily, it's probably safer to take something that
> you have tested.

Mark,

My apologies for not replying, but I was traveling. I had looked at your
resolution but only quickly and I did not notice the incorrect comparison
causing the issue.

I am attaching a diff that cleans up the conflict resolution.

Thanks!

-- 
Damien Le Moal
Western Digital Research

[-- Attachment #2: blk-zoned.diff --]
[-- Type: text/x-patch, Size: 1258 bytes --]

diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index bae48f0d3464..dd8a72e352c1 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -2031,11 +2031,16 @@ void disk_init_zone_resources(struct gendisk *disk)
 static unsigned int disk_get_nr_zones(struct gendisk *disk, sector_t capacity)
 {
 	struct queue_limits *lim = &disk->queue->limits;
+	unsigned long long nr_zones;
 
 	if (!capacity || !lim->chunk_sectors)
 		return 0;
 
-	return DIV_ROUND_UP_ULL(capacity, lim->chunk_sectors);
+	nr_zones = DIV_ROUND_UP_ULL(capacity, lim->chunk_sectors);
+	if (nr_zones > UINT_MAX)
+		return 0;
+
+	return nr_zones;
 }
 
 /*
@@ -2197,7 +2202,7 @@ static int disk_init_revalidate_args(struct gendisk *disk,
 {
 	args->disk = disk;
 	args->nr_zones = disk_get_nr_zones(disk, args->capacity);
-	if (args->nr_zones > UINT_MAX)
+	if (!args->nr_zones)
 		return -EINVAL;
 
 	/* Cached zone conditions: 1 byte per zone */
@@ -2329,8 +2334,6 @@ static int disk_revalidate_capacity(struct gendisk *disk,
 	nr_zones = disk_get_nr_zones(disk, args->capacity);
 	if (!args->capacity || !nr_zones)
 		goto drop_all_zwplugs;
-	if (args->capacity >= args->nr_zones)
-		goto drop_all_zwplugs;
 
 	/*
 	 * Check if the capacity has changed. If it did, assume that the device

^ permalink raw reply related	[flat|nested] 9+ messages in thread

* Re: next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
  2026-09-20 22:37     ` Damien Le Moal
@ 2026-09-21  0:49       ` Tao Cui
  2026-09-21  2:07         ` Damien Le Moal
  2026-09-21  2:31         ` Damien Le Moal
  2026-09-21  8:24       ` Mark Brown
  1 sibling, 2 replies; 9+ messages in thread
From: Tao Cui @ 2026-09-21  0:49 UTC (permalink / raw)
  To: Damien Le Moal, Mark Brown
  Cc: cui.tao, linux-next, Jens Axboe, linux-block, cuitao

Hi Mark, Damien,

在 2026/9/21 06:37, Damien Le Moal 写道:
> On 9/21/26 00:18, Mark Brown wrote:
>> I did say when I did that merge that I had no confidence in it but
>> nobody responded.  Can you please send me a commit on top of current
>> -next which fixes up the resolution appropriately?  I can drop extra
>> fixups in relatively easily, it's probably safer to take something that
>> you have tested.
> 
> Mark,
> 
> My apologies for not replying, but I was traveling. I had looked at your
> resolution but only quickly and I did not notice the incorrect comparison
> causing the issue.
> 
> I am attaching a diff that cleans up the conflict resolution.
> 

I tested Damien's patch on top of next-20260918.

It restores the zoned null_blk here: nullb0 is created again, reports
host-managed, and the queue chunk size is back to 131072 sectors
(CONFIG_BLK_DEV_NULL_BLK=y, "null_blk.zoned=1 null_blk.gb=1
null_blk.zone_size=64" on the kernel command line).  So this looks
good to me.

Tested-by: Tao Cui <cuitao@kylinos.cn>

One question for Damien while looking through the merged code: both
parents had the power-of-two check for zone_sectors, but that check
also disappears in the merged result.  Was that removal intentional,
or should it be restored as well?

Thanks.

> Thanks!
> 


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
  2026-09-21  0:49       ` Tao Cui
@ 2026-09-21  2:07         ` Damien Le Moal
  2026-09-21  2:31         ` Damien Le Moal
  1 sibling, 0 replies; 9+ messages in thread
From: Damien Le Moal @ 2026-09-21  2:07 UTC (permalink / raw)
  To: Tao Cui, Mark Brown; +Cc: linux-next, Jens Axboe, linux-block, cuitao

On 9/21/26 09:49, Tao Cui wrote:
> Hi Mark, Damien,
> 
> 在 2026/9/21 06:37, Damien Le Moal 写道:
>> On 9/21/26 00:18, Mark Brown wrote:
>>> I did say when I did that merge that I had no confidence in it but
>>> nobody responded.  Can you please send me a commit on top of current
>>> -next which fixes up the resolution appropriately?  I can drop extra
>>> fixups in relatively easily, it's probably safer to take something that
>>> you have tested.
>>
>> Mark,
>>
>> My apologies for not replying, but I was traveling. I had looked at your
>> resolution but only quickly and I did not notice the incorrect comparison
>> causing the issue.
>>
>> I am attaching a diff that cleans up the conflict resolution.
>>
> 
> I tested Damien's patch on top of next-20260918.
> 
> It restores the zoned null_blk here: nullb0 is created again, reports
> host-managed, and the queue chunk size is back to 131072 sectors
> (CONFIG_BLK_DEV_NULL_BLK=y, "null_blk.zoned=1 null_blk.gb=1
> null_blk.zone_size=64" on the kernel command line).  So this looks
> good to me.
> 
> Tested-by: Tao Cui <cuitao@kylinos.cn>
> 
> One question for Damien while looking through the merged code: both
> parents had the power-of-two check for zone_sectors, but that check
> also disappears in the merged result.  Was that removal intentional,
> or should it be restored as well?

Doh! Yes, you are absolutely correct.
I am going to do a manual merge to generate a proper conflict resolution.

> 
> Thanks.
> 
>> Thanks!
>>
> 


-- 
Damien Le Moal
Western Digital Research

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
  2026-09-21  0:49       ` Tao Cui
  2026-09-21  2:07         ` Damien Le Moal
@ 2026-09-21  2:31         ` Damien Le Moal
  2026-09-21 10:29           ` Mark Brown
  1 sibling, 1 reply; 9+ messages in thread
From: Damien Le Moal @ 2026-09-21  2:31 UTC (permalink / raw)
  To: Tao Cui, Mark Brown; +Cc: linux-next, Jens Axboe, linux-block, cuitao

[-- Attachment #1: Type: text/plain, Size: 1512 bytes --]

On 9/21/26 09:49, Tao Cui wrote:
> Hi Mark, Damien,
> 
> 在 2026/9/21 06:37, Damien Le Moal 写道:
>> On 9/21/26 00:18, Mark Brown wrote:
>>> I did say when I did that merge that I had no confidence in it but
>>> nobody responded.  Can you please send me a commit on top of current
>>> -next which fixes up the resolution appropriately?  I can drop extra
>>> fixups in relatively easily, it's probably safer to take something that
>>> you have tested.
>>
>> Mark,
>>
>> My apologies for not replying, but I was traveling. I had looked at your
>> resolution but only quickly and I did not notice the incorrect comparison
>> causing the issue.
>>
>> I am attaching a diff that cleans up the conflict resolution.
>>
> 
> I tested Damien's patch on top of next-20260918.
> 
> It restores the zoned null_blk here: nullb0 is created again, reports
> host-managed, and the queue chunk size is back to 131072 sectors
> (CONFIG_BLK_DEV_NULL_BLK=y, "null_blk.zoned=1 null_blk.gb=1
> null_blk.zone_size=64" on the kernel command line).  So this looks
> good to me.
> 
> Tested-by: Tao Cui <cuitao@kylinos.cn>
> 
> One question for Damien while looking through the merged code: both
> parents had the power-of-two check for zone_sectors, but that check
> also disappears in the merged result.  Was that removal intentional,
> or should it be restored as well?

Mark, Attaching a v2 diff for corrected the conflict resolution. This is on top
of the current linux-next.

Thanks.


-- 
Damien Le Moal
Western Digital Research

[-- Attachment #2: v2-blk-zoned.diff --]
[-- Type: text/x-patch, Size: 1907 bytes --]

diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index bae48f0d3464..19268afb8752 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -2031,11 +2031,16 @@ void disk_init_zone_resources(struct gendisk *disk)
 static unsigned int disk_get_nr_zones(struct gendisk *disk, sector_t capacity)
 {
 	struct queue_limits *lim = &disk->queue->limits;
+	unsigned long long nr_zones;
 
 	if (!capacity || !lim->chunk_sectors)
 		return 0;
 
-	return DIV_ROUND_UP_ULL(capacity, lim->chunk_sectors);
+	nr_zones = DIV_ROUND_UP_ULL(capacity, lim->chunk_sectors);
+	if (nr_zones > UINT_MAX)
+		return 0;
+
+	return nr_zones;
 }
 
 /*
@@ -2197,8 +2202,8 @@ static int disk_init_revalidate_args(struct gendisk *disk,
 {
 	args->disk = disk;
 	args->nr_zones = disk_get_nr_zones(disk, args->capacity);
-	if (args->nr_zones > UINT_MAX)
-		return -EINVAL;
+	if (!args->nr_zones)
+		return -ENODEV;
 
 	/* Cached zone conditions: 1 byte per zone */
 	args->zones_state = kzalloc(args->nr_zones, GFP_NOIO);
@@ -2322,15 +2327,22 @@ static void disk_drop_zone_wplug(struct blk_zone_wplug *zwplug, void *data)
 static int disk_revalidate_capacity(struct gendisk *disk,
 				    struct blk_revalidate_zone_args *args)
 {
+	struct queue_limits *lim = &disk->queue->limits;
+        sector_t zone_sectors = lim->chunk_sectors;
 	unsigned int nr_zones;
 	int ret = -ENODEV;
 
+	/* Checks that the device driver indicated a valid zone size. */
+	if (!zone_sectors || !is_power_of_2(zone_sectors)) {
+		pr_warn("%s: Invalid non power of two zone size (%llu)\n",
+			disk->disk_name, zone_sectors);
+		goto drop_all_zwplugs;
+	}
+
 	args->capacity = get_capacity(disk);
 	nr_zones = disk_get_nr_zones(disk, args->capacity);
 	if (!args->capacity || !nr_zones)
 		goto drop_all_zwplugs;
-	if (args->capacity >= args->nr_zones)
-		goto drop_all_zwplugs;
 
 	/*
 	 * Check if the capacity has changed. If it did, assume that the device

^ permalink raw reply related	[flat|nested] 9+ messages in thread

* Re: next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
  2026-09-20 22:37     ` Damien Le Moal
  2026-09-21  0:49       ` Tao Cui
@ 2026-09-21  8:24       ` Mark Brown
  1 sibling, 0 replies; 9+ messages in thread
From: Mark Brown @ 2026-09-21  8:24 UTC (permalink / raw)
  To: Damien Le Moal; +Cc: Tao Cui, linux-next, Jens Axboe, linux-block, cuitao

[-- Attachment #1: Type: text/plain, Size: 805 bytes --]

On Mon, Sep 21, 2026 at 07:37:58AM +0900, Damien Le Moal wrote:
> On 9/21/26 00:18, Mark Brown wrote:

> > I did say when I did that merge that I had no confidence in it but
> > nobody responded.  Can you please send me a commit on top of current
> > -next which fixes up the resolution appropriately?  I can drop extra
> > fixups in relatively easily, it's probably safer to take something that
> > you have tested.

> My apologies for not replying, but I was traveling. I had looked at your
> resolution but only quickly and I did not notice the incorrect comparison
> causing the issue.

No worries, stuff happens - I was just surprised nobody complained about
the merge TBH.

> I am attaching a diff that cleans up the conflict resolution.

Thanks, that should appear in today's -next all being well.

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 488 bytes --]

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: next-20260917 and next-20260918: merge of the block tree breaks all zoned devices
  2026-09-21  2:31         ` Damien Le Moal
@ 2026-09-21 10:29           ` Mark Brown
  0 siblings, 0 replies; 9+ messages in thread
From: Mark Brown @ 2026-09-21 10:29 UTC (permalink / raw)
  To: Damien Le Moal; +Cc: Tao Cui, linux-next, Jens Axboe, linux-block, cuitao

[-- Attachment #1: Type: text/plain, Size: 216 bytes --]

On Mon, Sep 21, 2026 at 11:31:11AM +0900, Damien Le Moal wrote:

> Mark, Attaching a v2 diff for corrected the conflict resolution. This is on top
> of the current linux-next.

I've replaced the patch with this one.

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 488 bytes --]

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-21 10:29 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-20 12:58 next-20260917 and next-20260918: merge of the block tree breaks all zoned devices Tao Cui
2026-09-20 13:12 ` Stephen Rothwell
2026-09-20 15:18   ` Mark Brown
2026-09-20 22:37     ` Damien Le Moal
2026-09-21  0:49       ` Tao Cui
2026-09-21  2:07         ` Damien Le Moal
2026-09-21  2:31         ` Damien Le Moal
2026-09-21 10:29           ` Mark Brown
2026-09-21  8:24       ` Mark Brown

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox