Linux-NVME Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known
@ 2026-09-14 10:57 Guixin Liu
  2026-09-15  6:36 ` Christoph Hellwig
  2026-09-22 15:13 ` Keith Busch
  0 siblings, 2 replies; 5+ messages in thread
From: Guixin Liu @ 2026-09-14 10:57 UTC (permalink / raw)
  To: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	Nilay Shroff, Daniel Wagner, John Garry, Hannes Reinecke
  Cc: linux-nvme

The namespace head is marked zoned at allocation time based only on
the command set identifier, before any zone information has been
queried.  If the zone info query fails on the first scan, the path
namespace is registered without zoned limits while the head still
advertises the zoned capability with a zone size of zero.  Reporting
zones or writing to the head then shifts by ilog2(0), triggering the
UBSAN shift-out-of-bounds report in the report-zones and write paths.

Drop the zoned feature from the head allocation and inherit it from
the path namespace: the head limits refresh already stacks the zoned
feature, the zone size and the zone resource limits from the path
queue, so the head matches the path namespace and becomes zoned once a
revalidation succeeds.  This also stops marking the head zoned when
CONFIG_BLK_DEV_ZONED is off, which used to fail the head allocation
with a WARN.

Found by code inspection while reviewing the nvme-7.3 branch.

Tested with a null_blk zoned namespace exported over two nvmet-tcp
ports, with the target patched to fail the command set specific
identify: the head no longer comes up zoned with zone size 0, the
UBSAN report is gone, and an ns-rescan once the identify succeeds again
transitions the head to zoned with the correct zone size.

Fixes: 28982ad73d6a ("nvme: set BLK_FEAT_ZONED for ZNS multipath disks")
Fixes: 3838e80fcfb3 ("nvme: skip the zoned limits update if the zone info query failed")
Cc: stable@vger.kernel.org
Signed-off-by: Guixin Liu <kanie@linux.alibaba.com>
---
 drivers/nvme/host/multipath.c | 2 --
 1 file changed, 2 deletions(-)

diff --git a/drivers/nvme/host/multipath.c b/drivers/nvme/host/multipath.c
index 75dbb58286a3..cdfaa04c25f8 100644
--- a/drivers/nvme/host/multipath.c
+++ b/drivers/nvme/host/multipath.c
@@ -763,8 +763,6 @@ int nvme_mpath_alloc_disk(struct nvme_ctrl *ctrl, struct nvme_ns_head *head)
 	lim.dma_alignment = 3;
 	lim.features |= BLK_FEAT_IO_STAT | BLK_FEAT_NOWAIT |
 		BLK_FEAT_POLL | BLK_FEAT_ATOMIC_WRITES | BLK_FEAT_PCI_P2PDMA;
-	if (head->ids.csi == NVME_CSI_ZNS)
-		lim.features |= BLK_FEAT_ZONED;
 
 	head->disk = blk_alloc_disk(&lim, ctrl->numa_node);
 	if (IS_ERR(head->disk))
-- 
2.43.7



^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH] nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known
  2026-09-14 10:57 [PATCH] nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known Guixin Liu
@ 2026-09-15  6:36 ` Christoph Hellwig
  2026-09-15  8:51   ` Guixin Liu
  2026-09-22 15:13 ` Keith Busch
  1 sibling, 1 reply; 5+ messages in thread
From: Christoph Hellwig @ 2026-09-15  6:36 UTC (permalink / raw)
  To: Guixin Liu
  Cc: Keith Busch, Jens Axboe, Christoph Hellwig, Sagi Grimberg,
	Nilay Shroff, Daniel Wagner, John Garry, Hannes Reinecke,
	linux-nvme

On Mon, Sep 14, 2026 at 06:57:13PM +0800, Guixin Liu wrote:
> The namespace head is marked zoned at allocation time based only on
> the command set identifier, before any zone information has been
> queried.  If the zone info query fails on the first scan, the path
> namespace is registered without zoned limits while the head still
> advertises the zoned capability with a zone size of zero.  Reporting
> zones or writing to the head then shifts by ilog2(0), triggering the
> UBSAN shift-out-of-bounds report in the report-zones and write paths.
> 
> Drop the zoned feature from the head allocation and inherit it from
> the path namespace: the head limits refresh already stacks the zoned
> feature, the zone size and the zone resource limits from the path
> queue, so the head matches the path namespace and becomes zoned once a
> revalidation succeeds.  This also stops marking the head zoned when
> CONFIG_BLK_DEV_ZONED is off, which used to fail the head allocation
> with a WARN.

Where do we actually stack it right now?



^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known
  2026-09-15  6:36 ` Christoph Hellwig
@ 2026-09-15  8:51   ` Guixin Liu
  2026-09-22  8:17     ` Guixin Liu
  0 siblings, 1 reply; 5+ messages in thread
From: Guixin Liu @ 2026-09-15  8:51 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Keith Busch, Jens Axboe, Sagi Grimberg, Nilay Shroff,
	Daniel Wagner, John Garry, Hannes Reinecke, linux-nvme



在 2026/9/15 14:36, Christoph Hellwig 写道:
> On Mon, Sep 14, 2026 at 06:57:13PM +0800, Guixin Liu wrote:
>> The namespace head is marked zoned at allocation time based only on
>> the command set identifier, before any zone information has been
>> queried.  If the zone info query fails on the first scan, the path
>> namespace is registered without zoned limits while the head still
>> advertises the zoned capability with a zone size of zero.  Reporting
>> zones or writing to the head then shifts by ilog2(0), triggering the
>> UBSAN shift-out-of-bounds report in the report-zones and write paths.
>>
>> Drop the zoned feature from the head allocation and inherit it from
>> the path namespace: the head limits refresh already stacks the zoned
>> feature, the zone size and the zone resource limits from the path
>> queue, so the head matches the path namespace and becomes zoned once a
>> revalidation succeeds.  This also stops marking the head zoned when
>> CONFIG_BLK_DEV_ZONED is off, which used to fail the head allocation
>> with a WARN.
> Where do we actually stack it right now?
The head stacks it in nvme_update_ns_info(): after the path namespace
update it's own limits, queue_limits_stack_bdev() stacks the path
queue into the head, and blk_stack_limits() inherits the BLK_FEAT_ZONED
through BLK_FEAT_INHERIT_MASK, along with chunk_sectors and the zone
resource limits. That happens before the head is registered, so
blk_revalidate_disk_zones() see the zoned feature in place.

Best Regards,
Guixin Liu



^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known
  2026-09-15  8:51   ` Guixin Liu
@ 2026-09-22  8:17     ` Guixin Liu
  0 siblings, 0 replies; 5+ messages in thread
From: Guixin Liu @ 2026-09-22  8:17 UTC (permalink / raw)
  To: Christoph Hellwig
  Cc: Keith Busch, Jens Axboe, Sagi Grimberg, Nilay Shroff,
	Daniel Wagner, John Garry, Hannes Reinecke, linux-nvme



在 2026/9/15 16:51, Guixin Liu 写道:
>
>
> 在 2026/9/15 14:36, Christoph Hellwig 写道:
>> On Mon, Sep 14, 2026 at 06:57:13PM +0800, Guixin Liu wrote:
>>> The namespace head is marked zoned at allocation time based only on
>>> the command set identifier, before any zone information has been
>>> queried.  If the zone info query fails on the first scan, the path
>>> namespace is registered without zoned limits while the head still
>>> advertises the zoned capability with a zone size of zero.  Reporting
>>> zones or writing to the head then shifts by ilog2(0), triggering the
>>> UBSAN shift-out-of-bounds report in the report-zones and write paths.
>>>
>>> Drop the zoned feature from the head allocation and inherit it from
>>> the path namespace: the head limits refresh already stacks the zoned
>>> feature, the zone size and the zone resource limits from the path
>>> queue, so the head matches the path namespace and becomes zoned once a
>>> revalidation succeeds.  This also stops marking the head zoned when
>>> CONFIG_BLK_DEV_ZONED is off, which used to fail the head allocation
>>> with a WARN.
>> Where do we actually stack it right now?
> The head stacks it in nvme_update_ns_info(): after the path namespace
> update it's own limits, queue_limits_stack_bdev() stacks the path
> queue into the head, and blk_stack_limits() inherits the BLK_FEAT_ZONED
> through BLK_FEAT_INHERIT_MASK, along with chunk_sectors and the zone
> resource limits. That happens before the head is registered, so
> blk_revalidate_disk_zones() see the zoned feature in place.
>
Hi Christoph, any feedback?

Best Regards,
Guixin Liu
> Best Regards,
> Guixin Liu




^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH] nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known
  2026-09-14 10:57 [PATCH] nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known Guixin Liu
  2026-09-15  6:36 ` Christoph Hellwig
@ 2026-09-22 15:13 ` Keith Busch
  1 sibling, 0 replies; 5+ messages in thread
From: Keith Busch @ 2026-09-22 15:13 UTC (permalink / raw)
  To: Guixin Liu
  Cc: Jens Axboe, Christoph Hellwig, Sagi Grimberg, Nilay Shroff,
	Daniel Wagner, John Garry, Hannes Reinecke, linux-nvme

On Mon, Sep 14, 2026 at 06:57:13PM +0800, Guixin Liu wrote:
> The namespace head is marked zoned at allocation time based only on
> the command set identifier, before any zone information has been
> queried.  If the zone info query fails on the first scan, the path
> namespace is registered without zoned limits while the head still
> advertises the zoned capability with a zone size of zero.  Reporting
> zones or writing to the head then shifts by ilog2(0), triggering the
> UBSAN shift-out-of-bounds report in the report-zones and write paths.
> 
> Drop the zoned feature from the head allocation and inherit it from
> the path namespace: the head limits refresh already stacks the zoned
> feature, the zone size and the zone resource limits from the path
> queue, so the head matches the path namespace and becomes zoned once a
> revalidation succeeds.  This also stops marking the head zoned when
> CONFIG_BLK_DEV_ZONED is off, which used to fail the head allocation
> with a WARN.

Looks correct to me. Applied to nvme-7.3.


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-22 15:13 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-14 10:57 [PATCH] nvme-multipath: set BLK_FEAT_ZONED only after the zone info is known Guixin Liu
2026-09-15  6:36 ` Christoph Hellwig
2026-09-15  8:51   ` Guixin Liu
2026-09-22  8:17     ` Guixin Liu
2026-09-22 15:13 ` Keith Busch

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox