* [PATCH 0/7] Add support for storage element depopulation
@ 2026-10-05 9:46 Damien Le Moal
2026-10-05 9:46 ` [PATCH 1/7] block: fail reads to offline zones early Damien Le Moal
` (7 more replies)
0 siblings, 8 replies; 26+ messages in thread
From: Damien Le Moal @ 2026-10-05 9:46 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Jens,
These patches define a new set of block device operations for generically
using from the block layer the storage element depopulation feature of
zoned block devices. Support for this feature is added to the SCSI disk
driver and an emulation of this feature added to the zloop driver.
The management operations can be accessed using block layer API, which is
intended for file systems (e.g. zonefs and XFS), as well as using ioctls
for users using zoned disks directly from user space.
Damien Le Moal (7):
block: fail reads to offline zones early
block: introduce storage element management
block: add storage element management ioctls
zloop: add storage element emulation
zloop: add degrade_element control command
scsi: sd_zbc: always revalidate zones for disks supporting head depopulation
scsi: sd_zbc: define storage element management operations
.../admin-guide/blockdev/zoned_loop.rst | 8 +-
block/blk-core.c | 5 +
block/blk-zoned.c | 300 ++++++++-
block/blk.h | 15 +
block/ioctl.c | 5 +
drivers/block/zloop.c | 616 +++++++++++++++++-
drivers/scsi/sd.c | 1 +
drivers/scsi/sd.h | 5 +
drivers/scsi/sd_zbc.c | 236 ++++++-
include/linux/blkdev.h | 16 +
include/uapi/linux/blkzoned.h | 81 +++
include/uapi/linux/fs.h | 1 +
12 files changed, 1257 insertions(+), 32 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 26+ messages in thread
* [PATCH 1/7] block: fail reads to offline zones early
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
@ 2026-10-05 9:46 ` Damien Le Moal
2026-10-05 10:01 ` sashiko-bot
2026-10-05 10:45 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 2/7] block: introduce storage element management Damien Le Moal
` (6 subsequent siblings)
7 siblings, 2 replies; 26+ messages in thread
From: Damien Le Moal @ 2026-10-05 9:46 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Any BIO targeting an offline zone of a zoned block device will fail,
including read commands. So there is no point in issuing read BIOs.
Fail these early if we ever see one, and be quiet about the error.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
block/blk-core.c | 5 +++++
block/blk-zoned.c | 18 ++++++++++++++++++
block/blk.h | 6 ++++++
3 files changed, 29 insertions(+)
diff --git a/block/blk-core.c b/block/blk-core.c
index 13dc70e8f55d..f098a31d19aa 100644
--- a/block/blk-core.c
+++ b/block/blk-core.c
@@ -866,6 +866,11 @@ void submit_bio_noacct(struct bio *bio)
switch (bio_op(bio)) {
case REQ_OP_READ:
+ if (bdev_is_zoned(bdev) &&
+ bdev_zone_is_offline(bdev, bio->bi_iter.bi_sector)) {
+ bio_set_flag(bio, BIO_QUIET);
+ goto end_io;
+ }
break;
case REQ_OP_WRITE:
if (bio->bi_opf & REQ_ATOMIC) {
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index 19268afb8752..131c9f50b3da 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -308,6 +308,24 @@ bool bdev_zone_is_seq(struct block_device *bdev, sector_t sector)
}
EXPORT_SYMBOL_GPL(bdev_zone_is_seq);
+/**
+ * bdev_zone_is_offline - check if a sector belongs to an offline zone
+ * @bdev: block device to check
+ * @sector: sector number
+ *
+ * Check if @sector on @bdev is contained in an offline zone.
+ */
+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector)
+{
+ enum blk_zone_cond cond;
+
+ if (!bdev_is_zoned(bdev))
+ return false;
+
+ cond = disk_zone_get_cond(bdev->bd_disk, sector);
+ return cond == BLK_ZONE_COND_OFFLINE;
+}
+
/**
* bdev_zone_mgmt_allowed - check if management operations are allowed on a zone
* @bdev: block device to check
diff --git a/block/blk.h b/block/blk.h
index 2cc03aa54c53..c2d07347a3ad 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -578,6 +578,7 @@ int blkdev_report_zones_ioctl(struct block_device *bdev, unsigned int cmd,
int blkdev_zone_mgmt_ioctl(struct block_device *bdev, blk_mode_t mode,
unsigned int cmd, unsigned long arg);
bool bdev_zone_mgmt_allowed(struct block_device *bdev, sector_t sector);
+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector);
#else /* CONFIG_BLK_DEV_ZONED */
static inline void disk_init_zone_resources(struct gendisk *disk)
{
@@ -625,6 +626,11 @@ static inline bool bdev_zone_mgmt_allowed(struct block_device *bdev,
{
return false;
}
+static inline bool bdev_zone_is_offline(struct block_device *bdev,
+ sector_t sector)
+{
+ return false;
+}
#endif /* CONFIG_BLK_DEV_ZONED */
struct block_device *bdev_alloc(struct gendisk *disk, u8 partno);
--
2.55.0
^ permalink raw reply related [flat|nested] 26+ messages in thread
* [PATCH 2/7] block: introduce storage element management
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
2026-10-05 9:46 ` [PATCH 1/7] block: fail reads to offline zones early Damien Le Moal
@ 2026-10-05 9:46 ` Damien Le Moal
2026-10-05 10:01 ` sashiko-bot
` (2 more replies)
2026-10-05 9:46 ` [PATCH 3/7] block: add storage element management ioctls Damien Le Moal
` (5 subsequent siblings)
7 siblings, 3 replies; 26+ messages in thread
From: Damien Le Moal @ 2026-10-05 9:46 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Recent SCSI (SBC) and ATA (ACS) standards define the storage element
depopulation feature. This feature is intended for managing hard-disks
heads, either combined read-write heads or pairs of read and write heads,
allowing to keep disks with defective heads longer in production by
allowing "depopulating" (removing) defective heads.
This feature comes in two different flavors:
- A destructive version which removes a head and reformats the disk at a
lower capacity point, restoaring a fully functional contiguous LBA
address space.
- A data preserving version restricted to host-managed zoned disks, which
marks the zones served by a removed head as offline or read-only.
In preparation for using the data-preserving flavor of the depopulation
feature in file systems natively supporting zoned block devices, introduce
a set of storage element management operations and functions to define
generic calls into block device drivers for managing storage elements.
The set of operations is defined with struct blk_storage_elements_ops and
includes three operations:
- report_elements: get information on a device storage elements state
- remove_element: depopulate a defective storage element
- restore_elements: restore depopulated storage elements
Each operation is called from the functions
bdev_report_storage_elements(), bdev_remove_storage_element() and
bdev_restore_storage_elements().
Removing a healthy storage element from a device is possible and useful
for testing. The storage element restoration operation restore_elements
allows repopulating such healthy element. Repopulating defective storage
elements is generally not allowed by devices.
The device drivers of zoned block devices can indicate support for the
storage element depopulation feature by specifying the storage element
management operations with the se_ops field of struct
block_device_operations.
Co-developed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
block/blk-zoned.c | 168 +++++++++++++++++++++++++++++++++-
include/linux/blkdev.h | 16 ++++
include/uapi/linux/blkzoned.h | 67 ++++++++++++++
3 files changed, 250 insertions(+), 1 deletion(-)
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index 131c9f50b3da..b713512f7d48 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -18,6 +18,7 @@
#include <linux/mempool.h>
#include <linux/kthread.h>
#include <linux/freezer.h>
+#include <linux/delay.h>
#include <trace/events/block.h>
@@ -2718,5 +2719,170 @@ int queue_zone_wplugs_show(void *data, struct seq_file *m)
return 0;
}
-
#endif
+
+static void disk_wait_for_se_mgmt_completion(struct gendisk *disk)
+{
+ struct blk_storage_element *elements, *e;
+ unsigned int i, nr_se, nr_elements = 0;
+ int ret;
+
+ ret = (disk->fops->se_ops->report_elements)(disk, &nr_elements, NULL);
+ if (ret) {
+ pr_err("Failed to get number of storage elements\n");
+ return;
+ }
+
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return;
+
+ while (1) {
+ /*
+ * Check if we have storage elements being removed or restored.
+ */
+ nr_se = nr_elements;
+ ret = (disk->fops->se_ops->report_elements)(disk,
+ &nr_elements, elements);
+ if (ret) {
+ pr_err("Failed to get storage elements\n");
+ break;
+ }
+
+ e = elements;
+ for (i = 0; i < nr_se; i++, e++) {
+ if (e->status == BLK_SE_STS_REMOVE_IN_PROGRESS ||
+ e->status == BLK_SE_STS_RESTORE_IN_PROGRESS)
+ break;
+ }
+ if (i >= nr_se)
+ break;
+
+ /* Not done yet: wait and retry. */
+ msleep(500);
+ }
+
+ kfree(elements);
+}
+
+/**
+ * bdev_report_storage_elements - report the storage elements of a block device
+ *
+ * Fill at most @nr_elements storage element descriptors in the array @elements.
+ * The number of storage elements filled in the array is returned using
+ * @nr_elements. If @elements is NULL, only @nr_elements is returned.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_report_storage_elements(struct block_device *bdev,
+ unsigned int *nr_elements,
+ struct blk_storage_element *elements)
+{
+ struct gendisk *disk = bdev->bd_disk;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ if (!nr_elements)
+ return -EINVAL;
+
+ if (*nr_elements && !elements)
+ return -EINVAL;
+
+ return (disk->fops->se_ops->report_elements)(disk, nr_elements,
+ elements);
+}
+EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
+
+/**
+ * bdev_remove_storage_element - Remove (depopulate) a storage element of a
+ * block device
+ *
+ * Remove (depopulate) the storage element identified by @element_id from the
+ * block device @bdev.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /* Zero is not a valid storage element ID. */
+ if (!element_id)
+ return -EINVAL;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding command.
+ * The caller is responsible for not queuing up more I/Os by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ /*
+ * Flush the volatile write cache if there is any. Note that this may
+ * lead to errors if degraded storage elements are used. So ignore
+ * errors here.
+ */
+ blkdev_issue_flush(disk->part0);
+
+ /* Invalidate all cached data. */
+ invalidate_inode_pages2_range(bdev->bd_mapping, 0,
+ get_capacity(disk) << SECTOR_SHIFT);
+
+ ret = (disk->fops->se_ops->remove_element)(disk, element_id);
+ if (ret)
+ return ret;
+
+ /* Revalidate the device zones once the opration completes. */
+ disk_wait_for_se_mgmt_completion(disk);
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
+
+/**
+ * bdev_restore_storage_elements - Restore all depopulated storage elements of a
+ * block device
+ *
+ * Restore all storage elements of @bdev that have been depopulated. Not all
+ * elements may be restored by this operation.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_restore_storage_elements(struct block_device *bdev)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding commands.
+ * The caller is responsible for not queuing up more I/O by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ /* Invalidate all cached data. */
+ invalidate_inode_pages2_range(bdev->bd_mapping, 0,
+ get_capacity(disk) << SECTOR_SHIFT);
+
+ ret = (disk->fops->se_ops->restore_elements)(disk);
+ if (ret)
+ return ret;
+
+ disk_wait_for_se_mgmt_completion(disk);
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index d003a9d2d1f6..0b2013c96cad 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1574,6 +1574,21 @@ enum blk_unique_id {
BLK_UID_NAA = 3,
};
+struct blk_storage_elements_ops {
+ int (*report_elements)(struct gendisk *disk,
+ unsigned int *nr_elements,
+ struct blk_storage_element *elements);
+ int (*remove_element)(struct gendisk *disk, unsigned int element_id);
+ int (*restore_elements)(struct gendisk *disk);
+};
+
+int bdev_report_storage_elements(struct block_device *bdev,
+ unsigned int *nr_elements,
+ struct blk_storage_element *elements);
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id);
+int bdev_restore_storage_elements(struct block_device *bdev);
+
struct block_device_operations {
void (*submit_bio)(struct bio *bio);
int (*poll_bio)(struct bio *bio, struct io_comp_batch *iob,
@@ -1601,6 +1616,7 @@ struct block_device_operations {
enum blk_unique_id id_type);
struct module *owner;
const struct pr_ops *pr_ops;
+ const struct blk_storage_elements_ops *se_ops;
/*
* Special callback for probing GPT entry at a given sector.
diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
index 663836120966..3a4bacfe7979 100644
--- a/include/uapi/linux/blkzoned.h
+++ b/include/uapi/linux/blkzoned.h
@@ -208,4 +208,71 @@ struct blk_zone_range {
#define BLKFINISHZONE _IOW(0x12, 136, struct blk_zone_range)
#define BLKREPORTZONEV2 _IOWR(0x12, 142, struct blk_zone_report)
+/**
+ * enum blk_storage_element_status - Statuc of a zoned device storage elements.
+ *
+ * @BLK_SE_RDWR: The storage element handles both reads and writes.
+ * @BLK_SE_READ: The storage element handles reads only.
+ * @BLK_SE_WRITE: The storage element handles writes only.
+ * @BLK_SE_WRITE: The storage element type is unknown.
+ */
+enum blk_storage_element_type {
+ BLK_SE_TYPE_RDWR = 0x01,
+ BLK_SE_TYPE_READ = 0x02,
+ BLK_SE_TYPE_WRITE = 0x03,
+ BLK_SE_TYPE_UNKNOWN = 0xFF,
+};
+
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_STS_OK: The storage element is operating normally.
+ * @BLK_SE_STS_DEGRADED: The storage element has degraded and is not operating
+ * normally.
+ * @BLK_SE_STS_REMOVE_IN_PROGRESS: The storage element is being removed.
+ * @BLK_SE_STS_REMOVE_ERROR: The storage element removal failed.
+ * @BLK_SE_STS_RESTORE_IN_PROGRESS: The storage element is being restored.
+ * @BLK_SE_STS_RESTORE_ERROR: The storage element restoration failed.
+ * @BLK_SE_STS_REMOVED: The storage element was removed.
+ * @BLK_SE_STS_UNKNOWN: The storage element status is unknown.
+ */
+enum blk_storage_element_status {
+ BLK_SE_STS_OK = 0x01,
+ BLK_SE_STS_DEGRADED = 0x02,
+ BLK_SE_STS_REMOVE_IN_PROGRESS = 0x03,
+ BLK_SE_STS_REMOVE_ERROR = 0x04,
+ BLK_SE_STS_RESTORE_IN_PROGRESS = 0x05,
+ BLK_SE_STS_RESTORE_ERROR = 0x06,
+ BLK_SE_STS_REMOVED = 0x07,
+ BLK_SE_STS_UNKOWN = 0xFF,
+};
+
+/**
+ * struct blk_storage_element - Zoned device storage element descriptor.
+ *
+ * @id: The ID of the element (cannot be 0).
+ * @paired_id: The ID of the paired element for an element that is not
+ * of type BLK_SE_TYPE_RDWR.
+ * @type: The type of the storage element (enum blk_storage_element_type).
+ * @status: The health status of the storage element
+ * (enum blk_storage_element_status).
+ * @restore_allowed: Indicate if the storage element can be restored.
+ * @nr_zones: The number of zones that the storage element handles.
+ */
+struct blk_storage_element {
+ __u32 id;
+ __u32 paired_id;
+ __u64 nr_zones;
+ __u8 type;
+ __u8 status;
+ __u8 restore_allowed;
+ __u8 reserved[5];
+};
+
+struct blk_storage_elements_report {
+ __u32 nr_elements;
+ __u32 reserved;
+ struct blk_storage_element elements[];
+};
+
#endif /* _UAPI_BLKZONED_H */
--
2.55.0
^ permalink raw reply related [flat|nested] 26+ messages in thread
* [PATCH 3/7] block: add storage element management ioctls
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
2026-10-05 9:46 ` [PATCH 1/7] block: fail reads to offline zones early Damien Le Moal
2026-10-05 9:46 ` [PATCH 2/7] block: introduce storage element management Damien Le Moal
@ 2026-10-05 9:46 ` Damien Le Moal
2026-10-05 10:00 ` sashiko-bot
2026-10-05 11:09 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 4/7] zloop: add storage element emulation Damien Le Moal
` (4 subsequent siblings)
7 siblings, 2 replies; 26+ messages in thread
From: Damien Le Moal @ 2026-10-05 9:46 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Define the new ioctl commands BLKGETNRSTORELEMS, BLKREPORTSTORELEMS,
BLKREMOVESTORELEM, and BLKRESTORESTORELEMS to provide users accessing
zoned block devices directly with an interface to the storage elements
management operations of the block layer.
The ioctls BLKGETNRSTORELEMS and BLKREPORTSTORELEMS are defined to
respectively get the number of storage elements of a zoned device and to
get an array of struct blk_storage_element describing the current state of
the device storage elements.
The ioctl BLKREMOVESTORELEM can be used to remove (depopulate) a storage
element that has a degraded status (i.e. BLK_SE_STS_DEGRADED).
Finally, the BLKRESTORESTORELEMS ioctl interfaces with
bdev_restore_storage_elements() to restore the depopulated storage
elements of a zoned device.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
block/blk-zoned.c | 114 ++++++++++++++++++++++++++++++++++
block/blk.h | 9 +++
block/ioctl.c | 5 ++
include/uapi/linux/blkzoned.h | 14 +++++
include/uapi/linux/fs.h | 1 +
5 files changed, 143 insertions(+)
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index b713512f7d48..62b9cf21e9ad 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -19,6 +19,7 @@
#include <linux/kthread.h>
#include <linux/freezer.h>
#include <linux/delay.h>
+#include <linux/uaccess.h>
#include <trace/events/block.h>
@@ -2794,6 +2795,69 @@ int bdev_report_storage_elements(struct block_device *bdev,
}
EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
+static int blkdev_get_nr_storage_elements_ioctl(struct block_device *bdev,
+ void __user *argp)
+{
+ unsigned int nr_elements = 0;
+ int ret;
+
+ ret = bdev_report_storage_elements(bdev, &nr_elements, NULL);
+ if (ret)
+ return ret;
+
+ if (put_user(nr_elements, (unsigned int __user *)argp))
+ return -EFAULT;
+
+ return 0;
+}
+
+static int blkdev_report_storage_elements_ioctl(struct block_device *bdev,
+ void __user *argp)
+{
+ struct blk_storage_elements_report rep;
+ struct blk_storage_element *elements;
+ unsigned int nr_elements;
+ unsigned long retc;
+ int ret;
+
+ if (!argp)
+ return -EINVAL;
+
+ if (copy_from_user(&rep, argp,
+ sizeof(struct blk_storage_elements_report)))
+ return -EFAULT;
+
+ if (!rep.nr_elements)
+ return -EINVAL;
+
+ nr_elements = rep.nr_elements;
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return -ENOMEM;
+
+ ret = bdev_report_storage_elements(bdev, &nr_elements, elements);
+ if (ret)
+ goto free_elements;
+
+ retc = copy_to_user(argp + sizeof(struct blk_storage_elements_report),
+ elements,
+ sizeof(struct blk_storage_element) * nr_elements);
+ if (retc) {
+ ret = -EFAULT;
+ goto free_elements;
+ }
+
+ rep.nr_elements = nr_elements;
+ retc = copy_to_user(argp, &rep,
+ sizeof(struct blk_storage_elements_report));
+ if (retc)
+ ret = -EFAULT;
+
+free_elements:
+ kfree(elements);
+ return ret;
+}
+
/**
* bdev_remove_storage_element - Remove (depopulate) a storage element of a
* block device
@@ -2847,6 +2911,22 @@ int bdev_remove_storage_element(struct block_device *bdev,
}
EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
+static int blkdev_remove_storage_element_ioctl(struct block_device *bdev,
+ blk_mode_t mode, void __user *argp)
+{
+ unsigned int element_id;
+
+ if (!(mode & BLK_OPEN_WRITE))
+ return -EBADF;
+ if (bdev_read_only(bdev))
+ return -EPERM;
+
+ if (get_user(element_id, (unsigned int __user *)argp))
+ return -EFAULT;
+
+ return bdev_remove_storage_element(bdev, element_id);
+}
+
/**
* bdev_restore_storage_elements - Restore all depopulated storage elements of a
* block device
@@ -2886,3 +2966,37 @@ int bdev_restore_storage_elements(struct block_device *bdev)
return blk_revalidate_disk_zones(disk);
}
EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
+
+static int blkdev_restore_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode)
+{
+ if (!(mode & BLK_OPEN_WRITE))
+ return -EBADF;
+ if (bdev_read_only(bdev))
+ return -EPERM;
+
+ return bdev_restore_storage_elements(bdev);
+}
+
+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp)
+{
+ if (!bdev_is_zoned(bdev) || !bdev->bd_disk->fops->se_ops)
+ return -ENOTTY;
+
+ switch (cmd) {
+ case BLKGETNRSTORELEMS:
+ return blkdev_get_nr_storage_elements_ioctl(bdev, argp);
+ case BLKREPORTSTORELEMS:
+ return blkdev_report_storage_elements_ioctl(bdev, argp);
+ case BLKREMOVESTORELEM:
+ return blkdev_remove_storage_element_ioctl(bdev, mode, argp);
+ case BLKRESTORESTORELEMS:
+ return blkdev_restore_storage_elements_ioctl(bdev, mode);
+ default:
+ break;
+ }
+
+ return -ENOTTY;
+}
diff --git a/block/blk.h b/block/blk.h
index c2d07347a3ad..274afb46a809 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -579,6 +579,9 @@ int blkdev_zone_mgmt_ioctl(struct block_device *bdev, blk_mode_t mode,
unsigned int cmd, unsigned long arg);
bool bdev_zone_mgmt_allowed(struct block_device *bdev, sector_t sector);
bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector);
+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp);
#else /* CONFIG_BLK_DEV_ZONED */
static inline void disk_init_zone_resources(struct gendisk *disk)
{
@@ -631,6 +634,12 @@ static inline bool bdev_zone_is_offline(struct block_device *bdev,
{
return false;
}
+static inline int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp)
+{
+ return -ENOTTY;
+}
#endif /* CONFIG_BLK_DEV_ZONED */
struct block_device *bdev_alloc(struct gendisk *disk, u8 partno);
diff --git a/block/ioctl.c b/block/ioctl.c
index 64b4e6c0f696..ccf806c2e37d 100644
--- a/block/ioctl.c
+++ b/block/ioctl.c
@@ -679,6 +679,11 @@ static int blkdev_common_ioctl(struct block_device *bdev, blk_mode_t mode,
return put_uint(argp, bdev_zone_sectors(bdev));
case BLKGETNRZONES:
return put_uint(argp, bdev_nr_zones(bdev));
+ case BLKGETNRSTORELEMS:
+ case BLKREPORTSTORELEMS:
+ case BLKREMOVESTORELEM:
+ case BLKRESTORESTORELEMS:
+ return blkdev_zone_storage_elements_ioctl(bdev, mode, cmd, argp);
case BLKROGET:
return put_int(argp, bdev_read_only(bdev) != 0);
case BLKSSZGET: /* get block device logical block size */
diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
index 3a4bacfe7979..bd2c55d1a0d2 100644
--- a/include/uapi/linux/blkzoned.h
+++ b/include/uapi/linux/blkzoned.h
@@ -275,4 +275,18 @@ struct blk_storage_elements_report {
struct blk_storage_element elements[];
};
+/**
+ * Zoned block device storage element management ioctl's:
+ *
+ * @BLKGETNRSTORELEMS: Get the number of storage elements of the device.
+ * @BLKREPORTSTORELEMS: Get the device storage elements.
+ * @BLKREMOVESTORELEM: Remove (depopulate) one storage element of a device.
+ * @BLKRESTORESTORELEM: Restore (repopulate if possible) all storage elements
+ * that have been removed.
+ */
+#define BLKGETNRSTORELEMS _IOR(0x12, 143, __u32)
+#define BLKREPORTSTORELEMS _IOWR(0x12, 144, struct blk_storage_elements_report)
+#define BLKREMOVESTORELEM _IOW(0x12, 145, __u32)
+#define BLKRESTORESTORELEMS _IO(0x12, 146)
+
#endif /* _UAPI_BLKZONED_H */
diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index 34c6f219462a..8a979326aa7f 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -309,6 +309,7 @@ struct file_attr {
/* 130-136 and 142 are used by zoned block device ioctls (uapi/linux/blkzoned.h) */
/* 137-141 are used by blk-crypto ioctls (uapi/linux/blk-crypto.h) */
#define BLKTRACESETUP2 _IOWR(0x12, 142, struct blk_user_trace_setup2)
+/* 143-146 are used by storage element management for zoned block devices. */
#define BMAP_IOCTL 1 /* obsolete - kept for compatibility */
#define FIBMAP _IO(0x00,1) /* bmap access */
--
2.55.0
^ permalink raw reply related [flat|nested] 26+ messages in thread
* [PATCH 4/7] zloop: add storage element emulation
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
` (2 preceding siblings ...)
2026-10-05 9:46 ` [PATCH 3/7] block: add storage element management ioctls Damien Le Moal
@ 2026-10-05 9:46 ` Damien Le Moal
2026-10-05 9:58 ` sashiko-bot
2026-10-05 11:14 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 5/7] zloop: add degrade_element control command Damien Le Moal
` (3 subsequent siblings)
7 siblings, 2 replies; 26+ messages in thread
From: Damien Le Moal @ 2026-10-05 9:46 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Emulate the storage element depopulation feature of SMR disks in the
zloop driver. This emulation is controlled using the new stor_elements
option.
This option can take several values:
- ZLOOP_STOR_ELEMENTS_NONE (0): no emulation (default)
- ZLOOP_STOR_ELEMENTS_RDWR (1): emulate all access storage elements
(e.g. read+write heads)
- ZLOOP_STOR_ELEMENTS_PAIRS (2): emulate fractional access storage
elements (e.g. pairs of read and write heads)
If enabled with the value 1 or 2, the number of storage elements, or of
pairs of fractional access storage elements, is automatically calculated
based on the number of zones of the device so that we have at least 2 and
at most 32 storage elements (or 2 pairs of fractional access storage
elements).
The mapping of zones to storage elements is defined simply as the modulo
of a zone number and a storage element ID. E.g, removing a storage
element from a set of 4 storage elements will offline 1 zone every 4
zones.
The storage element management operations are specified with the
zloop_se_ops (struct blk_storage_elements_ops).
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
.../admin-guide/blockdev/zoned_loop.rst | 8 +-
drivers/block/zloop.c | 510 +++++++++++++++++-
2 files changed, 490 insertions(+), 28 deletions(-)
diff --git a/Documentation/admin-guide/blockdev/zoned_loop.rst b/Documentation/admin-guide/blockdev/zoned_loop.rst
index 64277494fb36..ec39f141af8d 100644
--- a/Documentation/admin-guide/blockdev/zoned_loop.rst
+++ b/Documentation/admin-guide/blockdev/zoned_loop.rst
@@ -61,7 +61,7 @@ The options available for the add command can be listed by reading the
/dev/zloop-control device::
$ cat /dev/zloop-control
- add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache
+ add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache,stor_elements=%u
remove id=%d
In more details, the options that can be used with the "add" command are as
@@ -113,6 +113,12 @@ discard_write_cache Discard all data that was not explicitly persisted using a
each zone file to the size recorded during the last flush
operation. This simulates power fail events where
uncommitted data is lost.
+stor_elements Control storage element emulation. The default value is 0,
+ indicating no emulation. A value of 1 indicates that all
+ access storage elements (equivalent to read+write head of
+ a disk) are emulated. A value of 2 enables fractional
+ access storage element (equivalent to pairs of read and
+ write heads of a disk) emulation .
=================== =========================================================
3) Deleting a Zoned Device
diff --git a/drivers/block/zloop.c b/drivers/block/zloop.c
index f0ca221524db..7e6b5cc8017d 100644
--- a/drivers/block/zloop.c
+++ b/drivers/block/zloop.c
@@ -37,6 +37,7 @@ enum {
ZLOOP_OPT_ORDERED_ZONE_APPEND = (1 << 10),
ZLOOP_OPT_DISCARD_WRITE_CACHE = (1 << 11),
ZLOOP_OPT_MAX_OPEN_ZONES = (1 << 12),
+ ZLOOP_OPT_STOR_ELEMENTS = (1 << 13),
};
static const match_table_t zloop_opt_tokens = {
@@ -51,11 +52,22 @@ static const match_table_t zloop_opt_tokens = {
{ ZLOOP_OPT_BUFFERED_IO, "buffered_io" },
{ ZLOOP_OPT_ZONE_APPEND, "zone_append=%u" },
{ ZLOOP_OPT_ORDERED_ZONE_APPEND, "ordered_zone_append" },
- { ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
+ { ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
{ ZLOOP_OPT_MAX_OPEN_ZONES, "max_open_zones=%u" },
+ { ZLOOP_OPT_STOR_ELEMENTS, "stor_elements=%u" },
{ ZLOOP_OPT_ERR, NULL }
};
+/* Storage elements emulation types. */
+enum zloop_stor_elements {
+ /* No emulation. */
+ ZLOOP_STOR_ELEMENTS_NONE,
+ /* Emulate read+write storage elements. */
+ ZLOOP_STOR_ELEMENTS_RDWR,
+ /* Emulate pairs of associated read and write storage elements. */
+ ZLOOP_STOR_ELEMENTS_PAIRS,
+};
+
/* Default values for the "add" operation. */
#define ZLOOP_DEF_ID -1
#define ZLOOP_DEF_ZONE_SIZE ((256ULL * SZ_1M) >> SECTOR_SHIFT)
@@ -68,6 +80,7 @@ static const match_table_t zloop_opt_tokens = {
#define ZLOOP_DEF_BUFFERED_IO false
#define ZLOOP_DEF_ZONE_APPEND true
#define ZLOOP_DEF_ORDERED_ZONE_APPEND false
+#define ZLOOP_DEF_STOR_ELEMENTS ZLOOP_STOR_ELEMENTS_NONE
/* Arbitrary limit on the zone size (16GB). */
#define ZLOOP_MAX_ZONE_SIZE_MB 16384
@@ -87,6 +100,7 @@ struct zloop_options {
bool zone_append;
bool ordered_zone_append;
bool discard_write_cache;
+ enum zloop_stor_elements stor_elements;
};
/*
@@ -117,6 +131,8 @@ struct zloop_zone {
enum blk_zone_cond cond;
sector_t start;
sector_t wp;
+ unsigned int wr_se_id;
+ unsigned int rd_se_id;
gfp_t old_gfp_mask;
};
@@ -133,6 +149,7 @@ struct zloop_device {
bool zone_append;
bool ordered_zone_append;
bool discard_write_cache;
+ enum zloop_stor_elements stor_elements;
const char *base_dir;
struct file *data_dir;
@@ -150,6 +167,17 @@ struct zloop_device {
struct list_head open_zones_lru_list;
unsigned int nr_open_zones;
+ /* For storage elements emulation. */
+ struct mutex stor_elements_lock;
+ unsigned int nr_elements;
+ unsigned int max_nr_removed_elements;
+ unsigned int nr_removed_elements;
+ struct delayed_work remove_element_work;
+ unsigned int remove_element_id;
+ struct delayed_work restore_elements_work;
+ bool restore_in_progress;
+ struct blk_storage_element *elements;
+
struct zloop_zone zones[] __counted_by(nr_zones);
};
@@ -289,22 +317,30 @@ static bool zloop_do_open_zone(struct zloop_device *zlo,
}
}
-static void zloop_mark_full(struct zloop_device *zlo, struct zloop_zone *zone)
+static void zloop_set_zone_cond(struct zloop_device *zlo,
+ struct zloop_zone *zone,
+ enum blk_zone_cond cond)
{
lockdep_assert_held(&zone->wp_lock);
zloop_lru_remove_open_zone(zlo, zone);
- zone->cond = BLK_ZONE_COND_FULL;
- zone->wp = ULLONG_MAX;
+ zone->cond = cond;
+ if (cond == BLK_ZONE_COND_EMPTY)
+ zone->wp = zone->start;
+ else
+ zone->wp = ULLONG_MAX;
}
-static void zloop_mark_empty(struct zloop_device *zlo, struct zloop_zone *zone)
+static inline void zloop_set_zone_full(struct zloop_device *zlo,
+ struct zloop_zone *zone)
{
- lockdep_assert_held(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_FULL);
+}
- zloop_lru_remove_open_zone(zlo, zone);
- zone->cond = BLK_ZONE_COND_EMPTY;
- zone->wp = zone->start;
+static inline void zloop_set_zone_empty(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_EMPTY);
}
static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
@@ -339,9 +375,9 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
spin_lock(&zone->wp_lock);
if (!file_sectors) {
- zloop_mark_empty(zlo, zone);
+ zloop_set_zone_empty(zlo, zone);
} else if (file_sectors == zlo->zone_capacity) {
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
} else {
if (zone->cond != BLK_ZONE_COND_IMP_OPEN &&
zone->cond != BLK_ZONE_COND_EXP_OPEN)
@@ -363,6 +399,12 @@ static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zone->cond == BLK_ZONE_COND_OFFLINE ||
+ zone->cond == BLK_ZONE_COND_READONLY) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
ret = zloop_update_seq_zone(zlo, zone_no);
if (ret)
@@ -388,6 +430,12 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zone->cond == BLK_ZONE_COND_OFFLINE ||
+ zone->cond == BLK_ZONE_COND_READONLY) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
ret = zloop_update_seq_zone(zlo, zone_no);
if (ret)
@@ -420,7 +468,22 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)
return ret;
}
-static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)
+static int zloop_do_reset_zone(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ if (vfs_truncate(&zone->file->f_path, 0))
+ return -EIO;
+
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_empty(zlo, zone);
+ clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
+ spin_unlock(&zone->wp_lock);
+
+ return 0;
+}
+
+static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no,
+ bool all_zones)
{
struct zloop_zone *zone = &zlo->zones[zone_no];
int ret = 0;
@@ -430,20 +493,20 @@ static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zone->cond == BLK_ZONE_COND_OFFLINE ||
+ zone->cond == BLK_ZONE_COND_READONLY) {
+ if (!all_zones)
+ ret = -EIO;
+ goto unlock;
+ }
+
if (!test_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags) &&
zone->cond == BLK_ZONE_COND_EMPTY)
goto unlock;
- if (vfs_truncate(&zone->file->f_path, 0)) {
+ ret = zloop_do_reset_zone(zlo, zone);
+ if (ret)
set_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
- ret = -EIO;
- goto unlock;
- }
-
- spin_lock(&zone->wp_lock);
- zloop_mark_empty(zlo, zone);
- clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
- spin_unlock(&zone->wp_lock);
unlock:
mutex_unlock(&zone->lock);
@@ -457,7 +520,7 @@ static int zloop_reset_all_zones(struct zloop_device *zlo)
int ret;
for (i = zlo->nr_conv_zones; i < zlo->nr_zones; i++) {
- ret = zloop_reset_zone(zlo, i);
+ ret = zloop_reset_zone(zlo, i, true);
if (ret)
return ret;
}
@@ -475,6 +538,12 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zone->cond == BLK_ZONE_COND_OFFLINE ||
+ zone->cond == BLK_ZONE_COND_READONLY) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (!test_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags) &&
zone->cond == BLK_ZONE_COND_FULL)
goto unlock;
@@ -487,7 +556,7 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)
}
spin_lock(&zone->wp_lock);
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
spin_unlock(&zone->wp_lock);
@@ -624,7 +693,7 @@ static int zloop_seq_write_prep(struct zloop_cmd *cmd)
if (!is_append || !zlo->ordered_zone_append) {
zone->wp += nr_sectors;
if (zone->wp == zone_end)
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
}
out_unlock:
spin_unlock(&zone->wp_lock);
@@ -664,6 +733,11 @@ static void zloop_rw(struct zloop_cmd *cmd)
zone->start + zlo->zone_size))
goto out;
+ if (zone->cond == BLK_ZONE_COND_OFFLINE)
+ goto out;
+ if (zone->cond == BLK_ZONE_COND_READONLY && is_write)
+ goto out;
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
mutex_lock(&zone->lock);
ret = zloop_update_seq_zone(zlo, zone_no);
@@ -766,7 +840,7 @@ static void zloop_handle_cmd(struct zloop_cmd *cmd)
cmd->ret = zloop_flush(zlo);
break;
case REQ_OP_ZONE_RESET:
- cmd->ret = zloop_reset_zone(zlo, rq_zone_no(rq));
+ cmd->ret = zloop_reset_zone(zlo, rq_zone_no(rq), false);
break;
case REQ_OP_ZONE_RESET_ALL:
cmd->ret = zloop_reset_all_zones(zlo);
@@ -877,7 +951,7 @@ static bool zloop_set_zone_append_sector(struct request *rq)
rq->__sector = zone->wp;
zone->wp += blk_rq_sectors(rq);
if (zone->wp >= zone_end)
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
spin_unlock(&zone->wp_lock);
@@ -985,11 +1059,251 @@ static int zloop_report_zones(struct gendisk *disk, sector_t sector,
return nr_zones;
}
+static int zloop_report_elements(struct gendisk *disk,
+ unsigned int *nr_elements,
+ struct blk_storage_element *elements)
+{
+ struct zloop_device *zlo = disk->private_data;
+ unsigned int nr_report = *nr_elements;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ *nr_elements = zlo->nr_elements;
+ if (elements) {
+ struct blk_storage_element *se = zlo->elements;
+ unsigned int i;
+
+ for (i = 0; i < min(nr_report, zlo->nr_elements); i++, se++) {
+ switch (se->status) {
+ case BLK_SE_STS_REMOVED:
+ case BLK_SE_STS_RESTORE_ERROR:
+ se->restore_allowed = 1;
+ break;
+ default:
+ se->restore_allowed = 0;
+ }
+ memcpy(&elements[i], se,
+ sizeof(struct blk_storage_element));
+ }
+ }
+
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return 0;
+}
+
+static void zloop_remove_element_work(struct work_struct *work)
+{
+ struct zloop_device *zlo = container_of(work, struct zloop_device,
+ remove_element_work.work);
+ struct blk_storage_element *se, *paired_se = NULL;
+ enum blk_zone_cond cond;
+ struct zloop_zone *zone;
+ unsigned int i;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ /*
+ * If the element to remove is a read element, zones must go offline and
+ * the associated write element is also removed.
+ */
+ se = &zlo->elements[zlo->remove_element_id - 1];
+ switch (se->type) {
+ case BLK_SE_TYPE_RDWR:
+ cond = BLK_ZONE_COND_OFFLINE;
+ break;
+ case BLK_SE_TYPE_READ:
+ cond = BLK_ZONE_COND_OFFLINE;
+ paired_se = &zlo->elements[se->paired_id - 1];
+ break;
+ case BLK_SE_TYPE_WRITE:
+ cond = BLK_ZONE_COND_READONLY;
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ }
+
+ /*
+ * Change the condition of the zones owned by the (pair of) elements
+ * being removed and mark the elements removed.
+ */
+ for (i = 0, zone = zlo->zones; i < zlo->nr_zones; i++, zone++) {
+ if (zone->wr_se_id != se->id && zone->rd_se_id != se->id)
+ continue;
+ mutex_lock(&zone->lock);
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, cond);
+ spin_unlock(&zone->wp_lock);
+ mutex_unlock(&zone->lock);
+ }
+
+ se->status = BLK_SE_STS_REMOVED;
+
+ zlo->nr_removed_elements++;
+ if (paired_se) {
+ paired_se->status = BLK_SE_STS_REMOVED;
+ zlo->nr_removed_elements++;
+ }
+
+ zlo->remove_element_id = 0;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+}
+
+static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)
+{
+ struct zloop_device *zlo = disk->private_data;
+ struct blk_storage_element *se, *paired_se = NULL;
+ unsigned int nr_remove = 1;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ if (element_id > zlo->nr_elements)
+ return -EINVAL;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ if (zlo->remove_element_id || zlo->restore_in_progress) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /*
+ * Get the element to remove. If it is already removed, we have nothing
+ * to do.
+ */
+ se = &zlo->elements[element_id - 1];
+ if (se->status == BLK_SE_STS_REMOVED)
+ goto unlock;
+
+ /*
+ * If the element to remove is a read element, the associated write
+ * element must also be removed.
+ */
+ if (se->type == BLK_SE_TYPE_READ) {
+ paired_se = &zlo->elements[se->paired_id - 1];
+ nr_remove = 2;
+ }
+
+ if (zlo->nr_removed_elements + nr_remove >
+ zlo->max_nr_removed_elements) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /*
+ * Schedule the element removal with a delay, to emulate the (generally
+ * short) time it takes for a real device to depopulate a head and
+ * modify the zones.
+ */
+ zlo->remove_element_id = element_id;
+ se->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+ if (paired_se)
+ paired_se->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+
+ schedule_delayed_work(&zlo->remove_element_work,
+ msecs_to_jiffies(2000));
+
+unlock:
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
+static void zloop_restore_elements_work(struct work_struct *work)
+{
+ struct zloop_device *zlo = container_of(work, struct zloop_device,
+ restore_elements_work.work);
+ struct zloop_zone *zone = zlo->zones;
+ struct blk_storage_element *se;
+ unsigned int i;
+ int ret = 0;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ /* Reset all zones. */
+ for (i = 0; i < zlo->nr_zones && ret == 0; i++, zone++) {
+ mutex_lock(&zone->lock);
+ if (test_bit(ZLOOP_ZONE_CONV, &zone->flags)) {
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_NOT_WP);
+ spin_unlock(&zone->wp_lock);
+ } else {
+ ret = zloop_do_reset_zone(zlo, zone);
+ }
+ mutex_unlock(&zone->lock);
+ }
+
+ /* Restore all removed elements. */
+ for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
+ if (se->status != BLK_SE_STS_RESTORE_IN_PROGRESS)
+ continue;
+ if (!ret)
+ se->status = BLK_SE_STS_OK;
+ else
+ se->status = BLK_SE_STS_RESTORE_ERROR;
+ }
+
+ zlo->nr_removed_elements = 0;
+ zlo->restore_in_progress = false;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+}
+
+static int zloop_restore_elements(struct gendisk *disk)
+{
+ struct zloop_device *zlo = disk->private_data;
+ struct blk_storage_element *se;
+ unsigned int i;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ if (zlo->remove_element_id || zlo->restore_in_progress) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /* If we have no removed elements, we have nothing to do. */
+ if (!zlo->nr_removed_elements)
+ goto unlock;
+
+ /*
+ * Schedule the elements restoration with a delay, to emulate the time
+ * it takes for a real device to restore all removed heads and reset
+ * all zones.
+ */
+ zlo->restore_in_progress = true;
+ for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
+ if (se->status == BLK_SE_STS_REMOVED)
+ se->status = BLK_SE_STS_RESTORE_IN_PROGRESS;
+ }
+
+ schedule_delayed_work(&zlo->restore_elements_work,
+ msecs_to_jiffies(5000));
+
+unlock:
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
static void zloop_free_disk(struct gendisk *disk)
{
struct zloop_device *zlo = disk->private_data;
unsigned int i;
+ cancel_delayed_work_sync(&zlo->remove_element_work);
+ cancel_delayed_work_sync(&zlo->restore_elements_work);
+
blk_mq_free_tag_set(&zlo->tag_set);
for (i = 0; i < zlo->nr_zones; i++) {
@@ -1002,15 +1316,24 @@ static void zloop_free_disk(struct gendisk *disk)
fput(zlo->data_dir);
destroy_workqueue(zlo->workqueue);
+ kfree(zlo->elements);
kfree(zlo->base_dir);
kvfree(zlo);
}
+
+static const struct blk_storage_elements_ops zloop_se_ops = {
+ .report_elements = zloop_report_elements,
+ .remove_element = zloop_remove_element,
+ .restore_elements = zloop_restore_elements,
+};
+
static const struct block_device_operations zloop_fops = {
.owner = THIS_MODULE,
.open = zloop_open,
.report_zones = zloop_report_zones,
.free_disk = zloop_free_disk,
+ .se_ops = &zloop_se_ops,
};
__printf(3, 4)
@@ -1095,6 +1418,17 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,
if (!opts->buffered_io)
oflags |= O_DIRECT;
+ if (zlo->stor_elements != ZLOOP_STOR_ELEMENTS_NONE) {
+ unsigned int nr_elems = zlo->nr_elements;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ nr_elems /= 2;
+ zone->wr_se_id = (zone_no % nr_elems) + 1;
+ zone->rd_se_id = zone->wr_se_id;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ zone->rd_se_id += nr_elems;
+ }
+
if (zone_no < zlo->nr_conv_zones) {
/* Conventional zone file. */
set_bit(ZLOOP_ZONE_CONV, &zone->flags);
@@ -1165,6 +1499,93 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,
return ret;
}
+#define ZLOOP_MIN_STOR_ELEMENTS 2
+#define ZLOOP_MAX_STOR_ELEMENTS 32
+#define ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS 32
+
+static int zloop_create_storage_elements(struct zloop_device *zlo)
+{
+ struct blk_storage_element *se, *paired_se;
+ unsigned int i, nr_elems, nr_elements;
+ unsigned int nr_zones_per_element, nrz = 0;
+ sector_t capacity_per_element;
+
+ /*
+ * Calculate the number of storage elements we are going to emulate.
+ * To achieve a somewhat realistic emulation, we want at least 2 storage
+ * elements, and no more than 32, targeting at least 32 zones per
+ * element.
+ */
+ if (zlo->nr_zones <= 64)
+ nr_elements = 2;
+ else
+ nr_elements =
+ min(ZLOOP_MAX_STOR_ELEMENTS,
+ zlo->nr_zones / ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS);
+ nr_zones_per_element = zlo->nr_zones / nr_elements;
+ capacity_per_element =
+ (sector_t)nr_zones_per_element << zlo->zone_shift;
+
+ /*
+ * If we are emulating pairs of read and write storage elements, we need
+ * double the number of storage element descriptors.
+ */
+ nr_elems = nr_elements;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ nr_elems *= 2;
+ zlo->elements = kzalloc_objs(struct blk_storage_element, nr_elems);
+ if (!zlo->elements)
+ return -ENOMEM;
+
+ for (i = 0, se = zlo->elements; i < nr_elements; i++, se++) {
+ se->id = i + 1;
+ if (nrz + nr_zones_per_element > zlo->nr_zones)
+ se->nr_zones = zlo->nr_zones - nrz;
+ else
+ se->nr_zones = nr_zones_per_element;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ se->type = BLK_SE_TYPE_WRITE;
+ else
+ se->type = BLK_SE_TYPE_RDWR;
+
+ se->status = BLK_SE_STS_OK;
+
+ nrz += nr_zones_per_element;
+ }
+
+ /*
+ * If we are emulating pairs of read and write storage elements,
+ * initialize the read elements paired with the write elements we just
+ * initialized.
+ */
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {
+ for (i = 0; i < nr_elements; i++, se++) {
+ se->id = nr_elements + i + 1;
+ paired_se = &zlo->elements[i];
+ se->paired_id = paired_se->id;
+ paired_se->paired_id = se->id;
+ se->nr_zones = paired_se->nr_zones;
+ se->type = BLK_SE_TYPE_READ;
+ se->status = BLK_SE_STS_OK;
+ }
+ }
+
+ /*
+ * Make sure we do not allow removing all storage elements as that does
+ * not make any sense. This is consistent with the device advertized
+ * limit of SCSI and ATA devices supporting the storage element
+ * depopulation feature.
+ */
+ zlo->nr_elements = nr_elems;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ zlo->max_nr_removed_elements = zlo->nr_elements - 2;
+ else
+ zlo->max_nr_removed_elements = zlo->nr_elements - 1;
+
+ return 0;
+}
+
static bool zloop_dev_exists(struct zloop_device *zlo)
{
struct file *cnv, *seq;
@@ -1220,6 +1641,10 @@ static int zloop_ctl_add(struct zloop_options *opts)
WRITE_ONCE(zlo->state, Zlo_creating);
spin_lock_init(&zlo->open_zones_lock);
INIT_LIST_HEAD(&zlo->open_zones_lru_list);
+ mutex_init(&zlo->stor_elements_lock);
+ INIT_DELAYED_WORK(&zlo->remove_element_work, zloop_remove_element_work);
+ INIT_DELAYED_WORK(&zlo->restore_elements_work,
+ zloop_restore_elements_work);
ret = mutex_lock_killable(&zloop_ctl_mutex);
if (ret)
@@ -1253,12 +1678,19 @@ static int zloop_ctl_add(struct zloop_options *opts)
if (zlo->zone_append)
zlo->ordered_zone_append = opts->ordered_zone_append;
zlo->discard_write_cache = opts->discard_write_cache;
+ zlo->stor_elements = opts->stor_elements;
+
+ if (zlo->stor_elements != ZLOOP_STOR_ELEMENTS_NONE) {
+ ret = zloop_create_storage_elements(zlo);
+ if (ret)
+ goto out_free_idr;
+ }
zlo->workqueue = alloc_workqueue("zloop%d", WQ_UNBOUND | WQ_FREEZABLE,
opts->nr_queues * opts->queue_depth, zlo->id);
if (!zlo->workqueue) {
ret = -ENOMEM;
- goto out_free_idr;
+ goto out_destroy_storage_elements;
}
if (opts->base_dir)
@@ -1344,6 +1776,10 @@ static int zloop_ctl_add(struct zloop_options *opts)
zlo->id, zlo->nr_zones,
((sector_t)zlo->zone_size << SECTOR_SHIFT) >> 20,
zlo->block_size);
+ if (zlo->nr_elements)
+ pr_info("zloop%d: %d storage elements\n",
+ zlo->id, zlo->nr_elements);
+
pr_info("zloop%d: using %s%s zone append\n",
zlo->id,
zlo->ordered_zone_append ? "ordered " : "",
@@ -1367,6 +1803,8 @@ static int zloop_ctl_add(struct zloop_options *opts)
kfree(zlo->base_dir);
out_destroy_workqueue:
destroy_workqueue(zlo->workqueue);
+out_destroy_storage_elements:
+ kfree(zlo->elements);
out_free_idr:
mutex_lock(&zloop_ctl_mutex);
idr_remove(&zloop_index_idr, zlo->id);
@@ -1485,6 +1923,7 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
opts->buffered_io = ZLOOP_DEF_BUFFERED_IO;
opts->zone_append = ZLOOP_DEF_ZONE_APPEND;
opts->ordered_zone_append = ZLOOP_DEF_ORDERED_ZONE_APPEND;
+ opts->stor_elements = ZLOOP_DEF_STOR_ELEMENTS;
if (!buf)
return 0;
@@ -1619,6 +2058,23 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
case ZLOOP_OPT_DISCARD_WRITE_CACHE:
opts->discard_write_cache = true;
break;
+ case ZLOOP_OPT_STOR_ELEMENTS:
+ if (match_uint(args, &token)) {
+ ret = -EINVAL;
+ goto out;
+ }
+ switch (token) {
+ case ZLOOP_STOR_ELEMENTS_NONE:
+ case ZLOOP_STOR_ELEMENTS_RDWR:
+ case ZLOOP_STOR_ELEMENTS_PAIRS:
+ break;
+ default:
+ pr_err("Invalid zone_append value\n");
+ ret = -EINVAL;
+ goto out;
+ }
+ opts->stor_elements = token;
+ break;
case ZLOOP_OPT_ERR:
default:
pr_warn("unknown parameter or missing value '%s'\n", p);
--
2.55.0
^ permalink raw reply related [flat|nested] 26+ messages in thread
* [PATCH 5/7] zloop: add degrade_element control command
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
` (3 preceding siblings ...)
2026-10-05 9:46 ` [PATCH 4/7] zloop: add storage element emulation Damien Le Moal
@ 2026-10-05 9:46 ` Damien Le Moal
2026-10-05 9:58 ` sashiko-bot
2026-10-05 11:17 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation Damien Le Moal
` (2 subsequent siblings)
7 siblings, 2 replies; 26+ messages in thread
From: Damien Le Moal @ 2026-10-05 9:46 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Allow users to mark storage elements of a zloop device as degraded using
the new "degrade_element" control command. The element to degrade is
indicated using the element_id option. Example:
echo "degrade_element id=0,element_id=2" > /dev/zloop-control
If the element ID identifies an all access storage element, read and write
operations targeting a zone served by the degraded lement are failed.
For a partial access storage element, read or write operations are failed
depending on the storage element type.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
drivers/block/zloop.c | 112 ++++++++++++++++++++++++++++++++++++++++--
1 file changed, 107 insertions(+), 5 deletions(-)
diff --git a/drivers/block/zloop.c b/drivers/block/zloop.c
index 7e6b5cc8017d..1d51060b832b 100644
--- a/drivers/block/zloop.c
+++ b/drivers/block/zloop.c
@@ -38,6 +38,7 @@ enum {
ZLOOP_OPT_DISCARD_WRITE_CACHE = (1 << 11),
ZLOOP_OPT_MAX_OPEN_ZONES = (1 << 12),
ZLOOP_OPT_STOR_ELEMENTS = (1 << 13),
+ ZLOOP_OPT_ELEMENT_ID = (1 << 14),
};
static const match_table_t zloop_opt_tokens = {
@@ -55,6 +56,7 @@ static const match_table_t zloop_opt_tokens = {
{ ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
{ ZLOOP_OPT_MAX_OPEN_ZONES, "max_open_zones=%u" },
{ ZLOOP_OPT_STOR_ELEMENTS, "stor_elements=%u" },
+ { ZLOOP_OPT_ELEMENT_ID, "element_id=%u" },
{ ZLOOP_OPT_ERR, NULL }
};
@@ -81,6 +83,7 @@ enum zloop_stor_elements {
#define ZLOOP_DEF_ZONE_APPEND true
#define ZLOOP_DEF_ORDERED_ZONE_APPEND false
#define ZLOOP_DEF_STOR_ELEMENTS ZLOOP_STOR_ELEMENTS_NONE
+#define ZLOOP_DEF_ELEMENT_ID 0
/* Arbitrary limit on the zone size (16GB). */
#define ZLOOP_MAX_ZONE_SIZE_MB 16384
@@ -101,6 +104,7 @@ struct zloop_options {
bool ordered_zone_append;
bool discard_write_cache;
enum zloop_stor_elements stor_elements;
+ unsigned int element_id;
};
/*
@@ -201,6 +205,29 @@ static unsigned int rq_zone_no(struct request *rq)
return blk_rq_pos(rq) >> zlo->zone_shift;
}
+static bool zloop_zone_healthy(struct zloop_device *zlo,
+ struct zloop_zone *zone, bool write)
+{
+ struct blk_storage_element *se;
+
+ if (!zlo->nr_elements)
+ return true;
+
+ /* Check the zone condition first. */
+ if (zone->cond == BLK_ZONE_COND_OFFLINE)
+ return false;
+ if (zone->cond == BLK_ZONE_COND_READONLY && write)
+ return false;
+
+ /* Check the health state of the storage element serving the zone. */
+ if (write)
+ se = &zlo->elements[zone->wr_se_id - 1];
+ else
+ se = &zlo->elements[zone->rd_se_id - 1];
+
+ return se->status == BLK_SE_STS_DEGRADED;
+}
+
/*
* Open an already open zone. This is mostly a no-op, except for the imp open ->
* exp open condition change that may happen. We also move a zone at the tail of
@@ -733,9 +760,7 @@ static void zloop_rw(struct zloop_cmd *cmd)
zone->start + zlo->zone_size))
goto out;
- if (zone->cond == BLK_ZONE_COND_OFFLINE)
- goto out;
- if (zone->cond == BLK_ZONE_COND_READONLY && is_write)
+ if (!zloop_zone_healthy(zlo, zone, is_write))
goto out;
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
@@ -1095,6 +1120,31 @@ static int zloop_report_elements(struct gendisk *disk,
return 0;
}
+static int zloop_degrade_element(struct zloop_device *zlo,
+ unsigned int element_id)
+{
+ struct blk_storage_element *se;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ if (!element_id || element_id > zlo->nr_elements)
+ return -EINVAL;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ se = &zlo->elements[element_id - 1];
+ if (se->status == BLK_SE_STS_OK)
+ se->status = BLK_SE_STS_DEGRADED;
+ else
+ ret = -EINVAL;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
static void zloop_remove_element_work(struct work_struct *work)
{
struct zloop_device *zlo = container_of(work, struct zloop_device,
@@ -1904,6 +1954,39 @@ static int zloop_ctl_remove(struct zloop_options *opts)
return 0;
}
+static int zloop_ctl_degrade_element(struct zloop_options *opts)
+{
+ struct zloop_device *zlo;
+ int ret = 0;
+
+ if (!(opts->mask & ZLOOP_OPT_ID)) {
+ pr_err("No ID specified for degrade_element\n");
+ return -EINVAL;
+ }
+
+ if (opts->mask & ~(ZLOOP_OPT_ID | ZLOOP_OPT_ELEMENT_ID)) {
+ pr_err("Invalid option specified for degrade_element\n");
+ return -EINVAL;
+ }
+
+ mutex_lock(&zloop_ctl_mutex);
+ zlo = idr_find(&zloop_index_idr, opts->id);
+ if (!zlo || zlo->state == Zlo_creating)
+ ret = -ENODEV;
+ else if (zlo->state == Zlo_deleting)
+ ret = -EINVAL;
+ mutex_unlock(&zloop_ctl_mutex);
+ if (ret)
+ return ret;
+
+ ret = zloop_degrade_element(zlo, opts->element_id);
+ if (!ret)
+ pr_info("Degraded element %u of device %u\n",
+ opts->id, opts->element_id);
+
+ return ret;
+}
+
static int zloop_parse_options(struct zloop_options *opts, const char *buf)
{
substring_t args[MAX_OPT_ARGS];
@@ -1924,6 +2007,7 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
opts->zone_append = ZLOOP_DEF_ZONE_APPEND;
opts->ordered_zone_append = ZLOOP_DEF_ORDERED_ZONE_APPEND;
opts->stor_elements = ZLOOP_DEF_STOR_ELEMENTS;
+ opts->element_id = ZLOOP_DEF_ELEMENT_ID;
if (!buf)
return 0;
@@ -2075,6 +2159,13 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
}
opts->stor_elements = token;
break;
+ case ZLOOP_OPT_ELEMENT_ID:
+ if (match_uint(args, &token)) {
+ ret = -EINVAL;
+ goto out;
+ }
+ opts->element_id = token;
+ break;
case ZLOOP_OPT_ERR:
default:
pr_warn("unknown parameter or missing value '%s'\n", p);
@@ -2103,14 +2194,16 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
enum {
ZLOOP_CTL_ADD,
ZLOOP_CTL_REMOVE,
+ ZLOOP_CTL_DEGRADE_ELEMENT,
};
static struct zloop_ctl_op {
int code;
const char *name;
} zloop_ctl_ops[] = {
- { ZLOOP_CTL_ADD, "add" },
- { ZLOOP_CTL_REMOVE, "remove" },
+ { ZLOOP_CTL_ADD, "add" },
+ { ZLOOP_CTL_REMOVE, "remove" },
+ { ZLOOP_CTL_DEGRADE_ELEMENT, "degrade_element" },
{ -1, NULL },
};
@@ -2158,6 +2251,9 @@ static ssize_t zloop_ctl_write(struct file *file, const char __user *ubuf,
case ZLOOP_CTL_REMOVE:
ret = zloop_ctl_remove(&opts);
break;
+ case ZLOOP_CTL_DEGRADE_ELEMENT:
+ ret = zloop_ctl_degrade_element(&opts);
+ break;
default:
pr_err("Invalid operation\n");
ret = -EINVAL;
@@ -2181,6 +2277,8 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)
tok = &zloop_opt_tokens[i];
if (!tok->pattern)
break;
+ if (tok->token == ZLOOP_OPT_ELEMENT_ID)
+ continue;
if (i)
seq_putc(seq_file, ',');
seq_puts(seq_file, tok->pattern);
@@ -2191,6 +2289,10 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)
seq_puts(seq_file, zloop_ctl_ops[1].name);
seq_puts(seq_file, " id=%d\n");
+ /* Degrade element operation */
+ seq_puts(seq_file, zloop_ctl_ops[2].name);
+ seq_puts(seq_file, " id=%d,element_id=%d\n");
+
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 26+ messages in thread
* [PATCH 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
` (4 preceding siblings ...)
2026-10-05 9:46 ` [PATCH 5/7] zloop: add degrade_element control command Damien Le Moal
@ 2026-10-05 9:46 ` Damien Le Moal
2026-10-05 11:19 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
2026-10-05 11:13 ` [PATCH 0/7] Add support for storage element depopulation Hannes Reinecke
7 siblings, 1 reply; 26+ messages in thread
From: Damien Le Moal @ 2026-10-05 9:46 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
sd_zbc_revalidate_zones() skips revalidating the zones of a ZBC device if
the zone size and total number of zones of the disk has not changed. This
is to avoid a call to the rather slow blk_revalidate_disk_zones().
However, for ZBC devices that support data preserving head depopulation
(REMOVE ELEMENT AND MODIFY ZONES command), a disk capacity and number of
zones does not change after a head is depopulated but the condition of
zones changes as the zones served by the head that was depopulated become
either read-only or offline. In this case, not calling
blk_revalidate_disk_zones() prevents the block layer from taking
appropriate actions on the zone write plugs of the disk for the zones that
became read-only or offline.
Avoid any issue with the block layer view of the zone conditions by not
skipping the call to blk_revalidate_disk_zones() for disks that support
the REMOVE ELEMENT AND MODIFY ZONES command. This check is done from
sd_zbc_read_zones() using the helper function sd_zbc_check_modify_zones().
The new scsi disk flag modify_zones_supported is defined to remember the
result of this check.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
drivers/scsi/sd.h | 1 +
drivers/scsi/sd_zbc.c | 26 +++++++++++++++++++++++++-
2 files changed, 26 insertions(+), 1 deletion(-)
diff --git a/drivers/scsi/sd.h b/drivers/scsi/sd.h
index 574af8243016..6a72371fee78 100644
--- a/drivers/scsi/sd.h
+++ b/drivers/scsi/sd.h
@@ -156,6 +156,7 @@ struct scsi_disk {
unsigned ignore_medium_access_errors : 1;
unsigned rscs : 1; /* reduced stream control support */
unsigned use_atomic_write_boundary : 1;
+ unsigned modify_zones_supported : 1;
};
#define to_scsi_disk(obj) container_of(obj, struct scsi_disk, disk_dev)
diff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c
index 456beaf2e769..2c77f878ab3f 100644
--- a/drivers/scsi/sd_zbc.c
+++ b/drivers/scsi/sd_zbc.c
@@ -516,6 +516,21 @@ static int sd_zbc_check_capacity(struct scsi_disk *sdkp, unsigned char *buf,
return 0;
}
+/*
+ * sd_zbc_check_modify_zones - Check if the device supports depopulation
+ * @sdkp: Target disk
+ * @buf: command buffer
+ *
+ * Check if the device supports the REMOVE ELEMENT AND MODIFY ZONES command.
+ */
+static inline bool sd_zbc_check_modify_zones(struct scsi_disk *sdkp,
+ unsigned char *buf)
+{
+ return scsi_report_opcode(sdkp->device, buf, SD_BUF_SIZE,
+ SERVICE_ACTION_IN_16,
+ SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES) == 1;
+}
+
static void sd_zbc_print_zones(struct scsi_disk *sdkp)
{
if (sdkp->device->type != TYPE_ZBC || !sdkp->capacity)
@@ -554,9 +569,15 @@ int sd_zbc_revalidate_zones(struct scsi_disk *sdkp)
if (!blk_queue_is_zoned(q))
return 0;
+ /*
+ * If the zone size and number of zones has not changed, and the disk
+ * does not support depopulating heads, skip the rather slow call to
+ * blk_revalidate_disk_zones().
+ */
if (sdkp->zone_info.zone_blocks == zone_blocks &&
sdkp->zone_info.nr_zones == nr_zones &&
- disk->nr_zones == nr_zones)
+ disk->nr_zones == nr_zones &&
+ !sdkp->modify_zones_supported)
return 0;
sdkp->zone_info.zone_blocks = zone_blocks;
@@ -620,6 +641,9 @@ int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim,
if (ret != 0)
goto err;
+ /* Check if REMOVE ELEMENT AND MODIFY ZONES is supported. */
+ sdkp->modify_zones_supported = sd_zbc_check_modify_zones(sdkp, buf);
+
nr_zones = round_up(sdkp->capacity, zone_blocks) >> ilog2(zone_blocks);
if (nr_zones > INT_MAX) {
sd_printk(KERN_ERR, sdkp, "Too many zones (%llu)\n",
--
2.55.0
^ permalink raw reply related [flat|nested] 26+ messages in thread
* [PATCH 7/7] scsi: sd_zbc: define storage element management operations
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
` (5 preceding siblings ...)
2026-10-05 9:46 ` [PATCH 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation Damien Le Moal
@ 2026-10-05 9:46 ` Damien Le Moal
2026-10-05 9:59 ` sashiko-bot
` (3 more replies)
2026-10-05 11:13 ` [PATCH 0/7] Add support for storage element depopulation Hannes Reinecke
7 siblings, 4 replies; 26+ messages in thread
From: Damien Le Moal @ 2026-10-05 9:46 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Define the storage element management operations using struct
blk_storage_elements_ops. These operations are valid only on SMR disks
supporting the storage element depopulation feature.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
drivers/scsi/sd.c | 1 +
drivers/scsi/sd.h | 4 +
drivers/scsi/sd_zbc.c | 210 ++++++++++++++++++++++++++++++++++++++++++
3 files changed, 215 insertions(+)
diff --git a/drivers/scsi/sd.c b/drivers/scsi/sd.c
index a1b21ea14e54..799d2093d667 100644
--- a/drivers/scsi/sd.c
+++ b/drivers/scsi/sd.c
@@ -3937,6 +3937,7 @@ static const struct block_device_operations sd_fops = {
.get_unique_id = sd_get_unique_id,
.free_disk = scsi_disk_free_disk,
.pr_ops = &sd_pr_ops,
+ .se_ops = &sd_zbc_se_ops,
};
/**
diff --git a/drivers/scsi/sd.h b/drivers/scsi/sd.h
index 6a72371fee78..16a004a1e6f1 100644
--- a/drivers/scsi/sd.h
+++ b/drivers/scsi/sd.h
@@ -243,6 +243,8 @@ unsigned int sd_zbc_complete(struct scsi_cmnd *cmd, unsigned int good_bytes,
int sd_zbc_report_zones(struct gendisk *disk, sector_t sector,
unsigned int nr_zones, struct blk_report_zones_args *args);
+extern const struct blk_storage_elements_ops sd_zbc_se_ops;
+
#else /* CONFIG_BLK_DEV_ZONED */
static inline int sd_zbc_read_zones(struct scsi_disk *sdkp,
@@ -271,6 +273,8 @@ static inline unsigned int sd_zbc_complete(struct scsi_cmnd *cmd,
#define sd_zbc_report_zones NULL
+#define sd_zbc_se_ops NULL
+
#endif /* CONFIG_BLK_DEV_ZONED */
void sd_print_sense_hdr(struct scsi_disk *sdkp, struct scsi_sense_hdr *sshdr);
diff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c
index 2c77f878ab3f..586e0ab508ed 100644
--- a/drivers/scsi/sd_zbc.c
+++ b/drivers/scsi/sd_zbc.c
@@ -548,6 +548,216 @@ static void sd_zbc_print_zones(struct scsi_disk *sdkp)
sdkp->zone_info.zone_blocks);
}
+#define SD_ZBC_STORAGE_ELEMENTS_BUF_SIZE 4096
+
+static void sd_zbc_parse_storage_element(struct scsi_disk *sdkp, u8 *desc,
+ struct blk_storage_element *element)
+{
+ struct scsi_device *sdp = sdkp->device;
+ u64 capacity;
+
+ element->id = get_unaligned_be32(&desc[4]);
+
+ switch (desc[14]) {
+ case SCSI_PHYS_ELEM_TYPE_ALL_ACCESS_STORAGE:
+ element->paired_id = 0;
+ element->type = BLK_SE_TYPE_RDWR;
+ capacity = get_unaligned_be64(&desc[16]);
+ element->nr_zones = logical_to_sectors(sdp, capacity) >>
+ ilog2(sd_zbc_zone_sectors(sdkp));
+ break;
+ case SCSI_PHYS_ELEM_TYPE_FRAC_ACCESS_STORAGE:
+ element->paired_id = get_unaligned_be32(&desc[16]);
+ if (desc[20] & 0x01)
+ element->type = BLK_SE_TYPE_READ;
+ else
+ element->type = BLK_SE_TYPE_WRITE;
+ element->nr_zones = get_unaligned_be64(&desc[24]);
+ break;
+ default:
+ element->type = BLK_SE_TYPE_UNKNOWN;
+ }
+
+ switch (desc[15]) {
+ case SCSI_PHYS_ELEM_HEALTH_WITHIN_SPEC_LIMITS:
+ case SCSI_PHYS_ELEM_HEALTH_AT_SPEC_LIMITS:
+ element->status = BLK_SE_STS_OK;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_OUTSIDE_SPEC_LIMITS:
+ element->status = BLK_SE_STS_DEGRADED;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_ERR:
+ element->status = BLK_SE_STS_RESTORE_ERROR;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_IN_PROGRESS:
+ element->status = BLK_SE_STS_RESTORE_IN_PROGRESS;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_ERR:
+ element->status = BLK_SE_STS_REMOVE_ERROR;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_IN_PROGRESS:
+ element->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_OK:
+ element->status = BLK_SE_STS_REMOVED;
+ element->restore_allowed = desc[13] & 0x01;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_NOT_REPORTED:
+ default:
+ element->status = BLK_SE_STS_UNKOWN;
+ break;
+ }
+}
+
+static int sd_zbc_report_storage_elements(struct gendisk *disk,
+ unsigned int *nr_elements,
+ struct blk_storage_element *elements)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ unsigned int nr_descs;
+ int i, ret = 0, result;
+ u8 *desc, *buf;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ buf = kzalloc(SD_ZBC_STORAGE_ELEMENTS_BUF_SIZE, GFP_KERNEL);
+ if (!buf)
+ return -ENOMEM;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_GET_PHYSICAL_ELEMENT_STATUS;
+ put_unaligned_be32(SD_ZBC_STORAGE_ELEMENTS_BUF_SIZE, &cmd[10]);
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, buf,
+ SD_ZBC_STORAGE_ELEMENTS_BUF_SIZE,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "GET PHYSICAL ELEMENT STATUS failed\n");
+ sd_print_result(sdkp, "GET PHYSICAL ELEMENT STATUS", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ nr_descs = get_unaligned_be32(&buf[0]);
+ if (!nr_descs) {
+ sd_printk(KERN_ERR, sdkp,
+ "Invalid number of phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ if (!elements) {
+ *nr_elements = nr_descs;
+ goto free_buf;
+ }
+
+ nr_descs = get_unaligned_be32(&buf[4]);
+ if (!nr_descs) {
+ sd_printk(KERN_ERR, sdkp,
+ "Invalid number of reported phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ desc = &buf[32];
+ for (i = 0; i < min(*nr_elements, nr_descs); i++, desc += 32)
+ sd_zbc_parse_storage_element(sdkp, desc, &elements[i]);
+ *nr_elements = nr_descs;
+
+free_buf:
+ kfree(buf);
+
+ return ret;
+}
+
+static int sd_zbc_remove_storage_element(struct gendisk *disk,
+ unsigned int element_id)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ int result;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES;
+ put_unaligned_be32(element_id, &cmd[10]);
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "REMOVE ELEMENT AND MODIFY ZONES failed\n");
+ sd_print_result(sdkp,
+ "REMOVE ELEMENT AND MODIFY ZONES", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+static int sd_zbc_restore_storage_elements(struct gendisk *disk)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ int result;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_RESTORE_ELEMENTS_AND_REBUILD;
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "RESTORE ELEMENTS AND REBUILD failed\n");
+ sd_print_result(sdkp,
+ "RESTORE ELEMENTS AND REBUILD", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+const struct blk_storage_elements_ops sd_zbc_se_ops = {
+ .report_elements = sd_zbc_report_storage_elements,
+ .remove_element = sd_zbc_remove_storage_element,
+ .restore_elements = sd_zbc_restore_storage_elements,
+};
+
/*
* Call blk_revalidate_disk_zones() if any of the zoned disk properties have
* changed that make it necessary to call that function. Called by
--
2.55.0
^ permalink raw reply related [flat|nested] 26+ messages in thread
* Re: [PATCH 4/7] zloop: add storage element emulation
2026-10-05 9:46 ` [PATCH 4/7] zloop: add storage element emulation Damien Le Moal
@ 2026-10-05 9:58 ` sashiko-bot
2026-10-05 11:14 ` Hannes Reinecke
1 sibling, 0 replies; 26+ messages in thread
From: sashiko-bot @ 2026-10-05 9:58 UTC (permalink / raw)
To: Damien Le Moal; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
- [High] block/zloop: integer overflow in ordered zone append causes wp corruption on OFFLINE zones
- [Medium] block/zloop: failed storage element restore permanently bricks device
- [Medium] block/zloop: storage element emulation reports incorrect zone counts due to division truncation
- [Medium] block/zloop: lockless TOCTOU race allows I/O on OFFLINE and READONLY storage elements
- [Low] block/zloop: double counting of nr_removed_elements when removing READ elements
- [Low] block/zloop: misleading error message for invalid stor_elements option
--
commit 2ec54a3b571357c5ded6a5a4c9a7aea473cb335d
Author: Damien Le Moal <dlemoal@kernel.org>
zloop: add storage element emulation
This commit introduces storage element emulation to the zloop driver. It
allows for emulating SMR disks' storage element depopulation feature,
controlled by a new stor_elements option. The operations to manage storage
elements are specified via zloop_se_ops.
> diff --git a/drivers/block/zloop.c b/drivers/block/zloop.c
> index f0ca221524db6..7e6b5cc8017d2 100644
> --- a/drivers/block/zloop.c
> +++ b/drivers/block/zloop.c
> @@ -664,6 +733,11 @@ static void zloop_rw(struct zloop_cmd *cmd)
> zone->start + zlo->zone_size))
> goto out;
>
> + if (zone->cond == BLK_ZONE_COND_OFFLINE)
> + goto out;
> + if (zone->cond == BLK_ZONE_COND_READONLY && is_write)
> + goto out;
> +
[Severity: Medium]
Could this lockless check introduce a TOCTOU race window?
Concurrently, zloop_remove_element_work() can acquire zone->lock and change
the state to OFFLINE or READONLY.
If the work function runs between this lockless check and the dispatch of the
I/O, zloop_do_rw() could be executed on a disabled storage element, as it
bypasses taking zone->lock for reads and conventional zone writes.
[ ... ]
> @@ -877,7 +951,7 @@ static bool zloop_set_zone_append_sector(struct request *rq)
> rq->__sector = zone->wp;
> zone->wp += blk_rq_sectors(rq);
> if (zone->wp >= zone_end)
> - zloop_mark_full(zlo, zone);
> + zloop_set_zone_full(zlo, zone);
>
> spin_unlock(&zone->wp_lock);
[Severity: High]
Does this bounds check properly handle OFFLINE/READONLY zones?
When a zone becomes OFFLINE or READONLY, its zone->wp is set to ULLONG_MAX.
If an ordered ZONE_APPEND arrives for this zone, the bounds check in
zloop_set_zone_append_sector():
if (zone->cond == BLK_ZONE_COND_FULL ||
zone->wp + nr_sectors > zone_end) {
could evaluate to false because ULLONG_MAX + nr_sectors can overflow to a
small sector number. This would bypass the bounds check and update zone->wp
to the overflowed value, corrupting the write pointer state.
[ ... ]
> @@ -985,11 +1059,251 @@ static int zloop_report_zones(struct gendisk *disk, sector_t sector,
[ ... ]
> +static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)
> +{
[ ... ]
> + /*
> + * If the element to remove is a read element, the associated write
> + * element must also be removed.
> + */
> + if (se->type == BLK_SE_TYPE_READ) {
> + paired_se = &zlo->elements[se->paired_id - 1];
> + nr_remove = 2;
> + }
[Severity: Low]
Does this code double-count nr_removed_elements if the paired WRITE element
was already manually removed?
zloop_remove_element() allows targeting a paired WRITE element and overwriting
paired_se->status to BLK_SE_STS_REMOVE_IN_PROGRESS without validating if it
was already removed.
Then zloop_remove_element_work() unconditionally increments
zlo->nr_removed_elements for both elements:
zlo->nr_removed_elements++;
if (paired_se) {
paired_se->status = BLK_SE_STS_REMOVED;
zlo->nr_removed_elements++;
}
This could artificially inflate nr_removed_elements and prematurely trigger
the max_nr_removed_elements limit.
[ ... ]
> +static void zloop_restore_elements_work(struct work_struct *work)
> +{
[ ... ]
> + /* Restore all removed elements. */
> + for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
> + if (se->status != BLK_SE_STS_RESTORE_IN_PROGRESS)
> + continue;
> + if (!ret)
> + se->status = BLK_SE_STS_OK;
> + else
> + se->status = BLK_SE_STS_RESTORE_ERROR;
> + }
> +
> + zlo->nr_removed_elements = 0;
> + zlo->restore_in_progress = false;
> +
> + mutex_unlock(&zlo->stor_elements_lock);
> +}
[Severity: Medium]
Will this permanently brick the device if a restore fails?
If zloop_do_reset_zone() fails, elements are put into
BLK_SE_STS_RESTORE_ERROR, but the code still unconditionally clears
zlo->nr_removed_elements.
A subsequent call to zloop_restore_elements() checks:
if (!zlo->nr_removed_elements)
goto unlock;
This prevents retrying the restore operation. Even if the check is bypassed,
the loop only targets BLK_SE_STS_REMOVED elements, leaving RESTORE_ERROR
elements permanently stuck.
[ ... ]
> +static int zloop_create_storage_elements(struct zloop_device *zlo)
> +{
[ ... ]
> + nr_zones_per_element = zlo->nr_zones / nr_elements;
> + capacity_per_element =
> + (sector_t)nr_zones_per_element << zlo->zone_shift;
[ ... ]
> + for (i = 0, se = zlo->elements; i < nr_elements; i++, se++) {
> + se->id = i + 1;
> + if (nrz + nr_zones_per_element > zlo->nr_zones)
> + se->nr_zones = zlo->nr_zones - nrz;
> + else
> + se->nr_zones = nr_zones_per_element;
[Severity: Medium]
Does this code lose the remainder zones in its calculation?
Because nr_zones_per_element is the integer floor of
zlo->nr_zones / nr_elements, the condition
(nrz + nr_zones_per_element > zlo->nr_zones) will never evaluate to true.
Furthermore, elements are assigned to zones in a round-robin fashion (e.g.,
zone_no % nr_elems), meaning the number of zones per element should be
computed based on modulo arithmetic, not simple contiguous division. This
causes se->nr_zones to under-report the actual zones managed by the element.
[ ... ]
> @@ -1619,6 +2058,23 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
[ ... ]
> + case ZLOOP_OPT_STOR_ELEMENTS:
> + if (match_uint(args, &token)) {
> + ret = -EINVAL;
> + goto out;
> + }
> + switch (token) {
> + case ZLOOP_STOR_ELEMENTS_NONE:
> + case ZLOOP_STOR_ELEMENTS_RDWR:
> + case ZLOOP_STOR_ELEMENTS_PAIRS:
> + break;
> + default:
> + pr_err("Invalid zone_append value\n");
> + ret = -EINVAL;
> + goto out;
> + }
> + opts->stor_elements = token;
> + break;
[Severity: Low]
This isn't a bug, but the error message here says "Invalid zone_append value\n"
for the stor_elements option. This looks like a copy-paste error.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261005094632.580753-1-dlemoal@kernel.org?part=4
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 5/7] zloop: add degrade_element control command
2026-10-05 9:46 ` [PATCH 5/7] zloop: add degrade_element control command Damien Le Moal
@ 2026-10-05 9:58 ` sashiko-bot
2026-10-05 11:17 ` Hannes Reinecke
1 sibling, 0 replies; 26+ messages in thread
From: sashiko-bot @ 2026-10-05 9:58 UTC (permalink / raw)
To: Damien Le Moal; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [High] block: logic inversion in zloop_zone_healthy() aborts I/O on healthy zones
- [High] block: Use-After-Free in zloop_ctl_degrade_element() due to missing refcount
- [Medium] block: data race on se->status without READ_ONCE/WRITE_ONCE wrappers
--
commit 003661eebb0c9d5db84ec05ab948b11a2ca35348
Author: Damien Le Moal <dlemoal@kernel.org>
zloop: add degrade_element control command
This patch adds a "degrade_element" control command to zloop devices, allowing
users to mark specific storage elements as degraded. Depending on the element's
access type, read or write operations targeting its zones will be failed.
> diff --git a/drivers/block/zloop.c b/drivers/block/zloop.c
> index 7e6b5cc8017d2..1d51060b832bc 100644
> --- a/drivers/block/zloop.c
> +++ b/drivers/block/zloop.c
[ ... ]
> @@ -201,6 +205,29 @@ static unsigned int rq_zone_no(struct request *rq)
> return blk_rq_pos(rq) >> zlo->zone_shift;
> }
>
> +static bool zloop_zone_healthy(struct zloop_device *zlo,
> + struct zloop_zone *zone, bool write)
> +{
> + struct blk_storage_element *se;
> +
> + if (!zlo->nr_elements)
> + return true;
> +
> + /* Check the zone condition first. */
> + if (zone->cond == BLK_ZONE_COND_OFFLINE)
> + return false;
> + if (zone->cond == BLK_ZONE_COND_READONLY && write)
> + return false;
> +
> + /* Check the health state of the storage element serving the zone. */
> + if (write)
> + se = &zlo->elements[zone->wr_se_id - 1];
> + else
> + se = &zlo->elements[zone->rd_se_id - 1];
> +
> + return se->status == BLK_SE_STS_DEGRADED;
[Severity: Medium]
Since this status field is read locklessly in the fast I/O path while it can
be concurrently updated by zloop_degrade_element(), should this use
READ_ONCE() to prevent compiler optimizations from tearing the read or
caching the variable?
[Severity: High]
Does this evaluate the health condition inversely? It returns true when the
element is degraded and false when it is healthy.
> +}
> +
[ ... ]
> @@ -733,9 +760,7 @@ static void zloop_rw(struct zloop_cmd *cmd)
> zone->start + zlo->zone_size))
> goto out;
>
> - if (zone->cond == BLK_ZONE_COND_OFFLINE)
> - goto out;
> - if (zone->cond == BLK_ZONE_COND_READONLY && is_write)
> + if (!zloop_zone_healthy(zlo, zone, is_write))
> goto out;
[Severity: High]
Because zloop_zone_healthy() returns false for healthy zones, does this
inverted check systematically abort valid I/Os to operating zones?
>
> if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
[ ... ]
> @@ -1095,6 +1120,31 @@ static int zloop_report_elements(struct gendisk *disk,
> return 0;
> }
>
> +static int zloop_degrade_element(struct zloop_device *zlo,
> + unsigned int element_id)
> +{
> + struct blk_storage_element *se;
> + int ret = 0;
> +
> + if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
> + return -EOPNOTSUPP;
> +
> + if (!element_id || element_id > zlo->nr_elements)
> + return -EINVAL;
> +
> + mutex_lock(&zlo->stor_elements_lock);
> +
> + se = &zlo->elements[element_id - 1];
> + if (se->status == BLK_SE_STS_OK)
> + se->status = BLK_SE_STS_DEGRADED;
[Severity: Medium]
To match the lockless read in zloop_zone_healthy(), should this update be
wrapped in WRITE_ONCE() to ensure the memory access ordering?
> + else
> + ret = -EINVAL;
> +
> + mutex_unlock(&zlo->stor_elements_lock);
> +
> + return ret;
> +}
> +
[ ... ]
> @@ -1904,6 +1954,39 @@ static int zloop_ctl_remove(struct zloop_options *opts)
> return 0;
> }
>
> +static int zloop_ctl_degrade_element(struct zloop_options *opts)
> +{
> + struct zloop_device *zlo;
> + int ret = 0;
> +
> + if (!(opts->mask & ZLOOP_OPT_ID)) {
> + pr_err("No ID specified for degrade_element\n");
> + return -EINVAL;
> + }
> +
> + if (opts->mask & ~(ZLOOP_OPT_ID | ZLOOP_OPT_ELEMENT_ID)) {
> + pr_err("Invalid option specified for degrade_element\n");
> + return -EINVAL;
> + }
> +
> + mutex_lock(&zloop_ctl_mutex);
> + zlo = idr_find(&zloop_index_idr, opts->id);
> + if (!zlo || zlo->state == Zlo_creating)
> + ret = -ENODEV;
> + else if (zlo->state == Zlo_deleting)
> + ret = -EINVAL;
> + mutex_unlock(&zloop_ctl_mutex);
> + if (ret)
> + return ret;
> +
> + ret = zloop_degrade_element(zlo, opts->element_id);
[Severity: High]
Is zlo exposed to a concurrent free here? The mutex is unlocked without holding
a reference to zlo->disk via get_disk().
This could allow a concurrent zloop_ctl_remove() operation to remove the IDR
entry, unlock the same mutex, drop the final disk reference, and free the
zlo structure right before it is used in zloop_degrade_element().
> + if (!ret)
> + pr_info("Degraded element %u of device %u\n",
> + opts->id, opts->element_id);
> +
> + return ret;
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261005094632.580753-1-dlemoal@kernel.org?part=5
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 7/7] scsi: sd_zbc: define storage element management operations
2026-10-05 9:46 ` [PATCH 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
@ 2026-10-05 9:59 ` sashiko-bot
2026-10-05 11:48 ` Hannes Reinecke
` (2 subsequent siblings)
3 siblings, 0 replies; 26+ messages in thread
From: sashiko-bot @ 2026-10-05 9:59 UTC (permalink / raw)
To: Damien Le Moal; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [High] scsi: sd_zbc: heap OOB read from buf via unvalidated nr_descs
- [High] scsi: sd_zbc: heap OOB leak caused by overwriting caller capacity with nr_descs
- [Medium] scsi: sd_zbc: hardcoded 4KB buffer permanently limits storage elements to 127
- [Medium] scsi: sd_zbc: undefined shift via ilog2(0) when zone_blocks is uninitialized
- [Low] build: invalid &NULL assignment for se_ops in non-zoned configs
--
commit dc87627d74d92a84dc8259e4c75a364fd39dae85
Author: Damien Le Moal <dlemoal@kernel.org>
scsi: sd_zbc: define storage element management operations
This commit defines the storage element management operations using
struct blk_storage_elements_ops. These operations are intended for SMR disks
that support the storage element depopulation feature.
> diff --git a/drivers/scsi/sd.c b/drivers/scsi/sd.c
> index b0d2542fd3172..0a252739f5945 100644
> --- a/drivers/scsi/sd.c
> +++ b/drivers/scsi/sd.c
> @@ -3939,6 +3939,7 @@ static const struct block_device_operations sd_fops = {
> .get_unique_id = sd_get_unique_id,
> .free_disk = scsi_disk_free_disk,
> .pr_ops = &sd_pr_ops,
> + .se_ops = &sd_zbc_se_ops,
> };
[Severity: Low]
Does this cause a build failure when CONFIG_BLK_DEV_ZONED is disabled? Since
sd_zbc_se_ops is defined as NULL in sd.h for non-zoned configs, this expands
to &NULL which is syntactically invalid.
[ ... ]
> diff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c
> index 1bf5f7ea0d9c5..56e8456cd0eca 100644
> --- a/drivers/scsi/sd_zbc.c
> +++ b/drivers/scsi/sd_zbc.c
> @@ -548,6 +548,216 @@ static void sd_zbc_print_zones(struct scsi_disk *sdkp)
[ ... ]
> +static void sd_zbc_parse_storage_element(struct scsi_disk *sdkp, u8 *desc,
> + struct blk_storage_element *element)
> +{
> + struct scsi_device *sdp = sdkp->device;
> + u64 capacity;
> +
> + element->id = get_unaligned_be32(&desc[4]);
> +
> + switch (desc[14]) {
> + case SCSI_PHYS_ELEM_TYPE_ALL_ACCESS_STORAGE:
> + element->paired_id = 0;
> + element->type = BLK_SE_TYPE_RDWR;
> + capacity = get_unaligned_be64(&desc[16]);
> + element->nr_zones = logical_to_sectors(sdp, capacity) >>
> + ilog2(sd_zbc_zone_sectors(sdkp));
[Severity: Medium]
Does this risk an undefined shift on architectures where ilog2(0) faults? If
this function is called on a device where modify_zones_supported is true but
zone_blocks is zero (such as following a zoned revalidation failure),
sd_zbc_zone_sectors(sdkp) would return 0.
[ ... ]
> +static int sd_zbc_report_storage_elements(struct gendisk *disk,
> + unsigned int *nr_elements,
> + struct blk_storage_element *elements)
> +{
[ ... ]
> + memset(cmd, 0, 16);
> + cmd[0] = SERVICE_ACTION_IN_16;
> + cmd[1] = SAI_GET_PHYSICAL_ELEMENT_STATUS;
> + put_unaligned_be32(SD_ZBC_STORAGE_ELEMENTS_BUF_SIZE, &cmd[10]);
[Severity: Medium]
Does this hardcoded 4KB allocation length permanently limit the number of
reported storage elements to 127? There appears to be no mechanism to query
elements iteratively, which limits functionality for devices with a larger
number of elements.
[ ... ]
> + nr_descs = get_unaligned_be32(&buf[4]);
> + if (!nr_descs) {
> + sd_printk(KERN_ERR, sdkp,
> + "Invalid number of reported phys element descriptors\n");
> + ret = -EIO;
> + goto free_buf;
> + }
> +
> + desc = &buf[32];
> + for (i = 0; i < min(*nr_elements, nr_descs); i++, desc += 32)
> + sd_zbc_parse_storage_element(sdkp, desc, &elements[i]);
[Severity: High]
Can this loop read out-of-bounds of the 4096-byte buf? The driver extracts
nr_descs directly from the device payload without validating it against the
buffer size limit. If a device returns a large nr_descs, this loop will
advance the desc pointer past the end of the allocated buf.
> + *nr_elements = nr_descs;
[Severity: High]
Can this unconditionally overwrite the caller-provided array capacity
(*nr_elements) with the device-reported count? If the caller provided a
smaller array capacity than nr_descs, returning the larger nr_descs can
trick the caller into copying uninitialized out-of-bounds heap memory to
userspace.
> +
> +free_buf:
> + kfree(buf);
> +
> + return ret;
> +}
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261005094632.580753-1-dlemoal@kernel.org?part=7
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 3/7] block: add storage element management ioctls
2026-10-05 9:46 ` [PATCH 3/7] block: add storage element management ioctls Damien Le Moal
@ 2026-10-05 10:00 ` sashiko-bot
2026-10-05 11:09 ` Hannes Reinecke
1 sibling, 0 replies; 26+ messages in thread
From: sashiko-bot @ 2026-10-05 10:00 UTC (permalink / raw)
To: Damien Le Moal; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
- [High] block: missing queue freeze during storage element reconfiguration
- [High] block: integer truncation in invalidate_inode_pages2_range() skips cache invalidation
- [High] block: heap out-of-bounds read and info leak in blkdev_report_storage_elements_ioctl()
- [High] block: loop-carried dependency in disk_wait_for_se_mgmt_completion() causes heap overflow
- [Medium] block: unbounded memory allocation in blkdev_report_storage_elements_ioctl()
- [Medium] block: silent OOM failure in disk_wait_for_se_mgmt_completion() bypasses wait
--
commit c30fd4a3def1740a027fbbe35b4c863893b60362
Author: Damien Le Moal <dlemoal@kernel.org>
block: add storage element management ioctls
This commit introduces new ioctl commands for zoned block devices to allow
userspace management of storage elements. It provides interfaces for getting
the number of elements, reporting their status, and removing or restoring
degraded elements.
> diff --git a/block/blk-zoned.c b/block/blk-zoned.c
> index b713512f7d48f..62b9cf21e9adf 100644
> --- a/block/blk-zoned.c
> +++ b/block/blk-zoned.c
[ ... ]
> @@ -2794,6 +2795,69 @@ int bdev_report_storage_elements(struct block_device *bdev,
[ ... ]
> +static int blkdev_report_storage_elements_ioctl(struct block_device *bdev,
> + void __user *argp)
> +{
> + struct blk_storage_elements_report rep;
> + struct blk_storage_element *elements;
> + unsigned int nr_elements;
> + unsigned long retc;
> + int ret;
> +
> + if (!argp)
> + return -EINVAL;
> +
> + if (copy_from_user(&rep, argp,
> + sizeof(struct blk_storage_elements_report)))
> + return -EFAULT;
> +
> + if (!rep.nr_elements)
> + return -EINVAL;
> +
> + nr_elements = rep.nr_elements;
> + elements = kzalloc_objs(struct blk_storage_element, nr_elements);
> + if (!elements)
> + return -ENOMEM;
[Severity: Medium]
Is it safe to use kzalloc_objs() here with an unbounded user-provided count?
A maliciously crafted ioctl call with a huge rep.nr_elements will request an
allocation exceeding KMALLOC_MAX_SIZE. This will trigger a kernel
WARN_ON_ONCE stack trace and fail with -ENOMEM. Repeated calls by an
unprivileged user with read access can spam dmesg and degrade system
performance.
> +
> + ret = bdev_report_storage_elements(bdev, &nr_elements, elements);
> + if (ret)
> + goto free_elements;
> +
> + retc = copy_to_user(argp + sizeof(struct blk_storage_elements_report),
> + elements,
> + sizeof(struct blk_storage_element) * nr_elements);
[Severity: High]
Can this copy_to_user() read out of bounds and leak kernel heap memory?
If the hardware reports more elements than the user originally requested,
bdev_report_storage_elements() updates nr_elements to the larger count.
This larger nr_elements is then used to determine the copy byte length,
which copies past the end of the elements allocation and leaks uninitialized
adjacent kernel heap memory to userspace.
[ ... ]
> @@ -2847,6 +2911,22 @@ int bdev_remove_storage_element(struct block_device *bdev,
> }
> EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
>
> +static int blkdev_remove_storage_element_ioctl(struct block_device *bdev,
> + blk_mode_t mode, void __user *argp)
> +{
> + unsigned int element_id;
> +
> + if (!(mode & BLK_OPEN_WRITE))
> + return -EBADF;
> + if (bdev_read_only(bdev))
> + return -EPERM;
> +
> + if (get_user(element_id, (unsigned int __user *)argp))
> + return -EFAULT;
> +
> + return bdev_remove_storage_element(bdev, element_id);
> +}
[Severity: High]
Does bdev_remove_storage_element() fail to hold the queue freeze?
In bdev_remove_storage_element(), the queue is unfrozen prematurely:
memflags = blk_mq_freeze_queue(disk->queue);
blk_mq_unfreeze_queue(disk->queue, memflags);
The block layer subsystem guide requires maintaining a queue freeze across
reconfiguration or teardown. blk_revalidate_disk_zones() assumes the queue
is actively frozen when called from a reconfiguration context.
Since the ioctl does not implicitly serialize I/O, concurrent userspace I/O
can be submitted while the zone structure is being invalidated and rewritten.
This allows in-flight BIOs to read torn configuration state or cause
use-after-free.
[Severity: High]
Does bdev_remove_storage_element() pass bytes instead of a page index to
invalidate_inode_pages2_range()?
invalidate_inode_pages2_range(bdev->bd_mapping, 0,
get_capacity(disk) << SECTOR_SHIFT);
The end parameter expects a page index (pgoff_t). On 32-bit systems where
pgoff_t is a 32-bit unsigned long, passing a byte capacity for any disk
exceeding 4GB overflows the parameter.
For example, exactly 4GB evaluates to 0, invalidating only the 0th page.
This leaves the rest of the page cache entirely stale, causing silent data
corruption when userspace later reads from the modified block device.
[Severity: High]
Does disk_wait_for_se_mgmt_completion() have a loop-carried dependency that
causes a heap overflow?
When called by bdev_remove_storage_element(), the elements buffer is
allocated once before entering the polling loop. Inside the loop:
while (1) {
/*
* Check if we have storage elements being removed or restored.
*/
nr_se = nr_elements;
ret = (disk->fops->se_ops->report_elements)(disk,
&nr_elements, elements);
If the device dynamically increases the number of reported elements,
report_elements updates nr_elements to the higher value.
In the next loop iteration, this larger nr_elements is passed back into
report_elements, which will write the larger number of elements into the
originally sized elements heap buffer, causing an out-of-bounds heap
overwrite.
[Severity: Medium]
Does disk_wait_for_se_mgmt_completion() bypass the wait on OOM?
elements = kzalloc_objs(struct blk_storage_element, nr_elements);
if (!elements)
return;
If kzalloc_objs() fails, the function silently returns without completing its
waiting loop.
Under memory pressure, this allows the caller to immediately proceed and
invoke blk_revalidate_disk_zones() while the device is actively busy
mutating its zone configuration in the background. This leads to I/O timeouts,
torn reads of zone statuses, or hardware errors.
[ ... ]
> @@ -2886,3 +2966,37 @@ int bdev_restore_storage_elements(struct block_device *bdev)
> return blk_revalidate_disk_zones(disk);
> }
> EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
> +
> +static int blkdev_restore_storage_elements_ioctl(struct block_device *bdev,
> + blk_mode_t mode)
> +{
> + if (!(mode & BLK_OPEN_WRITE))
> + return -EBADF;
> + if (bdev_read_only(bdev))
> + return -EPERM;
> +
> + return bdev_restore_storage_elements(bdev);
> +}
[Severity: High]
Does bdev_restore_storage_elements() also fail to hold the queue freeze?
Similar to the removal path, the queue is unfrozen prematurely:
memflags = blk_mq_freeze_queue(disk->queue);
blk_mq_unfreeze_queue(disk->queue, memflags);
This opens a window for concurrent I/Os during restoration.
[Severity: High]
Does bdev_restore_storage_elements() also skip cache invalidation due to
integer truncation?
Similar to the removal path, it passes a byte capacity instead of a page
index:
invalidate_inode_pages2_range(bdev->bd_mapping, 0,
get_capacity(disk) << SECTOR_SHIFT);
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261005094632.580753-1-dlemoal@kernel.org?part=3
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 1/7] block: fail reads to offline zones early
2026-10-05 9:46 ` [PATCH 1/7] block: fail reads to offline zones early Damien Le Moal
@ 2026-10-05 10:01 ` sashiko-bot
2026-10-05 10:45 ` Hannes Reinecke
1 sibling, 0 replies; 26+ messages in thread
From: sashiko-bot @ 2026-10-05 10:01 UTC (permalink / raw)
To: Damien Le Moal; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] block: split bios bypass early offline zone read check
--
commit 4d9e6f59ea5e83a002020d10f8ee92430c6bfa6f
Author: Damien Le Moal <dlemoal@kernel.org>
block: fail reads to offline zones early
This commit fails read commands targeting offline zones of a zoned block
device early. It checks if the sector belongs to an offline zone in
submit_bio_noacct() and sets the BIO_QUIET flag before completing the bio
with an error.
> diff --git a/block/blk-core.c b/block/blk-core.c
> index 8836c7028568b..a72aa5c18096e 100644
> --- a/block/blk-core.c
> +++ b/block/blk-core.c
> @@ -867,6 +867,11 @@ void submit_bio_noacct(struct bio *bio)
>
> switch (bio_op(bio)) {
> case REQ_OP_READ:
> + if (bdev_is_zoned(bdev) &&
> + bdev_zone_is_offline(bdev, bio->bi_iter.bi_sector)) {
> + bio_set_flag(bio, BIO_QUIET);
> + goto end_io;
> + }
[Severity: Medium]
Does this offline zone check get bypassed if a read crosses a zone boundary
into an offline zone?
If a read starts in an online zone but crosses into an offline zone, the
initial sector check passes here. The bio is later split at the zone
boundary by blk_mq_submit_bio() due to chunk_sectors alignment for zoned
devices.
The remainder of the split bio, which now starts exactly in the offline
zone, is then resubmitted by bio_submit_split_bioset() using:
block/blk-merge.c:bio_submit_split_bioset() {
...
else if (!blk_throtl_bio(bio))
submit_bio_noacct_nocheck(bio, true);
...
}
Because the remainder is submitted via submit_bio_noacct_nocheck(), it skips
submit_bio_noacct() completely. The offline remainder will be issued to
the hardware without the BIO_QUIET flag, potentially causing the noisy
device errors this code is intended to prevent.
> break;
> case REQ_OP_WRITE:
> if (bio->bi_opf & REQ_ATOMIC) {
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261005094632.580753-1-dlemoal@kernel.org?part=1
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 2/7] block: introduce storage element management
2026-10-05 9:46 ` [PATCH 2/7] block: introduce storage element management Damien Le Moal
@ 2026-10-05 10:01 ` sashiko-bot
2026-10-05 10:52 ` Hannes Reinecke
2026-10-05 22:15 ` kernel test robot
2 siblings, 0 replies; 26+ messages in thread
From: sashiko-bot @ 2026-10-05 10:01 UTC (permalink / raw)
To: Damien Le Moal; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
- [High] block: page index vs byte count mismatch in invalidate_inode_pages2_range
- [High] block: heap buffer overflow in disk_wait_for_se_mgmt_completion array reuse
- [High] block: dirty page writeback omission before storage element removal
- [Medium] block: silent error abort in disk_wait_for_se_mgmt_completion causes premature zone revalidation
- [Low] block: blk_storage_element_status typos and duplicate fields in UAPI header
--
commit 71418115a91f7676078cbe8877b24a7405924eb7
Author: Damien Le Moal <dlemoal@kernel.org>
block: introduce storage element management
This patch introduces storage element management features for managing hard
disk heads. It defines generic block device driver operations for reporting,
removing, and restoring storage elements, allowing file systems to use the
data-preserving flavor of the depopulation feature.
> diff --git a/block/blk-zoned.c b/block/blk-zoned.c
> index 131c9f50b3da5..b713512f7d48f 100644
> --- a/block/blk-zoned.c
> +++ b/block/blk-zoned.c
[ ... ]
> +static void disk_wait_for_se_mgmt_completion(struct gendisk *disk)
> +{
> + struct blk_storage_element *elements, *e;
> + unsigned int i, nr_se, nr_elements = 0;
> + int ret;
> +
> + ret = (disk->fops->se_ops->report_elements)(disk, &nr_elements, NULL);
> + if (ret) {
> + pr_err("Failed to get number of storage elements\n");
> + return;
> + }
> +
> + elements = kzalloc_objs(struct blk_storage_element, nr_elements);
> + if (!elements)
> + return;
[Severity: Medium]
Will this silently abort on failure instead of returning an error or retrying?
If report_elements() or the memory allocation fails, the function simply
returns. This means callers like bdev_remove_storage_element() will
immediately proceed to invoke blk_revalidate_disk_zones() while the hardware
is potentially still actively removing or restoring an element.
> +
> + while (1) {
> + /*
> + * Check if we have storage elements being removed or restored.
> + */
> + nr_se = nr_elements;
> + ret = (disk->fops->se_ops->report_elements)(disk,
> + &nr_elements, elements);
[Severity: High]
Does this code risk a heap buffer overflow if the hardware dynamically
changes the number of reported elements?
The driver's report_elements() callback modifies nr_elements to the actual
number of descriptors returned. In the next iteration, nr_se is set to this
updated count and passed back as the buffer capacity. If the hardware reports
a higher count during polling, nr_elements is treated as a larger capacity
than what was originally allocated via kzalloc_objs().
> + if (ret) {
> + pr_err("Failed to get storage elements\n");
> + break;
> + }
> +
> + e = elements;
> + for (i = 0; i < nr_se; i++, e++) {
> + if (e->status == BLK_SE_STS_REMOVE_IN_PROGRESS ||
> + e->status == BLK_SE_STS_RESTORE_IN_PROGRESS)
> + break;
> + }
[Severity: High]
Could this read uninitialized or stale memory if the driver returns fewer
elements than nr_se?
The loop iterates up to the pre-call capacity nr_se instead of the actual
returned count nr_elements.
[ ... ]
> +int bdev_remove_storage_element(struct block_device *bdev,
> + unsigned int element_id)
> +{
[ ... ]
> + /*
> + * Flush the volatile write cache if there is any. Note that this may
> + * lead to errors if degraded storage elements are used. So ignore
> + * errors here.
> + */
> + blkdev_issue_flush(disk->part0);
> +
> + /* Invalidate all cached data. */
> + invalidate_inode_pages2_range(bdev->bd_mapping, 0,
> + get_capacity(disk) << SECTOR_SHIFT);
[Severity: High]
Does this miss dirty pages in the block device's page cache?
invalidate_inode_pages2_range() skips dirty pages and returns -EBUSY when it
encounters them. Since sync_blockdev() or filemap_write_and_wait() is not
called before this invalidation, dirty pages remain in memory and will
eventually be written back by background threads to obsolete or offline zones
after the storage element is removed.
[Severity: High]
Does this pass a byte size instead of a page index, leading to incomplete
cache invalidation?
The invalidate_inode_pages2_range() function expects a page index. The
code here passes get_capacity(disk) << SECTOR_SHIFT, which is the capacity
in bytes. On 32-bit platforms where pgoff_t is a 32-bit unsigned integer,
passing a large byte count causes the value to be silently truncated modulo
2^32. For example, a 1TB disk truncates to 0, meaning only page 0 is
invalidated.
> +
> + ret = (disk->fops->se_ops->remove_element)(disk, element_id);
> + if (ret)
> + return ret;
[ ... ]
> +int bdev_restore_storage_elements(struct block_device *bdev)
> +{
[ ... ]
> + /* Invalidate all cached data. */
> + invalidate_inode_pages2_range(bdev->bd_mapping, 0,
> + get_capacity(disk) << SECTOR_SHIFT);
[Severity: High]
Does this suffer from the same page index vs byte count mismatch as above,
causing truncation on 32-bit systems?
> +
> + ret = (disk->fops->se_ops->restore_elements)(disk);
> + if (ret)
> + return ret;
[ ... ]
> diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
> index 6638361209667..3a4bacfe79797 100644
> --- a/include/uapi/linux/blkzoned.h
> +++ b/include/uapi/linux/blkzoned.h
> @@ -208,4 +208,71 @@ struct blk_zone_range {
> #define BLKFINISHZONE _IOW(0x12, 136, struct blk_zone_range)
> #define BLKREPORTZONEV2 _IOWR(0x12, 142, struct blk_zone_report)
>
> +/**
> + * enum blk_storage_element_status - Statuc of a zoned device storage elements.
> + *
> + * @BLK_SE_RDWR: The storage element handles both reads and writes.
> + * @BLK_SE_READ: The storage element handles reads only.
> + * @BLK_SE_WRITE: The storage element handles writes only.
> + * @BLK_SE_WRITE: The storage element type is unknown.
> + */
> +enum blk_storage_element_type {
[Severity: Low]
This isn't a bug, but there are some typos in this UAPI header.
The documentation has a typo in "Statuc" and names the wrong enum in the
description. Additionally, @BLK_SE_WRITE is duplicated in the docstring
instead of documenting @BLK_SE_TYPE_UNKNOWN.
[ ... ]
> +enum blk_storage_element_status {
> + BLK_SE_STS_OK = 0x01,
> + BLK_SE_STS_DEGRADED = 0x02,
> + BLK_SE_STS_REMOVE_IN_PROGRESS = 0x03,
> + BLK_SE_STS_REMOVE_ERROR = 0x04,
> + BLK_SE_STS_RESTORE_IN_PROGRESS = 0x05,
> + BLK_SE_STS_RESTORE_ERROR = 0x06,
> + BLK_SE_STS_REMOVED = 0x07,
> + BLK_SE_STS_UNKOWN = 0xFF,
> +};
[Severity: Low]
This isn't a bug, but the UAPI constant for the unknown state is misspelled
as BLK_SE_STS_UNKOWN. Correcting this before the UAPI freezes could prevent
compatibility issues later.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261005094632.580753-1-dlemoal@kernel.org?part=2
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 1/7] block: fail reads to offline zones early
2026-10-05 9:46 ` [PATCH 1/7] block: fail reads to offline zones early Damien Le Moal
2026-10-05 10:01 ` sashiko-bot
@ 2026-10-05 10:45 ` Hannes Reinecke
1 sibling, 0 replies; 26+ messages in thread
From: Hannes Reinecke @ 2026-10-05 10:45 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 10/5/26 11:46 AM, Damien Le Moal wrote:
> Any BIO targeting an offline zone of a zoned block device will fail,
> including read commands. So there is no point in issuing read BIOs.
> Fail these early if we ever see one, and be quiet about the error.
>
> Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
> ---
> block/blk-core.c | 5 +++++
> block/blk-zoned.c | 18 ++++++++++++++++++
> block/blk.h | 6 ++++++
> 3 files changed, 29 insertions(+)
>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 2/7] block: introduce storage element management
2026-10-05 9:46 ` [PATCH 2/7] block: introduce storage element management Damien Le Moal
2026-10-05 10:01 ` sashiko-bot
@ 2026-10-05 10:52 ` Hannes Reinecke
2026-10-05 22:15 ` kernel test robot
2 siblings, 0 replies; 26+ messages in thread
From: Hannes Reinecke @ 2026-10-05 10:52 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 10/5/26 11:46 AM, Damien Le Moal wrote:
> Recent SCSI (SBC) and ATA (ACS) standards define the storage element
> depopulation feature. This feature is intended for managing hard-disks
> heads, either combined read-write heads or pairs of read and write heads,
> allowing to keep disks with defective heads longer in production by
> allowing "depopulating" (removing) defective heads.
>
> This feature comes in two different flavors:
> - A destructive version which removes a head and reformats the disk at a
> lower capacity point, restoaring a fully functional contiguous LBA
> address space.
> - A data preserving version restricted to host-managed zoned disks, which
> marks the zones served by a removed head as offline or read-only.
>
> In preparation for using the data-preserving flavor of the depopulation
> feature in file systems natively supporting zoned block devices, introduce
> a set of storage element management operations and functions to define
> generic calls into block device drivers for managing storage elements.
>
> The set of operations is defined with struct blk_storage_elements_ops and
> includes three operations:
> - report_elements: get information on a device storage elements state
> - remove_element: depopulate a defective storage element
> - restore_elements: restore depopulated storage elements
>
> Each operation is called from the functions
> bdev_report_storage_elements(), bdev_remove_storage_element() and
> bdev_restore_storage_elements().
>
> Removing a healthy storage element from a device is possible and useful
> for testing. The storage element restoration operation restore_elements
> allows repopulating such healthy element. Repopulating defective storage
> elements is generally not allowed by devices.
>
> The device drivers of zoned block devices can indicate support for the
> storage element depopulation feature by specifying the storage element
> management operations with the se_ops field of struct
> block_device_operations.
>
> Co-developed-by: Christoph Hellwig <hch@lst.de>
> Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
> ---
> block/blk-zoned.c | 168 +++++++++++++++++++++++++++++++++-
> include/linux/blkdev.h | 16 ++++
> include/uapi/linux/blkzoned.h | 67 ++++++++++++++
> 3 files changed, 250 insertions(+), 1 deletion(-)
>
> diff --git a/block/blk-zoned.c b/block/blk-zoned.c
> index 131c9f50b3da..b713512f7d48 100644
> --- a/block/blk-zoned.c
> +++ b/block/blk-zoned.c
> @@ -18,6 +18,7 @@
> #include <linux/mempool.h>
> #include <linux/kthread.h>
> #include <linux/freezer.h>
> +#include <linux/delay.h>
>
> #include <trace/events/block.h>
>
> @@ -2718,5 +2719,170 @@ int queue_zone_wplugs_show(void *data, struct seq_file *m)
>
> return 0;
> }
> -
> #endif
> +
> +static void disk_wait_for_se_mgmt_completion(struct gendisk *disk)
> +{
> + struct blk_storage_element *elements, *e;
> + unsigned int i, nr_se, nr_elements = 0;
> + int ret;
> +
> + ret = (disk->fops->se_ops->report_elements)(disk, &nr_elements, NULL);
> + if (ret) {
> + pr_err("Failed to get number of storage elements\n");
> + return;
> + }
> +
> + elements = kzalloc_objs(struct blk_storage_element, nr_elements);
> + if (!elements)
> + return;
> +
> + while (1) {
> + /*
> + * Check if we have storage elements being removed or restored.
> + */
> + nr_se = nr_elements;
> + ret = (disk->fops->se_ops->report_elements)(disk,
> + &nr_elements, elements);
> + if (ret) {
> + pr_err("Failed to get storage elements\n");
> + break;
> + }
> +
> + e = elements;
> + for (i = 0; i < nr_se; i++, e++) {
> + if (e->status == BLK_SE_STS_REMOVE_IN_PROGRESS ||
> + e->status == BLK_SE_STS_RESTORE_IN_PROGRESS)
> + break;
> + }
> + if (i >= nr_se)
> + break;
> +
> + /* Not done yet: wait and retry. */
> + msleep(500);
> + }
> +
> + kfree(elements);
> +}
> +
> +/**
> + * bdev_report_storage_elements - report the storage elements of a block device
> + *
> + * Fill at most @nr_elements storage element descriptors in the array @elements.
> + * The number of storage elements filled in the array is returned using
> + * @nr_elements. If @elements is NULL, only @nr_elements is returned.
> + *
> + * Returns 0 on success and a negative error code on failure.
> + */
> +int bdev_report_storage_elements(struct block_device *bdev,
> + unsigned int *nr_elements,
> + struct blk_storage_element *elements)
> +{
> + struct gendisk *disk = bdev->bd_disk;
> +
> + if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
> + return -EOPNOTSUPP;
> +
> + if (!nr_elements)
> + return -EINVAL;
> +
> + if (*nr_elements && !elements)
> + return -EINVAL;
> +
> + return (disk->fops->se_ops->report_elements)(disk, nr_elements,
> + elements);
> +}
> +EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
> +
> +/**
> + * bdev_remove_storage_element - Remove (depopulate) a storage element of a
> + * block device
> + *
> + * Remove (depopulate) the storage element identified by @element_id from the
> + * block device @bdev.
> + *
> + * Returns 0 on success and a negative error code on failure.
> + */
> +int bdev_remove_storage_element(struct block_device *bdev,
> + unsigned int element_id)
> +{
> + struct gendisk *disk = bdev->bd_disk;
> + unsigned int memflags;
> + int ret;
> +
> + if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
> + return -EOPNOTSUPP;
> +
> + /* Zero is not a valid storage element ID. */
> + if (!element_id)
> + return -EINVAL;
> +
> + /*
> + * Freeze and unfreeze the queue to flush any outstanding command.
> + * The caller is responsible for not queuing up more I/Os by higher
> + * level means.
> + */
The 'caller' is the ioctl, so it can only issue I/O _after_ the ioctl.
Maybe 'user' is a better wording here.
> + memflags = blk_mq_freeze_queue(disk->queue);
> + blk_mq_unfreeze_queue(disk->queue, memflags);
> +
> + /*
> + * Flush the volatile write cache if there is any. Note that this may
> + * lead to errors if degraded storage elements are used. So ignore
> + * errors here.
> + */
> + blkdev_issue_flush(disk->part0);
Do you need a flush here?
I would have thought that the freeze/unfreeze above would take care
of it.
But what do I know.
> +
> + /* Invalidate all cached data. */
> + invalidate_inode_pages2_range(bdev->bd_mapping, 0,
> + get_capacity(disk) << SECTOR_SHIFT);
> +
> + ret = (disk->fops->se_ops->remove_element)(disk, element_id);
> + if (ret)
> + return ret;
> +
> + /* Revalidate the device zones once the opration completes. */
> + disk_wait_for_se_mgmt_completion(disk);
> +
> + return blk_revalidate_disk_zones(disk);
> +}
> +EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
> +
> +/**
> + * bdev_restore_storage_elements - Restore all depopulated storage elements of a
> + * block device
> + *
> + * Restore all storage elements of @bdev that have been depopulated. Not all
> + * elements may be restored by this operation.
> + *
> + * Returns 0 on success and a negative error code on failure.
> + */
> +int bdev_restore_storage_elements(struct block_device *bdev)
> +{
> + struct gendisk *disk = bdev->bd_disk;
> + unsigned int memflags;
> + int ret;
> +
> + if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
> + return -EOPNOTSUPP;
> +
> + /*
> + * Freeze and unfreeze the queue to flush any outstanding commands.
> + * The caller is responsible for not queuing up more I/O by higher
> + * level means.
> + */
See above for 'the caller is responsible'.
> + memflags = blk_mq_freeze_queue(disk->queue);
> + blk_mq_unfreeze_queue(disk->queue, memflags);
> +
> + /* Invalidate all cached data. */
> + invalidate_inode_pages2_range(bdev->bd_mapping, 0,
> + get_capacity(disk) << SECTOR_SHIFT);
> +
> + ret = (disk->fops->se_ops->restore_elements)(disk);
> + if (ret)
> + return ret;
> +
> + disk_wait_for_se_mgmt_completion(disk);
> +
> + return blk_revalidate_disk_zones(disk);
> +}
> +EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
> diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
> index d003a9d2d1f6..0b2013c96cad 100644
> --- a/include/linux/blkdev.h
> +++ b/include/linux/blkdev.h
> @@ -1574,6 +1574,21 @@ enum blk_unique_id {
> BLK_UID_NAA = 3,
> };
>
> +struct blk_storage_elements_ops {
> + int (*report_elements)(struct gendisk *disk,
> + unsigned int *nr_elements,
> + struct blk_storage_element *elements);
> + int (*remove_element)(struct gendisk *disk, unsigned int element_id);
> + int (*restore_elements)(struct gendisk *disk);
> +};
> +
> +int bdev_report_storage_elements(struct block_device *bdev,
> + unsigned int *nr_elements,
> + struct blk_storage_element *elements);
> +int bdev_remove_storage_element(struct block_device *bdev,
> + unsigned int element_id);
> +int bdev_restore_storage_elements(struct block_device *bdev);
> +
> struct block_device_operations {
> void (*submit_bio)(struct bio *bio);
> int (*poll_bio)(struct bio *bio, struct io_comp_batch *iob,
> @@ -1601,6 +1616,7 @@ struct block_device_operations {
> enum blk_unique_id id_type);
> struct module *owner;
> const struct pr_ops *pr_ops;
> + const struct blk_storage_elements_ops *se_ops;
Indent?
>
> /*
> * Special callback for probing GPT entry at a given sector.
> diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
> index 663836120966..3a4bacfe7979 100644
> --- a/include/uapi/linux/blkzoned.h
> +++ b/include/uapi/linux/blkzoned.h
> @@ -208,4 +208,71 @@ struct blk_zone_range {
> #define BLKFINISHZONE _IOW(0x12, 136, struct blk_zone_range)
> #define BLKREPORTZONEV2 _IOWR(0x12, 142, struct blk_zone_report)
>
> +/**
> + * enum blk_storage_element_status - Statuc of a zoned device storage elements.
> + *
> + * @BLK_SE_RDWR: The storage element handles both reads and writes.
> + * @BLK_SE_READ: The storage element handles reads only.
> + * @BLK_SE_WRITE: The storage element handles writes only.
> + * @BLK_SE_WRITE: The storage element type is unknown.
> + */
> +enum blk_storage_element_type {
> + BLK_SE_TYPE_RDWR = 0x01,
> + BLK_SE_TYPE_READ = 0x02,
> + BLK_SE_TYPE_WRITE = 0x03,
> + BLK_SE_TYPE_UNKNOWN = 0xFF,
> +};
> +
> +/**
> + * enum blk_storage_element_status - Status of a zoned device storage elements.
> + *
> + * @BLK_SE_STS_OK: The storage element is operating normally.
> + * @BLK_SE_STS_DEGRADED: The storage element has degraded and is not operating
> + * normally.
> + * @BLK_SE_STS_REMOVE_IN_PROGRESS: The storage element is being removed.
> + * @BLK_SE_STS_REMOVE_ERROR: The storage element removal failed.
> + * @BLK_SE_STS_RESTORE_IN_PROGRESS: The storage element is being restored.
> + * @BLK_SE_STS_RESTORE_ERROR: The storage element restoration failed.
> + * @BLK_SE_STS_REMOVED: The storage element was removed.
> + * @BLK_SE_STS_UNKNOWN: The storage element status is unknown.
> + */
> +enum blk_storage_element_status {
> + BLK_SE_STS_OK = 0x01,
> + BLK_SE_STS_DEGRADED = 0x02,
> + BLK_SE_STS_REMOVE_IN_PROGRESS = 0x03,
> + BLK_SE_STS_REMOVE_ERROR = 0x04,
> + BLK_SE_STS_RESTORE_IN_PROGRESS = 0x05,
> + BLK_SE_STS_RESTORE_ERROR = 0x06,
> + BLK_SE_STS_REMOVED = 0x07,
> + BLK_SE_STS_UNKOWN = 0xFF,
> +};
> +
> +/**
> + * struct blk_storage_element - Zoned device storage element descriptor.
> + *
> + * @id: The ID of the element (cannot be 0).
> + * @paired_id: The ID of the paired element for an element that is not
> + * of type BLK_SE_TYPE_RDWR.
> + * @type: The type of the storage element (enum blk_storage_element_type).
> + * @status: The health status of the storage element
> + * (enum blk_storage_element_status).
> + * @restore_allowed: Indicate if the storage element can be restored.
> + * @nr_zones: The number of zones that the storage element handles.
> + */
> +struct blk_storage_element {
> + __u32 id;
> + __u32 paired_id;
> + __u64 nr_zones;
> + __u8 type;
> + __u8 status;
> + __u8 restore_allowed;
> + __u8 reserved[5];
> +};
> +
> +struct blk_storage_elements_report {
> + __u32 nr_elements;
> + __u32 reserved;
> + struct blk_storage_element elements[];
> +};
> +
> #endif /* _UAPI_BLKZONED_H */
Otherwise looks good.
Cheers
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 3/7] block: add storage element management ioctls
2026-10-05 9:46 ` [PATCH 3/7] block: add storage element management ioctls Damien Le Moal
2026-10-05 10:00 ` sashiko-bot
@ 2026-10-05 11:09 ` Hannes Reinecke
1 sibling, 0 replies; 26+ messages in thread
From: Hannes Reinecke @ 2026-10-05 11:09 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 10/5/26 11:46 AM, Damien Le Moal wrote:
> Define the new ioctl commands BLKGETNRSTORELEMS, BLKREPORTSTORELEMS,
> BLKREMOVESTORELEM, and BLKRESTORESTORELEMS to provide users accessing
> zoned block devices directly with an interface to the storage elements
> management operations of the block layer.
>
> The ioctls BLKGETNRSTORELEMS and BLKREPORTSTORELEMS are defined to
> respectively get the number of storage elements of a zoned device and to
> get an array of struct blk_storage_element describing the current state of
> the device storage elements.
>
> The ioctl BLKREMOVESTORELEM can be used to remove (depopulate) a storage
> element that has a degraded status (i.e. BLK_SE_STS_DEGRADED).
>
> Finally, the BLKRESTORESTORELEMS ioctl interfaces with
> bdev_restore_storage_elements() to restore the depopulated storage
> elements of a zoned device.
>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 0/7] Add support for storage element depopulation
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
` (6 preceding siblings ...)
2026-10-05 9:46 ` [PATCH 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
@ 2026-10-05 11:13 ` Hannes Reinecke
2026-10-07 7:14 ` Damien Le Moal
7 siblings, 1 reply; 26+ messages in thread
From: Hannes Reinecke @ 2026-10-05 11:13 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 10/5/26 11:46 AM, Damien Le Moal wrote:
> Jens,
>
> These patches define a new set of block device operations for generically
> using from the block layer the storage element depopulation feature of
> zoned block devices. Support for this feature is added to the SCSI disk
> driver and an emulation of this feature added to the zloop driver.
>
> The management operations can be accessed using block layer API, which is
> intended for file systems (e.g. zonefs and XFS), as well as using ioctls
> for users using zoned disks directly from user space.
>
> Damien Le Moal (7):
> block: fail reads to offline zones early
> block: introduce storage element management
> block: add storage element management ioctls
> zloop: add storage element emulation
> zloop: add degrade_element control command
> scsi: sd_zbc: always revalidate zones for disks supporting head depopulation
> scsi: sd_zbc: define storage element management operations
>
I wonder: shouldn't we send a uevent after either operation?
I guess that this would be helpful for admins, and not forgetting
udev which might want to re-run any rules the admin might have
configured.
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 4/7] zloop: add storage element emulation
2026-10-05 9:46 ` [PATCH 4/7] zloop: add storage element emulation Damien Le Moal
2026-10-05 9:58 ` sashiko-bot
@ 2026-10-05 11:14 ` Hannes Reinecke
1 sibling, 0 replies; 26+ messages in thread
From: Hannes Reinecke @ 2026-10-05 11:14 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 10/5/26 11:46 AM, Damien Le Moal wrote:
> Emulate the storage element depopulation feature of SMR disks in the
> zloop driver. This emulation is controlled using the new stor_elements
> option.
>
> This option can take several values:
> - ZLOOP_STOR_ELEMENTS_NONE (0): no emulation (default)
> - ZLOOP_STOR_ELEMENTS_RDWR (1): emulate all access storage elements
> (e.g. read+write heads)
> - ZLOOP_STOR_ELEMENTS_PAIRS (2): emulate fractional access storage
> elements (e.g. pairs of read and write heads)
>
> If enabled with the value 1 or 2, the number of storage elements, or of
> pairs of fractional access storage elements, is automatically calculated
> based on the number of zones of the device so that we have at least 2 and
> at most 32 storage elements (or 2 pairs of fractional access storage
> elements).
>
> The mapping of zones to storage elements is defined simply as the modulo
> of a zone number and a storage element ID. E.g, removing a storage
> element from a set of 4 storage elements will offline 1 zone every 4
> zones.
>
> The storage element management operations are specified with the
> zloop_se_ops (struct blk_storage_elements_ops).
>
> Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
> ---
> .../admin-guide/blockdev/zoned_loop.rst | 8 +-
> drivers/block/zloop.c | 510 +++++++++++++++++-
> 2 files changed, 490 insertions(+), 28 deletions(-)
>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 5/7] zloop: add degrade_element control command
2026-10-05 9:46 ` [PATCH 5/7] zloop: add degrade_element control command Damien Le Moal
2026-10-05 9:58 ` sashiko-bot
@ 2026-10-05 11:17 ` Hannes Reinecke
1 sibling, 0 replies; 26+ messages in thread
From: Hannes Reinecke @ 2026-10-05 11:17 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 10/5/26 11:46 AM, Damien Le Moal wrote:
> Allow users to mark storage elements of a zloop device as degraded using
> the new "degrade_element" control command. The element to degrade is
> indicated using the element_id option. Example:
>
> echo "degrade_element id=0,element_id=2" > /dev/zloop-control
>
> If the element ID identifies an all access storage element, read and write
> operations targeting a zone served by the degraded lement are failed.
> For a partial access storage element, read or write operations are failed
> depending on the storage element type.
>
> Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
> ---
> drivers/block/zloop.c | 112 ++++++++++++++++++++++++++++++++++++++++--
> 1 file changed, 107 insertions(+), 5 deletions(-)
>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation
2026-10-05 9:46 ` [PATCH 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation Damien Le Moal
@ 2026-10-05 11:19 ` Hannes Reinecke
0 siblings, 0 replies; 26+ messages in thread
From: Hannes Reinecke @ 2026-10-05 11:19 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 10/5/26 11:46 AM, Damien Le Moal wrote:
> sd_zbc_revalidate_zones() skips revalidating the zones of a ZBC device if
> the zone size and total number of zones of the disk has not changed. This
> is to avoid a call to the rather slow blk_revalidate_disk_zones().
>
> However, for ZBC devices that support data preserving head depopulation
> (REMOVE ELEMENT AND MODIFY ZONES command), a disk capacity and number of
> zones does not change after a head is depopulated but the condition of
> zones changes as the zones served by the head that was depopulated become
> either read-only or offline. In this case, not calling
> blk_revalidate_disk_zones() prevents the block layer from taking
> appropriate actions on the zone write plugs of the disk for the zones that
> became read-only or offline.
>
> Avoid any issue with the block layer view of the zone conditions by not
> skipping the call to blk_revalidate_disk_zones() for disks that support
> the REMOVE ELEMENT AND MODIFY ZONES command. This check is done from
> sd_zbc_read_zones() using the helper function sd_zbc_check_modify_zones().
> The new scsi disk flag modify_zones_supported is defined to remember the
> result of this check.
>
> Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
> ---
> drivers/scsi/sd.h | 1 +
> drivers/scsi/sd_zbc.c | 26 +++++++++++++++++++++++++-
> 2 files changed, 26 insertions(+), 1 deletion(-)
>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 7/7] scsi: sd_zbc: define storage element management operations
2026-10-05 9:46 ` [PATCH 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
2026-10-05 9:59 ` sashiko-bot
@ 2026-10-05 11:48 ` Hannes Reinecke
2026-10-05 20:48 ` kernel test robot
2026-10-05 21:41 ` kernel test robot
3 siblings, 0 replies; 26+ messages in thread
From: Hannes Reinecke @ 2026-10-05 11:48 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 10/5/26 11:46 AM, Damien Le Moal wrote:
> Define the storage element management operations using struct
> blk_storage_elements_ops. These operations are valid only on SMR disks
> supporting the storage element depopulation feature.
>
> Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
> ---
> drivers/scsi/sd.c | 1 +
> drivers/scsi/sd.h | 4 +
> drivers/scsi/sd_zbc.c | 210 ++++++++++++++++++++++++++++++++++++++++++
> 3 files changed, 215 insertions(+)
>
Reviewed-by: Hannes Reinecke <hare@kernel.org>
Cheers,
Hannes
--
Dr. Hannes Reinecke Kernel Storage Architect
hare@suse.de +49 911 74053 688
SUSE Software Solutions GmbH, Frankenstr. 146, 90461 Nürnberg
HRB 36809 (AG Nürnberg), GF: I. Totev, A. McDonald, W. Knoblich
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 7/7] scsi: sd_zbc: define storage element management operations
2026-10-05 9:46 ` [PATCH 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
2026-10-05 9:59 ` sashiko-bot
2026-10-05 11:48 ` Hannes Reinecke
@ 2026-10-05 20:48 ` kernel test robot
2026-10-05 21:41 ` kernel test robot
3 siblings, 0 replies; 26+ messages in thread
From: kernel test robot @ 2026-10-05 20:48 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
Cc: oe-kbuild-all
Hi Damien,
kernel test robot noticed the following build errors:
[auto build test ERROR on axboe/for-next]
[also build test ERROR on next-20261005]
[cannot apply to mkp-scsi/for-next jejb-scsi/for-next linus/master v7.3-rc6]
[If your patch is applied to the wrong git tree, kindly drop us a note.
And when submitting patch, we suggest to use '--base' as documented in
https://git-scm.com/docs/git-format-patch#_base_tree_information]
url: https://github.com/intel-lab-lkp/linux/commits/Damien-Le-Moal/block-fail-reads-to-offline-zones-early/20261005-184626
base: https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git for-next
patch link: https://lore.kernel.org/r/20261005094632.580753-8-dlemoal%40kernel.org
patch subject: [PATCH 7/7] scsi: sd_zbc: define storage element management operations
config: sparc-defconfig (https://download.01.org/0day-ci/archive/20261006/202610060413.UQ1Jyndv-lkp@intel.com/config)
compiler: sparc-linux-gcc (GCC) 16.1.0
reproduce (this is a W=1 build): (https://download.01.org/0day-ci/archive/20261006/202610060413.UQ1Jyndv-lkp@intel.com/reproduce)
If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <lkp@intel.com>
| Closes: https://lore.kernel.org/oe-kbuild-all/202610060413.UQ1Jyndv-lkp@intel.com/
All errors (new ones prefixed by >>):
>> drivers/scsi/sd.c:3940:35: error: lvalue required as unary '&' operand
3940 | .se_ops = &sd_zbc_se_ops,
| ^
vim +3940 drivers/scsi/sd.c
3926
3927 static const struct block_device_operations sd_fops = {
3928 .owner = THIS_MODULE,
3929 .open = sd_open,
3930 .release = sd_release,
3931 .ioctl = sd_ioctl,
3932 .getgeo = sd_getgeo,
3933 .compat_ioctl = blkdev_compat_ptr_ioctl,
3934 .check_events = sd_check_events,
3935 .unlock_native_capacity = sd_unlock_native_capacity,
3936 .report_zones = sd_zbc_report_zones,
3937 .get_unique_id = sd_get_unique_id,
3938 .free_disk = scsi_disk_free_disk,
3939 .pr_ops = &sd_pr_ops,
> 3940 .se_ops = &sd_zbc_se_ops,
3941 };
3942
--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 7/7] scsi: sd_zbc: define storage element management operations
2026-10-05 9:46 ` [PATCH 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
` (2 preceding siblings ...)
2026-10-05 20:48 ` kernel test robot
@ 2026-10-05 21:41 ` kernel test robot
3 siblings, 0 replies; 26+ messages in thread
From: kernel test robot @ 2026-10-05 21:41 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
Cc: llvm, oe-kbuild-all
Hi Damien,
kernel test robot noticed the following build errors:
[auto build test ERROR on axboe/for-next]
[also build test ERROR on next-20261005]
[cannot apply to mkp-scsi/for-next jejb-scsi/for-next linus/master v7.3-rc6]
[If your patch is applied to the wrong git tree, kindly drop us a note.
And when submitting patch, we suggest to use '--base' as documented in
https://git-scm.com/docs/git-format-patch#_base_tree_information]
url: https://github.com/intel-lab-lkp/linux/commits/Damien-Le-Moal/block-fail-reads-to-offline-zones-early/20261005-184626
base: https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git for-next
patch link: https://lore.kernel.org/r/20261005094632.580753-8-dlemoal%40kernel.org
patch subject: [PATCH 7/7] scsi: sd_zbc: define storage element management operations
config: sparc64-defconfig (https://download.01.org/0day-ci/archive/20261006/202610060512.08JQbDV4-lkp@intel.com/config)
compiler: clang version 24.0.0git (https://github.com/llvm/llvm-project 80cd965b6323b208519f8ab1d5ee86e4ea9862f4)
reproduce (this is a W=1 build): (https://download.01.org/0day-ci/archive/20261006/202610060512.08JQbDV4-lkp@intel.com/reproduce)
If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <lkp@intel.com>
| Closes: https://lore.kernel.org/oe-kbuild-all/202610060512.08JQbDV4-lkp@intel.com/
All errors (new ones prefixed by >>):
>> drivers/scsi/sd.c:3940:14: error: cannot take the address of an rvalue of type 'void *'
3940 | .se_ops = &sd_zbc_se_ops,
| ^~~~~~~~~~~~~~
1 error generated.
vim +3940 drivers/scsi/sd.c
3926
3927 static const struct block_device_operations sd_fops = {
3928 .owner = THIS_MODULE,
3929 .open = sd_open,
3930 .release = sd_release,
3931 .ioctl = sd_ioctl,
3932 .getgeo = sd_getgeo,
3933 .compat_ioctl = blkdev_compat_ptr_ioctl,
3934 .check_events = sd_check_events,
3935 .unlock_native_capacity = sd_unlock_native_capacity,
3936 .report_zones = sd_zbc_report_zones,
3937 .get_unique_id = sd_get_unique_id,
3938 .free_disk = scsi_disk_free_disk,
3939 .pr_ops = &sd_pr_ops,
> 3940 .se_ops = &sd_zbc_se_ops,
3941 };
3942
--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 2/7] block: introduce storage element management
2026-10-05 9:46 ` [PATCH 2/7] block: introduce storage element management Damien Le Moal
2026-10-05 10:01 ` sashiko-bot
2026-10-05 10:52 ` Hannes Reinecke
@ 2026-10-05 22:15 ` kernel test robot
2 siblings, 0 replies; 26+ messages in thread
From: kernel test robot @ 2026-10-05 22:15 UTC (permalink / raw)
To: Damien Le Moal, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
Cc: oe-kbuild-all
Hi Damien,
kernel test robot noticed the following build warnings:
[auto build test WARNING on axboe/for-next]
[also build test WARNING on next-20261005]
[cannot apply to mkp-scsi/for-next jejb-scsi/for-next linus/master v7.3-rc6]
[If your patch is applied to the wrong git tree, kindly drop us a note.
And when submitting patch, we suggest to use '--base' as documented in
https://git-scm.com/docs/git-format-patch#_base_tree_information]
url: https://github.com/intel-lab-lkp/linux/commits/Damien-Le-Moal/block-fail-reads-to-offline-zones-early/20261005-184626
base: https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git for-next
patch link: https://lore.kernel.org/r/20261005094632.580753-3-dlemoal%40kernel.org
patch subject: [PATCH 2/7] block: introduce storage element management
config: riscv-randconfig-1001-20261006 (https://download.01.org/0day-ci/archive/20261006/202610060627.NVrcnIcL-lkp@intel.com/config)
compiler: riscv64-linux-gcc (GCC) 8.5.0
reproduce (this is a W=1 build): (https://download.01.org/0day-ci/archive/20261006/202610060627.NVrcnIcL-lkp@intel.com/reproduce)
If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <lkp@intel.com>
| Closes: https://lore.kernel.org/oe-kbuild-all/202610060627.NVrcnIcL-lkp@intel.com/
All warnings (new ones prefixed by >>):
>> Warning: block/blk-zoned.c:2779 function parameter 'bdev' not described in 'bdev_report_storage_elements'
>> Warning: block/blk-zoned.c:2779 function parameter 'nr_elements' not described in 'bdev_report_storage_elements'
>> Warning: block/blk-zoned.c:2779 function parameter 'elements' not described in 'bdev_report_storage_elements'
>> Warning: block/blk-zoned.c:2807 function parameter 'bdev' not described in 'bdev_remove_storage_element'
>> Warning: block/blk-zoned.c:2807 function parameter 'element_id' not described in 'bdev_remove_storage_element'
>> Warning: block/blk-zoned.c:2859 function parameter 'bdev' not described in 'bdev_restore_storage_elements'
>> Warning: block/blk-zoned.c:2779 function parameter 'bdev' not described in 'bdev_report_storage_elements'
>> Warning: block/blk-zoned.c:2779 function parameter 'nr_elements' not described in 'bdev_report_storage_elements'
>> Warning: block/blk-zoned.c:2779 function parameter 'elements' not described in 'bdev_report_storage_elements'
>> Warning: block/blk-zoned.c:2807 function parameter 'bdev' not described in 'bdev_remove_storage_element'
>> Warning: block/blk-zoned.c:2807 function parameter 'element_id' not described in 'bdev_remove_storage_element'
>> Warning: block/blk-zoned.c:2859 function parameter 'bdev' not described in 'bdev_restore_storage_elements'
--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
^ permalink raw reply [flat|nested] 26+ messages in thread
* Re: [PATCH 0/7] Add support for storage element depopulation
2026-10-05 11:13 ` [PATCH 0/7] Add support for storage element depopulation Hannes Reinecke
@ 2026-10-07 7:14 ` Damien Le Moal
0 siblings, 0 replies; 26+ messages in thread
From: Damien Le Moal @ 2026-10-07 7:14 UTC (permalink / raw)
To: Hannes Reinecke, Jens Axboe, linux-block, Christoph Hellwig,
linux-scsi, Martin K . Petersen
On 2026/10/05 13:13, Hannes Reinecke wrote:
> On 10/5/26 11:46 AM, Damien Le Moal wrote:
>> Jens,
>>
>> These patches define a new set of block device operations for generically
>> using from the block layer the storage element depopulation feature of
>> zoned block devices. Support for this feature is added to the SCSI disk
>> driver and an emulation of this feature added to the zloop driver.
>>
>> The management operations can be accessed using block layer API, which is
>> intended for file systems (e.g. zonefs and XFS), as well as using ioctls
>> for users using zoned disks directly from user space.
>>
>> Damien Le Moal (7):
>> block: fail reads to offline zones early
>> block: introduce storage element management
>> block: add storage element management ioctls
>> zloop: add storage element emulation
>> zloop: add degrade_element control command
>> scsi: sd_zbc: always revalidate zones for disks supporting head depopulation
>> scsi: sd_zbc: define storage element management operations
>>
> I wonder: shouldn't we send a uevent after either operation?
> I guess that this would be helpful for admins, and not forgetting
> udev which might want to re-run any rules the admin might have
> configured.
Maybe. But given that head depop is a user/sysadmin initiated operation, at
least for now, udev/revalidation is something that the user/sysadmin can trigger
too for now.
In the next round, once we get file systems (zonefs and xfs) to initiate depop,
we can add such event notification.
--
Damien Le Moal
Western Digital Research
^ permalink raw reply [flat|nested] 26+ messages in thread
end of thread, other threads:[~2026-10-07 7:14 UTC | newest]
Thread overview: 26+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
2026-10-05 9:46 ` [PATCH 1/7] block: fail reads to offline zones early Damien Le Moal
2026-10-05 10:01 ` sashiko-bot
2026-10-05 10:45 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 2/7] block: introduce storage element management Damien Le Moal
2026-10-05 10:01 ` sashiko-bot
2026-10-05 10:52 ` Hannes Reinecke
2026-10-05 22:15 ` kernel test robot
2026-10-05 9:46 ` [PATCH 3/7] block: add storage element management ioctls Damien Le Moal
2026-10-05 10:00 ` sashiko-bot
2026-10-05 11:09 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 4/7] zloop: add storage element emulation Damien Le Moal
2026-10-05 9:58 ` sashiko-bot
2026-10-05 11:14 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 5/7] zloop: add degrade_element control command Damien Le Moal
2026-10-05 9:58 ` sashiko-bot
2026-10-05 11:17 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation Damien Le Moal
2026-10-05 11:19 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
2026-10-05 9:59 ` sashiko-bot
2026-10-05 11:48 ` Hannes Reinecke
2026-10-05 20:48 ` kernel test robot
2026-10-05 21:41 ` kernel test robot
2026-10-05 11:13 ` [PATCH 0/7] Add support for storage element depopulation Hannes Reinecke
2026-10-07 7:14 ` Damien Le Moal
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox