* [PATCH v2 1/7] block: fail reads to offline zones early
2026-10-06 12:45 [PATCH v2 0/7] Add support for storage element depopulation Damien Le Moal
@ 2026-10-06 12:45 ` Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 2/7] block: introduce storage element management Damien Le Moal
` (5 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Damien Le Moal @ 2026-10-06 12:45 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Any BIO targeting an offline zone of a zoned block device will fail,
including read commands. So there is no point in issuing read BIOs.
Fail these read operations early if we ever see one.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
block/blk-core.c | 3 +++
block/blk-zoned.c | 18 ++++++++++++++++++
block/blk.h | 6 ++++++
3 files changed, 27 insertions(+)
diff --git a/block/blk-core.c b/block/blk-core.c
index 13dc70e8f55d..d420c80d2d93 100644
--- a/block/blk-core.c
+++ b/block/blk-core.c
@@ -866,6 +866,9 @@ void submit_bio_noacct(struct bio *bio)
switch (bio_op(bio)) {
case REQ_OP_READ:
+ if (bdev_is_zoned(bdev) &&
+ bdev_zone_is_offline(bdev, bio->bi_iter.bi_sector))
+ goto end_io;
break;
case REQ_OP_WRITE:
if (bio->bi_opf & REQ_ATOMIC) {
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index 19268afb8752..131c9f50b3da 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -308,6 +308,24 @@ bool bdev_zone_is_seq(struct block_device *bdev, sector_t sector)
}
EXPORT_SYMBOL_GPL(bdev_zone_is_seq);
+/**
+ * bdev_zone_is_offline - check if a sector belongs to an offline zone
+ * @bdev: block device to check
+ * @sector: sector number
+ *
+ * Check if @sector on @bdev is contained in an offline zone.
+ */
+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector)
+{
+ enum blk_zone_cond cond;
+
+ if (!bdev_is_zoned(bdev))
+ return false;
+
+ cond = disk_zone_get_cond(bdev->bd_disk, sector);
+ return cond == BLK_ZONE_COND_OFFLINE;
+}
+
/**
* bdev_zone_mgmt_allowed - check if management operations are allowed on a zone
* @bdev: block device to check
diff --git a/block/blk.h b/block/blk.h
index 2cc03aa54c53..c2d07347a3ad 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -578,6 +578,7 @@ int blkdev_report_zones_ioctl(struct block_device *bdev, unsigned int cmd,
int blkdev_zone_mgmt_ioctl(struct block_device *bdev, blk_mode_t mode,
unsigned int cmd, unsigned long arg);
bool bdev_zone_mgmt_allowed(struct block_device *bdev, sector_t sector);
+bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector);
#else /* CONFIG_BLK_DEV_ZONED */
static inline void disk_init_zone_resources(struct gendisk *disk)
{
@@ -625,6 +626,11 @@ static inline bool bdev_zone_mgmt_allowed(struct block_device *bdev,
{
return false;
}
+static inline bool bdev_zone_is_offline(struct block_device *bdev,
+ sector_t sector)
+{
+ return false;
+}
#endif /* CONFIG_BLK_DEV_ZONED */
struct block_device *bdev_alloc(struct gendisk *disk, u8 partno);
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v2 2/7] block: introduce storage element management
2026-10-06 12:45 [PATCH v2 0/7] Add support for storage element depopulation Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 1/7] block: fail reads to offline zones early Damien Le Moal
@ 2026-10-06 12:45 ` Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 3/7] block: add storage element management ioctls Damien Le Moal
` (4 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Damien Le Moal @ 2026-10-06 12:45 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Recent SCSI (SBC) and ATA (ACS) standards define the storage element
depopulation feature. This feature is intended for managing hard-disks
heads, either combined read-write heads or pairs of read and write heads,
allowing to keep disks with defective heads longer in production by
allowing "depopulating" (removing) defective heads.
This feature comes in two different flavors:
- A destructive version which removes a head and reformats the disk at a
lower capacity point, restoaring a fully functional contiguous LBA
address space.
- A data preserving version restricted to host-managed zoned disks, which
marks the zones served by a removed head as offline or read-only.
In preparation for using the data-preserving flavor of the depopulation
feature in file systems natively supporting zoned block devices, introduce
a set of storage element management operations and functions to define
generic calls into block device drivers for managing storage elements.
The set of operations is defined with struct blk_storage_elements_ops and
includes three operations:
- report_elements: get information on a device storage elements state
- remove_element: depopulate a defective storage element
- restore_elements: restore depopulated storage elements
Each operation is called from the functions
bdev_report_storage_elements(), bdev_remove_storage_element() and
bdev_restore_storage_elements().
Removing a healthy storage element from a device is possible and useful
for testing. The storage element restoration operation restore_elements
allows repopulating such healthy element. Repopulating defective storage
elements is generally not allowed by devices.
The device drivers of zoned block devices can indicate support for the
storage element depopulation feature by specifying the storage element
management operations with the se_ops field of struct
block_device_operations.
Co-developed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
block/blk-zoned.c | 160 +++++++++++++++++++++++++++++++++-
include/linux/blkdev.h | 16 ++++
include/uapi/linux/blkzoned.h | 67 ++++++++++++++
3 files changed, 242 insertions(+), 1 deletion(-)
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index 131c9f50b3da..be57644ee266 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -18,6 +18,7 @@
#include <linux/mempool.h>
#include <linux/kthread.h>
#include <linux/freezer.h>
+#include <linux/delay.h>
#include <trace/events/block.h>
@@ -2718,5 +2719,162 @@ int queue_zone_wplugs_show(void *data, struct seq_file *m)
return 0;
}
-
#endif
+
+static int disk_wait_for_se_mgmt_completion(struct gendisk *disk)
+{
+ struct blk_storage_element *elements, *e;
+ unsigned int i, nr_se, nr_elements = 0;
+ int ret;
+
+ ret = disk->fops->se_ops->report_elements(disk, NULL, &nr_elements);
+ if (ret) {
+ pr_err("Failed to get number of storage elements\n");
+ return ret;
+ }
+
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return -ENOMEM;
+
+ while (1) {
+ /*
+ * Check if we have storage elements being removed or restored.
+ */
+ nr_se = nr_elements;
+ ret = disk->fops->se_ops->report_elements(disk, elements,
+ &nr_se);
+ if (ret) {
+ pr_err("Failed to get storage elements\n");
+ break;
+ }
+
+ e = elements;
+ for (i = 0; i < nr_se; i++, e++) {
+ if (e->status == BLK_SE_STS_REMOVE_IN_PROGRESS ||
+ e->status == BLK_SE_STS_RESTORE_IN_PROGRESS)
+ break;
+ }
+ if (i >= nr_se)
+ break;
+
+ /* Not done yet: wait and retry. */
+ msleep(500);
+ }
+
+ kfree(elements);
+
+ return ret;
+}
+
+/**
+ * bdev_report_storage_elements - report the storage elements of a block device
+ *
+ * Fill at most @nr_elements storage element descriptors in the array @elements.
+ * The number of storage elements filled in the array is returned using
+ * @nr_elements. If @elements is NULL, only @nr_elements is returned.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_report_storage_elements(struct block_device *bdev,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct gendisk *disk = bdev->bd_disk;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ if (!nr_elements)
+ return -EINVAL;
+
+ if (*nr_elements && !elements)
+ return -EINVAL;
+
+ return disk->fops->se_ops->report_elements(disk, elements, nr_elements);
+}
+EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
+
+/**
+ * bdev_remove_storage_element - Remove (depopulate) a storage element of a
+ * block device
+ *
+ * Remove (depopulate) the storage element identified by @element_id from the
+ * block device @bdev. The caller is responsible for taking care of any
+ * necessary device write cache flush and invalidation of cached data for the
+ * zones that will be offlined.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /* Zero is not a valid storage element ID. */
+ if (!element_id)
+ return -EINVAL;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding command.
+ * The caller is responsible for not queuing up more I/Os by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ ret = disk->fops->se_ops->remove_element(disk, element_id);
+ if (ret)
+ return ret;
+
+ /* Revalidate the device zones once the opration completes. */
+ ret = disk_wait_for_se_mgmt_completion(disk);
+ if (ret)
+ return ret;
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
+
+/**
+ * bdev_restore_storage_elements - Restore all depopulated storage elements of a
+ * block device
+ *
+ * Restore all storage elements of @bdev that have been depopulated. Not all
+ * elements may be restored by this operation.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_restore_storage_elements(struct block_device *bdev)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding commands.
+ * The caller is responsible for not queuing up more I/O by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ ret = disk->fops->se_ops->restore_elements(disk);
+ if (ret)
+ return ret;
+
+ ret = disk_wait_for_se_mgmt_completion(disk);
+ if (ret)
+ return ret;
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index d003a9d2d1f6..859917b3b256 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1574,6 +1574,21 @@ enum blk_unique_id {
BLK_UID_NAA = 3,
};
+struct blk_storage_elements_ops {
+ int (*report_elements)(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements);
+ int (*remove_element)(struct gendisk *disk, unsigned int element_id);
+ int (*restore_elements)(struct gendisk *disk);
+};
+
+int bdev_report_storage_elements(struct block_device *bdev,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements);
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id);
+int bdev_restore_storage_elements(struct block_device *bdev);
+
struct block_device_operations {
void (*submit_bio)(struct bio *bio);
int (*poll_bio)(struct bio *bio, struct io_comp_batch *iob,
@@ -1601,6 +1616,7 @@ struct block_device_operations {
enum blk_unique_id id_type);
struct module *owner;
const struct pr_ops *pr_ops;
+ const struct blk_storage_elements_ops *se_ops;
/*
* Special callback for probing GPT entry at a given sector.
diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
index 663836120966..f8f2c2651dff 100644
--- a/include/uapi/linux/blkzoned.h
+++ b/include/uapi/linux/blkzoned.h
@@ -208,4 +208,71 @@ struct blk_zone_range {
#define BLKFINISHZONE _IOW(0x12, 136, struct blk_zone_range)
#define BLKREPORTZONEV2 _IOWR(0x12, 142, struct blk_zone_report)
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_TYPE_RDWR: The storage element handles both reads and writes.
+ * @BLK_SE_TYPE_READ: The storage element handles reads only.
+ * @BLK_SE_TYPE_WRITE: The storage element handles writes only.
+ * @BLK_SE_TYPE_UNKNOWN: The storage element type is not known.
+ */
+enum blk_storage_element_type {
+ BLK_SE_TYPE_RDWR = 0x01,
+ BLK_SE_TYPE_READ = 0x02,
+ BLK_SE_TYPE_WRITE = 0x03,
+ BLK_SE_TYPE_UNKNOWN = 0xFF,
+};
+
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_STS_OK: The storage element is operating normally.
+ * @BLK_SE_STS_DEGRADED: The storage element has degraded and is not operating
+ * normally.
+ * @BLK_SE_STS_REMOVE_IN_PROGRESS: The storage element is being removed.
+ * @BLK_SE_STS_REMOVE_ERROR: The storage element removal failed.
+ * @BLK_SE_STS_RESTORE_IN_PROGRESS: The storage element is being restored.
+ * @BLK_SE_STS_RESTORE_ERROR: The storage element restoration failed.
+ * @BLK_SE_STS_REMOVED: The storage element was removed.
+ * @BLK_SE_STS_UNKNOWN: The storage element status is unknown.
+ */
+enum blk_storage_element_status {
+ BLK_SE_STS_OK = 0x01,
+ BLK_SE_STS_DEGRADED = 0x02,
+ BLK_SE_STS_REMOVE_IN_PROGRESS = 0x03,
+ BLK_SE_STS_REMOVE_ERROR = 0x04,
+ BLK_SE_STS_RESTORE_IN_PROGRESS = 0x05,
+ BLK_SE_STS_RESTORE_ERROR = 0x06,
+ BLK_SE_STS_REMOVED = 0x07,
+ BLK_SE_STS_UNKNOWN = 0xFF,
+};
+
+/**
+ * struct blk_storage_element - Zoned device storage element descriptor.
+ *
+ * @id: The ID of the element (cannot be 0).
+ * @paired_id: The ID of the paired element for an element that is not
+ * of type BLK_SE_TYPE_RDWR.
+ * @type: The type of the storage element (enum blk_storage_element_type).
+ * @status: The health status of the storage element
+ * (enum blk_storage_element_status).
+ * @restore_allowed: Indicate if the storage element can be restored.
+ * @nr_zones: The number of zones that the storage element handles.
+ */
+struct blk_storage_element {
+ __u32 id;
+ __u32 paired_id;
+ __u64 nr_zones;
+ __u8 type;
+ __u8 status;
+ __u8 restore_allowed;
+ __u8 reserved[5];
+};
+
+struct blk_storage_elements_report {
+ __u32 nr_elements;
+ __u32 reserved;
+ struct blk_storage_element elements[];
+};
+
#endif /* _UAPI_BLKZONED_H */
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v2 3/7] block: add storage element management ioctls
2026-10-06 12:45 [PATCH v2 0/7] Add support for storage element depopulation Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 1/7] block: fail reads to offline zones early Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 2/7] block: introduce storage element management Damien Le Moal
@ 2026-10-06 12:45 ` Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 4/7] zloop: add storage element emulation Damien Le Moal
` (3 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Damien Le Moal @ 2026-10-06 12:45 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Define the new ioctl commands BLKGETNRSTORELEMS, BLKREPORTSTORELEMS,
BLKREMOVESTORELEM, and BLKRESTORESTORELEMS to provide users accessing
zoned block devices directly with an interface to the storage elements
management operations of the block layer.
The ioctls BLKGETNRSTORELEMS and BLKREPORTSTORELEMS are defined to
respectively get the number of storage elements of a zoned device and to
get an array of struct blk_storage_element describing the current state of
the device storage elements.
The ioctl BLKREMOVESTORELEM can be used to remove (depopulate) a storage
element that has a degraded status (i.e. BLK_SE_STS_DEGRADED).
Finally, the BLKRESTORESTORELEMS ioctl interfaces with
bdev_restore_storage_elements() to restore the depopulated storage
elements of a zoned device.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
block/blk-zoned.c | 165 ++++++++++++++++++++++++++++++++++
block/blk.h | 9 ++
block/ioctl.c | 5 ++
include/uapi/linux/blkzoned.h | 14 +++
include/uapi/linux/fs.h | 1 +
5 files changed, 194 insertions(+)
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index be57644ee266..6ebb04e0a159 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -19,6 +19,7 @@
#include <linux/kthread.h>
#include <linux/freezer.h>
#include <linux/delay.h>
+#include <linux/uaccess.h>
#include <trace/events/block.h>
@@ -2795,6 +2796,73 @@ int bdev_report_storage_elements(struct block_device *bdev,
}
EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
+static int blkdev_get_nr_storage_elements_ioctl(struct block_device *bdev,
+ void __user *argp)
+{
+ unsigned int nr_elements = 0;
+ int ret;
+
+ ret = bdev_report_storage_elements(bdev, NULL, &nr_elements);
+ if (ret)
+ return ret;
+
+ if (put_user(nr_elements, (unsigned int __user *)argp))
+ return -EFAULT;
+
+ return 0;
+}
+
+static int blkdev_report_storage_elements_ioctl(struct block_device *bdev,
+ void __user *argp)
+{
+ struct blk_storage_elements_report rep;
+ struct blk_storage_element *elements;
+ unsigned int nr_elements = 0;
+ unsigned long retc;
+ int ret;
+
+ if (!argp)
+ return -EINVAL;
+
+ if (copy_from_user(&rep, argp,
+ sizeof(struct blk_storage_elements_report)))
+ return -EFAULT;
+
+ ret = bdev_report_storage_elements(bdev, NULL, &nr_elements);
+ if (ret)
+ return ret;
+
+ nr_elements = min(rep.nr_elements, nr_elements);
+ if (!nr_elements)
+ return -EINVAL;
+
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return -ENOMEM;
+
+ ret = bdev_report_storage_elements(bdev, elements, &nr_elements);
+ if (ret)
+ goto free_elements;
+
+ retc = copy_to_user(argp + sizeof(struct blk_storage_elements_report),
+ elements,
+ sizeof(struct blk_storage_element) * nr_elements);
+ if (retc) {
+ ret = -EFAULT;
+ goto free_elements;
+ }
+
+ rep.nr_elements = nr_elements;
+ retc = copy_to_user(argp, &rep,
+ sizeof(struct blk_storage_elements_report));
+ if (retc)
+ ret = -EFAULT;
+
+free_elements:
+ kfree(elements);
+ return ret;
+}
+
/**
* bdev_remove_storage_element - Remove (depopulate) a storage element of a
* block device
@@ -2841,6 +2909,45 @@ int bdev_remove_storage_element(struct block_device *bdev,
}
EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
+static int blkdev_remove_storage_element_ioctl(struct block_device *bdev,
+ blk_mode_t mode, void __user *argp)
+{
+ unsigned int element_id;
+ int ret;
+
+ if (!(mode & BLK_OPEN_WRITE))
+ return -EBADF;
+ if (bdev_read_only(bdev))
+ return -EPERM;
+
+ if (get_user(element_id, (unsigned int __user *)argp))
+ return -EFAULT;
+
+ /*
+ * Flush the device volatile write cache and invalidate all cached data
+ * so that reads do not return old data for zones that went offline.
+ */
+ inode_lock(bdev->bd_mapping->host);
+ filemap_invalidate_lock(bdev->bd_mapping);
+
+ ret = blkdev_issue_flush(bdev->bd_disk->part0);
+ if (ret)
+ goto out_unlock;
+
+ ret = truncate_bdev_range(bdev, mode, 0,
+ (get_capacity(bdev->bd_disk) << SECTOR_SHIFT) - 1);
+ if (ret)
+ goto out_unlock;
+
+ ret = bdev_remove_storage_element(bdev, element_id);
+
+out_unlock:
+ filemap_invalidate_unlock(bdev->bd_mapping);
+ inode_unlock(bdev->bd_mapping->host);
+
+ return ret;
+}
+
/**
* bdev_restore_storage_elements - Restore all depopulated storage elements of a
* block device
@@ -2878,3 +2985,61 @@ int bdev_restore_storage_elements(struct block_device *bdev)
return blk_revalidate_disk_zones(disk);
}
EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
+
+static int blkdev_restore_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode)
+{
+ int ret;
+
+ if (!(mode & BLK_OPEN_WRITE))
+ return -EBADF;
+ if (bdev_read_only(bdev))
+ return -EPERM;
+
+ /*
+ * Flush the device volatile write cache and invalidate all cached data
+ * so that reads do not return old data for zones that went offline.
+ */
+ inode_lock(bdev->bd_mapping->host);
+ filemap_invalidate_lock(bdev->bd_mapping);
+
+ ret = blkdev_issue_flush(bdev->bd_disk->part0);
+ if (ret)
+ goto out_unlock;
+
+ ret = truncate_bdev_range(bdev, mode, 0,
+ (get_capacity(bdev->bd_disk) << SECTOR_SHIFT) - 1);
+ if (ret)
+ goto out_unlock;
+
+ ret = bdev_restore_storage_elements(bdev);
+
+out_unlock:
+ filemap_invalidate_unlock(bdev->bd_mapping);
+ inode_unlock(bdev->bd_mapping->host);
+
+ return ret;
+}
+
+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp)
+{
+ if (!bdev_is_zoned(bdev) || !bdev->bd_disk->fops->se_ops)
+ return -ENOTTY;
+
+ switch (cmd) {
+ case BLKGETNRSTORELEMS:
+ return blkdev_get_nr_storage_elements_ioctl(bdev, argp);
+ case BLKREPORTSTORELEMS:
+ return blkdev_report_storage_elements_ioctl(bdev, argp);
+ case BLKREMOVESTORELEM:
+ return blkdev_remove_storage_element_ioctl(bdev, mode, argp);
+ case BLKRESTORESTORELEMS:
+ return blkdev_restore_storage_elements_ioctl(bdev, mode);
+ default:
+ break;
+ }
+
+ return -ENOTTY;
+}
diff --git a/block/blk.h b/block/blk.h
index c2d07347a3ad..274afb46a809 100644
--- a/block/blk.h
+++ b/block/blk.h
@@ -579,6 +579,9 @@ int blkdev_zone_mgmt_ioctl(struct block_device *bdev, blk_mode_t mode,
unsigned int cmd, unsigned long arg);
bool bdev_zone_mgmt_allowed(struct block_device *bdev, sector_t sector);
bool bdev_zone_is_offline(struct block_device *bdev, sector_t sector);
+int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp);
#else /* CONFIG_BLK_DEV_ZONED */
static inline void disk_init_zone_resources(struct gendisk *disk)
{
@@ -631,6 +634,12 @@ static inline bool bdev_zone_is_offline(struct block_device *bdev,
{
return false;
}
+static inline int blkdev_zone_storage_elements_ioctl(struct block_device *bdev,
+ blk_mode_t mode, unsigned int cmd,
+ void __user *argp)
+{
+ return -ENOTTY;
+}
#endif /* CONFIG_BLK_DEV_ZONED */
struct block_device *bdev_alloc(struct gendisk *disk, u8 partno);
diff --git a/block/ioctl.c b/block/ioctl.c
index 64b4e6c0f696..ccf806c2e37d 100644
--- a/block/ioctl.c
+++ b/block/ioctl.c
@@ -679,6 +679,11 @@ static int blkdev_common_ioctl(struct block_device *bdev, blk_mode_t mode,
return put_uint(argp, bdev_zone_sectors(bdev));
case BLKGETNRZONES:
return put_uint(argp, bdev_nr_zones(bdev));
+ case BLKGETNRSTORELEMS:
+ case BLKREPORTSTORELEMS:
+ case BLKREMOVESTORELEM:
+ case BLKRESTORESTORELEMS:
+ return blkdev_zone_storage_elements_ioctl(bdev, mode, cmd, argp);
case BLKROGET:
return put_int(argp, bdev_read_only(bdev) != 0);
case BLKSSZGET: /* get block device logical block size */
diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
index f8f2c2651dff..532cf69b5b11 100644
--- a/include/uapi/linux/blkzoned.h
+++ b/include/uapi/linux/blkzoned.h
@@ -275,4 +275,18 @@ struct blk_storage_elements_report {
struct blk_storage_element elements[];
};
+/**
+ * Zoned block device storage element management ioctl's:
+ *
+ * @BLKGETNRSTORELEMS: Get the number of storage elements of the device.
+ * @BLKREPORTSTORELEMS: Get the device storage elements.
+ * @BLKREMOVESTORELEM: Remove (depopulate) one storage element of a device.
+ * @BLKRESTORESTORELEM: Restore (repopulate if possible) all storage elements
+ * that have been removed.
+ */
+#define BLKGETNRSTORELEMS _IOR(0x12, 143, __u32)
+#define BLKREPORTSTORELEMS _IOWR(0x12, 144, struct blk_storage_elements_report)
+#define BLKREMOVESTORELEM _IOW(0x12, 145, __u32)
+#define BLKRESTORESTORELEMS _IO(0x12, 146)
+
#endif /* _UAPI_BLKZONED_H */
diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index 34c6f219462a..8a979326aa7f 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -309,6 +309,7 @@ struct file_attr {
/* 130-136 and 142 are used by zoned block device ioctls (uapi/linux/blkzoned.h) */
/* 137-141 are used by blk-crypto ioctls (uapi/linux/blk-crypto.h) */
#define BLKTRACESETUP2 _IOWR(0x12, 142, struct blk_user_trace_setup2)
+/* 143-146 are used by storage element management for zoned block devices. */
#define BMAP_IOCTL 1 /* obsolete - kept for compatibility */
#define FIBMAP _IO(0x00,1) /* bmap access */
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v2 4/7] zloop: add storage element emulation
2026-10-06 12:45 [PATCH v2 0/7] Add support for storage element depopulation Damien Le Moal
` (2 preceding siblings ...)
2026-10-06 12:45 ` [PATCH v2 3/7] block: add storage element management ioctls Damien Le Moal
@ 2026-10-06 12:45 ` Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 5/7] zloop: add degrade_element control command Damien Le Moal
` (2 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Damien Le Moal @ 2026-10-06 12:45 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Emulate the storage element depopulation feature of SMR disks in the
zloop driver. This emulation is controlled using the new stor_elements
option.
This option can take several values:
- ZLOOP_STOR_ELEMENTS_NONE (0): no emulation (default)
- ZLOOP_STOR_ELEMENTS_RDWR (1): emulate all access storage elements
(e.g. read+write heads)
- ZLOOP_STOR_ELEMENTS_PAIRS (2): emulate fractional access storage
elements (e.g. pairs of read and write heads)
If enabled with the value 1 or 2, the number of storage elements, or of
pairs of fractional access storage elements, is automatically calculated
based on the number of zones of the device so that we have at least 2 and
at most 32 storage elements (or 2 pairs of fractional access storage
elements).
The mapping of zones to storage elements is defined simply as the modulo
of a zone number and a storage element ID. E.g, removing a storage
element from a set of 4 storage elements will offline 1 zone every 4
zones.
The storage element management operations are specified with the
zloop_se_ops (struct blk_storage_elements_ops).
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
.../admin-guide/blockdev/zoned_loop.rst | 8 +-
drivers/block/zloop.c | 567 ++++++++++++++++--
2 files changed, 530 insertions(+), 45 deletions(-)
diff --git a/Documentation/admin-guide/blockdev/zoned_loop.rst b/Documentation/admin-guide/blockdev/zoned_loop.rst
index 64277494fb36..ec39f141af8d 100644
--- a/Documentation/admin-guide/blockdev/zoned_loop.rst
+++ b/Documentation/admin-guide/blockdev/zoned_loop.rst
@@ -61,7 +61,7 @@ The options available for the add command can be listed by reading the
/dev/zloop-control device::
$ cat /dev/zloop-control
- add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache
+ add id=%d,capacity_mb=%u,zone_size_mb=%u,zone_capacity_mb=%u,conv_zones=%u,max_open_zones=%u,base_dir=%s,nr_queues=%u,queue_depth=%u,buffered_io,zone_append=%u,ordered_zone_append,discard_write_cache,stor_elements=%u
remove id=%d
In more details, the options that can be used with the "add" command are as
@@ -113,6 +113,12 @@ discard_write_cache Discard all data that was not explicitly persisted using a
each zone file to the size recorded during the last flush
operation. This simulates power fail events where
uncommitted data is lost.
+stor_elements Control storage element emulation. The default value is 0,
+ indicating no emulation. A value of 1 indicates that all
+ access storage elements (equivalent to read+write head of
+ a disk) are emulated. A value of 2 enables fractional
+ access storage element (equivalent to pairs of read and
+ write heads of a disk) emulation .
=================== =========================================================
3) Deleting a Zoned Device
diff --git a/drivers/block/zloop.c b/drivers/block/zloop.c
index 394dc2408ed7..d6a45f214857 100644
--- a/drivers/block/zloop.c
+++ b/drivers/block/zloop.c
@@ -37,6 +37,7 @@ enum {
ZLOOP_OPT_ORDERED_ZONE_APPEND = (1 << 10),
ZLOOP_OPT_DISCARD_WRITE_CACHE = (1 << 11),
ZLOOP_OPT_MAX_OPEN_ZONES = (1 << 12),
+ ZLOOP_OPT_STOR_ELEMENTS = (1 << 13),
};
static const match_table_t zloop_opt_tokens = {
@@ -51,11 +52,22 @@ static const match_table_t zloop_opt_tokens = {
{ ZLOOP_OPT_BUFFERED_IO, "buffered_io" },
{ ZLOOP_OPT_ZONE_APPEND, "zone_append=%u" },
{ ZLOOP_OPT_ORDERED_ZONE_APPEND, "ordered_zone_append" },
- { ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
+ { ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
{ ZLOOP_OPT_MAX_OPEN_ZONES, "max_open_zones=%u" },
+ { ZLOOP_OPT_STOR_ELEMENTS, "stor_elements=%u" },
{ ZLOOP_OPT_ERR, NULL }
};
+/* Storage elements emulation types. */
+enum zloop_stor_elements {
+ /* No emulation. */
+ ZLOOP_STOR_ELEMENTS_NONE,
+ /* Emulate read+write storage elements. */
+ ZLOOP_STOR_ELEMENTS_RDWR,
+ /* Emulate pairs of associated read and write storage elements. */
+ ZLOOP_STOR_ELEMENTS_PAIRS,
+};
+
/* Default values for the "add" operation. */
#define ZLOOP_DEF_ID -1
#define ZLOOP_DEF_ZONE_SIZE ((256ULL * SZ_1M) >> SECTOR_SHIFT)
@@ -68,6 +80,7 @@ static const match_table_t zloop_opt_tokens = {
#define ZLOOP_DEF_BUFFERED_IO false
#define ZLOOP_DEF_ZONE_APPEND true
#define ZLOOP_DEF_ORDERED_ZONE_APPEND false
+#define ZLOOP_DEF_STOR_ELEMENTS ZLOOP_STOR_ELEMENTS_NONE
/* Arbitrary limit on the zone size (16GB). */
#define ZLOOP_MAX_ZONE_SIZE_MB 16384
@@ -87,6 +100,7 @@ struct zloop_options {
bool zone_append;
bool ordered_zone_append;
bool discard_write_cache;
+ enum zloop_stor_elements stor_elements;
};
/*
@@ -117,6 +131,8 @@ struct zloop_zone {
enum blk_zone_cond cond;
sector_t start;
sector_t wp;
+ unsigned int wr_se_id;
+ unsigned int rd_se_id;
gfp_t old_gfp_mask;
};
@@ -133,6 +149,7 @@ struct zloop_device {
bool zone_append;
bool ordered_zone_append;
bool discard_write_cache;
+ enum zloop_stor_elements stor_elements;
const char *base_dir;
struct file *data_dir;
@@ -150,6 +167,17 @@ struct zloop_device {
struct list_head open_zones_lru_list;
unsigned int nr_open_zones;
+ /* For storage elements emulation. */
+ struct mutex stor_elements_lock;
+ unsigned int nr_elements;
+ unsigned int max_nr_removed_elements;
+ unsigned int nr_removed_elements;
+ struct delayed_work remove_element_work;
+ unsigned int remove_element_id;
+ struct delayed_work restore_elements_work;
+ bool restore_in_progress;
+ struct blk_storage_element *elements;
+
struct zloop_zone zones[] __counted_by(nr_zones);
};
@@ -289,22 +317,30 @@ static bool zloop_do_open_zone(struct zloop_device *zlo,
}
}
-static void zloop_mark_full(struct zloop_device *zlo, struct zloop_zone *zone)
+static void zloop_set_zone_cond(struct zloop_device *zlo,
+ struct zloop_zone *zone,
+ enum blk_zone_cond cond)
{
lockdep_assert_held(&zone->wp_lock);
zloop_lru_remove_open_zone(zlo, zone);
- zone->cond = BLK_ZONE_COND_FULL;
- zone->wp = ULLONG_MAX;
+ zone->cond = cond;
+ if (cond == BLK_ZONE_COND_EMPTY)
+ zone->wp = zone->start;
+ else
+ zone->wp = ULLONG_MAX;
}
-static void zloop_mark_empty(struct zloop_device *zlo, struct zloop_zone *zone)
+static inline void zloop_set_zone_full(struct zloop_device *zlo,
+ struct zloop_zone *zone)
{
- lockdep_assert_held(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_FULL);
+}
- zloop_lru_remove_open_zone(zlo, zone);
- zone->cond = BLK_ZONE_COND_EMPTY;
- zone->wp = zone->start;
+static inline void zloop_set_zone_empty(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_EMPTY);
}
static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
@@ -339,9 +375,9 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
spin_lock(&zone->wp_lock);
if (!file_sectors) {
- zloop_mark_empty(zlo, zone);
+ zloop_set_zone_empty(zlo, zone);
} else if (file_sectors == zlo->zone_capacity) {
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
} else {
if (zone->cond != BLK_ZONE_COND_IMP_OPEN &&
zone->cond != BLK_ZONE_COND_EXP_OPEN)
@@ -353,6 +389,19 @@ static int zloop_update_seq_zone(struct zloop_device *zlo, unsigned int zone_no)
return 0;
}
+static bool zloop_zone_is_offline_or_readonly(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ bool ret;
+
+ spin_lock(&zone->wp_lock);
+ ret = zone->cond == BLK_ZONE_COND_OFFLINE ||
+ zone->cond == BLK_ZONE_COND_READONLY;
+ spin_unlock(&zone->wp_lock);
+
+ return ret;
+}
+
static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)
{
struct zloop_zone *zone = &zlo->zones[zone_no];
@@ -363,6 +412,11 @@ static int zloop_open_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
ret = zloop_update_seq_zone(zlo, zone_no);
if (ret)
@@ -388,6 +442,11 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (test_and_clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags)) {
ret = zloop_update_seq_zone(zlo, zone_no);
if (ret)
@@ -420,7 +479,22 @@ static int zloop_close_zone(struct zloop_device *zlo, unsigned int zone_no)
return ret;
}
-static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)
+static int zloop_do_reset_zone(struct zloop_device *zlo,
+ struct zloop_zone *zone)
+{
+ if (vfs_truncate(&zone->file->f_path, 0))
+ return -EIO;
+
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_empty(zlo, zone);
+ clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
+ spin_unlock(&zone->wp_lock);
+
+ return 0;
+}
+
+static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no,
+ bool all_zones)
{
struct zloop_zone *zone = &zlo->zones[zone_no];
int ret = 0;
@@ -430,20 +504,19 @@ static int zloop_reset_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ if (!all_zones)
+ ret = -EIO;
+ goto unlock;
+ }
+
if (!test_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags) &&
zone->cond == BLK_ZONE_COND_EMPTY)
goto unlock;
- if (vfs_truncate(&zone->file->f_path, 0)) {
+ ret = zloop_do_reset_zone(zlo, zone);
+ if (ret)
set_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
- ret = -EIO;
- goto unlock;
- }
-
- spin_lock(&zone->wp_lock);
- zloop_mark_empty(zlo, zone);
- clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
- spin_unlock(&zone->wp_lock);
unlock:
mutex_unlock(&zone->lock);
@@ -457,7 +530,7 @@ static int zloop_reset_all_zones(struct zloop_device *zlo)
int ret;
for (i = zlo->nr_conv_zones; i < zlo->nr_zones; i++) {
- ret = zloop_reset_zone(zlo, i);
+ ret = zloop_reset_zone(zlo, i, true);
if (ret)
return ret;
}
@@ -475,6 +548,11 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)
mutex_lock(&zone->lock);
+ if (zloop_zone_is_offline_or_readonly(zlo, zone)) {
+ ret = -EIO;
+ goto unlock;
+ }
+
if (!test_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags) &&
zone->cond == BLK_ZONE_COND_FULL)
goto unlock;
@@ -487,7 +565,7 @@ static int zloop_finish_zone(struct zloop_device *zlo, unsigned int zone_no)
}
spin_lock(&zone->wp_lock);
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
clear_bit(ZLOOP_ZONE_SEQ_ERROR, &zone->flags);
spin_unlock(&zone->wp_lock);
@@ -624,7 +702,7 @@ static int zloop_seq_write_prep(struct zloop_cmd *cmd)
if (!is_append || !zlo->ordered_zone_append) {
zone->wp += nr_sectors;
if (zone->wp == zone_end)
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
}
out_unlock:
spin_unlock(&zone->wp_lock);
@@ -766,7 +844,7 @@ static void zloop_handle_cmd(struct zloop_cmd *cmd)
cmd->ret = zloop_flush(zlo);
break;
case REQ_OP_ZONE_RESET:
- cmd->ret = zloop_reset_zone(zlo, rq_zone_no(rq));
+ cmd->ret = zloop_reset_zone(zlo, rq_zone_no(rq), false);
break;
case REQ_OP_ZONE_RESET_ALL:
cmd->ret = zloop_reset_all_zones(zlo);
@@ -875,30 +953,54 @@ static void zloop_complete_rq(struct request *rq)
blk_mq_end_request(rq, sts);
}
-static bool zloop_set_zone_append_sector(struct request *rq)
+static bool zloop_set_zone_append_sector(struct zloop_device *zlo,
+ struct zloop_zone *zone,
+ struct request *rq)
{
- struct zloop_device *zlo = rq->q->queuedata;
- unsigned int zone_no = rq_zone_no(rq);
- struct zloop_zone *zone = &zlo->zones[zone_no];
sector_t zone_end = zone->start + zlo->zone_capacity;
sector_t nr_sectors = blk_rq_sectors(rq);
- spin_lock(&zone->wp_lock);
-
if (zone->cond == BLK_ZONE_COND_FULL ||
- zone->wp + nr_sectors > zone_end) {
- spin_unlock(&zone->wp_lock);
+ zone->wp + nr_sectors > zone_end)
return false;
- }
rq->__sector = zone->wp;
zone->wp += blk_rq_sectors(rq);
if (zone->wp >= zone_end)
- zloop_mark_full(zlo, zone);
+ zloop_set_zone_full(zlo, zone);
+ return true;
+}
+
+
+static bool zloop_prep_rq(struct zloop_device *zlo, struct request *rq)
+{
+ struct zloop_zone *zone = &zlo->zones[rq_zone_no(rq)];
+ bool is_write = op_is_write(req_op(rq));
+ bool ret = true;
+
+ spin_lock(&zone->wp_lock);
+
+ if (zlo->nr_elements) {
+ if (zone->cond == BLK_ZONE_COND_OFFLINE ||
+ (zone->cond == BLK_ZONE_COND_READONLY && is_write)) {
+ ret = false;
+ goto unlock;
+ }
+ }
+
+ /*
+ * If we need to strongly order zone append operations, set the request
+ * sector to the zone write pointer location now instead of when the
+ * command work runs.
+ */
+ if (zlo->ordered_zone_append && req_op(rq) == REQ_OP_ZONE_APPEND)
+ ret = zloop_set_zone_append_sector(zlo, zone, rq);
+
+unlock:
spin_unlock(&zone->wp_lock);
- return true;
+ return ret;
}
static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,
@@ -913,14 +1015,15 @@ static blk_status_t zloop_queue_rq(struct blk_mq_hw_ctx *hctx,
return BLK_STS_IOERR;
}
- /*
- * If we need to strongly order zone append operations, set the request
- * sector to the zone write pointer location now instead of when the
- * command work runs.
- */
- if (zlo->ordered_zone_append && req_op(rq) == REQ_OP_ZONE_APPEND) {
- if (!zloop_set_zone_append_sector(rq))
+ switch (req_op(rq)) {
+ case REQ_OP_READ:
+ case REQ_OP_WRITE:
+ case REQ_OP_ZONE_APPEND:
+ if (!zloop_prep_rq(zlo, rq))
return BLK_STS_IOERR;
+ break;
+ default:
+ break;
}
blk_mq_start_request(rq);
@@ -1002,11 +1105,252 @@ static int zloop_report_zones(struct gendisk *disk, sector_t sector,
return nr_zones;
}
+static int zloop_report_elements(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct zloop_device *zlo = disk->private_data;
+ unsigned int nr_report = *nr_elements;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ *nr_elements = zlo->nr_elements;
+ if (elements) {
+ struct blk_storage_element *se = zlo->elements;
+ unsigned int i;
+
+ for (i = 0; i < min(nr_report, zlo->nr_elements); i++, se++) {
+ switch (se->status) {
+ case BLK_SE_STS_REMOVED:
+ case BLK_SE_STS_RESTORE_ERROR:
+ se->restore_allowed = 1;
+ break;
+ default:
+ se->restore_allowed = 0;
+ }
+ memcpy(&elements[i], se,
+ sizeof(struct blk_storage_element));
+ }
+ }
+
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return 0;
+}
+
+static void zloop_remove_element_work(struct work_struct *work)
+{
+ struct zloop_device *zlo = container_of(work, struct zloop_device,
+ remove_element_work.work);
+ struct blk_storage_element *se, *paired_se = NULL;
+ enum blk_zone_cond cond;
+ struct zloop_zone *zone;
+ unsigned int i;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ /*
+ * If the element to remove is a read element, zones must go offline and
+ * the associated write element is also removed.
+ */
+ se = &zlo->elements[zlo->remove_element_id - 1];
+ switch (se->type) {
+ case BLK_SE_TYPE_RDWR:
+ cond = BLK_ZONE_COND_OFFLINE;
+ break;
+ case BLK_SE_TYPE_READ:
+ cond = BLK_ZONE_COND_OFFLINE;
+ paired_se = &zlo->elements[se->paired_id - 1];
+ break;
+ case BLK_SE_TYPE_WRITE:
+ cond = BLK_ZONE_COND_READONLY;
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ }
+
+ /*
+ * Change the condition of the zones owned by the (pair of) elements
+ * being removed and mark the elements removed.
+ */
+ for (i = 0, zone = zlo->zones; i < zlo->nr_zones; i++, zone++) {
+ if (zone->wr_se_id != se->id && zone->rd_se_id != se->id)
+ continue;
+ mutex_lock(&zone->lock);
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, cond);
+ spin_unlock(&zone->wp_lock);
+ mutex_unlock(&zone->lock);
+ }
+
+ se->status = BLK_SE_STS_REMOVED;
+
+ zlo->nr_removed_elements++;
+ if (paired_se && paired_se->status != BLK_SE_STS_REMOVED) {
+ paired_se->status = BLK_SE_STS_REMOVED;
+ zlo->nr_removed_elements++;
+ }
+
+ zlo->remove_element_id = 0;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+}
+
+static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)
+{
+ struct zloop_device *zlo = disk->private_data;
+ struct blk_storage_element *se, *paired_se = NULL;
+ unsigned int nr_remove = 1;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ if (element_id > zlo->nr_elements)
+ return -EINVAL;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ if (zlo->remove_element_id || zlo->restore_in_progress) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /*
+ * Get the element to remove. If it is already removed, we have nothing
+ * to do.
+ */
+ se = &zlo->elements[element_id - 1];
+ if (se->status == BLK_SE_STS_REMOVED)
+ goto unlock;
+
+ /*
+ * If the element to remove is a read element, the associated write
+ * element must also be removed.
+ */
+ if (se->type == BLK_SE_TYPE_READ) {
+ paired_se = &zlo->elements[se->paired_id - 1];
+ if (paired_se->status != BLK_SE_STS_REMOVED)
+ nr_remove = 2;
+ }
+
+ if (zlo->nr_removed_elements + nr_remove >
+ zlo->max_nr_removed_elements) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /*
+ * Schedule the element removal with a delay, to emulate the (generally
+ * short) time it takes for a real device to depopulate a head and
+ * modify the zones.
+ */
+ zlo->remove_element_id = element_id;
+ se->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+ if (paired_se && paired_se->status != BLK_SE_STS_REMOVED)
+ paired_se->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+
+ schedule_delayed_work(&zlo->remove_element_work,
+ msecs_to_jiffies(2000));
+
+unlock:
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
+static void zloop_restore_elements_work(struct work_struct *work)
+{
+ struct zloop_device *zlo = container_of(work, struct zloop_device,
+ restore_elements_work.work);
+ struct zloop_zone *zone = zlo->zones;
+ struct blk_storage_element *se;
+ unsigned int i;
+ int ret = 0;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ /* Reset all zones. */
+ for (i = 0; i < zlo->nr_zones && ret == 0; i++, zone++) {
+ mutex_lock(&zone->lock);
+ if (test_bit(ZLOOP_ZONE_CONV, &zone->flags)) {
+ spin_lock(&zone->wp_lock);
+ zloop_set_zone_cond(zlo, zone, BLK_ZONE_COND_NOT_WP);
+ spin_unlock(&zone->wp_lock);
+ } else {
+ ret = zloop_do_reset_zone(zlo, zone);
+ }
+ mutex_unlock(&zone->lock);
+ }
+
+ /* Restore all removed elements. */
+ for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
+ if (se->status != BLK_SE_STS_RESTORE_IN_PROGRESS)
+ continue;
+ if (!ret)
+ se->status = BLK_SE_STS_OK;
+ else
+ se->status = BLK_SE_STS_RESTORE_ERROR;
+ }
+
+ zlo->nr_removed_elements = 0;
+ zlo->restore_in_progress = false;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+}
+
+static int zloop_restore_elements(struct gendisk *disk)
+{
+ struct zloop_device *zlo = disk->private_data;
+ struct blk_storage_element *se;
+ unsigned int i;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ if (zlo->remove_element_id || zlo->restore_in_progress) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ /* If we have no removed elements, we have nothing to do. */
+ if (!zlo->nr_removed_elements)
+ goto unlock;
+
+ /*
+ * Schedule the elements restoration with a delay, to emulate the time
+ * it takes for a real device to restore all removed heads and reset
+ * all zones.
+ */
+ zlo->restore_in_progress = true;
+ for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
+ if (se->status == BLK_SE_STS_REMOVED)
+ se->status = BLK_SE_STS_RESTORE_IN_PROGRESS;
+ }
+
+ schedule_delayed_work(&zlo->restore_elements_work,
+ msecs_to_jiffies(5000));
+
+unlock:
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
static void zloop_free_disk(struct gendisk *disk)
{
struct zloop_device *zlo = disk->private_data;
unsigned int i;
+ cancel_delayed_work_sync(&zlo->remove_element_work);
+ cancel_delayed_work_sync(&zlo->restore_elements_work);
+
blk_mq_free_tag_set(&zlo->tag_set);
for (i = 0; i < zlo->nr_zones; i++) {
@@ -1019,15 +1363,24 @@ static void zloop_free_disk(struct gendisk *disk)
fput(zlo->data_dir);
destroy_workqueue(zlo->workqueue);
+ kfree(zlo->elements);
kfree(zlo->base_dir);
kvfree(zlo);
}
+
+static const struct blk_storage_elements_ops zloop_se_ops = {
+ .report_elements = zloop_report_elements,
+ .remove_element = zloop_remove_element,
+ .restore_elements = zloop_restore_elements,
+};
+
static const struct block_device_operations zloop_fops = {
.owner = THIS_MODULE,
.open = zloop_open,
.report_zones = zloop_report_zones,
.free_disk = zloop_free_disk,
+ .se_ops = &zloop_se_ops,
};
__printf(3, 4)
@@ -1112,6 +1465,24 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,
if (!opts->buffered_io)
oflags |= O_DIRECT;
+ if (zlo->stor_elements != ZLOOP_STOR_ELEMENTS_NONE) {
+ unsigned int nr_elems = zlo->nr_elements;
+ unsigned int se_idx;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ nr_elems /= 2;
+ se_idx = zone_no % nr_elems;
+ zlo->elements[se_idx].nr_zones++;
+
+ zone->wr_se_id = se_idx + 1;
+ zone->rd_se_id = zone->wr_se_id;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {
+ zone->rd_se_id += nr_elems;
+ zlo->elements[zone->rd_se_id - 1].nr_zones =
+ zlo->elements[se_idx].nr_zones;
+ }
+ }
+
if (zone_no < zlo->nr_conv_zones) {
/* Conventional zone file. */
set_bit(ZLOOP_ZONE_CONV, &zone->flags);
@@ -1182,6 +1553,79 @@ static int zloop_init_zone(struct zloop_device *zlo, struct zloop_options *opts,
return ret;
}
+#define ZLOOP_MIN_STOR_ELEMENTS 2
+#define ZLOOP_MAX_STOR_ELEMENTS 32
+#define ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS 32
+
+static int zloop_create_storage_elements(struct zloop_device *zlo)
+{
+ struct blk_storage_element *se, *paired_se;
+ unsigned int i, nr_elems, nr_elements;
+
+ /*
+ * Calculate the number of storage elements we are going to emulate.
+ * To achieve a somewhat realistic emulation, we want at least 2 storage
+ * elements, and no more than 32, targeting at least 32 zones per
+ * element.
+ */
+ if (zlo->nr_zones <= 64)
+ nr_elements = 2;
+ else
+ nr_elements =
+ min(ZLOOP_MAX_STOR_ELEMENTS,
+ zlo->nr_zones / ZLOOP_MIN_ZONES_PER_STOR_ELEMENTS);
+
+ /*
+ * If we are emulating pairs of read and write storage elements, we need
+ * double the number of storage element descriptors.
+ */
+ nr_elems = nr_elements;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ nr_elems *= 2;
+ zlo->elements = kzalloc_objs(struct blk_storage_element, nr_elems);
+ if (!zlo->elements)
+ return -ENOMEM;
+
+ for (i = 0, se = zlo->elements; i < nr_elements; i++, se++) {
+ se->id = i + 1;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ se->type = BLK_SE_TYPE_WRITE;
+ else
+ se->type = BLK_SE_TYPE_RDWR;
+ se->status = BLK_SE_STS_OK;
+ }
+
+ /*
+ * If we are emulating pairs of read and write storage elements,
+ * initialize the read elements paired with the write elements we just
+ * initialized.
+ */
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS) {
+ for (i = 0; i < nr_elements; i++, se++) {
+ se->id = nr_elements + i + 1;
+ paired_se = &zlo->elements[i];
+ se->paired_id = paired_se->id;
+ paired_se->paired_id = se->id;
+ se->type = BLK_SE_TYPE_READ;
+ se->status = BLK_SE_STS_OK;
+ }
+ }
+
+ /*
+ * Make sure we do not allow removing all storage elements as that does
+ * not make any sense. This is consistent with the device advertized
+ * limit of SCSI and ATA devices supporting the storage element
+ * depopulation feature.
+ */
+ zlo->nr_elements = nr_elems;
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_PAIRS)
+ zlo->max_nr_removed_elements = zlo->nr_elements - 2;
+ else
+ zlo->max_nr_removed_elements = zlo->nr_elements - 1;
+
+ return 0;
+}
+
static bool zloop_dev_exists(struct zloop_device *zlo)
{
struct file *cnv, *seq;
@@ -1237,6 +1681,10 @@ static int zloop_ctl_add(struct zloop_options *opts)
WRITE_ONCE(zlo->state, Zlo_creating);
spin_lock_init(&zlo->open_zones_lock);
INIT_LIST_HEAD(&zlo->open_zones_lru_list);
+ mutex_init(&zlo->stor_elements_lock);
+ INIT_DELAYED_WORK(&zlo->remove_element_work, zloop_remove_element_work);
+ INIT_DELAYED_WORK(&zlo->restore_elements_work,
+ zloop_restore_elements_work);
ret = mutex_lock_killable(&zloop_ctl_mutex);
if (ret)
@@ -1270,12 +1718,19 @@ static int zloop_ctl_add(struct zloop_options *opts)
if (zlo->zone_append)
zlo->ordered_zone_append = opts->ordered_zone_append;
zlo->discard_write_cache = opts->discard_write_cache;
+ zlo->stor_elements = opts->stor_elements;
+
+ if (zlo->stor_elements != ZLOOP_STOR_ELEMENTS_NONE) {
+ ret = zloop_create_storage_elements(zlo);
+ if (ret)
+ goto out_free_idr;
+ }
zlo->workqueue = alloc_workqueue("zloop%d", WQ_UNBOUND | WQ_FREEZABLE,
opts->nr_queues * opts->queue_depth, zlo->id);
if (!zlo->workqueue) {
ret = -ENOMEM;
- goto out_free_idr;
+ goto out_destroy_storage_elements;
}
if (opts->base_dir)
@@ -1361,6 +1816,10 @@ static int zloop_ctl_add(struct zloop_options *opts)
zlo->id, zlo->nr_zones,
((sector_t)zlo->zone_size << SECTOR_SHIFT) >> 20,
zlo->block_size);
+ if (zlo->nr_elements)
+ pr_info("zloop%d: %d storage elements\n",
+ zlo->id, zlo->nr_elements);
+
pr_info("zloop%d: using %s%s zone append\n",
zlo->id,
zlo->ordered_zone_append ? "ordered " : "",
@@ -1384,6 +1843,8 @@ static int zloop_ctl_add(struct zloop_options *opts)
kfree(zlo->base_dir);
out_destroy_workqueue:
destroy_workqueue(zlo->workqueue);
+out_destroy_storage_elements:
+ kfree(zlo->elements);
out_free_idr:
mutex_lock(&zloop_ctl_mutex);
idr_remove(&zloop_index_idr, zlo->id);
@@ -1502,6 +1963,7 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
opts->buffered_io = ZLOOP_DEF_BUFFERED_IO;
opts->zone_append = ZLOOP_DEF_ZONE_APPEND;
opts->ordered_zone_append = ZLOOP_DEF_ORDERED_ZONE_APPEND;
+ opts->stor_elements = ZLOOP_DEF_STOR_ELEMENTS;
if (!buf)
return 0;
@@ -1636,6 +2098,23 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
case ZLOOP_OPT_DISCARD_WRITE_CACHE:
opts->discard_write_cache = true;
break;
+ case ZLOOP_OPT_STOR_ELEMENTS:
+ if (match_uint(args, &token)) {
+ ret = -EINVAL;
+ goto out;
+ }
+ switch (token) {
+ case ZLOOP_STOR_ELEMENTS_NONE:
+ case ZLOOP_STOR_ELEMENTS_RDWR:
+ case ZLOOP_STOR_ELEMENTS_PAIRS:
+ break;
+ default:
+ pr_err("Invalid stor_elements value\n");
+ ret = -EINVAL;
+ goto out;
+ }
+ opts->stor_elements = token;
+ break;
case ZLOOP_OPT_ERR:
default:
pr_warn("unknown parameter or missing value '%s'\n", p);
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v2 5/7] zloop: add degrade_element control command
2026-10-06 12:45 [PATCH v2 0/7] Add support for storage element depopulation Damien Le Moal
` (3 preceding siblings ...)
2026-10-06 12:45 ` [PATCH v2 4/7] zloop: add storage element emulation Damien Le Moal
@ 2026-10-06 12:45 ` Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
6 siblings, 0 replies; 8+ messages in thread
From: Damien Le Moal @ 2026-10-06 12:45 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Allow users to mark storage elements of a zloop device as degraded using
the new "degrade_element" control command. The element to degrade is
indicated using the element_id option. Example:
echo "degrade_element id=0,element_id=2" > /dev/zloop-control
If the element ID identifies an all access storage element, read and write
operations targeting a zone served by the degraded lement are failed.
For a partial access storage element, read or write operations are failed
depending on the storage element type. This check for an element status
is done without taking the storage elements mutex lock, so in order to
avoid memory access ordering issues, storage element status access is
changed to use READ_ONCE() and WRITE_ONCE()
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
drivers/block/zloop.c | 131 ++++++++++++++++++++++++++++++++++++------
1 file changed, 115 insertions(+), 16 deletions(-)
diff --git a/drivers/block/zloop.c b/drivers/block/zloop.c
index d6a45f214857..3913bd229910 100644
--- a/drivers/block/zloop.c
+++ b/drivers/block/zloop.c
@@ -38,6 +38,7 @@ enum {
ZLOOP_OPT_DISCARD_WRITE_CACHE = (1 << 11),
ZLOOP_OPT_MAX_OPEN_ZONES = (1 << 12),
ZLOOP_OPT_STOR_ELEMENTS = (1 << 13),
+ ZLOOP_OPT_ELEMENT_ID = (1 << 14),
};
static const match_table_t zloop_opt_tokens = {
@@ -55,6 +56,7 @@ static const match_table_t zloop_opt_tokens = {
{ ZLOOP_OPT_DISCARD_WRITE_CACHE, "discard_write_cache" },
{ ZLOOP_OPT_MAX_OPEN_ZONES, "max_open_zones=%u" },
{ ZLOOP_OPT_STOR_ELEMENTS, "stor_elements=%u" },
+ { ZLOOP_OPT_ELEMENT_ID, "element_id=%u" },
{ ZLOOP_OPT_ERR, NULL }
};
@@ -81,6 +83,7 @@ enum zloop_stor_elements {
#define ZLOOP_DEF_ZONE_APPEND true
#define ZLOOP_DEF_ORDERED_ZONE_APPEND false
#define ZLOOP_DEF_STOR_ELEMENTS ZLOOP_STOR_ELEMENTS_NONE
+#define ZLOOP_DEF_ELEMENT_ID 0
/* Arbitrary limit on the zone size (16GB). */
#define ZLOOP_MAX_ZONE_SIZE_MB 16384
@@ -101,6 +104,7 @@ struct zloop_options {
bool ordered_zone_append;
bool discard_write_cache;
enum zloop_stor_elements stor_elements;
+ unsigned int element_id;
};
/*
@@ -982,11 +986,26 @@ static bool zloop_prep_rq(struct zloop_device *zlo, struct request *rq)
spin_lock(&zone->wp_lock);
if (zlo->nr_elements) {
+ struct blk_storage_element *se;
+
if (zone->cond == BLK_ZONE_COND_OFFLINE ||
(zone->cond == BLK_ZONE_COND_READONLY && is_write)) {
ret = false;
goto unlock;
}
+
+ /*
+ * Check the health state of the storage element serving the
+ * zone.
+ */
+ if (is_write)
+ se = &zlo->elements[zone->wr_se_id - 1];
+ else
+ se = &zlo->elements[zone->rd_se_id - 1];
+ if (READ_ONCE(se->status) == BLK_SE_STS_DEGRADED) {
+ ret = false;
+ goto unlock;
+ }
}
/*
@@ -1123,7 +1142,7 @@ static int zloop_report_elements(struct gendisk *disk,
unsigned int i;
for (i = 0; i < min(nr_report, zlo->nr_elements); i++, se++) {
- switch (se->status) {
+ switch (READ_ONCE(se->status)) {
case BLK_SE_STS_REMOVED:
case BLK_SE_STS_RESTORE_ERROR:
se->restore_allowed = 1;
@@ -1141,6 +1160,31 @@ static int zloop_report_elements(struct gendisk *disk,
return 0;
}
+static int zloop_degrade_element(struct zloop_device *zlo,
+ unsigned int element_id)
+{
+ struct blk_storage_element *se;
+ int ret = 0;
+
+ if (zlo->stor_elements == ZLOOP_STOR_ELEMENTS_NONE)
+ return -EOPNOTSUPP;
+
+ if (!element_id || element_id > zlo->nr_elements)
+ return -EINVAL;
+
+ mutex_lock(&zlo->stor_elements_lock);
+
+ se = &zlo->elements[element_id - 1];
+ if (READ_ONCE(se->status) == BLK_SE_STS_OK)
+ WRITE_ONCE(se->status, BLK_SE_STS_DEGRADED);
+ else
+ ret = -EINVAL;
+
+ mutex_unlock(&zlo->stor_elements_lock);
+
+ return ret;
+}
+
static void zloop_remove_element_work(struct work_struct *work)
{
struct zloop_device *zlo = container_of(work, struct zloop_device,
@@ -1186,11 +1230,11 @@ static void zloop_remove_element_work(struct work_struct *work)
mutex_unlock(&zone->lock);
}
- se->status = BLK_SE_STS_REMOVED;
+ WRITE_ONCE(se->status, BLK_SE_STS_REMOVED);
zlo->nr_removed_elements++;
- if (paired_se && paired_se->status != BLK_SE_STS_REMOVED) {
- paired_se->status = BLK_SE_STS_REMOVED;
+ if (paired_se && READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED) {
+ WRITE_ONCE(paired_se->status, BLK_SE_STS_REMOVED);
zlo->nr_removed_elements++;
}
@@ -1224,7 +1268,7 @@ static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)
* to do.
*/
se = &zlo->elements[element_id - 1];
- if (se->status == BLK_SE_STS_REMOVED)
+ if (READ_ONCE(se->status) == BLK_SE_STS_REMOVED)
goto unlock;
/*
@@ -1233,7 +1277,7 @@ static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)
*/
if (se->type == BLK_SE_TYPE_READ) {
paired_se = &zlo->elements[se->paired_id - 1];
- if (paired_se->status != BLK_SE_STS_REMOVED)
+ if (READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED)
nr_remove = 2;
}
@@ -1249,9 +1293,9 @@ static int zloop_remove_element(struct gendisk *disk, unsigned int element_id)
* modify the zones.
*/
zlo->remove_element_id = element_id;
- se->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
- if (paired_se && paired_se->status != BLK_SE_STS_REMOVED)
- paired_se->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+ WRITE_ONCE(se->status, BLK_SE_STS_REMOVE_IN_PROGRESS);
+ if (paired_se && READ_ONCE(paired_se->status) != BLK_SE_STS_REMOVED)
+ WRITE_ONCE(paired_se->status, BLK_SE_STS_REMOVE_IN_PROGRESS);
schedule_delayed_work(&zlo->remove_element_work,
msecs_to_jiffies(2000));
@@ -1288,12 +1332,12 @@ static void zloop_restore_elements_work(struct work_struct *work)
/* Restore all removed elements. */
for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
- if (se->status != BLK_SE_STS_RESTORE_IN_PROGRESS)
+ if (READ_ONCE(se->status) != BLK_SE_STS_RESTORE_IN_PROGRESS)
continue;
if (!ret)
- se->status = BLK_SE_STS_OK;
+ WRITE_ONCE(se->status, BLK_SE_STS_OK);
else
- se->status = BLK_SE_STS_RESTORE_ERROR;
+ WRITE_ONCE(se->status, BLK_SE_STS_RESTORE_ERROR);
}
zlo->nr_removed_elements = 0;
@@ -1330,8 +1374,8 @@ static int zloop_restore_elements(struct gendisk *disk)
*/
zlo->restore_in_progress = true;
for (i = 0, se = zlo->elements; i < zlo->nr_elements; i++, se++) {
- if (se->status == BLK_SE_STS_REMOVED)
- se->status = BLK_SE_STS_RESTORE_IN_PROGRESS;
+ if (READ_ONCE(se->status) == BLK_SE_STS_REMOVED)
+ WRITE_ONCE(se->status, BLK_SE_STS_RESTORE_IN_PROGRESS);
}
schedule_delayed_work(&zlo->restore_elements_work,
@@ -1944,6 +1988,42 @@ static int zloop_ctl_remove(struct zloop_options *opts)
return 0;
}
+static int zloop_ctl_degrade_element(struct zloop_options *opts)
+{
+ struct zloop_device *zlo;
+ int ret = 0;
+
+ if (!(opts->mask & ZLOOP_OPT_ID)) {
+ pr_err("No ID specified for degrade_element\n");
+ return -EINVAL;
+ }
+
+ if (opts->mask & ~(ZLOOP_OPT_ID | ZLOOP_OPT_ELEMENT_ID)) {
+ pr_err("Invalid option specified for degrade_element\n");
+ return -EINVAL;
+ }
+
+ mutex_lock(&zloop_ctl_mutex);
+
+ zlo = idr_find(&zloop_index_idr, opts->id);
+ if (!zlo || zlo->state == Zlo_creating)
+ ret = -ENODEV;
+ else if (zlo->state == Zlo_deleting)
+ ret = -EINVAL;
+ if (ret)
+ goto unlock;
+
+ ret = zloop_degrade_element(zlo, opts->element_id);
+ if (!ret)
+ pr_info("Degraded element %u of device %u\n",
+ opts->id, opts->element_id);
+
+unlock:
+ mutex_unlock(&zloop_ctl_mutex);
+
+ return ret;
+}
+
static int zloop_parse_options(struct zloop_options *opts, const char *buf)
{
substring_t args[MAX_OPT_ARGS];
@@ -1964,6 +2044,7 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
opts->zone_append = ZLOOP_DEF_ZONE_APPEND;
opts->ordered_zone_append = ZLOOP_DEF_ORDERED_ZONE_APPEND;
opts->stor_elements = ZLOOP_DEF_STOR_ELEMENTS;
+ opts->element_id = ZLOOP_DEF_ELEMENT_ID;
if (!buf)
return 0;
@@ -2115,6 +2196,13 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
}
opts->stor_elements = token;
break;
+ case ZLOOP_OPT_ELEMENT_ID:
+ if (match_uint(args, &token)) {
+ ret = -EINVAL;
+ goto out;
+ }
+ opts->element_id = token;
+ break;
case ZLOOP_OPT_ERR:
default:
pr_warn("unknown parameter or missing value '%s'\n", p);
@@ -2143,14 +2231,16 @@ static int zloop_parse_options(struct zloop_options *opts, const char *buf)
enum {
ZLOOP_CTL_ADD,
ZLOOP_CTL_REMOVE,
+ ZLOOP_CTL_DEGRADE_ELEMENT,
};
static struct zloop_ctl_op {
int code;
const char *name;
} zloop_ctl_ops[] = {
- { ZLOOP_CTL_ADD, "add" },
- { ZLOOP_CTL_REMOVE, "remove" },
+ { ZLOOP_CTL_ADD, "add" },
+ { ZLOOP_CTL_REMOVE, "remove" },
+ { ZLOOP_CTL_DEGRADE_ELEMENT, "degrade_element" },
{ -1, NULL },
};
@@ -2198,6 +2288,9 @@ static ssize_t zloop_ctl_write(struct file *file, const char __user *ubuf,
case ZLOOP_CTL_REMOVE:
ret = zloop_ctl_remove(&opts);
break;
+ case ZLOOP_CTL_DEGRADE_ELEMENT:
+ ret = zloop_ctl_degrade_element(&opts);
+ break;
default:
pr_err("Invalid operation\n");
ret = -EINVAL;
@@ -2221,6 +2314,8 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)
tok = &zloop_opt_tokens[i];
if (!tok->pattern)
break;
+ if (tok->token == ZLOOP_OPT_ELEMENT_ID)
+ continue;
if (i)
seq_putc(seq_file, ',');
seq_puts(seq_file, tok->pattern);
@@ -2231,6 +2326,10 @@ static int zloop_ctl_show(struct seq_file *seq_file, void *private)
seq_puts(seq_file, zloop_ctl_ops[1].name);
seq_puts(seq_file, " id=%d\n");
+ /* Degrade element operation */
+ seq_puts(seq_file, zloop_ctl_ops[2].name);
+ seq_puts(seq_file, " id=%d,element_id=%d\n");
+
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v2 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation
2026-10-06 12:45 [PATCH v2 0/7] Add support for storage element depopulation Damien Le Moal
` (4 preceding siblings ...)
2026-10-06 12:45 ` [PATCH v2 5/7] zloop: add degrade_element control command Damien Le Moal
@ 2026-10-06 12:45 ` Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
6 siblings, 0 replies; 8+ messages in thread
From: Damien Le Moal @ 2026-10-06 12:45 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
sd_zbc_revalidate_zones() skips revalidating the zones of a ZBC device if
the zone size and total number of zones of the disk has not changed. This
is to avoid a call to the rather slow blk_revalidate_disk_zones().
However, for ZBC devices that support data preserving head depopulation
(REMOVE ELEMENT AND MODIFY ZONES command), a disk capacity and number of
zones does not change after a head is depopulated but the condition of
zones changes as the zones served by the head that was depopulated become
either read-only or offline. In this case, not calling
blk_revalidate_disk_zones() prevents the block layer from taking
appropriate actions on the zone write plugs of the disk for the zones that
became read-only or offline.
Avoid any issue with the block layer view of the zone conditions by not
skipping the call to blk_revalidate_disk_zones() for disks that support
the REMOVE ELEMENT AND MODIFY ZONES command. This check is done from
sd_zbc_read_zones() using the helper function sd_zbc_check_modify_zones().
The new scsi disk flag modify_zones_supported is defined to remember the
result of this check.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
drivers/scsi/sd.h | 1 +
drivers/scsi/sd_zbc.c | 26 +++++++++++++++++++++++++-
2 files changed, 26 insertions(+), 1 deletion(-)
diff --git a/drivers/scsi/sd.h b/drivers/scsi/sd.h
index 574af8243016..6a72371fee78 100644
--- a/drivers/scsi/sd.h
+++ b/drivers/scsi/sd.h
@@ -156,6 +156,7 @@ struct scsi_disk {
unsigned ignore_medium_access_errors : 1;
unsigned rscs : 1; /* reduced stream control support */
unsigned use_atomic_write_boundary : 1;
+ unsigned modify_zones_supported : 1;
};
#define to_scsi_disk(obj) container_of(obj, struct scsi_disk, disk_dev)
diff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c
index 456beaf2e769..2c77f878ab3f 100644
--- a/drivers/scsi/sd_zbc.c
+++ b/drivers/scsi/sd_zbc.c
@@ -516,6 +516,21 @@ static int sd_zbc_check_capacity(struct scsi_disk *sdkp, unsigned char *buf,
return 0;
}
+/*
+ * sd_zbc_check_modify_zones - Check if the device supports depopulation
+ * @sdkp: Target disk
+ * @buf: command buffer
+ *
+ * Check if the device supports the REMOVE ELEMENT AND MODIFY ZONES command.
+ */
+static inline bool sd_zbc_check_modify_zones(struct scsi_disk *sdkp,
+ unsigned char *buf)
+{
+ return scsi_report_opcode(sdkp->device, buf, SD_BUF_SIZE,
+ SERVICE_ACTION_IN_16,
+ SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES) == 1;
+}
+
static void sd_zbc_print_zones(struct scsi_disk *sdkp)
{
if (sdkp->device->type != TYPE_ZBC || !sdkp->capacity)
@@ -554,9 +569,15 @@ int sd_zbc_revalidate_zones(struct scsi_disk *sdkp)
if (!blk_queue_is_zoned(q))
return 0;
+ /*
+ * If the zone size and number of zones has not changed, and the disk
+ * does not support depopulating heads, skip the rather slow call to
+ * blk_revalidate_disk_zones().
+ */
if (sdkp->zone_info.zone_blocks == zone_blocks &&
sdkp->zone_info.nr_zones == nr_zones &&
- disk->nr_zones == nr_zones)
+ disk->nr_zones == nr_zones &&
+ !sdkp->modify_zones_supported)
return 0;
sdkp->zone_info.zone_blocks = zone_blocks;
@@ -620,6 +641,9 @@ int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim,
if (ret != 0)
goto err;
+ /* Check if REMOVE ELEMENT AND MODIFY ZONES is supported. */
+ sdkp->modify_zones_supported = sd_zbc_check_modify_zones(sdkp, buf);
+
nr_zones = round_up(sdkp->capacity, zone_blocks) >> ilog2(zone_blocks);
if (nr_zones > INT_MAX) {
sd_printk(KERN_ERR, sdkp, "Too many zones (%llu)\n",
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v2 7/7] scsi: sd_zbc: define storage element management operations
2026-10-06 12:45 [PATCH v2 0/7] Add support for storage element depopulation Damien Le Moal
` (5 preceding siblings ...)
2026-10-06 12:45 ` [PATCH v2 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation Damien Le Moal
@ 2026-10-06 12:45 ` Damien Le Moal
6 siblings, 0 replies; 8+ messages in thread
From: Damien Le Moal @ 2026-10-06 12:45 UTC (permalink / raw)
To: Jens Axboe, linux-block, Christoph Hellwig, linux-scsi,
Martin K . Petersen
Define the storage element management operations using struct
blk_storage_elements_ops. These operations are valid only on SMR disks
supporting the storage element depopulation feature.
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
drivers/scsi/sd.c | 3 +
drivers/scsi/sd.h | 2 +
drivers/scsi/sd_zbc.c | 233 ++++++++++++++++++++++++++++++++++++++++++
3 files changed, 238 insertions(+)
diff --git a/drivers/scsi/sd.c b/drivers/scsi/sd.c
index a1b21ea14e54..4372fa80f792 100644
--- a/drivers/scsi/sd.c
+++ b/drivers/scsi/sd.c
@@ -3937,6 +3937,9 @@ static const struct block_device_operations sd_fops = {
.get_unique_id = sd_get_unique_id,
.free_disk = scsi_disk_free_disk,
.pr_ops = &sd_pr_ops,
+#ifdef CONFIG_BLK_DEV_ZONED
+ .se_ops = &sd_zbc_se_ops,
+#endif
};
/**
diff --git a/drivers/scsi/sd.h b/drivers/scsi/sd.h
index 6a72371fee78..b68b7ee5a5fa 100644
--- a/drivers/scsi/sd.h
+++ b/drivers/scsi/sd.h
@@ -243,6 +243,8 @@ unsigned int sd_zbc_complete(struct scsi_cmnd *cmd, unsigned int good_bytes,
int sd_zbc_report_zones(struct gendisk *disk, sector_t sector,
unsigned int nr_zones, struct blk_report_zones_args *args);
+extern const struct blk_storage_elements_ops sd_zbc_se_ops;
+
#else /* CONFIG_BLK_DEV_ZONED */
static inline int sd_zbc_read_zones(struct scsi_disk *sdkp,
diff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c
index 2c77f878ab3f..6698d156a0c3 100644
--- a/drivers/scsi/sd_zbc.c
+++ b/drivers/scsi/sd_zbc.c
@@ -548,6 +548,239 @@ static void sd_zbc_print_zones(struct scsi_disk *sdkp)
sdkp->zone_info.zone_blocks);
}
+static void sd_zbc_parse_storage_element(struct scsi_disk *sdkp, u8 *desc,
+ struct blk_storage_element *element)
+{
+ struct scsi_device *sdp = sdkp->device;
+ sector_t zone_sectors = sd_zbc_zone_sectors(sdkp);
+ u64 capacity;
+
+ memset(element, 0, sizeof(*element));
+
+ element->id = get_unaligned_be32(&desc[4]);
+
+ switch (desc[14]) {
+ case SCSI_PHYS_ELEM_TYPE_ALL_ACCESS_STORAGE:
+ element->type = BLK_SE_TYPE_RDWR;
+ capacity = get_unaligned_be64(&desc[16]);
+ if (!zone_sectors || capacity == ULLONG_MAX)
+ element->nr_zones = 0;
+ else
+ element->nr_zones =
+ logical_to_sectors(sdp, capacity) >>
+ ilog2(zone_sectors);
+ break;
+ case SCSI_PHYS_ELEM_TYPE_FRAC_ACCESS_STORAGE:
+ element->paired_id = get_unaligned_be32(&desc[16]);
+ if (desc[20] & 0x01)
+ element->type = BLK_SE_TYPE_READ;
+ else
+ element->type = BLK_SE_TYPE_WRITE;
+ element->nr_zones = get_unaligned_be64(&desc[24]);
+ break;
+ default:
+ element->type = BLK_SE_TYPE_UNKNOWN;
+ }
+
+ switch (desc[15]) {
+ case SCSI_PHYS_ELEM_HEALTH_WITHIN_SPEC_LIMITS:
+ case SCSI_PHYS_ELEM_HEALTH_AT_SPEC_LIMITS:
+ element->status = BLK_SE_STS_OK;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_OUTSIDE_SPEC_LIMITS:
+ element->status = BLK_SE_STS_DEGRADED;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_ERR:
+ element->status = BLK_SE_STS_RESTORE_ERROR;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_REVOKE_IN_PROGRESS:
+ element->status = BLK_SE_STS_RESTORE_IN_PROGRESS;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_ERR:
+ element->status = BLK_SE_STS_REMOVE_ERROR;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_IN_PROGRESS:
+ element->status = BLK_SE_STS_REMOVE_IN_PROGRESS;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_DEPOP_OK:
+ element->status = BLK_SE_STS_REMOVED;
+ element->restore_allowed = desc[13] & 0x01;
+ break;
+ case SCSI_PHYS_ELEM_HEALTH_NOT_REPORTED:
+ default:
+ element->status = BLK_SE_STS_UNKNOWN;
+ break;
+ }
+}
+
+/*
+ * Large hard limit on the number of storage elements. This accommodates all
+ * known devices today and likely forever :)
+ */
+#define SD_ZBC_MAX_STORAGE_ELEMENTS 255
+
+static int sd_zbc_report_storage_elements(struct gendisk *disk,
+ struct blk_storage_element *elements,
+ unsigned int *nr_elements)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ unsigned int nr_descs, nr_se;
+ unsigned int buf_size;
+ int i, ret = 0, result;
+ u8 *desc, *buf;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ /*
+ * We need at least 32B for the report header and 32B for each
+ * descriptor.
+ */
+ nr_se = min(SD_ZBC_MAX_STORAGE_ELEMENTS, *nr_elements);
+ buf_size = ALIGN((nr_se + 1) * 32, SECTOR_SIZE);
+
+ buf = kzalloc(buf_size, GFP_KERNEL);
+ if (!buf)
+ return -ENOMEM;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_GET_PHYSICAL_ELEMENT_STATUS;
+ put_unaligned_be32(buf_size, &cmd[10]);
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, buf, buf_size,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "GET PHYSICAL ELEMENT STATUS failed\n");
+ sd_print_result(sdkp, "GET PHYSICAL ELEMENT STATUS", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ nr_descs = get_unaligned_be32(&buf[0]);
+ if (!nr_descs) {
+ sd_printk(KERN_ERR, sdkp,
+ "Invalid number of phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+ if (nr_descs > SD_ZBC_MAX_STORAGE_ELEMENTS) {
+ sd_printk(KERN_ERR, sdkp,
+ "Unsupported number of phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ if (!elements) {
+ *nr_elements = nr_descs;
+ goto free_buf;
+ }
+
+ nr_descs = get_unaligned_be32(&buf[4]);
+ if (!nr_descs) {
+ sd_printk(KERN_ERR, sdkp,
+ "Invalid number of reported phys element descriptors\n");
+ ret = -EIO;
+ goto free_buf;
+ }
+
+ desc = &buf[32];
+ for (i = 0; i < min(nr_se, nr_descs); i++, desc += 32)
+ sd_zbc_parse_storage_element(sdkp, desc, &elements[i]);
+ *nr_elements = i;
+
+free_buf:
+ kfree(buf);
+
+ return ret;
+}
+
+static int sd_zbc_remove_storage_element(struct gendisk *disk,
+ unsigned int element_id)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ int result;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_REMOVE_ELEMENT_AND_MODIFY_ZONES;
+ put_unaligned_be32(element_id, &cmd[10]);
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "REMOVE ELEMENT AND MODIFY ZONES failed\n");
+ sd_print_result(sdkp,
+ "REMOVE ELEMENT AND MODIFY ZONES", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+static int sd_zbc_restore_storage_elements(struct gendisk *disk)
+{
+ struct scsi_disk *sdkp = scsi_disk(disk);
+ struct scsi_device *sdp = sdkp->device;
+ const int timeout = sdp->request_queue->rq_timeout;
+ struct scsi_sense_hdr sshdr;
+ const struct scsi_exec_args exec_args = {
+ .sshdr = &sshdr,
+ };
+ unsigned char cmd[16];
+ int result;
+
+ if (!sdkp->modify_zones_supported)
+ return -EOPNOTSUPP;
+
+ memset(cmd, 0, 16);
+ cmd[0] = SERVICE_ACTION_IN_16;
+ cmd[1] = SAI_RESTORE_ELEMENTS_AND_REBUILD;
+
+ result = scsi_execute_cmd(sdp, cmd, REQ_OP_DRV_IN, NULL, 0,
+ timeout, 1, &exec_args);
+ if (result) {
+ sd_printk(KERN_ERR, sdkp,
+ "RESTORE ELEMENTS AND REBUILD failed\n");
+ sd_print_result(sdkp,
+ "RESTORE ELEMENTS AND REBUILD", result);
+ if (result > 0 && scsi_sense_valid(&sshdr))
+ sd_print_sense_hdr(sdkp, &sshdr);
+ return -EIO;
+ }
+
+ return 0;
+}
+
+const struct blk_storage_elements_ops sd_zbc_se_ops = {
+ .report_elements = sd_zbc_report_storage_elements,
+ .remove_element = sd_zbc_remove_storage_element,
+ .restore_elements = sd_zbc_restore_storage_elements,
+};
+
/*
* Call blk_revalidate_disk_zones() if any of the zoned disk properties have
* changed that make it necessary to call that function. Called by
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread