From: Damien Le Moal <dlemoal@kernel.org>
To: Jens Axboe <axboe@kernel.dk>,
linux-block@vger.kernel.org, Christoph Hellwig <hch@lst.de>,
linux-scsi@vger.kernel.org,
"Martin K . Petersen" <martin.petersen@oracle.com>
Subject: [PATCH 2/7] block: introduce storage element management
Date: Mon, 5 Oct 2026 18:46:27 +0900 [thread overview]
Message-ID: <20261005094632.580753-3-dlemoal@kernel.org> (raw)
In-Reply-To: <20261005094632.580753-1-dlemoal@kernel.org>
Recent SCSI (SBC) and ATA (ACS) standards define the storage element
depopulation feature. This feature is intended for managing hard-disks
heads, either combined read-write heads or pairs of read and write heads,
allowing to keep disks with defective heads longer in production by
allowing "depopulating" (removing) defective heads.
This feature comes in two different flavors:
- A destructive version which removes a head and reformats the disk at a
lower capacity point, restoaring a fully functional contiguous LBA
address space.
- A data preserving version restricted to host-managed zoned disks, which
marks the zones served by a removed head as offline or read-only.
In preparation for using the data-preserving flavor of the depopulation
feature in file systems natively supporting zoned block devices, introduce
a set of storage element management operations and functions to define
generic calls into block device drivers for managing storage elements.
The set of operations is defined with struct blk_storage_elements_ops and
includes three operations:
- report_elements: get information on a device storage elements state
- remove_element: depopulate a defective storage element
- restore_elements: restore depopulated storage elements
Each operation is called from the functions
bdev_report_storage_elements(), bdev_remove_storage_element() and
bdev_restore_storage_elements().
Removing a healthy storage element from a device is possible and useful
for testing. The storage element restoration operation restore_elements
allows repopulating such healthy element. Repopulating defective storage
elements is generally not allowed by devices.
The device drivers of zoned block devices can indicate support for the
storage element depopulation feature by specifying the storage element
management operations with the se_ops field of struct
block_device_operations.
Co-developed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
block/blk-zoned.c | 168 +++++++++++++++++++++++++++++++++-
include/linux/blkdev.h | 16 ++++
include/uapi/linux/blkzoned.h | 67 ++++++++++++++
3 files changed, 250 insertions(+), 1 deletion(-)
diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index 131c9f50b3da..b713512f7d48 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -18,6 +18,7 @@
#include <linux/mempool.h>
#include <linux/kthread.h>
#include <linux/freezer.h>
+#include <linux/delay.h>
#include <trace/events/block.h>
@@ -2718,5 +2719,170 @@ int queue_zone_wplugs_show(void *data, struct seq_file *m)
return 0;
}
-
#endif
+
+static void disk_wait_for_se_mgmt_completion(struct gendisk *disk)
+{
+ struct blk_storage_element *elements, *e;
+ unsigned int i, nr_se, nr_elements = 0;
+ int ret;
+
+ ret = (disk->fops->se_ops->report_elements)(disk, &nr_elements, NULL);
+ if (ret) {
+ pr_err("Failed to get number of storage elements\n");
+ return;
+ }
+
+ elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+ if (!elements)
+ return;
+
+ while (1) {
+ /*
+ * Check if we have storage elements being removed or restored.
+ */
+ nr_se = nr_elements;
+ ret = (disk->fops->se_ops->report_elements)(disk,
+ &nr_elements, elements);
+ if (ret) {
+ pr_err("Failed to get storage elements\n");
+ break;
+ }
+
+ e = elements;
+ for (i = 0; i < nr_se; i++, e++) {
+ if (e->status == BLK_SE_STS_REMOVE_IN_PROGRESS ||
+ e->status == BLK_SE_STS_RESTORE_IN_PROGRESS)
+ break;
+ }
+ if (i >= nr_se)
+ break;
+
+ /* Not done yet: wait and retry. */
+ msleep(500);
+ }
+
+ kfree(elements);
+}
+
+/**
+ * bdev_report_storage_elements - report the storage elements of a block device
+ *
+ * Fill at most @nr_elements storage element descriptors in the array @elements.
+ * The number of storage elements filled in the array is returned using
+ * @nr_elements. If @elements is NULL, only @nr_elements is returned.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_report_storage_elements(struct block_device *bdev,
+ unsigned int *nr_elements,
+ struct blk_storage_element *elements)
+{
+ struct gendisk *disk = bdev->bd_disk;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ if (!nr_elements)
+ return -EINVAL;
+
+ if (*nr_elements && !elements)
+ return -EINVAL;
+
+ return (disk->fops->se_ops->report_elements)(disk, nr_elements,
+ elements);
+}
+EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
+
+/**
+ * bdev_remove_storage_element - Remove (depopulate) a storage element of a
+ * block device
+ *
+ * Remove (depopulate) the storage element identified by @element_id from the
+ * block device @bdev.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /* Zero is not a valid storage element ID. */
+ if (!element_id)
+ return -EINVAL;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding command.
+ * The caller is responsible for not queuing up more I/Os by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ /*
+ * Flush the volatile write cache if there is any. Note that this may
+ * lead to errors if degraded storage elements are used. So ignore
+ * errors here.
+ */
+ blkdev_issue_flush(disk->part0);
+
+ /* Invalidate all cached data. */
+ invalidate_inode_pages2_range(bdev->bd_mapping, 0,
+ get_capacity(disk) << SECTOR_SHIFT);
+
+ ret = (disk->fops->se_ops->remove_element)(disk, element_id);
+ if (ret)
+ return ret;
+
+ /* Revalidate the device zones once the opration completes. */
+ disk_wait_for_se_mgmt_completion(disk);
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
+
+/**
+ * bdev_restore_storage_elements - Restore all depopulated storage elements of a
+ * block device
+ *
+ * Restore all storage elements of @bdev that have been depopulated. Not all
+ * elements may be restored by this operation.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_restore_storage_elements(struct block_device *bdev)
+{
+ struct gendisk *disk = bdev->bd_disk;
+ unsigned int memflags;
+ int ret;
+
+ if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+ return -EOPNOTSUPP;
+
+ /*
+ * Freeze and unfreeze the queue to flush any outstanding commands.
+ * The caller is responsible for not queuing up more I/O by higher
+ * level means.
+ */
+ memflags = blk_mq_freeze_queue(disk->queue);
+ blk_mq_unfreeze_queue(disk->queue, memflags);
+
+ /* Invalidate all cached data. */
+ invalidate_inode_pages2_range(bdev->bd_mapping, 0,
+ get_capacity(disk) << SECTOR_SHIFT);
+
+ ret = (disk->fops->se_ops->restore_elements)(disk);
+ if (ret)
+ return ret;
+
+ disk_wait_for_se_mgmt_completion(disk);
+
+ return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index d003a9d2d1f6..0b2013c96cad 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1574,6 +1574,21 @@ enum blk_unique_id {
BLK_UID_NAA = 3,
};
+struct blk_storage_elements_ops {
+ int (*report_elements)(struct gendisk *disk,
+ unsigned int *nr_elements,
+ struct blk_storage_element *elements);
+ int (*remove_element)(struct gendisk *disk, unsigned int element_id);
+ int (*restore_elements)(struct gendisk *disk);
+};
+
+int bdev_report_storage_elements(struct block_device *bdev,
+ unsigned int *nr_elements,
+ struct blk_storage_element *elements);
+int bdev_remove_storage_element(struct block_device *bdev,
+ unsigned int element_id);
+int bdev_restore_storage_elements(struct block_device *bdev);
+
struct block_device_operations {
void (*submit_bio)(struct bio *bio);
int (*poll_bio)(struct bio *bio, struct io_comp_batch *iob,
@@ -1601,6 +1616,7 @@ struct block_device_operations {
enum blk_unique_id id_type);
struct module *owner;
const struct pr_ops *pr_ops;
+ const struct blk_storage_elements_ops *se_ops;
/*
* Special callback for probing GPT entry at a given sector.
diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
index 663836120966..3a4bacfe7979 100644
--- a/include/uapi/linux/blkzoned.h
+++ b/include/uapi/linux/blkzoned.h
@@ -208,4 +208,71 @@ struct blk_zone_range {
#define BLKFINISHZONE _IOW(0x12, 136, struct blk_zone_range)
#define BLKREPORTZONEV2 _IOWR(0x12, 142, struct blk_zone_report)
+/**
+ * enum blk_storage_element_status - Statuc of a zoned device storage elements.
+ *
+ * @BLK_SE_RDWR: The storage element handles both reads and writes.
+ * @BLK_SE_READ: The storage element handles reads only.
+ * @BLK_SE_WRITE: The storage element handles writes only.
+ * @BLK_SE_WRITE: The storage element type is unknown.
+ */
+enum blk_storage_element_type {
+ BLK_SE_TYPE_RDWR = 0x01,
+ BLK_SE_TYPE_READ = 0x02,
+ BLK_SE_TYPE_WRITE = 0x03,
+ BLK_SE_TYPE_UNKNOWN = 0xFF,
+};
+
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_STS_OK: The storage element is operating normally.
+ * @BLK_SE_STS_DEGRADED: The storage element has degraded and is not operating
+ * normally.
+ * @BLK_SE_STS_REMOVE_IN_PROGRESS: The storage element is being removed.
+ * @BLK_SE_STS_REMOVE_ERROR: The storage element removal failed.
+ * @BLK_SE_STS_RESTORE_IN_PROGRESS: The storage element is being restored.
+ * @BLK_SE_STS_RESTORE_ERROR: The storage element restoration failed.
+ * @BLK_SE_STS_REMOVED: The storage element was removed.
+ * @BLK_SE_STS_UNKNOWN: The storage element status is unknown.
+ */
+enum blk_storage_element_status {
+ BLK_SE_STS_OK = 0x01,
+ BLK_SE_STS_DEGRADED = 0x02,
+ BLK_SE_STS_REMOVE_IN_PROGRESS = 0x03,
+ BLK_SE_STS_REMOVE_ERROR = 0x04,
+ BLK_SE_STS_RESTORE_IN_PROGRESS = 0x05,
+ BLK_SE_STS_RESTORE_ERROR = 0x06,
+ BLK_SE_STS_REMOVED = 0x07,
+ BLK_SE_STS_UNKOWN = 0xFF,
+};
+
+/**
+ * struct blk_storage_element - Zoned device storage element descriptor.
+ *
+ * @id: The ID of the element (cannot be 0).
+ * @paired_id: The ID of the paired element for an element that is not
+ * of type BLK_SE_TYPE_RDWR.
+ * @type: The type of the storage element (enum blk_storage_element_type).
+ * @status: The health status of the storage element
+ * (enum blk_storage_element_status).
+ * @restore_allowed: Indicate if the storage element can be restored.
+ * @nr_zones: The number of zones that the storage element handles.
+ */
+struct blk_storage_element {
+ __u32 id;
+ __u32 paired_id;
+ __u64 nr_zones;
+ __u8 type;
+ __u8 status;
+ __u8 restore_allowed;
+ __u8 reserved[5];
+};
+
+struct blk_storage_elements_report {
+ __u32 nr_elements;
+ __u32 reserved;
+ struct blk_storage_element elements[];
+};
+
#endif /* _UAPI_BLKZONED_H */
--
2.55.0
next prev parent reply other threads:[~2026-10-05 9:46 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 9:46 [PATCH 0/7] Add support for storage element depopulation Damien Le Moal
2026-10-05 9:46 ` [PATCH 1/7] block: fail reads to offline zones early Damien Le Moal
2026-10-05 10:01 ` sashiko-bot
2026-10-05 10:45 ` Hannes Reinecke
2026-10-05 9:46 ` Damien Le Moal [this message]
2026-10-05 10:01 ` [PATCH 2/7] block: introduce storage element management sashiko-bot
2026-10-05 10:52 ` Hannes Reinecke
2026-10-05 22:15 ` kernel test robot
2026-10-05 9:46 ` [PATCH 3/7] block: add storage element management ioctls Damien Le Moal
2026-10-05 10:00 ` sashiko-bot
2026-10-05 11:09 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 4/7] zloop: add storage element emulation Damien Le Moal
2026-10-05 9:58 ` sashiko-bot
2026-10-05 11:14 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 5/7] zloop: add degrade_element control command Damien Le Moal
2026-10-05 9:58 ` sashiko-bot
2026-10-05 11:17 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation Damien Le Moal
2026-10-05 11:19 ` Hannes Reinecke
2026-10-05 9:46 ` [PATCH 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal
2026-10-05 9:59 ` sashiko-bot
2026-10-05 11:48 ` Hannes Reinecke
2026-10-05 20:48 ` kernel test robot
2026-10-05 21:41 ` kernel test robot
2026-10-05 11:13 ` [PATCH 0/7] Add support for storage element depopulation Hannes Reinecke
2026-10-07 7:14 ` Damien Le Moal
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261005094632.580753-3-dlemoal@kernel.org \
--to=dlemoal@kernel.org \
--cc=axboe@kernel.dk \
--cc=hch@lst.de \
--cc=linux-block@vger.kernel.org \
--cc=linux-scsi@vger.kernel.org \
--cc=martin.petersen@oracle.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox