Linux SCSI subsystem development
 help / color / mirror / Atom feed
From: Damien Le Moal <dlemoal@kernel.org>
To: Jens Axboe <axboe@kernel.dk>,
	linux-block@vger.kernel.org, Christoph Hellwig <hch@lst.de>,
	linux-scsi@vger.kernel.org,
	"Martin K . Petersen" <martin.petersen@oracle.com>
Subject: [PATCH v2 2/7] block: introduce storage element management
Date: Tue,  6 Oct 2026 21:45:05 +0900	[thread overview]
Message-ID: <20261006124510.882017-3-dlemoal@kernel.org> (raw)
In-Reply-To: <20261006124510.882017-1-dlemoal@kernel.org>

Recent SCSI (SBC) and ATA (ACS) standards define the storage element
depopulation feature. This feature is intended for managing hard-disks
heads, either combined read-write heads or pairs of read and write heads,
allowing to keep disks with defective heads longer in production by
allowing "depopulating" (removing) defective heads.

This feature comes in two different flavors:
 - A destructive version which removes a head and reformats the disk at a
   lower capacity point, restoaring a fully functional contiguous LBA
   address space.
 - A data preserving version restricted to host-managed zoned disks, which
   marks the zones served by a removed head as offline or read-only.

In preparation for using the data-preserving flavor of the depopulation
feature in file systems natively supporting zoned block devices, introduce
a set of storage element management operations and functions to define
generic calls into block device drivers for managing storage elements.

The set of operations is defined with struct blk_storage_elements_ops and
includes three operations:
 - report_elements: get information on a device storage elements state
 - remove_element: depopulate a defective storage element
 - restore_elements: restore depopulated storage elements

Each operation is called from the functions
bdev_report_storage_elements(), bdev_remove_storage_element() and
bdev_restore_storage_elements().

Removing a healthy storage element from a device is possible and useful
for testing. The storage element restoration operation restore_elements
allows repopulating such healthy element. Repopulating defective storage
elements is generally not allowed by devices.

The device drivers of zoned block devices can indicate support for the
storage element depopulation feature by specifying the storage element
management operations with the se_ops field of struct
block_device_operations.

Co-developed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Damien Le Moal <dlemoal@kernel.org>
---
 block/blk-zoned.c             | 160 +++++++++++++++++++++++++++++++++-
 include/linux/blkdev.h        |  16 ++++
 include/uapi/linux/blkzoned.h |  67 ++++++++++++++
 3 files changed, 242 insertions(+), 1 deletion(-)

diff --git a/block/blk-zoned.c b/block/blk-zoned.c
index 131c9f50b3da..be57644ee266 100644
--- a/block/blk-zoned.c
+++ b/block/blk-zoned.c
@@ -18,6 +18,7 @@
 #include <linux/mempool.h>
 #include <linux/kthread.h>
 #include <linux/freezer.h>
+#include <linux/delay.h>
 
 #include <trace/events/block.h>
 
@@ -2718,5 +2719,162 @@ int queue_zone_wplugs_show(void *data, struct seq_file *m)
 
 	return 0;
 }
-
 #endif
+
+static int disk_wait_for_se_mgmt_completion(struct gendisk *disk)
+{
+	struct blk_storage_element *elements, *e;
+	unsigned int i, nr_se, nr_elements = 0;
+	int ret;
+
+	ret = disk->fops->se_ops->report_elements(disk, NULL, &nr_elements);
+	if (ret) {
+		pr_err("Failed to get number of storage elements\n");
+		return ret;
+	}
+
+	elements = kzalloc_objs(struct blk_storage_element, nr_elements);
+	if (!elements)
+		return -ENOMEM;
+
+	while (1) {
+		/*
+		 * Check if we have storage elements being removed or restored.
+		 */
+		nr_se = nr_elements;
+		ret = disk->fops->se_ops->report_elements(disk, elements,
+							  &nr_se);
+		if (ret) {
+			pr_err("Failed to get storage elements\n");
+			break;
+		}
+
+		e = elements;
+		for (i = 0; i < nr_se; i++, e++) {
+			if (e->status == BLK_SE_STS_REMOVE_IN_PROGRESS ||
+			    e->status == BLK_SE_STS_RESTORE_IN_PROGRESS)
+				break;
+		}
+		if (i >= nr_se)
+			break;
+
+		/* Not done yet: wait and retry. */
+		msleep(500);
+	}
+
+	kfree(elements);
+
+	return ret;
+}
+
+/**
+ * bdev_report_storage_elements - report the storage elements of a block device
+ *
+ * Fill at most @nr_elements storage element descriptors in the array @elements.
+ * The number of storage elements filled in the array is returned using
+ * @nr_elements. If @elements is NULL, only @nr_elements is returned.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_report_storage_elements(struct block_device *bdev,
+				 struct blk_storage_element *elements,
+				 unsigned int *nr_elements)
+{
+	struct gendisk *disk = bdev->bd_disk;
+
+	if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+		return -EOPNOTSUPP;
+
+	if (!nr_elements)
+		return -EINVAL;
+
+	if (*nr_elements && !elements)
+		return -EINVAL;
+
+	return disk->fops->se_ops->report_elements(disk, elements, nr_elements);
+}
+EXPORT_SYMBOL_GPL(bdev_report_storage_elements);
+
+/**
+ * bdev_remove_storage_element - Remove (depopulate) a storage element of a
+ *				 block device
+ *
+ * Remove (depopulate) the storage element identified by @element_id from the
+ * block device @bdev. The caller is responsible for taking care of any
+ * necessary device write cache flush and invalidation of cached data for the
+ * zones that will be offlined.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_remove_storage_element(struct block_device *bdev,
+				unsigned int element_id)
+{
+	struct gendisk *disk = bdev->bd_disk;
+	unsigned int memflags;
+	int ret;
+
+	if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+		return -EOPNOTSUPP;
+
+	/* Zero is not a valid storage element ID. */
+	if (!element_id)
+		return -EINVAL;
+
+	/*
+	 * Freeze and unfreeze the queue to flush any outstanding command.
+	 * The caller is responsible for not queuing up more I/Os by higher
+	 * level means.
+	 */
+	memflags = blk_mq_freeze_queue(disk->queue);
+	blk_mq_unfreeze_queue(disk->queue, memflags);
+
+	ret = disk->fops->se_ops->remove_element(disk, element_id);
+	if (ret)
+		return ret;
+
+	/* Revalidate the device zones once the opration completes. */
+	ret = disk_wait_for_se_mgmt_completion(disk);
+	if (ret)
+		return ret;
+
+	return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_remove_storage_element);
+
+/**
+ * bdev_restore_storage_elements - Restore all depopulated storage elements of a
+ *				   block device
+ *
+ * Restore all storage elements of @bdev that have been depopulated. Not all
+ * elements may be restored by this operation.
+ *
+ * Returns 0 on success and a negative error code on failure.
+ */
+int bdev_restore_storage_elements(struct block_device *bdev)
+{
+	struct gendisk *disk = bdev->bd_disk;
+	unsigned int memflags;
+	int ret;
+
+	if (!bdev_is_zoned(bdev) || !disk->fops->se_ops)
+		return -EOPNOTSUPP;
+
+	/*
+	 * Freeze and unfreeze the queue to flush any outstanding commands.
+	 * The caller is responsible for not queuing up more I/O by higher
+	 * level means.
+	 */
+	memflags = blk_mq_freeze_queue(disk->queue);
+	blk_mq_unfreeze_queue(disk->queue, memflags);
+
+	ret = disk->fops->se_ops->restore_elements(disk);
+	if (ret)
+		return ret;
+
+	ret = disk_wait_for_se_mgmt_completion(disk);
+	if (ret)
+		return ret;
+
+	return blk_revalidate_disk_zones(disk);
+}
+EXPORT_SYMBOL_GPL(bdev_restore_storage_elements);
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index d003a9d2d1f6..859917b3b256 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1574,6 +1574,21 @@ enum blk_unique_id {
 	BLK_UID_NAA	= 3,
 };
 
+struct blk_storage_elements_ops {
+	int (*report_elements)(struct gendisk *disk,
+			       struct blk_storage_element *elements,
+			       unsigned int *nr_elements);
+	int (*remove_element)(struct gendisk *disk, unsigned int element_id);
+	int (*restore_elements)(struct gendisk *disk);
+};
+
+int bdev_report_storage_elements(struct block_device *bdev,
+				 struct blk_storage_element *elements,
+				 unsigned int *nr_elements);
+int bdev_remove_storage_element(struct block_device *bdev,
+				unsigned int element_id);
+int bdev_restore_storage_elements(struct block_device *bdev);
+
 struct block_device_operations {
 	void (*submit_bio)(struct bio *bio);
 	int (*poll_bio)(struct bio *bio, struct io_comp_batch *iob,
@@ -1601,6 +1616,7 @@ struct block_device_operations {
 			enum blk_unique_id id_type);
 	struct module *owner;
 	const struct pr_ops *pr_ops;
+	const struct blk_storage_elements_ops	*se_ops;
 
 	/*
 	 * Special callback for probing GPT entry at a given sector.
diff --git a/include/uapi/linux/blkzoned.h b/include/uapi/linux/blkzoned.h
index 663836120966..f8f2c2651dff 100644
--- a/include/uapi/linux/blkzoned.h
+++ b/include/uapi/linux/blkzoned.h
@@ -208,4 +208,71 @@ struct blk_zone_range {
 #define BLKFINISHZONE	_IOW(0x12, 136, struct blk_zone_range)
 #define BLKREPORTZONEV2	_IOWR(0x12, 142, struct blk_zone_report)
 
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_TYPE_RDWR: The storage element handles both reads and writes.
+ * @BLK_SE_TYPE_READ: The storage element handles reads only.
+ * @BLK_SE_TYPE_WRITE: The storage element handles writes only.
+ * @BLK_SE_TYPE_UNKNOWN: The storage element type is not known.
+ */
+enum blk_storage_element_type {
+	BLK_SE_TYPE_RDWR		= 0x01,
+	BLK_SE_TYPE_READ		= 0x02,
+	BLK_SE_TYPE_WRITE		= 0x03,
+	BLK_SE_TYPE_UNKNOWN		= 0xFF,
+};
+
+/**
+ * enum blk_storage_element_status - Status of a zoned device storage elements.
+ *
+ * @BLK_SE_STS_OK: The storage element is operating normally.
+ * @BLK_SE_STS_DEGRADED: The storage element has degraded and is not operating
+ *			 normally.
+ * @BLK_SE_STS_REMOVE_IN_PROGRESS: The storage element is being removed.
+ * @BLK_SE_STS_REMOVE_ERROR: The storage element removal failed.
+ * @BLK_SE_STS_RESTORE_IN_PROGRESS: The storage element is being restored.
+ * @BLK_SE_STS_RESTORE_ERROR: The storage element restoration failed.
+ * @BLK_SE_STS_REMOVED: The storage element was removed.
+ * @BLK_SE_STS_UNKNOWN: The storage element status is unknown.
+ */
+enum blk_storage_element_status {
+	BLK_SE_STS_OK			= 0x01,
+	BLK_SE_STS_DEGRADED		= 0x02,
+	BLK_SE_STS_REMOVE_IN_PROGRESS	= 0x03,
+	BLK_SE_STS_REMOVE_ERROR		= 0x04,
+	BLK_SE_STS_RESTORE_IN_PROGRESS	= 0x05,
+	BLK_SE_STS_RESTORE_ERROR	= 0x06,
+	BLK_SE_STS_REMOVED		= 0x07,
+	BLK_SE_STS_UNKNOWN		= 0xFF,
+};
+
+/**
+ * struct blk_storage_element - Zoned device storage element descriptor.
+ *
+ * @id: The ID of the element (cannot be 0).
+ * @paired_id: The ID of the paired element for an element that is not
+ *	       of type BLK_SE_TYPE_RDWR.
+ * @type: The type of the storage element (enum blk_storage_element_type).
+ * @status: The health status of the storage element
+ *	    (enum blk_storage_element_status).
+ * @restore_allowed: Indicate if the storage element can be restored.
+ * @nr_zones: The number of zones that the storage element handles.
+ */
+struct blk_storage_element {
+	__u32	id;
+	__u32	paired_id;
+	__u64	nr_zones;
+	__u8	type;
+	__u8	status;
+	__u8	restore_allowed;
+	__u8	reserved[5];
+};
+
+struct blk_storage_elements_report {
+	__u32				nr_elements;
+	__u32				reserved;
+	struct blk_storage_element	elements[];
+};
+
 #endif /* _UAPI_BLKZONED_H */
-- 
2.55.0


  parent reply	other threads:[~2026-10-06 12:45 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-06 12:45 [PATCH v2 0/7] Add support for storage element depopulation Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 1/7] block: fail reads to offline zones early Damien Le Moal
2026-10-06 12:45 ` Damien Le Moal [this message]
2026-10-06 13:00   ` [PATCH v2 2/7] block: introduce storage element management sashiko-bot
2026-10-06 12:45 ` [PATCH v2 3/7] block: add storage element management ioctls Damien Le Moal
2026-10-06 13:02   ` sashiko-bot
2026-10-06 12:45 ` [PATCH v2 4/7] zloop: add storage element emulation Damien Le Moal
2026-10-06 12:56   ` sashiko-bot
2026-10-06 12:45 ` [PATCH v2 5/7] zloop: add degrade_element control command Damien Le Moal
2026-10-06 12:56   ` sashiko-bot
2026-10-06 13:12     ` Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 6/7] scsi: sd_zbc: always revalidate zones for disks supporting head depopulation Damien Le Moal
2026-10-06 12:45 ` [PATCH v2 7/7] scsi: sd_zbc: define storage element management operations Damien Le Moal

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261006124510.882017-3-dlemoal@kernel.org \
    --to=dlemoal@kernel.org \
    --cc=axboe@kernel.dk \
    --cc=hch@lst.de \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-scsi@vger.kernel.org \
    --cc=martin.petersen@oracle.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox