From: Nilay Shroff <nilay@linux.ibm.com>
To: Zheng Qixing <zhengqixing@huaweicloud.com>, axboe@kernel.dk
Cc: linux-block@vger.kernel.org, linux-kernel@vger.kernel.org,
yukuai3@huawei.com, yi.zhang@huawei.com, yangerkun@huawei.com,
houtao1@huawei.com, zhengqixing@huawei.com
Subject: Re: [PATCH] block: fix kobject double initialization in add_disk
Date: Thu, 7 Aug 2025 17:17:05 +0530 [thread overview]
Message-ID: <470ab442-e5eb-4fa8-bde7-d6d2d1115a5a@linux.ibm.com> (raw)
In-Reply-To: <20250807072056.2627592-1-zhengqixing@huaweicloud.com>
On 8/7/25 12:50 PM, Zheng Qixing wrote:
> From: Zheng Qixing <zhengqixing@huawei.com>
>
> Device-mapper can call add_disk() multiple times for the same gendisk
> due to its two-phase creation process (dm create + dm load). This leads
> to kobject double initialization errors when the underlying iSCSI devices
> become temporarily unavailable and then reappear.
>
> However, if the first add_disk() call fails and is retried, the queue_kobj
> gets initialized twice, causing:
>
> kobject: kobject (ffff88810c27bb90): tried to init an initialized object,
> something is seriously wrong.
> Call Trace:
> <TASK>
> dump_stack_lvl+0x5b/0x80
> kobject_init.cold+0x43/0x51
> blk_register_queue+0x46/0x280
> add_disk_fwnode+0xb5/0x280
> dm_setup_md_queue+0x194/0x1c0
> table_load+0x297/0x2d0
> ctl_ioctl+0x2a2/0x480
> dm_ctl_ioctl+0xe/0x20
> __x64_sys_ioctl+0xc7/0x110
> do_syscall_64+0x72/0x390
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
>
> Fix this by separating kobject initialization from sysfs registration:
> - Initialize queue_kobj early during gendisk allocation
> - add_disk() only adds the already-initialized kobject to sysfs
> - del_gendisk() removes from sysfs but doesn't destroy the kobject
> - Final cleanup happens when the disk is released
>
> Fixes: 2bd85221a625 ("block: untangle request_queue refcounting from sysfs")
> Reported-by: Li Lingfeng <lilingfeng3@huawei.com>
> Closes: https://lore.kernel.org/all/83591d0b-2467-433c-bce0-5581298eb161@huawei.com/
> Signed-off-by: Zheng Qixing <zhengqixing@huawei.com>
> ---
> block/blk-sysfs.c | 4 +---
> block/blk.h | 1 +
> block/genhd.c | 2 ++
> 3 files changed, 4 insertions(+), 3 deletions(-)
>
> diff --git a/block/blk-sysfs.c b/block/blk-sysfs.c
> index 396cded255ea..37d8654faff9 100644
> --- a/block/blk-sysfs.c
> +++ b/block/blk-sysfs.c
> @@ -847,7 +847,7 @@ static void blk_queue_release(struct kobject *kobj)
> /* nothing to do here, all data is associated with the parent gendisk */
> }
>
> -static const struct kobj_type blk_queue_ktype = {
> +const struct kobj_type blk_queue_ktype = {
> .default_groups = blk_queue_attr_groups,
> .sysfs_ops = &queue_sysfs_ops,
> .release = blk_queue_release,
> @@ -875,7 +875,6 @@ int blk_register_queue(struct gendisk *disk)
> struct request_queue *q = disk->queue;
> int ret;
>
> - kobject_init(&disk->queue_kobj, &blk_queue_ktype);
> ret = kobject_add(&disk->queue_kobj, &disk_to_dev(disk)->kobj, "queue");
> if (ret < 0)
> goto out_put_queue_kobj;
If the kobject_add() fails here, then we jump to the label out_put_queue_kobj,
where we release/put disk->queue_kobj. That would decrement the kref of
disk->queue_kobj and possibly bring it to zero.
Next time, when we call add_disk() again without invoking kobject_init()
(because the initialization is now moved outside add_disk()), the refcount
of disk->queue_kobj — which was previously released — would now go for a
toss. Wouldn't that lead to use-after-free or inconsistent state?
> @@ -986,5 +985,4 @@ void blk_unregister_queue(struct gendisk *disk)
> elevator_set_none(q);
>
> blk_debugfs_remove(disk);
> - kobject_put(&disk->queue_kobj);
> }
I'm thinking a case where add_disk() fails after the queue is registered.
In that case, we call blk_unregister_queue() — which would ideally put()
the disk->queue_kobj.
But if we skip that put() in blk_unregister_queue() (and that's what we do
above), and then later retry add_disk(), wouldn’t kobject_add() from
blk_register_queue() complain loudly — since we’re trying to add a kobject
that was already added previously?
Thanks,
--Nilay
next prev parent reply other threads:[~2025-08-07 11:52 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-08-07 7:20 [PATCH] block: fix kobject double initialization in add_disk Zheng Qixing
2025-08-07 8:42 ` Yu Kuai
2025-08-07 11:47 ` Nilay Shroff [this message]
2025-08-07 13:44 ` Zheng Qixing
2025-08-08 0:48 ` Yu Kuai
2025-08-08 8:09 ` Nilay Shroff
2025-08-08 8:34 ` Yu Kuai
2025-08-08 9:42 ` Nilay Shroff
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=470ab442-e5eb-4fa8-bde7-d6d2d1115a5a@linux.ibm.com \
--to=nilay@linux.ibm.com \
--cc=axboe@kernel.dk \
--cc=houtao1@huawei.com \
--cc=linux-block@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=yangerkun@huawei.com \
--cc=yi.zhang@huawei.com \
--cc=yukuai3@huawei.com \
--cc=zhengqixing@huawei.com \
--cc=zhengqixing@huaweicloud.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox