* [PATCH v4 1/2] block: Add run_todo() operation
@ 2026-09-19 10:50 Tetsuo Handa
2026-09-19 10:51 ` [PATCH v4 2/2] loop: Perform __loop_clr_fd() from run_todo callback Tetsuo Handa
2026-09-21 17:27 ` [PATCH v4 1/2] block: Add run_todo() operation Bart Van Assche
0 siblings, 2 replies; 8+ messages in thread
From: Tetsuo Handa @ 2026-09-19 10:50 UTC (permalink / raw)
To: Bart Van Assche, Jens Axboe, Christoph Hellwig, Jan Kara,
linux-block
Cc: Al Viro, Andrew Morton, Brian Foster, Damien Le Moal,
Hillf Danton, Markus Elfring, Ming Lei, Qu Wenruo, Tao Cui,
kernel test robot, rust-for-linux, Linus Torvalds, Nilay Shroff
Add run_todo() block device operation which provides a hook for performing
synchronous cleanup without disk->open_mutex held, which is needed by the
loop devices.
Real-world container engines, test suites, and system utilities rely on
fput() from __loop_clr_fd() being completed when lo_release() returns.
But changes which went to the v7.1 merge window broke an assumption that
there is no outstanding I/O when __loop_clr_fd() is called, causing NULL
pointer dereference problem in lo_rw_aio().
In order to fix this regression, we want to allow __loop_clr_fd() to flush
outstanding I/O. But calling drain_workqueue() from __loop_clr_fd() with
disk->open_mutex held causes lockdep warnings. We need a mechanism which
can flush outstanding I/O without disk->open_mutex held.
But deferring __loop_clr_fd() to WQ context has a problem that there is no
way to wait for completion of __loop_clr_fd() before the calling thread
returns to the userspace, for there is no hook for calling flush_work().
Despite what LO_FLAGS_AUTOCLEAR can guarantee is to clear backing device
"eventually" after the last thread called lo_release(), abovementioned
programs are expecting "synchronously" when a thread who is going to call
umount() or open() as soon as returning from close() called close().
That is an unsatisfiable expectation because the former is an objective
behavior and the latter is an subjective dependency. Nonetheless, we need
to try to wait for completion of __loop_clr_fd() at best-effort basis.
Also, since __loop_clr_fd() calls module_put(THIS_MODULE) and there is no
API for waiting for completion of remote thread's task work context,
deferring __loop_clr_fd() to task work context has a problem (aside from
task_work_add() being not exported to loadable modules) that module unload
operation can unmap code/data segment before __loop_clr_fd() completes.
Therefore, allow the loop driver to safely know completion of
__loop_clr_fd(), by adding a hook which is called after disk->open_mutex
is released.
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
---
Changes in v4:
Renamed post_release() to run_todo(), and make it be also called when
bdev_open() failed.
block/bdev.c | 10 +++++++++-
include/linux/blkdev.h | 10 ++++++++++
rust/kernel/block/mq/gen_disk.rs | 1 +
3 files changed, 20 insertions(+), 1 deletion(-)
diff --git a/block/bdev.c b/block/bdev.c
index cd8323083740..8f09aafbae11 100644
--- a/block/bdev.c
+++ b/block/bdev.c
@@ -974,6 +974,8 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
bool unblock_events = true;
struct gendisk *disk = bdev->bd_disk;
int ret;
+ struct module *fops_owner = NULL;
+ void (*run_todo)(struct gendisk *disk) = NULL;
if (holder) {
mode |= BLK_OPEN_EXCL;
@@ -996,6 +998,7 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
ret = -EBUSY;
if (!bdev_may_open(bdev, mode))
goto put_module;
+ run_todo = disk->fops->run_todo;
if (bdev_is_partition(bdev))
ret = blkdev_get_part(bdev, mode);
else
@@ -1037,12 +1040,15 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
return 0;
put_module:
- module_put(disk->fops->owner);
+ fops_owner = disk->fops->owner;
abort_claiming:
if (holder)
bd_abort_claiming(bdev, holder);
mutex_unlock(&disk->open_mutex);
disk_unblock_events(disk);
+ if (run_todo)
+ run_todo(disk);
+ module_put(fops_owner);
return ret;
}
@@ -1188,6 +1194,8 @@ void bdev_release(struct file *bdev_file)
else
blkdev_put_whole(bdev);
mutex_unlock(&disk->open_mutex);
+ if (disk->fops->run_todo)
+ disk->fops->run_todo(disk);
module_put(disk->fops->owner);
put_no_open:
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index 4f7905c3412b..d14fc43de172 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1577,6 +1577,16 @@ struct block_device_operations {
unsigned int flags);
int (*open)(struct gendisk *disk, blk_mode_t mode);
void (*release)(struct gendisk *disk);
+ /*
+ * This operation is for performing synchronous cleanup without
+ * disk->open_mutex held when blkdev_get_whole() returned an error or
+ * blkdev_put_whole() was called.
+ * Since this operation is called after disk->open_mutex was released,
+ * users of this operation must implement appropriate serialization.
+ * Also, users of this operation must not expect that either open() or
+ * release() was called before this operation is called.
+ */
+ void (*run_todo)(struct gendisk *disk);
int (*ioctl)(struct block_device *bdev, blk_mode_t mode,
unsigned cmd, unsigned long arg);
int (*compat_ioctl)(struct block_device *bdev, blk_mode_t mode,
diff --git a/rust/kernel/block/mq/gen_disk.rs b/rust/kernel/block/mq/gen_disk.rs
index fc97dd873974..fd837d1e3b27 100644
--- a/rust/kernel/block/mq/gen_disk.rs
+++ b/rust/kernel/block/mq/gen_disk.rs
@@ -129,6 +129,7 @@ pub fn build<T: Operations>(
submit_bio: None,
open: None,
release: None,
+ run_todo: None,
ioctl: None,
compat_ioctl: None,
check_events: None,
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread
* [PATCH v4 2/2] loop: Perform __loop_clr_fd() from run_todo callback.
2026-09-19 10:50 [PATCH v4 1/2] block: Add run_todo() operation Tetsuo Handa
@ 2026-09-19 10:51 ` Tetsuo Handa
2026-09-21 17:30 ` Bart Van Assche
2026-09-21 17:27 ` [PATCH v4 1/2] block: Add run_todo() operation Bart Van Assche
1 sibling, 1 reply; 8+ messages in thread
From: Tetsuo Handa @ 2026-09-19 10:51 UTC (permalink / raw)
To: Bart Van Assche, Jens Axboe, Christoph Hellwig, Jan Kara,
linux-block
Cc: Al Viro, Andrew Morton, Brian Foster, Damien Le Moal,
Hillf Danton, Markus Elfring, Ming Lei, Qu Wenruo, Tao Cui,
kernel test robot, rust-for-linux, Linus Torvalds, Nilay Shroff
syzbot is reporting NULL pointer dereference in lo_rw_aio().
An analysis by the Gemini AI collaborator considers that this problem
is caused by a timing shift primarily exposed by commit 65565ca5f99b
("block: unify the synchronous bi_end_io callbacks"), along with helper
refactorings like commit 92c3737a2473 ("block: add a bio_submit_or_kill
helper").
But due to difficulty of reproducing this race, discussion about what is
happening and how to fix this problem is stalling. Also, we haven't
identified how many filesystems are subjected to this problem.
Therefore, introduce a grace period for flushing outstanding I/O
(which should be a good thing from the perspective of defensive
programming) so that we won't hit NULL pointer dereference problem.
Since calling drain_workqueue() from __loop_clr_fd() with disk->open_mutex
held causes lockdep warnings, call __loop_clr_fd() from lo_run_todo().
Use rundown_owner for remembering who is responsible for calling
__loop_clr_fd() from lo_run_todo().
Link: https://lkml.kernel.org/r/fbb3edda-f108-4e5b-acf2-266f043f8125@I-love.SAKURA.ne.jp
Reported-by: syzbot+cd8a9a308e879a4e2c28@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=cd8a9a308e879a4e2c28
Reported-by: syzbot+bc273027d5643e48e5b3@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=bc273027d5643e48e5b3
Depends-on: "block: Add run_todo() operation"
Fixes: 65565ca5f99b ("block: unify the synchronous bi_end_io callbacks")
Assisted-by: Gemini-Pro
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
---
Changes in v4:
Removed need_clear variable.
drivers/block/loop.c | 66 ++++++++++++++++++++++++++++++++++----------
1 file changed, 52 insertions(+), 14 deletions(-)
diff --git a/drivers/block/loop.c b/drivers/block/loop.c
index 758c20678bf6..86ddcd9dadf0 100644
--- a/drivers/block/loop.c
+++ b/drivers/block/loop.c
@@ -75,6 +75,7 @@ struct loop_device {
struct gendisk *lo_disk;
struct mutex lo_mutex;
bool idr_visible;
+ struct task_struct *rundown_owner; /* current or NULL */
};
struct loop_cmd {
@@ -1138,11 +1139,39 @@ static int loop_configure(struct loop_device *lo, blk_mode_t mode,
static void __loop_clr_fd(struct loop_device *lo)
{
+ struct gendisk *disk = lo->lo_disk;
struct queue_limits lim;
struct file *filp;
gfp_t gfp = lo->old_gfp_mask;
int err;
+ /* Step 1: Flush all outstanding I/O, without open_mutex held. */
+ /*
+ * Since loop_queue_rq() is called with RCU read lock, this synchronize_rcu()
+ * makes sure that no more queue_work() calls are made from loop_queue_work()
+ * from loop_queue_rq(). Subsequent loop_queue_rq() calls which are made after
+ * this synchronize_rcu() returned shall see lo->lo_state != Lo_bound and
+ * return with BLK_STS_IOERR.
+ */
+ synchronize_rcu();
+ /*
+ * This drain_workqueue() makes sure that no more loop_handle_cmd() calls are
+ * made from loop_process_work() from loop_workfn()/loop_rootcg_workfn().
+ */
+ drain_workqueue(lo->workqueue);
+ /*
+ * This blk_mq_freeze_queue() waits for completion of all outstanding I/O
+ * which has been scheduled via loop_queue_rq(), by waiting for q_usage_counter
+ * to reach 0. Since the lo->lo_state != Lo_bound check in loop_queue_rq()
+ * guarantees that no more new I/O requests are made, we can call
+ * blk_mq_unfreeze_queue() immediately after blk_mq_freeze_queue() returns.
+ */
+ blk_mq_unfreeze_queue(lo->lo_queue, blk_mq_freeze_queue(lo->lo_queue));
+
+ /* Step 2: Perform remaining cleanup, with open_mutex held. */
+ mutex_lock(&disk->open_mutex);
+ WARN_ON_ONCE(lo->lo_state != Lo_rundown);
+
spin_lock_irq(&lo->lo_lock);
filp = lo->lo_backing_file;
lo->lo_backing_file = NULL;
@@ -1153,12 +1182,7 @@ static void __loop_clr_fd(struct loop_device *lo)
lo->lo_sizelimit = 0;
memset(lo->lo_file_name, 0, LO_NAME_SIZE);
- /*
- * Reset the block size to the default.
- *
- * No queue freezing needed because this is called from the final
- * ->release call only, so there can't be any outstanding I/O.
- */
+ /* Reset the block size to the default. */
lim = queue_limits_start_update(lo->lo_queue);
lim.logical_block_size = SECTOR_SIZE;
lim.physical_block_size = SECTOR_SIZE;
@@ -1201,11 +1225,9 @@ static void __loop_clr_fd(struct loop_device *lo)
WRITE_ONCE(lo->lo_state, Lo_unbound);
mutex_unlock(&lo->lo_mutex);
- /*
- * Need not hold lo_mutex to fput backing file. Calling fput holding
- * lo_mutex triggers a circular lock dependency possibility warning as
- * fput can take open_mutex which is usually taken before lo_mutex.
- */
+ /* Step 3: Drop refcounts, without open_mutex held. */
+ mutex_unlock(&disk->open_mutex);
+
fput(filp);
}
@@ -1754,7 +1776,6 @@ static int lo_open(struct gendisk *disk, blk_mode_t mode)
static void lo_release(struct gendisk *disk)
{
struct loop_device *lo = disk->private_data;
- bool need_clear = false;
if (disk_openers(disk) > 0)
return;
@@ -1768,11 +1789,27 @@ static void lo_release(struct gendisk *disk)
if (lo->lo_state == Lo_bound && (lo->lo_flags & LO_FLAGS_AUTOCLEAR))
WRITE_ONCE(lo->lo_state, Lo_rundown);
- need_clear = (lo->lo_state == Lo_rundown);
+ /*
+ * In order to flush outstanding I/O (without open_mutex for deadlock
+ * avoidance) before clearing the backing device, defer __loop_clr_fd()
+ * to lo_run_todo().
+ * The Lo_rundown state guarantees that lo_open() will fail with -ENXIO.
+ * The failing lo_open() guarantees that lo_release() will not overwrite
+ * lo->rundown_owner until __loop_clr_fd() resets to the Lo_unbound state.
+ */
+ if (lo->lo_state == Lo_rundown)
+ WRITE_ONCE(lo->rundown_owner, current);
mutex_unlock(&lo->lo_mutex);
+}
- if (need_clear)
+static void lo_run_todo(struct gendisk *disk)
+{
+ struct loop_device *lo = disk->private_data;
+
+ if (READ_ONCE(lo->rundown_owner) == current) {
+ WRITE_ONCE(lo->rundown_owner, NULL);
__loop_clr_fd(lo);
+ }
}
static void lo_free_disk(struct gendisk *disk)
@@ -1791,6 +1828,7 @@ static const struct block_device_operations lo_fops = {
.owner = THIS_MODULE,
.open = lo_open,
.release = lo_release,
+ .run_todo = lo_run_todo,
.ioctl = lo_ioctl,
#ifdef CONFIG_COMPAT
.compat_ioctl = lo_compat_ioctl,
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread
* Re: [PATCH v4 1/2] block: Add run_todo() operation
2026-09-19 10:50 [PATCH v4 1/2] block: Add run_todo() operation Tetsuo Handa
2026-09-19 10:51 ` [PATCH v4 2/2] loop: Perform __loop_clr_fd() from run_todo callback Tetsuo Handa
@ 2026-09-21 17:27 ` Bart Van Assche
2026-09-22 5:08 ` Tetsuo Handa
1 sibling, 1 reply; 8+ messages in thread
From: Bart Van Assche @ 2026-09-21 17:27 UTC (permalink / raw)
To: Tetsuo Handa, Jens Axboe, Christoph Hellwig, Jan Kara,
linux-block
Cc: Al Viro, Andrew Morton, Brian Foster, Damien Le Moal,
Hillf Danton, Markus Elfring, Ming Lei, Qu Wenruo, Tao Cui,
kernel test robot, rust-for-linux, Linus Torvalds, Nilay Shroff
On 9/19/26 3:50 AM, Tetsuo Handa wrote:
> Add run_todo() block device operation which provides a hook for performing
> synchronous cleanup without disk->open_mutex held, which is needed by the
> loop devices.
I like the "post_release()" name much better. I think it is more clear
and more descriptive than "run_todo()".
> diff --git a/block/bdev.c b/block/bdev.c
> index cd8323083740..8f09aafbae11 100644
> --- a/block/bdev.c
> +++ b/block/bdev.c
> @@ -974,6 +974,8 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
> bool unblock_events = true;
> struct gendisk *disk = bdev->bd_disk;
> int ret;
> + struct module *fops_owner = NULL;
> + void (*run_todo)(struct gendisk *disk) = NULL;
>
> if (holder) {
> mode |= BLK_OPEN_EXCL;
> @@ -996,6 +998,7 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
> ret = -EBUSY;
> if (!bdev_may_open(bdev, mode))
> goto put_module;
> + run_todo = disk->fops->run_todo;
> if (bdev_is_partition(bdev))
> ret = blkdev_get_part(bdev, mode);
> else
> @@ -1037,12 +1040,15 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
>
> return 0;
> put_module:
> - module_put(disk->fops->owner);
> + fops_owner = disk->fops->owner;
> abort_claiming:
> if (holder)
> bd_abort_claiming(bdev, holder);
> mutex_unlock(&disk->open_mutex);
> disk_unblock_events(disk);
> + if (run_todo)
> + run_todo(disk);
> + module_put(fops_owner);
> return ret;
> }
The above changes introduce a bug: these changes cause module_put()
to be called if try_module_get() fails. I think that's wrong.
Thanks,
Bart.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v4 2/2] loop: Perform __loop_clr_fd() from run_todo callback.
2026-09-19 10:51 ` [PATCH v4 2/2] loop: Perform __loop_clr_fd() from run_todo callback Tetsuo Handa
@ 2026-09-21 17:30 ` Bart Van Assche
0 siblings, 0 replies; 8+ messages in thread
From: Bart Van Assche @ 2026-09-21 17:30 UTC (permalink / raw)
To: Tetsuo Handa, Jens Axboe, Christoph Hellwig, Jan Kara,
linux-block
Cc: Al Viro, Andrew Morton, Brian Foster, Damien Le Moal,
Hillf Danton, Markus Elfring, Ming Lei, Qu Wenruo, Tao Cui,
kernel test robot, rust-for-linux, Linus Torvalds, Nilay Shroff
On 9/19/26 3:51 AM, Tetsuo Handa wrote:
> diff --git a/drivers/block/loop.c b/drivers/block/loop.c
> index 758c20678bf6..86ddcd9dadf0 100644
> --- a/drivers/block/loop.c
> +++ b/drivers/block/loop.c
> @@ -75,6 +75,7 @@ struct loop_device {
> struct gendisk *lo_disk;
> struct mutex lo_mutex;
> bool idr_visible;
> + struct task_struct *rundown_owner; /* current or NULL */
The comment next to rundown_owner is not useful. Please improve it or
remove it. Once this issue has been addressed, feel free to add:
Reviewed-by: Bart Van Assche <bvanassche@acm.org>
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v4 1/2] block: Add run_todo() operation
2026-09-21 17:27 ` [PATCH v4 1/2] block: Add run_todo() operation Bart Van Assche
@ 2026-09-22 5:08 ` Tetsuo Handa
2026-09-22 18:09 ` Bart Van Assche
0 siblings, 1 reply; 8+ messages in thread
From: Tetsuo Handa @ 2026-09-22 5:08 UTC (permalink / raw)
To: Bart Van Assche, Jens Axboe, Christoph Hellwig, Jan Kara,
linux-block
Cc: Al Viro, Andrew Morton, Brian Foster, Damien Le Moal,
Hillf Danton, Markus Elfring, Ming Lei, Qu Wenruo, Tao Cui,
kernel test robot, rust-for-linux, Linus Torvalds, Nilay Shroff
On 2026/09/22 2:27, Bart Van Assche wrote:
> On 9/19/26 3:50 AM, Tetsuo Handa wrote:
>> Add run_todo() block device operation which provides a hook for performing
>> synchronous cleanup without disk->open_mutex held, which is needed by the
>> loop devices.
>
> I like the "post_release()" name much better. I think it is more clear
> and more descriptive than "run_todo()".
I think post_release() is misleading because post_release() is also called
when disk->fops->open() in blkdev_get_whole() failed (and therefore
bdev->bd_disk->fops->release() in blkdev_put_whole() from blkdev_get_whole()
was not called).
>
>> diff --git a/block/bdev.c b/block/bdev.c
>> index cd8323083740..8f09aafbae11 100644
>> --- a/block/bdev.c
>> +++ b/block/bdev.c
>> @@ -974,6 +974,8 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
>> bool unblock_events = true;
>> struct gendisk *disk = bdev->bd_disk;
>> int ret;
>> + struct module *fops_owner = NULL;
>> + void (*run_todo)(struct gendisk *disk) = NULL;
>> if (holder) {
>> mode |= BLK_OPEN_EXCL;
>> @@ -996,6 +998,7 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
>> ret = -EBUSY;
>> if (!bdev_may_open(bdev, mode))
>> goto put_module;
>> + run_todo = disk->fops->run_todo;
>> if (bdev_is_partition(bdev))
>> ret = blkdev_get_part(bdev, mode);
>> else
>> @@ -1037,12 +1040,15 @@ int bdev_open(struct block_device *bdev, blk_mode_t mode, void *holder,
>> return 0;
>> put_module:
>> - module_put(disk->fops->owner);
>> + fops_owner = disk->fops->owner;
>> abort_claiming:
>> if (holder)
>> bd_abort_claiming(bdev, holder);
>> mutex_unlock(&disk->open_mutex);
>> disk_unblock_events(disk);
>> + if (run_todo)
>> + run_todo(disk);
>> + module_put(fops_owner);
>> return ret;
>> }
> The above changes introduce a bug: these changes cause module_put()
> to be called if try_module_get() fails. I think that's wrong.
No, this change is safe. module_put(NULL) is no-op, fops_owner is initialized with NULL,
and fops_owner might become non-NULL at the put_module label.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v4 1/2] block: Add run_todo() operation
2026-09-22 5:08 ` Tetsuo Handa
@ 2026-09-22 18:09 ` Bart Van Assche
2026-09-23 5:26 ` Tetsuo Handa
0 siblings, 1 reply; 8+ messages in thread
From: Bart Van Assche @ 2026-09-22 18:09 UTC (permalink / raw)
To: Tetsuo Handa, Jens Axboe, Christoph Hellwig, Jan Kara,
linux-block
Cc: Al Viro, Andrew Morton, Brian Foster, Damien Le Moal,
Hillf Danton, Markus Elfring, Ming Lei, Qu Wenruo, Tao Cui,
kernel test robot, rust-for-linux, Linus Torvalds, Nilay Shroff
On 9/21/26 10:08 PM, Tetsuo Handa wrote:
> I think post_release() is misleading because post_release() is also called
> when disk->fops->open() in blkdev_get_whole() failed (and therefore
> bdev->bd_disk->fops->release() in blkdev_put_whole() from blkdev_get_whole()
> was not called).
Has it been considered to call the new callback only if fops->release()
has been invoked? Here is an (untested) example of how this could be
implemented:
diff --git a/block/bdev.c b/block/bdev.c
index cd8323083740..82be4812a7a0 100644
--- a/block/bdev.c
+++ b/block/bdev.c
@@ -766,7 +766,8 @@ static void blkdev_put_whole(struct block_device *bdev)
bdev->bd_disk->fops->release(bdev->bd_disk);
}
-static int blkdev_get_whole(struct block_device *bdev, blk_mode_t mode)
+static int blkdev_get_whole(struct block_device *bdev, blk_mode_t mode,
+ bool *called_release)
{
struct gendisk *disk = bdev->bd_disk;
int ret;
@@ -793,18 +794,20 @@ static int blkdev_get_whole(struct block_device
*bdev, blk_mode_t mode)
ret = bdev_disk_changed(disk, false);
if (ret && (mode & BLK_OPEN_STRICT_SCAN)) {
blkdev_put_whole(bdev);
+ *called_release = true;
return ret;
}
}
return 0;
}
-static int blkdev_get_part(struct block_device *part, blk_mode_t mode)
+static int blkdev_get_part(struct block_device *part, blk_mode_t mode,
+ bool *called_release)
{
struct gendisk *disk = part->bd_disk;
int ret;
- ret = blkdev_get_whole(bdev_whole(part), mode);
+ ret = blkdev_get_whole(bdev_whole(part), mode, called_release);
if (ret)
return ret;
@@ -821,6 +824,7 @@ static int blkdev_get_part(struct block_device
*part, blk_mode_t mode)
out_blkdev_put:
blkdev_put_whole(bdev_whole(part));
+ *called_release = true;
return ret;
}
@@ -973,6 +977,7 @@ int bdev_open(struct block_device *bdev, blk_mode_t
mode, void *holder,
{
bool unblock_events = true;
struct gendisk *disk = bdev->bd_disk;
+ bool called_release = false;
int ret;
if (holder) {
@@ -997,9 +1002,9 @@ int bdev_open(struct block_device *bdev, blk_mode_t
mode, void *holder,
if (!bdev_may_open(bdev, mode))
goto put_module;
if (bdev_is_partition(bdev))
- ret = blkdev_get_part(bdev, mode);
+ ret = blkdev_get_part(bdev, mode, &called_release);
else
- ret = blkdev_get_whole(bdev, mode);
+ ret = blkdev_get_whole(bdev, mode, &called_release);
if (ret)
goto put_module;
bdev_claim_write_access(bdev, mode);
@@ -1037,7 +1042,14 @@ int bdev_open(struct block_device *bdev,
blk_mode_t mode, void *holder,
return 0;
put_module:
+ if (holder)
+ bd_abort_claiming(bdev, holder);
+ mutex_unlock(&disk->open_mutex);
+ if (called_release && disk->fops->post_release)
+ disk->fops->post_release(disk);
+ disk_unblock_events(disk);
module_put(disk->fops->owner);
+ return ret;
abort_claiming:
if (holder)
bd_abort_claiming(bdev, holder);
@@ -1189,6 +1201,9 @@ void bdev_release(struct file *bdev_file)
blkdev_put_whole(bdev);
mutex_unlock(&disk->open_mutex);
+ if (disk->fops->post_release)
+ disk->fops->post_release(disk);
+
module_put(disk->fops->owner);
put_no_open:
blkdev_put_no_open(bdev);
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index 4f7905c3412b..f2a2f70a9582 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1577,6 +1577,7 @@ struct block_device_operations {
unsigned int flags);
int (*open)(struct gendisk *disk, blk_mode_t mode);
void (*release)(struct gendisk *disk);
+ void (*post_release)(struct gendisk *disk);
int (*ioctl)(struct block_device *bdev, blk_mode_t mode,
unsigned cmd, unsigned long arg);
int (*compat_ioctl)(struct block_device *bdev, blk_mode_t mode,
Thanks,
Bart.
^ permalink raw reply related [flat|nested] 8+ messages in thread
* Re: [PATCH v4 1/2] block: Add run_todo() operation
2026-09-22 18:09 ` Bart Van Assche
@ 2026-09-23 5:26 ` Tetsuo Handa
2026-09-23 17:05 ` Bart Van Assche
0 siblings, 1 reply; 8+ messages in thread
From: Tetsuo Handa @ 2026-09-23 5:26 UTC (permalink / raw)
To: Bart Van Assche, Jens Axboe, Christoph Hellwig, Jan Kara,
linux-block
Cc: Al Viro, Andrew Morton, Brian Foster, Damien Le Moal,
Hillf Danton, Markus Elfring, Ming Lei, Qu Wenruo, Tao Cui,
kernel test robot, rust-for-linux, Linus Torvalds, Nilay Shroff
On 2026/09/23 3:09, Bart Van Assche wrote:
> On 9/21/26 10:08 PM, Tetsuo Handa wrote:
>> I think post_release() is misleading because post_release() is also called
>> when disk->fops->open() in blkdev_get_whole() failed (and therefore
>> bdev->bd_disk->fops->release() in blkdev_put_whole() from blkdev_get_whole()
>> was not called).
>
> Has it been considered to call the new callback only if fops->release()
> has been invoked? Here is an (untested) example of how this could be
> implemented:
Yes, I have considered that approach. I didn't implement that change because
it is a needless bloat. But if you are OK with that approach, we can go with
your approach.
The v5 patchset passed sashiko's review. Can we apply the v5 ?
https://sashiko.dev/#/log/178321
https://sashiko.dev/#/log/178322
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH v4 1/2] block: Add run_todo() operation
2026-09-23 5:26 ` Tetsuo Handa
@ 2026-09-23 17:05 ` Bart Van Assche
0 siblings, 0 replies; 8+ messages in thread
From: Bart Van Assche @ 2026-09-23 17:05 UTC (permalink / raw)
To: Tetsuo Handa, Jens Axboe, Christoph Hellwig, Jan Kara,
linux-block
Cc: Al Viro, Andrew Morton, Brian Foster, Damien Le Moal,
Hillf Danton, Markus Elfring, Ming Lei, Qu Wenruo, Tao Cui,
kernel test robot, rust-for-linux, Linus Torvalds, Nilay Shroff
On 9/22/26 10:26 PM, Tetsuo Handa wrote:
> The v5 patchset passed sashiko's review. Can we apply the v5 ?
That's up to Jens to decide. I think reviews from other block layer
experts would make it easier for him to decide whether or not to queue
this patch series.
Thanks,
Bart.
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-09-23 17:05 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-19 10:50 [PATCH v4 1/2] block: Add run_todo() operation Tetsuo Handa
2026-09-19 10:51 ` [PATCH v4 2/2] loop: Perform __loop_clr_fd() from run_todo callback Tetsuo Handa
2026-09-21 17:30 ` Bart Van Assche
2026-09-21 17:27 ` [PATCH v4 1/2] block: Add run_todo() operation Bart Van Assche
2026-09-22 5:08 ` Tetsuo Handa
2026-09-22 18:09 ` Bart Van Assche
2026-09-23 5:26 ` Tetsuo Handa
2026-09-23 17:05 ` Bart Van Assche
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox