* [PATCH v2 0/2] loop: Fix teardown
@ 2026-09-17 21:20 Bart Van Assche
2026-09-17 21:20 ` [PATCH v2 1/2] block: Add post_release() operation Bart Van Assche
0 siblings, 1 reply; 3+ messages in thread
From: Bart Van Assche @ 2026-09-17 21:20 UTC (permalink / raw)
To: Jens Axboe
Cc: linux-block, Christoph Hellwig, Tetsuo Handa, Nilay Shroff,
Bart Van Assche
Hi Jens,
This series fixes loop driver teardown. When tearing down an autoclear loop
device, the loop driver must drain in-flight I/O and flush workqueues to
prevent NULL pointer dereferences in lo_rw_aio() and related I/O paths.
The .release() block device callback is invoked while holding
disk->open_mutex. Freezing the request queue or draining workqueues under
disk->open_mutex causes lock inversion and circular locking dependencies
(e.g., when worker threads or I/O completion paths also acquire open_mutex
or interact with request queue synchronization).
This patch series resolves the lock inversion by:
1. Adding a .post_release() block device operation that is called
synchronously from bdev_release() immediately after disk->open_mutex
is released.
2. Migrating __loop_clr_fd() to .post_release(), allowing loop device
teardown and queue freezing to happen outside of disk->open_mutex.
This series is inspired by Tetsuo's loop driver patch series.
Please consider applying this patch series.
Thanks,
Bart.
Bart Van Assche (1):
loop: Perform __loop_clr_fd() after disk->open_mutex is dropped
Tetsuo Handa (1):
block: Add post_release() operation
block/bdev.c | 3 ++
drivers/block/loop.c | 60 +++++++++++++++++++++++---------
include/linux/blkdev.h | 8 +++++
rust/kernel/block/mq/gen_disk.rs | 1 +
4 files changed, 56 insertions(+), 16 deletions(-)
^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH v2 1/2] block: Add post_release() operation
2026-09-17 21:20 [PATCH v2 0/2] loop: Fix teardown Bart Van Assche
@ 2026-09-17 21:20 ` Bart Van Assche
2026-09-17 21:20 ` [PATCH v2 2/2] loop: Perform __loop_clr_fd() after disk->open_mutex is dropped Bart Van Assche
0 siblings, 1 reply; 3+ messages in thread
From: Bart Van Assche @ 2026-09-17 21:20 UTC (permalink / raw)
To: Jens Axboe
Cc: linux-block, Christoph Hellwig, Tetsuo Handa, Nilay Shroff,
Bart Van Assche, Andreas Hindborg, Miguel Ojeda, Gary Guo,
Tamir Duberstein, Haoze Xie, Ke Sun
From: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
Add post_release() block device operation which provides a hook for
performing synchronous cleanup without disk->open_mutex held, which is
needed by the loop devices.
Real-world container engines, test suites, and system utilities rely on
fput() from __loop_clr_fd() being completed when lo_release() returns.
But changes which went to the v7.1 merge window broke an assumption that
there is no outstanding I/O when __loop_clr_fd() is called, causing NULL
pointer dereference problem in lo_rw_aio().
In order to fix this regression, we want to allow __loop_clr_fd() to flush
outstanding I/O. But calling drain_workqueue() from __loop_clr_fd() with
disk->open_mutex held causes lockdep warnings. We need a mechanism which
can flush outstanding I/O without disk->open_mutex held.
This post_release() operation may be called multiple times since multiple
threads may open the same path concurrently.
Also, this post_release() operation is called from only bdev_release()
path. This is because loop_configure() is not yet called (there is nothing
to clear) if something went wrong between an initialization lo_open() and
an error-unwinding lo_release() within the bdev_open() path.
Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
[ bvanassche: changed bdev->bd_disk into disk and removed those references
to loop driver internals that are no longer correct ]
Signed-off-by: Bart Van Assche <bvanassche@acm.org>
---
block/bdev.c | 3 +++
include/linux/blkdev.h | 8 ++++++++
rust/kernel/block/mq/gen_disk.rs | 1 +
3 files changed, 12 insertions(+)
diff --git a/block/bdev.c b/block/bdev.c
index fac74319e9fb..350f3c29d682 100644
--- a/block/bdev.c
+++ b/block/bdev.c
@@ -1189,6 +1189,9 @@ void bdev_release(struct file *bdev_file)
blkdev_put_whole(bdev);
mutex_unlock(&disk->open_mutex);
+ if (disk->fops->post_release)
+ disk->fops->post_release(disk);
+
module_put(disk->fops->owner);
put_no_open:
blkdev_put_no_open(bdev);
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index d003a9d2d1f6..5f51fe0d8bcc 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -1580,6 +1580,14 @@ struct block_device_operations {
unsigned int flags);
int (*open)(struct gendisk *disk, blk_mode_t mode);
void (*release)(struct gendisk *disk);
+ /*
+ * This operation is called after returned from release() and
+ * disk->open_mutex was released. But this operation is not called
+ * after an initialization open() has succeeded but something went
+ * wrong and an error-unwinding release() was called.
+ * This operation might sleep and has to be idempotent.
+ */
+ void (*post_release)(struct gendisk *disk);
int (*ioctl)(struct block_device *bdev, blk_mode_t mode,
unsigned cmd, unsigned long arg);
int (*compat_ioctl)(struct block_device *bdev, blk_mode_t mode,
diff --git a/rust/kernel/block/mq/gen_disk.rs b/rust/kernel/block/mq/gen_disk.rs
index fc97dd873974..2ff77ef49781 100644
--- a/rust/kernel/block/mq/gen_disk.rs
+++ b/rust/kernel/block/mq/gen_disk.rs
@@ -129,6 +129,7 @@ pub fn build<T: Operations>(
submit_bio: None,
open: None,
release: None,
+ post_release: None,
ioctl: None,
compat_ioctl: None,
check_events: None,
^ permalink raw reply related [flat|nested] 3+ messages in thread
* [PATCH v2 2/2] loop: Perform __loop_clr_fd() after disk->open_mutex is dropped
2026-09-17 21:20 ` [PATCH v2 1/2] block: Add post_release() operation Bart Van Assche
@ 2026-09-17 21:20 ` Bart Van Assche
0 siblings, 0 replies; 3+ messages in thread
From: Bart Van Assche @ 2026-09-17 21:20 UTC (permalink / raw)
To: Jens Axboe
Cc: linux-block, Christoph Hellwig, Tetsuo Handa, Nilay Shroff,
Bart Van Assche
In order to prevent NULL pointer dereferences in lo_rw_aio() when tearing
down a loop device, outstanding I/O must be flushed before clearing the
backing file and device state. However, calling blk_mq_wait_quiesce_done(),
drain_workqueue(), or blk_mq_freeze_queue() with disk->open_mutex held
causes lockdep warnings and potential deadlocks.
Use the .post_release() block device operation to execute __loop_clr_fd()
synchronously after disk->open_mutex has been released by the block layer.
Inside __loop_clr_fd(), outstanding I/O is flushed and the request queue
is frozen before acquiring disk->open_mutex to perform the remaining
device teardown and partition rescans.
Introduce a new loop device state to ensure that __loop_clr_fd() clears
a loop device once even if it is called multiple times concurrently.
Signed-off-by: Bart Van Assche <bvanassche@acm.org>
---
drivers/block/loop.c | 60 ++++++++++++++++++++++++++++++++------------
1 file changed, 44 insertions(+), 16 deletions(-)
diff --git a/drivers/block/loop.c b/drivers/block/loop.c
index 758c20678bf6..0ebee9a8816a 100644
--- a/drivers/block/loop.c
+++ b/drivers/block/loop.c
@@ -42,6 +42,7 @@ enum {
Lo_unbound,
Lo_bound,
Lo_rundown,
+ Lo_clearing,
Lo_deleting,
};
@@ -1138,11 +1139,38 @@ static int loop_configure(struct loop_device *lo, blk_mode_t mode,
static void __loop_clr_fd(struct loop_device *lo)
{
+ struct gendisk *disk = lo->lo_disk;
struct queue_limits lim;
struct file *filp;
gfp_t gfp = lo->old_gfp_mask;
+ unsigned int memflags;
int err;
+ scoped_guard(mutex, &lo->lo_mutex) {
+ if (READ_ONCE(lo->lo_state) != Lo_rundown)
+ return;
+ WRITE_ONCE(lo->lo_state, Lo_clearing);
+ }
+
+ /*
+ * Wait for ongoing loop_queue_rq() calls. Subsequent loop_queue_rq()
+ * calls which are made after this call returned will see lo->lo_state
+ * != Lo_bound and return with BLK_STS_IOERR.
+ */
+ blk_mq_quiesce_queue(lo->lo_queue);
+ blk_mq_unquiesce_queue(lo->lo_queue);
+
+ /* loop_queue_rq() queues work on lo->workqueue, hence drain it. */
+ drain_workqueue(lo->workqueue);
+
+ lim = queue_limits_start_update(lo->lo_queue);
+
+ /*
+ * Freeze the request queue while updating parameters used while
+ * processing requests.
+ */
+ memflags = blk_mq_freeze_queue(lo->lo_queue);
+
spin_lock_irq(&lo->lo_lock);
filp = lo->lo_backing_file;
lo->lo_backing_file = NULL;
@@ -1153,18 +1181,17 @@ static void __loop_clr_fd(struct loop_device *lo)
lo->lo_sizelimit = 0;
memset(lo->lo_file_name, 0, LO_NAME_SIZE);
- /*
- * Reset the block size to the default.
- *
- * No queue freezing needed because this is called from the final
- * ->release call only, so there can't be any outstanding I/O.
- */
- lim = queue_limits_start_update(lo->lo_queue);
+ /* Reset the block size to the default. */
lim.logical_block_size = SECTOR_SIZE;
lim.physical_block_size = SECTOR_SIZE;
lim.io_min = SECTOR_SIZE;
queue_limits_commit_update(lo->lo_queue, &lim);
+ blk_mq_unfreeze_queue(lo->lo_queue, memflags);
+
+ /* Serialize against concurrent bdev_open() calls. */
+ mutex_lock(&disk->open_mutex);
+
invalidate_disk(lo->lo_disk);
loop_sysfs_exit(lo);
/* let user-space know about this change */
@@ -1178,9 +1205,6 @@ static void __loop_clr_fd(struct loop_device *lo)
/*
* Remove all partitions, including partitions added manually with
* BLKPG, which may exist even if LO_FLAGS_PARTSCAN is not set.
- *
- * open_mutex has been held already in release path, so don't acquire
- * it here.
*/
err = bdev_disk_changed(lo->lo_disk, false);
if (err)
@@ -1197,6 +1221,8 @@ static void __loop_clr_fd(struct loop_device *lo)
lo->lo_flags = 0;
if (!part_shift)
set_bit(GD_SUPPRESS_PART_SCAN, &lo->lo_disk->state);
+ mutex_unlock(&disk->open_mutex);
+
mutex_lock(&lo->lo_mutex);
WRITE_ONCE(lo->lo_state, Lo_unbound);
mutex_unlock(&lo->lo_mutex);
@@ -1745,7 +1771,7 @@ static int lo_open(struct gendisk *disk, blk_mode_t mode)
if (err)
return err;
- if (lo->lo_state == Lo_deleting || lo->lo_state == Lo_rundown)
+ if (lo->lo_state != Lo_bound && lo->lo_state != Lo_unbound)
err = -ENXIO;
mutex_unlock(&lo->lo_mutex);
return err;
@@ -1754,7 +1780,6 @@ static int lo_open(struct gendisk *disk, blk_mode_t mode)
static void lo_release(struct gendisk *disk)
{
struct loop_device *lo = disk->private_data;
- bool need_clear = false;
if (disk_openers(disk) > 0)
return;
@@ -1767,12 +1792,14 @@ static void lo_release(struct gendisk *disk)
mutex_lock(&lo->lo_mutex);
if (lo->lo_state == Lo_bound && (lo->lo_flags & LO_FLAGS_AUTOCLEAR))
WRITE_ONCE(lo->lo_state, Lo_rundown);
-
- need_clear = (lo->lo_state == Lo_rundown);
mutex_unlock(&lo->lo_mutex);
+}
+
+static void lo_post_release(struct gendisk *disk)
+{
+ struct loop_device *lo = disk->private_data;
- if (need_clear)
- __loop_clr_fd(lo);
+ __loop_clr_fd(lo);
}
static void lo_free_disk(struct gendisk *disk)
@@ -1791,6 +1818,7 @@ static const struct block_device_operations lo_fops = {
.owner = THIS_MODULE,
.open = lo_open,
.release = lo_release,
+ .post_release = lo_post_release,
.ioctl = lo_ioctl,
#ifdef CONFIG_COMPAT
.compat_ioctl = lo_compat_ioctl,
^ permalink raw reply related [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-17 21:20 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-17 21:20 [PATCH v2 0/2] loop: Fix teardown Bart Van Assche
2026-09-17 21:20 ` [PATCH v2 1/2] block: Add post_release() operation Bart Van Assche
2026-09-17 21:20 ` [PATCH v2 2/2] loop: Perform __loop_clr_fd() after disk->open_mutex is dropped Bart Van Assche
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox