* [PATCH v2 0/2] loop: Fix teardown @ 2026-09-17 21:20 Bart Van Assche 2026-09-17 21:20 ` [PATCH v2 1/2] block: Add post_release() operation Bart Van Assche 0 siblings, 1 reply; 3+ messages in thread From: Bart Van Assche @ 2026-09-17 21:20 UTC (permalink / raw) To: Jens Axboe Cc: linux-block, Christoph Hellwig, Tetsuo Handa, Nilay Shroff, Bart Van Assche Hi Jens, This series fixes loop driver teardown. When tearing down an autoclear loop device, the loop driver must drain in-flight I/O and flush workqueues to prevent NULL pointer dereferences in lo_rw_aio() and related I/O paths. The .release() block device callback is invoked while holding disk->open_mutex. Freezing the request queue or draining workqueues under disk->open_mutex causes lock inversion and circular locking dependencies (e.g., when worker threads or I/O completion paths also acquire open_mutex or interact with request queue synchronization). This patch series resolves the lock inversion by: 1. Adding a .post_release() block device operation that is called synchronously from bdev_release() immediately after disk->open_mutex is released. 2. Migrating __loop_clr_fd() to .post_release(), allowing loop device teardown and queue freezing to happen outside of disk->open_mutex. This series is inspired by Tetsuo's loop driver patch series. Please consider applying this patch series. Thanks, Bart. Bart Van Assche (1): loop: Perform __loop_clr_fd() after disk->open_mutex is dropped Tetsuo Handa (1): block: Add post_release() operation block/bdev.c | 3 ++ drivers/block/loop.c | 60 +++++++++++++++++++++++--------- include/linux/blkdev.h | 8 +++++ rust/kernel/block/mq/gen_disk.rs | 1 + 4 files changed, 56 insertions(+), 16 deletions(-) ^ permalink raw reply [flat|nested] 3+ messages in thread
* [PATCH v2 1/2] block: Add post_release() operation 2026-09-17 21:20 [PATCH v2 0/2] loop: Fix teardown Bart Van Assche @ 2026-09-17 21:20 ` Bart Van Assche 2026-09-17 21:20 ` [PATCH v2 2/2] loop: Perform __loop_clr_fd() after disk->open_mutex is dropped Bart Van Assche 0 siblings, 1 reply; 3+ messages in thread From: Bart Van Assche @ 2026-09-17 21:20 UTC (permalink / raw) To: Jens Axboe Cc: linux-block, Christoph Hellwig, Tetsuo Handa, Nilay Shroff, Bart Van Assche, Andreas Hindborg, Miguel Ojeda, Gary Guo, Tamir Duberstein, Haoze Xie, Ke Sun From: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp> Add post_release() block device operation which provides a hook for performing synchronous cleanup without disk->open_mutex held, which is needed by the loop devices. Real-world container engines, test suites, and system utilities rely on fput() from __loop_clr_fd() being completed when lo_release() returns. But changes which went to the v7.1 merge window broke an assumption that there is no outstanding I/O when __loop_clr_fd() is called, causing NULL pointer dereference problem in lo_rw_aio(). In order to fix this regression, we want to allow __loop_clr_fd() to flush outstanding I/O. But calling drain_workqueue() from __loop_clr_fd() with disk->open_mutex held causes lockdep warnings. We need a mechanism which can flush outstanding I/O without disk->open_mutex held. This post_release() operation may be called multiple times since multiple threads may open the same path concurrently. Also, this post_release() operation is called from only bdev_release() path. This is because loop_configure() is not yet called (there is nothing to clear) if something went wrong between an initialization lo_open() and an error-unwinding lo_release() within the bdev_open() path. Signed-off-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp> [ bvanassche: changed bdev->bd_disk into disk and removed those references to loop driver internals that are no longer correct ] Signed-off-by: Bart Van Assche <bvanassche@acm.org> --- block/bdev.c | 3 +++ include/linux/blkdev.h | 8 ++++++++ rust/kernel/block/mq/gen_disk.rs | 1 + 3 files changed, 12 insertions(+) diff --git a/block/bdev.c b/block/bdev.c index fac74319e9fb..350f3c29d682 100644 --- a/block/bdev.c +++ b/block/bdev.c @@ -1189,6 +1189,9 @@ void bdev_release(struct file *bdev_file) blkdev_put_whole(bdev); mutex_unlock(&disk->open_mutex); + if (disk->fops->post_release) + disk->fops->post_release(disk); + module_put(disk->fops->owner); put_no_open: blkdev_put_no_open(bdev); diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h index d003a9d2d1f6..5f51fe0d8bcc 100644 --- a/include/linux/blkdev.h +++ b/include/linux/blkdev.h @@ -1580,6 +1580,14 @@ struct block_device_operations { unsigned int flags); int (*open)(struct gendisk *disk, blk_mode_t mode); void (*release)(struct gendisk *disk); + /* + * This operation is called after returned from release() and + * disk->open_mutex was released. But this operation is not called + * after an initialization open() has succeeded but something went + * wrong and an error-unwinding release() was called. + * This operation might sleep and has to be idempotent. + */ + void (*post_release)(struct gendisk *disk); int (*ioctl)(struct block_device *bdev, blk_mode_t mode, unsigned cmd, unsigned long arg); int (*compat_ioctl)(struct block_device *bdev, blk_mode_t mode, diff --git a/rust/kernel/block/mq/gen_disk.rs b/rust/kernel/block/mq/gen_disk.rs index fc97dd873974..2ff77ef49781 100644 --- a/rust/kernel/block/mq/gen_disk.rs +++ b/rust/kernel/block/mq/gen_disk.rs @@ -129,6 +129,7 @@ pub fn build<T: Operations>( submit_bio: None, open: None, release: None, + post_release: None, ioctl: None, compat_ioctl: None, check_events: None, ^ permalink raw reply related [flat|nested] 3+ messages in thread
* [PATCH v2 2/2] loop: Perform __loop_clr_fd() after disk->open_mutex is dropped 2026-09-17 21:20 ` [PATCH v2 1/2] block: Add post_release() operation Bart Van Assche @ 2026-09-17 21:20 ` Bart Van Assche 0 siblings, 0 replies; 3+ messages in thread From: Bart Van Assche @ 2026-09-17 21:20 UTC (permalink / raw) To: Jens Axboe Cc: linux-block, Christoph Hellwig, Tetsuo Handa, Nilay Shroff, Bart Van Assche In order to prevent NULL pointer dereferences in lo_rw_aio() when tearing down a loop device, outstanding I/O must be flushed before clearing the backing file and device state. However, calling blk_mq_wait_quiesce_done(), drain_workqueue(), or blk_mq_freeze_queue() with disk->open_mutex held causes lockdep warnings and potential deadlocks. Use the .post_release() block device operation to execute __loop_clr_fd() synchronously after disk->open_mutex has been released by the block layer. Inside __loop_clr_fd(), outstanding I/O is flushed and the request queue is frozen before acquiring disk->open_mutex to perform the remaining device teardown and partition rescans. Introduce a new loop device state to ensure that __loop_clr_fd() clears a loop device once even if it is called multiple times concurrently. Signed-off-by: Bart Van Assche <bvanassche@acm.org> --- drivers/block/loop.c | 60 ++++++++++++++++++++++++++++++++------------ 1 file changed, 44 insertions(+), 16 deletions(-) diff --git a/drivers/block/loop.c b/drivers/block/loop.c index 758c20678bf6..0ebee9a8816a 100644 --- a/drivers/block/loop.c +++ b/drivers/block/loop.c @@ -42,6 +42,7 @@ enum { Lo_unbound, Lo_bound, Lo_rundown, + Lo_clearing, Lo_deleting, }; @@ -1138,11 +1139,38 @@ static int loop_configure(struct loop_device *lo, blk_mode_t mode, static void __loop_clr_fd(struct loop_device *lo) { + struct gendisk *disk = lo->lo_disk; struct queue_limits lim; struct file *filp; gfp_t gfp = lo->old_gfp_mask; + unsigned int memflags; int err; + scoped_guard(mutex, &lo->lo_mutex) { + if (READ_ONCE(lo->lo_state) != Lo_rundown) + return; + WRITE_ONCE(lo->lo_state, Lo_clearing); + } + + /* + * Wait for ongoing loop_queue_rq() calls. Subsequent loop_queue_rq() + * calls which are made after this call returned will see lo->lo_state + * != Lo_bound and return with BLK_STS_IOERR. + */ + blk_mq_quiesce_queue(lo->lo_queue); + blk_mq_unquiesce_queue(lo->lo_queue); + + /* loop_queue_rq() queues work on lo->workqueue, hence drain it. */ + drain_workqueue(lo->workqueue); + + lim = queue_limits_start_update(lo->lo_queue); + + /* + * Freeze the request queue while updating parameters used while + * processing requests. + */ + memflags = blk_mq_freeze_queue(lo->lo_queue); + spin_lock_irq(&lo->lo_lock); filp = lo->lo_backing_file; lo->lo_backing_file = NULL; @@ -1153,18 +1181,17 @@ static void __loop_clr_fd(struct loop_device *lo) lo->lo_sizelimit = 0; memset(lo->lo_file_name, 0, LO_NAME_SIZE); - /* - * Reset the block size to the default. - * - * No queue freezing needed because this is called from the final - * ->release call only, so there can't be any outstanding I/O. - */ - lim = queue_limits_start_update(lo->lo_queue); + /* Reset the block size to the default. */ lim.logical_block_size = SECTOR_SIZE; lim.physical_block_size = SECTOR_SIZE; lim.io_min = SECTOR_SIZE; queue_limits_commit_update(lo->lo_queue, &lim); + blk_mq_unfreeze_queue(lo->lo_queue, memflags); + + /* Serialize against concurrent bdev_open() calls. */ + mutex_lock(&disk->open_mutex); + invalidate_disk(lo->lo_disk); loop_sysfs_exit(lo); /* let user-space know about this change */ @@ -1178,9 +1205,6 @@ static void __loop_clr_fd(struct loop_device *lo) /* * Remove all partitions, including partitions added manually with * BLKPG, which may exist even if LO_FLAGS_PARTSCAN is not set. - * - * open_mutex has been held already in release path, so don't acquire - * it here. */ err = bdev_disk_changed(lo->lo_disk, false); if (err) @@ -1197,6 +1221,8 @@ static void __loop_clr_fd(struct loop_device *lo) lo->lo_flags = 0; if (!part_shift) set_bit(GD_SUPPRESS_PART_SCAN, &lo->lo_disk->state); + mutex_unlock(&disk->open_mutex); + mutex_lock(&lo->lo_mutex); WRITE_ONCE(lo->lo_state, Lo_unbound); mutex_unlock(&lo->lo_mutex); @@ -1745,7 +1771,7 @@ static int lo_open(struct gendisk *disk, blk_mode_t mode) if (err) return err; - if (lo->lo_state == Lo_deleting || lo->lo_state == Lo_rundown) + if (lo->lo_state != Lo_bound && lo->lo_state != Lo_unbound) err = -ENXIO; mutex_unlock(&lo->lo_mutex); return err; @@ -1754,7 +1780,6 @@ static int lo_open(struct gendisk *disk, blk_mode_t mode) static void lo_release(struct gendisk *disk) { struct loop_device *lo = disk->private_data; - bool need_clear = false; if (disk_openers(disk) > 0) return; @@ -1767,12 +1792,14 @@ static void lo_release(struct gendisk *disk) mutex_lock(&lo->lo_mutex); if (lo->lo_state == Lo_bound && (lo->lo_flags & LO_FLAGS_AUTOCLEAR)) WRITE_ONCE(lo->lo_state, Lo_rundown); - - need_clear = (lo->lo_state == Lo_rundown); mutex_unlock(&lo->lo_mutex); +} + +static void lo_post_release(struct gendisk *disk) +{ + struct loop_device *lo = disk->private_data; - if (need_clear) - __loop_clr_fd(lo); + __loop_clr_fd(lo); } static void lo_free_disk(struct gendisk *disk) @@ -1791,6 +1818,7 @@ static const struct block_device_operations lo_fops = { .owner = THIS_MODULE, .open = lo_open, .release = lo_release, + .post_release = lo_post_release, .ioctl = lo_ioctl, #ifdef CONFIG_COMPAT .compat_ioctl = lo_compat_ioctl, ^ permalink raw reply related [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-17 21:20 UTC | newest] Thread overview: 3+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-09-17 21:20 [PATCH v2 0/2] loop: Fix teardown Bart Van Assche 2026-09-17 21:20 ` [PATCH v2 1/2] block: Add post_release() operation Bart Van Assche 2026-09-17 21:20 ` [PATCH v2 2/2] loop: Perform __loop_clr_fd() after disk->open_mutex is dropped Bart Van Assche
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox