From: Tal Zussman <tz2294@columbia.edu>
To: Jens Axboe <axboe@kernel.dk>, Christoph Hellwig <hch@lst.de>,
Johannes Thumshirn <johannes.thumshirn@wdc.com>,
Luis Chamberlain <mcgrof@kernel.org>,
Hannes Reinecke <hare@suse.de>,
"Matthew Wilcox (Oracle)" <willy@infradead.org>,
John Garry <john.g.garry@oracle.com>,
Christian Brauner <brauner@kernel.org>,
"Darrick J. Wong" <djwong@kernel.org>,
Keith Busch <kbusch@kernel.org>,
"Martin K. Petersen" <martin.petersen@oracle.com>
Cc: linux-block@vger.kernel.org, linux-kernel@vger.kernel.org,
Tal Zussman <tz2294@columbia.edu>
Subject: [PATCH v2 6/7] block: unpin all pages of a bvec in bio_iov_iter_align_down()
Date: Fri, 28 Aug 2026 09:49:55 -0400 [thread overview]
Message-ID: <20260828-blkdev-fixes-v2-6-32f3f40cebed@columbia.edu> (raw)
In-Reply-To: <20260828-blkdev-fixes-v2-0-32f3f40cebed@columbia.edu>
bio_iov_iter_align_down() drops trailing bvecs with unpin_user_page(),
but a bvec built by iov_iter_extract_bvecs() can span several pages of
one folio, each with its own pin. All but the first pin leak.
The partially trimmed bvec has the same problem. Shrinking bv_len does
not release the pins for the pages cut off by the trim, and
__bio_release_pages() only unpins the pages bv_len still covers at
completion.
Both issues occur only with a logical block size above PAGE_SIZE and a
large folio backing the user buffer. On a device with a 64K logical
block size, an O_DIRECT pwritev() from a hugetlb mapping that ends 16K
past a block boundary leaks one huge page per call, whether the
remainder is its own bvec or the tail of a larger one.
Unpin all pages of a dropped bvec with unpin_user_folio(), as
__bio_release_pages() does, and unpin the pages trimmed off the last
bvec as well.
Fixes: 20a0e6276edb ("block: align the bio after building it")
Assisted-by: Claude:claude-fable-5
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
block/bio.c | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
diff --git a/block/bio.c b/block/bio.c
index 898b2f5ef8c8..48fa6b9a6dba 100644
--- a/block/bio.c
+++ b/block/bio.c
@@ -1196,6 +1196,11 @@ bool bio_iov_iter_set(struct bio *bio, const struct iov_iter *iter)
return true;
}
+static unsigned int bvec_nr_pages(const struct bio_vec *bv)
+{
+ return DIV_ROUND_UP(bv->bv_offset + bv->bv_len, PAGE_SIZE);
+}
+
/*
* Aligns the bio size to the len_align_mask, releasing excessive bio vecs that
* __bio_iov_iter_get_pages may have inserted, and reverts the trimmed length
@@ -1205,6 +1210,7 @@ static int bio_iov_iter_align_down(struct bio *bio, struct iov_iter *iter,
struct bio_vec *bv, unsigned len_align_mask)
{
size_t nbytes = bio->bi_iter.bi_size & len_align_mask;
+ unsigned int npages;
if (!nbytes)
return 0;
@@ -1213,14 +1219,24 @@ static int bio_iov_iter_align_down(struct bio *bio, struct iov_iter *iter,
bio->bi_iter.bi_size -= nbytes;
while (nbytes >= bv->bv_len) {
if (bio_flagged(bio, BIO_PAGE_PINNED))
- unpin_user_page(bv->bv_page);
+ unpin_user_folio(bvec_folio(bv),
+ bvec_nr_pages(bv));
if (!--bio->bi_vcnt)
return -EFAULT;
nbytes -= bv->bv_len;
bv--;
}
+
+ /*
+ * __bio_release_pages() only unpins the pages still covered by
+ * bv_len, so drop the pins for the pages trimmed off here.
+ */
+ npages = bvec_nr_pages(bv);
bv->bv_len -= nbytes;
+ npages -= bvec_nr_pages(bv);
+ if (npages && bio_flagged(bio, BIO_PAGE_PINNED))
+ unpin_user_folio(bvec_folio(bv), npages);
return 0;
}
--
2.39.5
next prev parent reply other threads:[~2026-08-28 13:50 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-28 13:49 [PATCH v2 0/7] block device fixes for large block sizes, IOCB_NOWAIT, and direct I/O Tal Zussman
2026-08-28 13:49 ` [PATCH v2 1/7] block: use iomap_dirty_folio for block devices Tal Zussman
2026-08-28 13:49 ` [PATCH v2 2/7] block: take i_rwsem for the direct I/O write fallback Tal Zussman
2026-08-28 13:49 ` [PATCH v2 3/7] block: take i_rwsem for the splice read path Tal Zussman
2026-08-28 13:49 ` [PATCH v2 4/7] block: honor IOCB_NOWAIT in the block device buffered " Tal Zussman
2026-08-28 13:49 ` [PATCH v2 5/7] block: fail atomic writes instead of falling back to buffered I/O Tal Zussman
2026-08-28 13:49 ` Tal Zussman [this message]
2026-08-28 14:36 ` [PATCH v2 6/7] block: unpin all pages of a bvec in bio_iov_iter_align_down() Tal Zussman
2026-08-28 13:49 ` [PATCH v2 7/7] block: remove dead metadata handling from the async direct I/O path Tal Zussman
2026-08-28 15:32 ` [PATCH v2 0/7] block device fixes for large block sizes, IOCB_NOWAIT, and direct I/O Tal Zussman
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260828-blkdev-fixes-v2-6-32f3f40cebed@columbia.edu \
--to=tz2294@columbia.edu \
--cc=axboe@kernel.dk \
--cc=brauner@kernel.org \
--cc=djwong@kernel.org \
--cc=hare@suse.de \
--cc=hch@lst.de \
--cc=johannes.thumshirn@wdc.com \
--cc=john.g.garry@oracle.com \
--cc=kbusch@kernel.org \
--cc=linux-block@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=martin.petersen@oracle.com \
--cc=mcgrof@kernel.org \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.