FILESYSTEM IN USERSPACE (FUSE) development
 help / color / mirror / Atom feed
From: Joanne Koong <joannelkoong@gmail.com>
To: miklos@szeredi.hu
Cc: fuse-devel@lists.linux.dev, Sashiko <sashiko-bot@kernel.org>,
	stable@vger.kernel.org
Subject: [PATCH v1 1/2] fuse: don't shorten the folio descriptor at LLONG_MAX
Date: Thu,  3 Sep 2026 21:56:50 -0700	[thread overview]
Message-ID: <20260904045651.1442505-2-joannelkoong@gmail.com> (raw)
In-Reply-To: <20260904045651.1442505-1-joannelkoong@gmail.com>

fuse_send_readpages() and fuse_do_readfolio() both decrement the folio
descriptor length when handling the overflow case where a read would
exceed LLONG_MAX after incrementing the file position by the number of
bytes that need to be read in.

Shortening it is unnecessary (and for the virtio-fs paths, a bug), and
with fuse using iomap for handling reads, shortening it is buggy.

For fuse_send_readpages(), ap->descs[] is also what fuse_readpages_end()
reports back to iomap_finish_folio_read(). iomap accounted the full
length when the range was submitted, so the lengths reported on
completion have to add up to what was submitted. Reporting one byte less
leaves ifs->read_bytes_pending nonzero, folio_end_read() is never called,
and the folio stays locked.

This is currently reachable on fuseblk mounts configured with block
sizes smaller than the page size.

Fix this by leaving the descriptor length untouched. The reply will then
be one byte shorter than what the descriptor lengths add up to, and the
last byte will be zeroed in fuse_copy_folios().

This also fixes a bug in virtio-fs that existed before any iomap changes
were added to fuse. For virtio-fs, data there arrives by DMA rather than
through fuse_copy_folios(), so the only zeroing is in
virtio_fs_request_complete(), which compares the reply size against the
descriptor length. With the descriptor length shortened, the two are
equal and nothing gets zeroed, which means the unrequested byte is left
holding whatever was in the folio, when the folio is marked uptodate.
This is both in the fuse_do_readfolio() and fuse_send_readpages() paths.
This dates back to commit 2f1398291bf3 ("fuse: don't overflow LLONG_MAX
with end offset").

Fixes: 4ea907108a5c ("fuse: use iomap for readahead")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Cc: stable@vger.kernel.org
Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
---
 fs/fuse/file.c | 38 +++++++++++++++++++++++++++++---------
 1 file changed, 29 insertions(+), 9 deletions(-)

diff --git a/fs/fuse/file.c b/fs/fuse/file.c
index 4bdf5dec2cb9..3bf7cb590538 100644
--- a/fs/fuse/file.c
+++ b/fs/fuse/file.c
@@ -864,18 +864,29 @@ static int fuse_do_readfolio(struct file *file, struct folio *folio,
 
 	attr_ver = fuse_get_attr_version(fm->fc);
 
-	/* Don't overflow end offset */
-	if (pos + (desc.length - 1) == LLONG_MAX)
-		desc.length--;
+	/*
+	 * Don't overflow end offset.
+	 *
+	 * Ask the server for len - 1 bytes. desc.length still holds the full
+	 * length. When the reply comes back, it will be one byte shorter than
+	 * desc.length and fuse_copy_folios() will zero that last byte.
+	 *
+	 * For this reason, desc.length must not be decremented too. The caller
+	 * reports the full length to iomap_finish_folio_read(), which marks
+	 * every block it covers uptodate. Shortening the descriptor would
+	 * suppress zeroing and leave the last byte holding stale data.
+	 */
+	if (pos + (len - 1) == LLONG_MAX)
+		len--;
 
-	fuse_read_args_fill(&ia, file, pos, desc.length, FUSE_READ);
+	fuse_read_args_fill(&ia, file, pos, len, FUSE_READ);
 	res = fuse_simple_request(fm, &ia.ap.args);
 	if (res < 0)
 		return res;
 	/*
 	 * Short read means EOF.  If file size is larger, truncate it
 	 */
-	if (res < desc.length)
+	if (res < len)
 		fuse_short_read(inode, attr_ver, res, &ia.ap);
 
 	return 0;
@@ -1068,11 +1079,20 @@ static void fuse_send_readpages(struct fuse_io_args *ia, struct file *file,
 	ap->args.page_zeroing = true;
 	ap->args.page_replace = true;
 
-	/* Don't overflow end offset */
-	if (pos + (count - 1) == LLONG_MAX) {
+	/*
+	 * Don't overflow end offset.
+	 *
+	 * Ask the server for count - 1 bytes. The reply is then one byte
+	 * shorter than what the descriptor lengths add up to, so
+	 * fuse_copy_folios() zeroes the last byte when it walks the folios.
+	 *
+	 * ap->descs[] must not be decremented here. It is what
+	 * fuse_readpages_end() reports back to iomap_finish_folio_read(), and
+	 * iomap has already accounted the full descriptor length, so shortening
+	 * it would leave ifs->read_bytes_pending nonzero and the folio locked.
+	 */
+	if (pos + (count - 1) == LLONG_MAX)
 		count--;
-		ap->descs[ap->num_folios - 1].length--;
-	}
 	WARN_ON((loff_t) (pos + count) < 0);
 
 	fuse_read_args_fill(ia, file, pos, count, FUSE_READ);
-- 
2.52.0


  reply	other threads:[~2026-09-04  4:57 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04  4:56 [PATCH v1 0/2] fuse: fixes for iomap bugs reported by Sashiko Joanne Koong
2026-09-04  4:56 ` Joanne Koong [this message]
2026-09-04  4:56 ` [PATCH v1 2/2] fuse: zero the correct range on a short read reply Joanne Koong

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260904045651.1442505-2-joannelkoong@gmail.com \
    --to=joannelkoong@gmail.com \
    --cc=fuse-devel@lists.linux.dev \
    --cc=miklos@szeredi.hu \
    --cc=sashiko-bot@kernel.org \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox