* [PATCH v2 01/15] netfs: clear post-EOF pagecache when extending a file via write
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
@ 2026-09-30 0:38 ` Paulo Alcantara
2026-09-30 0:38 ` [PATCH v2 02/15] smb: client: clear post-EOF pagecache when extending a file via truncate Paulo Alcantara
` (13 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:38 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
Fix netfs to erase the contents of a hole created after the EOF by an
ordinary write if dirty data has been previously left there by writes
through an mmapped region. Neither the buffered nor the unbuffered/DIO
write path clears that stale pagecache.
Zero the tail of the folio straddling the EOF before an extending
write. That is the only folio that can hold data written past the EOF
through an mmap, as pages wholly beyond the EOF can't be faulted in.
The folio is zeroed rather than dropped so a concurrent extending write
can't lose data.
Both write paths downgrade the i_rwsem to shared, so extending writes
can run concurrently and the i_size read by the caller may be stale by
the time the folio is locked. Re-read i_size under the folio lock and
clamp the zeroed range up to it, so a racing write that already put
data into the folio isn't clobbered.
Wait for any writeback on the folio to finish before zeroing it so that
the pagecache isn't modified while it may still be read by the transport
during transmission. Honour IOCB_NOWAIT by returning -EAGAIN rather
than blocking on the folio lock, on writeback, or in folio_mkclean()'s
rmap walk when the folio is mapped.
truncate_pagecache() can't be used here: it must be called with the
i_rwsem held exclusively, but these write paths only hold it shared,
and it would block unconditionally, breaking IOCB_NOWAIT.
Callers that hold i_rwsem exclusively for the whole resize (truncate,
setattr, fallocate, clone) exclude any genuine concurrent buffered
writer, so staleness can instead be decided from the folio's dirty
state, as pagecache_isize_extended() already does for filesystems that
serialise writes against truncate/setattr via a single i_rwsem.
Export netfs_clear_stale_post_isize() helper to handle such case.
The helper is required by the CIFS client to fix generic/363.
Closes: https://sashiko.dev/#/patchset/20260921230755.1133425-1-pc%40manguebit.org
Fixes: 938e13a73b24 ("netfs: Implement buffered write API")
Fixes: 153a9961b551 ("netfs: Implement unbuffered/DIO write support")
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
Documentation/filesystems/netfs_library.rst | 26 ++++++
fs/netfs/buffered_write.c | 9 ++
fs/netfs/direct_write.c | 9 ++
fs/netfs/internal.h | 2 +
fs/netfs/misc.c | 99 +++++++++++++++++++++
include/linux/netfs.h | 2 +
6 files changed, 147 insertions(+)
diff --git a/Documentation/filesystems/netfs_library.rst b/Documentation/filesystems/netfs_library.rst
index ddd799df6ce3..a9de281db8ee 100644
--- a/Documentation/filesystems/netfs_library.rst
+++ b/Documentation/filesystems/netfs_library.rst
@@ -451,6 +451,32 @@ one.
The inode should be marked ``NETFS_ICTX_SINGLE_NO_UPLOAD`` if this API is to be
used. The writeback function requires the buffer to be of ITER_FOLIOQ type.
+Clearing Stale Post-EOF Pagecache
+---------------------------------
+
+When a file is extended, data left in the pagecache past the old EOF by a write
+through an mmap must not be exposed as file content. Netfslib clears this on
+its own write paths, and exports a helper so a filesystem can do the same from
+a resize path (truncate, setattr, fallocate and the like) that holds the
+inode's ``i_rwsem`` exclusively across the whole resize::
+
+ void netfs_clear_stale_post_isize(struct inode *inode, uoff_t from,
+ uoff_t to);
+
+This zeroes any such data within the ``[from, to)`` hole to be made, where
+@from is the old EOF and @to is the new one, and the caller must have already
+updated ``i_size`` to @to before calling it. Only the folio straddling @from
+can hold data written past the EOF through an mmap, as pages wholly beyond the
+EOF can't be faulted in, so the zeroing is limited to that folio. The folio is
+zeroed rather than dropped so that a concurrent extending write can't lose
+data.
+
+This overlaps with ``pagecache_isize_extended()`` but can't reuse it: that
+helper is keyed on a sub-page block size and is a no-op when the block size is
+``>= PAGE_SIZE`` (as on network filesystems), and it doesn't wait for
+writeback. As both address the same problem, a change to one should probably
+be reflected in the other to keep them in sync.
+
High-Level VM API
==================
diff --git a/fs/netfs/buffered_write.c b/fs/netfs/buffered_write.c
index 2cdb68e6b16f..ecf119b4f916 100644
--- a/fs/netfs/buffered_write.c
+++ b/fs/netfs/buffered_write.c
@@ -469,6 +469,7 @@ ssize_t netfs_buffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *fr
struct netfs_group *netfs_group)
{
struct file *file = iocb->ki_filp;
+ struct inode *inode = file_inode(file);
ssize_t ret;
trace_netfs_write_iter(iocb, from);
@@ -481,6 +482,14 @@ ssize_t netfs_buffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *fr
if (ret)
return ret;
+ if (iocb->ki_pos > i_size_read(inode)) {
+ ret = netfs_clear_stale_pre_isize(inode, i_size_read(inode),
+ iocb->ki_pos,
+ iocb->ki_flags & IOCB_NOWAIT);
+ if (ret)
+ return ret;
+ }
+
return netfs_perform_write(iocb, from, netfs_group);
}
EXPORT_SYMBOL(netfs_buffered_write_iter_locked);
diff --git a/fs/netfs/direct_write.c b/fs/netfs/direct_write.c
index 2361277416c7..8be74706d85e 100644
--- a/fs/netfs/direct_write.c
+++ b/fs/netfs/direct_write.c
@@ -359,6 +359,15 @@ ssize_t netfs_unbuffered_write_iter(struct kiocb *iocb, struct iov_iter *from)
ret = file_update_time(file);
if (ret < 0)
goto out;
+
+ if (iocb->ki_pos > i_size_read(inode)) {
+ ret = netfs_clear_stale_pre_isize(inode, i_size_read(inode),
+ iocb->ki_pos,
+ iocb->ki_flags & IOCB_NOWAIT);
+ if (ret < 0)
+ goto out;
+ }
+
if (iocb->ki_flags & IOCB_NOWAIT) {
/* We could block if there are any pages in the range. */
ret = -EAGAIN;
diff --git a/fs/netfs/internal.h b/fs/netfs/internal.h
index c79c8e69d60c..786af76da98c 100644
--- a/fs/netfs/internal.h
+++ b/fs/netfs/internal.h
@@ -80,6 +80,8 @@ ssize_t netfs_wait_for_write(struct netfs_io_request *rreq);
void netfs_wait_for_paused_read(struct netfs_io_request *rreq);
void netfs_wait_for_paused_write(struct netfs_io_request *rreq);
void netfs_wait_for_put_ra_refs(struct netfs_io_request *rreq);
+int netfs_clear_stale_pre_isize(struct inode *inode, uoff_t from,
+ uoff_t to, bool nowait);
/*
* objects.c
diff --git a/fs/netfs/misc.c b/fs/netfs/misc.c
index f5c1c463f4ff..523057390d6f 100644
--- a/fs/netfs/misc.c
+++ b/fs/netfs/misc.c
@@ -6,6 +6,7 @@
*/
#include <linux/swap.h>
+#include <linux/rmap.h>
#include "internal.h"
/**
@@ -582,3 +583,101 @@ void netfs_wait_for_put_ra_refs(struct netfs_io_request *rreq)
trace_netfs_rreq(rreq, netfs_rreq_trace_waited_put_ra_refs);
finish_wait(&rreq->waitq, &myself);
}
+
+/**
+ * netfs_clear_stale_isize - Clear stale pagecache in a to-be-created hole
+ * @inode: The inode to act upon.
+ * @from: The base of the hole to be made.
+ * @to: The top of the hole to be made.
+ * @nowait: True to return -EAGAIN rather than block.
+ * @exclusive: True if the caller holds i_rwsem exclusively for the resize.
+ *
+ * Zero any data left in the pagecache within the [@from, @to) hole by a
+ * write through an mmap so that it isn't exposed as file content once the
+ * file is extended. Only the uptodate folio straddling @from can hold such
+ * data as pages wholly beyond the EOF can't be faulted in, so the zeroing
+ * is limited to that folio. The folio is zeroed rather than dropped so
+ * that a concurrent extending write can't lose data.
+ *
+ * If @exclusive is false, @from is re-read from i_size and used to clamp
+ * the zeroed range, for callers that may race with another writer also
+ * extending the file (eg. multiple buffered writes extending the same file
+ * under a shared i_rwsem). If @exclusive is true, for callers that hold
+ * i_rwsem exclusively across the whole resize and have already updated
+ * i_size to @to, staleness is decided from the folio's dirty state instead:
+ * since no genuine concurrent buffered writer can be racing, a lockless
+ * stat() adopting a server-confirmed size mid-resize has no data behind it
+ * and never dirties the folio, so it can't fool this check into skipping
+ * the zeroing the way it could fool the @exclusive false clamp.
+ *
+ * pagecache_isize_extended() can't be reused here: it is keyed on a
+ * sub-page block size (a no-op when the block size is >= PAGE_SIZE, as on
+ * network filesystems), runs after i_size is updated, can't honour
+ * @nowait, and doesn't wait for writeback. Keep the two in sync if either
+ * is changed.
+ *
+ * Return: 0 on success, or -EAGAIN if @nowait is set and the folio is
+ * mapped or under writeback and so can't be cleaned without blocking.
+ */
+static int netfs_clear_stale_isize(struct inode *inode, uoff_t from,
+ uoff_t to, bool nowait, bool exclusive)
+{
+ struct address_space *mapping = inode->i_mapping;
+ fgf_t fgp = FGP_LOCK;
+ struct folio *folio;
+ int ret;
+
+ if (from >= to)
+ return 0;
+
+ if (nowait)
+ fgp |= FGP_NOWAIT;
+
+ folio = __filemap_get_folio(mapping, from >> PAGE_SHIFT, fgp, 0);
+ if (IS_ERR(folio))
+ return PTR_ERR(folio) == -EAGAIN ? -EAGAIN : 0;
+
+ ret = 0;
+ if (nowait && (folio_mapped(folio) || folio_test_writeback(folio))) {
+ ret = -EAGAIN;
+ goto out;
+ }
+
+ folio_wait_writeback(folio);
+
+ if (folio_mkclean(folio))
+ folio_mark_dirty(folio);
+
+ if (folio_test_uptodate(folio) &&
+ (!exclusive || folio_test_dirty(folio))) {
+ uoff_t fpos = folio_pos(folio);
+
+ if (!exclusive)
+ from = umax(from, i_size_read(inode));
+ if (from < to && from < fpos + folio_size(folio)) {
+ size_t end = umin(to - fpos, folio_size(folio));
+ size_t offset = from - fpos;
+
+ folio_zero_segment(folio, offset, end);
+ }
+ }
+out:
+ folio_unlock(folio);
+ folio_put(folio);
+ return ret;
+}
+
+/* Clear stale pagecache before an extending buffered/DIO write. */
+int netfs_clear_stale_pre_isize(struct inode *inode, uoff_t from,
+ uoff_t to, bool nowait)
+{
+ return netfs_clear_stale_isize(inode, from, to, nowait, false);
+}
+
+/* Clear stale pagecache when extending a file under an exclusive resize. */
+void netfs_clear_stale_post_isize(struct inode *inode, uoff_t from,
+ uoff_t to)
+{
+ netfs_clear_stale_isize(inode, from, to, false, true);
+}
+EXPORT_SYMBOL(netfs_clear_stale_post_isize);
diff --git a/include/linux/netfs.h b/include/linux/netfs.h
index b4dd32863dd4..fe6275e55fc2 100644
--- a/include/linux/netfs.h
+++ b/include/linux/netfs.h
@@ -397,6 +397,8 @@ ssize_t netfs_unbuffered_write_iter(struct kiocb *iocb, struct iov_iter *from);
ssize_t netfs_unbuffered_write_iter_locked(struct kiocb *iocb, struct iov_iter *iter,
struct netfs_group *netfs_group);
ssize_t netfs_file_write_iter(struct kiocb *iocb, struct iov_iter *from);
+void netfs_clear_stale_post_isize(struct inode *inode, uoff_t from,
+ uoff_t to);
/* Single, monolithic object read/write API. */
void netfs_single_mark_inode_dirty(struct inode *inode);
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 02/15] smb: client: clear post-EOF pagecache when extending a file via truncate
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
2026-09-30 0:38 ` [PATCH v2 01/15] netfs: clear post-EOF pagecache when extending a file via write Paulo Alcantara
@ 2026-09-30 0:38 ` Paulo Alcantara
2026-09-30 0:38 ` [PATCH v2 03/15] smb: client: discard post-EOF pagecache when extending a file via zero range Paulo Alcantara
` (12 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:38 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
cifs_setsize() relied on pagecache_isize_extended() to zero the tail of
the folio straddling the old EOF, but that helper is a no-op on CIFS
(i_blkbits is 14), so data dirtied past EOF through an mmap survived
and became visible once the file was extended.
Use netfs_clear_stale_post_isize() instead, called after i_size is
updated since cifs_setsize() callers hold i_rwsem exclusively for the
whole resize. truncate_pagecache() is then called as in
truncate_setsize(): a no-op on extend, and it drops the pagecache
beyond the new EOF on shrink.
Also thread old_size as an explicit parameter through cifs_setsize()
and cifs_resize_file_locked(), captured by each caller before its own
resize RPC. This closes a race where a concurrent stat() could adopt
the RPC's already-updated server size via is_size_safe_to_change()
before cifs_setsize() reads old_size itself.
Closes: https://sashiko.dev/#/patchset/20260921230755.1133425-1-pc%40manguebit.org
Fixes: c510edb9734a ("cifs: call pagecache_isize_extended() in cifs_setsize() when extending")
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/cifsfs.h | 5 +++--
fs/smb/client/file.c | 3 ++-
fs/smb/client/inode.c | 22 ++++++++++++++--------
fs/smb/client/smb2ops.c | 10 ++++++----
4 files changed, 25 insertions(+), 15 deletions(-)
diff --git a/fs/smb/client/cifsfs.h b/fs/smb/client/cifsfs.h
index 0c85daa8386e..e3d820c9778b 100644
--- a/fs/smb/client/cifsfs.h
+++ b/fs/smb/client/cifsfs.h
@@ -146,8 +146,9 @@ ssize_t cifs_file_copychunk_range(unsigned int xid, struct file *src_file,
unsigned int flags);
long cifs_ioctl(struct file *filep, unsigned int command, unsigned long arg);
-void cifs_setsize(struct inode *inode, loff_t offset);
-void cifs_resize_file_locked(struct inode *inode, loff_t offset);
+void cifs_setsize(struct inode *inode, loff_t old_size, loff_t offset);
+void cifs_resize_file_locked(struct inode *inode, loff_t old_size,
+ loff_t offset);
struct fs_context;
struct smb3_fs_context;
diff --git a/fs/smb/client/file.c b/fs/smb/client/file.c
index 0d428517f454..4bcb87610897 100644
--- a/fs/smb/client/file.c
+++ b/fs/smb/client/file.c
@@ -1017,6 +1017,7 @@ static int cifs_do_truncate(const unsigned int xid, struct dentry *dentry)
if (!rc) {
if (cfile) {
struct netfs_inode *ictx = netfs_inode(inode);
+ loff_t old_size = i_size_read(inode);
tcon = tlink_tcon(cfile->tlink);
server = tcon->ses->server;
@@ -1025,7 +1026,7 @@ static int cifs_do_truncate(const unsigned int xid, struct dentry *dentry)
cfile, 0, false);
if (!rc) {
netfs_resize_file(&cinode->netfs, 0, true);
- cifs_setsize(inode, 0);
+ cifs_setsize(inode, old_size, 0);
cifs_invalidate_cache(inode, 0);
}
netfs_wb_end(ictx);
diff --git a/fs/smb/client/inode.c b/fs/smb/client/inode.c
index 1fe0ef0a95db..5062474991cf 100644
--- a/fs/smb/client/inode.c
+++ b/fs/smb/client/inode.c
@@ -3053,15 +3053,12 @@ int cifs_fiemap(struct inode *inode, struct fiemap_extent_info *fei, u64 start,
return -EOPNOTSUPP;
}
-void cifs_setsize(struct inode *inode, loff_t offset)
+void cifs_setsize(struct inode *inode, loff_t old_size, loff_t offset)
{
- loff_t old_size;
u64 blocks = CIFS_INO_BLOCKS(offset);
spin_lock(&inode->i_lock);
- old_size = i_size_read(inode);
i_size_write(inode, offset);
-
/*
* Extending EOF does not allocate the intervening range. Only clamp
* i_blocks on shrink; allocation growth comes from writes or from the
@@ -3071,20 +3068,28 @@ void cifs_setsize(struct inode *inode, loff_t offset)
inode->i_blocks = blocks;
spin_unlock(&inode->i_lock);
inode_set_mtime_to_ts(inode, inode_set_ctime_current(inode));
+
+ /*
+ * Zero the tail of the folio straddling the old EOF so data dirtied
+ * past EOF through an mmap isn't exposed. truncate_pagecache() then
+ * drops any pagecache beyond the new EOF, as in truncate_setsize().
+ */
if (offset > old_size)
- pagecache_isize_extended(inode, old_size, offset);
+ netfs_clear_stale_post_isize(inode, old_size, offset);
+
truncate_pagecache(inode, offset);
netfs_wait_for_outstanding_io(inode);
}
-void cifs_resize_file_locked(struct inode *inode, loff_t offset)
+void cifs_resize_file_locked(struct inode *inode, loff_t old_size,
+ loff_t offset)
{
struct fscache_cookie *cookie = cifs_inode_cookie(inode);
lockdep_assert_held_write(&inode->i_rwsem);
netfs_resize_file(netfs_inode(inode), offset, true);
- cifs_setsize(inode, offset);
+ cifs_setsize(inode, old_size, offset);
if (!cookie)
return;
@@ -3101,6 +3106,7 @@ int cifs_file_set_size(const unsigned int xid, struct dentry *dentry,
struct inode *inode = d_inode(dentry);
struct cifs_sb_info *cifs_sb = CIFS_SB(inode->i_sb);
struct cifsInodeInfo *cifsInode = CIFS_I(inode);
+ loff_t old_size = i_size_read(inode);
struct tcon_link *tlink = NULL;
struct cifs_tcon *tcon = NULL;
struct TCP_Server_Info *server;
@@ -3159,7 +3165,7 @@ int cifs_file_set_size(const unsigned int xid, struct dentry *dentry,
set_size_out:
if (rc == 0)
- cifs_resize_file_locked(inode, size);
+ cifs_resize_file_locked(inode, old_size, size);
return rc;
}
diff --git a/fs/smb/client/smb2ops.c b/fs/smb/client/smb2ops.c
index aa142420dae2..1ada1c0a728d 100644
--- a/fs/smb/client/smb2ops.c
+++ b/fs/smb/client/smb2ops.c
@@ -2287,6 +2287,7 @@ smb2_duplicate_extents(const unsigned int xid,
struct duplicate_extents_to_file dup_ext_buf;
struct timespec64 ts;
struct cifs_tcon *tcon = tlink_tcon(trgtfile->tlink);
+ loff_t old_size;
u64 asize;
/* server fileays advertise duplicate extent support with this flag */
@@ -2305,11 +2306,12 @@ smb2_duplicate_extents(const unsigned int xid,
trgtfile->fid.volatile_fid, tcon->tid,
tcon->ses->Suid, src_off, dest_off, len);
inode = d_inode(trgtfile->dentry);
- if (i_size_read(inode) < dest_off + len) {
+ old_size = i_size_read(inode);
+ if (old_size < dest_off + len) {
rc = smb2_set_file_size(xid, tcon, trgtfile, dest_off + len, false);
if (rc)
goto duplicate_extents_out;
- cifs_resize_file_locked(inode, dest_off + len);
+ cifs_resize_file_locked(inode, old_size, dest_off + len);
}
rc = SMB2_ioctl(xid, tcon, trgtfile->fid.persistent_fid,
trgtfile->fid.volatile_fid,
@@ -3883,7 +3885,7 @@ static long smb3_simple_falloc(struct file *file, struct cifs_tcon *tcon,
}
new_eof = off + len;
- cifs_resize_file_locked(inode, new_eof);
+ cifs_resize_file_locked(inode, old_eof, new_eof);
qrc = SMB2_query_info(xid, tcon,
cfile->fid.persistent_fid,
@@ -3931,7 +3933,7 @@ static long smb3_simple_falloc(struct file *file, struct cifs_tcon *tcon,
if (rc)
goto out;
- cifs_resize_file_locked(inode, new_eof);
+ cifs_resize_file_locked(inode, old_eof, new_eof);
qrc = SMB2_query_info(xid, tcon,
cfile->fid.persistent_fid,
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 03/15] smb: client: discard post-EOF pagecache when extending a file via zero range
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
2026-09-30 0:38 ` [PATCH v2 01/15] netfs: clear post-EOF pagecache when extending a file via write Paulo Alcantara
2026-09-30 0:38 ` [PATCH v2 02/15] smb: client: clear post-EOF pagecache when extending a file via truncate Paulo Alcantara
@ 2026-09-30 0:38 ` Paulo Alcantara
2026-09-30 0:38 ` [PATCH v2 04/15] smb: client: discard post-EOF pagecache when extending a file via copy range Paulo Alcantara
` (11 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:38 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
smb3_zero_range() removed the pagecache only for the range being zeroed,
so when it extends the file the folio straddling the old EOF - which may
hold data dirtied past EOF through an mmap - was never dropped and became
visible as file content once the file was extended.
Lower the discard boundary to the smaller of the zero-range offset and the
old EOF.
Fixes: 72c419d9b073 ("cifs: fix smb3_zero_range so it can expand the file-size when required")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/smb2ops.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git a/fs/smb/client/smb2ops.c b/fs/smb/client/smb2ops.c
index 1ada1c0a728d..849cfd2de705 100644
--- a/fs/smb/client/smb2ops.c
+++ b/fs/smb/client/smb2ops.c
@@ -3560,8 +3560,11 @@ static long smb3_zero_range(struct file *file, struct cifs_tcon *tcon,
/*
* We zero the range through ioctl, so we need remove the page caches
* first, otherwise the data may be inconsistent with the server.
+ *
+ * Start at the old EOF when extending so the folio straddling it, which
+ * may hold data written past EOF through an mmap, is dropped too.
*/
- truncate_pagecache_range(inode, offset, offset + len - 1);
+ truncate_pagecache_range(inode, min(offset, i_size), offset + len - 1);
netfs_wait_for_outstanding_io(inode);
/* if file not oplocked can't be sure whether asking to extend size */
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 04/15] smb: client: discard post-EOF pagecache when extending a file via copy range
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (2 preceding siblings ...)
2026-09-30 0:38 ` [PATCH v2 03/15] smb: client: discard post-EOF pagecache when extending a file via zero range Paulo Alcantara
@ 2026-09-30 0:38 ` Paulo Alcantara
2026-09-30 0:38 ` [PATCH v2 05/15] smb: client: discard post-EOF pagecache when extending a file via clone range Paulo Alcantara
` (10 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:38 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
cifs_file_copychunk_range() invalidated the pagecache only for the
destination region, so when a copy extends the file the folio straddling
the old EOF - which may hold data dirtied past EOF through an mmap - was
never dropped and became visible as file content once the file was
extended.
Start the invalidation at the smaller of the destination offset and the
old EOF. filemap_invalidate_inode() writes back before invalidating
while i_size is still old, so nothing past the old EOF is flushed to the
server.
Fixes: 8101d6e112e2 ("cifs: Fix copy offload to flush destination region")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/cifsfs.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/fs/smb/client/cifsfs.c b/fs/smb/client/cifsfs.c
index 7ecd70efdfea..f1fb4b7b7b7c 100644
--- a/fs/smb/client/cifsfs.c
+++ b/fs/smb/client/cifsfs.c
@@ -1581,8 +1581,13 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
/* Flush and invalidate all the folios in the destination region. If
* the copy was successful, then some of the flush is extra overhead,
* but we need to allow for the copy failing in some way (eg. ENOSPC).
+ *
+ * Start at the old EOF when extending so the folio straddling it, which
+ * may hold data written past EOF through an mmap, is dropped too.
*/
- rc = filemap_invalidate_inode(target_inode, true, destoff, destoff + len - 1);
+ rc = filemap_invalidate_inode(target_inode, true,
+ min(destoff, i_size_read(target_inode)),
+ destoff + len - 1);
if (rc)
goto unlock;
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 05/15] smb: client: discard post-EOF pagecache when extending a file via clone range
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (3 preceding siblings ...)
2026-09-30 0:38 ` [PATCH v2 04/15] smb: client: discard post-EOF pagecache when extending a file via copy range Paulo Alcantara
@ 2026-09-30 0:38 ` Paulo Alcantara
2026-09-30 0:38 ` [PATCH v2 06/15] smb: client: flush and commit data before querying allocated ranges Paulo Alcantara
` (9 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:38 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
cifs_remap_file_range() discarded the pagecache only for the folios
overlapping the destination region, so when a clone extends the file the
folio straddling the old EOF - which may hold data dirtied past EOF
through an mmap - was never dropped and became visible as file content
once the file was extended.
Start the discard at the smaller of fstart and the old EOF. i_size is
still old here, so nothing past the old EOF is flushed to the server.
Fixes: c54fc3a4f375 ("cifs: Fix flushing, invalidation and file size with FICLONE")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/cifsfs.c | 9 +++++++--
1 file changed, 7 insertions(+), 2 deletions(-)
diff --git a/fs/smb/client/cifsfs.c b/fs/smb/client/cifsfs.c
index f1fb4b7b7b7c..6410249f529a 100644
--- a/fs/smb/client/cifsfs.c
+++ b/fs/smb/client/cifsfs.c
@@ -1467,9 +1467,14 @@ static loff_t cifs_remap_file_range(struct file *src_file, loff_t off,
i_size = target_inode->i_size;
spin_unlock(&target_inode->i_lock);
- /* Discard all the folios that overlap the destination region. */
+ /*
+ * Discard all the folios that overlap the destination region. Start at
+ * the old EOF when extending so the folio straddling it, which may hold
+ * data written past EOF through an mmap, is dropped too.
+ */
cifs_dbg(FYI, "about to discard pages %llx-%llx\n", fstart, fend);
- truncate_inode_pages_range(&target_inode->i_data, fstart, fend);
+ truncate_inode_pages_range(&target_inode->i_data,
+ min(fstart, i_size), fend);
fscache_invalidate(cifs_inode_cookie(target_inode), NULL, i_size, 0);
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 06/15] smb: client: flush and commit data before querying allocated ranges
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (4 preceding siblings ...)
2026-09-30 0:38 ` [PATCH v2 05/15] smb: client: discard post-EOF pagecache when extending a file via clone range Paulo Alcantara
@ 2026-09-30 0:38 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 07/15] smb: client: drain outstanding I/O before truncating on O_TRUNC open Paulo Alcantara
` (8 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:38 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
The mode 0 / FALLOC_FL_KEEP_SIZE fallocate emulation for small interior
ranges in smb3_simple_fallocate_range() issues
FSCTL_QUERY_ALLOCATED_RANGES and zero-fills any sub-range the server
reports as an unallocated hole, to force block allocation without
changing file contents.
This is unsafe against Windows servers. Per the Win32 documentation,
FSCTL_QUERY_ALLOCATED_RANGES only reports ranges that *may* contain
nonzero data and is explicitly not coherent with recently written data:
"a call to FSCTL_QUERY_ALLOCATED_RANGES [after writing to a network
file] would not necessarily return a correct list of allocated
regions. To ensure coherency [...] flush the data to the file"
fsx (generic/363) reproduces the resulting corruption against Windows
Server 2022: a range is written, an uncached read returns the data, yet
the next FSCTL_QUERY_ALLOCATED_RANGES on that range reports it as a whole
hole, so the emulation overwrites live data with zeroes. A network trace
confirmed the WRITE, the READ returning data and the query returning an
empty range list within microseconds. Samba does not exhibit this; its
query is coherent with the written data.
Write back the dirty pages, drain any outstanding I/O and issue an SMB2
FLUSH to force the server to commit the data before querying the
allocated ranges, so the query reflects the data actually present.
With this, generic/363 passes against both Windows Server 2022 and Samba
over 100000 fsx operations.
Fixes: 966a3cb7c7db ("cifs: improve fallocate emulation")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/smb2ops.c | 32 ++++++++++++++++++++++++++------
1 file changed, 26 insertions(+), 6 deletions(-)
diff --git a/fs/smb/client/smb2ops.c b/fs/smb/client/smb2ops.c
index 849cfd2de705..1c41d3b24c6d 100644
--- a/fs/smb/client/smb2ops.c
+++ b/fs/smb/client/smb2ops.c
@@ -3750,6 +3750,24 @@ static int smb3_simple_fallocate_range(unsigned int xid,
goto out;
}
+ filemap_invalidate_lock(inode->i_mapping);
+
+ /*
+ * Flush and commit the data to the server, otherwise
+ * FSCTL_QUERY_ALLOCATED_RANGES might report recently written data as
+ * unallocated holes on Windows Servers, and the loop below would
+ * then zero-fill them and corrupt the file.
+ */
+ rc = filemap_write_and_wait_range(inode->i_mapping, off,
+ off + len - 1);
+ if (rc)
+ goto out_unlock;
+ netfs_wait_for_outstanding_io(inode);
+ rc = SMB2_flush(xid, tcon, cfile->fid.persistent_fid,
+ cfile->fid.volatile_fid);
+ if (rc)
+ goto out_unlock;
+
in_data.file_offset = cpu_to_le64(off);
in_data.length = cpu_to_le64(len);
rc = SMB2_ioctl(xid, tcon, cfile->fid.persistent_fid,
@@ -3759,7 +3777,7 @@ static int smb3_simple_fallocate_range(unsigned int xid,
1024 * sizeof(struct file_allocated_range_buffer),
(char **)&out_data, &out_data_len);
if (rc)
- goto out;
+ goto out_unlock;
tmp_data = out_data;
while (len) {
@@ -3769,12 +3787,12 @@ static int smb3_simple_fallocate_range(unsigned int xid,
if (out_data_len == 0) {
rc = smb3_simple_fallocate_write_range(xid, tcon,
cfile, off, len, buf);
- goto out;
+ goto out_unlock;
}
if (out_data_len < sizeof(struct file_allocated_range_buffer)) {
rc = -EINVAL;
- goto out;
+ goto out_unlock;
}
range_start = le64_to_cpu(tmp_data->file_offset);
@@ -3782,7 +3800,7 @@ static int smb3_simple_fallocate_range(unsigned int xid,
if (check_add_overflow(range_start, range_len, &range_end) ||
range_end > S64_MAX) {
rc = -EINVAL;
- goto out;
+ goto out_unlock;
}
if (off < range_start) {
@@ -3797,11 +3815,11 @@ static int smb3_simple_fallocate_range(unsigned int xid,
rc = smb3_simple_fallocate_write_range(xid, tcon,
cfile, off, l, buf);
if (rc)
- goto out;
+ goto out_unlock;
off = off + l;
len = len - l;
if (len == 0)
- goto out;
+ goto out_unlock;
}
/*
* We are at a section of allocated data, just skip forward
@@ -3820,6 +3838,8 @@ static int smb3_simple_fallocate_range(unsigned int xid,
out_data_len -= sizeof(struct file_allocated_range_buffer);
}
+ out_unlock:
+ filemap_invalidate_unlock(inode->i_mapping);
out:
kfree(out_data);
kvfree(buf);
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 07/15] smb: client: drain outstanding I/O before truncating on O_TRUNC open
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (5 preceding siblings ...)
2026-09-30 0:38 ` [PATCH v2 06/15] smb: client: flush and commit data before querying allocated ranges Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 08/15] smb: client: flush dirty data before zeroing a range Paulo Alcantara
` (7 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM,
Frank Sorenson, stable
cifs_do_truncate() flushes the pagecache with filemap_write_and_wait()
before resetting the file size to zero, but that does not wait for
direct/async I/O that was already in flight, which can complete after the
truncate and reinstate stale data past the new EOF.
Drain outstanding netfs I/O with netfs_wait_for_outstanding_io()
before resizing the file.
Reported-by: Frank Sorenson <sorenson@redhat.com>
Fixes: a8603b52b39f ("smb: client: fix data corruption with concurrent writes and O_TRUNC")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/file.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/fs/smb/client/file.c b/fs/smb/client/file.c
index 4bcb87610897..11355d6bf89d 100644
--- a/fs/smb/client/file.c
+++ b/fs/smb/client/file.c
@@ -1012,6 +1012,8 @@ static int cifs_do_truncate(const unsigned int xid, struct dentry *dentry)
}
mapping_set_error(inode->i_mapping, rc);
+ netfs_wait_for_outstanding_io(inode);
+
cfile = find_writable_file(cinode, FIND_FSUID_ONLY);
rc = cifs_file_flush(xid, inode, cfile);
if (!rc) {
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 08/15] smb: client: flush dirty data before zeroing a range
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (6 preceding siblings ...)
2026-09-30 0:39 ` [PATCH v2 07/15] smb: client: drain outstanding I/O before truncating on O_TRUNC open Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 09/15] smb: client: drain and invalidate before server-side copy/clone Paulo Alcantara
` (6 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
smb3_zero_range() emulates FALLOC_FL_ZERO_RANGE by invalidating the
page cache over the target range with truncate_pagecache_range() and
then issuing FSCTL_SET_ZERO_DATA to the server.
Dirty data was only flushed conditionally, when the range reached or
extended EOF. For a purely interior zero range that does not reach
EOF, no flush happened, so a dirty folio overlapping the range could
be written back to the server after the FSCTL and refill the range
that was just zeroed with stale data.
Fix this by unconditionally flushing and waiting for dirty data in the
range before invalidating the page cache and issuing
FSCTL_SET_ZERO_DATA, exactly as was done for smb3_punch_hole() in
commit d7d2adcd022b ("smb/client: flush dirty data before punching a
hole"); the two paths are structurally identical here.
This is observed as generic/363 randomly reading stale data where a
zeroed range is expected against Windows Server.
Fixes: 91d1dfae4649 ("cifs: Fix FALLOC_FL_ZERO_RANGE to preflush buffered part of target region")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/smb2ops.c | 10 ++++------
1 file changed, 4 insertions(+), 6 deletions(-)
diff --git a/fs/smb/client/smb2ops.c b/fs/smb/client/smb2ops.c
index 1c41d3b24c6d..a0fcb564e7a3 100644
--- a/fs/smb/client/smb2ops.c
+++ b/fs/smb/client/smb2ops.c
@@ -3549,13 +3549,11 @@ static long smb3_zero_range(struct file *file, struct cifs_tcon *tcon,
filemap_invalidate_lock(inode->i_mapping);
netfs_read_sizes(inode, &i_size, &remote_i_size, &zero_point);
- if (offset + len >= remote_i_size && offset < i_size) {
- unsigned long long top = umin(offset + len, i_size);
- rc = filemap_write_and_wait_range(inode->i_mapping, offset, top - 1);
- if (rc < 0)
- goto zero_range_exit;
- }
+ rc = filemap_write_and_wait_range(inode->i_mapping, offset,
+ offset + len - 1);
+ if (rc < 0)
+ goto zero_range_exit;
/*
* We zero the range through ioctl, so we need remove the page caches
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 09/15] smb: client: drain and invalidate before server-side copy/clone
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (7 preceding siblings ...)
2026-09-30 0:39 ` [PATCH v2 08/15] smb: client: flush dirty data before zeroing a range Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 10/15] smb: client: only require read lease for size-extending zero range Paulo Alcantara
` (5 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
cifs_file_copychunk_range() and the clone (FICLONE) path of
cifs_remap_file_range() do not serialise the destination page cache
against the server-side copy the way the other server-side range
operations (smb3_zero_range(), smb3_punch_hole(), smb3_collapse_range()
and smb3_insert_range()) do.
cifs_file_copychunk_range() invalidates the destination range with
filemap_invalidate_inode(), which takes and drops the mapping's
invalidate_lock internally, so the lock is no longer held when the
copychunk ioctl is issued. It also never drains in-flight netfs I/O on
the target. The clone path never takes the invalidate_lock at all (only
i_rwsem), uses a bare truncate_inode_pages_range() and likewise does not
drain outstanding I/O.
As a result an asynchronous destination writeback can complete after the
server-side copy/clone has run and reinstate stale data over the region
just written by the server, corrupting the file. This is the same class
of corruption as commit d7d2adcd022b ("smb/client: flush dirty data
before punching a hole") and has been seen randomly in generic/363
against Windows Server.
Fix both paths to follow the established ordering: hold the target
mapping's invalidate_lock across the flush and invalidation of the
destination and the server ioctl, and call netfs_wait_for_outstanding_io()
on the target to drain in-flight writes before the ioctl is issued. Only
the target inode's invalidate_lock is required, as the source is merely
flushed and not invalidated; i_rwsem (already held via
lock_two_nondirectories()) is acquired before the invalidate_lock,
matching the VFS lock ordering. Since filemap_invalidate_inode() takes
that same lock internally, replace it with its own unmap/flush/invalidate
steps instead of calling it, and return early on a zero-length copy to
avoid a range underflow.
Fixes: 8101d6e112e2 ("cifs: Fix copy offload to flush destination region")
Fixes: c54fc3a4f375 ("cifs: Fix flushing, invalidation and file size with FICLONE")
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/cifsfs.c | 31 ++++++++++++++++++++++++++-----
1 file changed, 26 insertions(+), 5 deletions(-)
diff --git a/fs/smb/client/cifsfs.c b/fs/smb/client/cifsfs.c
index 6410249f529a..98b610b2a114 100644
--- a/fs/smb/client/cifsfs.c
+++ b/fs/smb/client/cifsfs.c
@@ -1412,6 +1412,7 @@ static loff_t cifs_remap_file_range(struct file *src_file, loff_t off,
* server could even support copy of range where source = target
*/
lock_two_nondirectories(target_inode, src_inode);
+ filemap_invalidate_lock(target_inode->i_mapping);
if (len == 0) {
loff_t src_size = i_size_read(src_inode);
@@ -1475,6 +1476,7 @@ static loff_t cifs_remap_file_range(struct file *src_file, loff_t off,
cifs_dbg(FYI, "about to discard pages %llx-%llx\n", fstart, fend);
truncate_inode_pages_range(&target_inode->i_data,
min(fstart, i_size), fend);
+ netfs_wait_for_outstanding_io(target_inode);
fscache_invalidate(cifs_inode_cookie(target_inode), NULL, i_size, 0);
@@ -1513,6 +1515,7 @@ static loff_t cifs_remap_file_range(struct file *src_file, loff_t off,
if (rc)
CIFS_I(target_inode)->time = 0;
unlock:
+ filemap_invalidate_unlock(target_inode->i_mapping);
/* although unlocking in the reverse order from locking is not
strictly necessary here it is a little cleaner to be consistent */
unlock_two_nondirectories(src_inode, target_inode);
@@ -1536,6 +1539,9 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
struct cifs_tcon *target_tcon;
ssize_t rc;
+ if (len == 0)
+ return 0;
+
cifs_dbg(FYI, "copychunk range\n");
if (!src_file->private_data || !dst_file->private_data) {
@@ -1565,6 +1571,7 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
* server could even support copy of range where source = target
*/
lock_two_nondirectories(target_inode, src_inode);
+ filemap_invalidate_lock(target_inode->i_mapping);
cifs_dbg(FYI, "about to flush pages\n");
@@ -1590,11 +1597,24 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
* Start at the old EOF when extending so the folio straddling it, which
* may hold data written past EOF through an mmap, is dropped too.
*/
- rc = filemap_invalidate_inode(target_inode, true,
- min(destoff, i_size_read(target_inode)),
- destoff + len - 1);
- if (rc)
- goto unlock;
+ if (target_inode->i_mapping->nrpages) {
+ loff_t fstart = min(destoff, i_size_read(target_inode));
+ loff_t fend = destoff + len - 1;
+
+ unmap_mapping_pages(target_inode->i_mapping,
+ fstart >> PAGE_SHIFT,
+ (fend >> PAGE_SHIFT) -
+ (fstart >> PAGE_SHIFT) + 1,
+ false);
+ rc = filemap_write_and_wait_range(target_inode->i_mapping,
+ fstart, fend);
+ if (rc)
+ goto unlock;
+ invalidate_inode_pages2_range(target_inode->i_mapping,
+ fstart >> PAGE_SHIFT,
+ fend >> PAGE_SHIFT);
+ }
+ netfs_wait_for_outstanding_io(target_inode);
fscache_invalidate(cifs_inode_cookie(target_inode), NULL,
i_size_read(target_inode), 0);
@@ -1626,6 +1646,7 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
CIFS_I(target_inode)->time = 0;
unlock:
+ filemap_invalidate_unlock(target_inode->i_mapping);
/* although unlocking in the reverse order from locking is not
* strictly necessary here it is a little cleaner to be consistent
*/
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 10/15] smb: client: only require read lease for size-extending zero range
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (8 preceding siblings ...)
2026-09-30 0:39 ` [PATCH v2 09/15] smb: client: drain and invalidate before server-side copy/clone Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 11/15] netfs: zero gaps in read-gaps folio to avoid writing back stale data Paulo Alcantara
` (4 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
smb3_zero_range() refuses any FALLOC_FL_ZERO_RANGE that clears
FALLOC_FL_KEEP_SIZE with -EOPNOTSUPP whenever the inode is not read
caching:
/* if file not oplocked can't be sure whether asking to extend size */
rc = -EOPNOTSUPP;
if (keep_size == false && !CIFS_CACHE_READ(cifsi))
goto zero_range_exit;
The read lease is only needed to trust the cached i_size when deciding
whether the range extends the file. When it is not held, the size can
instead be fetched from the server, which is authoritative, rather than
refusing the request outright: query the server's end of file and take
the larger of it and the cached size for the interior-vs-extend
decision. The larger of the two is used because the server's end of
file reflects another client's growth while the cached size reflects
this client's own writes that may not have reached the server yet;
using the server size alone would wrongly shrink the file when the
range flush above left an extending write unwritten.
The query isn't lease-protected either, so its result is only trusted to
confirm the range is interior, never to justify extending: if it still
shows the range going past EOF, refuse with -EOPNOTSUPP instead of
calling SMB2_set_eof(), which could otherwise shrink the file if another
client extended it further between the query and the call. For the same
reason, take the larger of the already-computed i_size and a fresh
i_size_read(inode) right before that call, instead of relying on either
alone: the local variable carries the query's result, which is never
written back to the inode, while a concurrent write on this client can
still extend the cached i_size during the flush, the query, or the
zero-data round trip that happen in between.
Rejecting interior ranges is observed as generic/363 randomly failing
against Windows Server with
do_zero_range: fallocate: Operation not supported
fsx issues an interior, non-KEEP_SIZE zero range while the inode is
transiently not read caching: the server had just downgraded the file's
lease from RWH to RH after breaking the write caching, and the ensuing
handle reopen/revalidation left CIFS_CACHE_READ momentarily clear. The
range sat well within the server's end of file, so no extend was
needed, yet the range was refused and fsx aborted. This keeps the
emulation correct even when a genuine lease break from another client
leaves the inode without read caching -- the case the -EOPNOTSUPP guard
turned into a hard failure. The extra round trip only happens on the
no-lease path; the common cached case is unchanged.
Fixes: 30175628bf7f ("[SMB3] Enable fallocate -z support for SMB3 mounts")
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/smb2ops.c | 31 ++++++++++++++++++++++++++-----
1 file changed, 26 insertions(+), 5 deletions(-)
diff --git a/fs/smb/client/smb2ops.c b/fs/smb/client/smb2ops.c
index a0fcb564e7a3..2d4cb54739ca 100644
--- a/fs/smb/client/smb2ops.c
+++ b/fs/smb/client/smb2ops.c
@@ -3522,6 +3522,21 @@ static long smb3_zero_data(struct file *file, struct cifs_tcon *tcon,
0, NULL, NULL);
}
+static long query_server_eof(const unsigned int xid,
+ struct cifs_tcon *tcon,
+ struct cifsFileInfo *cfile,
+ unsigned long long *eof)
+{
+ struct smb2_file_all_info file_inf = {};
+ long rc;
+
+ rc = SMB2_query_info(xid, tcon, cfile->fid.persistent_fid,
+ cfile->fid.volatile_fid, &file_inf);
+ if (!rc)
+ *eof = le64_to_cpu(file_inf.EndOfFile);
+ return rc;
+}
+
static long smb3_zero_range(struct file *file, struct cifs_tcon *tcon,
unsigned long long offset, unsigned long long len,
bool keep_size)
@@ -3565,10 +3580,16 @@ static long smb3_zero_range(struct file *file, struct cifs_tcon *tcon,
truncate_pagecache_range(inode, min(offset, i_size), offset + len - 1);
netfs_wait_for_outstanding_io(inode);
- /* if file not oplocked can't be sure whether asking to extend size */
- rc = -EOPNOTSUPP;
- if (keep_size == false && !CIFS_CACHE_READ(cifsi))
- goto zero_range_exit;
+ if (!keep_size && !CIFS_CACHE_READ(cifsi)) {
+ rc = query_server_eof(xid, tcon, cfile, &remote_i_size);
+ if (rc)
+ goto zero_range_exit;
+ i_size = max(i_size, remote_i_size);
+ if (i_size < new_size) {
+ rc = -EOPNOTSUPP;
+ goto zero_range_exit;
+ }
+ }
fscache_invalidate(cifs_inode_cookie(inode), NULL,
i_size_read(inode), 0);
@@ -3580,7 +3601,7 @@ static long smb3_zero_range(struct file *file, struct cifs_tcon *tcon,
/*
* do we also need to change the size of the file?
*/
- if (keep_size == false && (unsigned long long)i_size_read(inode) < new_size) {
+ if (!keep_size && umax(i_size, i_size_read(inode)) < new_size) {
rc = SMB2_set_eof(xid, tcon, cfile->fid.persistent_fid,
cfile->fid.volatile_fid, cfile->pid, new_size);
if (rc >= 0) {
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 11/15] netfs: zero gaps in read-gaps folio to avoid writing back stale data
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (9 preceding siblings ...)
2026-09-30 0:39 ` [PATCH v2 10/15] smb: client: only require read lease for size-extending zero range Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 12/15] smb: client: only require read lease for size-extending preallocate Paulo Alcantara
` (3 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
netfs_read_gaps() leaves gaps around a streaming write's dirty region
unread on the server side, so a short/EOF response there left stale
folio content that later got written back.
Zero it after the read completes, not before: netfs_wait_for_read()
returns rreq->transferred once ret >= 0, so build an iterator over the
same bvec array the read used, advance past ret, and zero the rest.
The dirty region already maps to sink pages in that array rather than
the real folio, so this can't touch it, and only the actual shortfall
gets zeroed instead of the whole gap upfront.
Fixes: ee4cdf7ba857 ("netfs: Speed up buffered reading")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/netfs/buffered_read.c | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/fs/netfs/buffered_read.c b/fs/netfs/buffered_read.c
index 105194de6e13..e6506942aeda 100644
--- a/fs/netfs/buffered_read.c
+++ b/fs/netfs/buffered_read.c
@@ -541,6 +541,13 @@ static int netfs_read_gaps(struct file *file, struct folio *folio)
ret = netfs_wait_for_read(rreq);
if (ret >= 0) {
+ if (ret < flen) {
+ struct iov_iter iter;
+
+ iov_iter_bvec(&iter, ITER_DEST, bvec, i, flen);
+ iov_iter_advance(&iter, ret);
+ iov_iter_zero(flen - ret, &iter);
+ }
if (group)
folio_change_private(folio, group);
else
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 12/15] smb: client: only require read lease for size-extending preallocate
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (10 preceding siblings ...)
2026-09-30 0:39 ` [PATCH v2 11/15] netfs: zero gaps in read-gaps folio to avoid writing back stale data Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 13/15] netfs: zero the tail of a short DIO/unbuffered read Paulo Alcantara
` (2 subsequent siblings)
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
smb3_simple_falloc() refuses any fallocate that clears FALLOC_FL_KEEP_SIZE
with -EOPNOTSUPP whenever the inode is not read caching:
/* if file not oplocked can't be sure whether asking to extend size */
if (!CIFS_CACHE_READ(cifsi))
if (!keep_size) {
...
return rc;
}
As with smb3_zero_range(), the read lease is only needed to trust the
cached size when deciding whether the request extends the file. When it
is not held, the size can instead be fetched from the server, which is
authoritative, rather than refusing the request outright: after
flushing, query the server's end of file and take the larger of it and
the cached size for the interior-vs-extend decision. The larger of the
two is used because the server's end of file reflects another client's
growth while the cached size reflects this client's own writes that may
not have reached the server yet; using the server size alone would
wrongly treat an interior request as extending when a range flush left
an extending write unwritten.
smb3_simple_fallocate_range(), which performs the actual interior
emulation, decided whether the range already lies past EOF from its own
i_size_read(inode) rather than the old_eof computed above. In the
leaseless case, that is exactly the stale, undershooting size this
patch works around: a genuinely interior range can read as past EOF by
that stale count, which skips the FSCTL_QUERY_ALLOCATED_RANGES check
entirely and overwrites already-allocated server data with zeroes
instead of only filling the holes. Pass old_eof into
smb3_simple_fallocate_range() and use it for that comparison instead of
re-deriving a second, inconsistent one.
Rejecting interior requests is observed as generic/363 randomly failing
against Windows Server with
do_preallocate: fallocate: Operation not supported
fsx issues an interior, non-KEEP_SIZE preallocate while the inode is
transiently not read caching: the server had just downgraded the file's
lease from RWH to RH after breaking the write caching, and the ensuing
handle reopen/revalidation left CIFS_CACHE_READ momentarily clear. The
range sat within the server's end of file, so no extend was needed, yet
it was refused and fsx aborted. This keeps the emulation correct even
when a genuine lease break from another client leaves the inode
without read caching -- the case the -EOPNOTSUPP guard turned into a
hard failure. The extra flush and round trip only happen when both
keep_size is false and no read lease is held; every other case is
unchanged.
Fixes: 9ccf3216238c ("Add support for original fallocate")
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/smb2ops.c | 33 ++++++++++++++++++++++++---------
1 file changed, 24 insertions(+), 9 deletions(-)
diff --git a/fs/smb/client/smb2ops.c b/fs/smb/client/smb2ops.c
index 2d4cb54739ca..c4e68d1f330f 100644
--- a/fs/smb/client/smb2ops.c
+++ b/fs/smb/client/smb2ops.c
@@ -3747,7 +3747,8 @@ static int smb3_simple_fallocate_write_range(unsigned int xid,
static int smb3_simple_fallocate_range(unsigned int xid,
struct cifs_tcon *tcon,
struct cifsFileInfo *cfile,
- loff_t off, loff_t len)
+ loff_t off, loff_t len,
+ loff_t old_eof)
{
struct file_allocated_range_buffer in_data, *out_data = NULL, *tmp_data;
struct inode *inode = d_inode(cfile->dentry);
@@ -3763,7 +3764,7 @@ static int smb3_simple_fallocate_range(unsigned int xid,
goto out;
}
- if (off >= i_size_read(inode)) {
+ if (off >= old_eof) {
rc = smb3_simple_fallocate_write_range(xid, tcon, cfile,
off, len, buf);
goto out;
@@ -3887,14 +3888,28 @@ static long smb3_simple_falloc(struct file *file, struct cifs_tcon *tcon,
trace_smb3_falloc_enter(xid, cfile->fid.persistent_fid, tcon->tid,
tcon->ses->Suid, off, len);
- /* if file not oplocked can't be sure whether asking to extend size */
- if (!CIFS_CACHE_READ(cifsi))
- if (!keep_size) {
+
+ if (!keep_size && !CIFS_CACHE_READ(cifsi)) {
+ unsigned long long server_eof;
+
+ rc = filemap_write_and_wait(inode->i_mapping);
+ if (rc) {
trace_smb3_falloc_err(xid, cfile->fid.persistent_fid,
tcon->tid, tcon->ses->Suid, off, len, rc);
free_xid(xid);
return rc;
}
+ netfs_wait_for_outstanding_io(inode);
+
+ rc = query_server_eof(xid, tcon, cfile, &server_eof);
+ if (rc) {
+ trace_smb3_falloc_err(xid, cfile->fid.persistent_fid,
+ tcon->tid, tcon->ses->Suid, off, len, rc);
+ free_xid(xid);
+ return rc;
+ }
+ old_eof = max_t(loff_t, old_eof, server_eof);
+ }
/*
* Extending the file
@@ -3918,7 +3933,7 @@ static long smb3_simple_falloc(struct file *file, struct cifs_tcon *tcon,
}
rc = smb3_simple_fallocate_range(xid, tcon, cfile,
- off, len);
+ off, len, old_eof);
if (rc) {
spin_lock(&inode->i_lock);
cifsi->time = 0;
@@ -4020,7 +4035,7 @@ static long smb3_simple_falloc(struct file *file, struct cifs_tcon *tcon,
}
}
- if ((keep_size == true) || (i_size_read(inode) >= off + len)) {
+ if (keep_size || old_eof >= off + len) {
/*
* At this point, we are trying to fallocate an internal
* regions of a sparse file. Since smb2 does not have a
@@ -4037,7 +4052,7 @@ static long smb3_simple_falloc(struct file *file, struct cifs_tcon *tcon,
*/
if (len <= 1024 * 1024) {
rc = smb3_simple_fallocate_range(xid, tcon, cfile,
- off, len);
+ off, len, old_eof);
goto out;
}
@@ -4049,7 +4064,7 @@ static long smb3_simple_falloc(struct file *file, struct cifs_tcon *tcon,
* ie potentially making a few extra pages at the beginning
* or end of the file non-sparse via set_sparse is harmless.
*/
- if ((off > 8192) || (off + len + 8192 < i_size_read(inode))) {
+ if (off > 8192 || off + len + 8192 < old_eof) {
rc = -EOPNOTSUPP;
goto out;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 13/15] netfs: zero the tail of a short DIO/unbuffered read
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (11 preceding siblings ...)
2026-09-30 0:39 ` [PATCH v2 12/15] smb: client: only require read lease for size-extending preallocate Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 14/15] smb: client: distinguish real EOF from a stale remote_i_size on read Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 15/15] smb: client: require stable pages for signed connections Paulo Alcantara
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
The buffered read collector zero-fills the tail of a short read that
stops below the inode's i_size (netfs_clear_unread()), so a read that
races an extending write still returns the expected number of bytes.
The non-buffered collector path does no such thing: it just records
how much was transferred.
Add the same zero-fill for the non-buffered case, gated on
NETFS_SREQ_CLEAR_TAIL: a subreq's source sets that flag when a short
result from it is known to be safe to treat as a hole, as opposed to
NETFS_SREQ_HIT_EOF, which means the read genuinely ran off the end of
the file and should be reported short as-is. Only CLEAR_TAIL should
zero-fill here; a real EOF must stay a real short read.
No source currently sets CLEAR_TAIL on an unbuffered/DIO subrequest,
so this is inert on its own -- a following change teaches cifs to set
it in the one case that needs it.
Fixes: e2d46f2ec332 ("netfs: Change the read result collector to only use one work item")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/netfs/read_collect.c | 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
diff --git a/fs/netfs/read_collect.c b/fs/netfs/read_collect.c
index a94197ef0181..c5bfaf2c6f25 100644
--- a/fs/netfs/read_collect.c
+++ b/fs/netfs/read_collect.c
@@ -33,6 +33,22 @@ static void netfs_clear_unread(struct netfs_io_subrequest *subreq)
__set_bit(NETFS_SREQ_HIT_EOF, &subreq->flags);
}
+static void netfs_clear_unread_dio(struct netfs_io_subrequest *subreq)
+{
+ uoff_t pos = subreq->start + subreq->transferred;
+ struct netfs_io_request *rreq = subreq->rreq;
+ size_t fill;
+
+ if (pos >= rreq->i_size)
+ return;
+
+ fill = min_t(uoff_t, rreq->i_size - pos,
+ subreq->len - subreq->transferred);
+
+ netfs_reset_iter(subreq);
+ subreq->transferred += iov_iter_zero(fill, &subreq->io_iter);
+}
+
/*
* Cancel the copy-to-cache mark on a folio.
*/
@@ -311,6 +327,14 @@ static void netfs_collect_read_results(struct netfs_io_request *rreq)
test_bit(NETFS_SREQ_HIT_EOF, &front->flags))
netfs_read_unlock_folios(rreq, ¬es);
} else {
+ if (!(notes & HIT_PENDING) &&
+ front->error == 0 &&
+ transferred < front->len &&
+ test_bit(NETFS_SREQ_CLEAR_TAIL, &front->flags)) {
+ netfs_clear_unread_dio(front);
+ transferred = front->transferred;
+ trace_netfs_sreq(front, netfs_sreq_trace_clear);
+ }
stream->collected_to = front->start + transferred;
rreq->collected_to = stream->collected_to;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 14/15] smb: client: distinguish real EOF from a stale remote_i_size on read
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (12 preceding siblings ...)
2026-09-30 0:39 ` [PATCH v2 13/15] netfs: zero the tail of a short DIO/unbuffered read Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
2026-09-30 0:39 ` [PATCH v2 15/15] smb: client: require stable pages for signed connections Paulo Alcantara
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
smb2_readv_callback() sets NETFS_SREQ_HIT_EOF whenever a short read
lines up with netfs_read_remote_i_size(inode), the server's EOF as the
client currently believes it. That belief can be stale: after a lease
downgrade and handle reopen, the tracked remote_i_size can sit below
the client's own i_size while an extending write hasn't reached the
server yet. A read in that gap comes back short for a reason that has
nothing to do with the file's real size, but was still marked HIT_EOF,
and netfs reports a short read for it as-is.
Only treat it as real EOF when the position is also at or past the
client's own i_size; otherwise mark it NETFS_SREQ_CLEAR_TAIL instead,
which tells netfs the shortfall is safe to zero-fill rather than
report as a short read.
This is what fsx (generic/363) sees as "short read: 0x0 bytes instead
of 0x<n>" against a Windows server.
Fixes: 1da29f2c39b6 ("netfs, cifs: Fix handling of short DIO read")
Reviewed-by: David Howells <dhowells@redhat.com>
Reviewed-by: Namjae Jeon <linkinjeon@kernel.org>
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/smb2pdu.c | 10 ++++++++--
1 file changed, 8 insertions(+), 2 deletions(-)
diff --git a/fs/smb/client/smb2pdu.c b/fs/smb/client/smb2pdu.c
index 3d7ead36d1a0..cfe3c4b4abe7 100644
--- a/fs/smb/client/smb2pdu.c
+++ b/fs/smb/client/smb2pdu.c
@@ -4796,14 +4796,20 @@ smb2_readv_callback(struct TCP_Server_Info *server, struct mid_q_entry *mid)
rdata->got_bytes);
if (rdata->result == -ENODATA) {
- __set_bit(NETFS_SREQ_HIT_EOF, &rdata->subreq.flags);
rdata->result = 0;
+ if (rdata->subreq.start + rdata->subreq.transferred >= i_size_read(inode))
+ __set_bit(NETFS_SREQ_HIT_EOF, &rdata->subreq.flags);
+ else
+ __set_bit(NETFS_SREQ_CLEAR_TAIL, &rdata->subreq.flags);
} else {
size_t trans = rdata->subreq.transferred + rdata->got_bytes;
if (trans < rdata->subreq.len &&
rdata->subreq.start + trans >= netfs_read_remote_i_size(inode)) {
- __set_bit(NETFS_SREQ_HIT_EOF, &rdata->subreq.flags);
rdata->result = 0;
+ if (rdata->subreq.start + trans >= i_size_read(inode))
+ __set_bit(NETFS_SREQ_HIT_EOF, &rdata->subreq.flags);
+ else
+ __set_bit(NETFS_SREQ_CLEAR_TAIL, &rdata->subreq.flags);
}
if (rdata->got_bytes)
__set_bit(NETFS_SREQ_MADE_PROGRESS, &rdata->subreq.flags);
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread* [PATCH v2 15/15] smb: client: require stable pages for signed connections
2026-09-30 0:38 [PATCH v2 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
` (13 preceding siblings ...)
2026-09-30 0:39 ` [PATCH v2 14/15] smb: client: distinguish real EOF from a stale remote_i_size on read Paulo Alcantara
@ 2026-09-30 0:39 ` Paulo Alcantara
14 siblings, 0 replies; 16+ messages in thread
From: Paulo Alcantara @ 2026-09-30 0:39 UTC (permalink / raw)
To: linux-cifs, netfs
Cc: Christian Brauner, David Howells, Matthew Wilcox, Namjae Jeon,
Ronnie Sahlberg, Shyam Prasad N, Tom Talpey, Bharath SM, stable
When signing, cifs computes the SMB signature over the pagecache folios
in place and then hands those same folios to the socket. If a buffered
write mutates a folio while a write subrequest is still in flight, the
signature no longer matches the data that follows it, the server rejects
the write with STATUS_ACCESS_DENIED (-EACCES), and the error is latched
in the mapping, so the next fsync()/fallocate() returns -EIO.
Mark the mapping for stable writes so netfs_perform_write() waits for
writeback to complete before modifying an in-flight folio. This is
only needed when the connection is signed.
Fixes: 3ee1a1fc3981 ("cifs: Cut over to using netfslib")
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: David Howells <dhowells@redhat.com>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/inode.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/fs/smb/client/inode.c b/fs/smb/client/inode.c
index 5062474991cf..f29555f3c7d0 100644
--- a/fs/smb/client/inode.c
+++ b/fs/smb/client/inode.c
@@ -88,6 +88,8 @@ static void cifs_set_ops(struct inode *inode)
else
inode->i_data.a_ops = &cifs_addr_ops;
mapping_set_large_folios(inode->i_mapping);
+ if (tcon->ses->server->sign)
+ mapping_set_stable_writes(inode->i_mapping);
break;
case S_IFDIR:
if (IS_AUTOMOUNT(inode)) {
--
2.55.0
^ permalink raw reply related [flat|nested] 16+ messages in thread