From: Paulo Alcantara <pc@manguebit.org>
To: linux-cifs@vger.kernel.org, netfs@lists.linux.dev
Cc: Christian Brauner <brauner@kernel.org>,
David Howells <dhowells@redhat.com>,
Matthew Wilcox <willy@infradead.org>,
Namjae Jeon <linkinjeon@kernel.org>,
Ronnie Sahlberg <ronniesahlberg@gmail.com>,
Shyam Prasad N <sprasad@microsoft.com>,
Tom Talpey <tom@talpey.com>, Bharath SM <bharathsm@microsoft.com>,
stable@vger.kernel.org
Subject: [PATCH v3 09/15] smb: client: drain and invalidate before server-side copy/clone
Date: Wed, 30 Sep 2026 00:28:16 -0300 [thread overview]
Message-ID: <20260930032822.1835287-10-pc@manguebit.org> (raw)
In-Reply-To: <20260930032822.1835287-1-pc@manguebit.org>
cifs_file_copychunk_range() and the clone (FICLONE) path of
cifs_remap_file_range() do not serialise the destination page cache
against the server-side copy the way the other server-side range
operations (smb3_zero_range(), smb3_punch_hole(), smb3_collapse_range()
and smb3_insert_range()) do.
cifs_file_copychunk_range() invalidates the destination range with
filemap_invalidate_inode(), which takes and drops the mapping's
invalidate_lock internally, so the lock is no longer held when the
copychunk ioctl is issued. It also never drains in-flight netfs I/O on
the target. The clone path never takes the invalidate_lock at all (only
i_rwsem), uses a bare truncate_inode_pages_range() and likewise does not
drain outstanding I/O.
As a result an asynchronous destination writeback can complete after the
server-side copy/clone has run and reinstate stale data over the region
just written by the server, corrupting the file. This is the same class
of corruption as commit d7d2adcd022b ("smb/client: flush dirty data
before punching a hole") and has been seen randomly in generic/363
against Windows Server.
Fix both paths to follow the established ordering: hold the target
mapping's invalidate_lock across the flush and invalidation of the
destination and the server ioctl, and call netfs_wait_for_outstanding_io()
on the target to drain in-flight writes before the ioctl is issued. Only
the target inode's invalidate_lock is required, as the source is merely
flushed and not invalidated; i_rwsem (already held via
lock_two_nondirectories()) is acquired before the invalidate_lock,
matching the VFS lock ordering. Since filemap_invalidate_inode() takes
that same lock internally, replace it with its own unmap/flush/invalidate
steps instead of calling it, and return early on a zero-length copy to
avoid a range underflow.
Fixes: 8101d6e112e2 ("cifs: Fix copy offload to flush destination region")
Fixes: c54fc3a4f375 ("cifs: Fix flushing, invalidation and file size with FICLONE")
Signed-off-by: Paulo Alcantara <pc@manguebit.org>
Cc: Christian Brauner <brauner@kernel.org>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com>
Cc: Shyam Prasad N <sprasad@microsoft.com>
Cc: Tom Talpey <tom@talpey.com>
Cc: Bharath SM <bharathsm@microsoft.com>
Cc: stable@vger.kernel.org
---
fs/smb/client/cifsfs.c | 31 ++++++++++++++++++++++++++-----
1 file changed, 26 insertions(+), 5 deletions(-)
diff --git a/fs/smb/client/cifsfs.c b/fs/smb/client/cifsfs.c
index 6410249f529a..98b610b2a114 100644
--- a/fs/smb/client/cifsfs.c
+++ b/fs/smb/client/cifsfs.c
@@ -1412,6 +1412,7 @@ static loff_t cifs_remap_file_range(struct file *src_file, loff_t off,
* server could even support copy of range where source = target
*/
lock_two_nondirectories(target_inode, src_inode);
+ filemap_invalidate_lock(target_inode->i_mapping);
if (len == 0) {
loff_t src_size = i_size_read(src_inode);
@@ -1475,6 +1476,7 @@ static loff_t cifs_remap_file_range(struct file *src_file, loff_t off,
cifs_dbg(FYI, "about to discard pages %llx-%llx\n", fstart, fend);
truncate_inode_pages_range(&target_inode->i_data,
min(fstart, i_size), fend);
+ netfs_wait_for_outstanding_io(target_inode);
fscache_invalidate(cifs_inode_cookie(target_inode), NULL, i_size, 0);
@@ -1513,6 +1515,7 @@ static loff_t cifs_remap_file_range(struct file *src_file, loff_t off,
if (rc)
CIFS_I(target_inode)->time = 0;
unlock:
+ filemap_invalidate_unlock(target_inode->i_mapping);
/* although unlocking in the reverse order from locking is not
strictly necessary here it is a little cleaner to be consistent */
unlock_two_nondirectories(src_inode, target_inode);
@@ -1536,6 +1539,9 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
struct cifs_tcon *target_tcon;
ssize_t rc;
+ if (len == 0)
+ return 0;
+
cifs_dbg(FYI, "copychunk range\n");
if (!src_file->private_data || !dst_file->private_data) {
@@ -1565,6 +1571,7 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
* server could even support copy of range where source = target
*/
lock_two_nondirectories(target_inode, src_inode);
+ filemap_invalidate_lock(target_inode->i_mapping);
cifs_dbg(FYI, "about to flush pages\n");
@@ -1590,11 +1597,24 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
* Start at the old EOF when extending so the folio straddling it, which
* may hold data written past EOF through an mmap, is dropped too.
*/
- rc = filemap_invalidate_inode(target_inode, true,
- min(destoff, i_size_read(target_inode)),
- destoff + len - 1);
- if (rc)
- goto unlock;
+ if (target_inode->i_mapping->nrpages) {
+ loff_t fstart = min(destoff, i_size_read(target_inode));
+ loff_t fend = destoff + len - 1;
+
+ unmap_mapping_pages(target_inode->i_mapping,
+ fstart >> PAGE_SHIFT,
+ (fend >> PAGE_SHIFT) -
+ (fstart >> PAGE_SHIFT) + 1,
+ false);
+ rc = filemap_write_and_wait_range(target_inode->i_mapping,
+ fstart, fend);
+ if (rc)
+ goto unlock;
+ invalidate_inode_pages2_range(target_inode->i_mapping,
+ fstart >> PAGE_SHIFT,
+ fend >> PAGE_SHIFT);
+ }
+ netfs_wait_for_outstanding_io(target_inode);
fscache_invalidate(cifs_inode_cookie(target_inode), NULL,
i_size_read(target_inode), 0);
@@ -1626,6 +1646,7 @@ ssize_t cifs_file_copychunk_range(unsigned int xid,
CIFS_I(target_inode)->time = 0;
unlock:
+ filemap_invalidate_unlock(target_inode->i_mapping);
/* although unlocking in the reverse order from locking is not
* strictly necessary here it is a little cleaner to be consistent
*/
--
2.55.0
next prev parent reply other threads:[~2026-09-30 3:28 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 3:28 [PATCH v3 00/15] netfs, cifs: data corruption fixes Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 01/15] netfs: clear post-EOF pagecache when extending a file via write Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 02/15] smb: client: clear post-EOF pagecache when extending a file via truncate Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 03/15] smb: client: discard post-EOF pagecache when extending a file via zero range Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 04/15] smb: client: discard post-EOF pagecache when extending a file via copy range Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 05/15] smb: client: discard post-EOF pagecache when extending a file via clone range Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 06/15] smb: client: flush and commit data before querying allocated ranges Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 07/15] smb: client: drain outstanding I/O before truncating on O_TRUNC open Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 08/15] smb: client: flush dirty data before zeroing a range Paulo Alcantara
2026-09-30 3:28 ` Paulo Alcantara [this message]
2026-09-30 3:28 ` [PATCH v3 10/15] smb: client: only require read lease for size-extending zero range Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 11/15] netfs: zero gaps in read-gaps folio to avoid writing back stale data Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 12/15] smb: client: only require read lease for size-extending preallocate Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 13/15] netfs: zero the tail of a short DIO/unbuffered read Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 14/15] smb: client: distinguish real EOF from a stale remote_i_size on read Paulo Alcantara
2026-09-30 3:28 ` [PATCH v3 15/15] smb: client: require stable pages for signed connections Paulo Alcantara
2026-09-30 15:17 ` [PATCH v3 00/15] netfs, cifs: data corruption fixes Namjae Jeon
2026-09-30 17:37 ` Paulo Alcantara
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260930032822.1835287-10-pc@manguebit.org \
--to=pc@manguebit.org \
--cc=bharathsm@microsoft.com \
--cc=brauner@kernel.org \
--cc=dhowells@redhat.com \
--cc=linkinjeon@kernel.org \
--cc=linux-cifs@vger.kernel.org \
--cc=netfs@lists.linux.dev \
--cc=ronniesahlberg@gmail.com \
--cc=sprasad@microsoft.com \
--cc=stable@vger.kernel.org \
--cc=tom@talpey.com \
--cc=willy@infradead.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox