* [PATCH 1/6] NFS/localio: fix nfs_local_dio_misaligned tracepoint
2026-09-28 15:54 [PATCH 0/6] NFS: LOCALIO fixes and FF_FLAGS_NO_IO_THRU_MDS fixes Mike Snitzer
@ 2026-09-28 15:54 ` Mike Snitzer
2026-09-28 15:54 ` [PATCH 2/6] NFS/localio: detect a short read or write before the iterator has moved Mike Snitzer
` (4 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Mike Snitzer @ 2026-09-28 15:54 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker; +Cc: linux-nfs
The intended focus of nfs_local_iters_setup_dio()'s call to
trace_nfs_local_dio_misaligned() is on the middle segment being
misaligned, yet the @offset passed in was local_dio->start_len.
It would appear this was a cut-n-paste bug from the preceding
nfs_local_iter_setup() call that passes local_dio->start_len.
Fix this by passing the @offset as local_dio->middle_offset and
calculate the start segment's offset rather than assume.
Example traces, before this fix:
python3-32744 [006] .l... 132946.352360:
nfs_local_dio_write: fileid=00:33:1286 fhandle=0xf1f7c10b
offset=1048759 count=1048576 mem_align=4 offset_align=512
start=1048759+329 middle=1049088+1048064 end=2097152+183
python3-32744 [006] .l... 132946.352360:
nfs_local_dio_misaligned: fileid=00:33:1286 fhandle=0xf1f7c10b
offset=329 count=1048064 mem_align=4 offset_align=512
start=329+329 middle=1049088+1048064 end=2097152+183
After this fix:
python3-32744 [006] .l... 132946.352360:
nfs_local_dio_write: fileid=00:33:1286 fhandle=0xf1f7c10b
offset=1048759 count=1048576 mem_align=4 offset_align=512
start=1048759+329 middle=1049088+1048064 end=2097152+183
python3-32744 [006] .l... 132946.352360:
nfs_local_dio_misaligned: fileid=00:33:1286 fhandle=0xf1f7c10b
offset=1049088 count=1048064 mem_align=4 offset_align=512
start=1048759+329 middle=1049088+1048064 end=2097152+183
Fixes: 6a218b9c3183e ("nfs/localio: do not issue misaligned DIO out-of-order")
Cc: stable@vger.kernel.org
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfs/localio.c | 2 +-
fs/nfs/nfstrace.h | 5 ++++-
2 files changed, 5 insertions(+), 2 deletions(-)
diff --git a/fs/nfs/localio.c b/fs/nfs/localio.c
index f42b6112a6139..63c38dea50cce 100644
--- a/fs/nfs/localio.c
+++ b/fs/nfs/localio.c
@@ -446,7 +446,7 @@ nfs_local_iters_setup_dio(struct nfs_local_kiocb *iocb, int rw,
if (unlikely(!iocb->iter_is_dio_aligned[n_iters])) {
trace_nfs_local_dio_misaligned(iocb->hdr->inode,
- local_dio->start_len, local_dio->middle_len, local_dio);
+ local_dio->middle_offset, local_dio->middle_len, local_dio);
return 0; /* no DIO-aligned IO possible */
}
iocb->end_iter_index = n_iters;
diff --git a/fs/nfs/nfstrace.h b/fs/nfs/nfstrace.h
index b15c1732c8692..a32c76df72db5 100644
--- a/fs/nfs/nfstrace.h
+++ b/fs/nfs/nfstrace.h
@@ -1772,7 +1772,10 @@ DECLARE_EVENT_CLASS(nfs_local_dio_class,
__entry->count = count;
__entry->mem_align = local_dio->mem_align;
__entry->offset_align = local_dio->offset_align;
- __entry->start = offset;
+ if (local_dio->start_len)
+ __entry->start = local_dio->middle_offset - local_dio->start_len;
+ else
+ __entry->start = 0;
__entry->start_len = local_dio->start_len;
__entry->middle = local_dio->middle_offset;
__entry->middle_len = local_dio->middle_len;
--
2.44.0
^ permalink raw reply related [flat|nested] 7+ messages in thread* [PATCH 2/6] NFS/localio: detect a short read or write before the iterator has moved
2026-09-28 15:54 [PATCH 0/6] NFS: LOCALIO fixes and FF_FLAGS_NO_IO_THRU_MDS fixes Mike Snitzer
2026-09-28 15:54 ` [PATCH 1/6] NFS/localio: fix nfs_local_dio_misaligned tracepoint Mike Snitzer
@ 2026-09-28 15:54 ` Mike Snitzer
2026-09-28 15:54 ` [PATCH 3/6] NFS/localio: report the stability a DIO WRITE actually has Mike Snitzer
` (3 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Mike Snitzer @ 2026-09-28 15:54 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker; +Cc: linux-nfs
nfs_local_call_read() and nfs_local_call_write() issue a misaligned
DIO as up to three segments and must stop at the first short one, or
the next segment lands at the wrong file offset. They test for that
by comparing the bytes transferred against iov_iter_count() of the
segment, read after read_iter() or write_iter() has advanced the
iterator, when the count is the residual rather than the request:
before the call: count == expected
after the call: count == expected - status
The test is therefore status < expected - status, and it fires only
when less than half of the segment was transferred. A short I/O that
covers between 50% and 99% of a segment slips through: the loop moves
on with ki_pos advanced by the short amount, the next segment is
issued at the wrong offset, and the header's byte count is
over-reported to the caller. A write that slips through also never
sets NFS_CONTEXT_WRITE_SYNC, which is what makes the writes that
follow a short one synchronous.
Take the segment's count before issuing it and compare against that.
This is the same defect that commit 250ec14932d5 ("nfsd: fix
partial-write detection in nfsd_direct_write") fixed in NFSD's version
of this loop.
Fixes: c817248fc831 ("nfs/localio: add proper O_DIRECT support for READ and WRITE")
Fixes: d0497dd27452 ("nfs/localio: backfill missing partial read support for misaligned DIO")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfs/localio.c | 15 ++++++++++-----
1 file changed, 10 insertions(+), 5 deletions(-)
diff --git a/fs/nfs/localio.c b/fs/nfs/localio.c
index 63c38dea50cce..9fcae2391b726 100644
--- a/fs/nfs/localio.c
+++ b/fs/nfs/localio.c
@@ -674,6 +674,8 @@ static void nfs_local_call_read(struct work_struct *work)
n_iters = atomic_read(&iocb->n_iters);
for (int i = 0; i < n_iters ; i++) {
+ size_t expected;
+
if (iocb->iter_is_dio_aligned[i]) {
iocb->kiocb.ki_flags |= IOCB_DIRECT;
/* Only use AIO completion if DIO-aligned segment is last */
@@ -684,6 +686,8 @@ static void nfs_local_call_read(struct work_struct *work)
} else
iocb->kiocb.ki_flags &= ~IOCB_DIRECT;
+ /* read_iter() advances the iterator: measure it beforehand */
+ expected = iov_iter_count(&iocb->iters[i]);
scoped_with_creds(filp->f_cred)
status = filp->f_op->read_iter(&iocb->kiocb, &iocb->iters[i]);
@@ -691,7 +695,7 @@ static void nfs_local_call_read(struct work_struct *work)
continue;
/* Break on completion, errors, or short reads */
if (nfs_local_pgio_done(iocb, status) || status < 0 ||
- (size_t)status < iov_iter_count(&iocb->iters[i])) {
+ (size_t)status < expected) {
nfs_local_read_iocb_done(iocb);
break;
}
@@ -890,7 +894,7 @@ static void nfs_local_call_write(struct work_struct *work)
file_start_write(filp);
n_iters = atomic_read(&iocb->n_iters);
for (int i = 0; i < n_iters ; i++) {
- size_t icount;
+ size_t expected;
if (iocb->iter_is_dio_aligned[i]) {
iocb->kiocb.ki_flags |= IOCB_DIRECT;
@@ -902,16 +906,17 @@ static void nfs_local_call_write(struct work_struct *work)
} else
iocb->kiocb.ki_flags &= ~IOCB_DIRECT;
+ /* write_iter() advances the iterator: measure it beforehand */
+ expected = iov_iter_count(&iocb->iters[i]);
scoped_with_creds(filp->f_cred)
status = filp->f_op->write_iter(&iocb->kiocb, &iocb->iters[i]);
if (status == -EIOCBQUEUED)
continue;
/* Break on completion, errors, or short writes */
- icount = iov_iter_count(&iocb->iters[i]);
if (nfs_local_pgio_done(iocb, status) || status < 0 ||
- (size_t)status < icount) {
- if ((size_t)status < icount) {
+ (size_t)status < expected) {
+ if ((size_t)status < expected) {
struct nfs_lock_context *ctx =
iocb->hdr->req->wb_lock_context;
--
2.44.0
^ permalink raw reply related [flat|nested] 7+ messages in thread* [PATCH 3/6] NFS/localio: report the stability a DIO WRITE actually has
2026-09-28 15:54 [PATCH 0/6] NFS: LOCALIO fixes and FF_FLAGS_NO_IO_THRU_MDS fixes Mike Snitzer
2026-09-28 15:54 ` [PATCH 1/6] NFS/localio: fix nfs_local_dio_misaligned tracepoint Mike Snitzer
2026-09-28 15:54 ` [PATCH 2/6] NFS/localio: detect a short read or write before the iterator has moved Mike Snitzer
@ 2026-09-28 15:54 ` Mike Snitzer
2026-09-28 15:54 ` [PATCH 4/6] pNFS/flexfiles: honor FF_FLAGS_NO_IO_THRU_MDS over the mdsthreshold hint Mike Snitzer
` (2 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Mike Snitzer @ 2026-09-28 15:54 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker; +Cc: linux-nfs
Since commit d32ddfeb5593 ("nfs/localio: Ensure DIO WRITE's IO on
stable storage upon completion") a DIO WRITE issued through LOCALIO
is persisted before it completes: nfs_local_iters_init() sets
IOCB_DSYNC|IOCB_SYNC on it despite whatever the caller asked for, so
that the buffered head and tail and the O_DIRECT middle of a
misaligned write cannot complete out of order. The reply still
reported the stability that was asked for, so an UNSTABLE write came
back UNSTABLE, the client kept its pages on the commit list, and the
COMMIT that followed ran an fsync for data that is already on stable
storage.
Report the stability the kiocb actually carries instead: FILE_SYNC
when IOCB_SYNC is set, DATA_SYNC when only IOCB_DSYNC is. A write
told FILE_SYNC has no reason to COMMIT and sends none. Buffered
writes, and any write that asked for at least what it got, are
unchanged.
Fixes: d32ddfeb5593 ("nfs/localio: Ensure DIO WRITE's IO on stable storage upon completion")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-fable-5-1
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfs/localio.c | 14 +++++++++++++-
1 file changed, 13 insertions(+), 1 deletion(-)
diff --git a/fs/nfs/localio.c b/fs/nfs/localio.c
index 9fcae2391b726..db7aed7254380 100644
--- a/fs/nfs/localio.c
+++ b/fs/nfs/localio.c
@@ -936,6 +936,7 @@ static void nfs_local_do_write(struct nfs_local_kiocb *iocb,
const struct rpc_call_ops *call_ops)
{
struct nfs_pgio_header *hdr = iocb->hdr;
+ enum nfs3_stable_how committed = hdr->args.stable;
dprintk("%s: vfs_write count=%u pos=%llu %s\n",
__func__, hdr->args.count, hdr->args.offset,
@@ -954,9 +955,20 @@ static void nfs_local_do_write(struct nfs_local_kiocb *iocb,
iocb->kiocb.ki_flags |= IOCB_DSYNC|IOCB_SYNC;
}
+ /*
+ * Report the stability the write will actually have. A DIO WRITE
+ * is persisted before it completes whatever was asked for, see
+ * nfs_local_iters_init(), and a caller told so has no reason to
+ * COMMIT data that is already on stable storage.
+ */
+ if (iocb->kiocb.ki_flags & IOCB_SYNC)
+ committed = NFS_FILE_SYNC;
+ else if (iocb->kiocb.ki_flags & IOCB_DSYNC)
+ committed = NFS_DATA_SYNC;
+
nfs_local_pgio_init(hdr, call_ops);
- nfs_set_local_verifier(hdr->inode, hdr->res.verf, hdr->args.stable);
+ nfs_set_local_verifier(hdr->inode, hdr->res.verf, committed);
INIT_WORK(&iocb->work, nfs_local_call_write);
if (nfs_local_defer_io())
--
2.44.0
^ permalink raw reply related [flat|nested] 7+ messages in thread* [PATCH 4/6] pNFS/flexfiles: honor FF_FLAGS_NO_IO_THRU_MDS over the mdsthreshold hint
2026-09-28 15:54 [PATCH 0/6] NFS: LOCALIO fixes and FF_FLAGS_NO_IO_THRU_MDS fixes Mike Snitzer
` (2 preceding siblings ...)
2026-09-28 15:54 ` [PATCH 3/6] NFS/localio: report the stability a DIO WRITE actually has Mike Snitzer
@ 2026-09-28 15:54 ` Mike Snitzer
2026-09-28 15:54 ` [PATCH 5/6] pNFS/flexfiles: don't read through MDS when previous layout forbid it Mike Snitzer
2026-09-28 15:54 ` [PATCH 6/6] pNFS/flexfiles: don't reset to MDS for v4 error " Mike Snitzer
5 siblings, 0 replies; 7+ messages in thread
From: Mike Snitzer @ 2026-09-28 15:54 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker; +Cc: linux-nfs
A flexfiles layout may carry FF_FLAGS_NO_IO_THRU_MDS, which tells the
client that this file's data must not be read or written through the
metadata server.
The RFC 5661 mdsthreshold hint pulls the other way. When a server returns
one, pnfs_within_mdsthreshold() answers "use MDS I/O" for every I/O below
the threshold, pnfs_update_layout() then hands back no segment, and
ff_layout_pg_init_read() reads through the MDS - the very thing the
layout forbids. A server that sets both is asking for two incompatible
things; the layout flag is the stronger statement, and the one whose
violation the client cannot recover from.
Record the policy on the inode, where it outlives the layout hdr that
carried it, and consult it before the hint:
- ff_layout_alloc_lseg() sets NFS_INO_NO_IO_THRU_MDS beside the existing
sticky hdr-level NFS4_FF_HDR_NO_IO_THRU_MDS bit.
- pnfs_within_mdsthreshold() returns false at once for such an inode, so
no I/O is diverted to the MDS on the strength of the hint.
The bit is never cleared for the life of the in-core inode. Servers are
assumed to be consistent in their no-fallback policy per file, which is
the assumption ff_layout_hdr_no_fallback_to_mds() already makes; a server
that was not would lose the effect of its mdsthreshold hint - a SHOULD -
on an inode that is already in core, and nothing else.
Fixes: 260074cd8413 ("pNFS/flexfiles: Add support for FF_FLAGS_NO_IO_THRU_MDS")
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfs/flexfilelayout/flexfilelayout.c | 5 ++++-
fs/nfs/internal.h | 29 ++++++++++++++++++++++++++
fs/nfs/pnfs.c | 12 ++++++++++-
3 files changed, 44 insertions(+), 2 deletions(-)
diff --git a/fs/nfs/flexfilelayout/flexfilelayout.c b/fs/nfs/flexfilelayout/flexfilelayout.c
index 94cc324b591f4..7d45401ee5b11 100644
--- a/fs/nfs/flexfilelayout/flexfilelayout.c
+++ b/fs/nfs/flexfilelayout/flexfilelayout.c
@@ -639,9 +639,12 @@ ff_layout_alloc_lseg(struct pnfs_layout_hdr *lh,
if (!p)
goto out_sort_mirrors;
fls->flags = be32_to_cpup(p);
- if (fls->flags & FF_FLAGS_NO_IO_THRU_MDS)
+ if (fls->flags & FF_FLAGS_NO_IO_THRU_MDS) {
set_bit(NFS4_FF_HDR_NO_IO_THRU_MDS,
&FF_LAYOUT_FROM_HDR(lh)->flags);
+ /* Outlives this layout hdr; see NFS_INO_NO_IO_THRU_MDS */
+ nfs_set_no_io_thru_mds(lh->plh_inode);
+ }
p = xdr_inline_decode(&stream, 4);
if (!p)
diff --git a/fs/nfs/internal.h b/fs/nfs/internal.h
index 4c6fb5a612252..50273e41f27df 100644
--- a/fs/nfs/internal.h
+++ b/fs/nfs/internal.h
@@ -543,6 +543,35 @@ static inline bool nfs_file_io_is_buffered(struct nfs_inode *nfsi)
return test_bit(NFS_INO_ODIRECT, &nfsi->flags) == 0;
}
+/*
+ * Module-private nfs_inode->flags bit (not in <linux/nfs_fs.h>, so no
+ * exported layout changes): set, and never cleared for the life of the
+ * in-core inode, once a layout driver has seen the server forbid I/O
+ * through the MDS for this file (flexfiles sets it from
+ * FF_FLAGS_NO_IO_THRU_MDS in ff_layout_alloc_lseg()). The RFC 5661
+ * mdsthreshold hint must not be acted on for such a file, since the only
+ * thing pnfs_within_mdsthreshold() can ask for is the one thing the
+ * layout forbids. Servers are assumed to be consistent in their
+ * no-fallback policy per file, the same assumption
+ * ff_layout_hdr_no_fallback_to_mds() already makes; if one were not, the
+ * only effect is that its mdsthreshold hint - a SHOULD - stops being
+ * honored for an inode that is already in core.
+ */
+#define NFS_INO_NO_IO_THRU_MDS (30)
+
+static inline void nfs_set_no_io_thru_mds(struct inode *inode)
+{
+ struct nfs_inode *nfsi = NFS_I(inode);
+
+ if (!test_bit(NFS_INO_NO_IO_THRU_MDS, &nfsi->flags))
+ set_bit(NFS_INO_NO_IO_THRU_MDS, &nfsi->flags);
+}
+
+static inline bool nfs_no_io_thru_mds(struct inode *inode)
+{
+ return test_bit(NFS_INO_NO_IO_THRU_MDS, &NFS_I(inode)->flags);
+}
+
/* Must be called with exclusively locked inode->i_rwsem */
static inline void nfs_file_block_o_direct(struct nfs_inode *nfsi)
{
diff --git a/fs/nfs/pnfs.c b/fs/nfs/pnfs.c
index 65b806a76f8de..264d0c8efb6b6 100644
--- a/fs/nfs/pnfs.c
+++ b/fs/nfs/pnfs.c
@@ -2021,7 +2021,9 @@ pnfs_find_lseg(struct pnfs_layout_hdr *lo,
* is set to IOMODE_READ for a READ request, and set to IOMODE_RW for a
* WRITE request.
*
- * A return of true means use MDS I/O.
+ * A return of true means use MDS I/O, so for a file whose layout forbids
+ * that (see NFS_INO_NO_IO_THRU_MDS) the hint is not evaluated at all:
+ * there is no answer it could give that this client may act on.
*
* From rfc 5661:
* If a file's size is smaller than the file size threshold, data accesses
@@ -2042,6 +2044,14 @@ static bool pnfs_within_mdsthreshold(struct nfs_open_context *ctx,
if (t == NULL)
return ret;
+ /*
+ * The server has told a layout driver that this file's I/O may not
+ * go through the MDS (flexfiles FF_FLAGS_NO_IO_THRU_MDS). Honor
+ * that over its own mdsthreshold hint.
+ */
+ if (nfs_no_io_thru_mds(ino))
+ return ret;
+
dprintk("%s bm=0x%x rd_sz=%llu wr_sz=%llu rd_io=%llu wr_io=%llu\n",
__func__, t->bm, t->rd_sz, t->wr_sz, t->rd_io_sz, t->wr_io_sz);
--
2.44.0
^ permalink raw reply related [flat|nested] 7+ messages in thread* [PATCH 5/6] pNFS/flexfiles: don't read through MDS when previous layout forbid it
2026-09-28 15:54 [PATCH 0/6] NFS: LOCALIO fixes and FF_FLAGS_NO_IO_THRU_MDS fixes Mike Snitzer
` (3 preceding siblings ...)
2026-09-28 15:54 ` [PATCH 4/6] pNFS/flexfiles: honor FF_FLAGS_NO_IO_THRU_MDS over the mdsthreshold hint Mike Snitzer
@ 2026-09-28 15:54 ` Mike Snitzer
2026-09-28 15:54 ` [PATCH 6/6] pNFS/flexfiles: don't reset to MDS for v4 error " Mike Snitzer
5 siblings, 0 replies; 7+ messages in thread
From: Mike Snitzer @ 2026-09-28 15:54 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker; +Cc: linux-nfs
ff_layout_pg_init_read() falls back to the MDS whenever
pnfs_update_layout() returns no segment and no error: a bulk
CB_LAYOUTRECALL (FSID, DEVICEID or ALL), LAYOUTGETs blocked behind a
layout being returned, or a LAYOUTGET refused with
NFS4ERR_LAYOUTUNAVAILABLE or NFS4ERR_TOOSMALL, which set the per-iomode
fail bit for PNFS_LAYOUTGET_RETRY_TIMEOUT. That is the read-side twin
of the write path fixed by commit 1d62e659c0bf ("NFSv4/flexfiles:
honor FF_FLAGS_NO_IO_THRU_MDS in pg_get_mirror_count_write"), and it
is reached in exactly the situation that matters most: the file had a
valid layout, the layout had to go back, and the next one has not
arrived yet.
The write side could answer -EAGAIN because writeback owns the retry and
redirties the page. A read has no such owner, so -EAGAIN here would
surface as an I/O error to the application. Wait for a layout instead,
on the same terms the function already uses when the layout lookup
itself returns -EAGAIN: sleep a second and retry, bounded by
pg_maxretrans, which ff_layout_pg_init_read() sets only for soft and
softerr mounts. A hard mount therefore waits, which is what a hard
mount is for, and matches what the FF_FLAGS_NO_IO_THRU_MDS check on the
data server selection a few lines above already does. That includes
waiting out NFS4ERR_LAYOUTUNAVAILABLE: a file that was given a layout
with this flag must not be read off the MDS, whatever the server says
afterwards.
The test asks the inode (NFS_INO_NO_IO_THRU_MDS) rather than the layout
hdr, because in this path there is no segment, and the hdr that carried
the flag may already have been destroyed. Where a segment is in hand -
the data server selection above, ff_layout_read_pagelist() - the
existing per-segment ff_layout_no_fallback_to_mds() test stays, since it
is the more precise question.
Fixes: 260074cd8413 ("pNFS/flexfiles: Add support for FF_FLAGS_NO_IO_THRU_MDS")
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfs/flexfilelayout/flexfilelayout.c | 38 ++++++++++++++++++--------
1 file changed, 26 insertions(+), 12 deletions(-)
diff --git a/fs/nfs/flexfilelayout/flexfilelayout.c b/fs/nfs/flexfilelayout/flexfilelayout.c
index 7d45401ee5b11..84d22f715653d 100644
--- a/fs/nfs/flexfilelayout/flexfilelayout.c
+++ b/fs/nfs/flexfilelayout/flexfilelayout.c
@@ -1073,19 +1073,33 @@ ff_layout_pg_init_read(struct nfs_pageio_descriptor *pgio,
if (pgio->pg_error < 0) {
if (pgio->pg_error != -EAGAIN)
return;
- /* Retry getting layout segment if lower layer returned -EAGAIN */
- if (pgio->pg_maxretrans && req->wb_nio++ > pgio->pg_maxretrans) {
- if (NFS_SERVER(pgio->pg_inode)->flags & NFS_MOUNT_SOFTERR)
- pgio->pg_error = -ETIMEDOUT;
- else
- pgio->pg_error = -EIO;
- return;
- }
- pgio->pg_error = 0;
- /* Sleep for 1 second before retrying */
- ssleep(1);
- goto retry;
+ goto retry_nolseg;
}
+ /*
+ * No segment, and no error to report either: pnfs_update_layout()
+ * simply has nothing to give (NFS_LAYOUT_BULK_RECALL, a failed
+ * pnfs_layout_io_test, blocked LAYOUTGETs, a layout being
+ * returned). If the server forbids reading this file through the
+ * MDS there is no fallback to take, so wait for a layout on the
+ * same terms as the -EAGAIN above. The layout hdr that carried
+ * FF_FLAGS_NO_IO_THRU_MDS may itself be gone by now, which is why
+ * this asks the inode and not the hdr.
+ */
+ if (!nfs_no_io_thru_mds(pgio->pg_inode))
+ goto out_mds;
+retry_nolseg:
+ /* Retry getting layout segment if lower layer returned -EAGAIN */
+ if (pgio->pg_maxretrans && req->wb_nio++ > pgio->pg_maxretrans) {
+ if (NFS_SERVER(pgio->pg_inode)->flags & NFS_MOUNT_SOFTERR)
+ pgio->pg_error = -ETIMEDOUT;
+ else
+ pgio->pg_error = -EIO;
+ return;
+ }
+ pgio->pg_error = 0;
+ /* Sleep for 1 second before retrying */
+ ssleep(1);
+ goto retry;
out_mds:
trace_pnfs_mds_fallback_pg_init_read(pgio->pg_inode,
0, NFS4_MAX_UINT64, IOMODE_READ,
--
2.44.0
^ permalink raw reply related [flat|nested] 7+ messages in thread* [PATCH 6/6] pNFS/flexfiles: don't reset to MDS for v4 error when previous layout forbid it
2026-09-28 15:54 [PATCH 0/6] NFS: LOCALIO fixes and FF_FLAGS_NO_IO_THRU_MDS fixes Mike Snitzer
` (4 preceding siblings ...)
2026-09-28 15:54 ` [PATCH 5/6] pNFS/flexfiles: don't read through MDS when previous layout forbid it Mike Snitzer
@ 2026-09-28 15:54 ` Mike Snitzer
5 siblings, 0 replies; 7+ messages in thread
From: Mike Snitzer @ 2026-09-28 15:54 UTC (permalink / raw)
To: Trond Myklebust, Anna Schumaker; +Cc: linux-nfs
ff_layout_async_handle_error_v4() ends in -NFS4ERR_RESET_TO_MDS, which
makes ff_layout_read_release() / ff_layout_write_release() call
ff_layout_reset_read() or ff_layout_reset_write(hdr, false), and those
call pnfs_read_done_resend_to_mds() / pnfs_write_done_resend_to_mds().
Those build a page I/O descriptor with force_mds set, so the flexfiles
->pg_init() never runs and no layout is consulted anywhere on the way:
the READ or WRITE goes to the metadata server unconditionally. For a
layout carrying FF_FLAGS_NO_IO_THRU_MDS that is exactly what the server
forbade.
The fall-through into reset: is already safe, because
ff_layout_avoid_mds_available_ds() returns true for a segment with the
flag and so answers -NFS4ERR_RESET_TO_PNFS first. What is not safe is
the goto: the invalid-layout errors (NFS4ERR_PNFS_NO_LAYOUT, STALE,
BADHANDLE, ISDIR, FHEXPIRED, WRONG_TYPE) call pnfs_destroy_layout() and
jump straight to reset:, over that check. That is the case where the
file had a perfectly good layout, something forced it to be given back,
and the client answers by writing to the MDS behind the server's back.
Test at reset: itself, so it covers the goto and any future one, and ask
the inode (NFS_INO_NO_IO_THRU_MDS): the callers that jump here have just
destroyed the layout hdr, so the hdr-level flag is no longer reachable,
while the inode-level one is set for the life of the in-core inode.
-NFS4ERR_RESET_TO_PNFS is the right answer rather than an error. It is
what the fall-through already returns for the same flag, and the
resends it drives - ff_layout_resend_pnfs_read() and
ff_layout_reset_write(hdr, true) - re-enter the pgio path with force_mds
clear, so a fresh LAYOUTGET is taken, which is precisely what a
destroyed layout needs. A LAYOUTGET that fails fatally makes
pnfs_update_layout() return an error, which ->pg_init() reports; a
non-fatal refusal (NFS4ERR_LAYOUTUNAVAILABLE, NFS4ERR_TOOSMALL) returns
no segment instead, and the client keeps retrying through pNFS rather
than falling back.
ff_layout_async_handle_error_v3() needs no such change: it returns only
-NFS4ERR_RESET_TO_PNFS or -EAGAIN.
Fixes: 260074cd8413 ("pNFS/flexfiles: Add support for FF_FLAGS_NO_IO_THRU_MDS")
Assisted-by: Claude:claude-opus-5[1m]
Signed-off-by: Mike Snitzer <snitzer@kernel.org>
---
fs/nfs/flexfilelayout/flexfilelayout.c | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/fs/nfs/flexfilelayout/flexfilelayout.c b/fs/nfs/flexfilelayout/flexfilelayout.c
index 84d22f715653d..d4f2b57e9d5b2 100644
--- a/fs/nfs/flexfilelayout/flexfilelayout.c
+++ b/fs/nfs/flexfilelayout/flexfilelayout.c
@@ -1423,6 +1423,23 @@ static int ff_layout_async_handle_error_v4(struct rpc_task *task,
if (ff_layout_avoid_mds_available_ds(lseg))
return -NFS4ERR_RESET_TO_PNFS;
reset:
+ /*
+ * FF_FLAGS_NO_IO_THRU_MDS: never resend through the MDS. The
+ * caller would do so with force_mds set (ff_layout_reset_read(),
+ * ff_layout_reset_write(hdr, false) -> pnfs_*_done_resend_to_mds()),
+ * which builds a descriptor out of the plain MDS page ops and so
+ * consults no layout at all -- this is the last point at which the
+ * policy can still be applied. Retry through pNFS instead, which
+ * takes a fresh LAYOUTGET; that is also the right answer for the
+ * invalid-layout cases that jump here, since they have just called
+ * pnfs_destroy_layout(). Having done so they can no longer ask the
+ * layout hdr about the flag, hence the inode.
+ */
+ if (nfs_no_io_thru_mds(inode)) {
+ dprintk("%s Retry through pNFS, no MDS fallback. Error %d\n",
+ __func__, task->tk_status);
+ return -NFS4ERR_RESET_TO_PNFS;
+ }
dprintk("%s Retry through MDS. Error %d\n", __func__,
task->tk_status);
return -NFS4ERR_RESET_TO_MDS;
--
2.44.0
^ permalink raw reply related [flat|nested] 7+ messages in thread