* [PATCH v1 1/7] nfsd: use direct I/O only for the last operation in a COMPOUND
2026-10-04 19:40 [PATCH v1 0/7] NFSD short subjects Chuck Lever
@ 2026-10-04 19:40 ` Chuck Lever
2026-10-04 19:40 ` [PATCH v1 2/7] nfsd: account for page_base when advancing rq_next_page Chuck Lever
` (5 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Chuck Lever @ 2026-10-04 19:40 UTC (permalink / raw)
To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey; +Cc: linux-nfs
A direct read widens the requested byte range to the file system's
alignment and leaves the payload at a nonzero rq_res.page_base. The
xdr_stream encoder does not account for page_base:
xdr_reserve_space_vec() sizes each chunk from page_len alone, and
xdr_get_next_encode_buffer() opens each new page at its first byte.
The stream's position lags the end of the payload by page_base
bytes.
An NFSv4 operation encoded after such a READ overwrites the tail of
the payload, and the transport sends alignment padding in place of
that operation's result. The client sees corrupted file data
followed by an undecodable reply. NFSv2 and NFSv3 are unaffected:
their READ encoders pass page_base explicitly, and nothing is
encoded into the pages after the payload.
Fall back to buffered or DONTCACHE I/O when the READ is not the last
operation in its COMPOUND, as nfsd4_read() already does for splice
reads.
Fixes: d686e64e931c ("NFSD: Implement NFSD_IO_DIRECT for NFS READ")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/vfs.c | 22 ++++++++++++++++++++--
1 file changed, 20 insertions(+), 2 deletions(-)
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 4584d5b94fee..e052b9163692 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -37,6 +37,7 @@
#ifdef CONFIG_NFSD_V4
#include "acl.h"
#include "idmap.h"
+#include "xdr4.h"
#endif /* CONFIG_NFSD_V4 */
#include "nfsd.h"
@@ -1164,6 +1165,24 @@ nfsd_direct_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
eof, host_err);
}
+static bool nfsd_direct_read_ok(struct svc_rqst *rqstp, struct nfsd_file *nf)
+{
+ /* When dio_read_offset_align is zero, dio is not supported */
+ if (!nf->nf_dio_read_offset_align)
+ return false;
+ if (rqstp->rq_res.page_len)
+ return false;
+#ifdef CONFIG_NFSD_V4
+ /*
+ * The xdr_stream encoder ignores rq_res.page_base, so an operation
+ * encoded after a direct read would overwrite the payload's tail.
+ */
+ if (rqstp->rq_vers == 4 && !nfsd4_last_compound_op(rqstp))
+ return false;
+#endif
+ return true;
+}
+
/**
* nfsd_iter_read - Perform a VFS read using an iterator
* @rqstp: RPC transaction context
@@ -1197,8 +1216,7 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
case NFSD_IO_BUFFERED:
break;
case NFSD_IO_DIRECT:
- /* When dio_read_offset_align is zero, dio is not supported */
- if (nf->nf_dio_read_offset_align && !rqstp->rq_res.page_len)
+ if (nfsd_direct_read_ok(rqstp, nf))
return nfsd_direct_read(rqstp, fhp, nf, offset,
count, eof);
fallthrough;
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v1 2/7] nfsd: account for page_base when advancing rq_next_page
2026-10-04 19:40 [PATCH v1 0/7] NFSD short subjects Chuck Lever
2026-10-04 19:40 ` [PATCH v1 1/7] nfsd: use direct I/O only for the last operation in a COMPOUND Chuck Lever
@ 2026-10-04 19:40 ` Chuck Lever
2026-10-04 19:40 ` [PATCH v1 3/7] nfsd: locate the READ sink page from the reply buffer Chuck Lever
` (4 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Chuck Lever @ 2026-10-04 19:40 UTC (permalink / raw)
To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey; +Cc: linux-nfs
nfsd4_encode_operation() sets rq_next_page from xdr->page_ptr, which
the xdr_stream advances from page_len alone. A direct read leaves the
payload at a nonzero page_base, so its last page can lie beyond the
page xdr->page_ptr names. svc_rqst_release_pages() then skips that
page, and svc_alloc_arg() hands it to the next request as a receive
buffer while the transport still references it. svc_tcp_sendmsg()
splices the page into the socket without copying it, and svcrdma
keeps only the pages below rq_next_page until Send completion. The
tail of the READ payload reaches the client overwritten with bytes
of the next request.
Derive rq_next_page from page_base and page_len, the accounting that
xdr_truncate_encode() and the transports use to locate the payload.
Currently nfsd_direct_read() stores its alignment pad in page_base
even when the read returns no payload, and the pad can exceed the
size of the rq_respages array. Set page_base only when the read
returns payload, so that the new calculation stays within the pages
offered to the read.
Fixes: d686e64e931c ("NFSD: Implement NFSD_IO_DIRECT for NFS READ")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfs4xdr.c | 12 +++++++++---
fs/nfsd/vfs.c | 7 ++++---
2 files changed, 13 insertions(+), 6 deletions(-)
diff --git a/fs/nfsd/nfs4xdr.c b/fs/nfsd/nfs4xdr.c
index 7062c84f96dd..89230b3206ac 100644
--- a/fs/nfsd/nfs4xdr.c
+++ b/fs/nfsd/nfs4xdr.c
@@ -6714,6 +6714,7 @@ nfsd4_encode_operation(struct nfsd4_compoundres *resp, struct nfsd4_op *op)
struct svc_rqst *rqstp = resp->rqstp;
const struct nfsd4_operation *opdesc = op->opdesc;
unsigned int op_status_offset;
+ struct page **next_page;
nfsd4_enc encoder;
/*
@@ -6800,10 +6801,15 @@ nfsd4_encode_operation(struct nfsd4_compoundres *resp, struct nfsd4_op *op)
&op->status, XDR_UNIT);
release:
/*
- * Account for pages consumed while encoding this operation.
- * The xdr_stream primitives don't manage rq_next_page.
+ * Account for pages consumed while encoding this operation. The
+ * xdr_stream primitives don't manage rq_next_page, and
+ * xdr->page_ptr does not account for page_base. XDR padding can
+ * carry page_base + page_len past rq_page_end.
*/
- rqstp->rq_next_page = xdr->page_ptr + 1;
+ next_page = xdr->buf->pages +
+ DIV_ROUND_UP(xdr->buf->page_base + xdr->buf->page_len,
+ PAGE_SIZE);
+ rqstp->rq_next_page = min(next_page, rqstp->rq_page_end);
}
/**
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index e052b9163692..3f328378c805 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1143,9 +1143,6 @@ nfsd_direct_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
if (host_err >= 0) {
unsigned int pad = offset - dio_start;
- /* The returned payload starts after the pad */
- rqstp->rq_res.page_base = pad;
-
/* Compute the count of bytes to be returned */
if (host_err > pad + *count)
host_err = *count;
@@ -1153,6 +1150,10 @@ nfsd_direct_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
host_err -= pad;
else
host_err = 0;
+
+ /* The returned payload starts after the pad */
+ if (host_err)
+ rqstp->rq_res.page_base = pad;
} else if (unlikely(host_err == -EINVAL)) {
struct inode *inode = d_inode(fhp->fh_dentry);
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v1 3/7] nfsd: locate the READ sink page from the reply buffer
2026-10-04 19:40 [PATCH v1 0/7] NFSD short subjects Chuck Lever
2026-10-04 19:40 ` [PATCH v1 1/7] nfsd: use direct I/O only for the last operation in a COMPOUND Chuck Lever
2026-10-04 19:40 ` [PATCH v1 2/7] nfsd: account for page_base when advancing rq_next_page Chuck Lever
@ 2026-10-04 19:40 ` Chuck Lever
2026-10-04 19:40 ` [PATCH v1 4/7] nfsd: keep spliced pages below rq_next_page when a READ fails Chuck Lever
` (3 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Chuck Lever @ 2026-10-04 19:40 UTC (permalink / raw)
To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey; +Cc: linux-nfs
nfsd_iter_read() places file data at *rq_next_page, at an in-page
offset its caller computes from the xdr_buf. The page and the offset
come from different sources, and nothing keeps them consistent.
nfsd4_encode_readv() sets rq_next_page from the xdr_buf before each
call to keep them so, and any other caller has to do the same.
Have nfsd_iter_read() compute both the first sink page and the
in-page offset from rq_res.page_len, the accounting that
xdr_reserve_space_vec() extends once the read completes. NFSv2 and
NFSv3 hand the reply pages to READ untouched, so page_len is zero
there and the sink page is still *rq_next_page. nfsd_iter_read() and
nfsd_direct_read() now write rq_next_page and never read it.
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfs4xdr.c | 7 +------
fs/nfsd/vfs.c | 38 ++++++++++++++++++++------------------
fs/nfsd/vfs.h | 3 +--
3 files changed, 22 insertions(+), 26 deletions(-)
diff --git a/fs/nfsd/nfs4xdr.c b/fs/nfsd/nfs4xdr.c
index 89230b3206ac..b86d181b58c3 100644
--- a/fs/nfsd/nfs4xdr.c
+++ b/fs/nfsd/nfs4xdr.c
@@ -5303,17 +5303,12 @@ static __be32 nfsd4_encode_readv(struct nfsd4_compoundres *resp,
unsigned long maxcount)
{
struct xdr_stream *xdr = resp->xdr;
- unsigned int base = xdr->buf->page_len & ~PAGE_MASK;
unsigned int starting_len = xdr->buf->len;
__be32 zero = xdr_zero;
__be32 nfserr;
- resp->rqstp->rq_next_page = xdr->buf->pages +
- (xdr->buf->page_len >> PAGE_SHIFT);
-
nfserr = nfsd_iter_read(resp->rqstp, read->rd_fhp, read->rd_nf,
- read->rd_offset, &maxcount, base,
- &read->rd_eof);
+ read->rd_offset, &maxcount, &read->rd_eof);
read->rd_length = maxcount;
if (nfserr)
return nfserr;
diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c
index 3f328378c805..c7dd94697f2f 100644
--- a/fs/nfsd/vfs.c
+++ b/fs/nfsd/vfs.c
@@ -1097,14 +1097,13 @@ __be32 nfsd_splice_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
*
* Note that a direct read can be done only when the xdr_buf containing
* the NFS READ reply does not already have contents in its .pages array.
- * This is due to potentially restrictive alignment requirements on the
- * read buffer. When .page_len and @base are zero, the .pages array is
- * guaranteed to be page-aligned.
+ * Direct I/O alignment can be restrictive, and with .page_len zero the
+ * read buffer starts on a page boundary.
*/
static noinline_for_stack __be32
nfsd_direct_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset, unsigned long *count,
- u32 *eof)
+ struct page **page, u32 *eof)
{
u64 dio_start, dio_end;
unsigned long v, total;
@@ -1124,16 +1123,15 @@ nfsd_direct_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
v = 0;
total = dio_end - dio_start;
- while (total && v < rqstp->rq_maxpages &&
- rqstp->rq_next_page < rqstp->rq_page_end) {
+ while (total && v < rqstp->rq_maxpages && page < rqstp->rq_page_end) {
len = min_t(size_t, total, PAGE_SIZE);
- bvec_set_page(&rqstp->rq_bvec[v], *rqstp->rq_next_page,
- len, 0);
+ bvec_set_page(&rqstp->rq_bvec[v], *page, len, 0);
total -= len;
- ++rqstp->rq_next_page;
+ ++page;
++v;
}
+ rqstp->rq_next_page = page;
trace_nfsd_read_direct(rqstp, fhp, offset, *count - total);
iov_iter_bvec(&iter, ITER_DEST, rqstp->rq_bvec, v,
@@ -1191,19 +1189,24 @@ static bool nfsd_direct_read_ok(struct svc_rqst *rqstp, struct nfsd_file *nf)
* @nf: opened struct nfsd_file of file to be read
* @offset: starting byte offset
* @count: IN: requested number of bytes; OUT: number of bytes read
- * @base: offset in first page of read buffer
* @eof: OUT: set non-zero if operation reached the end of the file
*
* Some filesystems or situations cannot use nfsd_splice_read. This
* function is the slightly less-performant fallback for those cases.
+ * File data lands in @rqstp->rq_res.pages at the offset .page_len
+ * records. On return, @rqstp->rq_next_page points past the last page
+ * offered to the read.
*
* Returns nfs_ok on success, otherwise an nfserr stat value is
* returned.
*/
__be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset, unsigned long *count,
- unsigned int base, u32 *eof)
+ u32 *eof)
{
+ struct xdr_buf *buf = &rqstp->rq_res;
+ struct page **page = buf->pages + (buf->page_len >> PAGE_SHIFT);
+ unsigned int base = buf->page_len & ~PAGE_MASK;
struct file *file = nf->nf_file;
unsigned long v, total;
struct iov_iter iter;
@@ -1219,7 +1222,7 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
case NFSD_IO_DIRECT:
if (nfsd_direct_read_ok(rqstp, nf))
return nfsd_direct_read(rqstp, fhp, nf, offset,
- count, eof);
+ count, page, eof);
fallthrough;
case NFSD_IO_DONTCACHE:
if (file->f_op->fop_flags & FOP_DONTCACHE)
@@ -1231,17 +1234,16 @@ __be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
v = 0;
total = *count;
- while (total && v < rqstp->rq_maxpages &&
- rqstp->rq_next_page < rqstp->rq_page_end) {
+ while (total && v < rqstp->rq_maxpages && page < rqstp->rq_page_end) {
len = min_t(size_t, total, PAGE_SIZE - base);
- bvec_set_page(&rqstp->rq_bvec[v], *rqstp->rq_next_page,
- len, base);
+ bvec_set_page(&rqstp->rq_bvec[v], *page, len, base);
total -= len;
- ++rqstp->rq_next_page;
+ ++page;
++v;
base = 0;
}
+ rqstp->rq_next_page = page;
trace_nfsd_read_vector(rqstp, fhp, offset, *count - total);
iov_iter_bvec(&iter, ITER_DEST, rqstp->rq_bvec, v, *count - total);
@@ -1598,7 +1600,7 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
if (file->f_op->splice_read && nfsd_read_splice_ok(rqstp))
err = nfsd_splice_read(rqstp, fhp, file, offset, count, eof);
else
- err = nfsd_iter_read(rqstp, fhp, nf, offset, count, 0, eof);
+ err = nfsd_iter_read(rqstp, fhp, nf, offset, count, eof);
nfsd_file_put(nf);
trace_nfsd_read_done(rqstp, fhp, offset, *count);
diff --git a/fs/nfsd/vfs.h b/fs/nfsd/vfs.h
index 38f7d36bd4da..8c130473aa79 100644
--- a/fs/nfsd/vfs.h
+++ b/fs/nfsd/vfs.h
@@ -181,8 +181,7 @@ __be32 nfsd_splice_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
u32 *eof);
__be32 nfsd_iter_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
struct nfsd_file *nf, loff_t offset,
- unsigned long *count, unsigned int base,
- u32 *eof);
+ unsigned long *count, u32 *eof);
bool nfsd_read_splice_ok(struct svc_rqst *rqstp);
__be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp,
loff_t offset, unsigned long *count,
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v1 4/7] nfsd: keep spliced pages below rq_next_page when a READ fails
2026-10-04 19:40 [PATCH v1 0/7] NFSD short subjects Chuck Lever
` (2 preceding siblings ...)
2026-10-04 19:40 ` [PATCH v1 3/7] nfsd: locate the READ sink page from the reply buffer Chuck Lever
@ 2026-10-04 19:40 ` Chuck Lever
2026-10-04 19:40 ` [PATCH v1 5/7] nfsd: remove unreachable cancel of the layout fence work Chuck Lever
` (2 subsequent siblings)
6 siblings, 0 replies; 8+ messages in thread
From: Chuck Lever @ 2026-10-04 19:40 UTC (permalink / raw)
To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey; +Cc: linux-nfs
nfsd_splice_actor() replaces rq_respages entries with page cache
pages and advances rq_next_page past them. When
svc_encode_result_payload() then fails, nfsd4_encode_splice_read()
zeroes page_len, and nfsd4_encode_operation() computes rq_next_page
from the empty payload. svc_rqst_release_pages() and svc_alloc_arg()
handle only the entries below rq_next_page, so the spliced pages
stay in rq_respages. A later reply that encodes into those entries
overwrites the page cache of the file that was read.
svc_rdma_result_payload() returns -E2BIG when the READ payload is
larger than the Write chunk the client provided, so an NFS/RDMA
client can trigger the failure.
When an operation fails, do not move rq_next_page backward. A
successful operation still sets rq_next_page from the payload
length, because a short read through nfsd_iter_read() leaves
rq_next_page past pages the reply does not use.
Fixes: 76e5492b161f ("NFSD: Invoke svc_encode_result_payload() in "read" NFSD encoders")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfs4xdr.c | 9 ++++++++-
1 file changed, 8 insertions(+), 1 deletion(-)
diff --git a/fs/nfsd/nfs4xdr.c b/fs/nfsd/nfs4xdr.c
index b86d181b58c3..d9417ebbe9db 100644
--- a/fs/nfsd/nfs4xdr.c
+++ b/fs/nfsd/nfs4xdr.c
@@ -6804,7 +6804,14 @@ nfsd4_encode_operation(struct nfsd4_compoundres *resp, struct nfsd4_op *op)
next_page = xdr->buf->pages +
DIV_ROUND_UP(xdr->buf->page_base + xdr->buf->page_len,
PAGE_SIZE);
- rqstp->rq_next_page = min(next_page, rqstp->rq_page_end);
+ next_page = min(next_page, rqstp->rq_page_end);
+ /*
+ * A failed splice read leaves page cache pages in rq_respages
+ * above the truncated payload. Keep them below rq_next_page so
+ * that svc_rqst_release_pages() releases them.
+ */
+ if (!op->status || next_page > rqstp->rq_next_page)
+ rqstp->rq_next_page = next_page;
}
/**
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v1 5/7] nfsd: remove unreachable cancel of the layout fence work
2026-10-04 19:40 [PATCH v1 0/7] NFSD short subjects Chuck Lever
` (3 preceding siblings ...)
2026-10-04 19:40 ` [PATCH v1 4/7] nfsd: keep spliced pages below rq_next_page when a READ fails Chuck Lever
@ 2026-10-04 19:40 ` Chuck Lever
2026-10-04 19:40 ` [PATCH v1 6/7] sunrpc: preserve rq_daddrlen across request deferral Chuck Lever
2026-10-04 19:40 ` [PATCH v1 7/7] sunrpc: assign RQ_LOCAL from the transport on every receive Chuck Lever
6 siblings, 0 replies; 8+ messages in thread
From: Chuck Lever @ 2026-10-04 19:40 UTC (permalink / raw)
To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey; +Cc: linux-nfs
Clean up: Queued fence work holds a reference on its layout stateid.
nfsd4_layout_lm_breaker_timedout() takes the reference under ls_lock
before it queues the work, and the fence worker still holds it when
it queues a retry. nfsd4_free_layout_stateid() runs only after the
last reference is dropped, so its delayed_work_pending() check is
never true. The check reads as though sc_free cancels fence work,
but nfsd4_stop_layout_fence() is the only teardown path that does.
Signed-off-by: Chuck Lever <cel@kernel.org>
---
fs/nfsd/nfs4layouts.c | 7 -------
1 file changed, 7 deletions(-)
diff --git a/fs/nfsd/nfs4layouts.c b/fs/nfsd/nfs4layouts.c
index 7ac379fd5868..898776c9d72f 100644
--- a/fs/nfsd/nfs4layouts.c
+++ b/fs/nfsd/nfs4layouts.c
@@ -169,13 +169,6 @@ nfsd4_free_layout_stateid(struct nfs4_stid *stid)
trace_nfsd_layoutstate_free(&ls->ls_stid.sc_stateid);
- spin_lock(&ls->ls_lock);
- if (delayed_work_pending(&ls->ls_fence_work)) {
- spin_unlock(&ls->ls_lock);
- cancel_delayed_work_sync(&ls->ls_fence_work);
- } else
- spin_unlock(&ls->ls_lock);
-
spin_lock(&clp->cl_lock);
list_del_init(&ls->ls_perclnt);
spin_unlock(&clp->cl_lock);
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v1 6/7] sunrpc: preserve rq_daddrlen across request deferral
2026-10-04 19:40 [PATCH v1 0/7] NFSD short subjects Chuck Lever
` (4 preceding siblings ...)
2026-10-04 19:40 ` [PATCH v1 5/7] nfsd: remove unreachable cancel of the layout fence work Chuck Lever
@ 2026-10-04 19:40 ` Chuck Lever
2026-10-04 19:40 ` [PATCH v1 7/7] sunrpc: assign RQ_LOCAL from the transport on every receive Chuck Lever
6 siblings, 0 replies; 8+ messages in thread
From: Chuck Lever @ 2026-10-04 19:40 UTC (permalink / raw)
To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey; +Cc: linux-nfs
svc_defer() saves rq_daddr in the svc_deferred_req but not
rq_daddrlen, and svc_deferred_recv() restores only the address. The
thread that replays a deferred request keeps the rq_daddrlen of the
last request it processed. A newly created thread has a length of
zero.
gen_callback() and nlmsvc_lookup_host() copy rq_daddrlen bytes of
rq_daddr to form the source address of an NFSv4.0 or NLM callback
transport. When a new thread replays a deferred SETCLIENTID that
arrived over IPv6, the copy is empty and the source address is left
as AF_UNSPEC. xs_bind() then fails with -EAFNOSUPPORT, and the
callback transport never connects. A length left over from an IPv4
request copies only part of an IPv6 address. Found by code
inspection.
Record rq_daddrlen in the existing daddrlen field of struct
svc_deferred_req when the request is first deferred, and restore it
in svc_deferred_recv().
Fixes: 849a1cf13d43 ("SUNRPC: Replace svc_addr_u by sockaddr_storage")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
net/sunrpc/svc_xprt.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index bf5b2cdad35e..031cbac2a612 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -1343,6 +1343,7 @@ static struct cache_deferred_req *svc_defer(struct cache_req *req)
memcpy(&dr->addr, &rqstp->rq_addr, rqstp->rq_addrlen);
dr->addrlen = rqstp->rq_addrlen;
dr->daddr = rqstp->rq_daddr;
+ dr->daddrlen = rqstp->rq_daddrlen;
dr->secure = test_bit(RQ_SECURE, &rqstp->rq_flags);
dr->argslen = rqstp->rq_arg.len >> 2;
@@ -1377,6 +1378,7 @@ static noinline int svc_deferred_recv(struct svc_rqst *rqstp)
memcpy(&rqstp->rq_addr, &dr->addr, dr->addrlen);
rqstp->rq_addrlen = dr->addrlen;
rqstp->rq_daddr = dr->daddr;
+ rqstp->rq_daddrlen = dr->daddrlen;
rqstp->rq_xprt_ctxt = dr->xprt_ctxt;
/*
* ->xpo_recvfrom() is bypassed for a deferred request, so the
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread* [PATCH v1 7/7] sunrpc: assign RQ_LOCAL from the transport on every receive
2026-10-04 19:40 [PATCH v1 0/7] NFSD short subjects Chuck Lever
` (5 preceding siblings ...)
2026-10-04 19:40 ` [PATCH v1 6/7] sunrpc: preserve rq_daddrlen across request deferral Chuck Lever
@ 2026-10-04 19:40 ` Chuck Lever
6 siblings, 0 replies; 8+ messages in thread
From: Chuck Lever @ 2026-10-04 19:40 UTC (permalink / raw)
To: NeilBrown, Jeff Layton, Olga Kornievskaia, Dai Ngo, Tom Talpey; +Cc: linux-nfs
Only svc_tcp_recvfrom() assigns RQ_LOCAL. A request that arrives
over UDP or RDMA, or that svc_deferred_recv() replays, leaves the
svc_rqst with the value from the last TCP request its thread
received. When RQ_LOCAL is set, nfsd_vfs_write() sets
PF_LOCAL_THROTTLE, and dirty throttling considers only the backing
device being written. A stale RQ_LOCAL relaxes throttling that way
for a WRITE from a remote client. A stale clear bit withholds the
relaxation from a replayed WRITE that did arrive over loopback, and
NFSD can lock up behind the NFS client's dirty pages.
Assign RQ_LOCAL from XPT_LOCAL in svc_handle_xprt() before the
request is received. A deferred request is replayed on the
transport it arrived on, so struct svc_deferred_req needs no copy
of the flag. Only svc_tcp_accept() sets XPT_LOCAL. After the
change, RQ_LOCAL is clear for every UDP and RDMA request, including
one that arrives over loopback.
Fixes: ef11ce24875a ("SUNRPC: track whether a request is coming from a loop-back interface.")
Signed-off-by: Chuck Lever <cel@kernel.org>
---
net/sunrpc/svc_xprt.c | 2 ++
net/sunrpc/svcsock.c | 2 --
2 files changed, 2 insertions(+), 2 deletions(-)
diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c
index 031cbac2a612..acb6e58dc522 100644
--- a/net/sunrpc/svc_xprt.c
+++ b/net/sunrpc/svc_xprt.c
@@ -884,6 +884,8 @@ static void svc_handle_xprt(struct svc_rqst *rqstp, struct svc_xprt *xprt)
svc_xprt_received(xprt);
} else if (svc_xprt_reserve_slot(rqstp, xprt)) {
/* XPT_DATA|XPT_DEFERRED case: */
+ assign_bit(RQ_LOCAL, &rqstp->rq_flags,
+ test_bit(XPT_LOCAL, &xprt->xpt_flags));
rqstp->rq_deferred = svc_deferred_dequeue(xprt);
if (rqstp->rq_deferred)
len = svc_deferred_recv(rqstp);
diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c
index 1814e49c671d..b226bb3cd287 100644
--- a/net/sunrpc/svcsock.c
+++ b/net/sunrpc/svcsock.c
@@ -1287,8 +1287,6 @@ static int svc_tcp_recvfrom(struct svc_rqst *rqstp)
rqstp->rq_xprt_ctxt = NULL;
rqstp->rq_prot = IPPROTO_TCP;
- assign_bit(RQ_LOCAL, &rqstp->rq_flags,
- test_bit(XPT_LOCAL, &svsk->sk_xprt.xpt_flags));
/* Completing one message stops ->read_sock with whatever
* follows still queued, and no path from here re-arms XPT_DATA.
--
2.55.0
^ permalink raw reply related [flat|nested] 8+ messages in thread