From: Tim Menninger <tmenninger@everpuredata.com>
To: Trond Myklebust <trondmy@kernel.org>, Anna Schumaker <anna@kernel.org>
Cc: Chuck Lever <cel@kernel.org>, Jeff Layton <jlayton@kernel.org>,
linux-nfs@vger.kernel.org, linux-kernel@vger.kernel.org,
netdev@vger.kernel.org,
Shiva Lingappa <slingappa@everpuredata.com>,
Eric Badger <ebadger@everpuredata.com>,
Jon Curley <jcurley@everpuredata.com>
Subject: [PATCH 0/2] SUNRPC/NFS: limit transport disconnects when cancelling flexfiles I/O
Date: Thu, 24 Sep 2026 20:45:23 +0000 [thread overview]
Message-ID: <20260924204525.3381590-1-tmenninger@everpuredata.com> (raw)
This series narrows the transport disconnects performed when flexfiles
cancels I/O for a recalled or revoked layout segment.
Today ff_layout_cancel_io() first cancels matching RPC tasks and then
calls rpc_clnt_disconnect() if any tasks matched. rpc_clnt_disconnect()
operates on the entire rpc_clnt, so for a data server client using
multiple transports it disconnects transports that may have had no I/O
associated with the affected layout segment.
I encountered this while testing pNFS over RPC/RDMA during data server
recovery. Restarting the NFS service on a data server while sustained
direct-read I/O was active could cause repeated reconnect activity and a
large throughput drop. The client-wide disconnect from
ff_layout_cancel_io() amplified the disruption by tearing down otherwise
unaffected transports.
The first patch adds a SUNRPC helper that combines task cancellation with
request-scoped conditional disconnects. Matching tasks are cancelled as
before. A subsequent transport-queue pass marks matching requests
belonging to the same rpc_clnt that remain on the transmit or receive
queues so that, when released, xprt_conditional_disconnect() is called
using the request's saved connection cookie.
The disconnect state is stored on struct rpc_rqst rather than struct
rpc_task. A task can release one request and later acquire another, so a
task-scoped indication could otherwise be consumed by a successor
request and applied to the wrong transport generation.
The request-marking pass also preserves the rpc_clnt selection used by
rpc_cancel_tasks(). An rpc_xprt may be shared by multiple rpc_clnt
instances, so requests belonging to other clients sharing the same
transport are not marked.
The second patch switches flexfiles to the new helper, preserving its
I/O cancellation behavior while avoiding client-wide transport
disconnects.
Testing:
The code in this series does not depend on:
https://lore.kernel.org/all/20260916232906.4101987-1-tmenninger@everpuredata.com/
However, I used that change while collecting the results below because,
without it, the client enters the earlier recovery livelock before the
transport-disconnect behavior can be observed cleanly.
As a control experiment, I also tested skipping the
ff_layout_cancel_io() task-cancellation and client-disconnect path
entirely for RPC/RDMA. Doing so eliminates the reconnect pathology. This
shows that the pathology depends on the cancellation/disconnect path
introduced by commit b739a5bd9d9f ("NFSv4/flexfiles: Cancel I/O if the
layout is recalled or revoked"), but simply opting RPC/RDMA out of that
behavior would also remove the cancellation that is intended to drain
I/O associated with a recalled or revoked layout segment.
This series instead preserves that cancellation while limiting transport
disconnects to matching requests that remain in flight.
I have repeated the data-server restart test many times and the behavior
is consistent. The counters below are intentionally collected after the
restart event, once throughput has partially recovered and the client has
entered the post-recovery state in which it remains until the workload
completes.
This distinction matters because some transport activity during the
data-server restart itself is expected. The problem addressed here is
the continued reconnect activity after that transition has completed:
once recovery has settled, the client should not continue cycling
otherwise healthy RPC/RDMA connections.
Over a 30-second post-recovery sampling window, this series reduces
sunrpc events from 4,526 to 289, a roughly
16-fold reduction. RPC/RDMA disconnect and connect events fall by about
6.6-fold. During the same interval, read throughput improves from
approximately 0.4-1.1 GB/s to 1.9-2.2 GB/s.
Before (throughput: approximately 0.4-1.1 GB/s):
Performance counter stats for 'system wide':
4,526 sunrpc:xprt_disconnect_force
1,911 rpcrdma:xprtrdma_disconnect
1,909 rpcrdma:xprtrdma_connect
30.021267692 seconds time elapsed
After (throughput: approximately 1.9-2.2 GB/s):
Performance counter stats for 'system wide':
289 sunrpc:xprt_disconnect_force
289 rpcrdma:xprtrdma_disconnect
291 rpcrdma:xprtrdma_connect
30.080790959 seconds time elapsed
Tim Menninger (2):
SUNRPC: add request-scoped disconnects when cancelling tasks
NFSv4/flexfiles: limit transport disconnects when cancelling I/O
fs/nfs/flexfilelayout/flexfilelayout.c | 7 +-
include/linux/sunrpc/sched.h | 4 ++
include/linux/sunrpc/xprt.h | 5 +-
net/sunrpc/clnt.c | 93 +++++++++++++++++++++-----
net/sunrpc/sunrpc.h | 6 ++
net/sunrpc/xprt.c | 75 +++++++++++++++++++++
6 files changed, 168 insertions(+), 22 deletions(-)
--
2.34.1
next reply other threads:[~2026-09-24 20:45 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-24 20:45 Tim Menninger [this message]
2026-09-24 20:45 ` [PATCH 1/2] SUNRPC: add request-scoped disconnects when cancelling tasks Tim Menninger
2026-09-24 20:45 ` [PATCH 2/2] NFSv4/flexfiles: limit transport disconnects when cancelling I/O Tim Menninger
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260924204525.3381590-1-tmenninger@everpuredata.com \
--to=tmenninger@everpuredata.com \
--cc=anna@kernel.org \
--cc=cel@kernel.org \
--cc=ebadger@everpuredata.com \
--cc=jcurley@everpuredata.com \
--cc=jlayton@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=slingappa@everpuredata.com \
--cc=trondmy@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox