Linux CIFS filesystem development
 help / color / mirror / Atom feed
From: Enzo Matsumiya <ematsumiya@suse.de>
To: linux-cifs@vger.kernel.org
Cc: pc@manguebit.org, linkinjeon@kernel.org,
	ronniesahlberg@gmail.com, sprasad@microsoft.com, tom@talpey.com,
	bharathsm@microsoft.com, henrique.carvalho@suse.com,
	Enzo Matsumiya <ematsumiya@suse.de>
Subject: [PATCH v3 2/2] smb: client: prevent premature discard of requests if reconnecting
Date: Tue, 29 Sep 2026 14:03:58 -0300	[thread overview]
Message-ID: <20260929170358.270612-2-ematsumiya@suse.de> (raw)
In-Reply-To: <20260929170358.270612-1-ematsumiya@suse.de>

Add TCP_Server_Info::reconnecting to track cifs_reconnect() lifetime.
Wait for it to be true in cifs_wait_for_server_reconnect(), and only
then wait for tcpStatus change.

Also, return -ECONNRESET (instead of -EHOSTDOWN) if reconnect is still
ongoing at the end, so requests are not discarded unnecessarily;
-ECONNRESET is not immediately retried by intermediate layers, e.g. VFS
or netfs, but still serves as an indication to userspace that retrying
the operation might be successful.

Changes (refactor cifs_wait_for_server_reconnect()):
- remove unnecessary do/while loop; use wait_event_interruptible()
  for hard mounts
- increase timeout, and base it on echo_interval to match user
  preferences
- remove useless debug log for interrupt errors

Signed-off-by: Enzo Matsumiya <ematsumiya@suse.de>
---
v2 -> v3:
- handle CifsExiting status on hard mounts (from sashiko report)

v1 -> v2 (fix issues detected by sashiko):
- handle condition changes post-wait_event timeouts
- return -ECONNRESET instead of -EAGAIN when leaving still reconnecting

(not from sashiko):
- add missing wake_up() after setting server->reconnecting to true in
  cifs_tcp_ses_needs_reconnect()
- handle tcpStatus == CifsExiting in cifs_wait_for_server_reconnect()
 fs/smb/client/cifsglob.h |  1 +
 fs/smb/client/connect.c  |  4 ++
 fs/smb/client/misc.c     | 94 ++++++++++++++++++++++++++++------------
 3 files changed, 72 insertions(+), 27 deletions(-)

diff --git a/fs/smb/client/cifsglob.h b/fs/smb/client/cifsglob.h
index e49140d5c4a4..61191a167b89 100644
--- a/fs/smb/client/cifsglob.h
+++ b/fs/smb/client/cifsglob.h
@@ -803,6 +803,7 @@ struct TCP_Server_Info {
 	struct delayed_work reconnect; /* reconnect workqueue job */
 	struct mutex reconnect_mutex; /* prevent simultaneous reconnects */
 	bool need_sock_shutdown; /* true when CifsNeedReconnect was first set by non-cifsd task */
+	bool reconnecting; /* if cifs_reconnect() is indeed running */
 	unsigned long echo_interval;
 
 	/*
diff --git a/fs/smb/client/connect.c b/fs/smb/client/connect.c
index 733d50b042fd..8ea60965ae06 100644
--- a/fs/smb/client/connect.c
+++ b/fs/smb/client/connect.c
@@ -374,6 +374,8 @@ static bool cifs_tcp_ses_needs_reconnect(struct TCP_Server_Info *server, int num
 	shutdown = (server->tcpStatus != CifsNeedReconnect || server->need_sock_shutdown);
 	server->need_sock_shutdown = false;
 	server->tcpStatus = CifsNeedReconnect;
+	server->reconnecting = true;
+	wake_up(&server->response_q);
 	spin_unlock(&server->srv_lock);
 
 	if (shutdown) {
@@ -457,6 +459,7 @@ static int __cifs_reconnect(struct TCP_Server_Info *server,
 	spin_lock(&server->srv_lock);
 	if (server->tcpStatus == CifsNeedNegotiate)
 		mod_delayed_work(cifsiod_wq, &server->echo, 0);
+	server->reconnecting = false;
 	spin_unlock(&server->srv_lock);
 
 	wake_up(&server->response_q);
@@ -599,6 +602,7 @@ static int reconnect_dfs_server(struct TCP_Server_Info *server)
 	spin_lock(&server->srv_lock);
 	if (server->tcpStatus == CifsNeedNegotiate)
 		mod_delayed_work(cifsiod_wq, &server->echo, 0);
+	server->reconnecting = false;
 	spin_unlock(&server->srv_lock);
 
 	wake_up(&server->response_q);
diff --git a/fs/smb/client/misc.c b/fs/smb/client/misc.c
index 05168284f205..3fa4e36aac32 100644
--- a/fs/smb/client/misc.c
+++ b/fs/smb/client/misc.c
@@ -1077,44 +1077,84 @@ int cifs_inval_name_dfs_link_error(const unsigned int xid,
 
 int cifs_wait_for_server_reconnect(struct TCP_Server_Info *server, bool retry)
 {
-	int timeout = 10;
-	int rc;
+	int timeout, rc = 0;
 
 	spin_lock(&server->srv_lock);
 	if (server->tcpStatus != CifsNeedReconnect) {
+		if (unlikely(server->tcpStatus == CifsExiting))
+			rc = -ESHUTDOWN;
 		spin_unlock(&server->srv_lock);
-		return 0;
+
+		return rc;
+	}
+
+	/* Hard mounts keep waiting until process is killed or server comes back on-line. */
+	if (retry) {
+		spin_unlock(&server->srv_lock);
+
+		return wait_event_interruptible(server->response_q,
+						(server->tcpStatus != CifsNeedReconnect));
 	}
-	timeout *= server->nr_targets;
+
+	/*
+	 * Soft mounts, wait with timeout.
+	 * Compute timeout based on echo_interval, to match user expectations wrt. time to recover
+	 * from network interruptions.
+	 *
+	 * Minimum is 17 seconds (per reconnect target):
+	 *   7s for cifs socket timeout
+	 *   3s sleep after a failed attempt
+	 *   7s spare, to cover at least a full attempt
+	 *
+	 * Maximum is 60s; might not be enough in some rare cases, but >60s is too much.
+	 */
+	timeout = max(server->echo_interval / HZ, 17) * server->nr_targets;
+	timeout = min(timeout, 60);
 	spin_unlock(&server->srv_lock);
 
 	/*
-	 * Give demultiplex thread up to 10 seconds to each target available for
-	 * reconnect -- should be greater than cifs socket timeout which is 7
-	 * seconds.
+	 * First, wait for cifs_reconnect() to start.
+	 * This is needed because tcpStatus == CifsNeedReconnect doesn't imply cifs_reconnect()
+	 * is actually running.
 	 *
-	 * On "soft" mounts we wait once. Hard mounts keep retrying until
-	 * process is killed or server comes back on-line.
+	 * Wait 3*timeout (max 180s) here, which should be more than enough.
 	 */
-	do {
-		rc = wait_event_interruptible_timeout(server->response_q,
-						      (server->tcpStatus != CifsNeedReconnect),
-						      timeout * HZ);
-		if (rc < 0) {
-			cifs_dbg(FYI, "%s: aborting reconnect due to received signal\n",
-				 __func__);
-			return -ERESTARTSYS;
-		}
+	rc = wait_event_interruptible_timeout(server->response_q,
+					      (server->reconnecting ||
+					       server->tcpStatus != CifsNeedReconnect),
+					      min(3 * timeout, 180) * HZ);
 
-		/* are we still trying to reconnect? */
-		spin_lock(&server->srv_lock);
-		if (server->tcpStatus != CifsNeedReconnect) {
-			spin_unlock(&server->srv_lock);
-			return 0;
+	/* cifs_reconnect() never started, don't wait any longer */
+	spin_lock(&server->srv_lock);
+	if (!rc && !server->reconnecting && server->tcpStatus == CifsNeedReconnect)
+		rc = -EHOSTDOWN;
+	spin_unlock(&server->srv_lock);
+
+	if (rc < 0)
+		return rc;
+
+	/* cifs_reconnect() started (maybe already succeeded), wait for status change */
+	rc = wait_event_interruptible_timeout(server->response_q,
+					      (server->tcpStatus != CifsNeedReconnect),
+					      timeout * HZ);
+	if (rc < 0)
+		return rc;
+
+	spin_lock(&server->srv_lock);
+	if (server->tcpStatus != CifsNeedReconnect) {
+		rc = 0;
+		if (server->tcpStatus == CifsExiting)
+			rc = -ESHUTDOWN;
+	} else {
+		rc = -ECONNRESET;
+
+		/* Are we still even trying to reconnect? */
+		if (!server->reconnecting) {
+			cifs_dbg(FYI, "%s: gave up waiting on reconnect\n", __func__);
+			rc = -EHOSTDOWN;
 		}
-		spin_unlock(&server->srv_lock);
-	} while (retry);
+	}
+	spin_unlock(&server->srv_lock);
 
-	cifs_dbg(FYI, "%s: gave up waiting on reconnect\n", __func__);
-	return -EHOSTDOWN;
+	return rc;
 }
-- 
2.55.0


  reply	other threads:[~2026-09-29 17:04 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29 17:03 [PATCH v3 1/2] smb: client: fast fail sends if need to reconnect Enzo Matsumiya
2026-09-29 17:03 ` Enzo Matsumiya [this message]
2026-10-06 19:48   ` [PATCH v3 2/2] smb: client: prevent premature discard of requests if reconnecting Paulo Alcantara
2026-10-06 21:00     ` Enzo Matsumiya
2026-10-06 22:51       ` Paulo Alcantara
2026-10-07 15:04         ` Enzo Matsumiya
2026-10-07 23:25           ` Paulo Alcantara
2026-10-07 17:02         ` Enzo Matsumiya
2026-10-07 23:35           ` Paulo Alcantara
2026-10-06 18:14 ` [PATCH v3 1/2] smb: client: fast fail sends if need to reconnect Paulo Alcantara
2026-10-06 18:52   ` Enzo Matsumiya
2026-10-08 18:30     ` Enzo Matsumiya

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260929170358.270612-2-ematsumiya@suse.de \
    --to=ematsumiya@suse.de \
    --cc=bharathsm@microsoft.com \
    --cc=henrique.carvalho@suse.com \
    --cc=linkinjeon@kernel.org \
    --cc=linux-cifs@vger.kernel.org \
    --cc=pc@manguebit.org \
    --cc=ronniesahlberg@gmail.com \
    --cc=sprasad@microsoft.com \
    --cc=tom@talpey.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox