From: Alexander Aring <aahringo@redhat.com>
To: edumazet@kernel.org
Cc: aahringo@redhat.com, gfs2@lists.linux.dev,
netdev@vger.kernel.org, davem@davemloft.net, kuba@kernel.org,
pabeni@redhat.com
Subject: [RFC 2/2] dlm: fix socket hang on send buffer limitation with socket error
Date: Fri, 2 Oct 2026 19:21:07 -0400 [thread overview]
Message-ID: <20261002232107.1578646-3-aahringo@redhat.com> (raw)
In-Reply-To: <20261002232107.1578646-1-aahringo@redhat.com>
As Sashiko AI bot mentioned [0], a socket connection can get stuck when
the send buffer limitation is reached and a socket error simultaneously
occurs.
This issue can be reproduced with the following setup:
1. Set sk_sndbuf to SOCK_MIN_SNDBUF.
2. Slow down DLM connections using netem.
3. Use tcpkill to randomly send TCP resets to DLM connections.
I instrumented debug printouts to confirm that the sk_write_space()
notifier callbacks were being executed, using step 3 to trigger random
socket errors.
With these changes, I can no longer reproduce the hang.
Changes included:
- Removed sk_write_pending counting, as this should not be modified
at the socket application layer (or is at least unnecessary).
- Moved clearing the CF_SEND_PENDING bit—which allows re-queuing
swork (send worker for sendmsg())—to the sk_write_space() callback,
since this callback notifies us that the underlying socket is no longer
constrained by its send buffer.
- Handled the race condition between sendmsg() and evaluating
SOCK_NOSPACE after sendmsg(), where sk_write_space() could be called
in between, using CF_APP_LIMITED:
- In sk_write_space(), queue swork again if CF_APP_LIMITED is set.
- If CF_APP_LIMITED is not set, do nothing as send_to_sock() will handle
it, confirming the race occurred.
- Introduced new handling in lowcomms_error_report() when a socket error
occurs during send buffer limitation. If CF_APP_LIMITED is set, swork
will be re-queued, which will fail and trigger a reconnect.
- Added various comments explaining the interaction with CF_APP_LIMITED.
[0] https://lore.kernel.org/netdev/179090395863.434549.3668493667259120759@kernel.org/
Signed-off-by: Alexander Aring <aahringo@redhat.com>
---
fs/dlm/lowcomms.c | 46 ++++++++++++++++++++++++++++++++++++----------
1 file changed, 36 insertions(+), 10 deletions(-)
diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c
index c3a414d2e32b..29ee1b52d0df 100644
--- a/fs/dlm/lowcomms.c
+++ b/fs/dlm/lowcomms.c
@@ -521,12 +521,20 @@ static void lowcomms_write_space(struct sock *sk)
sk_clear_nospace(sk);
- spin_lock_bh(&con->writequeue_lock);
- if (test_and_clear_bit(CF_APP_LIMITED, &con->flags))
- con->sock->sk->sk_write_pending--;
-
- lowcomms_queue_swork(con);
- spin_unlock_bh(&con->writequeue_lock);
+ if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) {
+ /* signal to send again by clearing
+ * CF_SEND_PENDING and queue swork.
+ */
+ spin_lock_bh(&con->writequeue_lock);
+ clear_bit(CF_SEND_PENDING, &con->flags);
+ lowcomms_queue_swork(con);
+ spin_unlock_bh(&con->writequeue_lock);
+ } else {
+ /* CF_APP_LIMITED is cleared, so send_to_sock() will
+ * simply reschedule work without hitting the
+ * CF_APP_LIMITED path.
+ */
+ }
}
static void lowcomms_state_change(struct sock *sk)
@@ -623,6 +631,23 @@ static void lowcomms_error_report(struct sock *sk)
break;
}
+ /* if waiting on sk_write_space() and an sk_err occurs, the callback
+ * won't fire. Clear CF_SEND_PENDING and if CF_APP_LIMITED was set
+ * so resend tasks can re-queue swork, triggering a sendmsg() failure
+ * to initiate reconnection.
+ */
+ if (test_and_clear_bit(CF_APP_LIMITED, &con->flags)) {
+ spin_lock_bh(&con->writequeue_lock);
+ clear_bit(CF_SEND_PENDING, &con->flags);
+ /* dlm_midcomms_unack_msg_resend() does not always
+ * trigger lowcomms_queue_swork() as it tries to
+ * avoid to put pending messages into the lowcomms
+ * sending buffer. Force it here again.
+ */
+ lowcomms_queue_swork(con);
+ spin_unlock_bh(&con->writequeue_lock);
+ }
+
dlm_midcomms_unack_msg_resend(con->nodeid);
listen_sock.sk_error_report(sk);
@@ -1391,18 +1416,19 @@ static int send_to_sock(struct connection *con)
spin_lock_bh(&con->writequeue_lock);
if (test_bit(SOCK_NOSPACE, &con->sock->flags) &&
!test_and_set_bit(CF_APP_LIMITED, &con->flags)) {
- con->sock->sk->sk_write_pending++;
-
- clear_bit(CF_SEND_PENDING, &con->flags);
spin_unlock_bh(&con->writequeue_lock);
release_sock(con->sock->sk);
- /* wait for write_space() event */
+ /* wait for sk_write_space() event */
return DLM_IO_END;
}
spin_unlock_bh(&con->writequeue_lock);
release_sock(con->sock->sk);
+ /* the sk_write_space() came in between sock_sendmsg()
+ * and check on SOCK_NOSPACE and the socket became
+ * writeable again so just resched swork.
+ */
return DLM_IO_RESCHED;
} else if (ret < 0) {
return ret;
--
2.43.0
next prev parent reply other threads:[~2026-10-02 23:21 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-02 23:21 [RFC 0/2] dlm: fix socket hang on send buffer limitation with socket error Alexander Aring
2026-10-02 23:21 ` [RFC 1/2] dlm: drop setting SOCK_NOSPACE Alexander Aring
2026-10-02 23:21 ` Alexander Aring [this message]
2026-10-05 14:53 ` [RFC 2/2] dlm: fix socket hang on send buffer limitation with socket error Alexander Aring
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261002232107.1578646-3-aahringo@redhat.com \
--to=aahringo@redhat.com \
--cc=davem@davemloft.net \
--cc=edumazet@kernel.org \
--cc=gfs2@lists.linux.dev \
--cc=kuba@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox