* [PATCH 0/3] smb: smbdirect: fix listener backlog leak and teardown races
@ 2026-10-05 18:35 Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 1/3] smb: smbdirect: don't wait for RDMA_CM_EVENT_DISCONNECTED in smbdirect_socket_destroy_sync() Stefan Metzmacher
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: Stefan Metzmacher @ 2026-10-05 18:35 UTC (permalink / raw)
To: linux-cifs, samba-technical
Cc: metze, Namjae Jeon, Paulo Alcantara, Tom Talpey,
레드팀하고싶어요
Hi Namjae,
These fix a set of problems in the shared smbdirect module
(fs/smb/smbdirect/), reached through the listener/accept path, which in
practice is driven by ksmbd for incoming SMB Direct connections.
The main one is a pre-authentication remote denial of service: the
listener keeps every accepted connection on a fixed-size backlog
(listen.pending, limit 10 for ksmbd). A connection is only removed from
that backlog on the success path (promotion to listen.ready after a
completed negotiate, or dequeue by accept()). A connection that is
accepted at the RDMA level but then fails before it is accepted by the
upper layer - negotiate timeout, peer disconnect, or an invalid
negotiate request - is never removed, so its backlog slot is leaked.
After about ten such events the listener rejects every new connection
with -EBUSY and SMB Direct stays unusable until the service is
restarted. An unauthenticated peer can trigger this with ~10 aborted
connections.
This was reported by 레드팀하고싶어요 <ihopenre@gmail.com>:
https://lore.kernel.org/linux-cifs/CANTrAmxL5sWU8Sn29JAGZvSq5PBB2OA+VKA0MHsZJMesagtfzA@mail.gmail.com/
The series:
- smbdirect_socket_destroy_sync() no longer waits for
RDMA_CM_EVENT_DISCONNECTED after rdma_disconnect(). That wait could
take very long (until the rdma cm gives up, e.g. a peer that just
vanished) or never complete if the event already happened and the
status was overwritten. rdma_destroy_id() in
smbdirect_socket_destroy() is enough to stop further events; the
rest of the disconnect protocol is handled by the rdma core
asynchronously. This also makes releasing orphaned sockets
(next patch) non-blocking.
- A failed, not-yet-accepted child of a listener now moves itself to a
new listen.orphaned list at the end of its cleanup work and queues
listen.purge_orphaned_work, which releases it, so it no longer leaks
its backlog slot (nor its struct sock / RDMA resources) for the
lifetime of the listener. sc->accept.listener is only changed under
listener->listen.lock, and whoever clears it owns the release;
the listener is freed via kfree_rcu() so a child can dereference it
under rcu_read_lock(). This is the DoS fix.
- smbdirect_socket_accept() no longer hands out a socket that failed
after it was put on the ready list (e.g. the peer disconnected in
the meantime); it checks first_error / the status under listen.lock,
orphans such a socket and tries the next one.
I tested the reproducer and the problem is fixed
and I run various xfstests.
I think these are important and should go into 7.3
Stefan Metzmacher (3):
smb: smbdirect: don't wait for RDMA_CM_EVENT_DISCONNECTED in
smbdirect_socket_destroy_sync()
smb: smbdirect: release failed pending sockets of a listener
smb: smbdirect: don't hand out already failed sockets in
smbdirect_socket_accept()
fs/smb/smbdirect/accept.c | 122 ++++++++++++++++++++++++++--------
fs/smb/smbdirect/connection.c | 12 ++--
fs/smb/smbdirect/internal.h | 2 +
fs/smb/smbdirect/listen.c | 116 ++++++++++++++++++++++++++++++--
fs/smb/smbdirect/socket.c | 113 +++++++++++++++++++++++++++----
fs/smb/smbdirect/socket.h | 37 +++++++++++
6 files changed, 352 insertions(+), 50 deletions(-)
--
2.43.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH 1/3] smb: smbdirect: don't wait for RDMA_CM_EVENT_DISCONNECTED in smbdirect_socket_destroy_sync()
2026-10-05 18:35 [PATCH 0/3] smb: smbdirect: fix listener backlog leak and teardown races Stefan Metzmacher
@ 2026-10-05 18:35 ` Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 2/3] smb: smbdirect: release failed pending sockets of a listener Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 3/3] smb: smbdirect: don't hand out already failed sockets in smbdirect_socket_accept() Stefan Metzmacher
2 siblings, 0 replies; 4+ messages in thread
From: Stefan Metzmacher @ 2026-10-05 18:35 UTC (permalink / raw)
To: linux-cifs, samba-technical
Cc: metze, Namjae Jeon, Paulo Alcantara, Tom Talpey, linux-rdma
smbdirect_socket_destroy_sync() waited for RDMA_CM_EVENT_DISCONNECTED
after smbdirect_socket_cleanup_work() called rdma_disconnect(). That
can take very long, e.g. if the peer just disappeared, until the
rdma cm gives up, or forever if the event already happened, but the
status was overwritten afterwards.
This is not needed: smbdirect_socket_destroy() drains the qp under
rdma_lock_handler(), destroys it and calls rdma_destroy_id(), which
only waits for a currently running event handler and makes sure no
further events are delivered. The rest of the disconnect protocol
(DREQ/DREP and timewait for IB, abrupt close for iWarp) is handled
by the rdma core asynchronously. drivers/nvme/host/rdma.c also just
calls rdma_disconnect() and ib_drain_qp() before rdma_destroy_id().
So smbdirect_socket_destroy() now also accepts
SMBDIRECT_SOCKET_DISCONNECTING and changes the status to
SMBDIRECT_SOCKET_DISCONNECTED itself under rdma_lock_handler().
We could still get RDMA_CM_EVENT_DISCONNECTED in the small
windows between rdma_unlock_handler() and rdma_destroy_id(),
but in that case smbdirect_connection_rdma_event_handler()
is basically a no-op.
Fixes: 422a2436697d ("smb: smbdirect: introduce smbdirect_socket_destroy[_sync]()")
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Paulo Alcantara <pc@manguebit.org>
Cc: Tom Talpey <tom@talpey.com>
Cc: linux-cifs@vger.kernel.org
Cc: samba-technical@lists.samba.org
Cc: linux-rdma@vger.kernel.org
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Metzmacher <metze@samba.org>
---
fs/smb/smbdirect/connection.c | 12 ++++----
fs/smb/smbdirect/socket.c | 56 ++++++++++++++++++++++++++++-------
2 files changed, 52 insertions(+), 16 deletions(-)
diff --git a/fs/smb/smbdirect/connection.c b/fs/smb/smbdirect/connection.c
index afd31fa12a36..f48c857bda5a 100644
--- a/fs/smb/smbdirect/connection.c
+++ b/fs/smb/smbdirect/connection.c
@@ -83,9 +83,9 @@ static int smbdirect_connection_rdma_event_handler(struct rdma_cm_id *id,
* smbdirect_socket_schedule_cleanup[_status]() =>
* smbdirect_socket_cleanup_work().
*
- * As otherwise we'd set SMBDIRECT_SOCKET_DISCONNECTING,
- * but never ever get RDMA_CM_EVENT_DISCONNECTED and
- * never reach SMBDIRECT_SOCKET_DISCONNECTED.
+ * As otherwise we'd set SMBDIRECT_SOCKET_DISCONNECTING
+ * and call rdma_disconnect(), but never ever get
+ * RDMA_CM_EVENT_DISCONNECTED.
*/
if (event->event == RDMA_CM_EVENT_DEVICE_REMOVAL)
smbdirect_socket_schedule_cleanup_status(sc,
@@ -113,9 +113,9 @@ static int smbdirect_connection_rdma_event_handler(struct rdma_cm_id *id,
* smbdirect_socket_schedule_cleanup_status() =>
* smbdirect_socket_cleanup_work().
*
- * As otherwise we'd set SMBDIRECT_SOCKET_DISCONNECTING,
- * but never ever get RDMA_CM_EVENT_DISCONNECTED and
- * never reach SMBDIRECT_SOCKET_DISCONNECTED.
+ * As otherwise we'd set SMBDIRECT_SOCKET_DISCONNECTING
+ * and call rdma_disconnect(), but never ever get
+ * RDMA_CM_EVENT_DISCONNECTED.
*
* This is also a normal disconnect so
* SMBDIRECT_LOG_INFO should be good enough
diff --git a/fs/smb/smbdirect/socket.c b/fs/smb/smbdirect/socket.c
index bb02df6158b9..c36cb7cc0088 100644
--- a/fs/smb/smbdirect/socket.c
+++ b/fs/smb/smbdirect/socket.c
@@ -511,7 +511,13 @@ static void smbdirect_socket_destroy(struct smbdirect_socket *sc)
if (sc->status == SMBDIRECT_SOCKET_DESTROYED)
return;
- WARN_ONCE(sc->status != SMBDIRECT_SOCKET_DISCONNECTED,
+ /*
+ * smbdirect_socket_destroy_sync() doesn't wait
+ * for RDMA_CM_EVENT_DISCONNECTED, so we may still
+ * be in SMBDIRECT_SOCKET_DISCONNECTING
+ * (or already reached SMBDIRECT_SOCKET_DISCONNECTED)
+ */
+ WARN_ONCE(sc->status < SMBDIRECT_SOCKET_DISCONNECTING,
"status=%s first_error=%1pe",
smbdirect_socket_status_string(sc->status),
SMBDIRECT_DEBUG_ERR_PTR(sc->first_error));
@@ -542,6 +548,32 @@ static void smbdirect_socket_destroy(struct smbdirect_socket *sc)
if (sc->rdma.cm_id)
rdma_lock_handler(sc->rdma.cm_id);
+ /*
+ * We hold the handler lock, so the rdma event
+ * handlers can't change the status anymore.
+ *
+ * If RDMA_CM_EVENT_DISCONNECTED didn't arrive yet,
+ * we just stop waiting for it here.
+ *
+ * We already disabled disconnect_work above
+ * and before we call rdma_unlock_handler below
+ * we call smbdirect_connection_destroy_qp which
+ * sets sc->ib.qp = NULL.
+ *
+ * Between rdma_unlock_handler() and
+ * rdma_destroy_id() below there's a small
+ * windows where RDMA_CM_EVENT_DISCONNECTED
+ * could still arrive.
+ *
+ * But smbdirect_connection_rdma_event_handler
+ * will be a noop when calling smbdirect_socket_schedule_cleanup*
+ * ib_drain_qp() also won't be called.
+ */
+ if (sc->status < SMBDIRECT_SOCKET_DISCONNECTED) {
+ sc->status = SMBDIRECT_SOCKET_DISCONNECTED;
+ smbdirect_socket_wake_up_all(sc);
+ }
+
if (sc->ib.qp) {
smbdirect_log_rdma_event(sc, SMBDIRECT_LOG_INFO,
"drain qp\n");
@@ -679,17 +711,21 @@ void smbdirect_socket_destroy_sync(struct smbdirect_socket *sc)
"destroying rdma session\n");
if (sc->status < SMBDIRECT_SOCKET_DISCONNECTING)
smbdirect_socket_cleanup_work(&sc->disconnect_work);
- if (sc->status < SMBDIRECT_SOCKET_DISCONNECTED) {
- smbdirect_log_rdma_event(sc, SMBDIRECT_LOG_INFO,
- "wait for transport being disconnected\n");
- wait_event(sc->status_wait, sc->status == SMBDIRECT_SOCKET_DISCONNECTED);
- smbdirect_log_rdma_event(sc, SMBDIRECT_LOG_INFO,
- "waited for transport being disconnected\n");
- }
/*
- * Once we reached SMBDIRECT_SOCKET_DISCONNECTED,
- * we should call smbdirect_socket_destroy()
+ * We don't wait for RDMA_CM_EVENT_DISCONNECTED,
+ * rdma_disconnect() was already called by
+ * smbdirect_socket_cleanup_work() if needed
+ * and smbdirect_socket_destroy() drains the qp,
+ * destroys it and calls rdma_destroy_id(), which
+ * only waits for a currently running event handler
+ * and makes sure no further events are delivered.
+ * The rest of the disconnect protocol is handled
+ * by the rdma core asynchronously.
+ *
+ * Waiting for RDMA_CM_EVENT_DISCONNECTED could
+ * take very long or forever, e.g. if the peer
+ * just disappeared.
*/
smbdirect_socket_destroy(sc);
smbdirect_log_rdma_event(sc, SMBDIRECT_LOG_INFO,
--
2.43.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* [PATCH 2/3] smb: smbdirect: release failed pending sockets of a listener
2026-10-05 18:35 [PATCH 0/3] smb: smbdirect: fix listener backlog leak and teardown races Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 1/3] smb: smbdirect: don't wait for RDMA_CM_EVENT_DISCONNECTED in smbdirect_socket_destroy_sync() Stefan Metzmacher
@ 2026-10-05 18:35 ` Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 3/3] smb: smbdirect: don't hand out already failed sockets in smbdirect_socket_accept() Stefan Metzmacher
2 siblings, 0 replies; 4+ messages in thread
From: Stefan Metzmacher @ 2026-10-05 18:35 UTC (permalink / raw)
To: linux-cifs, samba-technical
Cc: metze,
레드팀하고싶어요,
Namjae Jeon, Paulo Alcantara, Tom Talpey
A child of a listener only left the listen.pending or listen.ready
list via smbdirect_socket_accept() or when the listener was
destroyed. So every connection that failed before it was accepted
(negotiate timeout, peer disconnect, invalid negotiate request, ...)
stayed on the list and counted against the backlog until the listener
was destroyed. After about backlog (10 for ksmbd) such connections
smbdirect_listen_connect_request() rejected every new connection with
-EBUSY, so an unauthenticated peer was able to permanently stop
ksmbd from accepting SMB-Direct connections.
Now a failed child moves itself to a new listen.orphaned list at the
end of smbdirect_socket_cleanup_work(), after the disconnect was
started, and queues listen.purge_orphaned_work, which releases it
with smbdirect_socket_release(), without holding the listener's
rdma_lock_handler() lock, in the same way smbdirect_socket_destroy()
does for the remaining children. Orphaned sockets still count against
the backlog until they are released.
sc->accept.listener is now only changed under listener->listen.lock
and whoever clears it is responsible for releasing the socket:
smbdirect_socket_accept(), purge_orphaned_work or the destroy of the
listener.
In order to make that work:
- a struct smbdirect_socket that is a listener is now freed via
kfree_rcu(), so that a child can dereference sc->accept.listener
under rcu_read_lock(), even if the listener is released concurrently.
- smbdirect_listen_connect_request() only sets nsc->accept.listener
after smbdirect_accept_connect_request() succeeded, so that a
failure within it can't let nsc orphan itself before we release it.
- smbdirect_accept_negotiate_recv_work() no longer moves a socket to
the ready list if it already failed or no longer belongs to the
listener. Every accepting socket has a listener, so if
sc->accept.listener is already cleared, it no longer sends a
negotiate response, as the socket is about to be released.
- smbdirect_socket_destroy() of the listener disables
purge_orphaned_work and also releases the orphaned sockets.
Fixes: dc691b91ad16 ("smb: smbdirect: introduce smbdirect_socket_{listen,accept}()")
Reported-by: 레드팀하고싶어요 <ihopenre@gmail.com>
Closes: https://lore.kernel.org/linux-cifs/CANTrAmxL5sWU8Sn29JAGZvSq5PBB2OA+VKA0MHsZJMesagtfzA@mail.gmail.com/
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Paulo Alcantara <pc@manguebit.org>
Cc: Tom Talpey <tom@talpey.com>
Cc: linux-cifs@vger.kernel.org
Cc: samba-technical@lists.samba.org
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Metzmacher <metze@samba.org>
---
fs/smb/smbdirect/accept.c | 67 ++++++++++++++-------
fs/smb/smbdirect/internal.h | 2 +
fs/smb/smbdirect/listen.c | 116 ++++++++++++++++++++++++++++++++++--
fs/smb/smbdirect/socket.c | 57 +++++++++++++++++-
fs/smb/smbdirect/socket.h | 37 ++++++++++++
5 files changed, 252 insertions(+), 27 deletions(-)
diff --git a/fs/smb/smbdirect/accept.c b/fs/smb/smbdirect/accept.c
index 1b30ca8476c1..1da95cdd1ebd 100644
--- a/fs/smb/smbdirect/accept.c
+++ b/fs/smb/smbdirect/accept.c
@@ -302,6 +302,7 @@ static void smbdirect_accept_negotiate_recv_work(struct work_struct *work)
struct smbdirect_socket *sc =
container_of(work, struct smbdirect_socket, connect.work);
struct smbdirect_socket_parameters *sp = &sc->parameters;
+ struct smbdirect_socket *lsc;
struct smbdirect_recv_io *recv_io;
struct smbdirect_negotiate_req *nreq;
unsigned long flags;
@@ -464,29 +465,55 @@ static void smbdirect_accept_negotiate_recv_work(struct work_struct *work)
*/
sp->max_fragmented_send_size = max_fragmented_size;
- if (sc->accept.listener) {
- struct smbdirect_socket *lsc = sc->accept.listener;
- unsigned long flags;
+ /*
+ * Every accepting socket was created by
+ * smbdirect_listen_connect_request() and has a
+ * listener. If it's already cleared, the listener
+ * (or its purge_orphaned_work) is responsible for
+ * releasing us, so we must not send a negotiate
+ * response.
+ *
+ * The memory of the listener is freed via kfree_rcu(),
+ * see smbdirect_listen_orphan_socket().
+ */
+ rcu_read_lock();
+ lsc = READ_ONCE(sc->accept.listener);
+ if (!lsc) {
+ rcu_read_unlock();
+ return;
+ }
- spin_lock_irqsave(&lsc->listen.lock, flags);
- list_del(&sc->accept.list);
- list_add_tail(&sc->accept.list, &lsc->listen.ready);
+ spin_lock_irqsave(&lsc->listen.lock, flags);
+ /*
+ * The listener (or its purge_orphaned_work)
+ * may have cleared sc->accept.listener in the
+ * meantime and is responsible for releasing us.
+ *
+ * If we already failed, we're either
+ * already on the orphaned list or
+ * smbdirect_socket_cleanup_work() will
+ * move us there.
+ *
+ * In both cases we must not move us
+ * to the ready list.
+ */
+ if (sc->accept.listener == lsc && !READ_ONCE(sc->first_error)) {
+ list_move_tail(&sc->accept.list, &lsc->listen.ready);
wake_up(&lsc->listen.wait_queue);
- spin_unlock_irqrestore(&lsc->listen.lock, flags);
-
- /*
- * smbdirect_socket_accept() will call
- * smbdirect_accept_negotiate_finish(nsc, 0);
- *
- * So that we don't send the negotiation
- * response that grants credits to the peer
- * before the socket is accepted by the
- * application.
- */
- return;
}
+ spin_unlock_irqrestore(&lsc->listen.lock, flags);
+ rcu_read_unlock();
- ntstatus = le32_to_cpu(STATUS_SUCCESS);
+ /*
+ * smbdirect_socket_accept() will call
+ * smbdirect_accept_negotiate_finish(nsc, 0);
+ *
+ * So that we don't send the negotiation
+ * response that grants credits to the peer
+ * before the socket is accepted by the
+ * application.
+ */
+ return;
not_supported:
smbdirect_accept_negotiate_finish(sc, ntstatus);
@@ -839,7 +866,7 @@ struct smbdirect_socket *smbdirect_socket_accept(struct smbdirect_socket *lsc,
struct smbdirect_socket,
accept.list);
if (nsc) {
- nsc->accept.listener = NULL;
+ WRITE_ONCE(nsc->accept.listener, NULL);
list_del_init_careful(&nsc->accept.list);
arg->is_empty = list_empty_careful(&lsc->listen.ready);
}
diff --git a/fs/smb/smbdirect/internal.h b/fs/smb/smbdirect/internal.h
index e9959e6dc13a..69d0c9f899de 100644
--- a/fs/smb/smbdirect/internal.h
+++ b/fs/smb/smbdirect/internal.h
@@ -74,6 +74,8 @@ void __smbdirect_socket_schedule_cleanup(struct smbdirect_socket *sc,
void smbdirect_socket_destroy_sync(struct smbdirect_socket *sc);
+void smbdirect_listen_orphan_socket(struct smbdirect_socket *sc);
+
int smbdirect_socket_wait_for_credits(struct smbdirect_socket *sc,
enum smbdirect_socket_status expected_status,
int unexpected_errno,
diff --git a/fs/smb/smbdirect/listen.c b/fs/smb/smbdirect/listen.c
index 2f78bcaedbf8..b30c0aa37e3a 100644
--- a/fs/smb/smbdirect/listen.c
+++ b/fs/smb/smbdirect/listen.c
@@ -10,6 +10,78 @@
static int smbdirect_listen_rdma_event_handler(struct rdma_cm_id *id,
struct rdma_cm_event *event);
+/*
+ * This is called by a socket that failed while it
+ * is still on the pending or ready list of its
+ * listener, typically at the end of
+ * smbdirect_socket_cleanup_work(), when the
+ * disconnect was already started.
+ *
+ * It moves itself to the orphaned list and
+ * lets smbdirect_listen_purge_orphaned_work()
+ * release it. Otherwise it would stay on the
+ * pending or ready list until the listener is
+ * destroyed and fill up the backlog, so that
+ * no new connections would be accepted.
+ *
+ * It's fine to call this more than once.
+ */
+void smbdirect_listen_orphan_socket(struct smbdirect_socket *sc)
+{
+ struct smbdirect_socket *lsc;
+ unsigned long flags;
+
+ /*
+ * The memory of the listener is freed via
+ * kfree_rcu(), so it's safe to dereference it
+ * under rcu_read_lock(), even if sc->accept.listener
+ * is cleared and the listener is released concurrently.
+ */
+ rcu_read_lock();
+ lsc = READ_ONCE(sc->accept.listener);
+ if (!lsc) {
+ rcu_read_unlock();
+ return;
+ }
+
+ spin_lock_irqsave(&lsc->listen.lock, flags);
+ if (sc->accept.listener == lsc) {
+ list_move_tail(&sc->accept.list, &lsc->listen.orphaned);
+ queue_work(lsc->workqueues.cleanup, &lsc->listen.purge_orphaned_work);
+ }
+ spin_unlock_irqrestore(&lsc->listen.lock, flags);
+ rcu_read_unlock();
+}
+
+static void smbdirect_listen_purge_orphaned_work(struct work_struct *work)
+{
+ struct smbdirect_socket *lsc =
+ container_of(work, struct smbdirect_socket, listen.purge_orphaned_work);
+ struct smbdirect_socket *psc, *tsc;
+ LIST_HEAD(orphaned_list);
+ unsigned long flags;
+
+ /*
+ * Clearing accept.listener under listen.lock
+ * makes us responsible for releasing them.
+ */
+ spin_lock_irqsave(&lsc->listen.lock, flags);
+ list_splice_tail_init(&lsc->listen.orphaned, &orphaned_list);
+ list_for_each_entry(psc, &orphaned_list, accept.list)
+ WRITE_ONCE(psc->accept.listener, NULL);
+ spin_unlock_irqrestore(&lsc->listen.lock, flags);
+
+ /*
+ * We don't hold the listener's rdma_lock_handler()
+ * lock here, see smbdirect_socket_destroy()
+ * for why that's important.
+ */
+ list_for_each_entry_safe(psc, tsc, &orphaned_list, accept.list) {
+ list_del_init(&psc->accept.list);
+ smbdirect_socket_release(psc);
+ }
+}
+
int smbdirect_socket_listen(struct smbdirect_socket *sc, int backlog)
{
int ret;
@@ -47,6 +119,9 @@ int smbdirect_socket_listen(struct smbdirect_socket *sc, int backlog)
sc->rdma.cm_id->event_handler = smbdirect_listen_rdma_event_handler;
rdma_unlock_handler(sc->rdma.cm_id);
+ INIT_WORK(&sc->listen.purge_orphaned_work,
+ smbdirect_listen_purge_orphaned_work);
+
ret = rdma_listen(sc->rdma.cm_id, backlog);
if (ret) {
sc->first_error = ret;
@@ -214,6 +289,7 @@ static int smbdirect_listen_connect_request(struct smbdirect_socket *lsc,
size_t backlog = max_t(size_t, 1, lsc->listen.backlog);
size_t psockets;
size_t rsockets;
+ size_t osockets;
int ret;
if (!smbdirect_frwr_is_supported(&new_id->device->attrs)) {
@@ -254,14 +330,21 @@ static int smbdirect_listen_connect_request(struct smbdirect_socket *lsc,
spin_lock_irqsave(&lsc->listen.lock, flags);
psockets = list_count_nodes(&lsc->listen.pending);
rsockets = list_count_nodes(&lsc->listen.ready);
+ /*
+ * Orphaned sockets still count against
+ * the backlog until they are released by
+ * smbdirect_listen_purge_orphaned_work().
+ */
+ osockets = list_count_nodes(&lsc->listen.orphaned);
spin_unlock_irqrestore(&lsc->listen.lock, flags);
if (psockets > backlog ||
rsockets > backlog ||
- (psockets + rsockets) > backlog) {
+ osockets > backlog ||
+ (psockets + rsockets + osockets) > backlog) {
smbdirect_log_rdma_event(lsc, SMBDIRECT_LOG_ERR,
- "Backlog[%d][%zu] full pending[%zu] ready[%zu]\n",
- lsc->listen.backlog, backlog, psockets, rsockets);
+ "Backlog[%d][%zu] full pending[%zu] ready[%zu] orphaned[%zu]\n",
+ lsc->listen.backlog, backlog, psockets, rsockets, osockets);
return -EBUSY;
}
@@ -279,21 +362,44 @@ static int smbdirect_listen_connect_request(struct smbdirect_socket *lsc,
if (ret)
goto set_settings_failed;
+ /*
+ * Note that nsc->accept.listener is only set
+ * after smbdirect_accept_connect_request()
+ * succeeded, so that a failure within it
+ * can't let nsc orphan itself via
+ * smbdirect_listen_orphan_socket(), as we
+ * release it below in that case.
+ *
+ * The listener's handler_mutex is held while
+ * we're called, so smbdirect_socket_destroy()
+ * of the listener can't see nsc on the
+ * pending list before we're done.
+ */
spin_lock_irqsave(&lsc->listen.lock, flags);
list_add_tail(&nsc->accept.list, &lsc->listen.pending);
- nsc->accept.listener = lsc;
spin_unlock_irqrestore(&lsc->listen.lock, flags);
ret = smbdirect_accept_connect_request(nsc, &event->param.conn);
if (ret)
goto accept_connect_failed;
+ spin_lock_irqsave(&lsc->listen.lock, flags);
+ WRITE_ONCE(nsc->accept.listener, lsc);
+ spin_unlock_irqrestore(&lsc->listen.lock, flags);
+
+ /*
+ * If nsc already failed, its
+ * smbdirect_socket_cleanup_work() may
+ * have missed the listener.
+ */
+ if (READ_ONCE(nsc->first_error))
+ smbdirect_listen_orphan_socket(nsc);
+
return 0;
accept_connect_failed:
spin_lock_irqsave(&lsc->listen.lock, flags);
list_del_init(&nsc->accept.list);
- nsc->accept.listener = NULL;
spin_unlock_irqrestore(&lsc->listen.lock, flags);
set_settings_failed:
set_params_failed:
diff --git a/fs/smb/smbdirect/socket.c b/fs/smb/smbdirect/socket.c
index c36cb7cc0088..e270e36b3145 100644
--- a/fs/smb/smbdirect/socket.c
+++ b/fs/smb/smbdirect/socket.c
@@ -319,6 +319,15 @@ void __smbdirect_socket_schedule_cleanup(struct smbdirect_socket *sc,
* (smbdirect_socket_destroy) to reap.
*/
if (sc->listen.backlog != -1) { /* was a listener */
+ /*
+ * We don't move them to the orphaned list here,
+ * each child does that itself at the end of its
+ * smbdirect_socket_cleanup_work(), see
+ * smbdirect_listen_orphan_socket(). Only that checks
+ * accept.listener, which is still NULL for a child
+ * that smbdirect_listen_connect_request() is still
+ * setting up and will release itself on failure.
+ */
spin_lock_irqsave(&sc->listen.lock, flags);
list_splice_init(&sc->listen.ready, &sc->listen.pending);
list_for_each_entry_safe(psc, tsc, &sc->listen.pending, accept.list)
@@ -427,6 +436,15 @@ static void smbdirect_socket_cleanup_work(struct work_struct *work)
* instances of one class -- harmless, but lockdep cannot tell).
*/
if (sc->listen.backlog != -1) { /* was a listener */
+ /*
+ * We don't move them to the orphaned list here,
+ * each child does that itself at the end of its
+ * smbdirect_socket_cleanup_work(), see
+ * smbdirect_listen_orphan_socket(). Only that checks
+ * accept.listener, which is still NULL for a child
+ * that smbdirect_listen_connect_request() is still
+ * setting up and will release itself on failure.
+ */
spin_lock_irqsave(&sc->listen.lock, flags);
list_splice_init(&sc->listen.ready, &sc->listen.pending);
list_for_each_entry_safe(psc, tsc, &sc->listen.pending, accept.list)
@@ -486,6 +504,18 @@ static void smbdirect_socket_cleanup_work(struct work_struct *work)
* in order to notice the broken connection.
*/
smbdirect_socket_wake_up_all(sc);
+
+ /*
+ * If we're still on the pending or ready list
+ * of a listener, we started the disconnect
+ * as far as possible above, so we move ourself
+ * to the orphaned list of the listener,
+ * which will release us.
+ *
+ * This is a no-op internally if
+ * sc->accept.listener is NULL.
+ */
+ smbdirect_listen_orphan_socket(sc);
}
static void smbdirect_socket_destroy(struct smbdirect_socket *sc)
@@ -544,6 +574,7 @@ static void smbdirect_socket_destroy(struct smbdirect_socket *sc)
disable_work_sync(&sc->recv_io.posted.refill_work);
disable_work_sync(&sc->idle.immediate_work);
disable_delayed_work_sync(&sc->idle.timer_work);
+ disable_work_sync(&sc->listen.purge_orphaned_work);
if (sc->rdma.cm_id)
rdma_lock_handler(sc->rdma.cm_id);
@@ -607,6 +638,19 @@ static void smbdirect_socket_destroy(struct smbdirect_socket *sc)
spin_lock_irqsave(&sc->listen.lock, flags);
list_splice_tail_init(&sc->listen.ready, &pending_list);
list_splice_tail_init(&sc->listen.pending, &pending_list);
+ /*
+ * purge_orphaned_work is already disabled above,
+ * so we also need to release the orphaned sockets.
+ *
+ * Clearing accept.listener under listen.lock
+ * makes us responsible for releasing them and
+ * prevents them from moving themselves to
+ * the orphaned list via
+ * smbdirect_listen_orphan_socket().
+ */
+ list_splice_tail_init(&sc->listen.orphaned, &pending_list);
+ list_for_each_entry(psc, &pending_list, accept.list)
+ WRITE_ONCE(psc->accept.listener, NULL);
spin_unlock_irqrestore(&sc->listen.lock, flags);
/* It's not possible for upper layer to get to reassembly */
@@ -650,7 +694,6 @@ static void smbdirect_socket_destroy(struct smbdirect_socket *sc)
"release %zu pending sockets\n", psockets);
list_for_each_entry_safe(psc, tsc, &pending_list, accept.list) {
list_del_init(&psc->accept.list);
- psc->accept.listener = NULL;
smbdirect_socket_release(psc);
}
if (sc->listen.backlog != -1) /* was a listener */
@@ -777,7 +820,17 @@ static void smbdirect_socket_release_destroy(struct kref *kref)
* in DESTROYED state, before we free the memory.
*/
smbdirect_socket_destroy_sync(sc);
- kfree(sc);
+
+ /*
+ * Only a listener (backlog != -1) is ever dereferenced
+ * via sc->accept.listener under rcu_read_lock(), see
+ * smbdirect_listen_orphan_socket(). Other sockets can be
+ * freed immediately.
+ */
+ if (sc->listen.backlog != -1) /* was a listener */
+ kfree_rcu(sc, refs.rcu);
+ else
+ kfree(sc);
}
void smbdirect_socket_release(struct smbdirect_socket *sc)
diff --git a/fs/smb/smbdirect/socket.h b/fs/smb/smbdirect/socket.h
index c09eddd8ad16..36623980ec17 100644
--- a/fs/smb/smbdirect/socket.h
+++ b/fs/smb/smbdirect/socket.h
@@ -151,6 +151,12 @@ struct smbdirect_socket {
* the disconnect refcount.
*/
struct kref destroy;
+
+ /*
+ * smbdirect_socket_release_destroy() uses
+ * kfree_rcu(), see accept.listener.
+ */
+ struct rcu_head rcu;
} refs;
/* RDMA related */
@@ -216,6 +222,16 @@ struct smbdirect_socket {
* only be > 0.
*/
int backlog;
+ /*
+ * Sockets on pending or ready that failed
+ * move themselves to orphaned and queue
+ * purge_orphaned_work, which releases them.
+ * So that they are freed while the listener
+ * is still alive. They still count against
+ * the backlog until they are released.
+ */
+ struct list_head orphaned;
+ struct work_struct purge_orphaned_work;
} listen;
/*
@@ -226,7 +242,25 @@ struct smbdirect_socket {
* connection.
*/
struct {
+ /*
+ * This is only set, protected by
+ * listener->listen.lock, while the socket
+ * is owned by the listener (on its pending,
+ * ready or orphaned list). The one who clears
+ * it is responsible for releasing the socket.
+ *
+ * The memory of a struct smbdirect_socket is
+ * freed via kfree_rcu(), so the listener
+ * can be dereferenced under rcu_read_lock(),
+ * even if accept.listener is cleared and
+ * the listener is released concurrently.
+ */
struct smbdirect_socket *listener;
+ /*
+ * Note even with listener being NULL,
+ * the socket can be on listeners.listen.pending
+ * before smbdirect_accept_connect_request() finished.
+ */
struct list_head list;
} accept;
@@ -588,6 +622,9 @@ static __always_inline void smbdirect_socket_init(struct smbdirect_socket *sc)
spin_lock_init(&sc->listen.lock);
INIT_LIST_HEAD(&sc->listen.pending);
INIT_LIST_HEAD(&sc->listen.ready);
+ INIT_LIST_HEAD(&sc->listen.orphaned);
+ INIT_WORK(&sc->listen.purge_orphaned_work, __smbdirect_socket_disabled_work);
+ disable_work_sync(&sc->listen.purge_orphaned_work);
sc->listen.backlog = -1; /* not a listener */
init_waitqueue_head(&sc->listen.wait_queue);
--
2.43.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* [PATCH 3/3] smb: smbdirect: don't hand out already failed sockets in smbdirect_socket_accept()
2026-10-05 18:35 [PATCH 0/3] smb: smbdirect: fix listener backlog leak and teardown races Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 1/3] smb: smbdirect: don't wait for RDMA_CM_EVENT_DISCONNECTED in smbdirect_socket_destroy_sync() Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 2/3] smb: smbdirect: release failed pending sockets of a listener Stefan Metzmacher
@ 2026-10-05 18:35 ` Stefan Metzmacher
2 siblings, 0 replies; 4+ messages in thread
From: Stefan Metzmacher @ 2026-10-05 18:35 UTC (permalink / raw)
To: linux-cifs, samba-technical
Cc: metze, Namjae Jeon, Paulo Alcantara, Tom Talpey
A socket on the ready list is in SMBDIRECT_SOCKET_NEGOTIATE_RUNNING,
but it may fail before smbdirect_socket_accept() picks it up, e.g.
RDMA_CM_EVENT_DISCONNECTED already moved it to
SMBDIRECT_SOCKET_DISCONNECTED.
smbdirect_socket_accept() unconditionally overwrote the status with
SMBDIRECT_SOCKET_CONNECTED and handed out the already disconnected
socket. smbdirect_socket_cleanup_work() then called rdma_disconnect()
on it again, and before the recent commit
smbdirect_socket_destroy_sync() waited forever for an
RDMA_CM_EVENT_DISCONNECTED that already happened.
Now we check first_error and change the status to
SMBDIRECT_SOCKET_CONNECTED only via cmpxchg() from
SMBDIRECT_SOCKET_NEGOTIATE_RUNNING, while we still hold listen.lock.
If the socket already failed, we move it to the orphaned list and
queue purge_orphaned_work, like smbdirect_listen_orphan_socket() does,
and try the next one. If we only found failed sockets, we wait for
the next one.
Fixes: dc691b91ad16 ("smb: smbdirect: introduce smbdirect_socket_{listen,accept}()")
Cc: Namjae Jeon <linkinjeon@kernel.org>
Cc: Paulo Alcantara <pc@manguebit.org>
Cc: Tom Talpey <tom@talpey.com>
Cc: linux-cifs@vger.kernel.org
Cc: samba-technical@lists.samba.org
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Metzmacher <metze@samba.org>
---
fs/smb/smbdirect/accept.c | 55 ++++++++++++++++++++++++++++++++++-----
1 file changed, 48 insertions(+), 7 deletions(-)
diff --git a/fs/smb/smbdirect/accept.c b/fs/smb/smbdirect/accept.c
index 1da95cdd1ebd..ee16be95424d 100644
--- a/fs/smb/smbdirect/accept.c
+++ b/fs/smb/smbdirect/accept.c
@@ -834,8 +834,12 @@ struct smbdirect_socket *smbdirect_socket_accept(struct smbdirect_socket *lsc,
struct proto_accept_arg *arg)
{
struct smbdirect_socket *nsc;
+ bool orphaned;
unsigned long flags;
+again:
+ orphaned = false;
+
if (lsc->status != SMBDIRECT_SOCKET_LISTENING) {
arg->err = -EINVAL;
return NULL;
@@ -862,9 +866,38 @@ struct smbdirect_socket *smbdirect_socket_accept(struct smbdirect_socket *lsc,
}
spin_lock_irqsave(&lsc->listen.lock, flags);
- nsc = list_first_entry_or_null(&lsc->listen.ready,
- struct smbdirect_socket,
- accept.list);
+ while ((nsc = list_first_entry_or_null(&lsc->listen.ready,
+ struct smbdirect_socket,
+ accept.list))) {
+ /*
+ * nsc may have failed after it was moved
+ * to the ready list, e.g. RDMA_CM_EVENT_DISCONNECTED
+ * already moved it to SMBDIRECT_SOCKET_DISCONNECTED.
+ * We must not overwrite that with
+ * SMBDIRECT_SOCKET_CONNECTED and hand
+ * out an already disconnected socket.
+ *
+ * Doing this under listen.lock means nsc
+ * still belongs to us, so we can just move
+ * a failed socket to the orphaned list,
+ * like smbdirect_listen_orphan_socket() does,
+ * and try the next one.
+ */
+ if (!READ_ONCE(nsc->first_error) &&
+ cmpxchg(&nsc->status,
+ SMBDIRECT_SOCKET_NEGOTIATE_RUNNING,
+ SMBDIRECT_SOCKET_CONNECTED) ==
+ SMBDIRECT_SOCKET_NEGOTIATE_RUNNING)
+ break;
+
+ smbdirect_log_rdma_event(nsc, SMBDIRECT_LOG_INFO,
+ "orphaning failed socket status=%s first_error=%1pe\n",
+ smbdirect_socket_status_string(nsc->status),
+ SMBDIRECT_DEBUG_ERR_PTR(nsc->first_error));
+ list_move_tail(&nsc->accept.list, &lsc->listen.orphaned);
+ queue_work(lsc->workqueues.cleanup, &lsc->listen.purge_orphaned_work);
+ orphaned = true;
+ }
if (nsc) {
WRITE_ONCE(nsc->accept.listener, NULL);
list_del_init_careful(&nsc->accept.list);
@@ -872,6 +905,12 @@ struct smbdirect_socket *smbdirect_socket_accept(struct smbdirect_socket *lsc,
}
spin_unlock_irqrestore(&lsc->listen.lock, flags);
if (!nsc) {
+ /*
+ * If we only found failed sockets,
+ * we wait for the next one.
+ */
+ if (orphaned)
+ goto again;
arg->err = -EAGAIN;
return NULL;
}
@@ -882,12 +921,14 @@ struct smbdirect_socket *smbdirect_socket_accept(struct smbdirect_socket *lsc,
* so it didn't grant any credits to us.
*
* The caller expects a connected socket
- * now as there are no credits anyway.
+ * now as there are no credits anyway,
+ * above we already changed to SMBDIRECT_SOCKET_CONNECTED
+ * under the lsc->listen.lock and with cmpxchg.
*
- * Then we send the negotiation response in
- * order to grant credits to the peer.
+ * Now we send the negotiation response in
+ * order to grant credits to the peer,
+ * as the socket is now visible to the application layer.
*/
- nsc->status = SMBDIRECT_SOCKET_CONNECTED;
smbdirect_accept_negotiate_finish(nsc, 0);
return nsc;
--
2.43.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-10-05 18:37 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 18:35 [PATCH 0/3] smb: smbdirect: fix listener backlog leak and teardown races Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 1/3] smb: smbdirect: don't wait for RDMA_CM_EVENT_DISCONNECTED in smbdirect_socket_destroy_sync() Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 2/3] smb: smbdirect: release failed pending sockets of a listener Stefan Metzmacher
2026-10-05 18:35 ` [PATCH 3/3] smb: smbdirect: don't hand out already failed sockets in smbdirect_socket_accept() Stefan Metzmacher
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox