Linux CIFS filesystem development
 help / color / mirror / Atom feed
From: Stefan Metzmacher <metze@samba.org>
To: linux-cifs@vger.kernel.org, samba-technical@lists.samba.org
Cc: metze@samba.org, Namjae Jeon <linkinjeon@kernel.org>,
	Paulo Alcantara <pc@manguebit.org>, Tom Talpey <tom@talpey.com>,
	Lee Seong Hyeon <ihopenre@gmail.com>
Subject: [PATCH v3 0/3] smb: smbdirect: fix listener backlog leak and teardown races
Date: Tue,  6 Oct 2026 21:08:54 +0200	[thread overview]
Message-ID: <cover.1791311817.git.metze@samba.org> (raw)

Hi Namjae,

These fix a set of problems in the shared smbdirect module
(fs/smb/smbdirect/), reached through the listener/accept path, which in
practice is driven by ksmbd for incoming SMB Direct connections.

The main one is a pre-authentication remote denial of service: the
listener keeps every accepted connection on a fixed-size backlog
(listen.pending, limit 10 for ksmbd). A connection is only removed from
that backlog on the success path (promotion to listen.ready after a
completed negotiate, or dequeue by accept()). A connection that is
accepted at the RDMA level but then fails before it is accepted by the
upper layer - negotiate timeout, peer disconnect, or an invalid
negotiate request - is never removed, so its backlog slot is leaked.
After about ten such events the listener rejects every new connection
with -EBUSY and SMB Direct stays unusable until the service is
restarted. An unauthenticated peer can trigger this with ~10 aborted
connections.

This was reported by Lee Seong Hyeon <ihopenre@gmail.com>:

  https://lore.kernel.org/linux-cifs/CANTrAmxL5sWU8Sn29JAGZvSq5PBB2OA+VKA0MHsZJMesagtfzA@mail.gmail.com/

The series:

  - smbdirect_socket_destroy_sync() no longer waits for
    RDMA_CM_EVENT_DISCONNECTED after rdma_disconnect(). That wait could
    take very long (until the rdma cm gives up, e.g. a peer that just
    vanished) or never complete if the event already happened and the
    status was overwritten. rdma_destroy_id() in
    smbdirect_socket_destroy() is enough to stop further events; the
    rest of the disconnect protocol is handled by the rdma core
    asynchronously. This also makes releasing orphaned sockets
    (next patch) non-blocking.

  - A failed, not-yet-accepted child of a listener now moves itself to a
    new listen.orphaned list at the end of its cleanup work and queues
    listen.purge_orphaned_work, which releases it, so it no longer leaks
    its backlog slot (nor its struct sock / RDMA resources) for the
    lifetime of the listener. sc->accept.listener is only changed under
    listener->listen.lock, and whoever clears it owns the release;
    the listener is freed via kfree_rcu() so a child can dereference it
    under rcu_read_lock(). This is the DoS fix.

  - smbdirect_socket_accept() no longer hands out a socket that failed
    after it was put on the ready list (e.g. the peer disconnected in
    the meantime); it checks first_error / the status under listen.lock,
    orphans such a socket and tries the next one.

I tested the reproducer and the problem is fixed
and I run various xfstests.

I think these are important and should go into 7.3
Changes since v2:
(https://lore.kernel.org/r/cover.1791229120.git.metze@samba.org/)

- "smb: smbdirect: release failed pending sockets of a listener":
  reworked the connect-request error handling after more Sashiko
  review. smbdirect_accept_connect_request() no longer tears down any
  RDMA state or clears the cm_id on failure; it only returns a
  not-yet-posted recv_io to the pool and returns the error.
  smbdirect_listen_connect_request() now sets nsc->accept.listener and
  publishes nsc on the pending list *before* it calls
  smbdirect_accept_connect_request(), and on a non-zero return it
  schedules nsc's teardown with smbdirect_socket_schedule_cleanup() and
  always returns 0, so the rdma_cm core keeps the connection id and nsc
  owns it (and destroys it in its own deferred teardown). The teardown
  must be deferred because nsc's id_priv->handler_mutex is held by the
  rdma_cm core across the CONNECT_REQUEST handler; that same lock (and
  the listener's) also guarantees nsc can't go away under us. This
  closes the race Sashiko reported, where a connection that established
  between rdma_accept() and setting the listener could be left stranded
  on the pending list.

Changes since v1:
(https://lore.kernel.org/r/cover.1791224972.git.metze@samba.org)

- "smb: smbdirect: release failed pending sockets of a listener":
  fix a use-after-free and a missed-orphan race in
  smbdirect_listen_connect_request(), spotted by Sashiko. The
  nsc->first_error check and the orphaning are now done under
  listen.lock, together with setting nsc->accept.listener, so a
  concurrent smbdirect_socket_cleanup_work()/purge_orphaned_work()
  can neither free nsc while we still look at it, nor miss the
  orphaning and leak it on the pending list.

- "smb: smbdirect: don't hand out already failed sockets in
  smbdirect_socket_accept()": bound the accept() wait across the
  "skip a failed socket and try the next one" retries, also spotted
  by Sashiko. smbdirect_socket_wait_for_accept() now returns the
  remaining timeout and smbdirect_socket_accept() carries it over, so
  a stream of failed connections can no longer reset the caller's
  timeout. ksmbd passes MAX_SCHEDULE_TIMEOUT, so it is unaffected in
  practice.

- Added Reported-by:/Closes: for the Sashiko reviews.

Stefan Metzmacher (3):
  smb: smbdirect: don't wait for RDMA_CM_EVENT_DISCONNECTED in
    smbdirect_socket_destroy_sync()
  smb: smbdirect: release failed pending sockets of a listener
  smb: smbdirect: don't hand out already failed sockets in
    smbdirect_socket_accept()

 fs/smb/smbdirect/accept.c     | 194 +++++++++++++++++++++++-----------
 fs/smb/smbdirect/connection.c |  12 +--
 fs/smb/smbdirect/internal.h   |   2 +
 fs/smb/smbdirect/listen.c     | 137 ++++++++++++++++++++++--
 fs/smb/smbdirect/socket.c     | 113 +++++++++++++++++---
 fs/smb/smbdirect/socket.h     |  42 ++++++++
 6 files changed, 411 insertions(+), 89 deletions(-)

-- 
2.43.0


             reply	other threads:[~2026-10-06 19:09 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-06 19:08 Stefan Metzmacher [this message]
2026-10-06 19:08 ` [PATCH v3 1/3] smb: smbdirect: don't wait for RDMA_CM_EVENT_DISCONNECTED in smbdirect_socket_destroy_sync() Stefan Metzmacher
2026-10-06 19:08 ` [PATCH v3 2/3] smb: smbdirect: release failed pending sockets of a listener Stefan Metzmacher
2026-10-08  2:26   ` Namjae Jeon
     [not found]     ` <CANTrAmz-eDf4yacfb+LKEug8Asz3cEEGg2h_6w-pmT_6L4GTzg@mail.gmail.com>
2026-10-08  6:10       ` Stefan Metzmacher
2026-10-08  6:19         ` Namjae Jeon
2026-10-06 19:08 ` [PATCH v3 3/3] smb: smbdirect: don't hand out already failed sockets in smbdirect_socket_accept() Stefan Metzmacher

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=cover.1791311817.git.metze@samba.org \
    --to=metze@samba.org \
    --cc=ihopenre@gmail.com \
    --cc=linkinjeon@kernel.org \
    --cc=linux-cifs@vger.kernel.org \
    --cc=pc@manguebit.org \
    --cc=samba-technical@lists.samba.org \
    --cc=tom@talpey.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox