Netdev List
 help / color / mirror / Atom feed
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
To: netdev@vger.kernel.org, linux-rt-devel@lists.linux.dev
Cc: Nam Cao <namcao@linutronix.de>,
	Kuniyuki Iwashima <kuniyu@google.com>,
	"David S . Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>
Subject: [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg()
Date: Wed, 30 Sep 2026 18:24:57 +0200	[thread overview]
Message-ID: <20260930162457.4ajVJ-Qa@linutronix.de> (raw)

From: Nam Cao <namcao@linutronix.de>

AF_UNIX sockets' sendmsg() schedules and blocks on the garbage collector
if the user sent a file descriptor and the user has too many inflight
unix sockets and there is a cyclic reference in the system. The user
needs to have more "unix_inflight" FDs than UNIX_INFLIGHT_SANE_USER
(2024) in order to wait for the garbage collector to complete.

The reason for blocking on garbage collector goes back to 2008, when
it was reported that "Local/unprivileged users can cause soft lockups
and take out system processes by triggering the OOM killer":
https://bugzilla.redhat.com/show_bug.cgi?id=470201

The soft lockup was because a process can keep queuing AF_UNIX sockets to
another process that is exiting. Back in 2008, the garbage collector was
run synchronously by the exiting process, therefore keep queuing AF_UNIX
sockets blocks that process from exiting.

The solution to that issue was forcing sendmsg() to wait for ongoing
garbage collector, in commit 5f23b734963ec ("net: Fix soft lockups/OOM
issues w/ unix garbage collector").

The OOM killer issue was brought up again in 2010:
https://lore.kernel.org/lkml/AANLkTi=Q967xpX0KLMwX-=_4_1AKO5wjHEuJ1TrNjCj9@mail.gmail.com/

To resolve that report, beside blocking on the garbage collector, sendmsg()
also schedules the garbage collector if the number of inflight FDs
in the system is too high. This was done in commit 9915672d41273
("af_unix: limit unix_tot_inflight").

Then in 2015, once again, the OOM killer problem was brought up:
https://lore.kernel.org/lkml/20151228141435.GA13351@1wt.eu/

That time, the issue was resolved by disallowing a user from having more
inflight AF_UNIX sockets than their RLIMIT_NOFILE. That was done by commit
712f4aad406b ("unix: properly account for FDs passed over unix sockets")
and commit 415e3d3e90ce ("unix: correctly track in-flight fds in sending
process user_struct").

Now, sendmsg() does not have to block on the garbage collector anymore,
because:

  - The soft lockup issue is no longer relevant, because the garbage
    collector now runs asynchronously since commit d9f21b361333 ("af_unix:
    Try to run GC async."). Collecting happens with disabled preemption
    but the free happens in preemptible context.

  - The OOM killer issue is addressed by ensuring no more than
    RLIMIT_NOFILE FDs can be inflight.

  - A privileged user can bypass the RLIMIT_NOFILE limit. Should the
    user continue to enqueue FDs then waiting on GC does only slow down
    the user. It does not avoid the OOM if the GC does not find and dead
    FDs which can be cleaned up.

Therefore, don't penalise the user by blocking on the GC. No problems
were observed after running the reproducers from the mentioned bug
reports.

Signed-off-by: Nam Cao <namcao@linutronix.de>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---

I played with it a bit. The first scheduling of the GC via
unix_schedule_gc() does not wait for its completion because it requires
unix_graph_cyclic_sccs to be set which is only the case after GC run
completed. Once this occurred, the subsequent invocation will block.
If there are no dead sockets then blocking was all for nothing but this
is not known in advance.
The ordinary user should be rate limited by the RLIMIT_NOFILE and the
privileged one can OOM anyway.
Usually the kworker should get enough CPU time to clean up the dead
sockets. I tried to hook it up to the shrinker, so that there is one
synchronous GC invocation in the OOM case, but this hardly helped. It
can not be estimated if it will free something so most of the time it
was pointless.
If the task exists without consuming the passed FDs then they remain and
can only be cleaned up by the GC. The GC is triggered enough by random
tasks on close() so it does not seems necessary to have a shrinker or
another hook to enforce this at random intervals.

v4…v5: https://lore.kernel.org/all/cover.1785824313.git.namcao@linutronix.de/
  - Reworded the commit message a bit
  - Also removed the comment about penalty

v4: 
  - Downsize this series to just removing the priority
    inversion. Garbage collector scheduling should still be cleaned up,
    but that is non-trivial and I do not want that to stand in the way
    of this simple fix.

v3:
  - Move unix_schedule_gc() to be after exit_task_work()

v2:
  - Add patch [1/3]
  - Rebase the other two patches onto the new patch
  - Change commit message to be more precise

 net/unix/garbage.c | 6 ------
 1 file changed, 6 deletions(-)

diff --git a/net/unix/garbage.c b/net/unix/garbage.c
index da774f56ca648..e4c6f645fbd0f 100644
--- a/net/unix/garbage.c
+++ b/net/unix/garbage.c
@@ -648,16 +648,10 @@ void unix_schedule_gc(struct user_struct *user)
 	if (READ_ONCE(unix_graph_state) == UNIX_GRAPH_NOT_CYCLIC)
 		return;
 
-	/* Penalise users who want to send AF_UNIX sockets
-	 * but whose sockets have not been received yet.
-	 */
 	if (user &&
 	    READ_ONCE(user->unix_inflight) < UNIX_INFLIGHT_SANE_USER)
 		return;
 
 	if (!READ_ONCE(gc_in_progress))
 		queue_work(system_dfl_wq, &unix_gc_work);
-
-	if (user && READ_ONCE(unix_graph_cyclic_sccs))
-		flush_work(&unix_gc_work);
 }
-- 
2.55.0


             reply	other threads:[~2026-09-30 16:25 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 16:24 Sebastian Andrzej Siewior [this message]
2026-10-01 16:02 ` [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg() Sebastian Andrzej Siewior
2026-10-02 19:53 ` Jakub Kicinski
2026-10-02 21:54   ` Sebastian Andrzej Siewior
2026-10-03 17:39     ` Kuniyuki Iwashima

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930162457.4ajVJ-Qa@linutronix.de \
    --to=bigeasy@linutronix.de \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=linux-rt-devel@lists.linux.dev \
    --cc=namcao@linutronix.de \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox