Netdev List
 help / color / mirror / Atom feed
* [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg()
@ 2026-09-30 16:24 Sebastian Andrzej Siewior
  2026-10-01 16:02 ` Sebastian Andrzej Siewior
  2026-10-02 19:53 ` Jakub Kicinski
  0 siblings, 2 replies; 5+ messages in thread
From: Sebastian Andrzej Siewior @ 2026-09-30 16:24 UTC (permalink / raw)
  To: netdev, linux-rt-devel
  Cc: Nam Cao, Kuniyuki Iwashima, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman

From: Nam Cao <namcao@linutronix.de>

AF_UNIX sockets' sendmsg() schedules and blocks on the garbage collector
if the user sent a file descriptor and the user has too many inflight
unix sockets and there is a cyclic reference in the system. The user
needs to have more "unix_inflight" FDs than UNIX_INFLIGHT_SANE_USER
(2024) in order to wait for the garbage collector to complete.

The reason for blocking on garbage collector goes back to 2008, when
it was reported that "Local/unprivileged users can cause soft lockups
and take out system processes by triggering the OOM killer":
https://bugzilla.redhat.com/show_bug.cgi?id=470201

The soft lockup was because a process can keep queuing AF_UNIX sockets to
another process that is exiting. Back in 2008, the garbage collector was
run synchronously by the exiting process, therefore keep queuing AF_UNIX
sockets blocks that process from exiting.

The solution to that issue was forcing sendmsg() to wait for ongoing
garbage collector, in commit 5f23b734963ec ("net: Fix soft lockups/OOM
issues w/ unix garbage collector").

The OOM killer issue was brought up again in 2010:
https://lore.kernel.org/lkml/AANLkTi=Q967xpX0KLMwX-=_4_1AKO5wjHEuJ1TrNjCj9@mail.gmail.com/

To resolve that report, beside blocking on the garbage collector, sendmsg()
also schedules the garbage collector if the number of inflight FDs
in the system is too high. This was done in commit 9915672d41273
("af_unix: limit unix_tot_inflight").

Then in 2015, once again, the OOM killer problem was brought up:
https://lore.kernel.org/lkml/20151228141435.GA13351@1wt.eu/

That time, the issue was resolved by disallowing a user from having more
inflight AF_UNIX sockets than their RLIMIT_NOFILE. That was done by commit
712f4aad406b ("unix: properly account for FDs passed over unix sockets")
and commit 415e3d3e90ce ("unix: correctly track in-flight fds in sending
process user_struct").

Now, sendmsg() does not have to block on the garbage collector anymore,
because:

  - The soft lockup issue is no longer relevant, because the garbage
    collector now runs asynchronously since commit d9f21b361333 ("af_unix:
    Try to run GC async."). Collecting happens with disabled preemption
    but the free happens in preemptible context.

  - The OOM killer issue is addressed by ensuring no more than
    RLIMIT_NOFILE FDs can be inflight.

  - A privileged user can bypass the RLIMIT_NOFILE limit. Should the
    user continue to enqueue FDs then waiting on GC does only slow down
    the user. It does not avoid the OOM if the GC does not find and dead
    FDs which can be cleaned up.

Therefore, don't penalise the user by blocking on the GC. No problems
were observed after running the reproducers from the mentioned bug
reports.

Signed-off-by: Nam Cao <namcao@linutronix.de>
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---

I played with it a bit. The first scheduling of the GC via
unix_schedule_gc() does not wait for its completion because it requires
unix_graph_cyclic_sccs to be set which is only the case after GC run
completed. Once this occurred, the subsequent invocation will block.
If there are no dead sockets then blocking was all for nothing but this
is not known in advance.
The ordinary user should be rate limited by the RLIMIT_NOFILE and the
privileged one can OOM anyway.
Usually the kworker should get enough CPU time to clean up the dead
sockets. I tried to hook it up to the shrinker, so that there is one
synchronous GC invocation in the OOM case, but this hardly helped. It
can not be estimated if it will free something so most of the time it
was pointless.
If the task exists without consuming the passed FDs then they remain and
can only be cleaned up by the GC. The GC is triggered enough by random
tasks on close() so it does not seems necessary to have a shrinker or
another hook to enforce this at random intervals.

v4…v5: https://lore.kernel.org/all/cover.1785824313.git.namcao@linutronix.de/
  - Reworded the commit message a bit
  - Also removed the comment about penalty

v4: 
  - Downsize this series to just removing the priority
    inversion. Garbage collector scheduling should still be cleaned up,
    but that is non-trivial and I do not want that to stand in the way
    of this simple fix.

v3:
  - Move unix_schedule_gc() to be after exit_task_work()

v2:
  - Add patch [1/3]
  - Rebase the other two patches onto the new patch
  - Change commit message to be more precise

 net/unix/garbage.c | 6 ------
 1 file changed, 6 deletions(-)

diff --git a/net/unix/garbage.c b/net/unix/garbage.c
index da774f56ca648..e4c6f645fbd0f 100644
--- a/net/unix/garbage.c
+++ b/net/unix/garbage.c
@@ -648,16 +648,10 @@ void unix_schedule_gc(struct user_struct *user)
 	if (READ_ONCE(unix_graph_state) == UNIX_GRAPH_NOT_CYCLIC)
 		return;
 
-	/* Penalise users who want to send AF_UNIX sockets
-	 * but whose sockets have not been received yet.
-	 */
 	if (user &&
 	    READ_ONCE(user->unix_inflight) < UNIX_INFLIGHT_SANE_USER)
 		return;
 
 	if (!READ_ONCE(gc_in_progress))
 		queue_work(system_dfl_wq, &unix_gc_work);
-
-	if (user && READ_ONCE(unix_graph_cyclic_sccs))
-		flush_work(&unix_gc_work);
 }
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg()
  2026-09-30 16:24 [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg() Sebastian Andrzej Siewior
@ 2026-10-01 16:02 ` Sebastian Andrzej Siewior
  2026-10-02 19:53 ` Jakub Kicinski
  1 sibling, 0 replies; 5+ messages in thread
From: Sebastian Andrzej Siewior @ 2026-10-01 16:02 UTC (permalink / raw)
  To: netdev, linux-rt-devel
  Cc: Nam Cao, Kuniyuki Iwashima, David S . Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman

On 2026-09-30 18:24:59 [+0200], To netdev@vger.kernel.org wrote:
> I played with it a bit. The first scheduling of the GC via
> unix_schedule_gc() does not wait for its completion because it requires

I have a few other data points from my play time:
The "on the flight limit" is the FD limit. This one can be increased to
the hardlimit by an ordinary user, so
| $ ulimit -n 524288

next based on [0] socketpair() + sendmsg() in a loop until the FD limit
is hit, and sleep.
Before start of the program, "free -h" reported for used 667Mi. After it
was done allocating (and not terminated) 3,1Gi.

I added a few trace_printk()s:

--- a/net/unix/garbage.c
+++ b/net/unix/garbage.c
@@ -614,6 +614,7 @@ static void unix_gc(struct work_struct *work)
 	WRITE_ONCE(gc_in_progress, true);
 
 	spin_lock(&unix_gc_lock);
+	trace_printk("Start\n");
 
 	if (unix_graph_state == UNIX_GRAPH_NOT_CYCLIC) {
 		spin_unlock(&unix_gc_lock);
@@ -627,6 +628,7 @@ static void unix_gc(struct work_struct *work)
 	else
 		unix_walk_scc(&hitlist);
 
+	trace_printk("End\n");
 	spin_unlock(&unix_gc_lock);
 
 	skb_queue_walk(&hitlist, skb) {
@@ -634,7 +636,9 @@ static void unix_gc(struct work_struct *work)
 			UNIXCB(skb).fp->dead = true;
 	}
 
+	trace_printk("Purge %d\n", hitlist.qlen);
 	__skb_queue_purge_reason(&hitlist, SKB_DROP_REASON_SOCKET_CLOSE);
+	trace_printk("Purged\n");
 skip_gc:
 	WRITE_ONCE(gc_in_progress, false);
 }

and after the program was done allocating:

|  kworker/u150:0-231     [021] ...1.  1046.960384: unix_gc: Start
|  kworker/u150:0-231     [021] .B.1.  1046.991199: unix_gc: End
|  kworker/u150:0-231     [021] .....  1046.991206: unix_gc: Purge 0
|  kworker/u150:0-231     [021] .....  1046.991206: unix_gc: Purged

~30ms to iterate over the lists, nothing to purge since everything is in
use. This is what I mean, that flush_work() slows things down but does
help. A deferred work would make sense just to throttle that gc.

Now I trigged the OOM killer and saw:
|  Tasks state (memory values in pages):
|  [  pid  ]   uid  tgid total_vm      rss rss_anon rss_file rss_shmem pgtables_bytes swapents oom_score_adj name
|  [   2549]  1001  2549      643      431       24      407         0    45056        0             0 unix-fd-tc
|  [   2365]  1001  2365     2172     1177       93     1084         0    53248        0           200 dbus-daemon
|  oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0-1,global_oom,task_memcg=/user.slice/user-1001.slice/user@1001.service/session.slice/dbus.service,task=dbus-daemon,pid=2365,uid=1001
|  Out of memory: Killed process 2365 (dbus-daemon) total-vm:8688kB, anon-rss:372kB, file-rss:4336kB, shmem-rss:0kB, UID:1001 pgtables:52kB oom_score_adj:200

That "unix-fd-tc" program looks very thin (given that >2GiB are in use).
But that is probably okay. I guess that the skbs are just accounted on
the socket and I don't hit any limits here (maybe I should?).

Now, killing that program, the memory remains occupied. A few seconds
later a random close triggered the GC and then

|  kworker/u143:2-2451    [030] ...1.  1066.959556: unix_gc: Start
|  kworker/u143:2-2451    [030] ...1.  1066.989993: unix_gc: End
|  kworker/u143:2-2451    [030] .....  1067.057696: unix_gc: Purge 524289
|  kworker/u143:2-2451    [030] .l...  1067.705185: unix_gc: Purged

again, 30ms to iterate and then free 524289 items took a bit but it was
preemptible. After that `used' dropped back to 675Mi.

I would argue that too_many_unix_fds() could use UNIX_INFLIGHT_SANE_USER
or 2 * UNIX_INFLIGHT_SANE_USER as a hard limit. Having 1000 fd inflight
for a user sounds insane high amount but there might be legitime
use case…

[0] https://lore.kernel.org/all/ba4ed2717e5225b7b77ef928fb97ee5544632811.1784712370.git.namcao@linutronix.de/

Sebastian

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg()
  2026-09-30 16:24 [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg() Sebastian Andrzej Siewior
  2026-10-01 16:02 ` Sebastian Andrzej Siewior
@ 2026-10-02 19:53 ` Jakub Kicinski
  2026-10-02 21:54   ` Sebastian Andrzej Siewior
  1 sibling, 1 reply; 5+ messages in thread
From: Jakub Kicinski @ 2026-10-02 19:53 UTC (permalink / raw)
  To: Sebastian Andrzej Siewior
  Cc: netdev, linux-rt-devel, Nam Cao, Kuniyuki Iwashima,
	David S . Miller, Eric Dumazet, Paolo Abeni, Simon Horman

On Wed, 30 Sep 2026 18:24:57 +0200 Sebastian Andrzej Siewior wrote:
> From: Nam Cao <namcao@linutronix.de>

This looks effectively identical to v4. I'll put it into 'Needs ACK'.
If Kuniyuki is unconvinced please do not repost this again.

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg()
  2026-10-02 19:53 ` Jakub Kicinski
@ 2026-10-02 21:54   ` Sebastian Andrzej Siewior
  2026-10-03 17:39     ` Kuniyuki Iwashima
  0 siblings, 1 reply; 5+ messages in thread
From: Sebastian Andrzej Siewior @ 2026-10-02 21:54 UTC (permalink / raw)
  To: Jakub Kicinski
  Cc: netdev, linux-rt-devel, Nam Cao, Kuniyuki Iwashima,
	David S . Miller, Eric Dumazet, Paolo Abeni, Simon Horman

On 2026-10-02 12:53:57 [-0700], Jakub Kicinski wrote:
> On Wed, 30 Sep 2026 18:24:57 +0200 Sebastian Andrzej Siewior wrote:
> > From: Nam Cao <namcao@linutronix.de>
> 
> This looks effectively identical to v4. I'll put it into 'Needs ACK'.
> If Kuniyuki is unconvinced please do not repost this again.

basically the same, yes. I did try to explain why that flush does not
improve things. Also I added my notes (as a reply to the patch) where
the bad boy consumes a lot of memory and the GC+flush do not help. This
looks like current limits are too high. I am also not sure if the memory
account for sockets in this case works as expected.

Sebastian

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg()
  2026-10-02 21:54   ` Sebastian Andrzej Siewior
@ 2026-10-03 17:39     ` Kuniyuki Iwashima
  0 siblings, 0 replies; 5+ messages in thread
From: Kuniyuki Iwashima @ 2026-10-03 17:39 UTC (permalink / raw)
  To: Sebastian Andrzej Siewior
  Cc: Jakub Kicinski, netdev, linux-rt-devel, Nam Cao, David S . Miller,
	Eric Dumazet, Paolo Abeni, Simon Horman

On Fri, Oct 2, 2026 at 2:54 PM Sebastian Andrzej Siewior
<bigeasy@linutronix.de> wrote:
>
> On 2026-10-02 12:53:57 [-0700], Jakub Kicinski wrote:
> > On Wed, 30 Sep 2026 18:24:57 +0200 Sebastian Andrzej Siewior wrote:
> > > From: Nam Cao <namcao@linutronix.de>
> >
> > This looks effectively identical to v4. I'll put it into 'Needs ACK'.
> > If Kuniyuki is unconvinced please do not repost this again.
>
> basically the same, yes. I did try to explain why that flush does not
> improve things. Also I added my notes (as a reply to the patch) where
> the bad boy consumes a lot of memory and the GC+flush do not help. This
> looks like current limits are too high. I am also not sure if the memory
> account for sockets in this case works as expected.

It works, but only with memcg; AF_UNIX sets sk_allocation to
GFP_KERNEL_ACCOUNT and other objects are also allocated
with it (or SLAB_ACCOUNT).

The default RLIMIT_NOFILE hard limit is 4096 and fs.file-max is
scaled to only 10% of RAM, but systemd decided to blindly bump
RLIMIT_NOFILE to 512K regardless of host memory size and disable
fs.file-max (see systemd's 52d620757817, a8b627aaed40, and
09dad04c49ca), relying on memcg.

Also, I'm still not convinced.  As mentioned before, flush_work() is
neither for OOM nor for the case without dead cycles (any process can
keep RLIMIT_NOFILE FDs inflight without cyclic references anyway).

---
pw-bot: rejected

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-03 17:39 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 16:24 [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg() Sebastian Andrzej Siewior
2026-10-01 16:02 ` Sebastian Andrzej Siewior
2026-10-02 19:53 ` Jakub Kicinski
2026-10-02 21:54   ` Sebastian Andrzej Siewior
2026-10-03 17:39     ` Kuniyuki Iwashima

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox