Netdev List
 help / color / mirror / Atom feed
From: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
To: netdev@vger.kernel.org, linux-rt-devel@lists.linux.dev
Cc: Nam Cao <namcao@linutronix.de>,
	Kuniyuki Iwashima <kuniyu@google.com>,
	"David S . Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>
Subject: Re: [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg()
Date: Thu, 1 Oct 2026 18:02:39 +0200	[thread overview]
Message-ID: <20261001160239.j1stzfx6@linutronix.de> (raw)
In-Reply-To: <20260930162457.4ajVJ-Qa@linutronix.de>

On 2026-09-30 18:24:59 [+0200], To netdev@vger.kernel.org wrote:
> I played with it a bit. The first scheduling of the GC via
> unix_schedule_gc() does not wait for its completion because it requires

I have a few other data points from my play time:
The "on the flight limit" is the FD limit. This one can be increased to
the hardlimit by an ordinary user, so
| $ ulimit -n 524288

next based on [0] socketpair() + sendmsg() in a loop until the FD limit
is hit, and sleep.
Before start of the program, "free -h" reported for used 667Mi. After it
was done allocating (and not terminated) 3,1Gi.

I added a few trace_printk()s:

--- a/net/unix/garbage.c
+++ b/net/unix/garbage.c
@@ -614,6 +614,7 @@ static void unix_gc(struct work_struct *work)
 	WRITE_ONCE(gc_in_progress, true);
 
 	spin_lock(&unix_gc_lock);
+	trace_printk("Start\n");
 
 	if (unix_graph_state == UNIX_GRAPH_NOT_CYCLIC) {
 		spin_unlock(&unix_gc_lock);
@@ -627,6 +628,7 @@ static void unix_gc(struct work_struct *work)
 	else
 		unix_walk_scc(&hitlist);
 
+	trace_printk("End\n");
 	spin_unlock(&unix_gc_lock);
 
 	skb_queue_walk(&hitlist, skb) {
@@ -634,7 +636,9 @@ static void unix_gc(struct work_struct *work)
 			UNIXCB(skb).fp->dead = true;
 	}
 
+	trace_printk("Purge %d\n", hitlist.qlen);
 	__skb_queue_purge_reason(&hitlist, SKB_DROP_REASON_SOCKET_CLOSE);
+	trace_printk("Purged\n");
 skip_gc:
 	WRITE_ONCE(gc_in_progress, false);
 }

and after the program was done allocating:

|  kworker/u150:0-231     [021] ...1.  1046.960384: unix_gc: Start
|  kworker/u150:0-231     [021] .B.1.  1046.991199: unix_gc: End
|  kworker/u150:0-231     [021] .....  1046.991206: unix_gc: Purge 0
|  kworker/u150:0-231     [021] .....  1046.991206: unix_gc: Purged

~30ms to iterate over the lists, nothing to purge since everything is in
use. This is what I mean, that flush_work() slows things down but does
help. A deferred work would make sense just to throttle that gc.

Now I trigged the OOM killer and saw:
|  Tasks state (memory values in pages):
|  [  pid  ]   uid  tgid total_vm      rss rss_anon rss_file rss_shmem pgtables_bytes swapents oom_score_adj name
|  [   2549]  1001  2549      643      431       24      407         0    45056        0             0 unix-fd-tc
|  [   2365]  1001  2365     2172     1177       93     1084         0    53248        0           200 dbus-daemon
|  oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0-1,global_oom,task_memcg=/user.slice/user-1001.slice/user@1001.service/session.slice/dbus.service,task=dbus-daemon,pid=2365,uid=1001
|  Out of memory: Killed process 2365 (dbus-daemon) total-vm:8688kB, anon-rss:372kB, file-rss:4336kB, shmem-rss:0kB, UID:1001 pgtables:52kB oom_score_adj:200

That "unix-fd-tc" program looks very thin (given that >2GiB are in use).
But that is probably okay. I guess that the skbs are just accounted on
the socket and I don't hit any limits here (maybe I should?).

Now, killing that program, the memory remains occupied. A few seconds
later a random close triggered the GC and then

|  kworker/u143:2-2451    [030] ...1.  1066.959556: unix_gc: Start
|  kworker/u143:2-2451    [030] ...1.  1066.989993: unix_gc: End
|  kworker/u143:2-2451    [030] .....  1067.057696: unix_gc: Purge 524289
|  kworker/u143:2-2451    [030] .l...  1067.705185: unix_gc: Purged

again, 30ms to iterate and then free 524289 items took a bit but it was
preemptible. After that `used' dropped back to 675Mi.

I would argue that too_many_unix_fds() could use UNIX_INFLIGHT_SANE_USER
or 2 * UNIX_INFLIGHT_SANE_USER as a hard limit. Having 1000 fd inflight
for a user sounds insane high amount but there might be legitime
use case…

[0] https://lore.kernel.org/all/ba4ed2717e5225b7b77ef928fb97ee5544632811.1784712370.git.namcao@linutronix.de/

Sebastian

  reply	other threads:[~2026-10-01 16:02 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30 16:24 [PATCH net-next v5] af_unix: Do not wait for garbage collector in sendmsg() Sebastian Andrzej Siewior
2026-10-01 16:02 ` Sebastian Andrzej Siewior [this message]
2026-10-02 19:53 ` Jakub Kicinski
2026-10-02 21:54   ` Sebastian Andrzej Siewior
2026-10-03 17:39     ` Kuniyuki Iwashima

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261001160239.j1stzfx6@linutronix.de \
    --to=bigeasy@linutronix.de \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=linux-rt-devel@lists.linux.dev \
    --cc=namcao@linutronix.de \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox