From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f53.google.com (mail-wm1-f53.google.com [209.85.128.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DD37B3D0907 for ; Wed, 22 Jul 2026 07:26:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.53 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784705167; cv=none; b=GR0MlAq8LHHMp1vzUxCbTBthhOlGGPD3a6KFUv0+uCt5I8PduGjKg8hOZecngIw9J2JfsrgONjy3mzzOpk0J0YKewz2p9gctFo2EjqSzJYO89ulgsEDdvYxjBn6y5DxYnWPpFmyGfixH6ihM3SSbzbT2E+s4vZOuXXXQKrywo/w= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784705167; c=relaxed/simple; bh=uJIAB5aQK5warnNdya8/fwYBHVjoZDLyZkCYyd2WN6s=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=oTwLi12+yV7IH3in8MLvWMf5oR7wxRrln8IHsD0Uaram4bVXlmCOB77ZH7x9RBD1S+nalnLqksfE+v+oc2EPHTP8HgQ6oRqyJ/7oubSwPhISzfDDLHTEytKRCfueOccRhBx5BtM22QuTSRfpZ2GH1oGDVg6nFag9WZoLytDA+vU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=qWpYI3Fz; arc=none smtp.client-ip=209.85.128.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="qWpYI3Fz" Received: by mail-wm1-f53.google.com with SMTP id 5b1f17b1804b1-4955158f26aso24596275e9.3 for ; Wed, 22 Jul 2026 00:26:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1784705162; x=1785309962; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=yG1zyrx4ZGGVITfwIyi0juPEmhwYGA6Wx/gAK25GPak=; b=qWpYI3FzsSc4fl9eNit1vGBjmgtXAAo2AMeTHowX3HD/5a5GYS9DzsQG+H+zbM651e /aetiIvH7fHdxrCSehYgSOfr2XzAek1fl+kSb/oWXrBKpJAFQpT/ybliX19WgnDUTQKQ /g04gI/sHhbJ+CXK2ikiAFeagDcRhGyPktXL17xPjMrKN0s9zjSwV9km9ulOpDuHSh50 pwnjVFrlSShXKOztKBkpduCZhrqbaEA4+CI8icnflU6Feo4qwRDzdjLtae//3wTd3HyQ 9WGTpSXuVclR3eklCj8nm98E2P71LzXDXYLuzr56EgF/JtZ5duAgLl+qcE7GLz6v+6QN 4w8Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784705162; x=1785309962; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=yG1zyrx4ZGGVITfwIyi0juPEmhwYGA6Wx/gAK25GPak=; b=QOIGZKZS0YYqT+8IfwSrsqE6EW6fnekSFKmFM9TxBj1rOLzCbiHxleswdcP1p/geBM KZ2TBE8Pc+rwzWtYtRb0TGBd03r+YrpHVT/e514xFTqdtkS84ZNGWfjJKcpCgSOJoj1I OcpLNYxiCrJ3VA/Tag5b9Yg8h9QGLtheKyUUOvmi/33r7XNpF6D8QCtF3x9eQILGu4Bg FWK5Zw6wxrxswWZREwdFAEpSF3uAOa0r6+NDnZUYw5Fs+H1BHfhkJht3Cf19oNlHBXhj e/8kZLfuzNI3vi7uFrAJu7Or0zk4wmpC6hopcgq/fk6sUvFpYFZ12eu224INbo2hgH2l EQxg== X-Forwarded-Encrypted: i=1; AHgh+Rps4qP3DWWBVqw814m8n69beQoGdbxz4st/Dwg8s5UdYxCN6P4JxqkUjUe5wXPPhN+bUHRcm5Mx0zCh+b4=@vger.kernel.org X-Gm-Message-State: AOJu0YyH2CpQq6MSLhpLuMyjzBN2rUN1f6ZuRGHwh7j6Tu31DO6djqpt g/mcZG7qNWH+I4TSpKsIaSH9FFE4DGviuphDAWvklF1dKqGcEB6OS7pfXXJuPCHK X-Gm-Gg: AR+sD12nLqqgPbQRor9c0q1h2MrCAD1q52zdETL+ZkhdxYsNMKLL90fdxfnwEh1U5w+ Kw7MJqplc+q6+b62TnyViZImwFJIOaKYa0popRAecEggQES9E8aTvc+HsIkwrtqa00i8RG5kkOD aA6dbF2uYIf1vXMG9D4VN7+5vuXwwzgFXhhOsfdoGsqL/atmdbx8Dr9op7kSKtwGloIoNb55By3 y94rCf5WnS8W+pmlNsX2yEkDY91CRRJB5jvF59Kd1f6Bqsg8FjVEKThsOiq06EeA1OZGtzDAKpA q5O2hkXIz7+nEO8sVKCz+vlxPaaa4nIVFw3mIs9O40AQHSsjO+yu0ASwMjSCnCnugA6qYD6QBBx 48cnuFElrsJTfId0xnh2KnhC1lXOJntPLVS9fEJp5C3NxeZdoWYHy9x25gkOegDA9iIqjGHfG1k unW7+QcMVoajEG8upjFuhrMPAT0RaVADE6HbkCAA== X-Received: by 2002:a05:600c:c84:b0:495:4689:1e98 with SMTP id 5b1f17b1804b1-4954a3ed426mr236754125e9.10.1784705161604; Wed, 22 Jul 2026 00:26:01 -0700 (PDT) Received: from localhost (ip87-106-108-193.pbiaas.com. [87.106.108.193]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4956537c7desm176881885e9.7.2026.07.22.00.26.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jul 2026 00:26:00 -0700 (PDT) Date: Wed, 22 Jul 2026 09:25:59 +0200 From: =?iso-8859-1?Q?G=FCnther?= Noack To: John Ericson Cc: David Laight , Kuniyuki Iwashima , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Cong Wang , Simon Horman , Christian Brauner , David Rheinsberg , Andy Lutomirski , Sergei Zimmerman , network dev , =?iso-8859-1?Q?Micka=EBl_Sala=FCn?= , =?iso-8859-1?Q?G=FCnther?= Noack , Paul Moore , linux-security-module@vger.kernel.org, LKML Subject: Re: unix_stream_connect and socket address resolution Message-ID: <20260722.bfca37efd700@gnoack.org> References: <20260703073948.2541875-1-John.Ericson@Obsidian.Systems> <20260703073948.2541875-3-John.Ericson@Obsidian.Systems> <20260718215855.07284fb1@pumpkin> <85991dc3-6fa5-4466-a0cb-b407291cdb44@app.fastmail.com> <9c437c7c-7919-41e2-9161-fc94803a9b34@app.fastmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <9c437c7c-7919-41e2-9161-fc94803a9b34@app.fastmail.com> Hello John! On Tue, Jul 21, 2026 at 04:37:07PM -0400, John Ericson wrote: > In case this is useful or interesting to anyone, I deep a some history > spelunking, and the restart logic in question seems to date back to > Import 2.2.4pre6: > > https://github.com/tbodt/linux-history/commit/7d4fc34b9bbc0a14d3e5f6b2f978373422e1ca8a#diff-0553d076c243e06ae312480cb8cb52f1cebe1d80fc099d3842593e12c9e0d4f3 > > --- > @@ -673,9 +703,25 @@ static int unix_stream_connect(struct socket *sock, struct sockaddr *uaddr, > we will have to recheck all again in any case. > */ > > +restart: > /* Find listening sock */ > other=unix_find_other(sunaddr, addr_len, sk->type, hash, &err); > > + if (!other) > + return -ECONNREFUSED; > + > + while (other->ack_backlog >= other->max_ack_backlog) { > + unix_unlock(other); > + if (other->dead || other->state != TCP_LISTEN) > + return -ECONNREFUSED; > + if (flags & O_NONBLOCK) > + return -EAGAIN; > + interruptible_sleep_on(&unix_ack_wqueue); > + if (signal_pending(current)) > + return -ERESTARTSYS; > + goto restart; > + } > + > /* create new sock for complete connection */ > newsk = unix_create1(NULL, 1); > > @@ -704,7 +750,7 @@ static int unix_stream_connect(struct socket *sock, struct sockaddr *uaddr, > > /* Check that listener is in valid state. */ > err = -ECONNREFUSED; > - if (other == NULL || other->dead || other->state != TCP_LISTEN) > + if (other->dead || other->state != TCP_LISTEN) > goto out; > > err = -ENOMEM; > --- > > My question can be basically restated: why should the `restart` label > not go *after* the `unix_find_other` call? To elaborate on David Laights answer -- the classic way of restarting a Unix Domain socket server is that the server unlink(2)s the socket file and then bind(2)s the address again. In other words, when the server restarts, the new instance offers the service on a socket file with the same name but it's technically a *fresh* socket file. So a "normal" server restart is not technically that different to the example you gave in your earlier mail (the "File system version"). Another variant of this is one where the server uses rename(2) to switch out the socket file atomically during restart. I was not around when that af_unix code was written, and cannot guarantee that by interpretation is correct, but in the hope that someone will point it out if it's totally bogus: My interpretation is that the "goto restart" loop in af_unix.c is designed to smoothen the server restart cases in a way so that the client doesn't have to deal with manual system call restarts. That loop needs to include the repeated vfs lookup so that it actually gets the socket that belongs to the fresh socket file. (Otherwise it would just observe the same SOCK_DEAD socket from the old server process again. - The new server serves from a new struct socket.) Example Scenario, where a client connect(2) races with the shutdown of an old server during an "atomic" (rename(2)) server restart: * Old server process serves on /foo/bar.sock * New server process starts up * New server process binds (and creates) /foo/bar.sock.tmp * Client runs connect(2) to /foo/bar.sock, does the unix_find_other lookup, getting the old server's socket * New server process renames /foo/bar.sock.tmp to /foo/bar.sock * Old server process shuts down, socket transitions to SOCK_DEAD through unix_release() -> unix_release_sock() -> sock_orphan() * Client connect(2) syscall checks the socket and discovers the SOCK_DEAD state, because it still holds a pointer to the old server socket. To recover from this, connect(2) does the VFS lookup again, because the already looked up dead socket isn't coming back. The connect(2) function pretends that it ran a tiny bit later and only observed the new socket. In this scenario, the userspace server is already going to great lengths to make the switch atomic - it would be surprising IMHO if the connect(2) operation could still return errors to userspace due to race conditions in that case. Again, this is just my own interpretation. I also did not find any better documentation on this. If I am wrong, I am more than happy to be corrected. :) –Günther