Netdev List
 help / color / mirror / Atom feed
From: "Cen Zhang (Microsoft)" <cenzhang@linux.microsoft.com>
To: "David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@kernel.org>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>
Cc: Simon Horman <horms@kernel.org>, Kees Cook <kees@kernel.org>,
	Jeff Layton <jlayton@kernel.org>,
	Nicolas Dichtel <nicolas.dichtel@6wind.com>,
	Ilya Maximets <i.maximets@ovn.org>,
	Kexin Sun <kexinsun@smail.nju.edu.cn>,
	Enrico Pozzobon <enrico.pozzobon@dissecto.com>,
	Breno Leitao <leitao@debian.org>,
	David Rheinsberg <david@readahead.eu>,
	netdev@vger.kernel.org, linux-kernel@vger.kernel.org,
	stable@vger.kernel.org, AutonomousCodeSecurity@microsoft.com,
	tgopinath@linux.microsoft.com,
	Cen Zhang <cenzhang@linux.microsoft.com>
Subject: [PATCH net] netlink: do not copy to user space under netlink_lock_table()
Date: Thu,  8 Oct 2026 15:48:24 -0400	[thread overview]
Message-ID: <20261008194824.52351-1-cenzhang@linux.microsoft.com> (raw)

netlink_getsockopt() handles NETLINK_LIST_MEMBERSHIPS by copying the
socket's group bitmap to user space with copy_to_iter() under
netlink_lock_table(). However, copy_to_iter() writes to a user address
and can trigger a page fault, and how long that fault takes is up to the
owner of the page. An unprivileged user can make it take as long as they
want, e.g. by mapping a memfd page and keeping it hole punched from
another thread. While the fault is pending, the caller still counts as a
holder of nl_table_users.

nl_table_users is one global counter shared by every netlink protocol
and every network namespace, and netlink_table_grab() waits for it to
drop to zero in TASK_UNINTERRUPTIBLE with no timeout and no signal
check. Therefore, by holding one page fault open, an unprivileged user
blocks every task on the machine that needs netlink_table_grab(): e.g.,
close of a netlink socket, multicast group join, and network namespace
creation. The blocked tasks stay in an unkillable D state until the
attacker lets the fault finish, and hung task detector reports is as
following:

  INFO: task V1-nl-close:103 blocked for more than 122 seconds.
  task:V1-nl-close     state:D stack:14448 pid:103 ...
  Call Trace:
   schedule+0x36/0xf0
   netlink_table_grab.part.0+0x66/0xe0
   netlink_release+0x6d4/0x780
   __x64_sys_close+0x38/0x80

Fix this by taking a kmemdup() snapshot of the bitmap under
netlink_lock_table() and copying the snapshot to user space after the
lock is released.

Fixes: 47191d65b647 ("netlink: fix locking around NETLINK_LIST_MEMBERSHIPS")
Reported-by: AutonomousCodeSecurity@microsoft.com
Signed-off-by: Cen Zhang (Microsoft) <cenzhang@linux.microsoft.com>
---
 net/netlink/af_netlink.c | 20 ++++++++++++++++----
 1 file changed, 16 insertions(+), 4 deletions(-)

diff --git a/net/netlink/af_netlink.c b/net/netlink/af_netlink.c
index 9fdf964224ab..f61ff699761f 100644
--- a/net/netlink/af_netlink.c
+++ b/net/netlink/af_netlink.c
@@ -1767,24 +1767,36 @@ static int netlink_getsockopt(struct socket *sock, int level, int optname,
 		flag = NETLINK_F_RECV_NO_ENOBUFS;
 		break;
 	case NETLINK_LIST_MEMBERSHIPS: {
+		unsigned long *groups = NULL;
 		int pos, idx, shift, err = 0;
+		unsigned int ngroups;
 
+		/* copy_to_iter() may sleep in a page fault, do it unlocked. */
 		netlink_lock_table();
-		for (pos = 0; pos * 8 < nlk->ngroups; pos += sizeof(u32)) {
+		ngroups = nlk->ngroups;
+		if (ngroups)
+			groups = kmemdup(nlk->groups, NLGRPSZ(ngroups),
+					 GFP_KERNEL);
+		netlink_unlock_table();
+
+		if (ngroups && !groups)
+			return -ENOMEM;
+
+		for (pos = 0; pos * 8 < ngroups; pos += sizeof(u32)) {
 			if (len - pos < sizeof(u32))
 				break;
 
 			idx = pos / sizeof(unsigned long);
 			shift = (pos % sizeof(unsigned long)) * 8;
-			group = (u32)(nlk->groups[idx] >> shift);
+			group = (u32)(groups[idx] >> shift);
 			if (copy_to_iter(&group, sizeof(u32),
 					 &opt->iter_out) != sizeof(u32)) {
 				err = -EFAULT;
 				break;
 			}
 		}
-		opt->optlen = ALIGN(BITS_TO_BYTES(nlk->ngroups), sizeof(u32));
-		netlink_unlock_table();
+		opt->optlen = ALIGN(BITS_TO_BYTES(ngroups), sizeof(u32));
+		kfree(groups);
 		return err;
 	}
 	case NETLINK_LISTEN_ALL_NSID:
-- 
2.55.0

             reply	other threads:[~2026-10-08 19:48 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08 19:48 Cen Zhang (Microsoft) [this message]
2026-10-08 20:42 ` [PATCH net] netlink: do not copy to user space under netlink_lock_table() Eric Dumazet
2026-10-08 21:31   ` Cen Zhang (Microsoft)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261008194824.52351-1-cenzhang@linux.microsoft.com \
    --to=cenzhang@linux.microsoft.com \
    --cc=AutonomousCodeSecurity@microsoft.com \
    --cc=davem@davemloft.net \
    --cc=david@readahead.eu \
    --cc=edumazet@kernel.org \
    --cc=enrico.pozzobon@dissecto.com \
    --cc=horms@kernel.org \
    --cc=i.maximets@ovn.org \
    --cc=jlayton@kernel.org \
    --cc=kees@kernel.org \
    --cc=kexinsun@smail.nju.edu.cn \
    --cc=kuba@kernel.org \
    --cc=leitao@debian.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=nicolas.dichtel@6wind.com \
    --cc=pabeni@redhat.com \
    --cc=stable@vger.kernel.org \
    --cc=tgopinath@linux.microsoft.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox