All of lore.kernel.org
 help / color / mirror / Atom feed
From: Jakub Kicinski <kuba@kernel.org>
To: Jamal Hadi Salim <jhs@mojatatu.com>
Cc: netdev@vger.kernel.org, Jiri Pirko <jiri@resnulli.us>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Paolo Abeni <pabeni@redhat.com>, Simon Horman <horms@kernel.org>,
	Willem de Bruijn <willemdebruijn.kernel@gmail.com>,
	Jason Wang <jasowangio@gmail.com>,
	Andrew Lunn <andrew+netdev@lunn.ch>,
	stable@vger.kernel.org, vega@nebusec.ai,
	Victor Nogueira <victor@mojatatu.com>
Subject: Re: [PATCH net 1/2] net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations
Date: Mon, 31 Aug 2026 20:08:20 -0700	[thread overview]
Message-ID: <20260831200820.523ac597@kernel.org> (raw)
In-Reply-To: <20260828121902.66837-1-jhs@mojatatu.com>

On Fri, 28 Aug 2026 08:19:01 -0400 Jamal Hadi Salim wrote:
> Several subsystems allocate ring buffers sized by dev->tx_queue_len
> with no upper bound. An unprivileged user (via unshare -Urn) can set a
> huge tx_queue_len and create many devices/queues to exhaust global
> memory, causing a system-wide OOM:
> 
> - pfifo_fast: pfifo_fast_init() and pfifo_fast_change_tx_queue_len()
>   allocate 3 skb_array rings of tx_queue_len entries each.
> - tun: tun_queue_resize() and the queue-attach path resize ptr_rings
>   to tx_queue_len on the NETDEV_CHANGE_TX_QUEUE_LEN notifier.
> - tap (macvtap/ipvtap): tap_queue_resize() and tap_init() resize/init
>   ptr_rings to tx_queue_len on the same notifier.
> 
> netif_change_tx_queue_len() is the single entry point for IFLA_TXQLEN,
> sysfs, and the SIOCSIFTXQLEN ioctl. Cap new_len at S16_MAX (32767)
> there so the oversized value is rejected at set time. This will take 
> effect whether the device is up or down before dev->tx_queue_len is
> written or any notifier fires or any ring is allocated. So a good
> choke spot.
> 
> S16_MAX is the virtio virtqueue size limit: the virtio specification
> stores the queue size as a u16 with a maximum of 32768, so 32767 is
> the largest tx_queue_len any in-tree driver can meaningfully use.
> Values above that only serve to inflate ring allocations.
> 
> Conditions to recreate the bug:
> - CONFIG_NET_SCHED=y, CONFIG_VETH=y, CONFIG_USER_NS=y, CONFIG_NET_NS=y.
> - Unprivileged user in a fresh user+net namespace (unshare -Urn).
> - pfifo_fast: create veth pairs, set tx_queue_len to 500000, attach
>   mq+pfifo_fast. ~28 iterations OOMs a 2GB guest.
> - tun: create 50 tun devices with IFF_MULTI_QUEUE, set tx_queue_len to
>   500000, open 8 queues each. ~1.6GB of ptr_ring allocations OOMs a
>   512MB guest.
> - tap: same as tun with IFF_TAP. ~960MB OOMs a 512MB guest.
> - On the fixed kernel the oversized tx_queue_len is rejected with
>   -ERANGE at set time.
> 
> Fixes: 6a643ddb5624 ("net: introduce helper dev_change_tx_queue_len()")
> Reported-by: vega@nebusec.ai
> Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
> ---
>  net/core/dev.c | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
> 
> diff --git a/net/core/dev.c b/net/core/dev.c
> index 38336858c168..1d3fc0a268a5 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -9982,7 +9982,7 @@ int netif_change_tx_queue_len(struct net_device *dev, unsigned long new_len)
>  	unsigned int orig_len = dev->tx_queue_len;
>  	int res;
>  
> -	if (new_len != (unsigned int)new_len)
> +	if (new_len > S16_MAX)
>  		return -ERANGE;
>  
>  	if (new_len != orig_len) {

clashiko points out that we need to also cover the newlink path:

net/core/rtnetlink.c:rtnl_create_link() {
	...
	if (tb[IFLA_TXQLEN])
		dev->tx_queue_len = nla_get_u32(tb[IFLA_TXQLEN]);
	...

      parent reply	other threads:[~2026-09-01  3:08 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-28 12:19 [PATCH net 1/2] net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations Jamal Hadi Salim
2026-08-28 12:19 ` [PATCH net 2/2] selftests: tc-testing: add tx_queue_len cap regression tests Jamal Hadi Salim
2026-09-01  3:08 ` Jakub Kicinski [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260831200820.523ac597@kernel.org \
    --to=kuba@kernel.org \
    --cc=andrew+netdev@lunn.ch \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=jasowangio@gmail.com \
    --cc=jhs@mojatatu.com \
    --cc=jiri@resnulli.us \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=stable@vger.kernel.org \
    --cc=vega@nebusec.ai \
    --cc=victor@mojatatu.com \
    --cc=willemdebruijn.kernel@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.