All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH net v2 0/3] net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations
@ 2026-09-02 21:29   ` Jamal Hadi Salim
  0 siblings, 0 replies; 8+ messages in thread
From: Jamal Hadi Salim @ 2026-09-02 21:29 UTC (permalink / raw)
  To: netdev
  Cc: Jamal Hadi Salim, stable, Jiri Pirko, David S . Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Donald Hunter, Vega, Victor Nogueira

An unprivileged user (via unshare -Urn) can set a huge tx_queue_len
and exhaust global memory through ring allocations sized from it
(pfifo_fast skb_arrays, tun/tap ptr_rings).
The reproducer from vega@nebusec.ai set the following params for
illustration: txqlen of 500000 -> ~32 GiB/ring attempts, 1.6 GB tun,
~960 MB tap. Gets worse when you consider qdiscs like mq.

What we fix: every path an unprivileged user can use to install
an oversized tx_queue_len is rejected with -ERANGE before any ring is
allocated; per-ring memory is bounded at 256 KiB.

This is for you sashikos: What we deliberately _do not fix_
bound the NUMBER of rings. With the cap in place the worst case moves
from "one knob" to the aggregate of ring x queues x devices, example:

  ip link add v0 numtxqueues 4096 txqueuelen 32767 type veth
  tc qdisc add dev v0 root mq
    -> 4096 * 3 * 32767 * 8 = ~3.0 GiB (one command)
  50 tun devices x 256 queues x 32767 x 8 = ~3.1 GiB

Unfortunately tx_queue_len is a bit ambigious in meaning:
In some cases it means a ring size (which is pre-allocated, ex:
tun, tap, and pfifo_fast); a cap of 4096 seems reasonable here.
but in other cases it is used to indicate a queue limit ex:
the qdisc consumers that allocate nothing (pfifo/bfifo/gred/plug/sfb,
htb direct_qlen, qfq, teql). 32767 is a legitimate high-BDP queue
length, so we are going to keep that value.

Getting back to you sashikos, after this is merged and shows up
in net-next we will send followup patches as follows:
this series is not misread as "closes the OOM class"):

a) Per-site ring limits at six identified locations
    - pfifo_fast init/resize,
    - tun attach/resize,
    - tap minor/resize)

   if you can spot more in your review we will take care of those as well.

b) memcg accounting (GFP_KERNEL_ACCOUNT) for those ring
   allocations: contains a memcg-limited container's ring memory.
   Not GFP_KERNEL_ACCOUNT has no effect on the unshare attacker
   but will protect against containers  (memory.max in its cgroup)


Patches:
--------

  1/3 net: cap tx_queue_len at S16_MAX in netif_change_tx_queue_len()
      (netlink set, sysfs, SIOCSIFTXQLEN choke point)
  2/3 net: reject oversized tx_queue_len at netlink parse time
      (IFLA_TXQLEN policy: closes the create path + veth peer nest)
  3/3 selftests: tdc regression tests (netlink, sysfs, create paths)

Changes:
--------
v1 -> v2:
- new patch 2/3: close the device-creation path (Jakub Kicinski
  flagged that rtnl_create_link() bypasses the cap)
- rationale reworded: the 32767 ceiling citing NLA u32 range
  policy can express (s16 bounds), replacing the invalid virtio
  ring-depth claim
- rebuilt tdc coverage (sysfs path now tested; nondeterministic
  resize-rollback case dropped)

Sashiko v1 review links:
https://sashiko.dev/#/patchset/20260828121902.66837-1-jhs@mojatatu.com
https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260828121902.66837-1-jhs@mojatatu.com

v1: https://lore.kernel.org/netdev/20260828121902.66837-1-jhs@mojatatu.com/

Jamal Hadi Salim (3):
  net: cap tx_queue_len at S16_MAX to prevent oversized ring
    allocations
  net: reject oversized tx_queue_len at netlink parse time
  selftests: tc-testing: add tx_queue_len cap regression tests

 .../tc-testing/tc-tests/qdiscs/pfifo_fast.json | 209 +++++++++++++++++-
 net/core/dev.c                                 |   2 +-
 net/core/rtnetlink.c                           |   9 +-
 3 files changed, 211 insertions(+), 2 deletions(-)

-- 
2.43.0

^ permalink raw reply	[flat|nested] 8+ messages in thread

end of thread, other threads:[~2026-09-05  1:10 UTC | newest]

Thread overview: 8+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-02 21:29 [PATCH net v2 0/3] net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations Jamal Hadi Salim
2026-09-02 21:29 ` Jamal Hadi Salim
2026-09-02 21:29   ` [PATCH net v2 1/3] " Jamal Hadi Salim
2026-09-02 21:29   ` [PATCH net v2 2/3] net: reject oversized tx_queue_len at netlink parse time Jamal Hadi Salim
2026-09-05  1:10     ` netdev-bot+sashiko
2026-09-02 21:29   ` [PATCH net v2 3/3] selftests: tc-testing: add tx_queue_len cap regression tests Jamal Hadi Salim
2026-09-05  1:10     ` netdev-bot+sashiko
2026-09-04 23:40   ` [PATCH net v2 0/3] net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations patchwork-bot+netdevbpf
2026-09-05  1:10   ` [PATCH net v2 1/3] " netdev-bot+sashiko

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.