Netdev List
 help / color / mirror / Atom feed
* [PATCH net v2] net/sched: bound qdisc_pkt_len to prevent qdisc soft lockup
@ 2026-08-25  8:14 Jamal Hadi Salim
  2026-08-27 11:01 ` Paolo Abeni
  2026-08-27 19:40 ` patchwork-bot+netdevbpf
  0 siblings, 2 replies; 10+ messages in thread
From: Jamal Hadi Salim @ 2026-08-25  8:14 UTC (permalink / raw)
  To: netdev
  Cc: Jamal Hadi Salim, Jiri Pirko, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, stable, vega,
	Victor Nogueira

qdisc_get_stab() accepts a user-supplied size table, and
__qdisc_calculate_pkt_len() amplifies qdisc_pkt_len() through the
overhead, the size-table data (u16), and size_log (up to
STAB_SIZE_LOG_MAX). A crafted stab can therefore set qdisc_pkt_len()
to ~1 GiB for an ordinary skb. Per-flow deficit schedulers such as
DRR and ETS replenish one quantum per loop iteration; with a tiny
quantum (1) they spin billions of times under the qdisc lock,
producing a soft lockup / RCU stall as illustrated by vega@nebusec.ai.

Cap the final qdisc_pkt_len() to QDISC_PKT_LEN_MAX so the size-table
amplification cannot drive deficit schedulers into an unbounded loop.
A legitimate size table (e.g. qfq's overhead 999999999, which is
handled by dropping) is still accepted.

Introduce cap QDISC_PKT_LEN_MAX (1 << 20) = 1 MiB which is well above
any legitimate single-skb wire length: the largest current skb->len
is GSO_MAX_SIZE (524280), and an ATM-style size table (53/48 cell tax)
amplifies that to ~578 KB, both comfortably below 1 MiB. At the same
time, 1 MiB bounds the deficit refill loop to ~1M iterations per
packet with quantum=1, which completes in a few milliseconds well
under the demonstrated softlockup threshold (~10^9 iterations).

Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_NET_SCH_DRR=y (or CONFIG_NET_SCH_ETS=y).
- Attach a DRR (or ETS) root qdisc with a crafted TCA_STAB that
  amplifies qdisc_pkt_len to ~1 GiB (e.g. size_log=15, data=[32768]).
- Add a class with a tiny quantum of 1 and send one small packet; the
  deficit loop spins billions of times under the qdisc lock and trips
  the softlockup detector (panic with kernel.softlockup_panic=1).
- Reachable as root or from an unprivileged user in a fresh user+net
  namespace (unshare -Urn) with namespace-local CAP_NET_ADMIN.

Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Reported-by: vega@nebusec.ai
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
---
v1 -> v2:
- Use a dedicated QDISC_PKT_LEN_MAX define in include/net/pkt_sched.h
  instead of GSO_MAX_SIZE (Jakub Kicinski, Eric Dumazet).
  __qdisc_calculate_pkt_len() deals with single-skb wire-length
  accounting, not GSO segments; GSO_MAX_SIZE is semantically wrong.
  QDISC_PKT_LEN_MAX = (1 << 20) = 1 MiB avoids truncating legitimate
  stab accounting (BIG TCP + ATM-style 53/48 table ~ 578 KB) while
  still bounding the deficit loop to ~1M iterations at quantum=1
  (completes in a few ms, well under the softlockup threshold of
  ~10^9 iterations).
---
 include/net/pkt_sched.h | 1 +
 net/sched/sch_api.c     | 7 +++++--
 2 files changed, 6 insertions(+), 2 deletions(-)

diff --git a/include/net/pkt_sched.h b/include/net/pkt_sched.h
index 18a419cd9d94..90d3e7943b19 100644
--- a/include/net/pkt_sched.h
+++ b/include/net/pkt_sched.h
@@ -12,6 +12,7 @@
 
 #define DEFAULT_TX_QUEUE_LEN	1000
 #define STAB_SIZE_LOG_MAX	30
+#define QDISC_PKT_LEN_MAX	(1 << 20)	/* 1 MiB */
 
 struct qdisc_walker {
 	int	stop;
diff --git a/net/sched/sch_api.c b/net/sched/sch_api.c
index 65b35528d125..90503e59e6e3 100644
--- a/net/sched/sch_api.c
+++ b/net/sched/sch_api.c
@@ -610,8 +610,11 @@ void __qdisc_calculate_pkt_len(struct sk_buff *skb,
 
 	pkt_len <<= stab->szopts.size_log;
 out:
-	if (unlikely(pkt_len < 1))
-		pkt_len = 1;
+	/* A size table can inflate qdisc_pkt_len() beyond any real packet
+	 * (via overhead, the data table, or size_log); cap it so deficit
+	 * schedulers such as DRR/ETS terminate their refill loops.
+	 */
+	pkt_len = clamp_t(int, pkt_len, 1, QDISC_PKT_LEN_MAX);
 	qdisc_skb_cb(skb)->pkt_len = pkt_len;
 }
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread
* [PATCH net v2] net/sched: bound qdisc_pkt_len to prevent qdisc soft lockup
@ 2026-08-19 14:32 Jamal Hadi Salim
  2026-08-19 14:52 ` Eric Dumazet
  0 siblings, 1 reply; 10+ messages in thread
From: Jamal Hadi Salim @ 2026-08-19 14:32 UTC (permalink / raw)
  To: netdev
  Cc: Jamal Hadi Salim, Jiri Pirko, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, stable, vega,
	Victor Nogueira

qdisc_get_stab() accepts a user-supplied size table, and
__qdisc_calculate_pkt_len() amplifies qdisc_pkt_len() through the
overhead, the size-table data (u16), and size_log (up to
STAB_SIZE_LOG_MAX).  A crafted stab can therefore set qdisc_pkt_len()
to ~1 GiB for an ordinary skb.  Per-flow deficit schedulers such as
DRR and ETS replenish one quantum per loop iteration; with a tiny
quantum (1) they spin billions of times under the qdisc lock,
producing a soft lockup / RCU stall.

Cap the final qdisc_pkt_len() to GSO_MAX_SIZE so the size-table
amplification cannot drive deficit schedulers into an unbounded loop.
A legitimate size table (e.g. qfq's overhead 999999999, which is
handled by dropping) is still accepted.

Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_NET_SCH_DRR=y (or CONFIG_NET_SCH_ETS=y).
- Attach a DRR (or ETS) root qdisc with a crafted TCA_STAB that
  amplifies qdisc_pkt_len to ~1 GiB (e.g. size_log=15, data=[32768]).
- Add a class with a tiny quantum of 1 and send one small packet; the
  deficit loop spins billions of times under the qdisc lock and trips
  the softlockup detector (panic with kernel.softlockup_panic=1).
- Reachable as root or from an unprivileged user in a fresh user+net
  namespace (unshare -Urn) with namespace-local CAP_NET_ADMIN.

Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Reported-by: vega@nebusec.ai
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
---
v1 -> v2

1. Sigh. The v1 overhead cap broke the existing qfq tdc test 5993, caught by
   running the whole tdc.sh instead of affected qdiscs reported reported by
   poc. That test legitimately uses stab overhead 999999999 qfq and expects
   the qdisc to be accepted (exit 0) with packets dropped.

2. Better Fix: cap the final qdisc_pkt_len() to GSO_MAX_SIZE(524280) in
   __qdisc_calculate_pkt_len() per sashikos[1][2] suggestions

3. Given existence of tdc 5993 we dont need the tdc test created earlier
   since the essence of that test is covered in tdc 5993.

[1] https://sashiko.dev/#/patchset/20260818101735.16655-1-jhs@mojatatu.com
[2] https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260818101735.16655-1-jhs@mojatatu.com
---
 net/sched/sch_api.c | 7 +++++--
 1 file changed, 5 insertions(+), 2 deletions(-)

diff --git a/net/sched/sch_api.c b/net/sched/sch_api.c
index 65b35528d125..ad4f117ff55a 100644
--- a/net/sched/sch_api.c
+++ b/net/sched/sch_api.c
@@ -610,8 +610,11 @@ void __qdisc_calculate_pkt_len(struct sk_buff *skb,
 
 	pkt_len <<= stab->szopts.size_log;
 out:
-	if (unlikely(pkt_len < 1))
-		pkt_len = 1;
+	/* A size table can inflate qdisc_pkt_len() beyond any real packet
+	 * (via overhead, the data table, or size_log); cap it so deficit
+	 * schedulers such as DRR/ETS terminate their refill loops.
+	 */
+	pkt_len = clamp_t(int, pkt_len, 1, GSO_MAX_SIZE);
 	qdisc_skb_cb(skb)->pkt_len = pkt_len;
 }
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 10+ messages in thread

end of thread, other threads:[~2026-08-27 19:41 UTC | newest]

Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-25  8:14 [PATCH net v2] net/sched: bound qdisc_pkt_len to prevent qdisc soft lockup Jamal Hadi Salim
2026-08-27 11:01 ` Paolo Abeni
2026-08-27 17:48   ` Jamal Hadi Salim
2026-08-27 19:40 ` patchwork-bot+netdevbpf
  -- strict thread matches above, loose matches on Subject: below --
2026-08-19 14:32 Jamal Hadi Salim
2026-08-19 14:52 ` Eric Dumazet
2026-08-19 14:58   ` Jamal Hadi Salim
2026-08-19 15:15     ` Eric Dumazet
2026-08-19 15:30       ` Jamal Hadi Salim
2026-08-24 18:36         ` Jakub Kicinski

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox