Netdev List
 help / color / mirror / Atom feed
From: Jamal Hadi Salim <jhs@mojatatu.com>
To: netdev@vger.kernel.org
Cc: Jamal Hadi Salim <jhs@mojatatu.com>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Jiri Pirko <jiri@resnulli.us>,
	Vinicius Costa Gomes <vinicius.gomes@intel.com>,
	Simon Horman <horms@kernel.org>,
	"Victor Nogueira" <victor@mojatatu.com>,
	Tonghao Zhang <xiangxia.m.yue@gmail.com>,
	Zero Day Initiative <zdi-disclosures@trendmicro.com>,
	hybris <hybris@mojatatu.ai>,
	stable@vger.kernel.org
Subject: [PATCH net] net: cap skb->queue_mapping when the tx queue is picked
Date: Mon, 21 Sep 2026 10:02:17 -0400	[thread overview]
Message-ID: <QDISC-9R8V.v1.20260921065103@mojatatu.com> (raw)

skbedit can set skb->queue_mapping and __dev_queue_xmit() honors it
through the skip_txqueue flag; netdev_tx_queue_mapping() clamps the
index it uses to select the netdev_queue but leaves the out-of-range
value in skb->queue_mapping.

Every later consumer of skb_get_queue_mapping()/skb_get_tx_queue() on
that path then reads past the device's queues. Taprio's child array
q->qdiscs[] is sized to the device's queue count, so taprio_enqueue()
indexes past its allocation; qdisc_restart() likewise dereferences
dev->_tx[queue_mapping].

A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer (mirred action) after setting a
mapping valid only on the larger device. That reaches these reads and,
under KASAN, faults with "slab-out-of-bounds in taprio_enqueue".

Store the clamped value back into skb->queue_mapping, as
netdev_core_pick_tx() already does for the mapping it picks, so the
whole egress path observes an in-range queue index.

I have looked at other alternative places to put this "fix", none
appealing: an out-of-range queue_mapping is read by every consumer
on the xmit path, not just by taprio. For example, upon testing
an approach that only bounds-checked taprio_enqueue() I observed
the fault relocated to sch_direct_xmit()/qdisc_restart() instead
(because dev->_tx[queue_mapping] is still indexed with the raw value).
Another approach was to cap it in skbedit;  cannot work: the
redirect target, whose queue count bounds the mapping, is not known
when the action runs, and act_mirred sets skb->dev afterwards.
So the decision is to cap the value where it is first trusted and
result is it fixes all downstream readers at once.

Conditions to recreate the bug: with CONFIG_NET_SCH_TAPRIO=y,
CONFIG_NET_ACT_SKBEDIT=y, CONFIG_NET_ACT_MIRRED=y and KASAN enabled,
create a 3-queue dummy qa and a 2-queue dummy qb, put a taprio root on
qb, then on qa's clsact add matchall with "action skbedit queue_mapping
2 pipe action mirred egress redirect dev qb" and send one packet out
qa. Mapping 2 is valid for qa but past qb's two-entry taprio child
array.

Reproduction: reproducer ran on a KASAN build with panic_on_warn=1:
the unfixed control faults with "BUG: KASAN: slab-out-of-bounds in
taprio_enqueue", a read 0 bytes past a 16-byte taprio_init allocation,
and panics.
The fixed kernel runs the same reproducer without a report, only the
expected ratelimited "qb selects TX queue 2, but real number of TX
queues is 2" notice.

Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
---
 net/core/dev.c | 9 +++++++--
 1 file changed, 7 insertions(+), 2 deletions(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index c67900354fa6..736b3664b635 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -4404,9 +4404,14 @@ EXPORT_SYMBOL(dev_loopback_xmit);
 static struct netdev_queue *
 netdev_tx_queue_mapping(struct net_device *dev, struct sk_buff *skb)
 {
-	int qm = skb_get_queue_mapping(skb);
+	int queue = skb_get_queue_mapping(skb);
+	int capped;
 
-	return netdev_get_tx_queue(dev, netdev_cap_txqueue(dev, qm));
+	capped = netdev_cap_txqueue(dev, queue);
+	if (unlikely(capped != queue))
+		skb_set_queue_mapping(skb, capped);
+
+	return netdev_get_tx_queue(dev, capped);
 }
 
 #ifndef CONFIG_PREEMPT_RT
-- 
2.43.0


             reply	other threads:[~2026-09-21 14:02 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-21 14:02 Jamal Hadi Salim [this message]
2026-09-21 14:42 ` [PATCH net] net: cap skb->queue_mapping when the tx queue is picked Eric Dumazet
2026-09-21 19:57   ` Jamal Hadi Salim
2026-09-22 14:03 ` netdev-bot+sashiko
2026-09-22 21:33   ` Jamal Hadi Salim
2026-09-24  8:06     ` Jamal Hadi Salim

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=QDISC-9R8V.v1.20260921065103@mojatatu.com \
    --to=jhs@mojatatu.com \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=hybris@mojatatu.ai \
    --cc=jiri@resnulli.us \
    --cc=kuba@kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=stable@vger.kernel.org \
    --cc=victor@mojatatu.com \
    --cc=vinicius.gomes@intel.com \
    --cc=xiangxia.m.yue@gmail.com \
    --cc=zdi-disclosures@trendmicro.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox