* [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked
@ 2026-09-28 12:46 Jamal Hadi Salim
2026-09-28 12:55 ` Eric Dumazet
` (2 more replies)
0 siblings, 3 replies; 5+ messages in thread
From: Jamal Hadi Salim @ 2026-09-28 12:46 UTC (permalink / raw)
To: netdev
Cc: Jamal Hadi Salim, Pablo Neira Ayuso, Florian Westphal,
Phil Sutter, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Jiri Pirko, Vinicius Costa Gomes, Simon Horman,
Tonghao Zhang, Sebastian Andrzej Siewior, Clark Williams,
Steven Rostedt, Victor Nogueira, Zero Day Initiative, hybris,
netfilter-devel, coreteam, linux-rt-devel, stable, Eric Dumazet
skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue
flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
cleared the flag before sch_handle_egress() and only read it afterwards,
so the flag was not confined to the xmit that set it: a nested xmit
(mirred redirect or mirror, or a drop after skbedit) could set the flag
and the outer xmit would consume it for an skb that never went through
skbedit.
A forwarded packet still carries the ingress NIC's rx_queue + 1 in
skb->queue_mapping, so the outer device then indexes its tx queue state
with that stale value. Taprio's child array q->qdiscs[] is sized to the
device's queue count, so taprio_enqueue() indexes past its allocation
and dereferences the result as a struct Qdisc *.
We (ab)use the skb->nf_skip_egress which means "skip netfilter egress
for this packet" to tag to "am I in tc egress?". Despite the overload
I dont see it as a conflict since the marker is set only around the
single sch_handle_egress() call and ingress path is guarded by
tc_at_ingress.
I will send a followup(net-next) patch once this hits net-next to
rename the skb->nf_skip_egress bit/flag to skb->skip_egress
Arm the flag only from the egress classifier that can use it: raise
skip_txqueue from tcf_skbedit_act() only when it runs inside
sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc
classifier runs in q->enqueue(), after the tx queue has been picked,
so a mapping it sets cannot affect the current packet; arming the flag
there only pollutes it for a later xmit. Then own the flag for the xmit
frame the egress hook runs in: save the incoming value and clear it just
before sch_handle_egress(), and restore it after the hook - on the
consumed (drop) path, or, in the same call that reads it, on the
surviving path. The save and the restores stay inside the
egress_needed_key static branch, so a packet pays for them only when
egress hooks are active (2f1e85b1aee4).
Store the value netdev_cap_txqueue() selected back into skb->queue_mapping
in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
the skip_txqueue path never hands a later reader on the xmit path a
mapping the device cannot serve. A store made still later in the same
frame, by a tc BPF program attached to a transmit qdisc, is outside this
path and is not re-capped; a separate followup will resolve that path.
netdev_xmit_skip_txqueue() returns the previous flag value so the
save-and-clear is one call, and a no-op stub is provided when
CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under
CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
skb_at_tc_egress() is valid whenever the egress path is built.
A local user in a network namespace can redirect a packet from a device
with more TX queues to one with fewer after setting a mapping valid only
on the larger device. That reaches these reads and, under KASAN, faults
with "slab-out-of-bounds in taprio_enqueue".
Conditions to recreate the bug: the report's own trigger is a local user
with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
(2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
(num_tc 1, queues 2@0) and clsact on all three; then add an egress
matchall filter on every device. On qa: "action skbedit queue_mapping 2
pipe action mirred egress redirect dev qb". On qb: "action mirred egress
mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
flight; without the fix qb consumes it and reads its two-entry taprio
child array with the forwarded packet's stale mapping. A qc whose
skbedit is instead installed in a transmit-qdisc classifier (a matchall
filter on the qc root qdisc) reaches the same code path the same way
without the fix.
Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
past a 16-byte taprio_init() allocation, for the clsact-setter and the
transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
store; the fixed kernel runs all three with no report, and the BPF store
variant additionally shows the expected "selects TX queue" clamp notice
from the write-back.
Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
Suggested-by: Eric Dumazet <edumazet@google.com>
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Tested-by: hybris <hybris@mojatatu.ai>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
---
v3 -> v4:
- Arm skip_txqueue from tcf_skbedit_act() only inside sch_handle_egress()
(new skb_at_tc_egress() helper); a transmit-qdisc classifier runs after
the tx queue is picked, so its flag set could not affect the current
packet and only leaked to a later xmit (Eric Dumazet).
- Move the save/clear/restore back inside the egress_needed_key static
branch: with the only producer inside sch_handle_egress() the
branch-local lifetime is sufficient, and packets no longer pay a
per-CPU read/write when no egress hook is active (Eric Dumazet).
- Compile skb->nf_skip_egress under CONFIG_NET_EGRESS rather than
CONFIG_NETFILTER_SKIP_EGRESS.
- Keep the retval helper, the !CONFIG_NET_EGRESS stub and the
netdev_tx_queue_mapping() capped write-back.
- Add to Cc more stake holders.
v2 -> v3:
- netdev_xmit_skip_txqueue() returns the previous flag value, so the entry
save-and-clear is one call; a no-op stub is provided when
!CONFIG_NET_EGRESS, removing the three inline #ifdefs (Jakub Kicinski).
- Narrow the capability claim: the capped write-back covers the
skip_txqueue path only; a store made later in the frame by a tc BPF
program on a transmit qdisc is not re-capped and is a separate followup
(nipa Sashiko).
- Comment: "per-CPU (per-task on PREEMPT_RT)" (nipa Sashiko).
- Drop the now-redundant #ifdef in act_skbedit.c.
v1 -> v2:
- Rework the root cause to the per-CPU skip_txqueue flag lifetime: it was
cleared before sch_handle_egress() and read after, so a nested xmit could
set it and the outer xmit consume it for an skb that never went through
skbedit (Eric Dumazet, nipa Sashiko).
- Retain the v1 producer-side cap (netdev_tx_queue_mapping() writes the
clamped value back).
- Correct the description of the reproducer to a three-device chain
(qa 3q -> qb 2q taprio -> qc 1q).
include/linux/netfilter_netdev.h | 2 +-
include/linux/rtnetlink.h | 7 ++++++-
include/linux/skbuff.h | 2 +-
include/net/sch_generic.h | 9 +++++++++
net/core/dev.c | 31 ++++++++++++++++++++++++-------
net/sched/act_skbedit.c | 5 ++---
6 files changed, 43 insertions(+), 13 deletions(-)
diff --git a/include/linux/netfilter_netdev.h b/include/linux/netfilter_netdev.h
index 3175073a66ba..2a854a0bbd4b 100644
--- a/include/linux/netfilter_netdev.h
+++ b/include/linux/netfilter_netdev.h
@@ -133,7 +133,7 @@ static inline struct sk_buff *nf_hook_egress(struct sk_buff *skb, int *rc,
static inline void nf_skip_egress(struct sk_buff *skb, bool skip)
{
-#ifdef CONFIG_NETFILTER_SKIP_EGRESS
+#ifdef CONFIG_NET_EGRESS
skb->nf_skip_egress = skip;
#endif
}
diff --git a/include/linux/rtnetlink.h b/include/linux/rtnetlink.h
index 95729339e7a5..a54ec40d095c 100644
--- a/include/linux/rtnetlink.h
+++ b/include/linux/rtnetlink.h
@@ -186,7 +186,12 @@ void net_dec_ingress_queue(void);
#ifdef CONFIG_NET_EGRESS
void net_inc_egress_queue(void);
void net_dec_egress_queue(void);
-void netdev_xmit_skip_txqueue(bool skip);
+bool netdev_xmit_skip_txqueue(bool skip);
+#else
+static inline bool netdev_xmit_skip_txqueue(bool skip)
+{
+ return false;
+}
#endif
void rtnetlink_init(void);
diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
index 84308498a3a8..a3ff380d43f0 100644
--- a/include/linux/skbuff.h
+++ b/include/linux/skbuff.h
@@ -1020,7 +1020,7 @@ struct sk_buff {
#ifdef CONFIG_NET_REDIRECT
__u8 from_ingress:1;
#endif
-#ifdef CONFIG_NETFILTER_SKIP_EGRESS
+#ifdef CONFIG_NET_EGRESS
__u8 nf_skip_egress:1;
#endif
#ifdef CONFIG_SKB_DECRYPTED
diff --git a/include/net/sch_generic.h b/include/net/sch_generic.h
index f35bd06a6bad..1acaadb3e2cf 100644
--- a/include/net/sch_generic.h
+++ b/include/net/sch_generic.h
@@ -810,6 +810,15 @@ static inline bool skb_at_tc_ingress(const struct sk_buff *skb)
#endif
}
+static inline bool skb_at_tc_egress(const struct sk_buff *skb)
+{
+#ifdef CONFIG_NET_EGRESS
+ return skb->nf_skip_egress && !skb_at_tc_ingress(skb);
+#else
+ return false;
+#endif
+}
+
static inline bool skb_skip_tc_classify(struct sk_buff *skb)
{
#ifdef CONFIG_NET_CLS_ACT
diff --git a/net/core/dev.c b/net/core/dev.c
index f660fccfc0db..8638c994bbcb 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -4404,9 +4404,14 @@ EXPORT_SYMBOL(dev_loopback_xmit);
static struct netdev_queue *
netdev_tx_queue_mapping(struct net_device *dev, struct sk_buff *skb)
{
- int qm = skb_get_queue_mapping(skb);
+ int queue = skb_get_queue_mapping(skb);
+ int capped;
- return netdev_get_tx_queue(dev, netdev_cap_txqueue(dev, qm));
+ capped = netdev_cap_txqueue(dev, queue);
+ if (unlikely(capped != queue))
+ skb_set_queue_mapping(skb, capped);
+
+ return netdev_get_tx_queue(dev, capped);
}
#ifndef CONFIG_PREEMPT_RT
@@ -4415,9 +4420,13 @@ static bool netdev_xmit_txqueue_skipped(void)
return __this_cpu_read(softnet_data.xmit.skip_txqueue);
}
-void netdev_xmit_skip_txqueue(bool skip)
+bool netdev_xmit_skip_txqueue(bool skip)
{
+ bool prev = netdev_xmit_txqueue_skipped();
+
__this_cpu_write(softnet_data.xmit.skip_txqueue, skip);
+
+ return prev;
}
EXPORT_SYMBOL_GPL(netdev_xmit_skip_txqueue);
@@ -4427,9 +4436,13 @@ static bool netdev_xmit_txqueue_skipped(void)
return current->net_xmit.skip_txqueue;
}
-void netdev_xmit_skip_txqueue(bool skip)
+bool netdev_xmit_skip_txqueue(bool skip)
{
+ bool prev = netdev_xmit_txqueue_skipped();
+
current->net_xmit.skip_txqueue = skip;
+
+ return prev;
}
EXPORT_SYMBOL_GPL(netdev_xmit_skip_txqueue);
#endif
@@ -4848,21 +4861,25 @@ int __dev_queue_xmit(struct sk_buff *skb, struct net_device *sb_dev)
tcx_set_ingress(skb, false);
#ifdef CONFIG_NET_EGRESS
if (static_branch_unlikely(&egress_needed_key)) {
+ bool skip_txq;
+
if (nf_hook_egress_active()) {
skb = nf_hook_egress(skb, &rc, dev);
if (!skb)
goto out;
}
- netdev_xmit_skip_txqueue(false);
+ skip_txq = netdev_xmit_skip_txqueue(false);
nf_skip_egress(skb, true);
skb = sch_handle_egress(skb, &rc, dev);
- if (!skb)
+ if (!skb) {
+ netdev_xmit_skip_txqueue(skip_txq);
goto out;
+ }
nf_skip_egress(skb, false);
- if (netdev_xmit_txqueue_skipped())
+ if (netdev_xmit_skip_txqueue(skip_txq))
txq = netdev_tx_queue_mapping(dev, skb);
}
#endif
diff --git a/net/sched/act_skbedit.c b/net/sched/act_skbedit.c
index bfec6b668410..26258b353cde 100644
--- a/net/sched/act_skbedit.c
+++ b/net/sched/act_skbedit.c
@@ -72,9 +72,8 @@ TC_INDIRECT_SCOPE int tcf_skbedit_act(struct sk_buff *skb,
}
if (params->flags & SKBEDIT_F_QUEUE_MAPPING &&
skb->dev->real_num_tx_queues > params->queue_mapping) {
-#ifdef CONFIG_NET_EGRESS
- netdev_xmit_skip_txqueue(true);
-#endif
+ if (skb_at_tc_egress(skb))
+ netdev_xmit_skip_txqueue(true);
skb_set_queue_mapping(skb, tcf_skbedit_hash(params, skb));
}
if (params->flags & SKBEDIT_F_MARK) {
^ permalink raw reply related [flat|nested] 5+ messages in thread* Re: [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked
2026-09-28 12:46 [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked Jamal Hadi Salim
@ 2026-09-28 12:55 ` Eric Dumazet
2026-09-29 3:45 ` netdev-bot+sashiko
2026-09-30 1:00 ` patchwork-bot+netdevbpf
2 siblings, 0 replies; 5+ messages in thread
From: Eric Dumazet @ 2026-09-28 12:55 UTC (permalink / raw)
To: Jamal Hadi Salim
Cc: netdev, Pablo Neira Ayuso, Florian Westphal, Phil Sutter,
David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Jiri Pirko, Vinicius Costa Gomes, Simon Horman, Tonghao Zhang,
Sebastian Andrzej Siewior, Clark Williams, Steven Rostedt,
Victor Nogueira, Zero Day Initiative, hybris, netfilter-devel,
coreteam, linux-rt-devel, stable
On Mon, Sep 28, 2026 at 2:46 PM Jamal Hadi Salim <jhs@mojatatu.com> wrote:
>
> skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue
> flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
> cleared the flag before sch_handle_egress() and only read it afterwards,
> so the flag was not confined to the xmit that set it: a nested xmit
> (mirred redirect or mirror, or a drop after skbedit) could set the flag
> and the outer xmit would consume it for an skb that never went through
> skbedit.
>
> A forwarded packet still carries the ingress NIC's rx_queue + 1 in
> skb->queue_mapping, so the outer device then indexes its tx queue state
> with that stale value. Taprio's child array q->qdiscs[] is sized to the
> device's queue count, so taprio_enqueue() indexes past its allocation
> and dereferences the result as a struct Qdisc *.
>
> We (ab)use the skb->nf_skip_egress which means "skip netfilter egress
> for this packet" to tag to "am I in tc egress?". Despite the overload
> I dont see it as a conflict since the marker is set only around the
> single sch_handle_egress() call and ingress path is guarded by
> tc_at_ingress.
> I will send a followup(net-next) patch once this hits net-next to
> rename the skb->nf_skip_egress bit/flag to skb->skip_egress
>
> Arm the flag only from the egress classifier that can use it: raise
> skip_txqueue from tcf_skbedit_act() only when it runs inside
> sch_handle_egress(), thanks to skb->nf_skip_egress. An egress qdisc
> classifier runs in q->enqueue(), after the tx queue has been picked,
> so a mapping it sets cannot affect the current packet; arming the flag
> there only pollutes it for a later xmit. Then own the flag for the xmit
> frame the egress hook runs in: save the incoming value and clear it just
> before sch_handle_egress(), and restore it after the hook - on the
> consumed (drop) path, or, in the same call that reads it, on the
> surviving path. The save and the restores stay inside the
> egress_needed_key static branch, so a packet pays for them only when
> egress hooks are active (2f1e85b1aee4).
>
> Store the value netdev_cap_txqueue() selected back into skb->queue_mapping
> in netdev_tx_queue_mapping(), as netdev_core_pick_tx() already does, so
> the skip_txqueue path never hands a later reader on the xmit path a
> mapping the device cannot serve. A store made still later in the same
> frame, by a tc BPF program attached to a transmit qdisc, is outside this
> path and is not re-capped; a separate followup will resolve that path.
>
> netdev_xmit_skip_txqueue() returns the previous flag value so the
> save-and-clear is one call, and a no-op stub is provided when
> CONFIG_NET_EGRESS is disabled. skb->nf_skip_egress is compiled under
> CONFIG_NET_EGRESS rather than CONFIG_NETFILTER_SKIP_EGRESS, so
> skb_at_tc_egress() is valid whenever the egress path is built.
>
> A local user in a network namespace can redirect a packet from a device
> with more TX queues to one with fewer after setting a mapping valid only
> on the larger device. That reaches these reads and, under KASAN, faults
> with "slab-out-of-bounds in taprio_enqueue".
>
> Conditions to recreate the bug: the report's own trigger is a local user
> with CAP_NET_ADMIN in a network namespace, so no eBPF program is needed.
> With CONFIG_NET_SCH_TAPRIO=y, CONFIG_NET_ACT_SKBEDIT=y,
> CONFIG_NET_ACT_MIRRED=y, CONFIG_NET_CLS_MATCHALL=y,
> CONFIG_NET_SCH_PRIO=y and KASAN enabled, create qa (3 queues), qb
> (2 queues) and qc (1 queue) as dummy devices; put a taprio root on qb
> (num_tc 1, queues 2@0) and clsact on all three; then add an egress
> matchall filter on every device. On qa: "action skbedit queue_mapping 2
> pipe action mirred egress redirect dev qb". On qb: "action mirred egress
> mirror dev qc". On qc: "action skbedit queue_mapping 0 pipe". Send one
> packet out qa. qc's skbedit sets the flag while qb's outer xmit is in
> flight; without the fix qb consumes it and reads its two-entry taprio
> child array with the forwarded packet's stale mapping. A qc whose
> skbedit is instead installed in a transmit-qdisc classifier (a matchall
> filter on the qc root qdisc) reaches the same code path the same way
> without the fix.
>
> Testing: on a KASAN build with panic_on_warn=1 the unfixed kernel panics
> with "BUG: KASAN: slab-out-of-bounds in taprio_enqueue", a read 0 bytes
> past a 16-byte taprio_init() allocation, for the clsact-setter and the
> transmit-qdisc-classifier reproducers and for a clsact skbedit-then-tc-BPF
> store; the fixed kernel runs all three with no report, and the BPF store
> variant additionally shows the expected "selects TX queue" clamp notice
> from the write-back.
>
> Fixes: 2f1e85b1aee4 ("net: sched: use queue_mapping to pick tx queue")
> Reported-by: Zero Day Initiative <zdi-disclosures@trendmicro.com>
> Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
> Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
> Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
> Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
> Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
> Suggested-by: Eric Dumazet <edumazet@google.com>
> Suggested-by: Jakub Kicinski <kuba@kernel.org>
> Tested-by: hybris <hybris@mojatatu.ai>
> Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
SGTM, thanks Jamal!
Reviewed-by: Eric Dumazet <edumazet@google.com>
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked
2026-09-28 12:46 [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked Jamal Hadi Salim
2026-09-28 12:55 ` Eric Dumazet
@ 2026-09-29 3:45 ` netdev-bot+sashiko
2026-09-29 9:53 ` Jamal Hadi Salim
2026-09-30 1:00 ` patchwork-bot+netdevbpf
2 siblings, 1 reply; 5+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 3:45 UTC (permalink / raw)
To: jhs
Cc: netdev, pablo, fw, phil, davem, edumazet, kuba, pabeni, jiri,
vinicius.gomes, horms, xiangxia.m.yue, bigeasy, clrkwllms,
rostedt, victor, zdi-disclosures, hybris, netfilter-devel,
coreteam, linux-rt-devel, stable, edumazet
Thank you for your contribution! Sashiko AI review found 3 potential
issue(s) to consider.
Critical: 0 · High: 0 (1) · Medium: 2 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] skb_at_tc_egress() treats skb->nf_skip_egress &&
!tc_at_ingress as meaning "inside sch_handle_egress()".
- [Medium] The fix keeps skip_txqueue inside one xmit frame only if every
nested __dev_queue_xmit() takes the egress_needed_key branch.
- [Low] After this patch, skb->nf_skip_egress also gates tc behaviour:
skb_at_tc_egress() decides whether tcf_skbedit_act() arms skip_txqueue.
Pre-existing issues:
- [High] This problem wasn't introduced by this patch, but it is high
severity.
--
net: cap skb->queue_mapping when the tx queue is picked
This patch keeps the per-CPU skip_txqueue flag inside the xmit frame that
runs sch_handle_egress(). It saves the flag before the hook and restores
it afterwards. tcf_skbedit_act() may now arm the flag only when
skb_at_tc_egress() reports tc egress. The patch also writes the capped
value back into skb->queue_mapping in netdev_tx_queue_mapping(), and
builds skb->nf_skip_egress under CONFIG_NET_EGRESS.
Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
[ ... ]
> diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
> index 84308498a3a8..a3ff380d43f0 100644
> --- a/include/linux/skbuff.h
> +++ b/include/linux/skbuff.h
> @@ -1020,7 +1020,7 @@ struct sk_buff {
> #ifdef CONFIG_NET_REDIRECT
> __u8 from_ingress:1;
> #endif
> -#ifdef CONFIG_NETFILTER_SKIP_EGRESS
> +#ifdef CONFIG_NET_EGRESS
> __u8 nf_skip_egress:1;
> #endif
[Severity: Low]
This isn't a bug, but skb_at_tc_egress() now reads this bit to gate
tcf_skbedit_act(), and the documentation still describes a flag used only
by netfilter.
The struct sk_buff kernel-doc says:
* @nf_skip_egress: packet shall skip nf egress - see netfilter_netdev.h
The comment above nf_hook_egress() in include/linux/netfilter_netdev.h
says:
* The nf_skip_egress flag controls whether netfilter is applied on egress.
* It is updated by __netif_receive_skb_core() and __dev_queue_xmit() when the
Neither comment mentions the tc egress meaning. The second one also leaves
out ifb_ri_tasklet() and netkit_prep_forward(), which write the bit too.
The new skb_at_tc_egress() has no comment about reusing this bit or about
the precondition it depends on.
Could these comments be updated along with this change?
> #ifdef CONFIG_SKB_DECRYPTED
> diff --git a/include/net/sch_generic.h b/include/net/sch_generic.h
> index f35bd06a6bad..1acaadb3e2cf 100644
> --- a/include/net/sch_generic.h
> +++ b/include/net/sch_generic.h
> @@ -810,6 +810,15 @@ static inline bool skb_at_tc_ingress(const struct sk_buff *skb)
> #endif
> }
>
> +static inline bool skb_at_tc_egress(const struct sk_buff *skb)
> +{
> +#ifdef CONFIG_NET_EGRESS
> + return skb->nf_skip_egress && !skb_at_tc_ingress(skb);
[Severity: Medium]
Is nf_skip_egress only ever set around sch_handle_egress()? The commit
message says:
"the marker is set only around the single sch_handle_egress() call and
ingress path is guarded by tc_at_ingress"
Some other writers seem to leave the bit set on skbs that later reach
__dev_queue_xmit().
__netif_receive_skb_core() sets the bit before sch_handle_ingress(), but
clears it only on the path where the skb survives:
nf_skip_egress(skb, true);
skb = sch_handle_ingress(skb, &pt_prev, &ret, orig_dev,
&another);
...
if (!skb)
goto out;
nf_skip_egress(skb, false);
An ingress mirred redirect consumes the skb and passes it to
dev_queue_xmit() with the bit still set. A mirror does the same with a
clone, which copies the bit through the headers group.
ifb_ri_tasklet() sets the bit and then calls dev_queue_xmit(). On the
from_ingress path it reinjects the skb with tc_skip_classify set, so
__netif_receive_skb_core() never resets the bit:
skb->tc_skip_classify = 1;
...
nf_skip_egress(skb, true);
netkit_prep_forward() also sets the bit. After that, netkit_xmit() calls
skb_do_redirect(), which can reach dev_queue_xmit().
In __dev_queue_xmit(), tcx_set_ingress(skb, false) always runs. The only
reset of nf_skip_egress is inside the egress_needed_key static branch.
Suppose only an ingress clsact is installed, for example one doing
"mirred egress redirect dev qc". The skb then reaches qc's root qdisc
enqueue with nf_skip_egress=1 and tc_at_ingress=0. A skbedit queue_mapping
action in a prio, multiq or htb classifier passes skb_at_tc_egress() and
calls netdev_xmit_skip_txqueue(true). The v4 changelog says this kind of
arming from a transmit-qdisc classifier no longer happens.
Nothing seems to consume this stale value today. Every reader inside the
static branch saves, clears and restores the flag. The one exception is
the static key race described further down.
There is also a config effect. nf_skip_egress() is now built under
CONFIG_NET_EGRESS instead of CONFIG_NETFILTER_SKIP_EGRESS. With
CONFIG_NET_EGRESS=y and CONFIG_NETFILTER_EGRESS=n, the ingress, ifb and
netkit calls used to compile to nothing. Now they write a bit that
skb_at_tc_egress() reads.
Should the reasoning in the commit message be revisited? Otherwise the
planned rename to skb->skip_egress would carry the same assumption
forward.
> +#else
> + return false;
> +#endif
> +}
[ ... ]
> diff --git a/net/core/dev.c b/net/core/dev.c
> index f660fccfc0db..8638c994bbcb 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -4404,9 +4404,14 @@ EXPORT_SYMBOL(dev_loopback_xmit);
> static struct netdev_queue *
> netdev_tx_queue_mapping(struct net_device *dev, struct sk_buff *skb)
> {
> - int qm = skb_get_queue_mapping(skb);
> + int queue = skb_get_queue_mapping(skb);
> + int capped;
>
> - return netdev_get_tx_queue(dev, netdev_cap_txqueue(dev, qm));
> + capped = netdev_cap_txqueue(dev, queue);
> + if (unlikely(capped != queue))
> + skb_set_queue_mapping(skb, capped);
> +
> + return netdev_get_tx_queue(dev, capped);
> }
[Severity: High]
This is a pre-existing issue and was not introduced by this patch. The
commit message leaves it for a separate followup. Can a cls_bpf program on
a transmit qdisc still index past the end of dev->_tx[]?
tc_cls_act_is_valid_access() allows writes to __sk_buff queue_mapping.
bpf_convert_ctx_access() only drops stores of NO_QUEUE_MAPPING or larger:
net/core/filter.c:bpf_convert_ctx_access() {
...
if (BPF_CLASS(si->code) == BPF_ST && si->imm >= NO_QUEUE_MAPPING) {
*insn++ = BPF_JMP_A(0); /* noop */
break;
}
if (BPF_CLASS(si->code) == BPF_STX)
*insn++ = BPF_JMP_IMM(BPF_JGE, si->src_reg, NO_QUEUE_MAPPING, 1);
*insn++ = BPF_EMIT_STORE(BPF_H, si, offset);
...
}
The stored value is never compared with dev->num_tx_queues or
real_num_tx_queues. A program attached under a prio or htb root runs
inside q->enqueue(). That is after netdev_core_pick_tx() or
netdev_tx_queue_mapping() has already capped the mapping, so the program
can store any value up to 0xfffe.
On dequeue:
qdisc_restart()
txq = skb_get_tx_queue(dev, skb)
netdev_get_tx_queue()
DEBUG_NET_WARN_ON_ONCE(index >= dev->num_tx_queues);
return &dev->_tx[index];
sch_direct_xmit(skb, q, dev, txq, ...)
HARD_TX_LOCK(dev, txq, cpu)
The gso_skb and skb_bad_txq requeue paths in dequeue_skb() also call
skb_get_tx_queue().
That would take txq->_xmit_lock, write xmit_lock_owner and read txq->state
in memory past the end of the dev->_tx allocation. Loading the program
needs CAP_BPF and CAP_NET_ADMIN, or a delegated BPF token. Even so, the
program has passed the verifier and should not be able to corrupt kernel
memory.
[ ... ]
> @@ -4848,21 +4861,25 @@ int __dev_queue_xmit(struct sk_buff *skb, struct net_device *sb_dev)
> tcx_set_ingress(skb, false);
> #ifdef CONFIG_NET_EGRESS
> if (static_branch_unlikely(&egress_needed_key)) {
> + bool skip_txq;
> +
> if (nf_hook_egress_active()) {
> skb = nf_hook_egress(skb, &rc, dev);
> if (!skb)
> goto out;
> }
>
> - netdev_xmit_skip_txqueue(false);
> + skip_txq = netdev_xmit_skip_txqueue(false);
>
> nf_skip_egress(skb, true);
> skb = sch_handle_egress(skb, &rc, dev);
> - if (!skb)
> + if (!skb) {
> + netdev_xmit_skip_txqueue(skip_txq);
> goto out;
> + }
> nf_skip_egress(skb, false);
>
> - if (netdev_xmit_txqueue_skipped())
> + if (netdev_xmit_skip_txqueue(skip_txq))
> txq = netdev_tx_queue_mapping(dev, skb);
> }
> #endif
[Severity: Medium]
Does the flag stay confined to this frame if egress_needed_key is turned
off while an outer frame is inside sch_handle_egress()?
Each nested __dev_queue_xmit() reads the static key again.
net_dec_egress_queue() is a plain static_branch_dec() and does not wait
for RCU-bh readers that are already inside the branch. Consider:
CPU1 CPU2
__dev_queue_xmit(qa)
static key is on
skip_txq = netdev_xmit_skip_txqueue(false)
nf_skip_egress(skb, true)
sch_handle_egress()
clsact_destroy()
net_dec_egress_queue()
static_branch_dec()
tcf_mirred_forward()
(mirror clone has nf_skip_egress=1)
dev_queue_xmit(qc)
static key is off: no save/restore,
no nf_skip_egress reset
q->enqueue() -> prio classifier
tcf_skbedit_act()
skb_at_tc_egress() is true
netdev_xmit_skip_txqueue(true)
nf_skip_egress(skb, false)
netdev_xmit_skip_txqueue(skip_txq) returns true
txq = netdev_tx_queue_mapping(dev, skb)
In this sequence, qa's surviving skb picks its tx queue from its own
queue_mapping, even though it never went through skbedit. For a forwarded
packet that value is the ingress rx_queue + 1.
The capped write-back in netdev_tx_queue_mapping() still prevents the
taprio out-of-bounds access. The result would be a wrong but valid tx
queue, and possibly the ratelimited "selects TX queue" warning.
The commit message says the flag is now owned by the xmit frame the
egress hook runs in. Is this cross-frame leak meant to be covered as well?
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/QDISC-9R8V.v4.20260928081529%40mojatatu.com
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked
2026-09-29 3:45 ` netdev-bot+sashiko
@ 2026-09-29 9:53 ` Jamal Hadi Salim
0 siblings, 0 replies; 5+ messages in thread
From: Jamal Hadi Salim @ 2026-09-29 9:53 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: netdev, pablo, fw, phil, davem, edumazet, kuba, pabeni, jiri,
vinicius.gomes, horms, xiangxia.m.yue, bigeasy, clrkwllms,
rostedt, victor, zdi-disclosures, hybris, netfilter-devel,
coreteam, linux-rt-devel, stable, edumazet
On Mon, Sep 28, 2026 at 11:45 PM <netdev-bot+sashiko@kernel.org> wrote:
>
> Thank you for your contribution! Sashiko AI review found 3 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 (1) · Medium: 2 · Low: 1
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [Medium] skb_at_tc_egress() treats skb->nf_skip_egress &&
> !tc_at_ingress as meaning "inside sch_handle_egress()".
will be handled in the rename followup
> - [Medium] The fix keeps skip_txqueue inside one xmit frame only if every
> nested __dev_queue_xmit() takes the egress_needed_key branch.
This is a valid issue but not a regression introduced by us.
Prior to this patch - the entire egress block is inside if
(static_branch_unlikely(&egress_needed_key)), and each nested
__dev_queue_xmit() re reads the key. TBH, the patch makes it more
observable in that window, but the resulting wrong-queue consequence
already existed.
So i will add this as a followup improvement to the renaming thread we
already captured.
> - [Low] After this patch, skb->nf_skip_egress also gates tc behaviour:
> skb_at_tc_egress() decides whether tcf_skbedit_act() arms skip_txqueue.
This documentation nit will be fixed when we send the rename followup.
> Pre-existing issues:
> - [High] This problem wasn't introduced by this patch, but it is high
> severity.
We are tracking this issue as a followup.
Unfortunately the followup cant be sent because it is dependent on
this commit being merged.
For this reason I dont see any need to respin v5.
cheers,
jamal
> --
>
> net: cap skb->queue_mapping when the tx queue is picked
>
> This patch keeps the per-CPU skip_txqueue flag inside the xmit frame that
> runs sch_handle_egress(). It saves the flag before the hook and restores
> it afterwards. tcf_skbedit_act() may now arm the flag only when
> skb_at_tc_egress() reports tc egress. The patch also writes the capped
> value back into skb->queue_mapping in netdev_tx_queue_mapping(), and
> builds skb->nf_skip_egress under CONFIG_NET_EGRESS.
>
> Link: https://lore.kernel.org/netdev/CANn89iLwYx8nCVf0pCEk_MmEiyC6kQaMwCQT9WkQVeeNzNQHqQ@mail.gmail.com/
> Link: https://lore.kernel.org/netdev/179008581937.2160803.7117814290574262942@kernel.org/
> Link: https://lore.kernel.org/netdev/179033713973.2160803.4914570693994398206@kernel.org/
> Link: https://lore.kernel.org/netdev/20260925180407.63647514@kernel.org/
> Link: https://lore.kernel.org/netdev/CANn89i+k-mZKDQVtvws_MEXeuMTAdaCcOXFZE-RfhcGTu90sjA@mail.gmail.com/
>
> [ ... ]
>
> > diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
> > index 84308498a3a8..a3ff380d43f0 100644
> > --- a/include/linux/skbuff.h
> > +++ b/include/linux/skbuff.h
> > @@ -1020,7 +1020,7 @@ struct sk_buff {
> > #ifdef CONFIG_NET_REDIRECT
> > __u8 from_ingress:1;
> > #endif
> > -#ifdef CONFIG_NETFILTER_SKIP_EGRESS
> > +#ifdef CONFIG_NET_EGRESS
> > __u8 nf_skip_egress:1;
> > #endif
>
> [Severity: Low]
> This isn't a bug, but skb_at_tc_egress() now reads this bit to gate
> tcf_skbedit_act(), and the documentation still describes a flag used only
> by netfilter.
>
> The struct sk_buff kernel-doc says:
>
> * @nf_skip_egress: packet shall skip nf egress - see netfilter_netdev.h
>
> The comment above nf_hook_egress() in include/linux/netfilter_netdev.h
> says:
>
> * The nf_skip_egress flag controls whether netfilter is applied on egress.
> * It is updated by __netif_receive_skb_core() and __dev_queue_xmit() when the
>
> Neither comment mentions the tc egress meaning. The second one also leaves
> out ifb_ri_tasklet() and netkit_prep_forward(), which write the bit too.
> The new skb_at_tc_egress() has no comment about reusing this bit or about
> the precondition it depends on.
>
> Could these comments be updated along with this change?
>
> > #ifdef CONFIG_SKB_DECRYPTED
> > diff --git a/include/net/sch_generic.h b/include/net/sch_generic.h
> > index f35bd06a6bad..1acaadb3e2cf 100644
> > --- a/include/net/sch_generic.h
> > +++ b/include/net/sch_generic.h
> > @@ -810,6 +810,15 @@ static inline bool skb_at_tc_ingress(const struct sk_buff *skb)
> > #endif
> > }
> >
> > +static inline bool skb_at_tc_egress(const struct sk_buff *skb)
> > +{
> > +#ifdef CONFIG_NET_EGRESS
> > + return skb->nf_skip_egress && !skb_at_tc_ingress(skb);
>
> [Severity: Medium]
> Is nf_skip_egress only ever set around sch_handle_egress()? The commit
> message says:
>
> "the marker is set only around the single sch_handle_egress() call and
> ingress path is guarded by tc_at_ingress"
>
> Some other writers seem to leave the bit set on skbs that later reach
> __dev_queue_xmit().
>
> __netif_receive_skb_core() sets the bit before sch_handle_ingress(), but
> clears it only on the path where the skb survives:
>
> nf_skip_egress(skb, true);
> skb = sch_handle_ingress(skb, &pt_prev, &ret, orig_dev,
> &another);
> ...
> if (!skb)
> goto out;
>
> nf_skip_egress(skb, false);
>
> An ingress mirred redirect consumes the skb and passes it to
> dev_queue_xmit() with the bit still set. A mirror does the same with a
> clone, which copies the bit through the headers group.
>
> ifb_ri_tasklet() sets the bit and then calls dev_queue_xmit(). On the
> from_ingress path it reinjects the skb with tc_skip_classify set, so
> __netif_receive_skb_core() never resets the bit:
>
> skb->tc_skip_classify = 1;
> ...
> nf_skip_egress(skb, true);
>
> netkit_prep_forward() also sets the bit. After that, netkit_xmit() calls
> skb_do_redirect(), which can reach dev_queue_xmit().
>
> In __dev_queue_xmit(), tcx_set_ingress(skb, false) always runs. The only
> reset of nf_skip_egress is inside the egress_needed_key static branch.
>
> Suppose only an ingress clsact is installed, for example one doing
> "mirred egress redirect dev qc". The skb then reaches qc's root qdisc
> enqueue with nf_skip_egress=1 and tc_at_ingress=0. A skbedit queue_mapping
> action in a prio, multiq or htb classifier passes skb_at_tc_egress() and
> calls netdev_xmit_skip_txqueue(true). The v4 changelog says this kind of
> arming from a transmit-qdisc classifier no longer happens.
>
> Nothing seems to consume this stale value today. Every reader inside the
> static branch saves, clears and restores the flag. The one exception is
> the static key race described further down.
>
> There is also a config effect. nf_skip_egress() is now built under
> CONFIG_NET_EGRESS instead of CONFIG_NETFILTER_SKIP_EGRESS. With
> CONFIG_NET_EGRESS=y and CONFIG_NETFILTER_EGRESS=n, the ingress, ifb and
> netkit calls used to compile to nothing. Now they write a bit that
> skb_at_tc_egress() reads.
>
> Should the reasoning in the commit message be revisited? Otherwise the
> planned rename to skb->skip_egress would carry the same assumption
> forward.
>
> > +#else
> > + return false;
> > +#endif
> > +}
>
> [ ... ]
>
> > diff --git a/net/core/dev.c b/net/core/dev.c
> > index f660fccfc0db..8638c994bbcb 100644
> > --- a/net/core/dev.c
> > +++ b/net/core/dev.c
> > @@ -4404,9 +4404,14 @@ EXPORT_SYMBOL(dev_loopback_xmit);
> > static struct netdev_queue *
> > netdev_tx_queue_mapping(struct net_device *dev, struct sk_buff *skb)
> > {
> > - int qm = skb_get_queue_mapping(skb);
> > + int queue = skb_get_queue_mapping(skb);
> > + int capped;
> >
> > - return netdev_get_tx_queue(dev, netdev_cap_txqueue(dev, qm));
> > + capped = netdev_cap_txqueue(dev, queue);
> > + if (unlikely(capped != queue))
> > + skb_set_queue_mapping(skb, capped);
> > +
> > + return netdev_get_tx_queue(dev, capped);
> > }
>
> [Severity: High]
> This is a pre-existing issue and was not introduced by this patch. The
> commit message leaves it for a separate followup. Can a cls_bpf program on
> a transmit qdisc still index past the end of dev->_tx[]?
>
> tc_cls_act_is_valid_access() allows writes to __sk_buff queue_mapping.
> bpf_convert_ctx_access() only drops stores of NO_QUEUE_MAPPING or larger:
>
> net/core/filter.c:bpf_convert_ctx_access() {
> ...
> if (BPF_CLASS(si->code) == BPF_ST && si->imm >= NO_QUEUE_MAPPING) {
> *insn++ = BPF_JMP_A(0); /* noop */
> break;
> }
>
> if (BPF_CLASS(si->code) == BPF_STX)
> *insn++ = BPF_JMP_IMM(BPF_JGE, si->src_reg, NO_QUEUE_MAPPING, 1);
> *insn++ = BPF_EMIT_STORE(BPF_H, si, offset);
> ...
> }
>
> The stored value is never compared with dev->num_tx_queues or
> real_num_tx_queues. A program attached under a prio or htb root runs
> inside q->enqueue(). That is after netdev_core_pick_tx() or
> netdev_tx_queue_mapping() has already capped the mapping, so the program
> can store any value up to 0xfffe.
>
> On dequeue:
>
> qdisc_restart()
> txq = skb_get_tx_queue(dev, skb)
> netdev_get_tx_queue()
> DEBUG_NET_WARN_ON_ONCE(index >= dev->num_tx_queues);
> return &dev->_tx[index];
> sch_direct_xmit(skb, q, dev, txq, ...)
> HARD_TX_LOCK(dev, txq, cpu)
>
> The gso_skb and skb_bad_txq requeue paths in dequeue_skb() also call
> skb_get_tx_queue().
>
> That would take txq->_xmit_lock, write xmit_lock_owner and read txq->state
> in memory past the end of the dev->_tx allocation. Loading the program
> needs CAP_BPF and CAP_NET_ADMIN, or a delegated BPF token. Even so, the
> program has passed the verifier and should not be able to corrupt kernel
> memory.
>
> [ ... ]
>
> > @@ -4848,21 +4861,25 @@ int __dev_queue_xmit(struct sk_buff *skb, struct net_device *sb_dev)
> > tcx_set_ingress(skb, false);
> > #ifdef CONFIG_NET_EGRESS
> > if (static_branch_unlikely(&egress_needed_key)) {
> > + bool skip_txq;
> > +
> > if (nf_hook_egress_active()) {
> > skb = nf_hook_egress(skb, &rc, dev);
> > if (!skb)
> > goto out;
> > }
> >
> > - netdev_xmit_skip_txqueue(false);
> > + skip_txq = netdev_xmit_skip_txqueue(false);
> >
> > nf_skip_egress(skb, true);
> > skb = sch_handle_egress(skb, &rc, dev);
> > - if (!skb)
> > + if (!skb) {
> > + netdev_xmit_skip_txqueue(skip_txq);
> > goto out;
> > + }
> > nf_skip_egress(skb, false);
> >
> > - if (netdev_xmit_txqueue_skipped())
> > + if (netdev_xmit_skip_txqueue(skip_txq))
> > txq = netdev_tx_queue_mapping(dev, skb);
> > }
> > #endif
>
> [Severity: Medium]
> Does the flag stay confined to this frame if egress_needed_key is turned
> off while an outer frame is inside sch_handle_egress()?
>
> Each nested __dev_queue_xmit() reads the static key again.
> net_dec_egress_queue() is a plain static_branch_dec() and does not wait
> for RCU-bh readers that are already inside the branch. Consider:
>
> CPU1 CPU2
> __dev_queue_xmit(qa)
> static key is on
> skip_txq = netdev_xmit_skip_txqueue(false)
> nf_skip_egress(skb, true)
> sch_handle_egress()
> clsact_destroy()
> net_dec_egress_queue()
> static_branch_dec()
> tcf_mirred_forward()
> (mirror clone has nf_skip_egress=1)
> dev_queue_xmit(qc)
> static key is off: no save/restore,
> no nf_skip_egress reset
> q->enqueue() -> prio classifier
> tcf_skbedit_act()
> skb_at_tc_egress() is true
> netdev_xmit_skip_txqueue(true)
> nf_skip_egress(skb, false)
> netdev_xmit_skip_txqueue(skip_txq) returns true
> txq = netdev_tx_queue_mapping(dev, skb)
>
> In this sequence, qa's surviving skb picks its tx queue from its own
> queue_mapping, even though it never went through skbedit. For a forwarded
> packet that value is the ingress rx_queue + 1.
>
> The capped write-back in netdev_tx_queue_mapping() still prevents the
> taprio out-of-bounds access. The result would be a wrong but valid tx
> queue, and possibly the ratelimited "selects TX queue" warning.
>
> The commit message says the flag is now owned by the xmit frame the
> egress hook runs in. Is this cross-frame leak meant to be covered as well?
>
> [ ... ]
>
> --
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/QDISC-9R8V.v4.20260928081529%40mojatatu.com
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked
2026-09-28 12:46 [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked Jamal Hadi Salim
2026-09-28 12:55 ` Eric Dumazet
2026-09-29 3:45 ` netdev-bot+sashiko
@ 2026-09-30 1:00 ` patchwork-bot+netdevbpf
2 siblings, 0 replies; 5+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-09-30 1:00 UTC (permalink / raw)
To: Jamal Hadi Salim
Cc: netdev, pablo, fw, phil, davem, edumazet, kuba, pabeni, jiri,
vinicius.gomes, horms, xiangxia.m.yue, bigeasy, clrkwllms,
rostedt, victor, zdi-disclosures, hybris, netfilter-devel,
coreteam, linux-rt-devel, stable, edumazet
Hello:
This patch was applied to netdev/net.git (main)
by Jakub Kicinski <kuba@kernel.org>:
On Mon, 28 Sep 2026 08:46:16 -0400 you wrote:
> skbedit can set skb->queue_mapping and raise the per-CPU skip_txqueue
> flag so __dev_queue_xmit() honours the mapping. __dev_queue_xmit()
> cleared the flag before sch_handle_egress() and only read it afterwards,
> so the flag was not confined to the xmit that set it: a nested xmit
> (mirred redirect or mirror, or a drop after skbedit) could set the flag
> and the outer xmit would consume it for an skb that never went through
> skbedit.
>
> [...]
Here is the summary with links:
- [net,v4] net: cap skb->queue_mapping when the tx queue is picked
https://git.kernel.org/netdev/net/c/ea4d4b5dddb5
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-09-30 1:00 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 12:46 [PATCH net v4] net: cap skb->queue_mapping when the tx queue is picked Jamal Hadi Salim
2026-09-28 12:55 ` Eric Dumazet
2026-09-29 3:45 ` netdev-bot+sashiko
2026-09-29 9:53 ` Jamal Hadi Salim
2026-09-30 1:00 ` patchwork-bot+netdevbpf
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox