Netdev List
 help / color / mirror / Atom feed
From: Cong Wang <xiyou.wangcong@gmail.com>
To: Will <willsroot@protonmail.com>
Cc: "netdev@vger.kernel.org" <netdev@vger.kernel.org>,
	Savy <savy@syst3mfailure.io>,
	jhs@mojatatu.com, jiri@resnulli.us
Subject: Re: [BUG] net/sched: Race Condition and Null Dereference in codel_change, pie_change, fq_pie_change, fq_codel_change, hhf_change
Date: Mon, 28 Apr 2025 12:53:49 -0700	[thread overview]
Message-ID: <aA/czQYEtPmMim0G@pop-os.localdomain> (raw)
In-Reply-To: <B2ZSzsBR9rUWlLkrgrMrCzqOGeSFxXIkYImvul6994v5tDSqykWo1UaWKRV-SNkNKJurgVzRcnPN07ZAVYykRaYhADyIwTxQ18OQfKDpILQ=@protonmail.com>

On Sun, Apr 27, 2025 at 09:26:43PM +0000, Will wrote:
> Hi Cong,
> 
> Thank you for the reply. On further analysis, we realized that you are correct - it is not a race condition between xxx_change and xxx_dequeue. The root cause is more complicated and actually relates to the parent tbf qdisc. The bug is still a race condition though.
> 
> __qdisc_dequeue_head() can still return null even if sch->q.qlen is non-zero because of qdisc_peek_dequeued, which is the vulnerable qdiscs' peek handler, and tbf_dequeue calls it (https://elixir.bootlin.com/linux/v6.15-rc3/source/net/sched/sch_tbf.c#L280). There, the inner qdisc dequeues content before, adds it back to gso_skb, and increments qlen (https://elixir.bootlin.com/linux/v6.15-rc3/source/include/net/sch_generic.h#L1133). A queue state consistency issue arises when tbf does not have enough tokens (https://elixir.bootlin.com/linux/v6.15-rc3/source/net/sched/sch_tbf.c#L302) for dequeuing. The qlen value will be fixed when sufficient tokens exist and the watchdog fires again. However, there is a window for the inner qdisc to encounter this inconsistency and thus hit the null dereference.
> 
> Savy made this diagram below to showcase the interactions to trigger the bug.
> 
> Packet 1 is sent:
> 
>     tbf_enqueue()
>         qdisc_enqueue()
>             codel_qdisc_enqueue() // Codel qlen is 0
>                 qdisc_enqueue_tail()
>                 // Packet 1 is added to the queue
>                 // Codel qlen = 1
> 
>     tbf_dequeue()
>         qdisc_peek_dequeued()
>             skb_peek(&sch->gso_skb) // sch->gso_skb is empty
>             codel_qdisc_dequeue() // Codel qlen is 1
>                 qdisc_dequeue_head()
>                 // Packet 1 is removed from the queue
>                 // Codel qlen = 0
>             __skb_queue_head(&sch->gso_skb, skb); // Packet 1 is added to gso_skb list
>             sch->q.qlen++ // Codel qlen = 1
>         qdisc_dequeue_peeked()
>             skb = __skb_dequeue(&sch->gso_skb) // Packet 1 is removed from the gso_skb list
>             sch->q.qlen-- // Codel qlen = 0
> 
> Packet 2 is sent:
> 
>     tbf_enqueue()
>         qdisc_enqueue()
>             codel_qdisc_enqueue() // Codel qlen is 0
>                 qdisc_enqueue_tail()
>                 // Packet 2 is added to the queue
>                 // Codel qlen = 1
> 
>     tbf_dequeue()
>         qdisc_peek_dequeued()
>             skb_peek(&sch->gso_skb) // sch->gso_skb is empty
>             codel_qdisc_dequeue() // Codel qlen is 1
>                 qdisc_dequeue_head()
>                 // Packet 2 is removed from the queue
>                 // Codel qlen = 0
>             __skb_queue_head(&sch->gso_skb, skb); // Packet 2 is added to gso_skb list
>             sch->q.qlen++ // Codel qlen = 1
> 
>         // TBF runs out of tokens and reschedules itself for later
>         qdisc_watchdog_schedule_ns()
> 
> 
> Notice here how codel is left in an "inconsistent" state, as sch->q.qlen > 0, but there are no packets left in the codel queue (sch->q.head is NULL)
> 
> At this point, codel_change() can be used to update the limit to 0. However, even if  sch->q.qlen > 0, there are no packets in the queue, so __qdisc_dequeue_head() returns NULL and the null-ptr-deref occurs.
> 

Excellent analysis!

Do you mind testing the following patch?

Note:

1) We can't just test NULL, because otherwise we would leak the skb's
in gso_skb list. 

2) I am totally aware that _maybe_ there are some other cases need the
same fix, but I want to be conservative here since this will be
targeting for -stable. It is why I intentionally keep my patch minimum.

Thanks!

--------------->

diff --git a/include/net/sch_generic.h b/include/net/sch_generic.h
index d48c657191cd..5a4840678ce5 100644
--- a/include/net/sch_generic.h
+++ b/include/net/sch_generic.h
@@ -1031,6 +1031,25 @@ static inline struct sk_buff *__qdisc_dequeue_head(struct qdisc_skb_head *qh)
 	return skb;
 }
 
+static inline struct sk_buff *qdisc_dequeue_internal(struct Qdisc *sch)
+{
+	struct sk_buff *skb;
+
+	skb = __skb_dequeue(&sch->gso_skb);
+	if (skb != NULL) {
+		if (qdisc_is_percpu_stats(sch)) {
+			qdisc_qstats_cpu_backlog_dec(sch, skb);
+			qdisc_qstats_cpu_qlen_dec(sch);
+		} else {
+			qdisc_qstats_backlog_dec(sch, skb);
+			sch->q.qlen--;
+		}
+		return skb;
+	}
+	skb = __qdisc_dequeue_head(&sch->q);
+	return skb;
+}
+
 static inline struct sk_buff *qdisc_dequeue_head(struct Qdisc *sch)
 {
 	struct sk_buff *skb = __qdisc_dequeue_head(&sch->q);
diff --git a/net/sched/sch_codel.c b/net/sched/sch_codel.c
index 12dd71139da3..e1bf4919d258 100644
--- a/net/sched/sch_codel.c
+++ b/net/sched/sch_codel.c
@@ -144,10 +144,9 @@ static int codel_change(struct Qdisc *sch, struct nlattr *opt,
 
 	qlen = sch->q.qlen;
 	while (sch->q.qlen > sch->limit) {
-		struct sk_buff *skb = __qdisc_dequeue_head(&sch->q);
+		struct sk_buff *skb = qdisc_dequeue_internal(sch);
 
 		dropped += qdisc_pkt_len(skb);
-		qdisc_qstats_backlog_dec(sch, skb);
 		rtnl_qdisc_drop(skb, sch);
 	}
 	qdisc_tree_reduce_backlog(sch, qlen - sch->q.qlen, dropped);
diff --git a/net/sched/sch_pie.c b/net/sched/sch_pie.c
index 3771d000b30d..b6ed94976e69 100644
--- a/net/sched/sch_pie.c
+++ b/net/sched/sch_pie.c
@@ -195,10 +195,9 @@ static int pie_change(struct Qdisc *sch, struct nlattr *opt,
 	/* Drop excess packets if new limit is lower */
 	qlen = sch->q.qlen;
 	while (sch->q.qlen > sch->limit) {
-		struct sk_buff *skb = __qdisc_dequeue_head(&sch->q);
+		struct sk_buff *skb = qdisc_dequeue_internal(sch);
 
 		dropped += qdisc_pkt_len(skb);
-		qdisc_qstats_backlog_dec(sch, skb);
 		rtnl_qdisc_drop(skb, sch);
 	}
 	qdisc_tree_reduce_backlog(sch, qlen - sch->q.qlen, dropped);

  reply	other threads:[~2025-04-28 19:53 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2025-04-25 14:14 [BUG] net/sched: Race Condition and Null Dereference in codel_change, pie_change, fq_pie_change, fq_codel_change, hhf_change Will
2025-04-26 22:56 ` Cong Wang
2025-04-27 21:26   ` Will
2025-04-28 19:53     ` Cong Wang [this message]
2025-04-29 13:41       ` Savy
2025-05-04 15:35         ` Will
2025-05-05 19:44         ` Cong Wang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aA/czQYEtPmMim0G@pop-os.localdomain \
    --to=xiyou.wangcong@gmail.com \
    --cc=jhs@mojatatu.com \
    --cc=jiri@resnulli.us \
    --cc=netdev@vger.kernel.org \
    --cc=savy@syst3mfailure.io \
    --cc=willsroot@protonmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox