Netdev List
 help / color / mirror / Atom feed
From: Jamal Hadi Salim <jhs@mojatatu.com>
To: netdev@vger.kernel.org
Cc: "Jamal Hadi Salim" <jhs@mojatatu.com>,
	"Toke Høiland-Jørgensen" <toke@toke.dk>,
	"Jiri Pirko" <jiri@resnulli.us>,
	"David S . Miller" <davem@davemloft.net>,
	"Eric Dumazet" <edumazet@kernel.org>,
	"Jakub Kicinski" <kuba@kernel.org>,
	"Paolo Abeni" <pabeni@redhat.com>,
	"Simon Horman" <horms@kernel.org>,
	cake@lists.bufferbloat.net,
	"Victor Nogueira" <victor@mojatatu.com>,
	Sashiko <sashiko-bot@kernel.org>
Subject: [PATCH net-next 1/4] net/sched/sch_cake: serialize reconfiguration with the datapath
Date: Thu,  8 Oct 2026 03:47:09 -0400	[thread overview]
Message-ID: <QDISC-5RTM.v1.20261007130704-2@mojatatu.com> (raw)
In-Reply-To: <QDISC-5RTM.v1.20261007130704@mojatatu.com>

This is a follow-up to commit 7cbfb180945c ("net/sched: sch_cake: fix
autorate reconfiguration throttling").  That change only stored
last_reconfig_time; it did not address the concurrency between the
autorate rate update in cake_enqueue() and the netlink reconfiguration
path, which the original review flagged.

cake_change() -> cake_config_change() updates the fields of the shared
struct cake_sched_config under RTNL but NOT under the qdisc lock: it
only took sch_tree_lock() later, around cake_reconfigure().  The datapath
in turn serializes cake_enqueue() with itself under the qdisc lock, and
its autorate-ingress path stores the estimated rate into
q->config->rate_bps.  Two writers therefore update the same rate with no
common lock:

  CPU0: cake_enqueue() sees CAKE_FLAG_AUTORATE_INGRESS, computes an
        estimate, and is about to store it.
  CPU1: 'tc qdisc change ... bandwidth X' clears autorate in its local
        rate_flags, writes rate_bps = X, publishes the cleared flag, and
        waits in sch_tree_lock().
  CPU0: stores the estimate, calls cake_reconfigure(), and unlocks.
  CPU1: reconfigures from the estimate and returns.

The qdisc is left with autorate disabled but with the estimate installed
instead of X.  On 32-bit machines the unlocked 64-bit rate_bps store is
also not atomic, so a reader can observe a torn rate.  iproute2 reaches
this because parsing 'bandwidth X' sets autorate = 0 and emits both
TCA_CAKE_BASE_RATE64 and TCA_CAKE_AUTORATE.

Commit the parsed configuration and reconfigure while holding the same
qdisc lock that serializes the datapath, so the netlink writer and the
autorate writer can no longer interleave.  Make rate_bps an atomic64_t
so a 64-bit rate update is atomic on 32-bit machines as well, and keep
the READ_ONCE()/WRITE_ONCE() annotations for the remaining lockless
readers (cake_config_dump(), which runs without the lock, and the
cake_mq shared config, which readers observe under different child
locks).

Conditions to recreate the bug: build with CONFIG_KCSAN=y (the race is
otherwise not observable on 64-bit), put cake in autorate-ingress on a
device, drive bursty traffic with gaps so the 250ms autorate
reconfiguration fires, and concurrently loop 'tc qdisc change dev ...
cake bandwidth <N>kbit autorate-ingress'.  The lost update of the
installed rate is reproducible without a sanitizer by running bursty
senders while repeatedly changing an autorate qdisc to a fixed
'bandwidth X' and reading the installed rate back.

Reported-by: Sashiko (gemini) <sashiko-bot@kernel.org>
Link: https://sashiko.dev/#/patchset/20260816012109.2865223-1-ooonea@gmail.com
Link: https://lore.kernel.org/netdev/20260816012109.2865223-1-ooonea@gmail.com/
Reviewed-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
---
 net/sched/sch_cake.c | 42 +++++++++++++++++++++++++-----------------
 1 file changed, 25 insertions(+), 17 deletions(-)

diff --git a/net/sched/sch_cake.c b/net/sched/sch_cake.c
index dc93267029e7..1c29695dc928 100644
--- a/net/sched/sch_cake.c
+++ b/net/sched/sch_cake.c
@@ -199,7 +199,7 @@ struct cake_tin_data {
 }; /* number of tins is small, so size of this struct doesn't matter much */
 
 struct cake_sched_config {
-	u64		rate_bps;
+	atomic64_t	rate_bps;
 	u64		interval;
 	u64		target;
 	u64		sync_time;
@@ -1906,7 +1906,8 @@ static s32 cake_enqueue(struct sk_buff *skb, struct Qdisc *sch,
 			if (ktime_after(now,
 					ktime_add_ms(q->last_reconfig_time,
 						     250))) {
-				q->config->rate_bps = (q->avg_peak_bandwidth * 15) >> 4;
+				atomic64_set(&q->config->rate_bps,
+					     (q->avg_peak_bandwidth * 15) >> 4);
 				q->last_reconfig_time = now;
 				cake_reconfigure(sch);
 			}
@@ -2021,7 +2022,7 @@ static struct sk_buff *cake_dequeue(struct Qdisc *sch)
 	    now - q->last_checked_active >= q->config->sync_time) {
 		struct net_device *dev = qdisc_dev(sch);
 		struct cake_sched_data *other_priv;
-		u64 new_rate = q->config->rate_bps;
+		u64 new_rate = atomic64_read(&q->config->rate_bps);
 		u64 other_qlen, other_last_active;
 		struct Qdisc *other_sch;
 		u32 num_active_qs = 1;
@@ -2042,7 +2043,7 @@ static struct sk_buff *cake_dequeue(struct Qdisc *sch)
 		}
 
 		if (num_active_qs > 1)
-			new_rate = div64_u64(q->config->rate_bps, num_active_qs);
+			new_rate = div64_u64(atomic64_read(&q->config->rate_bps), num_active_qs);
 
 		cake_configure_rates(sch, new_rate, true);
 		q->last_checked_active = now;
@@ -2626,12 +2627,12 @@ static void cake_reconfigure(struct Qdisc *sch)
 	struct cake_sched_config *q = qd->config;
 	u32 buffer_limit;
 
-	cake_configure_rates(sch, qd->config->rate_bps, false);
+	cake_configure_rates(sch, atomic64_read(&qd->config->rate_bps), false);
 
 	if (q->buffer_config_limit) {
 		buffer_limit = q->buffer_config_limit;
-	} else if (q->rate_bps) {
-		u64 t = q->rate_bps * q->interval;
+	} else if (atomic64_read(&q->rate_bps)) {
+		u64 t = atomic64_read(&q->rate_bps) * q->interval;
 
 		do_div(t, USEC_PER_SEC / 4);
 		buffer_limit = max_t(u32, t, 4U << 20);
@@ -2686,8 +2687,8 @@ static int cake_config_change(struct cake_sched_config *q, struct nlattr *opt,
 	}
 
 	if (tb[TCA_CAKE_BASE_RATE64])
-		WRITE_ONCE(q->rate_bps,
-			   nla_get_u64(tb[TCA_CAKE_BASE_RATE64]));
+		atomic64_set(&q->rate_bps,
+			     nla_get_u64(tb[TCA_CAKE_BASE_RATE64]));
 
 	if (tb[TCA_CAKE_DIFFSERV_MODE])
 		WRITE_ONCE(q->tin_mode,
@@ -2784,9 +2785,16 @@ static int cake_change(struct Qdisc *sch, struct nlattr *opt,
 		return -EOPNOTSUPP;
 	}
 
+	/* Once the qdisc is live, commit and reconfigure under the lock that
+	 * serializes the datapath, so a concurrent 'tc qdisc change' and the
+	 * autorate update cannot lose the configured rate.
+	 */
+	if (qd->tins)
+		sch_tree_lock(sch);
+
 	ret = cake_config_change(q, opt, extack, &overhead_changed);
 	if (ret)
-		return ret;
+		goto unlock;
 
 	if (overhead_changed) {
 		WRITE_ONCE(qd->max_netlen, 0);
@@ -2795,13 +2803,13 @@ static int cake_change(struct Qdisc *sch, struct nlattr *opt,
 		WRITE_ONCE(qd->min_adjlen, ~0);
 	}
 
-	if (qd->tins) {
-		sch_tree_lock(sch);
+	if (qd->tins)
 		cake_reconfigure(sch);
+unlock:
+	if (qd->tins)
 		sch_tree_unlock(sch);
-	}
 
-	return 0;
+	return ret;
 }
 
 static void cake_destroy(struct Qdisc *sch)
@@ -2818,7 +2826,7 @@ static void cake_config_init(struct cake_sched_config *q, bool is_shared)
 	q->tin_mode = CAKE_DIFFSERV_DIFFSERV3;
 	q->flow_mode  = CAKE_FLOW_TRIPLE;
 
-	q->rate_bps = 0; /* unlimited by default */
+	atomic64_set(&q->rate_bps, 0); /* unlimited by default */
 
 	q->interval = 100000; /* 100ms default */
 	q->target   =   5000; /* 5ms: codel RFC argues
@@ -2889,7 +2897,7 @@ static int cake_init(struct Qdisc *sch, struct nlattr *opt,
 	}
 
 	cake_reconfigure(sch);
-	qd->avg_peak_bandwidth = q->rate_bps;
+	qd->avg_peak_bandwidth = atomic64_read(&q->rate_bps);
 	qd->min_netlen = ~0;
 	qd->min_adjlen = ~0;
 	qd->active_queues = 0;
@@ -2917,7 +2925,7 @@ static int cake_config_dump(struct cake_sched_config *q, struct sk_buff *skb)
 		goto nla_put_failure;
 
 	if (nla_put_u64_64bit(skb, TCA_CAKE_BASE_RATE64,
-			      READ_ONCE(q->rate_bps), TCA_CAKE_PAD))
+			      atomic64_read(&q->rate_bps), TCA_CAKE_PAD))
 		goto nla_put_failure;
 
 	flow_mode = READ_ONCE(q->flow_mode);
-- 
2.43.0


  reply	other threads:[~2026-10-08  7:47 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-08  7:47 [PATCH net-next 0/4] net/sched: sch_cake: fixes for accounting and reconfiguration Jamal Hadi Salim
2026-10-08  7:47 ` Jamal Hadi Salim [this message]
2026-10-08 10:06   ` [PATCH net-next 1/4] net/sched/sch_cake: serialize reconfiguration with the datapath Toke Høiland-Jørgensen
2026-10-08 11:53     ` Jamal Hadi Salim
2026-10-08 12:03       ` Toke Høiland-Jørgensen
2026-10-08  7:47 ` [PATCH net-next 2/4] net/sched/sch_cake: reduce parent backlog when clearing a tin Jamal Hadi Salim
2026-10-08 10:19   ` Toke Høiland-Jørgensen
2026-10-08  7:47 ` [PATCH net-next 3/4] net/sched/sch_cake: widen buffer_used to u64 Jamal Hadi Salim
2026-10-08 10:24   ` Toke Høiland-Jørgensen
2026-10-08  7:47 ` [PATCH net-next 4/4] net/sched/sch_cake: drop dead nla_nest_end() error check Jamal Hadi Salim
2026-10-08 10:24   ` Toke Høiland-Jørgensen

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=QDISC-5RTM.v1.20261007130704-2@mojatatu.com \
    --to=jhs@mojatatu.com \
    --cc=cake@lists.bufferbloat.net \
    --cc=davem@davemloft.net \
    --cc=edumazet@kernel.org \
    --cc=horms@kernel.org \
    --cc=jiri@resnulli.us \
    --cc=kuba@kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=sashiko-bot@kernel.org \
    --cc=toke@toke.dk \
    --cc=victor@mojatatu.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox