From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm2-f12.google.com (mail-wm2-f12.google.com [74.125.225.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DD3B3502797 for ; Fri, 18 Sep 2026 15:30:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.140 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789745405; cv=none; b=imjhQD24Eg8WfQnHFub2Iz8455qUYV51HLW/O+UHWhhMKC+Oz4WYvov/GodJ1muxLqA1K9dTVn8GtzhtEO+eox2orRATL7TdUMbK0WdY3burA37Gb37FfDmaBAenXgeDU95w4cVRxuNPYrj4/rJapG9EEZwvyjxnbAaVRvkeaWc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789745405; c=relaxed/simple; bh=vPdVL81FDVg0TO4uofdiBt+hhnNdQuJHDjmcN1sb5bs=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=s6uKMhqy4efvXCo33zni3EBcZA31Qi8eh1XvMY3wazwOBMuvVkG6yIvb7OMQkkQxgBhc8JK1AojBE2Ot1IRS56xn2MAQgrmdvdNH5ojirj+WRi8qJiqAv51w1r63Tp1KEX00r2y6U5TB5W5CiLpxxpM/b8w4x6HodjXr/N8idWE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=blackwall.org; spf=none smtp.mailfrom=blackwall.org; dkim=pass (2048-bit key) header.d=blackwall.org header.i=@blackwall.org header.b=a9S9S2Se; arc=none smtp.client-ip=74.125.225.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=blackwall.org Authentication-Results: smtp.subspace.kernel.org; spf=none smtp.mailfrom=blackwall.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=blackwall.org header.i=@blackwall.org header.b="a9S9S2Se" Received: by mail-wm2-f12.google.com with SMTP id 5b1f17b1804b1-49b91369d18so7391355e9.0 for ; Fri, 18 Sep 2026 08:30:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=blackwall.org; s=google; t=1789745402; x=1790350202; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=amEgsq9/x54qq2eWnuO6h/vQWTMXR2GmqJyXo88O4/k=; b=a9S9S2Se/cQm4aiZWY1KGLrMTgJp6Nq5KkfOWoXFUnfVxjZIL3b5qguMltKNWIlz1E Bu2HP8TapRRr7guQqnAycF+IqsjXuvyqp3OESKfiR5376K/OCscIryxNaWfEqC4fbBPa im4KpDGsKobT+QTHLtplIIAYBkxdiDH/tWtVQ3gx+pcpizMKOimrvJcXxTA/cTw1kc8j Qm9ZSE2bfl0Crgw3H8ta2Rex6hncRn+cNdd2NlxQmTfvPkMZLcglwDAkFL9sfcyn+a8T NREXMVx3yyhFoH5RuPpl9TBlvmmyn1s88Y8BvhguTlGoLQOIHI6pvNEcajQQ4S9zv3pp +cqg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789745402; x=1790350202; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=amEgsq9/x54qq2eWnuO6h/vQWTMXR2GmqJyXo88O4/k=; b=BwpbTvDy3uwJCXpgY9zZmkqFBKUQl6zF2l9+B/ArU82XzHnBEaVjxyAtVujS5zEjX1 C66sCmTyxEcgzzwtLrQff6UinkhYxKr/zymksx3e6Xp8kPtWteP+kycTY5EVgTatj9+a 5JzkojxZw/jR1z+oUtvyv7GqLYNrAf62+FLFRITZng9lwokxqHpl6Z6g5BAcNuUi1j/t JZyzXCSZ4EOHHhHT5tMSpJ7LfVzko1WaiOWoVlW8HwFiP3PxnTF0r5ofHWe0xVPvm/Po fVn3HM+9SFuMYQRIVvbaQLKDJm0U+k8jMafJdsVG3cX58yTpRCmtPTgxWkrfqFnmkqY0 wQJA== X-Gm-Message-State: AFuF++l848bH+Scmm40H0k+AmX8Tw6JMoQOY6O22DLpg06NCUw/skLwi cxw4XSaMrAvk+XBJrbu5+CSeiLA3Km3bPeTSTYkVLl6ww+evMRxXLhOc783G63AfvQlp0yWcf2W c4Wb7 X-Gm-Gg: AYBFou3AqW5qg3Kxy5SxfJtwmXufdq1z7luOH3tb0ifIKANAb5gtbGnjl7QxUzqGKwT 1FfWk4yy3Xp4+3nklithpffXS2FDLHVL83StkrS4NDqViscs15FO4GTBTJzj5dl2kCeJ5BDZGHq lg9tnuF9sd4xnqZAQimcmvj5euHvhW+Q1TalflhgCJs6kFuVzJvFWBTvAGQl/wT4lgpANGXgvhY Qi/9xXpDEKljZrtJOfPFZBdIthEwTivMzqNwoWbIglEFDaggbq4WMyr9vhr3Kg8eFR/x3V4bG9+ e5WskVE8bCbWv1M+3CeXN+J14PZzlRvQVePrMAXBVt7cpYCjfKTF0Qh+JvU/lHzInrF6E/WQiFW WGJEZiSFfYc3+TJhxuEKCmZEK+RPQnNLHicj6TGTAP/5fbSjOZXp4XVw8fHF4U7o/iqGPrVcEOf gX65s9r/8IgU8vKRcGBZ5uwXm1YjC5XYOPezynR0zGqWmOutKNVgEA6N81pG5u1BIzIU2OXFo9T 1Fs74hA6/tgRTxEkNS4Xw== X-Received: by 2002:a05:600c:6217:b0:49e:8418:389d with SMTP id 5b1f17b1804b1-49fc571423dmr37735005e9.9.1789745401448; Fri, 18 Sep 2026 08:30:01 -0700 (PDT) Received: from localhost (78-154-14-127.ip.btc-net.bg. [78.154.14.127]) by smtp.gmail.com with UTF8SMTPSA id 5b1f17b1804b1-49fc6ca43edsm63465895e9.0.2026.09.18.08.30.00 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 18 Sep 2026 08:30:00 -0700 (PDT) From: Nikolay Aleksandrov To: netdev@vger.kernel.org Cc: idosch@nvidia.com, davem@davemloft.net, edumazet@google.com, kuba@kernel.org, pabeni@redhat.com, horms@kernel.org, bridge@lists.linux.dev, Nikolay Aleksandrov Subject: [PATCH net-next 6/9] net: bridge: vlan: use an RCU array for large flood sets Date: Fri, 18 Sep 2026 18:29:47 +0300 Message-ID: <20260918152950.1938259-7-razor@blackwall.org> X-Mailer: git-send-email 2.47.3 In-Reply-To: <20260918152950.1938259-1-razor@blackwall.org> References: <20260918152950.1938259-1-razor@blackwall.org> Precedence: bulk X-Mailing-List: netdev@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Walking the master VLAN's port-VLAN list avoids considering ports outside the VLAN but its pointer chasing becomes more expensive than array when many ports participate. Add an rcu array of port-VLAN pointers for larger flood sets. Rebuild and publish the array under RTNL when VLAN membership changes. Continue using the list for small flood sets and as a fallback if the array allocation fails. Signed-off-by: Nikolay Aleksandrov --- From local sashiko run: [Severity: Medium] Could these rebuilds be amortized or batched for large VLAN memberships? In net/bridge/br_vlan.c, br_vlan_rebuild_port_array() allocates and copies the complete masterv->port_vlist whenever count exceeds BR_VLAN_PORT_ARRAY_THRESHOLD. It then defers freeing the previous complete array through kvfree_rcu(), so sustained updates can retain several full-array generations until their RCU grace periods finish. Every successful individual port-VLAN addition calls this helper from __vlan_add(), while every individual deletion calls it from __vlan_del(). Growing a flood set from nine entries to N therefore copies 9 + 10 + ... + N pointers, and shrinking it performs the same quadratic work while RTNL is held. Can this cause a control-plane CPU and transient-memory regression during large incremental bridge VLAN updates? The later patches in the series retain this rebuild-on-add/delete path in the final series state. Nik: Yes, that is well understood but it is control path and I have tested sustained 2k / sec VLAN add/delete with 64 VLAN ports in each VLAN. If it ever becomes a problem we can optimize it, I think for the initial implementation would be best to keep it simple. net/bridge/br_forward.c | 39 ++++++++++++++++++++++++++++++--------- net/bridge/br_private.h | 13 +++++++++++++ net/bridge/br_vlan.c | 37 ++++++++++++++++++++++++++++++++++++- 3 files changed, 79 insertions(+), 10 deletions(-) diff --git a/net/bridge/br_forward.c b/net/bridge/br_forward.c index 251d61e7c312..e8f30f2df1ed 100644 --- a/net/bridge/br_forward.c +++ b/net/bridge/br_forward.c @@ -261,6 +261,35 @@ static void br_flood_port(struct net_bridge_port **prev, *prev = maybe_deliver(*prev, p, skb, local_orig); } +static void br_flood_vlan(struct net_bridge_port **prev, + struct net_bridge_vlan *v, struct sk_buff *skb, + enum br_pkt_type pkt_type, bool local_orig) +{ + struct net_bridge_vlan_port_array *array; + struct net_bridge_vlan *masterv, *pv; + + masterv = br_vlan_is_master(v) ? v : v->brvlan; + array = rcu_dereference(masterv->port_array); + if (array) { + unsigned int i; + + for (i = 0; i < array->count; i++) { + pv = array->vlans[i]; + br_flood_port(prev, pv->port, skb, pkt_type, + local_orig, v->vid); + if (IS_ERR(*prev)) + break; + } + } else { + list_for_each_entry_rcu(pv, &masterv->port_vlist, port_vlist) { + br_flood_port(prev, pv->port, skb, pkt_type, + local_orig, v->vid); + if (IS_ERR(*prev)) + break; + } + } +} + /* called under rcu_read_lock */ void br_flood(struct net_bridge *br, struct net_bridge_vlan *v, struct sk_buff *skb, enum br_pkt_type pkt_type, @@ -271,15 +300,7 @@ void br_flood(struct net_bridge *br, struct net_bridge_vlan *v, br_tc_skb_miss_set(skb, pkt_type != BR_PKT_BROADCAST); if (v) { - struct net_bridge_vlan *masterv, *pv; - - masterv = br_vlan_is_master(v) ? v : v->brvlan; - list_for_each_entry_rcu(pv, &masterv->port_vlist, port_vlist) { - br_flood_port(&prev, pv->port, skb, pkt_type, - local_orig, v->vid); - if (IS_ERR(prev)) - break; - } + br_flood_vlan(&prev, v, skb, pkt_type, local_orig); } else { struct net_bridge_port *p; diff --git a/net/bridge/br_private.h b/net/bridge/br_private.h index 239cf58d2268..a33da6e9765f 100644 --- a/net/bridge/br_private.h +++ b/net/bridge/br_private.h @@ -190,6 +190,17 @@ enum { BR_VLFLAG_NEIGH_FORWARD_GRAT_ENABLED = BIT(6), }; +/* start publishing arrays when there're > BR_VLAN_PORT_ARRAY_THRESHOLD + * port-VLANs + */ +#define BR_VLAN_PORT_ARRAY_THRESHOLD 8 + +struct net_bridge_vlan_port_array { + struct rcu_head rcu; + unsigned int count; + struct net_bridge_vlan *vlans[]; +}; + /** * struct net_bridge_vlan - per-vlan entry * @@ -210,6 +221,7 @@ enum { * @port_mcast_ctx: if MASTER flag unset, this is the per-port/vlan multicast * context * @msti: if MASTER flag set, this holds the VLANs MST instance + * @port_array: if MASTER flag set, this is the port-VLAN array * @port_vlist: if MASTER flag set, this is the port-VLAN list * @vlist: sorted list of VLAN entries * @rcu: used for entry destruction @@ -245,6 +257,7 @@ struct net_bridge_vlan { u16 msti; + struct net_bridge_vlan_port_array __rcu *port_array; struct list_head port_vlist; struct list_head vlist; diff --git a/net/bridge/br_vlan.c b/net/bridge/br_vlan.c index 34d1df59d190..d750581df64d 100644 --- a/net/bridge/br_vlan.c +++ b/net/bridge/br_vlan.c @@ -258,6 +258,36 @@ static void br_vlan_init_state(struct net_bridge_vlan *v) v->msti = 0; } +static unsigned int br_vlan_num_ports(const struct net_bridge_vlan *masterv) +{ + return refcount_read(&masterv->refcnt) - br_vlan_is_brentry(masterv); +} + +static void br_vlan_rebuild_port_array(struct net_bridge_vlan *masterv, + unsigned int count) +{ + struct net_bridge_vlan_port_array *array = NULL, *old; + unsigned int i = 0; + + WARN_ON(!br_vlan_is_master(masterv)); + + if (count > BR_VLAN_PORT_ARRAY_THRESHOLD) + array = kvmalloc(struct_size(array, vlans, count), GFP_KERNEL); + + if (array) { + struct net_bridge_vlan *pv; + + array->count = count; + list_for_each_entry(pv, &masterv->port_vlist, port_vlist) + array->vlans[i++] = pv; + } + + old = rtnl_dereference(masterv->port_array); + rcu_assign_pointer(masterv->port_array, array); + if (old) + kvfree_rcu(old, rcu); +} + /* This is the shared VLAN add function which works for both ports and bridge * devices. There are four possible calls to this function in terms of the * vlan entry type: @@ -368,8 +398,10 @@ static int __vlan_add(struct net_bridge_vlan *v, u16 flags, __vlan_flags_commit(v, flags); br_multicast_toggle_one_vlan(v, true); - if (p) + if (p) { + br_vlan_rebuild_port_array(masterv, br_vlan_num_ports(masterv)); nbp_vlan_set_vlan_dev_state(p, v->vid); + } out: return err; @@ -438,6 +470,9 @@ static void __vlan_del(struct net_bridge_vlan *v) rhashtable_remove_fast(&vg->vlan_hash, &v->vnode, br_vlan_rht_params); __vlan_del_list(v); + /* -1 because br_vlan_put_master() is called later */ + br_vlan_rebuild_port_array(masterv, + br_vlan_num_ports(masterv) - 1); nbp_vlan_set_vlan_dev_state(p, v->vid); br_multicast_toggle_one_vlan(v, false); br_multicast_port_ctx_deinit(&v->port_mcast_ctx); -- 2.47.3