Netdev List
 help / color / mirror / Atom feed
* [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps
@ 2026-09-21 10:01 Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes Eric Dumazet
                   ` (7 more replies)
  0 siblings, 8 replies; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

This patch series converts struct vxlan_config to RCU protection, eliminates
redundant destination fields, and enables lockless RTNL-free link dumping
in vxlan_fill_info().

Changes in v5:
- Widen vxlan_multicast_leave() and vxlan_multicast_join() conditions in
  vxlan_changelink() to also run when VXLAN_F_VNIFILTER is set, so per-VNI
  multicast memberships are migrated across lower-device changes even when
  the device-default remote_ip is 0.0.0.0 (patches 1, 6, 7).
- Unwind already-updated per-VNI entries and the device-wide default FDB
  entry if vxlan_vnilist_update_group() fails partway through in
  vxlan_changelink() (patches 1, 7).
- Bump the vnifilter dump generation counter inside vxlan_vni_update_group()
  as soon as vninode->remote_ip is overwritten, so concurrent dumps still
  observe NLM_F_DUMP_INTR if a subsequent IGMP leave/join fails (patch 3).
- In vxlan_config_validate(), reject clearing IFLA_VXLAN_LINK when
  VXLAN_F_VNIFILTER has per-VNI multicast groups configured (patch 7).
- In vxlan_config_apply(), reset dev->needed_tailroom when detaching a
  lower device and use READ_ONCE()/WRITE_ONCE() for dev->mtu (patch 7).
- Link to v4: https://lore.kernel.org/netdev/20260915175501.391567-1-edumazet@google.com/

Changes in v4:
- Split out the default FDB update fix when the lower device changes into
  patch 1 with a Fixes: tag, and updated vxlan_vnilist_update_group() to
  refresh all VNIs (including those with explicit per-VNI multicast groups)
  so their FDB entries do not retain a stale remote_ifindex. Kept in this
  series rather than routing through net to avoid a cross-tree conflict
  with the default_dst removal in patch 7.
- Added patch 3 to signal interrupted RTM_GETTUNNEL dumps to user space via
  a per-netns atomic generation counter and nl_dump_check_consistent(), per
  Paolo's suggestion.
- Added an explicit VXLAN_DEV_F_MDB check in mlxsw_sp_nve_vxlan_can_offload()
  so moving VXLAN_F_MDB out of cfg->flags does not silently relax mlxsw's
  deny-by-default flag validation (patch 5).
- In vxlan_config_apply(), gated the MTU clamp on lowerdev_changed so that
  unrelated changelink operations do not shrink dev->mtu (patch 7).
- Documented in patch 7's changelog that cfg->remote_ifindex is now committed
  unconditionally and that IFLA_VXLAN_LINK=0 tears down the upper/lower
  adjacency.
- Link to v3: https://lore.kernel.org/netdev/20260911033945.164207-1-edumazet@google.com/

Changes in v3:
- Folded configuration conversion patches into a single patch to maintain
  clean bisectability and avoid intermediate mutable state.
- In vxlan_changelink(), pass lowerdev to vxlan_config_apply() to prevent
  truncating dev->needed_headroom and dev->needed_tailroom when lowerdev
  does not change.
- Cleaned up redundant NULL checks and unreachable error paths.

Eric Dumazet (8):
  vxlan: update default fdb entries when the lower device changes
  vxlan: vnifilter: use list_for_each_entry_rcu() in
    vxlan_vnifilter_dump_dev()
  vxlan: vnifilter: signal interrupted RTM_GETTUNNEL dumps
  vxlan: pass vxlan_config pointer to helper functions
  vxlan: move VXLAN_F_MDB to struct vxlan_dev flags
  vxlan: convert configuration to RCU protection
  vxlan: remove default_dst and use vxlan_config and lowerdev
  vxlan: no longer rely on RTNL in vxlan_fill_info()

 .../mellanox/mlx5/core/en/tc_tun_vxlan.c      |  11 +-
 .../mellanox/mlxsw/spectrum_nve_vxlan.c       |  19 +-
 .../mellanox/mlxsw/spectrum_switchdev.c       |  57 +-
 drivers/net/vxlan/vxlan_core.c                | 754 +++++++++++-------
 drivers/net/vxlan/vxlan_mdb.c                 |  45 +-
 drivers/net/vxlan/vxlan_multicast.c           |  76 +-
 drivers/net/vxlan/vxlan_private.h             |  34 +-
 drivers/net/vxlan/vxlan_vnifilter.c           | 191 +++--
 include/net/vxlan.h                           |  12 +-
 9 files changed, 769 insertions(+), 430 deletions(-)

-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply	[flat|nested] 12+ messages in thread

* [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes
  2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
@ 2026-09-21 10:01 ` Eric Dumazet
  2026-09-22 16:02   ` netdev-bot+sashiko
  2026-09-21 10:01 ` [PATCH v5 net-next 2/8] vxlan: vnifilter: use list_for_each_entry_rcu() in vxlan_vnifilter_dump_dev() Eric Dumazet
                   ` (6 subsequent siblings)
  7 siblings, 1 reply; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

vxlan_changelink() only refreshed the default fdb entries when the
remote IP changed, but vxlan_config_apply() also updates
default_dst.remote_ifindex when the lower device changes. A changelink
that only swaps the lower device therefore left the all zeros mac rdst
pointing at the old ifindex:

	ip link add vxlan0 type vxlan id 10 group 239.1.1.1 dev eth0
	ip link set dev vxlan0 type vxlan group 239.1.1.1 dev eth1

vxlan_xmit_one() uses rdst->remote_ifindex as the route oif, so traffic
kept leaving eth0.

VNI filter entries have the same problem, and are worse: their fdb
entries are keyed on the device remote_ifindex even when the vni
carries its own group, but vxlan_vnilist_update_group() only visited
the vnis without one. vxlan_vni_delete_group() later looks an entry up
with the current remote_ifindex, and vxlan_fdb_find_rdst() requires an
exact match, so the lookup failed and the entry survived the delete.
Re-adding the same vni then appended a second rdst, duplicating
transmitted BUM traffic.

Pass the old and new ifindex down to vxlan_update_default_fdb_entry()
so the append targets the new lower device and the delete still matches
the entry created for the old one, and refresh every vni rather than
only those inheriting the device group. If updating the vni list fails
partway through, unwind the already updated fdb entries so default_dst
and the fdb entries do not diverge.

Also run vxlan_multicast_leave() and vxlan_multicast_join() when
VXLAN_F_VNIFILTER is set so per-VNI multicast memberships are migrated
even when the device default remote_ip is not multicast.

The new ifindex is the one vxlan_config_apply() will commit, which is
the current one when lowerdev is NULL, so that default_dst and the fdb
entries can not diverge.

Fixes: 8bcdc4f3a20b ("vxlan: add changelink support")
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 drivers/net/vxlan/vxlan_core.c      | 31 ++++++++++----
 drivers/net/vxlan/vxlan_private.h   |  6 +++
 drivers/net/vxlan/vxlan_vnifilter.c | 64 +++++++++++++++++++++++------
 3 files changed, 80 insertions(+), 21 deletions(-)

diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index 347245cc1de4ea44f723176312b0b53869a1641c..c4e3e8e8eef57196c3b7120ff4b8ce71207f3bf8 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -4452,6 +4452,7 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 	struct net_device *lowerdev;
 	struct vxlan_config conf;
 	struct vxlan_rdst *dst;
+	u32 new_ifindex;
 	int err;
 
 	if (!rtnl_dev_link_net_capable(dev, vxlan->net))
@@ -4475,13 +4476,16 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 	if (err)
 		return err;
 
+	/* vxlan_config_apply() only commits remote_ifindex if lowerdev is set */
+	new_ifindex = lowerdev ? conf.remote_ifindex : dst->remote_ifindex;
+
 	rem_ip_changed = !vxlan_addr_equal(&conf.remote_ip, &dst->remote_ip);
 	change_igmp = vxlan->dev->flags & IFF_UP &&
 		      (rem_ip_changed ||
-		       dst->remote_ifindex != conf.remote_ifindex);
+		       dst->remote_ifindex != new_ifindex);
 
 	/* handle default dst entry */
-	if (rem_ip_changed) {
+	if (rem_ip_changed || dst->remote_ifindex != new_ifindex) {
 		spin_lock_bh(&vxlan->hash_lock);
 		if (!vxlan_addr_any(&conf.remote_ip)) {
 			err = vxlan_fdb_update(vxlan, all_zeros_mac,
@@ -4490,7 +4494,7 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 					       NLM_F_APPEND | NLM_F_CREATE,
 					       vxlan->cfg.dst_port,
 					       conf.vni, conf.vni,
-					       conf.remote_ifindex,
+					       new_ifindex,
 					       NTF_SELF, 0, true, extack);
 			if (err) {
 				spin_unlock_bh(&vxlan->hash_lock);
@@ -4509,13 +4513,21 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 					   true);
 		spin_unlock_bh(&vxlan->hash_lock);
 
-		/* If vni filtering device, also update fdb entries of
-		 * all vnis that were using default remote ip
+		/* If vni filtering device, also update default fdb entries of
+		 * all vnis
 		 */
 		if (vxlan->cfg.flags & VXLAN_F_VNIFILTER) {
 			err = vxlan_vnilist_update_group(vxlan, &dst->remote_ip,
-							 &conf.remote_ip, extack);
+							 &conf.remote_ip,
+							 dst->remote_ifindex,
+							 new_ifindex, extack);
 			if (err) {
+				vxlan_update_default_fdb_entry(vxlan, conf.vni,
+							       &conf.remote_ip,
+							       &dst->remote_ip,
+							       new_ifindex,
+							       dst->remote_ifindex,
+							       NULL);
 				netdev_adjacent_change_abort(dst->remote_dev,
 							     lowerdev, dev);
 				return err;
@@ -4523,7 +4535,9 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 		}
 	}
 
-	if (change_igmp && vxlan_addr_multicast(&dst->remote_ip))
+	if (change_igmp &&
+	    (vxlan_addr_multicast(&dst->remote_ip) ||
+	     (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
 		err = vxlan_multicast_leave(vxlan);
 
 	if (netif_running(dev) && conf.age_interval != vxlan->cfg.age_interval)
@@ -4535,7 +4549,8 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 	vxlan_config_apply(dev, &conf, lowerdev, vxlan->net, true);
 
 	if (!err && change_igmp &&
-	    vxlan_addr_multicast(&dst->remote_ip))
+	    (vxlan_addr_multicast(&dst->remote_ip) ||
+	     (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
 		err = vxlan_multicast_join(vxlan);
 
 	return err;
diff --git a/drivers/net/vxlan/vxlan_private.h b/drivers/net/vxlan/vxlan_private.h
index b1eec221636088aa1c1674221d5ef0f13698b53f..e4ceca925bde7bea909bd691fa0f4ebfebdefbd3 100644
--- a/drivers/net/vxlan/vxlan_private.h
+++ b/drivers/net/vxlan/vxlan_private.h
@@ -213,9 +213,15 @@ void vxlan_vs_add_vnigrp(struct vxlan_dev *vxlan,
 			 struct vxlan_sock *vs,
 			 bool ipv6);
 void vxlan_vs_del_vnigrp(struct vxlan_dev *vxlan);
+int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
+				   union vxlan_addr *old_remote_ip,
+				   union vxlan_addr *remote_ip,
+				   u32 old_ifindex, u32 new_ifindex,
+				   struct netlink_ext_ack *extack);
 int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
 			       union vxlan_addr *old_remote_ip,
 			       union vxlan_addr *new_remote_ip,
+			       u32 old_ifindex, u32 new_ifindex,
 			       struct netlink_ext_ack *extack);
 
 
diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
index dd94085e088656d27b62420a5c8c95c609510a4c..12fa11a31818456232682edb1c79042478832836 100644
--- a/drivers/net/vxlan/vxlan_vnifilter.c
+++ b/drivers/net/vxlan/vxlan_vnifilter.c
@@ -470,14 +470,19 @@ static const struct nla_policy vni_filter_policy[VXLAN_VNIFILTER_MAX + 1] = {
 	[VXLAN_VNIFILTER_ENTRY] = { .type = NLA_NESTED },
 };
 
-static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
-					  union vxlan_addr *old_remote_ip,
-					  union vxlan_addr *remote_ip,
-					  struct netlink_ext_ack *extack)
+int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
+				   union vxlan_addr *old_remote_ip,
+				   union vxlan_addr *remote_ip,
+				   u32 old_ifindex, u32 new_ifindex,
+				   struct netlink_ext_ack *extack)
 {
-	struct vxlan_rdst *dst = &vxlan->default_dst;
 	int err = 0;
 
+	if (old_remote_ip && remote_ip &&
+	    vxlan_addr_equal(old_remote_ip, remote_ip) &&
+	    old_ifindex == new_ifindex)
+		return 0;
+
 	spin_lock_bh(&vxlan->hash_lock);
 	if (remote_ip && !vxlan_addr_any(remote_ip)) {
 		err = vxlan_fdb_update(vxlan, all_zeros_mac,
@@ -487,7 +492,7 @@ static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
 				       vxlan->cfg.dst_port,
 				       vni,
 				       vni,
-				       dst->remote_ifindex,
+				       new_ifindex,
 				       NTF_SELF, 0, true, extack);
 		if (err) {
 			spin_unlock_bh(&vxlan->hash_lock);
@@ -500,7 +505,7 @@ static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
 				   *old_remote_ip,
 				   vxlan->cfg.dst_port,
 				   vni, vni,
-				   dst->remote_ifindex,
+				   old_ifindex,
 				   true);
 	}
 	spin_unlock_bh(&vxlan->hash_lock);
@@ -546,6 +551,8 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
 
 	ret = vxlan_update_default_fdb_entry(vxlan, vninode->vni,
 					     oldrip, newrip,
+					     dst->remote_ifindex,
+					     dst->remote_ifindex,
 					     extack);
 	if (ret)
 		goto out;
@@ -583,8 +590,10 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
 int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
 			       union vxlan_addr *old_remote_ip,
 			       union vxlan_addr *new_remote_ip,
+			       u32 old_ifindex, u32 new_ifindex,
 			       struct netlink_ext_ack *extack)
 {
+	union vxlan_addr *oldrip, *newrip;
 	struct list_head *headp, *hpos;
 	struct vxlan_vni_group *vg;
 	struct vxlan_vni_node *vent;
@@ -595,17 +604,46 @@ int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
 	headp = &vg->vni_list;
 	list_for_each_prev(hpos, headp) {
 		vent = list_entry(hpos, struct vxlan_vni_node, vlist);
+
 		if (vxlan_addr_any(&vent->remote_ip)) {
-			ret = vxlan_update_default_fdb_entry(vxlan, vent->vni,
-							     old_remote_ip,
-							     new_remote_ip,
-							     extack);
-			if (ret)
-				return ret;
+			oldrip = old_remote_ip;
+			newrip = new_remote_ip;
+		} else {
+			/* A vni with its own group keeps it, but its fdb entry
+			 * is still keyed on the device remote_ifindex.
+			 */
+			oldrip = &vent->remote_ip;
+			newrip = &vent->remote_ip;
 		}
+
+		ret = vxlan_update_default_fdb_entry(vxlan, vent->vni,
+						     oldrip, newrip,
+						     old_ifindex, new_ifindex,
+						     extack);
+		if (ret)
+			goto err_unwind;
 	}
 
 	return 0;
+
+err_unwind:
+	list_for_each_continue(hpos, headp) {
+		vent = list_entry(hpos, struct vxlan_vni_node, vlist);
+
+		if (vxlan_addr_any(&vent->remote_ip)) {
+			oldrip = old_remote_ip;
+			newrip = new_remote_ip;
+		} else {
+			oldrip = &vent->remote_ip;
+			newrip = &vent->remote_ip;
+		}
+
+		vxlan_update_default_fdb_entry(vxlan, vent->vni,
+					       newrip, oldrip,
+					       new_ifindex, old_ifindex,
+					       NULL);
+	}
+	return ret;
 }
 
 static void vxlan_vni_delete_group(struct vxlan_dev *vxlan,
-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 net-next 2/8] vxlan: vnifilter: use list_for_each_entry_rcu() in vxlan_vnifilter_dump_dev()
  2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes Eric Dumazet
@ 2026-09-21 10:01 ` Eric Dumazet
  2026-09-22 16:02   ` netdev-bot+sashiko
  2026-09-21 10:01 ` [PATCH v5 net-next 3/8] vxlan: vnifilter: signal interrupted RTM_GETTUNNEL dumps Eric Dumazet
                   ` (5 subsequent siblings)
  7 siblings, 1 reply; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

RTM_GETTUNNEL dumps currently run under RTNL lock, but
vxlan_vnifilter_dump() also acquires rcu_read_lock().

1) Currently vxlan_vnifilter_dump_dev() traverses vg->vni_list using
   list_for_each_entry_safe(). Even though RTNL is held today, writers
   modify vg->vni_list with list_add_rcu() and list_del_rcu().
   Switch to list_for_each_entry_rcu() for proper RCU traversal and
   as preparation for future lockless dump support.

2) During a paginated dump, RTNL is released between dump skbs.
   If vxlan_vnifilter_dump_dev() returns early because VXLAN_F_VNIFILTER
   is not set or vg has no VNIs, cb->args[1] was not cleared. This leaked
   a non-zero VNI offset to subsequent devices, silently skipping their
   first N VNIs.
   Furthermore, if devices are added or removed between dump calls,
   ordinal device indexes can shift. Track the current device ifindex
   in cb->args[2] and reset cb->args[1] if the device changes.

Fixes: f9c4bb0b245c ("vxlan: vni filtering support on collect metadata device")
Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 drivers/net/vxlan/vxlan_vnifilter.c | 23 ++++++++++++++++++-----
 1 file changed, 18 insertions(+), 5 deletions(-)

diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
index 12fa11a31818456232682edb1c79042478832836..fad7c76418e929866c5c48431b08d179909f5931 100644
--- a/drivers/net/vxlan/vxlan_vnifilter.c
+++ b/drivers/net/vxlan/vxlan_vnifilter.c
@@ -333,22 +333,34 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
 				    struct sk_buff *skb,
 				    struct netlink_callback *cb)
 {
-	struct vxlan_vni_node *tmp, *v, *vbegin = NULL, *vend = NULL;
+	struct vxlan_vni_node *v, *vbegin = NULL, *vend = NULL;
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct tunnel_msg *new_tmsg, *tmsg;
-	int idx = 0, s_idx = cb->args[1];
 	struct vxlan_vni_group *vg;
 	struct nlmsghdr *nlh;
+	int idx = 0, s_idx;
 	bool dump_stats;
 	int err = 0;
 
-	if (!(vxlan->cfg.flags & VXLAN_F_VNIFILTER))
+	if (cb->args[2] != dev->ifindex) {
+		cb->args[1] = 0;
+		cb->args[2] = dev->ifindex;
+	}
+	s_idx = cb->args[1];
+
+	if (!(vxlan->cfg.flags & VXLAN_F_VNIFILTER)) {
+		cb->args[1] = 0;
+		cb->args[2] = 0;
 		return -EINVAL;
+	}
 
 	/* RCU needed because of the vni locking rules (rcu || rtnl) */
 	vg = rcu_dereference(vxlan->vnigrp);
-	if (!vg || !vg->num_vnis)
+	if (!vg || !vg->num_vnis) {
+		cb->args[1] = 0;
+		cb->args[2] = 0;
 		return 0;
+	}
 
 	tmsg = nlmsg_data(cb->nlh);
 	dump_stats = !!(tmsg->flags & TUNNEL_MSG_FLAG_STATS);
@@ -362,7 +374,7 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
 	new_tmsg->family = PF_BRIDGE;
 	new_tmsg->ifindex = dev->ifindex;
 
-	list_for_each_entry_safe(v, tmp, &vg->vni_list, vlist) {
+	list_for_each_entry_rcu(v, &vg->vni_list, vlist) {
 		if (idx < s_idx) {
 			idx++;
 			continue;
@@ -394,6 +406,7 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
 	}
 
 	cb->args[1] = err ? idx : 0;
+	cb->args[2] = err ? dev->ifindex : 0;
 
 	nlmsg_end(skb, nlh);
 
-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 net-next 3/8] vxlan: vnifilter: signal interrupted RTM_GETTUNNEL dumps
  2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 2/8] vxlan: vnifilter: use list_for_each_entry_rcu() in vxlan_vnifilter_dump_dev() Eric Dumazet
@ 2026-09-21 10:01 ` Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 4/8] vxlan: pass vxlan_config pointer to helper functions Eric Dumazet
                   ` (4 subsequent siblings)
  7 siblings, 0 replies; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

vxlan_vnifilter_dump() walks all vxlan devices of a netns, and for each
one walks vg->vni_list. Both cursors are plain ordinals stored in
cb->args[0] and cb->args[1], and RTNL is released between dump skbs.

A concurrent __vxlan_vni_add_list(), which inserts sorted by VNI,
__vxlan_vni_del_list(), or an in-place group update in
vxlan_vni_update() (which splits or merges coalesced VNI ranges) shifts
the second cursor, duplicating or skipping entries. Unregistering a
vxlan device shifts the first one, and the partially dumped device is
then skipped altogether by the "if (idx < s_idx)" test, silently losing
the rest of its VNIs.

Add a per-netns generation counter, bumped whenever the set of vxlan
devices in the netns or any vni_list changes, and feed it to
nl_dump_check_consistent() so that user space gets NLM_F_DUMP_INTR and
can retry, as vxlan_mdb_dump() already does.

The counter is keyed on dev_net(vxlan->dev) rather than vxlan->net,
because the dump enumerates devices with for_each_netdev_rcu() in the
netns the netdevice lives in, which differs from the packet i/o netns
when the device was created with a separate link netns.

Fixes: f9c4bb0b245c ("vxlan: vni filtering support on collect metadata device")
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
 drivers/net/vxlan/vxlan_core.c      |  4 ++++
 drivers/net/vxlan/vxlan_private.h   |  9 ++++++++
 drivers/net/vxlan/vxlan_vnifilter.c | 35 ++++++++++++++++++++++++-----
 3 files changed, 42 insertions(+), 6 deletions(-)

diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index c4e3e8e8eef57196c3b7120ff4b8ce71207f3bf8..03272fe95fa66d535094adf4b4df92da53d5c417 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -4767,6 +4767,10 @@ static int vxlan_netdevice_event(struct notifier_block *unused,
 	struct net_device *dev = netdev_notifier_info_to_dev(ptr);
 	struct vxlan_net *vn = net_generic(dev_net(dev), vxlan_net_id);
 
+	if ((event == NETDEV_REGISTER || event == NETDEV_UNREGISTER) &&
+	    netif_is_vxlan(dev))
+		vxlan_vnifilter_seq_inc(dev_net(dev));
+
 	if (event == NETDEV_UNREGISTER)
 		vxlan_handle_lowerdev_unregister(vn, dev);
 	else if (event == NETDEV_UDP_TUNNEL_PUSH_INFO)
diff --git a/drivers/net/vxlan/vxlan_private.h b/drivers/net/vxlan/vxlan_private.h
index e4ceca925bde7bea909bd691fa0f4ebfebdefbd3..74d713717cd5da267fc0aef9812593b5ed3ef37e 100644
--- a/drivers/net/vxlan/vxlan_private.h
+++ b/drivers/net/vxlan/vxlan_private.h
@@ -22,6 +22,8 @@ struct vxlan_net {
 	/* sock_list is protected by rtnl lock */
 	struct hlist_head sock_list[PORT_HASH_SIZE];
 	struct notifier_block nexthop_notifier_block;
+	/* Generation counter for RTM_GETTUNNEL dumps */
+	atomic_t vnifilter_seq;
 };
 
 struct vxlan_fdb_key {
@@ -177,6 +179,13 @@ vxlan_vnifilter_lookup(struct vxlan_dev *vxlan, __be32 vni)
 				      vxlan_vni_rht_params);
 }
 
+static inline void vxlan_vnifilter_seq_inc(const struct net *net)
+{
+	struct vxlan_net *vn = net_generic(net, vxlan_net_id);
+
+	atomic_inc(&vn->vnifilter_seq);
+}
+
 /* vxlan_core.c */
 int vxlan_fdb_create(struct vxlan_dev *vxlan,
 		     const u8 *mac, union vxlan_addr *ip,
diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
index fad7c76418e929866c5c48431b08d179909f5931..93fca54749e76da8ae1ad1c15c0461a35d0ee194 100644
--- a/drivers/net/vxlan/vxlan_vnifilter.c
+++ b/drivers/net/vxlan/vxlan_vnifilter.c
@@ -410,9 +410,22 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
 
 	nlmsg_end(skb, nlh);
 
+	nl_dump_check_consistent(cb, nlh);
+
 	return err;
 }
 
+static u32 vxlan_vnifilter_base_seq(const struct net *net)
+{
+	const struct vxlan_net *vn = net_generic(net, vxlan_net_id);
+	u32 res = atomic_read(&vn->vnifilter_seq);
+
+	/* Must not return 0 (see nl_dump_check_consistent()) */
+	if (!res)
+		res = 0x80000000;
+	return res;
+}
+
 static int vxlan_vnifilter_dump(struct sk_buff *skb, struct netlink_callback *cb)
 {
 	int idx = 0, err = 0, s_idx = cb->args[0];
@@ -432,6 +445,9 @@ static int vxlan_vnifilter_dump(struct sk_buff *skb, struct netlink_callback *cb
 	}
 
 	rcu_read_lock();
+
+	cb->seq = vxlan_vnifilter_base_seq(net);
+
 	if (tmsg->ifindex) {
 		dev = dev_get_by_index_rcu(net, tmsg->ifindex);
 		if (!dev) {
@@ -570,8 +586,11 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
 	if (ret)
 		goto out;
 
-	if (group)
+	if (group) {
 		memcpy(&vninode->remote_ip, group, sizeof(vninode->remote_ip));
+		if (!create)
+			vxlan_vnifilter_seq_inc(dev_net(vxlan->dev));
+	}
 
 	if (vxlan->dev->flags & IFF_UP) {
 		if (vxlan_addr_multicast(&old_remote_ip) &&
@@ -716,7 +735,8 @@ static int vxlan_vni_update(struct vxlan_dev *vxlan,
 	return 0;
 }
 
-static void __vxlan_vni_add_list(struct vxlan_vni_group *vg,
+static void __vxlan_vni_add_list(struct vxlan_dev *vxlan,
+				 struct vxlan_vni_group *vg,
 				 struct vxlan_vni_node *v)
 {
 	struct list_head *headp, *hpos;
@@ -732,13 +752,16 @@ static void __vxlan_vni_add_list(struct vxlan_vni_group *vg,
 	}
 	list_add_rcu(&v->vlist, hpos);
 	vg->num_vnis++;
+	vxlan_vnifilter_seq_inc(dev_net(vxlan->dev));
 }
 
-static void __vxlan_vni_del_list(struct vxlan_vni_group *vg,
+static void __vxlan_vni_del_list(struct vxlan_dev *vxlan,
+				 struct vxlan_vni_group *vg,
 				 struct vxlan_vni_node *v)
 {
 	list_del_rcu(&v->vlist);
 	vg->num_vnis--;
+	vxlan_vnifilter_seq_inc(dev_net(vxlan->dev));
 }
 
 static struct vxlan_vni_node *vxlan_vni_alloc(struct vxlan_dev *vxlan,
@@ -800,7 +823,7 @@ static int vxlan_vni_add(struct vxlan_dev *vxlan,
 		return err;
 	}
 
-	__vxlan_vni_add_list(vg, vninode);
+	__vxlan_vni_add_list(vxlan, vg, vninode);
 
 	if (vxlan->dev->flags & IFF_UP)
 		vxlan_vs_add_del_vninode(vxlan, vninode, false);
@@ -846,7 +869,7 @@ static int vxlan_vni_del(struct vxlan_dev *vxlan,
 	if (err)
 		goto out;
 
-	__vxlan_vni_del_list(vg, vninode);
+	__vxlan_vni_del_list(vxlan, vg, vninode);
 
 	vxlan_vnifilter_notify(vxlan, vninode, RTM_DELTUNNEL);
 
@@ -960,7 +983,7 @@ void vxlan_vnigroup_uninit(struct vxlan_dev *vxlan)
 #if IS_ENABLED(CONFIG_IPV6)
 		hlist_del_init_rcu(&v->hlist6.hlist);
 #endif
-		__vxlan_vni_del_list(vg, v);
+		__vxlan_vni_del_list(vxlan, vg, v);
 		vxlan_vnifilter_notify(vxlan, v, RTM_DELTUNNEL);
 		call_rcu(&v->rcu, vxlan_vni_node_rcu_free);
 	}
-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 net-next 4/8] vxlan: pass vxlan_config pointer to helper functions
  2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
                   ` (2 preceding siblings ...)
  2026-09-21 10:01 ` [PATCH v5 net-next 3/8] vxlan: vnifilter: signal interrupted RTM_GETTUNNEL dumps Eric Dumazet
@ 2026-09-21 10:01 ` Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 5/8] vxlan: move VXLAN_F_MDB to struct vxlan_dev flags Eric Dumazet
                   ` (3 subsequent siblings)
  7 siblings, 0 replies; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

In preparation for converting vxlan->cfg to an RCU-protected pointer,
refactor internal helper functions in the RX, TX, MDB, and VNIFILTER
paths to accept a pointer to struct vxlan_config (or pass flags/
saddr_family where appropriate) rather than directly accessing
vxlan->cfg.

No functional changes.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 drivers/net/vxlan/vxlan_core.c      | 343 +++++++++++++++-------------
 drivers/net/vxlan/vxlan_mdb.c       |  21 +-
 drivers/net/vxlan/vxlan_private.h   |   8 +-
 drivers/net/vxlan/vxlan_vnifilter.c |   5 +-
 4 files changed, 206 insertions(+), 171 deletions(-)

diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index 03272fe95fa66d535094adf4b4df92da53d5c417..e834b58ea745ce8394ff2ae6cef05eec3a9f8d26 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -377,14 +377,15 @@ static void vxlan_fdb_miss(struct vxlan_dev *vxlan, const u8 eth_addr[ETH_ALEN])
 
 /* Look up Ethernet address in forwarding table */
 static struct vxlan_fdb *vxlan_find_mac_rcu(struct vxlan_dev *vxlan,
+					    const struct vxlan_config *cfg,
 					    const u8 *mac, __be32 vni)
 {
 	struct vxlan_fdb_key key;
 
 	memset(&key, 0, sizeof(key));
 	memcpy(key.eth_addr, mac, sizeof(key.eth_addr));
-	if (!(vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA))
-		key.vni = vxlan->default_dst.remote_vni;
+	if (!(cfg->flags & VXLAN_F_COLLECT_METADATA))
+		key.vni = cfg->vni;
 	else
 		key.vni = vni;
 
@@ -393,11 +394,12 @@ static struct vxlan_fdb *vxlan_find_mac_rcu(struct vxlan_dev *vxlan,
 }
 
 static struct vxlan_fdb *vxlan_find_mac_tx(struct vxlan_dev *vxlan,
+					   const struct vxlan_config *cfg,
 					   const u8 *mac, __be32 vni)
 {
 	struct vxlan_fdb *f;
 
-	f = vxlan_find_mac_rcu(vxlan, mac, vni);
+	f = vxlan_find_mac_rcu(vxlan, cfg, mac, vni);
 	if (f) {
 		unsigned long now = jiffies;
 
@@ -416,7 +418,7 @@ static struct vxlan_fdb *vxlan_find_mac(struct vxlan_dev *vxlan,
 	lockdep_assert_held_once(&vxlan->hash_lock);
 
 	rcu_read_lock();
-	f = vxlan_find_mac_rcu(vxlan, mac, vni);
+	f = vxlan_find_mac_rcu(vxlan, &vxlan->cfg, mac, vni);
 	rcu_read_unlock();
 
 	return f;
@@ -457,7 +459,7 @@ int vxlan_fdb_find_uc(struct net_device *dev, const u8 *mac, __be32 vni,
 
 	rcu_read_lock();
 
-	f = vxlan_find_mac_rcu(vxlan, eth_addr, vni);
+	f = vxlan_find_mac_rcu(vxlan, &vxlan->cfg, eth_addr, vni);
 	if (f)
 		rdst = first_remote_rcu(f);
 	if (!rdst) {
@@ -1416,7 +1418,7 @@ static int vxlan_fdb_get(struct sk_buff *skb,
 
 	rcu_read_lock();
 
-	f = vxlan_find_mac_rcu(vxlan, addr, vni);
+	f = vxlan_find_mac_rcu(vxlan, &vxlan->cfg, addr, vni);
 	if (!f) {
 		NL_SET_ERR_MSG(extack, "Fdb entry not found");
 		err = -ENOENT;
@@ -1434,6 +1436,7 @@ static int vxlan_fdb_get(struct sk_buff *skb,
  * and Tunnel endpoint.
  */
 static enum skb_drop_reason vxlan_snoop(struct net_device *dev,
+					const struct vxlan_config *cfg,
 					union vxlan_addr *src_ip,
 					const u8 *src_mac, u32 src_ifindex,
 					__be32 vni)
@@ -1452,7 +1455,7 @@ static enum skb_drop_reason vxlan_snoop(struct net_device *dev,
 		ifindex = src_ifindex;
 #endif
 
-	f = vxlan_find_mac_rcu(vxlan, src_mac, vni);
+	f = vxlan_find_mac_rcu(vxlan, cfg, src_mac, vni);
 	if (likely(f)) {
 		struct vxlan_rdst *rdst = first_remote_rcu(f);
 		unsigned long now = jiffies;
@@ -1488,9 +1491,9 @@ static enum skb_drop_reason vxlan_snoop(struct net_device *dev,
 			vxlan_fdb_update(vxlan, src_mac, src_ip,
 					 NUD_REACHABLE,
 					 NLM_F_EXCL|NLM_F_CREATE,
-					 vxlan->cfg.dst_port,
+					 cfg->dst_port,
 					 vni,
-					 vxlan->default_dst.remote_vni,
+					 cfg->vni,
 					 ifindex, NTF_SELF, 0, true, NULL);
 		spin_unlock(&vxlan->hash_lock);
 	}
@@ -1598,6 +1601,7 @@ static void vxlan_parse_gbp_hdr(struct sk_buff *skb, u32 vxflags,
 }
 
 static enum skb_drop_reason vxlan_set_mac(struct vxlan_dev *vxlan,
+					  const struct vxlan_config *cfg,
 					  struct vxlan_sock *vs,
 					  struct sk_buff *skb, __be32 vni)
 {
@@ -1623,10 +1627,10 @@ static enum skb_drop_reason vxlan_set_mac(struct vxlan_dev *vxlan,
 #endif
 	}
 
-	if (!(vxlan->cfg.flags & VXLAN_F_LEARN))
+	if (!(cfg->flags & VXLAN_F_LEARN))
 		return SKB_NOT_DROPPED_YET;
 
-	return vxlan_snoop(skb->dev, &saddr, eth_hdr(skb)->h_source,
+	return vxlan_snoop(skb->dev, cfg, &saddr, eth_hdr(skb)->h_source,
 			   ifindex, vni);
 }
 
@@ -1657,18 +1661,21 @@ static bool vxlan_ecn_decapsulate(struct vxlan_sock *vs, void *oiph,
 static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 {
 	struct vxlan_vni_node *vninode = NULL;
-	const struct vxlanhdr *vh;
-	struct vxlan_dev *vxlan;
-	struct vxlan_sock *vs;
-	struct vxlan_metadata _md;
-	struct vxlan_metadata *md = &_md;
 	__be16 protocol = htons(ETH_P_TEB);
+	const struct vxlan_config *cfg;
 	enum skb_drop_reason reason;
+	const struct vxlanhdr *vh;
+	struct vxlan_metadata *md;
+	struct vxlan_metadata _md;
+	struct vxlan_dev *vxlan;
 	bool raw_proto = false;
-	void *oiph;
+	struct vxlan_sock *vs;
 	__be32 vni = 0;
+	void *oiph;
 	int nh;
 
+	md = &_md;
+
 	/* Need UDP and VXLAN header to be present */
 	reason = pskb_may_pull_reason(skb, VXLAN_HLEN);
 	if (reason)
@@ -1696,8 +1703,9 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 		goto drop;
 	}
 
-	if (vh->vx_flags & vxlan->cfg.reserved_bits.vx_flags ||
-	    vh->vx_vni & vxlan->cfg.reserved_bits.vx_vni) {
+	cfg = &vxlan->cfg;
+	if (vh->vx_flags & cfg->reserved_bits.vx_flags ||
+	    vh->vx_vni & cfg->reserved_bits.vx_vni) {
 		/* If the header uses bits besides those enabled by the
 		 * netdevice configuration, treat this as a malformed packet.
 		 * This behavior diverges from VXLAN RFC (RFC7348) which
@@ -1709,12 +1717,12 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 		reason = SKB_DROP_REASON_VXLAN_INVALID_HDR;
 		DEV_STATS_INC(vxlan->dev, rx_frame_errors);
 		DEV_STATS_INC(vxlan->dev, rx_errors);
-		vxlan_vnifilter_count(vxlan, vni, vninode,
+		vxlan_vnifilter_count(vxlan, cfg, vni, vninode,
 				      VXLAN_VNI_STATS_RX_ERRORS, 0);
 		goto drop;
 	}
 
-	if (vxlan->cfg.flags & VXLAN_F_GPE) {
+	if (cfg->flags & VXLAN_F_GPE) {
 		if (!vxlan_parse_gpe_proto(vh, &protocol))
 			goto drop;
 		raw_proto = true;
@@ -1726,8 +1734,8 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 		goto drop;
 	}
 
-	if (vxlan->cfg.flags & VXLAN_F_REMCSUM_RX) {
-		reason = vxlan_remcsum(skb, vxlan->cfg.flags);
+	if (cfg->flags & VXLAN_F_REMCSUM_RX) {
+		reason = vxlan_remcsum(skb, cfg->flags);
 		if (unlikely(reason))
 			goto drop;
 	}
@@ -1752,14 +1760,14 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 		memset(md, 0, sizeof(*md));
 	}
 
-	if (vxlan->cfg.flags & VXLAN_F_GBP)
-		vxlan_parse_gbp_hdr(skb, vxlan->cfg.flags, md);
+	if (cfg->flags & VXLAN_F_GBP)
+		vxlan_parse_gbp_hdr(skb, cfg->flags, md);
 	/* Note that GBP and GPE can never be active together. This is
 	 * ensured in vxlan_dev_configure.
 	 */
 
 	if (!raw_proto) {
-		reason = vxlan_set_mac(vxlan, vs, skb, vni);
+		reason = vxlan_set_mac(vxlan, cfg, vs, skb, vni);
 		if (reason)
 			goto drop;
 	} else {
@@ -1780,7 +1788,7 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 	if (reason) {
 		DEV_STATS_INC(vxlan->dev, rx_length_errors);
 		DEV_STATS_INC(vxlan->dev, rx_errors);
-		vxlan_vnifilter_count(vxlan, vni, vninode,
+		vxlan_vnifilter_count(vxlan, cfg, vni, vninode,
 				      VXLAN_VNI_STATS_RX_ERRORS, 0);
 		goto drop;
 	}
@@ -1792,7 +1800,7 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 		reason = SKB_DROP_REASON_IP_TUNNEL_ECN;
 		DEV_STATS_INC(vxlan->dev, rx_frame_errors);
 		DEV_STATS_INC(vxlan->dev, rx_errors);
-		vxlan_vnifilter_count(vxlan, vni, vninode,
+		vxlan_vnifilter_count(vxlan, cfg, vni, vninode,
 				      VXLAN_VNI_STATS_RX_ERRORS, 0);
 		goto drop;
 	}
@@ -1802,14 +1810,15 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 	if (unlikely(!(vxlan->dev->flags & IFF_UP))) {
 		rcu_read_unlock();
 		dev_dstats_rx_dropped(vxlan->dev);
-		vxlan_vnifilter_count(vxlan, vni, vninode,
+		vxlan_vnifilter_count(vxlan, cfg, vni, vninode,
 				      VXLAN_VNI_STATS_RX_DROPS, 0);
 		reason = SKB_DROP_REASON_DEV_READY;
 		goto drop;
 	}
 
 	dev_dstats_rx_add(vxlan->dev, skb->len);
-	vxlan_vnifilter_count(vxlan, vni, vninode, VXLAN_VNI_STATS_RX, skb->len);
+	vxlan_vnifilter_count(vxlan, cfg, vni, vninode, VXLAN_VNI_STATS_RX,
+			      skb->len);
 	gro_cells_receive(&vxlan->gro_cells, skb);
 
 	rcu_read_unlock();
@@ -1850,7 +1859,7 @@ static int vxlan_err_lookup(struct sock *sk, struct sk_buff *skb)
 	return 0;
 }
 
-static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni)
+static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni, u32 flags)
 {
 	struct neigh_table *tbl = arp_table(dev_net(dev));
 	struct vxlan_dev *vxlan = netdev_priv(dev);
@@ -1864,7 +1873,7 @@ static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni)
 
 	if (!pskb_network_may_pull(skb, arp_hdr_len(dev))) {
 		dev_dstats_tx_dropped(dev);
-		vxlan_vnifilter_count(vxlan, vni, NULL,
+		vxlan_vnifilter_count(vxlan, &vxlan->cfg, vni, NULL,
 				      VXLAN_VNI_STATS_TX_DROPS, 0);
 		goto out;
 	}
@@ -1905,7 +1914,7 @@ static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni)
 		neigh_ha_snapshot(ha, n, n->dev);
 
 		rcu_read_lock();
-		f = vxlan_find_mac_tx(vxlan, ha, vni);
+		f = vxlan_find_mac_tx(vxlan, &vxlan->cfg, ha, vni);
 		if (f)
 			rdst = first_remote_rcu(f);
 		if (rdst && vxlan_addr_any(&rdst->remote_ip)) {
@@ -1931,11 +1940,11 @@ static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni)
 
 		if (netif_rx(reply) == NET_RX_DROP) {
 			dev_dstats_rx_dropped(dev);
-			vxlan_vnifilter_count(vxlan, vni, NULL,
+			vxlan_vnifilter_count(vxlan, &vxlan->cfg, vni, NULL,
 					      VXLAN_VNI_STATS_RX_DROPS, 0);
 		}
 
-	} else if (vxlan->cfg.flags & VXLAN_F_L3MISS) {
+	} else if (flags & VXLAN_F_L3MISS) {
 		union vxlan_addr ipa = {
 			.sin.sin_addr.s_addr = tip,
 			.sin.sin_family = AF_INET,
@@ -2043,7 +2052,7 @@ static struct sk_buff *vxlan_na_create(struct sk_buff *request,
 	return reply;
 }
 
-static int neigh_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni)
+static int neigh_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni, u32 flags)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	const struct in6_addr *daddr;
@@ -2077,7 +2086,7 @@ static int neigh_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni)
 		}
 
 		neigh_ha_snapshot(ha, n, n->dev);
-		f = vxlan_find_mac_tx(vxlan, ha, vni);
+		f = vxlan_find_mac_tx(vxlan, &vxlan->cfg, ha, vni);
 		if (f)
 			rdst = first_remote_rcu(f);
 		if (rdst && vxlan_addr_any(&rdst->remote_ip)) {
@@ -2096,10 +2105,10 @@ static int neigh_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni)
 
 		if (netif_rx(reply) == NET_RX_DROP) {
 			dev_dstats_rx_dropped(dev);
-			vxlan_vnifilter_count(vxlan, vni, NULL,
+			vxlan_vnifilter_count(vxlan, &vxlan->cfg, vni, NULL,
 					      VXLAN_VNI_STATS_RX_DROPS, 0);
 		}
-	} else if (vxlan->cfg.flags & VXLAN_F_L3MISS) {
+	} else if (flags & VXLAN_F_L3MISS) {
 		union vxlan_addr ipa = {
 			.sin6.sin6_addr = msg->target,
 			.sin6.sin6_family = AF_INET6,
@@ -2115,9 +2124,9 @@ static int neigh_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni)
 }
 #endif
 
-static bool route_shortcircuit(struct net_device *dev, struct sk_buff *skb)
+static bool route_shortcircuit(struct net_device *dev, struct sk_buff *skb,
+			       const struct vxlan_config *cfg)
 {
-	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct neigh_table *tbl;
 	struct neighbour *n;
 
@@ -2136,7 +2145,7 @@ static bool route_shortcircuit(struct net_device *dev, struct sk_buff *skb)
 		tbl = arp_table(dev_net(dev));
 		pip = ip_hdr(skb);
 		n = neigh_lookup(tbl, &pip->daddr, dev);
-		if (!n && (vxlan->cfg.flags & VXLAN_F_L3MISS)) {
+		if (!n && (cfg->flags & VXLAN_F_L3MISS)) {
 			union vxlan_addr ipa = {
 				.sin.sin_addr.s_addr = pip->daddr,
 				.sin.sin_family = AF_INET,
@@ -2164,7 +2173,7 @@ static bool route_shortcircuit(struct net_device *dev, struct sk_buff *skb)
 		tbl = nd_table(dev_net(dev));
 		pip6 = ipv6_hdr(skb);
 		n = neigh_lookup(tbl, &pip6->daddr, dev);
-		if (!n && (vxlan->cfg.flags & VXLAN_F_L3MISS)) {
+		if (!n && (cfg->flags & VXLAN_F_L3MISS)) {
 			union vxlan_addr ipa = {
 				.sin6.sin6_addr = pip6->daddr,
 				.sin6.sin6_family = AF_INET6,
@@ -2282,20 +2291,21 @@ static int vxlan_build_skb(struct sk_buff *skb, struct dst_entry *dst,
 
 /* Bypass encapsulation if the destination is local */
 static void vxlan_encap_bypass(struct sk_buff *skb, struct vxlan_dev *src_vxlan,
-			       struct vxlan_dev *dst_vxlan, __be32 vni,
-			       bool snoop)
+			       struct vxlan_dev *dst_vxlan,
+			       const struct vxlan_config *src_cfg,
+			       __be32 vni, bool snoop)
 {
+	const struct vxlan_config *dst_cfg = &dst_vxlan->cfg;
 	union vxlan_addr loopback;
-	union vxlan_addr *remote_ip = &dst_vxlan->default_dst.remote_ip;
 	unsigned int len = skb->len;
-	struct net_device *dev;
+	struct net_device *dev = dst_vxlan->dev;
 
 	skb->pkt_type = PACKET_HOST;
 	skb->encapsulation = 0;
-	skb->dev = dst_vxlan->dev;
+	skb->dev = dev;
 	__skb_pull(skb, skb_network_offset(skb));
 
-	if (remote_ip->sa.sa_family == AF_INET) {
+	if (dst_vxlan->default_dst.remote_ip.sa.sa_family == AF_INET) {
 		loopback.sin.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
 		loopback.sa.sa_family =  AF_INET;
 #if IS_ENABLED(CONFIG_IPV6)
@@ -2306,26 +2316,25 @@ static void vxlan_encap_bypass(struct sk_buff *skb, struct vxlan_dev *src_vxlan,
 	}
 
 	rcu_read_lock();
-	dev = skb->dev;
 	if (unlikely(!(dev->flags & IFF_UP))) {
 		kfree_skb_reason(skb, SKB_DROP_REASON_DEV_READY);
 		goto drop;
 	}
 
-	if ((dst_vxlan->cfg.flags & VXLAN_F_LEARN) && snoop)
-		vxlan_snoop(dev, &loopback, eth_hdr(skb)->h_source, 0, vni);
+	if ((dst_cfg->flags & VXLAN_F_LEARN) && snoop)
+		vxlan_snoop(dev, dst_cfg, &loopback, eth_hdr(skb)->h_source, 0, vni);
 
 	dev_dstats_tx_add(src_vxlan->dev, len);
-	vxlan_vnifilter_count(src_vxlan, vni, NULL, VXLAN_VNI_STATS_TX, len);
+	vxlan_vnifilter_count(src_vxlan, src_cfg, vni, NULL, VXLAN_VNI_STATS_TX, len);
 
 	if (__netif_rx(skb) == NET_RX_SUCCESS) {
 		dev_dstats_rx_add(dst_vxlan->dev, len);
-		vxlan_vnifilter_count(dst_vxlan, vni, NULL, VXLAN_VNI_STATS_RX,
+		vxlan_vnifilter_count(dst_vxlan, dst_cfg, vni, NULL, VXLAN_VNI_STATS_RX,
 				      len);
 	} else {
 drop:
 		dev_dstats_rx_dropped(dev);
-		vxlan_vnifilter_count(dst_vxlan, vni, NULL,
+		vxlan_vnifilter_count(dst_vxlan, dst_cfg, vni, NULL,
 				      VXLAN_VNI_STATS_RX_DROPS, 0);
 	}
 	rcu_read_unlock();
@@ -2333,6 +2342,7 @@ static void vxlan_encap_bypass(struct sk_buff *skb, struct vxlan_dev *src_vxlan,
 
 static int encap_bypass_if_local(struct sk_buff *skb, struct net_device *dev,
 				 struct vxlan_dev *vxlan,
+				 const struct vxlan_config *cfg,
 				 int addr_family,
 				 __be16 dst_port, int dst_ifindex, __be32 vni,
 				 struct dst_entry *dst,
@@ -2348,22 +2358,22 @@ static int encap_bypass_if_local(struct sk_buff *skb, struct net_device *dev,
 	/* Bypass encapsulation if the destination is local */
 	if (rt_flags & RTCF_LOCAL &&
 	    !(rt_flags & (RTCF_BROADCAST | RTCF_MULTICAST)) &&
-	    vxlan->cfg.flags & VXLAN_F_LOCALBYPASS) {
+	    cfg->flags & VXLAN_F_LOCALBYPASS) {
 		struct vxlan_dev *dst_vxlan;
 
 		dst_release(dst);
 		dst_vxlan = vxlan_find_vni(vxlan->net, dst_ifindex, vni,
 					   addr_family, dst_port,
-					   vxlan->cfg.flags);
+					   cfg->flags);
 		if (!dst_vxlan) {
 			DEV_STATS_INC(dev, tx_errors);
-			vxlan_vnifilter_count(vxlan, vni, NULL,
+			vxlan_vnifilter_count(vxlan, cfg, vni, NULL,
 					      VXLAN_VNI_STATS_TX_ERRORS, 0);
 			kfree_skb_reason(skb, SKB_DROP_REASON_VXLAN_VNI_NOT_FOUND);
 
 			return -ENOENT;
 		}
-		vxlan_encap_bypass(skb, vxlan, dst_vxlan, vni, true);
+		vxlan_encap_bypass(skb, vxlan, dst_vxlan, cfg, vni, true);
 		return 1;
 	}
 
@@ -2371,30 +2381,35 @@ static int encap_bypass_if_local(struct sk_buff *skb, struct net_device *dev,
 }
 
 void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
+		    const struct vxlan_config *cfg,
 		    __be32 default_vni, struct vxlan_rdst *rdst, bool did_rsc)
 {
+	unsigned int pkt_len = skb->len;
+	struct vxlan_metadata _md = {};
+	__be16 src_port = 0, dst_port;
+	struct dst_entry *ndst = NULL;
+	enum skb_drop_reason reason;
 	struct dst_cache *dst_cache;
+	const struct iphdr *old_iph;
 	struct ip_tunnel_info *info;
 	struct ip_tunnel_key *pkey;
+	struct vxlan_metadata *md;
 	struct ip_tunnel_key key;
-	struct vxlan_dev *vxlan = netdev_priv(dev);
-	const struct iphdr *old_iph;
-	struct vxlan_metadata _md = {};
-	struct vxlan_metadata *md = &_md;
-	unsigned int pkt_len = skb->len;
-	__be16 src_port = 0, dst_port;
-	struct dst_entry *ndst = NULL;
+	struct vxlan_dev *vxlan;
+	u32 flags = cfg->flags;
+	bool udp_sum = false;
+	bool no_eth_encap;
 	int addr_family;
+	bool use_cache;
+	__be32 vni = 0;
 	__u8 tos, ttl;
 	int ifindex;
 	int err = 0;
-	u32 flags = vxlan->cfg.flags;
-	bool use_cache;
-	bool udp_sum = false;
-	bool xnet = !net_eq(vxlan->net, dev_net(vxlan->dev));
-	enum skb_drop_reason reason;
-	bool no_eth_encap;
-	__be32 vni = 0;
+	bool xnet;
+
+	vxlan = netdev_priv(dev);
+	xnet = !net_eq(vxlan->net, dev_net(vxlan->dev));
+	md = &_md;
 
 	no_eth_encap = flags & VXLAN_F_GPE && skb->protocol != htons(ETH_P_TEB);
 	reason = skb_vlan_inet_prepare(skb, no_eth_encap);
@@ -2414,23 +2429,23 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 		if (vxlan_addr_any(&rdst->remote_ip)) {
 			if (did_rsc) {
 				/* short-circuited back to local bridge */
-				vxlan_encap_bypass(skb, vxlan, vxlan,
+				vxlan_encap_bypass(skb, vxlan, vxlan, cfg,
 						   default_vni, true);
 				return;
 			}
 			goto drop;
 		}
 
-		addr_family = vxlan->cfg.saddr.sa.sa_family;
-		dst_port = rdst->remote_port ? rdst->remote_port : vxlan->cfg.dst_port;
+		addr_family = cfg->saddr.sa.sa_family;
+		dst_port = rdst->remote_port ? rdst->remote_port : cfg->dst_port;
 		vni = (rdst->remote_vni) ? : default_vni;
 		ifindex = rdst->remote_ifindex;
 
 		if (addr_family == AF_INET) {
-			key.u.ipv4.src = vxlan->cfg.saddr.sin.sin_addr.s_addr;
+			key.u.ipv4.src = cfg->saddr.sin.sin_addr.s_addr;
 			key.u.ipv4.dst = rdst->remote_ip.sin.sin_addr.s_addr;
 		} else {
-			key.u.ipv6.src = vxlan->cfg.saddr.sin6.sin6_addr;
+			key.u.ipv6.src = cfg->saddr.sin6.sin6_addr;
 			key.u.ipv6.dst = rdst->remote_ip.sin6.sin6_addr;
 		}
 
@@ -2439,11 +2454,11 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 		if (flags & VXLAN_F_TTL_INHERIT) {
 			ttl = ip_tunnel_get_ttl(old_iph, skb);
 		} else {
-			ttl = vxlan->cfg.ttl;
+			ttl = cfg->ttl;
 			if (!ttl && vxlan_addr_multicast(&rdst->remote_ip))
 				ttl = 1;
 		}
-		tos = vxlan->cfg.tos;
+		tos = cfg->tos;
 		if (tos == 1)
 			tos = ip_tunnel_get_dsfield(old_iph, skb);
 		if (tos && !info)
@@ -2454,9 +2469,9 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 		else
 			udp_sum = !(flags & VXLAN_F_UDP_ZERO_CSUM6_TX);
 #if IS_ENABLED(CONFIG_IPV6)
-		switch (vxlan->cfg.label_policy) {
+		switch (cfg->label_policy) {
 		case VXLAN_LABEL_FIXED:
-			key.label = vxlan->cfg.label;
+			key.label = cfg->label;
 			break;
 		case VXLAN_LABEL_INHERIT:
 			key.label = ip_tunnel_get_flowlabel(old_iph, skb);
@@ -2474,7 +2489,7 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 		}
 		pkey = &info->key;
 		addr_family = ip_tunnel_info_af(info);
-		dst_port = info->key.tp_dst ? : vxlan->cfg.dst_port;
+		dst_port = info->key.tp_dst ? : cfg->dst_port;
 		vni = tunnel_id_to_key32(info->key.tun_id);
 		ifindex = 0;
 		dst_cache = &info->dst_cache;
@@ -2487,8 +2502,8 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 		tos = info->key.tos;
 		udp_sum = test_bit(IP_TUNNEL_CSUM_BIT, info->key.tun_flags);
 	}
-	src_port = udp_flow_src_port(dev_net(dev), skb, vxlan->cfg.port_min,
-				     vxlan->cfg.port_max, true);
+	src_port = udp_flow_src_port(dev_net(dev), skb, cfg->port_min,
+				     cfg->port_max, true);
 
 	rcu_read_lock();
 	if (addr_family == AF_INET) {
@@ -2521,15 +2536,15 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 
 		if (!info) {
 			/* Bypass encapsulation if the destination is local */
-			err = encap_bypass_if_local(skb, dev, vxlan, AF_INET,
+			err = encap_bypass_if_local(skb, dev, vxlan, cfg, AF_INET,
 						    dst_port, ifindex, vni,
 						    &rt->dst, rt->rt_flags);
 			if (err)
 				goto out_unlock;
 
-			if (vxlan->cfg.df == VXLAN_DF_SET) {
+			if (cfg->df == VXLAN_DF_SET) {
 				df = htons(IP_DF);
-			} else if (vxlan->cfg.df == VXLAN_DF_INHERIT) {
+			} else if (cfg->df == VXLAN_DF_INHERIT) {
 				struct ethhdr *eth = eth_hdr(skb);
 
 				if (ntohs(eth->h_proto) == ETH_P_IPV6 ||
@@ -2558,7 +2573,7 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 				unclone->key.u.ipv4.src = pkey->u.ipv4.dst;
 				unclone->key.u.ipv4.dst = saddr;
 			}
-			vxlan_encap_bypass(skb, vxlan, vxlan, vni, false);
+			vxlan_encap_bypass(skb, vxlan, vxlan, cfg, vni, false);
 			dst_release(ndst);
 			goto out_unlock;
 		}
@@ -2608,7 +2623,7 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 		if (!info) {
 			u32 rt6i_flags = dst_rt6_info(ndst)->rt6i_flags;
 
-			err = encap_bypass_if_local(skb, dev, vxlan, AF_INET6,
+			err = encap_bypass_if_local(skb, dev, vxlan, cfg, AF_INET6,
 						    dst_port, ifindex, vni,
 						    ndst, rt6i_flags);
 			if (err)
@@ -2632,7 +2647,7 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 				unclone->key.u.ipv6.dst = saddr;
 			}
 
-			vxlan_encap_bypass(skb, vxlan, vxlan, vni, false);
+			vxlan_encap_bypass(skb, vxlan, vxlan, cfg, vni, false);
 			dst_release(ndst);
 			goto out_unlock;
 		}
@@ -2653,14 +2668,14 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 				     ip6cb_flags);
 #endif
 	}
-	vxlan_vnifilter_count(vxlan, vni, NULL, VXLAN_VNI_STATS_TX, pkt_len);
+	vxlan_vnifilter_count(vxlan, cfg, vni, NULL, VXLAN_VNI_STATS_TX, pkt_len);
 out_unlock:
 	rcu_read_unlock();
 	return;
 
 drop:
 	dev_dstats_tx_dropped(dev);
-	vxlan_vnifilter_count(vxlan, vni, NULL, VXLAN_VNI_STATS_TX_DROPS, 0);
+	vxlan_vnifilter_count(vxlan, cfg, vni, NULL, VXLAN_VNI_STATS_TX_DROPS, 0);
 	kfree_skb_reason(skb, reason);
 	return;
 
@@ -2672,11 +2687,12 @@ void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
 		DEV_STATS_INC(dev, tx_carrier_errors);
 	dst_release(ndst);
 	DEV_STATS_INC(dev, tx_errors);
-	vxlan_vnifilter_count(vxlan, vni, NULL, VXLAN_VNI_STATS_TX_ERRORS, 0);
+	vxlan_vnifilter_count(vxlan, cfg, vni, NULL, VXLAN_VNI_STATS_TX_ERRORS, 0);
 	kfree_skb_reason(skb, reason);
 }
 
 static void vxlan_xmit_nh(struct sk_buff *skb, struct net_device *dev,
+			  const struct vxlan_config *cfg,
 			  struct vxlan_fdb *f, __be32 vni, bool did_rsc)
 {
 	struct vxlan_rdst nh_rdst;
@@ -2693,7 +2709,7 @@ static void vxlan_xmit_nh(struct sk_buff *skb, struct net_device *dev,
 	do_xmit = vxlan_fdb_nh_path_select(nh, hash, &nh_rdst);
 
 	if (likely(do_xmit))
-		vxlan_xmit_one(skb, dev, vni, &nh_rdst, did_rsc);
+		vxlan_xmit_one(skb, dev, cfg, vni, &nh_rdst, did_rsc);
 	else
 		goto drop;
 
@@ -2701,15 +2717,15 @@ static void vxlan_xmit_nh(struct sk_buff *skb, struct net_device *dev,
 
 drop:
 	dev_dstats_tx_dropped(dev);
-	vxlan_vnifilter_count(netdev_priv(dev), vni, NULL,
+	vxlan_vnifilter_count(netdev_priv(dev), cfg, vni, NULL,
 			      VXLAN_VNI_STATS_TX_DROPS, 0);
 	dev_kfree_skb(skb);
 }
 
 static netdev_tx_t vxlan_xmit_nhid(struct sk_buff *skb, struct net_device *dev,
-				   u32 nhid, __be32 vni)
+				   u32 nhid, __be32 vni,
+				   const struct vxlan_config *cfg)
 {
-	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct vxlan_rdst nh_rdst;
 	struct nexthop *nh;
 	bool do_xmit;
@@ -2727,11 +2743,11 @@ static netdev_tx_t vxlan_xmit_nhid(struct sk_buff *skb, struct net_device *dev,
 	do_xmit = vxlan_fdb_nh_path_select(nh, hash, &nh_rdst);
 	rcu_read_unlock();
 
-	if (vxlan->cfg.saddr.sa.sa_family != nh_rdst.remote_ip.sa.sa_family)
+	if (cfg->saddr.sa.sa_family != nh_rdst.remote_ip.sa.sa_family)
 		goto drop;
 
 	if (likely(do_xmit))
-		vxlan_xmit_one(skb, dev, vni, &nh_rdst, false);
+		vxlan_xmit_one(skb, dev, cfg, vni, &nh_rdst, false);
 	else
 		goto drop;
 
@@ -2739,7 +2755,7 @@ static netdev_tx_t vxlan_xmit_nhid(struct sk_buff *skb, struct net_device *dev,
 
 drop:
 	dev_dstats_tx_dropped(dev);
-	vxlan_vnifilter_count(netdev_priv(dev), vni, NULL,
+	vxlan_vnifilter_count(netdev_priv(dev), cfg, vni, NULL,
 			      VXLAN_VNI_STATS_TX_DROPS, 0);
 	dev_kfree_skb(skb);
 	return NETDEV_TX_OK;
@@ -2756,34 +2772,41 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct vxlan_rdst *rdst, *fdst = NULL;
 	const struct ip_tunnel_info *info;
+	const struct vxlan_config *cfg;
+	__be32 default_vni;
 	struct vxlan_fdb *f;
 	struct ethhdr *eth;
 	__be32 vni = 0;
-	u32 nhid = 0;
 	bool did_rsc;
+	u32 nhid = 0;
+	u32 flags;
+
+	cfg = &vxlan->cfg;
+	flags = cfg->flags;
+	default_vni = cfg->vni;
 
 	info = skb_tunnel_info(skb);
 
 	skb_reset_mac_header(skb);
 
-	if (vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA) {
+	if (flags & VXLAN_F_COLLECT_METADATA) {
 		if (info && info->mode & IP_TUNNEL_INFO_BRIDGE &&
 		    info->mode & IP_TUNNEL_INFO_TX) {
 			vni = tunnel_id_to_key32(info->key.tun_id);
 			nhid = info->key.nhid;
 		} else {
 			if (info && info->mode & IP_TUNNEL_INFO_TX)
-				vxlan_xmit_one(skb, dev, vni, NULL, false);
+				vxlan_xmit_one(skb, dev, cfg, vni, NULL, false);
 			else
 				kfree_skb_reason(skb, SKB_DROP_REASON_TUNNEL_TXINFO);
 			return NETDEV_TX_OK;
 		}
 	}
 
-	if (vxlan->cfg.flags & VXLAN_F_PROXY) {
+	if (flags & VXLAN_F_PROXY) {
 		eth = eth_hdr(skb);
 		if (ntohs(eth->h_proto) == ETH_P_ARP)
-			return arp_reduce(dev, skb, vni);
+			return arp_reduce(dev, skb, vni, flags);
 #if IS_ENABLED(CONFIG_IPV6)
 		else if (ntohs(eth->h_proto) == ETH_P_IPV6 &&
 			 pskb_network_may_pull(skb, sizeof(struct ipv6hdr) +
@@ -2793,23 +2816,23 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 
 			if (m->icmph.icmp6_code == 0 &&
 			    m->icmph.icmp6_type == NDISC_NEIGHBOUR_SOLICITATION)
-				return neigh_reduce(dev, skb, vni);
+				return neigh_reduce(dev, skb, vni, flags);
 		}
 #endif
 	}
 
 	if (nhid)
-		return vxlan_xmit_nhid(skb, dev, nhid, vni);
+		return vxlan_xmit_nhid(skb, dev, nhid, vni, cfg);
 
-	if (vxlan->cfg.flags & VXLAN_F_MDB) {
+	if (flags & VXLAN_F_MDB) {
 		struct vxlan_mdb_entry *mdb_entry;
 
 		rcu_read_lock();
-		mdb_entry = vxlan_mdb_entry_skb_get(vxlan, skb, vni);
+		mdb_entry = vxlan_mdb_entry_skb_get(vxlan, cfg, skb, vni);
 		if (mdb_entry) {
 			netdev_tx_t ret;
 
-			ret = vxlan_mdb_xmit(vxlan, mdb_entry, skb);
+			ret = vxlan_mdb_xmit(vxlan, cfg, mdb_entry, skb);
 			rcu_read_unlock();
 			return ret;
 		}
@@ -2818,27 +2841,27 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 
 	eth = eth_hdr(skb);
 	rcu_read_lock();
-	f = vxlan_find_mac_tx(vxlan, eth->h_dest, vni);
+	f = vxlan_find_mac_tx(vxlan, cfg, eth->h_dest, vni);
 	did_rsc = false;
 
-	if (f && (f->flags & NTF_ROUTER) && (vxlan->cfg.flags & VXLAN_F_RSC) &&
+	if (f && (f->flags & NTF_ROUTER) && (flags & VXLAN_F_RSC) &&
 	    (ntohs(eth->h_proto) == ETH_P_IP ||
 	     ntohs(eth->h_proto) == ETH_P_IPV6)) {
-		did_rsc = route_shortcircuit(dev, skb);
+		did_rsc = route_shortcircuit(dev, skb, cfg);
 		eth = eth_hdr(skb);
 		if (did_rsc)
-			f = vxlan_find_mac_tx(vxlan, eth->h_dest, vni);
+			f = vxlan_find_mac_tx(vxlan, cfg, eth->h_dest, vni);
 	}
 
 	if (f == NULL) {
-		f = vxlan_find_mac_tx(vxlan, all_zeros_mac, vni);
+		f = vxlan_find_mac_tx(vxlan, cfg, all_zeros_mac, vni);
 		if (f == NULL) {
-			if ((vxlan->cfg.flags & VXLAN_F_L2MISS) &&
+			if ((flags & VXLAN_F_L2MISS) &&
 			    !is_multicast_ether_addr(eth->h_dest))
 				vxlan_fdb_miss(vxlan, eth->h_dest);
 
 			dev_dstats_tx_dropped(dev);
-			vxlan_vnifilter_count(vxlan, vni, NULL,
+			vxlan_vnifilter_count(vxlan, cfg, vni, NULL,
 					      VXLAN_VNI_STATS_TX_DROPS, 0);
 			kfree_skb_reason(skb, SKB_DROP_REASON_NO_TX_TARGET);
 			goto out;
@@ -2846,8 +2869,8 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 	}
 
 	if (rcu_access_pointer(f->nh)) {
-		vxlan_xmit_nh(skb, dev, f,
-			      (vni ? : vxlan->default_dst.remote_vni), did_rsc);
+		vxlan_xmit_nh(skb, dev, cfg, f,
+			      (vni ? : default_vni), did_rsc);
 	} else {
 		list_for_each_entry_rcu(rdst, &f->remotes, list) {
 			struct sk_buff *skb1;
@@ -2858,10 +2881,10 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 			}
 			skb1 = skb_clone(skb, GFP_ATOMIC);
 			if (skb1)
-				vxlan_xmit_one(skb1, dev, vni, rdst, did_rsc);
+				vxlan_xmit_one(skb1, dev, cfg, vni, rdst, did_rsc);
 		}
 		if (fdst)
-			vxlan_xmit_one(skb, dev, vni, fdst, did_rsc);
+			vxlan_xmit_one(skb, dev, cfg, vni, fdst, did_rsc);
 		else
 			kfree_skb_reason(skb, SKB_DROP_REASON_NO_TX_TARGET);
 	}
@@ -3734,7 +3757,7 @@ static int vxlan_sock_add(struct vxlan_dev *vxlan)
 }
 
 int vxlan_vni_in_use(struct net *src_net, struct vxlan_dev *vxlan,
-		     struct vxlan_config *conf, __be32 vni)
+		     const struct vxlan_config *conf, __be32 vni)
 {
 	struct vxlan_net *vn = net_generic(src_net, vxlan_net_id);
 	struct vxlan_dev *tmp;
@@ -4611,10 +4634,10 @@ static int vxlan_fill_info(struct sk_buff *skb, const struct net_device *dev)
 {
 	const struct vxlan_dev *vxlan = netdev_priv(dev);
 	const struct vxlan_rdst *dst = &vxlan->default_dst;
-	struct ifla_vxlan_port_range ports = {
-		.low =  htons(vxlan->cfg.port_min),
-		.high = htons(vxlan->cfg.port_max),
-	};
+	struct ifla_vxlan_port_range ports;
+	const struct vxlan_config *cfg;
+
+	cfg = &vxlan->cfg;
 
 	if (nla_put_u32(skb, IFLA_VXLAN_ID, be32_to_cpu(dst->remote_vni)))
 		goto nla_put_failure;
@@ -4636,79 +4659,81 @@ static int vxlan_fill_info(struct sk_buff *skb, const struct net_device *dev)
 	if (dst->remote_ifindex && nla_put_u32(skb, IFLA_VXLAN_LINK, dst->remote_ifindex))
 		goto nla_put_failure;
 
-	if (!vxlan_addr_any(&vxlan->cfg.saddr)) {
-		if (vxlan->cfg.saddr.sa.sa_family == AF_INET) {
+	if (!vxlan_addr_any(&cfg->saddr)) {
+		if (cfg->saddr.sa.sa_family == AF_INET) {
 			if (nla_put_in_addr(skb, IFLA_VXLAN_LOCAL,
-					    vxlan->cfg.saddr.sin.sin_addr.s_addr))
+					    cfg->saddr.sin.sin_addr.s_addr))
 				goto nla_put_failure;
 #if IS_ENABLED(CONFIG_IPV6)
 		} else {
 			if (nla_put_in6_addr(skb, IFLA_VXLAN_LOCAL6,
-					     &vxlan->cfg.saddr.sin6.sin6_addr))
+					     &cfg->saddr.sin6.sin6_addr))
 				goto nla_put_failure;
 #endif
 		}
 	}
 
-	if (nla_put_u8(skb, IFLA_VXLAN_TTL, vxlan->cfg.ttl) ||
+	if (nla_put_u8(skb, IFLA_VXLAN_TTL, cfg->ttl) ||
 	    nla_put_u8(skb, IFLA_VXLAN_TTL_INHERIT,
-		       !!(vxlan->cfg.flags & VXLAN_F_TTL_INHERIT)) ||
-	    nla_put_u8(skb, IFLA_VXLAN_TOS, vxlan->cfg.tos) ||
-	    nla_put_u8(skb, IFLA_VXLAN_DF, vxlan->cfg.df) ||
-	    nla_put_be32(skb, IFLA_VXLAN_LABEL, vxlan->cfg.label) ||
-	    nla_put_u32(skb, IFLA_VXLAN_LABEL_POLICY, vxlan->cfg.label_policy) ||
+		       !!(cfg->flags & VXLAN_F_TTL_INHERIT)) ||
+	    nla_put_u8(skb, IFLA_VXLAN_TOS, cfg->tos) ||
+	    nla_put_u8(skb, IFLA_VXLAN_DF, cfg->df) ||
+	    nla_put_be32(skb, IFLA_VXLAN_LABEL, cfg->label) ||
+	    nla_put_u32(skb, IFLA_VXLAN_LABEL_POLICY, cfg->label_policy) ||
 	    nla_put_u8(skb, IFLA_VXLAN_LEARNING,
-		       !!(vxlan->cfg.flags & VXLAN_F_LEARN)) ||
+		       !!(cfg->flags & VXLAN_F_LEARN)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_PROXY,
-		       !!(vxlan->cfg.flags & VXLAN_F_PROXY)) ||
+		       !!(cfg->flags & VXLAN_F_PROXY)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_RSC,
-		       !!(vxlan->cfg.flags & VXLAN_F_RSC)) ||
+		       !!(cfg->flags & VXLAN_F_RSC)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_L2MISS,
-		       !!(vxlan->cfg.flags & VXLAN_F_L2MISS)) ||
+		       !!(cfg->flags & VXLAN_F_L2MISS)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_L3MISS,
-		       !!(vxlan->cfg.flags & VXLAN_F_L3MISS)) ||
+		       !!(cfg->flags & VXLAN_F_L3MISS)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_COLLECT_METADATA,
-		       !!(vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA)) ||
-	    nla_put_u32(skb, IFLA_VXLAN_AGEING, vxlan->cfg.age_interval) ||
-	    nla_put_u32(skb, IFLA_VXLAN_LIMIT, vxlan->cfg.addrmax) ||
-	    nla_put_be16(skb, IFLA_VXLAN_PORT, vxlan->cfg.dst_port) ||
+		       !!(cfg->flags & VXLAN_F_COLLECT_METADATA)) ||
+	    nla_put_u32(skb, IFLA_VXLAN_AGEING, cfg->age_interval) ||
+	    nla_put_u32(skb, IFLA_VXLAN_LIMIT, cfg->addrmax) ||
+	    nla_put_be16(skb, IFLA_VXLAN_PORT, cfg->dst_port) ||
 	    nla_put_u8(skb, IFLA_VXLAN_UDP_CSUM,
-		       !(vxlan->cfg.flags & VXLAN_F_UDP_ZERO_CSUM_TX)) ||
+		       !(cfg->flags & VXLAN_F_UDP_ZERO_CSUM_TX)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_UDP_ZERO_CSUM6_TX,
-		       !!(vxlan->cfg.flags & VXLAN_F_UDP_ZERO_CSUM6_TX)) ||
+		       !!(cfg->flags & VXLAN_F_UDP_ZERO_CSUM6_TX)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_UDP_ZERO_CSUM6_RX,
-		       !!(vxlan->cfg.flags & VXLAN_F_UDP_ZERO_CSUM6_RX)) ||
+		       !!(cfg->flags & VXLAN_F_UDP_ZERO_CSUM6_RX)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_REMCSUM_TX,
-		       !!(vxlan->cfg.flags & VXLAN_F_REMCSUM_TX)) ||
+		       !!(cfg->flags & VXLAN_F_REMCSUM_TX)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_REMCSUM_RX,
-		       !!(vxlan->cfg.flags & VXLAN_F_REMCSUM_RX)) ||
+		       !!(cfg->flags & VXLAN_F_REMCSUM_RX)) ||
 	    nla_put_u8(skb, IFLA_VXLAN_LOCALBYPASS,
-		       !!(vxlan->cfg.flags & VXLAN_F_LOCALBYPASS)))
+		       !!(cfg->flags & VXLAN_F_LOCALBYPASS)))
 		goto nla_put_failure;
 
+	ports.low = htons(cfg->port_min);
+	ports.high = htons(cfg->port_max);
 	if (nla_put(skb, IFLA_VXLAN_PORT_RANGE, sizeof(ports), &ports))
 		goto nla_put_failure;
 
-	if (vxlan->cfg.flags & VXLAN_F_GBP &&
+	if (cfg->flags & VXLAN_F_GBP &&
 	    nla_put_flag(skb, IFLA_VXLAN_GBP))
 		goto nla_put_failure;
 
-	if (vxlan->cfg.flags & VXLAN_F_GPE &&
+	if (cfg->flags & VXLAN_F_GPE &&
 	    nla_put_flag(skb, IFLA_VXLAN_GPE))
 		goto nla_put_failure;
 
-	if (vxlan->cfg.flags & VXLAN_F_REMCSUM_NOPARTIAL &&
+	if (cfg->flags & VXLAN_F_REMCSUM_NOPARTIAL &&
 	    nla_put_flag(skb, IFLA_VXLAN_REMCSUM_NOPARTIAL))
 		goto nla_put_failure;
 
-	if (vxlan->cfg.flags & VXLAN_F_VNIFILTER &&
+	if (cfg->flags & VXLAN_F_VNIFILTER &&
 	    nla_put_u8(skb, IFLA_VXLAN_VNIFILTER,
-		       !!(vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
+		       !!(cfg->flags & VXLAN_F_VNIFILTER)))
 		goto nla_put_failure;
 
 	if (nla_put(skb, IFLA_VXLAN_RESERVED_BITS,
-		    sizeof(vxlan->cfg.reserved_bits),
-		    &vxlan->cfg.reserved_bits))
+		    sizeof(cfg->reserved_bits),
+		    &cfg->reserved_bits))
 		goto nla_put_failure;
 
 	return 0;
diff --git a/drivers/net/vxlan/vxlan_mdb.c b/drivers/net/vxlan/vxlan_mdb.c
index 841f42ffecb9aa351b10cfadf3d423f7e0f9f290..56ca9283283307d8b1231c3b26724474e6c838d6 100644
--- a/drivers/net/vxlan/vxlan_mdb.c
+++ b/drivers/net/vxlan/vxlan_mdb.c
@@ -165,6 +165,7 @@ static int vxlan_mdb_entry_info_fill(const struct vxlan_dev *vxlan,
 				     const struct vxlan_mdb_entry *mdb_entry,
 				     const struct vxlan_mdb_remote *remote)
 {
+	const struct vxlan_config *cfg = &vxlan->cfg;
 	struct vxlan_rdst *rd = rtnl_dereference(remote->rd);
 	struct br_mdb_entry e;
 	struct nlattr *nest;
@@ -189,7 +190,7 @@ static int vxlan_mdb_entry_info_fill(const struct vxlan_dev *vxlan,
 	    vxlan_nla_put_addr(skb, MDBA_MDB_EATTR_DST, &rd->remote_ip))
 		goto nest_err;
 
-	if (rd->remote_port && rd->remote_port != vxlan->cfg.dst_port &&
+	if (rd->remote_port && rd->remote_port != cfg->dst_port &&
 	    nla_put_u16(skb, MDBA_MDB_EATTR_DST_PORT,
 			be16_to_cpu(rd->remote_port)))
 		goto nest_err;
@@ -202,7 +203,7 @@ static int vxlan_mdb_entry_info_fill(const struct vxlan_dev *vxlan,
 	    nla_put_u32(skb, MDBA_MDB_EATTR_IFINDEX, rd->remote_ifindex))
 		goto nest_err;
 
-	if ((vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA) &&
+	if ((cfg->flags & VXLAN_F_COLLECT_METADATA) &&
 	    mdb_entry->key.vni && nla_put_u32(skb, MDBA_MDB_EATTR_SRC_VNI,
 					      be32_to_cpu(mdb_entry->key.vni)))
 		goto nest_err;
@@ -613,6 +614,7 @@ static int vxlan_mdb_config_init(struct vxlan_mdb_config *cfg,
 {
 	struct br_mdb_entry *entry = nla_data(tb[MDBA_SET_ENTRY]);
 	struct vxlan_dev *vxlan = netdev_priv(dev);
+	const struct vxlan_config *vcfg = &vxlan->cfg;
 
 	memset(cfg, 0, sizeof(*cfg));
 	cfg->vxlan = vxlan;
@@ -622,7 +624,7 @@ static int vxlan_mdb_config_init(struct vxlan_mdb_config *cfg,
 	cfg->filter_mode = MCAST_EXCLUDE;
 	cfg->rt_protocol = RTPROT_STATIC;
 	cfg->remote_vni = vxlan->default_dst.remote_vni;
-	cfg->remote_port = vxlan->cfg.dst_port;
+	cfg->remote_port = vcfg->dst_port;
 
 	if (entry->ifindex != dev->ifindex) {
 		NL_SET_ERR_MSG_MOD(extack, "Port net device must be the VXLAN net device");
@@ -957,6 +959,7 @@ vxlan_mdb_nlmsg_remote_size(const struct vxlan_dev *vxlan,
 			    const struct vxlan_mdb_entry *mdb_entry,
 			    const struct vxlan_mdb_remote *remote)
 {
+	const struct vxlan_config *cfg = &vxlan->cfg;
 	const struct vxlan_mdb_entry_key *group = &mdb_entry->key;
 	struct vxlan_rdst *rd = rtnl_dereference(remote->rd);
 	size_t nlmsg_size;
@@ -978,7 +981,7 @@ vxlan_mdb_nlmsg_remote_size(const struct vxlan_dev *vxlan,
 	/* MDBA_MDB_EATTR_DST */
 	nlmsg_size += nla_total_size(vxlan_addr_size(&rd->remote_ip));
 	/* MDBA_MDB_EATTR_DST_PORT */
-	if (rd->remote_port && rd->remote_port != vxlan->cfg.dst_port)
+	if (rd->remote_port && rd->remote_port != cfg->dst_port)
 		nlmsg_size += nla_total_size(sizeof(u16));
 	/* MDBA_MDB_EATTR_VNI */
 	if (rd->remote_vni != vxlan->default_dst.remote_vni)
@@ -987,7 +990,7 @@ vxlan_mdb_nlmsg_remote_size(const struct vxlan_dev *vxlan,
 	if (rd->remote_ifindex)
 		nlmsg_size += nla_total_size(sizeof(u32));
 	/* MDBA_MDB_EATTR_SRC_VNI */
-	if ((vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA) && group->vni)
+	if ((cfg->flags & VXLAN_F_COLLECT_METADATA) && group->vni)
 		nlmsg_size += nla_total_size(sizeof(u32));
 
 	return nlmsg_size;
@@ -1621,6 +1624,7 @@ int vxlan_mdb_get(struct net_device *dev, struct nlattr *tb[], u32 portid,
 }
 
 struct vxlan_mdb_entry *vxlan_mdb_entry_skb_get(struct vxlan_dev *vxlan,
+						const struct vxlan_config *cfg,
 						struct sk_buff *skb,
 						__be32 src_vni)
 {
@@ -1634,7 +1638,7 @@ struct vxlan_mdb_entry *vxlan_mdb_entry_skb_get(struct vxlan_dev *vxlan,
 	/* When not in collect metadata mode, 'src_vni' is zero, but MDB
 	 * entries are stored with the VNI of the VXLAN device.
 	 */
-	if (!(vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA))
+	if (!(cfg->flags & VXLAN_F_COLLECT_METADATA))
 		src_vni = vxlan->default_dst.remote_vni;
 
 	memset(&group, 0, sizeof(group));
@@ -1700,6 +1704,7 @@ struct vxlan_mdb_entry *vxlan_mdb_entry_skb_get(struct vxlan_dev *vxlan,
 }
 
 netdev_tx_t vxlan_mdb_xmit(struct vxlan_dev *vxlan,
+			   const struct vxlan_config *cfg,
 			   const struct vxlan_mdb_entry *mdb_entry,
 			   struct sk_buff *skb)
 {
@@ -1721,12 +1726,12 @@ netdev_tx_t vxlan_mdb_xmit(struct vxlan_dev *vxlan,
 
 		skb1 = skb_clone(skb, GFP_ATOMIC);
 		if (skb1)
-			vxlan_xmit_one(skb1, vxlan->dev, src_vni,
+			vxlan_xmit_one(skb1, vxlan->dev, cfg, src_vni,
 				       rcu_dereference(remote->rd), false);
 	}
 
 	if (fremote)
-		vxlan_xmit_one(skb, vxlan->dev, src_vni,
+		vxlan_xmit_one(skb, vxlan->dev, cfg, src_vni,
 			       rcu_dereference(fremote->rd), false);
 	else
 		kfree_skb_reason(skb, SKB_DROP_REASON_NO_TX_TARGET);
diff --git a/drivers/net/vxlan/vxlan_private.h b/drivers/net/vxlan/vxlan_private.h
index 74d713717cd5da267fc0aef9812593b5ed3ef37e..a2034ad418f7f77f64d8d81cb2312fcc608c7a0f 100644
--- a/drivers/net/vxlan/vxlan_private.h
+++ b/drivers/net/vxlan/vxlan_private.h
@@ -204,9 +204,10 @@ int vxlan_fdb_update(struct vxlan_dev *vxlan,
 		     __u32 ifindex, __u16 ndm_flags, u32 nhid,
 		     bool swdev_notify, struct netlink_ext_ack *extack);
 void vxlan_xmit_one(struct sk_buff *skb, struct net_device *dev,
+		    const struct vxlan_config *cfg,
 		    __be32 default_vni, struct vxlan_rdst *rdst, bool did_rsc);
 int vxlan_vni_in_use(struct net *src_net, struct vxlan_dev *vxlan,
-		     struct vxlan_config *conf, __be32 vni);
+		     const struct vxlan_config *conf, __be32 vni);
 
 /* vxlan_vnifilter.c */
 int vxlan_vnigroup_init(struct vxlan_dev *vxlan);
@@ -214,7 +215,8 @@ void vxlan_vnigroup_uninit(struct vxlan_dev *vxlan);
 
 int vxlan_vnifilter_init(void);
 void vxlan_vnifilter_uninit(void);
-void vxlan_vnifilter_count(struct vxlan_dev *vxlan, __be32 vni,
+void vxlan_vnifilter_count(struct vxlan_dev *vxlan,
+			   const struct vxlan_config *cfg, __be32 vni,
 			   struct vxlan_vni_node *vninode,
 			   int type, unsigned int len);
 
@@ -256,9 +258,11 @@ int vxlan_mdb_del_bulk(struct net_device *dev, struct nlattr *tb[],
 int vxlan_mdb_get(struct net_device *dev, struct nlattr *tb[], u32 portid,
 		  u32 seq, struct netlink_ext_ack *extack);
 struct vxlan_mdb_entry *vxlan_mdb_entry_skb_get(struct vxlan_dev *vxlan,
+						const struct vxlan_config *cfg,
 						struct sk_buff *skb,
 						__be32 src_vni);
 netdev_tx_t vxlan_mdb_xmit(struct vxlan_dev *vxlan,
+			   const struct vxlan_config *cfg,
 			   const struct vxlan_mdb_entry *mdb_entry,
 			   struct sk_buff *skb);
 int vxlan_mdb_init(struct vxlan_dev *vxlan);
diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
index 93fca54749e76da8ae1ad1c15c0461a35d0ee194..dc495d5eceb318be3c29e57d7a8bd431354ccec0 100644
--- a/drivers/net/vxlan/vxlan_vnifilter.c
+++ b/drivers/net/vxlan/vxlan_vnifilter.c
@@ -171,13 +171,14 @@ static void vxlan_vnifilter_stats_add(struct vxlan_vni_node *vninode,
 	u64_stats_update_end(&pstats->syncp);
 }
 
-void vxlan_vnifilter_count(struct vxlan_dev *vxlan, __be32 vni,
+void vxlan_vnifilter_count(struct vxlan_dev *vxlan,
+			   const struct vxlan_config *cfg, __be32 vni,
 			   struct vxlan_vni_node *vninode,
 			   int type, unsigned int len)
 {
 	struct vxlan_vni_node *vnode;
 
-	if (!(vxlan->cfg.flags & VXLAN_F_VNIFILTER))
+	if (!cfg || !(cfg->flags & VXLAN_F_VNIFILTER))
 		return;
 
 	if (vninode) {
-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 net-next 5/8] vxlan: move VXLAN_F_MDB to struct vxlan_dev flags
  2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
                   ` (3 preceding siblings ...)
  2026-09-21 10:01 ` [PATCH v5 net-next 4/8] vxlan: pass vxlan_config pointer to helper functions Eric Dumazet
@ 2026-09-21 10:01 ` Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 6/8] vxlan: convert configuration to RCU protection Eric Dumazet
                   ` (2 subsequent siblings)
  7 siblings, 0 replies; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

VXLAN_F_MDB is an internal runtime state flag indicating whether any
MDB entries are configured on the device, rather than a netlink
configuration attribute.

In preparation for converting vxlan->cfg to an RCU-protected pointer,
move VXLAN_F_MDB from struct vxlan_config to a dedicated 'flags' field
in struct vxlan_dev as VXLAN_DEV_F_MDB, using atomic bitops (set_bit(),
clear_bit(), test_bit()) to avoid KCSAN data races between the TX path
and RTNL operations.

This avoids having to dynamically reallocate and publish a new
vxlan_config structure via RCU whenever the first MDB entry is added
or the last one is removed, and prevents potential memory allocation
failures during MDB teardown under memory pressure.

mlxsw validates cfg->flags against a deny-by-default mask, so
VXLAN_F_MDB used to make mlxsw_sp_nve_vxlan_can_offload() reject a
device with MDB entries as carrying an unsupported flag. Add an
explicit VXLAN_DEV_F_MDB test there to keep that rejection, with a
message naming the actual reason.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c | 5 +++++
 drivers/net/vxlan/vxlan_core.c                           | 2 +-
 drivers/net/vxlan/vxlan_mdb.c                            | 6 +++---
 include/net/vxlan.h                                      | 6 +++++-
 4 files changed, 14 insertions(+), 5 deletions(-)

diff --git a/drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c b/drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c
index 52c2fe3644d4b9b27f1d589d9f7f597748339782..b78aff31c98f2d8aedfdf36bbfc1cb61e1066edf 100644
--- a/drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c
+++ b/drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c
@@ -92,6 +92,11 @@ static bool mlxsw_sp_nve_vxlan_can_offload(const struct mlxsw_sp_nve *nve,
 		return false;
 	}
 
+	if (test_bit(VXLAN_DEV_F_MDB, &vxlan->flags)) {
+		NL_SET_ERR_MSG_MOD(extack, "VxLAN: MDB entries are not supported");
+		return false;
+	}
+
 	switch (cfg->saddr.sa.sa_family) {
 	case AF_INET:
 		if (!mlxsw_sp_nve_vxlan_ipv4_flags_check(cfg, extack))
diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index e834b58ea745ce8394ff2ae6cef05eec3a9f8d26..f98f1d802724df40d4f68ba7282dc377752b1160 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -2824,7 +2824,7 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 	if (nhid)
 		return vxlan_xmit_nhid(skb, dev, nhid, vni, cfg);
 
-	if (flags & VXLAN_F_MDB) {
+	if (test_bit(VXLAN_DEV_F_MDB, &vxlan->flags)) {
 		struct vxlan_mdb_entry *mdb_entry;
 
 		rcu_read_lock();
diff --git a/drivers/net/vxlan/vxlan_mdb.c b/drivers/net/vxlan/vxlan_mdb.c
index 56ca9283283307d8b1231c3b26724474e6c838d6..cf606256d0929c4dd356ec8aa343150c10edfb82 100644
--- a/drivers/net/vxlan/vxlan_mdb.c
+++ b/drivers/net/vxlan/vxlan_mdb.c
@@ -1219,7 +1219,7 @@ vxlan_mdb_entry_get(struct vxlan_dev *vxlan,
 		goto err_free_entry;
 
 	if (hlist_is_singular_node(&mdb_entry->mdb_node, &vxlan->mdb_list))
-		vxlan->cfg.flags |= VXLAN_F_MDB;
+		set_bit(VXLAN_DEV_F_MDB, &vxlan->flags);
 
 	return mdb_entry;
 
@@ -1236,7 +1236,7 @@ static void vxlan_mdb_entry_put(struct vxlan_dev *vxlan,
 		return;
 
 	if (hlist_is_singular_node(&mdb_entry->mdb_node, &vxlan->mdb_list))
-		vxlan->cfg.flags &= ~VXLAN_F_MDB;
+		clear_bit(VXLAN_DEV_F_MDB, &vxlan->flags);
 
 	rhashtable_remove_fast(&vxlan->mdb_tbl, &mdb_entry->rhnode,
 			       vxlan_mdb_rht_params);
@@ -1762,7 +1762,7 @@ void vxlan_mdb_fini(struct vxlan_dev *vxlan)
 	struct vxlan_mdb_flush_desc desc = {};
 
 	vxlan_mdb_flush(vxlan, &desc);
-	WARN_ON_ONCE(vxlan->cfg.flags & VXLAN_F_MDB);
+	WARN_ON_ONCE(test_bit(VXLAN_DEV_F_MDB, &vxlan->flags));
 	rhashtable_free_and_destroy(&vxlan->mdb_tbl, vxlan_mdb_check_empty,
 				    NULL);
 }
diff --git a/include/net/vxlan.h b/include/net/vxlan.h
index 7b82075055237058d231d636698f640c75c521af..d323f91af2364822e148310297c978fec7664010 100644
--- a/include/net/vxlan.h
+++ b/include/net/vxlan.h
@@ -300,6 +300,7 @@ struct vxlan_dev {
 	spinlock_t	  hash_lock;
 	unsigned int	  addrcnt;
 	struct gro_cells  gro_cells;
+	unsigned long	  flags;
 
 	struct vxlan_config	cfg;
 
@@ -313,6 +314,10 @@ struct vxlan_dev {
 	unsigned int mdb_seq;
 };
 
+enum vxlan_dev_flags {
+	VXLAN_DEV_F_MDB,
+};
+
 #define VXLAN_F_LEARN			0x01
 #define VXLAN_F_PROXY			0x02
 #define VXLAN_F_RSC			0x04
@@ -331,7 +336,6 @@ struct vxlan_dev {
 #define VXLAN_F_IPV6_LINKLOCAL		0x8000
 #define VXLAN_F_TTL_INHERIT		0x10000
 #define VXLAN_F_VNIFILTER               0x20000
-#define VXLAN_F_MDB			0x40000
 #define VXLAN_F_LOCALBYPASS		0x80000
 #define VXLAN_F_MC_ROUTE		0x100000
 
-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 net-next 6/8] vxlan: convert configuration to RCU protection
  2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
                   ` (4 preceding siblings ...)
  2026-09-21 10:01 ` [PATCH v5 net-next 5/8] vxlan: move VXLAN_F_MDB to struct vxlan_dev flags Eric Dumazet
@ 2026-09-21 10:01 ` Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 7/8] vxlan: remove default_dst and use vxlan_config and lowerdev Eric Dumazet
  2026-09-21 10:01 ` [PATCH v5 net-next 8/8] vxlan: no longer rely on RTNL in vxlan_fill_info() Eric Dumazet
  7 siblings, 0 replies; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

Move 'struct vxlan_config' from an embedded structure inside
'struct vxlan_dev' to a dynamically allocated RCU-protected pointer
'vxlan->cfg'.

Updating configuration via vxlan_changelink() or vxlan_dev_configure()
allocates a new struct vxlan_config and publishes it with
rcu_assign_pointer(), freeing the previous config with kfree_rcu().

Readers are converted to use rcu_dereference() or rtnl_dereference().
In vxlan_xmit(), acquire rcu_read_lock() to protect config access.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 .../mellanox/mlxsw/spectrum_nve_vxlan.c       |  14 +-
 .../mellanox/mlxsw/spectrum_switchdev.c       |  57 ++--
 drivers/net/vxlan/vxlan_core.c                | 274 ++++++++++++------
 drivers/net/vxlan/vxlan_mdb.c                 |  10 +-
 drivers/net/vxlan/vxlan_multicast.c           |  12 +-
 drivers/net/vxlan/vxlan_vnifilter.c           |  21 +-
 include/net/vxlan.h                           |   3 +-
 7 files changed, 263 insertions(+), 128 deletions(-)

diff --git a/drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c b/drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c
index b78aff31c98f2d8aedfdf36bbfc1cb61e1066edf..21e7b499f3879c050cc68ff024fcc62a809be840 100644
--- a/drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c
+++ b/drivers/net/ethernet/mellanox/mlxsw/spectrum_nve_vxlan.c
@@ -59,8 +59,11 @@ static bool mlxsw_sp_nve_vxlan_can_offload(const struct mlxsw_sp_nve *nve,
 					   const struct mlxsw_sp_nve_params *params,
 					   struct netlink_ext_ack *extack)
 {
-	struct vxlan_dev *vxlan = netdev_priv(params->dev);
-	struct vxlan_config *cfg = &vxlan->cfg;
+	const struct vxlan_config *cfg;
+	struct vxlan_dev *vxlan;
+
+	vxlan = netdev_priv(params->dev);
+	cfg = rtnl_dereference(vxlan->cfg);
 
 	if (vxlan_addr_multicast(&cfg->remote_ip)) {
 		NL_SET_ERR_MSG_MOD(extack, "VxLAN: Multicast destination IP is not supported");
@@ -153,8 +156,11 @@ static void mlxsw_sp_nve_vxlan_config(const struct mlxsw_sp_nve *nve,
 				      const struct mlxsw_sp_nve_params *params,
 				      struct mlxsw_sp_nve_config *config)
 {
-	struct vxlan_dev *vxlan = netdev_priv(params->dev);
-	struct vxlan_config *cfg = &vxlan->cfg;
+	const struct vxlan_config *cfg;
+	struct vxlan_dev *vxlan;
+
+	vxlan = netdev_priv(params->dev);
+	cfg = rtnl_dereference(vxlan->cfg);
 
 	config->type = MLXSW_SP_NVE_TYPE_VXLAN;
 	config->ttl = cfg->ttl;
diff --git a/drivers/net/ethernet/mellanox/mlxsw/spectrum_switchdev.c b/drivers/net/ethernet/mellanox/mlxsw/spectrum_switchdev.c
index fe45e533a4b2efb532b85960009c586a53ade340..de60b698bea982593f0f7550571e9a4f434dcceb 100644
--- a/drivers/net/ethernet/mellanox/mlxsw/spectrum_switchdev.c
+++ b/drivers/net/ethernet/mellanox/mlxsw/spectrum_switchdev.c
@@ -2513,15 +2513,17 @@ mlxsw_sp_bridge_vlan_aware_vxlan_join(struct mlxsw_sp_bridge_device *bridge_devi
 {
 	struct mlxsw_sp *mlxsw_sp = mlxsw_sp_lower_get(bridge_device->dev);
 	struct vxlan_dev *vxlan = netdev_priv(vxlan_dev);
-	struct mlxsw_sp_nve_params params = {
-		.type = MLXSW_SP_NVE_TYPE_VXLAN,
-		.vni = vxlan->cfg.vni,
-		.dev = vxlan_dev,
-		.ethertype = ethertype,
-	};
+	struct mlxsw_sp_nve_params params;
+	const struct vxlan_config *cfg;
 	struct mlxsw_sp_fid *fid;
 	int err;
 
+	cfg = rtnl_dereference(vxlan->cfg);
+	params.type = MLXSW_SP_NVE_TYPE_VXLAN;
+	params.vni = cfg->vni;
+	params.dev = vxlan_dev;
+	params.ethertype = ethertype;
+
 	/* If the VLAN is 0, we need to find the VLAN that is configured as
 	 * PVID and egress untagged on the bridge port of the VxLAN device.
 	 * It is possible no such VLAN exists
@@ -2704,15 +2706,17 @@ mlxsw_sp_bridge_8021d_vxlan_join(struct mlxsw_sp_bridge_device *bridge_device,
 {
 	struct mlxsw_sp *mlxsw_sp = mlxsw_sp_lower_get(bridge_device->dev);
 	struct vxlan_dev *vxlan = netdev_priv(vxlan_dev);
-	struct mlxsw_sp_nve_params params = {
-		.type = MLXSW_SP_NVE_TYPE_VXLAN,
-		.vni = vxlan->cfg.vni,
-		.dev = vxlan_dev,
-		.ethertype = ETH_P_8021Q,
-	};
+	struct mlxsw_sp_nve_params params;
+	const struct vxlan_config *cfg;
 	struct mlxsw_sp_fid *fid;
 	int err;
 
+	cfg = rtnl_dereference(vxlan->cfg);
+	params.type = MLXSW_SP_NVE_TYPE_VXLAN;
+	params.vni = cfg->vni;
+	params.dev = vxlan_dev;
+	params.ethertype = ETH_P_8021Q;
+
 	fid = mlxsw_sp_fid_8021d_get(mlxsw_sp, bridge_device->dev->ifindex);
 	if (IS_ERR(fid)) {
 		NL_SET_ERR_MSG_MOD(extack, "Failed to create 802.1D FID");
@@ -2933,10 +2937,13 @@ static void __mlxsw_sp_bridge_vxlan_leave(struct mlxsw_sp *mlxsw_sp,
 					  const struct net_device *vxlan_dev)
 {
 	struct vxlan_dev *vxlan = netdev_priv(vxlan_dev);
+	const struct vxlan_config *cfg;
 	struct mlxsw_sp_fid *fid;
 
+	cfg = rtnl_dereference(vxlan->cfg);
+
 	/* If the VxLAN device is down, then the FID does not have a VNI */
-	fid = mlxsw_sp_fid_lookup_by_vni(mlxsw_sp, vxlan->cfg.vni);
+	fid = mlxsw_sp_fid_lookup_by_vni(mlxsw_sp, cfg->vni);
 	if (!fid)
 		return;
 
@@ -3029,11 +3036,13 @@ static void mlxsw_sp_fdb_vxlan_call_notifiers(struct net_device *dev,
 	struct switchdev_notifier_vxlan_fdb_info info;
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	enum switchdev_notifier_type type;
+	const struct vxlan_config *cfg;
 
+	cfg = rtnl_dereference(vxlan->cfg);
 	type = adding ? SWITCHDEV_VXLAN_FDB_ADD_TO_BRIDGE :
 			SWITCHDEV_VXLAN_FDB_DEL_TO_BRIDGE;
 	mlxsw_sp_switchdev_addr_vxlan_convert(proto, addr, &info.remote_ip);
-	info.remote_port = vxlan->cfg.dst_port;
+	info.remote_port = cfg->dst_port;
 	info.remote_vni = vni;
 	info.remote_ifindex = 0;
 	ether_addr_copy(info.eth_addr, mac);
@@ -3236,8 +3245,10 @@ __mlxsw_sp_fdb_notify_mac_uc_tunnel_process(struct mlxsw_sp *mlxsw_sp,
 
 	if (adding && netif_is_vxlan(dev)) {
 		struct vxlan_dev *vxlan = netdev_priv(dev);
+		const struct vxlan_config *cfg;
 
-		if (!(vxlan->cfg.flags & VXLAN_F_LEARN))
+		cfg = rtnl_dereference(vxlan->cfg);
+		if (!(cfg->flags & VXLAN_F_LEARN))
 			return -EINVAL;
 	}
 
@@ -3722,9 +3733,11 @@ mlxsw_sp_switchdev_vxlan_work_prepare(struct mlxsw_sp_switchdev_event_work *
 {
 	struct vxlan_dev *vxlan = netdev_priv(switchdev_work->dev);
 	struct switchdev_notifier_vxlan_fdb_info *vxlan_fdb_info;
-	struct vxlan_config *cfg = &vxlan->cfg;
+	const struct vxlan_config *cfg;
 	struct netlink_ext_ack *extack;
 
+	cfg = rcu_dereference_rtnl(vxlan->cfg);
+
 	extack = switchdev_notifier_info_to_extack(info);
 	vxlan_fdb_info = container_of(info,
 				      struct switchdev_notifier_vxlan_fdb_info,
@@ -3851,11 +3864,15 @@ mlxsw_sp_switchdev_vxlan_vlan_add(struct mlxsw_sp *mlxsw_sp,
 				  struct netlink_ext_ack *extack)
 {
 	struct vxlan_dev *vxlan = netdev_priv(vxlan_dev);
-	__be32 vni = vxlan->cfg.vni;
+	const struct vxlan_config *cfg;
 	struct mlxsw_sp_fid *fid;
 	u16 old_vid;
+	__be32 vni;
 	int err;
 
+	cfg = rtnl_dereference(vxlan->cfg);
+	vni = cfg->vni;
+
 	/* We cannot have the same VLAN as PVID and egress untagged on multiple
 	 * VxLAN devices. Note that we get this notification before the VLAN is
 	 * actually added to the bridge's database, so it is not possible for
@@ -3935,12 +3952,16 @@ mlxsw_sp_switchdev_vxlan_vlan_del(struct mlxsw_sp *mlxsw_sp,
 				  const struct net_device *vxlan_dev, u16 vid)
 {
 	struct vxlan_dev *vxlan = netdev_priv(vxlan_dev);
-	__be32 vni = vxlan->cfg.vni;
+	const struct vxlan_config *cfg;
 	struct mlxsw_sp_fid *fid;
+	__be32 vni;
 
 	if (!netif_running(vxlan_dev))
 		return;
 
+	cfg = rtnl_dereference(vxlan->cfg);
+	vni = cfg->vni;
+
 	fid = mlxsw_sp_fid_lookup_by_vni(mlxsw_sp, vni);
 	if (!fid)
 		return;
diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index f98f1d802724df40d4f68ba7282dc377752b1160..04b2925d7c38aafea3c4ab1a1b34ac1feeb93afe 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -110,20 +110,23 @@ static struct vxlan_dev *vxlan_vs_find_vni(struct vxlan_sock *vs,
 		vni = 0;
 
 	hlist_for_each_entry_rcu(node, vni_head(vs, vni), hlist) {
+		const struct vxlan_config *cfg;
+
 		if (!node->vxlan)
 			continue;
+
+		cfg = rcu_dereference(node->vxlan->cfg);
+
 		vnode = NULL;
-		if (node->vxlan->cfg.flags & VXLAN_F_VNIFILTER) {
+		if (cfg->flags & VXLAN_F_VNIFILTER) {
 			vnode = vxlan_vnifilter_lookup(node->vxlan, vni);
 			if (!vnode)
 				continue;
-		} else if (node->vxlan->default_dst.remote_vni != vni) {
+		} else if (cfg->vni != vni) {
 			continue;
 		}
 
 		if (IS_ENABLED(CONFIG_IPV6)) {
-			const struct vxlan_config *cfg = &node->vxlan->cfg;
-
 			if ((cfg->flags & VXLAN_F_IPV6_LINKLOCAL) &&
 			    cfg->remote_ifindex != ifindex)
 				continue;
@@ -157,6 +160,7 @@ static int vxlan_fdb_info(struct sk_buff *skb, struct vxlan_dev *vxlan,
 			  u32 portid, u32 seq, int type, unsigned int flags,
 			  const struct vxlan_rdst *rdst)
 {
+	const struct vxlan_config *cfg = rcu_dereference_rtnl(vxlan->cfg);
 	unsigned long now = jiffies;
 	struct nda_cacheinfo ci;
 	bool send_ip, send_eth;
@@ -216,10 +220,10 @@ static int vxlan_fdb_info(struct sk_buff *skb, struct vxlan_dev *vxlan,
 			goto nla_put_failure;
 
 		if (rdst->remote_port &&
-		    rdst->remote_port != vxlan->cfg.dst_port &&
+		    rdst->remote_port != cfg->dst_port &&
 		    nla_put_be16(skb, NDA_PORT, rdst->remote_port))
 			goto nla_put_failure;
-		if (rdst->remote_vni != vxlan->default_dst.remote_vni &&
+		if (rdst->remote_vni != cfg->vni &&
 		    nla_put_u32(skb, NDA_VNI, be32_to_cpu(rdst->remote_vni)))
 			goto nla_put_failure;
 		if (rdst->remote_ifindex &&
@@ -227,7 +231,7 @@ static int vxlan_fdb_info(struct sk_buff *skb, struct vxlan_dev *vxlan,
 			goto nla_put_failure;
 	}
 
-	if ((vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA) && fdb->key.vni &&
+	if ((cfg->flags & VXLAN_F_COLLECT_METADATA) && fdb->key.vni &&
 	    nla_put_u32(skb, NDA_SRC_VNI,
 			be32_to_cpu(fdb->key.vni)))
 		goto nla_put_failure;
@@ -418,7 +422,7 @@ static struct vxlan_fdb *vxlan_find_mac(struct vxlan_dev *vxlan,
 	lockdep_assert_held_once(&vxlan->hash_lock);
 
 	rcu_read_lock();
-	f = vxlan_find_mac_rcu(vxlan, &vxlan->cfg, mac, vni);
+	f = vxlan_find_mac_rcu(vxlan, rcu_dereference(vxlan->cfg), mac, vni);
 	rcu_read_unlock();
 
 	return f;
@@ -459,7 +463,7 @@ int vxlan_fdb_find_uc(struct net_device *dev, const u8 *mac, __be32 vni,
 
 	rcu_read_lock();
 
-	f = vxlan_find_mac_rcu(vxlan, &vxlan->cfg, eth_addr, vni);
+	f = vxlan_find_mac_rcu(vxlan, rcu_dereference(vxlan->cfg), eth_addr, vni);
 	if (f)
 		rdst = first_remote_rcu(f);
 	if (!rdst) {
@@ -865,12 +869,13 @@ int vxlan_fdb_create(struct vxlan_dev *vxlan,
 		     u32 nhid, struct vxlan_fdb **fdb,
 		     struct netlink_ext_ack *extack)
 {
+	const struct vxlan_config *cfg = rcu_dereference_rtnl(vxlan->cfg);
 	struct vxlan_rdst *rd = NULL;
 	struct vxlan_fdb *f;
 	int rc;
 
-	if (vxlan->cfg.addrmax &&
-	    vxlan->addrcnt >= vxlan->cfg.addrmax)
+	if (cfg->addrmax &&
+	    vxlan->addrcnt >= cfg->addrmax)
 		return -ENOSPC;
 
 	netdev_dbg(vxlan->dev, "add %pM -> %pIS\n", mac, ip);
@@ -1156,6 +1161,7 @@ static int vxlan_fdb_parse(struct nlattr *tb[], struct vxlan_dev *vxlan,
 			   __be32 *vni, u32 *ifindex, u32 *nhid,
 			   struct netlink_ext_ack *extack)
 {
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	struct net *net = dev_net(vxlan->dev);
 	int err;
 
@@ -1172,7 +1178,7 @@ static int vxlan_fdb_parse(struct nlattr *tb[], struct vxlan_dev *vxlan,
 			return err;
 		}
 	} else {
-		union vxlan_addr *remote = &vxlan->default_dst.remote_ip;
+		const union vxlan_addr *remote = &cfg->remote_ip;
 
 		if (remote->sa.sa_family == AF_INET) {
 			ip->sin.sin_addr.s_addr = htonl(INADDR_ANY);
@@ -1192,7 +1198,7 @@ static int vxlan_fdb_parse(struct nlattr *tb[], struct vxlan_dev *vxlan,
 		}
 		*port = nla_get_be16(tb[NDA_PORT]);
 	} else {
-		*port = vxlan->cfg.dst_port;
+		*port = cfg->dst_port;
 	}
 
 	if (tb[NDA_VNI]) {
@@ -1202,7 +1208,7 @@ static int vxlan_fdb_parse(struct nlattr *tb[], struct vxlan_dev *vxlan,
 		}
 		*vni = cpu_to_be32(nla_get_u32(tb[NDA_VNI]));
 	} else {
-		*vni = vxlan->default_dst.remote_vni;
+		*vni = cfg->vni;
 	}
 
 	if (tb[NDA_SRC_VNI]) {
@@ -1212,7 +1218,7 @@ static int vxlan_fdb_parse(struct nlattr *tb[], struct vxlan_dev *vxlan,
 		}
 		*src_vni = cpu_to_be32(nla_get_u32(tb[NDA_SRC_VNI]));
 	} else {
-		*src_vni = vxlan->default_dst.remote_vni;
+		*src_vni = cfg->vni;
 	}
 
 	if (tb[NDA_IFINDEX]) {
@@ -1407,18 +1413,21 @@ static int vxlan_fdb_get(struct sk_buff *skb,
 			 struct netlink_ext_ack *extack)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
+	const struct vxlan_config *cfg;
 	struct vxlan_fdb *f;
 	__be32 vni;
 	int err;
 
+	cfg = rcu_dereference_rtnl(vxlan->cfg);
+
 	if (tb[NDA_VNI])
 		vni = cpu_to_be32(nla_get_u32(tb[NDA_VNI]));
 	else
-		vni = vxlan->default_dst.remote_vni;
+		vni = cfg->vni;
 
 	rcu_read_lock();
 
-	f = vxlan_find_mac_rcu(vxlan, &vxlan->cfg, addr, vni);
+	f = vxlan_find_mac_rcu(vxlan, cfg, addr, vni);
 	if (!f) {
 		NL_SET_ERR_MSG(extack, "Fdb entry not found");
 		err = -ENOENT;
@@ -1521,6 +1530,7 @@ static bool __vxlan_sock_release_prep(struct vxlan_sock *vs)
 
 static void vxlan_sock_release(struct vxlan_dev *vxlan)
 {
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	struct vxlan_sock *sock4 = rtnl_dereference(vxlan->vn4_sock);
 #if IS_ENABLED(CONFIG_IPV6)
 	struct vxlan_sock *sock6 = rtnl_dereference(vxlan->vn6_sock);
@@ -1530,7 +1540,7 @@ static void vxlan_sock_release(struct vxlan_dev *vxlan)
 
 	RCU_INIT_POINTER(vxlan->vn4_sock, NULL);
 
-	if (vxlan->cfg.flags & VXLAN_F_VNIFILTER)
+	if (cfg->flags & VXLAN_F_VNIFILTER)
 		vxlan_vs_del_vnigrp(vxlan);
 	else
 		vxlan_vs_del_dev(vxlan);
@@ -1703,7 +1713,8 @@ static int vxlan_rcv(struct sock *sk, struct sk_buff *skb)
 		goto drop;
 	}
 
-	cfg = &vxlan->cfg;
+	cfg = rcu_dereference(vxlan->cfg);
+
 	if (vh->vx_flags & cfg->reserved_bits.vx_flags ||
 	    vh->vx_vni & cfg->reserved_bits.vx_vni) {
 		/* If the header uses bits besides those enabled by the
@@ -1859,7 +1870,8 @@ static int vxlan_err_lookup(struct sock *sk, struct sk_buff *skb)
 	return 0;
 }
 
-static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni, u32 flags)
+static int arp_reduce(struct net_device *dev, struct sk_buff *skb,
+		      const struct vxlan_config *cfg, __be32 vni)
 {
 	struct neigh_table *tbl = arp_table(dev_net(dev));
 	struct vxlan_dev *vxlan = netdev_priv(dev);
@@ -1873,7 +1885,7 @@ static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni, u
 
 	if (!pskb_network_may_pull(skb, arp_hdr_len(dev))) {
 		dev_dstats_tx_dropped(dev);
-		vxlan_vnifilter_count(vxlan, &vxlan->cfg, vni, NULL,
+		vxlan_vnifilter_count(vxlan, cfg, vni, NULL,
 				      VXLAN_VNI_STATS_TX_DROPS, 0);
 		goto out;
 	}
@@ -1914,7 +1926,7 @@ static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni, u
 		neigh_ha_snapshot(ha, n, n->dev);
 
 		rcu_read_lock();
-		f = vxlan_find_mac_tx(vxlan, &vxlan->cfg, ha, vni);
+		f = vxlan_find_mac_tx(vxlan, cfg, ha, vni);
 		if (f)
 			rdst = first_remote_rcu(f);
 		if (rdst && vxlan_addr_any(&rdst->remote_ip)) {
@@ -1940,11 +1952,11 @@ static int arp_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni, u
 
 		if (netif_rx(reply) == NET_RX_DROP) {
 			dev_dstats_rx_dropped(dev);
-			vxlan_vnifilter_count(vxlan, &vxlan->cfg, vni, NULL,
+			vxlan_vnifilter_count(vxlan, cfg, vni, NULL,
 					      VXLAN_VNI_STATS_RX_DROPS, 0);
 		}
 
-	} else if (flags & VXLAN_F_L3MISS) {
+	} else if (cfg->flags & VXLAN_F_L3MISS) {
 		union vxlan_addr ipa = {
 			.sin.sin_addr.s_addr = tip,
 			.sin.sin_family = AF_INET,
@@ -2052,7 +2064,8 @@ static struct sk_buff *vxlan_na_create(struct sk_buff *request,
 	return reply;
 }
 
-static int neigh_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni, u32 flags)
+static int neigh_reduce(struct net_device *dev, struct sk_buff *skb,
+			const struct vxlan_config *cfg, __be32 vni)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	const struct in6_addr *daddr;
@@ -2086,7 +2099,7 @@ static int neigh_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni,
 		}
 
 		neigh_ha_snapshot(ha, n, n->dev);
-		f = vxlan_find_mac_tx(vxlan, &vxlan->cfg, ha, vni);
+		f = vxlan_find_mac_tx(vxlan, cfg, ha, vni);
 		if (f)
 			rdst = first_remote_rcu(f);
 		if (rdst && vxlan_addr_any(&rdst->remote_ip)) {
@@ -2105,10 +2118,10 @@ static int neigh_reduce(struct net_device *dev, struct sk_buff *skb, __be32 vni,
 
 		if (netif_rx(reply) == NET_RX_DROP) {
 			dev_dstats_rx_dropped(dev);
-			vxlan_vnifilter_count(vxlan, &vxlan->cfg, vni, NULL,
+			vxlan_vnifilter_count(vxlan, cfg, vni, NULL,
 					      VXLAN_VNI_STATS_RX_DROPS, 0);
 		}
-	} else if (flags & VXLAN_F_L3MISS) {
+	} else if (cfg->flags & VXLAN_F_L3MISS) {
 		union vxlan_addr ipa = {
 			.sin6.sin6_addr = msg->target,
 			.sin6.sin6_family = AF_INET6,
@@ -2295,7 +2308,7 @@ static void vxlan_encap_bypass(struct sk_buff *skb, struct vxlan_dev *src_vxlan,
 			       const struct vxlan_config *src_cfg,
 			       __be32 vni, bool snoop)
 {
-	const struct vxlan_config *dst_cfg = &dst_vxlan->cfg;
+	const struct vxlan_config *dst_cfg;
 	union vxlan_addr loopback;
 	unsigned int len = skb->len;
 	struct net_device *dev = dst_vxlan->dev;
@@ -2316,6 +2329,7 @@ static void vxlan_encap_bypass(struct sk_buff *skb, struct vxlan_dev *src_vxlan,
 	}
 
 	rcu_read_lock();
+	dst_cfg = rcu_dereference(dst_vxlan->cfg);
 	if (unlikely(!(dev->flags & IFF_UP))) {
 		kfree_skb_reason(skb, SKB_DROP_REASON_DEV_READY);
 		goto drop;
@@ -2781,7 +2795,8 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 	u32 nhid = 0;
 	u32 flags;
 
-	cfg = &vxlan->cfg;
+	rcu_read_lock();
+	cfg = rcu_dereference(vxlan->cfg);
 	flags = cfg->flags;
 	default_vni = cfg->vni;
 
@@ -2799,14 +2814,19 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 				vxlan_xmit_one(skb, dev, cfg, vni, NULL, false);
 			else
 				kfree_skb_reason(skb, SKB_DROP_REASON_TUNNEL_TXINFO);
+			rcu_read_unlock();
 			return NETDEV_TX_OK;
 		}
 	}
 
 	if (flags & VXLAN_F_PROXY) {
 		eth = eth_hdr(skb);
-		if (ntohs(eth->h_proto) == ETH_P_ARP)
-			return arp_reduce(dev, skb, vni, flags);
+		if (ntohs(eth->h_proto) == ETH_P_ARP) {
+			netdev_tx_t res = arp_reduce(dev, skb, cfg, vni);
+
+			rcu_read_unlock();
+			return res;
+		}
 #if IS_ENABLED(CONFIG_IPV6)
 		else if (ntohs(eth->h_proto) == ETH_P_IPV6 &&
 			 pskb_network_may_pull(skb, sizeof(struct ipv6hdr) +
@@ -2815,32 +2835,36 @@ static netdev_tx_t vxlan_xmit(struct sk_buff *skb, struct net_device *dev)
 			struct nd_msg *m = (struct nd_msg *)(ipv6_hdr(skb) + 1);
 
 			if (m->icmph.icmp6_code == 0 &&
-			    m->icmph.icmp6_type == NDISC_NEIGHBOUR_SOLICITATION)
-				return neigh_reduce(dev, skb, vni, flags);
+			    m->icmph.icmp6_type == NDISC_NEIGHBOUR_SOLICITATION) {
+				netdev_tx_t res = neigh_reduce(dev, skb, cfg, vni);
+
+				rcu_read_unlock();
+				return res;
+			}
 		}
 #endif
 	}
 
-	if (nhid)
-		return vxlan_xmit_nhid(skb, dev, nhid, vni, cfg);
+	if (nhid) {
+		netdev_tx_t res = vxlan_xmit_nhid(skb, dev, nhid, vni, cfg);
+
+		rcu_read_unlock();
+		return res;
+	}
 
 	if (test_bit(VXLAN_DEV_F_MDB, &vxlan->flags)) {
 		struct vxlan_mdb_entry *mdb_entry;
 
-		rcu_read_lock();
 		mdb_entry = vxlan_mdb_entry_skb_get(vxlan, cfg, skb, vni);
 		if (mdb_entry) {
-			netdev_tx_t ret;
+			netdev_tx_t ret = vxlan_mdb_xmit(vxlan, cfg, mdb_entry, skb);
 
-			ret = vxlan_mdb_xmit(vxlan, cfg, mdb_entry, skb);
 			rcu_read_unlock();
 			return ret;
 		}
-		rcu_read_unlock();
 	}
 
 	eth = eth_hdr(skb);
-	rcu_read_lock();
 	f = vxlan_find_mac_tx(vxlan, cfg, eth->h_dest, vni);
 	did_rsc = false;
 
@@ -2899,12 +2923,15 @@ static void vxlan_cleanup(struct timer_list *t)
 {
 	struct vxlan_dev *vxlan = timer_container_of(vxlan, t, age_timer);
 	unsigned long next_timer = jiffies + FDB_AGE_INTERVAL;
+	const struct vxlan_config *cfg;
 	struct vxlan_fdb *f;
 
 	if (!netif_running(vxlan->dev))
 		return;
 
 	rcu_read_lock();
+	cfg = rcu_dereference(vxlan->cfg);
+
 	hlist_for_each_entry_rcu(f, &vxlan->fdb_list, fdb_node) {
 		unsigned long timeout;
 
@@ -2914,7 +2941,7 @@ static void vxlan_cleanup(struct timer_list *t)
 		if (f->flags & NTF_EXT_LEARNED)
 			continue;
 
-		timeout = READ_ONCE(f->updated) + vxlan->cfg.age_interval * HZ;
+		timeout = READ_ONCE(f->updated) + cfg->age_interval * HZ;
 		if (time_before_eq(timeout, jiffies)) {
 			spin_lock(&vxlan->hash_lock);
 			if (!hlist_unhashed(&f->fdb_node)) {
@@ -2958,13 +2985,16 @@ static void vxlan_vs_add_dev(struct vxlan_sock *vs, struct vxlan_dev *vxlan,
 static int vxlan_init(struct net_device *dev)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
+	const struct vxlan_config *cfg;
 	int err;
 
+	cfg = rtnl_dereference(vxlan->cfg);
+
 	err = rhashtable_init(&vxlan->fdb_hash_tbl, &vxlan_fdb_rht_params);
 	if (err)
 		return err;
 
-	if (vxlan->cfg.flags & VXLAN_F_VNIFILTER) {
+	if (cfg->flags & VXLAN_F_VNIFILTER) {
 		err = vxlan_vnigroup_init(vxlan);
 		if (err)
 			goto err_rhashtable_destroy;
@@ -2984,7 +3014,7 @@ static int vxlan_init(struct net_device *dev)
 err_gro_cells_destroy:
 	gro_cells_destroy(&vxlan->gro_cells);
 err_vnigroup_uninit:
-	if (vxlan->cfg.flags & VXLAN_F_VNIFILTER)
+	if (cfg->flags & VXLAN_F_VNIFILTER)
 		vxlan_vnigroup_uninit(vxlan);
 err_rhashtable_destroy:
 	rhashtable_destroy(&vxlan->fdb_hash_tbl);
@@ -2994,10 +3024,13 @@ static int vxlan_init(struct net_device *dev)
 static void vxlan_uninit(struct net_device *dev)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
+	const struct vxlan_config *cfg;
+
+	cfg = rtnl_dereference(vxlan->cfg);
 
 	vxlan_mdb_fini(vxlan);
 
-	if (vxlan->cfg.flags & VXLAN_F_VNIFILTER)
+	if (cfg->flags & VXLAN_F_VNIFILTER)
 		vxlan_vnigroup_uninit(vxlan);
 
 	gro_cells_destroy(&vxlan->gro_cells);
@@ -3009,6 +3042,7 @@ static void vxlan_uninit(struct net_device *dev)
 static int vxlan_open(struct net_device *dev)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
+	const struct vxlan_config *cfg;
 	int ret;
 
 	ret = vxlan_sock_add(vxlan);
@@ -3021,7 +3055,8 @@ static int vxlan_open(struct net_device *dev)
 		return ret;
 	}
 
-	if (vxlan->cfg.age_interval)
+	cfg = rtnl_dereference(vxlan->cfg);
+	if (cfg->age_interval)
 		mod_timer(&vxlan->age_timer, jiffies + FDB_AGE_INTERVAL);
 
 	return ret;
@@ -3043,8 +3078,10 @@ struct vxlan_fdb_flush_desc {
 static bool vxlan_fdb_is_default_entry(const struct vxlan_fdb *f,
 				       const struct vxlan_dev *vxlan)
 {
+	const struct vxlan_config *cfg = rcu_dereference_rtnl(vxlan->cfg);
+
 	return is_zero_ether_addr(f->key.eth_addr) &&
-	       f->key.vni == vxlan->cfg.vni;
+	       f->key.vni == cfg->vni;
 }
 
 static bool vxlan_fdb_nhid_matches(const struct vxlan_fdb *f, u32 nhid)
@@ -3262,14 +3299,18 @@ static int vxlan_change_mtu(struct net_device *dev, int new_mtu)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct vxlan_rdst *dst = &vxlan->default_dst;
-	struct net_device *lowerdev = __dev_get_by_index(vxlan->net,
-							 dst->remote_ifindex);
+	const struct vxlan_config *cfg;
+	struct net_device *lowerdev;
+
+	cfg = rtnl_dereference(vxlan->cfg);
+
+	lowerdev = __dev_get_by_index(vxlan->net, dst->remote_ifindex);
 
 	/* This check is different than dev->max_mtu, because it looks at
 	 * the lowerdev->mtu, rather than the static dev->max_mtu
 	 */
 	if (lowerdev) {
-		int max_mtu = lowerdev->mtu - vxlan_headroom(vxlan->cfg.flags);
+		int max_mtu = lowerdev->mtu - vxlan_headroom(cfg->flags);
 		if (new_mtu > max_mtu)
 			return -EINVAL;
 	}
@@ -3282,11 +3323,14 @@ static int vxlan_fill_metadata_dst(struct net_device *dev, struct sk_buff *skb)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct ip_tunnel_info *info = skb_tunnel_info(skb);
+	const struct vxlan_config *cfg;
 	__be16 sport, dport;
 
-	sport = udp_flow_src_port(dev_net(dev), skb, vxlan->cfg.port_min,
-				  vxlan->cfg.port_max, true);
-	dport = info->key.tp_dst ? : vxlan->cfg.dst_port;
+	cfg = rcu_dereference(vxlan->cfg);
+
+	sport = udp_flow_src_port(dev_net(dev), skb, cfg->port_min,
+				  cfg->port_max, true);
+	dport = info->key.tp_dst ? : cfg->dst_port;
 
 	if (ip_tunnel_info_af(info) == AF_INET) {
 		struct vxlan_sock *sock4 = rcu_dereference(vxlan->vn4_sock);
@@ -3396,6 +3440,15 @@ static void vxlan_offload_rx_ports(struct net_device *dev, bool push)
 	}
 }
 
+static void vxlan_free_dev(struct net_device *dev)
+{
+	struct vxlan_dev *vxlan = netdev_priv(dev);
+	struct vxlan_config *cfg = rcu_dereference_protected(vxlan->cfg, 1);
+
+	RCU_INIT_POINTER(vxlan->cfg, NULL);
+	kfree(cfg);
+}
+
 /* Initialize the device structure. */
 static void vxlan_setup(struct net_device *dev)
 {
@@ -3404,6 +3457,8 @@ static void vxlan_setup(struct net_device *dev)
 	eth_hw_addr_random(dev);
 	ether_setup(dev);
 
+	dev->priv_destructor = vxlan_free_dev;
+
 	dev->needs_free_netdev = true;
 	SET_NETDEV_DEVTYPE(dev, &vxlan_type);
 
@@ -3686,21 +3741,22 @@ static struct vxlan_sock *vxlan_socket_create(struct net *net, bool ipv6,
 
 static int __vxlan_sock_add(struct vxlan_dev *vxlan, bool ipv6)
 {
-	bool metadata = vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA;
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
+	bool metadata = cfg->flags & VXLAN_F_COLLECT_METADATA;
 	struct vxlan_sock *vs = NULL;
 	struct vxlan_dev_node *node;
 	int l3mdev_index = 0;
 
 	ASSERT_RTNL();
 
-	if (vxlan->cfg.remote_ifindex)
+	if (cfg->remote_ifindex)
 		l3mdev_index = l3mdev_master_upper_ifindex_by_index(
-			vxlan->net, vxlan->cfg.remote_ifindex);
+			vxlan->net, cfg->remote_ifindex);
 
-	if (!vxlan->cfg.no_share) {
+	if (!cfg->no_share) {
 		rcu_read_lock();
 		vs = vxlan_find_sock(vxlan->net, ipv6 ? AF_INET6 : AF_INET,
-				     vxlan->cfg.dst_port, vxlan->cfg.flags,
+				     cfg->dst_port, cfg->flags,
 				     l3mdev_index);
 		if (vs && !refcount_inc_not_zero(&vs->refcnt)) {
 			rcu_read_unlock();
@@ -3710,7 +3766,7 @@ static int __vxlan_sock_add(struct vxlan_dev *vxlan, bool ipv6)
 	}
 	if (!vs)
 		vs = vxlan_socket_create(vxlan->net, ipv6,
-					 vxlan->cfg.dst_port, vxlan->cfg.flags,
+					 cfg->dst_port, cfg->flags,
 					 l3mdev_index);
 	if (IS_ERR(vs))
 		return PTR_ERR(vs);
@@ -3725,7 +3781,7 @@ static int __vxlan_sock_add(struct vxlan_dev *vxlan, bool ipv6)
 		node = &vxlan->hlist4;
 	}
 
-	if (metadata && (vxlan->cfg.flags & VXLAN_F_VNIFILTER))
+	if (metadata && (cfg->flags & VXLAN_F_VNIFILTER))
 		vxlan_vs_add_vnigrp(vxlan, vs, ipv6);
 	else
 		vxlan_vs_add_dev(vs, vxlan, node);
@@ -3735,11 +3791,14 @@ static int __vxlan_sock_add(struct vxlan_dev *vxlan, bool ipv6)
 
 static int vxlan_sock_add(struct vxlan_dev *vxlan)
 {
-	bool metadata = vxlan->cfg.flags & VXLAN_F_COLLECT_METADATA;
-	bool ipv6 = vxlan->cfg.flags & VXLAN_F_IPV6 || metadata;
-	bool ipv4 = !ipv6 || metadata;
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
+	bool metadata, ipv6, ipv4;
 	int ret = 0;
 
+	metadata = cfg->flags & VXLAN_F_COLLECT_METADATA;
+	ipv6 = (cfg->flags & VXLAN_F_IPV6) || metadata;
+	ipv4 = !ipv6 || metadata;
+
 	RCU_INIT_POINTER(vxlan->vn4_sock, NULL);
 #if IS_ENABLED(CONFIG_IPV6)
 	RCU_INIT_POINTER(vxlan->vn6_sock, NULL);
@@ -3763,22 +3822,27 @@ int vxlan_vni_in_use(struct net *src_net, struct vxlan_dev *vxlan,
 	struct vxlan_dev *tmp;
 
 	list_for_each_entry(tmp, &vn->vxlan_list, next) {
+		const struct vxlan_config *tmp_cfg;
+
 		if (tmp == vxlan)
 			continue;
-		if (tmp->cfg.flags & VXLAN_F_VNIFILTER) {
+
+		tmp_cfg = rtnl_dereference(tmp->cfg);
+
+		if (tmp_cfg->flags & VXLAN_F_VNIFILTER) {
 			if (!vxlan_vnifilter_lookup(tmp, vni))
 				continue;
-		} else if (tmp->cfg.vni != vni) {
+		} else if (tmp_cfg->vni != vni) {
 			continue;
 		}
-		if (tmp->cfg.dst_port != conf->dst_port)
+		if (tmp_cfg->dst_port != conf->dst_port)
 			continue;
-		if ((tmp->cfg.flags & (VXLAN_F_RCV_FLAGS | VXLAN_F_IPV6)) !=
+		if ((tmp_cfg->flags & (VXLAN_F_RCV_FLAGS | VXLAN_F_IPV6)) !=
 		    (conf->flags & (VXLAN_F_RCV_FLAGS | VXLAN_F_IPV6)))
 			continue;
 
 		if ((conf->flags & VXLAN_F_IPV6_LINKLOCAL) &&
-		    tmp->cfg.remote_ifindex != conf->remote_ifindex)
+		    tmp_cfg->remote_ifindex != conf->remote_ifindex)
 			continue;
 
 		return -EEXIST;
@@ -3940,7 +4004,7 @@ static int vxlan_config_validate(struct net *src_net, struct vxlan_config *conf,
 }
 
 static void vxlan_config_apply(struct net_device *dev,
-			       struct vxlan_config *conf,
+			       struct vxlan_config *new_cfg,
 			       struct net_device *lowerdev,
 			       struct net *src_net,
 			       bool changelink)
@@ -3948,8 +4012,9 @@ static void vxlan_config_apply(struct net_device *dev,
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct vxlan_rdst *dst = &vxlan->default_dst;
 	unsigned short needed_headroom = ETH_HLEN;
+	struct vxlan_config *old_cfg;
 	int max_mtu = ETH_MAX_MTU;
-	u32 flags = conf->flags;
+	u32 flags = new_cfg->flags;
 
 	if (!changelink) {
 		if (flags & VXLAN_F_GPE)
@@ -3957,18 +4022,18 @@ static void vxlan_config_apply(struct net_device *dev,
 		else
 			vxlan_ether_setup(dev);
 
-		if (conf->mtu)
-			dev->mtu = conf->mtu;
+		if (new_cfg->mtu)
+			dev->mtu = new_cfg->mtu;
 
 		vxlan->net = src_net;
 	}
 
-	dst->remote_vni = conf->vni;
+	dst->remote_vni = new_cfg->vni;
 
-	memcpy(&dst->remote_ip, &conf->remote_ip, sizeof(conf->remote_ip));
+	memcpy(&dst->remote_ip, &new_cfg->remote_ip, sizeof(new_cfg->remote_ip));
 
 	if (lowerdev) {
-		dst->remote_ifindex = conf->remote_ifindex;
+		dst->remote_ifindex = new_cfg->remote_ifindex;
 
 		netif_inherit_tso_max(dev, lowerdev);
 
@@ -3981,7 +4046,7 @@ static void vxlan_config_apply(struct net_device *dev,
 		if (max_mtu < ETH_MIN_MTU)
 			max_mtu = ETH_MIN_MTU;
 
-		if (!changelink && !conf->mtu)
+		if (!changelink && !new_cfg->mtu)
 			dev->mtu = max_mtu;
 	}
 
@@ -3993,7 +4058,10 @@ static void vxlan_config_apply(struct net_device *dev,
 	needed_headroom += vxlan_headroom(flags);
 	dev->needed_headroom = needed_headroom;
 
-	memcpy(&vxlan->cfg, conf, sizeof(*conf));
+	old_cfg = rtnl_dereference(vxlan->cfg);
+	rcu_assign_pointer(vxlan->cfg, new_cfg);
+	if (old_cfg)
+		kfree_rcu(old_cfg, rcu);
 }
 
 static int vxlan_dev_configure(struct net *src_net, struct net_device *dev,
@@ -4002,13 +4070,18 @@ static int vxlan_dev_configure(struct net *src_net, struct net_device *dev,
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct net_device *lowerdev;
+	struct vxlan_config *new_cfg;
 	int ret;
 
 	ret = vxlan_config_validate(src_net, conf, &lowerdev, vxlan, extack);
 	if (ret)
 		return ret;
 
-	vxlan_config_apply(dev, conf, lowerdev, src_net, false);
+	new_cfg = kmemdup(conf, sizeof(*conf), GFP_KERNEL);
+	if (!new_cfg)
+		return -ENOMEM;
+
+	vxlan_config_apply(dev, new_cfg, lowerdev, src_net, false);
 
 	return 0;
 }
@@ -4020,6 +4093,7 @@ static int vxlan_dev_create(struct net *net, struct net_device *dev,
 	struct vxlan_net *vn = net_generic(net, vxlan_net_id);
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct net_device *remote_dev = NULL;
+	const struct vxlan_config *cfg;
 	struct vxlan_rdst *dst;
 	int err;
 
@@ -4028,11 +4102,15 @@ static int vxlan_dev_create(struct net *net, struct net_device *dev,
 	if (err)
 		return err;
 
+	cfg = rtnl_dereference(vxlan->cfg);
+
 	dev->ethtool_ops = &vxlan_ethtool_ops;
 
 	err = register_netdevice(dev);
-	if (err)
+	if (err) {
+		vxlan_free_dev(dev);
 		return err;
+	}
 
 	if (dst->remote_ifindex) {
 		remote_dev = __dev_get_by_index(net, dst->remote_ifindex);
@@ -4059,7 +4137,7 @@ static int vxlan_dev_create(struct net *net, struct net_device *dev,
 				       &dst->remote_ip,
 				       NUD_REACHABLE | NUD_PERMANENT,
 				       NLM_F_EXCL | NLM_F_CREATE,
-				       vxlan->cfg.dst_port,
+				       cfg->dst_port,
 				       dst->remote_vni,
 				       dst->remote_vni,
 				       dst->remote_ifindex,
@@ -4123,8 +4201,12 @@ static int vxlan_nl2conf(struct nlattr *tb[], struct nlattr *data[],
 	memset(conf, 0, sizeof(*conf));
 
 	/* if changelink operation, start with old existing cfg */
-	if (changelink)
-		memcpy(conf, &vxlan->cfg, sizeof(*conf));
+	if (changelink) {
+		const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
+
+		if (cfg)
+			memcpy(conf, cfg, sizeof(*conf));
+	}
 
 	if (data[IFLA_VXLAN_ID]) {
 		__be32 vni = cpu_to_be32(nla_get_u32(data[IFLA_VXLAN_ID]));
@@ -4471,9 +4553,11 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 			    struct netlink_ext_ack *extack)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	bool rem_ip_changed, change_igmp;
 	struct net_device *lowerdev;
 	struct vxlan_config conf;
+	struct vxlan_config *new_cfg;
 	struct vxlan_rdst *dst;
 	u32 new_ifindex;
 	int err;
@@ -4491,13 +4575,19 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 	if (err)
 		return err;
 
+	new_cfg = kmemdup(&conf, sizeof(conf), GFP_KERNEL);
+	if (!new_cfg)
+		return -ENOMEM;
+
 	if (dst->remote_dev == lowerdev)
 		lowerdev = NULL;
 
 	err = netdev_adjacent_change_prepare(dst->remote_dev, lowerdev, dev,
 					     extack);
-	if (err)
+	if (err) {
+		kfree(new_cfg);
 		return err;
+	}
 
 	/* vxlan_config_apply() only commits remote_ifindex if lowerdev is set */
 	new_ifindex = lowerdev ? conf.remote_ifindex : dst->remote_ifindex;
@@ -4515,7 +4605,7 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 					       &conf.remote_ip,
 					       NUD_REACHABLE | NUD_PERMANENT,
 					       NLM_F_APPEND | NLM_F_CREATE,
-					       vxlan->cfg.dst_port,
+					       cfg->dst_port,
 					       conf.vni, conf.vni,
 					       new_ifindex,
 					       NTF_SELF, 0, true, extack);
@@ -4523,13 +4613,14 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 				spin_unlock_bh(&vxlan->hash_lock);
 				netdev_adjacent_change_abort(dst->remote_dev,
 							     lowerdev, dev);
+				kfree(new_cfg);
 				return err;
 			}
 		}
 		if (!vxlan_addr_any(&dst->remote_ip))
 			__vxlan_fdb_delete(vxlan, all_zeros_mac,
 					   dst->remote_ip,
-					   vxlan->cfg.dst_port,
+					   cfg->dst_port,
 					   dst->remote_vni,
 					   dst->remote_vni,
 					   dst->remote_ifindex,
@@ -4539,7 +4630,7 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 		/* If vni filtering device, also update default fdb entries of
 		 * all vnis
 		 */
-		if (vxlan->cfg.flags & VXLAN_F_VNIFILTER) {
+		if (cfg->flags & VXLAN_F_VNIFILTER) {
 			err = vxlan_vnilist_update_group(vxlan, &dst->remote_ip,
 							 &conf.remote_ip,
 							 dst->remote_ifindex,
@@ -4553,6 +4644,7 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 							       NULL);
 				netdev_adjacent_change_abort(dst->remote_dev,
 							     lowerdev, dev);
+				kfree(new_cfg);
 				return err;
 			}
 		}
@@ -4560,20 +4652,20 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 
 	if (change_igmp &&
 	    (vxlan_addr_multicast(&dst->remote_ip) ||
-	     (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
+	     (cfg->flags & VXLAN_F_VNIFILTER)))
 		err = vxlan_multicast_leave(vxlan);
 
-	if (netif_running(dev) && conf.age_interval != vxlan->cfg.age_interval)
+	if (netif_running(dev) && conf.age_interval != cfg->age_interval)
 		mod_timer(&vxlan->age_timer, jiffies);
 
 	netdev_adjacent_change_commit(dst->remote_dev, lowerdev, dev);
 	if (lowerdev && lowerdev != dst->remote_dev)
 		dst->remote_dev = lowerdev;
-	vxlan_config_apply(dev, &conf, lowerdev, vxlan->net, true);
+	vxlan_config_apply(dev, new_cfg, lowerdev, vxlan->net, true);
 
 	if (!err && change_igmp &&
 	    (vxlan_addr_multicast(&dst->remote_ip) ||
-	     (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
+	     (new_cfg->flags & VXLAN_F_VNIFILTER)))
 		err = vxlan_multicast_join(vxlan);
 
 	return err;
@@ -4637,7 +4729,7 @@ static int vxlan_fill_info(struct sk_buff *skb, const struct net_device *dev)
 	struct ifla_vxlan_port_range ports;
 	const struct vxlan_config *cfg;
 
-	cfg = &vxlan->cfg;
+	cfg = rtnl_dereference(vxlan->cfg);
 
 	if (nla_put_u32(skb, IFLA_VXLAN_ID, be32_to_cpu(dst->remote_vni)))
 		goto nla_put_failure;
diff --git a/drivers/net/vxlan/vxlan_mdb.c b/drivers/net/vxlan/vxlan_mdb.c
index cf606256d0929c4dd356ec8aa343150c10edfb82..4ae6369ed4e35e307565f91d0706970f460af552 100644
--- a/drivers/net/vxlan/vxlan_mdb.c
+++ b/drivers/net/vxlan/vxlan_mdb.c
@@ -165,7 +165,7 @@ static int vxlan_mdb_entry_info_fill(const struct vxlan_dev *vxlan,
 				     const struct vxlan_mdb_entry *mdb_entry,
 				     const struct vxlan_mdb_remote *remote)
 {
-	const struct vxlan_config *cfg = &vxlan->cfg;
+	const struct vxlan_config *cfg = rcu_dereference_rtnl(vxlan->cfg);
 	struct vxlan_rdst *rd = rtnl_dereference(remote->rd);
 	struct br_mdb_entry e;
 	struct nlattr *nest;
@@ -614,7 +614,9 @@ static int vxlan_mdb_config_init(struct vxlan_mdb_config *cfg,
 {
 	struct br_mdb_entry *entry = nla_data(tb[MDBA_SET_ENTRY]);
 	struct vxlan_dev *vxlan = netdev_priv(dev);
-	const struct vxlan_config *vcfg = &vxlan->cfg;
+	const struct vxlan_config *vcfg;
+
+	vcfg = rtnl_dereference(vxlan->cfg);
 
 	memset(cfg, 0, sizeof(*cfg));
 	cfg->vxlan = vxlan;
@@ -959,12 +961,12 @@ vxlan_mdb_nlmsg_remote_size(const struct vxlan_dev *vxlan,
 			    const struct vxlan_mdb_entry *mdb_entry,
 			    const struct vxlan_mdb_remote *remote)
 {
-	const struct vxlan_config *cfg = &vxlan->cfg;
+	const struct vxlan_config *cfg = rcu_dereference_rtnl(vxlan->cfg);
 	const struct vxlan_mdb_entry_key *group = &mdb_entry->key;
 	struct vxlan_rdst *rd = rtnl_dereference(remote->rd);
 	size_t nlmsg_size;
 
-		     /* MDBA_MDB_ENTRY_INFO */
+	/* MDBA_MDB_ENTRY_INFO */
 	nlmsg_size = nla_total_size(sizeof(struct br_mdb_entry)) +
 		     /* MDBA_MDB_EATTR_TIMER */
 		     nla_total_size(sizeof(u32));
diff --git a/drivers/net/vxlan/vxlan_multicast.c b/drivers/net/vxlan/vxlan_multicast.c
index 3b75b48dc726df40cebb233095a8a046ee274c30..e2cf10da274f1b608d8bb5020d2b87ebfedeff46 100644
--- a/drivers/net/vxlan/vxlan_multicast.c
+++ b/drivers/net/vxlan/vxlan_multicast.c
@@ -147,6 +147,8 @@ bool vxlan_group_used(struct vxlan_net *vn, struct vxlan_dev *dev,
 #endif
 
 	list_for_each_entry(vxlan, &vn->vxlan_list, next) {
+		const struct vxlan_config *cfg;
+
 		if (!netif_running(vxlan->dev) || vxlan == dev)
 			continue;
 
@@ -158,7 +160,9 @@ bool vxlan_group_used(struct vxlan_net *vn, struct vxlan_dev *dev,
 		    rtnl_dereference(vxlan->vn6_sock) != sock6)
 			continue;
 #endif
-		if (vxlan->cfg.flags & VXLAN_F_VNIFILTER) {
+		cfg = rtnl_dereference(vxlan->cfg);
+
+		if (cfg->flags & VXLAN_F_VNIFILTER) {
 			if (!vxlan_group_used_by_vnifilter(vxlan, ip, ifindex))
 				continue;
 		} else {
@@ -233,6 +237,7 @@ static int vxlan_multicast_leave_vnigrp(struct vxlan_dev *vxlan)
 
 int vxlan_multicast_join(struct vxlan_dev *vxlan)
 {
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	int ret = 0;
 
 	if (vxlan_addr_multicast(&vxlan->default_dst.remote_ip)) {
@@ -244,7 +249,7 @@ int vxlan_multicast_join(struct vxlan_dev *vxlan)
 			return ret;
 	}
 
-	if (vxlan->cfg.flags & VXLAN_F_VNIFILTER)
+	if (cfg->flags & VXLAN_F_VNIFILTER)
 		return vxlan_multicast_join_vnigrp(vxlan);
 
 	return 0;
@@ -252,6 +257,7 @@ int vxlan_multicast_join(struct vxlan_dev *vxlan)
 
 int vxlan_multicast_leave(struct vxlan_dev *vxlan)
 {
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	struct vxlan_net *vn = net_generic(vxlan->net, vxlan_net_id);
 	int ret = 0;
 
@@ -263,7 +269,7 @@ int vxlan_multicast_leave(struct vxlan_dev *vxlan)
 			return ret;
 	}
 
-	if (vxlan->cfg.flags & VXLAN_F_VNIFILTER)
+	if (cfg->flags & VXLAN_F_VNIFILTER)
 		return vxlan_multicast_leave_vnigrp(vxlan);
 
 	return 0;
diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
index dc495d5eceb318be3c29e57d7a8bd431354ccec0..678565c4f8e277f31a485f8f7a2b79e1054ac530 100644
--- a/drivers/net/vxlan/vxlan_vnifilter.c
+++ b/drivers/net/vxlan/vxlan_vnifilter.c
@@ -178,7 +178,7 @@ void vxlan_vnifilter_count(struct vxlan_dev *vxlan,
 {
 	struct vxlan_vni_node *vnode;
 
-	if (!cfg || !(cfg->flags & VXLAN_F_VNIFILTER))
+	if (!(cfg->flags & VXLAN_F_VNIFILTER))
 		return;
 
 	if (vninode) {
@@ -337,6 +337,7 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
 	struct vxlan_vni_node *v, *vbegin = NULL, *vend = NULL;
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct tunnel_msg *new_tmsg, *tmsg;
+	const struct vxlan_config *cfg;
 	struct vxlan_vni_group *vg;
 	struct nlmsghdr *nlh;
 	int idx = 0, s_idx;
@@ -349,7 +350,8 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
 	}
 	s_idx = cb->args[1];
 
-	if (!(vxlan->cfg.flags & VXLAN_F_VNIFILTER)) {
+	cfg = rcu_dereference(vxlan->cfg);
+	if (!(cfg->flags & VXLAN_F_VNIFILTER)) {
 		cb->args[1] = 0;
 		cb->args[2] = 0;
 		return -EINVAL;
@@ -506,6 +508,7 @@ int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
 				   u32 old_ifindex, u32 new_ifindex,
 				   struct netlink_ext_ack *extack)
 {
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	int err = 0;
 
 	if (old_remote_ip && remote_ip &&
@@ -519,7 +522,7 @@ int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
 				       remote_ip,
 				       NUD_REACHABLE | NUD_PERMANENT,
 				       NLM_F_APPEND | NLM_F_CREATE,
-				       vxlan->cfg.dst_port,
+				       cfg->dst_port,
 				       vni,
 				       vni,
 				       new_ifindex,
@@ -533,7 +536,7 @@ int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
 	if (old_remote_ip && !vxlan_addr_any(old_remote_ip)) {
 		__vxlan_fdb_delete(vxlan, all_zeros_mac,
 				   *old_remote_ip,
-				   vxlan->cfg.dst_port,
+				   cfg->dst_port,
 				   vni, vni,
 				   old_ifindex,
 				   true);
@@ -683,6 +686,7 @@ static void vxlan_vni_delete_group(struct vxlan_dev *vxlan,
 				   struct vxlan_vni_node *vninode)
 {
 	struct vxlan_net *vn = net_generic(vxlan->net, vxlan_net_id);
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	struct vxlan_rdst *dst = &vxlan->default_dst;
 
 	/* if per vni remote_ip not present, delete the
@@ -694,7 +698,7 @@ static void vxlan_vni_delete_group(struct vxlan_dev *vxlan,
 		__vxlan_fdb_delete(vxlan, all_zeros_mac,
 				   (vxlan_addr_any(&vninode->remote_ip) ?
 				   dst->remote_ip : vninode->remote_ip),
-				   vxlan->cfg.dst_port,
+				   cfg->dst_port,
 				   vninode->vni, vninode->vni,
 				   dst->remote_ifindex,
 				   true);
@@ -798,6 +802,7 @@ static int vxlan_vni_add(struct vxlan_dev *vxlan,
 			 u32 vni, union vxlan_addr *group,
 			 struct netlink_ext_ack *extack)
 {
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	struct vxlan_vni_node *vninode;
 	__be32 v = cpu_to_be32(vni);
 	bool changed = false;
@@ -806,7 +811,7 @@ static int vxlan_vni_add(struct vxlan_dev *vxlan,
 	if (vxlan_vnifilter_lookup(vxlan, v))
 		return vxlan_vni_update(vxlan, vg, v, group, &changed, extack);
 
-	err = vxlan_vni_in_use(vxlan->net, vxlan, &vxlan->cfg, v);
+	err = vxlan_vni_in_use(vxlan->net, vxlan, cfg, v);
 	if (err) {
 		NL_SET_ERR_MSG(extack, "VNI in use");
 		return err;
@@ -1015,6 +1020,7 @@ static int vxlan_vnifilter_process(struct sk_buff *skb, struct nlmsghdr *nlh,
 				   struct netlink_ext_ack *extack)
 {
 	struct net *net = sock_net(skb->sk);
+	const struct vxlan_config *cfg;
 	struct tunnel_msg *tmsg;
 	struct vxlan_dev *vxlan;
 	struct net_device *dev;
@@ -1039,8 +1045,9 @@ static int vxlan_vnifilter_process(struct sk_buff *skb, struct nlmsghdr *nlh,
 	}
 
 	vxlan = netdev_priv(dev);
+	cfg = rtnl_dereference(vxlan->cfg);
 
-	if (!(vxlan->cfg.flags & VXLAN_F_VNIFILTER))
+	if (!(cfg->flags & VXLAN_F_VNIFILTER))
 		return -EOPNOTSUPP;
 
 	nlmsg_for_each_attr_type(attr, VXLAN_VNIFILTER_ENTRY, nlh,
diff --git a/include/net/vxlan.h b/include/net/vxlan.h
index d323f91af2364822e148310297c978fec7664010..7ced743ec8816d412bb14ec7ee7b422e97c38895 100644
--- a/include/net/vxlan.h
+++ b/include/net/vxlan.h
@@ -229,6 +229,7 @@ struct vxlan_config {
 	bool				no_share;
 	enum ifla_vxlan_df		df;
 	struct vxlanhdr			reserved_bits;
+	struct rcu_head			rcu;
 };
 
 enum {
@@ -302,7 +303,7 @@ struct vxlan_dev {
 	struct gro_cells  gro_cells;
 	unsigned long	  flags;
 
-	struct vxlan_config	cfg;
+	struct vxlan_config __rcu	*cfg;
 
 	struct vxlan_vni_group  __rcu *vnigrp;
 
-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 net-next 7/8] vxlan: remove default_dst and use vxlan_config and lowerdev
  2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
                   ` (5 preceding siblings ...)
  2026-09-21 10:01 ` [PATCH v5 net-next 6/8] vxlan: convert configuration to RCU protection Eric Dumazet
@ 2026-09-21 10:01 ` Eric Dumazet
  2026-09-22 16:02   ` netdev-bot+sashiko
  2026-09-21 10:01 ` [PATCH v5 net-next 8/8] vxlan: no longer rely on RTNL in vxlan_fill_info() Eric Dumazet
  7 siblings, 1 reply; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

Now that vxlan->cfg is an RCU-protected pointer, storing default
destination attributes (remote_ip, remote_vni, remote_ifindex) in
vxlan->default_dst is redundant and creates potential data races for
lockless readers.

Furthermore, several fields of struct vxlan_rdst (remote_port,
offloaded, list, rcu, dst_cache) in default_dst were completely unused.
The remaining one, remote_dev, only tracked the lower device, a role
now taken by vxlan->lowerdev, so drop it from struct vxlan_rdst.

Replace vxlan->default_dst with a 'struct net_device *lowerdev' pointer
in struct vxlan_dev to track adjacent upper/lower netdev topology under
RTNL, and switch all remaining users over to reading configuration
attributes from vxlan->cfg.

Also update mlx5e_tc_tun_get_remote_ifindex() to read remote_ifindex
from vxlan->cfg under rcu_read_lock().

In vxlan_changelink(), pass lowerdev to vxlan_config_apply() to preserve
needed_headroom and needed_tailroom. The mtu is still only reconsidered
when the lower device changes, so that an unrelated changelink can not
shrink it.

cfg->remote_ifindex is now committed unconditionally, where
default_dst.remote_ifindex was left untouched when IFLA_VXLAN_LINK was
cleared, so a changelink dropping the lower device no longer leaves a
stale ifindex behind. For the same reason vxlan_changelink() no longer
needs to compute the ifindex that will actually be committed.

Setting IFLA_VXLAN_LINK to 0 now also tears down the upper/lower
adjacency, which netdev_adjacent_change_commit() used to skip for a NULL
new device, resets needed_tailroom, and is rejected if per-VNI multicast
groups are configured.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 .../mellanox/mlx5/core/en/tc_tun_vxlan.c      |  11 +-
 drivers/net/vxlan/vxlan_core.c                | 201 ++++++++++--------
 drivers/net/vxlan/vxlan_mdb.c                 |  14 +-
 drivers/net/vxlan/vxlan_multicast.c           |  64 +++---
 drivers/net/vxlan/vxlan_private.h             |  15 +-
 drivers/net/vxlan/vxlan_vnifilter.c           |  57 +++--
 include/net/vxlan.h                           |   3 +-
 7 files changed, 207 insertions(+), 158 deletions(-)

diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_vxlan.c b/drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_vxlan.c
index 7a18a469961db809890d69f7d6d8bc656e560946..467fbe43b89e9bc3d28047a3a17a84495c3875af 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_vxlan.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/en/tc_tun_vxlan.c
@@ -241,9 +241,16 @@ static bool mlx5e_tc_tun_encap_info_equal_vxlan(struct mlx5e_encap_key *a,
 static int mlx5e_tc_tun_get_remote_ifindex(struct net_device *mirred_dev)
 {
 	const struct vxlan_dev *vxlan = netdev_priv(mirred_dev);
-	const struct vxlan_rdst *dst = &vxlan->default_dst;
+	const struct vxlan_config *cfg;
+	int ifindex = 0;
 
-	return dst->remote_ifindex;
+	rcu_read_lock();
+	cfg = rcu_dereference(vxlan->cfg);
+	if (cfg)
+		ifindex = cfg->remote_ifindex;
+	rcu_read_unlock();
+
+	return ifindex;
 }
 
 struct mlx5e_tc_tunnel vxlan_tunnel = {
diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index 04b2925d7c38aafea3c4ab1a1b34ac1feeb93afe..b505fefcef70ae9e1704eb5ad5f50a3d5d7488cf 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -804,6 +804,7 @@ static int vxlan_fdb_nh_update(struct vxlan_dev *vxlan, struct vxlan_fdb *fdb,
 			       u32 nhid, struct netlink_ext_ack *extack)
 {
 	struct nexthop *old_nh = rtnl_dereference(fdb->nh);
+	const struct vxlan_config *cfg;
 	struct nexthop *nh;
 	int err = -EINVAL;
 
@@ -832,7 +833,8 @@ static int vxlan_fdb_nh_update(struct vxlan_dev *vxlan, struct vxlan_fdb *fdb,
 	}
 
 	/* check nexthop group family */
-	switch (vxlan->default_dst.remote_ip.sa.sa_family) {
+	cfg = rtnl_dereference(vxlan->cfg);
+	switch (cfg->remote_ip.sa.sa_family) {
 	case AF_INET:
 		if (!nexthop_has_v4(nh)) {
 			err = -EAFNOSUPPORT;
@@ -1249,6 +1251,7 @@ static int vxlan_fdb_add(struct ndmsg *ndm, struct nlattr *tb[],
 			 const unsigned char *addr, u16 vid, u16 flags,
 			 bool *notified, struct netlink_ext_ack *extack)
 {
+	const struct vxlan_config *cfg;
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	/* struct net *net = dev_net(vxlan->dev); */
 	union vxlan_addr ip;
@@ -1276,7 +1279,8 @@ static int vxlan_fdb_add(struct ndmsg *ndm, struct nlattr *tb[],
 		return -EINVAL;
 	}
 
-	if (vxlan->default_dst.remote_ip.sa.sa_family != ip.sa.sa_family)
+	cfg = rtnl_dereference(vxlan->cfg);
+	if (cfg->remote_ip.sa.sa_family != ip.sa.sa_family)
 		return -EAFNOSUPPORT;
 
 	spin_lock_bh(&vxlan->hash_lock);
@@ -2318,7 +2322,14 @@ static void vxlan_encap_bypass(struct sk_buff *skb, struct vxlan_dev *src_vxlan,
 	skb->dev = dev;
 	__skb_pull(skb, skb_network_offset(skb));
 
-	if (dst_vxlan->default_dst.remote_ip.sa.sa_family == AF_INET) {
+	rcu_read_lock();
+	dst_cfg = rcu_dereference(dst_vxlan->cfg);
+	if (unlikely(!(dev->flags & IFF_UP))) {
+		kfree_skb_reason(skb, SKB_DROP_REASON_DEV_READY);
+		goto drop;
+	}
+
+	if (dst_cfg->remote_ip.sa.sa_family == AF_INET) {
 		loopback.sin.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
 		loopback.sa.sa_family =  AF_INET;
 #if IS_ENABLED(CONFIG_IPV6)
@@ -2328,13 +2339,6 @@ static void vxlan_encap_bypass(struct sk_buff *skb, struct vxlan_dev *src_vxlan,
 #endif
 	}
 
-	rcu_read_lock();
-	dst_cfg = rcu_dereference(dst_vxlan->cfg);
-	if (unlikely(!(dev->flags & IFF_UP))) {
-		kfree_skb_reason(skb, SKB_DROP_REASON_DEV_READY);
-		goto drop;
-	}
-
 	if ((dst_cfg->flags & VXLAN_F_LEARN) && snoop)
 		vxlan_snoop(dev, dst_cfg, &loopback, eth_hdr(skb)->h_source, 0, vni);
 
@@ -2973,10 +2977,14 @@ static void vxlan_vs_del_dev(struct vxlan_dev *vxlan)
 static void vxlan_vs_add_dev(struct vxlan_sock *vs, struct vxlan_dev *vxlan,
 			     struct vxlan_dev_node *node)
 {
-	__be32 vni = vxlan->default_dst.remote_vni;
+	const struct vxlan_config *cfg;
+	__be32 vni;
 
 	ASSERT_RTNL();
 
+	cfg = rtnl_dereference(vxlan->cfg);
+	vni = cfg->vni;
+
 	node->vxlan = vxlan;
 	hlist_add_head_rcu(&node->hlist, vni_head(vs, vni));
 }
@@ -3298,13 +3306,12 @@ static void vxlan_set_multicast_list(struct net_device *dev)
 static int vxlan_change_mtu(struct net_device *dev, int new_mtu)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
-	struct vxlan_rdst *dst = &vxlan->default_dst;
 	const struct vxlan_config *cfg;
 	struct net_device *lowerdev;
 
 	cfg = rtnl_dereference(vxlan->cfg);
 
-	lowerdev = __dev_get_by_index(vxlan->net, dst->remote_ifindex);
+	lowerdev = __dev_get_by_index(vxlan->net, cfg->remote_ifindex);
 
 	/* This check is different than dev->max_mtu, because it looks at
 	 * the lowerdev->mtu, rather than the static dev->max_mtu
@@ -3633,9 +3640,11 @@ static int vxlan_get_link_ksettings(struct net_device *dev,
 				    struct ethtool_link_ksettings *cmd)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
-	struct vxlan_rdst *dst = &vxlan->default_dst;
-	struct net_device *lowerdev = __dev_get_by_index(vxlan->net,
-							 dst->remote_ifindex);
+	const struct vxlan_config *cfg;
+	struct net_device *lowerdev;
+
+	cfg = rtnl_dereference(vxlan->cfg);
+	lowerdev = __dev_get_by_index(vxlan->net, cfg->remote_ifindex);
 
 	if (!lowerdev) {
 		cmd->base.duplex = DUPLEX_UNKNOWN;
@@ -3973,6 +3982,13 @@ static int vxlan_config_validate(struct net *src_net, struct vxlan_config *conf,
 			return -EINVAL;
 		}
 
+		if ((conf->flags & VXLAN_F_VNIFILTER) && old &&
+		    vxlan_vnifilter_has_multicast(old)) {
+			NL_SET_ERR_MSG(extack,
+				       "Local interface required for multicast remote group");
+			return -EINVAL;
+		}
+
 #if IS_ENABLED(CONFIG_IPV6)
 		if (conf->flags & VXLAN_F_IPV6_LINKLOCAL) {
 			NL_SET_ERR_MSG(extack,
@@ -4007,10 +4023,9 @@ static void vxlan_config_apply(struct net_device *dev,
 			       struct vxlan_config *new_cfg,
 			       struct net_device *lowerdev,
 			       struct net *src_net,
-			       bool changelink)
+			       bool changelink, bool lowerdev_changed)
 {
 	struct vxlan_dev *vxlan = netdev_priv(dev);
-	struct vxlan_rdst *dst = &vxlan->default_dst;
 	unsigned short needed_headroom = ETH_HLEN;
 	struct vxlan_config *old_cfg;
 	int max_mtu = ETH_MAX_MTU;
@@ -4023,18 +4038,13 @@ static void vxlan_config_apply(struct net_device *dev,
 			vxlan_ether_setup(dev);
 
 		if (new_cfg->mtu)
-			dev->mtu = new_cfg->mtu;
+			WRITE_ONCE(dev->mtu, new_cfg->mtu);
 
 		vxlan->net = src_net;
 	}
 
-	dst->remote_vni = new_cfg->vni;
-
-	memcpy(&dst->remote_ip, &new_cfg->remote_ip, sizeof(new_cfg->remote_ip));
-
+	dev->needed_tailroom = 0;
 	if (lowerdev) {
-		dst->remote_ifindex = new_cfg->remote_ifindex;
-
 		netif_inherit_tso_max(dev, lowerdev);
 
 		needed_headroom = lowerdev->hard_header_len;
@@ -4042,16 +4052,17 @@ static void vxlan_config_apply(struct net_device *dev,
 
 		dev->needed_tailroom = lowerdev->needed_tailroom;
 
-		max_mtu = lowerdev->mtu - vxlan_headroom(flags);
+		max_mtu = READ_ONCE(lowerdev->mtu) - vxlan_headroom(flags);
 		if (max_mtu < ETH_MIN_MTU)
 			max_mtu = ETH_MIN_MTU;
 
 		if (!changelink && !new_cfg->mtu)
-			dev->mtu = max_mtu;
+			WRITE_ONCE(dev->mtu, max_mtu);
 	}
 
-	if (dev->mtu > max_mtu)
-		dev->mtu = max_mtu;
+	/* A changelink leaving the lower device alone must not shrink the mtu */
+	if (lowerdev_changed && READ_ONCE(dev->mtu) > max_mtu)
+		WRITE_ONCE(dev->mtu, max_mtu);
 
 	if (flags & VXLAN_F_COLLECT_METADATA)
 		flags |= VXLAN_F_IPV6;
@@ -4081,7 +4092,7 @@ static int vxlan_dev_configure(struct net *src_net, struct net_device *dev,
 	if (!new_cfg)
 		return -ENOMEM;
 
-	vxlan_config_apply(dev, new_cfg, lowerdev, src_net, false);
+	vxlan_config_apply(dev, new_cfg, lowerdev, src_net, false, true);
 
 	return 0;
 }
@@ -4094,10 +4105,8 @@ static int vxlan_dev_create(struct net *net, struct net_device *dev,
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	struct net_device *remote_dev = NULL;
 	const struct vxlan_config *cfg;
-	struct vxlan_rdst *dst;
 	int err;
 
-	dst = &vxlan->default_dst;
 	err = vxlan_dev_configure(net, dev, conf, extack);
 	if (err)
 		return err;
@@ -4112,8 +4121,8 @@ static int vxlan_dev_create(struct net *net, struct net_device *dev,
 		return err;
 	}
 
-	if (dst->remote_ifindex) {
-		remote_dev = __dev_get_by_index(net, dst->remote_ifindex);
+	if (cfg->remote_ifindex) {
+		remote_dev = __dev_get_by_index(net, cfg->remote_ifindex);
 		if (!remote_dev) {
 			err = -ENODEV;
 			goto unregister;
@@ -4123,7 +4132,7 @@ static int vxlan_dev_create(struct net *net, struct net_device *dev,
 		if (err)
 			goto unregister;
 
-		dst->remote_dev = remote_dev;
+		vxlan->lowerdev = remote_dev;
 	}
 
 	err = rtnl_configure_link(dev, NULL, 0, NULL);
@@ -4131,16 +4140,18 @@ static int vxlan_dev_create(struct net *net, struct net_device *dev,
 		goto unlink;
 
 	/* create an fdb entry for a valid default destination */
-	if (!vxlan_addr_any(&dst->remote_ip)) {
+	if (!vxlan_addr_any(&cfg->remote_ip)) {
+		union vxlan_addr rip = cfg->remote_ip;
+
 		spin_lock_bh(&vxlan->hash_lock);
 		err = vxlan_fdb_update(vxlan, all_zeros_mac,
-				       &dst->remote_ip,
+				       &rip,
 				       NUD_REACHABLE | NUD_PERMANENT,
 				       NLM_F_EXCL | NLM_F_CREATE,
 				       cfg->dst_port,
-				       dst->remote_vni,
-				       dst->remote_vni,
-				       dst->remote_ifindex,
+				       cfg->vni,
+				       cfg->vni,
+				       cfg->remote_ifindex,
 				       NTF_SELF, 0, true, extack);
 		spin_unlock_bh(&vxlan->hash_lock);
 		if (err)
@@ -4552,20 +4563,19 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 			    struct nlattr *data[],
 			    struct netlink_ext_ack *extack)
 {
+	bool lowerdev_changed, rem_ip_changed, change_igmp;
 	struct vxlan_dev *vxlan = netdev_priv(dev);
-	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
-	bool rem_ip_changed, change_igmp;
+	const struct vxlan_config *cfg;
+	struct vxlan_config *new_cfg;
 	struct net_device *lowerdev;
 	struct vxlan_config conf;
-	struct vxlan_config *new_cfg;
-	struct vxlan_rdst *dst;
-	u32 new_ifindex;
 	int err;
 
+	cfg = rtnl_dereference(vxlan->cfg);
+
 	if (!rtnl_dev_link_net_capable(dev, vxlan->net))
 		return -EPERM;
 
-	dst = &vxlan->default_dst;
 	err = vxlan_nl2conf(tb, data, dev, &conf, true, extack);
 	if (err)
 		return err;
@@ -4579,26 +4589,23 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 	if (!new_cfg)
 		return -ENOMEM;
 
-	if (dst->remote_dev == lowerdev)
-		lowerdev = NULL;
-
-	err = netdev_adjacent_change_prepare(dst->remote_dev, lowerdev, dev,
-					     extack);
-	if (err) {
-		kfree(new_cfg);
-		return err;
+	lowerdev_changed = vxlan->lowerdev != lowerdev;
+	if (lowerdev_changed) {
+		err = netdev_adjacent_change_prepare(vxlan->lowerdev, lowerdev,
+						     dev, extack);
+		if (err) {
+			kfree(new_cfg);
+			return err;
+		}
 	}
 
-	/* vxlan_config_apply() only commits remote_ifindex if lowerdev is set */
-	new_ifindex = lowerdev ? conf.remote_ifindex : dst->remote_ifindex;
-
-	rem_ip_changed = !vxlan_addr_equal(&conf.remote_ip, &dst->remote_ip);
+	rem_ip_changed = !vxlan_addr_equal(&conf.remote_ip, &cfg->remote_ip);
 	change_igmp = vxlan->dev->flags & IFF_UP &&
 		      (rem_ip_changed ||
-		       dst->remote_ifindex != new_ifindex);
+		       cfg->remote_ifindex != conf.remote_ifindex);
 
 	/* handle default dst entry */
-	if (rem_ip_changed || dst->remote_ifindex != new_ifindex) {
+	if (rem_ip_changed || cfg->remote_ifindex != conf.remote_ifindex) {
 		spin_lock_bh(&vxlan->hash_lock);
 		if (!vxlan_addr_any(&conf.remote_ip)) {
 			err = vxlan_fdb_update(vxlan, all_zeros_mac,
@@ -4607,23 +4614,24 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 					       NLM_F_APPEND | NLM_F_CREATE,
 					       cfg->dst_port,
 					       conf.vni, conf.vni,
-					       new_ifindex,
+					       conf.remote_ifindex,
 					       NTF_SELF, 0, true, extack);
 			if (err) {
 				spin_unlock_bh(&vxlan->hash_lock);
-				netdev_adjacent_change_abort(dst->remote_dev,
-							     lowerdev, dev);
+				if (lowerdev_changed)
+					netdev_adjacent_change_abort(vxlan->lowerdev,
+								     lowerdev, dev);
 				kfree(new_cfg);
 				return err;
 			}
 		}
-		if (!vxlan_addr_any(&dst->remote_ip))
+		if (!vxlan_addr_any(&cfg->remote_ip))
 			__vxlan_fdb_delete(vxlan, all_zeros_mac,
-					   dst->remote_ip,
+					   cfg->remote_ip,
 					   cfg->dst_port,
-					   dst->remote_vni,
-					   dst->remote_vni,
-					   dst->remote_ifindex,
+					   cfg->vni,
+					   cfg->vni,
+					   cfg->remote_ifindex,
 					   true);
 		spin_unlock_bh(&vxlan->hash_lock);
 
@@ -4631,19 +4639,21 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 		 * all vnis
 		 */
 		if (cfg->flags & VXLAN_F_VNIFILTER) {
-			err = vxlan_vnilist_update_group(vxlan, &dst->remote_ip,
+			err = vxlan_vnilist_update_group(vxlan, &cfg->remote_ip,
 							 &conf.remote_ip,
-							 dst->remote_ifindex,
-							 new_ifindex, extack);
+							 cfg->remote_ifindex,
+							 conf.remote_ifindex,
+							 extack);
 			if (err) {
 				vxlan_update_default_fdb_entry(vxlan, conf.vni,
 							       &conf.remote_ip,
-							       &dst->remote_ip,
-							       new_ifindex,
-							       dst->remote_ifindex,
+							       &cfg->remote_ip,
+							       conf.remote_ifindex,
+							       cfg->remote_ifindex,
 							       NULL);
-				netdev_adjacent_change_abort(dst->remote_dev,
-							     lowerdev, dev);
+				if (lowerdev_changed)
+					netdev_adjacent_change_abort(vxlan->lowerdev,
+								     lowerdev, dev);
 				kfree(new_cfg);
 				return err;
 			}
@@ -4651,20 +4661,26 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
 	}
 
 	if (change_igmp &&
-	    (vxlan_addr_multicast(&dst->remote_ip) ||
+	    (vxlan_addr_multicast(&cfg->remote_ip) ||
 	     (cfg->flags & VXLAN_F_VNIFILTER)))
 		err = vxlan_multicast_leave(vxlan);
 
 	if (netif_running(dev) && conf.age_interval != cfg->age_interval)
 		mod_timer(&vxlan->age_timer, jiffies);
 
-	netdev_adjacent_change_commit(dst->remote_dev, lowerdev, dev);
-	if (lowerdev && lowerdev != dst->remote_dev)
-		dst->remote_dev = lowerdev;
-	vxlan_config_apply(dev, new_cfg, lowerdev, vxlan->net, true);
+	if (lowerdev_changed) {
+		if (lowerdev)
+			netdev_adjacent_change_commit(vxlan->lowerdev, lowerdev,
+						      dev);
+		else
+			netdev_upper_dev_unlink(vxlan->lowerdev, dev);
+		vxlan->lowerdev = lowerdev;
+	}
+	vxlan_config_apply(dev, new_cfg, lowerdev, vxlan->net, true,
+			   lowerdev_changed);
 
 	if (!err && change_igmp &&
-	    (vxlan_addr_multicast(&dst->remote_ip) ||
+	    (vxlan_addr_multicast(&new_cfg->remote_ip) ||
 	     (new_cfg->flags & VXLAN_F_VNIFILTER)))
 		err = vxlan_multicast_join(vxlan);
 
@@ -4680,8 +4696,8 @@ static void vxlan_dellink(struct net_device *dev, struct list_head *head)
 
 	list_del(&vxlan->next);
 	unregister_netdevice_queue(dev, head);
-	if (vxlan->default_dst.remote_dev)
-		netdev_upper_dev_unlink(vxlan->default_dst.remote_dev, dev);
+	if (vxlan->lowerdev)
+		netdev_upper_dev_unlink(vxlan->lowerdev, dev);
 }
 
 static size_t vxlan_get_size(const struct net_device *dev)
@@ -4725,30 +4741,29 @@ static size_t vxlan_get_size(const struct net_device *dev)
 static int vxlan_fill_info(struct sk_buff *skb, const struct net_device *dev)
 {
 	const struct vxlan_dev *vxlan = netdev_priv(dev);
-	const struct vxlan_rdst *dst = &vxlan->default_dst;
 	struct ifla_vxlan_port_range ports;
 	const struct vxlan_config *cfg;
 
 	cfg = rtnl_dereference(vxlan->cfg);
 
-	if (nla_put_u32(skb, IFLA_VXLAN_ID, be32_to_cpu(dst->remote_vni)))
+	if (nla_put_u32(skb, IFLA_VXLAN_ID, be32_to_cpu(cfg->vni)))
 		goto nla_put_failure;
 
-	if (!vxlan_addr_any(&dst->remote_ip)) {
-		if (dst->remote_ip.sa.sa_family == AF_INET) {
+	if (!vxlan_addr_any(&cfg->remote_ip)) {
+		if (cfg->remote_ip.sa.sa_family == AF_INET) {
 			if (nla_put_in_addr(skb, IFLA_VXLAN_GROUP,
-					    dst->remote_ip.sin.sin_addr.s_addr))
+					    cfg->remote_ip.sin.sin_addr.s_addr))
 				goto nla_put_failure;
 #if IS_ENABLED(CONFIG_IPV6)
 		} else {
 			if (nla_put_in6_addr(skb, IFLA_VXLAN_GROUP6,
-					     &dst->remote_ip.sin6.sin6_addr))
+					     &cfg->remote_ip.sin6.sin6_addr))
 				goto nla_put_failure;
 #endif
 		}
 	}
 
-	if (dst->remote_ifindex && nla_put_u32(skb, IFLA_VXLAN_LINK, dst->remote_ifindex))
+	if (cfg->remote_ifindex && nla_put_u32(skb, IFLA_VXLAN_LINK, cfg->remote_ifindex))
 		goto nla_put_failure;
 
 	if (!vxlan_addr_any(&cfg->saddr)) {
@@ -4863,7 +4878,7 @@ static void vxlan_handle_lowerdev_unregister(struct vxlan_net *vn,
 	LIST_HEAD(list_kill);
 
 	list_for_each_entry_safe(vxlan, next, &vn->vxlan_list, next) {
-		struct vxlan_rdst *dst = &vxlan->default_dst;
+		const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 
 		/* In case we created vxlan device with carrier
 		 * and we loose the carrier due to module unload
@@ -4871,7 +4886,7 @@ static void vxlan_handle_lowerdev_unregister(struct vxlan_net *vn,
 		 * cases, it's not necessary and remote_ifindex
 		 * is 0 here, so no matches.
 		 */
-		if (dst->remote_ifindex == dev->ifindex)
+		if (cfg->remote_ifindex == dev->ifindex)
 			vxlan_dellink(vxlan->dev, &list_kill);
 	}
 
diff --git a/drivers/net/vxlan/vxlan_mdb.c b/drivers/net/vxlan/vxlan_mdb.c
index 4ae6369ed4e35e307565f91d0706970f460af552..c1a0551990555ce3e1dca81a9b5df7196a2a5fdf 100644
--- a/drivers/net/vxlan/vxlan_mdb.c
+++ b/drivers/net/vxlan/vxlan_mdb.c
@@ -195,7 +195,7 @@ static int vxlan_mdb_entry_info_fill(const struct vxlan_dev *vxlan,
 			be16_to_cpu(rd->remote_port)))
 		goto nest_err;
 
-	if (rd->remote_vni != vxlan->default_dst.remote_vni &&
+	if (rd->remote_vni != cfg->vni &&
 	    nla_put_u32(skb, MDBA_MDB_EATTR_VNI, be32_to_cpu(rd->remote_vni)))
 		goto nest_err;
 
@@ -620,12 +620,12 @@ static int vxlan_mdb_config_init(struct vxlan_mdb_config *cfg,
 
 	memset(cfg, 0, sizeof(*cfg));
 	cfg->vxlan = vxlan;
-	cfg->group.vni = vxlan->default_dst.remote_vni;
+	cfg->group.vni = vcfg->vni;
 	INIT_LIST_HEAD(&cfg->src_list);
 	cfg->nlflags = nlmsg_flags;
 	cfg->filter_mode = MCAST_EXCLUDE;
 	cfg->rt_protocol = RTPROT_STATIC;
-	cfg->remote_vni = vxlan->default_dst.remote_vni;
+	cfg->remote_vni = vcfg->vni;
 	cfg->remote_port = vcfg->dst_port;
 
 	if (entry->ifindex != dev->ifindex) {
@@ -986,7 +986,7 @@ vxlan_mdb_nlmsg_remote_size(const struct vxlan_dev *vxlan,
 	if (rd->remote_port && rd->remote_port != cfg->dst_port)
 		nlmsg_size += nla_total_size(sizeof(u16));
 	/* MDBA_MDB_EATTR_VNI */
-	if (rd->remote_vni != vxlan->default_dst.remote_vni)
+	if (rd->remote_vni != cfg->vni)
 		nlmsg_size += nla_total_size(sizeof(u32));
 	/* MDBA_MDB_EATTR_IFINDEX */
 	if (rd->remote_ifindex)
@@ -1488,11 +1488,13 @@ static int vxlan_mdb_get_parse(struct net_device *dev, struct nlattr *tb[],
 {
 	struct br_mdb_entry *entry = nla_data(tb[MDBA_GET_ENTRY]);
 	struct nlattr *mdbe_attrs[MDBE_ATTR_MAX + 1];
+	const struct vxlan_config *cfg;
 	struct vxlan_dev *vxlan = netdev_priv(dev);
 	int err;
 
+	cfg = rtnl_dereference(vxlan->cfg);
 	memset(group, 0, sizeof(*group));
-	group->vni = vxlan->default_dst.remote_vni;
+	group->vni = cfg->vni;
 
 	if (!tb[MDBA_GET_ENTRY_ATTRS]) {
 		vxlan_mdb_group_set(group, entry, NULL);
@@ -1641,7 +1643,7 @@ struct vxlan_mdb_entry *vxlan_mdb_entry_skb_get(struct vxlan_dev *vxlan,
 	 * entries are stored with the VNI of the VXLAN device.
 	 */
 	if (!(cfg->flags & VXLAN_F_COLLECT_METADATA))
-		src_vni = vxlan->default_dst.remote_vni;
+		src_vni = cfg->vni;
 
 	memset(&group, 0, sizeof(group));
 	group.vni = src_vni;
diff --git a/drivers/net/vxlan/vxlan_multicast.c b/drivers/net/vxlan/vxlan_multicast.c
index e2cf10da274f1b608d8bb5020d2b87ebfedeff46..6fcc4a36734361c28563d487d4ec108ca22d9e4f 100644
--- a/drivers/net/vxlan/vxlan_multicast.c
+++ b/drivers/net/vxlan/vxlan_multicast.c
@@ -14,11 +14,12 @@
 /* Update multicast group membership when first VNI on
  * multicast address is brought up
  */
-int vxlan_igmp_join(struct vxlan_dev *vxlan, union vxlan_addr *rip,
+int vxlan_igmp_join(struct vxlan_dev *vxlan, const union vxlan_addr *rip,
 		    int rifindex)
 {
-	union vxlan_addr *ip = (rip ? : &vxlan->default_dst.remote_ip);
-	int ifindex = (rifindex ? : vxlan->default_dst.remote_ifindex);
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
+	const union vxlan_addr *ip = (rip ? : &cfg->remote_ip);
+	int ifindex = (rifindex ? : cfg->remote_ifindex);
 	int ret = -EINVAL;
 	struct sock *sk;
 
@@ -47,11 +48,12 @@ int vxlan_igmp_join(struct vxlan_dev *vxlan, union vxlan_addr *rip,
 	return ret;
 }
 
-int vxlan_igmp_leave(struct vxlan_dev *vxlan, union vxlan_addr *rip,
+int vxlan_igmp_leave(struct vxlan_dev *vxlan, const union vxlan_addr *rip,
 		     int rifindex)
 {
-	union vxlan_addr *ip = (rip ? : &vxlan->default_dst.remote_ip);
-	int ifindex = (rifindex ? : vxlan->default_dst.remote_ifindex);
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
+	const union vxlan_addr *ip = (rip ? : &cfg->remote_ip);
+	int ifindex = (rifindex ? : cfg->remote_ifindex);
 	int ret = -EINVAL;
 	struct sock *sk;
 
@@ -80,8 +82,8 @@ int vxlan_igmp_leave(struct vxlan_dev *vxlan, union vxlan_addr *rip,
 	return ret;
 }
 
-static bool vxlan_group_used_match(union vxlan_addr *ip, int ifindex,
-				   union vxlan_addr *rip, int rifindex)
+static bool vxlan_group_used_match(const union vxlan_addr *ip, int ifindex,
+				   const union vxlan_addr *rip, int rifindex)
 {
 	if (!vxlan_addr_multicast(rip))
 		return false;
@@ -96,14 +98,16 @@ static bool vxlan_group_used_match(union vxlan_addr *ip, int ifindex,
 }
 
 static bool vxlan_group_used_by_vnifilter(struct vxlan_dev *vxlan,
-					  union vxlan_addr *ip, int ifindex)
+					  const struct vxlan_config *cfg,
+					  const union vxlan_addr *ip,
+					  int ifindex)
 {
 	struct vxlan_vni_group *vg = rtnl_dereference(vxlan->vnigrp);
 	struct vxlan_vni_node *v, *tmp;
 
 	if (vxlan_group_used_match(ip, ifindex,
-				   &vxlan->default_dst.remote_ip,
-				   vxlan->default_dst.remote_ifindex))
+				   &cfg->remote_ip,
+				   cfg->remote_ifindex))
 		return true;
 
 	list_for_each_entry_safe(v, tmp, &vg->vni_list, vlist) {
@@ -112,7 +116,7 @@ static bool vxlan_group_used_by_vnifilter(struct vxlan_dev *vxlan,
 
 		if (vxlan_group_used_match(ip, ifindex,
 					   &v->remote_ip,
-					   vxlan->default_dst.remote_ifindex))
+					   cfg->remote_ifindex))
 			return true;
 	}
 
@@ -121,16 +125,17 @@ static bool vxlan_group_used_by_vnifilter(struct vxlan_dev *vxlan,
 
 /* See if multicast group is already in use by other ID */
 bool vxlan_group_used(struct vxlan_net *vn, struct vxlan_dev *dev,
-		      __be32 vni, union vxlan_addr *rip, int rifindex)
+		      __be32 vni, const union vxlan_addr *rip, int rifindex)
 {
-	union vxlan_addr *ip = (rip ? : &dev->default_dst.remote_ip);
-	int ifindex = (rifindex ? : dev->default_dst.remote_ifindex);
+	const struct vxlan_config *dev_cfg = rtnl_dereference(dev->cfg);
+	const union vxlan_addr *ip = (rip ? : &dev_cfg->remote_ip);
+	int ifindex = (rifindex ? : dev_cfg->remote_ifindex);
 	struct vxlan_dev *vxlan;
 	struct vxlan_sock *sock4;
 #if IS_ENABLED(CONFIG_IPV6)
 	struct vxlan_sock *sock6;
 #endif
-	unsigned short family = dev->default_dst.remote_ip.sa.sa_family;
+	unsigned short family = dev_cfg->remote_ip.sa.sa_family;
 
 	sock4 = rtnl_dereference(dev->vn4_sock);
 
@@ -163,12 +168,12 @@ bool vxlan_group_used(struct vxlan_net *vn, struct vxlan_dev *dev,
 		cfg = rtnl_dereference(vxlan->cfg);
 
 		if (cfg->flags & VXLAN_F_VNIFILTER) {
-			if (!vxlan_group_used_by_vnifilter(vxlan, ip, ifindex))
+			if (!vxlan_group_used_by_vnifilter(vxlan, cfg, ip, ifindex))
 				continue;
 		} else {
 			if (!vxlan_group_used_match(ip, ifindex,
-						    &vxlan->default_dst.remote_ip,
-						    vxlan->default_dst.remote_ifindex))
+						    &cfg->remote_ip,
+						    cfg->remote_ifindex))
 				continue;
 		}
 
@@ -178,7 +183,8 @@ bool vxlan_group_used(struct vxlan_net *vn, struct vxlan_dev *dev,
 	return false;
 }
 
-static int vxlan_multicast_join_vnigrp(struct vxlan_dev *vxlan)
+static int vxlan_multicast_join_vnigrp(struct vxlan_dev *vxlan,
+				       const struct vxlan_config *cfg)
 {
 	struct vxlan_vni_group *vg = rtnl_dereference(vxlan->vnigrp);
 	struct vxlan_vni_node *v, *tmp, *vgood = NULL;
@@ -189,7 +195,7 @@ static int vxlan_multicast_join_vnigrp(struct vxlan_dev *vxlan)
 			continue;
 		/* skip if address is same as default address */
 		if (vxlan_addr_equal(&v->remote_ip,
-				     &vxlan->default_dst.remote_ip))
+				     &cfg->remote_ip))
 			continue;
 		ret = vxlan_igmp_join(vxlan, &v->remote_ip, 0);
 		if (ret == -EADDRINUSE)
@@ -204,7 +210,7 @@ static int vxlan_multicast_join_vnigrp(struct vxlan_dev *vxlan)
 			if (!vxlan_addr_multicast(&v->remote_ip))
 				continue;
 			if (vxlan_addr_equal(&v->remote_ip,
-					     &vxlan->default_dst.remote_ip))
+					     &cfg->remote_ip))
 				continue;
 			vxlan_igmp_leave(vxlan, &v->remote_ip, 0);
 			if (v == vgood)
@@ -240,9 +246,9 @@ int vxlan_multicast_join(struct vxlan_dev *vxlan)
 	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
 	int ret = 0;
 
-	if (vxlan_addr_multicast(&vxlan->default_dst.remote_ip)) {
-		ret = vxlan_igmp_join(vxlan, &vxlan->default_dst.remote_ip,
-				      vxlan->default_dst.remote_ifindex);
+	if (vxlan_addr_multicast(&cfg->remote_ip)) {
+		ret = vxlan_igmp_join(vxlan, &cfg->remote_ip,
+				      cfg->remote_ifindex);
 		if (ret == -EADDRINUSE)
 			ret = 0;
 		if (ret)
@@ -250,7 +256,7 @@ int vxlan_multicast_join(struct vxlan_dev *vxlan)
 	}
 
 	if (cfg->flags & VXLAN_F_VNIFILTER)
-		return vxlan_multicast_join_vnigrp(vxlan);
+		return vxlan_multicast_join_vnigrp(vxlan, cfg);
 
 	return 0;
 }
@@ -261,10 +267,10 @@ int vxlan_multicast_leave(struct vxlan_dev *vxlan)
 	struct vxlan_net *vn = net_generic(vxlan->net, vxlan_net_id);
 	int ret = 0;
 
-	if (vxlan_addr_multicast(&vxlan->default_dst.remote_ip) &&
+	if (vxlan_addr_multicast(&cfg->remote_ip) &&
 	    !vxlan_group_used(vn, vxlan, 0, NULL, 0)) {
-		ret = vxlan_igmp_leave(vxlan, &vxlan->default_dst.remote_ip,
-				       vxlan->default_dst.remote_ifindex);
+		ret = vxlan_igmp_leave(vxlan, &cfg->remote_ip,
+				       cfg->remote_ifindex);
 		if (ret)
 			return ret;
 	}
diff --git a/drivers/net/vxlan/vxlan_private.h b/drivers/net/vxlan/vxlan_private.h
index a2034ad418f7f77f64d8d81cb2312fcc608c7a0f..f0ee85f83732a2924236995015ae0d3a4f4cb872 100644
--- a/drivers/net/vxlan/vxlan_private.h
+++ b/drivers/net/vxlan/vxlan_private.h
@@ -224,14 +224,15 @@ void vxlan_vs_add_vnigrp(struct vxlan_dev *vxlan,
 			 struct vxlan_sock *vs,
 			 bool ipv6);
 void vxlan_vs_del_vnigrp(struct vxlan_dev *vxlan);
+bool vxlan_vnifilter_has_multicast(const struct vxlan_dev *vxlan);
 int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
-				   union vxlan_addr *old_remote_ip,
-				   union vxlan_addr *remote_ip,
+				   const union vxlan_addr *old_remote_ip,
+				   const union vxlan_addr *remote_ip,
 				   u32 old_ifindex, u32 new_ifindex,
 				   struct netlink_ext_ack *extack);
 int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
-			       union vxlan_addr *old_remote_ip,
-			       union vxlan_addr *new_remote_ip,
+			       const union vxlan_addr *old_remote_ip,
+			       const union vxlan_addr *new_remote_ip,
 			       u32 old_ifindex, u32 new_ifindex,
 			       struct netlink_ext_ack *extack);
 
@@ -240,10 +241,10 @@ int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
 int vxlan_multicast_join(struct vxlan_dev *vxlan);
 int vxlan_multicast_leave(struct vxlan_dev *vxlan);
 bool vxlan_group_used(struct vxlan_net *vn, struct vxlan_dev *dev,
-		      __be32 vni, union vxlan_addr *rip, int rifindex);
-int vxlan_igmp_join(struct vxlan_dev *vxlan, union vxlan_addr *rip,
+		      __be32 vni, const union vxlan_addr *rip, int rifindex);
+int vxlan_igmp_join(struct vxlan_dev *vxlan, const union vxlan_addr *rip,
 		    int rifindex);
-int vxlan_igmp_leave(struct vxlan_dev *vxlan, union vxlan_addr *rip,
+int vxlan_igmp_leave(struct vxlan_dev *vxlan, const union vxlan_addr *rip,
 		     int rifindex);
 
 /* vxlan_mdb.c */
diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
index 678565c4f8e277f31a485f8f7a2b79e1054ac530..48135ff9b95a1b2710b61f2e5e406816d4661367 100644
--- a/drivers/net/vxlan/vxlan_vnifilter.c
+++ b/drivers/net/vxlan/vxlan_vnifilter.c
@@ -502,9 +502,25 @@ static const struct nla_policy vni_filter_policy[VXLAN_VNIFILTER_MAX + 1] = {
 	[VXLAN_VNIFILTER_ENTRY] = { .type = NLA_NESTED },
 };
 
+bool vxlan_vnifilter_has_multicast(const struct vxlan_dev *vxlan)
+{
+	struct vxlan_vni_group *vg = rtnl_dereference(vxlan->vnigrp);
+	struct vxlan_vni_node *v;
+
+	if (!vg)
+		return false;
+
+	list_for_each_entry(v, &vg->vni_list, vlist) {
+		if (vxlan_addr_multicast(&v->remote_ip))
+			return true;
+	}
+
+	return false;
+}
+
 int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
-				   union vxlan_addr *old_remote_ip,
-				   union vxlan_addr *remote_ip,
+				   const union vxlan_addr *old_remote_ip,
+				   const union vxlan_addr *remote_ip,
 				   u32 old_ifindex, u32 new_ifindex,
 				   struct netlink_ext_ack *extack)
 {
@@ -518,8 +534,10 @@ int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
 
 	spin_lock_bh(&vxlan->hash_lock);
 	if (remote_ip && !vxlan_addr_any(remote_ip)) {
+		union vxlan_addr rip = *remote_ip;
+
 		err = vxlan_fdb_update(vxlan, all_zeros_mac,
-				       remote_ip,
+				       &rip,
 				       NUD_REACHABLE | NUD_PERMANENT,
 				       NLM_F_APPEND | NLM_F_CREATE,
 				       cfg->dst_port,
@@ -553,8 +571,8 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
 				  struct netlink_ext_ack *extack)
 {
 	struct vxlan_net *vn = net_generic(vxlan->net, vxlan_net_id);
-	struct vxlan_rdst *dst = &vxlan->default_dst;
-	union vxlan_addr *newrip = NULL, *oldrip = NULL;
+	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
+	const union vxlan_addr *newrip = NULL, *oldrip = NULL;
 	union vxlan_addr old_remote_ip;
 	int ret = 0;
 
@@ -566,8 +584,8 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
 	if (group && !vxlan_addr_any(group)) {
 		newrip = group;
 	} else {
-		if (!vxlan_addr_any(&dst->remote_ip))
-			newrip = &dst->remote_ip;
+		if (!vxlan_addr_any(&cfg->remote_ip))
+			newrip = &cfg->remote_ip;
 	}
 
 	/* if old rip exists, and no newrip,
@@ -584,8 +602,8 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
 
 	ret = vxlan_update_default_fdb_entry(vxlan, vninode->vni,
 					     oldrip, newrip,
-					     dst->remote_ifindex,
-					     dst->remote_ifindex,
+					     cfg->remote_ifindex,
+					     cfg->remote_ifindex,
 					     extack);
 	if (ret)
 		goto out;
@@ -600,7 +618,7 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
 		if (vxlan_addr_multicast(&old_remote_ip) &&
 		    !vxlan_group_used(vn, vxlan, vninode->vni,
 				      &old_remote_ip,
-				      vxlan->default_dst.remote_ifindex)) {
+				      cfg->remote_ifindex)) {
 			ret = vxlan_igmp_leave(vxlan, &old_remote_ip,
 					       0);
 			if (ret)
@@ -624,12 +642,12 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
 }
 
 int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
-			       union vxlan_addr *old_remote_ip,
-			       union vxlan_addr *new_remote_ip,
+			       const union vxlan_addr *old_remote_ip,
+			       const union vxlan_addr *new_remote_ip,
 			       u32 old_ifindex, u32 new_ifindex,
 			       struct netlink_ext_ack *extack)
 {
-	union vxlan_addr *oldrip, *newrip;
+	const union vxlan_addr *oldrip, *newrip;
 	struct list_head *headp, *hpos;
 	struct vxlan_vni_group *vg;
 	struct vxlan_vni_node *vent;
@@ -687,20 +705,19 @@ static void vxlan_vni_delete_group(struct vxlan_dev *vxlan,
 {
 	struct vxlan_net *vn = net_generic(vxlan->net, vxlan_net_id);
 	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
-	struct vxlan_rdst *dst = &vxlan->default_dst;
 
 	/* if per vni remote_ip not present, delete the
 	 * default dst remote_ip previously added for this vni
 	 */
 	if (!vxlan_addr_any(&vninode->remote_ip) ||
-	    !vxlan_addr_any(&dst->remote_ip)) {
+	    !vxlan_addr_any(&cfg->remote_ip)) {
 		spin_lock_bh(&vxlan->hash_lock);
 		__vxlan_fdb_delete(vxlan, all_zeros_mac,
 				   (vxlan_addr_any(&vninode->remote_ip) ?
-				   dst->remote_ip : vninode->remote_ip),
+				   cfg->remote_ip : vninode->remote_ip),
 				   cfg->dst_port,
 				   vninode->vni, vninode->vni,
-				   dst->remote_ifindex,
+				   cfg->remote_ifindex,
 				   true);
 		spin_unlock_bh(&vxlan->hash_lock);
 	}
@@ -709,7 +726,7 @@ static void vxlan_vni_delete_group(struct vxlan_dev *vxlan,
 		if (vxlan_addr_multicast(&vninode->remote_ip) &&
 		    !vxlan_group_used(vn, vxlan, vninode->vni,
 				      &vninode->remote_ip,
-				      dst->remote_ifindex)) {
+				      cfg->remote_ifindex)) {
 			vxlan_igmp_leave(vxlan, &vninode->remote_ip, 0);
 		}
 	}
@@ -924,6 +941,7 @@ static int vxlan_process_vni_filter(struct vxlan_dev *vxlan,
 				    int cmd, struct netlink_ext_ack *extack)
 {
 	struct nlattr *vattrs[VXLAN_VNIFILTER_ENTRY_MAX + 1];
+	const struct vxlan_config *cfg;
 	u32 vni_start = 0, vni_end = 0;
 	union vxlan_addr group;
 	int err;
@@ -961,7 +979,8 @@ static int vxlan_process_vni_filter(struct vxlan_dev *vxlan,
 		memset(&group, 0, sizeof(group));
 	}
 
-	if (vxlan_addr_multicast(&group) && !vxlan->default_dst.remote_ifindex) {
+	cfg = rtnl_dereference(vxlan->cfg);
+	if (vxlan_addr_multicast(&group) && !cfg->remote_ifindex) {
 		NL_SET_ERR_MSG(extack,
 			       "Local interface required for multicast remote group");
 
diff --git a/include/net/vxlan.h b/include/net/vxlan.h
index 7ced743ec8816d412bb14ec7ee7b422e97c38895..c3c9f2ccc3d1bfdf661312482856e0714e9ea3a6 100644
--- a/include/net/vxlan.h
+++ b/include/net/vxlan.h
@@ -204,7 +204,6 @@ struct vxlan_rdst {
 	u8			 offloaded:1;
 	__be32			 remote_vni;
 	u32			 remote_ifindex;
-	struct net_device	 *remote_dev;
 	struct list_head	 list;
 	struct rcu_head		 rcu;
 	struct dst_cache	 dst_cache;
@@ -295,7 +294,7 @@ struct vxlan_dev {
 #endif
 	struct net_device *dev;
 	struct net	  *net;		/* netns for packet i/o */
-	struct vxlan_rdst default_dst;	/* default destination */
+	struct net_device *lowerdev;
 
 	struct timer_list age_timer;
 	spinlock_t	  hash_lock;
-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* [PATCH v5 net-next 8/8] vxlan: no longer rely on RTNL in vxlan_fill_info()
  2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
                   ` (6 preceding siblings ...)
  2026-09-21 10:01 ` [PATCH v5 net-next 7/8] vxlan: remove default_dst and use vxlan_config and lowerdev Eric Dumazet
@ 2026-09-21 10:01 ` Eric Dumazet
  7 siblings, 0 replies; 12+ messages in thread
From: Eric Dumazet @ 2026-09-21 10:01 UTC (permalink / raw)
  To: David S . Miller, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, Kuniyuki Iwashima, netdev, eric.dumazet,
	Eric Dumazet

Now that vxlan->cfg is RCU-protected, we can update vxlan_fill_info()
to run under RCU read lock instead of relying on RTNL.

This completes the transition to RTNL-less link info dumping for VXLAN.

Signed-off-by: Eric Dumazet <edumazet@google.com>
Reviewed-by: Kuniyuki Iwashima <kuniyu@google.com>
---
 drivers/net/vxlan/vxlan_core.c | 5 ++++-
 1 file changed, 4 insertions(+), 1 deletion(-)

diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index b505fefcef70ae9e1704eb5ad5f50a3d5d7488cf..2e975294c69eca9d4400d51115d6164ff7265087 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -4744,7 +4744,8 @@ static int vxlan_fill_info(struct sk_buff *skb, const struct net_device *dev)
 	struct ifla_vxlan_port_range ports;
 	const struct vxlan_config *cfg;
 
-	cfg = rtnl_dereference(vxlan->cfg);
+	rcu_read_lock();
+	cfg = rcu_dereference(vxlan->cfg);
 
 	if (nla_put_u32(skb, IFLA_VXLAN_ID, be32_to_cpu(cfg->vni)))
 		goto nla_put_failure;
@@ -4843,9 +4844,11 @@ static int vxlan_fill_info(struct sk_buff *skb, const struct net_device *dev)
 		    &cfg->reserved_bits))
 		goto nla_put_failure;
 
+	rcu_read_unlock();
 	return 0;
 
 nla_put_failure:
+	rcu_read_unlock();
 	return -EMSGSIZE;
 }
 
-- 
2.55.0.1082.g2b9226bbc0-goog


^ permalink raw reply related	[flat|nested] 12+ messages in thread

* Re: [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes
  2026-09-21 10:01 ` [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes Eric Dumazet
@ 2026-09-22 16:02   ` netdev-bot+sashiko
  0 siblings, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22 16:02 UTC (permalink / raw)
  To: edumazet; +Cc: davem, kuba, pabeni, horms, kuniyu, netdev, eric.dumazet

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 1 (1) · Medium: 1 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] vxlan_changelink() (drivers/net/vxlan/vxlan_core.c:4538-4554)
  now routes every VXLAN_F_VNIFILTER device through…
- [Medium] Both rollback paths added by this patch ignore the int return
  value of vxlan_update_default_fdb_entry(): the err_unwind loop in…

Pre-existing issues:
- [High] This problem wasn't introduced by this patch, but it directly
  undermines the bug class the patch targets.

--

vxlan: update default fdb entries when the lower device changes

This refreshes the all zeros mac fdb entries when a changelink only swaps
the lower device, by passing the old and new ifindex down to
vxlan_update_default_fdb_entry() and visiting every vni rather than only
the ones inheriting the device group.  It also runs the multicast
leave/join pair for VXLAN_F_VNIFILTER devices and unwinds already updated
fdb entries when the vni list update fails partway through.

> diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
> index 347245cc1de4e..c4e3e8e8eef57 100644
> --- a/drivers/net/vxlan/vxlan_core.c
> +++ b/drivers/net/vxlan/vxlan_core.c

[ ... ]

> @@ -4475,13 +4476,16 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
>  	if (err)
>  		return err;
>  
> +	/* vxlan_config_apply() only commits remote_ifindex if lowerdev is set */
> +	new_ifindex = lowerdev ? conf.remote_ifindex : dst->remote_ifindex;
> +
>  	rem_ip_changed = !vxlan_addr_equal(&conf.remote_ip, &dst->remote_ip);
>  	change_igmp = vxlan->dev->flags & IFF_UP &&
>  		      (rem_ip_changed ||
> -		       dst->remote_ifindex != conf.remote_ifindex);
> +		       dst->remote_ifindex != new_ifindex);
>  
>  	/* handle default dst entry */
> -	if (rem_ip_changed) {
> +	if (rem_ip_changed || dst->remote_ifindex != new_ifindex) {

[ ... ]

> @@ -4509,13 +4513,21 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
>  					   true);
>  		spin_unlock_bh(&vxlan->hash_lock);
>  
> -		/* If vni filtering device, also update fdb entries of
> -		 * all vnis that were using default remote ip
> +		/* If vni filtering device, also update default fdb entries of
> +		 * all vnis
>  		 */
>  		if (vxlan->cfg.flags & VXLAN_F_VNIFILTER) {
>  			err = vxlan_vnilist_update_group(vxlan, &dst->remote_ip,
> -							 &conf.remote_ip, extack);
> +							 &conf.remote_ip,
> +							 dst->remote_ifindex,
> +							 new_ifindex, extack);
>  			if (err) {
> +				vxlan_update_default_fdb_entry(vxlan, conf.vni,
> +							       &conf.remote_ip,
> +							       &dst->remote_ip,
> +							       new_ifindex,
> +							       dst->remote_ifindex,
> +							       NULL);
>  				netdev_adjacent_change_abort(dst->remote_dev,
>  							     lowerdev, dev);
>  				return err;
> @@ -4523,7 +4535,9 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
>  		}
>  	}
>  
> -	if (change_igmp && vxlan_addr_multicast(&dst->remote_ip))
> +	if (change_igmp &&
> +	    (vxlan_addr_multicast(&dst->remote_ip) ||
> +	     (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
>  		err = vxlan_multicast_leave(vxlan);

[Severity: High]
Can this leave/join pair end up leaving a running vnifilter device with no
multicast memberships at all?

For VXLAN_F_VNIFILTER the leave reaches
vxlan_multicast_leave() -> vxlan_multicast_leave_vnigrp(), which issues one
vxlan_igmp_leave() per vni entry:

	list_for_each_entry_safe(v, tmp, &vg->vni_list, vlist) {
		if (vxlan_addr_multicast(&v->remote_ip) &&
		    !vxlan_group_used(vn, vxlan, v->vni, &v->remote_ip,
				      0)) {
			ret = vxlan_igmp_leave(vxlan, &v->remote_ip, 0);
			if (ret)
				last_err = ret;
		}
	}

A single "bridge vni add dev vxlan0 vni 100-200 group 239.1.1.1" gives every
vni in the range the same group, so all of those entries share one socket
membership.  The first vxlan_igmp_leave() drops it, and the remaining ones
return -EADDRNOTAVAIL (ip_mc_leave_group() starts with
int ret = -EADDRNOTAVAIL, and ipv6_sock_mc_drop() behaves the same), which
leave_vnigrp() records in last_err and returns.

vxlan_group_used() cannot absorb the duplicate here either: it skips
vxlan == dev and returns false early when refcount_read(&sock4->refcnt) == 1.

Note the asymmetry with the join side.  vxlan_multicast_join_vnigrp() both
skips entries whose address equals the device default remote_ip and maps
-EADDRINUSE to 0, while vxlan_multicast_leave_vnigrp() does neither.

With err set to -EADDRNOTAVAIL, netdev_adjacent_change_commit() and
vxlan_config_apply() still commit the new lower device, and then:

	if (!err && change_igmp &&
	    (vxlan_addr_multicast(&dst->remote_ip) ||
	     (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
		err = vxlan_multicast_join(vxlan);

skips the re-join entirely.  Repeating the same changelink does not repair
it, because lowerdev == dst->remote_dev makes lowerdev NULL, so
new_ifindex == dst->remote_ifindex and change_igmp becomes false.

Reproducer shape:

	ip link add vxlan0 type vxlan external vnifilter dstport 4789 \
		local 10.0.0.1 dev eth0
	bridge vni add dev vxlan0 vni 100-200 group 239.1.1.1
	ip link set vxlan0 up
	ip link set dev vxlan0 type vxlan ... dev eth1

Would it be better to rejoin unconditionally here, and/or make the per-vni
teardown act once per unique (group, ifindex) membership and tolerate
-EADDRNOTAVAIL the way the join side tolerates -EADDRINUSE?

>  
>  	if (netif_running(dev) && conf.age_interval != vxlan->cfg.age_interval)
> @@ -4535,7 +4549,8 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
>  	vxlan_config_apply(dev, &conf, lowerdev, vxlan->net, true);
>  
>  	if (!err && change_igmp &&
> -	    vxlan_addr_multicast(&dst->remote_ip))
> +	    (vxlan_addr_multicast(&dst->remote_ip) ||
> +	     (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
>  		err = vxlan_multicast_join(vxlan);
>  
>  	return err;

[ ... ]

> diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
> index dd94085e08865..12fa11a318184 100644
> --- a/drivers/net/vxlan/vxlan_vnifilter.c
> +++ b/drivers/net/vxlan/vxlan_vnifilter.c
> @@ -470,14 +470,19 @@ static const struct nla_policy vni_filter_policy[VXLAN_VNIFILTER_MAX + 1] = {
>  	[VXLAN_VNIFILTER_ENTRY] = { .type = NLA_NESTED },
>  };
>  
> -static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
> -					  union vxlan_addr *old_remote_ip,
> -					  union vxlan_addr *remote_ip,
> -					  struct netlink_ext_ack *extack)
> +int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
> +				   union vxlan_addr *old_remote_ip,
> +				   union vxlan_addr *remote_ip,
> +				   u32 old_ifindex, u32 new_ifindex,
> +				   struct netlink_ext_ack *extack)
>  {
> -	struct vxlan_rdst *dst = &vxlan->default_dst;
>  	int err = 0;
>  
> +	if (old_remote_ip && remote_ip &&
> +	    vxlan_addr_equal(old_remote_ip, remote_ip) &&
> +	    old_ifindex == new_ifindex)
> +		return 0;
> +
>  	spin_lock_bh(&vxlan->hash_lock);
>  	if (remote_ip && !vxlan_addr_any(remote_ip)) {
>  		err = vxlan_fdb_update(vxlan, all_zeros_mac,
> @@ -487,7 +492,7 @@ static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
>  				       vxlan->cfg.dst_port,
>  				       vni,
>  				       vni,
> -				       dst->remote_ifindex,
> +				       new_ifindex,
>  				       NTF_SELF, 0, true, extack);
>  		if (err) {
>  			spin_unlock_bh(&vxlan->hash_lock);
> @@ -500,7 +505,7 @@ static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
>  				   *old_remote_ip,
>  				   vxlan->cfg.dst_port,
>  				   vni, vni,
> -				   dst->remote_ifindex,
> +				   old_ifindex,
>  				   true);
>  	}
>  	spin_unlock_bh(&vxlan->hash_lock);
> @@ -546,6 +551,8 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
>  
>  	ret = vxlan_update_default_fdb_entry(vxlan, vninode->vni,
>  					     oldrip, newrip,
> +					     dst->remote_ifindex,
> +					     dst->remote_ifindex,
>  					     extack);
>  	if (ret)
>  		goto out;

[Severity: High]
This isn't a bug introduced by this patch, but the oldrip/newrip selection
just above this call makes the new equality guard dead for this caller, and
it also feeds the bug class the changelog describes.

In vxlan_vni_update_group() the two pointers are mutually exclusive:

	if (group && !vxlan_addr_any(group)) {
		newrip = group;
	} else {
		if (!vxlan_addr_any(&dst->remote_ip))
			newrip = &dst->remote_ip;
	}

	/* if old rip exists, and no newrip,
	 * explicitly delete old rip
	 */
	if (!newrip && !vxlan_addr_any(&old_remote_ip))
		oldrip = &old_remote_ip;

So when a vni's group is changed:

	bridge vni add dev vxlan0 vni 100 group 239.1.1.1
	bridge vni add dev vxlan0 vni 100 group 239.1.1.2

vxlan_vni_add() -> vxlan_vni_update() -> vxlan_vni_update_group(create=false)
passes newrip = 239.1.1.2 with oldrip still NULL, and in
vxlan_update_default_fdb_entry() the delete side is skipped:

	if (old_remote_ip && !vxlan_addr_any(old_remote_ip)) {
		__vxlan_fdb_delete(vxlan, all_zeros_mac,
				   *old_remote_ip,

The 239.1.1.1 rdst of the (all_zeros_mac, vni) entry therefore survives, and
vxlan_vni_delete_group() only deletes vninode->remote_ip, so it can never be
removed and each further group update appends one more rdst.

The new migration loop below only re-keys vent->remote_ip:

		} else {
			oldrip = &vent->remote_ip;
			newrip = &vent->remote_ip;
		}

so after a lower device change that stale rdst additionally keeps pointing
at the old ifindex.

Both the pre-existing "!create && oldrip && newrip &&
vxlan_addr_equal(oldrip, newrip)" test and the newly added
"old_remote_ip && remote_ip && vxlan_addr_equal(...) && old_ifindex ==
new_ifindex" test are unreachable for this caller, which suggests oldrip was
meant to be set on a group change.  Should vxlan_vni_update_group() set
oldrip = &old_remote_ip whenever it differs from newrip?

Verified via git blame as unchanged since f9c4bb0b245ce ("vxlan: vni
filtering support on collect metadata device") and still present at the end
of this series.

> @@ -583,8 +590,10 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
>  int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
>  			       union vxlan_addr *old_remote_ip,
>  			       union vxlan_addr *new_remote_ip,
> +			       u32 old_ifindex, u32 new_ifindex,
>  			       struct netlink_ext_ack *extack)
>  {
> +	union vxlan_addr *oldrip, *newrip;
>  	struct list_head *headp, *hpos;
>  	struct vxlan_vni_group *vg;
>  	struct vxlan_vni_node *vent;
> @@ -595,17 +604,46 @@ int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
>  	headp = &vg->vni_list;
>  	list_for_each_prev(hpos, headp) {
>  		vent = list_entry(hpos, struct vxlan_vni_node, vlist);
> +
>  		if (vxlan_addr_any(&vent->remote_ip)) {
> -			ret = vxlan_update_default_fdb_entry(vxlan, vent->vni,
> -							     old_remote_ip,
> -							     new_remote_ip,
> -							     extack);
> -			if (ret)
> -				return ret;
> +			oldrip = old_remote_ip;
> +			newrip = new_remote_ip;
> +		} else {
> +			/* A vni with its own group keeps it, but its fdb entry
> +			 * is still keyed on the device remote_ifindex.
> +			 */
> +			oldrip = &vent->remote_ip;
> +			newrip = &vent->remote_ip;
>  		}
> +
> +		ret = vxlan_update_default_fdb_entry(vxlan, vent->vni,
> +						     oldrip, newrip,
> +						     old_ifindex, new_ifindex,
> +						     extack);
> +		if (ret)
> +			goto err_unwind;
>  	}
>  
>  	return 0;
> +
> +err_unwind:
> +	list_for_each_continue(hpos, headp) {
> +		vent = list_entry(hpos, struct vxlan_vni_node, vlist);
> +
> +		if (vxlan_addr_any(&vent->remote_ip)) {
> +			oldrip = old_remote_ip;
> +			newrip = new_remote_ip;
> +		} else {
> +			oldrip = &vent->remote_ip;
> +			newrip = &vent->remote_ip;
> +		}
> +
> +		vxlan_update_default_fdb_entry(vxlan, vent->vni,
> +					       newrip, oldrip,
> +					       new_ifindex, old_ifindex,
> +					       NULL);
> +	}
> +	return ret;
>  }

[Severity: Medium]
Does discarding the return value here still guarantee the changelog's claim
that "default_dst and the fdb entries do not diverge"?

vxlan_update_default_fdb_entry() is ordered add-then-delete, and it returns
early if the add fails:

		if (err) {
			spin_unlock_bh(&vxlan->hash_lock);
			return err;
		}

The re-add performed by the unwind is fallible.  It goes
vxlan_fdb_update() -> vxlan_fdb_update_existing() -> vxlan_fdb_append(),
which does:

	rd = kmalloc_obj(*rd, GFP_ATOMIC);
	if (rd == NULL)
		return -ENOMEM;

and vxlan_fdb_notify() can also fail through
vxlan_fdb_switchdev_call_notifiers() on an offloading device.

When that happens the new-ifindex rdst stays installed, __vxlan_fdb_delete()
is never reached, and vxlan_changelink() returns the error without calling
vxlan_config_apply(), so default_dst.remote_ifindex keeps the old value.
Since vxlan_vni_delete_group() looks the entry up with the current (old)
dst->remote_ifindex and vxlan_fdb_find_rdst() requires an exact ifindex
match, that surviving rdst can no longer be deleted, and the next
"bridge vni add" of the same vni appends another one, which is the duplicated
BUM traffic this patch sets out to fix.

The same discarded return exists for the device default entry in
vxlan_changelink(), where the revert call is:

				vxlan_update_default_fdb_entry(vxlan, conf.vni,
							       &conf.remote_ip,
							       &dst->remote_ip,
							       new_ifindex,
							       dst->remote_ifindex,
							       NULL);

Would it make sense to attempt the delete even when the re-add fails, or to
restructure so the rollback needs no allocation, for example by keeping the
original destinations installed until every new destination has been added
successfully?

>  
>  static void vxlan_vni_delete_group(struct vxlan_dev *vxlan,

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260921100139.508191-1-edumazet%40google.com

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v5 net-next 2/8] vxlan: vnifilter: use list_for_each_entry_rcu() in vxlan_vnifilter_dump_dev()
  2026-09-21 10:01 ` [PATCH v5 net-next 2/8] vxlan: vnifilter: use list_for_each_entry_rcu() in vxlan_vnifilter_dump_dev() Eric Dumazet
@ 2026-09-22 16:02   ` netdev-bot+sashiko
  0 siblings, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22 16:02 UTC (permalink / raw)
  To: edumazet; +Cc: davem, kuba, pabeni, horms, kuniyu, netdev, eric.dumazet

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] vxlan_vnifilter_dump_dev() resumes a paginated RTM_GETTUNNEL dump
  inside one device by a positional ordinal (cb->args[1], applied via…

Pre-existing issues:
- [Low] vxlan_vnifilter_dump() still resumes the device walk by ordinal
  position (cb->args[0], 'if (idx < s_idx) goto skip;'), while the new…

--

vxlan: vnifilter: use list_for_each_entry_rcu() in vxlan_vnifilter_dump_dev()

This switches the vg->vni_list walk in vxlan_vnifilter_dump_dev() from
list_for_each_entry_safe() to list_for_each_entry_rcu(), and clears
cb->args[1] on the early return paths.  It also records the current
device ifindex in cb->args[2] so the per-device VNI offset is reset when
the device being dumped changes.

A couple of questions about the resume bookkeeping below.

> diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
> index 12fa11a318184..fad7c76418e92 100644
> --- a/drivers/net/vxlan/vxlan_vnifilter.c
> +++ b/drivers/net/vxlan/vxlan_vnifilter.c
> @@ -333,22 +333,34 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
>  				    struct sk_buff *skb,
>  				    struct netlink_callback *cb)
>  {
> -	struct vxlan_vni_node *tmp, *v, *vbegin = NULL, *vend = NULL;
> +	struct vxlan_vni_node *v, *vbegin = NULL, *vend = NULL;
>  	struct vxlan_dev *vxlan = netdev_priv(dev);
>  	struct tunnel_msg *new_tmsg, *tmsg;
> -	int idx = 0, s_idx = cb->args[1];
>  	struct vxlan_vni_group *vg;
>  	struct nlmsghdr *nlh;
> +	int idx = 0, s_idx;
>  	bool dump_stats;
>  	int err = 0;
>  
> -	if (!(vxlan->cfg.flags & VXLAN_F_VNIFILTER))
> +	if (cb->args[2] != dev->ifindex) {
> +		cb->args[1] = 0;
> +		cb->args[2] = dev->ifindex;
> +	}
> +	s_idx = cb->args[1];

[Severity: Low]
This isn't a bug introduced by this patch, but the two halves of the
resume state now use different encodings.  vxlan_vnifilter_dump_dev()
keys per-device progress on dev->ifindex in cb->args[2], while the
caller vxlan_vnifilter_dump() still resumes the device walk by ordinal:

	for_each_netdev_rcu(net, dev) {
		if (!netif_is_vxlan(dev))
			continue;
		if (idx < s_idx)
			goto skip;
		err = vxlan_vnifilter_dump_dev(dev, skb, cb);

If a vxlan device positioned before the partially dumped device is
unregistered between dump callbacks, the ordinals shift and the
in-progress device is now at idx < s_idx, so it is skipped before
vxlan_vnifilter_dump_dev() is ever reached and the new cb->args[2]
check cannot re-locate it.  The remainder of that device's VNIs is
then never emitted (and if it was the last device, the callback emits
an empty skb and the dump ends with NLMSG_DONE).

Would it make sense to key the device walk on ifindex as well, so the
resume point survives device removal?

For the record, the following patch in this series, "vxlan: vnifilter:
signal interrupted RTM_GETTUNNEL dumps", bumps a per-netns generation
counter on NETDEV_REGISTER/NETDEV_UNREGISTER of vxlan devices and
publishes it through cb->seq plus nl_dump_check_consistent(), and
netlink_dump_done() applies the same check to NLMSG_DONE, so after that
patch the truncation is at least reported to user space via
NLM_F_DUMP_INTR.

> +
> +	if (!(vxlan->cfg.flags & VXLAN_F_VNIFILTER)) {
> +		cb->args[1] = 0;
> +		cb->args[2] = 0;
>  		return -EINVAL;
> +	}
>  
>  	/* RCU needed because of the vni locking rules (rcu || rtnl) */
>  	vg = rcu_dereference(vxlan->vnigrp);
> -	if (!vg || !vg->num_vnis)
> +	if (!vg || !vg->num_vnis) {
> +		cb->args[1] = 0;
> +		cb->args[2] = 0;
>  		return 0;
> +	}
>  
>  	tmsg = nlmsg_data(cb->nlh);
>  	dump_stats = !!(tmsg->flags & TUNNEL_MSG_FLAG_STATS);
> @@ -362,7 +374,7 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
>  	new_tmsg->family = PF_BRIDGE;
>  	new_tmsg->ifindex = dev->ifindex;
>  
> -	list_for_each_entry_safe(v, tmp, &vg->vni_list, vlist) {
> +	list_for_each_entry_rcu(v, &vg->vni_list, vlist) {
>  		if (idx < s_idx) {
>  			idx++;
>  			continue;

[Severity: Low]
The new cb->args[2] check validates which device the saved offset
belongs to, but the offset itself is still a position in vg->vni_list,
and writers insert or remove entries anywhere in that sorted list:

drivers/net/vxlan/vxlan_vnifilter.c:__vxlan_vni_add_list() {
	...
	list_add_rcu(&v->vlist, hpos);
	vg->num_vnis++;
	...
}

drivers/net/vxlan/vxlan_vnifilter.c:__vxlan_vni_del_list() {
	...
	list_del_rcu(&v->vlist);
	vg->num_vnis--;
	...
}

Since RTNL is dropped between dump skbs, can a "bridge vni add/del" on
the device being dumped shift that cursor, so the next callback skips or
repeats VNI entries?

> @@ -394,6 +406,7 @@ static int vxlan_vnifilter_dump_dev(const struct net_device *dev,
>  	}
>  
>  	cb->args[1] = err ? idx : 0;
> +	cb->args[2] = err ? dev->ifindex : 0;

[Severity: Low]
At this commit no generation counter is recorded in cb->seq and
nl_dump_check_consistent() is not called here, so a skipped or duplicated
VNI is invisible to user space, i.e. NLM_F_DUMP_INTR is never set on the
affected messages.  The sibling dump in the same driver does maintain
such a counter:

drivers/net/vxlan/vxlan_mdb.c:vxlan_mdb_dump_dev() {
	...
	cb->seq = vxlan->mdb_seq;
	nl_dump_check_consistent(cb, nlh);
	...
}

Should this dump do the same?  Note the immediately following patch in
this series, "vxlan: vnifilter: signal interrupted RTM_GETTUNNEL dumps",
adds a per-netns atomic counter (vxlan_net.vnifilter_seq) incremented in
__vxlan_vni_add_list(), __vxlan_vni_del_list() and
vxlan_vni_update_group(), sets cb->seq in vxlan_vnifilter_dump() and
calls nl_dump_check_consistent(cb, nlh) here, which addresses this.  Is
there a reason to keep the two changes in separate patches, given the
positional resume is being reworked in this one?

>  
>  	nlmsg_end(skb, nlh);
>

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260921100139.508191-1-edumazet%40google.com

^ permalink raw reply	[flat|nested] 12+ messages in thread

* Re: [PATCH v5 net-next 7/8] vxlan: remove default_dst and use vxlan_config and lowerdev
  2026-09-21 10:01 ` [PATCH v5 net-next 7/8] vxlan: remove default_dst and use vxlan_config and lowerdev Eric Dumazet
@ 2026-09-22 16:02   ` netdev-bot+sashiko
  0 siblings, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-09-22 16:02 UTC (permalink / raw)
  To: edumazet; +Cc: davem, kuba, pabeni, horms, kuniyu, netdev, eric.dumazet

Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.

Low priority / nit picks for your consideration:

- [Low] The patch is advertised as a pure redundancy-removal refactor
  ("remove default_dst and use vxlan_config and lowerdev"), but it…

Pre-existing issues:
- [High] On a running VXLAN device with VXLAN_F_VNIFILTER, a changelink
  that swaps the lower device (or changes the remote) makes…
- [Medium] This problem wasn't introduced by this patch, but the patch
  touches the affected function: in vxlan_vni_update_group()…

--

vxlan: remove default_dst and use vxlan_config and lowerdev

This drops vxlan->default_dst and replaces it with a plain
struct net_device *lowerdev in struct vxlan_dev, moving the remaining
readers of remote_ip, remote_vni and remote_ifindex over to vxlan->cfg.
vxlan_config_apply() gains a lowerdev_changed argument, and
vxlan_config_validate() gains a new rejection when IFLA_VXLAN_LINK is
cleared while per-VNI multicast groups exist.

A few questions below about how the behaviour changes are packaged, plus
two older issues in code this patch touches.

(no Link: tags in the commit header)

> diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
> index 04b2925d7c38a..b505fefcef70a 100644
> --- a/drivers/net/vxlan/vxlan_core.c
> +++ b/drivers/net/vxlan/vxlan_core.c

[ ... ]

> @@ -3973,6 +3982,13 @@ static int vxlan_config_validate(struct net *src_net, struct vxlan_config *conf,
>  			return -EINVAL;
>  		}
>  
> +		if ((conf->flags & VXLAN_F_VNIFILTER) && old &&
> +		    vxlan_vnifilter_has_multicast(old)) {
> +			NL_SET_ERR_MSG(extack,
> +				       "Local interface required for multicast remote group");
> +			return -EINVAL;
> +		}
> +
>  #if IS_ENABLED(CONFIG_IPV6)
>  		if (conf->flags & VXLAN_F_IPV6_LINKLOCAL) {
>  			NL_SET_ERR_MSG(extack,
> @@ -4007,10 +4023,9 @@ static void vxlan_config_apply(struct net_device *dev,
>  			       struct vxlan_config *new_cfg,
>  			       struct net_device *lowerdev,
>  			       struct net *src_net,
> -			       bool changelink)
> +			       bool changelink, bool lowerdev_changed)
>  {
>  	struct vxlan_dev *vxlan = netdev_priv(dev);
> -	struct vxlan_rdst *dst = &vxlan->default_dst;
>  	unsigned short needed_headroom = ETH_HLEN;
>  	struct vxlan_config *old_cfg;
>  	int max_mtu = ETH_MAX_MTU;
> @@ -4023,18 +4038,13 @@ static void vxlan_config_apply(struct net_device *dev,
>  			vxlan_ether_setup(dev);
>  
>  		if (new_cfg->mtu)
> -			dev->mtu = new_cfg->mtu;
> +			WRITE_ONCE(dev->mtu, new_cfg->mtu);
>  
>  		vxlan->net = src_net;
>  	}
>  
> -	dst->remote_vni = new_cfg->vni;
> -
> -	memcpy(&dst->remote_ip, &new_cfg->remote_ip, sizeof(new_cfg->remote_ip));
> -
> +	dev->needed_tailroom = 0;
>  	if (lowerdev) {
> -		dst->remote_ifindex = new_cfg->remote_ifindex;
> -
>  		netif_inherit_tso_max(dev, lowerdev);
>  
>  		needed_headroom = lowerdev->hard_header_len;
> @@ -4042,16 +4052,17 @@ static void vxlan_config_apply(struct net_device *dev,
>  
>  		dev->needed_tailroom = lowerdev->needed_tailroom;
>  
> -		max_mtu = lowerdev->mtu - vxlan_headroom(flags);
> +		max_mtu = READ_ONCE(lowerdev->mtu) - vxlan_headroom(flags);
>  		if (max_mtu < ETH_MIN_MTU)
>  			max_mtu = ETH_MIN_MTU;
>  
>  		if (!changelink && !new_cfg->mtu)
> -			dev->mtu = max_mtu;
> +			WRITE_ONCE(dev->mtu, max_mtu);
>  	}
>  
> -	if (dev->mtu > max_mtu)
> -		dev->mtu = max_mtu;
> +	/* A changelink leaving the lower device alone must not shrink the mtu */
> +	if (lowerdev_changed && READ_ONCE(dev->mtu) > max_mtu)
> +		WRITE_ONCE(dev->mtu, max_mtu);

[Severity: Low]

The subject presents this as removing a redundant copy of the default
destination, but two of the changes here look like fixes for existing
user triggerable defects, and there is no Fixes: tag for either.  Would
it make sense to split them out?

Before this patch vxlan_changelink() forced lowerdev to NULL when it was
unchanged:

	if (dst->remote_dev == lowerdev)
		lowerdev = NULL;

and vxlan_config_apply() only derives the headroom inside if (lowerdev):

	needed_headroom = lowerdev->hard_header_len;
	needed_headroom += lowerdev->needed_headroom;

So any changelink that kept the same lower device, for example
ip link set vxlan0 type vxlan ttl 64, recomputed dev->needed_headroom as
ETH_HLEN + vxlan_headroom(flags) and dropped the lower device's
contribution.  Passing lowerdev unconditionally repairs that, and the
new lowerdev_changed argument keeps the previous mtu semantics.

The second one: vxlan_config_apply() used to commit
dst->remote_ifindex = new_cfg->remote_ifindex only when lowerdev was
non-NULL, while vxlan->cfg received remote_ifindex == 0 anyway, so
clearing IFLA_VXLAN_LINK left default_dst.remote_ifindex stale.
vxlan_fill_info(), vxlan_change_mtu(), vxlan_get_link_ksettings() and
vxlan_handle_lowerdev_unregister() all read that stale value.  This
patch removes the divergence, again with no Fixes: tag.

Finally, the new vxlan_vnifilter_has_multicast() check in
vxlan_config_validate() makes a changelink with IFLA_VXLAN_LINK=0 return
-EINVAL where it was previously accepted, and IFLA_VXLAN_LINK reporting
after such a request changes as well.

Could the two fixes be carried as separate patches with Fixes: tags so
they can be picked up for stable, and the new -EINVAL be called out as a
user visible change?

[ ... ]

> @@ -4579,26 +4589,23 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
>  	if (!new_cfg)
>  		return -ENOMEM;
>  
> -	if (dst->remote_dev == lowerdev)
> -		lowerdev = NULL;
> -
> -	err = netdev_adjacent_change_prepare(dst->remote_dev, lowerdev, dev,
> -					     extack);
> -	if (err) {
> -		kfree(new_cfg);
> -		return err;
> +	lowerdev_changed = vxlan->lowerdev != lowerdev;
> +	if (lowerdev_changed) {
> +		err = netdev_adjacent_change_prepare(vxlan->lowerdev, lowerdev,
> +						     dev, extack);
> +		if (err) {
> +			kfree(new_cfg);
> +			return err;
> +		}
>  	}

[ ... ]

> @@ -4651,20 +4661,26 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
>  	}
>  
>  	if (change_igmp &&
> -	    (vxlan_addr_multicast(&dst->remote_ip) ||
> +	    (vxlan_addr_multicast(&cfg->remote_ip) ||
>  	     (cfg->flags & VXLAN_F_VNIFILTER)))
>  		err = vxlan_multicast_leave(vxlan);
>  
>  	if (netif_running(dev) && conf.age_interval != cfg->age_interval)
>  		mod_timer(&vxlan->age_timer, jiffies);
>  
> -	netdev_adjacent_change_commit(dst->remote_dev, lowerdev, dev);
> -	if (lowerdev && lowerdev != dst->remote_dev)
> -		dst->remote_dev = lowerdev;
> -	vxlan_config_apply(dev, new_cfg, lowerdev, vxlan->net, true);
> +	if (lowerdev_changed) {
> +		if (lowerdev)
> +			netdev_adjacent_change_commit(vxlan->lowerdev, lowerdev,
> +						      dev);
> +		else
> +			netdev_upper_dev_unlink(vxlan->lowerdev, dev);
> +		vxlan->lowerdev = lowerdev;
> +	}
> +	vxlan_config_apply(dev, new_cfg, lowerdev, vxlan->net, true,
> +			   lowerdev_changed);
>  
>  	if (!err && change_igmp &&
> -	    (vxlan_addr_multicast(&dst->remote_ip) ||
> +	    (vxlan_addr_multicast(&new_cfg->remote_ip) ||
>  	     (new_cfg->flags & VXLAN_F_VNIFILTER)))
>  		err = vxlan_multicast_join(vxlan);

[Severity: High]

This is a pre-existing issue and not introduced by this patch, but since
the error handling around vxlan_multicast_leave() is being touched here,
can a failing leave leave a running device with no memberships at all?

With VXLAN_F_VNIFILTER, vxlan_multicast_leave() calls
vxlan_multicast_leave_vnigrp(), which walks every VNI without
deduplicating:

	list_for_each_entry_safe(v, tmp, &vg->vni_list, vlist) {
		if (vxlan_addr_multicast(&v->remote_ip) &&
		    !vxlan_group_used(vn, vxlan, v->vni, &v->remote_ip, 0)) {
			ret = vxlan_igmp_leave(vxlan, &v->remote_ip, 0);
			if (ret)
				last_err = ret;
		}
	}

The join side collapses duplicates, so only one membership exists:

	bridge vni add dev vx0 vni 100 group 239.1.1.1
	bridge vni add dev vx0 vni 200 group 239.1.1.1

vxlan_multicast_join_vnigrp() maps the second join's -EADDRINUSE to
success, and it also skips any per-VNI group equal to cfg->remote_ip,
which the leave path does not.

vxlan_group_used() cannot see the duplicate inside the same device:

	if (family == AF_INET && sock4 && refcount_read(&sock4->refcnt) == 1)
		return false;
	...
		if (!netif_running(vxlan->dev) || vxlan == dev)
			continue;

so the second vxlan_igmp_leave() reaches ip_mc_leave_group() with no
membership left and returns -EADDRNOTAVAIL, which is propagated as
last_err.

At that point the new adjacency and the new cfg are still committed, and
the

	if (!err && change_igmp && ...)
		err = vxlan_multicast_join(vxlan);

guard skips the rejoin, so the device keeps running with no group
membership and stops receiving flooded traffic.  Retrying the same
request does not help, since cfg->remote_ifindex now matches the
requested value and change_igmp is false, so only an administrative
down/up restores it.  Should the leave path skip groups it never joined,
or should the rejoin not be gated on the leave result?

> diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
> index 678565c4f8e27..48135ff9b95a1 100644
> --- a/drivers/net/vxlan/vxlan_vnifilter.c
> +++ b/drivers/net/vxlan/vxlan_vnifilter.c

[ ... ]

> @@ -518,8 +534,10 @@ int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
>  
>  	spin_lock_bh(&vxlan->hash_lock);
>  	if (remote_ip && !vxlan_addr_any(remote_ip)) {
> +		union vxlan_addr rip = *remote_ip;
> +
>  		err = vxlan_fdb_update(vxlan, all_zeros_mac,
> -				       remote_ip,
> +				       &rip,
>  				       NUD_REACHABLE | NUD_PERMANENT,
>  				       NLM_F_APPEND | NLM_F_CREATE,
>  				       cfg->dst_port,
> @@ -553,8 +571,8 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
>  				  struct netlink_ext_ack *extack)
>  {
>  	struct vxlan_net *vn = net_generic(vxlan->net, vxlan_net_id);
> -	struct vxlan_rdst *dst = &vxlan->default_dst;
> -	union vxlan_addr *newrip = NULL, *oldrip = NULL;
> +	const struct vxlan_config *cfg = rtnl_dereference(vxlan->cfg);
> +	const union vxlan_addr *newrip = NULL, *oldrip = NULL;
>  	union vxlan_addr old_remote_ip;
>  	int ret = 0;
>  
> @@ -566,8 +584,8 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
>  	if (group && !vxlan_addr_any(group)) {
>  		newrip = group;
>  	} else {
> -		if (!vxlan_addr_any(&dst->remote_ip))
> -			newrip = &dst->remote_ip;
> +		if (!vxlan_addr_any(&cfg->remote_ip))
> +			newrip = &cfg->remote_ip;
>  	}
>  
>  	/* if old rip exists, and no newrip,

[Severity: Medium]

This isn't a bug introduced by this patch either, but vxlan_vni_update_group()
is being converted here, so while looking at it: does replacing one non-zero
per-VNI group with another leave the old flood destination behind?

oldrip is only set when there is no newrip:

	if (!newrip && !vxlan_addr_any(&old_remote_ip))
		oldrip = &old_remote_ip;

so on a replacement vxlan_update_default_fdb_entry() is called with
old_remote_ip == NULL and only runs the NLM_F_APPEND | NLM_F_CREATE
update.  For the all-zeros MAC that goes through
vxlan_fdb_update_existing() -> vxlan_fdb_append(), which keeps the
existing rdst:

	if ((flags & NLM_F_APPEND) &&
	    (is_multicast_ether_addr(f->key.eth_addr) ||
	     is_zero_ether_addr(f->key.eth_addr))) {
		rc = vxlan_fdb_append(f, ip, port, vni, ifindex, &rd);

vninode->remote_ip is then overwritten with the new group and the request
succeeds, so the zero-MAC entry ends up with two remotes:

	bridge vni add dev vx0 vni 100 group 239.1.1.1
	bridge vni add dev vx0 vni 100 group 239.1.1.2

vxlan_vni_add() routes the second request to vxlan_vni_update() ->
vxlan_vni_update_group() with create == false.  A later VNI delete only
targets vninode->remote_ip in vxlan_vni_delete_group(), so the stale rdst
survives and re-adding the VNI accumulates more.  The same applies when a
per-VNI group is switched back to a non-zero device default.

Should the replacement case pass the previous address as old_remote_ip so
the stale destination is removed?

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260921100139.508191-1-edumazet%40google.com

^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-09-22 16:02 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes Eric Dumazet
2026-09-22 16:02   ` netdev-bot+sashiko
2026-09-21 10:01 ` [PATCH v5 net-next 2/8] vxlan: vnifilter: use list_for_each_entry_rcu() in vxlan_vnifilter_dump_dev() Eric Dumazet
2026-09-22 16:02   ` netdev-bot+sashiko
2026-09-21 10:01 ` [PATCH v5 net-next 3/8] vxlan: vnifilter: signal interrupted RTM_GETTUNNEL dumps Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 4/8] vxlan: pass vxlan_config pointer to helper functions Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 5/8] vxlan: move VXLAN_F_MDB to struct vxlan_dev flags Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 6/8] vxlan: convert configuration to RCU protection Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 7/8] vxlan: remove default_dst and use vxlan_config and lowerdev Eric Dumazet
2026-09-22 16:02   ` netdev-bot+sashiko
2026-09-21 10:01 ` [PATCH v5 net-next 8/8] vxlan: no longer rely on RTNL in vxlan_fill_info() Eric Dumazet

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox