From: Eric Dumazet <edumazet@google.com>
To: "David S . Miller" <davem@davemloft.net>,
Jakub Kicinski <kuba@kernel.org>,
Paolo Abeni <pabeni@redhat.com>
Cc: Simon Horman <horms@kernel.org>,
Kuniyuki Iwashima <kuniyu@google.com>,
netdev@vger.kernel.org, eric.dumazet@gmail.com,
Eric Dumazet <edumazet@google.com>
Subject: [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes
Date: Mon, 21 Sep 2026 10:01:32 +0000 [thread overview]
Message-ID: <20260921100139.508191-2-edumazet@google.com> (raw)
In-Reply-To: <20260921100139.508191-1-edumazet@google.com>
vxlan_changelink() only refreshed the default fdb entries when the
remote IP changed, but vxlan_config_apply() also updates
default_dst.remote_ifindex when the lower device changes. A changelink
that only swaps the lower device therefore left the all zeros mac rdst
pointing at the old ifindex:
ip link add vxlan0 type vxlan id 10 group 239.1.1.1 dev eth0
ip link set dev vxlan0 type vxlan group 239.1.1.1 dev eth1
vxlan_xmit_one() uses rdst->remote_ifindex as the route oif, so traffic
kept leaving eth0.
VNI filter entries have the same problem, and are worse: their fdb
entries are keyed on the device remote_ifindex even when the vni
carries its own group, but vxlan_vnilist_update_group() only visited
the vnis without one. vxlan_vni_delete_group() later looks an entry up
with the current remote_ifindex, and vxlan_fdb_find_rdst() requires an
exact match, so the lookup failed and the entry survived the delete.
Re-adding the same vni then appended a second rdst, duplicating
transmitted BUM traffic.
Pass the old and new ifindex down to vxlan_update_default_fdb_entry()
so the append targets the new lower device and the delete still matches
the entry created for the old one, and refresh every vni rather than
only those inheriting the device group. If updating the vni list fails
partway through, unwind the already updated fdb entries so default_dst
and the fdb entries do not diverge.
Also run vxlan_multicast_leave() and vxlan_multicast_join() when
VXLAN_F_VNIFILTER is set so per-VNI multicast memberships are migrated
even when the device default remote_ip is not multicast.
The new ifindex is the one vxlan_config_apply() will commit, which is
the current one when lowerdev is NULL, so that default_dst and the fdb
entries can not diverge.
Fixes: 8bcdc4f3a20b ("vxlan: add changelink support")
Signed-off-by: Eric Dumazet <edumazet@google.com>
---
drivers/net/vxlan/vxlan_core.c | 31 ++++++++++----
drivers/net/vxlan/vxlan_private.h | 6 +++
drivers/net/vxlan/vxlan_vnifilter.c | 64 +++++++++++++++++++++++------
3 files changed, 80 insertions(+), 21 deletions(-)
diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c
index 347245cc1de4ea44f723176312b0b53869a1641c..c4e3e8e8eef57196c3b7120ff4b8ce71207f3bf8 100644
--- a/drivers/net/vxlan/vxlan_core.c
+++ b/drivers/net/vxlan/vxlan_core.c
@@ -4452,6 +4452,7 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
struct net_device *lowerdev;
struct vxlan_config conf;
struct vxlan_rdst *dst;
+ u32 new_ifindex;
int err;
if (!rtnl_dev_link_net_capable(dev, vxlan->net))
@@ -4475,13 +4476,16 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
if (err)
return err;
+ /* vxlan_config_apply() only commits remote_ifindex if lowerdev is set */
+ new_ifindex = lowerdev ? conf.remote_ifindex : dst->remote_ifindex;
+
rem_ip_changed = !vxlan_addr_equal(&conf.remote_ip, &dst->remote_ip);
change_igmp = vxlan->dev->flags & IFF_UP &&
(rem_ip_changed ||
- dst->remote_ifindex != conf.remote_ifindex);
+ dst->remote_ifindex != new_ifindex);
/* handle default dst entry */
- if (rem_ip_changed) {
+ if (rem_ip_changed || dst->remote_ifindex != new_ifindex) {
spin_lock_bh(&vxlan->hash_lock);
if (!vxlan_addr_any(&conf.remote_ip)) {
err = vxlan_fdb_update(vxlan, all_zeros_mac,
@@ -4490,7 +4494,7 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
NLM_F_APPEND | NLM_F_CREATE,
vxlan->cfg.dst_port,
conf.vni, conf.vni,
- conf.remote_ifindex,
+ new_ifindex,
NTF_SELF, 0, true, extack);
if (err) {
spin_unlock_bh(&vxlan->hash_lock);
@@ -4509,13 +4513,21 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
true);
spin_unlock_bh(&vxlan->hash_lock);
- /* If vni filtering device, also update fdb entries of
- * all vnis that were using default remote ip
+ /* If vni filtering device, also update default fdb entries of
+ * all vnis
*/
if (vxlan->cfg.flags & VXLAN_F_VNIFILTER) {
err = vxlan_vnilist_update_group(vxlan, &dst->remote_ip,
- &conf.remote_ip, extack);
+ &conf.remote_ip,
+ dst->remote_ifindex,
+ new_ifindex, extack);
if (err) {
+ vxlan_update_default_fdb_entry(vxlan, conf.vni,
+ &conf.remote_ip,
+ &dst->remote_ip,
+ new_ifindex,
+ dst->remote_ifindex,
+ NULL);
netdev_adjacent_change_abort(dst->remote_dev,
lowerdev, dev);
return err;
@@ -4523,7 +4535,9 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
}
}
- if (change_igmp && vxlan_addr_multicast(&dst->remote_ip))
+ if (change_igmp &&
+ (vxlan_addr_multicast(&dst->remote_ip) ||
+ (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
err = vxlan_multicast_leave(vxlan);
if (netif_running(dev) && conf.age_interval != vxlan->cfg.age_interval)
@@ -4535,7 +4549,8 @@ static int vxlan_changelink(struct net_device *dev, struct nlattr *tb[],
vxlan_config_apply(dev, &conf, lowerdev, vxlan->net, true);
if (!err && change_igmp &&
- vxlan_addr_multicast(&dst->remote_ip))
+ (vxlan_addr_multicast(&dst->remote_ip) ||
+ (vxlan->cfg.flags & VXLAN_F_VNIFILTER)))
err = vxlan_multicast_join(vxlan);
return err;
diff --git a/drivers/net/vxlan/vxlan_private.h b/drivers/net/vxlan/vxlan_private.h
index b1eec221636088aa1c1674221d5ef0f13698b53f..e4ceca925bde7bea909bd691fa0f4ebfebdefbd3 100644
--- a/drivers/net/vxlan/vxlan_private.h
+++ b/drivers/net/vxlan/vxlan_private.h
@@ -213,9 +213,15 @@ void vxlan_vs_add_vnigrp(struct vxlan_dev *vxlan,
struct vxlan_sock *vs,
bool ipv6);
void vxlan_vs_del_vnigrp(struct vxlan_dev *vxlan);
+int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
+ union vxlan_addr *old_remote_ip,
+ union vxlan_addr *remote_ip,
+ u32 old_ifindex, u32 new_ifindex,
+ struct netlink_ext_ack *extack);
int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
union vxlan_addr *old_remote_ip,
union vxlan_addr *new_remote_ip,
+ u32 old_ifindex, u32 new_ifindex,
struct netlink_ext_ack *extack);
diff --git a/drivers/net/vxlan/vxlan_vnifilter.c b/drivers/net/vxlan/vxlan_vnifilter.c
index dd94085e088656d27b62420a5c8c95c609510a4c..12fa11a31818456232682edb1c79042478832836 100644
--- a/drivers/net/vxlan/vxlan_vnifilter.c
+++ b/drivers/net/vxlan/vxlan_vnifilter.c
@@ -470,14 +470,19 @@ static const struct nla_policy vni_filter_policy[VXLAN_VNIFILTER_MAX + 1] = {
[VXLAN_VNIFILTER_ENTRY] = { .type = NLA_NESTED },
};
-static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
- union vxlan_addr *old_remote_ip,
- union vxlan_addr *remote_ip,
- struct netlink_ext_ack *extack)
+int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
+ union vxlan_addr *old_remote_ip,
+ union vxlan_addr *remote_ip,
+ u32 old_ifindex, u32 new_ifindex,
+ struct netlink_ext_ack *extack)
{
- struct vxlan_rdst *dst = &vxlan->default_dst;
int err = 0;
+ if (old_remote_ip && remote_ip &&
+ vxlan_addr_equal(old_remote_ip, remote_ip) &&
+ old_ifindex == new_ifindex)
+ return 0;
+
spin_lock_bh(&vxlan->hash_lock);
if (remote_ip && !vxlan_addr_any(remote_ip)) {
err = vxlan_fdb_update(vxlan, all_zeros_mac,
@@ -487,7 +492,7 @@ static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
vxlan->cfg.dst_port,
vni,
vni,
- dst->remote_ifindex,
+ new_ifindex,
NTF_SELF, 0, true, extack);
if (err) {
spin_unlock_bh(&vxlan->hash_lock);
@@ -500,7 +505,7 @@ static int vxlan_update_default_fdb_entry(struct vxlan_dev *vxlan, __be32 vni,
*old_remote_ip,
vxlan->cfg.dst_port,
vni, vni,
- dst->remote_ifindex,
+ old_ifindex,
true);
}
spin_unlock_bh(&vxlan->hash_lock);
@@ -546,6 +551,8 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
ret = vxlan_update_default_fdb_entry(vxlan, vninode->vni,
oldrip, newrip,
+ dst->remote_ifindex,
+ dst->remote_ifindex,
extack);
if (ret)
goto out;
@@ -583,8 +590,10 @@ static int vxlan_vni_update_group(struct vxlan_dev *vxlan,
int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
union vxlan_addr *old_remote_ip,
union vxlan_addr *new_remote_ip,
+ u32 old_ifindex, u32 new_ifindex,
struct netlink_ext_ack *extack)
{
+ union vxlan_addr *oldrip, *newrip;
struct list_head *headp, *hpos;
struct vxlan_vni_group *vg;
struct vxlan_vni_node *vent;
@@ -595,17 +604,46 @@ int vxlan_vnilist_update_group(struct vxlan_dev *vxlan,
headp = &vg->vni_list;
list_for_each_prev(hpos, headp) {
vent = list_entry(hpos, struct vxlan_vni_node, vlist);
+
if (vxlan_addr_any(&vent->remote_ip)) {
- ret = vxlan_update_default_fdb_entry(vxlan, vent->vni,
- old_remote_ip,
- new_remote_ip,
- extack);
- if (ret)
- return ret;
+ oldrip = old_remote_ip;
+ newrip = new_remote_ip;
+ } else {
+ /* A vni with its own group keeps it, but its fdb entry
+ * is still keyed on the device remote_ifindex.
+ */
+ oldrip = &vent->remote_ip;
+ newrip = &vent->remote_ip;
}
+
+ ret = vxlan_update_default_fdb_entry(vxlan, vent->vni,
+ oldrip, newrip,
+ old_ifindex, new_ifindex,
+ extack);
+ if (ret)
+ goto err_unwind;
}
return 0;
+
+err_unwind:
+ list_for_each_continue(hpos, headp) {
+ vent = list_entry(hpos, struct vxlan_vni_node, vlist);
+
+ if (vxlan_addr_any(&vent->remote_ip)) {
+ oldrip = old_remote_ip;
+ newrip = new_remote_ip;
+ } else {
+ oldrip = &vent->remote_ip;
+ newrip = &vent->remote_ip;
+ }
+
+ vxlan_update_default_fdb_entry(vxlan, vent->vni,
+ newrip, oldrip,
+ new_ifindex, old_ifindex,
+ NULL);
+ }
+ return ret;
}
static void vxlan_vni_delete_group(struct vxlan_dev *vxlan,
--
2.55.0.1082.g2b9226bbc0-goog
next prev parent reply other threads:[~2026-09-21 10:01 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 10:01 [PATCH v5 net-next 0/8] vxlan: convert configuration to RCU and enable lockless dumps Eric Dumazet
2026-09-21 10:01 ` Eric Dumazet [this message]
2026-09-22 16:02 ` [PATCH v5 net-next 1/8] vxlan: update default fdb entries when the lower device changes netdev-bot+sashiko
2026-09-21 10:01 ` [PATCH v5 net-next 2/8] vxlan: vnifilter: use list_for_each_entry_rcu() in vxlan_vnifilter_dump_dev() Eric Dumazet
2026-09-22 16:02 ` netdev-bot+sashiko
2026-09-21 10:01 ` [PATCH v5 net-next 3/8] vxlan: vnifilter: signal interrupted RTM_GETTUNNEL dumps Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 4/8] vxlan: pass vxlan_config pointer to helper functions Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 5/8] vxlan: move VXLAN_F_MDB to struct vxlan_dev flags Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 6/8] vxlan: convert configuration to RCU protection Eric Dumazet
2026-09-21 10:01 ` [PATCH v5 net-next 7/8] vxlan: remove default_dst and use vxlan_config and lowerdev Eric Dumazet
2026-09-22 16:02 ` netdev-bot+sashiko
2026-09-21 10:01 ` [PATCH v5 net-next 8/8] vxlan: no longer rely on RTNL in vxlan_fill_info() Eric Dumazet
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260921100139.508191-2-edumazet@google.com \
--to=edumazet@google.com \
--cc=davem@davemloft.net \
--cc=eric.dumazet@gmail.com \
--cc=horms@kernel.org \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox