* [PATCH net-next v10 1/6] netlink: specs: rt-addr: fix the type of target-netnsid
2026-10-07 11:58 [PATCH net-next v10 0/6] rtnetlink: dump link-layer multicast addresses Yuyang Huang
@ 2026-10-07 11:58 ` Yuyang Huang
2026-10-09 11:58 ` netdev-bot+sashiko
2026-10-07 11:58 ` [PATCH net-next v10 2/6] net: change netdev_hw_addr_list count through helpers Yuyang Huang
` (4 subsequent siblings)
5 siblings, 1 reply; 12+ messages in thread
From: Yuyang Huang @ 2026-10-07 11:58 UTC (permalink / raw)
To: Yuyang Huang
Cc: Ajay Singh, Aleksandr Loktionov, Andrew Lunn, Claudiu Beznea,
David S. Miller, David Ahern, Donald Hunter, Eric Dumazet,
Ido Schimmel, Jacob Keller, Jakub Kicinski, Johannes Berg,
Kees Cook, Kory Maincent, Kuniyuki Iwashima, Nicolas Dichtel,
Nikolaos Gkarlis, Paolo Abeni, Sabrina Dubroca, Shuah Khan,
Simon Horman, Stanislav Fomichev, Vadim Fedorenko,
Willem de Bruijn, linux-kernel, linux-kselftest, linux-wireless,
netdev
The kernel parses IFA_TARGET_NETNSID as NLA_S32 and rt-link.yaml
declares its target-netnsid as s32, but rt-addr.yaml has it as binary.
Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
Reviewed-by: Nicolas Dichtel <nicolas.dichtel@6wind.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
---
Documentation/netlink/specs/rt-addr.yaml | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/Documentation/netlink/specs/rt-addr.yaml b/Documentation/netlink/specs/rt-addr.yaml
index 0ecbd24c890c6..17ead22034515 100644
--- a/Documentation/netlink/specs/rt-addr.yaml
+++ b/Documentation/netlink/specs/rt-addr.yaml
@@ -119,7 +119,7 @@ attribute-sets:
type: u32
-
name: target-netnsid
- type: binary
+ type: s32
-
name: proto
type: u8
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread* Re: [PATCH net-next v10 1/6] netlink: specs: rt-addr: fix the type of target-netnsid
2026-10-07 11:58 ` [PATCH net-next v10 1/6] netlink: specs: rt-addr: fix the type of target-netnsid Yuyang Huang
@ 2026-10-09 11:58 ` netdev-bot+sashiko
0 siblings, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-10-09 11:58 UTC (permalink / raw)
To: sigefriedhyy
Cc: ajay.kathat, aleksandr.loktionov, andrew+netdev, claudiu.beznea,
davem, dsahern, donald.hunter, edumazet, idosch, jacob.e.keller,
kuba, johannes, kees, kory.maincent, kuniyu, nicolas.dichtel,
nickgarlis, pabeni, sd, shuah, horms, sdf.kernel, vadim.fedorenko,
willemb, linux-kernel, linux-kselftest, linux-wireless, netdev
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] Commit "netlink: specs: rt-addr: fix the type of target-netnsid"
fixes a real defect in Documentation/netlink/specs/rt-addr.yaml but…
--
netlink: specs: rt-addr: fix the type of target-netnsid
This patch changes the type of the target-netnsid attribute in
rt-addr.yaml from binary to s32. The kernel parses IFA_TARGET_NETNSID
as NLA_S32, and rt-link.yaml already declares its own target-netnsid
as s32.
> The kernel parses IFA_TARGET_NETNSID as NLA_S32 and rt-link.yaml
> declares its target-netnsid as s32, but rt-addr.yaml has it as binary.
>
> Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
> Reviewed-by: Nicolas Dichtel <nicolas.dichtel@6wind.com>
> Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
[Severity: Low]
Should this have a Fixes: tag?
The binary type on target-netnsid goes back to the commit that first
added the spec. A later rename to target-netnsid kept the type as it
was. On the kernel side, IFA_TARGET_NETNSID has always been s32. It
is NLA_S32 in ifa_ipv4_policy in net/ipv4/devinet.c and in the IPv6
policy in net/ipv6/addrconf.c, and it is written with nla_put_s32()
and read with nla_get_s32().
Python YNL is affected on any tree without this fix. When an integer
is passed for target-netnsid, YnlFamily._add_attr() in
tools/net/ynl/pyynl/lib/ynl.py takes the binary branch and raises:
raise YnlException(f'Unknown type for binary attribute, value: {value}')
Replies that carry the attribute are also decoded as raw bytes, not
as a signed int.
Would something like this be appropriate?
Fixes: dfb0f7d9d979 ("doc/netlink: Add spec for rt addr messages")
> diff --git a/Documentation/netlink/specs/rt-addr.yaml b/Documentation/netlink/specs/rt-addr.yaml
> index 0ecbd24c890c..17ead2203451 100644
> --- a/Documentation/netlink/specs/rt-addr.yaml
> +++ b/Documentation/netlink/specs/rt-addr.yaml
> @@ -119,7 +119,7 @@ attribute-sets:
> type: u32
> -
> name: target-netnsid
> - type: binary
> + type: s32
> -
> name: proto
> type: u8
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261007115821.11332-1-sigefriedhyy%40gmail.com
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH net-next v10 2/6] net: change netdev_hw_addr_list count through helpers
2026-10-07 11:58 [PATCH net-next v10 0/6] rtnetlink: dump link-layer multicast addresses Yuyang Huang
2026-10-07 11:58 ` [PATCH net-next v10 1/6] netlink: specs: rt-addr: fix the type of target-netnsid Yuyang Huang
@ 2026-10-07 11:58 ` Yuyang Huang
2026-10-07 11:58 ` [PATCH net-next v10 3/6] net: add a generation counter for dev->mc changes Yuyang Huang
` (3 subsequent siblings)
5 siblings, 0 replies; 12+ messages in thread
From: Yuyang Huang @ 2026-10-07 11:58 UTC (permalink / raw)
To: Yuyang Huang
Cc: Ajay Singh, Aleksandr Loktionov, Andrew Lunn, Claudiu Beznea,
David S. Miller, David Ahern, Donald Hunter, Eric Dumazet,
Ido Schimmel, Jacob Keller, Jakub Kicinski, Johannes Berg,
Kees Cook, Kory Maincent, Kuniyuki Iwashima, Nicolas Dichtel,
Nikolaos Gkarlis, Paolo Abeni, Sabrina Dubroca, Shuah Khan,
Simon Horman, Stanislav Fomichev, Vadim Fedorenko,
Willem de Bruijn, linux-kernel, linux-kselftest, linux-wireless,
netdev
The count of a netdev_hw_addr_list is changed in several places of
dev_addr_lists.c and a few drivers read it directly. The next patch
needs to account every change of the count of dev->mc.
Add __hw_addr_count_add(), __hw_addr_count_inc(), __hw_addr_count_dec()
and __hw_addr_count_reset(), use them for every change of the count
and rename the field to _count so that a direct write stands out.
Readers keep using netdev_hw_addr_list_count() and the netdev_uc_count()
and netdev_mc_count() helpers, the few that read the field directly
are converted. No functional change.
Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
Reviewed-by: Nicolas Dichtel <nicolas.dichtel@6wind.com>
---
.../net/ethernet/cavium/octeon/octeon_mgmt.c | 4 +-
.../net/wireless/microchip/wilc1000/netdev.c | 8 ++--
include/linux/netdevice.h | 7 ++-
net/core/dev_addr_lists.c | 44 ++++++++++++++-----
net/core/dev_addr_lists_test.c | 18 ++++----
net/mac80211/driver-ops.h | 2 +-
6 files changed, 53 insertions(+), 30 deletions(-)
diff --git a/drivers/net/ethernet/cavium/octeon/octeon_mgmt.c b/drivers/net/ethernet/cavium/octeon/octeon_mgmt.c
index 2cf3365b96364..bb72c08337211 100644
--- a/drivers/net/ethernet/cavium/octeon/octeon_mgmt.c
+++ b/drivers/net/ethernet/cavium/octeon/octeon_mgmt.c
@@ -573,14 +573,14 @@ static void octeon_mgmt_set_rx_filtering(struct net_device *netdev)
memset(&cam_state, 0, sizeof(cam_state));
- if ((netdev->flags & IFF_PROMISC) || netdev->uc.count > 7) {
+ if ((netdev->flags & IFF_PROMISC) || netdev_uc_count(netdev) > 7) {
cam_mode = 0;
available_cam_entries = 8;
} else {
/* One CAM entry for the primary address, leaves seven
* for the secondary addresses.
*/
- available_cam_entries = 7 - netdev->uc.count;
+ available_cam_entries = 7 - netdev_uc_count(netdev);
}
if (netdev->flags & IFF_MULTICAST) {
diff --git a/drivers/net/wireless/microchip/wilc1000/netdev.c b/drivers/net/wireless/microchip/wilc1000/netdev.c
index 956cb578bf37c..d3343113cea50 100644
--- a/drivers/net/wireless/microchip/wilc1000/netdev.c
+++ b/drivers/net/wireless/microchip/wilc1000/netdev.c
@@ -704,17 +704,17 @@ static void wilc_set_multicast_list(struct net_device *dev)
return;
if (dev->flags & IFF_ALLMULTI ||
- dev->mc.count > WILC_MULTICAST_TABLE_SIZE) {
+ netdev_mc_count(dev) > WILC_MULTICAST_TABLE_SIZE) {
wilc_setup_multicast_filter(vif, 0, 0, NULL);
return;
}
- if (dev->mc.count == 0) {
+ if (netdev_mc_empty(dev)) {
wilc_setup_multicast_filter(vif, 1, 0, NULL);
return;
}
- mc_list = kmalloc_array(dev->mc.count, ETH_ALEN, GFP_ATOMIC);
+ mc_list = kmalloc_array(netdev_mc_count(dev), ETH_ALEN, GFP_ATOMIC);
if (!mc_list)
return;
@@ -727,7 +727,7 @@ static void wilc_set_multicast_list(struct net_device *dev)
cur_mc += ETH_ALEN;
}
- if (wilc_setup_multicast_filter(vif, 1, dev->mc.count, mc_list))
+ if (wilc_setup_multicast_filter(vif, 1, netdev_mc_count(dev), mc_list))
kfree(mc_list);
}
diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index f61c972313095..3d9e3a70ca5c9 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -252,13 +252,16 @@ struct netdev_hw_addr {
struct netdev_hw_addr_list {
struct list_head list;
- int count;
+ /* Only changed through the __hw_addr_count_* helpers after
+ * __hw_addr_init()
+ */
+ int _count;
/* Auxiliary tree for faster lookup on addition and deletion */
struct rb_root tree;
};
-#define netdev_hw_addr_list_count(l) ((l)->count)
+#define netdev_hw_addr_list_count(l) ((l)->_count)
#define netdev_hw_addr_list_empty(l) (netdev_hw_addr_list_count(l) == 0)
#define netdev_hw_addr_list_for_each(ha, l) \
list_for_each_entry(ha, &(l)->list, list)
diff --git a/net/core/dev_addr_lists.c b/net/core/dev_addr_lists.c
index 08528ca0a8b31..d615192d1c3b1 100644
--- a/net/core/dev_addr_lists.c
+++ b/net/core/dev_addr_lists.c
@@ -16,6 +16,26 @@
#include "dev.h"
+static void __hw_addr_count_add(struct netdev_hw_addr_list *list, int value)
+{
+ list->_count += value;
+}
+
+static void __hw_addr_count_inc(struct netdev_hw_addr_list *list)
+{
+ __hw_addr_count_add(list, 1);
+}
+
+static void __hw_addr_count_dec(struct netdev_hw_addr_list *list)
+{
+ __hw_addr_count_add(list, -1);
+}
+
+static void __hw_addr_count_reset(struct netdev_hw_addr_list *list)
+{
+ list->_count = 0;
+}
+
/*
* General list handling functions
*/
@@ -125,7 +145,7 @@ static int __hw_addr_add_ex(struct netdev_hw_addr_list *list,
rb_insert_color(&ha->node, &list->tree);
list_add_tail_rcu(&ha->list, &list->list);
- list->count++;
+ __hw_addr_count_inc(list);
return 0;
}
@@ -161,7 +181,7 @@ static int __hw_addr_del_entry(struct netdev_hw_addr_list *list,
list_del_rcu(&ha->list);
kfree_rcu(ha, rcu_head);
- list->count--;
+ __hw_addr_count_dec(list);
return 0;
}
@@ -492,14 +512,14 @@ void __hw_addr_flush(struct netdev_hw_addr_list *list)
list_del_rcu(&ha->list);
kfree_rcu(ha, rcu_head);
}
- list->count = 0;
+ __hw_addr_count_reset(list);
}
EXPORT_SYMBOL_IF_KUNIT(__hw_addr_flush);
void __hw_addr_init(struct netdev_hw_addr_list *list)
{
INIT_LIST_HEAD(&list->list);
- list->count = 0;
+ list->_count = 0;
list->tree = RB_ROOT;
}
EXPORT_SYMBOL(__hw_addr_init);
@@ -509,8 +529,8 @@ static void __hw_addr_splice(struct netdev_hw_addr_list *dst,
{
src->tree = RB_ROOT;
list_splice_init(&src->list, &dst->list);
- dst->count += src->count;
- src->count = 0;
+ __hw_addr_count_add(dst, netdev_hw_addr_list_count(src));
+ __hw_addr_count_reset(src);
}
/**
@@ -532,11 +552,11 @@ int __hw_addr_list_snapshot(struct netdev_hw_addr_list *snap,
struct netdev_hw_addr *ha, *entry;
list_for_each_entry(ha, &list->list, list) {
- if (cache->count) {
+ if (netdev_hw_addr_list_count(cache)) {
entry = list_first_entry(&cache->list,
struct netdev_hw_addr, list);
list_del(&entry->list);
- cache->count--;
+ __hw_addr_count_dec(cache);
memcpy(entry->addr, ha->addr, addr_len);
entry->type = ha->type;
entry->global_use = false;
@@ -554,7 +574,7 @@ int __hw_addr_list_snapshot(struct netdev_hw_addr_list *snap,
list_add_tail(&entry->list, &snap->list);
__hw_addr_insert(snap, entry, addr_len);
- snap->count++;
+ __hw_addr_count_inc(snap);
}
return 0;
@@ -604,14 +624,14 @@ void __hw_addr_list_reconcile(struct netdev_hw_addr_list *real_list,
if (delta > 0) {
rb_erase(&ref_ha->node, &ref->tree);
list_del(&ref_ha->list);
- ref->count--;
+ __hw_addr_count_dec(ref);
ref_ha->sync_cnt = delta;
ref_ha->refcount = delta;
list_add_tail_rcu(&ref_ha->list,
&real_list->list);
__hw_addr_insert(real_list, ref_ha,
addr_len);
- real_list->count++;
+ __hw_addr_count_inc(real_list);
}
continue;
}
@@ -622,7 +642,7 @@ void __hw_addr_list_reconcile(struct netdev_hw_addr_list *real_list,
rb_erase(&real_ha->node, &real_list->tree);
list_del_rcu(&real_ha->list);
kfree_rcu(real_ha, rcu_head);
- real_list->count--;
+ __hw_addr_count_dec(real_list);
}
}
diff --git a/net/core/dev_addr_lists_test.c b/net/core/dev_addr_lists_test.c
index 260e71a2399f3..07c35a0af2b4d 100644
--- a/net/core/dev_addr_lists_test.c
+++ b/net/core/dev_addr_lists_test.c
@@ -291,7 +291,7 @@ static void dev_addr_test_snapshot_sync(struct kunit *test)
netif_addr_unlock_bh(netdev);
/* Real entry should now reflect the sync: sync_cnt=1, refcount=2 */
- KUNIT_EXPECT_EQ(test, 1, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 1, netdev_uc_count(netdev));
ha = list_first_entry(&netdev->uc.list, struct netdev_hw_addr, list);
KUNIT_EXPECT_MEMEQ(test, ha->addr, addr, ETH_ALEN);
KUNIT_EXPECT_EQ(test, 1, ha->sync_cnt);
@@ -303,7 +303,7 @@ static void dev_addr_test_snapshot_sync(struct kunit *test)
dev_addr_test_unsync);
KUNIT_EXPECT_EQ(test, 0, datp->addr_synced);
KUNIT_EXPECT_EQ(test, 0, datp->addr_unsynced);
- KUNIT_EXPECT_EQ(test, 1, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 1, netdev_uc_count(netdev));
__hw_addr_flush(&cache);
rtnl_unlock();
@@ -351,7 +351,7 @@ static void dev_addr_test_snapshot_remove_during_sync(struct kunit *test)
/* Concurrent removal: user deletes ADDR_A while driver was working */
memset(addr, ADDR_A, sizeof(addr));
KUNIT_EXPECT_EQ(test, 0, dev_uc_del(netdev, addr));
- KUNIT_EXPECT_EQ(test, 0, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 0, netdev_uc_count(netdev));
/* Reconcile: ADDR_A gone from real list but driver synced it,
* so it gets re-inserted as stale (sync_cnt=1, refcount=1).
@@ -361,7 +361,7 @@ static void dev_addr_test_snapshot_remove_during_sync(struct kunit *test)
&cache);
netif_addr_unlock_bh(netdev);
- KUNIT_EXPECT_EQ(test, 1, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 1, netdev_uc_count(netdev));
ha = list_first_entry(&netdev->uc.list, struct netdev_hw_addr, list);
KUNIT_EXPECT_MEMEQ(test, ha->addr, addr, ETH_ALEN);
KUNIT_EXPECT_EQ(test, 1, ha->sync_cnt);
@@ -373,7 +373,7 @@ static void dev_addr_test_snapshot_remove_during_sync(struct kunit *test)
dev_addr_test_unsync);
KUNIT_EXPECT_EQ(test, 0, datp->addr_synced);
KUNIT_EXPECT_EQ(test, 1 << ADDR_A, datp->addr_unsynced);
- KUNIT_EXPECT_EQ(test, 0, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 0, netdev_uc_count(netdev));
__hw_addr_flush(&cache);
rtnl_unlock();
@@ -433,7 +433,7 @@ static void dev_addr_test_snapshot_readd_during_unsync(struct kunit *test)
* stale entry and bumps refcount from 1 -> 2. sync_cnt stays 1.
*/
KUNIT_EXPECT_EQ(test, 0, dev_uc_add(netdev, addr));
- KUNIT_EXPECT_EQ(test, 1, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 1, netdev_uc_count(netdev));
/* Reconcile: ref sync_cnt=1 matches real sync_cnt=1, delta=-1
* applied. Result: sync_cnt=0, refcount=1 (fresh).
@@ -444,7 +444,7 @@ static void dev_addr_test_snapshot_readd_during_unsync(struct kunit *test)
netif_addr_unlock_bh(netdev);
/* Entry survives as fresh: needs re-sync to HW */
- KUNIT_EXPECT_EQ(test, 1, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 1, netdev_uc_count(netdev));
ha = list_first_entry(&netdev->uc.list, struct netdev_hw_addr, list);
KUNIT_EXPECT_MEMEQ(test, ha->addr, addr, ETH_ALEN);
KUNIT_EXPECT_EQ(test, 0, ha->sync_cnt);
@@ -528,7 +528,7 @@ static void dev_addr_test_snapshot_add_and_remove(struct kunit *test)
* ADDR_B: refcount went from 2->1 via dev_uc_del (still present, stale)
* ADDR_C: sync propagated (sync_cnt=1, refcount=2)
*/
- KUNIT_EXPECT_EQ(test, 3, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 3, netdev_uc_count(netdev));
netdev_hw_addr_list_for_each(ha, &netdev->uc) {
u8 id = ha->addr[0];
@@ -553,7 +553,7 @@ static void dev_addr_test_snapshot_add_and_remove(struct kunit *test)
dev_addr_test_unsync);
KUNIT_EXPECT_EQ(test, 0, datp->addr_synced);
KUNIT_EXPECT_EQ(test, 1 << ADDR_B, datp->addr_unsynced);
- KUNIT_EXPECT_EQ(test, 2, netdev->uc.count);
+ KUNIT_EXPECT_EQ(test, 2, netdev_uc_count(netdev));
__hw_addr_flush(&cache);
rtnl_unlock();
diff --git a/net/mac80211/driver-ops.h b/net/mac80211/driver-ops.h
index f1c0b87fddd5f..e80731c59ef50 100644
--- a/net/mac80211/driver-ops.h
+++ b/net/mac80211/driver-ops.h
@@ -187,7 +187,7 @@ static inline u64 drv_prepare_multicast(struct ieee80211_local *local,
{
u64 ret = 0;
- trace_drv_prepare_multicast(local, mc_list->count);
+ trace_drv_prepare_multicast(local, netdev_hw_addr_list_count(mc_list));
if (local->ops->prepare_multicast)
ret = local->ops->prepare_multicast(&local->hw, mc_list);
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH net-next v10 3/6] net: add a generation counter for dev->mc changes
2026-10-07 11:58 [PATCH net-next v10 0/6] rtnetlink: dump link-layer multicast addresses Yuyang Huang
2026-10-07 11:58 ` [PATCH net-next v10 1/6] netlink: specs: rt-addr: fix the type of target-netnsid Yuyang Huang
2026-10-07 11:58 ` [PATCH net-next v10 2/6] net: change netdev_hw_addr_list count through helpers Yuyang Huang
@ 2026-10-07 11:58 ` Yuyang Huang
2026-10-09 9:50 ` Nicolas Dichtel
2026-10-09 11:58 ` netdev-bot+sashiko
2026-10-07 11:58 ` [PATCH net-next v10 4/6] net: add AF_PACKET multicast dumps Yuyang Huang
` (2 subsequent siblings)
5 siblings, 2 replies; 12+ messages in thread
From: Yuyang Huang @ 2026-10-07 11:58 UTC (permalink / raw)
To: Yuyang Huang
Cc: Ajay Singh, Aleksandr Loktionov, Andrew Lunn, Claudiu Beznea,
David S. Miller, David Ahern, Donald Hunter, Eric Dumazet,
Ido Schimmel, Jacob Keller, Jakub Kicinski, Johannes Berg,
Kees Cook, Kory Maincent, Kuniyuki Iwashima, Nicolas Dichtel,
Nikolaos Gkarlis, Paolo Abeni, Sabrina Dubroca, Shuah Khan,
Simon Horman, Stanislav Fomichev, Vadim Fedorenko,
Willem de Bruijn, linux-kernel, linux-kselftest, linux-wireless,
netdev
A multi-part RTM_GETMULTICAST dump of dev->mc resumes by position, so
entries added or removed between two dump rounds can be skipped or
repeated. The IPv4 and IPv6 dumps report that with NLM_F_DUMP_INTR by
stamping cb->seq from a per netns generation counter combined with
dev_base_seq, see inet_base_seq().
Add the equivalent for the device multicast lists: a per netns counter
bumped whenever an entry is added to or removed from any dev->mc. The
list helpers do not know which device a list belongs to, so give
netdev_hw_addr_list an mc_dev pointer, set for dev->mc only, and bump
the counter of dev_net(mc_dev) from the count helpers. That covers the
dev_mc_* helpers, both lists of a sync, the hardware sync helpers
drivers call from their rx mode callbacks or their own workers and the
reconciliation after an asynchronous rx mode update. It is atomic
since the writers only hold the address lock of their own device.
Used by the following patch for the AF_PACKET multicast dump.
Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
---
include/linux/netdevice.h | 3 +++
include/net/net_namespace.h | 1 +
net/core/dev_addr_lists.c | 15 +++++++++++++++
3 files changed, 19 insertions(+)
diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index 3d9e3a70ca5c9..6c837ea34748e 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -259,6 +259,9 @@ struct netdev_hw_addr_list {
/* Auxiliary tree for faster lookup on addition and deletion */
struct rb_root tree;
+
+ /* Set for dev->mc only, the device the list belongs to */
+ struct net_device *mc_dev;
};
#define netdev_hw_addr_list_count(l) ((l)->_count)
diff --git a/include/net/net_namespace.h b/include/net/net_namespace.h
index 46b4c67e2966f..c29a22ca178fa 100644
--- a/include/net/net_namespace.h
+++ b/include/net/net_namespace.h
@@ -71,6 +71,7 @@ struct net {
spinlock_t rules_mod_lock;
unsigned int dev_base_seq; /* protected by rtnl_mutex */
+ atomic_t dev_mc_genid; /* dev->mc membership changes */
u32 ifindex;
spinlock_t nsid_lock;
diff --git a/net/core/dev_addr_lists.c b/net/core/dev_addr_lists.c
index d615192d1c3b1..3ab053cbf8b99 100644
--- a/net/core/dev_addr_lists.c
+++ b/net/core/dev_addr_lists.c
@@ -16,9 +16,19 @@
#include "dev.h"
+/* RTM_GETMULTICAST dumps use the netns generation counter to detect
+ * changes of dev->mc between dump rounds.
+ */
+static void __hw_addr_changed(struct netdev_hw_addr_list *list)
+{
+ if (list->mc_dev)
+ atomic_inc(&dev_net(list->mc_dev)->dev_mc_genid);
+}
+
static void __hw_addr_count_add(struct netdev_hw_addr_list *list, int value)
{
list->_count += value;
+ __hw_addr_changed(list);
}
static void __hw_addr_count_inc(struct netdev_hw_addr_list *list)
@@ -33,7 +43,10 @@ static void __hw_addr_count_dec(struct netdev_hw_addr_list *list)
static void __hw_addr_count_reset(struct netdev_hw_addr_list *list)
{
+ if (!list->_count)
+ return;
list->_count = 0;
+ __hw_addr_changed(list);
}
/*
@@ -521,6 +534,7 @@ void __hw_addr_init(struct netdev_hw_addr_list *list)
INIT_LIST_HEAD(&list->list);
list->_count = 0;
list->tree = RB_ROOT;
+ list->mc_dev = NULL;
}
EXPORT_SYMBOL(__hw_addr_init);
@@ -1197,6 +1211,7 @@ EXPORT_SYMBOL(dev_mc_flush);
void dev_mc_init(struct net_device *dev)
{
__hw_addr_init(&dev->mc);
+ dev->mc.mc_dev = dev;
}
EXPORT_SYMBOL(dev_mc_init);
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread* Re: [PATCH net-next v10 3/6] net: add a generation counter for dev->mc changes
2026-10-07 11:58 ` [PATCH net-next v10 3/6] net: add a generation counter for dev->mc changes Yuyang Huang
@ 2026-10-09 9:50 ` Nicolas Dichtel
2026-10-09 11:58 ` netdev-bot+sashiko
1 sibling, 0 replies; 12+ messages in thread
From: Nicolas Dichtel @ 2026-10-09 9:50 UTC (permalink / raw)
To: Yuyang Huang
Cc: Ajay Singh, Aleksandr Loktionov, Andrew Lunn, Claudiu Beznea,
David S. Miller, David Ahern, Donald Hunter, Eric Dumazet,
Ido Schimmel, Jacob Keller, Jakub Kicinski, Johannes Berg,
Kees Cook, Kory Maincent, Kuniyuki Iwashima, Nikolaos Gkarlis,
Paolo Abeni, Sabrina Dubroca, Shuah Khan, Simon Horman,
Stanislav Fomichev, Vadim Fedorenko, Willem de Bruijn,
linux-kernel, linux-kselftest, linux-wireless, netdev
Le 07/10/2026 à 13:58, Yuyang Huang a écrit :
> A multi-part RTM_GETMULTICAST dump of dev->mc resumes by position, so
> entries added or removed between two dump rounds can be skipped or
> repeated. The IPv4 and IPv6 dumps report that with NLM_F_DUMP_INTR by
> stamping cb->seq from a per netns generation counter combined with
> dev_base_seq, see inet_base_seq().
>
> Add the equivalent for the device multicast lists: a per netns counter
> bumped whenever an entry is added to or removed from any dev->mc. The
> list helpers do not know which device a list belongs to, so give
> netdev_hw_addr_list an mc_dev pointer, set for dev->mc only, and bump
> the counter of dev_net(mc_dev) from the count helpers. That covers the
> dev_mc_* helpers, both lists of a sync, the hardware sync helpers
> drivers call from their rx mode callbacks or their own workers and the
> reconciliation after an asynchronous rx mode update. It is atomic
> since the writers only hold the address lock of their own device.
>
> Used by the following patch for the AF_PACKET multicast dump.
>
> Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
Reviewed-by: Nicolas Dichtel <nicolas.dichtel@6wind.com>
^ permalink raw reply [flat|nested] 12+ messages in thread
* Re: [PATCH net-next v10 3/6] net: add a generation counter for dev->mc changes
2026-10-07 11:58 ` [PATCH net-next v10 3/6] net: add a generation counter for dev->mc changes Yuyang Huang
2026-10-09 9:50 ` Nicolas Dichtel
@ 2026-10-09 11:58 ` netdev-bot+sashiko
1 sibling, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-10-09 11:58 UTC (permalink / raw)
To: sigefriedhyy
Cc: ajay.kathat, aleksandr.loktionov, andrew+netdev, claudiu.beznea,
davem, dsahern, donald.hunter, edumazet, idosch, jacob.e.keller,
kuba, johannes, kees, kory.maincent, kuniyu, nicolas.dichtel,
nickgarlis, pabeni, sd, shuah, horms, sdf.kernel, vadim.fedorenko,
willemb, linux-kernel, linux-kselftest, linux-wireless, netdev
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] netlink: AF_PACKET multicast dump misses changes due to
reading sequence after loop
- [Low] __hw_addr_changed() (net/core/dev_addr_lists.c) does
atomic_inc(&dev_net(list->mc_dev)->dev_mc_genid).
--
net: add a generation counter for dev->mc changes
This commit adds a per netns generation counter, dev_mc_genid, that is
bumped whenever an entry is added to or removed from any dev->mc list.
To find the netns, netdev_hw_addr_list gets an mc_dev pointer, which is
set only for dev->mc, and the count helpers bump
dev_net(mc_dev)->dev_mc_genid. The counter will be used by a following
patch for the AF_PACKET multicast dump.
> diff --git a/net/core/dev_addr_lists.c b/net/core/dev_addr_lists.c
> index d615192d1c3b1..3ab053cbf8b99 100644
> --- a/net/core/dev_addr_lists.c
> +++ b/net/core/dev_addr_lists.c
> @@ -16,9 +16,19 @@
>
> #include "dev.h"
>
> +/* RTM_GETMULTICAST dumps use the netns generation counter to detect
> + * changes of dev->mc between dump rounds.
> + */
> +static void __hw_addr_changed(struct netdev_hw_addr_list *list)
> +{
> + if (list->mc_dev)
> + atomic_inc(&dev_net(list->mc_dev)->dev_mc_genid);
[Severity: Medium]
This counter only catches changes if the dump samples it before it
walks the lists, but the user in the following patch seems to sample it
afterwards. dev_mc_dump() assigns cb->seq from
dev_mc_base_seq(filter.tgt_net) only after the for_each_netdev_dump()
loop has finished.
The reader holds netif_addr_lock_bh() while it walks each dev->mc, so
the ordering of the atomic_inc() against the list update is not the
problem. The problem is the window between releasing the address lock
of the last device dumped in a round and taking the sample. For
example:
- a round fills the skb partway through eth0's dev->mc and stops,
recording the resume position
- eth0's addr_list_lock is released
- another CPU deletes an eth0 entry that sits before the resume
position and bumps dev_mc_genid
- dev_mc_dump() then stores the already bumped value in cb->seq
- the next round resumes eth0 at the stored position and skips one
entry, but the genid it sees matches the value stamped at the end
of the previous round, so nl_dump_check_consistent() has nothing
to flag and NLM_F_DUMP_INTR is never set
An add in the same window gives a duplicated entry instead. Either way
user space receives an inconsistent dump with no hint that it should
retry, which is what this counter is meant to prevent.
The commit message names inet_base_seq() as the model. In
inet_dump_ifaddr() it is sampled under rcu_read_lock() before
for_each_netdev_dump() is entered, so any change made during or after
the traversal shows up as a different value in the next round. Should
dev_mc_dump() take its sample before the loop in the same way?
[Severity: Low]
Is dev_net() safe to use here? __hw_addr_changed() runs with only the
device's addr_list_lock held. __dev_change_net_namespace() switches
dev->nd_net under RTNL and netdev_lock(dev), and it never takes
addr_list_lock:
net/core/dev.c:__dev_change_net_namespace() {
...
netdev_lock(dev);
dev_net_set(dev, net);
netdev_unlock(dev);
...
}
So a dev->mc writer that holds neither RTNL nor the instance lock can
run at the same time as a netns move. The rxe driver looks like one such
caller:
rxe_attach_mcast()
rxe_get_mcg()
rxe_mcast_add()
dev_mc_add(ndev)
__dev_mc_add()
__hw_addr_add_ex()
__hw_addr_count_inc()
__hw_addr_changed()
dev_net(ndev)
rxe_mcast_add() only holds the reference it gets from
ib_device_get_netdev(). On NETDEV_UNREGISTER, rxe_notify() only queues
the ib device unregistration, and rxe_net_del() does not clear the port
netdev. That means this path can still run while something like
"ip link set <ndev> netns X" executes.
dev_net() is read_pnet(), which does a plain load through
rcu_dereference_protected(pnet->net, true). That load races with the
rcu_assign_pointer() in dev_net_set().
Would KCSAN report this as a data race? And could the increment then
land on the old netns's dev_mc_genid instead of the new one?
The old struct net can't be freed under the BH-disabled spinlock, and
dev_base_seq covers the device appearing in the new netns. So the impact
looks limited to the race report and one bump going to the wrong netns.
Would dev_net_rcu() be a better fit for this reader?
> +}
> +
> static void __hw_addr_count_add(struct netdev_hw_addr_list *list, int value)
> {
> list->_count += value;
> + __hw_addr_changed(list);
> }
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261007115821.11332-1-sigefriedhyy%40gmail.com
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH net-next v10 4/6] net: add AF_PACKET multicast dumps
2026-10-07 11:58 [PATCH net-next v10 0/6] rtnetlink: dump link-layer multicast addresses Yuyang Huang
` (2 preceding siblings ...)
2026-10-07 11:58 ` [PATCH net-next v10 3/6] net: add a generation counter for dev->mc changes Yuyang Huang
@ 2026-10-07 11:58 ` Yuyang Huang
2026-10-09 9:57 ` Nicolas Dichtel
2026-10-09 11:58 ` netdev-bot+sashiko
2026-10-07 11:58 ` [PATCH net-next v10 5/6] netlink: specs: rt-addr: document " Yuyang Huang
2026-10-07 11:58 ` [PATCH net-next v10 6/6] selftests: net: test " Yuyang Huang
5 siblings, 2 replies; 12+ messages in thread
From: Yuyang Huang @ 2026-10-07 11:58 UTC (permalink / raw)
To: Yuyang Huang
Cc: Ajay Singh, Aleksandr Loktionov, Andrew Lunn, Claudiu Beznea,
David S. Miller, David Ahern, Donald Hunter, Eric Dumazet,
Ido Schimmel, Jacob Keller, Jakub Kicinski, Johannes Berg,
Kees Cook, Kory Maincent, Kuniyuki Iwashima, Nicolas Dichtel,
Nikolaos Gkarlis, Paolo Abeni, Sabrina Dubroca, Shuah Khan,
Simon Horman, Stanislav Fomichev, Vadim Fedorenko,
Willem de Bruijn, linux-kernel, linux-kselftest, linux-wireless,
netdev
RTM_GETMULTICAST dumps IPv4 and IPv6 multicast group memberships, but
the device multicast list (dev->mc) is only available through
/proc/net/dev_mcast, so "ip maddr show" still has to parse procfs for
its link-layer entries.
Handle RTM_GETMULTICAST dumps with ifa_family set to AF_PACKET next to
the dev->mc helpers in dev_addr_lists.c and report every entry of
dev->mc in the existing ifaddrmsg format:
- IFA_MULTICAST carries the raw link-layer address
- IFA_MC_USERS carries the entry reference count
- IFA_F_GLOBAL in IFA_FLAGS reports netdev_hw_addr::global_use, set
by dev_mc_add_global() (SIOCADDMULTI) and dev_mc_add_excl()
("bridge fdb add ... self"), i.e. entries added explicitly rather
than by a protocol join. This is the static column of
/proc/net/dev_mcast
- ifa_scope is RT_SCOPE_LINK
This covers every column of /proc/net/dev_mcast. AF_PACKET is the
family iproute2 already uses for link-layer addresses ("ip -0").
The default FDB dump also walks dev->mc, but only for Ethernet devices
and only if the driver uses it, vxlan for example does not, and it has
no users count or global_use bit. Extending it would change "bridge fdb
show" output and add NDA_* attributes.
Requests are always validated, there are no legacy users: prefixlen,
flags and scope must be zero, ifa_index selects one device and
IFA_TARGET_NETNSID is the only attribute accepted. The dump runs under
RCU and netif_addr_lock_bh() without RTNL. After each round cb->seq is
set from the dev->mc generation counter and dev_base_seq and checked
like rtnl_dump_ifinfo() does, so an entry added or removed since the
previous round sets NLM_F_DUMP_INTR.
Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
---
include/linux/netdevice.h | 1 +
include/uapi/linux/if_addr.h | 1 +
net/core/dev_addr_lists.c | 186 +++++++++++++++++++++++++++++++++++
net/core/rtnetlink.c | 2 +
4 files changed, 190 insertions(+)
diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index 6c837ea34748e..296eef148ef63 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -5171,6 +5171,7 @@ int dev_mc_sync_multiple(struct net_device *to, struct net_device *from);
void dev_mc_unsync(struct net_device *to, struct net_device *from);
void dev_mc_flush(struct net_device *dev);
void dev_mc_init(struct net_device *dev);
+int dev_mc_dump(struct sk_buff *skb, struct netlink_callback *cb);
/**
* __dev_mc_sync - Synchronize device's multicast list
diff --git a/include/uapi/linux/if_addr.h b/include/uapi/linux/if_addr.h
index 7fb630b7fe311..0a1ad9ebb47be 100644
--- a/include/uapi/linux/if_addr.h
+++ b/include/uapi/linux/if_addr.h
@@ -57,6 +57,7 @@ enum {
#define IFA_F_NOPREFIXROUTE 0x200
#define IFA_F_MCAUTOJOIN 0x400
#define IFA_F_STABLE_PRIVACY 0x800
+#define IFA_F_GLOBAL 0x1000
struct ifa_cacheinfo {
__u32 ifa_prefered;
diff --git a/net/core/dev_addr_lists.c b/net/core/dev_addr_lists.c
index 3ab053cbf8b99..7b9abeca431f6 100644
--- a/net/core/dev_addr_lists.c
+++ b/net/core/dev_addr_lists.c
@@ -10,8 +10,12 @@
#include <linux/netdevice.h>
#include <linux/rtnetlink.h>
#include <linux/export.h>
+#include <linux/if_addr.h>
#include <linux/list.h>
#include <linux/spinlock.h>
+#include <net/netlink.h>
+#include <net/rtnetlink.h>
+#include <net/sock.h>
#include <kunit/visibility.h>
#include "dev.h"
@@ -1215,6 +1219,188 @@ void dev_mc_init(struct net_device *dev)
}
EXPORT_SYMBOL(dev_mc_init);
+static int dev_mc_fill_addr(struct sk_buff *skb, const struct net_device *dev,
+ const struct netdev_hw_addr *ha, u32 portid,
+ u32 seq, unsigned int flags, int netnsid)
+{
+ struct ifaddrmsg *ifm;
+ struct nlmsghdr *nlh;
+
+ nlh = nlmsg_put(skb, portid, seq, RTM_GETMULTICAST, sizeof(*ifm),
+ flags);
+ if (!nlh)
+ return -EMSGSIZE;
+
+ ifm = nlmsg_data(nlh);
+ ifm->ifa_family = AF_PACKET;
+ ifm->ifa_prefixlen = 0;
+ ifm->ifa_flags = 0;
+ ifm->ifa_scope = RT_SCOPE_LINK;
+ ifm->ifa_index = dev->ifindex;
+
+ if ((netnsid >= 0 &&
+ nla_put_s32(skb, IFA_TARGET_NETNSID, netnsid)) ||
+ nla_put(skb, IFA_MULTICAST, dev->addr_len, ha->addr) ||
+ nla_put_u32(skb, IFA_MC_USERS, ha->refcount) ||
+ nla_put_u32(skb, IFA_FLAGS, ha->global_use ? IFA_F_GLOBAL : 0)) {
+ nlmsg_cancel(skb, nlh);
+ return -EMSGSIZE;
+ }
+
+ nlmsg_end(skb, nlh);
+ return 0;
+}
+
+/* Combine dev_mc_genid and dev_base_seq to detect changes, like
+ * inet_base_seq().
+ */
+static u32 dev_mc_base_seq(const struct net *net)
+{
+ u32 res = atomic_read(&net->dev_mc_genid) +
+ READ_ONCE(net->dev_base_seq);
+
+ /* Must not return 0 (see nl_dump_check_consistent()). */
+ if (!res)
+ res = 0x80000000;
+ return res;
+}
+
+static int dev_mc_dump_dev(struct net_device *dev, struct sk_buff *skb,
+ struct netlink_callback *cb, int *s_addr_idx,
+ unsigned int flags, int netnsid)
+{
+ struct netdev_hw_addr *ha;
+ int addr_idx = 0;
+ int err = 0;
+
+ netif_addr_lock_bh(dev);
+ netdev_for_each_mc_addr(ha, dev) {
+ if (addr_idx < *s_addr_idx) {
+ addr_idx++;
+ continue;
+ }
+ err = dev_mc_fill_addr(skb, dev, ha, NETLINK_CB(cb->skb).portid,
+ cb->nlh->nlmsg_seq, flags, netnsid);
+ if (err < 0)
+ break;
+ addr_idx++;
+ }
+ netif_addr_unlock_bh(dev);
+
+ *s_addr_idx = err < 0 ? addr_idx : 0;
+
+ return err;
+}
+
+struct dev_mc_dump_filter {
+ struct net *tgt_net;
+ netns_tracker ns_tracker;
+ int netnsid;
+ int ifindex;
+};
+
+static const struct nla_policy dev_mc_dump_policy[IFA_MAX + 1] = {
+ [IFA_TARGET_NETNSID] = { .type = NLA_S32 },
+};
+
+static int dev_mc_valid_dump_req(const struct nlmsghdr *nlh, struct sock *sk,
+ struct dev_mc_dump_filter *filter,
+ struct netlink_ext_ack *extack)
+{
+ struct nlattr *tb[IFA_MAX + 1];
+ struct ifaddrmsg *ifm;
+ int err;
+
+ ifm = nlmsg_payload(nlh, sizeof(*ifm));
+ if (!ifm) {
+ NL_SET_ERR_MSG(extack,
+ "Invalid header for multicast dump request");
+ return -EINVAL;
+ }
+
+ if (ifm->ifa_prefixlen || ifm->ifa_flags || ifm->ifa_scope) {
+ NL_SET_ERR_MSG(extack,
+ "Invalid values in multicast dump header");
+ return -EINVAL;
+ }
+
+ err = nlmsg_parse(nlh, sizeof(*ifm), tb, IFA_MAX,
+ dev_mc_dump_policy, extack);
+ if (err < 0)
+ return err;
+
+ if (tb[IFA_TARGET_NETNSID]) {
+ struct net *net;
+
+ filter->netnsid = nla_get_s32(tb[IFA_TARGET_NETNSID]);
+ net = rtnl_get_net_ns_capable(sk, filter->netnsid);
+ if (IS_ERR(net)) {
+ NL_SET_ERR_MSG(extack,
+ "Invalid target network namespace id");
+ return PTR_ERR(net);
+ }
+ netns_tracker_alloc(net, &filter->ns_tracker, GFP_KERNEL);
+ filter->tgt_net = net;
+ }
+
+ filter->ifindex = ifm->ifa_index;
+
+ return 0;
+}
+
+int dev_mc_dump(struct sk_buff *skb, struct netlink_callback *cb)
+{
+ struct dev_mc_dump_filter filter = {
+ .tgt_net = sock_net(skb->sk),
+ .netnsid = -1,
+ };
+ unsigned int flags = NLM_F_MULTI;
+ struct {
+ unsigned long ifindex;
+ int addr_idx;
+ } *ctx = (void *)cb->ctx;
+ unsigned long s_ifindex;
+ struct net_device *dev;
+ int err;
+
+ err = dev_mc_valid_dump_req(cb->nlh, skb->sk, &filter, cb->extack);
+ if (err < 0)
+ return err;
+
+ rcu_read_lock();
+
+ if (filter.ifindex) {
+ cb->answer_flags |= NLM_F_DUMP_FILTERED;
+ flags |= NLM_F_DUMP_FILTERED;
+ dev = dev_get_by_index_rcu(filter.tgt_net, filter.ifindex);
+ if (dev)
+ err = dev_mc_dump_dev(dev, skb, cb, &ctx->addr_idx,
+ flags, filter.netnsid);
+ else
+ err = -ENODEV;
+ } else {
+ s_ifindex = ctx->ifindex;
+ for_each_netdev_dump(filter.tgt_net, dev, ctx->ifindex) {
+ /* The device the dump stopped at is gone, do not
+ * skip entries of the next one.
+ */
+ if (dev->ifindex != s_ifindex)
+ ctx->addr_idx = 0;
+ err = dev_mc_dump_dev(dev, skb, cb, &ctx->addr_idx,
+ flags, filter.netnsid);
+ if (err < 0)
+ break;
+ }
+ }
+ cb->seq = dev_mc_base_seq(filter.tgt_net);
+ nl_dump_check_consistent(cb, nlmsg_hdr(skb));
+
+ rcu_read_unlock();
+ if (filter.netnsid >= 0)
+ put_net_track(filter.tgt_net, &filter.ns_tracker);
+ return err;
+}
+
static int netif_addr_lists_snapshot(struct net_device *dev,
struct netdev_hw_addr_list *uc_snap,
struct netdev_hw_addr_list *mc_snap,
diff --git a/net/core/rtnetlink.c b/net/core/rtnetlink.c
index 5bbebce0c2f50..1ec59b439986f 100644
--- a/net/core/rtnetlink.c
+++ b/net/core/rtnetlink.c
@@ -7282,6 +7282,8 @@ static const struct rtnl_msg_handler rtnetlink_rtnl_msg_handlers[] __initconst =
{.msgtype = RTM_SETSTATS, .doit = rtnl_stats_set},
{.msgtype = RTM_NEWLINKPROP, .doit = rtnl_newlinkprop},
{.msgtype = RTM_DELLINKPROP, .doit = rtnl_dellinkprop},
+ {.protocol = PF_PACKET, .msgtype = RTM_GETMULTICAST,
+ .dumpit = dev_mc_dump, .flags = RTNL_FLAG_DUMP_UNLOCKED},
{.protocol = PF_BRIDGE, .msgtype = RTM_GETLINK,
.dumpit = rtnl_bridge_getlink},
{.protocol = PF_BRIDGE, .msgtype = RTM_DELLINK,
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread* Re: [PATCH net-next v10 4/6] net: add AF_PACKET multicast dumps
2026-10-07 11:58 ` [PATCH net-next v10 4/6] net: add AF_PACKET multicast dumps Yuyang Huang
@ 2026-10-09 9:57 ` Nicolas Dichtel
2026-10-09 11:58 ` netdev-bot+sashiko
1 sibling, 0 replies; 12+ messages in thread
From: Nicolas Dichtel @ 2026-10-09 9:57 UTC (permalink / raw)
To: Yuyang Huang
Cc: Ajay Singh, Aleksandr Loktionov, Andrew Lunn, Claudiu Beznea,
David S. Miller, David Ahern, Donald Hunter, Eric Dumazet,
Ido Schimmel, Jacob Keller, Jakub Kicinski, Johannes Berg,
Kees Cook, Kory Maincent, Kuniyuki Iwashima, Nikolaos Gkarlis,
Paolo Abeni, Sabrina Dubroca, Shuah Khan, Simon Horman,
Stanislav Fomichev, Vadim Fedorenko, Willem de Bruijn,
linux-kernel, linux-kselftest, linux-wireless, netdev
Le 07/10/2026 à 13:58, Yuyang Huang a écrit :
> RTM_GETMULTICAST dumps IPv4 and IPv6 multicast group memberships, but
> the device multicast list (dev->mc) is only available through
> /proc/net/dev_mcast, so "ip maddr show" still has to parse procfs for
> its link-layer entries.
>
> Handle RTM_GETMULTICAST dumps with ifa_family set to AF_PACKET next to
> the dev->mc helpers in dev_addr_lists.c and report every entry of
> dev->mc in the existing ifaddrmsg format:
>
> - IFA_MULTICAST carries the raw link-layer address
> - IFA_MC_USERS carries the entry reference count
> - IFA_F_GLOBAL in IFA_FLAGS reports netdev_hw_addr::global_use, set
> by dev_mc_add_global() (SIOCADDMULTI) and dev_mc_add_excl()
> ("bridge fdb add ... self"), i.e. entries added explicitly rather
> than by a protocol join. This is the static column of
> /proc/net/dev_mcast
> - ifa_scope is RT_SCOPE_LINK
>
> This covers every column of /proc/net/dev_mcast. AF_PACKET is the
> family iproute2 already uses for link-layer addresses ("ip -0").
>
> The default FDB dump also walks dev->mc, but only for Ethernet devices
> and only if the driver uses it, vxlan for example does not, and it has
> no users count or global_use bit. Extending it would change "bridge fdb
> show" output and add NDA_* attributes.
>
> Requests are always validated, there are no legacy users: prefixlen,
> flags and scope must be zero, ifa_index selects one device and
> IFA_TARGET_NETNSID is the only attribute accepted. The dump runs under
> RCU and netif_addr_lock_bh() without RTNL. After each round cb->seq is
> set from the dev->mc generation counter and dev_base_seq and checked
> like rtnl_dump_ifinfo() does, so an entry added or removed since the
> previous round sets NLM_F_DUMP_INTR.
>
> Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
Reviewed-by: Nicolas Dichtel <nicolas.dichtel@6wind.com>
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH net-next v10 4/6] net: add AF_PACKET multicast dumps
2026-10-07 11:58 ` [PATCH net-next v10 4/6] net: add AF_PACKET multicast dumps Yuyang Huang
2026-10-09 9:57 ` Nicolas Dichtel
@ 2026-10-09 11:58 ` netdev-bot+sashiko
1 sibling, 0 replies; 12+ messages in thread
From: netdev-bot+sashiko @ 2026-10-09 11:58 UTC (permalink / raw)
To: sigefriedhyy
Cc: ajay.kathat, aleksandr.loktionov, andrew+netdev, claudiu.beznea,
davem, dsahern, donald.hunter, edumazet, idosch, jacob.e.keller,
kuba, johannes, kees, kory.maincent, kuniyu, nicolas.dichtel,
nickgarlis, pabeni, sd, shuah, horms, sdf.kernel, vadim.fedorenko,
willemb, linux-kernel, linux-kselftest, linux-wireless, netdev
Thank you for your contribution! Sashiko AI review found 3 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 2 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] dev_mc_dump() (net/core/dev_addr_lists.c) calls
nl_dump_check_consistent(cb, nlmsg_hdr(skb)) once at the end of every…
- [Medium] dev_mc_dump() reads the generation (cb->seq =
dev_mc_base_seq()) only after the walk, once dev_mc_dump_dev() has…
- [Low] At this commit, Documentation/netlink/specs/rt-addr.yaml is not
updated for the new AF_PACKET RTM_GETMULTICAST dump.
--
net: add AF_PACKET multicast dumps
This adds RTM_GETMULTICAST dump support for ifa_family AF_PACKET. Each
dev->mc entry is reported in ifaddrmsg format with its link-layer
address, refcount and global_use flag. The dump runs under RCU and the
device addr_list_lock without RTNL, and uses a dev->mc generation counter
to set NLM_F_DUMP_INTR.
> diff --git a/include/uapi/linux/if_addr.h b/include/uapi/linux/if_addr.h
> index 7fb630b7fe311..0a1ad9ebb47be 100644
> --- a/include/uapi/linux/if_addr.h
> +++ b/include/uapi/linux/if_addr.h
> @@ -57,6 +57,7 @@ enum {
> #define IFA_F_NOPREFIXROUTE 0x200
> #define IFA_F_MCAUTOJOIN 0x400
> #define IFA_F_STABLE_PRIVACY 0x800
> +#define IFA_F_GLOBAL 0x1000
[Severity: Low]
This isn't a bug, but this commit does not update
Documentation/netlink/specs/rt-addr.yaml for the new flag or the
AF_PACKET dump:
- the ifa-flags enum has no entry for IFA_F_GLOBAL
- the getmulticast doc still says "Get / dump IPv4/IPv6 multicast
addresses."
- the reply attribute list does not include flags or target-netnsid
- the dump request lists no attributes
The next patch in the series, "netlink: specs: rt-addr: document
AF_PACKET multicast dumps", appears to fix all of these. The final patch,
"selftests: net: test AF_PACKET multicast dumps", adds a selftest. So the
mismatch only affects ynl decoding at this intermediate commit.
>
> struct ifa_cacheinfo {
> __u32 ifa_prefered;
> diff --git a/net/core/dev_addr_lists.c b/net/core/dev_addr_lists.c
> index 3ab053cbf8b99..7b9abeca431f6 100644
> --- a/net/core/dev_addr_lists.c
> +++ b/net/core/dev_addr_lists.c
[ ... ]
> +/* Combine dev_mc_genid and dev_base_seq to detect changes, like
> + * inet_base_seq().
> + */
> +static u32 dev_mc_base_seq(const struct net *net)
> +{
> + u32 res = atomic_read(&net->dev_mc_genid) +
> + READ_ONCE(net->dev_base_seq);
[ ... ]
> +static int dev_mc_dump_dev(struct net_device *dev, struct sk_buff *skb,
> + struct netlink_callback *cb, int *s_addr_idx,
> + unsigned int flags, int netnsid)
> +{
> + struct netdev_hw_addr *ha;
> + int addr_idx = 0;
> + int err = 0;
> +
> + netif_addr_lock_bh(dev);
> + netdev_for_each_mc_addr(ha, dev) {
> + if (addr_idx < *s_addr_idx) {
> + addr_idx++;
> + continue;
> + }
> + err = dev_mc_fill_addr(skb, dev, ha, NETLINK_CB(cb->skb).portid,
> + cb->nlh->nlmsg_seq, flags, netnsid);
> + if (err < 0)
> + break;
> + addr_idx++;
> + }
> + netif_addr_unlock_bh(dev);
> +
> + *s_addr_idx = err < 0 ? addr_idx : 0;
> +
> + return err;
> +}
[ ... ]
> +int dev_mc_dump(struct sk_buff *skb, struct netlink_callback *cb)
> +{
[ ... ]
> + rcu_read_lock();
> +
> + if (filter.ifindex) {
> + cb->answer_flags |= NLM_F_DUMP_FILTERED;
> + flags |= NLM_F_DUMP_FILTERED;
> + dev = dev_get_by_index_rcu(filter.tgt_net, filter.ifindex);
> + if (dev)
> + err = dev_mc_dump_dev(dev, skb, cb, &ctx->addr_idx,
> + flags, filter.netnsid);
> + else
> + err = -ENODEV;
> + } else {
> + s_ifindex = ctx->ifindex;
> + for_each_netdev_dump(filter.tgt_net, dev, ctx->ifindex) {
> + /* The device the dump stopped at is gone, do not
> + * skip entries of the next one.
> + */
> + if (dev->ifindex != s_ifindex)
> + ctx->addr_idx = 0;
> + err = dev_mc_dump_dev(dev, skb, cb, &ctx->addr_idx,
> + flags, filter.netnsid);
> + if (err < 0)
> + break;
> + }
> + }
> + cb->seq = dev_mc_base_seq(filter.tgt_net);
[Severity: Medium]
Is the generation read too late for the first round?
cb->seq is read here, after dev_mc_dump_dev() has already dropped
addr_list_lock with netif_addr_unlock_bh(). On the first round
cb->prev_seq is 0, so the value read here becomes the baseline.
Suppose another CPU changes dev->mc on the device where round 1 stopped,
between that unlock and the atomic_read() in dev_mc_base_seq(). That
change then becomes part of the baseline:
round 1 stops at index 5 of dev A
netif_addr_unlock_bh(A)
other CPU deletes entry 2 of A
__hw_addr_changed() bumps dev_mc_genid
round 1 reads G1, cb->prev_seq = G1
round 2 resumes at index 5, old entry 5 (now index 4) is skipped
round 2 reads G1, no NLM_F_DUMP_INTR
An addition in the same window would produce a duplicate entry instead.
The window covers local_bh_enable() in the unlock, which can run
softirqs. Under PREEMPT_RCU the task can also be preempted there.
This ordering matches rtnl_dump_ifinfo(), but that dump resumes by
ifindex, not by list position. inet_dump_addr() in net/ipv4/devinet.c,
which resumes by position, reads the generation before the walk:
cb->seq = inet_base_seq(tgt_net);
for_each_netdev_dump(tgt_net, dev, ctx->ifindex) {
Should cb->seq be read before the walk here too?
> + nl_dump_check_consistent(cb, nlmsg_hdr(skb));
[Severity: Medium]
Can NLM_F_DUMP_INTR be lost when a round emits no messages?
This check runs on every round, even when nothing was added to the skb.
Dumps resume by list position through ctx->addr_idx. Consider this
sequence:
round N-1 returns -EMSGSIZE at the last dev->mc entry of device D,
where D is the last device with entries or the one selected by
ifa_index
before the next recvmsg(), an earlier entry of D is deleted, for
example by a socket leaving an IP multicast group, which reaches
dev_mc_del() and bumps dev_mc_genid
round N resumes at the saved addr_idx, which is now past the end of
D's list, so it emits nothing and returns 0
In round N, cb->seq differs from cb->prev_seq, so
nl_dump_check_consistent() sets NLM_F_DUMP_INTR in nlmsg_hdr(skb). With
skb->len == 0 that address is unused tailroom, not a message. The call
then sets cb->prev_seq to cb->seq.
Next, netlink_dump() calls netlink_dump_done():
net/netlink/af_netlink.c:netlink_dump_done() {
...
nlh = nlmsg_put_answer(skb, cb, NLMSG_DONE, sizeof(nlk->dump_done_errno),
NLM_F_MULTI | cb->answer_flags);
...
nl_dump_check_consistent(cb, nlh);
...
}
NLMSG_DONE is written at the same address, and __nlmsg_put() overwrites
nlmsg_flags. The second check then finds seq == prev_seq and sets
nothing.
Wouldn't the entry that moved into the deleted slot be missing from the
dump, with no NLM_F_DUMP_INTR reported? That seems to contradict the
commit message:
"so an entry added or removed since the previous round sets
NLM_F_DUMP_INTR"
in_dev_dump_ifaddr() in net/ipv4/devinet.c only calls the check after it
has filled a message:
err = inet_fill_ifaddr(skb, ifa, fillargs);
if (err < 0)
goto done;
nl_dump_check_consistent(cb, nlmsg_hdr(skb));
Would the same approach work here?
> +
> + rcu_read_unlock();
> + if (filter.netnsid >= 0)
> + put_net_track(filter.tgt_net, &filter.ns_tracker);
> + return err;
> +}
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261007115821.11332-1-sigefriedhyy%40gmail.com
^ permalink raw reply [flat|nested] 12+ messages in thread
* [PATCH net-next v10 5/6] netlink: specs: rt-addr: document AF_PACKET multicast dumps
2026-10-07 11:58 [PATCH net-next v10 0/6] rtnetlink: dump link-layer multicast addresses Yuyang Huang
` (3 preceding siblings ...)
2026-10-07 11:58 ` [PATCH net-next v10 4/6] net: add AF_PACKET multicast dumps Yuyang Huang
@ 2026-10-07 11:58 ` Yuyang Huang
2026-10-07 11:58 ` [PATCH net-next v10 6/6] selftests: net: test " Yuyang Huang
5 siblings, 0 replies; 12+ messages in thread
From: Yuyang Huang @ 2026-10-07 11:58 UTC (permalink / raw)
To: Yuyang Huang
Cc: Ajay Singh, Aleksandr Loktionov, Andrew Lunn, Claudiu Beznea,
David S. Miller, David Ahern, Donald Hunter, Eric Dumazet,
Ido Schimmel, Jacob Keller, Jakub Kicinski, Johannes Berg,
Kees Cook, Kory Maincent, Kuniyuki Iwashima, Nicolas Dichtel,
Nikolaos Gkarlis, Paolo Abeni, Sabrina Dubroca, Shuah Khan,
Simon Horman, Stanislav Fomichev, Vadim Fedorenko,
Willem de Bruijn, linux-kernel, linux-kselftest, linux-wireless,
netdev
Add the global flag, list the attributes the AF_PACKET dump uses and
describe how ifa-family selects IPv4, IPv6 or link-layer output for
RTM_GETMULTICAST.
Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
Reviewed-by: Nicolas Dichtel <nicolas.dichtel@6wind.com>
---
Documentation/netlink/specs/rt-addr.yaml | 19 +++++++++++++++++--
1 file changed, 17 insertions(+), 2 deletions(-)
diff --git a/Documentation/netlink/specs/rt-addr.yaml b/Documentation/netlink/specs/rt-addr.yaml
index 17ead22034515..cccfccf49cc11 100644
--- a/Documentation/netlink/specs/rt-addr.yaml
+++ b/Documentation/netlink/specs/rt-addr.yaml
@@ -77,6 +77,8 @@ definitions:
name: mcautojoin
-
name: stable-privacy
+ -
+ name: global
attribute-sets:
-
@@ -168,7 +170,17 @@ operations:
attributes: *ifaddr-all
-
name: getmulticast
- doc: Get / dump IPv4/IPv6 multicast addresses.
+ doc: |
+ Get / dump multicast addresses. ifa-family must select the address
+ family: AF_INET or AF_INET6 for the IP multicast groups joined on
+ a device, AF_PACKET for the link-layer multicast addresses in the
+ device filter. Link-layer entries added explicitly, e.g. with
+ SIOCADDMULTI or "bridge fdb add ... self", rather than by a
+ protocol join are reported with the global flag set. For AF_PACKET
+ a non-zero ifa-index restricts the dump to that device and
+ ifa-prefixlen, ifa-flags and ifa-scope must be zero, AF_INET and
+ AF_INET6 apply this and target-netnsid with NETLINK_GET_STRICT_CHK
+ only.
attribute-set: addr-attrs
fixed-header: ifaddrmsg
do:
@@ -181,10 +193,13 @@ operations:
- multicast
- mc-users
- cacheinfo
+ - flags
+ - target-netnsid
dump:
request:
value: 58
- attributes: []
+ attributes:
+ - target-netnsid
reply:
value: 58
attributes: *mcaddr-attrs
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread* [PATCH net-next v10 6/6] selftests: net: test AF_PACKET multicast dumps
2026-10-07 11:58 [PATCH net-next v10 0/6] rtnetlink: dump link-layer multicast addresses Yuyang Huang
` (4 preceding siblings ...)
2026-10-07 11:58 ` [PATCH net-next v10 5/6] netlink: specs: rt-addr: document " Yuyang Huang
@ 2026-10-07 11:58 ` Yuyang Huang
5 siblings, 0 replies; 12+ messages in thread
From: Yuyang Huang @ 2026-10-07 11:58 UTC (permalink / raw)
To: Yuyang Huang
Cc: Ajay Singh, Aleksandr Loktionov, Andrew Lunn, Claudiu Beznea,
David S. Miller, David Ahern, Donald Hunter, Eric Dumazet,
Ido Schimmel, Jacob Keller, Jakub Kicinski, Johannes Berg,
Kees Cook, Kory Maincent, Kuniyuki Iwashima, Nicolas Dichtel,
Nikolaos Gkarlis, Paolo Abeni, Sabrina Dubroca, Shuah Khan,
Simon Horman, Stanislav Fomichev, Vadim Fedorenko,
Willem de Bruijn, linux-kernel, linux-kselftest, linux-wireless,
netdev
Dump the link-layer multicast addresses of a dummy device and verify
that ifa_index restricts the dump to that device, that the all-hosts
address joined on link up is listed without IFA_F_GLOBAL, that an
address added with SIOCADDMULTI is listed with IFA_F_GLOBAL and
IFA_MC_USERS, and that IFA_TARGET_NETNSID dumps another netns.
Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
Reviewed-by: Nicolas Dichtel <nicolas.dichtel@6wind.com>
---
tools/testing/selftests/net/rtnetlink.py | 73 +++++++++++++++++++++++-
1 file changed, 71 insertions(+), 2 deletions(-)
diff --git a/tools/testing/selftests/net/rtnetlink.py b/tools/testing/selftests/net/rtnetlink.py
index dc8c77db48974..5534ade056a5c 100755
--- a/tools/testing/selftests/net/rtnetlink.py
+++ b/tools/testing/selftests/net/rtnetlink.py
@@ -5,13 +5,18 @@ import socket
import struct
import time
from lib.py import bkg, ip, ksft_exit, ksft_run, ksft_eq, ksft_ge, ksft_true, KsftSkipEx
-from lib.py import ksft_not_in, ksft_not_none
+from lib.py import ksft_in, ksft_not_in, ksft_not_none
from lib.py import CmdExitFailure, NetNS, NetNSEnter, RtnlAddrFamily, RtnlRouteFamily
from lib.py import defer
IPV4_ALL_HOSTS_MULTICAST = b'\xe0\x00\x00\x01'
IPV4_TEST_MULTICAST = b'\xef\x01\x01\x01'
IPV6_TEST_MULTICAST = bytes.fromhex('ff020000000000000000000000000123')
+ETH_ALL_HOSTS_MULTICAST = bytes.fromhex('01005e000001')
+ETH_TEST_MULTICAST_STR = '01:00:5e:01:01:01'
+ETH_TEST_MULTICAST = bytes.fromhex(ETH_TEST_MULTICAST_STR.replace(':', ''))
+ETH_PEER_MULTICAST_STR = '01:00:5e:02:02:02'
+ETH_PEER_MULTICAST = bytes.fromhex(ETH_PEER_MULTICAST_STR.replace(':', ''))
def _users_for(rtnl: RtnlAddrFamily, family: int, grp: bytes, ifindex: int):
@@ -105,6 +110,69 @@ def dump_mcaddr6_check() -> None:
s2.close()
+def dump_mcaddr_l2_check() -> None:
+ """
+ Verify link-layer multicast addresses in an AF_PACKET RTM_GETMULTICAST
+ dump: the ifa-index filter, mc-users, the global flag and
+ target-netnsid.
+ """
+
+ with NetNS() as ns, NetNSEnter(str(ns)):
+ for ifname in ("dummy1", "dummy2"):
+ ip(f"link add name {ifname} type dummy")
+ ip(f"link set {ifname} up")
+ dev_idx = socket.if_nametoindex("dummy1")
+ ip(f"maddr add {ETH_TEST_MULTICAST_STR} dev dummy1")
+
+ rtnl = RtnlAddrFamily()
+ defer(rtnl.close)
+ addresses = rtnl.getmulticast(
+ {"ifa-family": socket.AF_PACKET, "ifa-index": dev_idx},
+ dump=True)
+
+ # dummy2 has entries as well, only dummy1 may be listed
+ ksft_eq({addr['ifa-index'] for addr in addresses}, {dev_idx},
+ "AF_PACKET multicast dump ignored ifa-index filter")
+
+ entries = {addr['multicast']: addr for addr in addresses}
+
+ # Bringing an Ethernet device up joins 224.0.0.1, which maps
+ # to 01:00:5e:00:00:01 in the device multicast list.
+ all_hosts = entries.get(ETH_ALL_HOSTS_MULTICAST)
+ ksft_not_none(all_hosts,
+ "dummy1 does not have the all-hosts link-layer address")
+ if all_hosts is not None:
+ ksft_not_in('global', all_hosts['flags'],
+ "protocol entry is global")
+
+ static = entries.get(ETH_TEST_MULTICAST)
+ ksft_not_none(static, "dummy1 does not have the SIOCADDMULTI address")
+ if static is not None:
+ ksft_eq(static['mc-users'], 1,
+ "unexpected mc-users for the SIOCADDMULTI address")
+ ksft_in('global', static['flags'],
+ "SIOCADDMULTI entry is not global")
+
+ # target-netnsid dumps another netns, ifa-index is relative to it
+ with NetNS() as peer:
+ ip(f"netns set {peer} 5")
+ ip("link add name dummy3 type dummy", ns=peer)
+ ip("link set dummy3 up", ns=peer)
+ ip(f"maddr add {ETH_PEER_MULTICAST_STR} dev dummy3", ns=peer)
+ peer_idx = ip("link show dummy3", json=True, ns=peer)[0]['ifindex']
+
+ addresses = rtnl.getmulticast(
+ {"ifa-family": socket.AF_PACKET, "target-netnsid": 5,
+ "ifa-index": peer_idx}, dump=True)
+ ksft_eq({(addr['ifa-index'], addr['target-netnsid'])
+ for addr in addresses}, {(peer_idx, 5)},
+ "target-netnsid did not dump the peer netns")
+ # dummy1 in this netns can have the same ifindex as dummy3
+ ksft_in(ETH_PEER_MULTICAST,
+ {addr['multicast'] for addr in addresses},
+ "target-netnsid did not dump the peer device")
+
+
def ipv4_devconf_notify() -> None:
"""
Configure an interface and set ipv4-devconf values through netlink
@@ -424,7 +492,8 @@ def ipv6_verify_inter_scope_addr_order() -> None:
def main() -> None:
- ksft_run([dump_mcaddr_check, dump_mcaddr6_check, ipv4_devconf_notify,
+ ksft_run([dump_mcaddr_check, dump_mcaddr6_check, dump_mcaddr_l2_check,
+ ipv4_devconf_notify,
ipv6_route_del_reason_expired,
ipv6_route_del_reason_ra_withdrawn,
ipv6_route_del_reason_absent,
--
2.43.0
^ permalink raw reply related [flat|nested] 12+ messages in thread