From: Yuyang Huang <sigefriedhyy@gmail.com>
To: Yuyang Huang <sigefriedhyy@gmail.com>
Cc: Ajay Singh <ajay.kathat@microchip.com>,
Aleksandr Loktionov <aleksandr.loktionov@intel.com>,
Andrew Lunn <andrew+netdev@lunn.ch>,
Claudiu Beznea <claudiu.beznea@tuxon.dev>,
"David S. Miller" <davem@davemloft.net>,
David Ahern <dsahern@kernel.org>,
Donald Hunter <donald.hunter@gmail.com>,
Eric Dumazet <edumazet@google.com>,
Ido Schimmel <idosch@nvidia.com>,
Jacob Keller <jacob.e.keller@intel.com>,
Jakub Kicinski <kuba@kernel.org>,
Johannes Berg <johannes@sipsolutions.net>,
Kees Cook <kees@kernel.org>,
Kory Maincent <kory.maincent@bootlin.com>,
Kuniyuki Iwashima <kuniyu@google.com>,
Nicolas Dichtel <nicolas.dichtel@6wind.com>,
Nikolaos Gkarlis <nickgarlis@gmail.com>,
Paolo Abeni <pabeni@redhat.com>,
Sabrina Dubroca <sd@queasysnail.net>,
Shuah Khan <shuah@kernel.org>, Simon Horman <horms@kernel.org>,
Stanislav Fomichev <sdf.kernel@gmail.com>,
Vadim Fedorenko <vadim.fedorenko@linux.dev>,
Willem de Bruijn <willemb@google.com>,
linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org,
linux-wireless@vger.kernel.org, netdev@vger.kernel.org
Subject: [PATCH net-next v9 4/6] net: add AF_PACKET multicast dumps
Date: Wed, 30 Sep 2026 20:28:40 +0900 [thread overview]
Message-ID: <20260930112842.21323-5-sigefriedhyy@gmail.com> (raw)
In-Reply-To: <20260930112842.21323-1-sigefriedhyy@gmail.com>
RTM_GETMULTICAST dumps IPv4 and IPv6 multicast group memberships, but
the device multicast list (dev->mc) is only available through
/proc/net/dev_mcast, so "ip maddr show" still has to parse procfs for
its link-layer entries.
Handle RTM_GETMULTICAST dumps with ifa_family set to AF_PACKET next to
the dev->mc helpers in dev_addr_lists.c and report every entry of
dev->mc in the existing ifaddrmsg format:
- IFA_MULTICAST carries the raw link-layer address
- IFA_MC_USERS carries the entry reference count
- IFA_F_GLOBAL in IFA_FLAGS reports netdev_hw_addr::global_use, set
by dev_mc_add_global() (SIOCADDMULTI) and dev_mc_add_excl()
("bridge fdb add ... self"), i.e. entries added explicitly rather
than by a protocol join. This is the static column of
/proc/net/dev_mcast
- ifa_scope is RT_SCOPE_LINK
This covers every column of /proc/net/dev_mcast. AF_PACKET is the
family iproute2 already uses for link-layer addresses ("ip -0").
The default FDB dump also walks dev->mc, but only for Ethernet devices
without an ndo_fdb_dump of their own, so bridge, vxlan or macvlan
devices never show their multicast filter there, and it has no users
count or global_use bit. Extending it would change "bridge fdb show"
output and add NDA_* attributes.
Requests are always validated, there are no legacy users: prefixlen,
flags and scope must be zero, ifa_index selects one device and
IFA_TARGET_NETNSID is the only attribute accepted. The dump runs under
RCU and netif_addr_lock_bh() without RTNL. cb->seq is sampled from the
dev->mc generation counter and dev_base_seq under the address lock of
each device and again when a round ends, so a change since the previous
device or dump round sets NLM_F_DUMP_INTR, also when the last round
dumps nothing because the device it stopped at is gone.
Signed-off-by: Yuyang Huang <sigefriedhyy@gmail.com>
Reviewed-by: Nicolas Dichtel <nicolas.dichtel@6wind.com>
---
include/linux/netdevice.h | 1 +
include/uapi/linux/if_addr.h | 1 +
net/core/dev_addr_lists.c | 195 +++++++++++++++++++++++++++++++++++
net/core/rtnetlink.c | 2 +
4 files changed, 199 insertions(+)
diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index d409c56a02459..6eabaa0b48e0c 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -5168,6 +5168,7 @@ int dev_mc_sync_multiple(struct net_device *to, struct net_device *from);
void dev_mc_unsync(struct net_device *to, struct net_device *from);
void dev_mc_flush(struct net_device *dev);
void dev_mc_init(struct net_device *dev);
+int dev_mc_dump(struct sk_buff *skb, struct netlink_callback *cb);
/**
* __dev_mc_sync - Synchronize device's multicast list
diff --git a/include/uapi/linux/if_addr.h b/include/uapi/linux/if_addr.h
index 7fb630b7fe311..0a1ad9ebb47be 100644
--- a/include/uapi/linux/if_addr.h
+++ b/include/uapi/linux/if_addr.h
@@ -57,6 +57,7 @@ enum {
#define IFA_F_NOPREFIXROUTE 0x200
#define IFA_F_MCAUTOJOIN 0x400
#define IFA_F_STABLE_PRIVACY 0x800
+#define IFA_F_GLOBAL 0x1000
struct ifa_cacheinfo {
__u32 ifa_prefered;
diff --git a/net/core/dev_addr_lists.c b/net/core/dev_addr_lists.c
index 783c62895249b..948dc25635938 100644
--- a/net/core/dev_addr_lists.c
+++ b/net/core/dev_addr_lists.c
@@ -10,8 +10,12 @@
#include <linux/netdevice.h>
#include <linux/rtnetlink.h>
#include <linux/export.h>
+#include <linux/if_addr.h>
#include <linux/list.h>
#include <linux/spinlock.h>
+#include <net/netlink.h>
+#include <net/rtnetlink.h>
+#include <net/sock.h>
#include <kunit/visibility.h>
#include "dev.h"
@@ -1219,6 +1223,197 @@ void dev_mc_init(struct net_device *dev)
}
EXPORT_SYMBOL(dev_mc_init);
+static int dev_mc_fill_addr(struct sk_buff *skb, const struct net_device *dev,
+ const struct netdev_hw_addr *ha, u32 portid,
+ u32 seq, unsigned int flags, int netnsid)
+{
+ u32 ifa_flags = ha->global_use ? IFA_F_GLOBAL : 0;
+ struct ifaddrmsg *ifm;
+ struct nlmsghdr *nlh;
+
+ nlh = nlmsg_put(skb, portid, seq, RTM_GETMULTICAST, sizeof(*ifm),
+ flags);
+ if (!nlh)
+ return -EMSGSIZE;
+
+ ifm = nlmsg_data(nlh);
+ ifm->ifa_family = AF_PACKET;
+ ifm->ifa_prefixlen = 0;
+ /* ifm->ifa_flags holds 8 bits, the full value is in IFA_FLAGS */
+ ifm->ifa_flags = (__u8)ifa_flags;
+ ifm->ifa_scope = RT_SCOPE_LINK;
+ ifm->ifa_index = dev->ifindex;
+
+ if ((netnsid >= 0 &&
+ nla_put_s32(skb, IFA_TARGET_NETNSID, netnsid)) ||
+ nla_put(skb, IFA_MULTICAST, dev->addr_len, ha->addr) ||
+ nla_put_u32(skb, IFA_MC_USERS, ha->refcount) ||
+ nla_put_u32(skb, IFA_FLAGS, ifa_flags)) {
+ nlmsg_cancel(skb, nlh);
+ return -EMSGSIZE;
+ }
+
+ nlmsg_end(skb, nlh);
+ return 0;
+}
+
+/* Combine dev_mc_genid and dev_base_seq to detect changes, like
+ * inet_base_seq().
+ */
+static u32 dev_mc_base_seq(const struct net *net)
+{
+ u32 res = atomic_read(&net->dev_mc_genid) +
+ READ_ONCE(net->dev_base_seq);
+
+ /* Must not return 0 (see nl_dump_check_consistent()). */
+ if (!res)
+ res = 0x80000000;
+ return res;
+}
+
+static int dev_mc_dump_dev(struct net_device *dev, struct sk_buff *skb,
+ struct netlink_callback *cb, int *s_addr_idx,
+ unsigned int flags, int netnsid)
+{
+ struct netdev_hw_addr *ha;
+ int addr_idx = 0;
+ int err = 0;
+
+ netif_addr_lock_bh(dev);
+ /* Sampled under the lock, see nl_dump_check_consistent() */
+ cb->seq = dev_mc_base_seq(dev_net(dev));
+ netdev_for_each_mc_addr(ha, dev) {
+ if (addr_idx < *s_addr_idx) {
+ addr_idx++;
+ continue;
+ }
+ err = dev_mc_fill_addr(skb, dev, ha, NETLINK_CB(cb->skb).portid,
+ cb->nlh->nlmsg_seq, flags, netnsid);
+ if (err < 0)
+ break;
+ nl_dump_check_consistent(cb, nlmsg_hdr(skb));
+ addr_idx++;
+ }
+ netif_addr_unlock_bh(dev);
+
+ *s_addr_idx = err < 0 ? addr_idx : 0;
+
+ return err;
+}
+
+struct dev_mc_dump_filter {
+ struct net *tgt_net;
+ netns_tracker ns_tracker;
+ int netnsid;
+ int ifindex;
+};
+
+static const struct nla_policy dev_mc_dump_policy[IFA_MAX + 1] = {
+ [IFA_TARGET_NETNSID] = { .type = NLA_S32 },
+};
+
+static int dev_mc_valid_dump_req(const struct nlmsghdr *nlh, struct sock *sk,
+ struct dev_mc_dump_filter *filter,
+ struct netlink_ext_ack *extack)
+{
+ struct nlattr *tb[IFA_MAX + 1];
+ struct ifaddrmsg *ifm;
+ int err;
+
+ ifm = nlmsg_payload(nlh, sizeof(*ifm));
+ if (!ifm) {
+ NL_SET_ERR_MSG(extack,
+ "Invalid header for multicast dump request");
+ return -EINVAL;
+ }
+
+ if (ifm->ifa_prefixlen || ifm->ifa_flags || ifm->ifa_scope) {
+ NL_SET_ERR_MSG(extack,
+ "Invalid values in multicast dump header");
+ return -EINVAL;
+ }
+
+ err = nlmsg_parse(nlh, sizeof(*ifm), tb, IFA_MAX,
+ dev_mc_dump_policy, extack);
+ if (err < 0)
+ return err;
+
+ if (tb[IFA_TARGET_NETNSID]) {
+ struct net *net;
+
+ filter->netnsid = nla_get_s32(tb[IFA_TARGET_NETNSID]);
+ net = rtnl_get_net_ns_capable(sk, filter->netnsid);
+ if (IS_ERR(net)) {
+ NL_SET_ERR_MSG(extack,
+ "Invalid target network namespace id");
+ return PTR_ERR(net);
+ }
+ netns_tracker_alloc(net, &filter->ns_tracker, GFP_KERNEL);
+ filter->tgt_net = net;
+ }
+
+ filter->ifindex = ifm->ifa_index;
+
+ return 0;
+}
+
+int dev_mc_dump(struct sk_buff *skb, struct netlink_callback *cb)
+{
+ struct dev_mc_dump_filter filter = {
+ .tgt_net = sock_net(skb->sk),
+ .netnsid = -1,
+ };
+ unsigned int flags = NLM_F_MULTI;
+ struct {
+ unsigned long ifindex;
+ int addr_idx;
+ } *ctx = (void *)cb->ctx;
+ unsigned long s_ifindex;
+ struct net_device *dev;
+ int err;
+
+ err = dev_mc_valid_dump_req(cb->nlh, skb->sk, &filter, cb->extack);
+ if (err < 0)
+ return err;
+
+ rcu_read_lock();
+
+ if (filter.ifindex) {
+ cb->answer_flags |= NLM_F_DUMP_FILTERED;
+ flags |= NLM_F_DUMP_FILTERED;
+ dev = dev_get_by_index_rcu(filter.tgt_net, filter.ifindex);
+ if (!dev) {
+ err = -ENODEV;
+ goto out;
+ }
+ err = dev_mc_dump_dev(dev, skb, cb, &ctx->addr_idx, flags,
+ filter.netnsid);
+ goto out;
+ }
+
+ s_ifindex = ctx->ifindex;
+ for_each_netdev_dump(filter.tgt_net, dev, ctx->ifindex) {
+ /* The device the dump stopped at is gone, do not skip
+ * entries of the next one.
+ */
+ if (dev->ifindex != s_ifindex)
+ ctx->addr_idx = 0;
+ err = dev_mc_dump_dev(dev, skb, cb, &ctx->addr_idx, flags,
+ filter.netnsid);
+ if (err < 0)
+ break;
+ }
+out:
+ /* A round that dumps no device, e.g. the one it stopped at is gone,
+ * still needs the NLMSG_DONE check to see the change.
+ */
+ cb->seq = dev_mc_base_seq(filter.tgt_net);
+ rcu_read_unlock();
+ if (filter.netnsid >= 0)
+ put_net_track(filter.tgt_net, &filter.ns_tracker);
+ return err;
+}
+
static int netif_addr_lists_snapshot(struct net_device *dev,
struct netdev_hw_addr_list *uc_snap,
struct netdev_hw_addr_list *mc_snap,
diff --git a/net/core/rtnetlink.c b/net/core/rtnetlink.c
index e3444fd240615..204dc9040e3cc 100644
--- a/net/core/rtnetlink.c
+++ b/net/core/rtnetlink.c
@@ -7278,6 +7278,8 @@ static const struct rtnl_msg_handler rtnetlink_rtnl_msg_handlers[] __initconst =
{.msgtype = RTM_SETSTATS, .doit = rtnl_stats_set},
{.msgtype = RTM_NEWLINKPROP, .doit = rtnl_newlinkprop},
{.msgtype = RTM_DELLINKPROP, .doit = rtnl_dellinkprop},
+ {.protocol = PF_PACKET, .msgtype = RTM_GETMULTICAST,
+ .dumpit = dev_mc_dump, .flags = RTNL_FLAG_DUMP_UNLOCKED},
{.protocol = PF_BRIDGE, .msgtype = RTM_GETLINK,
.dumpit = rtnl_bridge_getlink},
{.protocol = PF_BRIDGE, .msgtype = RTM_DELLINK,
--
2.43.0
next prev parent reply other threads:[~2026-09-30 11:29 UTC|newest]
Thread overview: 24+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 11:28 [PATCH net-next v9 0/6] rtnetlink: dump link-layer multicast addresses Yuyang Huang
2026-09-30 11:28 ` [PATCH net-next v9 1/6] netlink: specs: rt-addr: fix the type of target-netnsid Yuyang Huang
2026-10-01 23:31 ` netdev-bot+sashiko
2026-10-02 10:05 ` Yuyang Huang
2026-09-30 11:28 ` [PATCH net-next v9 2/6] net: change netdev_hw_addr_list count through helpers Yuyang Huang
2026-09-30 13:03 ` Nicolas Dichtel
2026-09-30 13:43 ` Yuyang Huang
2026-09-30 14:08 ` Nicolas Dichtel
2026-09-30 14:13 ` Yuyang Huang
2026-10-01 23:31 ` netdev-bot+sashiko
2026-10-02 10:06 ` Yuyang Huang
2026-09-30 11:28 ` [PATCH net-next v9 3/6] net: add a generation counter for dev->mc changes Yuyang Huang
2026-09-30 13:04 ` Nicolas Dichtel
2026-10-05 23:49 ` Jakub Kicinski
2026-10-06 0:54 ` Yuyang Huang
2026-09-30 11:28 ` Yuyang Huang [this message]
2026-10-01 23:31 ` [PATCH net-next v9 4/6] net: add AF_PACKET multicast dumps netdev-bot+sashiko
2026-10-02 10:12 ` Yuyang Huang
2026-10-05 23:49 ` Jakub Kicinski
2026-10-06 1:09 ` Yuyang Huang
2026-09-30 11:28 ` [PATCH net-next v9 5/6] netlink: specs: rt-addr: document " Yuyang Huang
2026-10-01 23:31 ` netdev-bot+sashiko
2026-10-02 10:13 ` Yuyang Huang
2026-09-30 11:28 ` [PATCH net-next v9 6/6] selftests: net: test " Yuyang Huang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260930112842.21323-5-sigefriedhyy@gmail.com \
--to=sigefriedhyy@gmail.com \
--cc=ajay.kathat@microchip.com \
--cc=aleksandr.loktionov@intel.com \
--cc=andrew+netdev@lunn.ch \
--cc=claudiu.beznea@tuxon.dev \
--cc=davem@davemloft.net \
--cc=donald.hunter@gmail.com \
--cc=dsahern@kernel.org \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=idosch@nvidia.com \
--cc=jacob.e.keller@intel.com \
--cc=johannes@sipsolutions.net \
--cc=kees@kernel.org \
--cc=kory.maincent@bootlin.com \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-kselftest@vger.kernel.org \
--cc=linux-wireless@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=nickgarlis@gmail.com \
--cc=nicolas.dichtel@6wind.com \
--cc=pabeni@redhat.com \
--cc=sd@queasysnail.net \
--cc=sdf.kernel@gmail.com \
--cc=shuah@kernel.org \
--cc=vadim.fedorenko@linux.dev \
--cc=willemb@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.