* [PATCH 0/4] Remove limitations coming from legacy VMDq
@ 2026-04-03 9:18 David Marchand
2026-04-03 9:18 ` [PATCH 1/4] ethdev: skip VMDq pools unless configured David Marchand
` (15 more replies)
0 siblings, 16 replies; 146+ messages in thread
From: David Marchand @ 2026-04-03 9:18 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain
Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
VMDq has been imposing a maximum number of mac addresses in the
mac_addr_add/del API.
Nowadays, new Intel drivers do not support the feature and few other
drivers implement this feature.
This series proposes to flag drivers that support the feature, and
remove the limit of number of mac addresses for others.
Next step could be to remove the VMDq pool notion from the generic API.
However I have some concern about this, as changing the quite stable
mac_addr_add/del API now seems a lot of noise for not much benefit.
--
David Marchand
David Marchand (4):
ethdev: skip VMDq pools unless configured
ethdev: announce VMDq capability
ethdev: hide VMDq internal sizes
net/iavf: accept up to 32k unicast MAC addresses
drivers/net/bnxt/bnxt_ethdev.c | 3 +-
drivers/net/bnxt/bnxt_reps.c | 1 +
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
drivers/net/intel/e1000/em_ethdev.c | 1 +
drivers/net/intel/e1000/igb_ethdev.c | 1 +
drivers/net/intel/fm10k/fm10k_ethdev.c | 1 +
drivers/net/intel/i40e/i40e_ethdev.c | 3 +-
drivers/net/intel/i40e/i40e_vf_representor.c | 1 +
drivers/net/intel/iavf/iavf.h | 5 ++-
drivers/net/intel/iavf/iavf_ethdev.c | 10 ++---
drivers/net/intel/iavf/iavf_vchnl.c | 6 +--
drivers/net/intel/ipn3ke/ipn3ke_representor.c | 3 +-
drivers/net/intel/ixgbe/ixgbe_ethdev.c | 2 +
drivers/net/txgbe/txgbe_ethdev.c | 1 +
drivers/net/txgbe/txgbe_ethdev_vf.c | 1 +
lib/ethdev/ethdev_driver.h | 8 +++-
lib/ethdev/rte_ethdev.c | 45 ++++++++++++++++---
lib/ethdev/rte_ethdev.h | 8 +---
18 files changed, 73 insertions(+), 28 deletions(-)
--
2.53.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH 1/4] ethdev: skip VMDq pools unless configured
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
@ 2026-04-03 9:18 ` David Marchand
2026-06-01 9:30 ` Andrew Rybchenko
2026-04-03 9:18 ` [PATCH 2/4] ethdev: announce VMDq capability David Marchand
` (14 subsequent siblings)
15 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-04-03 9:18 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Nithin Dabilpuram, Kiran Kumar K,
Sunil Kumar Kori, Satha Rao, Harman Kalra, Thomas Monjalon,
Andrew Rybchenko
The mac_addr_add API describes that only the 0 pool should be passed
unless VMDq has been enabled, though there was no validation so far.
Add such a check, then cleanup the related operations (adding, removing,
restoring).
As a side effect, the net/cnxk does not need to manually reset the
mac_pool_sel[] array.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
lib/ethdev/rte_ethdev.c | 28 +++++++++++++++++++++-------
2 files changed, 21 insertions(+), 8 deletions(-)
diff --git a/drivers/net/cnxk/cnxk_ethdev_ops.c b/drivers/net/cnxk/cnxk_ethdev_ops.c
index 49e77e49a6..75decf7098 100644
--- a/drivers/net/cnxk/cnxk_ethdev_ops.c
+++ b/drivers/net/cnxk/cnxk_ethdev_ops.c
@@ -1240,7 +1240,6 @@ cnxk_nix_mc_addr_list_configure(struct rte_eth_dev *eth_dev, struct rte_ether_ad
/* Update address in NIC data structure */
rte_ether_addr_copy(&mc_addr_set[i], &data->mac_addrs[j]);
rte_ether_addr_copy(&mc_addr_set[i], &dev->dmac_addrs[j]);
- data->mac_pool_sel[j] = RTE_BIT64(0);
}
roc_nix_npc_promisc_ena_dis(nix, true);
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 2edc7a362e..9577b7d848 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1680,7 +1680,10 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
continue;
pool = 0;
- pool_mask = dev->data->mac_pool_sel[i];
+ if ((dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0)
+ pool_mask = dev->data->mac_pool_sel[i];
+ else
+ pool_mask = 1;
do {
if (pool_mask & UINT64_C(1))
@@ -5390,8 +5393,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
uint32_t pool)
{
struct rte_eth_dev *dev;
- int index;
uint64_t pool_mask;
+ bool vmdq;
+ int index;
int ret;
RTE_ETH_VALID_PORTID_OR_ERR_RET(port_id, -ENODEV);
@@ -5416,6 +5420,12 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
RTE_ETHDEV_LOG_LINE(ERR, "Pool ID must be 0-%d", RTE_ETH_64_POOLS - 1);
return -EINVAL;
}
+ vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+ if (!vmdq && pool != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Port %u: VMDq is not configured (pool %d)",
+ port_id, pool);
+ return -EINVAL;
+ }
index = eth_dev_get_mac_addr_index(port_id, addr);
if (index < 0) {
@@ -5425,7 +5435,7 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
port_id);
return -ENOSPC;
}
- } else {
+ } else if (vmdq) {
pool_mask = dev->data->mac_pool_sel[index];
/* Check if both MAC address and pool is already there, and do nothing */
@@ -5440,8 +5450,10 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
/* Update address in NIC data structure */
rte_ether_addr_copy(addr, &dev->data->mac_addrs[index]);
- /* Update pool bitmap in NIC data structure */
- dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ if (vmdq) {
+ /* Update pool bitmap in NIC data structure */
+ dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ }
}
ret = eth_err(port_id, ret);
@@ -5486,8 +5498,10 @@ rte_eth_dev_mac_addr_remove(uint16_t port_id, struct rte_ether_addr *addr)
/* Update address in NIC data structure */
rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
- /* reset pool bitmap */
- dev->data->mac_pool_sel[index] = 0;
+ if ((dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ /* reset pool bitmap */
+ dev->data->mac_pool_sel[index] = 0;
+ }
rte_ethdev_trace_mac_addr_remove(port_id, addr);
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH 2/4] ethdev: announce VMDq capability
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
2026-04-03 9:18 ` [PATCH 1/4] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-04-03 9:18 ` David Marchand
2026-04-06 22:22 ` Kishore Padmanabha
2026-06-01 9:32 ` Andrew Rybchenko
2026-04-03 9:18 ` [PATCH 3/4] ethdev: hide VMDq internal sizes David Marchand
` (13 subsequent siblings)
15 siblings, 2 replies; 146+ messages in thread
From: David Marchand @ 2026-04-03 9:18 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Kishore Padmanabha, Ajit Khaparde,
Bruce Richardson, Rosen Xu, Anatoly Burakov, Vladimir Medvedkin,
Jiawen Wu, Zaiyu Wang, Thomas Monjalon, Andrew Rybchenko
Let's mark VMDq feature availability as a per device capability.
We can then enforce API calls related to this feature are done on device
with such capability.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/bnxt/bnxt_ethdev.c | 3 ++-
drivers/net/bnxt/bnxt_reps.c | 1 +
drivers/net/intel/e1000/em_ethdev.c | 1 +
drivers/net/intel/e1000/igb_ethdev.c | 1 +
drivers/net/intel/fm10k/fm10k_ethdev.c | 1 +
drivers/net/intel/i40e/i40e_ethdev.c | 3 ++-
drivers/net/intel/i40e/i40e_vf_representor.c | 1 +
drivers/net/intel/ipn3ke/ipn3ke_representor.c | 3 ++-
drivers/net/intel/ixgbe/ixgbe_ethdev.c | 2 ++
drivers/net/txgbe/txgbe_ethdev.c | 1 +
drivers/net/txgbe/txgbe_ethdev_vf.c | 1 +
lib/ethdev/rte_ethdev.c | 17 +++++++++++++++++
lib/ethdev/rte_ethdev.h | 2 ++
13 files changed, 34 insertions(+), 3 deletions(-)
diff --git a/drivers/net/bnxt/bnxt_ethdev.c b/drivers/net/bnxt/bnxt_ethdev.c
index b677f9491d..0f783b9e98 100644
--- a/drivers/net/bnxt/bnxt_ethdev.c
+++ b/drivers/net/bnxt/bnxt_ethdev.c
@@ -1214,7 +1214,8 @@ static int bnxt_dev_info_get_op(struct rte_eth_dev *eth_dev,
dev_info->speed_capa = bnxt_get_speed_capabilities(bp);
dev_info->dev_capa = RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
- RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
+ RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
+ RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->default_rxconf = (struct rte_eth_rxconf) {
diff --git a/drivers/net/bnxt/bnxt_reps.c b/drivers/net/bnxt/bnxt_reps.c
index e26a086f41..5e834830e2 100644
--- a/drivers/net/bnxt/bnxt_reps.c
+++ b/drivers/net/bnxt/bnxt_reps.c
@@ -649,6 +649,7 @@ int bnxt_rep_dev_info_get_op(struct rte_eth_dev *eth_dev,
dev_info->max_tx_queues = max_rx_rings;
dev_info->reta_size = bnxt_rss_hash_tbl_size(parent_bp);
dev_info->hash_key_size = 40;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
/* MTU specifics */
diff --git a/drivers/net/intel/e1000/em_ethdev.c b/drivers/net/intel/e1000/em_ethdev.c
index 9e15e882b9..389744ad5e 100644
--- a/drivers/net/intel/e1000/em_ethdev.c
+++ b/drivers/net/intel/e1000/em_ethdev.c
@@ -1175,6 +1175,7 @@ eth_em_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
RTE_ETH_LINK_SPEED_100M_HD | RTE_ETH_LINK_SPEED_100M |
RTE_ETH_LINK_SPEED_1G;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
/* Preferred queue parameters */
diff --git a/drivers/net/intel/e1000/igb_ethdev.c b/drivers/net/intel/e1000/igb_ethdev.c
index ef1599ac38..fe68c18417 100644
--- a/drivers/net/intel/e1000/igb_ethdev.c
+++ b/drivers/net/intel/e1000/igb_ethdev.c
@@ -2324,6 +2324,7 @@ eth_igb_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->tx_queue_offload_capa = igb_get_tx_queue_offloads_capa(dev);
dev_info->tx_offload_capa = igb_get_tx_port_offloads_capa(dev) |
dev_info->tx_queue_offload_capa;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
switch (hw->mac.type) {
diff --git a/drivers/net/intel/fm10k/fm10k_ethdev.c b/drivers/net/intel/fm10k/fm10k_ethdev.c
index 97f61afec2..037d2206fd 100644
--- a/drivers/net/intel/fm10k/fm10k_ethdev.c
+++ b/drivers/net/intel/fm10k/fm10k_ethdev.c
@@ -1444,6 +1444,7 @@ fm10k_dev_infos_get(struct rte_eth_dev *dev,
dev_info->speed_capa = RTE_ETH_LINK_SPEED_1G | RTE_ETH_LINK_SPEED_2_5G |
RTE_ETH_LINK_SPEED_10G | RTE_ETH_LINK_SPEED_25G |
RTE_ETH_LINK_SPEED_40G | RTE_ETH_LINK_SPEED_100G;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
return 0;
}
diff --git a/drivers/net/intel/i40e/i40e_ethdev.c b/drivers/net/intel/i40e/i40e_ethdev.c
index 100a751225..64c29c6e85 100644
--- a/drivers/net/intel/i40e/i40e_ethdev.c
+++ b/drivers/net/intel/i40e/i40e_ethdev.c
@@ -3878,7 +3878,8 @@ i40e_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->dev_capa =
RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
- RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
+ RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
+ RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->hash_key_size = (I40E_PFQF_HKEY_MAX_INDEX + 1) *
diff --git a/drivers/net/intel/i40e/i40e_vf_representor.c b/drivers/net/intel/i40e/i40e_vf_representor.c
index e8f0bb62a0..d31148acb5 100644
--- a/drivers/net/intel/i40e/i40e_vf_representor.c
+++ b/drivers/net/intel/i40e/i40e_vf_representor.c
@@ -33,6 +33,7 @@ i40e_vf_representor_dev_infos_get(struct rte_eth_dev *ethdev,
/* get dev info for the vdev */
dev_info->device = ethdev->device;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->max_rx_queues = ethdev->data->nb_rx_queues;
diff --git a/drivers/net/intel/ipn3ke/ipn3ke_representor.c b/drivers/net/intel/ipn3ke/ipn3ke_representor.c
index cd34d08055..d581ee3c37 100644
--- a/drivers/net/intel/ipn3ke/ipn3ke_representor.c
+++ b/drivers/net/intel/ipn3ke/ipn3ke_representor.c
@@ -95,7 +95,8 @@ ipn3ke_rpst_dev_infos_get(struct rte_eth_dev *ethdev,
dev_info->dev_capa =
RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
- RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
+ RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
+ RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->switch_info.name = ethdev->device->name;
diff --git a/drivers/net/intel/ixgbe/ixgbe_ethdev.c b/drivers/net/intel/ixgbe/ixgbe_ethdev.c
index 57d929cf2c..5d886b3e28 100644
--- a/drivers/net/intel/ixgbe/ixgbe_ethdev.c
+++ b/drivers/net/intel/ixgbe/ixgbe_ethdev.c
@@ -3997,6 +3997,7 @@ ixgbe_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->max_mtu = dev_info->max_rx_pktlen - IXGBE_ETH_OVERHEAD;
dev_info->min_mtu = RTE_ETHER_MIN_MTU;
dev_info->vmdq_queue_num = dev_info->max_rx_queues;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->rx_queue_offload_capa = ixgbe_get_rx_queue_offloads(dev);
dev_info->rx_offload_capa = (ixgbe_get_rx_port_offloads(dev) |
dev_info->rx_queue_offload_capa);
@@ -4115,6 +4116,7 @@ ixgbevf_dev_info_get(struct rte_eth_dev *dev,
dev_info->max_vmdq_pools = RTE_ETH_16_POOLS;
else
dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->rx_queue_offload_capa = ixgbe_get_rx_queue_offloads(dev);
dev_info->rx_offload_capa = (ixgbe_get_rx_port_offloads(dev) |
dev_info->rx_queue_offload_capa);
diff --git a/drivers/net/txgbe/txgbe_ethdev.c b/drivers/net/txgbe/txgbe_ethdev.c
index 5d360f8305..bd818e8269 100644
--- a/drivers/net/txgbe/txgbe_ethdev.c
+++ b/drivers/net/txgbe/txgbe_ethdev.c
@@ -2836,6 +2836,7 @@ txgbe_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->max_vfs = pci_dev->max_vfs;
dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
dev_info->vmdq_queue_num = dev_info->max_rx_queues;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->rx_queue_offload_capa = txgbe_get_rx_queue_offloads(dev);
dev_info->rx_offload_capa = (txgbe_get_rx_port_offloads(dev) |
diff --git a/drivers/net/txgbe/txgbe_ethdev_vf.c b/drivers/net/txgbe/txgbe_ethdev_vf.c
index 39a5fff65c..934763574c 100644
--- a/drivers/net/txgbe/txgbe_ethdev_vf.c
+++ b/drivers/net/txgbe/txgbe_ethdev_vf.c
@@ -572,6 +572,7 @@ txgbevf_dev_info_get(struct rte_eth_dev *dev,
dev_info->max_hash_mac_addrs = TXGBE_VMDQ_NUM_UC_MAC;
dev_info->max_vfs = pci_dev->max_vfs;
dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->rx_queue_offload_capa = txgbe_get_rx_queue_offloads(dev);
dev_info->rx_offload_capa = (txgbe_get_rx_port_offloads(dev) |
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 9577b7d848..7ba539e796 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -158,6 +158,7 @@ static const struct {
{RTE_ETH_DEV_CAPA_RXQ_SHARE, "RXQ_SHARE"},
{RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP, "FLOW_RULE_KEEP"},
{RTE_ETH_DEV_CAPA_FLOW_SHARED_OBJECT_KEEP, "FLOW_SHARED_OBJECT_KEEP"},
+ {RTE_ETH_DEV_CAPA_VMDQ, "VMDQ"},
};
enum {
@@ -1581,6 +1582,22 @@ rte_eth_dev_configure(uint16_t port_id, uint16_t nb_rx_q, uint16_t nb_tx_q,
goto rollback;
}
+ if (!(dev_info.dev_capa & RTE_ETH_DEV_CAPA_VMDQ)) {
+ if ((dev_conf->rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq rx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ if (dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_DCB ||
+ dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_ONLY) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq tx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ }
+
/*
* Setup new number of Rx/Tx queues and reconfigure device.
*/
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index 0d8e2d0236..62c72de0e5 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -1696,6 +1696,8 @@ struct rte_eth_conf {
#define RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP RTE_BIT64(3)
/** Device supports keeping shared flow objects across restart. */
#define RTE_ETH_DEV_CAPA_FLOW_SHARED_OBJECT_KEEP RTE_BIT64(4)
+/** Device supports VMDq. */
+#define RTE_ETH_DEV_CAPA_VMDQ RTE_BIT64(5)
/**@}*/
/*
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH 3/4] ethdev: hide VMDq internal sizes
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
2026-04-03 9:18 ` [PATCH 1/4] ethdev: skip VMDq pools unless configured David Marchand
2026-04-03 9:18 ` [PATCH 2/4] ethdev: announce VMDq capability David Marchand
@ 2026-04-03 9:18 ` David Marchand
2026-06-01 9:34 ` Andrew Rybchenko
2026-04-03 9:18 ` [PATCH 4/4] net/iavf: accept up to 32k unicast MAC addresses David Marchand
` (12 subsequent siblings)
15 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-04-03 9:18 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Thomas Monjalon, Andrew Rybchenko
Hide RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
in the driver API as those (ambiguous) macros are only a driver concern.
In practice, this is only used by the bnxt and ixgbe (+ clones) drivers.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
lib/ethdev/ethdev_driver.h | 8 +++++++-
lib/ethdev/rte_ethdev.h | 6 ------
2 files changed, 7 insertions(+), 7 deletions(-)
diff --git a/lib/ethdev/ethdev_driver.h b/lib/ethdev/ethdev_driver.h
index 1255cd6f2c..a4e9cf5b90 100644
--- a/lib/ethdev/ethdev_driver.h
+++ b/lib/ethdev/ethdev_driver.h
@@ -119,6 +119,12 @@ struct __rte_cache_aligned rte_eth_dev {
struct rte_eth_dev_sriov;
struct rte_eth_dev_owner;
+/* Definitions used for receive MAC address */
+#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
+
+/* Definitions used for unicast hash */
+#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
+
/**
* @internal
* The data part, with no function pointers, associated with each Ethernet
@@ -153,7 +159,7 @@ struct __rte_cache_aligned rte_eth_dev_data {
* The first entry (index zero) is the default address.
*/
struct rte_ether_addr *mac_addrs;
- /** Bitmap associating MAC addresses to pools */
+ /** Bitmap associating MAC addresses to VMDq pools */
uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
/**
* Device Ethernet MAC addresses of hash filtering.
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index 62c72de0e5..6c1984e679 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -903,12 +903,6 @@ rte_eth_rss_hf_refine(uint64_t rss_hf)
#define RTE_ETH_VLAN_ID_MAX 0x0FFF /**< VLAN ID is in lower 12 bits*/
/**@}*/
-/* Definitions used for receive MAC address */
-#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
-
-/* Definitions used for unicast hash */
-#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
-
/**@{@name VMDq Rx mode
* @see rte_eth_vmdq_rx_conf.rx_mode
*/
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH 4/4] net/iavf: accept up to 32k unicast MAC addresses
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (2 preceding siblings ...)
2026-04-03 9:18 ` [PATCH 3/4] ethdev: hide VMDq internal sizes David Marchand
@ 2026-04-03 9:18 ` David Marchand
2026-04-05 18:47 ` [PATCH 0/4] Remove limitations coming from legacy VMDq Stephen Hemminger
` (11 subsequent siblings)
15 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-04-03 9:18 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Vladimir Medvedkin
E810 hardware provides 32k switch lookups.
Thanks to this, it is possible to allow a lot more secondary mac
addresses than what is possible today.
In practice, the maximum number of macs available per port may be lower
and depends on usage by other VFs on the same PF.
There is no way to figure out this limit but to try adding a mac address
and get an error from the PF driver.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/intel/iavf/iavf.h | 5 +++--
drivers/net/intel/iavf/iavf_ethdev.c | 10 +++++-----
drivers/net/intel/iavf/iavf_vchnl.c | 6 +++---
3 files changed, 11 insertions(+), 10 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 403c61e2e8..f1dede0694 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -31,7 +31,8 @@
#define IAVF_IRQ_MAP_NUM_PER_BUF 128
#define IAVF_RXTX_QUEUE_CHUNKS_NUM 2
-#define IAVF_NUM_MACADDR_MAX 64
+#define IAVF_UC_MACADDR_MAX 32768
+#define IAVF_MC_MACADDR_MAX 64
#define IAVF_DEV_WATCHDOG_PERIOD 2000 /* microseconds, set 0 to disable*/
@@ -253,7 +254,7 @@ struct iavf_info {
uint32_t link_speed;
/* Multicast addrs */
- struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
+ struct rte_ether_addr mc_addrs[IAVF_MC_MACADDR_MAX];
uint16_t mc_addrs_num; /* Multicast mac addresses number */
struct iavf_vsi vsi;
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index 1eca20bc9a..c69a012d50 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -379,10 +379,10 @@ iavf_set_mc_addr_list(struct rte_eth_dev *dev,
IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
int err, ret;
- if (mc_addrs_num > IAVF_NUM_MACADDR_MAX) {
+ if (mc_addrs_num > IAVF_MC_MACADDR_MAX) {
PMD_DRV_LOG(ERR,
"can't add more than a limited number (%u) of addresses.",
- (uint32_t)IAVF_NUM_MACADDR_MAX);
+ (uint32_t)IAVF_MC_MACADDR_MAX);
return -EINVAL;
}
@@ -1120,7 +1120,7 @@ iavf_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->hash_key_size = vf->vf_res->rss_key_size;
dev_info->reta_size = vf->vf_res->rss_lut_size;
dev_info->flow_type_rss_offloads = IAVF_RSS_OFFLOAD_ALL;
- dev_info->max_mac_addrs = IAVF_NUM_MACADDR_MAX;
+ dev_info->max_mac_addrs = IAVF_UC_MACADDR_MAX;
dev_info->dev_capa =
RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
@@ -2822,11 +2822,11 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
/* copy mac addr */
eth_dev->data->mac_addrs = rte_zmalloc(
- "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX, 0);
+ "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_UC_MACADDR_MAX, 0);
if (!eth_dev->data->mac_addrs) {
PMD_INIT_LOG(ERR, "Failed to allocate %d bytes needed to"
" store MAC addresses",
- RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX);
+ RTE_ETHER_ADDR_LEN * IAVF_UC_MACADDR_MAX);
ret = -ENOMEM;
goto init_vf_err;
}
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index 08dd6f2d7f..c70f2fbbc0 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1442,7 +1442,7 @@ iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
{
struct {
struct virtchnl_ether_addr_list list;
- struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
+ struct virtchnl_ether_addr addr[IAVF_UC_MACADDR_MAX];
} list_req = {0};
struct virtchnl_ether_addr_list *list = &list_req.list;
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
@@ -1450,7 +1450,7 @@ iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
int err, i;
size_t buf_len;
- for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
+ for (i = 0; i < IAVF_UC_MACADDR_MAX; i++) {
struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
@@ -2060,7 +2060,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
{
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
- (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
+ (IAVF_MC_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
struct virtchnl_ether_addr_list *list;
struct iavf_cmd_info args;
uint32_t i;
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH 0/4] Remove limitations coming from legacy VMDq
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (3 preceding siblings ...)
2026-04-03 9:18 ` [PATCH 4/4] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-04-05 18:47 ` Stephen Hemminger
2026-04-29 14:22 ` David Marchand
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
` (10 subsequent siblings)
15 siblings, 1 reply; 146+ messages in thread
From: Stephen Hemminger @ 2026-04-05 18:47 UTC (permalink / raw)
To: David Marchand; +Cc: dev, rjarry, cfontain
On Fri, 3 Apr 2026 11:18:31 +0200
David Marchand <david.marchand@redhat.com> wrote:
> Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
> VMDq has been imposing a maximum number of mac addresses in the
> mac_addr_add/del API.
>
> Nowadays, new Intel drivers do not support the feature and few other
> drivers implement this feature.
>
> This series proposes to flag drivers that support the feature, and
> remove the limit of number of mac addresses for others.
>
> Next step could be to remove the VMDq pool notion from the generic API.
> However I have some concern about this, as changing the quite stable
> mac_addr_add/del API now seems a lot of noise for not much benefit.
>
>
Make sense. The AI review found a couple of things.
Had to poke at it to make a good description
Subject: Re: [PATCH 1/4] ethdev: skip VMDq pools unless configured
Patches 1/4 and 3/4 look good to me.
Patch 2/4 has two issues:
1) In bnxt_reps.c, the new line:
dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
overwrites any default capabilities that were previously set before
the driver callback. The other drivers in this patch had the same
pre-existing pattern (plain assignment before &= ~FLOW_RULE_KEEP),
so for them it's no worse. But bnxt_reps.c previously only did the
&= ~ clear, so this is a new regression. Should be |= instead of =.
2) Several drivers that receive RTE_ETH_DEV_CAPA_VMDQ don't actually
support VMDq: e1000/em, bnxt representors, and i40e VF representors
have no max_vmdq_pools or VMDq configuration. Marking them as
VMDq-capable seems incorrect and would allow users to attempt VMDq
configuration on devices that can't handle it.
Patch 4/4 has a stack overflow:
iavf_add_del_all_mac_addr() allocates list_req on the stack with
addr[IAVF_UC_MACADDR_MAX]. At 32768 entries of ~8 bytes each,
that's roughly 256 KiB on the stack. This will blow the stack
on most configurations. Needs to be heap-allocated.
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH 2/4] ethdev: announce VMDq capability
2026-04-03 9:18 ` [PATCH 2/4] ethdev: announce VMDq capability David Marchand
@ 2026-04-06 22:22 ` Kishore Padmanabha
2026-04-29 14:18 ` David Marchand
2026-06-01 9:32 ` Andrew Rybchenko
1 sibling, 1 reply; 146+ messages in thread
From: Kishore Padmanabha @ 2026-04-06 22:22 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Ajit Khaparde, Bruce Richardson, Rosen Xu,
Anatoly Burakov, Vladimir Medvedkin, Jiawen Wu, Zaiyu Wang,
Thomas Monjalon, Andrew Rybchenko
[-- Attachment #1.1: Type: text/plain, Size: 11777 bytes --]
On Fri, Apr 3, 2026 at 5:19 AM David Marchand <david.marchand@redhat.com>
wrote:
> Let's mark VMDq feature availability as a per device capability.
> We can then enforce API calls related to this feature are done on device
> with such capability.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
> drivers/net/bnxt/bnxt_ethdev.c | 3 ++-
> drivers/net/bnxt/bnxt_reps.c | 1 +
> drivers/net/intel/e1000/em_ethdev.c | 1 +
> drivers/net/intel/e1000/igb_ethdev.c | 1 +
> drivers/net/intel/fm10k/fm10k_ethdev.c | 1 +
> drivers/net/intel/i40e/i40e_ethdev.c | 3 ++-
> drivers/net/intel/i40e/i40e_vf_representor.c | 1 +
> drivers/net/intel/ipn3ke/ipn3ke_representor.c | 3 ++-
> drivers/net/intel/ixgbe/ixgbe_ethdev.c | 2 ++
> drivers/net/txgbe/txgbe_ethdev.c | 1 +
> drivers/net/txgbe/txgbe_ethdev_vf.c | 1 +
> lib/ethdev/rte_ethdev.c | 17 +++++++++++++++++
> lib/ethdev/rte_ethdev.h | 2 ++
> 13 files changed, 34 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/net/bnxt/bnxt_ethdev.c
> b/drivers/net/bnxt/bnxt_ethdev.c
> index b677f9491d..0f783b9e98 100644
> --- a/drivers/net/bnxt/bnxt_ethdev.c
> +++ b/drivers/net/bnxt/bnxt_ethdev.c
> @@ -1214,7 +1214,8 @@ static int bnxt_dev_info_get_op(struct rte_eth_dev
> *eth_dev,
>
> dev_info->speed_capa = bnxt_get_speed_capabilities(bp);
> dev_info->dev_capa = RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
> - RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
> + RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
> + RTE_ETH_DEV_CAPA_VMDQ;
>
We have not been testing the VMDq feature for sometime, planning to
deprecate this feature. Please remove this change.
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
>
> dev_info->default_rxconf = (struct rte_eth_rxconf) {
> diff --git a/drivers/net/bnxt/bnxt_reps.c b/drivers/net/bnxt/bnxt_reps.c
> index e26a086f41..5e834830e2 100644
> --- a/drivers/net/bnxt/bnxt_reps.c
> +++ b/drivers/net/bnxt/bnxt_reps.c
> @@ -649,6 +649,7 @@ int bnxt_rep_dev_info_get_op(struct rte_eth_dev
> *eth_dev,
> dev_info->max_tx_queues = max_rx_rings;
> dev_info->reta_size = bnxt_rss_hash_tbl_size(parent_bp);
> dev_info->hash_key_size = 40;
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
>
> We have not been testing the VMDq feature for sometime, planning to
deprecate this feature. Please remove this change.
> /* MTU specifics */
> diff --git a/drivers/net/intel/e1000/em_ethdev.c
> b/drivers/net/intel/e1000/em_ethdev.c
> index 9e15e882b9..389744ad5e 100644
> --- a/drivers/net/intel/e1000/em_ethdev.c
> +++ b/drivers/net/intel/e1000/em_ethdev.c
> @@ -1175,6 +1175,7 @@ eth_em_infos_get(struct rte_eth_dev *dev, struct
> rte_eth_dev_info *dev_info)
> RTE_ETH_LINK_SPEED_100M_HD |
> RTE_ETH_LINK_SPEED_100M |
> RTE_ETH_LINK_SPEED_1G;
>
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
>
> /* Preferred queue parameters */
> diff --git a/drivers/net/intel/e1000/igb_ethdev.c
> b/drivers/net/intel/e1000/igb_ethdev.c
> index ef1599ac38..fe68c18417 100644
> --- a/drivers/net/intel/e1000/igb_ethdev.c
> +++ b/drivers/net/intel/e1000/igb_ethdev.c
> @@ -2324,6 +2324,7 @@ eth_igb_infos_get(struct rte_eth_dev *dev, struct
> rte_eth_dev_info *dev_info)
> dev_info->tx_queue_offload_capa =
> igb_get_tx_queue_offloads_capa(dev);
> dev_info->tx_offload_capa = igb_get_tx_port_offloads_capa(dev) |
> dev_info->tx_queue_offload_capa;
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
>
> switch (hw->mac.type) {
> diff --git a/drivers/net/intel/fm10k/fm10k_ethdev.c
> b/drivers/net/intel/fm10k/fm10k_ethdev.c
> index 97f61afec2..037d2206fd 100644
> --- a/drivers/net/intel/fm10k/fm10k_ethdev.c
> +++ b/drivers/net/intel/fm10k/fm10k_ethdev.c
> @@ -1444,6 +1444,7 @@ fm10k_dev_infos_get(struct rte_eth_dev *dev,
> dev_info->speed_capa = RTE_ETH_LINK_SPEED_1G |
> RTE_ETH_LINK_SPEED_2_5G |
> RTE_ETH_LINK_SPEED_10G | RTE_ETH_LINK_SPEED_25G |
> RTE_ETH_LINK_SPEED_40G | RTE_ETH_LINK_SPEED_100G;
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
>
> return 0;
> }
> diff --git a/drivers/net/intel/i40e/i40e_ethdev.c
> b/drivers/net/intel/i40e/i40e_ethdev.c
> index 100a751225..64c29c6e85 100644
> --- a/drivers/net/intel/i40e/i40e_ethdev.c
> +++ b/drivers/net/intel/i40e/i40e_ethdev.c
> @@ -3878,7 +3878,8 @@ i40e_dev_info_get(struct rte_eth_dev *dev, struct
> rte_eth_dev_info *dev_info)
>
> dev_info->dev_capa =
> RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
> - RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
> + RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
> + RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
>
> dev_info->hash_key_size = (I40E_PFQF_HKEY_MAX_INDEX + 1) *
> diff --git a/drivers/net/intel/i40e/i40e_vf_representor.c
> b/drivers/net/intel/i40e/i40e_vf_representor.c
> index e8f0bb62a0..d31148acb5 100644
> --- a/drivers/net/intel/i40e/i40e_vf_representor.c
> +++ b/drivers/net/intel/i40e/i40e_vf_representor.c
> @@ -33,6 +33,7 @@ i40e_vf_representor_dev_infos_get(struct rte_eth_dev
> *ethdev,
> /* get dev info for the vdev */
> dev_info->device = ethdev->device;
>
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
>
> dev_info->max_rx_queues = ethdev->data->nb_rx_queues;
> diff --git a/drivers/net/intel/ipn3ke/ipn3ke_representor.c
> b/drivers/net/intel/ipn3ke/ipn3ke_representor.c
> index cd34d08055..d581ee3c37 100644
> --- a/drivers/net/intel/ipn3ke/ipn3ke_representor.c
> +++ b/drivers/net/intel/ipn3ke/ipn3ke_representor.c
> @@ -95,7 +95,8 @@ ipn3ke_rpst_dev_infos_get(struct rte_eth_dev *ethdev,
>
> dev_info->dev_capa =
> RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
> - RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
> + RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
> + RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
>
> dev_info->switch_info.name = ethdev->device->name;
> diff --git a/drivers/net/intel/ixgbe/ixgbe_ethdev.c
> b/drivers/net/intel/ixgbe/ixgbe_ethdev.c
> index 57d929cf2c..5d886b3e28 100644
> --- a/drivers/net/intel/ixgbe/ixgbe_ethdev.c
> +++ b/drivers/net/intel/ixgbe/ixgbe_ethdev.c
> @@ -3997,6 +3997,7 @@ ixgbe_dev_info_get(struct rte_eth_dev *dev, struct
> rte_eth_dev_info *dev_info)
> dev_info->max_mtu = dev_info->max_rx_pktlen - IXGBE_ETH_OVERHEAD;
> dev_info->min_mtu = RTE_ETHER_MIN_MTU;
> dev_info->vmdq_queue_num = dev_info->max_rx_queues;
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->rx_queue_offload_capa = ixgbe_get_rx_queue_offloads(dev);
> dev_info->rx_offload_capa = (ixgbe_get_rx_port_offloads(dev) |
> dev_info->rx_queue_offload_capa);
> @@ -4115,6 +4116,7 @@ ixgbevf_dev_info_get(struct rte_eth_dev *dev,
> dev_info->max_vmdq_pools = RTE_ETH_16_POOLS;
> else
> dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->rx_queue_offload_capa = ixgbe_get_rx_queue_offloads(dev);
> dev_info->rx_offload_capa = (ixgbe_get_rx_port_offloads(dev) |
> dev_info->rx_queue_offload_capa);
> diff --git a/drivers/net/txgbe/txgbe_ethdev.c
> b/drivers/net/txgbe/txgbe_ethdev.c
> index 5d360f8305..bd818e8269 100644
> --- a/drivers/net/txgbe/txgbe_ethdev.c
> +++ b/drivers/net/txgbe/txgbe_ethdev.c
> @@ -2836,6 +2836,7 @@ txgbe_dev_info_get(struct rte_eth_dev *dev, struct
> rte_eth_dev_info *dev_info)
> dev_info->max_vfs = pci_dev->max_vfs;
> dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
> dev_info->vmdq_queue_num = dev_info->max_rx_queues;
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
> dev_info->rx_queue_offload_capa = txgbe_get_rx_queue_offloads(dev);
> dev_info->rx_offload_capa = (txgbe_get_rx_port_offloads(dev) |
> diff --git a/drivers/net/txgbe/txgbe_ethdev_vf.c
> b/drivers/net/txgbe/txgbe_ethdev_vf.c
> index 39a5fff65c..934763574c 100644
> --- a/drivers/net/txgbe/txgbe_ethdev_vf.c
> +++ b/drivers/net/txgbe/txgbe_ethdev_vf.c
> @@ -572,6 +572,7 @@ txgbevf_dev_info_get(struct rte_eth_dev *dev,
> dev_info->max_hash_mac_addrs = TXGBE_VMDQ_NUM_UC_MAC;
> dev_info->max_vfs = pci_dev->max_vfs;
> dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
> + dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
> dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
> dev_info->rx_queue_offload_capa = txgbe_get_rx_queue_offloads(dev);
> dev_info->rx_offload_capa = (txgbe_get_rx_port_offloads(dev) |
> diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
> index 9577b7d848..7ba539e796 100644
> --- a/lib/ethdev/rte_ethdev.c
> +++ b/lib/ethdev/rte_ethdev.c
> @@ -158,6 +158,7 @@ static const struct {
> {RTE_ETH_DEV_CAPA_RXQ_SHARE, "RXQ_SHARE"},
> {RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP, "FLOW_RULE_KEEP"},
> {RTE_ETH_DEV_CAPA_FLOW_SHARED_OBJECT_KEEP,
> "FLOW_SHARED_OBJECT_KEEP"},
> + {RTE_ETH_DEV_CAPA_VMDQ, "VMDQ"},
> };
>
> enum {
> @@ -1581,6 +1582,22 @@ rte_eth_dev_configure(uint16_t port_id, uint16_t
> nb_rx_q, uint16_t nb_tx_q,
> goto rollback;
> }
>
> + if (!(dev_info.dev_capa & RTE_ETH_DEV_CAPA_VMDQ)) {
> + if ((dev_conf->rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG)
> != 0) {
> + RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does
> not support VMDq rx mode",
> + port_id);
> + ret = -EINVAL;
> + goto rollback;
> + }
> + if (dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_DCB ||
> + dev_conf->txmode.mq_mode ==
> RTE_ETH_MQ_TX_VMDQ_ONLY) {
> + RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does
> not support VMDq tx mode",
> + port_id);
> + ret = -EINVAL;
> + goto rollback;
> + }
> + }
> +
> /*
> * Setup new number of Rx/Tx queues and reconfigure device.
> */
> diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
> index 0d8e2d0236..62c72de0e5 100644
> --- a/lib/ethdev/rte_ethdev.h
> +++ b/lib/ethdev/rte_ethdev.h
> @@ -1696,6 +1696,8 @@ struct rte_eth_conf {
> #define RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP RTE_BIT64(3)
> /** Device supports keeping shared flow objects across restart. */
> #define RTE_ETH_DEV_CAPA_FLOW_SHARED_OBJECT_KEEP RTE_BIT64(4)
> +/** Device supports VMDq. */
> +#define RTE_ETH_DEV_CAPA_VMDQ RTE_BIT64(5)
> /**@}*/
>
> /*
> --
> 2.53.0
>
>
[-- Attachment #1.2: Type: text/html, Size: 13987 bytes --]
[-- Attachment #2: S/MIME Cryptographic Signature --]
[-- Type: application/pkcs7-signature, Size: 5493 bytes --]
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH 2/4] ethdev: announce VMDq capability
2026-04-06 22:22 ` Kishore Padmanabha
@ 2026-04-29 14:18 ` David Marchand
2026-05-18 22:12 ` Kishore Padmanabha
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-04-29 14:18 UTC (permalink / raw)
To: Kishore Padmanabha
Cc: dev, rjarry, cfontain, Ajit Khaparde, Bruce Richardson, Rosen Xu,
Anatoly Burakov, Vladimir Medvedkin, Jiawen Wu, Zaiyu Wang,
Thomas Monjalon, Andrew Rybchenko
On Tue, 7 Apr 2026 at 00:22, Kishore Padmanabha
<kishore.padmanabha@broadcom.com> wrote:
>> diff --git a/drivers/net/bnxt/bnxt_ethdev.c b/drivers/net/bnxt/bnxt_ethdev.c
>> index b677f9491d..0f783b9e98 100644
>> --- a/drivers/net/bnxt/bnxt_ethdev.c
>> +++ b/drivers/net/bnxt/bnxt_ethdev.c
>> @@ -1214,7 +1214,8 @@ static int bnxt_dev_info_get_op(struct rte_eth_dev *eth_dev,
>>
>> dev_info->speed_capa = bnxt_get_speed_capabilities(bp);
>> dev_info->dev_capa = RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
>> - RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
>> + RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
>> + RTE_ETH_DEV_CAPA_VMDQ;
>
> We have not been testing the VMDq feature for sometime, planning to deprecate this feature. Please remove this change.
Currently, the driver supports VMDq:
$ git grep -i vmdq doc/guides/nics/features/bnxt.ini
doc/guides/nics/features/bnxt.ini:VMDq = Y
I don't intend to own this feature drop with my series :-).
Please announce such a deprecation and flag it properly.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH 0/4] Remove limitations coming from legacy VMDq
2026-04-05 18:47 ` [PATCH 0/4] Remove limitations coming from legacy VMDq Stephen Hemminger
@ 2026-04-29 14:22 ` David Marchand
0 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-04-29 14:22 UTC (permalink / raw)
To: Stephen Hemminger; +Cc: dev, rjarry, cfontain
On Sun, 5 Apr 2026 at 20:47, Stephen Hemminger
<stephen@networkplumber.org> wrote:
>
> On Fri, 3 Apr 2026 11:18:31 +0200
> David Marchand <david.marchand@redhat.com> wrote:
>
> > Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
> > VMDq has been imposing a maximum number of mac addresses in the
> > mac_addr_add/del API.
> >
> > Nowadays, new Intel drivers do not support the feature and few other
> > drivers implement this feature.
> >
> > This series proposes to flag drivers that support the feature, and
> > remove the limit of number of mac addresses for others.
> >
> > Next step could be to remove the VMDq pool notion from the generic API.
> > However I have some concern about this, as changing the quite stable
> > mac_addr_add/del API now seems a lot of noise for not much benefit.
> >
> >
>
> Make sense. The AI review found a couple of things.
> Had to poke at it to make a good description
>
> Subject: Re: [PATCH 1/4] ethdev: skip VMDq pools unless configured
>
> Patches 1/4 and 3/4 look good to me.
>
> Patch 2/4 has two issues:
>
> 1) In bnxt_reps.c, the new line:
>
> dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
>
> overwrites any default capabilities that were previously set before
> the driver callback. The other drivers in this patch had the same
> pre-existing pattern (plain assignment before &= ~FLOW_RULE_KEEP),
> so for them it's no worse. But bnxt_reps.c previously only did the
> &= ~ clear, so this is a new regression. Should be |= instead of =.
The explicit &= ~FLOW_RULE_KEEP was added as a "documentation" hint
for maintainers.
2fe6f1b76279 ("drivers/net: advertise no support for keeping flow rules")
Though I don't see much activity on the topic, so this hint probably
missed the target.
In any case, dev_capa is set to 0 from ethdev before calling the
driver op, so nothing is broken with an explicit =.
>
> 2) Several drivers that receive RTE_ETH_DEV_CAPA_VMDQ don't actually
> support VMDq: e1000/em, bnxt representors, and i40e VF representors
> have no max_vmdq_pools or VMDq configuration. Marking them as
> VMDq-capable seems incorrect and would allow users to attempt VMDq
> configuration on devices that can't handle it.
This one is interesting and it seems to make sense, I'll double check.
>
> Patch 4/4 has a stack overflow:
>
> iavf_add_del_all_mac_addr() allocates list_req on the stack with
> addr[IAVF_UC_MACADDR_MAX]. At 32768 entries of ~8 bytes each,
> that's roughly 256 KiB on the stack. This will blow the stack
> on most configurations. Needs to be heap-allocated.
The heap allocations were removed recently, I'll see how I can rework.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v2 0/5] Remove limitations coming from legacy VMDq
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (4 preceding siblings ...)
2026-04-05 18:47 ` [PATCH 0/4] Remove limitations coming from legacy VMDq Stephen Hemminger
@ 2026-05-06 12:35 ` David Marchand
2026-05-06 12:35 ` [PATCH v2 1/5] ethdev: skip VMDq pools unless configured David Marchand
` (5 more replies)
2026-05-10 17:03 ` [PATCH v3 " David Marchand
` (9 subsequent siblings)
15 siblings, 6 replies; 146+ messages in thread
From: David Marchand @ 2026-05-06 12:35 UTC (permalink / raw)
To: dev
Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
VMDq has been imposing a maximum number of mac addresses in the
mac_addr_add/del API.
Nowadays, new Intel drivers do not support the feature and few other
drivers implement this feature.
This series proposes to flag drivers that support the feature, and
remove the limit of number of mac addresses for others.
Next step could be to remove the VMDq pool notion from the generic API.
However I have some concern about this, as changing the quite stable
mac_addr_add/del API now seems a lot of noise for not much benefit.
--
David Marchand
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
- fixed buffer overflow on mailbox messages during port restart/VF reset,
- fixed duplicate MAC address installation on port start/restart,
David Marchand (5):
ethdev: skip VMDq pools unless configured
ethdev: announce VMDq capability
ethdev: hide VMDq internal sizes
net/iavf: accept up to 32k unicast MAC addresses
net/iavf: fix duplicate MAC addresses install
drivers/net/bnxt/bnxt_ethdev.c | 3 +-
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
drivers/net/intel/e1000/igb_ethdev.c | 1 +
drivers/net/intel/fm10k/fm10k_ethdev.c | 1 +
drivers/net/intel/i40e/i40e_ethdev.c | 3 +-
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 41 ++++++---
drivers/net/intel/iavf/iavf_vchnl.c | 117 ++++++++++++++++++-------
drivers/net/intel/ixgbe/ixgbe_ethdev.c | 2 +
drivers/net/txgbe/txgbe_ethdev.c | 1 +
drivers/net/txgbe/txgbe_ethdev_vf.c | 1 +
lib/ethdev/ethdev_driver.h | 8 +-
lib/ethdev/rte_ethdev.c | 45 ++++++++--
lib/ethdev/rte_ethdev.h | 8 +-
14 files changed, 172 insertions(+), 65 deletions(-)
--
2.53.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v2 1/5] ethdev: skip VMDq pools unless configured
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
@ 2026-05-06 12:35 ` David Marchand
2026-06-01 9:35 ` Andrew Rybchenko
2026-05-06 12:35 ` [PATCH v2 2/5] ethdev: announce VMDq capability David Marchand
` (4 subsequent siblings)
5 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-05-06 12:35 UTC (permalink / raw)
To: dev
Cc: Nithin Dabilpuram, Kiran Kumar K, Sunil Kumar Kori, Satha Rao,
Harman Kalra, Thomas Monjalon, Andrew Rybchenko
The mac_addr_add API describes that only the 0 pool should be passed
unless VMDq has been enabled, though there was no validation so far.
Add such a check, then cleanup the related operations (adding, removing,
restoring).
As a side effect, the net/cnxk does not need to manually reset the
mac_pool_sel[] array.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
lib/ethdev/rte_ethdev.c | 28 +++++++++++++++++++++-------
2 files changed, 21 insertions(+), 8 deletions(-)
diff --git a/drivers/net/cnxk/cnxk_ethdev_ops.c b/drivers/net/cnxk/cnxk_ethdev_ops.c
index 49e77e49a6..75decf7098 100644
--- a/drivers/net/cnxk/cnxk_ethdev_ops.c
+++ b/drivers/net/cnxk/cnxk_ethdev_ops.c
@@ -1240,7 +1240,6 @@ cnxk_nix_mc_addr_list_configure(struct rte_eth_dev *eth_dev, struct rte_ether_ad
/* Update address in NIC data structure */
rte_ether_addr_copy(&mc_addr_set[i], &data->mac_addrs[j]);
rte_ether_addr_copy(&mc_addr_set[i], &dev->dmac_addrs[j]);
- data->mac_pool_sel[j] = RTE_BIT64(0);
}
roc_nix_npc_promisc_ena_dis(nix, true);
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 2edc7a362e..9577b7d848 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1680,7 +1680,10 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
continue;
pool = 0;
- pool_mask = dev->data->mac_pool_sel[i];
+ if ((dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0)
+ pool_mask = dev->data->mac_pool_sel[i];
+ else
+ pool_mask = 1;
do {
if (pool_mask & UINT64_C(1))
@@ -5390,8 +5393,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
uint32_t pool)
{
struct rte_eth_dev *dev;
- int index;
uint64_t pool_mask;
+ bool vmdq;
+ int index;
int ret;
RTE_ETH_VALID_PORTID_OR_ERR_RET(port_id, -ENODEV);
@@ -5416,6 +5420,12 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
RTE_ETHDEV_LOG_LINE(ERR, "Pool ID must be 0-%d", RTE_ETH_64_POOLS - 1);
return -EINVAL;
}
+ vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+ if (!vmdq && pool != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Port %u: VMDq is not configured (pool %d)",
+ port_id, pool);
+ return -EINVAL;
+ }
index = eth_dev_get_mac_addr_index(port_id, addr);
if (index < 0) {
@@ -5425,7 +5435,7 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
port_id);
return -ENOSPC;
}
- } else {
+ } else if (vmdq) {
pool_mask = dev->data->mac_pool_sel[index];
/* Check if both MAC address and pool is already there, and do nothing */
@@ -5440,8 +5450,10 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
/* Update address in NIC data structure */
rte_ether_addr_copy(addr, &dev->data->mac_addrs[index]);
- /* Update pool bitmap in NIC data structure */
- dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ if (vmdq) {
+ /* Update pool bitmap in NIC data structure */
+ dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ }
}
ret = eth_err(port_id, ret);
@@ -5486,8 +5498,10 @@ rte_eth_dev_mac_addr_remove(uint16_t port_id, struct rte_ether_addr *addr)
/* Update address in NIC data structure */
rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
- /* reset pool bitmap */
- dev->data->mac_pool_sel[index] = 0;
+ if ((dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ /* reset pool bitmap */
+ dev->data->mac_pool_sel[index] = 0;
+ }
rte_ethdev_trace_mac_addr_remove(port_id, addr);
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v2 2/5] ethdev: announce VMDq capability
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
2026-05-06 12:35 ` [PATCH v2 1/5] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-05-06 12:35 ` David Marchand
2026-06-01 9:36 ` Andrew Rybchenko
2026-05-06 12:35 ` [PATCH v2 3/5] ethdev: hide VMDq internal sizes David Marchand
` (3 subsequent siblings)
5 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-05-06 12:35 UTC (permalink / raw)
To: dev
Cc: Kishore Padmanabha, Ajit Khaparde, Bruce Richardson,
Anatoly Burakov, Vladimir Medvedkin, Jiawen Wu, Zaiyu Wang,
Thomas Monjalon, Andrew Rybchenko
Let's mark VMDq feature availability as a per device capability.
We can then enforce API calls related to this feature are done on device
with such capability.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
---
drivers/net/bnxt/bnxt_ethdev.c | 3 ++-
drivers/net/intel/e1000/igb_ethdev.c | 1 +
drivers/net/intel/fm10k/fm10k_ethdev.c | 1 +
drivers/net/intel/i40e/i40e_ethdev.c | 3 ++-
drivers/net/intel/ixgbe/ixgbe_ethdev.c | 2 ++
drivers/net/txgbe/txgbe_ethdev.c | 1 +
drivers/net/txgbe/txgbe_ethdev_vf.c | 1 +
lib/ethdev/rte_ethdev.c | 17 +++++++++++++++++
lib/ethdev/rte_ethdev.h | 2 ++
9 files changed, 29 insertions(+), 2 deletions(-)
diff --git a/drivers/net/bnxt/bnxt_ethdev.c b/drivers/net/bnxt/bnxt_ethdev.c
index b677f9491d..0f783b9e98 100644
--- a/drivers/net/bnxt/bnxt_ethdev.c
+++ b/drivers/net/bnxt/bnxt_ethdev.c
@@ -1214,7 +1214,8 @@ static int bnxt_dev_info_get_op(struct rte_eth_dev *eth_dev,
dev_info->speed_capa = bnxt_get_speed_capabilities(bp);
dev_info->dev_capa = RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
- RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
+ RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
+ RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->default_rxconf = (struct rte_eth_rxconf) {
diff --git a/drivers/net/intel/e1000/igb_ethdev.c b/drivers/net/intel/e1000/igb_ethdev.c
index ef1599ac38..fe68c18417 100644
--- a/drivers/net/intel/e1000/igb_ethdev.c
+++ b/drivers/net/intel/e1000/igb_ethdev.c
@@ -2324,6 +2324,7 @@ eth_igb_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->tx_queue_offload_capa = igb_get_tx_queue_offloads_capa(dev);
dev_info->tx_offload_capa = igb_get_tx_port_offloads_capa(dev) |
dev_info->tx_queue_offload_capa;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
switch (hw->mac.type) {
diff --git a/drivers/net/intel/fm10k/fm10k_ethdev.c b/drivers/net/intel/fm10k/fm10k_ethdev.c
index 97f61afec2..037d2206fd 100644
--- a/drivers/net/intel/fm10k/fm10k_ethdev.c
+++ b/drivers/net/intel/fm10k/fm10k_ethdev.c
@@ -1444,6 +1444,7 @@ fm10k_dev_infos_get(struct rte_eth_dev *dev,
dev_info->speed_capa = RTE_ETH_LINK_SPEED_1G | RTE_ETH_LINK_SPEED_2_5G |
RTE_ETH_LINK_SPEED_10G | RTE_ETH_LINK_SPEED_25G |
RTE_ETH_LINK_SPEED_40G | RTE_ETH_LINK_SPEED_100G;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
return 0;
}
diff --git a/drivers/net/intel/i40e/i40e_ethdev.c b/drivers/net/intel/i40e/i40e_ethdev.c
index 100a751225..64c29c6e85 100644
--- a/drivers/net/intel/i40e/i40e_ethdev.c
+++ b/drivers/net/intel/i40e/i40e_ethdev.c
@@ -3878,7 +3878,8 @@ i40e_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->dev_capa =
RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
- RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
+ RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
+ RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->hash_key_size = (I40E_PFQF_HKEY_MAX_INDEX + 1) *
diff --git a/drivers/net/intel/ixgbe/ixgbe_ethdev.c b/drivers/net/intel/ixgbe/ixgbe_ethdev.c
index 57d929cf2c..5d886b3e28 100644
--- a/drivers/net/intel/ixgbe/ixgbe_ethdev.c
+++ b/drivers/net/intel/ixgbe/ixgbe_ethdev.c
@@ -3997,6 +3997,7 @@ ixgbe_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->max_mtu = dev_info->max_rx_pktlen - IXGBE_ETH_OVERHEAD;
dev_info->min_mtu = RTE_ETHER_MIN_MTU;
dev_info->vmdq_queue_num = dev_info->max_rx_queues;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->rx_queue_offload_capa = ixgbe_get_rx_queue_offloads(dev);
dev_info->rx_offload_capa = (ixgbe_get_rx_port_offloads(dev) |
dev_info->rx_queue_offload_capa);
@@ -4115,6 +4116,7 @@ ixgbevf_dev_info_get(struct rte_eth_dev *dev,
dev_info->max_vmdq_pools = RTE_ETH_16_POOLS;
else
dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->rx_queue_offload_capa = ixgbe_get_rx_queue_offloads(dev);
dev_info->rx_offload_capa = (ixgbe_get_rx_port_offloads(dev) |
dev_info->rx_queue_offload_capa);
diff --git a/drivers/net/txgbe/txgbe_ethdev.c b/drivers/net/txgbe/txgbe_ethdev.c
index 5d360f8305..bd818e8269 100644
--- a/drivers/net/txgbe/txgbe_ethdev.c
+++ b/drivers/net/txgbe/txgbe_ethdev.c
@@ -2836,6 +2836,7 @@ txgbe_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->max_vfs = pci_dev->max_vfs;
dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
dev_info->vmdq_queue_num = dev_info->max_rx_queues;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->rx_queue_offload_capa = txgbe_get_rx_queue_offloads(dev);
dev_info->rx_offload_capa = (txgbe_get_rx_port_offloads(dev) |
diff --git a/drivers/net/txgbe/txgbe_ethdev_vf.c b/drivers/net/txgbe/txgbe_ethdev_vf.c
index 39a5fff65c..934763574c 100644
--- a/drivers/net/txgbe/txgbe_ethdev_vf.c
+++ b/drivers/net/txgbe/txgbe_ethdev_vf.c
@@ -572,6 +572,7 @@ txgbevf_dev_info_get(struct rte_eth_dev *dev,
dev_info->max_hash_mac_addrs = TXGBE_VMDQ_NUM_UC_MAC;
dev_info->max_vfs = pci_dev->max_vfs;
dev_info->max_vmdq_pools = RTE_ETH_64_POOLS;
+ dev_info->dev_capa = RTE_ETH_DEV_CAPA_VMDQ;
dev_info->dev_capa &= ~RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP;
dev_info->rx_queue_offload_capa = txgbe_get_rx_queue_offloads(dev);
dev_info->rx_offload_capa = (txgbe_get_rx_port_offloads(dev) |
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 9577b7d848..7ba539e796 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -158,6 +158,7 @@ static const struct {
{RTE_ETH_DEV_CAPA_RXQ_SHARE, "RXQ_SHARE"},
{RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP, "FLOW_RULE_KEEP"},
{RTE_ETH_DEV_CAPA_FLOW_SHARED_OBJECT_KEEP, "FLOW_SHARED_OBJECT_KEEP"},
+ {RTE_ETH_DEV_CAPA_VMDQ, "VMDQ"},
};
enum {
@@ -1581,6 +1582,22 @@ rte_eth_dev_configure(uint16_t port_id, uint16_t nb_rx_q, uint16_t nb_tx_q,
goto rollback;
}
+ if (!(dev_info.dev_capa & RTE_ETH_DEV_CAPA_VMDQ)) {
+ if ((dev_conf->rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq rx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ if (dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_DCB ||
+ dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_ONLY) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq tx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ }
+
/*
* Setup new number of Rx/Tx queues and reconfigure device.
*/
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index 0d8e2d0236..62c72de0e5 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -1696,6 +1696,8 @@ struct rte_eth_conf {
#define RTE_ETH_DEV_CAPA_FLOW_RULE_KEEP RTE_BIT64(3)
/** Device supports keeping shared flow objects across restart. */
#define RTE_ETH_DEV_CAPA_FLOW_SHARED_OBJECT_KEEP RTE_BIT64(4)
+/** Device supports VMDq. */
+#define RTE_ETH_DEV_CAPA_VMDQ RTE_BIT64(5)
/**@}*/
/*
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v2 3/5] ethdev: hide VMDq internal sizes
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
2026-05-06 12:35 ` [PATCH v2 1/5] ethdev: skip VMDq pools unless configured David Marchand
2026-05-06 12:35 ` [PATCH v2 2/5] ethdev: announce VMDq capability David Marchand
@ 2026-05-06 12:35 ` David Marchand
2026-05-06 12:35 ` [PATCH v2 4/5] net/iavf: accept up to 32k unicast MAC addresses David Marchand
` (2 subsequent siblings)
5 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-05-06 12:35 UTC (permalink / raw)
To: dev; +Cc: Thomas Monjalon, Andrew Rybchenko
Hide RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
in the driver API as those (ambiguous) macros are only a driver concern.
In practice, this is only used by the bnxt and ixgbe (+ clones) drivers.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
lib/ethdev/ethdev_driver.h | 8 +++++++-
lib/ethdev/rte_ethdev.h | 6 ------
2 files changed, 7 insertions(+), 7 deletions(-)
diff --git a/lib/ethdev/ethdev_driver.h b/lib/ethdev/ethdev_driver.h
index 1255cd6f2c..a4e9cf5b90 100644
--- a/lib/ethdev/ethdev_driver.h
+++ b/lib/ethdev/ethdev_driver.h
@@ -119,6 +119,12 @@ struct __rte_cache_aligned rte_eth_dev {
struct rte_eth_dev_sriov;
struct rte_eth_dev_owner;
+/* Definitions used for receive MAC address */
+#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
+
+/* Definitions used for unicast hash */
+#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
+
/**
* @internal
* The data part, with no function pointers, associated with each Ethernet
@@ -153,7 +159,7 @@ struct __rte_cache_aligned rte_eth_dev_data {
* The first entry (index zero) is the default address.
*/
struct rte_ether_addr *mac_addrs;
- /** Bitmap associating MAC addresses to pools */
+ /** Bitmap associating MAC addresses to VMDq pools */
uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
/**
* Device Ethernet MAC addresses of hash filtering.
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index 62c72de0e5..6c1984e679 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -903,12 +903,6 @@ rte_eth_rss_hf_refine(uint64_t rss_hf)
#define RTE_ETH_VLAN_ID_MAX 0x0FFF /**< VLAN ID is in lower 12 bits*/
/**@}*/
-/* Definitions used for receive MAC address */
-#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
-
-/* Definitions used for unicast hash */
-#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
-
/**@{@name VMDq Rx mode
* @see rte_eth_vmdq_rx_conf.rx_mode
*/
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v2 4/5] net/iavf: accept up to 32k unicast MAC addresses
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
` (2 preceding siblings ...)
2026-05-06 12:35 ` [PATCH v2 3/5] ethdev: hide VMDq internal sizes David Marchand
@ 2026-05-06 12:35 ` David Marchand
2026-05-06 12:35 ` [PATCH v2 5/5] net/iavf: fix duplicate MAC addresses install David Marchand
2026-05-07 2:51 ` [PATCH v2 0/5] Remove limitations coming from legacy VMDq Stephen Hemminger
5 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-05-06 12:35 UTC (permalink / raw)
To: dev; +Cc: Vladimir Medvedkin
E810 hardware provides 32k switch lookups.
Thanks to this, it is possible to allow a lot more secondary mac
addresses than what is possible today.
In practice, the maximum number of macs available per port may be lower
and depends on usage by other (trusted?) VFs on the same PF.
There is no way to figure out this limit but to try adding a mac address
and get an error from the PF driver.
Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
accordingly.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v1:
- fixed buffer overflow on mailbox messages during port restart/VF reset,
---
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 10 +--
drivers/net/intel/iavf/iavf_vchnl.c | 117 +++++++++++++++++++--------
3 files changed, 93 insertions(+), 39 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 403c61e2e8..f1dede0694 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -31,7 +31,8 @@
#define IAVF_IRQ_MAP_NUM_PER_BUF 128
#define IAVF_RXTX_QUEUE_CHUNKS_NUM 2
-#define IAVF_NUM_MACADDR_MAX 64
+#define IAVF_UC_MACADDR_MAX 32768
+#define IAVF_MC_MACADDR_MAX 64
#define IAVF_DEV_WATCHDOG_PERIOD 2000 /* microseconds, set 0 to disable*/
@@ -253,7 +254,7 @@ struct iavf_info {
uint32_t link_speed;
/* Multicast addrs */
- struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
+ struct rte_ether_addr mc_addrs[IAVF_MC_MACADDR_MAX];
uint16_t mc_addrs_num; /* Multicast mac addresses number */
struct iavf_vsi vsi;
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index 1eca20bc9a..c69a012d50 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -379,10 +379,10 @@ iavf_set_mc_addr_list(struct rte_eth_dev *dev,
IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
int err, ret;
- if (mc_addrs_num > IAVF_NUM_MACADDR_MAX) {
+ if (mc_addrs_num > IAVF_MC_MACADDR_MAX) {
PMD_DRV_LOG(ERR,
"can't add more than a limited number (%u) of addresses.",
- (uint32_t)IAVF_NUM_MACADDR_MAX);
+ (uint32_t)IAVF_MC_MACADDR_MAX);
return -EINVAL;
}
@@ -1120,7 +1120,7 @@ iavf_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->hash_key_size = vf->vf_res->rss_key_size;
dev_info->reta_size = vf->vf_res->rss_lut_size;
dev_info->flow_type_rss_offloads = IAVF_RSS_OFFLOAD_ALL;
- dev_info->max_mac_addrs = IAVF_NUM_MACADDR_MAX;
+ dev_info->max_mac_addrs = IAVF_UC_MACADDR_MAX;
dev_info->dev_capa =
RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
@@ -2822,11 +2822,11 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
/* copy mac addr */
eth_dev->data->mac_addrs = rte_zmalloc(
- "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX, 0);
+ "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_UC_MACADDR_MAX, 0);
if (!eth_dev->data->mac_addrs) {
PMD_INIT_LOG(ERR, "Failed to allocate %d bytes needed to"
" store MAC addresses",
- RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX);
+ RTE_ETHER_ADDR_LEN * IAVF_UC_MACADDR_MAX);
ret = -ENOMEM;
goto init_vf_err;
}
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index 08dd6f2d7f..144968db35 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1437,48 +1437,101 @@ iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num)
return 0;
}
-void
-iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+static int
+iavf_add_del_uc_addr_bulk(struct iavf_adapter *adapter, struct rte_ether_addr *addrs,
+ uint32_t nb_addrs, bool add)
{
+#define IAVF_ETH_ADDR_PER_REQ \
+ ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
+ sizeof(struct virtchnl_ether_addr))
struct {
struct virtchnl_ether_addr_list list;
- struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
- } list_req = {0};
- struct virtchnl_ether_addr_list *list = &list_req.list;
+ struct virtchnl_ether_addr addr[IAVF_ETH_ADDR_PER_REQ];
+ } cmd_buffer;
+#undef IAVF_ETH_ADDR_PER_REQ
+ struct virtchnl_ether_addr_list *list = &cmd_buffer.list;
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
- struct iavf_cmd_info args = {0};
- int err, i;
- size_t buf_len;
- for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
- struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
- struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
+ for (uint32_t i = 0; i < nb_addrs; i++) {
+ size_t buf_len = sizeof(struct virtchnl_ether_addr_list) +
+ sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ struct iavf_cmd_info args;
+ uint32_t batch;
+ int err;
- /* ignore empty addresses */
- if (rte_is_zero_ether_addr(addr))
- continue;
+ batch = i % RTE_DIM(cmd_buffer.addr);
+
+ if (batch == 0) {
+ memset(&cmd_buffer, 0, sizeof(cmd_buffer));
+ list->vsi_id = vf->vsi_res->vsi_id;
+ list->num_elements = 0;
+ }
+
+ rte_memcpy(list->list[batch].addr, addrs[i].addr_bytes,
+ sizeof(list->list[batch].addr));
+ list->list[batch].type = VIRTCHNL_ETHER_ADDR_EXTRA;
list->num_elements++;
- memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
- vc_addr->type = (list->num_elements == 1) ?
- VIRTCHNL_ETHER_ADDR_PRIMARY :
- VIRTCHNL_ETHER_ADDR_EXTRA;
+ if (batch != RTE_DIM(cmd_buffer.addr) - 1 && i != nb_addrs - 1)
+ continue;
+
+ buf_len = sizeof(struct virtchnl_ether_addr_list) +
+ sizeof(struct virtchnl_ether_addr) * list->num_elements;
+
+ memset(&args, 0, sizeof(args));
+ args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
+ args.in_args = (uint8_t *)list;
+ args.in_args_size = buf_len;
+ args.out_buffer = vf->aq_resp;
+ args.out_size = IAVF_AQ_BUF_SZ;
+ err = iavf_execute_vf_cmd_safe(adapter, &args, 0);
+ if (err != 0) {
+ PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
+ return err;
+ }
+
+ PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
}
- /* for some reason PF side checks for buffer being too big, so adjust it down */
- buf_len = sizeof(struct virtchnl_ether_addr_list) +
- sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ return 0;
+}
- list->vsi_id = vf->vsi_res->vsi_id;
- args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
- args.in_args = (uint8_t *)list;
- args.in_args_size = buf_len;
- args.out_buffer = vf->aq_resp;
- args.out_size = IAVF_AQ_BUF_SZ;
- err = iavf_execute_vf_cmd_safe(adapter, &args, 0);
- if (err)
- PMD_DRV_LOG(ERR, "fail to execute command %s",
- add ? "OP_ADD_ETHER_ADDRESS" : "OP_DEL_ETHER_ADDRESS");
+void
+iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+{
+ int start = -1;
+ int i;
+
+ /* Handle primary address (index 0) separately */
+ if (!rte_is_zero_ether_addr(&adapter->dev_data->mac_addrs[0]))
+ iavf_add_del_eth_addr(adapter, &adapter->dev_data->mac_addrs[0], add,
+ VIRTCHNL_ETHER_ADDR_PRIMARY);
+
+ /* Process secondary addresses in contiguous blocks */
+ for (i = 1; i < IAVF_UC_MACADDR_MAX; i++) {
+ struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
+
+ if (!rte_is_zero_ether_addr(addr)) {
+ if (start == -1)
+ start = i;
+ continue;
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ start = -1;
+ }
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ }
}
int
@@ -2060,7 +2113,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
{
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
- (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
+ (IAVF_MC_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
struct virtchnl_ether_addr_list *list;
struct iavf_cmd_info args;
uint32_t i;
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v2 5/5] net/iavf: fix duplicate MAC addresses install
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
` (3 preceding siblings ...)
2026-05-06 12:35 ` [PATCH v2 4/5] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-05-06 12:35 ` David Marchand
2026-05-07 2:51 ` [PATCH v2 0/5] Remove limitations coming from legacy VMDq Stephen Hemminger
5 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-05-06 12:35 UTC (permalink / raw)
To: dev; +Cc: stable, Vladimir Medvedkin, Bruce Richardson
On port restart, all MAC addresses get pushed *twice* to the hardware,
once by the driver and once by the eth_dev_mac_restore() in ethdev.
On the other hand, MAC address filters are reset in the hardware
by the PF only when a VF reset is triggered.
Strictly speaking, the mac restore on port (re)start is unneeded,
if no VF reset happened, so we can announce to ethdev that no mac
restoration is needed via a get_restore_flags callback.
Then, move the mac restoration to the VF reset handler.
Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux driver")
Cc: stable@dpdk.org
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/intel/iavf/iavf_ethdev.c | 31 ++++++++++++++++++++--------
1 file changed, 22 insertions(+), 9 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index c69a012d50..5871ed3539 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -126,6 +126,8 @@ static void iavf_dev_del_mac_addr(struct rte_eth_dev *dev, uint32_t index);
static int iavf_dev_vlan_filter_set(struct rte_eth_dev *dev,
uint16_t vlan_id, int on);
static int iavf_dev_vlan_offload_set(struct rte_eth_dev *dev, int mask);
+static uint64_t iavf_get_restore_flags(struct rte_eth_dev *dev,
+ enum rte_eth_dev_operation op);
static int iavf_dev_rss_reta_update(struct rte_eth_dev *dev,
struct rte_eth_rss_reta_entry64 *reta_conf,
uint16_t reta_size);
@@ -249,6 +251,7 @@ static const struct eth_dev_ops iavf_eth_dev_ops = {
.tx_done_cleanup = iavf_dev_tx_done_cleanup,
.get_monitor_addr = iavf_get_monitor_addr,
.tm_ops_get = iavf_tm_ops_get,
+ .get_restore_flags = iavf_get_restore_flags,
};
static int
@@ -269,6 +272,13 @@ iavf_tm_ops_get(struct rte_eth_dev *dev,
return 0;
}
+static uint64_t
+iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
+ __rte_unused enum rte_eth_dev_operation op)
+{
+ return RTE_ETH_RESTORE_ALL & ~RTE_ETH_RESTORE_MAC_ADDR;
+}
+
__rte_unused
static int
iavf_vfr_inprogress(struct iavf_hw *hw)
@@ -1039,15 +1049,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
rte_intr_enable(intr_handle);
}
- /* Set all mac addrs */
- iavf_add_del_all_mac_addr(adapter, true);
-
- if (!adapter->mac_primary_set)
- adapter->mac_primary_set = true;
-
- /* Set all multicast addresses */
- iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num,
- true);
+ if (!adapter->mac_primary_set) {
+ if (iavf_add_del_eth_addr(adapter, &dev->data->mac_addrs[0], true,
+ VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
+ PMD_DRV_LOG(ERR, "failed to add primary MAC:" RTE_ETHER_ADDR_PRT_FMT,
+ RTE_ETHER_ADDR_BYTES(&dev->data->mac_addrs[0]));
+ else
+ adapter->mac_primary_set = true;
+ }
rte_spinlock_init(&vf->phc_time_aq_lock);
@@ -3144,6 +3153,10 @@ iavf_handle_hw_reset(struct rte_eth_dev *dev, bool vf_initiated_reset)
if (ret)
goto error;
+ /* after a VF reset, all mac addresses got flushed, restore them */
+ iavf_add_del_all_mac_addr(adapter, true);
+ iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num, true);
+
dev->data->dev_started = 1;
}
goto exit;
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH v2 0/5] Remove limitations coming from legacy VMDq
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
` (4 preceding siblings ...)
2026-05-06 12:35 ` [PATCH v2 5/5] net/iavf: fix duplicate MAC addresses install David Marchand
@ 2026-05-07 2:51 ` Stephen Hemminger
2026-05-10 15:03 ` David Marchand
5 siblings, 1 reply; 146+ messages in thread
From: Stephen Hemminger @ 2026-05-07 2:51 UTC (permalink / raw)
To: David Marchand; +Cc: dev
On Wed, 6 May 2026 14:35:48 +0200
David Marchand <david.marchand@redhat.com> wrote:
> Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
> VMDq has been imposing a maximum number of mac addresses in the
> mac_addr_add/del API.
>
> Nowadays, new Intel drivers do not support the feature and few other
> drivers implement this feature.
>
> This series proposes to flag drivers that support the feature, and
> remove the limit of number of mac addresses for others.
>
> Next step could be to remove the VMDq pool notion from the generic API.
> However I have some concern about this, as changing the quite stable
> mac_addr_add/del API now seems a lot of noise for not much benefit.
>
AI review (manual not automated) saw these:
On Wed, 6 May 2026 14:35:49 +0200
David Marchand <david.marchand@redhat.com> wrote:
[PATCH v2 1/5] ethdev: skip VMDq pools unless configured
--------------------------------------------------------
Warning: duplicate-add no longer returns 0 in non-VMDq mode
} else if (vmdq) {
pool_mask = dev->data->mac_pool_sel[index];
/* Check if both MAC address and pool is already there, and do nothing */
if (pool_mask & RTE_BIT64(pool))
return 0;
}
When index >= 0 (MAC already present) and VMDq is not configured,
the early "already there" return that existed before is now skipped
entirely. The driver's mac_addr_add op gets re-invoked and
mac_pool_sel is no longer updated for !vmdq, so subsequent calls
keep re-invoking it. Pre-patch behaviour was idempotent (return 0).
Suggested fix:
} else if (vmdq) {
pool_mask = dev->data->mac_pool_sel[index];
if (pool_mask & RTE_BIT64(pool))
return 0;
} else {
return 0; /* MAC already installed, only one pool */
}
[PATCH v2 2/5] ethdev: announce VMDq capability
-----------------------------------------------
Warning: capability advertised even when max_vmdq_pools may be 0
The capability bit is set unconditionally in several drivers where
VMDq availability is conditional on hardware variant or runtime
configuration:
- net/intel/e1000/igb_ethdev.c: max_vmdq_pools = 0 for e1000_82575,
e1000_i354 (no setting), e1000_i210, e1000_i211.
- net/intel/i40e/i40e_ethdev.c: max_vmdq_pools is gated on
(pf->flags & I40E_FLAG_VMDQ).
- net/bnxt/bnxt_ethdev.c: max_vmdq_pools is set to 0 when there
are not enough resources (see "Not enough resources to support
VMDq" path).
With this patch a user can pass mq_mode |= RTE_ETH_MQ_RX_VMDQ_FLAG
through rte_eth_dev_configure() because the new capability check
passes, but the driver has no pools to honour the request.
Suggested fix: gate the capa bit on max_vmdq_pools > 0, or set it
per-MAC-type / per-flag in each driver.
Warning: VMDq capability advertised on VFs that do not configure VMDq
ixgbevf_dev_info_get() and txgbevf_dev_info_get() now advertise
RTE_ETH_DEV_CAPA_VMDQ, but ixgbevf_dev_configure() and
txgbevf_dev_configure() only handle RTE_ETH_MQ_RX_RSS_FLAG. The
feature matrices doc/guides/nics/features/ixgbe_vf.ini and
txgbe_vf.ini do not list "VMDq = Y". A VF is itself a member of a
PF VMDq pool; advertising the capa from the VF is misleading.
Warning: ipn3ke advertises VMDq=Y in its features file but is not
updated by the patch
doc/guides/nics/features/ipn3ke.ini has "VMDq = Y" but the patch
does not add RTE_ETH_DEV_CAPA_VMDQ to ipn3ke. Either update the
driver or correct the feature matrix.
Warning: missing release note
The new RTE_ETH_DEV_CAPA_VMDQ public bit and the new rejection in
rte_eth_dev_configure() (returning -EINVAL when an application sets
a VMDq mq_mode on a non-VMDq device) are user-visible API and
behavioural changes. release_26_07.rst should mention them under
"API Changes".
[PATCH v2 3/5] ethdev: hide VMDq internal sizes
-----------------------------------------------
Warning: public defines removed without deprecation cycle
RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
were declared in the application-facing rte_ethdev.h. Moving them
to ethdev_driver.h removes them from the public API. Per the DPDK
API/ABI policy this normally needs a prior deprecation notice; at
minimum it warrants a release_26_07.rst "API Changes" entry.
[PATCH v2 4/5] net/iavf: accept up to 32k unicast MAC addresses
---------------------------------------------------------------
Warning: dead store in iavf_add_del_uc_addr_bulk()
for (uint32_t i = 0; i < nb_addrs; i++) {
size_t buf_len = sizeof(struct virtchnl_ether_addr_list) +
sizeof(struct virtchnl_ether_addr) * list->num_elements;
...
buf_len = sizeof(struct virtchnl_ether_addr_list) +
sizeof(struct virtchnl_ether_addr) * list->num_elements;
The first initialiser is unconditionally overwritten by the second
assignment before any read. Drop the initialiser (or move the
declaration to the use site).
Warning: missing release note for max_mac_addrs jump
max_mac_addrs goes from 64 to 32768 and dev_data->mac_addrs is now
allocated at RTE_ETHER_ADDR_LEN * 32768 = 192 KiB per VF port. This
is a notable behaviour change for iavf users and worth a
release_26_07.rst entry.
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v2 0/5] Remove limitations coming from legacy VMDq
2026-05-07 2:51 ` [PATCH v2 0/5] Remove limitations coming from legacy VMDq Stephen Hemminger
@ 2026-05-10 15:03 ` David Marchand
0 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-05-10 15:03 UTC (permalink / raw)
To: Stephen Hemminger; +Cc: dev
Hello Stephen,
On Thu, 7 May 2026 at 04:51, Stephen Hemminger
<stephen@networkplumber.org> wrote:
>
> On Wed, 6 May 2026 14:35:48 +0200
> David Marchand <david.marchand@redhat.com> wrote:
>
> > Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
> > VMDq has been imposing a maximum number of mac addresses in the
> > mac_addr_add/del API.
> >
> > Nowadays, new Intel drivers do not support the feature and few other
> > drivers implement this feature.
> >
> > This series proposes to flag drivers that support the feature, and
> > remove the limit of number of mac addresses for others.
> >
> > Next step could be to remove the VMDq pool notion from the generic API.
> > However I have some concern about this, as changing the quite stable
> > mac_addr_add/del API now seems a lot of noise for not much benefit.
> >
>
>
> AI review (manual not automated) saw these:
>
> On Wed, 6 May 2026 14:35:49 +0200
> David Marchand <david.marchand@redhat.com> wrote:
>
> [PATCH v2 1/5] ethdev: skip VMDq pools unless configured
> --------------------------------------------------------
>
> Warning: duplicate-add no longer returns 0 in non-VMDq mode
>
> } else if (vmdq) {
> pool_mask = dev->data->mac_pool_sel[index];
> /* Check if both MAC address and pool is already there, and do nothing */
> if (pool_mask & RTE_BIT64(pool))
> return 0;
> }
>
> When index >= 0 (MAC already present) and VMDq is not configured,
> the early "already there" return that existed before is now skipped
> entirely. The driver's mac_addr_add op gets re-invoked and
> mac_pool_sel is no longer updated for !vmdq, so subsequent calls
> keep re-invoking it. Pre-patch behaviour was idempotent (return 0).
>
> Suggested fix:
>
> } else if (vmdq) {
> pool_mask = dev->data->mac_pool_sel[index];
> if (pool_mask & RTE_BIT64(pool))
> return 0;
> } else {
> return 0; /* MAC already installed, only one pool */
> }
Yes, good catch.
>
>
> [PATCH v2 2/5] ethdev: announce VMDq capability
> -----------------------------------------------
>
> Warning: capability advertised even when max_vmdq_pools may be 0
>
> The capability bit is set unconditionally in several drivers where
> VMDq availability is conditional on hardware variant or runtime
> configuration:
>
> - net/intel/e1000/igb_ethdev.c: max_vmdq_pools = 0 for e1000_82575,
> e1000_i354 (no setting), e1000_i210, e1000_i211.
> - net/intel/i40e/i40e_ethdev.c: max_vmdq_pools is gated on
> (pf->flags & I40E_FLAG_VMDQ).
> - net/bnxt/bnxt_ethdev.c: max_vmdq_pools is set to 0 when there
> are not enough resources (see "Not enough resources to support
> VMDq" path).
>
> With this patch a user can pass mq_mode |= RTE_ETH_MQ_RX_VMDQ_FLAG
> through rte_eth_dev_configure() because the new capability check
> passes, but the driver has no pools to honour the request.
>
> Suggested fix: gate the capa bit on max_vmdq_pools > 0, or set it
> per-MAC-type / per-flag in each driver.
I realise now that this new device capability would also break VMDQ
with failsafe on top of VMDQ-capable ports.
In the end, as suggested here, max_vmdq_pools is already a gate and no
new capability is needed.
>
> Warning: VMDq capability advertised on VFs that do not configure VMDq
>
> ixgbevf_dev_info_get() and txgbevf_dev_info_get() now advertise
> RTE_ETH_DEV_CAPA_VMDQ, but ixgbevf_dev_configure() and
> txgbevf_dev_configure() only handle RTE_ETH_MQ_RX_RSS_FLAG. The
> feature matrices doc/guides/nics/features/ixgbe_vf.ini and
> txgbe_vf.ini do not list "VMDq = Y". A VF is itself a member of a
> PF VMDq pool; advertising the capa from the VF is misleading.
>
> Warning: ipn3ke advertises VMDq=Y in its features file but is not
> updated by the patch
>
> doc/guides/nics/features/ipn3ke.ini has "VMDq = Y" but the patch
> does not add RTE_ETH_DEV_CAPA_VMDQ to ipn3ke. Either update the
> driver or correct the feature matrix.
>
> Warning: missing release note
>
> The new RTE_ETH_DEV_CAPA_VMDQ public bit and the new rejection in
> rte_eth_dev_configure() (returning -EINVAL when an application sets
> a VMDq mq_mode on a non-VMDq device) are user-visible API and
> behavioural changes. release_26_07.rst should mention them under
> "API Changes".
>
> [PATCH v2 3/5] ethdev: hide VMDq internal sizes
> -----------------------------------------------
>
> Warning: public defines removed without deprecation cycle
>
> RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
> were declared in the application-facing rte_ethdev.h. Moving them
> to ethdev_driver.h removes them from the public API. Per the DPDK
> API/ABI policy this normally needs a prior deprecation notice; at
> minimum it warrants a release_26_07.rst "API Changes" entry.
Mixed feeling about it.
I can add a note, but I think it is noise.
>
>
> [PATCH v2 4/5] net/iavf: accept up to 32k unicast MAC addresses
> ---------------------------------------------------------------
>
> Warning: dead store in iavf_add_del_uc_addr_bulk()
>
> for (uint32_t i = 0; i < nb_addrs; i++) {
> size_t buf_len = sizeof(struct virtchnl_ether_addr_list) +
> sizeof(struct virtchnl_ether_addr) * list->num_elements;
> ...
> buf_len = sizeof(struct virtchnl_ether_addr_list) +
> sizeof(struct virtchnl_ether_addr) * list->num_elements;
>
> The first initialiser is unconditionally overwritten by the second
> assignment before any read. Drop the initialiser (or move the
> declaration to the use site).
Good catch.. rebase damage on my side, I dropped some
experiments/changes but forgot to update after.
>
> Warning: missing release note for max_mac_addrs jump
>
> max_mac_addrs goes from 64 to 32768 and dev_data->mac_addrs is now
> allocated at RTE_ETHER_ADDR_LEN * 32768 = 192 KiB per VF port. This
> is a notable behaviour change for iavf users and worth a
> release_26_07.rst entry.
>
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v3 0/5] Remove limitations coming from legacy VMDq
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (5 preceding siblings ...)
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
@ 2026-05-10 17:03 ` David Marchand
2026-05-10 17:03 ` [PATCH v3 1/5] ethdev: check VMDq availability David Marchand
` (4 more replies)
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
` (8 subsequent siblings)
15 siblings, 5 replies; 146+ messages in thread
From: David Marchand @ 2026-05-10 17:03 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain
Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
VMDq has been imposing a maximum number of mac addresses in the
mac_addr_add/del API.
Nowadays, new Intel drivers do not support the feature and few other
drivers implement this feature.
This series enforces that the driver announces VMDq pools before
using VMDq related features, then remove the limit of number of
mac addresses for others.
Next step could be to remove the VMDq pool notion from the generic API.
However I have some concern about this, as changing the quite stable
mac_addr_add/del API now seems a lot of noise for not much benefit.
--
David Marchand
Changes since v2:
- changed approach: did not introduce a new device capability,
relied on already existing dev_info->max_vmdq_pools,
- fixed duplicate mac addition without VMDq,
- updated documentation,
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
- fixed buffer overflow on mailbox messages during port restart/VF reset,
- fixed duplicate MAC address installation on port start/restart,
David Marchand (5):
ethdev: check VMDq availability
ethdev: skip VMDq pools unless configured
ethdev: hide VMDq internal sizes
net/iavf: accept up to 32k unicast MAC addresses
net/iavf: fix duplicate MAC addresses install
doc/guides/rel_notes/release_26_07.rst | 15 ++++
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 43 ++++++----
drivers/net/intel/iavf/iavf_vchnl.c | 113 ++++++++++++++++++-------
lib/ethdev/ethdev_driver.h | 8 +-
lib/ethdev/rte_ethdev.c | 68 +++++++++++----
lib/ethdev/rte_ethdev.h | 6 --
8 files changed, 186 insertions(+), 73 deletions(-)
--
2.53.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v3 1/5] ethdev: check VMDq availability
2026-05-10 17:03 ` [PATCH v3 " David Marchand
@ 2026-05-10 17:03 ` David Marchand
2026-06-01 9:38 ` Andrew Rybchenko
2026-05-10 17:03 ` [PATCH v3 2/5] ethdev: skip VMDq pools unless configured David Marchand
` (3 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-05-10 17:03 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Thomas Monjalon, Andrew Rybchenko
Refuse VMDq related Rx/Tx modes when the driver do not announce VMDq
pools availability.
This will used later as a gate to ignore/reject VMDq related matters.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v2:
- added an entry in release notes,
- changed approach: relied on pre-existing dev_info->max_vmdq_pools
rather than introduce a new device capability,
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
---
doc/guides/rel_notes/release_26_07.rst | 6 ++++++
lib/ethdev/rte_ethdev.c | 16 ++++++++++++++++
2 files changed, 22 insertions(+)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index f012d47a4b..b59793f177 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -92,6 +92,12 @@ API Changes
Also, make sure to start the actual text at the margin.
=======================================================
+* **Updated VMDq related API in ethdev.**
+
+ * At port configuration time, the number of VMDq pools advertised by a driver is now used to
+ validate VMDq related Rx and Tx modes (``RTE_ETH_MQ_RX_VMDQ_FLAG``, ``RTE_ETH_MQ_TX_VMDQ_DCB``,
+ ``RTE_ETH_MQ_TX_VMDQ_ONLY``).
+
ABI Changes
-----------
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 2edc7a362e..eae34954e2 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1581,6 +1581,22 @@ rte_eth_dev_configure(uint16_t port_id, uint16_t nb_rx_q, uint16_t nb_tx_q,
goto rollback;
}
+ if (dev_info.max_vmdq_pools == 0) {
+ if ((dev_conf->rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq rx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ if (dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_DCB ||
+ dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_ONLY) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq tx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ }
+
/*
* Setup new number of Rx/Tx queues and reconfigure device.
*/
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v3 2/5] ethdev: skip VMDq pools unless configured
2026-05-10 17:03 ` [PATCH v3 " David Marchand
2026-05-10 17:03 ` [PATCH v3 1/5] ethdev: check VMDq availability David Marchand
@ 2026-05-10 17:03 ` David Marchand
2026-06-01 9:38 ` Andrew Rybchenko
2026-05-10 17:03 ` [PATCH v3 3/5] ethdev: hide VMDq internal sizes David Marchand
` (2 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-05-10 17:03 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Nithin Dabilpuram, Kiran Kumar K,
Sunil Kumar Kori, Satha Rao, Harman Kalra, Thomas Monjalon,
Andrew Rybchenko
The mac_addr_add API describes that only the 0 pool should be passed
unless VMDq has been enabled, though there was no validation so far.
Add such a check, then cleanup the MAC related operations (adding,
removing, restoring).
As a side effect, the net/cnxk does not need to manually reset the
mac_pool_sel[] array.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v2:
- added an entry in release notes,
- fixed duplicate mac address handling for !vmdq,
- rewrote update of eth_dev_mac_restore to isolate the !vmdq case,
---
doc/guides/rel_notes/release_26_07.rst | 2 +
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
lib/ethdev/rte_ethdev.c | 52 ++++++++++++++++++--------
3 files changed, 38 insertions(+), 17 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index b59793f177..3d2e71102b 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -97,6 +97,8 @@ API Changes
* At port configuration time, the number of VMDq pools advertised by a driver is now used to
validate VMDq related Rx and Tx modes (``RTE_ETH_MQ_RX_VMDQ_FLAG``, ``RTE_ETH_MQ_TX_VMDQ_DCB``,
``RTE_ETH_MQ_TX_VMDQ_ONLY``).
+ * A check was added in ``rte_eth_dev_mac_addr_add`` to validate that the ``pool`` parameter is 0
+ when VMDq is not configured.
ABI Changes
diff --git a/drivers/net/cnxk/cnxk_ethdev_ops.c b/drivers/net/cnxk/cnxk_ethdev_ops.c
index 49e77e49a6..75decf7098 100644
--- a/drivers/net/cnxk/cnxk_ethdev_ops.c
+++ b/drivers/net/cnxk/cnxk_ethdev_ops.c
@@ -1240,7 +1240,6 @@ cnxk_nix_mc_addr_list_configure(struct rte_eth_dev *eth_dev, struct rte_ether_ad
/* Update address in NIC data structure */
rte_ether_addr_copy(&mc_addr_set[i], &data->mac_addrs[j]);
rte_ether_addr_copy(&mc_addr_set[i], &dev->dmac_addrs[j]);
- data->mac_pool_sel[j] = RTE_BIT64(0);
}
roc_nix_npc_promisc_ena_dis(nix, true);
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index eae34954e2..a628a7661a 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1677,7 +1677,7 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
{
struct rte_ether_addr *addr;
uint16_t i;
- uint32_t pool = 0;
+ uint32_t pool;
uint64_t pool_mask;
/* replay MAC address configuration including default MAC */
@@ -1685,9 +1685,11 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
if (dev->dev_ops->mac_addr_set != NULL)
dev->dev_ops->mac_addr_set(dev, addr);
else if (dev->dev_ops->mac_addr_add != NULL)
- dev->dev_ops->mac_addr_add(dev, addr, 0, pool);
+ dev->dev_ops->mac_addr_add(dev, addr, 0, 0);
if (dev->dev_ops->mac_addr_add != NULL) {
+ bool vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+
for (i = 1; i < dev_info->max_mac_addrs; i++) {
addr = &dev->data->mac_addrs[i];
@@ -1695,15 +1697,19 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
if (rte_is_zero_ether_addr(addr))
continue;
- pool = 0;
- pool_mask = dev->data->mac_pool_sel[i];
-
- do {
- if (pool_mask & UINT64_C(1))
- dev->dev_ops->mac_addr_add(dev, addr, i, pool);
- pool_mask >>= 1;
- pool++;
- } while (pool_mask);
+ if (!vmdq) {
+ dev->dev_ops->mac_addr_add(dev, addr, i, 0);
+ } else {
+ pool = 0;
+ pool_mask = dev->data->mac_pool_sel[i];
+
+ do {
+ if (pool_mask & UINT64_C(1))
+ dev->dev_ops->mac_addr_add(dev, addr, i, pool);
+ pool_mask >>= 1;
+ pool++;
+ } while (pool_mask);
+ }
}
}
}
@@ -5406,8 +5412,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
uint32_t pool)
{
struct rte_eth_dev *dev;
- int index;
uint64_t pool_mask;
+ bool vmdq;
+ int index;
int ret;
RTE_ETH_VALID_PORTID_OR_ERR_RET(port_id, -ENODEV);
@@ -5432,6 +5439,12 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
RTE_ETHDEV_LOG_LINE(ERR, "Pool ID must be 0-%d", RTE_ETH_64_POOLS - 1);
return -EINVAL;
}
+ vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+ if (!vmdq && pool != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Port %u: VMDq is not configured (pool %d)",
+ port_id, pool);
+ return -EINVAL;
+ }
index = eth_dev_get_mac_addr_index(port_id, addr);
if (index < 0) {
@@ -5442,6 +5455,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
return -ENOSPC;
}
} else {
+ if (!vmdq)
+ return 0;
+
pool_mask = dev->data->mac_pool_sel[index];
/* Check if both MAC address and pool is already there, and do nothing */
@@ -5456,8 +5472,10 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
/* Update address in NIC data structure */
rte_ether_addr_copy(addr, &dev->data->mac_addrs[index]);
- /* Update pool bitmap in NIC data structure */
- dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ if (vmdq) {
+ /* Update pool bitmap in NIC data structure */
+ dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ }
}
ret = eth_err(port_id, ret);
@@ -5502,8 +5520,10 @@ rte_eth_dev_mac_addr_remove(uint16_t port_id, struct rte_ether_addr *addr)
/* Update address in NIC data structure */
rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
- /* reset pool bitmap */
- dev->data->mac_pool_sel[index] = 0;
+ if ((dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ /* reset pool bitmap */
+ dev->data->mac_pool_sel[index] = 0;
+ }
rte_ethdev_trace_mac_addr_remove(port_id, addr);
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v3 3/5] ethdev: hide VMDq internal sizes
2026-05-10 17:03 ` [PATCH v3 " David Marchand
2026-05-10 17:03 ` [PATCH v3 1/5] ethdev: check VMDq availability David Marchand
2026-05-10 17:03 ` [PATCH v3 2/5] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-05-10 17:03 ` David Marchand
2026-06-01 9:39 ` Andrew Rybchenko
2026-05-10 17:03 ` [PATCH v3 4/5] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-05-10 17:03 ` [PATCH v3 5/5] net/iavf: fix duplicate MAC addresses install David Marchand
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-05-10 17:03 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Thomas Monjalon, Andrew Rybchenko
Hide RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
in the driver API as those (ambiguous) macros are only a driver concern.
In practice, this is only used by the bnxt and ixgbe (+ clones) drivers.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v2:
- added an entry in release notes,
---
doc/guides/rel_notes/release_26_07.rst | 3 +++
lib/ethdev/ethdev_driver.h | 8 +++++++-
lib/ethdev/rte_ethdev.h | 6 ------
3 files changed, 10 insertions(+), 7 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index 3d2e71102b..b425e6f8cd 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -99,6 +99,9 @@ API Changes
``RTE_ETH_MQ_TX_VMDQ_ONLY``).
* A check was added in ``rte_eth_dev_mac_addr_add`` to validate that the ``pool`` parameter is 0
when VMDq is not configured.
+ * The ``RTE_ETH_NUM_RECEIVE_MAC_ADDR`` and ``RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY`` macros are VMDq
+ related and are sizes of internal arrays in ethdev that only drivers need to care about.
+ Those macros are moved to the driver only ethdev API.
ABI Changes
diff --git a/lib/ethdev/ethdev_driver.h b/lib/ethdev/ethdev_driver.h
index 1255cd6f2c..a4e9cf5b90 100644
--- a/lib/ethdev/ethdev_driver.h
+++ b/lib/ethdev/ethdev_driver.h
@@ -119,6 +119,12 @@ struct __rte_cache_aligned rte_eth_dev {
struct rte_eth_dev_sriov;
struct rte_eth_dev_owner;
+/* Definitions used for receive MAC address */
+#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
+
+/* Definitions used for unicast hash */
+#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
+
/**
* @internal
* The data part, with no function pointers, associated with each Ethernet
@@ -153,7 +159,7 @@ struct __rte_cache_aligned rte_eth_dev_data {
* The first entry (index zero) is the default address.
*/
struct rte_ether_addr *mac_addrs;
- /** Bitmap associating MAC addresses to pools */
+ /** Bitmap associating MAC addresses to VMDq pools */
uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
/**
* Device Ethernet MAC addresses of hash filtering.
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index 0d8e2d0236..27d2ddc0c1 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -903,12 +903,6 @@ rte_eth_rss_hf_refine(uint64_t rss_hf)
#define RTE_ETH_VLAN_ID_MAX 0x0FFF /**< VLAN ID is in lower 12 bits*/
/**@}*/
-/* Definitions used for receive MAC address */
-#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
-
-/* Definitions used for unicast hash */
-#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
-
/**@{@name VMDq Rx mode
* @see rte_eth_vmdq_rx_conf.rx_mode
*/
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v3 4/5] net/iavf: accept up to 32k unicast MAC addresses
2026-05-10 17:03 ` [PATCH v3 " David Marchand
` (2 preceding siblings ...)
2026-05-10 17:03 ` [PATCH v3 3/5] ethdev: hide VMDq internal sizes David Marchand
@ 2026-05-10 17:03 ` David Marchand
2026-05-12 14:41 ` Stephen Hemminger
2026-05-10 17:03 ` [PATCH v3 5/5] net/iavf: fix duplicate MAC addresses install David Marchand
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-05-10 17:03 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Vladimir Medvedkin
E810 hardware provides 32k switch lookups.
Thanks to this, it is possible to allow a lot more secondary mac
addresses than what is possible today.
In practice, the maximum number of macs available per port may be lower
and depends on usage by other (trusted?) VFs on the same PF.
There is no way to figure out this limit but to try adding a mac address
and get an error from the PF driver.
Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
accordingly.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v2:
- added an entry in release notes,
- removed unneeded temp variable,
Changes since v1:
- fixed buffer overflow on mailbox messages during port restart/VF reset,
---
doc/guides/rel_notes/release_26_07.rst | 4 +
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 12 +--
drivers/net/intel/iavf/iavf_vchnl.c | 113 ++++++++++++++++++-------
4 files changed, 94 insertions(+), 40 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index b425e6f8cd..9d461f2837 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -63,6 +63,10 @@ New Features
``rte_eal_init`` and the application is responsible for probing each device,
* ``--auto-probing`` enables the initial bus probing, which is the current default behavior.
+* **Updated IAVF ethernet driver.**
+
+ * Increased the maximum number of secondary MAC addresses from 64 to 32k.
+ This increases a VF port memory footprint by ~192kB.
Removed Items
-------------
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 403c61e2e8..f1dede0694 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -31,7 +31,8 @@
#define IAVF_IRQ_MAP_NUM_PER_BUF 128
#define IAVF_RXTX_QUEUE_CHUNKS_NUM 2
-#define IAVF_NUM_MACADDR_MAX 64
+#define IAVF_UC_MACADDR_MAX 32768
+#define IAVF_MC_MACADDR_MAX 64
#define IAVF_DEV_WATCHDOG_PERIOD 2000 /* microseconds, set 0 to disable*/
@@ -253,7 +254,7 @@ struct iavf_info {
uint32_t link_speed;
/* Multicast addrs */
- struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
+ struct rte_ether_addr mc_addrs[IAVF_MC_MACADDR_MAX];
uint16_t mc_addrs_num; /* Multicast mac addresses number */
struct iavf_vsi vsi;
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index 1eca20bc9a..edbbc34cc5 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -379,10 +379,10 @@ iavf_set_mc_addr_list(struct rte_eth_dev *dev,
IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
int err, ret;
- if (mc_addrs_num > IAVF_NUM_MACADDR_MAX) {
+ if (mc_addrs_num > IAVF_MC_MACADDR_MAX) {
PMD_DRV_LOG(ERR,
"can't add more than a limited number (%u) of addresses.",
- (uint32_t)IAVF_NUM_MACADDR_MAX);
+ (uint32_t)IAVF_MC_MACADDR_MAX);
return -EINVAL;
}
@@ -1120,7 +1120,7 @@ iavf_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->hash_key_size = vf->vf_res->rss_key_size;
dev_info->reta_size = vf->vf_res->rss_lut_size;
dev_info->flow_type_rss_offloads = IAVF_RSS_OFFLOAD_ALL;
- dev_info->max_mac_addrs = IAVF_NUM_MACADDR_MAX;
+ dev_info->max_mac_addrs = IAVF_UC_MACADDR_MAX;
dev_info->dev_capa =
RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
@@ -2821,12 +2821,12 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
iavf_set_default_ptype_table(eth_dev);
/* copy mac addr */
- eth_dev->data->mac_addrs = rte_zmalloc(
- "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX, 0);
+ eth_dev->data->mac_addrs = rte_calloc("iavf_mac", IAVF_UC_MACADDR_MAX,
+ RTE_ETHER_ADDR_LEN, 0);
if (!eth_dev->data->mac_addrs) {
PMD_INIT_LOG(ERR, "Failed to allocate %d bytes needed to"
" store MAC addresses",
- RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX);
+ RTE_ETHER_ADDR_LEN * IAVF_UC_MACADDR_MAX);
ret = -ENOMEM;
goto init_vf_err;
}
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index 08dd6f2d7f..d0fc8dd54b 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1437,48 +1437,97 @@ iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num)
return 0;
}
-void
-iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+static int
+iavf_add_del_uc_addr_bulk(struct iavf_adapter *adapter, struct rte_ether_addr *addrs,
+ uint32_t nb_addrs, bool add)
{
+#define IAVF_ETH_ADDR_PER_REQ \
+ ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
+ sizeof(struct virtchnl_ether_addr))
struct {
struct virtchnl_ether_addr_list list;
- struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
- } list_req = {0};
- struct virtchnl_ether_addr_list *list = &list_req.list;
+ struct virtchnl_ether_addr addr[IAVF_ETH_ADDR_PER_REQ];
+ } cmd_buffer;
+#undef IAVF_ETH_ADDR_PER_REQ
+ struct virtchnl_ether_addr_list *list = &cmd_buffer.list;
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
- struct iavf_cmd_info args = {0};
- int err, i;
- size_t buf_len;
- for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
- struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
- struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
+ for (uint32_t i = 0; i < nb_addrs; i++) {
+ struct iavf_cmd_info args;
+ uint32_t batch;
+ int err;
- /* ignore empty addresses */
- if (rte_is_zero_ether_addr(addr))
- continue;
+ batch = i % RTE_DIM(cmd_buffer.addr);
+
+ if (batch == 0) {
+ memset(&cmd_buffer, 0, sizeof(cmd_buffer));
+ list->vsi_id = vf->vsi_res->vsi_id;
+ list->num_elements = 0;
+ }
+
+ rte_memcpy(list->list[batch].addr, addrs[i].addr_bytes,
+ sizeof(list->list[batch].addr));
+ list->list[batch].type = VIRTCHNL_ETHER_ADDR_EXTRA;
list->num_elements++;
- memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
- vc_addr->type = (list->num_elements == 1) ?
- VIRTCHNL_ETHER_ADDR_PRIMARY :
- VIRTCHNL_ETHER_ADDR_EXTRA;
+ if (batch != RTE_DIM(cmd_buffer.addr) - 1 && i != nb_addrs - 1)
+ continue;
+
+ memset(&args, 0, sizeof(args));
+ args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
+ args.in_args = (uint8_t *)list;
+ args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
+ sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ args.out_buffer = vf->aq_resp;
+ args.out_size = IAVF_AQ_BUF_SZ;
+ err = iavf_execute_vf_cmd_safe(adapter, &args, 0);
+ if (err != 0) {
+ PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
+ return err;
+ }
+
+ PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
}
- /* for some reason PF side checks for buffer being too big, so adjust it down */
- buf_len = sizeof(struct virtchnl_ether_addr_list) +
- sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ return 0;
+}
- list->vsi_id = vf->vsi_res->vsi_id;
- args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
- args.in_args = (uint8_t *)list;
- args.in_args_size = buf_len;
- args.out_buffer = vf->aq_resp;
- args.out_size = IAVF_AQ_BUF_SZ;
- err = iavf_execute_vf_cmd_safe(adapter, &args, 0);
- if (err)
- PMD_DRV_LOG(ERR, "fail to execute command %s",
- add ? "OP_ADD_ETHER_ADDRESS" : "OP_DEL_ETHER_ADDRESS");
+void
+iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+{
+ int start = -1;
+ int i;
+
+ /* Handle primary address (index 0) separately */
+ if (!rte_is_zero_ether_addr(&adapter->dev_data->mac_addrs[0]))
+ iavf_add_del_eth_addr(adapter, &adapter->dev_data->mac_addrs[0], add,
+ VIRTCHNL_ETHER_ADDR_PRIMARY);
+
+ /* Process secondary addresses in contiguous blocks */
+ for (i = 1; i < IAVF_UC_MACADDR_MAX; i++) {
+ struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
+
+ if (!rte_is_zero_ether_addr(addr)) {
+ if (start == -1)
+ start = i;
+ continue;
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ start = -1;
+ }
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ }
}
int
@@ -2060,7 +2109,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
{
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
- (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
+ (IAVF_MC_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
struct virtchnl_ether_addr_list *list;
struct iavf_cmd_info args;
uint32_t i;
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v3 5/5] net/iavf: fix duplicate MAC addresses install
2026-05-10 17:03 ` [PATCH v3 " David Marchand
` (3 preceding siblings ...)
2026-05-10 17:03 ` [PATCH v3 4/5] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-05-10 17:03 ` David Marchand
4 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-05-10 17:03 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, stable, Vladimir Medvedkin, Bruce Richardson
On port restart, all MAC addresses get pushed *twice* to the hardware,
once by the driver and once by the eth_dev_mac_restore() in ethdev.
On the other hand, MAC address filters are reset in the hardware
by the PF only when a VF reset is triggered.
Strictly speaking, the mac restore on port (re)start is unneeded,
if no VF reset happened, so we can announce to ethdev that no mac
restoration is needed via a get_restore_flags callback.
Then, move the mac restoration to the VF reset handler.
Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux driver")
Cc: stable@dpdk.org
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/intel/iavf/iavf_ethdev.c | 31 ++++++++++++++++++++--------
1 file changed, 22 insertions(+), 9 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index edbbc34cc5..504eabebe7 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -126,6 +126,8 @@ static void iavf_dev_del_mac_addr(struct rte_eth_dev *dev, uint32_t index);
static int iavf_dev_vlan_filter_set(struct rte_eth_dev *dev,
uint16_t vlan_id, int on);
static int iavf_dev_vlan_offload_set(struct rte_eth_dev *dev, int mask);
+static uint64_t iavf_get_restore_flags(struct rte_eth_dev *dev,
+ enum rte_eth_dev_operation op);
static int iavf_dev_rss_reta_update(struct rte_eth_dev *dev,
struct rte_eth_rss_reta_entry64 *reta_conf,
uint16_t reta_size);
@@ -249,6 +251,7 @@ static const struct eth_dev_ops iavf_eth_dev_ops = {
.tx_done_cleanup = iavf_dev_tx_done_cleanup,
.get_monitor_addr = iavf_get_monitor_addr,
.tm_ops_get = iavf_tm_ops_get,
+ .get_restore_flags = iavf_get_restore_flags,
};
static int
@@ -269,6 +272,13 @@ iavf_tm_ops_get(struct rte_eth_dev *dev,
return 0;
}
+static uint64_t
+iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
+ __rte_unused enum rte_eth_dev_operation op)
+{
+ return RTE_ETH_RESTORE_ALL & ~RTE_ETH_RESTORE_MAC_ADDR;
+}
+
__rte_unused
static int
iavf_vfr_inprogress(struct iavf_hw *hw)
@@ -1039,15 +1049,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
rte_intr_enable(intr_handle);
}
- /* Set all mac addrs */
- iavf_add_del_all_mac_addr(adapter, true);
-
- if (!adapter->mac_primary_set)
- adapter->mac_primary_set = true;
-
- /* Set all multicast addresses */
- iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num,
- true);
+ if (!adapter->mac_primary_set) {
+ if (iavf_add_del_eth_addr(adapter, &dev->data->mac_addrs[0], true,
+ VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
+ PMD_DRV_LOG(ERR, "failed to add primary MAC:" RTE_ETHER_ADDR_PRT_FMT,
+ RTE_ETHER_ADDR_BYTES(&dev->data->mac_addrs[0]));
+ else
+ adapter->mac_primary_set = true;
+ }
rte_spinlock_init(&vf->phc_time_aq_lock);
@@ -3144,6 +3153,10 @@ iavf_handle_hw_reset(struct rte_eth_dev *dev, bool vf_initiated_reset)
if (ret)
goto error;
+ /* after a VF reset, all mac addresses got flushed, restore them */
+ iavf_add_del_all_mac_addr(adapter, true);
+ iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num, true);
+
dev->data->dev_started = 1;
}
goto exit;
--
2.53.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH v3 4/5] net/iavf: accept up to 32k unicast MAC addresses
2026-05-10 17:03 ` [PATCH v3 4/5] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-05-12 14:41 ` Stephen Hemminger
2026-05-27 13:25 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Stephen Hemminger @ 2026-05-12 14:41 UTC (permalink / raw)
To: David Marchand; +Cc: dev, rjarry, cfontain, Vladimir Medvedkin
On Sun, 10 May 2026 19:03:04 +0200
David Marchand <david.marchand@redhat.com> wrote:
> E810 hardware provides 32k switch lookups.
> Thanks to this, it is possible to allow a lot more secondary mac
> addresses than what is possible today.
>
> In practice, the maximum number of macs available per port may be lower
> and depends on usage by other (trusted?) VFs on the same PF.
> There is no way to figure out this limit but to try adding a mac address
> and get an error from the PF driver.
>
> Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
> accordingly.
>
> Signed-off
AI review feedback; nothing major but worth a look
Series: ethdev: VMDq cleanup and iavf: 32k MACs / duplicate install fix
Tree: dpdk main (DPDK 26.07)
Patch 2 - ethdev: skip VMDq pools unless configured
Info
Doxygen of rte_eth_dev_mac_addr_add() is not updated to describe the
new constraint that pool must be 0 when VMDq is not configured. The
function-level @param documentation and @return list should mention
that -EINVAL is now returned when VMDq is disabled and pool != 0.
Subtle behavior change: when VMDq is not configured, mac_pool_sel[index]
is no longer set to RTE_BIT64(0) on add. Before this patch every added
MAC ended up with bit 0 set; afterwards the array stays at 0. This is
consistent with the new restore path (which now uses the VMDq flag
rather than walking pool_mask in the !vmdq case), but the change in
observable state of dev->data->mac_pool_sel[] could surprise an out-of-
tree consumer. Worth a sentence in the API change note.
Patch 4 - net/iavf: accept up to 32k unicast MAC addresses
Info
Release notes wording: "secondary MAC addresses from 64 to 32k" is
slightly imprecise. IAVF_NUM_MACADDR_MAX (64) was the total cap including
the primary MAC; IAVF_UC_MACADDR_MAX (32768) is also a total cap
including index 0. Either drop "secondary" or say "secondary MAC
addresses from 63 to 32767".
iavf_add_del_uc_addr_bulk() returns an int but the only caller
(iavf_add_del_all_mac_addr) discards it. The pre-existing function was
already void so this is no regression, but having the helper return an
error code that is unconditionally dropped is misleading. Either
propagate the error to the caller (and make iavf_add_del_all_mac_addr
return int) or make the helper void.
eth_dev_mac_restore() now iterates dev_info->max_mac_addrs which for
iavf becomes 32768. With patch 5 applied this path is skipped via
get_restore_flags, but any driver that increases max_mac_addrs into the
thousands without setting RTE_ETH_RESTORE_MAC_ADDR off will pay a 32k
rte_is_zero_ether_addr() scan on every dev_start. Not a bug here, but
worth keeping in mind.
sizeof(struct virtchnl_ether_addr_list) already accounts for the
embedded list[1] slot, so
in_args_size = sizeof(struct virtchnl_ether_addr_list) +
sizeof(struct virtchnl_ether_addr) * list->num_elements
overcounts by sizeof(struct virtchnl_ether_addr). This matches what
iavf_add_del_eth_addr() already does, and the buffer is sized to hold
it, so it's not a bug -- noting it for completeness.
Patch 5 - net/iavf: fix duplicate MAC addresses install
Info
Recovery edge case: if iavf_add_del_eth_addr() for the primary fails
during the first iavf_dev_start(), mac_primary_set stays false. On a
subsequent VF reset, iavf_dev_start() (called from iavf_handle_hw_reset)
will re-attempt the primary add, and the new
iavf_add_del_all_mac_addr() call right after will add the primary again
via VIRTCHNL_ETHER_ADDR_PRIMARY. This re-introduces a single-MAC
double-install in a narrow failure path. Could be avoided by skipping
the primary in iavf_add_del_all_mac_addr() when mac_primary_set is
already true, or by having iavf_dev_start() skip the primary add when
in_reset_recovery is set.
No release notes entry. With Fixes: + Cc: stable this is acceptable as
a bug fix, but the patch also introduces a new dev_op
(get_restore_flags) for the driver and changes when MC addresses are
re-pushed (only on VF reset, not on every dev_start). A one-line note
in the iavf section of the release notes would help users tracking
behaviour changes between stop/start cycles.
Patches 1 and 3 have no findings.
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH 2/4] ethdev: announce VMDq capability
2026-04-29 14:18 ` David Marchand
@ 2026-05-18 22:12 ` Kishore Padmanabha
0 siblings, 0 replies; 146+ messages in thread
From: Kishore Padmanabha @ 2026-05-18 22:12 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Ajit Khaparde, Bruce Richardson, Rosen Xu,
Anatoly Burakov, Vladimir Medvedkin, Jiawen Wu, Zaiyu Wang,
Thomas Monjalon, Andrew Rybchenko
[-- Attachment #1.1: Type: text/plain, Size: 1380 bytes --]
On Wed, Apr 29, 2026 at 10:18 AM David Marchand <david.marchand@redhat.com>
wrote:
> On Tue, 7 Apr 2026 at 00:22, Kishore Padmanabha
> <kishore.padmanabha@broadcom.com> wrote:
> >> diff --git a/drivers/net/bnxt/bnxt_ethdev.c
> b/drivers/net/bnxt/bnxt_ethdev.c
> >> index b677f9491d..0f783b9e98 100644
> >> --- a/drivers/net/bnxt/bnxt_ethdev.c
> >> +++ b/drivers/net/bnxt/bnxt_ethdev.c
> >> @@ -1214,7 +1214,8 @@ static int bnxt_dev_info_get_op(struct
> rte_eth_dev *eth_dev,
> >>
> >> dev_info->speed_capa = bnxt_get_speed_capabilities(bp);
> >> dev_info->dev_capa = RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
> >> - RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
> >> + RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP |
> >> + RTE_ETH_DEV_CAPA_VMDQ;
> >
> > We have not been testing the VMDq feature for sometime, planning to
> deprecate this feature. Please remove this change.
>
> Currently, the driver supports VMDq:
> $ git grep -i vmdq doc/guides/nics/features/bnxt.ini
> doc/guides/nics/features/bnxt.ini:VMDq = Y
>
> I don't intend to own this feature drop with my series :-).
> Please announce such a deprecation and flag it properly.
>
> I will post another patch to remove VMDq feaure from bnxt driver.
>
> --
> David Marchand
>
>
[-- Attachment #1.2: Type: text/html, Size: 2118 bytes --]
[-- Attachment #2: S/MIME Cryptographic Signature --]
[-- Type: application/pkcs7-signature, Size: 5493 bytes --]
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v3 4/5] net/iavf: accept up to 32k unicast MAC addresses
2026-05-12 14:41 ` Stephen Hemminger
@ 2026-05-27 13:25 ` David Marchand
0 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-05-27 13:25 UTC (permalink / raw)
To: Stephen Hemminger; +Cc: dev, rjarry, cfontain, Vladimir Medvedkin
On Tue, 12 May 2026 at 16:42, Stephen Hemminger
<stephen@networkplumber.org> wrote:
>
> On Sun, 10 May 2026 19:03:04 +0200
> David Marchand <david.marchand@redhat.com> wrote:
>
> > E810 hardware provides 32k switch lookups.
> > Thanks to this, it is possible to allow a lot more secondary mac
> > addresses than what is possible today.
> >
> > In practice, the maximum number of macs available per port may be lower
> > and depends on usage by other (trusted?) VFs on the same PF.
> > There is no way to figure out this limit but to try adding a mac address
> > and get an error from the PF driver.
> >
> > Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
> > accordingly.
> >
> > Signed-off
>
> AI review feedback; nothing major but worth a look
In the first revision, it was interesting, but now, that's a lot of noise.
>
> Series: ethdev: VMDq cleanup and iavf: 32k MACs / duplicate install fix
> Tree: dpdk main (DPDK 26.07)
>
> Patch 2 - ethdev: skip VMDq pools unless configured
> Info
> Doxygen of rte_eth_dev_mac_addr_add() is not updated to describe the
> new constraint that pool must be 0 when VMDq is not configured. The
> function-level @param documentation and @return list should mention
> that -EINVAL is now returned when VMDq is disabled and pool != 0.
* @param pool
* VMDq pool index to associate address with (if VMDq is enabled).
If VMDq is
* not enabled, this should be set to 0.
It has been documented since the start that a 0 pool should be passed.
I can update the return code EINVAL.
>
> Subtle behavior change: when VMDq is not configured, mac_pool_sel[index]
> is no longer set to RTE_BIT64(0) on add. Before this patch every added
> MAC ended up with bit 0 set; afterwards the array stays at 0. This is
> consistent with the new restore path (which now uses the VMDq flag
> rather than walking pool_mask in the !vmdq case), but the change in
> observable state of dev->data->mac_pool_sel[] could surprise an out-of-
> tree consumer. Worth a sentence in the API change note.
This is really internal stuff, only drivers would access this, and we
don't care about such out of tree consumer.
>
> Patch 4 - net/iavf: accept up to 32k unicast MAC addresses
> Info
> Release notes wording: "secondary MAC addresses from 64 to 32k" is
> slightly imprecise. IAVF_NUM_MACADDR_MAX (64) was the total cap including
> the primary MAC; IAVF_UC_MACADDR_MAX (32768) is also a total cap
> including index 0. Either drop "secondary" or say "secondary MAC
> addresses from 63 to 32767".
>
> iavf_add_del_uc_addr_bulk() returns an int but the only caller
> (iavf_add_del_all_mac_addr) discards it. The pre-existing function was
> already void so this is no regression, but having the helper return an
> error code that is unconditionally dropped is misleading. Either
> propagate the error to the caller (and make iavf_add_del_all_mac_addr
> return int) or make the helper void.
>
> eth_dev_mac_restore() now iterates dev_info->max_mac_addrs which for
> iavf becomes 32768. With patch 5 applied this path is skipped via
> get_restore_flags, but any driver that increases max_mac_addrs into the
> thousands without setting RTE_ETH_RESTORE_MAC_ADDR off will pay a 32k
> rte_is_zero_ether_addr() scan on every dev_start. Not a bug here, but
> worth keeping in mind.
>
> sizeof(struct virtchnl_ether_addr_list) already accounts for the
> embedded list[1] slot, so
> in_args_size = sizeof(struct virtchnl_ether_addr_list) +
> sizeof(struct virtchnl_ether_addr) * list->num_elements
> overcounts by sizeof(struct virtchnl_ether_addr). This matches what
> iavf_add_del_eth_addr() already does, and the buffer is sized to hold
> it, so it's not a bug -- noting it for completeness.
Those points do reflect the pre-existing implementations, but that's
it, I won't fix all bugs in the world... for now.
> Patch 5 - net/iavf: fix duplicate MAC addresses install
> Info
> Recovery edge case: if iavf_add_del_eth_addr() for the primary fails
> during the first iavf_dev_start(), mac_primary_set stays false. On a
> subsequent VF reset, iavf_dev_start() (called from iavf_handle_hw_reset)
> will re-attempt the primary add, and the new
> iavf_add_del_all_mac_addr() call right after will add the primary again
> via VIRTCHNL_ETHER_ADDR_PRIMARY. This re-introduces a single-MAC
> double-install in a narrow failure path. Could be avoided by skipping
> the primary in iavf_add_del_all_mac_addr() when mac_primary_set is
> already true, or by having iavf_dev_start() skip the primary add when
> in_reset_recovery is set.
>
> No release notes entry. With Fixes: + Cc: stable this is acceptable as
> a bug fix, but the patch also introduces a new dev_op
> (get_restore_flags) for the driver and changes when MC addresses are
> re-pushed (only on VF reset, not on every dev_start). A one-line note
> in the iavf section of the release notes would help users tracking
> behaviour changes between stop/start cycles.
Mac addresses are installed only once and restored on reset, that's all we need.
IOW, this patch should have no visible impact on an application, and
nothing to write about in the RN.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH 1/4] ethdev: skip VMDq pools unless configured
2026-04-03 9:18 ` [PATCH 1/4] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-06-01 9:30 ` Andrew Rybchenko
0 siblings, 0 replies; 146+ messages in thread
From: Andrew Rybchenko @ 2026-06-01 9:30 UTC (permalink / raw)
To: David Marchand, dev
Cc: rjarry, cfontain, Nithin Dabilpuram, Kiran Kumar K,
Sunil Kumar Kori, Satha Rao, Harman Kalra, Thomas Monjalon
On 4/3/26 12:18 PM, David Marchand wrote:
> The mac_addr_add API describes that only the 0 pool should be passed
> unless VMDq has been enabled, though there was no validation so far.
> Add such a check, then cleanup the related operations (adding, removing,
> restoring).
>
> As a side effect, the net/cnxk does not need to manually reset the
> mac_pool_sel[] array.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH 2/4] ethdev: announce VMDq capability
2026-04-03 9:18 ` [PATCH 2/4] ethdev: announce VMDq capability David Marchand
2026-04-06 22:22 ` Kishore Padmanabha
@ 2026-06-01 9:32 ` Andrew Rybchenko
1 sibling, 0 replies; 146+ messages in thread
From: Andrew Rybchenko @ 2026-06-01 9:32 UTC (permalink / raw)
To: David Marchand, dev
Cc: rjarry, cfontain, Kishore Padmanabha, Ajit Khaparde,
Bruce Richardson, Rosen Xu, Anatoly Burakov, Vladimir Medvedkin,
Jiawen Wu, Zaiyu Wang, Thomas Monjalon
On 4/3/26 12:18 PM, David Marchand wrote:
> Let's mark VMDq feature availability as a per device capability.
> We can then enforce API calls related to this feature are done on device
> with such capability.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
for ethdev
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH 3/4] ethdev: hide VMDq internal sizes
2026-04-03 9:18 ` [PATCH 3/4] ethdev: hide VMDq internal sizes David Marchand
@ 2026-06-01 9:34 ` Andrew Rybchenko
0 siblings, 0 replies; 146+ messages in thread
From: Andrew Rybchenko @ 2026-06-01 9:34 UTC (permalink / raw)
To: David Marchand, dev; +Cc: rjarry, cfontain, Thomas Monjalon
On 4/3/26 12:18 PM, David Marchand wrote:
> Hide RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
> in the driver API as those (ambiguous) macros are only a driver concern.
>
> In practice, this is only used by the bnxt and ixgbe (+ clones) drivers.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v2 1/5] ethdev: skip VMDq pools unless configured
2026-05-06 12:35 ` [PATCH v2 1/5] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-06-01 9:35 ` Andrew Rybchenko
0 siblings, 0 replies; 146+ messages in thread
From: Andrew Rybchenko @ 2026-06-01 9:35 UTC (permalink / raw)
To: David Marchand, dev
Cc: Nithin Dabilpuram, Kiran Kumar K, Sunil Kumar Kori, Satha Rao,
Harman Kalra, Thomas Monjalon
On 5/6/26 3:35 PM, David Marchand wrote:
> The mac_addr_add API describes that only the 0 pool should be passed
> unless VMDq has been enabled, though there was no validation so far.
> Add such a check, then cleanup the related operations (adding, removing,
> restoring).
>
> As a side effect, the net/cnxk does not need to manually reset the
> mac_pool_sel[] array.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko@oktetlabs.ru>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v2 2/5] ethdev: announce VMDq capability
2026-05-06 12:35 ` [PATCH v2 2/5] ethdev: announce VMDq capability David Marchand
@ 2026-06-01 9:36 ` Andrew Rybchenko
0 siblings, 0 replies; 146+ messages in thread
From: Andrew Rybchenko @ 2026-06-01 9:36 UTC (permalink / raw)
To: David Marchand, dev
Cc: Kishore Padmanabha, Ajit Khaparde, Bruce Richardson,
Anatoly Burakov, Vladimir Medvedkin, Jiawen Wu, Zaiyu Wang,
Thomas Monjalon
On 5/6/26 3:35 PM, David Marchand wrote:
> Let's mark VMDq feature availability as a per device capability.
> We can then enforce API calls related to this feature are done on device
> with such capability.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
for ethdev
Acked-by: Andrew Rybchenko@oktetlabs.ru>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v3 1/5] ethdev: check VMDq availability
2026-05-10 17:03 ` [PATCH v3 1/5] ethdev: check VMDq availability David Marchand
@ 2026-06-01 9:38 ` Andrew Rybchenko
0 siblings, 0 replies; 146+ messages in thread
From: Andrew Rybchenko @ 2026-06-01 9:38 UTC (permalink / raw)
To: David Marchand, dev; +Cc: rjarry, cfontain, Thomas Monjalon
On 5/10/26 8:03 PM, David Marchand wrote:
> Refuse VMDq related Rx/Tx modes when the driver do not announce VMDq
> pools availability.
> This will used later as a gate to ignore/reject VMDq related matters.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
sorry for spam, I found the latest version too late
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v3 2/5] ethdev: skip VMDq pools unless configured
2026-05-10 17:03 ` [PATCH v3 2/5] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-06-01 9:38 ` Andrew Rybchenko
0 siblings, 0 replies; 146+ messages in thread
From: Andrew Rybchenko @ 2026-06-01 9:38 UTC (permalink / raw)
To: David Marchand, dev
Cc: rjarry, cfontain, Nithin Dabilpuram, Kiran Kumar K,
Sunil Kumar Kori, Satha Rao, Harman Kalra, Thomas Monjalon
On 5/10/26 8:03 PM, David Marchand wrote:
> The mac_addr_add API describes that only the 0 pool should be passed
> unless VMDq has been enabled, though there was no validation so far.
> Add such a check, then cleanup the MAC related operations (adding,
> removing, restoring).
>
> As a side effect, the net/cnxk does not need to manually reset the
> mac_pool_sel[] array.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
for ethdev
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v3 3/5] ethdev: hide VMDq internal sizes
2026-05-10 17:03 ` [PATCH v3 3/5] ethdev: hide VMDq internal sizes David Marchand
@ 2026-06-01 9:39 ` Andrew Rybchenko
0 siblings, 0 replies; 146+ messages in thread
From: Andrew Rybchenko @ 2026-06-01 9:39 UTC (permalink / raw)
To: David Marchand, dev; +Cc: rjarry, cfontain, Thomas Monjalon
On 5/10/26 8:03 PM, David Marchand wrote:
> Hide RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
> in the driver API as those (ambiguous) macros are only a driver concern.
>
> In practice, this is only used by the bnxt and ixgbe (+ clones) drivers.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v4 00/10] Remove limitations coming from legacy VMDq
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (6 preceding siblings ...)
2026-05-10 17:03 ` [PATCH v3 " David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 01/10] ethdev: check VMDq availability David Marchand
` (9 more replies)
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (7 subsequent siblings)
15 siblings, 10 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain
Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
VMDq has been imposing a maximum number of mac addresses in the
mac_addr_add/del API.
Nowadays, new Intel drivers do not support the feature and few other
drivers implement this feature.
This series enforces that the driver announces VMDq pools before
using VMDq related features, then remove the limit of number of
mac addresses for others.
Next step could be to remove the VMDq pool notion from the generic API.
However I have some concern about this, as changing the quite stable
mac_addr_add/del API now seems a lot of noise for not much benefit.
--
David Marchand
Changes since v3:
- rebased (this series is too late for 26.07),
- update some doxygen comments,
- added mlx5 changes,
Changes since v2:
- changed approach: did not introduce a new device capability,
relied on already existing dev_info->max_vmdq_pools,
- fixed duplicate mac addition without VMDq,
- updated documentation,
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
- fixed buffer overflow on mailbox messages during port restart/VF reset,
- fixed duplicate MAC address installation on port start/restart,
David Marchand (10):
ethdev: check VMDq availability
ethdev: skip VMDq pools unless configured
ethdev: hide VMDq internal sizes
net/iavf: accept up to 32k unicast MAC addresses
net/iavf: fix duplicate MAC addresses install
net/mlx5: remove MAC addresses flush helper on Linux
net/mlx5: remove redundant MAC address index checks
net/mlx5: pass maximum number of unicast MAC to common code
net/mlx5: use bitset for tracking MAC addresses
net/mlx5: accept more unicast MAC addresses
doc/guides/rel_notes/release_26_07.rst | 16 ++++
drivers/common/mlx5/linux/mlx5_nl.c | 90 +++++---------------
drivers/common/mlx5/linux/mlx5_nl.h | 11 +--
drivers/common/mlx5/mlx5_common.h | 24 ------
drivers/common/mlx5/mlx5_devx_cmds.c | 4 +
drivers/common/mlx5/mlx5_devx_cmds.h | 2 +
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 43 ++++++----
drivers/net/intel/iavf/iavf_vchnl.c | 113 ++++++++++++++++++-------
drivers/net/mlx5/linux/mlx5_os.c | 61 +++++++++----
drivers/net/mlx5/mlx5.c | 9 +-
drivers/net/mlx5/mlx5.h | 17 +++-
drivers/net/mlx5/mlx5_ethdev.c | 2 +-
drivers/net/mlx5/mlx5_mac.c | 22 +++--
drivers/net/mlx5/mlx5_trigger.c | 12 +--
drivers/net/mlx5/windows/mlx5_os.c | 48 ++++++++---
lib/ethdev/ethdev_driver.h | 8 +-
lib/ethdev/rte_ethdev.c | 68 +++++++++++----
lib/ethdev/rte_ethdev.h | 8 +-
20 files changed, 340 insertions(+), 224 deletions(-)
--
2.54.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v4 01/10] ethdev: check VMDq availability
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 02/10] ethdev: skip VMDq pools unless configured David Marchand
` (8 subsequent siblings)
9 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Andrew Rybchenko, Thomas Monjalon
Refuse VMDq related Rx/Tx modes when the driver do not announce VMDq
pools availability.
This will used later as a gate to ignore/reject VMDq related matters.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v2:
- added an entry in release notes,
- changed approach: relied on pre-existing dev_info->max_vmdq_pools
rather than introduce a new device capability,
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
---
doc/guides/rel_notes/release_26_07.rst | 6 ++++++
lib/ethdev/rte_ethdev.c | 16 ++++++++++++++++
2 files changed, 22 insertions(+)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index 6352ef27ab..5e9178d36b 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -314,6 +314,12 @@ API Changes
- ``rte_flow_dynf_metadata_get``
- ``rte_flow_dynf_metadata_set``
+* **ethdev: updated VMDq related API.**
+
+ * At port configuration time, the number of VMDq pools advertised by a driver is now used to
+ validate VMDq related Rx and Tx modes (``RTE_ETH_MQ_RX_VMDQ_FLAG``, ``RTE_ETH_MQ_TX_VMDQ_DCB``,
+ ``RTE_ETH_MQ_TX_VMDQ_ONLY``).
+
* **mlx5: promoted driver event and steering management APIs from experimental to stable.**
The following mlx5 functions are no longer marked experimental:
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 9efeaf77cb..9e305d98c1 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1581,6 +1581,22 @@ rte_eth_dev_configure(uint16_t port_id, uint16_t nb_rx_q, uint16_t nb_tx_q,
goto rollback;
}
+ if (dev_info.max_vmdq_pools == 0) {
+ if ((dev_conf->rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq rx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ if (dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_DCB ||
+ dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_ONLY) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq tx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ }
+
/*
* Setup new number of Rx/Tx queues and reconfigure device.
*/
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 02/10] ethdev: skip VMDq pools unless configured
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
2026-07-09 16:02 ` [PATCH v4 01/10] ethdev: check VMDq availability David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 03/10] ethdev: hide VMDq internal sizes David Marchand
` (7 subsequent siblings)
9 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Andrew Rybchenko, Nithin Dabilpuram,
Kiran Kumar K, Sunil Kumar Kori, Satha Rao, Harman Kalra,
Thomas Monjalon
The mac_addr_add API describes that only the 0 pool should be passed
unless VMDq has been enabled, though there was no validation so far.
Add such a check, then cleanup the MAC related operations (adding,
removing, restoring).
As a side effect, the net/cnxk does not need to manually reset the
mac_pool_sel[] array.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v3:
- updated doxygen comments,
Changes since v2:
- added an entry in release notes,
- fixed duplicate mac address handling for !vmdq,
- rewrote update of eth_dev_mac_restore to isolate the !vmdq case,
---
doc/guides/rel_notes/release_26_07.rst | 2 +
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
lib/ethdev/rte_ethdev.c | 52 ++++++++++++++++++--------
lib/ethdev/rte_ethdev.h | 2 +-
4 files changed, 39 insertions(+), 18 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index 5e9178d36b..8e4e02e587 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -319,6 +319,8 @@ API Changes
* At port configuration time, the number of VMDq pools advertised by a driver is now used to
validate VMDq related Rx and Tx modes (``RTE_ETH_MQ_RX_VMDQ_FLAG``, ``RTE_ETH_MQ_TX_VMDQ_DCB``,
``RTE_ETH_MQ_TX_VMDQ_ONLY``).
+ * A check was added in ``rte_eth_dev_mac_addr_add`` to validate that the ``pool`` parameter is 0
+ when VMDq is not configured.
* **mlx5: promoted driver event and steering management APIs from experimental to stable.**
diff --git a/drivers/net/cnxk/cnxk_ethdev_ops.c b/drivers/net/cnxk/cnxk_ethdev_ops.c
index 0ea3d7e89f..c002a93fe1 100644
--- a/drivers/net/cnxk/cnxk_ethdev_ops.c
+++ b/drivers/net/cnxk/cnxk_ethdev_ops.c
@@ -1240,7 +1240,6 @@ cnxk_nix_mc_addr_list_configure(struct rte_eth_dev *eth_dev, struct rte_ether_ad
/* Update address in NIC data structure */
rte_ether_addr_copy(&mc_addr_set[i], &data->mac_addrs[j]);
rte_ether_addr_copy(&mc_addr_set[i], &dev->dmac_addrs[j]);
- data->mac_pool_sel[j] = RTE_BIT64(0);
}
roc_nix_npc_promisc_ena_dis(nix, true);
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 9e305d98c1..f20cc514a3 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1677,7 +1677,7 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
{
struct rte_ether_addr *addr;
uint16_t i;
- uint32_t pool = 0;
+ uint32_t pool;
uint64_t pool_mask;
/* replay MAC address configuration including default MAC */
@@ -1685,9 +1685,11 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
if (dev->dev_ops->mac_addr_set != NULL)
dev->dev_ops->mac_addr_set(dev, addr);
else if (dev->dev_ops->mac_addr_add != NULL)
- dev->dev_ops->mac_addr_add(dev, addr, 0, pool);
+ dev->dev_ops->mac_addr_add(dev, addr, 0, 0);
if (dev->dev_ops->mac_addr_add != NULL) {
+ bool vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+
for (i = 1; i < dev_info->max_mac_addrs; i++) {
addr = &dev->data->mac_addrs[i];
@@ -1695,15 +1697,19 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
if (rte_is_zero_ether_addr(addr))
continue;
- pool = 0;
- pool_mask = dev->data->mac_pool_sel[i];
-
- do {
- if (pool_mask & UINT64_C(1))
- dev->dev_ops->mac_addr_add(dev, addr, i, pool);
- pool_mask >>= 1;
- pool++;
- } while (pool_mask);
+ if (!vmdq) {
+ dev->dev_ops->mac_addr_add(dev, addr, i, 0);
+ } else {
+ pool = 0;
+ pool_mask = dev->data->mac_pool_sel[i];
+
+ do {
+ if (pool_mask & UINT64_C(1))
+ dev->dev_ops->mac_addr_add(dev, addr, i, pool);
+ pool_mask >>= 1;
+ pool++;
+ } while (pool_mask);
+ }
}
}
}
@@ -5414,8 +5420,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
uint32_t pool)
{
struct rte_eth_dev *dev;
- int index;
uint64_t pool_mask;
+ bool vmdq;
+ int index;
int ret;
RTE_ETH_VALID_PORTID_OR_ERR_RET(port_id, -ENODEV);
@@ -5440,6 +5447,12 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
RTE_ETHDEV_LOG_LINE(ERR, "Pool ID must be 0-%d", RTE_ETH_64_POOLS - 1);
return -EINVAL;
}
+ vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+ if (!vmdq && pool != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Port %u: VMDq is not configured (pool %d)",
+ port_id, pool);
+ return -EINVAL;
+ }
index = eth_dev_get_mac_addr_index(port_id, addr);
if (index < 0) {
@@ -5450,6 +5463,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
return -ENOSPC;
}
} else {
+ if (!vmdq)
+ return 0;
+
pool_mask = dev->data->mac_pool_sel[index];
/* Check if both MAC address and pool is already there, and do nothing */
@@ -5464,8 +5480,10 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
/* Update address in NIC data structure */
rte_ether_addr_copy(addr, &dev->data->mac_addrs[index]);
- /* Update pool bitmap in NIC data structure */
- dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ if (vmdq) {
+ /* Update pool bitmap in NIC data structure */
+ dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ }
}
ret = eth_err(port_id, ret);
@@ -5510,8 +5528,10 @@ rte_eth_dev_mac_addr_remove(uint16_t port_id, struct rte_ether_addr *addr)
/* Update address in NIC data structure */
rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
- /* reset pool bitmap */
- dev->data->mac_pool_sel[index] = 0;
+ if ((dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ /* reset pool bitmap */
+ dev->data->mac_pool_sel[index] = 0;
+ }
rte_ethdev_trace_mac_addr_remove(port_id, addr);
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index ee400b386f..e2a5ba1549 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -4636,7 +4636,7 @@ int rte_eth_dev_priority_flow_ctrl_set(uint16_t port_id,
* - (-ENODEV) if *port* is invalid.
* - (-EIO) if device is removed.
* - (-ENOSPC) if no more MAC addresses can be added.
- * - (-EINVAL) if MAC address is invalid.
+ * - (-EINVAL) if MAC address is invalid or a non 0 pool was passed but VMDq is not enabled.
*/
int rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *mac_addr,
uint32_t pool);
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 03/10] ethdev: hide VMDq internal sizes
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
2026-07-09 16:02 ` [PATCH v4 01/10] ethdev: check VMDq availability David Marchand
2026-07-09 16:02 ` [PATCH v4 02/10] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 04/10] net/iavf: accept up to 32k unicast MAC addresses David Marchand
` (6 subsequent siblings)
9 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Andrew Rybchenko, Thomas Monjalon
Hide RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
in the driver API as those (ambiguous) macros are only a driver concern.
In practice, this is only used by the bnxt and ixgbe (+ clones) drivers.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v2:
- added an entry in release notes,
---
doc/guides/rel_notes/release_26_07.rst | 3 +++
lib/ethdev/ethdev_driver.h | 8 +++++++-
lib/ethdev/rte_ethdev.h | 6 ------
3 files changed, 10 insertions(+), 7 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index 8e4e02e587..6b6fbe0b1a 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -321,6 +321,9 @@ API Changes
``RTE_ETH_MQ_TX_VMDQ_ONLY``).
* A check was added in ``rte_eth_dev_mac_addr_add`` to validate that the ``pool`` parameter is 0
when VMDq is not configured.
+ * The ``RTE_ETH_NUM_RECEIVE_MAC_ADDR`` and ``RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY`` macros are VMDq
+ related and are sizes of internal arrays in ethdev that only drivers need to care about.
+ Those macros are moved to the driver only ethdev API.
* **mlx5: promoted driver event and steering management APIs from experimental to stable.**
diff --git a/lib/ethdev/ethdev_driver.h b/lib/ethdev/ethdev_driver.h
index 0f336f9567..294f68504b 100644
--- a/lib/ethdev/ethdev_driver.h
+++ b/lib/ethdev/ethdev_driver.h
@@ -119,6 +119,12 @@ struct __rte_cache_aligned rte_eth_dev {
struct rte_eth_dev_sriov;
struct rte_eth_dev_owner;
+/* Definitions used for receive MAC address */
+#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
+
+/* Definitions used for unicast hash */
+#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
+
/**
* @internal
* The data part, with no function pointers, associated with each Ethernet
@@ -153,7 +159,7 @@ struct __rte_cache_aligned rte_eth_dev_data {
* The first entry (index zero) is the default address.
*/
struct rte_ether_addr *mac_addrs;
- /** Bitmap associating MAC addresses to pools */
+ /** Bitmap associating MAC addresses to VMDq pools */
uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
/**
* Device Ethernet MAC addresses of hash filtering.
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index e2a5ba1549..bee47549cb 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -903,12 +903,6 @@ rte_eth_rss_hf_refine(uint64_t rss_hf)
#define RTE_ETH_VLAN_ID_MAX 0x0FFF /**< VLAN ID is in lower 12 bits*/
/**@}*/
-/* Definitions used for receive MAC address */
-#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
-
-/* Definitions used for unicast hash */
-#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
-
/**@{@name VMDq Rx mode
* @see rte_eth_vmdq_rx_conf.rx_mode
*/
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 04/10] net/iavf: accept up to 32k unicast MAC addresses
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
` (2 preceding siblings ...)
2026-07-09 16:02 ` [PATCH v4 03/10] ethdev: hide VMDq internal sizes David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install David Marchand
` (5 subsequent siblings)
9 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Vladimir Medvedkin
E810 hardware provides 32k switch lookups.
Thanks to this, it is possible to allow a lot more secondary mac
addresses than what is possible today.
In practice, the maximum number of macs available per port may be lower
and depends on usage by other (trusted?) VFs on the same PF.
There is no way to figure out this limit but to try adding a mac address
and get an error from the PF driver.
Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
accordingly.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v2:
- added an entry in release notes,
- removed unneeded temp variable,
Changes since v1:
- fixed buffer overflow on mailbox messages during port restart/VF reset,
---
doc/guides/rel_notes/release_26_07.rst | 5 ++
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 12 +--
drivers/net/intel/iavf/iavf_vchnl.c | 113 ++++++++++++++++++-------
4 files changed, 95 insertions(+), 40 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index 6b6fbe0b1a..4225bd43a0 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -259,6 +259,11 @@ New Features
Added AGENTS.md file for AI review
and supporting scripts to review patches and documentation.
+* **Updated IAVF ethernet driver.**
+
+ * Increased the maximum number of secondary unicast MAC addresses from 64 to 32k.
+ This increases a VF port memory footprint by ~192kB.
+
Removed Items
-------------
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 293adaf6c9..47cd1d6311 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -31,7 +31,8 @@
#define IAVF_IRQ_MAP_NUM_PER_BUF 128
#define IAVF_RXTX_QUEUE_CHUNKS_NUM 2
-#define IAVF_NUM_MACADDR_MAX 64
+#define IAVF_UC_MACADDR_MAX 32768
+#define IAVF_MC_MACADDR_MAX 64
#define IAVF_DEV_WATCHDOG_PERIOD 2000 /* microseconds, set 0 to disable*/
@@ -255,7 +256,7 @@ struct iavf_info {
uint32_t link_speed;
/* Multicast addrs */
- struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
+ struct rte_ether_addr mc_addrs[IAVF_MC_MACADDR_MAX];
uint16_t mc_addrs_num; /* Multicast mac addresses number */
struct iavf_vsi vsi;
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index 80e740ef29..179d90ec55 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -394,10 +394,10 @@ iavf_set_mc_addr_list(struct rte_eth_dev *dev,
IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
int err, ret;
- if (mc_addrs_num > IAVF_NUM_MACADDR_MAX) {
+ if (mc_addrs_num > IAVF_MC_MACADDR_MAX) {
PMD_DRV_LOG(ERR,
"can't add more than a limited number (%u) of addresses.",
- (uint32_t)IAVF_NUM_MACADDR_MAX);
+ (uint32_t)IAVF_MC_MACADDR_MAX);
return -EINVAL;
}
@@ -1159,7 +1159,7 @@ iavf_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->hash_key_size = vf->vf_res->rss_key_size;
dev_info->reta_size = vf->vf_res->rss_lut_size;
dev_info->flow_type_rss_offloads = IAVF_RSS_OFFLOAD_ALL;
- dev_info->max_mac_addrs = IAVF_NUM_MACADDR_MAX;
+ dev_info->max_mac_addrs = IAVF_UC_MACADDR_MAX;
dev_info->dev_capa =
RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
@@ -3051,12 +3051,12 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
iavf_set_default_ptype_table(eth_dev);
/* copy mac addr */
- eth_dev->data->mac_addrs = rte_zmalloc(
- "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX, 0);
+ eth_dev->data->mac_addrs = rte_calloc("iavf_mac", IAVF_UC_MACADDR_MAX,
+ RTE_ETHER_ADDR_LEN, 0);
if (!eth_dev->data->mac_addrs) {
PMD_INIT_LOG(ERR, "Failed to allocate %d bytes needed to"
" store MAC addresses",
- RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX);
+ RTE_ETHER_ADDR_LEN * IAVF_UC_MACADDR_MAX);
ret = -ENOMEM;
goto init_vf_err;
}
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index cd6e8325fc..6743183bf9 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1648,49 +1648,98 @@ iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num)
return 0;
}
-void
-iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+static int
+iavf_add_del_uc_addr_bulk(struct iavf_adapter *adapter, struct rte_ether_addr *addrs,
+ uint32_t nb_addrs, bool add)
{
+#define IAVF_ETH_ADDR_PER_REQ \
+ ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
+ sizeof(struct virtchnl_ether_addr))
struct {
struct virtchnl_ether_addr_list list;
- struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
- } list_req = {0};
- struct virtchnl_ether_addr_list *list = &list_req.list;
+ struct virtchnl_ether_addr addr[IAVF_ETH_ADDR_PER_REQ];
+ } cmd_buffer;
+#undef IAVF_ETH_ADDR_PER_REQ
+ struct virtchnl_ether_addr_list *list = &cmd_buffer.list;
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
- struct iavf_cmd_info args = {0};
- int err, i;
- size_t buf_len;
- for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
- struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
- struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
+ for (uint32_t i = 0; i < nb_addrs; i++) {
+ struct iavf_cmd_info args;
+ uint32_t batch;
+ int err;
- /* ignore empty addresses */
- if (rte_is_zero_ether_addr(addr))
- continue;
+ batch = i % RTE_DIM(cmd_buffer.addr);
+
+ if (batch == 0) {
+ memset(&cmd_buffer, 0, sizeof(cmd_buffer));
+ list->vsi_id = vf->vsi_res->vsi_id;
+ list->num_elements = 0;
+ }
+
+ rte_memcpy(list->list[batch].addr, addrs[i].addr_bytes,
+ sizeof(list->list[batch].addr));
+ list->list[batch].type = VIRTCHNL_ETHER_ADDR_EXTRA;
list->num_elements++;
- memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
- vc_addr->type = (list->num_elements == 1) ?
- VIRTCHNL_ETHER_ADDR_PRIMARY :
- VIRTCHNL_ETHER_ADDR_EXTRA;
+ if (batch != RTE_DIM(cmd_buffer.addr) - 1 && i != nb_addrs - 1)
+ continue;
+
+ memset(&args, 0, sizeof(args));
+ args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
+ args.in_args = (uint8_t *)list;
+ args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
+ sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ args.out_buffer = msg_buf;
+ args.out_size = IAVF_AQ_BUF_SZ;
+ err = iavf_execute_vf_cmd_safe(adapter, &args);
+ if (err != 0) {
+ PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
+ return err;
+ }
+
+ PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
}
- /* for some reason PF side checks for buffer being too big, so adjust it down */
- buf_len = sizeof(struct virtchnl_ether_addr_list) +
- sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ return 0;
+}
- list->vsi_id = vf->vsi_res->vsi_id;
- args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
- args.in_args = (uint8_t *)list;
- args.in_args_size = buf_len;
- args.out_buffer = msg_buf;
- args.out_size = IAVF_AQ_BUF_SZ;
- err = iavf_execute_vf_cmd_safe(adapter, &args);
- if (err)
- PMD_DRV_LOG(ERR, "fail to execute command %s",
- add ? "OP_ADD_ETHER_ADDRESS" : "OP_DEL_ETHER_ADDRESS");
+void
+iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+{
+ int start = -1;
+ int i;
+
+ /* Handle primary address (index 0) separately */
+ if (!rte_is_zero_ether_addr(&adapter->dev_data->mac_addrs[0]))
+ iavf_add_del_eth_addr(adapter, &adapter->dev_data->mac_addrs[0], add,
+ VIRTCHNL_ETHER_ADDR_PRIMARY);
+
+ /* Process secondary addresses in contiguous blocks */
+ for (i = 1; i < IAVF_UC_MACADDR_MAX; i++) {
+ struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
+
+ if (!rte_is_zero_ether_addr(addr)) {
+ if (start == -1)
+ start = i;
+ continue;
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ start = -1;
+ }
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ }
}
int
@@ -2281,7 +2330,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
- (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
+ (IAVF_MC_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
struct virtchnl_ether_addr_list *list;
struct iavf_cmd_info args;
uint32_t i;
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
` (3 preceding siblings ...)
2026-07-09 16:02 ` [PATCH v4 04/10] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-13 13:12 ` Loftus, Ciara
2026-07-09 16:02 ` [PATCH v4 06/10] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (4 subsequent siblings)
9 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, stable, Vladimir Medvedkin, Bruce Richardson
On port restart, all MAC addresses get pushed *twice* to the hardware,
once by the driver and once by the eth_dev_mac_restore() in ethdev.
On the other hand, MAC address filters are reset in the hardware
by the PF only when a VF reset is triggered.
Strictly speaking, the mac restore on port (re)start is unneeded,
if no VF reset happened, so we can announce to ethdev that no mac
restoration is needed via a get_restore_flags callback.
Then, move the mac restoration to the VF reset handler.
Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux driver")
Cc: stable@dpdk.org
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/intel/iavf/iavf_ethdev.c | 31 ++++++++++++++++++++--------
1 file changed, 22 insertions(+), 9 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index 179d90ec55..fe54df4b9f 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -134,6 +134,8 @@ static int iavf_dev_vlan_filter_set(struct rte_eth_dev *dev,
static int iavf_vlan_tpid_set(struct rte_eth_dev *dev,
enum rte_vlan_type vlan_type, uint16_t tpid);
static int iavf_dev_vlan_offload_set(struct rte_eth_dev *dev, int mask);
+static uint64_t iavf_get_restore_flags(struct rte_eth_dev *dev,
+ enum rte_eth_dev_operation op);
static int iavf_dev_rss_reta_update(struct rte_eth_dev *dev,
struct rte_eth_rss_reta_entry64 *reta_conf,
uint16_t reta_size);
@@ -264,6 +266,7 @@ static const struct eth_dev_ops iavf_eth_dev_ops = {
.tx_done_cleanup = iavf_dev_tx_done_cleanup,
.get_monitor_addr = iavf_get_monitor_addr,
.tm_ops_get = iavf_tm_ops_get,
+ .get_restore_flags = iavf_get_restore_flags,
};
static int
@@ -284,6 +287,13 @@ iavf_tm_ops_get(struct rte_eth_dev *dev,
return 0;
}
+static uint64_t
+iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
+ __rte_unused enum rte_eth_dev_operation op)
+{
+ return RTE_ETH_RESTORE_ALL & ~RTE_ETH_RESTORE_MAC_ADDR;
+}
+
__rte_unused
static int
iavf_vfr_inprogress(struct iavf_hw *hw)
@@ -1074,15 +1084,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
rte_intr_enable(intr_handle);
}
- /* Set all mac addrs */
- iavf_add_del_all_mac_addr(adapter, true);
-
- if (!adapter->mac_primary_set)
- adapter->mac_primary_set = true;
-
- /* Set all multicast addresses */
- iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num,
- true);
+ if (!adapter->mac_primary_set) {
+ if (iavf_add_del_eth_addr(adapter, &dev->data->mac_addrs[0], true,
+ VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
+ PMD_DRV_LOG(ERR, "failed to add primary MAC:" RTE_ETHER_ADDR_PRT_FMT,
+ RTE_ETHER_ADDR_BYTES(&dev->data->mac_addrs[0]));
+ else
+ adapter->mac_primary_set = true;
+ }
rte_spinlock_init(&vf->phc_time_aq_lock);
@@ -3424,6 +3433,10 @@ iavf_handle_hw_reset(struct rte_eth_dev *dev, bool vf_initiated_reset)
if (ret)
goto error;
+ /* after a VF reset, all mac addresses got flushed, restore them */
+ iavf_add_del_all_mac_addr(adapter, true);
+ iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num, true);
+
dev->data->dev_started = 1;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 06/10] net/mlx5: remove MAC addresses flush helper on Linux
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
` (4 preceding siblings ...)
2026-07-09 16:02 ` [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 07/10] net/mlx5: remove redundant MAC address index checks David Marchand
` (3 subsequent siblings)
9 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
This helper is exposing internals of net/mlx5 for no good reason.
All this code does is calling the remove helper.
Walk through the list in Linux implementation like the Windows
implementation.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 40 -----------------------------
drivers/common/mlx5/linux/mlx5_nl.h | 5 ----
drivers/net/mlx5/linux/mlx5_os.c | 13 +++++++---
3 files changed, 10 insertions(+), 48 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 8b19838a7e..85736738ac 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -825,46 +825,6 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
}
-/**
- * Flush all added MAC addresses.
- *
- * @param[in] nlsk_fd
- * Netlink socket file descriptor.
- * @param[in] iface_idx
- * Net device interface index.
- * @param[in] mac_addrs
- * Mac addresses array to flush.
- * @param n
- * @p mac_addrs array size.
- * @param mac_own
- * BITFIELD_DECLARE array to store the mac.
- * @param vf
- * Flag for a VF device.
- */
-RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_flush)
-void
-mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own, bool vf)
-{
- int i;
-
- if (n <= 0 || n > MLX5_MAX_MAC_ADDRESSES)
- return;
-
- for (i = n - 1; i >= 0; --i) {
- struct rte_ether_addr *m = &mac_addrs[i];
-
- if (BITFIELD_ISSET(mac_own, i)) {
- if (vf)
- mlx5_nl_mac_addr_remove(nlsk_fd,
- iface_idx,
- m, i);
- BITFIELD_RESET(mac_own, i);
- }
- }
-}
-
/**
* Enable promiscuous / all multicast mode through Netlink.
*
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 3f79a73c85..256ed7e2b7 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -65,11 +65,6 @@ __rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
__rte_internal
-void mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own,
- bool vf);
-__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
int mlx5_nl_allmulti(int nlsk_fd, unsigned int iface_idx, int enable);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index adc5878296..9fd366b10a 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3495,10 +3495,17 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
{
struct mlx5_priv *priv = dev->data->dev_private;
const int vf = priv->sh->dev_cap.vf;
+ int i;
- mlx5_nl_mac_addr_flush(priv->nl_socket_route, mlx5_ifindex(dev),
- dev->data->mac_addrs,
- MLX5_MAX_MAC_ADDRESSES, priv->mac_own, vf);
+ for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (vf)
+ mlx5_nl_mac_addr_remove(priv->nl_socket_route,
+ mlx5_ifindex(dev),
+ &dev->data->mac_addrs[i], i);
+ BITFIELD_RESET(priv->mac_own, i);
+ }
+ }
}
static bool
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 07/10] net/mlx5: remove redundant MAC address index checks
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
` (5 preceding siblings ...)
2026-07-09 16:02 ` [PATCH v4 06/10] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 08/10] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
` (2 subsequent siblings)
9 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
On the net/mlx5 side, the OS specific helpers do not need to implement
again checks on the MAC index, the cde from mlx5_mac.c already does this.
Cascading this consideration, validating the MAC index against
MLX5_MAX_MAC_ADDRESSES in common code is also redundant.
The common code only deals with netlink, remove any index concern.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 17 ++---------------
drivers/common/mlx5/linux/mlx5_nl.h | 4 ++--
drivers/net/mlx5/linux/mlx5_os.c | 9 ++++-----
3 files changed, 8 insertions(+), 22 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 85736738ac..ceb504f84d 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -719,8 +719,6 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to register.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -728,16 +726,11 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_add)
int
mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
int ret;
ret = mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 1);
- if (!ret) {
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
- }
if (ret == -EEXIST)
return 0;
return ret;
@@ -752,8 +745,6 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to remove.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -761,12 +752,8 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_remove)
int
mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
-
return mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 0);
}
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 256ed7e2b7..500198b654 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -57,10 +57,10 @@ __rte_internal
int mlx5_nl_init(int protocol, int groups);
__rte_internal
int mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 9fd366b10a..1b241ba9d2 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3382,9 +3382,8 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[index], index);
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ &dev->data->mac_addrs[index]);
+ BITFIELD_RESET(priv->mac_own, index);
}
/**
@@ -3411,7 +3410,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
if (vf)
ret = mlx5_nl_mac_addr_add(priv->nl_socket_route,
mlx5_ifindex(dev),
- mac, index);
+ mac);
if (!ret)
BITFIELD_SET(priv->mac_own, index);
@@ -3502,7 +3501,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[i], i);
+ &dev->data->mac_addrs[i]);
BITFIELD_RESET(priv->mac_own, i);
}
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 08/10] net/mlx5: pass maximum number of unicast MAC to common code
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
` (6 preceding siblings ...)
2026-07-09 16:02 ` [PATCH v4 07/10] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 09/10] net/mlx5: use bitset for tracking MAC addresses David Marchand
2026-07-09 16:02 ` [PATCH v4 10/10] net/mlx5: accept more unicast " David Marchand
9 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Isolate how the MAC addresses array is walked through in the common code
by passing the max index at which a unicast MAC address is stored in
dev->data->mac_addrs[].
In the sync callback, the size of the array allocated on the stack is
known by the caller, treat the mac_n field as an input parameter too.
With this change, only net/mlx5 knows about the max number of
unicast/multicast MAC addresses.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 33 ++++++++++++++++++-----------
drivers/common/mlx5/linux/mlx5_nl.h | 2 +-
drivers/common/mlx5/mlx5_common.h | 8 -------
drivers/net/mlx5/linux/mlx5_os.c | 1 +
drivers/net/mlx5/mlx5.h | 8 +++++++
5 files changed, 31 insertions(+), 21 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index ceb504f84d..152b4cdda1 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -468,19 +468,22 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
struct mlx5_nl_mac_addr *data = arg;
struct ndmsg *r = NLMSG_DATA(nh);
struct rtattr *attribute;
+ int mac_n = 0;
int len;
+ int ret;
len = nh->nlmsg_len - NLMSG_LENGTH(sizeof(*r));
for (attribute = MLX5_NDA_RTA(r);
RTA_OK(attribute, len);
attribute = RTA_NEXT(attribute, len)) {
if (attribute->rta_type == NDA_LLADDR) {
- if (data->mac_n == MLX5_MAX_MAC_ADDRESSES) {
+ if (mac_n == data->mac_n) {
DRV_LOG(WARNING,
"not enough room to finalize the"
" request");
rte_errno = ENOMEM;
- return -rte_errno;
+ ret = -rte_errno;
+ goto out;
}
#ifdef RTE_PMD_MLX5_DEBUG
char m[RTE_ETHER_ADDR_FMT_SIZE];
@@ -489,11 +492,15 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
RTA_DATA(attribute));
DRV_LOG(DEBUG, "bridge MAC address %s", m);
#endif
- memcpy(&(*data->mac)[data->mac_n++],
+ memcpy(&(*data->mac)[mac_n++],
RTA_DATA(attribute), RTE_ETHER_ADDR_LEN);
}
}
- return 0;
+ ret = 0;
+
+out:
+ data->mac_n = mac_n;
+ return ret;
}
/**
@@ -505,9 +512,9 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
* Net device interface index.
* @param mac[out]
* Pointer to the array table of MAC addresses to fill.
- * Its size should be of MLX5_MAX_MAC_ADDRESSES.
- * @param mac_n[out]
- * Number of entries filled in MAC array.
+ * @param mac_n[in,out]
+ * Size of the MAC array on input.
+ * Number of entries filled in MAC array on output.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -532,7 +539,7 @@ mlx5_nl_mac_addr_list(int nlsk_fd, unsigned int iface_idx,
};
struct mlx5_nl_mac_addr data = {
.mac = mac,
- .mac_n = 0,
+ .mac_n = *mac_n,
};
uint32_t sn = MLX5_NL_SN_GENERATE;
int ret;
@@ -766,16 +773,18 @@ mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac_addrs
* Mac addresses array to sync.
+ * @param uc_n
+ * Number of UC entries in @p mac_addrs.
* @param n
* @p mac_addrs array size.
*/
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_sync)
void
mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n)
+ struct rte_ether_addr *mac_addrs, int uc_n, int n)
{
struct rte_ether_addr macs[n];
- int macs_n = 0;
+ int macs_n = n;
int i;
int ret;
@@ -794,7 +803,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
continue;
if (rte_is_multicast_ether_addr(&macs[i])) {
/* Find the first entry available. */
- for (j = MLX5_MAX_UC_MAC_ADDRESSES; j != n; ++j) {
+ for (j = uc_n; j != n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
@@ -802,7 +811,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
} else {
/* Find the first entry available. */
- for (j = 0; j != MLX5_MAX_UC_MAC_ADDRESSES; ++j) {
+ for (j = 0; j != uc_n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 500198b654..d71eadeffb 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -63,7 +63,7 @@ int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n);
+ struct rte_ether_addr *mac_addrs, int uc_n, int n);
__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 3e66c9e6c8..dbc06aff7e 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -159,14 +159,6 @@ enum {
PCI_DEVICE_ID_MELLANOX_CONNECTX10 = 0x1027,
};
-/* Maximum number of simultaneous unicast MAC addresses. */
-#define MLX5_MAX_UC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous Multicast MAC addresses. */
-#define MLX5_MAX_MC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous MAC addresses. */
-#define MLX5_MAX_MAC_ADDRESSES \
- (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
-
/* Recognized Infiniband device physical port name types. */
enum mlx5_nl_phys_port_name_type {
MLX5_PHYS_PORT_NAME_TYPE_NOTSET = 0, /* Not set. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 1b241ba9d2..5b6e45df2a 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -1762,6 +1762,7 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
+ MLX5_MAX_UC_MAC_ADDRESSES,
MLX5_MAX_MAC_ADDRESSES);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 27e6f4e31a..3ad8bad02f 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -82,6 +82,14 @@
/* Maximum allowed MTU to be reported whenever PMD cannot query it from OS. */
#define MLX5_ETH_MAX_MTU (9978)
+/* Maximum number of simultaneous unicast MAC addresses. */
+#define MLX5_MAX_UC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous Multicast MAC addresses. */
+#define MLX5_MAX_MC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous MAC addresses. */
+#define MLX5_MAX_MAC_ADDRESSES \
+ (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
+
enum mlx5_ipool_index {
#if defined(HAVE_IBV_FLOW_DV_SUPPORT) || !defined(HAVE_INFINIBAND_VERBS_H)
MLX5_IPOOL_DECAP_ENCAP = 0, /* Pool for encap/decap resource. */
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 09/10] net/mlx5: use bitset for tracking MAC addresses
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
` (7 preceding siblings ...)
2026-07-09 16:02 ` [PATCH v4 08/10] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-09 16:02 ` [PATCH v4 10/10] net/mlx5: accept more unicast " David Marchand
9 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
EAL provides bitset that does the same as this set of mlx5 macros.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/mlx5_common.h | 16 ----------------
drivers/net/mlx5/linux/mlx5_os.c | 8 ++++----
drivers/net/mlx5/mlx5.h | 3 ++-
drivers/net/mlx5/mlx5_trigger.c | 2 +-
drivers/net/mlx5/windows/mlx5_os.c | 8 ++++----
5 files changed, 11 insertions(+), 26 deletions(-)
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index dbc06aff7e..71985794fa 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -30,22 +30,6 @@
#define MLX5_PCI_DRIVER_NAME "mlx5_pci"
#define MLX5_AUXILIARY_DRIVER_NAME "mlx5_auxiliary"
-/* Bit-field manipulation. */
-#define BITFIELD_DECLARE(bf, type, size) \
- type bf[(((size_t)(size) / (sizeof(type) * CHAR_BIT)) + \
- !!((size_t)(size) % (sizeof(type) * CHAR_BIT)))]
-#define BITFIELD_DEFINE(bf, type, size) \
- BITFIELD_DECLARE((bf), type, (size)) = { 0 }
-#define BITFIELD_SET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] |= \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_RESET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] &= \
- ~((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_ISSET(bf, b) \
- !!(((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] & \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT)))))
-
/*
* Helper macros to work around __VA_ARGS__ limitations in a C99 compliant
* manner.
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 5b6e45df2a..1cad6e1091 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3384,7 +3384,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[index]);
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -3413,7 +3413,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
mlx5_ifindex(dev),
mac);
if (!ret)
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return ret;
}
@@ -3498,12 +3498,12 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[i]);
- BITFIELD_RESET(priv->mac_own, i);
+ rte_bitset_clear(priv->mac_own, i);
}
}
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 3ad8bad02f..f7ba8df108 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -14,6 +14,7 @@
#include <rte_pci.h>
#include <rte_ether.h>
+#include <rte_bitset.h>
#include <ethdev_driver.h>
#include <rte_rwlock.h>
#include <rte_interrupts.h>
@@ -2018,7 +2019,7 @@ struct mlx5_priv {
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- BITFIELD_DECLARE(mac_own, uint64_t, MLX5_MAX_MAC_ADDRESSES);
+ RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
/* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 25847c8ba2..e41e4643e0 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1907,7 +1907,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
/* Add flows for unicast and multicast mac addresses added by API. */
if (!memcmp(mac, &cmp, sizeof(*mac)) ||
- !BITFIELD_ISSET(priv->mac_own, i) ||
+ !rte_bitset_test(priv->mac_own, i) ||
(dev->data->all_multicast && rte_is_multicast_ether_addr(mac)))
continue;
memcpy(&unicast.hdr.dst_addr.addr_bytes,
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index 9acfa8ec84..15de5c22a9 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -699,8 +699,8 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i))
- BITFIELD_RESET(priv->mac_own, i);
+ if (rte_bitset_test(priv->mac_own, i))
+ rte_bitset_clear(priv->mac_own, i);
}
}
@@ -719,7 +719,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
struct mlx5_priv *priv = dev->data->dev_private;
if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -757,7 +757,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
return -ENOTSUP;
}
/* Mark this MAC address as owned by the PMD */
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return 0;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v4 10/10] net/mlx5: accept more unicast MAC addresses
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
` (8 preceding siblings ...)
2026-07-09 16:02 ` [PATCH v4 09/10] net/mlx5: use bitset for tracking MAC addresses David Marchand
@ 2026-07-09 16:02 ` David Marchand
2026-07-10 6:44 ` David Marchand
2026-07-10 7:48 ` David Marchand
9 siblings, 2 replies; 146+ messages in thread
From: David Marchand @ 2026-07-09 16:02 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Starting firmware version 22.49.1014, the number of mac addresses
per VF is not capped to 128 anymore.
The value can be increased via devlink:
$ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
cmode driverinit
$ devlink dev reload pci/0000:3b:00.2
On the DPDK side, we must retrieve the maximum number of unicast
and multicast addresses supported with a query to the firmware.
Then, dynamically allocate the mac addresses arrays and report the
limit instead of the previous hardcoded value.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/mlx5_devx_cmds.c | 4 +++
drivers/common/mlx5/mlx5_devx_cmds.h | 2 ++
drivers/net/mlx5/linux/mlx5_os.c | 42 ++++++++++++++++++++++------
drivers/net/mlx5/mlx5.c | 9 ++----
drivers/net/mlx5/mlx5.h | 8 ++++--
drivers/net/mlx5/mlx5_ethdev.c | 2 +-
drivers/net/mlx5/mlx5_mac.c | 22 +++++++++------
drivers/net/mlx5/mlx5_trigger.c | 10 +++----
drivers/net/mlx5/windows/mlx5_os.c | 40 ++++++++++++++++++++------
9 files changed, 99 insertions(+), 40 deletions(-)
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.c b/drivers/common/mlx5/mlx5_devx_cmds.c
index 140b057ab4..e5d9c92779 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.c
+++ b/drivers/common/mlx5/mlx5_devx_cmds.c
@@ -1110,6 +1110,10 @@ mlx5_devx_cmd_query_hca_attr(void *ctx,
attr->log_max_pd = MLX5_GET(cmd_hca_cap, hcattr, log_max_pd);
attr->log_max_srq = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq);
attr->log_max_srq_sz = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq_sz);
+ attr->log_max_current_uc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_uc_list);
+ attr->log_max_current_mc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_mc_list);
attr->reg_c_preserve =
MLX5_GET(cmd_hca_cap, hcattr, reg_c_preserve);
attr->mmo_regex_qp_en = MLX5_GET(cmd_hca_cap, hcattr, regexp_mmo_qp);
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.h b/drivers/common/mlx5/mlx5_devx_cmds.h
index 90beb2e9e6..7fe89bc6a4 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.h
+++ b/drivers/common/mlx5/mlx5_devx_cmds.h
@@ -356,6 +356,8 @@ struct mlx5_hca_attr {
uint8_t tx_sw_owner_v2:1;
uint8_t esw_sw_owner:1;
uint8_t esw_sw_owner_v2:1;
+ uint32_t log_max_current_uc_list:5;
+ uint32_t log_max_current_mc_list:5;
};
/* LAG Context. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 1cad6e1091..23632f55af 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -389,6 +389,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
sh->dev_cap.esw_info.regc_mask = 0;
#endif
sh->dev_cap.esw_info.is_set = 1;
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = 1u << hca_attr->log_max_current_uc_list;
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = 1u << hca_attr->log_max_current_mc_list;
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -1464,6 +1474,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
/* Some internal functions rely on Netlink sockets, open them now. */
priv->nl_socket_rdma = nl_rdma;
priv->nl_socket_route = mlx5_nl_init(NETLINK_ROUTE, 0);
@@ -1762,8 +1788,8 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
- MLX5_MAX_UC_MAC_ADDRESSES,
- MLX5_MAX_MAC_ADDRESSES);
+ sh->dev_cap.max_mac_addrs,
+ sh->dev_cap.max_uc_mac_addrs);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
TAILQ_INIT(&priv->flow_meters);
@@ -1963,17 +1989,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_flex_item_port_cleanup(eth_dev);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ eth_dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
if (eth_dev != NULL)
eth_dev->data->dev_private = NULL;
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
if (nl_rdma >= 0)
@@ -3497,7 +3521,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
const int vf = priv->sh->dev_cap.vf;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
diff --git a/drivers/net/mlx5/mlx5.c b/drivers/net/mlx5/mlx5.c
index 61c26d1206..1323ba49ac 100644
--- a/drivers/net/mlx5/mlx5.c
+++ b/drivers/net/mlx5/mlx5.c
@@ -2546,6 +2546,9 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_list_destroy(priv->hrxqs);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
sh->port[priv->dev_port - 1].nl_ih_port_id = RTE_MAX_ETHPORTS;
/*
* The interrupt handler port id must be reset before priv is reset
@@ -2580,12 +2583,6 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_flow_pools_destroy(priv);
memset(priv, 0, sizeof(*priv));
priv->domain_id = RTE_ETH_DEV_SWITCH_DOMAIN_ID_INVALID;
- /*
- * Reset mac_addrs to NULL such that it is not freed as part of
- * rte_eth_dev_release_port(). mac_addrs is part of dev_private so
- * it is freed when dev_private is freed.
- */
- dev->data->mac_addrs = NULL;
return 0;
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index f7ba8df108..1fcc40fade 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -217,6 +217,9 @@ struct mlx5_dev_cap {
} mprq; /* Capability for Multi-Packet RQ. */
char fw_ver[64]; /* Firmware version of this device. */
struct flow_hw_port_info esw_info; /* E-switch manager reg_c0. */
+ uint16_t max_uc_mac_addrs; /* Maximum unicast MAC addresses. */
+ uint16_t max_mc_mac_addrs; /* Maximum multicast MAC addresses. */
+ uint16_t max_mac_addrs; /* Total maximum MAC addresses. */
};
#define MLX5_MPESW_PORT_INVALID (-1)
@@ -2018,9 +2021,8 @@ struct mlx5_priv {
struct mlx5_dev_ctx_shared *sh; /* Shared device context. */
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
- struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
- /* Bit-field of MAC addresses owned by the PMD. */
+ struct rte_ether_addr *mac; /* MAC addresses. */
+ uint64_t *mac_own; /* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
/* Device properties. */
diff --git a/drivers/net/mlx5/mlx5_ethdev.c b/drivers/net/mlx5/mlx5_ethdev.c
index 8160d10e7e..306c1cd734 100644
--- a/drivers/net/mlx5/mlx5_ethdev.c
+++ b/drivers/net/mlx5/mlx5_ethdev.c
@@ -392,7 +392,7 @@ mlx5_dev_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *info)
max = RTE_MIN(max, (unsigned int)UINT16_MAX);
info->max_rx_queues = max;
info->max_tx_queues = max;
- info->max_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ info->max_mac_addrs = priv->sh->dev_cap.max_uc_mac_addrs;
info->rx_queue_offload_capa = mlx5_get_rx_queue_offloads(dev);
info->rx_seg_capa.max_nseg = MLX5_MAX_RXQ_NSEG;
info->rx_seg_capa.multi_pools = !priv->config.mprq.enabled;
diff --git a/drivers/net/mlx5/mlx5_mac.c b/drivers/net/mlx5/mlx5_mac.c
index 0e5d2be530..2ce7cfd407 100644
--- a/drivers/net/mlx5/mlx5_mac.c
+++ b/drivers/net/mlx5/mlx5_mac.c
@@ -36,7 +36,9 @@ mlx5_internal_mac_addr_remove(struct rte_eth_dev *dev,
uint32_t index,
struct rte_ether_addr *addr)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ struct mlx5_priv *priv = dev->data->dev_private;
+
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(&dev->data->mac_addrs[index]))
return false;
mlx5_os_mac_addr_remove(dev, index);
@@ -63,16 +65,17 @@ static int
mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
unsigned int i;
int ret;
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(mac)) {
rte_errno = EINVAL;
return -rte_errno;
}
/* First, make sure this address isn't already configured. */
- for (i = 0; (i != MLX5_MAX_MAC_ADDRESSES); ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
/* Skip this index, it's going to be reconfigured. */
if (i == index)
continue;
@@ -101,10 +104,11 @@ mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
void
mlx5_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
struct rte_ether_addr addr = { 0 };
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES)
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs)
return;
if (mlx5_internal_mac_addr_remove(dev, index, &addr)) {
ret = mlx5_traffic_mac_remove(dev, &addr);
@@ -133,9 +137,10 @@ int
mlx5_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index, uint32_t vmdq __rte_unused)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES) {
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs) {
rte_errno = EINVAL;
return -rte_errno;
}
@@ -217,16 +222,17 @@ int
mlx5_set_mc_addr_list(struct rte_eth_dev *dev,
struct rte_ether_addr *mc_addr_set, uint32_t nb_mc_addr)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
uint32_t i;
int ret;
- if (nb_mc_addr >= MLX5_MAX_MC_MAC_ADDRESSES) {
+ if (nb_mc_addr >= priv->sh->dev_cap.max_mc_mac_addrs) {
rte_errno = ENOSPC;
return -rte_errno;
}
- for (i = MLX5_MAX_UC_MAC_ADDRESSES; i != MLX5_MAX_MAC_ADDRESSES; ++i)
+ for (i = priv->sh->dev_cap.max_uc_mac_addrs; i != priv->sh->dev_cap.max_mac_addrs; ++i)
mlx5_internal_mac_addr_remove(dev, i, NULL);
- i = MLX5_MAX_UC_MAC_ADDRESSES;
+ i = priv->sh->dev_cap.max_uc_mac_addrs;
while (nb_mc_addr--) {
ret = mlx5_internal_mac_addr_add(dev, mc_addr_set++, i++);
if (ret)
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index e41e4643e0..e535e3e5be 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1902,7 +1902,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
}
}
/* Add MAC address flows. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
/* Add flows for unicast and multicast mac addresses added by API. */
@@ -2172,7 +2172,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
return 0;
/* Add all unicast DMAC flow rules with new VLAN attached. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2189,7 +2189,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
* Removing after creating VLAN rules so that traffic "gap" is not introduced.
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2227,7 +2227,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
* Recreating first to ensure no traffic "gap".
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2240,7 +2240,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
}
/* Remove all unicast DMAC flow rules with this VLAN. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index 15de5c22a9..30d99b8b64 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -261,6 +261,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
MLX5_GET(initial_seg, pv_iseg, fw_rev_subminor));
DRV_LOG(DEBUG, "Packet pacing is not supported.");
mlx5_rt_timestamp_config(sh, hca_attr);
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = 1u << hca_attr->log_max_current_uc_list;
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = 1u << hca_attr->log_max_current_mc_list;
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -396,6 +406,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
priv->mp_id.port_id = port_id;
strlcpy(priv->mp_id.name, MLX5_MP_NAME, RTE_MP_MAX_NAME_LEN);
priv->representor = !!switch_info->representor;
@@ -612,17 +638,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_l3t_destroy(priv->mtr_profile_tbl);
if (own_domain_id)
claim_zero(rte_eth_switch_domain_free(priv->domain_id));
+ mlx5_free(priv->mac);
+ eth_dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
if (eth_dev != NULL)
eth_dev->data->dev_private = NULL;
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
MLX5_ASSERT(err > 0);
@@ -698,7 +722,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
struct mlx5_priv *priv = dev->data->dev_private;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i))
rte_bitset_clear(priv->mac_own, i);
}
@@ -718,7 +742,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
struct mlx5_priv *priv = dev->data->dev_private;
- if (index < MLX5_MAX_MAC_ADDRESSES)
+ if (index < priv->sh->dev_cap.max_mac_addrs)
rte_bitset_clear(priv->mac_own, index);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH v4 10/10] net/mlx5: accept more unicast MAC addresses
2026-07-09 16:02 ` [PATCH v4 10/10] net/mlx5: accept more unicast " David Marchand
@ 2026-07-10 6:44 ` David Marchand
2026-07-10 7:48 ` David Marchand
1 sibling, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-10 6:44 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad, Thomas Monjalon
On Thu, 9 Jul 2026 at 18:04, David Marchand <david.marchand@redhat.com> wrote:
> diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
> index 1cad6e1091..23632f55af 100644
> --- a/drivers/net/mlx5/linux/mlx5_os.c
> +++ b/drivers/net/mlx5/linux/mlx5_os.c
[snip]
> @@ -1762,8 +1788,8 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
> mlx5_nl_mac_addr_sync(priv->nl_socket_route,
> mlx5_ifindex(eth_dev),
> eth_dev->data->mac_addrs,
> - MLX5_MAX_UC_MAC_ADDRESSES,
> - MLX5_MAX_MAC_ADDRESSES);
> + sh->dev_cap.max_mac_addrs,
> + sh->dev_cap.max_uc_mac_addrs);
> priv->ctrl_flows = 0;
> rte_spinlock_init(&priv->flow_list_lock);
> TAILQ_INIT(&priv->flow_meters);
Oops, I swapped those two arguments when cleaning up the patches...
I'll need a new revision anyway when sending the series rebased on 26.11-rc0.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v4 10/10] net/mlx5: accept more unicast MAC addresses
2026-07-09 16:02 ` [PATCH v4 10/10] net/mlx5: accept more unicast " David Marchand
2026-07-10 6:44 ` David Marchand
@ 2026-07-10 7:48 ` David Marchand
1 sibling, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-10 7:48 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
On Thu, 9 Jul 2026 at 18:04, David Marchand <david.marchand@redhat.com> wrote:
> diff --git a/drivers/common/mlx5/mlx5_devx_cmds.h b/drivers/common/mlx5/mlx5_devx_cmds.h
> index 90beb2e9e6..7fe89bc6a4 100644
> --- a/drivers/common/mlx5/mlx5_devx_cmds.h
> +++ b/drivers/common/mlx5/mlx5_devx_cmds.h
> @@ -356,6 +356,8 @@ struct mlx5_hca_attr {
> uint8_t tx_sw_owner_v2:1;
> uint8_t esw_sw_owner:1;
> uint8_t esw_sw_owner_v2:1;
> + uint32_t log_max_current_uc_list:5;
> + uint32_t log_max_current_mc_list:5;
uint8_t is enough.
> };
>
> /* LAG Context. */
[...]
> diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
> index f7ba8df108..1fcc40fade 100644
> --- a/drivers/net/mlx5/mlx5.h
> +++ b/drivers/net/mlx5/mlx5.h
> @@ -217,6 +217,9 @@ struct mlx5_dev_cap {
> } mprq; /* Capability for Multi-Packet RQ. */
> char fw_ver[64]; /* Firmware version of this device. */
> struct flow_hw_port_info esw_info; /* E-switch manager reg_c0. */
> + uint16_t max_uc_mac_addrs; /* Maximum unicast MAC addresses. */
> + uint16_t max_mc_mac_addrs; /* Maximum multicast MAC addresses. */
> + uint16_t max_mac_addrs; /* Total maximum MAC addresses. */
> };
>
> #define MLX5_MPESW_PORT_INVALID (-1)
The firmware provides a log value on 5 bits, so max value for uc and
mc would be 1 << 31, and the sum would be up to 1 << 32.
In theory, ethdev caps max_mac_addrs to 1 << 32 - 1, but the firmware
only allows up to 4k uc macs and extending this limit has been said
not possible easily.
So reaching 1 << 32 as a total is not going to happen soon and
uint32_t is enough for those 3 fields.
In next revision.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* RE: [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install
2026-07-09 16:02 ` [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install David Marchand
@ 2026-07-13 13:12 ` Loftus, Ciara
2026-07-13 14:10 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Loftus, Ciara @ 2026-07-13 13:12 UTC (permalink / raw)
To: David Marchand, dev@dpdk.org
Cc: rjarry@redhat.com, cfontain@redhat.com, stable@dpdk.org,
Medvedkin, Vladimir, Richardson, Bruce
> Subject: [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install
>
> On port restart, all MAC addresses get pushed *twice* to the hardware,
> once by the driver and once by the eth_dev_mac_restore() in ethdev.
>
> On the other hand, MAC address filters are reset in the hardware
> by the PF only when a VF reset is triggered.
>
> Strictly speaking, the mac restore on port (re)start is unneeded,
> if no VF reset happened, so we can announce to ethdev that no mac
> restoration is needed via a get_restore_flags callback.
>
> Then, move the mac restoration to the VF reset handler.
>
> Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux
> driver")
> Cc: stable@dpdk.org
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
> drivers/net/intel/iavf/iavf_ethdev.c | 31 ++++++++++++++++++++--------
> 1 file changed, 22 insertions(+), 9 deletions(-)
>
> diff --git a/drivers/net/intel/iavf/iavf_ethdev.c
> b/drivers/net/intel/iavf/iavf_ethdev.c
> index 179d90ec55..fe54df4b9f 100644
> --- a/drivers/net/intel/iavf/iavf_ethdev.c
> +++ b/drivers/net/intel/iavf/iavf_ethdev.c
> @@ -134,6 +134,8 @@ static int iavf_dev_vlan_filter_set(struct rte_eth_dev
> *dev,
> static int iavf_vlan_tpid_set(struct rte_eth_dev *dev,
> enum rte_vlan_type vlan_type, uint16_t tpid);
> static int iavf_dev_vlan_offload_set(struct rte_eth_dev *dev, int mask);
> +static uint64_t iavf_get_restore_flags(struct rte_eth_dev *dev,
> + enum rte_eth_dev_operation op);
> static int iavf_dev_rss_reta_update(struct rte_eth_dev *dev,
> struct rte_eth_rss_reta_entry64 *reta_conf,
> uint16_t reta_size);
> @@ -264,6 +266,7 @@ static const struct eth_dev_ops iavf_eth_dev_ops = {
> .tx_done_cleanup = iavf_dev_tx_done_cleanup,
> .get_monitor_addr = iavf_get_monitor_addr,
> .tm_ops_get = iavf_tm_ops_get,
> + .get_restore_flags = iavf_get_restore_flags,
> };
>
> static int
> @@ -284,6 +287,13 @@ iavf_tm_ops_get(struct rte_eth_dev *dev,
> return 0;
> }
>
> +static uint64_t
> +iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
> + __rte_unused enum rte_eth_dev_operation op)
> +{
> + return RTE_ETH_RESTORE_ALL & ~RTE_ETH_RESTORE_MAC_ADDR;
> +}
> +
> __rte_unused
> static int
> iavf_vfr_inprogress(struct iavf_hw *hw)
> @@ -1074,15 +1084,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
> rte_intr_enable(intr_handle);
> }
>
> - /* Set all mac addrs */
> - iavf_add_del_all_mac_addr(adapter, true);
> -
> - if (!adapter->mac_primary_set)
> - adapter->mac_primary_set = true;
> -
> - /* Set all multicast addresses */
> - iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> >mc_addrs_num,
> - true);
> + if (!adapter->mac_primary_set) {
> + if (iavf_add_del_eth_addr(adapter, &dev->data-
> >mac_addrs[0], true,
> + VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
> + PMD_DRV_LOG(ERR, "failed to add primary MAC:"
> RTE_ETHER_ADDR_PRT_FMT,
> + RTE_ETHER_ADDR_BYTES(&dev->data-
> >mac_addrs[0]));
> + else
> + adapter->mac_primary_set = true;
> + }
>
> rte_spinlock_init(&vf->phc_time_aq_lock);
>
> @@ -3424,6 +3433,10 @@ iavf_handle_hw_reset(struct rte_eth_dev *dev,
> bool vf_initiated_reset)
> if (ret)
> goto error;
>
> + /* after a VF reset, all mac addresses got flushed, restore them
> */
> + iavf_add_del_all_mac_addr(adapter, true);
> + iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> >mc_addrs_num, true);
> +
This restore block only runs if the device was started when the handler was
entered. It might make more sense to move it into iavf_post_reset_reconfig
which will run for VF-initiated resets regardless of started state, and for
PF-initiated will run if the device was started on entry.
For PF-initiated if the device was not started we return from
iavf_handle_hw_reset() immediately, so there's no opportunity to restore
the MAC addresses at all. It needs to be considered how to restore the macs
in that case.
Need to consider as well how the restore_flags and auto_reconfig features
work with one another. If the auto_reconfig devarg is set to zero it means
the user doesn't want any restoration of settings upon reset.
> dev->data->dev_started = 1;
> }
>
> --
> 2.54.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install
2026-07-13 13:12 ` Loftus, Ciara
@ 2026-07-13 14:10 ` David Marchand
2026-07-14 9:23 ` Loftus, Ciara
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-07-13 14:10 UTC (permalink / raw)
To: Loftus, Ciara
Cc: dev@dpdk.org, rjarry@redhat.com, cfontain@redhat.com,
stable@dpdk.org, Medvedkin, Vladimir, Richardson, Bruce
On Mon, 13 Jul 2026 at 15:14, Loftus, Ciara <ciara.loftus@intel.com> wrote:
>
> > Subject: [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install
> >
> > On port restart, all MAC addresses get pushed *twice* to the hardware,
> > once by the driver and once by the eth_dev_mac_restore() in ethdev.
> >
> > On the other hand, MAC address filters are reset in the hardware
> > by the PF only when a VF reset is triggered.
> >
> > Strictly speaking, the mac restore on port (re)start is unneeded,
> > if no VF reset happened, so we can announce to ethdev that no mac
> > restoration is needed via a get_restore_flags callback.
> >
> > Then, move the mac restoration to the VF reset handler.
> >
> > Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux
> > driver")
> > Cc: stable@dpdk.org
> >
> > Signed-off-by: David Marchand <david.marchand@redhat.com>
> > ---
> > drivers/net/intel/iavf/iavf_ethdev.c | 31 ++++++++++++++++++++--------
> > 1 file changed, 22 insertions(+), 9 deletions(-)
> >
> > diff --git a/drivers/net/intel/iavf/iavf_ethdev.c
> > b/drivers/net/intel/iavf/iavf_ethdev.c
> > index 179d90ec55..fe54df4b9f 100644
> > --- a/drivers/net/intel/iavf/iavf_ethdev.c
> > +++ b/drivers/net/intel/iavf/iavf_ethdev.c
> > @@ -134,6 +134,8 @@ static int iavf_dev_vlan_filter_set(struct rte_eth_dev
> > *dev,
> > static int iavf_vlan_tpid_set(struct rte_eth_dev *dev,
> > enum rte_vlan_type vlan_type, uint16_t tpid);
> > static int iavf_dev_vlan_offload_set(struct rte_eth_dev *dev, int mask);
> > +static uint64_t iavf_get_restore_flags(struct rte_eth_dev *dev,
> > + enum rte_eth_dev_operation op);
> > static int iavf_dev_rss_reta_update(struct rte_eth_dev *dev,
> > struct rte_eth_rss_reta_entry64 *reta_conf,
> > uint16_t reta_size);
> > @@ -264,6 +266,7 @@ static const struct eth_dev_ops iavf_eth_dev_ops = {
> > .tx_done_cleanup = iavf_dev_tx_done_cleanup,
> > .get_monitor_addr = iavf_get_monitor_addr,
> > .tm_ops_get = iavf_tm_ops_get,
> > + .get_restore_flags = iavf_get_restore_flags,
> > };
> >
> > static int
> > @@ -284,6 +287,13 @@ iavf_tm_ops_get(struct rte_eth_dev *dev,
> > return 0;
> > }
> >
> > +static uint64_t
> > +iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
> > + __rte_unused enum rte_eth_dev_operation op)
> > +{
> > + return RTE_ETH_RESTORE_ALL & ~RTE_ETH_RESTORE_MAC_ADDR;
> > +}
> > +
> > __rte_unused
> > static int
> > iavf_vfr_inprogress(struct iavf_hw *hw)
> > @@ -1074,15 +1084,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
> > rte_intr_enable(intr_handle);
> > }
> >
> > - /* Set all mac addrs */
> > - iavf_add_del_all_mac_addr(adapter, true);
> > -
> > - if (!adapter->mac_primary_set)
> > - adapter->mac_primary_set = true;
> > -
> > - /* Set all multicast addresses */
> > - iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> > >mc_addrs_num,
> > - true);
> > + if (!adapter->mac_primary_set) {
> > + if (iavf_add_del_eth_addr(adapter, &dev->data-
> > >mac_addrs[0], true,
> > + VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
> > + PMD_DRV_LOG(ERR, "failed to add primary MAC:"
> > RTE_ETHER_ADDR_PRT_FMT,
> > + RTE_ETHER_ADDR_BYTES(&dev->data-
> > >mac_addrs[0]));
> > + else
> > + adapter->mac_primary_set = true;
> > + }
> >
> > rte_spinlock_init(&vf->phc_time_aq_lock);
> >
> > @@ -3424,6 +3433,10 @@ iavf_handle_hw_reset(struct rte_eth_dev *dev,
> > bool vf_initiated_reset)
> > if (ret)
> > goto error;
> >
> > + /* after a VF reset, all mac addresses got flushed, restore them
> > */
> > + iavf_add_del_all_mac_addr(adapter, true);
> > + iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> > >mc_addrs_num, true);
> > +
>
> This restore block only runs if the device was started when the handler was
> entered. It might make more sense to move it into iavf_post_reset_reconfig
> which will run for VF-initiated resets regardless of started state, and for
> PF-initiated will run if the device was started on entry.
> For PF-initiated if the device was not started we return from
> iavf_handle_hw_reset() immediately, so there's no opportunity to restore
> the MAC addresses at all. It needs to be considered how to restore the macs
> in that case.
> Need to consider as well how the restore_flags and auto_reconfig features
> work with one another. If the auto_reconfig devarg is set to zero it means
> the user doesn't want any restoration of settings upon reset.
I was not aware of this knob.
It seems crazy to me to have such specific flag to disable something
that I take foregranted during the life of a DPDK port.
I really hope no application relies on the mac addresses list being
flushed during a VF reset...
While moving the code to iavf_post_reset_reconfig kind of makes sense
to me, I honesly don't know what to do wrt restore_flags.
It seems orthogonal to me.
If ethdev decides to restore the mac addresses, I don't see why some
devargs would have something to say about it.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* RE: [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install
2026-07-13 14:10 ` David Marchand
@ 2026-07-14 9:23 ` Loftus, Ciara
0 siblings, 0 replies; 146+ messages in thread
From: Loftus, Ciara @ 2026-07-14 9:23 UTC (permalink / raw)
To: David Marchand
Cc: dev@dpdk.org, rjarry@redhat.com, cfontain@redhat.com,
stable@dpdk.org, Medvedkin, Vladimir, Richardson, Bruce
> Subject: Re: [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install
>
> On Mon, 13 Jul 2026 at 15:14, Loftus, Ciara <ciara.loftus@intel.com> wrote:
> >
> > > Subject: [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install
> > >
> > > On port restart, all MAC addresses get pushed *twice* to the hardware,
> > > once by the driver and once by the eth_dev_mac_restore() in ethdev.
> > >
> > > On the other hand, MAC address filters are reset in the hardware
> > > by the PF only when a VF reset is triggered.
> > >
> > > Strictly speaking, the mac restore on port (re)start is unneeded,
> > > if no VF reset happened, so we can announce to ethdev that no mac
> > > restoration is needed via a get_restore_flags callback.
> > >
> > > Then, move the mac restoration to the VF reset handler.
> > >
> > > Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux
> > > driver")
> > > Cc: stable@dpdk.org
> > >
> > > Signed-off-by: David Marchand <david.marchand@redhat.com>
> > > ---
> > > drivers/net/intel/iavf/iavf_ethdev.c | 31 ++++++++++++++++++++--------
> > > 1 file changed, 22 insertions(+), 9 deletions(-)
> > >
> > > diff --git a/drivers/net/intel/iavf/iavf_ethdev.c
> > > b/drivers/net/intel/iavf/iavf_ethdev.c
> > > index 179d90ec55..fe54df4b9f 100644
> > > --- a/drivers/net/intel/iavf/iavf_ethdev.c
> > > +++ b/drivers/net/intel/iavf/iavf_ethdev.c
> > > @@ -134,6 +134,8 @@ static int iavf_dev_vlan_filter_set(struct
> rte_eth_dev
> > > *dev,
> > > static int iavf_vlan_tpid_set(struct rte_eth_dev *dev,
> > > enum rte_vlan_type vlan_type, uint16_t tpid);
> > > static int iavf_dev_vlan_offload_set(struct rte_eth_dev *dev, int mask);
> > > +static uint64_t iavf_get_restore_flags(struct rte_eth_dev *dev,
> > > + enum rte_eth_dev_operation op);
> > > static int iavf_dev_rss_reta_update(struct rte_eth_dev *dev,
> > > struct rte_eth_rss_reta_entry64 *reta_conf,
> > > uint16_t reta_size);
> > > @@ -264,6 +266,7 @@ static const struct eth_dev_ops iavf_eth_dev_ops
> = {
> > > .tx_done_cleanup = iavf_dev_tx_done_cleanup,
> > > .get_monitor_addr = iavf_get_monitor_addr,
> > > .tm_ops_get = iavf_tm_ops_get,
> > > + .get_restore_flags = iavf_get_restore_flags,
> > > };
> > >
> > > static int
> > > @@ -284,6 +287,13 @@ iavf_tm_ops_get(struct rte_eth_dev *dev,
> > > return 0;
> > > }
> > >
> > > +static uint64_t
> > > +iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
> > > + __rte_unused enum rte_eth_dev_operation op)
> > > +{
> > > + return RTE_ETH_RESTORE_ALL & ~RTE_ETH_RESTORE_MAC_ADDR;
> > > +}
> > > +
> > > __rte_unused
> > > static int
> > > iavf_vfr_inprogress(struct iavf_hw *hw)
> > > @@ -1074,15 +1084,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
> > > rte_intr_enable(intr_handle);
> > > }
> > >
> > > - /* Set all mac addrs */
> > > - iavf_add_del_all_mac_addr(adapter, true);
> > > -
> > > - if (!adapter->mac_primary_set)
> > > - adapter->mac_primary_set = true;
> > > -
> > > - /* Set all multicast addresses */
> > > - iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> > > >mc_addrs_num,
> > > - true);
> > > + if (!adapter->mac_primary_set) {
> > > + if (iavf_add_del_eth_addr(adapter, &dev->data-
> > > >mac_addrs[0], true,
> > > + VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
> > > + PMD_DRV_LOG(ERR, "failed to add primary MAC:"
> > > RTE_ETHER_ADDR_PRT_FMT,
> > > + RTE_ETHER_ADDR_BYTES(&dev->data-
> > > >mac_addrs[0]));
> > > + else
> > > + adapter->mac_primary_set = true;
> > > + }
> > >
> > > rte_spinlock_init(&vf->phc_time_aq_lock);
> > >
> > > @@ -3424,6 +3433,10 @@ iavf_handle_hw_reset(struct rte_eth_dev
> *dev,
> > > bool vf_initiated_reset)
> > > if (ret)
> > > goto error;
> > >
> > > + /* after a VF reset, all mac addresses got flushed, restore them
> > > */
> > > + iavf_add_del_all_mac_addr(adapter, true);
> > > + iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> > > >mc_addrs_num, true);
> > > +
> >
> > This restore block only runs if the device was started when the handler was
> > entered. It might make more sense to move it into iavf_post_reset_reconfig
> > which will run for VF-initiated resets regardless of started state, and for
> > PF-initiated will run if the device was started on entry.
> > For PF-initiated if the device was not started we return from
> > iavf_handle_hw_reset() immediately, so there's no opportunity to restore
> > the MAC addresses at all. It needs to be considered how to restore the macs
> > in that case.
> > Need to consider as well how the restore_flags and auto_reconfig features
> > work with one another. If the auto_reconfig devarg is set to zero it means
> > the user doesn't want any restoration of settings upon reset.
>
> I was not aware of this knob.
> It seems crazy to me to have such specific flag to disable something
> that I take foregranted during the life of a DPDK port.
> I really hope no application relies on the mac addresses list being
> flushed during a VF reset...
I think it's reasonable to keep the MAC restore outside of auto_reconfig,
that also preserves the existing behaviour of restoring MACs on reset.
But with your patch the restore is skipped whenever the port was stopped at
reset entry. Will the MACs be restored on the next dev_start?
>
> While moving the code to iavf_post_reset_reconfig kind of makes sense
> to me, I honesly don't know what to do wrt restore_flags.
> It seems orthogonal to me.
> If ethdev decides to restore the mac addresses, I don't see why some
> devargs would have something to say about it.
I agree ethdev should take precedence over the devarg. One option is to
have iavf_post_reset_reconfig honour the same restore_flags mask the normal
start path uses, so the two paths don't diverge. I don't think you need to
address that in your set, it's something I can look into.
And I agree the flag is quite specific. I think its main use is during a VF
initiated reset, the user may want to reset the VF completely and not
restore anything. I think instead of a devarg I could address this by
adding a parameter to the dev-private rte_pmd_iavf_reinit API. I'll look
into this too. Definitely room for improvement here around the reset path.
>
>
> --
> David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v5 00/10] Remove limitations coming from legacy VMDq
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (7 preceding siblings ...)
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 01/10] ethdev: check VMDq availability David Marchand
` (10 more replies)
2026-08-24 11:42 ` [PATCH v6 0/3] " David Marchand
` (6 subsequent siblings)
15 siblings, 11 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain
Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
VMDq has been imposing a maximum number of mac addresses in the
mac_addr_add/del API.
Nowadays, new Intel drivers do not support the feature and few other
drivers implement this feature.
This series enforces that the driver announces VMDq pools before
using VMDq related features, then remove the limit of number of
mac addresses for others.
Next step could be to remove the VMDq pool notion from the generic API.
However I have some concern about this, as changing the quite stable
mac_addr_add/del API now seems a lot of noise for not much benefit.
--
David Marchand
Changes since v4:
- rebased,
- moved mac restoration in dedicated IAVF reset helper,
- fixed inverted arguments when calling mlx5 mac sync,
Changes since v3:
- rebased (this series is too late for 26.07),
- update some doxygen comments,
- added mlx5 changes,
Changes since v2:
- changed approach: did not introduce a new device capability,
relied on already existing dev_info->max_vmdq_pools,
- fixed duplicate mac addition without VMDq,
- updated documentation,
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
- fixed buffer overflow on mailbox messages during port restart/VF reset,
- fixed duplicate MAC address installation on port start/restart,
David Marchand (10):
ethdev: check VMDq availability
ethdev: skip VMDq pools unless configured
ethdev: hide VMDq internal sizes
net/iavf: accept up to 32k unicast MAC addresses
net/iavf: fix duplicate MAC addresses install
net/mlx5: remove MAC addresses flush helper on Linux
net/mlx5: remove redundant MAC address index checks
net/mlx5: pass maximum number of unicast MAC to common code
net/mlx5: use bitset for tracking MAC addresses
net/mlx5: accept more unicast MAC addresses
doc/guides/rel_notes/release_26_11.rst | 21 +++++
drivers/common/mlx5/linux/mlx5_nl.c | 90 +++++---------------
drivers/common/mlx5/linux/mlx5_nl.h | 11 +--
drivers/common/mlx5/mlx5_common.h | 24 ------
drivers/common/mlx5/mlx5_devx_cmds.c | 4 +
drivers/common/mlx5/mlx5_devx_cmds.h | 2 +
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 44 ++++++----
drivers/net/intel/iavf/iavf_vchnl.c | 113 ++++++++++++++++++-------
drivers/net/mlx5/linux/mlx5_os.c | 61 +++++++++----
drivers/net/mlx5/mlx5.c | 9 +-
drivers/net/mlx5/mlx5.h | 17 +++-
drivers/net/mlx5/mlx5_ethdev.c | 2 +-
drivers/net/mlx5/mlx5_mac.c | 22 +++--
drivers/net/mlx5/mlx5_trigger.c | 12 +--
drivers/net/mlx5/windows/mlx5_os.c | 48 ++++++++---
lib/ethdev/ethdev_driver.h | 8 +-
lib/ethdev/rte_ethdev.c | 68 +++++++++++----
lib/ethdev/rte_ethdev.h | 8 +-
20 files changed, 346 insertions(+), 224 deletions(-)
--
2.54.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v5 01/10] ethdev: check VMDq availability
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 02/10] ethdev: skip VMDq pools unless configured David Marchand
` (9 subsequent siblings)
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Andrew Rybchenko, Thomas Monjalon
Refuse VMDq related Rx/Tx modes when the driver do not announce VMDq
pools availability.
This will used later as a gate to ignore/reject VMDq related matters.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v4:
- rebased,
Changes since v2:
- added an entry in release notes,
- changed approach: relied on pre-existing dev_info->max_vmdq_pools
rather than introduce a new device capability,
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
---
doc/guides/rel_notes/release_26_11.rst | 6 ++++++
lib/ethdev/rte_ethdev.c | 16 ++++++++++++++++
2 files changed, 22 insertions(+)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 938617ca75..618b5acd2d 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -84,6 +84,12 @@ API Changes
Also, make sure to start the actual text at the margin.
=======================================================
+* **ethdev: updated VMDq related API.**
+
+ * At port configuration time, the number of VMDq pools advertised by a driver is now used to
+ validate VMDq related Rx and Tx modes (``RTE_ETH_MQ_RX_VMDQ_FLAG``, ``RTE_ETH_MQ_TX_VMDQ_DCB``,
+ ``RTE_ETH_MQ_TX_VMDQ_ONLY``).
+
ABI Changes
-----------
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 9efeaf77cb..9e305d98c1 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1581,6 +1581,22 @@ rte_eth_dev_configure(uint16_t port_id, uint16_t nb_rx_q, uint16_t nb_tx_q,
goto rollback;
}
+ if (dev_info.max_vmdq_pools == 0) {
+ if ((dev_conf->rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq rx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ if (dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_DCB ||
+ dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_ONLY) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq tx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ }
+
/*
* Setup new number of Rx/Tx queues and reconfigure device.
*/
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 02/10] ethdev: skip VMDq pools unless configured
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
2026-07-23 12:41 ` [PATCH v5 01/10] ethdev: check VMDq availability David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 03/10] ethdev: hide VMDq internal sizes David Marchand
` (8 subsequent siblings)
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Andrew Rybchenko, Nithin Dabilpuram,
Kiran Kumar K, Sunil Kumar Kori, Satha Rao, Harman Kalra,
Thomas Monjalon
The mac_addr_add API describes that only the 0 pool should be passed
unless VMDq has been enabled, though there was no validation so far.
Add such a check, then cleanup the MAC related operations (adding,
removing, restoring).
As a side effect, the net/cnxk does not need to manually reset the
mac_pool_sel[] array.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v4:
- rebased,
Changes since v3:
- updated doxygen comments,
Changes since v2:
- added an entry in release notes,
- fixed duplicate mac address handling for !vmdq,
- rewrote update of eth_dev_mac_restore to isolate the !vmdq case,
---
doc/guides/rel_notes/release_26_11.rst | 2 +
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
lib/ethdev/rte_ethdev.c | 52 ++++++++++++++++++--------
lib/ethdev/rte_ethdev.h | 2 +-
4 files changed, 39 insertions(+), 18 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 618b5acd2d..fe6860a47c 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -89,6 +89,8 @@ API Changes
* At port configuration time, the number of VMDq pools advertised by a driver is now used to
validate VMDq related Rx and Tx modes (``RTE_ETH_MQ_RX_VMDQ_FLAG``, ``RTE_ETH_MQ_TX_VMDQ_DCB``,
``RTE_ETH_MQ_TX_VMDQ_ONLY``).
+ * A check was added in ``rte_eth_dev_mac_addr_add`` to validate that the ``pool`` parameter is 0
+ when VMDq is not configured.
ABI Changes
diff --git a/drivers/net/cnxk/cnxk_ethdev_ops.c b/drivers/net/cnxk/cnxk_ethdev_ops.c
index 0ea3d7e89f..c002a93fe1 100644
--- a/drivers/net/cnxk/cnxk_ethdev_ops.c
+++ b/drivers/net/cnxk/cnxk_ethdev_ops.c
@@ -1240,7 +1240,6 @@ cnxk_nix_mc_addr_list_configure(struct rte_eth_dev *eth_dev, struct rte_ether_ad
/* Update address in NIC data structure */
rte_ether_addr_copy(&mc_addr_set[i], &data->mac_addrs[j]);
rte_ether_addr_copy(&mc_addr_set[i], &dev->dmac_addrs[j]);
- data->mac_pool_sel[j] = RTE_BIT64(0);
}
roc_nix_npc_promisc_ena_dis(nix, true);
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 9e305d98c1..f20cc514a3 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1677,7 +1677,7 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
{
struct rte_ether_addr *addr;
uint16_t i;
- uint32_t pool = 0;
+ uint32_t pool;
uint64_t pool_mask;
/* replay MAC address configuration including default MAC */
@@ -1685,9 +1685,11 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
if (dev->dev_ops->mac_addr_set != NULL)
dev->dev_ops->mac_addr_set(dev, addr);
else if (dev->dev_ops->mac_addr_add != NULL)
- dev->dev_ops->mac_addr_add(dev, addr, 0, pool);
+ dev->dev_ops->mac_addr_add(dev, addr, 0, 0);
if (dev->dev_ops->mac_addr_add != NULL) {
+ bool vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+
for (i = 1; i < dev_info->max_mac_addrs; i++) {
addr = &dev->data->mac_addrs[i];
@@ -1695,15 +1697,19 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
if (rte_is_zero_ether_addr(addr))
continue;
- pool = 0;
- pool_mask = dev->data->mac_pool_sel[i];
-
- do {
- if (pool_mask & UINT64_C(1))
- dev->dev_ops->mac_addr_add(dev, addr, i, pool);
- pool_mask >>= 1;
- pool++;
- } while (pool_mask);
+ if (!vmdq) {
+ dev->dev_ops->mac_addr_add(dev, addr, i, 0);
+ } else {
+ pool = 0;
+ pool_mask = dev->data->mac_pool_sel[i];
+
+ do {
+ if (pool_mask & UINT64_C(1))
+ dev->dev_ops->mac_addr_add(dev, addr, i, pool);
+ pool_mask >>= 1;
+ pool++;
+ } while (pool_mask);
+ }
}
}
}
@@ -5414,8 +5420,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
uint32_t pool)
{
struct rte_eth_dev *dev;
- int index;
uint64_t pool_mask;
+ bool vmdq;
+ int index;
int ret;
RTE_ETH_VALID_PORTID_OR_ERR_RET(port_id, -ENODEV);
@@ -5440,6 +5447,12 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
RTE_ETHDEV_LOG_LINE(ERR, "Pool ID must be 0-%d", RTE_ETH_64_POOLS - 1);
return -EINVAL;
}
+ vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+ if (!vmdq && pool != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Port %u: VMDq is not configured (pool %d)",
+ port_id, pool);
+ return -EINVAL;
+ }
index = eth_dev_get_mac_addr_index(port_id, addr);
if (index < 0) {
@@ -5450,6 +5463,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
return -ENOSPC;
}
} else {
+ if (!vmdq)
+ return 0;
+
pool_mask = dev->data->mac_pool_sel[index];
/* Check if both MAC address and pool is already there, and do nothing */
@@ -5464,8 +5480,10 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
/* Update address in NIC data structure */
rte_ether_addr_copy(addr, &dev->data->mac_addrs[index]);
- /* Update pool bitmap in NIC data structure */
- dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ if (vmdq) {
+ /* Update pool bitmap in NIC data structure */
+ dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ }
}
ret = eth_err(port_id, ret);
@@ -5510,8 +5528,10 @@ rte_eth_dev_mac_addr_remove(uint16_t port_id, struct rte_ether_addr *addr)
/* Update address in NIC data structure */
rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
- /* reset pool bitmap */
- dev->data->mac_pool_sel[index] = 0;
+ if ((dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ /* reset pool bitmap */
+ dev->data->mac_pool_sel[index] = 0;
+ }
rte_ethdev_trace_mac_addr_remove(port_id, addr);
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index ee400b386f..e2a5ba1549 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -4636,7 +4636,7 @@ int rte_eth_dev_priority_flow_ctrl_set(uint16_t port_id,
* - (-ENODEV) if *port* is invalid.
* - (-EIO) if device is removed.
* - (-ENOSPC) if no more MAC addresses can be added.
- * - (-EINVAL) if MAC address is invalid.
+ * - (-EINVAL) if MAC address is invalid or a non 0 pool was passed but VMDq is not enabled.
*/
int rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *mac_addr,
uint32_t pool);
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 03/10] ethdev: hide VMDq internal sizes
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
2026-07-23 12:41 ` [PATCH v5 01/10] ethdev: check VMDq availability David Marchand
2026-07-23 12:41 ` [PATCH v5 02/10] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 04/10] net/iavf: accept up to 32k unicast MAC addresses David Marchand
` (7 subsequent siblings)
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Andrew Rybchenko, Thomas Monjalon
Hide RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
in the driver API as those (ambiguous) macros are only a driver concern.
In practice, this is only used by the bnxt and ixgbe (+ clones) drivers.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v4:
- rebased,
Changes since v2:
- added an entry in release notes,
---
doc/guides/rel_notes/release_26_11.rst | 3 +++
lib/ethdev/ethdev_driver.h | 8 +++++++-
lib/ethdev/rte_ethdev.h | 6 ------
3 files changed, 10 insertions(+), 7 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index fe6860a47c..1a234d1be3 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -91,6 +91,9 @@ API Changes
``RTE_ETH_MQ_TX_VMDQ_ONLY``).
* A check was added in ``rte_eth_dev_mac_addr_add`` to validate that the ``pool`` parameter is 0
when VMDq is not configured.
+ * The ``RTE_ETH_NUM_RECEIVE_MAC_ADDR`` and ``RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY`` macros are VMDq
+ related and are sizes of internal arrays in ethdev that only drivers need to care about.
+ Those macros are moved to the driver only ethdev API.
ABI Changes
diff --git a/lib/ethdev/ethdev_driver.h b/lib/ethdev/ethdev_driver.h
index 0f336f9567..294f68504b 100644
--- a/lib/ethdev/ethdev_driver.h
+++ b/lib/ethdev/ethdev_driver.h
@@ -119,6 +119,12 @@ struct __rte_cache_aligned rte_eth_dev {
struct rte_eth_dev_sriov;
struct rte_eth_dev_owner;
+/* Definitions used for receive MAC address */
+#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
+
+/* Definitions used for unicast hash */
+#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
+
/**
* @internal
* The data part, with no function pointers, associated with each Ethernet
@@ -153,7 +159,7 @@ struct __rte_cache_aligned rte_eth_dev_data {
* The first entry (index zero) is the default address.
*/
struct rte_ether_addr *mac_addrs;
- /** Bitmap associating MAC addresses to pools */
+ /** Bitmap associating MAC addresses to VMDq pools */
uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
/**
* Device Ethernet MAC addresses of hash filtering.
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index e2a5ba1549..bee47549cb 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -903,12 +903,6 @@ rte_eth_rss_hf_refine(uint64_t rss_hf)
#define RTE_ETH_VLAN_ID_MAX 0x0FFF /**< VLAN ID is in lower 12 bits*/
/**@}*/
-/* Definitions used for receive MAC address */
-#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
-
-/* Definitions used for unicast hash */
-#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
-
/**@{@name VMDq Rx mode
* @see rte_eth_vmdq_rx_conf.rx_mode
*/
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 04/10] net/iavf: accept up to 32k unicast MAC addresses
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (2 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 03/10] ethdev: hide VMDq internal sizes David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 05/10] net/iavf: fix duplicate MAC addresses install David Marchand
` (6 subsequent siblings)
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Vladimir Medvedkin
E810 hardware provides 32k switch lookups.
Thanks to this, it is possible to allow a lot more secondary mac
addresses than what is possible today.
In practice, the maximum number of macs available per port may be lower
and depends on usage by other (trusted?) VFs on the same PF.
There is no way to figure out this limit but to try adding a mac address
and get an error from the PF driver.
Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
accordingly.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v4:
- rebased,
Changes since v2:
- added an entry in release notes,
- removed unneeded temp variable,
Changes since v1:
- fixed buffer overflow on mailbox messages during port restart/VF reset,
---
doc/guides/rel_notes/release_26_11.rst | 5 ++
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 12 +--
drivers/net/intel/iavf/iavf_vchnl.c | 113 ++++++++++++++++++-------
4 files changed, 95 insertions(+), 40 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 1a234d1be3..333bf7525c 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,11 @@ New Features
Also, make sure to start the actual text at the margin.
=======================================================
+* **Updated Intel iavf driver.**
+
+ * Increased the maximum number of secondary unicast MAC addresses from 64 to 32k.
+ This increases a VF port memory footprint by ~192kB.
+
Removed Items
-------------
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 293adaf6c9..47cd1d6311 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -31,7 +31,8 @@
#define IAVF_IRQ_MAP_NUM_PER_BUF 128
#define IAVF_RXTX_QUEUE_CHUNKS_NUM 2
-#define IAVF_NUM_MACADDR_MAX 64
+#define IAVF_UC_MACADDR_MAX 32768
+#define IAVF_MC_MACADDR_MAX 64
#define IAVF_DEV_WATCHDOG_PERIOD 2000 /* microseconds, set 0 to disable*/
@@ -255,7 +256,7 @@ struct iavf_info {
uint32_t link_speed;
/* Multicast addrs */
- struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
+ struct rte_ether_addr mc_addrs[IAVF_MC_MACADDR_MAX];
uint16_t mc_addrs_num; /* Multicast mac addresses number */
struct iavf_vsi vsi;
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index d601ec3b6a..93305a2bb1 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -394,10 +394,10 @@ iavf_set_mc_addr_list(struct rte_eth_dev *dev,
IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
int err, ret;
- if (mc_addrs_num > IAVF_NUM_MACADDR_MAX) {
+ if (mc_addrs_num > IAVF_MC_MACADDR_MAX) {
PMD_DRV_LOG(ERR,
"can't add more than a limited number (%u) of addresses.",
- (uint32_t)IAVF_NUM_MACADDR_MAX);
+ (uint32_t)IAVF_MC_MACADDR_MAX);
return -EINVAL;
}
@@ -1159,7 +1159,7 @@ iavf_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->hash_key_size = vf->vf_res->rss_key_size;
dev_info->reta_size = vf->vf_res->rss_lut_size;
dev_info->flow_type_rss_offloads = IAVF_RSS_OFFLOAD_ALL;
- dev_info->max_mac_addrs = IAVF_NUM_MACADDR_MAX;
+ dev_info->max_mac_addrs = IAVF_UC_MACADDR_MAX;
dev_info->dev_capa =
RTE_ETH_DEV_CAPA_RUNTIME_RX_QUEUE_SETUP |
RTE_ETH_DEV_CAPA_RUNTIME_TX_QUEUE_SETUP;
@@ -3055,12 +3055,12 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
iavf_set_default_ptype_table(eth_dev);
/* copy mac addr */
- eth_dev->data->mac_addrs = rte_zmalloc(
- "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX, 0);
+ eth_dev->data->mac_addrs = rte_calloc("iavf_mac", IAVF_UC_MACADDR_MAX,
+ RTE_ETHER_ADDR_LEN, 0);
if (!eth_dev->data->mac_addrs) {
PMD_INIT_LOG(ERR, "Failed to allocate %d bytes needed to"
" store MAC addresses",
- RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX);
+ RTE_ETHER_ADDR_LEN * IAVF_UC_MACADDR_MAX);
ret = -ENOMEM;
goto init_vf_err;
}
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index f8fffa7802..c5328a6df4 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1653,49 +1653,98 @@ iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num)
return 0;
}
-void
-iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+static int
+iavf_add_del_uc_addr_bulk(struct iavf_adapter *adapter, struct rte_ether_addr *addrs,
+ uint32_t nb_addrs, bool add)
{
+#define IAVF_ETH_ADDR_PER_REQ \
+ ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
+ sizeof(struct virtchnl_ether_addr))
struct {
struct virtchnl_ether_addr_list list;
- struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
- } list_req = {0};
- struct virtchnl_ether_addr_list *list = &list_req.list;
+ struct virtchnl_ether_addr addr[IAVF_ETH_ADDR_PER_REQ];
+ } cmd_buffer;
+#undef IAVF_ETH_ADDR_PER_REQ
+ struct virtchnl_ether_addr_list *list = &cmd_buffer.list;
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
- struct iavf_cmd_info args = {0};
- int err, i;
- size_t buf_len;
- for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
- struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
- struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
+ for (uint32_t i = 0; i < nb_addrs; i++) {
+ struct iavf_cmd_info args;
+ uint32_t batch;
+ int err;
- /* ignore empty addresses */
- if (rte_is_zero_ether_addr(addr))
- continue;
+ batch = i % RTE_DIM(cmd_buffer.addr);
+
+ if (batch == 0) {
+ memset(&cmd_buffer, 0, sizeof(cmd_buffer));
+ list->vsi_id = vf->vsi_res->vsi_id;
+ list->num_elements = 0;
+ }
+
+ rte_memcpy(list->list[batch].addr, addrs[i].addr_bytes,
+ sizeof(list->list[batch].addr));
+ list->list[batch].type = VIRTCHNL_ETHER_ADDR_EXTRA;
list->num_elements++;
- memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
- vc_addr->type = (list->num_elements == 1) ?
- VIRTCHNL_ETHER_ADDR_PRIMARY :
- VIRTCHNL_ETHER_ADDR_EXTRA;
+ if (batch != RTE_DIM(cmd_buffer.addr) - 1 && i != nb_addrs - 1)
+ continue;
+
+ memset(&args, 0, sizeof(args));
+ args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
+ args.in_args = (uint8_t *)list;
+ args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
+ sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ args.out_buffer = msg_buf;
+ args.out_size = IAVF_AQ_BUF_SZ;
+ err = iavf_execute_vf_cmd_safe(adapter, &args);
+ if (err != 0) {
+ PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
+ return err;
+ }
+
+ PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
}
- /* for some reason PF side checks for buffer being too big, so adjust it down */
- buf_len = sizeof(struct virtchnl_ether_addr_list) +
- sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ return 0;
+}
- list->vsi_id = vf->vsi_res->vsi_id;
- args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
- args.in_args = (uint8_t *)list;
- args.in_args_size = buf_len;
- args.out_buffer = msg_buf;
- args.out_size = IAVF_AQ_BUF_SZ;
- err = iavf_execute_vf_cmd_safe(adapter, &args);
- if (err)
- PMD_DRV_LOG(ERR, "fail to execute command %s",
- add ? "OP_ADD_ETHER_ADDRESS" : "OP_DEL_ETHER_ADDRESS");
+void
+iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+{
+ int start = -1;
+ int i;
+
+ /* Handle primary address (index 0) separately */
+ if (!rte_is_zero_ether_addr(&adapter->dev_data->mac_addrs[0]))
+ iavf_add_del_eth_addr(adapter, &adapter->dev_data->mac_addrs[0], add,
+ VIRTCHNL_ETHER_ADDR_PRIMARY);
+
+ /* Process secondary addresses in contiguous blocks */
+ for (i = 1; i < IAVF_UC_MACADDR_MAX; i++) {
+ struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
+
+ if (!rte_is_zero_ether_addr(addr)) {
+ if (start == -1)
+ start = i;
+ continue;
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ start = -1;
+ }
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ }
}
int
@@ -2286,7 +2335,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
- (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
+ (IAVF_MC_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
struct virtchnl_ether_addr_list *list;
struct iavf_cmd_info args;
uint32_t i;
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 05/10] net/iavf: fix duplicate MAC addresses install
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (3 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 04/10] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 06/10] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (5 subsequent siblings)
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, stable, Vladimir Medvedkin, Bruce Richardson
On port restart, all MAC addresses get pushed *twice* to the hardware,
once by the driver and once by the eth_dev_mac_restore() in ethdev.
On the other hand, MAC address filters are reset in the hardware
by the PF only when a VF reset is triggered.
Strictly speaking, the mac restore on port (re)start is unneeded,
if no VF reset happened, so we can announce to ethdev that no mac
restoration is needed via a get_restore_flags callback.
Then, move the mac restoration to the VF reset handler.
Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux driver")
Cc: stable@dpdk.org
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v4:
- moved mac restoration in iavf_post_reset_reconfig,
---
drivers/net/intel/iavf/iavf_ethdev.c | 32 ++++++++++++++++++++--------
1 file changed, 23 insertions(+), 9 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index 93305a2bb1..00e3a7c412 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -134,6 +134,8 @@ static int iavf_dev_vlan_filter_set(struct rte_eth_dev *dev,
static int iavf_vlan_tpid_set(struct rte_eth_dev *dev,
enum rte_vlan_type vlan_type, uint16_t tpid);
static int iavf_dev_vlan_offload_set(struct rte_eth_dev *dev, int mask);
+static uint64_t iavf_get_restore_flags(struct rte_eth_dev *dev,
+ enum rte_eth_dev_operation op);
static int iavf_dev_rss_reta_update(struct rte_eth_dev *dev,
struct rte_eth_rss_reta_entry64 *reta_conf,
uint16_t reta_size);
@@ -264,6 +266,7 @@ static const struct eth_dev_ops iavf_eth_dev_ops = {
.tx_done_cleanup = iavf_dev_tx_done_cleanup,
.get_monitor_addr = iavf_get_monitor_addr,
.tm_ops_get = iavf_tm_ops_get,
+ .get_restore_flags = iavf_get_restore_flags,
};
static int
@@ -284,6 +287,13 @@ iavf_tm_ops_get(struct rte_eth_dev *dev,
return 0;
}
+static uint64_t
+iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
+ __rte_unused enum rte_eth_dev_operation op)
+{
+ return RTE_ETH_RESTORE_ALL & ~RTE_ETH_RESTORE_MAC_ADDR;
+}
+
__rte_unused
static int
iavf_vfr_inprogress(struct iavf_hw *hw)
@@ -1074,15 +1084,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
rte_intr_enable(intr_handle);
}
- /* Set all mac addrs */
- iavf_add_del_all_mac_addr(adapter, true);
-
- if (!adapter->mac_primary_set)
- adapter->mac_primary_set = true;
-
- /* Set all multicast addresses */
- iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num,
- true);
+ if (!adapter->mac_primary_set) {
+ if (iavf_add_del_eth_addr(adapter, &dev->data->mac_addrs[0], true,
+ VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
+ PMD_DRV_LOG(ERR, "failed to add primary MAC:" RTE_ETHER_ADDR_PRT_FMT,
+ RTE_ETHER_ADDR_BYTES(&dev->data->mac_addrs[0]));
+ else
+ adapter->mac_primary_set = true;
+ }
rte_spinlock_init(&vf->phc_time_aq_lock);
@@ -3361,6 +3370,11 @@ iavf_post_reset_reconfig(struct rte_eth_dev *dev)
int ret, status = 0;
bool allmulti = false, allunicast = false;
struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
+ struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data->dev_private);
+
+ /* After a VF reset, all MAC addresses got flushed, restore them. */
+ iavf_add_del_all_mac_addr(adapter, true);
+ iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num, true);
/* Restore pre-reset unicast promiscuous and multicast promiscuous states */
if (dev->data->promiscuous)
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 06/10] net/mlx5: remove MAC addresses flush helper on Linux
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (4 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 05/10] net/iavf: fix duplicate MAC addresses install David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 07/10] net/mlx5: remove redundant MAC address index checks David Marchand
` (4 subsequent siblings)
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
This helper is exposing internals of net/mlx5 for no good reason.
All this code does is calling the remove helper.
Walk through the list in Linux implementation like the Windows
implementation.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 40 -----------------------------
drivers/common/mlx5/linux/mlx5_nl.h | 5 ----
drivers/net/mlx5/linux/mlx5_os.c | 13 +++++++---
3 files changed, 10 insertions(+), 48 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 8b19838a7e..85736738ac 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -825,46 +825,6 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
}
-/**
- * Flush all added MAC addresses.
- *
- * @param[in] nlsk_fd
- * Netlink socket file descriptor.
- * @param[in] iface_idx
- * Net device interface index.
- * @param[in] mac_addrs
- * Mac addresses array to flush.
- * @param n
- * @p mac_addrs array size.
- * @param mac_own
- * BITFIELD_DECLARE array to store the mac.
- * @param vf
- * Flag for a VF device.
- */
-RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_flush)
-void
-mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own, bool vf)
-{
- int i;
-
- if (n <= 0 || n > MLX5_MAX_MAC_ADDRESSES)
- return;
-
- for (i = n - 1; i >= 0; --i) {
- struct rte_ether_addr *m = &mac_addrs[i];
-
- if (BITFIELD_ISSET(mac_own, i)) {
- if (vf)
- mlx5_nl_mac_addr_remove(nlsk_fd,
- iface_idx,
- m, i);
- BITFIELD_RESET(mac_own, i);
- }
- }
-}
-
/**
* Enable promiscuous / all multicast mode through Netlink.
*
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 3f79a73c85..256ed7e2b7 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -65,11 +65,6 @@ __rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
__rte_internal
-void mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own,
- bool vf);
-__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
int mlx5_nl_allmulti(int nlsk_fd, unsigned int iface_idx, int enable);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index adc5878296..9fd366b10a 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3495,10 +3495,17 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
{
struct mlx5_priv *priv = dev->data->dev_private;
const int vf = priv->sh->dev_cap.vf;
+ int i;
- mlx5_nl_mac_addr_flush(priv->nl_socket_route, mlx5_ifindex(dev),
- dev->data->mac_addrs,
- MLX5_MAX_MAC_ADDRESSES, priv->mac_own, vf);
+ for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (vf)
+ mlx5_nl_mac_addr_remove(priv->nl_socket_route,
+ mlx5_ifindex(dev),
+ &dev->data->mac_addrs[i], i);
+ BITFIELD_RESET(priv->mac_own, i);
+ }
+ }
}
static bool
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 07/10] net/mlx5: remove redundant MAC address index checks
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (5 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 06/10] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 08/10] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
` (3 subsequent siblings)
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
On the net/mlx5 side, mlx5_mac.c validates that any MAC address and its
index is valid before calling the OS specific helpers.
So those OS helpers do not have to validate again the index.
Cascading this consideration, validating the MAC index against
MLX5_MAX_MAC_ADDRESSES in common code is also redundant.
The common code only deals with netlink, remove any index concern.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 17 ++---------------
drivers/common/mlx5/linux/mlx5_nl.h | 4 ++--
drivers/net/mlx5/linux/mlx5_os.c | 9 ++++-----
3 files changed, 8 insertions(+), 22 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 85736738ac..ceb504f84d 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -719,8 +719,6 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to register.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -728,16 +726,11 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_add)
int
mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
int ret;
ret = mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 1);
- if (!ret) {
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
- }
if (ret == -EEXIST)
return 0;
return ret;
@@ -752,8 +745,6 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to remove.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -761,12 +752,8 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_remove)
int
mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
-
return mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 0);
}
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 256ed7e2b7..500198b654 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -57,10 +57,10 @@ __rte_internal
int mlx5_nl_init(int protocol, int groups);
__rte_internal
int mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 9fd366b10a..1b241ba9d2 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3382,9 +3382,8 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[index], index);
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ &dev->data->mac_addrs[index]);
+ BITFIELD_RESET(priv->mac_own, index);
}
/**
@@ -3411,7 +3410,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
if (vf)
ret = mlx5_nl_mac_addr_add(priv->nl_socket_route,
mlx5_ifindex(dev),
- mac, index);
+ mac);
if (!ret)
BITFIELD_SET(priv->mac_own, index);
@@ -3502,7 +3501,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[i], i);
+ &dev->data->mac_addrs[i]);
BITFIELD_RESET(priv->mac_own, i);
}
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 08/10] net/mlx5: pass maximum number of unicast MAC to common code
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (6 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 07/10] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 17:35 ` Stephen Hemminger
2026-07-23 12:41 ` [PATCH v5 09/10] net/mlx5: use bitset for tracking MAC addresses David Marchand
` (2 subsequent siblings)
10 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Isolate how the MAC addresses array is walked through in the common code
by passing the max index at which a unicast MAC address is stored in
dev->data->mac_addrs[].
In the sync callback, the size of the array allocated on the stack is
known by the caller, treat the mac_n field as an input parameter too.
With this change, only net/mlx5 knows about the max number of
unicast/multicast MAC addresses.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 33 ++++++++++++++++++-----------
drivers/common/mlx5/linux/mlx5_nl.h | 2 +-
drivers/common/mlx5/mlx5_common.h | 8 -------
drivers/net/mlx5/linux/mlx5_os.c | 1 +
drivers/net/mlx5/mlx5.h | 8 +++++++
5 files changed, 31 insertions(+), 21 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index ceb504f84d..152b4cdda1 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -468,19 +468,22 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
struct mlx5_nl_mac_addr *data = arg;
struct ndmsg *r = NLMSG_DATA(nh);
struct rtattr *attribute;
+ int mac_n = 0;
int len;
+ int ret;
len = nh->nlmsg_len - NLMSG_LENGTH(sizeof(*r));
for (attribute = MLX5_NDA_RTA(r);
RTA_OK(attribute, len);
attribute = RTA_NEXT(attribute, len)) {
if (attribute->rta_type == NDA_LLADDR) {
- if (data->mac_n == MLX5_MAX_MAC_ADDRESSES) {
+ if (mac_n == data->mac_n) {
DRV_LOG(WARNING,
"not enough room to finalize the"
" request");
rte_errno = ENOMEM;
- return -rte_errno;
+ ret = -rte_errno;
+ goto out;
}
#ifdef RTE_PMD_MLX5_DEBUG
char m[RTE_ETHER_ADDR_FMT_SIZE];
@@ -489,11 +492,15 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
RTA_DATA(attribute));
DRV_LOG(DEBUG, "bridge MAC address %s", m);
#endif
- memcpy(&(*data->mac)[data->mac_n++],
+ memcpy(&(*data->mac)[mac_n++],
RTA_DATA(attribute), RTE_ETHER_ADDR_LEN);
}
}
- return 0;
+ ret = 0;
+
+out:
+ data->mac_n = mac_n;
+ return ret;
}
/**
@@ -505,9 +512,9 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
* Net device interface index.
* @param mac[out]
* Pointer to the array table of MAC addresses to fill.
- * Its size should be of MLX5_MAX_MAC_ADDRESSES.
- * @param mac_n[out]
- * Number of entries filled in MAC array.
+ * @param mac_n[in,out]
+ * Size of the MAC array on input.
+ * Number of entries filled in MAC array on output.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -532,7 +539,7 @@ mlx5_nl_mac_addr_list(int nlsk_fd, unsigned int iface_idx,
};
struct mlx5_nl_mac_addr data = {
.mac = mac,
- .mac_n = 0,
+ .mac_n = *mac_n,
};
uint32_t sn = MLX5_NL_SN_GENERATE;
int ret;
@@ -766,16 +773,18 @@ mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac_addrs
* Mac addresses array to sync.
+ * @param uc_n
+ * Number of UC entries in @p mac_addrs.
* @param n
* @p mac_addrs array size.
*/
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_sync)
void
mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n)
+ struct rte_ether_addr *mac_addrs, int uc_n, int n)
{
struct rte_ether_addr macs[n];
- int macs_n = 0;
+ int macs_n = n;
int i;
int ret;
@@ -794,7 +803,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
continue;
if (rte_is_multicast_ether_addr(&macs[i])) {
/* Find the first entry available. */
- for (j = MLX5_MAX_UC_MAC_ADDRESSES; j != n; ++j) {
+ for (j = uc_n; j != n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
@@ -802,7 +811,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
} else {
/* Find the first entry available. */
- for (j = 0; j != MLX5_MAX_UC_MAC_ADDRESSES; ++j) {
+ for (j = 0; j != uc_n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 500198b654..d71eadeffb 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -63,7 +63,7 @@ int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n);
+ struct rte_ether_addr *mac_addrs, int uc_n, int n);
__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 3767020823..8aed3dbabe 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -160,14 +160,6 @@ enum {
PCI_DEVICE_ID_MELLANOX_CONNECTX10C2C = 0x2101,
};
-/* Maximum number of simultaneous unicast MAC addresses. */
-#define MLX5_MAX_UC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous Multicast MAC addresses. */
-#define MLX5_MAX_MC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous MAC addresses. */
-#define MLX5_MAX_MAC_ADDRESSES \
- (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
-
/* Recognized Infiniband device physical port name types. */
enum mlx5_nl_phys_port_name_type {
MLX5_PHYS_PORT_NAME_TYPE_NOTSET = 0, /* Not set. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 1b241ba9d2..5b6e45df2a 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -1762,6 +1762,7 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
+ MLX5_MAX_UC_MAC_ADDRESSES,
MLX5_MAX_MAC_ADDRESSES);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 27e6f4e31a..3ad8bad02f 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -82,6 +82,14 @@
/* Maximum allowed MTU to be reported whenever PMD cannot query it from OS. */
#define MLX5_ETH_MAX_MTU (9978)
+/* Maximum number of simultaneous unicast MAC addresses. */
+#define MLX5_MAX_UC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous Multicast MAC addresses. */
+#define MLX5_MAX_MC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous MAC addresses. */
+#define MLX5_MAX_MAC_ADDRESSES \
+ (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
+
enum mlx5_ipool_index {
#if defined(HAVE_IBV_FLOW_DV_SUPPORT) || !defined(HAVE_INFINIBAND_VERBS_H)
MLX5_IPOOL_DECAP_ENCAP = 0, /* Pool for encap/decap resource. */
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 09/10] net/mlx5: use bitset for tracking MAC addresses
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (7 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 08/10] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 10/10] net/mlx5: accept more unicast " David Marchand
2026-07-27 7:20 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
EAL provides bitset that does the same as this set of mlx5 macros.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/mlx5_common.h | 16 ----------------
drivers/net/mlx5/linux/mlx5_os.c | 8 ++++----
drivers/net/mlx5/mlx5.h | 3 ++-
drivers/net/mlx5/mlx5_trigger.c | 2 +-
drivers/net/mlx5/windows/mlx5_os.c | 8 ++++----
5 files changed, 11 insertions(+), 26 deletions(-)
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 8aed3dbabe..8b8f20e2f4 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -30,22 +30,6 @@
#define MLX5_PCI_DRIVER_NAME "mlx5_pci"
#define MLX5_AUXILIARY_DRIVER_NAME "mlx5_auxiliary"
-/* Bit-field manipulation. */
-#define BITFIELD_DECLARE(bf, type, size) \
- type bf[(((size_t)(size) / (sizeof(type) * CHAR_BIT)) + \
- !!((size_t)(size) % (sizeof(type) * CHAR_BIT)))]
-#define BITFIELD_DEFINE(bf, type, size) \
- BITFIELD_DECLARE((bf), type, (size)) = { 0 }
-#define BITFIELD_SET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] |= \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_RESET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] &= \
- ~((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_ISSET(bf, b) \
- !!(((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] & \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT)))))
-
/*
* Helper macros to work around __VA_ARGS__ limitations in a C99 compliant
* manner.
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 5b6e45df2a..1cad6e1091 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3384,7 +3384,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[index]);
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -3413,7 +3413,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
mlx5_ifindex(dev),
mac);
if (!ret)
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return ret;
}
@@ -3498,12 +3498,12 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[i]);
- BITFIELD_RESET(priv->mac_own, i);
+ rte_bitset_clear(priv->mac_own, i);
}
}
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 3ad8bad02f..f7ba8df108 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -14,6 +14,7 @@
#include <rte_pci.h>
#include <rte_ether.h>
+#include <rte_bitset.h>
#include <ethdev_driver.h>
#include <rte_rwlock.h>
#include <rte_interrupts.h>
@@ -2018,7 +2019,7 @@ struct mlx5_priv {
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- BITFIELD_DECLARE(mac_own, uint64_t, MLX5_MAX_MAC_ADDRESSES);
+ RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
/* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 25847c8ba2..e41e4643e0 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1907,7 +1907,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
/* Add flows for unicast and multicast mac addresses added by API. */
if (!memcmp(mac, &cmp, sizeof(*mac)) ||
- !BITFIELD_ISSET(priv->mac_own, i) ||
+ !rte_bitset_test(priv->mac_own, i) ||
(dev->data->all_multicast && rte_is_multicast_ether_addr(mac)))
continue;
memcpy(&unicast.hdr.dst_addr.addr_bytes,
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index 9acfa8ec84..15de5c22a9 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -699,8 +699,8 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i))
- BITFIELD_RESET(priv->mac_own, i);
+ if (rte_bitset_test(priv->mac_own, i))
+ rte_bitset_clear(priv->mac_own, i);
}
}
@@ -719,7 +719,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
struct mlx5_priv *priv = dev->data->dev_private;
if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -757,7 +757,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
return -ENOTSUP;
}
/* Mark this MAC address as owned by the PMD */
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return 0;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v5 10/10] net/mlx5: accept more unicast MAC addresses
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (8 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 09/10] net/mlx5: use bitset for tracking MAC addresses David Marchand
@ 2026-07-23 12:41 ` David Marchand
2026-07-27 7:20 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-23 12:41 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Starting firmware version 22.49.1014, the number of mac addresses
per VF is not capped to 128 anymore.
The value can be increased via devlink:
$ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
cmode driverinit
$ devlink dev reload pci/0000:3b:00.2
On the DPDK side, we must retrieve the maximum number of unicast
and multicast addresses supported with a query to the firmware.
Then, dynamically allocate the mac addresses arrays and report the
limit instead of the previous hardcoded value.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v4:
- added RN update,
- fixed types of fields added to mlx5_hca_attr and mlx5_dev_cap,
- used RTE_BIT32,
- fixed mlx5_nl_mac_addr_sync inverted arguments,
---
doc/guides/rel_notes/release_26_11.rst | 5 +++
drivers/common/mlx5/mlx5_devx_cmds.c | 4 +++
drivers/common/mlx5/mlx5_devx_cmds.h | 2 ++
drivers/net/mlx5/linux/mlx5_os.c | 42 ++++++++++++++++++++------
drivers/net/mlx5/mlx5.c | 9 ++----
drivers/net/mlx5/mlx5.h | 8 +++--
drivers/net/mlx5/mlx5_ethdev.c | 2 +-
drivers/net/mlx5/mlx5_mac.c | 22 +++++++++-----
drivers/net/mlx5/mlx5_trigger.c | 10 +++---
drivers/net/mlx5/windows/mlx5_os.c | 40 +++++++++++++++++++-----
10 files changed, 104 insertions(+), 40 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 333bf7525c..c78ca9b63b 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -60,6 +60,11 @@ New Features
* Increased the maximum number of secondary unicast MAC addresses from 64 to 32k.
This increases a VF port memory footprint by ~192kB.
+* **Updated NVIDIA mlx5 ethernet driver.**
+
+ * Increased the maximum number of secondary unicast MAC addresses from 128 to up to 4096
+ (depending on devlink configuration on the associated kernel netdevice).
+
Removed Items
-------------
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.c b/drivers/common/mlx5/mlx5_devx_cmds.c
index 140b057ab4..e5d9c92779 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.c
+++ b/drivers/common/mlx5/mlx5_devx_cmds.c
@@ -1110,6 +1110,10 @@ mlx5_devx_cmd_query_hca_attr(void *ctx,
attr->log_max_pd = MLX5_GET(cmd_hca_cap, hcattr, log_max_pd);
attr->log_max_srq = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq);
attr->log_max_srq_sz = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq_sz);
+ attr->log_max_current_uc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_uc_list);
+ attr->log_max_current_mc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_mc_list);
attr->reg_c_preserve =
MLX5_GET(cmd_hca_cap, hcattr, reg_c_preserve);
attr->mmo_regex_qp_en = MLX5_GET(cmd_hca_cap, hcattr, regexp_mmo_qp);
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.h b/drivers/common/mlx5/mlx5_devx_cmds.h
index 90beb2e9e6..504b4a4f64 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.h
+++ b/drivers/common/mlx5/mlx5_devx_cmds.h
@@ -356,6 +356,8 @@ struct mlx5_hca_attr {
uint8_t tx_sw_owner_v2:1;
uint8_t esw_sw_owner:1;
uint8_t esw_sw_owner_v2:1;
+ uint8_t log_max_current_uc_list:5;
+ uint8_t log_max_current_mc_list:5;
};
/* LAG Context. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 1cad6e1091..13ddabd465 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -389,6 +389,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
sh->dev_cap.esw_info.regc_mask = 0;
#endif
sh->dev_cap.esw_info.is_set = 1;
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -1464,6 +1474,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
/* Some internal functions rely on Netlink sockets, open them now. */
priv->nl_socket_rdma = nl_rdma;
priv->nl_socket_route = mlx5_nl_init(NETLINK_ROUTE, 0);
@@ -1762,8 +1788,8 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
- MLX5_MAX_UC_MAC_ADDRESSES,
- MLX5_MAX_MAC_ADDRESSES);
+ sh->dev_cap.max_uc_mac_addrs,
+ sh->dev_cap.max_mac_addrs);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
TAILQ_INIT(&priv->flow_meters);
@@ -1963,17 +1989,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_flex_item_port_cleanup(eth_dev);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ eth_dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
if (eth_dev != NULL)
eth_dev->data->dev_private = NULL;
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
if (nl_rdma >= 0)
@@ -3497,7 +3521,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
const int vf = priv->sh->dev_cap.vf;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
diff --git a/drivers/net/mlx5/mlx5.c b/drivers/net/mlx5/mlx5.c
index 3fa38694e5..926cb88047 100644
--- a/drivers/net/mlx5/mlx5.c
+++ b/drivers/net/mlx5/mlx5.c
@@ -2546,6 +2546,9 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_list_destroy(priv->hrxqs);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
sh->port[priv->dev_port - 1].nl_ih_port_id = RTE_MAX_ETHPORTS;
/*
* The interrupt handler port id must be reset before priv is reset
@@ -2580,12 +2583,6 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_flow_pools_destroy(priv);
memset(priv, 0, sizeof(*priv));
priv->domain_id = RTE_ETH_DEV_SWITCH_DOMAIN_ID_INVALID;
- /*
- * Reset mac_addrs to NULL such that it is not freed as part of
- * rte_eth_dev_release_port(). mac_addrs is part of dev_private so
- * it is freed when dev_private is freed.
- */
- dev->data->mac_addrs = NULL;
return 0;
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index f7ba8df108..626cec3614 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -217,6 +217,9 @@ struct mlx5_dev_cap {
} mprq; /* Capability for Multi-Packet RQ. */
char fw_ver[64]; /* Firmware version of this device. */
struct flow_hw_port_info esw_info; /* E-switch manager reg_c0. */
+ uint32_t max_uc_mac_addrs; /* Maximum unicast MAC addresses. */
+ uint32_t max_mc_mac_addrs; /* Maximum multicast MAC addresses. */
+ uint32_t max_mac_addrs; /* Total maximum MAC addresses. */
};
#define MLX5_MPESW_PORT_INVALID (-1)
@@ -2018,9 +2021,8 @@ struct mlx5_priv {
struct mlx5_dev_ctx_shared *sh; /* Shared device context. */
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
- struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
- /* Bit-field of MAC addresses owned by the PMD. */
+ struct rte_ether_addr *mac; /* MAC addresses. */
+ uint64_t *mac_own; /* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
/* Device properties. */
diff --git a/drivers/net/mlx5/mlx5_ethdev.c b/drivers/net/mlx5/mlx5_ethdev.c
index 8160d10e7e..306c1cd734 100644
--- a/drivers/net/mlx5/mlx5_ethdev.c
+++ b/drivers/net/mlx5/mlx5_ethdev.c
@@ -392,7 +392,7 @@ mlx5_dev_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *info)
max = RTE_MIN(max, (unsigned int)UINT16_MAX);
info->max_rx_queues = max;
info->max_tx_queues = max;
- info->max_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ info->max_mac_addrs = priv->sh->dev_cap.max_uc_mac_addrs;
info->rx_queue_offload_capa = mlx5_get_rx_queue_offloads(dev);
info->rx_seg_capa.max_nseg = MLX5_MAX_RXQ_NSEG;
info->rx_seg_capa.multi_pools = !priv->config.mprq.enabled;
diff --git a/drivers/net/mlx5/mlx5_mac.c b/drivers/net/mlx5/mlx5_mac.c
index 0e5d2be530..2ce7cfd407 100644
--- a/drivers/net/mlx5/mlx5_mac.c
+++ b/drivers/net/mlx5/mlx5_mac.c
@@ -36,7 +36,9 @@ mlx5_internal_mac_addr_remove(struct rte_eth_dev *dev,
uint32_t index,
struct rte_ether_addr *addr)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ struct mlx5_priv *priv = dev->data->dev_private;
+
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(&dev->data->mac_addrs[index]))
return false;
mlx5_os_mac_addr_remove(dev, index);
@@ -63,16 +65,17 @@ static int
mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
unsigned int i;
int ret;
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(mac)) {
rte_errno = EINVAL;
return -rte_errno;
}
/* First, make sure this address isn't already configured. */
- for (i = 0; (i != MLX5_MAX_MAC_ADDRESSES); ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
/* Skip this index, it's going to be reconfigured. */
if (i == index)
continue;
@@ -101,10 +104,11 @@ mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
void
mlx5_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
struct rte_ether_addr addr = { 0 };
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES)
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs)
return;
if (mlx5_internal_mac_addr_remove(dev, index, &addr)) {
ret = mlx5_traffic_mac_remove(dev, &addr);
@@ -133,9 +137,10 @@ int
mlx5_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index, uint32_t vmdq __rte_unused)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES) {
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs) {
rte_errno = EINVAL;
return -rte_errno;
}
@@ -217,16 +222,17 @@ int
mlx5_set_mc_addr_list(struct rte_eth_dev *dev,
struct rte_ether_addr *mc_addr_set, uint32_t nb_mc_addr)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
uint32_t i;
int ret;
- if (nb_mc_addr >= MLX5_MAX_MC_MAC_ADDRESSES) {
+ if (nb_mc_addr >= priv->sh->dev_cap.max_mc_mac_addrs) {
rte_errno = ENOSPC;
return -rte_errno;
}
- for (i = MLX5_MAX_UC_MAC_ADDRESSES; i != MLX5_MAX_MAC_ADDRESSES; ++i)
+ for (i = priv->sh->dev_cap.max_uc_mac_addrs; i != priv->sh->dev_cap.max_mac_addrs; ++i)
mlx5_internal_mac_addr_remove(dev, i, NULL);
- i = MLX5_MAX_UC_MAC_ADDRESSES;
+ i = priv->sh->dev_cap.max_uc_mac_addrs;
while (nb_mc_addr--) {
ret = mlx5_internal_mac_addr_add(dev, mc_addr_set++, i++);
if (ret)
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index e41e4643e0..e535e3e5be 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1902,7 +1902,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
}
}
/* Add MAC address flows. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
/* Add flows for unicast and multicast mac addresses added by API. */
@@ -2172,7 +2172,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
return 0;
/* Add all unicast DMAC flow rules with new VLAN attached. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2189,7 +2189,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
* Removing after creating VLAN rules so that traffic "gap" is not introduced.
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2227,7 +2227,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
* Recreating first to ensure no traffic "gap".
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2240,7 +2240,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
}
/* Remove all unicast DMAC flow rules with this VLAN. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index 15de5c22a9..3dfc85aa09 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -261,6 +261,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
MLX5_GET(initial_seg, pv_iseg, fw_rev_subminor));
DRV_LOG(DEBUG, "Packet pacing is not supported.");
mlx5_rt_timestamp_config(sh, hca_attr);
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -396,6 +406,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
priv->mp_id.port_id = port_id;
strlcpy(priv->mp_id.name, MLX5_MP_NAME, RTE_MP_MAX_NAME_LEN);
priv->representor = !!switch_info->representor;
@@ -612,17 +638,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_l3t_destroy(priv->mtr_profile_tbl);
if (own_domain_id)
claim_zero(rte_eth_switch_domain_free(priv->domain_id));
+ mlx5_free(priv->mac);
+ eth_dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
if (eth_dev != NULL)
eth_dev->data->dev_private = NULL;
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
MLX5_ASSERT(err > 0);
@@ -698,7 +722,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
struct mlx5_priv *priv = dev->data->dev_private;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i))
rte_bitset_clear(priv->mac_own, i);
}
@@ -718,7 +742,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
struct mlx5_priv *priv = dev->data->dev_private;
- if (index < MLX5_MAX_MAC_ADDRESSES)
+ if (index < priv->sh->dev_cap.max_mac_addrs)
rte_bitset_clear(priv->mac_own, index);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH v5 08/10] net/mlx5: pass maximum number of unicast MAC to common code
2026-07-23 12:41 ` [PATCH v5 08/10] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-07-23 17:35 ` Stephen Hemminger
0 siblings, 0 replies; 146+ messages in thread
From: Stephen Hemminger @ 2026-07-23 17:35 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
On Thu, 23 Jul 2026 14:41:54 +0200
David Marchand <david.marchand@redhat.com> wrote:
> Isolate how the MAC addresses array is walked through in the common code
> by passing the max index at which a unicast MAC address is stored in
> dev->data->mac_addrs[].
>
> In the sync callback, the size of the array allocated on the stack is
> known by the caller, treat the mac_n field as an input parameter too.
>
> With this change, only net/mlx5 knows about the max number of
> unicast/multicast MAC addresses.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
Deeper AI review found this:
[PATCH v5 08/10] net/mlx5: mlx5_nl_mac_addr_cb no longer accumulates across netlink messages
The refactor makes mac_n a local reset to 0 on every callback invocation and writes data->mac_n = mac_n (this-call count) on exit, while data->mac_n on entry is repurposed as the array capacity.
mlx5_nl_recv() calls the callback once per nlmsghdr, and a neighbor dump delivers one RTM_NEWNEIGH (one NDA_LLADDR) per message. So with N MAC addresses on the kernel netdev, the callback runs N times, each time writing to mac[0] and reporting count 1 — every address except the last is overwritten. mlx5_nl_mac_addr_list() then returns 1, and mlx5_nl_mac_addr_sync() syncs only the last address. The previous code kept data->mac_n as a persistent running counter, so it accumulated correctly.
Confirmed with a standalone harness (3 addresses across 3 invocations): new logic reports count=1, mac[0] = the 3rd address; old logic reports count=3 at indices 0/1/2. Works only when the netdev has a single MAC, which is why it can pass basic testing — and it's precisely the many-address path that 10/10 then widens to 4096.
The capacity and the running count need to be separate. Keep an input capacity field (e.g. add mac_cap to struct mlx5_nl_mac_addr, set from *mac_n) and leave mac_n as the persistent counter that survives across invocations, checking data->mac_n == data->mac_cap for the full condition — rather than folding both roles into one field.
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v5 00/10] Remove limitations coming from legacy VMDq
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
` (9 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 10/10] net/mlx5: accept more unicast " David Marchand
@ 2026-07-27 7:20 ` David Marchand
10 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-07-27 7:20 UTC (permalink / raw)
To: Stephen Hemminger; +Cc: dev, rjarry, cfontain
Hello Stephen,
On Thu, 23 Jul 2026 at 14:42, David Marchand <david.marchand@redhat.com> wrote:
>
> Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
> VMDq has been imposing a maximum number of mac addresses in the
> mac_addr_add/del API.
>
> Nowadays, new Intel drivers do not support the feature and few other
> drivers implement this feature.
>
> This series enforces that the driver announces VMDq pools before
> using VMDq related features, then remove the limit of number of
> mac addresses for others.
>
> Next step could be to remove the VMDq pool notion from the generic API.
> However I have some concern about this, as changing the quite stable
> mac_addr_add/del API now seems a lot of noise for not much benefit.
>
>
> --
> David Marchand
>
> Changes since v4:
> - rebased,
> - moved mac restoration in dedicated IAVF reset helper,
> - fixed inverted arguments when calling mlx5 mac sync,
If there is no comment on the ethdev changes, could you take the first
3 patches?
I'll send the IAVF and mlx5 changes through the relevant subtrees.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v6 0/3] Remove limitations coming from legacy VMDq
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (8 preceding siblings ...)
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
@ 2026-08-24 11:42 ` David Marchand
2026-08-24 11:42 ` [PATCH v6 1/3] ethdev: check VMDq availability David Marchand
` (3 more replies)
2026-09-04 12:28 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses David Marchand
` (5 subsequent siblings)
15 siblings, 4 replies; 146+ messages in thread
From: David Marchand @ 2026-08-24 11:42 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain
Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
VMDq has been imposing a maximum number of mac addresses in the
mac_addr_add/del API.
Nowadays, new Intel drivers do not support the feature and few other
drivers implement this feature.
This series enforces that the driver announces VMDq pools before
using VMDq related features, then remove the limit of number of
mac addresses for others.
Next step could be to remove the VMDq pool notion from the generic API.
However I have some concern about this, as changing the quite stable
mac_addr_add/del API now seems a lot of noise for not much benefit.
--
David Marchand
Changes since v5:
- no change, simply split out (intel and mlx) drivers update in
separate series,
Changes since v4:
- rebased,
- moved mac restoration in dedicated IAVF reset helper,
- fixed inverted arguments when calling mlx5 mac sync,
Changes since v3:
- rebased (this series is too late for 26.07),
- update some doxygen comments,
- added mlx5 changes,
Changes since v2:
- changed approach: did not introduce a new device capability,
relied on already existing dev_info->max_vmdq_pools,
- fixed duplicate mac addition without VMDq,
- updated documentation,
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
- fixed buffer overflow on mailbox messages during port restart/VF reset,
- fixed duplicate MAC address installation on port start/restart,
David Marchand (3):
ethdev: check VMDq availability
ethdev: skip VMDq pools unless configured
ethdev: hide VMDq internal sizes
doc/guides/rel_notes/release_26_11.rst | 11 +++++
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
lib/ethdev/ethdev_driver.h | 8 ++-
lib/ethdev/rte_ethdev.c | 68 ++++++++++++++++++++------
lib/ethdev/rte_ethdev.h | 8 +--
5 files changed, 71 insertions(+), 25 deletions(-)
--
2.54.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v6 1/3] ethdev: check VMDq availability
2026-08-24 11:42 ` [PATCH v6 0/3] " David Marchand
@ 2026-08-24 11:42 ` David Marchand
2026-08-24 11:42 ` [PATCH v6 2/3] ethdev: skip VMDq pools unless configured David Marchand
` (2 subsequent siblings)
3 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-08-24 11:42 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Andrew Rybchenko, Thomas Monjalon
Refuse VMDq related Rx/Tx modes when the driver do not announce VMDq
pools availability.
This will used later as a gate to ignore/reject VMDq related matters.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v4:
- rebased,
Changes since v2:
- added an entry in release notes,
- changed approach: relied on pre-existing dev_info->max_vmdq_pools
rather than introduce a new device capability,
Changes since v1:
- dropped incorrect VMDq feature announce for bnxt representors, em,
i40e representors, ipn3ke representors,
---
doc/guides/rel_notes/release_26_11.rst | 6 ++++++
lib/ethdev/rte_ethdev.c | 16 ++++++++++++++++
2 files changed, 22 insertions(+)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index c8cc86295d..4b68a5319e 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -93,6 +93,12 @@ API Changes
Also, make sure to start the actual text at the margin.
=======================================================
+* **ethdev: updated VMDq related API.**
+
+ * At port configuration time, the number of VMDq pools advertised by a driver is now used to
+ validate VMDq related Rx and Tx modes (``RTE_ETH_MQ_RX_VMDQ_FLAG``, ``RTE_ETH_MQ_TX_VMDQ_DCB``,
+ ``RTE_ETH_MQ_TX_VMDQ_ONLY``).
+
ABI Changes
-----------
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 9efeaf77cb..9e305d98c1 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1581,6 +1581,22 @@ rte_eth_dev_configure(uint16_t port_id, uint16_t nb_rx_q, uint16_t nb_tx_q,
goto rollback;
}
+ if (dev_info.max_vmdq_pools == 0) {
+ if ((dev_conf->rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq rx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ if (dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_DCB ||
+ dev_conf->txmode.mq_mode == RTE_ETH_MQ_TX_VMDQ_ONLY) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Ethdev port_id=%u does not support VMDq tx mode",
+ port_id);
+ ret = -EINVAL;
+ goto rollback;
+ }
+ }
+
/*
* Setup new number of Rx/Tx queues and reconfigure device.
*/
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v6 2/3] ethdev: skip VMDq pools unless configured
2026-08-24 11:42 ` [PATCH v6 0/3] " David Marchand
2026-08-24 11:42 ` [PATCH v6 1/3] ethdev: check VMDq availability David Marchand
@ 2026-08-24 11:42 ` David Marchand
2026-08-24 16:21 ` Stephen Hemminger
2026-08-24 11:42 ` [PATCH v6 3/3] ethdev: hide VMDq internal sizes David Marchand
2026-08-24 17:01 ` [PATCH v6 0/3] Remove limitations coming from legacy VMDq Stephen Hemminger
3 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-08-24 11:42 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Andrew Rybchenko, Nithin Dabilpuram,
Kiran Kumar K, Sunil Kumar Kori, Satha Rao, Harman Kalra,
Thomas Monjalon
The mac_addr_add API describes that only the 0 pool should be passed
unless VMDq has been enabled, though there was no validation so far.
Add such a check, then cleanup the MAC related operations (adding,
removing, restoring).
As a side effect, the net/cnxk does not need to manually reset the
mac_pool_sel[] array.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v4:
- rebased,
Changes since v3:
- updated doxygen comments,
Changes since v2:
- added an entry in release notes,
- fixed duplicate mac address handling for !vmdq,
- rewrote update of eth_dev_mac_restore to isolate the !vmdq case,
---
doc/guides/rel_notes/release_26_11.rst | 2 +
drivers/net/cnxk/cnxk_ethdev_ops.c | 1 -
lib/ethdev/rte_ethdev.c | 52 ++++++++++++++++++--------
lib/ethdev/rte_ethdev.h | 2 +-
4 files changed, 39 insertions(+), 18 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 4b68a5319e..40a301acf9 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -98,6 +98,8 @@ API Changes
* At port configuration time, the number of VMDq pools advertised by a driver is now used to
validate VMDq related Rx and Tx modes (``RTE_ETH_MQ_RX_VMDQ_FLAG``, ``RTE_ETH_MQ_TX_VMDQ_DCB``,
``RTE_ETH_MQ_TX_VMDQ_ONLY``).
+ * A check was added in ``rte_eth_dev_mac_addr_add`` to validate that the ``pool`` parameter is 0
+ when VMDq is not configured.
ABI Changes
diff --git a/drivers/net/cnxk/cnxk_ethdev_ops.c b/drivers/net/cnxk/cnxk_ethdev_ops.c
index 0ea3d7e89f..c002a93fe1 100644
--- a/drivers/net/cnxk/cnxk_ethdev_ops.c
+++ b/drivers/net/cnxk/cnxk_ethdev_ops.c
@@ -1240,7 +1240,6 @@ cnxk_nix_mc_addr_list_configure(struct rte_eth_dev *eth_dev, struct rte_ether_ad
/* Update address in NIC data structure */
rte_ether_addr_copy(&mc_addr_set[i], &data->mac_addrs[j]);
rte_ether_addr_copy(&mc_addr_set[i], &dev->dmac_addrs[j]);
- data->mac_pool_sel[j] = RTE_BIT64(0);
}
roc_nix_npc_promisc_ena_dis(nix, true);
diff --git a/lib/ethdev/rte_ethdev.c b/lib/ethdev/rte_ethdev.c
index 9e305d98c1..f20cc514a3 100644
--- a/lib/ethdev/rte_ethdev.c
+++ b/lib/ethdev/rte_ethdev.c
@@ -1677,7 +1677,7 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
{
struct rte_ether_addr *addr;
uint16_t i;
- uint32_t pool = 0;
+ uint32_t pool;
uint64_t pool_mask;
/* replay MAC address configuration including default MAC */
@@ -1685,9 +1685,11 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
if (dev->dev_ops->mac_addr_set != NULL)
dev->dev_ops->mac_addr_set(dev, addr);
else if (dev->dev_ops->mac_addr_add != NULL)
- dev->dev_ops->mac_addr_add(dev, addr, 0, pool);
+ dev->dev_ops->mac_addr_add(dev, addr, 0, 0);
if (dev->dev_ops->mac_addr_add != NULL) {
+ bool vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+
for (i = 1; i < dev_info->max_mac_addrs; i++) {
addr = &dev->data->mac_addrs[i];
@@ -1695,15 +1697,19 @@ eth_dev_mac_restore(struct rte_eth_dev *dev,
if (rte_is_zero_ether_addr(addr))
continue;
- pool = 0;
- pool_mask = dev->data->mac_pool_sel[i];
-
- do {
- if (pool_mask & UINT64_C(1))
- dev->dev_ops->mac_addr_add(dev, addr, i, pool);
- pool_mask >>= 1;
- pool++;
- } while (pool_mask);
+ if (!vmdq) {
+ dev->dev_ops->mac_addr_add(dev, addr, i, 0);
+ } else {
+ pool = 0;
+ pool_mask = dev->data->mac_pool_sel[i];
+
+ do {
+ if (pool_mask & UINT64_C(1))
+ dev->dev_ops->mac_addr_add(dev, addr, i, pool);
+ pool_mask >>= 1;
+ pool++;
+ } while (pool_mask);
+ }
}
}
}
@@ -5414,8 +5420,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
uint32_t pool)
{
struct rte_eth_dev *dev;
- int index;
uint64_t pool_mask;
+ bool vmdq;
+ int index;
int ret;
RTE_ETH_VALID_PORTID_OR_ERR_RET(port_id, -ENODEV);
@@ -5440,6 +5447,12 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
RTE_ETHDEV_LOG_LINE(ERR, "Pool ID must be 0-%d", RTE_ETH_64_POOLS - 1);
return -EINVAL;
}
+ vmdq = (dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0;
+ if (!vmdq && pool != 0) {
+ RTE_ETHDEV_LOG_LINE(ERR, "Port %u: VMDq is not configured (pool %d)",
+ port_id, pool);
+ return -EINVAL;
+ }
index = eth_dev_get_mac_addr_index(port_id, addr);
if (index < 0) {
@@ -5450,6 +5463,9 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
return -ENOSPC;
}
} else {
+ if (!vmdq)
+ return 0;
+
pool_mask = dev->data->mac_pool_sel[index];
/* Check if both MAC address and pool is already there, and do nothing */
@@ -5464,8 +5480,10 @@ rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *addr,
/* Update address in NIC data structure */
rte_ether_addr_copy(addr, &dev->data->mac_addrs[index]);
- /* Update pool bitmap in NIC data structure */
- dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ if (vmdq) {
+ /* Update pool bitmap in NIC data structure */
+ dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
+ }
}
ret = eth_err(port_id, ret);
@@ -5510,8 +5528,10 @@ rte_eth_dev_mac_addr_remove(uint16_t port_id, struct rte_ether_addr *addr)
/* Update address in NIC data structure */
rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
- /* reset pool bitmap */
- dev->data->mac_pool_sel[index] = 0;
+ if ((dev->data->dev_conf.rxmode.mq_mode & RTE_ETH_MQ_RX_VMDQ_FLAG) != 0) {
+ /* reset pool bitmap */
+ dev->data->mac_pool_sel[index] = 0;
+ }
rte_ethdev_trace_mac_addr_remove(port_id, addr);
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index ee400b386f..e2a5ba1549 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -4636,7 +4636,7 @@ int rte_eth_dev_priority_flow_ctrl_set(uint16_t port_id,
* - (-ENODEV) if *port* is invalid.
* - (-EIO) if device is removed.
* - (-ENOSPC) if no more MAC addresses can be added.
- * - (-EINVAL) if MAC address is invalid.
+ * - (-EINVAL) if MAC address is invalid or a non 0 pool was passed but VMDq is not enabled.
*/
int rte_eth_dev_mac_addr_add(uint16_t port_id, struct rte_ether_addr *mac_addr,
uint32_t pool);
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v6 3/3] ethdev: hide VMDq internal sizes
2026-08-24 11:42 ` [PATCH v6 0/3] " David Marchand
2026-08-24 11:42 ` [PATCH v6 1/3] ethdev: check VMDq availability David Marchand
2026-08-24 11:42 ` [PATCH v6 2/3] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-08-24 11:42 ` David Marchand
2026-08-24 17:01 ` [PATCH v6 0/3] Remove limitations coming from legacy VMDq Stephen Hemminger
3 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-08-24 11:42 UTC (permalink / raw)
To: dev; +Cc: rjarry, cfontain, Andrew Rybchenko, Thomas Monjalon
Hide RTE_ETH_NUM_RECEIVE_MAC_ADDR and RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY
in the driver API as those (ambiguous) macros are only a driver concern.
In practice, this is only used by the bnxt and ixgbe (+ clones) drivers.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
---
Changes since v4:
- rebased,
Changes since v2:
- added an entry in release notes,
---
doc/guides/rel_notes/release_26_11.rst | 3 +++
lib/ethdev/ethdev_driver.h | 8 +++++++-
lib/ethdev/rte_ethdev.h | 6 ------
3 files changed, 10 insertions(+), 7 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 40a301acf9..1d2de68715 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -100,6 +100,9 @@ API Changes
``RTE_ETH_MQ_TX_VMDQ_ONLY``).
* A check was added in ``rte_eth_dev_mac_addr_add`` to validate that the ``pool`` parameter is 0
when VMDq is not configured.
+ * The ``RTE_ETH_NUM_RECEIVE_MAC_ADDR`` and ``RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY`` macros are VMDq
+ related and are sizes of internal arrays in ethdev that only drivers need to care about.
+ Those macros are moved to the driver only ethdev API.
ABI Changes
diff --git a/lib/ethdev/ethdev_driver.h b/lib/ethdev/ethdev_driver.h
index 0f336f9567..294f68504b 100644
--- a/lib/ethdev/ethdev_driver.h
+++ b/lib/ethdev/ethdev_driver.h
@@ -119,6 +119,12 @@ struct __rte_cache_aligned rte_eth_dev {
struct rte_eth_dev_sriov;
struct rte_eth_dev_owner;
+/* Definitions used for receive MAC address */
+#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
+
+/* Definitions used for unicast hash */
+#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
+
/**
* @internal
* The data part, with no function pointers, associated with each Ethernet
@@ -153,7 +159,7 @@ struct __rte_cache_aligned rte_eth_dev_data {
* The first entry (index zero) is the default address.
*/
struct rte_ether_addr *mac_addrs;
- /** Bitmap associating MAC addresses to pools */
+ /** Bitmap associating MAC addresses to VMDq pools */
uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
/**
* Device Ethernet MAC addresses of hash filtering.
diff --git a/lib/ethdev/rte_ethdev.h b/lib/ethdev/rte_ethdev.h
index e2a5ba1549..bee47549cb 100644
--- a/lib/ethdev/rte_ethdev.h
+++ b/lib/ethdev/rte_ethdev.h
@@ -903,12 +903,6 @@ rte_eth_rss_hf_refine(uint64_t rss_hf)
#define RTE_ETH_VLAN_ID_MAX 0x0FFF /**< VLAN ID is in lower 12 bits*/
/**@}*/
-/* Definitions used for receive MAC address */
-#define RTE_ETH_NUM_RECEIVE_MAC_ADDR 128 /**< Maximum nb. of receive mac addr. */
-
-/* Definitions used for unicast hash */
-#define RTE_ETH_VMDQ_NUM_UC_HASH_ARRAY 128 /**< Maximum nb. of UC hash array. */
-
/**@{@name VMDq Rx mode
* @see rte_eth_vmdq_rx_conf.rx_mode
*/
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH v6 2/3] ethdev: skip VMDq pools unless configured
2026-08-24 11:42 ` [PATCH v6 2/3] ethdev: skip VMDq pools unless configured David Marchand
@ 2026-08-24 16:21 ` Stephen Hemminger
2026-08-24 16:24 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Stephen Hemminger @ 2026-08-24 16:21 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Andrew Rybchenko, Nithin Dabilpuram,
Kiran Kumar K, Sunil Kumar Kori, Satha Rao, Harman Kalra,
Thomas Monjalon
On Mon, 24 Aug 2026 13:42:06 +0200
David Marchand <david.marchand@redhat.com> wrote:
> The mac_addr_add API describes that only the 0 pool should be passed
> unless VMDq has been enabled, though there was no validation so far.
> Add such a check, then cleanup the MAC related operations (adding,
> removing, restoring).
>
> As a side effect, the net/cnxk does not need to manually reset the
> mac_pool_sel[] array.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
> ---
Claude Fable AI review has some warnings.
Review of [PATCH v6 0/3] ethdev: VMDq cleanups
Applied cleanly on current main (c1a46b9). Build testing not done
here per your note; findings below are from reading the applied tree.
Patches 1/3 and 3/3: no findings.
Patch 2/3 - ethdev: skip VMDq pools unless configured
-----------------------------------------------------
Error: MAC removal no longer reaches hardware on i40e and bnxt when
VMDq is not configured.
Both drivers read dev->data->mac_pool_sel[index] inside their
mac_addr_remove callback to decide which VSI / VNIC to delete the
filter from:
drivers/net/intel/i40e/i40e_ethdev.c:4516 i40e_macaddr_remove()
drivers/net/bnxt/bnxt_ethdev.c:2027 bnxt_mac_addr_remove_op()
Before this patch rte_eth_dev_mac_addr_add() always did
mac_pool_sel[index] |= RTE_BIT64(pool), so a non-VMDq add left bit 0
set and the drivers deleted from pool 0 (main VSI / VNIC 0). After
this patch the bitmap update is guarded by "if (vmdq)", so in RSS or
NONE mode mac_pool_sel[index] stays 0. rte_eth_dev_mac_addr_remove()
then calls dev_ops->mac_addr_remove() while the bitmap is still 0,
both drivers iterate an empty mask, and the hardware filter is never
removed. The software copy in mac_addrs[] is cleared, so the address
looks gone from the API but still receives traffic.
bnxt_restore_mac_filters() (bnxt_ethdev.c:4566) has the same
dependency on the recovery path: with a zero mask it replays nothing.
A second, less common consequence of the guards: after a VMDq ->
non-VMDq reconfigure, rte_eth_dev_mac_addr_remove() no longer clears
mac_pool_sel[index], so stale pool bits survive. A later VMDq
reconfigure that places a new address at the same index does |= on
top of the stale bits and eth_dev_mac_restore() replays the address
into pools it was never added to.
Since ethdev_driver.h documents mac_pool_sel as "Bitmap associating
MAC addresses to VMDq pools" and drivers are entitled to consume it,
ethdev should keep it coherent in all modes. The minimal fix is to
drop the two guards and keep the bitmap always up to date (pool is
already forced to 0 when !vmdq, so bit 0 is recorded):
In rte_eth_dev_mac_addr_add():
if (ret == 0) {
/* Update address in NIC data structure */
rte_ether_addr_copy(addr, &dev->data->mac_addrs[index]);
/* Update pool bitmap in NIC data structure */
dev->data->mac_pool_sel[index] |= RTE_BIT64(pool);
}
In rte_eth_dev_mac_addr_remove():
/* reset pool bitmap */
dev->data->mac_pool_sel[index] = 0;
The new "if (!vmdq) return 0;" early-return for an already-present
address, the !vmdq branch in eth_dev_mac_restore(), and the cnxk
cleanup all remain valid with this change.
The alternative is to make i40e and bnxt stop depending on
mac_pool_sel[] in non-VMDq mode, but that would need to land in the
same series and is more invasive.
Info: "Port %u: VMDq is not configured (pool %d)" - pool is uint32_t,
use %u.
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 2/3] ethdev: skip VMDq pools unless configured
2026-08-24 16:21 ` Stephen Hemminger
@ 2026-08-24 16:24 ` David Marchand
2026-08-24 16:39 ` Stephen Hemminger
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-08-24 16:24 UTC (permalink / raw)
To: Stephen Hemminger
Cc: dev, rjarry, cfontain, Andrew Rybchenko, Nithin Dabilpuram,
Kiran Kumar K, Sunil Kumar Kori, Satha Rao, Harman Kalra,
Thomas Monjalon
On Mon, 24 Aug 2026 at 18:21, Stephen Hemminger
<stephen@networkplumber.org> wrote:
>
> On Mon, 24 Aug 2026 13:42:06 +0200
> David Marchand <david.marchand@redhat.com> wrote:
>
> > The mac_addr_add API describes that only the 0 pool should be passed
> > unless VMDq has been enabled, though there was no validation so far.
> > Add such a check, then cleanup the MAC related operations (adding,
> > removing, restoring).
> >
> > As a side effect, the net/cnxk does not need to manually reset the
> > mac_pool_sel[] array.
> >
> > Signed-off-by: David Marchand <david.marchand@redhat.com>
> > Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
> > ---
>
> Claude Fable AI review has some warnings.
Enough bikeshedding please.
I requested a merge one month ago and got ignored.
Those warnings below are pointless.
Thanks.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 2/3] ethdev: skip VMDq pools unless configured
2026-08-24 16:24 ` David Marchand
@ 2026-08-24 16:39 ` Stephen Hemminger
0 siblings, 0 replies; 146+ messages in thread
From: Stephen Hemminger @ 2026-08-24 16:39 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Andrew Rybchenko, Nithin Dabilpuram,
Kiran Kumar K, Sunil Kumar Kori, Satha Rao, Harman Kalra,
Thomas Monjalon
On Mon, 24 Aug 2026 18:24:29 +0200
David Marchand <david.marchand@redhat.com> wrote:
> On Mon, 24 Aug 2026 at 18:21, Stephen Hemminger
> <stephen@networkplumber.org> wrote:
> >
> > On Mon, 24 Aug 2026 13:42:06 +0200
> > David Marchand <david.marchand@redhat.com> wrote:
> >
> > > The mac_addr_add API describes that only the 0 pool should be passed
> > > unless VMDq has been enabled, though there was no validation so far.
> > > Add such a check, then cleanup the MAC related operations (adding,
> > > removing, restoring).
> > >
> > > As a side effect, the net/cnxk does not need to manually reset the
> > > mac_pool_sel[] array.
> > >
> > > Signed-off-by: David Marchand <david.marchand@redhat.com>
> > > Acked-by: Andrew Rybchenko <andrew.rybchenko@oktetlabs.ru>
> > > ---
> >
> > Claude Fable AI review has some warnings.
>
> Enough bikeshedding please.
> I requested a merge one month ago and got ignored.
>
> Those warnings below are pointless.
>
> Thanks.
I was going to merge. Just wanted to post the current AI reviews
to make sure.
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 0/3] Remove limitations coming from legacy VMDq
2026-08-24 11:42 ` [PATCH v6 0/3] " David Marchand
` (2 preceding siblings ...)
2026-08-24 11:42 ` [PATCH v6 3/3] ethdev: hide VMDq internal sizes David Marchand
@ 2026-08-24 17:01 ` Stephen Hemminger
3 siblings, 0 replies; 146+ messages in thread
From: Stephen Hemminger @ 2026-08-24 17:01 UTC (permalink / raw)
To: David Marchand; +Cc: dev, rjarry, cfontain
On Mon, 24 Aug 2026 13:42:04 +0200
David Marchand <david.marchand@redhat.com> wrote:
> Since the commit 88ac4396ad29 ("ethdev: add VMDq support"),
> VMDq has been imposing a maximum number of mac addresses in the
> mac_addr_add/del API.
>
> Nowadays, new Intel drivers do not support the feature and few other
> drivers implement this feature.
>
> This series enforces that the driver announces VMDq pools before
> using VMDq related features, then remove the limit of number of
> mac addresses for others.
>
> Next step could be to remove the VMDq pool notion from the generic API.
> However I have some concern about this, as changing the quite stable
> mac_addr_add/del API now seems a lot of noise for not much benefit.
>
Applied to next-net
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (9 preceding siblings ...)
2026-08-24 11:42 ` [PATCH v6 0/3] " David Marchand
@ 2026-09-04 12:28 ` David Marchand
2026-09-04 12:28 ` [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install David Marchand
` (2 more replies)
2026-09-08 9:27 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (4 subsequent siblings)
15 siblings, 3 replies; 146+ messages in thread
From: David Marchand @ 2026-09-04 12:28 UTC (permalink / raw)
To: dev; +Cc: bruce.richardson, Vladimir Medvedkin
E810 hardware provides 32k switch lookups.
Thanks to this, it is possible to allow a lot more secondary mac
addresses than what is possible today.
In practice, the maximum number of macs available per port may be lower
and depends on usage by other (trusted?) VFs on the same PF.
There is no way to figure out this limit but to try adding a mac address
and get an error from the PF driver.
Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
accordingly.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v5:
- separated from series that went in next-net,
- rebased,
Changes since v4:
- rebased,
Changes since v2:
- added an entry in release notes,
- removed unneeded temp variable,
Changes since v1:
- fixed buffer overflow on mailbox messages during port restart/VF reset,
---
doc/guides/rel_notes/release_26_11.rst | 4 +
drivers/net/intel/iavf/iavf.h | 5 +-
drivers/net/intel/iavf/iavf_ethdev.c | 12 +--
drivers/net/intel/iavf/iavf_vchnl.c | 113 ++++++++++++++++++-------
4 files changed, 94 insertions(+), 40 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 907f9013ff..03e361d3ad 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -64,6 +64,10 @@ New Features
* Renamed the ``enable_ptype_lldp`` devarg to ``enable_lldp``.
The old name is no longer accepted.
+ * Increased the maximum number of secondary unicast MAC addresses
+ from 64 to 32k.
+ This increases a VF port memory footprint by ~192kB.
+
Removed Items
-------------
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 044de3cfd4..5271e84f8f 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -32,7 +32,8 @@
#define IAVF_IRQ_MAP_NUM_PER_BUF 128
#define IAVF_RXTX_QUEUE_CHUNKS_NUM 2
-#define IAVF_NUM_MACADDR_MAX 64
+#define IAVF_UC_MACADDR_MAX 32768
+#define IAVF_MC_MACADDR_MAX 64
#define IAVF_DEV_WATCHDOG_PERIOD 2000 /* microseconds, set 0 to disable*/
@@ -256,7 +257,7 @@ struct iavf_info {
uint32_t link_speed;
/* Multicast addrs */
- struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
+ struct rte_ether_addr mc_addrs[IAVF_MC_MACADDR_MAX];
uint16_t mc_addrs_num; /* Multicast mac addresses number */
struct iavf_vsi vsi;
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index 2ac4dbdac4..bbd1f08ff0 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -410,10 +410,10 @@ iavf_set_mc_addr_list(struct rte_eth_dev *dev,
IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
int err, ret;
- if (mc_addrs_num > IAVF_NUM_MACADDR_MAX) {
+ if (mc_addrs_num > IAVF_MC_MACADDR_MAX) {
PMD_DRV_LOG(ERR,
"can't add more than a limited number (%u) of addresses.",
- (uint32_t)IAVF_NUM_MACADDR_MAX);
+ (uint32_t)IAVF_MC_MACADDR_MAX);
return -EINVAL;
}
@@ -1186,7 +1186,7 @@ iavf_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->hash_key_size = vf->vf_res->rss_key_size;
dev_info->reta_size = vf->vf_res->rss_lut_size;
dev_info->flow_type_rss_offloads = IAVF_RSS_OFFLOAD_ALL;
- dev_info->max_mac_addrs = IAVF_NUM_MACADDR_MAX;
+ dev_info->max_mac_addrs = IAVF_UC_MACADDR_MAX;
/*
* Runtime queue setup can race with the hardware Tx rate limiter on
* E810 VFs and corrupt queue state. Once a per-queue bandwidth rte_tm
@@ -3105,12 +3105,12 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
iavf_set_default_ptype_table(eth_dev);
/* copy mac addr */
- eth_dev->data->mac_addrs = rte_zmalloc(
- "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX, 0);
+ eth_dev->data->mac_addrs = rte_calloc("iavf_mac", IAVF_UC_MACADDR_MAX,
+ RTE_ETHER_ADDR_LEN, 0);
if (!eth_dev->data->mac_addrs) {
PMD_INIT_LOG(ERR, "Failed to allocate %d bytes needed to"
" store MAC addresses",
- RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX);
+ RTE_ETHER_ADDR_LEN * IAVF_UC_MACADDR_MAX);
ret = -ENOMEM;
goto init_vf_err;
}
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index e3120c655d..e8f2085d6e 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1671,49 +1671,98 @@ iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num)
return 0;
}
-void
-iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+static int
+iavf_add_del_uc_addr_bulk(struct iavf_adapter *adapter, struct rte_ether_addr *addrs,
+ uint32_t nb_addrs, bool add)
{
+#define IAVF_ETH_ADDR_PER_REQ \
+ ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
+ sizeof(struct virtchnl_ether_addr))
struct {
struct virtchnl_ether_addr_list list;
- struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
- } list_req = {0};
- struct virtchnl_ether_addr_list *list = &list_req.list;
+ struct virtchnl_ether_addr addr[IAVF_ETH_ADDR_PER_REQ];
+ } cmd_buffer;
+#undef IAVF_ETH_ADDR_PER_REQ
+ struct virtchnl_ether_addr_list *list = &cmd_buffer.list;
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
- struct iavf_cmd_info args = {0};
- int err, i;
- size_t buf_len;
- for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
- struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
- struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
+ for (uint32_t i = 0; i < nb_addrs; i++) {
+ struct iavf_cmd_info args;
+ uint32_t batch;
+ int err;
- /* ignore empty addresses */
- if (rte_is_zero_ether_addr(addr))
- continue;
+ batch = i % RTE_DIM(cmd_buffer.addr);
+
+ if (batch == 0) {
+ memset(&cmd_buffer, 0, sizeof(cmd_buffer));
+ list->vsi_id = vf->vsi_res->vsi_id;
+ list->num_elements = 0;
+ }
+
+ memcpy(list->list[batch].addr, addrs[i].addr_bytes,
+ sizeof(list->list[batch].addr));
+ list->list[batch].type = VIRTCHNL_ETHER_ADDR_EXTRA;
list->num_elements++;
- memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
- vc_addr->type = (list->num_elements == 1) ?
- VIRTCHNL_ETHER_ADDR_PRIMARY :
- VIRTCHNL_ETHER_ADDR_EXTRA;
+ if (batch != RTE_DIM(cmd_buffer.addr) - 1 && i != nb_addrs - 1)
+ continue;
+
+ memset(&args, 0, sizeof(args));
+ args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
+ args.in_args = (uint8_t *)list;
+ args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
+ sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ args.out_buffer = msg_buf;
+ args.out_size = IAVF_AQ_BUF_SZ;
+ err = iavf_execute_vf_cmd_safe(adapter, &args);
+ if (err != 0) {
+ PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
+ return err;
+ }
+
+ PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
+ list->num_elements);
}
- /* for some reason PF side checks for buffer being too big, so adjust it down */
- buf_len = sizeof(struct virtchnl_ether_addr_list) +
- sizeof(struct virtchnl_ether_addr) * list->num_elements;
+ return 0;
+}
- list->vsi_id = vf->vsi_res->vsi_id;
- args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
- args.in_args = (uint8_t *)list;
- args.in_args_size = buf_len;
- args.out_buffer = msg_buf;
- args.out_size = IAVF_AQ_BUF_SZ;
- err = iavf_execute_vf_cmd_safe(adapter, &args);
- if (err)
- PMD_DRV_LOG(ERR, "fail to execute command %s",
- add ? "OP_ADD_ETHER_ADDRESS" : "OP_DEL_ETHER_ADDRESS");
+void
+iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+{
+ int start = -1;
+ int i;
+
+ /* Handle primary address (index 0) separately */
+ if (!rte_is_zero_ether_addr(&adapter->dev_data->mac_addrs[0]))
+ iavf_add_del_eth_addr(adapter, &adapter->dev_data->mac_addrs[0], add,
+ VIRTCHNL_ETHER_ADDR_PRIMARY);
+
+ /* Process secondary addresses in contiguous blocks */
+ for (i = 1; i < IAVF_UC_MACADDR_MAX; i++) {
+ struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
+
+ if (!rte_is_zero_ether_addr(addr)) {
+ if (start == -1)
+ start = i;
+ continue;
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ start = -1;
+ }
+ }
+
+ if (start != -1) {
+ iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
+ i - start, add);
+ }
}
int
@@ -2304,7 +2353,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
- (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
+ (IAVF_MC_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
struct virtchnl_ether_addr_list *list;
struct iavf_cmd_info args;
uint32_t i;
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install
2026-09-04 12:28 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-09-04 12:28 ` David Marchand
2026-09-09 9:39 ` Loftus, Ciara
2026-09-10 10:24 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses Burakov, Anatoly
2026-09-10 12:13 ` Burakov, Anatoly
2 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-04 12:28 UTC (permalink / raw)
To: dev; +Cc: bruce.richardson, stable, Vladimir Medvedkin
On port restart, all MAC addresses get pushed *twice* to the hardware,
once by the driver and once by the eth_dev_mac_restore() in ethdev.
On the other hand, MAC address filters are reset in the hardware
by the PF only when a VF reset is triggered.
Strictly speaking, the mac restore on port (re)start is unneeded,
if no VF reset happened, so we can announce to ethdev that no mac
restoration is needed via a get_restore_flags callback.
Then, move the mac restoration to the VF reset handler.
Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux driver")
Cc: stable@dpdk.org
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v4:
- rebased on next-net-intel,
Changes since v4:
- moved mac restoration in iavf_post_reset_reconfig,
---
drivers/net/intel/iavf/iavf_ethdev.c | 27 ++++++++++++++++-----------
1 file changed, 16 insertions(+), 11 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index bbd1f08ff0..bec7b3b6d7 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -292,11 +292,12 @@ iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
__rte_unused enum rte_eth_dev_operation op)
{
/*
- * The unicast and multicast promiscuous settings persist across a
+ * The mac addresses, unicast and multicast promiscuous settings persist across a
* stop/start; they are only cleared by a VF reset, which the driver
* restores itself. So ethdev does not need to restore them on start.
*/
- return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_PROMISC |
+ return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_MAC_ADDR |
+ RTE_ETH_RESTORE_PROMISC |
RTE_ETH_RESTORE_ALLMULTI);
}
@@ -1095,15 +1096,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
rte_intr_enable(intr_handle);
}
- /* Set all mac addrs */
- iavf_add_del_all_mac_addr(adapter, true);
-
- if (!adapter->mac_primary_set)
- adapter->mac_primary_set = true;
-
- /* Set all multicast addresses */
- iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num,
- true);
+ if (!adapter->mac_primary_set) {
+ if (iavf_add_del_eth_addr(adapter, &dev->data->mac_addrs[0], true,
+ VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
+ PMD_DRV_LOG(ERR, "failed to add primary MAC:" RTE_ETHER_ADDR_PRT_FMT,
+ RTE_ETHER_ADDR_BYTES(&dev->data->mac_addrs[0]));
+ else
+ adapter->mac_primary_set = true;
+ }
rte_spinlock_init(&vf->phc_time_aq_lock);
@@ -3434,6 +3434,11 @@ iavf_post_reset_reconfig(struct rte_eth_dev *dev)
int ret = 0;
bool allmulti = false, allunicast = false;
struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
+ struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data->dev_private);
+
+ /* After a VF reset, all MAC addresses got flushed, restore them. */
+ iavf_add_del_all_mac_addr(adapter, true);
+ iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num, true);
/* Restore pre-reset unicast promiscuous and multicast promiscuous states */
if (dev->data->promiscuous)
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (10 preceding siblings ...)
2026-09-04 12:28 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-09-08 9:27 ` David Marchand
2026-09-08 9:27 ` [PATCH v6 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
` (4 more replies)
2026-09-14 8:17 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset David Marchand
` (3 subsequent siblings)
15 siblings, 5 replies; 146+ messages in thread
From: David Marchand @ 2026-09-08 9:27 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
This helper is exposing internals of net/mlx5 for no good reason.
All this code does is calling the remove helper.
Walk through the list in Linux implementation like the Windows
implementation.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 40 -----------------------------
drivers/common/mlx5/linux/mlx5_nl.h | 5 ----
drivers/net/mlx5/linux/mlx5_os.c | 13 +++++++---
3 files changed, 10 insertions(+), 48 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 42ccb73e36..40a8b8a2cc 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -825,46 +825,6 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
}
-/**
- * Flush all added MAC addresses.
- *
- * @param[in] nlsk_fd
- * Netlink socket file descriptor.
- * @param[in] iface_idx
- * Net device interface index.
- * @param[in] mac_addrs
- * Mac addresses array to flush.
- * @param n
- * @p mac_addrs array size.
- * @param mac_own
- * BITFIELD_DECLARE array to store the mac.
- * @param vf
- * Flag for a VF device.
- */
-RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_flush)
-void
-mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own, bool vf)
-{
- int i;
-
- if (n <= 0 || n > MLX5_MAX_MAC_ADDRESSES)
- return;
-
- for (i = n - 1; i >= 0; --i) {
- struct rte_ether_addr *m = &mac_addrs[i];
-
- if (BITFIELD_ISSET(mac_own, i)) {
- if (vf)
- mlx5_nl_mac_addr_remove(nlsk_fd,
- iface_idx,
- m, i);
- BITFIELD_RESET(mac_own, i);
- }
- }
-}
-
/**
* Enable promiscuous / all multicast mode through Netlink.
*
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 8ccdd244b0..0242342c47 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -65,11 +65,6 @@ __rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
__rte_internal
-void mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own,
- bool vf);
-__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
int mlx5_nl_allmulti(int nlsk_fd, unsigned int iface_idx, int enable);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 592e233844..49a8751761 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3529,10 +3529,17 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
{
struct mlx5_priv *priv = dev->data->dev_private;
const int vf = priv->sh->dev_cap.vf;
+ int i;
- mlx5_nl_mac_addr_flush(priv->nl_socket_route, mlx5_ifindex(dev),
- dev->data->mac_addrs,
- MLX5_MAX_MAC_ADDRESSES, priv->mac_own, vf);
+ for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (vf)
+ mlx5_nl_mac_addr_remove(priv->nl_socket_route,
+ mlx5_ifindex(dev),
+ &dev->data->mac_addrs[i], i);
+ BITFIELD_RESET(priv->mac_own, i);
+ }
+ }
}
static bool
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v6 2/5] net/mlx5: remove redundant MAC address index checks
2026-09-08 9:27 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
@ 2026-09-08 9:27 ` David Marchand
2026-09-11 8:36 ` Dariusz Sosnowski
2026-09-08 9:27 ` [PATCH v6 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
` (3 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-08 9:27 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
On the net/mlx5 side, mlx5_mac.c validates that any MAC address and its
index is valid before calling the OS specific helpers.
So those OS helpers do not have to validate again the index.
Cascading this consideration, validating the MAC index against
MLX5_MAX_MAC_ADDRESSES in common code is also redundant.
The common code only deals with netlink, remove any index concern.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 17 ++---------------
drivers/common/mlx5/linux/mlx5_nl.h | 4 ++--
drivers/net/mlx5/linux/mlx5_os.c | 9 ++++-----
3 files changed, 8 insertions(+), 22 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 40a8b8a2cc..3207eae563 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -719,8 +719,6 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to register.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -728,16 +726,11 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_add)
int
mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
int ret;
ret = mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 1);
- if (!ret) {
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
- }
if (ret == -EEXIST)
return 0;
return ret;
@@ -752,8 +745,6 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to remove.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -761,12 +752,8 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_remove)
int
mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
-
return mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 0);
}
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 0242342c47..0d6259f4ad 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -57,10 +57,10 @@ __rte_internal
int mlx5_nl_init(int protocol, int groups);
__rte_internal
int mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 49a8751761..c65293cb25 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3382,9 +3382,8 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[index], index);
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ &dev->data->mac_addrs[index]);
+ BITFIELD_RESET(priv->mac_own, index);
}
/**
@@ -3411,7 +3410,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
if (vf)
ret = mlx5_nl_mac_addr_add(priv->nl_socket_route,
mlx5_ifindex(dev),
- mac, index);
+ mac);
if (!ret)
BITFIELD_SET(priv->mac_own, index);
@@ -3536,7 +3535,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[i], i);
+ &dev->data->mac_addrs[i]);
BITFIELD_RESET(priv->mac_own, i);
}
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v6 3/5] net/mlx5: pass maximum number of unicast MAC to common code
2026-09-08 9:27 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-08 9:27 ` [PATCH v6 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-09-08 9:27 ` David Marchand
2026-09-11 8:38 ` Dariusz Sosnowski
2026-09-08 9:27 ` [PATCH v6 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
` (2 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-08 9:27 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Isolate how the MAC addresses array is walked through in the common code
by passing the max index at which a unicast MAC address is stored in
dev->data->mac_addrs[].
With this change, only net/mlx5 knows about the max number of
unicast/multicast MAC addresses.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v5:
- fixed mlx5_nl_mac_addr_cb,
---
drivers/common/mlx5/linux/mlx5_nl.c | 20 ++++++++++++--------
drivers/common/mlx5/linux/mlx5_nl.h | 2 +-
drivers/common/mlx5/mlx5_common.h | 8 --------
drivers/net/mlx5/linux/mlx5_os.c | 1 +
drivers/net/mlx5/mlx5.h | 8 ++++++++
5 files changed, 22 insertions(+), 17 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 3207eae563..12942eefa5 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -174,6 +174,7 @@ struct mlx5_nl_mac_addr {
struct rte_ether_addr (*mac)[];
/**< MAC address handled by the device. */
int mac_n; /**< Number of addresses in the array. */
+ int max_macs; /**< Size of the array. */
};
static RTE_ATOMIC(uint32_t) atomic_sn;
@@ -475,7 +476,7 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
RTA_OK(attribute, len);
attribute = RTA_NEXT(attribute, len)) {
if (attribute->rta_type == NDA_LLADDR) {
- if (data->mac_n == MLX5_MAX_MAC_ADDRESSES) {
+ if (data->mac_n == data->max_macs) {
DRV_LOG(WARNING,
"not enough room to finalize the"
" request");
@@ -505,9 +506,9 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
* Net device interface index.
* @param mac[out]
* Pointer to the array table of MAC addresses to fill.
- * Its size should be of MLX5_MAX_MAC_ADDRESSES.
- * @param mac_n[out]
- * Number of entries filled in MAC array.
+ * @param mac_n[in,out]
+ * Size of the MAC array on input.
+ * Number of entries filled in MAC array on output.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -533,6 +534,7 @@ mlx5_nl_mac_addr_list(int nlsk_fd, unsigned int iface_idx,
struct mlx5_nl_mac_addr data = {
.mac = mac,
.mac_n = 0,
+ .max_macs = *mac_n,
};
uint32_t sn = MLX5_NL_SN_GENERATE;
int ret;
@@ -766,16 +768,18 @@ mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac_addrs
* Mac addresses array to sync.
+ * @param uc_n
+ * Number of UC entries in @p mac_addrs.
* @param n
* @p mac_addrs array size.
*/
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_sync)
void
mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n)
+ struct rte_ether_addr *mac_addrs, int uc_n, int n)
{
struct rte_ether_addr macs[n];
- int macs_n = 0;
+ int macs_n = n;
int i;
int ret;
@@ -794,7 +798,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
continue;
if (rte_is_multicast_ether_addr(&macs[i])) {
/* Find the first entry available. */
- for (j = MLX5_MAX_UC_MAC_ADDRESSES; j != n; ++j) {
+ for (j = uc_n; j != n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
@@ -802,7 +806,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
} else {
/* Find the first entry available. */
- for (j = 0; j != MLX5_MAX_UC_MAC_ADDRESSES; ++j) {
+ for (j = 0; j != uc_n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 0d6259f4ad..07a3b531b5 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -63,7 +63,7 @@ int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n);
+ struct rte_ether_addr *mac_addrs, int uc_n, int n);
__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 3767020823..8aed3dbabe 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -160,14 +160,6 @@ enum {
PCI_DEVICE_ID_MELLANOX_CONNECTX10C2C = 0x2101,
};
-/* Maximum number of simultaneous unicast MAC addresses. */
-#define MLX5_MAX_UC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous Multicast MAC addresses. */
-#define MLX5_MAX_MC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous MAC addresses. */
-#define MLX5_MAX_MAC_ADDRESSES \
- (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
-
/* Recognized Infiniband device physical port name types. */
enum mlx5_nl_phys_port_name_type {
MLX5_PHYS_PORT_NAME_TYPE_NOTSET = 0, /* Not set. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index c65293cb25..0e59d3f2f5 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -1762,6 +1762,7 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
+ MLX5_MAX_UC_MAC_ADDRESSES,
MLX5_MAX_MAC_ADDRESSES);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 190d203c49..c77cdc48a8 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -82,6 +82,14 @@
/* Maximum allowed MTU to be reported whenever PMD cannot query it from OS. */
#define MLX5_ETH_MAX_MTU (9978)
+/* Maximum number of simultaneous unicast MAC addresses. */
+#define MLX5_MAX_UC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous Multicast MAC addresses. */
+#define MLX5_MAX_MC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous MAC addresses. */
+#define MLX5_MAX_MAC_ADDRESSES \
+ (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
+
enum mlx5_ipool_index {
#if defined(HAVE_IBV_FLOW_DV_SUPPORT) || !defined(HAVE_INFINIBAND_VERBS_H)
MLX5_IPOOL_DECAP_ENCAP = 0, /* Pool for encap/decap resource. */
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v6 4/5] net/mlx5: use bitset for tracking MAC addresses
2026-09-08 9:27 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-08 9:27 ` [PATCH v6 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
2026-09-08 9:27 ` [PATCH v6 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-09-08 9:27 ` David Marchand
2026-09-11 8:40 ` Dariusz Sosnowski
2026-09-08 9:27 ` [PATCH v6 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-11 8:35 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux Dariusz Sosnowski
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-08 9:27 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
EAL provides bitset that does the same as this set of mlx5 macros.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/common/mlx5/mlx5_common.h | 16 ----------------
drivers/net/mlx5/linux/mlx5_os.c | 8 ++++----
drivers/net/mlx5/mlx5.h | 3 ++-
drivers/net/mlx5/mlx5_trigger.c | 2 +-
drivers/net/mlx5/windows/mlx5_os.c | 8 ++++----
5 files changed, 11 insertions(+), 26 deletions(-)
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 8aed3dbabe..8b8f20e2f4 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -30,22 +30,6 @@
#define MLX5_PCI_DRIVER_NAME "mlx5_pci"
#define MLX5_AUXILIARY_DRIVER_NAME "mlx5_auxiliary"
-/* Bit-field manipulation. */
-#define BITFIELD_DECLARE(bf, type, size) \
- type bf[(((size_t)(size) / (sizeof(type) * CHAR_BIT)) + \
- !!((size_t)(size) % (sizeof(type) * CHAR_BIT)))]
-#define BITFIELD_DEFINE(bf, type, size) \
- BITFIELD_DECLARE((bf), type, (size)) = { 0 }
-#define BITFIELD_SET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] |= \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_RESET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] &= \
- ~((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_ISSET(bf, b) \
- !!(((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] & \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT)))))
-
/*
* Helper macros to work around __VA_ARGS__ limitations in a C99 compliant
* manner.
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 0e59d3f2f5..7d8ea4acba 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3384,7 +3384,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[index]);
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -3413,7 +3413,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
mlx5_ifindex(dev),
mac);
if (!ret)
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return ret;
}
@@ -3532,12 +3532,12 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[i]);
- BITFIELD_RESET(priv->mac_own, i);
+ rte_bitset_clear(priv->mac_own, i);
}
}
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index c77cdc48a8..7f20c811f3 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -14,6 +14,7 @@
#include <rte_pci.h>
#include <rte_ether.h>
+#include <rte_bitset.h>
#include <ethdev_driver.h>
#include <rte_rwlock.h>
#include <rte_interrupts.h>
@@ -2018,7 +2019,7 @@ struct mlx5_priv {
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- BITFIELD_DECLARE(mac_own, uint64_t, MLX5_MAX_MAC_ADDRESSES);
+ RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
/* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 2a91e02b45..7f6148f4e1 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1921,7 +1921,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
/* Add flows for unicast and multicast mac addresses added by API. */
if (!memcmp(mac, &cmp, sizeof(*mac)) ||
- !BITFIELD_ISSET(priv->mac_own, i) ||
+ !rte_bitset_test(priv->mac_own, i) ||
(dev->data->all_multicast && rte_is_multicast_ether_addr(mac)))
continue;
memcpy(&unicast.hdr.dst_addr.addr_bytes,
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index cf34e4e1d6..eaa6c25c09 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -699,8 +699,8 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i))
- BITFIELD_RESET(priv->mac_own, i);
+ if (rte_bitset_test(priv->mac_own, i))
+ rte_bitset_clear(priv->mac_own, i);
}
}
@@ -719,7 +719,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
struct mlx5_priv *priv = dev->data->dev_private;
if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -757,7 +757,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
return -ENOTSUP;
}
/* Mark this MAC address as owned by the PMD */
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return 0;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v6 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-08 9:27 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (2 preceding siblings ...)
2026-09-08 9:27 ` [PATCH v6 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
@ 2026-09-08 9:27 ` David Marchand
2026-09-11 8:59 ` Dariusz Sosnowski
2026-09-11 8:35 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux Dariusz Sosnowski
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-08 9:27 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Starting firmware version 22.49.1014, the number of mac addresses
per VF is not capped to 128 anymore.
The value can be increased via devlink:
$ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
cmode driverinit
$ devlink dev reload pci/0000:3b:00.2
On the DPDK side, we must retrieve the maximum number of unicast
and multicast addresses supported with a query to the firmware.
Then, dynamically allocate the mac addresses arrays and report the
limit instead of the previous hardcoded value.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v4:
- added RN update,
- fixed types of fields added to mlx5_hca_attr and mlx5_dev_cap,
- used RTE_BIT32,
- fixed mlx5_nl_mac_addr_sync inverted arguments,
---
doc/guides/rel_notes/release_26_11.rst | 5 +++
drivers/common/mlx5/mlx5_devx_cmds.c | 4 +++
drivers/common/mlx5/mlx5_devx_cmds.h | 2 ++
drivers/net/mlx5/linux/mlx5_os.c | 42 ++++++++++++++++++++------
drivers/net/mlx5/mlx5.c | 9 ++----
drivers/net/mlx5/mlx5.h | 8 +++--
drivers/net/mlx5/mlx5_ethdev.c | 2 +-
drivers/net/mlx5/mlx5_mac.c | 22 +++++++++-----
drivers/net/mlx5/mlx5_trigger.c | 10 +++---
drivers/net/mlx5/windows/mlx5_os.c | 40 +++++++++++++++++++-----
10 files changed, 104 insertions(+), 40 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 87c7e81bde..43043cd579 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,11 @@ New Features
Also, make sure to start the actual text at the margin.
=======================================================
+* **Updated NVIDIA mlx5 ethernet driver.**
+
+ * Increased the maximum number of secondary unicast MAC addresses from 128 to up to 4096
+ (depending on devlink configuration on the associated kernel netdevice).
+
Removed Items
-------------
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.c b/drivers/common/mlx5/mlx5_devx_cmds.c
index 140b057ab4..e5d9c92779 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.c
+++ b/drivers/common/mlx5/mlx5_devx_cmds.c
@@ -1110,6 +1110,10 @@ mlx5_devx_cmd_query_hca_attr(void *ctx,
attr->log_max_pd = MLX5_GET(cmd_hca_cap, hcattr, log_max_pd);
attr->log_max_srq = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq);
attr->log_max_srq_sz = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq_sz);
+ attr->log_max_current_uc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_uc_list);
+ attr->log_max_current_mc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_mc_list);
attr->reg_c_preserve =
MLX5_GET(cmd_hca_cap, hcattr, reg_c_preserve);
attr->mmo_regex_qp_en = MLX5_GET(cmd_hca_cap, hcattr, regexp_mmo_qp);
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.h b/drivers/common/mlx5/mlx5_devx_cmds.h
index 90beb2e9e6..504b4a4f64 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.h
+++ b/drivers/common/mlx5/mlx5_devx_cmds.h
@@ -356,6 +356,8 @@ struct mlx5_hca_attr {
uint8_t tx_sw_owner_v2:1;
uint8_t esw_sw_owner:1;
uint8_t esw_sw_owner_v2:1;
+ uint8_t log_max_current_uc_list:5;
+ uint8_t log_max_current_mc_list:5;
};
/* LAG Context. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 7d8ea4acba..35efb1d572 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -389,6 +389,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
sh->dev_cap.esw_info.regc_mask = 0;
#endif
sh->dev_cap.esw_info.is_set = 1;
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -1464,6 +1474,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
/* Some internal functions rely on Netlink sockets, open them now. */
priv->nl_socket_rdma = nl_rdma;
priv->nl_socket_route = mlx5_nl_init(NETLINK_ROUTE, 0);
@@ -1762,8 +1788,8 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
- MLX5_MAX_UC_MAC_ADDRESSES,
- MLX5_MAX_MAC_ADDRESSES);
+ sh->dev_cap.max_uc_mac_addrs,
+ sh->dev_cap.max_mac_addrs);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
TAILQ_INIT(&priv->flow_meters);
@@ -1963,17 +1989,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_flex_item_port_cleanup(eth_dev);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ eth_dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
if (eth_dev != NULL)
eth_dev->data->dev_private = NULL;
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
if (nl_rdma >= 0)
@@ -3531,7 +3555,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
const int vf = priv->sh->dev_cap.vf;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
diff --git a/drivers/net/mlx5/mlx5.c b/drivers/net/mlx5/mlx5.c
index c7b0d3ef8b..4bd1c6d1c0 100644
--- a/drivers/net/mlx5/mlx5.c
+++ b/drivers/net/mlx5/mlx5.c
@@ -2556,6 +2556,9 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_list_destroy(priv->hrxqs);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
sh->port[priv->dev_port - 1].nl_ih_port_id = RTE_MAX_ETHPORTS;
/*
* The interrupt handler port id must be reset before priv is reset
@@ -2590,12 +2593,6 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_flow_pools_destroy(priv);
memset(priv, 0, sizeof(*priv));
priv->domain_id = RTE_ETH_DEV_SWITCH_DOMAIN_ID_INVALID;
- /*
- * Reset mac_addrs to NULL such that it is not freed as part of
- * rte_eth_dev_release_port(). mac_addrs is part of dev_private so
- * it is freed when dev_private is freed.
- */
- dev->data->mac_addrs = NULL;
return 0;
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 7f20c811f3..70efcda9bb 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -217,6 +217,9 @@ struct mlx5_dev_cap {
} mprq; /* Capability for Multi-Packet RQ. */
char fw_ver[64]; /* Firmware version of this device. */
struct flow_hw_port_info esw_info; /* E-switch manager reg_c0. */
+ uint32_t max_uc_mac_addrs; /* Maximum unicast MAC addresses. */
+ uint32_t max_mc_mac_addrs; /* Maximum multicast MAC addresses. */
+ uint32_t max_mac_addrs; /* Total maximum MAC addresses. */
};
#define MLX5_MPESW_PORT_INVALID (-1)
@@ -2018,9 +2021,8 @@ struct mlx5_priv {
struct mlx5_dev_ctx_shared *sh; /* Shared device context. */
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
- struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
- /* Bit-field of MAC addresses owned by the PMD. */
+ struct rte_ether_addr *mac; /* MAC addresses. */
+ uint64_t *mac_own; /* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
/* Device properties. */
diff --git a/drivers/net/mlx5/mlx5_ethdev.c b/drivers/net/mlx5/mlx5_ethdev.c
index 8160d10e7e..306c1cd734 100644
--- a/drivers/net/mlx5/mlx5_ethdev.c
+++ b/drivers/net/mlx5/mlx5_ethdev.c
@@ -392,7 +392,7 @@ mlx5_dev_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *info)
max = RTE_MIN(max, (unsigned int)UINT16_MAX);
info->max_rx_queues = max;
info->max_tx_queues = max;
- info->max_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ info->max_mac_addrs = priv->sh->dev_cap.max_uc_mac_addrs;
info->rx_queue_offload_capa = mlx5_get_rx_queue_offloads(dev);
info->rx_seg_capa.max_nseg = MLX5_MAX_RXQ_NSEG;
info->rx_seg_capa.multi_pools = !priv->config.mprq.enabled;
diff --git a/drivers/net/mlx5/mlx5_mac.c b/drivers/net/mlx5/mlx5_mac.c
index 0e5d2be530..2ce7cfd407 100644
--- a/drivers/net/mlx5/mlx5_mac.c
+++ b/drivers/net/mlx5/mlx5_mac.c
@@ -36,7 +36,9 @@ mlx5_internal_mac_addr_remove(struct rte_eth_dev *dev,
uint32_t index,
struct rte_ether_addr *addr)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ struct mlx5_priv *priv = dev->data->dev_private;
+
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(&dev->data->mac_addrs[index]))
return false;
mlx5_os_mac_addr_remove(dev, index);
@@ -63,16 +65,17 @@ static int
mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
unsigned int i;
int ret;
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(mac)) {
rte_errno = EINVAL;
return -rte_errno;
}
/* First, make sure this address isn't already configured. */
- for (i = 0; (i != MLX5_MAX_MAC_ADDRESSES); ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
/* Skip this index, it's going to be reconfigured. */
if (i == index)
continue;
@@ -101,10 +104,11 @@ mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
void
mlx5_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
struct rte_ether_addr addr = { 0 };
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES)
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs)
return;
if (mlx5_internal_mac_addr_remove(dev, index, &addr)) {
ret = mlx5_traffic_mac_remove(dev, &addr);
@@ -133,9 +137,10 @@ int
mlx5_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index, uint32_t vmdq __rte_unused)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES) {
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs) {
rte_errno = EINVAL;
return -rte_errno;
}
@@ -217,16 +222,17 @@ int
mlx5_set_mc_addr_list(struct rte_eth_dev *dev,
struct rte_ether_addr *mc_addr_set, uint32_t nb_mc_addr)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
uint32_t i;
int ret;
- if (nb_mc_addr >= MLX5_MAX_MC_MAC_ADDRESSES) {
+ if (nb_mc_addr >= priv->sh->dev_cap.max_mc_mac_addrs) {
rte_errno = ENOSPC;
return -rte_errno;
}
- for (i = MLX5_MAX_UC_MAC_ADDRESSES; i != MLX5_MAX_MAC_ADDRESSES; ++i)
+ for (i = priv->sh->dev_cap.max_uc_mac_addrs; i != priv->sh->dev_cap.max_mac_addrs; ++i)
mlx5_internal_mac_addr_remove(dev, i, NULL);
- i = MLX5_MAX_UC_MAC_ADDRESSES;
+ i = priv->sh->dev_cap.max_uc_mac_addrs;
while (nb_mc_addr--) {
ret = mlx5_internal_mac_addr_add(dev, mc_addr_set++, i++);
if (ret)
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 7f6148f4e1..c5f493117b 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1916,7 +1916,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
}
}
/* Add MAC address flows. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
/* Add flows for unicast and multicast mac addresses added by API. */
@@ -2186,7 +2186,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
return 0;
/* Add all unicast DMAC flow rules with new VLAN attached. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2203,7 +2203,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
* Removing after creating VLAN rules so that traffic "gap" is not introduced.
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2241,7 +2241,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
* Recreating first to ensure no traffic "gap".
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2254,7 +2254,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
}
/* Remove all unicast DMAC flow rules with this VLAN. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index eaa6c25c09..fa429a824b 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -261,6 +261,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
MLX5_GET(initial_seg, pv_iseg, fw_rev_subminor));
DRV_LOG(DEBUG, "Packet pacing is not supported.");
mlx5_rt_timestamp_config(sh, hca_attr);
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -396,6 +406,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
priv->mp_id.port_id = port_id;
strlcpy(priv->mp_id.name, MLX5_MP_NAME, RTE_MP_MAX_NAME_LEN);
priv->representor = !!switch_info->representor;
@@ -612,17 +638,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_l3t_destroy(priv->mtr_profile_tbl);
if (own_domain_id)
claim_zero(rte_eth_switch_domain_free(priv->domain_id));
+ mlx5_free(priv->mac);
+ eth_dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
if (eth_dev != NULL)
eth_dev->data->dev_private = NULL;
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
MLX5_ASSERT(err > 0);
@@ -698,7 +722,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
struct mlx5_priv *priv = dev->data->dev_private;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i))
rte_bitset_clear(priv->mac_own, i);
}
@@ -718,7 +742,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
struct mlx5_priv *priv = dev->data->dev_private;
- if (index < MLX5_MAX_MAC_ADDRESSES)
+ if (index < priv->sh->dev_cap.max_mac_addrs)
rte_bitset_clear(priv->mac_own, index);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* RE: [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install
2026-09-04 12:28 ` [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install David Marchand
@ 2026-09-09 9:39 ` Loftus, Ciara
2026-09-11 14:14 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Loftus, Ciara @ 2026-09-09 9:39 UTC (permalink / raw)
To: David Marchand, dev@dpdk.org
Cc: Richardson, Bruce, stable@dpdk.org, Medvedkin, Vladimir
> Subject: [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install
>
> On port restart, all MAC addresses get pushed *twice* to the hardware,
> once by the driver and once by the eth_dev_mac_restore() in ethdev.
>
> On the other hand, MAC address filters are reset in the hardware
> by the PF only when a VF reset is triggered.
>
> Strictly speaking, the mac restore on port (re)start is unneeded,
> if no VF reset happened, so we can announce to ethdev that no mac
> restoration is needed via a get_restore_flags callback.
>
> Then, move the mac restoration to the VF reset handler.
>
> Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux
> driver")
> Cc: stable@dpdk.org
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
> Changes since v4:
> - rebased on next-net-intel,
>
> Changes since v4:
> - moved mac restoration in iavf_post_reset_reconfig,
>
> ---
> drivers/net/intel/iavf/iavf_ethdev.c | 27 ++++++++++++++++-----------
> 1 file changed, 16 insertions(+), 11 deletions(-)
>
> diff --git a/drivers/net/intel/iavf/iavf_ethdev.c
> b/drivers/net/intel/iavf/iavf_ethdev.c
> index bbd1f08ff0..bec7b3b6d7 100644
> --- a/drivers/net/intel/iavf/iavf_ethdev.c
> +++ b/drivers/net/intel/iavf/iavf_ethdev.c
> @@ -292,11 +292,12 @@ iavf_get_restore_flags(__rte_unused struct
> rte_eth_dev *dev,
> __rte_unused enum rte_eth_dev_operation op)
> {
> /*
> - * The unicast and multicast promiscuous settings persist across a
> + * The mac addresses, unicast and multicast promiscuous settings
> persist across a
> * stop/start; they are only cleared by a VF reset, which the driver
> * restores itself. So ethdev does not need to restore them on start.
> */
> - return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_PROMISC |
> + return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_MAC_ADDR |
> + RTE_ETH_RESTORE_PROMISC |
> RTE_ETH_RESTORE_ALLMULTI);
> }
>
> @@ -1095,15 +1096,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
> rte_intr_enable(intr_handle);
> }
>
> - /* Set all mac addrs */
> - iavf_add_del_all_mac_addr(adapter, true);
> -
> - if (!adapter->mac_primary_set)
> - adapter->mac_primary_set = true;
> -
> - /* Set all multicast addresses */
> - iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> >mc_addrs_num,
> - true);
> + if (!adapter->mac_primary_set) {
> + if (iavf_add_del_eth_addr(adapter, &dev->data-
> >mac_addrs[0], true,
> + VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
> + PMD_DRV_LOG(ERR, "failed to add primary MAC:"
> RTE_ETHER_ADDR_PRT_FMT,
> + RTE_ETHER_ADDR_BYTES(&dev->data-
> >mac_addrs[0]));
> + else
> + adapter->mac_primary_set = true;
> + }
>
> rte_spinlock_init(&vf->phc_time_aq_lock);
>
> @@ -3434,6 +3434,11 @@ iavf_post_reset_reconfig(struct rte_eth_dev
> *dev)
> int ret = 0;
> bool allmulti = false, allunicast = false;
> struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev-
> >data->dev_private);
> + struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data-
> >dev_private);
> +
> + /* After a VF reset, all MAC addresses got flushed, restore them. */
> + iavf_add_del_all_mac_addr(adapter, true);
Could this lead to a double install of the primary on the reset path? In
iavf_handle_hw_reset, dev_start can run before
iavf_post_reset_reconfig and can install the primary MAC and set
mac_primary_set. iavf_post_reset_reconfig then calls
iavf_add_del_all_mac_addr which can install the primary again as it
doesn't check mac_primary_set.
> + iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> >mc_addrs_num, true);
>
> /* Restore pre-reset unicast promiscuous and multicast promiscuous
> states */
> if (dev->data->promiscuous)
> --
> 2.54.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-09-04 12:28 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-09-04 12:28 ` [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install David Marchand
@ 2026-09-10 10:24 ` Burakov, Anatoly
2026-09-11 9:37 ` Burakov, Anatoly
2026-09-10 12:13 ` Burakov, Anatoly
2 siblings, 1 reply; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-10 10:24 UTC (permalink / raw)
To: David Marchand, dev; +Cc: bruce.richardson, Vladimir Medvedkin
On 9/4/2026 2:28 PM, David Marchand wrote:
> E810 hardware provides 32k switch lookups.
> Thanks to this, it is possible to allow a lot more secondary mac
> addresses than what is possible today.
>
> In practice, the maximum number of macs available per port may be lower
> and depends on usage by other (trusted?) VFs on the same PF.
> There is no way to figure out this limit but to try adding a mac address
> and get an error from the PF driver.
>
> Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
> accordingly.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
> Changes since v5:
> - separated from series that went in next-net,
> - rebased,
>
> Changes since v4:
> - rebased,
>
> Changes since v2:
> - added an entry in release notes,
> - removed unneeded temp variable,
>
> Changes since v1:
> - fixed buffer overflow on mailbox messages during port restart/VF reset,
>
Hi David,
<snip>
>
> -void
> -iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
> +static int
> +iavf_add_del_uc_addr_bulk(struct iavf_adapter *adapter, struct rte_ether_addr *addrs,
> + uint32_t nb_addrs, bool add)
> {
> +#define IAVF_ETH_ADDR_PER_REQ \
> + ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
> + sizeof(struct virtchnl_ether_addr))
> struct {
> struct virtchnl_ether_addr_list list;
> - struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
> - } list_req = {0};
> - struct virtchnl_ether_addr_list *list = &list_req.list;
> + struct virtchnl_ether_addr addr[IAVF_ETH_ADDR_PER_REQ];
> + } cmd_buffer;
> +#undef IAVF_ETH_ADDR_PER_REQ
> + struct virtchnl_ether_addr_list *list = &cmd_buffer.list;
> struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
> uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
> - struct iavf_cmd_info args = {0};
> - int err, i;
> - size_t buf_len;
>
> - for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
> - struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
> - struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
> + for (uint32_t i = 0; i < nb_addrs; i++) {
> + struct iavf_cmd_info args;
> + uint32_t batch;
> + int err;
>
> - /* ignore empty addresses */
> - if (rte_is_zero_ether_addr(addr))
> - continue;
> + batch = i % RTE_DIM(cmd_buffer.addr);
> +
> + if (batch == 0) {
> + memset(&cmd_buffer, 0, sizeof(cmd_buffer));
> + list->vsi_id = vf->vsi_res->vsi_id;
> + list->num_elements = 0;
> + }
> +
> + memcpy(list->list[batch].addr, addrs[i].addr_bytes,
> + sizeof(list->list[batch].addr));
> + list->list[batch].type = VIRTCHNL_ETHER_ADDR_EXTRA;
> list->num_elements++;
>
> - memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
> - vc_addr->type = (list->num_elements == 1) ?
> - VIRTCHNL_ETHER_ADDR_PRIMARY :
> - VIRTCHNL_ETHER_ADDR_EXTRA;
> + if (batch != RTE_DIM(cmd_buffer.addr) - 1 && i != nb_addrs - 1)
> + continue;
> +
> + memset(&args, 0, sizeof(args));
> + args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
> + args.in_args = (uint8_t *)list;
> + args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
> + sizeof(struct virtchnl_ether_addr) * list->num_elements;
> + args.out_buffer = msg_buf;
> + args.out_size = IAVF_AQ_BUF_SZ;
> + err = iavf_execute_vf_cmd_safe(adapter, &args);
> + if (err != 0) {
> + PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
> + add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
> + list->num_elements);
> + return err;
> + }
> +
> + PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
> + add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR",
> + list->num_elements);
> }
>
> - /* for some reason PF side checks for buffer being too big, so adjust it down */
> - buf_len = sizeof(struct virtchnl_ether_addr_list) +
> - sizeof(struct virtchnl_ether_addr) * list->num_elements;
> + return 0;
> +}
>
> - list->vsi_id = vf->vsi_res->vsi_id;
> - args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
> - args.in_args = (uint8_t *)list;
> - args.in_args_size = buf_len;
> - args.out_buffer = msg_buf;
> - args.out_size = IAVF_AQ_BUF_SZ;
> - err = iavf_execute_vf_cmd_safe(adapter, &args);
> - if (err)
> - PMD_DRV_LOG(ERR, "fail to execute command %s",
> - add ? "OP_ADD_ETHER_ADDRESS" : "OP_DEL_ETHER_ADDRESS");
> +void
> +iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
> +{
> + int start = -1;
> + int i;
> +
> + /* Handle primary address (index 0) separately */
> + if (!rte_is_zero_ether_addr(&adapter->dev_data->mac_addrs[0]))
> + iavf_add_del_eth_addr(adapter, &adapter->dev_data->mac_addrs[0], add,
> + VIRTCHNL_ETHER_ADDR_PRIMARY);
> +
> + /* Process secondary addresses in contiguous blocks */
> + for (i = 1; i < IAVF_UC_MACADDR_MAX; i++) {
> + struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
> +
> + if (!rte_is_zero_ether_addr(addr)) {
> + if (start == -1)
> + start = i;
> + continue;
> + }
> +
> + if (start != -1) {
> + iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
> + i - start, add);
> + start = -1;
> + }
> + }
> +
> + if (start != -1) {
> + iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data->mac_addrs[start],
> + i - start, add);
> + }
> }
>
> int
> @@ -2304,7 +2353,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
> struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
> uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
> uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
> - (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
> + (IAVF_MC_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
> struct virtchnl_ether_addr_list *list;
> struct iavf_cmd_info args;
> uint32_t i;
I would have preferred it if the caller managed the chunking, not the
"add_del_addr_bulk" function. There is precedent for this style of
refactor already [1], and I would like to keep things consistent - keep
the loop simple (without memsets etc.), and make the caller manage how
many addresses are being sent at once.
[1]
https://patches.dpdk.org/project/dpdk/patch/5e6a55afa2b45e3ee5ec17af7a6c548c96e9698b.1771945933.git.anatoly.burakov@intel.com/
This specific refactor is more about removing rte_malloc, but it does
also reorganize the loop in a way that I find to be more readable.
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-09-04 12:28 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-09-04 12:28 ` [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install David Marchand
2026-09-10 10:24 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses Burakov, Anatoly
@ 2026-09-10 12:13 ` Burakov, Anatoly
2026-09-10 12:20 ` Burakov, Anatoly
2 siblings, 1 reply; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-10 12:13 UTC (permalink / raw)
To: David Marchand, dev; +Cc: bruce.richardson, Vladimir Medvedkin
On 9/4/2026 2:28 PM, David Marchand wrote:
> E810 hardware provides 32k switch lookups.
> Thanks to this, it is possible to allow a lot more secondary mac
> addresses than what is possible today.
>
> In practice, the maximum number of macs available per port may be lower
> and depends on usage by other (trusted?) VFs on the same PF.
> There is no way to figure out this limit but to try adding a mac address
> and get an error from the PF driver.
>
> Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
> accordingly.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
On another note, I don't think this is even compatible with ethdev API.
In `ethdev_driver.h`:
struct __rte_cache_aligned rte_eth_dev_data {
...
/**
* Device Ethernet link addresses.
* All entries are unique.
* The first entry (index zero) is the default address.
*/
struct rte_ether_addr *mac_addrs;
/** Bitmap associating MAC addresses to pools */
uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
...
}
example of this bitmap in rte_eth_dev_mac_addr_remove:
/* Update NIC */
dev->dev_ops->mac_addr_remove(dev, index);
/* Update address in NIC data structure */
rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
/* reset pool bitmap */
dev->data->mac_pool_sel[index] = 0;
rte_ethdev_trace_mac_addr_remove(port_id, addr);
meaning, the mac_addrs array and the mac_pool_sel have the same
limitation because they are indexed by the same index.
RTE_ETH_NUM_RECEIVE_MAC_ADDR is defined as 128, so correct me if I'm
wrong here, but according to ethdev API one cannot have more than 128
MAC addresses?
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-09-10 12:13 ` Burakov, Anatoly
@ 2026-09-10 12:20 ` Burakov, Anatoly
2026-09-10 12:30 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-10 12:20 UTC (permalink / raw)
To: David Marchand, dev; +Cc: bruce.richardson, Vladimir Medvedkin
On 9/10/2026 2:13 PM, Burakov, Anatoly wrote:
> On 9/4/2026 2:28 PM, David Marchand wrote:
>> E810 hardware provides 32k switch lookups.
>> Thanks to this, it is possible to allow a lot more secondary mac
>> addresses than what is possible today.
>>
>> In practice, the maximum number of macs available per port may be lower
>> and depends on usage by other (trusted?) VFs on the same PF.
>> There is no way to figure out this limit but to try adding a mac address
>> and get an error from the PF driver.
>>
>> Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
>> accordingly.
>>
>> Signed-off-by: David Marchand <david.marchand@redhat.com>
>> ---
>
> On another note, I don't think this is even compatible with ethdev API.
>
> In `ethdev_driver.h`:
>
> struct __rte_cache_aligned rte_eth_dev_data {
> ...
> /**
> * Device Ethernet link addresses.
> * All entries are unique.
> * The first entry (index zero) is the default address.
> */
> struct rte_ether_addr *mac_addrs;
> /** Bitmap associating MAC addresses to pools */
> uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
> ...
> }
>
> example of this bitmap in rte_eth_dev_mac_addr_remove:
>
> /* Update NIC */
> dev->dev_ops->mac_addr_remove(dev, index);
>
> /* Update address in NIC data structure */
> rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
>
> /* reset pool bitmap */
> dev->data->mac_pool_sel[index] = 0;
>
> rte_ethdev_trace_mac_addr_remove(port_id, addr);
>
> meaning, the mac_addrs array and the mac_pool_sel have the same
> limitation because they are indexed by the same index.
>
> RTE_ETH_NUM_RECEIVE_MAC_ADDR is defined as 128, so correct me if I'm
> wrong here, but according to ethdev API one cannot have more than 128
> MAC addresses?
>
I would even go as far as to suggest that ethdev API should probably
check max MAC addrs number to make sure it doesn't exceed the size of
RTE_ETH_NUM_RECEIVE_MAC_ADDR, because otherwise that's a latent
potential buffer overrun?
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-09-10 12:20 ` Burakov, Anatoly
@ 2026-09-10 12:30 ` David Marchand
2026-09-10 12:38 ` Burakov, Anatoly
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-10 12:30 UTC (permalink / raw)
To: Burakov, Anatoly; +Cc: dev, bruce.richardson, Vladimir Medvedkin
On Thu, 10 Sept 2026 at 14:20, Burakov, Anatoly
<anatoly.burakov@intel.com> wrote:
>
> On 9/10/2026 2:13 PM, Burakov, Anatoly wrote:
> > On 9/4/2026 2:28 PM, David Marchand wrote:
> >> E810 hardware provides 32k switch lookups.
> >> Thanks to this, it is possible to allow a lot more secondary mac
> >> addresses than what is possible today.
> >>
> >> In practice, the maximum number of macs available per port may be lower
> >> and depends on usage by other (trusted?) VFs on the same PF.
> >> There is no way to figure out this limit but to try adding a mac address
> >> and get an error from the PF driver.
> >>
> >> Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
> >> accordingly.
> >>
> >> Signed-off-by: David Marchand <david.marchand@redhat.com>
> >> ---
> >
> > On another note, I don't think this is even compatible with ethdev API.
> >
> > In `ethdev_driver.h`:
> >
> > struct __rte_cache_aligned rte_eth_dev_data {
> > ...
> > /**
> > * Device Ethernet link addresses.
> > * All entries are unique.
> > * The first entry (index zero) is the default address.
> > */
> > struct rte_ether_addr *mac_addrs;
> > /** Bitmap associating MAC addresses to pools */
> > uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
> > ...
> > }
> >
> > example of this bitmap in rte_eth_dev_mac_addr_remove:
> >
> > /* Update NIC */
> > dev->dev_ops->mac_addr_remove(dev, index);
> >
> > /* Update address in NIC data structure */
> > rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
> >
> > /* reset pool bitmap */
> > dev->data->mac_pool_sel[index] = 0;
> >
> > rte_ethdev_trace_mac_addr_remove(port_id, addr);
> >
> > meaning, the mac_addrs array and the mac_pool_sel have the same
> > limitation because they are indexed by the same index.
> >
> > RTE_ETH_NUM_RECEIVE_MAC_ADDR is defined as 128, so correct me if I'm
> > wrong here, but according to ethdev API one cannot have more than 128
> > MAC addresses?
> >
>
> I would even go as far as to suggest that ethdev API should probably
> check max MAC addrs number to make sure it doesn't exceed the size of
> RTE_ETH_NUM_RECEIVE_MAC_ADDR, because otherwise that's a latent
> potential buffer overrun?
This limit is something that was put in place for VMDq.
I removed it in next-net (it did not hit main yet).
https://git.dpdk.org/next/dpdk-next-net/commit?id=f9ddb36e00655e38c115502d5ee180d4fda633c0
https://git.dpdk.org/next/dpdk-next-net/commit?id=31ea14ef354c8664482af27039ecd3c6130254b1
There may be some check missing in case a driver announces a
max_mac_addrs larger than RTE_ETH_NUM_RECEIVE_MAC_ADDR with VMDq
enabled.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-09-10 12:30 ` David Marchand
@ 2026-09-10 12:38 ` Burakov, Anatoly
0 siblings, 0 replies; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-10 12:38 UTC (permalink / raw)
To: David Marchand; +Cc: dev, bruce.richardson, Vladimir Medvedkin
On 9/10/2026 2:30 PM, David Marchand wrote:
> On Thu, 10 Sept 2026 at 14:20, Burakov, Anatoly
> <anatoly.burakov@intel.com> wrote:
>>
>> On 9/10/2026 2:13 PM, Burakov, Anatoly wrote:
>>> On 9/4/2026 2:28 PM, David Marchand wrote:
>>>> E810 hardware provides 32k switch lookups.
>>>> Thanks to this, it is possible to allow a lot more secondary mac
>>>> addresses than what is possible today.
>>>>
>>>> In practice, the maximum number of macs available per port may be lower
>>>> and depends on usage by other (trusted?) VFs on the same PF.
>>>> There is no way to figure out this limit but to try adding a mac address
>>>> and get an error from the PF driver.
>>>>
>>>> Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
>>>> accordingly.
>>>>
>>>> Signed-off-by: David Marchand <david.marchand@redhat.com>
>>>> ---
>>>
>>> On another note, I don't think this is even compatible with ethdev API.
>>>
>>> In `ethdev_driver.h`:
>>>
>>> struct __rte_cache_aligned rte_eth_dev_data {
>>> ...
>>> /**
>>> * Device Ethernet link addresses.
>>> * All entries are unique.
>>> * The first entry (index zero) is the default address.
>>> */
>>> struct rte_ether_addr *mac_addrs;
>>> /** Bitmap associating MAC addresses to pools */
>>> uint64_t mac_pool_sel[RTE_ETH_NUM_RECEIVE_MAC_ADDR];
>>> ...
>>> }
>>>
>>> example of this bitmap in rte_eth_dev_mac_addr_remove:
>>>
>>> /* Update NIC */
>>> dev->dev_ops->mac_addr_remove(dev, index);
>>>
>>> /* Update address in NIC data structure */
>>> rte_ether_addr_copy(&null_mac_addr, &dev->data->mac_addrs[index]);
>>>
>>> /* reset pool bitmap */
>>> dev->data->mac_pool_sel[index] = 0;
>>>
>>> rte_ethdev_trace_mac_addr_remove(port_id, addr);
>>>
>>> meaning, the mac_addrs array and the mac_pool_sel have the same
>>> limitation because they are indexed by the same index.
>>>
>>> RTE_ETH_NUM_RECEIVE_MAC_ADDR is defined as 128, so correct me if I'm
>>> wrong here, but according to ethdev API one cannot have more than 128
>>> MAC addresses?
>>>
>>
>> I would even go as far as to suggest that ethdev API should probably
>> check max MAC addrs number to make sure it doesn't exceed the size of
>> RTE_ETH_NUM_RECEIVE_MAC_ADDR, because otherwise that's a latent
>> potential buffer overrun?
>
> This limit is something that was put in place for VMDq.
>
> I removed it in next-net (it did not hit main yet).
> https://git.dpdk.org/next/dpdk-next-net/commit?id=f9ddb36e00655e38c115502d5ee180d4fda633c0
> https://git.dpdk.org/next/dpdk-next-net/commit?id=31ea14ef354c8664482af27039ecd3c6130254b1
>
> There may be some check missing in case a driver announces a
> max_mac_addrs larger than RTE_ETH_NUM_RECEIVE_MAC_ADDR with VMDq
> enabled.
>
>
OK, in that case I'll send AI to dig around ixgbe and see if we can
tighten enforcement of this limit :)
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-09-08 9:27 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (3 preceding siblings ...)
2026-09-08 9:27 ` [PATCH v6 5/5] net/mlx5: accept more unicast " David Marchand
@ 2026-09-11 8:35 ` Dariusz Sosnowski
4 siblings, 0 replies; 146+ messages in thread
From: Dariusz Sosnowski @ 2026-09-11 8:35 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao, Ori Kam,
Suanming Mou, Matan Azrad
On Tue, Sep 08, 2026 at 11:27:39AM +0200, David Marchand wrote:
> This helper is exposing internals of net/mlx5 for no good reason.
> All this code does is calling the remove helper.
> Walk through the list in Linux implementation like the Windows
> implementation.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 2/5] net/mlx5: remove redundant MAC address index checks
2026-09-08 9:27 ` [PATCH v6 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-09-11 8:36 ` Dariusz Sosnowski
0 siblings, 0 replies; 146+ messages in thread
From: Dariusz Sosnowski @ 2026-09-11 8:36 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao, Ori Kam,
Suanming Mou, Matan Azrad
On Tue, Sep 08, 2026 at 11:27:40AM +0200, David Marchand wrote:
> On the net/mlx5 side, mlx5_mac.c validates that any MAC address and its
> index is valid before calling the OS specific helpers.
> So those OS helpers do not have to validate again the index.
>
> Cascading this consideration, validating the MAC index against
> MLX5_MAX_MAC_ADDRESSES in common code is also redundant.
> The common code only deals with netlink, remove any index concern.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 3/5] net/mlx5: pass maximum number of unicast MAC to common code
2026-09-08 9:27 ` [PATCH v6 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-09-11 8:38 ` Dariusz Sosnowski
0 siblings, 0 replies; 146+ messages in thread
From: Dariusz Sosnowski @ 2026-09-11 8:38 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao, Ori Kam,
Suanming Mou, Matan Azrad
On Tue, Sep 08, 2026 at 11:27:41AM +0200, David Marchand wrote:
> Isolate how the MAC addresses array is walked through in the common code
> by passing the max index at which a unicast MAC address is stored in
> dev->data->mac_addrs[].
>
> With this change, only net/mlx5 knows about the max number of
> unicast/multicast MAC addresses.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 4/5] net/mlx5: use bitset for tracking MAC addresses
2026-09-08 9:27 ` [PATCH v6 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
@ 2026-09-11 8:40 ` Dariusz Sosnowski
0 siblings, 0 replies; 146+ messages in thread
From: Dariusz Sosnowski @ 2026-09-11 8:40 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao, Ori Kam,
Suanming Mou, Matan Azrad
On Tue, Sep 08, 2026 at 11:27:42AM +0200, David Marchand wrote:
> EAL provides bitset that does the same as this set of mlx5 macros.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-08 9:27 ` [PATCH v6 5/5] net/mlx5: accept more unicast " David Marchand
@ 2026-09-11 8:59 ` Dariusz Sosnowski
2026-09-11 9:55 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Dariusz Sosnowski @ 2026-09-11 8:59 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao, Ori Kam,
Suanming Mou, Matan Azrad
On Tue, Sep 08, 2026 at 11:27:43AM +0200, David Marchand wrote:
> Starting firmware version 22.49.1014, the number of mac addresses
> per VF is not capped to 128 anymore.
>
> The value can be increased via devlink:
> $ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
> cmode driverinit
> $ devlink dev reload pci/0000:3b:00.2
>
> On the DPDK side, we must retrieve the maximum number of unicast
> and multicast addresses supported with a query to the firmware.
>
> Then, dynamically allocate the mac addresses arrays and report the
> limit instead of the previous hardcoded value.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
> Changes since v4:
> - added RN update,
> - fixed types of fields added to mlx5_hca_attr and mlx5_dev_cap,
> - used RTE_BIT32,
> - fixed mlx5_nl_mac_addr_sync inverted arguments,
>
> ---
> doc/guides/rel_notes/release_26_11.rst | 5 +++
> drivers/common/mlx5/mlx5_devx_cmds.c | 4 +++
> drivers/common/mlx5/mlx5_devx_cmds.h | 2 ++
> drivers/net/mlx5/linux/mlx5_os.c | 42 ++++++++++++++++++++------
> drivers/net/mlx5/mlx5.c | 9 ++----
> drivers/net/mlx5/mlx5.h | 8 +++--
> drivers/net/mlx5/mlx5_ethdev.c | 2 +-
> drivers/net/mlx5/mlx5_mac.c | 22 +++++++++-----
> drivers/net/mlx5/mlx5_trigger.c | 10 +++---
> drivers/net/mlx5/windows/mlx5_os.c | 40 +++++++++++++++++++-----
> 10 files changed, 104 insertions(+), 40 deletions(-)
>
> diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
> index 87c7e81bde..43043cd579 100644
> --- a/doc/guides/rel_notes/release_26_11.rst
> +++ b/doc/guides/rel_notes/release_26_11.rst
> @@ -55,6 +55,11 @@ New Features
> Also, make sure to start the actual text at the margin.
> =======================================================
>
> +* **Updated NVIDIA mlx5 ethernet driver.**
> +
> + * Increased the maximum number of secondary unicast MAC addresses from 128 to up to 4096
> + (depending on devlink configuration on the associated kernel netdevice).
> +
>
> Removed Items
> -------------
> diff --git a/drivers/common/mlx5/mlx5_devx_cmds.c b/drivers/common/mlx5/mlx5_devx_cmds.c
> index 140b057ab4..e5d9c92779 100644
> --- a/drivers/common/mlx5/mlx5_devx_cmds.c
> +++ b/drivers/common/mlx5/mlx5_devx_cmds.c
> @@ -1110,6 +1110,10 @@ mlx5_devx_cmd_query_hca_attr(void *ctx,
> attr->log_max_pd = MLX5_GET(cmd_hca_cap, hcattr, log_max_pd);
> attr->log_max_srq = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq);
> attr->log_max_srq_sz = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq_sz);
> + attr->log_max_current_uc_list = MLX5_GET(cmd_hca_cap, hcattr,
> + log_max_current_uc_list);
> + attr->log_max_current_mc_list = MLX5_GET(cmd_hca_cap, hcattr,
> + log_max_current_mc_list);
> attr->reg_c_preserve =
> MLX5_GET(cmd_hca_cap, hcattr, reg_c_preserve);
> attr->mmo_regex_qp_en = MLX5_GET(cmd_hca_cap, hcattr, regexp_mmo_qp);
> diff --git a/drivers/common/mlx5/mlx5_devx_cmds.h b/drivers/common/mlx5/mlx5_devx_cmds.h
> index 90beb2e9e6..504b4a4f64 100644
> --- a/drivers/common/mlx5/mlx5_devx_cmds.h
> +++ b/drivers/common/mlx5/mlx5_devx_cmds.h
> @@ -356,6 +356,8 @@ struct mlx5_hca_attr {
> uint8_t tx_sw_owner_v2:1;
> uint8_t esw_sw_owner:1;
> uint8_t esw_sw_owner_v2:1;
> + uint8_t log_max_current_uc_list:5;
> + uint8_t log_max_current_mc_list:5;
> };
>
> /* LAG Context. */
> diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
> index 7d8ea4acba..35efb1d572 100644
> --- a/drivers/net/mlx5/linux/mlx5_os.c
> +++ b/drivers/net/mlx5/linux/mlx5_os.c
> @@ -389,6 +389,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
> sh->dev_cap.esw_info.regc_mask = 0;
> #endif
> sh->dev_cap.esw_info.is_set = 1;
> + if (hca_attr->log_max_current_uc_list > 0)
> + sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
> + else
> + sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
> + if (hca_attr->log_max_current_mc_list > 0)
> + sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
> + else
> + sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
> + sh->dev_cap.max_mac_addrs =
> + sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
> return 0;
> }
>
> @@ -1464,6 +1474,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
> priv->sh = sh;
> priv->dev_port = spawn->phys_port;
> priv->pci_dev = spawn->pci_dev;
> + priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
> + sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
> + RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
> + if (priv->mac == NULL) {
> + DRV_LOG(ERR, "Failed to allocate MAC address array.");
> + err = ENOMEM;
> + goto error;
> + }
> + priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
> + RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
> + RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
> + if (priv->mac_own == NULL) {
> + DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
> + err = ENOMEM;
> + goto error;
> + }
> /* Some internal functions rely on Netlink sockets, open them now. */
> priv->nl_socket_rdma = nl_rdma;
> priv->nl_socket_route = mlx5_nl_init(NETLINK_ROUTE, 0);
> @@ -1762,8 +1788,8 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
> mlx5_nl_mac_addr_sync(priv->nl_socket_route,
> mlx5_ifindex(eth_dev),
> eth_dev->data->mac_addrs,
> - MLX5_MAX_UC_MAC_ADDRESSES,
> - MLX5_MAX_MAC_ADDRESSES);
> + sh->dev_cap.max_uc_mac_addrs,
> + sh->dev_cap.max_mac_addrs);
> priv->ctrl_flows = 0;
> rte_spinlock_init(&priv->flow_list_lock);
> TAILQ_INIT(&priv->flow_meters);
> @@ -1963,17 +1989,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
> mlx5_flex_item_port_cleanup(eth_dev);
> mlx5_free(priv->ext_rxqs);
> mlx5_free(priv->ext_txqs);
> + mlx5_free(priv->mac);
> + eth_dev->data->mac_addrs = NULL;
This can segfault because eth_dev is allocated later than priv.
Previous version of the error rollback accounted for that
(where mac_addrs was reset only if eth_dev != NULL).
The same logic applies in Windows code.
Could you please revert that?
> + mlx5_free(priv->mac_own);
> mlx5_free(priv);
> if (eth_dev != NULL)
> eth_dev->data->dev_private = NULL;
> }
> - if (eth_dev != NULL) {
> - /* mac_addrs must not be freed alone because part of
> - * dev_private
> - **/
> - eth_dev->data->mac_addrs = NULL;
> + if (eth_dev != NULL)
> rte_eth_dev_release_port(eth_dev);
> - }
> if (sh)
> mlx5_free_shared_dev_ctx(sh);
> if (nl_rdma >= 0)
> @@ -3531,7 +3555,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
> const int vf = priv->sh->dev_cap.vf;
> int i;
>
> - for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
> + for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
> if (rte_bitset_test(priv->mac_own, i)) {
> if (vf)
> mlx5_nl_mac_addr_remove(priv->nl_socket_route,
> diff --git a/drivers/net/mlx5/mlx5.c b/drivers/net/mlx5/mlx5.c
> index c7b0d3ef8b..4bd1c6d1c0 100644
> --- a/drivers/net/mlx5/mlx5.c
> +++ b/drivers/net/mlx5/mlx5.c
> @@ -2556,6 +2556,9 @@ mlx5_dev_close(struct rte_eth_dev *dev)
> mlx5_list_destroy(priv->hrxqs);
> mlx5_free(priv->ext_rxqs);
> mlx5_free(priv->ext_txqs);
> + mlx5_free(priv->mac);
> + dev->data->mac_addrs = NULL;
> + mlx5_free(priv->mac_own);
> sh->port[priv->dev_port - 1].nl_ih_port_id = RTE_MAX_ETHPORTS;
> /*
> * The interrupt handler port id must be reset before priv is reset
> @@ -2590,12 +2593,6 @@ mlx5_dev_close(struct rte_eth_dev *dev)
> mlx5_flow_pools_destroy(priv);
> memset(priv, 0, sizeof(*priv));
> priv->domain_id = RTE_ETH_DEV_SWITCH_DOMAIN_ID_INVALID;
> - /*
> - * Reset mac_addrs to NULL such that it is not freed as part of
> - * rte_eth_dev_release_port(). mac_addrs is part of dev_private so
> - * it is freed when dev_private is freed.
> - */
> - dev->data->mac_addrs = NULL;
> return 0;
> }
>
> diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
> index 7f20c811f3..70efcda9bb 100644
> --- a/drivers/net/mlx5/mlx5.h
> +++ b/drivers/net/mlx5/mlx5.h
> @@ -217,6 +217,9 @@ struct mlx5_dev_cap {
> } mprq; /* Capability for Multi-Packet RQ. */
> char fw_ver[64]; /* Firmware version of this device. */
> struct flow_hw_port_info esw_info; /* E-switch manager reg_c0. */
> + uint32_t max_uc_mac_addrs; /* Maximum unicast MAC addresses. */
> + uint32_t max_mc_mac_addrs; /* Maximum multicast MAC addresses. */
> + uint32_t max_mac_addrs; /* Total maximum MAC addresses. */
> };
>
> #define MLX5_MPESW_PORT_INVALID (-1)
> @@ -2018,9 +2021,8 @@ struct mlx5_priv {
> struct mlx5_dev_ctx_shared *sh; /* Shared device context. */
> uint32_t dev_port; /* Device port number. */
> struct rte_pci_device *pci_dev; /* Backend PCI device. */
> - struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
> - RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
> - /* Bit-field of MAC addresses owned by the PMD. */
> + struct rte_ether_addr *mac; /* MAC addresses. */
> + uint64_t *mac_own; /* Bit-field of MAC addresses owned by the PMD. */
> uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
> unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
> /* Device properties. */
> diff --git a/drivers/net/mlx5/mlx5_ethdev.c b/drivers/net/mlx5/mlx5_ethdev.c
> index 8160d10e7e..306c1cd734 100644
> --- a/drivers/net/mlx5/mlx5_ethdev.c
> +++ b/drivers/net/mlx5/mlx5_ethdev.c
> @@ -392,7 +392,7 @@ mlx5_dev_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *info)
> max = RTE_MIN(max, (unsigned int)UINT16_MAX);
> info->max_rx_queues = max;
> info->max_tx_queues = max;
> - info->max_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
> + info->max_mac_addrs = priv->sh->dev_cap.max_uc_mac_addrs;
> info->rx_queue_offload_capa = mlx5_get_rx_queue_offloads(dev);
> info->rx_seg_capa.max_nseg = MLX5_MAX_RXQ_NSEG;
> info->rx_seg_capa.multi_pools = !priv->config.mprq.enabled;
> diff --git a/drivers/net/mlx5/mlx5_mac.c b/drivers/net/mlx5/mlx5_mac.c
> index 0e5d2be530..2ce7cfd407 100644
> --- a/drivers/net/mlx5/mlx5_mac.c
> +++ b/drivers/net/mlx5/mlx5_mac.c
> @@ -36,7 +36,9 @@ mlx5_internal_mac_addr_remove(struct rte_eth_dev *dev,
> uint32_t index,
> struct rte_ether_addr *addr)
> {
> - MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
> + struct mlx5_priv *priv = dev->data->dev_private;
> +
> + MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
Coverity in our internal CI reports the following:
drivers/net/mlx5/mlx5_mac.c:39:20: CID 910875 (#1 of 1):
Type: Parse warning (PW.SET_BUT_NOT_USED)
Classification: Unclassified
Severity: Unspecified
Action: Undecided
Owner: Unassigned
Defect only exists locally.
drivers/net/mlx5/mlx5_mac.c:39:20:
1. set_but_not_used: variable "priv" was set but never used
priv variable can be removed using MLX5_SH macro:
MLX5_ASSERT(index < MLX5_SH(dev)->dev_cap.max_mac_addrs);
> if (rte_is_zero_ether_addr(&dev->data->mac_addrs[index]))
> return false;
> mlx5_os_mac_addr_remove(dev, index);
> @@ -63,16 +65,17 @@ static int
> mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
> uint32_t index)
> {
> + struct mlx5_priv *priv = dev->data->dev_private;
> unsigned int i;
> int ret;
>
> - MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
> + MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
> if (rte_is_zero_ether_addr(mac)) {
> rte_errno = EINVAL;
> return -rte_errno;
> }
> /* First, make sure this address isn't already configured. */
> - for (i = 0; (i != MLX5_MAX_MAC_ADDRESSES); ++i) {
> + for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
> /* Skip this index, it's going to be reconfigured. */
> if (i == index)
> continue;
> @@ -101,10 +104,11 @@ mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
> void
> mlx5_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
> {
> + struct mlx5_priv *priv = dev->data->dev_private;
> struct rte_ether_addr addr = { 0 };
> int ret;
>
> - if (index >= MLX5_MAX_UC_MAC_ADDRESSES)
> + if (index >= priv->sh->dev_cap.max_uc_mac_addrs)
> return;
> if (mlx5_internal_mac_addr_remove(dev, index, &addr)) {
> ret = mlx5_traffic_mac_remove(dev, &addr);
> @@ -133,9 +137,10 @@ int
> mlx5_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
> uint32_t index, uint32_t vmdq __rte_unused)
> {
> + struct mlx5_priv *priv = dev->data->dev_private;
> int ret;
>
> - if (index >= MLX5_MAX_UC_MAC_ADDRESSES) {
> + if (index >= priv->sh->dev_cap.max_uc_mac_addrs) {
> rte_errno = EINVAL;
> return -rte_errno;
> }
> @@ -217,16 +222,17 @@ int
> mlx5_set_mc_addr_list(struct rte_eth_dev *dev,
> struct rte_ether_addr *mc_addr_set, uint32_t nb_mc_addr)
> {
> + struct mlx5_priv *priv = dev->data->dev_private;
> uint32_t i;
> int ret;
>
> - if (nb_mc_addr >= MLX5_MAX_MC_MAC_ADDRESSES) {
> + if (nb_mc_addr >= priv->sh->dev_cap.max_mc_mac_addrs) {
> rte_errno = ENOSPC;
> return -rte_errno;
> }
> - for (i = MLX5_MAX_UC_MAC_ADDRESSES; i != MLX5_MAX_MAC_ADDRESSES; ++i)
> + for (i = priv->sh->dev_cap.max_uc_mac_addrs; i != priv->sh->dev_cap.max_mac_addrs; ++i)
> mlx5_internal_mac_addr_remove(dev, i, NULL);
> - i = MLX5_MAX_UC_MAC_ADDRESSES;
> + i = priv->sh->dev_cap.max_uc_mac_addrs;
> while (nb_mc_addr--) {
> ret = mlx5_internal_mac_addr_add(dev, mc_addr_set++, i++);
> if (ret)
> diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
> index 7f6148f4e1..c5f493117b 100644
> --- a/drivers/net/mlx5/mlx5_trigger.c
> +++ b/drivers/net/mlx5/mlx5_trigger.c
> @@ -1916,7 +1916,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
> }
> }
> /* Add MAC address flows. */
> - for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
> + for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
> struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
>
> /* Add flows for unicast and multicast mac addresses added by API. */
> @@ -2186,7 +2186,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
> return 0;
>
> /* Add all unicast DMAC flow rules with new VLAN attached. */
> - for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
> + for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
> struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
>
> if (rte_is_zero_ether_addr(mac))
> @@ -2203,7 +2203,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
> * Removing after creating VLAN rules so that traffic "gap" is not introduced.
> */
>
> - for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
> + for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
> struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
>
> if (rte_is_zero_ether_addr(mac))
> @@ -2241,7 +2241,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
> * Recreating first to ensure no traffic "gap".
> */
>
> - for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
> + for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
> struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
>
> if (rte_is_zero_ether_addr(mac))
> @@ -2254,7 +2254,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
> }
>
> /* Remove all unicast DMAC flow rules with this VLAN. */
> - for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
> + for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
> struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
>
> if (rte_is_zero_ether_addr(mac))
> diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
> index eaa6c25c09..fa429a824b 100644
> --- a/drivers/net/mlx5/windows/mlx5_os.c
> +++ b/drivers/net/mlx5/windows/mlx5_os.c
> @@ -261,6 +261,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
> MLX5_GET(initial_seg, pv_iseg, fw_rev_subminor));
> DRV_LOG(DEBUG, "Packet pacing is not supported.");
> mlx5_rt_timestamp_config(sh, hca_attr);
> + if (hca_attr->log_max_current_uc_list > 0)
> + sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
> + else
> + sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
> + if (hca_attr->log_max_current_mc_list > 0)
> + sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
> + else
> + sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
> + sh->dev_cap.max_mac_addrs =
> + sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
> return 0;
> }
>
> @@ -396,6 +406,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
> priv->sh = sh;
> priv->dev_port = spawn->phys_port;
> priv->pci_dev = spawn->pci_dev;
> + priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
> + sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
> + RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
> + if (priv->mac == NULL) {
> + DRV_LOG(ERR, "Failed to allocate MAC address array.");
> + err = ENOMEM;
> + goto error;
> + }
> + priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
> + RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
> + RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
> + if (priv->mac_own == NULL) {
> + DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
> + err = ENOMEM;
> + goto error;
> + }
> priv->mp_id.port_id = port_id;
> strlcpy(priv->mp_id.name, MLX5_MP_NAME, RTE_MP_MAX_NAME_LEN);
> priv->representor = !!switch_info->representor;
> @@ -612,17 +638,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
> mlx5_l3t_destroy(priv->mtr_profile_tbl);
> if (own_domain_id)
> claim_zero(rte_eth_switch_domain_free(priv->domain_id));
> + mlx5_free(priv->mac);
> + eth_dev->data->mac_addrs = NULL;
> + mlx5_free(priv->mac_own);
> mlx5_free(priv);
> if (eth_dev != NULL)
> eth_dev->data->dev_private = NULL;
> }
> - if (eth_dev != NULL) {
> - /* mac_addrs must not be freed alone because part of
> - * dev_private
> - **/
> - eth_dev->data->mac_addrs = NULL;
> + if (eth_dev != NULL)
> rte_eth_dev_release_port(eth_dev);
> - }
> if (sh)
> mlx5_free_shared_dev_ctx(sh);
> MLX5_ASSERT(err > 0);
> @@ -698,7 +722,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
> struct mlx5_priv *priv = dev->data->dev_private;
> int i;
>
> - for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
> + for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
> if (rte_bitset_test(priv->mac_own, i))
> rte_bitset_clear(priv->mac_own, i);
> }
> @@ -718,7 +742,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
> {
> struct mlx5_priv *priv = dev->data->dev_private;
>
> - if (index < MLX5_MAX_MAC_ADDRESSES)
> + if (index < priv->sh->dev_cap.max_mac_addrs)
> rte_bitset_clear(priv->mac_own, index);
> }
>
> --
> 2.54.0
>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-09-10 10:24 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses Burakov, Anatoly
@ 2026-09-11 9:37 ` Burakov, Anatoly
2026-09-11 11:52 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-11 9:37 UTC (permalink / raw)
To: David Marchand, dev; +Cc: bruce.richardson, Vladimir Medvedkin
On 9/10/2026 12:24 PM, Burakov, Anatoly wrote:
> On 9/4/2026 2:28 PM, David Marchand wrote:
>> E810 hardware provides 32k switch lookups.
>> Thanks to this, it is possible to allow a lot more secondary mac
>> addresses than what is possible today.
>>
>> In practice, the maximum number of macs available per port may be lower
>> and depends on usage by other (trusted?) VFs on the same PF.
>> There is no way to figure out this limit but to try adding a mac address
>> and get an error from the PF driver.
>>
>> Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
>> accordingly.
>>
>> Signed-off-by: David Marchand <david.marchand@redhat.com>
>> ---
>> Changes since v5:
>> - separated from series that went in next-net,
>> - rebased,
>>
>> Changes since v4:
>> - rebased,
>>
>> Changes since v2:
>> - added an entry in release notes,
>> - removed unneeded temp variable,
>>
>> Changes since v1:
>> - fixed buffer overflow on mailbox messages during port restart/VF reset,
>>
>
> Hi David,
>
> <snip>
>
>> -void
>> -iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
>> +static int
>> +iavf_add_del_uc_addr_bulk(struct iavf_adapter *adapter, struct
>> rte_ether_addr *addrs,
>> + uint32_t nb_addrs, bool add)
>> {
>> +#define IAVF_ETH_ADDR_PER_REQ \
>> + ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
>> + sizeof(struct virtchnl_ether_addr))
>> struct {
>> struct virtchnl_ether_addr_list list;
>> - struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
>> - } list_req = {0};
>> - struct virtchnl_ether_addr_list *list = &list_req.list;
>> + struct virtchnl_ether_addr addr[IAVF_ETH_ADDR_PER_REQ];
>> + } cmd_buffer;
>> +#undef IAVF_ETH_ADDR_PER_REQ
>> + struct virtchnl_ether_addr_list *list = &cmd_buffer.list;
>> struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
>> uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
>> - struct iavf_cmd_info args = {0};
>> - int err, i;
>> - size_t buf_len;
>> - for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
>> - struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
>> - struct virtchnl_ether_addr *vc_addr = &list->list[list-
>> >num_elements];
>> + for (uint32_t i = 0; i < nb_addrs; i++) {
>> + struct iavf_cmd_info args;
>> + uint32_t batch;
>> + int err;
>> - /* ignore empty addresses */
>> - if (rte_is_zero_ether_addr(addr))
>> - continue;
>> + batch = i % RTE_DIM(cmd_buffer.addr);
>> +
>> + if (batch == 0) {
>> + memset(&cmd_buffer, 0, sizeof(cmd_buffer));
>> + list->vsi_id = vf->vsi_res->vsi_id;
>> + list->num_elements = 0;
>> + }
>> +
>> + memcpy(list->list[batch].addr, addrs[i].addr_bytes,
>> + sizeof(list->list[batch].addr));
>> + list->list[batch].type = VIRTCHNL_ETHER_ADDR_EXTRA;
>> list->num_elements++;
>> - memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr-
>> >addr_bytes));
>> - vc_addr->type = (list->num_elements == 1) ?
>> - VIRTCHNL_ETHER_ADDR_PRIMARY :
>> - VIRTCHNL_ETHER_ADDR_EXTRA;
>> + if (batch != RTE_DIM(cmd_buffer.addr) - 1 && i != nb_addrs - 1)
>> + continue;
>> +
>> + memset(&args, 0, sizeof(args));
>> + args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR :
>> VIRTCHNL_OP_DEL_ETH_ADDR;
>> + args.in_args = (uint8_t *)list;
>> + args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
>> + sizeof(struct virtchnl_ether_addr) * list->num_elements;
>> + args.out_buffer = msg_buf;
>> + args.out_size = IAVF_AQ_BUF_SZ;
>> + err = iavf_execute_vf_cmd_safe(adapter, &args);
>> + if (err != 0) {
>> + PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
>> + add ? "VIRTCHNL_OP_ADD_ETH_ADDR" :
>> "VIRTCHNL_OP_DEL_ETH_ADDR",
>> + list->num_elements);
>> + return err;
>> + }
>> +
>> + PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
>> + add ? "VIRTCHNL_OP_ADD_ETH_ADDR" :
>> "VIRTCHNL_OP_DEL_ETH_ADDR",
>> + list->num_elements);
>> }
>> - /* for some reason PF side checks for buffer being too big, so
>> adjust it down */
>> - buf_len = sizeof(struct virtchnl_ether_addr_list) +
>> - sizeof(struct virtchnl_ether_addr) * list->num_elements;
>> + return 0;
>> +}
>> - list->vsi_id = vf->vsi_res->vsi_id;
>> - args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR :
>> VIRTCHNL_OP_DEL_ETH_ADDR;
>> - args.in_args = (uint8_t *)list;
>> - args.in_args_size = buf_len;
>> - args.out_buffer = msg_buf;
>> - args.out_size = IAVF_AQ_BUF_SZ;
>> - err = iavf_execute_vf_cmd_safe(adapter, &args);
>> - if (err)
>> - PMD_DRV_LOG(ERR, "fail to execute command %s",
>> - add ? "OP_ADD_ETHER_ADDRESS" : "OP_DEL_ETHER_ADDRESS");
>> +void
>> +iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
>> +{
>> + int start = -1;
>> + int i;
>> +
>> + /* Handle primary address (index 0) separately */
>> + if (!rte_is_zero_ether_addr(&adapter->dev_data->mac_addrs[0]))
>> + iavf_add_del_eth_addr(adapter, &adapter->dev_data-
>> >mac_addrs[0], add,
>> + VIRTCHNL_ETHER_ADDR_PRIMARY);
>> +
>> + /* Process secondary addresses in contiguous blocks */
>> + for (i = 1; i < IAVF_UC_MACADDR_MAX; i++) {
>> + struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
>> +
>> + if (!rte_is_zero_ether_addr(addr)) {
>> + if (start == -1)
>> + start = i;
>> + continue;
>> + }
>> +
>> + if (start != -1) {
>> + iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data-
>> >mac_addrs[start],
>> + i - start, add);
>> + start = -1;
>> + }
>> + }
>> +
>> + if (start != -1) {
>> + iavf_add_del_uc_addr_bulk(adapter, &adapter->dev_data-
>> >mac_addrs[start],
>> + i - start, add);
>> + }
>> }
>> int
>> @@ -2304,7 +2353,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter
>> *adapter,
>> struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
>> uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
>> uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
>> - (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
>> + (IAVF_MC_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
>> struct virtchnl_ether_addr_list *list;
>> struct iavf_cmd_info args;
>> uint32_t i;
>
> I would have preferred it if the caller managed the chunking, not the
> "add_del_addr_bulk" function. There is precedent for this style of
> refactor already [1], and I would like to keep things consistent - keep
> the loop simple (without memsets etc.), and make the caller manage how
> many addresses are being sent at once.
>
> [1] https://patches.dpdk.org/project/dpdk/
> patch/5e6a55afa2b45e3ee5ec17af7a6c548c96e9698b.1771945933.git.anatoly.burakov@intel.com/
>
> This specific refactor is more about removing rte_malloc, but it does
> also reorganize the loop in a way that I find to be more readable.
>
I tried prototyping a loop, and realized that the fact that MAC address
list has holes in it is making things a little difficult, but here's
what I came up with as an alternative implementation, I think it's a
little clearer:
```
#define IAVF_ETH_ADDR_PER_REQ \
((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
sizeof(struct virtchnl_ether_addr))
struct iavf_eth_addr_cmd {
struct virtchnl_ether_addr_list list;
struct virtchnl_ether_addr extra[IAVF_ETH_ADDR_PER_REQ];
};
static int
iavf_send_uc_addr_list(struct iavf_adapter *adapter,
struct virtchnl_ether_addr_list *list, bool add)
{
const char *opname = add ? "VIRTCHNL_OP_ADD_ETH_ADDR" :
"VIRTCHNL_OP_DEL_ETH_ADDR";
uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
struct iavf_cmd_info args = {0};
int err;
args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
args.in_args = (uint8_t *)list;
args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
sizeof(struct virtchnl_ether_addr) * list->num_elements;
args.out_buffer = msg_buf;
args.out_size = IAVF_AQ_BUF_SZ;
err = iavf_execute_vf_cmd_safe(adapter, &args);
if (err != 0)
PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
opname, list->num_elements);
else
PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
opname, list->num_elements);
return err;
}
void
iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
{
struct rte_ether_addr *addrs = adapter->dev_data->mac_addrs;
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint32_t idx = 1;
/* Handle primary address (index 0) separately */
if (!rte_is_zero_ether_addr(&addrs[0]))
iavf_add_del_eth_addr(adapter, &addrs[0], add,
VIRTCHNL_ETHER_ADDR_PRIMARY);
/* the secondary address list is sparse, so gather it into full batches */
while (idx < IAVF_UC_MACADDR_MAX) {
struct iavf_eth_addr_cmd cmd = {0};
uint16_t nb_addrs = 0;
for (; idx < IAVF_UC_MACADDR_MAX && nb_addrs < IAVF_ETH_ADDR_PER_REQ;
idx++) {
if (rte_is_zero_ether_addr(&addrs[idx]))
continue;
memcpy(cmd.list.list[nb_addrs].addr, addrs[idx].addr_bytes,
sizeof(cmd.list.list[nb_addrs].addr));
cmd.list.list[nb_addrs].type = VIRTCHNL_ETHER_ADDR_EXTRA;
nb_addrs++;
}
if (nb_addrs == 0)
break;
cmd.list.vsi_id = vf->vsi_res->vsi_id;
cmd.list.num_elements = nb_addrs;
if (iavf_send_uc_addr_list(adapter, &cmd.list, add) != 0)
break;
}
}
```
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-11 8:59 ` Dariusz Sosnowski
@ 2026-09-11 9:55 ` David Marchand
2026-09-11 10:01 ` Dariusz Sosnowski
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-11 9:55 UTC (permalink / raw)
To: Dariusz Sosnowski
Cc: dev, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao, Ori Kam,
Suanming Mou, Matan Azrad
Hello Dariusz,
Thanks for the review.
On Fri, 11 Sept 2026 at 11:00, Dariusz Sosnowski <dsosnowski@nvidia.com> wrote:
> > @@ -1963,17 +1989,15 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
> > mlx5_flex_item_port_cleanup(eth_dev);
> > mlx5_free(priv->ext_rxqs);
> > mlx5_free(priv->ext_txqs);
> > + mlx5_free(priv->mac);
> > + eth_dev->data->mac_addrs = NULL;
>
> This can segfault because eth_dev is allocated later than priv.
> Previous version of the error rollback accounted for that
> (where mac_addrs was reset only if eth_dev != NULL).
> The same logic applies in Windows code.
>
> Could you please revert that?
Well, the code before looked fishy to me.
Resetting mac_addrs if (eth_dev != NULL) outside of the if (priv !=
NULL) block seems incorrect.
> > + mlx5_free(priv->mac_own);
> > mlx5_free(priv);
> > if (eth_dev != NULL)
> > eth_dev->data->dev_private = NULL;
I would rather reset mac_addrs to NULL along the dev_private reset in
the if (eth_dev != NULL) block right after.
The comment about mac_addrs (see below) being in dev_private can also
be removed.
WDYT?
> > }
> > - if (eth_dev != NULL) {
> > - /* mac_addrs must not be freed alone because part of
> > - * dev_private
> > - **/
> > - eth_dev->data->mac_addrs = NULL;
> > + if (eth_dev != NULL)
> > rte_eth_dev_release_port(eth_dev);
> > - }
> > if (sh)
> > mlx5_free_shared_dev_ctx(sh);
> > if (nl_rdma >= 0)
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* RE: [PATCH v6 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-11 9:55 ` David Marchand
@ 2026-09-11 10:01 ` Dariusz Sosnowski
0 siblings, 0 replies; 146+ messages in thread
From: Dariusz Sosnowski @ 2026-09-11 10:01 UTC (permalink / raw)
To: David Marchand
Cc: dev@dpdk.org, rjarry@redhat.com, cfontain@redhat.com,
Slava Ovsiienko, Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
> -----Original Message-----
> From: David Marchand <david.marchand@redhat.com>
> Sent: Friday, September 11, 2026 11:55 AM
> To: Dariusz Sosnowski <dsosnowski@nvidia.com>
> Cc: dev@dpdk.org; rjarry@redhat.com; cfontain@redhat.com; Slava
> Ovsiienko <viacheslavo@nvidia.com>; Bing Zhao <bingz@nvidia.com>; Ori
> Kam <orika@nvidia.com>; Suanming Mou <suanmingm@nvidia.com>; Matan
> Azrad <matan@nvidia.com>
> Subject: Re: [PATCH v6 5/5] net/mlx5: accept more unicast MAC addresses
>
> External email: Use caution opening links or attachments
>
>
> Hello Dariusz,
>
> Thanks for the review.
>
> On Fri, 11 Sept 2026 at 11:00, Dariusz Sosnowski <dsosnowski@nvidia.com>
> wrote:
> > > @@ -1963,17 +1989,15 @@ mlx5_dev_spawn(struct rte_device
> *dpdk_dev,
> > > mlx5_flex_item_port_cleanup(eth_dev);
> > > mlx5_free(priv->ext_rxqs);
> > > mlx5_free(priv->ext_txqs);
> > > + mlx5_free(priv->mac);
> > > + eth_dev->data->mac_addrs = NULL;
> >
> > This can segfault because eth_dev is allocated later than priv.
> > Previous version of the error rollback accounted for that (where
> > mac_addrs was reset only if eth_dev != NULL).
> > The same logic applies in Windows code.
> >
> > Could you please revert that?
>
> Well, the code before looked fishy to me.
> Resetting mac_addrs if (eth_dev != NULL) outside of the if (priv !=
> NULL) block seems incorrect.
>
>
> > > + mlx5_free(priv->mac_own);
> > > mlx5_free(priv);
> > > if (eth_dev != NULL)
> > > eth_dev->data->dev_private = NULL;
>
> I would rather reset mac_addrs to NULL along the dev_private reset in the if
> (eth_dev != NULL) block right after.
> The comment about mac_addrs (see below) being in dev_private can also be
> removed.
>
> WDYT?
Resetting it along with dev_private reset sounds good to me.
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-09-11 9:37 ` Burakov, Anatoly
@ 2026-09-11 11:52 ` David Marchand
2026-09-11 12:14 ` Burakov, Anatoly
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-11 11:52 UTC (permalink / raw)
To: Burakov, Anatoly; +Cc: dev, bruce.richardson, Vladimir Medvedkin
On Fri, 11 Sept 2026 at 11:37, Burakov, Anatoly
<anatoly.burakov@intel.com> wrote:
> > I would have preferred it if the caller managed the chunking, not the
> > "add_del_addr_bulk" function. There is precedent for this style of
> > refactor already [1], and I would like to keep things consistent - keep
> > the loop simple (without memsets etc.), and make the caller manage how
> > many addresses are being sent at once.
> >
> > [1] https://patches.dpdk.org/project/dpdk/
> > patch/5e6a55afa2b45e3ee5ec17af7a6c548c96e9698b.1771945933.git.anatoly.burakov@intel.com/
> >
> > This specific refactor is more about removing rte_malloc, but it does
> > also reorganize the loop in a way that I find to be more readable.
> >
>
> I tried prototyping a loop, and realized that the fact that MAC address
> list has holes in it is making things a little difficult, but here's
> what I came up with as an alternative implementation, I think it's a
> little clearer:
>
> ```
> #define IAVF_ETH_ADDR_PER_REQ \
> ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
> sizeof(struct virtchnl_ether_addr))
>
> struct iavf_eth_addr_cmd {
> struct virtchnl_ether_addr_list list;
> struct virtchnl_ether_addr extra[IAVF_ETH_ADDR_PER_REQ];
> };
>
> static int
> iavf_send_uc_addr_list(struct iavf_adapter *adapter,
> struct virtchnl_ether_addr_list *list, bool add)
Passing the list object means the function *assumes* that the mac
addresses array follows right after.
Idem, the sending function now assumes the size of the passed object.
If the filling happens at the caller, then I'd rather pass the full
object and its size.
> {
> const char *opname = add ? "VIRTCHNL_OP_ADD_ETH_ADDR" :
> "VIRTCHNL_OP_DEL_ETH_ADDR";
> uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
> struct iavf_cmd_info args = {0};
> int err;
>
> args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
> args.in_args = (uint8_t *)list;
> args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
> sizeof(struct virtchnl_ether_addr) * list->num_elements;
> args.out_buffer = msg_buf;
> args.out_size = IAVF_AQ_BUF_SZ;
>
> err = iavf_execute_vf_cmd_safe(adapter, &args);
> if (err != 0)
> PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
> opname, list->num_elements);
> else
> PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
> opname, list->num_elements);
>
> return err;
> }
>
> void
> iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
> {
> struct rte_ether_addr *addrs = adapter->dev_data->mac_addrs;
> struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
> uint32_t idx = 1;
>
> /* Handle primary address (index 0) separately */
> if (!rte_is_zero_ether_addr(&addrs[0]))
> iavf_add_del_eth_addr(adapter, &addrs[0], add,
> VIRTCHNL_ETHER_ADDR_PRIMARY);
>
> /* the secondary address list is sparse, so gather it into full batches */
> while (idx < IAVF_UC_MACADDR_MAX) {
> struct iavf_eth_addr_cmd cmd = {0};
> uint16_t nb_addrs = 0;
>
> for (; idx < IAVF_UC_MACADDR_MAX && nb_addrs < IAVF_ETH_ADDR_PER_REQ;
> idx++) {
> if (rte_is_zero_ether_addr(&addrs[idx]))
> continue;
>
> memcpy(cmd.list.list[nb_addrs].addr, addrs[idx].addr_bytes,
> sizeof(cmd.list.list[nb_addrs].addr));
> cmd.list.list[nb_addrs].type = VIRTCHNL_ETHER_ADDR_EXTRA;
> nb_addrs++;
> }
>
> if (nb_addrs == 0)
> break;
>
> cmd.list.vsi_id = vf->vsi_res->vsi_id;
> cmd.list.num_elements = nb_addrs;
> if (iavf_send_uc_addr_list(adapter, &cmd.list, add) != 0)
> break;
> }
> }
> ```
Well, if we go with such a refactoring, I am not a fan of the nested
loops, but I get the idea.
I'll have a try.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses
2026-09-11 11:52 ` David Marchand
@ 2026-09-11 12:14 ` Burakov, Anatoly
0 siblings, 0 replies; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-11 12:14 UTC (permalink / raw)
To: David Marchand; +Cc: dev, bruce.richardson, Vladimir Medvedkin
On 9/11/2026 1:52 PM, David Marchand wrote:
> On Fri, 11 Sept 2026 at 11:37, Burakov, Anatoly
> <anatoly.burakov@intel.com> wrote:
>>> I would have preferred it if the caller managed the chunking, not the
>>> "add_del_addr_bulk" function. There is precedent for this style of
>>> refactor already [1], and I would like to keep things consistent - keep
>>> the loop simple (without memsets etc.), and make the caller manage how
>>> many addresses are being sent at once.
>>>
>>> [1] https://patches.dpdk.org/project/dpdk/
>>> patch/5e6a55afa2b45e3ee5ec17af7a6c548c96e9698b.1771945933.git.anatoly.burakov@intel.com/
>>>
>>> This specific refactor is more about removing rte_malloc, but it does
>>> also reorganize the loop in a way that I find to be more readable.
>>>
>>
>> I tried prototyping a loop, and realized that the fact that MAC address
>> list has holes in it is making things a little difficult, but here's
>> what I came up with as an alternative implementation, I think it's a
>> little clearer:
>>
>> ```
>> #define IAVF_ETH_ADDR_PER_REQ \
>> ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
>> sizeof(struct virtchnl_ether_addr))
>>
>> struct iavf_eth_addr_cmd {
>> struct virtchnl_ether_addr_list list;
>> struct virtchnl_ether_addr extra[IAVF_ETH_ADDR_PER_REQ];
>> };
>>
>> static int
>> iavf_send_uc_addr_list(struct iavf_adapter *adapter,
>> struct virtchnl_ether_addr_list *list, bool add)
>
> Passing the list object means the function *assumes* that the mac
> addresses array follows right after.
> Idem, the sending function now assumes the size of the passed object.
>
> If the filling happens at the caller, then I'd rather pass the full
> object and its size.
Yes, agreed, although `list` will have information about list size so
IMO just passing the full object is enough.
>
>
>> {
>> const char *opname = add ? "VIRTCHNL_OP_ADD_ETH_ADDR" :
>> "VIRTCHNL_OP_DEL_ETH_ADDR";
>> uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
>> struct iavf_cmd_info args = {0};
>> int err;
>>
>> args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
>> args.in_args = (uint8_t *)list;
>> args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
>> sizeof(struct virtchnl_ether_addr) * list->num_elements;
>> args.out_buffer = msg_buf;
>> args.out_size = IAVF_AQ_BUF_SZ;
>>
>> err = iavf_execute_vf_cmd_safe(adapter, &args);
>> if (err != 0)
>> PMD_DRV_LOG(ERR, "fail to execute command %s for %u macs",
>> opname, list->num_elements);
>> else
>> PMD_DRV_LOG(DEBUG, "executed command %s for %u macs",
>> opname, list->num_elements);
>>
>> return err;
>> }
>>
>> void
>> iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
>> {
>> struct rte_ether_addr *addrs = adapter->dev_data->mac_addrs;
>> struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
>> uint32_t idx = 1;
>>
>> /* Handle primary address (index 0) separately */
>> if (!rte_is_zero_ether_addr(&addrs[0]))
>> iavf_add_del_eth_addr(adapter, &addrs[0], add,
>> VIRTCHNL_ETHER_ADDR_PRIMARY);
>>
>> /* the secondary address list is sparse, so gather it into full batches */
>> while (idx < IAVF_UC_MACADDR_MAX) {
>> struct iavf_eth_addr_cmd cmd = {0};
>> uint16_t nb_addrs = 0;
>>
>> for (; idx < IAVF_UC_MACADDR_MAX && nb_addrs < IAVF_ETH_ADDR_PER_REQ;
>> idx++) {
>> if (rte_is_zero_ether_addr(&addrs[idx]))
>> continue;
>>
>> memcpy(cmd.list.list[nb_addrs].addr, addrs[idx].addr_bytes,
>> sizeof(cmd.list.list[nb_addrs].addr));
>> cmd.list.list[nb_addrs].type = VIRTCHNL_ETHER_ADDR_EXTRA;
>> nb_addrs++;
>> }
>>
>> if (nb_addrs == 0)
>> break;
>>
>> cmd.list.vsi_id = vf->vsi_res->vsi_id;
>> cmd.list.num_elements = nb_addrs;
>> if (iavf_send_uc_addr_list(adapter, &cmd.list, add) != 0)
>> break;
>> }
>> }
>> ```
>
> Well, if we go with such a refactoring, I am not a fan of the nested
> loops, but I get the idea.
> I'll have a try.
>
I would argue that nested loop doing compaction is idiomatic - it's
naturally a nested loop operation, so we're going to have nested loops
either way. However, the control flow is IMO much cleaner that way,
because there is no special casing inside the "send the list" function,
and additionally, such an approach lends itself to much fewer virtchnl
calls - with your code, in a degenerate "valid addr in every other
slot", you'd essentially be spamming virtchnl on every addr, while with
a nested loop like mine, you'd just compact it into a list straight away
and get away with far fewer virtchnl call-ins. So, I'd really like to
keep this kind of flow, if you don't mind :)
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install
2026-09-09 9:39 ` Loftus, Ciara
@ 2026-09-11 14:14 ` David Marchand
2026-09-11 15:37 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-11 14:14 UTC (permalink / raw)
To: Loftus, Ciara
Cc: dev@dpdk.org, Richardson, Bruce, stable@dpdk.org,
Medvedkin, Vladimir
Hello,
On Wed, 9 Sept 2026 at 11:41, Loftus, Ciara <ciara.loftus@intel.com> wrote:
>
> > Subject: [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install
> >
> > On port restart, all MAC addresses get pushed *twice* to the hardware,
> > once by the driver and once by the eth_dev_mac_restore() in ethdev.
> >
> > On the other hand, MAC address filters are reset in the hardware
> > by the PF only when a VF reset is triggered.
> >
> > Strictly speaking, the mac restore on port (re)start is unneeded,
> > if no VF reset happened, so we can announce to ethdev that no mac
> > restoration is needed via a get_restore_flags callback.
> >
> > Then, move the mac restoration to the VF reset handler.
> >
> > Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux
> > driver")
> > Cc: stable@dpdk.org
> >
> > Signed-off-by: David Marchand <david.marchand@redhat.com>
> > ---
> > Changes since v4:
> > - rebased on next-net-intel,
> >
> > Changes since v4:
> > - moved mac restoration in iavf_post_reset_reconfig,
> >
> > ---
> > drivers/net/intel/iavf/iavf_ethdev.c | 27 ++++++++++++++++-----------
> > 1 file changed, 16 insertions(+), 11 deletions(-)
> >
> > diff --git a/drivers/net/intel/iavf/iavf_ethdev.c
> > b/drivers/net/intel/iavf/iavf_ethdev.c
> > index bbd1f08ff0..bec7b3b6d7 100644
> > --- a/drivers/net/intel/iavf/iavf_ethdev.c
> > +++ b/drivers/net/intel/iavf/iavf_ethdev.c
> > @@ -292,11 +292,12 @@ iavf_get_restore_flags(__rte_unused struct
> > rte_eth_dev *dev,
> > __rte_unused enum rte_eth_dev_operation op)
> > {
> > /*
> > - * The unicast and multicast promiscuous settings persist across a
> > + * The mac addresses, unicast and multicast promiscuous settings
> > persist across a
> > * stop/start; they are only cleared by a VF reset, which the driver
> > * restores itself. So ethdev does not need to restore them on start.
> > */
> > - return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_PROMISC |
> > + return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_MAC_ADDR |
> > + RTE_ETH_RESTORE_PROMISC |
> > RTE_ETH_RESTORE_ALLMULTI);
> > }
> >
> > @@ -1095,15 +1096,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
> > rte_intr_enable(intr_handle);
> > }
> >
> > - /* Set all mac addrs */
> > - iavf_add_del_all_mac_addr(adapter, true);
> > -
> > - if (!adapter->mac_primary_set)
> > - adapter->mac_primary_set = true;
> > -
> > - /* Set all multicast addresses */
> > - iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> > >mc_addrs_num,
> > - true);
> > + if (!adapter->mac_primary_set) {
> > + if (iavf_add_del_eth_addr(adapter, &dev->data-
> > >mac_addrs[0], true,
> > + VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
> > + PMD_DRV_LOG(ERR, "failed to add primary MAC:"
> > RTE_ETHER_ADDR_PRT_FMT,
> > + RTE_ETHER_ADDR_BYTES(&dev->data-
> > >mac_addrs[0]));
> > + else
> > + adapter->mac_primary_set = true;
> > + }
> >
> > rte_spinlock_init(&vf->phc_time_aq_lock);
> >
> > @@ -3434,6 +3434,11 @@ iavf_post_reset_reconfig(struct rte_eth_dev
> > *dev)
> > int ret = 0;
> > bool allmulti = false, allunicast = false;
> > struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev-
> > >data->dev_private);
> > + struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data-
> > >dev_private);
> > +
> > + /* After a VF reset, all MAC addresses got flushed, restore them. */
> > + iavf_add_del_all_mac_addr(adapter, true);
>
> Could this lead to a double install of the primary on the reset path? In
> iavf_handle_hw_reset, dev_start can run before
> iavf_post_reset_reconfig and can install the primary MAC and set
> mac_primary_set. iavf_post_reset_reconfig then calls
> iavf_add_del_all_mac_addr which can install the primary again as it
> doesn't check mac_primary_set.
Indeed good catch, I'll fix it in next revision.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install
2026-09-11 14:14 ` David Marchand
@ 2026-09-11 15:37 ` David Marchand
2026-09-11 16:26 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-11 15:37 UTC (permalink / raw)
To: Loftus, Ciara, Richardson, Bruce
Cc: dev@dpdk.org, stable@dpdk.org, Medvedkin, Vladimir,
Burakov, Anatoly
On Fri, 11 Sept 2026 at 16:14, David Marchand <david.marchand@redhat.com> wrote:
>
> Hello,
>
> On Wed, 9 Sept 2026 at 11:41, Loftus, Ciara <ciara.loftus@intel.com> wrote:
> >
> > > Subject: [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install
> > >
> > > On port restart, all MAC addresses get pushed *twice* to the hardware,
> > > once by the driver and once by the eth_dev_mac_restore() in ethdev.
> > >
> > > On the other hand, MAC address filters are reset in the hardware
> > > by the PF only when a VF reset is triggered.
> > >
> > > Strictly speaking, the mac restore on port (re)start is unneeded,
> > > if no VF reset happened, so we can announce to ethdev that no mac
> > > restoration is needed via a get_restore_flags callback.
> > >
> > > Then, move the mac restoration to the VF reset handler.
> > >
> > > Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux
> > > driver")
> > > Cc: stable@dpdk.org
> > >
> > > Signed-off-by: David Marchand <david.marchand@redhat.com>
> > > ---
> > > Changes since v4:
> > > - rebased on next-net-intel,
> > >
> > > Changes since v4:
> > > - moved mac restoration in iavf_post_reset_reconfig,
> > >
> > > ---
> > > drivers/net/intel/iavf/iavf_ethdev.c | 27 ++++++++++++++++-----------
> > > 1 file changed, 16 insertions(+), 11 deletions(-)
> > >
> > > diff --git a/drivers/net/intel/iavf/iavf_ethdev.c
> > > b/drivers/net/intel/iavf/iavf_ethdev.c
> > > index bbd1f08ff0..bec7b3b6d7 100644
> > > --- a/drivers/net/intel/iavf/iavf_ethdev.c
> > > +++ b/drivers/net/intel/iavf/iavf_ethdev.c
> > > @@ -292,11 +292,12 @@ iavf_get_restore_flags(__rte_unused struct
> > > rte_eth_dev *dev,
> > > __rte_unused enum rte_eth_dev_operation op)
> > > {
> > > /*
> > > - * The unicast and multicast promiscuous settings persist across a
> > > + * The mac addresses, unicast and multicast promiscuous settings
> > > persist across a
> > > * stop/start; they are only cleared by a VF reset, which the driver
> > > * restores itself. So ethdev does not need to restore them on start.
> > > */
> > > - return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_PROMISC |
> > > + return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_MAC_ADDR |
> > > + RTE_ETH_RESTORE_PROMISC |
> > > RTE_ETH_RESTORE_ALLMULTI);
> > > }
> > >
> > > @@ -1095,15 +1096,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
> > > rte_intr_enable(intr_handle);
> > > }
> > >
> > > - /* Set all mac addrs */
> > > - iavf_add_del_all_mac_addr(adapter, true);
> > > -
> > > - if (!adapter->mac_primary_set)
> > > - adapter->mac_primary_set = true;
> > > -
> > > - /* Set all multicast addresses */
> > > - iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> > > >mc_addrs_num,
> > > - true);
> > > + if (!adapter->mac_primary_set) {
> > > + if (iavf_add_del_eth_addr(adapter, &dev->data-
> > > >mac_addrs[0], true,
> > > + VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
> > > + PMD_DRV_LOG(ERR, "failed to add primary MAC:"
> > > RTE_ETHER_ADDR_PRT_FMT,
> > > + RTE_ETHER_ADDR_BYTES(&dev->data-
> > > >mac_addrs[0]));
> > > + else
> > > + adapter->mac_primary_set = true;
> > > + }
> > >
> > > rte_spinlock_init(&vf->phc_time_aq_lock);
> > >
> > > @@ -3434,6 +3434,11 @@ iavf_post_reset_reconfig(struct rte_eth_dev
> > > *dev)
> > > int ret = 0;
> > > bool allmulti = false, allunicast = false;
> > > struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev-
> > > >data->dev_private);
> > > + struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data-
> > > >dev_private);
> > > +
> > > + /* After a VF reset, all MAC addresses got flushed, restore them. */
> > > + iavf_add_del_all_mac_addr(adapter, true);
> >
> > Could this lead to a double install of the primary on the reset path? In
> > iavf_handle_hw_reset, dev_start can run before
> > iavf_post_reset_reconfig and can install the primary MAC and set
> > mac_primary_set. iavf_post_reset_reconfig then calls
> > iavf_add_del_all_mac_addr which can install the primary again as it
> > doesn't check mac_primary_set.
>
> Indeed good catch, I'll fix it in next revision.
Interesting..
It's been so long I started touching this code, I am not sure at what
I tested now...
It seems I found a leak of the mac address array on VF reset.
iavf_dev_uninit + iavf_dev_init results in overwriting
dev->data->mac_addrs right?
This should break the mac restoration during a VF reset, regardless of
who does it, the driver or ethdev.
The multicast addresses look unaffected, as those are stored in the
iavf_info struct.
I could copy the mac address and reinsert the mac addresses during
iavf_post_reset_reconfig...
But I think it is simpler to just check eth->data->mac_addrs state in
iavf_dev_init.
Opinions?
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install
2026-09-11 15:37 ` David Marchand
@ 2026-09-11 16:26 ` David Marchand
0 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-11 16:26 UTC (permalink / raw)
To: Loftus, Ciara, Richardson, Bruce
Cc: dev@dpdk.org, stable@dpdk.org, Medvedkin, Vladimir,
Burakov, Anatoly
On Fri, 11 Sept 2026 at 17:37, David Marchand <david.marchand@redhat.com> wrote:
> Interesting..
>
> It's been so long I started touching this code, I am not sure at what
> I tested now...
>
> It seems I found a leak of the mac address array on VF reset.
> iavf_dev_uninit + iavf_dev_init results in overwriting
> dev->data->mac_addrs right?
>
> This should break the mac restoration during a VF reset, regardless of
> who does it, the driver or ethdev.
>
> The multicast addresses look unaffected, as those are stored in the
> iavf_info struct.
>
> I could copy the mac address and reinsert the mac addresses during
> iavf_post_reset_reconfig...
> But I think it is simpler to just check eth->data->mac_addrs state in
> iavf_dev_init.
I think I'll just embed this array in iavf_info, which already has the
multicast mac addresses.
That will just work.
I just need to carefully sanitize dev->data->mac_addrs before
releasing the port.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (11 preceding siblings ...)
2026-09-08 9:27 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
@ 2026-09-14 8:17 ` David Marchand
2026-09-14 8:17 ` [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install David Marchand
` (4 more replies)
2026-09-14 14:42 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (2 subsequent siblings)
15 siblings, 5 replies; 146+ messages in thread
From: David Marchand @ 2026-09-14 8:17 UTC (permalink / raw)
To: dev
Cc: ciara.loftus, anatoly.burakov, rjarry, cfontain, stable,
Vladimir Medvedkin, Lunyuan Cui, Jingjing Wu, Xiaolong Ye
When resetting, calling iavf_dev_uninit + iavf_dev_init results in
leaking the previous dev->data->mac_addrs array.
As a consequence, secondary MAC addresses are lost during a VF reset.
Move the MAC addresses array in the private iavf_info structure, along
the multicast MAC addresses.
Set/clear dev->data->mac_addrs in iavf_dev_init/iavf_dev_uninit.
Consistently use RTE_DIM() to avoid mixing with the multicast addresses
array.
Fixes: e74e1bb6280d ("net/iavf: enable port reset")
Cc: stable@dpdk.org
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/intel/iavf/iavf.h | 3 +++
drivers/net/intel/iavf/iavf_ethdev.c | 20 +++++++-------------
drivers/net/intel/iavf/iavf_vchnl.c | 10 +++++-----
3 files changed, 15 insertions(+), 18 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 044de3cfd4..1d52b3d113 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -255,6 +255,9 @@ struct iavf_info {
bool link_up;
uint32_t link_speed;
+ /* Unicast addrs */
+ struct rte_ether_addr mac_addrs[IAVF_NUM_MACADDR_MAX];
+
/* Multicast addrs */
struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
uint16_t mc_addrs_num; /* Multicast mac addresses number */
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index 2ac4dbdac4..b498258d58 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -1186,7 +1186,7 @@ iavf_dev_info_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *dev_info)
dev_info->hash_key_size = vf->vf_res->rss_key_size;
dev_info->reta_size = vf->vf_res->rss_lut_size;
dev_info->flow_type_rss_offloads = IAVF_RSS_OFFLOAD_ALL;
- dev_info->max_mac_addrs = IAVF_NUM_MACADDR_MAX;
+ dev_info->max_mac_addrs = RTE_DIM(vf->mac_addrs);
/*
* Runtime queue setup can race with the hardware Tx rate limiter on
* E810 VFs and corrupt queue state. Once a per-queue bandwidth rte_tm
@@ -3104,16 +3104,9 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
/* set default ptype table */
iavf_set_default_ptype_table(eth_dev);
- /* copy mac addr */
- eth_dev->data->mac_addrs = rte_zmalloc(
- "iavf_mac", RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX, 0);
- if (!eth_dev->data->mac_addrs) {
- PMD_INIT_LOG(ERR, "Failed to allocate %d bytes needed to"
- " store MAC addresses",
- RTE_ETHER_ADDR_LEN * IAVF_NUM_MACADDR_MAX);
- ret = -ENOMEM;
- goto init_vf_err;
- }
+ /* Point at the MAC addresses array from priv */
+ eth_dev->data->mac_addrs = vf->mac_addrs;
+
/* If the MAC address is not configured by host,
* generate a random one.
*/
@@ -3123,7 +3116,6 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
rte_ether_addr_copy((struct rte_ether_addr *)hw->mac.addr,
ð_dev->data->mac_addrs[0]);
-
if (vf->vf_res->vf_cap_flags & VIRTCHNL_VF_OFFLOAD_WB_ON_ITR &&
/* register callback func to eal lib */
rte_intr_callback_register(pci_dev->intr_handle,
@@ -3199,7 +3191,6 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
iavf_phc_sync_alarm_stop(eth_dev);
rte_eal_alarm_cancel(iavf_dev_alarm_handler, eth_dev);
- rte_free(eth_dev->data->mac_addrs);
eth_dev->data->mac_addrs = NULL;
init_vf_err:
@@ -3256,6 +3247,9 @@ iavf_dev_close(struct rte_eth_dev *dev)
adapter->closed = true;
+ /* Clear reference to the MAC addresses in priv */
+ dev->data->mac_addrs = NULL;
+
/* free iAVF security device context all related resources */
iavf_security_ctx_destroy(adapter);
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index e3120c655d..df97ff2052 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1674,19 +1674,19 @@ iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num)
void
iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
{
+ struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
struct {
struct virtchnl_ether_addr_list list;
- struct virtchnl_ether_addr addr[IAVF_NUM_MACADDR_MAX];
+ struct virtchnl_ether_addr addr[RTE_DIM(vf->mac_addrs)];
} list_req = {0};
struct virtchnl_ether_addr_list *list = &list_req.list;
- struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
struct iavf_cmd_info args = {0};
- int err, i;
+ int err;
size_t buf_len;
- for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
- struct rte_ether_addr *addr = &adapter->dev_data->mac_addrs[i];
+ for (unsigned int i = 0; i < RTE_DIM(vf->mac_addrs); i++) {
+ struct rte_ether_addr *addr = &vf->mac_addrs[i];
struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
/* ignore empty addresses */
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install
2026-09-14 8:17 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset David Marchand
@ 2026-09-14 8:17 ` David Marchand
2026-09-14 10:19 ` Loftus, Ciara
2026-09-23 11:53 ` Burakov, Anatoly
2026-09-14 8:17 ` [PATCH v7 3/4] net/iavf: add a helper for sending MAC addresses to PF David Marchand
` (3 subsequent siblings)
4 siblings, 2 replies; 146+ messages in thread
From: David Marchand @ 2026-09-14 8:17 UTC (permalink / raw)
To: dev
Cc: ciara.loftus, anatoly.burakov, rjarry, cfontain, stable,
Vladimir Medvedkin, Bruce Richardson
On port restart, all MAC addresses get pushed *twice* to the hardware,
once by the driver and once by the eth_dev_mac_restore() in ethdev.
On the other hand, MAC address filters are reset in the hardware
by the PF only when a VF reset is triggered.
Strictly speaking, the mac restore on port (re)start is unneeded,
if no VF reset happened, so we can announce to ethdev that no mac
restoration is needed via a get_restore_flags callback.
Then, move the mac restoration to the VF reset handler.
Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux driver")
Cc: stable@dpdk.org
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v6:
- renamed iavf_add_del_all_mac_addr and removed primary mac address
handling out of this helper (avoids double primary mac installation),
Changes since v4:
- rebased on next-net-intel,
Changes since v4:
- moved mac restoration in iavf_post_reset_reconfig,
---
drivers/net/intel/iavf/iavf.h | 2 +-
drivers/net/intel/iavf/iavf_ethdev.c | 30 ++++++++++++++++++----------
drivers/net/intel/iavf/iavf_vchnl.c | 11 +++++-----
3 files changed, 26 insertions(+), 17 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 1d52b3d113..34b9b4ad94 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -482,7 +482,7 @@ int iavf_add_del_vlan_v2(struct iavf_adapter *adapter, uint16_t vlanid,
int iavf_get_vlan_offload_caps_v2(struct iavf_adapter *adapter);
int iavf_config_irq_map(struct iavf_adapter *adapter);
int iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num);
-void iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add);
+void iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool add);
int iavf_dev_link_update(struct rte_eth_dev *dev,
__rte_unused int wait_to_complete);
void iavf_dev_alarm_handler(void *param);
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index b498258d58..f7aeac8c83 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -292,11 +292,12 @@ iavf_get_restore_flags(__rte_unused struct rte_eth_dev *dev,
__rte_unused enum rte_eth_dev_operation op)
{
/*
- * The unicast and multicast promiscuous settings persist across a
+ * The mac addresses, unicast and multicast promiscuous settings persist across a
* stop/start; they are only cleared by a VF reset, which the driver
* restores itself. So ethdev does not need to restore them on start.
*/
- return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_PROMISC |
+ return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_MAC_ADDR |
+ RTE_ETH_RESTORE_PROMISC |
RTE_ETH_RESTORE_ALLMULTI);
}
@@ -1095,15 +1096,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
rte_intr_enable(intr_handle);
}
- /* Set all mac addrs */
- iavf_add_del_all_mac_addr(adapter, true);
-
- if (!adapter->mac_primary_set)
- adapter->mac_primary_set = true;
-
- /* Set all multicast addresses */
- iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num,
- true);
+ if (!adapter->mac_primary_set) {
+ if (iavf_add_del_eth_addr(adapter, &dev->data->mac_addrs[0], true,
+ VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
+ PMD_DRV_LOG(ERR, "failed to add primary MAC:" RTE_ETHER_ADDR_PRT_FMT,
+ RTE_ETHER_ADDR_BYTES(&dev->data->mac_addrs[0]));
+ else
+ adapter->mac_primary_set = true;
+ }
rte_spinlock_init(&vf->phc_time_aq_lock);
@@ -3428,6 +3428,14 @@ iavf_post_reset_reconfig(struct rte_eth_dev *dev)
int ret = 0;
bool allmulti = false, allunicast = false;
struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
+ struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data->dev_private);
+
+ /*
+ * After a VF reset, all MAC addresses got flushed.
+ * The primary MAC should have been set in iavf_dev_start, restore the rest.
+ */
+ iavf_add_del_secondary_mac_addr(adapter, true);
+ (void)iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf->mc_addrs_num, true);
/* Restore pre-reset unicast promiscuous and multicast promiscuous states */
if (dev->data->promiscuous)
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index df97ff2052..a918db5443 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1672,7 +1672,7 @@ iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num)
}
void
-iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
+iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool add)
{
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
struct {
@@ -1685,7 +1685,7 @@ iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
int err;
size_t buf_len;
- for (unsigned int i = 0; i < RTE_DIM(vf->mac_addrs); i++) {
+ for (unsigned int i = 1; i < RTE_DIM(vf->mac_addrs); i++) {
struct rte_ether_addr *addr = &vf->mac_addrs[i];
struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
@@ -1695,11 +1695,12 @@ iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
list->num_elements++;
memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
- vc_addr->type = (list->num_elements == 1) ?
- VIRTCHNL_ETHER_ADDR_PRIMARY :
- VIRTCHNL_ETHER_ADDR_EXTRA;
+ vc_addr->type = VIRTCHNL_ETHER_ADDR_EXTRA;
}
+ if (list->num_elements == 0)
+ return;
+
/* for some reason PF side checks for buffer being too big, so adjust it down */
buf_len = sizeof(struct virtchnl_ether_addr_list) +
sizeof(struct virtchnl_ether_addr) * list->num_elements;
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v7 3/4] net/iavf: add a helper for sending MAC addresses to PF
2026-09-14 8:17 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset David Marchand
2026-09-14 8:17 ` [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install David Marchand
@ 2026-09-14 8:17 ` David Marchand
2026-09-23 12:04 ` Burakov, Anatoly
2026-09-14 8:17 ` [PATCH v7 4/4] net/iavf: accept up to 32k unicast MAC addresses David Marchand
` (2 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-14 8:17 UTC (permalink / raw)
To: dev; +Cc: ciara.loftus, anatoly.burakov, rjarry, cfontain,
Vladimir Medvedkin
Rather than have multiple implementations of the same code,
define a single helper.
This is also a good place where to check if the driver is requesting too
many addresses in a single message to the PF.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
drivers/net/intel/iavf/iavf_vchnl.c | 124 +++++++++++++---------------
1 file changed, 56 insertions(+), 68 deletions(-)
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index a918db5443..decfae3182 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1671,50 +1671,69 @@ iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num)
return 0;
}
+#define IAVF_ETH_ADDR_PER_REQ \
+ ((IAVF_AQ_BUF_SZ - sizeof(struct virtchnl_ether_addr_list)) / \
+ sizeof(struct virtchnl_ether_addr))
+#define IAVF_ETH_ADDR_CMD_SIZE(nb_addrs) \
+ (sizeof(struct virtchnl_ether_addr_list) + sizeof(struct virtchnl_ether_addr) * (nb_addrs))
+
+static int
+iavf_send_eth_addr_list(struct iavf_adapter *adapter, const char *caller,
+ struct virtchnl_ether_addr_list *list, bool add)
+{
+ uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
+ struct iavf_cmd_info args;
+ int err;
+
+ if (list->num_elements > IAVF_ETH_ADDR_PER_REQ) {
+ PMD_DRV_LOG(ERR, "cannot fit this ethernet address list in a message to the PF");
+ return -EINVAL;
+ }
+
+ memset(&args, 0, sizeof(args));
+ args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
+ args.in_args = (uint8_t *)list;
+ args.in_args_size = IAVF_ETH_ADDR_CMD_SIZE(list->num_elements);
+ args.out_buffer = msg_buf;
+ args.out_size = IAVF_AQ_BUF_SZ;
+ err = iavf_execute_vf_cmd_safe(adapter, &args);
+ if (err != 0) {
+ PMD_DRV_LOG(ERR, "fail to execute command %s for %s",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR", caller);
+ return err;
+ }
+
+ PMD_DRV_LOG(DEBUG, "executed command %s for %s",
+ add ? "VIRTCHNL_OP_ADD_ETH_ADDR" : "VIRTCHNL_OP_DEL_ETH_ADDR", caller);
+
+ return 0;
+}
+
void
iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool add)
{
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
- struct {
- struct virtchnl_ether_addr_list list;
- struct virtchnl_ether_addr addr[RTE_DIM(vf->mac_addrs)];
- } list_req = {0};
- struct virtchnl_ether_addr_list *list = &list_req.list;
- uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
- struct iavf_cmd_info args = {0};
- int err;
- size_t buf_len;
+ uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(RTE_DIM(vf->mac_addrs))] = {0};
+ struct virtchnl_ether_addr_list *list;
+
+ list = (struct virtchnl_ether_addr_list *)cmd_buffer;
+ list->vsi_id = vf->vsi_res->vsi_id;
+ list->num_elements = 0;
for (unsigned int i = 1; i < RTE_DIM(vf->mac_addrs); i++) {
struct rte_ether_addr *addr = &vf->mac_addrs[i];
struct virtchnl_ether_addr *vc_addr = &list->list[list->num_elements];
- /* ignore empty addresses */
- if (rte_is_zero_ether_addr(addr))
- continue;
- list->num_elements++;
+ if (!rte_is_zero_ether_addr(addr)) {
+ list->num_elements++;
- memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
- vc_addr->type = VIRTCHNL_ETHER_ADDR_EXTRA;
+ memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
+ vc_addr->type = VIRTCHNL_ETHER_ADDR_EXTRA;
+ }
}
- if (list->num_elements == 0)
- return;
-
- /* for some reason PF side checks for buffer being too big, so adjust it down */
- buf_len = sizeof(struct virtchnl_ether_addr_list) +
- sizeof(struct virtchnl_ether_addr) * list->num_elements;
-
- list->vsi_id = vf->vsi_res->vsi_id;
- args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
- args.in_args = (uint8_t *)list;
- args.in_args_size = buf_len;
- args.out_buffer = msg_buf;
- args.out_size = IAVF_AQ_BUF_SZ;
- err = iavf_execute_vf_cmd_safe(adapter, &args);
- if (err)
- PMD_DRV_LOG(ERR, "fail to execute command %s",
- add ? "OP_ADD_ETHER_ADDRESS" : "OP_DEL_ETHER_ADDRESS");
+ if (list->num_elements != 0)
+ (void)iavf_send_eth_addr_list(adapter, __func__, list, add);
}
int
@@ -1797,13 +1816,9 @@ int
iavf_add_del_eth_addr(struct iavf_adapter *adapter, struct rte_ether_addr *addr,
bool add, uint8_t type)
{
- struct virtchnl_ether_addr_list *list;
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
- uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
- uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
- sizeof(struct virtchnl_ether_addr)];
- struct iavf_cmd_info args;
- int err;
+ uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(1)] = {0};
+ struct virtchnl_ether_addr_list *list;
if (adapter->closed)
return -EIO;
@@ -1815,16 +1830,7 @@ iavf_add_del_eth_addr(struct iavf_adapter *adapter, struct rte_ether_addr *addr,
memcpy(list->list[0].addr, addr->addr_bytes,
sizeof(addr->addr_bytes));
- args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
- args.in_args = cmd_buffer;
- args.in_args_size = sizeof(cmd_buffer);
- args.out_buffer = msg_buf;
- args.out_size = IAVF_AQ_BUF_SZ;
- err = iavf_execute_vf_cmd_safe(adapter, &args);
- if (err)
- PMD_DRV_LOG(ERR, "fail to execute command %s",
- add ? "OP_ADD_ETH_ADDR" : "OP_DEL_ETH_ADDR");
- return err;
+ return iavf_send_eth_addr_list(adapter, __func__, list, add);
}
int
@@ -2302,14 +2308,10 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
struct rte_ether_addr *mc_addrs,
uint32_t mc_addrs_num, bool add)
{
+ uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(IAVF_NUM_MACADDR_MAX)] = {0};
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
- uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
- uint8_t cmd_buffer[sizeof(struct virtchnl_ether_addr_list) +
- (IAVF_NUM_MACADDR_MAX * sizeof(struct virtchnl_ether_addr))];
struct virtchnl_ether_addr_list *list;
- struct iavf_cmd_info args;
uint32_t i;
- int err;
if (mc_addrs == NULL || mc_addrs_num == 0)
return 0;
@@ -2330,21 +2332,7 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
list->list[i].type = VIRTCHNL_ETHER_ADDR_EXTRA;
}
- args.ops = add ? VIRTCHNL_OP_ADD_ETH_ADDR : VIRTCHNL_OP_DEL_ETH_ADDR;
- args.in_args = cmd_buffer;
- args.in_args_size = sizeof(struct virtchnl_ether_addr_list) +
- i * sizeof(struct virtchnl_ether_addr);
- args.out_buffer = msg_buf;
- args.out_size = IAVF_AQ_BUF_SZ;
- err = iavf_execute_vf_cmd_safe(adapter, &args);
-
- if (err) {
- PMD_DRV_LOG(ERR, "fail to execute command %s",
- add ? "OP_ADD_ETH_ADDR" : "OP_DEL_ETH_ADDR");
- return err;
- }
-
- return 0;
+ return iavf_send_eth_addr_list(adapter, __func__, list, add);
}
int
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v7 4/4] net/iavf: accept up to 32k unicast MAC addresses
2026-09-14 8:17 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset David Marchand
2026-09-14 8:17 ` [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install David Marchand
2026-09-14 8:17 ` [PATCH v7 3/4] net/iavf: add a helper for sending MAC addresses to PF David Marchand
@ 2026-09-14 8:17 ` David Marchand
2026-09-23 12:14 ` Burakov, Anatoly
2026-09-14 10:15 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset Loftus, Ciara
2026-09-23 11:47 ` Burakov, Anatoly
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-14 8:17 UTC (permalink / raw)
To: dev; +Cc: ciara.loftus, anatoly.burakov, rjarry, cfontain,
Vladimir Medvedkin
E810 hardware provides 32k switch lookups.
Thanks to this, it is possible to allow a lot more secondary mac
addresses than what is possible today.
In practice, the maximum number of macs available per port may be lower
and depends on usage by other (trusted?) VFs on the same PF.
There is no way to figure out this limit but to try adding a mac address
and get an error from the PF driver.
Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
accordingly.
Since unicast and multicast addresses arrays are sized with two
different constants, prefer RTE_DIM() whenever possible.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v6:
- reused helper added in previous commit,
- used RTE_DIM() instead of macro constants,
Changes since v5:
- separated from series that went in next-net,
- rebased,
Changes since v4:
- rebased,
Changes since v2:
- added an entry in release notes,
- removed unneeded temp variable,
Changes since v1:
- fixed buffer overflow on mailbox messages during port restart/VF reset,
---
doc/guides/rel_notes/release_26_11.rst | 4 ++++
drivers/net/intel/iavf/iavf.h | 7 ++++---
drivers/net/intel/iavf/iavf_ethdev.c | 7 +++----
drivers/net/intel/iavf/iavf_vchnl.c | 10 ++++++++--
4 files changed, 19 insertions(+), 9 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 6c29513041..5a4c7d815d 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -64,6 +64,10 @@ New Features
* Renamed the ``enable_ptype_lldp`` devarg to ``enable_lldp``.
The old name is no longer accepted.
+ * Increased the maximum number of secondary unicast MAC addresses
+ from 64 to 32k.
+ This increases a VF port memory footprint by ~192kB.
+
* **Updated Intel ixgbe driver.**
Added ``fdir_buffer_size`` devarg to select the Flow Director table size
diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
index 34b9b4ad94..f3f2b4ae51 100644
--- a/drivers/net/intel/iavf/iavf.h
+++ b/drivers/net/intel/iavf/iavf.h
@@ -32,7 +32,8 @@
#define IAVF_IRQ_MAP_NUM_PER_BUF 128
#define IAVF_RXTX_QUEUE_CHUNKS_NUM 2
-#define IAVF_NUM_MACADDR_MAX 64
+#define IAVF_UC_MACADDR_MAX 32768
+#define IAVF_MC_MACADDR_MAX 64
#define IAVF_DEV_WATCHDOG_PERIOD 2000 /* microseconds, set 0 to disable*/
@@ -256,10 +257,10 @@ struct iavf_info {
uint32_t link_speed;
/* Unicast addrs */
- struct rte_ether_addr mac_addrs[IAVF_NUM_MACADDR_MAX];
+ struct rte_ether_addr mac_addrs[IAVF_UC_MACADDR_MAX];
/* Multicast addrs */
- struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
+ struct rte_ether_addr mc_addrs[IAVF_MC_MACADDR_MAX];
uint16_t mc_addrs_num; /* Multicast mac addresses number */
struct iavf_vsi vsi;
diff --git a/drivers/net/intel/iavf/iavf_ethdev.c b/drivers/net/intel/iavf/iavf_ethdev.c
index f7aeac8c83..2d055e6f5b 100644
--- a/drivers/net/intel/iavf/iavf_ethdev.c
+++ b/drivers/net/intel/iavf/iavf_ethdev.c
@@ -411,10 +411,9 @@ iavf_set_mc_addr_list(struct rte_eth_dev *dev,
IAVF_DEV_PRIVATE_TO_ADAPTER(dev->data->dev_private);
int err, ret;
- if (mc_addrs_num > IAVF_NUM_MACADDR_MAX) {
- PMD_DRV_LOG(ERR,
- "can't add more than a limited number (%u) of addresses.",
- (uint32_t)IAVF_NUM_MACADDR_MAX);
+ if (mc_addrs_num > RTE_DIM(vf->mc_addrs)) {
+ PMD_DRV_LOG(ERR, "can't add more than a limited number (%u) of addresses.",
+ (unsigned int)RTE_DIM(vf->mc_addrs));
return -EINVAL;
}
diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
index decfae3182..418a7e897e 100644
--- a/drivers/net/intel/iavf/iavf_vchnl.c
+++ b/drivers/net/intel/iavf/iavf_vchnl.c
@@ -1712,8 +1712,8 @@ iavf_send_eth_addr_list(struct iavf_adapter *adapter, const char *caller,
void
iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool add)
{
+ uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(IAVF_ETH_ADDR_PER_REQ)] = {0};
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
- uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(RTE_DIM(vf->mac_addrs))] = {0};
struct virtchnl_ether_addr_list *list;
list = (struct virtchnl_ether_addr_list *)cmd_buffer;
@@ -1730,6 +1730,12 @@ iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool add)
memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
vc_addr->type = VIRTCHNL_ETHER_ADDR_EXTRA;
}
+
+ if (list->num_elements == IAVF_ETH_ADDR_PER_REQ) {
+ if (iavf_send_eth_addr_list(adapter, __func__, list, add))
+ return;
+ list->num_elements = 0;
+ }
}
if (list->num_elements != 0)
@@ -2308,8 +2314,8 @@ iavf_add_del_mc_addr_list(struct iavf_adapter *adapter,
struct rte_ether_addr *mc_addrs,
uint32_t mc_addrs_num, bool add)
{
- uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(IAVF_NUM_MACADDR_MAX)] = {0};
struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
+ uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(RTE_DIM(vf->mc_addrs))] = {0};
struct virtchnl_ether_addr_list *list;
uint32_t i;
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* RE: [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset
2026-09-14 8:17 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset David Marchand
` (2 preceding siblings ...)
2026-09-14 8:17 ` [PATCH v7 4/4] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-09-14 10:15 ` Loftus, Ciara
2026-09-14 11:54 ` David Marchand
2026-09-23 11:47 ` Burakov, Anatoly
4 siblings, 1 reply; 146+ messages in thread
From: Loftus, Ciara @ 2026-09-14 10:15 UTC (permalink / raw)
To: David Marchand, dev@dpdk.org
Cc: Burakov, Anatoly, rjarry@redhat.com, cfontain@redhat.com,
stable@dpdk.org, Medvedkin, Vladimir, Lunyuan Cui, Wu, Jingjing,
Xiaolong Ye
> Subject: [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset
>
> When resetting, calling iavf_dev_uninit + iavf_dev_init results in
> leaking the previous dev->data->mac_addrs array.
> As a consequence, secondary MAC addresses are lost during a VF reset.
>
> Move the MAC addresses array in the private iavf_info structure, along
> the multicast MAC addresses.
> Set/clear dev->data->mac_addrs in iavf_dev_init/iavf_dev_uninit.
>
> Consistently use RTE_DIM() to avoid mixing with the multicast addresses
> array.
>
> Fixes: e74e1bb6280d ("net/iavf: enable port reset")
> Cc: stable@dpdk.org
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
> drivers/net/intel/iavf/iavf.h | 3 +++
> drivers/net/intel/iavf/iavf_ethdev.c | 20 +++++++-------------
> drivers/net/intel/iavf/iavf_vchnl.c | 10 +++++-----
> 3 files changed, 15 insertions(+), 18 deletions(-)
>
> diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
> index 044de3cfd4..1d52b3d113 100644
> --- a/drivers/net/intel/iavf/iavf.h
> +++ b/drivers/net/intel/iavf/iavf.h
> @@ -255,6 +255,9 @@ struct iavf_info {
> bool link_up;
> uint32_t link_speed;
>
> + /* Unicast addrs */
> + struct rte_ether_addr mac_addrs[IAVF_NUM_MACADDR_MAX];
> +
> /* Multicast addrs */
> struct rte_ether_addr mc_addrs[IAVF_NUM_MACADDR_MAX];
> uint16_t mc_addrs_num; /* Multicast mac addresses number */
> diff --git a/drivers/net/intel/iavf/iavf_ethdev.c
> b/drivers/net/intel/iavf/iavf_ethdev.c
> index 2ac4dbdac4..b498258d58 100644
> --- a/drivers/net/intel/iavf/iavf_ethdev.c
> +++ b/drivers/net/intel/iavf/iavf_ethdev.c
> @@ -1186,7 +1186,7 @@ iavf_dev_info_get(struct rte_eth_dev *dev, struct
> rte_eth_dev_info *dev_info)
> dev_info->hash_key_size = vf->vf_res->rss_key_size;
> dev_info->reta_size = vf->vf_res->rss_lut_size;
> dev_info->flow_type_rss_offloads = IAVF_RSS_OFFLOAD_ALL;
> - dev_info->max_mac_addrs = IAVF_NUM_MACADDR_MAX;
> + dev_info->max_mac_addrs = RTE_DIM(vf->mac_addrs);
> /*
> * Runtime queue setup can race with the hardware Tx rate limiter on
> * E810 VFs and corrupt queue state. Once a per-queue bandwidth
> rte_tm
> @@ -3104,16 +3104,9 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
> /* set default ptype table */
> iavf_set_default_ptype_table(eth_dev);
>
> - /* copy mac addr */
> - eth_dev->data->mac_addrs = rte_zmalloc(
> - "iavf_mac", RTE_ETHER_ADDR_LEN *
> IAVF_NUM_MACADDR_MAX, 0);
> - if (!eth_dev->data->mac_addrs) {
> - PMD_INIT_LOG(ERR, "Failed to allocate %d bytes needed to"
> - " store MAC addresses",
> - RTE_ETHER_ADDR_LEN *
> IAVF_NUM_MACADDR_MAX);
> - ret = -ENOMEM;
> - goto init_vf_err;
> - }
> + /* Point at the MAC addresses array from priv */
> + eth_dev->data->mac_addrs = vf->mac_addrs;
> +
> /* If the MAC address is not configured by host,
> * generate a random one.
> */
> @@ -3123,7 +3116,6 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
> rte_ether_addr_copy((struct rte_ether_addr *)hw->mac.addr,
> ð_dev->data->mac_addrs[0]);
>
> -
Nit: stray line removal. The rest LGTM.
Acked-by: Ciara Loftus <ciara.loftus@intel.com>
> if (vf->vf_res->vf_cap_flags & VIRTCHNL_VF_OFFLOAD_WB_ON_ITR
> &&
> /* register callback func to eal lib */
> rte_intr_callback_register(pci_dev->intr_handle,
> @@ -3199,7 +3191,6 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
> iavf_phc_sync_alarm_stop(eth_dev);
> rte_eal_alarm_cancel(iavf_dev_alarm_handler, eth_dev);
>
> - rte_free(eth_dev->data->mac_addrs);
> eth_dev->data->mac_addrs = NULL;
>
> init_vf_err:
> @@ -3256,6 +3247,9 @@ iavf_dev_close(struct rte_eth_dev *dev)
>
> adapter->closed = true;
>
> + /* Clear reference to the MAC addresses in priv */
> + dev->data->mac_addrs = NULL;
> +
> /* free iAVF security device context all related resources */
> iavf_security_ctx_destroy(adapter);
>
> diff --git a/drivers/net/intel/iavf/iavf_vchnl.c
> b/drivers/net/intel/iavf/iavf_vchnl.c
> index e3120c655d..df97ff2052 100644
> --- a/drivers/net/intel/iavf/iavf_vchnl.c
> +++ b/drivers/net/intel/iavf/iavf_vchnl.c
> @@ -1674,19 +1674,19 @@ iavf_config_irq_map_lv(struct iavf_adapter
> *adapter, uint16_t num)
> void
> iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
> {
> + struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
> struct {
> struct virtchnl_ether_addr_list list;
> - struct virtchnl_ether_addr
> addr[IAVF_NUM_MACADDR_MAX];
> + struct virtchnl_ether_addr addr[RTE_DIM(vf->mac_addrs)];
> } list_req = {0};
> struct virtchnl_ether_addr_list *list = &list_req.list;
> - struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
> uint8_t msg_buf[IAVF_AQ_BUF_SZ] = {0};
> struct iavf_cmd_info args = {0};
> - int err, i;
> + int err;
> size_t buf_len;
>
> - for (i = 0; i < IAVF_NUM_MACADDR_MAX; i++) {
> - struct rte_ether_addr *addr = &adapter->dev_data-
> >mac_addrs[i];
> + for (unsigned int i = 0; i < RTE_DIM(vf->mac_addrs); i++) {
> + struct rte_ether_addr *addr = &vf->mac_addrs[i];
> struct virtchnl_ether_addr *vc_addr = &list->list[list-
> >num_elements];
>
> /* ignore empty addresses */
> --
> 2.54.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* RE: [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install
2026-09-14 8:17 ` [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install David Marchand
@ 2026-09-14 10:19 ` Loftus, Ciara
2026-09-14 11:56 ` David Marchand
2026-09-23 11:53 ` Burakov, Anatoly
1 sibling, 1 reply; 146+ messages in thread
From: Loftus, Ciara @ 2026-09-14 10:19 UTC (permalink / raw)
To: David Marchand, dev@dpdk.org
Cc: Burakov, Anatoly, rjarry@redhat.com, cfontain@redhat.com,
stable@dpdk.org, Medvedkin, Vladimir, Richardson, Bruce
> Subject: [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install
>
> On port restart, all MAC addresses get pushed *twice* to the hardware,
> once by the driver and once by the eth_dev_mac_restore() in ethdev.
>
> On the other hand, MAC address filters are reset in the hardware
> by the PF only when a VF reset is triggered.
>
> Strictly speaking, the mac restore on port (re)start is unneeded,
> if no VF reset happened, so we can announce to ethdev that no mac
> restoration is needed via a get_restore_flags callback.
>
> Then, move the mac restoration to the VF reset handler.
>
> Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux
> driver")
> Cc: stable@dpdk.org
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
> Changes since v6:
> - renamed iavf_add_del_all_mac_addr and removed primary mac address
> handling out of this helper (avoids double primary mac installation),
>
> Changes since v4:
> - rebased on next-net-intel,
>
> Changes since v4:
> - moved mac restoration in iavf_post_reset_reconfig,
>
> ---
> drivers/net/intel/iavf/iavf.h | 2 +-
> drivers/net/intel/iavf/iavf_ethdev.c | 30 ++++++++++++++++++----------
> drivers/net/intel/iavf/iavf_vchnl.c | 11 +++++-----
> 3 files changed, 26 insertions(+), 17 deletions(-)
>
> diff --git a/drivers/net/intel/iavf/iavf.h b/drivers/net/intel/iavf/iavf.h
> index 1d52b3d113..34b9b4ad94 100644
> --- a/drivers/net/intel/iavf/iavf.h
> +++ b/drivers/net/intel/iavf/iavf.h
> @@ -482,7 +482,7 @@ int iavf_add_del_vlan_v2(struct iavf_adapter
> *adapter, uint16_t vlanid,
> int iavf_get_vlan_offload_caps_v2(struct iavf_adapter *adapter);
> int iavf_config_irq_map(struct iavf_adapter *adapter);
> int iavf_config_irq_map_lv(struct iavf_adapter *adapter, uint16_t num);
> -void iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add);
> +void iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool
> add);
> int iavf_dev_link_update(struct rte_eth_dev *dev,
> __rte_unused int wait_to_complete);
> void iavf_dev_alarm_handler(void *param);
> diff --git a/drivers/net/intel/iavf/iavf_ethdev.c
> b/drivers/net/intel/iavf/iavf_ethdev.c
> index b498258d58..f7aeac8c83 100644
> --- a/drivers/net/intel/iavf/iavf_ethdev.c
> +++ b/drivers/net/intel/iavf/iavf_ethdev.c
> @@ -292,11 +292,12 @@ iavf_get_restore_flags(__rte_unused struct
> rte_eth_dev *dev,
> __rte_unused enum rte_eth_dev_operation op)
> {
> /*
> - * The unicast and multicast promiscuous settings persist across a
> + * The mac addresses, unicast and multicast promiscuous settings
> persist across a
> * stop/start; they are only cleared by a VF reset, which the driver
> * restores itself. So ethdev does not need to restore them on start.
> */
> - return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_PROMISC |
> + return RTE_ETH_RESTORE_ALL & ~(RTE_ETH_RESTORE_MAC_ADDR |
> + RTE_ETH_RESTORE_PROMISC |
> RTE_ETH_RESTORE_ALLMULTI);
> }
>
> @@ -1095,15 +1096,14 @@ iavf_dev_start(struct rte_eth_dev *dev)
> rte_intr_enable(intr_handle);
> }
>
> - /* Set all mac addrs */
> - iavf_add_del_all_mac_addr(adapter, true);
> -
> - if (!adapter->mac_primary_set)
> - adapter->mac_primary_set = true;
> -
> - /* Set all multicast addresses */
> - iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> >mc_addrs_num,
> - true);
> + if (!adapter->mac_primary_set) {
> + if (iavf_add_del_eth_addr(adapter, &dev->data-
> >mac_addrs[0], true,
> + VIRTCHNL_ETHER_ADDR_PRIMARY) != 0)
> + PMD_DRV_LOG(ERR, "failed to add primary MAC:"
> RTE_ETHER_ADDR_PRT_FMT,
> + RTE_ETHER_ADDR_BYTES(&dev->data-
> >mac_addrs[0]));
> + else
> + adapter->mac_primary_set = true;
> + }
>
> rte_spinlock_init(&vf->phc_time_aq_lock);
>
> @@ -3428,6 +3428,14 @@ iavf_post_reset_reconfig(struct rte_eth_dev
> *dev)
> int ret = 0;
> bool allmulti = false, allunicast = false;
> struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev-
> >data->dev_private);
> + struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data-
> >dev_private);
> +
> + /*
> + * After a VF reset, all MAC addresses got flushed.
> + * The primary MAC should have been set in iavf_dev_start, restore
iavf_dev_start is not guaranteed to have been executed before this handler.
So I would change this comment to something like:
"The primary MAC has been or will be restored by iavf_dev_start".
Other than that:
Acked-by: Ciara Loftus <ciara.loftus@intel.com>
> the rest.
> + */
> + iavf_add_del_secondary_mac_addr(adapter, true);
> + (void)iavf_add_del_mc_addr_list(adapter, vf->mc_addrs, vf-
> >mc_addrs_num, true);
>
> /* Restore pre-reset unicast promiscuous and multicast promiscuous
> states */
> if (dev->data->promiscuous)
> diff --git a/drivers/net/intel/iavf/iavf_vchnl.c
> b/drivers/net/intel/iavf/iavf_vchnl.c
> index df97ff2052..a918db5443 100644
> --- a/drivers/net/intel/iavf/iavf_vchnl.c
> +++ b/drivers/net/intel/iavf/iavf_vchnl.c
> @@ -1672,7 +1672,7 @@ iavf_config_irq_map_lv(struct iavf_adapter
> *adapter, uint16_t num)
> }
>
> void
> -iavf_add_del_all_mac_addr(struct iavf_adapter *adapter, bool add)
> +iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool add)
> {
> struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
> struct {
> @@ -1685,7 +1685,7 @@ iavf_add_del_all_mac_addr(struct iavf_adapter
> *adapter, bool add)
> int err;
> size_t buf_len;
>
> - for (unsigned int i = 0; i < RTE_DIM(vf->mac_addrs); i++) {
> + for (unsigned int i = 1; i < RTE_DIM(vf->mac_addrs); i++) {
> struct rte_ether_addr *addr = &vf->mac_addrs[i];
> struct virtchnl_ether_addr *vc_addr = &list->list[list-
> >num_elements];
>
> @@ -1695,11 +1695,12 @@ iavf_add_del_all_mac_addr(struct iavf_adapter
> *adapter, bool add)
> list->num_elements++;
>
> memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr-
> >addr_bytes));
> - vc_addr->type = (list->num_elements == 1) ?
> - VIRTCHNL_ETHER_ADDR_PRIMARY :
> - VIRTCHNL_ETHER_ADDR_EXTRA;
> + vc_addr->type = VIRTCHNL_ETHER_ADDR_EXTRA;
> }
>
> + if (list->num_elements == 0)
> + return;
> +
> /* for some reason PF side checks for buffer being too big, so adjust it
> down */
> buf_len = sizeof(struct virtchnl_ether_addr_list) +
> sizeof(struct virtchnl_ether_addr) * list->num_elements;
> --
> 2.54.0
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset
2026-09-14 10:15 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset Loftus, Ciara
@ 2026-09-14 11:54 ` David Marchand
2026-09-14 11:57 ` Loftus, Ciara
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-14 11:54 UTC (permalink / raw)
To: Loftus, Ciara
Cc: dev@dpdk.org, Burakov, Anatoly, rjarry@redhat.com,
cfontain@redhat.com, stable@dpdk.org, Medvedkin, Vladimir,
Lunyuan Cui, Wu, Jingjing, Xiaolong Ye
On Mon, 14 Sept 2026 at 12:16, Loftus, Ciara <ciara.loftus@intel.com> wrote:
> > @@ -3123,7 +3116,6 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
> > rte_ether_addr_copy((struct rte_ether_addr *)hw->mac.addr,
> > ð_dev->data->mac_addrs[0]);
> >
> > -
>
> Nit: stray line removal. The rest LGTM.
I did this on purpose.
I was reviewing accesses to eth_dev->data->mac_addrs and noticed this
stray double line.
No strong opinion, but I would prefer it is removed.
> Acked-by: Ciara Loftus <ciara.loftus@intel.com>
Thanks.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install
2026-09-14 10:19 ` Loftus, Ciara
@ 2026-09-14 11:56 ` David Marchand
2026-09-14 12:02 ` Bruce Richardson
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-14 11:56 UTC (permalink / raw)
To: Loftus, Ciara, Richardson, Bruce
Cc: dev@dpdk.org, Burakov, Anatoly, rjarry@redhat.com,
cfontain@redhat.com, stable@dpdk.org, Medvedkin, Vladimir
On Mon, 14 Sept 2026 at 12:19, Loftus, Ciara <ciara.loftus@intel.com> wrote:
> > @@ -3428,6 +3428,14 @@ iavf_post_reset_reconfig(struct rte_eth_dev
> > *dev)
> > int ret = 0;
> > bool allmulti = false, allunicast = false;
> > struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev-
> > >data->dev_private);
> > + struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data-
> > >dev_private);
> > +
> > + /*
> > + * After a VF reset, all MAC addresses got flushed.
> > + * The primary MAC should have been set in iavf_dev_start, restore
>
> iavf_dev_start is not guaranteed to have been executed before this handler.
> So I would change this comment to something like:
> "The primary MAC has been or will be restored by iavf_dev_start".
>
> Other than that:
>
> Acked-by: Ciara Loftus <ciara.loftus@intel.com>
Yes, true.
I'll update in a new revision if needed, otherwise, could it be
updated when applying?
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* RE: [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset
2026-09-14 11:54 ` David Marchand
@ 2026-09-14 11:57 ` Loftus, Ciara
0 siblings, 0 replies; 146+ messages in thread
From: Loftus, Ciara @ 2026-09-14 11:57 UTC (permalink / raw)
To: David Marchand
Cc: dev@dpdk.org, Burakov, Anatoly, rjarry@redhat.com,
cfontain@redhat.com, stable@dpdk.org, Medvedkin, Vladimir,
Lunyuan Cui, Wu, Jingjing, Xiaolong Ye
> Subject: Re: [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset
>
> On Mon, 14 Sept 2026 at 12:16, Loftus, Ciara <ciara.loftus@intel.com> wrote:
> > > @@ -3123,7 +3116,6 @@ iavf_dev_init(struct rte_eth_dev *eth_dev)
> > > rte_ether_addr_copy((struct rte_ether_addr *)hw->mac.addr,
> > > ð_dev->data->mac_addrs[0]);
> > >
> > > -
> >
> > Nit: stray line removal. The rest LGTM.
>
> I did this on purpose.
>
> I was reviewing accesses to eth_dev->data->mac_addrs and noticed this
> stray double line.
> No strong opinion, but I would prefer it is removed.
No objections to the removal on my end. Thanks for clarifying.
>
> > Acked-by: Ciara Loftus <ciara.loftus@intel.com>
>
> Thanks.
>
>
> --
> David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install
2026-09-14 11:56 ` David Marchand
@ 2026-09-14 12:02 ` Bruce Richardson
2026-09-14 12:27 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Bruce Richardson @ 2026-09-14 12:02 UTC (permalink / raw)
To: David Marchand
Cc: Loftus, Ciara, dev@dpdk.org, Burakov, Anatoly, rjarry@redhat.com,
cfontain@redhat.com, stable@dpdk.org, Medvedkin, Vladimir
On Mon, Sep 14, 2026 at 01:56:36PM +0200, David Marchand wrote:
> On Mon, 14 Sept 2026 at 12:19, Loftus, Ciara <ciara.loftus@intel.com> wrote:
> > > @@ -3428,6 +3428,14 @@ iavf_post_reset_reconfig(struct rte_eth_dev
> > > *dev)
> > > int ret = 0;
> > > bool allmulti = false, allunicast = false;
> > > struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev-
> > > >data->dev_private);
> > > + struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data-
> > > >dev_private);
> > > +
> > > + /*
> > > + * After a VF reset, all MAC addresses got flushed.
> > > + * The primary MAC should have been set in iavf_dev_start, restore
> >
> > iavf_dev_start is not guaranteed to have been executed before this handler.
> > So I would change this comment to something like:
> > "The primary MAC has been or will be restored by iavf_dev_start".
> >
> > Other than that:
> >
> > Acked-by: Ciara Loftus <ciara.loftus@intel.com>
>
> Yes, true.
> I'll update in a new revision if needed, otherwise, could it be
> updated when applying?
>
Single line text changes I can handle on apply, no problem, so long as the
rest of the set is good to go.
/Bruce
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install
2026-09-14 12:02 ` Bruce Richardson
@ 2026-09-14 12:27 ` David Marchand
0 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-14 12:27 UTC (permalink / raw)
To: Bruce Richardson
Cc: Loftus, Ciara, dev@dpdk.org, Burakov, Anatoly, rjarry@redhat.com,
cfontain@redhat.com, stable@dpdk.org, Medvedkin, Vladimir
On Mon, 14 Sept 2026 at 14:02, Bruce Richardson
<bruce.richardson@intel.com> wrote:
> On Mon, Sep 14, 2026 at 01:56:36PM +0200, David Marchand wrote:
> > On Mon, 14 Sept 2026 at 12:19, Loftus, Ciara <ciara.loftus@intel.com> wrote:
> > > > @@ -3428,6 +3428,14 @@ iavf_post_reset_reconfig(struct rte_eth_dev
> > > > *dev)
> > > > int ret = 0;
> > > > bool allmulti = false, allunicast = false;
> > > > struct iavf_adapter *adapter = IAVF_DEV_PRIVATE_TO_ADAPTER(dev-
> > > > >data->dev_private);
> > > > + struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(dev->data-
> > > > >dev_private);
> > > > +
> > > > + /*
> > > > + * After a VF reset, all MAC addresses got flushed.
> > > > + * The primary MAC should have been set in iavf_dev_start, restore
> > >
> > > iavf_dev_start is not guaranteed to have been executed before this handler.
> > > So I would change this comment to something like:
> > > "The primary MAC has been or will be restored by iavf_dev_start".
> > >
> > > Other than that:
> > >
> > > Acked-by: Ciara Loftus <ciara.loftus@intel.com>
> >
> > Yes, true.
> > I'll update in a new revision if needed, otherwise, could it be
> > updated when applying?
> >
> Single line text changes I can handle on apply, no problem, so long as the
> rest of the set is good to go.
Yes, sure.
Thanks Bruce.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (12 preceding siblings ...)
2026-09-14 8:17 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset David Marchand
@ 2026-09-14 14:42 ` David Marchand
2026-09-14 14:42 ` [PATCH v7 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
` (4 more replies)
2026-09-21 11:50 ` [PATCH v8 " David Marchand
2026-09-24 6:38 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
15 siblings, 5 replies; 146+ messages in thread
From: David Marchand @ 2026-09-14 14:42 UTC (permalink / raw)
To: dev
Cc: dsosnowski, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao,
Ori Kam, Suanming Mou, Matan Azrad
This helper is exposing internals of net/mlx5 for no good reason.
All this code does is calling the remove helper.
Walk through the list in Linux implementation like the Windows
implementation.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 40 -----------------------------
drivers/common/mlx5/linux/mlx5_nl.h | 5 ----
drivers/net/mlx5/linux/mlx5_os.c | 13 +++++++---
3 files changed, 10 insertions(+), 48 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 42ccb73e36..40a8b8a2cc 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -825,46 +825,6 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
}
-/**
- * Flush all added MAC addresses.
- *
- * @param[in] nlsk_fd
- * Netlink socket file descriptor.
- * @param[in] iface_idx
- * Net device interface index.
- * @param[in] mac_addrs
- * Mac addresses array to flush.
- * @param n
- * @p mac_addrs array size.
- * @param mac_own
- * BITFIELD_DECLARE array to store the mac.
- * @param vf
- * Flag for a VF device.
- */
-RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_flush)
-void
-mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own, bool vf)
-{
- int i;
-
- if (n <= 0 || n > MLX5_MAX_MAC_ADDRESSES)
- return;
-
- for (i = n - 1; i >= 0; --i) {
- struct rte_ether_addr *m = &mac_addrs[i];
-
- if (BITFIELD_ISSET(mac_own, i)) {
- if (vf)
- mlx5_nl_mac_addr_remove(nlsk_fd,
- iface_idx,
- m, i);
- BITFIELD_RESET(mac_own, i);
- }
- }
-}
-
/**
* Enable promiscuous / all multicast mode through Netlink.
*
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 8ccdd244b0..0242342c47 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -65,11 +65,6 @@ __rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
__rte_internal
-void mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own,
- bool vf);
-__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
int mlx5_nl_allmulti(int nlsk_fd, unsigned int iface_idx, int enable);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 592e233844..49a8751761 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3529,10 +3529,17 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
{
struct mlx5_priv *priv = dev->data->dev_private;
const int vf = priv->sh->dev_cap.vf;
+ int i;
- mlx5_nl_mac_addr_flush(priv->nl_socket_route, mlx5_ifindex(dev),
- dev->data->mac_addrs,
- MLX5_MAX_MAC_ADDRESSES, priv->mac_own, vf);
+ for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (vf)
+ mlx5_nl_mac_addr_remove(priv->nl_socket_route,
+ mlx5_ifindex(dev),
+ &dev->data->mac_addrs[i], i);
+ BITFIELD_RESET(priv->mac_own, i);
+ }
+ }
}
static bool
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v7 2/5] net/mlx5: remove redundant MAC address index checks
2026-09-14 14:42 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
@ 2026-09-14 14:42 ` David Marchand
2026-09-21 8:31 ` Raslan Darawsheh
2026-09-14 14:42 ` [PATCH v7 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
` (3 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-14 14:42 UTC (permalink / raw)
To: dev
Cc: dsosnowski, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao,
Ori Kam, Suanming Mou, Matan Azrad
On the net/mlx5 side, mlx5_mac.c validates that any MAC address and its
index is valid before calling the OS specific helpers.
So those OS helpers do not have to validate again the index.
Cascading this consideration, validating the MAC index against
MLX5_MAX_MAC_ADDRESSES in common code is also redundant.
The common code only deals with netlink, remove any index concern.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 17 ++---------------
drivers/common/mlx5/linux/mlx5_nl.h | 4 ++--
drivers/net/mlx5/linux/mlx5_os.c | 9 ++++-----
3 files changed, 8 insertions(+), 22 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 40a8b8a2cc..3207eae563 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -719,8 +719,6 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to register.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -728,16 +726,11 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_add)
int
mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
int ret;
ret = mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 1);
- if (!ret) {
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
- }
if (ret == -EEXIST)
return 0;
return ret;
@@ -752,8 +745,6 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to remove.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -761,12 +752,8 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_remove)
int
mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
-
return mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 0);
}
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 0242342c47..0d6259f4ad 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -57,10 +57,10 @@ __rte_internal
int mlx5_nl_init(int protocol, int groups);
__rte_internal
int mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 49a8751761..c65293cb25 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3382,9 +3382,8 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[index], index);
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ &dev->data->mac_addrs[index]);
+ BITFIELD_RESET(priv->mac_own, index);
}
/**
@@ -3411,7 +3410,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
if (vf)
ret = mlx5_nl_mac_addr_add(priv->nl_socket_route,
mlx5_ifindex(dev),
- mac, index);
+ mac);
if (!ret)
BITFIELD_SET(priv->mac_own, index);
@@ -3536,7 +3535,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[i], i);
+ &dev->data->mac_addrs[i]);
BITFIELD_RESET(priv->mac_own, i);
}
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v7 3/5] net/mlx5: pass maximum number of unicast MAC to common code
2026-09-14 14:42 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-14 14:42 ` [PATCH v7 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-09-14 14:42 ` David Marchand
2026-09-21 8:31 ` Raslan Darawsheh
2026-09-14 14:42 ` [PATCH v7 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
` (2 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-14 14:42 UTC (permalink / raw)
To: dev
Cc: dsosnowski, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao,
Ori Kam, Suanming Mou, Matan Azrad
Isolate how the MAC addresses array is walked through in the common code
by passing the max index at which a unicast MAC address is stored in
dev->data->mac_addrs[].
With this change, only net/mlx5 knows about the max number of
unicast/multicast MAC addresses.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
Changes since v5:
- fixed mlx5_nl_mac_addr_cb,
---
drivers/common/mlx5/linux/mlx5_nl.c | 20 ++++++++++++--------
drivers/common/mlx5/linux/mlx5_nl.h | 2 +-
drivers/common/mlx5/mlx5_common.h | 8 --------
drivers/net/mlx5/linux/mlx5_os.c | 1 +
drivers/net/mlx5/mlx5.h | 8 ++++++++
5 files changed, 22 insertions(+), 17 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 3207eae563..12942eefa5 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -174,6 +174,7 @@ struct mlx5_nl_mac_addr {
struct rte_ether_addr (*mac)[];
/**< MAC address handled by the device. */
int mac_n; /**< Number of addresses in the array. */
+ int max_macs; /**< Size of the array. */
};
static RTE_ATOMIC(uint32_t) atomic_sn;
@@ -475,7 +476,7 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
RTA_OK(attribute, len);
attribute = RTA_NEXT(attribute, len)) {
if (attribute->rta_type == NDA_LLADDR) {
- if (data->mac_n == MLX5_MAX_MAC_ADDRESSES) {
+ if (data->mac_n == data->max_macs) {
DRV_LOG(WARNING,
"not enough room to finalize the"
" request");
@@ -505,9 +506,9 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
* Net device interface index.
* @param mac[out]
* Pointer to the array table of MAC addresses to fill.
- * Its size should be of MLX5_MAX_MAC_ADDRESSES.
- * @param mac_n[out]
- * Number of entries filled in MAC array.
+ * @param mac_n[in,out]
+ * Size of the MAC array on input.
+ * Number of entries filled in MAC array on output.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -533,6 +534,7 @@ mlx5_nl_mac_addr_list(int nlsk_fd, unsigned int iface_idx,
struct mlx5_nl_mac_addr data = {
.mac = mac,
.mac_n = 0,
+ .max_macs = *mac_n,
};
uint32_t sn = MLX5_NL_SN_GENERATE;
int ret;
@@ -766,16 +768,18 @@ mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac_addrs
* Mac addresses array to sync.
+ * @param uc_n
+ * Number of UC entries in @p mac_addrs.
* @param n
* @p mac_addrs array size.
*/
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_sync)
void
mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n)
+ struct rte_ether_addr *mac_addrs, int uc_n, int n)
{
struct rte_ether_addr macs[n];
- int macs_n = 0;
+ int macs_n = n;
int i;
int ret;
@@ -794,7 +798,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
continue;
if (rte_is_multicast_ether_addr(&macs[i])) {
/* Find the first entry available. */
- for (j = MLX5_MAX_UC_MAC_ADDRESSES; j != n; ++j) {
+ for (j = uc_n; j != n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
@@ -802,7 +806,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
} else {
/* Find the first entry available. */
- for (j = 0; j != MLX5_MAX_UC_MAC_ADDRESSES; ++j) {
+ for (j = 0; j != uc_n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 0d6259f4ad..07a3b531b5 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -63,7 +63,7 @@ int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n);
+ struct rte_ether_addr *mac_addrs, int uc_n, int n);
__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 3767020823..8aed3dbabe 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -160,14 +160,6 @@ enum {
PCI_DEVICE_ID_MELLANOX_CONNECTX10C2C = 0x2101,
};
-/* Maximum number of simultaneous unicast MAC addresses. */
-#define MLX5_MAX_UC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous Multicast MAC addresses. */
-#define MLX5_MAX_MC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous MAC addresses. */
-#define MLX5_MAX_MAC_ADDRESSES \
- (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
-
/* Recognized Infiniband device physical port name types. */
enum mlx5_nl_phys_port_name_type {
MLX5_PHYS_PORT_NAME_TYPE_NOTSET = 0, /* Not set. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index c65293cb25..0e59d3f2f5 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -1762,6 +1762,7 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
+ MLX5_MAX_UC_MAC_ADDRESSES,
MLX5_MAX_MAC_ADDRESSES);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 190d203c49..c77cdc48a8 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -82,6 +82,14 @@
/* Maximum allowed MTU to be reported whenever PMD cannot query it from OS. */
#define MLX5_ETH_MAX_MTU (9978)
+/* Maximum number of simultaneous unicast MAC addresses. */
+#define MLX5_MAX_UC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous Multicast MAC addresses. */
+#define MLX5_MAX_MC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous MAC addresses. */
+#define MLX5_MAX_MAC_ADDRESSES \
+ (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
+
enum mlx5_ipool_index {
#if defined(HAVE_IBV_FLOW_DV_SUPPORT) || !defined(HAVE_INFINIBAND_VERBS_H)
MLX5_IPOOL_DECAP_ENCAP = 0, /* Pool for encap/decap resource. */
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v7 4/5] net/mlx5: use bitset for tracking MAC addresses
2026-09-14 14:42 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-14 14:42 ` [PATCH v7 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
2026-09-14 14:42 ` [PATCH v7 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-09-14 14:42 ` David Marchand
2026-09-14 14:42 ` [PATCH v7 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-21 8:31 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux Raslan Darawsheh
4 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-14 14:42 UTC (permalink / raw)
To: dev
Cc: dsosnowski, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao,
Ori Kam, Suanming Mou, Matan Azrad
EAL provides bitset that does the same as this set of mlx5 macros.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
drivers/common/mlx5/mlx5_common.h | 16 ----------------
drivers/net/mlx5/linux/mlx5_os.c | 8 ++++----
drivers/net/mlx5/mlx5.h | 3 ++-
drivers/net/mlx5/mlx5_trigger.c | 2 +-
drivers/net/mlx5/windows/mlx5_os.c | 8 ++++----
5 files changed, 11 insertions(+), 26 deletions(-)
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 8aed3dbabe..8b8f20e2f4 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -30,22 +30,6 @@
#define MLX5_PCI_DRIVER_NAME "mlx5_pci"
#define MLX5_AUXILIARY_DRIVER_NAME "mlx5_auxiliary"
-/* Bit-field manipulation. */
-#define BITFIELD_DECLARE(bf, type, size) \
- type bf[(((size_t)(size) / (sizeof(type) * CHAR_BIT)) + \
- !!((size_t)(size) % (sizeof(type) * CHAR_BIT)))]
-#define BITFIELD_DEFINE(bf, type, size) \
- BITFIELD_DECLARE((bf), type, (size)) = { 0 }
-#define BITFIELD_SET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] |= \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_RESET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] &= \
- ~((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_ISSET(bf, b) \
- !!(((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] & \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT)))))
-
/*
* Helper macros to work around __VA_ARGS__ limitations in a C99 compliant
* manner.
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 0e59d3f2f5..7d8ea4acba 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3384,7 +3384,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[index]);
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -3413,7 +3413,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
mlx5_ifindex(dev),
mac);
if (!ret)
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return ret;
}
@@ -3532,12 +3532,12 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[i]);
- BITFIELD_RESET(priv->mac_own, i);
+ rte_bitset_clear(priv->mac_own, i);
}
}
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index c77cdc48a8..7f20c811f3 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -14,6 +14,7 @@
#include <rte_pci.h>
#include <rte_ether.h>
+#include <rte_bitset.h>
#include <ethdev_driver.h>
#include <rte_rwlock.h>
#include <rte_interrupts.h>
@@ -2018,7 +2019,7 @@ struct mlx5_priv {
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- BITFIELD_DECLARE(mac_own, uint64_t, MLX5_MAX_MAC_ADDRESSES);
+ RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
/* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 2a91e02b45..7f6148f4e1 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1921,7 +1921,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
/* Add flows for unicast and multicast mac addresses added by API. */
if (!memcmp(mac, &cmp, sizeof(*mac)) ||
- !BITFIELD_ISSET(priv->mac_own, i) ||
+ !rte_bitset_test(priv->mac_own, i) ||
(dev->data->all_multicast && rte_is_multicast_ether_addr(mac)))
continue;
memcpy(&unicast.hdr.dst_addr.addr_bytes,
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index cf34e4e1d6..eaa6c25c09 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -699,8 +699,8 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i))
- BITFIELD_RESET(priv->mac_own, i);
+ if (rte_bitset_test(priv->mac_own, i))
+ rte_bitset_clear(priv->mac_own, i);
}
}
@@ -719,7 +719,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
struct mlx5_priv *priv = dev->data->dev_private;
if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -757,7 +757,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
return -ENOTSUP;
}
/* Mark this MAC address as owned by the PMD */
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return 0;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v7 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-14 14:42 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (2 preceding siblings ...)
2026-09-14 14:42 ` [PATCH v7 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
@ 2026-09-14 14:42 ` David Marchand
2026-09-14 14:51 ` Dariusz Sosnowski
2026-09-21 8:31 ` Raslan Darawsheh
2026-09-21 8:31 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux Raslan Darawsheh
4 siblings, 2 replies; 146+ messages in thread
From: David Marchand @ 2026-09-14 14:42 UTC (permalink / raw)
To: dev
Cc: dsosnowski, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao,
Ori Kam, Suanming Mou, Matan Azrad
Starting firmware version 22.49.1014, the number of mac addresses
per VF is not capped to 128 anymore.
The value can be increased via devlink:
$ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
cmode driverinit
$ devlink dev reload pci/0000:3b:00.2
On the DPDK side, we must retrieve the maximum number of unicast
and multicast addresses supported with a query to the firmware.
Then, dynamically allocate the mac addresses arrays and report the
limit instead of the previous hardcoded value.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v6:
- fixed crash in case of partial init failure in mlx5_dev_spawn,
and cleaned mlx5_os_pci_probe_pf,
- fixed compilation without assert in mlx5_internal_mac_addr_remove,
Changes since v4:
- added RN update,
- fixed types of fields added to mlx5_hca_attr and mlx5_dev_cap,
- used RTE_BIT32,
- fixed mlx5_nl_mac_addr_sync inverted arguments,
---
doc/guides/rel_notes/release_26_11.rst | 5 +++
drivers/common/mlx5/mlx5_devx_cmds.c | 4 +++
drivers/common/mlx5/mlx5_devx_cmds.h | 2 ++
drivers/net/mlx5/linux/mlx5_os.c | 47 +++++++++++++++++++-------
drivers/net/mlx5/mlx5.c | 9 ++---
drivers/net/mlx5/mlx5.h | 8 +++--
drivers/net/mlx5/mlx5_ethdev.c | 2 +-
drivers/net/mlx5/mlx5_mac.c | 20 ++++++-----
drivers/net/mlx5/mlx5_trigger.c | 10 +++---
drivers/net/mlx5/windows/mlx5_os.c | 43 ++++++++++++++++++-----
10 files changed, 106 insertions(+), 44 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 87c7e81bde..43043cd579 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,11 @@ New Features
Also, make sure to start the actual text at the margin.
=======================================================
+* **Updated NVIDIA mlx5 ethernet driver.**
+
+ * Increased the maximum number of secondary unicast MAC addresses from 128 to up to 4096
+ (depending on devlink configuration on the associated kernel netdevice).
+
Removed Items
-------------
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.c b/drivers/common/mlx5/mlx5_devx_cmds.c
index 140b057ab4..e5d9c92779 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.c
+++ b/drivers/common/mlx5/mlx5_devx_cmds.c
@@ -1110,6 +1110,10 @@ mlx5_devx_cmd_query_hca_attr(void *ctx,
attr->log_max_pd = MLX5_GET(cmd_hca_cap, hcattr, log_max_pd);
attr->log_max_srq = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq);
attr->log_max_srq_sz = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq_sz);
+ attr->log_max_current_uc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_uc_list);
+ attr->log_max_current_mc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_mc_list);
attr->reg_c_preserve =
MLX5_GET(cmd_hca_cap, hcattr, reg_c_preserve);
attr->mmo_regex_qp_en = MLX5_GET(cmd_hca_cap, hcattr, regexp_mmo_qp);
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.h b/drivers/common/mlx5/mlx5_devx_cmds.h
index 90beb2e9e6..504b4a4f64 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.h
+++ b/drivers/common/mlx5/mlx5_devx_cmds.h
@@ -356,6 +356,8 @@ struct mlx5_hca_attr {
uint8_t tx_sw_owner_v2:1;
uint8_t esw_sw_owner:1;
uint8_t esw_sw_owner_v2:1;
+ uint8_t log_max_current_uc_list:5;
+ uint8_t log_max_current_mc_list:5;
};
/* LAG Context. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 7d8ea4acba..9180e9aa20 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -389,6 +389,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
sh->dev_cap.esw_info.regc_mask = 0;
#endif
sh->dev_cap.esw_info.is_set = 1;
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -1464,6 +1474,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
/* Some internal functions rely on Netlink sockets, open them now. */
priv->nl_socket_rdma = nl_rdma;
priv->nl_socket_route = mlx5_nl_init(NETLINK_ROUTE, 0);
@@ -1762,8 +1788,8 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
- MLX5_MAX_UC_MAC_ADDRESSES,
- MLX5_MAX_MAC_ADDRESSES);
+ sh->dev_cap.max_uc_mac_addrs,
+ sh->dev_cap.max_mac_addrs);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
TAILQ_INIT(&priv->flow_meters);
@@ -1963,17 +1989,16 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_flex_item_port_cleanup(eth_dev);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
- if (eth_dev != NULL)
+ if (eth_dev != NULL) {
+ eth_dev->data->mac_addrs = NULL;
eth_dev->data->dev_private = NULL;
+ }
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
if (nl_rdma >= 0)
@@ -2979,8 +3004,6 @@ mlx5_os_pci_probe_pf(struct mlx5_common_device *cdev,
if (!list[i].eth_dev)
continue;
mlx5_dev_close(list[i].eth_dev);
- /* mac_addrs must not be freed because in dev_private */
- list[i].eth_dev->data->mac_addrs = NULL;
claim_zero(rte_eth_dev_release_port(list[i].eth_dev));
}
/* Restore original error. */
@@ -3531,7 +3554,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
const int vf = priv->sh->dev_cap.vf;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
diff --git a/drivers/net/mlx5/mlx5.c b/drivers/net/mlx5/mlx5.c
index c7b0d3ef8b..4bd1c6d1c0 100644
--- a/drivers/net/mlx5/mlx5.c
+++ b/drivers/net/mlx5/mlx5.c
@@ -2556,6 +2556,9 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_list_destroy(priv->hrxqs);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
sh->port[priv->dev_port - 1].nl_ih_port_id = RTE_MAX_ETHPORTS;
/*
* The interrupt handler port id must be reset before priv is reset
@@ -2590,12 +2593,6 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_flow_pools_destroy(priv);
memset(priv, 0, sizeof(*priv));
priv->domain_id = RTE_ETH_DEV_SWITCH_DOMAIN_ID_INVALID;
- /*
- * Reset mac_addrs to NULL such that it is not freed as part of
- * rte_eth_dev_release_port(). mac_addrs is part of dev_private so
- * it is freed when dev_private is freed.
- */
- dev->data->mac_addrs = NULL;
return 0;
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 7f20c811f3..70efcda9bb 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -217,6 +217,9 @@ struct mlx5_dev_cap {
} mprq; /* Capability for Multi-Packet RQ. */
char fw_ver[64]; /* Firmware version of this device. */
struct flow_hw_port_info esw_info; /* E-switch manager reg_c0. */
+ uint32_t max_uc_mac_addrs; /* Maximum unicast MAC addresses. */
+ uint32_t max_mc_mac_addrs; /* Maximum multicast MAC addresses. */
+ uint32_t max_mac_addrs; /* Total maximum MAC addresses. */
};
#define MLX5_MPESW_PORT_INVALID (-1)
@@ -2018,9 +2021,8 @@ struct mlx5_priv {
struct mlx5_dev_ctx_shared *sh; /* Shared device context. */
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
- struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
- /* Bit-field of MAC addresses owned by the PMD. */
+ struct rte_ether_addr *mac; /* MAC addresses. */
+ uint64_t *mac_own; /* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
/* Device properties. */
diff --git a/drivers/net/mlx5/mlx5_ethdev.c b/drivers/net/mlx5/mlx5_ethdev.c
index 8160d10e7e..306c1cd734 100644
--- a/drivers/net/mlx5/mlx5_ethdev.c
+++ b/drivers/net/mlx5/mlx5_ethdev.c
@@ -392,7 +392,7 @@ mlx5_dev_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *info)
max = RTE_MIN(max, (unsigned int)UINT16_MAX);
info->max_rx_queues = max;
info->max_tx_queues = max;
- info->max_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ info->max_mac_addrs = priv->sh->dev_cap.max_uc_mac_addrs;
info->rx_queue_offload_capa = mlx5_get_rx_queue_offloads(dev);
info->rx_seg_capa.max_nseg = MLX5_MAX_RXQ_NSEG;
info->rx_seg_capa.multi_pools = !priv->config.mprq.enabled;
diff --git a/drivers/net/mlx5/mlx5_mac.c b/drivers/net/mlx5/mlx5_mac.c
index 0e5d2be530..9e7bf7f2ec 100644
--- a/drivers/net/mlx5/mlx5_mac.c
+++ b/drivers/net/mlx5/mlx5_mac.c
@@ -36,7 +36,7 @@ mlx5_internal_mac_addr_remove(struct rte_eth_dev *dev,
uint32_t index,
struct rte_ether_addr *addr)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < MLX5_SH(dev)->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(&dev->data->mac_addrs[index]))
return false;
mlx5_os_mac_addr_remove(dev, index);
@@ -63,16 +63,17 @@ static int
mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
unsigned int i;
int ret;
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(mac)) {
rte_errno = EINVAL;
return -rte_errno;
}
/* First, make sure this address isn't already configured. */
- for (i = 0; (i != MLX5_MAX_MAC_ADDRESSES); ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
/* Skip this index, it's going to be reconfigured. */
if (i == index)
continue;
@@ -101,10 +102,11 @@ mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
void
mlx5_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
struct rte_ether_addr addr = { 0 };
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES)
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs)
return;
if (mlx5_internal_mac_addr_remove(dev, index, &addr)) {
ret = mlx5_traffic_mac_remove(dev, &addr);
@@ -133,9 +135,10 @@ int
mlx5_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index, uint32_t vmdq __rte_unused)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES) {
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs) {
rte_errno = EINVAL;
return -rte_errno;
}
@@ -217,16 +220,17 @@ int
mlx5_set_mc_addr_list(struct rte_eth_dev *dev,
struct rte_ether_addr *mc_addr_set, uint32_t nb_mc_addr)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
uint32_t i;
int ret;
- if (nb_mc_addr >= MLX5_MAX_MC_MAC_ADDRESSES) {
+ if (nb_mc_addr >= priv->sh->dev_cap.max_mc_mac_addrs) {
rte_errno = ENOSPC;
return -rte_errno;
}
- for (i = MLX5_MAX_UC_MAC_ADDRESSES; i != MLX5_MAX_MAC_ADDRESSES; ++i)
+ for (i = priv->sh->dev_cap.max_uc_mac_addrs; i != priv->sh->dev_cap.max_mac_addrs; ++i)
mlx5_internal_mac_addr_remove(dev, i, NULL);
- i = MLX5_MAX_UC_MAC_ADDRESSES;
+ i = priv->sh->dev_cap.max_uc_mac_addrs;
while (nb_mc_addr--) {
ret = mlx5_internal_mac_addr_add(dev, mc_addr_set++, i++);
if (ret)
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 7f6148f4e1..c5f493117b 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1916,7 +1916,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
}
}
/* Add MAC address flows. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
/* Add flows for unicast and multicast mac addresses added by API. */
@@ -2186,7 +2186,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
return 0;
/* Add all unicast DMAC flow rules with new VLAN attached. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2203,7 +2203,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
* Removing after creating VLAN rules so that traffic "gap" is not introduced.
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2241,7 +2241,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
* Recreating first to ensure no traffic "gap".
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2254,7 +2254,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
}
/* Remove all unicast DMAC flow rules with this VLAN. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index eaa6c25c09..f712ddae4b 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -261,6 +261,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
MLX5_GET(initial_seg, pv_iseg, fw_rev_subminor));
DRV_LOG(DEBUG, "Packet pacing is not supported.");
mlx5_rt_timestamp_config(sh, hca_attr);
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -396,6 +406,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
priv->mp_id.port_id = port_id;
strlcpy(priv->mp_id.name, MLX5_MP_NAME, RTE_MP_MAX_NAME_LEN);
priv->representor = !!switch_info->representor;
@@ -612,17 +638,16 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_l3t_destroy(priv->mtr_profile_tbl);
if (own_domain_id)
claim_zero(rte_eth_switch_domain_free(priv->domain_id));
+ mlx5_free(priv->mac);
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
- if (eth_dev != NULL)
+ if (eth_dev != NULL) {
+ eth_dev->data->mac_addrs = NULL;
eth_dev->data->dev_private = NULL;
+ }
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
MLX5_ASSERT(err > 0);
@@ -698,7 +723,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
struct mlx5_priv *priv = dev->data->dev_private;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i))
rte_bitset_clear(priv->mac_own, i);
}
@@ -718,7 +743,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
struct mlx5_priv *priv = dev->data->dev_private;
- if (index < MLX5_MAX_MAC_ADDRESSES)
+ if (index < priv->sh->dev_cap.max_mac_addrs)
rte_bitset_clear(priv->mac_own, index);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH v7 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-14 14:42 ` [PATCH v7 5/5] net/mlx5: accept more unicast " David Marchand
@ 2026-09-14 14:51 ` Dariusz Sosnowski
2026-09-21 8:31 ` Raslan Darawsheh
1 sibling, 0 replies; 146+ messages in thread
From: Dariusz Sosnowski @ 2026-09-14 14:51 UTC (permalink / raw)
To: David Marchand
Cc: dev, rjarry, cfontain, Viacheslav Ovsiienko, Bing Zhao, Ori Kam,
Suanming Mou, Matan Azrad
On Mon, Sep 14, 2026 at 04:42:33PM +0200, David Marchand wrote:
> Starting firmware version 22.49.1014, the number of mac addresses
> per VF is not capped to 128 anymore.
>
> The value can be increased via devlink:
> $ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
> cmode driverinit
> $ devlink dev reload pci/0000:3b:00.2
>
> On the DPDK side, we must retrieve the maximum number of unicast
> and multicast addresses supported with a query to the firmware.
>
> Then, dynamically allocate the mac addresses arrays and report the
> limit instead of the previous hardcoded value.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-09-14 14:42 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (3 preceding siblings ...)
2026-09-14 14:42 ` [PATCH v7 5/5] net/mlx5: accept more unicast " David Marchand
@ 2026-09-21 8:31 ` Raslan Darawsheh
2026-09-21 10:31 ` David Marchand
4 siblings, 1 reply; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-21 8:31 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Hi David,
🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
No functional concerns with this patch in isolation. Flagging for context
going into the rest of the series: by the end of it, the fixed
MLX5_MAX_MAC_ADDRESSES / MLX5_MAX_UC_MAC_ADDRESSES bound is replaced by a
device-reported dynamic limit, but mlx5_flow_hw.c's HWS control-flow paths
are never updated to match it (see my reply on patch 5/5,
"net/mlx5: accept more unicast MAC addresses").
--
Raslan Darawsheh
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 2/5] net/mlx5: remove redundant MAC address index checks
2026-09-14 14:42 ` [PATCH v7 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-09-21 8:31 ` Raslan Darawsheh
2026-09-21 10:06 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-21 8:31 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Hi David,
🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
In mlx5_os_mac_addr_remove() (drivers/net/mlx5/linux/mlx5_os.c), this
patch drops both the netlink `index` arg and the Linux bounds guard:
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ BITFIELD_RESET(priv->mac_own, index);
but the Windows counterpart (windows/mlx5_os.c) still has an equivalent
check, so the two backends diverge in defensiveness after this patch.
The check is indeed redundant given mlx5_mac.c's index validation before
calling into the OS helper -- but for consistency the Windows-side check
should be dropped too, rather than left as the odd one out. Could you
remove it there as well in the next version?
--
Raslan Darawsheh
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 3/5] net/mlx5: pass maximum number of unicast MAC to common code
2026-09-14 14:42 ` [PATCH v7 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-09-21 8:31 ` Raslan Darawsheh
2026-09-21 10:07 ` David Marchand
0 siblings, 1 reply; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-21 8:31 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Hi David,
🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
In mlx5_nl_mac_addr_sync() (drivers/common/mlx5/linux/mlx5_nl.c), `n` is
now the caller-supplied array size rather than the old fixed
MLX5_MAX_MAC_ADDRESSES (256):
struct rte_ether_addr macs[n];
This VLA scales with whatever `n` the caller passes, and by the end of
the series that can be as large as the device-reported unicast MAC
capability (up to ~4096 per the release note added later in the series)
-- roughly 25KB on the stack versus ~1.5KB before this series.
mlx5_nl_mac_addr_sync() can be reached from interrupt/alarm-thread or
hotplug contexts where stack budgets are tighter than the main lcore
stack, so this can overflow the stack. Could you switch this to a heap
allocation (or a bounded scratch buffer) now that `n` is no longer a
small compile-time constant?
--
Raslan Darawsheh
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-14 14:42 ` [PATCH v7 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-14 14:51 ` Dariusz Sosnowski
@ 2026-09-21 8:31 ` Raslan Darawsheh
1 sibling, 0 replies; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-21 8:31 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Hi David,
🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
This is the patch that raises the effective unicast MAC limit
(priv->sh->dev_cap.max_mac_addrs, up to ~4096 per the FW capability), but
it misses two consumers in the HWS control-flow path that still use the
old fixed constants:
1. __flow_hw_ctrl_flows_unicast() / __flow_hw_ctrl_flows_unicast_vlan()
(drivers/net/mlx5/mlx5_flow_hw.c, ~line 16703 and ~16770) still loop
over indices 0..MLX5_MAX_MAC_ADDRESSES-1 (256) instead of
priv->sh->dev_cap.max_mac_addrs. On a device that now reports a larger
capability, a unicast MAC added at index >= 256 via
rte_eth_dev_mac_addr_add() succeeds at the mlx5_mac_addr_add() level,
but no HWS control-flow rule gets created for it -- traffic to that
MAC is silently not steered. The non-HWS path in mlx5_trigger.c was
updated in this same patch, so this looks like an oversight. This one
is the more important of the two to fix.
2. ctrl_rx_nb_flows_map[MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_DMAC]
(mlx5_flow_hw.c, ~line 11588) still sizes the DMAC control-flow
template table with the old fixed MLX5_MAX_UC_MAC_ADDRESSES (128).
With more than 128 unicast MACs configured on a capable device,
flow_hw_create_ctrl_flow() for the 129th+ MAC would fail even though
mlx5_mac_addr_add() reported success.
Could you address these in a v2?
--
Raslan Darawsheh
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 2/5] net/mlx5: remove redundant MAC address index checks
2026-09-21 8:31 ` Raslan Darawsheh
@ 2026-09-21 10:06 ` David Marchand
0 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-21 10:06 UTC (permalink / raw)
To: Raslan Darawsheh; +Cc: dev, Dariusz Sosnowski
Hello Raslan,
On Mon, 21 Sept 2026 at 10:32, Raslan Darawsheh <rasland@nvidia.com> wrote:
>
> Hi David,
>
> 🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
>
> In mlx5_os_mac_addr_remove() (drivers/net/mlx5/linux/mlx5_os.c), this
> patch drops both the netlink `index` arg and the Linux bounds guard:
>
> - if (index < MLX5_MAX_MAC_ADDRESSES)
> - BITFIELD_RESET(priv->mac_own, index);
> + BITFIELD_RESET(priv->mac_own, index);
>
> but the Windows counterpart (windows/mlx5_os.c) still has an equivalent
> check, so the two backends diverge in defensiveness after this patch.
> The check is indeed redundant given mlx5_mac.c's index validation before
> calling into the OS helper -- but for consistency the Windows-side check
> should be dropped too, rather than left as the odd one out. Could you
> remove it there as well in the next version?
Indeed, fixed.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 3/5] net/mlx5: pass maximum number of unicast MAC to common code
2026-09-21 8:31 ` Raslan Darawsheh
@ 2026-09-21 10:07 ` David Marchand
0 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-21 10:07 UTC (permalink / raw)
To: Raslan Darawsheh; +Cc: dev, Dariusz Sosnowski
On Mon, 21 Sept 2026 at 10:32, Raslan Darawsheh <rasland@nvidia.com> wrote:
> 🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
>
> In mlx5_nl_mac_addr_sync() (drivers/common/mlx5/linux/mlx5_nl.c), `n` is
> now the caller-supplied array size rather than the old fixed
> MLX5_MAX_MAC_ADDRESSES (256):
>
> struct rte_ether_addr macs[n];
>
> This VLA scales with whatever `n` the caller passes, and by the end of
> the series that can be as large as the device-reported unicast MAC
> capability (up to ~4096 per the release note added later in the series)
> -- roughly 25KB on the stack versus ~1.5KB before this series.
> mlx5_nl_mac_addr_sync() can be reached from interrupt/alarm-thread or
> hotplug contexts where stack budgets are tighter than the main lcore
> stack, so this can overflow the stack. Could you switch this to a heap
> allocation (or a bounded scratch buffer) now that `n` is no longer a
> small compile-time constant?
I don't mind switching to heap, this is control path with netlink
involved in any case.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-09-21 8:31 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux Raslan Darawsheh
@ 2026-09-21 10:31 ` David Marchand
2026-09-21 11:02 ` Raslan Darawsheh
0 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-21 10:31 UTC (permalink / raw)
To: Raslan Darawsheh; +Cc: dev, Dariusz Sosnowski
On Mon, 21 Sept 2026 at 10:32, Raslan Darawsheh <rasland@nvidia.com> wrote:
>
> Hi David,
>
> 🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
>
> No functional concerns with this patch in isolation. Flagging for context
> going into the rest of the series: by the end of it, the fixed
> MLX5_MAX_MAC_ADDRESSES / MLX5_MAX_UC_MAC_ADDRESSES bound is replaced by a
> device-reported dynamic limit, but mlx5_flow_hw.c's HWS control-flow paths
> are never updated to match it (see my reply on patch 5/5,
> "net/mlx5: accept more unicast MAC addresses").
While I understand the comment on the last patch, this comment here
seems gratuitous.. ?
I don't get what I should change at this point of the series.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-09-21 10:31 ` David Marchand
@ 2026-09-21 11:02 ` Raslan Darawsheh
0 siblings, 0 replies; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-21 11:02 UTC (permalink / raw)
To: David Marchand; +Cc: dev, Dariusz Sosnowski
Hi David,
On 21/09/2026 1:31 PM, David Marchand wrote:
> On Mon, 21 Sept 2026 at 10:32, Raslan Darawsheh <rasland@nvidia.com> wrote:
>>
>> Hi David,
>>
>> 🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
>>
>> No functional concerns with this patch in isolation. Flagging for context
>> going into the rest of the series: by the end of it, the fixed
>> MLX5_MAX_MAC_ADDRESSES / MLX5_MAX_UC_MAC_ADDRESSES bound is replaced by a
>> device-reported dynamic limit, but mlx5_flow_hw.c's HWS control-flow paths
>> are never updated to match it (see my reply on patch 5/5,
>> "net/mlx5: accept more unicast MAC addresses").
>
> While I understand the comment on the last patch, this comment here
> seems gratuitous.. ?
> I don't get what I should change at this point of the series.
>
>
Fair, that one was unnecessary on its own — the actual comment is on
patch 5/5. Apologies for the noise.
Kindest regards
Raslan Darawsheh
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v8 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (13 preceding siblings ...)
2026-09-14 14:42 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
@ 2026-09-21 11:50 ` David Marchand
2026-09-21 11:50 ` [PATCH v8 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
` (3 more replies)
2026-09-24 6:38 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
15 siblings, 4 replies; 146+ messages in thread
From: David Marchand @ 2026-09-21 11:50 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
This helper is exposing internals of net/mlx5 for no good reason.
All this code does is calling the remove helper.
Walk through the list in Linux implementation like the Windows
implementation.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 40 -----------------------------
drivers/common/mlx5/linux/mlx5_nl.h | 5 ----
drivers/net/mlx5/linux/mlx5_os.c | 13 +++++++---
3 files changed, 10 insertions(+), 48 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 42ccb73e36..40a8b8a2cc 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -825,46 +825,6 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
}
-/**
- * Flush all added MAC addresses.
- *
- * @param[in] nlsk_fd
- * Netlink socket file descriptor.
- * @param[in] iface_idx
- * Net device interface index.
- * @param[in] mac_addrs
- * Mac addresses array to flush.
- * @param n
- * @p mac_addrs array size.
- * @param mac_own
- * BITFIELD_DECLARE array to store the mac.
- * @param vf
- * Flag for a VF device.
- */
-RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_flush)
-void
-mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own, bool vf)
-{
- int i;
-
- if (n <= 0 || n > MLX5_MAX_MAC_ADDRESSES)
- return;
-
- for (i = n - 1; i >= 0; --i) {
- struct rte_ether_addr *m = &mac_addrs[i];
-
- if (BITFIELD_ISSET(mac_own, i)) {
- if (vf)
- mlx5_nl_mac_addr_remove(nlsk_fd,
- iface_idx,
- m, i);
- BITFIELD_RESET(mac_own, i);
- }
- }
-}
-
/**
* Enable promiscuous / all multicast mode through Netlink.
*
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 8ccdd244b0..0242342c47 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -65,11 +65,6 @@ __rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
__rte_internal
-void mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own,
- bool vf);
-__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
int mlx5_nl_allmulti(int nlsk_fd, unsigned int iface_idx, int enable);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 592e233844..49a8751761 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3529,10 +3529,17 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
{
struct mlx5_priv *priv = dev->data->dev_private;
const int vf = priv->sh->dev_cap.vf;
+ int i;
- mlx5_nl_mac_addr_flush(priv->nl_socket_route, mlx5_ifindex(dev),
- dev->data->mac_addrs,
- MLX5_MAX_MAC_ADDRESSES, priv->mac_own, vf);
+ for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (vf)
+ mlx5_nl_mac_addr_remove(priv->nl_socket_route,
+ mlx5_ifindex(dev),
+ &dev->data->mac_addrs[i], i);
+ BITFIELD_RESET(priv->mac_own, i);
+ }
+ }
}
static bool
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v8 2/5] net/mlx5: remove redundant MAC address index checks
2026-09-21 11:50 ` [PATCH v8 " David Marchand
@ 2026-09-21 11:50 ` David Marchand
2026-09-21 11:50 ` [PATCH v8 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
` (2 subsequent siblings)
3 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-21 11:50 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
On the net/mlx5 side, mlx5_mac.c validates that any MAC address and its
index is valid before calling the OS specific helpers.
So those OS helpers do not have to validate again the index.
Cascading this consideration, validating the MAC index against
MLX5_MAX_MAC_ADDRESSES in common code is also redundant.
The common code only deals with netlink, remove any index concern.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
Changes since v7:
- updated Windows implementation,
---
drivers/common/mlx5/linux/mlx5_nl.c | 17 ++---------------
drivers/common/mlx5/linux/mlx5_nl.h | 4 ++--
drivers/net/mlx5/linux/mlx5_os.c | 9 ++++-----
drivers/net/mlx5/windows/mlx5_os.c | 3 +--
4 files changed, 9 insertions(+), 24 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 40a8b8a2cc..3207eae563 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -719,8 +719,6 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to register.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -728,16 +726,11 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_add)
int
mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
int ret;
ret = mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 1);
- if (!ret) {
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
- }
if (ret == -EEXIST)
return 0;
return ret;
@@ -752,8 +745,6 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to remove.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -761,12 +752,8 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_remove)
int
mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
-
return mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 0);
}
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 0242342c47..0d6259f4ad 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -57,10 +57,10 @@ __rte_internal
int mlx5_nl_init(int protocol, int groups);
__rte_internal
int mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 49a8751761..c65293cb25 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3382,9 +3382,8 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[index], index);
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ &dev->data->mac_addrs[index]);
+ BITFIELD_RESET(priv->mac_own, index);
}
/**
@@ -3411,7 +3410,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
if (vf)
ret = mlx5_nl_mac_addr_add(priv->nl_socket_route,
mlx5_ifindex(dev),
- mac, index);
+ mac);
if (!ret)
BITFIELD_SET(priv->mac_own, index);
@@ -3536,7 +3535,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[i], i);
+ &dev->data->mac_addrs[i]);
BITFIELD_RESET(priv->mac_own, i);
}
}
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index cf34e4e1d6..ccc94d54af 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -718,8 +718,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
struct mlx5_priv *priv = dev->data->dev_private;
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ BITFIELD_RESET(priv->mac_own, index);
}
/**
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v8 3/5] net/mlx5: pass maximum number of unicast MAC to common code
2026-09-21 11:50 ` [PATCH v8 " David Marchand
2026-09-21 11:50 ` [PATCH v8 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-09-21 11:50 ` David Marchand
2026-09-21 11:50 ` [PATCH v8 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
2026-09-21 11:50 ` [PATCH v8 5/5] net/mlx5: accept more unicast " David Marchand
3 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-21 11:50 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Isolate how the MAC addresses array is walked through in the common code
by passing the max index at which a unicast MAC address is stored in
dev->data->mac_addrs[].
With this change, only net/mlx5 knows about the max number of
unicast/multicast MAC addresses.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
Changes since v7:
- switched to heap allocation for intermediate array,
Changes since v5:
- fixed mlx5_nl_mac_addr_cb,
---
drivers/common/mlx5/linux/mlx5_nl.c | 36 ++++++++++++++++++-----------
drivers/common/mlx5/linux/mlx5_nl.h | 2 +-
drivers/common/mlx5/mlx5_common.h | 8 -------
drivers/net/mlx5/linux/mlx5_os.c | 1 +
drivers/net/mlx5/mlx5.h | 8 +++++++
5 files changed, 33 insertions(+), 22 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 3207eae563..3fabcfaee0 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -171,9 +171,10 @@
/* Add/remove MAC address through Netlink */
struct mlx5_nl_mac_addr {
- struct rte_ether_addr (*mac)[];
+ struct rte_ether_addr **mac;
/**< MAC address handled by the device. */
int mac_n; /**< Number of addresses in the array. */
+ int max_macs; /**< Size of the array. */
};
static RTE_ATOMIC(uint32_t) atomic_sn;
@@ -475,7 +476,7 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
RTA_OK(attribute, len);
attribute = RTA_NEXT(attribute, len)) {
if (attribute->rta_type == NDA_LLADDR) {
- if (data->mac_n == MLX5_MAX_MAC_ADDRESSES) {
+ if (data->mac_n == data->max_macs) {
DRV_LOG(WARNING,
"not enough room to finalize the"
" request");
@@ -505,16 +506,16 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
* Net device interface index.
* @param mac[out]
* Pointer to the array table of MAC addresses to fill.
- * Its size should be of MLX5_MAX_MAC_ADDRESSES.
- * @param mac_n[out]
- * Number of entries filled in MAC array.
+ * @param mac_n[in,out]
+ * Size of the MAC array on input.
+ * Number of entries filled in MAC array on output.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
*/
static int
mlx5_nl_mac_addr_list(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr (*mac)[], int *mac_n)
+ struct rte_ether_addr **mac, int *mac_n)
{
struct {
struct nlmsghdr hdr;
@@ -533,6 +534,7 @@ mlx5_nl_mac_addr_list(int nlsk_fd, unsigned int iface_idx,
struct mlx5_nl_mac_addr data = {
.mac = mac,
.mac_n = 0,
+ .max_macs = *mac_n,
};
uint32_t sn = MLX5_NL_SN_GENERATE;
int ret;
@@ -766,23 +768,28 @@ mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac_addrs
* Mac addresses array to sync.
+ * @param uc_n
+ * Number of UC entries in @p mac_addrs.
* @param n
* @p mac_addrs array size.
*/
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_sync)
void
mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n)
+ struct rte_ether_addr *mac_addrs, int uc_n, int n)
{
- struct rte_ether_addr macs[n];
- int macs_n = 0;
+ struct rte_ether_addr *macs = NULL;
+ int macs_n = n;
int i;
int ret;
- memset(macs, 0, n * sizeof(macs[0]));
+ macs = calloc(n, sizeof(macs[0]));
+ if (macs == NULL)
+ goto out;
+
ret = mlx5_nl_mac_addr_list(nlsk_fd, iface_idx, &macs, &macs_n);
if (ret)
- return;
+ goto out;
for (i = 0; i != macs_n; ++i) {
int j;
@@ -794,7 +801,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
continue;
if (rte_is_multicast_ether_addr(&macs[i])) {
/* Find the first entry available. */
- for (j = MLX5_MAX_UC_MAC_ADDRESSES; j != n; ++j) {
+ for (j = uc_n; j != n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
@@ -802,7 +809,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
} else {
/* Find the first entry available. */
- for (j = 0; j != MLX5_MAX_UC_MAC_ADDRESSES; ++j) {
+ for (j = 0; j != uc_n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
@@ -810,6 +817,9 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
}
}
+
+out:
+ free(macs);
}
/**
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 0d6259f4ad..07a3b531b5 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -63,7 +63,7 @@ int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n);
+ struct rte_ether_addr *mac_addrs, int uc_n, int n);
__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 3767020823..8aed3dbabe 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -160,14 +160,6 @@ enum {
PCI_DEVICE_ID_MELLANOX_CONNECTX10C2C = 0x2101,
};
-/* Maximum number of simultaneous unicast MAC addresses. */
-#define MLX5_MAX_UC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous Multicast MAC addresses. */
-#define MLX5_MAX_MC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous MAC addresses. */
-#define MLX5_MAX_MAC_ADDRESSES \
- (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
-
/* Recognized Infiniband device physical port name types. */
enum mlx5_nl_phys_port_name_type {
MLX5_PHYS_PORT_NAME_TYPE_NOTSET = 0, /* Not set. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index c65293cb25..0e59d3f2f5 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -1762,6 +1762,7 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
+ MLX5_MAX_UC_MAC_ADDRESSES,
MLX5_MAX_MAC_ADDRESSES);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 190d203c49..c77cdc48a8 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -82,6 +82,14 @@
/* Maximum allowed MTU to be reported whenever PMD cannot query it from OS. */
#define MLX5_ETH_MAX_MTU (9978)
+/* Maximum number of simultaneous unicast MAC addresses. */
+#define MLX5_MAX_UC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous Multicast MAC addresses. */
+#define MLX5_MAX_MC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous MAC addresses. */
+#define MLX5_MAX_MAC_ADDRESSES \
+ (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
+
enum mlx5_ipool_index {
#if defined(HAVE_IBV_FLOW_DV_SUPPORT) || !defined(HAVE_INFINIBAND_VERBS_H)
MLX5_IPOOL_DECAP_ENCAP = 0, /* Pool for encap/decap resource. */
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v8 4/5] net/mlx5: use bitset for tracking MAC addresses
2026-09-21 11:50 ` [PATCH v8 " David Marchand
2026-09-21 11:50 ` [PATCH v8 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
2026-09-21 11:50 ` [PATCH v8 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-09-21 11:50 ` David Marchand
2026-09-21 11:50 ` [PATCH v8 5/5] net/mlx5: accept more unicast " David Marchand
3 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-21 11:50 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
EAL provides bitset that does the same as this set of mlx5 macros.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
drivers/common/mlx5/mlx5_common.h | 16 ----------------
drivers/net/mlx5/linux/mlx5_os.c | 8 ++++----
drivers/net/mlx5/mlx5.h | 3 ++-
drivers/net/mlx5/mlx5_trigger.c | 2 +-
drivers/net/mlx5/windows/mlx5_os.c | 8 ++++----
5 files changed, 11 insertions(+), 26 deletions(-)
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 8aed3dbabe..8b8f20e2f4 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -30,22 +30,6 @@
#define MLX5_PCI_DRIVER_NAME "mlx5_pci"
#define MLX5_AUXILIARY_DRIVER_NAME "mlx5_auxiliary"
-/* Bit-field manipulation. */
-#define BITFIELD_DECLARE(bf, type, size) \
- type bf[(((size_t)(size) / (sizeof(type) * CHAR_BIT)) + \
- !!((size_t)(size) % (sizeof(type) * CHAR_BIT)))]
-#define BITFIELD_DEFINE(bf, type, size) \
- BITFIELD_DECLARE((bf), type, (size)) = { 0 }
-#define BITFIELD_SET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] |= \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_RESET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] &= \
- ~((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_ISSET(bf, b) \
- !!(((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] & \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT)))))
-
/*
* Helper macros to work around __VA_ARGS__ limitations in a C99 compliant
* manner.
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 0e59d3f2f5..7d8ea4acba 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3384,7 +3384,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[index]);
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -3413,7 +3413,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
mlx5_ifindex(dev),
mac);
if (!ret)
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return ret;
}
@@ -3532,12 +3532,12 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[i]);
- BITFIELD_RESET(priv->mac_own, i);
+ rte_bitset_clear(priv->mac_own, i);
}
}
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index c77cdc48a8..7f20c811f3 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -14,6 +14,7 @@
#include <rte_pci.h>
#include <rte_ether.h>
+#include <rte_bitset.h>
#include <ethdev_driver.h>
#include <rte_rwlock.h>
#include <rte_interrupts.h>
@@ -2018,7 +2019,7 @@ struct mlx5_priv {
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- BITFIELD_DECLARE(mac_own, uint64_t, MLX5_MAX_MAC_ADDRESSES);
+ RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
/* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 2a91e02b45..7f6148f4e1 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1921,7 +1921,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
/* Add flows for unicast and multicast mac addresses added by API. */
if (!memcmp(mac, &cmp, sizeof(*mac)) ||
- !BITFIELD_ISSET(priv->mac_own, i) ||
+ !rte_bitset_test(priv->mac_own, i) ||
(dev->data->all_multicast && rte_is_multicast_ether_addr(mac)))
continue;
memcpy(&unicast.hdr.dst_addr.addr_bytes,
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index ccc94d54af..0aaed36ace 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -699,8 +699,8 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i))
- BITFIELD_RESET(priv->mac_own, i);
+ if (rte_bitset_test(priv->mac_own, i))
+ rte_bitset_clear(priv->mac_own, i);
}
}
@@ -718,7 +718,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
struct mlx5_priv *priv = dev->data->dev_private;
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -756,7 +756,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
return -ENOTSUP;
}
/* Mark this MAC address as owned by the PMD */
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return 0;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v8 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-21 11:50 ` [PATCH v8 " David Marchand
` (2 preceding siblings ...)
2026-09-21 11:50 ` [PATCH v8 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
@ 2026-09-21 11:50 ` David Marchand
2026-09-21 11:52 ` David Marchand
2026-09-23 11:55 ` Raslan Darawsheh
3 siblings, 2 replies; 146+ messages in thread
From: David Marchand @ 2026-09-21 11:50 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Starting firmware version 22.49.1014, the number of mac addresses
per VF is not capped to 128 anymore.
The value can be increased via devlink:
$ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
cmode driverinit
$ devlink dev reload pci/0000:3b:00.2
On the DPDK side, we must retrieve the maximum number of unicast
and multicast addresses supported with a query to the firmware.
Then, dynamically allocate the mac addresses arrays and report the
limit instead of the previous hardcoded value.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v7:
- updated HWS control path and removed MLX5_MAX_MAC_ADDRESSES constant,
Changes since v6:
- fixed crash in case of partial init failure in mlx5_dev_spawn,
and cleaned mlx5_os_pci_probe_pf,
- fixed compilation without assert in mlx5_internal_mac_addr_remove,
Changes since v4:
- added RN update,
- fixed types of fields added to mlx5_hca_attr and mlx5_dev_cap,
- used RTE_BIT32,
- fixed mlx5_nl_mac_addr_sync inverted arguments,
---
doc/guides/rel_notes/release_26_11.rst | 5 +++
drivers/common/mlx5/mlx5_devx_cmds.c | 4 +++
drivers/common/mlx5/mlx5_devx_cmds.h | 2 ++
drivers/net/mlx5/linux/mlx5_os.c | 47 +++++++++++++++++++-------
drivers/net/mlx5/mlx5.c | 9 ++---
drivers/net/mlx5/mlx5.h | 11 +++---
drivers/net/mlx5/mlx5_ethdev.c | 2 +-
drivers/net/mlx5/mlx5_flow_hw.c | 5 +--
drivers/net/mlx5/mlx5_mac.c | 20 ++++++-----
drivers/net/mlx5/mlx5_trigger.c | 10 +++---
drivers/net/mlx5/windows/mlx5_os.c | 41 +++++++++++++++++-----
11 files changed, 108 insertions(+), 48 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 87c7e81bde..43043cd579 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,11 @@ New Features
Also, make sure to start the actual text at the margin.
=======================================================
+* **Updated NVIDIA mlx5 ethernet driver.**
+
+ * Increased the maximum number of secondary unicast MAC addresses from 128 to up to 4096
+ (depending on devlink configuration on the associated kernel netdevice).
+
Removed Items
-------------
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.c b/drivers/common/mlx5/mlx5_devx_cmds.c
index 140b057ab4..e5d9c92779 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.c
+++ b/drivers/common/mlx5/mlx5_devx_cmds.c
@@ -1110,6 +1110,10 @@ mlx5_devx_cmd_query_hca_attr(void *ctx,
attr->log_max_pd = MLX5_GET(cmd_hca_cap, hcattr, log_max_pd);
attr->log_max_srq = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq);
attr->log_max_srq_sz = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq_sz);
+ attr->log_max_current_uc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_uc_list);
+ attr->log_max_current_mc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_mc_list);
attr->reg_c_preserve =
MLX5_GET(cmd_hca_cap, hcattr, reg_c_preserve);
attr->mmo_regex_qp_en = MLX5_GET(cmd_hca_cap, hcattr, regexp_mmo_qp);
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.h b/drivers/common/mlx5/mlx5_devx_cmds.h
index 90beb2e9e6..504b4a4f64 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.h
+++ b/drivers/common/mlx5/mlx5_devx_cmds.h
@@ -356,6 +356,8 @@ struct mlx5_hca_attr {
uint8_t tx_sw_owner_v2:1;
uint8_t esw_sw_owner:1;
uint8_t esw_sw_owner_v2:1;
+ uint8_t log_max_current_uc_list:5;
+ uint8_t log_max_current_mc_list:5;
};
/* LAG Context. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 7d8ea4acba..9180e9aa20 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -389,6 +389,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
sh->dev_cap.esw_info.regc_mask = 0;
#endif
sh->dev_cap.esw_info.is_set = 1;
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -1464,6 +1474,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
/* Some internal functions rely on Netlink sockets, open them now. */
priv->nl_socket_rdma = nl_rdma;
priv->nl_socket_route = mlx5_nl_init(NETLINK_ROUTE, 0);
@@ -1762,8 +1788,8 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
- MLX5_MAX_UC_MAC_ADDRESSES,
- MLX5_MAX_MAC_ADDRESSES);
+ sh->dev_cap.max_uc_mac_addrs,
+ sh->dev_cap.max_mac_addrs);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
TAILQ_INIT(&priv->flow_meters);
@@ -1963,17 +1989,16 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_flex_item_port_cleanup(eth_dev);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
- if (eth_dev != NULL)
+ if (eth_dev != NULL) {
+ eth_dev->data->mac_addrs = NULL;
eth_dev->data->dev_private = NULL;
+ }
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
if (nl_rdma >= 0)
@@ -2979,8 +3004,6 @@ mlx5_os_pci_probe_pf(struct mlx5_common_device *cdev,
if (!list[i].eth_dev)
continue;
mlx5_dev_close(list[i].eth_dev);
- /* mac_addrs must not be freed because in dev_private */
- list[i].eth_dev->data->mac_addrs = NULL;
claim_zero(rte_eth_dev_release_port(list[i].eth_dev));
}
/* Restore original error. */
@@ -3531,7 +3554,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
const int vf = priv->sh->dev_cap.vf;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
diff --git a/drivers/net/mlx5/mlx5.c b/drivers/net/mlx5/mlx5.c
index c7b0d3ef8b..4bd1c6d1c0 100644
--- a/drivers/net/mlx5/mlx5.c
+++ b/drivers/net/mlx5/mlx5.c
@@ -2556,6 +2556,9 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_list_destroy(priv->hrxqs);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
sh->port[priv->dev_port - 1].nl_ih_port_id = RTE_MAX_ETHPORTS;
/*
* The interrupt handler port id must be reset before priv is reset
@@ -2590,12 +2593,6 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_flow_pools_destroy(priv);
memset(priv, 0, sizeof(*priv));
priv->domain_id = RTE_ETH_DEV_SWITCH_DOMAIN_ID_INVALID;
- /*
- * Reset mac_addrs to NULL such that it is not freed as part of
- * rte_eth_dev_release_port(). mac_addrs is part of dev_private so
- * it is freed when dev_private is freed.
- */
- dev->data->mac_addrs = NULL;
return 0;
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 7f20c811f3..e3df458915 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -87,9 +87,6 @@
#define MLX5_MAX_UC_MAC_ADDRESSES 128
/* Maximum number of simultaneous Multicast MAC addresses. */
#define MLX5_MAX_MC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous MAC addresses. */
-#define MLX5_MAX_MAC_ADDRESSES \
- (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
enum mlx5_ipool_index {
#if defined(HAVE_IBV_FLOW_DV_SUPPORT) || !defined(HAVE_INFINIBAND_VERBS_H)
@@ -217,6 +214,9 @@ struct mlx5_dev_cap {
} mprq; /* Capability for Multi-Packet RQ. */
char fw_ver[64]; /* Firmware version of this device. */
struct flow_hw_port_info esw_info; /* E-switch manager reg_c0. */
+ uint32_t max_uc_mac_addrs; /* Maximum unicast MAC addresses. */
+ uint32_t max_mc_mac_addrs; /* Maximum multicast MAC addresses. */
+ uint32_t max_mac_addrs; /* Total maximum MAC addresses. */
};
#define MLX5_MPESW_PORT_INVALID (-1)
@@ -2018,9 +2018,8 @@ struct mlx5_priv {
struct mlx5_dev_ctx_shared *sh; /* Shared device context. */
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
- struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
- /* Bit-field of MAC addresses owned by the PMD. */
+ struct rte_ether_addr *mac; /* MAC addresses. */
+ uint64_t *mac_own; /* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
/* Device properties. */
diff --git a/drivers/net/mlx5/mlx5_ethdev.c b/drivers/net/mlx5/mlx5_ethdev.c
index 8160d10e7e..306c1cd734 100644
--- a/drivers/net/mlx5/mlx5_ethdev.c
+++ b/drivers/net/mlx5/mlx5_ethdev.c
@@ -392,7 +392,7 @@ mlx5_dev_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *info)
max = RTE_MIN(max, (unsigned int)UINT16_MAX);
info->max_rx_queues = max;
info->max_tx_queues = max;
- info->max_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ info->max_mac_addrs = priv->sh->dev_cap.max_uc_mac_addrs;
info->rx_queue_offload_capa = mlx5_get_rx_queue_offloads(dev);
info->rx_seg_capa.max_nseg = MLX5_MAX_RXQ_NSEG;
info->rx_seg_capa.multi_pools = !priv->config.mprq.enabled;
diff --git a/drivers/net/mlx5/mlx5_flow_hw.c b/drivers/net/mlx5/mlx5_flow_hw.c
index 30ef2f1c2d..887524da9c 100644
--- a/drivers/net/mlx5/mlx5_flow_hw.c
+++ b/drivers/net/mlx5/mlx5_flow_hw.c
@@ -16697,10 +16697,11 @@ __flow_hw_ctrl_flows_unicast(struct rte_eth_dev *dev,
struct rte_flow_template_table *tbl,
const enum mlx5_flow_ctrl_rx_expanded_rss_type rss_type)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
unsigned int i;
int ret;
- for (i = 0; i < MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i < priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -16767,7 +16768,7 @@ __flow_hw_ctrl_flows_unicast_vlan(struct rte_eth_dev *dev,
unsigned int i;
unsigned int j;
- for (i = 0; i < MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i < priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
diff --git a/drivers/net/mlx5/mlx5_mac.c b/drivers/net/mlx5/mlx5_mac.c
index 0e5d2be530..9e7bf7f2ec 100644
--- a/drivers/net/mlx5/mlx5_mac.c
+++ b/drivers/net/mlx5/mlx5_mac.c
@@ -36,7 +36,7 @@ mlx5_internal_mac_addr_remove(struct rte_eth_dev *dev,
uint32_t index,
struct rte_ether_addr *addr)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < MLX5_SH(dev)->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(&dev->data->mac_addrs[index]))
return false;
mlx5_os_mac_addr_remove(dev, index);
@@ -63,16 +63,17 @@ static int
mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
unsigned int i;
int ret;
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(mac)) {
rte_errno = EINVAL;
return -rte_errno;
}
/* First, make sure this address isn't already configured. */
- for (i = 0; (i != MLX5_MAX_MAC_ADDRESSES); ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
/* Skip this index, it's going to be reconfigured. */
if (i == index)
continue;
@@ -101,10 +102,11 @@ mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
void
mlx5_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
struct rte_ether_addr addr = { 0 };
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES)
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs)
return;
if (mlx5_internal_mac_addr_remove(dev, index, &addr)) {
ret = mlx5_traffic_mac_remove(dev, &addr);
@@ -133,9 +135,10 @@ int
mlx5_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index, uint32_t vmdq __rte_unused)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES) {
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs) {
rte_errno = EINVAL;
return -rte_errno;
}
@@ -217,16 +220,17 @@ int
mlx5_set_mc_addr_list(struct rte_eth_dev *dev,
struct rte_ether_addr *mc_addr_set, uint32_t nb_mc_addr)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
uint32_t i;
int ret;
- if (nb_mc_addr >= MLX5_MAX_MC_MAC_ADDRESSES) {
+ if (nb_mc_addr >= priv->sh->dev_cap.max_mc_mac_addrs) {
rte_errno = ENOSPC;
return -rte_errno;
}
- for (i = MLX5_MAX_UC_MAC_ADDRESSES; i != MLX5_MAX_MAC_ADDRESSES; ++i)
+ for (i = priv->sh->dev_cap.max_uc_mac_addrs; i != priv->sh->dev_cap.max_mac_addrs; ++i)
mlx5_internal_mac_addr_remove(dev, i, NULL);
- i = MLX5_MAX_UC_MAC_ADDRESSES;
+ i = priv->sh->dev_cap.max_uc_mac_addrs;
while (nb_mc_addr--) {
ret = mlx5_internal_mac_addr_add(dev, mc_addr_set++, i++);
if (ret)
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 7f6148f4e1..c5f493117b 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1916,7 +1916,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
}
}
/* Add MAC address flows. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
/* Add flows for unicast and multicast mac addresses added by API. */
@@ -2186,7 +2186,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
return 0;
/* Add all unicast DMAC flow rules with new VLAN attached. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2203,7 +2203,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
* Removing after creating VLAN rules so that traffic "gap" is not introduced.
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2241,7 +2241,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
* Recreating first to ensure no traffic "gap".
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2254,7 +2254,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
}
/* Remove all unicast DMAC flow rules with this VLAN. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index 0aaed36ace..cc60f07e17 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -261,6 +261,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
MLX5_GET(initial_seg, pv_iseg, fw_rev_subminor));
DRV_LOG(DEBUG, "Packet pacing is not supported.");
mlx5_rt_timestamp_config(sh, hca_attr);
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -396,6 +406,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
priv->mp_id.port_id = port_id;
strlcpy(priv->mp_id.name, MLX5_MP_NAME, RTE_MP_MAX_NAME_LEN);
priv->representor = !!switch_info->representor;
@@ -612,17 +638,16 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_l3t_destroy(priv->mtr_profile_tbl);
if (own_domain_id)
claim_zero(rte_eth_switch_domain_free(priv->domain_id));
+ mlx5_free(priv->mac);
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
- if (eth_dev != NULL)
+ if (eth_dev != NULL) {
+ eth_dev->data->mac_addrs = NULL;
eth_dev->data->dev_private = NULL;
+ }
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
MLX5_ASSERT(err > 0);
@@ -698,7 +723,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
struct mlx5_priv *priv = dev->data->dev_private;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i))
rte_bitset_clear(priv->mac_own, i);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH v8 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-21 11:50 ` [PATCH v8 5/5] net/mlx5: accept more unicast " David Marchand
@ 2026-09-21 11:52 ` David Marchand
2026-09-23 11:55 ` Raslan Darawsheh
1 sibling, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-21 11:52 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
On Mon, 21 Sept 2026 at 13:51, David Marchand <david.marchand@redhat.com> wrote:
>
> Starting firmware version 22.49.1014, the number of mac addresses
> per VF is not capped to 128 anymore.
>
> The value can be increased via devlink:
> $ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
> cmode driverinit
> $ devlink dev reload pci/0000:3b:00.2
>
> On the DPDK side, we must retrieve the maximum number of unicast
> and multicast addresses supported with a query to the firmware.
>
> Then, dynamically allocate the mac addresses arrays and report the
> limit instead of the previous hardcoded value.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Sorry, I forgot to add:
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset
2026-09-14 8:17 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset David Marchand
` (3 preceding siblings ...)
2026-09-14 10:15 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset Loftus, Ciara
@ 2026-09-23 11:47 ` Burakov, Anatoly
2026-09-23 14:37 ` Bruce Richardson
4 siblings, 1 reply; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-23 11:47 UTC (permalink / raw)
To: David Marchand, dev
Cc: ciara.loftus, rjarry, cfontain, stable, Vladimir Medvedkin,
Lunyuan Cui, Jingjing Wu, Xiaolong Ye
On 9/14/2026 10:17 AM, David Marchand wrote:
> When resetting, calling iavf_dev_uninit + iavf_dev_init results in
> leaking the previous dev->data->mac_addrs array.
> As a consequence, secondary MAC addresses are lost during a VF reset.
>
> Move the MAC addresses array in the private iavf_info structure, along
> the multicast MAC addresses.
> Set/clear dev->data->mac_addrs in iavf_dev_init/iavf_dev_uninit.
>
> Consistently use RTE_DIM() to avoid mixing with the multicast addresses
> array.
>
> Fixes: e74e1bb6280d ("net/iavf: enable port reset")
> Cc: stable@dpdk.org
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
Acked-by: Anatoly Burakov <anatoly.burakov@intel.com>
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install
2026-09-14 8:17 ` [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install David Marchand
2026-09-14 10:19 ` Loftus, Ciara
@ 2026-09-23 11:53 ` Burakov, Anatoly
1 sibling, 0 replies; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-23 11:53 UTC (permalink / raw)
To: David Marchand, dev
Cc: ciara.loftus, rjarry, cfontain, stable, Vladimir Medvedkin,
Bruce Richardson
On 9/14/2026 10:17 AM, David Marchand wrote:
> On port restart, all MAC addresses get pushed *twice* to the hardware,
> once by the driver and once by the eth_dev_mac_restore() in ethdev.
>
> On the other hand, MAC address filters are reset in the hardware
> by the PF only when a VF reset is triggered.
>
> Strictly speaking, the mac restore on port (re)start is unneeded,
> if no VF reset happened, so we can announce to ethdev that no mac
> restoration is needed via a get_restore_flags callback.
>
> Then, move the mac restoration to the VF reset handler.
>
> Fixes: 3d42086def30 ("net/iavf: preserve MAC address with i40e PF Linux driver")
> Cc: stable@dpdk.org
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
(with Ciara's suggested fix)
Acked-by: Anatoly Burakov <anatoly.burakov@intel.com>
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v8 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-21 11:50 ` [PATCH v8 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-21 11:52 ` David Marchand
@ 2026-09-23 11:55 ` Raslan Darawsheh
1 sibling, 0 replies; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-23 11:55 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Hi David,
🤖 This review was drafted with assistance from Claude (Anthropic) and reviewed by me before posting.
Thanks for fixing __flow_hw_ctrl_flows_unicast() / _vlan() to loop over
priv->sh->dev_cap.max_mac_addrs.
One more spot to check: ctrl_rx_nb_flows_map[MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_DMAC]
(mlx5_flow_hw.c, ~line 11588) is a static initializer still set to the
fixed MLX5_MAX_UC_MAC_ADDRESSES (128, unchanged by this series). That
value becomes nb_flows -> cfg.max_idx (mlx5_flow_hw.c, ~line 5335), which
hard-caps the ipool backing the DMAC/DMAC_VLAN control-flow template
table, and this table isn't on the resizable-table path.
So __flow_hw_ctrl_flows_unicast() now loops up to
priv->sh->dev_cap.max_mac_addrs (which can be up to ~4096 per this series)
and tries to insert one control-flow rule per configured unicast MAC into
that same 128-capacity table. On any device that now advertises more than
128 unicast MACs -- which is the whole point of this series -- configuring
more than 128 will make flow-rule insertion fail for the 129th+ MAC, even
though mlx5_mac_addr_add() itself succeeded. Rx breaks silently for those
addresses.
Could you take a look at this one too?
Thanks,
Raslan
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 3/4] net/iavf: add a helper for sending MAC addresses to PF
2026-09-14 8:17 ` [PATCH v7 3/4] net/iavf: add a helper for sending MAC addresses to PF David Marchand
@ 2026-09-23 12:04 ` Burakov, Anatoly
0 siblings, 0 replies; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-23 12:04 UTC (permalink / raw)
To: David Marchand, dev; +Cc: ciara.loftus, rjarry, cfontain, Vladimir Medvedkin
On 9/14/2026 10:17 AM, David Marchand wrote:
> Rather than have multiple implementations of the same code,
> define a single helper.
>
> This is also a good place where to check if the driver is requesting too
> many addresses in a single message to the PF.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
Acked-by: Anatoly Burakov <anatoly.burakov@intel.com>
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 4/4] net/iavf: accept up to 32k unicast MAC addresses
2026-09-14 8:17 ` [PATCH v7 4/4] net/iavf: accept up to 32k unicast MAC addresses David Marchand
@ 2026-09-23 12:14 ` Burakov, Anatoly
0 siblings, 0 replies; 146+ messages in thread
From: Burakov, Anatoly @ 2026-09-23 12:14 UTC (permalink / raw)
To: David Marchand, dev; +Cc: ciara.loftus, rjarry, cfontain, Vladimir Medvedkin
On 9/14/2026 10:17 AM, David Marchand wrote:
> E810 hardware provides 32k switch lookups.
> Thanks to this, it is possible to allow a lot more secondary mac
> addresses than what is possible today.
>
> In practice, the maximum number of macs available per port may be lower
> and depends on usage by other (trusted?) VFs on the same PF.
> There is no way to figure out this limit but to try adding a mac address
> and get an error from the PF driver.
>
> Mailbox exchanges are limited to IAVF_AQ_BUF_SZ, segment messages
> accordingly.
>
> Since unicast and multicast addresses arrays are sized with two
> different constants, prefer RTE_DIM() whenever possible.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
> ---
> Changes since v6:
> - reused helper added in previous commit,
> - used RTE_DIM() instead of macro constants,
>
> Changes since v5:
> - separated from series that went in next-net,
> - rebased,
>
> Changes since v4:
> - rebased,
>
> Changes since v2:
> - added an entry in release notes,
> - removed unneeded temp variable,
>
> Changes since v1:
> - fixed buffer overflow on mailbox messages during port restart/VF reset,
>
> ---
<snip>
>
> diff --git a/drivers/net/intel/iavf/iavf_vchnl.c b/drivers/net/intel/iavf/iavf_vchnl.c
> index decfae3182..418a7e897e 100644
> --- a/drivers/net/intel/iavf/iavf_vchnl.c
> +++ b/drivers/net/intel/iavf/iavf_vchnl.c
> @@ -1712,8 +1712,8 @@ iavf_send_eth_addr_list(struct iavf_adapter *adapter, const char *caller,
> void
> iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool add)
> {
> + uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(IAVF_ETH_ADDR_PER_REQ)] = {0};
> struct iavf_info *vf = IAVF_DEV_PRIVATE_TO_VF(adapter);
> - uint8_t cmd_buffer[IAVF_ETH_ADDR_CMD_SIZE(RTE_DIM(vf->mac_addrs))] = {0};
> struct virtchnl_ether_addr_list *list;
>
> list = (struct virtchnl_ether_addr_list *)cmd_buffer;
> @@ -1730,6 +1730,12 @@ iavf_add_del_secondary_mac_addr(struct iavf_adapter *adapter, bool add)
> memcpy(vc_addr->addr, addr->addr_bytes, sizeof(addr->addr_bytes));
> vc_addr->type = VIRTCHNL_ETHER_ADDR_EXTRA;
> }
> +
> + if (list->num_elements == IAVF_ETH_ADDR_PER_REQ) {
> + if (iavf_send_eth_addr_list(adapter, __func__, list, add))
> + return;
> + list->num_elements = 0;
> + }
> }
Nitpick over my previous comment, but I really don't understand why
resetting a list in the middle of a loop is 1) not a nested loop,
cognitively speaking, and 2) more clear than just having every inner
loop start at 0 and end at ETH_ADDR_PER_REQ while having an outer loop
go from zero until RTE_DIM(vf->mac_addrs) which is what ends up
happening when this problem is modeled anyway. This to me reads like an
attempt at avoiding having a nested loop by breaking the loop up in the
middle, all for the sake of not having nested loops.
However, it works and not worth a respin, so
Acked-by: Anatoly Burakov <anatoly.burakov@intel.com>
--
Thanks,
Anatoly
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset
2026-09-23 11:47 ` Burakov, Anatoly
@ 2026-09-23 14:37 ` Bruce Richardson
0 siblings, 0 replies; 146+ messages in thread
From: Bruce Richardson @ 2026-09-23 14:37 UTC (permalink / raw)
To: Burakov, Anatoly
Cc: David Marchand, dev, ciara.loftus, rjarry, cfontain, stable,
Vladimir Medvedkin, Lunyuan Cui, Jingjing Wu, Xiaolong Ye
On Wed, Sep 23, 2026 at 01:47:25PM +0200, Burakov, Anatoly wrote:
> On 9/14/2026 10:17 AM, David Marchand wrote:
> > When resetting, calling iavf_dev_uninit + iavf_dev_init results in
> > leaking the previous dev->data->mac_addrs array.
> > As a consequence, secondary MAC addresses are lost during a VF reset.
> >
> > Move the MAC addresses array in the private iavf_info structure, along
> > the multicast MAC addresses.
> > Set/clear dev->data->mac_addrs in iavf_dev_init/iavf_dev_uninit.
> >
> > Consistently use RTE_DIM() to avoid mixing with the multicast addresses
> > array.
> >
> > Fixes: e74e1bb6280d ("net/iavf: enable port reset")
> > Cc: stable@dpdk.org
> >
> > Signed-off-by: David Marchand <david.marchand@redhat.com>
> > ---
>
> Acked-by: Anatoly Burakov <anatoly.burakov@intel.com>
>
Series appled to dpdk-next-net-intel.
Thanks,
/Bruce
^ permalink raw reply [flat|nested] 146+ messages in thread
* [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
` (14 preceding siblings ...)
2026-09-21 11:50 ` [PATCH v8 " David Marchand
@ 2026-09-24 6:38 ` David Marchand
2026-09-24 6:38 ` [PATCH v9 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
` (4 more replies)
15 siblings, 5 replies; 146+ messages in thread
From: David Marchand @ 2026-09-24 6:38 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
This helper is exposing internals of net/mlx5 for no good reason.
All this code does is calling the remove helper.
Walk through the list in Linux implementation like the Windows
implementation.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
drivers/common/mlx5/linux/mlx5_nl.c | 40 -----------------------------
drivers/common/mlx5/linux/mlx5_nl.h | 5 ----
drivers/net/mlx5/linux/mlx5_os.c | 13 +++++++---
3 files changed, 10 insertions(+), 48 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 42ccb73e36..40a8b8a2cc 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -825,46 +825,6 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
}
-/**
- * Flush all added MAC addresses.
- *
- * @param[in] nlsk_fd
- * Netlink socket file descriptor.
- * @param[in] iface_idx
- * Net device interface index.
- * @param[in] mac_addrs
- * Mac addresses array to flush.
- * @param n
- * @p mac_addrs array size.
- * @param mac_own
- * BITFIELD_DECLARE array to store the mac.
- * @param vf
- * Flag for a VF device.
- */
-RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_flush)
-void
-mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own, bool vf)
-{
- int i;
-
- if (n <= 0 || n > MLX5_MAX_MAC_ADDRESSES)
- return;
-
- for (i = n - 1; i >= 0; --i) {
- struct rte_ether_addr *m = &mac_addrs[i];
-
- if (BITFIELD_ISSET(mac_own, i)) {
- if (vf)
- mlx5_nl_mac_addr_remove(nlsk_fd,
- iface_idx,
- m, i);
- BITFIELD_RESET(mac_own, i);
- }
- }
-}
-
/**
* Enable promiscuous / all multicast mode through Netlink.
*
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 8ccdd244b0..0242342c47 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -65,11 +65,6 @@ __rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
__rte_internal
-void mlx5_nl_mac_addr_flush(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n,
- uint64_t *mac_own,
- bool vf);
-__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
int mlx5_nl_allmulti(int nlsk_fd, unsigned int iface_idx, int enable);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 592e233844..49a8751761 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3529,10 +3529,17 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
{
struct mlx5_priv *priv = dev->data->dev_private;
const int vf = priv->sh->dev_cap.vf;
+ int i;
- mlx5_nl_mac_addr_flush(priv->nl_socket_route, mlx5_ifindex(dev),
- dev->data->mac_addrs,
- MLX5_MAX_MAC_ADDRESSES, priv->mac_own, vf);
+ for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (vf)
+ mlx5_nl_mac_addr_remove(priv->nl_socket_route,
+ mlx5_ifindex(dev),
+ &dev->data->mac_addrs[i], i);
+ BITFIELD_RESET(priv->mac_own, i);
+ }
+ }
}
static bool
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v9 2/5] net/mlx5: remove redundant MAC address index checks
2026-09-24 6:38 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
@ 2026-09-24 6:38 ` David Marchand
2026-09-29 12:50 ` Raslan Darawsheh
2026-09-24 6:38 ` [PATCH v9 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
` (3 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-24 6:38 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
On the net/mlx5 side, mlx5_mac.c validates that any MAC address and its
index is valid before calling the OS specific helpers.
So those OS helpers do not have to validate again the index.
Cascading this consideration, validating the MAC index against
MLX5_MAX_MAC_ADDRESSES in common code is also redundant.
The common code only deals with netlink, remove any index concern.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
Changes since v7:
- updated Windows implementation,
---
drivers/common/mlx5/linux/mlx5_nl.c | 17 ++---------------
drivers/common/mlx5/linux/mlx5_nl.h | 4 ++--
drivers/net/mlx5/linux/mlx5_os.c | 9 ++++-----
drivers/net/mlx5/windows/mlx5_os.c | 3 +--
4 files changed, 9 insertions(+), 24 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 40a8b8a2cc..3207eae563 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -719,8 +719,6 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to register.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -728,16 +726,11 @@ mlx5_nl_vf_mac_addr_modify(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_add)
int
mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
int ret;
ret = mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 1);
- if (!ret) {
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
- }
if (ret == -EEXIST)
return 0;
return ret;
@@ -752,8 +745,6 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac
* MAC address to remove.
- * @param index
- * MAC address index.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
@@ -761,12 +752,8 @@ mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_remove)
int
mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index)
+ struct rte_ether_addr *mac)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
- if (index >= MLX5_MAX_MAC_ADDRESSES)
- return -EINVAL;
-
return mlx5_nl_mac_addr_modify(nlsk_fd, iface_idx, mac, 0);
}
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 0242342c47..0d6259f4ad 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -57,10 +57,10 @@ __rte_internal
int mlx5_nl_init(int protocol, int groups);
__rte_internal
int mlx5_nl_mac_addr_add(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac, uint32_t index);
+ struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac_addrs, int n);
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 49a8751761..c65293cb25 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3382,9 +3382,8 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[index], index);
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ &dev->data->mac_addrs[index]);
+ BITFIELD_RESET(priv->mac_own, index);
}
/**
@@ -3411,7 +3410,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
if (vf)
ret = mlx5_nl_mac_addr_add(priv->nl_socket_route,
mlx5_ifindex(dev),
- mac, index);
+ mac);
if (!ret)
BITFIELD_SET(priv->mac_own, index);
@@ -3536,7 +3535,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
- &dev->data->mac_addrs[i], i);
+ &dev->data->mac_addrs[i]);
BITFIELD_RESET(priv->mac_own, i);
}
}
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index cf34e4e1d6..ccc94d54af 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -718,8 +718,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
struct mlx5_priv *priv = dev->data->dev_private;
- if (index < MLX5_MAX_MAC_ADDRESSES)
- BITFIELD_RESET(priv->mac_own, index);
+ BITFIELD_RESET(priv->mac_own, index);
}
/**
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v9 3/5] net/mlx5: pass maximum number of unicast MAC to common code
2026-09-24 6:38 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-24 6:38 ` [PATCH v9 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-09-24 6:38 ` David Marchand
2026-09-29 12:50 ` Raslan Darawsheh
2026-09-24 6:38 ` [PATCH v9 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
` (2 subsequent siblings)
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-24 6:38 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Isolate how the MAC addresses array is walked through in the common code
by passing the max index at which a unicast MAC address is stored in
dev->data->mac_addrs[].
With this change, only net/mlx5 knows about the max number of
unicast/multicast MAC addresses.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
Changes since v7:
- switched to heap allocation for intermediate array,
Changes since v5:
- fixed mlx5_nl_mac_addr_cb,
---
drivers/common/mlx5/linux/mlx5_nl.c | 36 ++++++++++++++++++-----------
drivers/common/mlx5/linux/mlx5_nl.h | 2 +-
drivers/common/mlx5/mlx5_common.h | 8 -------
drivers/net/mlx5/linux/mlx5_os.c | 1 +
drivers/net/mlx5/mlx5.h | 8 +++++++
5 files changed, 33 insertions(+), 22 deletions(-)
diff --git a/drivers/common/mlx5/linux/mlx5_nl.c b/drivers/common/mlx5/linux/mlx5_nl.c
index 3207eae563..3fabcfaee0 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.c
+++ b/drivers/common/mlx5/linux/mlx5_nl.c
@@ -171,9 +171,10 @@
/* Add/remove MAC address through Netlink */
struct mlx5_nl_mac_addr {
- struct rte_ether_addr (*mac)[];
+ struct rte_ether_addr **mac;
/**< MAC address handled by the device. */
int mac_n; /**< Number of addresses in the array. */
+ int max_macs; /**< Size of the array. */
};
static RTE_ATOMIC(uint32_t) atomic_sn;
@@ -475,7 +476,7 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
RTA_OK(attribute, len);
attribute = RTA_NEXT(attribute, len)) {
if (attribute->rta_type == NDA_LLADDR) {
- if (data->mac_n == MLX5_MAX_MAC_ADDRESSES) {
+ if (data->mac_n == data->max_macs) {
DRV_LOG(WARNING,
"not enough room to finalize the"
" request");
@@ -505,16 +506,16 @@ mlx5_nl_mac_addr_cb(struct nlmsghdr *nh, void *arg)
* Net device interface index.
* @param mac[out]
* Pointer to the array table of MAC addresses to fill.
- * Its size should be of MLX5_MAX_MAC_ADDRESSES.
- * @param mac_n[out]
- * Number of entries filled in MAC array.
+ * @param mac_n[in,out]
+ * Size of the MAC array on input.
+ * Number of entries filled in MAC array on output.
*
* @return
* 0 on success, a negative errno value otherwise and rte_errno is set.
*/
static int
mlx5_nl_mac_addr_list(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr (*mac)[], int *mac_n)
+ struct rte_ether_addr **mac, int *mac_n)
{
struct {
struct nlmsghdr hdr;
@@ -533,6 +534,7 @@ mlx5_nl_mac_addr_list(int nlsk_fd, unsigned int iface_idx,
struct mlx5_nl_mac_addr data = {
.mac = mac,
.mac_n = 0,
+ .max_macs = *mac_n,
};
uint32_t sn = MLX5_NL_SN_GENERATE;
int ret;
@@ -766,23 +768,28 @@ mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
* Net device interface index.
* @param mac_addrs
* Mac addresses array to sync.
+ * @param uc_n
+ * Number of UC entries in @p mac_addrs.
* @param n
* @p mac_addrs array size.
*/
RTE_EXPORT_INTERNAL_SYMBOL(mlx5_nl_mac_addr_sync)
void
mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n)
+ struct rte_ether_addr *mac_addrs, int uc_n, int n)
{
- struct rte_ether_addr macs[n];
- int macs_n = 0;
+ struct rte_ether_addr *macs = NULL;
+ int macs_n = n;
int i;
int ret;
- memset(macs, 0, n * sizeof(macs[0]));
+ macs = calloc(n, sizeof(macs[0]));
+ if (macs == NULL)
+ goto out;
+
ret = mlx5_nl_mac_addr_list(nlsk_fd, iface_idx, &macs, &macs_n);
if (ret)
- return;
+ goto out;
for (i = 0; i != macs_n; ++i) {
int j;
@@ -794,7 +801,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
continue;
if (rte_is_multicast_ether_addr(&macs[i])) {
/* Find the first entry available. */
- for (j = MLX5_MAX_UC_MAC_ADDRESSES; j != n; ++j) {
+ for (j = uc_n; j != n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
@@ -802,7 +809,7 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
} else {
/* Find the first entry available. */
- for (j = 0; j != MLX5_MAX_UC_MAC_ADDRESSES; ++j) {
+ for (j = 0; j != uc_n; ++j) {
if (rte_is_zero_ether_addr(&mac_addrs[j])) {
mac_addrs[j] = macs[i];
break;
@@ -810,6 +817,9 @@ mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
}
}
}
+
+out:
+ free(macs);
}
/**
diff --git a/drivers/common/mlx5/linux/mlx5_nl.h b/drivers/common/mlx5/linux/mlx5_nl.h
index 0d6259f4ad..07a3b531b5 100644
--- a/drivers/common/mlx5/linux/mlx5_nl.h
+++ b/drivers/common/mlx5/linux/mlx5_nl.h
@@ -63,7 +63,7 @@ int mlx5_nl_mac_addr_remove(int nlsk_fd, unsigned int iface_idx,
struct rte_ether_addr *mac);
__rte_internal
void mlx5_nl_mac_addr_sync(int nlsk_fd, unsigned int iface_idx,
- struct rte_ether_addr *mac_addrs, int n);
+ struct rte_ether_addr *mac_addrs, int uc_n, int n);
__rte_internal
int mlx5_nl_promisc(int nlsk_fd, unsigned int iface_idx, int enable);
__rte_internal
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 3767020823..8aed3dbabe 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -160,14 +160,6 @@ enum {
PCI_DEVICE_ID_MELLANOX_CONNECTX10C2C = 0x2101,
};
-/* Maximum number of simultaneous unicast MAC addresses. */
-#define MLX5_MAX_UC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous Multicast MAC addresses. */
-#define MLX5_MAX_MC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous MAC addresses. */
-#define MLX5_MAX_MAC_ADDRESSES \
- (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
-
/* Recognized Infiniband device physical port name types. */
enum mlx5_nl_phys_port_name_type {
MLX5_PHYS_PORT_NAME_TYPE_NOTSET = 0, /* Not set. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index c65293cb25..0e59d3f2f5 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -1762,6 +1762,7 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
+ MLX5_MAX_UC_MAC_ADDRESSES,
MLX5_MAX_MAC_ADDRESSES);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 190d203c49..c77cdc48a8 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -82,6 +82,14 @@
/* Maximum allowed MTU to be reported whenever PMD cannot query it from OS. */
#define MLX5_ETH_MAX_MTU (9978)
+/* Maximum number of simultaneous unicast MAC addresses. */
+#define MLX5_MAX_UC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous Multicast MAC addresses. */
+#define MLX5_MAX_MC_MAC_ADDRESSES 128
+/* Maximum number of simultaneous MAC addresses. */
+#define MLX5_MAX_MAC_ADDRESSES \
+ (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
+
enum mlx5_ipool_index {
#if defined(HAVE_IBV_FLOW_DV_SUPPORT) || !defined(HAVE_INFINIBAND_VERBS_H)
MLX5_IPOOL_DECAP_ENCAP = 0, /* Pool for encap/decap resource. */
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v9 4/5] net/mlx5: use bitset for tracking MAC addresses
2026-09-24 6:38 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-24 6:38 ` [PATCH v9 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
2026-09-24 6:38 ` [PATCH v9 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-09-24 6:38 ` David Marchand
2026-09-29 12:50 ` Raslan Darawsheh
2026-09-24 6:38 ` [PATCH v9 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-29 12:50 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux Raslan Darawsheh
4 siblings, 1 reply; 146+ messages in thread
From: David Marchand @ 2026-09-24 6:38 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
EAL provides bitset that does the same as this set of mlx5 macros.
Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
---
drivers/common/mlx5/mlx5_common.h | 16 ----------------
drivers/net/mlx5/linux/mlx5_os.c | 8 ++++----
drivers/net/mlx5/mlx5.h | 3 ++-
drivers/net/mlx5/mlx5_trigger.c | 2 +-
drivers/net/mlx5/windows/mlx5_os.c | 8 ++++----
5 files changed, 11 insertions(+), 26 deletions(-)
diff --git a/drivers/common/mlx5/mlx5_common.h b/drivers/common/mlx5/mlx5_common.h
index 8aed3dbabe..8b8f20e2f4 100644
--- a/drivers/common/mlx5/mlx5_common.h
+++ b/drivers/common/mlx5/mlx5_common.h
@@ -30,22 +30,6 @@
#define MLX5_PCI_DRIVER_NAME "mlx5_pci"
#define MLX5_AUXILIARY_DRIVER_NAME "mlx5_auxiliary"
-/* Bit-field manipulation. */
-#define BITFIELD_DECLARE(bf, type, size) \
- type bf[(((size_t)(size) / (sizeof(type) * CHAR_BIT)) + \
- !!((size_t)(size) % (sizeof(type) * CHAR_BIT)))]
-#define BITFIELD_DEFINE(bf, type, size) \
- BITFIELD_DECLARE((bf), type, (size)) = { 0 }
-#define BITFIELD_SET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] |= \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_RESET(bf, b) \
- (void)((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] &= \
- ~((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT))))
-#define BITFIELD_ISSET(bf, b) \
- !!(((bf)[((b) / (sizeof((bf)[0]) * CHAR_BIT))] & \
- ((size_t)1 << ((b) % (sizeof((bf)[0]) * CHAR_BIT)))))
-
/*
* Helper macros to work around __VA_ARGS__ limitations in a C99 compliant
* manner.
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 0e59d3f2f5..7d8ea4acba 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -3384,7 +3384,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[index]);
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -3413,7 +3413,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
mlx5_ifindex(dev),
mac);
if (!ret)
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return ret;
}
@@ -3532,12 +3532,12 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i)) {
+ if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
mlx5_ifindex(dev),
&dev->data->mac_addrs[i]);
- BITFIELD_RESET(priv->mac_own, i);
+ rte_bitset_clear(priv->mac_own, i);
}
}
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index c77cdc48a8..7f20c811f3 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -14,6 +14,7 @@
#include <rte_pci.h>
#include <rte_ether.h>
+#include <rte_bitset.h>
#include <ethdev_driver.h>
#include <rte_rwlock.h>
#include <rte_interrupts.h>
@@ -2018,7 +2019,7 @@ struct mlx5_priv {
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- BITFIELD_DECLARE(mac_own, uint64_t, MLX5_MAX_MAC_ADDRESSES);
+ RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
/* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 2a91e02b45..7f6148f4e1 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1921,7 +1921,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
/* Add flows for unicast and multicast mac addresses added by API. */
if (!memcmp(mac, &cmp, sizeof(*mac)) ||
- !BITFIELD_ISSET(priv->mac_own, i) ||
+ !rte_bitset_test(priv->mac_own, i) ||
(dev->data->all_multicast && rte_is_multicast_ether_addr(mac)))
continue;
memcpy(&unicast.hdr.dst_addr.addr_bytes,
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index ccc94d54af..0aaed36ace 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -699,8 +699,8 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
int i;
for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
- if (BITFIELD_ISSET(priv->mac_own, i))
- BITFIELD_RESET(priv->mac_own, i);
+ if (rte_bitset_test(priv->mac_own, i))
+ rte_bitset_clear(priv->mac_own, i);
}
}
@@ -718,7 +718,7 @@ mlx5_os_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
struct mlx5_priv *priv = dev->data->dev_private;
- BITFIELD_RESET(priv->mac_own, index);
+ rte_bitset_clear(priv->mac_own, index);
}
/**
@@ -756,7 +756,7 @@ mlx5_os_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
return -ENOTSUP;
}
/* Mark this MAC address as owned by the PMD */
- BITFIELD_SET(priv->mac_own, index);
+ rte_bitset_set(priv->mac_own, index);
return 0;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* [PATCH v9 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-24 6:38 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (2 preceding siblings ...)
2026-09-24 6:38 ` [PATCH v9 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
@ 2026-09-24 6:38 ` David Marchand
2026-09-28 14:40 ` Dariusz Sosnowski
2026-09-29 12:50 ` Raslan Darawsheh
2026-09-29 12:50 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux Raslan Darawsheh
4 siblings, 2 replies; 146+ messages in thread
From: David Marchand @ 2026-09-24 6:38 UTC (permalink / raw)
To: dev
Cc: rjarry, cfontain, Dariusz Sosnowski, Viacheslav Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
Starting firmware version 22.49.1014, the number of mac addresses
per VF is not capped to 128 anymore.
The value can be increased via devlink:
$ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
cmode driverinit
$ devlink dev reload pci/0000:3b:00.2
On the DPDK side, we must retrieve the maximum number of unicast
and multicast addresses supported with a query to the firmware.
Then, dynamically allocate the mac addresses arrays and report the
limit instead of the previous hardcoded value.
Signed-off-by: David Marchand <david.marchand@redhat.com>
---
Changes since v8:
- fixed (hopefully...) last thing in HWS stuff,
Changes since v7:
- updated HWS control path and removed MLX5_MAX_MAC_ADDRESSES constant,
Changes since v6:
- fixed crash in case of partial init failure in mlx5_dev_spawn,
and cleaned mlx5_os_pci_probe_pf,
- fixed compilation without assert in mlx5_internal_mac_addr_remove,
Changes since v4:
- added RN update,
- fixed types of fields added to mlx5_hca_attr and mlx5_dev_cap,
- used RTE_BIT32,
- fixed mlx5_nl_mac_addr_sync inverted arguments,
---
doc/guides/rel_notes/release_26_11.rst | 5 +++
drivers/common/mlx5/mlx5_devx_cmds.c | 4 ++
drivers/common/mlx5/mlx5_devx_cmds.h | 2 +
drivers/net/mlx5/linux/mlx5_os.c | 47 +++++++++++++++++------
drivers/net/mlx5/mlx5.c | 9 ++---
drivers/net/mlx5/mlx5.h | 11 +++---
drivers/net/mlx5/mlx5_ethdev.c | 2 +-
drivers/net/mlx5/mlx5_flow_hw.c | 53 +++++++++++++++++---------
drivers/net/mlx5/mlx5_mac.c | 20 ++++++----
drivers/net/mlx5/mlx5_trigger.c | 10 ++---
drivers/net/mlx5/windows/mlx5_os.c | 41 ++++++++++++++++----
11 files changed, 140 insertions(+), 64 deletions(-)
diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst
index 87c7e81bde..43043cd579 100644
--- a/doc/guides/rel_notes/release_26_11.rst
+++ b/doc/guides/rel_notes/release_26_11.rst
@@ -55,6 +55,11 @@ New Features
Also, make sure to start the actual text at the margin.
=======================================================
+* **Updated NVIDIA mlx5 ethernet driver.**
+
+ * Increased the maximum number of secondary unicast MAC addresses from 128 to up to 4096
+ (depending on devlink configuration on the associated kernel netdevice).
+
Removed Items
-------------
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.c b/drivers/common/mlx5/mlx5_devx_cmds.c
index 140b057ab4..e5d9c92779 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.c
+++ b/drivers/common/mlx5/mlx5_devx_cmds.c
@@ -1110,6 +1110,10 @@ mlx5_devx_cmd_query_hca_attr(void *ctx,
attr->log_max_pd = MLX5_GET(cmd_hca_cap, hcattr, log_max_pd);
attr->log_max_srq = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq);
attr->log_max_srq_sz = MLX5_GET(cmd_hca_cap, hcattr, log_max_srq_sz);
+ attr->log_max_current_uc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_uc_list);
+ attr->log_max_current_mc_list = MLX5_GET(cmd_hca_cap, hcattr,
+ log_max_current_mc_list);
attr->reg_c_preserve =
MLX5_GET(cmd_hca_cap, hcattr, reg_c_preserve);
attr->mmo_regex_qp_en = MLX5_GET(cmd_hca_cap, hcattr, regexp_mmo_qp);
diff --git a/drivers/common/mlx5/mlx5_devx_cmds.h b/drivers/common/mlx5/mlx5_devx_cmds.h
index 90beb2e9e6..504b4a4f64 100644
--- a/drivers/common/mlx5/mlx5_devx_cmds.h
+++ b/drivers/common/mlx5/mlx5_devx_cmds.h
@@ -356,6 +356,8 @@ struct mlx5_hca_attr {
uint8_t tx_sw_owner_v2:1;
uint8_t esw_sw_owner:1;
uint8_t esw_sw_owner_v2:1;
+ uint8_t log_max_current_uc_list:5;
+ uint8_t log_max_current_mc_list:5;
};
/* LAG Context. */
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 7d8ea4acba..9180e9aa20 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -389,6 +389,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
sh->dev_cap.esw_info.regc_mask = 0;
#endif
sh->dev_cap.esw_info.is_set = 1;
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -1464,6 +1474,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
/* Some internal functions rely on Netlink sockets, open them now. */
priv->nl_socket_rdma = nl_rdma;
priv->nl_socket_route = mlx5_nl_init(NETLINK_ROUTE, 0);
@@ -1762,8 +1788,8 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_nl_mac_addr_sync(priv->nl_socket_route,
mlx5_ifindex(eth_dev),
eth_dev->data->mac_addrs,
- MLX5_MAX_UC_MAC_ADDRESSES,
- MLX5_MAX_MAC_ADDRESSES);
+ sh->dev_cap.max_uc_mac_addrs,
+ sh->dev_cap.max_mac_addrs);
priv->ctrl_flows = 0;
rte_spinlock_init(&priv->flow_list_lock);
TAILQ_INIT(&priv->flow_meters);
@@ -1963,17 +1989,16 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_flex_item_port_cleanup(eth_dev);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
- if (eth_dev != NULL)
+ if (eth_dev != NULL) {
+ eth_dev->data->mac_addrs = NULL;
eth_dev->data->dev_private = NULL;
+ }
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
if (nl_rdma >= 0)
@@ -2979,8 +3004,6 @@ mlx5_os_pci_probe_pf(struct mlx5_common_device *cdev,
if (!list[i].eth_dev)
continue;
mlx5_dev_close(list[i].eth_dev);
- /* mac_addrs must not be freed because in dev_private */
- list[i].eth_dev->data->mac_addrs = NULL;
claim_zero(rte_eth_dev_release_port(list[i].eth_dev));
}
/* Restore original error. */
@@ -3531,7 +3554,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
const int vf = priv->sh->dev_cap.vf;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i)) {
if (vf)
mlx5_nl_mac_addr_remove(priv->nl_socket_route,
diff --git a/drivers/net/mlx5/mlx5.c b/drivers/net/mlx5/mlx5.c
index c7b0d3ef8b..4bd1c6d1c0 100644
--- a/drivers/net/mlx5/mlx5.c
+++ b/drivers/net/mlx5/mlx5.c
@@ -2556,6 +2556,9 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_list_destroy(priv->hrxqs);
mlx5_free(priv->ext_rxqs);
mlx5_free(priv->ext_txqs);
+ mlx5_free(priv->mac);
+ dev->data->mac_addrs = NULL;
+ mlx5_free(priv->mac_own);
sh->port[priv->dev_port - 1].nl_ih_port_id = RTE_MAX_ETHPORTS;
/*
* The interrupt handler port id must be reset before priv is reset
@@ -2590,12 +2593,6 @@ mlx5_dev_close(struct rte_eth_dev *dev)
mlx5_flow_pools_destroy(priv);
memset(priv, 0, sizeof(*priv));
priv->domain_id = RTE_ETH_DEV_SWITCH_DOMAIN_ID_INVALID;
- /*
- * Reset mac_addrs to NULL such that it is not freed as part of
- * rte_eth_dev_release_port(). mac_addrs is part of dev_private so
- * it is freed when dev_private is freed.
- */
- dev->data->mac_addrs = NULL;
return 0;
}
diff --git a/drivers/net/mlx5/mlx5.h b/drivers/net/mlx5/mlx5.h
index 7f20c811f3..e3df458915 100644
--- a/drivers/net/mlx5/mlx5.h
+++ b/drivers/net/mlx5/mlx5.h
@@ -87,9 +87,6 @@
#define MLX5_MAX_UC_MAC_ADDRESSES 128
/* Maximum number of simultaneous Multicast MAC addresses. */
#define MLX5_MAX_MC_MAC_ADDRESSES 128
-/* Maximum number of simultaneous MAC addresses. */
-#define MLX5_MAX_MAC_ADDRESSES \
- (MLX5_MAX_UC_MAC_ADDRESSES + MLX5_MAX_MC_MAC_ADDRESSES)
enum mlx5_ipool_index {
#if defined(HAVE_IBV_FLOW_DV_SUPPORT) || !defined(HAVE_INFINIBAND_VERBS_H)
@@ -217,6 +214,9 @@ struct mlx5_dev_cap {
} mprq; /* Capability for Multi-Packet RQ. */
char fw_ver[64]; /* Firmware version of this device. */
struct flow_hw_port_info esw_info; /* E-switch manager reg_c0. */
+ uint32_t max_uc_mac_addrs; /* Maximum unicast MAC addresses. */
+ uint32_t max_mc_mac_addrs; /* Maximum multicast MAC addresses. */
+ uint32_t max_mac_addrs; /* Total maximum MAC addresses. */
};
#define MLX5_MPESW_PORT_INVALID (-1)
@@ -2018,9 +2018,8 @@ struct mlx5_priv {
struct mlx5_dev_ctx_shared *sh; /* Shared device context. */
uint32_t dev_port; /* Device port number. */
struct rte_pci_device *pci_dev; /* Backend PCI device. */
- struct rte_ether_addr mac[MLX5_MAX_MAC_ADDRESSES]; /* MAC addresses. */
- RTE_BITSET_DECLARE(mac_own, MLX5_MAX_MAC_ADDRESSES);
- /* Bit-field of MAC addresses owned by the PMD. */
+ struct rte_ether_addr *mac; /* MAC addresses. */
+ uint64_t *mac_own; /* Bit-field of MAC addresses owned by the PMD. */
uint16_t vlan_filter[MLX5_MAX_VLAN_IDS]; /* VLAN filters table. */
unsigned int vlan_filter_n; /* Number of configured VLAN filters. */
/* Device properties. */
diff --git a/drivers/net/mlx5/mlx5_ethdev.c b/drivers/net/mlx5/mlx5_ethdev.c
index 8160d10e7e..306c1cd734 100644
--- a/drivers/net/mlx5/mlx5_ethdev.c
+++ b/drivers/net/mlx5/mlx5_ethdev.c
@@ -392,7 +392,7 @@ mlx5_dev_infos_get(struct rte_eth_dev *dev, struct rte_eth_dev_info *info)
max = RTE_MIN(max, (unsigned int)UINT16_MAX);
info->max_rx_queues = max;
info->max_tx_queues = max;
- info->max_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ info->max_mac_addrs = priv->sh->dev_cap.max_uc_mac_addrs;
info->rx_queue_offload_capa = mlx5_get_rx_queue_offloads(dev);
info->rx_seg_capa.max_nseg = MLX5_MAX_RXQ_NSEG;
info->rx_seg_capa.multi_pools = !priv->config.mprq.enabled;
diff --git a/drivers/net/mlx5/mlx5_flow_hw.c b/drivers/net/mlx5/mlx5_flow_hw.c
index 30ef2f1c2d..8dabf21b76 100644
--- a/drivers/net/mlx5/mlx5_flow_hw.c
+++ b/drivers/net/mlx5/mlx5_flow_hw.c
@@ -11576,22 +11576,37 @@ static uint32_t ctrl_rx_rss_priority_map[MLX5_FLOW_HW_CTRL_RX_EXPANDED_RSS_MAX]
[MLX5_FLOW_HW_CTRL_RX_EXPANDED_RSS_IPV6_TCP] = MLX5_HW_CTRL_RX_PRIO_L4,
};
-static uint32_t ctrl_rx_nb_flows_map[MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_MAX] = {
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_ALL] = 1,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_ALL_MCAST] = 1,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_BCAST] = 1,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_BCAST_VLAN] = MLX5_MAX_VLAN_IDS,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_IPV4_MCAST] = 1,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_IPV4_MCAST_VLAN] = MLX5_MAX_VLAN_IDS,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_IPV6_MCAST] = 1,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_IPV6_MCAST_VLAN] = MLX5_MAX_VLAN_IDS,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_DMAC] = MLX5_MAX_UC_MAC_ADDRESSES,
- [MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_DMAC_VLAN] =
- MLX5_MAX_UC_MAC_ADDRESSES * MLX5_MAX_VLAN_IDS,
-};
+static uint32_t
+flow_hw_get_ctrl_rx_nb_flows(struct rte_eth_dev *dev,
+ enum mlx5_flow_ctrl_rx_eth_pattern_type eth_pattern_type)
+{
+ struct mlx5_priv *priv = dev->data->dev_private;
+ uint32_t max_uc_mac_addrs = priv->sh->dev_cap.max_uc_mac_addrs;
+
+ switch (eth_pattern_type) {
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_ALL:
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_ALL_MCAST:
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_BCAST:
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_IPV4_MCAST:
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_IPV6_MCAST:
+ return 1;
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_BCAST_VLAN:
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_IPV4_MCAST_VLAN:
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_IPV6_MCAST_VLAN:
+ return MLX5_MAX_VLAN_IDS;
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_DMAC:
+ return max_uc_mac_addrs;
+ case MLX5_FLOW_HW_CTRL_RX_ETH_PATTERN_DMAC_VLAN:
+ return max_uc_mac_addrs * MLX5_MAX_VLAN_IDS;
+ default:
+ MLX5_ASSERT(false);
+ return 0;
+ }
+}
static struct rte_flow_template_table_attr
-flow_hw_get_ctrl_rx_table_attr(enum mlx5_flow_ctrl_rx_eth_pattern_type eth_pattern_type,
+flow_hw_get_ctrl_rx_table_attr(struct rte_eth_dev *dev,
+ enum mlx5_flow_ctrl_rx_eth_pattern_type eth_pattern_type,
const enum mlx5_flow_ctrl_rx_expanded_rss_type rss_type)
{
return (struct rte_flow_template_table_attr){
@@ -11600,7 +11615,7 @@ flow_hw_get_ctrl_rx_table_attr(enum mlx5_flow_ctrl_rx_eth_pattern_type eth_patte
.priority = ctrl_rx_rss_priority_map[rss_type],
.ingress = 1,
},
- .nb_flows = ctrl_rx_nb_flows_map[eth_pattern_type],
+ .nb_flows = flow_hw_get_ctrl_rx_nb_flows(dev, eth_pattern_type),
};
}
@@ -11763,7 +11778,8 @@ mlx5_flow_hw_create_ctrl_rx_tables(struct rte_eth_dev *dev)
struct rte_flow_template_table_attr attr;
struct rte_flow_pattern_template *pt;
- attr = flow_hw_get_ctrl_rx_table_attr(eth_pattern_type, rss_type);
+ attr = flow_hw_get_ctrl_rx_table_attr(dev, eth_pattern_type,
+ rss_type);
pt = flow_hw_create_ctrl_rx_pattern_template(dev, eth_pattern_type,
rss_type);
if (!pt)
@@ -16697,10 +16713,11 @@ __flow_hw_ctrl_flows_unicast(struct rte_eth_dev *dev,
struct rte_flow_template_table *tbl,
const enum mlx5_flow_ctrl_rx_expanded_rss_type rss_type)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
unsigned int i;
int ret;
- for (i = 0; i < MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i < priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -16767,7 +16784,7 @@ __flow_hw_ctrl_flows_unicast_vlan(struct rte_eth_dev *dev,
unsigned int i;
unsigned int j;
- for (i = 0; i < MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i < priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
diff --git a/drivers/net/mlx5/mlx5_mac.c b/drivers/net/mlx5/mlx5_mac.c
index 0e5d2be530..9e7bf7f2ec 100644
--- a/drivers/net/mlx5/mlx5_mac.c
+++ b/drivers/net/mlx5/mlx5_mac.c
@@ -36,7 +36,7 @@ mlx5_internal_mac_addr_remove(struct rte_eth_dev *dev,
uint32_t index,
struct rte_ether_addr *addr)
{
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < MLX5_SH(dev)->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(&dev->data->mac_addrs[index]))
return false;
mlx5_os_mac_addr_remove(dev, index);
@@ -63,16 +63,17 @@ static int
mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
unsigned int i;
int ret;
- MLX5_ASSERT(index < MLX5_MAX_MAC_ADDRESSES);
+ MLX5_ASSERT(index < priv->sh->dev_cap.max_mac_addrs);
if (rte_is_zero_ether_addr(mac)) {
rte_errno = EINVAL;
return -rte_errno;
}
/* First, make sure this address isn't already configured. */
- for (i = 0; (i != MLX5_MAX_MAC_ADDRESSES); ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
/* Skip this index, it's going to be reconfigured. */
if (i == index)
continue;
@@ -101,10 +102,11 @@ mlx5_internal_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
void
mlx5_mac_addr_remove(struct rte_eth_dev *dev, uint32_t index)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
struct rte_ether_addr addr = { 0 };
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES)
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs)
return;
if (mlx5_internal_mac_addr_remove(dev, index, &addr)) {
ret = mlx5_traffic_mac_remove(dev, &addr);
@@ -133,9 +135,10 @@ int
mlx5_mac_addr_add(struct rte_eth_dev *dev, struct rte_ether_addr *mac,
uint32_t index, uint32_t vmdq __rte_unused)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
int ret;
- if (index >= MLX5_MAX_UC_MAC_ADDRESSES) {
+ if (index >= priv->sh->dev_cap.max_uc_mac_addrs) {
rte_errno = EINVAL;
return -rte_errno;
}
@@ -217,16 +220,17 @@ int
mlx5_set_mc_addr_list(struct rte_eth_dev *dev,
struct rte_ether_addr *mc_addr_set, uint32_t nb_mc_addr)
{
+ struct mlx5_priv *priv = dev->data->dev_private;
uint32_t i;
int ret;
- if (nb_mc_addr >= MLX5_MAX_MC_MAC_ADDRESSES) {
+ if (nb_mc_addr >= priv->sh->dev_cap.max_mc_mac_addrs) {
rte_errno = ENOSPC;
return -rte_errno;
}
- for (i = MLX5_MAX_UC_MAC_ADDRESSES; i != MLX5_MAX_MAC_ADDRESSES; ++i)
+ for (i = priv->sh->dev_cap.max_uc_mac_addrs; i != priv->sh->dev_cap.max_mac_addrs; ++i)
mlx5_internal_mac_addr_remove(dev, i, NULL);
- i = MLX5_MAX_UC_MAC_ADDRESSES;
+ i = priv->sh->dev_cap.max_uc_mac_addrs;
while (nb_mc_addr--) {
ret = mlx5_internal_mac_addr_add(dev, mc_addr_set++, i++);
if (ret)
diff --git a/drivers/net/mlx5/mlx5_trigger.c b/drivers/net/mlx5/mlx5_trigger.c
index 7f6148f4e1..c5f493117b 100644
--- a/drivers/net/mlx5/mlx5_trigger.c
+++ b/drivers/net/mlx5/mlx5_trigger.c
@@ -1916,7 +1916,7 @@ mlx5_traffic_enable(struct rte_eth_dev *dev)
}
}
/* Add MAC address flows. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
/* Add flows for unicast and multicast mac addresses added by API. */
@@ -2186,7 +2186,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
return 0;
/* Add all unicast DMAC flow rules with new VLAN attached. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2203,7 +2203,7 @@ mlx5_traffic_vlan_add(struct rte_eth_dev *dev, const uint16_t vid)
* Removing after creating VLAN rules so that traffic "gap" is not introduced.
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2241,7 +2241,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
* Recreating first to ensure no traffic "gap".
*/
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
@@ -2254,7 +2254,7 @@ mlx5_traffic_vlan_remove(struct rte_eth_dev *dev, const uint16_t vid)
}
/* Remove all unicast DMAC flow rules with this VLAN. */
- for (i = 0; i != MLX5_MAX_MAC_ADDRESSES; ++i) {
+ for (i = 0; i != priv->sh->dev_cap.max_mac_addrs; ++i) {
struct rte_ether_addr *mac = &dev->data->mac_addrs[i];
if (rte_is_zero_ether_addr(mac))
diff --git a/drivers/net/mlx5/windows/mlx5_os.c b/drivers/net/mlx5/windows/mlx5_os.c
index 0aaed36ace..cc60f07e17 100644
--- a/drivers/net/mlx5/windows/mlx5_os.c
+++ b/drivers/net/mlx5/windows/mlx5_os.c
@@ -261,6 +261,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
MLX5_GET(initial_seg, pv_iseg, fw_rev_subminor));
DRV_LOG(DEBUG, "Packet pacing is not supported.");
mlx5_rt_timestamp_config(sh, hca_attr);
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
@@ -396,6 +406,22 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
priv->sh = sh;
priv->dev_port = spawn->phys_port;
priv->pci_dev = spawn->pci_dev;
+ priv->mac = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ sizeof(*priv->mac) * sh->dev_cap.max_mac_addrs,
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC address array.");
+ err = ENOMEM;
+ goto error;
+ }
+ priv->mac_own = mlx5_malloc(MLX5_MEM_ZERO | MLX5_MEM_RTE,
+ RTE_BITSET_SIZE(sh->dev_cap.max_mac_addrs),
+ RTE_CACHE_LINE_SIZE, SOCKET_ID_ANY);
+ if (priv->mac_own == NULL) {
+ DRV_LOG(ERR, "Failed to allocate MAC ownership bitmap.");
+ err = ENOMEM;
+ goto error;
+ }
priv->mp_id.port_id = port_id;
strlcpy(priv->mp_id.name, MLX5_MP_NAME, RTE_MP_MAX_NAME_LEN);
priv->representor = !!switch_info->representor;
@@ -612,17 +638,16 @@ mlx5_dev_spawn(struct rte_device *dpdk_dev,
mlx5_l3t_destroy(priv->mtr_profile_tbl);
if (own_domain_id)
claim_zero(rte_eth_switch_domain_free(priv->domain_id));
+ mlx5_free(priv->mac);
+ mlx5_free(priv->mac_own);
mlx5_free(priv);
- if (eth_dev != NULL)
+ if (eth_dev != NULL) {
+ eth_dev->data->mac_addrs = NULL;
eth_dev->data->dev_private = NULL;
+ }
}
- if (eth_dev != NULL) {
- /* mac_addrs must not be freed alone because part of
- * dev_private
- **/
- eth_dev->data->mac_addrs = NULL;
+ if (eth_dev != NULL)
rte_eth_dev_release_port(eth_dev);
- }
if (sh)
mlx5_free_shared_dev_ctx(sh);
MLX5_ASSERT(err > 0);
@@ -698,7 +723,7 @@ mlx5_os_mac_addr_flush(struct rte_eth_dev *dev)
struct mlx5_priv *priv = dev->data->dev_private;
int i;
- for (i = MLX5_MAX_MAC_ADDRESSES - 1; i >= 0; --i) {
+ for (i = priv->sh->dev_cap.max_mac_addrs - 1; i >= 0; --i) {
if (rte_bitset_test(priv->mac_own, i))
rte_bitset_clear(priv->mac_own, i);
}
--
2.54.0
^ permalink raw reply related [flat|nested] 146+ messages in thread
* RE: [PATCH v9 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-24 6:38 ` [PATCH v9 5/5] net/mlx5: accept more unicast " David Marchand
@ 2026-09-28 14:40 ` Dariusz Sosnowski
2026-09-29 12:50 ` Raslan Darawsheh
1 sibling, 0 replies; 146+ messages in thread
From: Dariusz Sosnowski @ 2026-09-28 14:40 UTC (permalink / raw)
To: David Marchand, dev@dpdk.org
Cc: rjarry@redhat.com, cfontain@redhat.com, Slava Ovsiienko,
Bing Zhao, Ori Kam, Suanming Mou, Matan Azrad
> -----Original Message-----
> From: David Marchand <david.marchand@redhat.com>
> Sent: Thursday, September 24, 2026 8:39 AM
> To: dev@dpdk.org
> Cc: rjarry@redhat.com; cfontain@redhat.com; Dariusz Sosnowski
> <dsosnowski@nvidia.com>; Slava Ovsiienko <viacheslavo@nvidia.com>; Bing
> Zhao <bingz@nvidia.com>; Ori Kam <orika@nvidia.com>; Suanming Mou
> <suanmingm@nvidia.com>; Matan Azrad <matan@nvidia.com>
> Subject: [PATCH v9 5/5] net/mlx5: accept more unicast MAC addresses
>
> Starting firmware version 22.49.1014, the number of mac addresses per VF is
> not capped to 128 anymore.
>
> The value can be increased via devlink:
> $ devlink dev param set pci/0000:3b:00.2 name max_macs value 4096 \
> cmode driverinit
> $ devlink dev reload pci/0000:3b:00.2
>
> On the DPDK side, we must retrieve the maximum number of unicast and
> multicast addresses supported with a query to the firmware.
>
> Then, dynamically allocate the mac addresses arrays and report the limit
> instead of the previous hardcoded value.
>
> Signed-off-by: David Marchand <david.marchand@redhat.com>
Acked-by: Dariusz Sosnowski <dsosnowski@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v9 2/5] net/mlx5: remove redundant MAC address index checks
2026-09-24 6:38 ` [PATCH v9 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
@ 2026-09-29 12:50 ` Raslan Darawsheh
0 siblings, 0 replies; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-29 12:50 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Acked-by: Raslan Darawsheh <rasland@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v9 3/5] net/mlx5: pass maximum number of unicast MAC to common code
2026-09-24 6:38 ` [PATCH v9 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
@ 2026-09-29 12:50 ` Raslan Darawsheh
0 siblings, 0 replies; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-29 12:50 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Acked-by: Raslan Darawsheh <rasland@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v9 4/5] net/mlx5: use bitset for tracking MAC addresses
2026-09-24 6:38 ` [PATCH v9 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
@ 2026-09-29 12:50 ` Raslan Darawsheh
0 siblings, 0 replies; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-29 12:50 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Acked-by: Raslan Darawsheh <rasland@nvidia.com>
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v9 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-24 6:38 ` [PATCH v9 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-28 14:40 ` Dariusz Sosnowski
@ 2026-09-29 12:50 ` Raslan Darawsheh
2026-09-29 13:24 ` David Marchand
1 sibling, 1 reply; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-29 12:50 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Acked-by: Raslan Darawsheh <rasland@nvidia.com>
I'll squash the following fixup into this patch during merge, to handle
non-DevX devices (where mlx5_os_capabilities_prepare() would otherwise
return before max_uc_mac_addrs/max_mc_mac_addrs/max_mac_addrs get set,
leaving them at 0):
diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
index 9180e9aa20..7c2c35c945 100644
--- a/drivers/net/mlx5/linux/mlx5_os.c
+++ b/drivers/net/mlx5/linux/mlx5_os.c
@@ -325,6 +325,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
DRV_LOG(WARNING,
"Tunnel offloading disabled due to old OFED/rdma-core version");
#endif
+ if (hca_attr->log_max_current_uc_list > 0)
+ sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
+ else
+ sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
+ if (hca_attr->log_max_current_mc_list > 0)
+ sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
+ else
+ sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
+ sh->dev_cap.max_mac_addrs =
+ sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
if (!sh->cdev->config.devx)
return 0;
/* Check capabilities for Packet Pacing. */
@@ -389,16 +399,6 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
sh->dev_cap.esw_info.regc_mask = 0;
#endif
sh->dev_cap.esw_info.is_set = 1;
- if (hca_attr->log_max_current_uc_list > 0)
- sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
- else
- sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
- if (hca_attr->log_max_current_mc_list > 0)
- sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
- else
- sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
- sh->dev_cap.max_mac_addrs =
- sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
return 0;
}
Thanks,
Raslan
^ permalink raw reply related [flat|nested] 146+ messages in thread
* Re: [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux
2026-09-24 6:38 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
` (3 preceding siblings ...)
2026-09-24 6:38 ` [PATCH v9 5/5] net/mlx5: accept more unicast " David Marchand
@ 2026-09-29 12:50 ` Raslan Darawsheh
4 siblings, 0 replies; 146+ messages in thread
From: Raslan Darawsheh @ 2026-09-29 12:50 UTC (permalink / raw)
To: dev; +Cc: David Marchand, Dariusz Sosnowski
Acked-by: Raslan Darawsheh <rasland@nvidia.com>
Series applied to next-net-mlx.
Thanks,
Raslan
^ permalink raw reply [flat|nested] 146+ messages in thread
* Re: [PATCH v9 5/5] net/mlx5: accept more unicast MAC addresses
2026-09-29 12:50 ` Raslan Darawsheh
@ 2026-09-29 13:24 ` David Marchand
0 siblings, 0 replies; 146+ messages in thread
From: David Marchand @ 2026-09-29 13:24 UTC (permalink / raw)
To: Raslan Darawsheh; +Cc: dev, Dariusz Sosnowski
On Tue, 29 Sept 2026 at 14:52, Raslan Darawsheh <rasland@nvidia.com> wrote:
>
> Acked-by: Raslan Darawsheh <rasland@nvidia.com>
>
> I'll squash the following fixup into this patch during merge, to handle
> non-DevX devices (where mlx5_os_capabilities_prepare() would otherwise
> return before max_uc_mac_addrs/max_mc_mac_addrs/max_mac_addrs get set,
> leaving them at 0):
>
> diff --git a/drivers/net/mlx5/linux/mlx5_os.c b/drivers/net/mlx5/linux/mlx5_os.c
> index 9180e9aa20..7c2c35c945 100644
> --- a/drivers/net/mlx5/linux/mlx5_os.c
> +++ b/drivers/net/mlx5/linux/mlx5_os.c
> @@ -325,6 +325,16 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
> DRV_LOG(WARNING,
> "Tunnel offloading disabled due to old OFED/rdma-core version");
> #endif
> + if (hca_attr->log_max_current_uc_list > 0)
> + sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
> + else
> + sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
> + if (hca_attr->log_max_current_mc_list > 0)
> + sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
> + else
> + sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
> + sh->dev_cap.max_mac_addrs =
> + sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
> if (!sh->cdev->config.devx)
> return 0;
> /* Check capabilities for Packet Pacing. */
> @@ -389,16 +399,6 @@ mlx5_os_capabilities_prepare(struct mlx5_dev_ctx_shared *sh)
> sh->dev_cap.esw_info.regc_mask = 0;
> #endif
> sh->dev_cap.esw_info.is_set = 1;
> - if (hca_attr->log_max_current_uc_list > 0)
> - sh->dev_cap.max_uc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_uc_list);
> - else
> - sh->dev_cap.max_uc_mac_addrs = MLX5_MAX_UC_MAC_ADDRESSES;
> - if (hca_attr->log_max_current_mc_list > 0)
> - sh->dev_cap.max_mc_mac_addrs = RTE_BIT32(hca_attr->log_max_current_mc_list);
> - else
> - sh->dev_cap.max_mc_mac_addrs = MLX5_MAX_MC_MAC_ADDRESSES;
> - sh->dev_cap.max_mac_addrs =
> - sh->dev_cap.max_uc_mac_addrs + sh->dev_cap.max_mc_mac_addrs;
> return 0;
> }
This lgtm.
Thank you.
--
David Marchand
^ permalink raw reply [flat|nested] 146+ messages in thread
end of thread, other threads:[~2026-09-29 13:24 UTC | newest]
Thread overview: 146+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-04-03 9:18 [PATCH 0/4] Remove limitations coming from legacy VMDq David Marchand
2026-04-03 9:18 ` [PATCH 1/4] ethdev: skip VMDq pools unless configured David Marchand
2026-06-01 9:30 ` Andrew Rybchenko
2026-04-03 9:18 ` [PATCH 2/4] ethdev: announce VMDq capability David Marchand
2026-04-06 22:22 ` Kishore Padmanabha
2026-04-29 14:18 ` David Marchand
2026-05-18 22:12 ` Kishore Padmanabha
2026-06-01 9:32 ` Andrew Rybchenko
2026-04-03 9:18 ` [PATCH 3/4] ethdev: hide VMDq internal sizes David Marchand
2026-06-01 9:34 ` Andrew Rybchenko
2026-04-03 9:18 ` [PATCH 4/4] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-04-05 18:47 ` [PATCH 0/4] Remove limitations coming from legacy VMDq Stephen Hemminger
2026-04-29 14:22 ` David Marchand
2026-05-06 12:35 ` [PATCH v2 0/5] " David Marchand
2026-05-06 12:35 ` [PATCH v2 1/5] ethdev: skip VMDq pools unless configured David Marchand
2026-06-01 9:35 ` Andrew Rybchenko
2026-05-06 12:35 ` [PATCH v2 2/5] ethdev: announce VMDq capability David Marchand
2026-06-01 9:36 ` Andrew Rybchenko
2026-05-06 12:35 ` [PATCH v2 3/5] ethdev: hide VMDq internal sizes David Marchand
2026-05-06 12:35 ` [PATCH v2 4/5] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-05-06 12:35 ` [PATCH v2 5/5] net/iavf: fix duplicate MAC addresses install David Marchand
2026-05-07 2:51 ` [PATCH v2 0/5] Remove limitations coming from legacy VMDq Stephen Hemminger
2026-05-10 15:03 ` David Marchand
2026-05-10 17:03 ` [PATCH v3 " David Marchand
2026-05-10 17:03 ` [PATCH v3 1/5] ethdev: check VMDq availability David Marchand
2026-06-01 9:38 ` Andrew Rybchenko
2026-05-10 17:03 ` [PATCH v3 2/5] ethdev: skip VMDq pools unless configured David Marchand
2026-06-01 9:38 ` Andrew Rybchenko
2026-05-10 17:03 ` [PATCH v3 3/5] ethdev: hide VMDq internal sizes David Marchand
2026-06-01 9:39 ` Andrew Rybchenko
2026-05-10 17:03 ` [PATCH v3 4/5] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-05-12 14:41 ` Stephen Hemminger
2026-05-27 13:25 ` David Marchand
2026-05-10 17:03 ` [PATCH v3 5/5] net/iavf: fix duplicate MAC addresses install David Marchand
2026-07-09 16:02 ` [PATCH v4 00/10] Remove limitations coming from legacy VMDq David Marchand
2026-07-09 16:02 ` [PATCH v4 01/10] ethdev: check VMDq availability David Marchand
2026-07-09 16:02 ` [PATCH v4 02/10] ethdev: skip VMDq pools unless configured David Marchand
2026-07-09 16:02 ` [PATCH v4 03/10] ethdev: hide VMDq internal sizes David Marchand
2026-07-09 16:02 ` [PATCH v4 04/10] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-07-09 16:02 ` [PATCH v4 05/10] net/iavf: fix duplicate MAC addresses install David Marchand
2026-07-13 13:12 ` Loftus, Ciara
2026-07-13 14:10 ` David Marchand
2026-07-14 9:23 ` Loftus, Ciara
2026-07-09 16:02 ` [PATCH v4 06/10] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-07-09 16:02 ` [PATCH v4 07/10] net/mlx5: remove redundant MAC address index checks David Marchand
2026-07-09 16:02 ` [PATCH v4 08/10] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
2026-07-09 16:02 ` [PATCH v4 09/10] net/mlx5: use bitset for tracking MAC addresses David Marchand
2026-07-09 16:02 ` [PATCH v4 10/10] net/mlx5: accept more unicast " David Marchand
2026-07-10 6:44 ` David Marchand
2026-07-10 7:48 ` David Marchand
2026-07-23 12:41 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
2026-07-23 12:41 ` [PATCH v5 01/10] ethdev: check VMDq availability David Marchand
2026-07-23 12:41 ` [PATCH v5 02/10] ethdev: skip VMDq pools unless configured David Marchand
2026-07-23 12:41 ` [PATCH v5 03/10] ethdev: hide VMDq internal sizes David Marchand
2026-07-23 12:41 ` [PATCH v5 04/10] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-07-23 12:41 ` [PATCH v5 05/10] net/iavf: fix duplicate MAC addresses install David Marchand
2026-07-23 12:41 ` [PATCH v5 06/10] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-07-23 12:41 ` [PATCH v5 07/10] net/mlx5: remove redundant MAC address index checks David Marchand
2026-07-23 12:41 ` [PATCH v5 08/10] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
2026-07-23 17:35 ` Stephen Hemminger
2026-07-23 12:41 ` [PATCH v5 09/10] net/mlx5: use bitset for tracking MAC addresses David Marchand
2026-07-23 12:41 ` [PATCH v5 10/10] net/mlx5: accept more unicast " David Marchand
2026-07-27 7:20 ` [PATCH v5 00/10] Remove limitations coming from legacy VMDq David Marchand
2026-08-24 11:42 ` [PATCH v6 0/3] " David Marchand
2026-08-24 11:42 ` [PATCH v6 1/3] ethdev: check VMDq availability David Marchand
2026-08-24 11:42 ` [PATCH v6 2/3] ethdev: skip VMDq pools unless configured David Marchand
2026-08-24 16:21 ` Stephen Hemminger
2026-08-24 16:24 ` David Marchand
2026-08-24 16:39 ` Stephen Hemminger
2026-08-24 11:42 ` [PATCH v6 3/3] ethdev: hide VMDq internal sizes David Marchand
2026-08-24 17:01 ` [PATCH v6 0/3] Remove limitations coming from legacy VMDq Stephen Hemminger
2026-09-04 12:28 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-09-04 12:28 ` [PATCH v6 2/2] net/iavf: fix duplicate MAC addresses install David Marchand
2026-09-09 9:39 ` Loftus, Ciara
2026-09-11 14:14 ` David Marchand
2026-09-11 15:37 ` David Marchand
2026-09-11 16:26 ` David Marchand
2026-09-10 10:24 ` [PATCH v6 1/2] net/iavf: accept up to 32k unicast MAC addresses Burakov, Anatoly
2026-09-11 9:37 ` Burakov, Anatoly
2026-09-11 11:52 ` David Marchand
2026-09-11 12:14 ` Burakov, Anatoly
2026-09-10 12:13 ` Burakov, Anatoly
2026-09-10 12:20 ` Burakov, Anatoly
2026-09-10 12:30 ` David Marchand
2026-09-10 12:38 ` Burakov, Anatoly
2026-09-08 9:27 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-08 9:27 ` [PATCH v6 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
2026-09-11 8:36 ` Dariusz Sosnowski
2026-09-08 9:27 ` [PATCH v6 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
2026-09-11 8:38 ` Dariusz Sosnowski
2026-09-08 9:27 ` [PATCH v6 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
2026-09-11 8:40 ` Dariusz Sosnowski
2026-09-08 9:27 ` [PATCH v6 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-11 8:59 ` Dariusz Sosnowski
2026-09-11 9:55 ` David Marchand
2026-09-11 10:01 ` Dariusz Sosnowski
2026-09-11 8:35 ` [PATCH v6 1/5] net/mlx5: remove MAC addresses flush helper on Linux Dariusz Sosnowski
2026-09-14 8:17 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset David Marchand
2026-09-14 8:17 ` [PATCH v7 2/4] net/iavf: fix duplicate MAC addresses install David Marchand
2026-09-14 10:19 ` Loftus, Ciara
2026-09-14 11:56 ` David Marchand
2026-09-14 12:02 ` Bruce Richardson
2026-09-14 12:27 ` David Marchand
2026-09-23 11:53 ` Burakov, Anatoly
2026-09-14 8:17 ` [PATCH v7 3/4] net/iavf: add a helper for sending MAC addresses to PF David Marchand
2026-09-23 12:04 ` Burakov, Anatoly
2026-09-14 8:17 ` [PATCH v7 4/4] net/iavf: accept up to 32k unicast MAC addresses David Marchand
2026-09-23 12:14 ` Burakov, Anatoly
2026-09-14 10:15 ` [PATCH v7 1/4] net/iavf: fix MAC addresses leak on reset Loftus, Ciara
2026-09-14 11:54 ` David Marchand
2026-09-14 11:57 ` Loftus, Ciara
2026-09-23 11:47 ` Burakov, Anatoly
2026-09-23 14:37 ` Bruce Richardson
2026-09-14 14:42 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-14 14:42 ` [PATCH v7 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
2026-09-21 8:31 ` Raslan Darawsheh
2026-09-21 10:06 ` David Marchand
2026-09-14 14:42 ` [PATCH v7 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
2026-09-21 8:31 ` Raslan Darawsheh
2026-09-21 10:07 ` David Marchand
2026-09-14 14:42 ` [PATCH v7 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
2026-09-14 14:42 ` [PATCH v7 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-14 14:51 ` Dariusz Sosnowski
2026-09-21 8:31 ` Raslan Darawsheh
2026-09-21 8:31 ` [PATCH v7 1/5] net/mlx5: remove MAC addresses flush helper on Linux Raslan Darawsheh
2026-09-21 10:31 ` David Marchand
2026-09-21 11:02 ` Raslan Darawsheh
2026-09-21 11:50 ` [PATCH v8 " David Marchand
2026-09-21 11:50 ` [PATCH v8 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
2026-09-21 11:50 ` [PATCH v8 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
2026-09-21 11:50 ` [PATCH v8 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
2026-09-21 11:50 ` [PATCH v8 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-21 11:52 ` David Marchand
2026-09-23 11:55 ` Raslan Darawsheh
2026-09-24 6:38 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux David Marchand
2026-09-24 6:38 ` [PATCH v9 2/5] net/mlx5: remove redundant MAC address index checks David Marchand
2026-09-29 12:50 ` Raslan Darawsheh
2026-09-24 6:38 ` [PATCH v9 3/5] net/mlx5: pass maximum number of unicast MAC to common code David Marchand
2026-09-29 12:50 ` Raslan Darawsheh
2026-09-24 6:38 ` [PATCH v9 4/5] net/mlx5: use bitset for tracking MAC addresses David Marchand
2026-09-29 12:50 ` Raslan Darawsheh
2026-09-24 6:38 ` [PATCH v9 5/5] net/mlx5: accept more unicast " David Marchand
2026-09-28 14:40 ` Dariusz Sosnowski
2026-09-29 12:50 ` Raslan Darawsheh
2026-09-29 13:24 ` David Marchand
2026-09-29 12:50 ` [PATCH v9 1/5] net/mlx5: remove MAC addresses flush helper on Linux Raslan Darawsheh
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.