* [PATCH v11 net-next] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers
@ 2026-09-02 1:55 Ratheesh Kannoth
2026-09-03 2:29 ` Ratheesh Kannoth
2026-09-07 7:58 ` netdev-bot+sashiko
0 siblings, 2 replies; 3+ messages in thread
From: Ratheesh Kannoth @ 2026-09-02 1:55 UTC (permalink / raw)
To: bpf, linux-kernel, netdev
Cc: andrew+netdev, ast, daniel, davem, edumazet, hawk, john.fastabend,
kuba, pabeni, sdf, sgoutham, Ratheesh Kannoth
Add TC_SETUP_QDISC_MQPRIO offload for channel-mode mqprio with
TC_MQPRIO_SHAPER_BW_RATE. Program per-queue MDQ CIR/PIR through the
NIX TX scheduler mailbox for each non-QoS transmit queue. When offload
is active, allocate one SMQ per such queue, parent every MDQ under
TL4[0], and map each traffic class min/max rate to the queue(s) in that
class.
Rebuild the TX scheduler hierarchy by bouncing the netdev through
ndo_stop() and ndo_open() on mqprio add, replace, and delete. Cache
the active rates and restore MDQ shapers from otx2_mqprio_up() during
ndo_open(); log and continue if restoration fails so a normal open is
not blocked.
Track mqprio configuration in mq_offload_snap snapshots (TC layout and
rates). On tc qdisc replace, stage the new configuration while keeping
the previous snapshot for rollback: failed setup restores the old
snapshot, successful graft is recorded through TC_ROOT_GRAFT, and
teardown of the replaced qdisc instance commits the staged snapshot
without tearing down the live offload.
Reject offload unless the interface is running and the device advertises
CIR+PIR support. Reject per-TC rates for traffic classes mapped to more
than one queue, SDP rep ports, and concurrent PFC or XDP use. Block
ethtool channel count changes while mqprio bandwidth offload is active.
Add a ratelimited AF debug message when validating TX scheduler queue
ownership to aid mqprio hierarchy setup failures.
Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
---
v10 -> v11: Addressed sashiko comments.
https://sashiko.dev/#/patchset/20260831131014.2639581-1-rkannoth%40marvell.com
v9 -> v10: Addressed sashiko/jacub comments.
https://sashiko.dev/#/message/20260817032747.1765883-1-rkannoth%40marvell.com
v8 -> v9: Addressed Sashiko comments
https://lore.kernel.org/netdev/aoJ6FhtWue0FHDQV@rkannoth-OptiPlex-7090/
v7 -> v8: Addressed Sashiko comments
https://sashiko.dev/#/patchset/20260811085050.3212280-1-rkannoth%40marvell.com
v6 -> v7: Addressed Sashiko comments
https://sashiko.dev/#/message/20260810034738.1786029-1-rkannoth%40marvell.com
v5 -> v6: Addressed Sashiko comments
https://lore.kernel.org/netdev/20260806095434.1144397-1-rkannoth@marvell.com/
v4 -> v5: Addressed sashiko comments
https://sashiko.dev/#/patchset/20260803042724.3380209-1-rkannoth%40marvell.com
v3 -> v4: Addressed sashiko comments
https://lore.kernel.org/netdev/20260729105139.2302908-1-rkannoth@marvell.com/
v2 -> v3: Addressed sashiko comments
https://lore.kernel.org/netdev/amnYX866mYx02cBe@rkannoth-OptiPlex-7090/T/#m67310cbec48b21c7720858ab3a1ea083a0f8dc10
v1 -> v2: Addressed sashiko comments
https://lore.kernel.org/netdev/20260724075010.2665758-1-rkannoth@marvell.com/
---
.../ethernet/marvell/octeontx2/af/rvu_nix.c | 6 +-
.../marvell/octeontx2/nic/otx2_common.c | 145 +++-
.../marvell/octeontx2/nic/otx2_common.h | 27 +
.../marvell/octeontx2/nic/otx2_dcbnl.c | 6 +
.../marvell/octeontx2/nic/otx2_ethtool.c | 8 +
.../ethernet/marvell/octeontx2/nic/otx2_pf.c | 12 +
.../ethernet/marvell/octeontx2/nic/otx2_tc.c | 702 ++++++++++++++++++
.../net/ethernet/marvell/octeontx2/nic/qos.c | 3 +
8 files changed, 907 insertions(+), 2 deletions(-)
diff --git a/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c b/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
index 153eb57bad06..c9a4326d725d 100644
--- a/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
+++ b/drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c
@@ -331,8 +331,12 @@ static bool is_valid_txschq(struct rvu *rvu, int blkaddr,
return true;
}
- if (map_func != pcifunc)
+ if (map_func != pcifunc) {
+ dev_err_ratelimited(rvu->dev,
+ "pcifunc %x map pcifunc %x not equal, lvl=%u schq=%u\n",
+ pcifunc, map_func, lvl, schq);
return false;
+ }
return true;
}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
index 175992188c18..5c502c9d7c83 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
@@ -615,6 +615,142 @@ void otx2_get_mac_from_af(struct net_device *netdev)
}
EXPORT_SYMBOL(otx2_get_mac_from_af);
+static int
+otx2_nix_tmq_reg_write(struct otx2_nic *pfvf, int cnt,
+ u64 reg_addr[MAX_REGS_PER_MBOX_MSG],
+ u64 reg_val[MAX_REGS_PER_MBOX_MSG])
+{
+ struct mbox *mbox = &pfvf->mbox;
+ struct nix_txschq_config *req;
+ int i, err;
+
+ mutex_lock(&mbox->lock);
+ req = otx2_mbox_alloc_msg_nix_txschq_cfg(mbox);
+ if (!req) {
+ mutex_unlock(&mbox->lock);
+ return -ENOMEM;
+ }
+
+ req->lvl = NIX_TXSCH_LVL_MDQ;
+ req->num_regs = cnt;
+
+ for (i = 0; i < cnt; i++) {
+ req->reg[i] = reg_addr[i];
+ req->regval[i] = reg_val[i];
+ }
+
+ err = otx2_sync_mbox_msg(mbox);
+ mutex_unlock(&mbox->lock);
+
+ return err;
+}
+
+int otx2_nix_tm_clear_queue_shaper(struct otx2_nic *pfvf)
+{
+ u64 reg_addr[MAX_REGS_PER_MBOX_MSG];
+ u64 reg_val[MAX_REGS_PER_MBOX_MSG];
+ int err, smq, i, cnt = 0;
+
+ for (i = 0; i < pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ]; i++) {
+ smq = pfvf->hw.txschq_list[NIX_TXSCH_LVL_SMQ][i];
+
+ reg_addr[cnt] = NIX_AF_MDQX_PIR(smq);
+ reg_val[cnt] = 0;
+ cnt++;
+
+ reg_addr[cnt] = NIX_AF_MDQX_CIR(smq);
+ reg_val[cnt] = 0;
+ cnt++;
+
+ if (cnt < MAX_REGS_PER_MBOX_MSG - 1)
+ continue;
+
+ err = otx2_nix_tmq_reg_write(pfvf, cnt,
+ reg_addr, reg_val);
+ if (err)
+ goto fail;
+ cnt = 0;
+ }
+
+ if (cnt) {
+ err = otx2_nix_tmq_reg_write(pfvf, cnt,
+ reg_addr, reg_val);
+ if (err)
+ goto fail;
+ }
+
+ return 0;
+fail:
+ return err;
+}
+
+int otx2_nix_tm_set_queue_shaper(struct otx2_nic *pfvf,
+ int txq, u64 minrate, u64 maxrate)
+{
+ struct mbox *mbox = &pfvf->mbox;
+ struct nix_txschq_config *req;
+ int err, smq, n = 0;
+ u64 reg_addr[2];
+ u64 reg_val[2];
+ u64 rate;
+
+ if (!maxrate && !minrate) {
+ smq = otx2_get_smq_idx(pfvf, txq);
+ reg_addr[0] = NIX_AF_MDQX_PIR(smq);
+ reg_val[0] = 0;
+ reg_addr[1] = NIX_AF_MDQX_CIR(smq);
+ reg_val[1] = 0;
+ return otx2_nix_tmq_reg_write(pfvf, 2, reg_addr, reg_val);
+ }
+
+ smq = otx2_get_smq_idx(pfvf, txq);
+
+ mutex_lock(&mbox->lock);
+ req = otx2_mbox_alloc_msg_nix_txschq_cfg(mbox);
+ if (!req) {
+ mutex_unlock(&mbox->lock);
+ return -ENOMEM;
+ }
+
+ req->lvl = NIX_TXSCH_LVL_MDQ;
+
+ /* MQPRIO exposes only min/max rate, not burst. Pass burst 0 so
+ * otx2_get_egress_burst_cfg() programmes the largest burst the NIX
+ * encoding supports (CN10K_MAX_BURST_SIZE on CN10K). This differs
+ * from the 65536 byte default used in the HTB path, which is a
+ * kernel-side default when no explicit burst is configured, not a
+ * hardware cap.
+ *
+ * mqprio setup restarts the netdev (otx2_mqprio_restart_netdev),
+ * which resets MDQ shapers to zero. Program both PIR and CIR on
+ * every update so omitted rates are applied explicitly rather than
+ * relying on stale hardware state.
+ */
+ req->reg[n] = NIX_AF_MDQX_PIR(smq);
+ if (maxrate) {
+ rate = otx2_convert_rate(maxrate);
+ req->regval[n] = otx2_get_txschq_rate_regval(pfvf, rate, 0);
+ } else {
+ req->regval[n] = 0;
+ }
+ n++;
+
+ /* CIR+PIR support is required and checked at mqprio setup. */
+ req->reg[n] = NIX_AF_MDQX_CIR(smq);
+ if (minrate) {
+ rate = otx2_convert_rate(minrate);
+ req->regval[n] = otx2_get_txschq_rate_regval(pfvf, rate, 0);
+ } else {
+ req->regval[n] = 0;
+ }
+ n++;
+ req->num_regs = n;
+
+ err = otx2_sync_mbox_msg(mbox);
+ mutex_unlock(&mbox->lock);
+ return err;
+}
+
int otx2_txschq_config(struct otx2_nic *pfvf, int lvl, int prio, bool txschq_for_pfc)
{
u16 (*schq_list)[MAX_TXSCHQ_PER_FUNC];
@@ -651,7 +787,11 @@ int otx2_txschq_config(struct otx2_nic *pfvf, int lvl, int prio, bool txschq_for
(u64)hw->smq_link_type);
req->num_regs++;
/* MDQ config */
- parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
+ if (pfvf->mqprio.rate_limit)
+ parent = schq_list[NIX_TXSCH_LVL_TL4][0];
+ else
+ parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
+
req->reg[1] = NIX_AF_MDQX_PARENT(schq);
req->regval[1] = parent << 16;
req->num_regs++;
@@ -779,6 +919,9 @@ int otx2_txsch_alloc(struct otx2_nic *pfvf)
req->schq[NIX_TXSCH_LVL_TL4] = chan_cnt;
}
+ if (pfvf->mqprio.rate_limit)
+ req->schq[NIX_TXSCH_LVL_SMQ] = pfvf->hw.non_qos_queues;
+
rc = otx2_sync_mbox_msg(&pfvf->mbox);
if (rc)
return rc;
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
index eecee612b7b2..ede7f1113b71 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
@@ -17,6 +17,7 @@
#include <linux/soc/marvell/silicons.h>
#include <linux/soc/marvell/octeontx2/asm.h>
#include <net/macsec.h>
+#include <uapi/linux/pkt_sched.h>
#include <net/pkt_cls.h>
#include <net/devlink.h>
#include <linux/time64.h>
@@ -483,6 +484,23 @@ struct pf_irq_data {
int mdevs;
};
+struct mq_offload_snap {
+ u64 min_rate[TC_QOPT_MAX_QUEUE];
+ u64 max_rate[TC_QOPT_MAX_QUEUE];
+ __u8 num_tc;
+ __u16 count[TC_QOPT_MAX_QUEUE];
+ __u16 offset[TC_QOPT_MAX_QUEUE];
+};
+
+struct otx2_mqprio {
+ u32 flags;
+ u64 *min_rate;
+ u64 *max_rate;
+ bool rate_limit;
+ bool replace_setup_done;
+ bool replace_graft_done;
+};
+
struct otx2_nic {
void __iomem *reg_base;
struct net_device *netdev;
@@ -515,6 +533,10 @@ struct otx2_nic {
u64 flags;
u64 *cq_op_addr;
+ struct otx2_mqprio mqprio;
+ struct mq_offload_snap *cur_mq_snap;
+ struct mq_offload_snap *old_mq_snap;
+
struct bpf_prog *xdp_prog;
struct otx2_qset qset;
struct otx2_hw hw;
@@ -1246,6 +1268,11 @@ dma_addr_t otx2_dma_map_skb_frag(struct otx2_nic *pfvf,
struct sk_buff *skb, int seg, int *len);
void otx2_dma_unmap_skb_frags(struct otx2_nic *pfvf, struct sg_list *sg);
int otx2_read_free_sqe(struct otx2_nic *pfvf, u16 qidx);
+int otx2_nix_tm_set_queue_shaper(struct otx2_nic *pfvf, int txq,
+ u64 minrate, u64 maxrate);
+int otx2_nix_tm_clear_queue_shaper(struct otx2_nic *pfvf);
+int otx2_mqprio_down(struct otx2_nic *pfvf);
+int otx2_mqprio_up(struct otx2_nic *pfvf);
void otx2_queue_vf_work(struct mbox *mw, struct workqueue_struct *mbox_wq,
int first, int mdevs, u64 intr);
int otx2_del_mcam_flow_entry(struct otx2_nic *nic, u16 entry,
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
index f110dfa42360..4a70abc230be 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
@@ -413,6 +413,12 @@ static int otx2_dcbnl_ieee_setpfc(struct net_device *dev, struct ieee_pfc *pfc)
u8 old_pfc_en;
int err;
+ if (pfvf->mqprio.rate_limit && pfc->pfc_en) {
+ netdev_err(dev,
+ "PFC: cannot enable while mqprio bandwidth offload is active\n");
+ return -EOPNOTSUPP;
+ }
+
old_pfc_en = pfvf->pfc_en;
pfvf->pfc_en = pfc->pfc_en;
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
index 9bee1b91eeaa..ec2601c6c255 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
@@ -287,6 +287,14 @@ static int otx2_set_channels(struct net_device *dev,
return -EINVAL;
}
+ if (pfvf->mqprio.rate_limit &&
+ (channel->tx_count != pfvf->hw.tx_queues ||
+ channel->rx_count != pfvf->hw.rx_queues)) {
+ netdev_info(dev,
+ "Not permitted to change channel count while MQ prio is active\n");
+ return -EINVAL;
+ }
+
if (if_up)
dev->netdev_ops->ndo_stop(dev);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
index c995f2900859..79a82e445454 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
@@ -1980,6 +1980,12 @@ int otx2_open(struct net_device *netdev)
if (err)
goto err_free_mem;
+ err = otx2_mqprio_up(pf);
+ if (err)
+ netdev_err(pf->netdev,
+ "mqprio: failed to restore shapers during open: %d; continuing without bandwidth limits\n",
+ err);
+
/* Register NAPI handler */
for (qidx = 0; qidx < pf->hw.cint_cnt; qidx++) {
cq_poll = &qset->napi[qidx];
@@ -2846,6 +2852,12 @@ static int otx2_xdp_setup(struct otx2_nic *pf, struct bpf_prog *prog)
bool if_up = netif_running(pf->netdev);
struct bpf_prog *old_prog;
+ if (prog && pf->mqprio.rate_limit) {
+ netdev_err(dev,
+ "XDP: cannot attach while mqprio bandwidth offload is active\n");
+ return -EOPNOTSUPP;
+ }
+
if (prog && dev->mtu > MAX_XDP_MTU) {
netdev_warn(dev, "Jumbo frames not yet supported with XDP\n");
return -EOPNOTSUPP;
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
index 039fd47ebf52..5efa0b4d27d5 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
@@ -16,6 +16,8 @@
#include <net/tc_act/tc_mirred.h>
#include <net/tc_act/tc_vlan.h>
#include <net/ipv6.h>
+#include <net/pkt_sched.h>
+#include <net/sch_generic.h>
#include "cn10k.h"
#include "otx2_common.h"
@@ -31,6 +33,10 @@
#define MCAST_INVALID_GRP (-1U)
#define RATE_MANTISSA_BITS 8
+/* Min per-queue egress shaping rate the NIX TLX encoder supports (2 Mbps). */
+#define OTX2_MQPRIO_MIN_RATE_BYTES_PS 250000ULL
+/* Max egress shaping rate the NIX TLX encoder supports (130816 Mbps). */
+#define OTX2_MQPRIO_MAX_RATE_BYTES_PS ((MAX_BURST_SIZE * 1000000ULL) / 8ULL)
static void otx2_get_egress_burst_cfg(struct otx2_nic *nic, u32 burst,
u32 *burst_exp, u32 *burst_mantissa)
@@ -61,6 +67,9 @@ static void otx2_get_egress_burst_cfg(struct otx2_nic *nic, u32 burst,
*burst_mantissa = tmp / (1ULL << (*burst_exp - 7));
}
} else {
+ /* burst 0: largest encodable burst (CN10K_MAX_BURST_SIZE on
+ * CN10K), not a minimal burst.
+ */
*burst_exp = MAX_BURST_EXPONENT;
*burst_mantissa = max_mantissa;
}
@@ -1600,14 +1609,706 @@ static int otx2_setup_tc_block(struct net_device *netdev,
nic, nic, ingress);
}
+/* Free the per-queue min/max rate caches. */
+static void otx2_mqprio_free_cache(struct otx2_nic *pfvf)
+{
+ devm_kfree(pfvf->dev, pfvf->mqprio.min_rate);
+ devm_kfree(pfvf->dev, pfvf->mqprio.max_rate);
+ pfvf->mqprio.min_rate = NULL;
+ pfvf->mqprio.max_rate = NULL;
+ pfvf->mqprio.flags = 0;
+}
+
+static int otx2_mqprio_alloc_cache(struct otx2_nic *pfvf, bool replacing)
+{
+ u16 num_txq = pfvf->hw.non_qos_queues;
+
+ if (replacing && pfvf->mqprio.min_rate && pfvf->mqprio.max_rate) {
+ memset(pfvf->mqprio.min_rate, 0,
+ num_txq * sizeof(*pfvf->mqprio.min_rate));
+ memset(pfvf->mqprio.max_rate, 0,
+ num_txq * sizeof(*pfvf->mqprio.max_rate));
+ pfvf->mqprio.flags = 0;
+ return 0;
+ }
+
+ otx2_mqprio_free_cache(pfvf);
+
+ pfvf->mqprio.min_rate = devm_kcalloc(pfvf->dev, num_txq,
+ sizeof(*pfvf->mqprio.min_rate),
+ GFP_KERNEL);
+ pfvf->mqprio.max_rate = devm_kcalloc(pfvf->dev, num_txq,
+ sizeof(*pfvf->mqprio.max_rate),
+ GFP_KERNEL);
+ if (!pfvf->mqprio.min_rate || !pfvf->mqprio.max_rate) {
+ otx2_mqprio_free_cache(pfvf);
+ return -ENOMEM;
+ }
+
+ return 0;
+}
+
+static void otx2_mqprio_snap_free(struct otx2_nic *pfvf,
+ struct mq_offload_snap **snap)
+{
+ if (!*snap)
+ return;
+
+ devm_kfree(pfvf->dev, *snap);
+ *snap = NULL;
+}
+
+static int otx2_mqprio_snap_copy(struct otx2_nic *pfvf,
+ struct mq_offload_snap **dst,
+ const struct tc_mqprio_qopt_offload *mqprio)
+{
+ const struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+ struct mq_offload_snap *snap;
+ int tc;
+
+ if (!*dst) {
+ snap = devm_kzalloc(pfvf->dev, sizeof(*snap), GFP_KERNEL);
+ if (!snap)
+ return -ENOMEM;
+ *dst = snap;
+ } else {
+ snap = *dst;
+ }
+
+ snap->num_tc = qopt->num_tc;
+ for (tc = 0; tc < TC_QOPT_MAX_QUEUE; tc++) {
+ snap->count[tc] = qopt->count[tc];
+ snap->offset[tc] = qopt->offset[tc];
+ snap->min_rate[tc] = 0;
+ snap->max_rate[tc] = 0;
+ }
+
+ for (tc = 0; tc < qopt->num_tc; tc++) {
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+ snap->min_rate[tc] = mqprio->min_rate[tc];
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+ snap->max_rate[tc] = mqprio->max_rate[tc];
+ }
+
+ return 0;
+}
+
+static int otx2_mqprio_stage_cur(struct otx2_nic *pfvf,
+ const struct tc_mqprio_qopt_offload *mqprio)
+{
+ return otx2_mqprio_snap_copy(pfvf, &pfvf->cur_mq_snap, mqprio);
+}
+
+static void otx2_mqprio_snap_commit(struct otx2_nic *pfvf)
+{
+ otx2_mqprio_snap_free(pfvf, &pfvf->old_mq_snap);
+ pfvf->old_mq_snap = pfvf->cur_mq_snap;
+ pfvf->cur_mq_snap = NULL;
+}
+
+static void otx2_mqprio_clear_replace_state(struct otx2_nic *pfvf)
+{
+ pfvf->mqprio.replace_setup_done = false;
+ pfvf->mqprio.replace_graft_done = false;
+}
+
+static bool otx2_mqprio_mdq_allocated(struct otx2_nic *pfvf)
+{
+ return pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_MDQ] != 0;
+}
+
+static int otx2_mqprio_restore_old(struct otx2_nic *pfvf)
+{
+ struct mq_offload_snap *snap = pfvf->old_mq_snap;
+ struct net_device *netdev = pfvf->netdev;
+ u16 num_txq = pfvf->hw.non_qos_queues;
+ int tc, txq, err;
+
+ if (!snap)
+ return 0;
+
+ err = otx2_mqprio_alloc_cache(pfvf, false);
+ if (err)
+ return err;
+
+ memset(pfvf->mqprio.min_rate, 0, num_txq * sizeof(*pfvf->mqprio.min_rate));
+ memset(pfvf->mqprio.max_rate, 0, num_txq * sizeof(*pfvf->mqprio.max_rate));
+ pfvf->mqprio.flags = 0;
+
+ for (tc = 0; tc < snap->num_tc; tc++) {
+ u64 min_rate = snap->min_rate[tc];
+ u64 max_rate = snap->max_rate[tc];
+
+ if (min_rate)
+ pfvf->mqprio.flags |= TC_MQPRIO_F_MIN_RATE;
+ if (max_rate)
+ pfvf->mqprio.flags |= TC_MQPRIO_F_MAX_RATE;
+
+ for (txq = snap->offset[tc];
+ txq < snap->offset[tc] + snap->count[tc]; txq++) {
+ pfvf->mqprio.min_rate[txq] = min_rate;
+ pfvf->mqprio.max_rate[txq] = max_rate;
+ }
+ }
+
+ netdev_set_num_tc(netdev, snap->num_tc);
+ for (tc = 0; tc < snap->num_tc; tc++)
+ netdev_set_tc_queue(netdev, tc, snap->count[tc],
+ snap->offset[tc]);
+
+ if (otx2_mqprio_mdq_allocated(pfvf)) {
+ err = otx2_nix_tm_clear_queue_shaper(pfvf);
+ if (err)
+ return err;
+ }
+
+ /* otx2_mqprio_restart_netdev() clears rate_limit when ndo_open() fails. */
+ pfvf->mqprio.rate_limit = true;
+
+ err = otx2_mqprio_up(pfvf);
+ if (err)
+ return err;
+
+ otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
+
+ return 0;
+}
+
+static void otx2_mqprio_snap_destroy(struct otx2_nic *pfvf)
+{
+ otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
+ otx2_mqprio_snap_free(pfvf, &pfvf->old_mq_snap);
+}
+
+static void otx2_mqprio_clear_sw(struct otx2_nic *pfvf)
+{
+ struct net_device *netdev = pfvf->netdev;
+
+ pfvf->mqprio.rate_limit = false;
+ otx2_mqprio_clear_replace_state(pfvf);
+ netdev_set_num_tc(netdev, 0);
+ otx2_mqprio_free_cache(pfvf);
+}
+
+/* Tear down mqprio bandwidth offload: clear per-queue shapers,
+ * mqprio_rate_limit, netdev TC mappings, and the cached rates. Called on
+ * explicit mqprio teardown (tc qdisc del) and error cleanup, not on
+ * routine netdev stop/open cycles where the offload stays active.
+ */
+int otx2_mqprio_down(struct otx2_nic *pfvf)
+{
+ int err = 0;
+
+ if (!pfvf->mqprio.rate_limit)
+ return 0;
+
+ if (netif_running(pfvf->netdev) &&
+ otx2_mqprio_mdq_allocated(pfvf))
+ err = otx2_nix_tm_clear_queue_shaper(pfvf);
+
+ if (err) {
+ netdev_err(pfvf->netdev,
+ "mqprio: failed to clear hardware shapers: %d; some TX queues may retain bandwidth limits\n",
+ err);
+ return err;
+ }
+
+ otx2_mqprio_clear_sw(pfvf);
+
+ return 0;
+}
+
+int otx2_mqprio_up(struct otx2_nic *pfvf)
+{
+ struct net_device *netdev = pfvf->netdev;
+ int txq, err;
+
+ if (!pfvf->mqprio.rate_limit)
+ return 0;
+
+ if (!pfvf->mqprio.min_rate || !pfvf->mqprio.max_rate)
+ return 0;
+
+ for (txq = 0; txq < pfvf->hw.non_qos_queues; txq++) {
+ u64 min_rate = 0, max_rate = 0;
+
+ if (pfvf->mqprio.flags & TC_MQPRIO_F_MIN_RATE)
+ min_rate = pfvf->mqprio.min_rate[txq];
+ if (pfvf->mqprio.flags & TC_MQPRIO_F_MAX_RATE)
+ max_rate = pfvf->mqprio.max_rate[txq];
+
+ if (!min_rate && !max_rate)
+ continue;
+
+ err = otx2_nix_tm_set_queue_shaper(pfvf, txq, min_rate,
+ max_rate);
+ if (err) {
+ netdev_err(netdev,
+ "mqprio: failed to restore shaper for txq %d: %d\n",
+ txq, err);
+ return err;
+ }
+ }
+
+ return 0;
+}
+
+/* Restart the netdev to reprogram the TX scheduler hierarchy for mqprio
+ * bandwidth offload. Both mqprio add and delete (when offload was active)
+ * take this path via ndo_stop()/ndo_open() so VF-specific open logic (e.g.
+ * LBK carrier on) runs correctly. The full stop/open cycle clears
+ * carrier, stops all TX queues, tears down IRQs/NAPI and drops in-flight
+ * traffic. If open fails, the interface is left administratively down
+ * without calling ndo_stop() again on resources already torn down by
+ * the open error path.
+ *
+ * Do not call dev_deactivate()/dev_activate() here: this runs from
+ * ndo_setup_tc() while qdisc_graft() may already hold the device
+ * deactivated and must perform the final dev_activate().
+ */
+static int otx2_mqprio_restart_netdev(struct net_device *netdev, bool rate_limit)
+{
+ struct otx2_nic *pfvf = netdev_priv(netdev);
+ const struct net_device_ops *ops = netdev->netdev_ops;
+ int err;
+
+ /* TODO: Explore live TX scheduler reprogramming to avoid a full
+ * ndo_stop()/ndo_open() bounce on every mqprio change.
+ */
+ netdev_info(netdev,
+ "mqprio: restarting interface to reprogram TX scheduler; in-flight traffic will be dropped\n");
+
+ err = ops->ndo_stop(netdev);
+ if (err)
+ return err;
+
+ /* Set before ndo_open() so otx2_txsch_alloc() widens SMQ allocation. */
+ if (rate_limit)
+ pfvf->mqprio.rate_limit = true;
+
+ err = ops->ndo_open(netdev);
+ if (err) {
+ netdev_err(netdev,
+ "Failed to restart device after mqprio change: %d\n",
+ err);
+ /* ndo_open() already freed the TX schedulers on failure while
+ * netif_running() may still be true; drop mqprio software state
+ * only instead of sending shaper clears to freed queues.
+ */
+ otx2_mqprio_clear_sw(pfvf);
+ /* ndo_open() rolls back on failure; mark the interface down so
+ * netif_close() does not invoke ndo_stop() on freed NAPI/queue
+ * state. Caller holds RTNL; dev_close() would deadlock.
+ */
+ pfvf->flags |= OTX2_FLAG_INTF_DOWN;
+ /* visible to otx2_stop() on other cpus */
+ smp_wmb();
+ netif_close(netdev);
+ }
+
+ return err;
+}
+
+static int otx2_mqprio_validate_tc_rate(struct net_device *netdev,
+ struct netlink_ext_ack *extack,
+ u64 rate, u32 qcount, int tc,
+ const char *name)
+{
+ if (!rate)
+ return 0;
+
+ if (qcount <= 1)
+ return 0;
+
+ /* TODO: mqprio min_rate/max_rate are per traffic class, but bandwidth
+ * offload shapes on per-queue MDQ nodes parented under a single TL4.
+ * Without per-TC TL4 shapers the driver cannot honor TC-level limits
+ * for a traffic class that spans multiple queues without either
+ * dividing the rate across queues (uAPI mismatch) or exceeding the TC
+ * cap when every member queue is active. Reject until per-TC TL4
+ * shaping can be implemented without allocating additional TL4 nodes
+ * beyond the existing hierarchy.
+ */
+ netdev_err(netdev,
+ "mqprio: %s rate for tc %d not supported with %u queues\n",
+ name, tc, qcount);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: %s rate for tc %d not supported with %u queues",
+ name, tc, qcount);
+ return -EOPNOTSUPP;
+}
+
+static int otx2_mqprio_validate_txqs(struct net_device *netdev,
+ struct netlink_ext_ack *extack,
+ struct tc_mqprio_qopt *qopt)
+{
+ struct otx2_nic *pfvf = netdev_priv(netdev);
+ u16 num_txq = pfvf->hw.non_qos_queues;
+ int tc, txq;
+
+ if (qopt->num_tc > num_txq) {
+ netdev_err(netdev, "Number of TCs (%u) exceeds hw queues %u\n",
+ qopt->num_tc, num_txq);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "Number of TCs (%u) exceeds hw queues %u",
+ qopt->num_tc, num_txq);
+ return -EINVAL;
+ }
+
+ if (num_txq > MAX_TXSCHQ_PER_FUNC) {
+ netdev_err(netdev,
+ "Number of queues (%u) exceeds max scheduler queues %u\n",
+ num_txq, MAX_TXSCHQ_PER_FUNC);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "Number of queues (%u) exceeds max scheduler queues %u",
+ num_txq, MAX_TXSCHQ_PER_FUNC);
+ return -EINVAL;
+ }
+
+ for (tc = 0; tc < qopt->num_tc; tc++) {
+ u32 qcount = qopt->count[tc];
+
+ for (txq = qopt->offset[tc];
+ txq < qopt->offset[tc] + qcount; txq++) {
+ if (txq >= num_txq) {
+ netdev_err(netdev,
+ "mqprio: txq %d exceeds offload queue count %u\n",
+ txq, num_txq);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: txq %d exceeds offload queue count %u",
+ txq, num_txq);
+ return -EINVAL;
+ }
+ }
+ }
+
+ return 0;
+}
+
+static bool otx2_mqprio_rate_valid(u64 rate_bytes_ps)
+{
+ u64 mbps;
+
+ if (!rate_bytes_ps)
+ return true;
+
+ if (rate_bytes_ps < OTX2_MQPRIO_MIN_RATE_BYTES_PS)
+ return false;
+
+ if (rate_bytes_ps > OTX2_MQPRIO_MAX_RATE_BYTES_PS)
+ return false;
+
+ if (rate_bytes_ps > div_u64(U64_MAX, 8))
+ return false;
+
+ mbps = otx2_convert_rate(rate_bytes_ps);
+ return ilog2(mbps / 2) <= MAX_RATE_EXPONENT;
+}
+
+static int otx2_teardown_tc_mqprio(struct otx2_nic *pfvf,
+ struct tc_mqprio_qopt_offload *mqprio)
+{
+ bool had_mqprio = pfvf->mqprio.rate_limit;
+ struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+ struct net_device *netdev = pfvf->netdev;
+ bool if_up = netif_running(netdev);
+ int err;
+
+ qopt->hw = 0;
+
+ /* tc qdisc replace runs setup on the new mqprio before destroying the
+ * old one. replace_setup_done and TC_ROOT_GRAFT distinguish stale
+ * old-instance teardown from graft failure after setup.
+ */
+ if (pfvf->mqprio.replace_setup_done && pfvf->cur_mq_snap) {
+ if (pfvf->mqprio.replace_graft_done) {
+ otx2_mqprio_snap_commit(pfvf);
+ } else {
+ err = otx2_mqprio_restore_old(pfvf);
+ if (err)
+ return err;
+ }
+ otx2_mqprio_clear_replace_state(pfvf);
+ return 0;
+ }
+
+ /* Skip the netdev restart when mqprio offload was not active. */
+ if (!had_mqprio)
+ return 0;
+
+ if (if_up) {
+ int down_err, err;
+
+ down_err = otx2_mqprio_down(pfvf);
+ err = otx2_mqprio_restart_netdev(netdev, false);
+ if (err)
+ return err;
+ return down_err;
+ }
+
+ /* ndo_stop() already freed the TX scheduler TL nodes; drop software
+ * state only.
+ */
+ otx2_mqprio_clear_sw(pfvf);
+ return 0;
+}
+
+static int otx2_setup_tc_mqprio(struct net_device *netdev,
+ struct tc_mqprio_qopt_offload *mqprio)
+{
+ struct otx2_nic *pfvf = netdev_priv(netdev);
+ struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+ struct netlink_ext_ack *extack = mqprio->extack;
+ bool replacing = pfvf->mqprio.rate_limit;
+ bool if_up = netif_running(netdev);
+ int tc, txq, err, i;
+
+ if (!qopt->hw)
+ return otx2_teardown_tc_mqprio(pfvf, mqprio);
+
+ if (!if_up) {
+ netdev_err(netdev, "mqprio: setup requires interface UP\n");
+ NL_SET_ERR_MSG_MOD(extack, "mqprio: setup requires interface UP");
+ return -EOPNOTSUPP;
+ }
+
+ if (mqprio->shaper != TC_MQPRIO_SHAPER_BW_RATE) {
+ netdev_err(netdev, "Unsupported mqprio shaper %#x\n", mqprio->shaper);
+ NL_SET_ERR_MSG_FMT_MOD(extack, "Unsupported mqprio shaper %#x",
+ mqprio->shaper);
+ return -EOPNOTSUPP;
+ }
+
+ if (!test_bit(QOS_CIR_PIR_SUPPORT, &pfvf->hw.cap_flag)) {
+ netdev_err(netdev,
+ "mqprio: bandwidth offload requires CIR+PIR support\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: bandwidth offload requires CIR+PIR support");
+ return -EOPNOTSUPP;
+ }
+
+ if (is_otx2_sdp_rep(pfvf->pdev)) {
+ netdev_err(netdev, "mqprio: bandwidth offload not supported on SDP rep\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: bandwidth offload not supported on SDP rep");
+ return -EOPNOTSUPP;
+ }
+
+ if (pfvf->pfc_en) {
+ netdev_err(netdev,
+ "mqprio: cannot enable offload while PFC is enabled\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: cannot enable offload while PFC is enabled");
+ return -EOPNOTSUPP;
+ }
+
+ if (pfvf->xdp_prog) {
+ netdev_err(netdev,
+ "mqprio: cannot enable offload while XDP is active\n");
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: cannot enable offload while XDP is active");
+ return -EOPNOTSUPP;
+ }
+
+ for (tc = 0; tc < qopt->num_tc; tc++) {
+ u64 min_rate = 0, max_rate = 0;
+ u32 qcount = qopt->count[tc];
+
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+ min_rate = mqprio->min_rate[tc];
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+ max_rate = mqprio->max_rate[tc];
+
+ if (min_rate && max_rate && min_rate > max_rate) {
+ netdev_err(netdev,
+ "min_rate %llu exceeds max_rate %llu for tc %d\n",
+ min_rate, max_rate, tc);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "min_rate %llu exceeds max_rate %llu for tc %d",
+ min_rate, max_rate, tc);
+ return -EINVAL;
+ }
+
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE) {
+ err = otx2_mqprio_validate_tc_rate(netdev, extack, min_rate,
+ qcount, tc, "min");
+ if (err)
+ return err;
+ }
+
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE) {
+ err = otx2_mqprio_validate_tc_rate(netdev, extack, max_rate,
+ qcount, tc, "max");
+ if (err)
+ return err;
+ }
+
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE &&
+ !otx2_mqprio_rate_valid(min_rate)) {
+ netdev_err(netdev,
+ "mqprio: min_rate %llu for tc %d is outside hardware limits\n",
+ min_rate, tc);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: min_rate %llu for tc %d is outside hardware limits",
+ min_rate, tc);
+ return -EINVAL;
+ }
+
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE &&
+ !otx2_mqprio_rate_valid(max_rate)) {
+ netdev_err(netdev,
+ "mqprio: max_rate %llu for tc %d is outside hardware limits\n",
+ max_rate, tc);
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: max_rate %llu for tc %d is outside hardware limits",
+ max_rate, tc);
+ return -EINVAL;
+ }
+ }
+
+ err = otx2_mqprio_validate_txqs(netdev, extack, qopt);
+ if (err)
+ return err;
+
+ err = otx2_mqprio_stage_cur(pfvf, mqprio);
+ if (err)
+ return err;
+
+ err = otx2_mqprio_restart_netdev(pfvf->netdev, true);
+ if (err)
+ goto cleanup;
+
+ err = otx2_mqprio_alloc_cache(pfvf, replacing);
+ if (err)
+ goto cleanup;
+
+ /* otx2_mqprio_up() may have restored the previous configuration during
+ * the restart above. Clear every MDQ shaper before applying the new
+ * mapping so queues dropped from the TC layout do not keep stale
+ * limits in hardware.
+ */
+ if (otx2_mqprio_mdq_allocated(pfvf)) {
+ err = otx2_nix_tm_clear_queue_shaper(pfvf);
+ if (err)
+ goto cleanup;
+ }
+
+ pfvf->mqprio.flags = mqprio->flags;
+
+ for (tc = 0; tc < qopt->num_tc; tc++) {
+ u64 min_rate = 0, max_rate = 0;
+ u32 qcount = qopt->count[tc];
+
+ /* Rates omitted from tc mqprio are passed as zero and both MDQ
+ * shaper registers are programmed; see
+ * otx2_nix_tm_set_queue_shaper(). Multi-queue TCs with rates
+ * are rejected above.
+ */
+ if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+ min_rate = mqprio->min_rate[tc];
+ if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+ max_rate = mqprio->max_rate[tc];
+
+ for (txq = qopt->offset[tc];
+ txq < qopt->offset[tc] + qcount; txq++) {
+ netdev_dbg(netdev,
+ "mqprio: tc %d txq %d min_rate %llu max_rate %llu\n",
+ tc, txq, min_rate, max_rate);
+
+ pfvf->mqprio.min_rate[txq] = min_rate;
+ pfvf->mqprio.max_rate[txq] = max_rate;
+
+ err = otx2_nix_tm_set_queue_shaper(pfvf, txq,
+ min_rate, max_rate);
+ if (err)
+ goto cleanup;
+ }
+ }
+
+ netdev_set_num_tc(netdev, pfvf->cur_mq_snap->num_tc);
+ for (i = 0; i < pfvf->cur_mq_snap->num_tc; i++)
+ netdev_set_tc_queue(netdev, i, pfvf->cur_mq_snap->count[i],
+ qopt->offset[i]);
+
+ qopt->hw = TC_MQPRIO_HW_OFFLOAD_TCS;
+
+ if (replacing) {
+ pfvf->mqprio.replace_setup_done = true;
+ pfvf->mqprio.replace_graft_done = false;
+ } else {
+ otx2_mqprio_snap_commit(pfvf);
+ }
+
+ return 0;
+
+cleanup:
+ qopt->hw = 0;
+ if (replacing) {
+ int restore_err = otx2_mqprio_restore_old(pfvf);
+
+ otx2_mqprio_clear_replace_state(pfvf);
+ if (restore_err) {
+ netdev_err(netdev,
+ "mqprio: replace failed and prior configuration rollback failed: %d\n",
+ restore_err);
+ if (extack)
+ NL_SET_ERR_MSG_FMT_MOD(extack,
+ "mqprio: replace failed and prior configuration rollback failed: %d",
+ restore_err);
+ } else {
+ netdev_err(netdev,
+ "mqprio: replace failed; prior configuration restored\n");
+ if (extack)
+ NL_SET_ERR_MSG_MOD(extack,
+ "mqprio: replace failed; prior configuration restored");
+ }
+ return err ? err : -EIO;
+ }
+ otx2_teardown_tc_mqprio(pfvf, mqprio);
+ return err;
+}
+
+static int otx2_setup_tc_root(struct otx2_nic *pfvf,
+ struct tc_root_qopt_offload *root)
+{
+ switch (root->command) {
+ case TC_ROOT_GRAFT:
+ if (pfvf->mqprio.replace_setup_done)
+ pfvf->mqprio.replace_graft_done = true;
+ return 0;
+ default:
+ return -EOPNOTSUPP;
+ }
+}
+
+static int otx2_setup_tc_query_caps(void *type_data)
+{
+ struct tc_query_caps_base *base = type_data;
+ struct tc_mqprio_caps *caps;
+
+ if (base->type != TC_SETUP_QDISC_MQPRIO)
+ return -EOPNOTSUPP;
+
+ caps = base->caps;
+ caps->validate_queue_counts = true;
+
+ return 0;
+}
+
int otx2_setup_tc(struct net_device *netdev, enum tc_setup_type type,
void *type_data)
{
switch (type) {
+ case TC_QUERY_CAPS:
+ return otx2_setup_tc_query_caps(type_data);
case TC_SETUP_BLOCK:
return otx2_setup_tc_block(netdev, type_data);
case TC_SETUP_QDISC_HTB:
return otx2_setup_tc_htb(netdev, type_data);
+ case TC_SETUP_QDISC_MQPRIO:
+ return otx2_setup_tc_mqprio(netdev, type_data);
+ case TC_SETUP_ROOT_QDISC:
+ return otx2_setup_tc_root(netdev_priv(netdev), type_data);
default:
return -EOPNOTSUPP;
}
@@ -1632,6 +2333,7 @@ EXPORT_SYMBOL(otx2_init_tc);
void otx2_shutdown_tc(struct otx2_nic *nic)
{
otx2_destroy_tc_flow_list(nic);
+ otx2_mqprio_snap_destroy(nic);
}
EXPORT_SYMBOL(otx2_shutdown_tc);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/qos.c b/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
index 69c0911e28e9..d9af182e5721 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
@@ -118,6 +118,9 @@ static void otx2_config_sched_shaping(struct otx2_nic *pfvf,
/* configure PIR */
maxrate = (node->rate > node->ceil) ? node->rate : node->ceil;
+ /* 65536 is the kernel-side default burst when HTB does not supply an
+ * explicit value, not the NIX hardware maximum (CN10K_MAX_BURST_SIZE).
+ */
cfg->regval[*num_regs] =
otx2_get_txschq_rate_regval(pfvf, maxrate, 65536);
(*num_regs)++;
--
2.43.0
^ permalink raw reply related [flat|nested] 3+ messages in thread* Re: [PATCH v11 net-next] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers
2026-09-02 1:55 [PATCH v11 net-next] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
@ 2026-09-03 2:29 ` Ratheesh Kannoth
2026-09-07 7:58 ` netdev-bot+sashiko
1 sibling, 0 replies; 3+ messages in thread
From: Ratheesh Kannoth @ 2026-09-03 2:29 UTC (permalink / raw)
To: bpf, linux-kernel, netdev
Cc: andrew+netdev, ast, daniel, davem, edumazet, hawk, john.fastabend,
kuba, pabeni, sdf, sgoutham
On 2026-09-02 at 07:25:00, Ratheesh Kannoth (rkannoth@marvell.com) wrote:
> Add TC_SETUP_QDISC_MQPRIO offload for channel-mode mqprio with
> TC_MQPRIO_SHAPER_BW_RATE. Program per-queue MDQ CIR/PIR through the
> NIX TX scheduler mailbox for each non-QoS transmit queue. When offload
> is active, allocate one SMQ per such queue, parent every MDQ under
> TL4[0], and map each traffic class min/max rate to the queue(s) in that
> class.
>
> Rebuild the TX scheduler hierarchy by bouncing the netdev through
> ndo_stop() and ndo_open() on mqprio add, replace, and delete. Cache
> the active rates and restore MDQ shapers from otx2_mqprio_up() during
> ndo_open(); log and continue if restoration fails so a normal open is
> not blocked.
>
> Track mqprio configuration in mq_offload_snap snapshots (TC layout and
> rates). On tc qdisc replace, stage the new configuration while keeping
> the previous snapshot for rollback: failed setup restores the old
> snapshot, successful graft is recorded through TC_ROOT_GRAFT, and
> teardown of the replaced qdisc instance commits the staged snapshot
> without tearing down the live offload.
>
> Reject offload unless the interface is running and the device advertises
> CIR+PIR support. Reject per-TC rates for traffic classes mapped to more
> than one queue, SDP rep ports, and concurrent PFC or XDP use. Block
> ethtool channel count changes while mqprio bandwidth offload is active.
>
> Add a ratelimited AF debug message when validating TX scheduler queue
> ownership to aid mqprio hierarchy setup failures.
>
> Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
Will address sashiko reviews in v12.
pw-bot: changes-requested
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH v11 net-next] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers
2026-09-02 1:55 [PATCH v11 net-next] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
2026-09-03 2:29 ` Ratheesh Kannoth
@ 2026-09-07 7:58 ` netdev-bot+sashiko
1 sibling, 0 replies; 3+ messages in thread
From: netdev-bot+sashiko @ 2026-09-07 7:58 UTC (permalink / raw)
To: rkannoth
Cc: bpf, linux-kernel, netdev, andrew+netdev, ast, daniel, davem,
edumazet, hawk, john.fastabend, kuba, pabeni, sdf, sgoutham
Thank you for your contribution! Sashiko AI review found 9 potential
issue(s) to consider:
Critical: 0 · High: 4 · Medium: 5 · Low: 0
- [High] The mqprio replace/rollback state machine depends on receiving
TC_SETUP_ROOT_QDISC/TC_ROOT_GRAFT, but that notification is gated on…
- [High] otx2_mqprio_restart_netdev() calls ndo_stop() and then
ndo_open(); if ndo_open() fails it sets OTX2_FLAG_INTF_DOWN and calls…
- [High] In otx2_teardown_tc_mqprio(), when otx2_mqprio_down() fails it
returns early *without* calling otx2_mqprio_clear_sw(), so…
- [High] In otx2_teardown_tc_mqprio(), when the replace-rollback branch
is taken and otx2_mqprio_restore_old() fails, the function returns…
- [Medium] The commit message claims 'failed setup restores the old
snapshot' and the driver logs 'mqprio: replace failed; prior…
- [Medium] The driver installs netdev TC state with netdev_set_num_tc() +
netdev_set_tc_queue() but tears it down with netdev_set_num_tc(netdev,…
- [Medium] otx2_mqprio_up() programs MDQ CIR/PIR with one mailbox
transaction per TX queue and returns on the first failure without…
- [Medium] otx2_mqprio_mdq_allocated() derives 'MDQ nodes exist' from
pfvf->hw.txschq_cnt[], but otx2_txschq_stop() only zeroes…
- [Medium] With bandwidth offload active the patch parents every per-
queue MDQ under a single TL4 (schq_list[NIX_TXSCH_LVL_TL4][0] in…
--
octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers
This adds TC_SETUP_QDISC_MQPRIO offload for channel-mode mqprio with
TC_MQPRIO_SHAPER_BW_RATE, programming per-queue MDQ CIR/PIR through the NIX
TX scheduler mailbox. The TX scheduler hierarchy is rebuilt by bouncing the
netdev through ndo_stop()/ndo_open() on mqprio add, replace and delete, and
the configuration is tracked in mq_offload_snap snapshots so a failed
replace can be rolled back.
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
> index 175992188c18a..5c502c9d7c83f 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
> @@ -615,6 +615,142 @@ void otx2_get_mac_from_af(struct net_device *netdev)
> }
> EXPORT_SYMBOL(otx2_get_mac_from_af);
>
> +static int
> +otx2_nix_tmq_reg_write(struct otx2_nic *pfvf, int cnt,
> + u64 reg_addr[MAX_REGS_PER_MBOX_MSG],
> + u64 reg_val[MAX_REGS_PER_MBOX_MSG])
> +{
[ ... ]
> +int otx2_nix_tm_clear_queue_shaper(struct otx2_nic *pfvf)
> +{
> + u64 reg_addr[MAX_REGS_PER_MBOX_MSG];
> + u64 reg_val[MAX_REGS_PER_MBOX_MSG];
> + int err, smq, i, cnt = 0;
> +
> + for (i = 0; i < pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ]; i++) {
> + smq = pfvf->hw.txschq_list[NIX_TXSCH_LVL_SMQ][i];
> +
> + reg_addr[cnt] = NIX_AF_MDQX_PIR(smq);
> + reg_val[cnt] = 0;
> + cnt++;
[Severity: Medium]
Is the txschq_cnt[] loop bound trustworthy after the TX schedulers have been
freed? otx2_txschq_stop() zeroes txschq_list[][] but leaves txschq_cnt[]
untouched:
drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c:otx2_txschq_stop() {
...
/* Clear the txschq list */
for (lvl = 0; lvl < NIX_TXSCH_LVL_CNT; lvl++) {
for (schq = 0; schq < MAX_TXSCHQ_PER_FUNC; schq++)
pfvf->hw.txschq_list[lvl][schq] = 0;
}
...
}
So otx2_mqprio_mdq_allocated() still returns true for freed nodes, and this
loop then runs txschq_cnt[SMQ] times with smq == 0 every iteration, writing
NIX_AF_MDQX_PIR(0)/NIX_AF_MDQX_CIR(0). otx2_get_smq_idx() has the same
problem via "qidx % pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ]" returning index 0
of the zeroed list.
The path reaching this: otx2_mqprio_restart_netdev() -> ndo_open() fails ->
otx2_open() error path -> otx2_free_hw_resources() -> otx2_txschq_stop(), then
back in otx2_setup_tc_mqprio() at "cleanup:" -> otx2_mqprio_restore_old(),
which calls otx2_nix_tm_clear_queue_shaper() with no netif_running() guard
(unlike otx2_mqprio_down()).
Normally the AF rejects schq 0 in is_valid_txschq() (firing the new
dev_err_ratelimited() added by this patch), and the rollback then reports a
self-inflicted "prior configuration rollback failed". But if schq 0 is mapped
back to this pcifunc as a PFC or HTB SMQ, would that unrelated queue's shaper
be overwritten with the mqprio rate? Should otx2_txschq_stop() also zero
txschq_cnt[]?
> +
> + reg_addr[cnt] = NIX_AF_MDQX_CIR(smq);
> + reg_val[cnt] = 0;
> + cnt++;
> +
> + if (cnt < MAX_REGS_PER_MBOX_MSG - 1)
> + continue;
> +
> + err = otx2_nix_tmq_reg_write(pfvf, cnt,
> + reg_addr, reg_val);
> + if (err)
> + goto fail;
> + cnt = 0;
> + }
[Severity: Medium]
This batch loop aborts on the first mailbox failure, leaving an unknown subset
of MDQ shapers still enabled with no way for the caller to know which.
otx2_mqprio_up() has the same shape - it returns on the first
otx2_nix_tm_set_queue_shaper() failure without undoing the shapers already
accepted by the AF.
Combined with the handling in otx2_open():
err = otx2_mqprio_up(pf);
if (err)
netdev_err(pf->netdev,
"mqprio: failed to restore shapers during open: %d; continuing without bandwidth limits\n",
err);
the device then enables packet I/O with low-numbered queues shaped and later
queues unshaped, while mqprio.rate_limit, the cached min_rate[]/max_rate[]
and the qdisc's TCA_HW_OFFLOAD state all claim the full configuration is in
force. Is the message "continuing without bandwidth limits" accurate here,
given limits are partially applied?
> +
> + if (cnt) {
> + err = otx2_nix_tmq_reg_write(pfvf, cnt,
> + reg_addr, reg_val);
> + if (err)
> + goto fail;
> + }
> +
> + return 0;
> +fail:
> + return err;
> +}
> +
[ ... ]
> @@ -651,7 +787,11 @@ int otx2_txschq_config(struct otx2_nic *pfvf, int lvl, int prio, bool txschq_for
> (u64)hw->smq_link_type);
> req->num_regs++;
> /* MDQ config */
> - parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
> + if (pfvf->mqprio.rate_limit)
> + parent = schq_list[NIX_TXSCH_LVL_TL4][0];
> + else
> + parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
> +
[Severity: Medium]
Parenting every MDQ under TL4[0] creates what looks like the first multi-child
TL4 in this non-QoS path, but nothing programs the parent's
NIX_AF_TL4X_TOPOLOGY (RR_PRIO / PRIO_ANCHOR). Each MDQ gets
NIX_AF_MDQX_SCHEDULE with only the DWRR quantum, so PRIO is 0; round-robin
sharing needs the parent's RR_PRIO to be 0 too.
Does the AF guarantee TL4 TOPOLOGY is zero on allocation? From what I can
see it does not - nix_reset_tx_schedule() only rewrites PARENT and SCHEDULE,
nix_reset_tx_shaping() only CIR/PIR, and the only TOPOLOGY write in
rvu_nix.c is NIX_AF_TL1X_TOPOLOGY. Meanwhile the HTB path does program TL4
TOPOLOGY with non-zero values:
drivers/net/ethernet/marvell/octeontx2/nic/qos.c:otx2_qos_txschq_set_parent_topology() {
...
if (parent->level == NIX_TXSCH_LVL_TL4)
cfg->reg[0] = NIX_AF_TL4X_TOPOLOGY(parent->schq);
...
cfg->regval[0] = (u64)parent->prio_anchor << 32;
cfg->regval[0] |= ((parent->child_dwrr_prio != OTX2_QOS_DEFAULT_PRIO) ?
parent->child_dwrr_prio : 0) << 1;
...
}
If a TL4 recycled from a prior HTB hierarchy carries RR_PRIO != 0, would all
the mqprio MDQs become static-priority children anchored at a stale
PRIO_ANCHOR, starving some TX queues despite correct CIR/PIR?
> req->reg[1] = NIX_AF_MDQX_PARENT(schq);
> req->regval[1] = parent << 16;
> req->num_regs++;
[ ... ]
> @@ -779,6 +919,9 @@ int otx2_txsch_alloc(struct otx2_nic *pfvf)
> req->schq[NIX_TXSCH_LVL_TL4] = chan_cnt;
> }
>
> + if (pfvf->mqprio.rate_limit)
> + req->schq[NIX_TXSCH_LVL_SMQ] = pfvf->hw.non_qos_queues;
> +
> rc = otx2_sync_mbox_msg(&pfvf->mbox);
> if (rc)
> return rc;
[Severity: High]
Widening the SMQ request from 1 to hw.non_qos_queues makes reopen failure much
more likely, and the AF rejects the whole request rather than allocating fewer:
drivers/net/ethernet/marvell/octeontx2/af/rvu_nix.c:nix_check_txschq_alloc_req() {
...
if (free_cnt < req_schq || req->schq[lvl] > MAX_TXSCHQ_PER_FUNC ||
req->schq_contig[lvl] > MAX_TXSCHQ_PER_FUNC)
return NIX_AF_ERR_TLX_ALLOC_FAIL;
...
}
On silicon with hw->cap.nix_fixed_txschq_mapping the reported free_cnt for the
SMQ level is at most 1, so any request > 1 always fails there. The SMQ pool is
also shared across all PFs/VFs of the NIX block, so exhaustion is a normal
runtime condition.
When that reopen fails, otx2_mqprio_restart_netdev() has already run ndo_stop()
and its error path does:
pfvf->flags |= OTX2_FLAG_INTF_DOWN;
/* visible to otx2_stop() on other cpus */
smp_wmb();
netif_close(netdev);
netif_close() clears IFF_UP, so an ordinary "tc qdisc add/replace/del dev X
root mqprio" can leave the interface administratively down with connectivity
lost until a manual "ip link set up". The commit message says only
"The full stop/open cycle clears carrier, stops all TX queues, tears
down IRQs/NAPI and drops in-flight traffic."
and
"in-flight traffic will be dropped"
Should the commit message also state that a failed mqprio add/replace, or even
a successful delete, can leave the netdev down?
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> index eecee612b7b2c..ede7f1113b714 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> @@ -483,6 +484,23 @@ struct pf_irq_data {
> int mdevs;
> };
>
> +struct mq_offload_snap {
> + u64 min_rate[TC_QOPT_MAX_QUEUE];
> + u64 max_rate[TC_QOPT_MAX_QUEUE];
> + __u8 num_tc;
> + __u16 count[TC_QOPT_MAX_QUEUE];
> + __u16 offset[TC_QOPT_MAX_QUEUE];
> +};
[Severity: Medium]
Should this snapshot also record qopt->prio_tc_map[]? The core installs the
new priority mapping right after the driver's setup succeeds:
net/sched/sch_mqprio.c:mqprio_init() {
...
/* Always use supplied priority mappings */
for (i = 0; i < TC_BITMASK + 1; i++)
netdev_set_prio_tc_map(dev, i, qopt->prio_tc_map[i]);
...
}
otx2_mqprio_restore_old() then restores only num_tc/count/offset, so a rollback
after that point leaves the new configuration's prio_tc_map in effect over the
old configuration's num_tc and tc_to_txq. skb_tx_hash() reads
dev->tc_to_txq[netdev_get_prio_tc_map(dev, prio)], so packets can be steered
into ranges that disagree with the MDQ shapers just reprogrammed, or into a tc
with count 0.
Relatedly, otx2_mqprio_clear_sw() tears down with netdev_set_num_tc(netdev, 0)
rather than netdev_reset_tc(). netdev_set_num_tc() only resets XPS and
sb-channels and writes dev->num_tc, so tc_to_txq[] and prio_tc_map[] keep stale
contents on the device. Would netdev_reset_tc() be the right pairing here?
> +
> +struct otx2_mqprio {
> + u32 flags;
> + u64 *min_rate;
> + u64 *max_rate;
> + bool rate_limit;
> + bool replace_setup_done;
> + bool replace_graft_done;
> +};
> +
[ ... ]
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> index c995f29008590..79a82e4454546 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> @@ -1980,6 +1980,12 @@ int otx2_open(struct net_device *netdev)
> if (err)
> goto err_free_mem;
>
> + err = otx2_mqprio_up(pf);
> + if (err)
> + netdev_err(pf->netdev,
> + "mqprio: failed to restore shapers during open: %d; continuing without bandwidth limits\n",
> + err);
> +
[ ... ]
> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> index 039fd47ebf522..5efa0b4d27d53 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> @@ -1600,14 +1609,706 @@ static int otx2_setup_tc_block(struct net_device *netdev,
[ ... ]
> +static void otx2_mqprio_clear_sw(struct otx2_nic *pfvf)
> +{
> + struct net_device *netdev = pfvf->netdev;
> +
> + pfvf->mqprio.rate_limit = false;
> + otx2_mqprio_clear_replace_state(pfvf);
> + netdev_set_num_tc(netdev, 0);
> + otx2_mqprio_free_cache(pfvf);
> +}
> +
[ ... ]
> +int otx2_mqprio_down(struct otx2_nic *pfvf)
> +{
> + int err = 0;
> +
> + if (!pfvf->mqprio.rate_limit)
> + return 0;
> +
> + if (netif_running(pfvf->netdev) &&
> + otx2_mqprio_mdq_allocated(pfvf))
> + err = otx2_nix_tm_clear_queue_shaper(pfvf);
> +
> + if (err) {
> + netdev_err(pfvf->netdev,
> + "mqprio: failed to clear hardware shapers: %d; some TX queues may retain bandwidth limits\n",
> + err);
> + return err;
> + }
> +
> + otx2_mqprio_clear_sw(pfvf);
> +
> + return 0;
> +}
[Severity: High]
This early return skips otx2_mqprio_clear_sw(), so mqprio.rate_limit stays
true and the rate caches stay populated. The caller then restarts anyway:
if (if_up) {
int down_err, err;
down_err = otx2_mqprio_down(pfvf);
err = otx2_mqprio_restart_netdev(netdev, false);
and otx2_mqprio_restart_netdev() only ever sets rate_limit, never clears it.
So ndo_open() sees rate_limit == true and:
- otx2_txsch_alloc() re-requests the wide mqprio SMQ hierarchy
- otx2_txschq_config() re-parents every MDQ under TL4[0]
- otx2_open() -> otx2_mqprio_up() re-programs every cached MDQ shaper
Meanwhile the qdisc goes away regardless, because mqprio_disable_offload()
discards the driver's return value:
net/sched/sch_mqprio.c:mqprio_disable_offload() {
...
dev->netdev_ops->ndo_setup_tc(dev, TC_SETUP_QDISC_MQPRIO,
&mqprio);
...
}
Does that leave the offload latched on with no qdisc behind it - dev->num_tc
still populated, and the new rate_limit gates in otx2_set_channels(),
otx2_xdp_setup() and otx2_dcbnl_ieee_setpfc() refusing those operations until
the driver is unbound? Should the software state be cleared even when the
shaper clear fails?
> +
> +int otx2_mqprio_up(struct otx2_nic *pfvf)
> +{
[ ... ]
> + for (txq = 0; txq < pfvf->hw.non_qos_queues; txq++) {
[ ... ]
> + err = otx2_nix_tm_set_queue_shaper(pfvf, txq, min_rate,
> + max_rate);
> + if (err) {
> + netdev_err(netdev,
> + "mqprio: failed to restore shaper for txq %d: %d\n",
> + txq, err);
> + return err;
> + }
> + }
> +
> + return 0;
> +}
> +
[ ... ]
> +static int otx2_teardown_tc_mqprio(struct otx2_nic *pfvf,
> + struct tc_mqprio_qopt_offload *mqprio)
> +{
[ ... ]
> + if (pfvf->mqprio.replace_setup_done && pfvf->cur_mq_snap) {
> + if (pfvf->mqprio.replace_graft_done) {
> + otx2_mqprio_snap_commit(pfvf);
> + } else {
> + err = otx2_mqprio_restore_old(pfvf);
> + if (err)
> + return err;
> + }
> + otx2_mqprio_clear_replace_state(pfvf);
> + return 0;
> + }
[Severity: High]
When otx2_mqprio_restore_old() fails, this "if (err) return err;" bypasses
otx2_mqprio_clear_replace_state() and never frees pfvf->cur_mq_snap.
replace_setup_done stays true and cur_mq_snap stays non-NULL, so does every
later teardown - including a genuine "tc qdisc del" - re-enter this branch
instead of tearing the offload down? rate_limit, the devm min_rate/max_rate
caches, the netdev TC mapping and the hardware MDQ shapers would then only be
released at device remove.
otx2_mqprio_restore_old() can fail from otx2_mqprio_alloc_cache() (-ENOMEM),
otx2_nix_tm_clear_queue_shaper() or otx2_mqprio_up() (mailbox errors).
The "cleanup:" label in otx2_setup_tc_mqprio() calls
otx2_mqprio_clear_replace_state() before inspecting restore_err; would the
same ordering work here?
[Severity: High]
Is the replace_graft_done signal ever delivered in the default configuration?
qdisc_offload_graft_helper() is gated on tc_can_offload():
net/sched/sch_api.c:qdisc_offload_graft_helper() {
...
if (!tc_can_offload(dev) || !dev->netdev_ops->ndo_setup_tc)
return;
...
}
tc_can_offload() tests dev->features & NETIF_F_HW_TC, but otx2_probe() adds
that bit to hw_features only after the last features assignment:
drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c:otx2_probe() {
...
netdev->features |= netdev->hw_features;
...
/* HW supports tc offload but mutually exclusive with n-tuple filters */
if (pf->flags & OTX2_FLAG_TC_FLOWER_SUPPORT)
netdev->hw_features |= NETIF_F_HW_TC;
...
}
so hw-tc-offload is off in dev->features by default (otx2_vf.c has the same
ordering). The mqprio offload itself has no such gate - mqprio_enable_offload()
calls ndo_setup_tc() unconditionally.
With that, otx2_setup_tc_root() is never invoked, replace_graft_done stays
false, and a *successful* "tc qdisc replace dev X root mqprio ... hw 1" over an
already-offloaded mqprio takes the else branch above and calls
otx2_mqprio_restore_old(). Would that silently revert the hardware MDQ CIR/PIR
values and the netdev TC layout to the previous configuration while the new
qdisc is live and userspace was told the replace succeeded?
The commit message states:
"successful graft is recorded through TC_ROOT_GRAFT"
Can that mechanism work without NETIF_F_HW_TC enabled?
> +
> + /* Skip the netdev restart when mqprio offload was not active. */
> + if (!had_mqprio)
> + return 0;
> +
> + if (if_up) {
> + int down_err, err;
> +
> + down_err = otx2_mqprio_down(pfvf);
> + err = otx2_mqprio_restart_netdev(netdev, false);
> + if (err)
> + return err;
> + return down_err;
> + }
[ ... ]
> +static int otx2_setup_tc_mqprio(struct net_device *netdev,
> + struct tc_mqprio_qopt_offload *mqprio)
> +{
[ ... ]
> +cleanup:
> + qopt->hw = 0;
> + if (replacing) {
> + int restore_err = otx2_mqprio_restore_old(pfvf);
> +
> + otx2_mqprio_clear_replace_state(pfvf);
> + if (restore_err) {
> + netdev_err(netdev,
> + "mqprio: replace failed and prior configuration rollback failed: %d\n",
> + restore_err);
[Severity: Medium]
Does the netdev TC layout restored here actually survive? Returning an error
from otx2_setup_tc_mqprio() makes mqprio_enable_offload() return before setting
priv->hw_offload, mqprio_init() fails, and qdisc_create() still runs destroy:
net/sched/sch_api.c:qdisc_create() {
...
err_out4:
/* Even if ops->init() failed, we call ops->destroy()
* like qdisc_create_dflt().
*/
if (ops->destroy)
ops->destroy(sch);
...
}
net/sched/sch_mqprio.c:mqprio_destroy() {
...
if (priv->hw_offload && dev->netdev_ops->ndo_setup_tc)
mqprio_disable_offload(sch);
else
netdev_set_num_tc(dev, 0);
...
}
With hw_offload == 0 the else branch runs, so the netdev_set_num_tc() and
netdev_set_tc_queue() calls that otx2_mqprio_restore_old() just made are wiped.
The old, still-grafted mqprio qdisc keeps reporting its TCs and the restored
per-queue MDQ shapers stay programmed, but dev->num_tc becomes 0 - so TX queue
selection stops honouring the TC ranges while per-queue rate limits remain in
force.
The commit message says "failed setup restores the old snapshot" and the log
says "prior configuration restored". Is that accurate for the netdev TC part?
> + if (extack)
> + NL_SET_ERR_MSG_FMT_MOD(extack,
> + "mqprio: replace failed and prior configuration rollback failed: %d",
> + restore_err);
> + } else {
> + netdev_err(netdev,
> + "mqprio: replace failed; prior configuration restored\n");
[ ... ]
> +static int otx2_setup_tc_root(struct otx2_nic *pfvf,
> + struct tc_root_qopt_offload *root)
> +{
> + switch (root->command) {
> + case TC_ROOT_GRAFT:
> + if (pfvf->mqprio.replace_setup_done)
> + pfvf->mqprio.replace_graft_done = true;
> + return 0;
> + default:
> + return -EOPNOTSUPP;
> + }
> +}
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260902015500.2985371-1-rkannoth%40marvell.com
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-07 7:58 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-02 1:55 [PATCH v11 net-next] octeontx2-pf: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
2026-09-03 2:29 ` Ratheesh Kannoth
2026-09-07 7:58 ` netdev-bot+sashiko
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox