Netdev List
 help / color / mirror / Atom feed
* [PATCH v19 net-next 0/2] octeontx2: mqprio bandwidth offload for NIX TX schedulers
@ 2026-10-06  9:43 Ratheesh Kannoth
  2026-10-06  9:43 ` [PATCH v19 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags Ratheesh Kannoth
  2026-10-06  9:43 ` [PATCH v19 net-next 2/2] octeontx2: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
  0 siblings, 2 replies; 5+ messages in thread
From: Ratheesh Kannoth @ 2026-10-06  9:43 UTC (permalink / raw)
  To: bpf, linux-kernel, netdev
  Cc: andrew+netdev, ast, daniel, davem, edumazet, hawk, john.fastabend,
	kuba, pabeni, sdf, sgoutham, Ratheesh Kannoth

This series adds hardware offload for channel-mode mqprio with
TC_MQPRIO_SHAPER_BW_RATE on Marvell octeontx2/cn10k PF and VF RVU
netdevices. Each non-QoS transmit queue is shaped by programming MDQ CIR/PIR
on the NIX TX scheduler. When bandwidth offload is active, the driver
allocates one SMQ per queue, parents every MDQ under TL4[0], and maps each
traffic class min/max rate to the queue(s) in that class.

The NIX TX scheduler hierarchy cannot be reprogrammed live today, so
mqprio add, replace, delete, and failed-replace rollback rebuild it by
bouncing the netdev through ndo_stop()/ndo_open(). That intentionally
drops in-flight traffic on each change. otx2_mqprio_restart_netdev() clears
__LINK_STATE_START before ndo_stop() and does not call
dev_deactivate()/dev_activate(); carrier and TX queues are restored after
ndo_open() via the normal link-event path when the link is up. Cache the
active rates and restore MDQ shapers from otx2_mqprio_up() during ndo_open();
fail closed if restoration fails, leaving ndo_open() unsuccessful and the
interface administratively down.

Track mqprio configuration in mq_offload_snap snapshots (TC layout and
rates). On tc qdisc replace, stage the new configuration while keeping
the previous snapshot for rollback: failed setup restores the old
snapshot via netdev restart when the interface is running, successful
graft is recorded through TC_ROOT_GRAFT, and teardown of the replaced
qdisc instance commits the staged snapshot without tearing down the live
offload.

Patch 1 converts PF/VF and representor flag access to atomic bitops.
Patch 2 depends on it for safe OTX2_FLAG_INTF_DOWN and OTX2_FLAG_PORT_UP
updates on asynchronous mbox paths and during the mqprio netdev bounce.

The driver rejects offload unless the interface is running and the device
advertises CIR+PIR support. PF and VF RVU netdevices share the same TC
offload path via ndo_setup_tc / otx2_open(); SDP representors are not
supported. Per-TC rates are rejected when a traffic class spans more than
one queue. Concurrent PFC, XDP, SDP rep, or HTB use is blocked, and ethtool
channel count changes are blocked while mqprio bandwidth offload is active.

Ratheesh Kannoth (2):
  octeontx2: use atomic bitops for PF/VF and rep flags
  octeontx2: add mqprio bandwidth offload for NIX TX schedulers

 .../ethernet/marvell/octeontx2/nic/cn10k_ipsec.c   |   8 +-
 .../ethernet/marvell/octeontx2/nic/otx2_common.c   | 154 +++-
 .../ethernet/marvell/octeontx2/nic/otx2_common.h   | 107 ++-
 .../ethernet/marvell/octeontx2/nic/otx2_dcbnl.c    |   6 +
 .../ethernet/marvell/octeontx2/nic/otx2_devlink.c  |   2 +-
 .../ethernet/marvell/octeontx2/nic/otx2_ethtool.c  |  29 +-
 .../ethernet/marvell/octeontx2/nic/otx2_flows.c    |  34 +-
 .../net/ethernet/marvell/octeontx2/nic/otx2_pf.c   |  95 ++-
 .../net/ethernet/marvell/octeontx2/nic/otx2_tc.c   | 838 ++++++++++++++++++++-
 .../net/ethernet/marvell/octeontx2/nic/otx2_txrx.c |  16 +-
 .../net/ethernet/marvell/octeontx2/nic/otx2_vf.c   |  10 +-
 .../net/ethernet/marvell/octeontx2/nic/otx2_xsk.c  |   4 +-
 drivers/net/ethernet/marvell/octeontx2/nic/qos.c   |  11 +
 .../net/ethernet/marvell/octeontx2/nic/qos_sq.c    |   4 +-
 drivers/net/ethernet/marvell/octeontx2/nic/rep.c   |  32 +-
 drivers/net/ethernet/marvell/octeontx2/nic/rep.h   |   3 +-
 16 files changed, 1203 insertions(+), 150 deletions(-)

---

v18 -> v19: Addressed sashiko and review nits on v18.
- Gate mqprio netdev_tc_work with defer_tc_work: set the flag before
  cancel_work_sync() in otx2_shutdown_tc_mqprio() and skip scheduling or
  running the deferred TC restore work once teardown starts (Sashiko #1).
= Address Jakub comment.
- Clarify the otx2_mqprio_restart_netdev() comment on failed ndo_open():
  __LINK_STATE_START stays clear until open succeeds, so netif_running() is
  false on that path
	https://lore.kernel.org/netdev/20260929022915.2704627-1-rkannoth@marvell.com/

v17 -> v18: Addressed sashiko comments on v17.
- Commit staged mqprio replace snapshots when TC_ROOT_GRAFT is skipped
  because hw-tc-offload is off in netdev->features (!tc_can_offload()),
  instead of rolling back a replace that actually succeeded.
- Keep NETIF_F_HW_TC in netdev->hw_features only (PF and VF); gate mqprio
  setup/teardown on otx2_tc_can_offload() so hw-tc-offload stays opt-in via
  ethtool -K and flower/matchall are not pushed to ndo_setup_tc by default.
- Split TC shutdown around netdev unregister: cancel mqprio deferred work and
  free mqprio snapshots before unregister_netdev(), but destroy the TC flower
  flow list only after unregister so clsact teardown can still run
  otx2_tc_del_flow() and free MCAM/mcast/policer state.
- Reject ethtool -L TX queue reduction while mqprio offload snapshots remain,
  so a later replace rollback cannot restore stale per-queue rates past the
  current queue count.
- Clarify mqprio shaper and netdev-restart comments: ndo_open() may restore
  cached MDQ shapers via otx2_mqprio_up() before setup clears them, and a
  failed mqprio restart relies on OTX2_FLAG_INTF_DOWN so otx2_stop() returns
  early through netif_close(), not on skipping ndo_stop().
	https://lore.kernel.org/netdev/20260923032217.1732753-1-rkannoth@marvell.com/

v16 -> v17: Addressed sashiko comments on v16.
- Replace the per-bit otx2_sync_flags_from_rep() loop with a masked
  READ_ONCE/WRITE_ONCE publish of OTX2_REP_SYNC_FLAGS_MASK so lockless NAPI
  readers never observe torn PF/representor flag combinations.
- Evaluate mqprio.rate_limit and old_mq_snap inside rtnl_lock in
  otx2_mqprio_netdev_tc_work() so a concurrent qdisc delete cannot leave
  stale netdev TC mappings after offload teardown.
- Advertise NETIF_F_HW_TC in netdev->features (PF and VF) when TC flower
  offload is supported, so tc_can_offload() succeeds without ethtool -K
  hw-tc-offload on; move otx2_init_tc() before register_netdev() and fix
  probe/remove teardown ordering.
- Reject mqprio add when a software mqprio root is already installed
  (otx2_mqprio_keep_netdev_tc()) and defer netdev TC restore from a new
  fail_validate path on failed replace validation before any hardware
  change.
- Fix otx2_mqprio_max_rate_bytes_ps() to cap against the NIX TLX maximum
  rate instead of the burst-bucket size; use the 65536 byte HTB default
  burst when programming MDQ shapers; guard otx2_get_smq_idx() when
  txschq_cnt[NIX_TXSCH_LVL_SMQ] is zero after otx2_txschq_stop().
- Rename patch 2 to octeontx2: (driver-wide PF/VF offload, not PF-only).
	https://lore.kernel.org/netdev/20260918015906.1255204-1-rkannoth@marvell.com/

v15 -> v16: Addressed sashiko comments on v15 and aligned documentation with code.
- Sync representor flags through OTX2_FLAG_MAX in otx2_sync_flags_from_rep()
  instead of hard-coding OTX2_REP_VF_INITIALIZED as the loop bound.
- Drop the rvu_nix.c is_valid_txschq() ratelimited error print from the mqprio
  patch; remove the misplaced atomic-bitops and AF-debug paragraphs from the
  mqprio commit message (they belong to patch 1 or are out of scope).
- Extend mq_offload_snap to record prio_tc_map[] and mqprio rate flags; restore
  the full netdev TC layout (num_tc, queue ranges, and priority map) via
  otx2_mqprio_apply_snap_netdev() on rollback paths.
- Defer netdev TC restore on failed replace (otx2_mqprio_netdev_tc_work) so
  rollback survives mqprio_destroy() clearing dev->num_tc after setup errors
  once the core unwinds the failed qdisc instance.
- Stop calling dev_deactivate()/dev_activate() from otx2_mqprio_restart_netdev();
  bounce the interface with ndo_stop()/ndo_open() only and restore carrier
  through the normal link-event path after ndo_open(), avoiding qdisc
  reentrancy during tc replace graft.
- Preserve netdev TC mappings when tearing down an offloaded instance that is
  replaced by a software mqprio graft (otx2_mqprio_keep_netdev_tc()) instead
  of always calling netdev_set_num_tc(0) and breaking the live replacement.
- Return an error from otx2_mqprio_down() when clearing hardware shapers fails
  and keep offload software state, instead of v15's behaviour of clearing
  rate_limit while stale MDQ limits may remain programmed.
- Rebuild the TX scheduler via netdev restart in otx2_mqprio_restore_old() on
  a running interface after failed-replace rollback so partially applied MDQ
  shapers are not left running with mismatched software state.
- Clear txschq_cnt[] in otx2_txschq_stop() after freeing scheduler nodes so
  post-stop shaper mailbox operations do not consult stale counts.
- Document fail-closed ndo_open() when otx2_mqprio_up() cannot restore shapers,
  and that PF/VF RVU netdevices share the ndo_setup_tc / otx2_open() offload
  path (SDP representors remain unsupported); downgrade the mqprio restart
  notice to netdev_dbg().
	https://lore.kernel.org/netdev/20260911105521.689565-1-rkannoth@marvell.com/

v14 -> v15: Addressed sashiko comments.
- Split atomic PF/VF and representor flag access into a preparatory patch
  so mqprio netdev-restart and mbox paths can update OTX2_FLAG_INTF_DOWN
  and OTX2_FLAG_PORT_UP without data races on the shared flags word.
- Clear mqprio software state when hardware shaper teardown fails, warn,
  and still bounce the netdev on delete so offload does not remain stuck
  active after a mailbox error.
	https://lore.kernel.org/netdev/20260904031553.3196916-1-rkannoth@marvell.com/

v13 -> v14: Addressed sashiko comments.
- Use atomic set_bit()/clear_bit() for OTX2_FLAG_INTF_DOWN and
  OTX2_FLAG_PORT_UP updates on netdev-restart and mbox paths.
- Block concurrent mqprio bandwidth offload and HTB shaping.
- Fail ndo_open() if otx2_mqprio_up() cannot restore MDQ shapers.
- Rebuild the TX scheduler via netdev restart in otx2_mqprio_restore_old()
  when rolling back a failed replace on a running interface.
- Return an error from otx2_mqprio_down() if clearing hardware shapers
  fails instead of clearing software state anyway.
	https://lore.kernel.org/netdev/20260904031553.3196916-1-rkannoth@marvell.com/

v12 -> v13: Addressed sashiko comments.
	https://sashiko.dev/#/patchset/20260903023324.3078284-1-rkannoth%40marvell.com

v11 -> v12: Addressed sashiko comments.
	https://sashiko.dev/#/patchset/20260902015500.2985371-1-rkannoth%40marvell.com
v10 -> v11: Addressed sashiko comments.
	https://sashiko.dev/#/patchset/20260831131014.2639581-1-rkannoth%40marvell.com

v9 -> v10: Addressed sashiko/jacub comments.
	https://sashiko.dev/#/message/20260817032747.1765883-1-rkannoth%40marvell.com

v8 -> v9: Addressed Sashiko comments
	https://lore.kernel.org/netdev/aoJ6FhtWue0FHDQV@rkannoth-OptiPlex-7090/
v7 -> v8: Addressed Sashiko comments
	https://sashiko.dev/#/patchset/20260811085050.3212280-1-rkannoth%40marvell.com
v6 -> v7: Addressed Sashiko comments
	https://sashiko.dev/#/message/20260810034738.1786029-1-rkannoth%40marvell.com
v5 -> v6: Addressed Sashiko comments
	https://lore.kernel.org/netdev/20260806095434.1144397-1-rkannoth@marvell.com/
v4 -> v5: Addressed sashiko comments
	https://sashiko.dev/#/patchset/20260803042724.3380209-1-rkannoth%40marvell.com
v3 -> v4: Addressed sashiko comments
	https://lore.kernel.org/netdev/20260729105139.2302908-1-rkannoth@marvell.com/
v2 -> v3: Addressed sashiko comments
	https://lore.kernel.org/netdev/amnYX866mYx02cBe@rkannoth-OptiPlex-7090/T/#m67310cbec48b21c7720858ab3a1ea083a0f8dc10
v1 -> v2: Addressed sashiko comments
	https://lore.kernel.org/netdev/20260724075010.2665758-1-rkannoth@marvell.com/

--
2.43.0

^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v19 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags
  2026-10-06  9:43 [PATCH v19 net-next 0/2] octeontx2: mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
@ 2026-10-06  9:43 ` Ratheesh Kannoth
  2026-10-10 10:01   ` netdev-bot+sashiko
  2026-10-06  9:43 ` [PATCH v19 net-next 2/2] octeontx2: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
  1 sibling, 1 reply; 5+ messages in thread
From: Ratheesh Kannoth @ 2026-10-06  9:43 UTC (permalink / raw)
  To: bpf, linux-kernel, netdev
  Cc: andrew+netdev, ast, daniel, davem, edumazet, hawk, john.fastabend,
	kuba, pabeni, sdf, sgoutham, Ratheesh Kannoth

Switch PF/VF and representor flags from non-atomic u64 masks to unsigned
long bitmaps with set_bit/clear_bit/test_bit.  Add helpers and
otx2_sync_flags_from_rep() to publish representor-owned flags onto the PF
mailbox context with a masked WRITE_ONCE, giving lockless NAPI readers a
consistent word without clearing PF-owned bits.

Relocate representor VF initialization to OTX2_FLAG_REP_VF_INITIALIZED
(bit 21).

Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
---
 .../marvell/octeontx2/nic/cn10k_ipsec.c       |  8 +-
 .../marvell/octeontx2/nic/otx2_common.c       |  8 +-
 .../marvell/octeontx2/nic/otx2_common.h       | 84 ++++++++++++++-----
 .../marvell/octeontx2/nic/otx2_devlink.c      |  2 +-
 .../marvell/octeontx2/nic/otx2_ethtool.c      | 21 ++---
 .../marvell/octeontx2/nic/otx2_flows.c        | 34 ++++----
 .../ethernet/marvell/octeontx2/nic/otx2_pf.c  | 78 ++++++++---------
 .../ethernet/marvell/octeontx2/nic/otx2_tc.c  | 30 +++----
 .../marvell/octeontx2/nic/otx2_txrx.c         | 16 ++--
 .../ethernet/marvell/octeontx2/nic/otx2_vf.c  | 10 +--
 .../ethernet/marvell/octeontx2/nic/otx2_xsk.c |  4 +-
 .../ethernet/marvell/octeontx2/nic/qos_sq.c   |  4 +-
 .../net/ethernet/marvell/octeontx2/nic/rep.c  | 32 +++----
 .../net/ethernet/marvell/octeontx2/nic/rep.h  |  3 +-
 14 files changed, 185 insertions(+), 149 deletions(-)

diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/cn10k_ipsec.c b/drivers/net/ethernet/marvell/octeontx2/nic/cn10k_ipsec.c
index 77543d472345..50ec4542c418 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/cn10k_ipsec.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/cn10k_ipsec.c
@@ -334,7 +334,7 @@ static int cn10k_outb_cpt_init(struct net_device *netdev)
 						CN10K_CPT_LF_NQX(0));
 
 	/* Set ipsec offload enabled for this device */
-	pf->flags |= OTX2_FLAG_IPSEC_OFFLOAD_ENABLED;
+	otx2_set_flag(pf, OTX2_FLAG_IPSEC_OFFLOAD_ENABLED);
 
 	cn10k_cpt_device_set_available(pf);
 	return 0;
@@ -356,7 +356,7 @@ static int cn10k_outb_cpt_clean(struct otx2_nic *pf)
 	}
 
 	/* Set ipsec offload disabled for this device */
-	pf->flags &= ~OTX2_FLAG_IPSEC_OFFLOAD_ENABLED;
+	otx2_clear_flag(pf, OTX2_FLAG_IPSEC_OFFLOAD_ENABLED);
 
 	/* Disable CPTLF Instruction Queue (IQ) */
 	cn10k_outb_cptlf_iq_disable(pf);
@@ -820,7 +820,7 @@ void cn10k_ipsec_clean(struct otx2_nic *pf)
 	if (!is_dev_support_ipsec_offload(pf->pdev))
 		return;
 
-	if (!(pf->flags & OTX2_FLAG_IPSEC_OFFLOAD_ENABLED))
+	if (!otx2_test_flag(pf, OTX2_FLAG_IPSEC_OFFLOAD_ENABLED))
 		return;
 
 	if (pf->ipsec.sa_workq) {
@@ -945,7 +945,7 @@ bool cn10k_ipsec_transmit(struct otx2_nic *pf, struct netdev_queue *txq,
 	u16 dlen;
 
 	/* Check for IPSEC offload enabled */
-	if (!(pf->flags & OTX2_FLAG_IPSEC_OFFLOAD_ENABLED))
+	if (!otx2_test_flag(pf, OTX2_FLAG_IPSEC_OFFLOAD_ENABLED))
 		goto drop;
 
 	sp = skb_sec_path(skb);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
index 77c5cfbe4ce4..836601fc10d8 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
@@ -220,10 +220,10 @@ int otx2_set_mac_address(struct net_device *netdev, void *p)
 		eth_hw_addr_set(netdev, addr->sa_data);
 		/* update dmac field in vlan offload rule */
 		if (netif_running(netdev) &&
-		    pfvf->flags & OTX2_FLAG_RX_VLAN_SUPPORT)
+		    otx2_test_flag(pfvf, OTX2_FLAG_RX_VLAN_SUPPORT))
 			otx2_install_rxvlan_offload_flow(pfvf);
 		/* update dmac address in ntuple and DMAC filter list */
-		if (pfvf->flags & OTX2_FLAG_DMACFLTR_SUPPORT)
+		if (otx2_test_flag(pfvf, OTX2_FLAG_DMACFLTR_SUPPORT))
 			otx2_dmacflt_update_pfmac_flow(pfvf);
 	} else {
 		return -EPERM;
@@ -275,8 +275,8 @@ int otx2_config_pause_frm(struct otx2_nic *pfvf)
 		goto unlock;
 	}
 
-	req->rx_pause = !!(pfvf->flags & OTX2_FLAG_RX_PAUSE_ENABLED);
-	req->tx_pause = !!(pfvf->flags & OTX2_FLAG_TX_PAUSE_ENABLED);
+	req->rx_pause = otx2_test_flag(pfvf, OTX2_FLAG_RX_PAUSE_ENABLED);
+	req->tx_pause = otx2_test_flag(pfvf, OTX2_FLAG_TX_PAUSE_ENABLED);
 	req->set = 1;
 
 	err = otx2_sync_mbox_msg(&pfvf->mbox);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
index fda6c4487c4f..3bc57dd10451 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
@@ -482,6 +482,41 @@ struct pf_irq_data {
 	int mdevs;
 };
 
+enum otx2_flag_bits {
+	OTX2_FLAG_RX_TSTAMP_ENABLED,
+	OTX2_FLAG_TX_TSTAMP_ENABLED,
+	OTX2_FLAG_INTF_DOWN,
+	OTX2_FLAG_MCAM_ENTRIES_ALLOC,
+	OTX2_FLAG_NTUPLE_SUPPORT,
+	OTX2_FLAG_UCAST_FLTR_SUPPORT,
+	OTX2_FLAG_RX_VLAN_SUPPORT,
+	OTX2_FLAG_VF_VLAN_SUPPORT,
+	OTX2_FLAG_PF_SHUTDOWN,
+	OTX2_FLAG_RX_PAUSE_ENABLED,
+	OTX2_FLAG_TX_PAUSE_ENABLED,
+	OTX2_FLAG_TC_FLOWER_SUPPORT,
+	OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED,
+	OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED,
+	OTX2_FLAG_DMACFLTR_SUPPORT,
+	OTX2_FLAG_PTP_ONESTEP_SYNC,
+	OTX2_FLAG_ADPTV_INT_COAL_ENABLED,
+	OTX2_FLAG_TC_MARK_ENABLED,
+	OTX2_FLAG_REP_MODE_ENABLED,
+	OTX2_FLAG_PORT_UP,
+	OTX2_FLAG_IPSEC_OFFLOAD_ENABLED,
+	OTX2_FLAG_REP_VF_INITIALIZED,
+	OTX2_FLAG_MAX,
+};
+
+/* Representor-owned flags copied onto the PF mailbox context in
+ * rvu_rep_setup_tc_cb().  All other bits are owned by the PF/VF netdev.
+ */
+#define OTX2_REP_SYNC_FLAGS_MASK				\
+	(BIT(OTX2_FLAG_MCAM_ENTRIES_ALLOC) |			\
+	 BIT(OTX2_FLAG_NTUPLE_SUPPORT) |			\
+	 BIT(OTX2_FLAG_TC_FLOWER_SUPPORT) |			\
+	 BIT(OTX2_FLAG_REP_VF_INITIALIZED))
+
 struct otx2_nic {
 	void __iomem		*reg_base;
 	struct net_device	*netdev;
@@ -490,28 +525,7 @@ struct otx2_nic {
 	u16			tx_max_pktlen;
 	u16			rbsize; /* Receive buffer size */
 
-#define OTX2_FLAG_RX_TSTAMP_ENABLED		BIT_ULL(0)
-#define OTX2_FLAG_TX_TSTAMP_ENABLED		BIT_ULL(1)
-#define OTX2_FLAG_INTF_DOWN			BIT_ULL(2)
-#define OTX2_FLAG_MCAM_ENTRIES_ALLOC		BIT_ULL(3)
-#define OTX2_FLAG_NTUPLE_SUPPORT		BIT_ULL(4)
-#define OTX2_FLAG_UCAST_FLTR_SUPPORT		BIT_ULL(5)
-#define OTX2_FLAG_RX_VLAN_SUPPORT		BIT_ULL(6)
-#define OTX2_FLAG_VF_VLAN_SUPPORT		BIT_ULL(7)
-#define OTX2_FLAG_PF_SHUTDOWN			BIT_ULL(8)
-#define OTX2_FLAG_RX_PAUSE_ENABLED		BIT_ULL(9)
-#define OTX2_FLAG_TX_PAUSE_ENABLED		BIT_ULL(10)
-#define OTX2_FLAG_TC_FLOWER_SUPPORT		BIT_ULL(11)
-#define OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED	BIT_ULL(12)
-#define OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED	BIT_ULL(13)
-#define OTX2_FLAG_DMACFLTR_SUPPORT		BIT_ULL(14)
-#define OTX2_FLAG_PTP_ONESTEP_SYNC		BIT_ULL(15)
-#define OTX2_FLAG_ADPTV_INT_COAL_ENABLED BIT_ULL(16)
-#define OTX2_FLAG_TC_MARK_ENABLED		BIT_ULL(17)
-#define OTX2_FLAG_REP_MODE_ENABLED		 BIT_ULL(18)
-#define OTX2_FLAG_PORT_UP			BIT_ULL(19)
-#define OTX2_FLAG_IPSEC_OFFLOAD_ENABLED		BIT_ULL(20)
-	u64			flags;
+	unsigned long		flags;
 	u64			*cq_op_addr;
 
 	struct bpf_prog		*xdp_prog;
@@ -593,6 +607,32 @@ struct otx2_nic {
 	unsigned long		*af_xdp_zc_qidx;
 };
 
+static inline void otx2_set_flag(struct otx2_nic *nic, unsigned int flag)
+{
+	set_bit(flag, &nic->flags);
+}
+
+static inline void otx2_clear_flag(struct otx2_nic *nic, unsigned int flag)
+{
+	clear_bit(flag, &nic->flags);
+}
+
+static inline bool otx2_test_flag(struct otx2_nic *nic, unsigned int flag)
+{
+	return test_bit(flag, &nic->flags);
+}
+
+static inline void otx2_sync_flags_from_rep(struct otx2_nic *dst,
+					    unsigned long *src_flags)
+{
+	unsigned long src = READ_ONCE(*src_flags);
+	unsigned long new_flags;
+
+	new_flags = (READ_ONCE(dst->flags) & ~OTX2_REP_SYNC_FLAGS_MASK) |
+		    (src & OTX2_REP_SYNC_FLAGS_MASK);
+	WRITE_ONCE(dst->flags, new_flags);
+}
+
 static inline bool is_otx2_lbkvf(struct pci_dev *pdev)
 {
 	return (pdev->device == PCI_DEVID_OCTEONTX2_RVU_AFVF) ||
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_devlink.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_devlink.c
index 4a5ce0e67dda..863a5ced9a26 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_devlink.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_devlink.c
@@ -104,7 +104,7 @@ static int otx2_dl_ucast_flt_cnt_validate(struct devlink *devlink, u32 id,
 	struct otx2_nic *pfvf = otx2_dl->pfvf;
 
 	/* Check for UNICAST filter support*/
-	if (!(pfvf->flags & OTX2_FLAG_UCAST_FLTR_SUPPORT)) {
+	if (!otx2_test_flag(pfvf, OTX2_FLAG_UCAST_FLTR_SUPPORT)) {
 		NL_SET_ERR_MSG_MOD(extack,
 				   "Unicast filter not enabled");
 		return -EINVAL;
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
index a05dee0085a3..4fe473d9ea0d 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
@@ -354,14 +354,14 @@ static int otx2_set_pauseparam(struct net_device *netdev,
 		return -EOPNOTSUPP;
 
 	if (pause->rx_pause)
-		pfvf->flags |= OTX2_FLAG_RX_PAUSE_ENABLED;
+		otx2_set_flag(pfvf, OTX2_FLAG_RX_PAUSE_ENABLED);
 	else
-		pfvf->flags &= ~OTX2_FLAG_RX_PAUSE_ENABLED;
+		otx2_clear_flag(pfvf, OTX2_FLAG_RX_PAUSE_ENABLED);
 
 	if (pause->tx_pause)
-		pfvf->flags |= OTX2_FLAG_TX_PAUSE_ENABLED;
+		otx2_set_flag(pfvf, OTX2_FLAG_TX_PAUSE_ENABLED);
 	else
-		pfvf->flags &= ~OTX2_FLAG_TX_PAUSE_ENABLED;
+		otx2_clear_flag(pfvf, OTX2_FLAG_TX_PAUSE_ENABLED);
 
 	return otx2_config_pause_frm(pfvf);
 }
@@ -470,8 +470,7 @@ static int otx2_get_coalesce(struct net_device *netdev,
 	cmd->rx_max_coalesced_frames = hw->cq_ecount_wait;
 	cmd->tx_coalesce_usecs = hw->cq_time_wait;
 	cmd->tx_max_coalesced_frames = hw->cq_ecount_wait;
-	if ((pfvf->flags & OTX2_FLAG_ADPTV_INT_COAL_ENABLED) ==
-			OTX2_FLAG_ADPTV_INT_COAL_ENABLED) {
+	if (otx2_test_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED)) {
 		cmd->use_adaptive_rx_coalesce = 1;
 		cmd->use_adaptive_tx_coalesce = 1;
 	} else {
@@ -502,15 +501,14 @@ static int otx2_set_coalesce(struct net_device *netdev,
 	}
 
 	/* Check and update coalesce status */
-	if ((pfvf->flags & OTX2_FLAG_ADPTV_INT_COAL_ENABLED) ==
-			OTX2_FLAG_ADPTV_INT_COAL_ENABLED) {
+	if (otx2_test_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED)) {
 		priv_coalesce_status = 1;
 		if (!ec->use_adaptive_rx_coalesce)
-			pfvf->flags &= ~OTX2_FLAG_ADPTV_INT_COAL_ENABLED;
+			otx2_clear_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED);
 	} else {
 		priv_coalesce_status = 0;
 		if (ec->use_adaptive_rx_coalesce)
-			pfvf->flags |= OTX2_FLAG_ADPTV_INT_COAL_ENABLED;
+			otx2_set_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED);
 	}
 
 	/* 'cq_time_wait' is 8bit and is in multiple of 100ns,
@@ -556,8 +554,7 @@ static int otx2_set_coalesce(struct net_device *netdev,
 	 * 'on' to 'off'.
 	 */
 	if (priv_coalesce_status &&
-	    ((pfvf->flags & OTX2_FLAG_ADPTV_INT_COAL_ENABLED) !=
-	     OTX2_FLAG_ADPTV_INT_COAL_ENABLED)) {
+	    (!otx2_test_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED))) {
 		hw->cq_time_wait = CQ_TIMER_THRESH_DEFAULT;
 		hw->cq_ecount_wait = CQ_CQE_THRESH_DEFAULT;
 	}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_flows.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_flows.c
index 6d095023d17b..0d8576ef822f 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_flows.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_flows.c
@@ -270,9 +270,9 @@ int otx2_alloc_mcam_entries(struct otx2_nic *pfvf, u16 count)
 	flow_cfg->max_flows = allocated;
 
 	if (allocated) {
-		pfvf->flags |= OTX2_FLAG_MCAM_ENTRIES_ALLOC;
-		pfvf->flags |= OTX2_FLAG_NTUPLE_SUPPORT;
-		pfvf->flags |= OTX2_FLAG_TC_FLOWER_SUPPORT;
+		otx2_set_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC);
+		otx2_set_flag(pfvf, OTX2_FLAG_NTUPLE_SUPPORT);
+		otx2_set_flag(pfvf, OTX2_FLAG_TC_FLOWER_SUPPORT);
 	}
 
 	if (allocated != count)
@@ -375,7 +375,7 @@ int otx2_mcam_entry_init(struct otx2_nic *pfvf)
 	flow_cfg->unicast_offset = vf_vlan_max_flows;
 	flow_cfg->rx_vlan_offset = flow_cfg->unicast_offset +
 					flow_cfg->ucast_flt_cnt;
-	pfvf->flags |= OTX2_FLAG_UCAST_FLTR_SUPPORT;
+	otx2_set_flag(pfvf, OTX2_FLAG_UCAST_FLTR_SUPPORT);
 
 	/* Check if NPC_DMAC field is supported
 	 * by the mkex profile before setting VLAN support flag.
@@ -400,11 +400,11 @@ int otx2_mcam_entry_init(struct otx2_nic *pfvf)
 	}
 
 	if (frsp->enable) {
-		pfvf->flags |= OTX2_FLAG_RX_VLAN_SUPPORT;
-		pfvf->flags |= OTX2_FLAG_VF_VLAN_SUPPORT;
+		otx2_set_flag(pfvf, OTX2_FLAG_RX_VLAN_SUPPORT);
+		otx2_set_flag(pfvf, OTX2_FLAG_VF_VLAN_SUPPORT);
 	}
 
-	pfvf->flags |= OTX2_FLAG_MCAM_ENTRIES_ALLOC;
+	otx2_set_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC);
 	mutex_unlock(&pfvf->mbox.lock);
 
 	/* Allocate entries for Ntuple filters */
@@ -414,7 +414,7 @@ int otx2_mcam_entry_init(struct otx2_nic *pfvf)
 		return 0;
 	}
 
-	pfvf->flags |= OTX2_FLAG_TC_FLOWER_SUPPORT;
+	otx2_set_flag(pfvf, OTX2_FLAG_TC_FLOWER_SUPPORT);
 
 	refcount_set(&flow_cfg->mark_flows, 1);
 	return 0;
@@ -477,7 +477,7 @@ int otx2_mcam_flow_init(struct otx2_nic *pf)
 		return err;
 
 	/* Check if MCAM entries are allocate or not */
-	if (!(pf->flags & OTX2_FLAG_UCAST_FLTR_SUPPORT))
+	if (!otx2_test_flag(pf, OTX2_FLAG_UCAST_FLTR_SUPPORT))
 		return 0;
 
 	pf->mac_table = devm_kzalloc(pf->dev, sizeof(struct otx2_mac_table)
@@ -499,7 +499,7 @@ int otx2_mcam_flow_init(struct otx2_nic *pf)
 	if (!pf->flow_cfg->bmap_to_dmacindex)
 		return -ENOMEM;
 
-	pf->flags |= OTX2_FLAG_DMACFLTR_SUPPORT;
+	otx2_set_flag(pf, OTX2_FLAG_DMACFLTR_SUPPORT);
 
 	return 0;
 }
@@ -519,7 +519,7 @@ static int otx2_do_add_macfilter(struct otx2_nic *pf, const u8 *mac)
 	struct npc_install_flow_req *req;
 	int err, i;
 
-	if (!(pf->flags & OTX2_FLAG_UCAST_FLTR_SUPPORT))
+	if (!otx2_test_flag(pf, OTX2_FLAG_UCAST_FLTR_SUPPORT))
 		return -ENOMEM;
 
 	/* dont have free mcam entries or uc list is greater than alloted */
@@ -1164,7 +1164,7 @@ static int otx2_is_flow_rule_dmacfilter(struct otx2_nic *pfvf,
 	u64 ring_cookie = fsp->ring_cookie;
 	u32 flow_type;
 
-	if (!(pfvf->flags & OTX2_FLAG_DMACFLTR_SUPPORT))
+	if (!otx2_test_flag(pfvf, OTX2_FLAG_DMACFLTR_SUPPORT))
 		return false;
 
 	flow_type = fsp->flow_type & ~(FLOW_EXT | FLOW_MAC_EXT | FLOW_RSS);
@@ -1361,7 +1361,7 @@ int otx2_add_flow(struct otx2_nic *pfvf, struct ethtool_rxnfc *nfc)
 	}
 
 	ring = ethtool_get_flow_spec_ring(fsp->ring_cookie);
-	if (!(pfvf->flags & OTX2_FLAG_NTUPLE_SUPPORT))
+	if (!otx2_test_flag(pfvf, OTX2_FLAG_NTUPLE_SUPPORT))
 		return -ENOMEM;
 
 	/* Number of queues on a VF can be greater or less than
@@ -1593,7 +1593,7 @@ int otx2_destroy_ntuple_flows(struct otx2_nic *pfvf)
 	struct otx2_flow *iter, *tmp;
 	int err;
 
-	if (!(pfvf->flags & OTX2_FLAG_NTUPLE_SUPPORT))
+	if (!otx2_test_flag(pfvf, OTX2_FLAG_NTUPLE_SUPPORT))
 		return 0;
 
 	if (!flow_cfg->max_flows)
@@ -1626,7 +1626,7 @@ int otx2_destroy_mcam_flows(struct otx2_nic *pfvf)
 	struct otx2_flow *iter, *tmp;
 	int err;
 
-	if (!(pfvf->flags & OTX2_FLAG_MCAM_ENTRIES_ALLOC))
+	if (!otx2_test_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC))
 		return 0;
 
 	/* remove all flows */
@@ -1655,7 +1655,7 @@ int otx2_destroy_mcam_flows(struct otx2_nic *pfvf)
 		return err;
 	}
 
-	pfvf->flags &= ~OTX2_FLAG_MCAM_ENTRIES_ALLOC;
+	otx2_clear_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC);
 	flow_cfg->max_flows = 0;
 	mutex_unlock(&pfvf->mbox.lock);
 
@@ -1718,7 +1718,7 @@ int otx2_enable_rxvlan(struct otx2_nic *pf, bool enable)
 	int err;
 
 	/* Dont have enough mcam entries */
-	if (!(pf->flags & OTX2_FLAG_RX_VLAN_SUPPORT))
+	if (!otx2_test_flag(pf, OTX2_FLAG_RX_VLAN_SUPPORT))
 		return -ENOMEM;
 
 	if (enable) {
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
index 7dc0ce669d0b..777e7156badb 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
@@ -879,7 +879,7 @@ static void otx2_handle_link_event(struct otx2_nic *pf)
 	struct cgx_link_user_info *linfo = &pf->linfo;
 	struct net_device *netdev = pf->netdev;
 
-	if (pf->flags & OTX2_FLAG_PORT_UP)
+	if (otx2_test_flag(pf, OTX2_FLAG_PORT_UP))
 		return;
 
 	pr_info("%s NIC Link is %s %d Mbps %s duplex\n", netdev->name,
@@ -907,11 +907,11 @@ static int otx2_mbox_up_handler_rep_event_up_notify(struct otx2_nic *pf,
 
 	if (info->event == RVU_EVENT_PORT_STATE) {
 		if (info->evt_data.port_state) {
-			pf->flags |= OTX2_FLAG_PORT_UP;
+			otx2_set_flag(pf, OTX2_FLAG_PORT_UP);
 			netif_carrier_on(netdev);
 			netif_tx_start_all_queues(netdev);
 		} else {
-			pf->flags &= ~OTX2_FLAG_PORT_UP;
+			otx2_clear_flag(pf, OTX2_FLAG_PORT_UP);
 			netif_tx_stop_all_queues(netdev);
 			netif_carrier_off(netdev);
 		}
@@ -953,7 +953,7 @@ int otx2_mbox_up_handler_cgx_link_event(struct otx2_nic *pf,
 	}
 
 	/* interface has not been fully configured yet */
-	if (pf->flags & OTX2_FLAG_INTF_DOWN)
+	if (otx2_test_flag(pf, OTX2_FLAG_INTF_DOWN))
 		return 0;
 
 	otx2_handle_link_event(pf);
@@ -1829,7 +1829,7 @@ void otx2_free_hw_resources(struct otx2_nic *pf)
 	free_req = otx2_mbox_alloc_msg_nix_lf_free(mbox);
 	if (free_req) {
 		free_req->flags = NIX_LF_DISABLE_FLOWS | NIX_LF_DONT_FREE_DFT_IDXS;
-		if (!(pf->flags & OTX2_FLAG_PF_SHUTDOWN))
+		if (!otx2_test_flag(pf, OTX2_FLAG_PF_SHUTDOWN))
 			free_req->flags |= NIX_LF_DONT_FREE_TX_VTAG;
 		if (otx2_sync_mbox_msg(mbox))
 			dev_err(pf->dev, "%s failed to free nixlf\n", __func__);
@@ -2136,21 +2136,21 @@ int otx2_open(struct net_device *netdev)
 	}
 	otx2_write64(pf, NIX_LF_RAS_ENA_W1S, NIX_LF_RAS_MASK);
 
-	if (pf->flags & OTX2_FLAG_RX_VLAN_SUPPORT)
+	if (otx2_test_flag(pf, OTX2_FLAG_RX_VLAN_SUPPORT))
 		otx2_enable_rxvlan(pf, true);
 
 	/* When reinitializing enable time stamping if it is enabled before */
-	if (pf->flags & OTX2_FLAG_TX_TSTAMP_ENABLED) {
-		pf->flags &= ~OTX2_FLAG_TX_TSTAMP_ENABLED;
+	if (otx2_test_flag(pf, OTX2_FLAG_TX_TSTAMP_ENABLED)) {
+		otx2_clear_flag(pf, OTX2_FLAG_TX_TSTAMP_ENABLED);
 		otx2_config_hw_tx_tstamp(pf, true);
 	}
-	if (pf->flags & OTX2_FLAG_RX_TSTAMP_ENABLED) {
-		pf->flags &= ~OTX2_FLAG_RX_TSTAMP_ENABLED;
+	if (otx2_test_flag(pf, OTX2_FLAG_RX_TSTAMP_ENABLED)) {
+		otx2_clear_flag(pf, OTX2_FLAG_RX_TSTAMP_ENABLED);
 		otx2_config_hw_rx_tstamp(pf, true);
 	}
 
-	pf->flags &= ~OTX2_FLAG_INTF_DOWN;
-	pf->flags &= ~OTX2_FLAG_PORT_UP;
+	otx2_clear_flag(pf, OTX2_FLAG_INTF_DOWN);
+	otx2_clear_flag(pf, OTX2_FLAG_PORT_UP);
 	/* 'intf_down' may be checked on any cpu */
 	smp_wmb();
 
@@ -2162,7 +2162,7 @@ int otx2_open(struct net_device *netdev)
 		otx2_handle_link_event(pf);
 
 	/* Install DMAC Filters */
-	if (pf->flags & OTX2_FLAG_DMACFLTR_SUPPORT)
+	if (otx2_test_flag(pf, OTX2_FLAG_DMACFLTR_SUPPORT))
 		otx2_dmacflt_reinstall_flows(pf);
 
 	otx2_tc_apply_ingress_police_rules(pf);
@@ -2187,7 +2187,7 @@ int otx2_open(struct net_device *netdev)
 err_tx_stop_queues:
 	netif_tx_stop_all_queues(netdev);
 	netif_carrier_off(netdev);
-	pf->flags |= OTX2_FLAG_INTF_DOWN;
+	otx2_set_flag(pf, OTX2_FLAG_INTF_DOWN);
 	/* free NIXLF POISON irq */
 	vec = pci_irq_vector(pf->pdev,
 			     pf->hw.nix_msixoff + NIX_LF_POISON_VEC);
@@ -2221,13 +2221,13 @@ int otx2_stop(struct net_device *netdev)
 	int qidx, vec, wrk;
 
 	/* If the DOWN flag is set resources are already freed */
-	if (pf->flags & OTX2_FLAG_INTF_DOWN)
+	if (otx2_test_flag(pf, OTX2_FLAG_INTF_DOWN))
 		return 0;
 
 	netif_carrier_off(netdev);
 	netif_tx_stop_all_queues(netdev);
 
-	pf->flags |= OTX2_FLAG_INTF_DOWN;
+	otx2_set_flag(pf, OTX2_FLAG_INTF_DOWN);
 	/* 'intf_down' may be checked on any cpu */
 	smp_wmb();
 
@@ -2458,7 +2458,7 @@ static int otx2_config_hw_rx_tstamp(struct otx2_nic *pfvf, bool enable)
 	struct msg_req *req;
 	int err;
 
-	if (pfvf->flags & OTX2_FLAG_RX_TSTAMP_ENABLED && enable)
+	if (otx2_test_flag(pfvf, OTX2_FLAG_RX_TSTAMP_ENABLED) && enable)
 		return 0;
 
 	mutex_lock(&pfvf->mbox.lock);
@@ -2479,9 +2479,9 @@ static int otx2_config_hw_rx_tstamp(struct otx2_nic *pfvf, bool enable)
 
 	mutex_unlock(&pfvf->mbox.lock);
 	if (enable)
-		pfvf->flags |= OTX2_FLAG_RX_TSTAMP_ENABLED;
+		otx2_set_flag(pfvf, OTX2_FLAG_RX_TSTAMP_ENABLED);
 	else
-		pfvf->flags &= ~OTX2_FLAG_RX_TSTAMP_ENABLED;
+		otx2_clear_flag(pfvf, OTX2_FLAG_RX_TSTAMP_ENABLED);
 	return 0;
 }
 
@@ -2490,7 +2490,7 @@ static int otx2_config_hw_tx_tstamp(struct otx2_nic *pfvf, bool enable)
 	struct msg_req *req;
 	int err;
 
-	if (pfvf->flags & OTX2_FLAG_TX_TSTAMP_ENABLED && enable)
+	if (otx2_test_flag(pfvf, OTX2_FLAG_TX_TSTAMP_ENABLED) && enable)
 		return 0;
 
 	mutex_lock(&pfvf->mbox.lock);
@@ -2511,9 +2511,9 @@ static int otx2_config_hw_tx_tstamp(struct otx2_nic *pfvf, bool enable)
 
 	mutex_unlock(&pfvf->mbox.lock);
 	if (enable)
-		pfvf->flags |= OTX2_FLAG_TX_TSTAMP_ENABLED;
+		otx2_set_flag(pfvf, OTX2_FLAG_TX_TSTAMP_ENABLED);
 	else
-		pfvf->flags &= ~OTX2_FLAG_TX_TSTAMP_ENABLED;
+		otx2_clear_flag(pfvf, OTX2_FLAG_TX_TSTAMP_ENABLED);
 	return 0;
 }
 
@@ -2538,8 +2538,8 @@ int otx2_config_hwtstamp_set(struct net_device *netdev,
 
 	switch (config->tx_type) {
 	case HWTSTAMP_TX_OFF:
-		if (pfvf->flags & OTX2_FLAG_PTP_ONESTEP_SYNC)
-			pfvf->flags &= ~OTX2_FLAG_PTP_ONESTEP_SYNC;
+		if (otx2_test_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC))
+			otx2_clear_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC);
 
 		cancel_delayed_work(&pfvf->ptp->synctstamp_work);
 		otx2_config_hw_tx_tstamp(pfvf, false);
@@ -2550,7 +2550,7 @@ int otx2_config_hwtstamp_set(struct net_device *netdev,
 					   "One-step time stamping is not supported");
 			return -ERANGE;
 		}
-		pfvf->flags |= OTX2_FLAG_PTP_ONESTEP_SYNC;
+		otx2_set_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC);
 		schedule_delayed_work(&pfvf->ptp->synctstamp_work,
 				      msecs_to_jiffies(500));
 		fallthrough;
@@ -2836,7 +2836,7 @@ static int otx2_set_vf_vlan(struct net_device *netdev, int vf, u16 vlan, u8 qos,
 	if (proto != htons(ETH_P_8021Q))
 		return -EPROTONOSUPPORT;
 
-	if (!(pf->flags & OTX2_FLAG_VF_VLAN_SUPPORT))
+	if (!otx2_test_flag(pf, OTX2_FLAG_VF_VLAN_SUPPORT))
 		return -EOPNOTSUPP;
 
 	return otx2_do_set_vf_vlan(pf, vf, vlan, qos, proto);
@@ -3087,7 +3087,7 @@ int otx2_realloc_msix_vectors(struct otx2_nic *pf)
 	 * interrupt range (QINT, CINT, GINT, ERR and POISON vectors).
 	 */
 	num_vec = hw->nix_msixoff;
-	if (pf->flags & OTX2_FLAG_REP_MODE_ENABLED)
+	if (otx2_test_flag(pf, OTX2_FLAG_REP_MODE_ENABLED))
 		num_vec += NIX_LF_CINT_VEC_START + hw->max_queues;
 	else
 		num_vec += NIX_LF_POISON_VEC + 1;
@@ -3273,7 +3273,7 @@ static int otx2_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 	pf->pdev = pdev;
 	pf->dev = dev;
 	pf->total_vfs = pci_sriov_get_totalvfs(pdev);
-	pf->flags |= OTX2_FLAG_INTF_DOWN;
+	otx2_set_flag(pf, OTX2_FLAG_INTF_DOWN);
 
 	hw = &pf->hw;
 	hw->pdev = pdev;
@@ -3328,23 +3328,23 @@ static int otx2_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 	if (err)
 		goto err_del_mcam_entries;
 
-	if (pf->flags & OTX2_FLAG_NTUPLE_SUPPORT)
+	if (otx2_test_flag(pf, OTX2_FLAG_NTUPLE_SUPPORT))
 		netdev->hw_features |= NETIF_F_NTUPLE;
 
-	if (pf->flags & OTX2_FLAG_UCAST_FLTR_SUPPORT)
+	if (otx2_test_flag(pf, OTX2_FLAG_UCAST_FLTR_SUPPORT))
 		netdev->priv_flags |= IFF_UNICAST_FLT;
 
 	/* Support TSO on tag interface */
 	netdev->vlan_features |= netdev->features;
 	netdev->hw_features  |= NETIF_F_HW_VLAN_CTAG_TX |
 				NETIF_F_HW_VLAN_STAG_TX;
-	if (pf->flags & OTX2_FLAG_RX_VLAN_SUPPORT)
+	if (otx2_test_flag(pf, OTX2_FLAG_RX_VLAN_SUPPORT))
 		netdev->hw_features |= NETIF_F_HW_VLAN_CTAG_RX |
 				       NETIF_F_HW_VLAN_STAG_RX;
 	netdev->features |= netdev->hw_features;
 
 	/* HW supports tc offload but mutually exclusive with n-tuple filters */
-	if (pf->flags & OTX2_FLAG_TC_FLOWER_SUPPORT)
+	if (otx2_test_flag(pf, OTX2_FLAG_TC_FLOWER_SUPPORT))
 		netdev->hw_features |= NETIF_F_HW_TC;
 
 	netdev->hw_features |= NETIF_F_LOOPBACK | NETIF_F_RXALL;
@@ -3595,18 +3595,18 @@ static void otx2_remove(struct pci_dev *pdev)
 
 	pf = netdev_priv(netdev);
 
-	pf->flags |= OTX2_FLAG_PF_SHUTDOWN;
+	otx2_set_flag(pf, OTX2_FLAG_PF_SHUTDOWN);
 
-	if (pf->flags & OTX2_FLAG_TX_TSTAMP_ENABLED)
+	if (otx2_test_flag(pf, OTX2_FLAG_TX_TSTAMP_ENABLED))
 		otx2_config_hw_tx_tstamp(pf, false);
-	if (pf->flags & OTX2_FLAG_RX_TSTAMP_ENABLED)
+	if (otx2_test_flag(pf, OTX2_FLAG_RX_TSTAMP_ENABLED))
 		otx2_config_hw_rx_tstamp(pf, false);
 
 	/* Disable 802.3x pause frames */
-	if (pf->flags & OTX2_FLAG_RX_PAUSE_ENABLED ||
-	    (pf->flags & OTX2_FLAG_TX_PAUSE_ENABLED)) {
-		pf->flags &= ~OTX2_FLAG_RX_PAUSE_ENABLED;
-		pf->flags &= ~OTX2_FLAG_TX_PAUSE_ENABLED;
+	if (otx2_test_flag(pf, OTX2_FLAG_RX_PAUSE_ENABLED) ||
+	    otx2_test_flag(pf, OTX2_FLAG_TX_PAUSE_ENABLED)) {
+		otx2_clear_flag(pf, OTX2_FLAG_RX_PAUSE_ENABLED);
+		otx2_clear_flag(pf, OTX2_FLAG_TX_PAUSE_ENABLED);
 		otx2_config_pause_frm(pf);
 	}
 
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
index dee38d5c7f65..8877af348a09 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
@@ -159,7 +159,7 @@ static int otx2_tc_validate_flow(struct otx2_nic *nic,
 				 struct flow_action *actions,
 				 struct netlink_ext_ack *extack)
 {
-	if (nic->flags & OTX2_FLAG_INTF_DOWN) {
+	if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN)) {
 		NL_SET_ERR_MSG_MOD(extack, "Interface not initialized");
 		return -EINVAL;
 	}
@@ -223,7 +223,7 @@ static int otx2_tc_egress_matchall_install(struct otx2_nic *nic,
 	if (err)
 		return err;
 
-	if (nic->flags & OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED) {
+	if (otx2_test_flag(nic, OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED)) {
 		NL_SET_ERR_MSG_MOD(extack,
 				   "Only one Egress MATCHALL ratelimiter can be offloaded");
 		return -ENOMEM;
@@ -244,7 +244,7 @@ static int otx2_tc_egress_matchall_install(struct otx2_nic *nic,
 						    otx2_convert_rate(entry->police.rate_bytes_ps));
 		if (err)
 			return err;
-		nic->flags |= OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED;
+		otx2_set_flag(nic, OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED);
 		break;
 	default:
 		NL_SET_ERR_MSG_MOD(extack,
@@ -261,13 +261,13 @@ static int otx2_tc_egress_matchall_delete(struct otx2_nic *nic,
 	struct netlink_ext_ack *extack = cls->common.extack;
 	int err;
 
-	if (nic->flags & OTX2_FLAG_INTF_DOWN) {
+	if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN)) {
 		NL_SET_ERR_MSG_MOD(extack, "Interface not initialized");
 		return -EINVAL;
 	}
 
 	err = otx2_set_matchall_egress_rate(nic, 0, 0);
-	nic->flags &= ~OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED;
+	otx2_clear_flag(nic, OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED);
 	return err;
 }
 
@@ -505,7 +505,7 @@ static int otx2_tc_parse_actions(struct otx2_nic *nic,
 			mark = act->mark;
 			req->match_id = mark & OTX2_RX_MATCH_ID_MASK;
 			req->op = NIX_RX_ACTION_DEFAULT;
-			nic->flags |= OTX2_FLAG_TC_MARK_ENABLED;
+			otx2_set_flag(nic, OTX2_FLAG_TC_MARK_ENABLED);
 			refcount_inc(&nic->flow_cfg->mark_flows);
 			break;
 
@@ -942,7 +942,7 @@ static void otx2_destroy_tc_flow_list(struct otx2_nic *pfvf)
 	struct otx2_flow_config *flow_cfg = pfvf->flow_cfg;
 	struct otx2_tc_flow *iter, *tmp;
 
-	if (!(pfvf->flags & OTX2_FLAG_MCAM_ENTRIES_ALLOC))
+	if (!otx2_test_flag(pfvf, OTX2_FLAG_MCAM_ENTRIES_ALLOC))
 		return;
 
 	list_for_each_entry_safe(iter, tmp, &flow_cfg->flow_list_tc, list) {
@@ -1195,12 +1195,12 @@ static int otx2_tc_del_flow(struct otx2_nic *nic,
 	/* Disable TC MARK flag if they are no rules with skbedit mark action */
 	if (flow_node->req.match_id)
 		if (!refcount_dec_and_test(&flow_cfg->mark_flows))
-			nic->flags &= ~OTX2_FLAG_TC_MARK_ENABLED;
+			otx2_clear_flag(nic, OTX2_FLAG_TC_MARK_ENABLED);
 
 	if (flow_node->is_act_police) {
 		__clear_bit(flow_node->rq, &nic->rq_bmap);
 
-		if (nic->flags & OTX2_FLAG_INTF_DOWN)
+		if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN))
 			goto free_mcam_flow;
 
 		mutex_lock(&nic->mbox.lock);
@@ -1246,10 +1246,10 @@ static int otx2_tc_add_flow(struct otx2_nic *nic,
 	struct npc_install_flow_req *req, dummy;
 	int rc, err, entry;
 
-	if (!(nic->flags & OTX2_FLAG_TC_FLOWER_SUPPORT))
+	if (!otx2_test_flag(nic, OTX2_FLAG_TC_FLOWER_SUPPORT))
 		return -ENOMEM;
 
-	if (nic->flags & OTX2_FLAG_INTF_DOWN) {
+	if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN)) {
 		NL_SET_ERR_MSG_MOD(extack, "Interface not initialized");
 		return -EINVAL;
 	}
@@ -1444,7 +1444,7 @@ static int otx2_tc_ingress_matchall_install(struct otx2_nic *nic,
 	if (err)
 		return err;
 
-	if (nic->flags & OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED) {
+	if (otx2_test_flag(nic, OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED)) {
 		NL_SET_ERR_MSG_MOD(extack,
 				   "Only one ingress MATCHALL ratelimitter can be offloaded");
 		return -ENOMEM;
@@ -1469,7 +1469,7 @@ static int otx2_tc_ingress_matchall_install(struct otx2_nic *nic,
 		err = cn10k_set_matchall_ipolicer_rate(nic, entry->police.burst, rate);
 		if (err)
 			return err;
-		nic->flags |= OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED;
+		otx2_set_flag(nic, OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED);
 		break;
 	default:
 		NL_SET_ERR_MSG_MOD(extack,
@@ -1486,13 +1486,13 @@ static int otx2_tc_ingress_matchall_delete(struct otx2_nic *nic,
 	struct netlink_ext_ack *extack = cls->common.extack;
 	int err;
 
-	if (nic->flags & OTX2_FLAG_INTF_DOWN) {
+	if (otx2_test_flag(nic, OTX2_FLAG_INTF_DOWN)) {
 		NL_SET_ERR_MSG_MOD(extack, "Interface not initialized");
 		return -EINVAL;
 	}
 
 	err = cn10k_free_matchall_ipolicer(nic);
-	nic->flags &= ~OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED;
+	otx2_clear_flag(nic, OTX2_FLAG_TC_MATCHALL_INGRESS_ENABLED);
 	return err;
 }
 
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
index 8d2d607bc92f..f65ba44db60b 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_txrx.c
@@ -171,7 +171,7 @@ static void otx2_set_rxtstamp(struct otx2_nic *pfvf,
 	u64 timestamp, tsns;
 	int err;
 
-	if (!(pfvf->flags & OTX2_FLAG_RX_TSTAMP_ENABLED))
+	if (!otx2_test_flag(pfvf, OTX2_FLAG_RX_TSTAMP_ENABLED))
 		return;
 
 	timestamp = pfvf->ptp->convert_rx_ptp_tstmp(*(u64 *)data);
@@ -374,13 +374,13 @@ static void otx2_rcv_pkt_handler(struct otx2_nic *pfvf,
 	}
 	otx2_set_rxhash(pfvf, cqe, skb);
 
-	if (!(pfvf->flags & OTX2_FLAG_REP_MODE_ENABLED)) {
+	if (!otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED)) {
 		skb_record_rx_queue(skb, cq->cq_idx);
 		if (pfvf->netdev->features & NETIF_F_RXCSUM)
 			skb->ip_summed = CHECKSUM_UNNECESSARY;
 	}
 
-	if (pfvf->flags & OTX2_FLAG_TC_MARK_ENABLED)
+	if (otx2_test_flag(pfvf, OTX2_FLAG_TC_MARK_ENABLED))
 		skb->mark = parse->match_id;
 
 	skb_mark_for_recycle(skb);
@@ -513,7 +513,7 @@ static int otx2_tx_napi_handler(struct otx2_nic *pfvf,
 		     ((u64)cq->cq_idx << 32) | processed_cqe);
 
 #if IS_ENABLED(CONFIG_RVU_ESWITCH)
-	if (pfvf->flags & OTX2_FLAG_REP_MODE_ENABLED)
+	if (otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED))
 		ndev = pfvf->reps[qidx]->netdev;
 	else
 #endif
@@ -526,7 +526,7 @@ static int otx2_tx_napi_handler(struct otx2_nic *pfvf,
 
 		if (qidx >= pfvf->hw.tx_queues)
 			qidx -= pfvf->hw.xdp_queues;
-		if (pfvf->flags & OTX2_FLAG_REP_MODE_ENABLED)
+		if (otx2_test_flag(pfvf, OTX2_FLAG_REP_MODE_ENABLED))
 			qidx = 0;
 		txq = netdev_get_tx_queue(ndev, qidx);
 		netdev_tx_completed_queue(txq, tx_pkts, tx_bytes);
@@ -599,11 +599,11 @@ int otx2_napi_handler(struct napi_struct *napi, int budget)
 
 	if (workdone < budget && napi_complete_done(napi, workdone)) {
 		/* If interface is going down, don't re-enable IRQ */
-		if (pfvf->flags & OTX2_FLAG_INTF_DOWN)
+		if (otx2_test_flag(pfvf, OTX2_FLAG_INTF_DOWN))
 			return workdone;
 
 		/* Adjust irq coalese using net_dim */
-		if (pfvf->flags & OTX2_FLAG_ADPTV_INT_COAL_ENABLED)
+		if (otx2_test_flag(pfvf, OTX2_FLAG_ADPTV_INT_COAL_ENABLED))
 			otx2_adjust_adaptive_coalese(pfvf, cq_poll);
 
 		if (likely(cq))
@@ -1137,7 +1137,7 @@ static void otx2_set_txtstamp(struct otx2_nic *pfvf, struct sk_buff *skb,
 
 	if (unlikely(!skb_shinfo(skb)->gso_size &&
 		     (skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP))) {
-		if (unlikely(pfvf->flags & OTX2_FLAG_PTP_ONESTEP_SYNC &&
+		if (unlikely(otx2_test_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC) &&
 			     otx2_ptp_is_sync(skb, &ptp_offset, &udp_csum_crt))) {
 			origin_tstamp = (struct ptpv2_tstamp *)
 					((u8 *)skb->data + ptp_offset +
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
index f7765e19d78a..5f7915231ca3 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
@@ -610,7 +610,7 @@ static int otx2vf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 	vf->dev = dev;
 	vf->iommu_domain = iommu_get_domain_for_dev(dev);
 
-	vf->flags |= OTX2_FLAG_INTF_DOWN;
+	otx2_set_flag(vf, OTX2_FLAG_INTF_DOWN);
 	hw = &vf->hw;
 	hw->pdev = vf->pdev;
 	hw->rx_queues = qcount;
@@ -824,10 +824,10 @@ static void otx2vf_remove(struct pci_dev *pdev)
 	vf = netdev_priv(netdev);
 
 	/* Disable 802.3x pause frames */
-	if (vf->flags & OTX2_FLAG_RX_PAUSE_ENABLED ||
-	    (vf->flags & OTX2_FLAG_TX_PAUSE_ENABLED)) {
-		vf->flags &= ~OTX2_FLAG_RX_PAUSE_ENABLED;
-		vf->flags &= ~OTX2_FLAG_TX_PAUSE_ENABLED;
+	if (otx2_test_flag(vf, OTX2_FLAG_RX_PAUSE_ENABLED) ||
+	    otx2_test_flag(vf, OTX2_FLAG_TX_PAUSE_ENABLED)) {
+		otx2_clear_flag(vf, OTX2_FLAG_RX_PAUSE_ENABLED);
+		otx2_clear_flag(vf, OTX2_FLAG_TX_PAUSE_ENABLED);
 		otx2_config_pause_frm(vf);
 	}
 
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
index 0e8a6a6486c4..7808588a0234 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_xsk.c
@@ -96,7 +96,7 @@ static void otx2_clean_up_rq(struct otx2_nic *pfvf, int qidx)
 	u64 iova;
 
 	/* If the DOWN flag is set SQs are already freed */
-	if (pfvf->flags & OTX2_FLAG_INTF_DOWN)
+	if (otx2_test_flag(pfvf, OTX2_FLAG_INTF_DOWN))
 		return;
 
 	cq = &qset->cq[qidx];
@@ -172,7 +172,7 @@ int otx2_xsk_wakeup(struct net_device *dev, u32 queue_id, u32 flags)
 	struct otx2_cq_poll *cq_poll = NULL;
 	struct otx2_qset *qset = &pf->qset;
 
-	if (pf->flags & OTX2_FLAG_INTF_DOWN)
+	if (otx2_test_flag(pf, OTX2_FLAG_INTF_DOWN))
 		return -ENETDOWN;
 
 	if (queue_id >= pf->hw.rx_queues || queue_id >= pf->hw.tx_queues)
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/qos_sq.c b/drivers/net/ethernet/marvell/octeontx2/nic/qos_sq.c
index 2872adabc830..5f09e2960144 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/qos_sq.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/qos_sq.c
@@ -238,7 +238,7 @@ int otx2_qos_enable_sq(struct otx2_nic *pfvf, int qidx)
 	struct otx2_hw *hw = &pfvf->hw;
 	int pool_id, sq_idx, err;
 
-	if (pfvf->flags & OTX2_FLAG_INTF_DOWN)
+	if (otx2_test_flag(pfvf, OTX2_FLAG_INTF_DOWN))
 		return -EPERM;
 
 	sq_idx = hw->non_qos_queues + qidx;
@@ -288,7 +288,7 @@ void otx2_qos_disable_sq(struct otx2_nic *pfvf, int qidx)
 	sq_idx = hw->non_qos_queues + qidx;
 
 	/* If the DOWN flag is set SQs are already freed */
-	if (pfvf->flags & OTX2_FLAG_INTF_DOWN)
+	if (otx2_test_flag(pfvf, OTX2_FLAG_INTF_DOWN))
 		return;
 
 	sq = &pfvf->qset.sq[sq_idx];
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/rep.c b/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
index 0f5d5642d3f7..7df82c22cc12 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
@@ -93,9 +93,9 @@ static int rvu_rep_mcam_flow_init(struct rep_dev *rep)
 	rep->flow_cfg->max_flows = allocated;
 
 	if (allocated) {
-		rep->flags |= OTX2_FLAG_MCAM_ENTRIES_ALLOC;
-		rep->flags |= OTX2_FLAG_NTUPLE_SUPPORT;
-		rep->flags |= OTX2_FLAG_TC_FLOWER_SUPPORT;
+		set_bit(OTX2_FLAG_MCAM_ENTRIES_ALLOC, &rep->flags);
+		set_bit(OTX2_FLAG_NTUPLE_SUPPORT, &rep->flags);
+		set_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags);
 	}
 
 	INIT_LIST_HEAD(&rep->flow_cfg->flow_list);
@@ -109,14 +109,14 @@ static int rvu_rep_setup_tc_cb(enum tc_setup_type type,
 	struct rep_dev *rep = cb_priv;
 	struct otx2_nic *priv = rep->mdev;
 
-	if (!(rep->flags & RVU_REP_VF_INITIALIZED))
+	if (!test_bit(OTX2_FLAG_REP_VF_INITIALIZED, &rep->flags))
 		return -EINVAL;
 
-	if (!(rep->flags & OTX2_FLAG_TC_FLOWER_SUPPORT))
+	if (!test_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags))
 		rvu_rep_mcam_flow_init(rep);
 
 	priv->netdev = rep->netdev;
-	priv->flags = rep->flags;
+	otx2_sync_flags_from_rep(priv, &rep->flags);
 	priv->pcifunc = rep->pcifunc;
 	priv->flow_cfg = rep->flow_cfg;
 
@@ -303,9 +303,9 @@ static void rvu_rep_state_evt_handler(struct otx2_nic *priv,
 	rep_id = rvu_rep_get_repid(priv, info->pcifunc);
 	rep = priv->reps[rep_id];
 	if (info->evt_data.vf_state)
-		rep->flags |= RVU_REP_VF_INITIALIZED;
+		set_bit(OTX2_FLAG_REP_VF_INITIALIZED, &rep->flags);
 	else
-		rep->flags &= ~RVU_REP_VF_INITIALIZED;
+		clear_bit(OTX2_FLAG_REP_VF_INITIALIZED, &rep->flags);
 }
 
 int rvu_event_up_notify(struct otx2_nic *pf, struct rep_event *info)
@@ -382,7 +382,7 @@ static void rvu_rep_get_stats64(struct net_device *dev,
 {
 	struct rep_dev *rep = netdev_priv(dev);
 
-	if (!(rep->flags & RVU_REP_VF_INITIALIZED))
+	if (!test_bit(OTX2_FLAG_REP_VF_INITIALIZED, &rep->flags))
 		return;
 
 	stats->rx_packets = rep->stats.rx_frames;
@@ -453,7 +453,7 @@ static int rvu_rep_open(struct net_device *dev)
 	struct otx2_nic *priv = rep->mdev;
 	struct rep_event evt = {0};
 
-	if (!(rep->flags & RVU_REP_VF_INITIALIZED))
+	if (!test_bit(OTX2_FLAG_REP_VF_INITIALIZED, &rep->flags))
 		return 0;
 
 	netif_carrier_on(dev);
@@ -472,7 +472,7 @@ static int rvu_rep_stop(struct net_device *dev)
 	struct otx2_nic *priv = rep->mdev;
 	struct rep_event evt = {0};
 
-	if (!(rep->flags & RVU_REP_VF_INITIALIZED))
+	if (!test_bit(OTX2_FLAG_REP_VF_INITIALIZED, &rep->flags))
 		return 0;
 
 	netif_carrier_off(dev);
@@ -547,7 +547,7 @@ static int rvu_rep_napi_init(struct otx2_nic *priv,
 		otx2_write64(priv, NIX_LF_CINTX_INT(qidx), BIT_ULL(0));
 		otx2_write64(priv, NIX_LF_CINTX_ENA_W1S(qidx), BIT_ULL(0));
 	}
-	priv->flags &= ~OTX2_FLAG_INTF_DOWN;
+	otx2_clear_flag(priv, OTX2_FLAG_INTF_DOWN);
 	return 0;
 
 err_free_cints:
@@ -632,7 +632,7 @@ void rvu_rep_destroy(struct otx2_nic *priv)
 	int rep_id;
 
 	rvu_eswitch_config(priv, false);
-	priv->flags |= OTX2_FLAG_INTF_DOWN;
+	otx2_set_flag(priv, OTX2_FLAG_INTF_DOWN);
 	rvu_rep_free_cq_rsrc(priv);
 	for (rep_id = 0; rep_id < priv->rep_cnt; rep_id++) {
 		rep = priv->reps[rep_id];
@@ -801,8 +801,8 @@ static int rvu_rep_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 	pci_set_drvdata(pdev, priv);
 	priv->pdev = pdev;
 	priv->dev = dev;
-	priv->flags |= OTX2_FLAG_INTF_DOWN;
-	priv->flags |= OTX2_FLAG_REP_MODE_ENABLED;
+	otx2_set_flag(priv, OTX2_FLAG_INTF_DOWN);
+	otx2_set_flag(priv, OTX2_FLAG_REP_MODE_ENABLED);
 
 	hw = &priv->hw;
 	hw->pdev = pdev;
@@ -845,7 +845,7 @@ static void rvu_rep_remove(struct pci_dev *pdev)
 	struct otx2_nic *priv = pci_get_drvdata(pdev);
 
 	otx2_unregister_dl(priv);
-	if (!(priv->flags & OTX2_FLAG_INTF_DOWN))
+	if (!otx2_test_flag(priv, OTX2_FLAG_INTF_DOWN))
 		rvu_rep_destroy(priv);
 	otx2_detach_resources(&priv->mbox);
 	if (priv->hw.lmt_info)
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/rep.h b/drivers/net/ethernet/marvell/octeontx2/nic/rep.h
index 5bc9e2c7d800..45707c434d89 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/rep.h
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/rep.h
@@ -37,8 +37,7 @@ struct rep_dev {
 	struct delayed_work stats_wrk;
 	struct devlink_port dl_port;
 	struct otx2_flow_config	*flow_cfg;
-#define RVU_REP_VF_INITIALIZED		BIT_ULL(0)
-	u64 flags;
+	unsigned long		flags;
 	u16 rep_id;
 	u16 pcifunc;
 	u8 mac[ETH_ALEN];
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v19 net-next 2/2] octeontx2: add mqprio bandwidth offload for NIX TX schedulers
  2026-10-06  9:43 [PATCH v19 net-next 0/2] octeontx2: mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
  2026-10-06  9:43 ` [PATCH v19 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags Ratheesh Kannoth
@ 2026-10-06  9:43 ` Ratheesh Kannoth
  2026-10-10 10:01   ` netdev-bot+sashiko
  1 sibling, 1 reply; 5+ messages in thread
From: Ratheesh Kannoth @ 2026-10-06  9:43 UTC (permalink / raw)
  To: bpf, linux-kernel, netdev
  Cc: andrew+netdev, ast, daniel, davem, edumazet, hawk, john.fastabend,
	kuba, pabeni, sdf, sgoutham, Ratheesh Kannoth

Add TC_SETUP_QDISC_MQPRIO offload with TC_MQPRIO_SHAPER_BW_RATE on PF and
VF netdevs. Program per-queue MDQ CIR/PIR via the NIX TX scheduler mailbox;
when active, allocate one SMQ per non-QoS queue under TL4[0].

The TX scheduler cannot be reprogrammed live, so add/replace/delete and
failed setup bounce the netdev through ndo_stop()/ndo_open(). Cache rates
in software and restore them from otx2_mqprio_up() on open, failing closed
on error.

Stage tc replace in mq_offload_snap snapshots committed on TC_ROOT_GRAFT or
replaced-qdisc teardown.

Require a running interface with CIR+PIR support. Reject offload with PFC,
XDP, HTB, SDP representors, per-TC rates on multi-queue classes, and
ethtool channel changes while active.

Signed-off-by: Ratheesh Kannoth <rkannoth@marvell.com>
---
 .../marvell/octeontx2/nic/otx2_common.c       | 143 ++-
 .../marvell/octeontx2/nic/otx2_common.h       |  34 +
 .../marvell/octeontx2/nic/otx2_dcbnl.c        |   6 +
 .../marvell/octeontx2/nic/otx2_ethtool.c      |  15 +
 .../ethernet/marvell/octeontx2/nic/otx2_pf.c  |  34 +-
 .../ethernet/marvell/octeontx2/nic/otx2_tc.c  | 927 ++++++++++++++++++
 .../ethernet/marvell/octeontx2/nic/otx2_vf.c  |  28 +-
 .../net/ethernet/marvell/octeontx2/nic/qos.c  |  11 +
 8 files changed, 1177 insertions(+), 21 deletions(-)

diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
index 836601fc10d8..c95177602bf9 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
@@ -615,6 +615,139 @@ void otx2_get_mac_from_af(struct net_device *netdev)
 }
 EXPORT_SYMBOL(otx2_get_mac_from_af);
 
+static int
+otx2_nix_tmq_reg_write(struct otx2_nic *pfvf, int cnt,
+		       u64 reg_addr[MAX_REGS_PER_MBOX_MSG],
+		       u64 reg_val[MAX_REGS_PER_MBOX_MSG])
+{
+	struct mbox *mbox = &pfvf->mbox;
+	struct nix_txschq_config *req;
+	int i, err;
+
+	mutex_lock(&mbox->lock);
+	req = otx2_mbox_alloc_msg_nix_txschq_cfg(mbox);
+	if (!req) {
+		mutex_unlock(&mbox->lock);
+		return -ENOMEM;
+	}
+
+	req->lvl = NIX_TXSCH_LVL_MDQ;
+	req->num_regs = cnt;
+
+	for (i = 0; i < cnt; i++) {
+		req->reg[i] = reg_addr[i];
+		req->regval[i] = reg_val[i];
+	}
+
+	err = otx2_sync_mbox_msg(mbox);
+	mutex_unlock(&mbox->lock);
+
+	return err;
+}
+
+int otx2_nix_tm_clear_queue_shaper(struct otx2_nic *pfvf)
+{
+	u64 reg_addr[MAX_REGS_PER_MBOX_MSG];
+	u64 reg_val[MAX_REGS_PER_MBOX_MSG];
+	int err, smq, i, cnt = 0;
+
+	for (i = 0; i < pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ]; i++) {
+		smq = pfvf->hw.txschq_list[NIX_TXSCH_LVL_SMQ][i];
+
+		reg_addr[cnt] = NIX_AF_MDQX_PIR(smq);
+		reg_val[cnt] = 0;
+		cnt++;
+
+		reg_addr[cnt] = NIX_AF_MDQX_CIR(smq);
+		reg_val[cnt] = 0;
+		cnt++;
+
+		if (cnt < MAX_REGS_PER_MBOX_MSG - 1)
+			continue;
+
+		err = otx2_nix_tmq_reg_write(pfvf, cnt,
+					     reg_addr, reg_val);
+		if (err)
+			goto fail;
+		cnt = 0;
+	}
+
+	if (cnt) {
+		err = otx2_nix_tmq_reg_write(pfvf, cnt,
+					     reg_addr, reg_val);
+		if (err)
+			goto fail;
+	}
+
+	return 0;
+fail:
+	return err;
+}
+
+int otx2_nix_tm_set_queue_shaper(struct otx2_nic *pfvf,
+				 int txq, u64 minrate, u64 maxrate)
+{
+	struct mbox *mbox = &pfvf->mbox;
+	struct nix_txschq_config *req;
+	int err, smq, n = 0;
+	u64 reg_addr[2];
+	u64 reg_val[2];
+	u64 rate;
+
+	if (!maxrate && !minrate) {
+		smq = otx2_get_smq_idx(pfvf, txq);
+		reg_addr[0] = NIX_AF_MDQX_PIR(smq);
+		reg_val[0] = 0;
+		reg_addr[1] = NIX_AF_MDQX_CIR(smq);
+		reg_val[1] = 0;
+		return otx2_nix_tmq_reg_write(pfvf, 2, reg_addr, reg_val);
+	}
+
+	smq = otx2_get_smq_idx(pfvf, txq);
+
+	mutex_lock(&mbox->lock);
+	req = otx2_mbox_alloc_msg_nix_txschq_cfg(mbox);
+	if (!req) {
+		mutex_unlock(&mbox->lock);
+		return -ENOMEM;
+	}
+
+	req->lvl = NIX_TXSCH_LVL_MDQ;
+
+	/* MQPRIO exposes only min/max rate, not burst.  Use the same 65536
+	 * byte default as the HTB shaper path.
+	 *
+	 * mqprio setup restarts the netdev (otx2_mqprio_restart_netdev);
+	 * ndo_open() may restore cached MDQ shapers via otx2_mqprio_up(), and
+	 * setup clears every MDQ shaper before applying the new mapping.
+	 * Program both PIR and CIR on every update so omitted rates are
+	 * applied explicitly rather than relying on stale hardware state.
+	 */
+	req->reg[n] = NIX_AF_MDQX_PIR(smq);
+	if (maxrate) {
+		rate = otx2_convert_rate(maxrate);
+		req->regval[n] = otx2_get_txschq_rate_regval(pfvf, rate, 65536);
+	} else {
+		req->regval[n] = 0;
+	}
+	n++;
+
+	/* CIR+PIR support is required and checked at mqprio setup. */
+	req->reg[n] = NIX_AF_MDQX_CIR(smq);
+	if (minrate) {
+		rate = otx2_convert_rate(minrate);
+		req->regval[n] = otx2_get_txschq_rate_regval(pfvf, rate, 65536);
+	} else {
+		req->regval[n] = 0;
+	}
+	n++;
+	req->num_regs = n;
+
+	err = otx2_sync_mbox_msg(mbox);
+	mutex_unlock(&mbox->lock);
+	return err;
+}
+
 int otx2_txschq_config(struct otx2_nic *pfvf, int lvl, int prio, bool txschq_for_pfc)
 {
 	u16 (*schq_list)[MAX_TXSCHQ_PER_FUNC];
@@ -651,7 +784,11 @@ int otx2_txschq_config(struct otx2_nic *pfvf, int lvl, int prio, bool txschq_for
 						(u64)hw->smq_link_type);
 		req->num_regs++;
 		/* MDQ config */
-		parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
+		if (pfvf->mqprio.rate_limit)
+			parent = schq_list[NIX_TXSCH_LVL_TL4][0];
+		else
+			parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
+
 		req->reg[1] = NIX_AF_MDQX_PARENT(schq);
 		req->regval[1] = parent << 16;
 		req->num_regs++;
@@ -777,6 +914,9 @@ int otx2_txsch_alloc(struct otx2_nic *pfvf)
 		req->schq[NIX_TXSCH_LVL_TL4] = chan_cnt;
 	}
 
+	if (pfvf->mqprio.rate_limit)
+		req->schq[NIX_TXSCH_LVL_SMQ] = pfvf->hw.non_qos_queues;
+
 	rc = otx2_sync_mbox_msg(&pfvf->mbox);
 	if (rc)
 		return rc;
@@ -841,6 +981,7 @@ void otx2_txschq_stop(struct otx2_nic *pfvf)
 
 	/* Clear the txschq list */
 	for (lvl = 0; lvl < NIX_TXSCH_LVL_CNT; lvl++) {
+		pfvf->hw.txschq_cnt[lvl] = 0;
 		for (schq = 0; schq < MAX_TXSCHQ_PER_FUNC; schq++)
 			pfvf->hw.txschq_list[lvl][schq] = 0;
 	}
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
index 3bc57dd10451..3f3e21c9d347 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
@@ -17,6 +17,7 @@
 #include <linux/soc/marvell/silicons.h>
 #include <linux/soc/marvell/octeontx2/asm.h>
 #include <net/macsec.h>
+#include <uapi/linux/pkt_sched.h>
 #include <net/pkt_cls.h>
 #include <net/devlink.h>
 #include <linux/time64.h>
@@ -517,6 +518,27 @@ enum otx2_flag_bits {
 	 BIT(OTX2_FLAG_TC_FLOWER_SUPPORT) |			\
 	 BIT(OTX2_FLAG_REP_VF_INITIALIZED))
 
+struct mq_offload_snap {
+	u64 min_rate[TC_QOPT_MAX_QUEUE];
+	u64 max_rate[TC_QOPT_MAX_QUEUE];
+	u32 flags;
+	__u8 num_tc;
+	__u16 count[TC_QOPT_MAX_QUEUE];
+	__u16 offset[TC_QOPT_MAX_QUEUE];
+	__u8 prio_tc_map[TC_QOPT_BITMASK + 1];
+};
+
+struct otx2_mqprio {
+	u32	flags;
+	u64	*min_rate;
+	u64	*max_rate;
+	bool	rate_limit;
+	bool	replace_setup_done;
+	bool	replace_graft_done;
+	bool	defer_tc_work;
+	struct work_struct netdev_tc_work;
+};
+
 struct otx2_nic {
 	void __iomem		*reg_base;
 	struct net_device	*netdev;
@@ -528,6 +550,10 @@ struct otx2_nic {
 	unsigned long		flags;
 	u64			*cq_op_addr;
 
+	struct otx2_mqprio	mqprio;
+	struct mq_offload_snap	*cur_mq_snap;
+	struct mq_offload_snap	*old_mq_snap;
+
 	struct bpf_prog		*xdp_prog;
 	struct otx2_qset	qset;
 	struct otx2_hw		hw;
@@ -1036,6 +1062,8 @@ static inline u16 otx2_get_smq_idx(struct otx2_nic *pfvf, u16 qidx)
 	if (qidx >= pfvf->hw.non_qos_queues) {
 		smq = pfvf->qos.qid_to_sqmap[qidx - pfvf->hw.non_qos_queues];
 	} else {
+		if (!pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ])
+			return 0;
 		idx = qidx % pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ];
 		smq = pfvf->hw.txschq_list[NIX_TXSCH_LVL_SMQ][idx];
 	}
@@ -1223,6 +1251,7 @@ int otx2_mcam_entry_init(struct otx2_nic *pfvf);
 
 /* tc support */
 int otx2_init_tc(struct otx2_nic *nic);
+void otx2_shutdown_tc_mqprio(struct otx2_nic *nic);
 void otx2_shutdown_tc(struct otx2_nic *nic);
 int otx2_setup_tc(struct net_device *netdev, enum tc_setup_type type,
 		  void *type_data);
@@ -1294,6 +1323,11 @@ dma_addr_t otx2_dma_map_skb_frag(struct otx2_nic *pfvf,
 				 struct sk_buff *skb, int seg, int *len);
 void otx2_dma_unmap_skb_frags(struct otx2_nic *pfvf, struct sg_list *sg);
 int otx2_read_free_sqe(struct otx2_nic *pfvf, u16 qidx);
+int otx2_nix_tm_set_queue_shaper(struct otx2_nic *pfvf, int txq,
+				 u64 minrate, u64 maxrate);
+int otx2_nix_tm_clear_queue_shaper(struct otx2_nic *pfvf);
+int otx2_mqprio_down(struct otx2_nic *pfvf);
+int otx2_mqprio_up(struct otx2_nic *pfvf);
 void otx2_queue_vf_work(struct mbox *mw, struct workqueue_struct *mbox_wq,
 			int first, int mdevs, u64 intr);
 int otx2_del_mcam_flow_entry(struct otx2_nic *nic, u16 entry,
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
index e3d54009180d..851bf3fd5bab 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_dcbnl.c
@@ -408,6 +408,12 @@ static int otx2_dcbnl_ieee_setpfc(struct net_device *dev, struct ieee_pfc *pfc)
 	u8 old_pfc_en;
 	int err;
 
+	if (pfvf->mqprio.rate_limit && pfc->pfc_en) {
+		netdev_err(dev,
+			   "PFC: cannot enable while mqprio bandwidth offload is active\n");
+		return -EOPNOTSUPP;
+	}
+
 	old_pfc_en = pfvf->pfc_en;
 	pfvf->pfc_en = pfc->pfc_en;
 
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
index 4fe473d9ea0d..836582631127 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
@@ -287,6 +287,21 @@ static int otx2_set_channels(struct net_device *dev,
 		return -EINVAL;
 	}
 
+	if (pfvf->mqprio.rate_limit &&
+	    (channel->tx_count != pfvf->hw.tx_queues ||
+	     channel->rx_count != pfvf->hw.rx_queues)) {
+		netdev_info(dev,
+			    "Not permitted to change channel count while MQ prio is active\n");
+		return -EINVAL;
+	}
+
+	if ((pfvf->old_mq_snap || pfvf->cur_mq_snap) &&
+	    channel->tx_count < pfvf->hw.tx_queues) {
+		netdev_err(dev,
+			   "Cannot reduce TX queues after mqprio bandwidth offload was configured\n");
+		return -EINVAL;
+	}
+
 	if (if_up)
 		dev->netdev_ops->ndo_stop(dev);
 
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
index 777e7156badb..f51161e8b594 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
@@ -2008,6 +2008,15 @@ int otx2_open(struct net_device *netdev)
 	if (err)
 		goto err_free_mem;
 
+	/* Fail closed: abort open if cached mqprio shapers cannot be restored. */
+	err = otx2_mqprio_up(pf);
+	if (err) {
+		netdev_err(pf->netdev,
+			   "mqprio: failed to restore shapers during open: %d\n",
+			   err);
+		goto err_free_hw;
+	}
+
 	/* Register NAPI handler */
 	for (qidx = 0; qidx < pf->hw.cint_cnt; qidx++) {
 		cq_poll = &qset->napi[qidx];
@@ -2206,6 +2215,7 @@ int otx2_open(struct net_device *netdev)
 	free_irq(vec, pf);
 err_disable_napi:
 	otx2_disable_napi(pf);
+err_free_hw:
 	otx2_free_hw_resources(pf);
 err_free_mem:
 	otx2_free_queue_mem(qset);
@@ -2281,6 +2291,7 @@ int otx2_stop(struct net_device *netdev)
 	for (qidx = 0; qidx < netdev->num_tx_queues; qidx++)
 		netdev_tx_reset_queue(netdev_get_tx_queue(netdev, qidx));
 
+	synchronize_net();
 	otx2_free_queue_mem(qset);
 	/* Do not clear RQ/SQ ringsize settings */
 	memset_startat(qset, 0, sqe_cnt);
@@ -2924,6 +2935,12 @@ static int otx2_xdp_setup(struct otx2_nic *pf, struct bpf_prog *prog)
 	bool if_up = netif_running(pf->netdev);
 	struct bpf_prog *old_prog;
 
+	if (prog && pf->mqprio.rate_limit) {
+		netdev_err(dev,
+			   "XDP: cannot attach while mqprio bandwidth offload is active\n");
+		return -EOPNOTSUPP;
+	}
+
 	if (prog && dev->mtu > MAX_XDP_MTU) {
 		netdev_warn(dev, "Jumbo frames not yet supported with XDP\n");
 		return -EOPNOTSUPP;
@@ -3368,10 +3385,14 @@ static int otx2_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 	if (err)
 		goto err_mcs_free;
 
+	err = otx2_init_tc(pf);
+	if (err)
+		goto err_ipsec_clean;
+
 	err = register_netdev(netdev);
 	if (err) {
 		dev_err(dev, "Failed to register netdevice\n");
-		goto err_ipsec_clean;
+		goto err_shutdown_tc;
 	}
 
 	err = otx2_wq_init(pf);
@@ -3380,10 +3401,6 @@ static int otx2_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 
 	otx2_set_ethtool_ops(netdev);
 
-	err = otx2_init_tc(pf);
-	if (err)
-		goto err_mcam_flow_del;
-
 	err = otx2_register_dl(pf);
 	if (err)
 		goto err_mcam_flow_del;
@@ -3420,11 +3437,13 @@ static int otx2_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 	otx2_sriov_vfcfg_cleanup(pf);
 err_pf_sriov_init:
 	otx2_unregister_dl(pf);
-	otx2_shutdown_tc(pf);
 err_mcam_flow_del:
 	otx2_mcam_flow_del(pf);
 err_unreg_netdev:
+	otx2_shutdown_tc_mqprio(pf);
 	unregister_netdev(netdev);
+err_shutdown_tc:
+	otx2_shutdown_tc(pf);
 err_ipsec_clean:
 	cn10k_ipsec_clean(pf);
 err_mcs_free:
@@ -3622,7 +3641,9 @@ static void otx2_remove(struct pci_dev *pdev)
 	otx2_cgx_config_linkevents(pf, false);
 
 	otx2_unregister_dl(pf);
+	otx2_shutdown_tc_mqprio(pf);
 	unregister_netdev(netdev);
+	otx2_shutdown_tc(pf);
 	cn10k_ipsec_clean(pf);
 	cn10k_mcs_free(pf);
 	otx2_sriov_disable(pf->pdev);
@@ -3632,7 +3653,6 @@ static void otx2_remove(struct pci_dev *pdev)
 
 	otx2_ptp_destroy(pf);
 	otx2_mcam_flow_del(pf);
-	otx2_shutdown_tc(pf);
 	otx2_shutdown_qos(pf);
 	otx2_ndc_sync(pf);
 	otx2_detach_resources(&pf->mbox);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
index 8877af348a09..c05a7009a863 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
@@ -6,6 +6,8 @@
  */
 
 #include <linux/netdevice.h>
+#include <linux/rtnetlink.h>
+#include <linux/string.h>
 #include <linux/etherdevice.h>
 #include <linux/inetdevice.h>
 #include <linux/rhashtable.h>
@@ -16,6 +18,7 @@
 #include <net/tc_act/tc_mirred.h>
 #include <net/tc_act/tc_vlan.h>
 #include <net/ipv6.h>
+#include <net/pkt_sched.h>
 
 #include "cn10k.h"
 #include "otx2_common.h"
@@ -31,6 +34,19 @@
 
 #define MCAST_INVALID_GRP		(-1U)
 #define RATE_MANTISSA_BITS		8
+/* Min per-queue egress shaping rate the NIX TLX encoder supports (2 Mbps). */
+#define OTX2_MQPRIO_MIN_RATE_BYTES_PS	250000ULL
+
+static u64 otx2_mqprio_max_rate_bytes_ps(void)
+{
+	u64 max_mbps;
+
+	/* NIX TLX maximum rate (Mbps), not the burst bucket cap. */
+	max_mbps = 2ULL * ((256ULL + MAX_RATE_MANTISSA) << MAX_RATE_EXPONENT) /
+		   256ULL;
+
+	return div_u64(max_mbps * 1000000ULL, 8ULL);
+}
 
 static void otx2_get_egress_burst_cfg(struct otx2_nic *nic, u32 burst,
 				      u32 *burst_exp, u32 *burst_mantissa)
@@ -61,6 +77,9 @@ static void otx2_get_egress_burst_cfg(struct otx2_nic *nic, u32 burst,
 			*burst_mantissa = tmp / (1ULL << (*burst_exp - 7));
 		}
 	} else {
+		/* burst 0: largest encodable burst (CN10K_MAX_BURST_SIZE on
+		 * CN10K), not a minimal burst.
+		 */
 		*burst_exp = MAX_BURST_EXPONENT;
 		*burst_mantissa = max_mantissa;
 	}
@@ -1600,14 +1619,907 @@ static int otx2_setup_tc_block(struct net_device *netdev,
 					  nic, nic, ingress);
 }
 
+/* Free the per-queue min/max rate caches. */
+static void otx2_mqprio_free_cache(struct otx2_nic *pfvf)
+{
+	devm_kfree(pfvf->dev, pfvf->mqprio.min_rate);
+	devm_kfree(pfvf->dev, pfvf->mqprio.max_rate);
+	pfvf->mqprio.min_rate = NULL;
+	pfvf->mqprio.max_rate = NULL;
+	pfvf->mqprio.flags = 0;
+}
+
+static int otx2_mqprio_alloc_cache(struct otx2_nic *pfvf, bool replacing)
+{
+	u16 num_txq = pfvf->hw.non_qos_queues;
+	u64 *min_rate, *max_rate;
+
+	if (replacing && pfvf->mqprio.min_rate && pfvf->mqprio.max_rate) {
+		memset(pfvf->mqprio.min_rate, 0,
+		       num_txq * sizeof(*pfvf->mqprio.min_rate));
+		memset(pfvf->mqprio.max_rate, 0,
+		       num_txq * sizeof(*pfvf->mqprio.max_rate));
+		pfvf->mqprio.flags = 0;
+		return 0;
+	}
+
+	min_rate = devm_kcalloc(pfvf->dev, num_txq, sizeof(*min_rate), GFP_KERNEL);
+	max_rate = devm_kcalloc(pfvf->dev, num_txq, sizeof(*max_rate), GFP_KERNEL);
+	if (!min_rate || !max_rate) {
+		devm_kfree(pfvf->dev, min_rate);
+		devm_kfree(pfvf->dev, max_rate);
+		return -ENOMEM;
+	}
+
+	otx2_mqprio_free_cache(pfvf);
+	pfvf->mqprio.min_rate = min_rate;
+	pfvf->mqprio.max_rate = max_rate;
+
+	return 0;
+}
+
+static void otx2_mqprio_snap_free(struct otx2_nic *pfvf,
+				  struct mq_offload_snap **snap)
+{
+	if (!*snap)
+		return;
+
+	devm_kfree(pfvf->dev, *snap);
+	*snap = NULL;
+}
+
+static int otx2_mqprio_snap_copy(struct otx2_nic *pfvf,
+				 struct mq_offload_snap **dst,
+				 const struct tc_mqprio_qopt_offload *mqprio)
+{
+	const struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+	struct mq_offload_snap *snap;
+	int tc;
+
+	if (!*dst) {
+		snap = devm_kzalloc(pfvf->dev, sizeof(*snap), GFP_KERNEL);
+		if (!snap)
+			return -ENOMEM;
+		*dst = snap;
+	} else {
+		snap = *dst;
+	}
+
+	snap->num_tc = qopt->num_tc;
+	snap->flags = mqprio->flags;
+	for (tc = 0; tc < TC_QOPT_MAX_QUEUE; tc++) {
+		snap->count[tc] = qopt->count[tc];
+		snap->offset[tc] = qopt->offset[tc];
+		snap->min_rate[tc] = 0;
+		snap->max_rate[tc] = 0;
+	}
+
+	for (tc = 0; tc < qopt->num_tc; tc++) {
+		if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+			snap->min_rate[tc] = mqprio->min_rate[tc];
+		if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+			snap->max_rate[tc] = mqprio->max_rate[tc];
+	}
+	memcpy(snap->prio_tc_map, qopt->prio_tc_map, sizeof(snap->prio_tc_map));
+
+	return 0;
+}
+
+static int otx2_mqprio_stage_cur(struct otx2_nic *pfvf,
+				 const struct tc_mqprio_qopt_offload *mqprio)
+{
+	return otx2_mqprio_snap_copy(pfvf, &pfvf->cur_mq_snap, mqprio);
+}
+
+static void otx2_mqprio_snap_commit(struct otx2_nic *pfvf)
+{
+	otx2_mqprio_snap_free(pfvf, &pfvf->old_mq_snap);
+	pfvf->old_mq_snap = pfvf->cur_mq_snap;
+	pfvf->cur_mq_snap = NULL;
+}
+
+static void otx2_mqprio_clear_replace_state(struct otx2_nic *pfvf)
+{
+	pfvf->mqprio.replace_setup_done = false;
+	pfvf->mqprio.replace_graft_done = false;
+}
+
+static bool otx2_mqprio_mdq_allocated(struct otx2_nic *pfvf)
+{
+	return pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_MDQ] != 0;
+}
+
+static int otx2_mqprio_restart_netdev(struct net_device *netdev, bool rate_limit);
+
+static void otx2_mqprio_apply_snap_netdev(struct net_device *netdev,
+					  const struct mq_offload_snap *snap)
+{
+	int tc;
+
+	if (!snap)
+		return;
+
+	netdev_set_num_tc(netdev, snap->num_tc);
+	for (tc = 0; tc < snap->num_tc; tc++)
+		netdev_set_tc_queue(netdev, tc, snap->count[tc],
+				    snap->offset[tc]);
+	for (tc = 0; tc < TC_QOPT_BITMASK + 1; tc++)
+		netdev_set_prio_tc_map(netdev, tc, snap->prio_tc_map[tc]);
+}
+
+static void otx2_mqprio_netdev_tc_work(struct work_struct *work)
+{
+	struct otx2_mqprio *mqprio = container_of(work, struct otx2_mqprio,
+						  netdev_tc_work);
+	struct otx2_nic *pfvf = container_of(mqprio, struct otx2_nic, mqprio);
+
+	if (READ_ONCE(pfvf->mqprio.defer_tc_work))
+		return;
+
+	rtnl_lock();
+	if (READ_ONCE(pfvf->mqprio.defer_tc_work))
+		goto out;
+	if (pfvf->mqprio.rate_limit && pfvf->old_mq_snap)
+		otx2_mqprio_apply_snap_netdev(pfvf->netdev, pfvf->old_mq_snap);
+out:
+	rtnl_unlock();
+}
+
+static void otx2_mqprio_defer_netdev_tc_restore(struct otx2_nic *pfvf)
+{
+	if (READ_ONCE(pfvf->mqprio.defer_tc_work))
+		return;
+	schedule_work(&pfvf->mqprio.netdev_tc_work);
+}
+
+static int otx2_mqprio_restore_old(struct otx2_nic *pfvf)
+{
+	struct mq_offload_snap *snap = pfvf->old_mq_snap;
+	struct net_device *netdev = pfvf->netdev;
+	u16 num_txq = pfvf->hw.non_qos_queues;
+	int tc, txq, err;
+
+	if (!snap)
+		return 0;
+
+	err = otx2_mqprio_alloc_cache(pfvf, false);
+	if (err) {
+		netdev_err(netdev,
+			   "mqprio: rollback failed to allocate rate cache: %d\n",
+			   err);
+		return err;
+	}
+
+	memset(pfvf->mqprio.min_rate, 0, num_txq * sizeof(*pfvf->mqprio.min_rate));
+	memset(pfvf->mqprio.max_rate, 0, num_txq * sizeof(*pfvf->mqprio.max_rate));
+	pfvf->mqprio.flags = snap->flags;
+
+	for (tc = 0; tc < snap->num_tc; tc++) {
+		u64 min_rate = snap->min_rate[tc];
+		u64 max_rate = snap->max_rate[tc];
+
+		for (txq = snap->offset[tc];
+		     txq < snap->offset[tc] + snap->count[tc]; txq++) {
+			pfvf->mqprio.min_rate[txq] = min_rate;
+			pfvf->mqprio.max_rate[txq] = max_rate;
+		}
+	}
+
+	otx2_mqprio_apply_snap_netdev(netdev, snap);
+
+	if (otx2_mqprio_mdq_allocated(pfvf)) {
+		err = otx2_nix_tm_clear_queue_shaper(pfvf);
+		if (err) {
+			netdev_err(netdev,
+				   "mqprio: rollback shaper clear failed: %d; hardware limits may not match qdisc\n",
+				   err);
+			return err;
+		}
+	}
+
+	/* Rebuild the TX scheduler via netdev restart when running; otx2_mqprio_up()
+	 * alone is insufficient after a failed replace that already bounced the
+	 * interface. If open failed, TX schedulers were freed; defer shaper restore
+	 * to the next successful ndo_open() via otx2_mqprio_up().
+	 */
+	pfvf->mqprio.rate_limit = true;
+
+	if (netif_running(netdev)) {
+		err = otx2_mqprio_restart_netdev(netdev, true);
+		if (err) {
+			netdev_err(netdev,
+				   "mqprio: rollback netdev restart failed: %d; qdisc rates may not be enforced\n",
+				   err);
+			return err;
+		}
+	} else if (pfvf->hw.txschq_cnt[NIX_TXSCH_LVL_SMQ]) {
+		err = otx2_mqprio_up(pfvf);
+		if (err) {
+			netdev_err(netdev,
+				   "mqprio: rollback shaper restore failed: %d; qdisc rates may not be enforced\n",
+				   err);
+			return err;
+		}
+	}
+
+	otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
+
+	return 0;
+}
+
+static void otx2_mqprio_snap_destroy(struct otx2_nic *pfvf)
+{
+	otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
+	otx2_mqprio_snap_free(pfvf, &pfvf->old_mq_snap);
+}
+
+/* Offloaded mqprio replaced by software mqprio installs netdev TC layout in
+ * mqprio_init() before the old offload instance is destroyed during graft.
+ */
+static bool otx2_mqprio_keep_netdev_tc(struct otx2_nic *pfvf)
+{
+	struct Qdisc *qdisc = rtnl_dereference(pfvf->netdev->qdisc);
+
+	return qdisc && qdisc->ops && !strcmp(qdisc->ops->id, "mqprio");
+}
+
+static void otx2_mqprio_clear_sw(struct otx2_nic *pfvf)
+{
+	struct net_device *netdev = pfvf->netdev;
+
+	pfvf->mqprio.rate_limit = false;
+	otx2_mqprio_clear_replace_state(pfvf);
+	if (!otx2_mqprio_keep_netdev_tc(pfvf))
+		netdev_set_num_tc(netdev, 0);
+	otx2_mqprio_free_cache(pfvf);
+}
+
+/* Tear down mqprio bandwidth offload: clear per-queue shapers,
+ * mqprio_rate_limit, netdev TC mappings, and the cached rates.  Called on
+ * explicit mqprio teardown (tc qdisc del) and error cleanup, not on
+ * routine netdev stop/open cycles where the offload stays active.
+ */
+int otx2_mqprio_down(struct otx2_nic *pfvf)
+{
+	int err = 0;
+
+	if (!pfvf->mqprio.rate_limit)
+		return 0;
+
+	if (netif_running(pfvf->netdev) &&
+	    otx2_mqprio_mdq_allocated(pfvf))
+		err = otx2_nix_tm_clear_queue_shaper(pfvf);
+
+	if (err) {
+		netdev_warn(pfvf->netdev,
+			    "mqprio: failed to clear hardware shapers: %d; keeping offload state\n",
+			    err);
+		return err;
+	}
+
+	otx2_mqprio_clear_sw(pfvf);
+
+	return 0;
+}
+
+/* Restore cached mqprio MDQ shapers after ndo_open() reprograms the TX
+ * scheduler. Called from otx2_open() when bandwidth offload stays active
+ * across admin down/up or an mqprio netdev bounce.
+ *
+ * Returns an error if any shaper mailbox operation fails. otx2_open()
+ * fail-closes on that error: it aborts open and leaves the interface down
+ * rather than running with partial or missing bandwidth limits.
+ */
+int otx2_mqprio_up(struct otx2_nic *pfvf)
+{
+	struct net_device *netdev = pfvf->netdev;
+	int txq, err;
+
+	if (!pfvf->mqprio.rate_limit)
+		return 0;
+
+	if (!pfvf->mqprio.min_rate || !pfvf->mqprio.max_rate)
+		return 0;
+
+	for (txq = 0; txq < pfvf->hw.non_qos_queues; txq++) {
+		u64 min_rate = 0, max_rate = 0;
+
+		if (pfvf->mqprio.flags & TC_MQPRIO_F_MIN_RATE)
+			min_rate = pfvf->mqprio.min_rate[txq];
+		if (pfvf->mqprio.flags & TC_MQPRIO_F_MAX_RATE)
+			max_rate = pfvf->mqprio.max_rate[txq];
+
+		if (!min_rate && !max_rate)
+			continue;
+
+		err = otx2_nix_tm_set_queue_shaper(pfvf, txq, min_rate,
+						   max_rate);
+		if (err) {
+			netdev_err(netdev,
+				   "mqprio: failed to restore shaper for txq %d: %d\n",
+				   txq, err);
+			if (otx2_mqprio_mdq_allocated(pfvf) &&
+			    otx2_nix_tm_clear_queue_shaper(pfvf))
+				netdev_warn(netdev,
+					    "mqprio: failed to clear shapers after partial restore\n");
+			return err;
+		}
+	}
+
+	return 0;
+}
+
+/* Restart the netdev to reprogram the TX scheduler hierarchy for mqprio
+ * bandwidth offload.  Both mqprio add and delete (when offload was active)
+ * take this path via ndo_stop()/ndo_open() so VF-specific open logic (e.g.
+ * LBK carrier on) runs correctly.
+ *
+ * Intentional behaviour: this full stop/open cycle drops in-flight traffic
+ * (carrier off, IRQ/NAPI teardown, queue drain).  The NIX TX scheduler must
+ * be reallocated (e.g. one SMQ per non-QoS queue) and cannot be reprogrammed
+ * live today, so a netdev bounce is required on every mqprio add, replace,
+ * delete, and rollback.  Users see a brief connectivity blip; this is not a
+ * bug to "fix" without implementing the live-reprogramming path noted below.
+ * If open fails, OTX2_FLAG_INTF_DOWN is set before netif_close() so
+ * otx2_stop() returns early and does not touch resources already torn
+ * down by the open error path.
+ *
+ * Do not call dev_deactivate()/dev_activate() here.  On replace,
+ * qdisc_graft() already deactivates qdiscs around offload teardown;
+ * dev_activate() from ndo_setup_tc() would republish qdiscs before graft
+ * completes and race __qdisc_run() on the old root qdisc.  After
+ * ndo_open(), carrier and TX queues are restored via otx2_handle_link_event()
+ * when link is up, same as otx2_change_mtu(), not via dev_activate().
+ *
+ * Clear __LINK_STATE_START before ndo_stop() so netif_running() is false
+ * for the duration of the bounce.
+ */
+static int otx2_mqprio_restart_netdev(struct net_device *netdev, bool rate_limit)
+{
+	struct otx2_nic *pfvf = netdev_priv(netdev);
+	const struct net_device_ops *ops = netdev->netdev_ops;
+	bool running = netif_running(netdev);
+	int err;
+
+	/* TODO: Explore live TX scheduler reprogramming to avoid a full
+	 * ndo_stop()/ndo_open() bounce on every mqprio change.
+	 */
+	netdev_dbg(netdev,
+		   "mqprio: restarting interface to reprogram TX scheduler; in-flight traffic will be dropped\n");
+
+	if (running) {
+		clear_bit(__LINK_STATE_START, &netdev->state);
+		smp_mb__after_atomic(); /* Commit netif_running(). */
+	}
+
+	err = ops->ndo_stop(netdev);
+	if (err) {
+		if (running)
+			set_bit(__LINK_STATE_START, &netdev->state);
+		return err;
+	}
+
+	/* Set before ndo_open() so otx2_txsch_alloc() widens SMQ allocation.
+	 * On teardown, drop mqprio software state so ndo_open() does not
+	 * re-apply bandwidth limits via otx2_mqprio_up() after the kernel
+	 * removed the qdisc.
+	 */
+	if (rate_limit)
+		pfvf->mqprio.rate_limit = true;
+	else
+		otx2_mqprio_clear_sw(pfvf);
+
+	err = ops->ndo_open(netdev);
+	if (!err && running) {
+		set_bit(__LINK_STATE_START, &netdev->state);
+	} else if (err) {
+		netdev_err(netdev,
+			   "Failed to restart device after mqprio change: %d\n",
+			   err);
+		/* ndo_open() already tore down TX scheduler resources on failure;
+		 * netif_running() is false here (__LINK_STATE_START stays clear
+		 * until open succeeds). Drop mqprio software state only instead
+		 * of sending shaper clears to freed queues.
+		 */
+		otx2_mqprio_clear_sw(pfvf);
+		/* ndo_open() rolls back on failure; mark the interface down so
+		 * otx2_stop() returns early when netif_close() runs.  Caller
+		 * holds RTNL; dev_close() would deadlock.
+		 */
+		otx2_set_flag(pfvf, OTX2_FLAG_INTF_DOWN);
+		/* visible to otx2_stop() on other cpus */
+		smp_wmb();
+		netif_close(netdev);
+	}
+
+	return err;
+}
+
+static int otx2_mqprio_validate_tc_rate(struct net_device *netdev,
+					struct netlink_ext_ack *extack,
+					u64 rate, u32 qcount, int tc,
+					const char *name)
+{
+	if (!rate)
+		return 0;
+
+	if (qcount <= 1)
+		return 0;
+
+	/* TODO: per-TC TL4 shapers or equal per-queue MDQ split for multi-queue TC rates. */
+	netdev_err(netdev,
+		   "mqprio: %s rate for tc %d not supported with %u queues\n",
+		   name, tc, qcount);
+	NL_SET_ERR_MSG_FMT_MOD(extack,
+			       "mqprio: tc %d rate needs 1 queue (has %u)",
+			       tc, qcount);
+	return -EOPNOTSUPP;
+}
+
+static int otx2_mqprio_validate_txqs(struct net_device *netdev,
+				     struct netlink_ext_ack *extack,
+				     struct tc_mqprio_qopt *qopt)
+{
+	struct otx2_nic *pfvf = netdev_priv(netdev);
+	u16 num_txq = pfvf->hw.non_qos_queues;
+	int tc, txq;
+
+	if (qopt->num_tc > num_txq) {
+		netdev_err(netdev, "Number of TCs (%u) exceeds hw queues %u\n",
+			   qopt->num_tc, num_txq);
+		NL_SET_ERR_MSG_FMT_MOD(extack,
+				       "TCs (%u) exceed hw queues %u",
+				       qopt->num_tc, num_txq);
+		return -EINVAL;
+	}
+
+	if (num_txq > MAX_TXSCHQ_PER_FUNC) {
+		netdev_err(netdev,
+			   "Number of queues (%u) exceeds max scheduler queues %u\n",
+			   num_txq, MAX_TXSCHQ_PER_FUNC);
+		NL_SET_ERR_MSG_FMT_MOD(extack,
+				       "queues (%u) exceed max schq %u",
+				       num_txq, MAX_TXSCHQ_PER_FUNC);
+		return -EINVAL;
+	}
+
+	for (tc = 0; tc < qopt->num_tc; tc++) {
+		u32 qcount = qopt->count[tc];
+
+		for (txq = qopt->offset[tc];
+		     txq < qopt->offset[tc] + qcount; txq++) {
+			if (txq >= num_txq) {
+				netdev_err(netdev,
+					   "mqprio: txq %d exceeds offload queue count %u\n",
+					   txq, num_txq);
+				NL_SET_ERR_MSG_FMT_MOD(extack,
+						       "mqprio: txq %d exceeds queues %u",
+						       txq, num_txq);
+				return -EINVAL;
+			}
+		}
+	}
+
+	return 0;
+}
+
+static bool otx2_mqprio_rate_valid(struct otx2_nic *pfvf, u64 rate_bytes_ps)
+{
+	u64 mbps;
+
+	if (!rate_bytes_ps)
+		return true;
+
+	if (rate_bytes_ps < OTX2_MQPRIO_MIN_RATE_BYTES_PS)
+		return false;
+
+	if (rate_bytes_ps > otx2_mqprio_max_rate_bytes_ps())
+		return false;
+
+	if (rate_bytes_ps > div_u64(U64_MAX, 8))
+		return false;
+
+	mbps = otx2_convert_rate(rate_bytes_ps);
+	return ilog2(mbps / 2) <= MAX_RATE_EXPONENT;
+}
+
+static bool otx2_tc_can_offload(struct net_device *netdev)
+{
+	return !!(netdev->hw_features & NETIF_F_HW_TC);
+}
+
+static int otx2_teardown_tc_mqprio(struct otx2_nic *pfvf,
+				   struct tc_mqprio_qopt_offload *mqprio)
+{
+	struct netlink_ext_ack *extack = mqprio->extack;
+	struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+	bool had_mqprio = pfvf->mqprio.rate_limit;
+	struct net_device *netdev = pfvf->netdev;
+	bool if_up = netif_running(netdev);
+	int err;
+
+	if (!otx2_tc_can_offload(netdev))
+		return -EOPNOTSUPP;
+
+	qopt->hw = 0;
+
+	/* tc qdisc replace runs setup on the new mqprio before destroying the
+	 * old one. replace_setup_done and TC_ROOT_GRAFT distinguish stale
+	 * old-instance teardown from graft failure after setup.
+	 * qdisc_offload_graft_helper() skips TC_ROOT_GRAFT unless hw-tc-offload
+	 * is enabled in netdev->features; commit without graft in that case.
+	 */
+	if (pfvf->mqprio.replace_setup_done && pfvf->cur_mq_snap) {
+		err = 0;
+		if (pfvf->mqprio.replace_graft_done || !tc_can_offload(netdev)) {
+			otx2_mqprio_snap_commit(pfvf);
+		} else {
+			err = otx2_mqprio_restore_old(pfvf);
+			if (err && extack)
+				NL_SET_ERR_MSG_FMT_MOD(extack,
+						       "mqprio: rollback failed (%d)",
+						       err);
+		}
+		otx2_mqprio_clear_replace_state(pfvf);
+		return err;
+	}
+
+	/* Skip the netdev restart when mqprio offload was not active. */
+	if (!had_mqprio)
+		return 0;
+
+	if (if_up) {
+		err = otx2_mqprio_down(pfvf);
+		if (err) {
+			if (extack)
+				NL_SET_ERR_MSG_FMT_MOD(extack,
+						       "mqprio: teardown incomplete (%d)",
+						       err);
+			netdev_err(netdev,
+				   "mqprio: offload teardown incomplete (%d); driver state may not match removed qdisc\n",
+				   err);
+			return err;
+		}
+
+		err = otx2_mqprio_restart_netdev(netdev, false);
+		if (err) {
+			if (extack)
+				NL_SET_ERR_MSG_FMT_MOD(extack,
+						       "mqprio: teardown restart failed (%d)",
+						       err);
+			netdev_err(netdev,
+				   "mqprio: offload teardown restart failed (%d); driver state may not match removed qdisc\n",
+				   err);
+			return err;
+		}
+		return 0;
+	}
+
+	/* ndo_stop() already freed the TX scheduler TL nodes; drop software
+	 * state only.
+	 */
+	otx2_mqprio_clear_sw(pfvf);
+	return 0;
+}
+
+static int otx2_setup_tc_mqprio(struct net_device *netdev,
+				struct tc_mqprio_qopt_offload *mqprio)
+{
+	struct netlink_ext_ack *extack = mqprio->extack;
+	struct otx2_nic *pfvf = netdev_priv(netdev);
+	struct tc_mqprio_qopt *qopt = &mqprio->qopt;
+	bool replacing = pfvf->mqprio.rate_limit;
+	bool if_up = netif_running(netdev);
+	int tc, txq, err, i;
+
+	if (!otx2_tc_can_offload(netdev))
+		return -EOPNOTSUPP;
+
+	if (!qopt->hw)
+		return otx2_teardown_tc_mqprio(pfvf, mqprio);
+
+	if (!if_up) {
+		netdev_err(netdev, "mqprio: setup requires interface UP\n");
+		NL_SET_ERR_MSG_MOD(extack, "mqprio: setup requires interface UP");
+		err = -EOPNOTSUPP;
+		goto fail_validate;
+	}
+
+	if (!replacing && otx2_mqprio_keep_netdev_tc(pfvf)) {
+		netdev_err(netdev,
+			   "mqprio: delete existing mqprio before re-enabling hw/sw offload\n");
+		NL_SET_ERR_MSG_MOD(extack,
+				   "mqprio: delete existing mqprio before re-enabling hw/sw offload");
+		return -EOPNOTSUPP;
+	}
+
+	if (mqprio->shaper != TC_MQPRIO_SHAPER_BW_RATE) {
+		netdev_err(netdev, "Unsupported mqprio shaper %#x\n", mqprio->shaper);
+		NL_SET_ERR_MSG_FMT_MOD(extack, "mqprio: bad shaper %#x",
+				       mqprio->shaper);
+		err = -EOPNOTSUPP;
+		goto fail_validate;
+	}
+
+	if (!test_bit(QOS_CIR_PIR_SUPPORT, &pfvf->hw.cap_flag)) {
+		netdev_err(netdev,
+			   "mqprio: bandwidth offload requires CIR+PIR support\n");
+		NL_SET_ERR_MSG_MOD(extack,
+				   "mqprio: bandwidth offload requires CIR+PIR support");
+		err = -EOPNOTSUPP;
+		goto fail_validate;
+	}
+
+	if (is_otx2_sdp_rep(pfvf->pdev)) {
+		netdev_err(netdev, "mqprio: bandwidth offload not supported on SDP rep\n");
+		NL_SET_ERR_MSG_MOD(extack,
+				   "mqprio: bandwidth offload not supported on SDP rep");
+		err = -EOPNOTSUPP;
+		goto fail_validate;
+	}
+
+	if (pfvf->pfc_en) {
+		netdev_err(netdev,
+			   "mqprio: cannot enable offload while PFC is enabled\n");
+		NL_SET_ERR_MSG_MOD(extack,
+				   "mqprio: cannot enable offload while PFC is enabled");
+		err = -EOPNOTSUPP;
+		goto fail_validate;
+	}
+
+	if (pfvf->xdp_prog) {
+		netdev_err(netdev,
+			   "mqprio: cannot enable offload while XDP is active\n");
+		NL_SET_ERR_MSG_MOD(extack,
+				   "mqprio: cannot enable offload while XDP is active");
+		err = -EOPNOTSUPP;
+		goto fail_validate;
+	}
+
+	if (!list_empty(&pfvf->qos.qos_tree)) {
+		netdev_err(netdev,
+			   "mqprio: cannot enable offload while HTB is active\n");
+		NL_SET_ERR_MSG_MOD(extack,
+				   "mqprio: cannot enable offload while HTB is active");
+		err = -EOPNOTSUPP;
+		goto fail_validate;
+	}
+
+	for (tc = 0; tc < qopt->num_tc; tc++) {
+		u64 min_rate = 0, max_rate = 0;
+		u32 qcount = qopt->count[tc];
+
+		if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+			min_rate = mqprio->min_rate[tc];
+		if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+			max_rate = mqprio->max_rate[tc];
+
+		if (min_rate && max_rate && min_rate > max_rate) {
+			netdev_err(netdev,
+				   "min_rate %llu exceeds max_rate %llu for tc %d\n",
+				   min_rate, max_rate, tc);
+			NL_SET_ERR_MSG_FMT_MOD(extack,
+					       "mqprio: min>max rate for tc %d",
+					       tc);
+			err = -EINVAL;
+			goto fail_validate;
+		}
+
+		if (mqprio->flags & TC_MQPRIO_F_MIN_RATE) {
+			err = otx2_mqprio_validate_tc_rate(netdev, extack, min_rate,
+							   qcount, tc, "min");
+			if (err)
+				goto fail_validate;
+		}
+
+		if (mqprio->flags & TC_MQPRIO_F_MAX_RATE) {
+			err = otx2_mqprio_validate_tc_rate(netdev, extack, max_rate,
+							   qcount, tc, "max");
+			if (err)
+				goto fail_validate;
+		}
+
+		if (mqprio->flags & TC_MQPRIO_F_MIN_RATE &&
+		    !otx2_mqprio_rate_valid(pfvf, min_rate)) {
+			netdev_err(netdev,
+				   "mqprio: min_rate %llu for tc %d is outside hardware limits\n",
+				   min_rate, tc);
+			NL_SET_ERR_MSG_FMT_MOD(extack,
+					       "mqprio: min_rate out of range (tc %d)",
+					       tc);
+			err = -EINVAL;
+			goto fail_validate;
+		}
+
+		if (mqprio->flags & TC_MQPRIO_F_MAX_RATE &&
+		    !otx2_mqprio_rate_valid(pfvf, max_rate)) {
+			netdev_err(netdev,
+				   "mqprio: max_rate %llu for tc %d is outside hardware limits\n",
+				   max_rate, tc);
+			NL_SET_ERR_MSG_FMT_MOD(extack,
+					       "mqprio: max_rate out of range (tc %d)",
+					       tc);
+			err = -EINVAL;
+			goto fail_validate;
+		}
+	}
+
+	err = otx2_mqprio_validate_txqs(netdev, extack, qopt);
+	if (err)
+		goto fail_validate;
+
+	err = otx2_mqprio_stage_cur(pfvf, mqprio);
+	if (err)
+		goto fail_validate;
+
+	err = otx2_mqprio_restart_netdev(pfvf->netdev, true);
+	if (err)
+		goto cleanup;
+
+	err = otx2_mqprio_alloc_cache(pfvf, replacing);
+	if (err)
+		goto cleanup;
+
+	/* otx2_mqprio_up() may have restored the previous configuration during
+	 * the restart above. Clear every MDQ shaper before applying the new
+	 * mapping so queues dropped from the TC layout do not keep stale
+	 * limits in hardware.
+	 */
+	if (otx2_mqprio_mdq_allocated(pfvf)) {
+		err = otx2_nix_tm_clear_queue_shaper(pfvf);
+		if (err) {
+			if (extack)
+				NL_SET_ERR_MSG_FMT_MOD(extack,
+						       "mqprio: shaper clear failed (%d)",
+						       err);
+			goto cleanup;
+		}
+	}
+
+	pfvf->mqprio.flags = mqprio->flags;
+
+	for (tc = 0; tc < qopt->num_tc; tc++) {
+		u64 min_rate = 0, max_rate = 0;
+		u32 qcount = qopt->count[tc];
+
+		/* Rates omitted from tc mqprio are passed as zero and both MDQ
+		 * shaper registers are programmed; see
+		 * otx2_nix_tm_set_queue_shaper().
+		 * TODO: multi-queue TC rates need per-queue split; see
+		 * otx2_mqprio_validate_tc_rate().
+		 */
+		if (mqprio->flags & TC_MQPRIO_F_MIN_RATE)
+			min_rate = mqprio->min_rate[tc];
+		if (mqprio->flags & TC_MQPRIO_F_MAX_RATE)
+			max_rate = mqprio->max_rate[tc];
+
+		for (txq = qopt->offset[tc];
+		     txq < qopt->offset[tc] + qcount; txq++) {
+			netdev_dbg(netdev,
+				   "mqprio: tc %d txq %d min_rate %llu max_rate %llu\n",
+				   tc, txq, min_rate, max_rate);
+
+			pfvf->mqprio.min_rate[txq] = min_rate;
+			pfvf->mqprio.max_rate[txq] = max_rate;
+
+			err = otx2_nix_tm_set_queue_shaper(pfvf, txq,
+							   min_rate, max_rate);
+			if (err) {
+				if (extack)
+					NL_SET_ERR_MSG_FMT_MOD(extack,
+							       "mqprio: shaper program failed (%d)",
+							       err);
+				netdev_err(netdev,
+					   "mqprio: programming shaper failed (%d); partial limits may be active\n",
+					   err);
+				goto cleanup;
+			}
+		}
+	}
+
+	netdev_set_num_tc(netdev, pfvf->cur_mq_snap->num_tc);
+	for (i = 0; i < pfvf->cur_mq_snap->num_tc; i++)
+		netdev_set_tc_queue(netdev, i, pfvf->cur_mq_snap->count[i],
+				    qopt->offset[i]);
+
+	qopt->hw = TC_MQPRIO_HW_OFFLOAD_TCS;
+
+	if (replacing) {
+		pfvf->mqprio.replace_setup_done = true;
+		pfvf->mqprio.replace_graft_done = false;
+	} else {
+		otx2_mqprio_snap_commit(pfvf);
+	}
+
+	netdev_warn_once(netdev,
+			 "mqprio: bandwidth offload uses TL4[0] without explicit topology programming; verify scheduling if limits look wrong\n");
+
+	return 0;
+
+fail_validate:
+	/* Failed replace destroys the new qdisc with hw_offload unset, so
+	 * mqprio_destroy() clears netdev TC after we return. Re-apply the
+	 * prior layout when validation fails before any hardware change.
+	 */
+	if (replacing)
+		otx2_mqprio_defer_netdev_tc_restore(pfvf);
+	return err;
+
+cleanup:
+	qopt->hw = 0;
+	if (replacing) {
+		int restore_err = otx2_mqprio_restore_old(pfvf);
+
+		otx2_mqprio_clear_replace_state(pfvf);
+		if (restore_err) {
+			netdev_err(netdev,
+				   "mqprio: replace failed and prior configuration rollback failed: %d\n",
+				   restore_err);
+			if (extack)
+				NL_SET_ERR_MSG_FMT_MOD(extack,
+						       "mqprio: rollback failed (%d)",
+						       restore_err);
+		} else {
+			netdev_err(netdev,
+				   "mqprio: replace failed; hardware limits restored, netdev TC layout restore deferred\n");
+			if (extack)
+				NL_SET_ERR_MSG_MOD(extack,
+						   "mqprio: replace failed; hardware limits restored, netdev TC layout restore deferred");
+			/* Failed replace destroys the new qdisc with hw_offload
+			 * unset, so mqprio_destroy() clears netdev TC after we
+			 * return. Re-apply the restored layout once that unwind
+			 * finishes.
+			 */
+			otx2_mqprio_defer_netdev_tc_restore(pfvf);
+		}
+		return err ? err : -EIO;
+	}
+	otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
+	otx2_teardown_tc_mqprio(pfvf, mqprio);
+	return err;
+}
+
+static int otx2_setup_tc_root(struct otx2_nic *pfvf,
+			      struct tc_root_qopt_offload *root)
+{
+	switch (root->command) {
+	case TC_ROOT_GRAFT:
+		if (pfvf->mqprio.replace_setup_done)
+			pfvf->mqprio.replace_graft_done = true;
+		return 0;
+	default:
+		return -EOPNOTSUPP;
+	}
+}
+
+static int otx2_setup_tc_query_caps(void *type_data)
+{
+	struct tc_query_caps_base *base = type_data;
+	struct tc_mqprio_caps *caps;
+
+	if (base->type != TC_SETUP_QDISC_MQPRIO)
+		return -EOPNOTSUPP;
+
+	caps = base->caps;
+	caps->validate_queue_counts = true;
+
+	return 0;
+}
+
 int otx2_setup_tc(struct net_device *netdev, enum tc_setup_type type,
 		  void *type_data)
 {
 	switch (type) {
+	case TC_QUERY_CAPS:
+		return otx2_setup_tc_query_caps(type_data);
 	case TC_SETUP_BLOCK:
 		return otx2_setup_tc_block(netdev, type_data);
 	case TC_SETUP_QDISC_HTB:
 		return otx2_setup_tc_htb(netdev, type_data);
+	case TC_SETUP_QDISC_MQPRIO:
+		return otx2_setup_tc_mqprio(netdev, type_data);
+	case TC_SETUP_ROOT_QDISC:
+		return otx2_setup_tc_root(netdev_priv(netdev), type_data);
 	default:
 		return -EOPNOTSUPP;
 	}
@@ -1625,10 +2537,25 @@ int otx2_init_tc(struct otx2_nic *nic)
 		return -EINVAL;
 	}
 
+	INIT_WORK(&nic->mqprio.netdev_tc_work, otx2_mqprio_netdev_tc_work);
+	WRITE_ONCE(nic->mqprio.defer_tc_work, false);
+
 	return 0;
 }
 EXPORT_SYMBOL(otx2_init_tc);
 
+void otx2_shutdown_tc_mqprio(struct otx2_nic *nic)
+{
+	WRITE_ONCE(nic->mqprio.defer_tc_work, true);
+	/* Publish defer_tc_work before cancel_work_sync(). */
+	smp_wmb();
+	cancel_work_sync(&nic->mqprio.netdev_tc_work);
+	rtnl_lock();
+	otx2_mqprio_snap_destroy(nic);
+	rtnl_unlock();
+}
+EXPORT_SYMBOL(otx2_shutdown_tc_mqprio);
+
 void otx2_shutdown_tc(struct otx2_nic *nic)
 {
 	otx2_destroy_tc_flow_list(nic);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
index 5f7915231ca3..6ca08b12520a 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_vf.c
@@ -740,25 +740,25 @@ static int otx2vf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 	if (err)
 		goto err_ipsec_clean;
 
+	err = otx2vf_mcam_flow_init(vf);
+	if (err)
+		goto err_wq_destroy;
+
+	err = otx2_init_tc(vf);
+	if (err)
+		goto err_wq_destroy;
+
 	err = register_netdev(netdev);
 	if (err) {
 		dev_err(dev, "Failed to register netdevice\n");
-		goto err_wq_destroy;
+		goto err_shutdown_tc;
 	}
 
 	otx2vf_set_ethtool_ops(netdev);
 
-	err = otx2vf_mcam_flow_init(vf);
-	if (err)
-		goto err_unreg_netdev;
-
-	err = otx2_init_tc(vf);
-	if (err)
-		goto err_unreg_netdev;
-
 	err = otx2_register_dl(vf);
 	if (err)
-		goto err_shutdown_tc;
+		goto err_unreg_netdev;
 
 	vf->af_xdp_zc_qidx = bitmap_zalloc(qcount, GFP_KERNEL);
 	if (!vf->af_xdp_zc_qidx) {
@@ -784,10 +784,11 @@ static int otx2vf_probe(struct pci_dev *pdev, const struct pci_device_id *id)
 #endif
 err_unreg_devlink:
 	otx2_unregister_dl(vf);
-err_shutdown_tc:
-	otx2_shutdown_tc(vf);
 err_unreg_netdev:
+	otx2_shutdown_tc_mqprio(vf);
 	unregister_netdev(netdev);
+err_shutdown_tc:
+	otx2_shutdown_tc(vf);
 err_wq_destroy:
 	cancel_work_sync(&vf->reset_task);
 	cancel_work_sync(&vf->rx_mode_work);
@@ -840,7 +841,9 @@ static void otx2vf_remove(struct pci_dev *pdev)
 #endif
 
 	otx2_unregister_dl(vf);
+	otx2_shutdown_tc_mqprio(vf);
 	unregister_netdev(netdev);
+	otx2_shutdown_tc(vf);
 	if (vf->otx2_wq) {
 		cancel_work_sync(&vf->reset_task);
 		cancel_work_sync(&vf->rx_mode_work);
@@ -849,7 +852,6 @@ static void otx2vf_remove(struct pci_dev *pdev)
 	cn10k_ipsec_clean(vf);
 	otx2_ptp_destroy(vf);
 	otx2_mcam_flow_del(vf);
-	otx2_shutdown_tc(vf);
 	otx2_shutdown_qos(vf);
 	otx2_detach_resources(&vf->mbox);
 	otx2vf_disable_mbox_intr(vf);
diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/qos.c b/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
index f160b1618efa..9ef55a6db50b 100644
--- a/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
+++ b/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
@@ -118,6 +118,9 @@ static void otx2_config_sched_shaping(struct otx2_nic *pfvf,
 	/* configure PIR */
 	maxrate = (node->rate > node->ceil) ? node->rate : node->ceil;
 
+	/* 65536 is the kernel-side default burst when HTB does not supply an
+	 * explicit value, not the NIX hardware maximum (CN10K_MAX_BURST_SIZE).
+	 */
 	cfg->regval[*num_regs] =
 		otx2_get_txschq_rate_regval(pfvf, maxrate, 65536);
 	(*num_regs)++;
@@ -1088,6 +1091,14 @@ static int otx2_qos_root_add(struct otx2_nic *pfvf, u16 htb_maj_id, u16 htb_defc
 		   "TC_HTB_CREATE: handle=0x%x defcls=0x%x\n",
 		   htb_maj_id, htb_defcls);
 
+	if (pfvf->mqprio.rate_limit) {
+		netdev_err(pfvf->netdev,
+			   "HTB: cannot enable while mqprio bandwidth offload is active\n");
+		NL_SET_ERR_MSG_MOD(extack,
+				   "HTB: cannot enable while mqprio bandwidth offload is active");
+		return -EOPNOTSUPP;
+	}
+
 	root = otx2_qos_alloc_root(pfvf);
 	if (IS_ERR(root)) {
 		err = PTR_ERR(root);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [PATCH v19 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags
  2026-10-06  9:43 ` [PATCH v19 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags Ratheesh Kannoth
@ 2026-10-10 10:01   ` netdev-bot+sashiko
  0 siblings, 0 replies; 5+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 10:01 UTC (permalink / raw)
  To: rkannoth
  Cc: bpf, linux-kernel, netdev, andrew+netdev, ast, daniel, davem,
	edumazet, hawk, john.fastabend, kuba, pabeni, sdf, sgoutham

Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.

Critical: 0 · High: 1 (3) · Medium: 1 (5) · Low: 0 (1)

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] otx2_sync_flags_from_rep() (otx2_common.h:625-634) updates the
  shared priv->flags word with a plain read-modify-write:…
- [Medium] The commit message presents this as a cleanup to atomic
  bitops, but the patch also fixes a user-visible representor bug and…

Pre-existing issues:
- [High] Use-after-free of the representor's delayed stats work.
- [High] In rvu_rep_destroy() (rep.c:641-642),
  `free_netdev(rep->netdev);` is followed by `kfree(rep->flow_cfg);`.
- [High] rvu_rep_mcam_flow_init() (rep.c:54-91) builds and sends
  NPC_MCAM_ALLOC_ENTRY messages in the shared PF<->AF mailbox…
- [Medium] otx2_tc_del_flow() (otx2_tc.c:1196-1198) clears
  OTX2_FLAG_TC_MARK_ENABLED when…
- [Medium] rvu_rep_setup_tc_cb() (rep.c:115-116) ignores the return value
  of rvu_rep_mcam_flow_init().
- [Medium] In otx2_config_hwtstamp_set() (otx2_pf.c), the
  HWTSTAMP_TX_ONESTEP_SYNC case sets OTX2_FLAG_PTP_ONESTEP_SYNC and…
- [Medium] Representor flow configuration memory leaks.
- [Medium] A representor's flow_cfg->mark_flows starts at 0 because
  rvu_rep_mcam_flow_init() allocates flow_cfg zeroed and never calls…
- [Low] otx2_config_hwtstamp_set() ignores the return values of
  otx2_config_hw_tx_tstamp() and otx2_config_hw_rx_tstamp().

--

octeontx2: use atomic bitops for PF/VF and rep flags

This converts the otx2_nic and rep_dev flag words from u64 masks to
unsigned long bitmaps that are accessed with set_bit/clear_bit/test_bit.
It adds otx2_sync_flags_from_rep() to copy a masked set of representor
flags onto the shared PF context. It also moves the representor VF
initialized flag to OTX2_FLAG_REP_VF_INITIALIZED (bit 21).

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> index fda6c4487c4fb..3bc57dd104515 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.h

[ ... ]

> @@ -593,6 +607,32 @@ struct otx2_nic {
>  	unsigned long		*af_xdp_zc_qidx;
>  };
>  

[ ... ]

> +static inline void otx2_sync_flags_from_rep(struct otx2_nic *dst,
> +					    unsigned long *src_flags)
> +{
> +	unsigned long src = READ_ONCE(*src_flags);
> +	unsigned long new_flags;
> +
> +	new_flags = (READ_ONCE(dst->flags) & ~OTX2_REP_SYNC_FLAGS_MASK) |
> +		    (src & OTX2_REP_SYNC_FLAGS_MASK);
> +	WRITE_ONCE(dst->flags, new_flags);
> +}

[Severity: High]
Can this read-modify-write lose concurrent updates to PF-owned bits?

Every other writer of priv->flags now uses atomic set_bit()/clear_bit().
This helper instead reads the whole word, masks it, and writes the whole
word back.

rvu_rep_setup_tc_cb() runs under rtnl. otx2_devlink_eswitch_mode_set()
holds only the devlink instance lock:

CPU0 (tc filter add ... dev <rep>, rtnl held)
rvu_rep_setup_tc_cb()
  otx2_sync_flags_from_rep()
    READ_ONCE(dst->flags)

CPU1 (devlink dev eswitch set ... mode legacy, no rtnl)
otx2_devlink_eswitch_mode_set()
  rvu_rep_destroy()
    otx2_set_flag(priv, OTX2_FLAG_INTF_DOWN)

CPU0
    WRITE_ONCE(dst->flags, new_flags)

INTF_DOWN is outside OTX2_REP_SYNC_FLAGS_MASK, but CPU0's stale word still
overwrites it.

If that happens, otx2_napi_handler() no longer sees INTF_DOWN during
teardown and can re-enable CQ interrupts. On a later PCI remove,
rvu_rep_remove() also sees INTF_DOWN clear and runs teardown again:

    if (!otx2_test_flag(priv, OTX2_FLAG_INTF_DOWN))
        rvu_rep_destroy(priv);

The second rvu_rep_destroy() would walk qset->napi in
rvu_rep_free_cq_rsrc() after otx2_free_queue_mem() has already set it to
NULL. It would then unregister and free priv->reps[] and priv->reps a
second time.

The reverse case also looks possible. rvu_rep_napi_init() runs
otx2_clear_flag(priv, OTX2_FLAG_INTF_DOWN) after the rep netdevs are
registered, so they can already get TC callbacks. If that clear is lost,
CQ interrupts are never re-enabled and the representor datapath stalls.

The commit message says this helper publishes the flags "without clearing
PF-owned bits". For that to hold, does this need a try_cmpxchg() loop on
dst->flags, or assign_bit() for each bit in OTX2_REP_SYNC_FLAGS_MASK?

[ ... ]

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> index 7dc0ce669d0b7..777e7156badb7 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c

[ ... ]

> @@ -2511,9 +2511,9 @@ static int otx2_config_hw_tx_tstamp(struct otx2_nic *pfvf, bool enable)
>  
>  	mutex_unlock(&pfvf->mbox.lock);
>  	if (enable)
> -		pfvf->flags |= OTX2_FLAG_TX_TSTAMP_ENABLED;
> +		otx2_set_flag(pfvf, OTX2_FLAG_TX_TSTAMP_ENABLED);
>  	else
> -		pfvf->flags &= ~OTX2_FLAG_TX_TSTAMP_ENABLED;
> +		otx2_clear_flag(pfvf, OTX2_FLAG_TX_TSTAMP_ENABLED);
>  	return 0;
>  }

[Severity: Low]
This isn't a bug introduced by this patch, but otx2_config_hw_tx_tstamp()
and otx2_config_hw_rx_tstamp() can return -ENOMEM or a mailbox error
before they update these flags. otx2_config_hwtstamp_set() ignores both
return values and finishes with:

    pfvf->tstamp = *config;

    return 0;

Does this report a failed enable to user space as success? The getter
would then return a configuration that was never applied.

[ ... ]

> @@ -2550,7 +2550,7 @@ int otx2_config_hwtstamp_set(struct net_device *netdev,
>  					   "One-step time stamping is not supported");
>  			return -ERANGE;
>  		}
> -		pfvf->flags |= OTX2_FLAG_PTP_ONESTEP_SYNC;
> +		otx2_set_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC);
>  		schedule_delayed_work(&pfvf->ptp->synctstamp_work,
>  				      msecs_to_jiffies(500));
>  		fallthrough;

[Severity: Medium]
This is a pre-existing issue, but only HWTSTAMP_TX_OFF clears
OTX2_FLAG_PTP_ONESTEP_SYNC.

Consider a switch straight from HWTSTAMP_TX_ONESTEP_SYNC to
HWTSTAMP_TX_ON. The flag stays set and synctstamp_work is not cancelled.
otx2_config_hw_tx_tstamp() also returns early, because TX timestamping is
already enabled.

otx2_set_txtstamp() then still takes the one-step path for PTP Sync
packets:

    if (unlikely(otx2_test_flag(pfvf, OTX2_FLAG_PTP_ONESTEP_SYNC) &&
                 otx2_ptp_is_sync(skb, &ptp_offset, &udp_csum_crt))) {

Would this rewrite the origin timestamp and skip SKBTX_IN_PROGRESS? If
so, the two-step TX timestamp requested with TX_ON would never be
delivered.

[ ... ]

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> index dee38d5c7f65b..8877af348a090 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c

[ ... ]

> @@ -505,7 +505,7 @@ static int otx2_tc_parse_actions(struct otx2_nic *nic,
>  			mark = act->mark;
>  			req->match_id = mark & OTX2_RX_MATCH_ID_MASK;
>  			req->op = NIX_RX_ACTION_DEFAULT;
> -			nic->flags |= OTX2_FLAG_TC_MARK_ENABLED;
> +			otx2_set_flag(nic, OTX2_FLAG_TC_MARK_ENABLED);
>  			refcount_inc(&nic->flow_cfg->mark_flows);
>  			break;

[Severity: Medium]
This is a pre-existing issue. When nic->flow_cfg comes from a
representor, rvu_rep_mcam_flow_init() allocates it zeroed. Unlike
otx2_mcam_entry_init(), it never calls refcount_set(&flow_cfg->mark_flows, 1).

Won't this refcount_inc() then run on 0 for the first skbedit mark rule
on a representor? That would print the "refcount_t: addition on 0;
use-after-free" warning and leave mark_flows saturated.

[ ... ]

> @@ -1195,12 +1195,12 @@ static int otx2_tc_del_flow(struct otx2_nic *nic,
>  	/* Disable TC MARK flag if they are no rules with skbedit mark action */
>  	if (flow_node->req.match_id)
>  		if (!refcount_dec_and_test(&flow_cfg->mark_flows))
> -			nic->flags &= ~OTX2_FLAG_TC_MARK_ENABLED;
> +			otx2_clear_flag(nic, OTX2_FLAG_TC_MARK_ENABLED);

[Severity: Medium]
This isn't a bug introduced by this patch, but is this condition inverted?

mark_flows starts at 1 in otx2_mcam_entry_init(), and
otx2_tc_parse_actions() increments it for each mark rule, so N mark rules
give a count of N + 1. Deleting any one of them leaves the count at
N >= 1. refcount_dec_and_test() then returns false and TC_MARK_ENABLED is
cleared.

After that, wouldn't the remaining mark rules stop setting skb->mark in
otx2_rcv_pkt_handler()?

This logic appears to come from commit 6a57f091622a12.

[ ... ]

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/rep.c b/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
> index 0f5d5642d3f7a..7df82c22cc12b 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/rep.c
> @@ -93,9 +93,9 @@ static int rvu_rep_mcam_flow_init(struct rep_dev *rep)
>  	rep->flow_cfg->max_flows = allocated;
>  
>  	if (allocated) {
> -		rep->flags |= OTX2_FLAG_MCAM_ENTRIES_ALLOC;
> -		rep->flags |= OTX2_FLAG_NTUPLE_SUPPORT;
> -		rep->flags |= OTX2_FLAG_TC_FLOWER_SUPPORT;
> +		set_bit(OTX2_FLAG_MCAM_ENTRIES_ALLOC, &rep->flags);
> +		set_bit(OTX2_FLAG_NTUPLE_SUPPORT, &rep->flags);
> +		set_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags);
>  	}
>  
>  	INIT_LIST_HEAD(&rep->flow_cfg->flow_list);

[Severity: High]
This isn't a bug introduced by this patch, but rvu_rep_mcam_flow_init()
builds and sends NPC_MCAM_ALLOC_ENTRY messages without taking
priv->mbox.lock. It then unlocks the mutex anyway:

    while (allocated < count) {
        req = otx2_mbox_alloc_msg_npc_mcam_alloc_entry(&priv->mbox);
    ...
exit:
    ...
    mutex_unlock(&priv->mbox.lock);

The caller, rvu_rep_setup_tc_cb(), doesn't hold the lock either. Other
users of the same mailbox do take it, for example rvu_rep_get_stats() in a
kworker and rvu_eswitch_config() from devlink.

Can this interleave messages in the shared AF mailbox? Can the unbalanced
mutex_unlock() release a lock that another task holds?

If otx2_mbox_get_rsp() then matches the wrong response, rsp->entry_list[]
could be copied into rep->flow_cfg->flow_ent[]. Later flow installs would
then target MCAM entries this pcifunc does not own.

Even without any concurrency, the first tc flower command on a
representor would trigger a bad unlock balance warning with lockdep or
DEBUG_MUTEXES. This code appears to come from commit 6c40ca957fe5.

[Severity: Medium]
This is also a pre-existing issue. rvu_rep_mcam_flow_init() always
allocates a new rep->flow_cfg and flow_cfg->flow_ent without freeing any
earlier allocation:

    rep->flow_cfg = kzalloc_objs(struct otx2_flow_config, 1);

rvu_rep_setup_tc_cb() calls it on every callback while TC_FLOWER_SUPPORT
is clear. The flag stays clear if the AF grants zero MCAM entries, the
mailbox fails, or the flow_ent allocation fails. Does each REPLACE,
DESTROY or STATS callback then leak the previous flow_cfg and flow_ent?

The teardown path has related leaks. rvu_rep_destroy() frees only
rep->flow_cfg, not flow_cfg->flow_ent. otx2_destroy_tc_flow_list() is
only called from the PF and VF drivers, so the otx2_tc_flow nodes on a
representor's flow_list_tc also look leaked.

> @@ -109,14 +109,14 @@ static int rvu_rep_setup_tc_cb(enum tc_setup_type type,
>  	struct rep_dev *rep = cb_priv;
>  	struct otx2_nic *priv = rep->mdev;
>  
> -	if (!(rep->flags & RVU_REP_VF_INITIALIZED))
> +	if (!test_bit(OTX2_FLAG_REP_VF_INITIALIZED, &rep->flags))
>  		return -EINVAL;
>  
> -	if (!(rep->flags & OTX2_FLAG_TC_FLOWER_SUPPORT))
> +	if (!test_bit(OTX2_FLAG_TC_FLOWER_SUPPORT, &rep->flags))
>  		rvu_rep_mcam_flow_init(rep);

[Severity: Medium]
This isn't a bug introduced by this patch, but the return value of
rvu_rep_mcam_flow_init() is ignored here.

If the kzalloc of rep->flow_cfg fails, rep->flow_cfg is NULL. If the
kcalloc of flow_ent fails, the function returns before INIT_LIST_HEAD()
runs on the list heads. In both cases priv->flow_cfg = rep->flow_cfg is
still assigned and the command is still dispatched.

FLOW_CLS_REPLACE returns -ENOMEM early here because TC_FLOWER_SUPPORT is
clear. With skip_sw, fl_hw_replace_filter() then calls
fl_hw_destroy_filter(). The FLOW_CLS_DESTROY callback retries the init.
If the allocation fails again, can
otx2_tc_del_flow()->otx2_tc_get_entry_by_cookie() walk
flow_cfg->flow_list_tc on a NULL or uninitialized flow_cfg?

>  
>  	priv->netdev = rep->netdev;
> -	priv->flags = rep->flags;
> +	otx2_sync_flags_from_rep(priv, &rep->flags);

[Severity: Medium]
Before this patch, this assignment replaced every flag on the
representor PF's shared otx2_nic with rep->flags. rep->flags only held
RVU_REP_VF_INITIALIZED plus MCAM_ENTRIES_ALLOC, NTUPLE_SUPPORT and
TC_FLOWER_SUPPORT. RVU_REP_VF_INITIALIZED was BIT_ULL(0), the same bit as
OTX2_FLAG_RX_TSTAMP_ENABLED.

As a result, the first tc flower callback on a representor did three
things:

- It cleared the OTX2_FLAG_REP_MODE_ENABLED set by rvu_rep_probe().
  otx2_tx_napi_handler() then used pfvf->netdev instead of
  pfvf->reps[qidx]->netdev and stopped forcing qidx = 0.
  otx2_rcv_pkt_handler() started recording RX queues and setting
  CHECKSUM_UNNECESSARY on rep traffic.

- It set RX_TSTAMP_ENABLED, so otx2_set_rxtstamp() would dereference
  pfvf->ptp, which is never set up for the rep device.

- It cleared TC_MARK_ENABLED and INTF_DOWN.

The masked copy plus moving the rep bit to bit 21 fixes this. The commit
message only hints at it, with "without clearing PF-owned bits" and
"Relocate representor VF initialization to OTX2_FLAG_REP_VF_INITIALIZED
(bit 21)".

Could this be split out as a separate minimal fix with:

Fixes: 6c40ca957fe5 ("octeontx2-pf: Adds TC offload support")

so that stable trees pick it up, rather than being part of a 14-file
conversion?

>  	priv->pcifunc = rep->pcifunc;
>  	priv->flow_cfg = rep->flow_cfg;
>  

[ ... ]

> @@ -382,7 +382,7 @@ static void rvu_rep_get_stats64(struct net_device *dev,
>  {
>  	struct rep_dev *rep = netdev_priv(dev);
>  
> -	if (!(rep->flags & RVU_REP_VF_INITIALIZED))
> +	if (!test_bit(OTX2_FLAG_REP_VF_INITIALIZED, &rep->flags))
>  		return;
>  
>  	stats->rx_packets = rep->stats.rx_frames;

[Severity: High]
This is a pre-existing issue, but each stats read with VF_INITIALIZED set
ends with:

    schedule_delayed_work(&rep->stats_wrk, msecs_to_jiffies(100));

rep_dev is embedded in the netdev private area. Neither rvu_rep_destroy()
nor the error unwind in rvu_rep_create() calls
cancel_delayed_work_sync(&rep->stats_wrk) before free_netdev(rep->netdev).
rvu_rep_rsrc_free() only cancels the pool refill work.

Suppose a stats read happens within 100ms of a switch to legacy eswitch
mode or a PCI remove. The read can even come from the notification that
unregister_netdev() itself sends. Can rvu_rep_get_stats() then run on the
freed rep_dev and write rep->stats into freed memory?

Separately, rvu_rep_create() calls
INIT_DELAYED_WORK(&rep->stats_wrk, rvu_rep_get_stats) only after
register_netdev(). Would a stats read in that window schedule an
uninitialized work item?

[ ... ]

> @@ -632,7 +632,7 @@ void rvu_rep_destroy(struct otx2_nic *priv)
>  	int rep_id;
>  
>  	rvu_eswitch_config(priv, false);
> -	priv->flags |= OTX2_FLAG_INTF_DOWN;
> +	otx2_set_flag(priv, OTX2_FLAG_INTF_DOWN);
>  	rvu_rep_free_cq_rsrc(priv);
>  	for (rep_id = 0; rep_id < priv->rep_cnt; rep_id++) {
>  		rep = priv->reps[rep_id];

[Severity: High]
This isn't a bug introduced by this patch, but further down in this loop:

        free_netdev(rep->netdev);
        kfree(rep->flow_cfg);

rep is netdev_priv(rep->netdev), so isn't rep->flow_cfg read from freed
memory here? If that memory has been reused, the kfree() could free an
arbitrary pointer.

Should flow_cfg be freed before free_netdev()? Also, priv->flow_cfg and
priv->netdev are copied from the rep in rvu_rep_setup_tc_cb(), so they
seem to be left pointing at freed objects after this.

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261006094357.840823-1-rkannoth%40marvell.com

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [PATCH v19 net-next 2/2] octeontx2: add mqprio bandwidth offload for NIX TX schedulers
  2026-10-06  9:43 ` [PATCH v19 net-next 2/2] octeontx2: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
@ 2026-10-10 10:01   ` netdev-bot+sashiko
  0 siblings, 0 replies; 5+ messages in thread
From: netdev-bot+sashiko @ 2026-10-10 10:01 UTC (permalink / raw)
  To: rkannoth
  Cc: bpf, linux-kernel, netdev, andrew+netdev, ast, daniel, davem,
	edumazet, hawk, john.fastabend, kuba, pabeni, sdf, sgoutham

Thank you for your contribution! Sashiko AI review found 12 potential
issue(s) to consider.

Critical: 0 · High: 1 · Medium: 9 (1) · Low: 2 (1)

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [High] There is a use-after-free race on the deferred netdev_tc_work
  during device removal.
- [Medium] Offload state is leaked when qdisc teardown fails.
- [Medium] The snapshot lifetime is not tied to the offload lifetime, so
  TX channels can never be reduced once mqprio offload has been used,…
- [Medium] A failed tc command can leave the interface administratively
  down, and the commit message does not say so.
- [Medium] The TX scheduler topology is not programmed for the shared
  TL4.
- [Medium] The driver works around core mqprio behaviour with
  asynchronous work, which leaves a window where num_tc is 0 while the…
- [Medium] otx2_mqprio_keep_netdev_tc() only checks whether the new root
  qdisc is "mqprio" (strcmp(qdisc->ops->id, "mqprio")).
- [Medium] The driver commits a rejected config on a failed qdisc replace
  when hw-tc-offload is disabled, which is the default.
- [Medium] If rollback fails before the restart, the interface keeps
  transmitting with the rejected replacement's limits.
- [Medium] If ndo_open() fails during a rollback, the installed qdisc's
  offload state is discarded, so the next open runs with no limits.
- [Low] Two comments in otx2_mqprio_restart_netdev() contradict the code
  and the core API.
- [Low] The comment added in otx2_config_sched_shaping() says 65536 is
  "the kernel-side default burst when HTB does not supply an explicit…

Pre-existing issues:
- [Medium] This problem wasn't introduced by this patch, but the patch
  adds new triggers for it.
- [Low] This problem wasn't introduced by this patch, but the patch adds
  a new path into it.

--

octeontx2: add mqprio bandwidth offload for NIX TX schedulers

This adds TC_SETUP_QDISC_MQPRIO offload with bw_rlimit shaping on octeontx2
PF and VF netdevs. Per-queue MDQ CIR/PIR is programmed through the AF
mailbox, and one SMQ is allocated per non-QoS queue under TL4[0]. Every
add, replace and delete bounces the netdev through ndo_stop()/ndo_open().
Rates are cached in software and restored on open, and a replace is staged
in mq_offload_snap snapshots that are committed on TC_ROOT_GRAFT or when
the replaced qdisc is torn down.

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
> index 836601fc10d8e..c95177602bf95 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_common.c

[ ... ]

> @@ -651,7 +784,11 @@ int otx2_txschq_config(struct otx2_nic *pfvf, int lvl, int prio, bool txschq_for
>  						(u64)hw->smq_link_type);
>  		req->num_regs++;
>  		/* MDQ config */
> -		parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
> +		if (pfvf->mqprio.rate_limit)
> +			parent = schq_list[NIX_TXSCH_LVL_TL4][0];
> +		else
> +			parent = schq_list[NIX_TXSCH_LVL_TL4][prio];
> +
>  		req->reg[1] = NIX_AF_MDQX_PARENT(schq);
>  		req->regval[1] = parent << 16;
>  		req->num_regs++;

[ ... ]

> @@ -777,6 +914,9 @@ int otx2_txsch_alloc(struct otx2_nic *pfvf)
>  		req->schq[NIX_TXSCH_LVL_TL4] = chan_cnt;
>  	}
>  
> +	if (pfvf->mqprio.rate_limit)
> +		req->schq[NIX_TXSCH_LVL_SMQ] = pfvf->hw.non_qos_queues;
> +
>  	rc = otx2_sync_mbox_msg(&pfvf->mbox);

[Severity: Medium]
With rate_limit set, otx2_txsch_alloc() asks for hw.non_qos_queues SMQs
but still only one TL4. otx2_txschq_config() then parents every MDQ to
schq_list[NIX_TXSCH_LVL_TL4][0] and writes NIX_AF_MDQX_SCHEDULE with only
the DWRR weight, so PRIO is 0.

Is NIX_AF_TL4X_TOPOLOGY ever programmed for that shared TL4? The HTB path
writes the parent topology for multi-child nodes in
otx2_qos_txschq_set_parent_topology():

    cfg->reg[0] = NIX_AF_TL4X_TOPOLOGY(parent->schq);
    cfg->regval[0] = (u64)parent->prio_anchor << 32;

The AF's nix_reset_tx_schedule() clears only PARENT and SCHEDULE on
allocation. It does not clear TOPOLOGY.

TL4[0] could still hold a non-zero PRIO_ANCHOR/RR_PRIO, for example after
a failed HTB topology reset, which otx2_qos_reset_schq_topology() only
warns about. In that case, could the sibling MDQs be arbitrated in a way
that doesn't give the per-queue min/max rates, while qopt->hw still
reports TC_MQPRIO_HW_OFFLOAD_TCS?

The netdev_warn_once() in otx2_setup_tc_mqprio() ("uses TL4[0] without
explicit topology programming") seems to acknowledge this.

[ ... ]

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
> index 4fe473d9ea0dd..8365826311279 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_ethtool.c
> @@ -287,6 +287,21 @@ static int otx2_set_channels(struct net_device *dev,
>  		return -EINVAL;
>  	}
>  
> +	if (pfvf->mqprio.rate_limit &&
> +	    (channel->tx_count != pfvf->hw.tx_queues ||
> +	     channel->rx_count != pfvf->hw.rx_queues)) {
> +		netdev_info(dev,
> +			    "Not permitted to change channel count while MQ prio is active\n");
> +		return -EINVAL;
> +	}
> +
> +	if ((pfvf->old_mq_snap || pfvf->cur_mq_snap) &&
> +	    channel->tx_count < pfvf->hw.tx_queues) {
> +		netdev_err(dev,
> +			   "Cannot reduce TX queues after mqprio bandwidth offload was configured\n");
> +		return -EINVAL;
> +	}
> +

[Severity: Medium]
Can TX channels ever be reduced again once mqprio offload has been used?

After a successful setup, otx2_mqprio_snap_commit() leaves old_mq_snap
allocated. otx2_mqprio_clear_sw() and otx2_mqprio_down() only drop
rate_limit and the rate cache. Only otx2_shutdown_tc_mqprio() frees the
snapshots, and otx2_mqprio_restore_old() can also return early without
freeing cur_mq_snap.

So after this sequence:

    tc qdisc add dev ethX root mqprio ... hw 1 ...
    tc qdisc del dev ethX root
    ethtool -L ethX tx <fewer>

the ethtool command keeps failing with -EINVAL until the driver is
unbound.

The snapshot is only used when rate_limit is true, and the rate_limit
check above already blocks channel changes in that case. What does this
second check protect?

The commit message says channel changes are rejected "while active",
which doesn't match this behaviour. It also doesn't mention that the
!replacing && otx2_mqprio_keep_netdev_tc() check in otx2_setup_tc_mqprio()
rejects replacing a software (hw 0) mqprio with an offloaded one.

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> index 777e7156badb7..f51161e8b5949 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_pf.c
> @@ -2008,6 +2008,15 @@ int otx2_open(struct net_device *netdev)
>  	if (err)
>  		goto err_free_mem;
>  
> +	/* Fail closed: abort open if cached mqprio shapers cannot be restored. */
> +	err = otx2_mqprio_up(pf);
> +	if (err) {
> +		netdev_err(pf->netdev,
> +			   "mqprio: failed to restore shapers during open: %d\n",
> +			   err);
> +		goto err_free_hw;
> +	}
> +

[Severity: Low]
This is a pre-existing issue, but this patch adds another path into it.
otx2_init_hw_resources() -> otx2_config_nix_queues() allocates
pf->refill_wrk with devm_kcalloc() on every open. The only free is in
otx2_stop():

    cancel_delayed_work_sync(&pf->refill_wrk[wrk].pool_refill_work);
    devm_kfree(pf->dev, pf->refill_wrk);

The new err_free_hw path unwinds through otx2_free_hw_resources(), which
doesn't free it. After a failed open, otx2_stop() returns early because
OTX2_FLAG_INTF_DOWN is set.

Does each otx2_mqprio_up() failure during open then leak one refill_wrk
array until unbind? The existing err_disable_napi and later labels leak
it the same way.

[ ... ]

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> index 8877af348a090..c05a7009a863b 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/otx2_tc.c

[ ... ]

> @@ -1600,14 +1619,907 @@ static int otx2_setup_tc_block(struct net_device *netdev,

[ ... ]

> +static int otx2_mqprio_restore_old(struct otx2_nic *pfvf)
> +{
> +	struct mq_offload_snap *snap = pfvf->old_mq_snap;
> +	struct net_device *netdev = pfvf->netdev;
> +	u16 num_txq = pfvf->hw.non_qos_queues;
> +	int tc, txq, err;
> +
> +	if (!snap)
> +		return 0;
> +
> +	err = otx2_mqprio_alloc_cache(pfvf, false);
> +	if (err) {
> +		netdev_err(netdev,
> +			   "mqprio: rollback failed to allocate rate cache: %d\n",
> +			   err);
> +		return err;
> +	}

[ ... ]

> +	otx2_mqprio_apply_snap_netdev(netdev, snap);
> +
> +	if (otx2_mqprio_mdq_allocated(pfvf)) {
> +		err = otx2_nix_tm_clear_queue_shaper(pfvf);
> +		if (err) {
> +			netdev_err(netdev,
> +				   "mqprio: rollback shaper clear failed: %d; hardware limits may not match qdisc\n",
> +				   err);
> +			return err;
> +		}
> +	}

[Severity: Medium]
If either of these early returns is taken, what state is the running
interface left in?

If otx2_mqprio_alloc_cache() fails, the existing cache and the MDQ
shapers still hold the rejected replacement's rates.

If otx2_nix_tm_clear_queue_shaper() fails, the cache and netdev layout
have already been switched back to old_mq_snap, but the hardware shapers
are only partly cleared.

Neither path restarts the device, reapplies the old rates or frees
cur_mq_snap. The caller in otx2_teardown_tc_mqprio() just returns the
error, and mqprio_disable_offload() ignores it. The cleanup caller in
otx2_setup_tc_mqprio() only logs it.

Can the old qdisc then stay installed while the interface keeps
transmitting with the new or partially cleared limits?

> +
> +	/* Rebuild the TX scheduler via netdev restart when running; otx2_mqprio_up()
> +	 * alone is insufficient after a failed replace that already bounced the
> +	 * interface. If open failed, TX schedulers were freed; defer shaper restore
> +	 * to the next successful ndo_open() via otx2_mqprio_up().
> +	 */
> +	pfvf->mqprio.rate_limit = true;
> +
> +	if (netif_running(netdev)) {
> +		err = otx2_mqprio_restart_netdev(netdev, true);
> +		if (err) {
> +			netdev_err(netdev,
> +				   "mqprio: rollback netdev restart failed: %d; qdisc rates may not be enforced\n",
> +				   err);
> +			return err;
> +		}

[Severity: Medium]
What happens to the restored state if ndo_open() fails inside this
restart? otx2_mqprio_restart_netdev() handles an open failure with:

    otx2_mqprio_clear_sw(pfvf);
    ...
    netif_close(netdev);

That sets rate_limit to false and frees the rate cache that was just
rebuilt from old_mq_snap. restore_old() then returns without rebuilding
it, while the old offloaded qdisc stays installed.

On the next ip link set up, otx2_open() -> otx2_mqprio_up() sees
!rate_limit and returns 0, and otx2_txsch_alloc() asks for a single SMQ.

Doesn't the interface then come up and pass traffic without any of the
qdisc's shapers? That seems to contradict the "failing closed on error"
wording in the commit message.

[ ... ]

> +/* Offloaded mqprio replaced by software mqprio installs netdev TC layout in
> + * mqprio_init() before the old offload instance is destroyed during graft.
> + */
> +static bool otx2_mqprio_keep_netdev_tc(struct otx2_nic *pfvf)
> +{
> +	struct Qdisc *qdisc = rtnl_dereference(pfvf->netdev->qdisc);
> +
> +	return qdisc && qdisc->ops && !strcmp(qdisc->ops->id, "mqprio");
> +}
> +
> +static void otx2_mqprio_clear_sw(struct otx2_nic *pfvf)
> +{
> +	struct net_device *netdev = pfvf->netdev;
> +
> +	pfvf->mqprio.rate_limit = false;
> +	otx2_mqprio_clear_replace_state(pfvf);
> +	if (!otx2_mqprio_keep_netdev_tc(pfvf))
> +		netdev_set_num_tc(netdev, 0);
> +	otx2_mqprio_free_cache(pfvf);
> +}

[Severity: Medium]
Is mqprio the only root qdisc that installs a netdev TC layout before the
old root is destroyed? taprio_change(), called from taprio_init(), also
does netdev_set_num_tc(), netdev_set_tc_queue() and
netdev_set_prio_tc_map() for a software taprio.

Take this command run over an offloaded mqprio:

    tc qdisc replace dev ethX root taprio ...   (software mode)

qdisc_graft() publishes dev->qdisc as the taprio and then destroys the old
mqprio:

mqprio_destroy()
  mqprio_disable_offload()
    otx2_setup_tc_mqprio()        /* qopt->hw == 0 */
      otx2_teardown_tc_mqprio()
        otx2_mqprio_down()
          otx2_mqprio_clear_sw()
            netdev_set_num_tc(netdev, 0)

Since ops->id is "taprio" here, wouldn't this wipe the layout the live
taprio just installed and break its TC based classification?

> +
> +/* Tear down mqprio bandwidth offload: clear per-queue shapers,
> + * mqprio_rate_limit, netdev TC mappings, and the cached rates.  Called on
> + * explicit mqprio teardown (tc qdisc del) and error cleanup, not on
> + * routine netdev stop/open cycles where the offload stays active.
> + */
> +int otx2_mqprio_down(struct otx2_nic *pfvf)
> +{
> +	int err = 0;
> +
> +	if (!pfvf->mqprio.rate_limit)
> +		return 0;
> +
> +	if (netif_running(pfvf->netdev) &&
> +	    otx2_mqprio_mdq_allocated(pfvf))
> +		err = otx2_nix_tm_clear_queue_shaper(pfvf);
> +
> +	if (err) {
> +		netdev_warn(pfvf->netdev,
> +			    "mqprio: failed to clear hardware shapers: %d; keeping offload state\n",
> +			    err);
> +		return err;
> +	}
> +
> +	otx2_mqprio_clear_sw(pfvf);
> +
> +	return 0;
> +}

[Severity: Medium]
Can the driver keep this offload state with no qdisc owning it? On
tc qdisc del the path is:

mqprio_destroy()
  mqprio_disable_offload()
    otx2_setup_tc_mqprio()        /* qopt->hw == 0 */
      otx2_teardown_tc_mqprio()
        otx2_mqprio_down()

mqprio_disable_offload() is void and ignores the return value. The qdisc
is destroyed even if otx2_nix_tm_clear_queue_shaper() gets a mailbox
error here.

That leaves rate_limit set, the rate cache populated, the SMQ allocation
widened and the MDQ shapers programmed. otx2_open() -> otx2_mqprio_up()
then reapplies the shapers on every open, and XDP, HTB, PFC and channel
changes stay rejected.

The next tc qdisc add ... mqprio hw 1 computes replacing = rate_limit =
true and sets replace_setup_done. The old root is now the default mq,
whose destroy only sends TC_SETUP_QDISC_MQ, so cur_mq_snap is never
committed.

When that new mqprio is deleted, otx2_teardown_tc_mqprio() takes the
replace_setup_done && cur_mq_snap branch. It commits and returns 0
without clearing shapers or restarting.

The non-replacing cleanup path in otx2_setup_tc_mqprio() also ignores the
result of the teardown:

    otx2_mqprio_snap_free(pfvf, &pfvf->cur_mq_snap);
    otx2_teardown_tc_mqprio(pfvf, mqprio);
    return err;

The core can't veto this teardown, so should the software state be
released unconditionally?

[ ... ]

> +static int otx2_mqprio_restart_netdev(struct net_device *netdev, bool rate_limit)
> +{
> +	struct otx2_nic *pfvf = netdev_priv(netdev);
> +	const struct net_device_ops *ops = netdev->netdev_ops;
> +	bool running = netif_running(netdev);
> +	int err;

[ ... ]

> +	if (running) {
> +		clear_bit(__LINK_STATE_START, &netdev->state);
> +		smp_mb__after_atomic(); /* Commit netif_running(). */
> +	}

[ ... ]

> +	err = ops->ndo_open(netdev);
> +	if (!err && running) {
> +		set_bit(__LINK_STATE_START, &netdev->state);
> +	} else if (err) {
> +		netdev_err(netdev,
> +			   "Failed to restart device after mqprio change: %d\n",
> +			   err);
> +		/* ndo_open() already tore down TX scheduler resources on failure;
> +		 * netif_running() is false here (__LINK_STATE_START stays clear
> +		 * until open succeeds). Drop mqprio software state only instead
> +		 * of sending shaper clears to freed queues.
> +		 */
> +		otx2_mqprio_clear_sw(pfvf);
> +		/* ndo_open() rolls back on failure; mark the interface down so
> +		 * otx2_stop() returns early when netif_close() runs.  Caller
> +		 * holds RTNL; dev_close() would deadlock.
> +		 */
> +		otx2_set_flag(pfvf, OTX2_FLAG_INTF_DOWN);
> +		/* visible to otx2_stop() on other cpus */
> +		smp_wmb();
> +		netif_close(netdev);
> +	}

[Severity: Medium]
Is it intended that a failed tc command can take the interface
administratively down? netif_close() clears IFF_UP and sends
NETDEV_GOING_DOWN and NETDEV_DOWN. A failed tc qdisc add, replace, del or
rollback therefore leaves the link down until an admin brings it back.

The open can fail for reasons unrelated to the request. For example, the
AF may refuse the hw.non_qos_queues SMQs that otx2_txsch_alloc() now
requests, or otx2_mqprio_up() may fail.

This function also sets and clears the core-owned __LINK_STATE_START bit
itself from inside ndo_setup_tc.

The commit message describes "failing closed" only for shaper restore in
ndo_open(). Could it also mention that tc setup, replace and delete can
close the device?

[Severity: Low]
Is the dev_close() part of this comment accurate? In this tree,
dev_close() is:

net/core/dev_api.c:dev_close() {
	netdev_lock_ops(dev);
	netif_close(dev);
	netdev_unlock_ops(dev);
}

It doesn't take RTNL, so there is no RTNL deadlock to avoid.

The block comment above this function also says "Do not call
dev_deactivate()/dev_activate() here". However, netif_close() ->
__dev_close_many() calls dev_deactivate_many() and runs the down
notifiers from inside ndo_setup_tc. Could these comments be brought in
line with the code?

[ ... ]

> +static int otx2_teardown_tc_mqprio(struct otx2_nic *pfvf,
> +				   struct tc_mqprio_qopt_offload *mqprio)
> +{

[ ... ]

> +	if (pfvf->mqprio.replace_setup_done && pfvf->cur_mq_snap) {
> +		err = 0;
> +		if (pfvf->mqprio.replace_graft_done || !tc_can_offload(netdev)) {
> +			otx2_mqprio_snap_commit(pfvf);
> +		} else {
> +			err = otx2_mqprio_restore_old(pfvf);

[Severity: Medium]
Can this commit a configuration that the core rejected?

otx2_probe() and otx2vf_probe() add NETIF_F_HW_TC to hw_features only
after features |= hw_features. By default, tc_can_offload() is false and
TC_ROOT_GRAFT is never delivered. otx2_tc_can_offload() checks
hw_features, so mqprio setup still runs.

In that state, this branch can't tell two cases apart: teardown of the
old instance after a graft, and destruction of a new instance that failed
after init. For example:

    tc qdisc replace dev ethX root handle 2: mqprio hw 1 ... \
        shaper bw_rlimit ... estimator 1sec 8sec

otx2_setup_tc_mqprio() succeeds, programs the new MDQ shapers and sets
replace_setup_done. qdisc_create() then rejects TCA_RATE for a
TCQ_F_MQROOT qdisc:

    if (tca[TCA_RATE]) {
        err = -EOPNOTSUPP;
        if (sch->flags & TCQ_F_MQROOT) {
            ...
            goto err_out4;

err_out4 -> mqprio_destroy(new) then lands here and commits.

Doesn't the old qdisc then stay root while the hardware shapers, rate
cache, old_mq_snap and netdev TC mapping all hold the rejected config?
otx2_mqprio_up() would also reapply it on every open.

[ ... ]

> +static int otx2_setup_tc_mqprio(struct net_device *netdev,
> +				struct tc_mqprio_qopt_offload *mqprio)
> +{

[ ... ]

> +	err = otx2_mqprio_stage_cur(pfvf, mqprio);
> +	if (err)
> +		goto fail_validate;
> +
> +	err = otx2_mqprio_restart_netdev(pfvf->netdev, true);
> +	if (err)
> +		goto cleanup;

[Severity: Medium]
This isn't a bug introduced by this patch, but the new restart adds more
ways to trigger it. An egress matchall police filter programs
NIX_AF_TL4X_PIR on txschq_list[NIX_TXSCH_LVL_TL4][0] in
otx2_set_matchall_egress_rate() and sets
OTX2_FLAG_TC_MATCHALL_EGRESS_ENABLED. The rate and burst are not cached
for replay.

Every mqprio add, replace, delete or rollback bounces the netdev here.
ndo_stop() frees the TL4, and the AF clears its shaping through
nix_reset_tx_shaping() when it is reallocated. Neither otx2_open() nor
otx2_mqprio_up() replays the matchall rate, and otx2_setup_tc_mqprio()
doesn't reject the combination.

Does the filter then stay marked as offloaded with no hardware limit
behind it? The same loss already happens on ip link down/up, MTU changes
and ethtool -L.

[ ... ]

> +fail_validate:
> +	/* Failed replace destroys the new qdisc with hw_offload unset, so
> +	 * mqprio_destroy() clears netdev TC after we return. Re-apply the
> +	 * prior layout when validation fails before any hardware change.
> +	 */
> +	if (replacing)
> +		otx2_mqprio_defer_netdev_tc_restore(pfvf);
> +	return err;

[Severity: Medium]
Is there a window here where the old offload is still active but
dev->num_tc is 0? When a replace fails here, qdisc_create() unwinds via
err_out4 -> mqprio_destroy() on the new instance. With hw_offload unset,
that runs:

    if (priv->hw_offload && dev->netdev_ops->ndo_setup_tc)
        mqprio_disable_offload(sch);
    else
        netdev_set_num_tc(dev, 0);

The restore is queued to system_wq and can only run after the caller
drops RTNL. Until then, rate_limit and the old shapers are still in
place, but skb_tx_hash() ignores the prio to TC mapping. Traffic can then
go out on queues shaped for other TCs.

Other RTNL holders that get the lock first (another tc command, a dump, a
teardown) also see num_tc == 0. otx2_mqprio_netdev_tc_work() later
reapplies old_mq_snap without checking whether the layout changed in the
meantime. The successful-rollback branch in the cleanup path uses the
same deferral.

Could the commit message explain this approach and the window it leaves?

[ ... ]

> @@ -1625,10 +2537,25 @@ int otx2_init_tc(struct otx2_nic *nic)

[ ... ]

> +void otx2_shutdown_tc_mqprio(struct otx2_nic *nic)
> +{
> +	WRITE_ONCE(nic->mqprio.defer_tc_work, true);
> +	/* Publish defer_tc_work before cancel_work_sync(). */
> +	smp_wmb();
> +	cancel_work_sync(&nic->mqprio.netdev_tc_work);
> +	rtnl_lock();
> +	otx2_mqprio_snap_destroy(nic);
> +	rtnl_unlock();
> +}
> +EXPORT_SYMBOL(otx2_shutdown_tc_mqprio);

[Severity: High]
Can the work be queued again after this cancel_work_sync()? The flag is
written and the work cancelled without holding RTNL.
otx2_mqprio_defer_netdev_tc_restore() runs under RTNL and does a separate
check followed by schedule_work():

CPU A (failing tc qdisc replace, holds RTNL)
otx2_mqprio_defer_netdev_tc_restore()
    READ_ONCE(pfvf->mqprio.defer_tc_work) == false

CPU B (otx2_remove() or otx2vf_remove())
otx2_shutdown_tc_mqprio()
    WRITE_ONCE(nic->mqprio.defer_tc_work, true)
    cancel_work_sync()     /* nothing queued yet */
    rtnl_lock()            /* blocks */

CPU A
    schedule_work(&pfvf->mqprio.netdev_tc_work)
    /* returns, RTNL dropped */

CPU B then destroys the snapshots, calls unregister_netdev() and later
free_netdev(). It never calls cancel_work_sync() or flush_work() again.

If system_wq hasn't run the item by then, isn't the work_struct in
netdev_priv freed while still queued? otx2_mqprio_netdev_tc_work() would
then read defer_tc_work from freed memory.

Setting the flag under rtnl_lock(), or cancelling the work again after
unregister_netdev(), would close this.

[ ... ]

> diff --git a/drivers/net/ethernet/marvell/octeontx2/nic/qos.c b/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
> index f160b1618efa2..9ef55a6db50b7 100644
> --- a/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
> +++ b/drivers/net/ethernet/marvell/octeontx2/nic/qos.c
> @@ -118,6 +118,9 @@ static void otx2_config_sched_shaping(struct otx2_nic *pfvf,
>  	/* configure PIR */
>  	maxrate = (node->rate > node->ceil) ? node->rate : node->ceil;
>  
> +	/* 65536 is the kernel-side default burst when HTB does not supply an
> +	 * explicit value, not the NIX hardware maximum (CN10K_MAX_BURST_SIZE).
> +	 */
>  	cfg->regval[*num_regs] =
>  		otx2_get_txschq_rate_regval(pfvf, maxrate, 65536);
>  	(*num_regs)++;

[Severity: Low]
Is this comment accurate? struct tc_htb_qopt_offload carries quantum,
rate, ceil and prio but no burst. HTB offload can never supply a burst,
so 65536 is a driver constant passed every time rather than a fallback.

The new comment in otx2_nix_tm_set_queue_shaper() describes it the same
way ("the same 65536 byte default as the HTB shaper path").

[ ... ]

-- 
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261006094357.840823-1-rkannoth%40marvell.com

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-10 10:01 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-06  9:43 [PATCH v19 net-next 0/2] octeontx2: mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
2026-10-06  9:43 ` [PATCH v19 net-next 1/2] octeontx2: use atomic bitops for PF/VF and rep flags Ratheesh Kannoth
2026-10-10 10:01   ` netdev-bot+sashiko
2026-10-06  9:43 ` [PATCH v19 net-next 2/2] octeontx2: add mqprio bandwidth offload for NIX TX schedulers Ratheesh Kannoth
2026-10-10 10:01   ` netdev-bot+sashiko

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox