* [PATCH net-next v4 1/2] net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path
2026-10-01 9:26 [PATCH net-next v4 0/2] net: stmmac: introduce XDP multi-buff support for TX path Lorenzo Bianconi
@ 2026-10-01 9:26 ` Lorenzo Bianconi
2026-10-05 9:47 ` netdev-bot+sashiko
2026-10-01 9:26 ` [PATCH net-next v4 2/2] net: stmmac: add XDP multi-buff support for TX side Lorenzo Bianconi
2026-10-07 23:40 ` [PATCH net-next v4 0/2] net: stmmac: introduce XDP multi-buff support for TX path patchwork-bot+netdevbpf
2 siblings, 1 reply; 6+ messages in thread
From: Lorenzo Bianconi @ 2026-10-01 9:26 UTC (permalink / raw)
To: Maxime Chevallier, Andrew Lunn, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Maxime Coquelin, Alexandre Torgue,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Ong Boon Leong
Cc: netdev, linux-stm32, linux-arm-kernel, bpf, Lorenzo Bianconi
tx_coal_frames[] is written by ethtool (stmmac_set_coalesce) and read
locklessly from the TX fast path (stmmac_xmit, stmmac_tso_xmit,
stmmac_xdp_xmit_zc and stmmac_xdp_xmit_xdpf), so pair every access
with READ_ONCE()/WRITE_ONCE() to mark the intentional data race and
read a consistent value into a local variable.
Moreover, ethtool allows disabling TX coalescing by setting
tx_max_coalesced_frames to 0: guard the modulo operation in
stmmac_xdp_xmit_xdpf() against a zero tx_coal_frames[queue] to avoid
a division by zero.
Both issues were found by code inspection while extending the XDP TX
path with multi-buff support, and neither has ever been triggered, so
they are only theoretical. Tested on the Qualcomm RB3-gen2 board.
Fixes: db2f2842e6f5 ("net: stmmac: add per-queue TX & RX coalesce ethtool support")
Fixes: be8b38a722e6 ("net: stmmac: Add support for XDP_TX action")
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
---
.../net/ethernet/stmicro/stmmac/stmmac_ethtool.c | 11 +++---
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 42 ++++++++++------------
2 files changed, 24 insertions(+), 29 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
index 1cf0f8820b33..30e0d499326d 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
@@ -798,7 +798,8 @@ static int __stmmac_get_coalesce(struct net_device *dev,
if (queue < tx_cnt) {
ec->tx_coalesce_usecs = priv->tx_coal_timer[queue];
- ec->tx_max_coalesced_frames = priv->tx_coal_frames[queue];
+ ec->tx_max_coalesced_frames =
+ READ_ONCE(priv->tx_coal_frames[queue]);
} else {
ec->tx_coalesce_usecs = 0;
ec->tx_max_coalesced_frames = 0;
@@ -887,14 +888,14 @@ static int __stmmac_set_coalesce(struct net_device *dev,
int i;
for (i = 0; i < tx_cnt; i++) {
- priv->tx_coal_frames[i] =
- ec->tx_max_coalesced_frames;
+ WRITE_ONCE(priv->tx_coal_frames[i],
+ ec->tx_max_coalesced_frames);
priv->tx_coal_timer[i] =
ec->tx_coalesce_usecs;
}
} else if (queue < tx_cnt) {
- priv->tx_coal_frames[queue] =
- ec->tx_max_coalesced_frames;
+ WRITE_ONCE(priv->tx_coal_frames[queue],
+ ec->tx_max_coalesced_frames);
priv->tx_coal_timer[queue] =
ec->tx_coalesce_usecs;
}
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index d95482b7f6e5..df6a329fdb80 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -2731,7 +2731,8 @@ static bool stmmac_xdp_xmit_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
struct stmmac_metadata_request meta_req;
struct xsk_tx_metadata *meta = NULL;
dma_addr_t dma_addr;
- bool set_ic;
+ bool set_ic = false;
+ u32 tx_coal;
/* We are sharing with slow path and stop XSK TX desc submission when
* available TX ring is less than threshold.
@@ -2772,12 +2773,9 @@ static bool stmmac_xdp_xmit_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
tx_q->tx_count_frames++;
- if (!priv->tx_coal_frames[queue])
- set_ic = false;
- else if (tx_q->tx_count_frames % priv->tx_coal_frames[queue] == 0)
+ tx_coal = READ_ONCE(priv->tx_coal_frames[queue]);
+ if (tx_coal && !(tx_q->tx_count_frames % tx_coal))
set_ic = true;
- else
- set_ic = false;
meta_req.priv = priv;
meta_req.tx_desc = tx_desc;
@@ -3419,7 +3417,7 @@ static void stmmac_init_coalesce(struct stmmac_priv *priv)
for (chan = 0; chan < tx_channel_count; chan++) {
struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[chan];
- priv->tx_coal_frames[chan] = STMMAC_TX_FRAMES;
+ WRITE_ONCE(priv->tx_coal_frames[chan], STMMAC_TX_FRAMES);
priv->tx_coal_timer[chan] = STMMAC_COAL_TX_TIMER;
hrtimer_setup(&tx_q->txtimer, stmmac_tx_timer, CLOCK_MONOTONIC, HRTIMER_MODE_REL);
@@ -4557,10 +4555,10 @@ static netdev_tx_t stmmac_tso_xmit(struct sk_buff *skb, struct net_device *dev)
struct dma_desc *desc, *first, *mss_desc = NULL;
struct stmmac_priv *priv = netdev_priv(dev);
struct stmmac_txq_stats *txq_stats;
+ u32 tx_coal, pay_len, mss, queue;
int i, first_tx, nfrags, ndesc;
struct stmmac_tx_queue *tx_q;
bool set_ic, is_last_segment;
- u32 pay_len, mss, queue;
dma_addr_t des;
u8 hdr;
@@ -4679,16 +4677,16 @@ static netdev_tx_t stmmac_tso_xmit(struct sk_buff *skb, struct net_device *dev)
priv->dma_conf.dma_tx_size);
tx_q->tx_count_frames += tx_packets;
+ tx_coal = READ_ONCE(priv->tx_coal_frames[queue]);
if ((skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP) && priv->hwts_tx_en)
set_ic = true;
- else if (!priv->tx_coal_frames[queue])
+ else if (!tx_coal)
set_ic = false;
else if (!netdev_xmit_more())
set_ic = true;
- else if (tx_packets > priv->tx_coal_frames[queue])
+ else if (tx_packets > tx_coal)
set_ic = true;
- else if ((tx_q->tx_count_frames %
- priv->tx_coal_frames[queue]) < tx_packets)
+ else if ((tx_q->tx_count_frames % tx_coal) < tx_packets)
set_ic = true;
else
set_ic = false;
@@ -4843,9 +4841,9 @@ static netdev_tx_t stmmac_xmit(struct sk_buff *skb, struct net_device *dev)
struct dma_desc *desc, *first_desc;
struct stmmac_tx_queue *tx_q;
int i, csum_insertion = 0;
+ u32 tx_coal, sdu_len;
int entry, first_tx;
dma_addr_t dma_addr;
- u32 sdu_len;
if (priv->tx_path_in_lpi_mode && priv->eee_sw_timer_en)
stmmac_stop_sw_lpi(priv);
@@ -4986,16 +4984,16 @@ static netdev_tx_t stmmac_xmit(struct sk_buff *skb, struct net_device *dev)
tx_packets = CIRC_CNT(entry + 1, first_tx, priv->dma_conf.dma_tx_size);
tx_q->tx_count_frames += tx_packets;
+ tx_coal = READ_ONCE(priv->tx_coal_frames[queue]);
if ((skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP) && priv->hwts_tx_en)
set_ic = true;
- else if (!priv->tx_coal_frames[queue])
+ else if (!tx_coal)
set_ic = false;
else if (!netdev_xmit_more())
set_ic = true;
- else if (tx_packets > priv->tx_coal_frames[queue])
+ else if (tx_packets > tx_coal)
set_ic = true;
- else if ((tx_q->tx_count_frames %
- priv->tx_coal_frames[queue]) < tx_packets)
+ else if ((tx_q->tx_count_frames % tx_coal) < tx_packets)
set_ic = true;
else
set_ic = false;
@@ -5258,7 +5256,7 @@ static int stmmac_xdp_xmit_xdpf(struct stmmac_priv *priv, int queue,
enum stmmac_txbuf_type buf_type;
struct dma_desc *tx_desc;
dma_addr_t dma_addr;
- bool set_ic;
+ u32 tx_coal;
if (stmmac_tx_avail(priv, queue) < STMMAC_TX_THRESH(priv))
return STMMAC_XDP_CONSUMED;
@@ -5303,12 +5301,8 @@ static int stmmac_xdp_xmit_xdpf(struct stmmac_priv *priv, int queue,
tx_q->tx_count_frames++;
- if (tx_q->tx_count_frames % priv->tx_coal_frames[queue] == 0)
- set_ic = true;
- else
- set_ic = false;
-
- if (set_ic) {
+ tx_coal = READ_ONCE(priv->tx_coal_frames[queue]);
+ if (tx_coal && !(tx_q->tx_count_frames % tx_coal)) {
tx_q->tx_count_frames = 0;
stmmac_set_tx_ic(priv, tx_desc);
u64_stats_update_begin(&txq_stats->q_syncp);
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* Re: [PATCH net-next v4 1/2] net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path
2026-10-01 9:26 ` [PATCH net-next v4 1/2] net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path Lorenzo Bianconi
@ 2026-10-05 9:47 ` netdev-bot+sashiko
2026-10-05 10:12 ` Lorenzo Bianconi
0 siblings, 1 reply; 6+ messages in thread
From: netdev-bot+sashiko @ 2026-10-05 9:47 UTC (permalink / raw)
To: lorenzo.bianconi
Cc: maxime.chevallier, andrew+netdev, davem, edumazet, kuba, pabeni,
mcoquelin.stm32, alexandre.torgue, ast, daniel, hawk,
john.fastabend, sdf, boon.leong.ong, netdev, linux-stm32,
linux-arm-kernel, bpf
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0 (1)
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] The commit message says the stmmac_xdp_xmit_xdpf() divide-by-
zero is "only theoretical", but a privileged user can trigger it and…
Pre-existing issues:
- [Low] __stmmac_set_coalesce() publishes tx_coal_frames[q] and
tx_coal_timer[q] as two separate stores and shares no synchronisation…
--
net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path
This patch wraps every access to tx_coal_frames[] in READ_ONCE() or
WRITE_ONCE(), since ethtool writes the array while the TX fast paths read
it without a lock. It also guards the modulo in stmmac_xdp_xmit_xdpf()
against a zero tx_coal_frames[queue].
> Both issues were found by code inspection while extending the XDP TX
> path with multi-buff support, and neither has ever been triggered, so
> they are only theoretical. Tested on the Qualcomm RB3-gen2 board.
[Severity: Medium]
Is the division by zero in stmmac_xdp_xmit_xdpf() really only
theoretical?
__stmmac_set_coalesce() only rejects the case where both values are zero:
if ((ec->tx_coalesce_usecs == 0) &&
(ec->tx_max_coalesced_frames == 0))
return -EINVAL;
So "ethtool -C ethX tx-frames 0" is accepted. tx-usecs keeps its default
of STMMAC_COAL_TX_TIMER, and stmmac_ethtool_ops advertises
ETHTOOL_COALESCE_MAX_FRAMES. After that, tx_coal_frames[queue] is 0.
>From then on, before this patch, every XDP_TX verdict
(stmmac_xdp_xmit_back()->stmmac_xdp_xmit_xdpf()) and every ndo_xdp_xmit
redirect into the device would evaluate:
if (tx_q->tx_count_frames % priv->tx_coal_frames[queue] == 0)
On x86, for example the dwmac-intel platforms that advertise XDP
basic and redirect, this raises a divide error in NAPI softirq context.
On arm64, UDIV with a zero divisor returns 0 instead of trapping. That
may be why the RB3-gen2 board never hit it.
CAP_NET_ADMIN can trigger this with standard tools. Could the commit
message describe it as a reachable crash rather than a theoretical one?
Would it also make sense to send the div-by-zero fix separately from the
READ_ONCE()/WRITE_ONCE() annotations?
> diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
> index 1cf0f8820b337..30e0d499326d8 100644
> --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
> +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
[ ... ]
> @@ -887,14 +888,14 @@ static int __stmmac_set_coalesce(struct net_device *dev,
> int i;
>
> for (i = 0; i < tx_cnt; i++) {
> - priv->tx_coal_frames[i] =
> - ec->tx_max_coalesced_frames;
> + WRITE_ONCE(priv->tx_coal_frames[i],
> + ec->tx_max_coalesced_frames);
> priv->tx_coal_timer[i] =
> ec->tx_coalesce_usecs;
> }
> } else if (queue < tx_cnt) {
> - priv->tx_coal_frames[queue] =
> - ec->tx_max_coalesced_frames;
> + WRITE_ONCE(priv->tx_coal_frames[queue],
> + ec->tx_max_coalesced_frames);
> priv->tx_coal_timer[queue] =
> ec->tx_coalesce_usecs;
> }
[Severity: Low]
This is a pre-existing issue and was not introduced by this patch.
tx_coal_frames[] and tx_coal_timer[] are still published as two separate
stores, with no synchronisation against the TX path.
Can a concurrent stmmac_xmit() see a mixed pair? Take a change from
(frames=N, usecs=0) to (frames=0, usecs=U). The TX path could read the
new frames value:
tx_coal = READ_ONCE(priv->tx_coal_frames[queue]);
...
else if (!tx_coal)
set_ic = false;
It could then read the old timer value in stmmac_tx_timer_arm():
u32 tx_coal_timer = priv->tx_coal_timer[queue];
if (!tx_coal_timer)
return;
That frame would then get neither an IC interrupt nor the completion
timer. Its completion would wait for the next TX or NAPI activity on that
queue. The same window exists for stmmac_tso_xmit() and
stmmac_xdp_xmit_xdpf().
READ_ONCE()/WRITE_ONCE() make each field consistent on its own, which is
all the commit message claims. The window across the two fields is still
there.
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261001-b4-stmmac-xmit-mb-v4-0-664e0c3eb289%40oss.qualcomm.com
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH net-next v4 1/2] net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path
2026-10-05 9:47 ` netdev-bot+sashiko
@ 2026-10-05 10:12 ` Lorenzo Bianconi
0 siblings, 0 replies; 6+ messages in thread
From: Lorenzo Bianconi @ 2026-10-05 10:12 UTC (permalink / raw)
To: netdev-bot+sashiko
Cc: maxime.chevallier, andrew+netdev, davem, edumazet, kuba, pabeni,
mcoquelin.stm32, alexandre.torgue, ast, daniel, hawk,
john.fastabend, sdf, boon.leong.ong, netdev, linux-stm32,
linux-arm-kernel, bpf
[-- Attachment #1: Type: text/plain, Size: 5277 bytes --]
On Oct 05, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 0 (1)
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [Medium] The commit message says the stmmac_xdp_xmit_xdpf() divide-by-
> zero is "only theoretical", but a privileged user can trigger it and…
>
> Pre-existing issues:
> - [Low] __stmmac_set_coalesce() publishes tx_coal_frames[q] and
> tx_coal_timer[q] as two separate stores and shares no synchronisation…
>
> --
>
> net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path
>
> This patch wraps every access to tx_coal_frames[] in READ_ONCE() or
> WRITE_ONCE(), since ethtool writes the array while the TX fast paths read
> it without a lock. It also guards the modulo in stmmac_xdp_xmit_xdpf()
> against a zero tx_coal_frames[queue].
>
> > Both issues were found by code inspection while extending the XDP TX
> > path with multi-buff support, and neither has ever been triggered, so
> > they are only theoretical. Tested on the Qualcomm RB3-gen2 board.
>
> [Severity: Medium]
>
> Is the division by zero in stmmac_xdp_xmit_xdpf() really only
> theoretical?
with 'theoretical' I mean I spotted the issue with code inspection and I have
never observed it in practice. I agree this is a real issue.
>
> __stmmac_set_coalesce() only rejects the case where both values are zero:
>
> if ((ec->tx_coalesce_usecs == 0) &&
> (ec->tx_max_coalesced_frames == 0))
> return -EINVAL;
>
> So "ethtool -C ethX tx-frames 0" is accepted. tx-usecs keeps its default
> of STMMAC_COAL_TX_TIMER, and stmmac_ethtool_ops advertises
> ETHTOOL_COALESCE_MAX_FRAMES. After that, tx_coal_frames[queue] is 0.
>
> >From then on, before this patch, every XDP_TX verdict
> (stmmac_xdp_xmit_back()->stmmac_xdp_xmit_xdpf()) and every ndo_xdp_xmit
> redirect into the device would evaluate:
>
> if (tx_q->tx_count_frames % priv->tx_coal_frames[queue] == 0)
>
> On x86, for example the dwmac-intel platforms that advertise XDP
> basic and redirect, this raises a divide error in NAPI softirq context.
>
> On arm64, UDIV with a zero divisor returns 0 instead of trapping. That
> may be why the RB3-gen2 board never hit it.
>
> CAP_NET_ADMIN can trigger this with standard tools. Could the commit
> message describe it as a reachable crash rather than a theoretical one?
> Would it also make sense to send the div-by-zero fix separately from the
> READ_ONCE()/WRITE_ONCE() annotations?
>
> > diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
> > index 1cf0f8820b337..30e0d499326d8 100644
> > --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
> > +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_ethtool.c
>
> [ ... ]
>
> > @@ -887,14 +888,14 @@ static int __stmmac_set_coalesce(struct net_device *dev,
> > int i;
> >
> > for (i = 0; i < tx_cnt; i++) {
> > - priv->tx_coal_frames[i] =
> > - ec->tx_max_coalesced_frames;
> > + WRITE_ONCE(priv->tx_coal_frames[i],
> > + ec->tx_max_coalesced_frames);
> > priv->tx_coal_timer[i] =
> > ec->tx_coalesce_usecs;
> > }
> > } else if (queue < tx_cnt) {
> > - priv->tx_coal_frames[queue] =
> > - ec->tx_max_coalesced_frames;
> > + WRITE_ONCE(priv->tx_coal_frames[queue],
> > + ec->tx_max_coalesced_frames);
> > priv->tx_coal_timer[queue] =
> > ec->tx_coalesce_usecs;
> > }
>
> [Severity: Low]
>
> This is a pre-existing issue and was not introduced by this patch.
> tx_coal_frames[] and tx_coal_timer[] are still published as two separate
> stores, with no synchronisation against the TX path.
as pointed out by sashiko, this is a pre-existing issue and it needs to be
fixed with a dedicated patch.
Regards,
Lorenzo
>
> Can a concurrent stmmac_xmit() see a mixed pair? Take a change from
> (frames=N, usecs=0) to (frames=0, usecs=U). The TX path could read the
> new frames value:
>
> tx_coal = READ_ONCE(priv->tx_coal_frames[queue]);
> ...
> else if (!tx_coal)
> set_ic = false;
>
> It could then read the old timer value in stmmac_tx_timer_arm():
>
> u32 tx_coal_timer = priv->tx_coal_timer[queue];
>
> if (!tx_coal_timer)
> return;
>
> That frame would then get neither an IC interrupt nor the completion
> timer. Its completion would wait for the next TX or NAPI activity on that
> queue. The same window exists for stmmac_tso_xmit() and
> stmmac_xdp_xmit_xdpf().
>
> READ_ONCE()/WRITE_ONCE() make each field consistent on its own, which is
> all the commit message claims. The window across the two fields is still
> there.
>
> [ ... ]
>
> --
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261001-b4-stmmac-xmit-mb-v4-0-664e0c3eb289%40oss.qualcomm.com
[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 228 bytes --]
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH net-next v4 2/2] net: stmmac: add XDP multi-buff support for TX side
2026-10-01 9:26 [PATCH net-next v4 0/2] net: stmmac: introduce XDP multi-buff support for TX path Lorenzo Bianconi
2026-10-01 9:26 ` [PATCH net-next v4 1/2] net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path Lorenzo Bianconi
@ 2026-10-01 9:26 ` Lorenzo Bianconi
2026-10-07 23:40 ` [PATCH net-next v4 0/2] net: stmmac: introduce XDP multi-buff support for TX path patchwork-bot+netdevbpf
2 siblings, 0 replies; 6+ messages in thread
From: Lorenzo Bianconi @ 2026-10-01 9:26 UTC (permalink / raw)
To: Maxime Chevallier, Andrew Lunn, David S. Miller, Eric Dumazet,
Jakub Kicinski, Paolo Abeni, Maxime Coquelin, Alexandre Torgue,
Alexei Starovoitov, Daniel Borkmann, Jesper Dangaard Brouer,
John Fastabend, Stanislav Fomichev, Ong Boon Leong
Cc: netdev, linux-stm32, linux-arm-kernel, bpf, Lorenzo Bianconi
Extend stmmac_xdp_xmit_xdpf() to transmit XDP frames with fragments
(multi-buff). Each buffer, i.e. the frame head and every frag, is mapped
and programmed into a dedicated TX descriptor.
Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
---
drivers/net/ethernet/stmicro/stmmac/stmmac_main.c | 122 +++++++++++++++-------
drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c | 2 +-
2 files changed, 85 insertions(+), 39 deletions(-)
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
index df6a329fdb80..508e865f9387 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c
@@ -5249,73 +5249,119 @@ static unsigned int stmmac_rx_buf2_len(struct stmmac_priv *priv,
static int stmmac_xdp_xmit_xdpf(struct stmmac_priv *priv, int queue,
struct xdp_frame *xdpf, bool dma_map)
{
- struct stmmac_txq_stats *txq_stats = &priv->xstats.txq_stats[queue];
+ struct skb_shared_info *sinfo = xdp_get_shared_info_from_frame(xdpf);
struct stmmac_tx_queue *tx_q = &priv->dma_conf.tx_queue[queue];
bool csum = !priv->plat->tx_queues_cfg[queue].coe_unsupported;
+ unsigned int txq_thr = STMMAC_TX_THRESH(priv);
+ unsigned int first_entry = tx_q->cur_tx;
unsigned int entry = tx_q->cur_tx;
- enum stmmac_txbuf_type buf_type;
- struct dma_desc *tx_desc;
- dma_addr_t dma_addr;
+ unsigned int num_frames = 1;
+ struct dma_desc *desc;
u32 tx_coal;
+ int i = 0;
- if (stmmac_tx_avail(priv, queue) < STMMAC_TX_THRESH(priv))
+ if (unlikely(xdp_frame_has_frags(xdpf)))
+ num_frames += sinfo->nr_frags;
+
+ if (stmmac_tx_avail(priv, queue) < max(num_frames, txq_thr))
return STMMAC_XDP_CONSUMED;
if (priv->est && priv->est->enable &&
priv->est->max_sdu[queue] &&
- xdpf->len > priv->est->max_sdu[queue]) {
+ xdp_get_frame_len(xdpf) > priv->est->max_sdu[queue]) {
priv->xstats.max_sdu_txq_drop[queue]++;
return STMMAC_XDP_CONSUMED;
}
- tx_desc = stmmac_get_tx_desc(priv, tx_q, entry);
- if (dma_map) {
- dma_addr = dma_map_single(priv->device, xdpf->data,
- xdpf->len, DMA_TO_DEVICE);
- if (dma_mapping_error(priv->device, dma_addr))
- return STMMAC_XDP_CONSUMED;
-
- buf_type = STMMAC_TXBUF_T_XDP_NDO;
- } else {
- struct page *page = virt_to_page(xdpf->data);
-
- dma_addr = page_pool_get_dma_addr(page) + sizeof(*xdpf) +
- xdpf->headroom;
- dma_sync_single_for_device(priv->device, dma_addr,
- xdpf->len, DMA_BIDIRECTIONAL);
-
- buf_type = STMMAC_TXBUF_T_XDP_TX;
- }
+ while (true) {
+ skb_frag_t *frag = i ? &sinfo->frags[i - 1] : NULL;
+ int len = frag ? skb_frag_size(frag) : xdpf->len;
+ bool last_frame = i == num_frames - 1;
+ enum stmmac_txbuf_type buf_type;
+ dma_addr_t dma_addr;
- stmmac_set_tx_dma_entry(tx_q, entry, buf_type, dma_addr, xdpf->len,
- false);
- stmmac_set_tx_dma_last_segment(tx_q, entry);
+ desc = stmmac_get_tx_desc(priv, tx_q, entry);
+ if (dma_map) {
+ if (frag)
+ dma_addr = skb_frag_dma_map(priv->device,
+ frag, 0, len,
+ DMA_TO_DEVICE);
+ else
+ dma_addr = dma_map_single(priv->device,
+ xdpf->data, len,
+ DMA_TO_DEVICE);
+ if (dma_mapping_error(priv->device, dma_addr))
+ goto error_dma_unmap;
- tx_q->xdpf[entry] = xdpf;
+ buf_type = STMMAC_TXBUF_T_XDP_NDO;
+ } else {
+ struct page *page;
- stmmac_set_desc_addr(priv, tx_desc, dma_addr);
+ page = frag ? skb_frag_page(frag)
+ : virt_to_page(xdpf->data);
+ dma_addr = page_pool_get_dma_addr(page);
+ if (frag)
+ dma_addr += skb_frag_off(frag);
+ else
+ dma_addr += sizeof(*xdpf) + xdpf->headroom;
+ dma_sync_single_for_device(priv->device, dma_addr,
+ len, DMA_BIDIRECTIONAL);
+ buf_type = STMMAC_TXBUF_T_XDP_TX;
+ }
- stmmac_prepare_tx_desc(priv, tx_desc, 1, xdpf->len,
- csum, priv->descriptor_mode, true, true,
- xdpf->len);
+ stmmac_set_tx_dma_entry(tx_q, entry, buf_type, dma_addr, len,
+ dma_map && frag);
+ stmmac_set_desc_addr(priv, desc, dma_addr);
+ stmmac_prepare_tx_desc(priv, desc, !i, len, csum,
+ priv->descriptor_mode, !!i, last_frame,
+ xdp_get_frame_len(xdpf));
+ tx_q->xdpf[entry] = last_frame ? xdpf : NULL;
+ if (last_frame) {
+ stmmac_set_tx_dma_last_segment(tx_q, entry);
+ break;
+ }
- tx_q->tx_count_frames++;
+ entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
+ i++;
+ }
+ tx_q->tx_count_frames += num_frames;
tx_coal = READ_ONCE(priv->tx_coal_frames[queue]);
- if (tx_coal && !(tx_q->tx_count_frames % tx_coal)) {
+ if (tx_coal && (tx_q->tx_count_frames % tx_coal) < num_frames) {
+ struct stmmac_txq_stats *txq_stats;
+
+ desc = stmmac_get_tx_desc(priv, tx_q, entry);
+ stmmac_set_tx_ic(priv, desc);
tx_q->tx_count_frames = 0;
- stmmac_set_tx_ic(priv, tx_desc);
+
+ txq_stats = &priv->xstats.txq_stats[queue];
u64_stats_update_begin(&txq_stats->q_syncp);
u64_stats_inc(&txq_stats->q.tx_set_ic_bit);
u64_stats_update_end(&txq_stats->q_syncp);
}
+ /* Set the OWN bit on the first descriptor now that all descriptors
+ * for this xdp_frame are populated.
+ */
+ desc = stmmac_get_tx_desc(priv, tx_q, first_entry);
+ dma_wmb();
+ stmmac_set_tx_owner(priv, desc);
+ tx_q->cur_tx = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
stmmac_enable_dma_transmission(priv, priv->ioaddr, queue);
- entry = STMMAC_NEXT_ENTRY(entry, priv->dma_conf.dma_tx_size);
- tx_q->cur_tx = entry;
-
return STMMAC_XDP_TX;
+
+error_dma_unmap:
+ while (first_entry != entry) {
+ desc = stmmac_get_tx_desc(priv, tx_q, first_entry);
+ stmmac_release_tx_desc(priv, desc, priv->descriptor_mode);
+ stmmac_free_tx_buffer(priv, &priv->dma_conf, queue,
+ first_entry);
+ first_entry = STMMAC_NEXT_ENTRY(first_entry,
+ priv->dma_conf.dma_tx_size);
+ }
+
+ return STMMAC_XDP_CONSUMED;
}
static int stmmac_xdp_get_tx_queue(struct stmmac_priv *priv,
diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
index d7e4db7224b0..9af9e0cb366c 100644
--- a/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
+++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_xdp.c
@@ -135,7 +135,7 @@ int stmmac_xdp_set_prog(struct stmmac_priv *priv, struct bpf_prog *prog,
stmmac_xdp_open(dev);
if (prog)
- xdp_features_set_redirect_target(dev, false);
+ xdp_features_set_redirect_target(dev, true);
return 0;
}
--
2.55.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* Re: [PATCH net-next v4 0/2] net: stmmac: introduce XDP multi-buff support for TX path
2026-10-01 9:26 [PATCH net-next v4 0/2] net: stmmac: introduce XDP multi-buff support for TX path Lorenzo Bianconi
2026-10-01 9:26 ` [PATCH net-next v4 1/2] net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path Lorenzo Bianconi
2026-10-01 9:26 ` [PATCH net-next v4 2/2] net: stmmac: add XDP multi-buff support for TX side Lorenzo Bianconi
@ 2026-10-07 23:40 ` patchwork-bot+netdevbpf
2 siblings, 0 replies; 6+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-10-07 23:40 UTC (permalink / raw)
To: Lorenzo Bianconi
Cc: maxime.chevallier, andrew+netdev, davem, edumazet, kuba, pabeni,
mcoquelin.stm32, alexandre.torgue, ast, daniel, hawk,
john.fastabend, sdf, boon.leong.ong, netdev, linux-stm32,
linux-arm-kernel, bpf
Hello:
This series was applied to netdev/net-next.git (main)
by Jakub Kicinski <kuba@kernel.org>:
On Thu, 01 Oct 2026 11:26:11 +0200 you wrote:
> Extend stmmac_xdp_xmit_xdpf() to transmit XDP frames with fragments.
> Fix TX coalesce race and div-by-zero in XDP/legacy path
>
> ---
> Changes in v4:
> - Add patch 1/2 to fix theoretical division by zero issue in xdp/legacy
> tx path.
> - Link to v3: https://lore.kernel.org/r/20260925-b4-stmmac-xmit-mb-v3-1-ca08f029e81c@oss.qualcomm.com
>
> [...]
Here is the summary with links:
- [net-next,v4,1/2] net: stmmac: fix TX coalesce race and div-by-zero in XDP/legacy path
https://git.kernel.org/netdev/net-next/c/237782c3d416
- [net-next,v4,2/2] net: stmmac: add XDP multi-buff support for TX side
https://git.kernel.org/netdev/net-next/c/74c553648fb8
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 6+ messages in thread