BPF List
 help / color / mirror / Atom feed
From: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
To: <netdev@vger.kernel.org>
Cc: <bpf@vger.kernel.org>, <magnus.karlsson@intel.com>,
	<stfomichev@gmail.com>, <kuba@kernel.org>, <pabeni@redhat.com>,
	<horms@kernel.org>, <bjorn@kernel.org>,
	<kerneljasonxing@gmail.com>,
	"Jason Xing" <kernelxing@tencent.com>
Subject: Re: [PATCH v4 net 2/6] xsk: drain continuation descs after overflow in xsk_build_skb()
Date: Tue, 21 Jul 2026 20:04:24 +0200	[thread overview]
Message-ID: <al+0qKu1U7kutrih@boxer> (raw)
In-Reply-To: <20260719135609.147823-3-maciej.fijalkowski@intel.com>

On Sun, Jul 19, 2026 at 03:56:05PM +0200, Maciej Fijalkowski wrote:
> From: Jason Xing <kernelxing@tencent.com>
> 
> Fix generic xmit path multi-buffer logic when packets are either too big
> (count of descriptors exceed MAX_SKB_FRAGS) or an invalid descriptor is
> included in fragmented packet. Introduce xdp_sock::drain_cont and act
> upon this flag - when it is set, keep on consuming descriptors from
> AF_XDP Tx ring and put them directly onto Cq. Previously these
> descriptors were silently lost and could never be reached again.
> 
> Fixes: cf24f5a5feea ("xsk: add support for AF_XDP multi-buffer on Tx path")
> Closes: https://lore.kernel.org/all/20260425041726.85FB3C2BCB2@smtp.kernel.org/
> Reviewed-by: Jason Xing <kernelxing@tencent.com>
> Co-developed-by: Maciej Fijalkowski <maciej.fijalkowski@intel.com> # wrapped cq addr submission onto routine
> Signed-off-by: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
> Signed-off-by: Jason Xing <kernelxing@tencent.com>
> ---
>  include/net/xdp_sock.h |  1 +
>  net/xdp/xsk.c          | 45 +++++++++++++++++++++++++++++++++++++++---
>  2 files changed, 43 insertions(+), 3 deletions(-)
> 
> diff --git a/include/net/xdp_sock.h b/include/net/xdp_sock.h
> index ebac60a3d8a1..8b51876efbed 100644
> --- a/include/net/xdp_sock.h
> +++ b/include/net/xdp_sock.h
> @@ -80,6 +80,7 @@ struct xdp_sock {
>  	 * call of __xsk_generic_xmit().
>  	 */
>  	struct sk_buff *skb;
> +	bool drain_cont;
>  
>  	struct list_head map_list;
>  	/* Protects map_list */
> diff --git a/net/xdp/xsk.c b/net/xdp/xsk.c
> index a7a83dc4546a..12a845d012f6 100644
> --- a/net/xdp/xsk.c
> +++ b/net/xdp/xsk.c
> @@ -737,6 +737,19 @@ static void xsk_cq_submit_addr_locked(struct xsk_buff_pool *pool,
>  	spin_unlock_irqrestore(&pool->cq_prod_lock, flags);
>  }
>  
> +static void xsk_cq_submit_addr_single_locked(struct xsk_buff_pool *pool,
> +					     struct xdp_desc *desc)
> +{
> +	unsigned long flags;
> +	u32 idx;
> +
> +	spin_lock_irqsave(&pool->cq_prod_lock, flags);
> +	idx = xskq_get_prod(pool->cq);
> +	xskq_prod_write_addr(pool->cq, idx, desc->addr);
> +	xskq_prod_submit_n(pool->cq, 1);
> +	spin_unlock_irqrestore(&pool->cq_prod_lock, flags);
> +}
> +
>  static void xsk_cq_cancel_locked(struct xsk_buff_pool *pool, u32 n)
>  {
>  	spin_lock(&pool->cq->cq_cached_prod_lock);
> @@ -1028,13 +1041,14 @@ static struct sk_buff *xsk_build_skb(struct xdp_sock *xs,
>  static int __xsk_generic_xmit(struct sock *sk)
>  {
>  	struct xdp_sock *xs = xdp_sk(sk);
> -	bool sent_frame = false;
>  	struct xdp_desc desc;
>  	struct sk_buff *skb;
> +	u32 cached_cons;
>  	u32 max_batch;
>  	int err = 0;
>  
>  	mutex_lock(&xs->mutex);
> +	cached_cons = xs->tx->cached_cons;
>  
>  	/* Since we dropped the RCU read lock, the socket state might have changed. */
>  	if (unlikely(!xsk_is_bound(xs))) {
> @@ -1063,11 +1077,21 @@ static int __xsk_generic_xmit(struct sock *sk)
>  			goto out;
>  		}
>  
> +		if (unlikely(xs->drain_cont)) {
> +			xsk_cq_submit_addr_single_locked(xs->pool, &desc);
> +			xs->tx->invalid_descs++;
> +			xskq_cons_release(xs->tx);
> +			xs->drain_cont = xp_mb_desc(&desc);
> +			continue;
> +		}
> +
>  		skb = xsk_build_skb(xs, &desc);
>  		if (IS_ERR(skb)) {
>  			err = PTR_ERR(skb);
>  			if (err != -EOVERFLOW)
>  				goto out;
> +			if (xp_mb_desc(&desc))
> +				xs->drain_cont = true;
>  			err = 0;
>  			continue;
>  		}
> @@ -1096,18 +1120,33 @@ static int __xsk_generic_xmit(struct sock *sk)
>  			goto out;
>  		}
>  
> -		sent_frame = true;
>  		xs->skb = NULL;
>  	}
>  
>  	if (xskq_has_descs(xs->tx)) {
> +		bool drain = xs->skb || xs->drain_cont || xp_mb_desc(&desc);
> +
> +		err = xsk_cq_reserve_locked(xs->pool);
> +		if (err) {
> +			xs->tx->invalid_descs--;
> +			if (xs->skb)
> +				xsk_drop_skb(xs->skb);
> +			xs->drain_cont = drain;
> +			err = -EAGAIN;
> +			goto out;
> +		}
> +
>  		if (xs->skb)
>  			xsk_drop_skb(xs->skb);
> +
> +		xsk_cq_submit_addr_single_locked(xs->pool, &desc);
> +
>  		xskq_cons_release(xs->tx);
> +		xs->drain_cont = xp_mb_desc(&desc);
>  	}
>  
>  out:
> -	if (sent_frame)
> +	if (xs->tx->cached_cons != cached_cons)

Sashiko says:

[Severity: High]
This isn't a bug introduced by this patch, but does this change to the
release condition expose a pre-existing issue where rolling back a partial
multi-buffer packet corrupts the Tx ring?

During multi-buffer packet assembly, if the max_batch limit is reached,
__xsk_generic_xmit() exits and commits the partially consumed descriptors
to ring->consumer. User-space observes this and may overwrite those slots.

On the next call to __xsk_generic_xmit(), packet assembly resumes. If
__dev_direct_xmit() subsequently returns NETDEV_TX_BUSY, the code calls
xskq_cons_cancel_n() to rewind the local cached_cons by the full packet
length:

	if (err == NETDEV_TX_BUSY) {
		/* Tell MAC device to consume more and retry later. */
		xskq_cons_cancel_n(xs->tx, xsk_get_num_desc(skb));
		...

Because cached_cons is rewound by the full length (including fragments
consumed in the previous call), it becomes smaller than the cached_cons
saved at the start of this current call.

Will this new condition (xs->tx->cached_cons != cached_cons) then evaluate
to true and commit the rewound index to the globally visible ring->consumer?

If ring->consumer jumps backwards, wouldn't the kernel fetch descriptors
that user-space may have already overwritten with new data upon retry?

Maciej says:

So generic xmit is still not bullet-proof, sigh. It's a problem that was
present even when `sent_frame` based consumer pointer update was used, so
sashiko correctly classified it as pre-existing issue.

I think this can be addressed after current set lands, as no new bugs are
introduced and seems it got acks from Stan and Jason.

>  		__xsk_tx_release(xs);
>  
>  	mutex_unlock(&xs->mutex);
> -- 
> 2.43.0
> 

  parent reply	other threads:[~2026-07-21 18:04 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-19 13:56 [PATCH v4 net 0/6] xsk: fix AF_XDP multi-buffer Tx descriptor reclaim Maciej Fijalkowski
2026-07-19 13:56 ` [PATCH v4 net 1/6] xsk: fix buffer leak in xsk_drop_skb() for AF_XDP multi-buffer Tx Maciej Fijalkowski
2026-07-19 13:56 ` [PATCH v4 net 2/6] xsk: drain continuation descs after overflow in xsk_build_skb() Maciej Fijalkowski
2026-07-20 13:58   ` sashiko-bot
2026-07-21 18:04   ` Maciej Fijalkowski [this message]
2026-07-19 13:56 ` [PATCH v4 net 3/6] xsk: provide sufficient space in pool->tx_descs Maciej Fijalkowski
2026-07-20 13:58   ` sashiko-bot
2026-07-19 13:56 ` [PATCH v4 net 4/6] xsk: reclaim invalid Tx descriptors in ZC batch path Maciej Fijalkowski
2026-07-19 13:56 ` [PATCH v4 net 5/6] selftests/xsk: fix too-many-frags multi-buffer Tx test Maciej Fijalkowski
2026-07-19 13:56 ` [PATCH v4 net 6/6] selftests/xsk: account reclaimed invalid Tx descriptors Maciej Fijalkowski
2026-07-20 19:30 ` [PATCH v4 net 0/6] xsk: fix AF_XDP multi-buffer Tx descriptor reclaim Stanislav Fomichev

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=al+0qKu1U7kutrih@boxer \
    --to=maciej.fijalkowski@intel.com \
    --cc=bjorn@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=horms@kernel.org \
    --cc=kerneljasonxing@gmail.com \
    --cc=kernelxing@tencent.com \
    --cc=kuba@kernel.org \
    --cc=magnus.karlsson@intel.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=stfomichev@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox