From: Mohsin Bashir <mohsin.bashr@gmail.com>
To: Chenguang Zhao <chenguang.zhao@linux.dev>,
anthony.l.nguyen@intel.com, przemyslaw.kitszel@intel.com,
andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com,
kuba@kernel.org, pabeni@redhat.com
Cc: intel-wired-lan@lists.osuosl.org, netdev@vger.kernel.org,
kerneljasonxing@gmail.com,
Chenguang Zhao <zhaochenguang@kylinos.cn>
Subject: Re: [PATCH net-next v3] i40e: xsk: use xdp_build_skb_from_zc() for XDP_PASS
Date: Fri, 31 Jul 2026 08:22:33 -0700 [thread overview]
Message-ID: <34c8d95f-3897-498a-921c-df7c6868af61@gmail.com> (raw)
In-Reply-To: <20260731020736.170485-1-chenguang.zhao@linux.dev>
On 7/30/26 7:07 PM, Chenguang Zhao wrote:
> From: Chenguang Zhao <zhaochenguang@kylinos.cn>
>
> Replace the driver-local i40e_construct_skb_zc() with the common
> helper xdp_build_skb_from_zc(). On failure, free the xdp buff in
> the caller.
>
> xdp_build_skb_from_zc() already calls skb_record_rx_queue() and
> eth_type_trans(), so pull the remaining descriptor field setup into
> __i40e_process_skb_fields() and use that on the XDP_PASS path.
>
> Signed-off-by: Chenguang Zhao <zhaochenguang@kylinos.cn>
> ---
> v3:
> As suggested by Larysa:
> - Drop eth_skb_pad() on the XDP_PASS path, along with the
> temporary __skb_push()/__skb_pull() around it. The stack
> handles short skbs and other vendors do not pad on Rx.
>
> v2:
> - https://lore.kernel.org/all/20260729054916.720750-1-chenguang.zhao@linux.dev/
>
> v1:
> - https://lore.kernel.org/all/20260724020125.246333-1-chenguang.zhao@linux.dev/
> ---
> drivers/net/ethernet/intel/i40e/i40e_txrx.c | 32 ++++++--
> .../ethernet/intel/i40e/i40e_txrx_common.h | 2 +
> drivers/net/ethernet/intel/i40e/i40e_xsk.c | 79 ++-----------------
> 3 files changed, 33 insertions(+), 80 deletions(-)
>
> diff --git a/drivers/net/ethernet/intel/i40e/i40e_txrx.c b/drivers/net/ethernet/intel/i40e/i40e_txrx.c
> index ef5e657816f0..87553b14d34a 100644
> --- a/drivers/net/ethernet/intel/i40e/i40e_txrx.c
> +++ b/drivers/net/ethernet/intel/i40e/i40e_txrx.c
> @@ -1833,17 +1833,17 @@ static inline void i40e_rx_hash(struct i40e_ring *ring,
> }
>
> /**
> - * i40e_process_skb_fields - Populate skb header fields from Rx descriptor
> + * __i40e_process_skb_fields - Populate skb fields from Rx descriptor
> * @rx_ring: rx descriptor ring packet is being transacted on
> * @rx_desc: pointer to the EOP Rx descriptor
> * @skb: pointer to current skb being populated
> *
> - * This function checks the ring, descriptor, and packet information in
> - * order to populate the hash, checksum, VLAN, protocol, and
> - * other fields within the skb.
> + * Populate hash, checksum, PTP timestamp and VLAN from @rx_desc. Does not
> + * call skb_record_rx_queue() or eth_type_trans(); callers that already got
> + * those from xdp_build_skb_from_zc() should use this helper.
> **/
> -void i40e_process_skb_fields(struct i40e_ring *rx_ring,
> - union i40e_rx_desc *rx_desc, struct sk_buff *skb)
> +void __i40e_process_skb_fields(struct i40e_ring *rx_ring,
> + union i40e_rx_desc *rx_desc, struct sk_buff *skb)
> {
> u64 qword = le64_to_cpu(rx_desc->wb.qword1.status_error_len);
> u32 rx_status = FIELD_GET(I40E_RXD_QW1_STATUS_MASK, qword);
> @@ -1858,14 +1858,30 @@ void i40e_process_skb_fields(struct i40e_ring *rx_ring,
>
> i40e_rx_checksum(rx_ring->vsi, skb, rx_desc);
>
> - skb_record_rx_queue(skb, rx_ring->queue_index);
> -
> if (qword & BIT(I40E_RX_DESC_STATUS_L2TAG1P_SHIFT)) {
> __le16 vlan_tag = rx_desc->wb.qword0.lo_dword.l2tag1;
>
> __vlan_hwaccel_put_tag(skb, htons(ETH_P_8021Q),
> le16_to_cpu(vlan_tag));
> }
> +}
> +
> +/**
> + * i40e_process_skb_fields - Populate skb header fields from Rx descriptor
> + * @rx_ring: rx descriptor ring packet is being transacted on
> + * @rx_desc: pointer to the EOP Rx descriptor
> + * @skb: pointer to current skb being populated
> + *
> + * This function checks the ring, descriptor, and packet information in
> + * order to populate the hash, checksum, VLAN, protocol, and
> + * other fields within the skb.
> + **/
> +void i40e_process_skb_fields(struct i40e_ring *rx_ring,
> + union i40e_rx_desc *rx_desc, struct sk_buff *skb)
> +{
> + __i40e_process_skb_fields(rx_ring, rx_desc, skb);
> +
> + skb_record_rx_queue(skb, rx_ring->queue_index);
>
> /* modifies the skb - consumes the enet header */
> skb->protocol = eth_type_trans(skb, rx_ring->netdev);
> diff --git a/drivers/net/ethernet/intel/i40e/i40e_txrx_common.h b/drivers/net/ethernet/intel/i40e/i40e_txrx_common.h
> index e26807fd2123..3c52e0e60487 100644
> --- a/drivers/net/ethernet/intel/i40e/i40e_txrx_common.h
> +++ b/drivers/net/ethernet/intel/i40e/i40e_txrx_common.h
> @@ -9,6 +9,8 @@
> int i40e_xmit_xdp_tx_ring(struct xdp_buff *xdp, struct i40e_ring *xdp_ring);
> void i40e_clean_programming_status(struct i40e_ring *rx_ring, u64 qword0_raw,
> u64 qword1);
> +void __i40e_process_skb_fields(struct i40e_ring *rx_ring,
> + union i40e_rx_desc *rx_desc, struct sk_buff *skb);
> void i40e_process_skb_fields(struct i40e_ring *rx_ring,
> union i40e_rx_desc *rx_desc, struct sk_buff *skb);
> void i40e_xdp_ring_update_tail(struct i40e_ring *xdp_ring);
> diff --git a/drivers/net/ethernet/intel/i40e/i40e_xsk.c b/drivers/net/ethernet/intel/i40e/i40e_xsk.c
> index 9f47388eaba5..dfa519c30a0d 100644
> --- a/drivers/net/ethernet/intel/i40e/i40e_xsk.c
> +++ b/drivers/net/ethernet/intel/i40e/i40e_xsk.c
> @@ -3,6 +3,7 @@
>
> #include <linux/bpf_trace.h>
> #include <linux/unroll.h>
> +#include <net/xdp.h>
> #include <net/xdp_sock_drv.h>
> #include "i40e_txrx_common.h"
> #include "i40e_xsk.h"
> @@ -277,70 +278,6 @@ bool i40e_alloc_rx_buffers_zc(struct i40e_ring *rx_ring, u16 count)
> return count == nb_buffs;
> }
>
> -/**
> - * i40e_construct_skb_zc - Create skbuff from zero-copy Rx buffer
> - * @rx_ring: Rx ring
> - * @xdp: xdp_buff
> - *
> - * This functions allocates a new skb from a zero-copy Rx buffer.
> - *
> - * Returns the skb, or NULL on failure.
> - **/
> -static struct sk_buff *i40e_construct_skb_zc(struct i40e_ring *rx_ring,
> - struct xdp_buff *xdp)
> -{
> - unsigned int totalsize = xdp->data_end - xdp->data_meta;
> - unsigned int metasize = xdp->data - xdp->data_meta;
> - struct skb_shared_info *sinfo = NULL;
> - struct sk_buff *skb;
> - u32 nr_frags = 0;
> -
> - if (unlikely(xdp_buff_has_frags(xdp))) {
> - sinfo = xdp_get_shared_info_from_buff(xdp);
> - nr_frags = sinfo->nr_frags;
> - }
> - net_prefetch(xdp->data_meta);
> -
> - /* allocate a skb to store the frags */
> - skb = napi_alloc_skb(&rx_ring->q_vector->napi, totalsize);
> - if (unlikely(!skb))
> - goto out;
> -
> - memcpy(__skb_put(skb, totalsize), xdp->data_meta,
> - ALIGN(totalsize, sizeof(long)));
> -
> - if (metasize) {
> - skb_metadata_set(skb, metasize);
> - __skb_pull(skb, metasize);
> - }
> -
> - if (likely(!xdp_buff_has_frags(xdp)))
> - goto out;
> -
> - for (int i = 0; i < nr_frags; i++) {
> - struct skb_shared_info *skinfo = skb_shinfo(skb);
> - skb_frag_t *frag = &sinfo->frags[i];
> - struct page *page;
> - void *addr;
> -
> - page = dev_alloc_page();
> - if (!page) {
> - dev_kfree_skb(skb);
> - return NULL;
> - }
> - addr = page_to_virt(page);
> -
> - memcpy(addr, skb_frag_page(frag), skb_frag_size(frag));
> -
> - __skb_fill_page_desc_noacc(skinfo, skinfo->nr_frags++,
> - addr, 0, skb_frag_size(frag));
> - }
> -
> -out:
> - xsk_buff_free(xdp);
> - return skb;
> -}
> -
> static void i40e_handle_xdp_result_zc(struct i40e_ring *rx_ring,
> struct xdp_buff *xdp_buff,
> union i40e_rx_desc *rx_desc,
> @@ -372,21 +309,19 @@ static void i40e_handle_xdp_result_zc(struct i40e_ring *rx_ring,
> * BIT(I40E_RXD_QW1_ERROR_SHIFT). This is due to that
> * SBP is *not* set in PRT_SBPVSI (default not set).
> */
> - skb = i40e_construct_skb_zc(rx_ring, xdp_buff);
> + skb = xdp_build_skb_from_zc(xdp_buff);
> if (!skb) {
> + xsk_buff_free(xdp_buff);
> rx_ring->rx_stats.alloc_buff_failed++;
> *rx_packets = 0;
> *rx_bytes = 0;
> return;
> }
>
> - if (eth_skb_pad(skb)) {
> - *rx_packets = 0;
> - *rx_bytes = 0;
> - return;
> - }
> -
> - i40e_process_skb_fields(rx_ring, rx_desc, skb);
> + /* xdp_build_skb_from_zc() already ran eth_type_trans() and
> + * skb_record_rx_queue().
> + */
> + __i40e_process_skb_fields(rx_ring, rx_desc, skb);
> napi_gro_receive(&rx_ring->q_vector->napi, skb);
> return;
> }
The cleanup makes sense. I think you would also need to update kconfig.
xdp_build_skb_from_zc() verifies page pool support and if
CONFIG_PAGE_POOL=n, it would return NULL. In that case, we would end-up
dropping packets in XDP_PASS.
With that addressed:
Reviewed-by: Mohsin Bashir <hmohsin@meta.com>
prev parent reply other threads:[~2026-07-31 15:22 UTC|newest]
Thread overview: 3+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-31 2:07 [PATCH net-next v3] i40e: xsk: use xdp_build_skb_from_zc() for XDP_PASS Chenguang Zhao
2026-07-31 5:46 ` [Intel-wired-lan] " Loktionov, Aleksandr
2026-07-31 15:22 ` Mohsin Bashir [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=34c8d95f-3897-498a-921c-df7c6868af61@gmail.com \
--to=mohsin.bashr@gmail.com \
--cc=andrew+netdev@lunn.ch \
--cc=anthony.l.nguyen@intel.com \
--cc=chenguang.zhao@linux.dev \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=intel-wired-lan@lists.osuosl.org \
--cc=kerneljasonxing@gmail.com \
--cc=kuba@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=przemyslaw.kitszel@intel.com \
--cc=zhaochenguang@kylinos.cn \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox