From: Xuan Zhuo <xuanzhuo@linux.alibaba.com>
To: Jason Wang <jasowang@redhat.com>
Cc: netdev@vger.kernel.org, "Michael S. Tsirkin" <mst@redhat.com>,
"Eugenio Pérez" <eperezma@redhat.com>,
"David S. Miller" <davem@davemloft.net>,
"Eric Dumazet" <edumazet@google.com>,
"Jakub Kicinski" <kuba@kernel.org>,
"Paolo Abeni" <pabeni@redhat.com>,
"Alexei Starovoitov" <ast@kernel.org>,
"Daniel Borkmann" <daniel@iogearbox.net>,
"Jesper Dangaard Brouer" <hawk@kernel.org>,
"John Fastabend" <john.fastabend@gmail.com>,
virtualization@lists.linux.dev, bpf@vger.kernel.org
Subject: Re: [PATCH net-next v7 09/10] virtio_net: xsk: rx: support recv small mode
Date: Mon, 8 Jul 2024 16:09:48 +0800 [thread overview]
Message-ID: <1720426188.2428002-3-xuanzhuo@linux.alibaba.com> (raw)
In-Reply-To: <CACGkMEukkp9FxLfBGTXvSGso48Ugy2-m3rWNFiVGuEa52LT_-Q@mail.gmail.com>
On Mon, 8 Jul 2024 16:08:44 +0800, Jason Wang <jasowang@redhat.com> wrote:
> On Mon, Jul 8, 2024 at 3:47 PM Xuan Zhuo <xuanzhuo@linux.alibaba.com> wrote:
> >
> > On Mon, 8 Jul 2024 15:00:50 +0800, Jason Wang <jasowang@redhat.com> wrote:
> > > On Fri, Jul 5, 2024 at 3:38 PM Xuan Zhuo <xuanzhuo@linux.alibaba.com> wrote:
> > > >
> > > > In the process:
> > > > 1. We may need to copy data to create skb for XDP_PASS.
> > > > 2. We may need to call xsk_buff_free() to release the buffer.
> > > > 3. The handle for xdp_buff is difference from the buffer.
> > > >
> > > > If we pushed this logic into existing receive handle(merge and small),
> > > > we would have to maintain code scattered inside merge and small (and big).
> > > > So I think it is a good choice for us to put the xsk code into an
> > > > independent function.
> > > >
> > > > Signed-off-by: Xuan Zhuo <xuanzhuo@linux.alibaba.com>
> > > > ---
> > > >
> > > > v7:
> > > > 1. rename xdp_construct_skb to xsk_construct_skb
> > > > 2. refactor virtnet_receive()
> > > >
> > > > drivers/net/virtio_net.c | 176 +++++++++++++++++++++++++++++++++++++--
> > > > 1 file changed, 168 insertions(+), 8 deletions(-)
> > > >
> > > > diff --git a/drivers/net/virtio_net.c b/drivers/net/virtio_net.c
> > > > index 2b27f5ada64a..64d8cd481890 100644
> > > > --- a/drivers/net/virtio_net.c
> > > > +++ b/drivers/net/virtio_net.c
> > > > @@ -498,6 +498,12 @@ struct virtio_net_common_hdr {
> > > > };
> > > >
> > > > static void virtnet_sq_free_unused_buf(struct virtqueue *vq, void *buf);
> > > > +static int virtnet_xdp_handler(struct bpf_prog *xdp_prog, struct xdp_buff *xdp,
> > > > + struct net_device *dev,
> > > > + unsigned int *xdp_xmit,
> > > > + struct virtnet_rq_stats *stats);
> > > > +static void virtnet_receive_done(struct virtnet_info *vi, struct receive_queue *rq,
> > > > + struct sk_buff *skb, u8 flags);
> > > >
> > > > static bool is_xdp_frame(void *ptr)
> > > > {
> > > > @@ -1062,6 +1068,124 @@ static void sg_fill_dma(struct scatterlist *sg, dma_addr_t addr, u32 len)
> > > > sg->length = len;
> > > > }
> > > >
> > > > +static struct xdp_buff *buf_to_xdp(struct virtnet_info *vi,
> > > > + struct receive_queue *rq, void *buf, u32 len)
> > > > +{
> > > > + struct xdp_buff *xdp;
> > > > + u32 bufsize;
> > > > +
> > > > + xdp = (struct xdp_buff *)buf;
> > > > +
> > > > + bufsize = xsk_pool_get_rx_frame_size(rq->xsk_pool) + vi->hdr_len;
> > > > +
> > > > + if (unlikely(len > bufsize)) {
> > > > + pr_debug("%s: rx error: len %u exceeds truesize %u\n",
> > > > + vi->dev->name, len, bufsize);
> > > > + DEV_STATS_INC(vi->dev, rx_length_errors);
> > > > + xsk_buff_free(xdp);
> > > > + return NULL;
> > > > + }
> > > > +
> > > > + xsk_buff_set_size(xdp, len);
> > > > + xsk_buff_dma_sync_for_cpu(xdp);
> > > > +
> > > > + return xdp;
> > > > +}
> > > > +
> > > > +static struct sk_buff *xsk_construct_skb(struct receive_queue *rq,
> > > > + struct xdp_buff *xdp)
> > > > +{
> > > > + unsigned int metasize = xdp->data - xdp->data_meta;
> > > > + struct sk_buff *skb;
> > > > + unsigned int size;
> > > > +
> > > > + size = xdp->data_end - xdp->data_hard_start;
> > > > + skb = napi_alloc_skb(&rq->napi, size);
> > > > + if (unlikely(!skb)) {
> > > > + xsk_buff_free(xdp);
> > > > + return NULL;
> > > > + }
> > > > +
> > > > + skb_reserve(skb, xdp->data_meta - xdp->data_hard_start);
> > > > +
> > > > + size = xdp->data_end - xdp->data_meta;
> > > > + memcpy(__skb_put(skb, size), xdp->data_meta, size);
> > > > +
> > > > + if (metasize) {
> > > > + __skb_pull(skb, metasize);
> > > > + skb_metadata_set(skb, metasize);
> > > > + }
> > > > +
> > > > + xsk_buff_free(xdp);
> > > > +
> > > > + return skb;
> > > > +}
> > > > +
> > > > +static struct sk_buff *virtnet_receive_xsk_small(struct net_device *dev, struct virtnet_info *vi,
> > > > + struct receive_queue *rq, struct xdp_buff *xdp,
> > > > + unsigned int *xdp_xmit,
> > > > + struct virtnet_rq_stats *stats)
> > > > +{
> > > > + struct bpf_prog *prog;
> > > > + u32 ret;
> > > > +
> > > > + ret = XDP_PASS;
> > > > + rcu_read_lock();
> > > > + prog = rcu_dereference(rq->xdp_prog);
> > > > + if (prog)
> > > > + ret = virtnet_xdp_handler(prog, xdp, dev, xdp_xmit, stats);
> > > > + rcu_read_unlock();
> > > > +
> > > > + switch (ret) {
> > > > + case XDP_PASS:
> > > > + return xsk_construct_skb(rq, xdp);
> > > > +
> > > > + case XDP_TX:
> > > > + case XDP_REDIRECT:
> > > > + return NULL;
> > > > +
> > > > + default:
> > > > + /* drop packet */
> > > > + xsk_buff_free(xdp);
> > > > + u64_stats_inc(&stats->drops);
> > > > + return NULL;
> > > > + }
> > > > +}
> > > > +
> > > > +static void virtnet_receive_xsk_buf(struct virtnet_info *vi, struct receive_queue *rq,
> > > > + void *buf, u32 len,
> > > > + unsigned int *xdp_xmit,
> > > > + struct virtnet_rq_stats *stats)
> > > > +{
> > > > + struct net_device *dev = vi->dev;
> > > > + struct sk_buff *skb = NULL;
> > > > + struct xdp_buff *xdp;
> > > > + u8 flags;
> > > > +
> > > > + len -= vi->hdr_len;
> > > > +
> > > > + u64_stats_add(&stats->bytes, len);
> > > > +
> > > > + xdp = buf_to_xdp(vi, rq, buf, len);
> > > > + if (!xdp)
> > > > + return;
> > > > +
> > > > + if (unlikely(len < ETH_HLEN)) {
> > > > + pr_debug("%s: short packet %i\n", dev->name, len);
> > > > + DEV_STATS_INC(dev, rx_length_errors);
> > > > + xsk_buff_free(xdp);
> > > > + return;
> > > > + }
> > > > +
> > > > + flags = ((struct virtio_net_common_hdr *)(xdp->data - vi->hdr_len))->hdr.flags;
> > > > +
> > > > + if (!vi->mergeable_rx_bufs)
> > > > + skb = virtnet_receive_xsk_small(dev, vi, rq, xdp, xdp_xmit, stats);
> > >
> > > I wonder if we add the mergeable support in the next patch would it be
> > > better to re-order the patch? For example, the xsk binding needs to be
> > > moved to the last patch, otherwise we break xsk with a mergeable
> > > buffer here?
> >
> > If you worry that the user works with this commit, I want to say you do not
> > worry.
> >
> > Because the flags NETDEV_XDP_ACT_XSK_ZEROCOPY is not added. I plan to add that
> > after the tx is completed.
>
> Ok, this is something I missed, it would be better to mention it
> somewhere (or it is already there but I miss it).
OK. I will add it to next version cover.
Thanks.
>
> >
> > I do test by adding this flags locally.
> >
> > Thanks.
>
> Acked-by: Jason Wang <jasowang@redhat.com>
>
> Thanks
>
> >
> > >
> > > Or anything I missed here?
> > >
> > > Thanks
> > >
> >
>
next prev parent reply other threads:[~2024-07-08 8:11 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2024-07-05 7:37 [PATCH net-next v7 00/10] virtio-net: support AF_XDP zero copy Xuan Zhuo
2024-07-05 7:37 ` [PATCH net-next v7 01/10] virtio_net: replace VIRTIO_XDP_HEADROOM by XDP_PACKET_HEADROOM Xuan Zhuo
2024-07-08 6:18 ` Jason Wang
2024-07-05 7:37 ` [PATCH net-next v7 02/10] virtio_net: separate virtnet_rx_resize() Xuan Zhuo
2024-07-05 7:37 ` [PATCH net-next v7 03/10] virtio_net: separate virtnet_tx_resize() Xuan Zhuo
2024-07-05 7:37 ` [PATCH net-next v7 04/10] virtio_net: separate receive_buf Xuan Zhuo
2024-07-05 7:37 ` [PATCH net-next v7 05/10] virtio_net: separate receive_mergeable Xuan Zhuo
2024-07-05 7:37 ` [PATCH net-next v7 06/10] virtio_net: xsk: bind/unbind xsk for rx Xuan Zhuo
2024-07-08 6:36 ` Jason Wang
2024-07-05 7:37 ` [PATCH net-next v7 07/10] virtio_net: xsk: support wakeup Xuan Zhuo
2024-07-05 7:37 ` [PATCH net-next v7 08/10] virtio_net: xsk: rx: support fill with xsk buffer Xuan Zhuo
2024-07-08 6:49 ` Jason Wang
2024-07-08 7:57 ` Xuan Zhuo
2024-07-05 7:37 ` [PATCH net-next v7 09/10] virtio_net: xsk: rx: support recv small mode Xuan Zhuo
2024-07-08 7:00 ` Jason Wang
2024-07-08 7:42 ` Xuan Zhuo
2024-07-08 8:08 ` Jason Wang
2024-07-08 8:09 ` Xuan Zhuo [this message]
2024-07-05 7:37 ` [PATCH net-next v7 10/10] virtio_net: xsk: rx: support recv merge mode Xuan Zhuo
2024-07-08 8:10 ` Jason Wang
2024-07-05 14:14 ` [PATCH net-next v7 00/10] virtio-net: support AF_XDP zero copy Michal Kubiak
2024-07-08 1:11 ` Xuan Zhuo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1720426188.2428002-3-xuanzhuo@linux.alibaba.com \
--to=xuanzhuo@linux.alibaba.com \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=eperezma@redhat.com \
--cc=hawk@kernel.org \
--cc=jasowang@redhat.com \
--cc=john.fastabend@gmail.com \
--cc=kuba@kernel.org \
--cc=mst@redhat.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=virtualization@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.