Netdev List
 help / color / mirror / Atom feed
From: Jiayuan Chen <jiayuan.chen@linux.dev>
To: netdev@vger.kernel.org
Cc: Jiayuan Chen <jiayuan.chen@linux.dev>,
	Andrew Lunn <andrew+netdev@lunn.ch>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Toshiaki Makita <toshiaki.makita1@gmail.com>,
	John Fastabend <john.fastabend@gmail.com>,
	Daniel Borkmann <daniel@iogearbox.net>,
	linux-kernel@vger.kernel.org
Subject: [PATCH net-next] veth: clear rx queue hint in veth_xmit
Date: Fri,  9 Oct 2026 12:04:11 +0800	[thread overview]
Message-ID: <20261009040412.14571-1-jiayuan.chen@linux.dev> (raw)

With more tx queues on one end of a veth pair than rx queues on the
other, and RPS or generic XDP enabled on the receiving end, we get:

    veth1 received packet on queue 2, but number of RX queues is 2
    WARNING: net/core/dev.c:5212 at get_rps_cpu+0x560/0x1360
    Call Trace:
     <TASK>
     netif_rx_internal+0x1af/0x4c0
     __netif_rx+0x99/0x350
     veth_xmit+0x713/0xca0
     dev_hard_start_xmit+0x166/0x5f0
     __dev_queue_xmit+0x1797/0x42d0
     ip_finish_output2+0xa34/0x1f40
     __ip_finish_output+0x510/0x7e0
     ip_finish_output+0x2f/0x320
     ip_output+0x17a/0x3f0
     ip_send_skb+0x1bc/0x220
     ......

Easy to hit with "ethtool -L veth0 tx 4", "ethtool -L veth1 rx 2",
rps_cpus set on veth1 and a few flows sent over the pair [1].

veth_xmit() uses skb->queue_mapping to pick the peer rq, but never
clears it before veth_forward_skb(), so the rx side still sees the tx
queue index. Everything on the rx side that goes through
skb_get_rx_queue(), like get_rps_cpu() and netif_get_rxqueue() for
generic XDP, takes it as a recorded rx queue and warns once it is out
of range.

Clear it before handing the skb to the peer:

- It is the tx queue index of this device, it says nothing about the
  rx queue of the peer.

- The two sides don't even agree on the encoding. The tx side stores
  the index as is, the rx side stores index + 1 so that 0 can mean
  "not recorded". So tx queue k is read back as rx queue k - 1, and
  tx queue 0 as "not recorded".

- Commit 710ad98c363a ("veth: Do not record rx queue hint in
  veth_xmit") already decided that veth should not pass any queue
  hint to the peer. There is no tx->rx queue mapping to preserve, so
  nothing is lost by clearing it.

On NETDEV_TX_BUSY the skb goes back to the qdisc, which looks up the
txq from skb->queue_mapping to decide when to retry, so restore it
there, next to the existing __skb_push().

[1]: https://lore.kernel.org/netdev/156834bb-8e40-496e-9443-9d515fa18eab@linux.dev/

Fixes: 710ad98c363a ("veth: Do not record rx queue hint in veth_xmit")
Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
---
Target net-next since it is moderate.
Full reproducer, veth1 lives in netns ns1:

  ip netns add ns1
  ip link add veth0 type veth peer name veth1
  ip link set veth1 netns ns1
  ethtool -L veth0  tx 4
  ip netns exec ns1 ethtool -L veth1 rx 2
  ip netns exec ns1 sh -c  'echo f > /sys/class/net/veth1/queues/rx-0/rps_cpus'
  ip addr add 10.9.9.1/24 dev veth0
  ip link set veth0 up
  ip netns exec ns1 ip addr add 10.9.9.2/24 dev veth1
  ip netns exec ns1 ip link set veth1 up
  for i in $(seq 200); do echo hi > /dev/udp/10.9.9.2/9999; done
---
 drivers/net/veth.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/drivers/net/veth.c b/drivers/net/veth.c
index 71227d0389aa..7d181c5e387b 100644
--- a/drivers/net/veth.c
+++ b/drivers/net/veth.c
@@ -377,6 +377,9 @@ static netdev_tx_t veth_xmit(struct sk_buff *skb, struct net_device *dev)
 
 	skb_tx_timestamp(skb);
 
+	/* tx queue index of this device, meaningless as rx queue of the peer */
+	skb_set_queue_mapping(skb, 0);
+
 	ret = veth_forward_skb(rcv, skb, rq, use_napi);
 	switch (ret) {
 	case NET_RX_SUCCESS: /* same as NETDEV_TX_OK */
@@ -397,6 +400,8 @@ static netdev_tx_t veth_xmit(struct sk_buff *skb, struct net_device *dev)
 		}
 		/* Restore Eth hdr pulled by dev_forward_skb/eth_type_trans */
 		__skb_push(skb, ETH_HLEN);
+		/* qdisc requeues the skb on the txq it points to */
+		skb_set_queue_mapping(skb, rxq);
 		netif_tx_stop_queue(txq);
 		/* Makes sure NAPI peer consumer runs. Consumer is responsible
 		 * for starting txq again, until then ndo_start_xmit (this
-- 
2.43.0


             reply	other threads:[~2026-10-09  4:04 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-09  4:04 Jiayuan Chen [this message]
2026-10-09  6:52 ` [PATCH net-next] veth: clear rx queue hint in veth_xmit Björn Töpel

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261009040412.14571-1-jiayuan.chen@linux.dev \
    --to=jiayuan.chen@linux.dev \
    --cc=andrew+netdev@lunn.ch \
    --cc=daniel@iogearbox.net \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=john.fastabend@gmail.com \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=toshiaki.makita1@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox