From: Mingming Cao <mmc@linux.ibm.com>
To: netdev@vger.kernel.org
Cc: davem@davemloft.net, kuba@kernel.org, edumazet@google.com,
pabeni@redhat.com, andrew+netdev@lunn.ch, nnac123@linux.ibm.com,
maddy@linux.ibm.com, mpe@ellerman.id.au,
linuxppc-dev@lists.ozlabs.org, haren@linux.ibm.com,
ricklind@linux.ibm.com, davemarq@linux.ibm.com,
bjking1@linux.ibm.com, shaik.abdulla1@ibm.com,
Mingming Cao <mmc@linux.ibm.com>
Subject: [PATCH net-next v5 00/15] ibmveth: Add multi-queue RX support
Date: Fri, 14 Aug 2026 00:36:27 -0700 [thread overview]
Message-ID: <20260814073642.24630-1-mmc@linux.ibm.com> (raw)
Hi,
Power11 PHYP adds Virtual Ethernet multi-queue (MQ) RX: multiple
logical-LAN RX queues, per-queue buffer posting, and completion
delivery. Guest Linux did not use that; ibmveth still registered one
RX queue even when PHYP was MQ-capable.
This series adds the ibmveth MQ client for net-next. When PHYP
advertises IBMVETH_ILLAN_RX_MULTI_QUEUE_SUPPORT via H_ILLAN_ATTRIBUTES,
probe enables MQ with a default RX count of min(num_online_cpus(), 8)
(same cap as TX today); ethtool -L can raise RX up to 16. Packets are
received on per-queue NAPI. Older firmware without the bit is unchanged.
Queue selection remains firmware-defined (PHYP hash). Ethtool RSS hash
get/set for that algorithm is deferred to a follow-up series so this
one stays MQ datapath only.
User-visible bits: ethtool -l/-L (channels), ethtool -S and
ndo_get_stats64 (per-queue + aggregate), and a read-only debugfs
buffer_pools dump (v3's multi-line sysfs dump moved to debugfs; the
historical queue-0 poolN/ sysfs ABI is unchanged).
Background:
ibmveth today uses one logical LAN, one set of buffer pools, and one
NAPI context. PHYP MQ mode gives each RX queue its own handle (post via
H_ADD_LOGICAL_LAN_BUFFERS_QUEUE, subordinate register via
H_REG_LOGICAL_LAN_QUEUE); traffic can land on any active queue. The
driver needs per-queue pools, IRQs, and NAPI to match. Legacy firmware
keeps the original hcall path.
Series layout (15 patches):
1-2 Hypercall wrappers; MQ adapter layout (MAX_RX_QUEUES stays 1)
3-9 Queue-aware helpers (still SQ runtime): RX, per-queue pools,
IRQ, TX, PHYP, buffer submit (open/close 3-8); poll harden (9)
10 Enable MQ datapath at probe/open (subordinate register helpers
land here with first use)
11-13 Per-queue RX/TX stats; get_channels MQ counts; debugfs buffer_pools
14-15 Incremental RX resize (+ resize_rx_channels); wire ethtool
set_channels
- Helper patches (3-8) reshape ibmveth_open()/close() into
queue-aware helpers. Patch 9 hardens the SQ poll path with the same
queue-index helpers; it does not change open/close. Runtime
behaviour is unchanged through 3-9: num_rx_queues stays 1 and
multi_queue is false until patch 10.
- Patch 10 is the switch: probe sets multi_queue from firmware, raises
num_rx_queues, registers subordinates, and replenishes every active
queue.
- Statistics types land with first use (hcall_stats in 7, rx/tx qstats
in 11).
Testing:
ppc64le PowerVM LPAR, MQ-capable firmware:
* ethtool -L cycling (16/1/8/11/1/3/16/8/1) with ping - no hangs
* ethtool -L under iperf3; link down/up during traffic
* ifdown/ifup under iperf3 RX+TX (MQ and ethtool -L rx 1)
* Legacy firmware (no MQ bit): open/close/stress on helper path
* Bisect-friendly restack; W=1 clean at tip
Changes in v5:
* Restack mailed v4 (14 patches) to v5 (15):
v4 1-8 helpers -> v5 1-8
(new) SQ poll harden -> v5 9 (before MQ enable)
v4 9 MQ enable -> v5 10
v4 10 stats -> v5 11
(new) get_channels -> v5 12 (peeled from stats)
v4 11 debugfs -> v5 13
v4 12 resize -> v5 14
v4 13 set_channels -> v5 15
v4 14 trailing poll/shutdown -> folded into v5 5/9/10/14
(mailed "P14" was that trailer, not v5 14)
* Teardown-first resize after aggressive ethtool -L; thin defensive
poll skip remains; no correlator generation field this series
* opened / rx_irq_setup; set_channels keys on opened (not IFF_UP)
* filter_list_dma=0 on map error; restore default-active 64 KiB pool;
unwind pools by allocation presence; probe_cleanup clears vio
drvdata; remove: unregister then cancel_work
* TX quiesce before freeing bounce buffers; guard start_xmit if LTB gone
* MQ H_FUNCTION recovery (reset + SQ fallback); no printk under
replenish_lock; lock harvest with replenish; resume kicks all queues
* Per-queue update_rx_no_buffer; publish-before-free on resize;
CMO refresh; IRQ helpers return errno
* Harvest abort (no fake GRO / UAF); poll refuses PHYP re-arm on close;
wrap-safe skb_put; atomic set_channels; monotonic stats across shrink
* Keep mask -> sync -> napi_disable on teardown; open stays
request_irq -> napi_enable while PHYP masked; scale-up/recovery keep
napi_enable before enable_irq
* Pool geometry kept on free; restart_rx_queue after open/scale-down;
remask after napi_disable; schedule_rx_queue masks only when
napi_schedule_prep succeeds (STOP + poll no-rearm for storms)
Changes in v4:
Addresses Simon's v3 review and related fixes:
* First-use helpers/includes (irqdomain.h with first dispose); no
unused statics; dropped orphan open/close pipeline patch
* Open/close unwind (free LAN before RX pools); no double TX teardown
* MQ open: replenish all queues before PHYP unmask; H_FUNCTION on
subordinate register is a hard open failure
* Resize/set_channels hardenings; stats probe-lifetime + sum-on-read;
buffer_pools diagnostic on debugfs
* Patch 9: put already-created pool kobjects on probe failure paths
* Patch 14: correlator skip, skb tailroom check, napi_complete_done
shutdown return < budget
* Bisect-friendly restack (helpers with first use)
Changes in v3:
* Dropped RFC; addressed style / DMA feedback from earlier revisions
* Early MQ enablement iterations (see lore links below)
Comments welcome.
---
Previous versions
v4: https://lore.kernel.org/r/cover.1785457143.git.mmc@linux.ibm.com
v3: https://lore.kernel.org/r/20260706193603.8039-1-mmc@linux.ibm.com
v2: https://lore.kernel.org/r/20260701222327.61325-1-mmc@linux.ibm.com
v1: https://lore.kernel.org/r/cover.1782758799.git.mmc@linux.ibm.com
v4 review (Jakub Kicinski):
https://lore.kernel.org/r/20260806183614.3171785-1-kuba@kernel.org
v3 review (Simon Horman):
https://lore.kernel.org/r/20260714124327.GJ1364329@horms.kernel.org
Mingming Cao (15):
ibmveth: Add MQ RX hypercall wrappers and call definitions
ibmveth: Prepare MQ RX adapter data structures
ibmveth: Refactor RX resource allocation for MQ RX bring-up
ibmveth: Refactor buffer pool management for per-queue MQ RX
ibmveth: Refactor RX interrupt control for MQ RX queues
ibmveth: Refactor TX resource allocation in open/close paths
ibmveth: Add RX queue register helpers for MQ
ibmveth: Add queue-aware RX buffer submit helper for MQ
ibmveth: Harden RX poll path with helpers
ibmveth: Enable multi-queue RX receive path
ibmveth: Add per-queue RX and TX statistics collection
ibmveth: Report MQ-aware RX counts in ethtool get_channels
ibmveth: Expose per-queue buffer pool details via debugfs
ibmveth: Implement incremental MQ RX queue resize
ibmveth: Wire ethtool set_channels to MQ RX queue resize
arch/powerpc/include/asm/hvcall.h | 6 +-
drivers/net/ethernet/ibm/ibmveth.c | 3950 ++++++++++++++++++++++------
drivers/net/ethernet/ibm/ibmveth.h | 218 +-
3 files changed, 3432 insertions(+), 742 deletions(-)
--
2.50.1 (Apple Git-155)
next reply other threads:[~2026-08-14 7:38 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-14 7:36 Mingming Cao [this message]
2026-08-14 7:36 ` [PATCH net-next v5 01/15] ibmveth: Add MQ RX hypercall wrappers and call definitions Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 02/15] ibmveth: Prepare MQ RX adapter data structures Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 03/15] ibmveth: Refactor RX resource allocation for MQ RX bring-up Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 04/15] ibmveth: Refactor buffer pool management for per-queue MQ RX Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 05/15] ibmveth: Refactor RX interrupt control for MQ RX queues Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 06/15] ibmveth: Refactor TX resource allocation in open/close paths Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 07/15] ibmveth: Add RX queue register helpers for MQ Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 08/15] ibmveth: Add queue-aware RX buffer submit helper " Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 09/15] ibmveth: Harden RX poll path with helpers Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 10/15] ibmveth: Enable multi-queue RX receive path Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 11/15] ibmveth: Add per-queue RX and TX statistics collection Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 12/15] ibmveth: Report MQ-aware RX counts in ethtool get_channels Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 13/15] ibmveth: Expose per-queue buffer pool details via debugfs Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 14/15] ibmveth: Implement incremental MQ RX queue resize Mingming Cao
2026-08-14 7:36 ` [PATCH net-next v5 15/15] ibmveth: Wire ethtool set_channels to " Mingming Cao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260814073642.24630-1-mmc@linux.ibm.com \
--to=mmc@linux.ibm.com \
--cc=andrew+netdev@lunn.ch \
--cc=bjking1@linux.ibm.com \
--cc=davem@davemloft.net \
--cc=davemarq@linux.ibm.com \
--cc=edumazet@google.com \
--cc=haren@linux.ibm.com \
--cc=kuba@kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=maddy@linux.ibm.com \
--cc=mpe@ellerman.id.au \
--cc=netdev@vger.kernel.org \
--cc=nnac123@linux.ibm.com \
--cc=pabeni@redhat.com \
--cc=ricklind@linux.ibm.com \
--cc=shaik.abdulla1@ibm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.