* [PATCH net-next v7 01/15] ibmveth: Add MQ RX hypercall wrappers and call definitions
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-25 18:38 ` [PATCH net-next v7 02/15] ibmveth: Prepare MQ RX adapter data structures Mingming Cao
` (14 subsequent siblings)
15 siblings, 0 replies; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
[-- Warning: decoded text below may be mangled, UTF-8 assumed --]
[-- Attachment #1: Type: text/plain; charset=true, Size: 10639 bytes --]
Single-queue ibmveth only needs h_register_logical_lan() plus legacy
buffer add/free calls. MQ RX uses per-queue handles, so the driver must
also be able to register/deregister subordinate queues and post
buffers against a specific queue handle.
Add the PHYP call IDs for:
H_REG_LOGICAL_LAN_QUEUE (0x49C)
H_ADD_LOGICAL_LAN_BUFFERS_QUEUE (0x4A0)
H_FREE_LOGICAL_LAN_QUEUE (0x4A8)
as defined in PAPR 11.20.00. The gaps at 0x498 (reserved) and 0x4A4
(reserved for H_FREE_LOGICAL_LAN_BUFFER_QUEUE) are intentional; this
series tears queues down with H_FREE_LOGICAL_LAN_QUEUE and does not add
the buffer-queue free hcall.
Raising MAX_HCALL_OPCODE for these IDs also widens KVM's
kvm_arch.enabled_hcalls bitmap and the range KVM_CAP_PPC_ENABLE_HCALL
accepts. Neither is user-visible: KVM implements none of the three, so
the ioctl still rejects them, and the bitmap rounds to the same five
unsigned longs, so struct kvm_arch does not change size.
Add ibmveth.h wrapper helpers (h_register_logical_lan_queue(),
h_add_logical_lan_buffers_queue(), h_free_logical_lan_queue()) with
argument ordering and return semantics matching the existing ibmveth
hcall wrappers. h_free_logical_lan_queue() uses plpar_hcall_norets()
like h_free_logical_lan(). Also add h_register_logical_lan_with_handle()
so queue 0 can capture the PHYP queue handle in MQ mode. Both new
registration wrappers use plpar_hcall() rather than plpar_hcall9(), so
they do not read unwritten stack slots.
This patch is intentionally plumbing only: no runtime behavior change
yet. Legacy firmware keeps H_REGISTER_LOGICAL_LAN and the existing
buffer hcalls. The new wrappers are used only when a later commit sets
multi_queue from H_ILLAN_ATTRIBUTES.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- Drop the duplicated "Both new registration wrappers use
plpar_hcall()…" sentence
- Register and free-queue Return: also H_FUNCTION. Add-buffers
already had it in v6; all three PAPR 11.20.00 opcodes return that
status on older firmware
- h_register_logical_lan_queue Return: also
H_BUSY / H_LONG_BUSY_*; the caller retries those,
and free-queue already documented them
- tools/perf powerpc-hcalls.py: name 0x49C / 0x4A0 / 0x4A8.
0x498 and 0x4A4 stay unnamed (reserved holes)
Changes in v6:
- both new registration wrappers use plpar_hcall() instead of
plpar_hcall9(). They do not need nine args, and plpar_hcall9() was
reading three unwritten stack slots
- Widen add-buffers-queue and free-queue @queue_handle kdoc: handle may
come from h_register_logical_lan_queue() or
h_register_logical_lan_with_handle() (queue 0)
- add-buffers-queue Return: also H_FUNCTION (MQ firmware missing that
hcall)
Changes in v5:
- Cite PAPR 11.20.00 for MQ hcall IDs 0x49C / 0x4A0 / 0x4A8 (v4 named
the opcodes only)
- Call out reserved gaps at 0x498 and 0x4A4 (0x4A4 =
H_FREE_LOGICAL_LAN_BUFFER_QUEUE); decision: do not add that hcall -
queue teardown uses H_FREE_LOGICAL_LAN_QUEUE
- Note MAX_HCALL_OPCODE raise also widens KVM enabled_hcalls /
KVM_CAP_PPC_ENABLE_HCALL accepted range (cross-subsystem)
- Rename h_reg_logical_lan_queue() -> h_register_logical_lan_queue() to
match h_register_logical_lan() / h_free_logical_lan()
- Unify queue_handle out-params as unsigned long * on both
h_register_logical_lan_queue() and h_register_logical_lan_with_handle()
(v4 used u64 * on with_handle)
- h_free_logical_lan_queue() uses plpar_hcall_norets() like
h_free_logical_lan() (v4 used plpar_hcall9 with unused retbuf)
Changes in v4:
- Document @queue_handle and @irq in h_reg_logical_lan_queue() kdoc.
- Wrap h_register_logical_lan_with_handle() prototype for readability /
checkpatch.
- Drop unused H_FREE_LOGICAL_LAN_BUFFER_QUEUE wrapper/opcode (no caller;
buffer return is local harvest + queue free).
arch/powerpc/include/asm/hvcall.h | 6 +-
drivers/net/ethernet/ibm/ibmveth.h | 142 ++++++++++++++++++++
tools/perf/scripts/python/powerpc-hcalls.py | 4 +
3 files changed, 151 insertions(+), 1 deletion(-)
diff --git a/arch/powerpc/include/asm/hvcall.h b/arch/powerpc/include/asm/hvcall.h
index dff90a7d7f70..cb0ea53491e6 100644
--- a/arch/powerpc/include/asm/hvcall.h
+++ b/arch/powerpc/include/asm/hvcall.h
@@ -362,7 +362,11 @@
#define H_GUEST_DELETE 0x488
#define H_PKS_WRAP_OBJECT 0x490
#define H_PKS_UNWRAP_OBJECT 0x494
-#define MAX_HCALL_OPCODE H_PKS_UNWRAP_OBJECT
+/* 0x498 reserved; 0x4A4 = H_FREE_LOGICAL_LAN_BUFFER_QUEUE (unused here) */
+#define H_REG_LOGICAL_LAN_QUEUE 0x49C
+#define H_ADD_LOGICAL_LAN_BUFFERS_QUEUE 0x4A0
+#define H_FREE_LOGICAL_LAN_QUEUE 0x4A8
+#define MAX_HCALL_OPCODE H_FREE_LOGICAL_LAN_QUEUE
/* Scope args for H_SCM_UNBIND_ALL */
#define H_UNBIND_SCOPE_ALL (0x1)
diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h
index d87713668ed3..5c0b9f61eec9 100644
--- a/drivers/net/ethernet/ibm/ibmveth.h
+++ b/drivers/net/ethernet/ibm/ibmveth.h
@@ -66,6 +66,148 @@ static inline long h_add_logical_lan_buffers(unsigned long unit_address,
desc5, desc6, desc7, desc8);
}
+/**
+ * h_register_logical_lan_queue - Register a subordinate receive queue
+ * @unit_address: Device unit address
+ * @buffer_list: DMA address of 4KB page for tracking registered buffers
+ * @rec_queue: Buffer descriptor of receive queue
+ * @queue_handle: Output queue handle on success (may be NULL)
+ * @irq: Output hypervisor IRQ number on success (may be NULL)
+ *
+ * Registers a subordinate receive queue with the hypervisor.
+ *
+ * Return:
+ * H_SUCCESS (0) on success
+ * H_PARAMETER if parameters are invalid
+ * H_BUSY / H_LONG_BUSY_* if the resource is busy; retry
+ * H_FUNCTION if firmware does not support this hcall
+ *
+ * On success, hypervisor returns:
+ * R3: H_SUCCESS
+ * R4: Queue handle
+ * R5: IRQ number for this queue
+ */
+static inline long
+h_register_logical_lan_queue(unsigned long unit_address,
+ unsigned long buffer_list,
+ unsigned long rec_queue,
+ unsigned long *queue_handle,
+ unsigned long *irq)
+{
+ unsigned long retbuf[PLPAR_HCALL_BUFSIZE];
+ long rc;
+
+ rc = plpar_hcall(H_REG_LOGICAL_LAN_QUEUE,
+ retbuf, unit_address,
+ buffer_list, rec_queue);
+
+ if (rc == H_SUCCESS) {
+ if (queue_handle)
+ *queue_handle = retbuf[0];
+ if (irq)
+ *irq = retbuf[1];
+ }
+
+ return rc;
+}
+
+/**
+ * h_add_logical_lan_buffers_queue - Add buffers to subordinate queue
+ * @unit_address: Device unit address
+ * @queue_handle: Queue handle from h_register_logical_lan_queue() or
+ * h_register_logical_lan_with_handle() (queue 0)
+ * @buffersznum: Buffer size (upper 32 bits) | count (lower 32 bits)
+ * @ioba12: Buffer addresses 1 and 2 packed ((addr1 << 32) | addr2)
+ * @ioba34: Buffer addresses 3 and 4 packed
+ * @ioba56: Buffer addresses 5 and 6 packed
+ * @ioba78: Buffer addresses 7 and 8 packed
+ * @ioba910: Buffer addresses 9 and 10 packed
+ * @ioba1112: Buffer addresses 11 and 12 packed
+ *
+ * Return:
+ * H_SUCCESS - All buffers added successfully
+ * H_PARAMETER - Invalid parameters
+ * H_HARDWARE - Hardware error
+ * H_FUNCTION - Firmware does not support this hcall
+ */
+static inline long h_add_logical_lan_buffers_queue(unsigned long unit_address,
+ unsigned long queue_handle,
+ unsigned long buffersznum,
+ unsigned long ioba12,
+ unsigned long ioba34,
+ unsigned long ioba56,
+ unsigned long ioba78,
+ unsigned long ioba910,
+ unsigned long ioba1112)
+{
+ unsigned long retbuf[PLPAR_HCALL9_BUFSIZE];
+
+ return plpar_hcall9(H_ADD_LOGICAL_LAN_BUFFERS_QUEUE,
+ retbuf, unit_address,
+ queue_handle, buffersznum,
+ ioba12, ioba34, ioba56,
+ ioba78, ioba910, ioba1112);
+}
+
+/**
+ * h_free_logical_lan_queue - Deregister subordinate receive queue
+ * @unit_address: Device unit address
+ * @queue_handle: Queue handle from h_register_logical_lan_queue() or
+ * h_register_logical_lan_with_handle() (queue 0)
+ *
+ * Deregisters and frees all structures associated with the subordinate queue.
+ *
+ * Return:
+ * H_SUCCESS - Queue freed successfully
+ * H_PARAMETER - Invalid parameters
+ * H_HARDWARE - Hardware error
+ * H_STATE - VIOA not in valid state
+ * H_BUSY / H_LONG_BUSY_* - Resource busy, retry
+ * H_FUNCTION - Firmware does not support this hcall
+ */
+static inline long h_free_logical_lan_queue(unsigned long unit_address,
+ unsigned long queue_handle)
+{
+ return plpar_hcall_norets(H_FREE_LOGICAL_LAN_QUEUE,
+ unit_address, queue_handle);
+}
+
+/**
+ * h_register_logical_lan_with_handle - Register primary queue and get handle
+ * @unit_address: Device unit address
+ * @buffer_list: DMA address of buffer list
+ * @rec_queue: Buffer descriptor of receive queue
+ * @filter_list: DMA address of filter list
+ * @mac_address: MAC address
+ * @queue_handle: Output parameter for queue handle (may be NULL)
+ *
+ * Registers the primary receive queue (queue 0) with the hypervisor and
+ * returns the queue handle. This is needed in multi-queue mode to use
+ * h_add_logical_lan_buffers_queue() for all queues including queue 0.
+ *
+ * Return: H_SUCCESS (0) on success, error code otherwise
+ */
+static inline long
+h_register_logical_lan_with_handle(unsigned long unit_address,
+ unsigned long buffer_list,
+ unsigned long rec_queue,
+ unsigned long filter_list,
+ unsigned long mac_address,
+ unsigned long *queue_handle)
+{
+ unsigned long retbuf[PLPAR_HCALL_BUFSIZE];
+ long rc;
+
+ rc = plpar_hcall(H_REGISTER_LOGICAL_LAN, retbuf,
+ unit_address, buffer_list, rec_queue,
+ filter_list, mac_address);
+
+ if (rc == H_SUCCESS && queue_handle)
+ *queue_handle = retbuf[0];
+
+ return rc;
+}
+
/* FW allows us to send 6 descriptors but we only use one so mark
* the other 5 as unused (0)
*/
diff --git a/tools/perf/scripts/python/powerpc-hcalls.py b/tools/perf/scripts/python/powerpc-hcalls.py
index fedce1b68cad..afc354583b7b 100644
--- a/tools/perf/scripts/python/powerpc-hcalls.py
+++ b/tools/perf/scripts/python/powerpc-hcalls.py
@@ -217,6 +217,10 @@ hcall_table = {
# Key wrapping hcalls
1168: 'H_PKS_WRAP_OBJECT',
1172: 'H_PKS_UNWRAP_OBJECT',
+ # Logical LAN multi-queue RX (PAPR 11.20.00)
+ 1180: 'H_REG_LOGICAL_LAN_QUEUE',
+ 1184: 'H_ADD_LOGICAL_LAN_BUFFERS_QUEUE',
+ 1192: 'H_FREE_LOGICAL_LAN_QUEUE',
# Platform-specific hcalls used by the Ultravisor
61184: 'H_SVM_PAGE_IN',
61188: 'H_SVM_PAGE_OUT',
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* [PATCH net-next v7 02/15] ibmveth: Prepare MQ RX adapter data structures
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
2026-09-25 18:38 ` [PATCH net-next v7 01/15] ibmveth: Add MQ RX hypercall wrappers and call definitions Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-25 18:38 ` [PATCH net-next v7 03/15] ibmveth: Refactor RX resource allocation for MQ RX bring-up Mingming Cao
` (13 subsequent siblings)
15 siblings, 0 replies; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
MQ RX needs per-queue state for NAPI, queue handles/IRQs, RX rings,
buffer-list DMA mappings, and buffer pools. The current driver stores
most of this as single instances tied to queue 0.
Convert those fields to queue-indexed layouts sized by
IBMVETH_MAX_RX_QUEUES:
rx_queue[]
napi[]
queue_handle[] / queue_irq[]
buffer_list_addr[] / buffer_list_dma[]
rx_buff_pool[queue][pool]
and add multi_queue / num_rx_queues to track MQ capability and how
many RX queues are active. Keep IBMVETH_MAX_RX_QUEUES at 1 for now so
this remains a structural preparation patch; later enablement raises
the limit when multi-queue RX is actually turned on.
This patch keeps behavior unchanged by mechanically switching existing
references to index 0 (for example rx_queue -> rx_queue[0],
rx_buff_pool[pool] -> rx_buff_pool[0][pool], napi -> napi[0]).
open/poll/close still drive a single RX queue only.
First use of the new fields: queue_irq[] and multi_queue in the IRQ
control patch; queue_handle[] in the register-helpers patch, which
captures queue 0's handle from H_REGISTER_LOGICAL_LAN; num_rx_queues
in the RX resource-allocation patch. multi_queue only begins
selecting between code paths once enablement raises
IBMVETH_MAX_RX_QUEUES above 1.
Probe kobject / drvdata cleanup on register failure is pre-existing.
Inline kobject_put lands in the MQ enablement patch;
ibmveth_probe_cleanup() in the statistics patch.
Queue-0 pool sysfs (poolN/) remains the shared geometry template for
all RX queues: later patches clone that metadata per queue. Per-queue
runtime visibility is added via debugfs later, not per-queue sysfs.
Per-queue statistics structs are introduced later with their first use
(statistics collection); the replenish_* counters stay adapter-wide
plain u64 here and move into them at that point.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- no P02 code change
Changes in v5:
- Drop unrelated style nits that v4 mixed into the index-0 conversion
(ibmveth_rxq_get_buffer() prototype reflow, get_desired_dma
"Return:" kdoc, blank line after rx_large_packets)
- Document queue-0 poolN/ sysfs as the shared geometry/template for all
RX queues (buff_size/size/active); queues 1..N get no separate pool
sysfs in this series
- Decision: keep full ibmveth_buff_pool per queue for this series
(unused embedded kobjects on rows 1..N); a config-vs-runtime layout
split is out of scope here
Changes in v4:
- Keep IBMVETH_MAX_RX_QUEUES at 1 until MQ enablement (same idea as v3,
but v3 also planted unused stats types here).
- Layout-only: queue-indexed adapter fields only. Do not introduce
the rx/tx qstats here (first-use); they land with the per-queue
statistics patch.
- Subject: "Prepare MQ RX adapter data structures" (was "...and
statistics structures" in earlier drafts).
drivers/net/ethernet/ibm/ibmveth.c | 206 ++++++++++++++++-------------
drivers/net/ethernet/ibm/ibmveth.h | 17 ++-
2 files changed, 124 insertions(+), 99 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 73e051d26b9d..7cb828b476c1 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -101,7 +101,9 @@ static struct ibmveth_stat ibmveth_stats[] = {
/* simple methods of getting data from the current rxq entry */
static inline u32 ibmveth_rxq_flags(struct ibmveth_adapter *adapter)
{
- return be32_to_cpu(adapter->rx_queue.queue_addr[adapter->rx_queue.index].flags_off);
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[0];
+
+ return be32_to_cpu(rxq->queue_addr[rxq->index].flags_off);
}
static inline int ibmveth_rxq_toggle(struct ibmveth_adapter *adapter)
@@ -112,7 +114,7 @@ static inline int ibmveth_rxq_toggle(struct ibmveth_adapter *adapter)
static inline int ibmveth_rxq_pending_buffer(struct ibmveth_adapter *adapter)
{
- return ibmveth_rxq_toggle(adapter) == adapter->rx_queue.toggle;
+ return ibmveth_rxq_toggle(adapter) == adapter->rx_queue[0].toggle;
}
static inline int ibmveth_rxq_buffer_valid(struct ibmveth_adapter *adapter)
@@ -132,7 +134,9 @@ static inline int ibmveth_rxq_large_packet(struct ibmveth_adapter *adapter)
static inline int ibmveth_rxq_frame_length(struct ibmveth_adapter *adapter)
{
- return be32_to_cpu(adapter->rx_queue.queue_addr[adapter->rx_queue.index].length);
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[0];
+
+ return be32_to_cpu(rxq->queue_addr[rxq->index].length);
}
static inline int ibmveth_rxq_csum_good(struct ibmveth_adapter *adapter)
@@ -386,7 +390,7 @@ static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
*/
static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter)
{
- __be64 *p = adapter->buffer_list_addr + 4096 - 8;
+ __be64 *p = adapter->buffer_list_addr[0] + 4096 - 8;
adapter->rx_no_buffer = be64_to_cpup(p);
}
@@ -399,7 +403,7 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter)
adapter->replenish_task_cycles++;
for (i = (IBMVETH_NUM_BUFF_POOLS - 1); i >= 0; i--) {
- struct ibmveth_buff_pool *pool = &adapter->rx_buff_pool[i];
+ struct ibmveth_buff_pool *pool = &adapter->rx_buff_pool[0][i];
if (pool->active &&
(atomic_read(&pool->available) < pool->threshold))
@@ -463,12 +467,12 @@ static int ibmveth_remove_buffer_from_pool(struct ibmveth_adapter *adapter,
struct sk_buff *skb;
if (WARN_ON(pool >= IBMVETH_NUM_BUFF_POOLS) ||
- WARN_ON(index >= adapter->rx_buff_pool[pool].size)) {
+ WARN_ON(index >= adapter->rx_buff_pool[0][pool].size)) {
schedule_work(&adapter->work);
return -EINVAL;
}
- skb = adapter->rx_buff_pool[pool].skbuff[index];
+ skb = adapter->rx_buff_pool[0][pool].skbuff[index];
if (WARN_ON(!skb)) {
schedule_work(&adapter->work);
return -EFAULT;
@@ -482,24 +486,24 @@ static int ibmveth_remove_buffer_from_pool(struct ibmveth_adapter *adapter,
/* remove the skb pointer to mark free. actual freeing is done
* by upper level networking after gro_receive
*/
- adapter->rx_buff_pool[pool].skbuff[index] = NULL;
+ adapter->rx_buff_pool[0][pool].skbuff[index] = NULL;
dma_unmap_single(&adapter->vdev->dev,
- adapter->rx_buff_pool[pool].dma_addr[index],
- adapter->rx_buff_pool[pool].buff_size,
+ adapter->rx_buff_pool[0][pool].dma_addr[index],
+ adapter->rx_buff_pool[0][pool].buff_size,
DMA_FROM_DEVICE);
}
- free_index = adapter->rx_buff_pool[pool].producer_index;
- adapter->rx_buff_pool[pool].producer_index++;
- if (adapter->rx_buff_pool[pool].producer_index >=
- adapter->rx_buff_pool[pool].size)
- adapter->rx_buff_pool[pool].producer_index = 0;
- adapter->rx_buff_pool[pool].free_map[free_index] = index;
+ free_index = adapter->rx_buff_pool[0][pool].producer_index;
+ adapter->rx_buff_pool[0][pool].producer_index++;
+ if (adapter->rx_buff_pool[0][pool].producer_index >=
+ adapter->rx_buff_pool[0][pool].size)
+ adapter->rx_buff_pool[0][pool].producer_index = 0;
+ adapter->rx_buff_pool[0][pool].free_map[free_index] = index;
mb();
- atomic_dec(&(adapter->rx_buff_pool[pool].available));
+ atomic_dec(&adapter->rx_buff_pool[0][pool].available);
return 0;
}
@@ -507,17 +511,18 @@ static int ibmveth_remove_buffer_from_pool(struct ibmveth_adapter *adapter,
/* get the current buffer on the rx queue */
static inline struct sk_buff *ibmveth_rxq_get_buffer(struct ibmveth_adapter *adapter)
{
- u64 correlator = adapter->rx_queue.queue_addr[adapter->rx_queue.index].correlator;
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[0];
+ u64 correlator = rxq->queue_addr[rxq->index].correlator;
unsigned int pool = correlator >> 32;
unsigned int index = correlator & 0xffffffffUL;
if (WARN_ON(pool >= IBMVETH_NUM_BUFF_POOLS) ||
- WARN_ON(index >= adapter->rx_buff_pool[pool].size)) {
+ WARN_ON(index >= adapter->rx_buff_pool[0][pool].size)) {
schedule_work(&adapter->work);
return NULL;
}
- return adapter->rx_buff_pool[pool].skbuff[index];
+ return adapter->rx_buff_pool[0][pool].skbuff[index];
}
/**
@@ -538,14 +543,16 @@ static int ibmveth_rxq_harvest_buffer(struct ibmveth_adapter *adapter,
u64 cor;
int rc;
- cor = adapter->rx_queue.queue_addr[adapter->rx_queue.index].correlator;
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[0];
+
+ cor = rxq->queue_addr[rxq->index].correlator;
rc = ibmveth_remove_buffer_from_pool(adapter, cor, reuse);
if (unlikely(rc))
return rc;
- if (++adapter->rx_queue.index == adapter->rx_queue.num_slots) {
- adapter->rx_queue.index = 0;
- adapter->rx_queue.toggle = !adapter->rx_queue.toggle;
+ if (++adapter->rx_queue[0].index == adapter->rx_queue[0].num_slots) {
+ adapter->rx_queue[0].index = 0;
+ adapter->rx_queue[0].toggle = !adapter->rx_queue[0].toggle;
}
return 0;
@@ -595,7 +602,7 @@ static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
*/
retry:
rc = h_register_logical_lan(adapter->vdev->unit_address,
- adapter->buffer_list_dma, rxq_desc.desc,
+ adapter->buffer_list_dma[0], rxq_desc.desc,
adapter->filter_list_dma, mac_address);
if (rc != H_SUCCESS && try_again) {
@@ -623,14 +630,14 @@ static int ibmveth_open(struct net_device *netdev)
netdev_dbg(netdev, "open starting\n");
- napi_enable(&adapter->napi);
+ napi_enable(&adapter->napi[0]);
for(i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
- rxq_entries += adapter->rx_buff_pool[i].size;
+ rxq_entries += adapter->rx_buff_pool[0][i].size;
rc = -ENOMEM;
- adapter->buffer_list_addr = (void*) get_zeroed_page(GFP_KERNEL);
- if (!adapter->buffer_list_addr) {
+ adapter->buffer_list_addr[0] = (void *)get_zeroed_page(GFP_KERNEL);
+ if (!adapter->buffer_list_addr[0]) {
netdev_err(netdev, "unable to allocate list pages\n");
goto out;
}
@@ -643,17 +650,18 @@ static int ibmveth_open(struct net_device *netdev)
dev = &adapter->vdev->dev;
- adapter->rx_queue.queue_len = sizeof(struct ibmveth_rx_q_entry) *
+ adapter->rx_queue[0].queue_len = sizeof(struct ibmveth_rx_q_entry) *
rxq_entries;
- adapter->rx_queue.queue_addr =
- dma_alloc_coherent(dev, adapter->rx_queue.queue_len,
- &adapter->rx_queue.queue_dma, GFP_KERNEL);
- if (!adapter->rx_queue.queue_addr)
+ adapter->rx_queue[0].queue_addr =
+ dma_alloc_coherent(dev, adapter->rx_queue[0].queue_len,
+ &adapter->rx_queue[0].queue_dma, GFP_KERNEL);
+ if (!adapter->rx_queue[0].queue_addr)
goto out_free_filter_list;
- adapter->buffer_list_dma = dma_map_single(dev,
- adapter->buffer_list_addr, 4096, DMA_BIDIRECTIONAL);
- if (dma_mapping_error(dev, adapter->buffer_list_dma)) {
+ adapter->buffer_list_dma[0] =
+ dma_map_single(dev, adapter->buffer_list_addr[0],
+ 4096, DMA_BIDIRECTIONAL);
+ if (dma_mapping_error(dev, adapter->buffer_list_dma[0])) {
netdev_err(netdev, "unable to map buffer list pages\n");
goto out_free_queue_mem;
}
@@ -670,19 +678,21 @@ static int ibmveth_open(struct net_device *netdev)
goto out_free_tx_ltb;
}
- adapter->rx_queue.index = 0;
- adapter->rx_queue.num_slots = rxq_entries;
- adapter->rx_queue.toggle = 1;
+ adapter->rx_queue[0].index = 0;
+ adapter->rx_queue[0].num_slots = rxq_entries;
+ adapter->rx_queue[0].toggle = 1;
mac_address = ether_addr_to_u64(netdev->dev_addr);
rxq_desc.fields.flags_len = IBMVETH_BUF_VALID |
- adapter->rx_queue.queue_len;
- rxq_desc.fields.address = adapter->rx_queue.queue_dma;
+ adapter->rx_queue[0].queue_len;
+ rxq_desc.fields.address = adapter->rx_queue[0].queue_dma;
- netdev_dbg(netdev, "buffer list @ 0x%p\n", adapter->buffer_list_addr);
+ netdev_dbg(netdev, "buffer list @ 0x%p\n",
+ adapter->buffer_list_addr[0]);
netdev_dbg(netdev, "filter list @ 0x%p\n", adapter->filter_list_addr);
- netdev_dbg(netdev, "receive q @ 0x%p\n", adapter->rx_queue.queue_addr);
+ netdev_dbg(netdev, "receive q @ 0x%p\n",
+ adapter->rx_queue[0].queue_addr);
h_vio_signal(adapter->vdev->unit_address, VIO_IRQ_DISABLE);
@@ -693,7 +703,7 @@ static int ibmveth_open(struct net_device *netdev)
lpar_rc);
netdev_err(netdev, "buffer TCE:0x%llx filter TCE:0x%llx rxq "
"desc:0x%llx MAC:0x%llx\n",
- adapter->buffer_list_dma,
+ adapter->buffer_list_dma[0],
adapter->filter_list_dma,
rxq_desc.desc,
mac_address);
@@ -702,11 +712,11 @@ static int ibmveth_open(struct net_device *netdev)
}
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
- if (!adapter->rx_buff_pool[i].active)
+ if (!adapter->rx_buff_pool[0][i].active)
continue;
- if (ibmveth_alloc_buffer_pool(&adapter->rx_buff_pool[i])) {
+ if (ibmveth_alloc_buffer_pool(&adapter->rx_buff_pool[0][i])) {
netdev_err(netdev, "unable to alloc pool\n");
- adapter->rx_buff_pool[i].active = 0;
+ adapter->rx_buff_pool[0][i].active = 0;
rc = -ENOMEM;
goto out_free_buffer_pools;
}
@@ -738,9 +748,9 @@ static int ibmveth_open(struct net_device *netdev)
out_free_buffer_pools:
while (--i >= 0) {
- if (adapter->rx_buff_pool[i].active)
+ if (adapter->rx_buff_pool[0][i].active)
ibmveth_free_buffer_pool(adapter,
- &adapter->rx_buff_pool[i]);
+ &adapter->rx_buff_pool[0][i]);
}
out_unmap_filter_list:
dma_unmap_single(dev, adapter->filter_list_dma, 4096,
@@ -752,18 +762,18 @@ static int ibmveth_open(struct net_device *netdev)
}
out_unmap_buffer_list:
- dma_unmap_single(dev, adapter->buffer_list_dma, 4096,
+ dma_unmap_single(dev, adapter->buffer_list_dma[0], 4096,
DMA_BIDIRECTIONAL);
out_free_queue_mem:
- dma_free_coherent(dev, adapter->rx_queue.queue_len,
- adapter->rx_queue.queue_addr,
- adapter->rx_queue.queue_dma);
+ dma_free_coherent(dev, adapter->rx_queue[0].queue_len,
+ adapter->rx_queue[0].queue_addr,
+ adapter->rx_queue[0].queue_dma);
out_free_filter_list:
free_page((unsigned long)adapter->filter_list_addr);
out_free_buffer_list:
- free_page((unsigned long)adapter->buffer_list_addr);
+ free_page((unsigned long)adapter->buffer_list_addr[0]);
out:
- napi_disable(&adapter->napi);
+ napi_disable(&adapter->napi[0]);
return rc;
}
@@ -776,7 +786,7 @@ static int ibmveth_close(struct net_device *netdev)
netdev_dbg(netdev, "close starting\n");
- napi_disable(&adapter->napi);
+ napi_disable(&adapter->napi[0]);
netif_tx_stop_all_queues(netdev);
@@ -795,22 +805,22 @@ static int ibmveth_close(struct net_device *netdev)
ibmveth_update_rx_no_buffer(adapter);
- dma_unmap_single(dev, adapter->buffer_list_dma, 4096,
+ dma_unmap_single(dev, adapter->buffer_list_dma[0], 4096,
DMA_BIDIRECTIONAL);
- free_page((unsigned long)adapter->buffer_list_addr);
+ free_page((unsigned long)adapter->buffer_list_addr[0]);
dma_unmap_single(dev, adapter->filter_list_dma, 4096,
DMA_BIDIRECTIONAL);
free_page((unsigned long)adapter->filter_list_addr);
- dma_free_coherent(dev, adapter->rx_queue.queue_len,
- adapter->rx_queue.queue_addr,
- adapter->rx_queue.queue_dma);
+ dma_free_coherent(dev, adapter->rx_queue[0].queue_len,
+ adapter->rx_queue[0].queue_addr,
+ adapter->rx_queue[0].queue_dma);
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
- if (adapter->rx_buff_pool[i].active)
+ if (adapter->rx_buff_pool[0][i].active)
ibmveth_free_buffer_pool(adapter,
- &adapter->rx_buff_pool[i]);
+ &adapter->rx_buff_pool[0][i]);
for (i = 0; i < netdev->real_num_tx_queues; i++)
ibmveth_free_tx_ltb(adapter, i);
@@ -1448,7 +1458,7 @@ static void ibmveth_rx_csum_helper(struct sk_buff *skb,
static int ibmveth_poll(struct napi_struct *napi, int budget)
{
struct ibmveth_adapter *adapter =
- container_of(napi, struct ibmveth_adapter, napi);
+ container_of(napi, struct ibmveth_adapter, napi[0]);
struct net_device *netdev = adapter->netdev;
int frames_processed = 0;
unsigned long lpar_rc;
@@ -1573,11 +1583,11 @@ static irqreturn_t ibmveth_interrupt(int irq, void *dev_instance)
struct ibmveth_adapter *adapter = netdev_priv(netdev);
unsigned long lpar_rc;
- if (napi_schedule_prep(&adapter->napi)) {
+ if (napi_schedule_prep(&adapter->napi[0])) {
lpar_rc = h_vio_signal(adapter->vdev->unit_address,
VIO_IRQ_DISABLE);
WARN_ON(lpar_rc != H_SUCCESS);
- __napi_schedule(&adapter->napi);
+ __napi_schedule(&adapter->napi[0]);
}
return IRQ_HANDLED;
}
@@ -1645,7 +1655,7 @@ static int ibmveth_change_mtu(struct net_device *dev, int new_mtu)
int need_restart = 0;
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
- if (new_mtu_oh <= adapter->rx_buff_pool[i].buff_size)
+ if (new_mtu_oh <= adapter->rx_buff_pool[0][i].buff_size)
break;
if (i == IBMVETH_NUM_BUFF_POOLS)
@@ -1660,9 +1670,9 @@ static int ibmveth_change_mtu(struct net_device *dev, int new_mtu)
/* Look for an active buffer pool that can hold the new MTU */
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
- adapter->rx_buff_pool[i].active = 1;
+ adapter->rx_buff_pool[0][i].active = 1;
- if (new_mtu_oh <= adapter->rx_buff_pool[i].buff_size) {
+ if (new_mtu_oh <= adapter->rx_buff_pool[0][i].buff_size) {
WRITE_ONCE(dev->mtu, new_mtu);
vio_cmo_set_dev_desired(viodev,
ibmveth_get_desired_dma
@@ -1720,12 +1730,12 @@ static unsigned long ibmveth_get_desired_dma(struct vio_dev *vdev)
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
/* add the size of the active receive buffers */
- if (adapter->rx_buff_pool[i].active)
+ if (adapter->rx_buff_pool[0][i].active)
ret +=
- adapter->rx_buff_pool[i].size *
- IOMMU_PAGE_ALIGN(adapter->rx_buff_pool[i].
+ adapter->rx_buff_pool[0][i].size *
+ IOMMU_PAGE_ALIGN(adapter->rx_buff_pool[0][i].
buff_size, tbl);
- rxqentries += adapter->rx_buff_pool[i].size;
+ rxqentries += adapter->rx_buff_pool[0][i].size;
}
/* add the size of the receive queue entries */
ret += IOMMU_PAGE_ALIGN(
@@ -1844,7 +1854,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
adapter->mcastFilterSize = be32_to_cpu(*mcastFilterSize_p);
ibmveth_init_link_settings(netdev);
- netif_napi_add_weight(netdev, &adapter->napi, ibmveth_poll, 16);
+ netif_napi_add_weight(netdev, &adapter->napi[0], ibmveth_poll, 16);
netdev->irq = dev->irq;
netdev->netdev_ops = &ibmveth_netdev_ops;
@@ -1876,6 +1886,10 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
netdev->features |= NETIF_F_FRAGLIST;
}
+ /* Initialize queue count - always 1 for now */
+ adapter->multi_queue = 0;
+ adapter->num_rx_queues = IBMVETH_DEFAULT_RX_QUEUES;
+
if (ret == H_SUCCESS &&
(ret_attr & IBMVETH_ILLAN_RX_MULTI_BUFF_SUPPORT)) {
adapter->rx_buffers_per_hcall = IBMVETH_MAX_RX_PER_HCALL;
@@ -1898,10 +1912,10 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
memcpy(pool_count, pool_count_cmo, sizeof(pool_count));
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
- struct kobject *kobj = &adapter->rx_buff_pool[i].kobj;
+ struct kobject *kobj = &adapter->rx_buff_pool[0][i].kobj;
int error;
- ibmveth_init_buffer_pool(&adapter->rx_buff_pool[i], i,
+ ibmveth_init_buffer_pool(&adapter->rx_buff_pool[0][i], i,
pool_count[i], pool_size[i],
pool_active[i]);
error = kobject_init_and_add(kobj, &ktype_veth_pool,
@@ -1949,7 +1963,7 @@ static void ibmveth_remove(struct vio_dev *dev)
cancel_work_sync(&adapter->work);
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
- kobject_put(&adapter->rx_buff_pool[i].kobj);
+ kobject_put(&adapter->rx_buff_pool[0][i].kobj);
unregister_netdev(netdev);
@@ -2035,11 +2049,12 @@ static ssize_t veth_pool_store(struct kobject *kobj, struct attribute *attr,
/* Make sure there is a buffer pool with buffers that
can hold a packet of the size of the MTU */
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
- if (pool == &adapter->rx_buff_pool[i])
+ if (pool == &adapter->rx_buff_pool[0][i])
continue;
- if (!adapter->rx_buff_pool[i].active)
+ if (!adapter->rx_buff_pool[0][i].active)
continue;
- if (mtu <= adapter->rx_buff_pool[i].buff_size)
+ if (mtu <=
+ adapter->rx_buff_pool[0][i].buff_size)
break;
}
@@ -2213,11 +2228,11 @@ static void ibmveth_remove_buffer_from_pool_test(struct kunit *test)
/* Set sane values for buffer pools */
for (int i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
- ibmveth_init_buffer_pool(&adapter->rx_buff_pool[i], i,
+ ibmveth_init_buffer_pool(&adapter->rx_buff_pool[0][i], i,
pool_count[i], pool_size[i],
pool_active[i]);
- pool = &adapter->rx_buff_pool[0];
+ pool = &adapter->rx_buff_pool[0][0];
pool->skbuff = kunit_kcalloc(test, pool->size, sizeof(void *), GFP_KERNEL);
KUNIT_ASSERT_NOT_ERR_OR_NULL(test, pool->skbuff);
@@ -2225,7 +2240,7 @@ static void ibmveth_remove_buffer_from_pool_test(struct kunit *test)
KUNIT_EXPECT_EQ(test, -EINVAL, ibmveth_remove_buffer_from_pool(adapter, correlator, false));
KUNIT_EXPECT_EQ(test, -EINVAL, ibmveth_remove_buffer_from_pool(adapter, correlator, true));
- correlator = ((u64)0 << 32) | adapter->rx_buff_pool[0].size;
+ correlator = ((u64)0 << 32) | adapter->rx_buff_pool[0][0].size;
KUNIT_EXPECT_EQ(test, -EINVAL, ibmveth_remove_buffer_from_pool(adapter, correlator, false));
KUNIT_EXPECT_EQ(test, -EINVAL, ibmveth_remove_buffer_from_pool(adapter, correlator, true));
@@ -2258,30 +2273,33 @@ static void ibmveth_rxq_get_buffer_test(struct kunit *test)
INIT_WORK(&adapter->work, ibmveth_reset_kunit);
- adapter->rx_queue.queue_len = 1;
- adapter->rx_queue.index = 0;
- adapter->rx_queue.queue_addr = kunit_kzalloc(test, sizeof(struct ibmveth_rx_q_entry),
- GFP_KERNEL);
- KUNIT_ASSERT_NOT_ERR_OR_NULL(test, adapter->rx_queue.queue_addr);
+ adapter->rx_queue[0].queue_len = 1;
+ adapter->rx_queue[0].index = 0;
+ adapter->rx_queue[0].queue_addr =
+ kunit_kzalloc(test, sizeof(struct ibmveth_rx_q_entry),
+ GFP_KERNEL);
+ KUNIT_ASSERT_NOT_ERR_OR_NULL(test, adapter->rx_queue[0].queue_addr);
/* Set sane values for buffer pools */
for (int i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
- ibmveth_init_buffer_pool(&adapter->rx_buff_pool[i], i,
+ ibmveth_init_buffer_pool(&adapter->rx_buff_pool[0][i], i,
pool_count[i], pool_size[i],
pool_active[i]);
- pool = &adapter->rx_buff_pool[0];
+ pool = &adapter->rx_buff_pool[0][0];
pool->skbuff = kunit_kcalloc(test, pool->size, sizeof(void *), GFP_KERNEL);
KUNIT_ASSERT_NOT_ERR_OR_NULL(test, pool->skbuff);
- adapter->rx_queue.queue_addr[0].correlator = (u64)IBMVETH_NUM_BUFF_POOLS << 32 | 0;
+ adapter->rx_queue[0].queue_addr[0].correlator =
+ (u64)IBMVETH_NUM_BUFF_POOLS << 32 | 0;
KUNIT_EXPECT_PTR_EQ(test, NULL, ibmveth_rxq_get_buffer(adapter));
- adapter->rx_queue.queue_addr[0].correlator = (u64)0 << 32 | adapter->rx_buff_pool[0].size;
+ adapter->rx_queue[0].queue_addr[0].correlator =
+ (u64)0 << 32 | adapter->rx_buff_pool[0][0].size;
KUNIT_EXPECT_PTR_EQ(test, NULL, ibmveth_rxq_get_buffer(adapter));
pool->skbuff[0] = skb;
- adapter->rx_queue.queue_addr[0].correlator = (u64)0 << 32 | 0;
+ adapter->rx_queue[0].queue_addr[0].correlator = (u64)0 << 32 | 0;
KUNIT_EXPECT_PTR_EQ(test, skb, ibmveth_rxq_get_buffer(adapter));
flush_work(&adapter->work);
diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h
index 5c0b9f61eec9..7e956278e005 100644
--- a/drivers/net/ethernet/ibm/ibmveth.h
+++ b/drivers/net/ethernet/ibm/ibmveth.h
@@ -263,6 +263,8 @@ static inline long h_illan_attributes(unsigned long unit_address,
#define IBMVETH_MAX_TX_BUF_SIZE (1024 * 64)
#define IBMVETH_MAX_QUEUES 16U
#define IBMVETH_DEFAULT_QUEUES 8U
+#define IBMVETH_MAX_RX_QUEUES 1U
+#define IBMVETH_DEFAULT_RX_QUEUES 1U
#define IBMVETH_MAX_RX_PER_HCALL 8U
static int pool_size[] = { 512, 1024 * 2, 1024 * 16, 1024 * 32, 1024 * 64 };
@@ -299,18 +301,23 @@ struct ibmveth_rx_q {
struct ibmveth_adapter {
struct vio_dev *vdev;
struct net_device *netdev;
- struct napi_struct napi;
+ struct napi_struct napi[IBMVETH_MAX_RX_QUEUES];
struct work_struct work;
unsigned int mcastFilterSize;
- void *buffer_list_addr;
+ void *buffer_list_addr[IBMVETH_MAX_RX_QUEUES];
void *filter_list_addr;
void *tx_ltb_ptr[IBMVETH_MAX_QUEUES];
unsigned int tx_ltb_size;
dma_addr_t tx_ltb_dma[IBMVETH_MAX_QUEUES];
- dma_addr_t buffer_list_dma;
+ dma_addr_t buffer_list_dma[IBMVETH_MAX_RX_QUEUES];
dma_addr_t filter_list_dma;
- struct ibmveth_buff_pool rx_buff_pool[IBMVETH_NUM_BUFF_POOLS];
- struct ibmveth_rx_q rx_queue;
+ struct ibmveth_buff_pool
+ rx_buff_pool[IBMVETH_MAX_RX_QUEUES][IBMVETH_NUM_BUFF_POOLS];
+ struct ibmveth_rx_q rx_queue[IBMVETH_MAX_RX_QUEUES];
+ u64 queue_handle[IBMVETH_MAX_RX_QUEUES];
+ unsigned int queue_irq[IBMVETH_MAX_RX_QUEUES];
+ int multi_queue;
+ unsigned int num_rx_queues;
int rx_csum;
int large_send;
bool is_active_trunk;
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* [PATCH net-next v7 03/15] ibmveth: Refactor RX resource allocation for MQ RX bring-up
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
2026-09-25 18:38 ` [PATCH net-next v7 01/15] ibmveth: Add MQ RX hypercall wrappers and call definitions Mingming Cao
2026-09-25 18:38 ` [PATCH net-next v7 02/15] ibmveth: Prepare MQ RX adapter data structures Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 04/15] ibmveth: Refactor buffer pool management for per-queue MQ RX Mingming Cao
` (12 subsequent siblings)
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
ibmveth_open() allocates the filter list and every RX queue inline.
That is already a long sequence and would get uglier once we loop over
num_rx_queues, especially on error unwind.
Pull the RX bits into helpers and wire them into open()/close() in the
same patch:
ibmveth_alloc_filter_list() / ibmveth_free_filter_list()
- shared multicast filter list (one per adapter, not per queue)
ibmveth_alloc_rx_queues() / ibmveth_cleanup_rx_resources()
- per-queue buffer lists and RX rings, looping [0, num_rx_queues)
alloc_rx_queues() rolls back on failure so open() does not need nested
goto chains for every queue index. open-failure and close release the
same resources through the same helpers.
Cleanup NULLs each slot as it frees. A filter-list map error zeros
filter_list_dma so a later free path cannot unmap DMA_MAPPING_ERROR,
and unmap is gated on the CPU page because SPAPR can return DMA
address 0.
Runtime behavior stays single-queue (num_rx_queues is still 1). Buffer
pools, IRQ, TX LTB, and PHYP registration remain inline for later
helper patches.
Also set rc = -ENOMEM before the TX LTB allocation loop so a failed
ibmveth_allocate_tx_ltb() still returns a useful errno after the RX
allocation blocks move into helpers.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- NULL-check buffer_list_addr[0] in update_rx_no_buffer()
- reword err_cleanup comment for the mixed map-failure state
- drop the failed-reopen then ndo_stop double-free claim from the body
- noted: no Fixes: peel of the DMA-handle checks; [PATCH net 0/2]
<cover.1790357373.git.mmc@linux.ibm.com>
- noted: TX LTB stale handle predates this; P06 (pointer check, zero dma)
- noted: shared-i TX LTB leak predates this; P06 (own counter, TX
after pools)
- noted: pool-fail without h_free predates this; P06 adds it, P07
allocates pools first
- noted: napi imbalance on failed reopen predates this; P05
- noted: replenish_lock stays in P08; poll_controller !opened
stays in P15; leftover netpoll/irqsave after this series
Changes in v6:
- on RX-queue dma_alloc or map failure, free the buffer-list page
before cleanup so the addr-based unmap does not dma_unmap an
address that was never mapped
- unmap RX filter/buffer lists by CPU pointer, not dma_addr==0
- use %u for the two num_rx_queues dbg prints
- noted: NULL buffer_list_addr[0] window: close() in P05, netpoll
through replenish_task() until the per-queue guard in P10
- noted: shared-i TX LTB leak on pool-fail predates this; P04/P06
- noted: pool-fail after register without h_free predates this; P06/P07
Changes in v5:
- On filter_list DMA map failure: free_page and zero filter_list_dma so
a later free path cannot dma_unmap the DMA_MAPPING_ERROR sentinel
(match buffer_list_dma convention). Closes reopen->ifdown WARN path
after set_csum/set_tso/change_mtu/pool_store close+open while running
- v4 left the sentinel
- Call out rc = -ENOMEM before the TX LTB loop after RX helper extract
(already in v4; still required so allocate_tx_ltb failure returns a
useful errno)
Changes in v4:
- Introduce RX/filter allocation helpers in the same patch that wires
their first open/close callers; v3 left unused statics ahead of the
old open/close pipeline patch.
- Preserve correct -ENOMEM return on TX LTB allocation failure after
the RX helper extract.
- Drop reliance on v3's separate "open/close pipeline" patch for this
wiring.
drivers/net/ethernet/ibm/ibmveth.c | 284 +++++++++++++++++++++--------
1 file changed, 205 insertions(+), 79 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 7cb828b476c1..01efd318baab 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -151,6 +151,194 @@ static unsigned int ibmveth_real_max_tx_queues(void)
return min(n_cpu, IBMVETH_MAX_QUEUES);
}
+/**
+ * ibmveth_alloc_filter_list - Allocate and map filter list
+ * @adapter: ibmveth adapter structure
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_alloc_filter_list(struct ibmveth_adapter *adapter)
+{
+ struct device *dev = &adapter->vdev->dev;
+ struct net_device *netdev = adapter->netdev;
+
+ adapter->filter_list_addr = (void *)get_zeroed_page(GFP_KERNEL);
+ if (!adapter->filter_list_addr) {
+ netdev_err(netdev, "unable to allocate filter pages\n");
+ return -ENOMEM;
+ }
+
+ adapter->filter_list_dma = dma_map_single(dev,
+ adapter->filter_list_addr,
+ 4096, DMA_BIDIRECTIONAL);
+ if (dma_mapping_error(dev, adapter->filter_list_dma)) {
+ netdev_err(netdev, "unable to map filter list pages\n");
+ free_page((unsigned long)adapter->filter_list_addr);
+ adapter->filter_list_addr = NULL;
+ /* Do not leave DMA_MAPPING_ERROR for free_filter_list(). */
+ adapter->filter_list_dma = 0;
+ return -ENOMEM;
+ }
+
+ netdev_dbg(netdev, "filter list @ 0x%p (DMA: 0x%llx)\n",
+ adapter->filter_list_addr,
+ (unsigned long long)adapter->filter_list_dma);
+
+ return 0;
+}
+
+/**
+ * ibmveth_free_filter_list - Free filter list resources
+ * @adapter: ibmveth adapter structure
+ */
+static void
+ibmveth_free_filter_list(struct ibmveth_adapter *adapter)
+{
+ struct device *dev = &adapter->vdev->dev;
+
+ /* Unmap by CPU pointer: SPAPR can return DMA address 0. */
+ if (adapter->filter_list_addr) {
+ dma_unmap_single(dev, adapter->filter_list_dma, 4096,
+ DMA_BIDIRECTIONAL);
+ adapter->filter_list_dma = 0;
+ free_page((unsigned long)adapter->filter_list_addr);
+ adapter->filter_list_addr = NULL;
+ }
+}
+
+/**
+ * ibmveth_alloc_rx_queues - Allocate per-queue RX resources
+ * @adapter: ibmveth adapter structure
+ * @rxq_entries: Number of entries per RX queue
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_alloc_rx_queues(struct ibmveth_adapter *adapter, int rxq_entries)
+{
+ struct device *dev = &adapter->vdev->dev;
+ struct net_device *netdev = adapter->netdev;
+ int i;
+
+ for (i = 0; i < adapter->num_rx_queues; i++) {
+ adapter->buffer_list_addr[i] =
+ (void *)get_zeroed_page(GFP_KERNEL);
+ if (!adapter->buffer_list_addr[i]) {
+ netdev_err(netdev,
+ "unable to allocate buffer list for queue %d\n",
+ i);
+ goto err_cleanup;
+ }
+
+ adapter->rx_queue[i].queue_len =
+ sizeof(struct ibmveth_rx_q_entry) * rxq_entries;
+ adapter->rx_queue[i].queue_addr =
+ dma_alloc_coherent(dev, adapter->rx_queue[i].queue_len,
+ &adapter->rx_queue[i].queue_dma,
+ GFP_KERNEL);
+ if (!adapter->rx_queue[i].queue_addr) {
+ netdev_err(netdev,
+ "unable to allocate RX queue for queue %d\n",
+ i);
+ free_page((unsigned long)adapter->buffer_list_addr[i]);
+ adapter->buffer_list_addr[i] = NULL;
+ goto err_cleanup;
+ }
+
+ adapter->buffer_list_dma[i] =
+ dma_map_single(dev, adapter->buffer_list_addr[i],
+ 4096, DMA_BIDIRECTIONAL);
+ if (dma_mapping_error(dev, adapter->buffer_list_dma[i])) {
+ netdev_err(netdev,
+ "unable to map buffer list for queue %d\n",
+ i);
+ free_page((unsigned long)adapter->buffer_list_addr[i]);
+ adapter->buffer_list_addr[i] = NULL;
+ adapter->buffer_list_dma[i] = 0;
+ goto err_cleanup;
+ }
+
+ adapter->rx_queue[i].index = 0;
+ adapter->rx_queue[i].num_slots = rxq_entries;
+ adapter->rx_queue[i].toggle = 1;
+
+ netdev_dbg(netdev, "queue %d: buffer_list @ 0x%p (DMA: 0x%llx), rx_queue @ 0x%p (DMA: 0x%llx), %llu entries\n",
+ i, adapter->buffer_list_addr[i],
+ (unsigned long long)adapter->buffer_list_dma[i],
+ adapter->rx_queue[i].queue_addr,
+ (unsigned long long)adapter->rx_queue[i].queue_dma,
+ (unsigned long long)rxq_entries);
+ }
+
+ netdev_dbg(netdev, "allocated %u RX queue(s) with %d entries each\n",
+ adapter->num_rx_queues, rxq_entries);
+
+ return 0;
+
+err_cleanup:
+ /*
+ * Free each leftover resource by pointer presence. A
+ * dma_mapping_error on queue i has already dropped
+ * buffer_list_addr[i] but can leave rx_queue[i].queue_addr
+ * live, so this index may be mixed. Do not collapse the
+ * buffer_list and queue_addr checks.
+ */
+ for (; i >= 0; i--) {
+ if (adapter->buffer_list_addr[i]) {
+ dma_unmap_single(dev, adapter->buffer_list_dma[i],
+ 4096, DMA_BIDIRECTIONAL);
+ adapter->buffer_list_dma[i] = 0;
+ }
+ if (adapter->rx_queue[i].queue_addr) {
+ dma_free_coherent(dev, adapter->rx_queue[i].queue_len,
+ adapter->rx_queue[i].queue_addr,
+ adapter->rx_queue[i].queue_dma);
+ adapter->rx_queue[i].queue_addr = NULL;
+ }
+ if (adapter->buffer_list_addr[i]) {
+ free_page((unsigned long)adapter->buffer_list_addr[i]);
+ adapter->buffer_list_addr[i] = NULL;
+ }
+ }
+
+ return -ENOMEM;
+}
+
+/**
+ * ibmveth_cleanup_rx_resources - Free all RX queue resources
+ * @adapter: ibmveth adapter structure
+ */
+static void
+ibmveth_cleanup_rx_resources(struct ibmveth_adapter *adapter)
+{
+ struct device *dev = &adapter->vdev->dev;
+ int i;
+
+ netdev_dbg(adapter->netdev, "cleaning up %u RX queue(s)\n",
+ adapter->num_rx_queues);
+
+ for (i = 0; i < adapter->num_rx_queues; i++) {
+ if (adapter->buffer_list_addr[i]) {
+ dma_unmap_single(dev, adapter->buffer_list_dma[i],
+ 4096, DMA_BIDIRECTIONAL);
+ adapter->buffer_list_dma[i] = 0;
+ }
+
+ if (adapter->rx_queue[i].queue_addr) {
+ dma_free_coherent(dev, adapter->rx_queue[i].queue_len,
+ adapter->rx_queue[i].queue_addr,
+ adapter->rx_queue[i].queue_dma);
+ adapter->rx_queue[i].queue_addr = NULL;
+ }
+
+ if (adapter->buffer_list_addr[i]) {
+ free_page((unsigned long)adapter->buffer_list_addr[i]);
+ adapter->buffer_list_addr[i] = NULL;
+ }
+ }
+}
+
/* setup the initial settings for a buffer pool */
static void ibmveth_init_buffer_pool(struct ibmveth_buff_pool *pool,
u32 pool_index, u32 pool_size,
@@ -390,8 +578,12 @@ static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
*/
static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter)
{
- __be64 *p = adapter->buffer_list_addr[0] + 4096 - 8;
+ __be64 *p;
+ if (!adapter->buffer_list_addr[0])
+ return;
+
+ p = adapter->buffer_list_addr[0] + 4096 - 8;
adapter->rx_no_buffer = be64_to_cpup(p);
}
@@ -626,74 +818,34 @@ static int ibmveth_open(struct net_device *netdev)
int rc;
union ibmveth_buf_desc rxq_desc;
int i;
- struct device *dev;
netdev_dbg(netdev, "open starting\n");
napi_enable(&adapter->napi[0]);
- for(i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
+ for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
rxq_entries += adapter->rx_buff_pool[0][i].size;
- rc = -ENOMEM;
- adapter->buffer_list_addr[0] = (void *)get_zeroed_page(GFP_KERNEL);
- if (!adapter->buffer_list_addr[0]) {
- netdev_err(netdev, "unable to allocate list pages\n");
+ rc = ibmveth_alloc_filter_list(adapter);
+ if (rc)
goto out;
- }
- adapter->filter_list_addr = (void*) get_zeroed_page(GFP_KERNEL);
- if (!adapter->filter_list_addr) {
- netdev_err(netdev, "unable to allocate filter pages\n");
- goto out_free_buffer_list;
- }
-
- dev = &adapter->vdev->dev;
-
- adapter->rx_queue[0].queue_len = sizeof(struct ibmveth_rx_q_entry) *
- rxq_entries;
- adapter->rx_queue[0].queue_addr =
- dma_alloc_coherent(dev, adapter->rx_queue[0].queue_len,
- &adapter->rx_queue[0].queue_dma, GFP_KERNEL);
- if (!adapter->rx_queue[0].queue_addr)
+ rc = ibmveth_alloc_rx_queues(adapter, rxq_entries);
+ if (rc)
goto out_free_filter_list;
- adapter->buffer_list_dma[0] =
- dma_map_single(dev, adapter->buffer_list_addr[0],
- 4096, DMA_BIDIRECTIONAL);
- if (dma_mapping_error(dev, adapter->buffer_list_dma[0])) {
- netdev_err(netdev, "unable to map buffer list pages\n");
- goto out_free_queue_mem;
- }
-
- adapter->filter_list_dma = dma_map_single(dev,
- adapter->filter_list_addr, 4096, DMA_BIDIRECTIONAL);
- if (dma_mapping_error(dev, adapter->filter_list_dma)) {
- netdev_err(netdev, "unable to map filter list pages\n");
- goto out_unmap_buffer_list;
- }
-
+ rc = -ENOMEM;
for (i = 0; i < netdev->real_num_tx_queues; i++) {
if (ibmveth_allocate_tx_ltb(adapter, i))
goto out_free_tx_ltb;
}
- adapter->rx_queue[0].index = 0;
- adapter->rx_queue[0].num_slots = rxq_entries;
- adapter->rx_queue[0].toggle = 1;
-
mac_address = ether_addr_to_u64(netdev->dev_addr);
rxq_desc.fields.flags_len = IBMVETH_BUF_VALID |
adapter->rx_queue[0].queue_len;
rxq_desc.fields.address = adapter->rx_queue[0].queue_dma;
- netdev_dbg(netdev, "buffer list @ 0x%p\n",
- adapter->buffer_list_addr[0]);
- netdev_dbg(netdev, "filter list @ 0x%p\n", adapter->filter_list_addr);
- netdev_dbg(netdev, "receive q @ 0x%p\n",
- adapter->rx_queue[0].queue_addr);
-
h_vio_signal(adapter->vdev->unit_address, VIO_IRQ_DISABLE);
lpar_rc = ibmveth_register_logical_lan(adapter, rxq_desc, mac_address);
@@ -708,7 +860,7 @@ static int ibmveth_open(struct net_device *netdev)
rxq_desc.desc,
mac_address);
rc = -ENONET;
- goto out_unmap_filter_list;
+ goto out_free_tx_ltb;
}
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
@@ -735,8 +887,6 @@ static int ibmveth_open(struct net_device *netdev)
goto out_free_buffer_pools;
}
- rc = -ENOMEM;
-
netdev_dbg(netdev, "initial replenish cycle\n");
ibmveth_interrupt(netdev->irq, netdev);
@@ -752,26 +902,12 @@ static int ibmveth_open(struct net_device *netdev)
ibmveth_free_buffer_pool(adapter,
&adapter->rx_buff_pool[0][i]);
}
-out_unmap_filter_list:
- dma_unmap_single(dev, adapter->filter_list_dma, 4096,
- DMA_BIDIRECTIONAL);
-
out_free_tx_ltb:
- while (--i >= 0) {
+ while (--i >= 0)
ibmveth_free_tx_ltb(adapter, i);
- }
-
-out_unmap_buffer_list:
- dma_unmap_single(dev, adapter->buffer_list_dma[0], 4096,
- DMA_BIDIRECTIONAL);
-out_free_queue_mem:
- dma_free_coherent(dev, adapter->rx_queue[0].queue_len,
- adapter->rx_queue[0].queue_addr,
- adapter->rx_queue[0].queue_dma);
+ ibmveth_cleanup_rx_resources(adapter);
out_free_filter_list:
- free_page((unsigned long)adapter->filter_list_addr);
-out_free_buffer_list:
- free_page((unsigned long)adapter->buffer_list_addr[0]);
+ ibmveth_free_filter_list(adapter);
out:
napi_disable(&adapter->napi[0]);
return rc;
@@ -780,7 +916,6 @@ static int ibmveth_open(struct net_device *netdev)
static int ibmveth_close(struct net_device *netdev)
{
struct ibmveth_adapter *adapter = netdev_priv(netdev);
- struct device *dev = &adapter->vdev->dev;
long lpar_rc;
int i;
@@ -805,17 +940,8 @@ static int ibmveth_close(struct net_device *netdev)
ibmveth_update_rx_no_buffer(adapter);
- dma_unmap_single(dev, adapter->buffer_list_dma[0], 4096,
- DMA_BIDIRECTIONAL);
- free_page((unsigned long)adapter->buffer_list_addr[0]);
-
- dma_unmap_single(dev, adapter->filter_list_dma, 4096,
- DMA_BIDIRECTIONAL);
- free_page((unsigned long)adapter->filter_list_addr);
-
- dma_free_coherent(dev, adapter->rx_queue[0].queue_len,
- adapter->rx_queue[0].queue_addr,
- adapter->rx_queue[0].queue_dma);
+ ibmveth_cleanup_rx_resources(adapter);
+ ibmveth_free_filter_list(adapter);
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
if (adapter->rx_buff_pool[0][i].active)
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 03/15] ibmveth: Refactor RX resource allocation for MQ RX bring-up
2026-09-25 18:38 ` [PATCH net-next v7 03/15] ibmveth: Refactor RX resource allocation for MQ RX bring-up Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] This patch silently fixes an existing DMA mapping leak in
ibmveth_open().
Pre-existing issues:
- [High] If ibmveth_alloc_buffer_pool() fails after
ibmveth_register_logical_lan() has succeeded, ibmveth_open() jumps to…
- [Medium] In ibmveth_open(), the buffer-pool unwind and the TX LTB
unwind share the loop variable `i`.
- [Medium] ibmveth_update_rx_no_buffer() checks
adapter->buffer_list_addr[0] and then loads it again, with no lock and…
--
ibmveth: Refactor RX resource allocation for MQ RX bring-up
This patch moves the filter list and per-queue RX buffer list and RX ring
allocation out of ibmveth_open() into new helpers, and makes both the
open() error path and close() free them through the same cleanup helpers.
It also NULL-checks buffer_list_addr[0] in ibmveth_update_rx_no_buffer()
and sets rc = -ENOMEM before the TX LTB allocation loop.
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index 7cb828b476c1b..01efd318baabb 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
[ ... ]
> @@ -390,8 +578,12 @@ static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
> */
> static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter)
> {
> - __be64 *p = adapter->buffer_list_addr[0] + 4096 - 8;
> + __be64 *p;
>
> + if (!adapter->buffer_list_addr[0])
> + return;
> +
> + p = adapter->buffer_list_addr[0] + 4096 - 8;
> adapter->rx_no_buffer = be64_to_cpup(p);
> }
[Severity: Medium]
This is a pre-existing race and was not introduced by this patch, but the
new check makes it a little more visible.
set_csum, set_tso, change_mtu and pool_store all do a driver-internal
close/open. netif_running() stays true during that, and netpoll is not
disabled. That means this path:
ibmveth_poll_controller()->ibmveth_replenish_task()
->ibmveth_update_rx_no_buffer()
can run at the same time as:
ibmveth_close()->ibmveth_cleanup_rx_resources()
free_page((unsigned long)adapter->buffer_list_addr[i]);
adapter->buffer_list_addr[i] = NULL;
The NULL check and the pointer load are two separate plain reads of
buffer_list_addr[0], with no lock and no READ_ONCE().
Could the reader pass the check, then reload NULL and fault at offset
0xff8? Could it also read the page after free_page() has run, as the
baseline code already could?
As far as I can tell this is dealt with later in the series.
ibmveth_close() clears adapter->opened first, calls synchronize_net(), and
reads the no-buffer count under replenish_lock before freeing anything.
ibmveth_poll_controller() returns early when !opened.
[ ... ]
> @@ -752,26 +902,12 @@ static int ibmveth_open(struct net_device *netdev)
> ibmveth_free_buffer_pool(adapter,
> &adapter->rx_buff_pool[0][i]);
> }
> -out_unmap_filter_list:
> - dma_unmap_single(dev, adapter->filter_list_dma, 4096,
> - DMA_BIDIRECTIONAL);
> -
> out_free_tx_ltb:
> - while (--i >= 0) {
> + while (--i >= 0)
> ibmveth_free_tx_ltb(adapter, i);
[Severity: Medium]
This isn't a bug introduced by this patch, but the buffer pool unwind and
the TX LTB unwind share the loop variable i.
If ibmveth_alloc_buffer_pool() or request_irq() fails, out_free_buffer_pools
runs its own while (--i >= 0) loop, which leaves i == -1. Execution then
falls through to this loop.
Doesn't this loop then run zero times and leak every tx_ltb_ptr[] buffer
and its DMA mapping?
This looks like it gets fixed later in the series.
"ibmveth: Refactor buffer pool management for per-queue MQ RX" moves the
pool loop into a helper, and "ibmveth: Refactor TX resource allocation in
open/close paths" moves the TX allocation into ibmveth_alloc_tx_resources()
at the end of open(). After those patches the two unwinds no longer share
an index.
> - }
> -
> -out_unmap_buffer_list:
> - dma_unmap_single(dev, adapter->buffer_list_dma[0], 4096,
> - DMA_BIDIRECTIONAL);
> -out_free_queue_mem:
> - dma_free_coherent(dev, adapter->rx_queue[0].queue_len,
> - adapter->rx_queue[0].queue_addr,
> - adapter->rx_queue[0].queue_dma);
> + ibmveth_cleanup_rx_resources(adapter);
[Severity: High]
This is a pre-existing issue, but it is on the unwind path this patch
rewires.
The buffer pool failure branch in ibmveth_open() runs after
ibmveth_register_logical_lan() has already succeeded:
if (ibmveth_alloc_buffer_pool(&adapter->rx_buff_pool[0][i])) {
netdev_err(netdev, "unable to alloc pool\n");
adapter->rx_buff_pool[0][i].active = 0;
rc = -ENOMEM;
goto out_free_buffer_pools;
}
The request_irq() failure path calls h_free_logical_lan() before its goto.
This branch does not.
Execution then reaches ibmveth_cleanup_rx_resources() and
ibmveth_free_filter_list(). Those call dma_free_coherent() on the RX queue
and unmap and free the buffer list and filter list pages, while PHYP still
has all of them registered.
Can the hypervisor then write RX descriptors or the no-buffer counter into
freed memory?
The later patches in the series seem to address this.
"ibmveth: Refactor TX resource allocation in open/close paths" adds
h_free_logical_lan() on this path. "ibmveth: Add RX queue register helpers
for MQ" allocates the buffer pools before registration.
> out_free_filter_list:
> - free_page((unsigned long)adapter->filter_list_addr);
> -out_free_buffer_list:
> - free_page((unsigned long)adapter->buffer_list_addr[0]);
> + ibmveth_free_filter_list(adapter);
[Severity: Low]
Does this change behavior on the TX LTB failure path in a way the commit
message doesn't mention?
In the baseline, the filter list was mapped before the TX LTB loop, and a
failed ibmveth_allocate_tx_ltb() jumped to out_free_tx_ltb. That label sits
below out_unmap_filter_list, so the unmap was skipped. The page was then
freed by free_page() in out_free_filter_list while its DMA/TCE mapping was
still live.
After this patch the same failure reaches ibmveth_free_filter_list(), which
unmaps before freeing.
The commit message describes this as a helper extraction:
"Runtime behavior stays single-queue (num_rx_queues is still 1)."
Should the commit message mention this fix? Should it also be split out
with a Fixes: tag, so that stable trees pick up the filter list unmap fix?
> out:
> napi_disable(&adapter->napi[0]);
> return rc;
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 04/15] ibmveth: Refactor buffer pool management for per-queue MQ RX
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (2 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 03/15] ibmveth: Refactor RX resource allocation for MQ RX bring-up Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-25 18:38 ` [PATCH net-next v7 05/15] ibmveth: Refactor RX interrupt control for MQ RX queues Mingming Cao
` (11 subsequent siblings)
15 siblings, 0 replies; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
Legacy ibmveth uses five adapter-level RX buffer pools (512 B through
64 KiB). pool_active[] enables the standard-MTU pools by default;
larger pools activate when MTU requires them. With single-queue RX
that set is shared on one completion path.
MQ requires the same pool model per queue: buffers post with
H_ADD_LOGICAL_LAN_BUFFERS_QUEUE against a queue handle and completions
return on that queue. Sharing pools across queues would mix ownership
and break queue-local replenish/drain/teardown.
Refactor around queue-local pools:
rx_buff_pool[queue][pool]
ibmveth_alloc_queue_buffer_pools()
ibmveth_free_queue_buffer_pools()
ibmveth_alloc_buffer_pools() / ibmveth_free_buffer_pools()
Queue 0 remains the template for pool geometry and activation policy
(size, buff_size, threshold, index, active). Queues 1..N copy that
metadata from queue 0, then allocate backing arrays/skbs per queue.
The existing poolN/ sysfs knobs stay adapter-wide on queue 0: writing
a pool size multiplies real memory by num_rx_queues (still 1 here).
Per-queue runtime state is exposed later via debugfs, not separate
per-queue pool sysfs nodes.
Wire the helpers into open()/close() in the same patch. Runtime
remains single-queue (num_rx_queues is still 1).
Pulling the pool loop out has one side effect worth naming: it no
longer consumes open()'s loop index, so a pool failure or a
request_irq() failure reaches out_free_tx_ltb with i still at
real_num_tx_queues and the TX LTBs actually get freed. The shared
index that swallowed them was pre-existing, so there is no
standalone Fixes: tag; the TX side gets its own unwind two patches
later.
Error handling is queue-safe: allocation failure unwinds only what
that queue allocated (then prior queues in the caller).
alloc_buffer_pool() already undoes its own partials, so this open-fail
slot is empty. Free paths still release by real allocations
(free_map/dma_addr/skbuff), not only pool->active, which later resize
needs when a pool can hold memory after active was cleared.
close() also reorders pool teardown ahead of cleanup_rx_resources()
and filter-list free. That is safe here because h_free_logical_lan(),
napi_disable(), and free_irq() have already run, so neither PHYP nor
NAPI still reference the pool buffers or RX completion queue.
Keep the legacy 64 KiB pool enabled by default at standard MTU (same
as single-queue policy).
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- no P04 code change
- commit message: request_irq() failure also reached
out_free_tx_ltb with i spent
- noted: err_cleanup comment reworded in P03
- noted: no Fixes: peel of the shared-i TX LTB leak (named
in the v6 body after v5); [PATCH net 2/2]
<cover.1790357373.git.mmc@linux.ibm.com>
- noted: no pool-before-register reorder here; P06 issues h_free
on the pool-fail path, P07 allocates pools before register
Changes in v6:
- retarget the free-by-presence comment to the later resize paths;
alloc_buffer_pool() already undoes its own partials, so this
open-fail slot is empty
- use %u for the two num_rx_queues dbg prints
- noted: pool-fail after register without h_free_logical_lan predates
this patch; the inline pool loop reached the same labels. P06
issues h_free before the pool DMA; P07 moves pools ahead of
register_rx_queues
Changes in v5:
- vs mailed v4: keep pool_active[] = {1,1,0,0,1} (v4 set the 64 KiB
pool inactive). No code delta vs prior patches here; SQ large-receive /
CMO stay at the historical baseline. change_mtu does not re-enable
pool4 at MTU 1500; MQ memory pressure belongs at scale-up, not SQ
default
- Unwind / free pools by real allocations (free_map/dma_addr/skbuff),
not only pool->active, so open-fail cannot leak partially allocated
pools (v4 fail path freed by active and skipped the failing pool)
- Clarify kdoc: free paths use allocation presence (not "all active");
queue 1..N metadata copy from queue 0 (v4 said "queues 1-15" while
MAX is still 1 here)
- Document poolN sysfs size multiplies by num_rx_queues (queue 0 is the
shared template; no per-queue pool sysfs in this series)
- Call out close() pool-free before cleanup_rx_resources, after LAN/
NAPI/IRQ teardown (order already in v4; spell it for later drain)
Changes in v4:
- Introduce the pool helpers in the same patch that wires their first
open/close callers, instead of leaving unused statics.
- Copy pool->index when cloning queue-0 geometry to later queues
(needed for correlators; also required by incremental resize).
drivers/net/ethernet/ibm/ibmveth.c | 163 +++++++++++++++++++++++++----
1 file changed, 143 insertions(+), 20 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 01efd318baab..5813352943fb 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -639,6 +639,144 @@ static void ibmveth_free_buffer_pool(struct ibmveth_adapter *adapter,
}
}
+/**
+ * ibmveth_free_queue_buffer_pools - Free buffer pools for a single queue
+ * @adapter: ibmveth adapter structure
+ * @queue: queue index
+ *
+ * Frees buffer pools that still hold allocations for the specified
+ * queue (by free_map / dma_addr / skbuff presence), regardless of the
+ * active flag.
+ */
+static void ibmveth_free_queue_buffer_pools(struct ibmveth_adapter *adapter,
+ int queue)
+{
+ int i;
+
+ for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
+ struct ibmveth_buff_pool *pool =
+ &adapter->rx_buff_pool[queue][i];
+
+ /* Free pool if it has allocated memory, regardless of
+ * active flag. Allocation and active can diverge on failure
+ * paths, so check for actual allocations.
+ */
+ if (pool->free_map || pool->dma_addr || pool->skbuff)
+ ibmveth_free_buffer_pool(adapter, pool);
+ }
+}
+
+/**
+ * ibmveth_alloc_queue_buffer_pools - Allocate buffer pools for a single queue
+ * @adapter: ibmveth adapter structure
+ * @queue: queue index
+ *
+ * Allocates backing storage for each active pool on @queue.
+ * Inactive pools (!active) are skipped. Pool metadata must be
+ * initialized before calling this function.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int ibmveth_alloc_queue_buffer_pools(struct ibmveth_adapter *adapter,
+ int queue)
+{
+ struct net_device *netdev = adapter->netdev;
+ int i;
+
+ for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
+ struct ibmveth_buff_pool *bpool =
+ &adapter->rx_buff_pool[queue][i];
+
+ if (!bpool->active)
+ continue;
+
+ if (ibmveth_alloc_buffer_pool(bpool)) {
+ netdev_err(netdev,
+ "queue %d pool %d alloc failed (size=%u count=%u)\n",
+ queue, i,
+ bpool->buff_size,
+ bpool->size);
+ bpool->active = 0;
+ /* Free by allocation presence, not pool->active.
+ * Needed on later resize paths (a pool can still
+ * hold memory after active is cleared). This
+ * open-fail slot is already empty:
+ * alloc_buffer_pool() undoes its own partials.
+ */
+ ibmveth_free_queue_buffer_pools(adapter, queue);
+ return -ENOMEM;
+ }
+ }
+
+ return 0;
+}
+
+/**
+ * ibmveth_alloc_buffer_pools - Allocate buffer pools for all queues
+ * @adapter: ibmveth adapter structure
+ *
+ * Initializes pool metadata for queues 1-N from queue 0 settings,
+ * then allocates buffer pools for all queues using the helper function.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_alloc_buffer_pools(struct ibmveth_adapter *adapter)
+{
+ struct net_device *netdev = adapter->netdev;
+ int i, q, rc;
+
+ /* Initialize pool metadata for queues 1..N from queue 0 settings */
+ for (q = 1; q < adapter->num_rx_queues; q++) {
+ for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
+ struct ibmveth_buff_pool *src =
+ &adapter->rx_buff_pool[0][i];
+ struct ibmveth_buff_pool *dst =
+ &adapter->rx_buff_pool[q][i];
+
+ dst->size = src->size;
+ dst->index = src->index;
+ dst->buff_size = src->buff_size;
+ dst->threshold = src->threshold;
+ dst->active = src->active;
+ }
+ }
+
+ /* Allocate actual buffers for all queues */
+ for (q = 0; q < adapter->num_rx_queues; q++) {
+ rc = ibmveth_alloc_queue_buffer_pools(adapter, q);
+ if (rc) {
+ /* Free pools for all previous queues */
+ while (--q >= 0)
+ ibmveth_free_queue_buffer_pools(adapter, q);
+ return rc;
+ }
+ }
+
+ netdev_dbg(netdev, "allocated buffer pools for %u queue(s)\n",
+ adapter->num_rx_queues);
+ return 0;
+}
+
+/**
+ * ibmveth_free_buffer_pools - Free buffer pools for all queues
+ * @adapter: ibmveth adapter structure
+ *
+ * Frees buffer pools for all queues using the helper function.
+ */
+static void
+ibmveth_free_buffer_pools(struct ibmveth_adapter *adapter)
+{
+ int q;
+
+ /* Free buffer pools for all queues */
+ for (q = 0; q < adapter->num_rx_queues; q++)
+ ibmveth_free_queue_buffer_pools(adapter, q);
+
+ netdev_dbg(adapter->netdev, "freed buffer pools for %u queue(s)\n",
+ adapter->num_rx_queues);
+}
+
/**
* ibmveth_remove_buffer_from_pool - remove a buffer from a pool
* @adapter: adapter instance
@@ -863,16 +1001,9 @@ static int ibmveth_open(struct net_device *netdev)
goto out_free_tx_ltb;
}
- for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
- if (!adapter->rx_buff_pool[0][i].active)
- continue;
- if (ibmveth_alloc_buffer_pool(&adapter->rx_buff_pool[0][i])) {
- netdev_err(netdev, "unable to alloc pool\n");
- adapter->rx_buff_pool[0][i].active = 0;
- rc = -ENOMEM;
- goto out_free_buffer_pools;
- }
- }
+ rc = ibmveth_alloc_buffer_pools(adapter);
+ if (rc)
+ goto out_free_tx_ltb;
netdev_dbg(netdev, "registering irq 0x%x\n", netdev->irq);
rc = request_irq(netdev->irq, ibmveth_interrupt, 0, netdev->name,
@@ -897,11 +1028,7 @@ static int ibmveth_open(struct net_device *netdev)
return 0;
out_free_buffer_pools:
- while (--i >= 0) {
- if (adapter->rx_buff_pool[0][i].active)
- ibmveth_free_buffer_pool(adapter,
- &adapter->rx_buff_pool[0][i]);
- }
+ ibmveth_free_buffer_pools(adapter);
out_free_tx_ltb:
while (--i >= 0)
ibmveth_free_tx_ltb(adapter, i);
@@ -940,14 +1067,10 @@ static int ibmveth_close(struct net_device *netdev)
ibmveth_update_rx_no_buffer(adapter);
+ ibmveth_free_buffer_pools(adapter);
ibmveth_cleanup_rx_resources(adapter);
ibmveth_free_filter_list(adapter);
- for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
- if (adapter->rx_buff_pool[0][i].active)
- ibmveth_free_buffer_pool(adapter,
- &adapter->rx_buff_pool[0][i]);
-
for (i = 0; i < netdev->real_num_tx_queues; i++)
ibmveth_free_tx_ltb(adapter, i);
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* [PATCH net-next v7 05/15] ibmveth: Refactor RX interrupt control for MQ RX queues
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (3 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 04/15] ibmveth: Refactor buffer pool management for per-queue MQ RX Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 06/15] ibmveth: Refactor TX resource allocation in open/close paths Mingming Cao
` (10 subsequent siblings)
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
Queue 0 and subordinate RX queues use different interrupt control
interfaces in PHYP:
- queue 0: h_vio_signal() after h_register_logical_lan()
- queue N: H_VIOCTL against the queue's mapped hwirq
The current code is single-queue oriented and cannot safely scale to
multiple RX queues in poll completion and open/close IRQ setup.
Introduce queue-indexed interrupt helpers and wire them into
open()/close()/poll()/interrupt in the same patch:
ibmveth_toggle_irq() / enable_irq() / disable_irq()
ibmveth_setup_rx_interrupts() / ibmveth_cleanup_rx_interrupts()
ibmveth_schedule_rx_queue()
These helpers centralize queue0-vs-subordinate dispatch.
request_irq() uses &adapter->napi[i] as the per-queue cookie so the
handler can resolve the queue index.
Move napi_enable() into setup_rx_interrupts() (after LAN registration
and buffer-pool allocation): request_irq -> napi_enable. In this
single-queue tree, setup does not yet unmask PHYP; schedule_rx_queue()
masks queue 0 and schedules NAPI, and ibmveth_poll() is what unmasks
it on completion. That order matches the later scale-up rule (NAPI
live before PHYP unmask), not an inverted window relative to it.
Factor process-context RX kicks (open, resume, pool sysfs, netpoll)
into ibmveth_schedule_rx_queue(); keep ibmveth_interrupt() as a thin
IRQ-only wrapper.
cleanup_rx_interrupts() masks PHYP and synchronizes IRQs before
napi_disable, remasks and synchronizes again after it. The second
remask only catches a re-arm that lands before it; napi_disable
does not wait for poll to return, so enable_irq can still run
after free_irq. Close then proceeds to
h_free_logical_lan(): free_irq before free_lan is intentional once
PHYP delivery is masked.
Close harvests the PHYP no-buffer count before h_free so the
read still hits a live buffer-list page. synchronize_net()
after IRQ/NAPI teardown waits for a poll that already passed
the shutdown checks.
On setup_rx enable-fail (MQ path), if enable_irq() fails for queue i,
remask+sync queues 0..i, including the one that failed, before
napi_disable/free_irq; the rollback loop used while (--i) and skipped
it. H_PARAMETER stays an error on enable, so PHYP may already be
unmasked; an unmasked queue must not drive schedule_rx, which would
prep-fail without mask during the napi_disable wait (STOP storm).
err_disable_napi mirrors cleanup remask after napi_disable.
opened / rx_irq_setup gate whether cleanup walks IRQ/NAPI state.
Opened / rx_irq_setup also closes a pre-existing hang: after a
failed reopen, a later ndo_stop used to napi_disable and free_irq
a second time (rtnl spin + already-free IRQ). That depends on the
helpers in this patch, so there is no standalone Fixes: tag.
schedule_rx_queue() masks PHYP only when napi_schedule_prep() succeeds.
Masking on prep failure can race a completing poll that already
re-enabled PHYP and leave NAPI idle with the queue masked (TX OK, RX
stalled until reload). Teardown storm control stays on STOP
(disable_irq + synchronize_irq before napi_disable) and the
poll_stopping() re-arm guard added in P09, not on the schedule helper
failure path.
IRQ helpers return 0 or negative errno only (never raw H_* to
ethtool/resize). H_PARAMETER is folded to success only on disable
(idempotent mask). On enable it remains an error so a stuck-masked
queue stays visible to poll/resize recovery.
Runtime remains single-queue (num_rx_queues is still 1).
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- reword schedule_rx_queue() kdoc: out-of-range is
WARN_ON, then false
- commit message: remask only catches a re-arm that lands
before it; name harvest before h_free and synchronize_net()
after teardown
- noted: no Fixes: peel of the opened-gate hang; stays in
this patch, not the unwind standalone
Changes in v6:
- setup_rx enable-fail remasks queues 0..i, including the one that
failed (the rollback used while (--i) and left queue i unmasked
across napi_disable); err_disable_napi remasks after napi_disable
- drop WARN_ON() on enable/disable_irq() returns (schedule_rx and
poll); the helper already logs the hcall rc, and rate-limit that
print
- reword schedule_rx_queue() kdoc: true means NAPI was scheduled and
the mask attempted; a failed disable_irq() does not change the
return, and an out-of-range qindex returns false
- noted: pool-fail without h_free predates this; the opened gate
also drops the accidental h_free a later ndo_stop used to give.
P06/P07 close it
- noted: set_channels() IFF_UP vs opened TX LTB window closes in P14/P15
- noted: poll() takes queue_index in P08; range check / skip helpers
in P09
Changes in v5:
- Remask+sync after napi_disable in cleanup (in-flight poll can re-arm)
- Interrupt: quiet IRQ_NONE on out-of-range qindex (no WARN storm)
- Opened / rx_irq_setup gate cleanup so close after a failed open
cannot napi_disable / free_irq without a prior enable/request; set
opened on successful open; rx_irq_setup only on full setup success
- H_PARAMETER fold disable-only; enable stays error (stuck-masked visible)
- IRQ helpers return 0 / negative errno only (never raw H_* to ethtool)
- Decision: keep open IRQ order request_irq -> napi_enable while PHYP
stays masked until schedule/enable (coherent with later scale-up)
- schedule_rx_queue returns bool (napi_schedule_prep success)
- Keep mask-only-on-prep-success (no else-mask; avoids idle+masked race)
- Call out free_irq-before-free_lan as intentional once PHYP is masked
- Poll re-arm during teardown lands with SQ poll-harden (not claimed here)
- synchronize_net() after RX IRQ/NAPI teardown in close
- Add subordinate IRQ dispose helpers (per-queue + bulk 1..N; bound to MAX)
Changes in v4:
- Include irq.h / irqdomain.h with first irq_dispose_mapping() use.
- Introduce IRQ helpers in the same patch that wires open/close/poll
callers, instead of leaving unused statics.
- Factor process-context RX kicks into ibmveth_schedule_rx_queue();
keep ibmveth_interrupt() as the IRQ-only wrapper.
- On cleanup, mask PHYP and synchronize_irq before napi_disable
(storm-safety; not fully behavior-preserving vs classic close).
- Leave queue_irq[0] set after cleanup (queue 0 uses netdev->irq;
next open reuses it). Only subordinate virqs are disposed.
drivers/net/ethernet/ibm/ibmveth.c | 404 ++++++++++++++++++++++++++---
drivers/net/ethernet/ibm/ibmveth.h | 4 +
2 files changed, 367 insertions(+), 41 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 5813352943fb..a22a17e05ae1 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -21,6 +21,8 @@
#include <linux/skbuff.h>
#include <linux/init.h>
#include <linux/interrupt.h>
+#include <linux/irq.h>
+#include <linux/irqdomain.h>
#include <linux/mm.h>
#include <linux/pm.h>
#include <linux/ethtool.h>
@@ -339,6 +341,320 @@ ibmveth_cleanup_rx_resources(struct ibmveth_adapter *adapter)
}
}
+/**
+ * ibmveth_toggle_irq - Common helper to enable/disable queue interrupts
+ * @adapter: ibmveth adapter structure
+ * @queue_index: Index of the queue (0 for primary, 1+ for subordinate)
+ * @enable: true to enable, false to disable
+ *
+ * For queue 0 (primary), uses h_vio_signal() as it's registered via
+ * h_register_logical_lan(). For subordinate queues (1+), uses H_VIOCTL
+ * with H_ENABLE/DISABLE_VIO_INTERRUPT for per-queue interrupt control.
+ *
+ * Return: 0 on success, negative errno on failure (never raw H_*).
+ */
+static int
+ibmveth_toggle_irq(struct ibmveth_adapter *adapter, int queue_index,
+ bool enable)
+{
+ unsigned long h_rc;
+ unsigned long irq = adapter->queue_irq[queue_index];
+ const char *action = enable ? "enable" : "disable";
+
+ if (queue_index == 0) {
+ /* Primary queue: use h_vio_signal() */
+ h_rc = h_vio_signal(adapter->vdev->unit_address,
+ enable ? VIO_IRQ_ENABLE : VIO_IRQ_DISABLE);
+ } else {
+ /* Subordinate queues: use H_VIOCTL with hardware IRQ */
+ struct irq_data *irq_data = irq_get_irq_data(irq);
+ irq_hw_number_t hwirq;
+ u64 vioctl_cmd = enable ? H_ENABLE_VIO_INTERRUPT :
+ H_DISABLE_VIO_INTERRUPT;
+
+ if (!irq_data) {
+ netdev_err(adapter->netdev,
+ "Failed to get IRQ data for queue %d (virq=%lu)\n",
+ queue_index, irq);
+ return -EINVAL;
+ }
+
+ hwirq = irqd_to_hwirq(irq_data);
+ h_rc = plpar_hcall_norets(H_VIOCTL,
+ adapter->vdev->unit_address,
+ vioctl_cmd,
+ hwirq, 0, 0);
+
+ /*
+ * H_PARAMETER is ambiguous (already in requested state vs bad
+ * args). Fold only on disable as an idempotent mask. On enable
+ * keep it an error so a stuck-masked queue stays visible to
+ * poll/resize recovery.
+ */
+ if (h_rc == H_PARAMETER && !enable) {
+ dev_warn_ratelimited(&adapter->netdev->dev,
+ "H_VIOCTL %s IRQ returned H_PARAMETER for queue %d (hwirq=%lu)\n",
+ action, queue_index, hwirq);
+ return 0;
+ }
+ }
+
+ if (h_rc) {
+ dev_err_ratelimited(&adapter->netdev->dev,
+ "Failed to %s IRQ for queue %d, rc=0x%lx\n",
+ action, queue_index, h_rc);
+ return -EIO;
+ }
+ return 0;
+}
+
+/**
+ * ibmveth_disable_irq - Disable interrupt for a specific queue
+ * @adapter: ibmveth adapter structure
+ * @queue_index: Index of the queue (0 for primary, 1+ for subordinate)
+ *
+ * Return: 0 on success, negative errno on failure
+ */
+static int
+ibmveth_disable_irq(struct ibmveth_adapter *adapter, int queue_index)
+{
+ return ibmveth_toggle_irq(adapter, queue_index, false);
+}
+
+/**
+ * ibmveth_enable_irq - Enable interrupt for a specific queue
+ * @adapter: ibmveth adapter structure
+ * @queue_index: Index of the queue (0 for primary, 1+ for subordinate)
+ *
+ * Return: 0 on success, negative errno on failure
+ */
+static int
+ibmveth_enable_irq(struct ibmveth_adapter *adapter, int queue_index)
+{
+ return ibmveth_toggle_irq(adapter, queue_index, true);
+}
+
+/**
+ * ibmveth_dispose_subordinate_irq_mapping - Drop one subordinate virq mapping
+ * @adapter: ibmveth adapter structure
+ * @queue_idx: RX queue index (1..N)
+ *
+ * Subordinate queues get mappings from irq_create_mapping() during PHYP
+ * registration. Queue 0 uses netdev->irq from device tree and is left alone.
+ *
+ * Bound against IBMVETH_MAX_RX_QUEUES, not num_rx_queues: a caller may
+ * dispose a queue that is no longer in the published live set but still
+ * owns a virq in queue_irq[]. Contrast with the bulk helper, which only
+ * walks 1..num_rx_queues-1 (close / open-fail cleanup of the live set).
+ *
+ * Linux virq lifetime is owned by interrupt cleanup helpers. Call this only
+ * after free_irq() when a handler was installed, or from registration failure
+ * cleanup before request_irq().
+ */
+static void
+ibmveth_dispose_subordinate_irq_mapping(struct ibmveth_adapter *adapter,
+ int queue_idx)
+{
+ if (queue_idx <= 0 || queue_idx >= IBMVETH_MAX_RX_QUEUES)
+ return;
+
+ if (adapter->queue_irq[queue_idx]) {
+ irq_dispose_mapping(adapter->queue_irq[queue_idx]);
+ adapter->queue_irq[queue_idx] = 0;
+ }
+}
+
+/**
+ * ibmveth_dispose_subordinate_irq_mappings - Drop virq mappings for queues 1..N
+ * @adapter: ibmveth adapter structure
+ *
+ * Bulk helper for close / open-fail cleanup of the published live set
+ * (queues 1..num_rx_queues-1). Paths that need a retired or not-yet-published
+ * queue must call ibmveth_dispose_subordinate_irq_mapping() directly.
+ */
+static void
+ibmveth_dispose_subordinate_irq_mappings(struct ibmveth_adapter *adapter)
+{
+ int i;
+
+ for (i = 1; i < adapter->num_rx_queues; i++)
+ ibmveth_dispose_subordinate_irq_mapping(adapter, i);
+}
+
+/**
+ * ibmveth_setup_rx_interrupts - Register IRQs and enable NAPI
+ * @adapter: ibmveth adapter structure
+ *
+ * Registers interrupt handlers for all RX queues, enables NAPI, then
+ * enables hypervisor interrupt delivery for multi-queue mode after
+ * every queue has a Linux handler installed. For multi-queue open the
+ * caller should replenish RX buffers before this helper so traffic
+ * during open is not dropped (PHYP only interrupts after a successful
+ * enqueue, which needs buffers). Single-queue open leaves PHYP masked
+ * here and kicks NAPI afterward (classic path: first poll posts then
+ * enables).
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_setup_rx_interrupts(struct ibmveth_adapter *adapter)
+{
+ struct net_device *netdev = adapter->netdev;
+ int i, rc, num = adapter->num_rx_queues;
+
+ for (i = 0; i < num; i++) {
+ if (!adapter->queue_irq[i]) {
+ netdev_err(netdev, "queue %d has invalid IRQ (0)\n", i);
+ rc = -EINVAL;
+ goto err_free_irqs;
+ }
+
+ rc = request_irq(adapter->queue_irq[i], ibmveth_interrupt,
+ 0, netdev->name, &adapter->napi[i]);
+ if (rc) {
+ netdev_err(netdev,
+ "request_irq() failed for irq 0x%x queue %d: %d\n",
+ adapter->queue_irq[i], i, rc);
+ goto err_free_irqs;
+ }
+ }
+
+ for (i = 0; i < num; i++)
+ napi_enable(&adapter->napi[i]);
+
+ if (adapter->multi_queue && num > 1) {
+ for (i = 0; i < num; i++) {
+ rc = ibmveth_enable_irq(adapter, i);
+ if (rc) {
+ netdev_err(netdev,
+ "Failed to enable IRQ for queue %d, rc=%d\n",
+ i, rc);
+ for (; i >= 0; i--) {
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ }
+ rc = -EIO;
+ goto err_disable_napi;
+ }
+ }
+ }
+
+ /* Set only on full success; fail paths leave this false so a later
+ * close() / cleanup is a no-op.
+ */
+ adapter->rx_irq_setup = true;
+ return 0;
+
+err_disable_napi:
+ /* STOP: remask after napi_disable; an in-flight poll can re-arm. */
+ for (i = 0; i < num; i++)
+ napi_disable(&adapter->napi[i]);
+ for (i = 0; i < num; i++) {
+ if (!adapter->queue_irq[i])
+ continue;
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ }
+ for (i = 0; i < num; i++) {
+ if (adapter->queue_irq[i])
+ free_irq(adapter->queue_irq[i], &adapter->napi[i]);
+ }
+ goto err_dispose_mappings;
+
+err_free_irqs:
+ while (--i >= 0)
+ free_irq(adapter->queue_irq[i], &adapter->napi[i]);
+err_dispose_mappings:
+ /* Both setup failure paths own subordinate virq disposal. */
+ ibmveth_dispose_subordinate_irq_mappings(adapter);
+ return rc;
+}
+
+/**
+ * ibmveth_cleanup_rx_interrupts - Mask PHYP IRQs, stop NAPI, and free IRQs
+ * @adapter: ibmveth adapter structure
+ *
+ * Mask and synchronize each queue IRQ before napi_disable() so the handler
+ * cannot miss a PHYP mask while NAPI is already dead. Remask after
+ * napi_disable() in case an in-flight poll re-armed PHYP while we waited.
+ * free_irq() runs only after that. Safe for close and for open failure after
+ * setup_rx_interrupts() already unmasked PHYP. No-op if setup never
+ * succeeded (avoids double napi_disable / free_irq after a failed close+open
+ * while IFF_UP remains set).
+ */
+static void
+ibmveth_cleanup_rx_interrupts(struct ibmveth_adapter *adapter)
+{
+ int i;
+
+ if (!adapter->rx_irq_setup)
+ return;
+
+ for (i = 0; i < adapter->num_rx_queues; i++) {
+ if (!adapter->queue_irq[i])
+ continue;
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ }
+
+ for (i = 0; i < adapter->num_rx_queues; i++)
+ napi_disable(&adapter->napi[i]);
+
+ for (i = 0; i < adapter->num_rx_queues; i++) {
+ if (!adapter->queue_irq[i])
+ continue;
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ }
+
+ for (i = 0; i < adapter->num_rx_queues; i++) {
+ if (adapter->queue_irq[i])
+ free_irq(adapter->queue_irq[i], &adapter->napi[i]);
+ }
+
+ ibmveth_dispose_subordinate_irq_mappings(adapter);
+
+ /* Queue 0 uses netdev->irq; leave queue_irq[0] for next open. */
+ adapter->rx_irq_setup = false;
+}
+
+/**
+ * ibmveth_schedule_rx_queue - Mask PHYP IRQ and schedule NAPI for one RX queue
+ * @adapter: ibmveth adapter structure
+ * @qindex: RX queue index
+ *
+ * Shared by the IRQ handler and process-context kick sites (open, resume,
+ * pool sysfs, poll_controller).
+ *
+ * Return: true if napi_schedule_prep() succeeded and NAPI was scheduled.
+ * Mask is attempted in that case; a failed disable_irq() is logged by the
+ * helper and does not change the return (queue may still be unmasked).
+ * false if the index is out of range (WARN_ON, then return) or
+ * prep failed (including NAPI already scheduled).
+ */
+static bool ibmveth_schedule_rx_queue(struct ibmveth_adapter *adapter,
+ int qindex)
+{
+ struct napi_struct *napi = &adapter->napi[qindex];
+
+ if (WARN_ON(qindex < 0 || qindex >= adapter->num_rx_queues))
+ return false;
+
+ /*
+ * Only mask PHYP when NAPI will run. Masking on prep failure can
+ * race a completing poll that already re-enabled the queue, leaving
+ * NAPI idle with the IRQ masked (TX works, RX stalls) until reload.
+ * Storm prevention on teardown remains in cleanup/disable paths.
+ */
+ if (napi_schedule_prep(napi)) {
+ /* Failure is already logged with the hcall rc by the helper. */
+ ibmveth_disable_irq(adapter, qindex);
+ __napi_schedule(napi);
+ return true;
+ }
+ return false;
+}
+
/* setup the initial settings for a buffer pool */
static void ibmveth_init_buffer_pool(struct ibmveth_buff_pool *pool,
u32 pool_index, u32 pool_size,
@@ -959,8 +1275,6 @@ static int ibmveth_open(struct net_device *netdev)
netdev_dbg(netdev, "open starting\n");
- napi_enable(&adapter->napi[0]);
-
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
rxq_entries += adapter->rx_buff_pool[0][i].size;
@@ -984,7 +1298,8 @@ static int ibmveth_open(struct net_device *netdev)
adapter->rx_queue[0].queue_len;
rxq_desc.fields.address = adapter->rx_queue[0].queue_dma;
- h_vio_signal(adapter->vdev->unit_address, VIO_IRQ_DISABLE);
+ adapter->queue_irq[0] = netdev->irq;
+ ibmveth_disable_irq(adapter, 0);
lpar_rc = ibmveth_register_logical_lan(adapter, rxq_desc, mac_address);
@@ -1005,24 +1320,20 @@ static int ibmveth_open(struct net_device *netdev)
if (rc)
goto out_free_tx_ltb;
- netdev_dbg(netdev, "registering irq 0x%x\n", netdev->irq);
- rc = request_irq(netdev->irq, ibmveth_interrupt, 0, netdev->name,
- netdev);
- if (rc != 0) {
- netdev_err(netdev, "unable to request irq 0x%x, rc %d\n",
- netdev->irq, rc);
+ rc = ibmveth_setup_rx_interrupts(adapter);
+ if (rc) {
do {
lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
} while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
-
goto out_free_buffer_pools;
}
netdev_dbg(netdev, "initial replenish cycle\n");
- ibmveth_interrupt(netdev->irq, netdev);
+ ibmveth_schedule_rx_queue(adapter, 0);
netif_tx_start_all_queues(netdev);
+ adapter->opened = true;
netdev_dbg(netdev, "open complete\n");
return 0;
@@ -1036,7 +1347,6 @@ static int ibmveth_open(struct net_device *netdev)
out_free_filter_list:
ibmveth_free_filter_list(adapter);
out:
- napi_disable(&adapter->napi[0]);
return rc;
}
@@ -1046,27 +1356,32 @@ static int ibmveth_close(struct net_device *netdev)
long lpar_rc;
int i;
- netdev_dbg(netdev, "close starting\n");
+ /* Gate on opened, not IFF_UP: pool_store/change_mtu close+open can
+ * leave IFF_UP set after a failed reopen.
+ */
+ if (!adapter->opened)
+ return 0;
- napi_disable(&adapter->napi[0]);
+ adapter->opened = false;
+
+ netdev_dbg(netdev, "close starting\n");
netif_tx_stop_all_queues(netdev);
- h_vio_signal(adapter->vdev->unit_address, VIO_IRQ_DISABLE);
+ ibmveth_cleanup_rx_interrupts(adapter);
+ /* Wait for softirq/poll that already passed shutdown checks. */
+ synchronize_net();
+ ibmveth_update_rx_no_buffer(adapter);
+ /* Full LAN teardown (subordinates arrive with register helpers). */
do {
lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
} while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
-
if (lpar_rc != H_SUCCESS) {
- netdev_err(netdev, "h_free_logical_lan failed with %lx, "
- "continuing with close\n", lpar_rc);
+ netdev_err(adapter->netdev,
+ "h_free_logical_lan failed with %lx, continuing\n",
+ lpar_rc);
}
-
- free_irq(netdev->irq, netdev);
-
- ibmveth_update_rx_no_buffer(adapter);
-
ibmveth_free_buffer_pools(adapter);
ibmveth_cleanup_rx_resources(adapter);
ibmveth_free_filter_list(adapter);
@@ -1710,7 +2025,7 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
container_of(napi, struct ibmveth_adapter, napi[0]);
struct net_device *netdev = adapter->netdev;
int frames_processed = 0;
- unsigned long lpar_rc;
+ int rc;
u16 mss = 0;
restart_poll:
@@ -1810,15 +2125,14 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
/* We think we are done - reenable interrupts,
* then check once more to make sure we are done.
*/
- lpar_rc = h_vio_signal(adapter->vdev->unit_address, VIO_IRQ_ENABLE);
- if (WARN_ON(lpar_rc != H_SUCCESS)) {
+ rc = ibmveth_enable_irq(adapter, 0);
+ if (rc) {
schedule_work(&adapter->work);
goto out;
}
if (ibmveth_rxq_pending_buffer(adapter) && napi_schedule(napi)) {
- lpar_rc = h_vio_signal(adapter->vdev->unit_address,
- VIO_IRQ_DISABLE);
+ ibmveth_disable_irq(adapter, 0);
goto restart_poll;
}
@@ -1828,16 +2142,20 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
static irqreturn_t ibmveth_interrupt(int irq, void *dev_instance)
{
- struct net_device *netdev = dev_instance;
+ struct napi_struct *napi = dev_instance;
+ struct net_device *netdev = napi->dev;
struct ibmveth_adapter *adapter = netdev_priv(netdev);
- unsigned long lpar_rc;
+ int qindex;
- if (napi_schedule_prep(&adapter->napi[0])) {
- lpar_rc = h_vio_signal(adapter->vdev->unit_address,
- VIO_IRQ_DISABLE);
- WARN_ON(lpar_rc != H_SUCCESS);
- __napi_schedule(&adapter->napi[0]);
- }
+ qindex = napi - adapter->napi;
+ /*
+ * Quiet on out-of-range: teardown can leave a residual IRQ after the
+ * live count drops. Do not WARN-storm; return IRQ_NONE until free_irq.
+ */
+ if (qindex < 0 || qindex >= adapter->num_rx_queues)
+ return IRQ_NONE;
+
+ ibmveth_schedule_rx_queue(adapter, qindex);
return IRQ_HANDLED;
}
@@ -1942,8 +2260,10 @@ static int ibmveth_change_mtu(struct net_device *dev, int new_mtu)
#ifdef CONFIG_NET_POLL_CONTROLLER
static void ibmveth_poll_controller(struct net_device *dev)
{
- ibmveth_replenish_task(netdev_priv(dev));
- ibmveth_interrupt(dev->irq, dev);
+ struct ibmveth_adapter *adapter = netdev_priv(dev);
+
+ ibmveth_replenish_task(adapter);
+ ibmveth_schedule_rx_queue(adapter, 0);
}
#endif
@@ -2356,8 +2676,8 @@ static ssize_t veth_pool_store(struct kobject *kobj, struct attribute *attr,
}
rtnl_unlock();
- /* kick the interrupt handler to allocate/deallocate pools */
- ibmveth_interrupt(netdev->irq, netdev);
+ /* kick RX processing to allocate/deallocate pools */
+ ibmveth_schedule_rx_queue(adapter, 0);
return count;
unlock_err:
@@ -2397,7 +2717,9 @@ static struct kobj_type ktype_veth_pool = {
static int ibmveth_resume(struct device *dev)
{
struct net_device *netdev = dev_get_drvdata(dev);
- ibmveth_interrupt(netdev->irq, netdev);
+ struct ibmveth_adapter *adapter = netdev_priv(netdev);
+
+ ibmveth_schedule_rx_queue(adapter, 0);
return 0;
}
diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h
index 7e956278e005..a3178032c6f2 100644
--- a/drivers/net/ethernet/ibm/ibmveth.h
+++ b/drivers/net/ethernet/ibm/ibmveth.h
@@ -318,6 +318,10 @@ struct ibmveth_adapter {
unsigned int queue_irq[IBMVETH_MAX_RX_QUEUES];
int multi_queue;
unsigned int num_rx_queues;
+ /* Lifetime: true after successful ndo_open until close clears it. */
+ bool opened;
+ /* Lifetime: true while RX IRQ handlers / NAPI are installed. */
+ bool rx_irq_setup;
int rx_csum;
int large_send;
bool is_active_trunk;
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 05/15] ibmveth: Refactor RX interrupt control for MQ RX queues
2026-09-25 18:38 ` [PATCH net-next v7 05/15] ibmveth: Refactor RX interrupt control for MQ RX queues Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found 5 potential
issue(s) to consider.
Critical: 0 · High: 0 (2) · Medium: 3 · Low: 2
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] This patch fixes a hang that already exists in mainline, but
the commit message's reason for leaving out Fixes:/Cc: stable is wrong.
- [Medium] ibmveth_cleanup_rx_interrupts() masks PHYP and calls
synchronize_irq() (lines 593-598), calls napi_disable() (600-601),…
- [Medium] After a failed internal reopen, ibmveth_set_channels() (gated
on IFF_UP) can allocate TX LTBs that the now-gated ibmveth_close()…
- [Low] ibmveth_close() now calls ibmveth_update_rx_no_buffer() (line
1375) before h_free_logical_lan() (lines 1377-1379).
- [Low] The commit message argues for this commit's safety using code
that does not exist at this commit.
Pre-existing issues:
- [High] If ibmveth_alloc_buffer_pools() fails after a successful
ibmveth_register_logical_lan(), ibmveth_open() jumps to…
- [High] ibmveth_remove() calls cancel_work_sync(&adapter->work) before
unregister_netdev().
--
ibmveth: Refactor RX interrupt control for MQ RX queues
This commit adds queue-indexed RX interrupt helpers to ibmveth and wires
them into open, close, poll, the IRQ handler and the process-context RX
kick sites. It also adds opened and rx_irq_setup flags so that a close
after a failed reopen doesn't repeat napi_disable() and free_irq().
Runtime stays single queue.
> Opened / rx_irq_setup also closes a pre-existing hang: after a
> failed reopen, a later ndo_stop used to napi_disable and free_irq
> a second time (rtnl spin + already-free IRQ). That depends on the
> helpers in this patch, so there is no standalone Fixes: tag.
[Severity: Medium]
Is the "depends on the helpers" part accurate? The hang fix appears to
be only the opened flag:
ibmveth_close()
if (!adapter->opened)
return 0;
ibmveth_open()
adapter->opened = true; /* success path only */
Neither of these uses the new IRQ helpers, and the same bool would
apply to the parent tree unchanged.
In the parent tree, every ibmveth_open() failure ends at the out: label
with napi_disable(), and ibmveth_close() also starts with napi_disable().
ibmveth_change_mtu(), ibmveth_set_csum_offload(), ibmveth_set_tso() and
veth_pool_store() call ibmveth_close()/ibmveth_open() directly. When the
reopen fails, IFF_UP stays set.
The next ndo_stop (ifdown, or unregister_netdev() from ibmveth_remove())
then calls napi_disable() on a NAPI that is already disabled. It spins
under RTNL, and would then free_irq() an IRQ that was already freed.
Could the opened gate go into a small standalone patch ahead of the
series, with something like:
Fixes: 88426f2acae0 ("ibmveth: Cleanup error handling inside ibmveth_open")
Cc: stable@vger.kernel.org
so stable kernels get the hang fix?
> cleanup_rx_interrupts() masks PHYP and synchronizes IRQs before
> napi_disable, remasks and synchronizes again after it. The second
> remask only catches a re-arm that lands before it; napi_disable
> does not wait for poll to return, so enable_irq can still run
> after free_irq. Close then proceeds to
> h_free_logical_lan(): free_irq before free_lan is intentional once
> PHYP delivery is masked.
[ ... ]
> On setup_rx enable-fail (MQ path), if enable_irq() fails for queue i,
> remask+sync queues 0..i, including the one that failed, before
> napi_disable/free_irq; the rollback loop used while (--i) and skipped
> it.
[ ... ]
> stalled until reload). Teardown storm control stays on STOP
> (disable_irq + synchronize_irq before napi_disable) and the
> poll_stopping() re-arm guard added in P09, not on the schedule helper
> failure path.
[Severity: Low]
These parts of the message describe code that isn't in this commit.
There is no poll_stopping() helper yet. At this commit, ibmveth_poll()
calls ibmveth_enable_irq(adapter, 0) unconditionally after
napi_complete_done(). The guard only arrives with the later "ibmveth:
Harden RX poll path with helpers".
The first paragraph says enable_irq can still run after free_irq. It
then says free_irq before free_lan is intentional once PHYP delivery is
masked. At this commit, don't those two statements contradict each
other?
The while (--i) rollback loop isn't in the parent tree. It looks like
it came from an earlier revision of this patch.
The MQ enable-fail rollback also can't run at this commit. Probe sets
multi_queue = 0 and num_rx_queues = IBMVETH_DEFAULT_RX_QUEUES, and
IBMVETH_MAX_RX_QUEUES is 1U.
Series-relative labels such as P09 also lose their meaning once the
patch is applied.
Could the version history and series-relative notes move below the ---
line or into the cover letter? The body could then state plainly that
the re-arm window after free_irq stays open until "ibmveth: Harden RX
poll path with helpers".
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index 5813352943fb1..a22a17e05ae1c 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
[ ... ]
> @@ -339,6 +341,320 @@ ibmveth_cleanup_rx_resources(struct ibmveth_adapter *adapter)
[ ... ]
> +static void
> +ibmveth_cleanup_rx_interrupts(struct ibmveth_adapter *adapter)
> +{
> + int i;
> +
> + if (!adapter->rx_irq_setup)
> + return;
> +
> + for (i = 0; i < adapter->num_rx_queues; i++) {
> + if (!adapter->queue_irq[i])
> + continue;
> + ibmveth_disable_irq(adapter, i);
> + synchronize_irq(adapter->queue_irq[i]);
> + }
> +
> + for (i = 0; i < adapter->num_rx_queues; i++)
> + napi_disable(&adapter->napi[i]);
> +
> + for (i = 0; i < adapter->num_rx_queues; i++) {
> + if (!adapter->queue_irq[i])
> + continue;
> + ibmveth_disable_irq(adapter, i);
> + synchronize_irq(adapter->queue_irq[i]);
> + }
> +
> + for (i = 0; i < adapter->num_rx_queues; i++) {
> + if (adapter->queue_irq[i])
> + free_irq(adapter->queue_irq[i], &adapter->napi[i]);
> + }
[Severity: Medium]
Is the second mask here actually ordered after the tail of an in-flight
ibmveth_poll()?
napi_disable() returns as soon as it can claim NAPI_STATE_SCHED. That
happens right after the poll's napi_complete_done(), not when the poll
returns. The poll then continues:
ibmveth_poll()
if (!napi_complete_done(napi, frames_processed))
goto out;
...
rc = ibmveth_enable_irq(adapter, 0); /* H_VIO_SIGNAL enable */
...
napi_schedule(napi) /* fails, DISABLE is set */
At this commit there is no teardown check in that path. The re-arm can
land after the second ibmveth_disable_irq() above and after free_irq().
Queue 0 would then stay unmasked in PHYP with no Linux handler until
h_free_logical_lan().
The only wait for the poll tail is the synchronize_net() in
ibmveth_close(), and it runs after free_irq(). The parent tree called
h_free_logical_lan() before free_irq(), so it never had a window with
PHYP live and no handler installed.
The kernel-doc ("Remask after napi_disable() in case an in-flight poll
re-armed PHYP ... free_irq() runs only after that") assumes this
ordering. The err_disable_napi rollback in
ibmveth_setup_rx_interrupts() uses the same pattern.
Later in the series, "ibmveth: Harden RX poll path with helpers" adds an
ibmveth_poll_stopping() check after napi_complete_done(). That covers
dev_close(), where netif_running() is already false.
It doesn't cover the direct ibmveth_close() callers
(ibmveth_change_mtu(), veth_pool_store(), ibmveth_set_csum_offload(),
ibmveth_set_tso()), where netif_running() is still true.
Would it close the window to wait for the poll tail (for example with
synchronize_net()) before the second mask and synchronize_irq(), and
only then call free_irq()?
[ ... ]
> @@ -1005,24 +1320,20 @@ static int ibmveth_open(struct net_device *netdev)
> if (rc)
> goto out_free_tx_ltb;
>
[Severity: High]
This is a pre-existing issue, not one introduced by this patch. This
is the ibmveth_alloc_buffer_pools() failure path, which runs after a
successful ibmveth_register_logical_lan(). Does it leave the logical
LAN registered with PHYP?
out_free_tx_ltb never calls h_free_logical_lan():
out_free_tx_ltb:
while (--i >= 0)
ibmveth_free_tx_ltb(adapter, i);
ibmveth_cleanup_rx_resources(adapter);
So the buffer list, RX queue and filter list are unmapped and freed
while PHYP may still hold the registration.
With the new opened gate, no later ndo_stop reaches h_free_logical_lan()
either. In the parent tree, the later close hung in the second
napi_disable() before getting that far.
This looks fixed later in the series by "ibmveth: Add RX queue register
helpers for MQ", which moves ibmveth_alloc_buffer_pools() ahead of
ibmveth_register_rx_queues().
[ ... ]
> @@ -1046,27 +1356,32 @@ static int ibmveth_close(struct net_device *netdev)
> long lpar_rc;
> int i;
>
> - netdev_dbg(netdev, "close starting\n");
> + /* Gate on opened, not IFF_UP: pool_store/change_mtu close+open can
> + * leave IFF_UP set after a failed reopen.
> + */
> + if (!adapter->opened)
> + return 0;
[Severity: Medium]
Can this early return leak TX LTBs allocated by ibmveth_set_channels()?
After a failed internal reopen, IFF_UP stays set but opened is false.
ibmveth_set_channels() only checks IFF_UP:
ibmveth_set_channels() {
if (!(netdev->flags & IFF_UP))
return netif_set_real_num_tx_queues(netdev, goal);
...
for (i = old; i < goal; i++) {
if (adapter->tx_ltb_ptr[i])
continue;
rc = ibmveth_allocate_tx_ltb(adapter, i);
...
}
So ethtool -L in that state allocates new LTBs. ibmveth_close() then
returns here and skips the ibmveth_free_tx_ltb() loop. The next
ibmveth_open() overwrites tx_ltb_ptr[]/tx_ltb_dma[], which leaks the
buffer and its DMA mapping.
Before this patch the same sequence hung in the second napi_disable(),
so the gate trades the hang for a leak.
Later in the series this goes away: ibmveth_set_channels() checks
!adapter->opened and then only calls netif_set_real_num_tx_queues(),
without allocating LTBs. Could that check go into this patch alongside
the gate?
>
> - napi_disable(&adapter->napi[0]);
> + adapter->opened = false;
> +
> + netdev_dbg(netdev, "close starting\n");
>
> netif_tx_stop_all_queues(netdev);
>
> - h_vio_signal(adapter->vdev->unit_address, VIO_IRQ_DISABLE);
> + ibmveth_cleanup_rx_interrupts(adapter);
> + /* Wait for softirq/poll that already passed shutdown checks. */
> + synchronize_net();
>
> + ibmveth_update_rx_no_buffer(adapter);
> + /* Full LAN teardown (subordinates arrive with register helpers). */
> do {
> lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
> } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
> -
> if (lpar_rc != H_SUCCESS) {
> - netdev_err(netdev, "h_free_logical_lan failed with %lx, "
> - "continuing with close\n", lpar_rc);
> + netdev_err(adapter->netdev,
> + "h_free_logical_lan failed with %lx, continuing\n",
> + lpar_rc);
> }
> -
> - free_irq(netdev->irq, netdev);
> -
> - ibmveth_update_rx_no_buffer(adapter);
> -
[Severity: Low]
Does moving ibmveth_update_rx_no_buffer() ahead of h_free_logical_lan()
drop counts from rx_no_buffer?
While the logical LAN is still registered, PHYP keeps consuming posted
buffers. Once they run out, it keeps incrementing the no-buffer counter
in the last 8 bytes of buffer_list_addr[0]. NAPI is already disabled at
this point, so nothing replenishes the buffers.
Any increments between this read and H_FREE_LOGICAL_LAN never reach
adapter->rx_no_buffer, which is the value kept across close.
The commit message says "Close harvests the PHYP no-buffer count before
h_free so the read still hits a live buffer-list page". In either
order, though, buffer_list_addr[0] is only unmapped and freed later, in
ibmveth_cleanup_rx_resources().
Wouldn't reading after h_free_logical_lan(), as the old code did, be
just as safe and give the final value?
The same order is still there at the end of the series: the final
ibmveth_close() reads the counter under replenish_lock and then calls
ibmveth_free_all_queues().
[ ... ]
> @@ -1810,15 +2125,14 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
> /* We think we are done - reenable interrupts,
> * then check once more to make sure we are done.
> */
> - lpar_rc = h_vio_signal(adapter->vdev->unit_address, VIO_IRQ_ENABLE);
> - if (WARN_ON(lpar_rc != H_SUCCESS)) {
> + rc = ibmveth_enable_irq(adapter, 0);
> + if (rc) {
> schedule_work(&adapter->work);
> goto out;
> }
[Severity: High]
This isn't a bug introduced by this patch, but can this schedule_work()
race with ibmveth_remove()?
ibmveth_remove() {
...
cancel_work_sync(&adapter->work);
...
unregister_netdev(netdev);
free_netdev(netdev);
...
}
NAPI poll stays live until unregister_netdev() closes the device. If
ibmveth_enable_irq() fails in that window, adapter->work can be queued
again after cancel_work_sync(). free_netdev() then frees the embedded
work_struct while it may still be queued.
The parent tree has the same order. "ibmveth: Enable multi-queue RX
receive path" later reorders ibmveth_remove() so that unregister_netdev()
runs before cancel_work_sync().
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 06/15] ibmveth: Refactor TX resource allocation in open/close paths
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (4 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 05/15] ibmveth: Refactor RX interrupt control for MQ RX queues Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 07/15] ibmveth: Add RX queue register helpers for MQ Mingming Cao
` (9 subsequent siblings)
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
Same story as the RX refactor: pull TX LTB alloc/free out of open/close
into helpers and wire them in this patch.
ibmveth_alloc_tx_resources()
ibmveth_free_tx_resources()
They wrap the existing per-queue allocate_tx_ltb() / free_tx_ltb()
primitives. alloc_tx_resources() allocates every TX queue and unwinds
partial failure itself; free_tx_resources() walks real_num_tx_queues.
The helpers remove dependence on shared open/close loop indices and
match the RX helper structure. TX was already multi-queue capable via
ethtool -L.
Also tighten TX LTB lifetime: free_tx_ltb() returns early if
tx_ltb_ptr[] is already NULL, then clears both tx_ltb_ptr[] and
tx_ltb_dma[] before unmapping and freeing, so start_xmit() cannot pick
up a slot that is mid-teardown. allocate_tx_ltb() clears tx_ltb_dma[]
on the DMA-map failure path.
Move TX LTB allocation to the end of open(), after LAN registration,
RX pools, RX interrupt setup, and the initial replenish kick. A late
alloc_tx_resources() failure jumps to out_cleanup_rx_interrupts and
must not call free_tx_resources() again: alloc already freed any
partial TX LTBs. RX is live by then, so that label also calls
synchronize_net() after cleanup_rx_interrupts(), as close() does,
before the RX ring and pools are freed. start_xmit() bails if
tx_ltb_ptr[] is gone, counting
the drop in tx_dropped and falling into the existing out: label like
the function's other drop paths, so RX can be live while TX LTB alloc
still runs and close/failed-reopen cannot race a live mapping. The
close path is quiesced by netif_tx_disable(); NULL-first in
free_tx_ltb() only closes the check-then-use window, it is not itself
a UAF barrier. set_channels() IFF_UP vs opened is later (P14/P15).
After LAN registration, open-fail teardown issues h_free_logical_lan()
before RX pool DMA teardown on the pool-fail path that previously never
issued that hcall (missing deregistration, not a preference reorder).
close() quiesces TX with netif_tx_disable() (stop_all_queues does not
wait for in-flight ndo_start_xmit), then frees LTBs after
h_free_logical_lan() via free_tx_resources() - required because direct
close() callers bypass synchronize_net().
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- late alloc_tx_resources() failure: synchronize_net()
after cleanup_rx_interrupts(), as close() does, before
the RX ring and pools are freed
- noted: no Fixes: peel of the pool-fail h_free; stay
at 15; [PATCH net 1/2]
<cover.1790357373.git.mmc@linux.ibm.com>
(Fixes: d43732ce021f)
Changes in v6:
- NULL tx_ltb_ptr[idx] and zero tx_ltb_dma[idx] before unmap/free, so a
racing start_xmit() fails the pointer check. Close-path safety is
still netif_tx_disable()
- the NULL-LTB start_xmit drop increments tx_dropped and falls into
the existing out: label
- comment on allocate_tx_ltb(): caller must leave tx_ltb_ptr[idx] NULL
- kdoc alloc/free_tx_resources says real_num_tx_queues
- noted: ethtool -L TX shrink still stop-then-free; cover leftovers
Changes in v5:
- Quiesce TX with netif_tx_disable before free (stop_all_queues does not
wait for in-flight xmit); free LTBs after h_free_logical_lan - direct
close() callers bypass synchronize_net()
- Guard start_xmit if tx_ltb_ptr gone so open can leave RX live while TX
LTB alloc still runs (also covers close/failed-reopen with IFF_UP set)
- Drop fake mid-open TX-leak / Fixes: motivation; reword as helper
extraction matching RX (shared loop-index independence)
- Free TX LTB by pointer presence (drop dma==0 sentinel; dma_mapping_error
already cleared the slot on map failure)
- Document intentional open-fail LAN-first unwind (free_lan before RX
pool/DMA teardown) rather than leaving it silent in a TX-only refactor
- Drop drive-by blank-line cosmetics (header / start_xmit)
Changes in v4:
- Introduce the TX resource helpers in the same patch that wires their
first open/close callers.
- Do not free TX LTBs again after a failed alloc_tx_resources();
harden free_tx_ltb() against unset slots.
- Move TX allocation after RX IRQ setup / replenish kick so open()
failure unwind no longer depends on a shared loop index (also fixes
a mid-open TX LTB leak).
drivers/net/ethernet/ibm/ibmveth.c | 117 ++++++++++++++++++++++-------
1 file changed, 91 insertions(+), 26 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index a22a17e05ae1..011082db1e08 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -1206,12 +1206,27 @@ static int ibmveth_rxq_harvest_buffer(struct ibmveth_adapter *adapter,
static void ibmveth_free_tx_ltb(struct ibmveth_adapter *adapter, int idx)
{
- dma_unmap_single(&adapter->vdev->dev, adapter->tx_ltb_dma[idx],
- adapter->tx_ltb_size, DMA_TO_DEVICE);
- kfree(adapter->tx_ltb_ptr[idx]);
+ void *ltb = adapter->tx_ltb_ptr[idx];
+ dma_addr_t dma = adapter->tx_ltb_dma[idx];
+
+ if (!ltb)
+ return;
+
+ /*
+ * Clear the slot before releasing it. start_xmit() tests
+ * tx_ltb_ptr[idx] to decide whether the LTB is usable.
+ */
adapter->tx_ltb_ptr[idx] = NULL;
+ adapter->tx_ltb_dma[idx] = 0;
+
+ dma_unmap_single(&adapter->vdev->dev, dma, adapter->tx_ltb_size,
+ DMA_TO_DEVICE);
+ kfree(ltb);
}
+/* Caller must ensure tx_ltb_ptr[idx] is NULL. open() runs on
+ * probe-zeroed slots; set_channels() skips populated indices.
+ */
static int ibmveth_allocate_tx_ltb(struct ibmveth_adapter *adapter, int idx)
{
adapter->tx_ltb_ptr[idx] = kzalloc(adapter->tx_ltb_size,
@@ -1230,12 +1245,54 @@ static int ibmveth_allocate_tx_ltb(struct ibmveth_adapter *adapter, int idx)
"unable to DMA map tx long term buffer\n");
kfree(adapter->tx_ltb_ptr[idx]);
adapter->tx_ltb_ptr[idx] = NULL;
+ adapter->tx_ltb_dma[idx] = 0;
return -ENOMEM;
}
return 0;
}
+/**
+ * ibmveth_alloc_tx_resources - Allocate TX LTBs for real_num_tx_queues
+ * @adapter: ibmveth adapter structure
+ *
+ * Allocates TX Long Term Buffers (LTBs) for real_num_tx_queues.
+ *
+ * Return: 0 on success, -ENOMEM on failure
+ */
+static int ibmveth_alloc_tx_resources(struct ibmveth_adapter *adapter)
+{
+ struct net_device *netdev = adapter->netdev;
+ int i;
+
+ for (i = 0; i < netdev->real_num_tx_queues; i++) {
+ if (ibmveth_allocate_tx_ltb(adapter, i))
+ goto err_free_ltbs;
+ }
+
+ return 0;
+
+err_free_ltbs:
+ while (--i >= 0)
+ ibmveth_free_tx_ltb(adapter, i);
+ return -ENOMEM;
+}
+
+/**
+ * ibmveth_free_tx_resources - Free TX LTBs for real_num_tx_queues
+ * @adapter: ibmveth adapter structure
+ *
+ * Frees TX Long Term Buffers (LTBs) for real_num_tx_queues.
+ */
+static void ibmveth_free_tx_resources(struct ibmveth_adapter *adapter)
+{
+ struct net_device *netdev = adapter->netdev;
+ int i;
+
+ for (i = 0; i < netdev->real_num_tx_queues; i++)
+ ibmveth_free_tx_ltb(adapter, i);
+}
+
static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
union ibmveth_buf_desc rxq_desc, u64 mac_address)
{
@@ -1286,12 +1343,6 @@ static int ibmveth_open(struct net_device *netdev)
if (rc)
goto out_free_filter_list;
- rc = -ENOMEM;
- for (i = 0; i < netdev->real_num_tx_queues; i++) {
- if (ibmveth_allocate_tx_ltb(adapter, i))
- goto out_free_tx_ltb;
- }
-
mac_address = ether_addr_to_u64(netdev->dev_addr);
rxq_desc.fields.flags_len = IBMVETH_BUF_VALID |
@@ -1313,24 +1364,24 @@ static int ibmveth_open(struct net_device *netdev)
rxq_desc.desc,
mac_address);
rc = -ENONET;
- goto out_free_tx_ltb;
+ goto out_free_queue_mem;
}
rc = ibmveth_alloc_buffer_pools(adapter);
if (rc)
- goto out_free_tx_ltb;
+ goto out_unregister_lan;
rc = ibmveth_setup_rx_interrupts(adapter);
- if (rc) {
- do {
- lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
- } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
- goto out_free_buffer_pools;
- }
+ if (rc)
+ goto out_unregister_lan;
netdev_dbg(netdev, "initial replenish cycle\n");
ibmveth_schedule_rx_queue(adapter, 0);
+ rc = ibmveth_alloc_tx_resources(adapter);
+ if (rc)
+ goto out_cleanup_rx_interrupts;
+
netif_tx_start_all_queues(netdev);
adapter->opened = true;
@@ -1338,11 +1389,16 @@ static int ibmveth_open(struct net_device *netdev)
return 0;
-out_free_buffer_pools:
+out_cleanup_rx_interrupts:
+ ibmveth_cleanup_rx_interrupts(adapter);
+ /* As in close(): a poll past napi_complete_done() may still run. */
+ synchronize_net();
+out_unregister_lan:
+ do {
+ lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
+ } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
ibmveth_free_buffer_pools(adapter);
-out_free_tx_ltb:
- while (--i >= 0)
- ibmveth_free_tx_ltb(adapter, i);
+out_free_queue_mem:
ibmveth_cleanup_rx_resources(adapter);
out_free_filter_list:
ibmveth_free_filter_list(adapter);
@@ -1354,7 +1410,6 @@ static int ibmveth_close(struct net_device *netdev)
{
struct ibmveth_adapter *adapter = netdev_priv(netdev);
long lpar_rc;
- int i;
/* Gate on opened, not IFF_UP: pool_store/change_mtu close+open can
* leave IFF_UP set after a failed reopen.
@@ -1366,7 +1421,10 @@ static int ibmveth_close(struct net_device *netdev)
netdev_dbg(netdev, "close starting\n");
- netif_tx_stop_all_queues(netdev);
+ /* Disable and wait for in-flight ndo_start_xmit (stop_all_queues
+ * alone does not). Direct close() callers bypass synchronize_net().
+ */
+ netif_tx_disable(netdev);
ibmveth_cleanup_rx_interrupts(adapter);
/* Wait for softirq/poll that already passed shutdown checks. */
@@ -1382,13 +1440,14 @@ static int ibmveth_close(struct net_device *netdev)
"h_free_logical_lan failed with %lx, continuing\n",
lpar_rc);
}
+ /* Free TX LTBs after quiesce and after H_FREE_LOGICAL_LAN so xmit
+ * cannot touch unmapped bounce buffers while the LAN is live.
+ */
+ ibmveth_free_tx_resources(adapter);
ibmveth_free_buffer_pools(adapter);
ibmveth_cleanup_rx_resources(adapter);
ibmveth_free_filter_list(adapter);
- for (i = 0; i < netdev->real_num_tx_queues; i++)
- ibmveth_free_tx_ltb(adapter, i);
-
netdev_dbg(netdev, "close complete\n");
return 0;
@@ -1812,6 +1871,12 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
int i, queue_num = skb_get_queue_mapping(skb);
unsigned long mss = 0;
+ /* Close / failed reopen can free LTBs while IFF_UP is still set. */
+ if (unlikely(!adapter->tx_ltb_ptr[queue_num])) {
+ netdev->stats.tx_dropped++;
+ goto out;
+ }
+
if (ibmveth_is_packet_unsupported(skb, netdev))
goto out;
/* veth can't checksum offload UDP */
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 06/15] ibmveth: Refactor TX resource allocation in open/close paths
2026-09-25 18:38 ` [PATCH net-next v7 06/15] ibmveth: Refactor TX resource allocation in open/close paths Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found 2 potential
issue(s) to consider.
Critical: 0 · High: 1 · Medium: 0 (2) · Low: 1 (1)
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] This patch moves ibmveth_alloc_tx_resources() to the end of
ibmveth_open(), after h_register_logical_lan,…
- [Low] The new NULL-LTB guard in ibmveth_start_xmit()
(drivers/net/ethernet/ibm/ibmveth.c:1874-1878) fixes a NULL-pointer…
Pre-existing issues:
- [Medium] At this commit, the comment added to ibmveth_allocate_tx_ltb()
('Caller must ensure tx_ltb_ptr[idx] is NULL.
- [Medium] The retry loop at out_unregister_lan in ibmveth_open() (lines
1396-1399) discards the final H_FREE_LOGICAL_LAN result.
- [Low] The new drop path in ibmveth_start_xmit() does
`netdev->stats.tx_dropped++`, a non-atomic read-modify-write on a…
--
ibmveth: Refactor TX resource allocation in open/close paths
This patch moves TX long term buffer (LTB) allocation and freeing in
ibmveth_open() and ibmveth_close() into two new helpers,
ibmveth_alloc_tx_resources() and ibmveth_free_tx_resources(). TX LTB
allocation now happens at the end of open(), after LAN registration and RX
setup, and ibmveth_start_xmit() gains a NULL tx_ltb_ptr[] check. close()
now uses netif_tx_disable() and frees the TX LTBs after
h_free_logical_lan().
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index a22a17e05ae1c..011082db1e085 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
> @@ -1206,12 +1206,27 @@ static int ibmveth_rxq_harvest_buffer(struct ibmveth_adapter *adapter,
[ ... ]
> +/* Caller must ensure tx_ltb_ptr[idx] is NULL. open() runs on
> + * probe-zeroed slots; set_channels() skips populated indices.
> + */
> static int ibmveth_allocate_tx_ltb(struct ibmveth_adapter *adapter, int idx)
> {
> adapter->tx_ltb_ptr[idx] = kzalloc(adapter->tx_ltb_size,
[Severity: Medium]
This is a pre-existing issue and was not introduced by this patch. Does the
precondition in this comment still hold after a failed direct reopen?
ibmveth_set_channels() gates on IFF_UP:
if (!(netdev->flags & IFF_UP))
return netif_set_real_num_tx_queues(netdev, goal);
ibmveth_close() gates on adapter->opened:
if (!adapter->opened)
return 0;
Say a close()+open() cycle fails, for example from ibmveth_change_mtu() or
pool_store(). IFF_UP stays set and opened is false. If ethtool -L runs
next, it allocates LTBs for [old, goal) and calls
netif_tx_wake_all_queues().
A later close() returns early without freeing those LTBs. The next open()
then calls ibmveth_alloc_tx_resources(), which overwrites the populated
tx_ltb_ptr[] slots.
Does that leak the kzalloc() buffer and its DMA mapping? While the queues
are awake, xmit on them also calls h_send_logical_lan() on a LAN that has
been deregistered.
The commit message defers this ("set_channels() IFF_UP vs opened is later
(P14/P15)"). The later patch "ibmveth: Complete set_channels down-path and
mq_fallback max_rx cap" makes ibmveth_set_channels() check
adapter->opened, which removes the trigger.
[ ... ]
> @@ -1313,24 +1364,24 @@ static int ibmveth_open(struct net_device *netdev)
[ ... ]
> netdev_dbg(netdev, "initial replenish cycle\n");
> ibmveth_schedule_rx_queue(adapter, 0);
>
> + rc = ibmveth_alloc_tx_resources(adapter);
> + if (rc)
> + goto out_cleanup_rx_interrupts;
> +
> netif_tx_start_all_queues(netdev);
[Severity: High]
After this move, RX is live (LAN registered, NAPI enabled, replenish
kicked) while the TX LTBs are still being allocated. Can
ibmveth_start_xmit() run at the same time as ibmveth_alloc_tx_resources()?
On the first open after probe, the TX queues were never marked DRV_XOFF.
netdev_init_one_queue() leaves the queue state zeroed, and probe does not
stop the queues. generic_xdp_tx() in net/core/dev.c checks only the queue
state, not IFF_UP:
HARD_TX_LOCK(dev, txq, cpu);
if (!netif_xmit_frozen_or_drv_stopped(txq)) {
rc = netdev_start_xmit(skb, dev, txq, 0);
If an inbound frame hits a generic XDP program that returns XDP_TX,
ibmveth_start_xmit() can run from NAPI on another CPU during this window.
The only guard is the new plain load of tx_ltb_ptr[queue_num] in
ibmveth_start_xmit(), and the stores on the allocation side are not
ordered against it.
ibmveth_allocate_tx_ltb() stores the pointer before the DMA address:
adapter->tx_ltb_ptr[idx] = kzalloc(adapter->tx_ltb_size,
GFP_KERNEL);
...
adapter->tx_ltb_dma[idx] = dma_map_single(&adapter->vdev->dev,
Could an xmit that passes the NULL check build a descriptor from
tx_ltb_dma[queue_num] == 0 and pass it to the already registered LAN
through h_send_logical_lan()?
On a DMA map failure, the same function frees the buffer before clearing
the slot:
kfree(adapter->tx_ltb_ptr[idx]);
adapter->tx_ltb_ptr[idx] = NULL;
If a later queue fails, the err_free_ltbs unwind in
ibmveth_alloc_tx_resources() calls ibmveth_free_tx_ltb() on slots that
are already published, and nothing quiesces TX first. The
synchronize_net() under out_cleanup_rx_interrupts runs only after that
unwind.
Could an xmit that is already inside the memcpy() into
tx_ltb_ptr[queue_num] write packet data into freed memory? If it re-reads
the slot after the NULL store, could it oops on a NULL pointer instead?
The commit message itself says NULL-first "is not itself a UAF barrier".
This commit also has a second path in. A reopen fails, then ethtool -L
wakes the queues, then a direct close() skips netif_tx_disable() because
opened is false. open() then runs with TX awake under an active qdisc.
The final patch in the series removes that second path. The first-open
generic XDP path remains at the end of the series, because ibmveth_open()
still starts RX before ibmveth_alloc_tx_resources() and does not stop TX
first.
Would allocating the TX LTBs before registering the LAN and enabling RX,
or stopping TX before RX is enabled, close this window?
> @@ -1338,11 +1389,16 @@ static int ibmveth_open(struct net_device *netdev)
>
> return 0;
>
> -out_free_buffer_pools:
> +out_cleanup_rx_interrupts:
> + ibmveth_cleanup_rx_interrupts(adapter);
> + /* As in close(): a poll past napi_complete_done() may still run. */
> + synchronize_net();
> +out_unregister_lan:
> + do {
> + lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
> + } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
> ibmveth_free_buffer_pools(adapter);
[Severity: Medium]
This isn't a bug introduced by this patch. The final h_free_logical_lan()
result is thrown away here. The RX pools, RX queue, buffer list and filter
list are then unmapped and freed regardless.
ibmveth_close() does the same thing and only logs the failure:
if (lpar_rc != H_SUCCESS) {
netdev_err(adapter->netdev,
"h_free_logical_lan failed with %lx, continuing\n",
lpar_rc);
}
If deregistration fails with a non-busy error, could the hypervisor still
be using those buffers after they are unmapped? The answer depends on PHYP
behaviour that the driver code alone cannot confirm.
This patch improves the pool-fail case, which previously made no
deregistration call at all. The later ibmveth_free_all_queues() helper
also documents that callers continue with RX pool and DMA teardown
afterward.
[ ... ]
> @@ -1812,6 +1871,12 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
> int i, queue_num = skb_get_queue_mapping(skb);
> unsigned long mss = 0;
>
> + /* Close / failed reopen can free LTBs while IFF_UP is still set. */
> + if (unlikely(!adapter->tx_ltb_ptr[queue_num])) {
> + netdev->stats.tx_dropped++;
[Severity: Low]
This is a pre-existing pattern and was not introduced by this patch.
netdev->stats.tx_dropped++ is a non-atomic update to a counter that all TX
queues share, and each queue is serialised only by its own txq lock.
Can drops on different queues at the same time lose increments here?
tx_packets, tx_bytes and tx_dropped already use the same pattern elsewhere
in ibmveth_start_xmit() and ibmveth_is_packet_unsupported(). The later
patch "ibmveth: Add per-queue RX and TX statistics collection" moves these
drops to adapter->tx_qstats[queue_num].dropped_packets.
> + goto out;
> + }
[Severity: Low]
This check appears to fix a NULL pointer dereference that exists before
this patch. The patch is described as a refactor and carries no Fixes:
tag.
Before this patch, the sequence is:
ibmveth_change_mtu() / pool_store() / set_csum_offload() / set_tso()
ibmveth_close()
ibmveth_open()
fails (TX LTB kzalloc or DMA map, buffer pool alloc, or
h_register_logical_lan())
out_free_tx_ltb: frees every tx_ltb_ptr[]
IFF_UP stays set. ibmveth_set_channels() checks only IFF_UP. It allocates
LTBs for [old, goal), or none at all when shrinking, and then calls
netif_tx_wake_all_queues().
The next ibmveth_start_xmit() on queue 0 then does a memcpy() into the
NULL tx_ltb_ptr[0].
The v7 changelog points to a separate [PATCH net] series carrying Fixes:
d43732ce021f for a related open-fail fix. Should this guard be handled
the same way, so that the crash fix reaches stable trees?
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 07/15] ibmveth: Add RX queue register helpers for MQ
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (5 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 06/15] ibmveth: Refactor TX resource allocation in open/close paths Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-25 18:38 ` [PATCH net-next v7 08/15] ibmveth: Add queue-aware RX buffer submit helper " Mingming Cao
` (8 subsequent siblings)
15 siblings, 0 replies; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
MQ RX changes the RX queue lifecycle from one adapter-level
register/free pair into a mixed model: queue 0 is registered through
h_register_logical_lan*(), queues 1..N through
H_REG_LOGICAL_LAN_QUEUE. The queue-0 control plane that open() and
close() use today has to come out into helpers first, so the
subordinate paths have somewhere to attach when MQ RX is enabled.
Extract logical LAN registration and queue teardown into dedicated
helpers:
ibmveth_register_logical_lan()
ibmveth_register_rx_queues()
ibmveth_free_all_queues()
Runtime stays single-queue here. adapter->multi_queue is still false
until MQ RX is enabled, so queue 0 keeps the legacy
h_register_logical_lan() flow, the with_handle arm is not yet
reachable, and no subordinate queue is registered.
ibmveth_register_logical_lan() registers queue 0 with PHYP. When
multi-queue mode is active it uses h_register_logical_lan_with_handle()
and stores the handle only on H_SUCCESS, so a failed attempt leaves no
stale handle behind.
ibmveth_register_rx_queues() is the open()-side entry point: it builds
queue 0's buffer descriptor, records queue 0's virq in queue_irq[0],
masks that IRQ before registration, and calls
ibmveth_register_logical_lan(). It registers queue 0 only.
ibmveth_free_all_queues() issues one H_FREE_LOGICAL_LAN and clears all
queue handles. One hypercall is enough because H_FREE_LOGICAL_LAN is a
full teardown: per PAPR/PHYP it drops the primary LAN and any
subordinate queues registered under it. Close and the open-fail unwind
both rely on that. Incremental scale-down cannot, and issues
H_FREE_LOGICAL_LAN_QUEUE per queue instead; that path arrives with
resize.
open() allocates buffer pools before PHYP registration so a pool-fail
path never has a live LAN. Post-register errors still unwind through
ibmveth_free_all_queues().
open() also calls netif_set_real_num_rx_queues() with num_rx_queues
after registration (one queue until multi-queue is enabled); a failure
there unwinds the same way.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- no P07 code change
- commit message: name netif_set_real_num_rx_queues()
in open()
Changes in v6:
- drop the eight hypercall counters, struct ibmveth_hcall_stats, and
the adapter field. Only reg_lan and free_lan had writers here;
none of the eight had a reader. The ethtool -S keys v5 added in
patch 11 are gone with them. This patch no longer touches ibmveth.h
- trace LAN registration and the free retry with netdev_dbg() (open/
close/resize, not per packet). Teardown already uses netdev_err()
- document free_all_queues() log-and-continue in its kdoc, including
that queue_handle[] is cleared whatever the hypercall returned
- noted: queue 0 handle is stored only on H_SUCCESS
- noted: body drops the out_unregister_queues walk and the
unconditional PHYP-vs-pools guarantee
Changes in v5:
- Document H_FREE_LOGICAL_LAN full-teardown contract in kdoc/changelog:
one free_lan drops primary + any subordinate queues
(H_FREE_LOGICAL_LAN_QUEUE is for incremental scale-down only) - v4
helper text did not spell the subordinate semantics
- Call out close() update_rx_no_buffer() before free_lan as intentional
last-glimpse accounting while the LAN is still registered (order
already in v4; v4 step list omitted it); NULL-safe per-queue form
lands with MQ enablement
- opened / rx_irq_setup idempotent close is owned by the IRQ-helper
patch (same failed-reopen / second napi_disable hang raised here)
- Poll re-arm during teardown not claimed here (same race as IRQ patch;
lands with poll harden)
Changes in v4:
- Introduce register/free helpers in the same patch that wires their
first open/close callers; keep subordinate-only helpers deferred
until MQ enablement.
- Introduce adapter->hcall_stats here (first use); not in patch 2.
- Correct open/close unwind so free_all_queues() precedes
free_buffer_pools().
- Drop the orphaned big-bang "open/close pipeline" patch from v3; that
wiring is incremental across helper patches 3-7 instead.
drivers/net/ethernet/ibm/ibmveth.c | 174 +++++++++++++++++++++--------
1 file changed, 128 insertions(+), 46 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 011082db1e08..cc01b03d0e4b 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -1294,9 +1294,14 @@ static void ibmveth_free_tx_resources(struct ibmveth_adapter *adapter)
}
static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
- union ibmveth_buf_desc rxq_desc, u64 mac_address)
+ union ibmveth_buf_desc rxq_desc,
+ u64 mac_address)
{
int rc, try_again = 1;
+ unsigned long ua = adapter->vdev->unit_address;
+ unsigned long buf_dma = adapter->buffer_list_dma[0];
+ unsigned long filter_dma = adapter->filter_list_dma;
+ unsigned long qh0;
/*
* After a kexec the adapter will still be open, so our attempt to
@@ -1304,14 +1309,29 @@ static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
* try again, but only once.
*/
retry:
- rc = h_register_logical_lan(adapter->vdev->unit_address,
- adapter->buffer_list_dma[0], rxq_desc.desc,
- adapter->filter_list_dma, mac_address);
+ /* In multi-queue mode, obtain a queue handle for queue 0 so all RX
+ * queues can use the same per-queue buffer hypercalls.
+ */
+ if (adapter->multi_queue) {
+ rc = h_register_logical_lan_with_handle(ua, buf_dma,
+ rxq_desc.desc,
+ filter_dma,
+ mac_address,
+ &qh0);
+ if (rc == H_SUCCESS)
+ adapter->queue_handle[0] = qh0;
+ } else {
+ rc = h_register_logical_lan(ua, buf_dma, rxq_desc.desc,
+ filter_dma, mac_address);
+ }
+ netdev_dbg(adapter->netdev, "h_register_logical_lan%s rc=%d\n",
+ adapter->multi_queue ? "_with_handle" : "", rc);
if (rc != H_SUCCESS && try_again) {
do {
rc = h_free_logical_lan(adapter->vdev->unit_address);
} while (H_IS_LONG_BUSY(rc) || (rc == H_BUSY));
+ netdev_dbg(adapter->netdev, "h_free_logical_lan rc=%d\n", rc);
try_again = 0;
goto retry;
@@ -1320,14 +1340,97 @@ static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
return rc;
}
+/**
+ * ibmveth_free_all_queues - Free all RX queues at once
+ * @adapter: ibmveth adapter structure
+ *
+ * Issues one H_FREE_LOGICAL_LAN for full adapter teardown. Per PAPR/PHYP,
+ * that drops the primary LAN and any subordinate queues registered under
+ * it. Incremental scale-down uses H_FREE_LOGICAL_LAN_QUEUE per queue
+ * instead; do not use this helper for partial live-set shrink.
+ *
+ * Used during interface close and registration error cleanup.
+ *
+ * Retries only H_BUSY and H_IS_LONG_BUSY. On other failures, logs and
+ * returns; callers cannot observe hypercall status. queue_handle[] is
+ * cleared regardless. Callers still run RX pool and DMA teardown
+ * afterward (same as pre-helper close()).
+ *
+ * Clears queue handles only; queue_irq[] is released by
+ * ibmveth_cleanup_rx_interrupts().
+ */
+static void ibmveth_free_all_queues(struct ibmveth_adapter *adapter)
+{
+ unsigned long lpar_rc;
+ int i;
+
+ netdev_dbg(adapter->netdev, "freeing all RX queues at once\n");
+
+ do {
+ lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
+ } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
+
+ if (lpar_rc != H_SUCCESS) {
+ netdev_err(adapter->netdev,
+ "h_free_logical_lan failed: %ld\n", lpar_rc);
+ }
+
+ for (i = 0; i < adapter->num_rx_queues; i++)
+ adapter->queue_handle[i] = 0;
+}
+
+/**
+ * ibmveth_register_rx_queues - Register RX queues with hypervisor
+ * @adapter: ibmveth adapter structure
+ * @mac_address: MAC address for device registration
+ *
+ * Registers queue 0 via ibmveth_register_logical_lan(). Subordinate queue
+ * registration is added when multi-queue RX is enabled.
+ *
+ * Return: 0 on success, -ENONET if queue 0 registration fails
+ */
+static int
+ibmveth_register_rx_queues(struct ibmveth_adapter *adapter, u64 mac_address)
+{
+ struct net_device *netdev = adapter->netdev;
+ union ibmveth_buf_desc rxq_desc;
+ unsigned long lpar_rc;
+ int rc;
+
+ rxq_desc.fields.flags_len = IBMVETH_BUF_VALID |
+ adapter->rx_queue[0].queue_len;
+ rxq_desc.fields.address = adapter->rx_queue[0].queue_dma;
+ adapter->queue_irq[0] = netdev->irq;
+
+ rc = ibmveth_disable_irq(adapter, 0);
+ if (rc)
+ netdev_dbg(netdev,
+ "Failed to disable IRQ for queue 0 before registration, rc=%d\n",
+ rc);
+
+ lpar_rc = ibmveth_register_logical_lan(adapter, rxq_desc, mac_address);
+ if (lpar_rc != H_SUCCESS) {
+ netdev_err(netdev,
+ "h_register_logical_lan failed: %ld\n", lpar_rc);
+ netdev_err(netdev,
+ "buffer TCE:0x%llx filter TCE:0x%llx rxq desc:0x%llx MAC:0x%llx\n",
+ adapter->buffer_list_dma[0],
+ adapter->filter_list_dma,
+ rxq_desc.desc, mac_address);
+ return -ENONET;
+ }
+
+ netdev_dbg(netdev,
+ "registered 1 RX queue with hypervisor (single-queue mode)\n");
+ return 0;
+}
+
static int ibmveth_open(struct net_device *netdev)
{
struct ibmveth_adapter *adapter = netdev_priv(netdev);
- u64 mac_address;
+ u64 mac_address = ether_addr_to_u64(netdev->dev_addr);
int rxq_entries = 1;
- unsigned long lpar_rc;
int rc;
- union ibmveth_buf_desc rxq_desc;
int i;
netdev_dbg(netdev, "open starting\n");
@@ -1343,37 +1446,23 @@ static int ibmveth_open(struct net_device *netdev)
if (rc)
goto out_free_filter_list;
- mac_address = ether_addr_to_u64(netdev->dev_addr);
-
- rxq_desc.fields.flags_len = IBMVETH_BUF_VALID |
- adapter->rx_queue[0].queue_len;
- rxq_desc.fields.address = adapter->rx_queue[0].queue_dma;
-
- adapter->queue_irq[0] = netdev->irq;
- ibmveth_disable_irq(adapter, 0);
-
- lpar_rc = ibmveth_register_logical_lan(adapter, rxq_desc, mac_address);
-
- if (lpar_rc != H_SUCCESS) {
- netdev_err(netdev, "h_register_logical_lan failed with %ld\n",
- lpar_rc);
- netdev_err(netdev, "buffer TCE:0x%llx filter TCE:0x%llx rxq "
- "desc:0x%llx MAC:0x%llx\n",
- adapter->buffer_list_dma[0],
- adapter->filter_list_dma,
- rxq_desc.desc,
- mac_address);
- rc = -ENONET;
+ rc = ibmveth_alloc_buffer_pools(adapter);
+ if (rc)
goto out_free_queue_mem;
- }
- rc = ibmveth_alloc_buffer_pools(adapter);
+ rc = ibmveth_register_rx_queues(adapter, mac_address);
if (rc)
- goto out_unregister_lan;
+ goto out_free_buffer_pools;
+
+ rc = netif_set_real_num_rx_queues(netdev, adapter->num_rx_queues);
+ if (rc) {
+ netdev_err(netdev, "failed to set number of rx queues\n");
+ goto out_unregister_queues;
+ }
rc = ibmveth_setup_rx_interrupts(adapter);
if (rc)
- goto out_unregister_lan;
+ goto out_free_all_queues; /* setup already disposed IRQs */
netdev_dbg(netdev, "initial replenish cycle\n");
ibmveth_schedule_rx_queue(adapter, 0);
@@ -1393,10 +1482,12 @@ static int ibmveth_open(struct net_device *netdev)
ibmveth_cleanup_rx_interrupts(adapter);
/* As in close(): a poll past napi_complete_done() may still run. */
synchronize_net();
-out_unregister_lan:
- do {
- lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
- } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
+ goto out_free_all_queues; /* cleanup already disposed IRQs */
+out_unregister_queues:
+ ibmveth_dispose_subordinate_irq_mappings(adapter);
+out_free_all_queues:
+ ibmveth_free_all_queues(adapter);
+out_free_buffer_pools:
ibmveth_free_buffer_pools(adapter);
out_free_queue_mem:
ibmveth_cleanup_rx_resources(adapter);
@@ -1409,7 +1500,6 @@ static int ibmveth_open(struct net_device *netdev)
static int ibmveth_close(struct net_device *netdev)
{
struct ibmveth_adapter *adapter = netdev_priv(netdev);
- long lpar_rc;
/* Gate on opened, not IFF_UP: pool_store/change_mtu close+open can
* leave IFF_UP set after a failed reopen.
@@ -1431,15 +1521,7 @@ static int ibmveth_close(struct net_device *netdev)
synchronize_net();
ibmveth_update_rx_no_buffer(adapter);
- /* Full LAN teardown (subordinates arrive with register helpers). */
- do {
- lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
- } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
- if (lpar_rc != H_SUCCESS) {
- netdev_err(adapter->netdev,
- "h_free_logical_lan failed with %lx, continuing\n",
- lpar_rc);
- }
+ ibmveth_free_all_queues(adapter);
/* Free TX LTBs after quiesce and after H_FREE_LOGICAL_LAN so xmit
* cannot touch unmapped bounce buffers while the LAN is live.
*/
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* [PATCH net-next v7 08/15] ibmveth: Add queue-aware RX buffer submit helper for MQ
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (6 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 07/15] ibmveth: Add RX queue register helpers for MQ Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 09/15] ibmveth: Harden RX poll path with helpers Mingming Cao
` (7 subsequent siblings)
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
Replenish is the last open-path hypercall that still assumes queue 0.
Registration is already queue-aware, and multi-queue posts buffers
through H_ADD_LOGICAL_LAN_BUFFERS_QUEUE against adapter->queue_handle[],
but ibmveth_replenish_buffer_pool() still calls
h_add_logical_lan_buffer() or h_add_logical_lan_buffers() directly.
Add ibmveth_add_logical_lan_buffers() to route RX buffer pool
replenishment through H_ADD_LOGICAL_LAN_BUFFERS_QUEUE in multi-queue
mode, falling back to the legacy 8-buffer and single-buffer hypercalls
in single-queue mode.
Supporting that helper means the RX side can no longer assume queue 0.
The queue index is 0 everywhere until multi-queue RX is enabled, so
the queue plumbing below does not change behaviour yet. The correlator
and locking changes do take effect immediately, on the single queue.
Parameterise the RX accessors by queue. The eight ibmveth_rxq_*
helpers took only the adapter and hardcoded rx_queue[0]; they now take
a queue_index, as do ibmveth_remove_buffer_from_pool() and
ibmveth_rxq_get_buffer().
Add a per-queue replenish_lock to struct ibmveth_rx_q, initialised in
probe. Replenish is not the only writer of a pool's free_map: the
harvest and remove consumer runs from that queue's NAPI instance, and
ndo_poll_controller() runs replenish on the same queue from outside
NAPI, so producer and consumer can run at once on one queue.
Give ibmveth_replenish_buffer_pool() a return value. It was void and
logged from inside the critical section, where a printk can re-enter
replenish through netconsole on the same device. It now returns one of
IBMVETH_REPLENISH_OK, _RESET_MAP, _RESET_MQ, _HCALL_FAIL or
_BATCH_FALLBACK, and ibmveth_replenish_task() does the logging and any
schedule_work() after dropping replenish_lock. The replenish map uses
DMA_ATTR_NO_WARN so the iommu path cannot printk under that lock.
Replace the correlator WARN_ON()s with ibmveth_rxq_correlator_valid().
The pool and buffer index come from a hypervisor-supplied correlator,
which is not a kernel invariant, so WARN_ON() was the wrong tool: with
panic_on_warn set, a malformed correlator would take the partition
down. The helper returns false instead, callers propagate -EINVAL or
-EFAULT, and ibmveth_rxq_advance() still advances the ring so poll
makes progress. On a NULL buffer poll advances past the slot so it
does not spin on an invalid completion at this commit.
remove_buffer_from_pool() no longer schedules the reset itself;
rxq_get_buffer() schedules the reset on a bad correlator at this
commit, and the next patch moves that escalation into poll's skip
helper. The KUnit expectations are updated to match.
Redefine IBMVETH_MAX_RX_PER_HCALL from 8 to 12, the argument-list
capacity of H_ADD_LOGICAL_LAN_BUFFERS_QUEUE, and add
IBMVETH_MAX_RX_REGULAR (8) for the legacy hypercall.
free_buffer_pool() keeps probe/sysfs geometry (active/size/threshold)
and zeroes pool->available so teardown cannot leave a stale count.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- extend remove_buffer_from_pool Return: -EINVAL also
inactive / not live (skbuff or free_map NULL)
- replenish fail log names the wrapper from
filled, not batch
- advance past slot on NULL buffer so poll does not spin
at this commit; P09 skip_bad refactors error
escalation into helper
- noted: irqsave CS and free_buffer_pool vs lock stay
cover leftovers (lock-protocol rewrite, not
two conditions)
Changes in v6:
- drop inline on the queue-index rxq_* accessors (plain static)
- initialise the replenish locks once in probe; open() no longer
reinitialises them in alloc_rx_queues(), where it could reset a
lock another CPU was holding through poll_controller()
- pass DMA_ATTR_NO_WARN to dma_map_single_attrs() in the replenish
loop. The iommu warning could re-enter via netconsole and deadlock
on the replenish_lock this CPU already holds
- static_assert IBMVETH_MAX_RX_REGULAR == 8 and
IBMVETH_MAX_RX_PER_HCALL == 12 against the hand-written
legacy descs[0..7] and MQ ioba[0..5] argument lists
- add clarifying comment on the legacy 8-buffer hcall batch bound
- drop the three buffer-submit hypercall counters along with the rest
of hcall_stats; see patch 7
- fix the KUnit fixtures to allocate a dummy free_map
- noted: irqsave CS and free_buffer_pool vs replenish stay cover
- noted: skip_bad_correlator lands in P09
Changes in v5:
- On MQ buffer-add H_FUNCTION: schedule adapter reset after dropping
replenish_lock (v4 logged/broke with no recovery; can permanently dry
the pool)
- Move replenish fail logging / reset scheduling out from under
replenish_lock so netconsole cannot deadlock re-entering replenish
- Serialize harvest/remove with per-queue replenish_lock (netpoll
replenish vs NAPI consumer; v4 locked producer only)
- Fail logs use real wrapper names: h_add_logical_lan_buffers[_queue] /
h_add_logical_lan_buffer (v4 interpolated broken lan[_queue] strings)
- Document replenish_lock + no-printk-under-lock for netconsole (first
lock use); outcomes enum + ibmveth_replenish_fail defined here
- Defer adapter-global counter atomics and irqsave critical-section
shorten to cover follow-up
- Bad queue_index poll path: napi_complete before return lands with
poll harden (not claimed fully here)
- Keep pool active/size/threshold across free_buffer_pool (probe/sysfs
geometry); only clear runtime allocations + available (ifdown/up
reopen must still see active pools)
- get_buffer: use correlator_valid (drop WARN_ON; keep schedule_work
until poll skip owns reset)
- Introduce ibmveth_rxq_correlator_valid / ibmveth_rxq_advance at first
remove/harvest use; init replenish_lock in remove_buffer KUnit
- On RESET_MAP/RESET_MQ, stop remaining pool walks (goto unlock)
Changes in v4:
- Introduce queue-aware replenish/poll helpers with their first callers
in the same patch; do not leave a 2-arg replenish call ahead of the
signature change.
- Restore the pre-MQ LPM H_FUNCTION break instead of continue; do not
loop forever on a stale local batch size.
- Fold per-queue replenish_lock into this patch.
- Update kdoc for MQ parameters on remove_buffer_from_pool /
rxq_harvest_buffer.
drivers/net/ethernet/ibm/ibmveth.c | 563 +++++++++++++++++++++--------
drivers/net/ethernet/ibm/ibmveth.h | 6 +-
2 files changed, 423 insertions(+), 146 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index cc01b03d0e4b..ed75dea90a95 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -30,6 +30,7 @@
#include <linux/ip.h>
#include <linux/ipv6.h>
#include <linux/slab.h>
+#include <linux/spinlock.h>
#include <asm/hvcall.h>
#include <linux/atomic.h>
#include <asm/vio.h>
@@ -101,49 +102,58 @@ static struct ibmveth_stat ibmveth_stats[] = {
};
/* simple methods of getting data from the current rxq entry */
-static inline u32 ibmveth_rxq_flags(struct ibmveth_adapter *adapter)
+static u32 ibmveth_rxq_flags(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- struct ibmveth_rx_q *rxq = &adapter->rx_queue[0];
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[queue_index];
return be32_to_cpu(rxq->queue_addr[rxq->index].flags_off);
}
-static inline int ibmveth_rxq_toggle(struct ibmveth_adapter *adapter)
+static int ibmveth_rxq_toggle(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- return (ibmveth_rxq_flags(adapter) & IBMVETH_RXQ_TOGGLE) >>
- IBMVETH_RXQ_TOGGLE_SHIFT;
+ return (ibmveth_rxq_flags(adapter, queue_index) & IBMVETH_RXQ_TOGGLE) >>
+ IBMVETH_RXQ_TOGGLE_SHIFT;
}
-static inline int ibmveth_rxq_pending_buffer(struct ibmveth_adapter *adapter)
+static int ibmveth_rxq_pending_buffer(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- return ibmveth_rxq_toggle(adapter) == adapter->rx_queue[0].toggle;
+ return ibmveth_rxq_toggle(adapter, queue_index) ==
+ adapter->rx_queue[queue_index].toggle;
}
-static inline int ibmveth_rxq_buffer_valid(struct ibmveth_adapter *adapter)
+static int ibmveth_rxq_buffer_valid(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- return ibmveth_rxq_flags(adapter) & IBMVETH_RXQ_VALID;
+ return ibmveth_rxq_flags(adapter, queue_index) & IBMVETH_RXQ_VALID;
}
-static inline int ibmveth_rxq_frame_offset(struct ibmveth_adapter *adapter)
+static int ibmveth_rxq_frame_offset(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- return ibmveth_rxq_flags(adapter) & IBMVETH_RXQ_OFF_MASK;
+ return ibmveth_rxq_flags(adapter, queue_index) & IBMVETH_RXQ_OFF_MASK;
}
-static inline int ibmveth_rxq_large_packet(struct ibmveth_adapter *adapter)
+static int ibmveth_rxq_large_packet(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- return ibmveth_rxq_flags(adapter) & IBMVETH_RXQ_LRG_PKT;
+ return ibmveth_rxq_flags(adapter, queue_index) & IBMVETH_RXQ_LRG_PKT;
}
-static inline int ibmveth_rxq_frame_length(struct ibmveth_adapter *adapter)
+static int ibmveth_rxq_frame_length(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- struct ibmveth_rx_q *rxq = &adapter->rx_queue[0];
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[queue_index];
return be32_to_cpu(rxq->queue_addr[rxq->index].length);
}
-static inline int ibmveth_rxq_csum_good(struct ibmveth_adapter *adapter)
+static int ibmveth_rxq_csum_good(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- return ibmveth_rxq_flags(adapter) & IBMVETH_RXQ_CSUM_GOOD;
+ return ibmveth_rxq_flags(adapter, queue_index) & IBMVETH_RXQ_CSUM_GOOD;
}
static unsigned int ibmveth_real_max_tx_queues(void)
@@ -713,11 +723,100 @@ static inline void ibmveth_flush_buffer(void *addr, unsigned long length)
asm("dcbf %0,%1,1" :: "b" (addr), "r" (offset));
}
-/* replenish the buffers for a pool. note that we don't need to
- * skb_reserve these since they are used for incoming...
+/**
+ * ibmveth_add_logical_lan_buffers - Add receive buffers to hypervisor
+ * @adapter: ibmveth adapter structure
+ * @descs: array of buffer descriptors to add
+ * @filled: number of valid descriptors in the array
+ * @buff_size: size of each buffer (multi-queue mode only)
+ * @queue_index: RX queue index
+ *
+ * Return: hypervisor return code
+ */
+static long ibmveth_add_logical_lan_buffers(struct ibmveth_adapter *adapter,
+ union ibmveth_buf_desc *descs,
+ int filled,
+ unsigned long buff_size,
+ int queue_index)
+{
+ struct vio_dev *vdev = adapter->vdev;
+ unsigned long rc;
+
+ /*
+ * The MQ hcall takes six ioba words (12 packed addresses). The
+ * legacy hcall takes eight descriptors. The argument lists below
+ * are written out by hand; keep the defines matched to those lists.
+ */
+ static_assert(IBMVETH_MAX_RX_PER_HCALL == 12);
+ static_assert(IBMVETH_MAX_RX_REGULAR == 8);
+
+ if (adapter->multi_queue) {
+ unsigned long buffersznum = (buff_size << 32) | filled;
+ unsigned long ioba[IBMVETH_MAX_RX_PER_HCALL / 2] = {0};
+ unsigned long handle = adapter->queue_handle[queue_index];
+ int i;
+
+ /* Pack descriptor addresses into ioba pairs.
+ * Each ioba holds two 32-bit addresses packed into 64 bits:
+ * - Even descriptors (0,2,4...) go in high 32 bits
+ * - Odd descriptors (1,3,5...) go in low 32 bits
+ */
+ for (i = 0; i < filled && i < IBMVETH_MAX_RX_PER_HCALL; i++) {
+ int pair_idx = i / 2;
+ int is_high = (i % 2 == 0);
+
+ if (is_high)
+ ioba[pair_idx] = (unsigned long)
+ descs[i].fields.address << 32;
+ else
+ ioba[pair_idx] |= descs[i].fields.address;
+ }
+
+ rc = h_add_logical_lan_buffers_queue(vdev->unit_address,
+ handle,
+ buffersznum,
+ ioba[0], ioba[1], ioba[2],
+ ioba[3], ioba[4], ioba[5]);
+ } else if (filled == 1) {
+ rc = h_add_logical_lan_buffer(vdev->unit_address,
+ descs[0].desc);
+ } else {
+ /* Legacy 8-desc hcall; probe/mq_fallback keep batch <=
+ * IBMVETH_MAX_RX_REGULAR.
+ */
+ rc = h_add_logical_lan_buffers(vdev->unit_address,
+ descs[0].desc, descs[1].desc,
+ descs[2].desc, descs[3].desc,
+ descs[4].desc, descs[5].desc,
+ descs[6].desc, descs[7].desc);
+ }
+
+ return rc;
+}
+
+/* Outcomes for ibmveth_replenish_buffer_pool(); logged after unlock. */
+enum {
+ IBMVETH_REPLENISH_OK = 0,
+ IBMVETH_REPLENISH_RESET_MAP,
+ IBMVETH_REPLENISH_RESET_MQ,
+ IBMVETH_REPLENISH_HCALL_FAIL,
+ IBMVETH_REPLENISH_BATCH_FALLBACK,
+};
+
+struct ibmveth_replenish_fail {
+ unsigned long lpar_rc;
+ u32 filled;
+ u32 batch;
+};
+
+/* Replenish the buffers for a pool.
+ * Caller must hold the per-queue replenish_lock. Do not printk here:
+ * netconsole on the same device can re-enter replenish_task.
*/
-static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
- struct ibmveth_buff_pool *pool)
+static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
+ struct ibmveth_buff_pool *pool,
+ int queue_index,
+ struct ibmveth_replenish_fail *fail)
{
union ibmveth_buf_desc descs[IBMVETH_MAX_RX_PER_HCALL] = {0};
u32 remaining = pool->size - atomic_read(&pool->available);
@@ -729,6 +828,7 @@ static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
dma_addr_t dma_addr;
struct device *dev;
u32 index;
+ int outcome = IBMVETH_REPLENISH_OK;
vdev = adapter->vdev;
dev = &vdev->dev;
@@ -743,12 +843,9 @@ static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
/* Fill a batch of descriptors */
for (filled = 0; filled < min(remaining, batch); filled++) {
index = pool->free_map[free_index];
- if (WARN_ON(index == IBM_VETH_INVALID_MAP)) {
+ if (index == IBM_VETH_INVALID_MAP) {
adapter->replenish_add_buff_failure++;
- netdev_info(adapter->netdev,
- "Invalid map index %u, reset\n",
- index);
- schedule_work(&adapter->work);
+ outcome = IBMVETH_REPLENISH_RESET_MAP;
break;
}
@@ -763,9 +860,15 @@ static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
break;
}
- dma_addr = dma_map_single(dev, skb->data,
- pool->buff_size,
- DMA_FROM_DEVICE);
+ /* NO_WARN: hold replenish_lock; iommu
+ * printk can re-enter via netconsole.
+ */
+ dma_addr =
+ dma_map_single_attrs(dev,
+ skb->data,
+ pool->buff_size,
+ DMA_FROM_DEVICE,
+ DMA_ATTR_NO_WARN);
if (dma_mapping_error(dev, dma_addr)) {
dev_kfree_skb_any(skb);
adapter->replenish_add_buff_failure++;
@@ -800,28 +903,21 @@ static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
free_index = 0;
}
+ if (outcome != IBMVETH_REPLENISH_OK)
+ break;
+
if (!filled)
break;
- /* single buffer case*/
- if (filled == 1)
- lpar_rc = h_add_logical_lan_buffer(vdev->unit_address,
- descs[0].desc);
- else
- /* Multi-buffer hcall */
- lpar_rc = h_add_logical_lan_buffers(vdev->unit_address,
- descs[0].desc,
- descs[1].desc,
- descs[2].desc,
- descs[3].desc,
- descs[4].desc,
- descs[5].desc,
- descs[6].desc,
- descs[7].desc);
+ lpar_rc = ibmveth_add_logical_lan_buffers(adapter, descs,
+ filled,
+ pool->buff_size,
+ queue_index);
+
if (lpar_rc != H_SUCCESS) {
- dev_warn_ratelimited(dev,
- "RX h_add_logical_lan failed: filled=%u, rc=%lu, batch=%u\n",
- filled, lpar_rc, batch);
+ fail->lpar_rc = lpar_rc;
+ fail->filled = filled;
+ fail->batch = batch;
goto hcall_failure;
}
@@ -861,30 +957,35 @@ static void ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
}
adapter->replenish_add_buff_failure += filled;
- /*
- * If multi rx buffers hcall is no longer supported by FW
- * e.g. in the case of Live Partition Migration
- */
- if (batch > 1 && lpar_rc == H_FUNCTION) {
- /*
- * Instead of retry submit single buffer individually
- * here just set the max rx buffer per hcall to 1
- * buffers will be respleshed next time
- * when ibmveth_replenish_buffer_pool() is called again
- * with single-buffer case
- */
- netdev_info(adapter->netdev,
- "RX Multi buffers not supported by FW, rc=%lu\n",
- lpar_rc);
- adapter->rx_buffers_per_hcall = 1;
- netdev_info(adapter->netdev,
- "Next rx replesh will fall back to single-buffer hcall\n");
+ if (lpar_rc == H_FUNCTION) {
+ if (adapter->multi_queue) {
+ /*
+ * LPM / firmware may drop MQ buffer hcalls.
+ * Schedule reset so we do not sit forever in
+ * no-buffer with the link still up.
+ */
+ outcome = IBMVETH_REPLENISH_RESET_MQ;
+ } else if (batch > 1) {
+ /*
+ * Live Partition Migration may drop multi-
+ * buffer support. Fall back to single-buffer
+ * on the next replenish; do not continue with
+ * a stale local batch size (infinite loop).
+ */
+ adapter->rx_buffers_per_hcall = 1;
+ outcome = IBMVETH_REPLENISH_BATCH_FALLBACK;
+ } else {
+ outcome = IBMVETH_REPLENISH_HCALL_FAIL;
+ }
+ } else {
+ outcome = IBMVETH_REPLENISH_HCALL_FAIL;
}
break;
}
mb();
atomic_add(buffers_added, &(pool->available));
+ return outcome;
}
/*
@@ -904,21 +1005,85 @@ static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter)
}
/* replenish routine */
-static void ibmveth_replenish_task(struct ibmveth_adapter *adapter)
+static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- int i;
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[queue_index];
+ struct ibmveth_replenish_fail fail = {};
+ unsigned long flags;
+ int i, rc;
+ int need_reset = 0;
+ int batch_fallback = 0;
+ int hcall_fail = 0;
+
+ if (queue_index >= adapter->num_rx_queues) {
+ netdev_dbg(adapter->netdev,
+ "Skipping replenish for freed queue %d (num_queues=%u)\n",
+ queue_index, adapter->num_rx_queues);
+ return;
+ }
adapter->replenish_task_cycles++;
- for (i = (IBMVETH_NUM_BUFF_POOLS - 1); i >= 0; i--) {
- struct ibmveth_buff_pool *pool = &adapter->rx_buff_pool[0][i];
+ spin_lock_irqsave(&rxq->replenish_lock, flags);
- if (pool->active &&
- (atomic_read(&pool->available) < pool->threshold))
- ibmveth_replenish_buffer_pool(adapter, pool);
+ for (i = (IBMVETH_NUM_BUFF_POOLS - 1); i >= 0; i--) {
+ struct ibmveth_buff_pool *pool =
+ &adapter->rx_buff_pool[queue_index][i];
+
+ if (pool->active && pool->free_map &&
+ (atomic_read(&pool->available) < pool->threshold)) {
+ rc = ibmveth_replenish_buffer_pool(adapter, pool,
+ queue_index, &fail);
+ switch (rc) {
+ case IBMVETH_REPLENISH_RESET_MAP:
+ case IBMVETH_REPLENISH_RESET_MQ:
+ need_reset = rc;
+ goto out_unlock;
+ case IBMVETH_REPLENISH_BATCH_FALLBACK:
+ batch_fallback = 1;
+ break;
+ case IBMVETH_REPLENISH_HCALL_FAIL:
+ hcall_fail = 1;
+ break;
+ default:
+ break;
+ }
+ }
}
+out_unlock:
ibmveth_update_rx_no_buffer(adapter);
+
+ spin_unlock_irqrestore(&rxq->replenish_lock, flags);
+
+ /* Log and schedule reset only after dropping replenish_lock. */
+ if (need_reset == IBMVETH_REPLENISH_RESET_MAP) {
+ netdev_info(adapter->netdev,
+ "Invalid RX free_map entry on queue %d, reset\n",
+ queue_index);
+ schedule_work(&adapter->work);
+ } else if (need_reset == IBMVETH_REPLENISH_RESET_MQ) {
+ dev_err_ratelimited(&adapter->netdev->dev,
+ "MQ buffer add H_FUNCTION (q=%d, batch=%u), reset\n",
+ queue_index, fail.batch);
+ schedule_work(&adapter->work);
+ }
+
+ if (batch_fallback)
+ dev_warn_ratelimited(&adapter->netdev->dev,
+ "Legacy batch add H_FUNCTION (batch=%u), fallback\n",
+ fail.batch);
+
+ if (hcall_fail)
+ dev_warn_ratelimited(&adapter->netdev->dev,
+ "RX %s failed: filled=%u, rc=%lu, batch=%u\n",
+ adapter->multi_queue ?
+ "h_add_logical_lan_buffers_queue" :
+ (fail.filled == 1 ?
+ "h_add_logical_lan_buffer" :
+ "h_add_logical_lan_buffers"),
+ fail.filled, fail.lpar_rc, fail.batch);
}
/* empty and free ana buffer pool - also used to do cleanup in error paths */
@@ -953,6 +1118,12 @@ static void ibmveth_free_buffer_pool(struct ibmveth_adapter *adapter,
kfree(pool->skbuff);
pool->skbuff = NULL;
}
+
+ /*
+ * Keep probe/sysfs geometry (active, size, buff_size, threshold).
+ * Only tear down runtime allocations; open reuses active pools.
+ */
+ atomic_set(&pool->available, 0);
}
/**
@@ -1093,35 +1264,75 @@ ibmveth_free_buffer_pools(struct ibmveth_adapter *adapter)
adapter->num_rx_queues);
}
+static bool ibmveth_rxq_correlator_valid(struct ibmveth_adapter *adapter,
+ int queue_index, u64 correlator)
+{
+ unsigned int pool = correlator >> 32;
+ unsigned int index = correlator & 0xffffffffUL;
+ struct ibmveth_buff_pool *bpool;
+
+ if (pool >= IBMVETH_NUM_BUFF_POOLS)
+ return false;
+
+ bpool = &adapter->rx_buff_pool[queue_index][pool];
+
+ /* Require a live pool with allocated arrays before indexing.
+ * Inactive pools still have size from init; free clears skbuff.
+ */
+ if (!bpool->active || !bpool->skbuff || !bpool->free_map)
+ return false;
+
+ return index < bpool->size;
+}
+
+static void ibmveth_rxq_advance(struct ibmveth_rx_q *rxq)
+{
+ if (++rxq->index == rxq->num_slots) {
+ rxq->index = 0;
+ rxq->toggle = !rxq->toggle;
+ }
+}
+
/**
* ibmveth_remove_buffer_from_pool - remove a buffer from a pool
* @adapter: adapter instance
* @correlator: identifies pool and index
+ * @queue_index: RX queue index (0..num_rx_queues-1)
* @reuse: whether to reuse buffer
*
+ * Context: may run concurrently with netpoll replenish_task on the same
+ * queue; takes per-queue replenish_lock to serialize free_map /
+ * producer_index / available against the producer.
+ *
* Return:
* * %0 - success
- * * %-EINVAL - correlator maps to pool or index out of range
+ * * %-EINVAL - pool or index out of range, or the pool is
+ * inactive / not live (skbuff or free_map NULL)
* * %-EFAULT - pool and index map to null skb
*/
static int ibmveth_remove_buffer_from_pool(struct ibmveth_adapter *adapter,
- u64 correlator, bool reuse)
+ u64 correlator, int queue_index,
+ bool reuse)
{
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[queue_index];
unsigned int pool = correlator >> 32;
unsigned int index = correlator & 0xffffffffUL;
unsigned int free_index;
struct sk_buff *skb;
+ unsigned long flags;
+ int rc = 0;
- if (WARN_ON(pool >= IBMVETH_NUM_BUFF_POOLS) ||
- WARN_ON(index >= adapter->rx_buff_pool[0][pool].size)) {
- schedule_work(&adapter->work);
- return -EINVAL;
+ spin_lock_irqsave(&rxq->replenish_lock, flags);
+
+ if (!ibmveth_rxq_correlator_valid(adapter, queue_index, correlator)) {
+ rc = -EINVAL;
+ goto out_unlock;
}
- skb = adapter->rx_buff_pool[0][pool].skbuff[index];
- if (WARN_ON(!skb)) {
- schedule_work(&adapter->work);
- return -EFAULT;
+ skb = adapter->rx_buff_pool[queue_index][pool].skbuff[index];
+ if (!skb) {
+ rc = -EFAULT;
+ goto out_unlock;
}
/* if we are going to reuse the buffer then keep the pointers around
@@ -1132,75 +1343,88 @@ static int ibmveth_remove_buffer_from_pool(struct ibmveth_adapter *adapter,
/* remove the skb pointer to mark free. actual freeing is done
* by upper level networking after gro_receive
*/
- adapter->rx_buff_pool[0][pool].skbuff[index] = NULL;
+ struct ibmveth_buff_pool *bpool =
+ &adapter->rx_buff_pool[queue_index][pool];
+
+ bpool->skbuff[index] = NULL;
dma_unmap_single(&adapter->vdev->dev,
- adapter->rx_buff_pool[0][pool].dma_addr[index],
- adapter->rx_buff_pool[0][pool].buff_size,
+ bpool->dma_addr[index],
+ bpool->buff_size,
DMA_FROM_DEVICE);
}
- free_index = adapter->rx_buff_pool[0][pool].producer_index;
- adapter->rx_buff_pool[0][pool].producer_index++;
- if (adapter->rx_buff_pool[0][pool].producer_index >=
- adapter->rx_buff_pool[0][pool].size)
- adapter->rx_buff_pool[0][pool].producer_index = 0;
- adapter->rx_buff_pool[0][pool].free_map[free_index] = index;
+ free_index = adapter->rx_buff_pool[queue_index][pool].producer_index;
+ adapter->rx_buff_pool[queue_index][pool].producer_index++;
+ if (adapter->rx_buff_pool[queue_index][pool].producer_index >=
+ adapter->rx_buff_pool[queue_index][pool].size)
+ adapter->rx_buff_pool[queue_index][pool].producer_index = 0;
+ adapter->rx_buff_pool[queue_index][pool].free_map[free_index] = index;
mb();
- atomic_dec(&adapter->rx_buff_pool[0][pool].available);
+ atomic_dec(&adapter->rx_buff_pool[queue_index][pool].available);
- return 0;
+out_unlock:
+ spin_unlock_irqrestore(&rxq->replenish_lock, flags);
+ return rc;
}
/* get the current buffer on the rx queue */
-static inline struct sk_buff *ibmveth_rxq_get_buffer(struct ibmveth_adapter *adapter)
+static struct sk_buff *
+ibmveth_rxq_get_buffer(struct ibmveth_adapter *adapter,
+ int queue_index)
{
- struct ibmveth_rx_q *rxq = &adapter->rx_queue[0];
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[queue_index];
u64 correlator = rxq->queue_addr[rxq->index].correlator;
unsigned int pool = correlator >> 32;
unsigned int index = correlator & 0xffffffffUL;
- if (WARN_ON(pool >= IBMVETH_NUM_BUFF_POOLS) ||
- WARN_ON(index >= adapter->rx_buff_pool[0][pool].size)) {
+ if (!ibmveth_rxq_correlator_valid(adapter, queue_index, correlator)) {
schedule_work(&adapter->work);
return NULL;
}
- return adapter->rx_buff_pool[0][pool].skbuff[index];
+ return adapter->rx_buff_pool[queue_index][pool].skbuff[index];
}
/**
* ibmveth_rxq_harvest_buffer - Harvest buffer from pool
*
* @adapter: pointer to adapter
+ * @queue_index: RX queue index to harvest from
* @reuse: whether to reuse buffer
*
* Context: called from ibmveth_poll
*
+ * On a bad correlator (-EINVAL/-EFAULT) the ring is still advanced so poll
+ * cannot spin forever on one slot. The error is still returned: callers must
+ * not treat it as a successful take from the pool (especially reuse=false,
+ * which would hand the SKB to the stack while it remains pool-owned).
+ *
* Return:
- * * %0 - success
- * * other - non-zero return from ibmveth_remove_buffer_from_pool
+ * * %0 - buffer removed from pool (or marked for reuse) and ring advanced
+ * * other - non-zero return from ibmveth_remove_buffer_from_pool; ring has
+ * still been advanced for -EINVAL/-EFAULT
*/
static int ibmveth_rxq_harvest_buffer(struct ibmveth_adapter *adapter,
- bool reuse)
+ int queue_index, bool reuse)
{
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[queue_index];
u64 cor;
int rc;
- struct ibmveth_rx_q *rxq = &adapter->rx_queue[0];
-
cor = rxq->queue_addr[rxq->index].correlator;
- rc = ibmveth_remove_buffer_from_pool(adapter, cor, reuse);
- if (unlikely(rc))
+ rc = ibmveth_remove_buffer_from_pool(adapter, cor, queue_index, reuse);
+ if (unlikely(rc)) {
+ /* Skip a corrupt slot without claiming pool ownership. */
+ if (rc == -EINVAL || rc == -EFAULT)
+ ibmveth_rxq_advance(rxq);
return rc;
-
- if (++adapter->rx_queue[0].index == adapter->rx_queue[0].num_slots) {
- adapter->rx_queue[0].index = 0;
- adapter->rx_queue[0].toggle = !adapter->rx_queue[0].toggle;
}
+ ibmveth_rxq_advance(rxq);
+
return 0;
}
@@ -2168,36 +2392,48 @@ static void ibmveth_rx_csum_helper(struct sk_buff *skb,
static int ibmveth_poll(struct napi_struct *napi, int budget)
{
- struct ibmveth_adapter *adapter =
- container_of(napi, struct ibmveth_adapter, napi[0]);
- struct net_device *netdev = adapter->netdev;
+ struct net_device *netdev = napi->dev;
+ struct ibmveth_adapter *adapter = netdev_priv(netdev);
int frames_processed = 0;
- int rc;
+ int queue_index, rc;
u16 mss = 0;
+ queue_index = napi - adapter->napi;
+
restart_poll:
while (frames_processed < budget) {
- if (!ibmveth_rxq_pending_buffer(adapter))
+ if (!ibmveth_rxq_pending_buffer(adapter, queue_index))
break;
smp_rmb();
- if (!ibmveth_rxq_buffer_valid(adapter)) {
+ if (!ibmveth_rxq_buffer_valid(adapter, queue_index)) {
wmb(); /* suggested by larson1 */
adapter->rx_invalid_buffer++;
netdev_dbg(netdev, "recycling invalid buffer\n");
- if (unlikely(ibmveth_rxq_harvest_buffer(adapter, true)))
+ rc = ibmveth_rxq_harvest_buffer(adapter,
+ queue_index, true);
+ if (unlikely(rc))
break;
} else {
struct sk_buff *skb, *new_skb;
- int length = ibmveth_rxq_frame_length(adapter);
- int offset = ibmveth_rxq_frame_offset(adapter);
- int csum_good = ibmveth_rxq_csum_good(adapter);
- int lrg_pkt = ibmveth_rxq_large_packet(adapter);
+ int length = ibmveth_rxq_frame_length(adapter,
+ queue_index);
+ int offset = ibmveth_rxq_frame_offset(adapter,
+ queue_index);
+ int csum_good = ibmveth_rxq_csum_good(adapter,
+ queue_index);
+ int lrg_pkt = ibmveth_rxq_large_packet(adapter,
+ queue_index);
__sum16 iph_check = 0;
- skb = ibmveth_rxq_get_buffer(adapter);
- if (unlikely(!skb))
+ skb = ibmveth_rxq_get_buffer(adapter, queue_index);
+ if (unlikely(!skb)) {
+ struct ibmveth_rx_q *rxq =
+ &adapter->rx_queue[queue_index];
+
+ ibmveth_rxq_advance(rxq);
break;
+ }
/* if the large packet bit is set in the rx queue
* descriptor, the mss will be written by PHYP eight
@@ -2220,12 +2456,18 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
length);
if (rx_flush)
ibmveth_flush_buffer(skb->data,
- length + offset);
- if (unlikely(ibmveth_rxq_harvest_buffer(adapter, true)))
+ length + offset);
+ rc = ibmveth_rxq_harvest_buffer(adapter,
+ queue_index,
+ true);
+ if (unlikely(rc))
break;
skb = new_skb;
} else {
- if (unlikely(ibmveth_rxq_harvest_buffer(adapter, false)))
+ rc = ibmveth_rxq_harvest_buffer(adapter,
+ queue_index,
+ false);
+ if (unlikely(rc))
break;
skb_reserve(skb, offset);
}
@@ -2261,7 +2503,7 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
}
}
- ibmveth_replenish_task(adapter);
+ ibmveth_replenish_task(adapter, queue_index);
if (frames_processed == budget)
goto out;
@@ -2272,14 +2514,18 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
/* We think we are done - reenable interrupts,
* then check once more to make sure we are done.
*/
- rc = ibmveth_enable_irq(adapter, 0);
+ rc = ibmveth_enable_irq(adapter, queue_index);
if (rc) {
+ netdev_err(netdev,
+ "Failed to enable IRQ for queue %d (rc=%d), scheduling reset\n",
+ queue_index, rc);
schedule_work(&adapter->work);
goto out;
}
- if (ibmveth_rxq_pending_buffer(adapter) && napi_schedule(napi)) {
- ibmveth_disable_irq(adapter, 0);
+ if (ibmveth_rxq_pending_buffer(adapter, queue_index) &&
+ napi_schedule(napi)) {
+ ibmveth_disable_irq(adapter, queue_index);
goto restart_poll;
}
@@ -2409,7 +2655,7 @@ static void ibmveth_poll_controller(struct net_device *dev)
{
struct ibmveth_adapter *adapter = netdev_priv(dev);
- ibmveth_replenish_task(adapter);
+ ibmveth_replenish_task(adapter, 0);
ibmveth_schedule_rx_queue(adapter, 0);
}
#endif
@@ -2567,6 +2813,16 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
adapter->vdev = dev;
adapter->netdev = netdev;
INIT_WORK(&adapter->work, ibmveth_reset);
+
+ /*
+ * Initialise the replenish locks once. open() is re-entered on
+ * MTU and offload changes without netpoll_poll_disable(), so a
+ * lock set up there could be reinitialised while poll_controller()
+ * holds it.
+ */
+ for (i = 0; i < IBMVETH_MAX_RX_QUEUES; i++)
+ spin_lock_init(&adapter->rx_queue[i].replenish_lock);
+
adapter->mcastFilterSize = be32_to_cpu(*mcastFilterSize_p);
ibmveth_init_link_settings(netdev);
@@ -2608,7 +2864,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
if (ret == H_SUCCESS &&
(ret_attr & IBMVETH_ILLAN_RX_MULTI_BUFF_SUPPORT)) {
- adapter->rx_buffers_per_hcall = IBMVETH_MAX_RX_PER_HCALL;
+ adapter->rx_buffers_per_hcall = IBMVETH_MAX_RX_REGULAR;
netdev_dbg(netdev,
"RX Multi-buffer hcall supported by FW, batch set to %u\n",
adapter->rx_buffers_per_hcall);
@@ -2929,8 +3185,7 @@ static void ibmveth_reset_kunit(struct work_struct *w)
* @test: pointer to kunit structure
*
* Tests the error returns from ibmveth_remove_buffer_from_pool.
- * ibmveth_remove_buffer_from_pool also calls WARN_ON, so dmesg should be
- * checked to see that these warnings happened.
+ * Bad correlators return -EINVAL/-EFAULT (no WARN_ON).
*
* Return: void
*/
@@ -2944,6 +3199,8 @@ static void ibmveth_remove_buffer_from_pool_test(struct kunit *test)
INIT_WORK(&adapter->work, ibmveth_reset_kunit);
+ spin_lock_init(&adapter->rx_queue[0].replenish_lock);
+
/* Set sane values for buffer pools */
for (int i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
ibmveth_init_buffer_pool(&adapter->rx_buff_pool[0][i], i,
@@ -2953,19 +3210,34 @@ static void ibmveth_remove_buffer_from_pool_test(struct kunit *test)
pool = &adapter->rx_buff_pool[0][0];
pool->skbuff = kunit_kcalloc(test, pool->size, sizeof(void *), GFP_KERNEL);
KUNIT_ASSERT_NOT_ERR_OR_NULL(test, pool->skbuff);
+ pool->free_map = kunit_kcalloc(test, pool->size, sizeof(u16),
+ GFP_KERNEL);
+ KUNIT_ASSERT_NOT_ERR_OR_NULL(test, pool->free_map);
correlator = ((u64)IBMVETH_NUM_BUFF_POOLS << 32) | 0;
- KUNIT_EXPECT_EQ(test, -EINVAL, ibmveth_remove_buffer_from_pool(adapter, correlator, false));
- KUNIT_EXPECT_EQ(test, -EINVAL, ibmveth_remove_buffer_from_pool(adapter, correlator, true));
+ KUNIT_EXPECT_EQ(test, -EINVAL,
+ ibmveth_remove_buffer_from_pool(adapter,
+ correlator, 0, false));
+ KUNIT_EXPECT_EQ(test, -EINVAL,
+ ibmveth_remove_buffer_from_pool(adapter,
+ correlator, 0, true));
correlator = ((u64)0 << 32) | adapter->rx_buff_pool[0][0].size;
- KUNIT_EXPECT_EQ(test, -EINVAL, ibmveth_remove_buffer_from_pool(adapter, correlator, false));
- KUNIT_EXPECT_EQ(test, -EINVAL, ibmveth_remove_buffer_from_pool(adapter, correlator, true));
+ KUNIT_EXPECT_EQ(test, -EINVAL,
+ ibmveth_remove_buffer_from_pool(adapter,
+ correlator, 0, false));
+ KUNIT_EXPECT_EQ(test, -EINVAL,
+ ibmveth_remove_buffer_from_pool(adapter,
+ correlator, 0, true));
correlator = (u64)0 | 0;
pool->skbuff[0] = NULL;
- KUNIT_EXPECT_EQ(test, -EFAULT, ibmveth_remove_buffer_from_pool(adapter, correlator, false));
- KUNIT_EXPECT_EQ(test, -EFAULT, ibmveth_remove_buffer_from_pool(adapter, correlator, true));
+ KUNIT_EXPECT_EQ(test, -EFAULT,
+ ibmveth_remove_buffer_from_pool(adapter,
+ correlator, 0, false));
+ KUNIT_EXPECT_EQ(test, -EFAULT,
+ ibmveth_remove_buffer_from_pool(adapter,
+ correlator, 0, true));
flush_work(&adapter->work);
}
@@ -2974,9 +3246,7 @@ static void ibmveth_remove_buffer_from_pool_test(struct kunit *test)
* ibmveth_rxq_get_buffer_test - unit test for ibmveth_rxq_get_buffer
* @test: pointer to kunit structure
*
- * Tests ibmveth_rxq_get_buffer. ibmveth_rxq_get_buffer also calls WARN_ON for
- * the NULL returns, so dmesg should be checked to see that these warnings
- * happened.
+ * Tests ibmveth_rxq_get_buffer invalid correlator returns NULL without WARN.
*
* Return: void
*/
@@ -3007,18 +3277,21 @@ static void ibmveth_rxq_get_buffer_test(struct kunit *test)
pool = &adapter->rx_buff_pool[0][0];
pool->skbuff = kunit_kcalloc(test, pool->size, sizeof(void *), GFP_KERNEL);
KUNIT_ASSERT_NOT_ERR_OR_NULL(test, pool->skbuff);
+ pool->free_map = kunit_kcalloc(test, pool->size, sizeof(u16),
+ GFP_KERNEL);
+ KUNIT_ASSERT_NOT_ERR_OR_NULL(test, pool->free_map);
adapter->rx_queue[0].queue_addr[0].correlator =
(u64)IBMVETH_NUM_BUFF_POOLS << 32 | 0;
- KUNIT_EXPECT_PTR_EQ(test, NULL, ibmveth_rxq_get_buffer(adapter));
+ KUNIT_EXPECT_PTR_EQ(test, NULL, ibmveth_rxq_get_buffer(adapter, 0));
adapter->rx_queue[0].queue_addr[0].correlator =
(u64)0 << 32 | adapter->rx_buff_pool[0][0].size;
- KUNIT_EXPECT_PTR_EQ(test, NULL, ibmveth_rxq_get_buffer(adapter));
+ KUNIT_EXPECT_PTR_EQ(test, NULL, ibmveth_rxq_get_buffer(adapter, 0));
pool->skbuff[0] = skb;
adapter->rx_queue[0].queue_addr[0].correlator = (u64)0 << 32 | 0;
- KUNIT_EXPECT_PTR_EQ(test, skb, ibmveth_rxq_get_buffer(adapter));
+ KUNIT_EXPECT_PTR_EQ(test, skb, ibmveth_rxq_get_buffer(adapter, 0));
flush_work(&adapter->work);
}
diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h
index a3178032c6f2..dc0e5c25209a 100644
--- a/drivers/net/ethernet/ibm/ibmveth.h
+++ b/drivers/net/ethernet/ibm/ibmveth.h
@@ -14,6 +14,8 @@
#ifndef _IBMVETH_H
#define _IBMVETH_H
+#include <linux/spinlock_types.h>
+
/* constants for H_MULTICAST_CTRL */
#define IbmVethMcastReceptionModifyBit 0x80000UL
#define IbmVethMcastReceptionEnableBit 0x20000UL
@@ -265,7 +267,8 @@ static inline long h_illan_attributes(unsigned long unit_address,
#define IBMVETH_DEFAULT_QUEUES 8U
#define IBMVETH_MAX_RX_QUEUES 1U
#define IBMVETH_DEFAULT_RX_QUEUES 1U
-#define IBMVETH_MAX_RX_PER_HCALL 8U
+#define IBMVETH_MAX_RX_REGULAR 8U
+#define IBMVETH_MAX_RX_PER_HCALL 12U
static int pool_size[] = { 512, 1024 * 2, 1024 * 16, 1024 * 32, 1024 * 64 };
static int pool_count[] = { 256, 512, 256, 256, 256 };
@@ -296,6 +299,7 @@ struct ibmveth_rx_q {
dma_addr_t queue_dma;
u32 queue_len;
struct ibmveth_rx_q_entry *queue_addr;
+ spinlock_t replenish_lock; /* per-queue buffer replenish */
};
struct ibmveth_adapter {
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 08/15] ibmveth: Add queue-aware RX buffer submit helper for MQ
2026-09-25 18:38 ` [PATCH net-next v7 08/15] ibmveth: Add queue-aware RX buffer submit helper " Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found 5 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 3 · Low: 2 (3)
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] This patch fixes three real bugs that already exist in the
tree, but it presents itself as a feature ('Add queue-aware RX buffer…
- [Medium] This patch removes schedule_work(&adapter->work) from the
-EINVAL and -EFAULT exits of ibmveth_remove_buffer_from_pool().
- [Medium] ibmveth_replenish_task() now holds
spin_lock_irqsave(&rxq->replenish_lock) across the whole walk of all…
- [Low] ibmveth_replenish_task() fills a single struct
ibmveth_replenish_fail across every pool it walks.
- [Low] Reported concern: ibmveth_close() frees the buffer pools
(free_map, skbuff, dma_addr, available) and the buffer-list pages…
Pre-existing issues:
- [Low] Reported concern: on the rx_copybreak path of ibmveth_poll(),
new_skb is leaked when ibmveth_rxq_harvest_buffer() fails after the…
- [Low] Reported concern: adapter->replenish_task_cycles++ runs before
replenish_lock is taken, so concurrent NAPI and netpoll replenish…
- [Low] Reported concern: the frame offset and length supplied by the
hypervisor are not checked against the RX buffer size before…
--
ibmveth: Add queue-aware RX buffer submit helper for MQ
This patch adds ibmveth_add_logical_lan_buffers() so RX buffer pool
replenishment goes through H_ADD_LOGICAL_LAN_BUFFERS_QUEUE in multi-queue
mode, and gives the RX accessors a queue index. It also adds a per-queue
replenish_lock, moves replenish logging out of the critical section, and
replaces the correlator WARN_ON()s with ibmveth_rxq_correlator_valid().
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index cc01b03d0e4b0..ed75dea90a957 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
[ ... ]
> @@ -904,21 +1005,85 @@ static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter)
> }
>
> /* replenish routine */
> -static void ibmveth_replenish_task(struct ibmveth_adapter *adapter)
> +static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
> + int queue_index)
> {
[ ... ]
> adapter->replenish_task_cycles++;
[Severity: Low]
This isn't a bug introduced by this patch, but
adapter->replenish_task_cycles++ still runs before replenish_lock is
taken. NAPI and ndo_poll_controller() can both call
ibmveth_replenish_task() on the same queue. Can increments of this
counter be lost when they do?
A later patch in the series ("ibmveth: Add per-queue RX and TX statistics
collection") moves the per-queue replenish_task_cycles increment under
the queue's replenish_lock. That means only the intermediate commits are
affected, and only the statistics.
>
> - for (i = (IBMVETH_NUM_BUFF_POOLS - 1); i >= 0; i--) {
> - struct ibmveth_buff_pool *pool = &adapter->rx_buff_pool[0][i];
> + spin_lock_irqsave(&rxq->replenish_lock, flags);
[Severity: Medium]
This lock is held with hard IRQs off for the whole walk over all
IBMVETH_NUM_BUFF_POOLS pools. For each pool, ibmveth_replenish_buffer_pool()
loops on while (remaining > 0) until the full deficit is filled. That
work includes netdev_alloc_skb(), dma_map_single_attrs(), optional dcbf
flushes, and one hcall per batch.
On the first replenish after open every pool is empty. With the default
active pools (256 + 512 + 256 buffers, batch 8) that comes to roughly
1024 allocations and mappings and about 128 hcalls with IRQs off. Larger
pools set through sysfs, or the fallback to single-buffer hcalls, make
this much longer.
ibmveth_remove_buffer_from_pool() takes the same lock with irqsave from
NAPI, so a concurrent harvest on the same queue would spin with IRQs off
for that whole time.
Before this patch the same work ran in NAPI softirq with IRQs enabled.
Is this IRQ-off latency acceptable, or could the lock be dropped between
batches? The v7 notes list the irqsave section as a leftover, and it
keeps this shape at the end of the series.
>
> - if (pool->active &&
> - (atomic_read(&pool->available) < pool->threshold))
> - ibmveth_replenish_buffer_pool(adapter, pool);
> + for (i = (IBMVETH_NUM_BUFF_POOLS - 1); i >= 0; i--) {
> + struct ibmveth_buff_pool *pool =
> + &adapter->rx_buff_pool[queue_index][i];
> +
> + if (pool->active && pool->free_map &&
> + (atomic_read(&pool->available) < pool->threshold)) {
[Severity: Low]
Can this pool->free_map check race with ibmveth_close()?
ibmveth_close() calls ibmveth_free_buffer_pools() without taking
replenish_lock. Meanwhile ibmveth_poll_controller() ->
ibmveth_replenish_task() can run from netpoll. If that happens,
free_map, skbuff and dma_addr could be freed right after this check
passes.
The last patch in the series ("ibmveth: Complete set_channels down-path
and mq_fallback max_rx cap") closes this window. With it,
ibmveth_poll_controller() returns early when !adapter->opened, and
ibmveth_close() clears opened and calls synchronize_net() before freeing
the pools. Netpoll callers run with IRQs disabled, so that
synchronize_net() waits for them. That guard does not exist yet at this
commit.
> + rc = ibmveth_replenish_buffer_pool(adapter, pool,
> + queue_index, &fail);
> + switch (rc) {
> + case IBMVETH_REPLENISH_RESET_MAP:
> + case IBMVETH_REPLENISH_RESET_MQ:
> + need_reset = rc;
> + goto out_unlock;
> + case IBMVETH_REPLENISH_BATCH_FALLBACK:
> + batch_fallback = 1;
> + break;
> + case IBMVETH_REPLENISH_HCALL_FAIL:
> + hcall_fail = 1;
> + break;
> + default:
> + break;
> + }
> + }
> }
[ ... ]
> + if (batch_fallback)
> + dev_warn_ratelimited(&adapter->netdev->dev,
> + "Legacy batch add H_FUNCTION (batch=%u), fallback\n",
> + fail.batch);
> +
> + if (hcall_fail)
> + dev_warn_ratelimited(&adapter->netdev->dev,
> + "RX %s failed: filled=%u, rc=%lu, batch=%u\n",
> + adapter->multi_queue ?
> + "h_add_logical_lan_buffers_queue" :
> + (fail.filled == 1 ?
> + "h_add_logical_lan_buffer" :
> + "h_add_logical_lan_buffers"),
> + fail.filled, fail.lpar_rc, fail.batch);
> }
[Severity: Low]
Can these two messages print the wrong values? All pools in the loop
share one fail record. The BATCH_FALLBACK and HCALL_FAIL cases only break
out of the switch, so later pools still run. Each hcall failure makes
ibmveth_replenish_buffer_pool() overwrite fail->lpar_rc, fail->filled and
fail->batch.
Say one pool hits BATCH_FALLBACK with batch=8, which sets
rx_buffers_per_hcall to 1, and a later pool then has a hcall failure.
The "Legacy batch add H_FUNCTION (batch=%u), fallback" message would
print batch=1.
In the reverse order, the "RX %s failed" message would show the rc,
filled count and wrapper name of the H_FUNCTION fallback, not those of
the pool that actually failed.
The same structure is still there at the end of the series.
[ ... ]
> @@ -1093,35 +1264,75 @@ ibmveth_free_buffer_pools(struct ibmveth_adapter *adapter)
> adapter->num_rx_queues);
> }
>
> +static bool ibmveth_rxq_correlator_valid(struct ibmveth_adapter *adapter,
> + int queue_index, u64 correlator)
> +{
> + unsigned int pool = correlator >> 32;
> + unsigned int index = correlator & 0xffffffffUL;
> + struct ibmveth_buff_pool *bpool;
> +
> + if (pool >= IBMVETH_NUM_BUFF_POOLS)
> + return false;
> +
> + bpool = &adapter->rx_buff_pool[queue_index][pool];
> +
> + /* Require a live pool with allocated arrays before indexing.
> + * Inactive pools still have size from init; free clears skbuff.
> + */
> + if (!bpool->active || !bpool->skbuff || !bpool->free_map)
> + return false;
> +
> + return index < bpool->size;
> +}
[Severity: Medium]
This helper, plus the removed WARN_ON()s and the ring advance on a NULL
buffer in ibmveth_poll(), fixes three bugs that exist in the current
tree:
- ibmveth_remove_buffer_from_pool() and ibmveth_rxq_get_buffer() call
WARN_ON() on a hypervisor-supplied correlator. With panic_on_warn set,
that takes the partition down.
- In ibmveth_poll(), a valid correlator whose skbuff[index] is NULL used
to break without advancing the ring and without scheduling a reset.
The poll tail then calls ibmveth_enable_irq(), sees
ibmveth_rxq_pending_buffer() still true, calls napi_schedule() and
jumps to restart_poll. It polls the same slot forever in softirq.
- A correlator that names an inactive pool (pool_active[] =
{1,1,0,0,1}) passes the old size-only check. skbuff[index] is then
read through a NULL skbuff array.
These fixes sit inside a patch titled as a multi-queue feature. Neither
this patch nor any later one in the series has a Fixes: tag. Could they
be split out into a separate patch with Fixes: tags, so stable can take
them without the MQ refactor?
[ ... ]
> static int ibmveth_remove_buffer_from_pool(struct ibmveth_adapter *adapter,
> - u64 correlator, bool reuse)
> + u64 correlator, int queue_index,
> + bool reuse)
> {
[ ... ]
> - if (WARN_ON(pool >= IBMVETH_NUM_BUFF_POOLS) ||
> - WARN_ON(index >= adapter->rx_buff_pool[0][pool].size)) {
> - schedule_work(&adapter->work);
> - return -EINVAL;
> + spin_lock_irqsave(&rxq->replenish_lock, flags);
> +
> + if (!ibmveth_rxq_correlator_valid(adapter, queue_index, correlator)) {
> + rc = -EINVAL;
> + goto out_unlock;
> }
>
> - skb = adapter->rx_buff_pool[0][pool].skbuff[index];
> - if (WARN_ON(!skb)) {
> - schedule_work(&adapter->work);
> - return -EFAULT;
> + skb = adapter->rx_buff_pool[queue_index][pool].skbuff[index];
> + if (!skb) {
> + rc = -EFAULT;
> + goto out_unlock;
> }
[Severity: Medium]
Is dropping schedule_work() from both error exits intended? The commit
message says rxq_get_buffer() still schedules the reset. The
recycle-invalid branch of ibmveth_poll() never calls
ibmveth_rxq_get_buffer(), though:
ibmveth_poll()
if (!ibmveth_rxq_buffer_valid(adapter, queue_index))
ibmveth_rxq_harvest_buffer(adapter, queue_index, true)
ibmveth_remove_buffer_from_pool() returns -EINVAL/-EFAULT
ibmveth_rxq_advance(rxq)
break;
So a descriptor without IBMVETH_RXQ_VALID that carries a malformed
correlator is now skipped and no reset is scheduled.
The buffer PHYP actually consumed keeps its skbuff[] entry and its
IBM_VETH_INVALID_MAP free_map slot, and pool->available still counts it.
ibmveth_replenish_task() only refills when available < threshold. Can
repeated events like this leave PHYP short of RX buffers, with nothing
to resync the pool until an unrelated reset or ifdown/up?
The problem is still there at the end of the series.
ibmveth_poll_recycle_invalid() -> ibmveth_poll_harvest_slot() treats
-EINVAL/-EFAULT as handled and does not escalate. Only
ibmveth_poll_skip_bad_correlator(), on the valid-descriptor path, calls
schedule_work().
[ ... ]
> @@ -2168,36 +2392,48 @@ static void ibmveth_rx_csum_helper(struct sk_buff *skb,
[ ... ]
> } else {
> struct sk_buff *skb, *new_skb;
> - int length = ibmveth_rxq_frame_length(adapter);
> - int offset = ibmveth_rxq_frame_offset(adapter);
> - int csum_good = ibmveth_rxq_csum_good(adapter);
> - int lrg_pkt = ibmveth_rxq_large_packet(adapter);
> + int length = ibmveth_rxq_frame_length(adapter,
> + queue_index);
> + int offset = ibmveth_rxq_frame_offset(adapter,
> + queue_index);
[Severity: Low]
This isn't a bug introduced by this patch, but length and offset come
straight from the hypervisor descriptor. Nothing checks them against
the RX buffer size before skb_copy_to_linear_data(),
ibmveth_flush_buffer(), skb_reserve() and skb_put() later in
ibmveth_poll(). Could a bad descriptor cause an out-of-bounds read or
an skb_over_panic()?
A later patch in the series ("ibmveth: Harden RX poll path with
helpers") fixes this. It adds a check in ibmveth_poll_deliver_frame()
against min(skb_tailroom(skb), buff_size) that rejects
off >= room || len > room - off.
[ ... ]
> @@ -2220,12 +2456,18 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
> length);
> if (rx_flush)
> ibmveth_flush_buffer(skb->data,
> - length + offset);
> - if (unlikely(ibmveth_rxq_harvest_buffer(adapter, true)))
> + length + offset);
> + rc = ibmveth_rxq_harvest_buffer(adapter,
> + queue_index,
> + true);
> + if (unlikely(rc))
> break;
> skb = new_skb;
[Severity: Low]
This isn't a bug introduced by this patch, but on this rx_copybreak
path new_skb has already been allocated and filled. If
ibmveth_rxq_harvest_buffer() fails, the break skips it. Does this leak
new_skb?
At this commit the harvest can only fail here if pool state changed
after ibmveth_rxq_get_buffer() succeeded, so the window is small. A
later patch in the series fixes it by calling kfree_skb(new_skb) on the
error path in ibmveth_poll_deliver_frame().
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 09/15] ibmveth: Harden RX poll path with helpers
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (7 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 08/15] ibmveth: Add queue-aware RX buffer submit helper " Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 10/15] ibmveth: Enable multi-queue RX receive path Mingming Cao
` (6 subsequent siblings)
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
ibmveth_poll() must handle several distinct skip/fail outcomes on each
RX slot, not only the happy path:
- interface close / napi_disable must not re-arm PHYP delivery
- bad correlators and harvest errors must not look like successful GRO
- oversize or wrap-around offset+length must not skb_put() past the
buffer
- an out-of-range queue index must napi_complete_done rather than
fall through and keep polling
Doing all of that inline turns the NAPI callback into a deeply nested
switchyard. Split each outcome into a small helper so ibmveth_poll()
stays a thin budget loop:
ibmveth_poll_stopping()
ibmveth_poll_harvest_slot() / recycle_invalid / skip_bad_correlator
ibmveth_poll_drop_oversize()
ibmveth_poll_deliver_frame()
ibmveth_poll_bump_invalid()
Not pure motion: ibmveth_poll_bump_invalid() also counts oversize
frames and skipped slots (intentional). skip_bad_correlator escalates
a valid correlator with a NULL skb (-EFAULT) to reset, not only an
out-of-range correlator. Deliver also rejects a PHYP offset+length
beyond min(skb_tailroom, pool->buff_size);
skb_tailroom is the skb_put bound, the DMA
map is pool->buff_size. Skipped and dropped slots do not
count against the NAPI budget; only a delivered frame does.
Two further behaviour changes come with the split. On the rx_copybreak
path a harvest failure, after the frame has already been copied into
new_skb, used to break out of the loop and leak that skb;
deliver_frame() kfree_skb()s it before returning an error. And mss
becomes a local of deliver_frame() rather than living across
ibmveth_poll()'s whole budget loop; it was never read stale, because
ibmveth_rx_mss_helper() only uses it under lrg_pkt, so that part is
scoping hygiene and not a fix.
Drop schedule_work from rxq_get_buffer() (skip_bad_correlator owns
reset escalation). ibmveth_poll_stopping() ensures close/napi_disable
does not re-arm PHYP. Runtime is still single-queue: helpers are defined
and called from the existing SQ poll path in this same patch (first
use). The next patch turns on MQ and reuses this loop; subordinate
register helpers stay there.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- bound oversize against min(skb_tailroom, pool->buff_size);
name that bound in the commit message
- noted: poll_stopping stays mask-first; skip vs budget stays;
cancel_work stays in P10
Changes in v6:
- read the IPv4 header check through skb->data, as mainline does.
v5 used ip_hdr(), which reads the network header offset that
eth_type_trans() does not set on this path
- do not complete NAPI or consume budget when budget is 0. That is
the netpoll case, where v5 completed and returned budget - 1,
underflowing to -1
- noted: the split is not pure motion; the body lists what changed
Changes in v5:
- Wrap poll_stopping return after complete so it stays under 80 cols
- On poll_stopping after the budget loop: return < budget after
napi_complete_done (was frames_processed-1 even when under budget)
- New in v5: peel SQ poll harden + helpers before MQ enable so tip P10
stays bring-up focused (mailed v4 09/14 MQ enable -> tip P10; also
14->15). Kitchen-sink / enable-path poll fixes land here on the live
single-queue path first.
- After napi_complete_done, check poll_stopping again before enable_irq
(avoid re-arm while resize/close waits on napi_disable)
drivers/net/ethernet/ibm/ibmveth.c | 297 ++++++++++++++++++++---------
1 file changed, 202 insertions(+), 95 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index ed75dea90a95..7b6c0283e5c3 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -1380,10 +1380,8 @@ ibmveth_rxq_get_buffer(struct ibmveth_adapter *adapter,
unsigned int pool = correlator >> 32;
unsigned int index = correlator & 0xffffffffUL;
- if (!ibmveth_rxq_correlator_valid(adapter, queue_index, correlator)) {
- schedule_work(&adapter->work);
+ if (!ibmveth_rxq_correlator_valid(adapter, queue_index, correlator))
return NULL;
- }
return adapter->rx_buff_pool[queue_index][pool].skbuff[index];
}
@@ -2390,130 +2388,239 @@ static void ibmveth_rx_csum_helper(struct sk_buff *skb,
}
}
+static void ibmveth_poll_bump_invalid(struct ibmveth_adapter *adapter,
+ int queue_index)
+{
+ adapter->rx_invalid_buffer++;
+}
+
+static bool ibmveth_poll_stopping(struct net_device *netdev,
+ struct napi_struct *napi)
+{
+ return !netif_running(netdev) || napi_disable_pending(napi);
+}
+
+static bool ibmveth_poll_harvest_slot(struct ibmveth_adapter *adapter,
+ int queue_index, bool reuse)
+{
+ int rc = ibmveth_rxq_harvest_buffer(adapter, queue_index, reuse);
+
+ return !rc || rc == -EINVAL || rc == -EFAULT;
+}
+
+static bool ibmveth_poll_recycle_invalid(struct net_device *netdev,
+ struct ibmveth_adapter *adapter,
+ int queue_index)
+{
+ netdev_dbg(netdev, "recycling invalid buffer\n");
+ ibmveth_poll_bump_invalid(adapter, queue_index);
+ return ibmveth_poll_harvest_slot(adapter, queue_index, true);
+}
+
+static bool ibmveth_poll_skip_bad_correlator(struct net_device *netdev,
+ struct ibmveth_adapter *adapter,
+ int queue_index)
+{
+ if (net_ratelimit())
+ netdev_err(netdev,
+ "bad correlator on queue %d, skipping slot\n",
+ queue_index);
+ /* Residual stale slot after resize: recover via reset rather
+ * than spinning forever. Always escalate; only the log is
+ * rate-limited.
+ */
+ schedule_work(&adapter->work);
+ ibmveth_poll_bump_invalid(adapter, queue_index);
+ return ibmveth_poll_harvest_slot(adapter, queue_index, true);
+}
+
+static bool ibmveth_poll_drop_oversize(struct net_device *netdev,
+ struct ibmveth_adapter *adapter,
+ int queue_index, unsigned int off,
+ unsigned int len, unsigned int room)
+{
+ if (net_ratelimit())
+ netdev_err(netdev,
+ "RX frame %u+%u exceeds buffer %u on queue %d, dropping\n",
+ off, len, room, queue_index);
+ ibmveth_poll_bump_invalid(adapter, queue_index);
+ return ibmveth_poll_harvest_slot(adapter, queue_index, true);
+}
+
+/**
+ * ibmveth_poll_deliver_frame - Build SKB from one valid RX slot and GRO it
+ * @napi: NAPI context for this RX queue
+ * @adapter: ibmveth adapter
+ * @netdev: net_device for @adapter
+ * @queue_index: RX queue index
+ *
+ * Return: 1 frame delivered, 0 if the slot was skipped cleanly, -1 on error.
+ */
+static int ibmveth_poll_deliver_frame(struct napi_struct *napi,
+ struct ibmveth_adapter *adapter,
+ struct net_device *netdev,
+ int queue_index)
+{
+ struct sk_buff *skb, *new_skb;
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[queue_index];
+ unsigned int room, off, len, pool;
+ int length, offset, csum_good, lrg_pkt;
+ __sum16 iph_check = 0;
+ u16 mss = 0;
+ int rc;
+
+ length = ibmveth_rxq_frame_length(adapter, queue_index);
+ offset = ibmveth_rxq_frame_offset(adapter, queue_index);
+ csum_good = ibmveth_rxq_csum_good(adapter, queue_index);
+ lrg_pkt = ibmveth_rxq_large_packet(adapter, queue_index);
+
+ skb = ibmveth_rxq_get_buffer(adapter, queue_index);
+ if (unlikely(!skb)) {
+ if (!ibmveth_poll_skip_bad_correlator(netdev, adapter,
+ queue_index))
+ return -1;
+ return 0;
+ }
+
+ pool = rxq->queue_addr[rxq->index].correlator >> 32;
+ room = min_t(unsigned int, skb_tailroom(skb),
+ adapter->rx_buff_pool[queue_index][pool].buff_size);
+ off = offset;
+ len = length;
+ if (unlikely(off >= room || len > room - off)) {
+ if (!ibmveth_poll_drop_oversize(netdev, adapter, queue_index,
+ off, len, room))
+ return -1;
+ return 0;
+ }
+
+ if (lrg_pkt) {
+ __be64 *rxmss = (__be64 *)(skb->data + 8);
+
+ mss = (u16)be64_to_cpu(*rxmss);
+ }
+
+ new_skb = NULL;
+ if (length < rx_copybreak)
+ new_skb = netdev_alloc_skb(netdev, length);
+
+ if (new_skb) {
+ skb_copy_to_linear_data(new_skb, skb->data + offset, length);
+ if (rx_flush)
+ ibmveth_flush_buffer(skb->data, length + offset);
+ rc = ibmveth_rxq_harvest_buffer(adapter, queue_index, true);
+ if (unlikely(rc)) {
+ kfree_skb(new_skb);
+ return -1;
+ }
+ skb = new_skb;
+ } else {
+ rc = ibmveth_rxq_harvest_buffer(adapter, queue_index, false);
+ if (unlikely(rc))
+ return -1;
+ skb_reserve(skb, offset);
+ }
+
+ skb_put(skb, length);
+ skb->protocol = eth_type_trans(skb, netdev);
+
+ if (skb->protocol == cpu_to_be16(ETH_P_IP))
+ iph_check = ((struct iphdr *)skb->data)->check;
+
+ if ((length > netdev->mtu + ETH_HLEN) || lrg_pkt ||
+ iph_check == 0xffff) {
+ ibmveth_rx_mss_helper(skb, mss, lrg_pkt);
+ adapter->rx_large_packets++;
+ }
+
+ if (csum_good) {
+ skb->ip_summed = CHECKSUM_UNNECESSARY;
+ ibmveth_rx_csum_helper(skb, adapter);
+ }
+
+ napi_gro_receive(napi, skb);
+
+ netdev->stats.rx_packets++;
+ netdev->stats.rx_bytes += length;
+
+ return 1;
+}
+
static int ibmveth_poll(struct napi_struct *napi, int budget)
{
struct net_device *netdev = napi->dev;
struct ibmveth_adapter *adapter = netdev_priv(netdev);
int frames_processed = 0;
int queue_index, rc;
- u16 mss = 0;
queue_index = napi - adapter->napi;
+ if (WARN_ON(queue_index < 0 ||
+ queue_index >= adapter->num_rx_queues)) {
+ if (budget)
+ napi_complete_done(napi, 0);
+ return 0;
+ }
+
+ if (ibmveth_poll_stopping(netdev, napi)) {
+ if (budget)
+ napi_complete_done(napi, 0);
+ return 0;
+ }
+
restart_poll:
while (frames_processed < budget) {
+ if (ibmveth_poll_stopping(netdev, napi))
+ break;
+
if (!ibmveth_rxq_pending_buffer(adapter, queue_index))
break;
smp_rmb();
if (!ibmveth_rxq_buffer_valid(adapter, queue_index)) {
wmb(); /* suggested by larson1 */
- adapter->rx_invalid_buffer++;
- netdev_dbg(netdev, "recycling invalid buffer\n");
- rc = ibmveth_rxq_harvest_buffer(adapter,
- queue_index, true);
- if (unlikely(rc))
+ if (!ibmveth_poll_recycle_invalid(netdev, adapter,
+ queue_index))
break;
} else {
- struct sk_buff *skb, *new_skb;
- int length = ibmveth_rxq_frame_length(adapter,
- queue_index);
- int offset = ibmveth_rxq_frame_offset(adapter,
- queue_index);
- int csum_good = ibmveth_rxq_csum_good(adapter,
- queue_index);
- int lrg_pkt = ibmveth_rxq_large_packet(adapter,
- queue_index);
- __sum16 iph_check = 0;
-
- skb = ibmveth_rxq_get_buffer(adapter, queue_index);
- if (unlikely(!skb)) {
- struct ibmveth_rx_q *rxq =
- &adapter->rx_queue[queue_index];
-
- ibmveth_rxq_advance(rxq);
+ rc = ibmveth_poll_deliver_frame(napi, adapter, netdev,
+ queue_index);
+ if (rc < 0)
break;
- }
-
- /* if the large packet bit is set in the rx queue
- * descriptor, the mss will be written by PHYP eight
- * bytes from the start of the rx buffer, which is
- * skb->data at this stage
- */
- if (lrg_pkt) {
- __be64 *rxmss = (__be64 *)(skb->data + 8);
-
- mss = (u16)be64_to_cpu(*rxmss);
- }
-
- new_skb = NULL;
- if (length < rx_copybreak)
- new_skb = netdev_alloc_skb(netdev, length);
-
- if (new_skb) {
- skb_copy_to_linear_data(new_skb,
- skb->data + offset,
- length);
- if (rx_flush)
- ibmveth_flush_buffer(skb->data,
- length + offset);
- rc = ibmveth_rxq_harvest_buffer(adapter,
- queue_index,
- true);
- if (unlikely(rc))
- break;
- skb = new_skb;
- } else {
- rc = ibmveth_rxq_harvest_buffer(adapter,
- queue_index,
- false);
- if (unlikely(rc))
- break;
- skb_reserve(skb, offset);
- }
-
- skb_put(skb, length);
- skb->protocol = eth_type_trans(skb, netdev);
-
- /* PHYP without PLSO support places a -1 in the ip
- * checksum for large send frames.
- */
- if (skb->protocol == cpu_to_be16(ETH_P_IP)) {
- struct iphdr *iph = (struct iphdr *)skb->data;
-
- iph_check = iph->check;
- }
-
- if ((length > netdev->mtu + ETH_HLEN) ||
- lrg_pkt || iph_check == 0xffff) {
- ibmveth_rx_mss_helper(skb, mss, lrg_pkt);
- adapter->rx_large_packets++;
- }
-
- if (csum_good) {
- skb->ip_summed = CHECKSUM_UNNECESSARY;
- ibmveth_rx_csum_helper(skb, adapter);
- }
-
- napi_gro_receive(napi, skb); /* send it up */
-
- netdev->stats.rx_packets++;
- netdev->stats.rx_bytes += length;
- frames_processed++;
+ if (rc > 0)
+ frames_processed++;
}
}
ibmveth_replenish_task(adapter, queue_index);
+ if (ibmveth_poll_stopping(netdev, napi)) {
+ /* budget 0 is netpoll, which must not complete NAPI.
+ * Otherwise returning budget after completing would ask
+ * NAPI to reschedule, so cap the return at budget - 1.
+ */
+ if (budget) {
+ napi_complete_done(napi, frames_processed);
+ return min(frames_processed, budget - 1);
+ }
+ return 0;
+ }
+
if (frames_processed == budget)
goto out;
if (!napi_complete_done(napi, frames_processed))
goto out;
- /* We think we are done - reenable interrupts,
- * then check once more to make sure we are done.
+ /*
+ * napi_disable() sets DISABLE then waits for this poll. Without a
+ * second stopping check here, enable_irq() can re-arm PHYP after
+ * resize already masked the queue; late IRQs then hit the handler
+ * after num_rx_queues was published lower (lab WARN at interrupt).
*/
+ if (ibmveth_poll_stopping(netdev, napi))
+ goto out;
+
rc = ibmveth_enable_irq(adapter, queue_index);
if (rc) {
netdev_err(netdev,
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 09/15] ibmveth: Harden RX poll path with helpers
2026-09-25 18:38 ` [PATCH net-next v7 09/15] ibmveth: Harden RX poll path with helpers Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found 5 potential
issue(s) to consider.
Critical: 0 · High: 2 (2) · Medium: 1 · Low: 2
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] The re-arm guard in ibmveth_poll()
(drivers/net/ethernet/ibm/ibmveth.c) is a check-then-act race, so it…
- [High] net: ibmveth: kernel crash via OOB array read due to TOCTOU
double fetch
- [Medium] Reset escalation for bad correlators now covers only one of
the three harvest callers in ibmveth_poll().
- [Low] The new comment above the post-completion ibmveth_poll_stopping()
check gives a rationale the code does not support.
- [Low] When the RX code was moved into ibmveth_poll_deliver_frame(), two
comments that document PHYP behaviour were dropped.
Pre-existing issues:
- [High] Crafted IPv4 frames with the checksum-offload flag and IHL=0 can
crash the receiving partition.
- [High] The concern: ibmveth_remove() calls
cancel_work_sync(&adapter->work) before unregister_netdev().
--
ibmveth: Harden RX poll path with helpers
This splits the per-slot handling in ibmveth_poll() into small helpers
for invalid slots, bad correlators, oversize frames and frame delivery.
It adds stopping checks meant to keep close and napi_disable from
re-arming PHYP. It also bounds the PHYP offset+length against the RX
buffer, and moves bad-correlator reset escalation out of
ibmveth_rxq_get_buffer().
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index ed75dea90a957..7b6c0283e5c32 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
[ ... ]
> +static bool ibmveth_poll_harvest_slot(struct ibmveth_adapter *adapter,
> + int queue_index, bool reuse)
> +{
> + int rc = ibmveth_rxq_harvest_buffer(adapter, queue_index, reuse);
> +
> + return !rc || rc == -EINVAL || rc == -EFAULT;
> +}
> +
> +static bool ibmveth_poll_recycle_invalid(struct net_device *netdev,
> + struct ibmveth_adapter *adapter,
> + int queue_index)
> +{
> + netdev_dbg(netdev, "recycling invalid buffer\n");
> + ibmveth_poll_bump_invalid(adapter, queue_index);
> + return ibmveth_poll_harvest_slot(adapter, queue_index, true);
> +}
[Severity: Medium]
Does the invalid-buffer path lose reset escalation for bad correlators?
In mainline, ibmveth_remove_buffer_from_pool() called
schedule_work(&adapter->work) for an out-of-range correlator (-EINVAL)
or a NULL skb (-EFAULT), for every harvest caller. The previous patch
("ibmveth: Add queue-aware RX buffer submit helper for MQ") removed
that call. This patch then removes schedule_work() from
ibmveth_rxq_get_buffer() too.
The only schedule_work() left in the poll slot handling is in
ibmveth_poll_skip_bad_correlator(). That is reached only from the
valid-slot path in ibmveth_poll_deliver_frame().
Take a slot with the VALID bit clear and a stale or out-of-range
correlator. It goes through:
ibmveth_poll()
ibmveth_poll_recycle_invalid()
ibmveth_poll_harvest_slot() /* -EINVAL/-EFAULT -> true */
The slot is counted, logged only with netdev_dbg(), skipped, and the
loop continues. No reset is scheduled.
The buffer PHYP returned is never given back to free_map, and
pool->available is never decremented for it. That capacity looks lost
until some unrelated reset.
skip_bad_correlator's own comment says a "residual stale slot after
resize" should be recovered by reset. Should recycle_invalid escalate
in the same way?
The commit message says "skip_bad_correlator owns reset escalation" but
doesn't mention that the invalid-slot path loses it. This is still the
case at the end of the series.
> +
> +static bool ibmveth_poll_skip_bad_correlator(struct net_device *netdev,
> + struct ibmveth_adapter *adapter,
> + int queue_index)
> +{
> + if (net_ratelimit())
> + netdev_err(netdev,
> + "bad correlator on queue %d, skipping slot\n",
> + queue_index);
> + /* Residual stale slot after resize: recover via reset rather
> + * than spinning forever. Always escalate; only the log is
> + * rate-limited.
> + */
> + schedule_work(&adapter->work);
[Severity: High]
This is a pre-existing issue and was not introduced by this patch.
At this commit, ibmveth_remove() cancels the work before the netdev is
unregistered:
ibmveth_remove() {
...
cancel_work_sync(&adapter->work);
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
kobject_put(&adapter->rx_buff_pool[0][i].kobj);
unregister_netdev(netdev);
free_netdev(netdev);
...
}
Can a NAPI poll queue adapter->work from here in that window? If so,
ibmveth_reset() would run after free_netdev().
Mainline poll could already call schedule_work() through
remove_buffer_from_pool() and get_buffer().
This ordering is fixed later in the series by "ibmveth: Enable
multi-queue RX receive path". There, ibmveth_remove() calls
unregister_netdev() before cancel_work_sync().
> + ibmveth_poll_bump_invalid(adapter, queue_index);
> + return ibmveth_poll_harvest_slot(adapter, queue_index, true);
> +}
[ ... ]
> + if (lrg_pkt) {
> + __be64 *rxmss = (__be64 *)(skb->data + 8);
> +
> + mss = (u16)be64_to_cpu(*rxmss);
> + }
[ ... ]
> + skb_put(skb, length);
> + skb->protocol = eth_type_trans(skb, netdev);
> +
> + if (skb->protocol == cpu_to_be16(ETH_P_IP))
> + iph_check = ((struct iphdr *)skb->data)->check;
[Severity: Low]
This isn't a bug, but two comments about PHYP behaviour were dropped
when this code moved into ibmveth_poll_deliver_frame():
/* if the large packet bit is set in the rx queue
* descriptor, the mss will be written by PHYP eight
* bytes from the start of the rx buffer, which is
* skb->data at this stage
*/
/* PHYP without PLSO support places a -1 in the ip
* checksum for large send frames.
*/
Without them, the skb->data + 8 read and the 0xffff test have no
explanation. Could the comments be kept?
> +
> + if ((length > netdev->mtu + ETH_HLEN) || lrg_pkt ||
> + iph_check == 0xffff) {
> + ibmveth_rx_mss_helper(skb, mss, lrg_pkt);
> + adapter->rx_large_packets++;
> + }
> +
> + if (csum_good) {
> + skb->ip_summed = CHECKSUM_UNNECESSARY;
> + ibmveth_rx_csum_helper(skb, adapter);
> + }
[Severity: High]
This isn't a bug introduced by this patch, but can a crafted IPv4 frame
from another partition crash the receiver here? When csum_good is set,
ibmveth_rx_csum_helper() does:
if (iph->check == 0xffff) {
iph->check = 0;
iph->check = ip_fast_csum((unsigned char *)iph,
iph->ihl);
}
iph->ihl is not validated first. On powerpc64, ip_fast_csum_nofold()
has this loop, where ihl is unsigned:
for (i = 0; i < ihl - 1; i++, ptr++)
s += *ptr;
With IHL=0 that is about 4 billion 32-bit reads (about 16 GiB) past the
skb, in softirq context.
The RX IBMVETH_RXQ_CSUM_GOOD bit (0x01000000) mirrors the sender's TX
IBMVETH_BUF_CSUM_GOOD bit (0x01000000). So a peer on the same virtual
LAN controls both csum_good and the IP header.
The new offset+length bound in ibmveth_poll_deliver_frame() doesn't
cover this. Should iph->ihl be validated before ip_fast_csum() is
called?
[ ... ]
> if (!napi_complete_done(napi, frames_processed))
> goto out;
>
> - /* We think we are done - reenable interrupts,
> - * then check once more to make sure we are done.
> + /*
> + * napi_disable() sets DISABLE then waits for this poll. Without a
> + * second stopping check here, enable_irq() can re-arm PHYP after
> + * resize already masked the queue; late IRQs then hit the handler
> + * after num_rx_queues was published lower (lab WARN at interrupt).
> */
[Severity: Low]
Is this comment accurate for this commit? It refers to a resize that
publishes num_rx_queues lower, and to a "lab WARN at interrupt".
At this commit, num_rx_queues is only assigned in probe. Also,
ibmveth_interrupt() explicitly avoids a WARN:
/*
* Quiet on out-of-range: teardown can leave a residual IRQ after the
* live count drops. Do not WARN-storm; return IRQ_NONE until free_irq.
*/
if (qindex < 0 || qindex >= adapter->num_rx_queues)
return IRQ_NONE;
Resize arrives later in the series. However, the WARN-at-interrupt text
and "napi_disable() sets DISABLE then waits for this poll" are still
wrong in the final tree (see the next comment).
> + if (ibmveth_poll_stopping(netdev, napi))
> + goto out;
> +
> rc = ibmveth_enable_irq(adapter, queue_index);
[Severity: High]
Does this check close the re-arm race the commit message describes? The
message says:
ibmveth_poll_stopping() ensures close/napi_disable
does not re-arm PHYP.
By this point, napi_complete_done() has already cleared
NAPI_STATE_SCHED. napi_disable_locked() only waits while SCHED or NPSVC
is set, and it clears NAPI_STATE_DISABLE again before returning.
So a napi_disable() that starts after napi_complete_done() does not
wait for this poll, and napi_disable_pending() is only true briefly.
Documentation/networking/napi.rst also says napi_disable() waits only
for ownership to be released, not for the poll method to exit.
pool_store and change_mtu call ibmveth_close() directly, so
netif_running() stays true on those paths. That allows:
CPU0 (ibmveth_poll)
napi_complete_done()
ibmveth_poll_stopping() /* false: DISABLE not set yet */
<delayed, e.g. vCPU preempted>
CPU1 (ibmveth_close() from pool_store or change_mtu)
ibmveth_cleanup_rx_interrupts()
ibmveth_disable_irq(); synchronize_irq();
napi_disable(); /* SCHED clear, returns at once */
ibmveth_disable_irq(); synchronize_irq();
free_irq();
synchronize_net();
CPU0
ibmveth_enable_irq() /* unmasks PHYP after the last mask */
The later reschedule attempt in the poll fails because the disabled
NAPI now owns SCHED, so nothing masks the queue again.
ibmveth_close() only calls synchronize_net() after
ibmveth_cleanup_rx_interrupts() and does not remask:
ibmveth_cleanup_rx_interrupts(adapter);
/* Wait for softirq/poll that already passed shutdown checks. */
synchronize_net();
dev_close() callers look safe, because __dev_close_many() clears
__LINK_STATE_START and synchronizes before ndo_stop.
With the later "ibmveth: Implement incremental MQ RX queue resize"
patch, the same window seems to let a retiring queue stay unmasked
until H_FREE_LOGICAL_LAN_QUEUE.
Would this need real synchronization, for example a remask after
synchronize_net(), rather than another lockless flag test?
[ ... ]
> @@ -2390,130 +2388,239 @@ static void ibmveth_rx_csum_helper(struct sk_buff *skb,
> }
> }
>
> +static void ibmveth_poll_bump_invalid(struct ibmveth_adapter *adapter,
> + int queue_index)
> +{
> + adapter->rx_invalid_buffer++;
> +}
> +
> +static bool ibmveth_poll_stopping(struct net_device *netdev,
> + struct napi_struct *napi)
> +{
> + return !netif_running(netdev) || napi_disable_pending(napi);
> +}
> +
> +static bool ibmveth_poll_harvest_slot(struct ibmveth_adapter *adapter,
> + int queue_index, bool reuse)
> +{
> + int rc = ibmveth_rxq_harvest_buffer(adapter, queue_index, reuse);
> +
> + return !rc || rc == -EINVAL || rc == -EFAULT;
> +}
> +
> +static bool ibmveth_poll_recycle_invalid(struct net_device *netdev,
> + struct ibmveth_adapter *adapter,
> + int queue_index)
> +{
> + netdev_dbg(netdev, "recycling invalid buffer\n");
> + ibmveth_poll_bump_invalid(adapter, queue_index);
> + return ibmveth_poll_harvest_slot(adapter, queue_index, true);
> +}
> +
> +static bool ibmveth_poll_skip_bad_correlator(struct net_device *netdev,
> + struct ibmveth_adapter *adapter,
> + int queue_index)
> +{
> + if (net_ratelimit())
> + netdev_err(netdev,
> + "bad correlator on queue %d, skipping slot\n",
> + queue_index);
> + /* Residual stale slot after resize: recover via reset rather
> + * than spinning forever. Always escalate; only the log is
> + * rate-limited.
> + */
> + schedule_work(&adapter->work);
> + ibmveth_poll_bump_invalid(adapter, queue_index);
> + return ibmveth_poll_harvest_slot(adapter, queue_index, true);
> +}
> +
> +static bool ibmveth_poll_drop_oversize(struct net_device *netdev,
> + struct ibmveth_adapter *adapter,
> + int queue_index, unsigned int off,
> + unsigned int len, unsigned int room)
> +{
> + if (net_ratelimit())
> + netdev_err(netdev,
> + "RX frame %u+%u exceeds buffer %u on queue %d, dropping\n",
> + off, len, room, queue_index);
> + ibmveth_poll_bump_invalid(adapter, queue_index);
> + return ibmveth_poll_harvest_slot(adapter, queue_index, true);
> +}
> +
> +/**
> + * ibmveth_poll_deliver_frame - Build SKB from one valid RX slot and GRO it
> + * @napi: NAPI context for this RX queue
> + * @adapter: ibmveth adapter
> + * @netdev: net_device for @adapter
> + * @queue_index: RX queue index
> + *
> + * Return: 1 frame delivered, 0 if the slot was skipped cleanly, -1 on error.
> + */
> +static int ibmveth_poll_deliver_frame(struct napi_struct *napi,
> + struct ibmveth_adapter *adapter,
> + struct net_device *netdev,
> + int queue_index)
> +{
> + struct sk_buff *skb, *new_skb;
> + struct ibmveth_rx_q *rxq = &adapter->rx_queue[queue_index];
> + unsigned int room, off, len, pool;
> + int length, offset, csum_good, lrg_pkt;
> + __sum16 iph_check = 0;
> + u16 mss = 0;
> + int rc;
> +
> + length = ibmveth_rxq_frame_length(adapter, queue_index);
> + offset = ibmveth_rxq_frame_offset(adapter, queue_index);
> + csum_good = ibmveth_rxq_csum_good(adapter, queue_index);
> + lrg_pkt = ibmveth_rxq_large_packet(adapter, queue_index);
> +
> + skb = ibmveth_rxq_get_buffer(adapter, queue_index);
> + if (unlikely(!skb)) {
> + if (!ibmveth_poll_skip_bad_correlator(netdev, adapter,
> + queue_index))
> + return -1;
> + return 0;
> + }
> +
> + pool = rxq->queue_addr[rxq->index].correlator >> 32;
> + room = min_t(unsigned int, skb_tailroom(skb),
> + adapter->rx_buff_pool[queue_index][pool].buff_size);
> + off = offset;
> + len = length;
> + if (unlikely(off >= room || len > room - off)) {
> + if (!ibmveth_poll_drop_oversize(netdev, adapter, queue_index,
> + off, len, room))
> + return -1;
> + return 0;
> + }
> +
> + if (lrg_pkt) {
> + __be64 *rxmss = (__be64 *)(skb->data + 8);
> +
> + mss = (u16)be64_to_cpu(*rxmss);
> + }
> +
> + new_skb = NULL;
> + if (length < rx_copybreak)
> + new_skb = netdev_alloc_skb(netdev, length);
> +
> + if (new_skb) {
> + skb_copy_to_linear_data(new_skb, skb->data + offset, length);
> + if (rx_flush)
> + ibmveth_flush_buffer(skb->data, length + offset);
> + rc = ibmveth_rxq_harvest_buffer(adapter, queue_index, true);
> + if (unlikely(rc)) {
> + kfree_skb(new_skb);
> + return -1;
> + }
> + skb = new_skb;
> + } else {
> + rc = ibmveth_rxq_harvest_buffer(adapter, queue_index, false);
> + if (unlikely(rc))
> + return -1;
> + skb_reserve(skb, offset);
> + }
> +
> + skb_put(skb, length);
> + skb->protocol = eth_type_trans(skb, netdev);
> +
> + if (skb->protocol == cpu_to_be16(ETH_P_IP))
> + iph_check = ((struct iphdr *)skb->data)->check;
> +
> + if ((length > netdev->mtu + ETH_HLEN) || lrg_pkt ||
> + iph_check == 0xffff) {
> + ibmveth_rx_mss_helper(skb, mss, lrg_pkt);
> + adapter->rx_large_packets++;
> + }
> +
> + if (csum_good) {
> + skb->ip_summed = CHECKSUM_UNNECESSARY;
> + ibmveth_rx_csum_helper(skb, adapter);
> + }
> +
> + napi_gro_receive(napi, skb);
> +
> + netdev->stats.rx_packets++;
> + netdev->stats.rx_bytes += length;
> +
> + return 1;
> +}
> +
> static int ibmveth_poll(struct napi_struct *napi, int budget)
> {
> struct net_device *netdev = napi->dev;
> struct ibmveth_adapter *adapter = netdev_priv(netdev);
> int frames_processed = 0;
> int queue_index, rc;
> - u16 mss = 0;
>
> queue_index = napi - adapter->napi;
>
> + if (WARN_ON(queue_index < 0 ||
> + queue_index >= adapter->num_rx_queues)) {
> + if (budget)
> + napi_complete_done(napi, 0);
> + return 0;
> + }
> +
> + if (ibmveth_poll_stopping(netdev, napi)) {
> + if (budget)
> + napi_complete_done(napi, 0);
> + return 0;
> + }
> +
> restart_poll:
> while (frames_processed < budget) {
> + if (ibmveth_poll_stopping(netdev, napi))
> + break;
> +
> if (!ibmveth_rxq_pending_buffer(adapter, queue_index))
> break;
>
> smp_rmb();
> if (!ibmveth_rxq_buffer_valid(adapter, queue_index)) {
> wmb(); /* suggested by larson1 */
> - adapter->rx_invalid_buffer++;
> - netdev_dbg(netdev, "recycling invalid buffer\n");
> - rc = ibmveth_rxq_harvest_buffer(adapter,
> - queue_index, true);
> - if (unlikely(rc))
> + if (!ibmveth_poll_recycle_invalid(netdev, adapter,
> + queue_index))
> break;
> } else {
> - struct sk_buff *skb, *new_skb;
> - int length = ibmveth_rxq_frame_length(adapter,
> - queue_index);
> - int offset = ibmveth_rxq_frame_offset(adapter,
> - queue_index);
> - int csum_good = ibmveth_rxq_csum_good(adapter,
> - queue_index);
> - int lrg_pkt = ibmveth_rxq_large_packet(adapter,
> - queue_index);
> - __sum16 iph_check = 0;
> -
> - skb = ibmveth_rxq_get_buffer(adapter, queue_index);
> - if (unlikely(!skb)) {
> - struct ibmveth_rx_q *rxq =
> - &adapter->rx_queue[queue_index];
> -
> - ibmveth_rxq_advance(rxq);
> + rc = ibmveth_poll_deliver_frame(napi, adapter, netdev,
> + queue_index);
> + if (rc < 0)
> break;
> - }
> -
> - /* if the large packet bit is set in the rx queue
> - * descriptor, the mss will be written by PHYP eight
> - * bytes from the start of the rx buffer, which is
> - * skb->data at this stage
> - */
> - if (lrg_pkt) {
> - __be64 *rxmss = (__be64 *)(skb->data + 8);
> -
> - mss = (u16)be64_to_cpu(*rxmss);
> - }
> -
> - new_skb = NULL;
> - if (length < rx_copybreak)
> - new_skb = netdev_alloc_skb(netdev, length);
> -
> - if (new_skb) {
> - skb_copy_to_linear_data(new_skb,
> - skb->data + offset,
> - length);
> - if (rx_flush)
> - ibmveth_flush_buffer(skb->data,
> - length + offset);
> - rc = ibmveth_rxq_harvest_buffer(adapter,
> - queue_index,
> - true);
> - if (unlikely(rc))
> - break;
> - skb = new_skb;
> - } else {
> - rc = ibmveth_rxq_harvest_buffer(adapter,
> - queue_index,
> - false);
> - if (unlikely(rc))
> - break;
> - skb_reserve(skb, offset);
> - }
> -
> - skb_put(skb, length);
> - skb->protocol = eth_type_trans(skb, netdev);
> -
> - /* PHYP without PLSO support places a -1 in the ip
> - * checksum for large send frames.
> - */
> - if (skb->protocol == cpu_to_be16(ETH_P_IP)) {
> - struct iphdr *iph = (struct iphdr *)skb->data;
> -
> - iph_check = iph->check;
> - }
> -
> - if ((length > netdev->mtu + ETH_HLEN) ||
> - lrg_pkt || iph_check == 0xffff) {
> - ibmveth_rx_mss_helper(skb, mss, lrg_pkt);
> - adapter->rx_large_packets++;
> - }
> -
> - if (csum_good) {
> - skb->ip_summed = CHECKSUM_UNNECESSARY;
> - ibmveth_rx_csum_helper(skb, adapter);
> - }
> -
> - napi_gro_receive(napi, skb); /* send it up */
> -
> - netdev->stats.rx_packets++;
> - netdev->stats.rx_bytes += length;
> - frames_processed++;
> + if (rc > 0)
> + frames_processed++;
> }
> }
>
> ibmveth_replenish_task(adapter, queue_index);
>
> + if (ibmveth_poll_stopping(netdev, napi)) {
> + /* budget 0 is netpoll, which must not complete NAPI.
> + * Otherwise returning budget after completing would ask
> + * NAPI to reschedule, so cap the return at budget - 1.
> + */
> + if (budget) {
> + napi_complete_done(napi, frames_processed);
> + return min(frames_processed, budget - 1);
> + }
> + return 0;
> + }
> +
> if (frames_processed == budget)
> goto out;
>
> if (!napi_complete_done(napi, frames_processed))
> goto out;
>
> - /* We think we are done - reenable interrupts,
> - * then check once more to make sure we are done.
> + /*
> + * napi_disable() sets DISABLE then waits for this poll. Without a
> + * second stopping check here, enable_irq() can re-arm PHYP after
> + * resize already masked the queue; late IRQs then hit the handler
> + * after num_rx_queues was published lower (lab WARN at interrupt).
> */
> + if (ibmveth_poll_stopping(netdev, napi))
> + goto out;
> +
> rc = ibmveth_enable_irq(adapter, queue_index);
> if (rc) {
> netdev_err(netdev,
[Severity: High]
Can the pool index computed here go out of bounds?
ibmveth_rxq_get_buffer() loads the correlator from the RX ring once,
checks it with ibmveth_rxq_correlator_valid() and returns the skb.
This hunk then recomputes pool from a fresh load of
rxq->queue_addr[rxq->index].correlator. That field lives in DMA
memory that PHYP writes. The new value is never checked against
IBMVETH_NUM_BUFF_POOLS before it indexes
adapter->rx_buff_pool[queue_index][pool].
pool is an unsigned int taken from the top 32 bits. If the ring entry
no longer matches what get_buffer() validated, the buff_size read can
land far past the end of struct ibmveth_adapter and fault in NAPI
softirq context. None of these loads use READ_ONCE(), so the compiler
is also free to refetch.
ibmveth_rxq_harvest_buffer() then reads the correlator a third time
and hands it to ibmveth_remove_buffer_from_pool(). That read predates
this patch. If it disagrees with get_buffer(), the driver removes a
different pool entry from the one whose skb it gives to
napi_gro_receive(). With reuse false, the delivered skb would stay in
pool->skbuff[] and stay DMA-mapped. A later recycle or
ibmveth_free_buffer_pool() could then touch an skb the stack already
owns.
PHYP is not expected to rewrite a slot it has handed back. But the
commit message presents the new offset+length bound as hardening
against PHYP-supplied values, and this adds an unchecked index right
next to it.
Would it be better to read the correlator once per slot and pass the
validated value (or pool and index) through get_buffer, the room
calculation and the harvest, instead of going back to the ring each
time?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 10/15] ibmveth: Enable multi-queue RX receive path
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (8 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 09/15] ibmveth: Harden RX poll path with helpers Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 11/15] ibmveth: Add per-queue RX and TX statistics collection Mingming Cao
` (5 subsequent siblings)
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
Turn on the live MQ RX path when firmware sets
IBMVETH_ILLAN_RX_MULTI_QUEUE_SUPPORT in H_ILLAN_ATTRIBUTES. Probe sets
multi_queue and num_rx_queues to min(num_online_cpus(),
IBMVETH_DEFAULT_QUEUES), matching the existing TX default (cap 8).
IBMVETH_MAX_RX_QUEUES goes from 1 to 16, which is what actually gives
the adapter more than one RX ring: every rx_queue[], queue_handle[]
and queue_irq[] array sized by it grows to sixteen entries.
multi_queue also becomes a bool, now that it is a flag and not a count.
MQ firmware also raises the replenish batch from eight buffers per
hypercall to twelve, the argument-list capacity of
H_ADD_LOGICAL_LAN_BUFFERS_QUEUE, by setting rx_buffers_per_hcall to
IBMVETH_MAX_RX_PER_HCALL instead of IBMVETH_MAX_RX_REGULAR.
Patch 14 wires live ethtool -L rx (set_channels reads rx_count and
calls resize_rx_channels). Patch 15 completes the down-path
publish/rollback and caps max_rx at the live RX count once mq_fallback
latches.
This commit does not implement set_channels / rx_count. Without the bit,
behaviour stays single-queue.
Wire subordinate queues through H_REG_LOGICAL_LAN_QUEUE (with
irq_create_mapping for subordinate virqs), request_irq/napi_enable,
and PHYP enable_irq when multi_queue && num_rx_queues > 1. Queue 0
continues to use netdev->irq and is never disposed with the
subordinates.
Restate open() here (supersedes the register-helpers patch kick):
... register_rx_queues / set_real_num_rx_queues ...
replenish_task() for every live RX queue (before IRQ setup)
setup_rx_interrupts() (MQ also unmasks PHYP)
restart_rx_queue() for every live queue (schedule NAPI, or
enable_irq if prep fails)
... alloc_tx_resources / tx_start / opened ...
That path is the same for SQ and MQ; it replaces the prior
setup-then-schedule_rx_queue(0) kick. Close shape is unchanged.
H_FUNCTION means firmware withdrew MQ. Both the subordinate-register
and the buffer-add path latch mq_fallback, so the next open comes up
single-queue (apply_mq_fallback at open entry); they differ in what
happens to the open in progress. Register fails it outright rather
than dropping to single-queue silently mid-open, and buffer-add
schedules a reset. Reset does not retry if that close/open fails.
Both setup failure paths dispose subordinate virq mappings, so a
request_irq failure after successful registration cannot leak Linux
mappings.
On probe failure after pool kobjects were created, put them before
free_netdev(). The leak is pre-existing and unrelated to multi-queue;
the probe_cleanup helper lands in patch 11.
get_desired_dma() sizes RX from queue-0 pool metadata across
num_rx_queues. CMO (Power9 and earlier) and MQ firmware (Power11+)
do not coexist, so probe does not refresh CMO desired and a CMO
partition keeps its single-queue default.
This commit adds schedule_work() producers on buffer-add H_FUNCTION,
so the remove-path unregister / cancel_work_sync reorder and reset
reg_state gate land here to prevent a queued reset from racing
device teardown. RX counters are knowingly left racy for patch 11:
adapter->rx_no_buffer is assigned rather than summed from one queue's
buffer-list page, so it reports whichever queue replenished last and
can go backwards, while rx_packets, rx_bytes, rx_invalid_buffer and
rx_large_packets are plain read-modify-writes now reached from several
NAPI instances at once, so they can lose counts; patch 11 moves both
to per-queue storage summed on read.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- register_single_rx_queue logs lpar_rc with %ld
- setup_rx_interrupts kdoc: open posts first in both modes
- wait for pool kobject release before free_netdev
- remove-path unregister/cancel reorder and reset
reg_state gate land here alongside the reset producer
- drop the probe CMO desired refresh: CMO (Power9 and
earlier) and MQ (Power11+) do not coexist; say so in
the get_desired_dma() sizing comment
- commit message: drop the stale net sentence
- noted: restart enable_irq on prep fail stays;
counters stay in P11; poll_controller
!opened stays in P15
- noted: no CMO+MQ partition (CMO is P9, MQ is P11+);
default-8 is the TX match, not a CMO tradeoff;
get_desired_dma TX term is CMO desired-hint only;
real TX LTBs already scale; those two CMO leftovers
drop from the v7 cover
Changes in v6:
- ibmveth_get_num_rx_queues / ibmveth_publish_num_rx_queues are
static, not static inline
- wrap the rx_buffers_per_hcall assignment (81 cols)
- smp_store_release / smp_load_acquire for num_rx_queues, not
smp_wmb() + WRITE_ONCE()
- skb_record_rx_queue() before GRO
- drop the IBMVETH_MAX_RX_QUEUE alias (duplicated MAX_RX_PER_HCALL)
- drop WARN_ON on restart enable_irq; the helper already logs.
Enabling on prep failure is deliberate
- drop unused kick_rx_queue_if_pending() (WERROR); patch 14 uses
restart_rx_queue() at both sites
- register_rx_queues() kdoc: this function latches mq_fallback on
-EOPNOTSUPP
- noted: pool kobj vs DEBUG_KOBJECT_RELEASE stays cover leftovers
Changes in v5:
- RTNL/serialized readers use get_num_rx_queues() once the helper exists
(open/close/register/apply_mq_fallback/alloc/cleanup/IRQ setup)
- Interrupt: quiet IRQ_NONE on out-of-range qindex vs published live
count (scale-down residual IRQ; no WARN storm)
- Series renumber: mailed v4 09/14 MQ enable -> tip P10 (P09 peel; 14->15)
- Open: replenish all queues before setup; restart_rx_queue after setup
for SQ and MQ - replaces mailed MQ-only replenish + kick_if_pending vs
SQ schedule_rx_queue(0)
- update_rx_no_buffer(queue_index): NULL-safe, no all-queue walk under one
lock (fixes the resize/race class). Still stores adapter->rx_no_buffer
for now; per-queue slot + monotonic sum lands with qstats next
- H_FUNCTION: subordinate register fails open + stash mq_fallback;
buffer-add schedules reset + mq_fallback so next open drops to SQ
- Introduce get_num_rx_queues / publish_num_rx_queues (READ/WRITE_ONCE)
at first use; lockless IRQ/poll/replenish/interrupt bounds use the getter
- Introduce kick_rx_queue_if_pending() for pending-after-unmask; open uses
restart_rx_queue, resize calls the helper later
- resume() schedules every live RX queue (not only queue 0)
- Call out enable_irq-fail synchronize_irq and unused mac register-arg
cleanup (mailed enable-path review; code already earlier in tip)
Changes in v4:
- Fold subordinate register helpers and their review fixes into the MQ
enablement patch that first uses them.
- Prefer request_irq -> napi_enable -> PHYP enable on MQ open.
- Preserve open unwind so set_real_num_rx / IRQ failures free LAN
before buffer pools.
- MQ open replenishes every queue before setup_rx_interrupts() unmasks
PHYP (drop avoidance during open; PHYP only interrupts after a
successful enqueue). SQ keeps classic setup-then-schedule kick.
- Dispose subordinate virq mappings on setup_rx_interrupts()
request_irq failure (err_free_irqs), matching err_disable_napi.
- Open unwind: setup/cleanup own subordinate dispose; skip duplicate
dispose on those paths.
- H_FUNCTION on subordinate register is a hard open failure (no blind
retry / no fake single-queue fallback).
- Note: hot-path netdev->stats accounting moves to the next patch (qstats).
- Put already-created pool kobjects on probe kobject_init_and_add /
set_real_num_tx_queues / register_netdev failure (bisect-safe).
drivers/net/ethernet/ibm/ibmveth.c | 549 ++++++++++++++++++++++++-----
drivers/net/ethernet/ibm/ibmveth.h | 8 +-
2 files changed, 473 insertions(+), 84 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 7b6c0283e5c3..3f31793645a3 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -156,6 +156,28 @@ static int ibmveth_rxq_csum_good(struct ibmveth_adapter *adapter,
return ibmveth_rxq_flags(adapter, queue_index) & IBMVETH_RXQ_CSUM_GOOD;
}
+/* Lockless IRQ/poll readers vs resize publishers. */
+static unsigned int
+ibmveth_get_num_rx_queues(const struct ibmveth_adapter *adapter)
+{
+ /*
+ * Pairs with the release in ibmveth_publish_num_rx_queues(): a reader
+ * that sees the new count also sees the per-queue state behind it.
+ */
+ return smp_load_acquire(&adapter->num_rx_queues);
+}
+
+static void
+ibmveth_publish_num_rx_queues(struct ibmveth_adapter *adapter,
+ unsigned int num)
+{
+ /*
+ * Pairs with the acquire in ibmveth_get_num_rx_queues(): per-queue
+ * state must be visible to a reader before it observes the new count.
+ */
+ smp_store_release(&adapter->num_rx_queues, num);
+}
+
static unsigned int ibmveth_real_max_tx_queues(void)
{
unsigned int n_cpu = num_online_cpus();
@@ -233,7 +255,7 @@ ibmveth_alloc_rx_queues(struct ibmveth_adapter *adapter, int rxq_entries)
struct net_device *netdev = adapter->netdev;
int i;
- for (i = 0; i < adapter->num_rx_queues; i++) {
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
adapter->buffer_list_addr[i] =
(void *)get_zeroed_page(GFP_KERNEL);
if (!adapter->buffer_list_addr[i]) {
@@ -284,7 +306,7 @@ ibmveth_alloc_rx_queues(struct ibmveth_adapter *adapter, int rxq_entries)
}
netdev_dbg(netdev, "allocated %u RX queue(s) with %d entries each\n",
- adapter->num_rx_queues, rxq_entries);
+ ibmveth_get_num_rx_queues(adapter), rxq_entries);
return 0;
@@ -328,9 +350,9 @@ ibmveth_cleanup_rx_resources(struct ibmveth_adapter *adapter)
int i;
netdev_dbg(adapter->netdev, "cleaning up %u RX queue(s)\n",
- adapter->num_rx_queues);
+ ibmveth_get_num_rx_queues(adapter));
- for (i = 0; i < adapter->num_rx_queues; i++) {
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
if (adapter->buffer_list_addr[i]) {
dma_unmap_single(dev, adapter->buffer_list_dma[i],
4096, DMA_BIDIRECTIONAL);
@@ -487,7 +509,7 @@ ibmveth_dispose_subordinate_irq_mappings(struct ibmveth_adapter *adapter)
{
int i;
- for (i = 1; i < adapter->num_rx_queues; i++)
+ for (i = 1; i < ibmveth_get_num_rx_queues(adapter); i++)
ibmveth_dispose_subordinate_irq_mapping(adapter, i);
}
@@ -497,12 +519,11 @@ ibmveth_dispose_subordinate_irq_mappings(struct ibmveth_adapter *adapter)
*
* Registers interrupt handlers for all RX queues, enables NAPI, then
* enables hypervisor interrupt delivery for multi-queue mode after
- * every queue has a Linux handler installed. For multi-queue open the
- * caller should replenish RX buffers before this helper so traffic
- * during open is not dropped (PHYP only interrupts after a successful
- * enqueue, which needs buffers). Single-queue open leaves PHYP masked
- * here and kicks NAPI afterward (classic path: first poll posts then
- * enables).
+ * every queue has a Linux handler installed. The caller replenishes
+ * every live queue before this helper so traffic during open is not
+ * dropped (PHYP only interrupts after a successful enqueue, which
+ * needs buffers). Single-queue open still leaves PHYP masked here;
+ * restart_rx_queue then kicks NAPI.
*
* Return: 0 on success, negative error code on failure
*/
@@ -510,7 +531,7 @@ static int
ibmveth_setup_rx_interrupts(struct ibmveth_adapter *adapter)
{
struct net_device *netdev = adapter->netdev;
- int i, rc, num = adapter->num_rx_queues;
+ int i, rc, num = ibmveth_get_num_rx_queues(adapter);
for (i = 0; i < num; i++) {
if (!adapter->queue_irq[i]) {
@@ -600,24 +621,24 @@ ibmveth_cleanup_rx_interrupts(struct ibmveth_adapter *adapter)
if (!adapter->rx_irq_setup)
return;
- for (i = 0; i < adapter->num_rx_queues; i++) {
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
if (!adapter->queue_irq[i])
continue;
ibmveth_disable_irq(adapter, i);
synchronize_irq(adapter->queue_irq[i]);
}
- for (i = 0; i < adapter->num_rx_queues; i++)
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
napi_disable(&adapter->napi[i]);
- for (i = 0; i < adapter->num_rx_queues; i++) {
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
if (!adapter->queue_irq[i])
continue;
ibmveth_disable_irq(adapter, i);
synchronize_irq(adapter->queue_irq[i]);
}
- for (i = 0; i < adapter->num_rx_queues; i++) {
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
if (adapter->queue_irq[i])
free_irq(adapter->queue_irq[i], &adapter->napi[i]);
}
@@ -647,7 +668,7 @@ static bool ibmveth_schedule_rx_queue(struct ibmveth_adapter *adapter,
{
struct napi_struct *napi = &adapter->napi[qindex];
- if (WARN_ON(qindex < 0 || qindex >= adapter->num_rx_queues))
+ if (WARN_ON(qindex < 0 || qindex >= ibmveth_get_num_rx_queues(adapter)))
return false;
/*
@@ -993,15 +1014,21 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
* because there was not a buffer in the buffer list capable of holding
* the frame.
*/
-static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter)
+static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter,
+ int queue_index)
{
__be64 *p;
+ u64 drops;
- if (!adapter->buffer_list_addr[0])
+ if (queue_index < 0 ||
+ queue_index >= ibmveth_get_num_rx_queues(adapter) ||
+ !adapter->buffer_list_addr[queue_index])
return;
- p = adapter->buffer_list_addr[0] + 4096 - 8;
- adapter->rx_no_buffer = be64_to_cpup(p);
+ p = adapter->buffer_list_addr[queue_index] + 4096 - 8;
+ drops = be64_to_cpup(p);
+
+ adapter->rx_no_buffer = drops;
}
/* replenish routine */
@@ -1016,10 +1043,10 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
int batch_fallback = 0;
int hcall_fail = 0;
- if (queue_index >= adapter->num_rx_queues) {
+ if (queue_index >= ibmveth_get_num_rx_queues(adapter)) {
netdev_dbg(adapter->netdev,
"Skipping replenish for freed queue %d (num_queues=%u)\n",
- queue_index, adapter->num_rx_queues);
+ queue_index, ibmveth_get_num_rx_queues(adapter));
return;
}
@@ -1053,7 +1080,7 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
}
out_unlock:
- ibmveth_update_rx_no_buffer(adapter);
+ ibmveth_update_rx_no_buffer(adapter, queue_index);
spin_unlock_irqrestore(&rxq->replenish_lock, flags);
@@ -1067,6 +1094,7 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
dev_err_ratelimited(&adapter->netdev->dev,
"MQ buffer add H_FUNCTION (q=%d, batch=%u), reset\n",
queue_index, fail.batch);
+ adapter->mq_fallback = true;
schedule_work(&adapter->work);
}
@@ -1086,6 +1114,27 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
fail.filled, fail.lpar_rc, fail.batch);
}
+/**
+ * ibmveth_restart_rx_queue - Post buffers and ensure Q can take RX
+ * @adapter: ibmveth adapter
+ * @qindex: RX queue index
+ *
+ * SQ open leaves PHYP masked until the first poll. If schedule_prep fails,
+ * NAPI never runs and the queue stays masked (TX OK, RX/ARP dead) until
+ * reload. Replenish first so an enable_irq fallback can actually deliver.
+ * Also used after every open (SQ and MQ) and after scale-down so a
+ * queue is not left idle+masked.
+ */
+static void ibmveth_restart_rx_queue(struct ibmveth_adapter *adapter,
+ int qindex)
+{
+ ibmveth_replenish_task(adapter, qindex);
+ if (ibmveth_schedule_rx_queue(adapter, qindex))
+ return;
+
+ ibmveth_enable_irq(adapter, qindex);
+}
+
/* empty and free ana buffer pool - also used to do cleanup in error paths */
static void ibmveth_free_buffer_pool(struct ibmveth_adapter *adapter,
struct ibmveth_buff_pool *pool)
@@ -1214,7 +1263,7 @@ ibmveth_alloc_buffer_pools(struct ibmveth_adapter *adapter)
int i, q, rc;
/* Initialize pool metadata for queues 1..N from queue 0 settings */
- for (q = 1; q < adapter->num_rx_queues; q++) {
+ for (q = 1; q < ibmveth_get_num_rx_queues(adapter); q++) {
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
struct ibmveth_buff_pool *src =
&adapter->rx_buff_pool[0][i];
@@ -1230,7 +1279,7 @@ ibmveth_alloc_buffer_pools(struct ibmveth_adapter *adapter)
}
/* Allocate actual buffers for all queues */
- for (q = 0; q < adapter->num_rx_queues; q++) {
+ for (q = 0; q < ibmveth_get_num_rx_queues(adapter); q++) {
rc = ibmveth_alloc_queue_buffer_pools(adapter, q);
if (rc) {
/* Free pools for all previous queues */
@@ -1241,7 +1290,7 @@ ibmveth_alloc_buffer_pools(struct ibmveth_adapter *adapter)
}
netdev_dbg(netdev, "allocated buffer pools for %u queue(s)\n",
- adapter->num_rx_queues);
+ ibmveth_get_num_rx_queues(adapter));
return 0;
}
@@ -1257,11 +1306,11 @@ ibmveth_free_buffer_pools(struct ibmveth_adapter *adapter)
int q;
/* Free buffer pools for all queues */
- for (q = 0; q < adapter->num_rx_queues; q++)
+ for (q = 0; q < ibmveth_get_num_rx_queues(adapter); q++)
ibmveth_free_queue_buffer_pools(adapter, q);
netdev_dbg(adapter->netdev, "freed buffer pools for %u queue(s)\n",
- adapter->num_rx_queues);
+ ibmveth_get_num_rx_queues(adapter));
}
static bool ibmveth_rxq_correlator_valid(struct ibmveth_adapter *adapter,
@@ -1562,6 +1611,138 @@ static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
return rc;
}
+/**
+ * ibmveth_register_logical_lan_queue - Register subordinate queue with
+ * hypervisor
+ * @adapter: ibmveth adapter structure
+ * @rxq_desc: Receive queue descriptor
+ * @queue_index: RX queue index (1..N for subordinate queues)
+ *
+ * Registers a subordinate receive queue using H_REG_LOGICAL_LAN_QUEUE.
+ * On success, stores the queue handle and virtual IRQ in the adapter.
+ * If IRQ mapping fails after a successful hypervisor registration, the
+ * queue is freed before returning.
+ *
+ * Return: H_SUCCESS on success, negative errno on IRQ mapping failure,
+ * hypervisor error code otherwise
+ */
+static int
+ibmveth_register_logical_lan_queue(struct ibmveth_adapter *adapter,
+ union ibmveth_buf_desc rxq_desc,
+ int queue_index)
+{
+ unsigned long handle, hwirq;
+ unsigned int virq;
+ long lpar_rc;
+ unsigned long ua = adapter->vdev->unit_address;
+ unsigned long bl = adapter->buffer_list_dma[queue_index];
+
+ netdev_dbg(adapter->netdev,
+ "register queue %d: ua=0x%lx bl=0x%lx rxq=0x%llx\n",
+ queue_index, ua, bl, rxq_desc.desc);
+ do {
+ lpar_rc = h_register_logical_lan_queue(ua, bl,
+ rxq_desc.desc, &handle,
+ &hwirq);
+ } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
+ netdev_dbg(adapter->netdev,
+ "h_register_logical_lan_queue queue %d rc=%ld\n",
+ queue_index, lpar_rc);
+
+ if (lpar_rc == H_SUCCESS) {
+ virq = irq_create_mapping(NULL, hwirq);
+ if (!virq) {
+ unsigned long free_rc;
+
+ netdev_err(adapter->netdev,
+ "Failed to map IRQ for queue %d (hwirq=%lu)\n",
+ queue_index, hwirq);
+ do {
+ free_rc = h_free_logical_lan_queue(ua, handle);
+ } while (H_IS_LONG_BUSY(free_rc) ||
+ (free_rc == H_BUSY));
+ if (free_rc != H_SUCCESS)
+ netdev_err(adapter->netdev,
+ "h_free_logical_lan_queue failed for queue %d after IRQ map failure: rc=0x%lx\n",
+ queue_index, free_rc);
+ return -EINVAL;
+ }
+
+ adapter->queue_handle[queue_index] = handle;
+ adapter->queue_irq[queue_index] = virq;
+
+ netdev_dbg(adapter->netdev,
+ "queue %d registered: handle=0x%llx irq=%u\n",
+ queue_index, adapter->queue_handle[queue_index],
+ adapter->queue_irq[queue_index]);
+ return H_SUCCESS;
+ }
+
+ /*
+ * H_FUNCTION means firmware rejected this subordinate register
+ * (MQ unsupported / dropped after LPM). Caller fails this open and
+ * latches mq_fallback so the next open applies SQ; keep a specific
+ * log then the generic failure lines below.
+ */
+ if (lpar_rc == H_FUNCTION)
+ netdev_err(adapter->netdev,
+ "h_register_logical_lan_queue H_FUNCTION for queue %d (firmware MQ unsupported)\n",
+ queue_index);
+
+ netdev_err(adapter->netdev,
+ "h_register_logical_lan_queue failed for queue %d with %ld\n",
+ queue_index, lpar_rc);
+ netdev_err(adapter->netdev,
+ "queue %d params: unit_addr=0x%x buffer_list_dma=0x%llx rxq_desc=0x%llx\n",
+ queue_index, adapter->vdev->unit_address,
+ adapter->buffer_list_dma[queue_index],
+ rxq_desc.desc);
+
+ return lpar_rc;
+}
+
+/**
+ * ibmveth_register_single_rx_queue - Register one subordinate RX queue
+ * @adapter: ibmveth adapter structure
+ * @queue_idx: Queue index to register (1..N)
+ *
+ * Builds the queue descriptor and registers with the hypervisor via
+ * ibmveth_register_logical_lan_queue().
+ *
+ * Return: 0 on success, -EINVAL if @queue_idx is invalid, -EOPNOTSUPP if
+ * firmware rejects MQ (H_FUNCTION), -EIO on other failures
+ */
+static int
+ibmveth_register_single_rx_queue(struct ibmveth_adapter *adapter,
+ int queue_idx)
+{
+ struct net_device *netdev = adapter->netdev;
+ union ibmveth_buf_desc rxq_desc;
+ long lpar_rc;
+
+ if (WARN_ON(queue_idx < 1 || queue_idx >= IBMVETH_MAX_RX_QUEUES))
+ return -EINVAL;
+
+ rxq_desc.fields.flags_len = IBMVETH_BUF_VALID |
+ adapter->rx_queue[queue_idx].queue_len;
+ rxq_desc.fields.address = adapter->rx_queue[queue_idx].queue_dma;
+
+ lpar_rc = ibmveth_register_logical_lan_queue(adapter, rxq_desc,
+ queue_idx);
+ if (lpar_rc != H_SUCCESS) {
+ netdev_err(netdev, "Failed to register queue %d: rc=%ld\n",
+ queue_idx, lpar_rc);
+ if (lpar_rc == H_FUNCTION)
+ return -EOPNOTSUPP;
+ return -EIO;
+ }
+
+ netdev_dbg(netdev, "Registered queue %d with handle 0x%llx\n",
+ queue_idx, adapter->queue_handle[queue_idx]);
+
+ return 0;
+}
+
/**
* ibmveth_free_all_queues - Free all RX queues at once
* @adapter: ibmveth adapter structure
@@ -1579,7 +1760,8 @@ static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
* afterward (same as pre-helper close()).
*
* Clears queue handles only; queue_irq[] is released by
- * ibmveth_cleanup_rx_interrupts().
+ * ibmveth_cleanup_rx_interrupts() on close, or by
+ * ibmveth_dispose_subordinate_irq_mappings() on partial register failure.
*/
static void ibmveth_free_all_queues(struct ibmveth_adapter *adapter)
{
@@ -1597,7 +1779,7 @@ static void ibmveth_free_all_queues(struct ibmveth_adapter *adapter)
"h_free_logical_lan failed: %ld\n", lpar_rc);
}
- for (i = 0; i < adapter->num_rx_queues; i++)
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
adapter->queue_handle[i] = 0;
}
@@ -1606,10 +1788,13 @@ static void ibmveth_free_all_queues(struct ibmveth_adapter *adapter)
* @adapter: ibmveth adapter structure
* @mac_address: MAC address for device registration
*
- * Registers queue 0 via ibmveth_register_logical_lan(). Subordinate queue
- * registration is added when multi-queue RX is enabled.
+ * Registers queue 0 via ibmveth_register_logical_lan(), then subordinate
+ * queues 1..N when multi-queue mode is enabled.
*
- * Return: 0 on success, -ENONET if queue 0 registration fails
+ * Return: 0 on success, -ENONET if queue 0 registration fails,
+ * -EOPNOTSUPP if firmware rejects a subordinate queue (H_FUNCTION;
+ * this function latches mq_fallback), -EIO on other subordinate
+ * failures
*/
static int
ibmveth_register_rx_queues(struct ibmveth_adapter *adapter, u64 mac_address)
@@ -1617,7 +1802,8 @@ ibmveth_register_rx_queues(struct ibmveth_adapter *adapter, u64 mac_address)
struct net_device *netdev = adapter->netdev;
union ibmveth_buf_desc rxq_desc;
unsigned long lpar_rc;
- int rc;
+ unsigned int num;
+ int i, rc;
rxq_desc.fields.flags_len = IBMVETH_BUF_VALID |
adapter->rx_queue[0].queue_len;
@@ -1642,9 +1828,67 @@ ibmveth_register_rx_queues(struct ibmveth_adapter *adapter, u64 mac_address)
return -ENONET;
}
+ num = ibmveth_get_num_rx_queues(adapter);
+ if (num == 1 || !adapter->multi_queue) {
+ netdev_dbg(netdev,
+ "registered 1 RX queue with hypervisor (single-queue mode)\n");
+ return 0;
+ }
+
+ netdev_dbg(netdev, "Registering %u subordinate queues (1-%u)\n",
+ num - 1, num - 1);
+
+ for (i = 1; i < num; i++) {
+ rc = ibmveth_register_single_rx_queue(adapter, i);
+ if (rc) {
+ /* Firmware MQ gone: fall back to SQ on next open. */
+ if (rc == -EOPNOTSUPP)
+ adapter->mq_fallback = true;
+ goto err_unregister;
+ }
+ }
+
netdev_dbg(netdev,
- "registered 1 RX queue with hypervisor (single-queue mode)\n");
+ "registered %u RX queues with hypervisor (multi-queue mode)\n",
+ num);
+
return 0;
+
+err_unregister:
+ ibmveth_dispose_subordinate_irq_mappings(adapter);
+ ibmveth_free_all_queues(adapter);
+ return rc;
+}
+
+/**
+ * ibmveth_apply_mq_fallback - Drop multi-queue mode after firmware rejection
+ * @adapter: ibmveth adapter
+ *
+ * mq_fallback is set when firmware rejects MQ (subordinate register or
+ * buffer-add H_FUNCTION). Apply only at the start of open after teardown so
+ * num_rx_queues is not shrunk while IRQ/NAPI still reference higher queues.
+ * Consumes the flag and clears multi_queue, which is what makes the
+ * single-queue decision permanent for this device.
+ */
+static void ibmveth_apply_mq_fallback(struct ibmveth_adapter *adapter)
+{
+ struct net_device *netdev = adapter->netdev;
+
+ if (!adapter->mq_fallback)
+ return;
+
+ adapter->mq_fallback = false;
+
+ if (!adapter->multi_queue && ibmveth_get_num_rx_queues(adapter) == 1)
+ return;
+
+ netdev_warn(netdev,
+ "Falling back to single RX queue (firmware MQ unavailable)\n");
+ adapter->multi_queue = false;
+ ibmveth_publish_num_rx_queues(adapter, 1);
+ /* real_num_rx_queues is set later in open after resources exist. */
+ if (adapter->rx_buffers_per_hcall > IBMVETH_MAX_RX_REGULAR)
+ adapter->rx_buffers_per_hcall = IBMVETH_MAX_RX_REGULAR;
}
static int ibmveth_open(struct net_device *netdev)
@@ -1657,6 +1901,8 @@ static int ibmveth_open(struct net_device *netdev)
netdev_dbg(netdev, "open starting\n");
+ ibmveth_apply_mq_fallback(adapter);
+
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
rxq_entries += adapter->rx_buff_pool[0][i].size;
@@ -1676,18 +1922,34 @@ static int ibmveth_open(struct net_device *netdev)
if (rc)
goto out_free_buffer_pools;
- rc = netif_set_real_num_rx_queues(netdev, adapter->num_rx_queues);
+ rc = netif_set_real_num_rx_queues(netdev,
+ ibmveth_get_num_rx_queues(adapter));
+
if (rc) {
netdev_err(netdev, "failed to set number of rx queues\n");
goto out_unregister_queues;
}
+ /*
+ * Post buffers before setup_rx_interrupts(). MQ setup then unmasks
+ * PHYP; SQ setup leaves PHYP masked. Scheduling NAPI only when a
+ * descriptor is already pending is not enough: after ifdown/up
+ * (RX=8, no -L) NAPI can be idle with nothing pending and the
+ * queue stays dead (TX OK, ARP/RX fail).
+ * restart_rx_queue() replenishes, schedules NAPI, and unmasks if
+ * prep fails.
+ */
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
+ netdev_dbg(netdev, "initial replenish cycle for queue %d\n", i);
+ ibmveth_replenish_task(adapter, i);
+ }
+
rc = ibmveth_setup_rx_interrupts(adapter);
if (rc)
goto out_free_all_queues; /* setup already disposed IRQs */
- netdev_dbg(netdev, "initial replenish cycle\n");
- ibmveth_schedule_rx_queue(adapter, 0);
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
+ ibmveth_restart_rx_queue(adapter, i);
rc = ibmveth_alloc_tx_resources(adapter);
if (rc)
@@ -1722,6 +1984,7 @@ static int ibmveth_open(struct net_device *netdev)
static int ibmveth_close(struct net_device *netdev)
{
struct ibmveth_adapter *adapter = netdev_priv(netdev);
+ int i;
/* Gate on opened, not IFF_UP: pool_store/change_mtu close+open can
* leave IFF_UP set after a failed reopen.
@@ -1742,7 +2005,8 @@ static int ibmveth_close(struct net_device *netdev)
/* Wait for softirq/poll that already passed shutdown checks. */
synchronize_net();
- ibmveth_update_rx_no_buffer(adapter);
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
+ ibmveth_update_rx_no_buffer(adapter, i);
ibmveth_free_all_queues(adapter);
/* Free TX LTBs after quiesce and after H_FREE_LOGICAL_LAN so xmit
* cannot touch unmapped bounce buffers while the LAN is live.
@@ -1777,6 +2041,10 @@ static void ibmveth_reset(struct work_struct *w)
netdev_dbg(netdev, "reset starting\n");
rtnl_lock();
+ if (netdev->reg_state != NETREG_REGISTERED) {
+ rtnl_unlock();
+ return;
+ }
dev_close(adapter->netdev);
dev_open(adapter->netdev, NULL);
@@ -2538,6 +2806,7 @@ static int ibmveth_poll_deliver_frame(struct napi_struct *napi,
ibmveth_rx_csum_helper(skb, adapter);
}
+ skb_record_rx_queue(skb, queue_index);
napi_gro_receive(napi, skb);
netdev->stats.rx_packets++;
@@ -2556,7 +2825,7 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
queue_index = napi - adapter->napi;
if (WARN_ON(queue_index < 0 ||
- queue_index >= adapter->num_rx_queues)) {
+ queue_index >= ibmveth_get_num_rx_queues(adapter))) {
if (budget)
napi_complete_done(napi, 0);
return 0;
@@ -2649,10 +2918,11 @@ static irqreturn_t ibmveth_interrupt(int irq, void *dev_instance)
qindex = napi - adapter->napi;
/*
- * Quiet on out-of-range: teardown can leave a residual IRQ after the
- * live count drops. Do not WARN-storm; return IRQ_NONE until free_irq.
+ * Quiet on out-of-range: scale-down publishes a lower live count
+ * before free_irq(). A residual IRQ must not WARN-storm; return
+ * IRQ_NONE until the handler is removed.
*/
- if (qindex < 0 || qindex >= adapter->num_rx_queues)
+ if (qindex < 0 || qindex >= ibmveth_get_num_rx_queues(adapter))
return IRQ_NONE;
ibmveth_schedule_rx_queue(adapter, qindex);
@@ -2761,9 +3031,14 @@ static int ibmveth_change_mtu(struct net_device *dev, int new_mtu)
static void ibmveth_poll_controller(struct net_device *dev)
{
struct ibmveth_adapter *adapter = netdev_priv(dev);
+ unsigned int num = ibmveth_get_num_rx_queues(adapter);
+ int i;
- ibmveth_replenish_task(adapter, 0);
- ibmveth_schedule_rx_queue(adapter, 0);
+ for (i = 0; i < num; i++)
+ ibmveth_replenish_task(adapter, i);
+
+ for (i = 0; i < num; i++)
+ ibmveth_schedule_rx_queue(adapter, i);
}
#endif
@@ -2781,8 +3056,7 @@ static unsigned long ibmveth_get_desired_dma(struct vio_dev *vdev)
struct ibmveth_adapter *adapter;
struct iommu_table *tbl;
unsigned long ret;
- int i;
- int rxqentries = 1;
+ int i, q;
tbl = get_iommu_table_base(&vdev->dev);
@@ -2792,23 +3066,38 @@ static unsigned long ibmveth_get_desired_dma(struct vio_dev *vdev)
adapter = netdev_priv(netdev);
- ret = IBMVETH_BUFF_LIST_SIZE + IBMVETH_FILT_LIST_SIZE;
+ /* One buffer list page per RX queue; filter list is shared. */
+ ret = IBMVETH_BUFF_LIST_SIZE * ibmveth_get_num_rx_queues(adapter) +
+ IBMVETH_FILT_LIST_SIZE;
ret += IOMMU_PAGE_ALIGN(netdev->mtu, tbl);
/* add size of mapped tx buffers */
ret += IOMMU_PAGE_ALIGN(IBMVETH_MAX_TX_BUF_SIZE, tbl);
- for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
- /* add the size of the active receive buffers */
- if (adapter->rx_buff_pool[0][i].active)
- ret +=
- adapter->rx_buff_pool[0][i].size *
- IOMMU_PAGE_ALIGN(adapter->rx_buff_pool[0][i].
- buff_size, tbl);
- rxqentries += adapter->rx_buff_pool[0][i].size;
- }
- /* add the size of the receive queue entries */
- ret += IOMMU_PAGE_ALIGN(
- rxqentries * sizeof(struct ibmveth_rx_q_entry), tbl);
+ /*
+ * Pool metadata for queues 1+ is copied from queue 0 at open.
+ * Always size from pool 0 x num_rx_queues.
+ *
+ * CMO (Power9 and earlier) and MQ firmware (Power11+) do not
+ * coexist, so a CMO partition always sizes one queue here. The
+ * MQ terms are defensive only.
+ */
+ for (q = 0; q < ibmveth_get_num_rx_queues(adapter); q++) {
+ int rxqentries = 1;
+
+ for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
+ struct ibmveth_buff_pool *bpool =
+ &adapter->rx_buff_pool[0][i];
+
+ if (bpool->active)
+ ret += bpool->size *
+ IOMMU_PAGE_ALIGN(bpool->buff_size, tbl);
+ rxqentries += bpool->size;
+ }
+
+ /* add the size of the receive queue entries */
+ ret += IOMMU_PAGE_ALIGN(rxqentries *
+ sizeof(struct ibmveth_rx_q_entry), tbl);
+ }
return ret;
}
@@ -2873,9 +3162,45 @@ static const struct net_device_ops ibmveth_netdev_ops = {
#endif
};
+/**
+ * ibmveth_pool_kobj_release - Mark a pool kobject finished
+ * @kobj: kobject embedded in the pool
+ *
+ * Pool kobjects live in netdev_priv(). Last put waits for this
+ * before free_netdev() so DEBUG_KOBJECT_RELEASE delayed cleanup
+ * does not run on freed memory.
+ */
+static void ibmveth_pool_kobj_release(struct kobject *kobj)
+{
+ struct ibmveth_buff_pool *pool = container_of(kobj,
+ struct ibmveth_buff_pool,
+ kobj);
+
+ complete(&pool->released);
+}
+
+/**
+ * ibmveth_put_pool_kobjs - Drop pool kobjects and wait for release
+ * @adapter: ibmveth adapter
+ * @pools_ready: number of queue-0 pools that were initialized
+ *
+ * Put every initialized pool kobject, then wait so free_netdev()
+ * cannot race DEBUG_KOBJECT_RELEASE delayed cleanup.
+ */
+static void ibmveth_put_pool_kobjs(struct ibmveth_adapter *adapter,
+ int pools_ready)
+{
+ int i;
+
+ for (i = 0; i < pools_ready; i++)
+ kobject_put(&adapter->rx_buff_pool[0][i].kobj);
+ for (i = 0; i < pools_ready; i++)
+ wait_for_completion(&adapter->rx_buff_pool[0][i].released);
+}
+
static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
{
- int rc, i, mac_len;
+ int rc, i, mac_len, pools_ready = 0;
struct net_device *netdev;
struct ibmveth_adapter *adapter;
unsigned char *mac_addr_p;
@@ -2910,7 +3235,8 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
return -EINVAL;
}
- netdev = alloc_etherdev_mqs(sizeof(struct ibmveth_adapter), IBMVETH_MAX_QUEUES, 1);
+ netdev = alloc_etherdev_mqs(sizeof(struct ibmveth_adapter),
+ IBMVETH_MAX_QUEUES, IBMVETH_MAX_RX_QUEUES);
if (!netdev)
return -ENOMEM;
@@ -2933,7 +3259,9 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
adapter->mcastFilterSize = be32_to_cpu(*mcastFilterSize_p);
ibmveth_init_link_settings(netdev);
- netif_napi_add_weight(netdev, &adapter->napi[0], ibmveth_poll, 16);
+ for (i = 0; i < IBMVETH_MAX_RX_QUEUES; i++)
+ netif_napi_add_weight(netdev, &adapter->napi[i],
+ ibmveth_poll, 16);
netdev->irq = dev->irq;
netdev->netdev_ops = &ibmveth_netdev_ops;
@@ -2965,16 +3293,30 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
netdev->features |= NETIF_F_FRAGLIST;
}
- /* Initialize queue count - always 1 for now */
- adapter->multi_queue = 0;
- adapter->num_rx_queues = IBMVETH_DEFAULT_RX_QUEUES;
+ if (ret == H_SUCCESS &&
+ (ret_attr & IBMVETH_ILLAN_RX_MULTI_QUEUE_SUPPORT)) {
+ adapter->multi_queue = true;
+ ibmveth_publish_num_rx_queues(adapter,
+ min(num_online_cpus(),
+ IBMVETH_DEFAULT_QUEUES));
+ netdev_dbg(netdev, "RX multi queue mode enabled: %u queues\n",
+ ibmveth_get_num_rx_queues(adapter));
+ } else {
+ adapter->multi_queue = false;
+ ibmveth_publish_num_rx_queues(adapter,
+ IBMVETH_DEFAULT_RX_QUEUES);
+ }
if (ret == H_SUCCESS &&
(ret_attr & IBMVETH_ILLAN_RX_MULTI_BUFF_SUPPORT)) {
- adapter->rx_buffers_per_hcall = IBMVETH_MAX_RX_REGULAR;
+ if (adapter->multi_queue)
+ adapter->rx_buffers_per_hcall =
+ IBMVETH_MAX_RX_PER_HCALL;
+ else
+ adapter->rx_buffers_per_hcall = IBMVETH_MAX_RX_REGULAR;
netdev_dbg(netdev,
"RX Multi-buffer hcall supported by FW, batch set to %u\n",
- adapter->rx_buffers_per_hcall);
+ adapter->rx_buffers_per_hcall);
} else {
adapter->rx_buffers_per_hcall = 1;
netdev_dbg(netdev,
@@ -2992,15 +3334,30 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
struct kobject *kobj = &adapter->rx_buff_pool[0][i].kobj;
- int error;
ibmveth_init_buffer_pool(&adapter->rx_buff_pool[0][i], i,
pool_count[i], pool_size[i],
pool_active[i]);
- error = kobject_init_and_add(kobj, &ktype_veth_pool,
- &dev->dev.kobj, "pool%d", i);
- if (!error)
- kobject_uevent(kobj, KOBJ_ADD);
+ init_completion(&adapter->rx_buff_pool[0][i].released);
+ rc = kobject_init_and_add(kobj, &ktype_veth_pool,
+ &dev->dev.kobj, "pool%d", i);
+ if (rc) {
+ struct ibmveth_buff_pool *pool =
+ &adapter->rx_buff_pool[0][i];
+
+ dev_err(&dev->dev,
+ "failed to create pool%d kobject: %d\n", i, rc);
+ /* init_and_add takes a ref even on failure */
+ kobject_put(kobj);
+ wait_for_completion(&pool->released);
+ ibmveth_put_pool_kobjs(adapter, pools_ready);
+ dev_set_drvdata(&dev->dev, NULL);
+ free_netdev(netdev);
+ return rc;
+ }
+
+ pools_ready++;
+ kobject_uevent(kobj, KOBJ_ADD);
}
rc = netif_set_real_num_tx_queues(netdev, min(num_online_cpus(),
@@ -3008,9 +3365,29 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
if (rc) {
netdev_dbg(netdev, "failed to set number of tx queues rc=%d\n",
rc);
+ ibmveth_put_pool_kobjs(adapter, pools_ready);
+ dev_set_drvdata(&dev->dev, NULL);
free_netdev(netdev);
return rc;
}
+
+ /*
+ * alloc_etherdev_mqs() sized real_num_rx_queues to MAX. Match the
+ * advertised default (or SQ 1) before register_netdev so down-state
+ * readers agree with adapter->num_rx_queues / ethtool -l.
+ */
+ rc = netif_set_real_num_rx_queues(netdev,
+ ibmveth_get_num_rx_queues(adapter));
+
+ if (rc) {
+ netdev_dbg(netdev, "failed to set number of rx queues rc=%d\n",
+ rc);
+ ibmveth_put_pool_kobjs(adapter, pools_ready);
+ dev_set_drvdata(&dev->dev, NULL);
+ free_netdev(netdev);
+ return rc;
+ }
+
adapter->tx_ltb_size = PAGE_ALIGN(IBMVETH_MAX_TX_BUF_SIZE);
for (i = 0; i < IBMVETH_MAX_QUEUES; i++)
adapter->tx_ltb_ptr[i] = NULL;
@@ -3024,6 +3401,8 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
if (rc) {
netdev_dbg(netdev, "failed to register netdev rc=%d\n", rc);
+ ibmveth_put_pool_kobjs(adapter, pools_ready);
+ dev_set_drvdata(&dev->dev, NULL);
free_netdev(netdev);
return rc;
}
@@ -3037,14 +3416,17 @@ static void ibmveth_remove(struct vio_dev *dev)
{
struct net_device *netdev = dev_get_drvdata(&dev->dev);
struct ibmveth_adapter *adapter = netdev_priv(netdev);
- int i;
- cancel_work_sync(&adapter->work);
-
- for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
- kobject_put(&adapter->rx_buff_pool[0][i].kobj);
+ ibmveth_put_pool_kobjs(adapter, IBMVETH_NUM_BUFF_POOLS);
+ /*
+ * Unregister first so NAPI/xmit cannot re-arm reset work after we
+ * cancel it. cancel_work_sync() before unregister left a window
+ * where poll could schedule_work() and the worker ran after
+ * free_netdev().
+ */
unregister_netdev(netdev);
+ cancel_work_sync(&adapter->work);
free_netdev(netdev);
dev_set_drvdata(&dev->dev, NULL);
@@ -3195,7 +3577,6 @@ static ssize_t veth_pool_store(struct kobject *kobj, struct attribute *attr,
return rc;
}
-
#define ATTR(_name, _mode) \
struct attribute veth_##_name##_attr = { \
.name = __stringify(_name), .mode = _mode, \
@@ -3219,7 +3600,7 @@ static const struct sysfs_ops veth_pool_ops = {
};
static struct kobj_type ktype_veth_pool = {
- .release = NULL,
+ .release = ibmveth_pool_kobj_release,
.sysfs_ops = &veth_pool_ops,
.default_groups = veth_pool_groups,
};
@@ -3228,8 +3609,12 @@ static int ibmveth_resume(struct device *dev)
{
struct net_device *netdev = dev_get_drvdata(dev);
struct ibmveth_adapter *adapter = netdev_priv(netdev);
+ unsigned int num = ibmveth_get_num_rx_queues(adapter);
+ int i;
+
+ for (i = 0; i < num; i++)
+ ibmveth_schedule_rx_queue(adapter, i);
- ibmveth_schedule_rx_queue(adapter, 0);
return 0;
}
diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h
index dc0e5c25209a..84e703692aae 100644
--- a/drivers/net/ethernet/ibm/ibmveth.h
+++ b/drivers/net/ethernet/ibm/ibmveth.h
@@ -15,6 +15,7 @@
#define _IBMVETH_H
#include <linux/spinlock_types.h>
+#include <linux/completion.h>
/* constants for H_MULTICAST_CTRL */
#define IbmVethMcastReceptionModifyBit 0x80000UL
@@ -30,6 +31,7 @@
#define IbmVethMcastRemoveFilter 0x2UL
#define IbmVethMcastClearFilterTable 0x3UL
+#define IBMVETH_ILLAN_RX_MULTI_QUEUE_SUPPORT 0x0000000000080000UL
#define IBMVETH_ILLAN_RX_MULTI_BUFF_SUPPORT 0x0000000000040000UL
#define IBMVETH_ILLAN_LRG_SR_ENABLED 0x0000000000010000UL
#define IBMVETH_ILLAN_LRG_SND_SUPPORT 0x0000000000008000UL
@@ -265,7 +267,7 @@ static inline long h_illan_attributes(unsigned long unit_address,
#define IBMVETH_MAX_TX_BUF_SIZE (1024 * 64)
#define IBMVETH_MAX_QUEUES 16U
#define IBMVETH_DEFAULT_QUEUES 8U
-#define IBMVETH_MAX_RX_QUEUES 1U
+#define IBMVETH_MAX_RX_QUEUES 16U
#define IBMVETH_DEFAULT_RX_QUEUES 1U
#define IBMVETH_MAX_RX_REGULAR 8U
#define IBMVETH_MAX_RX_PER_HCALL 12U
@@ -290,6 +292,7 @@ struct ibmveth_buff_pool {
struct sk_buff **skbuff;
int active;
struct kobject kobj;
+ struct completion released;
};
struct ibmveth_rx_q {
@@ -320,8 +323,9 @@ struct ibmveth_adapter {
struct ibmveth_rx_q rx_queue[IBMVETH_MAX_RX_QUEUES];
u64 queue_handle[IBMVETH_MAX_RX_QUEUES];
unsigned int queue_irq[IBMVETH_MAX_RX_QUEUES];
- int multi_queue;
+ bool multi_queue;
unsigned int num_rx_queues;
+ bool mq_fallback;
/* Lifetime: true after successful ndo_open until close clears it. */
bool opened;
/* Lifetime: true while RX IRQ handlers / NAPI are installed. */
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 10/15] ibmveth: Enable multi-queue RX receive path
2026-09-25 18:38 ` [PATCH net-next v7 10/15] ibmveth: Enable multi-queue RX receive path Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found 7 potential
issue(s) to consider.
Critical: 0 · High: 0 (2) · Medium: 4 · Low: 3
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] The commit message says the ibmveth_remove()
unregister_netdev()/cancel_work_sync() reorder and the ibmveth_reset()…
- [Medium] With several MQ NAPI instances running at once, the adapter-
wide RX and replenish counters are still updated with plain read-…
- [Medium] In MQ open, ibmveth_setup_rx_interrupts() unmasks PHYP on
every queue (ibmveth_enable_irq) before ibmveth_open() calls…
- [Medium] When MQ interrupt setup fails partway, RX rings can be freed
while a poll that has already completed NAPI still reads them.
- [Low] At this commit, ibmveth_get_channels() reports max_rx = rx_count
= netdev->real_num_rx_queues, which is now min(num_online_cpus(), 8)…
- [Low] ibmveth_apply_mq_fallback() calls
ibmveth_publish_num_rx_queues(adapter, 1) but leaves…
- [Low] The new retry loops in ibmveth_register_logical_lan_queue() re-
issue H_REG_LOGICAL_LAN_QUEUE, and H_FREE_LOGICAL_LAN_QUEUE on the…
Pre-existing issues:
- [High] At this commit, ibmveth_poll_controller() replenishes and
schedules every live RX queue without checking adapter->opened.
- [High] This problem wasn't introduced by this patch, but it is a
remotely triggerable crash on the RX path that this patch extends to…
--
ibmveth: Enable multi-queue RX receive path
This commit turns on the multi-queue RX path when firmware advertises
IBMVETH_ILLAN_RX_MULTI_QUEUE_SUPPORT. It raises IBMVETH_MAX_RX_QUEUES to 16,
defaults to min(num_online_cpus(), 8) queues, registers the subordinate
queues through H_REG_LOGICAL_LAN_QUEUE, and wires up per-queue IRQs and
NAPI. It also adds an mq_fallback latch for H_FUNCTION rejections, reorders
ibmveth_remove(), gates ibmveth_reset() on reg_state, and reworks the pool
kobject lifetime.
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index 7b6c0283e5c32..3f31793645a31 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
[ ... ]
> @@ -993,15 +1014,21 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
> * because there was not a buffer in the buffer list capable of holding
> * the frame.
> */
> -static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter)
> +static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter,
> + int queue_index)
> {
> __be64 *p;
> + u64 drops;
>
> - if (!adapter->buffer_list_addr[0])
> + if (queue_index < 0 ||
> + queue_index >= ibmveth_get_num_rx_queues(adapter) ||
> + !adapter->buffer_list_addr[queue_index])
> return;
>
> - p = adapter->buffer_list_addr[0] + 4096 - 8;
> - adapter->rx_no_buffer = be64_to_cpup(p);
> + p = adapter->buffer_list_addr[queue_index] + 4096 - 8;
> + drops = be64_to_cpup(p);
> +
> + adapter->rx_no_buffer = drops;
> }
[Severity: Medium]
With several MQ NAPI instances running at once, adapter->rx_no_buffer is
overwritten by whichever queue ran ibmveth_update_rx_no_buffer() last. Can
it go backwards as a result?
Other adapter-wide counters have the same problem. These are plain
read-modify-writes with no lock, or with only a per-queue replenish_lock,
so concurrent updates can be lost:
- netdev->stats.rx_packets, rx_bytes and rx_large_packets in
ibmveth_poll_deliver_frame()
- rx_invalid_buffer in ibmveth_poll_bump_invalid()
- the replenish_* counters in ibmveth_replenish_task() and
ibmveth_replenish_buffer_pool()
The commit message acknowledges the rx_* and rx_no_buffer cases. The next
patch, "ibmveth: Add per-queue RX and TX statistics collection", moves all
of these counters, including replenish_*, into per-queue rx_qstats[]. The
commit message doesn't mention the replenish_* counters, though.
[ ... ]
> @@ -1086,6 +1114,27 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
> fail.filled, fail.lpar_rc, fail.batch);
> }
>
> +/**
> + * ibmveth_restart_rx_queue - Post buffers and ensure Q can take RX
[ ... ]
> +static void ibmveth_restart_rx_queue(struct ibmveth_adapter *adapter,
> + int qindex)
> +{
> + ibmveth_replenish_task(adapter, qindex);
> + if (ibmveth_schedule_rx_queue(adapter, qindex))
> + return;
> +
> + ibmveth_enable_irq(adapter, qindex);
> +}
[Severity: Medium]
In MQ mode, ibmveth_setup_rx_interrupts() has already unmasked PHYP on every
queue, with buffers posted, by the time ibmveth_open() calls this function.
Suppose a frame arrives in between:
ibmveth_interrupt()
ibmveth_schedule_rx_queue()
napi_schedule_prep() succeeds
ibmveth_disable_irq()
__napi_schedule()
ibmveth_restart_rx_queue()
ibmveth_schedule_rx_queue()
napi_schedule_prep() fails, NAPI already scheduled
ibmveth_enable_irq()
Doesn't this leave the queue unmasked while NAPI still owns it?
ibmveth_schedule_rx_queue() only masks when napi_schedule_prep() succeeds:
if (napi_schedule_prep(napi)) {
ibmveth_disable_irq(adapter, qindex);
__napi_schedule(napi);
return true;
}
return false;
Later interrupts during that NAPI run only set NAPIF_STATE_MISSED. MISSED
keeps napi_complete_done() returning false, so under sustained traffic the
queue could take an interrupt per event until a poll completes cleanly.
Could the fallback tell "NAPI already scheduled" apart from "NAPI will never
run" before it unmasks? This code is unchanged at the end of the series.
> @@ -1562,6 +1611,138 @@ static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
> return rc;
> }
>
> +/**
> + * ibmveth_register_logical_lan_queue - Register subordinate queue with
[ ... ]
> + do {
> + lpar_rc = h_register_logical_lan_queue(ua, bl,
> + rxq_desc.desc, &handle,
> + &hwirq);
> + } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
[ ... ]
> + do {
> + free_rc = h_free_logical_lan_queue(ua, handle);
> + } while (H_IS_LONG_BUSY(free_rc) ||
> + (free_rc == H_BUSY));
[Severity: Low]
Both loops re-issue the hcall immediately on H_IS_LONG_BUSY(). Should they
honour the requested delay through get_longbusy_msecs() and sleep, or at
least call cond_resched()?
This runs in ndo_open under RTNL, once per subordinate queue. If firmware
keeps returning H_LONG_BUSY_ORDER_*, a CPU would spin with RTNL held. The
same busy-loop shape already exists for h_free_logical_lan() in this
driver.
[ ... ]
> @@ -1642,9 +1828,67 @@ ibmveth_register_rx_queues(struct ibmveth_adapter *adapter, u64 mac_address)
[ ... ]
> +static void ibmveth_apply_mq_fallback(struct ibmveth_adapter *adapter)
> +{
[ ... ]
> + netdev_warn(netdev,
> + "Falling back to single RX queue (firmware MQ unavailable)\n");
> + adapter->multi_queue = false;
> + ibmveth_publish_num_rx_queues(adapter, 1);
> + /* real_num_rx_queues is set later in open after resources exist. */
> + if (adapter->rx_buffers_per_hcall > IBMVETH_MAX_RX_REGULAR)
> + adapter->rx_buffers_per_hcall = IBMVETH_MAX_RX_REGULAR;
> }
[Severity: Low]
This publishes 1 and consumes mq_fallback, but netdev->real_num_rx_queues
keeps the old MQ count until ibmveth_open() reaches
netif_set_real_num_rx_queues(). That only happens after these succeed:
- ibmveth_alloc_filter_list()
- ibmveth_alloc_rx_queues()
- ibmveth_alloc_buffer_pools()
- ibmveth_register_rx_queues()
If one of them fails, the device stays down with real_num_rx_queues at N
while num_rx_queues is 1, and the flag is already cleared. Doesn't that
break the invariant stated in the new comment in ibmveth_probe()?
Match the advertised default (or SQ 1) before register_netdev so
down-state readers agree with adapter->num_rx_queues / ethtool -l.
At this commit, ethtool -l would report N queues.
At the end of the series, get_channels() reads num_rx_queues, but the
queues/rx-* sysfs entries and the netdev-genl queue lists would still be
stale. A later down-state ethtool -L rx 1 would not resync them, because
goal_rx equals num_rx_queues.
Would calling netif_set_real_num_rx_queues() here help?
[ ... ]
> @@ -1676,18 +1922,34 @@ static int ibmveth_open(struct net_device *netdev)
[ ... ]
> + for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
> + netdev_dbg(netdev, "initial replenish cycle for queue %d\n", i);
> + ibmveth_replenish_task(adapter, i);
> + }
> +
> rc = ibmveth_setup_rx_interrupts(adapter);
> if (rc)
> goto out_free_all_queues; /* setup already disposed IRQs */
[Severity: Medium]
Suppose ibmveth_enable_irq() fails for queue k in
ibmveth_setup_rx_interrupts(). Queues 0 to k-1 are already unmasked, so
their NAPI may be polling.
Can a poll that has already passed napi_complete_done() still be reading
rx_queue[q].queue_addr when this path frees it?
ibmveth_poll() on queue q < k:
napi_complete_done()
passes ibmveth_poll_stopping()
err_disable_napi:
napi_disable() returns, SCHED clear
synchronize_irq() / free_irq()
ibmveth_open():
goto out_free_all_queues
ibmveth_free_buffer_pools()
ibmveth_cleanup_rx_resources()
ibmveth_enable_irq()
ibmveth_rxq_pending_buffer()
reads rx_queue[q].queue_addr[index]
The TX allocation failure branch in ibmveth_open() handles exactly this
window:
out_cleanup_rx_interrupts:
ibmveth_cleanup_rx_interrupts(adapter);
/* As in close(): a poll past napi_complete_done() may still run. */
synchronize_net();
Should the setup_rx_interrupts() failure path also call synchronize_net()
before freeing the rings? This code is unchanged at the end of the series.
[ ... ]
> @@ -2538,6 +2806,7 @@ static int ibmveth_poll_deliver_frame(struct napi_struct *napi,
> ibmveth_rx_csum_helper(skb, adapter);
> }
>
> + skb_record_rx_queue(skb, queue_index);
> napi_gro_receive(napi, skb);
[Severity: High]
This isn't a bug introduced by this patch, but it sits on the RX path that
this patch now enables for every new queue.
ibmveth_poll_deliver_frame() calls ibmveth_rx_csum_helper() for any frame
carrying IBMVETH_RXQ_CSUM_GOOD, and the sending partition controls that
flag. For ETH_P_IP with iph->check == 0xffff, the helper does:
iph->check = 0;
iph->check = ip_fast_csum((unsigned char *)iph,
iph->ihl);
It never checks that iph->ihl >= 5. On powerpc, ip_fast_csum_nofold() does:
for (i = 0; i < ihl - 1; i++, ptr++)
s += *ptr;
ihl is unsigned, so ihl == 0 makes the limit UINT_MAX.
Can a frame from another partition on the same VLAN, with ihl == 0 and
check == 0xffff, make this read far past the packet in softirq context?
ip_rcv_core() would reject such a header, but only after this helper has
run.
[ ... ]
> @@ -2761,9 +3031,14 @@ static int ibmveth_change_mtu(struct net_device *dev, int new_mtu)
> static void ibmveth_poll_controller(struct net_device *dev)
> {
> struct ibmveth_adapter *adapter = netdev_priv(dev);
> + unsigned int num = ibmveth_get_num_rx_queues(adapter);
> + int i;
>
> - ibmveth_replenish_task(adapter, 0);
> - ibmveth_schedule_rx_queue(adapter, 0);
> + for (i = 0; i < num; i++)
> + ibmveth_replenish_task(adapter, i);
> +
> + for (i = 0; i < num; i++)
> + ibmveth_schedule_rx_queue(adapter, i);
> }
[Severity: High]
This is a pre-existing issue for queue 0, but this change extends it to
every RX queue.
ibmveth_poll_controller() does not check adapter->opened. These callers
invoke ibmveth_close() directly and leave netif_running() true:
- ibmveth_set_tso()
- ibmveth_set_csum_offload()
- ibmveth_change_mtu()
- veth_pool_store()
That means netconsole's netpoll can run during their teardown:
netpoll
ibmveth_poll_controller()
ibmveth_replenish_task()
reads pool->free_map, skbuff, dma_addr
ibmveth_close()
ibmveth_free_buffer_pools()
ibmveth_free_buffer_pool()
kfree(pool->free_map)
ibmveth_free_buffer_pool() frees these arrays without taking
replenish_lock. Can this be a use-after-free?
Later in the series, "ibmveth: Complete set_channels down-path and
mq_fallback max_rx cap" adds an early return on !adapter->opened in
ibmveth_poll_controller(), which seems to close this. Would it be better to
add that gate in this patch, where the loop over every queue is introduced?
[ ... ]
> @@ -2965,16 +3293,30 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
> netdev->features |= NETIF_F_FRAGLIST;
> }
>
> - /* Initialize queue count - always 1 for now */
> - adapter->multi_queue = 0;
> - adapter->num_rx_queues = IBMVETH_DEFAULT_RX_QUEUES;
> + if (ret == H_SUCCESS &&
> + (ret_attr & IBMVETH_ILLAN_RX_MULTI_QUEUE_SUPPORT)) {
> + adapter->multi_queue = true;
> + ibmveth_publish_num_rx_queues(adapter,
> + min(num_online_cpus(),
> + IBMVETH_DEFAULT_QUEUES));
[Severity: Low]
On MQ firmware, real_num_rx_queues now becomes min(num_online_cpus(), 8),
and ibmveth_get_channels() reports it as both max_rx and rx_count. The
ethtool core will therefore accept a smaller rx_count. However,
ibmveth_set_channels() only reads channels->tx_count:
unsigned int old = netdev->real_num_tx_queues,
goal = channels->tx_count;
At this commit, does ethtool -L rx 2 return success without changing
anything?
The commit message notes that set_channels does not handle rx_count yet.
The next patch, "ibmveth: Report MQ-aware RX counts in ethtool
get_channels", returns -EOPNOTSUPP when rx_count changes. "ibmveth:
Implement incremental MQ RX queue resize" and "ibmveth: Complete
set_channels down-path and mq_fallback max_rx cap" then implement the
resize.
[ ... ]
> @@ -3037,14 +3416,17 @@ static void ibmveth_remove(struct vio_dev *dev)
> {
> struct net_device *netdev = dev_get_drvdata(&dev->dev);
> struct ibmveth_adapter *adapter = netdev_priv(netdev);
> - int i;
>
> - cancel_work_sync(&adapter->work);
> -
> - for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++)
> - kobject_put(&adapter->rx_buff_pool[0][i].kobj);
> + ibmveth_put_pool_kobjs(adapter, IBMVETH_NUM_BUFF_POOLS);
>
> + /*
> + * Unregister first so NAPI/xmit cannot re-arm reset work after we
> + * cancel it. cancel_work_sync() before unregister left a window
> + * where poll could schedule_work() and the worker ran after
> + * free_netdev().
> + */
> unregister_netdev(netdev);
> + cancel_work_sync(&adapter->work);
[Severity: Medium]
The commit message says:
This commit adds schedule_work() producers on buffer-add H_FUNCTION,
so the remove-path unregister / cancel_work_sync reorder and reset
reg_state gate land here to prevent a queued reset from racing
device teardown.
Is this attribution accurate?
Before this series, ibmveth_remove() already called cancel_work_sync()
before unregister_netdev(). In that gap, adapter->work could already be
queued from NAPI or replenish context by:
- the invalid free_map path in ibmveth_replenish_task()
- the WARN paths in ibmveth_remove_buffer_from_pool() and
ibmveth_rxq_get_buffer()
- the h_vio_signal() failure in ibmveth_poll()
The MQ schedule_work() on buffer-add H_FUNCTION came from the earlier
patch "ibmveth: Add queue-aware RX buffer submit helper for MQ". This
commit only makes it reachable and sets mq_fallback.
So the reset worker running after free_netdev() looks like a use-after-free
that predates this series. Two other lifetime bugs that predate it are also
fixed here:
- ktype_veth_pool had .release = NULL; it now has
ibmveth_pool_kobj_release() plus a completion wait before
free_netdev()
- the pool kobjects leaked on probe failure, which the commit message
already calls pre-existing
Could these three fixes be split into their own patches with Fixes: tags,
so they can reach stable?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 11/15] ibmveth: Add per-queue RX and TX statistics collection
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (9 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 10/15] ibmveth: Enable multi-queue RX receive path Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 12/15] ibmveth: Report MQ-aware RX counts in ethtool get_channels Mingming Cao
` (4 subsequent siblings)
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
MQ RX points several queues at the same adapter-wide counters, which
races the updates and leaves no way to attribute a count to a queue.
Move every counter that has more than one writer into per-queue
structs allocated at probe and freed at remove:
struct ibmveth_rx_queue_stats
struct ibmveth_tx_queue_stats
Each slot has a single writer — replenish_* under that queue's
replenish_lock, the other RX fields from that queue's NAPI, TX under
the stack's per-queue TX lock — so plain u64 is enough on this
PPC64-only driver. No atomic and no u64_stats_sync.
packets, bytes and drops go through struct netdev_stat_ops
(.get_queue_stats_rx, .get_queue_stats_tx, .get_base_stats).
.ndo_get_stats64() sums the per-queue packet and byte counters into
64-bit device totals. ethtool -S keeps only the driver-specific keys
that have no standard equivalent: interrupts, polls, large_packets,
invalid_buffers and no_buffer_drops per RX queue; large_packets,
send_failures and checksum_offload per TX queue. ETH_SS_STATS becomes
variable-length because that block scales with the live queue count.
Hypercall counters and pool%d_ keys are not added: which hcall a
batch picks is not ABI, and size/active already have sysfs (available
for every queue is patch 13).
The thirteen existing ethtool -S keys keep their exact names, their
order and their adapter-wide values, summed from the per-queue slots
on read. The storage moved; that ABI did not. The four replenish_*
counters get per-queue storage but no per-queue key of their own.
Holding that ABI while the storage moves needs the ethtool -S table
to record where each key lives. IBMVETH_STAT_OFF() could only express
an offset into struct ibmveth_adapter. Tag every entry with an enum
ibmveth_stat_src naming the struct it indexes: adapter-wide keys are
read directly, per-queue keys are summed across the slots by one pair
of offset-keyed helpers. That is what lets the field names change
while the key names do not (rx_invalid_buffer now reads
invalid_buffers, tx_send_failed reads send_failures, and the two
large_packets fields live in different structs). tx_map_failed still
reads from the adapter — it has no writer, here or in mainline —
and the three fw_enabled_* keys are capability flags, not counters.
The per-queue keys come from their own tables with the counts derived
by ARRAY_SIZE(), so get_strings(), get_ethtool_stats() and
get_sset_count() cannot drift apart.
ndo_get_stats64() walks every allocated slot rather than only the live
queues, so device totals cannot go backwards when ethtool -L shrinks
the queue count. get_base_stats() therefore reports the retired-queue
remainder rather than zero; the core sums it with the live queues it
iterates itself. Zeroing would assert that the live-queue sum is
already complete. Every field the per-queue callbacks fill is also
initialised there, because netdev_nl_stats_add() drops a field from
the device total unless both sides set it.
Give every queue a no_buffer_retired carry. PHYP's drop counter is
absolute for the buffer-list page currently mapped, so a reopen or a
queue reuse restarts it near zero. Storing only the newest absolute in
adapter->rx_no_buffer meant whichever queue ran last won, and the
value could go backwards. The carry sits beside the no_buffer_drops it
accumulates from, so both belong to one queue. The rx%d_no_buffer_drops
key reports that live page absolute on its own, so it is the one
exported value that is not monotonic; the adapter-wide rx_no_buffer
sums the two and per-queue rx-hw-drops includes both.
The remove-path unregister / cancel_work_sync reorder and reset
reg_state gate landed in the prior patch when the reset producer was
added; this patch frees the per-queue statistics arrays in remove()
between cancel_work_sync() and free_netdev().
ibmveth_probe_cleanup() also clears the vio drvdata before
free_netdev(). A probe failure never reaches ibmveth_remove(), and
CMO get_desired_dma() reads that pointer on a later rebind.
Readers do not test the arrays for NULL: both exist from before
register_netdev() until after unregister_netdev() and
cancel_work_sync(), and probe fails -ENOMEM if either allocation
does.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- close() harvests no_buffer under replenish_lock
- remove() uses put_pool_kobjs so the wait lands
before free_netdev
- remove-path unregister/cancel reorder and reset gate
moved to P10 with the reset producer
- noted: -S / get_stats64 / hw_drops mapping stays;
irqsave / free_buffer_pool / netpoll_poll_disable
stay leftover
Changes in v6:
- wrap the qstats local in replenish (81 cols)
- replenish_* per-queue u64, summed on the existing adapter-wide keys;
no per-queue replenish key; no atomics
- netdev_stat_ops for packets/bytes/drops, not private -S strings
- enum ibmveth_stat_src so existing -S keys keep names while storage
moves; ARRAY_SIZE() for the per-queue key counts
- no hcall_* keys (buffer-submit and H_SEND_LOGICAL_LAN)
- drop the three pool%d_ keys (15 ethtool entries)
- drop the fifteen qstats NULL checks outside the allocators
(five in the RX hot path)
- per-queue no_buffer_retired carry
- get_base_stats() reports the retired-queue remainder
- gate reset on NETREG_REGISTERED
- noted: harvest no_buffer on -L shrink is patch 14
Changes in v5:
- Series renumber: mailed v4 10/14 stats -> tip P11 (P09 peel;
get_channels -> P12)
- rx_no_buffer_retired + sum MAX_* slots so adapter no-buffer / qstat
totals stay monotonic across reopen and channel shrink
- probe_cleanup: clear vio drvdata before free_netdev (CMO cannot see a
freed netdev on rebind)
- remove: unregister_netdev then cancel_work_sync (no UAF reset worker)
Changes in v4:
- Merge v3's separate RX and TX stats commits into one patch.
- Introduce rx_queue_stats / tx_qstats / NUM macros here (first use).
- Allocate/free qstats at probe/remove instead of open/close.
- Report adapter-level ethtool strings by summing per-queue counters on
read; drop aggregate_* helpers.
- Sum global rx_no_buffer across MQ queues into this statistics patch.
- Cacheline-align per-queue stats; derive field counts with offsetof so
alignment padding is not counted as a statistic.
- probe_cleanup() cancels reset work, puts pool kobjects via helper from
the prior patch, and frees qstats on probe failure paths.
- Keep plain u64 qstats like existing ibmveth / ibmvnic (PPC_PSERIES).
drivers/net/ethernet/ibm/ibmveth.c | 502 +++++++++++++++++++++++++----
drivers/net/ethernet/ibm/ibmveth.h | 57 +++-
2 files changed, 488 insertions(+), 71 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 3f31793645a3..2575a7c11afb 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -38,6 +38,7 @@
#include <asm/firmware.h>
#include <net/tcp.h>
#include <net/ip6_checksum.h>
+#include <net/netdev_queues.h>
#include "ibmveth.h"
@@ -75,32 +76,101 @@ module_param(old_large_send, bool, 0444);
MODULE_PARM_DESC(old_large_send,
"Use old large send method on firmware that supports the new method");
+/**
+ * enum ibmveth_stat_src - where an ethtool -S counter is stored
+ * @IBMVETH_STAT_ADAPTER: plain u64 in struct ibmveth_adapter
+ * @IBMVETH_STAT_RX_QSUM: per-queue u64, summed over rx_qstats[]
+ * @IBMVETH_STAT_TX_QSUM: per-queue u64, summed over tx_qstats[]
+ * @IBMVETH_STAT_RX_NO_BUFFER: rx_qstats[] live-page absolute plus the
+ * absolutes carried over from pages the queue has already retired
+ *
+ * Counters live per-queue so multi-queue writers never share a field.
+ * The adapter is only ever read from ethtool, so summing there is free.
+ */
+enum ibmveth_stat_src {
+ IBMVETH_STAT_ADAPTER,
+ IBMVETH_STAT_RX_QSUM,
+ IBMVETH_STAT_TX_QSUM,
+ IBMVETH_STAT_RX_NO_BUFFER,
+};
+
struct ibmveth_stat {
char name[ETH_GSTRING_LEN];
- int offset;
+ enum ibmveth_stat_src src;
+ /* Offset into the struct named by @src. */
+ size_t off;
};
#define IBMVETH_STAT_OFF(stat) offsetof(struct ibmveth_adapter, stat)
+#define IBMVETH_RXQ_OFF(stat) offsetof(struct ibmveth_rx_queue_stats, stat)
+#define IBMVETH_TXQ_OFF(stat) offsetof(struct ibmveth_tx_queue_stats, stat)
#define IBMVETH_GET_STAT(a, off) *((u64 *)(((unsigned long)(a)) + off))
+#define IBMVETH_ADAPTER_STAT(key, field) \
+ { key, IBMVETH_STAT_ADAPTER, IBMVETH_STAT_OFF(field) }
+#define IBMVETH_RXQ_STAT(key, field) \
+ { key, IBMVETH_STAT_RX_QSUM, IBMVETH_RXQ_OFF(field) }
+#define IBMVETH_TXQ_STAT(key, field) \
+ { key, IBMVETH_STAT_TX_QSUM, IBMVETH_TXQ_OFF(field) }
+
+/*
+ * Key names and their order are ABI. Do not reorder or rename; append
+ * only, and only when the counter is worth a permanent interface.
+ */
static struct ibmveth_stat ibmveth_stats[] = {
- { "replenish_task_cycles", IBMVETH_STAT_OFF(replenish_task_cycles) },
- { "replenish_no_mem", IBMVETH_STAT_OFF(replenish_no_mem) },
- { "replenish_add_buff_failure",
- IBMVETH_STAT_OFF(replenish_add_buff_failure) },
- { "replenish_add_buff_success",
- IBMVETH_STAT_OFF(replenish_add_buff_success) },
- { "rx_invalid_buffer", IBMVETH_STAT_OFF(rx_invalid_buffer) },
- { "rx_no_buffer", IBMVETH_STAT_OFF(rx_no_buffer) },
- { "tx_map_failed", IBMVETH_STAT_OFF(tx_map_failed) },
- { "tx_send_failed", IBMVETH_STAT_OFF(tx_send_failed) },
- { "fw_enabled_ipv4_csum", IBMVETH_STAT_OFF(fw_ipv4_csum_support) },
- { "fw_enabled_ipv6_csum", IBMVETH_STAT_OFF(fw_ipv6_csum_support) },
- { "tx_large_packets", IBMVETH_STAT_OFF(tx_large_packets) },
- { "rx_large_packets", IBMVETH_STAT_OFF(rx_large_packets) },
- { "fw_enabled_large_send", IBMVETH_STAT_OFF(fw_large_send_support) }
+ IBMVETH_RXQ_STAT("replenish_task_cycles", replenish_task_cycles),
+ IBMVETH_RXQ_STAT("replenish_no_mem", replenish_no_mem),
+ IBMVETH_RXQ_STAT("replenish_add_buff_failure",
+ replenish_add_buff_failure),
+ IBMVETH_RXQ_STAT("replenish_add_buff_success",
+ replenish_add_buff_success),
+ IBMVETH_RXQ_STAT("rx_invalid_buffer", invalid_buffers),
+ { "rx_no_buffer", IBMVETH_STAT_RX_NO_BUFFER,
+ IBMVETH_RXQ_OFF(no_buffer_drops) },
+ IBMVETH_ADAPTER_STAT("tx_map_failed", tx_map_failed),
+ IBMVETH_TXQ_STAT("tx_send_failed", send_failures),
+ IBMVETH_ADAPTER_STAT("fw_enabled_ipv4_csum", fw_ipv4_csum_support),
+ IBMVETH_ADAPTER_STAT("fw_enabled_ipv6_csum", fw_ipv6_csum_support),
+ IBMVETH_TXQ_STAT("tx_large_packets", large_packets),
+ IBMVETH_RXQ_STAT("rx_large_packets", large_packets),
+ IBMVETH_ADAPTER_STAT("fw_enabled_large_send", fw_large_send_support),
+};
+
+/**
+ * struct ibmveth_qstat - a per-queue counter exposed through ethtool -S
+ * @fmt: key name, taking the queue index as its only argument
+ * @off: offset into the matching per-queue stats struct
+ *
+ * Driving the strings and the values from one table keeps the two in
+ * step; get_sset_count() derives its length from ARRAY_SIZE() so the
+ * three cannot drift apart.
+ */
+struct ibmveth_qstat {
+ const char *fmt;
+ size_t off;
+};
+
+/*
+ * Only counters with no home in the standard interfaces belong here.
+ * packets, bytes and drops are reported through netdev_stat_ops.
+ */
+static const struct ibmveth_qstat ibmveth_rx_qstat_keys[] = {
+ { "rx%d_interrupts", IBMVETH_RXQ_OFF(interrupts) },
+ { "rx%d_polls", IBMVETH_RXQ_OFF(polls) },
+ { "rx%d_large_packets", IBMVETH_RXQ_OFF(large_packets) },
+ { "rx%d_invalid_buffers", IBMVETH_RXQ_OFF(invalid_buffers) },
+ { "rx%d_no_buffer_drops", IBMVETH_RXQ_OFF(no_buffer_drops) },
};
+static const struct ibmveth_qstat ibmveth_tx_qstat_keys[] = {
+ { "tx%d_large_packets", IBMVETH_TXQ_OFF(large_packets) },
+ { "tx%d_send_failures", IBMVETH_TXQ_OFF(send_failures) },
+ { "tx%d_checksum_offload", IBMVETH_TXQ_OFF(checksum_offload) },
+};
+
+#define IBMVETH_NUM_RX_QSTATS ARRAY_SIZE(ibmveth_rx_qstat_keys)
+#define IBMVETH_NUM_TX_QSTATS ARRAY_SIZE(ibmveth_tx_qstat_keys)
+
/* simple methods of getting data from the current rxq entry */
static u32 ibmveth_rxq_flags(struct ibmveth_adapter *adapter,
int queue_index)
@@ -241,6 +311,60 @@ ibmveth_free_filter_list(struct ibmveth_adapter *adapter)
}
}
+/**
+ * ibmveth_alloc_rx_qstats - Allocate per-queue RX statistics
+ * @adapter: ibmveth adapter structure
+ *
+ * Return: 0 on success, -ENOMEM on failure
+ */
+static int ibmveth_alloc_rx_qstats(struct ibmveth_adapter *adapter)
+{
+ adapter->rx_qstats = kcalloc(IBMVETH_MAX_RX_QUEUES,
+ sizeof(*adapter->rx_qstats),
+ GFP_KERNEL);
+ if (!adapter->rx_qstats)
+ return -ENOMEM;
+
+ return 0;
+}
+
+/**
+ * ibmveth_free_rx_qstats - Free per-queue RX statistics
+ * @adapter: ibmveth adapter structure
+ */
+static void ibmveth_free_rx_qstats(struct ibmveth_adapter *adapter)
+{
+ kfree(adapter->rx_qstats);
+ adapter->rx_qstats = NULL;
+}
+
+/**
+ * ibmveth_alloc_tx_qstats - Allocate per-queue TX statistics
+ * @adapter: ibmveth adapter structure
+ *
+ * Return: 0 on success, -ENOMEM on failure
+ */
+static int ibmveth_alloc_tx_qstats(struct ibmveth_adapter *adapter)
+{
+ adapter->tx_qstats = kcalloc(IBMVETH_MAX_QUEUES,
+ sizeof(*adapter->tx_qstats),
+ GFP_KERNEL);
+ if (!adapter->tx_qstats)
+ return -ENOMEM;
+
+ return 0;
+}
+
+/**
+ * ibmveth_free_tx_qstats - Free per-queue TX statistics
+ * @adapter: ibmveth adapter structure
+ */
+static void ibmveth_free_tx_qstats(struct ibmveth_adapter *adapter)
+{
+ kfree(adapter->tx_qstats);
+ adapter->tx_qstats = NULL;
+}
+
/**
* ibmveth_alloc_rx_queues - Allocate per-queue RX resources
* @adapter: ibmveth adapter structure
@@ -839,6 +963,8 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
int queue_index,
struct ibmveth_replenish_fail *fail)
{
+ struct ibmveth_rx_queue_stats *qstats =
+ &adapter->rx_qstats[queue_index];
union ibmveth_buf_desc descs[IBMVETH_MAX_RX_PER_HCALL] = {0};
u32 remaining = pool->size - atomic_read(&pool->available);
u64 correlators[IBMVETH_MAX_RX_PER_HCALL] = {0};
@@ -865,7 +991,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
for (filled = 0; filled < min(remaining, batch); filled++) {
index = pool->free_map[free_index];
if (index == IBM_VETH_INVALID_MAP) {
- adapter->replenish_add_buff_failure++;
+ qstats->replenish_add_buff_failure++;
outcome = IBMVETH_REPLENISH_RESET_MAP;
break;
}
@@ -876,8 +1002,8 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
skb = netdev_alloc_skb(adapter->netdev,
pool->buff_size);
if (!skb) {
- adapter->replenish_no_mem++;
- adapter->replenish_add_buff_failure++;
+ qstats->replenish_no_mem++;
+ qstats->replenish_add_buff_failure++;
break;
}
@@ -892,7 +1018,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
DMA_ATTR_NO_WARN);
if (dma_mapping_error(dev, dma_addr)) {
dev_kfree_skb_any(skb);
- adapter->replenish_add_buff_failure++;
+ qstats->replenish_add_buff_failure++;
break;
}
@@ -953,7 +1079,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
}
buffers_added += filled;
- adapter->replenish_add_buff_success += filled;
+ qstats->replenish_add_buff_success += filled;
remaining -= filled;
memset(&descs, 0, sizeof(descs));
@@ -976,7 +1102,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
pool->skbuff[index] = NULL;
}
}
- adapter->replenish_add_buff_failure += filled;
+ qstats->replenish_add_buff_failure += filled;
if (lpar_rc == H_FUNCTION) {
if (adapter->multi_queue) {
@@ -1017,6 +1143,7 @@ static int ibmveth_replenish_buffer_pool(struct ibmveth_adapter *adapter,
static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter,
int queue_index)
{
+ struct ibmveth_rx_queue_stats *qstats;
__be64 *p;
u64 drops;
@@ -1028,7 +1155,18 @@ static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter,
p = adapter->buffer_list_addr[queue_index] + 4096 - 8;
drops = be64_to_cpup(p);
- adapter->rx_no_buffer = drops;
+ /*
+ * PHYP's buffer-list page counter is absolute for that page. A new
+ * page (reopen / queue reuse after -L) starts near zero; fold the
+ * previous absolute into this queue's retired carry so sums stay
+ * monotonic. Both fields belong to the queue being updated, so this
+ * stays single-writer under the queue's replenish_lock.
+ */
+ qstats = &adapter->rx_qstats[queue_index];
+
+ if (drops < qstats->no_buffer_drops)
+ qstats->no_buffer_retired += qstats->no_buffer_drops;
+ qstats->no_buffer_drops = drops;
}
/* replenish routine */
@@ -1050,10 +1188,10 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
return;
}
- adapter->replenish_task_cycles++;
-
spin_lock_irqsave(&rxq->replenish_lock, flags);
+ adapter->rx_qstats[queue_index].replenish_task_cycles++;
+
for (i = (IBMVETH_NUM_BUFF_POOLS - 1); i >= 0; i--) {
struct ibmveth_buff_pool *pool =
&adapter->rx_buff_pool[queue_index][i];
@@ -2005,8 +2143,14 @@ static int ibmveth_close(struct net_device *netdev)
/* Wait for softirq/poll that already passed shutdown checks. */
synchronize_net();
- for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[i];
+ unsigned long flags;
+
+ spin_lock_irqsave(&rxq->replenish_lock, flags);
ibmveth_update_rx_no_buffer(adapter, i);
+ spin_unlock_irqrestore(&rxq->replenish_lock, flags);
+ }
ibmveth_free_all_queues(adapter);
/* Free TX LTBs after quiesce and after H_FREE_LOGICAL_LAN so xmit
* cannot touch unmapped bounce buffers while the LAN is live.
@@ -2278,22 +2422,96 @@ static int ibmveth_set_features(struct net_device *dev,
return rc1 ? rc1 : rc2;
}
-static void ibmveth_get_strings(struct net_device *dev, u32 stringset, u8 *data)
+/*
+ * Sum per-queue counters for rare ethtool reads. The hot paths only ever
+ * touch their own queue's slot, so nothing here needs an atomic; the cost
+ * of aggregation is paid by the reader instead (ibmvnic-style).
+ *
+ * Every slot is summed, not just the live ones, so that shrinking the
+ * queue count with ethtool -L cannot make a counter go backwards.
+ */
+static u64 ibmveth_sum_rx_qstat(struct ibmveth_adapter *adapter, size_t off)
+{
+ u64 total = 0;
+ int i;
+
+ for (i = 0; i < IBMVETH_MAX_RX_QUEUES; i++)
+ total += *(u64 *)((u8 *)&adapter->rx_qstats[i] + off);
+
+ return total;
+}
+
+static u64 ibmveth_sum_tx_qstat(struct ibmveth_adapter *adapter, size_t off)
{
+ u64 total = 0;
int i;
+ for (i = 0; i < IBMVETH_MAX_QUEUES; i++)
+ total += *(u64 *)((u8 *)&adapter->tx_qstats[i] + off);
+
+ return total;
+}
+
+static u64 ibmveth_ethtool_adapter_stat(struct ibmveth_adapter *adapter,
+ int index)
+{
+ const struct ibmveth_stat *stat = &ibmveth_stats[index];
+
+ switch (stat->src) {
+ case IBMVETH_STAT_RX_QSUM:
+ return ibmveth_sum_rx_qstat(adapter, stat->off);
+ case IBMVETH_STAT_TX_QSUM:
+ return ibmveth_sum_tx_qstat(adapter, stat->off);
+ case IBMVETH_STAT_RX_NO_BUFFER:
+ /*
+ * PHYP's page counter is absolute for the page currently
+ * mapped, so a reopen or queue reuse restarts it near zero.
+ * ibmveth_update_rx_no_buffer() folds each decrease into the
+ * queue's retired carry; add both back to stay monotonic.
+ */
+ return ibmveth_sum_rx_qstat(adapter, stat->off) +
+ ibmveth_sum_rx_qstat(adapter,
+ IBMVETH_RXQ_OFF(no_buffer_retired));
+ case IBMVETH_STAT_ADAPTER:
+ break;
+ }
+
+ return IBMVETH_GET_STAT(adapter, stat->off);
+}
+
+static void ibmveth_get_strings(struct net_device *dev, u32 stringset, u8 *data)
+{
+ struct ibmveth_adapter *adapter = netdev_priv(dev);
+ u8 *p = data;
+ int i, j;
+
if (stringset != ETH_SS_STATS)
return;
- for (i = 0; i < ARRAY_SIZE(ibmveth_stats); i++, data += ETH_GSTRING_LEN)
- memcpy(data, ibmveth_stats[i].name, ETH_GSTRING_LEN);
+ for (i = 0; i < ARRAY_SIZE(ibmveth_stats); i++) {
+ memcpy(p, ibmveth_stats[i].name, ETH_GSTRING_LEN);
+ p += ETH_GSTRING_LEN;
+ }
+
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
+ for (j = 0; j < IBMVETH_NUM_RX_QSTATS; j++)
+ ethtool_sprintf(&p, ibmveth_rx_qstat_keys[j].fmt, i);
+
+ for (i = 0; i < dev->real_num_tx_queues; i++)
+ for (j = 0; j < IBMVETH_NUM_TX_QSTATS; j++)
+ ethtool_sprintf(&p, ibmveth_tx_qstat_keys[j].fmt, i);
}
static int ibmveth_get_sset_count(struct net_device *dev, int sset)
{
+ struct ibmveth_adapter *adapter = netdev_priv(dev);
+
switch (sset) {
case ETH_SS_STATS:
- return ARRAY_SIZE(ibmveth_stats);
+ return ARRAY_SIZE(ibmveth_stats) +
+ ibmveth_get_num_rx_queues(adapter) *
+ IBMVETH_NUM_RX_QSTATS +
+ dev->real_num_tx_queues * IBMVETH_NUM_TX_QSTATS;
default:
return -EOPNOTSUPP;
}
@@ -2302,11 +2520,27 @@ static int ibmveth_get_sset_count(struct net_device *dev, int sset)
static void ibmveth_get_ethtool_stats(struct net_device *dev,
struct ethtool_stats *stats, u64 *data)
{
- int i;
struct ibmveth_adapter *adapter = netdev_priv(dev);
+ int i, j, k;
for (i = 0; i < ARRAY_SIZE(ibmveth_stats); i++)
- data[i] = IBMVETH_GET_STAT(adapter, ibmveth_stats[i].offset);
+ data[i] = ibmveth_ethtool_adapter_stat(adapter, i);
+
+ for (j = 0; j < ibmveth_get_num_rx_queues(adapter); j++) {
+ const u8 *q = (const u8 *)&adapter->rx_qstats[j];
+
+ for (k = 0; k < IBMVETH_NUM_RX_QSTATS; k++)
+ data[i++] = *(const u64 *)
+ (q + ibmveth_rx_qstat_keys[k].off);
+ }
+
+ for (j = 0; j < dev->real_num_tx_queues; j++) {
+ const u8 *q = (const u8 *)&adapter->tx_qstats[j];
+
+ for (k = 0; k < IBMVETH_NUM_TX_QSTATS; k++)
+ data[i++] = *(const u64 *)
+ (q + ibmveth_tx_qstat_keys[k].off);
+ }
}
static void ibmveth_get_channels(struct net_device *netdev,
@@ -2418,8 +2652,10 @@ static int ibmveth_send(struct ibmveth_adapter *adapter,
}
static int ibmveth_is_packet_unsupported(struct sk_buff *skb,
- struct net_device *netdev)
+ struct ibmveth_adapter *adapter,
+ int queue_num)
{
+ struct net_device *netdev = adapter->netdev;
struct ethhdr *ether_header;
int ret = 0;
@@ -2427,7 +2663,7 @@ static int ibmveth_is_packet_unsupported(struct sk_buff *skb,
if (ether_addr_equal(ether_header->h_dest, netdev->dev_addr)) {
netdev_dbg(netdev, "veth doesn't support loopback packets, dropping packet.\n");
- netdev->stats.tx_dropped++;
+ adapter->tx_qstats[queue_num].dropped_packets++;
ret = -EOPNOTSUPP;
}
@@ -2445,11 +2681,11 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
/* Close / failed reopen can free LTBs while IFF_UP is still set. */
if (unlikely(!adapter->tx_ltb_ptr[queue_num])) {
- netdev->stats.tx_dropped++;
+ adapter->tx_qstats[queue_num].dropped_packets++;
goto out;
}
- if (ibmveth_is_packet_unsupported(skb, netdev))
+ if (ibmveth_is_packet_unsupported(skb, adapter, queue_num))
goto out;
/* veth can't checksum offload UDP */
if (skb->ip_summed == CHECKSUM_PARTIAL &&
@@ -2460,7 +2696,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
skb_checksum_help(skb)) {
netdev_err(netdev, "tx: failed to checksum packet\n");
- netdev->stats.tx_dropped++;
+ adapter->tx_qstats[queue_num].dropped_packets++;
goto out;
}
@@ -2472,6 +2708,8 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
desc_flags |= (IBMVETH_BUF_NO_CSUM | IBMVETH_BUF_CSUM_GOOD);
+ adapter->tx_qstats[queue_num].checksum_offload++;
+
/* Need to zero out the checksum */
buf[0] = 0;
buf[1] = 0;
@@ -2483,7 +2721,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
if (skb->ip_summed == CHECKSUM_PARTIAL && skb_is_gso(skb)) {
if (adapter->fw_large_send_support) {
mss = (unsigned long)skb_shinfo(skb)->gso_size;
- adapter->tx_large_packets++;
+ adapter->tx_qstats[queue_num].large_packets++;
} else if (!skb_is_gso_v6(skb)) {
/* Put -1 in the IP checksum to tell phyp it
* is a largesend packet. Put the mss in
@@ -2492,7 +2730,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
ip_hdr(skb)->check = 0xffff;
tcp_hdr(skb)->check =
cpu_to_be16(skb_shinfo(skb)->gso_size);
- adapter->tx_large_packets++;
+ adapter->tx_qstats[queue_num].large_packets++;
}
}
@@ -2500,7 +2738,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
if (unlikely(skb->len > adapter->tx_ltb_size)) {
netdev_err(adapter->netdev, "tx: packet size (%u) exceeds ltb (%u)\n",
skb->len, adapter->tx_ltb_size);
- netdev->stats.tx_dropped++;
+ adapter->tx_qstats[queue_num].dropped_packets++;
goto out;
}
memcpy(adapter->tx_ltb_ptr[queue_num], skb->data, skb_headlen(skb));
@@ -2517,7 +2755,7 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
if (unlikely(total_bytes != skb->len)) {
netdev_err(adapter->netdev, "tx: incorrect packet len copied into ltb (%u != %u)\n",
skb->len, total_bytes);
- netdev->stats.tx_dropped++;
+ adapter->tx_qstats[queue_num].dropped_packets++;
goto out;
}
desc.fields.flags_len = desc_flags | skb->len;
@@ -2526,11 +2764,11 @@ static netdev_tx_t ibmveth_start_xmit(struct sk_buff *skb,
dma_wmb();
if (ibmveth_send(adapter, desc.desc, mss)) {
- adapter->tx_send_failed++;
- netdev->stats.tx_dropped++;
+ adapter->tx_qstats[queue_num].send_failures++;
+ adapter->tx_qstats[queue_num].dropped_packets++;
} else {
- netdev->stats.tx_packets++;
- netdev->stats.tx_bytes += skb->len;
+ adapter->tx_qstats[queue_num].packets++;
+ adapter->tx_qstats[queue_num].bytes += skb->len;
}
out:
@@ -2659,7 +2897,7 @@ static void ibmveth_rx_csum_helper(struct sk_buff *skb,
static void ibmveth_poll_bump_invalid(struct ibmveth_adapter *adapter,
int queue_index)
{
- adapter->rx_invalid_buffer++;
+ adapter->rx_qstats[queue_index].invalid_buffers++;
}
static bool ibmveth_poll_stopping(struct net_device *netdev,
@@ -2798,7 +3036,7 @@ static int ibmveth_poll_deliver_frame(struct napi_struct *napi,
if ((length > netdev->mtu + ETH_HLEN) || lrg_pkt ||
iph_check == 0xffff) {
ibmveth_rx_mss_helper(skb, mss, lrg_pkt);
- adapter->rx_large_packets++;
+ adapter->rx_qstats[queue_index].large_packets++;
}
if (csum_good) {
@@ -2809,8 +3047,8 @@ static int ibmveth_poll_deliver_frame(struct napi_struct *napi,
skb_record_rx_queue(skb, queue_index);
napi_gro_receive(napi, skb);
- netdev->stats.rx_packets++;
- netdev->stats.rx_bytes += length;
+ adapter->rx_qstats[queue_index].packets++;
+ adapter->rx_qstats[queue_index].bytes += length;
return 1;
}
@@ -2837,6 +3075,8 @@ static int ibmveth_poll(struct napi_struct *napi, int budget)
return 0;
}
+ adapter->rx_qstats[queue_index].polls++;
+
restart_poll:
while (frames_processed < budget) {
if (ibmveth_poll_stopping(netdev, napi))
@@ -2925,6 +3165,8 @@ static irqreturn_t ibmveth_interrupt(int irq, void *dev_instance)
if (qindex < 0 || qindex >= ibmveth_get_num_rx_queues(adapter))
return IRQ_NONE;
+ adapter->rx_qstats[qindex].interrupts++;
+
ibmveth_schedule_rx_queue(adapter, qindex);
return IRQ_HANDLED;
}
@@ -3145,6 +3387,124 @@ static netdev_features_t ibmveth_features_check(struct sk_buff *skb,
return vlan_features_check(skb, features);
}
+/**
+ * ibmveth_get_stats64 - Return aggregated per-queue statistics
+ * @dev: network device
+ * @stats: rtnl link statistics storage
+ *
+ * Sums per-queue rx_qstats and tx_qstats into the rtnl counters.
+ * Walk the full allocated arrays (not the live queue count) so shrinking
+ * channels cannot make the totals go backwards.
+ * Callers use ndo_get_stats64(); avoid updating netdev->stats on the
+ * xmit/poll paths to keep per-queue counters off the hot cache line.
+ */
+static void ibmveth_get_stats64(struct net_device *dev,
+ struct rtnl_link_stats64 *stats)
+{
+ struct ibmveth_adapter *adapter = netdev_priv(dev);
+ int i;
+
+ for (i = 0; i < IBMVETH_MAX_RX_QUEUES; i++) {
+ stats->rx_packets += adapter->rx_qstats[i].packets;
+ stats->rx_bytes += adapter->rx_qstats[i].bytes;
+ }
+
+ for (i = 0; i < IBMVETH_MAX_QUEUES; i++) {
+ stats->tx_packets += adapter->tx_qstats[i].packets;
+ stats->tx_bytes += adapter->tx_qstats[i].bytes;
+ stats->tx_dropped += adapter->tx_qstats[i].dropped_packets;
+ }
+}
+
+static void ibmveth_get_queue_stats_rx(struct net_device *dev, int idx,
+ struct netdev_queue_stats_rx *stats)
+{
+ struct ibmveth_adapter *adapter = netdev_priv(dev);
+
+ stats->packets = adapter->rx_qstats[idx].packets;
+ stats->bytes = adapter->rx_qstats[idx].bytes;
+ /*
+ * All three are frames that entered the device and never left it,
+ * which is what rx-hw-drops is specified to cover: no_buffer_drops
+ * is PHYP dropping for lack of buffer space on the page mapped now,
+ * no_buffer_retired the same for pages this queue has already
+ * released, and invalid_buffers is a processing error.
+ */
+ stats->hw_drops = adapter->rx_qstats[idx].no_buffer_drops +
+ adapter->rx_qstats[idx].no_buffer_retired +
+ adapter->rx_qstats[idx].invalid_buffers;
+ stats->alloc_fail = adapter->rx_qstats[idx].replenish_no_mem;
+}
+
+static void ibmveth_get_queue_stats_tx(struct net_device *dev, int idx,
+ struct netdev_queue_stats_tx *stats)
+{
+ struct ibmveth_adapter *adapter = netdev_priv(dev);
+
+ stats->packets = adapter->tx_qstats[idx].packets;
+ stats->bytes = adapter->tx_qstats[idx].bytes;
+ stats->hw_drops = adapter->tx_qstats[idx].dropped_packets;
+}
+
+/**
+ * ibmveth_get_base_stats - account for traffic not on a live queue
+ * @dev: network device
+ * @rx: RX base statistics storage
+ * @tx: TX base statistics storage
+ *
+ * get_queue_stats_{rx,tx}() only report queues the core still iterates,
+ * i.e. below real_num_{rx,tx}_queues, while ibmveth_get_stats64() walks
+ * the full arrays so device totals stay monotonic across a shrink.
+ * Report the retired-queue remainder here, otherwise qstats and
+ * rtnl_link_stats64 disagree by a delta that grows with every shrink.
+ * Zeroing would not be neutral: per netdev_stat_ops it asserts the
+ * per-queue sum is already exact.
+ *
+ * Bound the live side with real_num_*_queues rather than the adapter's
+ * own count, so the split lines up with the core's iteration exactly.
+ *
+ * Every field the per-queue callbacks fill must also be initialised
+ * here: netdev_nl_stats_add() starts the sum at NETDEV_STAT_NOT_SET and
+ * only accumulates while both sides are set, so a field left unset here
+ * is dropped from the device total even though the queues report it.
+ */
+static void ibmveth_get_base_stats(struct net_device *dev,
+ struct netdev_queue_stats_rx *rx,
+ struct netdev_queue_stats_tx *tx)
+{
+ struct ibmveth_adapter *adapter = netdev_priv(dev);
+ unsigned int i;
+
+ rx->packets = 0;
+ rx->bytes = 0;
+ rx->alloc_fail = 0;
+ rx->hw_drops = 0;
+ tx->packets = 0;
+ tx->bytes = 0;
+ tx->hw_drops = 0;
+
+ for (i = dev->real_num_rx_queues; i < IBMVETH_MAX_RX_QUEUES; i++) {
+ rx->packets += adapter->rx_qstats[i].packets;
+ rx->bytes += adapter->rx_qstats[i].bytes;
+ rx->hw_drops += adapter->rx_qstats[i].no_buffer_drops +
+ adapter->rx_qstats[i].no_buffer_retired +
+ adapter->rx_qstats[i].invalid_buffers;
+ rx->alloc_fail += adapter->rx_qstats[i].replenish_no_mem;
+ }
+
+ for (i = dev->real_num_tx_queues; i < IBMVETH_MAX_QUEUES; i++) {
+ tx->packets += adapter->tx_qstats[i].packets;
+ tx->bytes += adapter->tx_qstats[i].bytes;
+ tx->hw_drops += adapter->tx_qstats[i].dropped_packets;
+ }
+}
+
+static const struct netdev_stat_ops ibmveth_stat_ops = {
+ .get_queue_stats_rx = ibmveth_get_queue_stats_rx,
+ .get_queue_stats_tx = ibmveth_get_queue_stats_tx,
+ .get_base_stats = ibmveth_get_base_stats,
+};
+
static const struct net_device_ops ibmveth_netdev_ops = {
.ndo_open = ibmveth_open,
.ndo_stop = ibmveth_close,
@@ -3157,6 +3517,7 @@ static const struct net_device_ops ibmveth_netdev_ops = {
.ndo_validate_addr = eth_validate_addr,
.ndo_set_mac_address = ibmveth_set_mac_addr,
.ndo_features_check = ibmveth_features_check,
+ .ndo_get_stats64 = ibmveth_get_stats64,
#ifdef CONFIG_NET_POLL_CONTROLLER
.ndo_poll_controller = ibmveth_poll_controller,
#endif
@@ -3198,6 +3559,23 @@ static void ibmveth_put_pool_kobjs(struct ibmveth_adapter *adapter,
wait_for_completion(&adapter->rx_buff_pool[0][i].released);
}
+static void ibmveth_probe_cleanup(struct ibmveth_adapter *adapter,
+ int pools_ready)
+{
+ struct net_device *netdev = adapter->netdev;
+
+ cancel_work_sync(&adapter->work);
+ ibmveth_put_pool_kobjs(adapter, pools_ready);
+
+ ibmveth_free_tx_qstats(adapter);
+ ibmveth_free_rx_qstats(adapter);
+ /* Probe failure never reaches ibmveth_remove(); clear before free so
+ * CMO get_desired_dma() cannot see a freed netdev on rebind.
+ */
+ dev_set_drvdata(&adapter->vdev->dev, NULL);
+ free_netdev(netdev);
+}
+
static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
{
int rc, i, mac_len, pools_ready = 0;
@@ -3263,9 +3641,16 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
netif_napi_add_weight(netdev, &adapter->napi[i],
ibmveth_poll, 16);
+ if (ibmveth_alloc_rx_qstats(adapter) ||
+ ibmveth_alloc_tx_qstats(adapter)) {
+ ibmveth_probe_cleanup(adapter, 0);
+ return -ENOMEM;
+ }
+
netdev->irq = dev->irq;
netdev->netdev_ops = &ibmveth_netdev_ops;
netdev->ethtool_ops = &netdev_ethtool_ops;
+ netdev->stat_ops = &ibmveth_stat_ops;
SET_NETDEV_DEV(netdev, &dev->dev);
netdev->hw_features = NETIF_F_SG;
if (vio_get_attribute(dev, "ibm,illan-options", NULL) != NULL) {
@@ -3350,9 +3735,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
/* init_and_add takes a ref even on failure */
kobject_put(kobj);
wait_for_completion(&pool->released);
- ibmveth_put_pool_kobjs(adapter, pools_ready);
- dev_set_drvdata(&dev->dev, NULL);
- free_netdev(netdev);
+ ibmveth_probe_cleanup(adapter, pools_ready);
return rc;
}
@@ -3365,9 +3748,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
if (rc) {
netdev_dbg(netdev, "failed to set number of tx queues rc=%d\n",
rc);
- ibmveth_put_pool_kobjs(adapter, pools_ready);
- dev_set_drvdata(&dev->dev, NULL);
- free_netdev(netdev);
+ ibmveth_probe_cleanup(adapter, pools_ready);
return rc;
}
@@ -3382,9 +3763,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
if (rc) {
netdev_dbg(netdev, "failed to set number of rx queues rc=%d\n",
rc);
- ibmveth_put_pool_kobjs(adapter, pools_ready);
- dev_set_drvdata(&dev->dev, NULL);
- free_netdev(netdev);
+ ibmveth_probe_cleanup(adapter, pools_ready);
return rc;
}
@@ -3401,9 +3780,7 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
if (rc) {
netdev_dbg(netdev, "failed to register netdev rc=%d\n", rc);
- ibmveth_put_pool_kobjs(adapter, pools_ready);
- dev_set_drvdata(&dev->dev, NULL);
- free_netdev(netdev);
+ ibmveth_probe_cleanup(adapter, pools_ready);
return rc;
}
@@ -3428,6 +3805,9 @@ static void ibmveth_remove(struct vio_dev *dev)
unregister_netdev(netdev);
cancel_work_sync(&adapter->work);
+ ibmveth_free_tx_qstats(adapter);
+ ibmveth_free_rx_qstats(adapter);
+
free_netdev(netdev);
dev_set_drvdata(&dev->dev, NULL);
}
diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h
index 84e703692aae..b35da8bdce5f 100644
--- a/drivers/net/ethernet/ibm/ibmveth.h
+++ b/drivers/net/ethernet/ibm/ibmveth.h
@@ -279,6 +279,43 @@ static int pool_active[] = { 1, 1, 0, 0, 1};
#define IBM_VETH_INVALID_MAP ((u16)0xffff)
+/*
+ * Per-queue RX counters. No field has two concurrent writers:
+ * interrupts is written only from this queue's IRQ handler; polls,
+ * packets, bytes, large_packets and invalid_buffers only from its NAPI
+ * poll; replenish_* only under its replenish_lock; and no_buffer_drops
+ * and no_buffer_retired under that lock or from a teardown path already
+ * quiesced by napi_disable()/synchronize_irq(). Plain u64 is therefore
+ * sufficient and no atomic or u64_stats_sync is needed: the driver is
+ * PPC64-only, so 64-bit loads and stores do not tear.
+ */
+struct ibmveth_rx_queue_stats {
+ u64 packets;
+ u64 bytes;
+ u64 interrupts;
+ u64 polls;
+ u64 large_packets;
+ u64 invalid_buffers;
+ /* PHYP's per-page absolute drop count for the live page. */
+ u64 no_buffer_drops;
+ /* Absolutes from pages this queue has already retired. */
+ u64 no_buffer_retired;
+ u64 replenish_task_cycles;
+ u64 replenish_no_mem;
+ u64 replenish_add_buff_failure;
+ u64 replenish_add_buff_success;
+} ____cacheline_aligned_in_smp;
+
+/* Per-queue TX counters; serialized by the stack's per-queue TX lock. */
+struct ibmveth_tx_queue_stats {
+ u64 packets;
+ u64 bytes;
+ u64 large_packets;
+ u64 dropped_packets;
+ u64 send_failures;
+ u64 checksum_offload;
+} ____cacheline_aligned_in_smp;
+
struct ibmveth_buff_pool {
u32 size;
u32 index;
@@ -338,17 +375,17 @@ struct ibmveth_adapter {
u64 fw_ipv6_csum_support;
u64 fw_ipv4_csum_support;
u64 fw_large_send_support;
- /* adapter specific stats */
- u64 replenish_task_cycles;
- u64 replenish_no_mem;
- u64 replenish_add_buff_failure;
- u64 replenish_add_buff_success;
- u64 rx_invalid_buffer;
- u64 rx_no_buffer;
+ /*
+ * Every other ethtool -S counter lives in rx_qstats/tx_qstats and is
+ * summed on read. tx_map_failed predates multi-queue, has never been
+ * updated by any code path, and is kept only so the key keeps
+ * reporting the zero userspace already sees.
+ */
u64 tx_map_failed;
- u64 tx_send_failed;
- u64 tx_large_packets;
- u64 rx_large_packets;
+
+ struct ibmveth_rx_queue_stats *rx_qstats;
+ struct ibmveth_tx_queue_stats *tx_qstats;
+
/* Ethtool settings */
u8 duplex;
u32 speed;
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 11/15] ibmveth: Add per-queue RX and TX statistics collection
2026-09-25 18:38 ` [PATCH net-next v7 11/15] ibmveth: Add per-queue RX and TX statistics collection Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 4 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] Three of the new per-queue ethtool -S keys duplicate
statistics that already have a standard uAPI, which the networking-…
- [Medium] The new no_buffer_retired carry in
ibmveth_update_rx_no_buffer() only notices a page change when the…
- [Medium] ibmveth_get_queue_stats_tx() (line 3446) and
ibmveth_get_base_stats() (line 3498) report…
- [Medium] net: stats: Torn reads of related drop counters cause massive
spikes
--
ibmveth: Add per-queue RX and TX statistics collection
This patch moves the ibmveth RX and TX counters into per-queue structs
that are allocated at probe. Packets, bytes and drops are reported
through netdev_stat_ops and ndo_get_stats64(), and new per-queue
ethtool -S keys are added. A per-queue no_buffer_retired carry is also
added so the PHYP no-buffer drop count stays monotonic across reopen and
queue reuse.
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index 3f31793645a31..2575a7c11afba 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
[ ... ]
> @@ -75,32 +76,101 @@ module_param(old_large_send, bool, 0444);
[ ... ]
> +/*
> + * Only counters with no home in the standard interfaces belong here.
> + * packets, bytes and drops are reported through netdev_stat_ops.
> + */
> +static const struct ibmveth_qstat ibmveth_rx_qstat_keys[] = {
> + { "rx%d_interrupts", IBMVETH_RXQ_OFF(interrupts) },
> + { "rx%d_polls", IBMVETH_RXQ_OFF(polls) },
> + { "rx%d_large_packets", IBMVETH_RXQ_OFF(large_packets) },
> + { "rx%d_invalid_buffers", IBMVETH_RXQ_OFF(invalid_buffers) },
> + { "rx%d_no_buffer_drops", IBMVETH_RXQ_OFF(no_buffer_drops) },
> };
>
> +static const struct ibmveth_qstat ibmveth_tx_qstat_keys[] = {
> + { "tx%d_large_packets", IBMVETH_TXQ_OFF(large_packets) },
> + { "tx%d_send_failures", IBMVETH_TXQ_OFF(send_failures) },
> + { "tx%d_checksum_offload", IBMVETH_TXQ_OFF(checksum_offload) },
> +};
[Severity: Medium]
Do three of these new per-queue keys duplicate counters that already
have a standard qstats uAPI?
ibmveth_start_xmit() only bumps tx%d_checksum_offload for
CHECKSUM_PARTIAL skbs where the device is asked to fill in the checksum
(IBMVETH_BUF_NO_CSUM | IBMVETH_BUF_CSUM_GOOD). That looks like
tx-needs-csum (netdev_queue_stats_tx.needs_csum).
tx%d_large_packets counts GSO skbs handed to PHYP for large send. That
looks like tx-hw-gso-packets (hw_gso_packets).
rx%d_no_buffer_drops counts frames PHYP dropped because no receive
buffers were posted. That matches rx-hw-drop-overruns ("transient lack
of resources, such as buffer space").
ibmveth_get_queue_stats_rx(), ibmveth_get_queue_stats_tx() and
ibmveth_get_base_stats() only fill packets, bytes, hw_drops and
alloc_fail. So needs_csum, hw_gso_packets and hw_drop_overruns stay
unset, and the same data goes out as private strings instead.
Could these be reported through netdev_stat_ops instead? The comment
above ibmveth_stats[] says ethtool -S key names are ABI. This is still
the case at the end of the series.
The commit message also says:
ethtool -S keeps only the driver-specific keys
that have no standard equivalent: interrupts, polls, large_packets,
invalid_buffers and no_buffer_drops per RX queue; large_packets,
send_failures and checksum_offload per TX queue.
Is that accurate, given the mappings above?
[ ... ]
> @@ -1028,7 +1155,18 @@ static void ibmveth_update_rx_no_buffer(struct ibmveth_adapter *adapter,
> p = adapter->buffer_list_addr[queue_index] + 4096 - 8;
> drops = be64_to_cpup(p);
>
> - adapter->rx_no_buffer = drops;
> + /*
> + * PHYP's buffer-list page counter is absolute for that page. A new
> + * page (reopen / queue reuse after -L) starts near zero; fold the
> + * previous absolute into this queue's retired carry so sums stay
> + * monotonic. Both fields belong to the queue being updated, so this
> + * stays single-writer under the queue's replenish_lock.
> + */
> + qstats = &adapter->rx_qstats[queue_index];
> +
> + if (drops < qstats->no_buffer_drops)
> + qstats->no_buffer_retired += qstats->no_buffer_drops;
> + qstats->no_buffer_drops = drops;
> }
[Severity: Medium]
Can a concurrent reader count the same drops twice here?
The rx_no_buffer case sums no_buffer_drops over every rx_qstats[] slot
in one pass. It then sums no_buffer_retired over every slot in a
second pass. ibmveth_update_rx_no_buffer() does the carry as two
separate stores: it adds the old absolute X to no_buffer_retired, and
then overwrites no_buffer_drops with the new, smaller reading.
Suppose that carry runs for a queue after the first pass has read that
queue's no_buffer_drops, but before the second pass reads its
no_buffer_retired. Then X lands in both sums, and rx_no_buffer jumps
by a whole page's worth of drops. The next ethtool -S read falls back
by the same amount. A tool computing a drop rate therefore sees a
burst followed by a negative rate.
Both passes walk all IBMVETH_MAX_RX_QUEUES slots. The window for each
queue is therefore a full walk of the array, not just two adjacent
loads.
The reader takes no replenish_lock, and no seqcount pairs the two
fields. The plain u64 argument in the commit message covers tearing of
a single field on PPC64. It does not cover consistency between two
fields that together form one counter.
ibmveth_get_queue_stats_rx() and ibmveth_get_base_stats() add the same
pair in one expression. That narrows the window without closing it,
and depending on load order the error can go either way.
Is the patch relying on rtnl to keep every carry away from these
readers? If so, could that be stated in the struct
ibmveth_rx_queue_stats comment? Replenish from ibmveth_poll() runs
without rtnl. If not, would it be simpler to read the pair together
under the queue's replenish_lock? Another option is to have the writer
maintain a single monotonic total that the readers load once.
[Severity: Medium]
Can the old page's drops be lost when the first reading on the new page
is not smaller than the old final value?
Each open gets a fresh get_zeroed_page() buffer list from
ibmveth_alloc_rx_queues(), so PHYP's counter in the last 8 bytes of the
new page starts at zero. On close, the harvest stores the old page's
final value X in no_buffer_drops but does not move it into
no_buffer_retired.
On the next open:
ibmveth_open()
ibmveth_alloc_rx_queues() /* new zeroed page */
ibmveth_register_rx_queues() /* PHYP owns it, no buffers posted */
ibmveth_replenish_task()
ibmveth_update_rx_no_buffer() /* first read, Y */
Frames that arrive between registration and the first replenish count as
no-buffer drops on the new page. If Y >= X, the decrease test does not
fire and X is overwritten. rx_no_buffer and rx-hw-drops then report
R + Y instead of R + X + Y. When X is small, this seems likely on every
reopen.
A queue that ethtool -L removes and then adds back hits the same
problem. ibmveth_scale_down_rx_queues() (from "ibmveth: Implement
incremental MQ RX queue resize" in this series) harvests the queue the
same way, and the re-added queue gets a new zeroed page.
There also seems to be a second window. ibmveth_close() harvests before
the page is released:
spin_lock_irqsave(&rxq->replenish_lock, flags);
ibmveth_update_rx_no_buffer(adapter, i);
spin_unlock_irqrestore(&rxq->replenish_lock, flags);
}
ibmveth_free_all_queues(adapter);
Scale-down likewise harvests before h_free_logical_lan_queue(). PHYP
still owns the page at that point, so any drops it adds after the
harvest read are never seen.
The driver allocates and retires these pages itself. Could the carry be
applied explicitly at retirement, after deregistration (something like
retired += drops; drops = 0), instead of being inferred from a decrease?
This code is unchanged at the end of the series.
[ ... ]
> @@ -3145,6 +3387,124 @@ static netdev_features_t ibmveth_features_check(struct sk_buff *skb,
[ ... ]
> +static void ibmveth_get_queue_stats_tx(struct net_device *dev, int idx,
> + struct netdev_queue_stats_tx *stats)
> +{
> + struct ibmveth_adapter *adapter = netdev_priv(dev);
> +
> + stats->packets = adapter->tx_qstats[idx].packets;
> + stats->bytes = adapter->tx_qstats[idx].bytes;
> + stats->hw_drops = adapter->tx_qstats[idx].dropped_packets;
> +}
[Severity: Medium]
Is dropped_packets the right source for tx-hw-drops?
netdev.yaml defines tx-hw-drops as packets that "arrived at the device
but never left it". Most dropped_packets increments come from host-side
paths that never reach ibmveth_send():
- missing TX LTB in ibmveth_start_xmit()
- destination MAC equal to dev_addr in ibmveth_is_packet_unsupported()
- skb_checksum_help() failure
- skb->len > tx_ltb_size
- short copy into the LTB
These look like they belong in rtnl tx_dropped, which
ibmveth_get_stats64() already fills from the same counter.
ibmveth_get_base_stats() uses the same mapping for retired queues:
tx->hw_drops += adapter->tx_qstats[i].dropped_packets;
Separately, ibmveth_send() treats H_DROPPED as success:
if (ret != H_SUCCESS && ret != H_DROPPED) {
...
return 1;
}
return 0;
As a result, ibmveth_start_xmit() adds frames that PHYP reports as
dropped to packets and bytes. If H_DROPPED means the hypervisor
discarded the frame, would the new tx-packets and tx-hw-drops pair
misattribute it?
The H_DROPPED handling is older than this patch. Exporting it through
these qstat fields is new.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 12/15] ibmveth: Report MQ-aware RX counts in ethtool get_channels
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (10 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 11/15] ibmveth: Add per-queue RX and TX statistics collection Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 13/15] ibmveth: Expose per-queue buffer pool details via debugfs Mingming Cao
` (3 subsequent siblings)
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
Report MQ-aware channel limits and counts from get_channels():
- max_rx is IBMVETH_MAX_RX_QUEUES when MQ firmware is enabled, else 1
- rx_count tracks the driver's published num_rx_queues
Previously get_channels() mirrored netdev->real_num_rx_queues for both
fields, which under-reports max_rx on MQ firmware and can disagree with
the driver's private RX count before the next open.
This patch only reports. set_channels() still only adjusts tx_count, so
an RX channel request returns -EOPNOTSUPP if rx_count changes, while
read-modify-write TX requests pass through with the live rx_count.
Patch 14 implements live RX resizing via ibmveth_resize_rx_channels(),
and patch 15 completes the down-path publish/rollback and caps max_rx
once mq_fallback latches.
Keep this out of the stats patch: channel reporting is ethtool -l ABI,
independent of per-queue counters, so it can be reviewed and blamed on
its own.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- reject rx_count change with -EOPNOTSUPP in
set_channels() until patch 14 implements resize
- noted: TX-only read-modify-write still passes;
max_rx stays here (reject, not reorder);
get_channels uses the adapter count,
get_base_stats uses real_num; RX max is
MAX_RX_QUEUES, not the TX CPU cap; P14
wires live rx_count resize; P15 is the
tip ABI (live rx_count, max_rx cap on
mq_fallback, max_tx at least the live tx_count)
Changes in v5:
- get_channels rx_count via get_num_rx_queues() (READ_ONCE consistency)
- New peel: get_channels from v4 stats -> tip P12 (14->15)
- get_channels: max_rx = MAX_RX_QUEUES when MQ else 1; rx_count =
get_num_rx_queues() (was mirroring real_num_rx_queues for both)
drivers/net/ethernet/ibm/ibmveth.c | 17 +++++++++++++++--
1 file changed, 15 insertions(+), 2 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 2575a7c11afb..cd2e5e4a81fe 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -2546,11 +2546,16 @@ static void ibmveth_get_ethtool_stats(struct net_device *dev,
static void ibmveth_get_channels(struct net_device *netdev,
struct ethtool_channels *channels)
{
+ struct ibmveth_adapter *adapter = netdev_priv(netdev);
+
channels->max_tx = ibmveth_real_max_tx_queues();
channels->tx_count = netdev->real_num_tx_queues;
- channels->max_rx = netdev->real_num_rx_queues;
- channels->rx_count = netdev->real_num_rx_queues;
+ if (adapter->multi_queue)
+ channels->max_rx = IBMVETH_MAX_RX_QUEUES;
+ else
+ channels->max_rx = 1;
+ channels->rx_count = ibmveth_get_num_rx_queues(adapter);
}
static int ibmveth_set_channels(struct net_device *netdev,
@@ -2561,6 +2566,14 @@ static int ibmveth_set_channels(struct net_device *netdev,
goal = channels->tx_count;
int rc, i;
+ /*
+ * RX channel resize is implemented in a later patch; reject any
+ * request that changes rx_count. Read-modify-write TX adjustments
+ * submit the current rx_count and proceed.
+ */
+ if (channels->rx_count != ibmveth_get_num_rx_queues(adapter))
+ return -EOPNOTSUPP;
+
/* If ndo_open has not been called yet then don't allocate, just set
* desired netdev_queue's and return
*/
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 12/15] ibmveth: Report MQ-aware RX counts in ethtool get_channels
2026-09-25 18:38 ` [PATCH net-next v7 12/15] ibmveth: Report MQ-aware RX counts in ethtool get_channels Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found more severe
issue(s) in other patches of this series. This mail is only an FYI, in case
you want to address the nit picks below when you respin.
Low priority / nit picks for your consideration:
- [Low] The commit message says "This patch only reports", and the
subject covers only get_channels().
- [Low] This patch switches ibmveth_get_channels() to report rx_count
from adapter->num_rx_queues (line 2558) instead of…
- [Low] At this commit, ibmveth_get_channels() computes max_rx only from
adapter->multi_queue (line 2554-2557) and ignores adapter->mq_fallback.
Pre-existing issues:
- [Low] In ibmveth_set_channels() at this commit, when
ibmveth_allocate_tx_ltb() fails in the grow loop (line 2593), the code…
--
ibmveth: Report MQ-aware RX counts in ethtool get_channels
This commit changes ibmveth_get_channels() to report max_rx as
IBMVETH_MAX_RX_QUEUES when multi_queue is set (otherwise 1). It also
reports rx_count from the driver's published num_rx_queues. A check is
added to ibmveth_set_channels() that rejects any request changing rx_count
with -EOPNOTSUPP.
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index 2575a7c11afba..cd2e5e4a81fea 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
> @@ -2546,11 +2546,16 @@ static void ibmveth_get_ethtool_stats(struct net_device *dev,
> static void ibmveth_get_channels(struct net_device *netdev,
> struct ethtool_channels *channels)
> {
> + struct ibmveth_adapter *adapter = netdev_priv(netdev);
> +
> channels->max_tx = ibmveth_real_max_tx_queues();
> channels->tx_count = netdev->real_num_tx_queues;
>
> - channels->max_rx = netdev->real_num_rx_queues;
> - channels->rx_count = netdev->real_num_rx_queues;
> + if (adapter->multi_queue)
> + channels->max_rx = IBMVETH_MAX_RX_QUEUES;
> + else
> + channels->max_rx = 1;
[Severity: Low]
Should max_rx also check adapter->mq_fallback here?
Firmware can reject MQ in two places. One is a subordinate queue
registration in ibmveth_register_all_queues() returning -EOPNOTSUPP. The
other is ibmveth_replenish_task() seeing H_FUNCTION. In both cases only
mq_fallback is latched:
if (rc == -EOPNOTSUPP)
adapter->mq_fallback = true;
multi_queue and num_rx_queues are only updated when the next
ibmveth_open() calls ibmveth_apply_mq_fallback(). Until then, ethtool -l
keeps reporting max_rx = IBMVETH_MAX_RX_QUEUES and the pre-fallback
rx_count.
The commit message says a later patch in the series caps max_rx once
mq_fallback latches. The final version of ibmveth_get_channels() does
check !adapter->mq_fallback, so this is resolved by the end of the series.
> + channels->rx_count = ibmveth_get_num_rx_queues(adapter);
[Severity: Low]
Can this rx_count disagree with netdev->real_num_rx_queues after a failed
open?
ibmveth_open() applies the fallback before it allocates anything:
ibmveth_open()
ibmveth_apply_mq_fallback(adapter);
adapter->multi_queue = false;
ibmveth_publish_num_rx_queues(adapter, 1);
/* real_num_rx_queues is set later in open after resources exist. */
...
rc = ibmveth_alloc_filter_list(adapter);
if (rc)
goto out;
...
rc = ibmveth_register_rx_queues(adapter, mac_address);
if (rc)
goto out_free_buffer_pools;
rc = netif_set_real_num_rx_queues(netdev,
ibmveth_get_num_rx_queues(adapter));
If any allocation or registration step fails, open returns with
num_rx_queues = 1 and multi_queue = false. real_num_rx_queues keeps its
earlier MQ value (for example 8 from probe).
In that state ethtool -l reports rx_count = 1 and max_rx = 1. The sysfs
rx-* queues and the per-queue qstats iteration still show 8 RX queues.
This lasts until the next successful open.
The comment in ibmveth_probe() says the counts are synced "so down-state
readers agree with adapter->num_rx_queues / ethtool -l". This path does
not seem to keep that promise.
The qstats device totals stay exact, because ibmveth_get_base_stats() and
ibmveth_get_queue_stats_rx() both split on real_num_rx_queues. The root
cause is the error handling in ibmveth_open() from earlier in the series.
This patch is what makes the mismatch visible through ethtool -l.
> }
>
> static int ibmveth_set_channels(struct net_device *netdev,
> @@ -2561,6 +2566,14 @@ static int ibmveth_set_channels(struct net_device *netdev,
> goal = channels->tx_count;
> int rc, i;
>
> + /*
> + * RX channel resize is implemented in a later patch; reject any
> + * request that changes rx_count. Read-modify-write TX adjustments
> + * submit the current rx_count and proceed.
> + */
> + if (channels->rx_count != ibmveth_get_num_rx_queues(adapter))
> + return -EOPNOTSUPP;
[Severity: Low]
The commit message says:
This patch only reports. set_channels() still only adjusts tx_count, so
an RX channel request returns -EOPNOTSUPP if rx_count changes, while
read-modify-write TX requests pass through with the live rx_count.
Is "This patch only reports" accurate? This check is new in this commit.
Before it, ibmveth_set_channels() did not look at rx_count at all.
On MQ firmware the old get_channels() reported max_rx as
real_num_rx_queues (for example 8). A request like "ethtool -L eth0 rx 4
tx 4" passed the core range check. The driver ignored rx, applied tx = 4
and returned 0.
With this patch the whole request fails with -EOPNOTSUPP, so the TX part
is not applied either.
Only the v7 changelog below the --- line says the check is new, and that
text is dropped when the patch is applied. A later patch replaces this
check with ibmveth_resize_rx_channels(). Still, bisect and blame will show
this commit's message.
Could the subject and body say that this patch also changes how
ibmveth_set_channels() handles rx_count?
> +
> /* If ndo_open has not been called yet then don't allocate, just set
> * desired netdev_queue's and return
> */
[Severity: Low]
This is a pre-existing issue and was not introduced by this patch. Further
down in ibmveth_set_channels(), is the error from a failed TX LTB
allocation lost?
rc = ibmveth_allocate_tx_ltb(adapter, i);
if (!rc)
continue;
...
goal = old;
old = i;
break;
}
rc = netif_set_real_num_tx_queues(netdev, goal);
At that point goal is back to the current count. So
netif_set_real_num_tx_queues() returns 0 and overwrites rc.
The LTBs allocated so far are freed. ethtool -L tx N then reports success
even though tx_count did not change.
These lines come from commit 10c2aba89cc053 (2022). Later in this series,
the final ibmveth_set_channels() returns rc ? rc : -ENOMEM when
real_num_tx_queues != want_tx, which fixes this.
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 13/15] ibmveth: Expose per-queue buffer pool details via debugfs
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (11 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 12/15] ibmveth: Report MQ-aware RX counts in ethtool get_channels Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-25 18:38 ` [PATCH net-next v7 14/15] ibmveth: Implement incremental MQ RX queue resize Mingming Cao
` (2 subsequent siblings)
15 siblings, 0 replies; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
With multi-queue RX each queue owns its own set of five buffer pools,
so a 16-queue adapter has 80 of them. Nothing reports their runtime
state: sysfs exposes queue 0 only, and only as configuration, and no
ethtool key is per-pool. When RX drops under load, rx%d_no_buffer_drops
names the queue but not which of its pools ran dry, nor how close the
others are.
Add a read-only buffer_pools debugfs file, one row per RX queue and
buffer pool:
/sys/kernel/debug/ibmveth/<dev_name>/buffer_pools
(e.g. /sys/kernel/debug/ibmveth/30000002/buffer_pools)
Queue Pool Count BuffSize Active Available
Active is live allocation (skbuff && free_map), not the sysfs
poolN/active configuration flag.
The root is driver-owned so each adapter directory can use its stable
vio name rather than the mutable netdev->name. It is created in
module_init() and unwound if vio_register_driver() fails.
A multi-line table does not belong in sysfs, so the historical queue-0
ABI is left alone:
.../poolN/{active,num,size}
Those stay one-value configuration for queue-0 pool classes. Open
copies that geometry to queues 1..N. This series does not add
per-queue pool sysfs dirs.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- no P13 code change
Changes in v6:
- create the debugfs root in module_init() instead of lazily on
first probe, which raced concurrent probes and could orphan the
directory on ERR_PTR(-EEXIST)
- rename the pool buffer-count column from Size to Count, so the
debugfs table stops reusing the word sysfs poolN/size spells as a
byte length on the same pool object
- widen the down banner: geometry above queue 0 is only populated
once open copies the queue-0 template
- scope the dump RTNL comment to geometry/pool->active; available is
atomic_read
Changes in v5:
- debugfs buffer_pools_show walks get_num_rx_queues()
- Series renumber: mailed v4 11/14 debugfs -> tip P13 (14->15)
- Path uses stable vio dev_name under a driver-owned root (not netdev
name - avoids rename/collide)
- rtnl_lock around dump (writers are under RTNL)
- Show Active/Available as 0 when pool !live (debugfs view; free-path
available clear already in the buffer-submit patch)
Changes in v4:
- Move the all-queue buffer_pools diagnostic from sysfs to debugfs;
subject updated to match.
- Keep historical queue-0 poolN/{active,num,size} sysfs as one-value
config (template for MQ); do not add per-queue pool sysfs dirs.
drivers/net/ethernet/ibm/ibmveth.c | 80 +++++++++++++++++++++++++++++-
drivers/net/ethernet/ibm/ibmveth.h | 2 +
2 files changed, 81 insertions(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index cd2e5e4a81fe..0ac0359bb71e 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -31,6 +31,7 @@
#include <linux/ipv6.h>
#include <linux/slab.h>
#include <linux/spinlock.h>
+#include <linux/debugfs.h>
#include <asm/hvcall.h>
#include <linux/atomic.h>
#include <asm/vio.h>
@@ -3536,6 +3537,67 @@ static const struct net_device_ops ibmveth_netdev_ops = {
#endif
};
+static int ibmveth_buffer_pools_show(struct seq_file *m, void *v)
+{
+ struct ibmveth_adapter *adapter = m->private;
+ int i, j;
+
+ /*
+ * size / buff_size / pool->active are written under RTNL
+ * (veth_pool_store, open template copy). Take the same lock so
+ * those columns are not a torn snapshot. available is updated
+ * from NAPI/softirq; only atomic_read() keeps it from tearing.
+ * Not required for memory safety; embedded arrays only.
+ */
+ rtnl_lock();
+
+ seq_puts(m, "Queue Pool Count BuffSize Active Available\n");
+ seq_puts(m, "----- ---- ----- -------- ------ ---------\n");
+ if (!adapter->opened) {
+ seq_puts(m, "# down: Active/Available 0 unless allocated\n");
+ seq_puts(m, "# down: geometry above queue 0 set at open\n");
+ }
+
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
+ for (j = 0; j < IBMVETH_NUM_BUFF_POOLS; j++) {
+ struct ibmveth_buff_pool *pool =
+ &adapter->rx_buff_pool[i][j];
+ bool live = pool->skbuff && pool->free_map;
+ int active = live ? pool->active : 0;
+ int available = live ? atomic_read(&pool->available)
+ : 0;
+
+ seq_printf(m, "%5d %4d %5u %8u %6d %9d\n",
+ i, j, pool->size, pool->buff_size,
+ active, available);
+ }
+ }
+
+ rtnl_unlock();
+ return 0;
+}
+DEFINE_SHOW_ATTRIBUTE(ibmveth_buffer_pools);
+
+/* Driver-owned root so per-adapter dirs use a stable vio name, not the
+ * mutable netdev->name (avoids stale names / eth0 collisions after rename).
+ */
+static struct dentry *ibmveth_dbg_root;
+
+static void ibmveth_debugfs_init(struct ibmveth_adapter *adapter)
+{
+ adapter->debugfs_dir =
+ debugfs_create_dir(dev_name(&adapter->vdev->dev),
+ ibmveth_dbg_root);
+ debugfs_create_file("buffer_pools", 0400, adapter->debugfs_dir,
+ adapter, &ibmveth_buffer_pools_fops);
+}
+
+static void ibmveth_debugfs_exit(struct ibmveth_adapter *adapter)
+{
+ debugfs_remove_recursive(adapter->debugfs_dir);
+ adapter->debugfs_dir = NULL;
+}
+
/**
* ibmveth_pool_kobj_release - Mark a pool kobject finished
* @kobj: kobject embedded in the pool
@@ -3799,6 +3861,8 @@ static int ibmveth_probe(struct vio_dev *dev, const struct vio_device_id *id)
netdev_dbg(netdev, "registered\n");
+ ibmveth_debugfs_init(adapter);
+
return 0;
}
@@ -3807,6 +3871,8 @@ static void ibmveth_remove(struct vio_dev *dev)
struct net_device *netdev = dev_get_drvdata(&dev->dev);
struct ibmveth_adapter *adapter = netdev_priv(netdev);
+ ibmveth_debugfs_exit(adapter);
+
ibmveth_put_pool_kobjs(adapter, IBMVETH_NUM_BUFF_POOLS);
/*
@@ -4032,15 +4098,27 @@ static struct vio_driver ibmveth_driver = {
static int __init ibmveth_module_init(void)
{
+ int rc;
+
printk(KERN_DEBUG "%s: %s %s\n", ibmveth_driver_name,
ibmveth_driver_string, ibmveth_driver_version);
- return vio_register_driver(&ibmveth_driver);
+ ibmveth_dbg_root = debugfs_create_dir(ibmveth_driver_name, NULL);
+
+ rc = vio_register_driver(&ibmveth_driver);
+ if (rc) {
+ debugfs_remove_recursive(ibmveth_dbg_root);
+ ibmveth_dbg_root = NULL;
+ }
+
+ return rc;
}
static void __exit ibmveth_module_exit(void)
{
vio_unregister_driver(&ibmveth_driver);
+ debugfs_remove_recursive(ibmveth_dbg_root);
+ ibmveth_dbg_root = NULL;
}
module_init(ibmveth_module_init);
diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h
index b35da8bdce5f..11eda224a680 100644
--- a/drivers/net/ethernet/ibm/ibmveth.h
+++ b/drivers/net/ethernet/ibm/ibmveth.h
@@ -386,6 +386,8 @@ struct ibmveth_adapter {
struct ibmveth_rx_queue_stats *rx_qstats;
struct ibmveth_tx_queue_stats *tx_qstats;
+ struct dentry *debugfs_dir;
+
/* Ethtool settings */
u8 duplex;
u32 speed;
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* [PATCH net-next v7 14/15] ibmveth: Implement incremental MQ RX queue resize
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (12 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 13/15] ibmveth: Expose per-queue buffer pool details via debugfs Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-25 18:38 ` [PATCH net-next v7 15/15] ibmveth: Complete set_channels down-path and mq_fallback max_rx cap Mingming Cao
2026-09-26 17:40 ` [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support mingming cao
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
Add ibmveth_resize_rx_queues_incremental() so the live RX channel
count can change without a full device tear-down, and
ibmveth_resize_rx_channels() as the entry point set_channels() calls.
The two directions live in ibmveth_scale_up_rx_queues() and
ibmveth_scale_down_rx_queues(); the incremental helper validates the
request, dispatches to one of them, and refreshes CMO on success.
Scale-up allocates DMA buffer lists and pool memory for the new
queues, registers them with PHYP via H_REG_LOGICAL_LAN_QUEUE and maps
subordinate IRQs, then brings each queue up: publish the count,
replenish, enable NAPI, and only then unmask PHYP, so the handler
cannot run on an unpublished, empty or NAPI-disabled queue. It
finishes with the same restart helper open() uses, so a total
replenish miss can still get a poll to retry.
Widen netif_set_real_num_rx_queues() to the target count before any
new queue is unmasked, so a frame on a new queue cannot carry an rx
queue index above real_num_rx_queues. Scale-up failure restores
real_num to the surviving count.
Scale-down masks PHYP on the retiring queues first, then disables NAPI,
then masks and synchronises again, because an in-flight poll can re-arm
PHYP while napi_disable() is waiting. It then drains the ring, harvests
the queue's final no_buffer count under that queue's replenish_lock,
since netpoll can still reach it until the count is lowered, publishes
the surviving count with smp_store_release() and calls
synchronize_net(), then netif_set_real_num_rx_queues(). On success
it then deregisters via H_FREE_LOGICAL_LAN_QUEUE, releases the IRQ
mapping and frees the DMA and pool memory. Deregistration has to
precede the unmap so PHYP releases ownership of the buffers while the
queue metadata its correlators index is still valid. If the free
hcall fails, leave the handle set and skip the unmap. Destroy walks
high-to-low so a failed free keeps 0..i live, publishes that count,
and restores NAPI/IRQ on the surviving new queues.
Scale-up register failure matches that skip-unmap rule: if IRQ
mapping fails after H_REG and the matching H_FREE fails, keep the
handle so the caller does not unmap. Scale-up enable_irq failure
that cannot destroy the new queue keeps that queue in the live set
with the queues this -L already added, instead of publishing below
it and orphaning the mapping.
A scale-up register or IRQ-setup failure whose H_FREE also fails
leaves that queue above the live count with its handle set. close()
frees it after successful H_FREE_LOGICAL_LAN, which drops every
subordinate queue; until then a scale-up that reaches that index
returns -EBUSY.
Both directions unwind. Scale-up failures jump to cleanup_new_queues:,
which restores real_num to the surviving count and destroys only the
queues this call created. Scale-down failures republish the remaining
count, then replenish, napi_enable and unmask each surviving new
queue (and restart if the IRQ comes back). If the IRQ cannot be
re-enabled they schedule a reset.
Use buffer-list allocation presence, not DMA-address truthiness, to
decide whether an incrementally removed queue needs
dma_unmap_single(): DMA address zero is valid.
Scale-up register failure returning -EOPNOTSUPP latches
mq_fallback so future scale-ups are blocked immediately. Reject
rx > 1 with -EOPNOTSUPP when firmware lacks MQ support or
mq_fallback is set. That rejection lives here rather than in
set_channels(), and it sits below the no-op check: ethtool -L is
read-modify-write, so a TX-only request arrives carrying the current
RX count and must not be rejected once mq_fallback is set. An rx_count
outside 1..IBMVETH_MAX_RX_QUEUES is rejected with -EINVAL ahead of both
checks; the ethtool core already range-checks against the max_rx that
get_channels() reports, and this is the driver's own guard.
Refresh CMO entitlement across the resize. ibmveth_desired_dma_for_rxqs()
is factored out of ibmveth_get_desired_dma() so the same math can be
applied to a prospective queue count before the driver commits to it.
Key the resize path's down check on adapter->opened rather than
IFF_UP: RX buffers, mappings and IRQs exist only after a successful
open, so IFF_UP alone does not mean there is anything to resize. An
RX count change while down is rejected with -EOPNOTSUPP here rather
than reported as success and discarded; patch 15 publishes the
down-path count and lifts the rejection.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- drain_rx_queue kdoc and debug log: slots
processed, not buffers recycled
- skip unmap when H_FREE_LOGICAL_LAN_QUEUE fails;
keep the handle and restore remaining queues
- widen set_real before scale-up unmask;
restore real_num on scale-up failure
- scale-up register fail: keep handle and skip
unmap if H_FREE after IRQ map fails
- scale-up enable_irq fail: if destroy fails,
keep that queue in the live set
- scale-up register / IRQ-setup failure with a failed
H_FREE: close() frees that queue after successful
H_FREE_LOGICAL_LAN; a scale-up to it returns -EBUSY
- scale-up register site latches mq_fallback
on -EOPNOTSUPP
- scale-up enable_irq failure: remask PHYP after
napi_disable(), as scale-down does
- CMO desired sites share ibmveth_refresh_cmo_desired();
desired_dma_for_rxqs comment notes CMO and MQ do not
coexist
- free_single_rx_queue() keys buffer-list unmap on
allocation presence because DMA address zero is valid
- reject an RX count change while down with -EOPNOTSUPP
until patch 15 publishes the down-path count
- split resize_rx_queues_incremental() into
scale_up_rx_queues() / scale_down_rx_queues() plus a
dispatcher; code motion only
- schedule_rx_queue kdoc: past the live count is quiet
- noted: mask-then-napi_disable stays;
no PHYP mask ack; no reset on failed
H_FREE (skip unmap); shrinks to 2..live-1
stay -EOPNOTSUPP once mq_fallback; P15
owns TX-fail RX rollback
Changes in v6:
- harvest no_buffer on retiring queues under that queue's
replenish_lock before publish: netpoll still reaches them until
the surviving count is stored, and update_rx_no_buffer() ignores
indexes at or above the live count
- publish num_rx_queues, then synchronize_net(), then destroy.
Scale-down runs set_real_num_rx_queues() between sync and destroy
(destroy only if that succeeds); cleanup_new_queues does not harvest
- napi_disable() before publishing the lowered count on scale-up
enable_irq failure, matching scale-down
- restart_rx_queue() after the checked enable_irq() on scale-up and
on scale-down rollback; kick_rx_queue_if_pending() cannot retry a
queue whose replenish posted nothing
- drop the scale-down survivor restart_rx_queue() loop: those queues
were never quiesced, and enable_irq() on prep-fail unmasks behind
a live poll (subordinate H_PARAMETER on enable is -EIO and reset)
- schedule_rx_queue() returns quietly when only the upper bound
trips: netpoll walks a snapshot of the old count, so an ethtool -L
shrink raced it into a WARN userspace can trigger at will
- drop leftover spin_lock_init() from alloc_single_rx_queue(); scale-up
runs on a live adapter and probe already inits the lock
- zero buffer_list_dma[] on the alloc_single_rx_queue() map-failure
path; free_single_rx_queue() keys its unmap on that field
- skip PHYP mask and remask on scale-down when queue_irq[i] is 0
- put the rx > 1 / mq_fallback gate below the no-op check, and refuse
any rx > 1 once mq_fallback is set, so a TX-only read-modify-write
is not rejected
- noted: down-path stash/CMO and RX rollback on TX fail are patch 15
- noted: KUnit free_map fixtures are patch 8
Changes in v5:
- resize / get_desired_dma read live count via get_num_rx_queues()
- Scale-down / scale-up-fail: remask+sync after napi_disable (poll may
re-arm while disable waits)
- Series renumber: mailed v4 12/14 resize -> tip P14 (14->15)
- Absorb kitchen-sink resize ownership: teardown-first scale-down
(drain/deregister before unmap) + live-pool correlator_valid /
schedule_work (cover has peel map)
- Ordered num_rx_queues publish + publish-before-destroy
- After unmask: kick if pending; enable_irq errno + rollback check;
raise CMO before scale-up allocs
- Scale-up enable_irq-fail: publish-down / napi_disable / drain / destroy
- After scale-down destroy, restart remaining RX queues so Q0 cannot stay
masked with NAPI idle
- resize_rx_channels: reject rx>1 without MQ / validate 1..MAX; call
before !opened early return (stash still tip P15)
Changes in v4:
- Copy pool->index when cloning buffer pools for incrementally added
queues.
- Scale-up order: publish -> replenish -> napi_enable -> enable_irq
(vs older enable-before-publish drafts); mirror NAPI-before-unmask on
set_real_num_rx rollback.
- Linux-owned subordinate virq disposal; deregister is PHYP-only;
dispose bound to MAX_RX_QUEUES; dispose on request_irq failure.
- Drain-path smp_rmb() before harvest.
- Mask PHYP before napi_disable/drain on scale-down and scale-up fail
cleanup.
- Replenish before re-enabling IRQ/NAPI on set_real_num_rx scale-down
rollback.
- Keep set_channels wiring as the following patch (same split as v3)
but call resize_rx_channels() here when IFF_UP so the helper is not
an unused static.
drivers/net/ethernet/ibm/ibmveth.c | 1003 ++++++++++++++++++++++++++--
1 file changed, 942 insertions(+), 61 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 0ac0359bb71e..1b1dd89dadf7 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -774,6 +774,58 @@ ibmveth_cleanup_rx_interrupts(struct ibmveth_adapter *adapter)
adapter->rx_irq_setup = false;
}
+/**
+ * ibmveth_setup_single_rx_interrupt - Setup interrupt for a single RX queue
+ * @adapter: ibmveth adapter structure
+ * @queue_idx: Queue index to setup
+ *
+ * Registers the IRQ handler for one queue. Used during incremental
+ * scale-up when adding new RX queues. The caller publishes the queue,
+ * replenishes buffers, enables NAPI, then unmasks PHYP delivery.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_setup_single_rx_interrupt(struct ibmveth_adapter *adapter,
+ int queue_idx)
+{
+ struct net_device *netdev = adapter->netdev;
+ int rc;
+
+ rc = request_irq(adapter->queue_irq[queue_idx], ibmveth_interrupt,
+ 0, netdev->name, &adapter->napi[queue_idx]);
+ if (rc) {
+ netdev_err(netdev, "request_irq() failed for queue %d: %d\n",
+ queue_idx, rc);
+ return rc;
+ }
+
+ netdev_dbg(netdev, "Setup IRQ %d for queue %d\n",
+ adapter->queue_irq[queue_idx], queue_idx);
+ return 0;
+}
+
+/**
+ * ibmveth_cleanup_single_rx_interrupt - Cleanup interrupt for a single RX queue
+ * @adapter: ibmveth adapter structure
+ * @queue_idx: Queue index to cleanup
+ *
+ * Frees the IRQ handler for one queue and releases the subordinate virq
+ * mapping. Used during incremental scale-down.
+ */
+static void
+ibmveth_cleanup_single_rx_interrupt(struct ibmveth_adapter *adapter,
+ int queue_idx)
+{
+ if (adapter->queue_irq[queue_idx]) {
+ free_irq(adapter->queue_irq[queue_idx],
+ &adapter->napi[queue_idx]);
+ ibmveth_dispose_subordinate_irq_mapping(adapter, queue_idx);
+ netdev_dbg(adapter->netdev,
+ "Freed IRQ for queue %d\n", queue_idx);
+ }
+}
+
/**
* ibmveth_schedule_rx_queue - Mask PHYP IRQ and schedule NAPI for one RX queue
* @adapter: ibmveth adapter structure
@@ -785,15 +837,25 @@ ibmveth_cleanup_rx_interrupts(struct ibmveth_adapter *adapter)
* Return: true if napi_schedule_prep() succeeded and NAPI was scheduled.
* Mask is attempted in that case; a failed disable_irq() is logged by the
* helper and does not change the return (queue may still be unmasked).
- * false if the index is out of range (WARN_ON, then return) or
- * prep failed (including NAPI already scheduled).
+ * false if the index is negative (WARN_ON), past the live count (quiet:
+ * a shrink can race netpoll), or prep failed (including NAPI already
+ * scheduled).
*/
static bool ibmveth_schedule_rx_queue(struct ibmveth_adapter *adapter,
int qindex)
{
struct napi_struct *napi = &adapter->napi[qindex];
- if (WARN_ON(qindex < 0 || qindex >= ibmveth_get_num_rx_queues(adapter)))
+ if (WARN_ON(qindex < 0))
+ return false;
+
+ /*
+ * A live shrink can publish a lower count while netpoll walks a
+ * snapshot of the old one, so an index past the end is expected
+ * here and must not splat. ibmveth_replenish_task() skips the
+ * same way; callers already treat false as "queue is gone".
+ */
+ if (qindex >= ibmveth_get_num_rx_queues(adapter))
return false;
/*
@@ -1261,8 +1323,7 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
* SQ open leaves PHYP masked until the first poll. If schedule_prep fails,
* NAPI never runs and the queue stays masked (TX OK, RX/ARP dead) until
* reload. Replenish first so an enable_irq fallback can actually deliver.
- * Also used after every open (SQ and MQ) and after scale-down so a
- * queue is not left idle+masked.
+ * Also used after every open (SQ and MQ) and after scale-down rollback.
*/
static void ibmveth_restart_rx_queue(struct ibmveth_adapter *adapter,
int qindex)
@@ -1452,6 +1513,141 @@ ibmveth_free_buffer_pools(struct ibmveth_adapter *adapter)
ibmveth_get_num_rx_queues(adapter));
}
+/**
+ * ibmveth_alloc_single_rx_queue - Allocate resources for a single RX queue
+ * @adapter: ibmveth adapter structure
+ * @queue_idx: Queue index to allocate
+ * @rxq_entries: Number of RX queue entries
+ *
+ * Allocates buffer list, RX queue, and per-queue buffer pools for one queue.
+ * Used during incremental scale-up without affecting existing queues.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_alloc_single_rx_queue(struct ibmveth_adapter *adapter, int queue_idx,
+ int rxq_entries)
+{
+ struct device *dev = &adapter->vdev->dev;
+ struct net_device *netdev = adapter->netdev;
+ int i, rc = -ENOMEM;
+
+ adapter->buffer_list_addr[queue_idx] =
+ (void *)get_zeroed_page(GFP_KERNEL);
+ if (!adapter->buffer_list_addr[queue_idx]) {
+ netdev_err(netdev, "unable to allocate buffer list for queue %d\n",
+ queue_idx);
+ return -ENOMEM;
+ }
+
+ adapter->rx_queue[queue_idx].queue_len =
+ sizeof(struct ibmveth_rx_q_entry) * rxq_entries;
+ adapter->rx_queue[queue_idx].queue_addr =
+ dma_alloc_coherent(dev, adapter->rx_queue[queue_idx].queue_len,
+ &adapter->rx_queue[queue_idx].queue_dma,
+ GFP_KERNEL);
+ if (!adapter->rx_queue[queue_idx].queue_addr) {
+ netdev_err(netdev, "unable to allocate RX queue for queue %d\n",
+ queue_idx);
+ goto out_free_buflist;
+ }
+
+ adapter->buffer_list_dma[queue_idx] =
+ dma_map_single(dev, adapter->buffer_list_addr[queue_idx],
+ 4096, DMA_BIDIRECTIONAL);
+ if (dma_mapping_error(dev, adapter->buffer_list_dma[queue_idx])) {
+ netdev_err(netdev, "unable to map buffer list for queue %d\n",
+ queue_idx);
+ adapter->buffer_list_dma[queue_idx] = 0;
+ goto out_free_rxq;
+ }
+
+ for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
+ struct ibmveth_buff_pool *src =
+ &adapter->rx_buff_pool[0][i];
+ struct ibmveth_buff_pool *dst =
+ &adapter->rx_buff_pool[queue_idx][i];
+
+ dst->size = src->size;
+ dst->index = src->index;
+ dst->buff_size = src->buff_size;
+ dst->threshold = src->threshold;
+ dst->active = src->active;
+ }
+
+ rc = ibmveth_alloc_queue_buffer_pools(adapter, queue_idx);
+ if (rc) {
+ netdev_err(netdev,
+ "Failed to allocate buffer pools for queue %d\n",
+ queue_idx);
+ goto out_unmap_buflist;
+ }
+
+ adapter->rx_queue[queue_idx].index = 0;
+ adapter->rx_queue[queue_idx].num_slots = rxq_entries;
+ adapter->rx_queue[queue_idx].toggle = 1;
+
+ netdev_dbg(netdev,
+ "Allocated queue %d: buffer_list @ %p (DMA: 0x%llx), rx_queue @ %p (DMA: 0x%llx), %d entries\n",
+ queue_idx, adapter->buffer_list_addr[queue_idx],
+ (unsigned long long)adapter->buffer_list_dma[queue_idx],
+ adapter->rx_queue[queue_idx].queue_addr,
+ (unsigned long long)adapter->rx_queue[queue_idx].queue_dma,
+ rxq_entries);
+
+ return 0;
+
+out_unmap_buflist:
+ dma_unmap_single(dev, adapter->buffer_list_dma[queue_idx],
+ 4096, DMA_BIDIRECTIONAL);
+ adapter->buffer_list_dma[queue_idx] = 0;
+out_free_rxq:
+ dma_free_coherent(dev, adapter->rx_queue[queue_idx].queue_len,
+ adapter->rx_queue[queue_idx].queue_addr,
+ adapter->rx_queue[queue_idx].queue_dma);
+ adapter->rx_queue[queue_idx].queue_addr = NULL;
+out_free_buflist:
+ free_page((unsigned long)adapter->buffer_list_addr[queue_idx]);
+ adapter->buffer_list_addr[queue_idx] = NULL;
+ return rc;
+}
+
+/**
+ * ibmveth_free_single_rx_queue - Free resources for a single RX queue
+ * @adapter: ibmveth adapter structure
+ * @queue_idx: Queue index to free
+ *
+ * Frees buffer list, RX queue, and per-queue buffer pools for one queue.
+ * Used during incremental scale-down without affecting remaining queues.
+ */
+static void
+ibmveth_free_single_rx_queue(struct ibmveth_adapter *adapter, int queue_idx)
+{
+ struct device *dev = &adapter->vdev->dev;
+
+ ibmveth_free_queue_buffer_pools(adapter, queue_idx);
+
+ if (adapter->buffer_list_addr[queue_idx]) {
+ dma_unmap_single(dev, adapter->buffer_list_dma[queue_idx],
+ 4096, DMA_BIDIRECTIONAL);
+ adapter->buffer_list_dma[queue_idx] = 0;
+ }
+
+ if (adapter->rx_queue[queue_idx].queue_addr) {
+ dma_free_coherent(dev, adapter->rx_queue[queue_idx].queue_len,
+ adapter->rx_queue[queue_idx].queue_addr,
+ adapter->rx_queue[queue_idx].queue_dma);
+ adapter->rx_queue[queue_idx].queue_addr = NULL;
+ }
+
+ if (adapter->buffer_list_addr[queue_idx]) {
+ free_page((unsigned long)adapter->buffer_list_addr[queue_idx]);
+ adapter->buffer_list_addr[queue_idx] = NULL;
+ }
+
+ netdev_dbg(adapter->netdev, "Freed queue %d resources\n", queue_idx);
+}
+
static bool ibmveth_rxq_correlator_valid(struct ibmveth_adapter *adapter,
int queue_index, u64 correlator)
{
@@ -1614,6 +1810,58 @@ static int ibmveth_rxq_harvest_buffer(struct ibmveth_adapter *adapter,
return 0;
}
+/**
+ * ibmveth_drain_rx_queue - Drain pending buffers from an RX queue
+ * @adapter: ibmveth adapter structure
+ * @queue_index: Queue index to drain
+ *
+ * Harvests pending completions back to the per-queue buffer pools.
+ * A corrupt slot still advances the ring, so the return is slots
+ * processed, not buffers recycled.
+ * Must be called with NAPI disabled for this queue.
+ *
+ * Return: number of slots processed
+ */
+static int
+ibmveth_drain_rx_queue(struct ibmveth_adapter *adapter, int queue_index)
+{
+ struct net_device *netdev = adapter->netdev;
+ int drained = 0;
+ int limit = adapter->rx_queue[queue_index].num_slots;
+ int rc;
+
+ netdev_dbg(netdev, "Draining RX queue %d (limit: %d slots)\n",
+ queue_index, limit);
+
+ while (drained < limit &&
+ ibmveth_rxq_pending_buffer(adapter, queue_index)) {
+ /* Match poll-side order before harvesting completion state. */
+ smp_rmb();
+ rc = ibmveth_rxq_harvest_buffer(adapter, queue_index, true);
+ if (rc) {
+ /* -EINVAL/-EFAULT already advanced past the slot. */
+ if (rc == -EINVAL || rc == -EFAULT) {
+ drained++;
+ continue;
+ }
+ netdev_err(netdev,
+ "Failed to harvest buffer from queue %d during drain: %d\n",
+ queue_index, rc);
+ break;
+ }
+ drained++;
+ }
+
+ if (drained > 0)
+ netdev_dbg(netdev, "Drained %d slot(s) from RX queue %d\n",
+ drained, queue_index);
+ else
+ netdev_dbg(netdev, "No slots to drain from RX queue %d\n",
+ queue_index);
+
+ return drained;
+}
+
static void ibmveth_free_tx_ltb(struct ibmveth_adapter *adapter, int idx)
{
void *ltb = adapter->tx_ltb_ptr[idx];
@@ -1760,7 +2008,8 @@ static int ibmveth_register_logical_lan(struct ibmveth_adapter *adapter,
* Registers a subordinate receive queue using H_REG_LOGICAL_LAN_QUEUE.
* On success, stores the queue handle and virtual IRQ in the adapter.
* If IRQ mapping fails after a successful hypervisor registration, the
- * queue is freed before returning.
+ * queue is freed before returning. If that free fails, the handle is
+ * kept so the caller can skip unmap.
*
* Return: H_SUCCESS on success, negative errno on IRQ mapping failure,
* hypervisor error code otherwise
@@ -1800,10 +2049,17 @@ ibmveth_register_logical_lan_queue(struct ibmveth_adapter *adapter,
free_rc = h_free_logical_lan_queue(ua, handle);
} while (H_IS_LONG_BUSY(free_rc) ||
(free_rc == H_BUSY));
- if (free_rc != H_SUCCESS)
+ if (free_rc != H_SUCCESS) {
netdev_err(adapter->netdev,
"h_free_logical_lan_queue failed for queue %d after IRQ map failure: rc=0x%lx\n",
queue_index, free_rc);
+ /*
+ * PHYP still owns this queue. Keep the
+ * handle so the caller skips unmap.
+ */
+ adapter->queue_handle[queue_index] = handle;
+ return -EIO;
+ }
return -EINVAL;
}
@@ -1882,6 +2138,596 @@ ibmveth_register_single_rx_queue(struct ibmveth_adapter *adapter,
return 0;
}
+/**
+ * ibmveth_deregister_single_rx_queue - Deregister one subordinate RX queue
+ * @adapter: ibmveth adapter structure
+ * @queue_idx: Queue index to deregister (1..N)
+ *
+ * Deregisters a single queue via H_FREE_LOGICAL_LAN_QUEUE. Linux IRQ handler
+ * teardown and subordinate virq mapping disposal are owned by interrupt
+ * cleanup helpers; queue 0 is freed only through ibmveth_free_all_queues()
+ * (H_FREE_LOGICAL_LAN).
+ *
+ * Return: 0 if the queue was freed or had no handle, -EIO if
+ * H_FREE_LOGICAL_LAN_QUEUE failed. On -EIO the handle is
+ * left set so the caller can skip unmap.
+ */
+static int
+ibmveth_deregister_single_rx_queue(struct ibmveth_adapter *adapter,
+ int queue_idx)
+{
+ unsigned long lpar_rc;
+ unsigned long ua = adapter->vdev->unit_address;
+ unsigned long qh = adapter->queue_handle[queue_idx];
+
+ if (!qh)
+ return 0;
+
+ do {
+ lpar_rc = h_free_logical_lan_queue(ua, qh);
+ } while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
+
+ if (lpar_rc != H_SUCCESS) {
+ netdev_err(adapter->netdev,
+ "h_free_logical_lan_queue failed, queue %d rc=%ld\n",
+ queue_idx, lpar_rc);
+ return -EIO;
+ }
+
+ adapter->queue_handle[queue_idx] = 0;
+
+ netdev_dbg(adapter->netdev, "Deregistered queue %d\n", queue_idx);
+ return 0;
+}
+
+/**
+ * ibmveth_destroy_subordinate_rx_queue - Tear down one subordinate RX queue
+ * @adapter: ibmveth adapter structure
+ * @queue_idx: Queue index to destroy (1..N)
+ *
+ * Deregister with PHYP before unmapping buffer pools so hypervisor buffer
+ * ownership is released while queue metadata is still valid. If the
+ * hcall fails, leave the queue mapped and keep the handle.
+ *
+ * Return: 0 on success, -EIO if deregister failed
+ */
+static int
+ibmveth_destroy_subordinate_rx_queue(struct ibmveth_adapter *adapter,
+ int queue_idx)
+{
+ int rc;
+
+ rc = ibmveth_deregister_single_rx_queue(adapter, queue_idx);
+ if (rc)
+ return rc;
+ ibmveth_cleanup_single_rx_interrupt(adapter, queue_idx);
+ ibmveth_free_single_rx_queue(adapter, queue_idx);
+ return 0;
+}
+
+/**
+ * ibmveth_desired_dma_for_rxqs - CMO entitlement for a given RX queue count
+ * @adapter: ibmveth adapter
+ * @rxqs: number of RX queues to size for
+ *
+ * Same math as ibmveth_get_desired_dma(), but uses @rxqs instead of the
+ * live adapter->num_rx_queues. Scale-up raises desired for the *target*
+ * count before allocating so vio_cmo_alloc cannot fail mid-resize.
+ *
+ * Return: bytes of IO memory desired for @rxqs RX queues
+ */
+static unsigned long
+ibmveth_desired_dma_for_rxqs(struct ibmveth_adapter *adapter,
+ unsigned int rxqs)
+{
+ struct net_device *netdev = adapter->netdev;
+ struct iommu_table *tbl;
+ unsigned long ret;
+ int i, q;
+
+ tbl = get_iommu_table_base(&adapter->vdev->dev);
+
+ ret = IBMVETH_BUFF_LIST_SIZE * rxqs + IBMVETH_FILT_LIST_SIZE;
+ ret += IOMMU_PAGE_ALIGN(netdev->mtu, tbl);
+ ret += IOMMU_PAGE_ALIGN(IBMVETH_MAX_TX_BUF_SIZE, tbl);
+
+ /*
+ * Pool metadata for queues 1+ is copied from queue 0 at open.
+ * Always size from pool 0 x @rxqs.
+ *
+ * CMO (Power9 and earlier) and MQ firmware (Power11+) do not
+ * coexist, so a CMO partition always sizes one queue here. The
+ * MQ terms and the ethtool -L CMO refreshes are defensive only.
+ */
+ for (q = 0; q < rxqs; q++) {
+ int rxqentries = 1;
+
+ for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
+ struct ibmveth_buff_pool *bpool =
+ &adapter->rx_buff_pool[0][i];
+
+ if (bpool->active)
+ ret += bpool->size *
+ IOMMU_PAGE_ALIGN(bpool->buff_size, tbl);
+ rxqentries += bpool->size;
+ }
+
+ ret += IOMMU_PAGE_ALIGN(rxqentries *
+ sizeof(struct ibmveth_rx_q_entry), tbl);
+ }
+
+ return ret;
+}
+
+/**
+ * ibmveth_free_stranded_rx_queues - Free queues stranded above the live count
+ * @adapter: ibmveth adapter structure
+ *
+ * A scale-up register or IRQ-setup failure whose H_FREE_LOGICAL_LAN_QUEUE
+ * also fails leaves that queue out of the live set with its handle set
+ * and its memory still mapped. Call only after successful H_FREE_LOGICAL_LAN,
+ * which drops every subordinate queue, so the memory can be released.
+ */
+static void ibmveth_free_stranded_rx_queues(struct ibmveth_adapter *adapter)
+{
+ int i;
+
+ for (i = ibmveth_get_num_rx_queues(adapter);
+ i < IBMVETH_MAX_RX_QUEUES; i++) {
+ if (!adapter->queue_handle[i])
+ continue;
+ adapter->queue_handle[i] = 0;
+ ibmveth_free_single_rx_queue(adapter, i);
+ }
+}
+
+/* Set CMO desired entitlement from the live RX queue count. */
+static void ibmveth_refresh_cmo_desired(struct ibmveth_adapter *adapter)
+{
+ if (firmware_has_feature(FW_FEATURE_CMO))
+ vio_cmo_set_dev_desired(adapter->vdev,
+ ibmveth_get_desired_dma(adapter->vdev));
+}
+
+/**
+ * ibmveth_scale_up_rx_queues - Add RX queues to a live adapter
+ * @adapter: ibmveth adapter structure
+ * @old_count: Current number of RX queues
+ * @new_count: Target number of RX queues, above @old_count
+ * @rxq_entries: Number of entries per RX queue
+ *
+ * Raises CMO entitlement and widens real_num_rx_queues before any new
+ * queue is unmasked, then brings each queue up in turn. A failure
+ * unwinds only the queues this call added; a queue PHYP refuses to
+ * release stays in the live set instead.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_scale_up_rx_queues(struct ibmveth_adapter *adapter, int old_count,
+ int new_count, int rxq_entries)
+{
+ struct net_device *netdev = adapter->netdev;
+ int failed_queue;
+ int rc, i;
+
+ netdev_dbg(netdev, "Scale-up: adding queues %d-%d\n",
+ old_count, new_count - 1);
+
+ /*
+ * Raise CMO desired for the target count before dma_map /
+ * dma_alloc_coherent / replenish (same order as change_mtu).
+ * Do not bump live num_rx_queues here, only entitlement.
+ */
+ if (firmware_has_feature(FW_FEATURE_CMO)) {
+ unsigned long dma;
+
+ dma = ibmveth_desired_dma_for_rxqs(adapter, new_count);
+ vio_cmo_set_dev_desired(adapter->vdev, dma);
+ }
+
+ /*
+ * Widen real_num before any new queue is unmasked so
+ * skb_record_rx_queue() cannot trip get_rps_cpu()
+ * WARN_ONCE. Probe already allocated num_rx_queues.
+ */
+ rc = netif_set_real_num_rx_queues(netdev, new_count);
+ if (rc) {
+ netdev_err(netdev,
+ "Failed to set real RX queues to %d: %d\n",
+ new_count, rc);
+ ibmveth_refresh_cmo_desired(adapter);
+ return rc;
+ }
+
+ for (i = old_count; i < new_count; i++) {
+ if (adapter->queue_handle[i]) {
+ /* Left by a failed H_FREE; close frees it. */
+ netdev_err(netdev,
+ "RX queue %d still held by PHYP, reset pending\n",
+ i);
+ rc = -EBUSY;
+ goto cleanup_new_queues;
+ }
+ rc = ibmveth_alloc_single_rx_queue(adapter, i,
+ rxq_entries);
+ if (rc) {
+ netdev_err(netdev, "Failed to allocate queue %d: %d\n",
+ i, rc);
+ goto cleanup_new_queues;
+ }
+
+ rc = ibmveth_register_single_rx_queue(adapter, i);
+ if (rc) {
+ netdev_err(netdev, "Failed to register queue %d: %d\n",
+ i, rc);
+ if (rc == -EOPNOTSUPP)
+ adapter->mq_fallback = true;
+ if (!ibmveth_deregister_single_rx_queue(adapter,
+ i))
+ ibmveth_free_single_rx_queue(adapter,
+ i);
+ else
+ schedule_work(&adapter->work);
+ goto cleanup_new_queues;
+ }
+
+ rc = ibmveth_setup_single_rx_interrupt(adapter, i);
+ if (rc) {
+ netdev_err(netdev,
+ "Failed to setup IRQ for queue %d: %d\n",
+ i, rc);
+ /* request_irq failed: mapped but no handler */
+ ibmveth_dispose_subordinate_irq_mapping(adapter,
+ i);
+ if (!ibmveth_deregister_single_rx_queue(adapter,
+ i))
+ ibmveth_free_single_rx_queue(adapter,
+ i);
+ else
+ schedule_work(&adapter->work);
+ goto cleanup_new_queues;
+ }
+
+ /*
+ * Fully ready before PHYP delivery, matching open():
+ * publish -> replenish -> napi_enable -> enable_irq.
+ * That way ibmveth_interrupt() cannot run on an
+ * unpublished, empty, or NAPI-disabled queue.
+ */
+ ibmveth_publish_num_rx_queues(adapter, i + 1);
+ ibmveth_replenish_task(adapter, i);
+ napi_enable(&adapter->napi[i]);
+
+ rc = ibmveth_enable_irq(adapter, i);
+ if (rc) {
+ netdev_err(netdev,
+ "Failed to enable IRQ for queue %d: %d\n",
+ i, rc);
+ /*
+ * Published, replenished, and NAPI-enabled,
+ * but PHYP never unmasked. Match scale-down /
+ * shared cleanup: drain posted buffers, then
+ * deregister before unmap via
+ * destroy_subordinate.
+ *
+ * napi_disable() must come BEFORE the count
+ * is lowered, matching scale-down and
+ * cleanup_new_queues. Lowering it first does
+ * not hide queue i from netpoll: after
+ * ndo_poll_controller, netpoll_poll_dev()
+ * calls poll_napi(), which walks dev->napi_list
+ * unbounded by the queue count and skips a NAPI
+ * only once NAPI_STATE_NPSVC is set. Queue i is
+ * enabled here, so ibmveth_poll() would run and
+ * trip its queue_index >= num_rx_queues
+ * WARN_ON. napi_disable() sets NPSVC, so
+ * poll_napi() skips the queue instead.
+ */
+ napi_disable(&adapter->napi[i]);
+ /* A poll may have unmasked PHYP; remask. */
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ ibmveth_publish_num_rx_queues(adapter, i);
+ ibmveth_drain_rx_queue(adapter, i);
+ synchronize_net();
+ }
+ if (rc &&
+ ibmveth_destroy_subordinate_rx_queue(adapter, i)) {
+ int keep = i + 1;
+
+ /*
+ * PHYP still owns queue i. Keep it in the
+ * live set with the queues this -L already
+ * added. cleanup_new_queues would publish
+ * below i and orphan this mapping.
+ */
+ ibmveth_publish_num_rx_queues(adapter, keep);
+ if (netif_set_real_num_rx_queues(netdev,
+ keep))
+ schedule_work(&adapter->work);
+ ibmveth_replenish_task(adapter, i);
+ napi_enable(&adapter->napi[i]);
+ rc = ibmveth_enable_irq(adapter, i);
+ if (rc) {
+ netdev_err(netdev,
+ "IRQ %d rc=%d\n",
+ i, rc);
+ schedule_work(&adapter->work);
+ } else {
+ ibmveth_restart_rx_queue(adapter, i);
+ }
+ return -EIO;
+ }
+ if (rc) {
+ /* enable_irq errno; keep -EIO. */
+ rc = -EIO;
+ goto cleanup_new_queues;
+ }
+ ibmveth_restart_rx_queue(adapter, i);
+ }
+
+ return 0;
+
+cleanup_new_queues:
+ failed_queue = i;
+ if (failed_queue > old_count)
+ netdev_err(netdev,
+ "Scale-up failed at queue %d, cleaning up queues %d-%d\n",
+ failed_queue, old_count, failed_queue - 1);
+ else
+ netdev_err(netdev,
+ "Scale-up failed at queue %d, nothing to clean up\n",
+ failed_queue);
+
+ for (i = old_count; i < failed_queue; i++) {
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ }
+
+ for (i = old_count; i < failed_queue; i++)
+ napi_disable(&adapter->napi[i]);
+
+ /* Same remask as scale-down: poll may have re-armed during disable. */
+ for (i = old_count; i < failed_queue; i++) {
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ }
+
+ for (i = old_count; i < failed_queue; i++)
+ ibmveth_drain_rx_queue(adapter, i);
+
+ /* Drop the live count before freeing the half-added queues. */
+ ibmveth_publish_num_rx_queues(adapter, old_count);
+ if (netif_set_real_num_rx_queues(netdev, old_count))
+ schedule_work(&adapter->work);
+ synchronize_net();
+
+ for (i = failed_queue - 1; i >= old_count; i--) {
+ int keep;
+
+ if (!ibmveth_destroy_subordinate_rx_queue(adapter, i))
+ continue;
+
+ keep = i + 1;
+ ibmveth_publish_num_rx_queues(adapter, keep);
+ if (netif_set_real_num_rx_queues(netdev, keep))
+ schedule_work(&adapter->work);
+ for (i = old_count; i < keep; i++) {
+ int irq_rc;
+
+ ibmveth_replenish_task(adapter, i);
+ /* START: NAPI before PHYP unmask. */
+ napi_enable(&adapter->napi[i]);
+ irq_rc = ibmveth_enable_irq(adapter, i);
+ if (irq_rc) {
+ netdev_err(netdev,
+ "IRQ %d rc=%d\n",
+ i, irq_rc);
+ schedule_work(&adapter->work);
+ continue;
+ }
+ ibmveth_restart_rx_queue(adapter, i);
+ }
+ ibmveth_refresh_cmo_desired(adapter);
+ netdev_warn(netdev,
+ "Keeping %d queues after scale-up failure\n",
+ keep);
+ return rc;
+ }
+
+ /* Roll CMO desired back to the surviving queue count. */
+ ibmveth_refresh_cmo_desired(adapter);
+
+ netdev_warn(netdev, "Keeping %d queues after scale-up failure\n",
+ old_count);
+ return rc;
+}
+
+/**
+ * ibmveth_scale_down_rx_queues - Remove RX queues from a live adapter
+ * @adapter: ibmveth adapter structure
+ * @old_count: Current number of RX queues
+ * @new_count: Target number of RX queues, below @old_count
+ *
+ * Quiesces the retiring queues, drains them and harvests their final
+ * no_buffer counts, publishes the surviving count, and only then
+ * deregisters them. Walks high to low so a failed free leaves 0..i
+ * live and republishes that count.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_scale_down_rx_queues(struct ibmveth_adapter *adapter, int old_count,
+ int new_count)
+{
+ struct net_device *netdev = adapter->netdev;
+ int rc, i;
+
+ netdev_dbg(netdev, "Scale-down: removing queues %d-%d\n",
+ new_count, old_count - 1);
+
+ /*
+ * Mask PHYP before napi_disable so the handler cannot miss
+ * a mask while NAPI is already dead. An in-flight poll can
+ * still re-arm PHYP while napi_disable() waits, so remask
+ * and sync again after NAPI is stopped. Then drain, harvest
+ * final no_buffer stats, publish the surviving count, and
+ * synchronize_net() before destroy so netpoll cannot walk
+ * dying queues (handler may still be registered until destroy).
+ */
+ for (i = new_count; i < old_count; i++) {
+ if (!adapter->queue_irq[i])
+ continue;
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ }
+
+ for (i = new_count; i < old_count; i++)
+ napi_disable(&adapter->napi[i]);
+
+ for (i = new_count; i < old_count; i++) {
+ if (!adapter->queue_irq[i])
+ continue;
+ ibmveth_disable_irq(adapter, i);
+ synchronize_irq(adapter->queue_irq[i]);
+ }
+
+ for (i = new_count; i < old_count; i++)
+ ibmveth_drain_rx_queue(adapter, i);
+
+ /*
+ * netpoll still reaches these queues until the publish
+ * below, and ibmveth_replenish_task() folds the same
+ * counters under the queue's replenish_lock. Take it here
+ * so the harvest stays the single writer it documents.
+ */
+ for (i = new_count; i < old_count; i++) {
+ struct ibmveth_rx_q *rxq = &adapter->rx_queue[i];
+ unsigned long flags;
+
+ spin_lock_irqsave(&rxq->replenish_lock, flags);
+ ibmveth_update_rx_no_buffer(adapter, i);
+ spin_unlock_irqrestore(&rxq->replenish_lock, flags);
+ }
+
+ ibmveth_publish_num_rx_queues(adapter, new_count);
+ synchronize_net();
+
+ rc = netif_set_real_num_rx_queues(netdev, new_count);
+ if (rc) {
+ netdev_err(netdev, "Failed to set real RX queues to %d: %d\n",
+ new_count, rc);
+ ibmveth_publish_num_rx_queues(adapter, old_count);
+ for (i = new_count; i < old_count; i++) {
+ int irq_rc;
+
+ ibmveth_replenish_task(adapter, i);
+ /* START: NAPI before PHYP unmask. */
+ napi_enable(&adapter->napi[i]);
+ irq_rc = ibmveth_enable_irq(adapter, i);
+ if (irq_rc) {
+ netdev_err(netdev,
+ "Failed to re-enable IRQ for queue %d during scale-down rollback (rc=%d), scheduling reset\n",
+ i, irq_rc);
+ schedule_work(&adapter->work);
+ continue;
+ }
+ ibmveth_restart_rx_queue(adapter, i);
+ }
+ return rc;
+ }
+
+ for (i = old_count - 1; i >= new_count; i--) {
+ int keep;
+
+ rc = ibmveth_destroy_subordinate_rx_queue(adapter, i);
+ if (!rc)
+ continue;
+
+ /*
+ * High-to-low so a failed free leaves 0..i live.
+ * Queues i+1..old_count-1 are already gone.
+ */
+ keep = i + 1;
+ ibmveth_publish_num_rx_queues(adapter, keep);
+ if (netif_set_real_num_rx_queues(netdev, keep))
+ schedule_work(&adapter->work);
+ for (i = new_count; i < keep; i++) {
+ int irq_rc;
+
+ ibmveth_replenish_task(adapter, i);
+ /* START: NAPI before PHYP unmask. */
+ napi_enable(&adapter->napi[i]);
+ irq_rc = ibmveth_enable_irq(adapter, i);
+ if (irq_rc) {
+ netdev_err(netdev,
+ "IRQ %d rc=%d\n",
+ i, irq_rc);
+ schedule_work(&adapter->work);
+ continue;
+ }
+ ibmveth_restart_rx_queue(adapter, i);
+ }
+ return rc;
+ }
+
+ return 0;
+}
+
+/**
+ * ibmveth_resize_rx_queues_incremental - Resize RX queue count incrementally
+ * @adapter: ibmveth adapter structure
+ * @new_count: Target number of RX queues
+ * @rxq_entries: Number of entries per RX queue
+ *
+ * Adds or removes RX queues without tearing down the entire adapter.
+ * Active queues continue receiving during scale-up. Scale-up widens
+ * real_num_rx_queues before unmasking a new queue. Scale-down drains
+ * excess queues before deregistering them with the hypervisor.
+ *
+ * Return: 0 on success, negative error code on failure
+ */
+static int
+ibmveth_resize_rx_queues_incremental(struct ibmveth_adapter *adapter,
+ int new_count, int rxq_entries)
+{
+ struct net_device *netdev = adapter->netdev;
+ int old_count = ibmveth_get_num_rx_queues(adapter);
+ int rc;
+
+ if (old_count == new_count) {
+ netdev_dbg(netdev, "RX queue count unchanged (%d), nothing to do\n",
+ old_count);
+ return 0;
+ }
+
+ if (new_count < 1 || new_count > IBMVETH_MAX_RX_QUEUES) {
+ netdev_err(netdev, "Invalid RX queue count %d (must be 1-%d)\n",
+ new_count, IBMVETH_MAX_RX_QUEUES);
+ return -EINVAL;
+ }
+
+ netdev_info(netdev, "Incrementally resizing RX queues: %d to %d\n",
+ old_count, new_count);
+
+ if (new_count > old_count)
+ rc = ibmveth_scale_up_rx_queues(adapter, old_count, new_count,
+ rxq_entries);
+ else
+ rc = ibmveth_scale_down_rx_queues(adapter, old_count,
+ new_count);
+ if (rc)
+ return rc;
+
+ netdev_info(netdev, "Successfully resized to %u RX queues (incremental)\n",
+ ibmveth_get_num_rx_queues(adapter));
+
+ ibmveth_refresh_cmo_desired(adapter);
+
+ return 0;
+}
+
/**
* ibmveth_free_all_queues - Free all RX queues at once
* @adapter: ibmveth adapter structure
@@ -1894,15 +2740,16 @@ ibmveth_register_single_rx_queue(struct ibmveth_adapter *adapter,
* Used during interface close and registration error cleanup.
*
* Retries only H_BUSY and H_IS_LONG_BUSY. On other failures, logs and
- * returns; callers cannot observe hypercall status. queue_handle[] is
- * cleared regardless. Callers still run RX pool and DMA teardown
- * afterward (same as pre-helper close()).
+ * returns -EIO; queue_handle[] is cleared regardless. Callers still run
+ * RX pool and DMA teardown afterward (same as pre-helper close()).
*
* Clears queue handles only; queue_irq[] is released by
* ibmveth_cleanup_rx_interrupts() on close, or by
* ibmveth_dispose_subordinate_irq_mappings() on partial register failure.
+ *
+ * Return: 0 on success, -EIO if H_FREE_LOGICAL_LAN failed
*/
-static void ibmveth_free_all_queues(struct ibmveth_adapter *adapter)
+static int ibmveth_free_all_queues(struct ibmveth_adapter *adapter)
{
unsigned long lpar_rc;
int i;
@@ -1913,13 +2760,16 @@ static void ibmveth_free_all_queues(struct ibmveth_adapter *adapter)
lpar_rc = h_free_logical_lan(adapter->vdev->unit_address);
} while (H_IS_LONG_BUSY(lpar_rc) || (lpar_rc == H_BUSY));
+ for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
+ adapter->queue_handle[i] = 0;
+
if (lpar_rc != H_SUCCESS) {
netdev_err(adapter->netdev,
"h_free_logical_lan failed: %ld\n", lpar_rc);
+ return -EIO;
}
- for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++)
- adapter->queue_handle[i] = 0;
+ return 0;
}
/**
@@ -2152,7 +3002,8 @@ static int ibmveth_close(struct net_device *netdev)
ibmveth_update_rx_no_buffer(adapter, i);
spin_unlock_irqrestore(&rxq->replenish_lock, flags);
}
- ibmveth_free_all_queues(adapter);
+ if (!ibmveth_free_all_queues(adapter))
+ ibmveth_free_stranded_rx_queues(adapter);
/* Free TX LTBs after quiesce and after H_FREE_LOGICAL_LAN so xmit
* cannot touch unmapped bounce buffers while the LAN is live.
*/
@@ -2559,6 +3410,74 @@ static void ibmveth_get_channels(struct net_device *netdev,
channels->rx_count = ibmveth_get_num_rx_queues(adapter);
}
+/**
+ * ibmveth_resize_rx_channels - Validate and apply a new RX queue count
+ * @adapter: ibmveth adapter
+ * @goal_rx: requested RX queue count
+ *
+ * Rejects rx > 1 without MQ firmware (-EOPNOTSUPP) and rx outside
+ * 1..IBMVETH_MAX_RX_QUEUES (-EINVAL). An RX count change while the
+ * device is down is rejected (-EOPNOTSUPP); publishing it without
+ * allocating arrives in the next patch. When up, apply via
+ * ibmveth_resize_rx_queues_incremental().
+ *
+ * Return: 0 or negative errno
+ */
+static int ibmveth_resize_rx_channels(struct ibmveth_adapter *adapter,
+ unsigned int goal_rx)
+{
+ struct net_device *netdev = adapter->netdev;
+ unsigned int old_rx = ibmveth_get_num_rx_queues(adapter);
+ int rxq_entries;
+ int rc;
+
+ if (goal_rx < 1 || goal_rx > IBMVETH_MAX_RX_QUEUES) {
+ netdev_err(netdev,
+ "Invalid RX queue count %u (must be 1-%d)\n",
+ goal_rx, IBMVETH_MAX_RX_QUEUES);
+ return -EINVAL;
+ }
+
+ /*
+ * Check for a no-op before the capability gate. ethtool -L is
+ * read-modify-write, so a TX-only request arrives carrying the
+ * current RX count; gating first would fail those with
+ * -EOPNOTSUPP once mq_fallback is set.
+ */
+ if (goal_rx == old_rx)
+ return 0;
+
+ /*
+ * Refuse any rx > 1, not just growth: once mq_fallback is set the
+ * next open comes up single-queue, so an intermediate count could
+ * not be honoured either, and accepting it would only repeat the
+ * silent clamp at open. max_rx stays at the live count so that
+ * read-modify-write TX-only requests still clear the core.
+ */
+ if (goal_rx > 1 && (!adapter->multi_queue || adapter->mq_fallback)) {
+ netdev_err(netdev,
+ "Cannot resize to %u RX queues: multi-queue mode not supported by firmware\n",
+ goal_rx);
+ return -EOPNOTSUPP;
+ }
+
+ /*
+ * Down / failed-open: there is nothing to resize, and publishing
+ * the desired count without allocating arrives in the next
+ * patch. Refuse rather than report success for a request that
+ * would be discarded.
+ */
+ if (!adapter->opened)
+ return -EOPNOTSUPP;
+
+ rxq_entries = adapter->rx_queue[0].num_slots;
+ rc = ibmveth_resize_rx_queues_incremental(adapter, goal_rx,
+ rxq_entries);
+ if (rc)
+ netdev_err(netdev, "Failed to resize RX queues: %d\n", rc);
+ return rc;
+}
+
static int ibmveth_set_channels(struct net_device *netdev,
struct ethtool_channels *channels)
{
@@ -2567,18 +3486,15 @@ static int ibmveth_set_channels(struct net_device *netdev,
goal = channels->tx_count;
int rc, i;
- /*
- * RX channel resize is implemented in a later patch; reject any
- * request that changes rx_count. Read-modify-write TX adjustments
- * submit the current rx_count and proceed.
+ /* Validate RX (and resize when opened) before the down-path
+ * early return so MQ/range errors are reported here. Publishing
+ * the desired RX count and CMO while down is the next patch.
*/
- if (channels->rx_count != ibmveth_get_num_rx_queues(adapter))
- return -EOPNOTSUPP;
+ rc = ibmveth_resize_rx_channels(adapter, channels->rx_count);
+ if (rc)
+ return rc;
- /* If ndo_open has not been called yet then don't allocate, just set
- * desired netdev_queue's and return
- */
- if (!(netdev->flags & IFF_UP))
+ if (!adapter->opened)
return netif_set_real_num_tx_queues(netdev, goal);
/* We have IBMVETH_MAX_QUEUES netdev_queue's allocated
@@ -3311,8 +4227,6 @@ static unsigned long ibmveth_get_desired_dma(struct vio_dev *vdev)
struct net_device *netdev = dev_get_drvdata(&vdev->dev);
struct ibmveth_adapter *adapter;
struct iommu_table *tbl;
- unsigned long ret;
- int i, q;
tbl = get_iommu_table_base(&vdev->dev);
@@ -3321,41 +4235,8 @@ static unsigned long ibmveth_get_desired_dma(struct vio_dev *vdev)
return IOMMU_PAGE_ALIGN(IBMVETH_IO_ENTITLEMENT_DEFAULT, tbl);
adapter = netdev_priv(netdev);
-
- /* One buffer list page per RX queue; filter list is shared. */
- ret = IBMVETH_BUFF_LIST_SIZE * ibmveth_get_num_rx_queues(adapter) +
- IBMVETH_FILT_LIST_SIZE;
- ret += IOMMU_PAGE_ALIGN(netdev->mtu, tbl);
- /* add size of mapped tx buffers */
- ret += IOMMU_PAGE_ALIGN(IBMVETH_MAX_TX_BUF_SIZE, tbl);
-
- /*
- * Pool metadata for queues 1+ is copied from queue 0 at open.
- * Always size from pool 0 x num_rx_queues.
- *
- * CMO (Power9 and earlier) and MQ firmware (Power11+) do not
- * coexist, so a CMO partition always sizes one queue here. The
- * MQ terms are defensive only.
- */
- for (q = 0; q < ibmveth_get_num_rx_queues(adapter); q++) {
- int rxqentries = 1;
-
- for (i = 0; i < IBMVETH_NUM_BUFF_POOLS; i++) {
- struct ibmveth_buff_pool *bpool =
- &adapter->rx_buff_pool[0][i];
-
- if (bpool->active)
- ret += bpool->size *
- IOMMU_PAGE_ALIGN(bpool->buff_size, tbl);
- rxqentries += bpool->size;
- }
-
- /* add the size of the receive queue entries */
- ret += IOMMU_PAGE_ALIGN(rxqentries *
- sizeof(struct ibmveth_rx_q_entry), tbl);
- }
-
- return ret;
+ return ibmveth_desired_dma_for_rxqs(adapter,
+ ibmveth_get_num_rx_queues(adapter));
}
static int ibmveth_set_mac_addr(struct net_device *dev, void *p)
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 14/15] ibmveth: Implement incremental MQ RX queue resize
2026-09-25 18:38 ` [PATCH net-next v7 14/15] ibmveth: Implement incremental MQ RX queue resize Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found 5 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 3 · Low: 2
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] Resolved later in this series; not a finding at the end of the
series.
- [Medium] Two rollback paths bring queues back into service even when
widening real_num_rx_queues back up fails.
- [Medium] Four new resize paths call ibmveth_enable_irq(adapter, i) to
unmask the subordinate queue with H_VIOCTL H_ENABLE_VIO_INTERRUPT,…
- [Low] Three comments added or edited by this patch do not match the
code.
- [Low] A subordinate RX queue stranded above the live count can leak.
--
ibmveth: Implement incremental MQ RX queue resize
This patch adds ibmveth_resize_rx_queues_incremental() and
ibmveth_resize_rx_channels(), so ethtool -L can change the live RX queue
count without tearing down the device. Scale-up allocates, registers and
unmasks the new subordinate queues. Scale-down quiesces, drains and
deregisters the retiring ones, and both directions roll back on failure.
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index 0ac0359bb71e6..1b1dd89dadf79 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
[ ... ]
> @@ -785,15 +837,25 @@ ibmveth_cleanup_rx_interrupts(struct ibmveth_adapter *adapter)
[ ... ]
> static bool ibmveth_schedule_rx_queue(struct ibmveth_adapter *adapter,
> int qindex)
> {
> struct napi_struct *napi = &adapter->napi[qindex];
>
> - if (WARN_ON(qindex < 0 || qindex >= ibmveth_get_num_rx_queues(adapter)))
> + if (WARN_ON(qindex < 0))
> + return false;
> +
> + /*
> + * A live shrink can publish a lower count while netpoll walks a
> + * snapshot of the old one, so an index past the end is expected
> + * here and must not splat. ibmveth_replenish_task() skips the
> + * same way; callers already treat false as "queue is gone".
> + */
> + if (qindex >= ibmveth_get_num_rx_queues(adapter))
> return false;
[ ... ]
> @@ -1261,8 +1323,7 @@ static void ibmveth_replenish_task(struct ibmveth_adapter *adapter,
> * SQ open leaves PHYP masked until the first poll. If schedule_prep fails,
> * NAPI never runs and the queue stays masked (TX OK, RX/ARP dead) until
> * reload. Replenish first so an enable_irq fallback can actually deliver.
> - * Also used after every open (SQ and MQ) and after scale-down so a
> - * queue is not left idle+masked.
> + * Also used after every open (SQ and MQ) and after scale-down rollback.
> */
> static void ibmveth_restart_rx_queue(struct ibmveth_adapter *adapter,
> int qindex)
[Severity: Low]
This isn't a bug, but three of the new comments don't match the code.
The updated ibmveth_restart_rx_queue() comment lists open and scale-down
rollback as its callers. ibmveth_scale_up_rx_queues() also calls it in
three places: after each successful queue bring-up, on the keep path after
a failed destroy, and on the cleanup_new_queues keep path.
The new ibmveth_schedule_rx_queue() comment says callers treat false as
"queue is gone". ibmveth_restart_rx_queue() does the opposite on false:
if (ibmveth_schedule_rx_queue(adapter, qindex))
return;
ibmveth_enable_irq(adapter, qindex);
A caller that follows the comment would unmask PHYP on a retired queue.
The ibmveth_cleanup_single_rx_interrupt() kdoc says "Used during
incremental scale-down". It also runs on the scale-up failure paths via
ibmveth_destroy_subordinate_rx_queue().
Could these comments be updated?
[ ... ]
> +static int
> +ibmveth_scale_up_rx_queues(struct ibmveth_adapter *adapter, int old_count,
> + int new_count, int rxq_entries)
> +{
[ ... ]
> + for (i = old_count; i < new_count; i++) {
> + if (adapter->queue_handle[i]) {
> + /* Left by a failed H_FREE; close frees it. */
> + netdev_err(netdev,
> + "RX queue %d still held by PHYP, reset pending\n",
> + i);
> + rc = -EBUSY;
> + goto cleanup_new_queues;
> + }
> + rc = ibmveth_alloc_single_rx_queue(adapter, i,
> + rxq_entries);
> + if (rc) {
> + netdev_err(netdev, "Failed to allocate queue %d: %d\n",
> + i, rc);
> + goto cleanup_new_queues;
> + }
> +
> + rc = ibmveth_register_single_rx_queue(adapter, i);
> + if (rc) {
> + netdev_err(netdev, "Failed to register queue %d: %d\n",
> + i, rc);
> + if (rc == -EOPNOTSUPP)
> + adapter->mq_fallback = true;
> + if (!ibmveth_deregister_single_rx_queue(adapter,
> + i))
> + ibmveth_free_single_rx_queue(adapter,
> + i);
> + else
> + schedule_work(&adapter->work);
> + goto cleanup_new_queues;
> + }
[Severity: Low]
If ibmveth_deregister_single_rx_queue() fails here, queue i keeps its
buffer list page, the dma_map_single() of that page, the
dma_alloc_coherent() RX ring and its pool arrays. queue_handle[i] also
stays set above the live count. The IRQ setup failure branch below does
the same.
Only ibmveth_free_stranded_rx_queues() frees such a queue, and it runs
only from ibmveth_close() when H_FREE_LOGICAL_LAN succeeds:
if (!ibmveth_free_all_queues(adapter))
ibmveth_free_stranded_rx_queues(adapter);
If H_FREE_LOGICAL_LAN fails in close, does this leak the stranded queue's
page, RX ring and DMA mapping? The out_free_all_queues unwind in
ibmveth_open() and ibmveth_remove() also don't seem to release queues
above the live count.
After a failed close, queue_handle[i] also stays set. Every later scale-up
that reaches index i then returns -EBUSY with the "reset pending" message,
but that path doesn't schedule a reset.
In the same failed-close case, the live queues' memory is freed anyway.
Keeping only the stranded queue's memory looks inconsistent.
[ ... ]
> + ibmveth_publish_num_rx_queues(adapter, i + 1);
> + ibmveth_replenish_task(adapter, i);
> + napi_enable(&adapter->napi[i]);
> +
> + rc = ibmveth_enable_irq(adapter, i);
> + if (rc) {
[ ... ]
> + if (rc) {
> + /* enable_irq errno; keep -EIO. */
> + rc = -EIO;
> + goto cleanup_new_queues;
> + }
> + ibmveth_restart_rx_queue(adapter, i);
> + }
[Severity: Medium]
Can this unmask PHYP behind a live poll?
Once ibmveth_enable_irq() unmasks queue i, an interrupt on another CPU can
run ibmveth_interrupt(), win napi_schedule_prep(), mask PHYP and schedule
NAPI. ibmveth_restart_rx_queue() then sees its own prep fail and falls
back to:
ibmveth_enable_irq(adapter, qindex);
That unmasks PHYP while the poll owns the queue. When the poll completes,
it calls ibmveth_enable_irq() again on a source that is already enabled.
ibmveth_toggle_irq() folds H_PARAMETER to success only on disable:
if (h_rc == H_PARAMETER && !enable) {
So on enable it returns -EIO, and ibmveth_poll() schedules a reset.
The v6 changelog gives this case as the reason the survivor restart loop
was dropped: "enable_irq() on prep-fail unmasks behind a live poll
(subordinate H_PARAMETER on enable is -EIO and reset)".
The same enable-then-restart sequence appears in four places:
- here
- the keep branch after a failed destroy
- the cleanup_new_queues keep loop
- both rollback loops in ibmveth_scale_down_rx_queues()
On a busy link, would an ethtool -L that otherwise succeeded end in an
adapter reset?
[ ... ]
> + /* Drop the live count before freeing the half-added queues. */
> + ibmveth_publish_num_rx_queues(adapter, old_count);
> + if (netif_set_real_num_rx_queues(netdev, old_count))
> + schedule_work(&adapter->work);
> + synchronize_net();
> +
> + for (i = failed_queue - 1; i >= old_count; i--) {
> + int keep;
> +
> + if (!ibmveth_destroy_subordinate_rx_queue(adapter, i))
> + continue;
> +
> + keep = i + 1;
> + ibmveth_publish_num_rx_queues(adapter, keep);
> + if (netif_set_real_num_rx_queues(netdev, keep))
> + schedule_work(&adapter->work);
> + for (i = old_count; i < keep; i++) {
> + int irq_rc;
> +
> + ibmveth_replenish_task(adapter, i);
> + /* START: NAPI before PHYP unmask. */
> + napi_enable(&adapter->napi[i]);
> + irq_rc = ibmveth_enable_irq(adapter, i);
[Severity: Medium]
At this point real_num_rx_queues has already been lowered to old_count,
and keep is above old_count. Growing the count goes through
net_rx_queue_update_kobjects(), which can fail (for example with -ENOMEM).
netif_set_real_num_rx_queues() then leaves dev->real_num_rx_queues
unchanged.
If that happens, is it safe to still replenish, napi_enable() and unmask
queues old_count..keep-1? Packets on those queues would carry
skb_record_rx_queue() indexes at or above real_num_rx_queues. That hits
the WARN_ONCE in get_rps_cpu() and the bypass in netif_get_rxqueue(),
which is what widening real_num up front in this function is meant to
avoid.
The destroy failure path in ibmveth_scale_down_rx_queues() has the same
pattern, after real_num has been lowered to new_count:
if (netif_set_real_num_rx_queues(netdev, keep))
schedule_work(&adapter->work);
for (i = new_count; i < keep; i++) {
[ ... ]
> @@ -2567,18 +3486,15 @@ static int ibmveth_set_channels(struct net_device *netdev,
> goal = channels->tx_count;
> int rc, i;
>
> - /*
> - * RX channel resize is implemented in a later patch; reject any
> - * request that changes rx_count. Read-modify-write TX adjustments
> - * submit the current rx_count and proceed.
> + /* Validate RX (and resize when opened) before the down-path
> + * early return so MQ/range errors are reported here. Publishing
> + * the desired RX count and CMO while down is the next patch.
> */
> - if (channels->rx_count != ibmveth_get_num_rx_queues(adapter))
> - return -EOPNOTSUPP;
> + rc = ibmveth_resize_rx_channels(adapter, channels->rx_count);
> + if (rc)
> + return rc;
[Severity: Medium]
The RX resize is committed here, before any TX work. If a TX LTB
allocation later fails in ibmveth_set_channels(), the code does:
goal = old;
old = i;
break;
}
rc = netif_set_real_num_tx_queues(netdev, goal);
This overwrites the allocation error with the result of reverting to the
old TX count, which is normally 0.
For a combined ethtool -L rx N tx M on an opened MQ adapter, would this
change RX, leave TX unchanged and still report success? Nothing at this
commit rolls RX back.
The last patch in the series, "ibmveth: Complete set_channels down-path
and mq_fallback max_rx cap", adds rx_changed tracking, a rollback through
ibmveth_resize_rx_queues_incremental(adapter, old_rx, ...), and a return
of rc ? rc : -ENOMEM when real_num_tx_queues doesn't match the request.
That fixes it by the end of the series, but the behavior above is still
present at this commit.
> - /* If ndo_open has not been called yet then don't allocate, just set
> - * desired netdev_queue's and return
> - */
> - if (!(netdev->flags & IFF_UP))
> + if (!adapter->opened)
> return netif_set_real_num_tx_queues(netdev, goal);
[ ... ]
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* [PATCH net-next v7 15/15] ibmveth: Complete set_channels down-path and mq_fallback max_rx cap
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (13 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 14/15] ibmveth: Implement incremental MQ RX queue resize Mingming Cao
@ 2026-09-25 18:38 ` Mingming Cao
2026-09-29 19:33 ` netdev-bot+sashiko
2026-09-26 17:40 ` [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support mingming cao
15 siblings, 1 reply; 29+ messages in thread
From: Mingming Cao @ 2026-09-25 18:38 UTC (permalink / raw)
To: netdev
Cc: davem, kuba, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1, Mingming Cao
Patch 14 wires live ethtool -L rx. This patch completes the down-path
publish/rollback and get_channels() once mq_fallback is set, replacing
patch 14's temporary -EOPNOTSUPP for an RX count change while down.
When down: set TX queues first, then publish the desired RX count in
adapter->num_rx_queues and netdev->real_num_rx_queues (rx-N sysfs and
CMO desired change immediately). Do not allocate RX mappings, buffers
or IRQs while down. If RX set_real fails, restore the old TX real_num.
When up: resize RX, then the existing TX LTB
stop/alloc/set_real_num_tx/free/wake path. Skip that path when the TX
count is unchanged. If TX cannot reach the requested count, roll RX
back; that rollback is best-effort.
Validate tx_count before touching RX so a request the driver will
reject does not resize RX first. The ethtool core already range-checks
against the max_tx get_channels() reports; this is the driver's own
guard.
get_channels() always reports the live rx_count. When mq_fallback is
set it caps max_rx at that count. Understating rx_count would turn a
TX-only ethtool -L into a silent RX shrink. Advertising max_rx = 1
while rx_count is still live fails the core the same way (rx_count >
max_rx) and blocks that TX-only request. Capping max_rx at the live
count blocks growth without misreporting what is configured.
max_tx is at least the live tx_count. After CPU offline,
ibmveth_real_max_tx_queues() can drop below real_num_tx_queues;
ethtool -L resubmits that count and the core rejects tx_count >
max_tx, which also blocked an RX-only request. set_channels uses
the same ceiling (live count or the online-CPU cap, whichever is
larger) and refuses only a request above that.
i = old_tx before the TX alloc loop is readability; the old
for-initializer already defined i for the free walk.
Guard poll_controller() with adapter->opened so netpoll cannot walk
unallocated queue state while closed.
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Shaik Abdulla <shaik.abdulla1@ibm.com>
---
Changes in v7:
- refresh the TX-fail RX rollback comment
- get_channels max_tx is at least the live
tx_count; set_channels uses that same
ceiling and refuses only a request above it
- lift patch 14's temporary -EOPNOTSUPP for an RX
count change while down
- noted: get_channels keeps the live rx_count;
mq_fallback caps max_rx at that count; do
not advertise max_rx = 1 and do not clamp
rx_count
Changes in v6:
- get_channels() caps max_rx at the live rx_count when mq_fallback
is set; rx_count stays live
- roll back TX real_num if down-path RX set_real fails
- poll_controller() returns if !opened
- noted: rx > 1 reject once mq_fallback is set is patch 14
Changes in v5:
- set_channels / resize_rx_channels read via get_num_rx_queues()
- Series renumber: mailed v4 13/14 set_channels -> tip P15 (end; 14->15)
- set_channels: gate on adapter->opened (not IFF_UP) for live RX resize
vs while-down stash
- Down-path RX stash also refreshes CMO (publish + set_real_num_rx)
- When up: resize RX then TX; roll RX back if TX cannot reach goal
Changes in v4:
- On !IFF_UP, stash num_rx_queues only after TX set succeeds; do not
allocate live subordinate IRQs/buffers while down.
- Initialize i = old_tx on the TX adjust path.
- Always return rc from set_channels().
- Split from the resize-helper patch (same split as v3) while keeping
a live caller of resize_rx_channels() in the previous patch.
drivers/net/ethernet/ibm/ibmveth.c | 170 +++++++++++++++++++++++------
1 file changed, 134 insertions(+), 36 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 1b1dd89dadf7..da14c6915211 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -3399,15 +3399,32 @@ static void ibmveth_get_channels(struct net_device *netdev,
struct ethtool_channels *channels)
{
struct ibmveth_adapter *adapter = netdev_priv(netdev);
+ unsigned int rx_count = ibmveth_get_num_rx_queues(adapter);
- channels->max_tx = ibmveth_real_max_tx_queues();
channels->tx_count = netdev->real_num_tx_queues;
+ /*
+ * Never advertise max_tx below the live count. After CPU
+ * offline, real_max can drop below real_num_tx_queues;
+ * ethtool -L is read-modify-write and the core rejects
+ * tx_count > max_tx, which would also block an RX-only
+ * request. set_channels uses this same ceiling.
+ */
+ channels->max_tx = max_t(unsigned int, channels->tx_count,
+ ibmveth_real_max_tx_queues());
- if (adapter->multi_queue)
+ /*
+ * Always report the live RX count. ethtool -L is read-modify-
+ * write, so a TX-only request echoes rx_count back at us; an
+ * understated value would be applied as a silent RX shrink.
+ * mq_fallback instead caps max_rx at the live count, which
+ * blocks growth in the core without misreporting what is
+ * currently configured.
+ */
+ channels->rx_count = rx_count;
+ if (adapter->multi_queue && !adapter->mq_fallback)
channels->max_rx = IBMVETH_MAX_RX_QUEUES;
else
- channels->max_rx = 1;
- channels->rx_count = ibmveth_get_num_rx_queues(adapter);
+ channels->max_rx = rx_count;
}
/**
@@ -3416,9 +3433,8 @@ static void ibmveth_get_channels(struct net_device *netdev,
* @goal_rx: requested RX queue count
*
* Rejects rx > 1 without MQ firmware (-EOPNOTSUPP) and rx outside
- * 1..IBMVETH_MAX_RX_QUEUES (-EINVAL). An RX count change while the
- * device is down is rejected (-EOPNOTSUPP); publishing it without
- * allocating arrives in the next patch. When up, apply via
+ * 1..IBMVETH_MAX_RX_QUEUES (-EINVAL). When RX resources are not live
+ * (!opened), only validate; do not allocate. When up, apply via
* ibmveth_resize_rx_queues_incremental().
*
* Return: 0 or negative errno
@@ -3461,14 +3477,9 @@ static int ibmveth_resize_rx_channels(struct ibmveth_adapter *adapter,
return -EOPNOTSUPP;
}
- /*
- * Down / failed-open: there is nothing to resize, and publishing
- * the desired count without allocating arrives in the next
- * patch. Refuse rather than report success for a request that
- * would be discarded.
- */
+ /* Down / failed-open: do not allocate. */
if (!adapter->opened)
- return -EOPNOTSUPP;
+ return 0;
rxq_entries = adapter->rx_queue[0].num_slots;
rc = ibmveth_resize_rx_queues_incremental(adapter, goal_rx,
@@ -3482,28 +3493,90 @@ static int ibmveth_set_channels(struct net_device *netdev,
struct ethtool_channels *channels)
{
struct ibmveth_adapter *adapter = netdev_priv(netdev);
- unsigned int old = netdev->real_num_tx_queues,
- goal = channels->tx_count;
+ unsigned int old_rx = ibmveth_get_num_rx_queues(adapter);
+ unsigned int goal_rx = channels->rx_count;
+ unsigned int old_tx = netdev->real_num_tx_queues;
+ unsigned int goal_tx = channels->tx_count;
+ unsigned int want_tx = goal_tx;
+ unsigned int max_tx;
+ bool rx_changed = false;
int rc, i;
- /* Validate RX (and resize when opened) before the down-path
- * early return so MQ/range errors are reported here. Publishing
- * the desired RX count and CMO while down is the next patch.
+ /*
+ * Same ceiling get_channels reports: live count or the
+ * online-CPU cap, whichever is larger.
*/
- rc = ibmveth_resize_rx_channels(adapter, channels->rx_count);
+ max_tx = max_t(unsigned int, old_tx,
+ ibmveth_real_max_tx_queues());
+ if (goal_tx < 1 || goal_tx > max_tx) {
+ netdev_err(netdev,
+ "Invalid TX queue count %u (must be 1-%u)\n",
+ goal_tx, max_tx);
+ return -EINVAL;
+ }
+
+ /* RX range / MQ checks live in ibmveth_resize_rx_channels(). */
+ rc = ibmveth_resize_rx_channels(adapter, goal_rx);
if (rc)
return rc;
- if (!adapter->opened)
- return netif_set_real_num_tx_queues(netdev, goal);
+ /* If RX resources are not live (never opened, or close+open failed
+ * while IFF_UP stayed set), publish desired queue counts without
+ * allocating.
+ */
+ if (!adapter->opened) {
+ /* Apply TX first so a failure leaves the published RX
+ * count unchanged.
+ */
+ rc = netif_set_real_num_tx_queues(netdev, goal_tx);
+ if (rc)
+ return rc;
+
+ /* Publish desired RX count for next open() and refresh CMO;
+ * do not allocate while down.
+ */
+ if (goal_rx != ibmveth_get_num_rx_queues(adapter)) {
+ ibmveth_publish_num_rx_queues(adapter, goal_rx);
+ rc = netif_set_real_num_rx_queues(netdev, goal_rx);
+ if (rc) {
+ int tx_rc;
+
+ ibmveth_publish_num_rx_queues(adapter, old_rx);
+ tx_rc = netif_set_real_num_tx_queues(netdev,
+ old_tx);
+ if (tx_rc)
+ netdev_err(netdev,
+ "Failed to restore TX queues to %u after RX failure: %d\n",
+ old_tx, tx_rc);
+ return rc;
+ }
+ if (firmware_has_feature(FW_FEATURE_CMO)) {
+ unsigned long dma;
+
+ dma = ibmveth_get_desired_dma(adapter->vdev);
+ vio_cmo_set_dev_desired(adapter->vdev, dma);
+ }
+ }
+ return 0;
+ }
+
+ if (goal_rx != old_rx)
+ rx_changed = true;
/* We have IBMVETH_MAX_QUEUES netdev_queue's allocated
* but we may need to alloc/free the ltb's.
*/
+ if (goal_tx == old_tx)
+ return 0;
+
netif_tx_stop_all_queues(netdev);
- /* Allocate any queue that we need */
- for (i = old; i < goal; i++) {
+ /* Allocate any new TX LTBs. i starts at old_tx for the free walk
+ * below when this loop body never runs (goal_tx == old_tx already
+ * returned; goal_tx < old_tx is scale-down).
+ */
+ i = old_tx;
+ for (; i < goal_tx; i++) {
if (adapter->tx_ltb_ptr[i])
continue;
@@ -3512,28 +3585,48 @@ static int ibmveth_set_channels(struct net_device *netdev,
continue;
/* if something goes wrong, free everything we just allocated */
- netdev_err(netdev, "Failed to allocate more tx queues, returning to %d queues\n",
- old);
- goal = old;
- old = i;
+ netdev_err(netdev, "Failed to allocate more tx queues, returning to %u queues\n",
+ old_tx);
+ goal_tx = old_tx;
+ old_tx = i;
break;
}
- rc = netif_set_real_num_tx_queues(netdev, goal);
+ rc = netif_set_real_num_tx_queues(netdev, goal_tx);
if (rc) {
- netdev_err(netdev, "Failed to set real tx queues, returning to %d queues\n",
- old);
- goal = old;
- old = i;
+ netdev_err(netdev, "Failed to set real tx queues, returning to %u queues\n",
+ old_tx);
+ goal_tx = old_tx;
+ old_tx = i;
}
/* Free any that are no longer needed */
- for (i = old; i > goal; i--) {
+ for (i = old_tx; i > goal_tx; i--) {
if (adapter->tx_ltb_ptr[i - 1])
ibmveth_free_tx_ltb(adapter, i - 1);
}
netif_tx_wake_all_queues(netdev);
- return rc;
+ if (netdev->real_num_tx_queues != want_tx) {
+ if (rx_changed) {
+ /*
+ * Restore the RX count from before this -L.
+ * num_slots is the live size after that resize.
+ */
+ int rxq_entries = adapter->rx_queue[0].num_slots;
+ int rb;
+
+ rb = ibmveth_resize_rx_queues_incremental(adapter,
+ old_rx,
+ rxq_entries);
+ if (rb)
+ netdev_err(netdev,
+ "Failed to roll back RX queues to %u after TX failure: %d\n",
+ old_rx, rb);
+ }
+ return rc ? rc : -ENOMEM;
+ }
+
+ return 0;
}
static const struct ethtool_ops netdev_ethtool_ops = {
@@ -4203,9 +4296,14 @@ static int ibmveth_change_mtu(struct net_device *dev, int new_mtu)
static void ibmveth_poll_controller(struct net_device *dev)
{
struct ibmveth_adapter *adapter = netdev_priv(dev);
- unsigned int num = ibmveth_get_num_rx_queues(adapter);
+ unsigned int num;
int i;
+ if (!adapter->opened)
+ return;
+
+ num = ibmveth_get_num_rx_queues(adapter);
+
for (i = 0; i < num; i++)
ibmveth_replenish_task(adapter, i);
--
2.50.1 (Apple Git-155)
^ permalink raw reply related [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 15/15] ibmveth: Complete set_channels down-path and mq_fallback max_rx cap
2026-09-25 18:38 ` [PATCH net-next v7 15/15] ibmveth: Complete set_channels down-path and mq_fallback max_rx cap Mingming Cao
@ 2026-09-29 19:33 ` netdev-bot+sashiko
0 siblings, 0 replies; 29+ messages in thread
From: netdev-bot+sashiko @ 2026-09-29 19:33 UTC (permalink / raw)
To: mmc
Cc: netdev, davem, kuba, horms, edumazet, pabeni, andrew+netdev,
nnac123, maddy, mpe, linuxppc-dev, haren, ricklind, davemarq,
bjking1, shaik.abdulla1
Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 3
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] The new !opened branch of ibmveth_set_channels() raises
adapter->num_rx_queues with `ibmveth_publish_num_rx_queues(adapter,…
- [Low] The commit log refers to other patches by series position: "Patch
14 wires live ethtool -L rx" and "replacing patch 14's temporary…
- [Low] When mq_fallback is set and the live RX count N is greater than
1, ibmveth_get_channels() reports max_rx = rx_count = N.
- [Low] ibmveth_poll_controller() now reads adapter->opened as a plain
bool from netpoll context, which runs without RTNL and concurrently…
--
ibmveth: Complete set_channels down-path and mq_fallback max_rx cap
This commit allows ethtool -L to change the RX count while the device is
down. The new count is published without allocating anything, and on the
live path RX is rolled back if TX fails. It also caps max_rx at the live
count under mq_fallback, raises the max_tx ceiling to at least the live
TX count, and guards ibmveth_poll_controller() with adapter->opened.
> Patch 14 wires live ethtool -L rx. This patch completes the down-path
> publish/rollback and get_channels() once mq_fallback is set, replacing
> patch 14's temporary -EOPNOTSUPP for an RX count change while down.
[Severity: Low]
This isn't a bug, but "Patch 14" and "patch 14's temporary -EOPNOTSUPP"
are series positions. They won't mean anything once this lands in git
history. Could they refer to the earlier commit by its subject instead?
The subject covers only the down path and the mq_fallback max_rx cap.
The patch also carries two further independent changes:
> max_tx is at least the live tx_count. After CPU offline,
[ ... ]
> Guard poll_controller() with adapter->opened so netpoll cannot walk
> unallocated queue state while closed.
After CPU offline, the max_tx change in ibmveth_get_channels() and
ibmveth_set_channels() lets ethtool -l report a max_tx above the online
CPU count. The ibmveth_poll_controller() guard is a separate netpoll
change.
Would splitting these out make bisecting and selective backporting
easier?
> diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
> index 1b1dd89dadf79..da14c69152110 100644
> --- a/drivers/net/ethernet/ibm/ibmveth.c
> +++ b/drivers/net/ethernet/ibm/ibmveth.c
> @@ -3399,15 +3399,32 @@ static void ibmveth_get_channels(struct net_device *netdev,
> struct ethtool_channels *channels)
> {
> struct ibmveth_adapter *adapter = netdev_priv(netdev);
> + unsigned int rx_count = ibmveth_get_num_rx_queues(adapter);
[ ... ]
> + channels->rx_count = rx_count;
> + if (adapter->multi_queue && !adapter->mq_fallback)
> channels->max_rx = IBMVETH_MAX_RX_QUEUES;
> else
> - channels->max_rx = 1;
> - channels->rx_count = ibmveth_get_num_rx_queues(adapter);
> + channels->max_rx = rx_count;
> }
[Severity: Low]
When mq_fallback is set and the live RX count N is above 1, this reports
max_rx = N. The ethtool core only compares rx_count against max_rx:
net/ethtool/channels.c:ethnl_set_channels() {
...
if (channels.rx_count > channels.max_rx)
err_attr = ETHTOOL_A_CHANNELS_RX_COUNT;
...
}
A request for any value in 2..N-1 therefore reaches the driver.
ibmveth_resize_rx_channels() then rejects it:
if (goal_rx > 1 && (!adapter->multi_queue || adapter->mq_fallback)) {
...
return -EOPNOTSUPP;
}
That leaves 1 and N as the only RX counts set_channels accepts, while
get_channels advertises 1..N. For example, ethtool -l can show a max of
8, but ethtool -L dev rx 4 fails with "Operation not supported".
This state can be reached after ibmveth_replenish_task() or a scale-up
register failure sets mq_fallback. It lasts until the next
ibmveth_open() applies the fallback.
The commit message says the cap "blocks growth". Doesn't it also block
shrinking within the advertised range? Could get_channels and
set_channels be made to agree on what is settable?
[ ... ]
> @@ -3482,28 +3493,90 @@ static int ibmveth_set_channels(struct net_device *netdev,
[ ... ]
> + if (!adapter->opened) {
> + /* Apply TX first so a failure leaves the published RX
> + * count unchanged.
> + */
> + rc = netif_set_real_num_tx_queues(netdev, goal_tx);
> + if (rc)
> + return rc;
> +
> + /* Publish desired RX count for next open() and refresh CMO;
> + * do not allocate while down.
> + */
> + if (goal_rx != ibmveth_get_num_rx_queues(adapter)) {
> + ibmveth_publish_num_rx_queues(adapter, goal_rx);
> + rc = netif_set_real_num_rx_queues(netdev, goal_rx);
[Severity: Medium]
Can raising num_rx_queues here leak a stranded subordinate RX queue?
Stranded queues (queue_handle[i] still set above the live count) are
reclaimed only in ibmveth_close(), and only if H_FREE_LOGICAL_LAN
succeeds:
if (!ibmveth_free_all_queues(adapter))
ibmveth_free_stranded_rx_queues(adapter);
That helper also only scans indices at or above the live count:
for (i = ibmveth_get_num_rx_queues(adapter);
i < IBMVETH_MAX_RX_QUEUES; i++) {
if (!adapter->queue_handle[i])
continue;
Consider a live scale-up where H_REG or IRQ setup fails and the
per-queue H_FREE also fails. If the following close then fails
H_FREE_LOGICAL_LAN, the strand survives with opened == false. This new
down path can then lift num_rx_queues above that index, which moves it
out of the scan range.
On the next open, ibmveth_alloc_rx_queues() overwrites the pointers
without checking them:
for (i = 0; i < ibmveth_get_num_rx_queues(adapter); i++) {
adapter->buffer_list_addr[i] =
(void *)get_zeroed_page(GFP_KERNEL);
The same applies to rx_queue[i].queue_addr/queue_dma and
buffer_list_dma[i]. Wouldn't the old page, the coherent RX ring, the DMA
mappings and the pool buffers for that index be leaked?
The live path guards against this in ibmveth_scale_up_rx_queues():
if (adapter->queue_handle[i]) {
/* Left by a failed H_FREE; close frees it. */
netdev_err(netdev,
"RX queue %d still held by PHYP, reset pending\n",
i);
rc = -EBUSY;
Should the down path apply the same queue_handle[] check before
publishing a larger count? Before this patch, an RX count change while
down returned -EOPNOTSUPP, so this path was not reachable.
[ ... ]
> @@ -4203,9 +4296,14 @@ static int ibmveth_change_mtu(struct net_device *dev, int new_mtu)
> static void ibmveth_poll_controller(struct net_device *dev)
> {
> struct ibmveth_adapter *adapter = netdev_priv(dev);
> - unsigned int num = ibmveth_get_num_rx_queues(adapter);
> + unsigned int num;
> int i;
>
> + if (!adapter->opened)
> + return;
[Severity: Low]
This reads adapter->opened as a plain bool from netpoll context. That
context can run concurrently with direct ibmveth_close()/ibmveth_open()
calls from ibmveth_change_mtu() and veth_pool_store(), which don't call
netpoll_poll_disable().
The writers are plain stores as well:
ibmveth_open():
adapter->opened = true;
ibmveth_close():
adapter->opened = false;
Isn't this an unannotated data race that KCSAN would report? Could the
accesses use READ_ONCE()/WRITE_ONCE(), or smp_store_release() and
smp_load_acquire()?
The close side looks fine, because opened is cleared before
synchronize_net() and netpoll runs with IRQs disabled or under RCU-bh.
On the open side, though, nothing explicitly orders the queue, pool and
NAPI setup before opened = true. Today it relies on the replenish_lock
unlock in open and on the acquire of num_rx_queues in the reader. Would
it be worth making that pairing explicit, so that reordering open()
later can't break the guard?
> for (i = 0; i < num; i++)
> ibmveth_replenish_task(adapter, i);
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1790319558.git.mmc%40linux.ibm.com
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support
2026-09-25 18:38 [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support Mingming Cao
` (14 preceding siblings ...)
2026-09-25 18:38 ` [PATCH net-next v7 15/15] ibmveth: Complete set_channels down-path and mq_fallback max_rx cap Mingming Cao
@ 2026-09-26 17:40 ` mingming cao
2026-10-01 22:56 ` Jakub Kicinski
15 siblings, 1 reply; 29+ messages in thread
From: mingming cao @ 2026-09-26 17:40 UTC (permalink / raw)
To: netdev, Jakub Kicinski
Cc: davem, horms, edumazet, pabeni, andrew+netdev, nnac123, maddy,
mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1
Hi Jacub,
I noticed this v7 series is flagged red on patchwork. The apply failure
is due to a dependency on my own [PATCH net 0/2] (Message-ID:
cover.1790357373.git.mmc@linux.ibm.com) sent the same day — kept
separate per your earlier feedback to peel fixes out of feature series.
Both touch ibmveth_open() in the same region, and the net pair is not
yet in net-next.
The v7 MQ series is based on net-next 161ea2d4f2a7 (2026-09-24). The
fixes are already subsumed by MQ patches 3 and 6.
Shall I wait and rebase MQ once the net pair lands in net-next, or send
a v8 now with the net pair folded in? Happy to do either.
Thanks, Mingming
On 9/25/26 11:38 AM, Mingming Cao wrote:
> Hi,
>
> Power11 PHYP adds Virtual Ethernet multi-queue (MQ) RX: multiple
> logical-LAN RX queues, per-queue buffer posting, and completion
> delivery. Guest Linux did not use that; ibmveth still registered one
> RX queue even when PHYP was MQ-capable.
>
> This series adds the ibmveth MQ client for net-next. When PHYP
> advertises IBMVETH_ILLAN_RX_MULTI_QUEUE_SUPPORT via H_ILLAN_ATTRIBUTES,
> probe enables MQ with a default RX count of min(num_online_cpus(), 8)
> (same cap as TX today); ethtool -L can raise RX up to 16. Packets are
> received on per-queue NAPI. Older firmware without the bit is unchanged.
> Queue selection remains firmware-defined (PHYP hash). Ethtool RSS hash
> get/set for that algorithm is deferred to a follow-up series so this
> one stays MQ datapath only.
>
> User-visible bits: ethtool -l/-L (channels); standard per-queue
> packets/bytes/drops via netdev_stat_ops (ethtool -S keeps only
> driver-specific counters; ndo_get_stats64 is the aggregate, including
> retired-queue history); and a read-only debugfs buffer_pools dump
> (v3's multi-line sysfs dump moved to debugfs; the historical queue-0
> poolN/ sysfs ABI is unchanged).
>
> Background:
>
> ibmveth today uses one logical LAN, one set of buffer pools, and one
> NAPI context. PHYP MQ mode gives each RX queue its own handle (post via
> H_ADD_LOGICAL_LAN_BUFFERS_QUEUE, subordinate register via
> H_REG_LOGICAL_LAN_QUEUE); traffic can land on any active queue. The
> driver needs per-queue pools, IRQs, and NAPI to match. Legacy firmware
> keeps the original hcall path.
>
> Series layout (15 patches):
>
> 1-2 Hypercall wrappers; MQ adapter layout (MAX_RX_QUEUES stays 1)
> 3-9 Queue-aware helpers (still SQ runtime): RX, per-queue pools,
> IRQ, TX, PHYP, buffer submit (open/close 3-8); poll harden (9)
> 10 Enable MQ datapath at probe/open (subordinate register helpers
> land here with first use)
> 11-13 Per-queue RX/TX stats; get_channels MQ counts; debugfs buffer_pools
> 14 Incremental RX resize; live ethtool -L rx
> 15 Down-path rollback and mq_fallback max_rx cap
>
> - Helper patches (3-8) reshape ibmveth_open()/close() into
> queue-aware helpers. Patch 9 hardens the SQ poll path with the same
> queue-index helpers; it does not change open/close. MQ stays off
> through 3-9: num_rx_queues stays 1 and multi_queue is false until
> patch 10. The live single-queue path still changes where the review
> required it (open/close unwind, IRQ remask, replenish lock, poll
> harden).
> - Patch 10 is the switch: probe sets multi_queue from firmware, raises
> num_rx_queues, registers subordinates, and replenishes every active
> queue.
> - Patch 11 moves counters per-queue and exports packets/bytes/drops
> through netdev_stat_ops. The thirteen existing -S keys stay; no
> hcall_* or pool%d_ keys.
>
> Testing:
>
> ppc64le PowerVM LPAR, MQ-capable firmware:
> * ethtool -L cycling (16/1/8/11/1/3/16/8/1) with ping - no hangs
> * ethtool -L under iperf3; link down/up during traffic
> * ifdown/ifup under iperf3 RX+TX (MQ and ethtool -L rx 1)
> * Legacy firmware (no MQ bit): open/close/stress on helper path
> * Bisect-safe build and boot at every commit; W=1 clean at tip
>
> Changes in v7:
>
> Same 15 patches as v6. We followed up the v6 netdev-bot review
> with replies; this v7 is the series after that, plus a few items
> from our own re-review.
>
> * Patch 3: update_rx_no_buffer() returns if buffer_list_addr[0] is
> NULL (the per-queue form stays in patch 10).
> * Patch 6: synchronize_net() on the late TX-alloc unwind before RX is
> freed.
> * Patch 8: replenish failure log names the wrapper from the filled
> count; advance ring on NULL buffer so poll does not spin.
> * Patch 9: oversize bound is min(skb_tailroom, pool->buff_size).
> * Patch 10: unregister_netdev before cancel_work_sync and gate reset
> on NETREG_REGISTERED (moved from patch 11); wait for pool kobject
> release before free_netdev(). Drop the probe CMO refresh and the two
> CMO follow-ups: CMO (Power9 and earlier) and MQ firmware (Power11+)
> do not coexist.
> * Patch 11: replenish_lock on the close harvest (remove-path
> unregister/cancel reorder moved to patch 10 with the reset producer).
> * Patch 12: set_channels() returns -EOPNOTSUPP on rx_count changes until
> patch 14 implements live resize.
> * Patch 14: key buffer-list unmap on allocation presence because
> DMA address zero is valid; reject an RX count change while down
> with -EOPNOTSUPP. Failed H_FREE skips unmap and restores the
> surviving count; widen real_num before scale-up unmask. Scale-up
> register -EOPNOTSUPP latches mq_fallback. The scale-up /
> scale-down helper split is code motion only.
> * Patch 15: publish the down-path RX count, which lifts patch 14's
> temporary rejection. get_channels max_tx is at least the live
> tx_count, and set_channels uses the same ceiling, so CPU offline
> cannot block an RX-only ethtool -L.
> * Commit message / kdoc / comment / debug-log updates on 1, 3, 4, 5,
> 7, 8, 9, 10, 11, 12, 14 and 15 (patch 1 also names the new hcalls in the perf
> powerpc-hcalls script and documents H_BUSY on the register-queue
> wrapper; patch 10 prints the register-queue failure with %ld).
> * Kept: enable_irq on schedule_prep failure; mask PHYP before
> napi_disable; get_channels reports the live rx_count (no clamp).
> * Reopen unwind in patches 3 and 6 is pre-existing. No Fixes: tag
> here. The SQ open-fail path is already on the list as
> [PATCH net 0/2] (Message-ID:
> <cover.1790357373.git.mmc@linux.ibm.com>). This series does
> not depend on it. If both land, keep the helper versions in
> patches 3/4/6; the net pair is the current single-queue path
> only.
>
> Known leftovers (not this series):
>
> * Single-queue: replenish vs free_buffer_pool is not serialized,
> irqsave still covers the whole fill, and close skips
> netpoll_poll_disable. That is a lock-protocol rewrite, not this
> series.
> * RX IRQ teardown: teardown masks PHYP, disables NAPI, then masks
> again, but a poll tail that already passed the shutdown checks can
> still re-enable PHYP after that second mask and after free_irq, and
> a mask hcall that failed is never acknowledged. Closing this needs a
> poll/teardown handshake rather than another remask, so the ordering
> is unchanged here.
>
> Changes in v6:
>
> Same 15 patches as v5. Jakub v5 review folded in; per-patch detail is
> below --- on each commit.
>
> * Both new registration wrappers use plpar_hcall(), not plpar_hcall9().
> * Poll: IPv4 check through skb->data; budget 0 does not complete NAPI.
> * Scale-down: publish the surviving count, then synchronize_net(),
> then destroy. num_rx_queues uses smp_store_release / smp_load_acquire.
> * packets/bytes/drops through netdev_stat_ops, not private -S strings.
> Thirteen existing -S keys kept. No hcall_* or pool%d_ keys.
> replenish_* are per-queue u64; no atomics. get_base_stats() is the
> retired-queue remainder.
> * Reset worker gated on NETREG_REGISTERED (cannot reopen after
> unregister).
> * get_channels() keeps the live rx_count; mq_fallback caps max_rx so
> a TX-only ethtool -L is not a silent RX shrink.
>
> Changes in v5:
>
> * Restack mailed v4 (14 patches) to v5 (15):
>
> v4 1-8 helpers -> v5 1-8
> (new) SQ poll harden -> v5 9 (before MQ enable)
> v4 9 MQ enable -> v5 10
> v4 10 stats -> v5 11
> (new) get_channels -> v5 12 (peeled from stats)
> v4 11 debugfs -> v5 13
> v4 12 resize -> v5 14
> v4 13 set_channels -> v5 15
> v4 14 trailing poll/shutdown -> folded into v5 5/9/10/14
> (mailed "P14" was that trailer, not v5 14)
> * Teardown-first resize after aggressive ethtool -L; thin defensive
> poll skip remains; no correlator generation field this series
> * opened / rx_irq_setup; set_channels keys on opened (not IFF_UP)
> * filter_list_dma=0 on map error; restore default-active 64 KiB pool;
> unwind pools by allocation presence; probe_cleanup clears vio
> drvdata; remove: unregister then cancel_work
> * TX quiesce before freeing bounce buffers; guard start_xmit if LTB gone
> * MQ H_FUNCTION recovery (reset + SQ fallback); no printk under
> replenish_lock; lock harvest with replenish; resume kicks all queues
> * Per-queue update_rx_no_buffer; publish-before-free on resize;
> CMO refresh; IRQ helpers return errno
> * Harvest abort (no fake GRO / UAF); poll refuses PHYP re-arm on close;
> wrap-safe skb_put; atomic set_channels; monotonic stats across shrink
> * Keep mask -> sync -> napi_disable on teardown; open stays
> request_irq -> napi_enable while PHYP masked; scale-up/recovery keep
> napi_enable before enable_irq
> * Pool geometry kept on free; restart_rx_queue after open/scale-down;
> remask after napi_disable; schedule_rx_queue masks only when
> napi_schedule_prep succeeds (STOP + poll no-rearm for storms)
>
> Changes in v4:
>
> Addresses Simon's v3 review and related fixes:
> * First-use helpers/includes (irqdomain.h with first dispose); no
> unused statics; dropped orphan open/close pipeline patch
> * Open/close unwind (free LAN before RX pools); no double TX teardown
> * MQ open: replenish all queues before PHYP unmask; H_FUNCTION on
> subordinate register is a hard open failure
> * Resize/set_channels hardenings; stats probe-lifetime + sum-on-read;
> buffer_pools diagnostic on debugfs
> * Patch 9: put already-created pool kobjects on probe failure paths
> * Patch 14: correlator skip, skb tailroom check, napi_complete_done
> shutdown return < budget
> * Bisect-friendly restack (helpers with first use)
>
> Changes in v3:
>
> * Dropped RFC; addressed style / DMA feedback from earlier revisions
> * Early MQ enablement iterations (see lore links below)
>
> Comments welcome.
>
> ---
> v6 lore:
> https://lore.kernel.org/r/cover.1788102125.git.mmc@linux.ibm.com
> Sashiko NIPA (v6):
> https://netdev-ai.bots.linux.dev/sashiko/#/patchset/cover.1788102125.git.mmc@linux.ibm.com
> v5 lore:
> https://lore.kernel.org/r/20260814073642.24630-1-mmc@linux.ibm.com
> v5 review (Jakub Kicinski):
> https://lore.kernel.org/r/20260818014710.3853684-1-kuba@kernel.org
> Sashiko Gemini (sashiko.dev):
> https://sashiko.dev/#/patchset/20260814073642.24630-1-mmc@linux.ibm.com
> Sashiko NIPA (v5):
> https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260814073642.24630-1-mmc@linux.ibm.com
>
> Previous versions
> v6: https://lore.kernel.org/r/cover.1788102125.git.mmc@linux.ibm.com
> v5: https://lore.kernel.org/r/20260814073642.24630-1-mmc@linux.ibm.com
> v4: https://lore.kernel.org/r/cover.1785457143.git.mmc@linux.ibm.com
> v3: https://lore.kernel.org/r/20260706193603.8039-1-mmc@linux.ibm.com
> v2: https://lore.kernel.org/r/20260701222327.61325-1-mmc@linux.ibm.com
> v1: https://lore.kernel.org/r/cover.1782758799.git.mmc@linux.ibm.com
> v4 review (Jakub Kicinski):
> https://lore.kernel.org/r/20260806183614.3171785-1-kuba@kernel.org
> v3 review (Simon Horman):
> https://lore.kernel.org/r/20260714124327.GJ1364329@horms.kernel.org
>
> Mingming Cao (15):
> ibmveth: Add MQ RX hypercall wrappers and call definitions
> ibmveth: Prepare MQ RX adapter data structures
> ibmveth: Refactor RX resource allocation for MQ RX bring-up
> ibmveth: Refactor buffer pool management for per-queue MQ RX
> ibmveth: Refactor RX interrupt control for MQ RX queues
> ibmveth: Refactor TX resource allocation in open/close paths
> ibmveth: Add RX queue register helpers for MQ
> ibmveth: Add queue-aware RX buffer submit helper for MQ
> ibmveth: Harden RX poll path with helpers
> ibmveth: Enable multi-queue RX receive path
> ibmveth: Add per-queue RX and TX statistics collection
> ibmveth: Report MQ-aware RX counts in ethtool get_channels
> ibmveth: Expose per-queue buffer pool details via debugfs
> ibmveth: Implement incremental MQ RX queue resize
> ibmveth: Complete set_channels down-path and mq_fallback max_rx cap
>
> arch/powerpc/include/asm/hvcall.h | 6 +-
> drivers/net/ethernet/ibm/ibmveth.c | 4441 +++++++++++++++----
> drivers/net/ethernet/ibm/ibmveth.h | 232 +-
> tools/perf/scripts/python/powerpc-hcalls.py | 4 +
> 4 files changed, 3921 insertions(+), 762 deletions(-)
>
>
> base-commit: 161ea2d4f2a7e784f14b5b0548fcef3e05fc34f8
^ permalink raw reply [flat|nested] 29+ messages in thread* Re: [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support
2026-09-26 17:40 ` [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support mingming cao
@ 2026-10-01 22:56 ` Jakub Kicinski
2026-10-03 2:11 ` mingming cao
0 siblings, 1 reply; 29+ messages in thread
From: Jakub Kicinski @ 2026-10-01 22:56 UTC (permalink / raw)
To: mingming cao
Cc: netdev, davem, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1
On Sat, 26 Sep 2026 10:40:56 -0700 mingming cao wrote:
> Hi Jacub,
There's no c in my name, sir. I wonder why people do that.
> I noticed this v7 series is flagged red on patchwork. The apply failure
> is due to a dependency on my own [PATCH net 0/2] (Message-ID:
> cover.1790357373.git.mmc@linux.ibm.com) sent the same day — kept
> separate per your earlier feedback to peel fixes out of feature series.
> Both touch ibmveth_open() in the same region, and the net pair is not
> yet in net-next.
>
> The v7 MQ series is based on net-next 161ea2d4f2a7 (2026-09-24). The
> fixes are already subsumed by MQ patches 3 and 6.
>
> Shall I wait and rebase MQ once the net pair lands in net-next, or send
> a v8 now with the net pair folded in? Happy to do either.
Sorry for the delay, hint: please put just me in the To: if you want me
to spot the email sooner. We get 400+ messages to the list every day,
I'm only reading the list as I review :(
To answer your question - please hold off with reposting until the fixes
are in net-next. We won't merge the patches if they conflict with fixes.
^ permalink raw reply [flat|nested] 29+ messages in thread
* Re: [PATCH net-next v7 00/15] ibmveth: Add multi-queue RX support
2026-10-01 22:56 ` Jakub Kicinski
@ 2026-10-03 2:11 ` mingming cao
0 siblings, 0 replies; 29+ messages in thread
From: mingming cao @ 2026-10-03 2:11 UTC (permalink / raw)
To: Jakub Kicinski
Cc: netdev, davem, horms, edumazet, pabeni, andrew+netdev, nnac123,
maddy, mpe, linuxppc-dev, haren, ricklind, davemarq, bjking1,
shaik.abdulla1
On 10/1/26 3:56 PM, Jakub Kicinski wrote:
> On Sat, 26 Sep 2026 10:40:56 -0700 mingming cao wrote:
>> Hi Jacub,
> There's no c in my name, sir. I wonder why people do that.
Sorry about that, Jakub. Noted, and noted on To: as well.
>> I noticed this v7 series is flagged red on patchwork. The apply failure
>> is due to a dependency on my own [PATCH net 0/2] (Message-ID:
>> cover.1790357373.git.mmc@linux.ibm.com) sent the same day — kept
>> separate per your earlier feedback to peel fixes out of feature series.
>> Both touch ibmveth_open() in the same region, and the net pair is not
>> yet in net-next.
>>
>> The v7 MQ series is based on net-next 161ea2d4f2a7 (2026-09-24). The
>> fixes are already subsumed by MQ patches 3 and 6.
>>
>> Shall I wait and rebase MQ once the net pair lands in net-next, or send
>> a v8 now with the net pair folded in? Happy to do either.
> Sorry for the delay, hint: please put just me in the To: if you want me
> to spot the email sooner. We get 400+ messages to the list every day,
> I'm only reading the list as I review :(
>
> To answer your question - please hold off with reposting until the fixes
> are in net-next. We won't merge the patches if they conflict with fixes.
Will do. Thanks for applying the net pair (af0524bf4ce1, 84bec0bf0352).
I will not repost the multi-queue series until they show up in net-next.
A heads-up so the next series is not a surprise: I posted seven more
ibmveth fixes, split out of this series:
[PATCH net-next 0/7] ibmveth: fix hangs, use-after-frees and
netpoll races
https://lore.kernel.org/netdev/cover.1790991039.git.mmc@linux.ibm.com/
<https://lore.kernel.org/netdev/cover.1790991039.git.mmc@linux.ibm.com/>
They fix serious pre-existing bugs in the driver (two hangs, three
use-after-frees and memory corruption), found by AI-assisted review
of this series. They are based on net-next and merge cleanly with
the net pair; they also apply to net if you would rather take them
there.
I am working on v8. My thinking for v8 of the multi-queue series is to
base it on
net-next once both the net pair and these fixes are in. With the
fixes split out it should be smaller than v7, with the MQ patches
mostly extending that handling per queue, and it would address the
review comments on v7. Hope that works for you. Happy to adjust as needed.
Regards,
Mingming
^ permalink raw reply [flat|nested] 29+ messages in thread