All of lore.kernel.org
 help / color / mirror / Atom feed
From: Mingming Cao <mmc@linux.ibm.com>
To: netdev@vger.kernel.org
Cc: davem@davemloft.net, kuba@kernel.org, edumazet@google.com,
	pabeni@redhat.com, andrew+netdev@lunn.ch, haren@linux.ibm.com,
	ricklind@linux.ibm.com, nnac123@linux.ibm.com,
	davemarq@linux.ibm.com, vaishnavi@linux.ibm.com,
	bjking1@linux.ibm.com, linuxppc-dev@lists.ozlabs.org,
	mmc@linux.ibm.com
Subject: [PATCH net-next v1 1/6] ibmvnic: cap rx pool entries against the real buffer size
Date: Wed,  5 Aug 2026 15:43:56 -0700	[thread overview]
Message-ID: <20260805224401.58791-2-mmc@linux.ibm.com> (raw)
In-Reply-To: <20260805224401.58791-1-mmc@linux.ibm.com>

Changing the mtu from 1500 to a value in the 2000..8000 range can fail
pool allocation and leave the interface down:

  ibmvnic 30000008: Couldn't alloc long term buffer
  __alloc_pages: ... order:N, mode:0x...

(9000 happens to succeed, which is how the middle of the range stood out.)

That is the non-bucket case this series exists to allow: the guest mtu is
not one of the sizes the backing buffers were sized for. Until the next
patch, req_mtu was always forced back to a size the device carries, so
the pool guard happened to measure the right buffer. Once arbitrary
in-range mtus are honoured, that luck goes away.

send_request_cap() limits how many entries a pool may hold so the whole
pool fits in one long term buffer set:

  max_entries = IBMVNIC_LTB_SET_SIZE /
                (adapter->req_mtu + IBMVNIC_BUFFER_HLEN);

The rx buffers are not sized from req_mtu. init_rx_pools() sizes them
from cur_rx_buf_sz, which comes from the login response and describes
the backing device. When the VIOS has the device at jumbo (cur_rx_buf_sz
around 9014) but the guest mtu is still smaller, the guard measures a
buffer much smaller than the one really allocated. The entry count
passes through untouched and the pool asks for several times the space
the guard assumed - enough that dma_alloc_coherent() / __alloc_pages()
refuses the request on kernels whose per-pool LTB budget is a single
allocation rather than a large set.

With a backing size of 9014 the buffers are ALIGN(9014, L1_CACHE_BYTES)
= 9088 bytes, while an mtu of 2000 leaves the guard working from 2514.
The overshoot only shrinks as req_mtu approaches the backing size, so
the largest mtus are the safe ones and everything below drifts.

Cap the count against the size the buffers are really allocated at.
reuse_rx_pools() has to compare against the same number or it would
see a difference on every reset and reallocate pools that are already
the right shape. clean_rx_pools() must walk rx_pool->size for the same
reason, or a clamped pool would be cleaned past its end.

Comes ahead of the next patch, which is what lets req_mtu sit below
the backing size. Keep the two together when backporting.

Fixes: c26eba03e407 ("ibmvnic: Update reset infrastructure to support tunable parameters")
Reviewed-by: Dave Marquardt <davemarq@linux.ibm.com>
Tested-by: Vaishnavi Bhat <vaishnavi@linux.ibm.com>
Signed-off-by: Mingming Cao <mmc@linux.ibm.com>
---
 drivers/net/ethernet/ibm/ibmvnic.c | 26 +++++++++++++++++++++-----
 1 file changed, 21 insertions(+), 5 deletions(-)

diff --git a/drivers/net/ethernet/ibm/ibmvnic.c b/drivers/net/ethernet/ibm/ibmvnic.c
index 5a510eed335e..86e643ee6b3b 100644
--- a/drivers/net/ethernet/ibm/ibmvnic.c
+++ b/drivers/net/ethernet/ibm/ibmvnic.c
@@ -1021,6 +1021,24 @@ static void release_rx_pools(struct ibmvnic_adapter *adapter)
 	adapter->prev_rx_pool_size = 0;
 }
 
+/**
+ * rx_pool_entries() - Number of buffers one rx pool may hold
+ * @adapter: ibmvnic adapter
+ *
+ * send_request_cap() budgets the entry count with req_mtu, but the
+ * buffers are sized from cur_rx_buf_sz. Cap against that size here.
+ */
+static u64 rx_pool_entries(struct ibmvnic_adapter *adapter)
+{
+	u64 buff_size = ALIGN(adapter->cur_rx_buf_sz, L1_CACHE_BYTES);
+
+	if (!buff_size)
+		return adapter->req_rx_add_entries_per_subcrq;
+
+	return min_t(u64, adapter->req_rx_add_entries_per_subcrq,
+		     IBMVNIC_LTB_SET_SIZE / buff_size);
+}
+
 /**
  * reuse_rx_pools() - Check if the existing rx pools can be reused.
  * @adapter: ibmvnic adapter
@@ -1048,7 +1066,7 @@ static bool reuse_rx_pools(struct ibmvnic_adapter *adapter)
 	new_num_pools = adapter->req_rx_queues;
 
 	old_pool_size = adapter->prev_rx_pool_size;
-	new_pool_size = adapter->req_rx_add_entries_per_subcrq;
+	new_pool_size = rx_pool_entries(adapter);
 
 	old_buff_size = adapter->prev_rx_buf_sz;
 	new_buff_size = adapter->cur_rx_buf_sz;
@@ -1082,7 +1100,7 @@ static int init_rx_pools(struct net_device *netdev)
 	u64 buff_size;
 	int i, j, rc;
 
-	pool_size = adapter->req_rx_add_entries_per_subcrq;
+	pool_size = rx_pool_entries(adapter);
 	num_pools = adapter->req_rx_queues;
 	buff_size = adapter->cur_rx_buf_sz;
 
@@ -2000,7 +2018,6 @@ static void clean_rx_pools(struct ibmvnic_adapter *adapter)
 {
 	struct ibmvnic_rx_pool *rx_pool;
 	struct ibmvnic_rx_buff *rx_buff;
-	u64 rx_entries;
 	int rx_scrqs;
 	int i, j;
 
@@ -2008,7 +2025,6 @@ static void clean_rx_pools(struct ibmvnic_adapter *adapter)
 		return;
 
 	rx_scrqs = adapter->num_active_rx_pools;
-	rx_entries = adapter->req_rx_add_entries_per_subcrq;
 
 	/* Free any remaining skbs in the rx buffer pools */
 	for (i = 0; i < rx_scrqs; i++) {
@@ -2017,7 +2033,7 @@ static void clean_rx_pools(struct ibmvnic_adapter *adapter)
 			continue;
 
 		netdev_dbg(adapter->netdev, "Cleaning rx_pool[%d]\n", i);
-		for (j = 0; j < rx_entries; j++) {
+		for (j = 0; j < rx_pool->size; j++) {
 			rx_buff = &rx_pool->rx_buff[j];
 			if (rx_buff && rx_buff->skb) {
 				dev_kfree_skb_any(rx_buff->skb);
-- 
2.50.1 (Apple Git-155)



  reply	other threads:[~2026-08-05 22:44 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-05 22:43 [PATCH net-next v1 0/6] ibmvnic: honour flexible in-range MTU values Mingming Cao
2026-08-05 22:43 ` Mingming Cao [this message]
2026-08-05 22:43 ` [PATCH net-next v1 2/6] ibmvnic: honour the requested mtu instead of reverting to a fallback Mingming Cao
2026-08-05 22:43 ` [PATCH net-next v1 3/6] ibmvnic: do not unmap long term buffers across a crq reconnect Mingming Cao
2026-08-05 22:43 ` [PATCH net-next v1 4/6] ibmvnic: allocate new buffer pools before releasing the old ones Mingming Cao
2026-08-05 22:44 ` [PATCH net-next v1 5/6] ibmvnic: allocate a new long term buffer before freeing the old one Mingming Cao
2026-08-05 22:44 ` [PATCH net-next v1 6/6] ibmvnic: change the mtu without a reset where the buffers allow it Mingming Cao
2026-08-06  2:38 ` [PATCH net-next v1 0/6] ibmvnic: honour flexible in-range MTU values Jakub Kicinski
2026-08-06  5:28   ` mingming cao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260805224401.58791-2-mmc@linux.ibm.com \
    --to=mmc@linux.ibm.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=bjking1@linux.ibm.com \
    --cc=davem@davemloft.net \
    --cc=davemarq@linux.ibm.com \
    --cc=edumazet@google.com \
    --cc=haren@linux.ibm.com \
    --cc=kuba@kernel.org \
    --cc=linuxppc-dev@lists.ozlabs.org \
    --cc=netdev@vger.kernel.org \
    --cc=nnac123@linux.ibm.com \
    --cc=pabeni@redhat.com \
    --cc=ricklind@linux.ibm.com \
    --cc=vaishnavi@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.