* [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations
@ 2026-09-11 15:45 Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
Replace devmem's gen_pool based fixed-size allocator with a binding-level
freelist similar to the one used by io_uring zero-copy receive.
This is motivated by allocation latency observed in the NAPI receive path:
[ 1036.228913] ? gen_pool_create+0x90/0x90
[ 1036.228915] net_devmem_alloc_dmabuf+0x1f/0x60
[ 1036.228918] mp_dmabuf_devmem_alloc_netmems+0x17/0x80
[ 1036.228920] mlx5e_post_rx_mpwqes+0xdbe/0xdd0
[ 1036.228926] mlx5e_napi_poll+0x113/0x830
[ 1036.228928] ? sched_clock+0x5/0x10
[ 1036.228931] ? wake_up_process+0x778/0x14b0
[ 1036.228933] net_rx_action+0x15d/0x570
[ 1036.228934] ? update_rq_clock+0x31/0x240
[ 1036.228937] ? __napi_schedule+0x55/0xa0
[ 1036.228938] ? mlx5_eq_comp_int+0x137/0x230
[ 1036.228940] ? atomic_notifier_call_chain+0x36/0x90
[ 1036.228943] ? sched_clock+0x5/0x10
[ 1036.228944] ? sched_clock_cpu+0xc/0x170
[ 1036.228947] irq_exit_rcu+0x12b/0x370
[ 1036.228950] common_interrupt+0x85/0x90
udmabuf can create a very large number of SG entries. In the worst case,
devmem ends up adding one gen_pool chunk for each net_iov allocation
unit backed by those entries. The gen_pool allocation path then has to
traverse a linked list that can become too long for this hot path.
Patch 1 removes the gen_pool and replaces it with a simple freelist of
net_iov pointers protected by the same spin_lock_bh() pattern used by
io_uring zcrx. Patch 2 removes the now-unnecessary chunk owner wrapper by
embedding the net_iov_area directly in the dma-buf binding. Patch 3
batches freelist allocations.
= Performance:
kperf/client ... \
--num-rx-queues 4 \
--dmabuf-rx-size-mb 2048 \
--dmabuf-tx-size-mb 2048 \
--validate no \
--time 60 \
--read-size 67108864 \
--write-size 67108864 \
--num-connections 4 \
--tcp-cc dctcp \
--pin-off 4 \
--devmem-rx \
--devmem-tx \
--devmem-rx-memory cuda \
--devmem-tx-memory cuda
With 4 queues, 4 flows, 2GB BB, cuda for both rx and tx I see no difference
in throughput or cpu utilization (see selective runs below).
== Before
10 runs: 206.031 243.546 293.931 319.015 319.923 323.189 324.321 325.499 327.134 349.450 Gbps
Sample:
client: == Source <redacted>
client: Tx 48.170 Gbps (361716776960 bytes in 60072872 usec)
client: Tx101.256 Gbps (760343429120 bytes in 60072872 usec)
client: Tx101.077 Gbps (759001251840 bytes in 60072872 usec)
client: Tx 69.440 Gbps (521435873280 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: Rx 0.000 Gbps (0 bytes in 60072872 usec)
client: == Target <redacted>
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Tx 0.000 Gbps (0 bytes in 60074846 usec)
client: Rx 48.158 Gbps (361638901920 bytes in 60074846 usec)
client: Rx101.253 Gbps (760343429120 bytes in 60074846 usec)
client: Rx101.074 Gbps (759001251840 bytes in 60074846 usec)
client: Rx 69.438 Gbps (521435873280 bytes in 60074846 usec)
client: net CPU 1: usr: 0.00% sys: 0.01% idle:39.42% iow: 0.00% irq: 0.64% sirq:59.91%
client: app CPU 5: usr: 1.21% sys:97.71% idle: 0.09% iow: 0.00% irq: 0.24% sirq: 0.71%
client: net CPU 2: usr: 0.00% sys: 0.00% idle:38.19% iow: 0.00% irq: 0.84% sirq:60.96%
client: app CPU 6: usr: 1.24% sys:97.73% idle: 0.04% iow: 0.00% irq: 0.24% sirq: 0.71%
client: net CPU 0: usr: 0.05% sys: 0.27% idle:69.92% iow: 0.00% irq: 7.71% sirq:22.04%
client: app CPU 4: usr: 1.78% sys:80.71% idle:16.79% iow: 0.00% irq: 0.38% sirq: 0.32%
client: net CPU 3: usr: 0.00% sys: 0.00% idle: 0.00% iow: 0.00% irq: 0.39% sirq:99.60%
client: app CPU 7: usr: 0.21% sys: 6.64% idle:90.75% iow: 0.00% irq: 0.05% sirq: 2.32%
== After
10 runs: 213.032 226.471 235.052 257.113 323.233 331.597 342.893 348.188 351.866 354.311 Gbps
Sample:
client: == Source <redacted>
client: Tx104.829 Gbps (786515886080 bytes in 60022934 usec)
client: Tx105.455 Gbps (791213506560 bytes in 60022934 usec)
client: Tx 61.717 Gbps (463051161600 bytes in 60022934 usec)
client: Tx 51.430 Gbps (385875968000 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: Rx 0.000 Gbps (0 bytes in 60022934 usec)
client: == Target <redacted>
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Tx 0.000 Gbps (0 bytes in 60058283 usec)
client: Rx104.767 Gbps (786515886080 bytes in 60058283 usec)
client: Rx105.393 Gbps (791213506560 bytes in 60058283 usec)
client: Rx 61.673 Gbps (462998508768 bytes in 60058283 usec)
client: Rx 51.400 Gbps (385875968000 bytes in 60058283 usec)
client: net CPU 0: usr: 0.01% sys: 0.21% idle:65.30% iow: 0.00% irq: 9.15% sirq:25.29%
client: app CPU 4: usr: 1.19% sys:98.00% idle: 0.03% iow: 0.00% irq: 0.21% sirq: 0.54%
client: net CPU 2: usr: 0.00% sys: 0.00% idle: 0.03% iow: 0.00% irq: 0.44% sirq:99.51%
client: app CPU 6: usr: 2.14% sys:77.48% idle:19.92% iow: 0.00% irq: 0.36% sirq: 0.07%
client: net CPU 3: usr: 0.00% sys: 0.00% idle:43.01% iow: 0.00% irq: 0.69% sirq:56.28%
client: app CPU 7: usr: 2.03% sys:71.04% idle:26.45% iow: 0.00% irq: 0.37% sirq: 0.08%
client: net CPU 1: usr: 0.00% sys: 0.01% idle:46.66% iow: 0.00% irq: 0.56% sirq:52.75%
client: app CPU 5: usr: 0.05% sys: 0.08% idle:99.21% iow: 0.00% irq: 0.03% sirq: 0.61%
== Comparison, over 10 runs
Median Target RX: 321.556 Gbps vs 327.415 Gbps
Mean Target RX: 303.204 Gbps vs 298.376 Gbps
Range: 206.031-349.450 Gbps vs 213.032-354.311 Gbps
v2:
- xmas tree (Jakub)
- batching (Mina)
- perf numbers (Mina & Jakub)
Stanislav Fomichev (3):
net: devmem: replace gen_pool with freelist
net: devmem: embed net_iov_area in binding
net: devmem: batch net_iov allocations into the page_pool cache
net/Kconfig | 1 -
net/core/devmem.c | 208 +++++++++++++++++++++-------------------------
net/core/devmem.h | 46 +++-------
3 files changed, 105 insertions(+), 150 deletions(-)
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist
2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev
@ 2026-09-11 15:45 ` Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev
2 siblings, 0 replies; 4+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
devmem only needs fixed-size net_iov allocations for each dma-buf binding.
The gen_pool tracks the same free set indirectly through DMA addresses,
which makes devmem depend on the generic allocator even though the users
are fixed-size net_iov chunks.
Mirror the io_uring zcrx model more closely by keeping a binding-level
freelist protected by spin_lock_bh(). Use a single net_iov_area owner for
the binding, populate each net_iov's DMA address while walking the SG
table, and check at teardown that all net_iovs have returned to the
freelist.
Drop the NET_DEVMEM select of GENERIC_ALLOCATOR now that devmem no longer
calls gen_pool APIs.
Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
---
net/Kconfig | 1 -
net/core/devmem.c | 172 +++++++++++++++++++++-------------------------
net/core/devmem.h | 16 ++---
3 files changed, 85 insertions(+), 104 deletions(-)
diff --git a/net/Kconfig b/net/Kconfig
index e38477393551..76ab44aa439a 100644
--- a/net/Kconfig
+++ b/net/Kconfig
@@ -68,7 +68,6 @@ config SKB_EXTENSIONS
config NET_DEVMEM
def_bool y
- select GENERIC_ALLOCATOR
depends on DMA_SHARED_BUFFER
depends on PAGE_POOL
diff --git a/net/core/devmem.c b/net/core/devmem.c
index f4d60654ce7f..4883eb7f3a95 100644
--- a/net/core/devmem.c
+++ b/net/core/devmem.c
@@ -8,7 +8,6 @@
*/
#include <linux/dma-buf.h>
-#include <linux/genalloc.h>
#include <linux/mm.h>
#include <linux/netdevice.h>
#include <linux/types.h>
@@ -30,23 +29,13 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1);
static const struct memory_provider_ops dmabuf_devmem_ops;
-static void net_devmem_dmabuf_free_chunk_owner(struct gen_pool *genpool,
- struct gen_pool_chunk *chunk,
- void *not_used)
+static void
+net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner)
{
- struct dmabuf_genpool_chunk_owner *owner = chunk->owner;
-
- kvfree(owner->area.niovs);
- kfree(owner);
-}
-
-static dma_addr_t net_devmem_get_dma_addr(const struct net_iov *niov)
-{
- struct dmabuf_genpool_chunk_owner *owner;
-
- owner = net_devmem_iov_to_chunk_owner(niov);
- return owner->base_dma_addr +
- ((dma_addr_t)net_iov_idx(niov) << owner->binding->niov_shift);
+ if (owner) {
+ kvfree(owner->area.niovs);
+ kfree(owner);
+ }
}
static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref)
@@ -62,24 +51,18 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
{
struct net_devmem_dmabuf_binding *binding = container_of(wq, typeof(*binding), unbind_w);
- size_t size, avail;
-
- gen_pool_for_each_chunk(binding->chunk_pool,
- net_devmem_dmabuf_free_chunk_owner, NULL);
-
- size = gen_pool_size(binding->chunk_pool);
- avail = gen_pool_avail(binding->chunk_pool);
-
- if (!WARN(size != avail, "can't destroy genpool. size=%zu, avail=%zu",
- size, avail))
- gen_pool_destroy(binding->chunk_pool);
+ WARN(binding->free_count != binding->total_niovs,
+ "can't destroy dmabuf binding. total=%zu, free=%zu",
+ binding->total_niovs, binding->free_count);
+ net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt,
binding->direction);
dma_buf_detach(binding->dmabuf, binding->attachment);
dma_buf_put(binding->dmabuf);
xa_destroy(&binding->bound_rxqs);
percpu_ref_exit(&binding->ref);
+ kvfree(binding->freelist);
kvfree(binding->tx_vec);
kfree(binding);
}
@@ -87,21 +70,16 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
struct net_iov *
net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
{
- struct dmabuf_genpool_chunk_owner *owner;
- unsigned long dma_addr;
struct net_iov *niov;
- ssize_t offset;
- ssize_t index;
-
- dma_addr = gen_pool_alloc_owner(binding->chunk_pool,
- 1UL << binding->niov_shift,
- (void **)&owner);
- if (!dma_addr)
+ spin_lock_bh(&binding->freelist_lock);
+ if (unlikely(!binding->free_count)) {
+ spin_unlock_bh(&binding->freelist_lock);
return NULL;
+ }
- offset = dma_addr - owner->base_dma_addr;
- index = offset >> binding->niov_shift;
- niov = &owner->area.niovs[index];
+ niov = binding->freelist[--binding->free_count];
+ binding->freelist[binding->free_count] = NULL;
+ spin_unlock_bh(&binding->freelist_lock);
niov->desc.pp_magic = 0;
niov->desc.pp = NULL;
@@ -113,14 +91,15 @@ net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
void net_devmem_free_dmabuf(struct net_iov *niov)
{
struct net_devmem_dmabuf_binding *binding = net_devmem_iov_binding(niov);
- unsigned long dma_addr = net_devmem_get_dma_addr(niov);
- size_t niov_size = 1UL << binding->niov_shift;
- if (WARN_ON(!gen_pool_has_addr(binding->chunk_pool, dma_addr,
- niov_size)))
+ spin_lock_bh(&binding->freelist_lock);
+ if (WARN_ON_ONCE(binding->free_count >= binding->total_niovs)) {
+ spin_unlock_bh(&binding->freelist_lock);
return;
+ }
- gen_pool_free(binding->chunk_pool, dma_addr, niov_size);
+ binding->freelist[binding->free_count++] = niov;
+ spin_unlock_bh(&binding->freelist_lock);
}
void net_devmem_unbind_dmabuf(struct net_devmem_dmabuf_binding *binding)
@@ -194,12 +173,15 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
struct netlink_ext_ack *extack)
{
struct net_devmem_dmabuf_binding *binding;
+ struct dmabuf_genpool_chunk_owner *owner;
size_t niov_size = 1UL << niov_shift;
static u32 id_alloc_next;
struct scatterlist *sg;
struct dma_buf *dmabuf;
- unsigned int sg_idx, i;
- unsigned long virtual;
+ unsigned int sg_idx;
+ size_t total_niovs;
+ size_t niov_idx;
+ size_t i;
int err;
if (!dma_dev) {
@@ -230,6 +212,7 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
goto err_free_binding;
mutex_init(&binding->lock);
+ spin_lock_init(&binding->freelist_lock);
binding->dmabuf = dmabuf;
binding->direction = direction;
@@ -262,20 +245,10 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
goto err_unmap;
}
}
-
- binding->chunk_pool = gen_pool_create(niov_shift,
- dev_to_node(&dev->dev));
- if (!binding->chunk_pool) {
- err = -ENOMEM;
- goto err_tx_vec;
- }
-
- virtual = 0;
+ total_niovs = 0;
for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
dma_addr_t dma_addr = sg_dma_address(sg);
- struct dmabuf_genpool_chunk_owner *owner;
size_t len = sg_dma_len(sg);
- struct net_iov *niov;
if (!IS_ALIGNED(dma_addr, niov_size) ||
!IS_ALIGNED(len, niov_size)) {
@@ -283,63 +256,74 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
NL_SET_ERR_MSG_FMT(extack,
"dmabuf sg entry (addr=%pad, len=%zu) not aligned to niov size %zu",
&dma_addr, len, niov_size);
- goto err_free_chunks;
+ goto err_tx_vec;
}
- owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
- dev_to_node(&dev->dev));
- if (!owner) {
- err = -ENOMEM;
- goto err_free_chunks;
- }
+ total_niovs += len >> niov_shift;
+ }
- owner->area.base_virtual = virtual;
- owner->base_dma_addr = dma_addr;
- owner->area.num_niovs = len >> niov_shift;
- owner->binding = binding;
+ binding->freelist = kvmalloc_array(total_niovs,
+ sizeof(binding->freelist[0]),
+ GFP_KERNEL);
+ if (!binding->freelist) {
+ err = -ENOMEM;
+ goto err_tx_vec;
+ }
+ binding->total_niovs = total_niovs;
- err = gen_pool_add_owner(binding->chunk_pool, dma_addr,
- dma_addr, len, dev_to_node(&dev->dev),
- owner);
- if (err) {
- kfree(owner);
- err = -EINVAL;
- goto err_free_chunks;
- }
+ owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
+ dev_to_node(&dev->dev));
+ if (!owner) {
+ err = -ENOMEM;
+ goto err_free_freelist;
+ }
- owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
- owner->area.num_niovs);
- if (!owner->area.niovs) {
- err = -ENOMEM;
- goto err_free_chunks;
- }
+ owner->area.num_niovs = total_niovs;
+ owner->binding = binding;
+ owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
+ owner->area.num_niovs);
+ if (!owner->area.niovs) {
+ err = -ENOMEM;
+ goto err_free_owner;
+ }
+ binding->chunk_owner = owner;
+
+ niov_idx = 0;
+ for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
+ dma_addr_t dma_addr = sg_dma_address(sg);
+ size_t len = sg_dma_len(sg);
+ struct net_iov *niov;
+ size_t nr_niovs = len >> niov_shift;
- for (i = 0; i < owner->area.num_niovs; i++) {
- niov = &owner->area.niovs[i];
+ for (i = 0; i < nr_niovs; i++, niov_idx++) {
+ niov = &owner->area.niovs[niov_idx];
net_iov_init(niov, &owner->area, NET_IOV_DMABUF);
page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov),
- net_devmem_get_dma_addr(niov));
+ dma_addr);
if (direction == DMA_TO_DEVICE)
- binding->tx_vec[owner->area.base_virtual / PAGE_SIZE + i] = niov;
+ binding->tx_vec[niov_idx] = niov;
+ binding->freelist[binding->free_count++] = niov;
+ dma_addr += niov_size;
}
-
- virtual += len;
}
err = xa_alloc_cyclic(&net_devmem_dmabuf_bindings, &binding->id,
binding, xa_limit_32b, &id_alloc_next,
GFP_KERNEL);
if (err < 0)
- goto err_free_chunks;
+ goto err_free_chunk_owner;
list_add(&binding->list, &priv->bindings);
return binding;
-err_free_chunks:
- gen_pool_for_each_chunk(binding->chunk_pool,
- net_devmem_dmabuf_free_chunk_owner, NULL);
- gen_pool_destroy(binding->chunk_pool);
+err_free_chunk_owner:
+ net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
+ goto err_free_freelist;
+err_free_owner:
+ kfree(owner);
+err_free_freelist:
+ kvfree(binding->freelist);
err_tx_vec:
kvfree(binding->tx_vec);
err_unmap:
diff --git a/net/core/devmem.h b/net/core/devmem.h
index 4a293a7d1149..a5ee2d8d9169 100644
--- a/net/core/devmem.h
+++ b/net/core/devmem.h
@@ -14,6 +14,7 @@
#include <net/netdev_netlink.h>
struct netlink_ext_ack;
+struct dmabuf_genpool_chunk_owner;
struct net_devmem_dmabuf_binding {
struct dma_buf *dmabuf;
@@ -26,7 +27,7 @@ struct net_devmem_dmabuf_binding {
* dereferenced.
*/
void *vdev;
- struct gen_pool *chunk_pool;
+ struct dmabuf_genpool_chunk_owner *chunk_owner;
/* Protect dev */
struct mutex lock;
@@ -57,6 +58,11 @@ struct net_devmem_dmabuf_binding {
/* rxq's this binding is active on. */
struct xarray bound_rxqs;
+ spinlock_t freelist_lock ____cacheline_aligned_in_smp;
+ size_t free_count;
+ size_t total_niovs;
+ struct net_iov **freelist;
+
/* ID of this binding. Globally unique to all bindings currently
* active.
*/
@@ -77,17 +83,9 @@ struct net_devmem_dmabuf_binding {
};
#if defined(CONFIG_NET_DEVMEM)
-/* Owner of the dma-buf chunks inserted into the gen pool. Each scatterlist
- * entry from the dmabuf is inserted into the genpool as a chunk, and needs
- * this owner struct to keep track of some metadata necessary to create
- * allocations from this chunk.
- */
struct dmabuf_genpool_chunk_owner {
struct net_iov_area area;
struct net_devmem_dmabuf_binding *binding;
-
- /* dma_addr of the start of the chunk. */
- dma_addr_t base_dma_addr;
};
void __net_devmem_dmabuf_binding_free(struct work_struct *wq);
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 4+ messages in thread
* [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding
2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
@ 2026-09-11 15:45 ` Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev
2 siblings, 0 replies; 4+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
After replacing the gen_pool with a binding-level freelist, devmem no
longer needs a separate chunk owner object. There is only one
net_iov_area for the binding, so store it directly in struct
net_devmem_dmabuf_binding.
Derive the binding from net_iov_owner() with container_of(), matching the
pattern used by io_uring zcrx. This removes the leftover
dmabuf_genpool_chunk_owner wrapper and its allocation/free path.
Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
---
net/core/devmem.c | 41 ++++++++++-------------------------------
net/core/devmem.h | 26 +++++++-------------------
2 files changed, 17 insertions(+), 50 deletions(-)
diff --git a/net/core/devmem.c b/net/core/devmem.c
index 4883eb7f3a95..7949f8425bcd 100644
--- a/net/core/devmem.c
+++ b/net/core/devmem.c
@@ -29,15 +29,6 @@ static DEFINE_XARRAY_FLAGS(net_devmem_dmabuf_bindings, XA_FLAGS_ALLOC1);
static const struct memory_provider_ops dmabuf_devmem_ops;
-static void
-net_devmem_dmabuf_free_chunk_owner(struct dmabuf_genpool_chunk_owner *owner)
-{
- if (owner) {
- kvfree(owner->area.niovs);
- kfree(owner);
- }
-}
-
static void net_devmem_dmabuf_binding_release(struct percpu_ref *ref)
{
struct net_devmem_dmabuf_binding *binding =
@@ -55,7 +46,7 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
"can't destroy dmabuf binding. total=%zu, free=%zu",
binding->total_niovs, binding->free_count);
- net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
+ kvfree(binding->area.niovs);
dma_buf_unmap_attachment_unlocked(binding->attachment, binding->sgt,
binding->direction);
dma_buf_detach(binding->dmabuf, binding->attachment);
@@ -271,23 +262,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
}
binding->total_niovs = total_niovs;
- owner = kzalloc_node(sizeof(*owner), GFP_KERNEL,
- dev_to_node(&dev->dev));
- if (!owner) {
+ binding->area.num_niovs = total_niovs;
+ binding->area.niovs = kvmalloc_objs(*binding->area.niovs,
+ binding->area.num_niovs);
+ if (!binding->area.niovs) {
err = -ENOMEM;
goto err_free_freelist;
}
- owner->area.num_niovs = total_niovs;
- owner->binding = binding;
- owner->area.niovs = kvmalloc_objs(*owner->area.niovs,
- owner->area.num_niovs);
- if (!owner->area.niovs) {
- err = -ENOMEM;
- goto err_free_owner;
- }
- binding->chunk_owner = owner;
-
niov_idx = 0;
for_each_sgtable_dma_sg(binding->sgt, sg, sg_idx) {
dma_addr_t dma_addr = sg_dma_address(sg);
@@ -296,8 +278,8 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
size_t nr_niovs = len >> niov_shift;
for (i = 0; i < nr_niovs; i++, niov_idx++) {
- niov = &owner->area.niovs[niov_idx];
- net_iov_init(niov, &owner->area, NET_IOV_DMABUF);
+ niov = &binding->area.niovs[niov_idx];
+ net_iov_init(niov, &binding->area, NET_IOV_DMABUF);
page_pool_set_dma_addr_netmem(net_iov_to_netmem(niov),
dma_addr);
if (direction == DMA_TO_DEVICE)
@@ -311,17 +293,14 @@ net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
binding, xa_limit_32b, &id_alloc_next,
GFP_KERNEL);
if (err < 0)
- goto err_free_chunk_owner;
+ goto err_free_niovs;
list_add(&binding->list, &priv->bindings);
return binding;
-err_free_chunk_owner:
- net_devmem_dmabuf_free_chunk_owner(binding->chunk_owner);
- goto err_free_freelist;
-err_free_owner:
- kfree(owner);
+err_free_niovs:
+ kvfree(binding->area.niovs);
err_free_freelist:
kvfree(binding->freelist);
err_tx_vec:
diff --git a/net/core/devmem.h b/net/core/devmem.h
index a5ee2d8d9169..20a3eb90ea7f 100644
--- a/net/core/devmem.h
+++ b/net/core/devmem.h
@@ -14,9 +14,9 @@
#include <net/netdev_netlink.h>
struct netlink_ext_ack;
-struct dmabuf_genpool_chunk_owner;
struct net_devmem_dmabuf_binding {
+ struct net_iov_area area;
struct dma_buf *dmabuf;
struct dma_buf_attachment *attachment;
struct sg_table *sgt;
@@ -27,7 +27,6 @@ struct net_devmem_dmabuf_binding {
* dereferenced.
*/
void *vdev;
- struct dmabuf_genpool_chunk_owner *chunk_owner;
/* Protect dev */
struct mutex lock;
@@ -83,11 +82,6 @@ struct net_devmem_dmabuf_binding {
};
#if defined(CONFIG_NET_DEVMEM)
-struct dmabuf_genpool_chunk_owner {
- struct net_iov_area area;
- struct net_devmem_dmabuf_binding *binding;
-};
-
void __net_devmem_dmabuf_binding_free(struct work_struct *wq);
struct net_devmem_dmabuf_binding *
net_devmem_bind_dmabuf(struct net_device *dev, void *vdev,
@@ -102,18 +96,12 @@ int net_devmem_bind_dmabuf_to_queue(struct net_device *dev, u32 rxq_idx,
struct net_devmem_dmabuf_binding *binding,
struct netlink_ext_ack *extack);
-static inline struct dmabuf_genpool_chunk_owner *
-net_devmem_iov_to_chunk_owner(const struct net_iov *niov)
-{
- struct net_iov_area *owner = net_iov_owner(niov);
-
- return container_of(owner, struct dmabuf_genpool_chunk_owner, area);
-}
-
static inline struct net_devmem_dmabuf_binding *
net_devmem_iov_binding(const struct net_iov *niov)
{
- return net_devmem_iov_to_chunk_owner(niov)->binding;
+ struct net_iov_area *owner = net_iov_owner(niov);
+
+ return container_of(owner, struct net_devmem_dmabuf_binding, area);
}
static inline u32 net_devmem_iov_binding_id(const struct net_iov *niov)
@@ -123,11 +111,11 @@ static inline u32 net_devmem_iov_binding_id(const struct net_iov *niov)
static inline unsigned long net_iov_virtual_addr(const struct net_iov *niov)
{
- struct dmabuf_genpool_chunk_owner *co =
- net_devmem_iov_to_chunk_owner(niov);
+ struct net_devmem_dmabuf_binding *binding =
+ net_devmem_iov_binding(niov);
return net_iov_owner(niov)->base_virtual +
- ((unsigned long)net_iov_idx(niov) << co->binding->niov_shift);
+ ((unsigned long)net_iov_idx(niov) << binding->niov_shift);
}
static inline bool
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 4+ messages in thread
* [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache
2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev
@ 2026-09-11 15:45 ` Stanislav Fomichev
2 siblings, 0 replies; 4+ messages in thread
From: Stanislav Fomichev @ 2026-09-11 15:45 UTC (permalink / raw)
To: netdev
Cc: davem, edumazet, kuba, pabeni, horms, sdf, bobbyeshleman,
almasrymina, linux-kernel
Rename net_devmem_alloc_dmabuf() into net_devmem_alloc_dmabuf_bulk() and
make it refill page pool with up to PP_ALLOC_CACHE_REFILL NIOVs,
similar to io_pp_zc_alloc_netmems(). That should amortize recently
introduced freelist_lock.
Signed-off-by: Stanislav Fomichev <sdf@fomichev.me>
---
net/core/devmem.c | 59 +++++++++++++++++++++++++++++------------------
net/core/devmem.h | 8 -------
2 files changed, 37 insertions(+), 30 deletions(-)
diff --git a/net/core/devmem.c b/net/core/devmem.c
index 7949f8425bcd..a0dcc896dd12 100644
--- a/net/core/devmem.c
+++ b/net/core/devmem.c
@@ -58,25 +58,25 @@ void __net_devmem_dmabuf_binding_free(struct work_struct *wq)
kfree(binding);
}
-struct net_iov *
-net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
+static unsigned int
+net_devmem_alloc_dmabuf_bulk(struct net_devmem_dmabuf_binding *binding,
+ netmem_ref *netmems, unsigned int count)
{
- struct net_iov *niov;
+ unsigned int i;
+
spin_lock_bh(&binding->freelist_lock);
- if (unlikely(!binding->free_count)) {
- spin_unlock_bh(&binding->freelist_lock);
- return NULL;
+
+ count = min_t(size_t, count, binding->free_count);
+ for (i = 0; i < count; i++) {
+ struct net_iov *niov = binding->freelist[--binding->free_count];
+
+ binding->freelist[binding->free_count] = NULL;
+ netmems[i] = net_iov_to_netmem(niov);
}
- niov = binding->freelist[--binding->free_count];
- binding->freelist[binding->free_count] = NULL;
spin_unlock_bh(&binding->freelist_lock);
- niov->desc.pp_magic = 0;
- niov->desc.pp = NULL;
- atomic_long_set(&niov->desc.pp_ref_count, 0);
-
- return niov;
+ return count;
}
void net_devmem_free_dmabuf(struct net_iov *niov)
@@ -434,20 +434,35 @@ int mp_dmabuf_devmem_init(struct page_pool *pool)
netmem_ref mp_dmabuf_devmem_alloc_netmems(struct page_pool *pool, gfp_t gfp)
{
struct net_devmem_dmabuf_binding *binding = pool->mp_priv;
- struct net_iov *niov;
- netmem_ref netmem;
+ netmem_ref *netmems = pool->alloc.cache;
+ unsigned int allocated, i;
+
+ if (WARN_ON_ONCE(pool->alloc.count))
+ return 0;
- niov = net_devmem_alloc_dmabuf(binding);
- if (!niov)
+ allocated = net_devmem_alloc_dmabuf_bulk(binding, netmems,
+ PP_ALLOC_CACHE_REFILL);
+ if (unlikely(!allocated))
return 0;
- netmem = net_iov_to_netmem(niov);
+ for (i = 0; i < allocated; i++) {
+ struct net_iov *niov = netmem_to_net_iov(netmems[i]);
- page_pool_set_pp_info(pool, netmem);
+ niov->desc.pp_magic = 0;
+ niov->desc.pp = NULL;
+ atomic_long_set(&niov->desc.pp_ref_count, 0);
+
+ page_pool_set_pp_info(pool, netmems[i]);
+
+ pool->pages_state_hold_cnt++;
+ trace_page_pool_state_hold(pool, netmems[i],
+ pool->pages_state_hold_cnt);
+ }
- pool->pages_state_hold_cnt++;
- trace_page_pool_state_hold(pool, netmem, pool->pages_state_hold_cnt);
- return netmem;
+ /* Return the last one, the rest stay in the page_pool cache. */
+ allocated--;
+ pool->alloc.count = allocated;
+ return netmems[allocated];
}
void mp_dmabuf_devmem_destroy(struct page_pool *pool)
diff --git a/net/core/devmem.h b/net/core/devmem.h
index 20a3eb90ea7f..7195769b8bd1 100644
--- a/net/core/devmem.h
+++ b/net/core/devmem.h
@@ -133,8 +133,6 @@ net_devmem_dmabuf_binding_put(struct net_devmem_dmabuf_binding *binding)
void net_devmem_get_net_iov(struct net_iov *niov);
void net_devmem_put_net_iov(struct net_iov *niov);
-struct net_iov *
-net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding);
void net_devmem_free_dmabuf(struct net_iov *ppiov);
@@ -191,12 +189,6 @@ net_devmem_bind_dmabuf_to_queue(struct net_device *dev, u32 rxq_idx,
return -EOPNOTSUPP;
}
-static inline struct net_iov *
-net_devmem_alloc_dmabuf(struct net_devmem_dmabuf_binding *binding)
-{
- return NULL;
-}
-
static inline void net_devmem_free_dmabuf(struct net_iov *ppiov)
{
}
--
2.53.0-Meta
^ permalink raw reply related [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-11 15:46 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-11 15:45 [PATCH net-next v2 0/3] net: devmem: remove gen_pool from dma-buf allocations Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 1/3] net: devmem: replace gen_pool with freelist Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 2/3] net: devmem: embed net_iov_area in binding Stanislav Fomichev
2026-09-11 15:45 ` [PATCH net-next v2 3/3] net: devmem: batch net_iov allocations into the page_pool cache Stanislav Fomichev
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox