DPDK-dev Archive on lore.kernel.org
 help / color / mirror / Atom feed
* RE: DPDK release candidate 26.07-rc3
From: Xu, HailinX @ 2026-07-20  9:58 UTC (permalink / raw)
  To: Thomas Monjalon, dpdk
  Cc: Hosamani, Manjunathgouda, Puttaswamy, Rajesh T, Richardson, Bruce,
	Mcnamara, John
In-Reply-To: <74T0wdLsQL6DTJtar0xt1g@monjalon.net>

> -----Original Message-----
> From: Thomas Monjalon <thomas@monjalon.net>
> Sent: Monday, July 13, 2026 4:12 AM
> To: announce@dpdk.org
> Subject: DPDK release candidate 26.07-rc3
> 
> A new DPDK release candidate is ready for testing:
> 	https://git.dpdk.org/dpdk/tag/?id=v26.07-rc3
> 
> There are 88 new patches in this snapshot.
> 
> Release notes:
> 	https://doc.dpdk.org/guides/rel_notes/release_26_07.html
> 
> Please test and report issues on bugs.dpdk.org.
> 
> You may share some release validation results by replying to this message at
> dev@dpdk.org and by adding tested hardware in the release notes.
> 
> Only documentation changes and critical fixes are expected at this stage.
> 
> Thank you everyone
> 
Update the test status for Intel part. dpdk26.07-rc3 all validation test done. found 1 new issue.

New issues:
  1. vf_vlan/test_vf_vlan_strip_avx512: Failed to strip vlan packet!!!            -> has fix patch
      https://patches.dpdk.org/project/dpdk/patch/20260715225234.482186-1-anurag.mandal@intel.com

* Build & CFLAG compile: cover the build test with latest GCC/Clang version on the following OS(all passed)
  - Ubuntu25.10/Ubuntu26.04
  - RHEL9.6/RHEL10
  - Fedora43
  - FreeBSD15.0
  - SUSE16
  - OpenAnolis8.10
  - OpenEuler24.04-SP2
  - AzureLinux3.0

* Function tests: All test done and found 1 issues.
  - ICE-(E810, E825, E830, E835, E2100) PF/VF: test scenarios including basic/RTE_FLOW/TSO/Jumboframe/checksum offload/mac_filter/VLAN/VXLAN/RSS/Switch/Package Management/Flow Director/Advanced Tx/Advanced RSS/ACL/DCF/Flexible Descriptor, etc.
  - i40E-(XXV710, X722) PF/VF: test scenarios including basic/RTE_FLOW/TSO/Jumboframe/checksum offload/mac_filter/VLAN/VXLAN/RSS, etc. 
  - IXGBE-(E610, X550) PF/VF: test scenarios including basic/TSO/Jumboframe/checksum offload/mac_filter/VLAN/VXLAN/RSS, etc. 
  - IGC-(i226) PF: test scenarios including basic/RTE_FLOW/TSO/Jumboframe/checksum offload/mac_filter/VLAN/VXLAN/RSS, etc.
  - IPsec: test scenarios including ipsec/ipsec-gw/ipsec library basic test - QAT&SW/FIB library, etc.
  - Virtio: both function and performance test are covered. Such as PVP/Virtio_loopback/virtio-user loopback/virtio-net VM2VM perf testing/VMAWARE ESXI 9.0, etc.
  - Cryptodev: test scenarios including Cryptodev API testing/CompressDev ISA-L/QAT/ZLIB PMD Testing/FIPS, etc.
  - Other: test scenarios including AF_XDP, Power, CBDMA, DSA

* Performance test: All test done and passed
  - Thoughput Performance
  - Cryptodev Latency
  - PF/VF NIC single core
  - XXV710/E810/E825/E2100 NIC Performance


Regards,
Xu, Hailin

^ permalink raw reply

* Recall: DPDK release candidate 26.07-rc3
From: Xu, HailinX @ 2026-07-20  9:57 UTC (permalink / raw)
  To: Thomas Monjalon, dpdk
  Cc: Hosamani, Manjunathgouda, Puttaswamy, Rajesh T, Richardson, Bruce,
	Mcnamara, John

Xu, HailinX would like to recall the message, "DPDK release candidate 26.07-rc3".

^ permalink raw reply

* RE: DPDK release candidate 26.07-rc3
From: Xu, HailinX @ 2026-07-20  9:56 UTC (permalink / raw)
  To: Thomas Monjalon, dpdk
  Cc: Hosamani, Manjunathgouda, Puttaswamy, Rajesh T, Richardson, Bruce,
	Mcnamara, John
In-Reply-To: <74T0wdLsQL6DTJtar0xt1g@monjalon.net>

> -----Original Message-----
> From: Thomas Monjalon <thomas@monjalon.net>
> Sent: Monday, July 13, 2026 4:12 AM
> To: announce@dpdk.org
> Subject: DPDK release candidate 26.07-rc3
> 
> A new DPDK release candidate is ready for testing:
> 	https://git.dpdk.org/dpdk/tag/?id=v26.07-rc3
> 
> There are 88 new patches in this snapshot.
> 
> Release notes:
> 	https://doc.dpdk.org/guides/rel_notes/release_26_07.html
> 
> Please test and report issues on bugs.dpdk.org.
> 
> You may share some release validation results by replying to this message at
> dev@dpdk.org and by adding tested hardware in the release notes.
> 
> Only documentation changes and critical fixes are expected at this stage.
> 
> Thank you everyone
> 
Update the test status for Intel part. dpdk26.07-rc3 all validation test done. found 1 new issue.

New issues:
  1. vf_vlan/test_vf_vlan_strip_avx512: Failed to strip vlan packet!!!            -> has fix patch
      https://patches.dpdk.org/project/dpdk/patch/20260715225234.482186-1-anurag.mandal@intel.com

* Build & CFLAG compile: cover the build test with latest GCC/Clang version on the following OS(all passed)
  - Ubuntu25.10/Ubuntu26.04
  - RHEL9.6/RHEL10
  - Fedora43
  - FreeBSD15.0
  - SUSE16
  - OpenAnolis8.10
  - OpenEuler24.04-SP2
  - AzureLinux3.0

* Function tests: All test done and found 4 issues.
  - ICE-(E810, E825, E830, E835, E2100) PF/VF: test scenarios including basic/RTE_FLOW/TSO/Jumboframe/checksum offload/mac_filter/VLAN/VXLAN/RSS/Switch/Package Management/Flow Director/Advanced Tx/Advanced RSS/ACL/DCF/Flexible Descriptor, etc.
  - i40E-(XXV710, X722) PF/VF: test scenarios including basic/RTE_FLOW/TSO/Jumboframe/checksum offload/mac_filter/VLAN/VXLAN/RSS, etc. 
  - IXGBE-(E610, X550) PF/VF: test scenarios including basic/TSO/Jumboframe/checksum offload/mac_filter/VLAN/VXLAN/RSS, etc. 
  - IGC-(i226) PF: test scenarios including basic/RTE_FLOW/TSO/Jumboframe/checksum offload/mac_filter/VLAN/VXLAN/RSS, etc.
  - IPsec: test scenarios including ipsec/ipsec-gw/ipsec library basic test - QAT&SW/FIB library, etc.
  - Virtio: both function and performance test are covered. Such as PVP/Virtio_loopback/virtio-user loopback/virtio-net VM2VM perf testing/VMAWARE ESXI 9.0, etc.
  - Cryptodev: test scenarios including Cryptodev API testing/CompressDev ISA-L/QAT/ZLIB PMD Testing/FIPS, etc.
  - Other: test scenarios including AF_XDP, Power, CBDMA, DSA

* Performance test: All test done and passed
  - Thoughput Performance
  - Cryptodev Latency
  - PF/VF NIC single core
  - XXV710/E810/E825/E2100 NIC Performance


Regards,
Xu, Hailin

^ permalink raw reply

* [PATCH v3] net/enic: check notify set return value during init
From: Alexey Simakov @ 2026-07-20  9:33 UTC (permalink / raw)
  To: hyonkim; +Cc: david.marchand, dev, johndale, ssujith, stable, Alexey Simakov
In-Reply-To: <IA3PR11MB89877F0FA4E9EF63B478A201BFC42@IA3PR11MB8987.namprd11.prod.outlook.com>

The return value of vnic_dev_notify_set() is silently ignored in
enic_dev_init(), so a memory allocation failure or hardware command
error goes unnoticed and the driver continues with uninitialized
notification state.

Check the return value and propagate the error to abort probe when
notification setup fails.

Fixes: fefed3d1e62c ("enic: new driver")
Cc: stable@dpdk.org

Signed-off-by: Alexey Simakov <bigalex934@gmail.com>
---

v3 changes: add log message

v2 link: https://patches.dpdk.org/project/dpdk/patch/20260715103105.39417-1-bigalex934@gmail.com/
v2 changes: validate return code of vnic_dev_notify_set() in driver init section

v1 link: https://patches.dpdk.org/project/dpdk/patch/20260707112014.82821-1-bigalex934@gmail.com/

 drivers/net/enic/enic_main.c | 6 +++++-
 1 file changed, 5 insertions(+), 1 deletion(-)

diff --git a/drivers/net/enic/enic_main.c b/drivers/net/enic/enic_main.c
index 2696fa77d4..2a1a65d8da 100644
--- a/drivers/net/enic/enic_main.c
+++ b/drivers/net/enic/enic_main.c
@@ -1887,7 +1887,11 @@ static int enic_dev_init(struct enic *enic)
 	LIST_INIT(&enic->flows);
 
 	/* set up link status checking */
-	vnic_dev_notify_set(enic->vdev, -1); /* No Intr for notify */
+	err = vnic_dev_notify_set(enic->vdev, -1); /* No Intr for notify */
+	if (err) {
+		dev_err(enic, "failed to enable notify buffer\n");
+		return err;
+	}
 
 	enic->overlay_offload = false;
 	/*
-- 
2.53.0


^ permalink raw reply related

* [DPDK/core Bug 1970] fib: trie leaks a tbl8 group and can corrupt the dataplane  when a tbl8 allocation fails during insert
From: bugzilla @ 2026-07-20  9:27 UTC (permalink / raw)
  To: dev

http://bugs.dpdk.org/show_bug.cgi?id=1970

            Bug ID: 1970
           Summary: fib: trie leaks a tbl8 group and can corrupt the
                    dataplane  when a tbl8 allocation fails during insert
           Product: DPDK
           Version: unspecified
          Hardware: All
                OS: All
            Status: UNCONFIRMED
          Severity: normal
          Priority: Normal
         Component: core
          Assignee: dev@dpdk.org
          Reporter: maxime@leroys.fr
  Target Milestone: ---

Created attachment 356
  --> http://bugs.dpdk.org/attachment.cgi?id=356&action=edit
Patch to reproduce the bug

On a full or tight tbl8 pool, an add whose install runs out of tbl8 groups
mid-way
is not rolled back: the group(s) already taken are leaked and the tbl24 entry
already written is left dangling.
trie_modify() undoes only the RIB node (rte_rib6_remove), never the dataplane
or the tbl8 pool. A later, valid add is then wrongly refused with -ENOSPC, and
a multi-interval update can leave the dataplane inconsistent.


Steps to reproduce (fixed pool, no RCU)
---------------------------------------

IPv6 trie, 2-byte next hop, config.trie.num_tbl8 = 1:

    rte_fib6_add(fib, 2001:db8::, 32, nh);   ->  -ENOSPC
    rte_fib6_add(fib, fc00::,     25, nh);   ->  -ENOSPC   (WRONG)

Expected: the /25 needs a single tbl8 group and the pool has capacity
1, so it must succeed. It is refused because the failed /32 leaked the
only group.

Instrumenting tbl8_pool_pos at the modify_dp() call confirms the leak:

    add /32: ret=-ENOSPC, tbl8_pool_pos 0 -> 1  (group taken, not returned)
    add /25: ret=-ENOSPC, tbl8_pool_pos 1 -> 1  (no free group left)

QSBR defer mode (pool of 2, a reader kept non-quiescent):

    add /25 ; delete /25    (group deferred; rsvd_tbl8s back to 0)
    add /32                 ->  passes the reservation check, then leaks
    add /25 (other prefix)  ->  -ENOSPC (WRONG: a group was leaked)

A reproducer is attached (test/fib6). Both new cases fail on current
main:
    0001-test-fib6-reproduce-tbl8-leak-on-failed-trie-insert.patch


Root cause
----------

install_to_dp() (build_common_root()/write_edge()) allocates and
publishes tbl8 groups incrementally and has no failure path: on
tbl8_alloc() failure it returns an error, leaving the partial writes
and the taken groups in place. trie_modify() only removes the RIB node.

Two reasons the pre-check does not prevent it:

1. Transient peak. trie_modify() checks the steady-state footprint:

       if (dp->rsvd_tbl8s + count_empty_levels(node) > dp->number_tbl8s)

   but install_to_dp() transiently needs footprint + 1 for a
   byte-aligned prefix: build_common_root() descends one byte level
   past the prefix and allocates a group there; recycle_root_path()
   frees it only at the end. So the check passes with a pool of one,
   the install takes the first group, then tbl8_alloc() for the
   transient group fails.

2. Logical vs physical, under QSBR defer. rsvd_tbl8s tracks the
   logical reservation, not physical occupancy. In
   RTE_FIB6_QSBR_MODE_DQ a recycled group stays out of the pool until
   a grace period elapses, so rsvd_tbl8s drops while the group is
   still allocated, and the check accepts an add that then fails.

modify_dp() also calls install_to_dp() once per gap around covering
more-specifics; under defer mode each interval's transient group
accumulates, so the peak can exceed footprint + 1, and a failure on a
late interval leaves earlier intervals already written (a lookup in
one range returns the new next hop while another returns the old one).

-- 
You are receiving this mail because:
You are the assignee for the bug.

^ permalink raw reply

* Re: [PATCH 2/2] net/tap: drain queue FD in Rx interrupt mode
From: Maxime Leroy @ 2026-07-20  8:33 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: dev, stable
In-Reply-To: <CAOaVG16_NrF=PNTzCVEuC_BWzt+4+4WHP6yXnHH=QbzitP2xiQ@mail.gmail.com>

On Fri, Jul 17, 2026 at 2:34 PM Stephen Hemminger
<stephen@networkplumber.org> wrote:
>
> Don't use int as a boolean. Use bool. Ideally find unused pad hole for it

Agreed int-as-a-boolean is not great. But bool in a structure is
discouraged by coding_style.rst:

   "Uses of bool in structures are not preferred as is wastes space and
    it's also not clear as to what type size the bool is."  (Ref: LKML)

and that very LKML reference is Linus recommending the opposite of
bool for struct members:

   "please don't use bool in structures at all. [...] Use bool mainly as a
    return type from functions [...]. [...] just make sure the base type is
    unsigned [...] you can specify the base type as you wish [...] for
    packing."

The current tap PMD follows that too: the existing flags (persist,
flow_init, flow_isolate) are all int, and stdbool.h is not even
included in the headers.

So for v2 I'll use uint8_t for intr_mode and intr_mode_set: fixed
1-byte size, unambiguous representation, and it fits the existing
padding so neither struct grows (pmd_internals stays 144 bytes,
rx_queue stays 80). That matches both the guide and the referenced
LKML guidance.

I can also add a third patch at the end of the series converting the
existing int flags (persist, flow_init, flow_isolate) to uint8_t, to
make the driver consistent

>
> On Fri, Jul 17, 2026, 11:59 AM Maxime Leroy <maxime@leroys.fr> wrote:
>>
>> The tap Rx path is driven by a SIGIO trigger: pmd_rx_burst() returns
>> without reading the queue fd unless tap_trigger, bumped by the O_ASYNC
>> signal handler, has advanced. That avoids a readv() on every empty poll.
>>
>> In Rx interrupt mode the application blocks on the queue fd through epoll
>> and polls only after a wakeup. The epoll wakeup and the trigger are
>> distinct signals, so the burst can return 0 right after a wakeup because
>> the trigger has not advanced, leaving the fd readable with no further
>> edge: traffic stalls.
>>
>> When the port is configured for Rx interrupts, drain the fd
>> unconditionally in the burst and do not arm the SIGIO trigger on the data
>> queue fd. A data queue fd can be closed and recreated on a later setup,
>> so the mode must stay constant to keep every fd on the same SIGIO policy:
>> it is fixed at configure time and a later change is rejected. Close and
>> reopen the port to switch modes.
>>
>> Fixes: 4870a8cdd968 ("net/tap: support Rx interrupt")
>> Cc: stable@dpdk.org
>>
>> Signed-off-by: Maxime Leroy <maxime@leroys.fr>
>> ---
>>  drivers/net/tap/rte_eth_tap.c | 18 +++++++++++++++++-
>>  drivers/net/tap/rte_eth_tap.h |  2 ++
>>  2 files changed, 19 insertions(+), 1 deletion(-)
>>
>> diff --git a/drivers/net/tap/rte_eth_tap.c b/drivers/net/tap/rte_eth_tap.c
>> index 99ede19e49..ea9ef76335 100644
>> --- a/drivers/net/tap/rte_eth_tap.c
>> +++ b/drivers/net/tap/rte_eth_tap.c
>> @@ -254,6 +254,10 @@ tun_alloc(struct pmd_internals *pmd, int is_keepalive, int persistent)
>>                 goto error;
>>         }
>>
>> +       /* interrupt mode wakes through epoll on the data queue fd, not the SIGIO trigger */
>> +       if (pmd->intr_mode)
>> +               return fd;
>> +
>>         /* Find a free realtime signal */
>>         for (signo = SIGRTMIN + 1; signo < SIGRTMAX; signo++) {
>>                 struct sigaction sa;
>> @@ -477,7 +481,7 @@ pmd_rx_burst(void *queue, struct rte_mbuf **bufs, uint16_t nb_pkts)
>>         unsigned long num_rx_bytes = 0;
>>         uint32_t trigger = tap_trigger;
>>
>> -       if (trigger == rxq->trigger_seen)
>> +       if (!rxq->intr_mode && trigger == rxq->trigger_seen)
>>                 return 0;
>>
>>         process_private = rte_eth_devices[rxq->in_port].process_private;
>> @@ -937,6 +941,16 @@ tap_dev_configure(struct rte_eth_dev *dev)
>>         struct pmd_internals *pmd = dev->data->dev_private;
>>         int intr_mode = !!dev->data->dev_conf.intr_conf.rxq;
>>
>> +       /* The queue fd is created once and its SIGIO trigger is armed for poll
>> +        * mode only; the interrupt mode cannot be toggled on an existing port.
>> +        */
>> +       if (pmd->intr_mode_set && pmd->intr_mode != intr_mode) {
>> +               TAP_LOG(ERR,
>> +                       "%s: Rx interrupt mode is fixed after configure, close and reopen the port to change it",
>> +                       dev->device->name);
>> +               return -ENOTSUP;
>> +       }
>> +
>>         if (dev->data->nb_rx_queues != dev->data->nb_tx_queues) {
>>                 TAP_LOG(ERR,
>>                         "%s: number of rx queues %d must be equal to number of tx queues %d",
>> @@ -947,6 +961,7 @@ tap_dev_configure(struct rte_eth_dev *dev)
>>         }
>>
>>         pmd->intr_mode = intr_mode;
>> +       pmd->intr_mode_set = 1;
>>
>>         TAP_LOG(INFO, "%s: %s: TX configured queues number: %u",
>>                 dev->device->name, pmd->name, dev->data->nb_tx_queues);
>> @@ -1620,6 +1635,7 @@ tap_rx_queue_setup(struct rte_eth_dev *dev,
>>         rxq->queue_id = rx_queue_id;
>>         rxq->max_rx_segs = max_rx_segs;
>>         rxq->rxmode = &dev->data->dev_conf.rxmode;
>> +       rxq->intr_mode = internals->intr_mode;
>>
>>         dev->data->rx_queues[rx_queue_id] = rxq;
>>         int fd = tap_setup_queue(dev, rx_queue_id, 1);
>> diff --git a/drivers/net/tap/rte_eth_tap.h b/drivers/net/tap/rte_eth_tap.h
>> index 3180719c34..abe22aac9f 100644
>> --- a/drivers/net/tap/rte_eth_tap.h
>> +++ b/drivers/net/tap/rte_eth_tap.h
>> @@ -50,6 +50,7 @@ struct rx_queue {
>>         uint16_t queue_id;              /* queue ID*/
>>         struct queue_stats stats;        /* Stats for this RX queue */
>>         uint16_t max_rx_segs;           /* max scatter segments per packet */
>> +       uint16_t intr_mode;             /* 1 when Rx queue interrupts are used */
>>         struct rte_eth_rxmode *rxmode;  /* RX features */
>>         struct rte_mbuf *pool;          /* mbufs pool for this queue */
>>         struct tun_pi pi;               /* packet info for iovecs */
>> @@ -94,6 +95,7 @@ struct pmd_internals {
>>
>>         struct rte_intr_handle *intr_handle;         /* LSC interrupt handle. */
>>         int intr_mode;                    /* Rx queue interrupt mode */
>> +       int intr_mode_set;                /* intr_mode locked after configure */
>>         int ka_fd;                        /* keep-alive file descriptor */
>>         struct rte_mempool *gso_ctx_mp;     /* Mempool for GSO packets */
>>  };
>> --
>> 2.43.0
>>


-- 
-------------------------------
Maxime Leroy
maxime@leroys.fr

^ permalink raw reply

* Re: mempool cache change
From: Bruce Richardson @ 2026-07-20  8:12 UTC (permalink / raw)
  To: Morten Brørup
  Cc: Kishore Padmanabha, fengchengwen, Thomas Monjalon, dev,
	Wisam Jaddo, Andrew Rybchenko
In-Reply-To: <98CBD80474FA8B44BF855DF32C47DC35F65979@smartserver.smartshare.dk>

On Sat, Jul 18, 2026 at 04:07:05PM +0200, Morten Brørup wrote:
> > From: Kishore Padmanabha [mailto:kishore.padmanabha@broadcom.com]
> > Sent: Friday, 17 July 2026 19.10
> > 
> > Hi Bruce,
> > 
> > The below patch works fine for us. We tested all the different packet
> > sizes.
> > Thanks for the patch. Do you want to push this patch since it is not
> > changing the ABI/API?
> 
> Too late in the release process.
> 
> Let's postpone the discussion for DPDK 26.11, where API/ABI breakage is allowed.

Agreed. Let's not unnecessarily rush this.

> 
> I'm not strongly opposed to Bruce's algorithm, targeting a fill level of 25 % from the edges and flushing/refilling up to 75 % of the cache when necessary. It does have its advantages for some mempool access patterns (which are not exotic).
> I just prefer the current algorithm, targeting a fill level of 50 % and only flushing/refilling up to 50 % of the cache when necessary. It performs better at random get/put access patterns, and the backend transactions are smaller.
> 
> For DPDK 26.11, where we can break the API/ABI, we can simply double RTE_MEMPOOL_CACHE_MAX_SIZE to 1024, to compensate for reducing the effective cache size from 150 % to 100 %.
> The mempool cache objs array will no longer be [RTE_MEMPOOL_CACHE_MAX_SIZE * 2], but only [RTE_MEMPOOL_CACHE_MAX_SIZE], so doubling RTE_MEMPOOL_CACHE_MAX_SIZE will not increase the memory footprint, but allow using a cache size up to 1024.

My concern with this approach is that it won't automatically fix the
problem if we have users who experience a performance regression due to the
mempool changes. While testpmd allows the mbcache size to be provided via
parameter, end applications are likely to have it hardcoded. That means
that if an app does experience a regression, the author/user has to be
either aware of the mempool changes, or has to debug it down to the mempool
and then know to increase the mempool cache size in the app.

It's not an insurmountable problem, but one that needs to be very clearly
called out in our documentation, what the change is, how it may affect
things and how to fix it.

On the other hand, in realworld, i.e. not just testpmd/l3fwd cases doing
little packet processesing, I would be fairly hopeful that regressions are
going to be few and very small.

> 
> Please also note that the current implementation is carefully designed to keep the transfers to/from the mempool backend CPU cache aligned (assuming cache->size is 2^N and large enough).
> Refer to the parameters passed to rte_mempool_ops_enqueue/dequeue_bulk().
> E.g. with mempool cache size 256, backend transfers are 128 objects, 16 full cache lines.
> Using CPU cache aligned transfers has a few advantages:
> - There are no cache line ownership issues across different CPU cores repeatedly accessing the backend.
> - The mempool backend drivers can be performance optimized for transferring full CPU cache lines. (Both source and destination addresses, and number of objects copied, are CPU cache aligned. Assuming all transfers go via the mempool cache.)
> 
> These details should be fine tuned in the implementation, if we do proceed with Bruce's algorithm.
> 
Yep, good points.

/Bruce

^ permalink raw reply

* [PATCH 5/6] examples/l3fwd-power: recheck Rx queues before sleeping
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy
In-Reply-To: <20260720080035.673782-1-maxime@leroys.fr>

After arming the Rx interrupts rx_interrupt_wait() sleeps in
rte_epoll_wait(). A packet that arrived between the last poll and the arm
can be missed on PMDs whose interrupt is edge-triggered on an empty to
non-empty transition, such as dpaa2: arming a queue that is already
non-empty raises no notification, so the lcore sleeps until the next
packet.

Before sleeping, check whether any queue already holds traffic and, if
so, skip the wait and go back to polling. NICs with level-triggered
interrupts are unaffected: a pending packet keeps the interrupt asserted,
so rte_epoll_wait() would return immediately anyway.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 27 ++++++++++++++++++++++++++-
 1 file changed, 26 insertions(+), 1 deletion(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index fa64e28e5c..c423b0dba5 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -949,6 +949,26 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
+static bool
+rx_queue_pending(struct lcore_conf *qconf)
+{
+	struct lcore_rx_queue *rx_queue;
+	uint16_t queue_id;
+	uint16_t port_id;
+	int i;
+
+	for (i = 0; i < qconf->n_rx_queue; ++i) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		if (rte_eth_rx_queue_count(port_id, queue_id) > 0)
+			return true;
+	}
+
+	return false;
+}
+
 static int
 rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
 {
@@ -968,7 +988,12 @@ rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
 		*intr_registered = 1;
 	}
 
-	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
+	/*
+	 * A packet that arrived during the arm window raises no new wakeup on
+	 * edge-triggered PMDs, so skip the sleep if a queue already has traffic.
+	 */
+	if (!rx_queue_pending(qconf))
+		sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
 	rx_intr_disable_all(qconf);
 	return 0;
 }
-- 
2.43.0


^ permalink raw reply related

* [PATCH 6/6] examples/l3fwd-power: block until Rx interrupt or exit
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy
In-Reply-To: <20260720080035.673782-1-maxime@leroys.fr>

In interrupt and legacy modes an idle worker arms its Rx queues and
sleeps in rte_epoll_wait() with a 10 ms timeout. The timeout lets the
worker wake periodically to re-check the quit flag, but it also hides a
broken interrupt path: with no interrupt delivered the worker still wakes
every 10 ms and polls, so traffic keeps flowing and interrupt mode looks
like it works when it does not.

Block indefinitely instead, so only a real Rx interrupt can wake the
worker. To still stop the workers on SIGINT, add an eventfd to each
worker epoll set; the signal handler writes it to wake every worker,
which then observes the quit flag and leaves its loop. Both loops share
this sleep path, so the eventfd is created for the interrupt and legacy
modes; it is skipped in the wakeup log so it is not reported as an Rx
interrupt.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 40 +++++++++++++++++++++++++++++++++----
 1 file changed, 36 insertions(+), 4 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index c423b0dba5..73ee2d2def 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -9,6 +9,8 @@
 #include <sys/types.h>
 #include <string.h>
 #include <sys/queue.h>
+#include <sys/epoll.h>
+#include <sys/eventfd.h>
 #include <stdarg.h>
 #include <errno.h>
 #include <getopt.h>
@@ -143,6 +145,9 @@ static int promiscuous_on = 0;
 /* NUMA is enabled by default. */
 static int numa_on = 1;
 volatile bool quit_signal;
+/* eventfd to wake idle workers out of a blocking rte_epoll_wait() at exit */
+static int wakeup_fd = -1;
+static struct rte_epoll_event wakeup_event[RTE_MAX_LCORE];
 /* timer to update telemetry every 500ms */
 static struct rte_timer telemetry_timer;
 
@@ -420,10 +425,17 @@ static int is_done(void)
 static void
 signal_exit_now(int sigtype)
 {
+	uint64_t one = 1;
+	ssize_t nbytes;
 
-	if (sigtype == SIGINT)
+	if (sigtype == SIGINT) {
 		quit_signal = true;
-
+		/* interrupt mode only: wake a worker blocked in rte_epoll_wait() */
+		if (wakeup_fd >= 0) {
+			nbytes = write(wakeup_fd, &one, sizeof(one));
+			RTE_SET_USED(nbytes);
+		}
+	}
 }
 
 /*  Frequency scale down timer callback */
@@ -835,7 +847,7 @@ sleep_until_rx_interrupt(int num, int lcore)
 	static alignas(RTE_CACHE_LINE_SIZE) struct {
 		bool wakeup;
 	} status[RTE_MAX_LCORE];
-	struct rte_epoll_event event[num];
+	struct rte_epoll_event event[num + 1];
 	int n, i;
 	uint16_t port_id;
 	uint16_t queue_id;
@@ -847,8 +859,10 @@ sleep_until_rx_interrupt(int num, int lcore)
 				rte_lcore_id());
 	}
 
-	n = rte_epoll_wait(RTE_EPOLL_PER_THREAD, event, num, 10);
+	n = rte_epoll_wait(RTE_EPOLL_PER_THREAD, event, num + 1, -1);
 	for (i = 0; i < n; i++) {
+		if (event[i].fd == wakeup_fd)
+			continue;
 		data = event[i].epdata.data;
 		port_id = ((uintptr_t)data) >> (sizeof(uint16_t) * CHAR_BIT);
 		queue_id = ((uintptr_t)data) &
@@ -921,6 +935,7 @@ rx_intr_enable_all(struct lcore_conf *qconf)
 static int event_register(struct lcore_conf *qconf)
 {
 	struct lcore_rx_queue *rx_queue;
+	struct rte_epoll_event *wev;
 	uint16_t queueid;
 	uint16_t portid;
 	uint32_t data;
@@ -946,6 +961,16 @@ static int event_register(struct lcore_conf *qconf)
 			return ret;
 	}
 
+	/* also watch the shutdown eventfd so a blocking wait wakes at exit */
+	if (wakeup_fd >= 0) {
+		wev = &wakeup_event[rte_lcore_id()];
+		wev->epdata.event = EPOLLIN | EPOLLPRI;
+		ret = rte_epoll_ctl(RTE_EPOLL_PER_THREAD, EPOLL_CTL_ADD,
+				    wakeup_fd, wev);
+		if (ret && ret != -EEXIST)
+			return ret;
+	}
+
 	return 0;
 }
 
@@ -2989,6 +3014,13 @@ main(int argc, char **argv)
 
 	check_all_ports_link_status(enabled_port_mask);
 
+	if (app_mode == APP_MODE_LEGACY || app_mode == APP_MODE_INTERRUPT) {
+		wakeup_fd = eventfd(0, EFD_NONBLOCK);
+		if (wakeup_fd < 0)
+			rte_exit(EXIT_FAILURE, "eventfd() failed: %s\n",
+				 strerror(errno));
+	}
+
 	/* launch per-lcore init on every lcore */
 	if (app_mode == APP_MODE_LEGACY) {
 		rte_eal_mp_remote_launch(main_legacy_loop, NULL, CALL_MAIN);
-- 
2.43.0


^ permalink raw reply related

* [PATCH 4/6] examples/l3fwd-power: accept shared Rx interrupt FD
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy
In-Reply-To: <20260720080035.673782-1-maxime@leroys.fr>

Some PMDs deliver the Rx interrupts of several queues through a single
shared fd; dpaa2, for instance, uses one QBMan portal per lcore.
Registering the fd a second time then returns -EEXIST, which
event_register() treated as fatal and disabled interrupt mode for the
whole lcore.

Accept -EEXIST as success. NICs with a dedicated fd per queue never
return it, so their behaviour is unchanged.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 4874a55de1..fa64e28e5c 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -937,7 +937,12 @@ static int event_register(struct lcore_conf *qconf)
 						RTE_EPOLL_PER_THREAD,
 						RTE_INTR_EVENT_ADD,
 						(void *)((uintptr_t)data));
-		if (ret)
+		/*
+		 * Queues polled on the same lcore may share one interrupt fd,
+		 * so registering the second onward returns -EEXIST; treat it
+		 * as success.
+		 */
+		if (ret && ret != -EEXIST)
 			return ret;
 	}
 
-- 
2.43.0


^ permalink raw reply related

* [PATCH 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy
In-Reply-To: <20260720080035.673782-1-maxime@leroys.fr>

event_register() added each Rx queue interrupt fd to the per-thread epoll
set before the worker loop, while the queues were still disarmed. This
works when the fd is assigned at queue setup and stays fixed, as on MSI-X
NICs.

The dpaa2 PMD delivers Rx interrupts through a per-lcore QBMan portal and
binds the fd when the interrupt is enabled, not at queue setup. Adding a
queue to the epoll set before arming it then watches an fd that is not
bound yet, so the lcore never wakes from rte_epoll_wait().

Register from rx_interrupt_wait(), after the queues are armed, and only
once: keep the epoll entry installed for the lifetime of the loop. NICs
with a fixed per-queue fd are unaffected.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 34 +++++++++++++++++-----------------
 1 file changed, 17 insertions(+), 17 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 70f76d202e..4874a55de1 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -945,7 +945,7 @@ static int event_register(struct lcore_conf *qconf)
 }
 
 static int
-rx_interrupt_wait(struct lcore_conf *qconf)
+rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
 {
 	int ret;
 
@@ -953,6 +953,16 @@ rx_interrupt_wait(struct lcore_conf *qconf)
 	if (ret != 0)
 		return ret;
 
+	/* some PMDs expose the interrupt fd only once the queue is armed */
+	if (!*intr_registered) {
+		ret = event_register(qconf);
+		if (ret != 0) {
+			rx_intr_disable_all(qconf);
+			return ret;
+		}
+		*intr_registered = 1;
+	}
+
 	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
 	rx_intr_disable_all(qconf);
 	return 0;
@@ -970,7 +980,8 @@ static int main_intr_loop(__rte_unused void *dummy)
 	struct lcore_rx_queue *rx_queue;
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
-	int intr_en = 0;
+	int intr_registered = 0;
+	int intr_en = 1;
 	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) /
@@ -998,12 +1009,6 @@ static int main_intr_loop(__rte_unused void *dummy)
 				lcore_id, portid, queueid);
 	}
 
-	/* add into event wait list */
-	if (event_register(qconf) == 0)
-		intr_en = 1;
-	else
-		RTE_LOG(INFO, L3FWD_POWER, "RX interrupt won't enable.\n");
-
 	while (!is_done()) {
 		stats[lcore_id].nb_iteration_looped++;
 
@@ -1103,7 +1108,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					ret = rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf, &intr_registered);
 					if (ret == -EAGAIN)
 						goto start_rx;
 					if (ret != 0) {
@@ -1260,7 +1265,8 @@ main_legacy_loop(__rte_unused void *dummy)
 	enum freq_scale_hint_t lcore_scaleup_hint;
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
-	int intr_en = 0;
+	int intr_registered = 0;
+	int intr_en = 1;
 	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) / US_PER_S * BURST_TX_DRAIN_US;
@@ -1286,12 +1292,6 @@ main_legacy_loop(__rte_unused void *dummy)
 			"rxqueueid=%" PRIu16 "\n", lcore_id, portid, queueid);
 	}
 
-	/* add into event wait list */
-	if (event_register(qconf) == 0)
-		intr_en = 1;
-	else
-		RTE_LOG(INFO, L3FWD_POWER, "RX interrupt won't enable.\n");
-
 	while (!is_done()) {
 		stats[lcore_id].nb_iteration_looped++;
 
@@ -1424,7 +1424,7 @@ main_legacy_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					ret = rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf, &intr_registered);
 					if (ret == -EAGAIN)
 						goto start_rx;
 					if (ret != 0) {
-- 
2.43.0


^ permalink raw reply related

* [PATCH 2/6] examples/l3fwd-power: check Rx interrupt enable errors
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy
In-Reply-To: <20260720080035.673782-1-maxime@leroys.fr>

rx_interrupt_wait() enabled the Rx queue interrupts of an lcore through
turn_on_off_intr() and ignored the return of
rte_eth_dev_rx_intr_enable(). Some PMDs report a transient condition
there: for example dpaa2 returns -EAGAIN when it finds traffic already
queued while arming, meaning the queue was not armed and the lcore must
poll rather than sleep.

Split the helper into rx_intr_enable_all(), which returns the error and
unwinds the queues it already armed, and rx_intr_disable_all(). Return
the error from rx_interrupt_wait(); in the sleep path, on -EAGAIN go back
to polling, and on any other error give up interrupt mode for the lcore
instead of sleeping on unarmed queues.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 74 ++++++++++++++++++++++++++++++++-----
 1 file changed, 64 insertions(+), 10 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 764899217c..70f76d202e 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -863,12 +863,32 @@ sleep_until_rx_interrupt(int num, int lcore)
 	return 0;
 }
 
-static void turn_on_off_intr(struct lcore_conf *qconf, bool on)
+static void
+rx_intr_disable_all(struct lcore_conf *qconf)
 {
+	struct lcore_rx_queue *rx_queue;
+	uint16_t queue_id;
+	uint16_t port_id;
 	int i;
+
+	for (i = 0; i < qconf->n_rx_queue; ++i) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		rte_spinlock_lock(&(locks[port_id]));
+		rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		rte_spinlock_unlock(&(locks[port_id]));
+	}
+}
+
+static int
+rx_intr_enable_all(struct lcore_conf *qconf)
+{
 	struct lcore_rx_queue *rx_queue;
 	uint16_t queue_id;
 	uint16_t port_id;
+	int i, ret;
 
 	for (i = 0; i < qconf->n_rx_queue; ++i) {
 		rx_queue = &(qconf->rx_queue_list[i]);
@@ -876,12 +896,26 @@ static void turn_on_off_intr(struct lcore_conf *qconf, bool on)
 		queue_id = rx_queue->queue_id;
 
 		rte_spinlock_lock(&(locks[port_id]));
-		if (on)
-			rte_eth_dev_rx_intr_enable(port_id, queue_id);
-		else
-			rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		ret = rte_eth_dev_rx_intr_enable(port_id, queue_id);
 		rte_spinlock_unlock(&(locks[port_id]));
+		if (ret != 0)
+			goto fail;
 	}
+
+	return 0;
+
+fail:
+	while (--i >= 0) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		rte_spinlock_lock(&(locks[port_id]));
+		rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		rte_spinlock_unlock(&(locks[port_id]));
+	}
+
+	return ret;
 }
 
 static int event_register(struct lcore_conf *qconf)
@@ -910,12 +944,18 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
-static void
+static int
 rx_interrupt_wait(struct lcore_conf *qconf)
 {
-	turn_on_off_intr(qconf, 1);
+	int ret;
+
+	ret = rx_intr_enable_all(qconf);
+	if (ret != 0)
+		return ret;
+
 	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
-	turn_on_off_intr(qconf, 0);
+	rx_intr_disable_all(qconf);
+	return 0;
 }
 
 /* Main processing loop. 8< */
@@ -931,6 +971,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
 	int intr_en = 0;
+	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) /
 				   US_PER_S * BURST_TX_DRAIN_US;
@@ -1062,7 +1103,13 @@ static int main_intr_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf);
+					if (ret == -EAGAIN)
+						goto start_rx;
+					if (ret != 0) {
+						intr_en = 0;
+						continue;
+					}
 					/**
 					 * start receiving packets immediately
 					 */
@@ -1214,6 +1261,7 @@ main_legacy_loop(__rte_unused void *dummy)
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
 	int intr_en = 0;
+	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) / US_PER_S * BURST_TX_DRAIN_US;
 
@@ -1376,7 +1424,13 @@ main_legacy_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf);
+					if (ret == -EAGAIN)
+						goto start_rx;
+					if (ret != 0) {
+						intr_en = 0;
+						continue;
+					}
 					/**
 					 * start receiving packets immediately
 					 */
-- 
2.43.0


^ permalink raw reply related

* [PATCH 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy
In-Reply-To: <20260720080035.673782-1-maxime@leroys.fr>

The interrupt and legacy main loops carry an identical block that arms
the Rx queue interrupts of the lcore, sleeps until one triggers and
disarms them again. Move it to a helper so the following fixes touch a
single place.

No functional change.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 705cab8f2d..764899217c 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -910,6 +910,14 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
+static void
+rx_interrupt_wait(struct lcore_conf *qconf)
+{
+	turn_on_off_intr(qconf, 1);
+	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
+	turn_on_off_intr(qconf, 0);
+}
+
 /* Main processing loop. 8< */
 static int main_intr_loop(__rte_unused void *dummy)
 {
@@ -1054,11 +1062,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					turn_on_off_intr(qconf, 1);
-					sleep_until_rx_interrupt(
-							qconf->n_rx_queue,
-							lcore_id);
-					turn_on_off_intr(qconf, 0);
+					rx_interrupt_wait(qconf);
 					/**
 					 * start receiving packets immediately
 					 */
@@ -1372,11 +1376,7 @@ main_legacy_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					turn_on_off_intr(qconf, 1);
-					sleep_until_rx_interrupt(
-							qconf->n_rx_queue,
-							lcore_id);
-					turn_on_off_intr(qconf, 0);
+					rx_interrupt_wait(qconf);
 					/**
 					 * start receiving packets immediately
 					 */
-- 
2.43.0


^ permalink raw reply related

* [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy

l3fwd-power's Rx interrupt (one-shot) mode was written around NICs whose
interrupt fd is assigned per queue at setup, stays fixed and is
level-triggered, as on MSI-X hardware. PMDs that do not fit that model
never sleep on interrupts: they either fail to arm, watch an unbound fd,
or miss the wake, and the 10 ms poll timeout hides all of it.

dpaa2 is the motivating example: it delivers Rx interrupts through one
QBMan portal per lcore, so the queues of a lcore share a single fd that
is bound when the interrupt is enabled, not at queue setup; the
notification is edge-triggered on an empty to non-empty transition; and
arming a queue that already holds traffic returns -EAGAIN with the queue
left unarmed.

This series makes the interrupt and legacy loops handle those cases
while leaving MSI-X NICs unchanged:

  1/6 factor the duplicated arm/sleep/disarm block into a helper.
  2/6 check the arm return: on -EAGAIN poll and retry, on any other
      error drop interrupt mode for the lcore instead of sleeping on
      unarmed queues.
  3/6 register the fd in the epoll set after arming, not before, so a
      fd bound at enable time is watched only once it exists.
  4/6 accept -EEXIST when several queues share one fd.
  5/6 recheck the queues before sleeping so a packet that arrived during
      the arm window is not missed under edge-triggered interrupts.
  6/6 block indefinitely instead of on a 10 ms timeout, and wake the
      workers through an eventfd on exit, so a broken interrupt path no
      longer hides behind periodic polling.

Maxime Leroy (6):
  examples/l3fwd-power: factor out Rx interrupt sleep path
  examples/l3fwd-power: check Rx interrupt enable errors
  examples/l3fwd-power: enable Rx interrupt before epoll add
  examples/l3fwd-power: accept shared Rx interrupt FD
  examples/l3fwd-power: recheck Rx queues before sleeping
  examples/l3fwd-power: block until Rx interrupt or exit

 examples/l3fwd-power/main.c | 184 +++++++++++++++++++++++++++++-------
 1 file changed, 150 insertions(+), 34 deletions(-)

-- 
2.43.0


^ permalink raw reply

* RE: [EXTERNAL] [PATCH] app/test-crypto-perf: reset mbuf state for IPsec outbound iterations
From: Nithinsen Kaithakadan @ 2026-07-20  7:44 UTC (permalink / raw)
  To: Gagandeep Singh, dev@dpdk.org; +Cc: hemant.agrawal@nxp.com, stable@dpdk.org
In-Reply-To: <20260716071620.1550810-1-g.singh@nxp.com>



> In cperf_set_ops_security_ipsec(), the same mbuf is reused across multiple
> throughput iterations.  For outbound (encap) IPsec the PMD may modify
> mbuf metadata such as data_off, data_len and pkt_len as part of in-place
> protocol processing.  When the mbuf is recycled for the next burst those
> modifications must be undone so the PMD always sees a clean, freshly-
> allocated-like buffer.
> 
> Use rte_pktmbuf_reset() to restore the mbuf to its post-alloc state in a PMD-
> agnostic manner, rather than patching individual fields with driver-specific
> knowledge of the expected offset values.
> 
> Without this reset, drivers that back-adjust data_off after encap (e.g. DPAA2
> SEC, which advances data_off by SEC_FLC_DHR_OUTBOUND on each
> dequeue) will receive a buffer with insufficient headroom on the second and
> subsequent enqueues, causing hardware errors such as:
> 
>   DPAA2_SEC: SEC encap returned Error - 50000045
>     (QI error 0x45: DHR correction underflow, reuse mode)
>   DPAA2_SEC: SEC returned Error - 50000020
>     (QI error 0x20: FD format error)
> 
> Fixes: 3b7d9f2bc6 ("app/test-crypto-perf: add ipsec lookaside support")
> Cc: stable@dpdk.org
> 
> Signed-off-by: Gagandeep Singh <g.singh@nxp.com>
> ---


Acked-by: Nithinsen Kaithakadan <nkaithakadan@marvell.com>

^ permalink raw reply

* RE: [PATCH v2] net/mlx5: avoid wildcard drop flow in priority discovery
From: Suanming Mou @ 2026-07-20  5:42 UTC (permalink / raw)
  To: Gavin Hu, dev@dpdk.org
  Cc: Wei Yan, stable@dpdk.org, Dariusz Sosnowski, Slava Ovsiienko,
	Bing Zhao, Ori Kam, Matan Azrad, Ophir Munk, Dmitry Kozlyuk
In-Reply-To: <20260714060751.44814-1-gahu@nvidia.com>



> -----Original Message-----
> From: Gavin Hu <gahu@nvidia.com>
> Sent: Tuesday, July 14, 2026 2:08 PM
> To: dev@dpdk.org
> Cc: Wei Yan <yanwei.wy@bytedance.com>; stable@dpdk.org; Dariusz
> Sosnowski <dsosnowski@nvidia.com>; Slava Ovsiienko
> <viacheslavo@nvidia.com>; Bing Zhao <bingz@nvidia.com>; Ori Kam
> <orika@nvidia.com>; Suanming Mou <suanmingm@nvidia.com>; Matan
> Azrad <matan@nvidia.com>; Ophir Munk <ophirmu@nvidia.com>; Dmitry
> Kozlyuk <dmitry.kozliuk@gmail.com>
> Subject: [PATCH v2] net/mlx5: avoid wildcard drop flow in priority discovery
> 
> From: Wei Yan <yanwei.wy@bytedance.com>
> 
> Priority discovery installs temporary drop flows to probe supported flow
> priorities. These flows match a wildcard Ethernet pattern, so they can also
> match real traffic while discovery is running. On shared devices this may
> temporarily drop kernel traffic and cause TCP connection timeouts.
> 
> Constrain the probe flows to match both source and destination MAC
> addresses as ff:ff:ff:ff:ff:ff. This keeps the rule valid for discovery while
> avoiding matches on normal Ethernet traffic.
> 
> Fixes: 3eca5f8a610e ("net/mlx5: move flow priority discovery to Verbs file")
> Fixes: c5042f93a425 ("net/mlx5: discover max flow priority using DevX")
> Cc: stable@dpdk.org
> 
> Signed-off-by: Wei Yan <yanwei.wy@bytedance.com>
> Signed-off-by: Gavin Hu <gahu@nvidia.com>
Acked-by: Suanming Mou <suanmingm@nvidia.com>

Thanks

^ permalink raw reply

* [PATCH] net/iavf: fix null pointer dereference in eCPRI FDIR pattern
From: sandeep.penigalapati @ 2026-07-20  0:34 UTC (permalink / raw)
  To: dev; +Cc: vladimir.medvedkin, bruce.richardson, stable,
	Sandeep Penigalapati

From: Sandeep Penigalapati <sandeep.penigalapati@intel.com>

In iavf_fdir_parse_pattern(), the RTE_FLOW_ITEM_TYPE_ECPRI case reads
ecpri_spec->hdr.common.u32 before checking that ecpri_spec is non-null.
A wildcard eCPRI flow item (no spec) with a queue or drop action, which
routes the rule to the FDIR engine, leaves item->spec as null, so the
unconditional dereference crashes the application at rule creation.

Move the ecpri_common read inside the existing
"if (ecpri_spec && ecpri_mask)" guard, mirroring the other pattern items
and the RSS-hash parser in iavf_hash.c which already handles a null
eCPRI spec. The rule is then rejected cleanly with -EINVAL.

Fixes: 6c37927479be ("net/iavf: support eCPRI message type 0 for flow director")
Cc: stable@dpdk.org
Signed-off-by: Sandeep Penigalapati <sandeep.penigalapati@intel.com>
---
 drivers/net/intel/iavf/iavf_fdir.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/drivers/net/intel/iavf/iavf_fdir.c b/drivers/net/intel/iavf/iavf_fdir.c
index ea620863ca..67003183bb 100644
--- a/drivers/net/intel/iavf/iavf_fdir.c
+++ b/drivers/net/intel/iavf/iavf_fdir.c
@@ -1271,13 +1271,13 @@ iavf_fdir_parse_pattern(__rte_unused struct iavf_adapter *ad,
 			ecpri_spec = item->spec;
 			ecpri_mask = item->mask;
 
-			ecpri_common.u32 = rte_be_to_cpu_32(ecpri_spec->hdr.common.u32);
-
 			hdr = &hdrs->proto_hdr[layer];
 
 			VIRTCHNL_SET_PROTO_HDR_TYPE(hdr, ECPRI);
 
 			if (ecpri_spec && ecpri_mask) {
+				ecpri_common.u32 = rte_be_to_cpu_32(ecpri_spec->hdr.common.u32);
+
 				if (ecpri_common.type == RTE_ECPRI_MSG_TYPE_IQ_DATA &&
 						ecpri_mask->hdr.type0.pc_id == UINT16_MAX) {
 					input_set |= IAVF_ECPRI_PC_RTC_ID;
-- 
2.27.0


^ permalink raw reply related

* [PATCH] doc: add tested platforms with NVIDIA NICs
From: Raslan Darawsheh @ 2026-07-19 13:21 UTC (permalink / raw)
  To: thomas; +Cc: dev

Add tested platforms with NVIDIA NICs to the 26.07 release notes.

Signed-off-by: Raslan Darawsheh <rasland@nvidia.com>
---
 doc/guides/rel_notes/release_26_07.rst | 118 +++++++++++++++++++++++++
 1 file changed, 118 insertions(+)

diff --git a/doc/guides/rel_notes/release_26_07.rst b/doc/guides/rel_notes/release_26_07.rst
index 6badd6d91b..fc97790116 100644
--- a/doc/guides/rel_notes/release_26_07.rst
+++ b/doc/guides/rel_notes/release_26_07.rst
@@ -403,3 +403,121 @@ Tested Platforms
    This section is a comment. Do not overwrite or remove it.
    Also, make sure to start the actual text at the margin.
    =======================================================
+
+* Intel\ |reg| platforms with NVIDIA\ |reg| NICs combinations
+
+  * CPU:
+
+    * Intel\ |reg| Xeon\ |reg| Gold 6154 CPU @ 3.00GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2697A v4 @ 2.60GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2697 v3 @ 2.60GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2680 v2 @ 2.80GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2670 0 @ 2.60GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2650 v4 @ 2.20GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2650 v3 @ 2.30GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2640 @ 2.50GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2650 0 @ 2.00GHz
+    * Intel\ |reg| Xeon\ |reg| CPU E5-2620 v4 @ 2.10GHz
+
+  * OS:
+
+    * Red Hat Enterprise Linux release 9.2 (Plow)
+    * Red Hat Enterprise Linux release 9.1 (Plow)
+    * Red Hat Enterprise Linux release 8.6 (Ootpa)
+    * Ubuntu 24.04
+    * Ubuntu 22.04
+    * SUSE Enterprise Linux 15 SP4
+
+  * DOCA:
+
+    * DOCA 3.5.0-035000 and above.
+
+  * upstream kernel:
+
+    * Linux 6.18.0 and above
+
+  * rdma-core:
+
+    * rdma-core-60.0 and above
+
+  * NICs
+
+    * NVIDIA\ |reg| ConnectX\ |reg|-6 Dx EN 100G MCX623106AN-CDAT (2x100G)
+
+      * Host interface: PCI Express 4.0 x16
+      * Device ID: 15b3:101d
+      * Firmware version: 22.50.0394 and above
+
+    * NVIDIA\ |reg| ConnectX\ |reg|-6 Lx EN 25G MCX631102AN-ADAT (2x25G)
+
+      * Host interface: PCI Express 4.0 x8
+      * Device ID: 15b3:101f
+      * Firmware version: 26.50.0394 and above
+
+    * NVIDIA\ |reg| ConnectX\ |reg|-7 200G CX713106AE-HEA_QP1_Ax (2x200G)
+
+      * Host interface: PCI Express 5.0 x16
+      * Device ID: 15b3:1021
+      * Firmware version: 28.50.0394 and above
+
+    * NVIDIA\ |reg| ConnectX\ |reg|-8 SuperNIC 400G  MT4131 - 900-9X81Q-00CN-STA (2x400G)
+
+      * Host interface: PCI Express 6.0 x16
+      * Device ID: 15b3:1023
+      * Firmware version: 40.50.0394 and above
+
+    * NVIDIA\ |reg| ConnectX\ |reg|-9 SuperNIC 800G  MT4133 - 900-9X91E-00EB-STA_Ax (1x800G)
+
+      * Host interface: PCI Express 6.0 x16
+      * Device ID: 15b3:1025
+      * Firmware version: 82.50.0394 and above
+
+* NVIDIA\ |reg| BlueField\ |reg| SmartNIC
+
+  * NVIDIA\ |reg| BlueField\ |reg|-2 SmartNIC MT41686 - MBF2H332A-AEEOT_A1 (2x25G)
+
+    * Host interface: PCI Express 3.0 x16
+    * Device ID: 15b3:a2d6
+    * Firmware version: 24.50.0394 and above
+
+  * NVIDIA\ |reg| BlueField\ |reg|-3 P-Series DPU MT41692 - 900-9D3B6-00CV-AAB (2x200G)
+
+    * Host interface: PCI Express 5.0 x16
+    * Device ID: 15b3:a2dc
+    * Firmware version: 32.49.1014 and above
+
+  * Embedded software:
+
+    * Ubuntu 24.04
+    * MLNX_OFED 26.07-0.3.5.0
+    * DOCA 3.5.0-035000
+    * bf-bundle-3.5.0-47_26.7_ubuntu-24.04
+    * DPDK application running on ARM cores
+
+* IBM Power 9 platforms with NVIDIA\ |reg| NICs combinations
+
+  * CPU:
+
+    * POWER9 2.2 (pvr 004e 1202)
+
+  * OS:
+
+    * Ubuntu 24.04
+
+  * NICs:
+
+    * NVIDIA\ |reg| ConnectX\ |reg|-6 Dx 100G MCX623106AN-CDAT (2x100G)
+
+      * Host interface: PCI Express 4.0 x16
+      * Device ID: 15b3:101d
+      * Firmware version: 22.50.0394 and above
+
+    * NVIDIA\ |reg| ConnectX\ |reg|-7 200G CX713106AE-HEA_QP1_Ax (2x200G)
+
+      * Host interface: PCI Express 5.0 x16
+      * Device ID: 15b3:1021
+      * Firmware version: 28.50.0394 and above
+
+  * DOCA:
+
+    * DOCA 3.5.0-035000 and above
-- 
2.55.0


^ permalink raw reply related

* RE: [PATCH] net/iavf: harden reset recovery and data path on link flap
From: Mandal, Anurag @ 2026-07-19 11:17 UTC (permalink / raw)
  To: dev@dpdk.org; +Cc: Mandal, Anurag
In-Reply-To: <20260717211526.716220-1-anurag.mandal@intel.com>

Recheck-request: iol-mellanox-Performance



^ permalink raw reply

* Re: [PATCH] maintainers: fix app/test attributions
From: Thomas Monjalon @ 2026-07-19  9:42 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: dev
In-Reply-To: <20260429162913.272259-1-stephen@networkplumber.org>

29/04/2026 18:29, Stephen Hemminger:
> Two app/test entries were attributed to the wrong section:
> 
> * test_event_ring.c exercises rte_event_ring
>   not the standalone app/test-eventdev/  binary.
>   Move it from "Eventdev test application" to "Eventdev API"
>   alongside test_eventdev.c, matching how test_event_eth_*_adapter.c and
>   test_event_*_adapter.c are grouped with their corresponding
>   lib/eventdev/*adapter* sub-libraries.
> 
> * The "test_*hash*" glob in the Hashes section unintentionally matched
>   three crypto hash test vector headers which belong only
>   to the Crypto API section. Replace with the two narrower globs
>   "test_hash*" and "test_thash*" which together cover every legitimate
>   hash test source (test_hash*.c and test_thash*.c) without leaking into
>   cryptodev.
> 
> Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>

Applied with fix lines, thanks.




^ permalink raw reply

* Re: [PATCH] MAINTAINERS: update for ena
From: Thomas Monjalon @ 2026-07-19  9:22 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: dev, shaibran, evgenys, amitbern, atrwajee
In-Reply-To: <20260620164001.129129-1-stephen@networkplumber.org>

20/06/2026 18:40, Stephen Hemminger:
> Email to Ron Beider <rbeider@amazon.com> bounced.
> I assume he is no longer available and other listed maintainers
> for ena driver will be enough.
> 
> Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>

Applied




^ permalink raw reply

* Re: [PATCH v1 1/3] updated i40e latest recommended matching list
From: Thomas Monjalon @ 2026-07-19  8:26 UTC (permalink / raw)
  To: Hailin Xu; +Cc: dev, bruce.richardson, manjunathgouda.hosamani
In-Reply-To: <20260717091233.12102-1-hailinx.xu@intel.com>

17/07/2026 11:12, Hailin Xu:
> Signed-off-by: Hailin Xu <hailinx.xu@intel.com>

Series applied, thanks.




^ permalink raw reply

* RE: [PATCH v2] net/enic: check notify set return value during init
From: Hyong Youb Kim (hyonkim) @ 2026-07-19  3:04 UTC (permalink / raw)
  To: Alexey Simakov
  Cc: John Daley (johndale), Neil Horman, Sujith Sankar, David Marchand,
	dev@dpdk.org, stable@dpdk.org
In-Reply-To: <20260717092259.vfwlgwahsquw3z2x@home-server>



> -----Original Message-----
> From: Alexey Simakov <bigalex934@gmail.com>
> Sent: Friday, July 17, 2026 6:23 PM
> To: Hyong Youb Kim (hyonkim) <hyonkim@cisco.com>
> Cc: John Daley (johndale) <johndale@cisco.com>; Neil Horman
> <nhorman@tuxdriver.com>; Sujith Sankar <ssujith@cisco.com>; David
> Marchand <david.marchand@redhat.com>; dev@dpdk.org; stable@dpdk.org
> Subject: Re: [PATCH v2] net/enic: check notify set return value during init
> 
> Sure, also dpdk ai review higlightet potential problems in memory
> management
> 
> For example:
> 
> enic->cq = rte_zmalloc("enic_vnic_cq", sizeof(struct vnic_cq) *
> 			       enic->conf_cq_count, 8);
> enic->intr = rte_zmalloc("enic_vnic_intr", sizeof(struct vnic_intr) *
> 			enic->conf_intr_count, 8);
> enic->rq = rte_zmalloc("enic_vnic_rq", sizeof(struct vnic_rq) *
> 			enic->conf_rq_count, 8);
> enic->wq = rte_zmalloc("enic_vnic_wq", sizeof(struct vnic_wq) *
> 			enic->conf_wq_count, 8);
> 
> if (enic->conf_cq_count > 0 && enic->cq == NULL) {
> 	dev_err(enic, "failed to allocate vnic_cq, aborting.\n");
> 	return -1;
> }
> if (enic->conf_intr_count > 0 && enic->intr == NULL) {
> 	dev_err(enic, "failed to allocate vnic_intr, aborting.\n");
> 	return -1;
> }
> 
> Lets imagine that enic->cq is allocated successfully, after that
> when we trying to allocate enic->intr we are failing, so in this
> case we leaking in second if statement, this related to all error
> handling paths in driver, since as far as I understand resource
> free happening only in enic_dev_deinit(...) function.
> 
> So my question is it problem? If yes, should I address this issue
> in scope of my patch or better to make it in another?

Why don't you create one cleanup patch that addresses both issues,
since they are in the same function?

Thanks.
-Hyong

^ permalink raw reply

* Re: [PATCH] doc: using container to build and run applications
From: Thomas Monjalon @ 2026-07-18 20:44 UTC (permalink / raw)
  To: Andrea Panattoni; +Cc: dev
In-Reply-To: <20260603140421.3023794-1-apanatto@redhat.com>

> +DPDK can be built inside a container to provide a reproducible build environment.

Changed "reproducible" to "isolated" as it is not always a reproducible build,
especially with "FROM fedora:latest".


> +Once the build completes, verify that the helloworld example runs:
> +
> +.. code-block:: console
> +
> +   podman run --rm dpdk-builder /dpdk/build/examples/dpdk-helloworld

Hugepages are not installed at this stage, so adding --no-huge is required.


They are small changes, done while merging after a quick test.

^ permalink raw reply

* Re: [PATCH] doc: using container to build and run applications
From: Thomas Monjalon @ 2026-07-18 20:22 UTC (permalink / raw)
  To: Andrea Panattoni; +Cc: dev
In-Reply-To: <20260603140421.3023794-1-apanatto@redhat.com>

03/06/2026 16:04, Andrea Panattoni:
> Add explanation about how container runtimes like podman
> or docker can be used to build and run DPDK application.
> 
> Signed-off-by: Andrea Panattoni <apanatto@redhat.com>
[...]
> +.. _building_dpdk_in_container:

Adding a prefix linux_gsg_ to this label.

> +
> +Building Applications in a Container
> +------------------------------------

Actually it should be "Building DPDK in a Container"


> +Running Sample Application in a Container
> +-----------------------------------------

It should be "Running an Application in a Container" as testpmd is not a sample app.

Applied with adjustments, thanks.



^ permalink raw reply


This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox