DPDK-dev Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs
@ 2026-07-20  8:00 Maxime Leroy
  2026-07-20  8:00 ` [PATCH 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path Maxime Leroy
                   ` (8 more replies)
  0 siblings, 9 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy

l3fwd-power's Rx interrupt (one-shot) mode was written around NICs whose
interrupt fd is assigned per queue at setup, stays fixed and is
level-triggered, as on MSI-X hardware. PMDs that do not fit that model
never sleep on interrupts: they either fail to arm, watch an unbound fd,
or miss the wake, and the 10 ms poll timeout hides all of it.

dpaa2 is the motivating example: it delivers Rx interrupts through one
QBMan portal per lcore, so the queues of a lcore share a single fd that
is bound when the interrupt is enabled, not at queue setup; the
notification is edge-triggered on an empty to non-empty transition; and
arming a queue that already holds traffic returns -EAGAIN with the queue
left unarmed.

This series makes the interrupt and legacy loops handle those cases
while leaving MSI-X NICs unchanged:

  1/6 factor the duplicated arm/sleep/disarm block into a helper.
  2/6 check the arm return: on -EAGAIN poll and retry, on any other
      error drop interrupt mode for the lcore instead of sleeping on
      unarmed queues.
  3/6 register the fd in the epoll set after arming, not before, so a
      fd bound at enable time is watched only once it exists.
  4/6 accept -EEXIST when several queues share one fd.
  5/6 recheck the queues before sleeping so a packet that arrived during
      the arm window is not missed under edge-triggered interrupts.
  6/6 block indefinitely instead of on a 10 ms timeout, and wake the
      workers through an eventfd on exit, so a broken interrupt path no
      longer hides behind periodic polling.

Maxime Leroy (6):
  examples/l3fwd-power: factor out Rx interrupt sleep path
  examples/l3fwd-power: check Rx interrupt enable errors
  examples/l3fwd-power: enable Rx interrupt before epoll add
  examples/l3fwd-power: accept shared Rx interrupt FD
  examples/l3fwd-power: recheck Rx queues before sleeping
  examples/l3fwd-power: block until Rx interrupt or exit

 examples/l3fwd-power/main.c | 184 +++++++++++++++++++++++++++++-------
 1 file changed, 150 insertions(+), 34 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 20+ messages in thread

* [PATCH 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
@ 2026-07-20  8:00 ` Maxime Leroy
  2026-07-20  8:00 ` [PATCH 2/6] examples/l3fwd-power: check Rx interrupt enable errors Maxime Leroy
                   ` (7 subsequent siblings)
  8 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy

The interrupt and legacy main loops carry an identical block that arms
the Rx queue interrupts of the lcore, sleeps until one triggers and
disarms them again. Move it to a helper so the following fixes touch a
single place.

No functional change.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 705cab8f2d..764899217c 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -910,6 +910,14 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
+static void
+rx_interrupt_wait(struct lcore_conf *qconf)
+{
+	turn_on_off_intr(qconf, 1);
+	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
+	turn_on_off_intr(qconf, 0);
+}
+
 /* Main processing loop. 8< */
 static int main_intr_loop(__rte_unused void *dummy)
 {
@@ -1054,11 +1062,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					turn_on_off_intr(qconf, 1);
-					sleep_until_rx_interrupt(
-							qconf->n_rx_queue,
-							lcore_id);
-					turn_on_off_intr(qconf, 0);
+					rx_interrupt_wait(qconf);
 					/**
 					 * start receiving packets immediately
 					 */
@@ -1372,11 +1376,7 @@ main_legacy_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					turn_on_off_intr(qconf, 1);
-					sleep_until_rx_interrupt(
-							qconf->n_rx_queue,
-							lcore_id);
-					turn_on_off_intr(qconf, 0);
+					rx_interrupt_wait(qconf);
 					/**
 					 * start receiving packets immediately
 					 */
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH 2/6] examples/l3fwd-power: check Rx interrupt enable errors
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
  2026-07-20  8:00 ` [PATCH 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path Maxime Leroy
@ 2026-07-20  8:00 ` Maxime Leroy
  2026-07-20  8:00 ` [PATCH 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add Maxime Leroy
                   ` (6 subsequent siblings)
  8 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy

rx_interrupt_wait() enabled the Rx queue interrupts of an lcore through
turn_on_off_intr() and ignored the return of
rte_eth_dev_rx_intr_enable(). Some PMDs report a transient condition
there: for example dpaa2 returns -EAGAIN when it finds traffic already
queued while arming, meaning the queue was not armed and the lcore must
poll rather than sleep.

Split the helper into rx_intr_enable_all(), which returns the error and
unwinds the queues it already armed, and rx_intr_disable_all(). Return
the error from rx_interrupt_wait(); in the sleep path, on -EAGAIN go back
to polling, and on any other error give up interrupt mode for the lcore
instead of sleeping on unarmed queues.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 74 ++++++++++++++++++++++++++++++++-----
 1 file changed, 64 insertions(+), 10 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 764899217c..70f76d202e 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -863,12 +863,32 @@ sleep_until_rx_interrupt(int num, int lcore)
 	return 0;
 }
 
-static void turn_on_off_intr(struct lcore_conf *qconf, bool on)
+static void
+rx_intr_disable_all(struct lcore_conf *qconf)
 {
+	struct lcore_rx_queue *rx_queue;
+	uint16_t queue_id;
+	uint16_t port_id;
 	int i;
+
+	for (i = 0; i < qconf->n_rx_queue; ++i) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		rte_spinlock_lock(&(locks[port_id]));
+		rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		rte_spinlock_unlock(&(locks[port_id]));
+	}
+}
+
+static int
+rx_intr_enable_all(struct lcore_conf *qconf)
+{
 	struct lcore_rx_queue *rx_queue;
 	uint16_t queue_id;
 	uint16_t port_id;
+	int i, ret;
 
 	for (i = 0; i < qconf->n_rx_queue; ++i) {
 		rx_queue = &(qconf->rx_queue_list[i]);
@@ -876,12 +896,26 @@ static void turn_on_off_intr(struct lcore_conf *qconf, bool on)
 		queue_id = rx_queue->queue_id;
 
 		rte_spinlock_lock(&(locks[port_id]));
-		if (on)
-			rte_eth_dev_rx_intr_enable(port_id, queue_id);
-		else
-			rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		ret = rte_eth_dev_rx_intr_enable(port_id, queue_id);
 		rte_spinlock_unlock(&(locks[port_id]));
+		if (ret != 0)
+			goto fail;
 	}
+
+	return 0;
+
+fail:
+	while (--i >= 0) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		rte_spinlock_lock(&(locks[port_id]));
+		rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		rte_spinlock_unlock(&(locks[port_id]));
+	}
+
+	return ret;
 }
 
 static int event_register(struct lcore_conf *qconf)
@@ -910,12 +944,18 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
-static void
+static int
 rx_interrupt_wait(struct lcore_conf *qconf)
 {
-	turn_on_off_intr(qconf, 1);
+	int ret;
+
+	ret = rx_intr_enable_all(qconf);
+	if (ret != 0)
+		return ret;
+
 	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
-	turn_on_off_intr(qconf, 0);
+	rx_intr_disable_all(qconf);
+	return 0;
 }
 
 /* Main processing loop. 8< */
@@ -931,6 +971,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
 	int intr_en = 0;
+	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) /
 				   US_PER_S * BURST_TX_DRAIN_US;
@@ -1062,7 +1103,13 @@ static int main_intr_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf);
+					if (ret == -EAGAIN)
+						goto start_rx;
+					if (ret != 0) {
+						intr_en = 0;
+						continue;
+					}
 					/**
 					 * start receiving packets immediately
 					 */
@@ -1214,6 +1261,7 @@ main_legacy_loop(__rte_unused void *dummy)
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
 	int intr_en = 0;
+	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) / US_PER_S * BURST_TX_DRAIN_US;
 
@@ -1376,7 +1424,13 @@ main_legacy_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf);
+					if (ret == -EAGAIN)
+						goto start_rx;
+					if (ret != 0) {
+						intr_en = 0;
+						continue;
+					}
 					/**
 					 * start receiving packets immediately
 					 */
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
  2026-07-20  8:00 ` [PATCH 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path Maxime Leroy
  2026-07-20  8:00 ` [PATCH 2/6] examples/l3fwd-power: check Rx interrupt enable errors Maxime Leroy
@ 2026-07-20  8:00 ` Maxime Leroy
  2026-09-04  9:15   ` David Marchand
  2026-07-20  8:00 ` [PATCH 4/6] examples/l3fwd-power: accept shared Rx interrupt FD Maxime Leroy
                   ` (5 subsequent siblings)
  8 siblings, 1 reply; 20+ messages in thread
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy

event_register() added each Rx queue interrupt fd to the per-thread epoll
set before the worker loop, while the queues were still disarmed. This
works when the fd is assigned at queue setup and stays fixed, as on MSI-X
NICs.

The dpaa2 PMD delivers Rx interrupts through a per-lcore QBMan portal and
binds the fd when the interrupt is enabled, not at queue setup. Adding a
queue to the epoll set before arming it then watches an fd that is not
bound yet, so the lcore never wakes from rte_epoll_wait().

Register from rx_interrupt_wait(), after the queues are armed, and only
once: keep the epoll entry installed for the lifetime of the loop. NICs
with a fixed per-queue fd are unaffected.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 34 +++++++++++++++++-----------------
 1 file changed, 17 insertions(+), 17 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 70f76d202e..4874a55de1 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -945,7 +945,7 @@ static int event_register(struct lcore_conf *qconf)
 }
 
 static int
-rx_interrupt_wait(struct lcore_conf *qconf)
+rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
 {
 	int ret;
 
@@ -953,6 +953,16 @@ rx_interrupt_wait(struct lcore_conf *qconf)
 	if (ret != 0)
 		return ret;
 
+	/* some PMDs expose the interrupt fd only once the queue is armed */
+	if (!*intr_registered) {
+		ret = event_register(qconf);
+		if (ret != 0) {
+			rx_intr_disable_all(qconf);
+			return ret;
+		}
+		*intr_registered = 1;
+	}
+
 	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
 	rx_intr_disable_all(qconf);
 	return 0;
@@ -970,7 +980,8 @@ static int main_intr_loop(__rte_unused void *dummy)
 	struct lcore_rx_queue *rx_queue;
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
-	int intr_en = 0;
+	int intr_registered = 0;
+	int intr_en = 1;
 	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) /
@@ -998,12 +1009,6 @@ static int main_intr_loop(__rte_unused void *dummy)
 				lcore_id, portid, queueid);
 	}
 
-	/* add into event wait list */
-	if (event_register(qconf) == 0)
-		intr_en = 1;
-	else
-		RTE_LOG(INFO, L3FWD_POWER, "RX interrupt won't enable.\n");
-
 	while (!is_done()) {
 		stats[lcore_id].nb_iteration_looped++;
 
@@ -1103,7 +1108,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					ret = rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf, &intr_registered);
 					if (ret == -EAGAIN)
 						goto start_rx;
 					if (ret != 0) {
@@ -1260,7 +1265,8 @@ main_legacy_loop(__rte_unused void *dummy)
 	enum freq_scale_hint_t lcore_scaleup_hint;
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
-	int intr_en = 0;
+	int intr_registered = 0;
+	int intr_en = 1;
 	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) / US_PER_S * BURST_TX_DRAIN_US;
@@ -1286,12 +1292,6 @@ main_legacy_loop(__rte_unused void *dummy)
 			"rxqueueid=%" PRIu16 "\n", lcore_id, portid, queueid);
 	}
 
-	/* add into event wait list */
-	if (event_register(qconf) == 0)
-		intr_en = 1;
-	else
-		RTE_LOG(INFO, L3FWD_POWER, "RX interrupt won't enable.\n");
-
 	while (!is_done()) {
 		stats[lcore_id].nb_iteration_looped++;
 
@@ -1424,7 +1424,7 @@ main_legacy_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					ret = rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf, &intr_registered);
 					if (ret == -EAGAIN)
 						goto start_rx;
 					if (ret != 0) {
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH 4/6] examples/l3fwd-power: accept shared Rx interrupt FD
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
                   ` (2 preceding siblings ...)
  2026-07-20  8:00 ` [PATCH 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add Maxime Leroy
@ 2026-07-20  8:00 ` Maxime Leroy
  2026-07-20  8:00 ` [PATCH 5/6] examples/l3fwd-power: recheck Rx queues before sleeping Maxime Leroy
                   ` (4 subsequent siblings)
  8 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy

Some PMDs deliver the Rx interrupts of several queues through a single
shared fd; dpaa2, for instance, uses one QBMan portal per lcore.
Registering the fd a second time then returns -EEXIST, which
event_register() treated as fatal and disabled interrupt mode for the
whole lcore.

Accept -EEXIST as success. NICs with a dedicated fd per queue never
return it, so their behaviour is unchanged.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 4874a55de1..fa64e28e5c 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -937,7 +937,12 @@ static int event_register(struct lcore_conf *qconf)
 						RTE_EPOLL_PER_THREAD,
 						RTE_INTR_EVENT_ADD,
 						(void *)((uintptr_t)data));
-		if (ret)
+		/*
+		 * Queues polled on the same lcore may share one interrupt fd,
+		 * so registering the second onward returns -EEXIST; treat it
+		 * as success.
+		 */
+		if (ret && ret != -EEXIST)
 			return ret;
 	}
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH 5/6] examples/l3fwd-power: recheck Rx queues before sleeping
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
                   ` (3 preceding siblings ...)
  2026-07-20  8:00 ` [PATCH 4/6] examples/l3fwd-power: accept shared Rx interrupt FD Maxime Leroy
@ 2026-07-20  8:00 ` Maxime Leroy
  2026-09-04  9:20   ` David Marchand
  2026-07-20  8:00 ` [PATCH 6/6] examples/l3fwd-power: block until Rx interrupt or exit Maxime Leroy
                   ` (3 subsequent siblings)
  8 siblings, 1 reply; 20+ messages in thread
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy

After arming the Rx interrupts rx_interrupt_wait() sleeps in
rte_epoll_wait(). A packet that arrived between the last poll and the arm
can be missed on PMDs whose interrupt is edge-triggered on an empty to
non-empty transition, such as dpaa2: arming a queue that is already
non-empty raises no notification, so the lcore sleeps until the next
packet.

Before sleeping, check whether any queue already holds traffic and, if
so, skip the wait and go back to polling. NICs with level-triggered
interrupts are unaffected: a pending packet keeps the interrupt asserted,
so rte_epoll_wait() would return immediately anyway.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 27 ++++++++++++++++++++++++++-
 1 file changed, 26 insertions(+), 1 deletion(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index fa64e28e5c..c423b0dba5 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -949,6 +949,26 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
+static bool
+rx_queue_pending(struct lcore_conf *qconf)
+{
+	struct lcore_rx_queue *rx_queue;
+	uint16_t queue_id;
+	uint16_t port_id;
+	int i;
+
+	for (i = 0; i < qconf->n_rx_queue; ++i) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		if (rte_eth_rx_queue_count(port_id, queue_id) > 0)
+			return true;
+	}
+
+	return false;
+}
+
 static int
 rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
 {
@@ -968,7 +988,12 @@ rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
 		*intr_registered = 1;
 	}
 
-	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
+	/*
+	 * A packet that arrived during the arm window raises no new wakeup on
+	 * edge-triggered PMDs, so skip the sleep if a queue already has traffic.
+	 */
+	if (!rx_queue_pending(qconf))
+		sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
 	rx_intr_disable_all(qconf);
 	return 0;
 }
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH 6/6] examples/l3fwd-power: block until Rx interrupt or exit
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
                   ` (4 preceding siblings ...)
  2026-07-20  8:00 ` [PATCH 5/6] examples/l3fwd-power: recheck Rx queues before sleeping Maxime Leroy
@ 2026-07-20  8:00 ` Maxime Leroy
  2026-08-25  8:07 ` [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
                   ` (2 subsequent siblings)
  8 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-07-20  8:00 UTC (permalink / raw)
  To: dev; +Cc: anatoly.burakov, sivaprasad.tummala, Maxime Leroy

In interrupt and legacy modes an idle worker arms its Rx queues and
sleeps in rte_epoll_wait() with a 10 ms timeout. The timeout lets the
worker wake periodically to re-check the quit flag, but it also hides a
broken interrupt path: with no interrupt delivered the worker still wakes
every 10 ms and polls, so traffic keeps flowing and interrupt mode looks
like it works when it does not.

Block indefinitely instead, so only a real Rx interrupt can wake the
worker. To still stop the workers on SIGINT, add an eventfd to each
worker epoll set; the signal handler writes it to wake every worker,
which then observes the quit flag and leaves its loop. Both loops share
this sleep path, so the eventfd is created for the interrupt and legacy
modes; it is skipped in the wakeup log so it is not reported as an Rx
interrupt.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 40 +++++++++++++++++++++++++++++++++----
 1 file changed, 36 insertions(+), 4 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index c423b0dba5..73ee2d2def 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -9,6 +9,8 @@
 #include <sys/types.h>
 #include <string.h>
 #include <sys/queue.h>
+#include <sys/epoll.h>
+#include <sys/eventfd.h>
 #include <stdarg.h>
 #include <errno.h>
 #include <getopt.h>
@@ -143,6 +145,9 @@ static int promiscuous_on = 0;
 /* NUMA is enabled by default. */
 static int numa_on = 1;
 volatile bool quit_signal;
+/* eventfd to wake idle workers out of a blocking rte_epoll_wait() at exit */
+static int wakeup_fd = -1;
+static struct rte_epoll_event wakeup_event[RTE_MAX_LCORE];
 /* timer to update telemetry every 500ms */
 static struct rte_timer telemetry_timer;
 
@@ -420,10 +425,17 @@ static int is_done(void)
 static void
 signal_exit_now(int sigtype)
 {
+	uint64_t one = 1;
+	ssize_t nbytes;
 
-	if (sigtype == SIGINT)
+	if (sigtype == SIGINT) {
 		quit_signal = true;
-
+		/* interrupt mode only: wake a worker blocked in rte_epoll_wait() */
+		if (wakeup_fd >= 0) {
+			nbytes = write(wakeup_fd, &one, sizeof(one));
+			RTE_SET_USED(nbytes);
+		}
+	}
 }
 
 /*  Frequency scale down timer callback */
@@ -835,7 +847,7 @@ sleep_until_rx_interrupt(int num, int lcore)
 	static alignas(RTE_CACHE_LINE_SIZE) struct {
 		bool wakeup;
 	} status[RTE_MAX_LCORE];
-	struct rte_epoll_event event[num];
+	struct rte_epoll_event event[num + 1];
 	int n, i;
 	uint16_t port_id;
 	uint16_t queue_id;
@@ -847,8 +859,10 @@ sleep_until_rx_interrupt(int num, int lcore)
 				rte_lcore_id());
 	}
 
-	n = rte_epoll_wait(RTE_EPOLL_PER_THREAD, event, num, 10);
+	n = rte_epoll_wait(RTE_EPOLL_PER_THREAD, event, num + 1, -1);
 	for (i = 0; i < n; i++) {
+		if (event[i].fd == wakeup_fd)
+			continue;
 		data = event[i].epdata.data;
 		port_id = ((uintptr_t)data) >> (sizeof(uint16_t) * CHAR_BIT);
 		queue_id = ((uintptr_t)data) &
@@ -921,6 +935,7 @@ rx_intr_enable_all(struct lcore_conf *qconf)
 static int event_register(struct lcore_conf *qconf)
 {
 	struct lcore_rx_queue *rx_queue;
+	struct rte_epoll_event *wev;
 	uint16_t queueid;
 	uint16_t portid;
 	uint32_t data;
@@ -946,6 +961,16 @@ static int event_register(struct lcore_conf *qconf)
 			return ret;
 	}
 
+	/* also watch the shutdown eventfd so a blocking wait wakes at exit */
+	if (wakeup_fd >= 0) {
+		wev = &wakeup_event[rte_lcore_id()];
+		wev->epdata.event = EPOLLIN | EPOLLPRI;
+		ret = rte_epoll_ctl(RTE_EPOLL_PER_THREAD, EPOLL_CTL_ADD,
+				    wakeup_fd, wev);
+		if (ret && ret != -EEXIST)
+			return ret;
+	}
+
 	return 0;
 }
 
@@ -2989,6 +3014,13 @@ main(int argc, char **argv)
 
 	check_all_ports_link_status(enabled_port_mask);
 
+	if (app_mode == APP_MODE_LEGACY || app_mode == APP_MODE_INTERRUPT) {
+		wakeup_fd = eventfd(0, EFD_NONBLOCK);
+		if (wakeup_fd < 0)
+			rte_exit(EXIT_FAILURE, "eventfd() failed: %s\n",
+				 strerror(errno));
+	}
+
 	/* launch per-lcore init on every lcore */
 	if (app_mode == APP_MODE_LEGACY) {
 		rte_eal_mp_remote_launch(main_legacy_loop, NULL, CALL_MAIN);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* Re: [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
                   ` (5 preceding siblings ...)
  2026-07-20  8:00 ` [PATCH 6/6] examples/l3fwd-power: block until Rx interrupt or exit Maxime Leroy
@ 2026-08-25  8:07 ` Maxime Leroy
  2026-09-04  9:30 ` David Marchand
  2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
  8 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-08-25  8:07 UTC (permalink / raw)
  To: anatoly.burakov, sivaprasad.tummala; +Cc: dev

Hi Anatoly, hi Sivaprasad,

On Mon, Jul 20, 2026 at 10:00 AM Maxime Leroy <maxime@leroys.fr> wrote:
>
> l3fwd-power's Rx interrupt (one-shot) mode was written around NICs whose
> interrupt fd is assigned per queue at setup, stays fixed and is
> level-triggered, as on MSI-X hardware. PMDs that do not fit that model
> never sleep on interrupts: they either fail to arm, watch an unbound fd,
> or miss the wake, and the 10 ms poll timeout hides all of it.
>
> dpaa2 is the motivating example: it delivers Rx interrupts through one
> QBMan portal per lcore, so the queues of a lcore share a single fd that
> is bound when the interrupt is enabled, not at queue setup; the
> notification is edge-triggered on an empty to non-empty transition; and
> arming a queue that already holds traffic returns -EAGAIN with the queue
> left unarmed.
>
> This series makes the interrupt and legacy loops handle those cases
> while leaving MSI-X NICs unchanged:
>
>   1/6 factor the duplicated arm/sleep/disarm block into a helper.
>   2/6 check the arm return: on -EAGAIN poll and retry, on any other
>       error drop interrupt mode for the lcore instead of sleeping on
>       unarmed queues.
>   3/6 register the fd in the epoll set after arming, not before, so a
>       fd bound at enable time is watched only once it exists.
>   4/6 accept -EEXIST when several queues share one fd.
>   5/6 recheck the queues before sleeping so a packet that arrived during
>       the arm window is not missed under edge-triggered interrupts.
>   6/6 block indefinitely instead of on a 10 ms timeout, and wake the
>       workers through an eventfd on exit, so a broken interrupt path no
>       longer hides behind periodic polling.
>
> Maxime Leroy (6):
>   examples/l3fwd-power: factor out Rx interrupt sleep path
>   examples/l3fwd-power: check Rx interrupt enable errors
>   examples/l3fwd-power: enable Rx interrupt before epoll add
>   examples/l3fwd-power: accept shared Rx interrupt FD
>   examples/l3fwd-power: recheck Rx queues before sleeping
>   examples/l3fwd-power: block until Rx interrupt or exit
>
>  examples/l3fwd-power/main.c | 184 +++++++++++++++++++++++++++++-------
>  1 file changed, 150 insertions(+), 34 deletions(-)
>
> --
> 2.43.0
>

Gentle ping on this series.

Thanks,

Maxime

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add
  2026-07-20  8:00 ` [PATCH 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add Maxime Leroy
@ 2026-09-04  9:15   ` David Marchand
  0 siblings, 0 replies; 20+ messages in thread
From: David Marchand @ 2026-09-04  9:15 UTC (permalink / raw)
  To: Maxime Leroy; +Cc: dev, anatoly.burakov, sivaprasad.tummala

Hello Maxime,

On Mon, 20 Jul 2026 at 10:01, Maxime Leroy <maxime@leroys.fr> wrote:
>
> event_register() added each Rx queue interrupt fd to the per-thread epoll
> set before the worker loop, while the queues were still disarmed. This
> works when the fd is assigned at queue setup and stays fixed, as on MSI-X
> NICs.
>
> The dpaa2 PMD delivers Rx interrupts through a per-lcore QBMan portal and
> binds the fd when the interrupt is enabled, not at queue setup. Adding a
> queue to the epoll set before arming it then watches an fd that is not
> bound yet, so the lcore never wakes from rte_epoll_wait().
>
> Register from rx_interrupt_wait(), after the queues are armed, and only
> once: keep the epoll entry installed for the lifetime of the loop. NICs
> with a fixed per-queue fd are unaffected.
>
> Signed-off-by: Maxime Leroy <maxime@leroys.fr>
> ---
>  examples/l3fwd-power/main.c | 34 +++++++++++++++++-----------------
>  1 file changed, 17 insertions(+), 17 deletions(-)
>
> diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
> index 70f76d202e..4874a55de1 100644
> --- a/examples/l3fwd-power/main.c
> +++ b/examples/l3fwd-power/main.c
> @@ -945,7 +945,7 @@ static int event_register(struct lcore_conf *qconf)
>  }
>
>  static int
> -rx_interrupt_wait(struct lcore_conf *qconf)
> +rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)

On the principle, passing intr_registered is ugly.

intr_registered is an internal state that does not need to be exported
in the caller.
Maybe move this to struct lcore_conf?

This change also drops a log that flagged that Rx interrupts were not
available, please restore it (maybe add a "once" boolean to only raise
the log on the first error).

>  {
>         int ret;
>

-- 
David Marchand


^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH 5/6] examples/l3fwd-power: recheck Rx queues before sleeping
  2026-07-20  8:00 ` [PATCH 5/6] examples/l3fwd-power: recheck Rx queues before sleeping Maxime Leroy
@ 2026-09-04  9:20   ` David Marchand
  2026-09-04 20:25     ` Maxime Leroy
  0 siblings, 1 reply; 20+ messages in thread
From: David Marchand @ 2026-09-04  9:20 UTC (permalink / raw)
  To: Maxime Leroy; +Cc: dev, anatoly.burakov, sivaprasad.tummala

On Mon, 20 Jul 2026 at 10:01, Maxime Leroy <maxime@leroys.fr> wrote:
>
> After arming the Rx interrupts rx_interrupt_wait() sleeps in
> rte_epoll_wait(). A packet that arrived between the last poll and the arm
> can be missed on PMDs whose interrupt is edge-triggered on an empty to
> non-empty transition, such as dpaa2: arming a queue that is already
> non-empty raises no notification, so the lcore sleeps until the next
> packet.
>
> Before sleeping, check whether any queue already holds traffic and, if
> so, skip the wait and go back to polling. NICs with level-triggered
> interrupts are unaffected: a pending packet keeps the interrupt asserted,
> so rte_epoll_wait() would return immediately anyway.
>
> Signed-off-by: Maxime Leroy <maxime@leroys.fr>
> ---
>  examples/l3fwd-power/main.c | 27 ++++++++++++++++++++++++++-
>  1 file changed, 26 insertions(+), 1 deletion(-)
>
> diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
> index fa64e28e5c..c423b0dba5 100644
> --- a/examples/l3fwd-power/main.c
> +++ b/examples/l3fwd-power/main.c
> @@ -949,6 +949,26 @@ static int event_register(struct lcore_conf *qconf)
>         return 0;
>  }
>
> +static bool
> +rx_queue_pending(struct lcore_conf *qconf)
> +{
> +       struct lcore_rx_queue *rx_queue;
> +       uint16_t queue_id;
> +       uint16_t port_id;
> +       int i;
> +
> +       for (i = 0; i < qconf->n_rx_queue; ++i) {
> +               rx_queue = &(qconf->rx_queue_list[i]);
> +               port_id = rx_queue->port_id;
> +               queue_id = rx_queue->queue_id;
> +
> +               if (rte_eth_rx_queue_count(port_id, queue_id) > 0)
> +                       return true;
> +       }
> +
> +       return false;
> +}
> +
>  static int
>  rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
>  {
> @@ -968,7 +988,12 @@ rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
>                 *intr_registered = 1;
>         }
>
> -       sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
> +       /*
> +        * A packet that arrived during the arm window raises no new wakeup on
> +        * edge-triggered PMDs, so skip the sleep if a queue already has traffic.
> +        */
> +       if (!rx_queue_pending(qconf))
> +               sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
>         rx_intr_disable_all(qconf);
>         return 0;
>  }

I don't see how this solves the race.
There is stil a window between checking the queue count and sleeping.

Looks like a bug that should be solved in the driver.


-- 
David Marchand


^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
                   ` (6 preceding siblings ...)
  2026-08-25  8:07 ` [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
@ 2026-09-04  9:30 ` David Marchand
  2026-09-04 21:10   ` Maxime Leroy
  2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
  8 siblings, 1 reply; 20+ messages in thread
From: David Marchand @ 2026-09-04  9:30 UTC (permalink / raw)
  To: Maxime Leroy, anatoly.burakov, sivaprasad.tummala; +Cc: dev, Mcnamara, John

On Mon, 20 Jul 2026 at 10:01, Maxime Leroy <maxime@leroys.fr> wrote:
>
> l3fwd-power's Rx interrupt (one-shot) mode was written around NICs whose
> interrupt fd is assigned per queue at setup, stays fixed and is
> level-triggered, as on MSI-X hardware. PMDs that do not fit that model
> never sleep on interrupts: they either fail to arm, watch an unbound fd,
> or miss the wake, and the 10 ms poll timeout hides all of it.
>
> dpaa2 is the motivating example: it delivers Rx interrupts through one
> QBMan portal per lcore, so the queues of a lcore share a single fd that
> is bound when the interrupt is enabled, not at queue setup; the
> notification is edge-triggered on an empty to non-empty transition; and
> arming a queue that already holds traffic returns -EAGAIN with the queue
> left unarmed.
>
> This series makes the interrupt and legacy loops handle those cases
> while leaving MSI-X NICs unchanged:
>
>   1/6 factor the duplicated arm/sleep/disarm block into a helper.
>   2/6 check the arm return: on -EAGAIN poll and retry, on any other
>       error drop interrupt mode for the lcore instead of sleeping on
>       unarmed queues.
>   3/6 register the fd in the epoll set after arming, not before, so a
>       fd bound at enable time is watched only once it exists.
>   4/6 accept -EEXIST when several queues share one fd.
>   5/6 recheck the queues before sleeping so a packet that arrived during
>       the arm window is not missed under edge-triggered interrupts.
>   6/6 block indefinitely instead of on a 10 ms timeout, and wake the
>       workers through an eventfd on exit, so a broken interrupt path no
>       longer hides behind periodic polling.
>
> Maxime Leroy (6):
>   examples/l3fwd-power: factor out Rx interrupt sleep path
>   examples/l3fwd-power: check Rx interrupt enable errors
>   examples/l3fwd-power: enable Rx interrupt before epoll add
>   examples/l3fwd-power: accept shared Rx interrupt FD
>   examples/l3fwd-power: recheck Rx queues before sleeping
>   examples/l3fwd-power: block until Rx interrupt or exit
>
>  examples/l3fwd-power/main.c | 184 +++++++++++++++++++++++++++++-------
>  1 file changed, 150 insertions(+), 34 deletions(-)

I had a quick look.
I am not a fan of changing applications because drivers behave
differently, but I think most of those changes are acceptable.

I sent one comment on the implementation of patch 3.
Patch 5 is a concern to me, as I don't think it solves anything, just
make the issue harder to reproduce maybe?

Anatoly, Sivaprasad, please review.


-- 
David Marchand


^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH 5/6] examples/l3fwd-power: recheck Rx queues before sleeping
  2026-09-04  9:20   ` David Marchand
@ 2026-09-04 20:25     ` Maxime Leroy
  0 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-04 20:25 UTC (permalink / raw)
  To: David Marchand; +Cc: dev, anatoly.burakov, sivaprasad.tummala

Hi David,

On Fri, Sep 4, 2026 at 11:20 AM David Marchand
<david.marchand@redhat.com> wrote:
>
> On Mon, 20 Jul 2026 at 10:01, Maxime Leroy <maxime@leroys.fr> wrote:
> >
> > After arming the Rx interrupts rx_interrupt_wait() sleeps in
> > rte_epoll_wait(). A packet that arrived between the last poll and the arm
> > can be missed on PMDs whose interrupt is edge-triggered on an empty to
> > non-empty transition, such as dpaa2: arming a queue that is already
> > non-empty raises no notification, so the lcore sleeps until the next
> > packet.
> >
> > Before sleeping, check whether any queue already holds traffic and, if
> > so, skip the wait and go back to polling. NICs with level-triggered
> > interrupts are unaffected: a pending packet keeps the interrupt asserted,
> > so rte_epoll_wait() would return immediately anyway.
> >
> > Signed-off-by: Maxime Leroy <maxime@leroys.fr>
> > ---
> >  examples/l3fwd-power/main.c | 27 ++++++++++++++++++++++++++-
> >  1 file changed, 26 insertions(+), 1 deletion(-)
> >
> > diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
> > index fa64e28e5c..c423b0dba5 100644
> > --- a/examples/l3fwd-power/main.c
> > +++ b/examples/l3fwd-power/main.c
> > @@ -949,6 +949,26 @@ static int event_register(struct lcore_conf *qconf)
> >         return 0;
> >  }
> >
> > +static bool
> > +rx_queue_pending(struct lcore_conf *qconf)
> > +{
> > +       struct lcore_rx_queue *rx_queue;
> > +       uint16_t queue_id;
> > +       uint16_t port_id;
> > +       int i;
> > +
> > +       for (i = 0; i < qconf->n_rx_queue; ++i) {
> > +               rx_queue = &(qconf->rx_queue_list[i]);
> > +               port_id = rx_queue->port_id;
> > +               queue_id = rx_queue->queue_id;
> > +
> > +               if (rte_eth_rx_queue_count(port_id, queue_id) > 0)
> > +                       return true;
> > +       }
> > +
> > +       return false;
> > +}
> > +
> >  static int
> >  rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
> >  {
> > @@ -968,7 +988,12 @@ rx_interrupt_wait(struct lcore_conf *qconf, int *intr_registered)
> >                 *intr_registered = 1;
> >         }
> >
> > -       sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
> > +       /*
> > +        * A packet that arrived during the arm window raises no new wakeup on
> > +        * edge-triggered PMDs, so skip the sleep if a queue already has traffic.
> > +        */
> > +       if (!rx_queue_pending(qconf))
> > +               sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
> >         rx_intr_disable_all(qconf);
> >         return 0;
> >  }
>
> I don't see how this solves the race.
> There is stil a window between checking the queue count and sleeping.

The check runs after the arm, not before:

        rx_intr_enable_all(qconf);              /* arm */
        if (!rx_queue_pending(qconf))           /* check */
                sleep_until_rx_interrupt();     /* sleep */

A packet arriving in the window you point at arrives with the interrupt
armed: the event is raised and latched in the event fd, so rte_epoll_wait()
returns immediately. Only a packet that arrived before the arm can be lost,
and that is what the check catches.

>
> Looks like a bug that should be solved in the driver.

This is not a dpaa2 workaround. The loop currently sleeps with a 10 ms
timeout, so a missed wakeup costs one poll round; patch 6/6 makes the wait
blocking, and from then on it is a permanent stall. This patch is what makes
that safe.

rte_ethdev.h does not define whether rx_queue_intr_enable() must raise when
the queue is already non-empty, and PMDs differ: mlx5 arms the CQ for the
next completion (mlx5_arm_cq(), the model whose documentation tells you to
re-poll after arming), ice sets GLINT_DYN_CTL_CLEARPBA_M while re-enabling,
and virtio just clears VRING_AVAIL_F_NO_INTERRUPT so the backend notifies
only on the next used buffer. I have only tested dpaa2, so correct me on the
first two if the hardware does re-raise.

Fixing this in drivers means defining that contract in the API first and
auditing every PMD that implements rx_queue_intr_enable.

>
>
> --
> David Marchand
>

--
Maxime Leroy

^ permalink raw reply	[flat|nested] 20+ messages in thread

* Re: [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs
  2026-09-04  9:30 ` David Marchand
@ 2026-09-04 21:10   ` Maxime Leroy
  0 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-04 21:10 UTC (permalink / raw)
  To: David Marchand; +Cc: anatoly.burakov, sivaprasad.tummala, dev, Mcnamara, John

On Fri, Sep 4, 2026 at 11:30 AM David Marchand
<david.marchand@redhat.com> wrote:
>
> On Mon, 20 Jul 2026 at 10:01, Maxime Leroy <maxime@leroys.fr> wrote:
> >
> > l3fwd-power's Rx interrupt (one-shot) mode was written around NICs whose
> > interrupt fd is assigned per queue at setup, stays fixed and is
> > level-triggered, as on MSI-X hardware. PMDs that do not fit that model
> > never sleep on interrupts: they either fail to arm, watch an unbound fd,
> > or miss the wake, and the 10 ms poll timeout hides all of it.
> >
> > dpaa2 is the motivating example: it delivers Rx interrupts through one
> > QBMan portal per lcore, so the queues of a lcore share a single fd that
> > is bound when the interrupt is enabled, not at queue setup; the
> > notification is edge-triggered on an empty to non-empty transition; and
> > arming a queue that already holds traffic returns -EAGAIN with the queue
> > left unarmed.
> >
> > This series makes the interrupt and legacy loops handle those cases
> > while leaving MSI-X NICs unchanged:
> >
> >   1/6 factor the duplicated arm/sleep/disarm block into a helper.
> >   2/6 check the arm return: on -EAGAIN poll and retry, on any other
> >       error drop interrupt mode for the lcore instead of sleeping on
> >       unarmed queues.
> >   3/6 register the fd in the epoll set after arming, not before, so a
> >       fd bound at enable time is watched only once it exists.
> >   4/6 accept -EEXIST when several queues share one fd.
> >   5/6 recheck the queues before sleeping so a packet that arrived during
> >       the arm window is not missed under edge-triggered interrupts.
> >   6/6 block indefinitely instead of on a 10 ms timeout, and wake the
> >       workers through an eventfd on exit, so a broken interrupt path no
> >       longer hides behind periodic polling.
> >
> > Maxime Leroy (6):
> >   examples/l3fwd-power: factor out Rx interrupt sleep path
> >   examples/l3fwd-power: check Rx interrupt enable errors
> >   examples/l3fwd-power: enable Rx interrupt before epoll add
> >   examples/l3fwd-power: accept shared Rx interrupt FD
> >   examples/l3fwd-power: recheck Rx queues before sleeping
> >   examples/l3fwd-power: block until Rx interrupt or exit
> >
> >  examples/l3fwd-power/main.c | 184 +++++++++++++++++++++++++++++-------
> >  1 file changed, 150 insertions(+), 34 deletions(-)
>
> I had a quick look.
> I am not a fan of changing applications because drivers behave
> differently, but I think most of those changes are acceptable.

What differs here is not driver behaviour, it is where the interrupt fd
comes from.

On MSI-X NICs there is one fd per Rx queue, created at queue setup, and it
never changes. l3fwd-power is built on that: it adds every queue fd to its
epoll set once, before the loop, then only arms and sleeps.

On dpaa2 the interrupt arrives on the QBMan portal of the lcore, and a
portal has a single fd. So the fd belongs to the lcore, not to the queue,
and which fd a queue uses is decided when a lcore arms it, because that is
when the driver learns which lcore polls that queue.

3/6 and 4/6 are exactly these two consequences:

  - before the first arm, the queue has no fd yet, so the epoll
    registration has to happen after the arm, not before;
  - all the queues of one lcore share that single fd, so registering the
    second one returns -EEXIST, which the application treats as an error.

The ethdev API never promised one fd per queue, the application assumed it,
and no driver can fix that from its side since the epoll set belongs to the
application.

Where a driver can help, dpaa2 already does: it returns -EAGAIN when the
queue could not be armed, instead of pretending it was. That only works if
the caller checks, and today turn_on_off_intr() ignores the return value and
sleeps anyway. That is 2/6.

>
> I sent one comment on the implementation of patch 3.

I will fix it in V2 thanks.

> Patch 5 is a concern to me, as I don't think it solves anything, just
> make the issue harder to reproduce maybe?

Patch 5 answered in its own thread.

>
> Anatoly, Sivaprasad, please review.
>
>
> --
> David Marchand
>

--
Maxime Leroy

^ permalink raw reply	[flat|nested] 20+ messages in thread

* [PATCH v2 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs
  2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
                   ` (7 preceding siblings ...)
  2026-09-04  9:30 ` David Marchand
@ 2026-09-09 15:04 ` Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path Maxime Leroy
                     ` (5 more replies)
  8 siblings, 6 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-09 15:04 UTC (permalink / raw)
  To: dev; +Cc: david.marchand, anatoly.burakov, sivaprasad.tummala, Maxime Leroy

l3fwd-power's Rx interrupt (one-shot) mode was written around NICs whose
interrupt fd is assigned per queue at setup, stays fixed and is
level-triggered, as on MSI-X hardware. PMDs that do not fit that model
never sleep on interrupts: they either fail to arm, watch an unbound fd,
or miss the wake, and the 10 ms poll timeout hides all of it.

dpaa2 is the motivating example: it delivers Rx interrupts through one
QBMan portal per lcore, so the queues of a lcore share a single fd that
is bound when the interrupt is enabled, not at queue setup; the
notification is edge-triggered on an empty to non-empty transition; and
arming a queue that already holds traffic returns -EAGAIN with the queue
left unarmed.

This series makes the interrupt and legacy loops handle those cases
while leaving MSI-X NICs unchanged:

  1/6 factor the duplicated arm/sleep/disarm block into a helper.
  2/6 check the arm return: on -EAGAIN poll and retry, on any other
      error drop interrupt mode for the lcore instead of sleeping on
      unarmed queues.
  3/6 register the fd in the epoll set after arming, not before, so a
      fd bound at enable time is watched only once it exists.
  4/6 accept -EEXIST when several queues share one fd.
  5/6 recheck the queues before sleeping so a packet that arrived during
      the arm window is not missed under edge-triggered interrupts.
  6/6 block indefinitely instead of on a 10 ms timeout, and wake the
      workers through an eventfd on exit, so a broken interrupt path no
      longer hides behind periodic polling.

v2:
  - 3/6: keep the registration state in struct lcore_conf instead of
    passing it to rx_interrupt_wait(), and restore the "RX interrupt
    won't enable." notice, now logged from the arm error path
    (David Marchand)

Maxime Leroy (6):
  examples/l3fwd-power: factor out Rx interrupt sleep path
  examples/l3fwd-power: check Rx interrupt enable errors
  examples/l3fwd-power: enable Rx interrupt before epoll add
  examples/l3fwd-power: accept shared Rx interrupt FD
  examples/l3fwd-power: recheck Rx queues before sleeping
  examples/l3fwd-power: block until Rx interrupt or exit

 examples/l3fwd-power/main.c | 187 +++++++++++++++++++++++++++++-------
 1 file changed, 153 insertions(+), 34 deletions(-)


base-commit: d55ccd4e6de64e3f797f60de9e81f1d60f849775
-- 
2.43.0


^ permalink raw reply	[flat|nested] 20+ messages in thread

* [PATCH v2 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path
  2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
@ 2026-09-09 15:04   ` Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 2/6] examples/l3fwd-power: check Rx interrupt enable errors Maxime Leroy
                     ` (4 subsequent siblings)
  5 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-09 15:04 UTC (permalink / raw)
  To: dev; +Cc: david.marchand, anatoly.burakov, sivaprasad.tummala, Maxime Leroy

The interrupt and legacy main loops carry an identical block that arms
the Rx queue interrupts of the lcore, sleeps until one triggers and
disarms them again. Move it to a helper so the following fixes touch a
single place.

No functional change.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 705cab8f2d..764899217c 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -910,6 +910,14 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
+static void
+rx_interrupt_wait(struct lcore_conf *qconf)
+{
+	turn_on_off_intr(qconf, 1);
+	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
+	turn_on_off_intr(qconf, 0);
+}
+
 /* Main processing loop. 8< */
 static int main_intr_loop(__rte_unused void *dummy)
 {
@@ -1054,11 +1062,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					turn_on_off_intr(qconf, 1);
-					sleep_until_rx_interrupt(
-							qconf->n_rx_queue,
-							lcore_id);
-					turn_on_off_intr(qconf, 0);
+					rx_interrupt_wait(qconf);
 					/**
 					 * start receiving packets immediately
 					 */
@@ -1372,11 +1376,7 @@ main_legacy_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					turn_on_off_intr(qconf, 1);
-					sleep_until_rx_interrupt(
-							qconf->n_rx_queue,
-							lcore_id);
-					turn_on_off_intr(qconf, 0);
+					rx_interrupt_wait(qconf);
 					/**
 					 * start receiving packets immediately
 					 */
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v2 2/6] examples/l3fwd-power: check Rx interrupt enable errors
  2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path Maxime Leroy
@ 2026-09-09 15:04   ` Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add Maxime Leroy
                     ` (3 subsequent siblings)
  5 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-09 15:04 UTC (permalink / raw)
  To: dev; +Cc: david.marchand, anatoly.burakov, sivaprasad.tummala, Maxime Leroy

rx_interrupt_wait() enabled the Rx queue interrupts of an lcore through
turn_on_off_intr() and ignored the return of
rte_eth_dev_rx_intr_enable(). Some PMDs report a transient condition
there: for example dpaa2 returns -EAGAIN when it finds traffic already
queued while arming, meaning the queue was not armed and the lcore must
poll rather than sleep.

Split the helper into rx_intr_enable_all(), which returns the error and
unwinds the queues it already armed, and rx_intr_disable_all(). Return
the error from rx_interrupt_wait(); in the sleep path, on -EAGAIN go back
to polling, and on any other error give up interrupt mode for the lcore
instead of sleeping on unarmed queues.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 74 ++++++++++++++++++++++++++++++++-----
 1 file changed, 64 insertions(+), 10 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 764899217c..70f76d202e 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -863,12 +863,32 @@ sleep_until_rx_interrupt(int num, int lcore)
 	return 0;
 }
 
-static void turn_on_off_intr(struct lcore_conf *qconf, bool on)
+static void
+rx_intr_disable_all(struct lcore_conf *qconf)
 {
+	struct lcore_rx_queue *rx_queue;
+	uint16_t queue_id;
+	uint16_t port_id;
 	int i;
+
+	for (i = 0; i < qconf->n_rx_queue; ++i) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		rte_spinlock_lock(&(locks[port_id]));
+		rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		rte_spinlock_unlock(&(locks[port_id]));
+	}
+}
+
+static int
+rx_intr_enable_all(struct lcore_conf *qconf)
+{
 	struct lcore_rx_queue *rx_queue;
 	uint16_t queue_id;
 	uint16_t port_id;
+	int i, ret;
 
 	for (i = 0; i < qconf->n_rx_queue; ++i) {
 		rx_queue = &(qconf->rx_queue_list[i]);
@@ -876,12 +896,26 @@ static void turn_on_off_intr(struct lcore_conf *qconf, bool on)
 		queue_id = rx_queue->queue_id;
 
 		rte_spinlock_lock(&(locks[port_id]));
-		if (on)
-			rte_eth_dev_rx_intr_enable(port_id, queue_id);
-		else
-			rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		ret = rte_eth_dev_rx_intr_enable(port_id, queue_id);
 		rte_spinlock_unlock(&(locks[port_id]));
+		if (ret != 0)
+			goto fail;
 	}
+
+	return 0;
+
+fail:
+	while (--i >= 0) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		rte_spinlock_lock(&(locks[port_id]));
+		rte_eth_dev_rx_intr_disable(port_id, queue_id);
+		rte_spinlock_unlock(&(locks[port_id]));
+	}
+
+	return ret;
 }
 
 static int event_register(struct lcore_conf *qconf)
@@ -910,12 +944,18 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
-static void
+static int
 rx_interrupt_wait(struct lcore_conf *qconf)
 {
-	turn_on_off_intr(qconf, 1);
+	int ret;
+
+	ret = rx_intr_enable_all(qconf);
+	if (ret != 0)
+		return ret;
+
 	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
-	turn_on_off_intr(qconf, 0);
+	rx_intr_disable_all(qconf);
+	return 0;
 }
 
 /* Main processing loop. 8< */
@@ -931,6 +971,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
 	int intr_en = 0;
+	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) /
 				   US_PER_S * BURST_TX_DRAIN_US;
@@ -1062,7 +1103,13 @@ static int main_intr_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf);
+					if (ret == -EAGAIN)
+						goto start_rx;
+					if (ret != 0) {
+						intr_en = 0;
+						continue;
+					}
 					/**
 					 * start receiving packets immediately
 					 */
@@ -1214,6 +1261,7 @@ main_legacy_loop(__rte_unused void *dummy)
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
 	int intr_en = 0;
+	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) / US_PER_S * BURST_TX_DRAIN_US;
 
@@ -1376,7 +1424,13 @@ main_legacy_loop(__rte_unused void *dummy)
 			else {
 				/* suspend until rx interrupt triggers */
 				if (intr_en) {
-					rx_interrupt_wait(qconf);
+					ret = rx_interrupt_wait(qconf);
+					if (ret == -EAGAIN)
+						goto start_rx;
+					if (ret != 0) {
+						intr_en = 0;
+						continue;
+					}
 					/**
 					 * start receiving packets immediately
 					 */
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v2 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add
  2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 2/6] examples/l3fwd-power: check Rx interrupt enable errors Maxime Leroy
@ 2026-09-09 15:04   ` Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 4/6] examples/l3fwd-power: accept shared Rx interrupt FD Maxime Leroy
                     ` (2 subsequent siblings)
  5 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-09 15:04 UTC (permalink / raw)
  To: dev; +Cc: david.marchand, anatoly.burakov, sivaprasad.tummala, Maxime Leroy

event_register() added each Rx queue interrupt fd to the per-thread epoll
set before the worker loop, while the queues were still disarmed. This
works when the fd is assigned at queue setup and stays fixed, as on MSI-X
NICs.

The dpaa2 PMD delivers Rx interrupts through a per-lcore QBMan portal and
binds the fd when the interrupt is enabled, not at queue setup. Adding a
queue to the epoll set before arming it then watches an fd that is not
bound yet, so the lcore never wakes from rte_epoll_wait().

Register from rx_interrupt_wait(), after the queues are armed, and only
once: keep the epoll entry installed for the lifetime of the loop, with
the state in struct lcore_conf. NICs with a fixed per-queue fd are
unaffected.

The "RX interrupt won't enable" notice moves to the loop error path, where
it is still logged once per lcore since intr_en latches off.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 31 +++++++++++++++++--------------
 1 file changed, 17 insertions(+), 14 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 70f76d202e..43a0509bd5 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -368,6 +368,7 @@ struct __rte_cache_aligned lcore_conf {
 	struct rte_eth_dev_tx_buffer *tx_buffer[RTE_MAX_ETHPORTS];
 	lookup_struct_t * ipv4_lookup_struct;
 	lookup_struct_t * ipv6_lookup_struct;
+	bool intr_registered;
 };
 
 struct __rte_cache_aligned lcore_stats {
@@ -953,6 +954,16 @@ rx_interrupt_wait(struct lcore_conf *qconf)
 	if (ret != 0)
 		return ret;
 
+	/* some PMDs expose the interrupt fd only once the queue is armed */
+	if (!qconf->intr_registered) {
+		ret = event_register(qconf);
+		if (ret != 0) {
+			rx_intr_disable_all(qconf);
+			return ret;
+		}
+		qconf->intr_registered = true;
+	}
+
 	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
 	rx_intr_disable_all(qconf);
 	return 0;
@@ -970,7 +981,7 @@ static int main_intr_loop(__rte_unused void *dummy)
 	struct lcore_rx_queue *rx_queue;
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
-	int intr_en = 0;
+	int intr_en = 1;
 	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) /
@@ -998,12 +1009,6 @@ static int main_intr_loop(__rte_unused void *dummy)
 				lcore_id, portid, queueid);
 	}
 
-	/* add into event wait list */
-	if (event_register(qconf) == 0)
-		intr_en = 1;
-	else
-		RTE_LOG(INFO, L3FWD_POWER, "RX interrupt won't enable.\n");
-
 	while (!is_done()) {
 		stats[lcore_id].nb_iteration_looped++;
 
@@ -1107,6 +1112,8 @@ static int main_intr_loop(__rte_unused void *dummy)
 					if (ret == -EAGAIN)
 						goto start_rx;
 					if (ret != 0) {
+						RTE_LOG(INFO, L3FWD_POWER,
+							"RX interrupt won't enable.\n");
 						intr_en = 0;
 						continue;
 					}
@@ -1260,7 +1267,7 @@ main_legacy_loop(__rte_unused void *dummy)
 	enum freq_scale_hint_t lcore_scaleup_hint;
 	uint32_t lcore_rx_idle_count = 0;
 	uint32_t lcore_idle_hint = 0;
-	int intr_en = 0;
+	int intr_en = 1;
 	int ret;
 
 	const uint64_t drain_tsc = (rte_get_tsc_hz() + US_PER_S - 1) / US_PER_S * BURST_TX_DRAIN_US;
@@ -1286,12 +1293,6 @@ main_legacy_loop(__rte_unused void *dummy)
 			"rxqueueid=%" PRIu16 "\n", lcore_id, portid, queueid);
 	}
 
-	/* add into event wait list */
-	if (event_register(qconf) == 0)
-		intr_en = 1;
-	else
-		RTE_LOG(INFO, L3FWD_POWER, "RX interrupt won't enable.\n");
-
 	while (!is_done()) {
 		stats[lcore_id].nb_iteration_looped++;
 
@@ -1428,6 +1429,8 @@ main_legacy_loop(__rte_unused void *dummy)
 					if (ret == -EAGAIN)
 						goto start_rx;
 					if (ret != 0) {
+						RTE_LOG(INFO, L3FWD_POWER,
+							"RX interrupt won't enable.\n");
 						intr_en = 0;
 						continue;
 					}
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v2 4/6] examples/l3fwd-power: accept shared Rx interrupt FD
  2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
                     ` (2 preceding siblings ...)
  2026-09-09 15:04   ` [PATCH v2 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add Maxime Leroy
@ 2026-09-09 15:04   ` Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 5/6] examples/l3fwd-power: recheck Rx queues before sleeping Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 6/6] examples/l3fwd-power: block until Rx interrupt or exit Maxime Leroy
  5 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-09 15:04 UTC (permalink / raw)
  To: dev; +Cc: david.marchand, anatoly.burakov, sivaprasad.tummala, Maxime Leroy

Some PMDs deliver the Rx interrupts of several queues through a single
shared fd; dpaa2, for instance, uses one QBMan portal per lcore.
Registering the fd a second time then returns -EEXIST, which
event_register() treated as fatal and disabled interrupt mode for the
whole lcore.

Accept -EEXIST as success. NICs with a dedicated fd per queue never
return it, so their behaviour is unchanged.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 43a0509bd5..5638d43826 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -938,7 +938,12 @@ static int event_register(struct lcore_conf *qconf)
 						RTE_EPOLL_PER_THREAD,
 						RTE_INTR_EVENT_ADD,
 						(void *)((uintptr_t)data));
-		if (ret)
+		/*
+		 * Queues polled on the same lcore may share one interrupt fd,
+		 * so registering the second onward returns -EEXIST; treat it
+		 * as success.
+		 */
+		if (ret && ret != -EEXIST)
 			return ret;
 	}
 
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v2 5/6] examples/l3fwd-power: recheck Rx queues before sleeping
  2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
                     ` (3 preceding siblings ...)
  2026-09-09 15:04   ` [PATCH v2 4/6] examples/l3fwd-power: accept shared Rx interrupt FD Maxime Leroy
@ 2026-09-09 15:04   ` Maxime Leroy
  2026-09-09 15:04   ` [PATCH v2 6/6] examples/l3fwd-power: block until Rx interrupt or exit Maxime Leroy
  5 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-09 15:04 UTC (permalink / raw)
  To: dev; +Cc: david.marchand, anatoly.burakov, sivaprasad.tummala, Maxime Leroy

After arming the Rx interrupts rx_interrupt_wait() sleeps in
rte_epoll_wait(). A packet that arrived between the last poll and the arm
can be missed on PMDs whose interrupt is edge-triggered on an empty to
non-empty transition, such as dpaa2: arming a queue that is already
non-empty raises no notification, so the lcore sleeps until the next
packet.

Before sleeping, check whether any queue already holds traffic and, if
so, skip the wait and go back to polling. NICs with level-triggered
interrupts are unaffected: a pending packet keeps the interrupt asserted,
so rte_epoll_wait() would return immediately anyway.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 27 ++++++++++++++++++++++++++-
 1 file changed, 26 insertions(+), 1 deletion(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index 5638d43826..c1ecf118f5 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -950,6 +950,26 @@ static int event_register(struct lcore_conf *qconf)
 	return 0;
 }
 
+static bool
+rx_queue_pending(struct lcore_conf *qconf)
+{
+	struct lcore_rx_queue *rx_queue;
+	uint16_t queue_id;
+	uint16_t port_id;
+	int i;
+
+	for (i = 0; i < qconf->n_rx_queue; ++i) {
+		rx_queue = &(qconf->rx_queue_list[i]);
+		port_id = rx_queue->port_id;
+		queue_id = rx_queue->queue_id;
+
+		if (rte_eth_rx_queue_count(port_id, queue_id) > 0)
+			return true;
+	}
+
+	return false;
+}
+
 static int
 rx_interrupt_wait(struct lcore_conf *qconf)
 {
@@ -969,7 +989,12 @@ rx_interrupt_wait(struct lcore_conf *qconf)
 		qconf->intr_registered = true;
 	}
 
-	sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
+	/*
+	 * A packet that arrived during the arm window raises no new wakeup on
+	 * edge-triggered PMDs, so skip the sleep if a queue already has traffic.
+	 */
+	if (!rx_queue_pending(qconf))
+		sleep_until_rx_interrupt(qconf->n_rx_queue, rte_lcore_id());
 	rx_intr_disable_all(qconf);
 	return 0;
 }
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

* [PATCH v2 6/6] examples/l3fwd-power: block until Rx interrupt or exit
  2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
                     ` (4 preceding siblings ...)
  2026-09-09 15:04   ` [PATCH v2 5/6] examples/l3fwd-power: recheck Rx queues before sleeping Maxime Leroy
@ 2026-09-09 15:04   ` Maxime Leroy
  5 siblings, 0 replies; 20+ messages in thread
From: Maxime Leroy @ 2026-09-09 15:04 UTC (permalink / raw)
  To: dev; +Cc: david.marchand, anatoly.burakov, sivaprasad.tummala, Maxime Leroy

In interrupt and legacy modes an idle worker arms its Rx queues and
sleeps in rte_epoll_wait() with a 10 ms timeout. The timeout lets the
worker wake periodically to re-check the quit flag, but it also hides a
broken interrupt path: with no interrupt delivered the worker still wakes
every 10 ms and polls, so traffic keeps flowing and interrupt mode looks
like it works when it does not.

Block indefinitely instead, so only a real Rx interrupt can wake the
worker. To still stop the workers on SIGINT, add an eventfd to each
worker epoll set; the signal handler writes it to wake every worker,
which then observes the quit flag and leaves its loop. Both loops share
this sleep path, so the eventfd is created for the interrupt and legacy
modes; it is skipped in the wakeup log so it is not reported as an Rx
interrupt.

Signed-off-by: Maxime Leroy <maxime@leroys.fr>
---
 examples/l3fwd-power/main.c | 40 +++++++++++++++++++++++++++++++++----
 1 file changed, 36 insertions(+), 4 deletions(-)

diff --git a/examples/l3fwd-power/main.c b/examples/l3fwd-power/main.c
index c1ecf118f5..5ee8e5217e 100644
--- a/examples/l3fwd-power/main.c
+++ b/examples/l3fwd-power/main.c
@@ -9,6 +9,8 @@
 #include <sys/types.h>
 #include <string.h>
 #include <sys/queue.h>
+#include <sys/epoll.h>
+#include <sys/eventfd.h>
 #include <stdarg.h>
 #include <errno.h>
 #include <getopt.h>
@@ -143,6 +145,9 @@ static int promiscuous_on = 0;
 /* NUMA is enabled by default. */
 static int numa_on = 1;
 volatile bool quit_signal;
+/* eventfd to wake idle workers out of a blocking rte_epoll_wait() at exit */
+static int wakeup_fd = -1;
+static struct rte_epoll_event wakeup_event[RTE_MAX_LCORE];
 /* timer to update telemetry every 500ms */
 static struct rte_timer telemetry_timer;
 
@@ -421,10 +426,17 @@ static int is_done(void)
 static void
 signal_exit_now(int sigtype)
 {
+	uint64_t one = 1;
+	ssize_t nbytes;
 
-	if (sigtype == SIGINT)
+	if (sigtype == SIGINT) {
 		quit_signal = true;
-
+		/* interrupt mode only: wake a worker blocked in rte_epoll_wait() */
+		if (wakeup_fd >= 0) {
+			nbytes = write(wakeup_fd, &one, sizeof(one));
+			RTE_SET_USED(nbytes);
+		}
+	}
 }
 
 /*  Frequency scale down timer callback */
@@ -836,7 +848,7 @@ sleep_until_rx_interrupt(int num, int lcore)
 	static alignas(RTE_CACHE_LINE_SIZE) struct {
 		bool wakeup;
 	} status[RTE_MAX_LCORE];
-	struct rte_epoll_event event[num];
+	struct rte_epoll_event event[num + 1];
 	int n, i;
 	uint16_t port_id;
 	uint16_t queue_id;
@@ -848,8 +860,10 @@ sleep_until_rx_interrupt(int num, int lcore)
 				rte_lcore_id());
 	}
 
-	n = rte_epoll_wait(RTE_EPOLL_PER_THREAD, event, num, 10);
+	n = rte_epoll_wait(RTE_EPOLL_PER_THREAD, event, num + 1, -1);
 	for (i = 0; i < n; i++) {
+		if (event[i].fd == wakeup_fd)
+			continue;
 		data = event[i].epdata.data;
 		port_id = ((uintptr_t)data) >> (sizeof(uint16_t) * CHAR_BIT);
 		queue_id = ((uintptr_t)data) &
@@ -922,6 +936,7 @@ rx_intr_enable_all(struct lcore_conf *qconf)
 static int event_register(struct lcore_conf *qconf)
 {
 	struct lcore_rx_queue *rx_queue;
+	struct rte_epoll_event *wev;
 	uint16_t queueid;
 	uint16_t portid;
 	uint32_t data;
@@ -947,6 +962,16 @@ static int event_register(struct lcore_conf *qconf)
 			return ret;
 	}
 
+	/* also watch the shutdown eventfd so a blocking wait wakes at exit */
+	if (wakeup_fd >= 0) {
+		wev = &wakeup_event[rte_lcore_id()];
+		wev->epdata.event = EPOLLIN | EPOLLPRI;
+		ret = rte_epoll_ctl(RTE_EPOLL_PER_THREAD, EPOLL_CTL_ADD,
+				    wakeup_fd, wev);
+		if (ret && ret != -EEXIST)
+			return ret;
+	}
+
 	return 0;
 }
 
@@ -2992,6 +3017,13 @@ main(int argc, char **argv)
 
 	check_all_ports_link_status(enabled_port_mask);
 
+	if (app_mode == APP_MODE_LEGACY || app_mode == APP_MODE_INTERRUPT) {
+		wakeup_fd = eventfd(0, EFD_NONBLOCK);
+		if (wakeup_fd < 0)
+			rte_exit(EXIT_FAILURE, "eventfd() failed: %s\n",
+				 strerror(errno));
+	}
+
 	/* launch per-lcore init on every lcore */
 	if (app_mode == APP_MODE_LEGACY) {
 		rte_eal_mp_remote_launch(main_legacy_loop, NULL, CALL_MAIN);
-- 
2.43.0


^ permalink raw reply related	[flat|nested] 20+ messages in thread

end of thread, other threads:[~2026-09-09 15:05 UTC | newest]

Thread overview: 20+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-20  8:00 [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
2026-07-20  8:00 ` [PATCH 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path Maxime Leroy
2026-07-20  8:00 ` [PATCH 2/6] examples/l3fwd-power: check Rx interrupt enable errors Maxime Leroy
2026-07-20  8:00 ` [PATCH 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add Maxime Leroy
2026-09-04  9:15   ` David Marchand
2026-07-20  8:00 ` [PATCH 4/6] examples/l3fwd-power: accept shared Rx interrupt FD Maxime Leroy
2026-07-20  8:00 ` [PATCH 5/6] examples/l3fwd-power: recheck Rx queues before sleeping Maxime Leroy
2026-09-04  9:20   ` David Marchand
2026-09-04 20:25     ` Maxime Leroy
2026-07-20  8:00 ` [PATCH 6/6] examples/l3fwd-power: block until Rx interrupt or exit Maxime Leroy
2026-08-25  8:07 ` [PATCH 0/6] examples/l3fwd-power: fix Rx interrupt mode for shared FD PMDs Maxime Leroy
2026-09-04  9:30 ` David Marchand
2026-09-04 21:10   ` Maxime Leroy
2026-09-09 15:04 ` [PATCH v2 " Maxime Leroy
2026-09-09 15:04   ` [PATCH v2 1/6] examples/l3fwd-power: factor out Rx interrupt sleep path Maxime Leroy
2026-09-09 15:04   ` [PATCH v2 2/6] examples/l3fwd-power: check Rx interrupt enable errors Maxime Leroy
2026-09-09 15:04   ` [PATCH v2 3/6] examples/l3fwd-power: enable Rx interrupt before epoll add Maxime Leroy
2026-09-09 15:04   ` [PATCH v2 4/6] examples/l3fwd-power: accept shared Rx interrupt FD Maxime Leroy
2026-09-09 15:04   ` [PATCH v2 5/6] examples/l3fwd-power: recheck Rx queues before sleeping Maxime Leroy
2026-09-09 15:04   ` [PATCH v2 6/6] examples/l3fwd-power: block until Rx interrupt or exit Maxime Leroy

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox