linux-kselftest.vger.kernel.org archive mirror
 help / color / mirror / Atom feed
* [PATCH nf-next] selftests: netfilter: nft_queue.sh: only queue icmp echo request/reply
@ 2026-09-18 18:31 Jakub Kicinski
  2026-09-19  6:24 ` Florian Westphal
  0 siblings, 1 reply; 2+ messages in thread
From: Jakub Kicinski @ 2026-09-18 18:31 UTC (permalink / raw)
  To: pablo, fw
  Cc: netdev, Jakub Kicinski, phil, shuah, netfilter-devel, coreteam,
	linux-kselftest

The bridge test flakes on debug kernels:

  FAIL: Expected 10 packets total, but got 24 packets total
  hook 3 packets 00000008
  hook 4 packets 00000010

The surplus are icmp fragment reassembly timeouts.  The udp flood in the
stress test leaves incomplete datagrams behind in ns2 and ns3, 30 seconds
later their reassembly queues expire and both namespaces send icmp time
exceeded to ns1's pre-bridge address.  ns3 is not reconfigured when the
router is turned into a bridge, so its messages arrive via veth2 and are
then forwarded out of br0. Such packets are locally originated from the
bridge point of view and are queued from the bridge output and
postrouting hooks, which is why only those two counters are off.

Restrict the ipv4 rule to echo request/reply, the icmpv6 rule already
does this.

We used to see 1 flake a day in NIPA before locally queuing this change,
zero flakes since (over 9 days)

Signed-off-by: Jakub Kicinski <kuba@kernel.org>
---
The difference between v4 and v6 has been there from day one,
which makes it seem intentional, but I don't understand nft
well enough to come up with any theories why..

CC: pablo@netfilter.org
CC: fw@strlen.de
CC: phil@nwl.cc
CC: shuah@kernel.org
CC: netfilter-devel@vger.kernel.org
CC: coreteam@netfilter.org
CC: linux-kselftest@vger.kernel.org
---
 tools/testing/selftests/net/netfilter/nft_queue.sh | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/tools/testing/selftests/net/netfilter/nft_queue.sh b/tools/testing/selftests/net/netfilter/nft_queue.sh
index 7c857a2e0f34..c8d1f2eb4133 100755
--- a/tools/testing/selftests/net/netfilter/nft_queue.sh
+++ b/tools/testing/selftests/net/netfilter/nft_queue.sh
@@ -92,7 +92,7 @@ load_ruleset() {
 ip netns exec "$nsrouter" nft -f /dev/stdin <<EOF
 table $family $name {
 	chain nfq {
-		ip protocol icmp queue bypass
+		icmp type { "echo-request", "echo-reply" } queue bypass
 		icmpv6 type { "echo-request", "echo-reply" } queue num 1 bypass
 	}
 	chain pre {
-- 
2.55.0


^ permalink raw reply related	[flat|nested] 2+ messages in thread

* Re: [PATCH nf-next] selftests: netfilter: nft_queue.sh: only queue icmp echo request/reply
  2026-09-18 18:31 [PATCH nf-next] selftests: netfilter: nft_queue.sh: only queue icmp echo request/reply Jakub Kicinski
@ 2026-09-19  6:24 ` Florian Westphal
  0 siblings, 0 replies; 2+ messages in thread
From: Florian Westphal @ 2026-09-19  6:24 UTC (permalink / raw)
  To: Jakub Kicinski
  Cc: pablo, netdev, phil, shuah, netfilter-devel, coreteam,
	linux-kselftest

Jakub Kicinski <kuba@kernel.org> wrote:
> The bridge test flakes on debug kernels:
> 
>   FAIL: Expected 10 packets total, but got 24 packets total
>   hook 3 packets 00000008
>   hook 4 packets 00000010
> 
> The surplus are icmp fragment reassembly timeouts.  The udp flood in the
> stress test leaves incomplete datagrams behind in ns2 and ns3, 30 seconds
> later their reassembly queues expire and both namespaces send icmp time
> exceeded to ns1's pre-bridge address.  ns3 is not reconfigured when the
> router is turned into a bridge, so its messages arrive via veth2 and are
> then forwarded out of br0. Such packets are locally originated from the
> bridge point of view and are queued from the bridge output and
> postrouting hooks, which is why only those two counters are off.
> 
> Restrict the ipv4 rule to echo request/reply, the icmpv6 rule already
> does this.
> 
> We used to see 1 flake a day in NIPA before locally queuing this change,
> zero flakes since (over 9 days)

Thanks for debugging and fixing this!

Reviewed-by: Florian Westphal <fw@strlen.de>

> The difference between v4 and v6 has been there from day one,
> which makes it seem intentional, but I don't understand nft
> well enough to come up with any theories why..

Without the restriction this would also queue IPv6 neighbour
discovery messages.

ARP is not seen by NFPROTO_IPV4 hooks, so the restriction
was not needed. I simply did not think of icmp reassembly
timeout errors getting sent after some time when I wrote
this.

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-09-19  6:24 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-18 18:31 [PATCH nf-next] selftests: netfilter: nft_queue.sh: only queue icmp echo request/reply Jakub Kicinski
2026-09-19  6:24 ` Florian Westphal

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).