Linux Netfilter development
 help / color / mirror / Atom feed
* nf_queue: hook changes drop or re-queue packets waiting in any nfqueue
@ 2026-09-25 12:58 Antonio Ojea
  2026-09-25 14:09 ` Florian Westphal
  0 siblings, 1 reply; 5+ messages in thread
From: Antonio Ojea @ 2026-09-25 12:58 UTC (permalink / raw)
  To: netfilter-devel
  Cc: Pablo Neira Ayuso, Florian Westphal, Phil Sutter, Dan Winship

[-- Attachment #1: Type: text/plain, Size: 5335 bytes --]

Hi,

deleting an nftables base chain drops every packet that is waiting for
a verdict in any nfqueue of the network namespace, including packets
queued by hooks of other tables and other families. Adding a base chain
in front of the one that queued a packet makes the packet traverse the
queuing hook again after the verdict, so userspace sees it twice.

We found the first problem in kube-network-policies and kindnet
(Kubernetes network policy and CNI agents built on nfqueue) [1]. Their
nftables sync did "add table; delete table; add table" to replace the
ruleset, and every sync lost the connections being evaluated at that
moment. The userspace symptom is the verdict for the dropped packet
failing asynchronously with -ENOENT:

  "Could not receive message"
  error="netlink receive: no such file or directory"

The add/delete/add idiom was my mistake. The documented way to replace
a ruleset atomically is "flush table" (or "flush chain") in the same
transaction, which keeps the base chains and their hooks, and we need
to change that. However, it does not cover upgrades. A table written by an
older version may have a chain and/or sets the new version
no longer uses. Removing them still requires deleting the table, or
the chain, at least once at startup, with the same effect on every
nfqueue in the namespace. Suggestions on how to handle that case are
welcome.

Independently of fixing it in our projects, any other component will
trigger the same problem that applies to any nfqueue consumer.

Analysis assisted by AI (v6.18-rc1, unchanged in current mainline)
--------------------------------------------------

Deleting a base chain ends in __nf_unregister_net_hook(), which after
shrinking the hook array calls nf_queue_nf_hook_drop(net)
(net/netfilter/core.c:522). nfqnl_nf_hook_drop()
(net/netfilter/nfnetlink_queue.c:1184) walks every nfqnl_instance of
the namespace and calls nfqnl_flush(inst, NULL, 0), which reinjects all
entries with NF_DROP. Nothing in that path looks at which hook was
removed, which table owned it, or which hook queued each entry.

The reason is nf_queue_entry::hook_index: the entry stores its position
in the nf_hook_entries array of its hook point, and nf_reinject()
resumes the traversal from that index. Any change to the array makes
the index stale, so on unregister everything queued is dropped.

Registering a hook does not drop anything, but it also changes the
array. When the new hook sorts before the queuing one, nf_reinject()
finds the new hook at hook_index, does ++i, and nf_iterate() runs the
queuing hook again. The packet is queued a second time.

History:

  039b40ee5854 ("netfilter: nf_queue: only call synchronize_net twice
  if nf_queue is active"), v4.12, removed the nf_hook_cmp() filter from
  nfqnl_nf_hook_drop() so that only entries queued by the removed hook
  were dropped. The commit message says: "For the rare case of base
  chain being unregistered or module removal while nfqueue is in use
  the extra hiccup due to the packet drops isn't a big deal."

  960632ece694 ("netfilter: convert hook list to an array"), v4.14,
  replaced the hook pointer in the entry with hook_index.

With nftables base chains managed by userspace daemons, unregistering
a hook is no longer rare, at least not as rare as before, and it happens in
namespaces where unrelated nfqueue consumers are running.

Reproducer
----------

The attached script (AI assisted) uses the nf_queue helper from
tools/testing/selftests/net/netfilter (gcc -o nf_queue nf_queue.c -lmnl).
It queues 20 UDP packets to queue 1 from an inet output chain, the
helper delays every verdict by 100 ms so most packets are waiting when
the ruleset is changed, and it checks that all 20 are delivered and that
the helper received each one exactly once.

  $ sudo NF_QUEUE=./nf_queue ./nfqueue-hook-drop.sh
  PASS: no ruleset change
        sent 20, queued 20, delivered 20
  PASS: delete regular chain (no hook) of unrelated table
        sent 20, queued 20, delivered 20
  PASS: add base chain to unrelated table (ip6 input)
        sent 20, queued 20, delivered 20
  PASS: add base chain after the queuing one (output prio 100)
        sent 20, queued 20, delivered 20
  FAIL: add base chain before the queuing one (output prio raw)
        sent 20, queued 37, delivered 20
  FAIL: delete base chain of unrelated table (ip6 prerouting)
        sent 20, queued ?, delivered 3
        nf_queue (exit status 1): mnl_cb_run: No such file or directory
  FAIL: delete unrelated table (ip6, no queue rule)
        sent 20, queued ?, delivered 3
        nf_queue (exit status 1): mnl_cb_run: No such file or directory

The "unrelated" table is ip6 with a prerouting base chain and no queue
rule. Deleting it drops the packets queued by the inet output chain of
another table. The "before" case shows the 17 packets that were waiting
being queued a second time. Reproduced on 7.1.6; the Go reproducer in
[1] shows the same drops on 7.0.12, also when the table being recreated
is not the one with the queue rule.

I can turn this into a test case for nft_queue.sh if that helps.

Happy to test patches or to send a selftest.

Thanks,
Antonio

[1] https://github.com/kubernetes-sigs/kube-network-policies/issues/402
    (kernel call path and a standalone Go reproducer by Ben Dwyer,
    who found the root cause)

[-- Attachment #2: nfqueue-hook-drop.sh.txt --]
[-- Type: text/plain, Size: 4423 bytes --]

#!/bin/bash
# SPDX-License-Identifier: GPL-2.0
#
# Packets waiting for a verdict in an nfqueue do not survive changes to the
# netfilter hook array of the network namespace:
#
#  - unregistering any hook (deleting any nftables base chain, even one of
#    an unrelated table or family) drops every packet queued in every
#    nfqueue instance of the namespace (nfqnl_nf_hook_drop()).
#  - registering a hook in front of the one that queued the packet makes
#    nf_reinject() resume at a stale index, so the packet traverses the
#    queuing hook again and is queued a second time.
#
# Needs the nf_queue helper from tools/testing/selftests/net/netfilter:
#   gcc -o nf_queue nf_queue.c -lmnl
#
# Usage: NF_QUEUE=./nf_queue ./nfqueue-hook-drop.sh

NF_QUEUE=${NF_QUEUE:-./nf_queue}
PACKETS=20
ret=0

ns="nfq-$(mktemp -u XXXXXX)"
tmp=$(mktemp)

cleanup() {
	ip netns pids "$ns" 2>/dev/null | xargs -r kill 2>/dev/null
	ip netns del "$ns" 2>/dev/null
	rm -f "$tmp"
}
trap cleanup EXIT

if [ ! -x "$NF_QUEUE" ]; then
	echo "SKIP: $NF_QUEUE not found, build tools/testing/selftests/net/netfilter/nf_queue.c" >&2
	exit 4
fi

ip netns add "$ns" || exit 4
ip -net "$ns" link set lo up

# The UDP packets to port 12345 that the namespace sends to itself are queued
# to queue 1 at the IPv4 output hook, without bypass so they wait for the
# verdict. The counter in the input chain tells how many were delivered.
# "unrelated" is a table of another family with a base chain and no queue rule.
setup_ruleset() {
	ip netns exec "$ns" nft -f - <<EOF
flush ruleset
table inet nfqtest {
	chain output {
		type filter hook output priority filter; policy accept;
		udp dport 12345 queue to 1
	}
	chain input {
		type filter hook input priority filter; policy accept;
		udp dport 12345 counter name delivered
	}
	counter delivered {}
}
table ip6 unrelated {
	chain prerouting {
		type filter hook prerouting priority filter; policy accept;
	}
	chain regular {
	}
}
EOF
}

delivered() {
	ip netns exec "$ns" nft list counter inet nfqtest delivered |
		sed -n 's/.*packets \([0-9]*\) .*/\1/p'
}

# queued is how many packets the helper received, printed by nf_queue -c on
# a clean exit. It exits early on ENOENT, a verdict for a packet the kernel
# already dropped, so the count is not available in that case.
queued() {
	sed -n 's/^\([0-9]*\) packets total$/\1/p' "$tmp"
}

nf_queue_wait() {
	for _ in $(seq 50); do
		if ip netns exec "$ns" grep -q '^ *1 ' /proc/net/netfilter/nfnetlink_queue; then
			return 0
		fi
		sleep 0.1
	done
	echo "FAIL: nf_queue did not bind queue 1" >&2
	exit 1
}

# run_case <description> <nft command...>
#
# Sends PACKETS packets while the helper delays every verdict by 100ms, so
# most of them are still waiting when the nft command runs, and checks that
# all of them are delivered and that each one was queued exactly once.
run_case() {
	local desc="$1"
	shift

	setup_ruleset || exit 1

	ip netns exec "$ns" "$NF_QUEUE" -c -q 1 -d 100 -t 2 >"$tmp" 2>&1 &
	local pid=$!
	nf_queue_wait

	ip netns exec "$ns" bash -c "
		for i in \$(seq $PACKETS); do echo x >/dev/udp/127.0.0.1/12345; done"
	sleep 0.3 # a few verdicts land, the rest of the packets wait in the queue

	ip netns exec "$ns" nft "$@" >/dev/null

	wait "$pid"
	local status=$? got_delivered got_queued
	got_delivered=$(delivered)
	got_queued=$(queued)

	if [ "$got_delivered" = "$PACKETS" ] && [ "$got_queued" = "$PACKETS" ]; then
		echo "PASS: $desc"
	else
		echo "FAIL: $desc"
		ret=1
	fi
	echo "      sent $PACKETS, queued ${got_queued:-?}, delivered $got_delivered"
	grep -v '^$' "$tmp" | grep -v '^hook\|packets total' | sed "s/^/      nf_queue (exit status $status): /"
}

run_case "no ruleset change" \
	list ruleset
run_case "delete regular chain (no hook) of unrelated table" \
	delete chain ip6 unrelated regular
run_case "add base chain to unrelated table (ip6 input)" \
	add chain ip6 unrelated input '{ type filter hook input priority filter; }'
run_case "add base chain after the queuing one (output prio 100)" \
	add chain inet nfqtest late '{ type filter hook output priority 100; }'
run_case "add base chain before the queuing one (output prio raw)" \
	add chain inet nfqtest early '{ type filter hook output priority raw; }'
run_case "delete base chain of unrelated table (ip6 prerouting)" \
	delete chain ip6 unrelated prerouting
run_case "delete unrelated table (ip6, no queue rule)" \
	delete table ip6 unrelated

exit $ret

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-28 14:11 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-25 12:58 nf_queue: hook changes drop or re-queue packets waiting in any nfqueue Antonio Ojea
2026-09-25 14:09 ` Florian Westphal
2026-09-26 10:03   ` [PATCH nf-next 1/2] netfilter: nf_queue: limit the hook drop flush to the unregistered hook point Antonio Ojea
2026-09-28 14:11     ` Florian Westphal
2026-09-26 10:03   ` [PATCH nf-next 2/2] selftests: netfilter: nft_queue: check the scope of the hook drop flush Antonio Ojea

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox