Linux Netfilter development
 help / color / mirror / Atom feed
From: Florian Westphal <fw@strlen.de>
To: Antonio Ojea <antonio.ojea.garcia@gmail.com>
Cc: netfilter-devel@vger.kernel.org,
	Pablo Neira Ayuso <pablo@netfilter.org>,
	Phil Sutter <phil@nwl.cc>, Dan Winship <danwinship@redhat.com>
Subject: Re: nf_queue: hook changes drop or re-queue packets waiting in any nfqueue
Date: Fri, 25 Sep 2026 16:09:10 +0200	[thread overview]
Message-ID: <araAhumrNKQ48-Re@strlen.de> (raw)
In-Reply-To: <CABhP=tYtCyr-qzL==98nkmDVZr6x+f8E26V1AAOpMXgZgt0vAQ@mail.gmail.com>

Antonio Ojea <antonio.ojea.garcia@gmail.com> wrote:
> deleting an nftables base chain drops every packet that is waiting for
> a verdict in any nfqueue of the network namespace, including packets
> queued by hooks of other tables and other families. Adding a base chain
> in front of the one that queued a packet makes the packet traverse the
> queuing hook again after the verdict, so userspace sees it twice.
> 
> We found the first problem in kube-network-policies and kindnet
> (Kubernetes network policy and CNI agents built on nfqueue) [1]. Their
> nftables sync did "add table; delete table; add table" to replace the
> ruleset, and every sync lost the connections being evaluated at that
> moment. The userspace symptom is the verdict for the dropped packet
> failing asynchronously with -ENOENT:
> 
>   "Could not receive message"
>   error="netlink receive: no such file or directory"
> 
> The add/delete/add idiom was my mistake. The documented way to replace
> a ruleset atomically is "flush table" (or "flush chain") in the same
> transaction, which keeps the base chains and their hooks, and we need
> to change that. However, it does not cover upgrades. A table written by an
> older version may have a chain and/or sets the new version
> no longer uses. Removing them still requires deleting the table, or
> the chain, at least once at startup, with the same effect on every
> nfqueue in the namespace. Suggestions on how to handle that case are
> welcome.
> 
> Independently of fixing it in our projects, any other component will
> trigger the same problem that applies to any nfqueue consumer.

I don't think this is fixable.  You could tell your coding assistant
to pass const struct nf_hook_ops *reg down into nf_queue_nf_hook_drop().

If thats not NULL, then pass it to a new nfqnl_cmpfn() that only
drops skbs when the struct nf_queue_entry shares same pf and
same hooknum as the one in nf_hook_ops arg.

That would limit the impact a bit.

  reply	other threads:[~2026-09-25 14:09 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25 12:58 nf_queue: hook changes drop or re-queue packets waiting in any nfqueue Antonio Ojea
2026-09-25 14:09 ` Florian Westphal [this message]
2026-09-26 10:03   ` [PATCH nf-next 1/2] netfilter: nf_queue: limit the hook drop flush to the unregistered hook point Antonio Ojea
2026-09-28 14:11     ` Florian Westphal
2026-09-26 10:03   ` [PATCH nf-next 2/2] selftests: netfilter: nft_queue: check the scope of the hook drop flush Antonio Ojea

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=araAhumrNKQ48-Re@strlen.de \
    --to=fw@strlen.de \
    --cc=antonio.ojea.garcia@gmail.com \
    --cc=danwinship@redhat.com \
    --cc=netfilter-devel@vger.kernel.org \
    --cc=pablo@netfilter.org \
    --cc=phil@nwl.cc \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox