From: Pablo Neira Ayuso <pablo@netfilter.org>
To: Martino Dell'Ambrogio <tillo@tillo.ch>
Cc: netfilter-devel@vger.kernel.org, fw@strlen.de, phil@nwl.cc
Subject: Re: hardware flows are left behind when a flowtable is deleted (nf_tables unbinds the offload block at commit time)
Date: Tue, 6 Oct 2026 13:28:26 +0200 [thread overview]
Message-ID: <asTbWsJFJjPMlSud@chamomile> (raw)
In-Reply-To: <20261005134810.312747-1-tillo@tillo.ch>
Hi,
On Mon, Oct 05, 2026 at 01:48:10PM +0200, Martino Dell'Ambrogio wrote:
> Deleting a flowtable leaves the flows its devices offloaded behind in the
> hardware. Reported against stable v6.12.108 on a Banana Pi R4 (MT7988,
> mtk_eth_soc PPE) running OpenWrt, and reproduced there; the analysis below is
> from that source and from net.git main, where the function is identical.
I have restored my Pi R64 testbed with OpenWrt.
Let me reproduce this issue and review all pending fixes for -stable
kernels too for the flowtable hardware offload.
I will get back to you with updates, thanks for reporting.
> Symptom
>
> After the flowtable is deleted, the offloaded flows stay BND in the PPE and
> keep forwarding. Their conntrack entries are not released either: they stop
> being refreshed and expire (60 s for UDP), after which the device is
> forwarding flows that exist in hardware only and that conntrack no longer
> knows about. On the router this is visible as entries in
> /sys/kernel/debug/ppe{1,2}/bind with no matching entry in `conntrack -L`.
>
> OpenWrt's fw4 deletes and recreates its flowtable in one transaction on every
> firewall reload, so this happens on every reload rather than in some rare
> teardown path.
>
> Why
>
> - Deleting a flowtable moves FLOW_BLOCK_UNBIND to commit time.
> __nft_unregister_flowtable_net_hooks() asks the driver to unbind, and
> nf_flow_table_block_setup() frees the driver callback there.
> - The flowtable itself is freed later, from the transaction destroy work
> after synchronize_rcu(): nft_commit_release() ->
> nf_tables_flowtable_destroy() -> type->free() -> nf_flow_table_free().
> That last step tears the remaining flows down and queues
> FLOW_CLS_DESTROY for each one.
> - nf_flow_offload_tuple() delivers those commands by walking
> flowtable->flow_block.cb_list, which has been empty since the commit. It
> reports no error for an empty list, so nothing notices.
>
> This is not a race between the last GC run and the unbind: the unbind always
> happens first, at commit time, so the hardware entries are never removed.
>
> It is a regression from 13210fc63f35 ("netfilter: nf_tables: imbalance in
> flowtable binding"). That commit removed the late FLOW_BLOCK_UNBIND from the
> flowtable destroy path, which 5acab91458ce ("netfilter: nf_tables: unbind
> callbacks from flowtable destroy path") had placed after
> nf_flow_table_free() on purpose:
>
> Callback unbinding needs to be done after nf_flow_table_free(),
> otherwise entries are not removed from the hardware.
>
> 13210fc63f35's thread only checked BIND/UNBIND balance; nothing there is
> about hardware entries left behind. It is in 6.12.y as 2e87c203b72f.
>
> All of the above is read from the sources rather than instrumented at run
> time: nothing here traces cb_list directly. What the reproducer below shows
> is the consequence, and that part is measured.
>
> Reproducer (~2 minutes, one host)
>
> 1. Drive a busy UDP flow (~100 pps) through the flowtable so it binds in
> hardware: confirm a BND entry in /sys/kernel/debug/ppe*/bind and
> [HW_OFFLOAD] in `conntrack -L`.
> 2. Delete the flowtable (on OpenWrt: `fw4 reload`, which deletes and
> recreates it).
> 3. Keep the same socket sending, at a low rate, for longer than the
> conntrack timeout. The window has to exceed it: a stranded flow still has
> a conntrack entry until that expires, so a shorter window shows nothing.
>
> Unpatched: the tuple stays BND and keeps forwarding while conntrack has no
> entry for it any more.
> The discriminator has to be per tuple -- bound in hardware AND absent from
> conntrack. An aggregate "no conntrack entries left" test never fires, because
> other hosts on the LAN are talking to the same destination.
>
> Scope
>
> Drivers that offload through TC_SETUP_FT and rely on FLOW_CLS_DESTROY to
> release hardware state look affected by inspection: mtk_eth_soc (measured),
> mtk_wed, airoha, mt76 npu, mlx5 representor FT. Drivers that flush their own
> entries when they remove their callback (mlx5 tc_ct, nfp, sfc) are not. Only
> mtk_eth_soc was tested.
>
> A possible fix, and what I measured
>
> Remove the flows using the device before the block is unbound:
> nf_flow_table_offload_setup() on FLOW_BLOCK_UNBIND could call
> nf_flow_table_gc_cleanup(flowtable, dev) first, so the FLOW_CLS_DESTROY is
> delivered while the callback is still on cb_list. That is what the kernel
> already does for indirect devices (nf_flow_table_indr_cleanup()) and on
> NETDEV_DOWN, and one site covers the commit, update, netdev-event and abort
> paths. I am aware that "generalize pending status bit" reworks this area.
>
> I have not sent a patch, but I have built that change and run it on the
> hardware: two images differing by that one hunk and nothing else, the same
> reproducer on each.
>
> - Unmodified kernel: the reproducer strands the flows. All 4 flows stay
> bound in the PPE, in 26 of 26 samples taken after the reload -- first
> sustained 22.7 s in and still true at 167 s, past both the 20 s and the
> 60 s UDP timeouts.
> - With the change: no flow is left bound with no conntrack entry behind
> it -- 0 of 4, clean in all 25 samples, and the flows re-bind under the
> new flowtable.
>
> An orphan check on the patched image (every bound entry cross-referenced
> against conntrack, with a synthetic control entry injected to show the check
> can fire) found none at +10 minutes or +1 hour. The change is carried as a
> downstream patch in OpenWrt.
>
> - Martino Dell'Ambrogio <tillo@tillo.ch>
>
> Assisted-by: LLM
>
prev parent reply other threads:[~2026-10-06 11:28 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-05 11:48 hardware flows are left behind when a flowtable is deleted (nf_tables unbinds the offload block at commit time) Martino Dell'Ambrogio
2026-10-06 11:28 ` Pablo Neira Ayuso [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asTbWsJFJjPMlSud@chamomile \
--to=pablo@netfilter.org \
--cc=fw@strlen.de \
--cc=netfilter-devel@vger.kernel.org \
--cc=phil@nwl.cc \
--cc=tillo@tillo.ch \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox