* hardware flows are left behind when a flowtable is deleted (nf_tables unbinds the offload block at commit time)
@ 2026-10-05 11:48 Martino Dell'Ambrogio
2026-10-06 11:28 ` Pablo Neira Ayuso
0 siblings, 1 reply; 2+ messages in thread
From: Martino Dell'Ambrogio @ 2026-10-05 11:48 UTC (permalink / raw)
To: netfilter-devel; +Cc: pablo, fw, phil
Deleting a flowtable leaves the flows its devices offloaded behind in the
hardware. Reported against stable v6.12.108 on a Banana Pi R4 (MT7988,
mtk_eth_soc PPE) running OpenWrt, and reproduced there; the analysis below is
from that source and from net.git main, where the function is identical.
Symptom
After the flowtable is deleted, the offloaded flows stay BND in the PPE and
keep forwarding. Their conntrack entries are not released either: they stop
being refreshed and expire (60 s for UDP), after which the device is
forwarding flows that exist in hardware only and that conntrack no longer
knows about. On the router this is visible as entries in
/sys/kernel/debug/ppe{1,2}/bind with no matching entry in `conntrack -L`.
OpenWrt's fw4 deletes and recreates its flowtable in one transaction on every
firewall reload, so this happens on every reload rather than in some rare
teardown path.
Why
- Deleting a flowtable moves FLOW_BLOCK_UNBIND to commit time.
__nft_unregister_flowtable_net_hooks() asks the driver to unbind, and
nf_flow_table_block_setup() frees the driver callback there.
- The flowtable itself is freed later, from the transaction destroy work
after synchronize_rcu(): nft_commit_release() ->
nf_tables_flowtable_destroy() -> type->free() -> nf_flow_table_free().
That last step tears the remaining flows down and queues
FLOW_CLS_DESTROY for each one.
- nf_flow_offload_tuple() delivers those commands by walking
flowtable->flow_block.cb_list, which has been empty since the commit. It
reports no error for an empty list, so nothing notices.
This is not a race between the last GC run and the unbind: the unbind always
happens first, at commit time, so the hardware entries are never removed.
It is a regression from 13210fc63f35 ("netfilter: nf_tables: imbalance in
flowtable binding"). That commit removed the late FLOW_BLOCK_UNBIND from the
flowtable destroy path, which 5acab91458ce ("netfilter: nf_tables: unbind
callbacks from flowtable destroy path") had placed after
nf_flow_table_free() on purpose:
Callback unbinding needs to be done after nf_flow_table_free(),
otherwise entries are not removed from the hardware.
13210fc63f35's thread only checked BIND/UNBIND balance; nothing there is
about hardware entries left behind. It is in 6.12.y as 2e87c203b72f.
All of the above is read from the sources rather than instrumented at run
time: nothing here traces cb_list directly. What the reproducer below shows
is the consequence, and that part is measured.
Reproducer (~2 minutes, one host)
1. Drive a busy UDP flow (~100 pps) through the flowtable so it binds in
hardware: confirm a BND entry in /sys/kernel/debug/ppe*/bind and
[HW_OFFLOAD] in `conntrack -L`.
2. Delete the flowtable (on OpenWrt: `fw4 reload`, which deletes and
recreates it).
3. Keep the same socket sending, at a low rate, for longer than the
conntrack timeout. The window has to exceed it: a stranded flow still has
a conntrack entry until that expires, so a shorter window shows nothing.
Unpatched: the tuple stays BND and keeps forwarding while conntrack has no
entry for it any more.
The discriminator has to be per tuple -- bound in hardware AND absent from
conntrack. An aggregate "no conntrack entries left" test never fires, because
other hosts on the LAN are talking to the same destination.
Scope
Drivers that offload through TC_SETUP_FT and rely on FLOW_CLS_DESTROY to
release hardware state look affected by inspection: mtk_eth_soc (measured),
mtk_wed, airoha, mt76 npu, mlx5 representor FT. Drivers that flush their own
entries when they remove their callback (mlx5 tc_ct, nfp, sfc) are not. Only
mtk_eth_soc was tested.
A possible fix, and what I measured
Remove the flows using the device before the block is unbound:
nf_flow_table_offload_setup() on FLOW_BLOCK_UNBIND could call
nf_flow_table_gc_cleanup(flowtable, dev) first, so the FLOW_CLS_DESTROY is
delivered while the callback is still on cb_list. That is what the kernel
already does for indirect devices (nf_flow_table_indr_cleanup()) and on
NETDEV_DOWN, and one site covers the commit, update, netdev-event and abort
paths. I am aware that "generalize pending status bit" reworks this area.
I have not sent a patch, but I have built that change and run it on the
hardware: two images differing by that one hunk and nothing else, the same
reproducer on each.
- Unmodified kernel: the reproducer strands the flows. All 4 flows stay
bound in the PPE, in 26 of 26 samples taken after the reload -- first
sustained 22.7 s in and still true at 167 s, past both the 20 s and the
60 s UDP timeouts.
- With the change: no flow is left bound with no conntrack entry behind
it -- 0 of 4, clean in all 25 samples, and the flows re-bind under the
new flowtable.
An orphan check on the patched image (every bound entry cross-referenced
against conntrack, with a synthetic control entry injected to show the check
can fire) found none at +10 minutes or +1 hour. The change is carried as a
downstream patch in OpenWrt.
- Martino Dell'Ambrogio <tillo@tillo.ch>
Assisted-by: LLM
^ permalink raw reply [flat|nested] 2+ messages in thread* Re: hardware flows are left behind when a flowtable is deleted (nf_tables unbinds the offload block at commit time)
2026-10-05 11:48 hardware flows are left behind when a flowtable is deleted (nf_tables unbinds the offload block at commit time) Martino Dell'Ambrogio
@ 2026-10-06 11:28 ` Pablo Neira Ayuso
0 siblings, 0 replies; 2+ messages in thread
From: Pablo Neira Ayuso @ 2026-10-06 11:28 UTC (permalink / raw)
To: Martino Dell'Ambrogio; +Cc: netfilter-devel, fw, phil
Hi,
On Mon, Oct 05, 2026 at 01:48:10PM +0200, Martino Dell'Ambrogio wrote:
> Deleting a flowtable leaves the flows its devices offloaded behind in the
> hardware. Reported against stable v6.12.108 on a Banana Pi R4 (MT7988,
> mtk_eth_soc PPE) running OpenWrt, and reproduced there; the analysis below is
> from that source and from net.git main, where the function is identical.
I have restored my Pi R64 testbed with OpenWrt.
Let me reproduce this issue and review all pending fixes for -stable
kernels too for the flowtable hardware offload.
I will get back to you with updates, thanks for reporting.
> Symptom
>
> After the flowtable is deleted, the offloaded flows stay BND in the PPE and
> keep forwarding. Their conntrack entries are not released either: they stop
> being refreshed and expire (60 s for UDP), after which the device is
> forwarding flows that exist in hardware only and that conntrack no longer
> knows about. On the router this is visible as entries in
> /sys/kernel/debug/ppe{1,2}/bind with no matching entry in `conntrack -L`.
>
> OpenWrt's fw4 deletes and recreates its flowtable in one transaction on every
> firewall reload, so this happens on every reload rather than in some rare
> teardown path.
>
> Why
>
> - Deleting a flowtable moves FLOW_BLOCK_UNBIND to commit time.
> __nft_unregister_flowtable_net_hooks() asks the driver to unbind, and
> nf_flow_table_block_setup() frees the driver callback there.
> - The flowtable itself is freed later, from the transaction destroy work
> after synchronize_rcu(): nft_commit_release() ->
> nf_tables_flowtable_destroy() -> type->free() -> nf_flow_table_free().
> That last step tears the remaining flows down and queues
> FLOW_CLS_DESTROY for each one.
> - nf_flow_offload_tuple() delivers those commands by walking
> flowtable->flow_block.cb_list, which has been empty since the commit. It
> reports no error for an empty list, so nothing notices.
>
> This is not a race between the last GC run and the unbind: the unbind always
> happens first, at commit time, so the hardware entries are never removed.
>
> It is a regression from 13210fc63f35 ("netfilter: nf_tables: imbalance in
> flowtable binding"). That commit removed the late FLOW_BLOCK_UNBIND from the
> flowtable destroy path, which 5acab91458ce ("netfilter: nf_tables: unbind
> callbacks from flowtable destroy path") had placed after
> nf_flow_table_free() on purpose:
>
> Callback unbinding needs to be done after nf_flow_table_free(),
> otherwise entries are not removed from the hardware.
>
> 13210fc63f35's thread only checked BIND/UNBIND balance; nothing there is
> about hardware entries left behind. It is in 6.12.y as 2e87c203b72f.
>
> All of the above is read from the sources rather than instrumented at run
> time: nothing here traces cb_list directly. What the reproducer below shows
> is the consequence, and that part is measured.
>
> Reproducer (~2 minutes, one host)
>
> 1. Drive a busy UDP flow (~100 pps) through the flowtable so it binds in
> hardware: confirm a BND entry in /sys/kernel/debug/ppe*/bind and
> [HW_OFFLOAD] in `conntrack -L`.
> 2. Delete the flowtable (on OpenWrt: `fw4 reload`, which deletes and
> recreates it).
> 3. Keep the same socket sending, at a low rate, for longer than the
> conntrack timeout. The window has to exceed it: a stranded flow still has
> a conntrack entry until that expires, so a shorter window shows nothing.
>
> Unpatched: the tuple stays BND and keeps forwarding while conntrack has no
> entry for it any more.
> The discriminator has to be per tuple -- bound in hardware AND absent from
> conntrack. An aggregate "no conntrack entries left" test never fires, because
> other hosts on the LAN are talking to the same destination.
>
> Scope
>
> Drivers that offload through TC_SETUP_FT and rely on FLOW_CLS_DESTROY to
> release hardware state look affected by inspection: mtk_eth_soc (measured),
> mtk_wed, airoha, mt76 npu, mlx5 representor FT. Drivers that flush their own
> entries when they remove their callback (mlx5 tc_ct, nfp, sfc) are not. Only
> mtk_eth_soc was tested.
>
> A possible fix, and what I measured
>
> Remove the flows using the device before the block is unbound:
> nf_flow_table_offload_setup() on FLOW_BLOCK_UNBIND could call
> nf_flow_table_gc_cleanup(flowtable, dev) first, so the FLOW_CLS_DESTROY is
> delivered while the callback is still on cb_list. That is what the kernel
> already does for indirect devices (nf_flow_table_indr_cleanup()) and on
> NETDEV_DOWN, and one site covers the commit, update, netdev-event and abort
> paths. I am aware that "generalize pending status bit" reworks this area.
>
> I have not sent a patch, but I have built that change and run it on the
> hardware: two images differing by that one hunk and nothing else, the same
> reproducer on each.
>
> - Unmodified kernel: the reproducer strands the flows. All 4 flows stay
> bound in the PPE, in 26 of 26 samples taken after the reload -- first
> sustained 22.7 s in and still true at 167 s, past both the 20 s and the
> 60 s UDP timeouts.
> - With the change: no flow is left bound with no conntrack entry behind
> it -- 0 of 4, clean in all 25 samples, and the flows re-bind under the
> new flowtable.
>
> An orphan check on the patched image (every bound entry cross-referenced
> against conntrack, with a synthetic control entry injected to show the check
> can fire) found none at +10 minutes or +1 hour. The change is carried as a
> downstream patch in OpenWrt.
>
> - Martino Dell'Ambrogio <tillo@tillo.ch>
>
> Assisted-by: LLM
>
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-06 11:28 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 11:48 hardware flows are left behind when a flowtable is deleted (nf_tables unbinds the offload block at commit time) Martino Dell'Ambrogio
2026-10-06 11:28 ` Pablo Neira Ayuso
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox