From: Pablo Neira Ayuso <pablo@netfilter.org>
To: kuba@kernel.org
Cc: netdev-bot+sashiko@kernel.org, netfilter-devel@vger.kernel.org,
davem@davemloft.net, netdev@vger.kernel.org, pabeni@redhat.com,
edumazet@google.com, horms@kernel.org, fw@strlen.de, ja@ssi.bg
Subject: Re: [PATCH net 8/8] netfilter: nf_tables: skip expired catchall elements on insert and delete
Date: Sat, 19 Sep 2026 16:57:18 +0200 [thread overview]
Message-ID: <aq6iznIv5wHnIhDg@chamomile> (raw)
In-Reply-To: <178982489117.22033.1574367831667981015@kernel.org>
Hi Jakub,
This comments below look like brainstorming, nothing really that
should prevent this patch from being merged.
On Sat, Sep 19, 2026 at 01:34:51PM +0000, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider:
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 0
>
> - [Medium] The relaxed duplicate predicate in
> nft_setelem_catchall_insert() (nf_tables_api.c:7003-7016) removes the…
>
> --
>
> netfilter: nf_tables: skip expired catchall elements on insert and delete
>
> This aligns nft_setelem_catchall_insert() and
> nft_setelem_catchall_deactivate() with the lookup and dump predicates, so
> an expired catchall element no longer counts as a duplicate and no longer
> gets picked as the delete candidate. Both walks now also test
> __nft_set_elem_expired() with the per-netns transaction timestamp and
> nft_set_elem_is_dead().
>
> > diff --git a/net/netfilter/nf_tables_api.c b/net/netfilter/nf_tables_api.c
> > index c0b754a2d45b..b59628e6240c 100644
> > --- a/net/netfilter/nf_tables_api.c
> > +++ b/net/netfilter/nf_tables_api.c
> > @@ -6995,11 +6995,14 @@ static int nft_setelem_catchall_insert(const struct net *net,
> > {
> > struct nft_set_elem_catchall *catchall;
> > u8 genmask = nft_genmask_next(net);
> > + u64 tstamp = nft_net_tstamp(net);
> > struct nft_set_ext *ext;
> >
> > list_for_each_entry(catchall, &set->catchall_list, list) {
> > ext = nft_set_elem_ext(set, catchall->elem);
> > - if (nft_set_elem_active(ext, genmask)) {
> > + if (nft_set_elem_active(ext, genmask) &&
> > + !__nft_set_elem_expired(ext, tstamp) &&
> > + !nft_set_elem_is_dead(ext)) {
> > *priv = catchall->elem;
> > return -EEXIST;
> > }
> > }
>
> [Severity: Medium]
> Can repeated expire-then-readd cycles grow set->catchall_list without
> bound here?
>
> With the relaxed predicate, the expired entry is skipped but stays linked
> and genmask-active until GC runs, so the tail of
> nft_setelem_catchall_insert() appends a second entry alongside it:
>
> catchall = kmalloc_obj(*catchall, GFP_KERNEL_ACCOUNT);
> if (!catchall)
> return -ENOMEM;
>
> catchall->elem = elem->priv;
> list_add_tail_rcu(&catchall->list, &set->catchall_list);
>
> Unlike ordinary elements, catchall elements are excluded from the size
> accounting in nft_add_set_elem():
Set size is specified by user... User can create a set without a set
size... This comment makes no sense.
> if (!(flags & NFT_SET_ELEM_CATCHALL)) {
> unsigned int max = nft_set_maxsize(set), nelems;
>
> nelems = atomic_inc_return(&set->nelems);
> if (nelems > max)
> set_full = true;
> }
>
> so the "return set_full ? -ENFILE : 0;" limiter that caps accumulation of
> ordinary elements does not apply to this path. Is there anything else that
> bounds the number of stale catchall entries a set can hold?
We could sets a cap for maximum number of items in a transaction, but
noone requested this so far. There is also memcg which should provide
a cap.
> Reclamation looks like it depends entirely on GC, and the interval is
> user-tunable with no upper bound (NFTA_SET_GC_INTERVAL is accepted as a u32
> in milliseconds and nft_set_gc_interval() just converts it):
>
> include/net/netfilter/nf_tables.h:nft_set_gc_interval() {
> u32 gc_int = READ_ONCE(set->gc_int);
>
> return gc_int ? msecs_to_jiffies(gc_int) : HZ;
> }
>
> For nft_set_rhash and nft_set_rbtree the catchall reclaim runs only from
> the delayed gc work requeued with that value, so with a large gc-interval
> the window for accumulation is correspondingly large.
>
> There is also a packet-path effect: nft_set_catchall_lookup() walks the
> whole list, and new elements are appended at the tail, so every accumulated
> stale entry is dereferenced before the live one is found:
>
> list_for_each_entry_rcu(catchall, &set->catchall_list, list) {
> ext = nft_set_elem_ext(set, catchall->elem);
> if (nft_set_elem_active(ext, genmask) &&
> !nft_set_elem_expired(ext) &&
> !nft_set_elem_is_dead(ext))
> return ext;
> }
>
> Would it be preferable to reuse or force-reclaim the stale catchall element
> on insert, rather than only ignoring it?
Maybe, but that is unrelated to this patch.
> > @@ -7092,11 +7095,14 @@ static int nft_setelem_catchall_deactivate(const struct net *net,
> > struct nft_set_elem *elem)
> > {
> > struct nft_set_elem_catchall *catchall;
> > + u64 tstamp = nft_net_tstamp(net);
> > struct nft_set_ext *ext;
> >
> > list_for_each_entry(catchall, &set->catchall_list, list) {
> > ext = nft_set_elem_ext(set, catchall->elem);
> > - if (!nft_is_active_next(net, ext))
> > + if (!nft_is_active_next(net, ext) ||
> > + __nft_set_elem_expired(ext, tstamp) ||
> > + nft_set_elem_is_dead(ext))
> > continue;
> >
> > kfree(elem->priv);
>
> [Severity: Medium]
> Related to the above: with stale entries now skipped and -ENOENT returned,
> per-element DELSETELEM can no longer retire the accumulated entries, so a
> full set flush or set deletion appears to be the only way to drop them
> before GC fires. Is that the intended behaviour?
>
> --
> Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260918112844.194503-2-pablo%40netfilter.org
prev parent reply other threads:[~2026-09-19 14:57 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-18 11:28 [PATCH net,v2 0/8] Netfilter/IPVS fixes for net Pablo Neira Ayuso
2026-09-18 11:28 ` [PATCH net 1/8] netfilter: flowtable: publish HW_DEAD after worker is done Pablo Neira Ayuso
2026-09-19 13:34 ` netdev-bot+sashiko
2026-09-19 14:47 ` Pablo Neira Ayuso
2026-09-21 22:20 ` patchwork-bot+netdevbpf
2026-09-18 11:28 ` [PATCH net 2/8] netfilter: nfnetlink_queue: hold nfnl mutex in event notifier Pablo Neira Ayuso
2026-09-18 11:28 ` [PATCH net 3/8] netfilter: ip6t_rpfilter: reject routes without inet6_dev Pablo Neira Ayuso
2026-09-18 11:28 ` [PATCH net 4/8] netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read Pablo Neira Ayuso
2026-09-19 13:34 ` netdev-bot+sashiko
2026-09-19 15:02 ` Pablo Neira Ayuso
2026-09-18 11:28 ` [PATCH net 5/8] netfilter: nft_synproxy: use the family-aware checksum helper Pablo Neira Ayuso
2026-09-18 11:28 ` [PATCH net 6/8] ipvs: revalidate ihl before icmp_send Pablo Neira Ayuso
2026-09-18 11:28 ` [PATCH net 7/8] netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name Pablo Neira Ayuso
2026-09-18 11:28 ` [PATCH net 8/8] netfilter: nf_tables: skip expired catchall elements on insert and delete Pablo Neira Ayuso
2026-09-19 13:34 ` netdev-bot+sashiko
2026-09-19 14:57 ` Pablo Neira Ayuso [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aq6iznIv5wHnIhDg@chamomile \
--to=pablo@netfilter.org \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=fw@strlen.de \
--cc=horms@kernel.org \
--cc=ja@ssi.bg \
--cc=kuba@kernel.org \
--cc=netdev-bot+sashiko@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=netfilter-devel@vger.kernel.org \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox