* [PATCH net 01/10] netfilter: ipset: do not update comments from kernel-side adds
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:44 ` netdev-bot+sinfo
2026-10-01 10:20 ` patchwork-bot+netdevbpf
2026-09-30 7:41 ` [PATCH net 02/10] netfilter: nft_flow_offload: drop flowtable reference on init error path Pablo Neira Ayuso
` (8 subsequent siblings)
9 siblings, 2 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
From: Florian Westphal <fw@strlen.de>
'Fixes' commit stopped calling ip_set_init_comment() for hash types
from kernel-side-adds (xtables .. -j SET). ip_set_init_comment() says:
"The kadt functions don't use the comment extensions in any way."
But bitmap set type calls the function from kadt cb too.
While this appears to be safe (serialized via the set spinlock), it seems
better to not call the init function either, least of all to keep
behaviour consistent.
ip_set_list calls ip_set_init_comment() only from uadt cb, it can be
kept as-is.
This was triggered by yet another LLM review, hinting that the existing
rcu_dereference_protected() cannot be downgraded to only check if the
nfnl mutex is held.
Fixes: f30415929be8 ("netfilter: ipset: do not update comments from kernel-side hash adds")
Signed-off-by: Florian Westphal <fw@strlen.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/ipset/ip_set_bitmap_gen.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/net/netfilter/ipset/ip_set_bitmap_gen.h b/net/netfilter/ipset/ip_set_bitmap_gen.h
index d6a7e6604542..ae376fa3e7a3 100644
--- a/net/netfilter/ipset/ip_set_bitmap_gen.h
+++ b/net/netfilter/ipset/ip_set_bitmap_gen.h
@@ -159,7 +159,7 @@ mtype_add(struct ip_set *set, void *value, const struct ip_set_ext *ext,
if (SET_WITH_COUNTER(set))
ip_set_init_counter(ext_counter(x, set), ext);
- if (SET_WITH_COMMENT(set))
+ if (SET_WITH_COMMENT(set) && !ext->target)
ip_set_init_comment(set, ext_comment(x, set), ext);
if (SET_WITH_SKBINFO(set))
ip_set_init_skbinfo(ext_skbinfo(x, set), ext);
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* Re: [PATCH net 01/10] netfilter: ipset: do not update comments from kernel-side adds
2026-09-30 7:41 ` [PATCH net 01/10] netfilter: ipset: do not update comments from kernel-side adds Pablo Neira Ayuso
@ 2026-09-30 7:44 ` netdev-bot+sinfo
2026-10-01 10:20 ` patchwork-bot+netdevbpf
1 sibling, 0 replies; 13+ messages in thread
From: netdev-bot+sinfo @ 2026-09-30 7:44 UTC (permalink / raw)
To: Pablo Neira Ayuso
Cc: netfilter-devel, davem, netdev, kuba, pabeni, edumazet, horms, fw,
ja
Hi!
This is an automated message. This series looks like a fix, but its
commit messages seem to be missing some information:
- How the issue was discovered, e.g. hit in production, hit during
development, syzbot report, manual code inspection, LLM or static
analysis tool scan.
- Whether the issue was actually triggered, or is only theoretical
(e.g. found by code inspection). If it was triggered please include
the symptoms, like the stack trace or error messages.
Please do not repost the series just to address the above. Instead,
reply to this email with the missing information, so that reviewers
can take it into account. If the series needs another revision for
other reasons, please include the information in the commit messages
then.
The evaluation is done by an LLM so it may be wrong, if you think
that is the case please reply and explain.
^ permalink raw reply [flat|nested] 13+ messages in thread
* Re: [PATCH net 01/10] netfilter: ipset: do not update comments from kernel-side adds
2026-09-30 7:41 ` [PATCH net 01/10] netfilter: ipset: do not update comments from kernel-side adds Pablo Neira Ayuso
2026-09-30 7:44 ` netdev-bot+sinfo
@ 2026-10-01 10:20 ` patchwork-bot+netdevbpf
1 sibling, 0 replies; 13+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-10-01 10:20 UTC (permalink / raw)
To: Pablo Neira Ayuso
Cc: netfilter-devel, davem, netdev, kuba, pabeni, edumazet, horms, fw,
ja
Hello:
This series was applied to netdev/net.git (main)
by Pablo Neira Ayuso <pablo@netfilter.org>:
On Wed, 30 Sep 2026 09:41:32 +0200 you wrote:
> From: Florian Westphal <fw@strlen.de>
>
> 'Fixes' commit stopped calling ip_set_init_comment() for hash types
> from kernel-side-adds (xtables .. -j SET). ip_set_init_comment() says:
>
> "The kadt functions don't use the comment extensions in any way."
>
> [...]
Here is the summary with links:
- [net,01/10] netfilter: ipset: do not update comments from kernel-side adds
https://git.kernel.org/netdev/net/c/a442c8a89fb9
- [net,02/10] netfilter: nft_flow_offload: drop flowtable reference on init error path
https://git.kernel.org/netdev/net/c/747928b5d4ff
- [net,03/10] ipvs: fix missing counter decrement in lblc
https://git.kernel.org/netdev/net/c/bdac17779f46
- [net,04/10] ipvs: bound LBLCR and LBLC cache growth
https://git.kernel.org/netdev/net/c/2a3c3de660f9
- [net,05/10] ipvs: do not create invisible templates
https://git.kernel.org/netdev/net/c/616cf5c06934
- [net,06/10] ipvs: filter some flags received in the backup server
https://git.kernel.org/netdev/net/c/f96e91748f63
- [net,07/10] netfilter: nft_set_rbtree: skip transaction elements during GC
https://git.kernel.org/netdev/net/c/86b6471e690c
- [net,08/10] netfilter: bpf: reject invalid NAT manipulation types
https://git.kernel.org/netdev/net/c/16d464013ec2
- [net,09/10] netfilter: flowtable: generalize pending status bit
https://git.kernel.org/netdev/net/c/7c549fb7eecd
- [net,10/10] netfilter: flowtable: restore ieee80211 forward path
https://git.kernel.org/netdev/net/c/3ae37eafd366
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 13+ messages in thread
* [PATCH net 02/10] netfilter: nft_flow_offload: drop flowtable reference on init error path
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 01/10] netfilter: ipset: do not update comments from kernel-side adds Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 03/10] ipvs: fix missing counter decrement in lblc Pablo Neira Ayuso
` (7 subsequent siblings)
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
From: Aohan Mei <henrymei@tencent.com>
nft_flow_offload_init() bumps the flowtable use count with
nft_use_inc() before calling nf_ct_netns_get(). When the latter
fails, the error is returned as-is and the reference is leaked.
The upper layers do not balance it either: nf_tables_newexpr()
clears expr->ops when the expression init callback fails, so the
nft_expr_more() iteration in nft_rule_expr_deactivate() and
nf_tables_rule_destroy() stops right before the failed expression
and its ->destroy callback, which would drop the reference, never
runs.
Each failed rule addition therefore leaks one flowtable reference
and the flowtable can no longer be removed: NFT_MSG_DELFLOWTABLE
keeps reporting -EBUSY even though no rule references it.
Save the nf_ct_netns_get() return value and undo the nft_use_inc()
when it fails, restoring the inc/dec pairing within
nft_flow_offload_init() itself.
Fixes: a3c90f7a2323 ("netfilter: nf_tables: flow offload expression")
Reported-by: TencentOS Corvus AI <corvus@tencent.com>
Cc: stable@vger.kernel.org
Assisted-by: CodeBuddy:Kimi-K3
Signed-off-by: Aohan Mei <henrymei@tencent.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/nft_flow_offload.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/net/netfilter/nft_flow_offload.c b/net/netfilter/nft_flow_offload.c
index 32b4281038dd..d3c5651dd699 100644
--- a/net/netfilter/nft_flow_offload.c
+++ b/net/netfilter/nft_flow_offload.c
@@ -160,6 +160,7 @@ static int nft_flow_offload_init(const struct nft_ctx *ctx,
struct nft_flow_offload *priv = nft_expr_priv(expr);
u8 genmask = nft_genmask_next(ctx->net);
struct nft_flowtable *flowtable;
+ int err;
if (!tb[NFTA_FLOW_TABLE_NAME])
return -EINVAL;
@@ -174,7 +175,11 @@ static int nft_flow_offload_init(const struct nft_ctx *ctx,
priv->flowtable = flowtable;
- return nf_ct_netns_get(ctx->net, ctx->family);
+ err = nf_ct_netns_get(ctx->net, ctx->family);
+ if (err < 0)
+ nft_use_dec(&flowtable->use);
+
+ return err;
}
static void nft_flow_offload_deactivate(const struct nft_ctx *ctx,
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH net 03/10] ipvs: fix missing counter decrement in lblc
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 01/10] netfilter: ipset: do not update comments from kernel-side adds Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 02/10] netfilter: nft_flow_offload: drop flowtable reference on init error path Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 04/10] ipvs: bound LBLCR and LBLC cache growth Pablo Neira Ayuso
` (6 subsequent siblings)
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
From: Julian Anastasov <ja@ssi.bg>
LBLC may delete cache entries for destinations that are
removed or overloaded and replace them with available ones.
But ip_vs_lblc_new() forgets to decrement the tbl->entries
counter after calling ip_vs_lblc_del(). This can lead to
increased shrinking of the cache with every new garbage
collection.
Fixes: 2f3d771a35fe ("ipvs: do not use dest after ip_vs_dest_put in LBLC")
Link: https://sashiko.dev/#/patchset/0bdd5abe9968ded7ca2b9cb6844ba83d94cc8d53.1787318053.git.zhilinz%40nebusec.ai
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/ipvs/ip_vs_lblc.c | 1 +
1 file changed, 1 insertion(+)
diff --git a/net/netfilter/ipvs/ip_vs_lblc.c b/net/netfilter/ipvs/ip_vs_lblc.c
index 693bcc82ccb7..a2b574904f13 100644
--- a/net/netfilter/ipvs/ip_vs_lblc.c
+++ b/net/netfilter/ipvs/ip_vs_lblc.c
@@ -203,6 +203,7 @@ ip_vs_lblc_new(struct ip_vs_lblc_table *tbl, const union nf_inet_addr *daddr,
if (en->dest == dest)
return en;
ip_vs_lblc_del(en);
+ atomic_dec(&tbl->entries);
}
en = kmalloc_obj(*en, GFP_ATOMIC);
if (!en)
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH net 04/10] ipvs: bound LBLCR and LBLC cache growth
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
` (2 preceding siblings ...)
2026-09-30 7:41 ` [PATCH net 03/10] ipvs: fix missing counter decrement in lblc Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 05/10] ipvs: do not create invisible templates Pablo Neira Ayuso
` (5 subsequent siblings)
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
From: Zhiling Zou <zhilinz@nebusec.ai>
ip_vs_lblcr_new() and ip_vs_lblc_new() create cache entries for
every previously unseen destination address. The table max_size only
tells the periodic collector to reclaim entries after the cache has
already exceeded the limit. It does not reclaim entries that the
attacker continues to use.
Reject new cache entries once either table reaches max_size * 3 / 2.
The extra headroom lets the periodic collector catch up while the
existing scheduler fallback continues to use the selected destination
when cache creation fails. New traffic therefore stays serviceable
without growing the tables further.
Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Suggested-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Zhiling Zou <zhilinz@nebusec.ai>
Acked-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/ipvs/ip_vs_lblc.c | 3 +++
net/netfilter/ipvs/ip_vs_lblcr.c | 3 +++
2 files changed, 6 insertions(+)
diff --git a/net/netfilter/ipvs/ip_vs_lblc.c b/net/netfilter/ipvs/ip_vs_lblc.c
index a2b574904f13..55af77a56929 100644
--- a/net/netfilter/ipvs/ip_vs_lblc.c
+++ b/net/netfilter/ipvs/ip_vs_lblc.c
@@ -205,6 +205,9 @@ ip_vs_lblc_new(struct ip_vs_lblc_table *tbl, const union nf_inet_addr *daddr,
ip_vs_lblc_del(en);
atomic_dec(&tbl->entries);
}
+ if (atomic_read(&tbl->entries) >= tbl->max_size * 3 / 2)
+ return NULL;
+
en = kmalloc_obj(*en, GFP_ATOMIC);
if (!en)
return NULL;
diff --git a/net/netfilter/ipvs/ip_vs_lblcr.c b/net/netfilter/ipvs/ip_vs_lblcr.c
index f53f05ceea36..858393b1d2d1 100644
--- a/net/netfilter/ipvs/ip_vs_lblcr.c
+++ b/net/netfilter/ipvs/ip_vs_lblcr.c
@@ -363,6 +363,9 @@ ip_vs_lblcr_new(struct ip_vs_lblcr_table *tbl, const union nf_inet_addr *daddr,
en = ip_vs_lblcr_get(af, tbl, daddr);
if (!en) {
+ if (atomic_read(&tbl->entries) >= tbl->max_size * 3 / 2)
+ return NULL;
+
en = kmalloc_obj(*en, GFP_ATOMIC);
if (!en)
return NULL;
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH net 05/10] ipvs: do not create invisible templates
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
` (3 preceding siblings ...)
2026-09-30 7:41 ` [PATCH net 04/10] ipvs: bound LBLCR and LBLC cache growth Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 06/10] ipvs: filter some flags received in the backup server Pablo Neira Ayuso
` (4 subsequent siblings)
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
From: Julian Anastasov <ja@ssi.bg>
The IP_VS_CONN_F_ONE_PACKET flag was implemented for normal
connections. When conn template inherits this flag from
dest->conn_flags it will not be hashed. As result, we will
create new template for every new normal connection.
Fix it to allow one template to be used by many normal
connections.
Fixes: 26ec037f9841 ("IPVS: one-packet scheduling")
Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260916231652.127456-1-pablo%40netfilter.org
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/ipvs/ip_vs_conn.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/net/netfilter/ipvs/ip_vs_conn.c b/net/netfilter/ipvs/ip_vs_conn.c
index 6fa3e1dc534c..cb009208826f 100644
--- a/net/netfilter/ipvs/ip_vs_conn.c
+++ b/net/netfilter/ipvs/ip_vs_conn.c
@@ -1102,6 +1102,9 @@ ip_vs_bind_dest(struct ip_vs_conn *cp, struct ip_vs_dest *dest)
if (cp->protocol != IPPROTO_UDP)
conn_flags &= ~IP_VS_CONN_F_ONE_PACKET;
flags = cp->flags;
+ /* Only visible templates can control multiple connections */
+ if (flags & IP_VS_CONN_F_TEMPLATE)
+ conn_flags &= ~IP_VS_CONN_F_ONE_PACKET;
/* Bind with the destination and its corresponding transmitter */
if (flags & IP_VS_CONN_F_SYNC) {
/* Synced conns are hashed, so they can not get this flag */
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH net 06/10] ipvs: filter some flags received in the backup server
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
` (4 preceding siblings ...)
2026-09-30 7:41 ` [PATCH net 05/10] ipvs: do not create invisible templates Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 07/10] netfilter: nft_set_rbtree: skip transaction elements during GC Pablo Neira Ayuso
` (3 subsequent siblings)
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
From: Julian Anastasov <ja@ssi.bg>
While the IPVS SYNC protocol is not secure by design
we can still protect the backup server from messages that
can wreak havoc.
This commit addresses problems from received connection flags
or their combinations. We now drop messages as follows:
1. the NO_CPORT+TEMPLATE combination allows lookups for normal
connections to hit template which can break in many ways.
While the master does not sync connections with NO_CPORT flag,
i.e. before they are established, we still accept NO_CPORT
without TEMPLATE.
2. ONE_PACKET: it is not sent by master, so we do not
expect it in backup. Before now it was ignored by
IP_VS_CONN_F_BACKUP_MASK for protocol v1 while protocol
v0 created connections that are not hashed and dropped
immediately. Better to apply the IP_VS_CONN_F_BACKUP_MASK
also to the flags from v0 messages for consistency with v1.
Fixes: 87375ab47cd0 ("[IPVS]: ip_vs_ftp breaks connections using persistence")
Signed-off-by: Julian Anastasov <ja@ssi.bg>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/ipvs/ip_vs_sync.c | 33 ++++++++++++++++++++++++++++++---
1 file changed, 30 insertions(+), 3 deletions(-)
diff --git a/net/netfilter/ipvs/ip_vs_sync.c b/net/netfilter/ipvs/ip_vs_sync.c
index 5383aeafb0ae..69dc28153ec1 100644
--- a/net/netfilter/ipvs/ip_vs_sync.c
+++ b/net/netfilter/ipvs/ip_vs_sync.c
@@ -949,6 +949,21 @@ static void ip_vs_proc_conn(struct netns_ipvs *ipvs, struct ip_vs_conn_param *pa
ip_vs_conn_put(cp);
}
+/* Check for incompatible flags */
+static bool ip_vs_sync_validate_flags(u32 flags)
+{
+ /* We do not expect NO_CPORT, especially to allow lookups
+ * to hit templates
+ */
+ if (flags & IP_VS_CONN_F_NO_CPORT) {
+ if (flags & IP_VS_CONN_F_TEMPLATE)
+ return false;
+ }
+ if (flags & IP_VS_CONN_F_ONE_PACKET)
+ return false;
+ return true;
+}
+
/*
* Process received multicast message for Version 0
*/
@@ -972,8 +987,7 @@ static void ip_vs_process_message_v0(struct netns_ipvs *ipvs, const char *buffer
return;
}
s = (struct ip_vs_sync_conn_v0 *) p;
- flags = ntohs(s->flags) | IP_VS_CONN_F_SYNC;
- flags &= ~IP_VS_CONN_F_HASHED;
+ flags = ntohs(s->flags);
if (flags & IP_VS_CONN_F_SEQ_MASK) {
opt = (struct ip_vs_sync_conn_options *)&s[1];
p += FULL_CONN_SIZE;
@@ -986,6 +1000,13 @@ static void ip_vs_process_message_v0(struct netns_ipvs *ipvs, const char *buffer
p += SIMPLE_CONN_SIZE;
}
+ if (!ip_vs_sync_validate_flags(flags)) {
+ IP_VS_DBG(2, "BACKUP v0, Invalid flags 0x%X\n", flags);
+ continue;
+ }
+ flags &= IP_VS_CONN_F_BACKUP_MASK;
+ flags |= IP_VS_CONN_F_SYNC;
+
state = ntohs(s->state);
if (!(flags & IP_VS_CONN_F_TEMPLATE)) {
pp = ip_vs_proto_get(s->protocol);
@@ -1141,7 +1162,13 @@ static inline int ip_vs_proc_sync_conn(struct netns_ipvs *ipvs, __u8 *p, __u8 *m
}
/* Get flags and Mask off unsupported */
- flags = ntohl(s->v4.flags) & IP_VS_CONN_F_BACKUP_MASK;
+ flags = ntohl(s->v4.flags);
+ if (!ip_vs_sync_validate_flags(flags)) {
+ IP_VS_DBG(3, "BACKUP, Invalid flags 0x%X\n", flags);
+ retc = 25;
+ goto out;
+ }
+ flags &= IP_VS_CONN_F_BACKUP_MASK;
flags |= IP_VS_CONN_F_SYNC;
state = ntohs(s->v4.state);
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH net 07/10] netfilter: nft_set_rbtree: skip transaction elements during GC
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
` (5 preceding siblings ...)
2026-09-30 7:41 ` [PATCH net 06/10] ipvs: filter some flags received in the backup server Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 08/10] netfilter: bpf: reject invalid NAT manipulation types Pablo Neira Ayuso
` (2 subsequent siblings)
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
From: Weiming Shi <bestswngs@gmail.com>
Since nft_set_commit_update() runs set commit callbacks before processing
NEWSETELEM transactions, nft_rbtree_gc_scan() can observe elements added by
the transaction being committed.
The scan records an interval end in rbe_end without checking the element's
transaction state. A later, unrelated expired start then moves both
elements to the expired list. The synchronous GC queue can free the new end
element before the transaction subsequently activates it, causing a
use-after-free.
Only consider elements that are fully active in both generations. This
keeps transaction-state elements out of the GC scan and preserves interval
pairing across skipped elements.
KASAN reports:
BUG: KASAN: slab-use-after-free in nft_setelem_activate
nft_setelem_activate net/netfilter/nf_tables_api.c:7047
nf_tables_commit net/netfilter/nf_tables_api.c:11137
Allocated by task 130:
nft_set_elem_init net/netfilter/nf_tables_api.c:6794
nft_add_set_elem net/netfilter/nf_tables_api.c:7523
Freed by task 130:
nft_trans_gc_trans_free net/netfilter/nf_tables_api.c:10506
rcu_core kernel/rcu/tree.c:2919
Fixes: 1e3b9e1c77fe ("netfilter: nf_tables: call set ops .commit when building new ruleset blob")
Reported-by: <co+ee5e50ef2670e5f4@bugs.sh>
Assisted-by: LLM
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/nft_set_rbtree.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/net/netfilter/nft_set_rbtree.c b/net/netfilter/nft_set_rbtree.c
index 9894832281c4..12431b55752f 100644
--- a/net/netfilter/nft_set_rbtree.c
+++ b/net/netfilter/nft_set_rbtree.c
@@ -900,6 +900,8 @@ static void nft_rbtree_gc_scan(struct nft_set *set)
next = rb_next(node);
rbe = rb_entry(node, struct nft_rbtree_elem, node);
+ if (!nft_set_elem_active(&rbe->ext, NFT_GENMASK_ANY))
+ continue;
/* elements are reversed in the rbtree for historical reasons,
* from highest to lowest value, that is why end element is
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH net 08/10] netfilter: bpf: reject invalid NAT manipulation types
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
` (6 preceding siblings ...)
2026-09-30 7:41 ` [PATCH net 07/10] netfilter: nft_set_rbtree: skip transaction elements during GC Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 09/10] netfilter: flowtable: generalize pending status bit Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 10/10] netfilter: flowtable: restore ieee80211 forward path Pablo Neira Ayuso
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
From: Fernando Fernandez Mancera <fmancera@suse.de>
As bpf_ct_set_nat_info() is not validating the NAT manipulation type a
wrong value can be passed directly to nf_nat_setup_info(). This triggers
the WARN_ON() at nf_nat_setup_info() and if panic_on_warn isn't set,
then IPS_SRC_NAT_DONE is set without adding nat_bysource and conntrack
cleanup tries to unlink an uninitialized hlist node.
Fix this by checking that NAT manipulation type is correct before
calling nf_nat_setup_info(). In addition, if the WARN_ON is hit, return
NF_DROP instead of continuing with the processing to avoid similar
situations in the future.
Reported-by: VEGA <vega@nebusec.ai>
Fixes: 0fabd2aa199f ("net: netfilter: add bpf_ct_set_nat_info kfunc helper")
Signed-off-by: Fernando Fernandez Mancera <fmancera@suse.de>
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
net/netfilter/nf_nat_bpf.c | 3 +++
net/netfilter/nf_nat_core.c | 5 +++--
2 files changed, 6 insertions(+), 2 deletions(-)
diff --git a/net/netfilter/nf_nat_bpf.c b/net/netfilter/nf_nat_bpf.c
index f9dd85ccea01..7572b58c448b 100644
--- a/net/netfilter/nf_nat_bpf.c
+++ b/net/netfilter/nf_nat_bpf.c
@@ -39,6 +39,9 @@ __bpf_kfunc int bpf_ct_set_nat_info(struct nf_conn___init *nfct,
if (proto != NFPROTO_IPV4 && proto != NFPROTO_IPV6)
return -EINVAL;
+ if (manip != NF_NAT_MANIP_SRC && manip != NF_NAT_MANIP_DST)
+ return -EINVAL;
+
memset(&range, 0, sizeof(struct nf_nat_range2));
range.flags = NF_NAT_RANGE_MAP_IPS;
range.min_addr = *addr;
diff --git a/net/netfilter/nf_nat_core.c b/net/netfilter/nf_nat_core.c
index a4858c2b2d65..cc8e1e81006d 100644
--- a/net/netfilter/nf_nat_core.c
+++ b/net/netfilter/nf_nat_core.c
@@ -767,8 +767,9 @@ nf_nat_setup_info(struct nf_conn *ct,
if (nf_ct_is_confirmed(ct))
return NF_ACCEPT;
- WARN_ON(maniptype != NF_NAT_MANIP_SRC &&
- maniptype != NF_NAT_MANIP_DST);
+ if (WARN_ON(maniptype != NF_NAT_MANIP_SRC &&
+ maniptype != NF_NAT_MANIP_DST))
+ return NF_DROP;
if (WARN_ON(nf_nat_initialized(ct, maniptype)))
return NF_DROP;
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH net 09/10] netfilter: flowtable: generalize pending status bit
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
` (7 preceding siblings ...)
2026-09-30 7:41 ` [PATCH net 08/10] netfilter: bpf: reject invalid NAT manipulation types Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
2026-09-30 7:41 ` [PATCH net 10/10] netfilter: flowtable: restore ieee80211 forward path Pablo Neira Ayuso
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
Rename NF_FLOW_HW_PENDING to NF_FLOW_PENDING and use it to inhibit the
flowtable GC worker until pending hw offload work has been completed.
Apparently, nf_flow_offload_stats() can schedule work to retrieve stats
while the flow is being removed by GC.
And this bit can also be used in a follow up patch to disable GC until
the flow has been fully added in both directions.
Revert the reordering done in commit d644b23afe1e ("netfilter:
flowtable: publish HW_DEAD after worker is done") to prevent a race
between GC and hw offload handler.
Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/net/netfilter/nf_flow_table.h | 2 +-
net/netfilter/nf_flow_table_core.c | 7 ++++++-
net/netfilter/nf_flow_table_offload.c | 14 +++++---------
net/sched/act_ct.c | 2 +-
4 files changed, 13 insertions(+), 12 deletions(-)
diff --git a/include/net/netfilter/nf_flow_table.h b/include/net/netfilter/nf_flow_table.h
index f2e2771f188f..5b611efaa3cd 100644
--- a/include/net/netfilter/nf_flow_table.h
+++ b/include/net/netfilter/nf_flow_table.h
@@ -183,10 +183,10 @@ enum nf_flow_flags {
NF_FLOW_DNAT,
NF_FLOW_CLOSING,
NF_FLOW_TEARDOWN,
+ NF_FLOW_PENDING,
NF_FLOW_HW,
NF_FLOW_HW_DYING,
NF_FLOW_HW_DEAD,
- NF_FLOW_HW_PENDING,
NF_FLOW_HW_BIDIRECTIONAL,
NF_FLOW_HW_ESTABLISHED,
};
diff --git a/net/netfilter/nf_flow_table_core.c b/net/netfilter/nf_flow_table_core.c
index 934c6151f558..36bbc7be2f74 100644
--- a/net/netfilter/nf_flow_table_core.c
+++ b/net/netfilter/nf_flow_table_core.c
@@ -575,7 +575,12 @@ static void nf_flow_table_extend_ct_timeout(struct nf_conn *ct)
static void nf_flow_offload_gc_step(struct nf_flowtable *flow_table,
struct flow_offload *flow, void *data)
{
- bool teardown = test_bit(NF_FLOW_TEARDOWN, &flow->flags);
+ bool teardown;
+
+ if (test_bit(NF_FLOW_PENDING, &flow->flags))
+ return;
+
+ teardown = test_bit(NF_FLOW_TEARDOWN, &flow->flags);
if (nf_flow_has_expired(flow) ||
nf_ct_is_dying(flow->ct) ||
diff --git a/net/netfilter/nf_flow_table_offload.c b/net/netfilter/nf_flow_table_offload.c
index 6757fd89c1f1..4365859220e6 100644
--- a/net/netfilter/nf_flow_table_offload.c
+++ b/net/netfilter/nf_flow_table_offload.c
@@ -995,6 +995,7 @@ static void flow_offload_work_del(struct flow_offload_work *offload)
flow_offload_tuple_del(offload, FLOW_OFFLOAD_DIR_ORIGINAL);
if (test_bit(NF_FLOW_HW_BIDIRECTIONAL, &offload->flow->flags))
flow_offload_tuple_del(offload, FLOW_OFFLOAD_DIR_REPLY);
+ set_bit(NF_FLOW_HW_DEAD, &offload->flow->flags);
}
static void flow_offload_tuple_stats(struct flow_offload_work *offload,
@@ -1056,13 +1057,8 @@ static void flow_offload_work_handler(struct work_struct *work)
default:
WARN_ON_ONCE(1);
}
-
- clear_bit(NF_FLOW_HW_PENDING, &offload->flow->flags);
- if (offload->cmd == FLOW_CLS_DESTROY) {
- /* Publish after the worker's last flow access. */
- smp_mb__before_atomic();
- set_bit(NF_FLOW_HW_DEAD, &offload->flow->flags);
- }
+ smp_mb__before_atomic();
+ clear_bit(NF_FLOW_PENDING, &offload->flow->flags);
kfree(offload);
}
@@ -1089,12 +1085,12 @@ nf_flow_offload_work_alloc(struct nf_flowtable *flowtable,
{
struct flow_offload_work *offload;
- if (test_and_set_bit(NF_FLOW_HW_PENDING, &flow->flags))
+ if (test_and_set_bit(NF_FLOW_PENDING, &flow->flags))
return NULL;
offload = kmalloc_obj(struct flow_offload_work, GFP_ATOMIC);
if (!offload) {
- clear_bit(NF_FLOW_HW_PENDING, &flow->flags);
+ clear_bit(NF_FLOW_PENDING, &flow->flags);
return NULL;
}
diff --git a/net/sched/act_ct.c b/net/sched/act_ct.c
index 55f3521edb4c..626c9a5af0ef 100644
--- a/net/sched/act_ct.c
+++ b/net/sched/act_ct.c
@@ -289,7 +289,7 @@ static bool tcf_ct_flow_is_outdated(const struct flow_offload *flow)
{
return test_bit(IPS_SEEN_REPLY_BIT, &flow->ct->status) &&
test_bit(IPS_HW_OFFLOAD_BIT, &flow->ct->status) &&
- !test_bit(NF_FLOW_HW_PENDING, &flow->flags) &&
+ !test_bit(NF_FLOW_PENDING, &flow->flags) &&
!test_bit(NF_FLOW_HW_ESTABLISHED, &flow->flags);
}
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread* [PATCH net 10/10] netfilter: flowtable: restore ieee80211 forward path
2026-09-30 7:41 [PATCH net,v2 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
` (8 preceding siblings ...)
2026-09-30 7:41 ` [PATCH net 09/10] netfilter: flowtable: generalize pending status bit Pablo Neira Ayuso
@ 2026-09-30 7:41 ` Pablo Neira Ayuso
9 siblings, 0 replies; 13+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-30 7:41 UTC (permalink / raw)
To: netfilter-devel; +Cc: davem, netdev, kuba, pabeni, edumazet, horms, fw, ja
Before commit 871df5007eda ("netfilter: flowtable: bail out if forward
path cannot be discovered"), there was a fallback to set up a forward
path in case .ndo_fill_forward_path fails or DEV_PATH_MTK_WDMA was used.
Such fallback was used by commit d787a3e38f01 ("mac80211: add support
for .ndo_fill_forward_path").
One possibility is to handle DEV_PATH_MTK_WDMA from the flowtable
forward path discovery. However, this is only used internally by drivers
to retrieve mtk_wdma information to set up hardware offload. Felix
decided to use the .fill_forward_path interface for this purpose due to
the lack of a better interface at that time.
Add a new DEV_PATH_IEEE80211 path which is offered if the new ieee80211
flag is set on in the struct net_device_path_ctx to restore the
flowtable with a ieee80211 netdevice. Handle this new DEV_PATH_IEEE80211
path just like DEV_PATH_ETHERNET and DEV_PATH_DSA, ie. this is the last
netdevice in the stack.
This new ieee80211 flag is implicitly unset for mtk_ppe and airoha which
call dev_fill_forward_path() to retrieve a DEV_PATH_MTK_WDMA path.
Fixes: 871df5007eda ("netfilter: flowtable: bail out if forward path cannot be discovered")
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/linux/netdevice.h | 3 +++
net/mac80211/iface.c | 7 +++++++
net/netfilter/nf_flow_table_path.c | 3 +++
3 files changed, 13 insertions(+)
diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index 87cafc932e9e..3cff2174dc03 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -887,6 +887,7 @@ enum net_device_path_type {
DEV_PATH_DSA,
DEV_PATH_MTK_WDMA,
DEV_PATH_TUN,
+ DEV_PATH_IEEE80211,
};
struct net_device_path {
@@ -953,6 +954,8 @@ struct net_device_path_ctx {
u16 id;
__be16 proto;
} vlan[NET_DEVICE_PATH_VLAN_MAX];
+
+ bool ieee80211;
};
enum tc_setup_type {
diff --git a/net/mac80211/iface.c b/net/mac80211/iface.c
index 889c32fd8de1..c5584435fde9 100644
--- a/net/mac80211/iface.c
+++ b/net/mac80211/iface.c
@@ -1023,6 +1023,13 @@ static int ieee80211_netdev_fill_forward_path(struct net_device_path_ctx *ctx,
struct sta_info *sta;
int ret = -ENOENT;
+ if (ctx->ieee80211) {
+ path->type = DEV_PATH_IEEE80211;
+ path->dev = ctx->dev;
+ ctx->dev = NULL;
+ return 0;
+ }
+
sdata = IEEE80211_DEV_TO_SUB_IF(ctx->dev);
local = sdata->local;
diff --git a/net/netfilter/nf_flow_table_path.c b/net/netfilter/nf_flow_table_path.c
index 1e55644f2edb..d90013685bf1 100644
--- a/net/netfilter/nf_flow_table_path.c
+++ b/net/netfilter/nf_flow_table_path.c
@@ -53,6 +53,7 @@ static int nft_dev_fill_forward_path(const struct dst_entry *dst_cache,
struct net_device_path_ctx ctx = {
.dev = dev,
.ether_type = ether_type,
+ .ieee80211 = true,
};
struct neighbour *n;
u8 nud_state;
@@ -114,6 +115,7 @@ static int nft_dev_path_info(struct net_device_path_stack *stack,
path = &stack->path[i];
switch (path->type) {
case DEV_PATH_ETHERNET:
+ case DEV_PATH_IEEE80211:
case DEV_PATH_DSA:
case DEV_PATH_VLAN:
case DEV_PATH_PPPOE:
@@ -123,6 +125,7 @@ static int nft_dev_path_info(struct net_device_path_stack *stack,
memcpy(info->h_source, path->dev->dev_addr, ETH_ALEN);
if (path->type == DEV_PATH_ETHERNET ||
+ path->type == DEV_PATH_IEEE80211 ||
path->type == DEV_PATH_DSA)
break;
--
2.47.3
^ permalink raw reply related [flat|nested] 13+ messages in thread