Netdev List
 help / color / mirror / Atom feed
From: Jakub Kicinski <kuba@kernel.org>
To: pablo@netfilter.org
Cc: Jakub Kicinski <kuba@kernel.org>,
	netfilter-devel@vger.kernel.org, davem@davemloft.net,
	netdev@vger.kernel.org, pabeni@redhat.com, edumazet@google.com,
	horms@kernel.org, fw@strlen.de, ja@ssi.bg
Subject: Re: [PATCH net-next 01/12] netfilter: add DEBUG_NET_WARN_ON_ONCE to skb_set_nfct()
Date: Tue, 11 Aug 2026 18:34:41 -0700	[thread overview]
Message-ID: <20260812013441.1673114-1-kuba@kernel.org> (raw)
In-Reply-To: <20260810194015.932627-2-pablo@netfilter.org>

This is an AI-generated review of your patch. The human sending this
email has considered the AI review valid, or at least plausible.
---
netfilter: add DEBUG_NET_WARN_ON_ONCE to skb_set_nfct()

This adds a DEBUG_NET_WARN_ON_ONCE() to skb_set_nfct() that fires when
skb->_nfct already holds a conntrack pointer, so that nf_ct_set() calls
overwriting a live ct object are caught.  Several existing nf_ct_set()
callers in ipvs, conntrack clash resolution, openvswitch and act_ct are
converted to call nf_reset_ct() first.

> diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
> index 22eda1d54a0e8..95184183180f6 100644
> --- a/include/linux/skbuff.h
> +++ b/include/linux/skbuff.h
> @@ -5004,6 +5004,7 @@ static inline unsigned long skb_get_nfct(const struct sk_buff *skb)
>  static inline void skb_set_nfct(struct sk_buff *skb, unsigned long nfct)
>  {
>  #if IS_ENABLED(CONFIG_NF_CONNTRACK)
> +	DEBUG_NET_WARN_ON_ONCE(skb->_nfct & NFCT_PTRMASK);
>  	skb->slow_gro |= !!nfct;
>  	skb->_nfct = nfct;
>  #endif

[Severity: High]
Was nf_ct_bridge_pre() in net/bridge/netfilter/nf_conntrack_bridge.c meant
to be converted too?

That hook deliberately lets an skb carrying a conntrack template through its
entry check:

net/bridge/netfilter/nf_conntrack_bridge.c:nf_ct_bridge_pre() {
	ct = nf_ct_get(skb, &ctinfo);
	if ((ct && !nf_ct_is_template(ct)) ||
	    ctinfo == IP_CT_UNTRACKED)
		return NF_ACCEPT;

The ETH_P_IP and ETH_P_IPV6 arms then hand the template to
nf_conntrack_in(), which zeroes skb->_nfct and drops the template reference
with nf_ct_put(tmpl) at its out: label, so those paths satisfy the new
precondition.  The default: arm does not:

	default:
		nf_ct_set(skb, NULL, IP_CT_UNTRACKED);
		return NF_ACCEPT;
	}

Since the new warning tests the old value of skb->_nfct regardless of the
new value, does nf_ct_set(skb, NULL, IP_CT_UNTRACKED) here trip
DEBUG_NET_WARN_ON_ONCE() whenever a template is still attached?

Can a template be attached before this hook runs?  An nft bridge base chain
in NF_BR_PRE_ROUTING at a priority below NF_IP_PRI_CONNTRACK (-200) appears
to be accepted, as nf_tables_api.c only restricts priorities at or below
NF_IP_PRI_CONNTRACK for NAT-type chains.  A rule with ct zone set in such a
chain reaches nft_ct_set_zone_eval():

net/netfilter/nft_ct.c:nft_ct_set_zone_eval() {
	...
	nf_ct_set(skb, ct, IP_CT_NEW);

which takes a reference on the per-CPU template (or allocates a fresh one
via nf_ct_tmpl_alloc() when that one is busy).  nft_do_chain_bridge()
evaluates rules for every ethertype, including ARP, PPPoE and LLDP frames,
which then land in the default: arm above.

With CONFIG_DEBUG_NET=y, does this turn any received non-IPv4/IPv6 frame on
such a bridge into a warning splat, and a panic when panic_on_warn=1?

Separately, the reference drop at that site is not new: the overwrite has
always discarded the template reference without nf_conntrack_put().  Does
that pin the per-CPU template refcount above 1 permanently, so
nft_ct_set_zone_eval() takes the nf_ct_tmpl_alloc() path for every later
frame, and each of those allocations is leaked by the same
nf_ct_set(skb, NULL, IP_CT_UNTRACKED)?

For reference, the same shape was converted in this patch:

> diff --git a/include/net/ip_vs.h b/include/net/ip_vs.h
> index b3bb228ad75cb..3dca7d387dd08 100644
> --- a/include/net/ip_vs.h
> +++ b/include/net/ip_vs.h
> @@ -2121,7 +2121,7 @@ static inline void ip_vs_notrack(struct sk_buff *skb)
>  	struct nf_conn *ct = nf_ct_get(skb, &ctinfo);
>  
>  	if (ct) {
> -		nf_conntrack_put(&ct->ct_general);
> +		nf_reset_ct(skb);
>  		nf_ct_set(skb, NULL, IP_CT_UNTRACKED);
>  	}
>  #endif

[ ... ]

Would adding nf_reset_ct(skb) before the nf_ct_set() in the default: arm of
nf_ct_bridge_pre() be the right fix?  If it is intentionally left for a
separate change, could the commit message mention it so that a new warning
in bridge plus conntrack zone setups is not a surprise?

Checking the tip of the series (839e6caf13c7332db9b40817696bb7a326b810d5),
nf_ct_bridge_pre() is still unchanged and the warning in skb_set_nfct() is
still present.

Two nearby nf_ct_set() users also set a ct on an skb without clearing a
possible pre-existing one: nf_conntrack_attach() in
net/netfilter/nf_conntrack_core.c and tcf_ct_flow_table_restore_skb() in
include/net/tc_act/tc_ct.h.  Both rely on the caller supplying a fresh skb.
Is that assumption worth documenting alongside the new invariant?

  reply	other threads:[~2026-08-12  1:34 UTC|newest]

Thread overview: 19+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-10 19:40 [PATCH net-next 00/12] Netfilter updates for net Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 01/12] netfilter: add DEBUG_NET_WARN_ON_ONCE to skb_set_nfct() Pablo Neira Ayuso
2026-08-12  1:34   ` Jakub Kicinski [this message]
2026-08-12  7:20     ` Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 02/12] net: pass net_device_path_ctx to dev_fill_forward_path() Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 03/12] net: netfilter: add ether_type to net_device_path_ctx and use it Pablo Neira Ayuso
2026-08-12  1:34   ` Jakub Kicinski
2026-08-10 19:40 ` [PATCH net-next 04/12] netfilter: flowtable: rename tun.l3_proto to tun.inner_proto Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 05/12] netfilter: flowtable: rename ctx.tun.proto to ctx.tun.inner_proto Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 06/12] netfilter: flowtable: store ethertype in flowtable context Pablo Neira Ayuso
2026-08-12  1:34   ` Jakub Kicinski
2026-08-10 19:40 ` [PATCH net-next 07/12] netfilter: flowtable: move ipv4 and ipv6 xmit path to function Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 08/12] netfilter: flowtable: detach layer 2 encapsulation parser from lookup Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 09/12] netfilter: nft_ct: move custom expectation support to helper Pablo Neira Ayuso
2026-08-12  1:34   ` Jakub Kicinski
2026-08-10 19:40 ` [PATCH net-next 10/12] netfilter: conntrack: always lower timeout for non-closing RST packets Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 11/12] netfilter: nf_conntrack_expect: bail out on insert dead expectations Pablo Neira Ayuso
2026-08-12  1:34   ` Jakub Kicinski
2026-08-10 19:40 ` [PATCH net-next 12/12] selftests: netfilter: conntrack_dump_flush: remove unused variables and fix typo Pablo Neira Ayuso

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260812013441.1673114-1-kuba@kernel.org \
    --to=kuba@kernel.org \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=fw@strlen.de \
    --cc=horms@kernel.org \
    --cc=ja@ssi.bg \
    --cc=netdev@vger.kernel.org \
    --cc=netfilter-devel@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=pablo@netfilter.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox