From: Pablo Neira Ayuso <pablo@netfilter.org>
To: netfilter-devel@vger.kernel.org
Cc: davem@davemloft.net, netdev@vger.kernel.org, kuba@kernel.org,
pabeni@redhat.com, edumazet@google.com, horms@kernel.org,
fw@strlen.de, ja@ssi.bg
Subject: [PATCH net-next 09/12] netfilter: nft_ct: move custom expectation support to helper
Date: Mon, 10 Aug 2026 21:40:12 +0200 [thread overview]
Message-ID: <20260810194015.932627-10-pablo@netfilter.org> (raw)
In-Reply-To: <20260810194015.932627-1-pablo@netfilter.org>
Originally, the ct expectation support called nf_ct_helper_ext_add() for
confirmed conntracks, which is invalid, triggering a splat. This was
fixed by commit 1710eb913bdc ("netfilter: nft_ct: skip expectations for
confirmed conntrack") which restricted it to unconfirmed conntracks.
However, early insertion of expectations into the expectations list when
the conntrack is unconfirmed leads to stale entries pointing to the
wrong hlist_head through .pprev due to ct extension reallocation.
Commit 7c9664351980 ("netfilter: move nat hlist_head to nf_conn") moved
the nat hlist_head to nf_conn for this reason:
1. ...
2. When reallocation of extension area occurs we need to fixup the
bysource hash head via hlist_replace_rcu.
I'd rather not increase the size of the struct nf_conn for this feature
has very limited scope: only one expectation can be created at a time
given expect_clash() will make nf_ct_expect_related() reports EBUSY.
For this reason, relax nf_ct_expect_related() not to drop packets in
case expectation creation fails, therefore, expectation creation becomes
best effort.
To address this issue, add an internal ct helper and attach it to the
conntrack entry to streamline the custom ct expectation support with
existing ct helpers.
Expose a new nf_conntrack_helper_release() function to release the
internal helper that is allocated and attached to the conntrack entry to
create the custom expectations. The nft_ct module removal always waits
for rcu grace period, then the NULL helper callback is observed after
this.
This patch also restricts the creation of expectations to different
helpers other than this custom helper that is created for this type of
expectations.
Fixes: 857b46027d6f ("netfilter: nft_ct: add ct expectations support")
Reported-by: Jaeyeong Lee <iostreampy@proton.me>
Link: https://patch.msgid.link/20260715144755.00ea7dfcd9f@proton.me
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/net/netfilter/nf_conntrack_helper.h | 1 +
net/netfilter/nf_conntrack_helper.c | 14 +-
net/netfilter/nft_ct.c | 167 +++++++++++++++-----
3 files changed, 135 insertions(+), 47 deletions(-)
diff --git a/include/net/netfilter/nf_conntrack_helper.h b/include/net/netfilter/nf_conntrack_helper.h
index bc5427d239f4..335b8c43694f 100644
--- a/include/net/netfilter/nf_conntrack_helper.h
+++ b/include/net/netfilter/nf_conntrack_helper.h
@@ -106,6 +106,7 @@ void nf_ct_helper_init(struct nf_conntrack_helper *helper,
int nf_conntrack_helper_register(struct nf_conntrack_helper *, struct nf_conntrack_helper **);
int __nf_conntrack_helper_register(struct nf_conntrack_helper *);
void nf_conntrack_helper_unregister(struct nf_conntrack_helper *);
+void nf_conntrack_helper_release(struct nf_conntrack_helper *);
int nf_conntrack_helpers_register(struct nf_conntrack_helper *, unsigned int,
struct nf_conntrack_helper **);
diff --git a/net/netfilter/nf_conntrack_helper.c b/net/netfilter/nf_conntrack_helper.c
index 506c58034761..c30ae3f203be 100644
--- a/net/netfilter/nf_conntrack_helper.c
+++ b/net/netfilter/nf_conntrack_helper.c
@@ -448,6 +448,15 @@ static bool expect_iter_me(struct nf_conntrack_expect *exp, void *data)
return this == me;
}
+void nf_conntrack_helper_release(struct nf_conntrack_helper *me)
+{
+ nf_ct_expect_iterate_destroy(expect_iter_me, me);
+
+ if (refcount_dec_and_test(&me->ct_refcnt))
+ kfree_rcu(me, rcu);
+}
+EXPORT_SYMBOL_GPL(nf_conntrack_helper_release);
+
void nf_conntrack_helper_unregister(struct nf_conntrack_helper *me)
{
mutex_lock(&nf_ct_helper_mutex);
@@ -463,10 +472,7 @@ void nf_conntrack_helper_unregister(struct nf_conntrack_helper *me)
*/
synchronize_rcu();
- nf_ct_expect_iterate_destroy(expect_iter_me, me);
-
- if (refcount_dec_and_test(&me->ct_refcnt))
- kfree_rcu(me, rcu);
+ nf_conntrack_helper_release(me);
}
EXPORT_SYMBOL_GPL(nf_conntrack_helper_unregister);
diff --git a/net/netfilter/nft_ct.c b/net/netfilter/nft_ct.c
index 358b9287e12e..9dbf127df9c8 100644
--- a/net/netfilter/nft_ct.c
+++ b/net/netfilter/nft_ct.c
@@ -1213,6 +1213,8 @@ struct nft_ct_expect_obj {
u8 l4proto;
u8 size;
u32 timeout;
+
+ struct nf_conntrack_helper *helper;
};
static int nft_ct_expect_timeout_get(const struct nlattr *attr, u32 *val)
@@ -1226,6 +1228,93 @@ static int nft_ct_expect_timeout_get(const struct nlattr *attr, u32 *val)
return 0;
}
+#if IS_ENABLED(CONFIG_NF_NAT)
+static void nft_ct_nat_follow_master(struct nf_conn *ct, struct nf_conntrack_expect *this)
+{
+ const struct nf_ct_helper_expectfn *expfn;
+
+ expfn = nf_ct_helper_expectfn_find_by_name("nat-follow-master");
+ if (expfn)
+ expfn->expectfn(ct, this);
+}
+#endif
+
+struct nft_ct_expect_data {
+ struct nft_ct_expect_obj obj;
+ enum ip_conntrack_dir dir;
+};
+
+static int ct_expect_help(struct sk_buff *skb, unsigned int protoff,
+ struct nf_conn *ct, enum ip_conntrack_info ctinfo)
+{
+ enum ip_conntrack_dir dir = CTINFO2DIR(ctinfo);
+ struct nft_ct_expect_data *expect_data;
+ struct nf_conntrack_expect *exp;
+ int ret = NF_ACCEPT;
+ u16 l3num;
+
+ if (nf_ct_is_confirmed(ct))
+ return NF_ACCEPT;
+
+ expect_data = nfct_help_data(ct);
+ if (!expect_data)
+ return NF_ACCEPT;
+
+ if (expect_data->dir != dir)
+ return NF_ACCEPT;
+
+ exp = nf_ct_expect_alloc(ct);
+ if (!exp)
+ return NF_DROP;
+
+ if (expect_data->obj.l3num == NFPROTO_INET)
+ l3num = nf_ct_l3num(ct);
+ else
+ l3num = expect_data->obj.l3num;
+
+ nf_ct_expect_init(exp, NF_CT_EXPECT_CLASS_DEFAULT, l3num,
+ &ct->tuplehash[!dir].tuple.src.u3,
+ &ct->tuplehash[!dir].tuple.dst.u3,
+ expect_data->obj.l4proto, NULL, &expect_data->obj.dport);
+ exp->timeout += expect_data->obj.timeout;
+
+#if IS_ENABLED(CONFIG_NF_NAT)
+ if (ct->status & IPS_NAT_MASK) {
+ exp->saved_proto.tcp.port = expect_data->obj.dport;
+ exp->dir = !dir;
+ exp->expectfn = nft_ct_nat_follow_master;
+ }
+#endif
+ if (nf_ct_expect_related(exp, 0) != 0)
+ ret = NF_ACCEPT;
+
+ nf_ct_expect_put(exp);
+
+ return ret;
+}
+
+static int nft_ct_expect_helper_alloc(struct nft_ct_expect_obj *priv)
+{
+ struct nf_conntrack_helper *ct_expect_helper;
+
+ ct_expect_helper = kzalloc_obj(struct nf_conntrack_helper,
+ GFP_KERNEL_ACCOUNT);
+ if (!ct_expect_helper)
+ return -ENOMEM;
+
+ snprintf(ct_expect_helper->name, sizeof(ct_expect_helper->name), "%s",
+ "nft_ct_expect");
+ ct_expect_helper->me = THIS_MODULE;
+ ct_expect_helper->expect_policy[NF_CT_EXPECT_CLASS_DEFAULT].max_expected = priv->size;
+ rcu_assign_pointer(ct_expect_helper->help, ct_expect_help);
+ refcount_set(&ct_expect_helper->ct_refcnt, 1);
+
+ /* No need to register this helper, this is internal. */
+ priv->helper = ct_expect_helper;
+
+ return 0;
+}
+
static int nft_ct_expect_obj_init(const struct nft_ctx *ctx,
const struct nlattr * const tb[],
struct nft_object *obj)
@@ -1233,6 +1322,8 @@ static int nft_ct_expect_obj_init(const struct nft_ctx *ctx,
struct nft_ct_expect_obj *priv = nft_obj_data(obj);
int err;
+ NF_CT_HELPER_BUILD_BUG_ON(sizeof(struct nft_ct_expect_data));
+
if (!tb[NFTA_CT_EXPECT_L4PROTO] ||
!tb[NFTA_CT_EXPECT_DPORT] ||
!tb[NFTA_CT_EXPECT_TIMEOUT] ||
@@ -1272,13 +1363,31 @@ static int nft_ct_expect_obj_init(const struct nft_ctx *ctx,
priv->dport = nla_get_be16(tb[NFTA_CT_EXPECT_DPORT]);
priv->size = nla_get_u8(tb[NFTA_CT_EXPECT_SIZE]);
+ if (!priv->size)
+ priv->size = NF_CT_EXPECT_MAX_CNT;
+
+ err = nf_ct_netns_get(ctx->net, ctx->family);
+ if (err < 0)
+ return err;
- return nf_ct_netns_get(ctx->net, ctx->family);
+ err = nft_ct_expect_helper_alloc(priv);
+ if (err < 0) {
+ nf_ct_netns_put(ctx->net, ctx->family);
+ return err;
+ }
+
+ return err;
}
static void nft_ct_expect_obj_destroy(const struct nft_ctx *ctx,
- struct nft_object *obj)
+ struct nft_object *obj)
{
+ const struct nft_ct_expect_obj *priv = nft_obj_data(obj);
+ struct nf_conntrack_helper *me = priv->helper;
+
+ /* This helper is going away, disable it. */
+ rcu_assign_pointer(me->help, NULL);
+ nf_conntrack_helper_release(me);
nf_ct_netns_put(ctx->net, ctx->family);
}
@@ -1297,27 +1406,14 @@ static int nft_ct_expect_obj_dump(struct sk_buff *skb,
return 0;
}
-#if IS_ENABLED(CONFIG_NF_NAT)
-static void nft_ct_nat_follow_master(struct nf_conn *ct, struct nf_conntrack_expect *this)
-{
- const struct nf_ct_helper_expectfn *expfn;
-
- expfn = nf_ct_helper_expectfn_find_by_name("nat-follow-master");
- if (expfn)
- expfn->expectfn(ct, this);
-}
-#endif
-
static void nft_ct_expect_obj_eval(struct nft_object *obj,
struct nft_regs *regs,
const struct nft_pktinfo *pkt)
{
const struct nft_ct_expect_obj *priv = nft_obj_data(obj);
- struct nf_conntrack_expect *exp;
+ struct nft_ct_expect_data *expect_data;
enum ip_conntrack_info ctinfo;
struct nf_conn_help *help;
- enum ip_conntrack_dir dir;
- u16 l3num = priv->l3num;
struct nf_conn *ct;
ct = nf_ct_get(pkt->skb, &ctinfo);
@@ -1325,45 +1421,30 @@ static void nft_ct_expect_obj_eval(struct nft_object *obj,
regs->verdict.code = NFT_BREAK;
return;
}
- dir = CTINFO2DIR(ctinfo);
help = nfct_help(ct);
- if (!help)
- help = nf_ct_helper_ext_add(ct, GFP_ATOMIC);
- if (!help) {
- regs->verdict.code = NF_DROP;
- return;
- }
-
- if (help->expecting[NF_CT_EXPECT_CLASS_DEFAULT] >= priv->size) {
+ if (help) {
regs->verdict.code = NFT_BREAK;
return;
}
- if (l3num == NFPROTO_INET)
- l3num = nf_ct_l3num(ct);
- exp = nf_ct_expect_alloc(ct);
- if (exp == NULL) {
+ help = nf_ct_helper_ext_add(ct, GFP_ATOMIC);
+ if (!help) {
regs->verdict.code = NF_DROP;
return;
}
- nf_ct_expect_init(exp, NF_CT_EXPECT_CLASS_DEFAULT, l3num,
- &ct->tuplehash[!dir].tuple.src.u3,
- &ct->tuplehash[!dir].tuple.dst.u3,
- priv->l4proto, NULL, &priv->dport);
- exp->timeout += priv->timeout;
-#if IS_ENABLED(CONFIG_NF_NAT)
- if (ct->status & IPS_NAT_MASK) {
- exp->saved_proto.tcp.port = priv->dport;
- exp->dir = !dir;
- exp->expectfn = nft_ct_nat_follow_master;
+ expect_data = nfct_help_data(ct);
+ if (!expect_data) {
+ regs->verdict.code = NFT_BREAK;
+ return;
}
-#endif
- if (nf_ct_expect_related(exp, 0) != 0)
- regs->verdict.code = NF_DROP;
+ expect_data->obj = *priv;
+ expect_data->obj.helper = NULL;
+ expect_data->dir = CTINFO2DIR(ctinfo);
- nf_ct_expect_put(exp);
+ if (help && refcount_inc_not_zero(&priv->helper->ct_refcnt))
+ rcu_assign_pointer(help->helper, priv->helper);
}
static const struct nla_policy nft_ct_expect_policy[NFTA_CT_EXPECT_MAX + 1] = {
--
2.47.3
next prev parent reply other threads:[~2026-08-10 19:40 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-10 19:40 [PATCH net-next 00/12] Netfilter updates for net Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 01/12] netfilter: add DEBUG_NET_WARN_ON_ONCE to skb_set_nfct() Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 02/12] net: pass net_device_path_ctx to dev_fill_forward_path() Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 03/12] net: netfilter: add ether_type to net_device_path_ctx and use it Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 04/12] netfilter: flowtable: rename tun.l3_proto to tun.inner_proto Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 05/12] netfilter: flowtable: rename ctx.tun.proto to ctx.tun.inner_proto Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 06/12] netfilter: flowtable: store ethertype in flowtable context Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 07/12] netfilter: flowtable: move ipv4 and ipv6 xmit path to function Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 08/12] netfilter: flowtable: detach layer 2 encapsulation parser from lookup Pablo Neira Ayuso
2026-08-10 19:40 ` Pablo Neira Ayuso [this message]
2026-08-10 19:40 ` [PATCH net-next 10/12] netfilter: conntrack: always lower timeout for non-closing RST packets Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 11/12] netfilter: nf_conntrack_expect: bail out on insert dead expectations Pablo Neira Ayuso
2026-08-10 19:40 ` [PATCH net-next 12/12] selftests: netfilter: conntrack_dump_flush: remove unused variables and fix typo Pablo Neira Ayuso
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260810194015.932627-10-pablo@netfilter.org \
--to=pablo@netfilter.org \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=fw@strlen.de \
--cc=horms@kernel.org \
--cc=ja@ssi.bg \
--cc=kuba@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=netfilter-devel@vger.kernel.org \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox