From: Pablo Neira Ayuso <pablo@netfilter.org>
To: netfilter-devel@vger.kernel.org
Cc: davem@davemloft.net, netdev@vger.kernel.org, kuba@kernel.org,
pabeni@redhat.com, edumazet@google.com, fw@strlen.de,
horms@kernel.org
Subject: [PATCH net 10/10] netfilter: nft_ct: move custom expectation support to helper
Date: Fri, 31 Jul 2026 17:18:06 +0200 [thread overview]
Message-ID: <20260731151806.849724-11-pablo@netfilter.org> (raw)
In-Reply-To: <20260731151806.849724-1-pablo@netfilter.org>
Originally, the ct expectation support called nf_ct_helper_ext_add() for
confirmed conntracks, which is invalid, triggering a splat. This was
fixed by commit 1710eb913bdc ("netfilter: nft_ct: skip expectations for
confirmed conntrack") which restricted it to confirmed conntracks.
However, early insertion of expectations into the expectations list when
the conntrack is unconfirmed leads to stale entries pointing to the
wrong hlist_head through .pprev due to ct extension reallocation.
Commit 7c9664351980 ("netfilter: move nat hlist_head to nf_conn") moved
the nat hlist_head to nf_conn for this reason:
1. ...
2. When reallocation of extension area occurs we need to fixup the
bysource hash head via hlist_replace_rcu.
But I'd rather not increase the size of the struct nf_conn for this
feature, it only supports for creating expectations in the other
direction and it was broken with DNAT too.
The existing feature has very limited scope because of a pre-existing
issue: two different connections can create the same expectation leading
to expect_clash(), resulting in packet drops.
To address this issue, add an internal ct helper and attach it to the
conntrack entry to streamline the custom ct expectation support with
existing ct helpers.
Expose a new nf_conntrack_helper_free() function to safely release the
internal helper that is allocated and attached to the conntrack entry to
create the custom expectations.
Fixes: 857b46027d6f ("netfilter: nft_ct: add ct expectations support")
Reported-by: Jaeyeong Lee <iostreampy@proton.me>
Link: https://patch.msgid.link/20260715144755.00ea7dfcd9f@proton.me
Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org>
---
include/net/netfilter/nf_conntrack_helper.h | 2 +
net/netfilter/nf_conntrack_helper.c | 18 ++-
net/netfilter/nft_ct.c | 127 +++++++++++++++-----
3 files changed, 114 insertions(+), 33 deletions(-)
diff --git a/include/net/netfilter/nf_conntrack_helper.h b/include/net/netfilter/nf_conntrack_helper.h
index c761cd8158b2..4420846e41f2 100644
--- a/include/net/netfilter/nf_conntrack_helper.h
+++ b/include/net/netfilter/nf_conntrack_helper.h
@@ -114,6 +114,8 @@ int nf_conntrack_helpers_register(struct nf_conntrack_helper *, unsigned int,
void nf_conntrack_helpers_unregister(struct nf_conntrack_helper **,
unsigned int);
+void nf_conntrack_helper_free(struct nf_conntrack_helper *me);
+
#define nf_conntrack_helper_deprecated(name) \
pr_warn("The %s conntrack helper is scheduled for removal.\n" \
"Please contact the netfilter-devel mailing list if you still need this.\n", name)
diff --git a/net/netfilter/nf_conntrack_helper.c b/net/netfilter/nf_conntrack_helper.c
index 500509b17663..1197e8793494 100644
--- a/net/netfilter/nf_conntrack_helper.c
+++ b/net/netfilter/nf_conntrack_helper.c
@@ -456,13 +456,8 @@ static bool expect_iter_me(struct nf_conntrack_expect *exp, void *data)
return this == me;
}
-void nf_conntrack_helper_unregister(struct nf_conntrack_helper *me)
+void nf_conntrack_helper_free(struct nf_conntrack_helper *me)
{
- mutex_lock(&nf_ct_helper_mutex);
- hlist_del_rcu(&me->hnode);
- nf_ct_helper_count--;
- mutex_unlock(&nf_ct_helper_mutex);
-
/* This helper is going away, disable it. */
rcu_assign_pointer(me->help, NULL);
@@ -476,6 +471,17 @@ void nf_conntrack_helper_unregister(struct nf_conntrack_helper *me)
if (refcount_dec_and_test(&me->ct_refcnt))
kfree_rcu(me, rcu);
}
+EXPORT_SYMBOL_GPL(nf_conntrack_helper_free);
+
+void nf_conntrack_helper_unregister(struct nf_conntrack_helper *me)
+{
+ mutex_lock(&nf_ct_helper_mutex);
+ hlist_del_rcu(&me->hnode);
+ nf_ct_helper_count--;
+ mutex_unlock(&nf_ct_helper_mutex);
+
+ nf_conntrack_helper_free(me);
+}
EXPORT_SYMBOL_GPL(nf_conntrack_helper_unregister);
void nf_ct_helper_init(struct nf_conntrack_helper *helper,
diff --git a/net/netfilter/nft_ct.c b/net/netfilter/nft_ct.c
index 03a88c77e0f0..30c9358dbf48 100644
--- a/net/netfilter/nft_ct.c
+++ b/net/netfilter/nft_ct.c
@@ -1213,6 +1213,8 @@ struct nft_ct_expect_obj {
u8 l4proto;
u8 size;
u32 timeout;
+
+ struct nf_conntrack_helper *helper;
};
static int nft_ct_expect_timeout_get(const struct nlattr *attr, u32 *val)
@@ -1226,6 +1228,73 @@ static int nft_ct_expect_timeout_get(const struct nlattr *attr, u32 *val)
return 0;
}
+struct nft_ct_expect_data {
+ struct nft_ct_expect_obj obj;
+ enum ip_conntrack_dir dir;
+ atomic_t num_expects;
+};
+
+static int ct_expect_help(struct sk_buff *skb, unsigned int protoff,
+ struct nf_conn *ct, enum ip_conntrack_info ctinfo)
+{
+ enum ip_conntrack_dir dir = CTINFO2DIR(ctinfo);
+ struct nft_ct_expect_data *expect_data;
+ struct nf_conntrack_expect *exp;
+ int ret = NF_ACCEPT;
+
+ expect_data = nfct_help_data(ct);
+ if (!expect_data)
+ return NF_ACCEPT;
+
+ if (expect_data->dir != dir)
+ return NF_ACCEPT;
+
+ if (!atomic_add_unless(&expect_data->num_expects, 1, expect_data->obj.size))
+ return NF_ACCEPT;
+
+ exp = nf_ct_expect_alloc(ct);
+ if (!exp) {
+ atomic_dec(&expect_data->num_expects);
+ return NF_DROP;
+ }
+
+ nf_ct_expect_init(exp, NF_CT_EXPECT_CLASS_DEFAULT, nf_ct_l3num(ct),
+ &ct->tuplehash[!dir].tuple.src.u3,
+ &ct->tuplehash[!dir].tuple.dst.u3,
+ expect_data->obj.l4proto, NULL, &expect_data->obj.dport);
+ exp->timeout += expect_data->obj.timeout;
+
+ if (nf_ct_expect_related(exp, 0) != 0) {
+ atomic_dec(&expect_data->num_expects);
+ ret = NF_DROP;
+ }
+
+ nf_ct_expect_put(exp);
+
+ return ret;
+}
+
+static int nft_ct_expect_helper_alloc(struct nft_ct_expect_obj *priv)
+{
+ struct nf_conntrack_helper *ct_expect_helper;
+
+ ct_expect_helper = kzalloc_obj(struct nf_conntrack_helper);
+ if (!ct_expect_helper)
+ return -ENOMEM;
+
+ snprintf(ct_expect_helper->name, sizeof(ct_expect_helper->name), "%s",
+ "nft_ct_expect");
+ ct_expect_helper->me = THIS_MODULE;
+ ct_expect_helper->expect_policy[NF_CT_EXPECT_CLASS_DEFAULT].max_expected = priv->size;
+ rcu_assign_pointer(ct_expect_helper->help, ct_expect_help);
+ refcount_set(&ct_expect_helper->ct_refcnt, 1);
+
+ /* No need to register this helper, this is internal. */
+ priv->helper = ct_expect_helper;
+
+ return 0;
+}
+
static int nft_ct_expect_obj_init(const struct nft_ctx *ctx,
const struct nlattr * const tb[],
struct nft_object *obj)
@@ -1233,6 +1302,8 @@ static int nft_ct_expect_obj_init(const struct nft_ctx *ctx,
struct nft_ct_expect_obj *priv = nft_obj_data(obj);
int err;
+ NF_CT_HELPER_BUILD_BUG_ON(sizeof(struct nft_ct_expect_data));
+
if (!tb[NFTA_CT_EXPECT_L4PROTO] ||
!tb[NFTA_CT_EXPECT_DPORT] ||
!tb[NFTA_CT_EXPECT_TIMEOUT] ||
@@ -1273,13 +1344,26 @@ static int nft_ct_expect_obj_init(const struct nft_ctx *ctx,
priv->dport = nla_get_be16(tb[NFTA_CT_EXPECT_DPORT]);
priv->size = nla_get_u8(tb[NFTA_CT_EXPECT_SIZE]);
- return nf_ct_netns_get(ctx->net, ctx->family);
+ err = nf_ct_netns_get(ctx->net, ctx->family);
+ if (err < 0)
+ return err;
+
+ err = nft_ct_expect_helper_alloc(priv);
+ if (err < 0) {
+ nf_ct_netns_put(ctx->net, ctx->family);
+ return err;
+ }
+
+ return err;
}
static void nft_ct_expect_obj_destroy(const struct nft_ctx *ctx,
- struct nft_object *obj)
+ struct nft_object *obj)
{
+ const struct nft_ct_expect_obj *priv = nft_obj_data(obj);
+
nf_ct_netns_put(ctx->net, ctx->family);
+ nf_conntrack_helper_free(priv->helper);
}
static int nft_ct_expect_obj_dump(struct sk_buff *skb,
@@ -1302,50 +1386,39 @@ static void nft_ct_expect_obj_eval(struct nft_object *obj,
const struct nft_pktinfo *pkt)
{
const struct nft_ct_expect_obj *priv = nft_obj_data(obj);
- struct nf_conntrack_expect *exp;
+ struct nft_ct_expect_data *expect_data;
enum ip_conntrack_info ctinfo;
struct nf_conn_help *help;
- enum ip_conntrack_dir dir;
- u16 l3num = priv->l3num;
struct nf_conn *ct;
ct = nf_ct_get(pkt->skb, &ctinfo);
- if (!ct || nf_ct_is_confirmed(ct) || nf_ct_is_template(ct)) {
+ if (!ct || nf_ct_is_template(ct) || nf_ct_is_confirmed(ct)) {
regs->verdict.code = NFT_BREAK;
return;
}
- dir = CTINFO2DIR(ctinfo);
help = nfct_help(ct);
- if (!help)
- help = nf_ct_helper_ext_add(ct, GFP_ATOMIC);
- if (!help) {
- regs->verdict.code = NF_DROP;
- return;
- }
-
- if (help->expecting[NF_CT_EXPECT_CLASS_DEFAULT] >= priv->size) {
+ if (help) {
regs->verdict.code = NFT_BREAK;
return;
}
- if (l3num == NFPROTO_INET)
- l3num = nf_ct_l3num(ct);
- exp = nf_ct_expect_alloc(ct);
- if (exp == NULL) {
+ help = nf_ct_helper_ext_add(ct, GFP_ATOMIC);
+ if (!help) {
regs->verdict.code = NF_DROP;
return;
}
- nf_ct_expect_init(exp, NF_CT_EXPECT_CLASS_DEFAULT, l3num,
- &ct->tuplehash[!dir].tuple.src.u3,
- &ct->tuplehash[!dir].tuple.dst.u3,
- priv->l4proto, NULL, &priv->dport);
- exp->timeout += priv->timeout;
- if (nf_ct_expect_related(exp, 0) != 0)
- regs->verdict.code = NF_DROP;
+ expect_data = nfct_help_data(ct);
+ if (!expect_data) {
+ regs->verdict.code = NFT_BREAK;
+ return;
+ }
+ expect_data->obj = *priv;
+ expect_data->dir = CTINFO2DIR(ctinfo);
- nf_ct_expect_put(exp);
+ if (help && refcount_inc_not_zero(&priv->helper->ct_refcnt))
+ rcu_assign_pointer(help->helper, priv->helper);
}
static const struct nla_policy nft_ct_expect_policy[NFTA_CT_EXPECT_MAX + 1] = {
--
2.47.3
prev parent reply other threads:[~2026-07-31 15:18 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-31 15:17 [PATCH net 00/10] Netfilter/IPVS fixes for net Pablo Neira Ayuso
2026-07-31 15:17 ` [PATCH net 01/10] ipvs: stop estimator after disabled calc phase Pablo Neira Ayuso
2026-07-31 15:17 ` [PATCH net 02/10] netfilter: ebt_nflog: pin the NFLOG backend Pablo Neira Ayuso
2026-07-31 15:17 ` [PATCH net 03/10] netfilter: ipset: rework cidr bookkeeping Pablo Neira Ayuso
2026-07-31 15:18 ` [PATCH net 04/10] netfilter: ipset: switch ext_size to atomic64_t Pablo Neira Ayuso
2026-07-31 15:18 ` [PATCH net 05/10] netfilter: ipset: add small wrappers for hash and bucket sizes Pablo Neira Ayuso
2026-07-31 15:18 ` [PATCH net 06/10] netfilter: ipset: add and use mtype_del_cidr_all helper Pablo Neira Ayuso
2026-07-31 15:18 ` [PATCH net 07/10] netfilter: ipset: switch to rcu work Pablo Neira Ayuso
2026-07-31 15:18 ` [PATCH net 08/10] ipvs: avoid out-of-bounds write in ip_vs_nat_icmp Pablo Neira Ayuso
2026-07-31 15:18 ` [PATCH net 09/10] ipvs: return the csum validation for forward hook Pablo Neira Ayuso
2026-07-31 15:18 ` Pablo Neira Ayuso [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260731151806.849724-11-pablo@netfilter.org \
--to=pablo@netfilter.org \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=fw@strlen.de \
--cc=horms@kernel.org \
--cc=kuba@kernel.org \
--cc=netdev@vger.kernel.org \
--cc=netfilter-devel@vger.kernel.org \
--cc=pabeni@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox