* [PATCH nf-next v2 0/4] netfilter: offload a TCP flow whose reply is never seen
@ 2026-10-04 17:16 Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 1/4] netfilter: conntrack: pick up a TCP flow whose SYN was never answered Julius Bairaktaris
` (3 more replies)
0 siblings, 4 replies; 5+ messages in thread
From: Julius Bairaktaris @ 2026-10-04 17:16 UTC (permalink / raw)
To: pablo, netfilter-devel
Cc: kadlec, fw, coreteam, netdev, geldot, shuah, linux-kselftest
A router that sees only one direction of a TCP connection, because the
reply takes another path, cannot offload it: conntrack marks the
client's ACK after the SYN invalid, and the flowtable only takes assured
TCP connections. The setup this comes from is two subnets on one L2
segment, with the server answering over the shared segment.
The series adds no TCP state:
- patch 1 hands that ACK to the existing loose pickup in tcp_new(),
gated by nf_conntrack_tcp_loose, as if the SYN had not been seen
- patch 2 keeps the reply direction of such a flow on the classic path
until the connection is assured, as act_ct does for UDP
- patch 3 offloads the connection in the original direction and caps
the timeout handed back to conntrack at UNACK, as
nf_conntrack_tcp_packet() does
- patch 4 adds a selftest, which fails without patches 1 to 3
Pablo, on v1 you asked how TCP tracking copes without the reply
direction, and guessed right that the goal is flowtable offload. The
connection is tracked like a loose mid-stream pickup today: window
checks are off in both directions, and a FIN or RST still takes the
flow off the flowtable.
This series was written with the help of an AI coding assistant. I have
reviewed and tested it.
Changes in v2:
- patch 2: the XDP flowtable lookup does not return the held reply tuple
- patch 3: cap the conntrack timeout at UNACK while no reply was seen
- rebased onto nf-next
v1: https://lore.kernel.org/netfilter-devel/20260910090052.2034970-1-julius@bairaktaris.de/
Gary Dotzler (2):
netfilter: conntrack: pick up a TCP flow whose SYN was never answered
netfilter: nft_flow_offload: offload a TCP flow that has no reply
Julius Bairaktaris (2):
netfilter: flowtable: promote a flow offloaded in one direction only
selftests: netfilter: cover a TCP flow whose reply is never seen
include/net/netfilter/nf_conntrack_l4proto.h | 9 +++
net/netfilter/nf_conntrack_proto_tcp.c | 16 ++++
net/netfilter/nf_flow_table_bpf.c | 4 +
net/netfilter/nf_flow_table_core.c | 5 ++
net/netfilter/nf_flow_table_ip.c | 26 +++++++
net/netfilter/nft_flow_offload.c | 6 +-
.../selftests/net/netfilter/nft_flowtable.sh | 75 +++++++++++++++++++
7 files changed, 139 insertions(+), 2 deletions(-)
base-commit: 87b80c2f6b05cad9f0ff9136709c62a0f59923e3
--
2.53.0
^ permalink raw reply [flat|nested] 5+ messages in thread
* [PATCH nf-next v2 1/4] netfilter: conntrack: pick up a TCP flow whose SYN was never answered
2026-10-04 17:16 [PATCH nf-next v2 0/4] netfilter: offload a TCP flow whose reply is never seen Julius Bairaktaris
@ 2026-10-04 17:16 ` Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 2/4] netfilter: flowtable: promote a flow offloaded in one direction only Julius Bairaktaris
` (2 subsequent siblings)
3 siblings, 0 replies; 5+ messages in thread
From: Julius Bairaktaris @ 2026-10-04 17:16 UTC (permalink / raw)
To: pablo, netfilter-devel
Cc: kadlec, fw, coreteam, netdev, geldot, shuah, linux-kselftest
From: Gary Dotzler <geldot@protonmail.com>
When only one direction of a connection passes the host, conntrack sees
the SYN but not the SYN/ACK. The client's next ACK is invalid in
SYN_SENT, as is every packet after it.
Had the SYN not been seen, the loose pickup in tcp_new() would have
taken that ACK, with window checks off in both directions. Delete the
SYN_SENT entry and look the packet up again so that tcp_new() picks it
up. This applies only with nf_conntrack_tcp_loose, without synproxy,
before any reply, and to an original direction ACK at the expected
sequence number. No TCP state or transition is added.
Signed-off-by: Gary Dotzler <geldot@protonmail.com>
Assisted-by: Claude:claude-opus-5
Signed-off-by: Julius Bairaktaris <julius@bairaktaris.de>
---
net/netfilter/nf_conntrack_proto_tcp.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/net/netfilter/nf_conntrack_proto_tcp.c b/net/netfilter/nf_conntrack_proto_tcp.c
index ad6f1986d52a..31bd4b926e48 100644
--- a/net/netfilter/nf_conntrack_proto_tcp.c
+++ b/net/netfilter/nf_conntrack_proto_tcp.c
@@ -1173,6 +1173,22 @@ int nf_conntrack_tcp_packet(struct nf_conn *ct,
return NF_ACCEPT;
}
+ /* No reply seen and the client continues after its SYN: the
+ * reply takes another path. Recreate the entry through the
+ * loose pickup in tcp_new().
+ */
+ if (tn->tcp_loose && !nfct_synproxy(ct) &&
+ old_state == TCP_CONNTRACK_SYN_SENT &&
+ index == TCP_ACK_SET && dir == IP_CT_DIR_ORIGINAL &&
+ !test_bit(IPS_SEEN_REPLY_BIT, &ct->status) &&
+ ntohl(th->seq) == ct->proto.tcp.seen[dir].td_end) {
+ spin_unlock_bh(&ct->lock);
+
+ if (nf_ct_kill(ct))
+ return -NF_REPEAT;
+ return NF_DROP;
+ }
+
/* Invalid packet */
spin_unlock_bh(&ct->lock);
nf_ct_l4proto_log_invalid(skb, ct, state,
--
2.53.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH nf-next v2 2/4] netfilter: flowtable: promote a flow offloaded in one direction only
2026-10-04 17:16 [PATCH nf-next v2 0/4] netfilter: offload a TCP flow whose reply is never seen Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 1/4] netfilter: conntrack: pick up a TCP flow whose SYN was never answered Julius Bairaktaris
@ 2026-10-04 17:16 ` Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 3/4] netfilter: nft_flow_offload: offload a TCP flow that has no reply Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 4/4] selftests: netfilter: cover a TCP flow whose reply is never seen Julius Bairaktaris
3 siblings, 0 replies; 5+ messages in thread
From: Julius Bairaktaris @ 2026-10-04 17:16 UTC (permalink / raw)
To: pablo, netfilter-devel
Cc: kadlec, fw, coreteam, netdev, geldot, shuah, linux-kselftest
The next patch offloads TCP flows without a reply in the original
direction only. Keep the reply direction of such a flow on the classic
path until the connection is assured, then offload it too, as act_ct
does for UDP. The XDP flowtable lookup does not return that reply tuple
either.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Julius Bairaktaris <julius@bairaktaris.de>
---
net/netfilter/nf_flow_table_bpf.c | 4 ++++
net/netfilter/nf_flow_table_ip.c | 26 ++++++++++++++++++++++++++
2 files changed, 30 insertions(+)
diff --git a/net/netfilter/nf_flow_table_bpf.c b/net/netfilter/nf_flow_table_bpf.c
index cbd5b97a6329..7070bc600db1 100644
--- a/net/netfilter/nf_flow_table_bpf.c
+++ b/net/netfilter/nf_flow_table_bpf.c
@@ -50,6 +50,10 @@ bpf_xdp_flow_tuple_lookup(struct net_device *dev,
nf_flow = container_of(tuplehash, struct flow_offload,
tuplehash[tuplehash->tuple.dir]);
+ if (tuplehash->tuple.dir == FLOW_OFFLOAD_DIR_REPLY &&
+ !test_bit(NF_FLOW_HW_BIDIRECTIONAL, &nf_flow->flags))
+ return ERR_PTR(-ENOENT);
+
flow_offload_refresh(nf_flow_table, nf_flow, false);
return tuplehash;
diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
index c8c29a9a1684..289fe6c9e4f5 100644
--- a/net/netfilter/nf_flow_table_ip.c
+++ b/net/netfilter/nf_flow_table_ip.c
@@ -465,6 +465,26 @@ nf_flow_offload_lookup(struct nf_flowtable_ctx *ctx,
return flow_offload_lookup(flow_table, &tuple);
}
+/* The reply direction of a flow offloaded in one direction only stays on the
+ * classic path so that conntrack sees it. Once the connection is assured, that
+ * direction is offloaded too.
+ */
+static bool nf_flow_reply_unoffloaded(struct nf_flowtable *flow_table,
+ struct flow_offload *flow,
+ enum flow_offload_tuple_dir dir)
+{
+ if (dir != FLOW_OFFLOAD_DIR_REPLY ||
+ test_bit(NF_FLOW_HW_BIDIRECTIONAL, &flow->flags))
+ return false;
+
+ if (test_bit(IPS_ASSURED_BIT, &flow->ct->status)) {
+ set_bit(NF_FLOW_HW_BIDIRECTIONAL, &flow->flags);
+ flow_offload_refresh(flow_table, flow, true);
+ }
+
+ return true;
+}
+
static int nf_flow_offload_forward(struct nf_flowtable_ctx *ctx,
struct nf_flowtable *flow_table,
struct flow_offload_tuple_rhash *tuplehash,
@@ -478,6 +498,9 @@ static int nf_flow_offload_forward(struct nf_flowtable_ctx *ctx,
dir = tuplehash->tuple.dir;
flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
+ if (nf_flow_reply_unoffloaded(flow_table, flow, dir))
+ return 0;
+
mtu = flow->tuplehash[dir].tuple.mtu + ctx->offset;
if (flow->tuplehash[!dir].tuple.tun_num)
mtu -= sizeof(*iph);
@@ -1074,6 +1097,9 @@ static int nf_flow_offload_ipv6_forward(struct nf_flowtable_ctx *ctx,
dir = tuplehash->tuple.dir;
flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
+ if (nf_flow_reply_unoffloaded(flow_table, flow, dir))
+ return 0;
+
mtu = flow->tuplehash[dir].tuple.mtu + ctx->offset;
if (flow->tuplehash[!dir].tuple.tun_num)
mtu -= sizeof(*ip6h);
--
2.53.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH nf-next v2 3/4] netfilter: nft_flow_offload: offload a TCP flow that has no reply
2026-10-04 17:16 [PATCH nf-next v2 0/4] netfilter: offload a TCP flow whose reply is never seen Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 1/4] netfilter: conntrack: pick up a TCP flow whose SYN was never answered Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 2/4] netfilter: flowtable: promote a flow offloaded in one direction only Julius Bairaktaris
@ 2026-10-04 17:16 ` Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 4/4] selftests: netfilter: cover a TCP flow whose reply is never seen Julius Bairaktaris
3 siblings, 0 replies; 5+ messages in thread
From: Julius Bairaktaris @ 2026-10-04 17:16 UTC (permalink / raw)
To: pablo, netfilter-devel
Cc: kadlec, fw, coreteam, netdev, geldot, shuah, linux-kselftest
From: Gary Dotzler <geldot@protonmail.com>
A TCP connection picked up without a reply is not assured, so it is not
offloaded. Offload it in the original direction; if a reply arrives and
the connection becomes assured, the reply direction follows.
When such a flow leaves the flowtable, cap the conntrack timeout at
UNACK, as nf_conntrack_tcp_packet() does without a reply.
Signed-off-by: Gary Dotzler <geldot@protonmail.com>
Assisted-by: Claude:claude-opus-5
Co-developed-by: Julius Bairaktaris <julius@bairaktaris.de>
Signed-off-by: Julius Bairaktaris <julius@bairaktaris.de>
---
With nf_conntrack_tcp_timeout_established=7440, an idle unreplied flow
that aged out of the flowtable had 7390 s left without the cap and
250 s with it (3 runs each, virtme-ng).
include/net/netfilter/nf_conntrack_l4proto.h | 9 +++++++++
net/netfilter/nf_flow_table_core.c | 5 +++++
net/netfilter/nft_flow_offload.c | 6 ++++--
3 files changed, 18 insertions(+), 2 deletions(-)
diff --git a/include/net/netfilter/nf_conntrack_l4proto.h b/include/net/netfilter/nf_conntrack_l4proto.h
index fde2427ceb8f..c251ee862bb6 100644
--- a/include/net/netfilter/nf_conntrack_l4proto.h
+++ b/include/net/netfilter/nf_conntrack_l4proto.h
@@ -208,6 +208,15 @@ static inline bool nf_conntrack_tcp_established(const struct nf_conn *ct)
return ct->proto.tcp.state == TCP_CONNTRACK_ESTABLISHED &&
test_bit(IPS_ASSURED_BIT, &ct->status);
}
+
+/* Picked up mid-stream, no reply seen yet. Caller must check
+ * nf_ct_protonum(ct) is IPPROTO_TCP.
+ */
+static inline bool nf_conntrack_tcp_unreplied(const struct nf_conn *ct)
+{
+ return ct->proto.tcp.state == TCP_CONNTRACK_ESTABLISHED &&
+ !test_bit(IPS_SEEN_REPLY_BIT, &ct->status);
+}
#endif
#ifdef CONFIG_NF_CT_PROTO_SCTP
diff --git a/net/netfilter/nf_flow_table_core.c b/net/netfilter/nf_flow_table_core.c
index 03241d4bfd5e..2f2a6ecc9195 100644
--- a/net/netfilter/nf_flow_table_core.c
+++ b/net/netfilter/nf_flow_table_core.c
@@ -224,6 +224,11 @@ static void flow_offload_fixup_ct(struct flow_offload *flow)
tcp_state = READ_ONCE(ct->proto.tcp.state);
flow_offload_fixup_tcp(ct, tcp_state);
timeout = READ_ONCE(tn->timeouts[tcp_state]);
+ if (nf_conntrack_tcp_unreplied(ct)) {
+ u32 unack = READ_ONCE(tn->timeouts[TCP_CONNTRACK_UNACK]);
+
+ timeout = min_t(s32, timeout, unack);
+ }
expired = nf_flow_has_expired(flow);
}
offload_timeout = READ_ONCE(tn->offload_timeout);
diff --git a/net/netfilter/nft_flow_offload.c b/net/netfilter/nft_flow_offload.c
index 32b4281038dd..b37590c3dac0 100644
--- a/net/netfilter/nft_flow_offload.c
+++ b/net/netfilter/nft_flow_offload.c
@@ -73,7 +73,8 @@ static void nft_flow_offload_eval(const struct nft_expr *expr,
tcph = skb_header_pointer(pkt->skb, nft_thoff(pkt),
sizeof(_tcph), &_tcph);
if (unlikely(!tcph || tcph->fin || tcph->rst ||
- !nf_conntrack_tcp_established(ct)))
+ (!nf_conntrack_tcp_established(ct) &&
+ !nf_conntrack_tcp_unreplied(ct))))
goto out;
break;
case IPPROTO_UDP:
@@ -117,7 +118,8 @@ static void nft_flow_offload_eval(const struct nft_expr *expr,
if (tcph)
flow_offload_ct_tcp(ct);
- __set_bit(NF_FLOW_HW_BIDIRECTIONAL, &flow->flags);
+ if (!tcph || test_bit(IPS_ASSURED_BIT, &ct->status))
+ __set_bit(NF_FLOW_HW_BIDIRECTIONAL, &flow->flags);
ret = flow_offload_add(flowtable, flow);
if (ret < 0)
goto err_flow_add;
--
2.53.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH nf-next v2 4/4] selftests: netfilter: cover a TCP flow whose reply is never seen
2026-10-04 17:16 [PATCH nf-next v2 0/4] netfilter: offload a TCP flow whose reply is never seen Julius Bairaktaris
` (2 preceding siblings ...)
2026-10-04 17:16 ` [PATCH nf-next v2 3/4] netfilter: nft_flow_offload: offload a TCP flow that has no reply Julius Bairaktaris
@ 2026-10-04 17:16 ` Julius Bairaktaris
3 siblings, 0 replies; 5+ messages in thread
From: Julius Bairaktaris @ 2026-10-04 17:16 UTC (permalink / raw)
To: pablo, netfilter-devel
Cc: kadlec, fw, coreteam, netdev, geldot, shuah, linux-kselftest
Add a case where ns2 answers over a direct link, so nsr1 only sees the
original direction. The forward hook must count less than half of the
transferred bytes, which only happens when the flowtable takes over the
connection.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Julius Bairaktaris <julius@bairaktaris.de>
---
.../selftests/net/netfilter/nft_flowtable.sh | 75 +++++++++++++++++++
1 file changed, 75 insertions(+)
diff --git a/tools/testing/selftests/net/netfilter/nft_flowtable.sh b/tools/testing/selftests/net/netfilter/nft_flowtable.sh
index 449c518bd947..2c2669b8fccf 100755
--- a/tools/testing/selftests/net/netfilter/nft_flowtable.sh
+++ b/tools/testing/selftests/net/netfilter/nft_flowtable.sh
@@ -516,6 +516,81 @@ else
ret=1
fi
+# Asymmetric path test:
+# ns2 answers over a direct link, so nsr1 sees the original direction only.
+# Such a connection never becomes assured, but the flowtable is expected to
+# take over the direction that nsr1 does see.
+check_orig_offloaded()
+{
+ local what=$1
+
+ local orig
+ orig=$(ip netns exec "$nsr1" nft reset counter inet filter routed_orig | grep packets)
+ local orig_cnt=${orig#*bytes}
+
+ local fs
+ fs=$(du -sb "$nsin")
+ local max_orig=$(( ${fs%%/*} / 2 ))
+
+ # the flowtable takes over after the first few packets, so the forward
+ # hook must see a small fraction of the transferred file.
+ if [ "$orig_cnt" -gt "$max_orig" ];then
+ echo "FAIL: $what: original counter $orig_cnt exceeds expected value $max_orig" 1>&2
+ ret=1
+ return 1
+ fi
+
+ echo "PASS: $what"
+}
+
+test_asymmetric_path()
+{
+ ip link add name eth1 netns "$ns1" type veth peer name eth1 netns "$ns2"
+ ip -net "$ns1" addr add 10.0.9.99/24 dev eth1
+ ip -net "$ns2" addr add 10.0.9.98/24 dev eth1
+ ip -net "$ns1" addr add dead:9::99/64 dev eth1 nodad
+ ip -net "$ns2" addr add dead:9::98/64 dev eth1 nodad
+ ip -net "$ns1" link set eth1 up
+ ip -net "$ns2" link set eth1 up
+
+ # ns1 keeps sending through nsr1, ns2 answers on the direct link.
+ ip -net "$ns2" route add 10.0.1.99 via 10.0.9.99 dev eth1
+ ip -6 -net "$ns2" route add dead:1::99 via dead:9::99 dev eth1
+
+ # with PMTU discovery the endpoints size their packets for the
+ # router's link, so the fast path forwards them unfragmented
+ ip netns exec "$ns1" sysctl -q net.ipv4.ip_no_pmtu_disc=0
+ ip netns exec "$ns2" sysctl -q net.ipv4.ip_no_pmtu_disc=0
+
+ ip netns exec "$nsr1" nft reset counters table inet filter >/dev/null
+
+ if test_tcp_forwarding "$ns1" "$ns2" 1 4 10.0.2.99 12345; then
+ check_orig_offloaded "flow offloaded for ns1/ns2 without reply"
+ else
+ echo "FAIL: flow offload for ns1/ns2 without reply" 1>&2
+ ip netns exec "$nsr1" nft list ruleset 1>&2
+ ret=1
+ fi
+
+ ip netns exec "$nsr1" nft reset counters table inet filter >/dev/null
+
+ if test_tcp_forwarding "$ns1" "$ns2" 1 6 "[dead:2::99]" 12345; then
+ check_orig_offloaded "IPv6 flow offloaded for ns1/ns2 without reply"
+ else
+ echo "FAIL: IPv6 flow offload for ns1/ns2 without reply" 1>&2
+ ip netns exec "$nsr1" nft list ruleset 1>&2
+ ret=1
+ fi
+
+ ip netns exec "$ns1" sysctl -q net.ipv4.ip_no_pmtu_disc=1
+ ip netns exec "$ns2" sysctl -q net.ipv4.ip_no_pmtu_disc=1
+
+ ip -net "$ns1" link del eth1
+ ip netns exec "$nsr1" nft reset counters table inet filter >/dev/null
+}
+
+test_asymmetric_path
+
# delete default route, i.e. ns2 won't be able to reach ns1 and
# will depend on ns1 being masqueraded in nsr1.
# expect ns1 has nsr1 address.
--
2.53.0
^ permalink raw reply related [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-04 17:16 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-04 17:16 [PATCH nf-next v2 0/4] netfilter: offload a TCP flow whose reply is never seen Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 1/4] netfilter: conntrack: pick up a TCP flow whose SYN was never answered Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 2/4] netfilter: flowtable: promote a flow offloaded in one direction only Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 3/4] netfilter: nft_flow_offload: offload a TCP flow that has no reply Julius Bairaktaris
2026-10-04 17:16 ` [PATCH nf-next v2 4/4] selftests: netfilter: cover a TCP flow whose reply is never seen Julius Bairaktaris
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox