Linux Netfilter development
 help / color / mirror / Atom feed
* [PATCH nf-next v2 0/6] net: netfilter: preliminary support for IPv4 over IPv6 and SIT flowtable offload
@ 2026-09-07  7:33 Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 1/6] net: netfilter: nf_flow_table: recognize IPv4/IPv6 in tunnel proto matching Lorenzo Bianconi
                   ` (5 more replies)
  0 siblings, 6 replies; 9+ messages in thread
From: Lorenzo Bianconi @ 2026-09-07  7:33 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Florian Westphal, Phil Sutter, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Andrew Lunn, David Ahern, Ido Schimmel
  Cc: netfilter-devel, coreteam, netdev, Lorenzo Bianconi,
	Lorenzo Bianconi

Introduce some preliminary changes needed to support flowtable offload
for IPv4 over IPv6 and SIT traffic.

---
Changes in v2:
- Fix dir configuration in nf_flow_offload_check_mtu() to compute tunnel
  MTU overhead.
- Link to v1: https://lore.kernel.org/r/20260901-nf-flowtable-sw-accel-ip6ip-sit-preliminary-v1-0-72e49be8c31f@oss.qualcomm.com

---
Lorenzo Bianconi (6):
      net: netfilter: nf_flow_table: recognize IPv4/IPv6 in tunnel proto matching
      net: netfilter: nf_flow_table: set inner protocol when popping tunnel header
      net: netfilter: nf_flow_table: populate tunnel tuple regardless of inner protocol
      net: netfilter: add encap_proto to flow_offload_tunnel
      net: netfilter: nf_flow_table: refactor MTU check for tunnel offload
      net: netfilter: nf_flow_table: unify tunnel push for IPv4 and IPv6

 include/linux/netdevice.h             |   1 +
 include/net/netfilter/nf_flow_table.h |   1 +
 net/ipv4/ipip.c                       |   1 +
 net/ipv6/ip6_tunnel.c                 |   1 +
 net/netfilter/nf_flow_table_ip.c      | 201 ++++++++++++++++++++++------------
 net/netfilter/nf_flow_table_path.c    |   2 +
 6 files changed, 135 insertions(+), 72 deletions(-)
---
base-commit: 6ebcf5074cff0402730c6981d2397139fee6322d
change-id: 20260901-nf-flowtable-sw-accel-ip6ip-sit-preliminary-5635fae03f3e

Best regards,
-- 
Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>


^ permalink raw reply	[flat|nested] 9+ messages in thread

* [PATCH nf-next v2 1/6] net: netfilter: nf_flow_table: recognize IPv4/IPv6 in tunnel proto matching
  2026-09-07  7:33 [PATCH nf-next v2 0/6] net: netfilter: preliminary support for IPv4 over IPv6 and SIT flowtable offload Lorenzo Bianconi
@ 2026-09-07  7:33 ` Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 2/6] net: netfilter: nf_flow_table: set inner protocol when popping tunnel header Lorenzo Bianconi
                   ` (4 subsequent siblings)
  5 siblings, 0 replies; 9+ messages in thread
From: Lorenzo Bianconi @ 2026-09-07  7:33 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Florian Westphal, Phil Sutter, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Andrew Lunn, David Ahern, Ido Schimmel
  Cc: netfilter-devel, coreteam, netdev, Lorenzo Bianconi

Allow the flowtable to parse IPv4-in-IPv6 and IPv6-in-IPv4 tunnels by
accepting both IPPROTO_IPIP and IPPROTO_IPV6 as the inner protocol when
walking the outer header in nf_flow_ip4_tunnel_proto() and
nf_flow_ip6_tunnel_proto(). This is a preliminary patch to support IPv4
over IPv6 and SIT tunnel flowtable offload.
Please note IPv4 over IPv6 and SIT tunnel flowtable offloading is not
enabled yet.

Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
---
 net/netfilter/nf_flow_table_ip.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
index c8c29a9a1684..42e8de696474 100644
--- a/net/netfilter/nf_flow_table_ip.c
+++ b/net/netfilter/nf_flow_table_ip.c
@@ -326,7 +326,7 @@ static bool nf_flow_ip4_tunnel_proto(struct nf_flowtable_ctx *ctx,
 	if (iph->ttl <= 1)
 		return false;
 
-	if (iph->protocol == IPPROTO_IPIP) {
+	if (iph->protocol == IPPROTO_IPIP || iph->protocol == IPPROTO_IPV6) {
 		ctx->tun.inner_proto = iph->protocol;
 		ctx->tun.hdr_size = size;
 		ctx->offset += ctx->tun.hdr_size;
@@ -351,7 +351,7 @@ static bool nf_flow_ip6_tunnel_proto(struct nf_flowtable_ctx *ctx,
 	if (ipv6_ext_hdr(ip6h->nexthdr))
 		return false;
 
-	if (ip6h->nexthdr == IPPROTO_IPV6) {
+	if (ip6h->nexthdr == IPPROTO_IPIP || ip6h->nexthdr == IPPROTO_IPV6) {
 		ctx->tun.inner_proto = ip6h->nexthdr;
 		ctx->tun.hdr_size = sizeof(*ip6h);
 		ctx->offset += ctx->tun.hdr_size;

-- 
2.55.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH nf-next v2 2/6] net: netfilter: nf_flow_table: set inner protocol when popping tunnel header
  2026-09-07  7:33 [PATCH nf-next v2 0/6] net: netfilter: preliminary support for IPv4 over IPv6 and SIT flowtable offload Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 1/6] net: netfilter: nf_flow_table: recognize IPv4/IPv6 in tunnel proto matching Lorenzo Bianconi
@ 2026-09-07  7:33 ` Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 3/6] net: netfilter: nf_flow_table: populate tunnel tuple regardless of inner protocol Lorenzo Bianconi
                   ` (3 subsequent siblings)
  5 siblings, 0 replies; 9+ messages in thread
From: Lorenzo Bianconi @ 2026-09-07  7:33 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Florian Westphal, Phil Sutter, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Andrew Lunn, David Ahern, Ido Schimmel
  Cc: netfilter-devel, coreteam, netdev, Lorenzo Bianconi

Pop the outer tunnel header and set skb->protocol to the inner protocol
derived from ctx->tun.inner_proto in nf_flow_ip_tunnel_pop(), instead of
bailing out unless the tunnel carries IP-in-IP or IPv6-in-IPv6.
This is a preliminary patch to support IPv4 over IPv6 and SIT tunnel
flowtable offload.

Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
---
 net/netfilter/nf_flow_table_ip.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
index 42e8de696474..cd636b4cf74b 100644
--- a/net/netfilter/nf_flow_table_ip.c
+++ b/net/netfilter/nf_flow_table_ip.c
@@ -372,6 +372,11 @@ static void nf_flow_ip_tunnel_pop(struct nf_flowtable_ctx *ctx,
 
 	skb_pull(skb, ctx->tun.hdr_size);
 	skb_reset_network_header(skb);
+
+	if (ctx->tun.inner_proto == IPPROTO_IPIP)
+		skb->protocol = htons(ETH_P_IP);
+	else
+		skb->protocol = htons(ETH_P_IPV6);
 }
 
 static bool nf_flow_skb_encap_protocol(struct nf_flowtable_ctx *ctx,

-- 
2.55.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH nf-next v2 3/6] net: netfilter: nf_flow_table: populate tunnel tuple regardless of inner protocol
  2026-09-07  7:33 [PATCH nf-next v2 0/6] net: netfilter: preliminary support for IPv4 over IPv6 and SIT flowtable offload Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 1/6] net: netfilter: nf_flow_table: recognize IPv4/IPv6 in tunnel proto matching Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 2/6] net: netfilter: nf_flow_table: set inner protocol when popping tunnel header Lorenzo Bianconi
@ 2026-09-07  7:33 ` Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 4/6] net: netfilter: add encap_proto to flow_offload_tunnel Lorenzo Bianconi
                   ` (2 subsequent siblings)
  5 siblings, 0 replies; 9+ messages in thread
From: Lorenzo Bianconi @ 2026-09-07  7:33 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Florian Westphal, Phil Sutter, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Andrew Lunn, David Ahern, Ido Schimmel
  Cc: netfilter-devel, coreteam, netdev, Lorenzo Bianconi

Store the outer tunnel addresses and inner protocol in the flow tuple
unconditionally in nf_flow_tuple_encap(), instead of only when the
tunnel is IP-in-IP or IPv6-in-IPv6. Bail out early when the inner
protocol is neither IPPROTO_IPIP nor IPPROTO_IPV6.
This makes the tuple usable for cross-family tunnels such as
IPv4-over-IPv6 and SIT, where the inner protocol differs from the
outer address family.
This is a preliminary patch to support IPv4 over IPv6 and SIT tunnel
flowtable offload.
Please note IPv4 over IPv6 and SIT tunnel flowtable offloading is not
enabled yet.

Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
---
 net/netfilter/nf_flow_table_ip.c | 20 ++++++++++----------
 1 file changed, 10 insertions(+), 10 deletions(-)

diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
index cd636b4cf74b..a53393f8eda0 100644
--- a/net/netfilter/nf_flow_table_ip.c
+++ b/net/netfilter/nf_flow_table_ip.c
@@ -189,22 +189,22 @@ static void nf_flow_tuple_encap(struct nf_flowtable_ctx *ctx,
 		break;
 	}
 
+	if (ctx->tun.inner_proto != IPPROTO_IPIP &&
+	    ctx->tun.inner_proto != IPPROTO_IPV6)
+		return;
+
 	switch (ctx->ether_type) {
 	case htons(ETH_P_IP):
 		iph = (struct iphdr *)(skb_network_header(skb) + offset);
-		if (ctx->tun.inner_proto == IPPROTO_IPIP) {
-			tuple->tun.dst_v4.s_addr = iph->daddr;
-			tuple->tun.src_v4.s_addr = iph->saddr;
-			tuple->tun.inner_proto = IPPROTO_IPIP;
-		}
+		tuple->tun.dst_v4.s_addr = iph->daddr;
+		tuple->tun.src_v4.s_addr = iph->saddr;
+		tuple->tun.inner_proto = ctx->tun.inner_proto;
 		break;
 	case htons(ETH_P_IPV6):
 		ip6h = (struct ipv6hdr *)(skb_network_header(skb) + offset);
-		if (ctx->tun.inner_proto == IPPROTO_IPV6) {
-			tuple->tun.dst_v6 = ip6h->daddr;
-			tuple->tun.src_v6 = ip6h->saddr;
-			tuple->tun.inner_proto = IPPROTO_IPV6;
-		}
+		tuple->tun.dst_v6 = ip6h->daddr;
+		tuple->tun.src_v6 = ip6h->saddr;
+		tuple->tun.inner_proto = ctx->tun.inner_proto;
 		break;
 	default:
 		break;

-- 
2.55.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH nf-next v2 4/6] net: netfilter: add encap_proto to flow_offload_tunnel
  2026-09-07  7:33 [PATCH nf-next v2 0/6] net: netfilter: preliminary support for IPv4 over IPv6 and SIT flowtable offload Lorenzo Bianconi
                   ` (2 preceding siblings ...)
  2026-09-07  7:33 ` [PATCH nf-next v2 3/6] net: netfilter: nf_flow_table: populate tunnel tuple regardless of inner protocol Lorenzo Bianconi
@ 2026-09-07  7:33 ` Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 5/6] net: netfilter: nf_flow_table: refactor MTU check for tunnel offload Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 6/6] net: netfilter: nf_flow_table: unify tunnel push for IPv4 and IPv6 Lorenzo Bianconi
  5 siblings, 0 replies; 9+ messages in thread
From: Lorenzo Bianconi @ 2026-09-07  7:33 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Florian Westphal, Phil Sutter, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Andrew Lunn, David Ahern, Ido Schimmel
  Cc: netfilter-devel, coreteam, netdev, Lorenzo Bianconi,
	Lorenzo Bianconi

From: Lorenzo Bianconi <lorenzo@kernel.org>

Add encap_proto (AF_INET or AF_INET6) to struct flow_offload_tunnel
to allow its use as part of the hash table key during flowtable entry
lookup.
This is a preliminary change to support IPv4 over IPv6 tunneling via
the flowtable infrastructure for software acceleration.

Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
---
 include/linux/netdevice.h             | 1 +
 include/net/netfilter/nf_flow_table.h | 1 +
 net/ipv4/ipip.c                       | 1 +
 net/ipv6/ip6_tunnel.c                 | 1 +
 net/netfilter/nf_flow_table_ip.c      | 2 ++
 net/netfilter/nf_flow_table_path.c    | 2 ++
 6 files changed, 8 insertions(+)

diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index 8454646d6a45..4531e4de8416 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -910,6 +910,7 @@ struct net_device_path {
 			};
 
 			u8	inner_proto;
+			u8	encap_proto;
 		} tun;
 		struct {
 			enum {
diff --git a/include/net/netfilter/nf_flow_table.h b/include/net/netfilter/nf_flow_table.h
index f2e2771f188f..306e7d9ec3ec 100644
--- a/include/net/netfilter/nf_flow_table.h
+++ b/include/net/netfilter/nf_flow_table.h
@@ -118,6 +118,7 @@ struct flow_offload_tunnel {
 	};
 
 	u8	inner_proto;
+	u8	encap_proto;
 };
 
 struct flow_offload_tuple {
diff --git a/net/ipv4/ipip.c b/net/ipv4/ipip.c
index f684baf8e58f..c06ea32ec1df 100644
--- a/net/ipv4/ipip.c
+++ b/net/ipv4/ipip.c
@@ -379,6 +379,7 @@ static int ipip_fill_forward_path(struct net_device_path_ctx *ctx,
 	path->tun.src_v4.s_addr = tiph->saddr;
 	path->tun.dst_v4.s_addr = tiph->daddr;
 	path->tun.inner_proto = IPPROTO_IPIP;
+	path->tun.encap_proto = AF_INET;
 	path->tun.dst = &rt->dst;
 	path->dev = ctx->dev;
 
diff --git a/net/ipv6/ip6_tunnel.c b/net/ipv6/ip6_tunnel.c
index d5ff50a2ac01..64a184bca467 100644
--- a/net/ipv6/ip6_tunnel.c
+++ b/net/ipv6/ip6_tunnel.c
@@ -1868,6 +1868,7 @@ static int ip6_tnl_fill_forward_path(struct net_device_path_ctx *ctx,
 		path->tun.src_v6 = fl6.saddr;
 		path->tun.dst_v6 = fl6.daddr;
 		path->tun.inner_proto = IPPROTO_IPV6;
+		path->tun.encap_proto = AF_INET6;
 		path->tun.dst = dst;
 		path->dev = ctx->dev;
 		ctx->dev = dst->dev;
diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
index a53393f8eda0..1b9360d54dfc 100644
--- a/net/netfilter/nf_flow_table_ip.c
+++ b/net/netfilter/nf_flow_table_ip.c
@@ -199,12 +199,14 @@ static void nf_flow_tuple_encap(struct nf_flowtable_ctx *ctx,
 		tuple->tun.dst_v4.s_addr = iph->daddr;
 		tuple->tun.src_v4.s_addr = iph->saddr;
 		tuple->tun.inner_proto = ctx->tun.inner_proto;
+		tuple->tun.encap_proto = AF_INET;
 		break;
 	case htons(ETH_P_IPV6):
 		ip6h = (struct ipv6hdr *)(skb_network_header(skb) + offset);
 		tuple->tun.dst_v6 = ip6h->daddr;
 		tuple->tun.src_v6 = ip6h->saddr;
 		tuple->tun.inner_proto = ctx->tun.inner_proto;
+		tuple->tun.encap_proto = AF_INET6;
 		break;
 	default:
 		break;
diff --git a/net/netfilter/nf_flow_table_path.c b/net/netfilter/nf_flow_table_path.c
index 1e55644f2edb..90360406d11e 100644
--- a/net/netfilter/nf_flow_table_path.c
+++ b/net/netfilter/nf_flow_table_path.c
@@ -135,6 +135,7 @@ static int nft_dev_path_info(struct net_device_path_stack *stack,
 				info->tun.dst_v6 = path->tun.dst_v6;
 				info->tun.inner_proto = path->tun.inner_proto;
 				info->tun_dst = path->tun.dst;
+				info->tun.encap_proto = path->tun.encap_proto;
 				info->num_tuns++;
 			} else {
 				if (info->num_encaps >= NF_FLOW_TABLE_ENCAP_MAX)
@@ -246,6 +247,7 @@ static int nft_dev_forward_path(const struct nft_pktinfo *pkt,
 		route->tuple[!dir].in.tun.src_v6 = info.tun.dst_v6;
 		route->tuple[!dir].in.tun.dst_v6 = info.tun.src_v6;
 		route->tuple[!dir].in.tun.inner_proto = info.tun.inner_proto;
+		route->tuple[!dir].in.tun.encap_proto = info.tun.encap_proto;
 		route->tuple[!dir].in.num_tuns = info.num_tuns;
 		dst_release(route->tuple[dir].dst);
 		route->tuple[dir].dst = info.tun_dst;

-- 
2.55.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH nf-next v2 5/6] net: netfilter: nf_flow_table: refactor MTU check for tunnel offload
  2026-09-07  7:33 [PATCH nf-next v2 0/6] net: netfilter: preliminary support for IPv4 over IPv6 and SIT flowtable offload Lorenzo Bianconi
                   ` (3 preceding siblings ...)
  2026-09-07  7:33 ` [PATCH nf-next v2 4/6] net: netfilter: add encap_proto to flow_offload_tunnel Lorenzo Bianconi
@ 2026-09-07  7:33 ` Lorenzo Bianconi
  2026-09-07  7:33 ` [PATCH nf-next v2 6/6] net: netfilter: nf_flow_table: unify tunnel push for IPv4 and IPv6 Lorenzo Bianconi
  5 siblings, 0 replies; 9+ messages in thread
From: Lorenzo Bianconi @ 2026-09-07  7:33 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Florian Westphal, Phil Sutter, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Andrew Lunn, David Ahern, Ido Schimmel
  Cc: netfilter-devel, coreteam, netdev, Lorenzo Bianconi

Introduce nf_flow_offload_check_mtu() helper and use the encapsulated
protocol (tun.encap_proto) to compute the tunnel overhead instead of
relying on tun_num, so the correct inner header size (IPv4 vs IPv6) is
accounted for. Use it in both the IPv4 and IPv6 forward paths.
This is a preliminary patch to support IPv4 over IPv6 and SIT flowtable
tunnel offload.
Please note IPv4 over IPv6 and SIT tunnel flowtable offloading is not
enabled yet.

Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
---
 net/netfilter/nf_flow_table_ip.c | 41 ++++++++++++++++++++++++++++------------
 1 file changed, 29 insertions(+), 12 deletions(-)

diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
index 1b9360d54dfc..96dbdadba4e7 100644
--- a/net/netfilter/nf_flow_table_ip.c
+++ b/net/netfilter/nf_flow_table_ip.c
@@ -472,6 +472,31 @@ nf_flow_offload_lookup(struct nf_flowtable_ctx *ctx,
 	return flow_offload_lookup(flow_table, &tuple);
 }
 
+static int nf_flow_offload_check_mtu(struct nf_flowtable_ctx *ctx,
+				     struct flow_offload *flow,
+				     enum flow_offload_tuple_dir dir,
+				     struct sk_buff *skb)
+{
+	unsigned int mtu;
+
+	mtu = flow->tuplehash[dir].tuple.mtu + ctx->offset;
+	switch (flow->tuplehash[!dir].tuple.tun.encap_proto) {
+	case AF_INET:
+		mtu -= sizeof(struct iphdr);
+		break;
+	case AF_INET6:
+		mtu -= sizeof(struct ipv6hdr);
+		break;
+	default:
+		break;
+	}
+
+	if (unlikely(nf_flow_exceeds_mtu(skb, mtu)))
+		return -EINVAL;
+
+	return 0;
+}
+
 static int nf_flow_offload_forward(struct nf_flowtable_ctx *ctx,
 				   struct nf_flowtable *flow_table,
 				   struct flow_offload_tuple_rhash *tuplehash,
@@ -479,17 +504,13 @@ static int nf_flow_offload_forward(struct nf_flowtable_ctx *ctx,
 {
 	enum flow_offload_tuple_dir dir;
 	struct flow_offload *flow;
-	unsigned int thoff, mtu;
+	unsigned int thoff;
 	struct iphdr *iph;
 
 	dir = tuplehash->tuple.dir;
 	flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
 
-	mtu = flow->tuplehash[dir].tuple.mtu + ctx->offset;
-	if (flow->tuplehash[!dir].tuple.tun_num)
-		mtu -= sizeof(*iph);
-
-	if (unlikely(nf_flow_exceeds_mtu(skb, mtu)))
+	if (nf_flow_offload_check_mtu(ctx, flow, dir, skb))
 		return 0;
 
 	iph = (struct iphdr *)(skb_network_header(skb) + ctx->offset);
@@ -1075,17 +1096,13 @@ static int nf_flow_offload_ipv6_forward(struct nf_flowtable_ctx *ctx,
 {
 	enum flow_offload_tuple_dir dir;
 	struct flow_offload *flow;
-	unsigned int thoff, mtu;
 	struct ipv6hdr *ip6h;
+	unsigned int thoff;
 
 	dir = tuplehash->tuple.dir;
 	flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
 
-	mtu = flow->tuplehash[dir].tuple.mtu + ctx->offset;
-	if (flow->tuplehash[!dir].tuple.tun_num)
-		mtu -= sizeof(*ip6h);
-
-	if (unlikely(nf_flow_exceeds_mtu(skb, mtu)))
+	if (nf_flow_offload_check_mtu(ctx, flow, dir, skb))
 		return 0;
 
 	ip6h = (struct ipv6hdr *)(skb_network_header(skb) + ctx->offset);

-- 
2.55.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH nf-next v2 6/6] net: netfilter: nf_flow_table: unify tunnel push for IPv4 and IPv6
  2026-09-07  7:33 [PATCH nf-next v2 0/6] net: netfilter: preliminary support for IPv4 over IPv6 and SIT flowtable offload Lorenzo Bianconi
                   ` (4 preceding siblings ...)
  2026-09-07  7:33 ` [PATCH nf-next v2 5/6] net: netfilter: nf_flow_table: refactor MTU check for tunnel offload Lorenzo Bianconi
@ 2026-09-07  7:33 ` Lorenzo Bianconi
  2026-09-16 22:14   ` Pablo Neira Ayuso
  5 siblings, 1 reply; 9+ messages in thread
From: Lorenzo Bianconi @ 2026-09-07  7:33 UTC (permalink / raw)
  To: Pablo Neira Ayuso, Florian Westphal, Phil Sutter, David S. Miller,
	Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
	Andrew Lunn, David Ahern, Ido Schimmel
  Cc: netfilter-devel, coreteam, netdev, Lorenzo Bianconi

Refactor nf_flow_tunnel_ipip_push() and nf_flow_tunnel_ip6ip6_push()
into nf_flow_tunnel_ip_push() and nf_flow_tunnel_ip6_push(), keying the
inner header handling off tuple->tun.inner_proto so both IP-in-IP and
IPv6-in-IPv6 inner protocols are supported regardless of the outer
address family. Replace nf_flow_tunnel_v4_push() and
nf_flow_tunnel_v6_push() with a single nf_flow_tunnel_push() that
dispatches on tuple->tun.encap_proto, and set skb->protocol explicitly
after pushing the outer header.
This is a preliminary patch to support IPv4 over IPv6 and SIT tunnel
flowtable offload.
Please note IPv4 over IPv6 and SIT tunnel flowtable offloading is not
enabled yet.

Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
---
 net/netfilter/nf_flow_table_ip.c | 129 ++++++++++++++++++++++++---------------
 1 file changed, 81 insertions(+), 48 deletions(-)

diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
index 96dbdadba4e7..25875cc97bda 100644
--- a/net/netfilter/nf_flow_table_ip.c
+++ b/net/netfilter/nf_flow_table_ip.c
@@ -608,23 +608,41 @@ static int nf_flow_pppoe_push(struct sk_buff *skb, u16 id,
 	return 0;
 }
 
-static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
-				    struct flow_offload_tuple *tuple,
-				    struct dst_entry *dst, __be32 *ip_daddr)
+static int nf_flow_tunnel_ip_push(struct net *net, struct sk_buff *skb,
+				  struct flow_offload_tuple *tuple,
+				  struct dst_entry *dst, __be32 *ip_daddr)
 {
-	struct iphdr *iph = (struct iphdr *)skb_network_header(skb);
-	struct rtable *rt = dst_rtable(dst);
-	u8 tos = iph->tos, ttl = iph->ttl;
-	__be16 frag_off = iph->frag_off;
-	u32 headroom = sizeof(*iph);
+	__be16 frag_off = 0;
+	struct iphdr *iph;
+	u8 tos = 0, ttl;
+	u32 headroom;
 	int err;
 
+	switch (tuple->tun.inner_proto) {
+	case IPPROTO_IPV6: {
+		struct ipv6hdr *ip6h;
+
+		ip6h = (struct ipv6hdr *)skb_network_header(skb);
+		tos = ipv6_get_dsfield(ip6h);
+		ttl = ip6h->hop_limit;
+		frag_off = htons(IP_DF);
+		break;
+	}
+	default:
+		iph = (struct iphdr *)skb_network_header(skb);
+		frag_off = iph->frag_off;
+		tos = iph->tos;
+		ttl = iph->ttl;
+		break;
+	}
+
 	err = iptunnel_handle_offloads(skb, SKB_GSO_IPXIP4);
 	if (err)
 		return err;
 
-	skb_set_inner_ipproto(skb, IPPROTO_IPIP);
-	headroom += LL_RESERVED_SPACE(rt->dst.dev) + rt->dst.header_len;
+	skb_set_inner_ipproto(skb, tuple->tun.inner_proto);
+	headroom = sizeof(*iph) + LL_RESERVED_SPACE(dst->dev) +
+		   dst->header_len;
 	err = skb_cow_head(skb, headroom);
 	if (err)
 		return err;
@@ -635,11 +653,12 @@ static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
 	/* Push down and install the IP header. */
 	skb_push(skb, sizeof(*iph));
 	skb_reset_network_header(skb);
+	skb->protocol = htons(ETH_P_IP);
 
 	iph = ip_hdr(skb);
 	iph->version	= 4;
 	iph->ihl	= sizeof(*iph) >> 2;
-	iph->frag_off	= ip_mtu_locked(&rt->dst) ? 0 : frag_off;
+	iph->frag_off	= ip_mtu_locked(dst) ? 0 : frag_off;
 	iph->protocol	= tuple->tun.inner_proto;
 	iph->tos	= tos;
 	iph->daddr	= tuple->tun.src_v4.s_addr;
@@ -654,57 +673,61 @@ static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
 	return 0;
 }
 
-static int nf_flow_tunnel_v4_push(struct net *net, struct sk_buff *skb,
-				  struct flow_offload_tuple *tuple,
-				  struct dst_entry *dst,  __be32 *ip_daddr)
+static int nf_flow_tunnel_ip6_push(struct net *net, struct sk_buff *skb,
+				   struct flow_offload_tuple *tuple,
+				   struct dst_entry *dst,
+				   struct in6_addr **ip6_daddr)
 {
-	if (tuple->tun_num)
-		return nf_flow_tunnel_ipip_push(net, skb, tuple, dst, ip_daddr);
-
-	return 0;
-}
-
-static int nf_flow_tunnel_ip6ip6_push(struct net *net, struct sk_buff *skb,
-				      struct flow_offload_tuple *tuple,
-				      struct dst_entry *dst,
-				      struct in6_addr **ip6_daddr)
-{
-	struct ipv6hdr *ip6h = (struct ipv6hdr *)skb_network_header(skb);
-	__u8 dsfield = ipv6_get_dsfield(ip6h);
-	struct rtable *rt = dst_rtable(dst);
 	struct flowi6 fl6 = {
 		.daddr = tuple->tun.src_v6,
 		.saddr = tuple->tun.dst_v6,
-		.flowi6_proto = IPPROTO_IPV6,
+		.flowi6_proto = tuple->tun.inner_proto,
 	};
-	u8 hop_limit = ip6h->hop_limit;
+	u8 hop_limit, dsfield;
+	struct ipv6hdr *ip6h;
 	int err, mtu;
 	u32 headroom;
 
+	switch (tuple->tun.inner_proto) {
+	case IPPROTO_IPIP: {
+		struct iphdr *iph = (struct iphdr *)skb_network_header(skb);
+
+		dsfield = ipv4_get_dsfield(iph);
+		hop_limit = iph->ttl;
+		break;
+	}
+	default:
+		ip6h = (struct ipv6hdr *)skb_network_header(skb);
+		dsfield = ipv6_get_dsfield(ip6h);
+		hop_limit = ip6h->hop_limit;
+		break;
+	}
+
 	err = iptunnel_handle_offloads(skb, SKB_GSO_IPXIP6);
 	if (err)
 		return err;
 
-	skb_set_inner_ipproto(skb, IPPROTO_IPV6);
-	headroom = sizeof(*ip6h) + LL_RESERVED_SPACE(rt->dst.dev) +
-		   rt->dst.header_len;
+	skb_set_inner_ipproto(skb, tuple->tun.inner_proto);
+	headroom = sizeof(*ip6h) + LL_RESERVED_SPACE(dst->dev) +
+		   dst->header_len;
 	err = skb_cow_head(skb, headroom);
 	if (err)
 		return err;
 
 	skb_scrub_packet(skb, true);
-	mtu = dst_mtu(&rt->dst) - sizeof(*ip6h);
+	mtu = dst_mtu(dst) - sizeof(*ip6h);
 	mtu = max(mtu, IPV6_MIN_MTU);
 	skb_dst_update_pmtu_no_confirm(skb, mtu);
 
 	skb_push(skb, sizeof(*ip6h));
 	skb_reset_network_header(skb);
+	skb->protocol = htons(ETH_P_IPV6);
 
 	ip6h = ipv6_hdr(skb);
 	ip6_flow_hdr(ip6h, dsfield,
 		     ip6_make_flowlabel(net, skb, fl6.flowlabel, true, &fl6));
 	ip6h->hop_limit = hop_limit;
-	ip6h->nexthdr = IPPROTO_IPV6;
+	ip6h->nexthdr = tuple->tun.inner_proto;
 	ip6h->daddr = tuple->tun.src_v6;
 	ip6h->saddr = tuple->tun.dst_v6;
 	ipv6_hdr(skb)->payload_len = htons(skb->len - sizeof(*ip6h));
@@ -715,15 +738,20 @@ static int nf_flow_tunnel_ip6ip6_push(struct net *net, struct sk_buff *skb,
 	return 0;
 }
 
-static int nf_flow_tunnel_v6_push(struct net *net, struct sk_buff *skb,
-				  struct flow_offload_tuple *tuple,
-				  struct dst_entry *dst,
-				  struct in6_addr **ip6_daddr)
+static int nf_flow_tunnel_push(struct net *net, struct sk_buff *skb,
+			       struct flow_offload_tuple *tuple,
+			       struct dst_entry *dst, __be32 *ip_daddr,
+			       struct in6_addr **ip6_daddr)
 {
-	if (tuple->tun_num)
-		return nf_flow_tunnel_ip6ip6_push(net, skb, tuple, dst, ip6_daddr);
-
-	return 0;
+	switch (tuple->tun.encap_proto) {
+	case AF_INET:
+		return nf_flow_tunnel_ip_push(net, skb, tuple, dst, ip_daddr);
+	case AF_INET6:
+		return nf_flow_tunnel_ip6_push(net, skb, tuple, dst,
+					       ip6_daddr);
+	default:
+		return 0;
+	}
 }
 
 static int nf_flow_encap_push(struct sk_buff *skb,
@@ -830,6 +858,7 @@ static int nf_flow_queue_xmit4(struct sk_buff *skb,
 	struct flow_offload_tuple *other_tuple;
 	enum flow_offload_tuple_dir dir;
 	struct nf_flow_xmit xmit = {};
+	struct in6_addr *ip6_daddr;
 	struct flow_offload *flow;
 	struct neighbour *neigh;
 	struct rtable *rt;
@@ -847,9 +876,11 @@ static int nf_flow_queue_xmit4(struct sk_buff *skb,
 	flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
 	other_tuple = &flow->tuplehash[!dir].tuple;
 	ip_daddr = other_tuple->src_v4.s_addr;
+	ip6_daddr = &other_tuple->src_v6;
 
-	if (nf_flow_tunnel_v4_push(state->net, skb, other_tuple,
-				   tuplehash->tuple.dst_cache, &ip_daddr) < 0)
+	if (nf_flow_tunnel_push(state->net, skb, other_tuple,
+				tuplehash->tuple.dst_cache,
+				&ip_daddr, &ip6_daddr) < 0)
 		return NF_DROP;
 
 	switch (tuplehash->tuple.xmit_type) {
@@ -1158,6 +1189,7 @@ static int nf_flow_queue_xmit6(struct sk_buff *skb,
 	struct flow_offload *flow;
 	struct neighbour *neigh;
 	struct rt6_info *rt;
+	__be32 ip_daddr;
 
 	if (unlikely(tuplehash->tuple.xmit_type == FLOW_OFFLOAD_XMIT_XFRM)) {
 		rt = dst_rt6_info(tuplehash->tuple.dst_cache);
@@ -1170,11 +1202,12 @@ static int nf_flow_queue_xmit6(struct sk_buff *skb,
 	dir = tuplehash->tuple.dir;
 	flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
 	other_tuple = &flow->tuplehash[!dir].tuple;
+	ip_daddr = other_tuple->src_v4.s_addr;
 	ip6_daddr = &other_tuple->src_v6;
 
-	if (nf_flow_tunnel_v6_push(state->net, skb, other_tuple,
-				   tuplehash->tuple.dst_cache,
-				   &ip6_daddr) < 0)
+	if (nf_flow_tunnel_push(state->net, skb, other_tuple,
+				tuplehash->tuple.dst_cache,
+				&ip_daddr, &ip6_daddr) < 0)
 		return NF_DROP;
 
 	switch (tuplehash->tuple.xmit_type) {

-- 
2.55.0


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* Re: [PATCH nf-next v2 6/6] net: netfilter: nf_flow_table: unify tunnel push for IPv4 and IPv6
  2026-09-07  7:33 ` [PATCH nf-next v2 6/6] net: netfilter: nf_flow_table: unify tunnel push for IPv4 and IPv6 Lorenzo Bianconi
@ 2026-09-16 22:14   ` Pablo Neira Ayuso
  2026-09-17  8:26     ` Lorenzo Bianconi
  0 siblings, 1 reply; 9+ messages in thread
From: Pablo Neira Ayuso @ 2026-09-16 22:14 UTC (permalink / raw)
  To: Lorenzo Bianconi
  Cc: Florian Westphal, Phil Sutter, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, Andrew Lunn,
	David Ahern, Ido Schimmel, netfilter-devel, coreteam, netdev

On Mon, Sep 07, 2026 at 09:33:41AM +0200, Lorenzo Bianconi wrote:
> Refactor nf_flow_tunnel_ipip_push() and nf_flow_tunnel_ip6ip6_push()
> into nf_flow_tunnel_ip_push() and nf_flow_tunnel_ip6_push(), keying the
> inner header handling off tuple->tun.inner_proto so both IP-in-IP and
> IPv6-in-IPv6 inner protocols are supported regardless of the outer
> address family. Replace nf_flow_tunnel_v4_push() and
> nf_flow_tunnel_v6_push() with a single nf_flow_tunnel_push() that
> dispatches on tuple->tun.encap_proto, and set skb->protocol explicitly
> after pushing the outer header.
> This is a preliminary patch to support IPv4 over IPv6 and SIT tunnel
> flowtable offload.
> Please note IPv4 over IPv6 and SIT tunnel flowtable offloading is not
> enabled yet.
> 
> Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
> ---
>  net/netfilter/nf_flow_table_ip.c | 129 ++++++++++++++++++++++++---------------
>  1 file changed, 81 insertions(+), 48 deletions(-)
> 
> diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
> index 96dbdadba4e7..25875cc97bda 100644
> --- a/net/netfilter/nf_flow_table_ip.c
> +++ b/net/netfilter/nf_flow_table_ip.c
> @@ -608,23 +608,41 @@ static int nf_flow_pppoe_push(struct sk_buff *skb, u16 id,
>  	return 0;
>  }
>  
> -static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
> -				    struct flow_offload_tuple *tuple,
> -				    struct dst_entry *dst, __be32 *ip_daddr)
> +static int nf_flow_tunnel_ip_push(struct net *net, struct sk_buff *skb,
> +				  struct flow_offload_tuple *tuple,
> +				  struct dst_entry *dst, __be32 *ip_daddr)
>  {
> -	struct iphdr *iph = (struct iphdr *)skb_network_header(skb);
> -	struct rtable *rt = dst_rtable(dst);
> -	u8 tos = iph->tos, ttl = iph->ttl;
> -	__be16 frag_off = iph->frag_off;
> -	u32 headroom = sizeof(*iph);
> +	__be16 frag_off = 0;
> +	struct iphdr *iph;
> +	u8 tos = 0, ttl;
> +	u32 headroom;
>  	int err;
>  
> +	switch (tuple->tun.inner_proto) {
> +	case IPPROTO_IPV6: {
> +		struct ipv6hdr *ip6h;
> +
> +		ip6h = (struct ipv6hdr *)skb_network_header(skb);
> +		tos = ipv6_get_dsfield(ip6h);
> +		ttl = ip6h->hop_limit;
> +		frag_off = htons(IP_DF);
> +		break;
> +	}
> +	default:

Please add an explicit case to check for IPv4 here, ie. no default:

> +		iph = (struct iphdr *)skb_network_header(skb);
> +		frag_off = iph->frag_off;
> +		tos = iph->tos;
> +		ttl = iph->ttl;
> +		break;
> +	}
> +
>  	err = iptunnel_handle_offloads(skb, SKB_GSO_IPXIP4);
>  	if (err)
>  		return err;
>  
> -	skb_set_inner_ipproto(skb, IPPROTO_IPIP);
> -	headroom += LL_RESERVED_SPACE(rt->dst.dev) + rt->dst.header_len;
> +	skb_set_inner_ipproto(skb, tuple->tun.inner_proto);
> +	headroom = sizeof(*iph) + LL_RESERVED_SPACE(dst->dev) +
> +		   dst->header_len;
>  	err = skb_cow_head(skb, headroom);
>  	if (err)
>  		return err;
> @@ -635,11 +653,12 @@ static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
>  	/* Push down and install the IP header. */
>  	skb_push(skb, sizeof(*iph));
>  	skb_reset_network_header(skb);
> +	skb->protocol = htons(ETH_P_IP);
>  
>  	iph = ip_hdr(skb);
>  	iph->version	= 4;
>  	iph->ihl	= sizeof(*iph) >> 2;
> -	iph->frag_off	= ip_mtu_locked(&rt->dst) ? 0 : frag_off;
> +	iph->frag_off	= ip_mtu_locked(dst) ? 0 : frag_off;
>  	iph->protocol	= tuple->tun.inner_proto;
>  	iph->tos	= tos;
>  	iph->daddr	= tuple->tun.src_v4.s_addr;
> @@ -654,57 +673,61 @@ static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
>  	return 0;
>  }
>  
> -static int nf_flow_tunnel_v4_push(struct net *net, struct sk_buff *skb,
> -				  struct flow_offload_tuple *tuple,
> -				  struct dst_entry *dst,  __be32 *ip_daddr)
> +static int nf_flow_tunnel_ip6_push(struct net *net, struct sk_buff *skb,
> +				   struct flow_offload_tuple *tuple,
> +				   struct dst_entry *dst,
> +				   struct in6_addr **ip6_daddr)
>  {
> -	if (tuple->tun_num)
> -		return nf_flow_tunnel_ipip_push(net, skb, tuple, dst, ip_daddr);
> -
> -	return 0;
> -}
> -
> -static int nf_flow_tunnel_ip6ip6_push(struct net *net, struct sk_buff *skb,
> -				      struct flow_offload_tuple *tuple,
> -				      struct dst_entry *dst,
> -				      struct in6_addr **ip6_daddr)
> -{
> -	struct ipv6hdr *ip6h = (struct ipv6hdr *)skb_network_header(skb);
> -	__u8 dsfield = ipv6_get_dsfield(ip6h);
> -	struct rtable *rt = dst_rtable(dst);
>  	struct flowi6 fl6 = {
>  		.daddr = tuple->tun.src_v6,
>  		.saddr = tuple->tun.dst_v6,
> -		.flowi6_proto = IPPROTO_IPV6,
> +		.flowi6_proto = tuple->tun.inner_proto,
>  	};
> -	u8 hop_limit = ip6h->hop_limit;
> +	u8 hop_limit, dsfield;
> +	struct ipv6hdr *ip6h;
>  	int err, mtu;
>  	u32 headroom;
>  
> +	switch (tuple->tun.inner_proto) {
> +	case IPPROTO_IPIP: {
> +		struct iphdr *iph = (struct iphdr *)skb_network_header(skb);
> +
> +		dsfield = ipv4_get_dsfield(iph);
> +		hop_limit = iph->ttl;
> +		break;
> +	}
> +	default:

Same here.

> +		ip6h = (struct ipv6hdr *)skb_network_header(skb);
> +		dsfield = ipv6_get_dsfield(ip6h);
> +		hop_limit = ip6h->hop_limit;
> +		break;
> +	}
> +
>  	err = iptunnel_handle_offloads(skb, SKB_GSO_IPXIP6);
>  	if (err)
>  		return err;
>  
> -	skb_set_inner_ipproto(skb, IPPROTO_IPV6);
> -	headroom = sizeof(*ip6h) + LL_RESERVED_SPACE(rt->dst.dev) +
> -		   rt->dst.header_len;
> +	skb_set_inner_ipproto(skb, tuple->tun.inner_proto);
> +	headroom = sizeof(*ip6h) + LL_RESERVED_SPACE(dst->dev) +
> +		   dst->header_len;
>  	err = skb_cow_head(skb, headroom);
>  	if (err)
>  		return err;
>  
>  	skb_scrub_packet(skb, true);
> -	mtu = dst_mtu(&rt->dst) - sizeof(*ip6h);
> +	mtu = dst_mtu(dst) - sizeof(*ip6h);
>  	mtu = max(mtu, IPV6_MIN_MTU);
>  	skb_dst_update_pmtu_no_confirm(skb, mtu);
>  
>  	skb_push(skb, sizeof(*ip6h));
>  	skb_reset_network_header(skb);
> +	skb->protocol = htons(ETH_P_IPV6);
>  
>  	ip6h = ipv6_hdr(skb);
>  	ip6_flow_hdr(ip6h, dsfield,
>  		     ip6_make_flowlabel(net, skb, fl6.flowlabel, true, &fl6));
>  	ip6h->hop_limit = hop_limit;
> -	ip6h->nexthdr = IPPROTO_IPV6;
> +	ip6h->nexthdr = tuple->tun.inner_proto;
>  	ip6h->daddr = tuple->tun.src_v6;
>  	ip6h->saddr = tuple->tun.dst_v6;
>  	ipv6_hdr(skb)->payload_len = htons(skb->len - sizeof(*ip6h));
> @@ -715,15 +738,20 @@ static int nf_flow_tunnel_ip6ip6_push(struct net *net, struct sk_buff *skb,
>  	return 0;
>  }
>  
> -static int nf_flow_tunnel_v6_push(struct net *net, struct sk_buff *skb,
> -				  struct flow_offload_tuple *tuple,
> -				  struct dst_entry *dst,
> -				  struct in6_addr **ip6_daddr)
> +static int nf_flow_tunnel_push(struct net *net, struct sk_buff *skb,
> +			       struct flow_offload_tuple *tuple,
> +			       struct dst_entry *dst, __be32 *ip_daddr,
> +			       struct in6_addr **ip6_daddr)
>  {
> -	if (tuple->tun_num)
> -		return nf_flow_tunnel_ip6ip6_push(net, skb, tuple, dst, ip6_daddr);
> -
> -	return 0;
> +	switch (tuple->tun.encap_proto) {
> +	case AF_INET:
> +		return nf_flow_tunnel_ip_push(net, skb, tuple, dst, ip_daddr);
> +	case AF_INET6:
> +		return nf_flow_tunnel_ip6_push(net, skb, tuple, dst,
> +					       ip6_daddr);
> +	default:
> +		return 0;
> +	}
>  }
>  
>  static int nf_flow_encap_push(struct sk_buff *skb,
> @@ -830,6 +858,7 @@ static int nf_flow_queue_xmit4(struct sk_buff *skb,
>  	struct flow_offload_tuple *other_tuple;
>  	enum flow_offload_tuple_dir dir;
>  	struct nf_flow_xmit xmit = {};
> +	struct in6_addr *ip6_daddr;
>  	struct flow_offload *flow;
>  	struct neighbour *neigh;
>  	struct rtable *rt;
> @@ -847,9 +876,11 @@ static int nf_flow_queue_xmit4(struct sk_buff *skb,
>  	flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
>  	other_tuple = &flow->tuplehash[!dir].tuple;
>  	ip_daddr = other_tuple->src_v4.s_addr;
> +	ip6_daddr = &other_tuple->src_v6;
>  
> -	if (nf_flow_tunnel_v4_push(state->net, skb, other_tuple,
> -				   tuplehash->tuple.dst_cache, &ip_daddr) < 0)
> +	if (nf_flow_tunnel_push(state->net, skb, other_tuple,
> +				tuplehash->tuple.dst_cache,
> +				&ip_daddr, &ip6_daddr) < 0)

See comment below regarding this.

>  		return NF_DROP;
>  
>  	switch (tuplehash->tuple.xmit_type) {
> @@ -1158,6 +1189,7 @@ static int nf_flow_queue_xmit6(struct sk_buff *skb,
>  	struct flow_offload *flow;
>  	struct neighbour *neigh;
>  	struct rt6_info *rt;
> +	__be32 ip_daddr;
>  
>  	if (unlikely(tuplehash->tuple.xmit_type == FLOW_OFFLOAD_XMIT_XFRM)) {
>  		rt = dst_rt6_info(tuplehash->tuple.dst_cache);
> @@ -1170,11 +1202,12 @@ static int nf_flow_queue_xmit6(struct sk_buff *skb,
>  	dir = tuplehash->tuple.dir;
>  	flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
>  	other_tuple = &flow->tuplehash[!dir].tuple;
> +	ip_daddr = other_tuple->src_v4.s_addr;
>  	ip6_daddr = &other_tuple->src_v6;

IIRC this is pointing to the same address, it is a double fetch of the
same pointer? See below:

>  
> -	if (nf_flow_tunnel_v6_push(state->net, skb, other_tuple,
> -				   tuplehash->tuple.dst_cache,
> -				   &ip6_daddr) < 0)
> +	if (nf_flow_tunnel_push(state->net, skb, other_tuple,
> +				tuplehash->tuple.dst_cache,
> +				&ip_daddr, &ip6_daddr) < 0)

... time to use union nf_inet_addr here instead of these two ip_daddr
and ip6_daddr?

>  		return NF_DROP;
>  
>  	switch (tuplehash->tuple.xmit_type) {
> 
> -- 
> 2.55.0
> 
> 

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH nf-next v2 6/6] net: netfilter: nf_flow_table: unify tunnel push for IPv4 and IPv6
  2026-09-16 22:14   ` Pablo Neira Ayuso
@ 2026-09-17  8:26     ` Lorenzo Bianconi
  0 siblings, 0 replies; 9+ messages in thread
From: Lorenzo Bianconi @ 2026-09-17  8:26 UTC (permalink / raw)
  To: Pablo Neira Ayuso
  Cc: Florian Westphal, Phil Sutter, David S. Miller, Eric Dumazet,
	Jakub Kicinski, Paolo Abeni, Simon Horman, Andrew Lunn,
	David Ahern, Ido Schimmel, netfilter-devel, coreteam, netdev

[-- Attachment #1: Type: text/plain, Size: 10227 bytes --]

> On Mon, Sep 07, 2026 at 09:33:41AM +0200, Lorenzo Bianconi wrote:
> > Refactor nf_flow_tunnel_ipip_push() and nf_flow_tunnel_ip6ip6_push()
> > into nf_flow_tunnel_ip_push() and nf_flow_tunnel_ip6_push(), keying the
> > inner header handling off tuple->tun.inner_proto so both IP-in-IP and
> > IPv6-in-IPv6 inner protocols are supported regardless of the outer
> > address family. Replace nf_flow_tunnel_v4_push() and
> > nf_flow_tunnel_v6_push() with a single nf_flow_tunnel_push() that
> > dispatches on tuple->tun.encap_proto, and set skb->protocol explicitly
> > after pushing the outer header.
> > This is a preliminary patch to support IPv4 over IPv6 and SIT tunnel
> > flowtable offload.
> > Please note IPv4 over IPv6 and SIT tunnel flowtable offloading is not
> > enabled yet.
> > 
> > Signed-off-by: Lorenzo Bianconi <lorenzo.bianconi@oss.qualcomm.com>
> > ---
> >  net/netfilter/nf_flow_table_ip.c | 129 ++++++++++++++++++++++++---------------
> >  1 file changed, 81 insertions(+), 48 deletions(-)
> > 
> > diff --git a/net/netfilter/nf_flow_table_ip.c b/net/netfilter/nf_flow_table_ip.c
> > index 96dbdadba4e7..25875cc97bda 100644
> > --- a/net/netfilter/nf_flow_table_ip.c
> > +++ b/net/netfilter/nf_flow_table_ip.c
> > @@ -608,23 +608,41 @@ static int nf_flow_pppoe_push(struct sk_buff *skb, u16 id,
> >  	return 0;
> >  }
> >  
> > -static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
> > -				    struct flow_offload_tuple *tuple,
> > -				    struct dst_entry *dst, __be32 *ip_daddr)
> > +static int nf_flow_tunnel_ip_push(struct net *net, struct sk_buff *skb,
> > +				  struct flow_offload_tuple *tuple,
> > +				  struct dst_entry *dst, __be32 *ip_daddr)
> >  {
> > -	struct iphdr *iph = (struct iphdr *)skb_network_header(skb);
> > -	struct rtable *rt = dst_rtable(dst);
> > -	u8 tos = iph->tos, ttl = iph->ttl;
> > -	__be16 frag_off = iph->frag_off;
> > -	u32 headroom = sizeof(*iph);
> > +	__be16 frag_off = 0;
> > +	struct iphdr *iph;
> > +	u8 tos = 0, ttl;
> > +	u32 headroom;
> >  	int err;
> >  
> > +	switch (tuple->tun.inner_proto) {
> > +	case IPPROTO_IPV6: {
> > +		struct ipv6hdr *ip6h;
> > +
> > +		ip6h = (struct ipv6hdr *)skb_network_header(skb);
> > +		tos = ipv6_get_dsfield(ip6h);
> > +		ttl = ip6h->hop_limit;
> > +		frag_off = htons(IP_DF);
> > +		break;
> > +	}
> > +	default:
> 
> Please add an explicit case to check for IPv4 here, ie. no default:

ack, I will fix it in v3.

> 
> > +		iph = (struct iphdr *)skb_network_header(skb);
> > +		frag_off = iph->frag_off;
> > +		tos = iph->tos;
> > +		ttl = iph->ttl;
> > +		break;
> > +	}
> > +
> >  	err = iptunnel_handle_offloads(skb, SKB_GSO_IPXIP4);
> >  	if (err)
> >  		return err;
> >  
> > -	skb_set_inner_ipproto(skb, IPPROTO_IPIP);
> > -	headroom += LL_RESERVED_SPACE(rt->dst.dev) + rt->dst.header_len;
> > +	skb_set_inner_ipproto(skb, tuple->tun.inner_proto);
> > +	headroom = sizeof(*iph) + LL_RESERVED_SPACE(dst->dev) +
> > +		   dst->header_len;
> >  	err = skb_cow_head(skb, headroom);
> >  	if (err)
> >  		return err;
> > @@ -635,11 +653,12 @@ static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
> >  	/* Push down and install the IP header. */
> >  	skb_push(skb, sizeof(*iph));
> >  	skb_reset_network_header(skb);
> > +	skb->protocol = htons(ETH_P_IP);
> >  
> >  	iph = ip_hdr(skb);
> >  	iph->version	= 4;
> >  	iph->ihl	= sizeof(*iph) >> 2;
> > -	iph->frag_off	= ip_mtu_locked(&rt->dst) ? 0 : frag_off;
> > +	iph->frag_off	= ip_mtu_locked(dst) ? 0 : frag_off;
> >  	iph->protocol	= tuple->tun.inner_proto;
> >  	iph->tos	= tos;
> >  	iph->daddr	= tuple->tun.src_v4.s_addr;
> > @@ -654,57 +673,61 @@ static int nf_flow_tunnel_ipip_push(struct net *net, struct sk_buff *skb,
> >  	return 0;
> >  }
> >  
> > -static int nf_flow_tunnel_v4_push(struct net *net, struct sk_buff *skb,
> > -				  struct flow_offload_tuple *tuple,
> > -				  struct dst_entry *dst,  __be32 *ip_daddr)
> > +static int nf_flow_tunnel_ip6_push(struct net *net, struct sk_buff *skb,
> > +				   struct flow_offload_tuple *tuple,
> > +				   struct dst_entry *dst,
> > +				   struct in6_addr **ip6_daddr)
> >  {
> > -	if (tuple->tun_num)
> > -		return nf_flow_tunnel_ipip_push(net, skb, tuple, dst, ip_daddr);
> > -
> > -	return 0;
> > -}
> > -
> > -static int nf_flow_tunnel_ip6ip6_push(struct net *net, struct sk_buff *skb,
> > -				      struct flow_offload_tuple *tuple,
> > -				      struct dst_entry *dst,
> > -				      struct in6_addr **ip6_daddr)
> > -{
> > -	struct ipv6hdr *ip6h = (struct ipv6hdr *)skb_network_header(skb);
> > -	__u8 dsfield = ipv6_get_dsfield(ip6h);
> > -	struct rtable *rt = dst_rtable(dst);
> >  	struct flowi6 fl6 = {
> >  		.daddr = tuple->tun.src_v6,
> >  		.saddr = tuple->tun.dst_v6,
> > -		.flowi6_proto = IPPROTO_IPV6,
> > +		.flowi6_proto = tuple->tun.inner_proto,
> >  	};
> > -	u8 hop_limit = ip6h->hop_limit;
> > +	u8 hop_limit, dsfield;
> > +	struct ipv6hdr *ip6h;
> >  	int err, mtu;
> >  	u32 headroom;
> >  
> > +	switch (tuple->tun.inner_proto) {
> > +	case IPPROTO_IPIP: {
> > +		struct iphdr *iph = (struct iphdr *)skb_network_header(skb);
> > +
> > +		dsfield = ipv4_get_dsfield(iph);
> > +		hop_limit = iph->ttl;
> > +		break;
> > +	}
> > +	default:
> 
> Same here.

ack, I will fix it in v3.

> 
> > +		ip6h = (struct ipv6hdr *)skb_network_header(skb);
> > +		dsfield = ipv6_get_dsfield(ip6h);
> > +		hop_limit = ip6h->hop_limit;
> > +		break;
> > +	}
> > +
> >  	err = iptunnel_handle_offloads(skb, SKB_GSO_IPXIP6);
> >  	if (err)
> >  		return err;
> >  
> > -	skb_set_inner_ipproto(skb, IPPROTO_IPV6);
> > -	headroom = sizeof(*ip6h) + LL_RESERVED_SPACE(rt->dst.dev) +
> > -		   rt->dst.header_len;
> > +	skb_set_inner_ipproto(skb, tuple->tun.inner_proto);
> > +	headroom = sizeof(*ip6h) + LL_RESERVED_SPACE(dst->dev) +
> > +		   dst->header_len;
> >  	err = skb_cow_head(skb, headroom);
> >  	if (err)
> >  		return err;
> >  
> >  	skb_scrub_packet(skb, true);
> > -	mtu = dst_mtu(&rt->dst) - sizeof(*ip6h);
> > +	mtu = dst_mtu(dst) - sizeof(*ip6h);
> >  	mtu = max(mtu, IPV6_MIN_MTU);
> >  	skb_dst_update_pmtu_no_confirm(skb, mtu);
> >  
> >  	skb_push(skb, sizeof(*ip6h));
> >  	skb_reset_network_header(skb);
> > +	skb->protocol = htons(ETH_P_IPV6);
> >  
> >  	ip6h = ipv6_hdr(skb);
> >  	ip6_flow_hdr(ip6h, dsfield,
> >  		     ip6_make_flowlabel(net, skb, fl6.flowlabel, true, &fl6));
> >  	ip6h->hop_limit = hop_limit;
> > -	ip6h->nexthdr = IPPROTO_IPV6;
> > +	ip6h->nexthdr = tuple->tun.inner_proto;
> >  	ip6h->daddr = tuple->tun.src_v6;
> >  	ip6h->saddr = tuple->tun.dst_v6;
> >  	ipv6_hdr(skb)->payload_len = htons(skb->len - sizeof(*ip6h));
> > @@ -715,15 +738,20 @@ static int nf_flow_tunnel_ip6ip6_push(struct net *net, struct sk_buff *skb,
> >  	return 0;
> >  }
> >  
> > -static int nf_flow_tunnel_v6_push(struct net *net, struct sk_buff *skb,
> > -				  struct flow_offload_tuple *tuple,
> > -				  struct dst_entry *dst,
> > -				  struct in6_addr **ip6_daddr)
> > +static int nf_flow_tunnel_push(struct net *net, struct sk_buff *skb,
> > +			       struct flow_offload_tuple *tuple,
> > +			       struct dst_entry *dst, __be32 *ip_daddr,
> > +			       struct in6_addr **ip6_daddr)
> >  {
> > -	if (tuple->tun_num)
> > -		return nf_flow_tunnel_ip6ip6_push(net, skb, tuple, dst, ip6_daddr);
> > -
> > -	return 0;
> > +	switch (tuple->tun.encap_proto) {
> > +	case AF_INET:
> > +		return nf_flow_tunnel_ip_push(net, skb, tuple, dst, ip_daddr);
> > +	case AF_INET6:
> > +		return nf_flow_tunnel_ip6_push(net, skb, tuple, dst,
> > +					       ip6_daddr);
> > +	default:
> > +		return 0;
> > +	}
> >  }
> >  
> >  static int nf_flow_encap_push(struct sk_buff *skb,
> > @@ -830,6 +858,7 @@ static int nf_flow_queue_xmit4(struct sk_buff *skb,
> >  	struct flow_offload_tuple *other_tuple;
> >  	enum flow_offload_tuple_dir dir;
> >  	struct nf_flow_xmit xmit = {};
> > +	struct in6_addr *ip6_daddr;
> >  	struct flow_offload *flow;
> >  	struct neighbour *neigh;
> >  	struct rtable *rt;
> > @@ -847,9 +876,11 @@ static int nf_flow_queue_xmit4(struct sk_buff *skb,
> >  	flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
> >  	other_tuple = &flow->tuplehash[!dir].tuple;
> >  	ip_daddr = other_tuple->src_v4.s_addr;
> > +	ip6_daddr = &other_tuple->src_v6;
> >  
> > -	if (nf_flow_tunnel_v4_push(state->net, skb, other_tuple,
> > -				   tuplehash->tuple.dst_cache, &ip_daddr) < 0)
> > +	if (nf_flow_tunnel_push(state->net, skb, other_tuple,
> > +				tuplehash->tuple.dst_cache,
> > +				&ip_daddr, &ip6_daddr) < 0)
> 
> See comment below regarding this.
> 
> >  		return NF_DROP;
> >  
> >  	switch (tuplehash->tuple.xmit_type) {
> > @@ -1158,6 +1189,7 @@ static int nf_flow_queue_xmit6(struct sk_buff *skb,
> >  	struct flow_offload *flow;
> >  	struct neighbour *neigh;
> >  	struct rt6_info *rt;
> > +	__be32 ip_daddr;
> >  
> >  	if (unlikely(tuplehash->tuple.xmit_type == FLOW_OFFLOAD_XMIT_XFRM)) {
> >  		rt = dst_rt6_info(tuplehash->tuple.dst_cache);
> > @@ -1170,11 +1202,12 @@ static int nf_flow_queue_xmit6(struct sk_buff *skb,
> >  	dir = tuplehash->tuple.dir;
> >  	flow = container_of(tuplehash, struct flow_offload, tuplehash[dir]);
> >  	other_tuple = &flow->tuplehash[!dir].tuple;
> > +	ip_daddr = other_tuple->src_v4.s_addr;
> >  	ip6_daddr = &other_tuple->src_v6;
> 
> IIRC this is pointing to the same address, it is a double fetch of the
> same pointer? See below:

right.

> 
> >  
> > -	if (nf_flow_tunnel_v6_push(state->net, skb, other_tuple,
> > -				   tuplehash->tuple.dst_cache,
> > -				   &ip6_daddr) < 0)
> > +	if (nf_flow_tunnel_push(state->net, skb, other_tuple,
> > +				tuplehash->tuple.dst_cache,
> > +				&ip_daddr, &ip6_daddr) < 0)
> 
> ... time to use union nf_inet_addr here instead of these two ip_daddr
> and ip6_daddr?

ack, I will fix it in v3.

Regards,
Lorenzo

> 
> >  		return NF_DROP;
> >  
> >  	switch (tuplehash->tuple.xmit_type) {
> > 
> > -- 
> > 2.55.0
> > 
> > 

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 228 bytes --]

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-09-17  8:26 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-07  7:33 [PATCH nf-next v2 0/6] net: netfilter: preliminary support for IPv4 over IPv6 and SIT flowtable offload Lorenzo Bianconi
2026-09-07  7:33 ` [PATCH nf-next v2 1/6] net: netfilter: nf_flow_table: recognize IPv4/IPv6 in tunnel proto matching Lorenzo Bianconi
2026-09-07  7:33 ` [PATCH nf-next v2 2/6] net: netfilter: nf_flow_table: set inner protocol when popping tunnel header Lorenzo Bianconi
2026-09-07  7:33 ` [PATCH nf-next v2 3/6] net: netfilter: nf_flow_table: populate tunnel tuple regardless of inner protocol Lorenzo Bianconi
2026-09-07  7:33 ` [PATCH nf-next v2 4/6] net: netfilter: add encap_proto to flow_offload_tunnel Lorenzo Bianconi
2026-09-07  7:33 ` [PATCH nf-next v2 5/6] net: netfilter: nf_flow_table: refactor MTU check for tunnel offload Lorenzo Bianconi
2026-09-07  7:33 ` [PATCH nf-next v2 6/6] net: netfilter: nf_flow_table: unify tunnel push for IPv4 and IPv6 Lorenzo Bianconi
2026-09-16 22:14   ` Pablo Neira Ayuso
2026-09-17  8:26     ` Lorenzo Bianconi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox