* [PATCH net-next v2] net: dropreason: add SKB_DROP_REASON_IP_TTL_EXCEEDED
@ 2026-09-01 2:06 Junjie Cao
2026-09-01 2:23 ` David Ahern
0 siblings, 1 reply; 3+ messages in thread
From: Junjie Cao @ 2026-09-01 2:06 UTC (permalink / raw)
To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
Cc: Simon Horman, David Ahern, Ido Schimmel,
Fernando Fernandez Mancera, netdev, linux-kernel
The forwarding paths report an expired TTL or hop limit as
SKB_DROP_REASON_IP_INHDR, the reason otherwise used for a header that is
malformed (ip_input.c, exthdrs.c, br_netfilter). Nothing else in the drop
path separates the two: IPSTATS_MIB_INHDRERRORS covers both, and the TTL
check runs before NF_INET_FORWARD, so netfilter tracing stops at
PREROUTING and never sees the drop.
The Fedora bug linked below shows how that reads in practice. The
reporter took kfree_skb(reason=IP_INHDR, loc=ip_forward) to mean the
software header checksum check had failed, and worked through RX checksum
offload, tc csum actions and both libvirt firewall backends before the
drops turned out to be replies arriving with TTL 1. ip_forward() never
verifies the header checksum; that runs earlier, in ip_rcv_core(), and
reports IP_CSUM.
TTL expiry is not a corner case -- every traceroute through a Linux
router goes through too_many_hops.
The three loopback hop limit checks in exthdrs.c drop with no reason at
all; give them the new one.
IPSTATS_MIB_INHDRERRORS stays as it is: RFC 1213 counts time-to-live
exceeded under ipInHdrErrors. The drop reason has no such constraint.
Link: https://bugzilla.redhat.com/show_bug.cgi?id=2517131
Signed-off-by: Junjie Cao <junjie.cao@intel.com>
---
v2: repost after net-next reopened; no code change.
v1: https://lore.kernel.org/netdev/20260825073906.336072-1-junjie.cao@intel.com/
include/net/dropreason-core.h | 6 ++++++
net/ipv4/ip_forward.c | 2 +-
net/ipv6/exthdrs.c | 6 +++---
net/ipv6/ip6_output.c | 2 +-
4 files changed, 11 insertions(+), 5 deletions(-)
diff --git a/include/net/dropreason-core.h b/include/net/dropreason-core.h
index 2f312d1f67d6..3046a2699479 100644
--- a/include/net/dropreason-core.h
+++ b/include/net/dropreason-core.h
@@ -128,6 +128,7 @@
FN(PSP_INPUT) \
FN(PSP_OUTPUT) \
FN(RECURSION_LIMIT) \
+ FN(IP_TTL_EXCEEDED) \
FNe(MAX)
/**
@@ -606,6 +607,11 @@ enum skb_drop_reason {
SKB_DROP_REASON_PSP_OUTPUT,
/** @SKB_DROP_REASON_RECURSION_LIMIT: Dead loop on virtual device. */
SKB_DROP_REASON_RECURSION_LIMIT,
+ /**
+ * @SKB_DROP_REASON_IP_TTL_EXCEEDED: IPv4 TTL or IPv6 hop limit hit
+ * zero on a packet being forwarded (see IPSTATS_MIB_INHDRERRORS)
+ */
+ SKB_DROP_REASON_IP_TTL_EXCEEDED,
/**
* @SKB_DROP_REASON_MAX: the maximum of core drop reasons, which
* shouldn't be used as a real 'reason' - only for tracing code gen
diff --git a/net/ipv4/ip_forward.c b/net/ipv4/ip_forward.c
index 8b65f12583eb..b242561d37e7 100644
--- a/net/ipv4/ip_forward.c
+++ b/net/ipv4/ip_forward.c
@@ -174,7 +174,7 @@ int ip_forward(struct sk_buff *skb)
/* Tell the sender its packet died... */
__IP_INC_STATS(net, IPSTATS_MIB_INHDRERRORS);
icmp_send(skb, ICMP_TIME_EXCEEDED, ICMP_EXC_TTL, 0);
- SKB_DR_SET(reason, IP_INHDR);
+ SKB_DR_SET(reason, IP_TTL_EXCEEDED);
drop:
kfree_skb_reason(skb, reason);
return NET_RX_DROP;
diff --git a/net/ipv6/exthdrs.c b/net/ipv6/exthdrs.c
index 51941ad656a3..74caaf8746ff 100644
--- a/net/ipv6/exthdrs.c
+++ b/net/ipv6/exthdrs.c
@@ -464,7 +464,7 @@ static int ipv6_srh_rcv(struct sk_buff *skb, struct inet6_dev *idev)
__IP6_INC_STATS(net, idev, IPSTATS_MIB_INHDRERRORS);
icmpv6_send(skb, ICMPV6_TIME_EXCEED,
ICMPV6_EXC_HOPLIMIT, 0);
- kfree_skb(skb);
+ kfree_skb_reason(skb, SKB_DROP_REASON_IP_TTL_EXCEEDED);
return -1;
}
ipv6_hdr(skb)->hop_limit--;
@@ -623,7 +623,7 @@ static int ipv6_rpl_srh_rcv(struct sk_buff *skb, struct inet6_dev *idev)
__IP6_INC_STATS(net, idev, IPSTATS_MIB_INHDRERRORS);
icmpv6_send(skb, ICMPV6_TIME_EXCEED,
ICMPV6_EXC_HOPLIMIT, 0);
- kfree_skb(skb);
+ kfree_skb_reason(skb, SKB_DROP_REASON_IP_TTL_EXCEEDED);
return -1;
}
ipv6_hdr(skb)->hop_limit--;
@@ -815,7 +815,7 @@ static int ipv6_rthdr_rcv(struct sk_buff *skb)
__IP6_INC_STATS(net, idev, IPSTATS_MIB_INHDRERRORS);
icmpv6_send(skb, ICMPV6_TIME_EXCEED, ICMPV6_EXC_HOPLIMIT,
0);
- kfree_skb(skb);
+ kfree_skb_reason(skb, SKB_DROP_REASON_IP_TTL_EXCEEDED);
return -1;
}
ipv6_hdr(skb)->hop_limit--;
diff --git a/net/ipv6/ip6_output.c b/net/ipv6/ip6_output.c
index 8fc4766c8da9..0b6d78c8b6be 100644
--- a/net/ipv6/ip6_output.c
+++ b/net/ipv6/ip6_output.c
@@ -577,7 +577,7 @@ int ip6_forward(struct sk_buff *skb)
icmpv6_send(skb, ICMPV6_TIME_EXCEED, ICMPV6_EXC_HOPLIMIT, 0);
__IP6_INC_STATS(net, idev, IPSTATS_MIB_INHDRERRORS);
- kfree_skb_reason(skb, SKB_DROP_REASON_IP_INHDR);
+ kfree_skb_reason(skb, SKB_DROP_REASON_IP_TTL_EXCEEDED);
return -ETIMEDOUT;
}
--
2.43.0
^ permalink raw reply related [flat|nested] 3+ messages in thread
* Re: [PATCH net-next v2] net: dropreason: add SKB_DROP_REASON_IP_TTL_EXCEEDED
2026-09-01 2:06 [PATCH net-next v2] net: dropreason: add SKB_DROP_REASON_IP_TTL_EXCEEDED Junjie Cao
@ 2026-09-01 2:23 ` David Ahern
2026-09-01 2:50 ` Junjie Cao
0 siblings, 1 reply; 3+ messages in thread
From: David Ahern @ 2026-09-01 2:23 UTC (permalink / raw)
To: Junjie Cao, David S. Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni
Cc: Simon Horman, Ido Schimmel, Fernando Fernandez Mancera, netdev,
linux-kernel
On 8/31/26 8:06 PM, Junjie Cao wrote:
> diff --git a/include/net/dropreason-core.h b/include/net/dropreason-core.h
> index 2f312d1f67d6..3046a2699479 100644
> --- a/include/net/dropreason-core.h
> +++ b/include/net/dropreason-core.h
> @@ -128,6 +128,7 @@
> FN(PSP_INPUT) \
> FN(PSP_OUTPUT) \
> FN(RECURSION_LIMIT) \
> + FN(IP_TTL_EXCEEDED) \
> FNe(MAX)
>
> /**
> @@ -606,6 +607,11 @@ enum skb_drop_reason {
> SKB_DROP_REASON_PSP_OUTPUT,
> /** @SKB_DROP_REASON_RECURSION_LIMIT: Dead loop on virtual device. */
> SKB_DROP_REASON_RECURSION_LIMIT,
> + /**
> + * @SKB_DROP_REASON_IP_TTL_EXCEEDED: IPv4 TTL or IPv6 hop limit hit
> + * zero on a packet being forwarded (see IPSTATS_MIB_INHDRERRORS)
> + */
> + SKB_DROP_REASON_IP_TTL_EXCEEDED,
> /**
> * @SKB_DROP_REASON_MAX: the maximum of core drop reasons, which
> * shouldn't be used as a real 'reason' - only for tracing code gen
> diff --git a/net/ipv4/ip_forward.c b/net/ipv4/ip_forward.c
> index 8b65f12583eb..b242561d37e7 100644
> --- a/net/ipv4/ip_forward.c
> +++ b/net/ipv4/ip_forward.c
> @@ -174,7 +174,7 @@ int ip_forward(struct sk_buff *skb)
> /* Tell the sender its packet died... */
> __IP_INC_STATS(net, IPSTATS_MIB_INHDRERRORS);
> icmp_send(skb, ICMP_TIME_EXCEEDED, ICMP_EXC_TTL, 0);
> - SKB_DR_SET(reason, IP_INHDR);
> + SKB_DR_SET(reason, IP_TTL_EXCEEDED);
> drop:
> kfree_skb_reason(skb, reason);
> return NET_RX_DROP;
what about icmp_unreach, ip_expire, and other sources of
ICMP_TIME_EXCEEDED? The reason code applies to more than just forwarding
paths.
> diff --git a/net/ipv6/exthdrs.c b/net/ipv6/exthdrs.c
> index 51941ad656a3..74caaf8746ff 100644
> --- a/net/ipv6/exthdrs.c
> +++ b/net/ipv6/exthdrs.c
> @@ -464,7 +464,7 @@ static int ipv6_srh_rcv(struct sk_buff *skb, struct inet6_dev *idev)
> __IP6_INC_STATS(net, idev, IPSTATS_MIB_INHDRERRORS);
> icmpv6_send(skb, ICMPV6_TIME_EXCEED,
> ICMPV6_EXC_HOPLIMIT, 0);
> - kfree_skb(skb);
> + kfree_skb_reason(skb, SKB_DROP_REASON_IP_TTL_EXCEEDED);
> return -1;
> }
> ipv6_hdr(skb)->hop_limit--;
> @@ -623,7 +623,7 @@ static int ipv6_rpl_srh_rcv(struct sk_buff *skb, struct inet6_dev *idev)
> __IP6_INC_STATS(net, idev, IPSTATS_MIB_INHDRERRORS);
> icmpv6_send(skb, ICMPV6_TIME_EXCEED,
> ICMPV6_EXC_HOPLIMIT, 0);
> - kfree_skb(skb);
> + kfree_skb_reason(skb, SKB_DROP_REASON_IP_TTL_EXCEEDED);
> return -1;
> }
> ipv6_hdr(skb)->hop_limit--;
> @@ -815,7 +815,7 @@ static int ipv6_rthdr_rcv(struct sk_buff *skb)
> __IP6_INC_STATS(net, idev, IPSTATS_MIB_INHDRERRORS);
> icmpv6_send(skb, ICMPV6_TIME_EXCEED, ICMPV6_EXC_HOPLIMIT,
> 0);
> - kfree_skb(skb);
> + kfree_skb_reason(skb, SKB_DROP_REASON_IP_TTL_EXCEEDED);
> return -1;
> }
> ipv6_hdr(skb)->hop_limit--;
> diff --git a/net/ipv6/ip6_output.c b/net/ipv6/ip6_output.c
> index 8fc4766c8da9..0b6d78c8b6be 100644
> --- a/net/ipv6/ip6_output.c
> +++ b/net/ipv6/ip6_output.c
> @@ -577,7 +577,7 @@ int ip6_forward(struct sk_buff *skb)
> icmpv6_send(skb, ICMPV6_TIME_EXCEED, ICMPV6_EXC_HOPLIMIT, 0);
> __IP6_INC_STATS(net, idev, IPSTATS_MIB_INHDRERRORS);
>
> - kfree_skb_reason(skb, SKB_DROP_REASON_IP_INHDR);
> + kfree_skb_reason(skb, SKB_DROP_REASON_IP_TTL_EXCEEDED);
> return -ETIMEDOUT;
> }
>
similarly for IPv6, there are more ICMPV6_TIME_EXCEED cases than just
the forwarding path.
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH net-next v2] net: dropreason: add SKB_DROP_REASON_IP_TTL_EXCEEDED
2026-09-01 2:23 ` David Ahern
@ 2026-09-01 2:50 ` Junjie Cao
0 siblings, 0 replies; 3+ messages in thread
From: Junjie Cao @ 2026-09-01 2:50 UTC (permalink / raw)
To: David Ahern
Cc: davem, edumazet, kuba, pabeni, horms, idosch, fmancera, netdev,
linux-kernel
On Mon, Aug 31, 2026 at 8:23 PM David Ahern <dsahern@kernel.org> wrote:
> what about icmp_unreach, ip_expire, and other sources of
> ICMP_TIME_EXCEEDED? The reason code applies to more than just forwarding
> paths.
ip_expire and ip6frag_expire_frag_queue already report
FRAG_REASM_TIMEOUT. That is the reassembly timer (code FRAGTIME), not a
TTL, so it stays.
icmp_unreach is the receiver of a time-exceeded; no packet expires
there.
That leaves IPVS decrement_ttl(): it sends the ICMP, and the callers
then free the skb at their tx_error labels with plain kfree_skb(). A
reason there has to come out of __ip_vs_get_out_rt(), so that is a
separate patch.
> similarly for IPv6, there are more ICMPV6_TIME_EXCEED cases than just
> the forwarding path.
Same picture: frag expiry has FRAG_REASM_TIMEOUT, IPVS is the gap, and
ip6_err_gen_icmpv6_unreach only translates a received v4 error for sit
and gre.
The kernel-doc says "on a packet being forwarded"; v3 drops the
qualifier.
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-09-01 2:50 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-01 2:06 [PATCH net-next v2] net: dropreason: add SKB_DROP_REASON_IP_TTL_EXCEEDED Junjie Cao
2026-09-01 2:23 ` David Ahern
2026-09-01 2:50 ` Junjie Cao
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox