BPF List
 help / color / mirror / Atom feed
* [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
@ 2026-09-19 14:37 Jason Xing
  2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
                   ` (10 more replies)
  0 siblings, 11 replies; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

From: Jason Xing <kernelxing@gmail.com>

Greeting,

It's BPF Timestamping 2.0 that aims to observe the packet latency
more efficiently and simply for different protocols. The current
series is only focused on TCP protocol.

At Netdev 0x1a/Netconf 2026, the history, background, motivation and
rough implementation of the feature were exhaustively introduced[1].

History
=======
- In 2009, Patrick Ohly implemented the basic infrastructure
- In 2014, Willem de Bruijn enhanced the TCP latency observation
- In 2024, Jason Xing proposed its lightweight BPF version
Detailed slides from 28 to 32 [1].

Background
==========
Even though BPF Timetamping 1.0 is comparatively low-overhead,
transparent, it's still complicated due to a few points inherited
from the design:
- Inflexible/fixed reporting phases (qdisc/driver/ack)
  When we confirm the issue arises from the kernel by using attribution
  ability of timestamping feature, we need to further minimize the scope
  until the issue is fixed. That means, we then have to resort to write
  a few complex BPF progs with the similar functionalities (like skb
  level tag) which should not happen.
- Not enough low-overhead
  Serving the sensitive applications, an always-on latency observation
  platform should mitigate the self-impact as much as possible. As we
  can conclude from BPF Timestamping selftests, there are some blocking
  and time-consuming points like where reading/writing BPF maps happen
  in the extremely hot paths.
- Minor flaws
  There are a few minor flaws inherited from the initial design, like
  missing tagging the last packet[2][3][4].
Detailed slides from 33 to 39 [1].

Motivation
==========
During the process of the large scale deployment over the last few years,
we eventually realized timestamping feature doesn't support container
scenario and we need a finer-grained and flexible tracing tool (packet
basis) after a few rounds of attribution of issues.

Design
======
- Start time
  For the specific protocol, we need to accurately set the start time of
  each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
  and the driver time as the start point, so that any BPF program is
  capable of computing the delta between start time and current time.
- Simplicity
  Previous BPF program (like selftests) is too complex to implement. The
  core idea is to make everything as simple as possible. And it should be
  decoupled from BPF area and previous timestamping feature as much as
  possible.
- Flexibility
  BPF program hooking any function with skb parameter can get the latency
  value, which means it's no longer bound to the pre-embeded reporting
  phases (see __skb_tstamp_tx)
- Efficiency
  Avoid the previous BPF operations as much as possible. Make sure the
  feature achieves the lowest performance impact, which means only time
  operations remain.
Detailed slides from 40 to 62 [1].

Implementations
===============
in-kernel
- Find a suitable place to timestamp for each packet
- Pick the right start time for TCP
- Handle the split skb due to various reasons
BPF prog
- Hook any functions that carry skb parameter
- Read out the start time from the skb
- Generate the current time and then compute the latency

Discussion?
===========
- Do we need a kfunc to allow users to reset the start time of each skb?
  What I had in mind is if someone tries to observe the latency between
  two specific functions (rather than tcp_sendmsg_locked).
- Current implementation is real hardware timestamp always wins, which
  means BPF prog possibly gets the hardware time that is not aligned
  with bpf_ktime_get_real_ns.
- After the series, do we need to implement the same logic for
  SYN/FIN/PROBE... As far as I know according to numerous user reports,
  a small handful of issues came from 3-way handshake.
- Reusing the slot of hwtstamp might bring potential problems or make the
  code hard to maintain. Can we add a timestamping specific field in
  skb to deal with the latency observation?
- netdev_data conflict in IGC driver. It seems unavoidable to pollute
  start time when it's enabled. Should V2 feature coexist with hardware
  timestamping?
- Should V2 coexist with net timestamping and BPF timestamping? If not,
  the maintenance should be easier.

Any suggestions are greatly welcome!

After we set how to use it from the perspective of users, I would add
a corresponding selftest for this.

[1]: https://netdevconf.info/0x1A/sessions/bof/network-observability-bof.html
[2]: https://lore.kernel.org/all/20260404150452.83904-1-kerneljasonxing@gmail.com/
[3]: commit 838eb9687691 ("tcp: tcp_tx_timestamp() must look at the rtx queue")
[4]: https://lore.kernel.org/all/20260915214450.2882680-1-dw@davidwei.uk/


Jason Xing (9):
  net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2
  bpf: add bpf_ktime_get_real_ns() kfunc
  tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2
  net: reuse skb_shared_hwtstamps for BPF Timestamping v2
  net-timestamp: use pskb_copy to avoid polluting the orig skb's start
    time
  bpf-timestamping: restore skb hwtstamp if it is used by start time
  tcp: propagate the start time onto every skb in the tx path
  net: generate the start time for every skb in the rx path
  tcp: handle the start time of each split skb in the tx path

 include/linux/skbuff.h         | 17 +++++++++
 include/net/sock.h             |  2 +
 include/uapi/linux/bpf.h       |  5 ++-
 kernel/bpf/helpers.c           |  6 +++
 net/core/dev.c                 | 69 +++++++++++++++++++++++++++++++---
 net/core/filter.c              |  7 ++++
 net/core/skbuff.c              | 10 ++++-
 net/core/sock.c                |  5 +++
 net/ipv4/tcp.c                 |  7 ++++
 net/ipv4/tcp_offload.c         |  3 ++
 net/ipv4/tcp_output.c          |  5 +++
 tools/include/uapi/linux/bpf.h |  5 ++-
 12 files changed, 131 insertions(+), 10 deletions(-)

-- 
2.43.7


^ permalink raw reply	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 14:37 ` [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc Jason Xing
                   ` (9 subsequent siblings)
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

Introduce a per-socket bit SK_BPF_CB_TIMESTAMPING_V2 for BPF Timestamping
V2, that allows a socket opt into recording a "start time" on every
skb it produces or receives, which should be reflected in the rest of
patches.

Add a static key to avoid normal traffic suffered from the performance
affect. It works exactly like netstamp_needed_key.

Use bpf_setsockopt(SK_BPF_CB_FLAGS) to toggle the feature that is the
same mechanism used by SK_BPF_CB_TX_TIMESTAMPING which is actually
version 1.

The whole feature is only reachable from BPF, not from user-space
setsockopt().

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
 include/linux/skbuff.h         |  4 +++
 include/uapi/linux/bpf.h       |  5 ++--
 net/core/dev.c                 | 53 ++++++++++++++++++++++++++++++++++
 net/core/filter.c              |  7 +++++
 net/core/sock.c                |  5 ++++
 tools/include/uapi/linux/bpf.h |  5 ++--
 6 files changed, 75 insertions(+), 4 deletions(-)

diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
index 671c13494566..55ae1653a1cf 100644
--- a/include/linux/skbuff.h
+++ b/include/linux/skbuff.h
@@ -4514,6 +4514,10 @@ static inline void skb_set_delivery_type_by_clockid(struct sk_buff *skb,
 
 DECLARE_STATIC_KEY_FALSE(netstamp_needed_key);
 
+DECLARE_STATIC_KEY_FALSE(bpfts_v2_needed_key);
+void bpfts_v2_enable(void);
+void bpfts_v2_disable(void);
+
 /* It is used in the ingress path to clear the delivery_time.
  * If needed, set the skb->tstamp to the (rcv) timestamp.
  */
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 732b35cc08d1..40feea175a10 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -7143,8 +7143,9 @@ enum {
 
 enum {
 	SK_BPF_CB_TX_TIMESTAMPING	= 1<<0,
-	SK_BPF_CB_MASK			= (SK_BPF_CB_TX_TIMESTAMPING - 1) |
-					   SK_BPF_CB_TX_TIMESTAMPING
+	SK_BPF_CB_TIMESTAMPING_V2		= 1<<1,
+	SK_BPF_CB_MASK			= (SK_BPF_CB_TIMESTAMPING_V2 - 1) |
+					   SK_BPF_CB_TIMESTAMPING_V2
 };
 
 /* List of known BPF sock_ops operators.
diff --git a/net/core/dev.c b/net/core/dev.c
index 38336858c168..4479fc87b789 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -2440,6 +2440,59 @@ void net_disable_timestamp(void)
 }
 EXPORT_SYMBOL(net_disable_timestamp);
 
+DEFINE_STATIC_KEY_FALSE(bpfts_v2_needed_key);
+EXPORT_SYMBOL(bpfts_v2_needed_key);
+#ifdef CONFIG_JUMP_LABEL
+static atomic_t bpfts_v2_needed_deferred;
+static atomic_t bpfts_v2_wanted;
+static void bpfts_v2_clear(struct work_struct *work)
+{
+	int deferred = atomic_xchg(&bpfts_v2_needed_deferred, 0);
+	int wanted;
+
+	wanted = atomic_add_return(deferred, &bpfts_v2_wanted);
+	if (wanted > 0)
+		static_branch_enable(&bpfts_v2_needed_key);
+	else
+		static_branch_disable(&bpfts_v2_needed_key);
+}
+static DECLARE_WORK(bpfts_v2_work, bpfts_v2_clear);
+#endif
+
+void bpfts_v2_enable(void)
+{
+#ifdef CONFIG_JUMP_LABEL
+	int wanted = atomic_read(&bpfts_v2_wanted);
+
+	while (wanted > 0) {
+		if (atomic_try_cmpxchg(&bpfts_v2_wanted, &wanted, wanted + 1))
+			return;
+	}
+	atomic_inc(&bpfts_v2_needed_deferred);
+	schedule_work(&bpfts_v2_work);
+#else
+	static_branch_inc(&bpfts_v2_needed_key);
+#endif
+}
+EXPORT_SYMBOL(bpfts_v2_enable);
+
+void bpfts_v2_disable(void)
+{
+#ifdef CONFIG_JUMP_LABEL
+	int wanted = atomic_read(&bpfts_v2_wanted);
+
+	while (wanted > 1) {
+		if (atomic_try_cmpxchg(&bpfts_v2_wanted, &wanted, wanted - 1))
+			return;
+	}
+	atomic_dec(&bpfts_v2_needed_deferred);
+	schedule_work(&bpfts_v2_work);
+#else
+	static_branch_dec(&bpfts_v2_needed_key);
+#endif
+}
+EXPORT_SYMBOL(bpfts_v2_disable);
+
 static inline void net_timestamp_set(struct sk_buff *skb)
 {
 	skb->tstamp = 0;
diff --git a/net/core/filter.c b/net/core/filter.c
index 61940e753552..3486c5c2d170 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -5468,6 +5468,13 @@ static int sk_bpf_set_get_cb_flags(struct sock *sk, char *optval, bool getopt)
 	if (sk_bpf_cb_flags & ~SK_BPF_CB_MASK)
 		return -EINVAL;
 
+	if ((sk_bpf_cb_flags ^ sk->sk_bpf_cb_flags) & SK_BPF_CB_TIMESTAMPING_V2) {
+		if (sk_bpf_cb_flags & SK_BPF_CB_TIMESTAMPING_V2)
+			bpfts_v2_enable();
+		else
+			bpfts_v2_disable();
+	}
+
 	sk->sk_bpf_cb_flags = sk_bpf_cb_flags;
 
 	return 0;
diff --git a/net/core/sock.c b/net/core/sock.c
index 1ad41904db25..251c88e2ee45 100644
--- a/net/core/sock.c
+++ b/net/core/sock.c
@@ -2364,6 +2364,9 @@ static void __sk_destruct(struct rcu_head *head)
 
 	sock_disable_timestamp(sk, SK_FLAGS_TIMESTAMP);
 
+	if (sk->sk_bpf_cb_flags & SK_BPF_CB_TIMESTAMPING_V2)
+		bpfts_v2_disable();
+
 #ifdef CONFIG_BPF_SYSCALL
 	bpf_sk_storage_free(sk);
 #endif
@@ -2550,6 +2553,8 @@ struct sock *sk_clone(const struct sock *sk, const gfp_t priority,
 
 	if (sock_needs_netstamp(sk) && newsk->sk_flags & SK_FLAGS_TIMESTAMP)
 		net_enable_timestamp();
+	if (newsk->sk_bpf_cb_flags & SK_BPF_CB_TIMESTAMPING_V2)
+		bpfts_v2_enable();
 
 	rcu_read_lock();
 	filter = rcu_dereference(sk->sk_filter);
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 732b35cc08d1..40feea175a10 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -7143,8 +7143,9 @@ enum {
 
 enum {
 	SK_BPF_CB_TX_TIMESTAMPING	= 1<<0,
-	SK_BPF_CB_MASK			= (SK_BPF_CB_TX_TIMESTAMPING - 1) |
-					   SK_BPF_CB_TX_TIMESTAMPING
+	SK_BPF_CB_TIMESTAMPING_V2		= 1<<1,
+	SK_BPF_CB_MASK			= (SK_BPF_CB_TIMESTAMPING_V2 - 1) |
+					   SK_BPF_CB_TIMESTAMPING_V2
 };
 
 /* List of known BPF sock_ops operators.
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
  2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 14:37 ` [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
                   ` (8 subsequent siblings)
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

Introduce a new BPF kfunc bpf_ktime_get_real_ns() that returns the
current CLOCK_REALTIME time. This allows BPF programs to perform
accurate delay calculations without clock domain mismatch issue.

Previously BPF programs only had access to monotonic, boot and
TAI clocks (bpf_ktime_get_ns/boot_ns/tai_ns), none of which can be
subtracted from a realtime start time.

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
https://lore.kernel.org/all/20260518082344.96647-2-kerneljasonxing@gmail.com/
---
 kernel/bpf/helpers.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
index b3cc5c8fc875..7a0db620a273 100644
--- a/kernel/bpf/helpers.c
+++ b/kernel/bpf/helpers.c
@@ -2335,6 +2335,11 @@ void bpf_rb_root_free(const struct btf_field *field, void *rb_root,
 
 __bpf_kfunc_start_defs();
 
+__bpf_kfunc u64 bpf_ktime_get_real_ns(void)
+{
+	return ktime_to_ns(ktime_get_real());
+}
+
 /**
  * bpf_obj_new() - allocate an object described by program BTF
  * @local_type_id__k: type ID in program BTF
@@ -4972,6 +4977,7 @@ BTF_ID_FLAGS(func, bpf_task_work_schedule_resume, KF_IMPLICIT_ARGS)
 BTF_ID_FLAGS(func, bpf_dynptr_from_file)
 BTF_ID_FLAGS(func, bpf_dynptr_file_discard, KF_RELEASE)
 BTF_ID_FLAGS(func, bpf_timer_cancel_async)
+BTF_ID_FLAGS(func, bpf_ktime_get_real_ns)
 BTF_KFUNCS_END(common_btf_ids)
 
 static const struct btf_kfunc_id_set common_kfunc_set = {
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
  2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
  2026-09-19 14:37 ` [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 14:37 ` [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2 Jason Xing
                   ` (7 subsequent siblings)
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

Add a new field sk_start_time to 'struct sock' to support different
types of sockets (TCP/UDP/ICMP...) to record the pre-set start time
value and then propagates it into skbs.

Start time might vary across different protocols, so for TCP the
patch picks the entry of tcp_sendmsg_locked as the start time because
1) it already holds the socket lock,
2) it nearly reflects when to trigger a sendmsg syscall from applications.

The value for TCP is a real (wall-clock) time from ktime_get_real() so
it stays comparable with the timestamps produced elsewhere in the stack
and can be consumed directly by a BPF prog using newly added
bpf_ktime_get_real_ns()).

A later patch copies this value onto every skb entailed by the TCP
write queue.

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
 include/net/sock.h | 2 ++
 net/ipv4/tcp.c     | 4 ++++
 2 files changed, 6 insertions(+)

diff --git a/include/net/sock.h b/include/net/sock.h
index 51185222aac2..782ad5fc8ff5 100644
--- a/include/net/sock.h
+++ b/include/net/sock.h
@@ -314,6 +314,7 @@ struct sk_filter;
   *	@mptcp_retransmit_timer: mptcp retransmit timer
   *	@sk_stamp: time stamp of last packet received
   *	@sk_stamp_seq: lock for accessing sk_stamp on 32 bit architectures only
+  *	@sk_start_time: start time recorded for SK_BPF_CB_TIMESTAMPING_V2
   *	@sk_tsflags: SO_TIMESTAMPING flags
   *	@sk_bpf_cb_flags: used in bpf_setsockopt()
   *	@sk_use_task_frag: allow sk_page_frag() to use current->task_frag.
@@ -552,6 +553,7 @@ struct sock {
 #if BITS_PER_LONG==32
 	seqlock_t		sk_stamp_seq;
 #endif
+	ktime_t			sk_start_time;
 	int			sk_disconnects;
 
 	union {
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index b4237d0e994d..28730a3e4473 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -1129,6 +1129,10 @@ int tcp_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_t size)
 
 	flags = msg->msg_flags;
 
+	if (static_branch_unlikely(&bpfts_v2_needed_key) &&
+	    SK_BPF_CB_FLAG_TEST(sk, SK_BPF_CB_TIMESTAMPING_V2))
+		sk->sk_start_time = ktime_get_real();
+
 	sockc = (struct sockcm_cookie){ .tsflags = READ_ONCE(sk->sk_tsflags) };
 	if (msg->msg_controllen) {
 		sockc_err = sock_cmsg_send(sk, msg, &sockc);
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
                   ` (2 preceding siblings ...)
  2026-09-19 14:37 ` [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 14:37 ` [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time Jason Xing
                   ` (6 subsequent siblings)
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

The ultimate goal is to find a place to store the start time per-packet
basis for different protocols.

It's almost unacceptable to add a new field for latency tracker (BPF
Timestamping V2) for every nornal skb because the space is so precious
nowadays. So what left we can do is to try to find a field to reuse.

The principles are 1) it should be a safe place to store the start
time, 2) it can traverse across netns (candidates like skb->tstamp
are not appropriate due to the clearance by skb_scrub_packet() in
the container scenario), 3) it works in dual directions. Detailed
comparisons and reasons can be found on slides 53-60[1].

As we discussed at Netconf 2026 in Rome, we eventually chose to reuse
hwtstamp with considering the above reasons. Now this reused field has
three meanings.

The two uses between 'hwtstamp' and 'start time' are deliberately ordered:
a real hardware time stamp always wins. The start time is only generated
when the hwtstamp slot is still empty. When hardware timestamping is present,
that means we have a better value from real hardware than the software one,
so choose the hardware timestamp as the start time in the rx path.

[1]: https://netdevconf.info/0x1A/sessions/bof/network-observability-bof.html

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
 include/linux/skbuff.h | 7 +++++++
 1 file changed, 7 insertions(+)

diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
index 55ae1653a1cf..c5347be72ff9 100644
--- a/include/linux/skbuff.h
+++ b/include/linux/skbuff.h
@@ -447,6 +447,12 @@ static inline bool skb_frag_must_loop(struct page *p)
  * struct skb_shared_hwtstamps - hardware time stamps
  * @hwtstamp:		hardware time stamp transformed into duration
  *			since arbitrary point in time
+ * @start_time:		Start time stamped when SK_BPF_CB_TIMESTAMPING_V2
+ *			is enabled; shares storage with @hwtstamp. A real
+ *			hardware time stamp always takes precedence: the start
+ *			time is only recorded when @hwtstamp is still empty
+ *			(no hardware time stamp available), so the two are
+ *			never needed at the same time on a given skb.
  * @netdev_data:	address/cookie of network device driver used as
  *			reference to actual hardware time stamp
  *
@@ -462,6 +468,7 @@ static inline bool skb_frag_must_loop(struct page *p)
 struct skb_shared_hwtstamps {
 	union {
 		ktime_t	hwtstamp;
+		ktime_t	start_time;
 		void *netdev_data;
 	};
 };
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
                   ` (3 preceding siblings ...)
  2026-09-19 14:37 ` [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2 Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 14:37 ` [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by " Jason Xing
                   ` (5 subsequent siblings)
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

When BPF Timestamping V2 is active, start_time is stored in the
skb_shared_hwtstamps union alongside hwtstamp. In __skb_tstamp_tx(), the
!tsonly path uses skb_clone() to create the error-queue skb, which shares
skb_shared_info with the original.

If net timestamping feature and BPF timestamping v2 are both enabled,
When the driver reports a hardware TX timestamp, the subsequent write

    *skb_hwtstamps(skb) = *hwtstamps;

clobbers start_time on the retransmit-queue original through the shared
memory, corrupting the value seen by later BPF hooks and retransmissions.

Apply pskb_copy() instead of skb_clone() in this case. Then pskb_copy()
gives the error-queue skb its own skb_shared_info while still sharing
page frags.

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
 net/core/skbuff.c | 5 ++++-
 1 file changed, 4 insertions(+), 1 deletion(-)

diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 966af3beed94..44aeea713da7 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -5749,7 +5749,10 @@ void __skb_tstamp_tx(struct sk_buff *orig_skb,
 #endif
 			skb = alloc_skb(0, GFP_ATOMIC);
 	} else {
-		skb = skb_clone(orig_skb, GFP_ATOMIC);
+		if (static_branch_unlikely(&bpfts_v2_needed_key) && hwtstamps)
+			skb = pskb_copy(orig_skb, GFP_ATOMIC);
+		else
+			skb = skb_clone(orig_skb, GFP_ATOMIC);
 
 		if (skb_orphan_frags_rx(skb, GFP_ATOMIC)) {
 			kfree_skb(skb);
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by start time
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
                   ` (4 preceding siblings ...)
  2026-09-19 14:37 ` [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 14:37 ` [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path Jason Xing
                   ` (4 subsequent siblings)
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

When BPF Timestamping V1 (SKBTX_BPF) and V2 are both active,
skb_tstamp_tx_report_bpf_timestamping() writes the hardware
timestamp into orig_skb:

    *skb_hwtstamps(skb) = *hwtstamps;

Since orig_skb shares skb_shared_info with the retransmit-queue
original via skb_clone(), this clobbers V2's start_time on the
retransmit-queue skb.

Save start_time before the write and restore it after the V1 callback
returns, so both timestamping generations can coexist.

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
 net/core/skbuff.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index 44aeea713da7..98cd813006b8 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -5686,6 +5686,7 @@ static void skb_tstamp_tx_report_bpf_timestamping(struct sk_buff *skb,
 						  struct sock *sk,
 						  int tstype)
 {
+	ktime_t start = 0;
 	int op;
 
 	switch (tstype) {
@@ -5695,6 +5696,8 @@ static void skb_tstamp_tx_report_bpf_timestamping(struct sk_buff *skb,
 	case SCM_TSTAMP_SND:
 		if (hwtstamps) {
 			op = BPF_SOCK_OPS_TSTAMP_SND_HW_CB;
+			if (static_branch_unlikely(&bpfts_v2_needed_key))
+				start = skb_hwtstamps(skb)->start_time;
 			*skb_hwtstamps(skb) = *hwtstamps;
 		} else {
 			op = BPF_SOCK_OPS_TSTAMP_SND_SW_CB;
@@ -5708,6 +5711,8 @@ static void skb_tstamp_tx_report_bpf_timestamping(struct sk_buff *skb,
 	}
 
 	bpf_skops_tx_timestamping(sk, skb, op);
+	if (static_branch_unlikely(&bpfts_v2_needed_key) && start)
+		skb_hwtstamps(skb)->start_time = start;
 }
 
 void __skb_tstamp_tx(struct sk_buff *orig_skb,
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
                   ` (5 preceding siblings ...)
  2026-09-19 14:37 ` [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by " Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 14:37 ` [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path Jason Xing
                   ` (3 subsequent siblings)
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

Now that sendmsg() records a start time in sk->sk_start_time within
the socket lock protection, copy that value into the shared hwtstamps
of each skb.

The timing is tcp_skb_entail() runs once per newly allocated skb
in the sendmsg() loop, every segment produced by a single sendmsg()
carries the same start time.

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
 net/ipv4/tcp.c | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index 28730a3e4473..85e3bae3569d 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -702,6 +702,9 @@ void tcp_skb_entail(struct sock *sk, struct sk_buff *skb)
 	tcb->tcp_flags = TCPHDR_ACK;
 	__skb_header_release(skb);
 	psp_enqueue_set_decrypted(sk, skb);
+	if (static_branch_unlikely(&bpfts_v2_needed_key) &&
+	    SK_BPF_CB_FLAG_TEST(sk, SK_BPF_CB_TIMESTAMPING_V2))
+		skb_hwtstamps(skb)->start_time = sk->sk_start_time;
 	tcp_add_write_queue_tail(sk, skb);
 	sk_wmem_queued_add(sk, skb->truesize);
 	sk_mem_charge(sk, skb->truesize);
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
                   ` (6 preceding siblings ...)
  2026-09-19 14:37 ` [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 14:37 ` [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path Jason Xing
                   ` (2 subsequent siblings)
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

Extend net_timestamp_check(), so that it also records an in-kernel start
time in the skb's shared hwtstamps. This gives received skbs the same
start_time semantics that TCP already gives to transmitted skbs, letting
a BPF program compute how long an skb has spent in the stack.

In the receiving path, net_timestamp_check() will 1) reuse the hardware
timestamp because the hardware time always takes precedence, or 2) uses
skb->tstamp if any, or 3) generate a current timestamp.

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
 net/core/dev.c | 16 +++++++++++-----
 1 file changed, 11 insertions(+), 5 deletions(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index 4479fc87b789..8a9d2205a884 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -2501,11 +2501,17 @@ static inline void net_timestamp_set(struct sk_buff *skb)
 		skb->tstamp = ktime_get_real();
 }
 
-#define net_timestamp_check(COND, SKB)				\
-	if (static_branch_unlikely(&netstamp_needed_key)) {	\
-		if ((COND) && !(SKB)->tstamp)			\
-			(SKB)->tstamp = ktime_get_real();	\
-	}							\
+#define net_timestamp_check(COND, SKB)					\
+	if (static_branch_unlikely(&netstamp_needed_key) ||		\
+	    static_branch_unlikely(&bpfts_v2_needed_key)) {		\
+		if (static_branch_unlikely(&netstamp_needed_key) &&	\
+		    (COND) && !(SKB)->tstamp)				\
+			(SKB)->tstamp = ktime_get_real();		\
+		if (static_branch_unlikely(&bpfts_v2_needed_key) &&	\
+		    !skb_hwtstamps(SKB)->hwtstamp)			\
+			skb_hwtstamps(SKB)->start_time =		\
+				(SKB)->tstamp ?: ktime_get_real();	\
+	}								\
 
 bool is_skb_forwardable(const struct net_device *dev, const struct sk_buff *skb)
 {
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
                   ` (7 preceding siblings ...)
  2026-09-19 14:37 ` [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path Jason Xing
@ 2026-09-19 14:37 ` Jason Xing
  2026-09-20 14:39   ` sashiko-bot
  2026-09-19 18:12 ` [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Alexei Starovoitov
  2026-09-21 18:45 ` Stanislav Fomichev
  10 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

Segmentation and collapsing create or reuse skbs with a fresh
shared_info, which would otherwise drop the start time recorded
on the original skb.

Handle every TCP path that manipulates the shared hwtstamps:

  - tcp_gso_segment(): handle the gso case. Propagate the same
    gso_skb's start time to every resulting segment.

  - tcp_fragment/tso_fragment(): handle the tso/recovery.. cases.
    Copy the start time onto the second half of the split so both
    fragments keep it.

  - tcp_skb_collapse_tstamp(): handle collapse case. Copy the eaten
    skb's start time onto the survivor.

  - __tcp_retransmit_skb(): handle retrans case - the retransmit
    path __tcp_retransmit_skb->__pskb_copy.

Signed-off-by: Jason Xing <kerneljasonxing@gmail.com>
---
rx path logic: https://lore.kernel.org/all/20260917132742.87117-1-kerneljasonxing@gmail.com/
---
 include/linux/skbuff.h | 6 ++++++
 net/ipv4/tcp_offload.c | 3 +++
 net/ipv4/tcp_output.c  | 5 +++++
 3 files changed, 14 insertions(+)

diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
index c5347be72ff9..5f01239af191 100644
--- a/include/linux/skbuff.h
+++ b/include/linux/skbuff.h
@@ -5491,6 +5491,12 @@ static inline void skb_mark_for_recycle(struct sk_buff *skb)
 #endif
 }
 
+static inline void skb_copy_start_time(struct sk_buff *dst, struct sk_buff *src)
+{
+	if (static_branch_unlikely(&bpfts_v2_needed_key))
+		skb_hwtstamps(dst)->start_time = skb_hwtstamps(src)->start_time;
+}
+
 ssize_t skb_splice_from_iter(struct sk_buff *skb, struct iov_iter *iter,
 			     ssize_t maxsize);
 
diff --git a/net/ipv4/tcp_offload.c b/net/ipv4/tcp_offload.c
index 3b1fdcd3cb29..3b33e9f4fbc9 100644
--- a/net/ipv4/tcp_offload.c
+++ b/net/ipv4/tcp_offload.c
@@ -201,6 +201,8 @@ struct sk_buff *tcp_gso_segment(struct sk_buff *skb,
 	if (unlikely(skb_shinfo(gso_skb)->tx_flags & SKBTX_ANY_TSTAMP))
 		tcp_gso_tstamp(segs, gso_skb, seq, mss);
 
+	skb_copy_start_time(segs, gso_skb);
+
 	newcheck = ~csum_fold(csum_add(csum_unfold(th->check), delta));
 
 	ecn_cwr_mask = !!(skb_shinfo(gso_skb)->gso_type & SKB_GSO_TCP_ACCECN);
@@ -226,6 +228,7 @@ struct sk_buff *tcp_gso_segment(struct sk_buff *skb,
 		th->seq = htonl(seq);
 
 		th->cwr &= ecn_cwr_mask;
+		skb_copy_start_time(skb, gso_skb);
 	}
 
 	/* Following permits TCP Small Queues to work well with GSO :
diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
index 6f4dca4a4de9..006157321f49 100644
--- a/net/ipv4/tcp_output.c
+++ b/net/ipv4/tcp_output.c
@@ -1900,6 +1900,7 @@ int tcp_fragment(struct sock *sk, enum tcp_queue tcp_queue,
 	tcp_skb_fragment_eor(skb, buff);
 
 	skb_split(skb, buff, len);
+	skb_copy_start_time(buff, skb);
 
 	skb_set_delivery_time(buff, skb->tstamp, SKB_CLOCK_MONOTONIC);
 	tcp_fragment_tstamp(skb, buff);
@@ -2431,6 +2432,7 @@ static int tso_fragment(struct sock *sk, struct sk_buff *skb, unsigned int len,
 	tcp_skb_fragment_eor(skb, buff);
 
 	skb_split(skb, buff, len);
+	skb_copy_start_time(buff, skb);
 	tcp_fragment_tstamp(skb, buff);
 
 	/* Fix up tso_factor for both original and new SKB.  */
@@ -3445,6 +3447,8 @@ void tcp_skb_collapse_tstamp(struct sk_buff *skb,
 		TCP_SKB_CB(skb)->txstamp_ack |=
 			TCP_SKB_CB(next_skb)->txstamp_ack;
 	}
+	if (!skb_hwtstamps(skb)->start_time)
+		skb_copy_start_time(skb, (struct sk_buff *)next_skb);
 }
 
 /* Collapses two adjacent SKB's during retransmission. */
@@ -3661,6 +3665,7 @@ int __tcp_retransmit_skb(struct sock *sk, struct sk_buff *skb, int segs)
 			nskb = __pskb_copy(skb, MAX_TCP_HEADER, GFP_ATOMIC);
 			if (nskb) {
 				nskb->dev = NULL;
+				skb_copy_start_time(nskb, skb);
 				err = tcp_transmit_skb(sk, nskb, 0, GFP_ATOMIC);
 			} else {
 				err = -ENOBUFS;
-- 
2.43.7


^ permalink raw reply related	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
                   ` (8 preceding siblings ...)
  2026-09-19 14:37 ` [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path Jason Xing
@ 2026-09-19 18:12 ` Alexei Starovoitov
  2026-09-20  0:41   ` Jason Xing
  2026-09-21 18:45 ` Stanislav Fomichev
  10 siblings, 1 reply; 27+ messages in thread
From: Alexei Starovoitov @ 2026-09-19 18:12 UTC (permalink / raw)
  To: Jason Xing, davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

On Sat Sep 19, 2026 at 2:37 PM UTC, Jason Xing wrote:
> From: Jason Xing <kernelxing@gmail.com>
>
> Greeting,
>
> It's BPF Timestamping 2.0 that aims to observe the packet latency
> more efficiently and simply for different protocols. The current
> series is only focused on TCP protocol.
>
> At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> rough implementation of the feature were exhaustively introduced[1].
>
> History
> =======
> - In 2009, Patrick Ohly implemented the basic infrastructure
> - In 2014, Willem de Bruijn enhanced the TCP latency observation
> - In 2024, Jason Xing proposed its lightweight BPF version
> Detailed slides from 28 to 32 [1].
>
> Background
> ==========
> Even though BPF Timetamping 1.0 is comparatively low-overhead,
> transparent, it's still complicated due to a few points inherited
> from the design:
> - Inflexible/fixed reporting phases (qdisc/driver/ack)
>   When we confirm the issue arises from the kernel by using attribution
>   ability of timestamping feature, we need to further minimize the scope
>   until the issue is fixed. That means, we then have to resort to write
>   a few complex BPF progs with the similar functionalities (like skb
>   level tag) which should not happen.
> - Not enough low-overhead
>   Serving the sensitive applications, an always-on latency observation
>   platform should mitigate the self-impact as much as possible. As we
>   can conclude from BPF Timestamping selftests, there are some blocking
>   and time-consuming points like where reading/writing BPF maps happen
>   in the extremely hot paths.
> - Minor flaws
>   There are a few minor flaws inherited from the initial design, like
>   missing tagging the last packet[2][3][4].


All these "issues" sound minor and on their own do not warrant new uapi.

This RFC don't even have an example of how it could be used end-to-end.

So soft nack from me for now.

pw-bot: cr

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
  2026-09-19 18:12 ` [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Alexei Starovoitov
@ 2026-09-20  0:41   ` Jason Xing
  0 siblings, 0 replies; 27+ messages in thread
From: Jason Xing @ 2026-09-20  0:41 UTC (permalink / raw)
  To: Alexei Starovoitov
  Cc: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu, netdev,
	bpf, Jason Xing

[-- Attachment #1: Type: text/plain, Size: 2967 bytes --]

On Sun, Sep 20, 2026 at 2:12 AM Alexei Starovoitov
<alexei.starovoitov@gmail.com> wrote:
>
> On Sat Sep 19, 2026 at 2:37 PM UTC, Jason Xing wrote:
> > From: Jason Xing <kernelxing@gmail.com>
> >
> > Greeting,
> >
> > It's BPF Timestamping 2.0 that aims to observe the packet latency
> > more efficiently and simply for different protocols. The current
> > series is only focused on TCP protocol.
> >
> > At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> > rough implementation of the feature were exhaustively introduced[1].
> >
> > History
> > =======
> > - In 2009, Patrick Ohly implemented the basic infrastructure
> > - In 2014, Willem de Bruijn enhanced the TCP latency observation
> > - In 2024, Jason Xing proposed its lightweight BPF version
> > Detailed slides from 28 to 32 [1].
> >
> > Background
> > ==========
> > Even though BPF Timetamping 1.0 is comparatively low-overhead,
> > transparent, it's still complicated due to a few points inherited
> > from the design:
> > - Inflexible/fixed reporting phases (qdisc/driver/ack)
> >   When we confirm the issue arises from the kernel by using attribution
> >   ability of timestamping feature, we need to further minimize the scope
> >   until the issue is fixed. That means, we then have to resort to write
> >   a few complex BPF progs with the similar functionalities (like skb
> >   level tag) which should not happen.
> > - Not enough low-overhead
> >   Serving the sensitive applications, an always-on latency observation
> >   platform should mitigate the self-impact as much as possible. As we
> >   can conclude from BPF Timestamping selftests, there are some blocking
> >   and time-consuming points like where reading/writing BPF maps happen
> >   in the extremely hot paths.
> > - Minor flaws
> >   There are a few minor flaws inherited from the initial design, like
> >   missing tagging the last packet[2][3][4].
>
>
> All these "issues" sound minor and on their own do not warrant new uapi.

Some of them are minor, some definitely not.

Current BPF timestamping 1.0 is not perfect and __doesn't support
container scenario__ which is a big thing.

Besides, performance impact and unflexibility are real problems when
we're trying so hard to deploy the platform. The experience isn't
limited to us; several of my friends at other companies widely deploy
the version 1.0 feature.

>
> This RFC don't even have an example of how it could be used end-to-end.

Sorry, I should've uploaded the examples. Please see the attachment.
# make && ./bpfts_v2

The usage is simple:
1. bpf_setsockopt to enable the feature
2. read out the start time of skb
3. compute the delta

Please note: my intention here is to initiate a discussion on how it
looks like in the future. Reusing hwtstamp potentially risks breaking
some existing applications, which is the point I'm most open to
discussing.

Thanks,
Jason

[-- Attachment #2: Makefile --]
[-- Type: application/octet-stream, Size: 761 bytes --]

CLANG   ?= clang
CC      ?= gcc
BPFTOOL ?= bpftool
ARCH    ?= $(shell uname -m | sed 's/x86_64/x86/' \
			| sed 's/aarch64/arm64/' \
			| sed 's/ppc64le/powerpc/' \
			| sed 's/mips.*/mips/' \
			| sed 's/riscv64/riscv/' \
			| sed 's/loongarch.*/loongarch/')

CFLAGS   := -g -O2 -Wall
INCLUDES := -I.

.PHONY: all clean

all: bpfts_v2

vmlinux.h:
	$(BPFTOOL) btf dump file /sys/kernel/btf/vmlinux format c > $@

bpfts_v2.bpf.o: bpfts_v2.bpf.c vmlinux.h
	$(CLANG) $(CFLAGS) -target bpf -D__TARGET_ARCH_$(ARCH) \
		$(INCLUDES) -c $< -o $@

bpfts_v2.skel.h: bpfts_v2.bpf.o
	$(BPFTOOL) gen skeleton $< > $@

bpfts_v2: bpfts_v2.c bpfts_v2.skel.h
	$(CC) $(CFLAGS) $(INCLUDES) -o $@ $< -lbpf -lelf -lz

clean:
	rm -f bpfts_v2 bpfts_v2.bpf.o bpfts_v2.skel.h vmlinux.h

[-- Attachment #3: bpfts_v2.bpf.c --]
[-- Type: application/octet-stream, Size: 3639 bytes --]

// SPDX-License-Identifier: GPL-2.0
/*
 * BPF Timestamping V2 — standalone BPF program
 *
 * sockops: enable SK_BPF_CB_TIMESTAMPING_V2 on new TCP connections
 * fentry:  read skb start_time and subtract bpf_ktime_get_real_ns()
 *
 * No BPF map on the timing path. Compile outside the kernel tree:
 *   make && sudo ./bpfts_v2
 */
#include "vmlinux.h"
#include <bpf/bpf_helpers.h>
#include <bpf/bpf_tracing.h>
#include <bpf/bpf_core_read.h>

/*
 * bpf_core_cast() landed in newer libbpf. Distro headers (libbpf-dev)
 * often do not have it yet; provide the same expansion.
 */
#ifndef bpf_core_type_id_kernel
#define bpf_core_type_id_kernel(type) \
	__builtin_btf_type_id(*(typeof(type) *)0, 1)
#endif
#ifndef bpf_core_cast
extern void *bpf_rdonly_cast(const void *obj__, __u32 btf_id__) __ksym __weak;
#define bpf_core_cast(ptr, type) \
	((typeof(type) *)bpf_rdonly_cast((ptr), bpf_core_type_id_kernel(type)))
#endif

#define SOL_SOCKET			1
#define SK_BPF_CB_FLAGS			1009
#define SK_BPF_CB_TIMESTAMPING_V2	(1 << 1)

#define DIR_TX	0
#define DIR_RX	1

extern __u64 bpf_ktime_get_real_ns(void) __ksym;

struct lat_event {
	__u64 delta_ns;
	__u32 pid;
	__u8  dir;
	char  comm[16];
};

struct {
	__uint(type, BPF_MAP_TYPE_RINGBUF);
	__uint(max_entries, 256 * 1024);
} events SEC(".maps");

__u64 tx_cnt, rx_cnt;
__u64 tx_sum_ns, rx_sum_ns;
__u64 tx_max_ns, rx_max_ns;
__u64 skip_cnt;

/* Software start_time is CLOCK_REALTIME. A PHC/hardware stamp in the
 * shared union is a different clock; subtracting ktime_get_real()
 * yields ~epoch-sized garbage (1e15+ us). Drop those.
 */
#define MAX_LAT_NS	(10ULL * 1000 * 1000 * 1000)

static __always_inline __u64 skb_start_time(const struct sk_buff *skb)
{
	struct skb_shared_info *shinfo;

	if (!skb)
		return 0;
	/* same storage as start_time (union with hwtstamp) */
	shinfo = bpf_core_cast(skb->head + skb->end, struct skb_shared_info);
	return (__u64)shinfo->hwtstamps.hwtstamp;
}

static __always_inline void record_lat(const struct sk_buff *skb, __u8 dir)
{
	struct lat_event *e;
	__u64 start, delta;

	start = skb_start_time(skb);
	if (!start)
		return;

	delta = bpf_ktime_get_real_ns() - start;
	/* start > now underflows; PHC/hw stamp is usually a small value */
	if (delta > MAX_LAT_NS) {
		__sync_fetch_and_add(&skip_cnt, 1);
		return;
	}

	if (dir == DIR_TX) {
		__sync_fetch_and_add(&tx_cnt, 1);
		__sync_fetch_and_add(&tx_sum_ns, delta);
		if (delta > tx_max_ns)
			tx_max_ns = delta;
	} else {
		__sync_fetch_and_add(&rx_cnt, 1);
		__sync_fetch_and_add(&rx_sum_ns, delta);
		if (delta > rx_max_ns)
			rx_max_ns = delta;
	}

	e = bpf_ringbuf_reserve(&events, sizeof(*e), 0);
	if (!e)
		return;
	e->delta_ns = delta;
	e->pid = bpf_get_current_pid_tgid() >> 32;
	e->dir = dir;
	bpf_get_current_comm(&e->comm, sizeof(e->comm));
	bpf_ringbuf_submit(e, 0);
}

SEC("sockops")
int enable_v2(struct bpf_sock_ops *skops)
{
	int flags = SK_BPF_CB_TIMESTAMPING_V2;

	switch (skops->op) {
	case BPF_SOCK_OPS_ACTIVE_ESTABLISHED_CB:
	case BPF_SOCK_OPS_PASSIVE_ESTABLISHED_CB:
		bpf_setsockopt(skops, SOL_SOCKET, SK_BPF_CB_FLAGS,
			       &flags, sizeof(flags));
		break;
	}
	return 1;
}

/* TX: sendmsg -> device queue */
SEC("fentry/__dev_queue_xmit")
int BPF_PROG(on_dev_xmit, struct sk_buff *skb, struct net_device *sb_dev)
{
	record_lat(skb, DIR_TX);
	return 0;
}

/* RX IPv4: arrival -> TCP receive */
SEC("fentry/tcp_v4_rcv")
int BPF_PROG(on_tcp_v4_rcv, struct sk_buff *skb)
{
	record_lat(skb, DIR_RX);
	return 0;
}

/* RX IPv6 */
SEC("fentry/tcp_v6_rcv")
int BPF_PROG(on_tcp_v6_rcv, struct sk_buff *skb)
{
	record_lat(skb, DIR_RX);
	return 0;
}

char LICENSE[] SEC("license") = "GPL";

[-- Attachment #4: bpfts_v2.c --]
[-- Type: application/octet-stream, Size: 4079 bytes --]

// SPDX-License-Identifier: GPL-2.0
/*
 * Userspace loader for BPF Timestamping V2.
 *
 *   make
 *   sudo ./bpfts_v2
 *
 * Then open a NEW TCP connection (V2 is enabled on ESTABLISHED):
 *   curl -4 http://127.0.0.1/
 *   ping is ICMP — it will not produce events.
 */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <fcntl.h>
#include <signal.h>
#include <unistd.h>
#include <time.h>
#include <bpf/libbpf.h>
#include "bpfts_v2.skel.h"

#define DIR_TX 0
#define DIR_RX 1

static volatile sig_atomic_t exiting;

struct lat_event {
	__u64 delta_ns;
	__u32 pid;
	__u8  dir;
	char  comm[16];
};

static void on_signal(int sig)
{
	exiting = 1;
}

static int libbpf_print(enum libbpf_print_level level, const char *fmt,
			va_list args)
{
	if (level == LIBBPF_DEBUG)
		return 0;
	return vfprintf(stderr, fmt, args);
}

static int handle_event(void *ctx, void *data, size_t len)
{
	const struct lat_event *e = data;

	printf("%s  %-16s pid=%-6u  lat=%8.3f us\n",
	       e->dir == DIR_TX ? "TX" : "RX",
	       e->comm, e->pid, e->delta_ns / 1000.0);
	fflush(stdout);
	return 0;
}

static void print_summary(const struct bpfts_v2_bpf *skel)
{
	__u64 tx = skel->bss->tx_cnt;
	__u64 rx = skel->bss->rx_cnt;

	printf("\n----- summary -----\n");
	if (tx)
		printf("TX  %llu pkts  avg %.3f us  max %.3f us\n",
		       (unsigned long long)tx,
		       skel->bss->tx_sum_ns / (double)tx / 1000.0,
		       skel->bss->tx_max_ns / 1000.0);
	else
		printf("TX  0 pkts\n");
	if (rx)
		printf("RX  %llu pkts  avg %.3f us  max %.3f us\n",
		       (unsigned long long)rx,
		       skel->bss->rx_sum_ns / (double)rx / 1000.0,
		       skel->bss->rx_max_ns / 1000.0);
	else
		printf("RX  0 pkts\n");
	if (skel->bss->skip_cnt)
		printf("skipped %llu (hw/PHC stamp, not CLOCK_REALTIME)\n",
		       (unsigned long long)skel->bss->skip_cnt);
	printf("-------------------\n");
}

static void usage(const char *argv0)
{
	fprintf(stderr,
		"Usage: %s [-c cgroup_path]\n"
		"  -c   cgroup to attach sockops (default: /sys/fs/cgroup)\n"
		"\n"
		"V2 is enabled on TCP connections established AFTER attach.\n"
		"Put the traffic process into that cgroup, or attach the root\n"
		"cgroup (default) so every new TCP socket is covered.\n",
		argv0);
}

int main(int argc, char **argv)
{
	const char *cg_path = "/sys/fs/cgroup";
	struct ring_buffer *rb = NULL;
	struct bpfts_v2_bpf *skel = NULL;
	struct bpf_link *link;
	int opt, err, cg_fd = -1;

	while ((opt = getopt(argc, argv, "c:h")) != -1) {
		switch (opt) {
		case 'c':
			cg_path = optarg;
			break;
		case 'h':
		default:
			usage(argv[0]);
			return opt == 'h' ? 0 : 1;
		}
	}

	libbpf_set_print(libbpf_print);
	signal(SIGINT, on_signal);
	signal(SIGTERM, on_signal);

	skel = bpfts_v2_bpf__open_and_load();
	if (!skel) {
		fprintf(stderr, "failed to open/load BPF object: %s\n",
			strerror(errno));
		return 1;
	}

	/* fentry auto-attach; sockops needs a cgroup fd */
	bpf_program__set_autoattach(skel->progs.enable_v2, false);
	err = bpfts_v2_bpf__attach(skel);
	if (err) {
		fprintf(stderr, "failed to attach fentry: %d\n", err);
		goto out;
	}

	cg_fd = open(cg_path, O_RDONLY);
	if (cg_fd < 0) {
		fprintf(stderr, "open cgroup %s: %s\n", cg_path, strerror(errno));
		goto out;
	}

	link = bpf_program__attach_cgroup(skel->progs.enable_v2, cg_fd);
	if (!link) {
		fprintf(stderr, "attach sockops to %s: %s\n",
			cg_path, strerror(errno));
		goto out;
	}
	skel->links.enable_v2 = link;

	rb = ring_buffer__new(bpf_map__fd(skel->maps.events),
			      handle_event, NULL, NULL);
	if (!rb) {
		fprintf(stderr, "failed to create ring buffer\n");
		goto out;
	}

	printf("BPF Timestamping V2 running (cgroup %s)\n", cg_path);
	printf("Start a NEW TCP connection now, Ctrl-C to stop.\n\n");

	while (!exiting) {
		err = ring_buffer__poll(rb, 200);
		if (err == -EINTR)
			break;
		if (err < 0) {
			fprintf(stderr, "ring_buffer__poll: %d\n", err);
			break;
		}
	}

	print_summary(skel);
	err = 0;
out:
	ring_buffer__free(rb);
	bpfts_v2_bpf__destroy(skel);
	if (cg_fd >= 0)
		close(cg_fd);
	return err != 0;
}

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time
  2026-09-19 14:37 ` [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] net: missing tx_flags and tskey initialization after pskb_copy()
--

commit 25f6ea681d5fbb36cc3d7e6a5d0bca7daa16f7a1
Author: Jason Xing <kerneljasonxing@gmail.com>

net-timestamp: use pskb_copy to avoid polluting the orig skb's start time

This commit changes the behavior of __skb_tstamp_tx() when both net
timestamping and BPF timestamping v2 are enabled. It uses pskb_copy()
instead of skb_clone() to create the error-queue skb, ensuring it gets a
separate skb_shared_info block to prevent hardware TX timestamp writes from
clobbering start_time.

> diff --git a/net/core/skbuff.c b/net/core/skbuff.c
> index dbbe10277d51d..a38c98554ff82 100644
> --- a/net/core/skbuff.c
> +++ b/net/core/skbuff.c
> @@ -5749,7 +5749,10 @@ void __skb_tstamp_tx(struct sk_buff *orig_skb,
[ ... ]
>  			skb = alloc_skb(0, GFP_ATOMIC);
>  	} else {
> -		skb = skb_clone(orig_skb, GFP_ATOMIC);
> +		if (static_branch_unlikely(&bpfts_v2_needed_key) && hwtstamps)
> +			skb = pskb_copy(orig_skb, GFP_ATOMIC);
> +		else
> +			skb = skb_clone(orig_skb, GFP_ATOMIC);

[Severity: High]
Will this leave tx_flags and tskey uninitialized when the pskb_copy() path
is taken?

Unlike skb_clone(), pskb_copy() creates a new zero-initialized
skb_shared_info block, omitting the copying of metadata like tskey and
tx_flags.

Later in __skb_tstamp_tx(), the explicit copying of these fields is bypassed
because the pskb_copy() fallback path means tsonly is false:

        if (tsonly) {
                skb_shinfo(skb)->tx_flags |= skb_shinfo(orig_skb)->tx_flags &
                                             SKBTX_ANY_TSTAMP;
                skb_shinfo(skb)->tskey = skb_shinfo(orig_skb)->tskey;
        }

Because these fields remain zero on the newly copied SKB, will this break
the SOF_TIMESTAMPING_OPT_ID socket option functionality by reporting an
invalid or zero timestamp ID (ee_data) to userspace when a hardware
timestamp is generated?

>  		if (skb_orphan_frags_rx(skb, GFP_ATOMIC)) {
>  			kfree_skb(skb);
>  			return;

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=5

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2
  2026-09-19 14:37 ` [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [Medium] tcp: concurrent sendmsg corrupts per-syscall sk_start_time
--

commit c712b000347bc5edb0a47ae0ad843a63f0d66b24
Author: Jason Xing <kerneljasonxing@gmail.com>

tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2

This commit adds a new sk_start_time field to struct sock to record the
syscall start time for SK_BPF_CB_TIMESTAMPING_V2. The time is recorded on
entry to tcp_sendmsg_locked() for later propagation into skbs.

> diff --git a/include/net/sock.h b/include/net/sock.h
> index 60ea55dc18854..30e7894c15b43 100644
> --- a/include/net/sock.h
> +++ b/include/net/sock.h
[ ... ]
> @@ -552,6 +553,7 @@ struct sock {
>  #if BITS_PER_LONG==32
>  	seqlock_t		sk_stamp_seq;
>  #endif
> +	ktime_t			sk_start_time;

[Severity: Medium]
Is struct sock the appropriate place to store a per-syscall state?

Since this field is in a shared structure, it could be vulnerable to being
overwritten by concurrent calls. Could this time be passed down via a local
execution context like struct sockcm_cookie instead?

>  	int			sk_disconnects;
>  
>  	union {
> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index 3ac4856852794..f07abd53ed27f 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -1132,6 +1132,10 @@ int tcp_sendmsg_locked(struct sock *sk, struct msghdr *msg, size_t size)
>  
>  	flags = msg->msg_flags;
>  
> +	if (static_branch_unlikely(&bpfts_v2_needed_key) &&
> +	    SK_BPF_CB_FLAG_TEST(sk, SK_BPF_CB_TIMESTAMPING_V2))
> +		sk->sk_start_time = ktime_get_real();
> +

[Severity: Medium]
Can concurrent sendmsg() calls corrupt this per-syscall start time?

If the current sendmsg() blocks later in tcp_sendmsg_locked() (for example,
while waiting for memory in sk_stream_wait_memory() or for a connection in
sk_stream_wait_connect()), the socket lock is temporarily released.

If another thread executes a concurrent sendmsg() on the same socket, it
will acquire the lock and overwrite sk->sk_start_time here. When the first
thread resumes execution, won't any new SKBs it allocates incorrectly
inherit the second thread's start time?

>  	sockc = (struct sockcm_cookie){ .tsflags = READ_ONCE(sk->sk_tsflags) };
>  	if (msg->msg_controllen) {

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=3

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by start time
  2026-09-19 14:37 ` [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by " Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: start_time clobbered by hwtstamp when start_time is zero
--

commit 0169de67fe5c3f53f4f477e65c35152b1fba307d
Author: Jason Xing <kerneljasonxing@gmail.com>

bpf-timestamping: restore skb hwtstamp if it is used by start time

When BPF Timestamping V1 (SKBTX_BPF) and V2 are both active, writing the
hardware timestamp into orig_skb clobbers V2's start_time on the
retransmit-queue skb. This commit saves start_time before the write and
restores it after the V1 callback returns, so both timestamping generations
can coexist.

> diff --git a/net/core/skbuff.c b/net/core/skbuff.c
> --- a/net/core/skbuff.c
> +++ b/net/core/skbuff.c
> @@ -5686,6 +5686,7 @@ static void skb_tstamp_tx_report_bpf_timestamping(struct sk_buff *skb,
>  						  struct sock *sk,
>  						  int tstype)
>  {
> +	ktime_t start = 0;
>  	int op;
>  
>  	switch (tstype) {
> @@ -5695,6 +5696,8 @@ static void skb_tstamp_tx_report_bpf_timestamping(struct sk_buff *skb,
>  	case SCM_TSTAMP_SND:
>  		if (hwtstamps) {
>  			op = BPF_SOCK_OPS_TSTAMP_SND_HW_CB;
> +			if (static_branch_unlikely(&bpfts_v2_needed_key))
> +				start = skb_hwtstamps(skb)->start_time;
>  			*skb_hwtstamps(skb) = *hwtstamps;
>  		} else {
>  			op = BPF_SOCK_OPS_TSTAMP_SND_SW_CB;
[ ... ]
>  	bpf_skops_tx_timestamping(sk, skb, op);
> +	if (static_branch_unlikely(&bpfts_v2_needed_key) && start)
> +		skb_hwtstamps(skb)->start_time = start;

[Severity: High]
Does this code permanently overwrite start_time with hwtstamps when
start_time is initially 0?

The && start condition skips the restoration when start_time was 0.
This leaves the skb_shared_hwtstamps union corrupted with hwtstamps instead
of 0 for start_time on the retransmit queue, causing downstream logic to read
a massive hardware timestamp instead of 0.

Also, could re-evaluating the global static key bpfts_v2_needed_key here
introduce a race condition that skips the restore if the key is disabled
concurrently? A local boolean flag might be needed to track whether
start_time was saved.

>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=6

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path
  2026-09-19 14:37 ` [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] bpf: MPTCP subflows leave sk_start_time uninitialized causing hardware timestamp leak
- [High] net: global pskb_copy performance regression for hardware timestamping
- [Medium] bpf: start_time leaks to receiver as hardware timestamp over loopback
- [Medium] tcp: TFO SYN-data segments bypass sk_start_time assignment
--

commit 1c42878ae739db740438d18e58461996447a2443
Author: Jason Xing <kerneljasonxing@gmail.com>

tcp: propagate the start time onto every skb in the tx path
    
Now that sendmsg() records a start time in sk->sk_start_time within
the socket lock protection, copy that value into the shared hwtstamps
of each skb.

The timing is tcp_skb_entail() runs once per newly allocated skb
in the sendmsg() loop, every segment produced by a single sendmsg()
carries the same start time.

> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index f07abd53ed27f..040f612608476 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -705,6 +705,9 @@ void tcp_skb_entail(struct sock *sk, struct sk_buff *skb)
>  	tcb->tcp_flags = TCPHDR_ACK;
>  	__skb_header_release(skb);
>  	psp_enqueue_set_decrypted(sk, skb);
> +	if (static_branch_unlikely(&bpfts_v2_needed_key) &&
> +	    SK_BPF_CB_FLAG_TEST(sk, SK_BPF_CB_TIMESTAMPING_V2))
> +		skb_hwtstamps(skb)->start_time = sk->sk_start_time;

[Severity: Medium]
Does this leak the internal timestamp to the receiver as a hardware timestamp
over local virtual interfaces?

When sending traffic over loopback or veth, the SKB enters the receive path.
skb_scrub_packet() clears skb->tstamp but not skb_shinfo(skb)->hwtstamps.

Since start_time and hwtstamp share a union, does the receiving socket
erroneously read the sender's software start_time as a hardware RX timestamp?

[Severity: Medium]
Does this code miss assigning sk_start_time for Fast Open segments?

tcp_sendmsg_fastopen() delegates to tcp_connect() to allocate and queue the
SYN segment carrying payload:

net/ipv4/tcp_output.c:tcp_connect() {
    ...
    buff = tcp_stream_alloc_skb(sk, sk->sk_allocation, true);
    ...
    tcp_init_nondata_skb(buff, sk, tp->write_seq, TCPHDR_SYN);
    ...
    tcp_connect_queue_skb(sk, buff);
}

Both tcp_connect() and tcp_send_syn_data() bypass tcp_skb_entail().
Will this result in a loss of BPF timestamping functionality for any
application using MSG_FASTOPEN?

[Severity: High]
Can this corrupt the shared SKB in the write queue for MPTCP sockets?

MPTCP bypasses tcp_sendmsg_locked(), which leaves the subflow socket's
sk_start_time uninitialized (0). When __mptcp_alloc_tx_skb() calls
tcp_skb_entail(), the SKB's start_time is set to 0.

Upon TX completion with a hardware timestamp, the callback in
skb_tstamp_tx_report_bpf_timestamping() reads the 0:

net/core/skbuff.c:skb_tstamp_tx_report_bpf_timestamping() {
    ...
    if (static_branch_unlikely(&bpfts_v2_needed_key) && start)
        skb_hwtstamps(skb)->start_time = start;
}

Since the check fails when start is 0, does this permanently leak the
hardware timestamp into the shared orig_skb in the transmit queue?

[Severity: High]
Does this series introduce a system-wide performance regression for all
hardware TX timestamping users?

When the static key is enabled, all hardware timestamps fall into this path:

net/core/skbuff.c:__skb_tstamp_tx() {
    ...
    if (static_branch_unlikely(&bpfts_v2_needed_key) && hwtstamps)
        skb = pskb_copy(orig_skb, GFP_ATOMIC);
    else
        skb = skb_clone(orig_skb, GFP_ATOMIC);
}

Because the static key is globally enabled when BPF timestamping is used,
will unrelated sockets using normal SO_TIMESTAMPING suffer the high overhead
of copying the entire SKB head via pskb_copy() instead of skb_clone()?

>  	tcp_add_write_queue_tail(sk, skb);
>  	sk_wmem_queued_add(sk, skb->truesize);
>  	sk_mem_charge(sk, skb->truesize);

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=7

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2
  2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: TOCTOU race in sk_bpf_set_get_cb_flags() corrupts global static key
--

commit 1e7a759ab84f855625d2d275769f450b5e750f84
Author: Jason Xing <kerneljasonxing@gmail.com>

net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2

Introduce a per-socket bit SK_BPF_CB_TIMESTAMPING_V2 for BPF Timestamping
V2, that allows a socket opt into recording a "start time" on every
skb it produces or receives.

> diff --git a/net/core/filter.c b/net/core/filter.c
> index 61940e7535523..3486c5c2d1703 100644
> --- a/net/core/filter.c
> +++ b/net/core/filter.c
> @@ -5468,6 +5468,13 @@ static int sk_bpf_set_get_cb_flags(struct sock *sk, char *optval, bool getopt)
>  	if (sk_bpf_cb_flags & ~SK_BPF_CB_MASK)
>  		return -EINVAL;
>  
> +	if ((sk_bpf_cb_flags ^ sk->sk_bpf_cb_flags) & SK_BPF_CB_TIMESTAMPING_V2) {
> +		if (sk_bpf_cb_flags & SK_BPF_CB_TIMESTAMPING_V2)
> +			bpfts_v2_enable();
> +		else
> +			bpfts_v2_disable();
> +	}
> +
>  	sk->sk_bpf_cb_flags = sk_bpf_cb_flags;
>  
>  	return 0;

[Severity: High]
Does this unprotected compare-and-branch allow concurrent threads to trigger
unmatched global enable/disable calls, permanently corrupting the reference
state?

If multiple threads concurrently call bpf_unlocked_sk_setsockopt(SK_BPF_CB_FLAGS)
on a shared socket, they evaluate (sk_bpf_cb_flags ^ sk->sk_bpf_cb_flags)
without locks. Both threads could observe the same initial state and invoke
bpfts_v2_enable() twice, but the socket will only decrement it once during
__sk_destruct().

Could this permanently degrade system-wide performance by leaking the static
key, or underflow it and break the feature for other active users?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=1

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc
  2026-09-19 14:37 ` [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] bpf: bpf_ktime_get_real_ns() causes hard lockup in NMI context
--

commit 5767d1a435f8b04a603eb9c7fef3df289f64be7b
Author: Jason Xing <kerneljasonxing@gmail.com>

bpf: add bpf_ktime_get_real_ns() kfunc

Introduce a new BPF kfunc bpf_ktime_get_real_ns() that returns the
current CLOCK_REALTIME time.

> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c
> index b3cc5c8fc8756..7a0db620a2730 100644
> --- a/kernel/bpf/helpers.c
> +++ b/kernel/bpf/helpers.c
> @@ -2335,6 +2335,11 @@ void bpf_rb_root_free(const struct btf_field *field, void *rb_root,
>  
>  __bpf_kfunc_start_defs();
>  
> +__bpf_kfunc u64 bpf_ktime_get_real_ns(void)
> +{
> +	return ktime_to_ns(ktime_get_real());

[Severity: High]
Can this cause a hard lockup if called from NMI context?

The ktime_get_real() function uses read_seqcount_begin(&tk_core.seq)
internally to read the time. If a BPF program running in NMI context (such as
a perf event or kprobe) calls bpf_ktime_get_real_ns() while the CPU happened
to interrupt a timekeeper update like do_settimeofday64, the seqcount will be
odd. The reader will execute a busy-wait loop indefinitely because the
preempted writer can never complete.

Could this be made NMI-safe by using the latch-based fast timekeeper API
ktime_get_real_fast_ns(), analogous to how bpf_ktime_get_ns() uses
ktime_get_mono_fast_ns()?

> +}
> +
>  /**
>   * bpf_obj_new() - allocate an object described by program BTF
>   * @local_type_id__k: type ID in program BTF

[ ... ]

> @@ -4972,6 +4977,7 @@ BTF_ID_FLAGS(func, bpf_task_work_schedule_resume, KF_IMPLICIT_ARGS)
>  BTF_ID_FLAGS(func, bpf_dynptr_from_file)
>  BTF_ID_FLAGS(func, bpf_dynptr_file_discard, KF_RELEASE)
>  BTF_ID_FLAGS(func, bpf_timer_cancel_async)
> +BTF_ID_FLAGS(func, bpf_ktime_get_real_ns)

[Severity: High]
Does registering this kfunc in the common_btf_ids set (BPF_PROG_TYPE_UNSPEC)
expose it to NMI-context BPF programs, leading to the potential lockup
described above?

>  BTF_KFUNCS_END(common_btf_ids)
>  
>  static const struct btf_kfunc_id_set common_kfunc_set = {

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=2

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2
  2026-09-19 14:37 ` [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2 Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [High] net: BPF start_time breaks SOF_TIMESTAMPING_RAW_HARDWARE and namespace isolation
- [Medium] net: hardware TX timestamps overwrite start_time breaking BPF latency tracking
--

commit a3500c380740df14fa2b49c448a409fd50a64d25
Author: Jason Xing <kerneljasonxing@gmail.com>

net: reuse skb_shared_hwtstamps for BPF Timestamping v2

This commit reuses the hwtstamp field in skb_shared_hwtstamps to store a
software start time for BPF Timestamping V2. A start_time union member is
added that shares storage with hwtstamp.

> diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h
> index 0f4a4f4bafb2a..275c797f6a4b5 100644
> --- a/include/linux/skbuff.h
> +++ b/include/linux/skbuff.h
[ ... ]
> @@ -462,6 +468,7 @@ static inline bool skb_frag_must_loop(struct page *p)
>  struct skb_shared_hwtstamps {
>  	union {
>  		ktime_t	hwtstamp;
> +		ktime_t	start_time;

[Severity: High]
Could unioning start_time and hwtstamp impact SOF_TIMESTAMPING_RAW_HARDWARE
and namespace isolation?

When BPF Timestamping v2 is enabled via bpfts_v2_needed_key,
net_timestamp_check() populates start_time for received packets lacking a
hardware timestamp. Because they share memory, __sock_recv_timestamp() sees
this as a non-zero hwtstamp.

Applications requesting SOF_TIMESTAMPING_RAW_HARDWARE (such as PTP daemons)
will erroneously receive the software-generated start_time as a real hardware
timestamp in tss.ts[2], which could corrupt clock servos.

Furthermore, skb_scrub_packet() clears skb->tstamp to prevent timing leaks
across namespaces, but it does not clear skb_hwtstamps(). Smuggling a
software timestamp in the hardware timestamp union appears to bypass
namespace isolation.

[Severity: Medium]
Will hardware TX timestamps overwrite start_time and affect BPF latency
tracking?

BPF Timestamping V2 intends to measure latency by stamping start_time upon
packet enqueue. When a hardware TX timestamp is received, __skb_tstamp_tx()
calls skb_tstamp_tx_report_bpf_timestamping(), which assigns:

    *skb_hwtstamps(skb) = *hwtstamps;

before invoking the BPF program via bpf_skops_tx_timestamping().

Because start_time shares the same union as hwtstamp, this assignment
overwrites the enqueue start_time with the TX completion time before the BPF
program runs on the BPF_SOCK_OPS_TSTAMP_SND_HW_CB hook.

Could this cause the BPF program to read the hardware timestamp instead of
the original start_time, preventing correct latency calculations?

>  		void *netdev_data;
>  	};
>  };

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=4

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path
  2026-09-19 14:37 ` [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] bpf: __tcp_transmit_skb() drops start_time when copying cloned SKBs
- [Medium] bpf: tcp_mtu_probe() misses start_time propagation for partially consumed SKBs
--

commit 29685480e121f376d52c9ec47c6328d8844a8e0f
Author: Jason Xing <kerneljasonxing@gmail.com>

tcp: handle the start time of each split skb in the tx path
    
Segmentation and collapsing create or reuse skbs with a fresh
shared_info, which would otherwise drop the start time recorded
on the original skb.
    
Handle every TCP path that manipulates the shared hwtstamps:
    
  - tcp_gso_segment(): handle the gso case. Propagate the same
    gso_skb's start time to every resulting segment.
    
  - tcp_fragment/tso_fragment(): handle the tso/recovery.. cases.
    Copy the start time onto the second half of the split so both
    fragments keep it.
    
  - tcp_skb_collapse_tstamp(): handle collapse case. Copy the eaten
    skb's start time onto the survivor.
    
  - __tcp_retransmit_skb(): handle retrans case - the retransmit
    path __tcp_retransmit_skb->__pskb_copy.

> diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
> index 00417a4292225..cb43215e8b3c9 100644
> --- a/net/ipv4/tcp_output.c
> +++ b/net/ipv4/tcp_output.c

[ ... ]

> @@ -3669,6 +3673,7 @@ int __tcp_retransmit_skb(struct sock *sk, struct sk_buff *skb, int segs)
>  			nskb = __pskb_copy(skb, MAX_TCP_HEADER, GFP_ATOMIC);
>  			if (nskb) {
>  				nskb->dev = NULL;
> +				skb_copy_start_time(nskb, skb);
>  				err = tcp_transmit_skb(sk, nskb, 0, GFP_ATOMIC);
>  			} else {
>  				err = -ENOBUFS;

[Severity: Medium]
Does this code miss propagating the start_time when copying cloned SKBs in
__tcp_transmit_skb()?

In __tcp_transmit_skb(), if clone_it is true and the original SKB is
already cloned (e.g., still held by Qdisc), it calls pskb_copy():

net/ipv4/tcp_output.c:__tcp_transmit_skb() {
    ...
		tcp_skb_tsorted_save(oskb) {
			if (unlikely(skb_cloned(oskb)))
				skb = pskb_copy(oskb, gfp_mask);
			else
				skb = skb_clone(oskb, gfp_mask);
		} tcp_skb_tsorted_restore(oskb);
    ...
}

When pskb_copy() is executed, it fails to propagate the original SKB's
start_time to the new transmit SKB. This allocates a new skb_shared_info
with a zeroed start_time, silently dropping timestamps for these packets.

[Severity: Medium]
Could there be another missed start_time propagation for partially
consumed SKBs in tcp_mtu_probe()?

tcp_mtu_probe() allocates a new MTU probe SKB (nskb) with a zeroed
start_time, and then absorbs payload from queued SKBs. If the first
SKB is larger than the probe size, it takes the else branch to
partially consume it:

net/ipv4/tcp_output.c:tcp_mtu_probe() {
    ...
		} else {
			TCP_SKB_CB(nskb)->tcp_flags |= TCP_SKB_CB(skb)->tcp_flags &
						   ~(TCPHDR_FIN|TCPHDR_PSH);
			__pskb_trim_head(skb, copy);
			tcp_set_skb_tso_segs(skb, mss_now);
			TCP_SKB_CB(skb)->seq += copy;
		}
    ...
}

In this partial consumption branch, tcp_eat_one_skb() is not invoked, and
start_time is never explicitly copied, leaving the MTU probe packet with
start_time = 0.

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=9

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path
  2026-09-19 14:37 ` [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path Jason Xing
@ 2026-09-20 14:39   ` sashiko-bot
  0 siblings, 0 replies; 27+ messages in thread
From: sashiko-bot @ 2026-09-20 14:39 UTC (permalink / raw)
  To: Jason Xing; +Cc: bpf

Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Critical] net: start_time and hwtstamp union aliasing causes UAPI breakage
- [High] net: modifying shared skb_shinfo in net_timestamp_check corrupts cloned skbs
--

commit f67f60b2a43b1183b5caab4725cd26a0b5760176
Author: Jason Xing <kerneljasonxing@gmail.com>

net: generate the start time for every skb in the rx path

This commit extends net_timestamp_check() to record an in-kernel start time
in the skb's shared hwtstamps. It enables computing how long an skb has
spent in the stack during the RX path.

> diff --git a/net/core/dev.c b/net/core/dev.c
> index 1da3250f25d2c..e358d28ae6c3f 100644
> --- a/net/core/dev.c
> +++ b/net/core/dev.c
> @@ -2501,11 +2501,17 @@ static inline void net_timestamp_set(struct sk_buff *skb)
>  		skb->tstamp = ktime_get_real();
>  }
>  
> -#define net_timestamp_check(COND, SKB)				\
> -	if (static_branch_unlikely(&netstamp_needed_key)) {	\
> -		if ((COND) && !(SKB)->tstamp)			\
> -			(SKB)->tstamp = ktime_get_real();	\
> -	}							\
> +#define net_timestamp_check(COND, SKB)					\
> +	if (static_branch_unlikely(&netstamp_needed_key) ||		\
> +	    static_branch_unlikely(&bpfts_v2_needed_key)) {		\
> +		if (static_branch_unlikely(&netstamp_needed_key) &&	\
> +		    (COND) && !(SKB)->tstamp)				\
> +			(SKB)->tstamp = ktime_get_real();		\
> +		if (static_branch_unlikely(&bpfts_v2_needed_key) &&	\
> +		    !skb_hwtstamps(SKB)->hwtstamp)			\
> +			skb_hwtstamps(SKB)->start_time =		\

[Severity: High]
Does this code modify shared skb_shinfo without verifying SKB exclusivity?

The skb_hwtstamps(SKB) macro resolves to &skb_shinfo(skb)->hwtstamps, which
is explicitly shared across all clones of an SKB.

In the RX path, such as in netif_rx_internal() where net_timestamp_check()
is called, packets may be clones. For example, TCP transmits a clone while
retaining the original SKB in its retransmit queue, and if sent over a
loopback interface, this clone is directly passed to netif_rx().

Writing directly to shared SKB metadata without calling skb_unshare() or
verifying exclusivity via skb_cloned() can cause data races and silent
corruption across all concurrent users of the cloned SKB.

[Severity: Critical]
Does this code also inadvertently overwrite hardware timestamps delivered to
userspace?

In include/linux/skbuff.h, start_time and hwtstamp are aliases in the same
union:

struct skb_shared_hwtstamps {
	union {
		ktime_t hwtstamp;
		ktime_t start_time;
		void *netdev_data;
	};
};

When a packet without a hardware timestamp comes through, this assignment
writes a software timestamp into start_time, silently overwriting the empty
hwtstamp field with a non-zero value.

Later, when the packet propagates up the stack to a socket with
SOF_TIMESTAMPING_RAW_HARDWARE enabled, __sock_recv_timestamp() checks the
timestamp:

net/socket.c:__sock_recv_timestamp() {
    ...
    if (shhwtstamps->hwtstamp) {
        tss.ts[2] = shhwtstamps->hwtstamp;
    ...
}

Will this cause applications like PTP daemons to falsely receive a software
timestamp as a hardware timestamp, breaking the SO_TIMESTAMPING UAPI?

> +				(SKB)->tstamp ?: ktime_get_real();	\
> +	}								\
>  
>  bool is_skb_forwardable(const struct net_device *dev, const struct sk_buff *skb)
>  {

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260919143732.11772-1-kerneljasonxing@gmail.com?part=8

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
  2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
                   ` (9 preceding siblings ...)
  2026-09-19 18:12 ` [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Alexei Starovoitov
@ 2026-09-21 18:45 ` Stanislav Fomichev
  2026-09-22  1:20   ` Jason Xing
  10 siblings, 1 reply; 27+ messages in thread
From: Stanislav Fomichev @ 2026-09-21 18:45 UTC (permalink / raw)
  To: Jason Xing
  Cc: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu, netdev,
	bpf, Jason Xing

On 09/19, Jason Xing wrote:
> From: Jason Xing <kernelxing@gmail.com>
> 
> Greeting,
> 
> It's BPF Timestamping 2.0 that aims to observe the packet latency
> more efficiently and simply for different protocols. The current
> series is only focused on TCP protocol.
> 
> At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> rough implementation of the feature were exhaustively introduced[1].
> 
> History
> =======
> - In 2009, Patrick Ohly implemented the basic infrastructure
> - In 2014, Willem de Bruijn enhanced the TCP latency observation
> - In 2024, Jason Xing proposed its lightweight BPF version
> Detailed slides from 28 to 32 [1].
> 
> Background
> ==========
> Even though BPF Timetamping 1.0 is comparatively low-overhead,
> transparent, it's still complicated due to a few points inherited
> from the design:
> - Inflexible/fixed reporting phases (qdisc/driver/ack)
>   When we confirm the issue arises from the kernel by using attribution
>   ability of timestamping feature, we need to further minimize the scope
>   until the issue is fixed. That means, we then have to resort to write
>   a few complex BPF progs with the similar functionalities (like skb
>   level tag) which should not happen.
> - Not enough low-overhead
>   Serving the sensitive applications, an always-on latency observation
>   platform should mitigate the self-impact as much as possible. As we
>   can conclude from BPF Timestamping selftests, there are some blocking
>   and time-consuming points like where reading/writing BPF maps happen
>   in the extremely hot paths.
> - Minor flaws
>   There are a few minor flaws inherited from the initial design, like
>   missing tagging the last packet[2][3][4].
> Detailed slides from 33 to 39 [1].
> 
> Motivation
> ==========
> During the process of the large scale deployment over the last few years,
> we eventually realized timestamping feature doesn't support container
> scenario and we need a finer-grained and flexible tracing tool (packet
> basis) after a few rounds of attribution of issues.
> 
> Design
> ======
> - Start time
>   For the specific protocol, we need to accurately set the start time of
>   each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
>   and the driver time as the start point, so that any BPF program is
>   capable of computing the delta between start time and current time.
> - Simplicity
>   Previous BPF program (like selftests) is too complex to implement. The
>   core idea is to make everything as simple as possible. And it should be
>   decoupled from BPF area and previous timestamping feature as much as
>   possible.
> - Flexibility
>   BPF program hooking any function with skb parameter can get the latency
>   value, which means it's no longer bound to the pre-embeded reporting
>   phases (see __skb_tstamp_tx)
> - Efficiency
>   Avoid the previous BPF operations as much as possible. Make sure the
>   feature achieves the lowest performance impact, which means only time
>   operations remain.
> Detailed slides from 40 to 62 [1].
> 
> Implementations
> ===============
> in-kernel
> - Find a suitable place to timestamp for each packet
> - Pick the right start time for TCP
> - Handle the split skb due to various reasons
> BPF prog
> - Hook any functions that carry skb parameter
> - Read out the start time from the skb
> - Generate the current time and then compute the latency
> 
> Discussion?
> ===========
> - Do we need a kfunc to allow users to reset the start time of each skb?
>   What I had in mind is if someone tries to observe the latency between
>   two specific functions (rather than tcp_sendmsg_locked).
> - Current implementation is real hardware timestamp always wins, which
>   means BPF prog possibly gets the hardware time that is not aligned
>   with bpf_ktime_get_real_ns.
> - After the series, do we need to implement the same logic for
>   SYN/FIN/PROBE... As far as I know according to numerous user reports,
>   a small handful of issues came from 3-way handshake.
> - Reusing the slot of hwtstamp might bring potential problems or make the
>   code hard to maintain. Can we add a timestamping specific field in
>   skb to deal with the latency observation?
> - netdev_data conflict in IGC driver. It seems unavoidable to pollute
>   start time when it's enabled. Should V2 feature coexist with hardware
>   timestamping?
> - Should V2 coexist with net timestamping and BPF timestamping? If not,
>   the maintenance should be easier.

After netconf discussion, I was under the impressions that no kernel
changes are needed, so what changed? Is it hard to track start_time
from tcp_sendmsg_locked on the bpf side that we need kernel support?

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
  2026-09-21 18:45 ` Stanislav Fomichev
@ 2026-09-22  1:20   ` Jason Xing
  2026-09-22 20:53     ` Stanislav Fomichev
  0 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-22  1:20 UTC (permalink / raw)
  To: Stanislav Fomichev
  Cc: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu, netdev,
	bpf, Jason Xing

On Tue, Sep 22, 2026 at 2:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
>
> On 09/19, Jason Xing wrote:
> > From: Jason Xing <kernelxing@gmail.com>
> >
> > Greeting,
> >
> > It's BPF Timestamping 2.0 that aims to observe the packet latency
> > more efficiently and simply for different protocols. The current
> > series is only focused on TCP protocol.
> >
> > At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> > rough implementation of the feature were exhaustively introduced[1].
> >
> > History
> > =======
> > - In 2009, Patrick Ohly implemented the basic infrastructure
> > - In 2014, Willem de Bruijn enhanced the TCP latency observation
> > - In 2024, Jason Xing proposed its lightweight BPF version
> > Detailed slides from 28 to 32 [1].
> >
> > Background
> > ==========
> > Even though BPF Timetamping 1.0 is comparatively low-overhead,
> > transparent, it's still complicated due to a few points inherited
> > from the design:
> > - Inflexible/fixed reporting phases (qdisc/driver/ack)
> >   When we confirm the issue arises from the kernel by using attribution
> >   ability of timestamping feature, we need to further minimize the scope
> >   until the issue is fixed. That means, we then have to resort to write
> >   a few complex BPF progs with the similar functionalities (like skb
> >   level tag) which should not happen.
> > - Not enough low-overhead
> >   Serving the sensitive applications, an always-on latency observation
> >   platform should mitigate the self-impact as much as possible. As we
> >   can conclude from BPF Timestamping selftests, there are some blocking
> >   and time-consuming points like where reading/writing BPF maps happen
> >   in the extremely hot paths.
> > - Minor flaws
> >   There are a few minor flaws inherited from the initial design, like
> >   missing tagging the last packet[2][3][4].
> > Detailed slides from 33 to 39 [1].
> >
> > Motivation
> > ==========
> > During the process of the large scale deployment over the last few years,
> > we eventually realized timestamping feature doesn't support container
> > scenario and we need a finer-grained and flexible tracing tool (packet
> > basis) after a few rounds of attribution of issues.
> >
> > Design
> > ======
> > - Start time
> >   For the specific protocol, we need to accurately set the start time of
> >   each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
> >   and the driver time as the start point, so that any BPF program is
> >   capable of computing the delta between start time and current time.
> > - Simplicity
> >   Previous BPF program (like selftests) is too complex to implement. The
> >   core idea is to make everything as simple as possible. And it should be
> >   decoupled from BPF area and previous timestamping feature as much as
> >   possible.
> > - Flexibility
> >   BPF program hooking any function with skb parameter can get the latency
> >   value, which means it's no longer bound to the pre-embeded reporting
> >   phases (see __skb_tstamp_tx)
> > - Efficiency
> >   Avoid the previous BPF operations as much as possible. Make sure the
> >   feature achieves the lowest performance impact, which means only time
> >   operations remain.
> > Detailed slides from 40 to 62 [1].
> >
> > Implementations
> > ===============
> > in-kernel
> > - Find a suitable place to timestamp for each packet
> > - Pick the right start time for TCP
> > - Handle the split skb due to various reasons
> > BPF prog
> > - Hook any functions that carry skb parameter
> > - Read out the start time from the skb
> > - Generate the current time and then compute the latency
> >
> > Discussion?
> > ===========
> > - Do we need a kfunc to allow users to reset the start time of each skb?
> >   What I had in mind is if someone tries to observe the latency between
> >   two specific functions (rather than tcp_sendmsg_locked).
> > - Current implementation is real hardware timestamp always wins, which
> >   means BPF prog possibly gets the hardware time that is not aligned
> >   with bpf_ktime_get_real_ns.
> > - After the series, do we need to implement the same logic for
> >   SYN/FIN/PROBE... As far as I know according to numerous user reports,
> >   a small handful of issues came from 3-way handshake.
> > - Reusing the slot of hwtstamp might bring potential problems or make the
> >   code hard to maintain. Can we add a timestamping specific field in
> >   skb to deal with the latency observation?
> > - netdev_data conflict in IGC driver. It seems unavoidable to pollute
> >   start time when it's enabled. Should V2 feature coexist with hardware
> >   timestamping?
> > - Should V2 coexist with net timestamping and BPF timestamping? If not,
> >   the maintenance should be easier.
>
> After netconf discussion, I was under the impressions that no kernel
> changes are needed, so what changed? Is it hard to track start_time

Ah, you refered to the internal version, right? We wrote a kernel
module implementing similar logic which differently finds/borrows an
unused field of socket to store the start time. It's quite similar to
the series actually.

After deploying it at a small scale in production, I think it's
meaningful to upstream it. But as you noticed, there are remaining
discussion points on which I hope we can share opinions, especially
the future shape.

> from tcp_sendmsg_locked on the bpf side that we need kernel support?

Sure, we need kernel support that why I'm trying to introduce the
sk_start_time to help.
https://lore.kernel.org/all/20260919143732.11772-4-kerneljasonxing@gmail.com/

Thanks,
Jason

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
  2026-09-22  1:20   ` Jason Xing
@ 2026-09-22 20:53     ` Stanislav Fomichev
  2026-09-23  9:40       ` Jason Xing
  0 siblings, 1 reply; 27+ messages in thread
From: Stanislav Fomichev @ 2026-09-22 20:53 UTC (permalink / raw)
  To: Jason Xing
  Cc: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu, netdev,
	bpf

On 09/22, Jason Xing wrote:
> On Tue, Sep 22, 2026 at 2:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> >
> > On 09/19, Jason Xing wrote:
> > > From: Jason Xing <kernelxing@gmail.com>
> > >
> > > Greeting,
> > >
> > > It's BPF Timestamping 2.0 that aims to observe the packet latency
> > > more efficiently and simply for different protocols. The current
> > > series is only focused on TCP protocol.
> > >
> > > At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> > > rough implementation of the feature were exhaustively introduced[1].
> > >
> > > History
> > > =======
> > > - In 2009, Patrick Ohly implemented the basic infrastructure
> > > - In 2014, Willem de Bruijn enhanced the TCP latency observation
> > > - In 2024, Jason Xing proposed its lightweight BPF version
> > > Detailed slides from 28 to 32 [1].
> > >
> > > Background
> > > ==========
> > > Even though BPF Timetamping 1.0 is comparatively low-overhead,
> > > transparent, it's still complicated due to a few points inherited
> > > from the design:
> > > - Inflexible/fixed reporting phases (qdisc/driver/ack)
> > >   When we confirm the issue arises from the kernel by using attribution
> > >   ability of timestamping feature, we need to further minimize the scope
> > >   until the issue is fixed. That means, we then have to resort to write
> > >   a few complex BPF progs with the similar functionalities (like skb
> > >   level tag) which should not happen.
> > > - Not enough low-overhead
> > >   Serving the sensitive applications, an always-on latency observation
> > >   platform should mitigate the self-impact as much as possible. As we
> > >   can conclude from BPF Timestamping selftests, there are some blocking
> > >   and time-consuming points like where reading/writing BPF maps happen
> > >   in the extremely hot paths.
> > > - Minor flaws
> > >   There are a few minor flaws inherited from the initial design, like
> > >   missing tagging the last packet[2][3][4].
> > > Detailed slides from 33 to 39 [1].
> > >
> > > Motivation
> > > ==========
> > > During the process of the large scale deployment over the last few years,
> > > we eventually realized timestamping feature doesn't support container
> > > scenario and we need a finer-grained and flexible tracing tool (packet
> > > basis) after a few rounds of attribution of issues.
> > >
> > > Design
> > > ======
> > > - Start time
> > >   For the specific protocol, we need to accurately set the start time of
> > >   each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
> > >   and the driver time as the start point, so that any BPF program is
> > >   capable of computing the delta between start time and current time.
> > > - Simplicity
> > >   Previous BPF program (like selftests) is too complex to implement. The
> > >   core idea is to make everything as simple as possible. And it should be
> > >   decoupled from BPF area and previous timestamping feature as much as
> > >   possible.
> > > - Flexibility
> > >   BPF program hooking any function with skb parameter can get the latency
> > >   value, which means it's no longer bound to the pre-embeded reporting
> > >   phases (see __skb_tstamp_tx)
> > > - Efficiency
> > >   Avoid the previous BPF operations as much as possible. Make sure the
> > >   feature achieves the lowest performance impact, which means only time
> > >   operations remain.
> > > Detailed slides from 40 to 62 [1].
> > >
> > > Implementations
> > > ===============
> > > in-kernel
> > > - Find a suitable place to timestamp for each packet
> > > - Pick the right start time for TCP
> > > - Handle the split skb due to various reasons
> > > BPF prog
> > > - Hook any functions that carry skb parameter
> > > - Read out the start time from the skb
> > > - Generate the current time and then compute the latency
> > >
> > > Discussion?
> > > ===========
> > > - Do we need a kfunc to allow users to reset the start time of each skb?
> > >   What I had in mind is if someone tries to observe the latency between
> > >   two specific functions (rather than tcp_sendmsg_locked).
> > > - Current implementation is real hardware timestamp always wins, which
> > >   means BPF prog possibly gets the hardware time that is not aligned
> > >   with bpf_ktime_get_real_ns.
> > > - After the series, do we need to implement the same logic for
> > >   SYN/FIN/PROBE... As far as I know according to numerous user reports,
> > >   a small handful of issues came from 3-way handshake.
> > > - Reusing the slot of hwtstamp might bring potential problems or make the
> > >   code hard to maintain. Can we add a timestamping specific field in
> > >   skb to deal with the latency observation?
> > > - netdev_data conflict in IGC driver. It seems unavoidable to pollute
> > >   start time when it's enabled. Should V2 feature coexist with hardware
> > >   timestamping?
> > > - Should V2 coexist with net timestamping and BPF timestamping? If not,
> > >   the maintenance should be easier.
> >
> > After netconf discussion, I was under the impressions that no kernel
> > changes are needed, so what changed? Is it hard to track start_time
> 
> Ah, you refered to the internal version, right? We wrote a kernel
> module implementing similar logic which differently finds/borrows an
> unused field of socket to store the start time. It's quite similar to
> the series actually.
> 
> After deploying it at a small scale in production, I think it's
> meaningful to upstream it. But as you noticed, there are remaining
> discussion points on which I hope we can share opinions, especially
> the future shape.
> 
> > from tcp_sendmsg_locked on the bpf side that we need kernel support?

[..]

> Sure, we need kernel support that why I'm trying to introduce the
> sk_start_time to help.
> https://lore.kernel.org/all/20260919143732.11772-4-kerneljasonxing@gmail.com/

Why can you not do this on the bpf side? There is even now a sendmsg_locked
tracepoint with skb/sk argument. Or is it mostly because you can't
distinguish between cgroups? And looking at your example [1] and don't
see why a tracepoint won't be enough. Or is it too much overhead? In this
case, it needs to have some numbers attached...

1: https://lore.kernel.org/all/CAL+tcoD36zA=TYSzKSNWV0Ypo_HRBRU_Qh6rQtPL51aXR54p7Q@mail.gmail.com/

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
  2026-09-22 20:53     ` Stanislav Fomichev
@ 2026-09-23  9:40       ` Jason Xing
  2026-09-24 16:11         ` Stanislav Fomichev
  0 siblings, 1 reply; 27+ messages in thread
From: Jason Xing @ 2026-09-23  9:40 UTC (permalink / raw)
  To: Stanislav Fomichev
  Cc: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu, netdev,
	bpf

On Wed, Sep 23, 2026 at 4:53 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
>
> On 09/22, Jason Xing wrote:
> > On Tue, Sep 22, 2026 at 2:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> > >
> > > On 09/19, Jason Xing wrote:
> > > > From: Jason Xing <kernelxing@gmail.com>
> > > >
> > > > Greeting,
> > > >
> > > > It's BPF Timestamping 2.0 that aims to observe the packet latency
> > > > more efficiently and simply for different protocols. The current
> > > > series is only focused on TCP protocol.
> > > >
> > > > At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> > > > rough implementation of the feature were exhaustively introduced[1].
> > > >
> > > > History
> > > > =======
> > > > - In 2009, Patrick Ohly implemented the basic infrastructure
> > > > - In 2014, Willem de Bruijn enhanced the TCP latency observation
> > > > - In 2024, Jason Xing proposed its lightweight BPF version
> > > > Detailed slides from 28 to 32 [1].
> > > >
> > > > Background
> > > > ==========
> > > > Even though BPF Timetamping 1.0 is comparatively low-overhead,
> > > > transparent, it's still complicated due to a few points inherited
> > > > from the design:
> > > > - Inflexible/fixed reporting phases (qdisc/driver/ack)
> > > >   When we confirm the issue arises from the kernel by using attribution
> > > >   ability of timestamping feature, we need to further minimize the scope
> > > >   until the issue is fixed. That means, we then have to resort to write
> > > >   a few complex BPF progs with the similar functionalities (like skb
> > > >   level tag) which should not happen.
> > > > - Not enough low-overhead
> > > >   Serving the sensitive applications, an always-on latency observation
> > > >   platform should mitigate the self-impact as much as possible. As we
> > > >   can conclude from BPF Timestamping selftests, there are some blocking
> > > >   and time-consuming points like where reading/writing BPF maps happen
> > > >   in the extremely hot paths.
> > > > - Minor flaws
> > > >   There are a few minor flaws inherited from the initial design, like
> > > >   missing tagging the last packet[2][3][4].
> > > > Detailed slides from 33 to 39 [1].
> > > >
> > > > Motivation
> > > > ==========
> > > > During the process of the large scale deployment over the last few years,
> > > > we eventually realized timestamping feature doesn't support container
> > > > scenario and we need a finer-grained and flexible tracing tool (packet
> > > > basis) after a few rounds of attribution of issues.
> > > >
> > > > Design
> > > > ======
> > > > - Start time
> > > >   For the specific protocol, we need to accurately set the start time of
> > > >   each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
> > > >   and the driver time as the start point, so that any BPF program is
> > > >   capable of computing the delta between start time and current time.
> > > > - Simplicity
> > > >   Previous BPF program (like selftests) is too complex to implement. The
> > > >   core idea is to make everything as simple as possible. And it should be
> > > >   decoupled from BPF area and previous timestamping feature as much as
> > > >   possible.
> > > > - Flexibility
> > > >   BPF program hooking any function with skb parameter can get the latency
> > > >   value, which means it's no longer bound to the pre-embeded reporting
> > > >   phases (see __skb_tstamp_tx)
> > > > - Efficiency
> > > >   Avoid the previous BPF operations as much as possible. Make sure the
> > > >   feature achieves the lowest performance impact, which means only time
> > > >   operations remain.
> > > > Detailed slides from 40 to 62 [1].
> > > >
> > > > Implementations
> > > > ===============
> > > > in-kernel
> > > > - Find a suitable place to timestamp for each packet
> > > > - Pick the right start time for TCP
> > > > - Handle the split skb due to various reasons
> > > > BPF prog
> > > > - Hook any functions that carry skb parameter
> > > > - Read out the start time from the skb
> > > > - Generate the current time and then compute the latency
> > > >
> > > > Discussion?
> > > > ===========
> > > > - Do we need a kfunc to allow users to reset the start time of each skb?
> > > >   What I had in mind is if someone tries to observe the latency between
> > > >   two specific functions (rather than tcp_sendmsg_locked).
> > > > - Current implementation is real hardware timestamp always wins, which
> > > >   means BPF prog possibly gets the hardware time that is not aligned
> > > >   with bpf_ktime_get_real_ns.
> > > > - After the series, do we need to implement the same logic for
> > > >   SYN/FIN/PROBE... As far as I know according to numerous user reports,
> > > >   a small handful of issues came from 3-way handshake.
> > > > - Reusing the slot of hwtstamp might bring potential problems or make the
> > > >   code hard to maintain. Can we add a timestamping specific field in
> > > >   skb to deal with the latency observation?
> > > > - netdev_data conflict in IGC driver. It seems unavoidable to pollute
> > > >   start time when it's enabled. Should V2 feature coexist with hardware
> > > >   timestamping?
> > > > - Should V2 coexist with net timestamping and BPF timestamping? If not,
> > > >   the maintenance should be easier.
> > >
> > > After netconf discussion, I was under the impressions that no kernel
> > > changes are needed, so what changed? Is it hard to track start_time
> >
> > Ah, you refered to the internal version, right? We wrote a kernel
> > module implementing similar logic which differently finds/borrows an
> > unused field of socket to store the start time. It's quite similar to
> > the series actually.
> >
> > After deploying it at a small scale in production, I think it's
> > meaningful to upstream it. But as you noticed, there are remaining
> > discussion points on which I hope we can share opinions, especially
> > the future shape.
> >
> > > from tcp_sendmsg_locked on the bpf side that we need kernel support?
>
> [..]
>
> > Sure, we need kernel support that why I'm trying to introduce the
> > sk_start_time to help.
> > https://lore.kernel.org/all/20260919143732.11772-4-kerneljasonxing@gmail.com/
>
> Why can you not do this on the bpf side? There is even now a sendmsg_locked
> tracepoint with skb/sk argument. Or is it mostly because you can't
> distinguish between cgroups? And looking at your example [1] and don't
> see why a tracepoint won't be enough. Or is it too much overhead? In this
> case, it needs to have some numbers attached...

I understand what you meant. Sure, we can generate the initial time in
the tcp_sendmsg_locked() and try to pass it on to each skb in
skb_entail(), which means 1) we need at least two hooks, which brings
obvious overhead, 2) it's still not that easy to use as I expect it to
be super easy to use/write/deploy, 3) I try to decouple it from BPF
infra or complex use for the convenience. It looks like we now go back
to BPF Timestamping V1.0.

Adding hooks does harm to the performance to those real
latency-sensitive users who were actually yelling at us. There are
some interesting numbers I collected previously (Sorry, I will not be
able to collect more real data because I left that company:( ):
1) The kernel module introduces around 5-15% performance impact in the
real workload. After we removed the hook in tcp_sendmsg_locked, even
though the module didn't have the full function to calculate the
latency, the cost was decreased by ~3-5%.
2) Fentry has ~4% impact on some workloads. It's very similar to what
we use netperf on loopback [1].

The crucial idea behind the feature is that we're trying so hard to
deploy the latency platform 7x24 without any selective sampling. It
now looks like an advanced/21-century tcpdump and it can be
_always-on_. If someone is just looking for one-shot tools, of course
there are a few alternatives that are not good enough though.

[1]: sides 25
https://netdevconf.info/0x19/sessions/talk/the-future-of-so_timestamping.html

Thanks,
Jason

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
  2026-09-23  9:40       ` Jason Xing
@ 2026-09-24 16:11         ` Stanislav Fomichev
  2026-09-30 10:21           ` Jason Xing
  0 siblings, 1 reply; 27+ messages in thread
From: Stanislav Fomichev @ 2026-09-24 16:11 UTC (permalink / raw)
  To: Jason Xing
  Cc: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu, netdev,
	bpf

On 09/23, Jason Xing wrote:
> On Wed, Sep 23, 2026 at 4:53 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> >
> > On 09/22, Jason Xing wrote:
> > > On Tue, Sep 22, 2026 at 2:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> > > >
> > > > On 09/19, Jason Xing wrote:
> > > > > From: Jason Xing <kernelxing@gmail.com>
> > > > >
> > > > > Greeting,
> > > > >
> > > > > It's BPF Timestamping 2.0 that aims to observe the packet latency
> > > > > more efficiently and simply for different protocols. The current
> > > > > series is only focused on TCP protocol.
> > > > >
> > > > > At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> > > > > rough implementation of the feature were exhaustively introduced[1].
> > > > >
> > > > > History
> > > > > =======
> > > > > - In 2009, Patrick Ohly implemented the basic infrastructure
> > > > > - In 2014, Willem de Bruijn enhanced the TCP latency observation
> > > > > - In 2024, Jason Xing proposed its lightweight BPF version
> > > > > Detailed slides from 28 to 32 [1].
> > > > >
> > > > > Background
> > > > > ==========
> > > > > Even though BPF Timetamping 1.0 is comparatively low-overhead,
> > > > > transparent, it's still complicated due to a few points inherited
> > > > > from the design:
> > > > > - Inflexible/fixed reporting phases (qdisc/driver/ack)
> > > > >   When we confirm the issue arises from the kernel by using attribution
> > > > >   ability of timestamping feature, we need to further minimize the scope
> > > > >   until the issue is fixed. That means, we then have to resort to write
> > > > >   a few complex BPF progs with the similar functionalities (like skb
> > > > >   level tag) which should not happen.
> > > > > - Not enough low-overhead
> > > > >   Serving the sensitive applications, an always-on latency observation
> > > > >   platform should mitigate the self-impact as much as possible. As we
> > > > >   can conclude from BPF Timestamping selftests, there are some blocking
> > > > >   and time-consuming points like where reading/writing BPF maps happen
> > > > >   in the extremely hot paths.
> > > > > - Minor flaws
> > > > >   There are a few minor flaws inherited from the initial design, like
> > > > >   missing tagging the last packet[2][3][4].
> > > > > Detailed slides from 33 to 39 [1].
> > > > >
> > > > > Motivation
> > > > > ==========
> > > > > During the process of the large scale deployment over the last few years,
> > > > > we eventually realized timestamping feature doesn't support container
> > > > > scenario and we need a finer-grained and flexible tracing tool (packet
> > > > > basis) after a few rounds of attribution of issues.
> > > > >
> > > > > Design
> > > > > ======
> > > > > - Start time
> > > > >   For the specific protocol, we need to accurately set the start time of
> > > > >   each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
> > > > >   and the driver time as the start point, so that any BPF program is
> > > > >   capable of computing the delta between start time and current time.
> > > > > - Simplicity
> > > > >   Previous BPF program (like selftests) is too complex to implement. The
> > > > >   core idea is to make everything as simple as possible. And it should be
> > > > >   decoupled from BPF area and previous timestamping feature as much as
> > > > >   possible.
> > > > > - Flexibility
> > > > >   BPF program hooking any function with skb parameter can get the latency
> > > > >   value, which means it's no longer bound to the pre-embeded reporting
> > > > >   phases (see __skb_tstamp_tx)
> > > > > - Efficiency
> > > > >   Avoid the previous BPF operations as much as possible. Make sure the
> > > > >   feature achieves the lowest performance impact, which means only time
> > > > >   operations remain.
> > > > > Detailed slides from 40 to 62 [1].
> > > > >
> > > > > Implementations
> > > > > ===============
> > > > > in-kernel
> > > > > - Find a suitable place to timestamp for each packet
> > > > > - Pick the right start time for TCP
> > > > > - Handle the split skb due to various reasons
> > > > > BPF prog
> > > > > - Hook any functions that carry skb parameter
> > > > > - Read out the start time from the skb
> > > > > - Generate the current time and then compute the latency
> > > > >
> > > > > Discussion?
> > > > > ===========
> > > > > - Do we need a kfunc to allow users to reset the start time of each skb?
> > > > >   What I had in mind is if someone tries to observe the latency between
> > > > >   two specific functions (rather than tcp_sendmsg_locked).
> > > > > - Current implementation is real hardware timestamp always wins, which
> > > > >   means BPF prog possibly gets the hardware time that is not aligned
> > > > >   with bpf_ktime_get_real_ns.
> > > > > - After the series, do we need to implement the same logic for
> > > > >   SYN/FIN/PROBE... As far as I know according to numerous user reports,
> > > > >   a small handful of issues came from 3-way handshake.
> > > > > - Reusing the slot of hwtstamp might bring potential problems or make the
> > > > >   code hard to maintain. Can we add a timestamping specific field in
> > > > >   skb to deal with the latency observation?
> > > > > - netdev_data conflict in IGC driver. It seems unavoidable to pollute
> > > > >   start time when it's enabled. Should V2 feature coexist with hardware
> > > > >   timestamping?
> > > > > - Should V2 coexist with net timestamping and BPF timestamping? If not,
> > > > >   the maintenance should be easier.
> > > >
> > > > After netconf discussion, I was under the impressions that no kernel
> > > > changes are needed, so what changed? Is it hard to track start_time
> > >
> > > Ah, you refered to the internal version, right? We wrote a kernel
> > > module implementing similar logic which differently finds/borrows an
> > > unused field of socket to store the start time. It's quite similar to
> > > the series actually.
> > >
> > > After deploying it at a small scale in production, I think it's
> > > meaningful to upstream it. But as you noticed, there are remaining
> > > discussion points on which I hope we can share opinions, especially
> > > the future shape.
> > >
> > > > from tcp_sendmsg_locked on the bpf side that we need kernel support?
> >
> > [..]
> >
> > > Sure, we need kernel support that why I'm trying to introduce the
> > > sk_start_time to help.
> > > https://lore.kernel.org/all/20260919143732.11772-4-kerneljasonxing@gmail.com/
> >
> > Why can you not do this on the bpf side? There is even now a sendmsg_locked
> > tracepoint with skb/sk argument. Or is it mostly because you can't
> > distinguish between cgroups? And looking at your example [1] and don't
> > see why a tracepoint won't be enough. Or is it too much overhead? In this
> > case, it needs to have some numbers attached...
> 
> I understand what you meant. Sure, we can generate the initial time in
> the tcp_sendmsg_locked() and try to pass it on to each skb in
> skb_entail(), which means 1) we need at least two hooks, which brings
> obvious overhead, 2) it's still not that easy to use as I expect it to
> be super easy to use/write/deploy, 3) I try to decouple it from BPF
> infra or complex use for the convenience. It looks like we now go back
> to BPF Timestamping V1.0.
> 
> Adding hooks does harm to the performance to those real
> latency-sensitive users who were actually yelling at us. There are
> some interesting numbers I collected previously (Sorry, I will not be
> able to collect more real data because I left that company:( ):
> 1) The kernel module introduces around 5-15% performance impact in the
> real workload. After we removed the hook in tcp_sendmsg_locked, even
> though the module didn't have the full function to calculate the
> latency, the cost was decreased by ~3-5%.
> 2) Fentry has ~4% impact on some workloads. It's very similar to what
> we use netperf on loopback [1].
> 
> The crucial idea behind the feature is that we're trying so hard to
> deploy the latency platform 7x24 without any selective sampling. It
> now looks like an advanced/21-century tcpdump and it can be
> _always-on_. If someone is just looking for one-shot tools, of course
> there are a few alternatives that are not good enough though.

IMO the feature as posted looks very tailored to a specific/narrow
use case. You save the sendmsg time in the socket and then use
it during skb allocation (with a few quirks here and there to account
for fragmentation/tso/cloning/etc).

For the very least, If you're looking to turn it into a non-rfc submission, 
I'd add numbers: tracepoint based implementation vs this kernel
accelerated path. And then we can discuss how much slower the tracepoints
are and whether you're using the proper ones... (and have an actual selftest
example as part of the series)

^ permalink raw reply	[flat|nested] 27+ messages in thread

* Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
  2026-09-24 16:11         ` Stanislav Fomichev
@ 2026-09-30 10:21           ` Jason Xing
  0 siblings, 0 replies; 27+ messages in thread
From: Jason Xing @ 2026-09-30 10:21 UTC (permalink / raw)
  To: Stanislav Fomichev
  Cc: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu, netdev,
	bpf

On Fri, Sep 25, 2026 at 12:11 AM Stanislav Fomichev
<sdf.kernel@gmail.com> wrote:
>
> On 09/23, Jason Xing wrote:
> > On Wed, Sep 23, 2026 at 4:53 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> > >
> > > On 09/22, Jason Xing wrote:
> > > > On Tue, Sep 22, 2026 at 2:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> > > > >
> > > > > On 09/19, Jason Xing wrote:
> > > > > > From: Jason Xing <kernelxing@gmail.com>
> > > > > >
> > > > > > Greeting,
> > > > > >
> > > > > > It's BPF Timestamping 2.0 that aims to observe the packet latency
> > > > > > more efficiently and simply for different protocols. The current
> > > > > > series is only focused on TCP protocol.
> > > > > >
> > > > > > At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> > > > > > rough implementation of the feature were exhaustively introduced[1].
> > > > > >
> > > > > > History
> > > > > > =======
> > > > > > - In 2009, Patrick Ohly implemented the basic infrastructure
> > > > > > - In 2014, Willem de Bruijn enhanced the TCP latency observation
> > > > > > - In 2024, Jason Xing proposed its lightweight BPF version
> > > > > > Detailed slides from 28 to 32 [1].
> > > > > >
> > > > > > Background
> > > > > > ==========
> > > > > > Even though BPF Timetamping 1.0 is comparatively low-overhead,
> > > > > > transparent, it's still complicated due to a few points inherited
> > > > > > from the design:
> > > > > > - Inflexible/fixed reporting phases (qdisc/driver/ack)
> > > > > >   When we confirm the issue arises from the kernel by using attribution
> > > > > >   ability of timestamping feature, we need to further minimize the scope
> > > > > >   until the issue is fixed. That means, we then have to resort to write
> > > > > >   a few complex BPF progs with the similar functionalities (like skb
> > > > > >   level tag) which should not happen.
> > > > > > - Not enough low-overhead
> > > > > >   Serving the sensitive applications, an always-on latency observation
> > > > > >   platform should mitigate the self-impact as much as possible. As we
> > > > > >   can conclude from BPF Timestamping selftests, there are some blocking
> > > > > >   and time-consuming points like where reading/writing BPF maps happen
> > > > > >   in the extremely hot paths.
> > > > > > - Minor flaws
> > > > > >   There are a few minor flaws inherited from the initial design, like
> > > > > >   missing tagging the last packet[2][3][4].
> > > > > > Detailed slides from 33 to 39 [1].
> > > > > >
> > > > > > Motivation
> > > > > > ==========
> > > > > > During the process of the large scale deployment over the last few years,
> > > > > > we eventually realized timestamping feature doesn't support container
> > > > > > scenario and we need a finer-grained and flexible tracing tool (packet
> > > > > > basis) after a few rounds of attribution of issues.
> > > > > >
> > > > > > Design
> > > > > > ======
> > > > > > - Start time
> > > > > >   For the specific protocol, we need to accurately set the start time of
> > > > > >   each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
> > > > > >   and the driver time as the start point, so that any BPF program is
> > > > > >   capable of computing the delta between start time and current time.
> > > > > > - Simplicity
> > > > > >   Previous BPF program (like selftests) is too complex to implement. The
> > > > > >   core idea is to make everything as simple as possible. And it should be
> > > > > >   decoupled from BPF area and previous timestamping feature as much as
> > > > > >   possible.
> > > > > > - Flexibility
> > > > > >   BPF program hooking any function with skb parameter can get the latency
> > > > > >   value, which means it's no longer bound to the pre-embeded reporting
> > > > > >   phases (see __skb_tstamp_tx)
> > > > > > - Efficiency
> > > > > >   Avoid the previous BPF operations as much as possible. Make sure the
> > > > > >   feature achieves the lowest performance impact, which means only time
> > > > > >   operations remain.
> > > > > > Detailed slides from 40 to 62 [1].
> > > > > >
> > > > > > Implementations
> > > > > > ===============
> > > > > > in-kernel
> > > > > > - Find a suitable place to timestamp for each packet
> > > > > > - Pick the right start time for TCP
> > > > > > - Handle the split skb due to various reasons
> > > > > > BPF prog
> > > > > > - Hook any functions that carry skb parameter
> > > > > > - Read out the start time from the skb
> > > > > > - Generate the current time and then compute the latency
> > > > > >
> > > > > > Discussion?
> > > > > > ===========
> > > > > > - Do we need a kfunc to allow users to reset the start time of each skb?
> > > > > >   What I had in mind is if someone tries to observe the latency between
> > > > > >   two specific functions (rather than tcp_sendmsg_locked).
> > > > > > - Current implementation is real hardware timestamp always wins, which
> > > > > >   means BPF prog possibly gets the hardware time that is not aligned
> > > > > >   with bpf_ktime_get_real_ns.
> > > > > > - After the series, do we need to implement the same logic for
> > > > > >   SYN/FIN/PROBE... As far as I know according to numerous user reports,
> > > > > >   a small handful of issues came from 3-way handshake.
> > > > > > - Reusing the slot of hwtstamp might bring potential problems or make the
> > > > > >   code hard to maintain. Can we add a timestamping specific field in
> > > > > >   skb to deal with the latency observation?
> > > > > > - netdev_data conflict in IGC driver. It seems unavoidable to pollute
> > > > > >   start time when it's enabled. Should V2 feature coexist with hardware
> > > > > >   timestamping?
> > > > > > - Should V2 coexist with net timestamping and BPF timestamping? If not,
> > > > > >   the maintenance should be easier.
> > > > >
> > > > > After netconf discussion, I was under the impressions that no kernel
> > > > > changes are needed, so what changed? Is it hard to track start_time
> > > >
> > > > Ah, you refered to the internal version, right? We wrote a kernel
> > > > module implementing similar logic which differently finds/borrows an
> > > > unused field of socket to store the start time. It's quite similar to
> > > > the series actually.
> > > >
> > > > After deploying it at a small scale in production, I think it's
> > > > meaningful to upstream it. But as you noticed, there are remaining
> > > > discussion points on which I hope we can share opinions, especially
> > > > the future shape.
> > > >
> > > > > from tcp_sendmsg_locked on the bpf side that we need kernel support?
> > >
> > > [..]
> > >
> > > > Sure, we need kernel support that why I'm trying to introduce the
> > > > sk_start_time to help.
> > > > https://lore.kernel.org/all/20260919143732.11772-4-kerneljasonxing@gmail.com/
> > >
> > > Why can you not do this on the bpf side? There is even now a sendmsg_locked
> > > tracepoint with skb/sk argument. Or is it mostly because you can't
> > > distinguish between cgroups? And looking at your example [1] and don't
> > > see why a tracepoint won't be enough. Or is it too much overhead? In this
> > > case, it needs to have some numbers attached...
> >
> > I understand what you meant. Sure, we can generate the initial time in
> > the tcp_sendmsg_locked() and try to pass it on to each skb in
> > skb_entail(), which means 1) we need at least two hooks, which brings
> > obvious overhead, 2) it's still not that easy to use as I expect it to
> > be super easy to use/write/deploy, 3) I try to decouple it from BPF
> > infra or complex use for the convenience. It looks like we now go back
> > to BPF Timestamping V1.0.
> >
> > Adding hooks does harm to the performance to those real
> > latency-sensitive users who were actually yelling at us. There are
> > some interesting numbers I collected previously (Sorry, I will not be
> > able to collect more real data because I left that company:( ):
> > 1) The kernel module introduces around 5-15% performance impact in the
> > real workload. After we removed the hook in tcp_sendmsg_locked, even
> > though the module didn't have the full function to calculate the
> > latency, the cost was decreased by ~3-5%.
> > 2) Fentry has ~4% impact on some workloads. It's very similar to what
> > we use netperf on loopback [1].
> >
> > The crucial idea behind the feature is that we're trying so hard to
> > deploy the latency platform 7x24 without any selective sampling. It
> > now looks like an advanced/21-century tcpdump and it can be
> > _always-on_. If someone is just looking for one-shot tools, of course
> > there are a few alternatives that are not good enough though.
>
> IMO the feature as posted looks very tailored to a specific/narrow

Quite opposite.  Perhaps you missed reading the key part about the
"lack of container scenario support." And the fact is that there was a
wifi guy talking to me that he also wanted this to land and support
the wifi.

The new arch serves all kinds of complex scenarios that BPF
timestamping v1 doesn't have the ability to support at all :)

And BPF timestamping v2 is per-packet-based, which makes everything
different: allows admins/users to investigate stall details if they
want.

I believe the full explaination/theory/motivation... have been added
in the commit message. Please if you could reread all of them
(including the slides), I would appreciate it :)

> use case. You save the sendmsg time in the socket and then use
> it during skb allocation (with a few quirks here and there to account
> for fragmentation/tso/cloning/etc).
>
> For the very least, If you're looking to turn it into a non-rfc submission,
> I'd add numbers: tracepoint based implementation vs this kernel
> accelerated path. And then we can discuss how much slower the tracepoints
> are and whether you're using the proper ones... (and have an actual selftest
> example as part of the series)

Sorry, as said, this is only one part of the new arch.

Sure thing, as you suggested I'm going to add a few numbers with BPF.
Also a selftest!

Thanks,
Jason

^ permalink raw reply	[flat|nested] 27+ messages in thread

end of thread, other threads:[~2026-09-30 10:21 UTC | newest]

Thread overview: 27+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by " Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 18:12 ` [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Alexei Starovoitov
2026-09-20  0:41   ` Jason Xing
2026-09-21 18:45 ` Stanislav Fomichev
2026-09-22  1:20   ` Jason Xing
2026-09-22 20:53     ` Stanislav Fomichev
2026-09-23  9:40       ` Jason Xing
2026-09-24 16:11         ` Stanislav Fomichev
2026-09-30 10:21           ` Jason Xing

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox