BPF List
 help / color / mirror / Atom feed
* [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
@ 2026-09-19 14:37 Jason Xing
  2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
                   ` (10 more replies)
  0 siblings, 11 replies; 27+ messages in thread
From: Jason Xing @ 2026-09-19 14:37 UTC (permalink / raw)
  To: davem, edumazet, kuba, pabeni, horms, willemb, kuniyu
  Cc: netdev, bpf, Jason Xing

From: Jason Xing <kernelxing@gmail.com>

Greeting,

It's BPF Timestamping 2.0 that aims to observe the packet latency
more efficiently and simply for different protocols. The current
series is only focused on TCP protocol.

At Netdev 0x1a/Netconf 2026, the history, background, motivation and
rough implementation of the feature were exhaustively introduced[1].

History
=======
- In 2009, Patrick Ohly implemented the basic infrastructure
- In 2014, Willem de Bruijn enhanced the TCP latency observation
- In 2024, Jason Xing proposed its lightweight BPF version
Detailed slides from 28 to 32 [1].

Background
==========
Even though BPF Timetamping 1.0 is comparatively low-overhead,
transparent, it's still complicated due to a few points inherited
from the design:
- Inflexible/fixed reporting phases (qdisc/driver/ack)
  When we confirm the issue arises from the kernel by using attribution
  ability of timestamping feature, we need to further minimize the scope
  until the issue is fixed. That means, we then have to resort to write
  a few complex BPF progs with the similar functionalities (like skb
  level tag) which should not happen.
- Not enough low-overhead
  Serving the sensitive applications, an always-on latency observation
  platform should mitigate the self-impact as much as possible. As we
  can conclude from BPF Timestamping selftests, there are some blocking
  and time-consuming points like where reading/writing BPF maps happen
  in the extremely hot paths.
- Minor flaws
  There are a few minor flaws inherited from the initial design, like
  missing tagging the last packet[2][3][4].
Detailed slides from 33 to 39 [1].

Motivation
==========
During the process of the large scale deployment over the last few years,
we eventually realized timestamping feature doesn't support container
scenario and we need a finer-grained and flexible tracing tool (packet
basis) after a few rounds of attribution of issues.

Design
======
- Start time
  For the specific protocol, we need to accurately set the start time of
  each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
  and the driver time as the start point, so that any BPF program is
  capable of computing the delta between start time and current time.
- Simplicity
  Previous BPF program (like selftests) is too complex to implement. The
  core idea is to make everything as simple as possible. And it should be
  decoupled from BPF area and previous timestamping feature as much as
  possible.
- Flexibility
  BPF program hooking any function with skb parameter can get the latency
  value, which means it's no longer bound to the pre-embeded reporting
  phases (see __skb_tstamp_tx)
- Efficiency
  Avoid the previous BPF operations as much as possible. Make sure the
  feature achieves the lowest performance impact, which means only time
  operations remain.
Detailed slides from 40 to 62 [1].

Implementations
===============
in-kernel
- Find a suitable place to timestamp for each packet
- Pick the right start time for TCP
- Handle the split skb due to various reasons
BPF prog
- Hook any functions that carry skb parameter
- Read out the start time from the skb
- Generate the current time and then compute the latency

Discussion?
===========
- Do we need a kfunc to allow users to reset the start time of each skb?
  What I had in mind is if someone tries to observe the latency between
  two specific functions (rather than tcp_sendmsg_locked).
- Current implementation is real hardware timestamp always wins, which
  means BPF prog possibly gets the hardware time that is not aligned
  with bpf_ktime_get_real_ns.
- After the series, do we need to implement the same logic for
  SYN/FIN/PROBE... As far as I know according to numerous user reports,
  a small handful of issues came from 3-way handshake.
- Reusing the slot of hwtstamp might bring potential problems or make the
  code hard to maintain. Can we add a timestamping specific field in
  skb to deal with the latency observation?
- netdev_data conflict in IGC driver. It seems unavoidable to pollute
  start time when it's enabled. Should V2 feature coexist with hardware
  timestamping?
- Should V2 coexist with net timestamping and BPF timestamping? If not,
  the maintenance should be easier.

Any suggestions are greatly welcome!

After we set how to use it from the perspective of users, I would add
a corresponding selftest for this.

[1]: https://netdevconf.info/0x1A/sessions/bof/network-observability-bof.html
[2]: https://lore.kernel.org/all/20260404150452.83904-1-kerneljasonxing@gmail.com/
[3]: commit 838eb9687691 ("tcp: tcp_tx_timestamp() must look at the rtx queue")
[4]: https://lore.kernel.org/all/20260915214450.2882680-1-dw@davidwei.uk/


Jason Xing (9):
  net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2
  bpf: add bpf_ktime_get_real_ns() kfunc
  tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2
  net: reuse skb_shared_hwtstamps for BPF Timestamping v2
  net-timestamp: use pskb_copy to avoid polluting the orig skb's start
    time
  bpf-timestamping: restore skb hwtstamp if it is used by start time
  tcp: propagate the start time onto every skb in the tx path
  net: generate the start time for every skb in the rx path
  tcp: handle the start time of each split skb in the tx path

 include/linux/skbuff.h         | 17 +++++++++
 include/net/sock.h             |  2 +
 include/uapi/linux/bpf.h       |  5 ++-
 kernel/bpf/helpers.c           |  6 +++
 net/core/dev.c                 | 69 +++++++++++++++++++++++++++++++---
 net/core/filter.c              |  7 ++++
 net/core/skbuff.c              | 10 ++++-
 net/core/sock.c                |  5 +++
 net/ipv4/tcp.c                 |  7 ++++
 net/ipv4/tcp_offload.c         |  3 ++
 net/ipv4/tcp_output.c          |  5 +++
 tools/include/uapi/linux/bpf.h |  5 ++-
 12 files changed, 131 insertions(+), 10 deletions(-)

-- 
2.43.7


^ permalink raw reply	[flat|nested] 27+ messages in thread

end of thread, other threads:[~2026-09-30 10:21 UTC | newest]

Thread overview: 27+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by " Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 18:12 ` [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Alexei Starovoitov
2026-09-20  0:41   ` Jason Xing
2026-09-21 18:45 ` Stanislav Fomichev
2026-09-22  1:20   ` Jason Xing
2026-09-22 20:53     ` Stanislav Fomichev
2026-09-23  9:40       ` Jason Xing
2026-09-24 16:11         ` Stanislav Fomichev
2026-09-30 10:21           ` Jason Xing

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox