From: Stanislav Fomichev <sdf.kernel@gmail.com>
To: Jason Xing <kerneljasonxing@gmail.com>
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
pabeni@redhat.com, horms@kernel.org, willemb@google.com,
kuniyu@google.com, netdev@vger.kernel.org, bpf@vger.kernel.org
Subject: Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
Date: Tue, 22 Sep 2026 13:53:04 -0700 [thread overview]
Message-ID: <arLpc9BBv_W6nw8c@devvm7509.cco0.facebook.com> (raw)
In-Reply-To: <CAL+tcoDossUOJCVwZ-6U519L_bNMYngnd82e1_oQB-czA6rL3A@mail.gmail.com>
On 09/22, Jason Xing wrote:
> On Tue, Sep 22, 2026 at 2:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> >
> > On 09/19, Jason Xing wrote:
> > > From: Jason Xing <kernelxing@gmail.com>
> > >
> > > Greeting,
> > >
> > > It's BPF Timestamping 2.0 that aims to observe the packet latency
> > > more efficiently and simply for different protocols. The current
> > > series is only focused on TCP protocol.
> > >
> > > At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> > > rough implementation of the feature were exhaustively introduced[1].
> > >
> > > History
> > > =======
> > > - In 2009, Patrick Ohly implemented the basic infrastructure
> > > - In 2014, Willem de Bruijn enhanced the TCP latency observation
> > > - In 2024, Jason Xing proposed its lightweight BPF version
> > > Detailed slides from 28 to 32 [1].
> > >
> > > Background
> > > ==========
> > > Even though BPF Timetamping 1.0 is comparatively low-overhead,
> > > transparent, it's still complicated due to a few points inherited
> > > from the design:
> > > - Inflexible/fixed reporting phases (qdisc/driver/ack)
> > > When we confirm the issue arises from the kernel by using attribution
> > > ability of timestamping feature, we need to further minimize the scope
> > > until the issue is fixed. That means, we then have to resort to write
> > > a few complex BPF progs with the similar functionalities (like skb
> > > level tag) which should not happen.
> > > - Not enough low-overhead
> > > Serving the sensitive applications, an always-on latency observation
> > > platform should mitigate the self-impact as much as possible. As we
> > > can conclude from BPF Timestamping selftests, there are some blocking
> > > and time-consuming points like where reading/writing BPF maps happen
> > > in the extremely hot paths.
> > > - Minor flaws
> > > There are a few minor flaws inherited from the initial design, like
> > > missing tagging the last packet[2][3][4].
> > > Detailed slides from 33 to 39 [1].
> > >
> > > Motivation
> > > ==========
> > > During the process of the large scale deployment over the last few years,
> > > we eventually realized timestamping feature doesn't support container
> > > scenario and we need a finer-grained and flexible tracing tool (packet
> > > basis) after a few rounds of attribution of issues.
> > >
> > > Design
> > > ======
> > > - Start time
> > > For the specific protocol, we need to accurately set the start time of
> > > each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
> > > and the driver time as the start point, so that any BPF program is
> > > capable of computing the delta between start time and current time.
> > > - Simplicity
> > > Previous BPF program (like selftests) is too complex to implement. The
> > > core idea is to make everything as simple as possible. And it should be
> > > decoupled from BPF area and previous timestamping feature as much as
> > > possible.
> > > - Flexibility
> > > BPF program hooking any function with skb parameter can get the latency
> > > value, which means it's no longer bound to the pre-embeded reporting
> > > phases (see __skb_tstamp_tx)
> > > - Efficiency
> > > Avoid the previous BPF operations as much as possible. Make sure the
> > > feature achieves the lowest performance impact, which means only time
> > > operations remain.
> > > Detailed slides from 40 to 62 [1].
> > >
> > > Implementations
> > > ===============
> > > in-kernel
> > > - Find a suitable place to timestamp for each packet
> > > - Pick the right start time for TCP
> > > - Handle the split skb due to various reasons
> > > BPF prog
> > > - Hook any functions that carry skb parameter
> > > - Read out the start time from the skb
> > > - Generate the current time and then compute the latency
> > >
> > > Discussion?
> > > ===========
> > > - Do we need a kfunc to allow users to reset the start time of each skb?
> > > What I had in mind is if someone tries to observe the latency between
> > > two specific functions (rather than tcp_sendmsg_locked).
> > > - Current implementation is real hardware timestamp always wins, which
> > > means BPF prog possibly gets the hardware time that is not aligned
> > > with bpf_ktime_get_real_ns.
> > > - After the series, do we need to implement the same logic for
> > > SYN/FIN/PROBE... As far as I know according to numerous user reports,
> > > a small handful of issues came from 3-way handshake.
> > > - Reusing the slot of hwtstamp might bring potential problems or make the
> > > code hard to maintain. Can we add a timestamping specific field in
> > > skb to deal with the latency observation?
> > > - netdev_data conflict in IGC driver. It seems unavoidable to pollute
> > > start time when it's enabled. Should V2 feature coexist with hardware
> > > timestamping?
> > > - Should V2 coexist with net timestamping and BPF timestamping? If not,
> > > the maintenance should be easier.
> >
> > After netconf discussion, I was under the impressions that no kernel
> > changes are needed, so what changed? Is it hard to track start_time
>
> Ah, you refered to the internal version, right? We wrote a kernel
> module implementing similar logic which differently finds/borrows an
> unused field of socket to store the start time. It's quite similar to
> the series actually.
>
> After deploying it at a small scale in production, I think it's
> meaningful to upstream it. But as you noticed, there are remaining
> discussion points on which I hope we can share opinions, especially
> the future shape.
>
> > from tcp_sendmsg_locked on the bpf side that we need kernel support?
[..]
> Sure, we need kernel support that why I'm trying to introduce the
> sk_start_time to help.
> https://lore.kernel.org/all/20260919143732.11772-4-kerneljasonxing@gmail.com/
Why can you not do this on the bpf side? There is even now a sendmsg_locked
tracepoint with skb/sk argument. Or is it mostly because you can't
distinguish between cgroups? And looking at your example [1] and don't
see why a tracepoint won't be enough. Or is it too much overhead? In this
case, it needs to have some numbers attached...
1: https://lore.kernel.org/all/CAL+tcoD36zA=TYSzKSNWV0Ypo_HRBRU_Qh6rQtPL51aXR54p7Q@mail.gmail.com/
next prev parent reply other threads:[~2026-09-22 20:53 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2 Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by " Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path Jason Xing
2026-09-19 18:12 ` [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Alexei Starovoitov
2026-09-20 0:41 ` Jason Xing
2026-09-21 18:45 ` Stanislav Fomichev
2026-09-22 1:20 ` Jason Xing
2026-09-22 20:53 ` Stanislav Fomichev [this message]
2026-09-23 9:40 ` Jason Xing
2026-09-24 16:11 ` Stanislav Fomichev
2026-09-30 10:21 ` Jason Xing
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=arLpc9BBv_W6nw8c@devvm7509.cco0.facebook.com \
--to=sdf.kernel@gmail.com \
--cc=bpf@vger.kernel.org \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=kerneljasonxing@gmail.com \
--cc=kuba@kernel.org \
--cc=kuniyu@google.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=willemb@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox