BPF List
 help / color / mirror / Atom feed
From: Stanislav Fomichev <sdf.kernel@gmail.com>
To: Jason Xing <kerneljasonxing@gmail.com>
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	 pabeni@redhat.com, horms@kernel.org, willemb@google.com,
	kuniyu@google.com,  netdev@vger.kernel.org, bpf@vger.kernel.org
Subject: Re: [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP
Date: Thu, 24 Sep 2026 09:11:04 -0700	[thread overview]
Message-ID: <arVLfFuu02eS2Txs@devvm7509.cco0.facebook.com> (raw)
In-Reply-To: <CAL+tcoCE1MmEQKZRX2ct4zrWxZiSj53ztVuRWU_o1+9AmO7B_g@mail.gmail.com>

On 09/23, Jason Xing wrote:
> On Wed, Sep 23, 2026 at 4:53 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> >
> > On 09/22, Jason Xing wrote:
> > > On Tue, Sep 22, 2026 at 2:45 AM Stanislav Fomichev <sdf.kernel@gmail.com> wrote:
> > > >
> > > > On 09/19, Jason Xing wrote:
> > > > > From: Jason Xing <kernelxing@gmail.com>
> > > > >
> > > > > Greeting,
> > > > >
> > > > > It's BPF Timestamping 2.0 that aims to observe the packet latency
> > > > > more efficiently and simply for different protocols. The current
> > > > > series is only focused on TCP protocol.
> > > > >
> > > > > At Netdev 0x1a/Netconf 2026, the history, background, motivation and
> > > > > rough implementation of the feature were exhaustively introduced[1].
> > > > >
> > > > > History
> > > > > =======
> > > > > - In 2009, Patrick Ohly implemented the basic infrastructure
> > > > > - In 2014, Willem de Bruijn enhanced the TCP latency observation
> > > > > - In 2024, Jason Xing proposed its lightweight BPF version
> > > > > Detailed slides from 28 to 32 [1].
> > > > >
> > > > > Background
> > > > > ==========
> > > > > Even though BPF Timetamping 1.0 is comparatively low-overhead,
> > > > > transparent, it's still complicated due to a few points inherited
> > > > > from the design:
> > > > > - Inflexible/fixed reporting phases (qdisc/driver/ack)
> > > > >   When we confirm the issue arises from the kernel by using attribution
> > > > >   ability of timestamping feature, we need to further minimize the scope
> > > > >   until the issue is fixed. That means, we then have to resort to write
> > > > >   a few complex BPF progs with the similar functionalities (like skb
> > > > >   level tag) which should not happen.
> > > > > - Not enough low-overhead
> > > > >   Serving the sensitive applications, an always-on latency observation
> > > > >   platform should mitigate the self-impact as much as possible. As we
> > > > >   can conclude from BPF Timestamping selftests, there are some blocking
> > > > >   and time-consuming points like where reading/writing BPF maps happen
> > > > >   in the extremely hot paths.
> > > > > - Minor flaws
> > > > >   There are a few minor flaws inherited from the initial design, like
> > > > >   missing tagging the last packet[2][3][4].
> > > > > Detailed slides from 33 to 39 [1].
> > > > >
> > > > > Motivation
> > > > > ==========
> > > > > During the process of the large scale deployment over the last few years,
> > > > > we eventually realized timestamping feature doesn't support container
> > > > > scenario and we need a finer-grained and flexible tracing tool (packet
> > > > > basis) after a few rounds of attribution of issues.
> > > > >
> > > > > Design
> > > > > ======
> > > > > - Start time
> > > > >   For the specific protocol, we need to accurately set the start time of
> > > > >   each packet first. For TCP, we chose the entry of tcp_sendmsg_locked
> > > > >   and the driver time as the start point, so that any BPF program is
> > > > >   capable of computing the delta between start time and current time.
> > > > > - Simplicity
> > > > >   Previous BPF program (like selftests) is too complex to implement. The
> > > > >   core idea is to make everything as simple as possible. And it should be
> > > > >   decoupled from BPF area and previous timestamping feature as much as
> > > > >   possible.
> > > > > - Flexibility
> > > > >   BPF program hooking any function with skb parameter can get the latency
> > > > >   value, which means it's no longer bound to the pre-embeded reporting
> > > > >   phases (see __skb_tstamp_tx)
> > > > > - Efficiency
> > > > >   Avoid the previous BPF operations as much as possible. Make sure the
> > > > >   feature achieves the lowest performance impact, which means only time
> > > > >   operations remain.
> > > > > Detailed slides from 40 to 62 [1].
> > > > >
> > > > > Implementations
> > > > > ===============
> > > > > in-kernel
> > > > > - Find a suitable place to timestamp for each packet
> > > > > - Pick the right start time for TCP
> > > > > - Handle the split skb due to various reasons
> > > > > BPF prog
> > > > > - Hook any functions that carry skb parameter
> > > > > - Read out the start time from the skb
> > > > > - Generate the current time and then compute the latency
> > > > >
> > > > > Discussion?
> > > > > ===========
> > > > > - Do we need a kfunc to allow users to reset the start time of each skb?
> > > > >   What I had in mind is if someone tries to observe the latency between
> > > > >   two specific functions (rather than tcp_sendmsg_locked).
> > > > > - Current implementation is real hardware timestamp always wins, which
> > > > >   means BPF prog possibly gets the hardware time that is not aligned
> > > > >   with bpf_ktime_get_real_ns.
> > > > > - After the series, do we need to implement the same logic for
> > > > >   SYN/FIN/PROBE... As far as I know according to numerous user reports,
> > > > >   a small handful of issues came from 3-way handshake.
> > > > > - Reusing the slot of hwtstamp might bring potential problems or make the
> > > > >   code hard to maintain. Can we add a timestamping specific field in
> > > > >   skb to deal with the latency observation?
> > > > > - netdev_data conflict in IGC driver. It seems unavoidable to pollute
> > > > >   start time when it's enabled. Should V2 feature coexist with hardware
> > > > >   timestamping?
> > > > > - Should V2 coexist with net timestamping and BPF timestamping? If not,
> > > > >   the maintenance should be easier.
> > > >
> > > > After netconf discussion, I was under the impressions that no kernel
> > > > changes are needed, so what changed? Is it hard to track start_time
> > >
> > > Ah, you refered to the internal version, right? We wrote a kernel
> > > module implementing similar logic which differently finds/borrows an
> > > unused field of socket to store the start time. It's quite similar to
> > > the series actually.
> > >
> > > After deploying it at a small scale in production, I think it's
> > > meaningful to upstream it. But as you noticed, there are remaining
> > > discussion points on which I hope we can share opinions, especially
> > > the future shape.
> > >
> > > > from tcp_sendmsg_locked on the bpf side that we need kernel support?
> >
> > [..]
> >
> > > Sure, we need kernel support that why I'm trying to introduce the
> > > sk_start_time to help.
> > > https://lore.kernel.org/all/20260919143732.11772-4-kerneljasonxing@gmail.com/
> >
> > Why can you not do this on the bpf side? There is even now a sendmsg_locked
> > tracepoint with skb/sk argument. Or is it mostly because you can't
> > distinguish between cgroups? And looking at your example [1] and don't
> > see why a tracepoint won't be enough. Or is it too much overhead? In this
> > case, it needs to have some numbers attached...
> 
> I understand what you meant. Sure, we can generate the initial time in
> the tcp_sendmsg_locked() and try to pass it on to each skb in
> skb_entail(), which means 1) we need at least two hooks, which brings
> obvious overhead, 2) it's still not that easy to use as I expect it to
> be super easy to use/write/deploy, 3) I try to decouple it from BPF
> infra or complex use for the convenience. It looks like we now go back
> to BPF Timestamping V1.0.
> 
> Adding hooks does harm to the performance to those real
> latency-sensitive users who were actually yelling at us. There are
> some interesting numbers I collected previously (Sorry, I will not be
> able to collect more real data because I left that company:( ):
> 1) The kernel module introduces around 5-15% performance impact in the
> real workload. After we removed the hook in tcp_sendmsg_locked, even
> though the module didn't have the full function to calculate the
> latency, the cost was decreased by ~3-5%.
> 2) Fentry has ~4% impact on some workloads. It's very similar to what
> we use netperf on loopback [1].
> 
> The crucial idea behind the feature is that we're trying so hard to
> deploy the latency platform 7x24 without any selective sampling. It
> now looks like an advanced/21-century tcpdump and it can be
> _always-on_. If someone is just looking for one-shot tools, of course
> there are a few alternatives that are not good enough though.

IMO the feature as posted looks very tailored to a specific/narrow
use case. You save the sendmsg time in the socket and then use
it during skb allocation (with a few quirks here and there to account
for fragmentation/tso/cloning/etc).

For the very least, If you're looking to turn it into a non-rfc submission, 
I'd add numbers: tracepoint based implementation vs this kernel
accelerated path. And then we can discuss how much slower the tracepoints
are and whether you're using the proper ones... (and have an actual selftest
example as part of the series)

  reply	other threads:[~2026-09-24 16:11 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-19 14:37 [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Jason Xing
2026-09-19 14:37 ` [PATCH RFC net-next 1/9] net: add bpf_setsockopt for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 2/9] bpf: add bpf_ktime_get_real_ns() kfunc Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 3/9] tcp: record a start time in the tx path for SK_BPF_CB_TIMESTAMPING_V2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 4/9] net: reuse skb_shared_hwtstamps for BPF Timestamping v2 Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 5/9] net-timestamp: use pskb_copy to avoid polluting the orig skb's start time Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 6/9] bpf-timestamping: restore skb hwtstamp if it is used by " Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 7/9] tcp: propagate the start time onto every skb in the tx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 8/9] net: generate the start time for every skb in the rx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 14:37 ` [PATCH RFC net-next 9/9] tcp: handle the start time of each split skb in the tx path Jason Xing
2026-09-20 14:39   ` sashiko-bot
2026-09-19 18:12 ` [PATCH RFC net-next 0/9] net: BPF Timestamping 2.0 for TCP Alexei Starovoitov
2026-09-20  0:41   ` Jason Xing
2026-09-21 18:45 ` Stanislav Fomichev
2026-09-22  1:20   ` Jason Xing
2026-09-22 20:53     ` Stanislav Fomichev
2026-09-23  9:40       ` Jason Xing
2026-09-24 16:11         ` Stanislav Fomichev [this message]
2026-09-30 10:21           ` Jason Xing

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arVLfFuu02eS2Txs@devvm7509.cco0.facebook.com \
    --to=sdf.kernel@gmail.com \
    --cc=bpf@vger.kernel.org \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kerneljasonxing@gmail.com \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=willemb@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox