From: Simon Schippers <simon.schippers@tu-dortmund.de>
To: "Jesper Dangaard Brouer" <hawk@kernel.org>,
netdev@vger.kernel.org,
"Jonas Köppeler" <j.koeppeler@tu-berlin.de>
Cc: kernel-team@cloudflare.com,
"David S. Miller" <davem@davemloft.net>,
"Eric Dumazet" <edumazet@google.com>,
"Jakub Kicinski" <kuba@kernel.org>,
"Paolo Abeni" <pabeni@redhat.com>,
"Simon Horman" <horms@kernel.org>,
"Chris Arges" <carges@cloudflare.com>,
"Mike Freemon" <mfreemon@cloudflare.com>,
"Toke Høiland-Jørgensen" <toke@toke.dk>,
"Breno Leitao" <leitao@debian.org>,
"Alexei Starovoitov" <ast@kernel.org>,
"Daniel Borkmann" <daniel@iogearbox.net>,
"John Fastabend" <john.fastabend@gmail.com>,
"Stanislav Fomichev" <sdf@fomichev.me>,
bpf@vger.kernel.org
Subject: Re: [PATCH net-next v7 0/5] veth: add Byte Queue Limits (BQL) support
Date: Wed, 16 Sep 2026 10:14:12 +0200 [thread overview]
Message-ID: <f4509145-1148-4ec1-b1aa-9bf944de7263@tu-dortmund.de> (raw)
In-Reply-To: <a184d9c0-aa3a-4d00-9a0e-e7a1d9d344b0@kernel.org>
On 9/16/26 08:45, Jesper Dangaard Brouer wrote:
>
>
> On 9/15/26 16:30, Simon Schippers wrote:
>> On 8/10/26 15:25, Simon Schippers wrote:
>>> On 6/12/26 10:35, hawk@kernel.org wrote:
>>>> From: Jesper Dangaard Brouer <hawk@kernel.org>
>>>>
>>>> This series adds BQL (Byte Queue Limits) to the veth driver, reducing
>>>> latency by dynamically limiting in-flight packets in the ptr_ring and
>>>> moving buffering into the qdisc where AQM algorithms can act on it.
>>>
>>> Hi :)
>>>
>>> I worked on my implementation of DQL coalescing that lives in
>>> dynamic_queue_limits.{h,c} and wanted to share it so we can consider it
>>> for the next cycle. This is because in the next cycle I would
>>> like to add BQL support for tun/tap as well, in addition to veth.
>>> Is that fine for you?
>>>
>>> I think it is in good shape. It uses the same logic as the v7, but
>>> every new field fits inside the existing dql struct, and drivers only
>>> need to call the usual netdev_tx_sent_queue() and
>>> netdev_tx_completed_queue() to use it. Patch 3, 5 and 6 are the same
>>> as before, only 1, 2 and 4 are new. Benchmarks looked fine for me.
>>>
>>> There should be no regressions for other DQL/BQL users. I paid close
>>> attention not to break the dql cache lines or other logic.
>>> coal_usecs is now configurable per queue via sysfs and also via ethtool
>>> as usual.
>>>
>>> While working on this, I found a missing barrier in v7:
>>> There was no smp_rmb() pairing the smp_wmb() in __ptr_ring_produce()
>>> before dql_completed() reads dql->num_queued. This happens to be safe
>>> on x86, but on other platforms the read of dql->num_queued could be
>>> reordered before __ptr_ring_consume(), triggering a BUG_ON() in
>>> dql_completed(). Fixed by adding the missing smp_rmb() in veth_xdp_rcv()
>>> before completing.
>>>
>>> Would love to hear your thoughts on the implementation!
>>>
>>> Thanks,
>>> Simon
>>
>> Hi! Any thoughts on this? Do you want to continue this series?
>
> Appreciate getting poked :-)
>
> I will not have time to work on this until after October 5th.
Noted, no problem.
>
> If you Simon have time, feel free to submit a V8 patchset with the
> barrier fix mentioned above. I should have cycles to review and ACK
> (except between 26 sep to Oct 4).
I would then also move the coalescing logic to
dynamic_queue_limits.{h,c}, ok?
I am convinced that it is the right decision, because it:
- saves us the new struct veth_bql_state, everything fits nicely into
the dql struct in the existing cachelines
- allows other software interfaces like tun/tap to use the same logic,
which is my goal :-) others could be e.g. wireguard and ifb
- allows to change coal_usecs per queue
>
> We are still interested in getting this merged. Notice the bug fix from
> Jonas 60db47f02bfa ("veth: fix queue index used to wake the peer txq in
> veth_poll"). We are running XDP on our veth production interfaces, so
> we didn't notice this. Still, we are currently waiting for this fix to
> get fully rolled out, before proceeding with the BQL variant.
Yes, I saw that. Probably did not happen to me because I was always
using single queue..
Thanks.
prev parent reply other threads:[~2026-09-16 8:14 UTC|newest]
Thread overview: 18+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-06-12 8:35 [PATCH net-next v7 0/5] veth: add Byte Queue Limits (BQL) support hawk
2026-06-12 8:35 ` [PATCH net-next v7 1/5] net: add dev->bql flag to allow BQL sysfs for IFF_NO_QUEUE devices hawk
2026-06-12 8:35 ` [PATCH net-next v7 2/5] veth: implement Byte Queue Limits (BQL) for latency reduction hawk
2026-06-12 8:35 ` [PATCH net-next v7 3/5] veth: add tx_timeout watchdog as BQL safety net hawk
2026-06-12 8:35 ` [PATCH net-next v7 4/5] net: sched: add timeout count to NETDEV WATCHDOG message hawk
2026-06-12 8:35 ` [PATCH net-next v7 5/5] veth: time-based BQL completion coalescing via ethtool tx-usecs hawk
2026-06-13 14:14 ` Simon Schippers
2026-06-30 14:00 ` Jonas Köppeler
2026-06-30 19:07 ` Simon Schippers
2026-07-09 10:03 ` Jonas Köppeler
2026-06-12 14:10 ` [PATCH net-next v7 0/5] veth: add Byte Queue Limits (BQL) support Simon Schippers
2026-06-12 17:21 ` Jonas Köppeler
2026-06-13 13:57 ` Simon Schippers
2026-06-16 1:53 ` Jakub Kicinski
2026-08-10 13:25 ` Simon Schippers
2026-09-15 14:30 ` Simon Schippers
2026-09-16 6:45 ` Jesper Dangaard Brouer
2026-09-16 8:14 ` Simon Schippers [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=f4509145-1148-4ec1-b1aa-9bf944de7263@tu-dortmund.de \
--to=simon.schippers@tu-dortmund.de \
--cc=ast@kernel.org \
--cc=bpf@vger.kernel.org \
--cc=carges@cloudflare.com \
--cc=daniel@iogearbox.net \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=hawk@kernel.org \
--cc=horms@kernel.org \
--cc=j.koeppeler@tu-berlin.de \
--cc=john.fastabend@gmail.com \
--cc=kernel-team@cloudflare.com \
--cc=kuba@kernel.org \
--cc=leitao@debian.org \
--cc=mfreemon@cloudflare.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=sdf@fomichev.me \
--cc=toke@toke.dk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox