* [PATCH net-next] net: inline eth_type_trans() fast path
@ 2026-09-30 17:29 Eric Dumazet
2026-10-01 15:57 ` Alexander Lobakin
2026-10-02 18:52 ` Eric Dumazet
0 siblings, 2 replies; 6+ messages in thread
From: Eric Dumazet @ 2026-09-30 17:29 UTC (permalink / raw)
To: David S . Miller, Jakub Kicinski, Paolo Abeni
Cc: Simon Horman, netdev, edumazet, Eric Dumazet
eth_type_trans() is called once per received packet by most
Ethernet drivers, and by core helpers (napi_gro_frags(),
xdp_build_skb_from_*(), veth, tun, tunnels, loopback...).
With CONFIG_MITIGATION_RETHUNK / SRSO, the call and return
are not free anymore.
Add eth_type_trans_inline(), handling the common case inline:
unicast frame sent to dev->dev_addr, with a real ethertype,
on a device which is not a DSA conduit. All other frames
(multicast, broadcast, otherhost, 802.2, runts, DSA)
are handled by the out-of-line eth_type_trans().
eth_type_trans() is now a macro calling eth_type_trans_inline(),
so that all existing callers get the fast path. The out-of-line
version remains exported, and can be called with
(eth_type_trans)(skb, dev). bpf_prog_test_run_skb() uses this
to keep tools/testing/selftests/bpf fentry/fexit tests working.
The inline part is about 25 instructions on x86_64, and the whole
function would be almost twice as big when fully inlined.
Signed-off-by: Eric Dumazet <edumazet@kernel.org>
---
include/linux/etherdevice.h | 43 +++++++++++++++++++++++++++++++++++++
net/bpf/test_run.c | 3 ++-
net/ethernet/eth.c | 8 +++++++
3 files changed, 53 insertions(+), 1 deletion(-)
diff --git a/include/linux/etherdevice.h b/include/linux/etherdevice.h
index df8f88f63a7063fbd1df5248d2fc02c859a7bc74..a45c9bcca7dd61e0d696be5a6f2e7a6203d548d0 100644
--- a/include/linux/etherdevice.h
+++ b/include/linux/etherdevice.h
@@ -645,6 +645,49 @@ static inline struct ethhdr *eth_skb_pull_mac(struct sk_buff *skb)
return eth;
}
+static inline bool eth_dev_maybe_dsa(const struct net_device *dev)
+{
+#if IS_ENABLED(CONFIG_NET_DSA)
+ return !!dev->dsa_ptr;
+#else
+ return false;
+#endif
+}
+
+/**
+ * eth_type_trans_inline - determine the packet's protocol ID.
+ * @skb: received socket data
+ * @dev: receiving network device
+ *
+ * Inline fast path of eth_type_trans(), handling unicast frames
+ * sent to @dev address, with a real ethertype.
+ * Other frames are handled by the out-of-line eth_type_trans().
+ *
+ * Return: the packet protocol ID, in network byte order.
+ */
+static __always_inline __be16 eth_type_trans_inline(struct sk_buff *skb,
+ struct net_device *dev)
+{
+ const struct ethhdr *eth = (const struct ethhdr *)skb->data;
+
+ if (likely(skb->len >= ETH_HLEN &&
+ ether_addr_equal_64bits(eth->h_dest, dev->dev_addr) &&
+ eth_proto_is_802_3(eth->h_proto) &&
+ !eth_dev_maybe_dsa(dev))) {
+ skb->dev = dev;
+ skb_reset_mac_header(skb);
+ __skb_pull(skb, ETH_HLEN);
+ /* Like eth_skb_pkt_type(), leave skb->pkt_type unchanged. */
+ return eth->h_proto;
+ }
+ return (eth_type_trans)(skb, dev);
+}
+
+/* eth_type_trans() is called for every received packet by most drivers.
+ * Use (eth_type_trans)(skb, dev) to force a call to the out-of-line version.
+ */
+#define eth_type_trans(skb, dev) eth_type_trans_inline(skb, dev)
+
/**
* eth_skb_pad - Pad buffer to minimum number of octets for Ethernet frame
* @skb: Buffer to pad
diff --git a/net/bpf/test_run.c b/net/bpf/test_run.c
index 513354e928cb58f837fdc93a77f2e896de1ef916..f3eed8c0a767b7c141117f30c0c2c12dc207d9a9 100644
--- a/net/bpf/test_run.c
+++ b/net/bpf/test_run.c
@@ -1165,7 +1165,8 @@ int bpf_prog_test_run_skb(struct bpf_prog *prog, const union bpf_attr *kattr,
goto out;
}
}
- skb->protocol = eth_type_trans(skb, dev);
+ /* Keep calling the out-of-line version, for fentry/fexit selftests. */
+ skb->protocol = (eth_type_trans)(skb, dev);
skb_reset_network_header(skb);
switch (skb->protocol) {
diff --git a/net/ethernet/eth.c b/net/ethernet/eth.c
index d9faadbe9b6c86a746cace6d7a7cfffdb84e4519..476c39fb6a80f4e913bd3631bf1783473bface34 100644
--- a/net/ethernet/eth.c
+++ b/net/ethernet/eth.c
@@ -143,6 +143,9 @@ u32 eth_get_headlen(const struct net_device *dev, const void *data, u32 len)
}
EXPORT_SYMBOL(eth_get_headlen);
+/* Define the out-of-line version, not the eth_type_trans_inline() wrapper. */
+#undef eth_type_trans
+
/**
* eth_type_trans - determine the packet's protocol ID.
* @skb: received socket data
@@ -151,6 +154,11 @@ EXPORT_SYMBOL(eth_get_headlen);
* The rule here is that we
* assume 802.3 if the type field is short enough to be a length.
* This is normal practice and works for any 'now in use' protocol.
+ *
+ * Most callers use the eth_type_trans_inline() fast path,
+ * which calls this function for less common frames.
+ *
+ * Return: the packet protocol ID, in network byte order.
*/
__be16 eth_type_trans(struct sk_buff *skb, struct net_device *dev)
{
--
2.53.0
^ permalink raw reply related [flat|nested] 6+ messages in thread* Re: [PATCH net-next] net: inline eth_type_trans() fast path
2026-09-30 17:29 [PATCH net-next] net: inline eth_type_trans() fast path Eric Dumazet
@ 2026-10-01 15:57 ` Alexander Lobakin
2026-10-01 17:10 ` Eric Dumazet
2026-10-02 18:52 ` Eric Dumazet
1 sibling, 1 reply; 6+ messages in thread
From: Alexander Lobakin @ 2026-10-01 15:57 UTC (permalink / raw)
To: Eric Dumazet
Cc: David S . Miller, Jakub Kicinski, Paolo Abeni, Simon Horman,
netdev, edumazet
From: Edumazet@kernel.org <edumazet@kernel.org>
Date: Wed, 30 Sep 2026 19:29:57 +0200
> eth_type_trans() is called once per received packet by most
> Ethernet drivers, and by core helpers (napi_gro_frags(),
> xdp_build_skb_from_*(), veth, tun, tunnels, loopback...).
>
> With CONFIG_MITIGATION_RETHUNK / SRSO, the call and return
> are not free anymore.
>
> Add eth_type_trans_inline(), handling the common case inline:
> unicast frame sent to dev->dev_addr, with a real ethertype,
> on a device which is not a DSA conduit. All other frames
> (multicast, broadcast, otherhost, 802.2, runts, DSA)
> are handled by the out-of-line eth_type_trans().
>
> eth_type_trans() is now a macro calling eth_type_trans_inline(),
> so that all existing callers get the fast path. The out-of-line
> version remains exported, and can be called with
> (eth_type_trans)(skb, dev). bpf_prog_test_run_skb() uses this
Shouldn't it get renamed to e.g. eth_type_trans_slow() to avoid this
confusion?
> to keep tools/testing/selftests/bpf fentry/fexit tests working.
>
> The inline part is about 25 instructions on x86_64, and the whole
Nice!
> function would be almost twice as big when fully inlined.
>
> Signed-off-by: Eric Dumazet <edumazet@kernel.org>
> ---
> include/linux/etherdevice.h | 43 +++++++++++++++++++++++++++++++++++++
> net/bpf/test_run.c | 3 ++-
> net/ethernet/eth.c | 8 +++++++
> 3 files changed, 53 insertions(+), 1 deletion(-)
Thanks,
Olek
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net-next] net: inline eth_type_trans() fast path
2026-10-01 15:57 ` Alexander Lobakin
@ 2026-10-01 17:10 ` Eric Dumazet
2026-10-02 15:26 ` Stanislav Fomichev
2026-10-02 16:54 ` Alexander Lobakin
0 siblings, 2 replies; 6+ messages in thread
From: Eric Dumazet @ 2026-10-01 17:10 UTC (permalink / raw)
To: Alexander Lobakin
Cc: David S . Miller, Jakub Kicinski, Paolo Abeni, Simon Horman,
netdev, edumazet
On Thu, Oct 1, 2026 at 5:58 PM Alexander Lobakin
<aleksander.lobakin@intel.com> wrote:
>
> From: Edumazet@kernel.org <edumazet@kernel.org>
> Date: Wed, 30 Sep 2026 19:29:57 +0200
>
> > eth_type_trans() is called once per received packet by most
> > Ethernet drivers, and by core helpers (napi_gro_frags(),
> > xdp_build_skb_from_*(), veth, tun, tunnels, loopback...).
> >
> > With CONFIG_MITIGATION_RETHUNK / SRSO, the call and return
> > are not free anymore.
> >
> > Add eth_type_trans_inline(), handling the common case inline:
> > unicast frame sent to dev->dev_addr, with a real ethertype,
> > on a device which is not a DSA conduit. All other frames
> > (multicast, broadcast, otherhost, 802.2, runts, DSA)
> > are handled by the out-of-line eth_type_trans().
> >
> > eth_type_trans() is now a macro calling eth_type_trans_inline(),
> > so that all existing callers get the fast path. The out-of-line
> > version remains exported, and can be called with
> > (eth_type_trans)(skb, dev). bpf_prog_test_run_skb() uses this
>
> Shouldn't it get renamed to e.g. eth_type_trans_slow() to avoid this
> confusion?
This was my initial idea, but this would make the patch a bit more
invasive and touch bpf selftests, something like:
Let me know if I should send a V2 and CC bpf maintainers, thanks!
diff --git a/net/bpf/test_run.c b/net/bpf/test_run.c
index 513354e928cb58f837fdc93a77f2e896de1ef916..8f4e9dd18001a4c32455fefd74fe38db270756d1
100644
--- a/net/bpf/test_run.c
+++ b/net/bpf/test_run.c
@@ -1165,7 +1165,8 @@ int bpf_prog_test_run_skb(struct bpf_prog *prog,
const union bpf_attr *kattr,
goto out;
}
}
- skb->protocol = eth_type_trans(skb, dev);
+ /* Always call the out-of-line version, for fentry/fexit selftests. */
+ skb->protocol = eth_type_trans_slow(skb, dev);
skb_reset_network_header(skb);
switch (skb->protocol) {
diff --git a/tools/testing/selftests/bpf/progs/core_kern.c
b/tools/testing/selftests/bpf/progs/core_kern.c
index 004f2acef2eb0b0d80184423a3e8ba399feee1f3..fca26af98a7b92898a4adfde19f63b342f29db74
100644
--- a/tools/testing/selftests/bpf/progs/core_kern.c
+++ b/tools/testing/selftests/bpf/progs/core_kern.c
@@ -47,14 +47,14 @@ int BPF_PROG(tp_xdp_devmap_xmit_multi, const
struct net_device
return randmap(from_dev->ifindex, from_dev);
}
-SEC("fentry/eth_type_trans")
+SEC("fentry/eth_type_trans_slow")
int BPF_PROG(fentry_eth_type_trans, struct sk_buff *skb,
struct net_device *dev, unsigned short protocol)
{
return randmap(dev->ifindex + skb->len, dev);
}
-SEC("fexit/eth_type_trans")
+SEC("fexit/eth_type_trans_slow")
int BPF_PROG(fexit_eth_type_trans, struct sk_buff *skb,
struct net_device *dev, unsigned short protocol)
{
diff --git a/tools/testing/selftests/bpf/progs/kfree_skb.c
b/tools/testing/selftests/bpf/progs/kfree_skb.c
index 7236da72ce8055f9a0165d3b4334a04d17d110b8..678c15c151944b44e276cf1280c77cb436ff7b10
100644
--- a/tools/testing/selftests/bpf/progs/kfree_skb.c
+++ b/tools/testing/selftests/bpf/progs/kfree_skb.c
@@ -114,7 +114,7 @@ struct {
bool fexit_test_ok;
} result = {};
-SEC("fentry/eth_type_trans")
+SEC("fentry/eth_type_trans_slow")
int BPF_PROG(fentry_eth_type_trans, struct sk_buff *skb, struct
net_device *dev,
unsigned short protocol)
{
@@ -132,7 +132,7 @@ int BPF_PROG(fentry_eth_type_trans, struct sk_buff
*skb, struct net_device *dev,
return 0;
}
-SEC("fexit/eth_type_trans")
+SEC("fexit/eth_type_trans_slow")
int BPF_PROG(fexit_eth_type_trans, struct sk_buff *skb, struct net_device *dev,
unsigned short protocol)
{
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH net-next] net: inline eth_type_trans() fast path
2026-10-01 17:10 ` Eric Dumazet
@ 2026-10-02 15:26 ` Stanislav Fomichev
2026-10-02 16:54 ` Alexander Lobakin
1 sibling, 0 replies; 6+ messages in thread
From: Stanislav Fomichev @ 2026-10-02 15:26 UTC (permalink / raw)
To: Eric Dumazet
Cc: Alexander Lobakin, David S . Miller, Jakub Kicinski, Paolo Abeni,
Simon Horman, netdev, edumazet
On 10/01, Eric Dumazet wrote:
> On Thu, Oct 1, 2026 at 5:58 PM Alexander Lobakin
> <aleksander.lobakin@intel.com> wrote:
> >
> > From: Edumazet@kernel.org <edumazet@kernel.org>
> > Date: Wed, 30 Sep 2026 19:29:57 +0200
> >
> > > eth_type_trans() is called once per received packet by most
> > > Ethernet drivers, and by core helpers (napi_gro_frags(),
> > > xdp_build_skb_from_*(), veth, tun, tunnels, loopback...).
> > >
> > > With CONFIG_MITIGATION_RETHUNK / SRSO, the call and return
> > > are not free anymore.
> > >
> > > Add eth_type_trans_inline(), handling the common case inline:
> > > unicast frame sent to dev->dev_addr, with a real ethertype,
> > > on a device which is not a DSA conduit. All other frames
> > > (multicast, broadcast, otherhost, 802.2, runts, DSA)
> > > are handled by the out-of-line eth_type_trans().
> > >
> > > eth_type_trans() is now a macro calling eth_type_trans_inline(),
> > > so that all existing callers get the fast path. The out-of-line
> > > version remains exported, and can be called with
> > > (eth_type_trans)(skb, dev). bpf_prog_test_run_skb() uses this
> >
> > Shouldn't it get renamed to e.g. eth_type_trans_slow() to avoid this
> > confusion?
>
> This was my initial idea, but this would make the patch a bit more
> invasive and touch bpf selftests, something like:
>
> Let me know if I should send a V2 and CC bpf maintainers, thanks!
I prefer this _slow version as well, less magic.
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net-next] net: inline eth_type_trans() fast path
2026-10-01 17:10 ` Eric Dumazet
2026-10-02 15:26 ` Stanislav Fomichev
@ 2026-10-02 16:54 ` Alexander Lobakin
1 sibling, 0 replies; 6+ messages in thread
From: Alexander Lobakin @ 2026-10-02 16:54 UTC (permalink / raw)
To: Eric Dumazet, bpf
Cc: David S . Miller, Jakub Kicinski, Paolo Abeni, Simon Horman,
netdev, edumazet
From: Edumazet@kernel.org <edumazet@kernel.org>
Date: Thu, 1 Oct 2026 19:10:46 +0200
> On Thu, Oct 1, 2026 at 5:58 PM Alexander Lobakin
> <aleksander.lobakin@intel.com> wrote:
>>
>> From: Edumazet@kernel.org <edumazet@kernel.org>
>> Date: Wed, 30 Sep 2026 19:29:57 +0200
>>
>>> eth_type_trans() is called once per received packet by most
>>> Ethernet drivers, and by core helpers (napi_gro_frags(),
>>> xdp_build_skb_from_*(), veth, tun, tunnels, loopback...).
>>>
>>> With CONFIG_MITIGATION_RETHUNK / SRSO, the call and return
>>> are not free anymore.
>>>
>>> Add eth_type_trans_inline(), handling the common case inline:
>>> unicast frame sent to dev->dev_addr, with a real ethertype,
>>> on a device which is not a DSA conduit. All other frames
>>> (multicast, broadcast, otherhost, 802.2, runts, DSA)
>>> are handled by the out-of-line eth_type_trans().
>>>
>>> eth_type_trans() is now a macro calling eth_type_trans_inline(),
>>> so that all existing callers get the fast path. The out-of-line
>>> version remains exported, and can be called with
>>> (eth_type_trans)(skb, dev). bpf_prog_test_run_skb() uses this
>>
>> Shouldn't it get renamed to e.g. eth_type_trans_slow() to avoid this
>> confusion?
>
> This was my initial idea, but this would make the patch a bit more
> invasive and touch bpf selftests, something like:
>
> Let me know if I should send a V2 and CC bpf maintainers, thanks!
>
> diff --git a/net/bpf/test_run.c b/net/bpf/test_run.c
> index 513354e928cb58f837fdc93a77f2e896de1ef916..8f4e9dd18001a4c32455fefd74fe38db270756d1
> 100644
> --- a/net/bpf/test_run.c
> +++ b/net/bpf/test_run.c
> @@ -1165,7 +1165,8 @@ int bpf_prog_test_run_skb(struct bpf_prog *prog,
> const union bpf_attr *kattr,
> goto out;
> }
> }
> - skb->protocol = eth_type_trans(skb, dev);
> + /* Always call the out-of-line version, for fentry/fexit selftests. */
> + skb->protocol = eth_type_trans_slow(skb, dev);
> skb_reset_network_header(skb);
>
> switch (skb->protocol) {
> diff --git a/tools/testing/selftests/bpf/progs/core_kern.c
> b/tools/testing/selftests/bpf/progs/core_kern.c
> index 004f2acef2eb0b0d80184423a3e8ba399feee1f3..fca26af98a7b92898a4adfde19f63b342f29db74
> 100644
> --- a/tools/testing/selftests/bpf/progs/core_kern.c
> +++ b/tools/testing/selftests/bpf/progs/core_kern.c
> @@ -47,14 +47,14 @@ int BPF_PROG(tp_xdp_devmap_xmit_multi, const
> struct net_device
> return randmap(from_dev->ifindex, from_dev);
> }
>
> -SEC("fentry/eth_type_trans")
> +SEC("fentry/eth_type_trans_slow")
> int BPF_PROG(fentry_eth_type_trans, struct sk_buff *skb,
> struct net_device *dev, unsigned short protocol)
> {
> return randmap(dev->ifindex + skb->len, dev);
> }
>
> -SEC("fexit/eth_type_trans")
> +SEC("fexit/eth_type_trans_slow")
> int BPF_PROG(fexit_eth_type_trans, struct sk_buff *skb,
> struct net_device *dev, unsigned short protocol)
> {
> diff --git a/tools/testing/selftests/bpf/progs/kfree_skb.c
> b/tools/testing/selftests/bpf/progs/kfree_skb.c
> index 7236da72ce8055f9a0165d3b4334a04d17d110b8..678c15c151944b44e276cf1280c77cb436ff7b10
> 100644
> --- a/tools/testing/selftests/bpf/progs/kfree_skb.c
> +++ b/tools/testing/selftests/bpf/progs/kfree_skb.c
> @@ -114,7 +114,7 @@ struct {
> bool fexit_test_ok;
> } result = {};
>
> -SEC("fentry/eth_type_trans")
> +SEC("fentry/eth_type_trans_slow")
> int BPF_PROG(fentry_eth_type_trans, struct sk_buff *skb, struct
> net_device *dev,
> unsigned short protocol)
> {
> @@ -132,7 +132,7 @@ int BPF_PROG(fentry_eth_type_trans, struct sk_buff
> *skb, struct net_device *dev,
> return 0;
> }
>
> -SEC("fexit/eth_type_trans")
> +SEC("fexit/eth_type_trans_slow")
> int BPF_PROG(fexit_eth_type_trans, struct sk_buff *skb, struct net_device *dev,
> unsigned short protocol)
> {
This shouldn't break anything, but CCing bpf just in case.
Thanks,
Olek
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net-next] net: inline eth_type_trans() fast path
2026-09-30 17:29 [PATCH net-next] net: inline eth_type_trans() fast path Eric Dumazet
2026-10-01 15:57 ` Alexander Lobakin
@ 2026-10-02 18:52 ` Eric Dumazet
1 sibling, 0 replies; 6+ messages in thread
From: Eric Dumazet @ 2026-10-02 18:52 UTC (permalink / raw)
To: David S . Miller, Jakub Kicinski, Paolo Abeni
Cc: Simon Horman, netdev, edumazet
Le mer. 30 sept. 2026 à 19:30, Eric Dumazet <edumazet@kernel.org> a écrit :
>
> eth_type_trans() is called once per received packet by most
> Ethernet drivers, and by core helpers (napi_gro_frags(),
> xdp_build_skb_from_*(), veth, tun, tunnels, loopback...).
>
> With CONFIG_MITIGATION_RETHUNK / SRSO, the call and return
> are not free anymore.
>
> Add eth_type_trans_inline(), handling the common case inline:
> unicast frame sent to dev->dev_addr, with a real ethertype,
> on a device which is not a DSA conduit. All other frames
> (multicast, broadcast, otherhost, 802.2, runts, DSA)
> are handled by the out-of-line eth_type_trans().
>
> eth_type_trans() is now a macro calling eth_type_trans_inline(),
> so that all existing callers get the fast path. The out-of-line
> version remains exported, and can be called with
> (eth_type_trans)(skb, dev). bpf_prog_test_run_skb() uses this
> to keep tools/testing/selftests/bpf fentry/fexit tests working.
>
> The inline part is about 25 instructions on x86_64, and the whole
> function would be almost twice as big when fully inlined.
I think we can fully inline it.
I will remove PX raw 802.3 detection from eth_type_trans() first.
pw-bot: cr
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-10-02 18:52 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-30 17:29 [PATCH net-next] net: inline eth_type_trans() fast path Eric Dumazet
2026-10-01 15:57 ` Alexander Lobakin
2026-10-01 17:10 ` Eric Dumazet
2026-10-02 15:26 ` Stanislav Fomichev
2026-10-02 16:54 ` Alexander Lobakin
2026-10-02 18:52 ` Eric Dumazet
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox