Netdev List
 help / color / mirror / Atom feed
* Performance question: af_packet with bpf filter vs TX path skb_clone
@ 2023-07-21 17:55 Maciej Żenczykowski
  2023-07-21 18:14 ` Eric Dumazet
  0 siblings, 1 reply; 11+ messages in thread
From: Maciej Żenczykowski @ 2023-07-21 17:55 UTC (permalink / raw)
  To: Linux NetDev, Jesper Dangaard Brouer, Eric Dumazet, Pengtao He,
	Willem Bruijn, Stanislav Fomichev, Xiao Ma, Patrick Rohr,
	Alexei Starovoitov

I've been asked to review:
  https://android-review.googlesource.com/c/platform/packages/modules/NetworkStack/+/2648779

where it comes to light that in Android due to background debugging of
connectivity problems
(of which there are *plenty* due to various types of buggy [primarily]
wifi networks)
we have a permanent AF_PACKET, ETH_P_ALL socket with a cBPF filter:

   arp or (ip and udp port 68) or (icmp6 and ip6[40] >= 133 and ip6[40] <= 136)

ie. it catches ARP, IPv4 DHCP and IPv6 ND (NS/NA/RS/RA)

If I'm reading the kernel code right this appears to cause skb_clone()
to be called on *every* outgoing packet,
even though most packets will not be accepted by the filter.

(In the TX path the filter appears to get called *after* the clone,
I think that's unlike the RX path where the filter is called first)

Unfortunately, I don't think it's possible to eliminate the
functionality this socket provides.
We need to be able to log RX & TX of ARP/DHCP/ND for debugging /
bugreports / etc.
and they *really* should be in order wrt. to each other.
(and yeah, that means last few minutes history when an issue happens,
so not possible to simply enable it on demand)

We could of course split the socket into 3 separate ones:
- ETH_P_ARP
- ETH_P_IP + cbpf udp dport=dhcp
- ETH_P_IPV6 + cbpf icmpv6 type=NS/NA/RS/RA

But I don't think that will help - I believe we'll still get
skb_clone() for every outbound ipv4/ipv6 packet.

I have some ideas for what could be done to avoid the clone (with
existing kernel functionality)... but none of it is pretty...
Anyone have any smart ideas?

Perhaps a way to move the clone past the af_packet packet_rcv run_filter?
Unfortunately packet_rcv() does a little bit of 'setup' before it
calls the filter - so this may be hard.

Or an 'extra' early pre-filter hook [prot_hook.prefilter()] that has
very minimal
functionality... like match 2 bytes at an offset into the packet?
Maybe even not a hook at all, just adding a
prot_hook.prefilter{1,2}_u64_{offset,mask,value}
It doesn't have to be perfect, but if it could discard 99% of the
packets we don't care about...
(and leave filtering of the remaining 1% to the existing cbpf program)
that would already be a huge win?

Thoughts?

Thanks,
Maciej

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2023-08-05 10:09 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2023-07-21 17:55 Performance question: af_packet with bpf filter vs TX path skb_clone Maciej Żenczykowski
2023-07-21 18:14 ` Eric Dumazet
2023-07-21 18:18   ` Stanislav Fomichev
2023-07-21 18:24     ` Eric Dumazet
2023-07-21 18:24     ` Maciej Żenczykowski
2023-07-21 18:56       ` Willem de Bruijn
2023-07-21 19:12         ` Maciej Żenczykowski
2023-08-02 14:30   ` Jesper Dangaard Brouer
2023-08-03  8:46     ` Maciej Żenczykowski
2023-08-05  8:54       ` Vincent Bernat
2023-08-05 10:09         ` Maciej Żenczykowski

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox