Netdev List
 help / color / mirror / Atom feed
* [RFC bpf-next v2 0/3] bpf: Add PPPoE encap/decap support to bpf_skb_adjust_room
@ 2026-10-03 19:33 ThisSeanZhang
  2026-10-03 19:33 ` [RFC bpf-next v2 1/3] bpf: Add PPPoE encap " ThisSeanZhang
                   ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: ThisSeanZhang @ 2026-10-03 19:33 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, netdev, Nick Hudson, Felix Fietkau,
	Qingfang Deng

Hello,

A concrete use case for this support is a TC BPF PPPoE forwarding
path. When such a path forwards a GSO packet larger than the negotiated
PPPoE MTU, inserting the PPPoE header manually leaves skb->protocol set
to the original IP protocol. The stack then cannot select the PPPoE
segmentation path, and the packet can be dropped.

The kernel already provides PPPoE GRO/GSO support for ETH_P_PPP_SES
(added by commit 55a5d8fca836 ("net: pppoe: implement GRO/GSO support")),
but a TC BPF program currently cannot update skb->protocol and related
metadata when it adds or removes the PPPoE header. As a result, later
networking code does not see the packet as a PPPoE frame and cannot use
those handlers correctly.

The current workaround uses
BPF_F_ADJ_ROOM_ENCAP_L3_IPV4/IPV6, which marks the skb as IP-in-IP
while the wire format is PPPoE. This requires the program to adjust
gso_size itself and leaves the skb metadata describing an encapsulation
that is not on the wire [1].

This series adds PPPoE session encapsulation and decapsulation support
to bpf_skb_adjust_room(). The new operations update the packet layout
and skb metadata together, making the skb compatible with the kernel's
existing PPPoE GRO/GSO handlers.

Examples:

    bpf_skb_adjust_room(skb, PPPOE_SES_HLEN, BPF_ADJ_ROOM_MAC,
                        BPF_F_ADJ_ROOM_ENCAP_PPPOE);
    bpf_skb_adjust_room(skb, -PPPOE_SES_HLEN, BPF_ADJ_ROOM_MAC,
                        BPF_F_ADJ_ROOM_DECAP_PPPOE);

Encapsulation changes skb->protocol to ETH_P_PPP_SES. Decapsulation
accepts only packets whose skb->protocol is ETH_P_PPP_SES, reads the
PPP protocol field, and restores ETH_P_IP or ETH_P_IPV6. Both flags
require BPF_ADJ_ROOM_MAC and a fixed eight-byte adjustment.

The PPPoE GRO/GSO support must be built in or loaded before this path
is used for receive-side GRO or software GSO; loading pppoe.ko remains
the responsibility of the user-space setup.

The interface follows the existing BPF_F_ADJ_ROOM_ENCAP_* family. I
also considered a kfunc-based interface and would appreciate
feedback on whether that would be preferred.

The selftest covers IPv4 and IPv6 round trips, protocol transitions,
and invalid flag, size, mode, protocol, and truncated-packet cases. It
passes in a QEMU/KVM boot of the resulting kernel (11/11 subtests).

[1] https://github.com/ThisSeanZhang/landscape/blob/80bc43eef7143ad029283fe2a8d510718c22af12/landscape-ebpf/src/bpf/tc_chain/tc_pppoe.bpf.c#L38-L47

ThisSeanZhang (3):
  bpf: Add PPPoE encap support to bpf_skb_adjust_room
  bpf: Add PPPoE decap support to bpf_skb_adjust_room
  selftests/bpf: Add a test for the PPPoE encap/decap adjust_room flags

 include/uapi/linux/bpf.h                      |  22 ++
 net/core/filter.c                             |  85 +++++-
 tools/include/uapi/linux/bpf.h                |  22 ++
 .../selftests/bpf/prog_tests/tc_pppoe.c       | 254 ++++++++++++++++++
 tools/testing/selftests/bpf/progs/tc_pppoe.c  | 163 +++++++++++
 5 files changed, 543 insertions(+), 3 deletions(-)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/tc_pppoe.c
 create mode 100644 tools/testing/selftests/bpf/progs/tc_pppoe.c

---

Changes since v1:
- Rework the cover letter to explain the forwarding use case and the
  PPPoE skb metadata transition.
- Clarify the PPPoE module prerequisite and GRO/GSO scope.
- Replace the reviewer questions with a single question about the
  interface choice.
- Fix multi-line comment style in the selftests.

v1: https://lore.kernel.org/bpf/20260926210757.2152159-1-thisseanzhang@gmail.com/

-- 
2.47.3

^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC bpf-next v2 1/3] bpf: Add PPPoE encap support to bpf_skb_adjust_room
  2026-10-03 19:33 [RFC bpf-next v2 0/3] bpf: Add PPPoE encap/decap support to bpf_skb_adjust_room ThisSeanZhang
@ 2026-10-03 19:33 ` ThisSeanZhang
  2026-10-03 19:33 ` [RFC bpf-next v2 2/3] bpf: Add PPPoE decap " ThisSeanZhang
  2026-10-03 19:33 ` [RFC bpf-next v2 3/3] selftests/bpf: Add a test for the PPPoE encap/decap adjust_room flags ThisSeanZhang
  2 siblings, 0 replies; 5+ messages in thread
From: ThisSeanZhang @ 2026-10-03 19:33 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, netdev, Nick Hudson, Felix Fietkau,
	Qingfang Deng

Add a new BPF_F_ADJ_ROOM_ENCAP_PPPOE flag to bpf_skb_adjust_room() so
that a TC BPF program can encapsulate an IPv4 or IPv6 packet into a
PPPoE session header:

    bpf_skb_adjust_room(skb, PPPOE_SES_HLEN, BPF_ADJ_ROOM_MAC,
                        BPF_F_ADJ_ROOM_ENCAP_PPPOE);

The flag inserts eight bytes between the MAC and network headers: the
six-byte PPPoE session header followed by the two-byte PPP protocol
field. It then updates the skb metadata, including skb->protocol and
skb->mac_len. The BPF program remains responsible for filling in the
PPPoE header and the Ethernet ethertype, for example with
bpf_skb_store_bytes().

Previously, a TC program could add the header bytes manually, but the
skb would continue to be treated as a plain IP packet because its
protocol metadata was unchanged. With the stale skb->protocol, software
GSO may select a non-PPPoE segmentation handler and process the packet
incorrectly.

This makes the skb metadata compatible with the PPPoE GRO/GSO handlers
registered for ETH_P_PPP_SES. The PPPoE GRO/GSO support must be built in
or loaded before receive-side GRO or software GSO uses those handlers.

Signed-off-by: ThisSeanZhang <thisseanzhang@gmail.com>
---
 include/uapi/linux/bpf.h       | 12 ++++++++++++
 net/core/filter.c              | 22 ++++++++++++++++++++++
 tools/include/uapi/linux/bpf.h | 12 ++++++++++++
 3 files changed, 46 insertions(+)

diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 732b35cc0..30481d040 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -3036,6 +3036,17 @@ union bpf_attr {
  *		  Use with BPF_F_ADJ_ROOM_ENCAP_L2 flag to further specify the
  *		  L2 type as Ethernet.
  *
+ *		* **BPF_F_ADJ_ROOM_ENCAP_PPPOE**:
+ *		  Encapsulate the packet in a PPPoE session header. Must be
+ *		  used with **BPF_ADJ_ROOM_MAC** mode and *len_diff* equal to
+ *		  the size of the PPPoE session header plus the PPP protocol
+ *		  field (8 bytes in total). The room is inserted between the
+ *		  MAC header and the network header, *skb->protocol* is set
+ *		  to **ETH_P_PPP_SES** and *skb->mac_len* is updated
+ *		  accordingly. The PPPoE header itself and the ethertype of
+ *		  the MAC header are filled in by the BPF program, e.g. via
+ *		  **bpf_skb_store_bytes**.
+ *
  *		* **BPF_F_ADJ_ROOM_DECAP_L3_IPV4**,
  *		  **BPF_F_ADJ_ROOM_DECAP_L3_IPV6**:
  *		  Indicate the new IP header version after decapsulating the
@@ -6323,6 +6334,7 @@ enum bpf_adj_room_flags {
 	BPF_F_ADJ_ROOM_DECAP_L4_UDP	= (1ULL << 10),
 	BPF_F_ADJ_ROOM_DECAP_IPXIP4	= (1ULL << 11),
 	BPF_F_ADJ_ROOM_DECAP_IPXIP6	= (1ULL << 12),
+	BPF_F_ADJ_ROOM_ENCAP_PPPOE	= (1ULL << 13),
 };
 
 enum {
diff --git a/net/core/filter.c b/net/core/filter.c
index 70dc62167..5b204e316 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -47,6 +47,7 @@
 #include <linux/ratelimit.h>
 #include <linux/seccomp.h>
 #include <linux/if_vlan.h>
+#include <linux/if_pppox.h>
 #include <linux/bpf.h>
 #include <linux/btf.h>
 #include <net/sch_generic.h>
@@ -3583,6 +3584,7 @@ static u32 bpf_skb_net_base_len(const struct sk_buff *skb)
 					 BPF_F_ADJ_ROOM_ENCAP_L4_GRE | \
 					 BPF_F_ADJ_ROOM_ENCAP_L4_UDP | \
 					 BPF_F_ADJ_ROOM_ENCAP_L2_ETH | \
+					 BPF_F_ADJ_ROOM_ENCAP_PPPOE | \
 					 BPF_F_ADJ_ROOM_ENCAP_L2( \
 					  BPF_ADJ_ROOM_ENCAP_L2_MASK))
 
@@ -3688,6 +3690,14 @@ static int bpf_skb_net_grow(struct sk_buff *skb, u32 off, u32 len_diff,
 			skb_dst_drop(skb);
 	}
 
+	if (flags & BPF_F_ADJ_ROOM_ENCAP_PPPOE) {
+		/* Network header points at the inserted PPPoE header. */
+		skb->protocol = htons(ETH_P_PPP_SES);
+		skb_reset_mac_len(skb);
+		if (skb_valid_dst(skb))
+			skb_dst_drop(skb);
+	}
+
 	if (skb_is_gso(skb)) {
 		struct skb_shared_info *shinfo = skb_shinfo(skb);
 
@@ -3874,6 +3884,18 @@ BPF_CALL_4(bpf_skb_adjust_room, struct sk_buff *, skb, s32, len_diff,
 		return -ENOTSUPP;
 	}
 
+	if (flags & BPF_F_ADJ_ROOM_ENCAP_PPPOE) {
+		/* The PPPoE session header has a fixed size and is
+		 * inserted directly after the MAC header.
+		 */
+		if (shrink || mode != BPF_ADJ_ROOM_MAC ||
+		    len_diff != PPPOE_SES_HLEN ||
+		    flags & ((BPF_F_ADJ_ROOM_ENCAP_MASK |
+			      BPF_F_ADJ_ROOM_DECAP_MASK) &
+			     ~BPF_F_ADJ_ROOM_ENCAP_PPPOE))
+			return -EINVAL;
+	}
+
 	if (flags & BPF_F_ADJ_ROOM_DECAP_MASK) {
 		u32 len_decap_min = 0;
 
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 732b35cc0..30481d040 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -3036,6 +3036,17 @@ union bpf_attr {
  *		  Use with BPF_F_ADJ_ROOM_ENCAP_L2 flag to further specify the
  *		  L2 type as Ethernet.
  *
+ *		* **BPF_F_ADJ_ROOM_ENCAP_PPPOE**:
+ *		  Encapsulate the packet in a PPPoE session header. Must be
+ *		  used with **BPF_ADJ_ROOM_MAC** mode and *len_diff* equal to
+ *		  the size of the PPPoE session header plus the PPP protocol
+ *		  field (8 bytes in total). The room is inserted between the
+ *		  MAC header and the network header, *skb->protocol* is set
+ *		  to **ETH_P_PPP_SES** and *skb->mac_len* is updated
+ *		  accordingly. The PPPoE header itself and the ethertype of
+ *		  the MAC header are filled in by the BPF program, e.g. via
+ *		  **bpf_skb_store_bytes**.
+ *
  *		* **BPF_F_ADJ_ROOM_DECAP_L3_IPV4**,
  *		  **BPF_F_ADJ_ROOM_DECAP_L3_IPV6**:
  *		  Indicate the new IP header version after decapsulating the
@@ -6323,6 +6334,7 @@ enum bpf_adj_room_flags {
 	BPF_F_ADJ_ROOM_DECAP_L4_UDP	= (1ULL << 10),
 	BPF_F_ADJ_ROOM_DECAP_IPXIP4	= (1ULL << 11),
 	BPF_F_ADJ_ROOM_DECAP_IPXIP6	= (1ULL << 12),
+	BPF_F_ADJ_ROOM_ENCAP_PPPOE	= (1ULL << 13),
 };
 
 enum {
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [RFC bpf-next v2 2/3] bpf: Add PPPoE decap support to bpf_skb_adjust_room
  2026-10-03 19:33 [RFC bpf-next v2 0/3] bpf: Add PPPoE encap/decap support to bpf_skb_adjust_room ThisSeanZhang
  2026-10-03 19:33 ` [RFC bpf-next v2 1/3] bpf: Add PPPoE encap " ThisSeanZhang
@ 2026-10-03 19:33 ` ThisSeanZhang
  2026-10-03 19:33 ` [RFC bpf-next v2 3/3] selftests/bpf: Add a test for the PPPoE encap/decap adjust_room flags ThisSeanZhang
  2 siblings, 0 replies; 5+ messages in thread
From: ThisSeanZhang @ 2026-10-03 19:33 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, netdev, Nick Hudson, Felix Fietkau,
	Qingfang Deng

Add a new BPF_F_ADJ_ROOM_DECAP_PPPOE flag to bpf_skb_adjust_room() so
that a TC BPF program can decapsulate a PPPoE session packet:

    bpf_skb_adjust_room(skb, -PPPOE_SES_HLEN, BPF_ADJ_ROOM_MAC,
                        BPF_F_ADJ_ROOM_DECAP_PPPOE);

The flag removes the six-byte PPPoE session header and the following
two-byte PPP protocol field between the MAC and network headers. It
restores consistent skb metadata for later consumers: the PPP protocol
field selects ETH_P_IP or ETH_P_IPV6, and any other PPP protocol is
rejected.

Before this patch, bpf_skb_adjust_room() rejected PPPoE packets because
it only accepted ETH_P_IP or ETH_P_IPV6 in skb->protocol. A TC program
therefore could not remove the PPPoE header and restore the inner
protocol metadata with the existing helper.

This is the counterpart to BPF_F_ADJ_ROOM_ENCAP_PPPOE introduced in
the previous patch and restores plain IPv4 or IPv6 skb metadata after
PPPoE decapsulation.

The flag also resets mac_len after removing the header. On flows where
a packet encapsulated on the same host re-enters TC ingress without
passing through the receive path, mac_len may still include the
removed PPPoE header. The reset mirrors what pppoe_gso_segment() does
for its segments.

Signed-off-by: ThisSeanZhang <thisseanzhang@gmail.com>
---
 include/uapi/linux/bpf.h       | 10 +++++
 net/core/filter.c              | 69 +++++++++++++++++++++++++++++++---
 tools/include/uapi/linux/bpf.h | 10 +++++
 3 files changed, 83 insertions(+), 6 deletions(-)

diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 30481d040..52a509060 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -3071,6 +3071,15 @@ union bpf_attr {
  *		  decapsulating a tunnel with an outer IPv6 header (IPv6-in-IPv6
  *		  or IPv4-in-IPv6).
  *
+ *		* **BPF_F_ADJ_ROOM_DECAP_PPPOE**:
+ *		  Decapsulate a PPPoE session header. Must be used with
+ *		  **BPF_ADJ_ROOM_MAC** mode and a negative *len_diff* equal to
+ *		  the size of the PPPoE session header plus the PPP protocol
+ *		  field (8 bytes in total). The PPP protocol field determines
+ *		  the new *skb->protocol* (**ETH_P_IP** or **ETH_P_IPV6**);
+ *		  a payload with any other PPP protocol is rejected. The Ethernet
+ *		  type in the MAC header is left for the BPF program to restore.
+ *
  *		When using the decapsulation flags above, the skb->encapsulation
  *		flag is automatically cleared if all tunnel-specific GSO flags
  *		(SKB_GSO_UDP_TUNNEL, SKB_GSO_UDP_TUNNEL_CSUM, SKB_GSO_GRE,
@@ -6335,6 +6344,7 @@ enum bpf_adj_room_flags {
 	BPF_F_ADJ_ROOM_DECAP_IPXIP4	= (1ULL << 11),
 	BPF_F_ADJ_ROOM_DECAP_IPXIP6	= (1ULL << 12),
 	BPF_F_ADJ_ROOM_ENCAP_PPPOE	= (1ULL << 13),
+	BPF_F_ADJ_ROOM_DECAP_PPPOE	= (1ULL << 14),
 };
 
 enum {
diff --git a/net/core/filter.c b/net/core/filter.c
index 5b204e316..46986a84e 100644
--- a/net/core/filter.c
+++ b/net/core/filter.c
@@ -48,6 +48,7 @@
 #include <linux/seccomp.h>
 #include <linux/if_vlan.h>
 #include <linux/if_pppox.h>
+#include <linux/ppp_defs.h>
 #include <linux/bpf.h>
 #include <linux/btf.h>
 #include <net/sch_generic.h>
@@ -3590,7 +3591,8 @@ static u32 bpf_skb_net_base_len(const struct sk_buff *skb)
 
 #define BPF_F_ADJ_ROOM_DECAP_MASK	(BPF_F_ADJ_ROOM_DECAP_L3_MASK | \
 					 BPF_F_ADJ_ROOM_DECAP_L4_MASK | \
-					 BPF_F_ADJ_ROOM_DECAP_IPXIP_MASK)
+					 BPF_F_ADJ_ROOM_DECAP_IPXIP_MASK | \
+					 BPF_F_ADJ_ROOM_DECAP_PPPOE)
 
 #define BPF_F_ADJ_ROOM_MASK		(BPF_F_ADJ_ROOM_FIXED_GSO | \
 					 BPF_F_ADJ_ROOM_ENCAP_MASK | \
@@ -3724,6 +3726,8 @@ static int bpf_skb_net_shrink(struct sk_buff *skb, u32 off, u32 len_diff,
 			      u64 flags)
 {
 	bool decap = flags & BPF_F_ADJ_ROOM_DECAP_L3_MASK;
+	__be16 inner_proto = 0;
+	u32 inner_len = 0;
 	int ret;
 
 	if (unlikely(flags & ~(BPF_F_ADJ_ROOM_DECAP_MASK |
@@ -3742,6 +3746,33 @@ static int bpf_skb_net_shrink(struct sk_buff *skb, u32 off, u32 len_diff,
 	if (unlikely(ret < 0))
 		return ret;
 
+	if (flags & BPF_F_ADJ_ROOM_DECAP_PPPOE) {
+		u16 ppp_proto;
+
+		if (unlikely(!pskb_may_pull(skb, off + PPPOE_SES_HLEN)))
+			return -ENOMEM;
+
+		/* PPP protocol field follows the PPPoE session header. */
+		ppp_proto = get_unaligned_be16(skb->data + off +
+					       sizeof(struct pppoe_hdr));
+		switch (ppp_proto) {
+		case PPP_IP:
+			inner_proto = htons(ETH_P_IP);
+			inner_len = sizeof(struct iphdr);
+			break;
+		case PPP_IPV6:
+			inner_proto = htons(ETH_P_IPV6);
+			inner_len = sizeof(struct ipv6hdr);
+			break;
+		default:
+			return -ENOTSUPP;
+		}
+
+		/* A full inner L3 header must remain after decapsulation. */
+		if (skb->len - off - PPPOE_SES_HLEN < inner_len)
+			return -EINVAL;
+	}
+
 	ret = bpf_skb_net_hdr_pop(skb, off, len_diff);
 	if (unlikely(ret < 0))
 		return ret;
@@ -3757,6 +3788,13 @@ static int bpf_skb_net_shrink(struct sk_buff *skb, u32 off, u32 len_diff,
 			skb_dst_drop(skb);
 	}
 
+	if (flags & BPF_F_ADJ_ROOM_DECAP_PPPOE) {
+		skb->protocol = inner_proto;
+		skb_reset_mac_len(skb);
+		if (skb_valid_dst(skb))
+			skb_dst_drop(skb);
+	}
+
 	if (skb_is_gso(skb)) {
 		struct skb_shared_info *shinfo = skb_shinfo(skb);
 
@@ -3869,9 +3907,13 @@ BPF_CALL_4(bpf_skb_adjust_room, struct sk_buff *, skb, s32, len_diff,
 		return -EINVAL;
 	if (unlikely(len_diff_abs > 0xfffU))
 		return -EFAULT;
-	if (unlikely(proto != htons(ETH_P_IP) &&
-		     proto != htons(ETH_P_IPV6)))
+	if (unlikely(flags & BPF_F_ADJ_ROOM_DECAP_PPPOE)) {
+		if (proto != htons(ETH_P_PPP_SES))
+			return -ENOTSUPP;
+	} else if (unlikely(proto != htons(ETH_P_IP) &&
+			    proto != htons(ETH_P_IPV6))) {
 		return -ENOTSUPP;
+	}
 
 	off = skb_mac_header_len(skb);
 	switch (mode) {
@@ -3885,9 +3927,7 @@ BPF_CALL_4(bpf_skb_adjust_room, struct sk_buff *, skb, s32, len_diff,
 	}
 
 	if (flags & BPF_F_ADJ_ROOM_ENCAP_PPPOE) {
-		/* The PPPoE session header has a fixed size and is
-		 * inserted directly after the MAC header.
-		 */
+		/* Fixed-size header, inserted after the MAC header. */
 		if (shrink || mode != BPF_ADJ_ROOM_MAC ||
 		    len_diff != PPPOE_SES_HLEN ||
 		    flags & ((BPF_F_ADJ_ROOM_ENCAP_MASK |
@@ -3920,6 +3960,23 @@ BPF_CALL_4(bpf_skb_adjust_room, struct sk_buff *, skb, s32, len_diff,
 		    (flags & BPF_F_ADJ_ROOM_DECAP_IPXIP_MASK))
 			return -EINVAL;
 
+		/* PPPoE decapsulation is mutually exclusive with the
+		 * other decapsulation types.
+		 */
+		if ((flags & BPF_F_ADJ_ROOM_DECAP_PPPOE) &&
+		    (flags & (BPF_F_ADJ_ROOM_DECAP_MASK &
+			      ~BPF_F_ADJ_ROOM_DECAP_PPPOE)))
+			return -EINVAL;
+
+		if (flags & BPF_F_ADJ_ROOM_DECAP_PPPOE) {
+			/* Fixed-size header; require a full inner L3 header. */
+			if (mode != BPF_ADJ_ROOM_MAC ||
+			    len_diff_abs != PPPOE_SES_HLEN)
+				return -EINVAL;
+
+			len_min = sizeof(struct iphdr);
+		}
+
 		if (flags & BPF_F_ADJ_ROOM_DECAP_L4_MASK)
 			len_decap_min += bpf_skb_net_base_len(skb);
 
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 30481d040..52a509060 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -3071,6 +3071,15 @@ union bpf_attr {
  *		  decapsulating a tunnel with an outer IPv6 header (IPv6-in-IPv6
  *		  or IPv4-in-IPv6).
  *
+ *		* **BPF_F_ADJ_ROOM_DECAP_PPPOE**:
+ *		  Decapsulate a PPPoE session header. Must be used with
+ *		  **BPF_ADJ_ROOM_MAC** mode and a negative *len_diff* equal to
+ *		  the size of the PPPoE session header plus the PPP protocol
+ *		  field (8 bytes in total). The PPP protocol field determines
+ *		  the new *skb->protocol* (**ETH_P_IP** or **ETH_P_IPV6**);
+ *		  a payload with any other PPP protocol is rejected. The Ethernet
+ *		  type in the MAC header is left for the BPF program to restore.
+ *
  *		When using the decapsulation flags above, the skb->encapsulation
  *		flag is automatically cleared if all tunnel-specific GSO flags
  *		(SKB_GSO_UDP_TUNNEL, SKB_GSO_UDP_TUNNEL_CSUM, SKB_GSO_GRE,
@@ -6335,6 +6344,7 @@ enum bpf_adj_room_flags {
 	BPF_F_ADJ_ROOM_DECAP_IPXIP4	= (1ULL << 11),
 	BPF_F_ADJ_ROOM_DECAP_IPXIP6	= (1ULL << 12),
 	BPF_F_ADJ_ROOM_ENCAP_PPPOE	= (1ULL << 13),
+	BPF_F_ADJ_ROOM_DECAP_PPPOE	= (1ULL << 14),
 };
 
 enum {
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [RFC bpf-next v2 3/3] selftests/bpf: Add a test for the PPPoE encap/decap adjust_room flags
  2026-10-03 19:33 [RFC bpf-next v2 0/3] bpf: Add PPPoE encap/decap support to bpf_skb_adjust_room ThisSeanZhang
  2026-10-03 19:33 ` [RFC bpf-next v2 1/3] bpf: Add PPPoE encap " ThisSeanZhang
  2026-10-03 19:33 ` [RFC bpf-next v2 2/3] bpf: Add PPPoE decap " ThisSeanZhang
@ 2026-10-03 19:33 ` ThisSeanZhang
  2026-10-03 20:11   ` bot+bpf-ci
  2 siblings, 1 reply; 5+ messages in thread
From: ThisSeanZhang @ 2026-10-03 19:33 UTC (permalink / raw)
  To: bpf
  Cc: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
	Eduard Zingerman, netdev, Nick Hudson, Felix Fietkau,
	Qingfang Deng

Add a tc test that exercises the new BPF_F_ADJ_ROOM_ENCAP_PPPOE and
BPF_F_ADJ_ROOM_DECAP_PPPOE flags of bpf_skb_adjust_room().

The encap program is run over plain Ethernet/IPv4 and Ethernet/IPv6
packets. The test checks that the PPPoE session header is inserted
between the Ethernet and network headers, that the PPP protocol and
skb->protocol are correct, and that the payload is preserved. It then
runs the decap program and checks that the original packet and
protocol metadata are restored.

A rejection program exercises calls to bpf_skb_adjust_room() that
the helper must reject: an encapsulation length other than
PPPOE_SES_HLEN, encapsulation in BPF_ADJ_ROOM_NET mode, a negative
encapsulation length, the encap flag combined with a decap flag,
shrinking a PPPoE packet without the PPPoE flag, and decapsulation of a
packet whose skb->protocol is not ETH_P_PPP_SES. Decapsulation of a PPPoE
packet with an unsupported PPP protocol and of a packet too short to
still hold an IP header is rejected as well.

Signed-off-by: ThisSeanZhang <thisseanzhang@gmail.com>
---
 .../selftests/bpf/prog_tests/tc_pppoe.c       | 254 ++++++++++++++++++
 tools/testing/selftests/bpf/progs/tc_pppoe.c  | 163 +++++++++++
 2 files changed, 417 insertions(+)
 create mode 100644 tools/testing/selftests/bpf/prog_tests/tc_pppoe.c
 create mode 100644 tools/testing/selftests/bpf/progs/tc_pppoe.c

diff --git a/tools/testing/selftests/bpf/prog_tests/tc_pppoe.c b/tools/testing/selftests/bpf/prog_tests/tc_pppoe.c
new file mode 100644
index 000000000..73e602c91
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/tc_pppoe.c
@@ -0,0 +1,254 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 ThisSeanZhang */
+
+#include <arpa/inet.h>
+#include <test_progs.h>
+#include "tc_pppoe.skel.h"
+
+/* Ethernet + IPv4 + TCP, 54 bytes in total. */
+static const __u8 ip4_pkt[] = {
+	/* Ethernet header. */
+	0x11, 0x22, 0x33, 0x44, 0x55, 0x66,
+	0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff,
+	0x08, 0x00,
+	/* IPv4 header. */
+	0x45, 0x00, 0x00, 0x28,
+	0x12, 0x34, 0x40, 0x00,
+	0x40, 0x06, 0x00, 0x00,
+	0xc0, 0xa8, 0x01, 0x01,
+	0xc0, 0xa8, 0x01, 0x02,
+	/* TCP header. */
+	0x00, 0x50, 0x1f, 0x90,
+	0x00, 0x00, 0x00, 0x01,
+	0x00, 0x00, 0x00, 0x00,
+	0x50, 0x02, 0x10, 0x00,
+	0x00, 0x00, 0x00, 0x00,
+};
+
+/* Ethernet + IPv6 + TCP, 74 bytes in total. */
+static const __u8 ip6_pkt[] = {
+	/* Ethernet header. */
+	0x11, 0x22, 0x33, 0x44, 0x55, 0x66,
+	0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff,
+	0x86, 0xdd,
+	/* IPv6 header. */
+	0x60, 0x00, 0x00, 0x00,
+	0x00, 0x14, 0x06, 0x40,
+	0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00,
+	0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01,
+	0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00,
+	0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x02,
+	/* TCP header. */
+	0x00, 0x50, 0x1f, 0x90,
+	0x00, 0x00, 0x00, 0x01,
+	0x00, 0x00, 0x00, 0x00,
+	0x50, 0x02, 0x10, 0x00,
+	0x00, 0x00, 0x00, 0x00,
+};
+
+#define PPP_SES_HLEN	8
+#define TC_ACT_SHOT	2
+
+static int run_prog(int prog_fd, const void *data_in, __u32 size_in,
+		    void *data_out, __u32 size_out, __u32 *retval,
+		    __u32 *size_out_actual)
+{
+	LIBBPF_OPTS(bpf_test_run_opts, opts,
+		    .data_in = (void *)data_in,
+		    .data_size_in = size_in,
+		    .data_out = data_out,
+		    .data_size_out = size_out,
+	);
+	int ret;
+
+	ret = bpf_prog_test_run_opts(prog_fd, &opts);
+	if (!ret && retval)
+		*retval = opts.retval;
+	if (!ret && size_out_actual)
+		*size_out_actual = opts.data_size_out;
+
+	return ret;
+}
+
+static void test_encap_decap(struct tc_pppoe *skel, const char *subtest,
+			     const __u8 *pkt, __u32 pkt_len, int ethertype,
+			     __u8 ppp_proto)
+{
+	__u8 encap_pkt[128];
+	__u8 decap_pkt[128];
+	__u32 retval, out_len;
+	int ret;
+
+	if (!test__start_subtest(subtest))
+		return;
+
+	skel->bss->encap_proto = 0;
+	ret = run_prog(bpf_program__fd(skel->progs.tc_pppoe_encap),
+		       pkt, pkt_len, encap_pkt, sizeof(encap_pkt),
+		       &retval, &out_len);
+	ASSERT_OK(ret, "encap test_run");
+	ASSERT_OK(retval, "encap retval");
+	ASSERT_EQ(out_len, pkt_len + PPP_SES_HLEN, "encap pkt len");
+	ASSERT_EQ(skel->bss->encap_proto, htons(0x8864),
+		  "encap skb protocol");
+
+	/* Ethernet header, with the ethertype changed to PPPoE session. */
+	ASSERT_EQ(encap_pkt[12], 0x88, "encap eth h_proto");
+	ASSERT_EQ(encap_pkt[13], 0x64, "encap eth h_proto");
+	/*
+	 * PPPoE session header: ver/type/code, session id, length and
+	 * PPP protocol.
+	 */
+	ASSERT_EQ(encap_pkt[14], 0x11, "encap pppoe ver/type");
+	ASSERT_EQ(encap_pkt[15], 0x00, "encap pppoe code");
+	ASSERT_EQ(encap_pkt[16], 0xde, "encap pppoe sid");
+	ASSERT_EQ(encap_pkt[17], 0xad, "encap pppoe sid");
+	ASSERT_EQ(encap_pkt[18], 0x00, "encap pppoe length");
+	ASSERT_EQ(encap_pkt[19], pkt_len - 14 + 2, "encap pppoe length");
+	ASSERT_EQ(encap_pkt[20], 0x00, "encap ppp proto");
+	ASSERT_EQ(encap_pkt[21], ppp_proto, "encap ppp proto");
+	/*
+	 * The original packet must be shifted unchanged behind the
+	 * new header.
+	 */
+	ASSERT_MEMEQ(encap_pkt + 14 + PPP_SES_HLEN, pkt + 14,
+		     pkt_len - 14, "encap payload");
+
+	skel->bss->decap_proto = 0;
+	ret = run_prog(bpf_program__fd(skel->progs.tc_pppoe_decap),
+		       encap_pkt, pkt_len + PPP_SES_HLEN, decap_pkt,
+		       sizeof(decap_pkt), &retval, &out_len);
+	ASSERT_OK(ret, "decap test_run");
+	ASSERT_OK(retval, "decap retval");
+	ASSERT_EQ(out_len, pkt_len, "decap pkt len");
+	ASSERT_EQ(skel->bss->decap_proto, ethertype, "decap skb protocol");
+	ASSERT_MEMEQ(decap_pkt, pkt, pkt_len, "decap packet");
+}
+
+static void test_reject(struct tc_pppoe *skel, int case_id,
+			const void *data_in, __u32 size_in, const char *name)
+{
+	__u8 out[128];
+	__u32 retval = 0;
+	int ret;
+
+	if (!test__start_subtest(name))
+		return;
+
+	skel->bss->reject_case = case_id;
+	skel->bss->reject_unexpected = 0;
+	ret = run_prog(bpf_program__fd(skel->progs.tc_pppoe_reject),
+		       data_in, size_in, out, sizeof(out), &retval, NULL);
+	ASSERT_OK(ret, "reject test_run");
+	ASSERT_EQ(retval, TC_ACT_SHOT, "helper rejected the call");
+	ASSERT_EQ(skel->bss->reject_unexpected, 0, "no unexpected success");
+}
+
+static void test_decap_reject_input(struct tc_pppoe *skel,
+				    const __u8 *pkt, __u32 pkt_len,
+				    const char *subtest)
+{
+	__u8 out[128];
+	__u32 retval = 0;
+	int ret;
+
+	if (!test__start_subtest(subtest))
+		return;
+
+	skel->bss->decap_proto = 0;
+	ret = run_prog(bpf_program__fd(skel->progs.tc_pppoe_decap),
+		       pkt, pkt_len, out, sizeof(out), &retval, NULL);
+	ASSERT_OK(ret, "decap reject test_run");
+	ASSERT_EQ(retval, TC_ACT_SHOT, "helper rejected the call");
+	ASSERT_EQ(skel->bss->decap_proto, 0, "skb protocol unchanged");
+}
+
+void test_tc_pppoe(void)
+{
+	/*
+	 * A PPPoE packet whose PPP protocol is neither IPv4 nor IPv6:
+	 * IP control protocol (0x8021) in this case.
+	 */
+	static const __u8 bad_ppp_pkt[] = {
+		0x11, 0x22, 0x33, 0x44, 0x55, 0x66,
+		0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff,
+		0x88, 0x64,
+		0x11, 0x00, 0x00, 0x00, 0x00, 0x28, 0x80, 0x21,
+		0x45, 0x00, 0x00, 0x28,
+		0x12, 0x34, 0x40, 0x00,
+		0x40, 0x06, 0x00, 0x00,
+		0xc0, 0xa8, 0x01, 0x01,
+		0xc0, 0xa8, 0x01, 0x02,
+		0x00, 0x50, 0x1f, 0x90,
+		0x00, 0x00, 0x00, 0x01,
+		0x00, 0x00, 0x00, 0x00,
+		0x50, 0x02, 0x10, 0x00,
+		0x00, 0x00, 0x00, 0x00,
+	};
+	/*
+	 * A PPPoE packet whose payload is too short to still contain a
+	 * full IP header after decapsulation.
+	 */
+	static const __u8 truncated_pkt[] = {
+		0x11, 0x22, 0x33, 0x44, 0x55, 0x66,
+		0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff,
+		0x88, 0x64,
+		0x11, 0x00, 0x00, 0x00, 0x00, 0x02, 0x00, 0x21,
+		0x45, 0x00,
+	};
+	/* Non-PPPoE IPv4 packet whose bytes 20/21 decode as PPP_IP. */
+	static const __u8 fake_ppp_pkt[] = {
+		0x11, 0x22, 0x33, 0x44, 0x55, 0x66,
+		0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff,
+		0x08, 0x00,
+		0x45, 0x00, 0x00, 0x28,
+		0x12, 0x34, 0x00, 0x21,
+		0x40, 0x06, 0x00, 0x00,
+		0xc0, 0xa8, 0x01, 0x01,
+		0xc0, 0xa8, 0x01, 0x02,
+		0x00, 0x50, 0x1f, 0x90,
+		0x00, 0x00, 0x00, 0x01,
+		0x00, 0x00, 0x00, 0x00,
+		0x50, 0x02, 0x10, 0x00,
+		0x00, 0x00, 0x00, 0x00,
+	};
+	__u8 encap_pkt[128];
+	struct tc_pppoe *skel;
+	__u32 retval, out_len;
+
+	skel = tc_pppoe__open_and_load();
+	if (!ASSERT_OK_PTR(skel, "skel open_and_load"))
+		return;
+
+	test_encap_decap(skel, "encap-decap-v4", ip4_pkt, sizeof(ip4_pkt),
+			 htons(0x0800), 0x21);
+	test_encap_decap(skel, "encap-decap-v6", ip6_pkt, sizeof(ip6_pkt),
+			 htons(0x86dd), 0x57);
+
+	test_decap_reject_input(skel, bad_ppp_pkt, sizeof(bad_ppp_pkt),
+				"decap-bad-ppp-proto");
+	test_decap_reject_input(skel, truncated_pkt, sizeof(truncated_pkt),
+				"decap-truncated");
+	test_reject(skel, 6, fake_ppp_pkt, sizeof(fake_ppp_pkt),
+		    "reject-decap-fake-ppp-proto");
+
+	/*
+	 * Encapsulate a v4 packet once more to get a PPPoE packet as
+	 * input for the "decap without the flag" rejection case.
+	 */
+	ASSERT_OK(run_prog(bpf_program__fd(skel->progs.tc_pppoe_encap),
+			   ip4_pkt, sizeof(ip4_pkt), encap_pkt,
+			   sizeof(encap_pkt), &retval, &out_len),
+		  "encap v4 for reject input");
+
+	test_reject(skel, 1, ip4_pkt, sizeof(ip4_pkt), "reject-encap-len");
+	test_reject(skel, 2, ip4_pkt, sizeof(ip4_pkt), "reject-encap-mode");
+	test_reject(skel, 3, ip4_pkt, sizeof(ip4_pkt), "reject-encap-shrink");
+	test_reject(skel, 4, ip4_pkt, sizeof(ip4_pkt), "reject-flag-mix");
+	test_reject(skel, 5, encap_pkt, sizeof(ip4_pkt) + PPP_SES_HLEN,
+		    "reject-decap-no-flag");
+	test_reject(skel, 6, ip4_pkt, sizeof(ip4_pkt),
+		    "reject-decap-non-pppoe");
+
+	tc_pppoe__destroy(skel);
+}
diff --git a/tools/testing/selftests/bpf/progs/tc_pppoe.c b/tools/testing/selftests/bpf/progs/tc_pppoe.c
new file mode 100644
index 000000000..b1624f0f9
--- /dev/null
+++ b/tools/testing/selftests/bpf/progs/tc_pppoe.c
@@ -0,0 +1,163 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026 ThisSeanZhang */
+
+#include "vmlinux.h"
+#include <bpf/bpf_helpers.h>
+#include <bpf/bpf_endian.h>
+
+#define ETH_P_IP_TEST		0x0800
+#define ETH_P_IPV6_TEST		0x86dd
+#define ETH_P_PPP_SES_TEST	0x8864
+#define PPP_IP_TEST		0x21
+#define PPP_IPV6_TEST		0x57
+
+#define ETH_HLEN_TEST		14
+#define PPPOE_SES_HLEN_TEST	8
+
+#define TC_ACT_OK_TEST		0
+#define TC_ACT_SHOT_TEST	2
+
+/* Selects the bpf_skb_adjust_room() call made by tc_pppoe_reject. */
+int reject_case;
+
+/* skb->protocol as observed after the helper call. */
+int encap_proto;
+int decap_proto;
+
+/*
+ * Set when tc_pppoe_reject observes a call that should have been
+ * rejected by the helper.
+ */
+int reject_unexpected;
+
+SEC("tc")
+int tc_pppoe_encap(struct __sk_buff *skb)
+{
+	__u8 hdr[PPPOE_SES_HLEN_TEST] = {
+		0x11, 0x00,			/* ver, type, code */
+		0xde, 0xad,			/* session id */
+		0x00, 0x00,			/* length, set below */
+		0x00, 0x00,			/* PPP protocol, set below */
+	};
+	struct ethhdr eth;
+	/*
+	 * The PPPoE length field covers everything after the 6 byte
+	 * session header: the PPP protocol field plus the payload.
+	 */
+	__u16 plen = skb->len - ETH_HLEN_TEST + 2;
+
+	hdr[4] = (plen >> 8) & 0xff;
+	hdr[5] = plen & 0xff;
+
+	switch (skb->protocol) {
+	case bpf_htons(ETH_P_IP_TEST):
+		hdr[7] = PPP_IP_TEST;
+		break;
+	case bpf_htons(ETH_P_IPV6_TEST):
+		hdr[7] = PPP_IPV6_TEST;
+		break;
+	default:
+		return TC_ACT_SHOT_TEST;
+	}
+
+	if (bpf_skb_adjust_room(skb, PPPOE_SES_HLEN_TEST, BPF_ADJ_ROOM_MAC,
+				BPF_F_ADJ_ROOM_ENCAP_PPPOE))
+		return TC_ACT_SHOT_TEST;
+
+	encap_proto = skb->protocol;
+
+	if (bpf_skb_store_bytes(skb, ETH_HLEN_TEST, hdr, sizeof(hdr), 0))
+		return TC_ACT_SHOT_TEST;
+
+	if (bpf_skb_load_bytes(skb, 0, &eth, sizeof(eth)))
+		return TC_ACT_SHOT_TEST;
+	eth.h_proto = bpf_htons(ETH_P_PPP_SES_TEST);
+	if (bpf_skb_store_bytes(skb, 0, &eth, sizeof(eth), 0))
+		return TC_ACT_SHOT_TEST;
+
+	return TC_ACT_OK_TEST;
+}
+
+SEC("tc")
+int tc_pppoe_decap(struct __sk_buff *skb)
+{
+	struct ethhdr eth;
+
+	if (bpf_skb_load_bytes(skb, 0, &eth, sizeof(eth)))
+		return TC_ACT_SHOT_TEST;
+	if (eth.h_proto != bpf_htons(ETH_P_PPP_SES_TEST))
+		return TC_ACT_SHOT_TEST;
+
+	if (bpf_skb_adjust_room(skb, -PPPOE_SES_HLEN_TEST, BPF_ADJ_ROOM_MAC,
+				BPF_F_ADJ_ROOM_DECAP_PPPOE))
+		return TC_ACT_SHOT_TEST;
+
+	decap_proto = skb->protocol;
+
+	/*
+	 * Restore the ethertype to the protocol of the decapsulated
+	 * payload, as picked by the kernel from the PPP protocol field.
+	 */
+	eth.h_proto = (__be16)skb->protocol;
+	if (bpf_skb_store_bytes(skb, 0, &eth, sizeof(eth), 0))
+		return TC_ACT_SHOT_TEST;
+
+	return TC_ACT_OK_TEST;
+}
+
+/*
+ * Every bpf_skb_adjust_room() call below must be rejected by the
+ * helper; tc_pppoe_reject reports (and fails the test) if one of
+ * them unexpectedly succeeds.
+ */
+SEC("tc")
+int tc_pppoe_reject(struct __sk_buff *skb)
+{
+	int ret = 0;
+
+	switch (reject_case) {
+	case 1:
+		/* Encap with a wrong room size. */
+		ret = bpf_skb_adjust_room(skb, PPPOE_SES_HLEN_TEST - 4,
+					  BPF_ADJ_ROOM_MAC,
+					  BPF_F_ADJ_ROOM_ENCAP_PPPOE);
+		break;
+	case 2:
+		/* Encap at the wrong position. */
+		ret = bpf_skb_adjust_room(skb, PPPOE_SES_HLEN_TEST,
+					  BPF_ADJ_ROOM_NET,
+					  BPF_F_ADJ_ROOM_ENCAP_PPPOE);
+		break;
+	case 3:
+		/* Encap with a negative room size. */
+		ret = bpf_skb_adjust_room(skb, -PPPOE_SES_HLEN_TEST,
+					  BPF_ADJ_ROOM_MAC,
+					  BPF_F_ADJ_ROOM_ENCAP_PPPOE);
+		break;
+	case 4:
+		/* Encap flag combined with a decap flag. */
+		ret = bpf_skb_adjust_room(skb, PPPOE_SES_HLEN_TEST,
+					  BPF_ADJ_ROOM_MAC,
+					  BPF_F_ADJ_ROOM_ENCAP_PPPOE |
+					  BPF_F_ADJ_ROOM_DECAP_PPPOE);
+		break;
+	case 5:
+		/* Shrink of a PPPoE packet without the PPPoE flag. */
+		ret = bpf_skb_adjust_room(skb, -PPPOE_SES_HLEN_TEST,
+					  BPF_ADJ_ROOM_MAC, 0);
+		break;
+	case 6:
+		/* Decap of a non-PPPoE packet (plain IP or fake PPP proto). */
+		ret = bpf_skb_adjust_room(skb, -PPPOE_SES_HLEN_TEST,
+					  BPF_ADJ_ROOM_MAC,
+					  BPF_F_ADJ_ROOM_DECAP_PPPOE);
+		break;
+	}
+
+	if (!ret)
+		reject_unexpected = 1;
+
+	return TC_ACT_SHOT_TEST;
+}
+
+char _license[] SEC("license") = "GPL";
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 5+ messages in thread

* Re: [RFC bpf-next v2 3/3] selftests/bpf: Add a test for the PPPoE encap/decap adjust_room flags
  2026-10-03 19:33 ` [RFC bpf-next v2 3/3] selftests/bpf: Add a test for the PPPoE encap/decap adjust_room flags ThisSeanZhang
@ 2026-10-03 20:11   ` bot+bpf-ci
  0 siblings, 0 replies; 5+ messages in thread
From: bot+bpf-ci @ 2026-10-03 20:11 UTC (permalink / raw)
  To: thisseanzhang, bpf
  Cc: ast, daniel, andrii, eddyz87, netdev, nhudson, nbd, dqfext, ast,
	andrii, daniel, martin.lau, eddyz87, yonghong.song, mason,
	ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 6141 bytes --]

> selftests/bpf: Add a test for the PPPoE encap/decap adjust_room flags

TL;DR:
- Test issue: reject-flag-mix and reject-encap-shrink pass even without the
  new ENCAP_PPPOE checks
- No IPv6 truncated-decap subtest, so the inner_len check goes untested

> diff --git a/tools/testing/selftests/bpf/prog_tests/tc_pppoe.c b/tools/testing/selftests/bpf/prog_tests/tc_pppoe.c
> new file mode 100644
> index 000000000000..73e602c91230
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/tc_pppoe.c
> @@ -0,0 +1,254 @@

[ ... ]

> +void test_tc_pppoe(void)
> +{

[ ... ]

> +	/*
> +	 * A PPPoE packet whose payload is too short to still contain a
> +	 * full IP header after decapsulation.
> +	 */
> +	static const __u8 truncated_pkt[] = {
> +		0x11, 0x22, 0x33, 0x44, 0x55, 0x66,
> +		0xaa, 0xbb, 0xcc, 0xdd, 0xee, 0xff,
> +		0x88, 0x64,
> +		0x11, 0x00, 0x00, 0x00, 0x00, 0x02, 0x00, 0x21,
> +		0x45, 0x00,
> +	};

[ ... ]

> +	test_decap_reject_input(skel, bad_ppp_pkt, sizeof(bad_ppp_pkt),
> +				"decap-bad-ppp-proto");
> +	test_decap_reject_input(skel, truncated_pkt, sizeof(truncated_pkt),
> +				"decap-truncated");

Does decap-truncated reach the new "full inner L3 header" check in
bpf_skb_net_shrink()?

truncated_pkt is 24 bytes, so in bpf_skb_adjust_room() the generic length
check fires first:

    len_cur = 24 - 14 = 10
    len_cur - len_diff_abs = 2, which is < len_min (sizeof(struct iphdr))

and the call returns -ENOTSUPP before bpf_skb_net_shrink() runs.

The check added by the decap patch:

    /* A full inner L3 header must remain after decapsulation. */
    if (skb->len - off - PPPOE_SES_HLEN < inner_len)
        return -EINVAL;

is redundant for PPP_IP here, because inner_len == len_min == 20.  It is the
only guard for PPP_IPV6, where inner_len is 40, and no subtest feeds a
PPPoE/IPv6 packet with 20-39 bytes after the PPP header.

If that check were removed or broken, an skb could be marked ETH_P_IPV6
without a full IPv6 header, and none of the subtests would notice.

Would it make sense to add an IPv6 truncated case, for example PPP protocol
0x0057 with 20 bytes of payload, so the inner_len check is exercised?

> +	test_reject(skel, 6, fake_ppp_pkt, sizeof(fake_ppp_pkt),
> +		    "reject-decap-fake-ppp-proto");

[ ... ]

> +	test_reject(skel, 1, ip4_pkt, sizeof(ip4_pkt), "reject-encap-len");
> +	test_reject(skel, 2, ip4_pkt, sizeof(ip4_pkt), "reject-encap-mode");
> +	test_reject(skel, 3, ip4_pkt, sizeof(ip4_pkt), "reject-encap-shrink");
> +	test_reject(skel, 4, ip4_pkt, sizeof(ip4_pkt), "reject-flag-mix");

Can reject-flag-mix or reject-encap-shrink fail on a kernel that lacks the
new ENCAP_PPPOE checks?

For reject-flag-mix, case 4 passes ENCAP_PPPOE | DECAP_PPPOE on ip4_pkt, so
skb->protocol is ETH_P_IP.  bpf_skb_adjust_room() returns early at the
decap protocol gate:

    if (unlikely(flags & BPF_F_ADJ_ROOM_DECAP_PPPOE)) {
        if (proto != htons(ETH_P_PPP_SES))
            return -ENOTSUPP;

which is before the ENCAP_PPPOE flag combination check is reached.  Even with
a PPPoE input, a grow with any decap flag is already rejected by the existing
"if (!shrink) return -EINVAL;" in the BPF_F_ADJ_ROOM_DECAP_MASK block.

For reject-encap-shrink, without the "shrink ||" clause in the ENCAP_PPPOE
check, bpf_skb_net_shrink() would still return -EINVAL, because
BPF_F_ADJ_ROOM_ENCAP_PPPOE is outside its allowed flag mask.

tc_pppoe_reject only records !ret, so both subtests pass whether or not the
new checks exist.  A negative test should give a different result on a
kernel that is missing the check it targets.

Should case 4 use ENCAP_PPPOE together with another encap flag, for example
BPF_F_ADJ_ROOM_ENCAP_L3_IPV4 or BPF_F_ADJ_ROOM_ENCAP_L2_ETH, so the new
rejection is actually exercised?

[ ... ]

> diff --git a/tools/testing/selftests/bpf/progs/tc_pppoe.c b/tools/testing/selftests/bpf/progs/tc_pppoe.c
> new file mode 100644
> index 000000000000..b1624f0f9d8a
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/progs/tc_pppoe.c
> @@ -0,0 +1,163 @@

[ ... ]

> +SEC("tc")
> +int tc_pppoe_reject(struct __sk_buff *skb)
> +{
> +	int ret = 0;
> +
> +	switch (reject_case) {

[ ... ]

> +	case 4:
> +		/* Encap flag combined with a decap flag. */
> +		ret = bpf_skb_adjust_room(skb, PPPOE_SES_HLEN_TEST,
> +					  BPF_ADJ_ROOM_MAC,
> +					  BPF_F_ADJ_ROOM_ENCAP_PPPOE |
> +					  BPF_F_ADJ_ROOM_DECAP_PPPOE);
> +		break;

Does this case reach the ENCAP_PPPOE flag combination check it is meant to
exercise?

prog_tests/tc_pppoe.c runs it as test_reject(skel, 4, ip4_pkt, ...), so
eth_type_trans() in bpf_prog_test_run_skb() sets skb->protocol to
htons(ETH_P_IP).  Because BPF_F_ADJ_ROOM_DECAP_PPPOE is set,
bpf_skb_adjust_room() returns -ENOTSUPP at the protocol check, before the
new validation is reached:

    flags & ((BPF_F_ADJ_ROOM_ENCAP_MASK |
              BPF_F_ADJ_ROOM_DECAP_MASK) &
             ~BPF_F_ADJ_ROOM_ENCAP_PPPOE))
        return -EINVAL;

The call is rejected for the same reason as case 6 (reject-decap-non-pppoe
on the same ip4_pkt), so this subtest adds no coverage.  With a PPPoE input,
a positive len_diff combined with any decap flag would still be caught by
the existing "if (!shrink) return -EINVAL;" in the DECAP_MASK block.

The combination that only the new check rejects is not tested at all:
ENCAP_PPPOE together with another encap flag, for example:

    bpf_skb_adjust_room(skb, PPPOE_SES_HLEN_TEST, BPF_ADJ_ROOM_MAC,
                        BPF_F_ADJ_ROOM_ENCAP_PPPOE |
                        BPF_F_ADJ_ROOM_ENCAP_L3_IPV4);

If that check were missing, this call would go into bpf_skb_net_grow() with
encap == true and succeed, and reject_unexpected would catch it.  The test
already builds fake_ppp_pkt so that case 6 can fail on a kernel without the
protocol check.

Could case 4 use a second encap flag in the same way, so that it also fails
on a kernel without the check it targets?


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/37148983810

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-10-03 20:11 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-03 19:33 [RFC bpf-next v2 0/3] bpf: Add PPPoE encap/decap support to bpf_skb_adjust_room ThisSeanZhang
2026-10-03 19:33 ` [RFC bpf-next v2 1/3] bpf: Add PPPoE encap " ThisSeanZhang
2026-10-03 19:33 ` [RFC bpf-next v2 2/3] bpf: Add PPPoE decap " ThisSeanZhang
2026-10-03 19:33 ` [RFC bpf-next v2 3/3] selftests/bpf: Add a test for the PPPoE encap/decap adjust_room flags ThisSeanZhang
2026-10-03 20:11   ` bot+bpf-ci

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox