From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pz2-f41.google.com (mail-pz2-f41.google.com [74.125.228.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 949514E7816 for ; Mon, 21 Sep 2026 17:39:46 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.228.41 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790012388; cv=none; b=XR9cmE+fB3exe8cZY6QZ6VcfEFegzs00FASOIiaJL7mUOUO0YtM8QVRRbVMtATQoVFwrERg97agr/mELKyJ+R0M8jaVrumJIg67SKamcH9yQDTbyukMpk3YQnDu47L2vNHvBlwd+YS06EJagSifZH/D+AK44is5UfbyVkyauBO8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790012388; c=relaxed/simple; bh=U5EKdWBIbiSJU1dP09Omb4q7Mq8wk1havQycrFwGi+s=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=N/vV9kD2Y8j1TTJnIhgXfG79pjIwhFrzBHa6YMDE7mV17Q1b0tSv44rStL7WgzVZuK3IrYDi+jOdXyxlZKHbf+KXfGMxq+v0/E/JnOUtWqKjC4q3s8INUQS0tmB0SAMewtlc8u/r/qq1/RTQ8Yb2w8xmGGEEYqmF6F75HJJQIfw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=jXTAaKPs; arc=none smtp.client-ip=74.125.228.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="jXTAaKPs" Received: by mail-pz2-f41.google.com with SMTP id 41be03b00d2f7-cc50d1b048eso2592955a12.1 for ; Mon, 21 Sep 2026 10:39:46 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790012386; x=1790617186; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=w5FIqOvJ37TsMdMZxbiL8+8z7eHOaHEabZWsoVStaOQ=; b=jXTAaKPs74N5povwGroOLXHnRCwicKLScjR/cOolkBqb0m1nMJUoAhXmrw/kUfxa8q tEj+K80ADjNfAiv/6Yu/CO9K3RsiMflhgwZyy5iw0eQmxQBlKISNecdxLJdc2Rx8yW8c 86xKSWxWw37D0bOiMWTBqBWz+HedZZWEZBfxntJAzHmwKpMjJ10uKAXGsMxWrN5ISctA nWW/GxMrlv9OsBAnBqliBj6HrwDtFfU9VnaIS59SIb0soR2lseF+wap1lDdjyuqWEBWa Te45lDBiTH+veIaFvwaWEdI5Dr15KUqsedSC/q7laSnBcO2gs+vXnH3f72NF5SDdshgO t32Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790012386; x=1790617186; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=w5FIqOvJ37TsMdMZxbiL8+8z7eHOaHEabZWsoVStaOQ=; b=JhiOM/yOCC0iXN1EDRyWWFnBzOd0mN4iKdoLrHkM65dg1yzt/VucejOd0buODdcxZT t1uzAP4AlI7a7VdTmfRbGO2ci+KWOF/Y0enFql6wyTPU610ayKRtdUHb78L9mGqied9B 9inPRpgKLcAxyjX52AQdPd9MkI6ovuRiBitwmo2pWSvkhp/sOqMvIQVM0sHyCbg3A3Ht c4XjJHdbK3txv11O5mFXRyGvQ5zg1rur9qij6v125DT2h4sawR9wE1HLCHsIEHecCBC3 8oJ94A5TTtRb+9FhZm17GudpvLPUd6vjy8e/EMRX4JZZ6Sqx/u9/GI7K1/Rm+ErPssPj ozmA== X-Forwarded-Encrypted: i=1; AKwUvByaKY3O4iNYIhTv4OOCDFZMWzGB2WrzSl8eUyhYicS2zSQbeIamsLEId30w24tTP8PqN+fYQ0IEx+Q8httwqhs=@vger.kernel.org X-Gm-Message-State: AFuF++lOUfnOCt88NXv7iyk5ioH9s6TIqsA1/2b4bMgRQzJPbZtAq40g kLDSCyHfaNyY5WbtC3MKIhi1G74vV0BmdQabuH6t9rIo5f7xu4ltQPr4 X-Gm-Gg: AYBFou2aYmD2EXSP26SIEQGK1F4EdgS3/KIibnfE3dCxcd1Uq1bY1eHpYFwrN5sO2Sg p1J338QPoUB9RfmcPLUAI6pi6j/L5MZM0Vuo2pGSQj2x5nai0r3lzdp5vp7CHyBe2QbJXh7JhyX Pvf37SIW9p1tMnxqrOLgzre3gvCpdxp34b6fTVqgdxegFwjmViQuGZzjjdZqeA6Ldus9ZPKugKC 6vDxr0zhJhDUFMhdqLecMI5MBU5osPxBiwadlwoIaJ5MkeR0ytQyiDMMd4aMQQUAsB7r+e9ydgk 3qUQmIKKglkddY4iQB94YKBxIu1xRXqKV/SRq2QFTwAeevMSxTZSAQvJD26M8lW+qa6ErdPwDOA ZFykb/iUE5Tx1gxhY6WM6uZlAcCYCbuGxZypqTWgKXgcaMZHsygE2gcz99ZINAr6dgboMt3ULlC l74D0zn/Hilkk6czR/SEEZC4DdH8M4cqm7LYH7EZHkaND0oXnXVhWtSOfv4wL1TBRLri938yNEJ 3P70Zzd0Pr2ZmF8wh6wOiUXrYKtdpY2 X-Received: by 2002:a05:6300:618d:b0:3dd:a196:69e4 with SMTP id adf61e73a8af0-3dda196784emr12197593637.63.1790012385816; Mon, 21 Sep 2026 10:39:45 -0700 (PDT) Received: from 192.168.50.3 ([198.176.50.208]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-cc72af0955esm3986985a12.32.2026.09.21.10.39.31 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 10:39:45 -0700 (PDT) From: Weiming Shi To: Alexei Starovoitov , Daniel Borkmann , Andrii Nakryiko , Eduard Zingerman , Kumar Kartikeya Dwivedi , Martin KaFai Lau , Song Liu , Yonghong Song , Jiri Olsa , Emil Tsalapatis , Ihor Solodrai , John Fastabend , "David S . Miller" , Eric Dumazet , Jakub Kicinski , Paolo Abeni , Simon Horman , Shuah Khan Cc: bpf@vger.kernel.org, linux-kernel@vger.kernel.org, netdev@vger.kernel.org, linux-kselftest@vger.kernel.org, Xiang Mei , Weiming Shi , stable@vger.kernel.org, Alexei Starovoitov Subject: [PATCH v4 1/2] bpf: clear stale IPv4 options after LWT encapsulation Date: Tue, 22 Sep 2026 01:39:09 +0800 Message-ID: <20260921173910.90847-2-bestswngs@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260921173910.90847-1-bestswngs@gmail.com> References: <20260921173910.90847-1-bestswngs@gmail.com> Precedence: bulk X-Mailing-List: linux-kselftest@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit bpf_lwt_push_ip_encap() rebases the network header after prepending an IP header, but leaves IPCB(skb)->opt describing the inner IPv4 header. An ingress LWT route can consequently make an ICMP error use stale option offsets when constructing its reply. Mark completed LWT IP encapsulation in the run's BPF network context and clear stale metadata after bpf_prog_run_save_cb() restores the protocol control block. For same-family IPv4 encapsulation, clear only the cached options, as the IP tunnel transmit paths do. When the family changes, reset the destination family's control block and initialize its ingress interface, following seg6_do_srh_encap(). Save and restore the marker around nested LWT runs. Cc: stable@vger.kernel.org Fixes: 52f278774e79 ("bpf: implement BPF_LWT_ENCAP_IP mode in bpf_lwt_push_encap") Reported-by: Xiang Mei Suggested-by: Daniel Borkmann Suggested-by: Alexei Starovoitov Assisted-by: LLM Signed-off-by: Weiming Shi --- include/linux/filter.h | 1 + net/core/lwt_bpf.c | 30 ++++++++++++++++++++++++++++++ 2 files changed, 31 insertions(+) diff --git a/include/linux/filter.h b/include/linux/filter.h index 39decde7fc73..195898af9528 100644 --- a/include/linux/filter.h +++ b/include/linux/filter.h @@ -848,6 +848,7 @@ struct bpf_nh_params { #define BPF_RI_F_CPU_MAP_INIT BIT(2) #define BPF_RI_F_DEV_MAP_INIT BIT(3) #define BPF_RI_F_XSK_MAP_INIT BIT(4) +#define BPF_RI_F_LWT_IP_ENCAP BIT(5) struct bpf_redirect_info { u64 tgt_index; diff --git a/net/core/lwt_bpf.c b/net/core/lwt_bpf.c index da49364ec63d..88664382ca44 100644 --- a/net/core/lwt_bpf.c +++ b/net/core/lwt_bpf.c @@ -36,10 +36,30 @@ static inline struct bpf_lwt *bpf_lwt_lwtunnel(struct lwtunnel_state *lwt) #define NO_REDIRECT false #define CAN_REDIRECT true +static void bpf_lwt_reset_cb(struct sk_buff *skb, __be16 orig_proto) +{ + if (skb->protocol == orig_proto) { + if (skb->protocol == htons(ETH_P_IP)) + memset(&IPCB(skb)->opt, 0, sizeof(IPCB(skb)->opt)); + return; + } + + if (skb->protocol == htons(ETH_P_IP)) { + memset(IPCB(skb), 0, sizeof(*IPCB(skb))); + IPCB(skb)->iif = skb->skb_iif; + } else if (skb->protocol == htons(ETH_P_IPV6)) { + memset(IP6CB(skb), 0, sizeof(*IP6CB(skb))); + IP6CB(skb)->iif = skb->skb_iif; + } +} + static int run_lwt_bpf(struct sk_buff *skb, struct bpf_lwt_prog *lwt, struct dst_entry *dst, bool can_redirect) { struct bpf_net_context __bpf_net_ctx, *bpf_net_ctx; + struct bpf_redirect_info *ri; + bool lwt_ip_encap, saved_lwt_ip_encap; + __be16 orig_proto = skb->protocol; int ret; /* Disabling BH is needed to protect per-CPU bpf_redirect_info between @@ -47,8 +67,17 @@ static int run_lwt_bpf(struct sk_buff *skb, struct bpf_lwt_prog *lwt, */ local_bh_disable(); bpf_net_ctx = bpf_net_ctx_set(&__bpf_net_ctx); + ri = bpf_net_ctx_get_ri(); + saved_lwt_ip_encap = ri->kern_flags & BPF_RI_F_LWT_IP_ENCAP; + ri->kern_flags &= ~BPF_RI_F_LWT_IP_ENCAP; bpf_compute_data_pointers(skb); ret = bpf_prog_run_save_cb(lwt->prog, skb); + lwt_ip_encap = ri->kern_flags & BPF_RI_F_LWT_IP_ENCAP; + ri->kern_flags &= ~BPF_RI_F_LWT_IP_ENCAP; + if (saved_lwt_ip_encap) + ri->kern_flags |= BPF_RI_F_LWT_IP_ENCAP; + if (lwt_ip_encap) + bpf_lwt_reset_cb(skb, orig_proto); switch (ret) { case BPF_OK: @@ -668,6 +697,7 @@ int bpf_lwt_push_ip_encap(struct sk_buff *skb, void *hdr, u32 len, bool ingress) } else { skb->protocol = htons(ETH_P_IPV6); } + bpf_net_ctx_get_ri()->kern_flags |= BPF_RI_F_LWT_IP_ENCAP; if (skb_is_gso(skb)) return handle_gso_encap(skb, ipv4, len); -- 2.55.0