* [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
@ 2026-09-24 6:33 Chenguang Zhao
2026-09-25 12:42 ` Ido Schimmel
` (2 more replies)
0 siblings, 3 replies; 6+ messages in thread
From: Chenguang Zhao @ 2026-09-24 6:33 UTC (permalink / raw)
To: davem, edumazet, kuba, pabeni, horms, dsahern, idosch
Cc: kerneljasonxing, netdev, chenguang.zhao, Chenguang Zhao,
syzbot+2b120190d9e54ad8c65d
From: Chenguang Zhao <zhaochenguang@kylinos.cn>
Commit de2c211868b9 ("ipvs: Always clear ipvs_property flag in
skb_scrub_packet()") moved ipvs_reset() before the xnet check, making
the call unconditional. The intent was to fix a bpf_redirect case where
stale ipvs_property on an skb re-entering the RX path caused the SNAT
hook to be skipped. However the change is too broad: when IPVS NAT
sits above an ipvlan L3 interface in the same netns, the following
loop happens:
LOCAL_OUT -> IPVS DNAT (sets ipvs_property=1) -> dst_output -> ipvlan
-> skb_scrub_packet() -> ipvs_reset() clears the flag
-> ipvlan_process_v4_outbound() -> ip_local_out() -> LOCAL_OUT
-> IPVS sees ipvs_property=0, processes again -> infinite recursion
syzbot reported this as a stack overflow on a KASAN kernel where each
level burns ~3.3 KB of stack and XMIT_RECURSION_LIMIT falls short. On
non-KASAN kernels the dead-loop detector catches it and prints "Dead
loop on virtual device", but traffic is still broken.
Fixes: de2c211868b9 ("ipvs: Always clear ipvs_property flag in skb_scrub_packet()")
Reported-by: syzbot+2b120190d9e54ad8c65d@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=2b120190d9e54ad8c65d
Signed-off-by: Chenguang Zhao <zhaochenguang@kylinos.cn>
---
Fix it with two changes:
1. Move ipvs_reset() back inside the xnet guard in skb_scrub_packet(),
so ipvs_property is only cleared when the skb actually crosses a
netns boundary. This restores IPVS re-entry protection for the
ipvlan path.
2. To preserve the bpf_redirect fix, add ipvs_reset() in ip_rcv() and
ipv6_rcv() right before the NF_HOOK into PREROUTING. Every
redirected packet enters the stack through these points, so
clearing ipvs_property there covers the original use case without
affecting the ipvlan code path.
net/core/skbuff.c | 3 +--
net/ipv4/ip_input.c | 1 +
net/ipv6/ip6_input.c | 1 +
3 files changed, 3 insertions(+), 2 deletions(-)
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index cc3b4b70288b..6912ca0d2228 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -6303,11 +6303,10 @@ void skb_scrub_packet(struct sk_buff *skb, bool xnet)
skb->offload_fwd_mark = 0;
skb->offload_l3_fwd_mark = 0;
#endif
- ipvs_reset(skb);
-
if (!xnet)
return;
+ ipvs_reset(skb);
skb->mark = 0;
skb_clear_tstamp(skb);
}
diff --git a/net/ipv4/ip_input.c b/net/ipv4/ip_input.c
index 9860178752b8..00f3b328e90a 100644
--- a/net/ipv4/ip_input.c
+++ b/net/ipv4/ip_input.c
@@ -609,6 +609,7 @@ int ip_rcv(struct sk_buff *skb, struct net_device *dev, struct packet_type *pt,
if (skb == NULL)
return NET_RX_DROP;
+ ipvs_reset(skb);
return NF_HOOK(NFPROTO_IPV4, NF_INET_PRE_ROUTING,
net, NULL, skb, dev, NULL,
ip_rcv_finish);
diff --git a/net/ipv6/ip6_input.c b/net/ipv6/ip6_input.c
index d332ec60f915..05917095ef6d 100644
--- a/net/ipv6/ip6_input.c
+++ b/net/ipv6/ip6_input.c
@@ -348,6 +348,7 @@ int ipv6_rcv(struct sk_buff *skb, struct net_device *dev, struct packet_type *pt
skb = ip6_rcv_core(skb, dev, net);
if (skb == NULL)
return NET_RX_DROP;
+ ipvs_reset(skb);
return NF_HOOK(NFPROTO_IPV6, NF_INET_PRE_ROUTING,
net, NULL, skb, dev, NULL,
ip6_rcv_finish);
--
2.25.1
^ permalink raw reply related [flat|nested] 6+ messages in thread
* Re: [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
2026-09-24 6:33 [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear Chenguang Zhao
@ 2026-09-25 12:42 ` Ido Schimmel
2026-09-28 6:55 ` Chenguang Zhao
2026-09-26 20:12 ` [syzbot ci] " syzbot ci
2026-09-28 6:58 ` [PATCH net] " netdev-bot+sashiko
2 siblings, 1 reply; 6+ messages in thread
From: Ido Schimmel @ 2026-09-25 12:42 UTC (permalink / raw)
To: Chenguang Zhao
Cc: davem, edumazet, kuba, pabeni, horms, dsahern, kerneljasonxing,
netdev, Chenguang Zhao, syzbot+2b120190d9e54ad8c65d, ja
On Thu, Sep 24, 2026 at 02:33:12PM +0800, Chenguang Zhao wrote:
> From: Chenguang Zhao <zhaochenguang@kylinos.cn>
>
> Commit de2c211868b9 ("ipvs: Always clear ipvs_property flag in
> skb_scrub_packet()") moved ipvs_reset() before the xnet check, making
> the call unconditional. The intent was to fix a bpf_redirect case where
> stale ipvs_property on an skb re-entering the RX path caused the SNAT
> hook to be skipped. However the change is too broad: when IPVS NAT
> sits above an ipvlan L3 interface in the same netns, the following
> loop happens:
>
> LOCAL_OUT -> IPVS DNAT (sets ipvs_property=1) -> dst_output -> ipvlan
> -> skb_scrub_packet() -> ipvs_reset() clears the flag
> -> ipvlan_process_v4_outbound() -> ip_local_out() -> LOCAL_OUT
> -> IPVS sees ipvs_property=0, processes again -> infinite recursion
>
> syzbot reported this as a stack overflow on a KASAN kernel where each
> level burns ~3.3 KB of stack and XMIT_RECURSION_LIMIT falls short. On
> non-KASAN kernels the dead-loop detector catches it and prints "Dead
> loop on virtual device", but traffic is still broken.
>
> Fixes: de2c211868b9 ("ipvs: Always clear ipvs_property flag in skb_scrub_packet()")
> Reported-by: syzbot+2b120190d9e54ad8c65d@syzkaller.appspotmail.com
> Closes: https://syzkaller.appspot.com/bug?extid=2b120190d9e54ad8c65d
> Signed-off-by: Chenguang Zhao <zhaochenguang@kylinos.cn>
1. You need to copy IPVS maintainers on patches related to IPVS. Added
Julian.
2. Does it reproduce with commit 7f1de03e3103 ("net: reduce
XMIT_RECURSION_LIMIT under KASAN") in net-next?
^ permalink raw reply [flat|nested] 6+ messages in thread
* [syzbot ci] Re: ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
2026-09-24 6:33 [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear Chenguang Zhao
2026-09-25 12:42 ` Ido Schimmel
@ 2026-09-26 20:12 ` syzbot ci
2026-09-28 6:58 ` [PATCH net] " netdev-bot+sashiko
2 siblings, 0 replies; 6+ messages in thread
From: syzbot ci @ 2026-09-26 20:12 UTC (permalink / raw)
To: chenguang.zhao, davem, dsahern, edumazet, horms, idosch,
kerneljasonxing, kuba, netdev, pabeni, syzbot, zhaochenguang
Cc: syzbot, syzkaller-bugs
syzbot ci has tested the following series
[v1] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
https://lore.kernel.org/all/20260924063312.1194019-1-chenguang.zhao@linux.dev
* [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
and found the following issue:
INFO: task hung in netdev_run_todo
Full report is available here:
https://ci.syzbot.org/series/b6478e5d-791e-4d72-8b8d-0fa748c90168
***
INFO: task hung in netdev_run_todo
tree: net
URL: https://kernel.googlesource.com/pub/scm/linux/kernel/git/netdev/net.git
base: e7e0a54300a896e731dffd3e2e8dae5631d6243d
arch: amd64
compiler: Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
config: https://ci.syzbot.org/builds/9865431b-add2-4883-97e2-de4ec60d969e/config
INFO: task syz-executor:31961 blocked for more than 144 seconds.
Not tainted syzkaller #0
Blocked by coredump.
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
task:syz-executor state:D stack:26120 pid:31961 tgid:31961 ppid:1 task_flags:0x40054c flags:0x00080003
Call Trace:
<TASK>
__schedule+0x17db/0x58f0
schedule+0x164/0x2b0
schedule_preempt_disabled+0x13/0x30
__mutex_lock+0x7c1/0x1550
rcu_barrier+0x4c/0x530
netdev_run_todo+0x2fc/0x1060
tun_chr_close+0x13c/0x1c0
__fput+0x418/0xa50
task_work_run+0x1d9/0x270
do_exit+0x73a/0x2360
do_group_exit+0x22d/0x2f0
get_signal+0x121b/0x12c0
arch_do_signal_or_restart+0xbb/0x860
exit_to_user_mode_loop+0x10e/0x770
do_syscall_64+0x328/0x520
entry_SYSCALL_64_after_hwframe+0x77/0x7f
RIP: 0033:0x7f4b2cf9ddeb
RSP: 002b:00007ffedfbd6b60 EFLAGS: 00000246 ORIG_RAX: 0000000000000010
RAX: 0000000000000000 RBX: 00007ffedfbd6c10 RCX: 00007f4b2cf9ddeb
RDX: 00007ffedfbd6be0 RSI: 00000000400454ca RDI: 00000000000000c8
RBP: 00007f4b2d0359ad R08: 0000000000000000 R09: 0000000000000000
R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000008
R13: 0000000000000003 R14: 00007ffedfbd6f08 R15: 0000000000000000
</TASK>
Showing all locks held in the system:
locks held by kworker/u8:1/13: 3, on CPU#1:
#0: ffff88817721b940 ((wq_completion)ipv6_addrconf){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#1: ffffc90000127c40 ((work_completion)(&(&net->ipv6.addr_chk_work)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#2: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: addrconf_verify_work+0x19/0x30
locks held by khungtaskd/35: 1, last CPU#1:
#0: ffffffff8ed5c8a0 (rcu_read_lock){....}-{1:3}, at: debug_show_all_locks+0x2e/0x180
locks held by getty/5419: 2, on CPU#1:
#0: ffff88817a7b30a0 (&tty->ldisc_sem){++++}-{0:0}, at: tty_ldisc_ref_wait+0x25/0x70
#1: ffffc900027682e8 (&ldata->atomic_read_lock){+.+.}-{4:4}, at: n_tty_read+0x45a/0x1360
locks held by kworker/u8:4/9945: 2, last CPU#1:
#0: ffff888177d09940 ((wq_completion)bat_events){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#1: ffffc90006db7c40 ((work_completion)(&(&bat_priv->tt.work)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
locks held by kworker/u8:5/19085: 4, last CPU#1:
#0: ffff88817741f940 ((wq_completion)krds_cp_wq#12/0){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#1: ffffc90003a0fc40 ((work_completion)(&(&cp->cp_conn_w)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#2: ffff888169e6a180 (&tc->t_conn_path_lock){+.+.}-{4:4}, at: rds_tcp_conn_path_connect+0x1cc/0x920
#3: ffff888192303660 (k-sk_lock-AF_INET){+.+.}-{0:0}, at: rds_tcp_tune+0x220/0x920
locks held by kworker/u10:0/25329: 3, on CPU#1:
#0: ffff8881000ac140 ((wq_completion)events_unbound){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#1: ffffc9000418fc40 ((linkwatch_work).work){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#2: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: linkwatch_event+0xe/0x60
locks held by kworker/0:1/27790: 7, last CPU#0:
#0: ffff88810006a940 ((wq_completion)events_long){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#1: ffffc9000141fc40 ((work_completion)(&(&br->gc_work)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#2: ffffffff8ed5c8a0 (rcu_read_lock){....}-{1:3}, at: process_backlog+0x3de/0x18b0
#3: ffffffff8ed5c8a0 (rcu_read_lock){....}-{1:3}, at: NF_HOOK+0x9e/0x3c0
#4: ffffffff8ed5c8a0 (rcu_read_lock){....}-{1:3}, at: ip_route_output_key_hash+0xd8/0x2a0
#5: ffff888121028298 (hrtimer_bases.lock){-.-.}-{2:2}, at: __hrtimer_rearm_deferred+0x99/0x4b0
#6: ffffffff9aa29748 (&____s->seqcount#2){----}-{0:0}, at: __ip_vs_conn_in_get+0x2db/0x10f0
locks held by syz.0.7664/31917: 1, on CPU#1:
#0: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: tun_chr_close+0x3e/0x1c0
locks held by syz.3.7666/31925: 1, on CPU#1:
#0: ffffffff8ed62b38 (rcu_state.barrier_mutex){+.+.}-{4:4}, at: rcu_barrier+0x4c/0x530
locks held by syz.2.7668/31938: 1, on CPU#1:
#0: ffffffff8ed62b38 (rcu_state.barrier_mutex){+.+.}-{4:4}, at: rcu_barrier+0x4c/0x530
locks held by syz-executor/31955: 1, on CPU#1:
#0: ffffffff8ed62b38 (rcu_state.barrier_mutex){+.+.}-{4:4}, at: rcu_barrier+0x4c/0x530
locks held by syz-executor/31961: 1, on CPU#1:
#0: ffffffff8ed62b38 (rcu_state.barrier_mutex){+.+.}-{4:4}, at: rcu_barrier+0x4c/0x530
locks held by kworker/u8:7/31997: 4, on CPU#1:
#0: ffff8881012d5940 ((wq_completion)netns){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#1: ffffc900079cfc40 (net_cleanup_work){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630
#2: ffffffff90244da8 (pernet_ops_rwsem){++++}-{4:4}, at: cleanup_net+0xf5/0x810
#3: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: ops_undo_list+0x28e/0x8d0
locks held by syz-executor/31998: 1, on CPU#1:
#0: ffffffff8ed62b38 (rcu_state.barrier_mutex){+.+.}-{4:4}, at: rcu_barrier+0x4c/0x530
locks held by syz-executor/32005: 1, on CPU#1:
#0: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: tun_chr_close+0x3e/0x1c0
locks held by syz-executor/32009: 1, last CPU#1:
#0: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: tun_chr_close+0x3e/0x1c0
locks held by syz-executor/32074: 1, on CPU#1:
#0: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: rtnl_newlink+0xc10/0x1c30
locks held by syz-executor/32077: 1, on CPU#1:
#0: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: rtnl_newlink+0xc10/0x1c30
locks held by syz-executor/32095: 2, last CPU#1:
#0: ffff88816f5b1d38 (&mm->mmap_lock){++++}-{4:4}, at: exit_mmap+0x1a4/0x9f0
#1: ffffffff8ed5c8a0 (rcu_read_lock){....}-{1:3}, at: __pte_offset_map+0x29/0x240
locks held by syz-executor/32096: 1, on CPU#1:
#0: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: inet_rtm_newaddr+0x470/0x19f0
locks held by syz-executor/32103: 1, on CPU#1:
#0: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: inet_rtm_newaddr+0x470/0x19f0
locks held by syz-executor/32108: 1, on CPU#1:
#0: ffffffff90253600 (rtnl_mutex){+.+.}-{4:4}, at: inet_rtm_newaddr+0x470/0x19f0
locks held by syz-executor/32116: 2, last CPU#1:
#0: ffff88817a26d038 (&mm->mmap_lock){++++}-{4:4}, at: vm_mmap_pgoff+0x1dd/0x4e0
#1: ffffffff8ed5c8a0 (rcu_read_lock){....}-{1:3}, at: __pte_offset_map+0x29/0x240
=============================================
NMI backtrace for cpu 1
CPU: 1 UID: 0 PID: 35 Comm: khungtaskd Not tainted syzkaller #0 PREEMPT(full)
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014
Call Trace:
<TASK>
dump_stack_lvl+0xe8/0x150
nmi_cpu_backtrace+0x274/0x2d0
nmi_trigger_cpumask_backtrace+0x17d/0x390
sys_info+0x135/0x170
watchdog+0xfd7/0x1030
kthread+0x38b/0x480
ret_from_fork+0x514/0xb70
ret_from_fork_asm+0x1a/0x30
</TASK>
Sending NMI from CPU 1 to CPUs 0:
NMI backtrace for cpu 0
CPU: 0 UID: 0 PID: 27790 Comm: kworker/0:1 Not tainted syzkaller #0 PREEMPT(full)
Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-debian-1.16.2-1 04/01/2014
Workqueue: events_long br_fdb_cleanup
RIP: 0010:lock_acquire+0x232/0x350
Code: ff ff ff e8 50 28 42 0a f7 44 24 10 00 02 00 00 0f 84 38 ff ff ff 65 48 8b 05 a2 e3 f1 11 48 3b 44 24 50 75 33 fb 48 83 c4 58 <5b> 41 5c 41 5d 41 5e 41 5f 5d e9 bf 2c 45 0a cc 48 8d 3d 67 4e da
RSP: 0018:ffffc90000006dc0 EFLAGS: 00000296
RAX: 32f582729c8e9900 RBX: 0000000000000000 RCX: 8000000000000100
RDX: 00000000071bee9f RSI: ffffffff8e69ca38 RDI: ffffffff8c6da580
RBP: ffff888112a39e00 R08: ffffffff8177f1df R09: 0000000000000000
R10: ffffc90000006ef8 R11: ffffffff81b28790 R12: ffffffff8ed5c8a0
R13: 0000000000000002 R14: 0000000000000000 R15: 0000000000000246
FS: 0000000000000000(0000) GS:ffff88818d6cf000(0000) knlGS:0000000000000000
CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000200000000040 CR3: 000000000eb48000 CR4: 00000000000006f0
Call Trace:
<IRQ>
unwind_next_frame+0xac/0x2550
arch_stack_walk+0x11b/0x150
stack_trace_save+0xa9/0x100
kasan_save_track+0x3e/0x80
__kasan_slab_alloc+0x6c/0x80
kmem_cache_alloc_node_noprof+0x354/0x610
__alloc_skb+0x1db/0x7a0
synproxy_send_client_synack+0x175/0xde0
nft_synproxy_eval_v4+0x362/0x4e0
nft_synproxy_do_eval+0x335/0x550
nft_do_chain+0x132b/0x1b10
nft_do_chain_ipv4+0x195/0x290
nf_hook_slow+0xc5/0x220
NF_HOOK+0x21f/0x3c0
NF_HOOK+0x336/0x3c0
process_backlog+0xa6b/0x18b0
__napi_poll+0xaa/0x330
net_rx_action+0x61d/0xf50
handle_softirqs+0x226/0x860
__irq_exit_rcu+0xcb/0x220
irq_exit_rcu+0x9/0x30
sysvec_apic_timer_interrupt+0xa6/0xc0
</IRQ>
<TASK>
asm_sysvec_apic_timer_interrupt+0x1a/0x20
RIP: 0010:lock_acquire+0x232/0x350
Code: ff ff ff e8 50 28 42 0a f7 44 24 10 00 02 00 00 0f 84 38 ff ff ff 65 48 8b 05 a2 e3 f1 11 48 3b 44 24 50 75 33 fb 48 83 c4 58 <5b> 41 5c 41 5d 41 5e 41 5f 5d e9 bf 2c 45 0a cc 48 8d 3d 67 4e da
RSP: 0018:ffffc9000141faa0 EFLAGS: 00000296
RAX: 32f582729c8e9900 RBX: 0000000000000001 RCX: 0000000000000046
RDX: 000000004abe915d RSI: ffffffff8e69ca38 RDI: ffffffff8c6da580
RBP: ffff888112a39e00 R08: ffffffff818f9aba R09: 0000000000000000
R10: dffffc0000000000 R11: fffffbfff20f4830 R12: ffffc9000141fc40
R13: 0000000000000000 R14: 0000000000000000 R15: 0000000000000246
process_scheduled_works+0xc03/0x1630
worker_thread+0xa47/0xfb0
kthread+0x38b/0x480
ret_from_fork+0x514/0xb70
ret_from_fork_asm+0x1a/0x30
</TASK>
***
If these findings have caused you to resend the series or submit a
separate fix, please add the following tag to your commit message:
Tested-by: syzbot@syzkaller.appspotmail.com
---
This report is generated by a bot. It may contain errors.
syzbot ci engineers can be reached at syzkaller@googlegroups.com.
To test a fix for this bug, please reply with `#syz test`
(on a separate line) and attach the patch to the email.
Notes:
- The patch will be applied on top of the tested series (as an
incremental fix).
- To test a new version of the whole series, please send it directly
to syzbot@lists.linux.dev.
- Arguments like custom git repos and branches are not supported.
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
2026-09-25 12:42 ` Ido Schimmel
@ 2026-09-28 6:55 ` Chenguang Zhao
2026-09-28 8:10 ` Julian Anastasov
0 siblings, 1 reply; 6+ messages in thread
From: Chenguang Zhao @ 2026-09-28 6:55 UTC (permalink / raw)
To: Ido Schimmel
Cc: davem, edumazet, kuba, pabeni, horms, dsahern, kerneljasonxing,
netdev, Chenguang Zhao, syzbot+2b120190d9e54ad8c65d, ja
On Fri, Sep 25, 2026 at 03:42:42PM +0300, Ido Schimmel wrote:
> On Thu, Sep 24, 2026 at 02:33:12PM +0800, Chenguang Zhao wrote:
> > From: Chenguang Zhao <zhaochenguang@kylinos.cn>
> >
> > Commit de2c211868b9 ("ipvs: Always clear ipvs_property flag in
> > skb_scrub_packet()") moved ipvs_reset() before the xnet check, making
> > the call unconditional. The intent was to fix a bpf_redirect case where
> > stale ipvs_property on an skb re-entering the RX path caused the SNAT
> > hook to be skipped. However the change is too broad: when IPVS NAT
> > sits above an ipvlan L3 interface in the same netns, the following
> > loop happens:
> >
> > LOCAL_OUT -> IPVS DNAT (sets ipvs_property=1) -> dst_output -> ipvlan
> > -> skb_scrub_packet() -> ipvs_reset() clears the flag
> > -> ipvlan_process_v4_outbound() -> ip_local_out() -> LOCAL_OUT
> > -> IPVS sees ipvs_property=0, processes again -> infinite recursion
> >
> > syzbot reported this as a stack overflow on a KASAN kernel where each
> > level burns ~3.3 KB of stack and XMIT_RECURSION_LIMIT falls short. On
> > non-KASAN kernels the dead-loop detector catches it and prints "Dead
> > loop on virtual device", but traffic is still broken.
> >
> > Fixes: de2c211868b9 ("ipvs: Always clear ipvs_property flag in skb_scrub_packet()")
> > Reported-by: syzbot+2b120190d9e54ad8c65d@syzkaller.appspotmail.com
> > Closes: https://syzkaller.appspot.com/bug?extid=2b120190d9e54ad8c65d
> > Signed-off-by: Chenguang Zhao <zhaochenguang@kylinos.cn>
>
> 1. You need to copy IPVS maintainers on patches related to IPVS. Added
> Julian.
Thanks, will add Julian to CC in v2.
>
> 2. Does it reproduce with commit 7f1de03e3103 ("net: reduce
> XMIT_RECURSION_LIMIT under KASAN") in net-next?
The dead loop itself still reproduces — dmesg shows "Dead loop on virtual device" for
every connection, and IPVS InPkts keeps climbing. But the stack overflow crash does not
reproduce with 7f1de03e3103 on KASAN, since the lower recursion limit catches the loop
before the stack blows up. So 7f1de03e3103 turns the panic into a packet drop,
which is expected, but doesn't fix the underlying issue (ipvs_property being cleared
unconditionally in skb_scrub_packet).
Thanks
Chenguang
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
2026-09-24 6:33 [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear Chenguang Zhao
2026-09-25 12:42 ` Ido Schimmel
2026-09-26 20:12 ` [syzbot ci] " syzbot ci
@ 2026-09-28 6:58 ` netdev-bot+sashiko
2 siblings, 0 replies; 6+ messages in thread
From: netdev-bot+sashiko @ 2026-09-28 6:58 UTC (permalink / raw)
To: chenguang.zhao
Cc: davem, edumazet, kuba, pabeni, horms, dsahern, idosch,
kerneljasonxing, netdev, zhaochenguang,
syzbot+2b120190d9e54ad8c65d
Thank you for your contribution! Sashiko AI review found 4 potential
issue(s) to consider.
Critical: 0 · High: 2 · Medium: 1 · Low: 1
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [High] The replacement for de2c211868b9's reset only covers the single-
skb receive functions.
- [High] net: ipvs: SNAT bypass for bridged packets due to rx_handler
interception
- [Medium] The new unconditional ipvs_reset() in ip_rcv()/ipv6_rcv()
clears ipvs_property on every packet delivered through the single-…
- [Low] The recorded commit message does not describe the fix.
--
ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
This patch moves ipvs_reset() in skb_scrub_packet() back under the xnet
check, so ipvs_property is only cleared when an skb crosses a netns. That
fixes an infinite LOCAL_OUT recursion when IPVS NAT sits above ipvlan L3.
To keep the bpf_redirect fix from de2c211868b9, it also adds ipvs_reset()
to ip_rcv() and ipv6_rcv() right before the PRE_ROUTING hook.
Closes: https://syzkaller.appspot.com/bug?extid=2b120190d9e54ad8c65d
> Signed-off-by: Chenguang Zhao <zhaochenguang@kylinos.cn>
> ---
> Fix it with two changes:
>
> 1. Move ipvs_reset() back inside the xnet guard in skb_scrub_packet(),
> so ipvs_property is only cleared when the skb actually crosses a
> netns boundary. This restores IPVS re-entry protection for the
> ipvlan path.
>
> 2. To preserve the bpf_redirect fix, add ipvs_reset() in ip_rcv() and
> ipv6_rcv() right before the NF_HOOK into PREROUTING. Every
> redirected packet enters the stack through these points, so
> clearing ipvs_property there covers the original use case without
> affecting the ipvlan code path.
[Severity: Low]
This isn't a bug, but the "Fix it with two changes" text is below the ---
separator, so it gets dropped when the patch is applied.
The commit recorded in git ends at Signed-off-by. It never mentions the
new ipvs_reset() calls in ip_rcv() and ipv6_rcv().
Could this explanation be moved above the ---?
> diff --git a/net/core/skbuff.c b/net/core/skbuff.c
> index b4edbd06655e..4a5c9b69d8b7 100644
> --- a/net/core/skbuff.c
> +++ b/net/core/skbuff.c
> @@ -6306,11 +6306,10 @@ void skb_scrub_packet(struct sk_buff *skb, bool xnet)
> skb->offload_fwd_mark = 0;
> skb->offload_l3_fwd_mark = 0;
> #endif
> - ipvs_reset(skb);
> -
> if (!xnet)
> return;
>
> + ipvs_reset(skb);
> skb->mark = 0;
> skb_clear_tstamp(skb);
> }
> diff --git a/net/ipv4/ip_input.c b/net/ipv4/ip_input.c
> index 9860178752b8..00f3b328e90a 100644
> --- a/net/ipv4/ip_input.c
> +++ b/net/ipv4/ip_input.c
> @@ -609,6 +609,7 @@ int ip_rcv(struct sk_buff *skb, struct net_device *dev, struct packet_type *pt,
> if (skb == NULL)
> return NET_RX_DROP;
>
> + ipvs_reset(skb);
> return NF_HOOK(NFPROTO_IPV4, NF_INET_PRE_ROUTING,
> net, NULL, skb, dev, NULL,
> ip_rcv_finish);
[Severity: High]
Does this leave a stale flag on packets that never reach ip_rcv()
because a bridge rx_handler consumes them first?
Suppose IPVS has already marked a packet, and the packet is then
redirected into an interface that is a bridge port in the same netns.
Two examples are bpf_redirect() to ingress (__bpf_rx_skb() ->
____dev_forward_skb()) and a veth whose peer is enslaved to a bridge
in the same netns. In both cases the scrub now runs with xnet=false,
so ipvs_property is kept:
____dev_forward_skb()
skb_scrub_packet(skb, false) /* flag kept */
netif_rx()
...
__netif_receive_skb_core()
br_handle_frame() /* RX_HANDLER_CONSUMED */
ip_rcv() and ipv6_rcv() are never called on the port device, so the
new ipvs_reset() calls do not run.
With br_netfilter enabled, br_nf_pre_routing() and br_nf_forward_ip()
still pass the bridged packet through the IPv4/IPv6 netfilter hooks.
The IPVS FORWARD hook then bails out:
if (skb->ipvs_property)
return NF_ACCEPT;
That skips SNAT for the NAT reply. This is the same failure that
de2c211868b9 fixed, now for bridged traffic instead of routed
traffic, and it can leak real server addresses. In the parent tree
the unconditional reset in skb_scrub_packet() covered this case.
Frames the bridge passes up to the host via br_pass_frame_up() do
reach ip_rcv() on the bridge device and get reset. Frames the bridge
forwards do not.
Is relying on ip_rcv()/ipv6_rcv() enough here? Or should the reset
stay in the xnet=false paths that re-inject into RX, such as
____dev_forward_skb(), and be skipped only for the ipvlan outbound
case that causes the recursion?
[Severity: High]
Does this also cover the batched receive path? ip_packet_type and
ipv6_packet_type register .list_func handlers as well, and
__netif_receive_skb_list_ptype() calls those for list delivery:
net/core/dev.c:__netif_receive_skb_list_ptype() {
...
if (pt_prev->list_func != NULL)
INDIRECT_CALL_INET(pt_prev->list_func, ipv6_list_rcv,
ip_list_rcv, head, pt_prev, orig_dev);
...
}
ip_list_rcv() calls ip_rcv_core() and then enters PRE_ROUTING through
ip_sublist_rcv() without calling ipvs_reset():
net/ipv4/ip_input.c:ip_sublist_rcv() {
NF_HOOK_LIST(NFPROTO_IPV4, NF_INET_PRE_ROUTING, net, NULL,
head, dev, NULL, ip_rcv_finish);
...
}
ipv6_list_rcv()->ip6_sublist_rcv() behaves the same way.
Now that skb_scrub_packet() only resets the flag when xnet is true, it
looks like the stale flag survives this path:
IPVS NAT in netns N sets skb->ipvs_property = 1
veth_xmit() /* peer also in netns N */
veth_forward_skb()
__dev_forward_skb()
____dev_forward_skb()
skb_scrub_packet(skb, false) /* flag kept */
veth_poll() /* peer has GRO or XDP, use_napi */
napi_gro_receive()
... netif_receive_skb_list_internal()
ip_list_rcv() /* ipvs_property still 1 */
Then ip_vs_out_hook() returns early:
if (skb->ipvs_property)
return NF_ACCEPT;
ip_vs_in_hook() does the same. That brings back the skipped SNAT that
de2c211868b9 fixed, and whether it happens depends on the receiving
device's GRO/XDP state.
skb_do_redirect() with BPF_F_PEER looks affected as well:
net/core/filter.c:skb_do_redirect() {
...
if (unlikely(!dev ||
!(dev->flags & IFF_UP) ||
net_eq(net, dev_net(dev))))
goto out_drop;
skb_scrub_packet(skb, false);
...
}
This path only runs when the target is in another netns, yet it scrubs
with xnet=false. As a result, ipvs_property now crosses the netns
boundary and can reach ip_list_rcv() in the target netns.
The notes say "Every redirected packet enters the stack through these
points", but that doesn't seem to be true for NAPI/GRO list delivery.
Would it work better to do the reset in ip_rcv_core()/ip6_rcv_core(), to
add it to both list handlers, or to clear the flag in the xnet=false
forwarders that re-inject into RX?
> diff --git a/net/ipv6/ip6_input.c b/net/ipv6/ip6_input.c
> index d332ec60f915..05917095ef6d 100644
> --- a/net/ipv6/ip6_input.c
> +++ b/net/ipv6/ip6_input.c
> @@ -348,6 +348,7 @@ int ipv6_rcv(struct sk_buff *skb, struct net_device *dev, struct packet_type *pt
> skb = ip6_rcv_core(skb, dev, net);
> if (skb == NULL)
> return NET_RX_DROP;
> + ipvs_reset(skb);
> return NF_HOOK(NFPROTO_IPV6, NF_INET_PRE_ROUTING,
> net, NULL, skb, dev, NULL,
> ip6_rcv_finish);
[Severity: Medium]
Can the unconditional clear here, and the one in ip_rcv() above, drop
an ipvs_property that IPVS set on purpose?
Take a local client connecting to a VIP whose DR real server is on the
same host:
LOCAL_OUT ip_vs_in_hook()
ip_vs_dr_xmit() /* local route */
ip_vs_send_or_cont(..., 1)
skb->ipvs_property = 1; /* returns NF_ACCEPT */
loopback_xmit() /* no scrub */
__netif_rx()
process_backlog()
ip_rcv()
ipvs_reset(skb) /* flag cleared */
LOCAL_IN ip_vs_in_hook()
At LOCAL_IN the early return no longer fires:
/* Already marked as IPVS request or reply? */
if (skb->ipvs_property)
return NF_ACCEPT;
So ip_vs_in_stats(), ip_vs_set_state() and the in_pkts update run a
second time for the same packet:
ip_vs_in_stats(cp, skb);
ip_vs_set_state(cp, IP_VS_DIR_INPUT, skb, pd, iph.len);
...
pkts = atomic_inc_return(&cp->in_pkts);
if (ipvs->sync_state & IP_VS_STATE_MASTER)
ip_vs_sync_conn(ipvs, cp, pkts);
xt_ipvs --ipvs rules at INPUT would also stop matching these packets. In
the parent tree the flag survived loopback, and LOCAL_IN skipped them.
There is a similar concern for ignore_tunneled. ip_vs_in_hook() marks the
outer packet:
/* The only way we'll see this packet again is if it's
* encapsulated, so mark it with ipvs_property=1 so we
* skip it if we're ignoring tunneled packets
*/
if (sysctl_ignore_tunneled(ipvs))
skb->ipvs_property = 1;
With this patch, ip_tunnel_rcv() keeps the mark because its scrub uses
xnet=false. It then calls gro_cells_receive(), which falls back to
netif_rx() in three cases:
- there are no cells
- skb_cloned(skb) is true, for example when a packet tap is active
- netif_elide_gro(dev) is true
The netif_rx() path reaches ip_rcv() and clears the mark. The GRO path
reaches ip_list_rcv() and keeps it. So whether the sysctl works now
depends on GRO/XDP/tap state.
The parent tree already cleared the mark in ip_tunnel_rcv() through the
unconditional scrub from de2c211868b9, so the ignore_tunneled part is not
new. This patch also doesn't restore the pre-de2c211868b9 behaviour for
it.
Is there a way to clear only stale flags here, instead of every packet
that comes in through ip_rcv() and ipv6_rcv()?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260924063312.1194019-1-chenguang.zhao%40linux.dev
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear
2026-09-28 6:55 ` Chenguang Zhao
@ 2026-09-28 8:10 ` Julian Anastasov
0 siblings, 0 replies; 6+ messages in thread
From: Julian Anastasov @ 2026-09-28 8:10 UTC (permalink / raw)
To: Chenguang Zhao
Cc: Ido Schimmel, davem, edumazet, kuba, pabeni, horms, dsahern,
kerneljasonxing, netdev, Chenguang Zhao,
syzbot+2b120190d9e54ad8c65d
[-- Attachment #1: Type: text/plain, Size: 3031 bytes --]
Hello,
On Mon, 28 Sep 2026, Chenguang Zhao wrote:
> On Fri, Sep 25, 2026 at 03:42:42PM +0300, Ido Schimmel wrote:
> > On Thu, Sep 24, 2026 at 02:33:12PM +0800, Chenguang Zhao wrote:
> > > From: Chenguang Zhao <zhaochenguang@kylinos.cn>
> > >
> > > Commit de2c211868b9 ("ipvs: Always clear ipvs_property flag in
> > > skb_scrub_packet()") moved ipvs_reset() before the xnet check, making
> > > the call unconditional. The intent was to fix a bpf_redirect case where
> > > stale ipvs_property on an skb re-entering the RX path caused the SNAT
> > > hook to be skipped. However the change is too broad: when IPVS NAT
> > > sits above an ipvlan L3 interface in the same netns, the following
> > > loop happens:
> > >
> > > LOCAL_OUT -> IPVS DNAT (sets ipvs_property=1) -> dst_output -> ipvlan
> > > -> skb_scrub_packet() -> ipvs_reset() clears the flag
> > > -> ipvlan_process_v4_outbound() -> ip_local_out() -> LOCAL_OUT
> > > -> IPVS sees ipvs_property=0, processes again -> infinite recursion
> > >
> > > syzbot reported this as a stack overflow on a KASAN kernel where each
> > > level burns ~3.3 KB of stack and XMIT_RECURSION_LIMIT falls short. On
> > > non-KASAN kernels the dead-loop detector catches it and prints "Dead
> > > loop on virtual device", but traffic is still broken.
> > >
> > > Fixes: de2c211868b9 ("ipvs: Always clear ipvs_property flag in skb_scrub_packet()")
> > > Reported-by: syzbot+2b120190d9e54ad8c65d@syzkaller.appspotmail.com
> > > Closes: https://syzkaller.appspot.com/bug?extid=2b120190d9e54ad8c65d
> > > Signed-off-by: Chenguang Zhao <zhaochenguang@kylinos.cn>
> >
> > 1. You need to copy IPVS maintainers on patches related to IPVS. Added
> > Julian.
>
> Thanks, will add Julian to CC in v2.
I saw it the first time, not needed. But I was busy
with other changes, sorry for that. May be the packet scrubing
needs 3 states, not just 2. But as this is not my area I have
to dig more. SCRUB_REROUTE/SCRUB_TUNNEL (xdev=false, no ipvs_reset)
used by IPVLAN L3, SCRUB_FORWARD (xdev=false, ipvs_reset) used
by bpf_redirect in same netns, SCRUB_NET (xdev=true). But there
are many call sites to check. The problem is to know how far
a packet can reach without reset.
Another option is to add new skb_scrub_packet_XXX() for the
needed sites, I'm not sure. But adding IPVS code in the IP layer
does not look nice.
> >
> > 2. Does it reproduce with commit 7f1de03e3103 ("net: reduce
> > XMIT_RECURSION_LIMIT under KASAN") in net-next?
>
> The dead loop itself still reproduces — dmesg shows "Dead loop on virtual device" for
> every connection, and IPVS InPkts keeps climbing. But the stack overflow crash does not
> reproduce with 7f1de03e3103 on KASAN, since the lower recursion limit catches the loop
> before the stack blows up. So 7f1de03e3103 turns the panic into a packet drop,
> which is expected, but doesn't fix the underlying issue (ipvs_property being cleared
> unconditionally in skb_scrub_packet).
>
> Thanks
> Chenguang
Regards
--
Julian Anastasov <ja@ssi.bg>
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-28 8:11 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 6:33 [PATCH net] ipvs: fix infinite loop with ipvlan L3 from unconditional ipvs_property clear Chenguang Zhao
2026-09-25 12:42 ` Ido Schimmel
2026-09-28 6:55 ` Chenguang Zhao
2026-09-28 8:10 ` Julian Anastasov
2026-09-26 20:12 ` [syzbot ci] " syzbot ci
2026-09-28 6:58 ` [PATCH net] " netdev-bot+sashiko
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox