From: Ilya Maximets <i.maximets@ovn.org>
To: Ilya Maximets <i.maximets@ovn.org>, netdev@vger.kernel.org
Cc: Aaron Conole <aconole@redhat.com>,
Eelco Chaudron <echaudro@redhat.com>,
"David S. Miller" <davem@davemloft.net>,
Eric Dumazet <edumazet@google.com>,
Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
Simon Horman <horms@kernel.org>,
dev@openvswitch.org, linux-kernel@vger.kernel.org,
stable@vger.kernel.org
Subject: Re: [PATCH net] net: openvswitch: fix flow mask use-after-free on flow deletion
Date: Mon, 17 Aug 2026 20:37:24 +0200 [thread overview]
Message-ID: <2a44f5b3-d0cb-4333-829d-6a1f77c811ff@ovn.org> (raw)
In-Reply-To: <20260815005915.1097270-1-i.maximets@ovn.org>
On 8/15/26 2:58 AM, Ilya Maximets wrote:
> The commit in the Fixes tag below made so flow->mask free is scheduled
> via RCU right after it is removed from the flow table. The pointer
> stays in the flow structure and it can be accessible while in the same
> RCU critical section. This is done to avoid requiring ovs_mutex for
> the ovs_flow_free().
>
> However, while removing the flow during processing of CMD_DEL, we do
> not take RCU read lock before the removal, and ovs_flow_cmd_fill_info()
> uses the flow->mask pointer afterwards. The RCU read lock is taken,
> but it's already late at that point. The comment on that line
> acknowledges that the lock is cosmetic and doesn't serve a real purpose.
>
> This leads to use-after-free if the RCU grace period passes between
> removal and the filling. It is a short race window, but it is there
> and can lead to a real crash in case memory allocation for the info
> takes a bit longer:
>
> BUG: KASAN: slab-use-after-free in __ovs_nla_put_key
> net/openvswitch/flow_netlink.c:1996
> BUG: KASAN: slab-use-after-free in ovs_nla_put_key+0x2463/0x2e30
> net/openvswitch/flow_netlink.c:2250
> Read of size 4 at addr ffff88801ee89970 by task ovs_flow_del_ec/9487
>
> Call Trace:
> <TASK>
> __ovs_nla_put_key net/openvswitch/flow_netlink.c:1996
> ovs_nla_put_key+0x2463/0x2e30 net/openvswitch/flow_netlink.c:2250
> ovs_flow_cmd_fill_info+0x420/0x9c0 net/openvswitch/datapath.c:930
> ovs_flow_cmd_del+0x53a/0x970 net/openvswitch/datapath.c:1467
> ...
> netlink_rcv_skb+0x156/0x420 net/netlink/af_netlink.c:2556
> </TASK>
>
> Allocated by task 9487:
> mask_alloc net/openvswitch/flow_table.c:967
> flow_mask_insert net/openvswitch/flow_table.c:1012
> ovs_flow_tbl_insert+0xea2/0x1a90 net/openvswitch/flow_table.c:1084
> ovs_flow_cmd_new+0x7e3/0xd90 net/openvswitch/datapath.c:1086
> ...
> netlink_rcv_skb+0x156/0x420 net/netlink/af_netlink.c:2556
>
> Freed by task 9485:
> rcu_free_sheaf+0x1e/0x100 mm/slub.c:5978
> rcu_do_batch kernel/rcu/tree.c:2645
> rcu_core+0x59c/0x10c0 kernel/rcu/tree.c:2897
> handle_softirqs+0x1e4/0x9a0 kernel/softirq.c:622
> ...
> instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1062
>
> ovs_flow_tbl_remove() must be called after the ovs_flow_cmd_fill_info()
> to avoid this race. This also helps with cleaning up the forced cast
> and the cosmetic RCU read lock. Before the commit in the Fixes tag the
> order did not matter as long as the flow object itself was not freed.
>
> A wider RCU critical section could be another option, but we have a
> GFP_KERNEL allocation in the way.
>
> Reported by Trend Micro's Zero Day Initiative as ZDI-CAN-32042.
>
> Fixes: 56c19868e115 ("openvswitch: Make flow mask removal symmetric.")
> Cc: stable@vger.kernel.org
> Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
> ---
Sashiko complains:
> Does this reordering drop the final packet and byte counts for packets
> that hit the flow between the snapshot and the unlink?
> ovs_flow_cmd_fill_info() -> ovs_flow_cmd_fill_stats() -> ovs_flow_stats_get()
> now serializes the counters while the flow is still linked in dp->table,
> and ovs_flow_tbl_remove() only runs afterwards. The datapath writer runs
> in softirq context and takes only the per-CPU stats lock, never ovs_mutex:
> net/openvswitch/flow.c:ovs_flow_stats_update() {
> ...
> stats = rcu_dereference(flow->stats[cpu]);
> /* Check if already have CPU-specific stats. */
> if (likely(stats)) {
> spin_lock(&stats->lock);
> ...
> stats->packet_count++;
> stats->byte_count += len;
> ...
> }
> So ovs_dp_process_packet() -> ovs_flow_tbl_lookup_stats() still finds the
> flow and bumps flow->stats[cpu] during that window. Those increments are
> then discarded by:
> ovs_flow_free(flow, true);
> Since OVS_FLOW_ATTR_STATS in the DEL reply/notification is the last place
> user space can collect a flow's final counters, would those packets be
> lost from accounting? Before the patch the unlink preceded the snapshot,
> so no new lookup could match the flow after the counters were read.
> The window here is bounded by the remaining nla_put work in
> ovs_flow_cmd_fill_actions() plus any preemption of the deleting task, not
> by a sleeping allocation, since ovs_flow_cmd_alloc_info() with GFP_KERNEL
> now runs before ovs_flow_cmd_fill_info(). Would it be worth mentioning
> this trade-off in the commit message?
This is not a new issue. The race window is a bit different, but it was
there before the change. The datapath processing is only protected by RCU
and we're not synchronizing it between removal and reading the stats.
So, there will always be a chance to not account for some of the packets.
That said, this is also not a concern for any real setup as ovs-vswitchd
doesn't delete active flows under normal circumstances.
Best regards, Ilya Maximets.
next prev parent reply other threads:[~2026-08-17 18:37 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-15 0:58 [PATCH net] net: openvswitch: fix flow mask use-after-free on flow deletion Ilya Maximets
2026-08-16 14:52 ` Aaron Conole
2026-08-17 18:37 ` Ilya Maximets [this message]
2026-08-18 17:10 ` patchwork-bot+netdevbpf
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=2a44f5b3-d0cb-4333-829d-6a1f77c811ff@ovn.org \
--to=i.maximets@ovn.org \
--cc=aconole@redhat.com \
--cc=davem@davemloft.net \
--cc=dev@openvswitch.org \
--cc=echaudro@redhat.com \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox