Netdev List
 help / color / mirror / Atom feed
From: Ilya Maximets <i.maximets@ovn.org>
To: "Denis V. Lunev" <den@openvz.org>, netdev@vger.kernel.org
Cc: dev@openvswitch.org, Aaron Conole <aconole@redhat.com>,
	Eelco Chaudron <echaudro@redhat.com>,
	Ilya Maximets <i.maximets@ovn.org>,
	"David S. Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@google.com>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>,
	stable@vger.kernel.org
Subject: Re: [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump
Date: Fri, 2 Oct 2026 22:50:42 +0200	[thread overview]
Message-ID: <892a3d9c-65f7-4216-95c2-e55938254e96@ovn.org> (raw)
In-Reply-To: <20260929072519.2803304-2-den@openvz.org>

On 9/29/26 9:25 AM, Denis V. Lunev wrote:
> From: Denis V. Lunev <den@openvz.org>
> 
> A production compute node carrying a few thousand datapath flows hit a
> soft lockup inside a single netlink flow dump and panicked.
> 
> ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it
> emits, and that releases stats->lock with spin_unlock_bh() once per
> CPU that has touched the flow. Every release is a local_bh_enable(),
> and each one runs the pending softirq backlog in the dumping thread's
> own context.
> 
> The skb bounds the flows one callback emits, but not the softirq work
> it absorbs. On a CPU that carries the box's packet load the backlog
> refills as fast as it drains, so the dumping thread becomes that CPU's
> softirq engine. It never sleeps and it has no reschedule point, so
> under voluntary preemption nothing can take the CPU away from it:
> neither the ksoftirqd the kernel woke to take the work over, nor the
> stopper thread the softlockup detector dispatches to refresh its
> timestamp.
> 
> Hold BH off across the whole callback instead, the way
> ctnetlink_dump_table() does, so the nested spin_unlock_bh() stop
> draining softirqs. The loop already runs under rcu_read_lock() and
> cannot sleep. What it gives up is preemption under CONFIG_PREEMPT,
> since a BH-off region is not preemptible outside PREEMPT_RT. The
> region stays short: the skb caps the flows one callback emits, and
> empty buckets cost no skb space but are each visited once per dump, as
> the cursor only moves forward. The table grows on insert and shrinks
> only on flush, so the walk is bounded by the largest flow count the
> datapath has held. The softirq backlog the callback used to absorb has
> no bound at all.
> 
> Fixes: 63e7959c4b9b ("openvswitch: Per NUMA node flow stats.")
> Cc: stable@vger.kernel.org
> Signed-off-by: Denis V. Lunev <den@openvz.org>
> ---
> v3:
> - add the net prefix, Fixes tag and Cc stable
> - explain why the BH-off walk stays bounded: the cursor visits each
>   empty bucket once per dump and the table shrinks only on flush
> - drop "here" from the comment, add blank lines around the
>   local_bh_disable()/local_bh_enable() pair
> v2: https://lore.kernel.org/netdev/20260915122401.3910188-1-den@openvz.org/
> - leave ovs_vport_cmd_dump() alone: nsid_lock has not been BH-safe
>   since commit aed4969f2bdf ("net: net->nsid_lock does not need BH
>   safety"), so the vport dump never drained softirqs
> - disable BH before the table dereference and say in a comment that
>   the region is not there for safety
> - drop the ovs_flow_stats_get() history, note the empty-bucket walk
>   and the lost CONFIG_PREEMPT preemption in the message
> - move the Cc list out of the commit message
> v1: https://lore.kernel.org/netdev/20260909171920.1001074-1-den@openvz.org/
> 
>  net/openvswitch/datapath.c | 8 ++++++++
>  1 file changed, 8 insertions(+)
> 
> diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
> index 631a03136fa1..4fc5d0bebd85 100644
> --- a/net/openvswitch/datapath.c
> +++ b/net/openvswitch/datapath.c
> @@ -1532,6 +1532,12 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
>  		return -ENODEV;
>  	}
>  
> +	/*
> +	 * Not needed for safety. Stops every spin_unlock_bh() in
> +	 * ovs_flow_stats_get() from running the softirq backlog.
> +	 */
> +	local_bh_disable();
> +
>  	ti = rcu_dereference(dp->table.ti);
>  	for (;;) {
>  		struct sw_flow *flow;
> @@ -1552,6 +1558,8 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
>  		cb->args[0] = bucket;
>  		cb->args[1] = obj;
>  	}
> +
> +	local_bh_enable();
>  	rcu_read_unlock();
>  	return skb->len;
>  }

Reviewed-by: Ilya Maximets <i.maximets@ovn.org>

  parent reply	other threads:[~2026-10-02 20:50 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-29  7:25 [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump Denis V. Lunev
2026-09-29  7:25 ` Denis V. Lunev
2026-10-01 22:27   ` netdev-bot+sashiko
2026-10-02 20:41     ` Ilya Maximets
2026-10-02 20:50   ` Ilya Maximets [this message]
2026-10-05  3:00   ` patchwork-bot+netdevbpf

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=892a3d9c-65f7-4216-95c2-e55938254e96@ovn.org \
    --to=i.maximets@ovn.org \
    --cc=aconole@redhat.com \
    --cc=davem@davemloft.net \
    --cc=den@openvz.org \
    --cc=dev@openvswitch.org \
    --cc=echaudro@redhat.com \
    --cc=edumazet@google.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox