* [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump
@ 2026-09-29 7:25 Denis V. Lunev
2026-09-29 7:25 ` Denis V. Lunev
0 siblings, 1 reply; 6+ messages in thread
From: Denis V. Lunev @ 2026-09-29 7:25 UTC (permalink / raw)
To: netdev
Cc: dev, Aaron Conole, Eelco Chaudron, Ilya Maximets, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
Denis V. Lunev, stable
From: Denis V. Lunev <den@openvz.org>
A production compute node carrying a few thousand datapath flows hit a
soft lockup inside a single netlink flow dump and panicked.
ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it
emits, and that releases stats->lock with spin_unlock_bh() once per
CPU that has touched the flow. Every release is a local_bh_enable(),
and each one runs the pending softirq backlog in the dumping thread's
own context.
The skb bounds the flows one callback emits, but not the softirq work
it absorbs. On a CPU that carries the box's packet load the backlog
refills as fast as it drains, so the dumping thread becomes that CPU's
softirq engine. It never sleeps and it has no reschedule point, so
under voluntary preemption nothing can take the CPU away from it:
neither the ksoftirqd the kernel woke to take the work over, nor the
stopper thread the softlockup detector dispatches to refresh its
timestamp.
Hold BH off across the whole callback instead, the way
ctnetlink_dump_table() does, so the nested spin_unlock_bh() stop
draining softirqs. The loop already runs under rcu_read_lock() and
cannot sleep. What it gives up is preemption under CONFIG_PREEMPT,
since a BH-off region is not preemptible outside PREEMPT_RT. The
region stays short: the skb caps the flows one callback emits, and
empty buckets cost no skb space but are each visited once per dump, as
the cursor only moves forward. The table grows on insert and shrinks
only on flush, so the walk is bounded by the largest flow count the
datapath has held. The softirq backlog the callback used to absorb has
no bound at all.
Fixes: 63e7959c4b9b ("openvswitch: Per NUMA node flow stats.")
Cc: stable@vger.kernel.org
Signed-off-by: Denis V. Lunev <den@openvz.org>
---
v3:
- add the net prefix, Fixes tag and Cc stable
- explain why the BH-off walk stays bounded: the cursor visits each
empty bucket once per dump and the table shrinks only on flush
- drop "here" from the comment, add blank lines around the
local_bh_disable()/local_bh_enable() pair
v2: https://lore.kernel.org/netdev/20260915122401.3910188-1-den@openvz.org/
- leave ovs_vport_cmd_dump() alone: nsid_lock has not been BH-safe
since commit aed4969f2bdf ("net: net->nsid_lock does not need BH
safety"), so the vport dump never drained softirqs
- disable BH before the table dereference and say in a comment that
the region is not there for safety
- drop the ovs_flow_stats_get() history, note the empty-bucket walk
and the lost CONFIG_PREEMPT preemption in the message
- move the Cc list out of the commit message
v1: https://lore.kernel.org/netdev/20260909171920.1001074-1-den@openvz.org/
net/openvswitch/datapath.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
index 631a03136fa1..4fc5d0bebd85 100644
--- a/net/openvswitch/datapath.c
+++ b/net/openvswitch/datapath.c
@@ -1532,6 +1532,12 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
return -ENODEV;
}
+ /*
+ * Not needed for safety. Stops every spin_unlock_bh() in
+ * ovs_flow_stats_get() from running the softirq backlog.
+ */
+ local_bh_disable();
+
ti = rcu_dereference(dp->table.ti);
for (;;) {
struct sw_flow *flow;
@@ -1552,6 +1558,8 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
cb->args[0] = bucket;
cb->args[1] = obj;
}
+
+ local_bh_enable();
rcu_read_unlock();
return skb->len;
}
--
2.53.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump
2026-09-29 7:25 [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump Denis V. Lunev
@ 2026-09-29 7:25 ` Denis V. Lunev
2026-10-01 22:27 ` netdev-bot+sashiko
` (2 more replies)
0 siblings, 3 replies; 6+ messages in thread
From: Denis V. Lunev @ 2026-09-29 7:25 UTC (permalink / raw)
To: netdev
Cc: dev, Aaron Conole, Eelco Chaudron, Ilya Maximets, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman,
Denis V. Lunev, stable
From: Denis V. Lunev <den@openvz.org>
A production compute node carrying a few thousand datapath flows hit a
soft lockup inside a single netlink flow dump and panicked.
ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it
emits, and that releases stats->lock with spin_unlock_bh() once per
CPU that has touched the flow. Every release is a local_bh_enable(),
and each one runs the pending softirq backlog in the dumping thread's
own context.
The skb bounds the flows one callback emits, but not the softirq work
it absorbs. On a CPU that carries the box's packet load the backlog
refills as fast as it drains, so the dumping thread becomes that CPU's
softirq engine. It never sleeps and it has no reschedule point, so
under voluntary preemption nothing can take the CPU away from it:
neither the ksoftirqd the kernel woke to take the work over, nor the
stopper thread the softlockup detector dispatches to refresh its
timestamp.
Hold BH off across the whole callback instead, the way
ctnetlink_dump_table() does, so the nested spin_unlock_bh() stop
draining softirqs. The loop already runs under rcu_read_lock() and
cannot sleep. What it gives up is preemption under CONFIG_PREEMPT,
since a BH-off region is not preemptible outside PREEMPT_RT. The
region stays short: the skb caps the flows one callback emits, and
empty buckets cost no skb space but are each visited once per dump, as
the cursor only moves forward. The table grows on insert and shrinks
only on flush, so the walk is bounded by the largest flow count the
datapath has held. The softirq backlog the callback used to absorb has
no bound at all.
Fixes: 63e7959c4b9b ("openvswitch: Per NUMA node flow stats.")
Cc: stable@vger.kernel.org
Signed-off-by: Denis V. Lunev <den@openvz.org>
---
v3:
- add the net prefix, Fixes tag and Cc stable
- explain why the BH-off walk stays bounded: the cursor visits each
empty bucket once per dump and the table shrinks only on flush
- drop "here" from the comment, add blank lines around the
local_bh_disable()/local_bh_enable() pair
v2: https://lore.kernel.org/netdev/20260915122401.3910188-1-den@openvz.org/
- leave ovs_vport_cmd_dump() alone: nsid_lock has not been BH-safe
since commit aed4969f2bdf ("net: net->nsid_lock does not need BH
safety"), so the vport dump never drained softirqs
- disable BH before the table dereference and say in a comment that
the region is not there for safety
- drop the ovs_flow_stats_get() history, note the empty-bucket walk
and the lost CONFIG_PREEMPT preemption in the message
- move the Cc list out of the commit message
v1: https://lore.kernel.org/netdev/20260909171920.1001074-1-den@openvz.org/
net/openvswitch/datapath.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
index 631a03136fa1..4fc5d0bebd85 100644
--- a/net/openvswitch/datapath.c
+++ b/net/openvswitch/datapath.c
@@ -1532,6 +1532,12 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
return -ENODEV;
}
+ /*
+ * Not needed for safety. Stops every spin_unlock_bh() in
+ * ovs_flow_stats_get() from running the softirq backlog.
+ */
+ local_bh_disable();
+
ti = rcu_dereference(dp->table.ti);
for (;;) {
struct sw_flow *flow;
@@ -1552,6 +1558,8 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
cb->args[0] = bucket;
cb->args[1] = obj;
}
+
+ local_bh_enable();
rcu_read_unlock();
return skb->len;
}
--
2.53.0
^ permalink raw reply related [flat|nested] 6+ messages in thread
* Re: [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump
2026-09-29 7:25 ` Denis V. Lunev
@ 2026-10-01 22:27 ` netdev-bot+sashiko
2026-10-02 20:41 ` Ilya Maximets
2026-10-02 20:50 ` Ilya Maximets
2026-10-05 3:00 ` patchwork-bot+netdevbpf
2 siblings, 1 reply; 6+ messages in thread
From: netdev-bot+sashiko @ 2026-10-01 22:27 UTC (permalink / raw)
To: den
Cc: netdev, dev, aconole, echaudro, i.maximets, davem, edumazet, kuba,
pabeni, horms, stable
Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.
Critical: 0 · High: 0 · Medium: 1 · Low: 0
If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).
- [Medium] ovs_flow_cmd_dump() (net/openvswitch/datapath.c) now holds BH
disabled across the whole dump callback.
--
openvswitch: fix soft lockup in the netlink flow dump
This patch disables BH across the whole table walk in ovs_flow_cmd_dump().
The goal is to stop the per-CPU spin_unlock_bh() calls in
ovs_flow_stats_get() from running the softirq backlog in the dumping
thread's context. That backlog processing led to a soft lockup on a
heavily loaded compute node.
> diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
> index 631a03136fa14..4fc5d0bebd85e 100644
> --- a/net/openvswitch/datapath.c
> +++ b/net/openvswitch/datapath.c
> @@ -1532,6 +1532,12 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
> return -ENODEV;
> }
>
> + /*
> + * Not needed for safety. Stops every spin_unlock_bh() in
> + * ovs_flow_stats_get() from running the softirq backlog.
> + */
> + local_bh_disable();
> +
> ti = rcu_dereference(dp->table.ti);
> for (;;) {
> struct sw_flow *flow;
[ ... ]
> @@ -1552,6 +1558,8 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
> cb->args[0] = bucket;
> cb->args[1] = obj;
> }
> +
> + local_bh_enable();
> rcu_read_unlock();
> return skb->len;
> }
[Severity: Medium]
Can a namespace user control how long this BH-off region lasts?
The region now covers the whole walk in ovs_flow_tbl_dump_next(). When
buckets are empty, that loop has no exit that depends on skb space:
net/openvswitch/flow_table.c:ovs_flow_tbl_dump_next() {
...
while (*bucket < ti->n_buckets) {
i = 0;
head = &ti->buckets[*bucket];
hlist_for_each_entry_rcu(flow, head, flow_table.node[ver]) {
...
}
(*bucket)++;
*last = 0;
}
...
}
This means one callback can scan the entire bucket array with softirqs held
off.
The only cap on n_buckets is the kvmalloc limit, which allows about 2^27
buckets:
net/openvswitch/flow_table.c:flow_key_insert() {
...
if (table->count > ti->n_buckets)
new_ti = table_instance_expand(ti, false);
else if (time_after(jiffies, table->last_rehash + REHASH_INTERVAL))
new_ti = table_instance_rehash(ti, ti->n_buckets, false);
...
}
The table never gets smaller in normal use:
- ovs_flow_tbl_remove() never shrinks it.
- The periodic rehash keeps the same size.
- Only ovs_flow_tbl_flush() goes back to TBL_MIN_BUCKETS.
OVS_DP_CMD_NEW and OVS_FLOW_CMD_NEW/DEL only need GENL_UNS_ADMIN_PERM. The
dump needs no privilege at all:
net/openvswitch/datapath.c:dp_flow_genl_ops[] {
...
{ .cmd = OVS_FLOW_CMD_GET,
.validate = GENL_DONT_VALIDATE_STRICT | GENL_DONT_VALIDATE_DUMP,
.flags = 0, /* OK for unprivileged users. */
.doit = ovs_flow_cmd_get,
.dumpit = ovs_flow_cmd_dump
},
...
}
Root in an unprivileged user+net namespace could do the following:
1. Create a datapath.
2. Insert N flows so that n_buckets grows to between N and 2N.
3. Delete every flow except one in a high bucket.
4. Run the dump repeatedly from threads pinned to each CPU.
Each dump could then walk up to 1 GB of hlist_heads. During that walk,
NET_RX, timers, tasklets and RCU callbacks are all deferred on that CPU. On
CONFIG_PREEMPT (non-RT) kernels, the walk also can't be preempted.
Before this patch, softirqs could still run on hardirq exit during the walk,
and PREEMPT_RCU kernels could preempt it. On PREEMPT_NONE/VOLUNTARY kernels
the walk was already non-preemptible under rcu_read_lock(), so the only new
cost there is the softirq deferral.
Container memcg limits don't cap how far the table can grow:
- The sw_flow kmem_cache is created in ovs_flow_init() without
SLAB_ACCOUNT.
- The bucket array from table_instance_alloc() comes from unaccounted
kvmalloc memory.
The commit message says:
The region stays short: the skb caps the flows one callback emits, and
empty buckets cost no skb space but are each visited once per dump, as
the cursor only moves forward. The table grows on insert and shrinks
only on flush, so the walk is bounded by the largest flow count the
datapath has held.
That bound is real, but a namespace user chooses it. Is "stays short"
accurate here?
Two ways to keep the soft lockup fix without an unbounded BH-off window:
- Stop after a fixed number of buckets per callback and resume from
cb->args.
- Re-enable and disable BH after each batch of buckets.
Would either of these work?
--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260929072519.2803304-2-den%40openvz.org
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump
2026-10-01 22:27 ` netdev-bot+sashiko
@ 2026-10-02 20:41 ` Ilya Maximets
0 siblings, 0 replies; 6+ messages in thread
From: Ilya Maximets @ 2026-10-02 20:41 UTC (permalink / raw)
To: netdev-bot+sashiko, den
Cc: netdev, dev, aconole, echaudro, i.maximets, davem, edumazet, kuba,
pabeni, horms, stable
On 10/2/26 12:27 AM, netdev-bot+sashiko@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential
> issue(s) to consider.
>
> Critical: 0 · High: 0 · Medium: 1 · Low: 0
>
> If you believe none of the issues are legitimate, please reply to each
> issue of Medium severity or higher and explain why the AI is wrong.
> If any of them are legitimate and you plan to respin, please let patchwork
> know by including "pw-bot: cr" as a separate line at the end of your reply
> (one such reply per series is enough).
>
> - [Medium] ovs_flow_cmd_dump() (net/openvswitch/datapath.c) now holds BH
> disabled across the whole dump callback.
>
> --
>
> openvswitch: fix soft lockup in the netlink flow dump
>
> This patch disables BH across the whole table walk in ovs_flow_cmd_dump().
> The goal is to stop the per-CPU spin_unlock_bh() calls in
> ovs_flow_stats_get() from running the softirq backlog in the dumping
> thread's context. That backlog processing led to a soft lockup on a
> heavily loaded compute node.
>
>> diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
>> index 631a03136fa14..4fc5d0bebd85e 100644
>> --- a/net/openvswitch/datapath.c
>> +++ b/net/openvswitch/datapath.c
>> @@ -1532,6 +1532,12 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
>> return -ENODEV;
>> }
>>
>> + /*
>> + * Not needed for safety. Stops every spin_unlock_bh() in
>> + * ovs_flow_stats_get() from running the softirq backlog.
>> + */
>> + local_bh_disable();
>> +
>> ti = rcu_dereference(dp->table.ti);
>> for (;;) {
>> struct sw_flow *flow;
>
> [ ... ]
>
>> @@ -1552,6 +1558,8 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
>> cb->args[0] = bucket;
>> cb->args[1] = obj;
>> }
>> +
>> + local_bh_enable();
>> rcu_read_unlock();
>> return skb->len;
>> }
>
> [Severity: Medium]
> Can a namespace user control how long this BH-off region lasts?
>
> The region now covers the whole walk in ovs_flow_tbl_dump_next(). When
> buckets are empty, that loop has no exit that depends on skb space:
>
> net/openvswitch/flow_table.c:ovs_flow_tbl_dump_next() {
> ...
> while (*bucket < ti->n_buckets) {
> i = 0;
> head = &ti->buckets[*bucket];
> hlist_for_each_entry_rcu(flow, head, flow_table.node[ver]) {
> ...
> }
> (*bucket)++;
> *last = 0;
> }
> ...
> }
>
> This means one callback can scan the entire bucket array with softirqs held
> off.
>
> The only cap on n_buckets is the kvmalloc limit, which allows about 2^27
> buckets:
>
> net/openvswitch/flow_table.c:flow_key_insert() {
> ...
> if (table->count > ti->n_buckets)
> new_ti = table_instance_expand(ti, false);
> else if (time_after(jiffies, table->last_rehash + REHASH_INTERVAL))
> new_ti = table_instance_rehash(ti, ti->n_buckets, false);
> ...
> }
>
> The table never gets smaller in normal use:
>
> - ovs_flow_tbl_remove() never shrinks it.
> - The periodic rehash keeps the same size.
> - Only ovs_flow_tbl_flush() goes back to TBL_MIN_BUCKETS.
>
> OVS_DP_CMD_NEW and OVS_FLOW_CMD_NEW/DEL only need GENL_UNS_ADMIN_PERM. The
> dump needs no privilege at all:
>
> net/openvswitch/datapath.c:dp_flow_genl_ops[] {
> ...
> { .cmd = OVS_FLOW_CMD_GET,
> .validate = GENL_DONT_VALIDATE_STRICT | GENL_DONT_VALIDATE_DUMP,
> .flags = 0, /* OK for unprivileged users. */
> .doit = ovs_flow_cmd_get,
> .dumpit = ovs_flow_cmd_dump
> },
> ...
> }
>
> Root in an unprivileged user+net namespace could do the following:
>
> 1. Create a datapath.
> 2. Insert N flows so that n_buckets grows to between N and 2N.
> 3. Delete every flow except one in a high bucket.
> 4. Run the dump repeatedly from threads pinned to each CPU.
>
> Each dump could then walk up to 1 GB of hlist_heads.
I don't think this is a concern, as the user will have to allocate
hundreds of GBs of flows in the first place before freeing them,
which is likely a larger problem.
> During that walk,
> NET_RX, timers, tasklets and RCU callbacks are all deferred on that CPU. On
> CONFIG_PREEMPT (non-RT) kernels, the walk also can't be preempted.
>
> Before this patch, softirqs could still run on hardirq exit during the walk,
> and PREEMPT_RCU kernels could preempt it. On PREEMPT_NONE/VOLUNTARY kernels
> the walk was already non-preemptible under rcu_read_lock(), so the only new
> cost there is the softirq deferral.
>
> Container memcg limits don't cap how far the table can grow:
>
> - The sw_flow kmem_cache is created in ovs_flow_init() without
> SLAB_ACCOUNT.
> - The bucket array from table_instance_alloc() comes from unaccounted
> kvmalloc memory.
I have a patch set in works to properly account memory towards the memcg
of the user. That will cover this case. However, it's likely a net-next
material as it changes how users need to manage their memory.
Best regards, Ilya Maximets.
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump
2026-09-29 7:25 ` Denis V. Lunev
2026-10-01 22:27 ` netdev-bot+sashiko
@ 2026-10-02 20:50 ` Ilya Maximets
2026-10-05 3:00 ` patchwork-bot+netdevbpf
2 siblings, 0 replies; 6+ messages in thread
From: Ilya Maximets @ 2026-10-02 20:50 UTC (permalink / raw)
To: Denis V. Lunev, netdev
Cc: dev, Aaron Conole, Eelco Chaudron, Ilya Maximets, David S. Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Simon Horman, stable
On 9/29/26 9:25 AM, Denis V. Lunev wrote:
> From: Denis V. Lunev <den@openvz.org>
>
> A production compute node carrying a few thousand datapath flows hit a
> soft lockup inside a single netlink flow dump and panicked.
>
> ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it
> emits, and that releases stats->lock with spin_unlock_bh() once per
> CPU that has touched the flow. Every release is a local_bh_enable(),
> and each one runs the pending softirq backlog in the dumping thread's
> own context.
>
> The skb bounds the flows one callback emits, but not the softirq work
> it absorbs. On a CPU that carries the box's packet load the backlog
> refills as fast as it drains, so the dumping thread becomes that CPU's
> softirq engine. It never sleeps and it has no reschedule point, so
> under voluntary preemption nothing can take the CPU away from it:
> neither the ksoftirqd the kernel woke to take the work over, nor the
> stopper thread the softlockup detector dispatches to refresh its
> timestamp.
>
> Hold BH off across the whole callback instead, the way
> ctnetlink_dump_table() does, so the nested spin_unlock_bh() stop
> draining softirqs. The loop already runs under rcu_read_lock() and
> cannot sleep. What it gives up is preemption under CONFIG_PREEMPT,
> since a BH-off region is not preemptible outside PREEMPT_RT. The
> region stays short: the skb caps the flows one callback emits, and
> empty buckets cost no skb space but are each visited once per dump, as
> the cursor only moves forward. The table grows on insert and shrinks
> only on flush, so the walk is bounded by the largest flow count the
> datapath has held. The softirq backlog the callback used to absorb has
> no bound at all.
>
> Fixes: 63e7959c4b9b ("openvswitch: Per NUMA node flow stats.")
> Cc: stable@vger.kernel.org
> Signed-off-by: Denis V. Lunev <den@openvz.org>
> ---
> v3:
> - add the net prefix, Fixes tag and Cc stable
> - explain why the BH-off walk stays bounded: the cursor visits each
> empty bucket once per dump and the table shrinks only on flush
> - drop "here" from the comment, add blank lines around the
> local_bh_disable()/local_bh_enable() pair
> v2: https://lore.kernel.org/netdev/20260915122401.3910188-1-den@openvz.org/
> - leave ovs_vport_cmd_dump() alone: nsid_lock has not been BH-safe
> since commit aed4969f2bdf ("net: net->nsid_lock does not need BH
> safety"), so the vport dump never drained softirqs
> - disable BH before the table dereference and say in a comment that
> the region is not there for safety
> - drop the ovs_flow_stats_get() history, note the empty-bucket walk
> and the lost CONFIG_PREEMPT preemption in the message
> - move the Cc list out of the commit message
> v1: https://lore.kernel.org/netdev/20260909171920.1001074-1-den@openvz.org/
>
> net/openvswitch/datapath.c | 8 ++++++++
> 1 file changed, 8 insertions(+)
>
> diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
> index 631a03136fa1..4fc5d0bebd85 100644
> --- a/net/openvswitch/datapath.c
> +++ b/net/openvswitch/datapath.c
> @@ -1532,6 +1532,12 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
> return -ENODEV;
> }
>
> + /*
> + * Not needed for safety. Stops every spin_unlock_bh() in
> + * ovs_flow_stats_get() from running the softirq backlog.
> + */
> + local_bh_disable();
> +
> ti = rcu_dereference(dp->table.ti);
> for (;;) {
> struct sw_flow *flow;
> @@ -1552,6 +1558,8 @@ static int ovs_flow_cmd_dump(struct sk_buff *skb, struct netlink_callback *cb)
> cb->args[0] = bucket;
> cb->args[1] = obj;
> }
> +
> + local_bh_enable();
> rcu_read_unlock();
> return skb->len;
> }
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump
2026-09-29 7:25 ` Denis V. Lunev
2026-10-01 22:27 ` netdev-bot+sashiko
2026-10-02 20:50 ` Ilya Maximets
@ 2026-10-05 3:00 ` patchwork-bot+netdevbpf
2 siblings, 0 replies; 6+ messages in thread
From: patchwork-bot+netdevbpf @ 2026-10-05 3:00 UTC (permalink / raw)
To: Denis V. Lunev
Cc: netdev, dev, aconole, echaudro, i.maximets, davem, edumazet, kuba,
pabeni, horms, stable
Hello:
This patch was applied to netdev/net.git (main)
by David S. Miller <davem@davemloft.net>:
On Tue, 29 Sep 2026 09:25:19 +0200 you wrote:
> From: Denis V. Lunev <den@openvz.org>
>
> A production compute node carrying a few thousand datapath flows hit a
> soft lockup inside a single netlink flow dump and panicked.
>
> ovs_flow_cmd_dump() calls ovs_flow_stats_get() for every flow it
> emits, and that releases stats->lock with spin_unlock_bh() once per
> CPU that has touched the flow. Every release is a local_bh_enable(),
> and each one runs the pending softirq backlog in the dumping thread's
> own context.
>
> [...]
Here is the summary with links:
- [net,v3] openvswitch: fix soft lockup in the netlink flow dump
https://git.kernel.org/netdev/net/c/aaaaf87ea99b
You are awesome, thank you!
--
Deet-doot-dot, I am a bot.
https://korg.docs.kernel.org/patchwork/pwbot.html
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-10-05 3:00 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-29 7:25 [PATCH net v3] openvswitch: fix soft lockup in the netlink flow dump Denis V. Lunev
2026-09-29 7:25 ` Denis V. Lunev
2026-10-01 22:27 ` netdev-bot+sashiko
2026-10-02 20:41 ` Ilya Maximets
2026-10-02 20:50 ` Ilya Maximets
2026-10-05 3:00 ` patchwork-bot+netdevbpf
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox