* [PATCH v3 net-next 1/6] ibmvnic: Move long delayed work on system_dfl_long_wq
2026-07-20 10:08 [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Marco Crivellari
@ 2026-07-20 10:08 ` Marco Crivellari
2026-07-20 10:08 ` [PATCH v3 net-next 2/6] net: ti: icssg-stats: " Marco Crivellari
` (5 subsequent siblings)
6 siblings, 0 replies; 14+ messages in thread
From: Marco Crivellari @ 2026-07-20 10:08 UTC (permalink / raw)
To: linux-kernel, netdev
Cc: Tejun Heo, Lai Jiangshan, Frederic Weisbecker,
Sebastian Andrzej Siewior, Marco Crivellari, Michal Hocko,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Haren Myneni, Rick Lindsley, Nick Child,
Madhavan Srinivasan, Michael Ellerman, Nicholas Piggin,
Christophe Leroy (CS GROUP), linuxppc-dev
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: Haren Myneni <haren@linux.ibm.com>
Cc: Rick Lindsley <ricklind@linux.ibm.com>
Cc: Nick Child <nnac123@linux.ibm.com>
Cc: Madhavan Srinivasan <maddy@linux.ibm.com>
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Nicholas Piggin <npiggin@gmail.com>
Cc: Christophe Leroy (CS GROUP) <chleroy@kernel.org>
Cc: linuxppc-dev@lists.ozlabs.org
Signed-off-by: Marco Crivellari <marco.crivellari@suse.com>
---
drivers/net/ethernet/ibm/ibmvnic.c | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
diff --git a/drivers/net/ethernet/ibm/ibmvnic.c b/drivers/net/ethernet/ibm/ibmvnic.c
index 5a510eed335e..d4c284c8ef43 100644
--- a/drivers/net/ethernet/ibm/ibmvnic.c
+++ b/drivers/net/ethernet/ibm/ibmvnic.c
@@ -3229,7 +3229,7 @@ static void __ibmvnic_reset(struct work_struct *work)
if (adapter->state == VNIC_PROBING &&
!wait_for_completion_timeout(&adapter->probe_done, timeout)) {
dev_err(dev, "Reset thread timed out on probe");
- queue_delayed_work(system_long_wq,
+ queue_delayed_work(system_dfl_long_wq,
&adapter->ibmvnic_delayed_reset,
IBMVNIC_RESET_DELAY);
return;
@@ -3267,7 +3267,7 @@ static void __ibmvnic_reset(struct work_struct *work)
spin_lock(&adapter->rwi_lock);
if (!list_empty(&adapter->rwi_list)) {
if (test_and_set_bit_lock(0, &adapter->resetting)) {
- queue_delayed_work(system_long_wq,
+ queue_delayed_work(system_dfl_long_wq,
&adapter->ibmvnic_delayed_reset,
IBMVNIC_RESET_DELAY);
} else {
@@ -3454,7 +3454,7 @@ static int ibmvnic_reset(struct ibmvnic_adapter *adapter,
list_add_tail(&rwi->list, &adapter->rwi_list);
netdev_dbg(adapter->netdev, "Scheduling reset (reason %s)\n",
reset_reason_to_string(reason));
- queue_work(system_long_wq, &adapter->ibmvnic_reset);
+ queue_work(system_dfl_long_wq, &adapter->ibmvnic_reset);
ret = 0;
err:
--
2.54.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH v3 net-next 2/6] net: ti: icssg-stats: Move long delayed work on system_dfl_long_wq
2026-07-20 10:08 [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Marco Crivellari
2026-07-20 10:08 ` [PATCH v3 net-next 1/6] ibmvnic: Move long delayed work on system_dfl_long_wq Marco Crivellari
@ 2026-07-20 10:08 ` Marco Crivellari
2026-07-20 10:08 ` [PATCH v3 net-next 3/6] net: ti: icssg-prueth: " Marco Crivellari
` (4 subsequent siblings)
6 siblings, 0 replies; 14+ messages in thread
From: Marco Crivellari @ 2026-07-20 10:08 UTC (permalink / raw)
To: linux-kernel, netdev
Cc: Tejun Heo, Lai Jiangshan, Frederic Weisbecker,
Sebastian Andrzej Siewior, Marco Crivellari, Michal Hocko,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, MD Danish Anwar, Roger Quadros, linux-arm-kernel,
Richard Cheng
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: MD Danish Anwar <danishanwar@ti.com>
Cc: Roger Quadros <rogerq@kernel.org>
Cc: linux-arm-kernel@lists.infradead.org
Signed-off-by: Marco Crivellari <marco.crivellari@suse.com>
Reviewed-by: Richard Cheng <icheng@nvidia.com>
---
drivers/net/ethernet/ti/icssg/icssg_stats.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/ti/icssg/icssg_stats.c b/drivers/net/ethernet/ti/icssg/icssg_stats.c
index 7159baa0155c..7d6d6692d819 100644
--- a/drivers/net/ethernet/ti/icssg/icssg_stats.c
+++ b/drivers/net/ethernet/ti/icssg/icssg_stats.c
@@ -69,7 +69,7 @@ void icssg_stats_work_handler(struct work_struct *work)
stats_work.work);
emac_update_hardware_stats(emac);
- queue_delayed_work(system_long_wq, &emac->stats_work,
+ queue_delayed_work(system_dfl_long_wq, &emac->stats_work,
msecs_to_jiffies((STATS_TIME_LIMIT_1G_MS * 1000) / emac->speed));
}
EXPORT_SYMBOL_GPL(icssg_stats_work_handler);
--
2.54.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH v3 net-next 3/6] net: ti: icssg-prueth: Move long delayed work on system_dfl_long_wq
2026-07-20 10:08 [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Marco Crivellari
2026-07-20 10:08 ` [PATCH v3 net-next 1/6] ibmvnic: Move long delayed work on system_dfl_long_wq Marco Crivellari
2026-07-20 10:08 ` [PATCH v3 net-next 2/6] net: ti: icssg-stats: " Marco Crivellari
@ 2026-07-20 10:08 ` Marco Crivellari
2026-07-20 10:08 ` [PATCH v3 net-next 4/6] net: thunderbolt: " Marco Crivellari
` (3 subsequent siblings)
6 siblings, 0 replies; 14+ messages in thread
From: Marco Crivellari @ 2026-07-20 10:08 UTC (permalink / raw)
To: linux-kernel, netdev
Cc: Tejun Heo, Lai Jiangshan, Frederic Weisbecker,
Sebastian Andrzej Siewior, Marco Crivellari, Michal Hocko,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, MD Danish Anwar, Roger Quadros
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: MD Danish Anwar <danishanwar@ti.com>
Cc: Roger Quadros <rogerq@kernel.org>
Signed-off-by: Marco Crivellari <marco.crivellari@suse.com>
---
drivers/net/ethernet/ti/icssg/icssg_prueth.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/ti/icssg/icssg_prueth.c b/drivers/net/ethernet/ti/icssg/icssg_prueth.c
index 591be5c8056b..0fba22c1046e 100644
--- a/drivers/net/ethernet/ti/icssg/icssg_prueth.c
+++ b/drivers/net/ethernet/ti/icssg/icssg_prueth.c
@@ -1100,7 +1100,7 @@ static int emac_ndo_open(struct net_device *ndev)
prueth->emacs_initialized++;
- queue_work(system_long_wq, &emac->stats_work.work);
+ queue_work(system_dfl_long_wq, &emac->stats_work.work);
return 0;
--
2.54.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH v3 net-next 4/6] net: thunderbolt: Move long delayed work on system_dfl_long_wq
2026-07-20 10:08 [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Marco Crivellari
` (2 preceding siblings ...)
2026-07-20 10:08 ` [PATCH v3 net-next 3/6] net: ti: icssg-prueth: " Marco Crivellari
@ 2026-07-20 10:08 ` Marco Crivellari
2026-07-20 10:08 ` [PATCH v3 net-next 5/6] net: usb: pegasus: " Marco Crivellari
` (2 subsequent siblings)
6 siblings, 0 replies; 14+ messages in thread
From: Marco Crivellari @ 2026-07-20 10:08 UTC (permalink / raw)
To: linux-kernel, netdev
Cc: Tejun Heo, Lai Jiangshan, Frederic Weisbecker,
Sebastian Andrzej Siewior, Marco Crivellari, Michal Hocko,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Mika Westerberg, Yehezkel Bernat
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: Mika Westerberg <westeri@kernel.org>
Cc: Yehezkel Bernat <YehezkelShB@gmail.com>
Signed-off-by: Marco Crivellari <marco.crivellari@suse.com>
Acked-by: Mika Westerberg <westeri@kernel.org>
---
drivers/net/thunderbolt/main.c | 7 ++++---
1 file changed, 4 insertions(+), 3 deletions(-)
diff --git a/drivers/net/thunderbolt/main.c b/drivers/net/thunderbolt/main.c
index 02a91650561a..be27972bed18 100644
--- a/drivers/net/thunderbolt/main.c
+++ b/drivers/net/thunderbolt/main.c
@@ -316,7 +316,7 @@ static void start_login(struct tbnet *net)
net->login_received = false;
mutex_unlock(&net->connection_lock);
- queue_delayed_work(system_long_wq, &net->login_work,
+ queue_delayed_work(system_dfl_long_wq, &net->login_work,
msecs_to_jiffies(1000));
}
@@ -460,7 +460,7 @@ static int tbnet_handle_packet(const void *buf, size_t size, void *data)
if (net->login_retries >= TBNET_LOGIN_RETRIES ||
!net->login_sent) {
net->login_retries = 0;
- queue_delayed_work(system_long_wq,
+ queue_delayed_work(system_dfl_long_wq,
&net->login_work, 0);
}
mutex_unlock(&net->connection_lock);
@@ -700,7 +700,8 @@ static void tbnet_login_work(struct work_struct *work)
netdev_dbg(net->dev, "sending login request failed, ret=%d\n",
ret);
if (net->login_retries++ < TBNET_LOGIN_RETRIES) {
- queue_delayed_work(system_long_wq, &net->login_work,
+ queue_delayed_work(system_dfl_long_wq,
+ &net->login_work,
delay);
} else {
netdev_info(net->dev, "ThunderboltIP login timed out\n");
--
2.54.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH v3 net-next 5/6] net: usb: pegasus: Move long delayed work on system_dfl_long_wq
2026-07-20 10:08 [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Marco Crivellari
` (3 preceding siblings ...)
2026-07-20 10:08 ` [PATCH v3 net-next 4/6] net: thunderbolt: " Marco Crivellari
@ 2026-07-20 10:08 ` Marco Crivellari
2026-07-22 8:29 ` Oliver Neukum
2026-07-20 10:08 ` [PATCH v3 net-next 6/6] net: usb: r8152: " Marco Crivellari
2026-07-20 22:35 ` [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Jacob Keller
6 siblings, 1 reply; 14+ messages in thread
From: Marco Crivellari @ 2026-07-20 10:08 UTC (permalink / raw)
To: linux-kernel, netdev
Cc: Tejun Heo, Lai Jiangshan, Frederic Weisbecker,
Sebastian Andrzej Siewior, Marco Crivellari, Michal Hocko,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Petko Manolov, linux-usb
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: Petko Manolov <petkan@nucleusys.com>
Cc: linux-usb@vger.kernel.org
Signed-off-by: Marco Crivellari <marco.crivellari@suse.com>
---
drivers/net/usb/pegasus.c | 9 +++++----
1 file changed, 5 insertions(+), 4 deletions(-)
diff --git a/drivers/net/usb/pegasus.c b/drivers/net/usb/pegasus.c
index 8700eeb8e22d..c1798e14b224 100644
--- a/drivers/net/usb/pegasus.c
+++ b/drivers/net/usb/pegasus.c
@@ -1126,8 +1126,9 @@ static void check_carrier(struct work_struct *work)
pegasus_t *pegasus = container_of(work, pegasus_t, carrier_check.work);
set_carrier(pegasus->net);
if (!(pegasus->flags & PEGASUS_UNPLUG)) {
- queue_delayed_work(system_long_wq, &pegasus->carrier_check,
- CARRIER_CHECK_DELAY);
+ queue_delayed_work(system_dfl_long_wq,
+ &pegasus->carrier_check,
+ CARRIER_CHECK_DELAY);
}
}
@@ -1232,7 +1233,7 @@ static int pegasus_probe(struct usb_interface *intf,
res = register_netdev(net);
if (res)
goto out3;
- queue_delayed_work(system_long_wq, &pegasus->carrier_check,
+ queue_delayed_work(system_dfl_long_wq, &pegasus->carrier_check,
CARRIER_CHECK_DELAY);
dev_info(&intf->dev, "%s, %s, %pM\n", net->name,
usb_dev_id[dev_index].name, net->dev_addr);
@@ -1297,7 +1298,7 @@ static int pegasus_resume(struct usb_interface *intf)
pegasus->intr_urb->actual_length = 0;
intr_callback(pegasus->intr_urb);
}
- queue_delayed_work(system_long_wq, &pegasus->carrier_check,
+ queue_delayed_work(system_dfl_long_wq, &pegasus->carrier_check,
CARRIER_CHECK_DELAY);
return 0;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* Re: [PATCH v3 net-next 5/6] net: usb: pegasus: Move long delayed work on system_dfl_long_wq
2026-07-20 10:08 ` [PATCH v3 net-next 5/6] net: usb: pegasus: " Marco Crivellari
@ 2026-07-22 8:29 ` Oliver Neukum
2026-08-25 15:18 ` Sebastian Andrzej Siewior
0 siblings, 1 reply; 14+ messages in thread
From: Oliver Neukum @ 2026-07-22 8:29 UTC (permalink / raw)
To: Marco Crivellari, linux-kernel, netdev
Cc: Tejun Heo, Lai Jiangshan, Frederic Weisbecker,
Sebastian Andrzej Siewior, Michal Hocko, Andrew Lunn,
David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Petko Manolov, linux-usb
On 20.07.26 12:08, Marco Crivellari wrote:
Hi,
>
> Since the workqueue work doesn't rely on per-cpu variables, there is no
> obvious reason that justify the use of a per-cpu workqueue. So change
> system_long_wq with system_dfl_long_wq so that the work may benefit from
> scheduler task placement.
these changes are problematic, although they look like a good cleanup
in first place. But the test you are using to determine whether USB
devices need their own work queue is incomplete because you are not
considering the reason they allocate their own work queues.
These drivers have their own work queues because they are part of the block layer.
USB devices can share a device with a block device (storage & UAS) and
USB devices have common, per device operations, in particular reset
and runtime power management and disconnect handling. Because these operations
can be necessary to complete block IO neither they nor anything
they depend on can use IO to allocate memory. That is they need to
perform any memory allocation with GFP_NOIO or GFP_ATOMIC.
That is also true for any operation on a work queue they need to wait
for to make progress. That means you cannot limit your check to per-cpu
variables. You also need to check for such dependencies. In particular
any usage of flush_work() on such queues can deadlock, if you use
common queues.
Please refrain from making such changes unless you have fully analyzed
the dependencies.
Regards
Oliver
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH v3 net-next 5/6] net: usb: pegasus: Move long delayed work on system_dfl_long_wq
2026-07-22 8:29 ` Oliver Neukum
@ 2026-08-25 15:18 ` Sebastian Andrzej Siewior
2026-08-27 16:01 ` Alan Stern
0 siblings, 1 reply; 14+ messages in thread
From: Sebastian Andrzej Siewior @ 2026-08-25 15:18 UTC (permalink / raw)
To: Oliver Neukum
Cc: Marco Crivellari, linux-kernel, netdev, Tejun Heo, Lai Jiangshan,
Frederic Weisbecker, Michal Hocko, Andrew Lunn, David S . Miller,
Eric Dumazet, Jakub Kicinski, Paolo Abeni, Petko Manolov,
linux-usb
On 2026-07-22 10:29:50 [+0200], Oliver Neukum wrote:
> On 20.07.26 12:08, Marco Crivellari wrote:
> Hi,
Hi Oliver,
> > Since the workqueue work doesn't rely on per-cpu variables, there is no
> > obvious reason that justify the use of a per-cpu workqueue. So change
> > system_long_wq with system_dfl_long_wq so that the work may benefit from
> > scheduler task placement.
>
> these changes are problematic, although they look like a good cleanup
> in first place. But the test you are using to determine whether USB
> devices need their own work queue is incomplete because you are not
> considering the reason they allocate their own work queues.
This driver does not using its own workqueue.
> These drivers have their own work queues because they are part of the block layer.
> USB devices can share a device with a block device (storage & UAS) and
> USB devices have common, per device operations, in particular reset
> and runtime power management and disconnect handling. Because these operations
> can be necessary to complete block IO neither they nor anything
> they depend on can use IO to allocate memory. That is they need to
> perform any memory allocation with GFP_NOIO or GFP_ATOMIC.
This is a networking driver. It has nothing to do with storage and UAS.
It uses USB, yes.
> That is also true for any operation on a work queue they need to wait
> for to make progress. That means you cannot limit your check to per-cpu
> variables. You also need to check for such dependencies. In particular
> any usage of flush_work() on such queues can deadlock, if you use
> common queues.
It has nothing to do with per-CPU variables. In fact its
queue_delayed_work(,, CARRIER_CHECK_DELAY) usage already ensures that it
can be executed on a random and not on the submitting CPU.
> Please refrain from making such changes unless you have fully analyzed
> the dependencies.
Did you refer to the wrong patch? As far as this patch goes I can not
reason why you assume Marco did not fully analyze the dependencies. I
don't see anything wrong with this patch.
Reviewed-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
> Regards
> Oliver
>
Sebastian
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH v3 net-next 5/6] net: usb: pegasus: Move long delayed work on system_dfl_long_wq
2026-08-25 15:18 ` Sebastian Andrzej Siewior
@ 2026-08-27 16:01 ` Alan Stern
2026-08-28 9:43 ` Sebastian Andrzej Siewior
0 siblings, 1 reply; 14+ messages in thread
From: Alan Stern @ 2026-08-27 16:01 UTC (permalink / raw)
To: Sebastian Andrzej Siewior
Cc: Oliver Neukum, Marco Crivellari, linux-kernel, netdev, Tejun Heo,
Lai Jiangshan, Frederic Weisbecker, Michal Hocko, Andrew Lunn,
David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Petko Manolov, linux-usb
On Tue, Aug 25, 2026 at 05:18:12PM +0200, Sebastian Andrzej Siewior wrote:
> On 2026-07-22 10:29:50 [+0200], Oliver Neukum wrote:
> > On 20.07.26 12:08, Marco Crivellari wrote:
> > Hi,
> Hi Oliver,
>
> > > Since the workqueue work doesn't rely on per-cpu variables, there is no
> > > obvious reason that justify the use of a per-cpu workqueue. So change
> > > system_long_wq with system_dfl_long_wq so that the work may benefit from
> > > scheduler task placement.
> >
> > these changes are problematic, although they look like a good cleanup
> > in first place. But the test you are using to determine whether USB
> > devices need their own work queue is incomplete because you are not
> > considering the reason they allocate their own work queues.
>
> This driver does not using its own workqueue.
>
> > These drivers have their own work queues because they are part of the block layer.
> > USB devices can share a device with a block device (storage & UAS) and
> > USB devices have common, per device operations, in particular reset
> > and runtime power management and disconnect handling. Because these operations
> > can be necessary to complete block IO neither they nor anything
> > they depend on can use IO to allocate memory. That is they need to
> > perform any memory allocation with GFP_NOIO or GFP_ATOMIC.
>
> This is a networking driver. It has nothing to do with storage and UAS.
> It uses USB, yes.
Ah, but a composite USB device can have both a networking interface and
a mass-storage interface.
Suppose you have such a device, and suppose the disk attached to its
mass-storage interface contains a swap partition. Now suppose the
device is being reset, and as part of the preparation for that reset the
networking driver needs to flush its workqueue. This means waiting
until the work routines that are already running have completed.
Since it's a general-purpose workqueue, you don't know what those work
routines are going to do. One of them might try to allocate memory
using GFP_KERNEL. Suppose that in order to satisfy the memory request,
the kernel decides it needs to write some pages to the swap partition on
the USB mass-storage interface. But the mass-storage driver is stuck;
it can't do anything until the device reset finishes. Deadlock.
That's why USB drivers have to use their own workqueues.
Alan Stern
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH v3 net-next 5/6] net: usb: pegasus: Move long delayed work on system_dfl_long_wq
2026-08-27 16:01 ` Alan Stern
@ 2026-08-28 9:43 ` Sebastian Andrzej Siewior
2026-08-28 14:10 ` Alan Stern
0 siblings, 1 reply; 14+ messages in thread
From: Sebastian Andrzej Siewior @ 2026-08-28 9:43 UTC (permalink / raw)
To: Alan Stern
Cc: Oliver Neukum, Marco Crivellari, linux-kernel, netdev, Tejun Heo,
Lai Jiangshan, Frederic Weisbecker, Michal Hocko, Andrew Lunn,
David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Petko Manolov, linux-usb
On 2026-08-27 12:01:02 [-0400], Alan Stern wrote:
> > > These drivers have their own work queues because they are part of the block layer.
> > > USB devices can share a device with a block device (storage & UAS) and
> > > USB devices have common, per device operations, in particular reset
> > > and runtime power management and disconnect handling. Because these operations
> > > can be necessary to complete block IO neither they nor anything
> > > they depend on can use IO to allocate memory. That is they need to
> > > perform any memory allocation with GFP_NOIO or GFP_ATOMIC.
> >
> > This is a networking driver. It has nothing to do with storage and UAS.
> > It uses USB, yes.
>
> Ah, but a composite USB device can have both a networking interface and
> a mass-storage interface.
okay.
> Suppose you have such a device, and suppose the disk attached to its
> mass-storage interface contains a swap partition. Now suppose the
> device is being reset, and as part of the preparation for that reset the
> networking driver needs to flush its workqueue. This means waiting
> until the work routines that are already running have completed.
The individual functions are usually independent. But if the USB core
would reset the whole device it would reset each function.
> Since it's a general-purpose workqueue, you don't know what those work
> routines are going to do. One of them might try to allocate memory
> using GFP_KERNEL. Suppose that in order to satisfy the memory request,
> the kernel decides it needs to write some pages to the swap partition on
> the USB mass-storage interface. But the mass-storage driver is stuck;
> it can't do anything until the device reset finishes. Deadlock.
So you are saying, the USB-storage device is in reset and we wait until
the networking part finishes its workqueue flush. That flush is stuck
behind behind a memory allocation which waits on the storage device. So
any URB passed to usb_submit_urb() just waits for the reset to complete?
> That's why USB drivers have to use their own workqueues.
While this does make sense I don't see how this is related to this
patch. The pegasus driver uses `system_long_wq'. This is a system wide
workqueue_struct and is not limited to USB or this driver.
This workqueue is per-CPU meaning if you enqueue the work item on CPU3
it will be executed on CPU3. However pegasus uses a delayed work item
and the timer can fire on any CPU so even if it is enqueued on CPU3 it
could be executed on CPU1.
Therefore the suggested change system_long_wq -> system_dfl_long_wq
should not make a difference here: it is a different workqueue and it is
unbound (instead of per-CPU) but given the usage it is unchanged but
more obvious. Also its usage recommendations (use this for long running
items) is the same.
The plan is remove system_long_wq from the tree.
> Alan Stern
Sebastian
^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [PATCH v3 net-next 5/6] net: usb: pegasus: Move long delayed work on system_dfl_long_wq
2026-08-28 9:43 ` Sebastian Andrzej Siewior
@ 2026-08-28 14:10 ` Alan Stern
0 siblings, 0 replies; 14+ messages in thread
From: Alan Stern @ 2026-08-28 14:10 UTC (permalink / raw)
To: Sebastian Andrzej Siewior
Cc: Oliver Neukum, Marco Crivellari, linux-kernel, netdev, Tejun Heo,
Lai Jiangshan, Frederic Weisbecker, Michal Hocko, Andrew Lunn,
David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Petko Manolov, linux-usb
On Fri, Aug 28, 2026 at 11:43:18AM +0200, Sebastian Andrzej Siewior wrote:
> On 2026-08-27 12:01:02 [-0400], Alan Stern wrote:
> > > > These drivers have their own work queues because they are part of the block layer.
> > > > USB devices can share a device with a block device (storage & UAS) and
> > > > USB devices have common, per device operations, in particular reset
> > > > and runtime power management and disconnect handling. Because these operations
> > > > can be necessary to complete block IO neither they nor anything
> > > > they depend on can use IO to allocate memory. That is they need to
> > > > perform any memory allocation with GFP_NOIO or GFP_ATOMIC.
> > >
> > > This is a networking driver. It has nothing to do with storage and UAS.
> > > It uses USB, yes.
> >
> > Ah, but a composite USB device can have both a networking interface and
> > a mass-storage interface.
>
> okay.
>
> > Suppose you have such a device, and suppose the disk attached to its
> > mass-storage interface contains a swap partition. Now suppose the
> > device is being reset, and as part of the preparation for that reset the
> > networking driver needs to flush its workqueue. This means waiting
> > until the work routines that are already running have completed.
>
> The individual functions are usually independent. But if the USB core
> would reset the whole device it would reset each function.
>
> > Since it's a general-purpose workqueue, you don't know what those work
> > routines are going to do. One of them might try to allocate memory
> > using GFP_KERNEL. Suppose that in order to satisfy the memory request,
> > the kernel decides it needs to write some pages to the swap partition on
> > the USB mass-storage interface. But the mass-storage driver is stuck;
> > it can't do anything until the device reset finishes. Deadlock.
>
> So you are saying, the USB-storage device is in reset and we wait until
> the networking part finishes its workqueue flush. That flush is stuck
> behind behind a memory allocation which waits on the storage device. So
> any URB passed to usb_submit_urb() just waits for the reset to complete?
>
> > That's why USB drivers have to use their own workqueues.
>
> While this does make sense I don't see how this is related to this
> patch. The pegasus driver uses `system_long_wq'. This is a system wide
> workqueue_struct and is not limited to USB or this driver.
> This workqueue is per-CPU meaning if you enqueue the work item on CPU3
> it will be executed on CPU3. However pegasus uses a delayed work item
> and the timer can fire on any CPU so even if it is enqueued on CPU3 it
> could be executed on CPU1.
It doesn't matter what CPU the work item runs on. Here's the deadlock
sequence, in brief:
USB device reset cannot proceed until network interface's
->pre_reset() method returns.
The ->pre_reset() method cannot return until its call to
flush_workqueue() returns.
flush_workqueue() cannot return until the already executing
work item finishes.
The work item cannot finish until its kmalloc() call returns.
kmalloc() won't return until the kernel can free up memory by
writing some pages to the swap partition.
The write to the swap partition cannot take place until the
usb_storage/uas driver carries it out.
usb_storage/uas cannot do anything until the USB device reset
is finished.
> Therefore the suggested change system_long_wq -> system_dfl_long_wq
> should not make a difference here: it is a different workqueue and it is
> unbound (instead of per-CPU) but given the usage it is unchanged but
> more obvious. Also its usage recommendations (use this for long running
> items) is the same.
>
> The plan is remove system_long_wq from the tree.
The point Oliver was making is that the driver shouldn't be using a
general-purpose workqueue at all. Switching from one general-purpose
workqueue to another ignores this point; it's not the right thing to do.
Alan Stern
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH v3 net-next 6/6] net: usb: r8152: Move long delayed work on system_dfl_long_wq
2026-07-20 10:08 [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Marco Crivellari
` (4 preceding siblings ...)
2026-07-20 10:08 ` [PATCH v3 net-next 5/6] net: usb: pegasus: " Marco Crivellari
@ 2026-07-20 10:08 ` Marco Crivellari
2026-07-20 22:35 ` [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Jacob Keller
6 siblings, 0 replies; 14+ messages in thread
From: Marco Crivellari @ 2026-07-20 10:08 UTC (permalink / raw)
To: linux-kernel, netdev
Cc: Tejun Heo, Lai Jiangshan, Frederic Weisbecker,
Sebastian Andrzej Siewior, Marco Crivellari, Michal Hocko,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Ethan Nelson-Moore, linux-usb
Currently the code enqueue work items using {queue|mod}_delayed_work(),
using system_long_wq. This workqueue should be used when long works are
expected and it is a per-cpu workqueue.
The function(s) end up calling __queue_delayed_work(), which set a global
timer that could fire anywhere, enqueuing the work where the timer fired.
Unbound works could benefit from scheduler task placement, to optimize
performance and power consumption. Long work shouldn't stick to a single
CPU.
Recently, a new unbound workqueue specific for long running work has
been added:
c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
Since the workqueue work doesn't rely on per-cpu variables, there is no
obvious reason that justify the use of a per-cpu workqueue. So change
system_long_wq with system_dfl_long_wq so that the work may benefit from
scheduler task placement.
Cc: Ethan Nelson-Moore <enelsonmoore@gmail.com>
Cc: linux-usb@vger.kernel.org
Signed-off-by: Marco Crivellari <marco.crivellari@suse.com>
---
drivers/net/usb/r8152.c | 7 ++++---
1 file changed, 4 insertions(+), 3 deletions(-)
diff --git a/drivers/net/usb/r8152.c b/drivers/net/usb/r8152.c
index f61686433031..f6af66f294db 100644
--- a/drivers/net/usb/r8152.c
+++ b/drivers/net/usb/r8152.c
@@ -7072,7 +7072,8 @@ static void rtl_hw_phy_work_func_t(struct work_struct *work)
/* Delay execution in case request_firmware() is not ready yet.
*/
- queue_delayed_work(system_long_wq, &tp->hw_phy_work, HZ * 10);
+ queue_delayed_work(system_dfl_long_wq, &tp->hw_phy_work,
+ HZ * 10);
goto ignore_once;
}
@@ -8840,7 +8841,7 @@ static int rtl8152_reset_resume(struct usb_interface *intf)
clear_bit(SELECTIVE_SUSPEND, &tp->flags);
rtl_reset_ocp_base(tp);
tp->rtl_ops.init(tp);
- queue_delayed_work(system_long_wq, &tp->hw_phy_work, 0);
+ queue_delayed_work(system_dfl_long_wq, &tp->hw_phy_work, 0);
set_ethernet_addr(tp, true);
return rtl8152_resume(intf);
}
@@ -10295,7 +10296,7 @@ static int rtl8152_probe_once(struct usb_interface *intf,
/* Retry in case request_firmware() is not ready yet. */
tp->rtl_fw.retry = true;
#endif
- queue_delayed_work(system_long_wq, &tp->hw_phy_work, 0);
+ queue_delayed_work(system_dfl_long_wq, &tp->hw_phy_work, 0);
set_ethernet_addr(tp, false);
usb_set_intfdata(intf, tp);
--
2.54.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* Re: [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq
2026-07-20 10:08 [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Marco Crivellari
` (5 preceding siblings ...)
2026-07-20 10:08 ` [PATCH v3 net-next 6/6] net: usb: r8152: " Marco Crivellari
@ 2026-07-20 22:35 ` Jacob Keller
2026-07-21 8:20 ` Marco Crivellari
6 siblings, 1 reply; 14+ messages in thread
From: Jacob Keller @ 2026-07-20 22:35 UTC (permalink / raw)
To: Marco Crivellari, linux-kernel, netdev
Cc: Tejun Heo, Lai Jiangshan, Frederic Weisbecker,
Sebastian Andrzej Siewior, Michal Hocko, Andrew Lunn,
David S . Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Christophe Leroy (CS GROUP), Ethan Nelson-Moore, Haren Myneni,
Madhavan Srinivasan, MD Danish Anwar, Michael Ellerman,
Mika Westerberg, Nicholas Piggin, Nick Child, Petko Manolov,
Richard Cheng, Rick Lindsley, Roger Quadros, Yehezkel Bernat
On 7/20/2026 3:08 AM, Marco Crivellari wrote:
> Hello,
>
> Currently the code uses the per-cpu workqueue system_long_wq to schedule
> long running works.
>
> Unbound works could benefit from scheduler task placement, to optimize
> performance and power consumption. Another good reason to have this unbound,
> is the "queue_delayed_work()" function, used to enqueue the work item.
> More details on this will follow in the next section.
>
> Recently, a new unbound workqueue specific for long running work has been
> added:
>
> c116737e972e ("workqueue: Add system_dfl_long_wq for long unbound works")
>
> ~~~ Details about queue_delayed_work ~~~
>
> system_long_wq is a per-cpu workqueue and it is used as a parameter of
> queue_delayed_work(). This function schedule an item that it will later
> be enqueued (once the timer will fire). __queue_delayed_work() does the job
> receiving as "cpu" WORK_CPU_UNBOUND:
>
> if (housekeeping_enabled(HK_TYPE_TIMER)) {
> // [....]
> } else {
> if (likely(cpu == WORK_CPU_UNBOUND))
> add_timer_global(timer);
> else
> add_timer_on(timer, cpu);
> }
>
> The timer is global, so can fire everywhere, and the work item will be
> enqueued where the timer fired.
>
> Since the workqueue work doesn't rely on per-cpu variables, there is no
> obvious reason that justify the use of a per-cpu workqueue. So change the
> workqueue with the new system_dfl_long_wq, so that the used workqueue is
> now unbound and can benefit from scheduler task placement.
>
Ok. So if I am understanding this correctly, the current code uses
system_long_wq which is per-CPU, but is fired using an unbound timer. As
a result, whichever CPU the timer triggers on will be the one which
selects the work queue. From there, the work item will be enqueued to
that work queue and remain on that work queue until resolving with no
way for scheduler to adjust it?
With the new change, we schedule on the system_dfl_long_wq which *isn't*
per CPU, so the scheduler is free to move the task around and
reschedule. As a result we get better overall behavior with more input
from the scheduler, instead of effective randomness from the timer which
is then forced so that such long running task cannot migrate?
That sounds like a pretty good improvement for the cases where the
queued work doesn't depend on any per-cpu behavior. Nice!
I am not sure I can speak to any of the individual drivers here since I
wouldn't know whether moving that particular work item would be
affected.. so feel free to take this review with a grain of salt :)
Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
^ permalink raw reply [flat|nested] 14+ messages in thread* Re: [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq
2026-07-20 22:35 ` [PATCH v3 net-next 0/6] net: Move system_long_wq to system_dfl_long_wq Jacob Keller
@ 2026-07-21 8:20 ` Marco Crivellari
0 siblings, 0 replies; 14+ messages in thread
From: Marco Crivellari @ 2026-07-21 8:20 UTC (permalink / raw)
To: Jacob Keller
Cc: linux-kernel, netdev, Tejun Heo, Lai Jiangshan,
Frederic Weisbecker, Sebastian Andrzej Siewior, Michal Hocko,
Andrew Lunn, David S . Miller, Eric Dumazet, Jakub Kicinski,
Paolo Abeni, Christophe Leroy (CS GROUP), Ethan Nelson-Moore,
Haren Myneni, Madhavan Srinivasan, MD Danish Anwar,
Michael Ellerman, Mika Westerberg, Nicholas Piggin, Nick Child,
Petko Manolov, Richard Cheng, Rick Lindsley, Roger Quadros,
Yehezkel Bernat
Hi,
On Tue, Jul 21, 2026 at 12:35 AM Jacob Keller <jacob.e.keller@intel.com> wrote:
> [...]
> > system_long_wq is a per-cpu workqueue and it is used as a parameter of
> > queue_delayed_work(). This function schedule an item that it will later
> > be enqueued (once the timer will fire). __queue_delayed_work() does the job
> > receiving as "cpu" WORK_CPU_UNBOUND:
> >
> > if (housekeeping_enabled(HK_TYPE_TIMER)) {
> > // [....]
> > } else {
> > if (likely(cpu == WORK_CPU_UNBOUND))
> > add_timer_global(timer);
> > else
> > add_timer_on(timer, cpu);
> > }
> >
> > The timer is global, so can fire everywhere, and the work item will be
> > enqueued where the timer fired.
> >
> > Since the workqueue work doesn't rely on per-cpu variables, there is no
> > obvious reason that justify the use of a per-cpu workqueue. So change the
> > workqueue with the new system_dfl_long_wq, so that the used workqueue is
> > now unbound and can benefit from scheduler task placement.
> >
>
>
> Ok. So if I am understanding this correctly, the current code uses
> system_long_wq which is per-CPU, but is fired using an unbound timer. As
> a result, whichever CPU the timer triggers on will be the one which
> selects the work queue. From there, the work item will be enqueued to
> that work queue and remain on that work queue until resolving with no
> way for scheduler to adjust it?
>
> With the new change, we schedule on the system_dfl_long_wq which *isn't*
> per CPU, so the scheduler is free to move the task around and
> reschedule. As a result we get better overall behavior with more input
> from the scheduler, instead of effective randomness from the timer which
> is then forced so that such long running task cannot migrate?
>
> That sounds like a pretty good improvement for the cases where the
> queued work doesn't depend on any per-cpu behavior. Nice!
Yes, that's pretty much it!
> I am not sure I can speak to any of the individual drivers here since I
> wouldn't know whether moving that particular work item would be
> affected.. so feel free to take this review with a grain of salt :)
>
> Reviewed-by: Jacob Keller <jacob.e.keller@intel.com>
Sure, thank you!
--
Marco Crivellari
SUSE Labs
^ permalink raw reply [flat|nested] 14+ messages in thread