* [PATCH 5.10] RDMA/hns: Fix WQ_MEM_RECLAIM warning
@ 2026-10-06 8:18 Roman Demidov
2026-10-06 8:37 ` sashiko-bot
0 siblings, 1 reply; 2+ messages in thread
From: Roman Demidov @ 2026-10-06 8:18 UTC (permalink / raw)
To: stable, Greg Kroah-Hartman
Cc: Roman Demidov, Chengchang Tang, Junxian Huang, Jason Gunthorpe,
Leon Romanovsky, Yixian Liu, Salil Mehta, linux-rdma,
linux-kernel, lvc-project, Sasha Levin
From: Chengchang Tang <tangchengchang@huawei.com>
commit c0a26bbd3f99b7b03f072e3409aff4e6ec8af6f6 upstream.
When sunrpc is used, if a reset triggered, our wq may lead the
following trace:
workqueue: WQ_MEM_RECLAIM xprtiod:xprt_rdma_connect_worker [rpcrdma]
is flushing !WQ_MEM_RECLAIM hns_roce_irq_workq:flush_work_handle
[hns_roce_hw_v2]
WARNING: CPU: 0 PID: 8250 at kernel/workqueue.c:2644 check_flush_dependency+0xe0/0x144
Call trace:
check_flush_dependency+0xe0/0x144
start_flush_work.constprop.0+0x1d0/0x2f0
__flush_work.isra.0+0x40/0xb0
flush_work+0x14/0x30
hns_roce_v2_destroy_qp+0xac/0x1e0 [hns_roce_hw_v2]
ib_destroy_qp_user+0x9c/0x2b4
rdma_destroy_qp+0x34/0xb0
rpcrdma_ep_destroy+0x28/0xcc [rpcrdma]
rpcrdma_ep_put+0x74/0xb4 [rpcrdma]
rpcrdma_xprt_disconnect+0x1d8/0x260 [rpcrdma]
xprt_rdma_connect_worker+0xc0/0x120 [rpcrdma]
process_one_work+0x1cc/0x4d0
worker_thread+0x154/0x414
kthread+0x104/0x144
ret_from_fork+0x10/0x18
Since QP destruction frees memory, this wq should have the WQ_MEM_RECLAIM.
Fixes: ffd541d45726 ("RDMA/hns: Add the workqueue framework for flush cqe handler")
Signed-off-by: Chengchang Tang <tangchengchang@huawei.com>
Signed-off-by: Junxian Huang <huangjunxian6@hisilicon.com>
Link: https://patch.msgid.link/20260104064057.1582216-2-huangjunxian6@hisilicon.com
Signed-off-by: Leon Romanovsky <leon@kernel.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
Signed-off-by: Roman Demidov <roman.demidov.nn@gmail.com>
---
Backport fix for CVE-2026-46265
drivers/infiniband/hw/hns/hns_roce_hw_v2.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/drivers/infiniband/hw/hns/hns_roce_hw_v2.c b/drivers/infiniband/hw/hns/hns_roce_hw_v2.c
index bb744ba155e2..56c7af1c3b24 100644
--- a/drivers/infiniband/hw/hns/hns_roce_hw_v2.c
+++ b/drivers/infiniband/hw/hns/hns_roce_hw_v2.c
@@ -6087,7 +6087,8 @@ static int hns_roce_v2_init_eq_table(struct hns_roce_dev *hr_dev)
goto err_request_irq_fail;
}
- hr_dev->irq_workq = alloc_ordered_workqueue("hns_roce_irq_workq", 0);
+ hr_dev->irq_workq = alloc_ordered_workqueue("hns_roce_irq_workq",
+ WQ_MEM_RECLAIM);
if (!hr_dev->irq_workq) {
dev_err(dev, "Create irq workqueue failed!\n");
ret = -ENOMEM;
--
2.53.0
^ permalink raw reply related [flat|nested] 2+ messages in thread
* Re: [PATCH 5.10] RDMA/hns: Fix WQ_MEM_RECLAIM warning
2026-10-06 8:18 [PATCH 5.10] RDMA/hns: Fix WQ_MEM_RECLAIM warning Roman Demidov
@ 2026-10-06 8:37 ` sashiko-bot
0 siblings, 0 replies; 2+ messages in thread
From: sashiko-bot @ 2026-10-06 8:37 UTC (permalink / raw)
To: Roman Demidov; +Cc: linux-rdma
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] workqueue: WQ_MEM_RECLAIM rescuer deadlock via GFP_KERNEL
--
commit 324f8a5f04c3f5e308cb6f7d1d9cb0367924235e
Author: Chengchang Tang <tangchengchang@huawei.com>
RDMA/hns: Fix WQ_MEM_RECLAIM warning
This commit adds the WQ_MEM_RECLAIM flag to the hns_roce_irq_workq
workqueue to resolve a check_flush_dependency warning triggered during
QP destruction.
Link: https://patch.msgid.link/20260104064057.1582216-2-huangjunxian6@hisilicon.com
> diff --git a/drivers/infiniband/hw/hns/hns_roce_hw_v2.c b/drivers/infiniband/hw/hns/hns_roce_hw_v2.c
> index 0468028ffe390..3e49baf2abf61 100644
> --- a/drivers/infiniband/hw/hns/hns_roce_hw_v2.c
> +++ b/drivers/infiniband/hw/hns/hns_roce_hw_v2.c
> @@ -6082,7 +6082,8 @@ static int hns_roce_v2_init_eq_table(struct hns_roce_dev *hr_dev)
> goto err_request_irq_fail;
> }
>
> - hr_dev->irq_workq = alloc_ordered_workqueue("hns_roce_irq_workq", 0);
> + hr_dev->irq_workq = alloc_ordered_workqueue("hns_roce_irq_workq",
> + WQ_MEM_RECLAIM);
[Severity: High]
Adding WQ_MEM_RECLAIM to a workqueue creates a rescuer thread that
executes during memory pressure to guarantee forward progress. Can this
introduce a deadlock if the work items perform GFP_KERNEL allocations?
Looking at the work executed by this queue, flush_work_handle() calls
hns_roce_modify_qp(), which eventually calls hns_roce_alloc_cmd_mailbox().
In hns_roce_alloc_cmd_mailbox(), memory is allocated using GFP_KERNEL:
drivers/infiniband/hw/hns/hns_roce_cmd.c:hns_roce_alloc_cmd_mailbox() {
...
mailbox = kmalloc(sizeof(*mailbox), GFP_KERNEL);
if (!mailbox)
return ERR_PTR(-ENOMEM);
mailbox->buf = dma_pool_alloc(hr_dev->cmd.pool, GFP_KERNEL,
&mailbox->dma);
...
}
Since the rescuer thread is supposed to relieve memory pressure, blocking
on GFP_KERNEL allocations defeats its purpose and could permanently block
the rescuer. Does this patch trade a flush dependency warning for a
potential deadlock?
> if (!hr_dev->irq_workq) {
> dev_err(dev, "Create irq workqueue failed!\n");
> ret = -ENOMEM;
--
Sashiko AI review · https://sashiko.dev/#/patchset/20261006081840.125567-1-roman.demidov.nn@gmail.com?part=1
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-06 8:37 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-06 8:18 [PATCH 5.10] RDMA/hns: Fix WQ_MEM_RECLAIM warning Roman Demidov
2026-10-06 8:37 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox