* [PATCH v4] scsi: libsas: Fix SMP IO deadlock during HA resume
@ 2026-09-28 4:02 Xingui Yang
2026-09-28 4:14 ` sashiko-bot
2026-09-30 10:40 ` John Garry
0 siblings, 2 replies; 5+ messages in thread
From: Xingui Yang @ 2026-09-28 4:02 UTC (permalink / raw)
To: john.garry, yanaijie, jejb, mkp
Cc: linux-scsi, linux-kernel, linuxarm, yangxingui, liuyonglong,
kangfenglong
smp_execute_task_sg() calls pm_runtime_get_sync() on the host before
issuing an SMP command. When that command is itself issued from the
HA resume path, the get_sync() deadlocks: it waits for the ongoing
resume (the device is RPM_RESUMING), while the resume is blocked in
sas_drain_work() waiting for that same SMP IO to complete.
The deadlock needs an expander-attached SATA disk. During
sas_resume_ha() -> sas_drain_work(), DISCE_RESUME ->
sas_resume_sata() -> ata_sas_port_resume() requests ATA_EH_RESET,
and the hard reset for such a disk is done via SMP PHY CONTROL
(sas_ata_hard_reset() -> sas_phy_reset() -> sas_smp_phy_control() ->
smp_execute_task_sg()). Direct-attached SATA resets through
lldd_control_phy() and SSP devices use TMFs, so neither hits this.
Replace the get_sync()/put_sync() pair with
pm_runtime_get_noresume()/pm_runtime_put(). smp_execute_task_sg()
only needs to hold off autosuspend while the SMP is in flight, and it
must not try to resume the host: a sync resume issued from the HA
resume path itself is what deadlocks, and by the time sas_resume_ha()
runs, hw_init has already reinitialized the hardware, so the device
is accessible without one.
The usage reference is still required. Discovery work normally runs
inside an event worker's PM reference, taken at
sas_notify_port_event() notify time and held until the handler has
flushed the disco queue. sas_rediscover_ex_phy() however requeues
DISCE_REVALIDATE_DOMAIN from within the revalidation worker itself,
and flush_workqueue() does not wait for work items queued during
execution, so that chained revalidation runs with no outer PM
reference - without the get_noresume(), its SMP could race
autosuspend.
For the BSG path, sas_smp_handler() is the only caller which may
find the host autosuspended: expander SMP requests do not go through
any SCSI device request queue, so nothing else in that path holds
the host awake. Resume it there with pm_runtime_resume_and_get()
and check the result.
Fixes: 3dbbbf656b850 ("scsi: libsas: Fix HA resume deadlock and hisi_sas disk-wake race")
Signed-off-by: Xingui Yang <yangxingui@huawei.com>
---
Changes since v3:
- Move the host resume to sas_smp_handler(), the only caller which may
find the host autosuspended, as suggested by John Garry
- Replace get_sync()/put_sync() with get_noresume()/put() in
smp_execute_task_sg(): a blocking resume issued from the HA resume
path itself is what deadlocks
- Drop the racy SAS_HA_RESUMING check
Changes since v2:
- Drop the reference with pm_runtime_put() instead of
pm_runtime_put_noidle().
Changes since v1:
- Use pm_runtime_get_noresume()/put_noidle() during HA resume instead
of skipping the PM reference entirely, so an in-flight SMP IO always
keeps autosuspend away.
- Convert pm_runtime_get_sync() to pm_runtime_resume_and_get() and
check the result (pre-existing issue flagged by sashiko).
drivers/scsi/libsas/sas_expander.c | 14 ++++++++++++--
1 file changed, 12 insertions(+), 2 deletions(-)
diff --git a/drivers/scsi/libsas/sas_expander.c b/drivers/scsi/libsas/sas_expander.c
index 811c9eb4fef1..26c2099c28b9 100644
--- a/drivers/scsi/libsas/sas_expander.c
+++ b/drivers/scsi/libsas/sas_expander.c
@@ -62,7 +62,11 @@ static int smp_execute_task_sg(struct domain_device *dev,
to_sas_internal(dev->port->ha->shost->transportt);
struct sas_ha_struct *ha = dev->port->ha;
- pm_runtime_get_sync(ha->dev);
+ /*
+ * Non-blocking: a sync resume here would deadlock against
+ * sas_drain_work() during HA resume.
+ */
+ pm_runtime_get_noresume(ha->dev);
mutex_lock(&dev->ex_dev.cmd_mutex);
for (retry = 0; retry < 3; retry++) {
if (test_bit(SAS_DEV_GONE, &dev->state)) {
@@ -135,7 +139,7 @@ static int smp_execute_task_sg(struct domain_device *dev,
}
}
mutex_unlock(&dev->ex_dev.cmd_mutex);
- pm_runtime_put_sync(ha->dev);
+ pm_runtime_put(ha->dev);
BUG_ON(retry == 3 && task != NULL);
sas_free_task(task);
@@ -2222,6 +2226,11 @@ void sas_smp_handler(struct bsg_job *job, struct Scsi_Host *shost,
goto out;
}
+ /* The host may have autosuspended, resume it here. */
+ ret = pm_runtime_resume_and_get(dev->port->ha->dev);
+ if (ret)
+ goto out;
+
ret = smp_execute_task_sg(dev, job->request_payload.sg_list,
job->reply_payload.sg_list);
if (ret >= 0) {
@@ -2229,6 +2238,7 @@ void sas_smp_handler(struct bsg_job *job, struct Scsi_Host *shost,
rcvlen = job->reply_payload.payload_len - ret;
ret = 0;
}
+ pm_runtime_put(dev->port->ha->dev);
out:
bsg_job_done(job, ret, rcvlen);
--
2.43.0
^ permalink raw reply related [flat|nested] 5+ messages in thread* Re: [PATCH v4] scsi: libsas: Fix SMP IO deadlock during HA resume
2026-09-28 4:02 [PATCH v4] scsi: libsas: Fix SMP IO deadlock during HA resume Xingui Yang
@ 2026-09-28 4:14 ` sashiko-bot
2026-09-28 7:30 ` yangxingui
2026-09-30 10:40 ` John Garry
1 sibling, 1 reply; 5+ messages in thread
From: sashiko-bot @ 2026-09-28 4:14 UTC (permalink / raw)
To: Xingui Yang; +Cc: linux-scsi
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] scsi: libsas: hardware access on suspended host in smp_execute_task_sg()
--
commit b6ffaa3a7e18a8e3fec2c6c49ebb1c88915b9f8d
Author: Xingui Yang <yangxingui@huawei.com>
scsi: libsas: Fix SMP IO deadlock during HA resume
smp_execute_task_sg() calls pm_runtime_get_sync() on the host before
issuing an SMP command. When that command is itself issued from the
HA resume path, the get_sync() deadlocks: it waits for the ongoing
resume (the device is RPM_RESUMING), while the resume is blocked in
sas_drain_work() waiting for that same SMP IO to complete.
> The usage reference is still required. Discovery work normally runs
> inside an event worker's PM reference, taken at
> sas_notify_port_event() notify time and held until the handler has
> flushed the disco queue. sas_rediscover_ex_phy() however requeues
> DISCE_REVALIDATE_DOMAIN from within the revalidation worker itself,
> and flush_workqueue() does not wait for work items queued during
> execution, so that chained revalidation runs with no outer PM
> reference - without the get_noresume(), its SMP could race
> autosuspend.
[Severity: High]
Does pm_runtime_get_noresume() actually prevent this race if the device
has already autosuspended?
When chained revalidation occurs in sas_discover.c:sas_revalidate_domain(),
it runs without an outer PM reference. If the host autosuspends before
smp_execute_task_sg() runs, pm_runtime_get_noresume() will only increment
the usage counter and leave the device suspended.
> diff --git a/drivers/scsi/libsas/sas_expander.c b/drivers/scsi/libsas/sas_expander.c
> index 811c9eb4fef11..26c2099c28b97 100644
> --- a/drivers/scsi/libsas/sas_expander.c
> +++ b/drivers/scsi/libsas/sas_expander.c
> @@ -62,7 +62,11 @@ static int smp_execute_task_sg(struct domain_device *dev,
> to_sas_internal(dev->port->ha->shost->transportt);
> struct sas_ha_struct *ha = dev->port->ha;
>
> - pm_runtime_get_sync(ha->dev);
> + /*
> + * Non-blocking: a sync resume here would deadlock against
> + * sas_drain_work() during HA resume.
> + */
> + pm_runtime_get_noresume(ha->dev);
[Severity: High]
If the host has already autosuspended when this is called, will this cause
a crash or hardware fault?
Because pm_runtime_get_noresume() does not wake suspended devices,
when lldd_execute_task() is called later in this function:
drivers/scsi/libsas/sas_expander.c:smp_execute_task_sg() {
...
res = i->dft->lldd_execute_task(task, GFP_KERNEL);
...
}
hardware commands would be dispatched to a powered-down host.
> mutex_lock(&dev->ex_dev.cmd_mutex);
> for (retry = 0; retry < 3; retry++) {
> if (test_bit(SAS_DEV_GONE, &dev->state)) {
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260928040234.992912-1-yangxingui@huawei.com?part=1
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH v4] scsi: libsas: Fix SMP IO deadlock during HA resume
2026-09-28 4:14 ` sashiko-bot
@ 2026-09-28 7:30 ` yangxingui
0 siblings, 0 replies; 5+ messages in thread
From: yangxingui @ 2026-09-28 7:30 UTC (permalink / raw)
To: sashiko-reviews; +Cc: linux-scsi
On 2026/9/28 12:14, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
> - [High] scsi: libsas: hardware access on suspended host in smp_execute_task_sg()
> --
>
> commit b6ffaa3a7e18a8e3fec2c6c49ebb1c88915b9f8d
> Author: Xingui Yang <yangxingui@huawei.com>
>
> scsi: libsas: Fix SMP IO deadlock during HA resume
>
> smp_execute_task_sg() calls pm_runtime_get_sync() on the host before
> issuing an SMP command. When that command is itself issued from the
> HA resume path, the get_sync() deadlocks: it waits for the ongoing
> resume (the device is RPM_RESUMING), while the resume is blocked in
> sas_drain_work() waiting for that same SMP IO to complete.
>
>> The usage reference is still required. Discovery work normally runs
>> inside an event worker's PM reference, taken at
>> sas_notify_port_event() notify time and held until the handler has
>> flushed the disco queue. sas_rediscover_ex_phy() however requeues
>> DISCE_REVALIDATE_DOMAIN from within the revalidation worker itself,
>> and flush_workqueue() does not wait for work items queued during
>> execution, so that chained revalidation runs with no outer PM
>> reference - without the get_noresume(), its SMP could race
>> autosuspend.
>
> [Severity: High]
> Does pm_runtime_get_noresume() actually prevent this race if the device
> has already autosuspended?
>
> When chained revalidation occurs in sas_discover.c:sas_revalidate_domain(),
> it runs without an outer PM reference. If the host autosuspends before
> smp_execute_task_sg() runs, pm_runtime_get_noresume() will only increment
> the usage counter and leave the device suspended.
>
>> diff --git a/drivers/scsi/libsas/sas_expander.c b/drivers/scsi/libsas/sas_expander.c
>> index 811c9eb4fef11..26c2099c28b97 100644
>> --- a/drivers/scsi/libsas/sas_expander.c
>> +++ b/drivers/scsi/libsas/sas_expander.c
>> @@ -62,7 +62,11 @@ static int smp_execute_task_sg(struct domain_device *dev,
>> to_sas_internal(dev->port->ha->shost->transportt);
>> struct sas_ha_struct *ha = dev->port->ha;
>>
>> - pm_runtime_get_sync(ha->dev);
>> + /*
>> + * Non-blocking: a sync resume here would deadlock against
>> + * sas_drain_work() during HA resume.
>> + */
>> + pm_runtime_get_noresume(ha->dev);
>
> [Severity: High]
> If the host has already autosuspended when this is called, will this cause
> a crash or hardware fault?
For the chained revalidation the "already
autosuspended" state cannot arise: the host cannot finish
autosuspending while the chained round is queued, because
_suspend_v3_hw() -> sas_suspend_ha() -> sas_drain_work() drains the
disco queue, so the suspend callback blocks until that round, queued
on the same queue, has completed. The round can at worst observe
RPM_SUSPENDING, never RPM_SUSPENDED, and the PCI power state is only
lowered after the driver callback returns - which the drain prevents -
so commands are never dispatched to a powered-down host.
Within the RPM_SUSPENDING window the reference taken by
pm_runtime_get_noresume() is caught by the existing checks: the core
usage check rejects the suspend outright, and if the attempt has
already passed it, the check added by e368d38cb952 ("PM suspend: host
status cannot be suspended") aborts it. We verified this by fault
injection: without the reference the host suspends while the chained
revalidation has an SMP in progress and the command times out - a
recoverable discovery failure, with the reference in place the same
test shows the suspend attempt aborted by that check, with the round
running on an active host.
[162894.251646] hisi_sas_v3_hw 0000:74:04.0: end of resuming controller
[162894.251649] sas: broadcast received: 0
[162894.251664] sas: REVALIDATING DOMAIN on port 0, pid:894458
[162894.259057] sas: SMP 500e004aaaaaaa1f: usage=3 status=0
[162894.259063] sd 5:0:2:0: [sdh] Starting disk
[162894.259066] sd 5:0:3:0: [sdi] Starting disk
[162894.276354] sas: ex 500e004aaaaaaa1f phy00 change count has changed
[162894.352863] sas: INJECT: faking replacement on phy02 (real
5000c5008f23d735)
[162894.361300] sas: ex 500e004aaaaaaa1f phy02 replace 5000c5008f23d735
[162894.375076] smp_execute_task_sg: inject smp timeout
[162900.390107] sd 5:0:3:0: [sdi] Synchronizing SCSI cache
[162900.390110] sd 5:0:2:0: [sdh] Synchronizing SCSI cache
[162900.403138] sd 5:0:2:0: [sdh] Stopping disk
[162900.428158] sd 5:0:3:0: [sdi] Stopping disk
[162916.262081] sas: smp task timed out or aborted
[162916.267906] hisi_sas_v3_hw 0000:74:04.0: abort task: rc=5
[162916.274434] sas: SMP task aborted and not done
[162916.280002] sas: done REVALIDATING DOMAIN on port 0, pid:894458, res
0xffffffba
[162916.297048] hisi_sas_v3_hw 0000:74:04.0: dev[20:1] is gone
[162916.304433] sas: REVALIDATING DOMAIN on port 0, pid:894458
[162916.304437] sas: SMP 500e004aaaaaaa1f: usage=1 status=0
[162916.304455] hisi_sas_v3_hw 0000:74:04.0: entering suspend state
[162916.310802] smp_execute_task_sg: inject smp timeout
[162916.317864] hisi_sas_v3_hw 0000:74:04.0: PM suspend: host status
cannot be suspended // <============ cannot be suspended
[162936.742077] sas: smp task timed out or aborted
[162936.747971] hisi_sas_v3_hw 0000:74:04.0: abort task: rc=5
[162936.754523] sas: SMP task aborted and not done
[162936.760126] sas: done REVALIDATING DOMAIN on port 0, pid:894458, res
0xffffffba
[162936.768890] hisi_sas_v3_hw 0000:74:04.0: entering suspend state
[162937.187359] sas: Enter sas_scsi_recover_host busy: 0 failed: 0
[162937.194389] sas: ata76: end_device-5:0:5: dev error handler
[162937.194395] sas: ata77: end_device-5:0:7: dev error handler
[162937.194431] sas: --- Exit sas_scsi_recover_host: busy: 0 failed: 0
tries: 1
[162937.204475] hisi_sas_v3_hw 0000:74:04.0: dev[19:2] is gone
[162937.211372] hisi_sas_v3_hw 0000:74:04.0: dev[21:1] is gone
[162937.218221] hisi_sas_v3_hw 0000:74:04.0: dev[22:5] is gone
[162937.225065] hisi_sas_v3_hw 0000:74:04.0: dev[23:5] is gone
[162937.231850] hisi_sas_v3_hw 0000:74:04.0: dev[24:1] is gone
[162937.238602] hisi_sas_v3_hw 0000:74:04.0: dev[25:1] is gone
[162937.245397] hisi_sas_v3_hw 0000:74:04.0: dev[26:1] is gone
[162937.252142] hisi_sas_v3_hw 0000:74:04.0: dev[27:1] is gone
[162937.266997] hisi_sas_v3_hw 0000:74:04.0: end of suspending controller
The other callers hold the host active through their own context: the
BSG path resumes it first (sas_smp_handler() calls
pm_runtime_resume_and_get()), event-triggered discovery - including
the ata port probe of a newly found SATA device, which the discovery
work item waits for - runs inside the event workers' PM references,
EH commands run with failed commands still outstanding - which holds
the host active through the hisi_sas device links (16fd4a7c5917) - or
are issued from the resume path itself (RPM_RESUMING, hw_init already
done), and the sysfs PHY paths hold their own reference.
Thanks,
Xingui
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v4] scsi: libsas: Fix SMP IO deadlock during HA resume
2026-09-28 4:02 [PATCH v4] scsi: libsas: Fix SMP IO deadlock during HA resume Xingui Yang
2026-09-28 4:14 ` sashiko-bot
@ 2026-09-30 10:40 ` John Garry
2026-10-08 3:11 ` yangxingui
1 sibling, 1 reply; 5+ messages in thread
From: John Garry @ 2026-09-30 10:40 UTC (permalink / raw)
To: Xingui Yang, yanaijie, jejb, mkp
Cc: linux-scsi, linux-kernel, linuxarm, liuyonglong, kangfenglong
On 9/28/26 05:02, Xingui Yang wrote:
> smp_execute_task_sg() calls pm_runtime_get_sync() on the host before
> issuing an SMP command. When that command is itself issued from the
> HA resume path, the get_sync() deadlocks: it waits for the ongoing
> resume (the device is RPM_RESUMING), while the resume is blocked in
> sas_drain_work() waiting for that same SMP IO to complete.
>
> The deadlock needs an expander-attached SATA disk.
What do you mean by "deadlock needs an expander-attached SATA disk?
> During
> sas_resume_ha() -> sas_drain_work(), DISCE_RESUME ->
> sas_resume_sata() -> ata_sas_port_resume() requests ATA_EH_RESET,
> and the hard reset for such a disk is done via SMP PHY CONTROL
> (sas_ata_hard_reset() -> sas_phy_reset() -> sas_smp_phy_control() ->
> smp_execute_task_sg()). Direct-attached SATA resets through
> lldd_control_phy() and SSP devices use TMFs, so neither hits this.
>
> Replace the get_sync()/put_sync() pair with
> pm_runtime_get_noresume()/pm_runtime_put(). smp_execute_task_sg()
> only needs to hold off autosuspend while the SMP is in flight, and it
> must not try to resume the host: a sync resume issued from the HA
> resume path itself is what deadlocks, and by the time sas_resume_ha()
> runs, hw_init has already reinitialized the hardware, so the device
> is accessible without one.
>
> The usage reference is still required.
Do you mean that usage reference from pm_runtime_get_noresume() is still
required?
> Discovery work normally runs
> inside an event worker's PM reference, taken at
> sas_notify_port_event() notify time and held until the handler has
> flushed the disco queue. sas_rediscover_ex_phy() however requeues
> DISCE_REVALIDATE_DOMAIN from within the revalidation worker itself,
> and flush_workqueue() does not wait for work items queued during
> execution, so that chained revalidation runs with no outer PM
> reference - without the get_noresume(), its SMP could race
> autosuspend.
>
> For the BSG path, sas_smp_handler() is the only caller which may
> find the host autosuspended: expander SMP requests do not go through
> any SCSI device request queue, so nothing else in that path holds
> the host awake. Resume it there with pm_runtime_resume_and_get()
> and check the result.
>
> Fixes: 3dbbbf656b850 ("scsi: libsas: Fix HA resume deadlock and hisi_sas disk-wake race")
> Signed-off-by: Xingui Yang <yangxingui@huawei.com>
Question: do you have a (non-hisi_sas) SAS HBA card whose driver uses
libsas? pm8001 would be such an example. It would be nice to verify that
all these and other non-rpm libsas changes does cause regression there.
> ---
> Changes since v3:
> - Move the host resume to sas_smp_handler(), the only caller which may
> find the host autosuspended, as suggested by John Garry
> - Replace get_sync()/put_sync() with get_noresume()/put() in
> smp_execute_task_sg(): a blocking resume issued from the HA resume
> path itself is what deadlocks
> - Drop the racy SAS_HA_RESUMING check
>
> Changes since v2:
> - Drop the reference with pm_runtime_put() instead of
> pm_runtime_put_noidle().
>
> Changes since v1:
> - Use pm_runtime_get_noresume()/put_noidle() during HA resume instead
> of skipping the PM reference entirely, so an in-flight SMP IO always
> keeps autosuspend away.
> - Convert pm_runtime_get_sync() to pm_runtime_resume_and_get() and
> check the result (pre-existing issue flagged by sashiko).
>
> drivers/scsi/libsas/sas_expander.c | 14 ++++++++++++--
> 1 file changed, 12 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/scsi/libsas/sas_expander.c b/drivers/scsi/libsas/sas_expander.c
> index 811c9eb4fef1..26c2099c28b9 100644
> --- a/drivers/scsi/libsas/sas_expander.c
> +++ b/drivers/scsi/libsas/sas_expander.c
> @@ -62,7 +62,11 @@ static int smp_execute_task_sg(struct domain_device *dev,
> to_sas_internal(dev->port->ha->shost->transportt);
> struct sas_ha_struct *ha = dev->port->ha;
>
> - pm_runtime_get_sync(ha->dev);
> + /*
> + * Non-blocking: a sync resume here would deadlock against
> + * sas_drain_work() during HA resume.
> + */
> + pm_runtime_get_noresume(ha->dev);
> mutex_lock(&dev->ex_dev.cmd_mutex);
> for (retry = 0; retry < 3; retry++) {
> if (test_bit(SAS_DEV_GONE, &dev->state)) {
> @@ -135,7 +139,7 @@ static int smp_execute_task_sg(struct domain_device *dev,
> }
> }
> mutex_unlock(&dev->ex_dev.cmd_mutex);
> - pm_runtime_put_sync(ha->dev);
> + pm_runtime_put(ha->dev);
>
> BUG_ON(retry == 3 && task != NULL);
> sas_free_task(task);
> @@ -2222,6 +2226,11 @@ void sas_smp_handler(struct bsg_job *job, struct Scsi_Host *shost,
> goto out;
> }
>
> + /* The host may have autosuspended, resume it here. */
> + ret = pm_runtime_resume_and_get(dev->port->ha->dev);
> + if (ret)
> + goto out;
> +
> ret = smp_execute_task_sg(dev, job->request_payload.sg_list,
> job->reply_payload.sg_list);
> if (ret >= 0) {
> @@ -2229,6 +2238,7 @@ void sas_smp_handler(struct bsg_job *job, struct Scsi_Host *shost,
> rcvlen = job->reply_payload.payload_len - ret;
> ret = 0;
> }
> + pm_runtime_put(dev->port->ha->dev);
>
> out:
> bsg_job_done(job, ret, rcvlen);
^ permalink raw reply [flat|nested] 5+ messages in thread* Re: [PATCH v4] scsi: libsas: Fix SMP IO deadlock during HA resume
2026-09-30 10:40 ` John Garry
@ 2026-10-08 3:11 ` yangxingui
0 siblings, 0 replies; 5+ messages in thread
From: yangxingui @ 2026-10-08 3:11 UTC (permalink / raw)
To: John Garry, yanaijie, jejb, mkp
Cc: linux-scsi, linux-kernel, linuxarm, liuyonglong, kangfenglong
Hi, John
On 2026/9/30 18:40, John Garry wrote:
> On 9/28/26 05:02, Xingui Yang wrote:
>> smp_execute_task_sg() calls pm_runtime_get_sync() on the host before
>> issuing an SMP command. When that command is itself issued from the
>> HA resume path, the get_sync() deadlocks: it waits for the ongoing
>> resume (the device is RPM_RESUMING), while the resume is blocked in
>> sas_drain_work() waiting for that same SMP IO to complete.
>>
>> The deadlock needs an expander-attached SATA disk.
>
> What do you mean by "deadlock needs an expander-attached SATA disk?
Yes, that is the only configuration whose resume-time hard reset is
implemented as an SMP command. The disk's phy belongs to the
expander, so the ATA EH hard reset is an SMP PHY CONTROL sent to the
expander (sas_ata_hard_reset() -> sas_phy_reset() ->
sas_smp_phy_control() -> smp_execute_task_sg()), a directly attached
SATA disk resets the HBA's own phy through lldd_control_phy(), and
SSP devices are reset through TMFs, so neither issues SMP on resume.
An SMP request from outside the resume path cannot construct the
cycle either: a BSG request arriving while a resume is in progress
blocks waiting for that resume (get_sync() before this patch,
resume_and_get() at the new call site) before it takes the expander
cmd_mutex, and the resume waits for nothing in the BSG context - it
can only be a victim of the deadlock, never a participant.
>
>> During
>> sas_resume_ha() -> sas_drain_work(), DISCE_RESUME ->
>> sas_resume_sata() -> ata_sas_port_resume() requests ATA_EH_RESET,
>> and the hard reset for such a disk is done via SMP PHY CONTROL
>> (sas_ata_hard_reset() -> sas_phy_reset() -> sas_smp_phy_control() ->
>> smp_execute_task_sg()). Direct-attached SATA resets through
>> lldd_control_phy() and SSP devices use TMFs, so neither hits this.
>>
>> Replace the get_sync()/put_sync() pair with
>> pm_runtime_get_noresume()/pm_runtime_put(). smp_execute_task_sg()
>> only needs to hold off autosuspend while the SMP is in flight, and it
>> must not try to resume the host: a sync resume issued from the HA
>> resume path itself is what deadlocks, and by the time sas_resume_ha()
>> runs, hw_init has already reinitialized the hardware, so the device
>> is accessible without one.
>>
>> The usage reference is still required.
>
> Do you mean that usage reference from pm_runtime_get_noresume() is still
> required?
Yes, the reference taken by the pm_runtime_get_noresume() which
replaces the get_sync(). It is not an artefact of the conversion:
the chained revalidation described in the next paragraph runs
outside the event workers' PM references, so without it its SMP
could race autosuspend.
>> Discovery work normally runs
>> inside an event worker's PM reference, taken at
>> sas_notify_port_event() notify time and held until the handler has
>> flushed the disco queue. sas_rediscover_ex_phy() however requeues
>> DISCE_REVALIDATE_DOMAIN from within the revalidation worker itself,
>> and flush_workqueue() does not wait for work items queued during
>> execution, so that chained revalidation runs with no outer PM
>> reference - without the get_noresume(), its SMP could race
>> autosuspend.
>>
>> For the BSG path, sas_smp_handler() is the only caller which may
>> find the host autosuspended: expander SMP requests do not go through
>> any SCSI device request queue, so nothing else in that path holds
>> the host awake. Resume it there with pm_runtime_resume_and_get()
>> and check the result.
>>
>> Fixes: 3dbbbf656b850 ("scsi: libsas: Fix HA resume deadlock and
>> hisi_sas disk-wake race")
>> Signed-off-by: Xingui Yang <yangxingui@huawei.com>
>
> Question: do you have a (non-hisi_sas) SAS HBA card whose driver uses
> libsas? pm8001 would be such an example. It would be nice to verify that
> all these and other non-rpm libsas changes does cause regression there.
We don't have a pm8001 card at hand. From analysis the patch is a
no-op for the LLDDs without runtime PM support: pm8001 and isci
only register system suspend/resume through SIMPLE_DEV_PM_OPS(), and
aic94xx and mvsas register no PM ops at all. For every PCI device
the PCI core forbids runtime PM by default and marks the device
runtime-active (pci_pm_init()), and enables runtime PM later
(pci_bus_add_device()), those devices are thus permanently RPM_ACTIVE
with the usage counter held, and can never runtime-suspend. In that
state get_noresume()/put() only moves the usage counter - exactly
what the get_sync()/put_sync() they replace did - and
pm_runtime_resume_and_get() in sas_smp_handler() succeeds (runtime
PM is enabled and the device is already active), so the new error
check never rejects a request on them.
The only observable difference would be userspace enabling runtime
PM (power/control = auto) on a driver which does not support it, and
that configuration is broken independently of this patch.
Conversely, should an LLDD gain runtime PM support later, it would
need this patch: the deadlock is in the shared resume path - any
runtime-resuming LLDD with an expander-attached SATA disk in the
domain hits it identically - and the reference/resume split
completes the generic protection model, while the LLDD-side
obligations (device links, hw reinit before sas_resume_ha()) are
unchanged.
Thanks,
Xingui
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-10-08 3:11 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-28 4:02 [PATCH v4] scsi: libsas: Fix SMP IO deadlock during HA resume Xingui Yang
2026-09-28 4:14 ` sashiko-bot
2026-09-28 7:30 ` yangxingui
2026-09-30 10:40 ` John Garry
2026-10-08 3:11 ` yangxingui
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox