From: 李真能 <lizhenneng@kylinos.cn>
To: "Christian König" <ckoenig.leichtzumerken@gmail.com>,
"Alex Deucher" <alexander.deucher@amd.com>
Cc: "David Airlie" <airlied@linux.ie>,
"Pan Xinhui" <Xinhui.Pan@amd.com>,
linux-kernel@vger.kernel.org, amd-gfx@lists.freedesktop.org,
dri-devel@lists.freedesktop.org,
"Daniel Vetter" <daniel@ffwll.ch>,
"Christian König" <christian.koenig@amd.com>
Subject: Re: [PATCH] drm/radeon: add a force flush to delay work when radeon
Date: Mon, 15 Aug 2022 15:34:18 +0800 [thread overview]
Message-ID: <db2a43da-256d-402e-882b-c05436d4e83b@kylinos.cn> (raw)
In-Reply-To: <b23e4037-2030-32d0-d626-b5a846fcafeb@gmail.com>
在 2022/8/12 18:55, Christian König 写道:
> Am 11.08.22 um 09:25 schrieb Zhenneng Li:
>> Although radeon card fence and wait for gpu to finish processing
>> current batch rings,
>> there is still a corner case that radeon lockup work queue may not be
>> fully flushed,
>> and meanwhile the radeon_suspend_kms() function has called
>> pci_set_power_state() to
>> put device in D3hot state.
>
> If I'm not completely mistaken the reset worker uses the
> suspend/resume functionality as well to get the hardware into a
> working state again.
>
> So if I'm not completely mistaken this here would lead to a deadlock,
> please double check that.
We have tested many times, there are no deadlock.
In which situation, there would lead to a deadlock?
>
> Regards,
> Christian.
>
>> Per PCI spec rev 4.0 on 5.3.1.4.1 D3hot State.
>>> Configuration and Message requests are the only TLPs accepted by a
>>> Function in
>>> the D3hot state. All other received Requests must be handled as
>>> Unsupported Requests,
>>> and all received Completions may optionally be handled as Unexpected
>>> Completions.
>> This issue will happen in following logs:
>> Unable to handle kernel paging request at virtual address
>> 00008800e0008010
>> CPU 0 kworker/0:3(131): Oops 0
>> pc = [<ffffffff811bea5c>] ra = [<ffffffff81240844>] ps = 0000
>> Tainted: G W
>> pc is at si_gpu_check_soft_reset+0x3c/0x240
>> ra is at si_dma_is_lockup+0x34/0xd0
>> v0 = 0000000000000000 t0 = fff08800e0008010 t1 = 0000000000010000
>> t2 = 0000000000008010 t3 = fff00007e3c00000 t4 = fff00007e3c00258
>> t5 = 000000000000ffff t6 = 0000000000000001 t7 = fff00007ef078000
>> s0 = fff00007e3c016e8 s1 = fff00007e3c00000 s2 = fff00007e3c00018
>> s3 = fff00007e3c00000 s4 = fff00007fff59d80 s5 = 0000000000000000
>> s6 = fff00007ef07bd98
>> a0 = fff00007e3c00000 a1 = fff00007e3c016e8 a2 = 0000000000000008
>> a3 = 0000000000000001 a4 = 8f5c28f5c28f5c29 a5 = ffffffff810f4338
>> t8 = 0000000000000275 t9 = ffffffff809b66f8 t10 = ff6769c5d964b800
>> t11= 000000000000b886 pv = ffffffff811bea20 at = 0000000000000000
>> gp = ffffffff81d89690 sp = 00000000aa814126
>> Disabling lock debugging due to kernel taint
>> Trace:
>> [<ffffffff81240844>] si_dma_is_lockup+0x34/0xd0
>> [<ffffffff81119610>] radeon_fence_check_lockup+0xd0/0x290
>> [<ffffffff80977010>] process_one_work+0x280/0x550
>> [<ffffffff80977350>] worker_thread+0x70/0x7c0
>> [<ffffffff80977410>] worker_thread+0x130/0x7c0
>> [<ffffffff80982040>] kthread+0x200/0x210
>> [<ffffffff809772e0>] worker_thread+0x0/0x7c0
>> [<ffffffff80981f8c>] kthread+0x14c/0x210
>> [<ffffffff80911658>] ret_from_kernel_thread+0x18/0x20
>> [<ffffffff80981e40>] kthread+0x0/0x210
>> Code: ad3e0008 43f0074a ad7e0018 ad9e0020 8c3001e8 40230101
>> <88210000> 4821ed21
>> So force lockup work queue flush to fix this problem.
>>
>> Signed-off-by: Zhenneng Li <lizhenneng@kylinos.cn>
>> ---
>> drivers/gpu/drm/radeon/radeon_device.c | 3 +++
>> 1 file changed, 3 insertions(+)
>>
>> diff --git a/drivers/gpu/drm/radeon/radeon_device.c
>> b/drivers/gpu/drm/radeon/radeon_device.c
>> index 15692cb241fc..e608ca26780a 100644
>> --- a/drivers/gpu/drm/radeon/radeon_device.c
>> +++ b/drivers/gpu/drm/radeon/radeon_device.c
>> @@ -1604,6 +1604,9 @@ int radeon_suspend_kms(struct drm_device *dev,
>> bool suspend,
>> if (r) {
>> /* delay GPU reset to resume */
>> radeon_fence_driver_force_completion(rdev, i);
>> + } else {
>> + /* finish executing delayed work */
>> + flush_delayed_work(&rdev->fence_drv[i].lockup_work);
>> }
>> }
>
WARNING: multiple messages have this Message-ID (diff)
From: 李真能 <lizhenneng@kylinos.cn>
To: "Christian König" <ckoenig.leichtzumerken@gmail.com>,
"Alex Deucher" <alexander.deucher@amd.com>
Cc: "David Airlie" <airlied@linux.ie>,
"Pan Xinhui" <Xinhui.Pan@amd.com>,
linux-kernel@vger.kernel.org, amd-gfx@lists.freedesktop.org,
dri-devel@lists.freedesktop.org,
"Christian König" <christian.koenig@amd.com>
Subject: Re: [PATCH] drm/radeon: add a force flush to delay work when radeon
Date: Mon, 15 Aug 2022 15:34:18 +0800 [thread overview]
Message-ID: <db2a43da-256d-402e-882b-c05436d4e83b@kylinos.cn> (raw)
In-Reply-To: <b23e4037-2030-32d0-d626-b5a846fcafeb@gmail.com>
在 2022/8/12 18:55, Christian König 写道:
> Am 11.08.22 um 09:25 schrieb Zhenneng Li:
>> Although radeon card fence and wait for gpu to finish processing
>> current batch rings,
>> there is still a corner case that radeon lockup work queue may not be
>> fully flushed,
>> and meanwhile the radeon_suspend_kms() function has called
>> pci_set_power_state() to
>> put device in D3hot state.
>
> If I'm not completely mistaken the reset worker uses the
> suspend/resume functionality as well to get the hardware into a
> working state again.
>
> So if I'm not completely mistaken this here would lead to a deadlock,
> please double check that.
We have tested many times, there are no deadlock.
In which situation, there would lead to a deadlock?
>
> Regards,
> Christian.
>
>> Per PCI spec rev 4.0 on 5.3.1.4.1 D3hot State.
>>> Configuration and Message requests are the only TLPs accepted by a
>>> Function in
>>> the D3hot state. All other received Requests must be handled as
>>> Unsupported Requests,
>>> and all received Completions may optionally be handled as Unexpected
>>> Completions.
>> This issue will happen in following logs:
>> Unable to handle kernel paging request at virtual address
>> 00008800e0008010
>> CPU 0 kworker/0:3(131): Oops 0
>> pc = [<ffffffff811bea5c>] ra = [<ffffffff81240844>] ps = 0000
>> Tainted: G W
>> pc is at si_gpu_check_soft_reset+0x3c/0x240
>> ra is at si_dma_is_lockup+0x34/0xd0
>> v0 = 0000000000000000 t0 = fff08800e0008010 t1 = 0000000000010000
>> t2 = 0000000000008010 t3 = fff00007e3c00000 t4 = fff00007e3c00258
>> t5 = 000000000000ffff t6 = 0000000000000001 t7 = fff00007ef078000
>> s0 = fff00007e3c016e8 s1 = fff00007e3c00000 s2 = fff00007e3c00018
>> s3 = fff00007e3c00000 s4 = fff00007fff59d80 s5 = 0000000000000000
>> s6 = fff00007ef07bd98
>> a0 = fff00007e3c00000 a1 = fff00007e3c016e8 a2 = 0000000000000008
>> a3 = 0000000000000001 a4 = 8f5c28f5c28f5c29 a5 = ffffffff810f4338
>> t8 = 0000000000000275 t9 = ffffffff809b66f8 t10 = ff6769c5d964b800
>> t11= 000000000000b886 pv = ffffffff811bea20 at = 0000000000000000
>> gp = ffffffff81d89690 sp = 00000000aa814126
>> Disabling lock debugging due to kernel taint
>> Trace:
>> [<ffffffff81240844>] si_dma_is_lockup+0x34/0xd0
>> [<ffffffff81119610>] radeon_fence_check_lockup+0xd0/0x290
>> [<ffffffff80977010>] process_one_work+0x280/0x550
>> [<ffffffff80977350>] worker_thread+0x70/0x7c0
>> [<ffffffff80977410>] worker_thread+0x130/0x7c0
>> [<ffffffff80982040>] kthread+0x200/0x210
>> [<ffffffff809772e0>] worker_thread+0x0/0x7c0
>> [<ffffffff80981f8c>] kthread+0x14c/0x210
>> [<ffffffff80911658>] ret_from_kernel_thread+0x18/0x20
>> [<ffffffff80981e40>] kthread+0x0/0x210
>> Code: ad3e0008 43f0074a ad7e0018 ad9e0020 8c3001e8 40230101
>> <88210000> 4821ed21
>> So force lockup work queue flush to fix this problem.
>>
>> Signed-off-by: Zhenneng Li <lizhenneng@kylinos.cn>
>> ---
>> drivers/gpu/drm/radeon/radeon_device.c | 3 +++
>> 1 file changed, 3 insertions(+)
>>
>> diff --git a/drivers/gpu/drm/radeon/radeon_device.c
>> b/drivers/gpu/drm/radeon/radeon_device.c
>> index 15692cb241fc..e608ca26780a 100644
>> --- a/drivers/gpu/drm/radeon/radeon_device.c
>> +++ b/drivers/gpu/drm/radeon/radeon_device.c
>> @@ -1604,6 +1604,9 @@ int radeon_suspend_kms(struct drm_device *dev,
>> bool suspend,
>> if (r) {
>> /* delay GPU reset to resume */
>> radeon_fence_driver_force_completion(rdev, i);
>> + } else {
>> + /* finish executing delayed work */
>> + flush_delayed_work(&rdev->fence_drv[i].lockup_work);
>> }
>> }
>
WARNING: multiple messages have this Message-ID (diff)
From: 李真能 <lizhenneng@kylinos.cn>
To: "Christian König" <ckoenig.leichtzumerken@gmail.com>,
"Alex Deucher" <alexander.deucher@amd.com>
Cc: "David Airlie" <airlied@linux.ie>,
"Pan Xinhui" <Xinhui.Pan@amd.com>,
linux-kernel@vger.kernel.org, dri-devel@lists.freedesktop.org,
amd-gfx@lists.freedesktop.org, "Daniel Vetter" <daniel@ffwll.ch>,
"Christian König" <christian.koenig@amd.com>
Subject: Re: [PATCH] drm/radeon: add a force flush to delay work when radeon
Date: Mon, 15 Aug 2022 15:34:18 +0800 [thread overview]
Message-ID: <db2a43da-256d-402e-882b-c05436d4e83b@kylinos.cn> (raw)
In-Reply-To: <b23e4037-2030-32d0-d626-b5a846fcafeb@gmail.com>
在 2022/8/12 18:55, Christian König 写道:
> Am 11.08.22 um 09:25 schrieb Zhenneng Li:
>> Although radeon card fence and wait for gpu to finish processing
>> current batch rings,
>> there is still a corner case that radeon lockup work queue may not be
>> fully flushed,
>> and meanwhile the radeon_suspend_kms() function has called
>> pci_set_power_state() to
>> put device in D3hot state.
>
> If I'm not completely mistaken the reset worker uses the
> suspend/resume functionality as well to get the hardware into a
> working state again.
>
> So if I'm not completely mistaken this here would lead to a deadlock,
> please double check that.
We have tested many times, there are no deadlock.
In which situation, there would lead to a deadlock?
>
> Regards,
> Christian.
>
>> Per PCI spec rev 4.0 on 5.3.1.4.1 D3hot State.
>>> Configuration and Message requests are the only TLPs accepted by a
>>> Function in
>>> the D3hot state. All other received Requests must be handled as
>>> Unsupported Requests,
>>> and all received Completions may optionally be handled as Unexpected
>>> Completions.
>> This issue will happen in following logs:
>> Unable to handle kernel paging request at virtual address
>> 00008800e0008010
>> CPU 0 kworker/0:3(131): Oops 0
>> pc = [<ffffffff811bea5c>] ra = [<ffffffff81240844>] ps = 0000
>> Tainted: G W
>> pc is at si_gpu_check_soft_reset+0x3c/0x240
>> ra is at si_dma_is_lockup+0x34/0xd0
>> v0 = 0000000000000000 t0 = fff08800e0008010 t1 = 0000000000010000
>> t2 = 0000000000008010 t3 = fff00007e3c00000 t4 = fff00007e3c00258
>> t5 = 000000000000ffff t6 = 0000000000000001 t7 = fff00007ef078000
>> s0 = fff00007e3c016e8 s1 = fff00007e3c00000 s2 = fff00007e3c00018
>> s3 = fff00007e3c00000 s4 = fff00007fff59d80 s5 = 0000000000000000
>> s6 = fff00007ef07bd98
>> a0 = fff00007e3c00000 a1 = fff00007e3c016e8 a2 = 0000000000000008
>> a3 = 0000000000000001 a4 = 8f5c28f5c28f5c29 a5 = ffffffff810f4338
>> t8 = 0000000000000275 t9 = ffffffff809b66f8 t10 = ff6769c5d964b800
>> t11= 000000000000b886 pv = ffffffff811bea20 at = 0000000000000000
>> gp = ffffffff81d89690 sp = 00000000aa814126
>> Disabling lock debugging due to kernel taint
>> Trace:
>> [<ffffffff81240844>] si_dma_is_lockup+0x34/0xd0
>> [<ffffffff81119610>] radeon_fence_check_lockup+0xd0/0x290
>> [<ffffffff80977010>] process_one_work+0x280/0x550
>> [<ffffffff80977350>] worker_thread+0x70/0x7c0
>> [<ffffffff80977410>] worker_thread+0x130/0x7c0
>> [<ffffffff80982040>] kthread+0x200/0x210
>> [<ffffffff809772e0>] worker_thread+0x0/0x7c0
>> [<ffffffff80981f8c>] kthread+0x14c/0x210
>> [<ffffffff80911658>] ret_from_kernel_thread+0x18/0x20
>> [<ffffffff80981e40>] kthread+0x0/0x210
>> Code: ad3e0008 43f0074a ad7e0018 ad9e0020 8c3001e8 40230101
>> <88210000> 4821ed21
>> So force lockup work queue flush to fix this problem.
>>
>> Signed-off-by: Zhenneng Li <lizhenneng@kylinos.cn>
>> ---
>> drivers/gpu/drm/radeon/radeon_device.c | 3 +++
>> 1 file changed, 3 insertions(+)
>>
>> diff --git a/drivers/gpu/drm/radeon/radeon_device.c
>> b/drivers/gpu/drm/radeon/radeon_device.c
>> index 15692cb241fc..e608ca26780a 100644
>> --- a/drivers/gpu/drm/radeon/radeon_device.c
>> +++ b/drivers/gpu/drm/radeon/radeon_device.c
>> @@ -1604,6 +1604,9 @@ int radeon_suspend_kms(struct drm_device *dev,
>> bool suspend,
>> if (r) {
>> /* delay GPU reset to resume */
>> radeon_fence_driver_force_completion(rdev, i);
>> + } else {
>> + /* finish executing delayed work */
>> + flush_delayed_work(&rdev->fence_drv[i].lockup_work);
>> }
>> }
>
next prev parent reply other threads:[~2022-08-15 13:11 UTC|newest]
Thread overview: 24+ messages / expand[flat|nested] mbox.gz Atom feed top
2022-08-11 7:25 [PATCH] drm/radeon: add a force flush to delay work when radeon Zhenneng Li
2022-08-11 7:25 ` Zhenneng Li
2022-08-11 7:25 ` Zhenneng Li
2022-08-12 10:55 ` Christian König
2022-08-12 10:55 ` Christian König
2022-08-12 10:55 ` Christian König
2022-08-15 7:34 ` 李真能 [this message]
2022-08-15 7:34 ` 李真能
2022-08-15 7:34 ` 李真能
2022-08-15 13:12 ` Christian König
2022-08-15 13:12 ` Christian König
2022-08-15 13:12 ` Christian König
2022-08-17 7:31 ` 李真能
2022-08-17 7:31 ` 李真能
2022-08-17 7:31 ` 李真能
2022-08-17 11:40 ` Christian König
2022-08-17 11:40 ` Christian König
2022-08-17 11:40 ` Christian König
2022-08-19 9:34 ` 李真能
2022-08-19 9:34 ` 李真能
2022-08-19 10:07 ` Christian König
2022-08-19 10:07 ` Christian König
2022-08-19 15:55 ` Alex Deucher
2022-08-19 15:55 ` Alex Deucher
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=db2a43da-256d-402e-882b-c05436d4e83b@kylinos.cn \
--to=lizhenneng@kylinos.cn \
--cc=Xinhui.Pan@amd.com \
--cc=airlied@linux.ie \
--cc=alexander.deucher@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=christian.koenig@amd.com \
--cc=ckoenig.leichtzumerken@gmail.com \
--cc=daniel@ffwll.ch \
--cc=dri-devel@lists.freedesktop.org \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.