amd-gfx.lists.freedesktop.org archive mirror
 help / color / mirror / Atom feed
* Hard lockups with ROCM
@ 2019-05-13  1:44 Daniel Kasak
       [not found] ` <CAF73Y=QuYq3ALtP6xiPyqS+jm_TJCQQDyQ+WA5ZJG8EhWSKiTw-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  0 siblings, 1 reply; 6+ messages in thread
From: Daniel Kasak @ 2019-05-13  1:44 UTC (permalink / raw)
  To: amd-gfx-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW


[-- Attachment #1.1: Type: text/plain, Size: 917 bytes --]

Hi all. I had version 2.2.0 of the ROCM stack running on a 5.0.x and 5.1.0
kernel. Things were going great with various boinc GPU tasks. But there is
a setiathome GPU task which reliably gives me a hard lockup within about 30
minutes of running. I actually had to do *two* emergency re-installs over
the past week. Perhaps part of this was my fault ( running btrfs with lzo
compression on my root partition ... ). But absolutely part of this was the
hard lockups. I've tested all kinds of other things ( eg rebuilding lots of
stuff under Gentoo ) ... I don't have a general stability issue even under
hours of high load. But after restarting boinc with that same setiathome
task ... <bang>!

If someone wants me to sacrifice another installation, they can point me to
instructions for trying to gather more information.

Anyway ... perhaps more work around detecting and recovering from GPU
lockups is in order?

Dan

[-- Attachment #1.2: Type: text/html, Size: 1040 bytes --]

[-- Attachment #2: Type: text/plain, Size: 153 bytes --]

_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Hard lockups with ROCM
       [not found] ` <CAF73Y=QuYq3ALtP6xiPyqS+jm_TJCQQDyQ+WA5ZJG8EhWSKiTw-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2019-05-16  0:33   ` Daniel Kasak
       [not found]     ` <CAF73Y=R96zUxCAEKopSvGReqB+sEFWcHhXSKnR98rpetMbKf4Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  2019-05-16 19:03   ` Kuehling, Felix
  1 sibling, 1 reply; 6+ messages in thread
From: Daniel Kasak @ 2019-05-16  0:33 UTC (permalink / raw)
  To: amd-gfx-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW


[-- Attachment #1.1: Type: text/plain, Size: 1101 bytes --]

On Mon, May 13, 2019 at 11:44 AM Daniel Kasak <d.j.kasak.dk-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>
wrote:

> Hi all. I had version 2.2.0 of the ROCM stack running on a 5.0.x and 5.1.0
> kernel. Things were going great with various boinc GPU tasks. But there is
> a setiathome GPU task which reliably gives me a hard lockup within about 30
> minutes of running. I actually had to do *two* emergency re-installs over
> the past week. Perhaps part of this was my fault ( running btrfs with lzo
> compression on my root partition ... ). But absolutely part of this was the
> hard lockups. I've tested all kinds of other things ( eg rebuilding lots of
> stuff under Gentoo ) ... I don't have a general stability issue even under
> hours of high load. But after restarting boinc with that same setiathome
> task ... <bang>!
>
> If someone wants me to sacrifice another installation, they can point me
> to instructions for trying to gather more information.
>
> Anyway ... perhaps more work around detecting and recovering from GPU
> lockups is in order?
>
> Dan
>

<sigh>

That's what I was afraid of :(

[-- Attachment #1.2: Type: text/html, Size: 1559 bytes --]

[-- Attachment #2: Type: text/plain, Size: 153 bytes --]

_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Hard lockups with ROCM
       [not found]     ` <CAF73Y=R96zUxCAEKopSvGReqB+sEFWcHhXSKnR98rpetMbKf4Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2019-05-16  1:43       ` Alex Deucher
       [not found]         ` <CADnq5_OOEP+YsQx12cOBaM3NdjM=eGAbbbudEusqf5rNRz8C2Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  0 siblings, 1 reply; 6+ messages in thread
From: Alex Deucher @ 2019-05-16  1:43 UTC (permalink / raw)
  To: Daniel Kasak; +Cc: amd-gfx list

On Wed, May 15, 2019 at 8:33 PM Daniel Kasak <d.j.kasak.dk@gmail.com> wrote:
>
> On Mon, May 13, 2019 at 11:44 AM Daniel Kasak <d.j.kasak.dk@gmail.com> wrote:
>>
>> Hi all. I had version 2.2.0 of the ROCM stack running on a 5.0.x and 5.1.0 kernel. Things were going great with various boinc GPU tasks. But there is a setiathome GPU task which reliably gives me a hard lockup within about 30 minutes of running. I actually had to do *two* emergency re-installs over the past week. Perhaps part of this was my fault ( running btrfs with lzo compression on my root partition ... ). But absolutely part of this was the hard lockups. I've tested all kinds of other things ( eg rebuilding lots of stuff under Gentoo ) ... I don't have a general stability issue even under hours of high load. But after restarting boinc with that same setiathome task ... <bang>!
>>
>> If someone wants me to sacrifice another installation, they can point me to instructions for trying to gather more information.
>>
>> Anyway ... perhaps more work around detecting and recovering from GPU lockups is in order?
>>
>> Dan
>
>
> <sigh>
>
> That's what I was afraid of :(

Not sure what you were afraid of.  I don't think anyone has looked at
setiathome on ROCm.  I'd suggest filing a bug
(https://bugs.freedesktop.org) and attaching your dmesg output and
xorg log (if using X).  If there is a GPU reset, note that you will
need to restart your desktop environment because currently neither
glamor or any compositors support GL robustness extensions to reset
their contexts after a GPU reset.

Alex
_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Hard lockups with ROCM
       [not found]         ` <CADnq5_OOEP+YsQx12cOBaM3NdjM=eGAbbbudEusqf5rNRz8C2Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2019-05-16 11:52           ` Daniel Kasak
       [not found]             ` <CAF73Y=TfsjO+-cnvzZ8s7VnuCTcu3KT0ZG3KtzbYGTTYKK-PzA-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  0 siblings, 1 reply; 6+ messages in thread
From: Daniel Kasak @ 2019-05-16 11:52 UTC (permalink / raw)
  To: Alex Deucher; +Cc: amd-gfx list


[-- Attachment #1.1: Type: text/plain, Size: 2245 bytes --]

On Thu, May 16, 2019 at 11:43 AM Alex Deucher <alexdeucher-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:

> On Wed, May 15, 2019 at 8:33 PM Daniel Kasak <d.j.kasak.dk-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>
> wrote:
> >
> > On Mon, May 13, 2019 at 11:44 AM Daniel Kasak <d.j.kasak.dk-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>
> wrote:
> >>
> >> Hi all. I had version 2.2.0 of the ROCM stack running on a 5.0.x and
> 5.1.0 kernel. Things were going great with various boinc GPU tasks. But
> there is a setiathome GPU task which reliably gives me a hard lockup within
> about 30 minutes of running. I actually had to do *two* emergency
> re-installs over the past week. Perhaps part of this was my fault ( running
> btrfs with lzo compression on my root partition ... ). But absolutely part
> of this was the hard lockups. I've tested all kinds of other things ( eg
> rebuilding lots of stuff under Gentoo ) ... I don't have a general
> stability issue even under hours of high load. But after restarting boinc
> with that same setiathome task ... <bang>!
> >>
> >> If someone wants me to sacrifice another installation, they can point
> me to instructions for trying to gather more information.
> >>
> >> Anyway ... perhaps more work around detecting and recovering from GPU
> lockups is in order?
> >>
> >> Dan
> >
> >
> > <sigh>
> >
> > That's what I was afraid of :(
>
> Not sure what you were afraid of.  I don't think anyone has looked at
> setiathome on ROCm.  I'd suggest filing a bug
> (https://bugs.freedesktop.org) and attaching your dmesg output and
> xorg log (if using X).  If there is a GPU reset, note that you will
> need to restart your desktop environment because currently neither
> glamor or any compositors support GL robustness extensions to reset
> their contexts after a GPU reset.
>
> Alex
>

Hi Alex. dmesg output is not available ... this is a *hard* lockup. I need
to power-cycle after it happens ( ALT + SysRq + { S , U , B } doesn't even
work ). That's why I asked for instructions to possibly gather more info. I
did check the xorg log after I did an emergency export of my filesystem ...
nothing of interest in there. It seems like I currently don't really have
enough info to make a bug report worthwhile.

Dan

[-- Attachment #1.2: Type: text/html, Size: 3094 bytes --]

[-- Attachment #2: Type: text/plain, Size: 153 bytes --]

_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Hard lockups with ROCM
       [not found]             ` <CAF73Y=TfsjO+-cnvzZ8s7VnuCTcu3KT0ZG3KtzbYGTTYKK-PzA-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
@ 2019-05-16 15:56               ` Paul Menzel
  0 siblings, 0 replies; 6+ messages in thread
From: Paul Menzel @ 2019-05-16 15:56 UTC (permalink / raw)
  To: Daniel Kasak, Alex Deucher; +Cc: amd-gfx-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW


[-- Attachment #1.1: Type: text/plain, Size: 2676 bytes --]

Dear Daniel,


On 05/16/2019 01:52 PM, Daniel Kasak wrote:
> On Thu, May 16, 2019 at 11:43 AM Alex Deucher <alexdeucher-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
> 
>> On Wed, May 15, 2019 at 8:33 PM Daniel Kasak <d.j.kasak.dk-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>
>> wrote:
>>>
>>> On Mon, May 13, 2019 at 11:44 AM Daniel Kasak <d.j.kasak.dk-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>
>> wrote:
>>>>
>>>> Hi all. I had version 2.2.0 of the ROCM stack running on a 5.0.x and
>> 5.1.0 kernel. Things were going great with various boinc GPU tasks. But
>> there is a setiathome GPU task which reliably gives me a hard lockup within
>> about 30 minutes of running. I actually had to do *two* emergency
>> re-installs over the past week. Perhaps part of this was my fault ( running
>> btrfs with lzo compression on my root partition ... ). But absolutely part
>> of this was the hard lockups. I've tested all kinds of other things ( eg
>> rebuilding lots of stuff under Gentoo ) ... I don't have a general
>> stability issue even under hours of high load. But after restarting boinc
>> with that same setiathome task ... <bang>!
>>>>
>>>> If someone wants me to sacrifice another installation, they can point
>> me to instructions for trying to gather more information.
>>>>
>>>> Anyway ... perhaps more work around detecting and recovering from GPU
>> lockups is in order?

>>> <sigh>
>>>
>>> That's what I was afraid of :(
>>
>> Not sure what you were afraid of.  I don't think anyone has looked at
>> setiathome on ROCm.  I'd suggest filing a bug
>> (https://bugs.freedesktop.org) and attaching your dmesg output and
>> xorg log (if using X).  If there is a GPU reset, note that you will
>> need to restart your desktop environment because currently neither
>> glamor or any compositors support GL robustness extensions to reset
>> their contexts after a GPU reset.

> Hi Alex. dmesg output is not available ... this is a *hard* lockup. I need
> to power-cycle after it happens ( ALT + SysRq + { S , U , B } doesn't even
> work ). That's why I asked for instructions to possibly gather more info. I
> did check the xorg log after I did an emergency export of my filesystem ...
> nothing of interest in there. It seems like I currently don't really have
> enough info to make a bug report worthwhile.

Does your board have a serial port? If yes, please use the serial console to
gather the messages on another system.

Sometimes the netconsole [1] is also supposed to be able to send the last
Linux messages out.


Kind regards,

Paul


[1]: https://www.kernel.org/doc/Documentation/networking/netconsole.txt


[-- Attachment #1.2: S/MIME Cryptographic Signature --]
[-- Type: application/pkcs7-signature, Size: 5174 bytes --]

[-- Attachment #2: Type: text/plain, Size: 153 bytes --]

_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: Hard lockups with ROCM
       [not found] ` <CAF73Y=QuYq3ALtP6xiPyqS+jm_TJCQQDyQ+WA5ZJG8EhWSKiTw-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
  2019-05-16  0:33   ` Daniel Kasak
@ 2019-05-16 19:03   ` Kuehling, Felix
  1 sibling, 0 replies; 6+ messages in thread
From: Kuehling, Felix @ 2019-05-16 19:03 UTC (permalink / raw)
  To: amd-gfx-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org,
	Daniel Kasak

Hi Daniel,

On 2019-05-12 9:44 p.m., Daniel Kasak wrote:
> [CAUTION: External Email]
> Hi all. I had version 2.2.0 of the ROCM stack running on a 5.0.x and 
> 5.1.0 kernel. Things were going great with various boinc GPU tasks. 
> But there is a setiathome GPU task which reliably gives me a hard 
> lockup within about 30 minutes of running. I actually had to do *two* 
> emergency re-installs over the past week.

Sorry to hear about your trouble. Do you have a second computer you can 
use to remote login into your system? Chances are that it's still 
responsive and only the screen is frozen.

Also, you could try booting in console mode (without an xserver). The 
console usually still works even when the GPU compute units or SDMA 
engines are hanging.

If you manage to do an emergency reboot with sysrq (remount-RO and 
reboot), you should see the kernel log of your previous session in 
/var/log. On Ubuntu it's in /var/log/kern.log. Not sure where it is on 
Gentoo. There is a good chance the log contains helpful information 
(e.g. if the driver detected a hang but failed to reset the GPU, or 
maybe a driver bug that leads to a deadlock or kernel panic).

> Perhaps part of this was my fault ( running btrfs with lzo compression 
> on my root partition ... ). But absolutely part of this was the hard 
> lockups. I've tested all kinds of other things ( eg rebuilding lots of 
> stuff under Gentoo ) ... I don't have a general stability issue even 
> under hours of high load. But after restarting boinc with that same 
> setiathome task ... <bang>!
>
> If someone wants me to sacrifice another installation, they can point 
> me to instructions for trying to gather more information.

If you want to risk another installation, it may be a good idea to do it 
on a spare hard drive, or a spare partition on your existing hard drive. 
Also, use a more conventional choice of file system. A simple ext4 is 
pretty robust in my experience. We get hard lockups all the time. I 
usually only reinstall my system for big OS upgrades or if I'm stupid 
and mess something up myself.

Which GPU are you using?

There are some things you could try to narrow down the cause of your 
problem.

 1. Monitor GPU temperature while running setiathome
 2. If you're building your own kernel, enable some helpful kernel debug
    options that can provide very helpful diagnostic info: lock
    debugging, memory debugging, lockup/hang debugging
 3. Try running with lower GPU clocks (rocm-smi --setperflevel low). If
    that fixes it, you may have inadequate cooling or power supply
 4. Try running in console mode (without Xserver or other graphical UI
    running). If that fixes it, there may be a bad interaction between
    graphics and compute
 5. Try updating your firmware. The DKMS package included in our ROCm
    releases includes the latest firmware. You should be able to extract
    it from there and drop it into /lib/firmware/amdgpu
 6. Try to find a regression point. Is there any known version of ROCm
    or the kernel where it worked correctly?

Regards,
   Felix


>
> Anyway ... perhaps more work around detecting and recovering from GPU 
> lockups is in order?
>
> Dan
>
> _______________________________________________
> amd-gfx mailing list
> amd-gfx@lists.freedesktop.org
> https://lists.freedesktop.org/mailman/listinfo/amd-gfx
_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2019-05-16 19:03 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2019-05-13  1:44 Hard lockups with ROCM Daniel Kasak
     [not found] ` <CAF73Y=QuYq3ALtP6xiPyqS+jm_TJCQQDyQ+WA5ZJG8EhWSKiTw-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2019-05-16  0:33   ` Daniel Kasak
     [not found]     ` <CAF73Y=R96zUxCAEKopSvGReqB+sEFWcHhXSKnR98rpetMbKf4Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2019-05-16  1:43       ` Alex Deucher
     [not found]         ` <CADnq5_OOEP+YsQx12cOBaM3NdjM=eGAbbbudEusqf5rNRz8C2Q-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2019-05-16 11:52           ` Daniel Kasak
     [not found]             ` <CAF73Y=TfsjO+-cnvzZ8s7VnuCTcu3KT0ZG3KtzbYGTTYKK-PzA-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2019-05-16 15:56               ` Paul Menzel
2019-05-16 19:03   ` Kuehling, Felix

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).