From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 105113] [hawaii, radeonsi, clover] Running Piglit cl/program/execute/{, tail-}calls{, -struct, -workitem-id}.cl cause GPU VM error and ring stalled GPU lockup Date: Sun, 18 Nov 2018 19:24:54 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0122903116==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id CAF6989C56 for ; Sun, 18 Nov 2018 19:24:54 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0122903116== Content-Type: multipart/alternative; boundary="15425690944.51Ee8.31720" Content-Transfer-Encoding: 7bit --15425690944.51Ee8.31720 Date: Sun, 18 Nov 2018 19:24:54 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D105113 --- Comment #7 from Jan Vesely --- (In reply to Maciej S. Szmigiero from comment #6) > There are really two issues at play here: > 1) If the LLVM-generated code cannot be run properly then it should be si= mply > rejected by whatever is actually in charge of submitting it to the GPU (I > guess > this would be Mesa?). > This way an application will know it cannot use OpenCL for computation, at > least > not with this compute kernel. >=20 > Instead, it currently looks like many of these test run but give incorrect > results, which is obviously rather bad. Do you have an example of this? clover should return OUT_OF_RESOURCES error when the compute state creation fails (like in the presence of code relocations). It does not change the content of the buffer, so it will return whatever was stored in the buffer on creation. > 2) Some (previous) Mesa + LLVM versions generate a command stream that > crashes the GPU and, as far as I can remember, sometimes even lockup the > whole machine. >=20 > It should not be possible to crash the GPU, regardless how incorrect a > command stream that userspace sends to it is - because otherwise it is > possible for > an unprivileged user with GPU access to DoS the machine. This is a separate issue. GPU hangs are generally addressed via gpu reset w= hich should be enabled for gfx8/9 GPUs in recent amdgpu.ko [0] [0] https://patchwork.freedesktop.org/patch/257994/ --=20 You are receiving this mail because: You are the assignee for the bug.= --15425690944.51Ee8.31720 Date: Sun, 18 Nov 2018 19:24:54 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 7 on bug 10511= 3 from Jan Vesely
(In reply to Maciej S. Szmigiero from comment #6)
> There are really two issues at play here:
> 1) If the LLVM-generated code cannot be run properly then it should be=
 simply
> rejected by whatever is actually in charge of submitting it to the GPU=
 (I
> guess
> this would be Mesa?).
> This way an application will know it cannot use OpenCL for computation=
, at
> least
> not with this compute kernel.
>=20
> Instead, it currently looks like many of these test run but give incor=
rect
> results, which is obviously rather bad.

Do you have an example of this? clover should return OUT_OF_RESOURCES error
when the compute state creation fails (like in the presence of code
relocations).
It does not change the content of the buffer, so it will return whatever was
stored in the buffer on creation.

> 2) Some (previous) Mesa + LLVM versions generate=
 a command stream that
> crashes the GPU and, as far as I can remember, sometimes even lockup t=
he
> whole machine.
>=20
> It should not be possible to crash the GPU, regardless how incorrect a
> command stream that userspace sends to it is - because otherwise it is
> possible for
> an unprivileged user with GPU access to DoS the machine.

This is a separate issue. GPU hangs are generally addressed via gpu reset w=
hich
should be enabled for gfx8/9 GPUs in recent amdgpu.ko [0]

[0] https://pat=
chwork.freedesktop.org/patch/257994/


You are receiving this mail because:
  • You are the assignee for the bug.
= --15425690944.51Ee8.31720-- --===============0122903116== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============0122903116==--