From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 93264] Tonga VM Faults since llvm ScheduleDAGInstrs: Rework schedule graph builder. Date: Tue, 08 Dec 2015 22:12:14 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0812695595==" Return-path: Received: from culpepper.freedesktop.org (unknown [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id B68016E3E8 for ; Tue, 8 Dec 2015 14:12:14 -0800 (PST) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0812695595== Content-Type: multipart/alternative; boundary="1449612734.6B8601.1699"; charset="UTF-8" --1449612734.6B8601.1699 Date: Tue, 8 Dec 2015 22:12:14 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable https://bugs.freedesktop.org/show_bug.cgi?id=3D93264 --- Comment #3 from Nicolai H=C3=A4hnle --- Time to document some information I've gathered. I can now confirm that Mesa has nothing to do with it. Something must have = gone wrong with my builds initially, sorry for having caused confusion. I captured an apitrace for better reproducibility[0] and ran it with shader dumps enabled and flushing after each draw call in the "interesting" region. I am going to attach the R600_DEBUG=3Dcheck_vm dump which I've cross-refere= nced with R600_DEBUG=3Dvm to obtain the shaders that were active during the draw= call (file names with prefix llvm-c0a189c.mesa-caf12bebd). I then matched the shaders to those dumped by a run with a good version of LLVM (commit just before the bad one, file names with prefix llvm-26ddca1.mesa-caf12bebd). Clearly, the LLVM changes caused some significant re-ordering of the instruction schedule, and that somehow, surprisingly, seems to be responsib= le for the VM faults. Another aspect to note is that the shaders are compiled before draw call 174000, while the VM faults happen shortly after draw call 178000. This see= ms to suggest that the shaders alone only cause VM faults in conjunction with = some other state. However, the VM faults have always happened in exactly the same point so far, so it does appear to be deterministic. [0] The demo always causes VM faults, so I'm not going to upload the giant trace; however, the timing varies between runs, so it's cleaner to reproduce using a trace. --=20 You are receiving this mail because: You are the assignee for the bug. --1449612734.6B8601.1699 Date: Tue, 8 Dec 2015 22:12:14 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable

Comment= # 3 on bug 93264<= /a> from Nicolai H=C3=A4hnle
Time to document some information I've gathered.

I can now confirm that Mesa has nothing to do with it. Something must have =
gone
wrong with my builds initially, sorry for having caused confusion.

I captured an apitrace for better reproducibility[0] and ran it with shader
dumps enabled and flushing after each draw call in the "interesting&qu=
ot; region.

I am going to attach the R600_DEBUG=3Dcheck_vm dump which I've cross-refere=
nced
with R600_DEBUG=3Dvm to obtain the shaders that were active during the draw=
 call
(file names with prefix llvm-c0a189c.mesa-caf12bebd). I then matched the
shaders to those dumped by a run with a good version of LLVM (commit just
before the bad one, file names with prefix llvm-26ddca1.mesa-caf12bebd).

Clearly, the LLVM changes caused some significant re-ordering of the
instruction schedule, and that somehow, surprisingly, seems to be responsib=
le
for the VM faults.

Another aspect to note is that the shaders are compiled before draw call
174000, while the VM faults happen shortly after draw call 178000. This see=
ms
to suggest that the shaders alone only cause VM faults in conjunction with =
some
other state. However, the VM faults have always happened in exactly the same
point so far, so it does appear to be deterministic.

[0] The demo always causes VM faults, so I'm not going to upload the giant
trace; however, the timing varies between runs, so it's cleaner to reproduce
using a trace.


You are receiving this mail because: =20=20=20=20=20=20
  • You are the assignee for the bug.
--1449612734.6B8601.1699-- --===============0812695595== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHA6Ly9saXN0 cy5mcmVlZGVza3RvcC5vcmcvbWFpbG1hbi9saXN0aW5mby9kcmktZGV2ZWwK --===============0812695595==--