From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Tue, 17 Apr 2018 22:29:27 +0000 Message-ID: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0368113053==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id 5FEA76E1DF for ; Tue, 17 Apr 2018 22:29:28 +0000 (UTC) List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0368113053== Content-Type: multipart/alternative; boundary="15240041680.a55aB3b2.26691" Content-Transfer-Encoding: 7bit --15240041680.a55aB3b2.26691 Date: Tue, 17 Apr 2018 22:29:28 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 Bug ID: 106111 Summary: [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Product: DRI Version: unspecified Hardware: x86-64 (AMD64) OS: All Status: NEW Severity: normal Priority: medium Component: DRM/AMDgpu Assignee: dri-devel@lists.freedesktop.org Reporter: berillions@gmail.com CC: alexdeucher@gmail.com Created attachment 138887 --> https://bugs.freedesktop.org/attachment.cgi?id=3D138887&action=3Dedit xorg.conf Hi, My Setup : - AMD Ryzen 1600 - 16 Gb Memory RAM - Host (Debian Stable, kernel 4.16.2) : AMD Rx560 4Gb - Guest (Windows 10 / Archlinux Kernel 4.15.x-4.16.x) : AMD Rx580 - 8Gb Years ago there was an issue on Windows virtual machine with Qemu/VFIO and = AMD GPU. It was impossible to reboot or use a 2nde time the Guest because the G= PU was not reinitialized when the Host was shutdown. The only solution to re-u= se the VM was to reboot the Host OR use a Nvidia GPU. Actually, the issue is fixed on Windows VM + AMD GPU passed through (i don't know how), i can use more times my VM without reboot the Host.=20 But if i use my Linux VM with my Rx580, the issue still exist. The first la= unch works, i can use the Rx580 to play without problem. But if i shutdown/reboot the guest, the Rx580 is "blocked". I need to hard reboot because the system hangs after ~2-3 minutes. Thanks for your help, Maxime=20 (Sorry for my English, i'm French) --=20 You are receiving this mail because: You are the assignee for the bug.= --15240041680.a55aB3b2.26691 Date: Tue, 17 Apr 2018 22:29:28 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated
Bug ID 106111
Summary [GPU Passthrough]GPU (Polaris) not reinitialized with Linux V= M (Reset bug)
Product DRI
Version unspecified
Hardware x86-64 (AMD64)
OS All
Status NEW
Severity normal
Priority medium
Component DRM/AMDgpu
Assignee dri-devel@lists.freedesktop.org
Reporter berillions@gmail.com
CC alexdeucher@gmail.com

Created attachment 138887 [deta=
ils]
xorg.conf

Hi,

My Setup :
- AMD Ryzen 1600
- 16 Gb Memory RAM
- Host (Debian Stable, kernel 4.16.2) : AMD Rx560 4Gb
- Guest (Windows 10 / Archlinux Kernel 4.15.x-4.16.x) : AMD Rx580 - 8Gb

Years ago there was an issue on Windows virtual machine with Qemu/VFIO and =
AMD
GPU. It was impossible to reboot or use a 2nde time the Guest because the G=
PU
was not reinitialized when the Host was shutdown. The only solution to re-u=
se
the VM was to reboot the Host OR use a Nvidia GPU.

Actually, the issue is fixed on Windows VM + AMD GPU passed through (i don't
know how), i can use more times my VM without reboot the Host.=20

But if i use my Linux VM with my Rx580, the issue still exist. The first la=
unch
works, i can use the Rx580 to play without problem. But if i shutdown/reboot
the guest, the Rx580 is "blocked". I need to hard reboot because =
the system
hangs after ~2-3 minutes.

Thanks for your help,
Maxime=20

(Sorry for my English, i'm French)


You are receiving this mail because:
  • You are the assignee for the bug.
= --15240041680.a55aB3b2.26691-- --===============0368113053== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============0368113053==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Tue, 17 Apr 2018 22:30:13 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0049184112==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id D819389BD5 for ; Tue, 17 Apr 2018 22:30:12 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0049184112== Content-Type: multipart/alternative; boundary="15240042120.9eE72332.27378" Content-Transfer-Encoding: 7bit --15240042120.9eE72332.27378 Date: Tue, 17 Apr 2018 22:30:12 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 --- Comment #1 from Max --- Created attachment 138888 --> https://bugs.freedesktop.org/attachment.cgi?id=3D138888&action=3Dedit dmesg output after to launch the VM a second time --=20 You are receiving this mail because: You are the assignee for the bug.= --15240042120.9eE72332.27378 Date: Tue, 17 Apr 2018 22:30:12 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 1 on bug 10611= 1 from <= span class=3D"fn">Max
Created attachment 138888 [=
details]
dmesg output after to launch the VM a second time


You are receiving this mail because:
  • You are the assignee for the bug.
= --15240042120.9eE72332.27378-- --===============0049184112== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============0049184112==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Tue, 17 Apr 2018 22:55:23 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0114128205==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id C41D06E233 for ; Tue, 17 Apr 2018 22:55:23 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0114128205== Content-Type: multipart/alternative; boundary="15240057231.f6fb.1905" Content-Transfer-Encoding: 7bit --15240057231.f6fb.1905 Date: Tue, 17 Apr 2018 22:55:23 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 --- Comment #2 from Alex Williamson --- The IOMMU looks to be unhappy first: [ 40.201258] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270 [ 40.201271] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0 [ 40.201279] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370 [ 159.958402] AMD-Vi: Completion-Wait loop timed out [ 160.118777] AMD-Vi: Completion-Wait loop timed out [ 160.799864] AMD-Vi: Event logged [ [ 160.799868] IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8e8= 550] [ 160.799872] AMD-Vi: Event logged [ [ 160.799874] IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8e8= 570] [ 160.799876] AMD-Vi: Event logged [ [ 160.799878] IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8e8= 590] [ 161.801729] AMD-Vi: Event logged [ [ 161.801732] IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8e8= 5e0] [ 180.096365] AMD-Vi: Completion-Wait loop timed out [ 180.256758] AMD-Vi: Completion-Wait loop timed out [ 180.417182] AMD-Vi: Completion-Wait loop timed out [ 180.577636] AMD-Vi: Completion-Wait loop timed out Can you try a v4.17-rc1 kernel? Specifically, these two updates: 6bd06f5a486c vfio/type1: Adopt fast IOTLB flush interface when unmap IOVAs eb5ecd1a40e2 iommu/amd: Add support for fast IOTLB flushing Something about AMD GPUs get unhappy if the IOMMU sends out too many invalidations and the above two patches can reduce the number of those invalidations by up to a factor of 512. https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?= id=3D6bd06f5a486c06023a618a86e8153b91d26f75f4 https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?= id=3Deb5ecd1a40e2098f805fb63cb07817ac48826e40 --=20 You are receiving this mail because: You are the assignee for the bug.= --15240057231.f6fb.1905 Date: Tue, 17 Apr 2018 22:55:23 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 2 on bug 10611= 1 from Alex Williamson
The IOMMU looks to be unhappy first:

[   40.201258] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270
[   40.201271] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0
[   40.201279] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370
[  159.958402] AMD-Vi: Completion-Wait loop timed out
[  160.118777] AMD-Vi: Completion-Wait loop timed out
[  160.799864] AMD-Vi: Event logged [
[  160.799868] IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8e8=
550]
[  160.799872] AMD-Vi: Event logged [
[  160.799874] IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8e8=
570]
[  160.799876] AMD-Vi: Event logged [
[  160.799878] IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8e8=
590]
[  161.801729] AMD-Vi: Event logged [
[  161.801732] IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8e8=
5e0]
[  180.096365] AMD-Vi: Completion-Wait loop timed out
[  180.256758] AMD-Vi: Completion-Wait loop timed out
[  180.417182] AMD-Vi: Completion-Wait loop timed out
[  180.577636] AMD-Vi: Completion-Wait loop timed out

Can you try a v4.17-rc1 kernel?  Specifically, these two updates:

6bd06f5a486c vfio/type1: Adopt fast IOTLB flush interface when unmap IOVAs
eb5ecd1a40e2 iommu/amd: Add support for fast IOTLB flushing

Something about AMD GPUs get unhappy if the IOMMU sends out too many
invalidations and the above two patches can reduce the number of those
invalidations by up to a factor of 512.

https://git.kerne=
l.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=3D6bd06f5a486c=
06023a618a86e8153b91d26f75f4
https://git.kerne=
l.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=3Deb5ecd1a40e2=
098f805fb63cb07817ac48826e40


You are receiving this mail because:
  • You are the assignee for the bug.
= --15240057231.f6fb.1905-- --===============0114128205== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============0114128205==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Wed, 18 Apr 2018 06:11:21 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============1659959622==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id D27D26E1AE for ; Wed, 18 Apr 2018 06:11:21 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============1659959622== Content-Type: multipart/alternative; boundary="15240318811.C8564A94.16895" Content-Transfer-Encoding: 7bit --15240318811.C8564A94.16895 Date: Wed, 18 Apr 2018 06:11:21 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 --- Comment #3 from Max --- Created attachment 138893 --> https://bugs.freedesktop.org/attachment.cgi?id=3D138893&action=3Dedit dmesg after second launch + 4.17-rc1 Same problem with the Kernel 4.17-rc1. To be sure, i need to install this kernel only on the Host, no need to install it on the Linux Guest ? I use my own kernel 4.17 so maybe IOMMU/VFIO options are missing : odelpasso@debian-desktop:~/Bureau$ cat /boot/config-4.17.0-rc1 | grep VFIO CONFIG_VFIO_IOMMU_TYPE1=3Dm CONFIG_VFIO_VIRQFD=3Dm CONFIG_VFIO=3Dm # CONFIG_VFIO_NOIOMMU is not set CONFIG_VFIO_PCI=3Dm CONFIG_VFIO_PCI_VGA=3Dy CONFIG_VFIO_PCI_MMAP=3Dy CONFIG_VFIO_PCI_INTX=3Dy CONFIG_VFIO_PCI_IGD=3Dy # CONFIG_VFIO_MDEV is not set CONFIG_KVM_VFIO=3Dy odelpasso@debian-desktop:~/Bureau$ cat /boot/config-4.17.0-rc1 | grep IOMMU # CONFIG_GART_IOMMU is not set # CONFIG_CALGARY_IOMMU is not set CONFIG_IOMMU_HELPER=3Dy CONFIG_VFIO_IOMMU_TYPE1=3Dm # CONFIG_VFIO_NOIOMMU is not set CONFIG_IOMMU_API=3Dy CONFIG_IOMMU_SUPPORT=3Dy # Generic IOMMU Pagetable Support CONFIG_IOMMU_IOVA=3Dy CONFIG_AMD_IOMMU=3Dy CONFIG_AMD_IOMMU_V2=3Dy # CONFIG_INTEL_IOMMU is not set --=20 You are receiving this mail because: You are the assignee for the bug.= --15240318811.C8564A94.16895 Date: Wed, 18 Apr 2018 06:11:21 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 3 on bug 10611= 1 from <= span class=3D"fn">Max
Created att=
achment 138893 [details]
dmesg after second launch + 4.17-rc1

Same problem with the Kernel 4.17-rc1. To be sure, i need to install this
kernel only on the Host, no need to install it on the Linux Guest ?

I use my own kernel 4.17 so maybe IOMMU/VFIO options are missing :

odelpasso@debian-desktop:~/Bureau$ cat /boot/config-4.17.0-rc1 | grep V=
FIO
CONFIG_VFIO_IOMMU_TYPE1=3Dm
CONFIG_VFIO_VIRQFD=3Dm
CONFIG_VFIO=3Dm
# CONFIG_VFIO_NOIOMMU is not set
CONFIG_VFIO_PCI=3Dm
CONFIG_VFIO_PCI_VGA=3Dy
CONFIG_VFIO_PCI_MMAP=3Dy
CONFIG_VFIO_PCI_INTX=3Dy
CONFIG_VFIO_PCI_IGD=3Dy
# CONFIG_VFIO_MDEV is not set
CONFIG_KVM_VFIO=3Dy

odelpasso@debian-desktop:~/Bureau$ cat /boot/config-4.17.0-rc1 | grep I=
OMMU
# CONFIG_GART_IOMMU is not set
# CONFIG_CALGARY_IOMMU is not set
CONFIG_IOMMU_HELPER=3Dy
CONFIG_VFIO_IOMMU_TYPE1=3Dm
# CONFIG_VFIO_NOIOMMU is not set
CONFIG_IOMMU_API=3Dy
CONFIG_IOMMU_SUPPORT=3Dy
# Generic IOMMU Pagetable Support
CONFIG_IOMMU_IOVA=3Dy
CONFIG_AMD_IOMMU=3Dy
CONFIG_AMD_IOMMU_V2=3Dy
# CONFIG_INTEL_IOMMU is not set


You are receiving this mail because:
  • You are the assignee for the bug.
= --15240318811.C8564A94.16895-- --===============1659959622== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============1659959622==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Wed, 18 Apr 2018 16:13:52 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============2073211401==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id A52788986D for ; Wed, 18 Apr 2018 16:13:52 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============2073211401== Content-Type: multipart/alternative; boundary="15240680320.e62e.29602" Content-Transfer-Encoding: 7bit --15240680320.e62e.29602 Date: Wed, 18 Apr 2018 16:13:52 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 --- Comment #4 from Alex Williamson --- There is a difference, now we have: [ 84.997634] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270 [ 84.997645] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0 [ 84.997653] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370 [ 145.518307] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270 [ 145.518313] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0 [ 145.518318] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370 So prior to time 145.5 the VM was shutdown and started again and we could s= till read config space of the device. Previously we were already getting IOMMU faults before the second startup. But shortly after: [ 193.328586] AMD-Vi: Completion-Wait loop timed out [ 193.488711] AMD-Vi: Completion-Wait loop timed out [ 194.169913] iommu ivhd0: AMD-Vi: Event logged [ [ 194.169921] iommu ivhd0: IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8aaca0] [ 194.169924] iommu ivhd0: AMD-Vi: Event logged [ [ 194.169928] iommu ivhd0: IOTLB_INV_TIMEOUT device=3D0a:00.0 address=3D0x000000043e8aacc0] And the stuck in D3 state is evidence that the device is no longer accessib= le on the bus. So that only delayed the issue, some interaction between the I= OMMU and GPU is still failing. --=20 You are receiving this mail because: You are the assignee for the bug.= --15240680320.e62e.29602 Date: Wed, 18 Apr 2018 16:13:52 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 4 on bug 10611= 1 from Alex Williamson
There is a difference, now we have:

[   84.997634] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270
[   84.997645] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0
[   84.997653] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370
[  145.518307] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270
[  145.518313] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0
[  145.518318] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370

So prior to time 145.5 the VM was shutdown and started again and we could s=
till
read config space of the device.  Previously we were already getting IOMMU
faults before the second startup.  But shortly after:

[  193.328586] AMD-Vi: Completion-Wait loop timed out
[  193.488711] AMD-Vi: Completion-Wait loop timed out
[  194.169913] iommu ivhd0: AMD-Vi: Event logged [
[  194.169921] iommu ivhd0: IOTLB_INV_TIMEOUT device=3D0a:00.0
address=3D0x000000043e8aaca0]
[  194.169924] iommu ivhd0: AMD-Vi: Event logged [
[  194.169928] iommu ivhd0: IOTLB_INV_TIMEOUT device=3D0a:00.0
address=3D0x000000043e8aacc0]

And the stuck in D3 state is evidence that the device is no longer accessib=
le
on the bus.  So that only delayed the issue, some interaction between the I=
OMMU
and GPU is still failing.


You are receiving this mail because:
  • You are the assignee for the bug.
= --15240680320.e62e.29602-- --===============2073211401== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============2073211401==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Wed, 18 Apr 2018 16:33:36 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============1898757763==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id 1DEEE6E584 for ; Wed, 18 Apr 2018 16:33:36 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============1898757763== Content-Type: multipart/alternative; boundary="15240692151.aeB0E.7462" Content-Transfer-Encoding: 7bit --15240692151.aeB0E.7462 Date: Wed, 18 Apr 2018 16:33:35 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 --- Comment #5 from Max --- (In reply to Alex Williamson from comment #4) > There is a difference, now we have: >=20 > [ 84.997634] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270 > [ 84.997645] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0 > [ 84.997653] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370 > [ 145.518307] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270 > [ 145.518313] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0 > [ 145.518318] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370 >=20 > So prior to time 145.5 the VM was shutdown and started again and we could > still read config space of the device. Previously we were already getting > IOMMU faults before the second startup. But shortly after: >=20 > [ 193.328586] AMD-Vi: Completion-Wait loop timed out > [ 193.488711] AMD-Vi: Completion-Wait loop timed out > [ 194.169913] iommu ivhd0: AMD-Vi: Event logged [ > [ 194.169921] iommu ivhd0: IOTLB_INV_TIMEOUT device=3D0a:00.0 > address=3D0x000000043e8aaca0] > [ 194.169924] iommu ivhd0: AMD-Vi: Event logged [ > [ 194.169928] iommu ivhd0: IOTLB_INV_TIMEOUT device=3D0a:00.0 > address=3D0x000000043e8aacc0] >=20 > And the stuck in D3 state is evidence that the device is no longer > accessible on the bus. So that only delayed the issue, some interaction > between the IOMMU and GPU is still failing. Thanks for the explaination Alex. Something could be done ?=20 By AMD or VFIO mainteners ? --=20 You are receiving this mail because: You are the assignee for the bug.= --15240692151.aeB0E.7462 Date: Wed, 18 Apr 2018 16:33:35 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 5 on bug 10611= 1 from <= span class=3D"fn">Max
(In reply to Alex Williamson from comment #4)
> There is a difference, now we have:
>=20
> [   84.997634] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270
> [   84.997645] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0
> [   84.997653] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370
> [  145.518307] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x19@0x270
> [  145.518313] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1b@0x2d0
> [  145.518318] vfio_ecap_init: 0000:0a:00.0 hiding ecap 0x1e@0x370
>=20
> So prior to time 145.5 the VM was shutdown and started again and we co=
uld
> still read config space of the device.  Previously we were already get=
ting
> IOMMU faults before the second startup.  But shortly after:
>=20
> [  193.328586] AMD-Vi: Completion-Wait loop timed out
> [  193.488711] AMD-Vi: Completion-Wait loop timed out
> [  194.169913] iommu ivhd0: AMD-Vi: Event logged [
> [  194.169921] iommu ivhd0: IOTLB_INV_TIMEOUT device=3D0a:00.0
> address=3D0x000000043e8aaca0]
> [  194.169924] iommu ivhd0: AMD-Vi: Event logged [
> [  194.169928] iommu ivhd0: IOTLB_INV_TIMEOUT device=3D0a:00.0
> address=3D0x000000043e8aacc0]
>=20
> And the stuck in D3 state is evidence that the device is no longer
> accessible on the bus.  So that only delayed the issue, some interacti=
on
> between the IOMMU and GPU is still failing.

Thanks for the explaination Alex.
Something could be done ?=20
By AMD or VFIO mainteners ?


You are receiving this mail because:
  • You are the assignee for the bug.
= --15240692151.aeB0E.7462-- --===============1898757763== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============1898757763==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Sat, 18 Aug 2018 08:31:32 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============1143063492==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id 9E4CB6E037 for ; Sat, 18 Aug 2018 08:31:33 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============1143063492== Content-Type: multipart/alternative; boundary="15345810931.Ef81FE.23195" Content-Transfer-Encoding: 7bit --15345810931.Ef81FE.23195 Date: Sat, 18 Aug 2018 08:31:33 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 --- Comment #6 from Rados=C5=82aw Szkodzi=C5=84ski = --- This is still happening. It seems that these GPU need engine resets before = bus reset, similar to what was done for Fury and Polaris, but more extensive. Temporary workaround (yeah sure) is to eject the driver - rmmod in guest or eject in Windows. This resets the engines. Windows did the resets on shutdown until version 18.5.1 where they broke shutdown sequence again - read release notes on Radeon Pro Vega FE drivers where they actually slightly care. --=20 You are receiving this mail because: You are the assignee for the bug.= --15345810931.Ef81FE.23195 Date: Sat, 18 Aug 2018 08:31:33 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 6 on bug 10611= 1 from Rados=C5=82aw Szkodzi=C5=84s= ki
This is still happening. It seems that these GPU need engine r=
esets before bus
reset, similar to what was done for Fury and Polaris, but more extensive.

Temporary workaround (yeah sure) is to eject the driver - rmmod in guest or
eject in Windows. This resets the engines.

Windows did the resets on shutdown until version 18.5.1 where they broke
shutdown sequence again - read release notes on Radeon Pro Vega FE drivers
where they actually slightly care.


You are receiving this mail because:
  • You are the assignee for the bug.
= --15345810931.Ef81FE.23195-- --===============1143063492== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============1143063492==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Fri, 14 Sep 2018 10:30:30 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============1650329097==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [IPv6:2610:10:20:722:a800:ff:fe98:4b55]) by gabe.freedesktop.org (Postfix) with ESMTP id 8C8F76E865 for ; Fri, 14 Sep 2018 10:30:30 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============1650329097== Content-Type: multipart/alternative; boundary="15369210300.B1c2Dd5.23065" Content-Transfer-Encoding: 7bit --15369210300.B1c2Dd5.23065 Date: Fri, 14 Sep 2018 10:30:30 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 --- Comment #7 from Andrew Sheldon --- Another workaround that has worked for me with a Vega 56 is to suspend-to-r= am the host system before trying to start the guest again. --=20 You are receiving this mail because: You are the assignee for the bug.= --15369210300.B1c2Dd5.23065 Date: Fri, 14 Sep 2018 10:30:30 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Commen= t # 7 on bug 10611= 1 from Andrew Sheldon
Another workaround that has worked for me with a Vega 56 is to=
 suspend-to-ram
the host system before trying to start the guest again.


You are receiving this mail because:
  • You are the assignee for the bug.
= --15369210300.B1c2Dd5.23065-- --===============1650329097== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVsCg== --===============1650329097==-- From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon@freedesktop.org Subject: [Bug 106111] [GPU Passthrough]GPU (Polaris) not reinitialized with Linux VM (Reset bug) Date: Tue, 19 Nov 2019 08:35:24 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0693949196==" Return-path: Received: from culpepper.freedesktop.org (culpepper.freedesktop.org [131.252.210.165]) by gabe.freedesktop.org (Postfix) with ESMTP id D61EE6ED22 for ; Tue, 19 Nov 2019 08:35:24 +0000 (UTC) In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dri-devel-bounces@lists.freedesktop.org Sender: "dri-devel" To: dri-devel@lists.freedesktop.org List-Id: dri-devel@lists.freedesktop.org --===============0693949196== Content-Type: multipart/alternative; boundary="15741525245.EaAf6Cc8.18922" Content-Transfer-Encoding: 7bit --15741525245.EaAf6Cc8.18922 Date: Tue, 19 Nov 2019 08:35:24 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D106111 Martin Peres changed: What |Removed |Added ---------------------------------------------------------------------------- Resolution|--- |MOVED Status|NEW |RESOLVED --- Comment #8 from Martin Peres --- -- GitLab Migration Automatic Message -- This bug has been migrated to freedesktop.org's GitLab instance and has been closed from further activity. You can subscribe and participate further through the new bug through this = link to our GitLab instance: https://gitlab.freedesktop.org/drm/amd/issues/346. --=20 You are receiving this mail because: You are the assignee for the bug.= --15741525245.EaAf6Cc8.18922 Date: Tue, 19 Nov 2019 08:35:24 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated <= span class=3D"fn">Martin Peres changed bug 10611= 1
What Removed Added
Resolution --- MOVED
Status NEW RESOLVED

Commen= t # 8 on bug 10611= 1 from Martin Peres
-- GitLab Migration Automatic Message --

This bug has been migrated to freedesktop.org's GitLab instance and has been
closed from further activity.

You can subscribe and participate further through the new bug through this =
link
to our GitLab instance: https://gitlab.freedesktop.org/drm/amd/issues/346.


You are receiving this mail because:
  • You are the assignee for the bug.
= --15741525245.EaAf6Cc8.18922-- --===============0693949196== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KZHJpLWRldmVs IG1haWxpbmcgbGlzdApkcmktZGV2ZWxAbGlzdHMuZnJlZWRlc2t0b3Aub3JnCmh0dHBzOi8vbGlz dHMuZnJlZWRlc2t0b3Aub3JnL21haWxtYW4vbGlzdGluZm8vZHJpLWRldmVs --===============0693949196==--