From mboxrd@z Thu Jan 1 00:00:00 1970 From: bugzilla-daemon-CC+yJ3UmIYqDUpFQwHEjaQ@public.gmane.org Subject: [Bug 100567] Nouveau system freeze fifo: SCHED_ERROR 0a [CTXSW_TIMEOUT] Date: Mon, 06 May 2019 16:51:39 +0000 Message-ID: References: Mime-Version: 1.0 Content-Type: multipart/mixed; boundary="===============0339306775==" Return-path: In-Reply-To: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: nouveau-bounces-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org Sender: "Nouveau" To: nouveau-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org List-Id: nouveau.vger.kernel.org --===============0339306775== Content-Type: multipart/alternative; boundary="15571614992.ecfCE4cFa.27637" Content-Transfer-Encoding: 7bit --15571614992.ecfCE4cFa.27637 Date: Mon, 6 May 2019 16:51:39 +0000 MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated https://bugs.freedesktop.org/show_bug.cgi?id=3D100567 --- Comment #29 from Marc Burkhardt --- (In reply to Roy from comment #28) > Somewhat surprised that this particular report hasn't received any attent= ion > from a core dev. Sadly, I'm afraid my response will not be hugely satisfy= ing > eiter. >=20 > The message "fifo: SCHED_ERROR 0a [CTXSW_TIMEOUT]" is not the result of *a > bug* and does not point in the direction of one. Rather it is a symptom o= f a > wide variety of potential problems with nouveau, (all of) which result in= a > hang of the rendering part of the GPU. On itself it doesn't give meaningf= ul > information to any developer as to what may be the culprit. It's like > measuring a fever, it could be the result of many conditions. >=20 > Developers would be helped if there was a reliably reproducible situation= in > which this event can be triggered. If there is a list of steps that can be > followed that would result in this message and a hang, that somehow doesn= 't > involve words like "random" or "wait for possibly a few hours", and that = can > be traced with tools such as APITrace, we could get a step further in > analysing what goes wrong. However, it seems unlikely this is the case, > especially since we are also aware of multithreading-related issues that > make isolating such problems extremely difficult. Unfortunately, post-mor= tem > syslogs and dmesgs are unlikely to add any useful information to this or > similar bug reports. Hi Roy, first of all thank you very much for your time taken to answer this thread. Also thank you for the insights you gave. As there are "tons" of threads out there, from Ubuntu to Fedora bug tracker= s, as well as several threads here on freedesktop and the Linux bug tracker, t= hat are open created even 2 years ago, I guess it's time to get hands on to get= rid of this bug. You cannot imagine how annoying it is if you work on a machine that could easily just completely "stall" when you open the wrong menu item= or scroll a Twitter page just at the wrong time. Anyways, I'm still offering $100 for fixing this, so me and a good bunch of people around get this fixed upstream, so that it won't show up regardless = of distro or whatever kernel they run. This definitely needs a good amount of backporting though, as it bugs us for a long time. Maybe the one or other is willing two throw another $5 at it, so this would actually not benefit the = user only, but also the developer. If someone is familiar with a crowd-funding or something -> I am not. I'm willing to participate in testing patches or whatever I can do as a non-graphics-driver developer. I'm surely familiar with the processes testi= ng patches, however. I don't know what to suggest as a start to get this thing done. I have nouv= eau "debug" logs on my machine for a long time, but actually never saw anything relevant in it besides the "message of death". Just a blink of an eye after this happens the machine is completely unusable - no chance to interact any= how besides SysRq+REISSSUB. I would very much like to make a progress here. Maybe we could start and gather the already opened bugs and the people who participated in these. Just to have a couple of people really "hit" by the problem, that are willing to help and, moreover, are able to gather info on their machines. I'm pretty happy to run my machines without the "BLOB" and would happily ke= ep it this way. But this is a showstopper. A showstopper that has not gotten t= he right attention for a long time now. Thanks in advance - no matter what we will achieve. Marc --=20 You are receiving this mail because: You are the assignee for the bug.= --15571614992.ecfCE4cFa.27637 Date: Mon, 6 May 2019 16:51:39 +0000 MIME-Version: 1.0 Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-Bugzilla-URL: http://bugs.freedesktop.org/ Auto-Submitted: auto-generated

Comme= nt # 29 on bug 10056= 7 from Marc Burkhardt
(In reply to Roy from comment #28)
> Somewhat surprised that this particular report h=
asn't received any attention
> from a core dev. Sadly, I'm afraid my response will not be hugely sati=
sfying
> eiter.
>=20
> The message "fifo: SCHED_ERROR 0a [CTXSW_TIMEOUT]" is not th=
e result of *a
> bug* and does not point in the direction of one. Rather it is a sympto=
m of a
> wide variety of potential problems with nouveau, (all of) which result=
 in a
> hang of the rendering part of the GPU. On itself it doesn't give meani=
ngful
> information to any developer as to what may be the culprit. It's like
> measuring a fever, it could be the result of many conditions.
>=20
> Developers would be helped if there was a reliably reproducible situat=
ion in
> which this event can be triggered. If there is a list of steps that ca=
n be
> followed that would result in this message and a hang, that somehow do=
esn't
> involve words like "random" or "wait for possibly a few=
 hours", and that can
> be traced with tools such as APITrace, we could get a step further in
> analysing what goes wrong. However, it seems unlikely this is the case,
> especially since we are also aware of multithreading-related issues th=
at
> make isolating such problems extremely difficult. Unfortunately, post-=
mortem
> syslogs and dmesgs are unlikely to add any useful information to this =
or
> similar bug reports.

Hi Roy,

first of all thank you very much for your time taken to answer this thread.
Also thank you for the insights you gave.

As there are "tons" of threads out there, from Ubuntu to Fedora b=
ug trackers,
as well as several threads here on freedesktop and the Linux bug tracker, t=
hat
are open created even 2 years ago, I guess it's time to get hands on to get=
 rid
of this bug. You cannot imagine how annoying it is if you work on a machine
that could easily just completely "stall" when you open the wrong=
 menu item or
scroll a Twitter page just at the wrong time.

Anyways, I'm still offering $100 for fixing this, so me and a good bunch of
people around get this fixed upstream, so that it won't show up regardless =
of
distro or whatever kernel they run. This definitely needs a good amount of
backporting though, as it bugs us for a long time. Maybe the one or other is
willing two throw another $5 at it, so this would actually not benefit the =
user
only, but also the developer. If someone is familiar with a crowd-funding or
something -> I am not.

I'm willing to participate in testing patches or whatever I can do as a
non-graphics-driver developer. I'm surely familiar with the processes testi=
ng
patches, however.

I don't know what to suggest as a start to get this thing done. I have nouv=
eau
"debug" logs on my machine for a long time, but actually never sa=
w anything
relevant in it besides the "message of death". Just a blink of an=
 eye after
this happens the machine is completely unusable - no chance to interact any=
how
besides SysRq+REISSSUB.

I would very much like to make a progress here.

Maybe we could start and gather the already opened bugs and the people who
participated in these. Just to have a couple of people really "hit&quo=
t; by the
problem, that are willing to help and, moreover, are able to gather info on
their machines.

I'm pretty happy to run my machines without the "BLOB" and would =
happily keep
it this way. But this is a showstopper. A showstopper that has not gotten t=
he
right attention for a long time now.

Thanks in advance - no matter what we will achieve.

Marc


You are receiving this mail because:
  • You are the assignee for the bug.
= --15571614992.ecfCE4cFa.27637-- --===============0339306775== Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: base64 Content-Disposition: inline X19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX19fX18KTm91dmVhdSBt YWlsaW5nIGxpc3QKTm91dmVhdUBsaXN0cy5mcmVlZGVza3RvcC5vcmcKaHR0cHM6Ly9saXN0cy5m cmVlZGVza3RvcC5vcmcvbWFpbG1hbi9saXN0aW5mby9ub3V2ZWF1 --===============0339306775==--