All of lore.kernel.org
 help / color / mirror / Atom feed
From: Mario Limonciello <mario.limonciello@amd.com>
To: Mathieu Fluhr <mathieu.fluhr@gmail.com>,
	Lovekesh Solanki <lovekeshsolanki00@gmail.com>
Cc: Michal Pecio <michal.pecio@gmail.com>,
	Thorsten Leemhuis <regressions@leemhuis.info>,
	Mathias Nyman <mathias.nyman@linux.intel.com>,
	linux-usb@vger.kernel.org, regressions@lists.linux.dev,
	stable@vger.kernel.org, linux-kernel@vger.kernel.org,
	Forest <forestix@gaga.casa>,
	Slavik Dev <developer.slavik@gmail.com>
Subject: Re: [REGRESSION] 6.12.36+: usb: hub: post-resume delayed work triggers > uncorrected MCE / data fabric sync flood on Threadripper 7970X > (bisected to aec11e5f9c45)
Date: Mon, 24 Aug 2026 17:09:02 -0500	[thread overview]
Message-ID: <226b2f34-fa25-4ca6-8afb-0b3dc942715c@amd.com> (raw)
In-Reply-To: <CAPyJwA8m_Xd2oyV1i9k+-6f7A-j0FiC97yXL=Ci8ZAErs5W4vA@mail.gmail.com>



On 8/24/26 14:37, Mathieu Fluhr wrote:
>> The issue is obviously a severe HW malfunction (you mentioned MCEs, the
>> Ryzen CPUs simply totally locked up), triggered by poking certain xHCI
>> controllers on the I/O die of these CPUs in some wrong way.
> 
> Yes. As mentioned, I first thought that the emulator itself triggered that by
> doing something that the CPU did not like.To be honest, I barely play with old
> Android versions anymore, but seeing that I could reproduce it even with
> Android 13 or 14 made me suspicious.
> 
> I _guess_ Google implemented a workaround inside adb for version Android
> 15 since using this version, it remains stable for more than 2 hours.
> 
> But, in the end, the situation is that from a simple user account having access
> to some usb plugged in devices (I usually add my user account to the plugdev
> group and use some known udev rules to access my Android tests devices),
> you have a way to crash the complete system.
> 
>> Opinions seem to vary on whether CPU load must be present or absent.
> 
> On my side (and I am here only speaking about my TR. I don't know about
> other Ryzen CPUs), I can't reproduce it under load, and one condition to
> reproduce it is my CPU going in C2 state.
> -> I did a 2:30 hour test using several youtube videos playing at the same
> time on my desktop, also stressing the emulator with some 3D Mark runs
> (as mentioned, I first suspected the nvidia driver to be faulty). As long as my
> computer was busy everything went fine. But then I let it stand still for a few
> minutes, and it just crashed.
> 
> If you need me to do some further tests or experiments, let me know. I will
> be more than happy to play the guinea pig here.
> 
The behavior that is described here sounds like a platform firmware bug 
to me.  Are you on the latest BIOS available from your OEM?

Can you please confirm:

1. Your CPU model number/codename
2. Your OEM (from /sys/class/dmi/id)
3. OEM BIOS version (from /sys/class/dmi/id)
4. AGESA version (see commit bc91133e260c8113c1119073c03b93c12aa41738 if 
the OEM didn't tear it out.  Otherwise look in BIOS menus)?


  reply	other threads:[~2026-08-24 22:09 UTC|newest]

Thread overview: 24+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-23  9:55 Mathieu Fluhr
2026-08-23 10:12 ` Mathieu Fluhr
2026-08-23 10:15   ` [REGRESSION] 6.12.36+: usb: hub: post-resume delayed work triggers > uncorrected MCE / data fabric sync flood on Threadripper 7970X > (bisected to aec11e5f9c45) Mathieu Fluhr
2026-08-23 10:36     ` Thorsten Leemhuis
2026-08-23 11:44       ` Mathieu Fluhr
2026-08-23 16:10         ` Michal Pecio
2026-08-23 15:17       ` Lovekesh Solanki
2026-08-23 15:40         ` Michal Pecio
2026-08-23 17:05           ` Lovekesh Solanki
2026-08-24 19:37             ` Mathieu Fluhr
2026-08-24 22:09               ` Mario Limonciello [this message]
2026-08-25  8:57                 ` Mathieu Fluhr
2026-08-26  3:14                   ` Mario Limonciello
2026-08-26  6:27                     ` Mathieu Fluhr
2026-08-26 15:32                       ` Mario Limonciello
2026-08-26  6:59                     ` Michal Pecio
2026-08-25 10:15                 ` Lovekesh Solanki
2026-08-25 10:06               ` Lovekesh Solanki
2026-08-25 21:43                 ` Mathieu Fluhr
2026-08-26 12:35                   ` Lovekesh Solanki
2026-08-26 13:40                     ` Mathias Nyman
2026-08-26 16:03                       ` Mathieu Fluhr
2026-08-26 18:11                       ` Lovekesh Solanki
2026-08-27 19:34                         ` Mathias Nyman

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=226b2f34-fa25-4ca6-8afb-0b3dc942715c@amd.com \
    --to=mario.limonciello@amd.com \
    --cc=developer.slavik@gmail.com \
    --cc=forestix@gaga.casa \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-usb@vger.kernel.org \
    --cc=lovekeshsolanki00@gmail.com \
    --cc=mathias.nyman@linux.intel.com \
    --cc=mathieu.fluhr@gmail.com \
    --cc=michal.pecio@gmail.com \
    --cc=regressions@leemhuis.info \
    --cc=regressions@lists.linux.dev \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.