The Linux Kernel Mailing List
 help / color / mirror / Atom feed
From: Mario Limonciello <mario.limonciello@amd.com>
To: Mathieu Fluhr <mathieu.fluhr@gmail.com>,
	Lovekesh Solanki <lovekeshsolanki00@gmail.com>
Cc: Michal Pecio <michal.pecio@gmail.com>,
	Thorsten Leemhuis <regressions@leemhuis.info>,
	Mathias Nyman <mathias.nyman@linux.intel.com>,
	linux-usb@vger.kernel.org, regressions@lists.linux.dev,
	stable@vger.kernel.org, linux-kernel@vger.kernel.org,
	Forest <forestix@gaga.casa>,
	Slavik Dev <developer.slavik@gmail.com>
Subject: Re: [REGRESSION] 6.12.36+: usb: hub: post-resume delayed work triggers > uncorrected MCE / data fabric sync flood on Threadripper 7970X > (bisected to aec11e5f9c45)
Date: Mon, 24 Aug 2026 17:09:02 -0500	[thread overview]
Message-ID: <226b2f34-fa25-4ca6-8afb-0b3dc942715c@amd.com> (raw)
In-Reply-To: <CAPyJwA8m_Xd2oyV1i9k+-6f7A-j0FiC97yXL=Ci8ZAErs5W4vA@mail.gmail.com>



On 8/24/26 14:37, Mathieu Fluhr wrote:
>> The issue is obviously a severe HW malfunction (you mentioned MCEs, the
>> Ryzen CPUs simply totally locked up), triggered by poking certain xHCI
>> controllers on the I/O die of these CPUs in some wrong way.
> 
> Yes. As mentioned, I first thought that the emulator itself triggered that by
> doing something that the CPU did not like.To be honest, I barely play with old
> Android versions anymore, but seeing that I could reproduce it even with
> Android 13 or 14 made me suspicious.
> 
> I _guess_ Google implemented a workaround inside adb for version Android
> 15 since using this version, it remains stable for more than 2 hours.
> 
> But, in the end, the situation is that from a simple user account having access
> to some usb plugged in devices (I usually add my user account to the plugdev
> group and use some known udev rules to access my Android tests devices),
> you have a way to crash the complete system.
> 
>> Opinions seem to vary on whether CPU load must be present or absent.
> 
> On my side (and I am here only speaking about my TR. I don't know about
> other Ryzen CPUs), I can't reproduce it under load, and one condition to
> reproduce it is my CPU going in C2 state.
> -> I did a 2:30 hour test using several youtube videos playing at the same
> time on my desktop, also stressing the emulator with some 3D Mark runs
> (as mentioned, I first suspected the nvidia driver to be faulty). As long as my
> computer was busy everything went fine. But then I let it stand still for a few
> minutes, and it just crashed.
> 
> If you need me to do some further tests or experiments, let me know. I will
> be more than happy to play the guinea pig here.
> 
The behavior that is described here sounds like a platform firmware bug 
to me.  Are you on the latest BIOS available from your OEM?

Can you please confirm:

1. Your CPU model number/codename
2. Your OEM (from /sys/class/dmi/id)
3. OEM BIOS version (from /sys/class/dmi/id)
4. AGESA version (see commit bc91133e260c8113c1119073c03b93c12aa41738 if 
the OEM didn't tear it out.  Otherwise look in BIOS menus)?


  reply	other threads:[~2026-08-24 22:09 UTC|newest]

Thread overview: 17+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-23  9:55 Mathieu Fluhr
2026-08-23 10:12 ` Mathieu Fluhr
2026-08-23 10:15   ` [REGRESSION] 6.12.36+: usb: hub: post-resume delayed work triggers > uncorrected MCE / data fabric sync flood on Threadripper 7970X > (bisected to aec11e5f9c45) Mathieu Fluhr
     [not found]     ` <07435e6b-ee30-4c85-8c8b-0ce3a4ead1d9@leemhuis.info>
2026-08-23 11:44       ` Mathieu Fluhr
2026-08-23 16:10         ` Michal Pecio
2026-08-23 15:17       ` Lovekesh Solanki
2026-08-23 15:40         ` Michal Pecio
2026-08-23 17:05           ` Lovekesh Solanki
2026-08-24 19:37             ` Mathieu Fluhr
2026-08-24 22:09               ` Mario Limonciello [this message]
2026-08-25  8:57                 ` Mathieu Fluhr
2026-08-26  3:14                   ` Mario Limonciello
2026-08-26  6:27                     ` Mathieu Fluhr
2026-08-26  6:59                     ` Michal Pecio
2026-08-25 10:15                 ` Lovekesh Solanki
2026-08-25 10:06               ` Lovekesh Solanki
2026-08-25 21:43                 ` Mathieu Fluhr

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=226b2f34-fa25-4ca6-8afb-0b3dc942715c@amd.com \
    --to=mario.limonciello@amd.com \
    --cc=developer.slavik@gmail.com \
    --cc=forestix@gaga.casa \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-usb@vger.kernel.org \
    --cc=lovekeshsolanki00@gmail.com \
    --cc=mathias.nyman@linux.intel.com \
    --cc=mathieu.fluhr@gmail.com \
    --cc=michal.pecio@gmail.com \
    --cc=regressions@leemhuis.info \
    --cc=regressions@lists.linux.dev \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox