From: Andrew Cooper <andrew.cooper3@citrix.com>
To: Jan Beulich <jbeulich@suse.com>
Cc: "Andrew Cooper" <andrew.cooper3@citrix.com>,
"Roger Pau Monné" <roger@xenproject.org>,
"Teddy Astie" <teddy.astie@vates.tech>,
Xen-devel <xen-devel@lists.xenproject.org>
Subject: Re: [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()
Date: Wed, 19 Aug 2026 12:13:47 +0100 [thread overview]
Message-ID: <65789d8b-d18f-4f1c-8ffc-e1247daa4031@citrix.com> (raw)
In-Reply-To: <5fa222d8-af04-4b1a-a457-ee5004ace484@suse.com>
On 18/08/2026 3:31 pm, Jan Beulich wrote:
> On 18.08.2026 15:07, Andrew Cooper wrote:
>> Panicing in the case of a timeout turns out to be about the worst possible
>> action Xen can take. It leaves all other APs waiting on the condition
>> variable, some in NMI context. As a result, they fail to be shot down and
>> dump state for kexec crash analysis.
> At the same time there likely isn't much to be learned from a kexec dump, as
> the source of the issue is in the CPU, not in Xen.
The single most valuable print message I've ever added to Xen is the one
which reports which CPUs didn't respond to NMIs. Up until now, it has
always highlighted hardware issues.
Despite my wish to remove this panic specifically, there is still
information to be gained from kexec in a similar scenario.
> Further, this code runs with the watchdog disabled. If there's truly no
> progress anymore, how would one know from the outside whether the system is
> dead altogether vs the control CPU still kicking around?
The scenario you describe can only occur if the BSP accepts the
microcode successfully, and one of the APs locks up properly.
If the BSP locks up, we never get as far as deciding to panic(), and the
system hangs already.
If we have a bad microcode, it is far more likely for the BSP to hang
than for the BSP to work one of the APs hang.
>> Microcode Loading on Granite Rapids takes about 4.5s of wallclock time, far in
>> excess of the of the arbitrary 1s Xen allows. This time is spent in the WRMSR
>> to load the blob, and there's nothing the system can do but to sit and wait.
>> Despite the delay, the system as a whole does survive.
> For this I wonder whether the log message issues after 1 sec is adequate. If
> we know things can take this long, wouldn't we better issue the log message
> no earlier than, say, 5s * nr_sockets?
>
I have some work there not submitted yet, which would periodically print
a message. That at least gives you slight signs of life.
But nothing involving nr_sockets. It's far more often wrong than it is
right, and that's not how our algorithm scales.
GNR has gone from milliseconds to 4.5s. Putting the limit at 5s is just
going to need another change in a year or two.
>> Microcode loading occures through admin operation only, so get rid of the
>> timeout completely. It does nothing but make a bad sitaution worse.
> Worse when taking one perspective, yes, yet a silent hang also is worse than
> a panic() telling you what was wrong.
>
> This is a tough one, with - likely - no really good options. And establishing
> what's "least bad" may also be difficult.
As indicated, there's only one narrow and unlikely case where the
panic() is a true positive.
The likely true-hang case don't reach the panic(), and that only leaves
"timeout too short" which has proved to be the case on GNR.
~Andrew
next prev parent reply other threads:[~2026-08-19 11:14 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-18 13:07 [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic() Andrew Cooper
2026-08-18 14:31 ` Jan Beulich
2026-08-19 11:13 ` Andrew Cooper [this message]
2026-08-19 11:56 ` Jan Beulich
2026-08-19 14:04 ` Andrew Cooper
2026-08-19 14:53 ` Jan Beulich
2026-08-19 21:06 ` Andrew Cooper
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=65789d8b-d18f-4f1c-8ffc-e1247daa4031@citrix.com \
--to=andrew.cooper3@citrix.com \
--cc=jbeulich@suse.com \
--cc=roger@xenproject.org \
--cc=teddy.astie@vates.tech \
--cc=xen-devel@lists.xenproject.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.