All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()
@ 2026-08-18 13:07 Andrew Cooper
  2026-08-18 14:31 ` Jan Beulich
  0 siblings, 1 reply; 7+ messages in thread
From: Andrew Cooper @ 2026-08-18 13:07 UTC (permalink / raw)
  To: Xen-devel; +Cc: Andrew Cooper, Jan Beulich, Roger Pau Monné, Teddy Astie

Panicing in the case of a timeout turns out to be about the worst possible
action Xen can take.  It leaves all other APs waiting on the condition
variable, some in NMI context.  As a result, they fail to be shot down and
dump state for kexec crash analysis.

Microcode Loading on Granite Rapids takes about 4.5s of wallclock time, far in
excess of the of the arbitrary 1s Xen allows.  This time is spent in the WRMSR
to load the blob, and there's nothing the system can do but to sit and wait.
Despite the delay, the system as a whole does survive.

Microcode loading occures through admin operation only, so get rid of the
timeout completely.  It does nothing but make a bad sitaution worse.

Signed-off-by: Andrew Cooper <andrew.cooper3@citrix.com>
---
CC: Jan Beulich <jbeulich@suse.com>
CC: Roger Pau Monné <roger@xenproject.org>
CC: Teddy Astie <teddy.astie@vates.tech>
---
 xen/arch/x86/cpu/microcode/core.c | 18 +-----------------
 1 file changed, 1 insertion(+), 17 deletions(-)

diff --git a/xen/arch/x86/cpu/microcode/core.c b/xen/arch/x86/cpu/microcode/core.c
index 9b8d1e09cb98..12edd52fee87 100644
--- a/xen/arch/x86/cpu/microcode/core.c
+++ b/xen/arch/x86/cpu/microcode/core.c
@@ -52,12 +52,6 @@
  */
 #define MICROCODE_CALLIN_TIMEOUT_US 30000
 
-/*
- * Timeout for each thread to complete update is set to 1s. It is a
- * conservative choice considering all possible interference.
- */
-#define MICROCODE_UPDATE_TIMEOUT_US 1000000
-
 static bool __initdata __maybe_unused ucode_mod_forced;
 static unsigned int nr_cores;
 
@@ -422,17 +416,7 @@ static int control_thread_fn(const struct microcode_patch *patch,
     /* Wait for primary threads finishing update */
     while ( (done = atomic_read(&cpu_out)) != nr_cores )
     {
-        /*
-         * During each timeout interval, at least a CPU is expected to
-         * finish its update. Otherwise, something goes wrong.
-         *
-         * Note that RDTSC (in wait_for_condition()) is safe for threads to
-         * execute while waiting for completion of loading an update.
-         */
-        if ( wait_for_condition(wait_cpu_callout, (done + 1),
-                                MICROCODE_UPDATE_TIMEOUT_US) )
-            panic("Timeout when finished updating microcode (finished %u/%u)\n",
-                  done, nr_cores);
+        cpu_relax();
 
         /* Print warning message once if long time is spent here */
         if ( tick && rdtsc_ordered() - tick >= cpu_khz * 1000 )
-- 
2.39.5



^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()
  2026-08-18 13:07 [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic() Andrew Cooper
@ 2026-08-18 14:31 ` Jan Beulich
  2026-08-19 11:13   ` Andrew Cooper
  0 siblings, 1 reply; 7+ messages in thread
From: Jan Beulich @ 2026-08-18 14:31 UTC (permalink / raw)
  To: Andrew Cooper; +Cc: Roger Pau Monné, Teddy Astie, Xen-devel

On 18.08.2026 15:07, Andrew Cooper wrote:
> Panicing in the case of a timeout turns out to be about the worst possible
> action Xen can take.  It leaves all other APs waiting on the condition
> variable, some in NMI context.  As a result, they fail to be shot down and
> dump state for kexec crash analysis.

At the same time there likely isn't much to be learned from a kexec dump, as
the source of the issue is in the CPU, not in Xen.

Further, this code runs with the watchdog disabled. If there's truly no
progress anymore, how would one know from the outside whether the system is
dead altogether vs the control CPU still kicking around?

> Microcode Loading on Granite Rapids takes about 4.5s of wallclock time, far in
> excess of the of the arbitrary 1s Xen allows.  This time is spent in the WRMSR
> to load the blob, and there's nothing the system can do but to sit and wait.
> Despite the delay, the system as a whole does survive.

For this I wonder whether the log message issues after 1 sec is adequate. If
we know things can take this long, wouldn't we better issue the log message
no earlier than, say, 5s * nr_sockets?

> Microcode loading occures through admin operation only, so get rid of the
> timeout completely.  It does nothing but make a bad sitaution worse.

Worse when taking one perspective, yes, yet a silent hang also is worse than
a panic() telling you what was wrong.

This is a tough one, with - likely - no really good options. And establishing
what's "least bad" may also be difficult.

Jan


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()
  2026-08-18 14:31 ` Jan Beulich
@ 2026-08-19 11:13   ` Andrew Cooper
  2026-08-19 11:56     ` Jan Beulich
  0 siblings, 1 reply; 7+ messages in thread
From: Andrew Cooper @ 2026-08-19 11:13 UTC (permalink / raw)
  To: Jan Beulich; +Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Xen-devel

On 18/08/2026 3:31 pm, Jan Beulich wrote:
> On 18.08.2026 15:07, Andrew Cooper wrote:
>> Panicing in the case of a timeout turns out to be about the worst possible
>> action Xen can take.  It leaves all other APs waiting on the condition
>> variable, some in NMI context.  As a result, they fail to be shot down and
>> dump state for kexec crash analysis.
> At the same time there likely isn't much to be learned from a kexec dump, as
> the source of the issue is in the CPU, not in Xen.

The single most valuable print message I've ever added to Xen is the one
which reports which CPUs didn't respond to NMIs.  Up until now, it has
always highlighted hardware issues.

Despite my wish to remove this panic specifically, there is still
information to be gained from kexec in a similar scenario.

> Further, this code runs with the watchdog disabled. If there's truly no
> progress anymore, how would one know from the outside whether the system is
> dead altogether vs the control CPU still kicking around?

The scenario you describe can only occur if the BSP accepts the
microcode successfully, and one of the APs locks up properly.

If the BSP locks up, we never get as far as deciding to panic(), and the
system hangs already.

If we have a bad microcode, it is far more likely for the BSP to hang
than for the BSP to work one of the APs hang.

>> Microcode Loading on Granite Rapids takes about 4.5s of wallclock time, far in
>> excess of the of the arbitrary 1s Xen allows.  This time is spent in the WRMSR
>> to load the blob, and there's nothing the system can do but to sit and wait.
>> Despite the delay, the system as a whole does survive.
> For this I wonder whether the log message issues after 1 sec is adequate. If
> we know things can take this long, wouldn't we better issue the log message
> no earlier than, say, 5s * nr_sockets?
>

I have some work there not submitted yet, which would periodically print
a message.  That at least gives you slight signs of life.

But nothing involving nr_sockets.  It's far more often wrong than it is
right, and that's not how our algorithm scales.

GNR has gone from milliseconds to 4.5s.  Putting the limit at 5s is just
going to need another change in a year or two.

>> Microcode loading occures through admin operation only, so get rid of the
>> timeout completely.  It does nothing but make a bad sitaution worse.
> Worse when taking one perspective, yes, yet a silent hang also is worse than
> a panic() telling you what was wrong.
>
> This is a tough one, with - likely - no really good options. And establishing
> what's "least bad" may also be difficult.

As indicated, there's only one narrow and unlikely case where the
panic() is a true positive.

The likely true-hang case don't reach the panic(), and that only leaves
"timeout too short" which has proved to be the case on GNR.

~Andrew


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()
  2026-08-19 11:13   ` Andrew Cooper
@ 2026-08-19 11:56     ` Jan Beulich
  2026-08-19 14:04       ` Andrew Cooper
  0 siblings, 1 reply; 7+ messages in thread
From: Jan Beulich @ 2026-08-19 11:56 UTC (permalink / raw)
  To: Andrew Cooper; +Cc: Roger Pau Monné, Teddy Astie, Xen-devel

On 19.08.2026 13:13, Andrew Cooper wrote:
> On 18/08/2026 3:31 pm, Jan Beulich wrote:
>> On 18.08.2026 15:07, Andrew Cooper wrote:
>>> Panicing in the case of a timeout turns out to be about the worst possible
>>> action Xen can take.  It leaves all other APs waiting on the condition
>>> variable, some in NMI context.  As a result, they fail to be shot down and
>>> dump state for kexec crash analysis.
>> At the same time there likely isn't much to be learned from a kexec dump, as
>> the source of the issue is in the CPU, not in Xen.
> 
> The single most valuable print message I've ever added to Xen is the one
> which reports which CPUs didn't respond to NMIs.  Up until now, it has
> always highlighted hardware issues.
> 
> Despite my wish to remove this panic specifically, there is still
> information to be gained from kexec in a similar scenario.

That is you think of a system without console, where those log messages would
only be possible to fish out of the dump. Fair enough.

>> Further, this code runs with the watchdog disabled. If there's truly no
>> progress anymore, how would one know from the outside whether the system is
>> dead altogether vs the control CPU still kicking around?
> 
> The scenario you describe can only occur if the BSP accepts the
> microcode successfully, and one of the APs locks up properly.
> 
> If the BSP locks up, we never get as far as deciding to panic(), and the
> system hangs already.
> 
> If we have a bad microcode, it is far more likely for the BSP to hang
> than for the BSP to work one of the APs hang.

Hmm, probably. (I'm always having in mind the one old system I have where only
the primary cores on each socket get ucode updated by firmware, with secondary
cores needing us to deal with them.) Still somewhat hesitantly:
Acked-by: Jan Beulich <jbeulich@suse.com>

>>> Microcode Loading on Granite Rapids takes about 4.5s of wallclock time, far in
>>> excess of the of the arbitrary 1s Xen allows.  This time is spent in the WRMSR
>>> to load the blob, and there's nothing the system can do but to sit and wait.
>>> Despite the delay, the system as a whole does survive.
>> For this I wonder whether the log message issues after 1 sec is adequate. If
>> we know things can take this long, wouldn't we better issue the log message
>> no earlier than, say, 5s * nr_sockets?
> 
> I have some work there not submitted yet, which would periodically print
> a message.  That at least gives you slight signs of life.
> 
> But nothing involving nr_sockets.  It's far more often wrong than it is
> right, and that's not how our algorithm scales.
> 
> GNR has gone from milliseconds to 4.5s.  Putting the limit at 5s is just
> going to need another change in a year or two.

That's one possibility, yes. The other is that they realize that they have to
bring processing time back down.

Jan


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()
  2026-08-19 11:56     ` Jan Beulich
@ 2026-08-19 14:04       ` Andrew Cooper
  2026-08-19 14:53         ` Jan Beulich
  0 siblings, 1 reply; 7+ messages in thread
From: Andrew Cooper @ 2026-08-19 14:04 UTC (permalink / raw)
  To: Jan Beulich; +Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Xen-devel

On 19/08/2026 12:56 pm, Jan Beulich wrote:
> On 19.08.2026 13:13, Andrew Cooper wrote:
>> On 18/08/2026 3:31 pm, Jan Beulich wrote:
>>> On 18.08.2026 15:07, Andrew Cooper wrote:
>>>> Panicing in the case of a timeout turns out to be about the worst possible
>>>> action Xen can take.  It leaves all other APs waiting on the condition
>>>> variable, some in NMI context.  As a result, they fail to be shot down and
>>>> dump state for kexec crash analysis.
>>> At the same time there likely isn't much to be learned from a kexec dump, as
>>> the source of the issue is in the CPU, not in Xen.
>> The single most valuable print message I've ever added to Xen is the one
>> which reports which CPUs didn't respond to NMIs.  Up until now, it has
>> always highlighted hardware issues.
>>
>> Despite my wish to remove this panic specifically, there is still
>> information to be gained from kexec in a similar scenario.
> That is you think of a system without console, where those log messages would
> only be possible to fish out of the dump. Fair enough.

Yes.  This covers approximately all of XenServer deployments.

>
>>> Further, this code runs with the watchdog disabled. If there's truly no
>>> progress anymore, how would one know from the outside whether the system is
>>> dead altogether vs the control CPU still kicking around?
>> The scenario you describe can only occur if the BSP accepts the
>> microcode successfully, and one of the APs locks up properly.
>>
>> If the BSP locks up, we never get as far as deciding to panic(), and the
>> system hangs already.
>>
>> If we have a bad microcode, it is far more likely for the BSP to hang
>> than for the BSP to work one of the APs hang.
> Hmm, probably. (I'm always having in mind the one old system I have where only
> the primary cores on each socket get ucode updated by firmware, with secondary
> cores needing us to deal with them.) Still somewhat hesitantly:
> Acked-by: Jan Beulich <jbeulich@suse.com>

Thanks.

>
>>>> Microcode Loading on Granite Rapids takes about 4.5s of wallclock time, far in
>>>> excess of the of the arbitrary 1s Xen allows.  This time is spent in the WRMSR
>>>> to load the blob, and there's nothing the system can do but to sit and wait.
>>>> Despite the delay, the system as a whole does survive.
>>> For this I wonder whether the log message issues after 1 sec is adequate. If
>>> we know things can take this long, wouldn't we better issue the log message
>>> no earlier than, say, 5s * nr_sockets?
>> I have some work there not submitted yet, which would periodically print
>> a message.  That at least gives you slight signs of life.
>>
>> But nothing involving nr_sockets.  It's far more often wrong than it is
>> right, and that's not how our algorithm scales.
>>
>> GNR has gone from milliseconds to 4.5s.  Putting the limit at 5s is just
>> going to need another change in a year or two.
> That's one possibility, yes. The other is that they realize that they have to
> bring processing time back down.

I'm still trying to get details out of Intel, but I strongly suspect the
massive jump here is because of the introduction of the Ucode Staging
Buffer.

Even the architectural description is:

IA32_MCU_ENUMERATION (0x7b), bit 4: MCU_STAGING
"When set to 1, indicates that the microcode update staging capability
is supported by the processor. When supported, the use of the MCU
staging capability is recommended to reduce the latency of the
IA32_BIOS_UPDT_TRIG operation."

which all-but-says "high latency now expected".

When the staging buffer is primed properly, it gets back down to a
reasonable time, but I have no idea when we're going to be able to
support that, and in the meantime running xen-ucode needs to not crash.

One other idea I had:

If we rearrange the loop to have a periodic "taken $X seconds" (shows
liveness), and maybe at 40s decide to panic (by passing through the out:
label first so we release the APs, so they can crash more cleanly), then
that might be acceptable.

By 40s, the system isn't surviving even if the cores do all come back.

~Andrew


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()
  2026-08-19 14:04       ` Andrew Cooper
@ 2026-08-19 14:53         ` Jan Beulich
  2026-08-19 21:06           ` Andrew Cooper
  0 siblings, 1 reply; 7+ messages in thread
From: Jan Beulich @ 2026-08-19 14:53 UTC (permalink / raw)
  To: Andrew Cooper; +Cc: Roger Pau Monné, Teddy Astie, Xen-devel

On 19.08.2026 16:04, Andrew Cooper wrote:
> One other idea I had:
> 
> If we rearrange the loop to have a periodic "taken $X seconds" (shows
> liveness), and maybe at 40s decide to panic (by passing through the out:
> label first so we release the APs, so they can crash more cleanly), then
> that might be acceptable.
> 
> By 40s, the system isn't surviving even if the cores do all come back.

Maybe, yet I'm curious: How did you land at 40s as the "magic" boundary?
For a runtime load, even the 4.5s that you say GNR takes may already be
too much for the system (Dom0 and/or DomU-s) to properly survive as a
whole, depending on their requirements and/or configuration. Even our
time calibration rendezvous wants to occur once a second (albeit it may
not really need to run that often, plus iirc you have been saying that
we should get rid of it altogether).

Jan


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()
  2026-08-19 14:53         ` Jan Beulich
@ 2026-08-19 21:06           ` Andrew Cooper
  0 siblings, 0 replies; 7+ messages in thread
From: Andrew Cooper @ 2026-08-19 21:06 UTC (permalink / raw)
  To: Jan Beulich; +Cc: Andrew Cooper, Roger Pau Monné, Teddy Astie, Xen-devel

On 19/08/2026 3:53 pm, Jan Beulich wrote:
> On 19.08.2026 16:04, Andrew Cooper wrote:
>> One other idea I had:
>>
>> If we rearrange the loop to have a periodic "taken $X seconds" (shows
>> liveness), and maybe at 40s decide to panic (by passing through the out:
>> label first so we release the APs, so they can crash more cleanly), then
>> that might be acceptable.
>>
>> By 40s, the system isn't surviving even if the cores do all come back.
> Maybe, yet I'm curious: How did you land at 40s as the "magic" boundary?

Mostly arbitrary.  Linux has as least one timeouts around the 20s mark
(panic on soft lockup).  XenServer has some clustering timeouts around 30s.

> For a runtime load, even the 4.5s that you say GNR takes may already be
> too much for the system (Dom0 and/or DomU-s) to properly survive as a
> whole, depending on their requirements and/or configuration.

Its 9s all together, given our algorithm.  And I was a bit surprised,
but even a busy system doesn't seem to suffer any permanent issues as a
consequence.

> Even our
> time calibration rendezvous wants to occur once a second (albeit it may
> not really need to run that often, plus iirc you have been saying that
> we should get rid of it altogether).

Most hardware from the past 15y doesn't need it at all, because of
Invariant TSC.  Some >=8 socket Intel servers need a one-off calibration
during AP bringup, but that suffices AIUI.

There is one advantage to the tsc rendezvous.  If a core/thread/etc
really goes offline, it guarantees to get caught.  With the watchdog, 5
seconds after that, you'll notice.  System-wide TLB flushes have this
property too, but they're irregular and never the thing caught by the
watchdog.

~Andrew


^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-08-19 21:07 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-18 13:07 [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic() Andrew Cooper
2026-08-18 14:31 ` Jan Beulich
2026-08-19 11:13   ` Andrew Cooper
2026-08-19 11:56     ` Jan Beulich
2026-08-19 14:04       ` Andrew Cooper
2026-08-19 14:53         ` Jan Beulich
2026-08-19 21:06           ` Andrew Cooper

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.