From: Mario Limonciello <mario.limonciello@amd.com>
To: Fourhundred Thecat <400thecat@ik.me>,
platform-driver-x86@vger.kernel.org
Cc: linux-pm@vger.kernel.org, Shyam-sundar.S-k@amd.com,
hansg@kernel.org, ilpo.jarvinen@linux.intel.com,
rafael@kernel.org
Subject: Re: [BUG] s2idle: unrecoverable sleep on ThinkPad P16s Gen 4 AMD, (Strix Point) when more than 16 logical CPUs are online
Date: Wed, 23 Sep 2026 11:59:06 -0500 [thread overview]
Message-ID: <6d7773c8-0dc5-4e11-8bbb-088289dd508b@amd.com> (raw)
In-Reply-To: <4a161dd5-4dfe-81f4-e4d8-ba9d3280c25a@ik.me>
On 9/23/26 11:54, Fourhundred Thecat wrote:
> On 2026-09-23 18:41, Mario Limonciello wrote:
>>
>>
>> On 9/23/26 11:37, Fourhundred Thecat wrote:
>>> On 2026-09-23 18:20, Mario Limonciello wrote:
>>>>
>>>>
>>>> On 9/23/26 11:15, Mario Limonciello wrote:
>>>>>
>>>>>
>>>>> On 9/23/26 11:12, Fourhundred Thecat wrote:
>>>>>> On 2026-09-23 16:36, Mario Limonciello wrote:
>>>>>>>
>>>>>>>
>>>>>>> Some other thoughts that might be the root cause based on other
>>>>>>> historical issues.
>>>>>>>
>>>>>>> 1) Have you changed TPM policy or Pluton policy in BIOS setup?
>>>>>>> What did you change it from and to.
>>>>>>> 2) Have you enabled a storage security password? If you disable
>>>>>>> it does it help this issue?
>>>>>>> 3) Do you have WWAN in your device? If you disable it does it help?
>>>>>>> 4) Does booting with `amd_iommu=off` help?
>>>>>>
>>>>>> 4) amd_iommu=off
>>>>>>
>>>>>> Yes, that fixes it. With amd_iommu=off and all 24 CPUs online,
>>>>>> suspend/ resume works reliably.
>>>>>>
>>>>>> So this is not really a CPU count problem. nr_cpus=16 was only
>>>>>> masking an IOMMU interaction. But narrowing it further was
>>>>>> surprising: neither of the IOMMU's two functions is responsible on
>>>>>> its own. All of the following were run with all 24 CPUs online,
>>>>>> and I verified in each case that the parameter actually took effect:
>>>>>>
>>>>>> amd_iommu=off IOMMU off entirely WAKES
>>>>>> intremap=off IR off, DMA remapping on no wake
>>>>>> iommu=pt DMA passthrough, IR on no wake
>>>>>> amd_iommu_intr=legacy legacy GA mode, IR on, DMA on no wake
>>>>>>
>>>>>> Verification for each:
>>>>>>
>>>>>> intremap=off /proc/interrupts went from 58 IR- lines
>>>>>> to 0,
>>>>>> and irq 1 (i8042) is no longer IR-IO-APIC.
>>>>>> iommu still enabled, domain type
>>>>>> Translated.
>>>>>> iommu=pt "iommu: Default domain type: Passthrough
>>>>>> (set via
>>>>>> kernel command line)", all 40 PCI
>>>>>> devices in
>>>>>> identity domains, IR still on (58 IR-
>>>>>> lines).
>>>>>> amd_iommu_intr=legacy "AMD-Vi: Virtual APIC enabled" no longer
>>>>>> printed,
>>>>>> only "AMD-Vi: Interrupt remapping enabled".
>>>>>>
>>>>>> So only disabling the IOMMU outright helps. Turning off interrupt
>>>>>> remapping alone, bypassing DMA translation alone, or dropping out
>>>>>> of vAPIC/GA mode all still hang.
>>>>>>
>>>>>> The CPU dependency is still there on top of that. With the IOMMU
>>>>>> enabled, nr_cpus=16 works and 24 CPUs hangs. Offlining cpu16-23 by
>>>>>> hotplug after booting with all 24 does not help; the CPUs have to
>>>>>> never be brought up. cpu16-23 here are the second SMT thread of
>>>>>> the eight Zen5c cores (APIC ids 17,19..31).
>>>>>>
>>>>>> Both conditions appear to be required: the IOMMU enabled, and more
>>>>>> than 16 CPUs brought up at boot. Either one alone is fine.
>>>>>>
>>>>>>
>>>>>> 1) TPM / Pluton policy
>>>>>>
>>>>>> Current values:
>>>>>>
>>>>>> TpmSelection = DiscreteTPM2.0 (possible:
>>>>>> DiscreteTPM2.0;PlutonTPM2.0)
>>>>>> PlutonSecurityProcessor = Disable
>>>>>> SecurityChip = Enable
>>>>>>
>>>>>> The TPM that binds is a discrete STMicro part, tpm0 -> STM0925:00.
>>>>>>
>>>>>> I did change this. As I recall I disabled Microsoft Pluton, which
>>>>>> moves the TPM selection off the PlutonTPM2.0 default onto the
>>>>>> discrete part, so Enable -> Disable for Pluton and PlutonTPM2.0 ->
>>>>>> DiscreteTPM2.0 for the selection. I will confirm the exact
>>>>>> original values in setup. I have not yet tested whether restoring
>>>>>> the Pluton default changes the behaviour, since the IOMMU result
>>>>>> looked more promising.
>>>>>>
>>>>>>
>>>>>> 2) Storage security password
>>>>>>
>>>>>> None enrolled. HardDiskPasswordControl=Disable, and the HDD, NVMe,
>>>>>> Admin, System and Power-on authentication slots all report
>>>>>> is_enabled=0. BlockSIDAuthentication=Enable is the only non-
>>>>>> default setting in that area. Nothing to disable, so nothing to test.
>>>>>>
>>>>>>
>>>>>> 3) WWAN
>>>>>>
>>>>>> Yes: Quectel [1eac:1007] at 0000:c4:00.0, attached over MHI,
>>>>>> exposing wwan0. Its power/wakeup is enabled. Not yet tested with
>>>>>> WirelessWANAccess disabled in BIOS.
>>>>>>
>>>>>>
>>>>>> so please suggest which test I should do next, now that we have
>>>>>> more info
>>>>>
>>>>> Of your above the most likely cause is PlutonSecurityProcessor =
>>>>> Disable. Please try to re-enable that and then try with IOMMU
>>>>> enabled.
>>>>
>>>> BTW - what version of amd-s2idle didn't flag this? I am surprised,
>>>> we had a check for this that /should/ have failed prerequisites.
>>>
>>> Tested, and it does not help.
>>>
>>> PlutonSecurityProcessor Disable -> Enable
>>> TpmSelection DiscreteTPM2.0 -> PlutonTPM2.0
>>> SecurityChip Enable -> Active
>>>
>>> With those set, IOMMU enabled, no IOMMU boot parameters and all 24
>>> CPUs online, the machine still does not wake.
>>
>> That's interesting. We'll have to see what the report shows if it's
>> not the Pluton setting.
>>
>>>
>>> Worth noting what that test also covers: with Pluton selected the TPM
>>> presents through the CRB interface (MSFT0101:00, status=15), and this
>>> kernel has CONFIG_TCG_CRB=n. So during that suspend Linux had no TPM
>>> driver bound at all -- no /sys/class/tpm, no /dev/tpm0, and the
>>> discrete STM0925 was gone from the platform bus. Previously tpm_tis
>>> was bound to STM0925:00. So this rules out the tpm_tis driver as a
>>> factor as well as the Pluton policy.
>>>
>>> Current state of what is ruled out, all with the IOMMU enabled and 24
>>> CPUs online:
>>>
>>> amd_pmf initcall_blacklist=amd_pmf_driver_init no
>>> wake
>>> amdxdna (NPU)
>>> initcall_blacklist=amdxdna_pci_driver_init no wake
>>> TPM driver no driver bound at all (Pluton/CRB, no
>>> CONFIG_TCG_CRB) no wake
>>> Pluton policy Pluton enabled, PlutonTPM2.0 no wake
>>> interrupt remapping intremap=off (verified: 0 IR- lines) no wake
>>> DMA remapping iommu=pt (verified: Passthrough,
>>> identity) no wake
>>> vAPIC / GA mode amd_iommu_intr=legacy (verified) no wake
>>>
>>> The only two things that let it wake are amd_iommu=off with all 24
>>> CPUs, or the IOMMU enabled with nr_cpus=16. Offlining cpu16-23 by
>>> hotplug after booting with all 24 does not work; they have to never
>>> be brought up.
>>
>> No. nr_cpus=16 wasn't a pass. Don't treat it as such. You didn't
>> get to HW sleep. Let's please not conflate changing NR CPUs. Let's
>> figure out what's wrong with all CPUs enabled and IOMMU enabled, and
>> then peel it back if you need to turn off CPUs.
>>
>>>
>>> I am rebuilding now with CONFIG_DEBUG_FS, CONFIG_PM_DEBUG,
>>> CONFIG_DYNAMIC_DEBUG and CONFIG_AMD_MP2_STB so I can run amd-s2idle
>>> and send you the report from the working nr_cpus=16 configuration.
>>
>> So you didn't run it yet? I thought you said it failed.
>>
>
> > No. nr_cpus=16 wasn't a pass. Don't treat it as such. You didn't get
> > to HW sleep. Let's please not conflate changing NR CPUs.
>
> Agreed, I will drop it from the framing.
>
> That does leave a gap in my own data which I should close: I never
> measured whether amd_iommu=off reaches hardware sleep either. I only
> recorded total_hw_sleep=0 for the nr_cpus=16 case and did not check the
> counter after an amd_iommu=off resume. If that is also 0 then nothing on
> this machine has ever reached s0i3, and the wake failure is a second-
> order effect rather than the thing to chase. I will measure it.
>
> I would also like to confirm the counter is meaningful here before
> drawing conclusions from it. max_hw_sleep reads 18446744073709551615,
> which looks like an unpopulated value, so total_hw_sleep=0 may be a
> reporting gap rather than a real zero. With CONFIG_DEBUG_FS and
> CONFIG_AMD_MP2_STB in the new build I can read /sys/kernel/debug/
> amd_pmc/s0ix_stats directly instead of inferring it from suspend_stats.
I don't care about max_hw_sleep. It's a hardcoded value.
https://docs.kernel.org/admin-guide/abi-testing.html#abi-sys-power-suspend-stats-max-hw-sleep
>
> > So you didn't run it yet? I thought you said it failed.
>
> Correct, I have not run amd-s2idle yet.
That should have been your first debugging step.
next prev parent reply other threads:[~2026-09-23 16:59 UTC|newest]
Thread overview: 32+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-22 13:43 [BUG] s2idle: unrecoverable sleep on ThinkPad P16s Gen 4 AMD, (Strix Point) when more than 16 logical CPUs are online Fourhundred Thecat
2026-09-22 14:47 ` Mario Limonciello
2026-09-23 5:58 ` Fourhundred Thecat
2026-09-23 12:41 ` Mario Limonciello
2026-09-23 13:51 ` Fourhundred Thecat
2026-09-23 14:36 ` Mario Limonciello
2026-09-23 16:12 ` Fourhundred Thecat
2026-09-23 16:15 ` Mario Limonciello
2026-09-23 16:20 ` Mario Limonciello
2026-09-23 16:37 ` Fourhundred Thecat
2026-09-23 16:41 ` Mario Limonciello
2026-09-23 16:54 ` Fourhundred Thecat
2026-09-23 16:59 ` Mario Limonciello [this message]
2026-09-23 18:02 ` Fourhundred Thecat
2026-09-23 18:11 ` Mario Limonciello
2026-09-23 18:14 ` Fourhundred Thecat
2026-09-23 18:25 ` Mario Limonciello
2026-09-23 18:39 ` Fourhundred Thecat
2026-09-23 18:57 ` Mario Limonciello
2026-09-23 19:13 ` Fourhundred Thecat
2026-09-23 19:20 ` Mario Limonciello
2026-09-23 20:26 ` Fourhundred Thecat
2026-09-23 20:48 ` Fourhundred Thecat
2026-09-23 21:10 ` Mario Limonciello
2026-09-24 5:36 ` Fourhundred Thecat
2026-09-25 7:05 ` Fourhundred Thecat
2026-09-25 13:30 ` Mario Limonciello
2026-09-26 4:13 ` Fourhundred Thecat
2026-09-26 18:29 ` Mario Limonciello
2026-09-29 5:43 ` Fourhundred Thecat
2026-09-29 5:57 ` Fourhundred Thecat
2026-09-29 13:37 ` Mario Limonciello
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=6d7773c8-0dc5-4e11-8bbb-088289dd508b@amd.com \
--to=mario.limonciello@amd.com \
--cc=400thecat@ik.me \
--cc=Shyam-sundar.S-k@amd.com \
--cc=hansg@kernel.org \
--cc=ilpo.jarvinen@linux.intel.com \
--cc=linux-pm@vger.kernel.org \
--cc=platform-driver-x86@vger.kernel.org \
--cc=rafael@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox