Linux Power Management development
 help / color / mirror / Atom feed
From: Fourhundred Thecat <400thecat@ik.me>
To: Mario Limonciello <mario.limonciello@amd.com>,
	platform-driver-x86@vger.kernel.org
Cc: linux-pm@vger.kernel.org, Shyam-sundar.S-k@amd.com,
	hansg@kernel.org, ilpo.jarvinen@linux.intel.com,
	rafael@kernel.org
Subject: Re: [BUG] s2idle: unrecoverable sleep on ThinkPad P16s Gen 4 AMD, (Strix Point) when more than 16 logical CPUs are online
Date: Wed, 23 Sep 2026 22:48:33 +0200	[thread overview]
Message-ID: <afb1949a-8465-cfc8-bddb-bc27783af223@ik.me> (raw)
In-Reply-To: <c58f84c7-bfe6-58db-cad5-6859e96a7daf@ik.me>

On 2026-09-23 22:26, Fourhundred Thecat wrote:
> On 2026-09-23 21:20, Mario Limonciello wrote:
>>
>>
>> On 9/23/26 14:13, Fourhundred Thecat wrote:
>>> On 2026-09-23 20:57, Mario Limonciello wrote:
>>>>
>>>>
>>>> On 9/23/26 13:39, Fourhundred Thecat wrote:
>>>>> On 23/09/2026 20.25, Mario Limonciello wrote:
>>>>>
>>>>>>
>>>>>> Can you please share your kernel log from this run as well?  It seems
>>>>>> that your distro dmesg tool didn't pick it up in the tool run.
>>>>>>
>>>>>> And can I please see dmesg from a run with amd_iommu=on too.
>>>>>>
>>>>>>>    🚦 DMI data was not setup
>>>>>>
>>>>>> What is up with the missing data here?
>>>>>>
>>>>>> Does your BIOS offer anything to control D3 behavior for the storage?
>>>>>>
>>>>>> And this other point I mentioned: another useful data point will be
>>>>>> whether this can reproduce on 7.3-rc4 with IOMMU enabled to rule 
>>>>>> out a
>>>>>> backport issue.
>>>>>
>>>>>> Can you please share your kernel log from this run as well?
>>>>>
>>>>> Attached as dmesg-iommu-off.txt (full log, 1082 lines, 
>>>>> amd_iommu=off, 24
>>>>> CPUs, the cycle that succeeded). The suspend/resume portion:
>>>>>
>>>>>    PM: suspend entry (s2idle)
>>>>>    Filesystems sync: 0.010 seconds
>>>>>    Freezing user space processes
>>>>>    Freezing user space processes completed (elapsed 0.001 seconds)
>>>>>    OOM killer disabled.
>>>>>    Freezing remaining freezable tasks
>>>>>    Freezing remaining freezable tasks completed (elapsed 0.000 
>>>>> seconds)
>>>>>    PM: Triggering wakeup from IRQ 9
>>>>>    ACPI: PM: Rearming ACPI SCI for wakeup
>>>>>    amd_pmc: SMU idlemask s0i3: 0xffff9afd
>>>>>    PM: Triggering wakeup from IRQ 9
>>>>>    ACPI: PM: Rearming ACPI SCI for wakeup
>>>>>    PM: Triggering wakeup from IRQ 9
>>>>>    amd_pmc: SMU idlemask s0i3: 0xffff9abd
>>>>>    ACPI: PM: Rearming ACPI SCI for wakeup
>>>>>    amd_pmc: SMU idlemask s0i3: 0xffff9abd
>>>>>    PM: Triggering wakeup from IRQ 9
>>>>>    PM: Triggering wakeup from IRQ 7
>>>>>    ACPI: PM: Wakeup after ACPI Notify sync
>>>>>    OOM killer enabled.
>>>>>    Restarting tasks: Starting
>>>>>    Restarting tasks: Done
>>>>>    PM: suspend exit
>>>>>
>>>>> For reference, IRQ 9 is the ACPI SCI and IRQ 7 is pinctrl_amd, so the
>>>>> SCI fires and re-arms several times and the actual wake arrives 
>>>>> through
>>>>> the AMD GPIO controller.
>>>>>
>>>>>> And can I please see dmesg from a run with amd_iommu=on too.
>>>>>
>>>>> I cannot produce one. With the IOMMU enabled the machine never 
>>>>> resumes,
>>>>> so the ring buffer is lost to the forced power cycle. Streaming it out
>>>>> does not work either: dmesg -w and sshd are both frozen by the 
>>>>> freezer,
>>>>> so over ssh the log stops at
>>>>>
>>>>>    PM: suspend entry (s2idle)
>>>>>    Filesystems sync: 0.011 seconds
>>>>>
>>>>
>>>> I don't need the full run, I'm looking for how it sets up differently.
>>>>
>>>> I am specifically expecting the message from 
>>>> 51c33f333bbf7bdb6aa2a327e3a3e4bbb2591511 to come up and want to 
>>>> confirm that.
>>>>
>>>>> and nothing after that ever leaves the machine. There is no serial 
>>>>> port
>>>>> on this laptop, and since it hangs rather than panics, pstore captures
>>>>> nothing.
>>>>>
>>>>> What I can send instead is a full kernel log with the IOMMU enabled
>>>>> using /sys/power/pm_test=platform, which runs the whole suspend path
>>>>> including LPS0 _DSM entry and returns without entering the idle loop.
>>>>> That gives you every device callback and the platform prepare with the
>>>>> IOMMU active. I will send it in a follow-up unless you would rather 
>>>>> have
>>>>> something else.
>>>>>
>>>>>> Uh, the hardware does support a wakealarm.  You might have 
>>>>>> disabled it
>>>>> in your kernel.
>>>>>
>>>>> I checked, and the relevant options are all enabled:
>>>>>
>>>>>    CONFIG_RTC_CLASS=y
>>>>>    CONFIG_RTC_DRV_CMOS=y
>>>>>    CONFIG_RTC_INTF_SYSFS=y
>>>>>    CONFIG_RTC_INTF_DEV=y
>>>>>    CONFIG_HPET=y
>>>>>    CONFIG_HPET_TIMER=y
>>>>>    CONFIG_HPET_EMULATE_RTC=y
>>>>>
>>>>> What happens at boot is:
>>>>>
>>>>>    hpet0: at MMIO 0xfed00000, IRQs 2, 8, 0
>>>>>    hpet0: 3 comparators, 32-bit 14.318180 MHz counter
>>>>>    clocksource: hpet: mask: 0xffffffff max_cycles: 0xffffffff,
>>>>> max_idle_ns: 133484873504 ns
>>>>>    rtc_cmos PNP0B00:00: error -ENXIO: IRQ index 0 not found
>>>>>    rtc_cmos PNP0B00:00: RTC can wake from S4
>>>>>    rtc_cmos PNP0B00:00: registered as rtc0
>>>>>
>>>>> and /proc/driver/rtc reports HPET_emulated: no.
>>>>>
>>>>> As far as I can follow it, HPET is registered as a clocksource only 
>>>>> and
>>>>> legacy replacement is never enabled, so is_hpet_enabled()
>>>>> (is_hpet_capable() && hpet_legacy_int_enabled) is false. That makes
>>>>> use_acpi_alarm_quirks() return at its "if (!is_hpet_enabled()) 
>>>>> return;"
>>>>> check, so use_acpi_alarm stays false, and use_hpet_alarm() is false 
>>>>> too.
>>>>> ACPI does not give PNP0B00 an interrupt resource, so
>>>>> is_valid_irq(rtc_irq) fails and cmos_do_probe() takes the else branch
>>>>> that does clear_bit(RTC_FEATURE_ALARM, ...), which is why there is no
>>>>> wakealarm attribute.
>>>>
>>>> I think you're missing commit e9f850ba66cdf6b77fb4f005e46c4b605c4de434.
>>>>
>>>>>
>>>>>>   🚦 DMI data was not setup
>>>>>> What is up with the missing data here?
>>>>>
>>>>> CONFIG_DMIID is not set in my config, so /sys/class/dmi/id does not
>>>>> exist. DMI itself is scanned normally:
>>>>>
>>>>>    DMI: LENOVO 21RXS07D00/21RXS07D00, BIOS R2XET40W (1.20 ) 05/26/2026
>>>>>
>>>>> so dmi_check_system() quirks do apply. I will enable CONFIG_DMIID 
>>>>> in the
>>>>> next build so the tool stops reporting it.
>>>>>
>>>>>> Does your BIOS offer anything to control D3 behavior for the storage?
>>>>>
>>>>> No. I dumped all 96 attributes exposed by think-lmi and there is 
>>>>> nothing
>>>>> for storage power management or D3. The only storage related 
>>>>> entries are
>>>>> HardDiskPasswordControl and BlockSIDAuthentication, both access 
>>>>> control
>>>>> rather than power.
>>>>
>>>> I don't know for sure if Think LMI will export all BIOS options in 
>>>> the BIOS GUI.
>>>>
>>>>>
>>>>>> whether this can reproduce on 7.3-rc4 with IOMMU enabled to rule 
>>>>>> out a
>>>>> backport issue
>>>>>
>>>>> i will try to test 7.3-rc4  as you suggest
>>>>
>>>> OK.
>>>
>>>  > I am specifically expecting the message from 51c33f333bbf to come 
>>> up and want to confirm that.
>>>
>>> That commit is in my tree, but the message does not appear. Booted 
>>> with the IOMMU enabled, 24 CPUs, no IOMMU parameters, the only thing 
>>> matching is the unrelated generic ACPI one:
>>>
>>>    dmesg | grep -iE 'FW_BUG|Firmware Bug|matched UID|MSFT0201|acpihid'
>>>    AMD-Vi: ivrs, add hid:MSFT0201, uid:1, rdevid:0x60
>>>    ACPI: [Firmware Bug]: BIOS _OSI(Linux) query ignored
>>>    platform MSFT0201:00: Adding to iommu group 0
>>>
>>> No "No ACPI device matched UID, but N device(s) matched HID." The 
>>> UIDs agree on this machine:
>>>
>>>    IVRS:  hid:MSFT0201, uid:1, rdevid:0x60
>>>    ACPI:  MSFT0201:00, _UID = 1, path \_SB_.MHSP
>>>
>>> so get_acpihid_device_id() matches on the first pass and fw_bug is 
>>> never set. Same device path as in your commit, but this BIOS appears 
>>> to have consistent UIDs.
>>>
>>> The full AMD-Vi block for this boot:
>>>
>>>    ACPI: IVRS 0x000000006B1BF000 0001F6 (v02 LENOVO TP-R2X   00001200 
>>> PTEC 00000002)
>>>    AMD-Vi: ivrs, add hid:AMDI0020, uid:ID00, rdevid:0xa0
>>>    AMD-Vi: ivrs, add hid:AMDI0020, uid:ID01, rdevid:0xa0
>>>    AMD-Vi: ivrs, add hid:AMDI0020, uid:ID02, rdevid:0xa0
>>>    AMD-Vi: ivrs, add hid:AMDI0020, uid:ID03, rdevid:0x98
>>>    AMD-Vi: ivrs, add hid:MSFT0201, uid:1, rdevid:0x60
>>>    AMD-Vi: ivrs, add hid:AMDI0020, uid:ID04, rdevid:0x98
>>>    AMD-Vi: Using global IVHD EFR:0x246577efa2254afa, EFR2:0x10
>>>    pci 0000:00:00.2: AMD-Vi: IOMMU performance counters supported
>>>    AMD-Vi: Extended features (0x246577efa2254afa, 0x10): PPR NX GT 
>>> [5] IA GA PC GA_vAPIC
>>>    AMD-Vi: Interrupt remapping enabled
>>>    AMD-Vi: Virtual APIC enabled
>>>
>>> One thing that may be worth your attention anyway: MSFT0201:00 is the 
>>> only ACPI HID device in any IOMMU group on this system, alone in 
>>> group 0. It is not the TPM - tpm0 is MSFT0101:00. So the single 
>>> non-PCI device the IOMMU manages here is Pluton.
>>>
>>
>> Got it.  Then this is likely not a Pluton/TPM related issue as it 
>> originally seemed as the BIOS has that UID aligned.
>>
>>> For what it is worth, when I tested with PlutonSecurityProcessor set 
>>> to Disable earlier in this thread the machine still failed to wake, 
>>> though I did not check at the time whether MSFT0201 was still present 
>>> in the IVRS in that configuration. I can re-run that combination and 
>>> capture the IVRS block if it would help.
>>>
>>>  > I think you're missing commit e9f850ba66cd.
>>>
> 
> e9f850ba66cd works, thank you. With it applied I get 
> /sys/class/rtc/rtc0/wakealarm and the -ENXIO message is gone:
> 
>    rtc_cmos PNP0B00:00: RTC can wake from S4
>    rtc_cmos PNP0B00:00: registered as rtc0
>    rtc_cmos PNP0B00:00: alarms up to one month, y3k, 114 bytes nvram
>    8: ... IO-APIC 8-edge rtc0
> 
> With amd_iommu=off and all 24 CPUs I now get fully unattended timed cycles:
> 
>    pm_wakeup_irq:        9
>    Last S0i3 Status:     Success
>    Time (in us) to S0i3: 430540
>    Residency Time:       13282379
>    last_hw_sleep:        13282379
>    ff_rt_clk:            incremented on each cycle
> 
> One small platform note in case it is useful to you: IRQ 8 never fires 
> on this machine while the system is running. If I arm an alarm and poll, 
> the alarm counts down and expires exactly on schedule but the IRQ 8 
> counter stays at zero:
> 
>    t=0   arming +8       irq8=0
>    t=1-7 wakealarm=1790194169   irq8=0
>    t=8   wakealarm=''           irq8=0
>    t=15  wakealarm=''           irq8=0
> 
> Identical with the IOMMU on and off, so it is not an interrupt remapping 
> effect. The wake itself comes through the ACPI fixed-feature RTC event 
> rather than IRQ 8, which is why it works anyway: ff_rt_clk increments 
> and pm_wakeup_irq reports 9. IRQ 8 does register a single count at 
> resume time.
> 
> Now the part that matters. With the IOMMU enabled and all 24 CPUs, the 
> same timed suspend does not wake. Recorded before the cycle:
> 
>    === 2026-09-23T22:15:02 BEFORE
>    cmdline:     (no iommu parameters)
>    cpus online: 0-23
>    iommu:       ivhd0
>    ff_rt_clk:   0
>    S0ix Entry Time: 0
>    S0ix Exit Time: 0
>    Residency Time: 0
>    suspending with +30s alarm
> 
> Nothing was written after that. The machine never resumed and needed a 
> forced power off.
> 
> So with the IOMMU enabled, every independent wake mechanism on this 
> machine has now been tried and none of them work:
> 
>    lid switch (PNP0C0D)                          no wake
>    internal keyboard (i8042, via AMD GPIO)       no wake
>    power button (ACPI fixed feature)             no wake
>    USB mouse, wakeup armed on device and hub     no wake
>    ACPI fixed-feature RTC alarm                  no wake
> 
> The last one seems the most pointed to me. The RTC alarm wakes the 
> machine through the ACPI SCI on IRQ 9, and that is exactly the path that 
> works when amd_iommu=off - pm_wakeup_irq reports 9 on every successful 
> cycle. With the IOMMU enabled the same mechanism produces nothing at 
> all, and ff_rt_clk stays at 0 across the attempt.
> 
> Still to come: the 7.3-rc4 test, and a manual walk through BIOS setup 
> looking for storage link power or D3 controls that think-lmi does not 
> export.

the bug reproduces on 7.3-rc4.

Built 7.3-rc4 from the torvalds snapshot, migrated my 6.18.51 config 
with make olddefconfig, IOMMU enabled, all 24 CPUs online, no IOMMU boot 
parameters. Same result: the machine suspends and never wakes, forced 
power off required.

Both attempts recorded by the same script, neither produced a resume block:

   === 2026-09-23T22:15:02 BEFORE
   cmdline:     (6.18.51, no iommu parameters)
   cpus online: 0-23
   iommu:       ivhd0
   ff_rt_clk:   0
   S0ix Entry Time: 0 / Exit Time: 0 / Residency Time: 0
   suspending with +30s alarm

   === 2026-09-23T22:39:06 BEFORE
   cmdline:     (7.3.0-rc4, no iommu parameters)
   cpus online: 0-23
   iommu:       ivhd0
   ff_rt_clk:   0
   S0ix Entry Time: 0 / Exit Time: 0 / Residency Time: 0
   suspending with +30s alarm

Both used a programmed RTC wakealarm at +30s and were left untouched for 
several minutes.

Relevant details of the 7.3-rc4 build, so this is not a configuration 
difference:

   CONFIG_AMD_IOMMU=y, CONFIG_IOMMU_PT=y, CONFIG_IOMMU_PT_AMDV1=y
   CONFIG_NR_CPUS=24, CONFIG_MODULES not set
   CONFIG_DEBUG_FS=y, CONFIG_PM_DEBUG=y, CONFIG_DYNAMIC_DEBUG=y
   CONFIG_AMD_MP2_STB=y, CONFIG_X86_MSR=y, CONFIG_DMIID=y
   CONFIG_LOG_BUF_SHIFT=20, CONFIG_RANDSTRUCT_NONE=y
   kernel is not tainted, builds with zero warnings

Only 243 config lines differ between my 6.18.51 and 7.3-rc4 configs. 
Both commits you pointed me at are present in 7.3-rc4: rtc-cmos uses 
platform_get_irq_optional(), and iommu/amd carries the "No ACPI device 
matched UID" check. As on 6.18, that FW_BUG message does not appear here 
either.

Two things worth noting about the 7.3 build specifically. The new IOMMU 
page table layer is in use (CONFIG_IOMMU_PT / IOMMU_PT_AMDV1 replacing 
CONFIG_IOMMU_IO_PGTABLE), so that rework does not change the outcome. 
And AMD_PMF and DRM_ACCEL_AMDXDNA are not compiled into either kernel - 
DRM_ACCEL and AMD_SFH_HID are both off in my config - so neither driver 
is involved in any of these results at all, which is a stronger 
statement than the initcall_blacklist tests I reported earlier.

Where that leaves things. With the IOMMU enabled and all 24 CPUs, on 
both 6.18.51 and 7.3-rc4:

   lid switch, internal keyboard, power button, USB mouse with wakeup
   armed, and a programmed ACPI RTC alarm        all fail to wake

   intremap=off, iommu=pt, amd_iommu_intr=legacy  all still hang
   pm_test freezer / devices / platform           all pass
   PC6 and CC6 enabled per amd-s2idle             yes

   amd_iommu=off                                  wakes normally,
                                                  Last S0i3 Status Success,
                                                  13.3s residency,
                                                  pm_wakeup_irq 9

Since it fails identically on 6.18 and 7.3 there is no working kernel 
version to bisect against. The only variable that changes the outcome is 
whether the IOMMU is enabled, and secondarily whether more than 16 CPUs 
are brought up at boot.


  reply	other threads:[~2026-09-23 20:48 UTC|newest]

Thread overview: 32+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-22 13:43 [BUG] s2idle: unrecoverable sleep on ThinkPad P16s Gen 4 AMD, (Strix Point) when more than 16 logical CPUs are online Fourhundred Thecat
2026-09-22 14:47 ` Mario Limonciello
2026-09-23  5:58   ` Fourhundred Thecat
2026-09-23 12:41     ` Mario Limonciello
2026-09-23 13:51       ` Fourhundred Thecat
2026-09-23 14:36         ` Mario Limonciello
2026-09-23 16:12           ` Fourhundred Thecat
2026-09-23 16:15             ` Mario Limonciello
2026-09-23 16:20               ` Mario Limonciello
2026-09-23 16:37                 ` Fourhundred Thecat
2026-09-23 16:41                   ` Mario Limonciello
2026-09-23 16:54                     ` Fourhundred Thecat
2026-09-23 16:59                       ` Mario Limonciello
2026-09-23 18:02                         ` Fourhundred Thecat
2026-09-23 18:11                           ` Mario Limonciello
2026-09-23 18:14                             ` Fourhundred Thecat
2026-09-23 18:25                               ` Mario Limonciello
2026-09-23 18:39                                 ` Fourhundred Thecat
2026-09-23 18:57                                   ` Mario Limonciello
2026-09-23 19:13                                     ` Fourhundred Thecat
2026-09-23 19:20                                       ` Mario Limonciello
2026-09-23 20:26                                         ` Fourhundred Thecat
2026-09-23 20:48                                           ` Fourhundred Thecat [this message]
2026-09-23 21:10                                             ` Mario Limonciello
2026-09-24  5:36                                               ` Fourhundred Thecat
2026-09-25  7:05                                               ` Fourhundred Thecat
2026-09-25 13:30                                                 ` Mario Limonciello
2026-09-26  4:13                                                   ` Fourhundred Thecat
2026-09-26 18:29                                                     ` Mario Limonciello
2026-09-29  5:43                                                       ` Fourhundred Thecat
2026-09-29  5:57                                                         ` Fourhundred Thecat
2026-09-29 13:37                                                           ` Mario Limonciello

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=afb1949a-8465-cfc8-bddb-bc27783af223@ik.me \
    --to=400thecat@ik.me \
    --cc=Shyam-sundar.S-k@amd.com \
    --cc=hansg@kernel.org \
    --cc=ilpo.jarvinen@linux.intel.com \
    --cc=linux-pm@vger.kernel.org \
    --cc=mario.limonciello@amd.com \
    --cc=platform-driver-x86@vger.kernel.org \
    --cc=rafael@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox