All of lore.kernel.org
 help / color / mirror / Atom feed
From: Hans de Goede <hansg@kernel.org>
To: Sergey Lebedev <lsa.uz@pm.me>,
	Dmitry Torokhov <dmitry.torokhov@gmail.com>,
	Maximilian Luz <luzmaximilian@gmail.com>
Cc: linux-input@vger.kernel.org, platform-driver-x86@vger.kernel.org,
	linux-kernel@vger.kernel.org
Subject: Re: soc_button_array: ignored -EPROBE_DEFER loses the buttons on 68% of boots (Surface Pro 11)
Date: Tue, 8 Sep 2026 23:46:21 +0200	[thread overview]
Message-ID: <0fe7a8b6-f3df-4822-a03a-d072b86cdfa8@kernel.org> (raw)
In-Reply-To: <20260830141355.55898-1-lsa.uz@pm.me>

Hi,

On 30-Aug-26 16:14, Sergey Lebedev wrote:
> Hello,
> 
> On a Microsoft Surface Pro 11 for Business (Intel, Lunar Lake) the power and
> volume buttons fail to appear on most boots. The probe runs, gives up
> silently, and the device is lost until something rebinds it by hand.
> 
> Measured across the 40 boots in this machine's journal, using the monotonic
> offset of the first "input: gpio-keys" message:
> 
>   driver attached on its own   13 boots   at 1.28 - 1.78 s
>   did not attach               27 boots   68 %
> 
> The timing is binary. Either the inputs appear at about 1.3 s, or they never
> appear at all - no intermediate values in 40 boots.
> 
> Where it fails
> --------------
> 
> Nothing at all is logged on a failing boot, so I built with
> soc_button_array.dyndbg=+p on the kernel command line. On four consecutive
> failing boots the probe printed exactly one line:
> 
>   [    1.286567] soc_button_array MSHW0040:00: OEM Platform Revision 129
> 
> That eliminates most of the possibilities:
> 
>   - the probe does run
>   - it passes soc_device_check_MSHW0040() - oem_platform_rev is 129, not 0
>   - it passes gpiod_count() - the dev_dbg("no GPIO attached, ignoring...")
>     would have printed with dyndbg on, and did not
> 
> By elimination the buttons are lost inside soc_button_device_create(), at
> 
>         error = soc_button_lookup_gpio(&pdev->dev, info->acpi_index,
>                                        &gpio, &irq);
>         if (error || irq < 0) {
>                 /*
>                  * Skip GPIO if not present. Note we deliberately
>                  * ignore -EPROBE_DEFER errors here. ...
>                  */
>                 continue;
>         }
> 
> Every button is skipped, n_buttons stays 0, and the two returns that follow
> are both silent:
> 
>         if (n_buttons == 0) { error = -ENODEV; goto err_free_mem; }
>   and
>         if (!priv->children[0] && !priv->children[1]) return -ENODEV;
> 
> which is why nothing reaches the log.
> 
> The GPIO controller here is INTC105D, five gpiochips driven by
> pinctrl-intel-platform. Both that and soc_button_array are modules in this
> config:
> 
>   CONFIG_INPUT_SOC_BUTTON_ARRAY=m
>   CONFIG_PINCTRL_INTEL_PLATFORM=m
> 
> A manual bind always works: 25 unbind/bind cycles at runtime gave 25
> successes and 0 failures. The failure exists only in the early-boot window.
> 
> One hypothesis, tested and wrong
> ---------------------------------
> 
> Since both are modules, load order looked like the cause, so I tried
> 
>   softdep soc_button_array pre: pinctrl-intel-platform
> 
> verified applied (modprobe --show-depends inserts pinctrl first), with the
> rescue that normally binds it by hand disabled so nothing masked the result.
> Five boots: 2 successes, 3 failures - the same rate as before.
> 
> pinctrl_intel_platform is a dependency of fifteen other modules and is
> loaded long before either way. softdep orders modules; the race is against
> the asynchronous probe of the individual INTC105D:0N devices, which register
> their gpiochips at their own pace. Reporting this so nobody else spends the
> evening on it.
> 
> Why no patch
> ------------
> 
> The comment says -EPROBE_DEFER is discarded because Intel uses "virtual
> GPIOs" that never materialise, so deferring on them would hang the probe
> forever. A fix has to tell that case apart from a controller that simply has
> not probed yet, and I do not know which signal you would want to use for
> that. Guessing would risk reintroducing the problem the workaround exists
> for.
> > Happy to test whatever shape you would prefer - the machine is here and I
> can build and boot kernels on it.

Thank you for your detailed bug report. I've just submitted a patch which
I believe should fix this:

https://lore.kernel.org/platform-driver-x86/20260908214410.70422-1-johannes.goede@oss.qualcomm.com/

Please give this a try and confirm if it fixes things.

Regards,

Hans



  reply	other threads:[~2026-09-08 21:46 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-30 14:14 soc_button_array: ignored -EPROBE_DEFER loses the buttons on 68% of boots (Surface Pro 11) Sergey Lebedev
2026-09-08 21:46 ` Hans de Goede [this message]
2026-09-09  9:41   ` Hans de Goede

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=0fe7a8b6-f3df-4822-a03a-d072b86cdfa8@kernel.org \
    --to=hansg@kernel.org \
    --cc=dmitry.torokhov@gmail.com \
    --cc=linux-input@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lsa.uz@pm.me \
    --cc=luzmaximilian@gmail.com \
    --cc=platform-driver-x86@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.