Linux userland API discussions
 help / color / mirror / Atom feed
From: "Ilpo Järvinen" <ilpo.jarvinen@linux.intel.com>
To: Ferran Duarri <ferran.duarri@me.com>
Cc: Bjorn Helgaas <bhelgaas@google.com>,
	linux-pci@vger.kernel.org,  linux-api@vger.kernel.org,
	LKML <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v2] PCI/sysfs: document the link speed and width attributes
Date: Fri, 21 Aug 2026 11:09:17 +0300 (EEST)	[thread overview]
Message-ID: <fc897ad3-ed5e-8755-f5d7-4d6dbe7f57c0@linux.intel.com> (raw)
In-Reply-To: <20260820200359.283335-1-ferran.duarri@me.com>

On Thu, 20 Aug 2026, Ferran Duarri wrote:

> max_link_speed, max_link_width, current_link_speed and current_link_width
> have been exported under /sys/bus/pci/devices/.../ since 2018, by
> commit 56c1af4606f0 ("PCI: Add sysfs max_link_speed/width, current_link_speed/width, etc"),
> and none of the four appear anywhere in Documentation/ABI.
> 
> The gap matters most for current_link_speed. current_link_speed_show()
> performs a fresh PCI_EXP_LNKSTA read on every open, so the value reflects
> the link state at that instant. Modern GPUs retrain their link continuously
> as part of idle power management, which means a single read can legitimately
> return any speed the link supports, not the speed the link will use under
> load.
> 
> Observed on an RTX 5070 in a PCIe 4.0 x16 slot, same boot, no configuration
> change between the two reads: 5.0 GT/s while idle, 16.0 GT/s under load.
> Comparing current_link_speed against max_link_speed at idle is therefore not
> a valid test for a degraded link, though it reads like one.
> 
> Document all four. For the max_* pair, state that each reports the
> capability of the device it is read from and not a property of the link: a
> link trains at the lower of what its two ends support, so an endpoint
> capable of more than the port above it reports the higher figure while that
> port reports the lower one. Record where each value comes from, which
> differs between the two attributes. max_link_speed is derived from the
> Supported Link Speeds Vector in Link Capabilities 2, capped by Max Link
> Speed in Link Capabilities, synthesized from the latter alone on devices
> predating PCIe r3.0, and cached at enumeration. max_link_width is read from
> Maximum Link Width in Link Capabilities on each access.
> 
> For current_link_speed, state that it is instantaneous, that comparing it
> against max_link_speed at idle is not a valid degradation test, and that
> callers wanting what the link will actually deliver should sample under
> load -- noting that max_link_speed is not that figure either, being one
> end's capability rather than the link's.
> 
> No functional change.
> 
> Signed-off-by: Ferran Duarri <ferran.duarri@me.com>
> ---
> Changes in v2, all corrections to what v1 claimed rather than new material:
> 
>  - max_link_speed: v1 called it "the ceiling the link may negotiate, which
>    is the lower of what the two ends of the link support". That is wrong,
>    and v1 contradicted it one sentence later. max_link_speed_show() calls
>    pcie_get_speed_cap(), which returns the capability of the device being
>    read and never consults the other end of the link.
>  - max_link_speed: v1 said the value is read from the Max Link Speed field
>    of Link Capabilities. pcie_get_supported_speeds() derives it from the
>    Supported Link Speeds Vector in Link Capabilities 2, masks it against
>    Max Link Speed, and synthesizes from Max Link Speed alone only on
>    devices predating PCIe r3.0.
>  - max_link_speed: v1 did not say the value is read once at enumeration and
>    cached in pci_dev->supported_speeds. Since the current_link_speed entry
>    states that nothing is cached there, a reader could reasonably infer the
>    same of max_link_speed. It does not hold.
>  - current_link_speed: v1 advised callers wanting the ceiling to use
>    max_link_speed. That overestimates whenever the upstream port is the
>    slower end. v2 says to sample under load and warns that max_link_speed
>    is one end's capability, not the link's.
>  - max_link_width: register attribution was correct and is unchanged in
>    substance, reworded only for the same device-versus-link distinction.
>  - Dropped a private Forward-Port-Notes: trailer that should not have been
>    in the commit message.
> 
>  Documentation/ABI/testing/sysfs-bus-pci | 78 +++++++++++++++++++++++++
>  1 file changed, 78 insertions(+)
> 
> diff --git a/Documentation/ABI/testing/sysfs-bus-pci b/Documentation/ABI/testing/sysfs-bus-pci
> index b767db2c52cb..ee1846f3dafa 100644
> --- a/Documentation/ABI/testing/sysfs-bus-pci
> +++ b/Documentation/ABI/testing/sysfs-bus-pci
> @@ -174,6 +174,84 @@ Description:
>  		similiar to writing 1 to their individual "reset" file, so use
>  		with caution.
>  
> +What:		/sys/bus/pci/devices/.../max_link_speed
> +Date:		September 2018
> +Contact:	linux-pci@vger.kernel.org
> +Description:
> +		The maximum link speed this device is capable of, as a
> +		human-readable string such as "16.0 GT/s PCIe".
> +
> +		Derived from the Supported Link Speeds Vector in the device's
> +		Link Capabilities 2 register, capped by the Max Link Speed
> +		field in Link Capabilities. Devices predating PCIe r3.0 have no
> +		Link Capabilities 2, and for those the value is synthesized
> +		from Max Link Speed alone.

IMO it's unnecessary detail to say it's synthetized. For <= PCIe r3.0, 
kernel does indeed synthetize something, but that "supported speeds" which 
is a set of values. pcie_get_speed_cap() used in max_link_speed_show() 
only returns the max value out of that set which, for sane devices, should 
be same as Max Link Speed.

> Read once during enumeration and
> +		cached thereafter, so unlike current_link_speed it does not
> +		change between reads.
> +
> +		This is the device's own capability, not a property of the
> +		link. A link trains at the lower of what its two ends support,
> +		so an endpoint capable of a higher speed than the port above it
> +		reports that higher speed here while the port reports the lower
> +		one. Reading one end therefore does not tell you what the link
> +		will do; read both ends and take the lower.
> +
> +		Present only for PCI Express devices.
> +
> +What:		/sys/bus/pci/devices/.../max_link_width
> +Date:		September 2018
> +Contact:	linux-pci@vger.kernel.org
> +Description:
> +		The maximum link width this device is capable of, in lanes,
> +		e.g. "16". Read from the Maximum Link Width field of the
> +		device's Link Capabilities register.
> +
> +		As with max_link_speed this is the device's own capability, not
> +		a property of the link; a link trains at the lower of what its
> +		two ends support.
> +
> +		Present only for PCI Express devices.
> +
> +What:		/sys/bus/pci/devices/.../current_link_speed
> +Date:		September 2018
> +Contact:	linux-pci@vger.kernel.org
> +Description:
> +		The speed the link is operating at right now, as a
> +		human-readable string such as "16.0 GT/s PCIe". Read fresh from
> +		the device's Link Status register on every read of this file;
> +		nothing is cached.
> +
> +		This value is instantaneous and may change at any time. A link
> +		is permitted to retrain to a lower speed and back, and devices
> +		with aggressive link power management (GPUs in particular) do so
> +		routinely while idle. Two reads seconds apart, with no
> +		configuration change in between, can legitimately differ by
> +		several generations.
> +
> +		Consequently, comparing this attribute against max_link_speed is
> +		not by itself a test for a degraded link: an idle device will
> +		frequently report a lower speed and is working correctly.
> +		Callers that need a figure representing what the link will
> +		actually deliver should sample while the device is under load.
> +		max_link_speed is not that figure either: it reports one end's
> +		capability, and the link is limited by the lower of its two
> +		ends.

The speed may also be capped by Target Link Speed in Link Control 2 that 
is under control of PCIe BW Controller (bwctrl).

> +
> +		Present only for PCI Express devices.
> +
> +What:		/sys/bus/pci/devices/.../current_link_width
> +Date:		September 2018
> +Contact:	linux-pci@vger.kernel.org
> +Description:
> +		The width the link is operating at right now, in lanes, e.g.
> +		"16". Read fresh from the device's Link Status register on every
> +		read of this file.
> +
> +		As with current_link_speed, this is instantaneous. Links may
> +		also narrow and re-widen under link power management.
> +
> +		Present only for PCI Express devices.
> +
>  What:		/sys/bus/pci/devices/.../vpd
>  Date:		February 2008
>  Contact:	Ben Hutchings <bwh@kernel.org>
> 

-- 
 i.


  parent reply	other threads:[~2026-08-21  8:09 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-20 18:42 [PATCH] PCI/sysfs: document the link speed and width attributes Ferran Duarri
2026-08-20 19:53 ` [PATCH v2] " Ferran Duarri
2026-08-20 20:03 ` Ferran Duarri
2026-08-21  4:31   ` Greg KH
2026-08-21  8:09   ` Ilpo Järvinen [this message]
2026-08-21  8:33 ` [PATCH v3] " Ferran Duarri

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=fc897ad3-ed5e-8755-f5d7-4d6dbe7f57c0@linux.intel.com \
    --to=ilpo.jarvinen@linux.intel.com \
    --cc=bhelgaas@google.com \
    --cc=ferran.duarri@me.com \
    --cc=linux-api@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox