* [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
@ 2026-08-24 22:37 Priyank Rathod
2026-08-24 22:51 ` sashiko-bot
0 siblings, 1 reply; 5+ messages in thread
From: Priyank Rathod @ 2026-08-24 22:37 UTC (permalink / raw)
To: Bjorn Helgaas, Priyank Rathod, Shuah Khan, Kees Cook,
Gustavo A. R. Silva, Jonathan Corbet, Shuah Khan
Cc: linux-kernel, linux-pci, linux-kselftest, linux-hardening,
linux-doc, Ilpo Järvinen
Per PCIe Base Specification r6.0, sec 8.4.4 ("Lane Margining at
Receiver"), PCIe devices operating at 16.0 GT/s (Gen 4) or higher data
rates support the Lane Margining at Receiver Extended Capability
(ID 0x27), and it is mandatory for receivers operating at 64.0 GT/s
(Gen 6) or higher data rates. Lane Margining allows software to
evaluate high-speed link margins by measuring timing and voltage steps
for each individual physical lane and receiver.
Add driver and debugfs support for PCIe Lane Margining at Receiver:
- Add Lane Margining at Receiver Extended Capability register
definitions (PCI_EXT_CAP_ID_LMR, PCI_LMR_PORT_CAP, PCI_LMR_PORT_STS,
PCI_LMR_LANE_CTRL, PCI_LMR_LANE_STS) to <uapi/linux/pci_regs.h>.
- Add Kconfig option CONFIG_PCIE_LMR (under drivers/pci/pcie/Kconfig)
dependent on DEBUG_FS.
- Implement drivers/pci/pcie/margin.c to probe the capability on Gen4+
links and expose per-device debugfs entries under:
/sys/kernel/debug/pci/pcie_lmr_<pci_dev_name>/
providing control over margining enablement, receiver selection, and
execution of timing/voltage margin step commands. Distinguish
between missing mandatory LMR capability on Gen6+ vs optional on
Gen4/Gen5.
- Hook pci_lmr_init() into pci_init_capabilities() during device probe
in drivers/pci/probe.c and pci_lmr_exit() into drivers/pci/remove.c.
- Add kselftest script under tools/testing/selftests/pcie_lmt/pcie_lmt.sh
to test debugfs capability reads, enablement, and stepping.
- Add MAINTAINERS entry for PCIe Lane Margining at Receiver (LMR).
Signed-off-by: Priyank Rathod <rathodpriyank@google.com>
---
Changes in v6:
- Added kernel documentation under Documentation/PCI/pcie-lmr.rst and indexed in Documentation/PCI/index.rst (Ilpo Järvinen).
- Updated MAINTAINERS with Documentation/PCI/pcie-lmr.rst (Ilpo Järvinen).
- Aligned capability bit naming and comments with PCIe Base Specification r6.0 sec 8.4.4 Table "Report Margining Capabilities Payload" (Ilpo Järvinen).
- Clarified Sample Multiple Receivers concurrency verification and rules across physical lanes in kerneldoc and documentation (Ilpo Järvinen).
- Refactored pci_lmr_run_cmd() to pass struct pci_margin_dev *mdev directly, eliminating redundant NULL checks and using mdev->num_lanes (Ilpo Järvinen).
- Converted PCI config read/write return checking across all helpers to pcibios_err_to_errno() (Ilpo Järvinen).
- Reversed return logic in pci_lmr_demargin_lane() to return early on error (Ilpo Järvinen).
- Refactored margin_lane_step_write() to eliminate bool is_voltage parameter, using command type (LMR_TYPE_TIMING / LMR_TYPE_VOLTAGE) and switch/case with consolidated bounds checks (Ilpo Järvinen).
- Renamed __pci_suspend_lmr_locked() to pci_lmr_disable_locked() to avoid PM terminology confusion and added lockdep_assert_held(&mdev->lock) (Ilpo Järvinen).
- Replaced -EACCES with -EBUSY across debugfs show/write callbacks when margining is inactive (Ilpo Järvinen).
- Clarified comment for active operating link speed check (Gen4+ capability vs dynamically operating speed) in margin_enable_write() (Ilpo Järvinen).
- Added WARN_ON_ONCE(!dev) check in pci_lmr_init() (Ilpo Järvinen).
- Demoted capability detection log message from pci_info to pci_dbg to prevent boot log noise (Ilpo Järvinen).
- Fixed timing step mask extraction in pci_lmr_cache_rx_info() to use 6-bit LMR_TIMING_STEP_MASK (sashiko-bot).
- Resumed runtime PM via pm_runtime_resume_and_get() before performing config space reads in margin_enable_write() (sashiko-bot).
- Switched to pm_runtime_put_sync() during margining teardown (sashiko-bot).
- Link to v5: https://lore.kernel.org/r/20260820-pcie-lmt-v5-1-943b3b0e18bf@google.com
PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
Per PCIe Base Specification r6.0, section 8.4.4 ("Lane Margining at Receiver"),
PCIe devices operating at 16.0 GT/s (Gen 4) or higher data rates support the
Lane Margining at Receiver Extended Capability (ID 0x27), and it is mandatory
for receivers operating at 64.0 GT/s (Gen 6) or higher data rates.
Lane Margining allows system software to evaluate high-speed link signal
integrity and margins by measuring timing and voltage steps for each physical
lane and receiver independently.
This series introduces kernel driver support, debugfs controls, and a
kselftest automation script for PCIe Lane Margining at Receiver (LMR/LMT).
==============================================================================
1. How to Enable & Configure
==============================================================================
Enable the Kconfig option under PCI support:
CONFIG_PCIE_LMR=y (or =m)
(Depends on CONFIG_PCI and CONFIG_DEBUG_FS)
Upon boot or device hotplug on Gen4+ links (>= 16.0 GT/s), the driver probes
Extended Capability ID 0x27 and exposes per-device debugfs interfaces:
/sys/kernel/debug/pci/pcie_lmr_<domain>:<bus>:<dev>.<func>/
==============================================================================
2. How to Use the Debugfs Interface (Manual Margining)
==============================================================================
Inspect device-wide margining capabilities and port status:
# Inspect root device LMR capabilities & status
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/capabilities
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/port_status
Enable active Lane Margining on the device:
# Enable Lane Margining state machine
echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
Inspect and step individual lanes (e.g. lane0):
# Select target receiver (0 = local receiver, 1..6 = retimers/link partners)
echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/receiver
# Check available timing and voltage steps for this receiver
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/caps
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_timing_steps
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_voltage_steps
# Step timing margin or voltage margin offset
echo 2 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
# Reset margin offset back to nominal (0)
echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
Disable Lane Margining when finished:
echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
==============================================================================
3. How to Run Automated Kselftests Using the Test Script
==============================================================================
An automated kselftest script is included to test capability reads, receiver
selection, and margining commands across all enumerated LMR devices:
# Run directly as root
sudo ./tools/testing/selftests/pcie_lmt/pcie_lmt.sh
Or run via the kselftest Makefile harness:
make -C tools/testing/selftests TARGETS=pcie_lmt run_tests
Sample script output on an LMR-capable device:
pcie_lmt: testing PCIe LMR debugfs entries
pcie_lmt: probing device pcie_lmr_0000:01:00.0
pcie_lmr_0000:01:00.0: capabilities read OK
pcie_lmr_0000:01:00.0: port_status read OK
pcie_lmr_0000:01:00.0: margining enabled OK
pcie_lmr_0000:01:00.0: testing lane0
pcie_lmr_0000:01:00.0: testing lane1
pcie_lmr_0000:01:00.0: margining disabled OK
pcie_lmt [PASS]
To: Bjorn Helgaas <bhelgaas@google.com>
To: Shuah Khan <shuah@kernel.org>
Cc: linux-kernel@vger.kernel.org
Cc: linux-pci@vger.kernel.org
Cc: linux-kselftest@vger.kernel.org
Cc: Ilpo Järvinen <ilpo.jarvinen@linux.intel.com>
Changes in v5:
- Sorted #include directives alphabetically and added missing includes for bits.h, bitfield.h, cleanup.h, overflow.h, and slab.h (Ilpo Järvinen).
- Converted bitmasks to GENMASK() and BIT() macros and used FIELD_PREP() and FIELD_GET() instead of manual bit shifts (Ilpo Järvinen).
- Added pci_lmr_sts_payload() helper to cleanly extract the status payload byte before applying step and capability masks (Ilpo Järvinen).
- Replaced manual mutex locking sequences with guard(mutex)(&mdev->lock) across show and write callbacks to simplify control flow (Ilpo Järvinen).
- Documented mutex lock protection scope in kerneldoc for struct pci_margin_dev (Ilpo Järvinen).
- Used standard PCI_POSSIBLE_ERROR(), str_yes_no(), and scnprintf() helpers throughout the driver (Ilpo Järvinen).
- Clarified receiver range (0..6 per PCIe r6.0 sec 8.4.4; 7 reserved) in comments and validation checks (Ilpo Järvinen).
- Deduplicated timing and voltage show/write handlers using margin_lane_steps_show() and margin_lane_step_write() (Ilpo Järvinen).
- Placed speed check immediately following pcie_get_speed_cap() and handled PCI_SPEED_UNKNOWN (Ilpo Järvinen).
- Converted lanes in struct pci_margin_dev to a flexible array member with __counted_by(num_lanes) allocated via struct_size() (Ilpo Järvinen).
Changes in v4:
- Added Sample Multiple Receivers (Bit 5) concurrency verification in margin_lane_timing_write() and margin_lane_voltage_write() per PCIe r6.0 sec 8.4.4, returning -EBUSY if another lane on the same receiver is already margined when simultaneous lane margining is not supported.
- Added active operating link speed verification (PCI_EXP_LNKSTA_CLS >= 16.0 GT/s) in margin_enable_write() before enabling LMR, as LMR commands are physically undefined on links operating at Gen1/Gen2/Gen3 speeds.
- Added fast-path hardware NAK detection in pci_lmr_run_cmd() to return -EOPNOTSUPP immediately if a receiver echoes MTYPE == NO_CMD (0x7) after command issuance rather than waiting 150ms for a timeout.
- Added pci_reset_lmr() hooked into __pci_reset_function_locked() to synchronize software state and demargin on FLR or Secondary Bus Reset.
- Comprehensive NULL pointer checks and array/lane/receiver bounds checks added across all internal helpers and debugfs write handlers.
- Added MAINTAINERS entry for PCIe Lane Margining at Receiver (LMR).
Changes in v2:
- Fixed NO_CMD (0x7) clearing in pci_lmr_run_cmd() before issuing new commands per PCIe r6.0 sec 8.4.4.
- Protected plane->rx updates with mdev->lock in margin_lane_receiver_write().
- Corrected Margining Port Capabilities bit definition to PCI_LMR_PORT_CAP_USES_SW_READY (0x0001) in <uapi/linux/pci_regs.h>.
- Updated kselftest script (pcie_lmt.sh) to locate LMR debugfs entries.
- Validated integer bounds against LMR_MAX_TIMING_STEP / LMR_MAX_VOLTAGE_STEP before narrowing u8 cast.
- Moved mdev->enabled checks inside mutex_lock(&mdev->lock) to eliminate TOCTOU races.
- Checked return values of all pci_read_config_word() calls, propagating -EIO on failure.
- Eliminated dead store of cap in margin_enable_write().
- Explicitly checked speed == PCIE_SPEED_64_0GT in pci_lmr_init() to avoid misidentifying PCI_SPEED_UNKNOWN (0xFF) as Gen6.
---
Documentation/PCI/index.rst | 1 +
Documentation/PCI/pcie-lmr.rst | 171 +++++
MAINTAINERS | 8 +
drivers/pci/pci-driver.c | 1 +
drivers/pci/pci.c | 4 +-
drivers/pci/pci.h | 12 +
drivers/pci/pcie/Kconfig | 12 +
drivers/pci/pcie/Makefile | 1 +
drivers/pci/pcie/margin.c | 1061 ++++++++++++++++++++++++++
drivers/pci/probe.c | 1 +
drivers/pci/remove.c | 1 +
include/linux/pci.h | 6 +
include/uapi/linux/pci_regs.h | 18 +
tools/testing/selftests/Makefile | 1 +
tools/testing/selftests/pcie_lmt/Makefile | 3 +
tools/testing/selftests/pcie_lmt/pcie_lmt.sh | 105 +++
16 files changed, 1405 insertions(+), 1 deletion(-)
diff --git a/Documentation/PCI/index.rst b/Documentation/PCI/index.rst
index 5d720d2a415e..9170c98cbf3f 100644
--- a/Documentation/PCI/index.rst
+++ b/Documentation/PCI/index.rst
@@ -20,3 +20,4 @@ PCI Bus Subsystem
controller/index
boot-interrupts
tph
+ pcie-lmr
diff --git a/Documentation/PCI/pcie-lmr.rst b/Documentation/PCI/pcie-lmr.rst
new file mode 100644
index 000000000000..1ebb8317cc83
--- /dev/null
+++ b/Documentation/PCI/pcie-lmr.rst
@@ -0,0 +1,171 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+=======================================================
+PCI Express Lane Margining at Receiver (LMR) Subsystem
+=======================================================
+
+:Author: Priyank Rathod <rathodpriyank@google.com>
+:Copyright: 2026 Google LLC
+
+Overview
+========
+
+Lane Margining at Receiver (LMR), specified in the PCI Express Base
+Specification (r6.0+ sec 8.4.4), allows system software to evaluate high-speed
+link physical signal integrity and eye margins. LMR measures available timing
+(jitter/phase) and voltage margin offsets for each physical lane and receiver
+independently while the link is operating in active L0 state.
+
+Lane Margining Extended Capability (ID 0x27) is optional for links operating at
+16.0 GT/s (PCIe Gen 4) and 32.0 GT/s (Gen 5), and is mandatory for receivers
+operating at 64.0 GT/s (Gen 6) and higher.
+
+Target Receivers
+================
+
+Each physical lane can margin up to 7 distinct receivers per PCIe link:
+
+* **Receiver 0 (Local Receiver)**: The receiver in the immediate link partner.
+* **Receivers 1 to 6 (Retimers)**: Retimer pseudo-ports along the physical link
+ (up to 3 retimers, each with upstream and downstream pseudo-ports).
+* **Receiver 7**: Reserved per PCIe Base Specification.
+
+Kernel Configuration
+====================
+
+Enable the kernel configuration option under PCI support:
+
+.. code-block:: none
+
+ CONFIG_PCIE_LMR=y (or =m)
+
+Dependencies:
+* ``CONFIG_PCI``
+* ``CONFIG_DEBUG_FS``
+
+Debugfs Interface Guide
+=======================
+
+When an LMR-capable device is enumerated on a Gen4+ link, the kernel exposes
+per-device control and status files under debugfs:
+
+.. code-block:: none
+
+ /sys/kernel/debug/pci/pcie_lmr_<domain>:<bus>:<dev>.<func>/
+
+Device-Level Attributes
+-----------------------
+
+* ``capabilities`` (read-only):
+ Displays the 16-bit Margining Port Capabilities register and whether the
+ device uses the Software Ready handshake bit.
+
+* ``port_status`` (read-only):
+ Displays the Margining Port Status register, indicating Margining Ready and
+ SW Ready states.
+
+* ``enable`` (read-write):
+ Enables (``1``) or disables (``0``) Lane Margining on the device.
+ Enabling margining locks the link into D0, prevents runtime PM suspend,
+ disables ASPM L0s/L1, and verifies that the link is operating at >= 16.0 GT/s.
+ Disabling margining restores ASPM and runtime PM, and returns all lanes to
+ nominal (demargined) state.
+
+Lane-Level Attributes
+---------------------
+
+For each physical lane (``lane0``, ``lane1``, ...):
+
+* ``receiver`` (read-write):
+ Gets or sets the active target receiver number (``0`` for local receiver,
+ ``1..6`` for retimers). Switching receivers automatically demargins previous
+ offsets per PCIe single-receiver margining requirements.
+
+* ``caps`` (read-only):
+ Reports the target receiver's margining capabilities:
+ - Margining uses Driver Software (vs hardware autonomous)
+ - Independent Left/Right Timing Margining support
+ - Independent Up/Down Voltage Margining support
+ - Error Sampler vs Main Sampler
+ - Sample Multiple Receivers support
+
+* ``num_timing_steps`` (read-only):
+ Maximum timing margin steps supported by the receiver (0..63).
+
+* ``num_voltage_steps`` (read-only):
+ Maximum voltage margin steps supported by the receiver (0..127).
+
+* ``margin_timing`` (read-write):
+ Applies timing margin step offset (+/-). Writing ``0`` clears timing margin
+ back to nominal.
+
+* ``margin_voltage`` (read-write):
+ Applies voltage margin step offset (+/-). Writing ``0`` clears voltage margin
+ back to nominal.
+
+Manual Margining Example
+========================
+
+1. Inspect device capabilities and status:
+
+.. code-block:: sh
+
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/capabilities
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/port_status
+
+2. Enable Lane Margining mode:
+
+.. code-block:: sh
+
+ echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
+
+3. Configure target receiver and inspect step limits on lane 0:
+
+.. code-block:: sh
+
+ echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/receiver
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/caps
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_timing_steps
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_voltage_steps
+
+4. Apply timing and voltage margin steps:
+
+.. code-block:: sh
+
+ # Step timing margin +2 steps
+ echo 2 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
+
+ # Step voltage margin +1 step
+ echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
+
+5. Reset margins back to nominal:
+
+.. code-block:: sh
+
+ echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
+ echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
+
+6. Disable Lane Margining when complete:
+
+.. code-block:: sh
+
+ echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
+
+Automated Testing via Kselftest
+===============================
+
+The kernel includes an automated kselftest script under
+``tools/testing/selftests/pcie_lmt/pcie_lmt.sh`` to probe, validate, and exercise
+debugfs controls across all enumerated LMR devices.
+
+Run directly as root:
+
+.. code-block:: sh
+
+ sudo ./tools/testing/selftests/pcie_lmt/pcie_lmt.sh
+
+Or run via the kselftest test harness:
+
+.. code-block:: sh
+
+ make -C tools/testing/selftests TARGETS=pcie_lmt run_tests
diff --git a/MAINTAINERS b/MAINTAINERS
index b7094a616afd..b5deaae11bfe 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -21059,6 +21059,14 @@ F: Documentation/devicetree/bindings/pci/qcom,sa8255p-pcie-ep.yaml
F: drivers/pci/controller/dwc/pcie-qcom-common.c
F: drivers/pci/controller/dwc/pcie-qcom-ep.c
+PCIE LANE MARGINING AT RECEIVER (LMR)
+M: Priyank Rathod <rathodpriyank@google.com>
+L: linux-pci@vger.kernel.org
+S: Maintained
+F: Documentation/PCI/pcie-lmr.rst
+F: drivers/pci/pcie/margin.c
+F: tools/testing/selftests/pcie_lmt/
+
PCMCIA SUBSYSTEM
M: Dominik Brodowski <linux@dominikbrodowski.net>
S: Odd Fixes
diff --git a/drivers/pci/pci-driver.c b/drivers/pci/pci-driver.c
index f36778e62ac1..17544a7023fc 100644
--- a/drivers/pci/pci-driver.c
+++ b/drivers/pci/pci-driver.c
@@ -821,6 +821,7 @@ static int pci_pm_suspend(struct device *dev)
* since Coffee Lake, to enter a lower-power PM state.
*/
pci_suspend_ptm(pci_dev);
+ pci_suspend_lmr(pci_dev);
if (pci_has_legacy_pm_support(pci_dev))
return pci_legacy_suspend(dev, PMSG_SUSPEND);
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index 77b17b13ee61..dc9724cb7b4d 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -5145,8 +5145,10 @@ int __pci_reset_function_locked(struct pci_dev *dev)
method = &pci_reset_fn_methods[m];
pci_dbg(dev, "reset via %s\n", method->name);
rc = method->reset_fn(dev, PCI_RESET_DO_RESET);
- if (!rc)
+ if (!rc) {
+ pci_reset_lmr(dev);
return 0;
+ }
pci_dbg(dev, "%s failed with %d\n", method->name, rc);
if (rc != -ENOTTY)
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index 4469e1a77f3c..6322a81f9e50 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -1023,6 +1023,18 @@ static inline void pci_no_tph(void) { }
static inline void pci_tph_init(struct pci_dev *dev) { }
#endif
+#ifdef CONFIG_PCIE_LMR
+void pci_lmr_init(struct pci_dev *dev);
+void pci_lmr_exit(struct pci_dev *dev);
+void pci_suspend_lmr(struct pci_dev *dev);
+void pci_reset_lmr(struct pci_dev *dev);
+#else
+static inline void pci_lmr_init(struct pci_dev *dev) { }
+static inline void pci_lmr_exit(struct pci_dev *dev) { }
+static inline void pci_suspend_lmr(struct pci_dev *dev) { }
+static inline void pci_reset_lmr(struct pci_dev *dev) { }
+#endif
+
#ifdef CONFIG_PCIE_PTM
void pci_ptm_init(struct pci_dev *dev);
void pci_save_ptm_state(struct pci_dev *dev);
diff --git a/drivers/pci/pcie/Kconfig b/drivers/pci/pcie/Kconfig
index 207c2deae35f..3b021ca2fe84 100644
--- a/drivers/pci/pcie/Kconfig
+++ b/drivers/pci/pcie/Kconfig
@@ -137,6 +137,18 @@ config PCIE_PTM
This is only useful if you have devices that support PTM, but it
is safe to enable even if you don't.
+config PCIE_LMR
+ bool "PCI Express Lane Margining at Receiver Support"
+ depends on DEBUG_FS
+ help
+ This enables the PCI Express Lane Margining at Receiver support.
+ Lane Margining allows software to determine the voltage and
+ timing margin of each lane on a PCIe link (16.0 GT/s and above).
+ The margining data is exposed via debugfs.
+
+ This is only useful if you have devices that support lane
+ margining, but it is safe to enable even if you don't.
+
config PCIE_EDR
bool "PCI Express Error Disconnect Recover support"
depends on PCIE_DPC && ACPI
diff --git a/drivers/pci/pcie/Makefile b/drivers/pci/pcie/Makefile
index b0b43a18c304..aac45ae0402e 100644
--- a/drivers/pci/pcie/Makefile
+++ b/drivers/pci/pcie/Makefile
@@ -13,4 +13,5 @@ obj-$(CONFIG_PCIEAER_INJECT) += aer_inject.o
obj-$(CONFIG_PCIE_PME) += pme.o
obj-$(CONFIG_PCIE_DPC) += dpc.o
obj-$(CONFIG_PCIE_PTM) += ptm.o
+obj-$(CONFIG_PCIE_LMR) += margin.o
obj-$(CONFIG_PCIE_EDR) += edr.o
diff --git a/drivers/pci/pcie/margin.c b/drivers/pci/pcie/margin.c
new file mode 100644
index 000000000000..45646b5952d6
--- /dev/null
+++ b/drivers/pci/pcie/margin.c
@@ -0,0 +1,1061 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * PCI Express Lane Margining at Receiver
+ *
+ * Copyright (C) 2026 Google LLC
+ * Author: Priyank Rathod <rathodpriyank@google.com>
+ *
+ * Lane Margining at Receiver (PCIe Base Specification r6.0, sec 8.4.4)
+ * allows system software to determine the voltage and timing margins of
+ * each physical lane on a PCIe link. The Extended Capability (ID 0x27)
+ * is available for receivers operating at 16.0 GT/s (Gen4) or higher data
+ * rates, and is mandatory for receivers operating at 64.0 GT/s (Gen6) or
+ * higher data rates.
+ *
+ * This driver implements:
+ * - Probing Extended Capability ID 0x27 and Margining Port Capabilities.
+ * - Managing ASPM L0s/L1 link states during active margining with restoration.
+ * - PCIe r6.0 NO_CMD (0x7) clearing handshake per receiver and lane.
+ * - Caching receiver capabilities & step counts to avoid DEMARGIN side-effects.
+ * - Handling Symmetric vs Independent Left/Right & Up/Down margin steps.
+ * - Runtime PM protection (D0 enforcement) during active margining.
+ * - Exposing per-device debugfs interfaces under /sys/kernel/debug/pci/.
+ */
+
+#include <linux/bitfield.h>
+#include <linux/bits.h>
+#include <linux/cleanup.h>
+#include <linux/debugfs.h>
+#include <linux/delay.h>
+#include <linux/err.h>
+#include <linux/errno.h>
+#include <linux/jiffies.h>
+#include <linux/kstrtox.h>
+#include <linux/minmax.h>
+#include <linux/mutex.h>
+#include <linux/overflow.h>
+#include <linux/pci.h>
+#include <linux/pm_runtime.h>
+#include <linux/seq_file.h>
+#include <linux/slab.h>
+#include <linux/sprintf.h>
+#include <linux/string_choices.h>
+#include <linux/types.h>
+
+#include "../pci.h"
+
+/* Margin type encodings per PCIe Base Spec r6.0 sec 8.4.4 */
+#define LMR_TYPE_DEMARGIN 0x0
+#define LMR_TYPE_REPORT_CAPS 0x1
+#define LMR_TYPE_REPORT_VOLTAGE_STEPS 0x2
+#define LMR_TYPE_REPORT_TIMING_STEPS 0x3
+#define LMR_TYPE_TIMING 0x4
+#define LMR_TYPE_VOLTAGE 0x5
+#define LMR_TYPE_NO_CMD 0x7
+
+/* LMR command timing parameters */
+#define LMR_CMD_TIMEOUT_MS 150
+#define LMR_CMD_SLEEP_MIN_US 100
+#define LMR_CMD_SLEEP_MAX_US 250
+#define LMR_ENABLE_TIMEOUT_MS 150
+#define LMR_ENABLE_SLEEP_MIN_US 1000
+#define LMR_ENABLE_SLEEP_MAX_US 2000
+
+/*
+ * LMR limits:
+ * Valid receiver numbers are 0 (local receiver) to 6 (up to 3 retimers)
+ * per PCIe Base Specification r6.0 sec 8.4.4. Receiver number 7 is reserved.
+ */
+#define LMR_MAX_LANES 32
+#define LMR_MAX_RX_NUM 6
+#define LMR_MAX_TIMING_STEP 63
+#define LMR_MAX_VOLTAGE_STEP 127
+
+/* LMR PCIe generation numbers and helper */
+#define LMR_GEN6 6
+#define LMR_GEN5 5
+#define LMR_GEN4 4
+
+#define LMR_SPEED_TO_GEN(speed) \
+ ((speed) >= PCIE_SPEED_64_0GT ? LMR_GEN6 : \
+ (speed) >= PCIE_SPEED_32_0GT ? LMR_GEN5 : \
+ LMR_GEN4)
+
+/* LMR lane register stride */
+#define LMR_LANE_REG_STRIDE 4
+
+/* LMR receivers and directions */
+#define LMR_RX_LOCAL 0
+#define LMR_STEP_DIR_INCREASE 1
+#define LMR_STEP_DIR_DECREASE 0
+
+/* LMR payload field masks per PCIe Base Spec r6.0 sec 8.4.4 */
+#define LMR_STEPS_MASK GENMASK(6, 0)
+#define LMR_TIMING_STEP_MASK GENMASK(5, 0)
+#define LMR_TIMING_DIR_MASK BIT(6)
+#define LMR_VOLTAGE_STEP_MASK GENMASK(6, 0)
+#define LMR_VOLTAGE_DIR_MASK BIT(7)
+
+/*
+ * Margining Capabilities report bit fields (PCIe Base Spec r6.0 sec 8.4.4,
+ * Table "Report Margining Capabilities Payload"):
+ * Bit 0: Margining Uses Driver Software (1 = Driver software sequence; 0 = Hardware)
+ * Bit 2: Independent Left/Right Timing Margining Supported (1 = Supported; 0 = Symmetric)
+ * Bit 3: Independent Up/Down Voltage Margining Supported (1 = Supported; 0 = Symmetric)
+ * Bit 4: Margining Error Sampler (1 = Error Sampler; 0 = Main Sampler)
+ * Bit 5: Sample Multiple Receivers (1 = Multiple receivers; 0 = Single receiver only)
+ */
+#define LMR_CAP_USES_DRIVER_SW BIT(0)
+#define LMR_CAP_IND_LEFT_RIGHT_TIMING BIT(2)
+#define LMR_CAP_IND_UP_DOWN_VOLTAGE BIT(3)
+#define LMR_CAP_ERROR_SAMPLER BIT(4)
+#define LMR_CAP_SAMPLE_MULTIPLE_RX BIT(5)
+
+/**
+ * struct pci_margin_rx_info - Cached Lane Margining receiver capabilities
+ * @caps_cached: True if receiver capabilities and step limits are cached
+ * @caps: Margining capabilities byte reported by receiver
+ * @num_timing_steps: Maximum timing margin steps supported by receiver
+ * @num_voltage_steps: Maximum voltage margin steps supported by receiver
+ */
+struct pci_margin_rx_info {
+ bool caps_cached;
+ u8 caps;
+ u8 num_timing_steps;
+ u8 num_voltage_steps;
+};
+
+/**
+ * struct pci_margin_lane - Per-lane margining state
+ * @mdev: Parent LMR margin device
+ * @lane: Physical lane index (0..num_lanes - 1)
+ * @rx: Selected target receiver number (0 = local, 1..6 = retimers)
+ * @timing_val: Current applied timing margin step offset (+/-)
+ * @voltage_val: Current applied voltage margin step offset (+/-)
+ * @rx_info: Cached receiver capabilities per receiver number
+ */
+struct pci_margin_lane {
+ struct pci_margin_dev *mdev;
+ int lane;
+ u8 rx;
+ int timing_val;
+ int voltage_val;
+ struct pci_margin_rx_info rx_info[LMR_MAX_RX_NUM + 1];
+};
+
+/**
+ * struct pci_margin_dev - PCIe Lane Margining device instance
+ * @dev: Underlying PCI device
+ * @cap: Extended capability offset (PCI_EXT_CAP_ID_LMR)
+ * @debugfs: Root debugfs dentry for this device
+ * @lock: Mutex protecting LMR hardware access, active margining enablement,
+ * target receiver selection, lane margining steps, and ASPM state
+ * @enabled: True if Lane Margining is currently enabled
+ * @aspm_saved: True if original ASPM configuration has been saved
+ * @saved_aspm: Saved ASPM control register bits for the device
+ * @saved_parent_aspm: Saved ASPM control register bits for parent bridge
+ * @num_lanes: Number of lanes on the link
+ * @lanes: Flexible array of per-lane state structures
+ */
+struct pci_margin_dev {
+ struct pci_dev *dev;
+ u16 cap;
+ struct dentry *debugfs;
+ struct mutex lock;
+ bool enabled;
+ bool aspm_saved;
+ u16 saved_aspm;
+ u16 saved_parent_aspm;
+ int num_lanes;
+ struct pci_margin_lane lanes[] __counted_by(num_lanes);
+};
+
+#if IS_ENABLED(CONFIG_DEBUG_FS)
+static DEFINE_MUTEX(pci_debugfs_root_lock);
+static struct dentry *pci_debugfs_root_dir;
+
+static struct dentry *get_pci_debugfs_root(void)
+{
+ mutex_lock(&pci_debugfs_root_lock);
+ if (!pci_debugfs_root_dir)
+ pci_debugfs_root_dir = debugfs_lookup("pci", NULL);
+ if (!pci_debugfs_root_dir)
+ pci_debugfs_root_dir = debugfs_create_dir("pci", NULL);
+ mutex_unlock(&pci_debugfs_root_lock);
+ return pci_debugfs_root_dir;
+}
+#endif
+
+/*
+ * pci_lmr_disable_aspm() - Temporarily disable ASPM L0s/L1 during active
+ * margining per PCIe Base Spec r6.0 sec 8.4.4, saving original ASPMC bits.
+ */
+static void pci_lmr_disable_aspm(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev = mdev->dev;
+ struct pci_dev *parent = pci_upstream_bridge(dev);
+ u16 ctl;
+
+ if (mdev->aspm_saved)
+ return;
+
+ if (!pcie_capability_read_word(dev, PCI_EXP_LNKCTL, &ctl)) {
+ mdev->saved_aspm = ctl & PCI_EXP_LNKCTL_ASPMC;
+ pcie_capability_clear_word(dev, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_ASPMC);
+ }
+
+ if (parent && pci_is_pcie(parent)) {
+ if (!pcie_capability_read_word(parent, PCI_EXP_LNKCTL, &ctl)) {
+ mdev->saved_parent_aspm = ctl & PCI_EXP_LNKCTL_ASPMC;
+ pcie_capability_clear_word(parent, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_ASPMC);
+ }
+ }
+ mdev->aspm_saved = true;
+}
+
+/*
+ * pci_lmr_restore_aspm() - Restore original ASPM L0s/L1 state when margining
+ * is disabled or torn down.
+ */
+static void pci_lmr_restore_aspm(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev = mdev->dev;
+ struct pci_dev *parent = pci_upstream_bridge(dev);
+
+ if (!mdev->aspm_saved)
+ return;
+
+ pcie_capability_clear_and_set_word(dev, PCI_EXP_LNKCTL,
+ PCI_EXP_LNKCTL_ASPMC,
+ mdev->saved_aspm);
+ if (parent && pci_is_pcie(parent))
+ pcie_capability_clear_and_set_word(parent, PCI_EXP_LNKCTL,
+ PCI_EXP_LNKCTL_ASPMC,
+ mdev->saved_parent_aspm);
+ mdev->aspm_saved = false;
+}
+
+static inline u8 pci_lmr_sts_payload(u16 sts)
+{
+ return FIELD_GET(PCI_LMR_LANE_STS_PAYLOAD, sts);
+}
+
+/*
+ * pci_lmr_run_cmd() - Issue LMR command to Lane Control and wait for Status.
+ * Must be called with mdev->lock held.
+ */
+static int pci_lmr_run_cmd(struct pci_margin_dev *mdev, int lane, u8 rx, u8 type,
+ u8 usage, u8 payload, u16 *status_val)
+{
+ struct pci_dev *dev;
+ u16 lmr, ctrl_offset, sts_offset;
+ u16 ctrl, sts;
+ unsigned long timeout;
+ int ret;
+
+ if (!mdev || lane < 0 || lane >= mdev->num_lanes || rx > LMR_MAX_RX_NUM)
+ return -EINVAL;
+
+ dev = mdev->dev;
+ lmr = mdev->cap;
+ ctrl_offset = lmr + PCI_LMR_LANE_CTRL + LMR_LANE_REG_STRIDE * lane;
+ sts_offset = lmr + PCI_LMR_LANE_STS + LMR_LANE_REG_STRIDE * lane;
+
+ /*
+ * Per PCIe Base Spec r6.0 sec 8.4.4, software must issue NO_CMD (0x7)
+ * targeting the specific receiver (rx) to clear MTYPE in Lane Status
+ * before issuing a subsequent command.
+ */
+ if (type != LMR_TYPE_NO_CMD) {
+ ctrl = FIELD_PREP(PCI_LMR_LANE_CTRL_RX_NUM, rx) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_MTYPE, LMR_TYPE_NO_CMD) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_USAGE, 0) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_PAYLOAD, 0);
+
+ ret = pci_write_config_word(dev, ctrl_offset, ctrl);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+
+ timeout = jiffies + msecs_to_jiffies(LMR_CMD_TIMEOUT_MS);
+ while (1) {
+ ret = pci_read_config_word(dev, sts_offset, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+ if (PCI_POSSIBLE_ERROR(sts))
+ return -ENODEV;
+ if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) == LMR_TYPE_NO_CMD &&
+ FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx)
+ break;
+ if (time_after(jiffies, timeout))
+ return -ETIMEDOUT;
+ usleep_range(LMR_CMD_SLEEP_MIN_US, LMR_CMD_SLEEP_MAX_US);
+ }
+ }
+
+ ctrl = FIELD_PREP(PCI_LMR_LANE_CTRL_RX_NUM, rx) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_MTYPE, type) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_USAGE, usage) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_PAYLOAD, payload);
+
+ ret = pci_write_config_word(dev, ctrl_offset, ctrl);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+
+ timeout = jiffies + msecs_to_jiffies(LMR_CMD_TIMEOUT_MS);
+ while (1) {
+ ret = pci_read_config_word(dev, sts_offset, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+ if (PCI_POSSIBLE_ERROR(sts))
+ return -ENODEV;
+
+ if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) == type &&
+ FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx) {
+ if (status_val)
+ *status_val = sts;
+ return 0;
+ }
+
+ if (time_after(jiffies, timeout)) {
+ /*
+ * Per PCIe Base Spec r6.0 sec 8.4.4, if receiver echoes
+ * NO_CMD (0x7) after command issuance, it indicates NAK.
+ */
+ if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) == LMR_TYPE_NO_CMD &&
+ FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx)
+ return -EOPNOTSUPP;
+ break;
+ }
+
+ usleep_range(LMR_CMD_SLEEP_MIN_US, LMR_CMD_SLEEP_MAX_US);
+ }
+
+ return -ETIMEDOUT;
+}
+
+static int pci_lmr_demargin_lane(struct pci_margin_lane *plane)
+{
+ u16 sts;
+ int ret;
+
+ if (!plane || !plane->mdev)
+ return -EINVAL;
+
+ if (plane->timing_val == 0 && plane->voltage_val == 0)
+ return 0;
+
+ ret = pci_lmr_run_cmd(plane->mdev, plane->lane, plane->rx,
+ LMR_TYPE_DEMARGIN, 0, 0, &sts);
+ if (ret)
+ return ret;
+
+ plane->timing_val = 0;
+ plane->voltage_val = 0;
+ return 0;
+}
+
+static int pci_lmr_cache_rx_info(struct pci_margin_lane *plane, u8 rx)
+{
+ struct pci_margin_rx_info *info;
+ u16 sts;
+ int ret;
+
+ if (!plane || rx > LMR_MAX_RX_NUM)
+ return -EINVAL;
+
+ info = &plane->rx_info[rx];
+
+ if (info->caps_cached)
+ return 0;
+
+ /* Issuing REPORT_CAPS aborts any active margin per PCIe spec */
+ ret = pci_lmr_demargin_lane(plane);
+ if (ret)
+ return ret;
+
+ ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
+ LMR_TYPE_REPORT_CAPS, 0, 0, &sts);
+ if (ret)
+ return ret;
+ info->caps = pci_lmr_sts_payload(sts);
+
+ ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
+ LMR_TYPE_REPORT_TIMING_STEPS, 0, 0, &sts);
+ if (ret)
+ return ret;
+ info->num_timing_steps = FIELD_GET(LMR_TIMING_STEP_MASK, pci_lmr_sts_payload(sts));
+
+ ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
+ LMR_TYPE_REPORT_VOLTAGE_STEPS, 0, 0, &sts);
+ if (ret)
+ return ret;
+ info->num_voltage_steps = FIELD_GET(LMR_VOLTAGE_STEP_MASK, pci_lmr_sts_payload(sts));
+
+ info->caps_cached = true;
+ return 0;
+}
+
+#if IS_ENABLED(CONFIG_DEBUG_FS)
+
+static int margin_caps_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_dev *mdev = s->private;
+ struct pci_dev *dev = mdev->dev;
+ u16 cap;
+ int ret;
+
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_CAP, &cap);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+
+ seq_printf(s, "Port Capabilities: %#06x\n", cap);
+ seq_printf(s, " Uses SW Ready: %s\n",
+ str_yes_no(cap & PCI_LMR_PORT_CAP_USES_SW_READY));
+ return 0;
+}
+DEFINE_SHOW_ATTRIBUTE(margin_caps);
+
+static int margin_port_status_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_dev *mdev = s->private;
+ struct pci_dev *dev = mdev->dev;
+ u16 sts;
+ int ret;
+
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+
+ seq_printf(s, "Port Status: %#06x\n", sts);
+ seq_printf(s, " Margining Ready: %s\n",
+ str_yes_no(sts & PCI_LMR_PORT_STS_MARGIN_READY));
+ seq_printf(s, " SW Ready: %s\n",
+ str_yes_no(sts & PCI_LMR_PORT_STS_SW_READY));
+ return 0;
+}
+DEFINE_SHOW_ATTRIBUTE(margin_port_status);
+
+static int margin_enable_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_dev *mdev = s->private;
+
+ guard(mutex)(&mdev->lock);
+ seq_printf(s, "%d\n", mdev->enabled);
+ return 0;
+}
+
+static void pci_lmr_disable_locked(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev;
+ int i, ret;
+ u16 sts;
+
+ if (!mdev)
+ return;
+
+ lockdep_assert_held(&mdev->lock);
+
+ if (!mdev->enabled)
+ return;
+
+ dev = mdev->dev;
+
+ for (i = 0; i < mdev->num_lanes; i++)
+ pci_lmr_demargin_lane(&mdev->lanes[i]);
+
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret == PCIBIOS_SUCCESSFUL) {
+ sts &= ~PCI_LMR_PORT_STS_SW_READY;
+ pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
+ }
+ pci_lmr_restore_aspm(mdev);
+ pm_runtime_put_sync(&dev->dev);
+ mdev->enabled = false;
+}
+
+static ssize_t margin_enable_write(struct file *file, const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ struct seq_file *s = file->private_data;
+ struct pci_margin_dev *mdev = s->private;
+ struct pci_dev *dev = mdev->dev;
+ unsigned long timeout;
+ u16 sts, cap, lnksta;
+ bool enable;
+ int ret, i;
+
+ ret = kstrtobool_from_user(user_buf, count, &enable);
+ if (ret)
+ return ret;
+
+ guard(mutex)(&mdev->lock);
+
+ if (mdev->enabled == enable)
+ return count;
+
+ if (!enable) {
+ pci_lmr_disable_locked(mdev);
+ return count;
+ }
+
+ /* Ensure device is powered (D0) before reading configuration registers */
+ ret = pm_runtime_resume_and_get(&dev->dev);
+ if (ret < 0)
+ return ret;
+
+ /*
+ * PCIe r6.0 sec 8.4.4: LMR is physically undefined below 16.0 GT/s.
+ * Even if a device supports Gen4+, if the link is currently trained
+ * and operating at Gen1..Gen3 speeds (< 16.0 GT/s), reject margining.
+ */
+ pcie_capability_read_word(dev, PCI_EXP_LNKSTA, &lnksta);
+ if ((lnksta & PCI_EXP_LNKSTA_CLS) < PCI_EXP_LNKSTA_CLS_16_0GB) {
+ ret = -EOPNOTSUPP;
+ goto err_rpm;
+ }
+
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_CAP, &cap);
+ if (ret != PCIBIOS_SUCCESSFUL) {
+ ret = pcibios_err_to_errno(ret);
+ goto err_rpm;
+ }
+
+ /* Disable ASPM L0s/L1 during margining with restoration path */
+ pci_lmr_disable_aspm(mdev);
+
+ /* Ensure link is settled in L0 mode per PCIe r6.0 sec 8.4.4 */
+ usleep_range(2000, 3000);
+
+ if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL) {
+ ret = pcibios_err_to_errno(ret);
+ goto err_aspm;
+ }
+ sts |= PCI_LMR_PORT_STS_SW_READY;
+ pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
+ }
+
+ timeout = jiffies + msecs_to_jiffies(LMR_ENABLE_TIMEOUT_MS);
+ while (1) {
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL) {
+ ret = pcibios_err_to_errno(ret);
+ goto err_sw_ready;
+ }
+ if (PCI_POSSIBLE_ERROR(sts)) {
+ ret = -ENODEV;
+ goto err_sw_ready;
+ }
+ if (sts & PCI_LMR_PORT_STS_MARGIN_READY)
+ break;
+ if (time_after(jiffies, timeout)) {
+ ret = -ETIMEDOUT;
+ goto err_sw_ready;
+ }
+ usleep_range(LMR_ENABLE_SLEEP_MIN_US, LMR_ENABLE_SLEEP_MAX_US);
+ }
+
+ /* Cache capabilities for configured receiver on all lanes */
+ for (i = 0; i < mdev->num_lanes; i++) {
+ ret = pci_lmr_cache_rx_info(&mdev->lanes[i], mdev->lanes[i].rx);
+ if (ret)
+ goto err_sw_ready;
+ }
+ mdev->enabled = true;
+ return count;
+
+err_sw_ready:
+ if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret == PCIBIOS_SUCCESSFUL) {
+ sts &= ~PCI_LMR_PORT_STS_SW_READY;
+ pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
+ }
+ }
+err_aspm:
+ pci_lmr_restore_aspm(mdev);
+err_rpm:
+ pm_runtime_put_sync(&dev->dev);
+ return ret;
+}
+
+static int margin_enable_open(struct inode *inode, struct file *file)
+{
+ return single_open(file, margin_enable_show, inode->i_private);
+}
+
+static const struct file_operations margin_enable_fops = {
+ .open = margin_enable_open,
+ .read = seq_read,
+ .write = margin_enable_write,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static int margin_lane_receiver_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_lane *plane = s->private;
+
+ guard(mutex)(&plane->mdev->lock);
+ seq_printf(s, "%d\n", plane->rx);
+ return 0;
+}
+
+static ssize_t margin_lane_receiver_write(struct file *file, const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ struct seq_file *s = file->private_data;
+ struct pci_margin_lane *plane = s->private;
+ struct pci_margin_dev *mdev = plane->mdev;
+ int ret;
+ u8 rx;
+
+ ret = kstrtou8_from_user(user_buf, count, 0, &rx);
+ if (ret)
+ return ret;
+
+ /* Valid receiver numbers are 0..6 per PCIe r6.0 sec 8.4.4; 7 is reserved */
+ if (rx > LMR_MAX_RX_NUM)
+ return -EINVAL;
+
+ guard(mutex)(&mdev->lock);
+ if (plane->rx == rx)
+ return count;
+
+ if (mdev->enabled) {
+ /* Demargin previous receiver per single-receiver spec rule */
+ ret = pci_lmr_demargin_lane(plane);
+ if (ret)
+ return ret;
+ ret = pci_lmr_cache_rx_info(plane, rx);
+ if (ret)
+ return ret;
+ }
+
+ plane->rx = rx;
+ return count;
+}
+
+static int margin_lane_receiver_open(struct inode *inode, struct file *file)
+{
+ return single_open(file, margin_lane_receiver_show, inode->i_private);
+}
+
+static const struct file_operations margin_lane_receiver_fops = {
+ .open = margin_lane_receiver_open,
+ .read = seq_read,
+ .write = margin_lane_receiver_write,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static int margin_lane_caps_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_lane *plane = s->private;
+ struct pci_margin_dev *mdev = plane->mdev;
+ struct pci_margin_rx_info *info;
+ int ret;
+ u8 val;
+
+ guard(mutex)(&mdev->lock);
+ if (!mdev->enabled)
+ return -EBUSY;
+
+ ret = pci_lmr_cache_rx_info(plane, plane->rx);
+ if (ret)
+ return ret;
+
+ info = &plane->rx_info[plane->rx];
+ val = info->caps;
+ seq_printf(s, "Lane %d Rx %d Capabilities: %#02x\n", plane->lane, plane->rx, val);
+ seq_printf(s, " Uses Driver Software: %s\n",
+ str_yes_no(val & LMR_CAP_USES_DRIVER_SW));
+ seq_printf(s, " Left/Right: %s\n",
+ (val & LMR_CAP_IND_LEFT_RIGHT_TIMING) ? "independent" : "symmetric");
+ seq_printf(s, " Up/Down: %s\n",
+ (val & LMR_CAP_IND_UP_DOWN_VOLTAGE) ? "independent" : "symmetric");
+ seq_printf(s, " Error Sampler: %s\n",
+ (val & LMR_CAP_ERROR_SAMPLER) ? "yes" : "no (main sampler)");
+ seq_printf(s, " Sample Multiple Receivers: %s\n",
+ str_yes_no(val & LMR_CAP_SAMPLE_MULTIPLE_RX));
+ return 0;
+}
+DEFINE_SHOW_ATTRIBUTE(margin_lane_caps);
+
+static int margin_lane_steps_show(struct seq_file *s, u8 type)
+{
+ struct pci_margin_lane *plane = s->private;
+ struct pci_margin_dev *mdev = plane->mdev;
+ struct pci_margin_rx_info *info;
+ int ret;
+
+ guard(mutex)(&mdev->lock);
+ if (!mdev->enabled)
+ return -EBUSY;
+
+ ret = pci_lmr_cache_rx_info(plane, plane->rx);
+ if (ret)
+ return ret;
+
+ info = &plane->rx_info[plane->rx];
+ seq_printf(s, "%d\n", (type == LMR_TYPE_VOLTAGE) ?
+ info->num_voltage_steps : info->num_timing_steps);
+ return 0;
+}
+
+static int margin_lane_timing_steps_show(struct seq_file *s, void *v)
+{
+ return margin_lane_steps_show(s, LMR_TYPE_TIMING);
+}
+DEFINE_SHOW_ATTRIBUTE(margin_lane_timing_steps);
+
+static int margin_lane_voltage_steps_show(struct seq_file *s, void *v)
+{
+ return margin_lane_steps_show(s, LMR_TYPE_VOLTAGE);
+}
+DEFINE_SHOW_ATTRIBUTE(margin_lane_voltage_steps);
+
+/*
+ * pci_lmr_check_sample_multiple_rx() - Check multi-receiver concurrency.
+ * Per PCIe Base Spec r6.0 sec 8.4.4, if bit 5 (Sample Multiple Receivers) is 0,
+ * software must not margin more than one receiver at a time across the link.
+ */
+static bool pci_lmr_check_sample_multiple_rx(struct pci_margin_dev *mdev,
+ struct pci_margin_lane *plane)
+{
+ struct pci_margin_rx_info *info = &plane->rx_info[plane->rx];
+ int i;
+
+ if (info->caps & LMR_CAP_SAMPLE_MULTIPLE_RX)
+ return true;
+
+ for (i = 0; i < mdev->num_lanes; i++) {
+ struct pci_margin_lane *other = &mdev->lanes[i];
+
+ if (i == plane->lane)
+ continue;
+ if (other->rx == plane->rx &&
+ (other->timing_val != 0 || other->voltage_val != 0))
+ return false;
+ }
+ return true;
+}
+
+static ssize_t margin_lane_step_write(struct file *file, const char __user *user_buf,
+ size_t count, u8 type)
+{
+ struct seq_file *s = file->private_data;
+ struct pci_margin_lane *plane = s->private;
+ struct pci_margin_dev *mdev = plane->mdev;
+ struct pci_margin_rx_info *info;
+ u8 step, dir, payload;
+ int max_step, val, ret;
+ u16 sts;
+ u8 caps;
+
+ ret = kstrtoint_from_user(user_buf, count, 0, &val);
+ if (ret)
+ return ret;
+
+ guard(mutex)(&mdev->lock);
+ if (!mdev->enabled)
+ return -EBUSY;
+
+ if (val == 0) {
+ ret = pci_lmr_demargin_lane(plane);
+ return ret ? ret : count;
+ }
+
+ ret = pci_lmr_cache_rx_info(plane, plane->rx);
+ if (ret)
+ return ret;
+
+ if (!pci_lmr_check_sample_multiple_rx(mdev, plane))
+ return -EBUSY;
+
+ info = &plane->rx_info[plane->rx];
+ caps = info->caps;
+
+ switch (type) {
+ case LMR_TYPE_TIMING:
+ if (val < -LMR_MAX_TIMING_STEP || val > LMR_MAX_TIMING_STEP)
+ return -EINVAL;
+ if (val < 0) {
+ if (!(caps & LMR_CAP_IND_LEFT_RIGHT_TIMING))
+ return -EINVAL;
+ step = -val;
+ dir = LMR_STEP_DIR_DECREASE;
+ } else {
+ step = val;
+ dir = LMR_STEP_DIR_INCREASE;
+ }
+ max_step = info->num_timing_steps;
+ if (step > max_step)
+ return -EINVAL;
+
+ payload = FIELD_PREP(LMR_TIMING_DIR_MASK, dir) |
+ FIELD_PREP(LMR_TIMING_STEP_MASK, step);
+ ret = pci_lmr_run_cmd(mdev, plane->lane, plane->rx,
+ LMR_TYPE_TIMING, 0, payload, &sts);
+ if (ret)
+ return ret;
+ step = FIELD_GET(LMR_TIMING_STEP_MASK, pci_lmr_sts_payload(sts));
+ plane->timing_val = (dir == LMR_STEP_DIR_DECREASE) ? -step : step;
+ break;
+
+ case LMR_TYPE_VOLTAGE:
+ if (val < -LMR_MAX_VOLTAGE_STEP || val > LMR_MAX_VOLTAGE_STEP)
+ return -EINVAL;
+ if (val < 0) {
+ if (!(caps & LMR_CAP_IND_UP_DOWN_VOLTAGE))
+ return -EINVAL;
+ step = -val;
+ dir = 0;
+ } else {
+ step = val;
+ dir = 1;
+ }
+ max_step = info->num_voltage_steps;
+ if (step > max_step)
+ return -EINVAL;
+
+ payload = FIELD_PREP(LMR_VOLTAGE_DIR_MASK, dir) |
+ FIELD_PREP(LMR_VOLTAGE_STEP_MASK, step);
+ ret = pci_lmr_run_cmd(mdev, plane->lane, plane->rx,
+ LMR_TYPE_VOLTAGE, 0, payload, &sts);
+ if (ret)
+ return ret;
+ step = FIELD_GET(LMR_VOLTAGE_STEP_MASK, pci_lmr_sts_payload(sts));
+ plane->voltage_val = (dir == 0) ? -step : step;
+ break;
+
+ default:
+ return -EINVAL;
+ }
+
+ return count;
+}
+
+static ssize_t margin_lane_timing_write(struct file *file, const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ return margin_lane_step_write(file, user_buf, count, LMR_TYPE_TIMING);
+}
+
+static int margin_lane_step_show(struct seq_file *s, u8 type)
+{
+ struct pci_margin_lane *plane = s->private;
+
+ guard(mutex)(&plane->mdev->lock);
+ seq_printf(s, "%d\n", (type == LMR_TYPE_VOLTAGE) ?
+ plane->voltage_val : plane->timing_val);
+ return 0;
+}
+
+static int margin_lane_timing_show(struct seq_file *s, void *v)
+{
+ return margin_lane_step_show(s, LMR_TYPE_TIMING);
+}
+
+static int margin_lane_timing_open(struct inode *inode, struct file *file)
+{
+ return single_open(file, margin_lane_timing_show, inode->i_private);
+}
+
+static const struct file_operations margin_lane_timing_fops = {
+ .open = margin_lane_timing_open,
+ .read = seq_read,
+ .write = margin_lane_timing_write,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static ssize_t margin_lane_voltage_write(struct file *file, const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ return margin_lane_step_write(file, user_buf, count, LMR_TYPE_VOLTAGE);
+}
+
+static int margin_lane_voltage_show(struct seq_file *s, void *v)
+{
+ return margin_lane_step_show(s, LMR_TYPE_VOLTAGE);
+}
+
+static int margin_lane_voltage_open(struct inode *inode, struct file *file)
+{
+ return single_open(file, margin_lane_voltage_show, inode->i_private);
+}
+
+static const struct file_operations margin_lane_voltage_fops = {
+ .open = margin_lane_voltage_open,
+ .read = seq_read,
+ .write = margin_lane_voltage_write,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static void pci_margin_debugfs_init(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev = mdev->dev;
+ struct dentry *parent;
+ char dirname[64];
+ int i;
+
+ parent = get_pci_debugfs_root();
+ scnprintf(dirname, sizeof(dirname), "pcie_lmr_%s", dev_name(&dev->dev));
+ mdev->debugfs = debugfs_create_dir(dirname, parent);
+
+ debugfs_create_file("capabilities", 0444, mdev->debugfs, mdev, &margin_caps_fops);
+ debugfs_create_file("port_status", 0444, mdev->debugfs, mdev, &margin_port_status_fops);
+ debugfs_create_file("enable", 0644, mdev->debugfs, mdev, &margin_enable_fops);
+
+ for (i = 0; i < mdev->num_lanes; i++) {
+ struct pci_margin_lane *plane = &mdev->lanes[i];
+ struct dentry *lane_dir;
+ char lane_name[16];
+
+ scnprintf(lane_name, sizeof(lane_name), "lane%d", i);
+ lane_dir = debugfs_create_dir(lane_name, mdev->debugfs);
+
+ debugfs_create_file("receiver", 0644, lane_dir, plane, &margin_lane_receiver_fops);
+ debugfs_create_file("caps", 0444, lane_dir, plane, &margin_lane_caps_fops);
+ debugfs_create_file("num_timing_steps", 0444, lane_dir, plane,
+ &margin_lane_timing_steps_fops);
+ debugfs_create_file("num_voltage_steps", 0444, lane_dir, plane,
+ &margin_lane_voltage_steps_fops);
+ debugfs_create_file("margin_timing", 0644, lane_dir, plane,
+ &margin_lane_timing_fops);
+ debugfs_create_file("margin_voltage", 0644, lane_dir, plane,
+ &margin_lane_voltage_fops);
+ }
+}
+
+static void pci_margin_debugfs_remove(struct pci_margin_dev *mdev)
+{
+ debugfs_remove_recursive(mdev->debugfs);
+}
+
+#else
+static inline void pci_margin_debugfs_init(struct pci_margin_dev *mdev) { }
+static inline void pci_margin_debugfs_remove(struct pci_margin_dev *mdev) { }
+#endif
+
+void pci_lmr_init(struct pci_dev *dev)
+{
+ struct pci_margin_dev *mdev;
+ enum pci_bus_speed speed;
+ u16 lmr, lnksta;
+ int num_lanes, i;
+
+ if (WARN_ON_ONCE(!dev) || !pci_is_pcie(dev))
+ return;
+
+ speed = pcie_get_speed_cap(dev);
+ if (speed < PCIE_SPEED_16_0GT || speed == PCI_SPEED_UNKNOWN)
+ return;
+
+ lmr = pci_find_ext_capability(dev, PCI_EXT_CAP_ID_LMR);
+ if (!lmr) {
+ if (speed >= PCIE_SPEED_64_0GT)
+ pci_warn(dev,
+ "Missing Lane Margining at Receiver Capability (mandatory for Gen6+)\n");
+ else
+ pci_dbg(dev,
+ "Optional Lane Margining at Receiver Capability not found\n");
+ return;
+ }
+
+ pcie_capability_read_word(dev, PCI_EXP_LNKSTA, &lnksta);
+ num_lanes = FIELD_GET(PCI_EXP_LNKSTA_NLW, lnksta);
+ if (num_lanes == 0 || num_lanes > LMR_MAX_LANES) {
+ pci_warn(dev, "Invalid link width %d for LMR\n", num_lanes);
+ return;
+ }
+
+ dev->lmr_cap = lmr;
+
+ mdev = kzalloc(struct_size(mdev, lanes, num_lanes), GFP_KERNEL);
+ if (!mdev)
+ return;
+
+ mdev->num_lanes = num_lanes;
+ mdev->dev = dev;
+ mdev->cap = lmr;
+ mutex_init(&mdev->lock);
+
+ for (i = 0; i < num_lanes; i++) {
+ mdev->lanes[i].mdev = mdev;
+ mdev->lanes[i].lane = i;
+ mdev->lanes[i].rx = LMR_RX_LOCAL;
+ }
+
+ pci_margin_debugfs_init(mdev);
+
+ dev->lmr = mdev;
+
+ pci_dbg(dev, "Lane Margining at Receiver (Gen%u) Capability detected\n",
+ LMR_SPEED_TO_GEN(speed));
+}
+
+void pci_lmr_exit(struct pci_dev *dev)
+{
+ struct pci_margin_dev *mdev = dev->lmr;
+
+ if (!dev || !mdev)
+ return;
+
+ pci_suspend_lmr(dev);
+
+ pci_margin_debugfs_remove(mdev);
+ mutex_destroy(&mdev->lock);
+ kfree(mdev);
+ dev->lmr = NULL;
+}
+
+void pci_suspend_lmr(struct pci_dev *dev)
+{
+ struct pci_margin_dev *mdev = dev->lmr;
+
+ if (!dev || !mdev)
+ return;
+
+ guard(mutex)(&mdev->lock);
+ pci_lmr_disable_locked(mdev);
+}
+
+static void pci_lmr_reset_software_state_locked(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev;
+ int i;
+
+ if (!mdev)
+ return;
+
+ lockdep_assert_held(&mdev->lock);
+
+ if (!mdev->enabled)
+ return;
+
+ dev = mdev->dev;
+
+ for (i = 0; i < mdev->num_lanes; i++) {
+ mdev->lanes[i].timing_val = 0;
+ mdev->lanes[i].voltage_val = 0;
+ }
+
+ pci_lmr_restore_aspm(mdev);
+ pm_runtime_put_sync(&dev->dev);
+ mdev->enabled = false;
+}
+
+void pci_reset_lmr(struct pci_dev *dev)
+{
+ struct pci_margin_dev *mdev = dev->lmr;
+
+ if (!dev || !mdev)
+ return;
+
+ guard(mutex)(&mdev->lock);
+ pci_lmr_reset_software_state_locked(mdev);
+}
diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c
index dd0abbc63e18..352b95568ebf 100644
--- a/drivers/pci/probe.c
+++ b/drivers/pci/probe.c
@@ -2666,6 +2666,7 @@ static void pci_init_capabilities(struct pci_dev *dev)
pci_pasid_init(dev); /* Process Address Space ID */
pci_acs_init(dev); /* Access Control Services */
pci_ptm_init(dev); /* Precision Time Measurement */
+ pci_lmr_init(dev); /* Lane Margining at Receiver */
pci_aer_init(dev); /* Advanced Error Reporting */
pci_dpc_init(dev); /* Downstream Port Containment */
pci_rcec_init(dev); /* Root Complex Event Collector */
diff --git a/drivers/pci/remove.c b/drivers/pci/remove.c
index d8bffa21498a..6fba29040e44 100644
--- a/drivers/pci/remove.c
+++ b/drivers/pci/remove.c
@@ -36,6 +36,7 @@ static void pci_destroy_dev(struct pci_dev *dev)
pci_doe_sysfs_teardown(dev);
pci_npem_remove(dev);
+ pci_lmr_exit(dev);
/*
* While device is in D0 drop the device from TSM link operations
diff --git a/include/linux/pci.h b/include/linux/pci.h
index 64b308b6e61c..ef1275f4c5b6 100644
--- a/include/linux/pci.h
+++ b/include/linux/pci.h
@@ -349,6 +349,8 @@ struct rcec_ea;
* number resources to allow for hierarchy expansion.
* @is_pciehp: PCIe Hot-Plug Capable bridge.
*/
+struct pci_margin_dev;
+
struct pci_dev {
struct list_head bus_list; /* Node in per-bus list */
struct pci_bus *bus; /* Bus this device is on */
@@ -528,6 +530,10 @@ struct pci_dev {
atomic_t ptm_enable_cnt;
u8 ptm_granularity;
#endif
+#ifdef CONFIG_PCIE_LMR
+ u16 lmr_cap; /* Lane Margining Capability */
+ struct pci_margin_dev *lmr;
+#endif
#ifdef CONFIG_PCI_MSI
void __iomem *msix_base;
raw_spinlock_t msi_lock;
diff --git a/include/uapi/linux/pci_regs.h b/include/uapi/linux/pci_regs.h
index facaa324bd86..90cbe310e62f 100644
--- a/include/uapi/linux/pci_regs.h
+++ b/include/uapi/linux/pci_regs.h
@@ -757,6 +757,7 @@
#define PCI_EXT_CAP_ID_VF_REBAR 0x24 /* VF Resizable BAR */
#define PCI_EXT_CAP_ID_DLF 0x25 /* Data Link Feature */
#define PCI_EXT_CAP_ID_PL_16GT 0x26 /* Physical Layer 16.0 GT/s */
+#define PCI_EXT_CAP_ID_LMR 0x27 /* Lane Margining at Receiver */
#define PCI_EXT_CAP_ID_NPEM 0x29 /* Native PCIe Enclosure Management */
#define PCI_EXT_CAP_ID_PL_32GT 0x2A /* Physical Layer 32.0 GT/s */
#define PCI_EXT_CAP_ID_DOE 0x2E /* Data Object Exchange */
@@ -1181,6 +1182,23 @@
#define PCI_PL_16GT_LE_CTRL_USP_TX_PRESET_MASK 0x000000F0
#define PCI_PL_16GT_LE_CTRL_USP_TX_PRESET_SHIFT 4
+/* Lane Margining at Receiver */
+#define PCI_LMR_PORT_CAP 0x04 /* Margining Port Capabilities */
+#define PCI_LMR_PORT_CAP_USES_SW_READY 0x0001 /* Margining Uses Software Ready */
+#define PCI_LMR_PORT_STS 0x06 /* Margining Port Status */
+#define PCI_LMR_PORT_STS_MARGIN_READY 0x0001 /* Margining Ready */
+#define PCI_LMR_PORT_STS_SW_READY 0x0002 /* Margining SW Ready */
+#define PCI_LMR_LANE_CTRL 0x08 /* Margining Lane Control */
+#define PCI_LMR_LANE_CTRL_RX_NUM 0x0007 /* Receiver Number */
+#define PCI_LMR_LANE_CTRL_MTYPE 0x0038 /* Margining Type */
+#define PCI_LMR_LANE_CTRL_USAGE 0x0040 /* Margining Usage Model */
+#define PCI_LMR_LANE_CTRL_PAYLOAD 0xFF00 /* Margining Payload */
+#define PCI_LMR_LANE_STS 0x0A /* Margining Lane Status */
+#define PCI_LMR_LANE_STS_RX_NUM 0x0007 /* Receiver Number */
+#define PCI_LMR_LANE_STS_MTYPE 0x0038 /* Margining Type */
+#define PCI_LMR_LANE_STS_USAGE 0x0040 /* Margining Usage Model */
+#define PCI_LMR_LANE_STS_PAYLOAD 0xFF00 /* Margining Payload */
+
/* Physical Layer 32.0 GT/s */
#define PCI_PL_32GT_LE_CTRL 0x20 /* Lane Equalization Control Register */
diff --git a/tools/testing/selftests/Makefile b/tools/testing/selftests/Makefile
index 8a4b6ddc68df..6990d999388a 100644
--- a/tools/testing/selftests/Makefile
+++ b/tools/testing/selftests/Makefile
@@ -91,6 +91,7 @@ TARGETS += net/tcp_ao
TARGETS += nolibc
TARGETS += pci_endpoint
TARGETS += pcie_bwctrl
+TARGETS += pcie_lmt
TARGETS += perf_events
TARGETS += pidfd
TARGETS += pid_namespace
diff --git a/tools/testing/selftests/pcie_lmt/Makefile b/tools/testing/selftests/pcie_lmt/Makefile
new file mode 100644
index 000000000000..36ac85937d78
--- /dev/null
+++ b/tools/testing/selftests/pcie_lmt/Makefile
@@ -0,0 +1,3 @@
+# SPDX-License-Identifier: GPL-2.0
+TEST_PROGS = pcie_lmt.sh
+include ../lib.mk
diff --git a/tools/testing/selftests/pcie_lmt/pcie_lmt.sh b/tools/testing/selftests/pcie_lmt/pcie_lmt.sh
new file mode 100755
index 000000000000..22c00c2b8956
--- /dev/null
+++ b/tools/testing/selftests/pcie_lmt/pcie_lmt.sh
@@ -0,0 +1,105 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+#
+# Copyright (C) 2026 Google LLC
+# Author: Priyank Rathod <rathodpriyank@google.com>
+#
+# Kselftest for PCIe Lane Margining at Receiver (LMR / LMT)
+# Tests the debugfs interface exposed by drivers/pci/pcie/margin.c
+# (/sys/kernel/debug/pci/pcie_lmr_<pci_dev_name>/)
+
+set -e
+
+TESTNAME="pcie_lmt"
+
+# Kselftest framework requirement - SKIP code is 4.
+ksft_skip=4
+retval=0
+skipmsg="skip all tests:"
+
+if [ $UID != 0 ]; then
+ echo "$skipmsg must be run as root" >&2
+ exit $ksft_skip
+fi
+
+DEBUGFS=$(mount -t debugfs | head -1 | awk '{ print $3 }')
+if [ -z "$DEBUGFS" ]; then
+ if [ -d "/sys/kernel/debug" ]; then
+ DEBUGFS="/sys/kernel/debug"
+ else
+ echo "$skipmsg debugfs is not mounted" >&2
+ exit $ksft_skip
+ fi
+fi
+
+if [ ! -d "$DEBUGFS/pci" ]; then
+ # Allow searching debugfs root or pci directory
+ :
+fi
+
+LMR_DEVS=$(ls -d $DEBUGFS/pci/pcie_lmr_* $DEBUGFS/pcie_lmr_* 2>/dev/null || true)
+if [ -z "$LMR_DEVS" ]; then
+ echo "$skipmsg no PCIe LMR devices found in $DEBUGFS/" >&2
+ exit $ksft_skip
+fi
+
+cleanup_dev()
+{
+ local dev="$1"
+ echo 0 > "$dev/enable" 2>/dev/null || true
+}
+
+echo "$TESTNAME: testing PCIe LMR debugfs entries"
+
+for dev in $LMR_DEVS; do
+ dev_name=$(basename "$dev")
+ echo "$TESTNAME: probing device $dev_name"
+
+ if [ ! -r "$dev/capabilities" ] || [ ! -r "$dev/port_status" ] ||
+ [ ! -r "$dev/enable" ] || [ ! -w "$dev/enable" ]; then
+ echo "$TESTNAME: $dev_name missing mandatory root attributes"
+ retval=1
+ continue
+ fi
+
+ caps=$(cat "$dev/capabilities")
+ status=$(cat "$dev/port_status")
+ echo " $dev_name: capabilities read OK"
+ echo " $dev_name: port_status read OK"
+
+ trap 'cleanup_dev "$dev"' EXIT
+
+ if ! echo 1 > "$dev/enable" 2>/dev/null; then
+ echo " $dev_name: margining not ready by hardware (skipping active lanes)"
+ continue
+ fi
+
+ echo " $dev_name: margining enabled OK"
+
+ for lane_dir in $(ls -d "$dev"/lane* 2>/dev/null || true); do
+ lane=$(basename "$lane_dir")
+ echo " $dev_name: testing $lane"
+
+ # Test setting receiver (Rx 0 is always local receiver)
+ echo 0 > "$lane_dir/receiver"
+ cat "$lane_dir/caps" > /dev/null
+ cat "$lane_dir/num_timing_steps" > /dev/null
+ cat "$lane_dir/num_voltage_steps" > /dev/null
+
+ # Test resetting timing and voltage margin
+ echo 0 > "$lane_dir/margin_timing"
+ echo 0 > "$lane_dir/margin_voltage"
+ done
+
+ echo 0 > "$dev/enable"
+ trap - EXIT
+ echo " $dev_name: margining disabled OK"
+done
+
+if [ $retval -eq 0 ]; then
+ echo "$TESTNAME [PASS]"
+else
+ echo "$TESTNAME [FAIL]"
+fi
+
+exit $retval
---
base-commit: 0f23d56f17fdfc7db69d51f64c8b91bbab947aa9
change-id: 20260818-pcie-lmt-3044d586aaec
Best regards,
--
Priyank Rathod <rathodpriyank@google.com>
^ permalink raw reply related [flat|nested] 5+ messages in thread
* [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
@ 2026-08-24 22:38 Priyank Rathod
2026-08-24 22:51 ` sashiko-bot
2026-08-25 11:02 ` Ilpo Järvinen
0 siblings, 2 replies; 5+ messages in thread
From: Priyank Rathod @ 2026-08-24 22:38 UTC (permalink / raw)
To: Bjorn Helgaas, Priyank Rathod, Shuah Khan, Kees Cook,
Gustavo A. R. Silva, Jonathan Corbet, Shuah Khan
Cc: linux-kernel, linux-pci, linux-kselftest, linux-hardening,
linux-doc, Ilpo Järvinen
Per PCIe Base Specification r6.0, sec 8.4.4 ("Lane Margining at
Receiver"), PCIe devices operating at 16.0 GT/s (Gen 4) or higher data
rates support the Lane Margining at Receiver Extended Capability
(ID 0x27), and it is mandatory for receivers operating at 64.0 GT/s
(Gen 6) or higher data rates. Lane Margining allows software to
evaluate high-speed link margins by measuring timing and voltage steps
for each individual physical lane and receiver.
Add driver and debugfs support for PCIe Lane Margining at Receiver:
- Add Lane Margining at Receiver Extended Capability register
definitions (PCI_EXT_CAP_ID_LMR, PCI_LMR_PORT_CAP, PCI_LMR_PORT_STS,
PCI_LMR_LANE_CTRL, PCI_LMR_LANE_STS) to <uapi/linux/pci_regs.h>.
- Add Kconfig option CONFIG_PCIE_LMR (under drivers/pci/pcie/Kconfig)
dependent on DEBUG_FS.
- Implement drivers/pci/pcie/margin.c to probe the capability on Gen4+
links and expose per-device debugfs entries under:
/sys/kernel/debug/pci/pcie_lmr_<pci_dev_name>/
providing control over margining enablement, receiver selection, and
execution of timing/voltage margin step commands. Distinguish
between missing mandatory LMR capability on Gen6+ vs optional on
Gen4/Gen5.
- Hook pci_lmr_init() into pci_init_capabilities() during device probe
in drivers/pci/probe.c and pci_lmr_exit() into drivers/pci/remove.c.
- Add kselftest script under tools/testing/selftests/pcie_lmt/pcie_lmt.sh
to test debugfs capability reads, enablement, and stepping.
- Add MAINTAINERS entry for PCIe Lane Margining at Receiver (LMR).
Signed-off-by: Priyank Rathod <rathodpriyank@google.com>
---
Changes in v6:
- Added kernel documentation under Documentation/PCI/pcie-lmr.rst and indexed in Documentation/PCI/index.rst (Ilpo Järvinen).
- Updated MAINTAINERS with Documentation/PCI/pcie-lmr.rst (Ilpo Järvinen).
- Aligned capability bit naming and comments with PCIe Base Specification r6.0 sec 8.4.4 Table "Report Margining Capabilities Payload" (Ilpo Järvinen).
- Clarified Sample Multiple Receivers concurrency verification and rules across physical lanes in kerneldoc and documentation (Ilpo Järvinen).
- Refactored pci_lmr_run_cmd() to pass struct pci_margin_dev *mdev directly, eliminating redundant NULL checks and using mdev->num_lanes (Ilpo Järvinen).
- Converted PCI config read/write return checking across all helpers to pcibios_err_to_errno() (Ilpo Järvinen).
- Reversed return logic in pci_lmr_demargin_lane() to return early on error (Ilpo Järvinen).
- Refactored margin_lane_step_write() to eliminate bool is_voltage parameter, using command type (LMR_TYPE_TIMING / LMR_TYPE_VOLTAGE) and switch/case with consolidated bounds checks (Ilpo Järvinen).
- Renamed __pci_suspend_lmr_locked() to pci_lmr_disable_locked() to avoid PM terminology confusion and added lockdep_assert_held(&mdev->lock) (Ilpo Järvinen).
- Replaced -EACCES with -EBUSY across debugfs show/write callbacks when margining is inactive (Ilpo Järvinen).
- Clarified comment for active operating link speed check (Gen4+ capability vs dynamically operating speed) in margin_enable_write() (Ilpo Järvinen).
- Added WARN_ON_ONCE(!dev) check in pci_lmr_init() (Ilpo Järvinen).
- Demoted capability detection log message from pci_info to pci_dbg to prevent boot log noise (Ilpo Järvinen).
- Fixed timing step mask extraction in pci_lmr_cache_rx_info() to use 6-bit LMR_TIMING_STEP_MASK (sashiko-bot).
- Resumed runtime PM via pm_runtime_resume_and_get() before performing config space reads in margin_enable_write() (sashiko-bot).
- Switched to pm_runtime_put_sync() during margining teardown (sashiko-bot).
- Link to v5: https://lore.kernel.org/r/20260820-pcie-lmt-v5-1-943b3b0e18bf@google.com
PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
Per PCIe Base Specification r6.0, section 8.4.4 ("Lane Margining at Receiver"),
PCIe devices operating at 16.0 GT/s (Gen 4) or higher data rates support the
Lane Margining at Receiver Extended Capability (ID 0x27), and it is mandatory
for receivers operating at 64.0 GT/s (Gen 6) or higher data rates.
Lane Margining allows system software to evaluate high-speed link signal
integrity and margins by measuring timing and voltage steps for each physical
lane and receiver independently.
This series introduces kernel driver support, debugfs controls, and a
kselftest automation script for PCIe Lane Margining at Receiver (LMR/LMT).
==============================================================================
1. How to Enable & Configure
==============================================================================
Enable the Kconfig option under PCI support:
CONFIG_PCIE_LMR=y (or =m)
(Depends on CONFIG_PCI and CONFIG_DEBUG_FS)
Upon boot or device hotplug on Gen4+ links (>= 16.0 GT/s), the driver probes
Extended Capability ID 0x27 and exposes per-device debugfs interfaces:
/sys/kernel/debug/pci/pcie_lmr_<domain>:<bus>:<dev>.<func>/
==============================================================================
2. How to Use the Debugfs Interface (Manual Margining)
==============================================================================
Inspect device-wide margining capabilities and port status:
# Inspect root device LMR capabilities & status
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/capabilities
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/port_status
Enable active Lane Margining on the device:
# Enable Lane Margining state machine
echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
Inspect and step individual lanes (e.g. lane0):
# Select target receiver (0 = local receiver, 1..6 = retimers/link partners)
echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/receiver
# Check available timing and voltage steps for this receiver
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/caps
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_timing_steps
cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_voltage_steps
# Step timing margin or voltage margin offset
echo 2 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
# Reset margin offset back to nominal (0)
echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
Disable Lane Margining when finished:
echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
==============================================================================
3. How to Run Automated Kselftests Using the Test Script
==============================================================================
An automated kselftest script is included to test capability reads, receiver
selection, and margining commands across all enumerated LMR devices:
# Run directly as root
sudo ./tools/testing/selftests/pcie_lmt/pcie_lmt.sh
Or run via the kselftest Makefile harness:
make -C tools/testing/selftests TARGETS=pcie_lmt run_tests
Sample script output on an LMR-capable device:
pcie_lmt: testing PCIe LMR debugfs entries
pcie_lmt: probing device pcie_lmr_0000:01:00.0
pcie_lmr_0000:01:00.0: capabilities read OK
pcie_lmr_0000:01:00.0: port_status read OK
pcie_lmr_0000:01:00.0: margining enabled OK
pcie_lmr_0000:01:00.0: testing lane0
pcie_lmr_0000:01:00.0: testing lane1
pcie_lmr_0000:01:00.0: margining disabled OK
pcie_lmt [PASS]
To: Bjorn Helgaas <bhelgaas@google.com>
To: Shuah Khan <shuah@kernel.org>
Cc: linux-kernel@vger.kernel.org
Cc: linux-pci@vger.kernel.org
Cc: linux-kselftest@vger.kernel.org
Cc: Ilpo Järvinen <ilpo.jarvinen@linux.intel.com>
Changes in v5:
- Sorted #include directives alphabetically and added missing includes for bits.h, bitfield.h, cleanup.h, overflow.h, and slab.h (Ilpo Järvinen).
- Converted bitmasks to GENMASK() and BIT() macros and used FIELD_PREP() and FIELD_GET() instead of manual bit shifts (Ilpo Järvinen).
- Added pci_lmr_sts_payload() helper to cleanly extract the status payload byte before applying step and capability masks (Ilpo Järvinen).
- Replaced manual mutex locking sequences with guard(mutex)(&mdev->lock) across show and write callbacks to simplify control flow (Ilpo Järvinen).
- Documented mutex lock protection scope in kerneldoc for struct pci_margin_dev (Ilpo Järvinen).
- Used standard PCI_POSSIBLE_ERROR(), str_yes_no(), and scnprintf() helpers throughout the driver (Ilpo Järvinen).
- Clarified receiver range (0..6 per PCIe r6.0 sec 8.4.4; 7 reserved) in comments and validation checks (Ilpo Järvinen).
- Deduplicated timing and voltage show/write handlers using margin_lane_steps_show() and margin_lane_step_write() (Ilpo Järvinen).
- Placed speed check immediately following pcie_get_speed_cap() and handled PCI_SPEED_UNKNOWN (Ilpo Järvinen).
- Converted lanes in struct pci_margin_dev to a flexible array member with __counted_by(num_lanes) allocated via struct_size() (Ilpo Järvinen).
Changes in v4:
- Added Sample Multiple Receivers (Bit 5) concurrency verification in margin_lane_timing_write() and margin_lane_voltage_write() per PCIe r6.0 sec 8.4.4, returning -EBUSY if another lane on the same receiver is already margined when simultaneous lane margining is not supported.
- Added active operating link speed verification (PCI_EXP_LNKSTA_CLS >= 16.0 GT/s) in margin_enable_write() before enabling LMR, as LMR commands are physically undefined on links operating at Gen1/Gen2/Gen3 speeds.
- Added fast-path hardware NAK detection in pci_lmr_run_cmd() to return -EOPNOTSUPP immediately if a receiver echoes MTYPE == NO_CMD (0x7) after command issuance rather than waiting 150ms for a timeout.
- Added pci_reset_lmr() hooked into __pci_reset_function_locked() to synchronize software state and demargin on FLR or Secondary Bus Reset.
- Comprehensive NULL pointer checks and array/lane/receiver bounds checks added across all internal helpers and debugfs write handlers.
- Added MAINTAINERS entry for PCIe Lane Margining at Receiver (LMR).
Changes in v2:
- Fixed NO_CMD (0x7) clearing in pci_lmr_run_cmd() before issuing new commands per PCIe r6.0 sec 8.4.4.
- Protected plane->rx updates with mdev->lock in margin_lane_receiver_write().
- Corrected Margining Port Capabilities bit definition to PCI_LMR_PORT_CAP_USES_SW_READY (0x0001) in <uapi/linux/pci_regs.h>.
- Updated kselftest script (pcie_lmt.sh) to locate LMR debugfs entries.
- Validated integer bounds against LMR_MAX_TIMING_STEP / LMR_MAX_VOLTAGE_STEP before narrowing u8 cast.
- Moved mdev->enabled checks inside mutex_lock(&mdev->lock) to eliminate TOCTOU races.
- Checked return values of all pci_read_config_word() calls, propagating -EIO on failure.
- Eliminated dead store of cap in margin_enable_write().
- Explicitly checked speed == PCIE_SPEED_64_0GT in pci_lmr_init() to avoid misidentifying PCI_SPEED_UNKNOWN (0xFF) as Gen6.
---
Documentation/PCI/index.rst | 1 +
Documentation/PCI/pcie-lmr.rst | 171 +++++
MAINTAINERS | 8 +
drivers/pci/pci-driver.c | 1 +
drivers/pci/pci.c | 4 +-
drivers/pci/pci.h | 12 +
drivers/pci/pcie/Kconfig | 12 +
drivers/pci/pcie/Makefile | 1 +
drivers/pci/pcie/margin.c | 1061 ++++++++++++++++++++++++++
drivers/pci/probe.c | 1 +
drivers/pci/remove.c | 1 +
include/linux/pci.h | 6 +
include/uapi/linux/pci_regs.h | 18 +
tools/testing/selftests/Makefile | 1 +
tools/testing/selftests/pcie_lmt/Makefile | 3 +
tools/testing/selftests/pcie_lmt/pcie_lmt.sh | 105 +++
16 files changed, 1405 insertions(+), 1 deletion(-)
diff --git a/Documentation/PCI/index.rst b/Documentation/PCI/index.rst
index 5d720d2a415e..9170c98cbf3f 100644
--- a/Documentation/PCI/index.rst
+++ b/Documentation/PCI/index.rst
@@ -20,3 +20,4 @@ PCI Bus Subsystem
controller/index
boot-interrupts
tph
+ pcie-lmr
diff --git a/Documentation/PCI/pcie-lmr.rst b/Documentation/PCI/pcie-lmr.rst
new file mode 100644
index 000000000000..1ebb8317cc83
--- /dev/null
+++ b/Documentation/PCI/pcie-lmr.rst
@@ -0,0 +1,171 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+=======================================================
+PCI Express Lane Margining at Receiver (LMR) Subsystem
+=======================================================
+
+:Author: Priyank Rathod <rathodpriyank@google.com>
+:Copyright: 2026 Google LLC
+
+Overview
+========
+
+Lane Margining at Receiver (LMR), specified in the PCI Express Base
+Specification (r6.0+ sec 8.4.4), allows system software to evaluate high-speed
+link physical signal integrity and eye margins. LMR measures available timing
+(jitter/phase) and voltage margin offsets for each physical lane and receiver
+independently while the link is operating in active L0 state.
+
+Lane Margining Extended Capability (ID 0x27) is optional for links operating at
+16.0 GT/s (PCIe Gen 4) and 32.0 GT/s (Gen 5), and is mandatory for receivers
+operating at 64.0 GT/s (Gen 6) and higher.
+
+Target Receivers
+================
+
+Each physical lane can margin up to 7 distinct receivers per PCIe link:
+
+* **Receiver 0 (Local Receiver)**: The receiver in the immediate link partner.
+* **Receivers 1 to 6 (Retimers)**: Retimer pseudo-ports along the physical link
+ (up to 3 retimers, each with upstream and downstream pseudo-ports).
+* **Receiver 7**: Reserved per PCIe Base Specification.
+
+Kernel Configuration
+====================
+
+Enable the kernel configuration option under PCI support:
+
+.. code-block:: none
+
+ CONFIG_PCIE_LMR=y (or =m)
+
+Dependencies:
+* ``CONFIG_PCI``
+* ``CONFIG_DEBUG_FS``
+
+Debugfs Interface Guide
+=======================
+
+When an LMR-capable device is enumerated on a Gen4+ link, the kernel exposes
+per-device control and status files under debugfs:
+
+.. code-block:: none
+
+ /sys/kernel/debug/pci/pcie_lmr_<domain>:<bus>:<dev>.<func>/
+
+Device-Level Attributes
+-----------------------
+
+* ``capabilities`` (read-only):
+ Displays the 16-bit Margining Port Capabilities register and whether the
+ device uses the Software Ready handshake bit.
+
+* ``port_status`` (read-only):
+ Displays the Margining Port Status register, indicating Margining Ready and
+ SW Ready states.
+
+* ``enable`` (read-write):
+ Enables (``1``) or disables (``0``) Lane Margining on the device.
+ Enabling margining locks the link into D0, prevents runtime PM suspend,
+ disables ASPM L0s/L1, and verifies that the link is operating at >= 16.0 GT/s.
+ Disabling margining restores ASPM and runtime PM, and returns all lanes to
+ nominal (demargined) state.
+
+Lane-Level Attributes
+---------------------
+
+For each physical lane (``lane0``, ``lane1``, ...):
+
+* ``receiver`` (read-write):
+ Gets or sets the active target receiver number (``0`` for local receiver,
+ ``1..6`` for retimers). Switching receivers automatically demargins previous
+ offsets per PCIe single-receiver margining requirements.
+
+* ``caps`` (read-only):
+ Reports the target receiver's margining capabilities:
+ - Margining uses Driver Software (vs hardware autonomous)
+ - Independent Left/Right Timing Margining support
+ - Independent Up/Down Voltage Margining support
+ - Error Sampler vs Main Sampler
+ - Sample Multiple Receivers support
+
+* ``num_timing_steps`` (read-only):
+ Maximum timing margin steps supported by the receiver (0..63).
+
+* ``num_voltage_steps`` (read-only):
+ Maximum voltage margin steps supported by the receiver (0..127).
+
+* ``margin_timing`` (read-write):
+ Applies timing margin step offset (+/-). Writing ``0`` clears timing margin
+ back to nominal.
+
+* ``margin_voltage`` (read-write):
+ Applies voltage margin step offset (+/-). Writing ``0`` clears voltage margin
+ back to nominal.
+
+Manual Margining Example
+========================
+
+1. Inspect device capabilities and status:
+
+.. code-block:: sh
+
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/capabilities
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/port_status
+
+2. Enable Lane Margining mode:
+
+.. code-block:: sh
+
+ echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
+
+3. Configure target receiver and inspect step limits on lane 0:
+
+.. code-block:: sh
+
+ echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/receiver
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/caps
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_timing_steps
+ cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_voltage_steps
+
+4. Apply timing and voltage margin steps:
+
+.. code-block:: sh
+
+ # Step timing margin +2 steps
+ echo 2 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
+
+ # Step voltage margin +1 step
+ echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
+
+5. Reset margins back to nominal:
+
+.. code-block:: sh
+
+ echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
+ echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
+
+6. Disable Lane Margining when complete:
+
+.. code-block:: sh
+
+ echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
+
+Automated Testing via Kselftest
+===============================
+
+The kernel includes an automated kselftest script under
+``tools/testing/selftests/pcie_lmt/pcie_lmt.sh`` to probe, validate, and exercise
+debugfs controls across all enumerated LMR devices.
+
+Run directly as root:
+
+.. code-block:: sh
+
+ sudo ./tools/testing/selftests/pcie_lmt/pcie_lmt.sh
+
+Or run via the kselftest test harness:
+
+.. code-block:: sh
+
+ make -C tools/testing/selftests TARGETS=pcie_lmt run_tests
diff --git a/MAINTAINERS b/MAINTAINERS
index b7094a616afd..b5deaae11bfe 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -21059,6 +21059,14 @@ F: Documentation/devicetree/bindings/pci/qcom,sa8255p-pcie-ep.yaml
F: drivers/pci/controller/dwc/pcie-qcom-common.c
F: drivers/pci/controller/dwc/pcie-qcom-ep.c
+PCIE LANE MARGINING AT RECEIVER (LMR)
+M: Priyank Rathod <rathodpriyank@google.com>
+L: linux-pci@vger.kernel.org
+S: Maintained
+F: Documentation/PCI/pcie-lmr.rst
+F: drivers/pci/pcie/margin.c
+F: tools/testing/selftests/pcie_lmt/
+
PCMCIA SUBSYSTEM
M: Dominik Brodowski <linux@dominikbrodowski.net>
S: Odd Fixes
diff --git a/drivers/pci/pci-driver.c b/drivers/pci/pci-driver.c
index f36778e62ac1..17544a7023fc 100644
--- a/drivers/pci/pci-driver.c
+++ b/drivers/pci/pci-driver.c
@@ -821,6 +821,7 @@ static int pci_pm_suspend(struct device *dev)
* since Coffee Lake, to enter a lower-power PM state.
*/
pci_suspend_ptm(pci_dev);
+ pci_suspend_lmr(pci_dev);
if (pci_has_legacy_pm_support(pci_dev))
return pci_legacy_suspend(dev, PMSG_SUSPEND);
diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
index 77b17b13ee61..dc9724cb7b4d 100644
--- a/drivers/pci/pci.c
+++ b/drivers/pci/pci.c
@@ -5145,8 +5145,10 @@ int __pci_reset_function_locked(struct pci_dev *dev)
method = &pci_reset_fn_methods[m];
pci_dbg(dev, "reset via %s\n", method->name);
rc = method->reset_fn(dev, PCI_RESET_DO_RESET);
- if (!rc)
+ if (!rc) {
+ pci_reset_lmr(dev);
return 0;
+ }
pci_dbg(dev, "%s failed with %d\n", method->name, rc);
if (rc != -ENOTTY)
diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
index 4469e1a77f3c..6322a81f9e50 100644
--- a/drivers/pci/pci.h
+++ b/drivers/pci/pci.h
@@ -1023,6 +1023,18 @@ static inline void pci_no_tph(void) { }
static inline void pci_tph_init(struct pci_dev *dev) { }
#endif
+#ifdef CONFIG_PCIE_LMR
+void pci_lmr_init(struct pci_dev *dev);
+void pci_lmr_exit(struct pci_dev *dev);
+void pci_suspend_lmr(struct pci_dev *dev);
+void pci_reset_lmr(struct pci_dev *dev);
+#else
+static inline void pci_lmr_init(struct pci_dev *dev) { }
+static inline void pci_lmr_exit(struct pci_dev *dev) { }
+static inline void pci_suspend_lmr(struct pci_dev *dev) { }
+static inline void pci_reset_lmr(struct pci_dev *dev) { }
+#endif
+
#ifdef CONFIG_PCIE_PTM
void pci_ptm_init(struct pci_dev *dev);
void pci_save_ptm_state(struct pci_dev *dev);
diff --git a/drivers/pci/pcie/Kconfig b/drivers/pci/pcie/Kconfig
index 207c2deae35f..3b021ca2fe84 100644
--- a/drivers/pci/pcie/Kconfig
+++ b/drivers/pci/pcie/Kconfig
@@ -137,6 +137,18 @@ config PCIE_PTM
This is only useful if you have devices that support PTM, but it
is safe to enable even if you don't.
+config PCIE_LMR
+ bool "PCI Express Lane Margining at Receiver Support"
+ depends on DEBUG_FS
+ help
+ This enables the PCI Express Lane Margining at Receiver support.
+ Lane Margining allows software to determine the voltage and
+ timing margin of each lane on a PCIe link (16.0 GT/s and above).
+ The margining data is exposed via debugfs.
+
+ This is only useful if you have devices that support lane
+ margining, but it is safe to enable even if you don't.
+
config PCIE_EDR
bool "PCI Express Error Disconnect Recover support"
depends on PCIE_DPC && ACPI
diff --git a/drivers/pci/pcie/Makefile b/drivers/pci/pcie/Makefile
index b0b43a18c304..aac45ae0402e 100644
--- a/drivers/pci/pcie/Makefile
+++ b/drivers/pci/pcie/Makefile
@@ -13,4 +13,5 @@ obj-$(CONFIG_PCIEAER_INJECT) += aer_inject.o
obj-$(CONFIG_PCIE_PME) += pme.o
obj-$(CONFIG_PCIE_DPC) += dpc.o
obj-$(CONFIG_PCIE_PTM) += ptm.o
+obj-$(CONFIG_PCIE_LMR) += margin.o
obj-$(CONFIG_PCIE_EDR) += edr.o
diff --git a/drivers/pci/pcie/margin.c b/drivers/pci/pcie/margin.c
new file mode 100644
index 000000000000..45646b5952d6
--- /dev/null
+++ b/drivers/pci/pcie/margin.c
@@ -0,0 +1,1061 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * PCI Express Lane Margining at Receiver
+ *
+ * Copyright (C) 2026 Google LLC
+ * Author: Priyank Rathod <rathodpriyank@google.com>
+ *
+ * Lane Margining at Receiver (PCIe Base Specification r6.0, sec 8.4.4)
+ * allows system software to determine the voltage and timing margins of
+ * each physical lane on a PCIe link. The Extended Capability (ID 0x27)
+ * is available for receivers operating at 16.0 GT/s (Gen4) or higher data
+ * rates, and is mandatory for receivers operating at 64.0 GT/s (Gen6) or
+ * higher data rates.
+ *
+ * This driver implements:
+ * - Probing Extended Capability ID 0x27 and Margining Port Capabilities.
+ * - Managing ASPM L0s/L1 link states during active margining with restoration.
+ * - PCIe r6.0 NO_CMD (0x7) clearing handshake per receiver and lane.
+ * - Caching receiver capabilities & step counts to avoid DEMARGIN side-effects.
+ * - Handling Symmetric vs Independent Left/Right & Up/Down margin steps.
+ * - Runtime PM protection (D0 enforcement) during active margining.
+ * - Exposing per-device debugfs interfaces under /sys/kernel/debug/pci/.
+ */
+
+#include <linux/bitfield.h>
+#include <linux/bits.h>
+#include <linux/cleanup.h>
+#include <linux/debugfs.h>
+#include <linux/delay.h>
+#include <linux/err.h>
+#include <linux/errno.h>
+#include <linux/jiffies.h>
+#include <linux/kstrtox.h>
+#include <linux/minmax.h>
+#include <linux/mutex.h>
+#include <linux/overflow.h>
+#include <linux/pci.h>
+#include <linux/pm_runtime.h>
+#include <linux/seq_file.h>
+#include <linux/slab.h>
+#include <linux/sprintf.h>
+#include <linux/string_choices.h>
+#include <linux/types.h>
+
+#include "../pci.h"
+
+/* Margin type encodings per PCIe Base Spec r6.0 sec 8.4.4 */
+#define LMR_TYPE_DEMARGIN 0x0
+#define LMR_TYPE_REPORT_CAPS 0x1
+#define LMR_TYPE_REPORT_VOLTAGE_STEPS 0x2
+#define LMR_TYPE_REPORT_TIMING_STEPS 0x3
+#define LMR_TYPE_TIMING 0x4
+#define LMR_TYPE_VOLTAGE 0x5
+#define LMR_TYPE_NO_CMD 0x7
+
+/* LMR command timing parameters */
+#define LMR_CMD_TIMEOUT_MS 150
+#define LMR_CMD_SLEEP_MIN_US 100
+#define LMR_CMD_SLEEP_MAX_US 250
+#define LMR_ENABLE_TIMEOUT_MS 150
+#define LMR_ENABLE_SLEEP_MIN_US 1000
+#define LMR_ENABLE_SLEEP_MAX_US 2000
+
+/*
+ * LMR limits:
+ * Valid receiver numbers are 0 (local receiver) to 6 (up to 3 retimers)
+ * per PCIe Base Specification r6.0 sec 8.4.4. Receiver number 7 is reserved.
+ */
+#define LMR_MAX_LANES 32
+#define LMR_MAX_RX_NUM 6
+#define LMR_MAX_TIMING_STEP 63
+#define LMR_MAX_VOLTAGE_STEP 127
+
+/* LMR PCIe generation numbers and helper */
+#define LMR_GEN6 6
+#define LMR_GEN5 5
+#define LMR_GEN4 4
+
+#define LMR_SPEED_TO_GEN(speed) \
+ ((speed) >= PCIE_SPEED_64_0GT ? LMR_GEN6 : \
+ (speed) >= PCIE_SPEED_32_0GT ? LMR_GEN5 : \
+ LMR_GEN4)
+
+/* LMR lane register stride */
+#define LMR_LANE_REG_STRIDE 4
+
+/* LMR receivers and directions */
+#define LMR_RX_LOCAL 0
+#define LMR_STEP_DIR_INCREASE 1
+#define LMR_STEP_DIR_DECREASE 0
+
+/* LMR payload field masks per PCIe Base Spec r6.0 sec 8.4.4 */
+#define LMR_STEPS_MASK GENMASK(6, 0)
+#define LMR_TIMING_STEP_MASK GENMASK(5, 0)
+#define LMR_TIMING_DIR_MASK BIT(6)
+#define LMR_VOLTAGE_STEP_MASK GENMASK(6, 0)
+#define LMR_VOLTAGE_DIR_MASK BIT(7)
+
+/*
+ * Margining Capabilities report bit fields (PCIe Base Spec r6.0 sec 8.4.4,
+ * Table "Report Margining Capabilities Payload"):
+ * Bit 0: Margining Uses Driver Software (1 = Driver software sequence; 0 = Hardware)
+ * Bit 2: Independent Left/Right Timing Margining Supported (1 = Supported; 0 = Symmetric)
+ * Bit 3: Independent Up/Down Voltage Margining Supported (1 = Supported; 0 = Symmetric)
+ * Bit 4: Margining Error Sampler (1 = Error Sampler; 0 = Main Sampler)
+ * Bit 5: Sample Multiple Receivers (1 = Multiple receivers; 0 = Single receiver only)
+ */
+#define LMR_CAP_USES_DRIVER_SW BIT(0)
+#define LMR_CAP_IND_LEFT_RIGHT_TIMING BIT(2)
+#define LMR_CAP_IND_UP_DOWN_VOLTAGE BIT(3)
+#define LMR_CAP_ERROR_SAMPLER BIT(4)
+#define LMR_CAP_SAMPLE_MULTIPLE_RX BIT(5)
+
+/**
+ * struct pci_margin_rx_info - Cached Lane Margining receiver capabilities
+ * @caps_cached: True if receiver capabilities and step limits are cached
+ * @caps: Margining capabilities byte reported by receiver
+ * @num_timing_steps: Maximum timing margin steps supported by receiver
+ * @num_voltage_steps: Maximum voltage margin steps supported by receiver
+ */
+struct pci_margin_rx_info {
+ bool caps_cached;
+ u8 caps;
+ u8 num_timing_steps;
+ u8 num_voltage_steps;
+};
+
+/**
+ * struct pci_margin_lane - Per-lane margining state
+ * @mdev: Parent LMR margin device
+ * @lane: Physical lane index (0..num_lanes - 1)
+ * @rx: Selected target receiver number (0 = local, 1..6 = retimers)
+ * @timing_val: Current applied timing margin step offset (+/-)
+ * @voltage_val: Current applied voltage margin step offset (+/-)
+ * @rx_info: Cached receiver capabilities per receiver number
+ */
+struct pci_margin_lane {
+ struct pci_margin_dev *mdev;
+ int lane;
+ u8 rx;
+ int timing_val;
+ int voltage_val;
+ struct pci_margin_rx_info rx_info[LMR_MAX_RX_NUM + 1];
+};
+
+/**
+ * struct pci_margin_dev - PCIe Lane Margining device instance
+ * @dev: Underlying PCI device
+ * @cap: Extended capability offset (PCI_EXT_CAP_ID_LMR)
+ * @debugfs: Root debugfs dentry for this device
+ * @lock: Mutex protecting LMR hardware access, active margining enablement,
+ * target receiver selection, lane margining steps, and ASPM state
+ * @enabled: True if Lane Margining is currently enabled
+ * @aspm_saved: True if original ASPM configuration has been saved
+ * @saved_aspm: Saved ASPM control register bits for the device
+ * @saved_parent_aspm: Saved ASPM control register bits for parent bridge
+ * @num_lanes: Number of lanes on the link
+ * @lanes: Flexible array of per-lane state structures
+ */
+struct pci_margin_dev {
+ struct pci_dev *dev;
+ u16 cap;
+ struct dentry *debugfs;
+ struct mutex lock;
+ bool enabled;
+ bool aspm_saved;
+ u16 saved_aspm;
+ u16 saved_parent_aspm;
+ int num_lanes;
+ struct pci_margin_lane lanes[] __counted_by(num_lanes);
+};
+
+#if IS_ENABLED(CONFIG_DEBUG_FS)
+static DEFINE_MUTEX(pci_debugfs_root_lock);
+static struct dentry *pci_debugfs_root_dir;
+
+static struct dentry *get_pci_debugfs_root(void)
+{
+ mutex_lock(&pci_debugfs_root_lock);
+ if (!pci_debugfs_root_dir)
+ pci_debugfs_root_dir = debugfs_lookup("pci", NULL);
+ if (!pci_debugfs_root_dir)
+ pci_debugfs_root_dir = debugfs_create_dir("pci", NULL);
+ mutex_unlock(&pci_debugfs_root_lock);
+ return pci_debugfs_root_dir;
+}
+#endif
+
+/*
+ * pci_lmr_disable_aspm() - Temporarily disable ASPM L0s/L1 during active
+ * margining per PCIe Base Spec r6.0 sec 8.4.4, saving original ASPMC bits.
+ */
+static void pci_lmr_disable_aspm(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev = mdev->dev;
+ struct pci_dev *parent = pci_upstream_bridge(dev);
+ u16 ctl;
+
+ if (mdev->aspm_saved)
+ return;
+
+ if (!pcie_capability_read_word(dev, PCI_EXP_LNKCTL, &ctl)) {
+ mdev->saved_aspm = ctl & PCI_EXP_LNKCTL_ASPMC;
+ pcie_capability_clear_word(dev, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_ASPMC);
+ }
+
+ if (parent && pci_is_pcie(parent)) {
+ if (!pcie_capability_read_word(parent, PCI_EXP_LNKCTL, &ctl)) {
+ mdev->saved_parent_aspm = ctl & PCI_EXP_LNKCTL_ASPMC;
+ pcie_capability_clear_word(parent, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_ASPMC);
+ }
+ }
+ mdev->aspm_saved = true;
+}
+
+/*
+ * pci_lmr_restore_aspm() - Restore original ASPM L0s/L1 state when margining
+ * is disabled or torn down.
+ */
+static void pci_lmr_restore_aspm(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev = mdev->dev;
+ struct pci_dev *parent = pci_upstream_bridge(dev);
+
+ if (!mdev->aspm_saved)
+ return;
+
+ pcie_capability_clear_and_set_word(dev, PCI_EXP_LNKCTL,
+ PCI_EXP_LNKCTL_ASPMC,
+ mdev->saved_aspm);
+ if (parent && pci_is_pcie(parent))
+ pcie_capability_clear_and_set_word(parent, PCI_EXP_LNKCTL,
+ PCI_EXP_LNKCTL_ASPMC,
+ mdev->saved_parent_aspm);
+ mdev->aspm_saved = false;
+}
+
+static inline u8 pci_lmr_sts_payload(u16 sts)
+{
+ return FIELD_GET(PCI_LMR_LANE_STS_PAYLOAD, sts);
+}
+
+/*
+ * pci_lmr_run_cmd() - Issue LMR command to Lane Control and wait for Status.
+ * Must be called with mdev->lock held.
+ */
+static int pci_lmr_run_cmd(struct pci_margin_dev *mdev, int lane, u8 rx, u8 type,
+ u8 usage, u8 payload, u16 *status_val)
+{
+ struct pci_dev *dev;
+ u16 lmr, ctrl_offset, sts_offset;
+ u16 ctrl, sts;
+ unsigned long timeout;
+ int ret;
+
+ if (!mdev || lane < 0 || lane >= mdev->num_lanes || rx > LMR_MAX_RX_NUM)
+ return -EINVAL;
+
+ dev = mdev->dev;
+ lmr = mdev->cap;
+ ctrl_offset = lmr + PCI_LMR_LANE_CTRL + LMR_LANE_REG_STRIDE * lane;
+ sts_offset = lmr + PCI_LMR_LANE_STS + LMR_LANE_REG_STRIDE * lane;
+
+ /*
+ * Per PCIe Base Spec r6.0 sec 8.4.4, software must issue NO_CMD (0x7)
+ * targeting the specific receiver (rx) to clear MTYPE in Lane Status
+ * before issuing a subsequent command.
+ */
+ if (type != LMR_TYPE_NO_CMD) {
+ ctrl = FIELD_PREP(PCI_LMR_LANE_CTRL_RX_NUM, rx) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_MTYPE, LMR_TYPE_NO_CMD) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_USAGE, 0) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_PAYLOAD, 0);
+
+ ret = pci_write_config_word(dev, ctrl_offset, ctrl);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+
+ timeout = jiffies + msecs_to_jiffies(LMR_CMD_TIMEOUT_MS);
+ while (1) {
+ ret = pci_read_config_word(dev, sts_offset, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+ if (PCI_POSSIBLE_ERROR(sts))
+ return -ENODEV;
+ if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) == LMR_TYPE_NO_CMD &&
+ FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx)
+ break;
+ if (time_after(jiffies, timeout))
+ return -ETIMEDOUT;
+ usleep_range(LMR_CMD_SLEEP_MIN_US, LMR_CMD_SLEEP_MAX_US);
+ }
+ }
+
+ ctrl = FIELD_PREP(PCI_LMR_LANE_CTRL_RX_NUM, rx) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_MTYPE, type) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_USAGE, usage) |
+ FIELD_PREP(PCI_LMR_LANE_CTRL_PAYLOAD, payload);
+
+ ret = pci_write_config_word(dev, ctrl_offset, ctrl);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+
+ timeout = jiffies + msecs_to_jiffies(LMR_CMD_TIMEOUT_MS);
+ while (1) {
+ ret = pci_read_config_word(dev, sts_offset, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+ if (PCI_POSSIBLE_ERROR(sts))
+ return -ENODEV;
+
+ if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) == type &&
+ FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx) {
+ if (status_val)
+ *status_val = sts;
+ return 0;
+ }
+
+ if (time_after(jiffies, timeout)) {
+ /*
+ * Per PCIe Base Spec r6.0 sec 8.4.4, if receiver echoes
+ * NO_CMD (0x7) after command issuance, it indicates NAK.
+ */
+ if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) == LMR_TYPE_NO_CMD &&
+ FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx)
+ return -EOPNOTSUPP;
+ break;
+ }
+
+ usleep_range(LMR_CMD_SLEEP_MIN_US, LMR_CMD_SLEEP_MAX_US);
+ }
+
+ return -ETIMEDOUT;
+}
+
+static int pci_lmr_demargin_lane(struct pci_margin_lane *plane)
+{
+ u16 sts;
+ int ret;
+
+ if (!plane || !plane->mdev)
+ return -EINVAL;
+
+ if (plane->timing_val == 0 && plane->voltage_val == 0)
+ return 0;
+
+ ret = pci_lmr_run_cmd(plane->mdev, plane->lane, plane->rx,
+ LMR_TYPE_DEMARGIN, 0, 0, &sts);
+ if (ret)
+ return ret;
+
+ plane->timing_val = 0;
+ plane->voltage_val = 0;
+ return 0;
+}
+
+static int pci_lmr_cache_rx_info(struct pci_margin_lane *plane, u8 rx)
+{
+ struct pci_margin_rx_info *info;
+ u16 sts;
+ int ret;
+
+ if (!plane || rx > LMR_MAX_RX_NUM)
+ return -EINVAL;
+
+ info = &plane->rx_info[rx];
+
+ if (info->caps_cached)
+ return 0;
+
+ /* Issuing REPORT_CAPS aborts any active margin per PCIe spec */
+ ret = pci_lmr_demargin_lane(plane);
+ if (ret)
+ return ret;
+
+ ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
+ LMR_TYPE_REPORT_CAPS, 0, 0, &sts);
+ if (ret)
+ return ret;
+ info->caps = pci_lmr_sts_payload(sts);
+
+ ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
+ LMR_TYPE_REPORT_TIMING_STEPS, 0, 0, &sts);
+ if (ret)
+ return ret;
+ info->num_timing_steps = FIELD_GET(LMR_TIMING_STEP_MASK, pci_lmr_sts_payload(sts));
+
+ ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
+ LMR_TYPE_REPORT_VOLTAGE_STEPS, 0, 0, &sts);
+ if (ret)
+ return ret;
+ info->num_voltage_steps = FIELD_GET(LMR_VOLTAGE_STEP_MASK, pci_lmr_sts_payload(sts));
+
+ info->caps_cached = true;
+ return 0;
+}
+
+#if IS_ENABLED(CONFIG_DEBUG_FS)
+
+static int margin_caps_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_dev *mdev = s->private;
+ struct pci_dev *dev = mdev->dev;
+ u16 cap;
+ int ret;
+
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_CAP, &cap);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+
+ seq_printf(s, "Port Capabilities: %#06x\n", cap);
+ seq_printf(s, " Uses SW Ready: %s\n",
+ str_yes_no(cap & PCI_LMR_PORT_CAP_USES_SW_READY));
+ return 0;
+}
+DEFINE_SHOW_ATTRIBUTE(margin_caps);
+
+static int margin_port_status_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_dev *mdev = s->private;
+ struct pci_dev *dev = mdev->dev;
+ u16 sts;
+ int ret;
+
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL)
+ return pcibios_err_to_errno(ret);
+
+ seq_printf(s, "Port Status: %#06x\n", sts);
+ seq_printf(s, " Margining Ready: %s\n",
+ str_yes_no(sts & PCI_LMR_PORT_STS_MARGIN_READY));
+ seq_printf(s, " SW Ready: %s\n",
+ str_yes_no(sts & PCI_LMR_PORT_STS_SW_READY));
+ return 0;
+}
+DEFINE_SHOW_ATTRIBUTE(margin_port_status);
+
+static int margin_enable_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_dev *mdev = s->private;
+
+ guard(mutex)(&mdev->lock);
+ seq_printf(s, "%d\n", mdev->enabled);
+ return 0;
+}
+
+static void pci_lmr_disable_locked(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev;
+ int i, ret;
+ u16 sts;
+
+ if (!mdev)
+ return;
+
+ lockdep_assert_held(&mdev->lock);
+
+ if (!mdev->enabled)
+ return;
+
+ dev = mdev->dev;
+
+ for (i = 0; i < mdev->num_lanes; i++)
+ pci_lmr_demargin_lane(&mdev->lanes[i]);
+
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret == PCIBIOS_SUCCESSFUL) {
+ sts &= ~PCI_LMR_PORT_STS_SW_READY;
+ pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
+ }
+ pci_lmr_restore_aspm(mdev);
+ pm_runtime_put_sync(&dev->dev);
+ mdev->enabled = false;
+}
+
+static ssize_t margin_enable_write(struct file *file, const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ struct seq_file *s = file->private_data;
+ struct pci_margin_dev *mdev = s->private;
+ struct pci_dev *dev = mdev->dev;
+ unsigned long timeout;
+ u16 sts, cap, lnksta;
+ bool enable;
+ int ret, i;
+
+ ret = kstrtobool_from_user(user_buf, count, &enable);
+ if (ret)
+ return ret;
+
+ guard(mutex)(&mdev->lock);
+
+ if (mdev->enabled == enable)
+ return count;
+
+ if (!enable) {
+ pci_lmr_disable_locked(mdev);
+ return count;
+ }
+
+ /* Ensure device is powered (D0) before reading configuration registers */
+ ret = pm_runtime_resume_and_get(&dev->dev);
+ if (ret < 0)
+ return ret;
+
+ /*
+ * PCIe r6.0 sec 8.4.4: LMR is physically undefined below 16.0 GT/s.
+ * Even if a device supports Gen4+, if the link is currently trained
+ * and operating at Gen1..Gen3 speeds (< 16.0 GT/s), reject margining.
+ */
+ pcie_capability_read_word(dev, PCI_EXP_LNKSTA, &lnksta);
+ if ((lnksta & PCI_EXP_LNKSTA_CLS) < PCI_EXP_LNKSTA_CLS_16_0GB) {
+ ret = -EOPNOTSUPP;
+ goto err_rpm;
+ }
+
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_CAP, &cap);
+ if (ret != PCIBIOS_SUCCESSFUL) {
+ ret = pcibios_err_to_errno(ret);
+ goto err_rpm;
+ }
+
+ /* Disable ASPM L0s/L1 during margining with restoration path */
+ pci_lmr_disable_aspm(mdev);
+
+ /* Ensure link is settled in L0 mode per PCIe r6.0 sec 8.4.4 */
+ usleep_range(2000, 3000);
+
+ if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL) {
+ ret = pcibios_err_to_errno(ret);
+ goto err_aspm;
+ }
+ sts |= PCI_LMR_PORT_STS_SW_READY;
+ pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
+ }
+
+ timeout = jiffies + msecs_to_jiffies(LMR_ENABLE_TIMEOUT_MS);
+ while (1) {
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret != PCIBIOS_SUCCESSFUL) {
+ ret = pcibios_err_to_errno(ret);
+ goto err_sw_ready;
+ }
+ if (PCI_POSSIBLE_ERROR(sts)) {
+ ret = -ENODEV;
+ goto err_sw_ready;
+ }
+ if (sts & PCI_LMR_PORT_STS_MARGIN_READY)
+ break;
+ if (time_after(jiffies, timeout)) {
+ ret = -ETIMEDOUT;
+ goto err_sw_ready;
+ }
+ usleep_range(LMR_ENABLE_SLEEP_MIN_US, LMR_ENABLE_SLEEP_MAX_US);
+ }
+
+ /* Cache capabilities for configured receiver on all lanes */
+ for (i = 0; i < mdev->num_lanes; i++) {
+ ret = pci_lmr_cache_rx_info(&mdev->lanes[i], mdev->lanes[i].rx);
+ if (ret)
+ goto err_sw_ready;
+ }
+ mdev->enabled = true;
+ return count;
+
+err_sw_ready:
+ if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
+ ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
+ if (ret == PCIBIOS_SUCCESSFUL) {
+ sts &= ~PCI_LMR_PORT_STS_SW_READY;
+ pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
+ }
+ }
+err_aspm:
+ pci_lmr_restore_aspm(mdev);
+err_rpm:
+ pm_runtime_put_sync(&dev->dev);
+ return ret;
+}
+
+static int margin_enable_open(struct inode *inode, struct file *file)
+{
+ return single_open(file, margin_enable_show, inode->i_private);
+}
+
+static const struct file_operations margin_enable_fops = {
+ .open = margin_enable_open,
+ .read = seq_read,
+ .write = margin_enable_write,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static int margin_lane_receiver_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_lane *plane = s->private;
+
+ guard(mutex)(&plane->mdev->lock);
+ seq_printf(s, "%d\n", plane->rx);
+ return 0;
+}
+
+static ssize_t margin_lane_receiver_write(struct file *file, const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ struct seq_file *s = file->private_data;
+ struct pci_margin_lane *plane = s->private;
+ struct pci_margin_dev *mdev = plane->mdev;
+ int ret;
+ u8 rx;
+
+ ret = kstrtou8_from_user(user_buf, count, 0, &rx);
+ if (ret)
+ return ret;
+
+ /* Valid receiver numbers are 0..6 per PCIe r6.0 sec 8.4.4; 7 is reserved */
+ if (rx > LMR_MAX_RX_NUM)
+ return -EINVAL;
+
+ guard(mutex)(&mdev->lock);
+ if (plane->rx == rx)
+ return count;
+
+ if (mdev->enabled) {
+ /* Demargin previous receiver per single-receiver spec rule */
+ ret = pci_lmr_demargin_lane(plane);
+ if (ret)
+ return ret;
+ ret = pci_lmr_cache_rx_info(plane, rx);
+ if (ret)
+ return ret;
+ }
+
+ plane->rx = rx;
+ return count;
+}
+
+static int margin_lane_receiver_open(struct inode *inode, struct file *file)
+{
+ return single_open(file, margin_lane_receiver_show, inode->i_private);
+}
+
+static const struct file_operations margin_lane_receiver_fops = {
+ .open = margin_lane_receiver_open,
+ .read = seq_read,
+ .write = margin_lane_receiver_write,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static int margin_lane_caps_show(struct seq_file *s, void *v)
+{
+ struct pci_margin_lane *plane = s->private;
+ struct pci_margin_dev *mdev = plane->mdev;
+ struct pci_margin_rx_info *info;
+ int ret;
+ u8 val;
+
+ guard(mutex)(&mdev->lock);
+ if (!mdev->enabled)
+ return -EBUSY;
+
+ ret = pci_lmr_cache_rx_info(plane, plane->rx);
+ if (ret)
+ return ret;
+
+ info = &plane->rx_info[plane->rx];
+ val = info->caps;
+ seq_printf(s, "Lane %d Rx %d Capabilities: %#02x\n", plane->lane, plane->rx, val);
+ seq_printf(s, " Uses Driver Software: %s\n",
+ str_yes_no(val & LMR_CAP_USES_DRIVER_SW));
+ seq_printf(s, " Left/Right: %s\n",
+ (val & LMR_CAP_IND_LEFT_RIGHT_TIMING) ? "independent" : "symmetric");
+ seq_printf(s, " Up/Down: %s\n",
+ (val & LMR_CAP_IND_UP_DOWN_VOLTAGE) ? "independent" : "symmetric");
+ seq_printf(s, " Error Sampler: %s\n",
+ (val & LMR_CAP_ERROR_SAMPLER) ? "yes" : "no (main sampler)");
+ seq_printf(s, " Sample Multiple Receivers: %s\n",
+ str_yes_no(val & LMR_CAP_SAMPLE_MULTIPLE_RX));
+ return 0;
+}
+DEFINE_SHOW_ATTRIBUTE(margin_lane_caps);
+
+static int margin_lane_steps_show(struct seq_file *s, u8 type)
+{
+ struct pci_margin_lane *plane = s->private;
+ struct pci_margin_dev *mdev = plane->mdev;
+ struct pci_margin_rx_info *info;
+ int ret;
+
+ guard(mutex)(&mdev->lock);
+ if (!mdev->enabled)
+ return -EBUSY;
+
+ ret = pci_lmr_cache_rx_info(plane, plane->rx);
+ if (ret)
+ return ret;
+
+ info = &plane->rx_info[plane->rx];
+ seq_printf(s, "%d\n", (type == LMR_TYPE_VOLTAGE) ?
+ info->num_voltage_steps : info->num_timing_steps);
+ return 0;
+}
+
+static int margin_lane_timing_steps_show(struct seq_file *s, void *v)
+{
+ return margin_lane_steps_show(s, LMR_TYPE_TIMING);
+}
+DEFINE_SHOW_ATTRIBUTE(margin_lane_timing_steps);
+
+static int margin_lane_voltage_steps_show(struct seq_file *s, void *v)
+{
+ return margin_lane_steps_show(s, LMR_TYPE_VOLTAGE);
+}
+DEFINE_SHOW_ATTRIBUTE(margin_lane_voltage_steps);
+
+/*
+ * pci_lmr_check_sample_multiple_rx() - Check multi-receiver concurrency.
+ * Per PCIe Base Spec r6.0 sec 8.4.4, if bit 5 (Sample Multiple Receivers) is 0,
+ * software must not margin more than one receiver at a time across the link.
+ */
+static bool pci_lmr_check_sample_multiple_rx(struct pci_margin_dev *mdev,
+ struct pci_margin_lane *plane)
+{
+ struct pci_margin_rx_info *info = &plane->rx_info[plane->rx];
+ int i;
+
+ if (info->caps & LMR_CAP_SAMPLE_MULTIPLE_RX)
+ return true;
+
+ for (i = 0; i < mdev->num_lanes; i++) {
+ struct pci_margin_lane *other = &mdev->lanes[i];
+
+ if (i == plane->lane)
+ continue;
+ if (other->rx == plane->rx &&
+ (other->timing_val != 0 || other->voltage_val != 0))
+ return false;
+ }
+ return true;
+}
+
+static ssize_t margin_lane_step_write(struct file *file, const char __user *user_buf,
+ size_t count, u8 type)
+{
+ struct seq_file *s = file->private_data;
+ struct pci_margin_lane *plane = s->private;
+ struct pci_margin_dev *mdev = plane->mdev;
+ struct pci_margin_rx_info *info;
+ u8 step, dir, payload;
+ int max_step, val, ret;
+ u16 sts;
+ u8 caps;
+
+ ret = kstrtoint_from_user(user_buf, count, 0, &val);
+ if (ret)
+ return ret;
+
+ guard(mutex)(&mdev->lock);
+ if (!mdev->enabled)
+ return -EBUSY;
+
+ if (val == 0) {
+ ret = pci_lmr_demargin_lane(plane);
+ return ret ? ret : count;
+ }
+
+ ret = pci_lmr_cache_rx_info(plane, plane->rx);
+ if (ret)
+ return ret;
+
+ if (!pci_lmr_check_sample_multiple_rx(mdev, plane))
+ return -EBUSY;
+
+ info = &plane->rx_info[plane->rx];
+ caps = info->caps;
+
+ switch (type) {
+ case LMR_TYPE_TIMING:
+ if (val < -LMR_MAX_TIMING_STEP || val > LMR_MAX_TIMING_STEP)
+ return -EINVAL;
+ if (val < 0) {
+ if (!(caps & LMR_CAP_IND_LEFT_RIGHT_TIMING))
+ return -EINVAL;
+ step = -val;
+ dir = LMR_STEP_DIR_DECREASE;
+ } else {
+ step = val;
+ dir = LMR_STEP_DIR_INCREASE;
+ }
+ max_step = info->num_timing_steps;
+ if (step > max_step)
+ return -EINVAL;
+
+ payload = FIELD_PREP(LMR_TIMING_DIR_MASK, dir) |
+ FIELD_PREP(LMR_TIMING_STEP_MASK, step);
+ ret = pci_lmr_run_cmd(mdev, plane->lane, plane->rx,
+ LMR_TYPE_TIMING, 0, payload, &sts);
+ if (ret)
+ return ret;
+ step = FIELD_GET(LMR_TIMING_STEP_MASK, pci_lmr_sts_payload(sts));
+ plane->timing_val = (dir == LMR_STEP_DIR_DECREASE) ? -step : step;
+ break;
+
+ case LMR_TYPE_VOLTAGE:
+ if (val < -LMR_MAX_VOLTAGE_STEP || val > LMR_MAX_VOLTAGE_STEP)
+ return -EINVAL;
+ if (val < 0) {
+ if (!(caps & LMR_CAP_IND_UP_DOWN_VOLTAGE))
+ return -EINVAL;
+ step = -val;
+ dir = 0;
+ } else {
+ step = val;
+ dir = 1;
+ }
+ max_step = info->num_voltage_steps;
+ if (step > max_step)
+ return -EINVAL;
+
+ payload = FIELD_PREP(LMR_VOLTAGE_DIR_MASK, dir) |
+ FIELD_PREP(LMR_VOLTAGE_STEP_MASK, step);
+ ret = pci_lmr_run_cmd(mdev, plane->lane, plane->rx,
+ LMR_TYPE_VOLTAGE, 0, payload, &sts);
+ if (ret)
+ return ret;
+ step = FIELD_GET(LMR_VOLTAGE_STEP_MASK, pci_lmr_sts_payload(sts));
+ plane->voltage_val = (dir == 0) ? -step : step;
+ break;
+
+ default:
+ return -EINVAL;
+ }
+
+ return count;
+}
+
+static ssize_t margin_lane_timing_write(struct file *file, const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ return margin_lane_step_write(file, user_buf, count, LMR_TYPE_TIMING);
+}
+
+static int margin_lane_step_show(struct seq_file *s, u8 type)
+{
+ struct pci_margin_lane *plane = s->private;
+
+ guard(mutex)(&plane->mdev->lock);
+ seq_printf(s, "%d\n", (type == LMR_TYPE_VOLTAGE) ?
+ plane->voltage_val : plane->timing_val);
+ return 0;
+}
+
+static int margin_lane_timing_show(struct seq_file *s, void *v)
+{
+ return margin_lane_step_show(s, LMR_TYPE_TIMING);
+}
+
+static int margin_lane_timing_open(struct inode *inode, struct file *file)
+{
+ return single_open(file, margin_lane_timing_show, inode->i_private);
+}
+
+static const struct file_operations margin_lane_timing_fops = {
+ .open = margin_lane_timing_open,
+ .read = seq_read,
+ .write = margin_lane_timing_write,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static ssize_t margin_lane_voltage_write(struct file *file, const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ return margin_lane_step_write(file, user_buf, count, LMR_TYPE_VOLTAGE);
+}
+
+static int margin_lane_voltage_show(struct seq_file *s, void *v)
+{
+ return margin_lane_step_show(s, LMR_TYPE_VOLTAGE);
+}
+
+static int margin_lane_voltage_open(struct inode *inode, struct file *file)
+{
+ return single_open(file, margin_lane_voltage_show, inode->i_private);
+}
+
+static const struct file_operations margin_lane_voltage_fops = {
+ .open = margin_lane_voltage_open,
+ .read = seq_read,
+ .write = margin_lane_voltage_write,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static void pci_margin_debugfs_init(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev = mdev->dev;
+ struct dentry *parent;
+ char dirname[64];
+ int i;
+
+ parent = get_pci_debugfs_root();
+ scnprintf(dirname, sizeof(dirname), "pcie_lmr_%s", dev_name(&dev->dev));
+ mdev->debugfs = debugfs_create_dir(dirname, parent);
+
+ debugfs_create_file("capabilities", 0444, mdev->debugfs, mdev, &margin_caps_fops);
+ debugfs_create_file("port_status", 0444, mdev->debugfs, mdev, &margin_port_status_fops);
+ debugfs_create_file("enable", 0644, mdev->debugfs, mdev, &margin_enable_fops);
+
+ for (i = 0; i < mdev->num_lanes; i++) {
+ struct pci_margin_lane *plane = &mdev->lanes[i];
+ struct dentry *lane_dir;
+ char lane_name[16];
+
+ scnprintf(lane_name, sizeof(lane_name), "lane%d", i);
+ lane_dir = debugfs_create_dir(lane_name, mdev->debugfs);
+
+ debugfs_create_file("receiver", 0644, lane_dir, plane, &margin_lane_receiver_fops);
+ debugfs_create_file("caps", 0444, lane_dir, plane, &margin_lane_caps_fops);
+ debugfs_create_file("num_timing_steps", 0444, lane_dir, plane,
+ &margin_lane_timing_steps_fops);
+ debugfs_create_file("num_voltage_steps", 0444, lane_dir, plane,
+ &margin_lane_voltage_steps_fops);
+ debugfs_create_file("margin_timing", 0644, lane_dir, plane,
+ &margin_lane_timing_fops);
+ debugfs_create_file("margin_voltage", 0644, lane_dir, plane,
+ &margin_lane_voltage_fops);
+ }
+}
+
+static void pci_margin_debugfs_remove(struct pci_margin_dev *mdev)
+{
+ debugfs_remove_recursive(mdev->debugfs);
+}
+
+#else
+static inline void pci_margin_debugfs_init(struct pci_margin_dev *mdev) { }
+static inline void pci_margin_debugfs_remove(struct pci_margin_dev *mdev) { }
+#endif
+
+void pci_lmr_init(struct pci_dev *dev)
+{
+ struct pci_margin_dev *mdev;
+ enum pci_bus_speed speed;
+ u16 lmr, lnksta;
+ int num_lanes, i;
+
+ if (WARN_ON_ONCE(!dev) || !pci_is_pcie(dev))
+ return;
+
+ speed = pcie_get_speed_cap(dev);
+ if (speed < PCIE_SPEED_16_0GT || speed == PCI_SPEED_UNKNOWN)
+ return;
+
+ lmr = pci_find_ext_capability(dev, PCI_EXT_CAP_ID_LMR);
+ if (!lmr) {
+ if (speed >= PCIE_SPEED_64_0GT)
+ pci_warn(dev,
+ "Missing Lane Margining at Receiver Capability (mandatory for Gen6+)\n");
+ else
+ pci_dbg(dev,
+ "Optional Lane Margining at Receiver Capability not found\n");
+ return;
+ }
+
+ pcie_capability_read_word(dev, PCI_EXP_LNKSTA, &lnksta);
+ num_lanes = FIELD_GET(PCI_EXP_LNKSTA_NLW, lnksta);
+ if (num_lanes == 0 || num_lanes > LMR_MAX_LANES) {
+ pci_warn(dev, "Invalid link width %d for LMR\n", num_lanes);
+ return;
+ }
+
+ dev->lmr_cap = lmr;
+
+ mdev = kzalloc(struct_size(mdev, lanes, num_lanes), GFP_KERNEL);
+ if (!mdev)
+ return;
+
+ mdev->num_lanes = num_lanes;
+ mdev->dev = dev;
+ mdev->cap = lmr;
+ mutex_init(&mdev->lock);
+
+ for (i = 0; i < num_lanes; i++) {
+ mdev->lanes[i].mdev = mdev;
+ mdev->lanes[i].lane = i;
+ mdev->lanes[i].rx = LMR_RX_LOCAL;
+ }
+
+ pci_margin_debugfs_init(mdev);
+
+ dev->lmr = mdev;
+
+ pci_dbg(dev, "Lane Margining at Receiver (Gen%u) Capability detected\n",
+ LMR_SPEED_TO_GEN(speed));
+}
+
+void pci_lmr_exit(struct pci_dev *dev)
+{
+ struct pci_margin_dev *mdev = dev->lmr;
+
+ if (!dev || !mdev)
+ return;
+
+ pci_suspend_lmr(dev);
+
+ pci_margin_debugfs_remove(mdev);
+ mutex_destroy(&mdev->lock);
+ kfree(mdev);
+ dev->lmr = NULL;
+}
+
+void pci_suspend_lmr(struct pci_dev *dev)
+{
+ struct pci_margin_dev *mdev = dev->lmr;
+
+ if (!dev || !mdev)
+ return;
+
+ guard(mutex)(&mdev->lock);
+ pci_lmr_disable_locked(mdev);
+}
+
+static void pci_lmr_reset_software_state_locked(struct pci_margin_dev *mdev)
+{
+ struct pci_dev *dev;
+ int i;
+
+ if (!mdev)
+ return;
+
+ lockdep_assert_held(&mdev->lock);
+
+ if (!mdev->enabled)
+ return;
+
+ dev = mdev->dev;
+
+ for (i = 0; i < mdev->num_lanes; i++) {
+ mdev->lanes[i].timing_val = 0;
+ mdev->lanes[i].voltage_val = 0;
+ }
+
+ pci_lmr_restore_aspm(mdev);
+ pm_runtime_put_sync(&dev->dev);
+ mdev->enabled = false;
+}
+
+void pci_reset_lmr(struct pci_dev *dev)
+{
+ struct pci_margin_dev *mdev = dev->lmr;
+
+ if (!dev || !mdev)
+ return;
+
+ guard(mutex)(&mdev->lock);
+ pci_lmr_reset_software_state_locked(mdev);
+}
diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c
index dd0abbc63e18..352b95568ebf 100644
--- a/drivers/pci/probe.c
+++ b/drivers/pci/probe.c
@@ -2666,6 +2666,7 @@ static void pci_init_capabilities(struct pci_dev *dev)
pci_pasid_init(dev); /* Process Address Space ID */
pci_acs_init(dev); /* Access Control Services */
pci_ptm_init(dev); /* Precision Time Measurement */
+ pci_lmr_init(dev); /* Lane Margining at Receiver */
pci_aer_init(dev); /* Advanced Error Reporting */
pci_dpc_init(dev); /* Downstream Port Containment */
pci_rcec_init(dev); /* Root Complex Event Collector */
diff --git a/drivers/pci/remove.c b/drivers/pci/remove.c
index d8bffa21498a..6fba29040e44 100644
--- a/drivers/pci/remove.c
+++ b/drivers/pci/remove.c
@@ -36,6 +36,7 @@ static void pci_destroy_dev(struct pci_dev *dev)
pci_doe_sysfs_teardown(dev);
pci_npem_remove(dev);
+ pci_lmr_exit(dev);
/*
* While device is in D0 drop the device from TSM link operations
diff --git a/include/linux/pci.h b/include/linux/pci.h
index 64b308b6e61c..ef1275f4c5b6 100644
--- a/include/linux/pci.h
+++ b/include/linux/pci.h
@@ -349,6 +349,8 @@ struct rcec_ea;
* number resources to allow for hierarchy expansion.
* @is_pciehp: PCIe Hot-Plug Capable bridge.
*/
+struct pci_margin_dev;
+
struct pci_dev {
struct list_head bus_list; /* Node in per-bus list */
struct pci_bus *bus; /* Bus this device is on */
@@ -528,6 +530,10 @@ struct pci_dev {
atomic_t ptm_enable_cnt;
u8 ptm_granularity;
#endif
+#ifdef CONFIG_PCIE_LMR
+ u16 lmr_cap; /* Lane Margining Capability */
+ struct pci_margin_dev *lmr;
+#endif
#ifdef CONFIG_PCI_MSI
void __iomem *msix_base;
raw_spinlock_t msi_lock;
diff --git a/include/uapi/linux/pci_regs.h b/include/uapi/linux/pci_regs.h
index facaa324bd86..90cbe310e62f 100644
--- a/include/uapi/linux/pci_regs.h
+++ b/include/uapi/linux/pci_regs.h
@@ -757,6 +757,7 @@
#define PCI_EXT_CAP_ID_VF_REBAR 0x24 /* VF Resizable BAR */
#define PCI_EXT_CAP_ID_DLF 0x25 /* Data Link Feature */
#define PCI_EXT_CAP_ID_PL_16GT 0x26 /* Physical Layer 16.0 GT/s */
+#define PCI_EXT_CAP_ID_LMR 0x27 /* Lane Margining at Receiver */
#define PCI_EXT_CAP_ID_NPEM 0x29 /* Native PCIe Enclosure Management */
#define PCI_EXT_CAP_ID_PL_32GT 0x2A /* Physical Layer 32.0 GT/s */
#define PCI_EXT_CAP_ID_DOE 0x2E /* Data Object Exchange */
@@ -1181,6 +1182,23 @@
#define PCI_PL_16GT_LE_CTRL_USP_TX_PRESET_MASK 0x000000F0
#define PCI_PL_16GT_LE_CTRL_USP_TX_PRESET_SHIFT 4
+/* Lane Margining at Receiver */
+#define PCI_LMR_PORT_CAP 0x04 /* Margining Port Capabilities */
+#define PCI_LMR_PORT_CAP_USES_SW_READY 0x0001 /* Margining Uses Software Ready */
+#define PCI_LMR_PORT_STS 0x06 /* Margining Port Status */
+#define PCI_LMR_PORT_STS_MARGIN_READY 0x0001 /* Margining Ready */
+#define PCI_LMR_PORT_STS_SW_READY 0x0002 /* Margining SW Ready */
+#define PCI_LMR_LANE_CTRL 0x08 /* Margining Lane Control */
+#define PCI_LMR_LANE_CTRL_RX_NUM 0x0007 /* Receiver Number */
+#define PCI_LMR_LANE_CTRL_MTYPE 0x0038 /* Margining Type */
+#define PCI_LMR_LANE_CTRL_USAGE 0x0040 /* Margining Usage Model */
+#define PCI_LMR_LANE_CTRL_PAYLOAD 0xFF00 /* Margining Payload */
+#define PCI_LMR_LANE_STS 0x0A /* Margining Lane Status */
+#define PCI_LMR_LANE_STS_RX_NUM 0x0007 /* Receiver Number */
+#define PCI_LMR_LANE_STS_MTYPE 0x0038 /* Margining Type */
+#define PCI_LMR_LANE_STS_USAGE 0x0040 /* Margining Usage Model */
+#define PCI_LMR_LANE_STS_PAYLOAD 0xFF00 /* Margining Payload */
+
/* Physical Layer 32.0 GT/s */
#define PCI_PL_32GT_LE_CTRL 0x20 /* Lane Equalization Control Register */
diff --git a/tools/testing/selftests/Makefile b/tools/testing/selftests/Makefile
index 8a4b6ddc68df..6990d999388a 100644
--- a/tools/testing/selftests/Makefile
+++ b/tools/testing/selftests/Makefile
@@ -91,6 +91,7 @@ TARGETS += net/tcp_ao
TARGETS += nolibc
TARGETS += pci_endpoint
TARGETS += pcie_bwctrl
+TARGETS += pcie_lmt
TARGETS += perf_events
TARGETS += pidfd
TARGETS += pid_namespace
diff --git a/tools/testing/selftests/pcie_lmt/Makefile b/tools/testing/selftests/pcie_lmt/Makefile
new file mode 100644
index 000000000000..36ac85937d78
--- /dev/null
+++ b/tools/testing/selftests/pcie_lmt/Makefile
@@ -0,0 +1,3 @@
+# SPDX-License-Identifier: GPL-2.0
+TEST_PROGS = pcie_lmt.sh
+include ../lib.mk
diff --git a/tools/testing/selftests/pcie_lmt/pcie_lmt.sh b/tools/testing/selftests/pcie_lmt/pcie_lmt.sh
new file mode 100755
index 000000000000..22c00c2b8956
--- /dev/null
+++ b/tools/testing/selftests/pcie_lmt/pcie_lmt.sh
@@ -0,0 +1,105 @@
+#!/bin/bash
+# SPDX-License-Identifier: GPL-2.0
+#
+# Copyright (C) 2026 Google LLC
+# Author: Priyank Rathod <rathodpriyank@google.com>
+#
+# Kselftest for PCIe Lane Margining at Receiver (LMR / LMT)
+# Tests the debugfs interface exposed by drivers/pci/pcie/margin.c
+# (/sys/kernel/debug/pci/pcie_lmr_<pci_dev_name>/)
+
+set -e
+
+TESTNAME="pcie_lmt"
+
+# Kselftest framework requirement - SKIP code is 4.
+ksft_skip=4
+retval=0
+skipmsg="skip all tests:"
+
+if [ $UID != 0 ]; then
+ echo "$skipmsg must be run as root" >&2
+ exit $ksft_skip
+fi
+
+DEBUGFS=$(mount -t debugfs | head -1 | awk '{ print $3 }')
+if [ -z "$DEBUGFS" ]; then
+ if [ -d "/sys/kernel/debug" ]; then
+ DEBUGFS="/sys/kernel/debug"
+ else
+ echo "$skipmsg debugfs is not mounted" >&2
+ exit $ksft_skip
+ fi
+fi
+
+if [ ! -d "$DEBUGFS/pci" ]; then
+ # Allow searching debugfs root or pci directory
+ :
+fi
+
+LMR_DEVS=$(ls -d $DEBUGFS/pci/pcie_lmr_* $DEBUGFS/pcie_lmr_* 2>/dev/null || true)
+if [ -z "$LMR_DEVS" ]; then
+ echo "$skipmsg no PCIe LMR devices found in $DEBUGFS/" >&2
+ exit $ksft_skip
+fi
+
+cleanup_dev()
+{
+ local dev="$1"
+ echo 0 > "$dev/enable" 2>/dev/null || true
+}
+
+echo "$TESTNAME: testing PCIe LMR debugfs entries"
+
+for dev in $LMR_DEVS; do
+ dev_name=$(basename "$dev")
+ echo "$TESTNAME: probing device $dev_name"
+
+ if [ ! -r "$dev/capabilities" ] || [ ! -r "$dev/port_status" ] ||
+ [ ! -r "$dev/enable" ] || [ ! -w "$dev/enable" ]; then
+ echo "$TESTNAME: $dev_name missing mandatory root attributes"
+ retval=1
+ continue
+ fi
+
+ caps=$(cat "$dev/capabilities")
+ status=$(cat "$dev/port_status")
+ echo " $dev_name: capabilities read OK"
+ echo " $dev_name: port_status read OK"
+
+ trap 'cleanup_dev "$dev"' EXIT
+
+ if ! echo 1 > "$dev/enable" 2>/dev/null; then
+ echo " $dev_name: margining not ready by hardware (skipping active lanes)"
+ continue
+ fi
+
+ echo " $dev_name: margining enabled OK"
+
+ for lane_dir in $(ls -d "$dev"/lane* 2>/dev/null || true); do
+ lane=$(basename "$lane_dir")
+ echo " $dev_name: testing $lane"
+
+ # Test setting receiver (Rx 0 is always local receiver)
+ echo 0 > "$lane_dir/receiver"
+ cat "$lane_dir/caps" > /dev/null
+ cat "$lane_dir/num_timing_steps" > /dev/null
+ cat "$lane_dir/num_voltage_steps" > /dev/null
+
+ # Test resetting timing and voltage margin
+ echo 0 > "$lane_dir/margin_timing"
+ echo 0 > "$lane_dir/margin_voltage"
+ done
+
+ echo 0 > "$dev/enable"
+ trap - EXIT
+ echo " $dev_name: margining disabled OK"
+done
+
+if [ $retval -eq 0 ]; then
+ echo "$TESTNAME [PASS]"
+else
+ echo "$TESTNAME [FAIL]"
+fi
+
+exit $retval
---
base-commit: 0f23d56f17fdfc7db69d51f64c8b91bbab947aa9
change-id: 20260818-pcie-lmt-3044d586aaec
Best regards,
--
Priyank Rathod <rathodpriyank@google.com>
^ permalink raw reply related [flat|nested] 5+ messages in thread
* Re: [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
2026-08-24 22:37 Priyank Rathod
@ 2026-08-24 22:51 ` sashiko-bot
0 siblings, 0 replies; 5+ messages in thread
From: sashiko-bot @ 2026-08-24 22:51 UTC (permalink / raw)
To: Priyank Rathod; +Cc: linux-pci
Thank you for your contribution! Sashiko AI review found 8 potential issue(s) to consider:
- [High] Device ASPM state is permanently disabled after a PCI reset because `pci_restore_state()` clobbers the correctly restored LMR ASPM configuration.
- [High] Concurrent execution of sysfs device reset and device removal leads to a Use-After-Free of `dev->lmr` and its mutex.
- [High] Incorrect teardown ordering in `pci_lmr_exit()` allows concurrent debugfs writes to permanently leak an RPM reference and leave upstream ASPM disabled.
- [Medium] LMR software state is not suspended during hibernation, leading to a mismatch between software tracking and hardware state upon thaw/restore.
- [Medium] The driver violates PCIe specification requirements by failing to automatically issue a Demargin command before reversing margin directions.
- [Medium] The error cleanup path in `margin_enable_write()` clobbers the original negative error code, causing the function to silently return success.
- [Medium] Secondary Bus Resets (SBR) bypass device-specific reset handlers, permanently desynchronizing LMR software state and leaking the runtime PM reference.
- [Medium] The driver writes 1 to reserved direction bits when configuring devices that only support symmetric margining.
--
--- Patch [1]: [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support ---
commit 23e30598998c8cd47f12107ec01b8ff50f90586b
Author: Priyank Rathod <rathodpriyank@google.com>
Subject: PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
This patch introduces support for PCIe Lane Margining at Receiver (LMR). It
adds debugfs controls to allow manual execution of timing and voltage margin
steps on capable receivers, along with basic state management.
Link: https://lore.kernel.org/r/20260820-pcie-lmt-v5-1-943b3b0e18bf@google.com
> --- a/drivers/pci/pci-driver.c
> +++ b/drivers/pci/pci-driver.c
> @@ -821,6 +821,7 @@ static int pci_pm_suspend(struct device *dev)
> * since Coffee Lake, to enter a lower-power PM state.
> */
> pci_suspend_ptm(pci_dev);
> + pci_suspend_lmr(pci_dev);
>
> if (pci_has_legacy_pm_support(pci_dev))
> return pci_legacy_suspend(dev, PMSG_SUSPEND);
[Severity: Medium]
Does this miss hibernation (S4) support?
The LMR suspend is hooked into pci_pm_suspend() for S3 sleep, but
pci_pm_freeze() and pci_pm_poweroff() don't appear to call pci_suspend_lmr().
If the system hibernates, the device loses power and the margining state
resets, but upon resume, mdev->enabled would still be true, leaving ASPM
tracking desynchronized and a leaked runtime PM reference.
> --- a/drivers/pci/pci.c
> +++ b/drivers/pci/pci.c
> @@ -5145,8 +5145,10 @@ int __pci_reset_function_locked(struct pci_dev *dev)
> method = &pci_reset_fn_methods[m];
> pci_dbg(dev, "reset via %s\n", method->name);
> rc = method->reset_fn(dev, PCI_RESET_DO_RESET);
> - if (!rc)
> + if (!rc) {
> + pci_reset_lmr(dev);
> return 0;
> + }
>
> pci_dbg(dev, "%s failed with %d\n", method->name, rc);
[Severity: High]
Will the ASPM restoration in pci_reset_lmr() be immediately clobbered here?
The pci_reset_function() sequence calls pci_save_state() before this, and
pci_restore_state() afterwards. If pci_reset_lmr() restores the ASPM hardware
state to its original configuration, won't pci_restore_state() subsequently
blindly overwrite it with the saved config space (where ASPM was disabled by
LMR), leaving ASPM permanently disabled?
[Severity: Medium]
By hooking LMR reset cleanup exclusively here in __pci_reset_function_locked(),
do we miss Secondary Bus Resets (SBR)?
If a root user triggers a bus reset via the parent bridge, the downstream
devices are reset using a different path. The hardware LMR state would be
wiped, but pci_reset_lmr() is never invoked, leaving mdev->enabled as true
and failing to drop the runtime PM reference.
> --- a/drivers/pci/pcie/margin.c
> +++ b/drivers/pci/pcie/margin.c
[ ... ]
> @@ -557,6 +557,13 @@ static ssize_t margin_enable_write(struct file *file, const char __user *user_buf,
> mdev->enabled = true;
> return count;
>
> +err_sw_ready:
> + if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
> + if (ret == PCIBIOS_SUCCESSFUL) {
> + sts &= ~PCI_LMR_PORT_STS_SW_READY;
> + pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
> + }
> + }
> +err_aspm:
[Severity: Medium]
Does this error cleanup path accidentally clobber the original negative error
code?
If an error like -ETIMEDOUT occurs above, the jump to err_sw_ready issues a
pci_read_config_word(). If that read succeeds, ret is overwritten with
PCIBIOS_SUCCESSFUL (0). The function will then silently return 0 instead of
propagating the actual failure to userspace.
[ ... ]
> @@ -732,6 +732,10 @@ static ssize_t margin_lane_step_write(struct file *file, const char __user *user_buf,
> if (!mdev->enabled)
> return -EBUSY;
>
> + if (val == 0) {
> + ret = pci_lmr_demargin_lane(plane);
> + return ret ? ret : count;
> + }
> +
> + ret = pci_lmr_cache_rx_info(plane, plane->rx);
[Severity: Medium]
Are we missing an automatic demargin when changing margin directions?
The code issues a demargin command if the requested value is exactly 0, but
if a user sequentially writes a positive offset and then a negative offset (or
vice versa), the driver issues the new step command immediately. Per the PCIe
Base Specification, software must issue a Demargin (or Clear Margin) command
before changing margin directions to avoid undefined behavior or NAKs.
[ ... ]
> @@ -754,6 +758,11 @@ static ssize_t margin_lane_step_write(struct file *file, const char __user *user_buf,
> if (val < 0) {
> if (!(caps & LMR_CAP_IND_LEFT_RIGHT_TIMING))
> return -EINVAL;
> step = -val;
> dir = LMR_STEP_DIR_DECREASE;
> } else {
> step = val;
> dir = LMR_STEP_DIR_INCREASE;
> }
[Severity: Medium]
Does this write 1 to a reserved direction bit on symmetric-only receivers?
For positive margins, the code unconditionally assigns
dir = LMR_STEP_DIR_INCREASE (1). If the receiver doesn't support independent
left/right or up/down margining, the PCIe Base Specification designates the
direction bits as Reserved, so they should be written as 0.
[ ... ]
> @@ -983,6 +983,16 @@ void pci_lmr_exit(struct pci_dev *dev)
> {
> struct pci_margin_dev *mdev = dev->lmr;
>
> + if (!dev || !mdev)
> + return;
> +
> + pci_suspend_lmr(dev);
> +
> + pci_margin_debugfs_remove(mdev);
> + mutex_destroy(&mdev->lock);
> + kfree(mdev);
> + dev->lmr = NULL;
> }
[Severity: High]
Does this teardown ordering allow concurrent debugfs writes to leak a runtime
PM reference?
By calling pci_suspend_lmr() before pci_margin_debugfs_remove(), a concurrent
write to the enable file could be blocked on mdev->lock. Once
pci_suspend_lmr() sets enabled = false and releases the lock, the write thread
wakes up, sees enabled is false, enables LMR again, acquires a runtime PM
reference, and disables ASPM. Then pci_margin_debugfs_remove() completes and
the structure is freed, abandoning the newly created hardware and RPM state.
[ ... ]
> @@ -1029,6 +1039,11 @@ void pci_reset_lmr(struct pci_dev *dev)
> {
> struct pci_margin_dev *mdev = dev->lmr;
>
> + if (!dev || !mdev)
> + return;
> +
> + guard(mutex)(&mdev->lock);
> + pci_lmr_reset_software_state_locked(mdev);
> }
[Severity: High]
Can a concurrent device removal and sysfs reset trigger a use-after-free here?
In pci_lmr_exit(), mdev is freed and dev->lmr is set to NULL without holding
the device_lock. If a concurrent sysfs reset triggers pci_reset_function()
(which acquires device_lock) and calls pci_reset_lmr(), it might read the
dev->lmr pointer just before pci_lmr_exit() nullifies it. It would then
attempt to lock the mutex of the already freed mdev structure.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260824-pcie-lmt-v6-1-86ada768b0c8@google.com?part=1
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
2026-08-24 22:38 [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support Priyank Rathod
@ 2026-08-24 22:51 ` sashiko-bot
2026-08-25 11:02 ` Ilpo Järvinen
1 sibling, 0 replies; 5+ messages in thread
From: sashiko-bot @ 2026-08-24 22:51 UTC (permalink / raw)
To: Priyank Rathod; +Cc: linux-pci
Thank you for your contribution! Sashiko AI review found 9 potential issue(s) to consider:
- [High] Temporary ASPM disablement leaks into the permanently saved PCI config state during resets or system suspend.
- [High] Inverted logic in pci_lmr_check_sample_multiple_rx() allows invalid concurrent margining.
- [High] Teardown race in pci_lmr_exit() causes resource leaks due to concurrent debugfs accesses.
- [High] Direct modification of ASPM hardware registers bypasses the PCI ASPM subsystem and its locking.
- [High] ASPM L1 is restored in the wrong order, enabling the Downstream component before the Upstream component.
- [Medium] Missing runtime PM reference before reading config space in debugfs show callbacks.
- [Medium] Mixing goto-based error handling with guard() cleanup in margin_enable_write().
- [Medium] Error return code is clobbered in the cleanup path of margin_enable_write().
- [Medium] Lane array size and debugfs entries incorrectly initialized using Negotiated Link Width instead of Maximum Link Width.
--
--- Patch [1]: [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support ---
commit 2b24000364052ba88665b523eec643f2f9d276e4
Author: Priyank Rathod <rathodpriyank@google.com>
Subject: PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
This commit adds support for PCIe Lane Margining at Receiver (LMR), which
allows software to evaluate high-speed link margins by measuring timing
and voltage steps for each physical lane. It exposes capabilities and
margining controls via debugfs.
Link: https://lore.kernel.org/r/20260820-pcie-lmt-v5-1-943b3b0e18bf@google.com
> --- a/drivers/pci/pci.c
> +++ b/drivers/pci/pci.c
> @@ -5145,8 +5145,10 @@ int __pci_reset_function_locked(struct pci_dev *dev)
> method = &pci_reset_fn_methods[m];
> pci_dbg(dev, "reset via %s\n", method->name);
> rc = method->reset_fn(dev, PCI_RESET_DO_RESET);
> - if (!rc)
> + if (!rc) {
> + pci_reset_lmr(dev);
> return 0;
> + }
[Severity: High]
Is it possible that the temporary ASPM disablement leaks into the
permanently saved PCI configuration state here?
The PCI core calls pci_save_state() prior to this function, capturing the
state while LMR (and its temporary ASPM disablement) might be active. When
the reset finishes, pci_restore_state() might overwrite the hardware config
with this saved state, clobbering the correct restoration performed by
pci_reset_lmr().
> --- /dev/null
> +++ b/drivers/pci/pcie/margin.c
[ ... ]
> +static void pci_lmr_disable_aspm(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *dev = mdev->dev;
> + struct pci_dev *parent = pci_upstream_bridge(dev);
> + u16 ctl;
> +
> + if (mdev->aspm_saved)
> + return;
> +
> + if (!pcie_capability_read_word(dev, PCI_EXP_LNKCTL, &ctl)) {
> + mdev->saved_aspm = ctl & PCI_EXP_LNKCTL_ASPMC;
> + pcie_capability_clear_word(dev, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_ASPMC);
[Severity: High]
Does modifying the ASPM hardware registers directly bypass the PCI ASPM
subsystem and its locking?
This clears the PCI_EXP_LNKCTL_ASPMC bits without holding the ASPM lock or
using the official PCI core APIs like pci_disable_link_state(). If the
ASPM core concurrently re-evaluates the link, or if userspace changes ASPM
policies, these direct modifications could race with or be clobbered by
the core's changes.
[ ... ]
> +static void pci_lmr_restore_aspm(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *dev = mdev->dev;
> + struct pci_dev *parent = pci_upstream_bridge(dev);
> +
> + if (!mdev->aspm_saved)
> + return;
> +
> + pcie_capability_clear_and_set_word(dev, PCI_EXP_LNKCTL,
> + PCI_EXP_LNKCTL_ASPMC,
> + mdev->saved_aspm);
> + if (parent && pci_is_pcie(parent))
> + pcie_capability_clear_and_set_word(parent, PCI_EXP_LNKCTL,
> + PCI_EXP_LNKCTL_ASPMC,
> + mdev->saved_parent_aspm);
[Severity: High]
Does this restore the Downstream component before the Upstream component?
PCIe r6.0 sec 7.5.3.7 mandates that when enabling ASPM L1, software must
enable the Upstream component prior to enabling the Downstream component.
This sequence appears to execute the L1 enablement backward.
[Severity: High]
Could blindly restoring these cached ASPM bits erase legitimate ASPM policy
changes that might have occurred while margining was active?
[ ... ]
> +static int margin_caps_show(struct seq_file *s, void *v)
> +{
> + struct pci_margin_dev *mdev = s->private;
> + struct pci_dev *dev = mdev->dev;
> + u16 cap;
> + int ret;
> +
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_CAP, &cap);
[Severity: Medium]
Could this return bogus configuration data (like 0xFFFF) to user space
if the device is in a lower power state (e.g., D3hot)?
There does not appear to be a pm_runtime_resume_and_get() call here to
ensure the device is in D0 before reading the config space.
[ ... ]
> +static int margin_port_status_show(struct seq_file *s, void *v)
> +{
> + struct pci_margin_dev *mdev = s->private;
> + struct pci_dev *dev = mdev->dev;
> + u16 sts;
> + int ret;
> +
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
[Severity: Medium]
Similar to the capabilities check, should this function wake the device
before attempting to read the port status?
[ ... ]
> +static ssize_t margin_enable_write(struct file *file, const char __user *user_buf,
> + size_t count, loff_t *ppos)
> +{
> + struct seq_file *s = file->private_data;
> + struct pci_margin_dev *mdev = s->private;
> + struct pci_dev *dev = mdev->dev;
> + unsigned long timeout;
> + u16 sts, cap, lnksta;
> + bool enable;
> + int ret, i;
> +
> + ret = kstrtobool_from_user(user_buf, count, &enable);
> + if (ret)
> + return ret;
> +
> + guard(mutex)(&mdev->lock);
[Severity: Medium]
Does this function mix goto-based error handling with guard() cleanup?
The function uses guard(mutex)(&mdev->lock) at the start but relies on
goto labels (like goto err_sw_ready) for error handling below. The kernel
cleanup subsystem guidelines specify that goto labels and scope-based
cleanup should not be mixed in the same function to avoid confusing
ownership semantics.
[ ... ]
> +err_sw_ready:
> + if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
> + if (ret == PCIBIOS_SUCCESSFUL) {
> + sts &= ~PCI_LMR_PORT_STS_SW_READY;
> + pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
> + }
> + }
> +err_aspm:
> + pci_lmr_restore_aspm(mdev);
> +err_rpm:
> + pm_runtime_put_sync(&dev->dev);
> + return ret;
[Severity: Medium]
Is the error return code clobbered in this cleanup path?
If a failure occurs (like -ETIMEDOUT), the code jumps to err_sw_ready. If
pci_read_config_word() succeeds here, ret is overwritten with 0
(PCIBIOS_SUCCESSFUL), masking the original error and returning success to
user space.
[ ... ]
> +static bool pci_lmr_check_sample_multiple_rx(struct pci_margin_dev *mdev,
> + struct pci_margin_lane *plane)
> +{
> + struct pci_margin_rx_info *info = &plane->rx_info[plane->rx];
> + int i;
> +
> + if (info->caps & LMR_CAP_SAMPLE_MULTIPLE_RX)
> + return true;
> +
> + for (i = 0; i < mdev->num_lanes; i++) {
> + struct pci_margin_lane *other = &mdev->lanes[i];
> +
> + if (i == plane->lane)
> + continue;
> + if (other->rx == plane->rx &&
> + (other->timing_val != 0 || other->voltage_val != 0))
> + return false;
> + }
> + return true;
> +}
[Severity: High]
Does this logic allow invalid concurrent margining while blocking valid
concurrent margining?
Per PCIe r6.0 sec 8.4.4, if 'Sample Multiple Receivers' is 0, software may
only margin a single receiver at a time across the link, but may do so on
any or all lanes.
The check "if (other->rx == plane->rx)" appears inverted. It wrongly
blocks evaluating the same receiver on multiple lanes, but allows different
receivers to be margined simultaneously because they skip the return block.
[ ... ]
> +void pci_lmr_init(struct pci_dev *dev)
> +{
> + struct pci_margin_dev *mdev;
> + enum pci_bus_speed speed;
> + u16 lmr, lnksta;
> + int num_lanes, i;
> +
> + if (WARN_ON_ONCE(!dev) || !pci_is_pcie(dev))
> + return;
[ ... ]
> + pcie_capability_read_word(dev, PCI_EXP_LNKSTA, &lnksta);
> + num_lanes = FIELD_GET(PCI_EXP_LNKSTA_NLW, lnksta);
[Severity: Medium]
Should the lane array size and debugfs entries be initialized using the
Maximum Link Width (PCI_EXP_LNKCAP_MLW) instead of the Negotiated Link
Width (PCI_EXP_LNKSTA_NLW)?
If the link trains to a reduced width at boot (e.g., x4 instead of x16) due
to power saving or retrain, the driver will lack the structures to margin
the upper lanes if it later retrains to its full width.
[ ... ]
> +void pci_lmr_exit(struct pci_dev *dev)
> +{
> + struct pci_margin_dev *mdev = dev->lmr;
> +
> + if (!dev || !mdev)
> + return;
> +
> + pci_suspend_lmr(dev);
> +
> + pci_margin_debugfs_remove(mdev);
> + mutex_destroy(&mdev->lock);
> + kfree(mdev);
> + dev->lmr = NULL;
> +}
[Severity: High]
Does pci_suspend_lmr() drop the lock before the debugfs files are removed?
pci_suspend_lmr() drops the lock after disabling LMR. Before
pci_margin_debugfs_remove() finishes, a concurrent write to the 'enable'
debugfs file could acquire the lock and re-enable margining. When teardown
completes, the hardware state and PM references initialized by the racing
write might be left dangling. Debugfs files usually need to be removed
before tearing down the associated hardware state.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260824-pcie-lmt-v6-1-bab4ce233fa5@google.com?part=1
^ permalink raw reply [flat|nested] 5+ messages in thread
* Re: [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
2026-08-24 22:38 [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support Priyank Rathod
2026-08-24 22:51 ` sashiko-bot
@ 2026-08-25 11:02 ` Ilpo Järvinen
1 sibling, 0 replies; 5+ messages in thread
From: Ilpo Järvinen @ 2026-08-25 11:02 UTC (permalink / raw)
To: Priyank Rathod
Cc: Bjorn Helgaas, Shuah Khan, Kees Cook, Gustavo A. R. Silva,
Jonathan Corbet, Shuah Khan, LKML, linux-pci, linux-kselftest,
linux-hardening, linux-doc
[-- Attachment #1: Type: text/plain, Size: 36529 bytes --]
On Mon, 24 Aug 2026, Priyank Rathod wrote:
> Per PCIe Base Specification r6.0, sec 8.4.4 ("Lane Margining at
> Receiver"), PCIe devices operating at 16.0 GT/s (Gen 4) or higher data
> rates support the Lane Margining at Receiver Extended Capability
> (ID 0x27), and it is mandatory for receivers operating at 64.0 GT/s
> (Gen 6) or higher data rates. Lane Margining allows software to
> evaluate high-speed link margins by measuring timing and voltage steps
> for each individual physical lane and receiver.
>
> Add driver and debugfs support for PCIe Lane Margining at Receiver:
>
> - Add Lane Margining at Receiver Extended Capability register
> definitions (PCI_EXT_CAP_ID_LMR, PCI_LMR_PORT_CAP, PCI_LMR_PORT_STS,
> PCI_LMR_LANE_CTRL, PCI_LMR_LANE_STS) to <uapi/linux/pci_regs.h>.
> - Add Kconfig option CONFIG_PCIE_LMR (under drivers/pci/pcie/Kconfig)
> dependent on DEBUG_FS.
> - Implement drivers/pci/pcie/margin.c to probe the capability on Gen4+
> links and expose per-device debugfs entries under:
> /sys/kernel/debug/pci/pcie_lmr_<pci_dev_name>/
> providing control over margining enablement, receiver selection, and
> execution of timing/voltage margin step commands. Distinguish
> between missing mandatory LMR capability on Gen6+ vs optional on
> Gen4/Gen5.
> - Hook pci_lmr_init() into pci_init_capabilities() during device probe
> in drivers/pci/probe.c and pci_lmr_exit() into drivers/pci/remove.c.
> - Add kselftest script under tools/testing/selftests/pcie_lmt/pcie_lmt.sh
> to test debugfs capability reads, enablement, and stepping.
> - Add MAINTAINERS entry for PCIe Lane Margining at Receiver (LMR).
>
> Signed-off-by: Priyank Rathod <rathodpriyank@google.com>
> ---
> Changes in v6:
> - Added kernel documentation under Documentation/PCI/pcie-lmr.rst and indexed in Documentation/PCI/index.rst (Ilpo Järvinen).
> - Updated MAINTAINERS with Documentation/PCI/pcie-lmr.rst (Ilpo Järvinen).
> - Aligned capability bit naming and comments with PCIe Base Specification r6.0 sec 8.4.4 Table "Report Margining Capabilities Payload" (Ilpo Järvinen).
> - Clarified Sample Multiple Receivers concurrency verification and rules across physical lanes in kerneldoc and documentation (Ilpo Järvinen).
> - Refactored pci_lmr_run_cmd() to pass struct pci_margin_dev *mdev directly, eliminating redundant NULL checks and using mdev->num_lanes (Ilpo Järvinen).
> - Converted PCI config read/write return checking across all helpers to pcibios_err_to_errno() (Ilpo Järvinen).
> - Reversed return logic in pci_lmr_demargin_lane() to return early on error (Ilpo Järvinen).
> - Refactored margin_lane_step_write() to eliminate bool is_voltage parameter, using command type (LMR_TYPE_TIMING / LMR_TYPE_VOLTAGE) and switch/case with consolidated bounds checks (Ilpo Järvinen).
> - Renamed __pci_suspend_lmr_locked() to pci_lmr_disable_locked() to avoid PM terminology confusion and added lockdep_assert_held(&mdev->lock) (Ilpo Järvinen).
> - Replaced -EACCES with -EBUSY across debugfs show/write callbacks when margining is inactive (Ilpo Järvinen).
> - Clarified comment for active operating link speed check (Gen4+ capability vs dynamically operating speed) in margin_enable_write() (Ilpo Järvinen).
> - Added WARN_ON_ONCE(!dev) check in pci_lmr_init() (Ilpo Järvinen).
> - Demoted capability detection log message from pci_info to pci_dbg to prevent boot log noise (Ilpo Järvinen).
> - Fixed timing step mask extraction in pci_lmr_cache_rx_info() to use 6-bit LMR_TIMING_STEP_MASK (sashiko-bot).
> - Resumed runtime PM via pm_runtime_resume_and_get() before performing config space reads in margin_enable_write() (sashiko-bot).
> - Switched to pm_runtime_put_sync() during margining teardown (sashiko-bot).
> - Link to v5: https://lore.kernel.org/r/20260820-pcie-lmt-v5-1-943b3b0e18bf@google.com
>
> PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
>
> Per PCIe Base Specification r6.0, section 8.4.4 ("Lane Margining at Receiver"),
> PCIe devices operating at 16.0 GT/s (Gen 4) or higher data rates support the
> Lane Margining at Receiver Extended Capability (ID 0x27), and it is mandatory
> for receivers operating at 64.0 GT/s (Gen 6) or higher data rates.
>
> Lane Margining allows system software to evaluate high-speed link signal
> integrity and margins by measuring timing and voltage steps for each physical
> lane and receiver independently.
>
> This series introduces kernel driver support, debugfs controls, and a
> kselftest automation script for PCIe Lane Margining at Receiver (LMR/LMT).
>
> ==============================================================================
> 1. How to Enable & Configure
> ==============================================================================
> Enable the Kconfig option under PCI support:
> CONFIG_PCIE_LMR=y (or =m)
> (Depends on CONFIG_PCI and CONFIG_DEBUG_FS)
>
> Upon boot or device hotplug on Gen4+ links (>= 16.0 GT/s), the driver probes
> Extended Capability ID 0x27 and exposes per-device debugfs interfaces:
> /sys/kernel/debug/pci/pcie_lmr_<domain>:<bus>:<dev>.<func>/
>
> ==============================================================================
> 2. How to Use the Debugfs Interface (Manual Margining)
> ==============================================================================
> Inspect device-wide margining capabilities and port status:
> # Inspect root device LMR capabilities & status
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/capabilities
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/port_status
>
> Enable active Lane Margining on the device:
> # Enable Lane Margining state machine
> echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
>
> Inspect and step individual lanes (e.g. lane0):
> # Select target receiver (0 = local receiver, 1..6 = retimers/link partners)
> echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/receiver
>
> # Check available timing and voltage steps for this receiver
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/caps
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_timing_steps
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_voltage_steps
>
> # Step timing margin or voltage margin offset
> echo 2 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
> echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
>
> # Reset margin offset back to nominal (0)
> echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
> echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
>
> Disable Lane Margining when finished:
> echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
>
> ==============================================================================
> 3. How to Run Automated Kselftests Using the Test Script
> ==============================================================================
> An automated kselftest script is included to test capability reads, receiver
> selection, and margining commands across all enumerated LMR devices:
>
> # Run directly as root
> sudo ./tools/testing/selftests/pcie_lmt/pcie_lmt.sh
>
> Or run via the kselftest Makefile harness:
> make -C tools/testing/selftests TARGETS=pcie_lmt run_tests
>
> Sample script output on an LMR-capable device:
> pcie_lmt: testing PCIe LMR debugfs entries
> pcie_lmt: probing device pcie_lmr_0000:01:00.0
> pcie_lmr_0000:01:00.0: capabilities read OK
> pcie_lmr_0000:01:00.0: port_status read OK
> pcie_lmr_0000:01:00.0: margining enabled OK
> pcie_lmr_0000:01:00.0: testing lane0
> pcie_lmr_0000:01:00.0: testing lane1
> pcie_lmr_0000:01:00.0: margining disabled OK
> pcie_lmt [PASS]
>
> To: Bjorn Helgaas <bhelgaas@google.com>
> To: Shuah Khan <shuah@kernel.org>
> Cc: linux-kernel@vger.kernel.org
> Cc: linux-pci@vger.kernel.org
> Cc: linux-kselftest@vger.kernel.org
> Cc: Ilpo Järvinen <ilpo.jarvinen@linux.intel.com>
>
> Changes in v5:
> - Sorted #include directives alphabetically and added missing includes for bits.h, bitfield.h, cleanup.h, overflow.h, and slab.h (Ilpo Järvinen).
> - Converted bitmasks to GENMASK() and BIT() macros and used FIELD_PREP() and FIELD_GET() instead of manual bit shifts (Ilpo Järvinen).
> - Added pci_lmr_sts_payload() helper to cleanly extract the status payload byte before applying step and capability masks (Ilpo Järvinen).
> - Replaced manual mutex locking sequences with guard(mutex)(&mdev->lock) across show and write callbacks to simplify control flow (Ilpo Järvinen).
> - Documented mutex lock protection scope in kerneldoc for struct pci_margin_dev (Ilpo Järvinen).
> - Used standard PCI_POSSIBLE_ERROR(), str_yes_no(), and scnprintf() helpers throughout the driver (Ilpo Järvinen).
> - Clarified receiver range (0..6 per PCIe r6.0 sec 8.4.4; 7 reserved) in comments and validation checks (Ilpo Järvinen).
> - Deduplicated timing and voltage show/write handlers using margin_lane_steps_show() and margin_lane_step_write() (Ilpo Järvinen).
> - Placed speed check immediately following pcie_get_speed_cap() and handled PCI_SPEED_UNKNOWN (Ilpo Järvinen).
> - Converted lanes in struct pci_margin_dev to a flexible array member with __counted_by(num_lanes) allocated via struct_size() (Ilpo Järvinen).
>
> Changes in v4:
> - Added Sample Multiple Receivers (Bit 5) concurrency verification in margin_lane_timing_write() and margin_lane_voltage_write() per PCIe r6.0 sec 8.4.4, returning -EBUSY if another lane on the same receiver is already margined when simultaneous lane margining is not supported.
> - Added active operating link speed verification (PCI_EXP_LNKSTA_CLS >= 16.0 GT/s) in margin_enable_write() before enabling LMR, as LMR commands are physically undefined on links operating at Gen1/Gen2/Gen3 speeds.
> - Added fast-path hardware NAK detection in pci_lmr_run_cmd() to return -EOPNOTSUPP immediately if a receiver echoes MTYPE == NO_CMD (0x7) after command issuance rather than waiting 150ms for a timeout.
> - Added pci_reset_lmr() hooked into __pci_reset_function_locked() to synchronize software state and demargin on FLR or Secondary Bus Reset.
> - Comprehensive NULL pointer checks and array/lane/receiver bounds checks added across all internal helpers and debugfs write handlers.
> - Added MAINTAINERS entry for PCIe Lane Margining at Receiver (LMR).
>
> Changes in v2:
> - Fixed NO_CMD (0x7) clearing in pci_lmr_run_cmd() before issuing new commands per PCIe r6.0 sec 8.4.4.
> - Protected plane->rx updates with mdev->lock in margin_lane_receiver_write().
> - Corrected Margining Port Capabilities bit definition to PCI_LMR_PORT_CAP_USES_SW_READY (0x0001) in <uapi/linux/pci_regs.h>.
> - Updated kselftest script (pcie_lmt.sh) to locate LMR debugfs entries.
> - Validated integer bounds against LMR_MAX_TIMING_STEP / LMR_MAX_VOLTAGE_STEP before narrowing u8 cast.
> - Moved mdev->enabled checks inside mutex_lock(&mdev->lock) to eliminate TOCTOU races.
> - Checked return values of all pci_read_config_word() calls, propagating -EIO on failure.
> - Eliminated dead store of cap in margin_enable_write().
> - Explicitly checked speed == PCIE_SPEED_64_0GT in pci_lmr_init() to avoid misidentifying PCI_SPEED_UNKNOWN (0xFF) as Gen6.
> ---
> Documentation/PCI/index.rst | 1 +
> Documentation/PCI/pcie-lmr.rst | 171 +++++
> MAINTAINERS | 8 +
> drivers/pci/pci-driver.c | 1 +
> drivers/pci/pci.c | 4 +-
> drivers/pci/pci.h | 12 +
> drivers/pci/pcie/Kconfig | 12 +
> drivers/pci/pcie/Makefile | 1 +
> drivers/pci/pcie/margin.c | 1061 ++++++++++++++++++++++++++
> drivers/pci/probe.c | 1 +
> drivers/pci/remove.c | 1 +
> include/linux/pci.h | 6 +
> include/uapi/linux/pci_regs.h | 18 +
> tools/testing/selftests/Makefile | 1 +
> tools/testing/selftests/pcie_lmt/Makefile | 3 +
> tools/testing/selftests/pcie_lmt/pcie_lmt.sh | 105 +++
> 16 files changed, 1405 insertions(+), 1 deletion(-)
>
> diff --git a/Documentation/PCI/index.rst b/Documentation/PCI/index.rst
> index 5d720d2a415e..9170c98cbf3f 100644
> --- a/Documentation/PCI/index.rst
> +++ b/Documentation/PCI/index.rst
> @@ -20,3 +20,4 @@ PCI Bus Subsystem
> controller/index
> boot-interrupts
> tph
> + pcie-lmr
> diff --git a/Documentation/PCI/pcie-lmr.rst b/Documentation/PCI/pcie-lmr.rst
> new file mode 100644
> index 000000000000..1ebb8317cc83
> --- /dev/null
> +++ b/Documentation/PCI/pcie-lmr.rst
> @@ -0,0 +1,171 @@
> +.. SPDX-License-Identifier: GPL-2.0
> +
> +=======================================================
> +PCI Express Lane Margining at Receiver (LMR) Subsystem
> +=======================================================
> +
> +:Author: Priyank Rathod <rathodpriyank@google.com>
> +:Copyright: 2026 Google LLC
> +
> +Overview
> +========
> +
> +Lane Margining at Receiver (LMR), specified in the PCI Express Base
> +Specification (r6.0+ sec 8.4.4), allows system software to evaluate high-speed
> +link physical signal integrity and eye margins. LMR measures available timing
> +(jitter/phase) and voltage margin offsets for each physical lane and receiver
> +independently while the link is operating in active L0 state.
> +
> +Lane Margining Extended Capability (ID 0x27) is optional for links operating at
> +16.0 GT/s (PCIe Gen 4) and 32.0 GT/s (Gen 5), and is mandatory for receivers
> +operating at 64.0 GT/s (Gen 6) and higher.
> +
> +Target Receivers
> +================
> +
> +Each physical lane can margin up to 7 distinct receivers per PCIe link:
> +
> +* **Receiver 0 (Local Receiver)**: The receiver in the immediate link partner.
> +* **Receivers 1 to 6 (Retimers)**: Retimer pseudo-ports along the physical link
> + (up to 3 retimers, each with upstream and downstream pseudo-ports).
> +* **Receiver 7**: Reserved per PCIe Base Specification.
> +
> +Kernel Configuration
> +====================
> +
> +Enable the kernel configuration option under PCI support:
> +
> +.. code-block:: none
> +
> + CONFIG_PCIE_LMR=y (or =m)
> +
> +Dependencies:
> +* ``CONFIG_PCI``
> +* ``CONFIG_DEBUG_FS``
> +
> +Debugfs Interface Guide
> +=======================
> +
> +When an LMR-capable device is enumerated on a Gen4+ link, the kernel exposes
> +per-device control and status files under debugfs:
> +
> +.. code-block:: none
> +
> + /sys/kernel/debug/pci/pcie_lmr_<domain>:<bus>:<dev>.<func>/
> +
> +Device-Level Attributes
> +-----------------------
> +
> +* ``capabilities`` (read-only):
> + Displays the 16-bit Margining Port Capabilities register and whether the
> + device uses the Software Ready handshake bit.
> +
> +* ``port_status`` (read-only):
> + Displays the Margining Port Status register, indicating Margining Ready and
> + SW Ready states.
> +
> +* ``enable`` (read-write):
> + Enables (``1``) or disables (``0``) Lane Margining on the device.
> + Enabling margining locks the link into D0, prevents runtime PM suspend,
> + disables ASPM L0s/L1, and verifies that the link is operating at >= 16.0 GT/s.
> + Disabling margining restores ASPM and runtime PM, and returns all lanes to
> + nominal (demargined) state.
> +
> +Lane-Level Attributes
> +---------------------
> +
> +For each physical lane (``lane0``, ``lane1``, ...):
> +
> +* ``receiver`` (read-write):
> + Gets or sets the active target receiver number (``0`` for local receiver,
> + ``1..6`` for retimers). Switching receivers automatically demargins previous
> + offsets per PCIe single-receiver margining requirements.
> +
> +* ``caps`` (read-only):
> + Reports the target receiver's margining capabilities:
> + - Margining uses Driver Software (vs hardware autonomous)
> + - Independent Left/Right Timing Margining support
> + - Independent Up/Down Voltage Margining support
> + - Error Sampler vs Main Sampler
> + - Sample Multiple Receivers support
> +
> +* ``num_timing_steps`` (read-only):
> + Maximum timing margin steps supported by the receiver (0..63).
> +
> +* ``num_voltage_steps`` (read-only):
> + Maximum voltage margin steps supported by the receiver (0..127).
> +
> +* ``margin_timing`` (read-write):
> + Applies timing margin step offset (+/-). Writing ``0`` clears timing margin
> + back to nominal.
> +
> +* ``margin_voltage`` (read-write):
> + Applies voltage margin step offset (+/-). Writing ``0`` clears voltage margin
> + back to nominal.
> +
> +Manual Margining Example
> +========================
> +
> +1. Inspect device capabilities and status:
> +
> +.. code-block:: sh
> +
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/capabilities
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/port_status
> +
> +2. Enable Lane Margining mode:
> +
> +.. code-block:: sh
> +
> + echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
> +
> +3. Configure target receiver and inspect step limits on lane 0:
> +
> +.. code-block:: sh
> +
> + echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/receiver
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/caps
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_timing_steps
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_voltage_steps
> +
> +4. Apply timing and voltage margin steps:
> +
> +.. code-block:: sh
> +
> + # Step timing margin +2 steps
> + echo 2 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
> +
> + # Step voltage margin +1 step
> + echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
> +
> +5. Reset margins back to nominal:
> +
> +.. code-block:: sh
> +
> + echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
> + echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
> +
> +6. Disable Lane Margining when complete:
> +
> +.. code-block:: sh
> +
> + echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
> +
> +Automated Testing via Kselftest
> +===============================
> +
> +The kernel includes an automated kselftest script under
> +``tools/testing/selftests/pcie_lmt/pcie_lmt.sh`` to probe, validate, and exercise
> +debugfs controls across all enumerated LMR devices.
> +
> +Run directly as root:
> +
> +.. code-block:: sh
> +
> + sudo ./tools/testing/selftests/pcie_lmt/pcie_lmt.sh
> +
> +Or run via the kselftest test harness:
> +
> +.. code-block:: sh
> +
> + make -C tools/testing/selftests TARGETS=pcie_lmt run_tests
> diff --git a/MAINTAINERS b/MAINTAINERS
> index b7094a616afd..b5deaae11bfe 100644
> --- a/MAINTAINERS
> +++ b/MAINTAINERS
> @@ -21059,6 +21059,14 @@ F: Documentation/devicetree/bindings/pci/qcom,sa8255p-pcie-ep.yaml
> F: drivers/pci/controller/dwc/pcie-qcom-common.c
> F: drivers/pci/controller/dwc/pcie-qcom-ep.c
>
> +PCIE LANE MARGINING AT RECEIVER (LMR)
> +M: Priyank Rathod <rathodpriyank@google.com>
> +L: linux-pci@vger.kernel.org
> +S: Maintained
> +F: Documentation/PCI/pcie-lmr.rst
> +F: drivers/pci/pcie/margin.c
> +F: tools/testing/selftests/pcie_lmt/
> +
> PCMCIA SUBSYSTEM
> M: Dominik Brodowski <linux@dominikbrodowski.net>
> S: Odd Fixes
> diff --git a/drivers/pci/pci-driver.c b/drivers/pci/pci-driver.c
> index f36778e62ac1..17544a7023fc 100644
> --- a/drivers/pci/pci-driver.c
> +++ b/drivers/pci/pci-driver.c
> @@ -821,6 +821,7 @@ static int pci_pm_suspend(struct device *dev)
> * since Coffee Lake, to enter a lower-power PM state.
> */
> pci_suspend_ptm(pci_dev);
> + pci_suspend_lmr(pci_dev);
>
> if (pci_has_legacy_pm_support(pci_dev))
> return pci_legacy_suspend(dev, PMSG_SUSPEND);
> diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
> index 77b17b13ee61..dc9724cb7b4d 100644
> --- a/drivers/pci/pci.c
> +++ b/drivers/pci/pci.c
> @@ -5145,8 +5145,10 @@ int __pci_reset_function_locked(struct pci_dev *dev)
> method = &pci_reset_fn_methods[m];
> pci_dbg(dev, "reset via %s\n", method->name);
> rc = method->reset_fn(dev, PCI_RESET_DO_RESET);
> - if (!rc)
> + if (!rc) {
> + pci_reset_lmr(dev);
> return 0;
> + }
>
> pci_dbg(dev, "%s failed with %d\n", method->name, rc);
> if (rc != -ENOTTY)
> diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
> index 4469e1a77f3c..6322a81f9e50 100644
> --- a/drivers/pci/pci.h
> +++ b/drivers/pci/pci.h
> @@ -1023,6 +1023,18 @@ static inline void pci_no_tph(void) { }
> static inline void pci_tph_init(struct pci_dev *dev) { }
> #endif
>
> +#ifdef CONFIG_PCIE_LMR
> +void pci_lmr_init(struct pci_dev *dev);
> +void pci_lmr_exit(struct pci_dev *dev);
> +void pci_suspend_lmr(struct pci_dev *dev);
> +void pci_reset_lmr(struct pci_dev *dev);
> +#else
> +static inline void pci_lmr_init(struct pci_dev *dev) { }
> +static inline void pci_lmr_exit(struct pci_dev *dev) { }
> +static inline void pci_suspend_lmr(struct pci_dev *dev) { }
> +static inline void pci_reset_lmr(struct pci_dev *dev) { }
> +#endif
> +
> #ifdef CONFIG_PCIE_PTM
> void pci_ptm_init(struct pci_dev *dev);
> void pci_save_ptm_state(struct pci_dev *dev);
> diff --git a/drivers/pci/pcie/Kconfig b/drivers/pci/pcie/Kconfig
> index 207c2deae35f..3b021ca2fe84 100644
> --- a/drivers/pci/pcie/Kconfig
> +++ b/drivers/pci/pcie/Kconfig
> @@ -137,6 +137,18 @@ config PCIE_PTM
> This is only useful if you have devices that support PTM, but it
> is safe to enable even if you don't.
>
> +config PCIE_LMR
> + bool "PCI Express Lane Margining at Receiver Support"
> + depends on DEBUG_FS
> + help
> + This enables the PCI Express Lane Margining at Receiver support.
> + Lane Margining allows software to determine the voltage and
> + timing margin of each lane on a PCIe link (16.0 GT/s and above).
> + The margining data is exposed via debugfs.
> +
> + This is only useful if you have devices that support lane
> + margining, but it is safe to enable even if you don't.
> +
> config PCIE_EDR
> bool "PCI Express Error Disconnect Recover support"
> depends on PCIE_DPC && ACPI
> diff --git a/drivers/pci/pcie/Makefile b/drivers/pci/pcie/Makefile
> index b0b43a18c304..aac45ae0402e 100644
> --- a/drivers/pci/pcie/Makefile
> +++ b/drivers/pci/pcie/Makefile
> @@ -13,4 +13,5 @@ obj-$(CONFIG_PCIEAER_INJECT) += aer_inject.o
> obj-$(CONFIG_PCIE_PME) += pme.o
> obj-$(CONFIG_PCIE_DPC) += dpc.o
> obj-$(CONFIG_PCIE_PTM) += ptm.o
> +obj-$(CONFIG_PCIE_LMR) += margin.o
> obj-$(CONFIG_PCIE_EDR) += edr.o
> diff --git a/drivers/pci/pcie/margin.c b/drivers/pci/pcie/margin.c
> new file mode 100644
> index 000000000000..45646b5952d6
> --- /dev/null
> +++ b/drivers/pci/pcie/margin.c
> @@ -0,0 +1,1061 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * PCI Express Lane Margining at Receiver
> + *
> + * Copyright (C) 2026 Google LLC
> + * Author: Priyank Rathod <rathodpriyank@google.com>
> + *
> + * Lane Margining at Receiver (PCIe Base Specification r6.0, sec 8.4.4)
> + * allows system software to determine the voltage and timing margins of
> + * each physical lane on a PCIe link. The Extended Capability (ID 0x27)
> + * is available for receivers operating at 16.0 GT/s (Gen4) or higher data
> + * rates, and is mandatory for receivers operating at 64.0 GT/s (Gen6) or
> + * higher data rates.
> + *
> + * This driver implements:
> + * - Probing Extended Capability ID 0x27 and Margining Port Capabilities.
> + * - Managing ASPM L0s/L1 link states during active margining with restoration.
> + * - PCIe r6.0 NO_CMD (0x7) clearing handshake per receiver and lane.
> + * - Caching receiver capabilities & step counts to avoid DEMARGIN side-effects.
> + * - Handling Symmetric vs Independent Left/Right & Up/Down margin steps.
> + * - Runtime PM protection (D0 enforcement) during active margining.
> + * - Exposing per-device debugfs interfaces under /sys/kernel/debug/pci/.
> + */
> +
> +#include <linux/bitfield.h>
> +#include <linux/bits.h>
> +#include <linux/cleanup.h>
> +#include <linux/debugfs.h>
> +#include <linux/delay.h>
> +#include <linux/err.h>
> +#include <linux/errno.h>
> +#include <linux/jiffies.h>
> +#include <linux/kstrtox.h>
> +#include <linux/minmax.h>
> +#include <linux/mutex.h>
> +#include <linux/overflow.h>
> +#include <linux/pci.h>
> +#include <linux/pm_runtime.h>
> +#include <linux/seq_file.h>
> +#include <linux/slab.h>
> +#include <linux/sprintf.h>
> +#include <linux/string_choices.h>
> +#include <linux/types.h>
> +
> +#include "../pci.h"
> +
> +/* Margin type encodings per PCIe Base Spec r6.0 sec 8.4.4 */
> +#define LMR_TYPE_DEMARGIN 0x0
> +#define LMR_TYPE_REPORT_CAPS 0x1
> +#define LMR_TYPE_REPORT_VOLTAGE_STEPS 0x2
> +#define LMR_TYPE_REPORT_TIMING_STEPS 0x3
> +#define LMR_TYPE_TIMING 0x4
> +#define LMR_TYPE_VOLTAGE 0x5
> +#define LMR_TYPE_NO_CMD 0x7
> +
> +/* LMR command timing parameters */
> +#define LMR_CMD_TIMEOUT_MS 150
> +#define LMR_CMD_SLEEP_MIN_US 100
> +#define LMR_CMD_SLEEP_MAX_US 250
> +#define LMR_ENABLE_TIMEOUT_MS 150
> +#define LMR_ENABLE_SLEEP_MIN_US 1000
> +#define LMR_ENABLE_SLEEP_MAX_US 2000
> +
> +/*
> + * LMR limits:
> + * Valid receiver numbers are 0 (local receiver) to 6 (up to 3 retimers)
> + * per PCIe Base Specification r6.0 sec 8.4.4. Receiver number 7 is reserved.
> + */
> +#define LMR_MAX_LANES 32
> +#define LMR_MAX_RX_NUM 6
> +#define LMR_MAX_TIMING_STEP 63
> +#define LMR_MAX_VOLTAGE_STEP 127
> +
> +/* LMR PCIe generation numbers and helper */
> +#define LMR_GEN6 6
> +#define LMR_GEN5 5
> +#define LMR_GEN4 4
> +
> +#define LMR_SPEED_TO_GEN(speed) \
> + ((speed) >= PCIE_SPEED_64_0GT ? LMR_GEN6 : \
> + (speed) >= PCIE_SPEED_32_0GT ? LMR_GEN5 : \
> + LMR_GEN4)
> +
> +/* LMR lane register stride */
> +#define LMR_LANE_REG_STRIDE 4
> +
> +/* LMR receivers and directions */
> +#define LMR_RX_LOCAL 0
> +#define LMR_STEP_DIR_INCREASE 1
> +#define LMR_STEP_DIR_DECREASE 0
> +
> +/* LMR payload field masks per PCIe Base Spec r6.0 sec 8.4.4 */
> +#define LMR_STEPS_MASK GENMASK(6, 0)
> +#define LMR_TIMING_STEP_MASK GENMASK(5, 0)
> +#define LMR_TIMING_DIR_MASK BIT(6)
> +#define LMR_VOLTAGE_STEP_MASK GENMASK(6, 0)
> +#define LMR_VOLTAGE_DIR_MASK BIT(7)
> +
> +/*
> + * Margining Capabilities report bit fields (PCIe Base Spec r6.0 sec 8.4.4,
> + * Table "Report Margining Capabilities Payload"):
In 8.4.4, my copy of r6.0.1 (and same with r7.0) PCIe spec, I only have
one table and that is called:
"Table 8-13 Lane Margining"
And no search finds "Report Margining Capabilities Payload" table anywhere.
...So I still fail to find it.
(AI has tendencity to come up non-existing things, I hope it's not the
case here as it would be rather rude to waste reviewers time on chasing
non-existing things.)
> + * Bit 0: Margining Uses Driver Software (1 = Driver software sequence; 0 = Hardware)
> + * Bit 2: Independent Left/Right Timing Margining Supported (1 = Supported; 0 = Symmetric)
> + * Bit 3: Independent Up/Down Voltage Margining Supported (1 = Supported; 0 = Symmetric)
> + * Bit 4: Margining Error Sampler (1 = Error Sampler; 0 = Main Sampler)
> + * Bit 5: Sample Multiple Receivers (1 = Multiple receivers; 0 = Single receiver only)
> + */
> +#define LMR_CAP_USES_DRIVER_SW BIT(0)
> +#define LMR_CAP_IND_LEFT_RIGHT_TIMING BIT(2)
> +#define LMR_CAP_IND_UP_DOWN_VOLTAGE BIT(3)
> +#define LMR_CAP_ERROR_SAMPLER BIT(4)
> +#define LMR_CAP_SAMPLE_MULTIPLE_RX BIT(5)
> +
> +/**
> + * struct pci_margin_rx_info - Cached Lane Margining receiver capabilities
> + * @caps_cached: True if receiver capabilities and step limits are cached
> + * @caps: Margining capabilities byte reported by receiver
> + * @num_timing_steps: Maximum timing margin steps supported by receiver
> + * @num_voltage_steps: Maximum voltage margin steps supported by receiver
> + */
> +struct pci_margin_rx_info {
> + bool caps_cached;
> + u8 caps;
> + u8 num_timing_steps;
> + u8 num_voltage_steps;
> +};
> +
> +/**
> + * struct pci_margin_lane - Per-lane margining state
> + * @mdev: Parent LMR margin device
> + * @lane: Physical lane index (0..num_lanes - 1)
> + * @rx: Selected target receiver number (0 = local, 1..6 = retimers)
> + * @timing_val: Current applied timing margin step offset (+/-)
> + * @voltage_val: Current applied voltage margin step offset (+/-)
> + * @rx_info: Cached receiver capabilities per receiver number
> + */
> +struct pci_margin_lane {
> + struct pci_margin_dev *mdev;
> + int lane;
> + u8 rx;
> + int timing_val;
> + int voltage_val;
> + struct pci_margin_rx_info rx_info[LMR_MAX_RX_NUM + 1];
> +};
> +
> +/**
> + * struct pci_margin_dev - PCIe Lane Margining device instance
> + * @dev: Underlying PCI device
> + * @cap: Extended capability offset (PCI_EXT_CAP_ID_LMR)
> + * @debugfs: Root debugfs dentry for this device
> + * @lock: Mutex protecting LMR hardware access, active margining enablement,
> + * target receiver selection, lane margining steps, and ASPM state
> + * @enabled: True if Lane Margining is currently enabled
> + * @aspm_saved: True if original ASPM configuration has been saved
> + * @saved_aspm: Saved ASPM control register bits for the device
> + * @saved_parent_aspm: Saved ASPM control register bits for parent bridge
> + * @num_lanes: Number of lanes on the link
> + * @lanes: Flexible array of per-lane state structures
> + */
> +struct pci_margin_dev {
> + struct pci_dev *dev;
> + u16 cap;
> + struct dentry *debugfs;
> + struct mutex lock;
> + bool enabled;
> + bool aspm_saved;
> + u16 saved_aspm;
> + u16 saved_parent_aspm;
> + int num_lanes;
> + struct pci_margin_lane lanes[] __counted_by(num_lanes);
> +};
> +
> +#if IS_ENABLED(CONFIG_DEBUG_FS)
> +static DEFINE_MUTEX(pci_debugfs_root_lock);
> +static struct dentry *pci_debugfs_root_dir;
> +
> +static struct dentry *get_pci_debugfs_root(void)
> +{
> + mutex_lock(&pci_debugfs_root_lock);
> + if (!pci_debugfs_root_dir)
> + pci_debugfs_root_dir = debugfs_lookup("pci", NULL);
> + if (!pci_debugfs_root_dir)
> + pci_debugfs_root_dir = debugfs_create_dir("pci", NULL);
> + mutex_unlock(&pci_debugfs_root_lock);
> + return pci_debugfs_root_dir;
> +}
> +#endif
> +
> +/*
> + * pci_lmr_disable_aspm() - Temporarily disable ASPM L0s/L1 during active
> + * margining per PCIe Base Spec r6.0 sec 8.4.4, saving original ASPMC bits.
> + */
> +static void pci_lmr_disable_aspm(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *dev = mdev->dev;
> + struct pci_dev *parent = pci_upstream_bridge(dev);
> + u16 ctl;
> +
> + if (mdev->aspm_saved)
> + return;
> +
> + if (!pcie_capability_read_word(dev, PCI_EXP_LNKCTL, &ctl)) {
> + mdev->saved_aspm = ctl & PCI_EXP_LNKCTL_ASPMC;
> + pcie_capability_clear_word(dev, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_ASPMC);
> + }
> +
> + if (parent && pci_is_pcie(parent)) {
> + if (!pcie_capability_read_word(parent, PCI_EXP_LNKCTL, &ctl)) {
> + mdev->saved_parent_aspm = ctl & PCI_EXP_LNKCTL_ASPMC;
> + pcie_capability_clear_word(parent, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_ASPMC);
> + }
> + }
> + mdev->aspm_saved = true;
Since you ignored my previous inquiry, I'm asking again...
How exactly you intend to prevent the aspm driver from re-enabling ASPM
while this driver wants it to remain off?
> +}
> +
> +/*
> + * pci_lmr_restore_aspm() - Restore original ASPM L0s/L1 state when margining
> + * is disabled or torn down.
> + */
> +static void pci_lmr_restore_aspm(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *dev = mdev->dev;
> + struct pci_dev *parent = pci_upstream_bridge(dev);
> +
> + if (!mdev->aspm_saved)
> + return;
> +
> + pcie_capability_clear_and_set_word(dev, PCI_EXP_LNKCTL,
> + PCI_EXP_LNKCTL_ASPMC,
> + mdev->saved_aspm);
> + if (parent && pci_is_pcie(parent))
> + pcie_capability_clear_and_set_word(parent, PCI_EXP_LNKCTL,
> + PCI_EXP_LNKCTL_ASPMC,
> + mdev->saved_parent_aspm);
> + mdev->aspm_saved = false;
> +}
> +static ssize_t margin_enable_write(struct file *file, const char __user *user_buf,
> + size_t count, loff_t *ppos)
> +{
> + struct seq_file *s = file->private_data;
> + struct pci_margin_dev *mdev = s->private;
> + struct pci_dev *dev = mdev->dev;
> + unsigned long timeout;
> + u16 sts, cap, lnksta;
> + bool enable;
> + int ret, i;
> +
> + ret = kstrtobool_from_user(user_buf, count, &enable);
> + if (ret)
> + return ret;
> +
> + guard(mutex)(&mdev->lock);
> +
> + if (mdev->enabled == enable)
> + return count;
> +
> + if (!enable) {
> + pci_lmr_disable_locked(mdev);
> + return count;
> + }
> +
> + /* Ensure device is powered (D0) before reading configuration registers */
> + ret = pm_runtime_resume_and_get(&dev->dev);
> + if (ret < 0)
> + return ret;
> +
> + /*
> + * PCIe r6.0 sec 8.4.4: LMR is physically undefined below 16.0 GT/s.
> + * Even if a device supports Gen4+, if the link is currently trained
> + * and operating at Gen1..Gen3 speeds (< 16.0 GT/s), reject margining.
> + */
> + pcie_capability_read_word(dev, PCI_EXP_LNKSTA, &lnksta);
> + if ((lnksta & PCI_EXP_LNKSTA_CLS) < PCI_EXP_LNKSTA_CLS_16_0GB) {
> + ret = -EOPNOTSUPP;
> + goto err_rpm;
> + }
> +
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_CAP, &cap);
> + if (ret != PCIBIOS_SUCCESSFUL) {
> + ret = pcibios_err_to_errno(ret);
> + goto err_rpm;
> + }
> +
> + /* Disable ASPM L0s/L1 during margining with restoration path */
> + pci_lmr_disable_aspm(mdev);
What about the other steps besides ASPM that the spec required to be
disabled during Lane Margining?
> + /* Ensure link is settled in L0 mode per PCIe r6.0 sec 8.4.4 */
> + usleep_range(2000, 3000);
> +
> + if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
> + if (ret != PCIBIOS_SUCCESSFUL) {
> + ret = pcibios_err_to_errno(ret);
> + goto err_aspm;
> + }
> + sts |= PCI_LMR_PORT_STS_SW_READY;
> + pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
> + }
> +
> + timeout = jiffies + msecs_to_jiffies(LMR_ENABLE_TIMEOUT_MS);
> + while (1) {
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
> + if (ret != PCIBIOS_SUCCESSFUL) {
> + ret = pcibios_err_to_errno(ret);
> + goto err_sw_ready;
> + }
> + if (PCI_POSSIBLE_ERROR(sts)) {
> + ret = -ENODEV;
> + goto err_sw_ready;
> + }
> + if (sts & PCI_LMR_PORT_STS_MARGIN_READY)
> + break;
> + if (time_after(jiffies, timeout)) {
> + ret = -ETIMEDOUT;
> + goto err_sw_ready;
> + }
> + usleep_range(LMR_ENABLE_SLEEP_MIN_US, LMR_ENABLE_SLEEP_MAX_US);
> + }
> +
> + /* Cache capabilities for configured receiver on all lanes */
> + for (i = 0; i < mdev->num_lanes; i++) {
> + ret = pci_lmr_cache_rx_info(&mdev->lanes[i], mdev->lanes[i].rx);
> + if (ret)
> + goto err_sw_ready;
> + }
> + mdev->enabled = true;
> + return count;
> +
> +err_sw_ready:
> + if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
> + if (ret == PCIBIOS_SUCCESSFUL) {
> + sts &= ~PCI_LMR_PORT_STS_SW_READY;
> + pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
> + }
> + }
> +err_aspm:
> + pci_lmr_restore_aspm(mdev);
> +err_rpm:
> + pm_runtime_put_sync(&dev->dev);
> + return ret;
> +}
--
i.
^ permalink raw reply [flat|nested] 5+ messages in thread
end of thread, other threads:[~2026-08-25 11:02 UTC | newest]
Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-24 22:38 [PATCH v6] PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support Priyank Rathod
2026-08-24 22:51 ` sashiko-bot
2026-08-25 11:02 ` Ilpo Järvinen
-- strict thread matches above, loose matches on Subject: below --
2026-08-24 22:37 Priyank Rathod
2026-08-24 22:51 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox