All of lore.kernel.org
 help / color / mirror / Atom feed
From: Rick Warner <rick@microway.com>
To: helgaas@kernel.org
Cc: ilpo.jarvinen@linux.intel.com, linux-pci@vger.kernel.org,
	linux-kernel@vger.kernel.org, rhoughton@microway.com,
	ssingh@microway.com, rick@microway.com, lukas@wunner.de
Subject: [PATCH v2] PCI: Fix Intel Xeon 6 x2 quirk collateral damage on adjacent x4 endpoints
Date: Tue,  4 Aug 2026 16:48:45 -0400	[thread overview]
Message-ID: <20260804204845.171483-1-rick@microway.com> (raw)

commit a22250fe933d ("PCI: Add Extended Tag + MRRS quirk for Xeon 6")
introduced a traffic mitigation quirk for Intel Xeon 6 root ports that
negotiate down to an x2 lane width. However, that patch manipulated host
bridge properties globally via 'bridge->no_ext_tags = 1' and by hooking
'bridge->enable_device'.

Because multiple unrelated root ports can reside under the exact same
global pci_host_bridge domain structure, this overly aggressive mitigation
causes severe collateral damage. When a low-speed or bifurcated secondary
device (such as an onboard ASMedia SATA controller or BMC graphics link)
matches the x2 condition, the kernel strips away Extended Tags and forces
a 128B MRRS restriction across that ENTIRE host bridge. This instantly
breaks or starves adjacent high-performance, unrelated x4 endpoints
(such as NVMe drives), resulting in controller timeouts, initialization
failures, and missing drives at boot.

Fix this by refactoring the quirk logic to be completely per-device and
downstream-isolated. Remove the broad host bridge no_ext_tags and
enable_device = limit_mrrs_to_128 assignments, instead setting only
no_inc_mrrs = 1 to prevent increases from vfio, VM, etc usage.

DECLARE_PCI_FIXUP_HEADER then registers pci_xeon6_x2_local_endpoint_fixup
to check every pci device as it initializes to see if it is downstream of
an affected x2 port, and if so, the device has extended tags disabled and
mrrs set to 128. This enforces Intel's stability parameters locally on the
bottlenecked lane branches while fully protecting the parallel x4 channels.

Fixes: a22250fe933d ("PCI: Add Extended Tag + MRRS quirk for Xeon 6")
Signed-off-by: Rick Warner <rick@microway.com>
---
v2
Original submission did a single pci bus walk and didn't handle hot
plugged devices.  This has been rewritten to handle the quirk per
device during initialization by checking if it's a descendant of a x2 port.

This was tested successfully on a Gigabyte MS74-HB0 motherboard with a
Seagate ZP4000GM30063 M.2 drive in the 2nd slot. With the stock 7.0 kernel,
the drive fails to initialize and is unavailable once booted. Testing showed
that extended tags (vs mrrs) are the key to getting this drive working.

Concerns were raised about vfio/VM usage being able to re-enable extended 
tags if bridge->no_ext_tags is not set.  I'm not sure how to best address 
that if it's needed. The best option I've come up with is adding
DECLARE_PCI_FIXUP_FINAL calls for 0x0db0-0xdb9 that set no_ext_tags after
the initial pci bus walk is done. That would enable the already 
initialized devices to maintain their configured extended tag support 
but would block all newly hotplugged devices on that bridge from using
extended tags, even if they weren't downstream of a x2 port. That might be
good enough, as it would at least solve this issue for NVME drives not
initializing. Otherwise I think bigger changes might be required to 
handle it more cleanly.

 arch/x86/pci/fixup.c | 58 +++++++++++++++++++++++++++++++++++---------
 1 file changed, 46 insertions(+), 12 deletions(-)

diff --git a/arch/x86/pci/fixup.c b/arch/x86/pci/fixup.c
index b301c6c8df75..0ab92a5cc0d1 100644
--- a/arch/x86/pci/fixup.c
+++ b/arch/x86/pci/fixup.c
@@ -301,15 +301,6 @@ DECLARE_PCI_FIXUP_FINAL(PCI_VENDOR_ID_INTEL,	PCI_DEVICE_ID_INTEL_MCH_PC1,	pcie_r
  *
  * https://cdrdv2.intel.com/v1/dl/getContent/837176
  */
-static int limit_mrrs_to_128(struct pci_host_bridge *b, struct pci_dev *pdev)
-{
-	int readrq = pcie_get_readrq(pdev);
-
-	if (readrq > 128)
-		pcie_set_readrq(pdev, 128);
-
-	return 0;
-}
 
 static void pci_xeon_x2_bifurc_quirk(struct pci_dev *pdev)
 {
@@ -320,9 +311,8 @@ static void pci_xeon_x2_bifurc_quirk(struct pci_dev *pdev)
 	if (FIELD_GET(PCI_EXP_LNKCAP_MLW, linkcap) != 0x2)
 		return;
 
-	bridge->no_ext_tags = 1;
-	bridge->enable_device = limit_mrrs_to_128;
-	pci_info(pdev, "Disabling Extended Tags and limiting MRRS to 128B (performance reasons due to x2 PCIe link)\n");
+	bridge->no_inc_mrrs = 1;
+	pci_info(pdev, "Blocking devices on this bridge from increasing MRRS for performance reasons due to x2 PCIe link)\n");
 }
 
 DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x0db0, pci_xeon_x2_bifurc_quirk);
@@ -334,6 +324,50 @@ DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x0db7, pci_xeon_x2_bifurc_quirk);
 DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x0db8, pci_xeon_x2_bifurc_quirk);
 DECLARE_PCI_FIXUP_EARLY(PCI_VENDOR_ID_INTEL, 0x0db9, pci_xeon_x2_bifurc_quirk);
 
+/* Helper to check if a device descends from an affected Xeon 6 x2 Root Port */
+static bool is_descendant_of_xeon6_x2_rp(struct pci_dev *pdev)
+{
+	u32 linkcap;
+	struct pci_dev *upstream = pci_upstream_bridge(pdev);
+
+	while (upstream) {
+		if (upstream->vendor == PCI_VENDOR_ID_INTEL &&
+		    upstream->device >= 0x0db0 &&
+		    upstream->device <= 0x0db9) {
+			pcie_capability_read_dword(upstream, PCI_EXP_LNKCAP, &linkcap);
+			if (FIELD_GET(PCI_EXP_LNKCAP_MLW, linkcap) != 0x2)
+				return false;
+			else
+				return true;
+		}
+		upstream = pci_upstream_bridge(upstream);
+	}
+	return false;
+}
+
+static void pci_xeon6_x2_local_endpoint_fixup(struct pci_dev *pdev)
+{
+	/* Skip bridges/switches; only target actual endpoints */
+	if (pci_is_bridge(pdev))
+		return;
+
+	/* Only apply to devices under the x2 branch; leaves x4 branches completely untouched */
+	if (!is_descendant_of_xeon6_x2_rp(pdev))
+		return;
+
+	pci_info(pdev, "Applying local Xeon 6 x2 quirk: Disabling Extended Tags and locking MRRS to 128B\n");
+
+	pcie_capability_clear_word(pdev, PCI_EXP_DEVCTL, PCI_EXP_DEVCTL_EXT_TAG);
+	pcie_capability_clear_word(pdev, PCI_EXP_DEVCTL, PCI_EXP_DEVCTL_READRQ);
+}
+
+/*
+ * HEADER fixups run for EVERY endpoint during its initial discovery phase.
+ * This natively catches boot devices, hotplugged devices, and SR-IOV VFs.
+ */
+DECLARE_PCI_FIXUP_HEADER(PCI_ANY_ID, PCI_ANY_ID, pci_xeon6_x2_local_endpoint_fixup);
+
+
 /*
  * Fixup to mark boot BIOS video selected by BIOS before it changes
  *
-- 
2.52.0


             reply	other threads:[~2026-08-04 20:48 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-04 20:48 Rick Warner [this message]
2026-08-04 21:05 ` [PATCH v2] PCI: Fix Intel Xeon 6 x2 quirk collateral damage on adjacent x4 endpoints sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260804204845.171483-1-rick@microway.com \
    --to=rick@microway.com \
    --cc=helgaas@kernel.org \
    --cc=ilpo.jarvinen@linux.intel.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=lukas@wunner.de \
    --cc=rhoughton@microway.com \
    --cc=ssingh@microway.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.