Linux Power Management development
 help / color / mirror / Atom feed
From: Bjorn Helgaas <helgaas@kernel.org>
To: Rowan Cramer <rowan.cramer@gmail.com>
Cc: linux-pci@vger.kernel.org,
	Nirmal Patel <nirmal.patel@linux.intel.com>,
	Jonathan Derrick <jonathan.derrick@linux.dev>,
	Nam Cao <namcao@linutronix.de>, Zide Chen <zide.chen@intel.com>,
	Zhang Rui <rui.zhang@intel.com>,
	"Rafael J. Wysocki" <rafael@kernel.org>,
	Daniel Lezcano <daniel.lezcano@kernel.org>,
	Christian Loehle <christian.loehle@arm.com>,
	linux-pm@vger.kernel.org
Subject: Re: NVMe MSI-X completions lost behind VMD across package C-state exit (even PC2) on Arrow Lake client VMD [8086:7d0b]
Date: Tue, 14 Jul 2026 12:33:14 -0500	[thread overview]
Message-ID: <20260714173314.GA1365789@bhelgaas> (raw)
In-Reply-To: <CAPew1NKKWf7rr6eZXp=KooN43CKaTbcc-U9ZkKrBhcXaARTKUQ@mail.gmail.com>

[+cc VMD, cpuidle folks]

On Tue, Jul 14, 2026 at 08:57:17AM -0400, Rowan Cramer wrote:
> Hi,
> 
> I have what looks like a platform/VMD interrupt-delivery bug on a
> new Arrow Lake laptop, isolated with a reliable reproducer and a
> three-arm test matrix. Reporting here because the fault sits below
> the NVMe driver, in the VMD interrupt path, and is invisible unless
> the package is allowed to idle. Happy to test patches or run any
> further diagnostics.
> 
> Summary
> -------
> NVMe MSI-X completion interrupts are intermittently lost when they
> arrive through VMD's interrupt remapper while the CPU package is in
> (or exiting) ANY idle state deeper than C1 — including plain PC2.
> The drive completes the I/O correctly (the kernel's timeout handler
> always finds the completion already in the queue: "timeout,
> completion polled"); only the interrupt vanishes. Holding PM-QoS
> /dev/cpu_dma_latency at 100 us (C1-only idle) eliminates the fault
> completely; nothing else does.
> 
> System
> ------
> - Dell 14 Premium / XPS 14 (DA14250), BIOS 1.11.0 (latest; Dell
>   ships VMD force-enabled, no BIOS toggle on this model)
> - Intel Arrow Lake; VMD controller 8086:7d0b at 0000:00:0e.0
>   (client SKU -> VMD_FEATS_CLIENT, so MSI remapping is always on;
>   no VMD_FEAT_CAN_BYPASS_MSI_REMAP, no module parameter)
> - NVMe: Micron 2650 1TB (DRAM-less/HMB), FW 26500101, at
>   10000:e1:00.0; all queues route via VMD-PCI-MSIX per
>   /proc/interrupts
> - Fault is kernel-independent: reproduced identically on 6.18.33
>   (LTS), 7.0.9 and 7.0.10 (Arch builds, clean cmdline)
> - intel_idle in ACPI mode; exposed states POLL / C1_ACPI (1 us) /
>   C2_ACPI (127 us) / C3_ACPI (1048 us)
> 
> Symptom
> -------
>   nvme nvme0: I/O tag N QID Q timeout, completion polled
> 
> Present since the machine's very first boot (3 events in the first
> 90 s of its life). Baseline ~125 events/day idle; 60-190/hr under
> gaming I/O. Each event stalls the waiting I/O for io_timeout
> seconds (we run nvme_core.io_timeout=10 as mitigation), which
> manifests as multi-second system freezes. SMART is immaculate: 0
> media errors, 0 controller error-log entries against >500 host-side
> timeouts — the device completes the I/O; the notification is lost.
> 
> Key evidence
> ------------
> 1. Load is NOT the trigger. ~4 TB of worst-case saturating stress
>    (random reads at 2+ GiB/s, plus CPU, memory-bandwidth, GPU and
>    fsync-write pressure, combined) produced ZERO events. Saturating
>    I/O never lets the platform idle.
> 
> 2. Idle->active transitions ARE the trigger. A trivial reproducer
>    (bursty random reads separated by 1-8 s idle gaps, script below)
>    produces 13-16 lost interrupts per 15 min, every run.
> 
> 3. PM-QoS bisection puts the loss boundary exactly at C1|C2:
>      /dev/cpu_dma_latency hold   idle allowed        events/15min
>      none (stock)                C1+C2+C3            13-16
>      200 us                      C1+C2               19
>      100 us                      C1 only             0
>      0 us                        none (poll)         0
> 
> 4. Ruled out: drive APST (nvme set-feature 0x0c=0: no change), link
>    ASPM (runtime-disabled via sysfs on the VMD-domain link: no
>    change; note cmdline pcie_aspm=off does not govern VMD-managed
>    links, the vmd driver enables L1.1/L1.2 itself per
>    VMD_FEAT_BIOS_PM_QUIRK), kernel version, drive health, firmware
>    updates (none available).
> 
> 5. turbostat correlation (5 s intervals alongside the reproducer):
>    the package NEVER enters PC8/PC10 on this platform under bursty
>    I/O — deepest observed is PC6 — and every lost interrupt falls
>    in intervals showing only PC2/PC3/PC6 residency:
> 
>      arm (900 s each)        Pkg%pc2/pc3/pc6/pc8/pc10   events
>      stock                   13.2 / 0.3 / 4.8 / 0 / 0     13
>      endpoint LTR clamped    19.8 /  0  /  0  / 0 / 0     14
>        to 102.4 us
>      cpu_dma_latency=100us    ~0  /  0  /  0  / 0 / 0      0
> 
> 6. LTR is irrelevant. Firmware sets the endpoint LTR to 15.7 ms
>    (0x100f100f — nonzero, so vmd's BIOS PM quirk never touches it).
>    Clamping it to 102.4 us via setpci was demonstrably honored by
>    the platform (PC3/PC6 residency dropped to zero) yet the loss
>    rate was unchanged (14 vs 13). The fault does not need deep
>    package states — losses occur with nothing deeper than PC2, with
>    cores in C6/C7.
> 
> Current mitigation
> ------------------
> A systemd unit holding /dev/cpu_dma_latency at 100 us. Fully
> effective (a full day of gaming: 1 event vs the prior 60-190/hr),
> but it costs ~2.5-4 W package power versus letting the machine
> idle properly, since it must forbid core C6/C7 and PC2 everywhere.
> 
> Questions
> ---------
> 1. Is this a known Arrow Lake (or MTL-family) VMD erratum? The
>    machine shipped this way; Windows presumably masks it via its
>    own idle policy or is equally affected but silent.
> 
> 2. The client feature set forces MSI remapping on 8086:7d0b. Is the
>    VMCONFIG_MSI_REMAP bypass (as used for 28C0) architecturally
>    possible on client VMD? I am happy to test a patch enabling
>    bypass on 7d0b, or any other diagnostic/experimental patch —
>    the reproducer gives a decisive answer in 15 minutes.
> 
> 3. If the hardware genuinely cannot deliver remapped MSIs across
>    package C-state exit, should vmd.c be constraining PM (PM-QoS /
>    LTR / DMI-quirk) on affected platforms? As it stands, every
>    affected laptop loses NVMe interrupts silently at its default
>    idle settings.
> 
> Full journals since the machine's first boot, lspci -vvv, SMART
> dumps, raw turbostat logs and the complete phase-by-phase test
> history are preserved and available on request.

Can you make these available somewhere and include a URL?

> Reproducer
> ----------
> Point ROOT at any directory with a few multi-hundred-MB files on
> the VMD-attached NVMe, run for 900 s, and count kernel "completion
> polled" events (nvme_core.io_timeout=10 makes them visible within
> 10 s):
> 
>   #!/usr/bin/env python3
>   # bursty-I/O reproducer: idle gaps let the platform enter package
>   # idle; the first read of each burst rides the wake-up.
>   import os, sys, random, time
> 
>   ROOT = os.path.expanduser("~/BG3_Game")   # adjust
>   CHUNK = 65536
>   MIN_SIZE = 200 * 1024 * 1024
> 
>   files = []
>   for dirpath, _, names in os.walk(ROOT, followlinks=True):
>       for n in names:
>           p = os.path.join(dirpath, n)
>           try:
>               s = os.path.getsize(p)
>               if s >= MIN_SIZE:
>                   files.append((p, s))
>           except OSError:
>               pass
>   if not files:
>       sys.exit("no large files under " + ROOT)
> 
>   deadline = time.time() + int(sys.argv[1])
>   bursts = 0
>   fds = {}
>   while time.time() < deadline:
>       time.sleep(random.uniform(1.0, 8.0))
>       for _ in range(random.randint(4, 40)):
>           path, size = random.choice(files)
>           fd = fds.get(path)
>           if fd is None:
>               fd = fds[path] = os.open(path, os.O_RDONLY)
>           off = random.randrange(0, max(1, size - CHUNK)) & ~4095
>           os.pread(fd, CHUNK, off)
>       bursts += 1
>   print(f"bursts={bursts}")
> 
> (Drop page caches before each run — echo 3 >
> /proc/sys/vm/drop_caches — or cached reads will suppress the rate.)
> 
> Thanks,
> Rowan Cramer

       reply	other threads:[~2026-07-14 17:33 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <CAPew1NKKWf7rr6eZXp=KooN43CKaTbcc-U9ZkKrBhcXaARTKUQ@mail.gmail.com>
2026-07-14 17:33 ` Bjorn Helgaas [this message]
2026-07-15 22:09   ` NVMe MSI-X completions lost behind VMD across package C-state exit (even PC2) on Arrow Lake client VMD [8086:7d0b] Rowan Cramer

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260714173314.GA1365789@bhelgaas \
    --to=helgaas@kernel.org \
    --cc=christian.loehle@arm.com \
    --cc=daniel.lezcano@kernel.org \
    --cc=jonathan.derrick@linux.dev \
    --cc=linux-pci@vger.kernel.org \
    --cc=linux-pm@vger.kernel.org \
    --cc=namcao@linutronix.de \
    --cc=nirmal.patel@linux.intel.com \
    --cc=rafael@kernel.org \
    --cc=rowan.cramer@gmail.com \
    --cc=rui.zhang@intel.com \
    --cc=zide.chen@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox