* Re: NVMe MSI-X completions lost behind VMD across package C-state exit (even PC2) on Arrow Lake client VMD [8086:7d0b]
[not found] <CAPew1NKKWf7rr6eZXp=KooN43CKaTbcc-U9ZkKrBhcXaARTKUQ@mail.gmail.com>
@ 2026-07-14 17:33 ` Bjorn Helgaas
2026-07-15 22:09 ` Rowan Cramer
0 siblings, 1 reply; 2+ messages in thread
From: Bjorn Helgaas @ 2026-07-14 17:33 UTC (permalink / raw)
To: Rowan Cramer
Cc: linux-pci, Nirmal Patel, Jonathan Derrick, Nam Cao, Zide Chen,
Zhang Rui, Rafael J. Wysocki, Daniel Lezcano, Christian Loehle,
linux-pm
[+cc VMD, cpuidle folks]
On Tue, Jul 14, 2026 at 08:57:17AM -0400, Rowan Cramer wrote:
> Hi,
>
> I have what looks like a platform/VMD interrupt-delivery bug on a
> new Arrow Lake laptop, isolated with a reliable reproducer and a
> three-arm test matrix. Reporting here because the fault sits below
> the NVMe driver, in the VMD interrupt path, and is invisible unless
> the package is allowed to idle. Happy to test patches or run any
> further diagnostics.
>
> Summary
> -------
> NVMe MSI-X completion interrupts are intermittently lost when they
> arrive through VMD's interrupt remapper while the CPU package is in
> (or exiting) ANY idle state deeper than C1 — including plain PC2.
> The drive completes the I/O correctly (the kernel's timeout handler
> always finds the completion already in the queue: "timeout,
> completion polled"); only the interrupt vanishes. Holding PM-QoS
> /dev/cpu_dma_latency at 100 us (C1-only idle) eliminates the fault
> completely; nothing else does.
>
> System
> ------
> - Dell 14 Premium / XPS 14 (DA14250), BIOS 1.11.0 (latest; Dell
> ships VMD force-enabled, no BIOS toggle on this model)
> - Intel Arrow Lake; VMD controller 8086:7d0b at 0000:00:0e.0
> (client SKU -> VMD_FEATS_CLIENT, so MSI remapping is always on;
> no VMD_FEAT_CAN_BYPASS_MSI_REMAP, no module parameter)
> - NVMe: Micron 2650 1TB (DRAM-less/HMB), FW 26500101, at
> 10000:e1:00.0; all queues route via VMD-PCI-MSIX per
> /proc/interrupts
> - Fault is kernel-independent: reproduced identically on 6.18.33
> (LTS), 7.0.9 and 7.0.10 (Arch builds, clean cmdline)
> - intel_idle in ACPI mode; exposed states POLL / C1_ACPI (1 us) /
> C2_ACPI (127 us) / C3_ACPI (1048 us)
>
> Symptom
> -------
> nvme nvme0: I/O tag N QID Q timeout, completion polled
>
> Present since the machine's very first boot (3 events in the first
> 90 s of its life). Baseline ~125 events/day idle; 60-190/hr under
> gaming I/O. Each event stalls the waiting I/O for io_timeout
> seconds (we run nvme_core.io_timeout=10 as mitigation), which
> manifests as multi-second system freezes. SMART is immaculate: 0
> media errors, 0 controller error-log entries against >500 host-side
> timeouts — the device completes the I/O; the notification is lost.
>
> Key evidence
> ------------
> 1. Load is NOT the trigger. ~4 TB of worst-case saturating stress
> (random reads at 2+ GiB/s, plus CPU, memory-bandwidth, GPU and
> fsync-write pressure, combined) produced ZERO events. Saturating
> I/O never lets the platform idle.
>
> 2. Idle->active transitions ARE the trigger. A trivial reproducer
> (bursty random reads separated by 1-8 s idle gaps, script below)
> produces 13-16 lost interrupts per 15 min, every run.
>
> 3. PM-QoS bisection puts the loss boundary exactly at C1|C2:
> /dev/cpu_dma_latency hold idle allowed events/15min
> none (stock) C1+C2+C3 13-16
> 200 us C1+C2 19
> 100 us C1 only 0
> 0 us none (poll) 0
>
> 4. Ruled out: drive APST (nvme set-feature 0x0c=0: no change), link
> ASPM (runtime-disabled via sysfs on the VMD-domain link: no
> change; note cmdline pcie_aspm=off does not govern VMD-managed
> links, the vmd driver enables L1.1/L1.2 itself per
> VMD_FEAT_BIOS_PM_QUIRK), kernel version, drive health, firmware
> updates (none available).
>
> 5. turbostat correlation (5 s intervals alongside the reproducer):
> the package NEVER enters PC8/PC10 on this platform under bursty
> I/O — deepest observed is PC6 — and every lost interrupt falls
> in intervals showing only PC2/PC3/PC6 residency:
>
> arm (900 s each) Pkg%pc2/pc3/pc6/pc8/pc10 events
> stock 13.2 / 0.3 / 4.8 / 0 / 0 13
> endpoint LTR clamped 19.8 / 0 / 0 / 0 / 0 14
> to 102.4 us
> cpu_dma_latency=100us ~0 / 0 / 0 / 0 / 0 0
>
> 6. LTR is irrelevant. Firmware sets the endpoint LTR to 15.7 ms
> (0x100f100f — nonzero, so vmd's BIOS PM quirk never touches it).
> Clamping it to 102.4 us via setpci was demonstrably honored by
> the platform (PC3/PC6 residency dropped to zero) yet the loss
> rate was unchanged (14 vs 13). The fault does not need deep
> package states — losses occur with nothing deeper than PC2, with
> cores in C6/C7.
>
> Current mitigation
> ------------------
> A systemd unit holding /dev/cpu_dma_latency at 100 us. Fully
> effective (a full day of gaming: 1 event vs the prior 60-190/hr),
> but it costs ~2.5-4 W package power versus letting the machine
> idle properly, since it must forbid core C6/C7 and PC2 everywhere.
>
> Questions
> ---------
> 1. Is this a known Arrow Lake (or MTL-family) VMD erratum? The
> machine shipped this way; Windows presumably masks it via its
> own idle policy or is equally affected but silent.
>
> 2. The client feature set forces MSI remapping on 8086:7d0b. Is the
> VMCONFIG_MSI_REMAP bypass (as used for 28C0) architecturally
> possible on client VMD? I am happy to test a patch enabling
> bypass on 7d0b, or any other diagnostic/experimental patch —
> the reproducer gives a decisive answer in 15 minutes.
>
> 3. If the hardware genuinely cannot deliver remapped MSIs across
> package C-state exit, should vmd.c be constraining PM (PM-QoS /
> LTR / DMI-quirk) on affected platforms? As it stands, every
> affected laptop loses NVMe interrupts silently at its default
> idle settings.
>
> Full journals since the machine's first boot, lspci -vvv, SMART
> dumps, raw turbostat logs and the complete phase-by-phase test
> history are preserved and available on request.
Can you make these available somewhere and include a URL?
> Reproducer
> ----------
> Point ROOT at any directory with a few multi-hundred-MB files on
> the VMD-attached NVMe, run for 900 s, and count kernel "completion
> polled" events (nvme_core.io_timeout=10 makes them visible within
> 10 s):
>
> #!/usr/bin/env python3
> # bursty-I/O reproducer: idle gaps let the platform enter package
> # idle; the first read of each burst rides the wake-up.
> import os, sys, random, time
>
> ROOT = os.path.expanduser("~/BG3_Game") # adjust
> CHUNK = 65536
> MIN_SIZE = 200 * 1024 * 1024
>
> files = []
> for dirpath, _, names in os.walk(ROOT, followlinks=True):
> for n in names:
> p = os.path.join(dirpath, n)
> try:
> s = os.path.getsize(p)
> if s >= MIN_SIZE:
> files.append((p, s))
> except OSError:
> pass
> if not files:
> sys.exit("no large files under " + ROOT)
>
> deadline = time.time() + int(sys.argv[1])
> bursts = 0
> fds = {}
> while time.time() < deadline:
> time.sleep(random.uniform(1.0, 8.0))
> for _ in range(random.randint(4, 40)):
> path, size = random.choice(files)
> fd = fds.get(path)
> if fd is None:
> fd = fds[path] = os.open(path, os.O_RDONLY)
> off = random.randrange(0, max(1, size - CHUNK)) & ~4095
> os.pread(fd, CHUNK, off)
> bursts += 1
> print(f"bursts={bursts}")
>
> (Drop page caches before each run — echo 3 >
> /proc/sys/vm/drop_caches — or cached reads will suppress the rate.)
>
> Thanks,
> Rowan Cramer
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: NVMe MSI-X completions lost behind VMD across package C-state exit (even PC2) on Arrow Lake client VMD [8086:7d0b]
2026-07-14 17:33 ` NVMe MSI-X completions lost behind VMD across package C-state exit (even PC2) on Arrow Lake client VMD [8086:7d0b] Bjorn Helgaas
@ 2026-07-15 22:09 ` Rowan Cramer
0 siblings, 0 replies; 2+ messages in thread
From: Rowan Cramer @ 2026-07-15 22:09 UTC (permalink / raw)
To: Bjorn Helgaas
Cc: linux-pci, Nirmal Patel, Jonathan Derrick, Nam Cao, Zide Chen,
Zhang Rui, Rafael J. Wysocki, Daniel Lezcano, Christian Loehle,
linux-pm
Done:
https://github.com/Laif-added/nvme-vmd-interrupt-loss
Contents (README.md maps each file to the specific claims in my
report):
- logs/journal-kernel-all-boots.txt — complete kernel log of every
boot since the machine's first power-on (2026-06-26 onward, ~71k
lines; 9 boots — a 10th, 3 s aborted first-power-on boot logged
no kernel messages), including the 3 losses in the first ~90 s
of its life
- logs/timeout-counts-per-boot.txt — per-boot event counts with
kernel versions (6.18.33 / 7.0.9 / 7.0.10)
- logs/turbostat-test-*.log — raw 5 s turbostat for the PC-state
correlation and LTR-clamp arms
- hw/lspci-vvv.txt — full dump incl. the VMD domain (10000:), plus
the raw LTR register readout (0x100f100f)
- hw/smartctl-a.txt, hw/nvme-error-log.txt — 0 media errors, 0
controller error-log entries (serials redacted)
- report/test-history.txt — raw phase-by-phase results P0-P12 +
TEST A/B
- scripts/ — the reproducer (nvme-irq-bursty.py) and every test
driver used
One scope note: the published journals are kernel-transport only
(journalctl _TRANSPORT=kernel) to keep userspace log noise and
personal data out of a public repo. The full journals are preserved
locally; I'm happy to send them to any maintainer off-list if
something beyond the kernel messages is needed.
Standing offer from the original report still applies: the
reproducer gives a decisive answer in 15 minutes, so I can turn
around any diagnostic or experimental patch (e.g. a MSI-remap
bypass attempt on 7d0b) quickly.
Thanks,
Rowan Cramer
On Tue, Jul 14, 2026 at 1:33 PM Bjorn Helgaas <helgaas@kernel.org> wrote:
>
> [+cc VMD, cpuidle folks]
>
> On Tue, Jul 14, 2026 at 08:57:17AM -0400, Rowan Cramer wrote:
> > Hi,
> >
> > I have what looks like a platform/VMD interrupt-delivery bug on a
> > new Arrow Lake laptop, isolated with a reliable reproducer and a
> > three-arm test matrix. Reporting here because the fault sits below
> > the NVMe driver, in the VMD interrupt path, and is invisible unless
> > the package is allowed to idle. Happy to test patches or run any
> > further diagnostics.
> >
> > Summary
> > -------
> > NVMe MSI-X completion interrupts are intermittently lost when they
> > arrive through VMD's interrupt remapper while the CPU package is in
> > (or exiting) ANY idle state deeper than C1 — including plain PC2.
> > The drive completes the I/O correctly (the kernel's timeout handler
> > always finds the completion already in the queue: "timeout,
> > completion polled"); only the interrupt vanishes. Holding PM-QoS
> > /dev/cpu_dma_latency at 100 us (C1-only idle) eliminates the fault
> > completely; nothing else does.
> >
> > System
> > ------
> > - Dell 14 Premium / XPS 14 (DA14250), BIOS 1.11.0 (latest; Dell
> > ships VMD force-enabled, no BIOS toggle on this model)
> > - Intel Arrow Lake; VMD controller 8086:7d0b at 0000:00:0e.0
> > (client SKU -> VMD_FEATS_CLIENT, so MSI remapping is always on;
> > no VMD_FEAT_CAN_BYPASS_MSI_REMAP, no module parameter)
> > - NVMe: Micron 2650 1TB (DRAM-less/HMB), FW 26500101, at
> > 10000:e1:00.0; all queues route via VMD-PCI-MSIX per
> > /proc/interrupts
> > - Fault is kernel-independent: reproduced identically on 6.18.33
> > (LTS), 7.0.9 and 7.0.10 (Arch builds, clean cmdline)
> > - intel_idle in ACPI mode; exposed states POLL / C1_ACPI (1 us) /
> > C2_ACPI (127 us) / C3_ACPI (1048 us)
> >
> > Symptom
> > -------
> > nvme nvme0: I/O tag N QID Q timeout, completion polled
> >
> > Present since the machine's very first boot (3 events in the first
> > 90 s of its life). Baseline ~125 events/day idle; 60-190/hr under
> > gaming I/O. Each event stalls the waiting I/O for io_timeout
> > seconds (we run nvme_core.io_timeout=10 as mitigation), which
> > manifests as multi-second system freezes. SMART is immaculate: 0
> > media errors, 0 controller error-log entries against >500 host-side
> > timeouts — the device completes the I/O; the notification is lost.
> >
> > Key evidence
> > ------------
> > 1. Load is NOT the trigger. ~4 TB of worst-case saturating stress
> > (random reads at 2+ GiB/s, plus CPU, memory-bandwidth, GPU and
> > fsync-write pressure, combined) produced ZERO events. Saturating
> > I/O never lets the platform idle.
> >
> > 2. Idle->active transitions ARE the trigger. A trivial reproducer
> > (bursty random reads separated by 1-8 s idle gaps, script below)
> > produces 13-16 lost interrupts per 15 min, every run.
> >
> > 3. PM-QoS bisection puts the loss boundary exactly at C1|C2:
> > /dev/cpu_dma_latency hold idle allowed events/15min
> > none (stock) C1+C2+C3 13-16
> > 200 us C1+C2 19
> > 100 us C1 only 0
> > 0 us none (poll) 0
> >
> > 4. Ruled out: drive APST (nvme set-feature 0x0c=0: no change), link
> > ASPM (runtime-disabled via sysfs on the VMD-domain link: no
> > change; note cmdline pcie_aspm=off does not govern VMD-managed
> > links, the vmd driver enables L1.1/L1.2 itself per
> > VMD_FEAT_BIOS_PM_QUIRK), kernel version, drive health, firmware
> > updates (none available).
> >
> > 5. turbostat correlation (5 s intervals alongside the reproducer):
> > the package NEVER enters PC8/PC10 on this platform under bursty
> > I/O — deepest observed is PC6 — and every lost interrupt falls
> > in intervals showing only PC2/PC3/PC6 residency:
> >
> > arm (900 s each) Pkg%pc2/pc3/pc6/pc8/pc10 events
> > stock 13.2 / 0.3 / 4.8 / 0 / 0 13
> > endpoint LTR clamped 19.8 / 0 / 0 / 0 / 0 14
> > to 102.4 us
> > cpu_dma_latency=100us ~0 / 0 / 0 / 0 / 0 0
> >
> > 6. LTR is irrelevant. Firmware sets the endpoint LTR to 15.7 ms
> > (0x100f100f — nonzero, so vmd's BIOS PM quirk never touches it).
> > Clamping it to 102.4 us via setpci was demonstrably honored by
> > the platform (PC3/PC6 residency dropped to zero) yet the loss
> > rate was unchanged (14 vs 13). The fault does not need deep
> > package states — losses occur with nothing deeper than PC2, with
> > cores in C6/C7.
> >
> > Current mitigation
> > ------------------
> > A systemd unit holding /dev/cpu_dma_latency at 100 us. Fully
> > effective (a full day of gaming: 1 event vs the prior 60-190/hr),
> > but it costs ~2.5-4 W package power versus letting the machine
> > idle properly, since it must forbid core C6/C7 and PC2 everywhere.
> >
> > Questions
> > ---------
> > 1. Is this a known Arrow Lake (or MTL-family) VMD erratum? The
> > machine shipped this way; Windows presumably masks it via its
> > own idle policy or is equally affected but silent.
> >
> > 2. The client feature set forces MSI remapping on 8086:7d0b. Is the
> > VMCONFIG_MSI_REMAP bypass (as used for 28C0) architecturally
> > possible on client VMD? I am happy to test a patch enabling
> > bypass on 7d0b, or any other diagnostic/experimental patch —
> > the reproducer gives a decisive answer in 15 minutes.
> >
> > 3. If the hardware genuinely cannot deliver remapped MSIs across
> > package C-state exit, should vmd.c be constraining PM (PM-QoS /
> > LTR / DMI-quirk) on affected platforms? As it stands, every
> > affected laptop loses NVMe interrupts silently at its default
> > idle settings.
> >
> > Full journals since the machine's first boot, lspci -vvv, SMART
> > dumps, raw turbostat logs and the complete phase-by-phase test
> > history are preserved and available on request.
>
> Can you make these available somewhere and include a URL?
>
> > Reproducer
> > ----------
> > Point ROOT at any directory with a few multi-hundred-MB files on
> > the VMD-attached NVMe, run for 900 s, and count kernel "completion
> > polled" events (nvme_core.io_timeout=10 makes them visible within
> > 10 s):
> >
> > #!/usr/bin/env python3
> > # bursty-I/O reproducer: idle gaps let the platform enter package
> > # idle; the first read of each burst rides the wake-up.
> > import os, sys, random, time
> >
> > ROOT = os.path.expanduser("~/BG3_Game") # adjust
> > CHUNK = 65536
> > MIN_SIZE = 200 * 1024 * 1024
> >
> > files = []
> > for dirpath, _, names in os.walk(ROOT, followlinks=True):
> > for n in names:
> > p = os.path.join(dirpath, n)
> > try:
> > s = os.path.getsize(p)
> > if s >= MIN_SIZE:
> > files.append((p, s))
> > except OSError:
> > pass
> > if not files:
> > sys.exit("no large files under " + ROOT)
> >
> > deadline = time.time() + int(sys.argv[1])
> > bursts = 0
> > fds = {}
> > while time.time() < deadline:
> > time.sleep(random.uniform(1.0, 8.0))
> > for _ in range(random.randint(4, 40)):
> > path, size = random.choice(files)
> > fd = fds.get(path)
> > if fd is None:
> > fd = fds[path] = os.open(path, os.O_RDONLY)
> > off = random.randrange(0, max(1, size - CHUNK)) & ~4095
> > os.pread(fd, CHUNK, off)
> > bursts += 1
> > print(f"bursts={bursts}")
> >
> > (Drop page caches before each run — echo 3 >
> > /proc/sys/vm/drop_caches — or cached reads will suppress the rate.)
> >
> > Thanks,
> > Rowan Cramer
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-07-15 22:09 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <CAPew1NKKWf7rr6eZXp=KooN43CKaTbcc-U9ZkKrBhcXaARTKUQ@mail.gmail.com>
2026-07-14 17:33 ` NVMe MSI-X completions lost behind VMD across package C-state exit (even PC2) on Arrow Lake client VMD [8086:7d0b] Bjorn Helgaas
2026-07-15 22:09 ` Rowan Cramer
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox