Linux PCI subsystem development
 help / color / mirror / Atom feed
* [BUG] ALSA: hda: AE-7 behind ASM1083 rev 03 bridge: fatal MCE at probe
@ 2026-10-05 12:47 Munteanu Vlad
  2026-10-05 13:30 ` Takashi Iwai
  0 siblings, 1 reply; 2+ messages in thread
From: Munteanu Vlad @ 2026-10-05 12:47 UTC (permalink / raw)
  To: tiwai, linux-sound; +Cc: perex, conmanx360, bhelgaas, linux-pci

[-- Attachment #1: Type: text/plain, Size: 7495 bytes --]

Hi,

Newer Sound Blaster AE-7 cards (CA0132, PCI 1102:0010, SSID 1102:0081)
put the CA0132 behind an ASMedia ASM1083/1085 rev 03 PCIe-to-PCI bridge.
On my machine, binding snd_hda_intel to the card takes the whole system
down every time: a fatal machine check (or a silent hard hang) within
about a second of probe. This looks like the problem other AE-7 owners
report, where the card "never worked" on Linux [1].

I tracked it down to MMIO reads to the CA0132 that never complete. I
have a workaround (attached, not meant for merging as-is) with which the
card now works fully: probe, DSP firmware download, playback. I'd like
your advice on the proper fix, and I'm happy to test patches.


Hardware / software
-------------------

- Dell Precision 7920 Tower, 2x Xeon Gold 6152 (Skylake-SP),
BIOS 2.53.1, APEI firmware-first error handling.
- 44:00.0 Intel Sky Lake-E PCIe Root Port A [8086:2030]
-> 45:00.0 ASMedia ASM1083/1085 PCIe-to-PCI bridge [1b21:1080] rev 03
-> 46:00.0 Creative CA0132 [1102:0010] rev 01, subsystem [1102:0081]
- Ubuntu kernel 7.0.0-38-generic (based on 7.0.14). The code paths
involved (sound/hda/core/controller.c, core/stream.c,
codecs/ca0132.c) are the same in current mainline as far as I can
see. I haven't built mainline yet, but can if that helps. The same
hang happened with Ubuntu's 6.17 kernel at install time.


1. The failure
--------------

Identical in every run (CPU PPIN removed):

mce: [Hardware Error]: CPU 0: Machine Check Exception: 5 Bank 6:
bb80000000000e0b
mce: [Hardware Error]: RIP !INEXACT! 10:<ffffffffc1fdbf1b>
{snd_hdac_bus_init_cmd_io+0x1db/0x260 [snd_hda_core]}
mce: [Hardware Error]: TSC 2e59a247c66 MISC 44000000
mce: [Hardware Error]: PROCESSOR 0:50654 TIME 1791056639 SOCKET 0 APIC
0 microcode 2007006
mce: [Hardware Error]: Machine check: Processor context corrupt
Kernel panic - not syncing: Fatal machine check

Bank 6 is the IIO, and MCACOD 0x0e0b is a generic I/O bus error. The MCE
comes about 1.18 s after the last driver message ("codec_mask = 0x2"),
which matches the root port's completion timeout (260-900 ms). So a CPU
read to the CA0132 never got a completion.

In this build, +0x1db is the instruction right after
"mov 0x4a(%rax),%ax", i.e. the readw(CORBRP) in the first poll loop of
azx_clear_corbrp(), immediately after writew(CORBRP, AZX_CORBRP_RST).
CORB DMA isn't running yet at that point.


2. What the experiments showed
------------------------------

- Done from userspace with the card unbound (setpci / devmem), each
register access is fine on its own. That includes the controller reset
at driver timing, and the whole CORB/RIRB setup sequence done slowly
with a read after each step. Note that the CORB/RIRB base addresses
were 0 in that test, so no real DMA happened, and the IOMMU logged no
faults.

- In the driver, with the CORBRP readback avoided (write RST, wait
10 ms, write 0, wait 10 ms, no reads in between), the MCE moved to
snd_hdac_bus_init_cmd_io+0x186. That's readl(GCTL) in the final
updatel(GCTL, UNSOL): the first read after CORBCTL=RUN, the RIRB
base/size writes, RIRBWP=RST, RINTCNT and RIRBCTL. Adding 10 ms delays
between those writes did NOT help. Adding a GCTL read after each
group of writes did.

- With that in place, the probe got through codec enumeration and then
hung (no MCE record that time). The last trace point was
ca0132_mmio_init() returning. The next code is ae5_register_set() (the
AE-7 path): 19 BAR2 writes, then the first BAR2 read, in
ca0113_mmio_command_set_type2(). So presumably it was that read.

- Turning off AER/SERR/parity reporting on the card, the bridge and the
root port did not help either. The machine hung at the same point, but
left no MCE record and had to be power-cycled.

So the pattern seems to be: on this bridge, a read that follows a run
of posted writes to the CA0132 never completes. One data point doesn't
fit a simple count, though. azx_int_clear() does 13 writes to
registers 32 bytes apart, and the read after it was fine. So it may be
about writes to neighbouring registers being merged by the bridge. I
couldn't pin that down. ASM1083 rev 03 is reported as broken with other
PCI cards as well [2].


3. Workaround that works (attached, against the Ubuntu 7.0.0-38 tree)
--------------------------------------------------------------------

All of it applies only to Creative HDA controllers (PCI vendor check):

a) snd_hdac_reg_write{b,w,l}(): read GCTL after every register write.
b) snd_hdac_bus_init_cmd_io(): do the CORB read pointer reset without
polling CORBRP while RST is set (set, 10 ms, clear, 10 ms), and
settle + read GCTL after each group of CORB/RIRB writes.
c) snd_hdac_stream_reset(): keep the SRST handshake, but wait 5 ms
before each read-back.
d) ca0132.c: follow every write to spec->mem_base (45 sites) with the
same flush read.

With this everything works over several boots and hours of playback:
controller and codec probe, "ca0132 DSP downloaded and running", the
AE-7 post-DSP setup, and playback in 2.0 and 2.1 (6 channels). There
have been no MCEs since. NVIDIA HDMI audio, driven by the same
snd_hda_intel, is unaffected (the flushes are gated on the vendor).

(c) may not be needed. My first version did the stream reset blind (no
SRST read-back) because I suspected reads during reset. The DMAR faults
I then saw turned out to be the separate issue in section 4.


4. Second issue: the CA0132 reads past the end of the cyclic buffer
-------------------------------------------------------------------

With IOMMU translation (default DMA-FQ domain), every buffer wrap during
playback produces:

DMAR: [DMA Read NO_PASID] Request device [46:00.0] fault addr
0xffec0000 [fault reason 0x06] PTE Read access is not set

The fault repeats at the buffer-wrap period (0.683 s for 32768 frames
at 48 kHz). The address changes with the buffer, and it's consistent
with the address just past the end of a size-aligned IOVA allocation of
the current PCM buffer: 0xffec0000 for a 768 KiB buffer (6 ch, which
would sit at 0xffe00000), and 0xfff40000 / 0xfff80000 for 256 KiB
buffers. I haven't read the buffer's IOVA directly. If that's right,
the controller prefetches past the last BDL entry before wrapping. The
audio itself is fine. As a workaround I switch the card's IOMMU group
to identity before binding it.


5. Minor
--------

PipeWire first asks for buffer=1572864, period=49152, which gives
"Too many BDL entries" (because of AZX_DCAPS_4K_BDLE_BOUNDARY). It then
falls back to a smaller buffer and works.


Questions
---------

- Would a quirk keyed on an ASMedia 1b21:1080 (rev 03) bridge directly
upstream of the controller be acceptable? It would set something like
bus->flush_writes (read back after every write), plus a CORB reset
that doesn't poll while RST is set. Or would you rather this live in
the PCI layer?
- Are there any known ASM1083/1085 errata about posted writes or write
merging?
- For the buffer overread, what's the preferred fix: padding the
allocation, an extra BDL entry, or something else?

I can test patches on this machine. Failures leave crash dumps in
pstore/ERST, so each failed test is quick to diagnose.

[1] https://forum.endeavouros.com/t/unable-to-boot-after-installing-new-sound-card-ae-7/40457
[2] https://projects.osmocom.org/projects/retronetworking/wiki/PCIe-%3EPCI_bridges
[3] https://bugzilla.kernel.org/show_bug.cgi?id=208667 (ASM1083/1085 ASPM quirk)
[4] https://bugzilla.kernel.org/show_bug.cgi?id=217510 (AE-7, possibly related)

Regards,
gigiou_88

[-- Warning: decoded text below may be mangled, UTF-8 assumed --]
[-- Attachment #2: ae7-asm1083-workaround.patch --]
[-- Type: text/x-patch; charset="US-ASCII"; name="ae7-asm1083-workaround.patch", Size: 15022 bytes --]

--- a/include/sound/hdaudio.h
+++ b/include/sound/hdaudio.h
@@ -427,6 +427,24 @@
 #define snd_hdac_aligned_write(val, addr, mask) do {} while (0)
 #endif
 
+/*
+ * ae7 local patch: a Sound Blaster AE-7 (Creative CA0132) behind an ASMedia ASM1083/1085 rev 03 PCIe-to-PCI
+ * bridge loses a register read that follows a run of posted writes: the read never completes, the root port's
+ * completion timeout fires and the CPU takes a fatal machine check. Reading back after every write avoids it,
+ * so on Creative HDA controllers every register write is followed by a read of GCTL (offset 0x08).
+ */
+static inline bool snd_hdac_bus_flush_writes(struct hdac_bus *bus)
+{
+	return bus && bus->dev && dev_is_pci(bus->dev) &&
+	       to_pci_dev(bus->dev)->vendor == PCI_VENDOR_ID_CREATIVE;
+}
+
+static inline void snd_hdac_reg_flush(struct hdac_bus *bus)
+{
+	if (snd_hdac_bus_flush_writes(bus) && bus->remap_addr)
+		readl(bus->remap_addr + 0x08);	/* AZX_REG_GCTL */
+}
+
 static inline void snd_hdac_reg_writeb(struct hdac_bus *bus, void __iomem *addr,
 				       u8 val)
 {
@@ -434,6 +452,7 @@
 		snd_hdac_aligned_write(val, addr, 0xff);
 	else
 		writeb(val, addr);
+	snd_hdac_reg_flush(bus);
 }
 
 static inline void snd_hdac_reg_writew(struct hdac_bus *bus, void __iomem *addr,
@@ -443,6 +462,7 @@
 		snd_hdac_aligned_write(val, addr, 0xffff);
 	else
 		writew(val, addr);
+	snd_hdac_reg_flush(bus);
 }
 
 static inline u8 snd_hdac_reg_readb(struct hdac_bus *bus, void __iomem *addr)
@@ -457,7 +477,12 @@
 		snd_hdac_aligned_read(addr, 0xffff) : readw(addr);
 }
 
-#define snd_hdac_reg_writel(bus, addr, val)	writel(val, addr)
+static inline void snd_hdac_reg_writel(struct hdac_bus *bus, void __iomem *addr,
+				       u32 val)
+{
+	writel(val, addr);
+	snd_hdac_reg_flush(bus);
+}
 #define snd_hdac_reg_readl(bus, addr)	readl(addr)
 #define snd_hdac_reg_writeq(bus, addr, val)	writeq(val, addr)
 #define snd_hdac_reg_readq(bus, addr)		readq(addr)
--- a/sound/hda/core/controller.c
+++ b/sound/hda/core/controller.c
@@ -40,6 +40,52 @@
  * snd_hdac_bus_init_cmd_io - set up CORB/RIRB buffers
  * @bus: HD-audio core bus
  */
+/*
+ * ae7 local patch, Sound Blaster AE-7 behind an ASMedia ASM1083/1085 rev 03 bridge (see include/sound/hdaudio.h):
+ * the stock CORB/RIRB setup hangs the machine twice — reading CORBRP back while CORBRP_RST is set never completes,
+ * and neither does the first read after the back-to-back CORB start / RIRB setup writes. This is the sequence that
+ * was tested to work on that card: let the writes settle and read GCTL after each group, and do the CORB read
+ * pointer reset without reading it back.
+ */
+static void ae7_settle(struct hdac_bus *bus)
+{
+	mdelay(10);
+	readl(bus->remap_addr + AZX_REG_GCTL);
+}
+
+static void ae7_init_cmd_io(struct hdac_bus *bus)
+{
+	void __iomem *regs = bus->remap_addr;
+
+	ae7_settle(bus);
+	writel((u32)(bus->corb.addr + bus->addr_offset), regs + AZX_REG_CORBLBASE);
+	writel(upper_32_bits(bus->corb.addr + bus->addr_offset), regs + AZX_REG_CORBUBASE);
+	writeb(0x02, regs + AZX_REG_CORBSIZE);
+	writew(0, regs + AZX_REG_CORBWP);
+	ae7_settle(bus);
+	/* CORB read pointer reset: never read anything while CORBRP_RST is set */
+	writew(AZX_CORBRP_RST, regs + AZX_REG_CORBRP);
+	mdelay(10);
+	writew(0, regs + AZX_REG_CORBRP);
+	mdelay(10);
+	if (!bus->use_pio_for_commands)
+		writeb(AZX_CORBCTL_RUN, regs + AZX_REG_CORBCTL);
+	ae7_settle(bus);
+	writel((u32)(bus->rirb.addr + bus->addr_offset), regs + AZX_REG_RIRBLBASE);
+	writel(upper_32_bits(bus->rirb.addr + bus->addr_offset), regs + AZX_REG_RIRBUBASE);
+	ae7_settle(bus);
+	writeb(0x02, regs + AZX_REG_RIRBSIZE);
+	writew(AZX_RIRBWP_RST, regs + AZX_REG_RIRBWP);
+	ae7_settle(bus);
+	writew(1, regs + AZX_REG_RINTCNT);
+	ae7_settle(bus);
+	writeb(bus->not_use_interrupts ? AZX_RBCTL_DMA_EN : AZX_RBCTL_DMA_EN | AZX_RBCTL_IRQ_EN,
+	       regs + AZX_REG_RIRBCTL);
+	ae7_settle(bus);
+	writel(readl(regs + AZX_REG_GCTL) | AZX_GCTL_UNSOL, regs + AZX_REG_GCTL);
+	ae7_settle(bus);
+}
+
 void snd_hdac_bus_init_cmd_io(struct hdac_bus *bus)
 {
 	WARN_ON_ONCE(!bus->rb.area);
@@ -48,6 +94,14 @@
 	/* CORB set up */
 	bus->corb.addr = bus->rb.addr;
 	bus->corb.buf = (__le32 *)bus->rb.area;
+	if (snd_hdac_bus_flush_writes(bus)) {
+		bus->rirb.addr = bus->rb.addr + 2048;
+		bus->rirb.buf = (__le32 *)(bus->rb.area + 2048);
+		bus->rirb.wp = bus->rirb.rp = 0;
+		memset(bus->rirb.cmds, 0, sizeof(bus->rirb.cmds));
+		ae7_init_cmd_io(bus);
+		return;
+	}
 	snd_hdac_chip_writel(bus, CORBLBASE, (u32)(bus->corb.addr + bus->addr_offset));
 	snd_hdac_chip_writel(bus, CORBUBASE, upper_32_bits(bus->corb.addr + bus->addr_offset));
 
--- a/sound/hda/core/stream.c
+++ b/sound/hda/core/stream.c
@@ -230,6 +230,37 @@
 
 	dma_run_state = snd_hdac_stream_readb(azx_dev, SD_CTL) & SD_CTL_DMA_START;
 
+	if (snd_hdac_bus_flush_writes(azx_dev->bus)) {
+		/*
+		 * ae7 local patch (see include/sound/hdaudio.h): keep the spec handshake — SRST must be read back
+		 * as 1 and then as 0, otherwise stale DMA state survives the reset (a playback stream reused after
+		 * the DSP firmware upload kept reading the freed upload buffer). Give the hardware time before each
+		 * read-back instead of polling right after the write.
+		 */
+		static atomic_t ae7_logged = ATOMIC_INIT(0);
+		int in_polls = 0, out_polls = 0;
+
+		val = snd_hdac_stream_readb(azx_dev, SD_CTL);
+		writeb(val | SD_CTL_STREAM_RESET, azx_dev->sd_addr + AZX_REG_SD_CTL);
+		mdelay(5);
+		while (!(snd_hdac_stream_readb(azx_dev, SD_CTL) & SD_CTL_STREAM_RESET) && ++in_polls < 100)
+			udelay(10);
+		if (azx_dev->bus->dma_stop_delay && dma_run_state)
+			udelay(azx_dev->bus->dma_stop_delay);
+		writeb(val & ~SD_CTL_STREAM_RESET, azx_dev->sd_addr + AZX_REG_SD_CTL);
+		mdelay(5);
+		while ((snd_hdac_stream_readb(azx_dev, SD_CTL) & SD_CTL_STREAM_RESET) && ++out_polls < 100)
+			udelay(10);
+		if (atomic_inc_return(&ae7_logged) <= 30)
+			dev_info(azx_dev->bus->dev, "ae7: stream %d reset: SRST read 1 %s, read 0 %s\n",
+				 azx_dev->index,
+				 in_polls < 100 ? "ok" : "NEVER (timeout)",
+				 out_polls < 100 ? "ok" : "NEVER (timeout)");
+		if (azx_dev->posbuf)
+			*azx_dev->posbuf = 0;
+		return;
+	}
+
 	snd_hdac_stream_updateb(azx_dev, SD_CTL, 0, SD_CTL_STREAM_RESET);
 
 	/* wait for hardware to report that the stream entered reset */
--- a/sound/hda/codecs/ca0132.c
+++ b/sound/hda/codecs/ca0132.c
@@ -3627,6 +3627,29 @@
  * of the on-card LED. It seems to use pin 2 for data, then toggles 3 to on and
  * then off to send that bit.
  */
+/*
+ * ae7 local patch: flush every write to the card's second MMIO region (BAR 2) with a read, like the HDA register
+ * writes in include/sound/hdaudio.h — on a Sound Blaster AE-7 behind an ASMedia ASM1083/1085 rev 03 bridge a read
+ * that follows a run of posted writes never completes and takes the machine down.
+ */
+static void ca0132_mmio_writel(struct hda_codec *codec, u32 val, void __iomem *addr)
+{
+	writel(val, addr);
+	snd_hdac_reg_flush(&codec->bus->core);
+}
+
+static void ca0132_mmio_writew(struct hda_codec *codec, u16 val, void __iomem *addr)
+{
+	writew(val, addr);
+	snd_hdac_reg_flush(&codec->bus->core);
+}
+
+static void ca0132_mmio_writeb(struct hda_codec *codec, u8 val, void __iomem *addr)
+{
+	writeb(val, addr);
+	snd_hdac_reg_flush(&codec->bus->core);
+}
+
 static void ca0113_mmio_gpio_set(struct hda_codec *codec, unsigned int gpio_pin,
 		bool enable)
 {
@@ -3636,7 +3659,7 @@
 	gpio_data = gpio_pin & 0xF;
 	gpio_data |= ((enable << 8) & 0x100);
 
-	writew(gpio_data, spec->mem_base + 0x320);
+	ca0132_mmio_writew(codec, gpio_data, spec->mem_base + 0x320);
 }
 
 /*
@@ -3653,21 +3676,21 @@
 	struct ca0132_spec *spec = codec->spec;
 	unsigned int write_val;
 
-	writel(0x0000007e, spec->mem_base + 0x210);
+	ca0132_mmio_writel(codec, 0x0000007e, spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
-	writel(0x0000005a, spec->mem_base + 0x210);
+	ca0132_mmio_writel(codec, 0x0000005a, spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
 
-	writel(0x00800005, spec->mem_base + 0x20c);
-	writel(group, spec->mem_base + 0x804);
+	ca0132_mmio_writel(codec, 0x00800005, spec->mem_base + 0x20c);
+	ca0132_mmio_writel(codec, group, spec->mem_base + 0x804);
 
-	writel(0x00800005, spec->mem_base + 0x20c);
+	ca0132_mmio_writel(codec, 0x00800005, spec->mem_base + 0x20c);
 	write_val = (target & 0xff);
 	write_val |= (value << 8);
 
 
-	writel(write_val, spec->mem_base + 0x204);
+	ca0132_mmio_writel(codec, write_val, spec->mem_base + 0x204);
 	/*
 	 * Need delay here or else it goes too fast and works inconsistently.
 	 */
@@ -3677,8 +3700,8 @@
 	readl(spec->mem_base + 0x854);
 	readl(spec->mem_base + 0x840);
 
-	writel(0x00800004, spec->mem_base + 0x20c);
-	writel(0x00000000, spec->mem_base + 0x210);
+	ca0132_mmio_writel(codec, 0x00800004, spec->mem_base + 0x20c);
+	ca0132_mmio_writel(codec, 0x00000000, spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
 }
@@ -3692,28 +3715,28 @@
 	struct ca0132_spec *spec = codec->spec;
 	unsigned int write_val;
 
-	writel(0x0000007e, spec->mem_base + 0x210);
+	ca0132_mmio_writel(codec, 0x0000007e, spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
-	writel(0x0000005a, spec->mem_base + 0x210);
+	ca0132_mmio_writel(codec, 0x0000005a, spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
 
-	writel(0x00800003, spec->mem_base + 0x20c);
-	writel(group, spec->mem_base + 0x804);
+	ca0132_mmio_writel(codec, 0x00800003, spec->mem_base + 0x20c);
+	ca0132_mmio_writel(codec, group, spec->mem_base + 0x804);
 
-	writel(0x00800005, spec->mem_base + 0x20c);
+	ca0132_mmio_writel(codec, 0x00800005, spec->mem_base + 0x20c);
 	write_val = (target & 0xff);
 	write_val |= (value << 8);
 
 
-	writel(write_val, spec->mem_base + 0x204);
+	ca0132_mmio_writel(codec, write_val, spec->mem_base + 0x204);
 	msleep(20);
 	readl(spec->mem_base + 0x860);
 	readl(spec->mem_base + 0x854);
 	readl(spec->mem_base + 0x840);
 
-	writel(0x00800004, spec->mem_base + 0x20c);
-	writel(0x00000000, spec->mem_base + 0x210);
+	ca0132_mmio_writel(codec, 0x00800004, spec->mem_base + 0x20c);
+	ca0132_mmio_writel(codec, 0x00000000, spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
 	readl(spec->mem_base + 0x210);
 }
@@ -7894,18 +7917,18 @@
 	chipio_8051_write_direct(codec, 0x93, 0x10);
 	chipio_8051_write_pll_pmu(codec, 0x44, 0xc2);
 
-	writeb(0xff, spec->mem_base + 0x304);
-	writeb(0xff, spec->mem_base + 0x304);
-	writeb(0xff, spec->mem_base + 0x304);
-	writeb(0xff, spec->mem_base + 0x304);
-	writeb(0x00, spec->mem_base + 0x100);
-	writeb(0xff, spec->mem_base + 0x304);
-	writeb(0x00, spec->mem_base + 0x100);
-	writeb(0xff, spec->mem_base + 0x304);
-	writeb(0x00, spec->mem_base + 0x100);
-	writeb(0xff, spec->mem_base + 0x304);
-	writeb(0x00, spec->mem_base + 0x100);
-	writeb(0xff, spec->mem_base + 0x304);
+	ca0132_mmio_writeb(codec, 0xff, spec->mem_base + 0x304);
+	ca0132_mmio_writeb(codec, 0xff, spec->mem_base + 0x304);
+	ca0132_mmio_writeb(codec, 0xff, spec->mem_base + 0x304);
+	ca0132_mmio_writeb(codec, 0xff, spec->mem_base + 0x304);
+	ca0132_mmio_writeb(codec, 0x00, spec->mem_base + 0x100);
+	ca0132_mmio_writeb(codec, 0xff, spec->mem_base + 0x304);
+	ca0132_mmio_writeb(codec, 0x00, spec->mem_base + 0x100);
+	ca0132_mmio_writeb(codec, 0xff, spec->mem_base + 0x304);
+	ca0132_mmio_writeb(codec, 0x00, spec->mem_base + 0x100);
+	ca0132_mmio_writeb(codec, 0xff, spec->mem_base + 0x304);
+	ca0132_mmio_writeb(codec, 0x00, spec->mem_base + 0x100);
+	ca0132_mmio_writeb(codec, 0xff, spec->mem_base + 0x304);
 
 	ca0113_mmio_command_set(codec, 0x30, 0x2b, 0x3f);
 	ca0113_mmio_command_set(codec, 0x30, 0x2d, 0x3f);
@@ -8827,9 +8850,9 @@
 	unsigned int i;
 
 	for (i = 0; i < 4; i++)
-		writeb(0x0, spec->mem_base + 0x100);
+		ca0132_mmio_writeb(codec, 0x0, spec->mem_base + 0x100);
 	for (i = 0; i < 8; i++)
-		writeb(0xb3, spec->mem_base + 0x304);
+		ca0132_mmio_writeb(codec, 0xb3, spec->mem_base + 0x304);
 
 	ca0113_mmio_gpio_set(codec, 0, false);
 	ca0113_mmio_gpio_set(codec, 1, false);
@@ -9111,8 +9134,8 @@
 {
 	struct ca0132_spec *spec = codec->spec;
 
-	writel(0x00820680, spec->mem_base + 0x01C);
-	writel(0x00820680, spec->mem_base + 0x01C);
+	ca0132_mmio_writel(codec, 0x00820680, spec->mem_base + 0x01C);
+	ca0132_mmio_writel(codec, 0x00820680, spec->mem_base + 0x01C);
 
 	chipio_write(codec, 0x18b0a4, 0x000000c2);
 
@@ -9228,7 +9251,7 @@
 
 	addr = ca0113_mmio_init_address_sbz;
 	for (i = 0; i < 3; i++)
-		writel(0x00000000, spec->mem_base + addr[i]);
+		ca0132_mmio_writel(codec, 0x00000000, spec->mem_base + addr[i]);
 
 	cur_addr = i;
 	switch (ca0132_quirk(spec)) {
@@ -9251,7 +9274,7 @@
 	}
 
 	for (i = 0; i < 2; i++)
-		writel(tmp[i], spec->mem_base + addr[cur_addr + i]);
+		ca0132_mmio_writel(codec, tmp[i], spec->mem_base + addr[cur_addr + i]);
 
 	cur_addr += i;
 
@@ -9267,7 +9290,7 @@
 	}
 
 	for (i = 0; i < count; i++)
-		writel(data[i], spec->mem_base + addr[cur_addr + i]);
+		ca0132_mmio_writel(codec, data[i], spec->mem_base + addr[cur_addr + i]);
 }
 
 static void ca0132_mmio_init_ae5(struct hda_codec *codec)
@@ -9281,8 +9304,8 @@
 	count = ARRAY_SIZE(ca0113_mmio_init_data_ae5);
 
 	if (ca0132_quirk(spec) == QUIRK_AE7) {
-		writel(0x00000680, spec->mem_base + 0x1c);
-		writel(0x00880680, spec->mem_base + 0x1c);
+		ca0132_mmio_writel(codec, 0x00000680, spec->mem_base + 0x1c);
+		ca0132_mmio_writel(codec, 0x00880680, spec->mem_base + 0x1c);
 	}
 
 	for (i = 0; i < count; i++) {
@@ -9291,15 +9314,15 @@
 		 * a different value to 0x20c.
 		 */
 		if (i == 21 && ca0132_quirk(spec) == QUIRK_AE7) {
-			writel(0x00800001, spec->mem_base + addr[i]);
+			ca0132_mmio_writel(codec, 0x00800001, spec->mem_base + addr[i]);
 			continue;
 		}
 
-		writel(data[i], spec->mem_base + addr[i]);
+		ca0132_mmio_writel(codec, data[i], spec->mem_base + addr[i]);
 	}
 
 	if (ca0132_quirk(spec) == QUIRK_AE5)
-		writel(0x00880680, spec->mem_base + 0x1c);
+		ca0132_mmio_writel(codec, 0x00880680, spec->mem_base + 0x1c);
 }
 
 static void ca0132_mmio_init(struct hda_codec *codec)
@@ -9361,19 +9384,19 @@
 	}
 
 	for (i = cur_addr = 0; i < 3; i++, cur_addr++)
-		writeb(tmp[i], spec->mem_base + addr[cur_addr]);
+		ca0132_mmio_writeb(codec, tmp[i], spec->mem_base + addr[cur_addr]);
 
 	/*
 	 * First writes are in single bytes, final are in 4 bytes. So, we use
 	 * writeb, then writel.
 	 */
 	for (i = 0; cur_addr < 12; i++, cur_addr++)
-		writeb(data[i], spec->mem_base + addr[cur_addr]);
+		ca0132_mmio_writeb(codec, data[i], spec->mem_base + addr[cur_addr]);
 
 	for (; cur_addr < count; i++, cur_addr++)
-		writel(data[i], spec->mem_base + addr[cur_addr]);
+		ca0132_mmio_writel(codec, data[i], spec->mem_base + addr[cur_addr]);
 
-	writel(0x00800001, spec->mem_base + 0x20c);
+	ca0132_mmio_writel(codec, 0x00800001, spec->mem_base + 0x20c);
 
 	if (ca0132_quirk(spec) == QUIRK_AE7) {
 		ca0113_mmio_command_set_type2(codec, 0x48, 0x07, 0x83);

^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [BUG] ALSA: hda: AE-7 behind ASM1083 rev 03 bridge: fatal MCE at probe
  2026-10-05 12:47 [BUG] ALSA: hda: AE-7 behind ASM1083 rev 03 bridge: fatal MCE at probe Munteanu Vlad
@ 2026-10-05 13:30 ` Takashi Iwai
  0 siblings, 0 replies; 2+ messages in thread
From: Takashi Iwai @ 2026-10-05 13:30 UTC (permalink / raw)
  To: Munteanu Vlad; +Cc: tiwai, linux-sound, perex, conmanx360, bhelgaas, linux-pci

On Mon, 05 Oct 2026 14:47:15 +0200,
Munteanu Vlad wrote:
> 
> Hi,
> 
> Newer Sound Blaster AE-7 cards (CA0132, PCI 1102:0010, SSID 1102:0081)
> put the CA0132 behind an ASMedia ASM1083/1085 rev 03 PCIe-to-PCI bridge.
> On my machine, binding snd_hda_intel to the card takes the whole system
> down every time: a fatal machine check (or a silent hard hang) within
> about a second of probe. This looks like the problem other AE-7 owners
> report, where the card "never worked" on Linux [1].
> 
> I tracked it down to MMIO reads to the CA0132 that never complete. I
> have a workaround (attached, not meant for merging as-is) with which the
> card now works fully: probe, DSP firmware download, playback. I'd like
> your advice on the proper fix, and I'm happy to test patches.
> 
> 
> Hardware / software
> -------------------
> 
> - Dell Precision 7920 Tower, 2x Xeon Gold 6152 (Skylake-SP),
> BIOS 2.53.1, APEI firmware-first error handling.
> - 44:00.0 Intel Sky Lake-E PCIe Root Port A [8086:2030]
> -> 45:00.0 ASMedia ASM1083/1085 PCIe-to-PCI bridge [1b21:1080] rev 03
> -> 46:00.0 Creative CA0132 [1102:0010] rev 01, subsystem [1102:0081]
> - Ubuntu kernel 7.0.0-38-generic (based on 7.0.14). The code paths
> involved (sound/hda/core/controller.c, core/stream.c,
> codecs/ca0132.c) are the same in current mainline as far as I can
> see. I haven't built mainline yet, but can if that helps. The same
> hang happened with Ubuntu's 6.17 kernel at install time.
> 
> 
> 1. The failure
> --------------
> 
> Identical in every run (CPU PPIN removed):
> 
> mce: [Hardware Error]: CPU 0: Machine Check Exception: 5 Bank 6:
> bb80000000000e0b
> mce: [Hardware Error]: RIP !INEXACT! 10:<ffffffffc1fdbf1b>
> {snd_hdac_bus_init_cmd_io+0x1db/0x260 [snd_hda_core]}
> mce: [Hardware Error]: TSC 2e59a247c66 MISC 44000000
> mce: [Hardware Error]: PROCESSOR 0:50654 TIME 1791056639 SOCKET 0 APIC
> 0 microcode 2007006
> mce: [Hardware Error]: Machine check: Processor context corrupt
> Kernel panic - not syncing: Fatal machine check
> 
> Bank 6 is the IIO, and MCACOD 0x0e0b is a generic I/O bus error. The MCE
> comes about 1.18 s after the last driver message ("codec_mask = 0x2"),
> which matches the root port's completion timeout (260-900 ms). So a CPU
> read to the CA0132 never got a completion.
> 
> In this build, +0x1db is the instruction right after
> "mov 0x4a(%rax),%ax", i.e. the readw(CORBRP) in the first poll loop of
> azx_clear_corbrp(), immediately after writew(CORBRP, AZX_CORBRP_RST).
> CORB DMA isn't running yet at that point.
> 
> 
> 2. What the experiments showed
> ------------------------------
> 
> - Done from userspace with the card unbound (setpci / devmem), each
> register access is fine on its own. That includes the controller reset
> at driver timing, and the whole CORB/RIRB setup sequence done slowly
> with a read after each step. Note that the CORB/RIRB base addresses
> were 0 in that test, so no real DMA happened, and the IOMMU logged no
> faults.
> 
> - In the driver, with the CORBRP readback avoided (write RST, wait
> 10 ms, write 0, wait 10 ms, no reads in between), the MCE moved to
> snd_hdac_bus_init_cmd_io+0x186. That's readl(GCTL) in the final
> updatel(GCTL, UNSOL): the first read after CORBCTL=RUN, the RIRB
> base/size writes, RIRBWP=RST, RINTCNT and RIRBCTL. Adding 10 ms delays
> between those writes did NOT help. Adding a GCTL read after each
> group of writes did.
> 
> - With that in place, the probe got through codec enumeration and then
> hung (no MCE record that time). The last trace point was
> ca0132_mmio_init() returning. The next code is ae5_register_set() (the
> AE-7 path): 19 BAR2 writes, then the first BAR2 read, in
> ca0113_mmio_command_set_type2(). So presumably it was that read.
> 
> - Turning off AER/SERR/parity reporting on the card, the bridge and the
> root port did not help either. The machine hung at the same point, but
> left no MCE record and had to be power-cycled.
> 
> So the pattern seems to be: on this bridge, a read that follows a run
> of posted writes to the CA0132 never completes. One data point doesn't
> fit a simple count, though. azx_int_clear() does 13 writes to
> registers 32 bytes apart, and the read after it was fine. So it may be
> about writes to neighbouring registers being merged by the bridge. I
> couldn't pin that down. ASM1083 rev 03 is reported as broken with other
> PCI cards as well [2].
> 
> 
> 3. Workaround that works (attached, against the Ubuntu 7.0.0-38 tree)
> --------------------------------------------------------------------
> 
> All of it applies only to Creative HDA controllers (PCI vendor check):
> 
> a) snd_hdac_reg_write{b,w,l}(): read GCTL after every register write.
> b) snd_hdac_bus_init_cmd_io(): do the CORB read pointer reset without
> polling CORBRP while RST is set (set, 10 ms, clear, 10 ms), and
> settle + read GCTL after each group of CORB/RIRB writes.
> c) snd_hdac_stream_reset(): keep the SRST handshake, but wait 5 ms
> before each read-back.
> d) ca0132.c: follow every write to spec->mem_base (45 sites) with the
> same flush read.
> 
> With this everything works over several boots and hours of playback:
> controller and codec probe, "ca0132 DSP downloaded and running", the
> AE-7 post-DSP setup, and playback in 2.0 and 2.1 (6 channels). There
> have been no MCEs since. NVIDIA HDMI audio, driven by the same
> snd_hda_intel, is unaffected (the flushes are gated on the vendor).
> 
> (c) may not be needed. My first version did the stream reset blind (no
> SRST read-back) because I suspected reads during reset. The DMAR faults
> I then saw turned out to be the separate issue in section 4.
> 
> 
> 4. Second issue: the CA0132 reads past the end of the cyclic buffer
> -------------------------------------------------------------------
> 
> With IOMMU translation (default DMA-FQ domain), every buffer wrap during
> playback produces:
> 
> DMAR: [DMA Read NO_PASID] Request device [46:00.0] fault addr
> 0xffec0000 [fault reason 0x06] PTE Read access is not set
> 
> The fault repeats at the buffer-wrap period (0.683 s for 32768 frames
> at 48 kHz). The address changes with the buffer, and it's consistent
> with the address just past the end of a size-aligned IOVA allocation of
> the current PCM buffer: 0xffec0000 for a 768 KiB buffer (6 ch, which
> would sit at 0xffe00000), and 0xfff40000 / 0xfff80000 for 256 KiB
> buffers. I haven't read the buffer's IOVA directly. If that's right,
> the controller prefetches past the last BDL entry before wrapping. The
> audio itself is fine. As a workaround I switch the card's IOMMU group
> to identity before binding it.
> 
> 
> 5. Minor
> --------
> 
> PipeWire first asks for buffer=1572864, period=49152, which gives
> "Too many BDL entries" (because of AZX_DCAPS_4K_BDLE_BOUNDARY). It then
> falls back to a smaller buffer and works.
> 
> 
> Questions
> ---------
> 
> - Would a quirk keyed on an ASMedia 1b21:1080 (rev 03) bridge directly
> upstream of the controller be acceptable? It would set something like
> bus->flush_writes (read back after every write), plus a CORB reset
> that doesn't poll while RST is set. Or would you rather this live in
> the PCI layer?
> - Are there any known ASM1083/1085 errata about posted writes or write
> merging?
> - For the buffer overread, what's the preferred fix: padding the
> allocation, an extra BDL entry, or something else?
> 
> I can test patches on this machine. Failures leave crash dumps in
> pstore/ERST, so each failed test is quick to diagnose.
> 
> [1] https://forum.endeavouros.com/t/unable-to-boot-after-installing-new-sound-card-ae-7/40457
> [2] https://projects.osmocom.org/projects/retronetworking/wiki/PCIe-%3EPCI_bridges
> [3] https://bugzilla.kernel.org/show_bug.cgi?id=208667 (ASM1083/1085 ASPM quirk)
> [4] https://bugzilla.kernel.org/show_bug.cgi?id=217510 (AE-7, possibly related)
> 
> Regards,
> gigiou_88

Thanks for the report and the fix attempt.

My gut feeling is, though, it should be rather addressed in the PCI
core side or similar lower layer than tweaking too much on the
HD-audio driver side.


Takashi

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2026-10-05 13:30 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-05 12:47 [BUG] ALSA: hda: AE-7 behind ASM1083 rev 03 bridge: fatal MCE at probe Munteanu Vlad
2026-10-05 13:30 ` Takashi Iwai

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox