Linux PCI subsystem development
 help / color / mirror / Atom feed
From: Takashi Iwai <tiwai@suse.de>
To: Munteanu Vlad <m.vlad.stefan@gmail.com>
Cc: tiwai@suse.com, linux-sound@vger.kernel.org, perex@perex.cz,
	conmanx360@gmail.com, bhelgaas@google.com,
	linux-pci@vger.kernel.org
Subject: Re: [BUG] ALSA: hda: AE-7 behind ASM1083 rev 03 bridge: fatal MCE at probe
Date: Mon, 05 Oct 2026 15:30:35 +0200	[thread overview]
Message-ID: <878q4cs29w.wl-tiwai@suse.de> (raw)
In-Reply-To: <CAMx+sF_kr_WW4LENM2K-7XxDHN4vER+5qz07NAOhKC3A07Gvaw@mail.gmail.com>

On Mon, 05 Oct 2026 14:47:15 +0200,
Munteanu Vlad wrote:
> 
> Hi,
> 
> Newer Sound Blaster AE-7 cards (CA0132, PCI 1102:0010, SSID 1102:0081)
> put the CA0132 behind an ASMedia ASM1083/1085 rev 03 PCIe-to-PCI bridge.
> On my machine, binding snd_hda_intel to the card takes the whole system
> down every time: a fatal machine check (or a silent hard hang) within
> about a second of probe. This looks like the problem other AE-7 owners
> report, where the card "never worked" on Linux [1].
> 
> I tracked it down to MMIO reads to the CA0132 that never complete. I
> have a workaround (attached, not meant for merging as-is) with which the
> card now works fully: probe, DSP firmware download, playback. I'd like
> your advice on the proper fix, and I'm happy to test patches.
> 
> 
> Hardware / software
> -------------------
> 
> - Dell Precision 7920 Tower, 2x Xeon Gold 6152 (Skylake-SP),
> BIOS 2.53.1, APEI firmware-first error handling.
> - 44:00.0 Intel Sky Lake-E PCIe Root Port A [8086:2030]
> -> 45:00.0 ASMedia ASM1083/1085 PCIe-to-PCI bridge [1b21:1080] rev 03
> -> 46:00.0 Creative CA0132 [1102:0010] rev 01, subsystem [1102:0081]
> - Ubuntu kernel 7.0.0-38-generic (based on 7.0.14). The code paths
> involved (sound/hda/core/controller.c, core/stream.c,
> codecs/ca0132.c) are the same in current mainline as far as I can
> see. I haven't built mainline yet, but can if that helps. The same
> hang happened with Ubuntu's 6.17 kernel at install time.
> 
> 
> 1. The failure
> --------------
> 
> Identical in every run (CPU PPIN removed):
> 
> mce: [Hardware Error]: CPU 0: Machine Check Exception: 5 Bank 6:
> bb80000000000e0b
> mce: [Hardware Error]: RIP !INEXACT! 10:<ffffffffc1fdbf1b>
> {snd_hdac_bus_init_cmd_io+0x1db/0x260 [snd_hda_core]}
> mce: [Hardware Error]: TSC 2e59a247c66 MISC 44000000
> mce: [Hardware Error]: PROCESSOR 0:50654 TIME 1791056639 SOCKET 0 APIC
> 0 microcode 2007006
> mce: [Hardware Error]: Machine check: Processor context corrupt
> Kernel panic - not syncing: Fatal machine check
> 
> Bank 6 is the IIO, and MCACOD 0x0e0b is a generic I/O bus error. The MCE
> comes about 1.18 s after the last driver message ("codec_mask = 0x2"),
> which matches the root port's completion timeout (260-900 ms). So a CPU
> read to the CA0132 never got a completion.
> 
> In this build, +0x1db is the instruction right after
> "mov 0x4a(%rax),%ax", i.e. the readw(CORBRP) in the first poll loop of
> azx_clear_corbrp(), immediately after writew(CORBRP, AZX_CORBRP_RST).
> CORB DMA isn't running yet at that point.
> 
> 
> 2. What the experiments showed
> ------------------------------
> 
> - Done from userspace with the card unbound (setpci / devmem), each
> register access is fine on its own. That includes the controller reset
> at driver timing, and the whole CORB/RIRB setup sequence done slowly
> with a read after each step. Note that the CORB/RIRB base addresses
> were 0 in that test, so no real DMA happened, and the IOMMU logged no
> faults.
> 
> - In the driver, with the CORBRP readback avoided (write RST, wait
> 10 ms, write 0, wait 10 ms, no reads in between), the MCE moved to
> snd_hdac_bus_init_cmd_io+0x186. That's readl(GCTL) in the final
> updatel(GCTL, UNSOL): the first read after CORBCTL=RUN, the RIRB
> base/size writes, RIRBWP=RST, RINTCNT and RIRBCTL. Adding 10 ms delays
> between those writes did NOT help. Adding a GCTL read after each
> group of writes did.
> 
> - With that in place, the probe got through codec enumeration and then
> hung (no MCE record that time). The last trace point was
> ca0132_mmio_init() returning. The next code is ae5_register_set() (the
> AE-7 path): 19 BAR2 writes, then the first BAR2 read, in
> ca0113_mmio_command_set_type2(). So presumably it was that read.
> 
> - Turning off AER/SERR/parity reporting on the card, the bridge and the
> root port did not help either. The machine hung at the same point, but
> left no MCE record and had to be power-cycled.
> 
> So the pattern seems to be: on this bridge, a read that follows a run
> of posted writes to the CA0132 never completes. One data point doesn't
> fit a simple count, though. azx_int_clear() does 13 writes to
> registers 32 bytes apart, and the read after it was fine. So it may be
> about writes to neighbouring registers being merged by the bridge. I
> couldn't pin that down. ASM1083 rev 03 is reported as broken with other
> PCI cards as well [2].
> 
> 
> 3. Workaround that works (attached, against the Ubuntu 7.0.0-38 tree)
> --------------------------------------------------------------------
> 
> All of it applies only to Creative HDA controllers (PCI vendor check):
> 
> a) snd_hdac_reg_write{b,w,l}(): read GCTL after every register write.
> b) snd_hdac_bus_init_cmd_io(): do the CORB read pointer reset without
> polling CORBRP while RST is set (set, 10 ms, clear, 10 ms), and
> settle + read GCTL after each group of CORB/RIRB writes.
> c) snd_hdac_stream_reset(): keep the SRST handshake, but wait 5 ms
> before each read-back.
> d) ca0132.c: follow every write to spec->mem_base (45 sites) with the
> same flush read.
> 
> With this everything works over several boots and hours of playback:
> controller and codec probe, "ca0132 DSP downloaded and running", the
> AE-7 post-DSP setup, and playback in 2.0 and 2.1 (6 channels). There
> have been no MCEs since. NVIDIA HDMI audio, driven by the same
> snd_hda_intel, is unaffected (the flushes are gated on the vendor).
> 
> (c) may not be needed. My first version did the stream reset blind (no
> SRST read-back) because I suspected reads during reset. The DMAR faults
> I then saw turned out to be the separate issue in section 4.
> 
> 
> 4. Second issue: the CA0132 reads past the end of the cyclic buffer
> -------------------------------------------------------------------
> 
> With IOMMU translation (default DMA-FQ domain), every buffer wrap during
> playback produces:
> 
> DMAR: [DMA Read NO_PASID] Request device [46:00.0] fault addr
> 0xffec0000 [fault reason 0x06] PTE Read access is not set
> 
> The fault repeats at the buffer-wrap period (0.683 s for 32768 frames
> at 48 kHz). The address changes with the buffer, and it's consistent
> with the address just past the end of a size-aligned IOVA allocation of
> the current PCM buffer: 0xffec0000 for a 768 KiB buffer (6 ch, which
> would sit at 0xffe00000), and 0xfff40000 / 0xfff80000 for 256 KiB
> buffers. I haven't read the buffer's IOVA directly. If that's right,
> the controller prefetches past the last BDL entry before wrapping. The
> audio itself is fine. As a workaround I switch the card's IOMMU group
> to identity before binding it.
> 
> 
> 5. Minor
> --------
> 
> PipeWire first asks for buffer=1572864, period=49152, which gives
> "Too many BDL entries" (because of AZX_DCAPS_4K_BDLE_BOUNDARY). It then
> falls back to a smaller buffer and works.
> 
> 
> Questions
> ---------
> 
> - Would a quirk keyed on an ASMedia 1b21:1080 (rev 03) bridge directly
> upstream of the controller be acceptable? It would set something like
> bus->flush_writes (read back after every write), plus a CORB reset
> that doesn't poll while RST is set. Or would you rather this live in
> the PCI layer?
> - Are there any known ASM1083/1085 errata about posted writes or write
> merging?
> - For the buffer overread, what's the preferred fix: padding the
> allocation, an extra BDL entry, or something else?
> 
> I can test patches on this machine. Failures leave crash dumps in
> pstore/ERST, so each failed test is quick to diagnose.
> 
> [1] https://forum.endeavouros.com/t/unable-to-boot-after-installing-new-sound-card-ae-7/40457
> [2] https://projects.osmocom.org/projects/retronetworking/wiki/PCIe-%3EPCI_bridges
> [3] https://bugzilla.kernel.org/show_bug.cgi?id=208667 (ASM1083/1085 ASPM quirk)
> [4] https://bugzilla.kernel.org/show_bug.cgi?id=217510 (AE-7, possibly related)
> 
> Regards,
> gigiou_88

Thanks for the report and the fix attempt.

My gut feeling is, though, it should be rather addressed in the PCI
core side or similar lower layer than tweaking too much on the
HD-audio driver side.


Takashi

      reply	other threads:[~2026-10-05 13:30 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-05 12:47 [BUG] ALSA: hda: AE-7 behind ASM1083 rev 03 bridge: fatal MCE at probe Munteanu Vlad
2026-10-05 13:30 ` Takashi Iwai [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=878q4cs29w.wl-tiwai@suse.de \
    --to=tiwai@suse.de \
    --cc=bhelgaas@google.com \
    --cc=conmanx360@gmail.com \
    --cc=linux-pci@vger.kernel.org \
    --cc=linux-sound@vger.kernel.org \
    --cc=m.vlad.stefan@gmail.com \
    --cc=perex@perex.cz \
    --cc=tiwai@suse.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox