All of lore.kernel.org
 help / color / mirror / Atom feed
From: Alvin Lim <alvinwylim@gmail.com>
To: mikael1022bzh@gmail.com, cassel@kernel.org, mario.limonciello@amd.com
Cc: roland.waltersson@netinsight.net, artmoty@gmail.com,
	linux-ide@vger.kernel.org, david.laight.linux@gmail.com,
	kernel@wantstofly.org
Subject: Re: [PATCH] ahci: force 32-bit DMA for JMicron JMB582/JMB585
Date: Sat,  5 Sep 2026 19:38:48 +0800	[thread overview]
Message-ID: <20260905113848.1397882-1-alvinwylim@gmail.com> (raw)
In-Reply-To: <178854093896.541085.18395405788748617410@gmail.com>

On Fri, Sep 04, 2026 at 11:55:38PM +0700, Mikael Etienne wrote:
> That would be particularly worth it for Alvin, since iommu=pt is
> insufficient for him while it is completely clean for me -- that is the
> only divergence between our reports.

There is no divergence, and the fault is mine. I owe the thread an
apology.

My writeup stated that iommu=pt is insufficient. I never tested it. I have
no data on iommu=pt on this machine in either direction, and I am not now
claiming the opposite either -- only that the statement was not mine to
make and should be disregarded wherever it has been read.

The cause was poor discipline in my own record keeping. My incident notes
recorded some things I had measured and some things I had picked up from
other people's reports about other hardware, and by the time I wrote them
up for publication the distinction had been lost. The published page then
read as though all of it were my own result. That is how an untested
claim ended up being cited on this list, and it is a fair description of
what went wrong.

I am cleaning that up now. The rewritten page contains only what I
personally tested and observed on this machine, plus an explicit section
listing what I did not test, so nobody has to guess which is which. The
previous version stays in git history rather than being quietly replaced.

What I actually observed
------------------------

  Controller : ASMedia ASM1166 [1b21:1166] rev 02, all 6 SATA bays
  Kernel     : 7.0.6-2-pve (Proxmox) at the time of the corruption
  IVHD       : AMD-Vi: Using global IVHD EFR:0x246577efa2054ada, EFR2:0x0

(Full platform details are at the end of this mail, in case they help
with deciding who tests what.)

Symptom, with the IOMMU enabled: silent non-deterministic read corruption
on every SATA disk. A single isolated read was often clean; six concurrent
reads of one file returned six different md5s. SMART clean, no link
resets, no MCE. NVMe on the same host was unaffected. On-disk data turned
out to be intact -- it was entirely a read-path problem.

The fix I applied was amd_iommu=off, and nothing else. The machine's RAM
was later changed, so it has now run under that single parameter at two
different physical address ranges:

  2026-06-17 -> 2026-08-01   45 days   32 GB   (30 GiB visible)
  2026-08-01 -> 2026-09-05   35 days   96 GB   (92 GiB visible)

Kernel error lines on this host, same grep either side of the fix:

  2026-06-15 -> 06-17 16:00   (~2.5 days)    4240
  2026-06-17 16:00 -> now     (80 days)         0

Zero corruption to date in both memory configurations.

One further correction I have to make, because it wrongly closed off a
test. My writeup said this part lacks GIOSup, so amd_iommu=pgtbl_v2 could
not work. I read that off the "Extended features" boot line and treated
the absence of the name there as absence of the feature. Decoding the
register instead, bit 48 (GIOSUP) and bit 4 (GT) are both set here, so on
my reading of amd_iommu_v2_pgtbl_supported() it should be available on
this box. I have not tested pgtbl_v2 either -- I have only established
that my stated reason for ruling it out was invalid.

What I did NOT test on this machine
-----------------------------------

  - iommu=pt
  - amd_iommu=pgtbl_v2
  - iommu.forcedac=1
  - any other kernel version
  - the 32-bit DMA quirk itself; no kernel was ever built with it
  - which DMA addresses were actually in play; never instrumented, so I
    cannot say whether the failing addresses were above or below 4 GB
  - any controller other than this ASM1166, and any non-AMD platform

I know the symptom, I know amd_iommu=off removes it, and I know this
controller handles physical addresses far above 4 GB without error. I do
not know the mechanism, and I should not have implied otherwise.

Accordingly, please treat my ASM1166 patch as withdrawn:

  https://lore.kernel.org/linux-ide/20260621100844.1224301-1-alvinwylim@gmail.com/

I never built or ran a kernel with that quirk, and the submission did not
disclose that. Its premise -- that the controller cannot address above
4 GB -- is contradicted by the 80 days above.

The submission also carried a Fixes: tag naming commit 3bf614106094
("ata: ahci: add identifiers for ASM2116 series adapters"), together with
Cc: stable@vger.kernel.org. That Fixes tag was wrong, and I want to be
clear why rather than leave it as an assertion.

ahci_pci_tbl ends with a generic catch-all,
PCI_DEVICE_CLASS(PCI_CLASS_STORAGE_SATA_AHCI, 0xffffff) -> board_ahci,
which is present in v5.10 and v6.1, both well before that commit. So the
ASM1166 already bound to board_ahci through the catch-all, and adding its
explicit ID did not change how it was handled. Nothing about DMA differed
before and after, so there was no regression for that commit to have
introduced.

That matters because of what the two tags do together: Fixes: is what the
stable tooling reads to decide which series a backport applies to, so with
Cc: stable@ beside it the pair would have aimed an untested patch at every
stable series containing that commit. Nothing came of it, because the
patch was never applied. It should not have been there all the same.

What I can offer as a test platform
-----------------------------------

Setting this out in full, so nobody has to reconstruct it from earlier
messages when deciding who should test what.

  Board      : AOOSTAR WTR MAX (6-bay mini NAS)
  CPU        : AMD Ryzen 7 PRO 8845HS (Zen 4)
  IOMMU      : AMD. EFR 0x246577efa2054ada, EFR2 0x0
               GIOSup (bit 48) and GT (bit 4) both set
  SATA       : one ASMedia ASM1166 [1b21:1166] rev 02, 6 ports,
               6 HDDs, 98 TB total, 2-3 TB free per filesystem
  NVMe       : present, on separate controllers, never affected
  OS         : Proxmox VE (Debian). Kernel 7.0.14-14-pve now;
               7.0.6-2-pve when the corruption happened
  Current    : running amd_iommu=off since 2026-06-17

What I do NOT have, so please do not wait on me for it: no JMicron
JMB582/585, no Intel or ARM IOMMU machine, and no SATA controller other
than this ASM1166. The JMB585-plus-Intel/ARM test asked for earlier in
this thread is not one I can run.

What I can do:

  - any boot parameter combination: amd_iommu=pgtbl_v2, iommu.forcedac=1,
    iommu=pt, or back to amd_iommu=off
  - older kernels from the Proxmox archive
  - build and run a custom kernel, if a patch needs testing
  - run the fio crc32c canary; there is room for a 256 GB canary and I can
    leave verification looping overnight

One constraint, stated up front rather than discovered later: this is a
live NAS and the affected controller carries real data, including a Ceph
OSD. Any test with the IOMMU re-enabled needs a scheduled window with the
filesystems unmounted and that OSD stopped. I would write the canary first
under amd_iommu=off, through the path I know is good, and then run
verify-only passes in the test configuration, so nothing is ever written
through a path under suspicion.

So: tell me which test would actually be useful and I will run it and
report the result either way, including a negative one. I would rather
spend the window on the question you want answered than guess at it.

Alvin Lim

  reply	other threads:[~2026-09-05 11:38 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-03 13:35 [PATCH] ahci: force 32-bit DMA for JMicron JMB582/JMB585 Roland Waltersson
2026-09-03 21:40 ` Niklas Cassel
2026-09-03 22:36   ` Mario Limonciello
2026-09-04  5:20     ` Roland Waltersson
2026-09-04 11:14       ` Niklas Cassel
2026-09-04 12:38         ` Mario Limonciello
2026-09-04 12:50           ` Niklas Cassel
2026-09-04 16:55             ` Mikael Etienne
2026-09-05 11:38               ` Alvin Lim [this message]
2026-09-05 12:26                 ` Mario Limonciello
2026-09-06  5:05                   ` Alvin Lim
2026-09-06 19:27                     ` John Smith
2026-09-05 14:13               ` Lennert Buytenhek
     [not found] <AM8P193MB2659454F98AD53DEB002CF8AA9B42@AM8P193MB2659.EURP193.PROD.OUTLOOK.COM>
2026-09-07 10:28 ` Niklas Cassel
  -- strict thread matches above, loose matches on Subject: below --
2026-09-06 11:48 snoep
2026-04-03  5:04 Arthur Husband
2026-04-03  7:02 ` Damien Le Moal
2026-04-03  8:12 ` Niklas Cassel
2026-04-03  8:19 ` Niklas Cassel
2026-04-03  5:02 Arthur Husband

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260905113848.1397882-1-alvinwylim@gmail.com \
    --to=alvinwylim@gmail.com \
    --cc=artmoty@gmail.com \
    --cc=cassel@kernel.org \
    --cc=david.laight.linux@gmail.com \
    --cc=kernel@wantstofly.org \
    --cc=linux-ide@vger.kernel.org \
    --cc=mario.limonciello@amd.com \
    --cc=mikael1022bzh@gmail.com \
    --cc=roland.waltersson@netinsight.net \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.