All of lore.kernel.org
 help / color / mirror / Atom feed
From: Mario Limonciello <mario.limonciello@amd.com>
To: Mikael Etienne <mikael1022bzh@gmail.com>,
	vasant.hegde@amd.com, iommu@lists.linux.dev,
	linux-ide@vger.kernel.org, linux-block@vger.kernel.org
Cc: regressions@lists.linux.dev, joro@8bytes.org,
	suravee.suthikulpanit@amd.com
Subject: Re: [REGRESSION] Silent SATA read corruption with dma-iommu on AMD 600-series AHCI (6.19 good, 7.0+ bad)
Date: Fri, 28 Aug 2026 10:37:51 -0500	[thread overview]
Message-ID: <d49add14-c8de-4b6b-9049-4a07e5692ee9@amd.com> (raw)
In-Reply-To: <178793069172.1699234.13635367891277204598@gmail.com>



On 8/28/26 10:24, Mikael Etienne wrote:
> On 8/28/2026 6:22 PM, Limonciello, Mario wrote:
>> BUT the issue internally does show the issue is specificially once the
>> 32-bit IOVA space is exhausted.
>>
>> Unfortunately; the solution is currently a BIOS change in how the type
>> bytes of the IOVA is handled.
> 
> Hi Mario,
> 
> Thanks -- the 32-bit IOVA exhaustion detail is the first thing that explains
> the timing I see. My failure takes 3h25m of sustained reads to appear, and a
> reboot clears it completely with the same kernel and the same on-disk data.
> "Low IOVA space works, high IOVA space does not, and a reboot starts from an
> empty low space again" fits that exactly.
> 
> The ASM1166 write-up you linked is also a very close match: a controller that
> advertises CAP.S64A but cannot actually reach above 4 GB. My SATA controller is
> the AMD 600-series chipset one (1022:43f6), which is Promontory/ASMedia silicon,
> and it likewise advertises 64bit in its AHCI flags.
> 
> Firmware, since you mention a BIOS fix: I am already on the latest available for
> this board.
> 
>    Gigabyte X870I AORUS PRO ICE
>    BIOS FB1c, 2026-07-21, AMI, platform firmware revision 5.41
>    AGESA!V9 ComboAm5PI 1.3.0.1c
>    CPU microcode 0x0a70520a
> 
> So whatever BIOS-side change you have internally is either not in AGESA
> 1.3.0.1c, or not sufficient on this board.

It's not in any AGESA release yet.  This is very fresh information I am 
sharing that we have root caused the issue and have a proposed 
modification.  It will take a while to make it through the process 
machinery.
> 
> On the bisection: before committing to two weeks, I would like to try making the
> reproducer fast. If the trigger really is 32-bit IOVA exhaustion, then booting
> with iommu.forcedac=1 alone should hand out high IOVAs immediately and fail in
> minutes rather than after 3h25m. If that works, a v6.19..v7.0 bisection becomes
> an evening rather than a fortnight, and you would also have a reproducer that is
> practical to run in a lab.

We do have a reproducer in our lab environment that will rapidly 
allocate and trip this issue which is how we could analyze it and root 
cause it.

> 
> I will try that tonight, with all services stopped, then the
> amd_iommu=pgtbl_v2 iommu.forcedac=1 combination Vasant asked for. I will report
> bytes re-read, wall time and workload rather than a verdict.
> 
> If you already know that forcedac alone will not behave that way, please tell me
> and I will not waste the evening on it.
> 
> And to be straight about the rest: this is a production home server, so a
> two-week bisection is a real cost for me. The reproducer is three lines of fio.
> If you can run it on comparable hardware in your lab, that would very likely be
> faster than me doing it here -- and I am happy to run any specific kernel or
> debug patch you want in the meantime.

I don't yet have any confirmation we can patch this at runtime.  If I do 
come up with a way to do that which works will let you know.

By chance did this issue coincide with you switching from something 
different to the Ryzen 7 8700G?  For example switching from Raphael or 
Granite Ridge parts to that Phoenix part.

  reply	other threads:[~2026-08-28 15:38 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <CAHEgv3TiJUvS=Hustyem0ts8jkbFVAdzBazmPhCbv2sNnTsZkg@mail.gmail.com>
2026-08-28  5:43 ` [REGRESSION] Silent SATA read corruption with dma-iommu on AMD 600-series AHCI (6.19 good, 7.0+ bad) Vasant Hegde
2026-08-28 11:02   ` Mikael Etienne
2026-08-28 12:52     ` Mario Limonciello
2026-08-28 15:24       ` Mikael Etienne
2026-08-28 15:37         ` Mario Limonciello [this message]
2026-08-28 16:42           ` Mikael Etienne
2026-08-28  4:56 Mikael Etienne

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=d49add14-c8de-4b6b-9049-4a07e5692ee9@amd.com \
    --to=mario.limonciello@amd.com \
    --cc=iommu@lists.linux.dev \
    --cc=joro@8bytes.org \
    --cc=linux-block@vger.kernel.org \
    --cc=linux-ide@vger.kernel.org \
    --cc=mikael1022bzh@gmail.com \
    --cc=regressions@lists.linux.dev \
    --cc=suravee.suthikulpanit@amd.com \
    --cc=vasant.hegde@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.