Linux-RISC-V Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Aurelien Jarno <aurelien@aurel32.net>
To: Karl Mehltretter <kmehltretter@gmail.com>
Cc: Andy Chiu <tchiu@tenstorrent.com>,
	spacemit@lists.linux.dev, linux-riscv@lists.infradead.org
Subject: Re: Random corruption on SpacemiT K1 (and K3) with RVV
Date: Fri, 25 Sep 2026 06:40:31 +0200	[thread overview]
Message-ID: <arX7PwOnlz0N9rUc@aurel32.net> (raw)
In-Reply-To: <arS-yYLK8Df5knSE@gmail.com>

Hi Karl,

On 2026-09-24 08:15, Karl Mehltretter wrote:
> On Thu, Sep 24, 2026 at 06:52:21AM +0100, Aurelien Jarno wrote:
> > Thanks for your feedback. Note that at this stage I have not been able 
> > to reproduce the issue with QEMU. I guess it's very timing dependent, 
> > also I am not sure if QEMU simulates partially executed instructions 
> > (outside of page faults).
> > 
> 
> Hello Aurelien, Andy,
> 
> I tested this with a local TCG diagnostic change. It did not reproduce
> the K1 failure, but it answers the partial-instruction question.
> 
> My LLM agent helped me running these tests.
> 
> In stock TCG at 7074591d7954, vle8.v can leave partial state on a memory
> fault, but the vector helper runs atomically with respect to guest
> interrupts [1,2]. Stock QEMU therefore cannot take an asynchronous
> interrupt partway through this load.
> 
> In a bare-metal test at VLEN=256 and LMUL=8, a page fault after element
> 16 left vstart=16 and a snapshot containing 16 source bytes followed by
> 240 poison bytes. After the missing page was mapped, QEMU resumed the
> load and completed the copy correctly.
> 
> I then added a hook to QEMU vector-load which performs 16 elements, sets
> vstart=16, raises a timer interrupt, and resumes at the same vle8.v. The
> bare-metal test completed 262,145 such restarts and 67,108,864 copied
> bytes without a mismatch.
> 
> I also booted Linux 388b607d107c with:
> 
>   CONFIG_RISCV_ISA_V=y
>   CONFIG_RISCV_ISA_V_UCOPY_THRESHOLD=1
>   CONFIG_RISCV_ISA_V_PREEMPTIVE=n
> 
> With the hook restricted to S-mode loads with SUM and SIE set, a checked
> pipe test completed 90,308,608 bytes across copy_from_user() and
> copy_to_user(), including demand-faulting source and destination pages.
> The hook logged at least 327,680 forced interruptions at vstart=16,
> without a byte mismatch.
> 
> As a negative control, I made one load read 16 elements and then retire
> as if all 256 had completed. That produced the 16-source/240-poison
> signature in bare metal, and the Linux checker reported exactly 240
> differing bytes. This is an injected symptom, but confirms that the test
> detects the reported failure shape.
> 
> Correct QEMU fault recovery and forced interrupt restart therefore did
> not produce the corruption. Reaching the 16/240 result required
> deliberately modelling a load that completed early without a trap. That
> fits the observed first load/store pair, but does not establish its
> cause. Your IRQ-disabled result also makes a normal asynchronous restart
> a poor fit.

Thanks for all those extensive tests.

> The normal Linux load-fault path exits before vse8.v and falls back to
> the scalar copy [3,4]. Do you have the original trap PC, cause and fault
> address for the second-iteration fault? Those values could show whether
> an unexpected synchronous trap is involved.

Unfortunately the issue is quite rare, so I have not found a way to 
trace all that information when the issue happens.

That said Han Gao pointed me to this patch:
https://lore.kernel.org/linux-riscv/20260807-vector_fpu_regs_status_rmw_fix-v1-1-0c16848b60db@intel.com/

I have been testing it, and so far it seems to fix my issue with both 
CONFIG_RISCV_ISA_V_PREEMPTIVE enable and disabled. Or maybe it just 
hides the issue by changing the timing, as at this point, I haven't 
fully understood how the bug fixed by this patch can completely explain 
my observations.

I still observe a few GCC crashes, on both the K3, but more rarely, but 
without this corruption pattern, just as non reproducible ICE in GCC. It 
is likely a different issue.

Regards
Aurelien

-- 
Aurelien Jarno                          GPG: 4096R/1DDD8C9B
aurelien@aurel32.net                     http://aurel32.net

_______________________________________________
linux-riscv mailing list
linux-riscv@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-riscv

  reply	other threads:[~2026-09-25  4:40 UTC|newest]

Thread overview: 12+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-30 20:52 Random corruption on SpacemiT K1 with RVV and THP Aurelien Jarno
2026-09-08  4:42 ` Random corruption on SpacemiT K1 with RVV Aurelien Jarno
2026-09-09 16:45   ` Random corruption on SpacemiT K1 (and K3) " Aurelien Jarno
2026-09-15 21:14     ` Aurelien Jarno
2026-09-17 19:18       ` Palmer Dabbelt
2026-09-18  4:34         ` Aurelien Jarno
2026-09-22 21:33     ` Andy Chiu
2026-09-24  4:52       ` Aurelien Jarno
2026-09-24  6:15         ` Karl Mehltretter
2026-09-25  4:40           ` Aurelien Jarno [this message]
2026-09-28  4:32             ` Aurelien Jarno
2026-09-29  4:07               ` Vivian Wang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=arX7PwOnlz0N9rUc@aurel32.net \
    --to=aurelien@aurel32.net \
    --cc=kmehltretter@gmail.com \
    --cc=linux-riscv@lists.infradead.org \
    --cc=spacemit@lists.linux.dev \
    --cc=tchiu@tenstorrent.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox