From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 8E702C9832F for ; Mon, 28 Sep 2026 04:33:10 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender: Content-Transfer-Encoding:Content-Type:List-Subscribe:List-Help:List-Post: List-Archive:List-Unsubscribe:List-Id:In-Reply-To:MIME-Version:References: Message-ID:Subject:Cc:To:From:Date:Reply-To:Content-ID:Content-Description: Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID: List-Owner; bh=A6iwW+YRNUkvLpA1P40bX2bbwTtUqYSXyGxtxyeM17k=; b=L6cGgf0VgkRvbm E9hiSh4x7KUQ32kjV1HKsnOLv1o/pC5bMmHZXrcMIQLASeDsoSnPeotm/mVzA+KwybXVvrnaMly6d 77ZG3GsnRq5qTbo0Q+NDjqS50HfkZ8xftAtX2TStQ3p3je9gd/KEjNWqWxRBPcXM/zEWXQ+nO1rT6 m3J1pOVCnHhXViFk3zeLNeKhWeWT18JNBiEGsr+aL4bLHuBcu3kAnlTXl5gt9SYeq8rvrnCHhav2O VaLpgdd4H1LWAlxgxvrSvueLsXZzFe/wRlAw9Q/7NttM3oXHsExRGgFHCXH98glEaMqP9Xs85JOqk ZDkz8k8nBSFTrVIbsESQ==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.99.1 #2 (Red Hat Linux)) id 1xB32l-0000000HJsN-16yu; Mon, 28 Sep 2026 04:32:47 +0000 Received: from hall.aurel32.net ([2001:bc8:30d7:100::1]) by bombadil.infradead.org with esmtps (Exim 4.99.1 #2 (Red Hat Linux)) id 1xB32i-0000000HJrr-1gGT for linux-riscv@lists.infradead.org; Mon, 28 Sep 2026 04:32:45 +0000 DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=aurel32.net ; s=202004.hall; h=In-Reply-To:Content-Type:MIME-Version:References: Message-ID:Subject:Cc:To:From:Date:Content-Transfer-Encoding:From:Reply-To: Subject:Content-ID:Content-Description:X-Debbugs-Cc; bh=xmk5euR5PRSTE9/TBR12w66CsbiVK597hUhgaveMVSI=; b=qLyctAVp1Ua5wXeJGUXKJ3Gjmz HrVDu7pCr/x3SvuT9+XlwcYr+z67OYbjLv2TjlJStsEK+ry0yY+0Q4UUakGUf5qam31CSR9S28PJX Ql5XPnLiwKIfcb7itfyi2e5pZsZ4GqJXYIUSmV1d4yabKUYDlB9mY+40u85urWVSodj3BS9Zf9GdR 1btAfOS16NAuCEt5xQHIO96qMxIpLZEIt8q4LdyHnWSg5Zv/tVBpO97ZUSUCWAF3AUgD78M88PjF2 21oXle8SIFYmh5x436tn6eRajI1PJJIkxiMr/nt84t3WKbypyOog0+jru82kFoRCzNMbz6wBixSks nx7Q4ckg==; Received: from authenticated user by hall.aurel32.net with esmtpsa (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1xB32e-0000000E2Dv-3397; Mon, 28 Sep 2026 06:32:40 +0200 Date: Mon, 28 Sep 2026 06:32:40 +0200 From: Aurelien Jarno To: Karl Mehltretter Cc: Andy Chiu , spacemit@lists.linux.dev, linux-riscv@lists.infradead.org Subject: Re: Random corruption on SpacemiT K1 (and K3) with RVV Message-ID: References: MIME-Version: 1.0 Content-Disposition: inline In-Reply-To: User-Agent: Mutt/2.4.1 (2026-07-04) X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.9.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20260927_213244_442867_14F35D49 X-CRM114-Status: GOOD ( 29.31 ) X-BeenThere: linux-riscv@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: 7bit Sender: "linux-riscv" Errors-To: linux-riscv-bounces+linux-riscv=archiver.kernel.org@lists.infradead.org Hi, On 2026-09-25 06:40, Aurelien Jarno wrote: > Hi Karl, > > On 2026-09-24 08:15, Karl Mehltretter wrote: > > On Thu, Sep 24, 2026 at 06:52:21AM +0100, Aurelien Jarno wrote: > > > Thanks for your feedback. Note that at this stage I have not been able > > > to reproduce the issue with QEMU. I guess it's very timing dependent, > > > also I am not sure if QEMU simulates partially executed instructions > > > (outside of page faults). > > > > > > > Hello Aurelien, Andy, > > > > I tested this with a local TCG diagnostic change. It did not reproduce > > the K1 failure, but it answers the partial-instruction question. > > > > My LLM agent helped me running these tests. > > > > In stock TCG at 7074591d7954, vle8.v can leave partial state on a memory > > fault, but the vector helper runs atomically with respect to guest > > interrupts [1,2]. Stock QEMU therefore cannot take an asynchronous > > interrupt partway through this load. > > > > In a bare-metal test at VLEN=256 and LMUL=8, a page fault after element > > 16 left vstart=16 and a snapshot containing 16 source bytes followed by > > 240 poison bytes. After the missing page was mapped, QEMU resumed the > > load and completed the copy correctly. > > > > I then added a hook to QEMU vector-load which performs 16 elements, sets > > vstart=16, raises a timer interrupt, and resumes at the same vle8.v. The > > bare-metal test completed 262,145 such restarts and 67,108,864 copied > > bytes without a mismatch. > > > > I also booted Linux 388b607d107c with: > > > > CONFIG_RISCV_ISA_V=y > > CONFIG_RISCV_ISA_V_UCOPY_THRESHOLD=1 > > CONFIG_RISCV_ISA_V_PREEMPTIVE=n > > > > With the hook restricted to S-mode loads with SUM and SIE set, a checked > > pipe test completed 90,308,608 bytes across copy_from_user() and > > copy_to_user(), including demand-faulting source and destination pages. > > The hook logged at least 327,680 forced interruptions at vstart=16, > > without a byte mismatch. > > > > As a negative control, I made one load read 16 elements and then retire > > as if all 256 had completed. That produced the 16-source/240-poison > > signature in bare metal, and the Linux checker reported exactly 240 > > differing bytes. This is an injected symptom, but confirms that the test > > detects the reported failure shape. > > > > Correct QEMU fault recovery and forced interrupt restart therefore did > > not produce the corruption. Reaching the 16/240 result required > > deliberately modelling a load that completed early without a trap. That > > fits the observed first load/store pair, but does not establish its > > cause. Your IRQ-disabled result also makes a normal asynchronous restart > > a poor fit. > > Thanks for all those extensive tests. > > > The normal Linux load-fault path exits before vse8.v and falls back to > > the scalar copy [3,4]. Do you have the original trap PC, cause and fault > > address for the second-iteration fault? Those values could show whether > > an unexpected synchronous trap is involved. > > Unfortunately the issue is quite rare, so I have not found a way to > trace all that information when the issue happens. > > That said Han Gao pointed me to this patch: > https://lore.kernel.org/linux-riscv/20260807-vector_fpu_regs_status_rmw_fix-v1-1-0c16848b60db@intel.com/ > > I have been testing it, and so far it seems to fix my issue with both > CONFIG_RISCV_ISA_V_PREEMPTIVE enable and disabled. Or maybe it just > hides the issue by changing the timing, as at this point, I haven't > fully understood how the bug fixed by this patch can completely explain > my observations. Unfortunately, I have been able to reproduce the issue, even with this patch applied. It just significantly reduces the frequency of the issue. Regards Aurelien -- Aurelien Jarno GPG: 4096R/1DDD8C9B aurelien@aurel32.net http://aurel32.net _______________________________________________ linux-riscv mailing list linux-riscv@lists.infradead.org http://lists.infradead.org/mailman/listinfo/linux-riscv