From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from lists.ozlabs.org (lists.ozlabs.org [112.213.38.117]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 4ABCAC83F26 for ; Thu, 24 Jul 2025 08:00:23 +0000 (UTC) Received: from boromir.ozlabs.org (localhost [127.0.0.1]) by lists.ozlabs.org (Postfix) with ESMTP id 4bnk2G05ylz30T9; Thu, 24 Jul 2025 18:00:22 +1000 (AEST) Authentication-Results: lists.ozlabs.org; arc=none smtp.remote-ip=115.124.30.124 ARC-Seal: i=1; a=rsa-sha256; d=lists.ozlabs.org; s=201707; t=1753344021; cv=none; b=FQNtMtaOWugC9c4dJ6rvuGW98F0bWf2GaWFQk4qS5WFTyIRfN4IDMrOkFB9YhYpZt7yFUev/3SqwV8Vfmz1/EgvW/MDLJri0rsaEiTwC0ap70nPZSqFZpDVh9PrOm0nTsVdWAUesmMEXI2gcm1wQenadBr7b5diVNPhukwTNUNlisPMkNFyJND6T2kKjugt1tfC81TQFaQ/z4ZTN8+GlO6L2AMxRJyz6v7tL21cpdabtb6Zaju33bLhBTzKcjwWkYqsMBlCJfiY7Vv446c+tjzWEpjwHJZGytd7U+SyavDFfEHVZZzp/0iRmxCslRhDCIbs9/L0f4Y2j5WlkSkqgEw== ARC-Message-Signature: i=1; a=rsa-sha256; d=lists.ozlabs.org; s=201707; t=1753344021; c=relaxed/relaxed; bh=nhf3IJcOYSpRMKrPjWNbbcWsyyetTfh18hHD10/wjt4=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=iR6GEUqggkDah6qZZL9sChE9Ze8QDMQfOv3tlzQZ3BLtMDAWJ0IZ2ku8wJ5+MzJg0fkviocvMWJ/XHof23Hhmj+Ug7AkjPnvGA8bnB3h6/quowSLrkbO6dG9Pw52MTE3jeRIJb2vVWHLYajgL9EbMyP8tXbtYsaTNKC70tFHr7DjKMYNDChc23+ADQ2myamQ1dxwHXU1LpxWGEkW90nM7gyAF1Kxa2Yf1We5KBvz7WsKMKW5CoFkNXjygJMirK28Zu7LGQzFqbAG98DnyzUvihUKv8Ex9d4lIV1MRK1KpVojeoifJS65rFs0KBBVJXvOiRu0z8MbntpRbyJPrVPANQ== ARC-Authentication-Results: i=1; lists.ozlabs.org; dmarc=pass (p=none dis=none) header.from=linux.alibaba.com; dkim=pass (1024-bit key; unprotected) header.d=linux.alibaba.com header.i=@linux.alibaba.com header.a=rsa-sha256 header.s=default header.b=NPavVHCQ; dkim-atps=neutral; spf=pass (client-ip=115.124.30.124; helo=out30-124.freemail.mail.aliyun.com; envelope-from=xueshuai@linux.alibaba.com; receiver=lists.ozlabs.org) smtp.mailfrom=linux.alibaba.com Authentication-Results: lists.ozlabs.org; dmarc=pass (p=none dis=none) header.from=linux.alibaba.com Authentication-Results: lists.ozlabs.org; dkim=pass (1024-bit key; unprotected) header.d=linux.alibaba.com header.i=@linux.alibaba.com header.a=rsa-sha256 header.s=default header.b=NPavVHCQ; dkim-atps=neutral Authentication-Results: lists.ozlabs.org; spf=pass (sender SPF authorized) smtp.mailfrom=linux.alibaba.com (client-ip=115.124.30.124; helo=out30-124.freemail.mail.aliyun.com; envelope-from=xueshuai@linux.alibaba.com; receiver=lists.ozlabs.org) Received: from out30-124.freemail.mail.aliyun.com (out30-124.freemail.mail.aliyun.com [115.124.30.124]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by lists.ozlabs.org (Postfix) with ESMTPS id 4bnk2C6R49z30T8 for ; Thu, 24 Jul 2025 18:00:18 +1000 (AEST) DKIM-Signature:v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.alibaba.com; s=default; t=1753344014; h=Message-ID:Date:MIME-Version:Subject:To:From:Content-Type; bh=nhf3IJcOYSpRMKrPjWNbbcWsyyetTfh18hHD10/wjt4=; b=NPavVHCQ1VNVVQDWlB0oVFrocNwFe6TTnNsoK4RMISK18ozqDOoBMzvQsmovDdtgQGEpVtY4bt/Z6Ih9K7g3Pmp9H6ykuw+tiaFpqUXIT+lK/v0+0ZhxBey/8TTFlB3HgzFzj10rNxKQBHVh5TvP7u2wRtfuSkPjJeEv+U3bLSA= Received: from 30.246.181.19(mailfrom:xueshuai@linux.alibaba.com fp:SMTPD_---0Wjlb2GV_1753344010 cluster:ay36) by smtp.aliyun-inc.com; Thu, 24 Jul 2025 16:00:11 +0800 Message-ID: <7ce9731a-b212-4e27-8809-0559eb36c5f2@linux.alibaba.com> Date: Thu, 24 Jul 2025 16:00:09 +0800 X-Mailing-List: linuxppc-dev@lists.ozlabs.org List-Id: List-Help: List-Owner: List-Post: List-Archive: , List-Subscribe: , , List-Unsubscribe: Precedence: list MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v3] vmcoreinfo: Track and log recoverable hardware errors To: Breno Leitao , "Rafael J. Wysocki" , Len Brown , James Morse , Tony Luck , Borislav Petkov , Robert Moore , Thomas Gleixner , Ingo Molnar , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Hanjun Guo , Mauro Carvalho Chehab , Mahesh J Salgaonkar , Oliver O'Halloran , Bjorn Helgaas Cc: linux-acpi@vger.kernel.org, linux-kernel@vger.kernel.org, acpica-devel@lists.linux.dev, osandov@osandov.com, konrad.wilk@oracle.com, linux-edac@vger.kernel.org, linuxppc-dev@lists.ozlabs.org, linux-pci@vger.kernel.org, kernel-team@meta.com References: <20250722-vmcore_hw_error-v3-1-ff0683fc1f17@debian.org> From: Shuai Xue In-Reply-To: <20250722-vmcore_hw_error-v3-1-ff0683fc1f17@debian.org> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit Hi, Breno, 在 2025/7/23 00:56, Breno Leitao 写道: > Introduce a generic infrastructure for tracking recoverable hardware > errors (HW errors that did not cause a panic) and record them for vmcore > consumption. This aids post-mortem crash analysis tools by preserving > a count and timestamp for the last occurrence of such errors. > > Add centralized logging for three common sources of recoverable hardware > errors: The term "recoverable" is highly ambiguous. Even within the x86 architecture, different vendors define errors differently. I'm not trying to be pedantic about classification. As far as I know, for 2-bit memory errors detected by scrub, AMD defines them as deferred errors (DE) and handles them with log_error_deferred, while Intel uses machine_check_poll. For 2-bit memory errors consumed by processes, both Intel and AMD use MCE handling viado_machine_check(). Does your HWERR_RECOV_MCE only focus on synchronous UE errors handled in do_machine_check? What makes it special? > > - PCIe AER Correctable errors > - x86 Machine Check Exceptions (MCE) > - APEI/CPER GHES corrected or recoverable errors > > hwerror_data is write-only at kernel runtime, and it is meant to be > read from vmcore using tools like crash/drgn. For example, this is how > it looks like when opening the crashdump from drgn. > > >>> prog['hwerror_data'] > (struct hwerror_info[3]){ > { > .count = (int)844, > .timestamp = (time64_t)1752852018, > }, > ... > > This helps fleet operators quickly triage whether a crash may be > influenced by hardware recoverable errors (which executes a uncommon > code path in the kernel), especially when recoverable errors occurred > shortly before a panic, such as the bug fixed by > commit ee62ce7a1d90 ("page_pool: Track DMA-mapped pages and unmap them > when destroying the pool") > > This is not intended to replace full hardware diagnostics but provides > a fast way to correlate hardware events with kernel panics quickly. > > Suggested-by: Tony Luck > Signed-off-by: Breno Leitao > --- > Changes in v3: > - Add more information about this feature in the commit message > (Borislav Petkov) > - Renamed the function to hwerr_log_error_type() and use hwerr as > suffix (Borislav Petkov) > - Make the empty function static inline (kernel test robot) > - Link to v2: https://lore.kernel.org/r/20250721-vmcore_hw_error-v2-1-ab65a6b43c5a@debian.org > > Changes in v2: > - Split the counter by recoverable error (Tony Luck) > - Link to v1: https://lore.kernel.org/r/20250714-vmcore_hw_error-v1-1-8cf45edb6334@debian.org > --- > arch/x86/kernel/cpu/mce/core.c | 3 +++ > drivers/acpi/apei/ghes.c | 8 ++++++-- > drivers/pci/pcie/aer.c | 2 ++ > include/linux/vmcore_info.h | 14 ++++++++++++++ > kernel/vmcore_info.c | 18 ++++++++++++++++++ > 5 files changed, 43 insertions(+), 2 deletions(-) > > diff --git a/arch/x86/kernel/cpu/mce/core.c b/arch/x86/kernel/cpu/mce/core.c > index 4da4eab56c81d..cb225a42eebbb 100644 > --- a/arch/x86/kernel/cpu/mce/core.c > +++ b/arch/x86/kernel/cpu/mce/core.c > @@ -45,6 +45,7 @@ > #include > #include > #include > +#include > > #include > #include > @@ -1692,6 +1693,8 @@ noinstr void do_machine_check(struct pt_regs *regs) > out: > instrumentation_end(); > > + /* Given it didn't panic, mark it as recoverable */ > + hwerr_log_error_type(HWERR_RECOV_MCE); > clear: > mce_wrmsrq(MSR_IA32_MCG_STATUS, 0); > } > diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c > index a0d54993edb3b..ebda2aa3d68f2 100644 > --- a/drivers/acpi/apei/ghes.c > +++ b/drivers/acpi/apei/ghes.c > @@ -43,6 +43,7 @@ > #include > #include > #include > +#include > > #include > #include > @@ -1136,13 +1137,16 @@ static int ghes_proc(struct ghes *ghes) > { > struct acpi_hest_generic_status *estatus = ghes->estatus; > u64 buf_paddr; > - int rc; > + int rc, sev; > > rc = ghes_read_estatus(ghes, estatus, &buf_paddr, FIX_APEI_GHES_IRQ); > if (rc) > goto out; > > - if (ghes_severity(estatus->error_severity) >= GHES_SEV_PANIC) > + sev = ghes_severity(estatus->error_severity); > + if (sev == GHES_SEV_RECOVERABLE || sev == GHES_SEV_CORRECTED) > + hwerr_log_error_type(HWERR_RECOV_GHES); APEI does not define an error type named GHES. GHES is just a kernel driver name. Many hardware error types can be handled in GHES (see ghes_do_proc), for example, AER is routed by GHES when firmware-first mode is used. As far as I know, firmware-first mode is commonly used in production. Should GHES errors be categorized into AER, memory, and CXL memory instead? Thanks. Shuai