From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [192.198.163.13]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 405BF35C193 for ; Mon, 24 Aug 2026 21:32:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=192.198.163.13 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787607176; cv=none; b=Vx6fQQBsCa5JJ+2TmzvEIpPOuC4Nua6hVH6oIOxPn6cRuuxccnSistCwmgAkUCKcG7PPct9Z0z3/e/GqfZYJG4W9wh+2D8CDL7tdghzWu12XzgRxARdfIjJnMyrLJf6jVh2H0AqarUc66h34+RE2YsroXbiKPtpUaGp3NzXTTL0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787607176; c=relaxed/simple; bh=rUuAq1ugvnUq9BIrNxm33Yk0Ld+9rTjVp/7XAOUCIgU=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=R0U5c0Z17nh+GCYLcF70OwDNZpTlOpIg17y0ZHAR0lprej5kfHy4byO5oPqy1WfbDXHzh0nUzt6O++kSd0xWVLwrQ3yOqXmW/ek5i26ZzHFTGDg/LwTuoC2FtcqmjPPKTe51eqqavcWexO1xxCHYLdWJ9DIP2xSejqaEMaYPN30= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=Z2htSF8T; arc=none smtp.client-ip=192.198.163.13 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="Z2htSF8T" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1787607174; x=1819143174; h=message-id:date:mime-version:subject:to:cc:references: from:in-reply-to:content-transfer-encoding; bh=rUuAq1ugvnUq9BIrNxm33Yk0Ld+9rTjVp/7XAOUCIgU=; b=Z2htSF8TQOJA0lFZuFJuTYbXnJW1SEjCOY9xG0j152c0bQ5zWxYNu8l6 yyhQxcZc4z1WYUzd2C8JQTjp2xOfBOo7p/BYSlGEnQQUr0OABZeSzd/ks RK/NGaRNQ6sy5HiXXn5MlyiUxwzQh5u1ihSnqZoRc/JXQrpmDVKdR4yze kyY3n8zFu1HPjXkTygOB8Cz69ozzhzClzKv8OUcJS2JtmvsXny9+tlr9/ bgir75uCiTkRRWE33JCPneXK7ToKvgB2zdimJ69jHEppj+MYDDmJz73Hi +iuFPrQ7Tya3+jK61eZfPBwGjmeNXlo53E500xpM9PSITAEzZGXxTaw4K A==; X-CSE-ConnectionGUID: PSHv7ZJKSEC4ENWwH/4zVQ== X-CSE-MsgGUID: WGUeRDO8R5Kb69YjeMpv2g== X-IronPort-AV: E=McAfee;i="6800,10657,11885"; a="90584312" X-IronPort-AV: E=Sophos;i="6.25,241,1779174000"; d="scan'208";a="90584312" Received: from orviesa008.jf.intel.com ([10.64.159.148]) by fmvoesa107.fm.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 24 Aug 2026 14:32:53 -0700 X-CSE-ConnectionGUID: CW7u7kl2SKGyIKRWRkh/mA== X-CSE-MsgGUID: 652A0P5wQKG1X3F91NlRrQ== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,241,1779174000"; d="scan'208";a="266582054" Received: from jjgreens-desk24.amr.corp.intel.com (HELO [10.125.108.201]) ([10.125.108.201]) by orviesa008-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 24 Aug 2026 14:32:53 -0700 Message-ID: <3189ffc8-d3a5-44e4-9301-4c98189cbab0@intel.com> Date: Mon, 24 Aug 2026 14:32:52 -0700 Precedence: bulk X-Mailing-List: linux-cxl@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v4 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata To: sashiko-reviews@lists.linux.dev Cc: linux-cxl@vger.kernel.org References: <20260824174936.939059-1-dave.jiang@intel.com> <20260824174936.939059-13-dave.jiang@intel.com> <20260824181753.F3F681F000E9@smtp.kernel.org> From: Dave Jiang Content-Language: en-US In-Reply-To: <20260824181753.F3F681F000E9@smtp.kernel.org> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 8/24/26 11:17 AM, sashiko-bot@kernel.org wrote: > Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider: > > New issues: > - [High] Hardcoding `header_len = 0` and `flit = false` in `ghes_handle_aer()` actively breaks the parsing and printing of legitimate Flit mode TLP logs from CPER AER records. > > Pre-existing issues: > - [High] Memory allocated for `aer_capability_regs` from the fixed-size `ghes_estatus_pool` is permanently leaked when `aer_recover_queue()` fails to insert into the `aer_recover_ring`. > - [High] Memory allocated for `aer_capability_regs` is leaked in `aer_recover_work_func()` when the target PCI device cannot be found. > - [High] Direct `memcpy` of raw CPER AER hardware bytes into `struct aer_capability_regs` misaligns fields, causing hardware registers to be incorrectly parsed and printed as TLP Prefix Logs. > - [High] Uninitialized and unvalidated software-only fields in `cxl_rch_get_aer_info()` lead to a massive out-of-bounds kernel stack read in `pcie_print_tlp_log()`. > - [High] Direct `memcpy` of little-endian CPER AER info into native integer fields of `struct aer_capability_regs` corrupts values on big-endian architectures. > -- > > --- Patch [12]: [PATCH v4 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata --- > commit 91c35473296c2f22c740c76f30c6c146671ec39a > Author: Dave Jiang > > ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata > > ghes_handle_aer() copies sizeof(struct aer_capability_regs) out of the > fixed 96-byte pcie_err->aer_info. The struct is larger, so the copy reads > past the section. Validate error_data_length, zero the destination, > bound the copy to the 96-byte source, and clear header_len and flit. > > Closes: https://sashiko.dev/#/patchset/20260714231835.303081-1-dave.jiang@intel.com?part=3 > >> diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c >> index 0cc1e63836356..d6643127a2df0 100644 >> --- a/drivers/acpi/apei/ghes.c >> +++ b/drivers/acpi/apei/ghes.c > > [ ... ] > >> @@ -664,13 +667,22 @@ static void ghes_handle_aer(struct acpi_hest_generic_data *gdata) >> sizeof(struct aer_capability_regs)); >> if (!aer_info) >> return; >> - memcpy(aer_info, pcie_err->aer_info, sizeof(struct aer_capability_regs)); >> + >> + /* >> + * The CPER source is a fixed 96 bytes, shorter than struct >> + * aer_capability_regs, so bound the copy to it. header_len and >> + * flit are software-only and land inside those 96 bytes; clear >> + * them so firmware cannot drive the pcie_print_tlp_log() loop >> + * over dw[] out of bounds. >> + */ >> + memset(aer_info, 0, sizeof(struct aer_capability_regs)); >> + memcpy(aer_info, pcie_err->aer_info, sizeof(pcie_err->aer_info)); > > [Severity: High] > This isn't a bug introduced by this patch, but does this direct memcpy > misalign fields in struct aer_capability_regs? > > The CPER hardware payload expects Root Error Command and Root Error Status > at offset 44 and 48 respectively. However, struct aer_capability_regs > embeds the software struct pcie_tlp_log at offset 28. > > Because pcie_tlp_log is padded for Flit mode and takes up ~60 bytes, doesn't > this incorrectly place the hardware's Root Error Command into > aer_info->header_log.dw[4], which aliases to prefix[0]? > > This would cause pcie_print_tlp_log() to dump entirely unrelated hardware > registers as TLP Prefix logs. > > [Severity: High] > This is another pre-existing issue, but does this direct copy of little-endian > CPER AER info into native integer fields corrupt the values on big-endian > architectures? > > Without le32_to_cpu() conversions, parsing these fields on a big-endian > system will interpret the bytes in reverse order, breaking downstream > bitwise evaluations. > >> + aer_info->header_log.header_len = 0; >> + aer_info->header_log.flit = false; > > [Severity: High] > Does hardcoding header_len and flit to 0 and false unconditionally break the > parsing and printing of legitimate Flit mode TLP logs from CPER AER records? > > If the hardware actually experienced an error in Flit mode, indicated by > PCI_ERR_CAP_TLP_LOG_FLIT in aer_info->cap_control, this forcibly discards > that state. Won't pcie_print_tlp_log() then ignore the extended Flit mode > header DWORDS entirely? It will, but there is nothing to discard. Flit mode logs are not decoded on this path today, and this patch does not change that. What it changes is that the failure stops being an out of bounds read. DJ > >> >> aer_recover_queue(pcie_err->device_id.segment, >> pcie_err->device_id.bus, >> - devfn, aer_severity, >> - (struct aer_capability_regs *) >> - aer_info); >> + devfn, aer_severity, aer_info); >> } >> #endif >> } > > [Severity: High] > This is a pre-existing issue, but does this code leak the memory allocated > for aer_info from the fixed-size ghes_estatus_pool if aer_recover_queue() > fails to insert into the aer_recover_ring? > > drivers/pci/pcie/aer.c:aer_recover_queue() { > ... > if (kfifo_in_spinlocked(&aer_recover_ring, &entry, 1, > &aer_recover_ring_lock)) > schedule_work(&aer_recover_work); > else > pr_err("buffer overflow in recovery for ...\n", ...); > } > > Since aer_recover_queue() returns void, ghes_handle_aer() cannot detect > the failure to perform cleanup, permanently leaking the pool memory. > > [Severity: High] > This is also a pre-existing issue, but is the memory for aer_info leaked > in aer_recover_work_func() if the target PCI device cannot be found? > > drivers/pci/pcie/aer.c:aer_recover_work_func() { > ... > pdev = pci_get_domain_bus_and_slot(entry.domain, entry.bus, > entry.devfn); > if (!pdev) { > pr_err_ratelimited("... no pci_dev found\n"); > continue; > } > ... > } > > The continue statement skips the ghes_estatus_pool_region_free() cleanup > call at the bottom of the loop, permanently leaking the memory. > > [Severity: High] > This isn't a bug in this patch either, but does a similar issue with > uninitialized software-only fields exist in cxl_rch_get_aer_info()? > > drivers/cxl/core/ras_rch.c:cxl_rch_get_aer_info() { > ... > for (n = 0; n < read_cnt; n++) > aer_regs_buf[n] = readl(aer_base + n * sizeof(u32)); > ... > } > > If it reads from unimplemented MMIO space, it returns 0xFFFFFFFF, which > populates the software-only metadata fields on the stack such that flit > evaluates to true and header_len becomes 255. > > Doesn't this lead to an out-of-bounds stack read later? > > drivers/pci/pcie/tlp.c:pcie_print_tlp_log() { > ... > for (i = PCIE_STD_NUM_TLP_HEADERLOG; i < log->header_len; i++) { > len += scnprintf(buf + len, sizeof(buf) - len, > " %#010x", log->dw[i]); > ... > } >