From: sashiko-bot@kernel.org
To: "Dave Jiang" <dave.jiang@intel.com>
Cc: linux-cxl@vger.kernel.org
Subject: Re: [PATCH v6 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata
Date: Fri, 04 Sep 2026 18:08:41 +0000 [thread overview]
Message-ID: <20260904180842.15A9C1F00A3D@smtp.kernel.org> (raw)
In-Reply-To: <20260904172337.1409775-13-dave.jiang@intel.com>
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
Pre-existing issues:
- [High] Memory leak of `aer_capability_regs` allocated from `ghes_estatus_pool` in error paths of `aer_recover_queue` and `aer_recover_work_func`.
--
commit 639b51ace1216953e300ae9efcce13516dadf849
Author: Dave Jiang <dave.jiang@intel.com>
ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata
This commit validates the error_data_length in ghes_handle_aer() and ensures
that copying firmware error information into the larger aer_capability_regs
struct does not read out of bounds. It also zeroes out the destination and
properly maps the header log and prefix log to prevent firmware from
manipulating software-only fields.
Closes: https://sashiko.dev/#/patchset/20260714231835.303081-1-dave.jiang@intel.com?part=3
Fixes: 7e077e6707b3 ("PCI/ERR: Handle TLP Log in Flit mode")
> diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c
> index 2077cc90569a0..08c985e729d6c 100644
> --- a/drivers/acpi/apei/ghes.c
> +++ b/drivers/acpi/apei/ghes.c
[ ... ]
> @@ -664,13 +667,25 @@ static void ghes_handle_aer(struct acpi_hest_generic_data *gdata)
> sizeof(struct aer_capability_regs));
> if (!aer_info)
> return;
> - memcpy(aer_info, pcie_err->aer_info, sizeof(struct aer_capability_regs));
> +
> + /*
> + * Map aer_info onto the struct as extlog_print_pcie() does:
> + * copy up to the four Header Log DWORDs, then place the TLP
> + * Prefix Log from where the hardware keeps it. The rest stays
> + * zero, so firmware cannot drive the pcie_print_tlp_log() loop
> + * over dw[] out of bounds.
> + */
> + memset(aer_info, 0, sizeof(struct aer_capability_regs));
> + memcpy(aer_info, pcie_err->aer_info,
> + offsetof(struct aer_capability_regs, header_log) +
> + PCIE_STD_NUM_TLP_HEADERLOG * sizeof(u32));
> + memcpy(aer_info->header_log.prefix,
> + pcie_err->aer_info + PCI_ERR_PREFIX_LOG,
> + sizeof(aer_info->header_log.prefix));
>
> aer_recover_queue(pcie_err->device_id.segment,
> pcie_err->device_id.bus,
> - devfn, aer_severity,
> - (struct aer_capability_regs *)
> - aer_info);
> + devfn, aer_severity, aer_info);
[Severity: High]
This is a pre-existing issue, but does this code leak the newly allocated
aer_info pointer from ghes_estatus_pool if the recovery queue is full or if
the device is not found during workqueue execution?
The memory is allocated from the fixed-size pool earlier in ghes_handle_aer():
drivers/acpi/apei/ghes.c:ghes_handle_aer() {
...
aer_info = (void *)gen_pool_alloc(ghes_estatus_pool,
sizeof(struct aer_capability_regs));
...
}
However, in drivers/pci/pcie/aer.c:aer_recover_queue(), if the internal kfifo
is full, the pointer is dropped without being freed:
drivers/pci/pcie/aer.c:aer_recover_queue() {
...
if (kfifo_in_spinlocked(&aer_recover_ring, &entry, 1,
&aer_recover_ring_lock))
schedule_work(&aer_recover_work);
else
pr_err("buffer overflow in recovery for %04x:%02x:%02x.%x\n",
domain, bus, PCI_SLOT(devfn), PCI_FUNC(devfn));
...
}
Additionally, in drivers/pci/pcie/aer.c:aer_recover_work_func(), if the target
PCI device cannot be found (for instance, due to hot-removal), the function
bypasses the pool region free:
drivers/pci/pcie/aer.c:aer_recover_work_func() {
...
if (!pdev) {
pr_err_ratelimited("%04x:%02x:%02x.%x: no pci_dev found\n",
entry.domain, entry.bus,
PCI_SLOT(entry.devfn),
PCI_FUNC(entry.devfn));
continue;
}
...
}
The continue statement bypasses the subsequent call to
ghes_estatus_pool_region_free(), which leaks the entry.regs pointer. Can this
eventually exhaust the fixed-size ghes_estatus_pool during error storms and
cause subsequent hardware error reports to be silently dropped?
> }
> #endif
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260904172337.1409775-1-dave.jiang@intel.com?part=12
next prev parent reply other threads:[~2026-09-04 18:08 UTC|newest]
Thread overview: 25+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-04 17:23 [PATCH v6 00/13] ACPI: APEI: GHES: Collection of fixes for issues reported by sashiko Dave Jiang
2026-09-04 17:23 ` [PATCH v6 01/13] efi/cper: Reject CPER records with an out-of-range error_data_length Dave Jiang
2026-09-04 17:39 ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 02/13] efi/cper: Reject an error status block length that wraps a u32 Dave Jiang
2026-09-04 17:38 ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 03/13] ACPI: extlog: Validate elog record length before walking sections Dave Jiang
2026-09-04 17:40 ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 04/13] ACPI: extlog: Defer CXL protocol error handling to avoid lock inversion Dave Jiang
2026-09-04 17:23 ` [PATCH v6 05/13] ACPI: extlog: Avoid populating software AER metadata from raw hardware buffer Dave Jiang
2026-09-04 17:58 ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 06/13] ACPI: extlog: Validate PCIe error section length before payload access Dave Jiang
2026-09-04 17:23 ` [PATCH v6 07/13] ACPI: extlog: Fix CONFIG_ACPI_APEI_PCIEAER guard typo Dave Jiang
2026-09-04 17:23 ` [PATCH v6 08/13] ACPI: APEI: GHES: Bound CXL event record copy to the firmware section length Dave Jiang
2026-09-04 17:23 ` [PATCH v6 09/13] ACPI: APEI: GHES: Validate CXL protocol error section length before RAS cap copy Dave Jiang
2026-09-04 18:05 ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 10/13] efi/cper: Read only validated fields in cper_mem_err_pack() Dave Jiang
2026-09-04 18:09 ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 11/13] ACPI: APEI: GHES: Validate memory error section length before payload access Dave Jiang
2026-09-04 18:12 ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata Dave Jiang
2026-09-04 18:08 ` sashiko-bot [this message]
2026-09-04 17:23 ` [PATCH v6 13/13] cxl/ras: Make cxl_cper_handle_prot_err() static Dave Jiang
2026-09-04 18:16 ` sashiko-bot
2026-09-04 18:21 ` [PATCH v6 00/13] ACPI: APEI: GHES: Collection of fixes for issues reported by sashiko Rafael J. Wysocki (Intel)
2026-09-04 19:44 ` Dave Jiang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260904180842.15A9C1F00A3D@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=dave.jiang@intel.com \
--cc=linux-cxl@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox