From: sashiko-bot@kernel.org
To: "Dave Jiang" <dave.jiang@intel.com>
Cc: linux-cxl@vger.kernel.org
Subject: Re: [PATCH v5 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata
Date: Thu, 27 Aug 2026 21:01:03 +0000 [thread overview]
Message-ID: <20260827210103.A352D1F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260827203726.3027541-13-dave.jiang@intel.com>
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
Pre-existing issues:
- [High] Pre-existing memory leak of dynamically allocated `aer_info` in `aer_recover_queue()` when the recovery kfifo is full.
- [High] Pre-existing memory leak of pool region in `aer_recover_work_func()` when the PCI device cannot be found.
--
commit 3fa389697c596cc80036725c9cd2026267df97d9
Author: Dave Jiang <dave.jiang@intel.com>
ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata
ghes_handle_aer() copies sizeof(struct aer_capability_regs) out of the
fixed 96-byte pcie_err->aer_info. The struct is larger, so the copy reads
past the section, and it fills the software-only header_len and flit fields
of the embedded struct pcie_tlp_log from firmware bytes. This commit
validates error_data_length, zeroes the destination, and copies only what
maps onto the struct to prevent out-of-bounds loop execution.
Closes: https://sashiko.dev/#/patchset/20260714231835.303081-1-dave.jiang@intel.com?part=3
> diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c
> index 0cc1e63836356..3df62ec20cf96 100644
> --- a/drivers/acpi/apei/ghes.c
> +++ b/drivers/acpi/apei/ghes.c
> @@ -664,13 +667,21 @@ static void ghes_handle_aer(struct acpi_hest_generic_data *gdata)
> sizeof(struct aer_capability_regs));
> if (!aer_info)
> return;
> - memcpy(aer_info, pcie_err->aer_info, sizeof(struct aer_capability_regs));
> +
> + /*
> + * Copy only what maps onto the struct, as extlog_print_pcie()
> + * does: the leading registers and the four Header Log DWORDs.
> + * The rest stays zero, so firmware cannot drive the
> + * pcie_print_tlp_log() loop over dw[] out of bounds.
> + */
> + memset(aer_info, 0, sizeof(struct aer_capability_regs));
> + memcpy(aer_info, pcie_err->aer_info,
> + offsetof(struct aer_capability_regs, header_log) +
> + PCIE_STD_NUM_TLP_HEADERLOG * sizeof(u32));
>
> aer_recover_queue(pcie_err->device_id.segment,
> pcie_err->device_id.bus,
> - devfn, aer_severity,
> - (struct aer_capability_regs *)
> - aer_info);
> + devfn, aer_severity, aer_info);
[Severity: High]
This is a pre-existing issue, but does aer_recover_queue() leak the
dynamically allocated aer_info if the recovery kfifo is full?
In ghes_handle_aer(), aer_info is allocated from the ghes_estatus_pool and
passed to aer_recover_queue(). If the aer_recover_ring is full,
kfifo_in_spinlocked() returns 0 and aer_recover_queue() logs an error:
drivers/pci/pcie/aer.c:aer_recover_queue() {
...
if (kfifo_in_spinlocked(&aer_recover_ring, &entry, 1,
&aer_recover_ring_lock))
schedule_work(&aer_recover_work);
else
pr_err("buffer overflow in recovery for %04x:%02x:%02x.%x\n",
domain, bus, PCI_SLOT(devfn), PCI_FUNC(devfn));
}
The function returns without freeing aer_regs, permanently leaking memory
from the lockless GHES estatus pool.
[Severity: High]
This is also a pre-existing issue, but does aer_recover_work_func() leak
this pool region when a PCI device cannot be found?
When aer_recover_work_func() dequeues entries, it attempts to resolve the
pci_dev. If it fails, it executes continue, skipping the rest of the loop:
drivers/pci/pcie/aer.c:aer_recover_work_func() {
...
pdev = pci_get_domain_bus_and_slot(entry.domain, entry.bus,
entry.devfn);
if (!pdev) {
pr_err_ratelimited("%04x:%02x:%02x.%x: no pci_dev found\n",
entry.domain, entry.bus,
PCI_SLOT(entry.devfn),
PCI_FUNC(entry.devfn));
continue;
}
...
}
The call to ghes_estatus_pool_region_free() is at the end of the loop body,
so the continue statement bypasses it, permanently leaking the memory from
the GHES estatus pool.
> }
> #endif
> }
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260827203726.3027541-1-dave.jiang@intel.com?part=12
next prev parent reply other threads:[~2026-08-27 21:01 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 20:37 [PATCH v5 00/13] ACPI: APEI: GHES: Collection of fixes for issues reported by sashiko Dave Jiang
2026-08-27 20:37 ` [PATCH v5 01/13] efi/cper: Reject CPER records with an out-of-range error_data_length Dave Jiang
2026-08-27 20:52 ` sashiko-bot
2026-08-27 20:37 ` [PATCH v5 02/13] efi/cper: Reject an error status block length that wraps a u32 Dave Jiang
2026-08-27 20:52 ` sashiko-bot
2026-08-27 20:37 ` [PATCH v5 03/13] ACPI: extlog: Validate elog record length before walking sections Dave Jiang
2026-08-27 20:55 ` sashiko-bot
2026-08-27 20:37 ` [PATCH v5 04/13] ACPI: extlog: Defer CXL protocol error handling to avoid lock inversion Dave Jiang
2026-08-27 20:55 ` sashiko-bot
2026-08-27 20:37 ` [PATCH v5 05/13] ACPI: extlog: Avoid populating software AER metadata from raw hardware buffer Dave Jiang
2026-08-27 20:58 ` sashiko-bot
2026-08-27 23:06 ` Dave Jiang
2026-08-27 20:37 ` [PATCH v5 06/13] ACPI: extlog: Validate PCIe error section length before payload access Dave Jiang
2026-08-27 20:37 ` [PATCH v5 07/13] ACPI: extlog: Fix CONFIG_ACPI_APEI_PCIEAER guard typo Dave Jiang
2026-08-27 20:37 ` [PATCH v5 08/13] ACPI: APEI: GHES: Bound CXL event record copy to the firmware section length Dave Jiang
2026-08-27 20:55 ` sashiko-bot
2026-08-27 20:37 ` [PATCH v5 09/13] ACPI: APEI: GHES: Validate CXL protocol error section length before RAS cap copy Dave Jiang
2026-08-27 20:53 ` sashiko-bot
2026-08-27 20:37 ` [PATCH v5 10/13] efi/cper: Read only validated fields in cper_mem_err_pack() Dave Jiang
2026-08-27 20:55 ` sashiko-bot
2026-08-27 20:37 ` [PATCH v5 11/13] ACPI: APEI: GHES: Validate memory error section length before payload access Dave Jiang
2026-08-27 21:05 ` sashiko-bot
2026-08-27 20:37 ` [PATCH v5 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata Dave Jiang
2026-08-27 21:01 ` sashiko-bot [this message]
2026-08-27 20:37 ` [PATCH v5 13/13] cxl/ras: Make cxl_cper_handle_prot_err() static Dave Jiang
2026-08-27 20:59 ` sashiko-bot
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260827210103.A352D1F000E9@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=dave.jiang@intel.com \
--cc=linux-cxl@vger.kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.