Linux CXL
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Dave Jiang" <dave.jiang@intel.com>
Cc: linux-cxl@vger.kernel.org
Subject: Re: [PATCH v6 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata
Date: Fri, 04 Sep 2026 18:08:41 +0000	[thread overview]
Message-ID: <20260904180842.15A9C1F00A3D@smtp.kernel.org> (raw)
In-Reply-To: <20260904172337.1409775-13-dave.jiang@intel.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:

Pre-existing issues:
- [High] Memory leak of `aer_capability_regs` allocated from `ghes_estatus_pool` in error paths of `aer_recover_queue` and `aer_recover_work_func`.
--

commit 639b51ace1216953e300ae9efcce13516dadf849
Author: Dave Jiang <dave.jiang@intel.com>

ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata

This commit validates the error_data_length in ghes_handle_aer() and ensures
that copying firmware error information into the larger aer_capability_regs
struct does not read out of bounds. It also zeroes out the destination and
properly maps the header log and prefix log to prevent firmware from
manipulating software-only fields.

Closes: https://sashiko.dev/#/patchset/20260714231835.303081-1-dave.jiang@intel.com?part=3
Fixes: 7e077e6707b3 ("PCI/ERR: Handle TLP Log in Flit mode")

> diff --git a/drivers/acpi/apei/ghes.c b/drivers/acpi/apei/ghes.c
> index 2077cc90569a0..08c985e729d6c 100644
> --- a/drivers/acpi/apei/ghes.c
> +++ b/drivers/acpi/apei/ghes.c
[ ... ]
> @@ -664,13 +667,25 @@ static void ghes_handle_aer(struct acpi_hest_generic_data *gdata)
>  						  sizeof(struct aer_capability_regs));
>  		if (!aer_info)
>  			return;
> -		memcpy(aer_info, pcie_err->aer_info, sizeof(struct aer_capability_regs));
> +
> +		/*
> +		 * Map aer_info onto the struct as extlog_print_pcie() does:
> +		 * copy up to the four Header Log DWORDs, then place the TLP
> +		 * Prefix Log from where the hardware keeps it. The rest stays
> +		 * zero, so firmware cannot drive the pcie_print_tlp_log() loop
> +		 * over dw[] out of bounds.
> +		 */
> +		memset(aer_info, 0, sizeof(struct aer_capability_regs));
> +		memcpy(aer_info, pcie_err->aer_info,
> +		       offsetof(struct aer_capability_regs, header_log) +
> +		       PCIE_STD_NUM_TLP_HEADERLOG * sizeof(u32));
> +		memcpy(aer_info->header_log.prefix,
> +		       pcie_err->aer_info + PCI_ERR_PREFIX_LOG,
> +		       sizeof(aer_info->header_log.prefix));
>  
>  		aer_recover_queue(pcie_err->device_id.segment,
>  				  pcie_err->device_id.bus,
> -				  devfn, aer_severity,
> -				  (struct aer_capability_regs *)
> -				  aer_info);
> +				  devfn, aer_severity, aer_info);

[Severity: High]
This is a pre-existing issue, but does this code leak the newly allocated
aer_info pointer from ghes_estatus_pool if the recovery queue is full or if
the device is not found during workqueue execution?

The memory is allocated from the fixed-size pool earlier in ghes_handle_aer():

drivers/acpi/apei/ghes.c:ghes_handle_aer() {
    ...
		aer_info = (void *)gen_pool_alloc(ghes_estatus_pool,
						  sizeof(struct aer_capability_regs));
    ...
}

However, in drivers/pci/pcie/aer.c:aer_recover_queue(), if the internal kfifo
is full, the pointer is dropped without being freed:

drivers/pci/pcie/aer.c:aer_recover_queue() {
    ...
	if (kfifo_in_spinlocked(&aer_recover_ring, &entry, 1,
				 &aer_recover_ring_lock))
		schedule_work(&aer_recover_work);
	else
		pr_err("buffer overflow in recovery for %04x:%02x:%02x.%x\n",
		       domain, bus, PCI_SLOT(devfn), PCI_FUNC(devfn));
    ...
}

Additionally, in drivers/pci/pcie/aer.c:aer_recover_work_func(), if the target
PCI device cannot be found (for instance, due to hot-removal), the function
bypasses the pool region free:

drivers/pci/pcie/aer.c:aer_recover_work_func() {
    ...
		if (!pdev) {
			pr_err_ratelimited("%04x:%02x:%02x.%x: no pci_dev found\n",
					   entry.domain, entry.bus,
					   PCI_SLOT(entry.devfn),
					   PCI_FUNC(entry.devfn));
			continue;
		}
    ...
}

The continue statement bypasses the subsequent call to
ghes_estatus_pool_region_free(), which leaks the entry.regs pointer. Can this
eventually exhaust the fixed-size ghes_estatus_pool during error storms and
cause subsequent hardware error reports to be silently dropped?

>  	}
>  #endif
>  }

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260904172337.1409775-1-dave.jiang@intel.com?part=12

  reply	other threads:[~2026-09-04 18:08 UTC|newest]

Thread overview: 25+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-04 17:23 [PATCH v6 00/13] ACPI: APEI: GHES: Collection of fixes for issues reported by sashiko Dave Jiang
2026-09-04 17:23 ` [PATCH v6 01/13] efi/cper: Reject CPER records with an out-of-range error_data_length Dave Jiang
2026-09-04 17:39   ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 02/13] efi/cper: Reject an error status block length that wraps a u32 Dave Jiang
2026-09-04 17:38   ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 03/13] ACPI: extlog: Validate elog record length before walking sections Dave Jiang
2026-09-04 17:40   ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 04/13] ACPI: extlog: Defer CXL protocol error handling to avoid lock inversion Dave Jiang
2026-09-04 17:23 ` [PATCH v6 05/13] ACPI: extlog: Avoid populating software AER metadata from raw hardware buffer Dave Jiang
2026-09-04 17:58   ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 06/13] ACPI: extlog: Validate PCIe error section length before payload access Dave Jiang
2026-09-04 17:23 ` [PATCH v6 07/13] ACPI: extlog: Fix CONFIG_ACPI_APEI_PCIEAER guard typo Dave Jiang
2026-09-04 17:23 ` [PATCH v6 08/13] ACPI: APEI: GHES: Bound CXL event record copy to the firmware section length Dave Jiang
2026-09-04 17:23 ` [PATCH v6 09/13] ACPI: APEI: GHES: Validate CXL protocol error section length before RAS cap copy Dave Jiang
2026-09-04 18:05   ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 10/13] efi/cper: Read only validated fields in cper_mem_err_pack() Dave Jiang
2026-09-04 18:09   ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 11/13] ACPI: APEI: GHES: Validate memory error section length before payload access Dave Jiang
2026-09-04 18:12   ` sashiko-bot
2026-09-04 17:23 ` [PATCH v6 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata Dave Jiang
2026-09-04 18:08   ` sashiko-bot [this message]
2026-09-04 17:23 ` [PATCH v6 13/13] cxl/ras: Make cxl_cper_handle_prot_err() static Dave Jiang
2026-09-04 18:16   ` sashiko-bot
2026-09-04 18:21 ` [PATCH v6 00/13] ACPI: APEI: GHES: Collection of fixes for issues reported by sashiko Rafael J. Wysocki (Intel)
2026-09-04 19:44   ` Dave Jiang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260904180842.15A9C1F00A3D@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=dave.jiang@intel.com \
    --cc=linux-cxl@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox