Linux CXL
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Dave Jiang" <dave.jiang@intel.com>
Cc: linux-cxl@vger.kernel.org
Subject: Re: [PATCH v4 09/13] ACPI: APEI: GHES: Validate CXL protocol error section length before RAS cap copy
Date: Mon, 24 Aug 2026 18:10:20 +0000	[thread overview]
Message-ID: <20260824181020.87F621F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260824174936.939059-10-dave.jiang@intel.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:

Pre-existing issues:
- [High] `cxl_cper_print_prot_err()` reads out-of-bounds memory when dumping the CXL DVSEC and Error Log if `prot_err->dvsec_len` exceeds the actual error record data length.
--

--- Patch [9]: [PATCH v4 09/13] ACPI: APEI: GHES: Validate CXL protocol error section length before RAS cap copy ---
commit 86649d90c05d1f4d84b3cdb3e62e7482bd0800ac
Author: Dave Jiang <dave.jiang@intel.com>

ACPI: APEI: GHES: Validate CXL protocol error section length before RAS cap copy

cxl_cper_setup_prot_err_work_data() locates the RAS Capability block at
prot_err + sizeof(*prot_err) + dvsec_len and copies it, but dvsec_len is
firmware controlled and never validated, so it can point the copy outside
the section.

Extend cxl_cper_sec_prot_err_valid() to check that the section can hold
the header, and that the header, DVSEC and RAS Capability block together
fit the reported section length.

Link: https://sashiko.dev/#/patchset/20260617-topics-ahmtib01-ras_ffh_arm_internal_review-v6-0-91f725174aa0@arm.com?part=6
Link: https://lore.kernel.org/linux-cxl/20260709165457.8BA181F000E9@smtp.kernel.org/

> diff --git a/drivers/acpi/apei/ghes_helpers.c b/drivers/acpi/apei/ghes_helpers.c
> index bc7111b740afb..df41b993f4130 100644
> --- a/drivers/acpi/apei/ghes_helpers.c
> +++ b/drivers/acpi/apei/ghes_helpers.c

[ ... ]

> @@ -23,6 +30,15 @@ int cxl_cper_sec_prot_err_valid(struct cxl_cper_sec_prot_err *prot_err)
>  		return -EINVAL;
>  	}
>  
> +	/* The RAS Capability block sits after a firmware-sized DVSEC. */
> +	if (sizeof(*prot_err) + prot_err->dvsec_len +
> +	    sizeof(struct cxl_ras_capability_regs) > len) {
> +		pr_err_ratelimited(FW_WARN
> +				   "CXL CPER prot err DVSEC (%u) overruns section (%u)\n",
> +				   prot_err->dvsec_len, len);
> +		return -EINVAL;
> +	}
> +

[Severity: High]
This is a pre-existing issue, but does this new bounds validation cover the
synchronous print path during CXL protocol error handling?

While this patch adds bounds validation via cxl_cper_sec_prot_err_valid()
for the workqueue execution path, the synchronous print path appears to
completely bypass this new validation.

Earlier in the flow, ghes_proc() calls ghes_print_estatus(), which leads to
cper_estatus_print_section():

drivers/firmware/efi/cper.c:cper_estatus_print_section() {
    ...
		if (gdata->error_data_length >= sizeof(*prot_err))
			cxl_cper_print_prot_err(newpfx, prot_err);
		else
			goto err_section_too_small;
    ...
}

This checks that the error section is large enough to hold the base header
structure, but fails to validate that it can also hold the variable-length
DVSEC or the RAS capabilities before passing it to the print routine.

The untrusted dvsec_len is then used directly in cxl_cper_print_prot_err(),
leading to an out-of-bounds memory read:

drivers/firmware/efi/cper_cxl.c:cxl_cper_print_prot_err() {
    ...
		print_hex_dump(pfx, "", DUMP_PREFIX_OFFSET, 16, 4, (prot_err + 1),
			       prot_err->dvsec_len, 0);
    ...
		cxl_ras = (struct cxl_ras_capability_regs *)((long)prot_err + size);
		pr_info("%s cxl_ras_uncor_status: 0x%08x", pfx,
			cxl_ras->uncor_status);
    ...
}

Could firmware still trigger an unhandled page fault or leak kernel memory
to the system log by providing an artificially large dvsec_len that reaches
this print path?

>  	if ((prot_err->agent_type == RCD || prot_err->agent_type == DEVICE ||
>  	     prot_err->agent_type == LD || prot_err->agent_type == FMLD) &&
>  	    !(prot_err->valid_bits & PROT_ERR_VALID_SERIAL_NUMBER))

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260824174936.939059-1-dave.jiang@intel.com?part=9

  reply	other threads:[~2026-08-24 18:10 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-24 17:49 [PATCH v4 00/13] ACPI: APEI: GHES: Collection of fixes for issues reported by sashiko Dave Jiang
2026-08-24 17:49 ` [PATCH v4 01/13] efi/cper: Reject CPER records with an out-of-range error_data_length Dave Jiang
2026-08-24 18:10   ` sashiko-bot
2026-08-24 21:57   ` Jonathan Cameron
2026-08-25 16:31     ` Dave Jiang
2026-08-24 17:49 ` [PATCH v4 02/13] efi/cper: Reject an error status block length that wraps a u32 Dave Jiang
2026-08-24 18:03   ` sashiko-bot
2026-08-24 22:22   ` Jonathan Cameron
2026-08-24 17:49 ` [PATCH v4 03/13] ACPI: extlog: Validate elog record length before walking sections Dave Jiang
2026-08-24 18:07   ` sashiko-bot
2026-08-24 22:28   ` Jonathan Cameron
2026-08-25 17:15     ` Dave Jiang
2026-08-24 17:49 ` [PATCH v4 04/13] ACPI: extlog: Defer CXL protocol error handling to avoid lock inversion Dave Jiang
2026-08-24 18:03   ` sashiko-bot
2026-08-24 22:29   ` Jonathan Cameron
2026-08-24 17:49 ` [PATCH v4 05/13] ACPI: extlog: Avoid populating software AER metadata from raw hardware buffer Dave Jiang
2026-08-24 18:13   ` sashiko-bot
2026-08-24 23:05   ` Jonathan Cameron
2026-08-25 17:38     ` Dave Jiang
2026-08-24 17:49 ` [PATCH v4 06/13] ACPI: extlog: Validate PCIe error section length before payload access Dave Jiang
2026-08-24 18:07   ` sashiko-bot
2026-08-24 23:08   ` Jonathan Cameron
2026-08-24 17:49 ` [PATCH v4 07/13] ACPI: extlog: Fix CONFIG_ACPI_APEI_PCIEAER guard typo Dave Jiang
2026-08-24 18:15   ` sashiko-bot
2026-08-24 23:11   ` Jonathan Cameron
2026-08-24 17:49 ` [PATCH v4 08/13] ACPI: APEI: GHES: Bound CXL event record copy to the firmware section length Dave Jiang
2026-08-24 23:13   ` Jonathan Cameron
2026-08-24 17:49 ` [PATCH v4 09/13] ACPI: APEI: GHES: Validate CXL protocol error section length before RAS cap copy Dave Jiang
2026-08-24 18:10   ` sashiko-bot [this message]
2026-08-24 23:14   ` Jonathan Cameron
2026-08-24 17:49 ` [PATCH v4 10/13] efi/cper: Read only validated fields in cper_mem_err_pack() Dave Jiang
2026-08-24 18:08   ` sashiko-bot
2026-08-24 17:49 ` [PATCH v4 11/13] ACPI: APEI: GHES: Validate memory error section length before payload access Dave Jiang
2026-08-24 18:17   ` sashiko-bot
2026-08-24 17:49 ` [PATCH v4 12/13] ACPI: APEI: GHES: Bound AER info copy and sanitize software metadata Dave Jiang
2026-08-24 18:17   ` sashiko-bot
2026-08-24 21:32     ` Dave Jiang
2026-08-24 17:49 ` [PATCH v4 13/13] cxl/ras: Make cxl_cper_handle_prot_err() static Dave Jiang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260824181020.87F621F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=dave.jiang@intel.com \
    --cc=linux-cxl@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox