Linux PCI subsystem development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Terry Bowman" <terry.bowman@amd.com>
Cc: linux-pci@vger.kernel.org, linux-cxl@vger.kernel.org
Subject: Re: [PATCH v19 12/14] cxl: Add port and dport identifiers to CXL AER trace events
Date: Mon, 03 Aug 2026 22:42:23 +0000	[thread overview]
Message-ID: <20260803224224.5FF041F000E9@smtp.kernel.org> (raw)
In-Reply-To: <20260803221810.3685703-13-terry.bowman@amd.com>

Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Missing device lock in cxl_cper_handle_prot_err() leads to Use-After-Free of devres-managed objects like 'dport' during concurrent driver unbind.
--

--- Patch [12]: [PATCH v19 12/14] cxl: Add port and dport identifiers to CXL AER trace events ---
commit 3fa84eaf0343ed1c7fb39fd22762efda3659de2c
Author: Dan Williams <djbw@kernel.org>

cxl: Add port and dport identifiers to CXL AER trace events

This commit modifies CXL AER trace events to use struct cxl_port and struct
cxl_dport pointers instead of deriving device strings at the caller. It
unifies the correctable and uncorrectable error trace events and updates
cxl_cper_handle_prot_err() to use these new helpers while dropping a device
lock it considered redundant.

> Changes in v18->v19:
> - Drop redundant device lock in cxl_cper_handle_prot_err(); the port
>   reference already keeps the object alive and no RAS iomap is accessed.

[Severity: High]
Does dropping the device lock here leave devres-managed objects unprotected?

The commit message notes that the port reference keeps the object alive, but a
device reference only protects the struct device allocation itself, not the
driver binding state or devm actions.

[ ... ]

> diff --git a/drivers/cxl/core/ras.c b/drivers/cxl/core/ras.c
> index 5183b3c532952..0254b7ec64c30 100644
> --- a/drivers/cxl/core/ras.c
> +++ b/drivers/cxl/core/ras.c

[ ... ]

> @@ -108,44 +76,32 @@ static struct cxl_port *find_cxl_port_by_dev(struct device *dev, struct cxl_dpor
>  
>  void cxl_cper_handle_prot_err(struct cxl_cper_prot_err_work_data *data)
>  {

[ ... ]

> -	guard(device)(&pdev->dev);
> -	if (!pdev->dev.driver)
> +	struct cxl_port *port __free(put_cxl_port) = find_cxl_port_by_dev(&pdev->dev, NULL);
> +	if (!port) {
> +		dev_err_ratelimited(&pdev->dev,
> +				    "Failed to find parent port device in CXL topology\n");
>  		return;
> +	}
>  
> -	struct device *mem_dev __free(put_device) = bus_find_device(
> -		&cxl_bus_type, NULL, pdev, match_memdev_by_parent);
> -	if (!mem_dev)
> -		return;
> +	/* dport is NULL for Endpoint and Upstream Port devices */
> +	dport = cxl_find_dport_by_dev(port, &pdev->dev);
>  
> -	cxlmd = to_cxl_memdev(mem_dev);
>  	if (data->severity == AER_CORRECTABLE)
> -		cxl_cper_trace_corr_prot_err(cxlmd, data->ras_cap);
> +		cxl_cper_trace_corr_prot_err(port, dport, pdev->dsn,
> +					     &data->ras_cap);
>  	else
> -		cxl_cper_trace_uncorr_prot_err(cxlmd, data->ras_cap);
> +		cxl_cper_trace_uncorr_prot_err(port, dport, pdev->dsn,
> +					       &data->ras_cap);
>  }

[Severity: High]
Can accessing dport here race with a driver unbind?

If a concurrent CXL port driver unbind occurs (via sysfs or device hot-remove),
the devres release actions will execute while cxl_cper_handle_prot_err() is
still running.

This could result in the devm-managed dport being freed, leaving the dport
pointer referencing freed memory. This pointer is then passed into the trace
events where cxl_trace_dport_name() attempts to read dport->dport_dev:

drivers/cxl/core/trace.c:cxl_trace_dport_name() {
    ...
    if (dport)
        return dev_name(dport->dport_dev);
    ...
}

Could this result in a use-after-free?

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/20260803221810.3685703-1-terry.bowman@amd.com?part=12

  reply	other threads:[~2026-08-03 22:42 UTC|newest]

Thread overview: 40+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-03 22:17 [PATCH v19 00/14] Enable CXL PCIe Port Protocol Error handling and logging Terry Bowman
2026-08-03 22:17 ` [PATCH v19 01/14] cxl/ras: Fix cxl_rch_get_aer_info() out-of-bounds AER register read Terry Bowman
2026-08-03 22:42   ` sashiko-bot
2026-08-04  2:10   ` Alison Schofield
2026-08-03 22:17 ` [PATCH v19 02/14] cxl/ras: Fix cxl_rch_get_aer_severity() wrong severity register Terry Bowman
2026-08-03 22:35   ` sashiko-bot
2026-08-04  2:11   ` Alison Schofield
2026-08-03 22:17 ` [PATCH v19 03/14] acpi/apei/ghes: Use raw_spinlock_t for CXL CPER work locks Terry Bowman
2026-08-03 22:39   ` sashiko-bot
2026-08-03 22:18 ` [PATCH v19 04/14] cxl: Tighten CPER kfifo registration API and symbol visibility Terry Bowman
2026-08-03 22:30   ` sashiko-bot
2026-08-04  2:13   ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 05/14] cxl: Rename find_cxl_port() to find_cxl_port_by_dport() Terry Bowman
2026-08-03 22:29   ` sashiko-bot
2026-08-04  2:14   ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 06/14] PCI/AER: Introduce AER-CXL protocol error kfifo Terry Bowman
2026-08-03 22:28   ` sashiko-bot
2026-08-03 22:18 ` [PATCH v19 07/14] PCI: Establish common CXL Port protocol error flow Terry Bowman
2026-08-03 22:56   ` sashiko-bot
2026-08-03 22:18 ` [PATCH v19 08/14] cxl/ras: Handle RCH correctable and uncorrectable errors in one pass Terry Bowman
2026-08-03 22:29   ` sashiko-bot
2026-08-04  2:16   ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 09/14] cxl/pci: Thread port and dport through RAS handling helpers Terry Bowman
2026-08-03 22:33   ` sashiko-bot
2026-08-04  2:16   ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 10/14] cxl: Update CXL Endpoint AER handler Terry Bowman
2026-08-03 22:40   ` sashiko-bot
2026-08-04  2:17   ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 11/14] PCI: Cache PCI DSN into pci_dev->dsn during probe Terry Bowman
2026-08-03 22:29   ` sashiko-bot
2026-08-04  2:26   ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 12/14] cxl: Add port and dport identifiers to CXL AER trace events Terry Bowman
2026-08-03 22:42   ` sashiko-bot [this message]
2026-08-04  2:27   ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 13/14] PCI/CXL: Mask/Unmask CXL protocol errors Terry Bowman
2026-08-03 22:55   ` sashiko-bot
2026-08-04  2:29   ` Alison Schofield
2026-08-03 22:18 ` [PATCH v19 14/14] Documentation: cxl: Document CXL protocol error handling Terry Bowman
2026-08-03 22:31   ` sashiko-bot
2026-08-04  2:30   ` Alison Schofield

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260803224224.5FF041F000E9@smtp.kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=linux-pci@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    --cc=terry.bowman@amd.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox