From: Richard Cheng <icheng@nvidia.com>
To: Guixin Liu <kanie@linux.alibaba.com>
Cc: Davidlohr Bueso <dave@stgolabs.net>,
Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Alison Schofield <alison.schofield@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Dan Williams <djbw@kernel.org>, Ira Weiny <iweiny@kernel.org>,
Li Ming <ming.li@zohomail.com>,
linux-cxl@vger.kernel.org
Subject: Re: [PATCH v2] cxl/cdat: Fix uninitialized stack use in endpoint bandwidth gathering
Date: Wed, 12 Aug 2026 16:17:34 +0800 [thread overview]
Message-ID: <anwrGdUjKDCjr4He@MWDK4CY14F> (raw)
In-Reply-To: <20260812060912.54932-1-kanie@linux.alibaba.com>
On Wed, Aug 12, 2026 at 02:09:12PM +0800, Guixin Liu wrote:
> cxl_endpoint_gather_bandwidth() declares three access_coordinate arrays on
> the stack - pci_coord, sw_coord and ep_coord - without initializing them,
> and relies on its helpers to fill every member. None of them does.
> cxl_pci_get_bandwidth() and cxl_port_get_switch_dport_bandwidth() assign
> only read_bandwidth and write_bandwidth, and
> __cxl_coordinates_combine() assigns an output bandwidth only when both
> input bandwidths are non-zero. Every member a producer declines to set is
> read back as whatever was on the stack.
>
> Both cases occur on real topologies. The latency members of pci_coord and
> sw_coord are never written, yet __cxl_coordinates_combine() sums them
> unconditionally, so the latency reads are undefined on every call. And a
> device whose CDAT DSLBIS reports no bandwidth for an access class leaves
> perf->cdat_coord zero for that class, which is exactly the condition that
> makes __cxl_coordinates_combine() skip the bandwidth assignment and leave
> the ep_coord entry untouched.
>
> The ep_coord case escapes the function: cxl_bandwidth_add() accumulates it
> into the per-upstream-port aggregate that is published through the region's
> access coordinate sysfs attributes, so stack contents are reported to
> userspace as a bandwidth figure. The latency sums are discarded by
> cxl_bandwidth_add() rather than published, but they are still computed from
> uninitialized memory.
>
> Zero initialize the three arrays. Zero is already the value this code uses
> for "not reported" - both the __cxl_coordinates_combine() guard and
> coordinates_valid() test for it - so a member no producer sets now reads
> back as unknown rather than as a plausible number.
>
> Fixes: a5ab0de0ebaa ("cxl: Calculate region bandwidth of targets with shared upstream link")
> Signed-off-by: Guixin Liu <kanie@linux.alibaba.com>
> ---
> This was patch 6/8 of the "cxl: Assorted fixes" series [1]. Per review
> feedback that series is not being reworked as a whole; the fixes are resent
> individually instead. Patches 1, 2 and 7 of the series are dropped, as those
> issues are already fixed in cxl/next.
>
> v1->v2:
> - rebase onto cxl/next
> - rewrite the commit message to describe the behaviour rather than narrate
> the code change (Alison Schofield)
>
> [1] https://lore.kernel.org/linux-cxl/20260811113608.2815625-1-kanie@linux.alibaba.com/
>
> drivers/cxl/core/cdat.c | 6 +++---
> 1 file changed, 3 insertions(+), 3 deletions(-)
>
> diff --git a/drivers/cxl/core/cdat.c b/drivers/cxl/core/cdat.c
> index 5c9f07262513..3c6a1537f89b 100644
> --- a/drivers/cxl/core/cdat.c
> +++ b/drivers/cxl/core/cdat.c
> @@ -633,9 +633,9 @@ static int cxl_endpoint_gather_bandwidth(struct cxl_region *cxlr,
> struct cxl_port *endpoint = to_cxl_port(cxled->cxld.dev.parent);
> struct cxl_port *parent_port = to_cxl_port(endpoint->dev.parent);
> struct cxl_port *gp_port = to_cxl_port(parent_port->dev.parent);
> - struct access_coordinate pci_coord[ACCESS_COORDINATE_MAX];
> - struct access_coordinate sw_coord[ACCESS_COORDINATE_MAX];
> - struct access_coordinate ep_coord[ACCESS_COORDINATE_MAX];
> + struct access_coordinate pci_coord[ACCESS_COORDINATE_MAX] = { };
> + struct access_coordinate sw_coord[ACCESS_COORDINATE_MAX] = { };
> + struct access_coordinate ep_coord[ACCESS_COORDINATE_MAX] = { };
> struct cxl_memdev *cxlmd = cxled_to_memdev(cxled);
> struct cxl_dev_state *cxlds = cxlmd->cxlds;
> struct pci_dev *pdev = to_pci_dev(cxlds->dev);
>
> base-commit: 7098e9cd98a05c0c5de2fae0c2465f9d966fdd07
> --
> 2.43.7
>
>
Code changes looks sane to me, but I'm curious what compiler did you use
and what level of optimization did you open ?
Normally compiler can figure what whether to init them on their own.
But perhaps since the arrays are passed to the helpers in other source
files, in this scenario the compiler can't initialize them reliably, not sure
about this.
Otherwise I have no issue.
Reviewed-by: Richard Cheng <icheng@nvidia.com>
Best regards,
Richard Cheng.
next prev parent reply other threads:[~2026-08-12 8:17 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-12 6:09 [PATCH v2] cxl/cdat: Fix uninitialized stack use in endpoint bandwidth gathering Guixin Liu
2026-08-12 6:22 ` sashiko-bot
2026-08-12 6:44 ` Guixin Liu
2026-08-12 8:17 ` Richard Cheng [this message]
2026-08-12 8:45 ` Guixin Liu
2026-08-12 11:44 ` Li Ming
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=anwrGdUjKDCjr4He@MWDK4CY14F \
--to=icheng@nvidia.com \
--cc=alison.schofield@intel.com \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=djbw@kernel.org \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=kanie@linux.alibaba.com \
--cc=linux-cxl@vger.kernel.org \
--cc=ming.li@zohomail.com \
--cc=vishal.l.verma@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.