From: sashiko-bot@kernel.org
To: "David Zhang" <yidong.zhang@amd.com>
Cc: dri-devel@lists.freedesktop.org
Subject: Re: [PATCH V1 13/20] accel/amdxdna: Prepare for AIE4 command submission
Date: Wed, 30 Sep 2026 04:00:24 +0000 [thread overview]
Message-ID: <20260930040024.844E81F000FF@smtp.kernel.org> (raw)
In-Reply-To: <20260930033233.1727265-14-yidong.zhang@amd.com>
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Race in get_read_index() allows last_read_index to regress
--
commit b18a81613bb4ba1a66541441dd19a7be5abf5cef
Author: David Zhang <yidong.zhang@amd.com>
accel/amdxdna: Prepare for AIE4 command submission
This commit prepares the data structures and completion wait helpers required
for AIE4 command submission. It defines struct amdxdna_cmd_start_dpu, extends
union amdxdna_job_priv, and updates completion wait helpers to handle read
tearing and asynchronous device disconnects.
> diff --git a/drivers/accel/amdxdna/aie4_ctx.c b/drivers/accel/amdxdna/aie4_ctx.c
> index 3927c9fef05f9..59d37bd5c5a08 100644
> --- a/drivers/accel/amdxdna/aie4_ctx.c
> +++ b/drivers/accel/amdxdna/aie4_ctx.c
[ ... ]
> @@ -423,34 +423,92 @@ static inline bool valid_queue_index(u64 read, u64 write, u32 capacity)
>
> static u64 get_read_index(struct amdxdna_hwctx *hwctx)
> {
[ ... ]
> if (!valid_queue_index(ri, wi, CTX_MAX_CMDS)) {
> - XDNA_WARN(xdna, "Invalid index, ri %llu, wi %llu", ri, wi);
> - usleep_range(100, 200);
> - ri = READ_ONCE(*hwctx->priv->umq_read_index);
> + ri = READ_ONCE(*priv->umq_read_index);
> + /* Order the read_index sample before the write_index sample. */
> + smp_rmb();
> + wi = READ_ONCE(priv->write_index);
> if (!valid_queue_index(ri, wi, CTX_MAX_CMDS)) {
> - XDNA_ERR(xdna, "Invalid index after retry, ri %llu, wi %llu", ri, wi);
> - ri = 0;
> + /*
> + * Still invalid (torn 64-bit read, or a transient
> + * accounting skew). Return the last valid read_index
> + * instead of 0: read_index only advances, so the cached
> + * value is a safe lower bound -- it never reports a
> + * command complete that isn't, and never regresses the
> + * worker into falsely timing out a finished job.
> + */
> + XDNA_DBG(xdna, "Invalid index, ri %llu, wi %llu", ri, wi);
> + return READ_ONCE(priv->last_read_index);
> }
> }
>
> + WRITE_ONCE(priv->last_read_index, ri);
[Severity: High]
Can this lockless blind write allow an older read index to overwrite a newer
one when executed concurrently by multiple waiters?
If Thread A reads `ri = 10` from hardware and Thread B reads `ri = 12`, Thread B
could execute `WRITE_ONCE(priv->last_read_index, 12)` before Thread A executes
`WRITE_ONCE(priv->last_read_index, 10)`, causing last_read_index to regress.
This breaks the monotonicity required to safely prevent false timeouts.
> return ri;
> }
[ ... ]
> @@ -466,11 +524,14 @@ int aie4_cmd_wait(struct amdxdna_hwctx *hwctx, u64 seq, u32 timeout)
> wait_jifs = msecs_to_jiffies(timeout);
>
> ret = wait_event_interruptible_timeout(cert_comp->waitq,
> - (check_cmd_done(hwctx, seq)),
> + check_cmd_done(hwctx, seq, cert_comp),
> wait_jifs);
>
> if (!ret)
> ret = -ETIME;
> + else if (ret > 0 && get_read_index(hwctx) <= seq)
> + /* Woke on disconnect/reset, not on real completion. */
> + ret = -EAGAIN;
[Severity: High]
Could this disconnect-detection logic incorrectly trigger when get_read_index()
returns the corrupted, backward-jumped last_read_index during a torn read?
If last_read_index regresses due to the race condition above, a subsequent call
to get_read_index() experiencing a torn read would fall back to this stale
value.
This check would then falsely evaluate to true, causing successfully completed
jobs to falsely return -EAGAIN to userspace.
>
> aie4_put_cert_comp(cert_comp);
>
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260930033233.1727265-1-yidong.zhang@amd.com?part=13
next prev parent reply other threads:[~2026-09-30 4:00 UTC|newest]
Thread overview: 33+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-30 3:32 [PATCH V1 00/20] accel/amdxdna: Kernel submission and PM for AIE4 David Zhang
2026-09-30 3:32 ` [PATCH V1 01/20] accel/amdxdna: Rename NPU3 firmware files David Zhang
2026-09-30 3:32 ` [PATCH V1 02/20] accel/amdxdna: Remove mmap for doorbell David Zhang
2026-09-30 3:32 ` [PATCH V1 03/20] accel/amdxdna: Add CERT firmware version support David Zhang
2026-09-30 3:53 ` sashiko-bot
2026-09-30 3:32 ` [PATCH V1 04/20] accel/amdxdna: Upgrade firmware version to 6.0 David Zhang
2026-09-30 3:32 ` [PATCH V1 05/20] accel/amdxdna: Add NPU3 classic device support David Zhang
2026-09-30 3:32 ` [PATCH V1 06/20] accel/amdxdna: Add AIE version query to aie4_get_info David Zhang
2026-09-30 3:32 ` [PATCH V1 07/20] accel/amdxdna: Add get and set power_mode for AIE4 David Zhang
2026-09-30 3:56 ` sashiko-bot
2026-09-30 3:32 ` [PATCH V1 08/20] accel/amdxdna: Add clock, DPM frequency, and resource info queries " David Zhang
2026-09-30 3:32 ` [PATCH V1 09/20] accel/amdxdna: Add context switch hysteresis with debugfs control David Zhang
2026-09-30 3:52 ` sashiko-bot
2026-09-30 3:32 ` [PATCH V1 10/20] accel/amdxdna: Refactor AIE4 hardware initialization sequence David Zhang
2026-09-30 3:32 ` [PATCH V1 11/20] accel/amdxdna: Decouple AIE4 doorbell and MSI-X notification transport hooks David Zhang
2026-09-30 3:32 ` [PATCH V1 12/20] accel/amdxdna: Implement AIE4 kernel queue lifecycle and memory layout David Zhang
2026-09-30 4:00 ` sashiko-bot
2026-09-30 3:32 ` [PATCH V1 13/20] accel/amdxdna: Prepare for AIE4 command submission David Zhang
2026-09-30 4:00 ` sashiko-bot [this message]
2026-09-30 3:32 ` [PATCH V1 14/20] accel/amdxdna: Implement AIE4 command packet building and submission David Zhang
2026-09-30 4:04 ` sashiko-bot
2026-10-05 21:52 ` Zhang, Yidong (David)
2026-09-30 3:32 ` [PATCH V1 15/20] accel/amdxdna: Finalize runtime PM before acquiring dev_lock on removal David Zhang
2026-09-30 3:32 ` [PATCH V1 16/20] accel/amdxdna: Implement AIE4 suspend and resume David Zhang
2026-09-30 4:07 ` sashiko-bot
2026-10-05 21:49 ` Zhang, Yidong (David)
2026-09-30 3:32 ` [PATCH V1 17/20] accel/amdxdna: Link SR-IOV VFs for power management sequencing David Zhang
2026-09-30 3:59 ` sashiko-bot
2026-09-30 3:32 ` [PATCH V1 18/20] accel/amdxdna: Implement runtime suspend and resume support David Zhang
2026-09-30 4:00 ` sashiko-bot
2026-09-30 3:32 ` [PATCH V1 19/20] accel/amdxdna: Add stub hwctx_config for AIE4 David Zhang
2026-09-30 3:59 ` sashiko-bot
2026-09-30 3:32 ` [PATCH V1 20/20] accel/amdxdna: Enable AIE4 firmware logging to DRAM David Zhang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260930040024.844E81F000FF@smtp.kernel.org \
--to=sashiko-bot@kernel.org \
--cc=dri-devel@lists.freedesktop.org \
--cc=sashiko-reviews@lists.linux.dev \
--cc=yidong.zhang@amd.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox