From: Jiaxing Hu <gahing@gahingwoo.com>
To: royalnet026@gmail.com
Cc: tomeu@tomeuvizoso.net, linux-rockchip@lists.infradead.org,
dri-devel@lists.freedesktop.org
Subject: Re: [PATCH v13 03/14] accel/rocket: wait for a running IRQ handler before resetting a core
Date: Sat, 19 Sep 2026 21:17:25 +1200 [thread overview]
Message-ID: <20260919091725.8981-1-gahing@gahingwoo.com> (raw)
In-Reply-To: <20260916132824.13527-1-royalnet026@gmail.com>
Hi Igor,
Thank you for re-running it on the changed form. The tag carries to v14
with your comment unchanged.
The limit you put on it is sharper than the one the cover put on it, and
v14 will say yours instead of mine. A single 1x1 convolution is a
one-task job, so hw_submit() never runs from the IRQ thread and
drm_sched_stop() always fences it. That is the path the race needs. So
what your runs establish is that the lock scope adds no lockdep report,
no MMU fault and no hang on the path they do reach, and not that it
closes anything. The cover said the fix was an argument from the code;
your sentence says which part of the code the test never visited, which
is the useful half.
So that I am quoting you correctly: I read the tally as 8 of the 126
scored inferences missing on all 48 channels across the three runs, and
5 of those, in the two runs you traced, sitting on the 5 -ECANCELED
completions. Correct me if the 8 and the 5 are not nested that way.
If you ever want to reach the other path, it needs a job with more than
one task. On the Mesa side that is a graph whose weights do not fit the
CBUF: rkt_ml_subgraph_invoke() then submits one job per task rather than
one per operation, and MobileNet comes out as 34 tasks here. A 1x1
convolution will be one task whatever else changes.
And the mirror of that, from this end, since it is the reason your runs
are the only ones there are. JOB_TIMEOUT_MS=2 does not survive on this
RK3576. Running rocket_reset() at that rate takes the board's PMIC down
through its I2C:
rk3x-i2c 2ac40000.i2c: irq in STATE_IDLE
a big-core voltage transition then fails with -ETIMEDOUT and two CPUs
stop answering an NMI. I bisected it across four boots against a clean
next-20260914 and against the rail change on its own: it is the timeout
constant, not this series and not the fourteen patches. Whether that is
an RK3576 property or this board's PMIC I cannot say from one board. It
is why the RK3576 side of 2/14, 3/14 and 4/14 has no induced-reset
evidence at all.
4/14 may well change shape in v14, since the asynchronous put has a
[High] on it again and the fix would move the put out from under
job_lock. I will say so in the cover if it does, and take you up on the
re-run.
Regards,
Jiaxing
_______________________________________________
Linux-rockchip mailing list
Linux-rockchip@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-rockchip
WARNING: multiple messages have this Message-ID (diff)
From: Jiaxing Hu <gahing@gahingwoo.com>
To: royalnet026@gmail.com
Cc: tomeu@tomeuvizoso.net, linux-rockchip@lists.infradead.org,
dri-devel@lists.freedesktop.org
Subject: Re: [PATCH v13 03/14] accel/rocket: wait for a running IRQ handler before resetting a core
Date: Sat, 19 Sep 2026 21:17:25 +1200 [thread overview]
Message-ID: <20260919091725.8981-1-gahing@gahingwoo.com> (raw)
In-Reply-To: <20260916132824.13527-1-royalnet026@gmail.com>
Hi Igor,
Thank you for re-running it on the changed form. The tag carries to v14
with your comment unchanged.
The limit you put on it is sharper than the one the cover put on it, and
v14 will say yours instead of mine. A single 1x1 convolution is a
one-task job, so hw_submit() never runs from the IRQ thread and
drm_sched_stop() always fences it. That is the path the race needs. So
what your runs establish is that the lock scope adds no lockdep report,
no MMU fault and no hang on the path they do reach, and not that it
closes anything. The cover said the fix was an argument from the code;
your sentence says which part of the code the test never visited, which
is the useful half.
So that I am quoting you correctly: I read the tally as 8 of the 126
scored inferences missing on all 48 channels across the three runs, and
5 of those, in the two runs you traced, sitting on the 5 -ECANCELED
completions. Correct me if the 8 and the 5 are not nested that way.
If you ever want to reach the other path, it needs a job with more than
one task. On the Mesa side that is a graph whose weights do not fit the
CBUF: rkt_ml_subgraph_invoke() then submits one job per task rather than
one per operation, and MobileNet comes out as 34 tasks here. A 1x1
convolution will be one task whatever else changes.
And the mirror of that, from this end, since it is the reason your runs
are the only ones there are. JOB_TIMEOUT_MS=2 does not survive on this
RK3576. Running rocket_reset() at that rate takes the board's PMIC down
through its I2C:
rk3x-i2c 2ac40000.i2c: irq in STATE_IDLE
a big-core voltage transition then fails with -ETIMEDOUT and two CPUs
stop answering an NMI. I bisected it across four boots against a clean
next-20260914 and against the rail change on its own: it is the timeout
constant, not this series and not the fourteen patches. Whether that is
an RK3576 property or this board's PMIC I cannot say from one board. It
is why the RK3576 side of 2/14, 3/14 and 4/14 has no induced-reset
evidence at all.
4/14 may well change shape in v14, since the asynchronous put has a
[High] on it again and the fix would move the put out from under
job_lock. I will say so in the cover if it does, and take you up on the
re-run.
Regards,
Jiaxing
next prev parent reply other threads:[~2026-09-19 9:17 UTC|newest]
Thread overview: 66+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-15 10:43 [PATCH v13 00/14] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 01/14] accel/rocket: request the core clocks by name Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 02/14] accel/rocket: take the completion register writes under job_lock Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 03/14] accel/rocket: wait for a running IRQ handler before resetting a core Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:59 ` sashiko-bot
2026-09-16 13:28 ` Igor Paunovic
2026-09-16 13:28 ` Igor Paunovic
2026-09-19 9:17 ` Jiaxing Hu [this message]
2026-09-19 9:17 ` Jiaxing Hu
2026-09-19 10:34 ` Igor Paunovic
2026-09-19 10:34 ` Igor Paunovic
2026-09-15 10:43 ` [PATCH v13 04/14] accel/rocket: let the core suspend after a reset Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:58 ` sashiko-bot
2026-09-15 10:43 ` [PATCH v13 05/14] accel/rocket: factor the completion tail out of the IRQ handler Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 06/14] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 06/14] dt-bindings: npu: rockchip: add rockchip, rk3576-rknn-core Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 06/14] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 07/14] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-21 21:52 ` Heiko Stuebner
2026-09-21 21:52 ` Heiko Stuebner
2026-09-15 10:43 ` [PATCH v13 08/14] dt-bindings: iommu: rockchip: describe the RK3576 NPU MMU Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 09/14] pmdomain: rockchip: add optional per-domain power-on settle delay Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-21 12:41 ` Ulf Hansson
2026-09-21 12:41 ` Ulf Hansson
2026-09-21 22:06 ` Heiko Stuebner
2026-09-21 22:06 ` Heiko Stuebner
2026-09-22 1:28 ` Chaoyi Chen
2026-09-22 1:28 ` Chaoyi Chen
2026-09-24 9:08 ` Jiaxing Hu
2026-09-24 9:08 ` Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 10/14] pmdomain: rockchip: cycle optional power-domain resets on power-on Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:56 ` sashiko-bot
2026-09-21 12:43 ` Ulf Hansson
2026-09-21 12:43 ` Ulf Hansson
2026-09-23 9:38 ` Philipp Zabel
2026-09-23 9:38 ` Philipp Zabel
2026-09-15 10:43 ` [PATCH v13 11/14] accel/rocket: select the per-core clock and reset counts from match data Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 12/14] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 13/14] arm64: dts: rockchip: add NPU (RKNN) nodes to rk3576 Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-21 21:51 ` Heiko Stuebner
2026-09-21 21:51 ` Heiko Stuebner
2026-09-15 10:43 ` [PATCH v13 14/14] arm64: dts: rockchip: enable the NPU on rk3576-rock-4d Jiaxing Hu
2026-09-15 10:43 ` Jiaxing Hu
2026-09-19 7:32 ` [PATCH v13 00/14] accel/rocket: RK3576 NPU (RKNN) enablement Sidong Yang
2026-09-19 7:32 ` Sidong Yang
2026-09-19 9:17 ` Jiaxing Hu
2026-09-19 9:17 ` Jiaxing Hu
2026-09-21 12:46 ` Ulf Hansson
2026-09-21 12:46 ` Ulf Hansson
2026-09-24 9:08 ` Jiaxing Hu
2026-09-24 9:08 ` Jiaxing Hu
2026-09-24 13:48 ` Ulf Hansson
2026-09-24 13:48 ` Ulf Hansson
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260919091725.8981-1-gahing@gahingwoo.com \
--to=gahing@gahingwoo.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=linux-rockchip@lists.infradead.org \
--cc=royalnet026@gmail.com \
--cc=tomeu@tomeuvizoso.net \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.