dri-devel Archive on lore.kernel.org
 help / color / mirror / Atom feed
From: Igor Paunovic <royalnet026@gmail.com>
To: Jiaxing Hu <gahing@gahingwoo.com>
Cc: Igor Paunovic <royalnet026@gmail.com>,
	Tomeu Vizoso <tomeu@tomeuvizoso.net>,
	linux-rockchip@lists.infradead.org,
	dri-devel@lists.freedesktop.org
Subject: Re: [PATCH v13 03/14] accel/rocket: wait for a running IRQ handler before resetting a core
Date: Sat, 19 Sep 2026 12:34:22 +0200	[thread overview]
Message-ID: <20260919103422.148834-1-royalnet026@gmail.com> (raw)
In-Reply-To: <20260919091725.8981-1-gahing@gahingwoo.com>

Hi Jiaxing,

On Sat, Sep 19, 2026 at 09:17:25PM +1200, Jiaxing Hu wrote:
> I read the tally as 8 of the 126 scored inferences missing on all 48
> channels across the three runs, and 5 of those, in the two runs you
> traced, sitting on the 5 -ECANCELED completions.

Nested that way, yes. One correction: six runs, three per arm. Five
of the eight zeros are on 2+3 and three on 2+3+4. In the two traced
runs the -ECANCELED completions sit on exactly the inferences that
scored 0/48 there: rounds 1, 4, 8 and the one after suspend/resume
on 2+3, round 7 on 2+3+4. The other three zeros are in runs without
the kprobe.

After that mail, on the evening of 16 September, I reached the other
path. A 1x1 convolution gets there when the input, not the weights,
overflows the CBUF: 80x80x64 in, 48 out, so rkt_split_tasks() splits
on input rows and rkt_ml_subgraph_invoke() submits one job with
task_count 2 (rows 0-63 and 64-79). A graph whose weights do not fit
goes the other way: reuse_weights_cbuf is false and each task becomes
its own one-task job, back on the scheduler thread.

Same board and config (JOB_TIMEOUT_MS=2, PROVE_LOCKING,
DEBUG_ATOMIC_SLEEP), the induced-reset protocol at 60 inferences a
run. On 2+3: five runs, 310 inferences, 180 timeouts, one of the runs
with a 32-output variant that splits the same way. On 2+3+4: three
runs, 186 inferences, 181 timeouts. In the three traced runs on 2+3,
24 of 110 timeouts cut into a job after the IRQ thread had submitted
its next task (9, 9 and 6); in the one traced run on 2+3+4, 11 of
61. No lockdep report, warning or MMU fault in either boot, and
debug_locks 1 at the end of every run.

So the lock scope ran with the IRQ thread submitting tasks, but this
does not show the race is closed: the window in hw_submit() is
microseconds, one reset began while the IRQ thread was in it (in the
32-output run), and waiting on job_lock itself was not instrumented.

One result bears on 4/14. Without it the core stayed active after
every cancelled job (113 of 113) and no cancelled job was followed by
another (0 of 109). With it the core suspended after each (122 of
122) and 84 of 120 cancelled jobs were followed by another, the next
job waiting for resume against the same 2 ms. That is the timeout
feeding itself, not a defect, and it is why 181 of 186 inferences
timed out on that arm against 180 of 310. Core 0, one client. If
4/14 changes shape in v14 I will re-run that arm, with the two-task
job in it.

Scripts, counting and this draft were prepared with an LLM assistant;
I ran the tests, and every number came from the raw files of each run.

Regards,
Igor

  reply	other threads:[~2026-09-19 10:34 UTC|newest]

Thread overview: 33+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-15 10:43 [PATCH v13 00/14] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 01/14] accel/rocket: request the core clocks by name Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 02/14] accel/rocket: take the completion register writes under job_lock Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 03/14] accel/rocket: wait for a running IRQ handler before resetting a core Jiaxing Hu
2026-09-15 10:59   ` sashiko-bot
2026-09-16 13:28   ` Igor Paunovic
2026-09-19  9:17     ` Jiaxing Hu
2026-09-19 10:34       ` Igor Paunovic [this message]
2026-09-15 10:43 ` [PATCH v13 04/14] accel/rocket: let the core suspend after a reset Jiaxing Hu
2026-09-15 10:58   ` sashiko-bot
2026-09-15 10:43 ` [PATCH v13 05/14] accel/rocket: factor the completion tail out of the IRQ handler Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 06/14] dt-bindings: npu: rockchip: add rockchip, rk3576-rknn-core Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 07/14] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
2026-09-21 21:52   ` Heiko Stuebner
2026-09-15 10:43 ` [PATCH v13 08/14] dt-bindings: iommu: rockchip: describe the RK3576 NPU MMU Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 09/14] pmdomain: rockchip: add optional per-domain power-on settle delay Jiaxing Hu
2026-09-21 12:41   ` Ulf Hansson
2026-09-21 22:06   ` Heiko Stuebner
2026-09-22  1:28     ` Chaoyi Chen
2026-09-15 10:43 ` [PATCH v13 10/14] pmdomain: rockchip: cycle optional power-domain resets on power-on Jiaxing Hu
2026-09-15 10:56   ` sashiko-bot
2026-09-21 12:43   ` Ulf Hansson
2026-09-23  9:38   ` Philipp Zabel
2026-09-15 10:43 ` [PATCH v13 11/14] accel/rocket: select the per-core clock and reset counts from match data Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 12/14] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
2026-09-15 10:43 ` [PATCH v13 13/14] arm64: dts: rockchip: add NPU (RKNN) nodes to rk3576 Jiaxing Hu
2026-09-21 21:51   ` Heiko Stuebner
2026-09-15 10:43 ` [PATCH v13 14/14] arm64: dts: rockchip: enable the NPU on rk3576-rock-4d Jiaxing Hu
2026-09-19  7:32 ` [PATCH v13 00/14] accel/rocket: RK3576 NPU (RKNN) enablement Sidong Yang
2026-09-19  9:17   ` Jiaxing Hu
2026-09-21 12:46 ` Ulf Hansson
2026-09-24  9:08   ` Jiaxing Hu
2026-09-24 13:48     ` Ulf Hansson

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260919103422.148834-1-royalnet026@gmail.com \
    --to=royalnet026@gmail.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=gahing@gahingwoo.com \
    --cc=linux-rockchip@lists.infradead.org \
    --cc=tomeu@tomeuvizoso.net \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox