From: Jiaxing Hu <gahing@gahingwoo.com>
To: royalnet026@gmail.com
Cc: tomeu@tomeuvizoso.net, heiko@sntech.de,
chaoyi.chen@rock-chips.com, alchark@flipper.net,
dri-devel@lists.freedesktop.org,
linux-rockchip@lists.infradead.org,
linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support
Date: Mon, 17 Aug 2026 20:31:26 +1200 [thread overview]
Message-ID: <20260817083126.984173-1-gahing@gahingwoo.com> (raw)
In-Reply-To: <CAEWPSH7ETfw2TDjUfOvcVtDm19FC43x2vP=Lh7m+XPOGptnPWA@mail.gmail.com>
Hi Igor,
That is the measurement I could not make from here, and between your row and one
of mine the question is answered.
The control you propose is not one I can run in that form. Upstream Mesa has no
RK3576 path, so its encoder emits the RK3588 register layout, and this SoC does
not share that layout at the same offsets. I would not expect the result to be a
convolution at all, and I have not run it to find out.
What I can run is the vendor userspace, which is the same trade the other way
around, same silicon and an entirely different stack. Five models, output tensor
read as int8, the NPU interrupt count advancing by exactly one per run so
each of them reached the hardware.
model zero point values below it sitting exactly on it
a_lin 17 2478 of 4096 60.50% 32
a_lin2 -14 2420 of 4096 59.08% 41
g_cal -8 108564 of 204800 53.01% 2551
pq_oc 0 62720 of 128576 48.78% 3136
w_160 0 236287 of 409600 57.69% 2877
g_cal is conv2d-cal's geometry exactly, 16 input channels to 128 output over an
80x80 surface, 5x5 at stride 2. pq_oc and w_160 carry conv2d-cal's zero point
exactly, 0 in int8 being the 128 my tables report in uint8. So the geometry and
the zero point are each covered by a model that does not clamp. The counts
sitting on the zero point are 0.7 to 2.4 percent of the surface, which is close
to what you measured on the other side and close to what the distribution gives.
The grid now reads
RK3576, vendor userspace does not clamp
RK3588, upstream Mesa does not clamp, your run
RK3576, my Mesa clamps
The first and third rows are the same silicon. So the clamp is mine, and there
is no hardware behaviour left for me to appeal to. Your third control is the one
that makes your row carry weight, since fifteen channels off by one is something
a delegate falling back to the CPU could not produce.
I have not found it yet. The register stream is byte identical to the vendor's
at this geometry apart from addresses, the requantisation, the pad value and the
padding. A, B and C swapped one at a time from the vendor's records each leave
the floor where it is. The weight buffer is the last thing I have not compared.
The 88 and 120 sweep is queued for the next time the board is flashed, and I
will send what it says either way.
Jiaxing
next prev parent reply other threads:[~2026-08-17 8:31 UTC|newest]
Thread overview: 34+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-12 9:40 [PATCH v7 00/10] accel/rocket: RK3576 NPU (RKNN) enablement Jiaxing Hu
2026-08-12 9:40 ` [PATCH v7 01/10] accel/rocket: take the completion register writes under job_lock Jiaxing Hu
2026-08-12 12:47 ` Igor Paunovic
2026-08-12 9:40 ` [PATCH v7 02/10] dt-bindings: npu: rockchip: add rockchip,rk3576-rknn-core Jiaxing Hu
2026-08-13 7:04 ` Krzysztof Kozlowski
2026-08-12 9:40 ` [PATCH v7 03/10] dt-bindings: power: rockchip: allow resets in a power domain node Jiaxing Hu
2026-08-13 7:06 ` Krzysztof Kozlowski
2026-08-14 8:21 ` Jiaxing Hu
2026-08-12 9:40 ` [PATCH v7 04/10] dt-bindings: iommu: rockchip: allow the RK3576 NPU MMU clock set Jiaxing Hu
2026-08-12 10:45 ` Diederik de Haas
2026-08-13 9:27 ` Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 05/10] pmdomain/rockchip: add optional per-domain power-on settle delay Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 06/10] pmdomain/rockchip: cycle optional power-domain resets on power-on Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 07/10] accel/rocket: select the per-core clock and reset counts from match data Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 08/10] accel/rocket: add RK3576 NPU (RKNN) support Jiaxing Hu
2026-08-12 12:48 ` Igor Paunovic
2026-08-13 9:26 ` Jiaxing Hu
2026-08-13 9:56 ` Igor Paunovic
2026-08-14 8:26 ` Jiaxing Hu
2026-08-14 11:08 ` Igor Paunovic
[not found] ` <20260814110841.11238-1-royalnet026@gmail.com>
2026-08-15 3:12 ` Jiaxing Hu
2026-08-15 13:05 ` Igor Paunovic
2026-08-16 4:12 ` Jiaxing Hu
2026-08-16 18:53 ` Igor Paunovic
2026-08-16 19:58 ` Jiaxing Hu
2026-08-16 20:25 ` Igor Paunovic
2026-08-17 8:31 ` Jiaxing Hu [this message]
2026-08-17 9:45 ` Jiaxing Hu
2026-08-17 10:00 ` Igor Paunovic
2026-08-17 10:20 ` Jiaxing Hu
2026-08-17 11:05 ` Igor Paunovic
2026-08-12 9:41 ` [PATCH v7 09/10] arm64: dts: rockchip: rk3576: add NPU (RKNN) nodes Jiaxing Hu
2026-08-12 9:41 ` [PATCH v7 10/10] arm64: dts: rockchip: rk3576-rock-4d: enable NPU Jiaxing Hu
2026-08-12 10:20 ` Chaoyi Chen
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260817083126.984173-1-gahing@gahingwoo.com \
--to=gahing@gahingwoo.com \
--cc=alchark@flipper.net \
--cc=chaoyi.chen@rock-chips.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=heiko@sntech.de \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rockchip@lists.infradead.org \
--cc=royalnet026@gmail.com \
--cc=tomeu@tomeuvizoso.net \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox