* [PATCH v2 01/17] dt-bindings: media: add apple,avd
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-29 8:46 ` Krzysztof Kozlowski
2026-09-26 13:14 ` [PATCH v2 02/17] media: v4l2: Add P210 pixel format Sofus Forstreuter
` (15 subsequent siblings)
16 siblings, 1 reply; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Document the Apple Video Decoder. A hardware video decoder present on
Apple hardware.
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
.../devicetree/bindings/media/apple,avd.yaml | 127 +++++++++++++++++++++
MAINTAINERS | 1 +
2 files changed, 128 insertions(+)
diff --git a/Documentation/devicetree/bindings/media/apple,avd.yaml b/Documentation/devicetree/bindings/media/apple,avd.yaml
new file mode 100644
index 000000000000..0aa2cffae50f
--- /dev/null
+++ b/Documentation/devicetree/bindings/media/apple,avd.yaml
@@ -0,0 +1,127 @@
+# SPDX-License-Identifier: GPL-2.0-only OR BSD-2-Clause
+%YAML 1.2
+---
+$id: http://devicetree.org/schemas/media/apple,avd.yaml#
+$schema: http://devicetree.org/meta-schemas/core.yaml#
+
+title: Apple Video Decoder (AVD)
+
+description: |
+ AVD is a stateless video decode IP block present on Apple devices.
+ Notably it includes an ARM Cortex-M3 and associated DMA controller.
+ AVD is capable of decoding H264, H265, VP9 and AV1 on newer SoCs.
+
+ The IP does not change much, with changes being mostly the CM3 having more
+ memory and interrupts.
+
+maintainers:
+ - Sofus Forstreuter <sofus.c@icloud.com>
+
+properties:
+ compatible:
+ oneOf:
+ - items:
+ - enum:
+ - apple,t6030-avd
+ - apple,t6031-avd
+ - const: apple,t8122-avd
+ - enum:
+ - apple,t8103-avd
+ - apple,t8112-avd
+ - apple,t8122-avd
+ - apple,t6000-avd
+ - apple,t6020-avd
+
+ reg:
+ items:
+ - description: DMA control
+ - description: Cortex-M3 code (unsigned)
+ - description: Cortex-M3 sram
+ - description: Mailbox
+ - description: Video decode control
+
+ reg-names:
+ items:
+ - const: piodma
+ - const: code
+ - const: sram
+ - const: mbox
+ - const: ctrl
+
+ iommus:
+ items:
+ - description: Decode related
+ - description: DMA controller
+
+ power-domains:
+ maxItems: 1
+
+ resets:
+ maxItems: 1
+
+ interrupts:
+ items:
+ - description: mailbox 0
+ - description: mailbox 1
+ - description: mailbox 2
+ - description: mailbox 3
+ - description: flag 0
+ - description: flag 1
+ - description: DMA status
+ description: |
+ Notably there are no decode related interrupts. All decode related
+ interrupts are sent to the CM3.
+
+ interrupt-names:
+ items:
+ - const: mbox0
+ - const: mbox1
+ - const: mbox2
+ - const: mbox3
+ - const: flag0
+ - const: flag1
+ - const: piodma
+
+required:
+ - compatible
+ - reg
+ - interrupts
+ - reg-names
+ - iommus
+ - power-domains
+ - resets
+
+additionalProperties: false
+
+examples:
+ - |
+ #include <dt-bindings/interrupt-controller/apple-aic.h>
+ #include <dt-bindings/interrupt-controller/irq.h>
+
+ soc {
+ #address-cells = <2>;
+ #size-cells = <2>;
+
+ video-codec@289070000 {
+ compatible = "apple,t8122-avd";
+ reg = <0x2 0x89070000 0x0 0x4000>,
+ <0x2 0x89080000 0x0 0x10000>,
+ <0x2 0x89092000 0x0 0x14000>,
+ <0x2 0x890a4000 0x0 0x4000>,
+ <0x2 0x89100000 0x0 0x10000>;
+ reg-names = "piodma", "code", "sram", "mbox", "ctrl";
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 679 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 680 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 681 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 682 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 683 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 684 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 685 IRQ_TYPE_LEVEL_HIGH>;
+ interrupt-names = "mbox0", "mbox1", "mbox2", "mbox3",
+ "flag0", "flag1", "piodma";
+ power-domains = <&ps_avd_sys>;
+ resets = <&ps_avd_sys>;
+ iommus = <&avd_dart 0>, <&avd_dart 1>;
+ };
+ };
diff --git a/MAINTAINERS b/MAINTAINERS
index 4cc4a2dc6d30..3c7c5fe3f0bc 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -2643,6 +2643,7 @@ F: Documentation/devicetree/bindings/iommu/apple,dart.yaml
F: Documentation/devicetree/bindings/iommu/apple,sart.yaml
F: Documentation/devicetree/bindings/leds/backlight/apple,dwi-bl.yaml
F: Documentation/devicetree/bindings/mailbox/apple,mailbox.yaml
+F: Documentation/devicetree/bindings/media/apple,avd.yaml
F: Documentation/devicetree/bindings/mfd/apple,smc.yaml
F: Documentation/devicetree/bindings/net/bluetooth/brcm,bcm4377-bluetooth.yaml
F: Documentation/devicetree/bindings/nvme/apple,nvme-ans.yaml
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* Re: [PATCH v2 01/17] dt-bindings: media: add apple,avd
2026-09-26 13:14 ` [PATCH v2 01/17] dt-bindings: media: add apple,avd Sofus Forstreuter
@ 2026-09-29 8:46 ` Krzysztof Kozlowski
0 siblings, 0 replies; 19+ messages in thread
From: Krzysztof Kozlowski @ 2026-09-29 8:46 UTC (permalink / raw)
To: Sofus Forstreuter
Cc: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Philipp Zabel,
Heiko Stuebner, asahi, linux-arm-kernel, linux-media, devicetree,
linux-kernel, linux-rockchip
On Sat, Sep 26, 2026 at 03:14:40PM +0200, Sofus Forstreuter wrote:
> + compatible:
> + oneOf:
> + - items:
> + - enum:
> + - apple,t6030-avd
> + - apple,t6031-avd
> + - const: apple,t8122-avd
> + - enum:
> + - apple,t8103-avd
> + - apple,t8112-avd
> + - apple,t8122-avd
> + - apple,t6000-avd
> + - apple,t6020-avd
These are misordered. Order is alphabetical.
With this fixed:
Reviewed-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>
Best regards,
Krzysztof
^ permalink raw reply [flat|nested] 19+ messages in thread
* [PATCH v2 02/17] media: v4l2: Add P210 pixel format
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 01/17] dt-bindings: media: add apple,avd Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 03/17] media: v4l2: Add Apple interchange pixel formats Sofus Forstreuter
` (14 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
P210 is a YUV format with 10-bits per component with interleaved UV.
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
.../userspace-api/media/v4l/pixfmt-yuv-planar.rst | 14 ++++++++++++--
drivers/media/v4l2-core/v4l2-common.c | 1 +
drivers/media/v4l2-core/v4l2-ioctl.c | 1 +
include/uapi/linux/videodev2.h | 1 +
4 files changed, 15 insertions(+), 2 deletions(-)
diff --git a/Documentation/userspace-api/media/v4l/pixfmt-yuv-planar.rst b/Documentation/userspace-api/media/v4l/pixfmt-yuv-planar.rst
index 0631919bd667..f2d2e79698fa 100644
--- a/Documentation/userspace-api/media/v4l/pixfmt-yuv-planar.rst
+++ b/Documentation/userspace-api/media/v4l/pixfmt-yuv-planar.rst
@@ -138,6 +138,13 @@ All components are stored with the same number of bits per component.
- Cb, Cr
- No
- Linear
+ * - V4L2_PIX_FMT_P210
+ - 'P210'
+ - 10
+ - 4:2:2
+ - Cb, Cr
+ - Yes
+ - Linear
* - V4L2_PIX_FMT_NV15
- 'NV15'
- 10
@@ -834,12 +841,15 @@ number of lines as the luma plane.
.. _V4L2_PIX_FMT_P010:
.. _V4L2-PIX-FMT-P010-4L4:
+.. _V4L2-PIX-FMT-P210:
-P010 and tiled P010
--------------------
+P010, tiled P010 and P210
+-------------------------
P010 is like NV12 with 10 bits per component, expanded to 16 bits.
Data in the 10 high bits, zeros in the 6 low bits, arranged in little endian order.
+P210 is the same as P010 but the chroma plane is only subsampled by 2 in the
+horizontal direction.
.. flat-table:: Sample 4x4 P010 Image
:header-rows: 0
diff --git a/drivers/media/v4l2-core/v4l2-common.c b/drivers/media/v4l2-core/v4l2-common.c
index c825820fd481..b78fc0845a16 100644
--- a/drivers/media/v4l2-core/v4l2-common.c
+++ b/drivers/media/v4l2-core/v4l2-common.c
@@ -325,6 +325,7 @@ const struct v4l2_format_info *v4l2_format_info(u32 format)
{ .format = V4L2_PIX_FMT_NV42, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 1, 2, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 1, .vdiv = 1 },
{ .format = V4L2_PIX_FMT_P010, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 2, 2, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 2, .vdiv = 1 },
{ .format = V4L2_PIX_FMT_P012, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 2, 4, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 2, .vdiv = 2 },
+ { .format = V4L2_PIX_FMT_P210, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 2, 4, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 2, .vdiv = 1 },
{ .format = V4L2_PIX_FMT_YUV410, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 3, .bpp = { 1, 1, 1, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 4, .vdiv = 4 },
{ .format = V4L2_PIX_FMT_YVU410, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 3, .bpp = { 1, 1, 1, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 4, .vdiv = 4 },
diff --git a/drivers/media/v4l2-core/v4l2-ioctl.c b/drivers/media/v4l2-core/v4l2-ioctl.c
index 17ba1ae70735..1f2519ffeb2b 100644
--- a/drivers/media/v4l2-core/v4l2-ioctl.c
+++ b/drivers/media/v4l2-core/v4l2-ioctl.c
@@ -1369,6 +1369,7 @@ static void v4l_fill_fmtdesc(struct v4l2_fmtdesc *fmt)
case V4L2_PIX_FMT_NV42: descr = "Y/VU 4:4:4"; break;
case V4L2_PIX_FMT_P010: descr = "10-bit Y/UV 4:2:0"; break;
case V4L2_PIX_FMT_P012: descr = "12-bit Y/UV 4:2:0"; break;
+ case V4L2_PIX_FMT_P210: descr = "10-bit Y/UV 4:2:2"; break;
case V4L2_PIX_FMT_NV12_4L4: descr = "Y/UV 4:2:0 (4x4 Linear)"; break;
case V4L2_PIX_FMT_NV12_16L16: descr = "Y/UV 4:2:0 (16x16 Linear)"; break;
case V4L2_PIX_FMT_NV12_32L32: descr = "Y/UV 4:2:0 (32x32 Linear)"; break;
diff --git a/include/uapi/linux/videodev2.h b/include/uapi/linux/videodev2.h
index 5373dba640fa..5472e433ed06 100644
--- a/include/uapi/linux/videodev2.h
+++ b/include/uapi/linux/videodev2.h
@@ -659,6 +659,7 @@ struct v4l2_pix_format {
#define V4L2_PIX_FMT_NV42 v4l2_fourcc('N', 'V', '4', '2') /* 24 Y/CrCb 4:4:4 */
#define V4L2_PIX_FMT_P010 v4l2_fourcc('P', '0', '1', '0') /* 24 Y/CbCr 4:2:0 10-bit per component */
#define V4L2_PIX_FMT_P012 v4l2_fourcc('P', '0', '1', '2') /* 24 Y/CbCr 4:2:0 12-bit per component */
+#define V4L2_PIX_FMT_P210 v4l2_fourcc('P', '2', '1', '0') /* 32 Y/CbCr 4:2:2 10-bit per component */
/* two non contiguous planes - one Y, one Cr + Cb interleaved */
#define V4L2_PIX_FMT_NV12M v4l2_fourcc('N', 'M', '1', '2') /* 12 Y/CbCr 4:2:0 */
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 03/17] media: v4l2: Add Apple interchange pixel formats
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 01/17] dt-bindings: media: add apple,avd Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 02/17] media: v4l2: Add P210 pixel format Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 04/17] media: v4l2-ctrls: validate av1 tile info Sofus Forstreuter
` (13 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Apple Interchange is a compressed and tiled format modifier.
Its intended use is for sharing buffers between hardware blocks (GPU,
display controller, video codecs) on Apple platforms.
Add it to the v4l2_format_info table to help with calculating sizes.
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
drivers/media/v4l2-core/v4l2-common.c | 7 +++++++
drivers/media/v4l2-core/v4l2-ioctl.c | 6 ++++++
include/uapi/linux/videodev2.h | 8 ++++++++
3 files changed, 21 insertions(+)
diff --git a/drivers/media/v4l2-core/v4l2-common.c b/drivers/media/v4l2-core/v4l2-common.c
index b78fc0845a16..eb2eb1c59fdd 100644
--- a/drivers/media/v4l2-core/v4l2-common.c
+++ b/drivers/media/v4l2-core/v4l2-common.c
@@ -336,6 +336,13 @@ const struct v4l2_format_info *v4l2_format_info(u32 format)
{ .format = V4L2_PIX_FMT_GREY, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 1, .bpp = { 1, 0, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 1, .vdiv = 1 },
{ .format = V4L2_PIX_FMT_Y12, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 1, .bpp = { 2, 0, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 1, .vdiv = 1 },
+ { .format = V4L2_PIX_FMT_IC12, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 1, 2, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 2, .vdiv = 2 },
+ { .format = V4L2_PIX_FMT_IC16, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 1, 2, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 2, .vdiv = 1 },
+ { .format = V4L2_PIX_FMT_IC24, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 1, 2, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 1, .vdiv = 1 },
+ { .format = V4L2_PIX_FMT_IC03, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 5, 10, 0, 0 }, .bpp_div = { 4, 4, 1, 1 }, .hdiv = 2, .vdiv = 2 },
+ { .format = V4L2_PIX_FMT_IC23, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 5, 10, 0, 0 }, .bpp_div = { 4, 4, 1, 1 }, .hdiv = 2, .vdiv = 1 },
+ { .format = V4L2_PIX_FMT_IC43, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 5, 10, 0, 0 }, .bpp_div = { 4, 4, 1, 1 }, .hdiv = 1, .vdiv = 1 },
+
/* Tiled YUV formats */
{ .format = V4L2_PIX_FMT_NV12_4L4, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 1, 2, 0, 0 }, .bpp_div = { 1, 1, 1, 1 }, .hdiv = 2, .vdiv = 2 },
{ .format = V4L2_PIX_FMT_NV15_4L4, .pixel_enc = V4L2_PIXEL_ENC_YUV, .mem_planes = 1, .comp_planes = 2, .bpp = { 5, 10, 0, 0 }, .bpp_div = { 4, 4, 1, 1 }, .hdiv = 2, .vdiv = 2,
diff --git a/drivers/media/v4l2-core/v4l2-ioctl.c b/drivers/media/v4l2-core/v4l2-ioctl.c
index 1f2519ffeb2b..b3c49b08ef2f 100644
--- a/drivers/media/v4l2-core/v4l2-ioctl.c
+++ b/drivers/media/v4l2-core/v4l2-ioctl.c
@@ -1561,6 +1561,12 @@ static void v4l_fill_fmtdesc(struct v4l2_fmtdesc *fmt)
case V4L2_PIX_FMT_PISP_COMP2_GBRG: descr = "PiSP 8b GBGB/RGRG mode2 compr"; break;
case V4L2_PIX_FMT_PISP_COMP2_BGGR: descr = "PiSP 8b BGBG/GRGR mode2 compr"; break;
case V4L2_PIX_FMT_PISP_COMP2_MONO: descr = "PiSP 8b monochrome mode2 compr"; break;
+ case V4L2_PIX_FMT_IC12: descr = "Interchange NV12"; break;
+ case V4L2_PIX_FMT_IC16: descr = "Interchange NV16"; break;
+ case V4L2_PIX_FMT_IC24: descr = "Interchange NV24"; break;
+ case V4L2_PIX_FMT_IC03: descr = "Interchange P030"; break;
+ case V4L2_PIX_FMT_IC23: descr = "Interchange P230"; break;
+ case V4L2_PIX_FMT_IC43: descr = "Interchange P430"; break;
default:
if (fmt->description[0])
return;
diff --git a/include/uapi/linux/videodev2.h b/include/uapi/linux/videodev2.h
index 5472e433ed06..918e060b6fe0 100644
--- a/include/uapi/linux/videodev2.h
+++ b/include/uapi/linux/videodev2.h
@@ -842,6 +842,14 @@ struct v4l2_pix_format {
#define V4L2_PIX_FMT_PISP_COMP2_BGGR v4l2_fourcc('P', 'C', '2', 'B') /* PiSP 8-bit mode 2 compressed BGGR bayer */
#define V4L2_PIX_FMT_PISP_COMP2_MONO v4l2_fourcc('P', 'C', '2', 'M') /* PiSP 8-bit mode 2 compressed monochrome */
+/* Apple Interchange, compressed and tiled formats */
+#define V4L2_PIX_FMT_IC12 v4l2_fourcc('I', 'C', '1', '2') /* Interchange NV12 */
+#define V4L2_PIX_FMT_IC16 v4l2_fourcc('I', 'C', '1', '6') /* Interchange NV16 */
+#define V4L2_PIX_FMT_IC24 v4l2_fourcc('I', 'C', '2', '4') /* Interchange NV24 */
+#define V4L2_PIX_FMT_IC03 v4l2_fourcc('I', 'C', '0', '3') /* Interchange P030 */
+#define V4L2_PIX_FMT_IC23 v4l2_fourcc('I', 'C', '2', '3') /* Interchange P230 */
+#define V4L2_PIX_FMT_IC43 v4l2_fourcc('I', 'C', '4', '3') /* Interchange P430 */
+
/* Renesas RZ/V2H CRU packed formats. 64-bit units with contiguous pixels */
#define V4L2_PIX_FMT_RAW_CRU10 v4l2_fourcc('C', 'R', '1', '0')
#define V4L2_PIX_FMT_RAW_CRU12 v4l2_fourcc('C', 'R', '1', '2')
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 04/17] media: v4l2-ctrls: validate av1 tile info
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (2 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 03/17] media: v4l2: Add Apple interchange pixel formats Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 05/17] media: v4l2-ctrls: validate vp9 tile_rows_log2 Sofus Forstreuter
` (12 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Many drivers use tile_cols and/or tile_rows, so make sure they are
valid.
Increase V4L2_AV1_MAX_TILE_COUNT to MAX_TILE_COLS * MAX_TILE_ROWS
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
drivers/media/v4l2-core/v4l2-ctrls-core.c | 15 +++++++++++++++
include/uapi/linux/v4l2-controls.h | 2 +-
2 files changed, 16 insertions(+), 1 deletion(-)
diff --git a/drivers/media/v4l2-core/v4l2-ctrls-core.c b/drivers/media/v4l2-core/v4l2-ctrls-core.c
index 661a3a25da52..c3e7262ba4d3 100644
--- a/drivers/media/v4l2-core/v4l2-ctrls-core.c
+++ b/drivers/media/v4l2-core/v4l2-ctrls-core.c
@@ -800,10 +800,25 @@ static int validate_av1_film_grain(struct v4l2_ctrl_av1_film_grain *fg)
return 0;
}
+static int validate_av1_tile_info(struct v4l2_av1_tile_info *ti)
+{
+ if (ti->tile_cols > V4L2_AV1_MAX_TILE_COLS ||
+ ti->tile_rows > V4L2_AV1_MAX_TILE_ROWS)
+ return -EINVAL;
+
+ if (ti->context_update_tile_id > ti->tile_cols * ti->tile_rows)
+ return -EINVAL;
+
+ return 0;
+}
+
static int validate_av1_frame(struct v4l2_ctrl_av1_frame *f)
{
int ret = 0;
+ ret = validate_av1_tile_info(&f->tile_info);
+ if (ret)
+ return ret;
ret = validate_av1_quantization(&f->quantization);
if (ret)
return ret;
diff --git a/include/uapi/linux/v4l2-controls.h b/include/uapi/linux/v4l2-controls.h
index d17e41d51d2e..abdfa102a77b 100644
--- a/include/uapi/linux/v4l2-controls.h
+++ b/include/uapi/linux/v4l2-controls.h
@@ -2939,7 +2939,7 @@ struct v4l2_ctrl_vp9_compressed_hdr {
#define V4L2_AV1_MAX_NUM_PLANES 3
#define V4L2_AV1_MAX_TILE_COLS 64
#define V4L2_AV1_MAX_TILE_ROWS 64
-#define V4L2_AV1_MAX_TILE_COUNT 512
+#define V4L2_AV1_MAX_TILE_COUNT 4096
#define V4L2_AV1_SEQUENCE_FLAG_STILL_PICTURE 0x00000001
#define V4L2_AV1_SEQUENCE_FLAG_USE_128X128_SUPERBLOCK 0x00000002
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 05/17] media: v4l2-ctrls: validate vp9 tile_rows_log2
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (3 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 04/17] media: v4l2-ctrls: validate av1 tile info Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 06/17] media: apple: add avd driver Sofus Forstreuter
` (11 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
drivers/media/v4l2-core/v4l2-ctrls-core.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/media/v4l2-core/v4l2-ctrls-core.c b/drivers/media/v4l2-core/v4l2-ctrls-core.c
index c3e7262ba4d3..39f7cd9b0a2d 100644
--- a/drivers/media/v4l2-core/v4l2-ctrls-core.c
+++ b/drivers/media/v4l2-core/v4l2-ctrls-core.c
@@ -638,7 +638,7 @@ validate_vp9_frame(struct v4l2_ctrl_vp9_frame *frame)
* According to the spec, tile_cols_log2 shall be less than or equal
* to 6.
*/
- if (frame->tile_cols_log2 > 6)
+ if (frame->tile_cols_log2 > 6 || frame->tile_rows_log2 > 2)
return -EINVAL;
if (frame->reference_mode > V4L2_VP9_REFERENCE_MODE_SELECT)
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 06/17] media: apple: add avd driver
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (4 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 05/17] media: v4l2-ctrls: validate vp9 tile_rows_log2 Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 07/17] media: apple: avd: add h264 support Sofus Forstreuter
` (10 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Add the AVD (Apple Video Decoder) driver with V4L2 M2M stateless
support based largely on rockchips implementation.
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
MAINTAINERS | 1 +
drivers/media/platform/Kconfig | 1 +
drivers/media/platform/Makefile | 1 +
drivers/media/platform/apple/Kconfig | 5 +
drivers/media/platform/apple/Makefile | 3 +
drivers/media/platform/apple/avd/Kconfig | 16 +
drivers/media/platform/apple/avd/Makefile | 4 +
drivers/media/platform/apple/avd/avd-drv.c | 776 +++++++++++++++++++++++++++
drivers/media/platform/apple/avd/avd-inst.h | 210 ++++++++
drivers/media/platform/apple/avd/avd-v4l2.c | 780 ++++++++++++++++++++++++++++
drivers/media/platform/apple/avd/avd.h | 267 ++++++++++
11 files changed, 2064 insertions(+)
diff --git a/MAINTAINERS b/MAINTAINERS
index 3c7c5fe3f0bc..6903baa15955 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -2676,6 +2676,7 @@ F: drivers/input/touchscreen/apple_z2.c
F: drivers/iommu/apple-dart.c
F: drivers/iommu/io-pgtable-dart.c
F: drivers/irqchip/irq-apple-aic.c
+F: drivers/media/platform/apple/*
F: drivers/mfd/macsmc.c
F: drivers/nvme/host/apple.c
F: drivers/nvmem/apple-efuses.c
diff --git a/drivers/media/platform/Kconfig b/drivers/media/platform/Kconfig
index 2c7699b6610b..280a9db25935 100644
--- a/drivers/media/platform/Kconfig
+++ b/drivers/media/platform/Kconfig
@@ -66,6 +66,7 @@ source "drivers/media/platform/allegro-dvt/Kconfig"
source "drivers/media/platform/amd/Kconfig"
source "drivers/media/platform/amlogic/Kconfig"
source "drivers/media/platform/amphion/Kconfig"
+source "drivers/media/platform/apple/Kconfig"
source "drivers/media/platform/arm/Kconfig"
source "drivers/media/platform/aspeed/Kconfig"
source "drivers/media/platform/atmel/Kconfig"
diff --git a/drivers/media/platform/Makefile b/drivers/media/platform/Makefile
index d47c47d817da..aa82e189936b 100644
--- a/drivers/media/platform/Makefile
+++ b/drivers/media/platform/Makefile
@@ -9,6 +9,7 @@ obj-y += allegro-dvt/
obj-y += amd/
obj-y += amlogic/
obj-y += amphion/
+obj-y += apple/
obj-y += arm/
obj-y += aspeed/
obj-y += atmel/
diff --git a/drivers/media/platform/apple/Kconfig b/drivers/media/platform/apple/Kconfig
new file mode 100644
index 000000000000..43c0a56c36a8
--- /dev/null
+++ b/drivers/media/platform/apple/Kconfig
@@ -0,0 +1,5 @@
+# SPDX-License-Identifier: GPL-2.0-only
+
+comment "Apple media platform drivers"
+
+source "drivers/media/platform/apple/avd/Kconfig"
diff --git a/drivers/media/platform/apple/Makefile b/drivers/media/platform/apple/Makefile
new file mode 100644
index 000000000000..d502cab93970
--- /dev/null
+++ b/drivers/media/platform/apple/Makefile
@@ -0,0 +1,3 @@
+# SPDX-License-Identifier: GPL-2.0-only
+
+obj-y += avd/
diff --git a/drivers/media/platform/apple/avd/Kconfig b/drivers/media/platform/apple/avd/Kconfig
new file mode 100644
index 000000000000..68efa2df368a
--- /dev/null
+++ b/drivers/media/platform/apple/avd/Kconfig
@@ -0,0 +1,16 @@
+# SPDX-License-Identifier: GPL-2.0
+
+config VIDEO_APPLE_AVD
+ tristate "Apple Silicon Video Decoding driver"
+ depends on VIDEO_DEV
+ depends on MEDIA_CONTROLLER
+ depends on ARCH_APPLE || COMPILE_TEST
+ depends on OF_ADDRESS
+ depends on V4L_PLATFORM_DRIVERS
+ select V4L2_MEM2MEM_DEV
+ select VIDEOBUF2_DMA_CONTIG
+ help
+ Support for hardware video decoding on Apple Silicon devices using
+ the Apple Video Decoder (AVD).
+ To compile this driver as a module, choose M here: the module will
+ be called apple-avd.
diff --git a/drivers/media/platform/apple/avd/Makefile b/drivers/media/platform/apple/avd/Makefile
new file mode 100644
index 000000000000..b2f6736a790e
--- /dev/null
+++ b/drivers/media/platform/apple/avd/Makefile
@@ -0,0 +1,4 @@
+# SPDX-License-Identifier: GPL-2.0-only
+
+apple-avd-y := avd-drv.o avd-v4l2.o
+obj-$(CONFIG_VIDEO_APPLE_AVD) += apple-avd.o
diff --git a/drivers/media/platform/apple/avd/avd-drv.c b/drivers/media/platform/apple/avd/avd-drv.c
new file mode 100644
index 000000000000..68752c44d4ec
--- /dev/null
+++ b/drivers/media/platform/apple/avd/avd-drv.c
@@ -0,0 +1,776 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Apple Video Decoder driver
+ *
+ * Copyright (C) 2026 The Asahi Linux Contributors
+ * Copyright (C) 2026 Sofus Forstreuter <sofus.c@icloud.com>
+ *
+ * Based on rkvdec driver by Collabora, Ltd.
+ * Copyright (C) 2019 Collabora, Ltd.
+ * Based on rkvdec driver by Google LLC. (Tomasz Figa <tfiga@chromium.org>)
+ * Based on s5p-mfc driver by Samsung Electronics Co., Ltd.
+ * Copyright (C) 2011 Samsung Electronics Co., Ltd.
+ */
+
+#include <linux/pm_runtime.h>
+#include <linux/iommu.h>
+#include <linux/reset.h>
+#include <linux/dev_printk.h>
+#include <linux/iopoll.h>
+
+#include <media/videobuf2-dma-contig.h>
+#include <media/videobuf2-v4l2.h>
+
+#include "avd.h"
+#include "avd-inst.h"
+
+static void calc_tile_meta(u32 w, u32 h, u32 bpb, u32 tile_dim,
+ u32 meta_hdr_bytes, u32 *tile, u32 *meta)
+{
+ u32 tiles_width, tiles_height, meta_tile_w, meta_tile_h, tile_bytes;
+
+ tiles_width = DIV_ROUND_UP(w, tile_dim);
+ tiles_height = DIV_ROUND_UP(h, tile_dim);
+ tile_bytes = DIV_ROUND_UP(tile_dim * tile_dim * bpb, 8);
+ *tile = ALIGN(tiles_width * tiles_height * tile_bytes, 16);
+
+ meta_tile_w = roundup_pow_of_two(tiles_width);
+ meta_tile_h = roundup_pow_of_two(tiles_height);
+
+ *meta = ALIGN(meta_tile_w * meta_tile_h * meta_hdr_bytes, 16);
+}
+
+static inline u8 v4l2_format_info_bpp(const struct v4l2_format_info *info,
+ int plane)
+{
+ return 8 * info->bpp[plane] / info->bpp_div[plane];
+}
+
+void fill_comp(struct avd_comp *comp, enum avd_image_fmt fmt, u32 width,
+ u32 height)
+{
+ u32 y_meta, y, uv_meta, uv, fourcc;
+ const struct v4l2_format_info *info;
+
+ switch (fmt) {
+ case AVD_IMG_FMT_ANY:
+ case AVD_IMG_FMT_420_8BIT:
+ fourcc = V4L2_PIX_FMT_IC12;
+ break;
+ case AVD_IMG_FMT_420_10BIT:
+ fourcc = V4L2_PIX_FMT_IC03;
+ break;
+ case AVD_IMG_FMT_422_8BIT:
+ fourcc = V4L2_PIX_FMT_IC16;
+ break;
+ case AVD_IMG_FMT_422_10BIT:
+ fourcc = V4L2_PIX_FMT_IC23;
+ break;
+ }
+ info = v4l2_format_info(fourcc);
+
+ /* y has 32x32 tiles and 32 bytes of metadata per tile */
+ calc_tile_meta(width, height, v4l2_format_info_bpp(info, 0), 32, 32,
+ &y, &y_meta);
+ /* uv has 16x16 tiles and 8 bytes of metadata per tile */
+ calc_tile_meta(width / info->vdiv, height / info->hdiv,
+ v4l2_format_info_bpp(info, 1), 16, 8, &uv,
+ &uv_meta);
+
+ /* output like DCP driver expects */
+ comp->offsets[0] = y;
+ comp->offsets[1] = 0;
+ comp->offsets[2] = y + y_meta + uv;
+ comp->offsets[3] = y + y_meta;
+
+ comp->size = y_meta + y + uv_meta + uv;
+}
+
+int avd_buf_alloc(struct avd_dev *avd, struct avd_buf *buf, size_t size)
+{
+ if (buf->cpu && size < buf->size)
+ return 0;
+ else if (buf->cpu)
+ avd_buf_free(avd, buf);
+
+ if (size <= 0)
+ return -ENOMEM;
+
+ buf->size = size;
+ buf->cpu =
+ dma_alloc_coherent(avd->dev, buf->size, &buf->addr, GFP_KERNEL);
+ return buf->cpu ? 0 : -ENOMEM;
+}
+
+void avd_buf_free(struct avd_dev *avd, struct avd_buf *buf)
+{
+ if (buf->cpu)
+ dma_free_coherent(avd->dev, buf->size, buf->cpu, buf->addr);
+ memset(buf, 0, sizeof(*buf));
+}
+
+struct avd_decoded_buffer *
+avd_get_ref_buf(struct avd_ctx *ctx, struct vb2_v4l2_buffer *dst, u64 timestamp)
+{
+ struct v4l2_m2m_ctx *m2m_ctx = ctx->fh.m2m_ctx;
+ struct vb2_queue *cap_q = &m2m_ctx->cap_q_ctx.q;
+ struct vb2_buffer *buf;
+
+ /*
+ * If a ref is unused or invalid, address of current destination
+ * buffer is returned.
+ */
+ buf = vb2_find_buffer(cap_q, timestamp);
+ if (!buf)
+ buf = &dst->vb2_buf;
+
+ return vb2_to_avd_decoded_buf(buf);
+}
+
+int avd_end_segment(struct avd_ctx *ctx, bool update_submit)
+{
+ struct avd_job *job = &ctx->job;
+ struct avd_segment *seg = &job->segments[job->num];
+
+ /* avd_segment includes piodma_cmd which is not transferred */
+ seg->piodma_cmd =
+ AVD_PIODMA_CMD_SIZE((sizeof(struct avd_segment) - 8) / 4);
+ seg->piodma_cmd |= AVD_PIODMA_CMD_DEST(job->dest);
+ seg->piodma_cmd |= AVD_PIODMA_CMD_CONST;
+
+ job->num++;
+ if (update_submit)
+ job->num_submit++;
+ return job->num >= job->num_alloc;
+}
+
+int avd_init_job(struct avd_ctx *ctx, enum avd_codec codec, size_t segments)
+{
+ int ret = 0;
+ struct avd_job *job = &ctx->job;
+
+ job->codec = codec;
+ job->dest = 0x1000;
+ job->num = 0;
+ job->num_submit = 0;
+ job->num_alloc = segments;
+ ret = avd_buf_alloc(ctx->dev, &job->buf,
+ job->num_alloc * sizeof(*job->segments));
+ job->segments = job->buf.cpu;
+ memset(job->buf.cpu, 0, job->buf.size);
+ return ret;
+}
+
+struct avd_cm3_job {
+ enum avd_codec codec;
+ u32 dest;
+ u32 num;
+ u32 num_submit;
+ u64 iova;
+ u64 insn;
+};
+
+int avd_submit_job(struct avd_ctx *ctx)
+{
+ struct avd_dev *avd = ctx->dev;
+ struct avd_job *job = &ctx->job;
+ int submit_off = 0x100;
+ struct avd_cm3_job submit = (struct avd_cm3_job) {
+ .codec = job->codec,
+ .dest = job->dest,
+ .num = job->num,
+ .num_submit = job->num_submit,
+ .iova = job->buf.addr,
+ .insn = ctx->inst.addr,
+ };
+
+ schedule_delayed_work(&ctx->watchdog_work, msecs_to_jiffies(2000));
+ memcpy_toio(avd->sram + submit_off, &submit, sizeof(submit));
+ writel(submit_off, avd->mbox + AVD_REG_MBOX1_SUBMIT);
+
+ return 0;
+}
+
+static int avd_boot(struct avd_dev *avd)
+{
+ u32 val;
+ int ret;
+ char version[64];
+
+ if (avd->variant->revision != 3)
+ dev_info_once(avd->dev, "booting hw version: %04x",
+ readl_relaxed(avd->ctrl));
+
+ writel(avd->sram_start, avd->piodma + 0x24);
+ dev_info_once(avd->dev, "piodma version: %04x base: %08x",
+ readl_relaxed(avd->piodma + 0xb4),
+ readl_relaxed(avd->piodma + 0x24));
+
+ memcpy_toio(avd->code, avd->fw->data, avd->fw->size);
+
+ writel_relaxed(AVD_MBOX_ENABLE, avd->mbox + AVD_REG_MBOX1_STATUS);
+ writel_relaxed(AVD_MBOX_ENABLE, avd->mbox + AVD_REG_MBOX0_STATUS);
+ writel_relaxed(AVD_MBOX0_NOT_EMPTY,
+ avd->mbox + AVD_REG_MBOX_IRQ_ENABLE);
+ writel_relaxed(AVD_RUN_CTRL_UNK_RUN, avd->mbox + AVD_REG_RUN_CTRL);
+
+ /* wait for cm3 to boot */
+ ret = readl_poll_timeout(avd->mbox + AVD_REG_FLAG0_SET, val, val == 1,
+ 10, 10000);
+ if (ret)
+ return ret;
+
+ memcpy_fromio(version, avd->sram, sizeof(version));
+ dev_info_once(avd->dev, "fw version: %s\n", version);
+
+ return 0;
+}
+
+static void avd_shutdown(struct avd_dev *avd)
+{
+ writel_relaxed(AVD_RUN_CTRL_UNK_STOP, avd->mbox + AVD_REG_RUN_CTRL);
+ writel_relaxed(1, avd->mbox + AVD_REG_FLAG0_CLR);
+ writel_relaxed(0, avd->mbox + AVD_REG_MBOX_IRQ_ENABLE);
+}
+
+static int avd_reset(struct avd_dev *avd)
+{
+ int ret = 0;
+
+ ret = pm_runtime_resume_and_get(avd->dev);
+ if (ret < 0)
+ return ret;
+
+ ret = reset_control_reset(avd->rstc);
+ if (ret)
+ dev_err(avd->dev, "reset: failed: %d", ret);
+
+ iommu_attach_device(avd->empty_domain, avd->dev);
+ iommu_detach_device(avd->empty_domain, avd->dev);
+
+ ret = avd_boot(avd);
+ if (ret)
+ dev_err(avd->dev, "reset: failed to boot");
+
+ pm_runtime_put_autosuspend(avd->dev);
+
+ return ret;
+}
+
+static void avd_watchdog_func(struct work_struct *work)
+{
+ struct avd_dev *avd;
+ struct avd_ctx *ctx;
+ int ret;
+
+ ctx = container_of(to_delayed_work(work), struct avd_ctx,
+ watchdog_work);
+ if (!ctx)
+ return;
+
+ avd = ctx->dev;
+
+ dev_err(avd->dev, "Frame processing timed out!");
+
+ writel(0, avd->mbox + AVD_REG_MBOX_IRQ_ENABLE);
+ ret = avd_reset(avd);
+ if (ret)
+ dev_err(avd->dev, "failed to reset: %d", ret);
+
+ avd_job_finish(ctx, VB2_BUF_STATE_ERROR);
+}
+
+static irqreturn_t avd_irq_handler(int irq, void *data)
+{
+ struct avd_dev *avd = data;
+ struct avd_ctx *ctx = v4l2_m2m_get_curr_priv(avd->m2m_dev);
+ enum vb2_buffer_state state;
+ u32 status;
+
+ status = readl(avd->mbox + AVD_REG_MBOX0_RETRIEVE);
+ writel(AVD_MBOX0_NOT_EMPTY, avd->mbox + AVD_REG_MBOX_IRQ_CLR);
+
+ if (status & 0x10000) { /* dbg */
+ dev_warn(avd->dev, "no handler for IRQ: %3d",
+ status & ~0x10000);
+ writel_relaxed(0, avd->mbox + AVD_REG_MBOX_IRQ_ENABLE);
+ return IRQ_HANDLED;
+ }
+
+ if (!ctx)
+ return IRQ_HANDLED;
+
+ if (status & 0x1000) {
+ state = VB2_BUF_STATE_DONE;
+ } else {
+ dev_err(avd->dev, "error: fw says: %x", status);
+ /* let watchdog handle */
+ goto done;
+ }
+
+ /* if the watchdog_work has run the work has already been submitted */
+ if (cancel_delayed_work(&ctx->watchdog_work))
+ avd_job_finish(ctx, state);
+
+done:
+ return IRQ_HANDLED;
+}
+
+static void avd_device_run(void *priv)
+{
+ struct avd_ctx *ctx = priv;
+ struct avd_dev *avd = ctx->dev;
+ const struct avd_coded_fmt_desc *desc = ctx->coded_fmt_desc;
+ int ret;
+
+ if (WARN_ON(!desc))
+ return;
+
+ ret = pm_runtime_resume_and_get(avd->dev);
+ if (ret < 0) {
+ avd_job_finish_no_pm(ctx, VB2_BUF_STATE_ERROR);
+ return;
+ }
+
+ ret = desc->ops->run(ctx);
+ if (ret)
+ avd_job_finish(ctx, VB2_BUF_STATE_ERROR);
+}
+
+static int avd_queue_init(void *priv, struct vb2_queue *src_vq,
+ struct vb2_queue *dst_vq)
+{
+ struct avd_ctx *ctx = priv;
+ int ret;
+
+ src_vq->type = V4L2_BUF_TYPE_VIDEO_OUTPUT_MPLANE;
+ src_vq->io_modes = VB2_MMAP | VB2_DMABUF;
+ src_vq->drv_priv = ctx;
+ src_vq->ops = &avd_queue_ops;
+ src_vq->mem_ops = &vb2_dma_contig_memops;
+
+ src_vq->dma_attrs = 0;
+ src_vq->buf_struct_size = sizeof(struct v4l2_m2m_buffer);
+ src_vq->timestamp_flags = V4L2_BUF_FLAG_TIMESTAMP_COPY;
+ src_vq->lock = &ctx->dev->vdev_lock;
+ src_vq->dev = ctx->dev->v4l2_dev.dev;
+ src_vq->supports_requests = true;
+
+ ret = vb2_queue_init(src_vq);
+ if (ret)
+ return ret;
+
+ dst_vq->bidirectional = true;
+ dst_vq->mem_ops = &vb2_dma_contig_memops;
+ dst_vq->dma_attrs = 0;
+ dst_vq->type = V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE;
+ dst_vq->io_modes = VB2_MMAP | VB2_DMABUF;
+ dst_vq->drv_priv = ctx;
+ dst_vq->ops = &avd_queue_ops;
+ dst_vq->buf_struct_size = sizeof(struct avd_decoded_buffer);
+ dst_vq->timestamp_flags = V4L2_BUF_FLAG_TIMESTAMP_COPY;
+ dst_vq->lock = &ctx->dev->vdev_lock;
+ dst_vq->dev = ctx->dev->v4l2_dev.dev;
+
+ return vb2_queue_init(dst_vq);
+}
+
+static int avd_open(struct file *filp)
+{
+ struct avd_dev *avd = video_drvdata(filp);
+ struct avd_ctx *ctx;
+ int ret;
+
+ ctx = kzalloc_obj(*ctx, GFP_KERNEL);
+ if (!ctx)
+ return -ENOMEM;
+
+ ctx->dev = avd;
+
+ ret = avd_buf_alloc(avd, &ctx->inst, AVD_FIFO_SIZE);
+ if (ret)
+ goto err_free_ctx;
+
+ ret = avd_buf_alloc(avd, &ctx->pipe_state, 512);
+ if (ret)
+ goto err_free_ctx;
+
+ INIT_DELAYED_WORK(&ctx->watchdog_work, avd_watchdog_func);
+
+ avd_reset_coded_fmt(ctx);
+ avd_reset_decoded_fmt(ctx);
+
+ v4l2_fh_init(&ctx->fh, video_devdata(filp));
+
+ ctx->fh.m2m_ctx = v4l2_m2m_ctx_init(avd->m2m_dev, ctx, avd_queue_init);
+ if (IS_ERR(ctx->fh.m2m_ctx)) {
+ ret = PTR_ERR(ctx->fh.m2m_ctx);
+ goto err_free_ctx;
+ }
+
+ ret = avd_init_ctrls(ctx);
+ if (ret)
+ goto err_cleanup_m2m_ctx;
+
+ v4l2_fh_add(&ctx->fh, filp);
+
+ return 0;
+
+err_cleanup_m2m_ctx:
+ v4l2_m2m_ctx_release(ctx->fh.m2m_ctx);
+
+err_free_ctx:
+ avd_buf_free(avd, &ctx->pipe_state);
+ avd_buf_free(avd, &ctx->inst);
+ kfree(ctx);
+ return ret;
+}
+
+static int avd_release(struct file *filp)
+{
+ struct avd_ctx *ctx = file_to_ctx(filp);
+
+ cancel_delayed_work(&ctx->watchdog_work);
+
+ v4l2_fh_del(&ctx->fh, filp);
+ v4l2_m2m_ctx_release(ctx->fh.m2m_ctx);
+ v4l2_ctrl_handler_free(&ctx->ctrl_hdl);
+ v4l2_fh_exit(&ctx->fh);
+ avd_buf_free(ctx->dev, &ctx->inst);
+ avd_buf_free(ctx->dev, &ctx->pipe_state);
+ avd_buf_free(ctx->dev, &ctx->job.buf);
+ kfree(ctx);
+
+ return 0;
+}
+
+static const struct v4l2_file_operations avd_fops = {
+ .owner = THIS_MODULE,
+ .open = avd_open,
+ .release = avd_release,
+ .poll = v4l2_m2m_fop_poll,
+ .unlocked_ioctl = video_ioctl2,
+ .mmap = v4l2_m2m_fop_mmap,
+};
+
+static const struct v4l2_m2m_ops avd_m2m_ops = {
+ .device_run = avd_device_run,
+};
+
+static const struct media_device_ops avd_media_ops = {
+ .req_validate = vb2_request_validate,
+ .req_queue = v4l2_m2m_request_queue,
+};
+
+static int avd_v4l2_init(struct avd_dev *avd)
+{
+ int ret;
+
+ ret = v4l2_device_register(avd->dev, &avd->v4l2_dev);
+ if (ret) {
+ dev_err(avd->dev, "Failed to register V4L2 device\n");
+ return ret;
+ }
+
+ avd->m2m_dev = v4l2_m2m_init(&avd_m2m_ops);
+ if (IS_ERR(avd->m2m_dev)) {
+ v4l2_err(&avd->v4l2_dev, "Failed to init mem2mem device\n");
+ ret = PTR_ERR(avd->m2m_dev);
+ goto err_unregister_v4l2;
+ }
+
+ avd->mdev.dev = avd->dev;
+ strscpy(avd->mdev.model, "avd", sizeof(avd->mdev.model));
+ strscpy(avd->mdev.bus_info, "platform:avd", sizeof(avd->mdev.bus_info));
+ media_device_init(&avd->mdev);
+ avd->mdev.ops = &avd_media_ops;
+ avd->v4l2_dev.mdev = &avd->mdev;
+
+ avd->vdev.lock = &avd->vdev_lock;
+ avd->vdev.v4l2_dev = &avd->v4l2_dev;
+ avd->vdev.fops = &avd_fops;
+ avd->vdev.release = video_device_release_empty;
+ avd->vdev.vfl_dir = VFL_DIR_M2M;
+ avd->vdev.device_caps = V4L2_CAP_STREAMING | V4L2_CAP_VIDEO_M2M_MPLANE;
+ avd->vdev.ioctl_ops = &avd_ioctl_ops;
+ video_set_drvdata(&avd->vdev, avd);
+ strscpy(avd->vdev.name, "avd", sizeof(avd->vdev.name));
+
+ ret = video_register_device(&avd->vdev, VFL_TYPE_VIDEO, -1);
+ if (ret) {
+ v4l2_err(&avd->v4l2_dev, "Failed to register video device\n");
+ goto err_cleanup_mc;
+ }
+
+ ret = v4l2_m2m_register_media_controller(avd->m2m_dev, &avd->vdev,
+ MEDIA_ENT_F_PROC_VIDEO_DECODER);
+ if (ret) {
+ v4l2_err(&avd->v4l2_dev,
+ "Failed to initialize V4L2 M2M media controller\n");
+ goto err_unregister_vdev;
+ }
+
+ ret = media_device_register(&avd->mdev);
+ if (ret) {
+ v4l2_err(&avd->v4l2_dev, "Failed to register media device\n");
+ goto err_unregister_mc;
+ }
+
+ return 0;
+
+err_unregister_mc:
+ v4l2_m2m_unregister_media_controller(avd->m2m_dev);
+
+err_unregister_vdev:
+ video_unregister_device(&avd->vdev);
+
+err_cleanup_mc:
+ media_device_cleanup(&avd->mdev);
+ v4l2_m2m_release(avd->m2m_dev);
+
+err_unregister_v4l2:
+ v4l2_device_unregister(&avd->v4l2_dev);
+ return ret;
+}
+
+static void avd_v4l2_cleanup(struct avd_dev *avd)
+{
+ media_device_unregister(&avd->mdev);
+ v4l2_m2m_unregister_media_controller(avd->m2m_dev);
+ video_unregister_device(&avd->vdev);
+ media_device_cleanup(&avd->mdev);
+ v4l2_m2m_release(avd->m2m_dev);
+ v4l2_device_unregister(&avd->v4l2_dev);
+}
+
+static const struct avd_variant avd_t8103_variant = {
+ .capabilities = AVD_CAPABILITY_HEVC |
+ AVD_CAPABILITY_H264 |
+ AVD_CAPABILITY_VP9,
+ .fw_name = "apple/avd-fw-v2-t0.bin",
+ .revision = 3,
+ .quirks = AVD_QUIRK_LSR | AVD_QUIRK_NO_PIPE_STATE,
+};
+
+static const struct avd_variant avd_t6000_variant = {
+ .capabilities = AVD_CAPABILITY_HEVC |
+ AVD_CAPABILITY_H264 |
+ AVD_CAPABILITY_VP9,
+ .fw_name = "apple/avd-fw-v3-t0.bin",
+ .revision = 4,
+ .quirks = AVD_QUIRK_LSR | AVD_QUIRK_NO_PIPE_STATE,
+};
+
+static const struct avd_variant avd_t8112_variant = {
+ .capabilities = AVD_CAPABILITY_HEVC |
+ AVD_CAPABILITY_H264 |
+ AVD_CAPABILITY_VP9,
+ .fw_name = "apple/avd-fw-v3-t1.bin",
+ .revision = 4,
+ .quirks = AVD_QUIRK_LSR,
+};
+
+static const struct avd_variant avd_t6020_variant = {
+ .capabilities = AVD_CAPABILITY_HEVC |
+ AVD_CAPABILITY_H264 |
+ AVD_CAPABILITY_VP9,
+ .fw_name = "apple/avd-fw-v3-t2.bin",
+ .revision = 4,
+};
+
+static const struct avd_variant avd_t8122_variant = {
+ .capabilities = AVD_CAPABILITY_HEVC |
+ AVD_CAPABILITY_H264 |
+ AVD_CAPABILITY_VP9 |
+ AVD_CAPABILITY_AV1,
+ .fw_name = "apple/avd-fw-v4-t0.bin",
+ .revision = 4,
+};
+
+static const struct avd_variant avd_t8140_variant = {
+ .capabilities = AVD_CAPABILITY_HEVC |
+ AVD_CAPABILITY_H264 |
+ AVD_CAPABILITY_VP9 |
+ AVD_CAPABILITY_AV1,
+ .fw_name = "apple/avd-fw-v5-t0.bin",
+ .revision = 4,
+};
+
+static const struct avd_variant avd_t8132_variant = {
+ .capabilities = AVD_CAPABILITY_HEVC |
+ AVD_CAPABILITY_H264 |
+ AVD_CAPABILITY_VP9 |
+ AVD_CAPABILITY_AV1,
+ .fw_name = "apple/avd-fw-v5-t1.bin",
+ .revision = 4,
+};
+
+/* can also be derived from a version register */
+static const struct of_device_id avd_of_match[] = {
+ { .compatible = "apple,t8103-avd", .data = &avd_t8103_variant },
+ { .compatible = "apple,t6000-avd", .data = &avd_t6000_variant },
+ { .compatible = "apple,t8112-avd", .data = &avd_t8112_variant },
+ { .compatible = "apple,t6020-avd", .data = &avd_t6020_variant },
+ { .compatible = "apple,t8122-avd", .data = &avd_t8122_variant },
+ { .compatible = "apple,t8132-avd", .data = &avd_t8132_variant },
+ { .compatible = "apple,t8140-avd", .data = &avd_t8140_variant },
+ {},
+};
+
+MODULE_DEVICE_TABLE(of, avd_of_match);
+
+static int avd_probe(struct platform_device *pdev)
+{
+ struct avd_dev *avd;
+ int ret, irq;
+
+ avd = devm_kzalloc(&pdev->dev, sizeof(*avd), GFP_KERNEL);
+ if (!avd)
+ return -ENOMEM;
+
+ platform_set_drvdata(pdev, avd);
+ avd->dev = &pdev->dev;
+ avd->pdev = pdev;
+
+ mutex_init(&avd->vdev_lock);
+
+ avd->variant = of_device_get_match_data(&pdev->dev);
+
+ avd->rstc = devm_reset_control_get_exclusive(avd->dev, NULL);
+
+ avd->piodma = devm_platform_ioremap_resource_byname(pdev, "piodma");
+ if (IS_ERR(avd->piodma))
+ return PTR_ERR(avd->piodma);
+
+ avd->code = devm_platform_ioremap_resource_byname(pdev, "code");
+ if (IS_ERR(avd->code))
+ return PTR_ERR(avd->code);
+
+ avd->sram = devm_platform_ioremap_resource_byname(pdev, "sram");
+ if (IS_ERR(avd->sram))
+ return PTR_ERR(avd->sram);
+
+ avd->mbox = devm_platform_ioremap_resource_byname(pdev, "mbox");
+ if (IS_ERR(avd->mbox))
+ return PTR_ERR(avd->mbox);
+
+ avd->ctrl = devm_platform_ioremap_resource_byname(pdev, "ctrl");
+ if (IS_ERR(avd->ctrl))
+ return PTR_ERR(avd->ctrl);
+
+ avd->sram_start = platform_get_resource_byname(pdev, IORESOURCE_MEM,
+ "sram")->start >> 4;
+
+ ret = dma_set_mask_and_coherent(avd->dev,
+ DMA_BIT_MASK((avd->variant->quirks &
+ AVD_QUIRK_LSR) ? 38 : 64));
+ if (ret) {
+ dev_err(avd->dev, "Failed to set DMA mask");
+ return ret;
+ }
+
+ irq = platform_get_irq_byname(pdev, "mbox0");
+ if (irq < 0)
+ return irq;
+ ret = devm_request_threaded_irq(&pdev->dev, irq, NULL, avd_irq_handler,
+ IRQF_ONESHOT, dev_name(&pdev->dev),
+ avd);
+ if (ret) {
+ dev_err(avd->dev, "Could not request IRQ 0");
+ return ret;
+ }
+
+ avd->domain = iommu_get_domain_for_dev(avd->dev);
+ if (!avd->domain)
+ return -EINVAL;
+
+ avd->empty_domain = iommu_paging_domain_alloc(avd->dev);
+ if (IS_ERR(avd->empty_domain)) {
+ dev_err(avd->dev, "cannot alloc new empty domain");
+ return PTR_ERR(avd->empty_domain);
+ }
+
+ ret = request_firmware(&avd->fw, avd->variant->fw_name, avd->dev);
+ if (ret) {
+ dev_err(avd->dev, "failed to load firmware: %d", ret);
+ iommu_domain_free(avd->empty_domain);
+ return ret;
+ }
+
+ pm_runtime_set_autosuspend_delay(avd->dev, 100);
+ pm_runtime_use_autosuspend(avd->dev);
+ pm_runtime_enable(avd->dev);
+
+ ret = avd_v4l2_init(avd);
+ if (ret)
+ goto err_disable_runtime_pm;
+
+ return 0;
+
+err_disable_runtime_pm:
+ pm_runtime_dont_use_autosuspend(&pdev->dev);
+ pm_runtime_disable(&pdev->dev);
+ release_firmware(avd->fw);
+ iommu_domain_free(avd->empty_domain);
+ return ret;
+}
+
+static void avd_remove(struct platform_device *pdev)
+{
+ struct avd_dev *avd = platform_get_drvdata(pdev);
+
+ avd_v4l2_cleanup(avd);
+
+ iommu_domain_free(avd->empty_domain);
+
+ release_firmware(avd->fw);
+
+ pm_runtime_disable(avd->dev);
+ pm_runtime_dont_use_autosuspend(avd->dev);
+}
+
+static __maybe_unused int avd_runtime_resume(struct device *dev)
+{
+ int ret;
+ struct avd_dev *avd = platform_get_drvdata(to_platform_device(dev));
+
+ ret = avd_boot(avd);
+ if (ret)
+ dev_err(dev, "failed to boot");
+ return ret;
+}
+
+static __maybe_unused int avd_runtime_suspend(struct device *dev)
+{
+ struct avd_dev *avd = platform_get_drvdata(to_platform_device(dev));
+
+ avd_shutdown(avd);
+ return 0;
+}
+
+static const struct dev_pm_ops avd_pm_ops = {
+ SET_SYSTEM_SLEEP_PM_OPS(pm_runtime_force_suspend,
+ pm_runtime_force_resume)
+ SET_RUNTIME_PM_OPS(avd_runtime_suspend, avd_runtime_resume, NULL)
+};
+
+static struct platform_driver avd_driver = {
+ .probe = avd_probe,
+ .remove = avd_remove,
+ .driver = {
+ .name = "avd",
+ .of_match_table = avd_of_match,
+ .pm = pm_ptr(&avd_pm_ops),
+ },
+};
+module_platform_driver(avd_driver);
+
+MODULE_LICENSE("GPL");
+MODULE_DESCRIPTION("Apple Video Decoder driver");
+MODULE_FIRMWARE("apple/avd-fw-v2-t0.bin");
+MODULE_FIRMWARE("apple/avd-fw-v3-t0.bin");
+MODULE_FIRMWARE("apple/avd-fw-v3-t1.bin");
+MODULE_FIRMWARE("apple/avd-fw-v3-t2.bin");
+MODULE_FIRMWARE("apple/avd-fw-v4-t0.bin");
+MODULE_FIRMWARE("apple/avd-fw-v5-t0.bin");
+MODULE_FIRMWARE("apple/avd-fw-v5-t1.bin");
diff --git a/drivers/media/platform/apple/avd/avd-inst.h b/drivers/media/platform/apple/avd/avd-inst.h
new file mode 100644
index 000000000000..d44d1f7f66c2
--- /dev/null
+++ b/drivers/media/platform/apple/avd/avd-inst.h
@@ -0,0 +1,210 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Apple Video Decoder instruction stream definitions
+ *
+ * Copyright (C) 2026 The Asahi Linux Contributors
+ * Copyright (C) 2026 Sofus Forstreuter <sofus.c@icloud.com>
+ * Copyright (C) 2023 Eileen Yoon <eyn@gmx.com>
+ *
+ * The AVD block consists of one to four processors for each codecs
+ * called "VP". Its the VP's job to prepare codecs specifik syntax into
+ * something the shared processor(s?) (called PP, Pipeline?) understands. PP
+ * implements different DSP routines: QT, IPMC, LF, MV, TP, PC, SW, LR.
+ *
+ * When VP processes a u32, if the value has the two upper bits set and unset
+ * respectively, it will execute the opcode function stored in the next 8
+ * bits.
+ *
+ * Please refer to
+ * https://web.archive.org/web/20240105073842/https://eiln.net/avd-notes2.html
+ * for an excellent write-up by Eileen.
+ */
+
+#ifndef AVD_INST_H_
+#define AVD_INST_H_
+
+#include <linux/types.h>
+#include <linux/bitfield.h>
+
+#include "avd.h"
+
+#define AVD_OP_HDR FIELD_PREP(GENMASK(31, 20), 0x2db)
+#define AVD_OP_HDR_CONST FIELD_PREP(GENMASK(10, 0), 0x2e0)
+
+/*
+ * output in packed PX30
+ * 10 bits per component. With groups of tree packed into 4 bytes (little
+ * endian order)
+ * P 3 2 1
+ * [2:10:10:10]
+ *
+ * additionally, av1 and vp9 need an extra scratch buffer if this is set
+ */
+#define AVD_OP_HDR_FLAG_PACKED(v) FIELD_PREP(BIT(10), !!(v))
+/* decompress pixel data */
+#define AVD_OP_HDR_FLAG_DECOMP(v) FIELD_PREP(BIT(12), !!(v))
+#define AVD_OP_HDR_FLAG_INTRA(v) FIELD_PREP(BIT(13), !!(v))
+#define AVD_OP_HDR_FLAG_PIPE_STATE(v) FIELD_PREP(BIT(19), !!(v))
+
+#define AVD_OP_WEIGHTS_HDR FIELD_PREP(GENMASK(31, 20), 0x2dd)
+#define AVD_OP_WEIGHTS_HDR_CHROMA(v) FIELD_PREP(GENMASK(2, 0), v)
+#define AVD_OP_WEIGHTS_HDR_LUMA(v) FIELD_PREP(GENMASK(5, 3), v)
+#define AVD_OP_WEIGHTS_HDR_FLAG0(v) FIELD_PREP(BIT(6), !!(v))
+#define AVD_OP_WEIGHTS_HDR_FLAG1(v) FIELD_PREP(BIT(7), !!(v))
+
+#define AVD_OP_WEIGHTS FIELD_PREP(GENMASK(31, 20), 0x2de)
+#define AVD_OP_WEIGHTS_WEIGHT(v) FIELD_PREP(GENMASK(8, 0), v)
+#define AVD_OP_WEIGHTS_INDEX(v) FIELD_PREP(GENMASK(12, 9), v)
+#define AVD_OP_WEIGHTS_LIST_IDX(v) FIELD_PREP(BIT(13), v)
+/* 1 = luma, 2,3 = chroma[{0,1}] */
+#define AVD_OP_WEIGHTS_IDENT(v) FIELD_PREP(GENMASK(16, 14), v)
+
+#define AVD_OP_OFFSETS FIELD_PREP(GENMASK(31, 20), 0x2df)
+#define AVD_OP_OFFSETS_OFFSET(v) FIELD_PREP(GENMASK(15, 0), v)
+
+#define AVD_OP_CODED_DATA FIELD_PREP(GENMASK(31, 20), 0x2d8)
+#define AVD_OP_CODED_IN_HI(v) FIELD_PREP(GENMASK(12, 0), (u64)(v) >> 32)
+#define AVD_OP_CODED_IN_LO(v) ((u32)(v) & 0xffffffff)
+#define AVD_OP_CODED_DATA_FLAG0(v) FIELD_PREP(BIT(13), !!(v))
+#define AVD_OP_CODED_DATA_FLAG1(v) FIELD_PREP(BIT(14), !!(v))
+#define AVD_OP_CODED_DATA_BIT_OFF(v) FIELD_PREP(GENMASK(18, 15), v)
+
+#define AVD_OP_SL_LOC FIELD_PREP(GENMASK(31, 24), 0x2c)
+#define AVD_OP_SL_LOC_X(v) FIELD_PREP(GENMASK(11, 0), v)
+#define AVD_OP_SL_LOC_Y(v) FIELD_PREP(GENMASK(23, 12), v)
+
+#define AVD_OP_SL_DIM_START FIELD_PREP(GENMASK(31, 24), 0x2a)
+#define AVD_OP_SL_DIM_START_X(v) FIELD_PREP(GENMASK(11, 0), v)
+#define AVD_OP_SL_DIM_START_Y(v) FIELD_PREP(GENMASK(23, 12), v)
+
+/* end is not an op */
+#define AVD_SL_DIM_END_X(v) FIELD_PREP(GENMASK(11, 0), v)
+#define AVD_SL_DIM_END_Y(v) FIELD_PREP(GENMASK(23, 12), v)
+#define AVD_SL_DIM_END_COL(v) FIELD_PREP(GENMASK(27, 24), v)
+#define AVD_SL_DIM_END_ROW(v) FIELD_PREP(GENMASK(31, 28), v)
+
+#define AVD_OP_SL_REF FIELD_PREP(GENMASK(31, 24), 0x2d)
+#define AVD_OP_SL_REF_MAX_MERGE(v) FIELD_PREP(GENMASK(3, 1), v)
+
+#define AVD_OP_SL_REF_FLAG0(v) FIELD_PREP(BIT(4), !!(v))
+#define AVD_OP_SL_REF_FLAG_CABAC(v) FIELD_PREP(BIT(5), !!(v))
+#define AVD_OP_SL_REF_FLAG1(v) FIELD_PREP(BIT(6), !!(v))
+#define AVD_OP_SL_REF_NUM_L0(v) FIELD_PREP(GENMASK(15, 11), v)
+#define AVD_OP_SL_REF_NUM_L1(v) FIELD_PREP(GENMASK(10, 7), v)
+#define AVD_OP_SL_REF_FLAG2(v) FIELD_PREP(BIT(15), !!(v))
+#define AVD_OP_SL_REF_SLICE_P(v) FIELD_PREP(BIT(16), !!(v))
+#define AVD_OP_SL_REF_SLICE_I(v) FIELD_PREP(BIT(17), !!(v))
+/* not really kinda more like has_ref_and_ref_is_valid_ref */
+#define AVD_OP_SL_REF_SLICE_B(v) FIELD_PREP(BIT(18), !!(v))
+
+#define AVD_OP_QP FIELD_PREP(GENMASK(31, 20), 0x2d9)
+#define AVD_OP_QP_CR_OFF(v) FIELD_PREP(GENMASK(4, 0), v)
+#define AVD_OP_QP_CB_OFF(v) FIELD_PREP(GENMASK(9, 5), v)
+#define AVD_OP_QP_VAL(v) FIELD_PREP(GENMASK(17, 10), v)
+
+#define AVD_OP_DBLK FIELD_PREP(GENMASK(31, 20), 0x2da)
+#define AVD_OP_DBLK_FLAG_SAO_CHROMA(v) FIELD_PREP(BIT(6), !!(v))
+#define AVD_OP_DBLK_FLAG_SAO_LUMA(v) FIELD_PREP(BIT(7), !!(v))
+#define AVD_OP_DBLK_OFF0(v) FIELD_PREP(GENMASK(11, 8), v)
+#define AVD_OP_DBLK_OFF1(v) FIELD_PREP(GENMASK(16, 12), v)
+#define AVD_OP_DBLK_FLAG_EN(v) FIELD_PREP(BIT(16), !!(v))
+#define AVD_OP_DBLK_FLAG_FULL_EN(v) FIELD_PREP(BIT(17), !!(v))
+#define AVD_OP_DBLK_FLAG_TILES_EN(v) FIELD_PREP(BIT(18), !!(v))
+#define AVD_OP_DBLK_FLAG_PCM_EN(v) FIELD_PREP(BIT(19), !!(v))
+
+#define AVD_OP_REF FIELD_PREP(GENMASK(31, 20), 0x2dc)
+/* same order as they where submitted */
+#define AVD_OP_REF_DBP_IDX(v) FIELD_PREP(GENMASK(3, 0), v)
+#define AVD_OP_REF_LOOP_IDX(v) FIELD_PREP(GENMASK(7, 4), v)
+#define AVD_OP_REF_LIST_IDX(v) FIELD_PREP(GENMASK(11, 8), v)
+
+#define AVD_HDR_CODEC_MODE(v) FIELD_PREP(GENMASK(28, 24), v)
+#define AVD_HDR_WIDTH(v) FIELD_PREP(GENMASK(15, 0), v)
+#define AVD_HDR_HEIGHT(v) FIELD_PREP(GENMASK(31, 16), v)
+
+#define AVD_HDR_FEAT_H264 FIELD_PREP(GENMASK(3, 0), 10)
+#define AVD_HDR_FEAT_PIPE_STATE_EN(v) FIELD_PREP(GENMASK(7, 4), (v) ? 3 : 0)
+#define AVD_HDR_FEAT_H26X FIELD_PREP(BIT(20), 1)
+#define AVD_HDR_FEAT_COMMON FIELD_PREP(BIT(21), 1)
+#define AVD_HDR_FEAT_VP9 FIELD_PREP(BIT(17), 1)
+
+#define AVD_HDR_COMMON_FLAG0(v) FIELD_PREP(BIT(0), !!(v))
+#define AVD_HDR_COMMON_LUMA_TBS(v) FIELD_PREP(GENMASK(8, 7), v)
+#define AVD_HDR_COMMON_MIN_LUMA_TBS(v) FIELD_PREP(GENMASK(10, 9), v)
+#define AVD_HDR_COMMON_LUMA_CBS(v) FIELD_PREP(GENMASK(12, 11), v)
+#define AVD_HDR_COMMON_MIN_LUMA_CBS(v) FIELD_PREP(GENMASK(14, 13), v)
+#define AVD_HDR_COMMON_BIT_DEPTH_L(v) FIELD_PREP(GENMASK(18, 15), v)
+#define AVD_HDR_COMMON_BIT_DEPTH_C(v) FIELD_PREP(GENMASK(23, 19), v)
+#define AVD_HDR_COMMON_CHROMA_FORMAT(v) FIELD_PREP(GENMASK(26, 24), v)
+
+#define AVD_HDR_H26X_QP_OFFSET_CR(v) FIELD_PREP(GENMASK(4, 0), v)
+#define AVD_HDR_H26X_QP_OFFSET_CB(v) FIELD_PREP(GENMASK(9, 5), v)
+
+#define AVD_SCALING_I0(v) FIELD_PREP(GENMASK(7, 0), v)
+#define AVD_SCALING_I1(v) FIELD_PREP(GENMASK(15, 8), v)
+#define AVD_SCALING_I2(v) FIELD_PREP(GENMASK(23, 16), v)
+#define AVD_SCALING_I3(v) FIELD_PREP(GENMASK(31, 24), v)
+
+#define AVD_REF_NUM(v) FIELD_PREP(GENMASK(31, 28), v)
+#define AVD_REF_FLAG_CONST FIELD_PREP(BIT(24), 1)
+#define AVD_REF_FLAG_LONG(v) FIELD_PREP(BIT(17), !!(v))
+#define AVD_REF_DELTA_POC(v) FIELD_PREP(GENMASK(16, 0), v)
+
+#define AVD_FIFO_SIZE (0x100000 * 12)
+
+static inline void push(struct avd_ctx *ctx, u32 inst)
+{
+ struct avd_job *job = &ctx->job;
+ struct avd_segment *seg = &job->segments[job->num];
+
+ seg->instructions[seg->num++] = inst;
+}
+
+static inline void push_address(struct avd_ctx *ctx, dma_addr_t addr)
+{
+ if (ctx->dev->variant->quirks & AVD_QUIRK_LSR) {
+ push(ctx, addr >> 8);
+ } else {
+ push(ctx, (u32)(addr & 0xffffffff));
+ push(ctx, (u32)((u64)addr >> 32));
+ }
+}
+
+static inline void push_comp(struct avd_ctx *ctx, dma_addr_t addr,
+ u32 offsets[4])
+{
+ if (ctx->dev->variant->quirks & AVD_QUIRK_LSR) {
+ for (int i = 0; i < 4; i++)
+ push(ctx, (addr + offsets[i]) >> 7);
+ } else {
+ for (int i = 0; i < 4; i++)
+ push_address(ctx, (addr + offsets[i]));
+ }
+}
+
+#ifdef DEBUG_INST
+#define push(inst, name) \
+ do { \
+ dev_info(ctx->dev->dev, "%8x | %s", (inst), name); \
+ push(ctx, inst); \
+ } while (0)
+
+#else
+#define push(inst, name) push(ctx, inst)
+#endif
+
+#ifdef DEBUG_INST_ADDR
+#define pusha(inst, name, i) \
+ do { \
+ dev_info(ctx->dev->dev, "%8llx | %s[%d]", (inst) & 0xffffffff, \
+ name, i); \
+ dev_info(ctx->dev->dev, "%8llx | %s[%d] (high)", (inst) >> 32, \
+ name, i); \
+ push_address(ctx, inst); \
+ } while (0)
+
+#else
+#define pusha(inst, name, i) push_address(ctx, inst)
+#endif
+
+#endif /* AVD_INST_H_ */
diff --git a/drivers/media/platform/apple/avd/avd-v4l2.c b/drivers/media/platform/apple/avd/avd-v4l2.c
new file mode 100644
index 000000000000..88079da5ee27
--- /dev/null
+++ b/drivers/media/platform/apple/avd/avd-v4l2.c
@@ -0,0 +1,780 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Apple Video Decoder V4L2 M2M api functions
+ *
+ * Copyright (C) 2026 The Asahi Linux Contributors
+ * Copyright (C) 2026 Sofus Forstreuter <sofus.c@icloud.com>
+ *
+ * Based on rkvdec driver by Collabora, Ltd.
+ * Copyright (C) 2019 Collabora, Ltd.
+ * Based on rkvdec driver by Google LLC. (Tomasz Figa <tfiga@chromium.org>)
+ * Based on s5p-mfc driver by Samsung Electronics Co., Ltd.
+ * Copyright (C) 2011 Samsung Electronics Co., Ltd.
+ */
+
+#include <linux/pm_runtime.h>
+
+#include <media/v4l2-event.h>
+#include <media/v4l2-mem2mem.h>
+#include <media/videobuf2-dma-contig.h>
+
+#include "avd.h"
+
+static const struct avd_decoded_fmt_desc avd_decoded_fmts[] = {
+ {
+ .fourcc = V4L2_PIX_FMT_NV12,
+ .image_fmt = AVD_IMG_FMT_420_8BIT,
+ },
+ {
+ .fourcc = V4L2_PIX_FMT_P010,
+ .image_fmt = AVD_IMG_FMT_420_10BIT,
+ },
+ {
+ .fourcc = V4L2_PIX_FMT_NV16,
+ .image_fmt = AVD_IMG_FMT_422_8BIT,
+ },
+ {
+ .fourcc = V4L2_PIX_FMT_P210,
+ .image_fmt = AVD_IMG_FMT_422_10BIT,
+ },
+ {
+ .fourcc = V4L2_PIX_FMT_IC12,
+ .image_fmt = AVD_IMG_FMT_420_8BIT,
+ },
+ {
+ .fourcc = V4L2_PIX_FMT_IC03,
+ .image_fmt = AVD_IMG_FMT_420_10BIT,
+ },
+ {
+ .fourcc = V4L2_PIX_FMT_IC16,
+ .image_fmt = AVD_IMG_FMT_422_8BIT,
+ },
+ {
+ .fourcc = V4L2_PIX_FMT_IC23,
+ .image_fmt = AVD_IMG_FMT_422_10BIT,
+ },
+};
+
+static bool is_interchange(u32 fourcc)
+{
+ return fourcc == V4L2_PIX_FMT_IC12 ||
+ fourcc == V4L2_PIX_FMT_IC16 ||
+ fourcc == V4L2_PIX_FMT_IC03 ||
+ fourcc == V4L2_PIX_FMT_IC23;
+}
+
+static bool avd_image_fmt_match(enum avd_image_fmt fmt1,
+ enum avd_image_fmt fmt2)
+{
+ return fmt1 == fmt2 || fmt2 == AVD_IMG_FMT_ANY ||
+ fmt1 == AVD_IMG_FMT_ANY;
+}
+
+static bool avd_image_fmt_changed(struct avd_ctx *ctx,
+ enum avd_image_fmt image_fmt)
+{
+ if (image_fmt == AVD_IMG_FMT_ANY)
+ return false;
+
+ return ctx->image_fmt != image_fmt;
+}
+
+static u32 avd_enum_decoded_fmt(struct avd_ctx *ctx, int index,
+ enum avd_image_fmt image_fmt)
+{
+ const struct avd_coded_fmt_desc *desc = ctx->coded_fmt_desc;
+ int fmt_idx = -1;
+ unsigned int i;
+
+ if (WARN_ON(!desc))
+ return 0;
+
+ for (i = 0; i < ARRAY_SIZE(avd_decoded_fmts); i++) {
+ if (!avd_image_fmt_match(avd_decoded_fmts[i].image_fmt,
+ image_fmt))
+ continue;
+ fmt_idx++;
+ if (index == fmt_idx)
+ return avd_decoded_fmts[i].fourcc;
+ }
+
+ return 0;
+}
+
+static bool avd_is_valid_fmt(struct avd_ctx *ctx, u32 fourcc,
+ enum avd_image_fmt image_fmt)
+{
+ unsigned int i;
+
+ for (i = 0; i < ARRAY_SIZE(avd_decoded_fmts); i++) {
+ if (avd_image_fmt_match(avd_decoded_fmts[i].image_fmt,
+ image_fmt) &&
+ avd_decoded_fmts[i].fourcc == fourcc)
+ return true;
+ }
+
+ return false;
+}
+
+static void avd_fill_decoded_pixfmt(struct avd_ctx *ctx,
+ struct v4l2_pix_format_mplane *pix_mp)
+{
+ v4l2_fill_pixfmt_mp(pix_mp, pix_mp->pixelformat, pix_mp->width,
+ pix_mp->height);
+
+ if (is_interchange(pix_mp->pixelformat))
+ pix_mp->plane_fmt[0].sizeimage = 0;
+ ctx->comp.start_offset = pix_mp->plane_fmt[0].sizeimage;
+
+ fill_comp(&ctx->comp, ctx->image_fmt, pix_mp->width, pix_mp->height);
+ pix_mp->plane_fmt[0].sizeimage += ctx->comp.size;
+
+ if (ctx->coded_fmt_desc->ops->adjust_decoded_fmt)
+ ctx->coded_fmt_desc->ops->adjust_decoded_fmt(ctx, pix_mp);
+}
+
+static void avd_reset_fmt(struct avd_ctx *ctx, struct v4l2_format *f,
+ u32 fourcc)
+{
+ memset(f, 0, sizeof(*f));
+ f->fmt.pix_mp.pixelformat = fourcc;
+ f->fmt.pix_mp.field = V4L2_FIELD_NONE;
+ f->fmt.pix_mp.colorspace = V4L2_COLORSPACE_REC709;
+ f->fmt.pix_mp.ycbcr_enc = V4L2_YCBCR_ENC_DEFAULT;
+ f->fmt.pix_mp.quantization = V4L2_QUANTIZATION_DEFAULT;
+ f->fmt.pix_mp.xfer_func = V4L2_XFER_FUNC_DEFAULT;
+}
+
+void avd_reset_decoded_fmt(struct avd_ctx *ctx)
+{
+ struct v4l2_format *f = &ctx->decoded_fmt;
+ u32 fourcc;
+
+ fourcc = avd_enum_decoded_fmt(ctx, 0, ctx->image_fmt);
+ avd_reset_fmt(ctx, f, fourcc);
+ f->type = V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE;
+ f->fmt.pix_mp.width = ctx->coded_fmt.fmt.pix_mp.width;
+ f->fmt.pix_mp.height = ctx->coded_fmt.fmt.pix_mp.height;
+ avd_fill_decoded_pixfmt(ctx, &f->fmt.pix_mp);
+}
+
+static int avd_try_ctrl(struct v4l2_ctrl *ctrl)
+{
+ struct avd_ctx *ctx =
+ container_of(ctrl->handler, struct avd_ctx, ctrl_hdl);
+ const struct avd_coded_fmt_desc *desc = ctx->coded_fmt_desc;
+
+ if (desc->ops->try_ctrl)
+ return desc->ops->try_ctrl(ctx, ctrl);
+
+ return 0;
+}
+
+static int avd_s_ctrl(struct v4l2_ctrl *ctrl)
+{
+ struct avd_ctx *ctx =
+ container_of(ctrl->handler, struct avd_ctx, ctrl_hdl);
+ const struct avd_coded_fmt_desc *desc = ctx->coded_fmt_desc;
+ enum avd_image_fmt image_fmt;
+ struct vb2_queue *vq;
+
+ /* Check if this change requires a capture format reset */
+ if (!desc->ops->get_image_fmt)
+ return 0;
+
+ image_fmt = desc->ops->get_image_fmt(ctx, ctrl);
+ if (avd_image_fmt_changed(ctx, image_fmt)) {
+ vq = v4l2_m2m_get_vq(ctx->fh.m2m_ctx,
+ V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE);
+ if (vb2_is_busy(vq))
+ return -EBUSY;
+
+ ctx->image_fmt = image_fmt;
+ avd_reset_decoded_fmt(ctx);
+ }
+
+ return 0;
+}
+
+const struct v4l2_ctrl_ops avd_ctrl_ops = {
+ .try_ctrl = avd_try_ctrl,
+ .s_ctrl = avd_s_ctrl,
+};
+
+static const struct avd_coded_fmt_desc avd_coded_fmts[] = {};
+
+static bool avd_is_capable(struct avd_ctx *ctx, unsigned int capability)
+{
+ return (ctx->dev->variant->capabilities & capability) == capability;
+}
+
+static const struct avd_coded_fmt_desc *
+avd_enum_coded_fmt_desc(struct avd_ctx *ctx, int index)
+{
+ int fmt_idx = -1;
+ unsigned int i;
+
+ for (i = 0; i < ARRAY_SIZE(avd_coded_fmts); i++) {
+ if (!avd_is_capable(ctx, avd_coded_fmts[i].capability))
+ continue;
+ fmt_idx++;
+ if (index == fmt_idx)
+ return &avd_coded_fmts[i];
+ }
+
+ return NULL;
+}
+
+static const struct avd_coded_fmt_desc *
+avd_find_coded_fmt_desc(struct avd_ctx *ctx, u32 fourcc)
+{
+ unsigned int i;
+
+ for (i = 0; i < ARRAY_SIZE(avd_coded_fmts); i++) {
+ if (avd_is_capable(ctx, avd_coded_fmts[i].capability) &&
+ avd_coded_fmts[i].fourcc == fourcc)
+ return &avd_coded_fmts[i];
+ }
+
+ return NULL;
+}
+
+void avd_reset_coded_fmt(struct avd_ctx *ctx)
+{
+ struct v4l2_format *f = &ctx->coded_fmt;
+
+ ctx->coded_fmt_desc = avd_enum_coded_fmt_desc(ctx, 0);
+ avd_reset_fmt(ctx, f, ctx->coded_fmt_desc->fourcc);
+
+ f->type = V4L2_BUF_TYPE_VIDEO_OUTPUT_MPLANE;
+ f->fmt.pix_mp.width = ctx->coded_fmt_desc->frmsize.min_width;
+ f->fmt.pix_mp.height = ctx->coded_fmt_desc->frmsize.min_height;
+
+ f->fmt.pix_mp.num_planes = 1;
+ if (!f->fmt.pix_mp.plane_fmt[0].sizeimage)
+ f->fmt.pix_mp.plane_fmt[0].sizeimage =
+ f->fmt.pix_mp.width * f->fmt.pix_mp.height;
+}
+
+static int avd_enum_framesizes(struct file *file, void *priv,
+ struct v4l2_frmsizeenum *fsize)
+{
+ struct avd_ctx *ctx = file_to_ctx(file);
+ const struct avd_coded_fmt_desc *desc;
+
+ if (fsize->index != 0)
+ return -EINVAL;
+
+ desc = avd_find_coded_fmt_desc(ctx, fsize->pixel_format);
+ if (!desc)
+ return -EINVAL;
+
+ fsize->type = V4L2_FRMSIZE_TYPE_CONTINUOUS;
+ fsize->stepwise.min_width = 1;
+ fsize->stepwise.max_width = desc->frmsize.max_width;
+ fsize->stepwise.step_width = 1;
+ fsize->stepwise.min_height = 1;
+ fsize->stepwise.max_height = desc->frmsize.max_height;
+ fsize->stepwise.step_height = 1;
+
+ return 0;
+}
+
+static int avd_querycap(struct file *file, void *priv,
+ struct v4l2_capability *cap)
+{
+ struct avd_dev *avd = video_drvdata(file);
+ struct video_device *vdev = video_devdata(file);
+
+ strscpy(cap->driver, avd->dev->driver->name, sizeof(cap->driver));
+ strscpy(cap->card, vdev->name, sizeof(cap->card));
+ snprintf(cap->bus_info, sizeof(cap->bus_info), "platform:%s",
+ avd->dev->driver->name);
+ return 0;
+}
+
+static int avd_try_capture_fmt(struct file *file, void *priv,
+ struct v4l2_format *f)
+{
+ struct v4l2_pix_format_mplane *pix_mp = &f->fmt.pix_mp;
+ struct avd_ctx *ctx = file_to_ctx(file);
+ const struct avd_coded_fmt_desc *coded_desc;
+
+ /*
+ * The codec context should point to a coded format desc, if the format
+ * on the coded end has not been set yet, it should point to the
+ * default value.
+ */
+ coded_desc = ctx->coded_fmt_desc;
+ if (WARN_ON(!coded_desc)) {
+ dev_err(ctx->dev->dev, "no coded desc!");
+ return -EINVAL;
+ }
+
+ if (!avd_is_valid_fmt(ctx, pix_mp->pixelformat, ctx->image_fmt))
+ pix_mp->pixelformat =
+ avd_enum_decoded_fmt(ctx, 0, ctx->image_fmt);
+
+ /* Always apply the frmsize constraint of the coded end. */
+ pix_mp->width = max(pix_mp->width, ctx->coded_fmt.fmt.pix_mp.width);
+ pix_mp->height = max(pix_mp->height, ctx->coded_fmt.fmt.pix_mp.height);
+ v4l2_apply_frmsize_constraints(&pix_mp->width, &pix_mp->height,
+ &coded_desc->frmsize);
+
+ avd_fill_decoded_pixfmt(ctx, pix_mp);
+ pix_mp->field = V4L2_FIELD_NONE;
+
+ return 0;
+}
+
+static int avd_try_output_fmt(struct file *file, void *priv,
+ struct v4l2_format *f)
+{
+ struct v4l2_pix_format_mplane *pix_mp = &f->fmt.pix_mp;
+ struct avd_ctx *ctx = file_to_ctx(file);
+ const struct avd_coded_fmt_desc *desc;
+
+ desc = avd_find_coded_fmt_desc(ctx, pix_mp->pixelformat);
+ if (!desc) {
+ desc = avd_enum_coded_fmt_desc(ctx, 0);
+ pix_mp->pixelformat = desc->fourcc;
+ }
+
+ v4l2_apply_frmsize_constraints(&pix_mp->width, &pix_mp->height,
+ &desc->frmsize);
+
+ pix_mp->field = V4L2_FIELD_NONE;
+ /* All coded formats are considered single planar for now. */
+ pix_mp->num_planes = 1;
+
+ if (!pix_mp->plane_fmt[0].sizeimage)
+ pix_mp->plane_fmt[0].sizeimage =
+ pix_mp->width * f->fmt.pix_mp.height;
+
+ return 0;
+}
+
+static int avd_s_capture_fmt(struct file *file, void *priv,
+ struct v4l2_format *f)
+{
+ struct avd_ctx *ctx = file_to_ctx(file);
+ struct vb2_queue *vq;
+ int ret;
+
+ /* Change not allowed if queue is busy */
+ vq = v4l2_m2m_get_vq(ctx->fh.m2m_ctx,
+ V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE);
+ if (vb2_is_busy(vq))
+ return -EBUSY;
+
+ ret = avd_try_capture_fmt(file, priv, f);
+ if (ret)
+ return ret;
+
+ ctx->decoded_fmt = *f;
+ return 0;
+}
+
+static int avd_s_output_fmt(struct file *file, void *priv,
+ struct v4l2_format *f)
+{
+ struct avd_ctx *ctx = file_to_ctx(file);
+ struct v4l2_m2m_ctx *m2m_ctx = ctx->fh.m2m_ctx;
+ const struct avd_coded_fmt_desc *desc;
+ struct v4l2_format *cap_fmt;
+ struct vb2_queue *peer_vq, *vq;
+ int ret;
+
+ /*
+ * In order to support dynamic resolution change, the decoder admits
+ * a resolution change, as long as the pixelformat remains. Can't be
+ * done if streaming.
+ */
+ vq = v4l2_m2m_get_vq(m2m_ctx, V4L2_BUF_TYPE_VIDEO_OUTPUT_MPLANE);
+ if (vb2_is_streaming(vq) ||
+ (vb2_is_busy(vq) && f->fmt.pix_mp.pixelformat !=
+ ctx->coded_fmt.fmt.pix_mp.pixelformat))
+ return -EBUSY;
+
+ /*
+ * Since format change on the OUTPUT queue will reset the CAPTURE
+ * queue, we can't allow doing so when the CAPTURE queue has buffers
+ * allocated.
+ */
+ peer_vq = v4l2_m2m_get_vq(m2m_ctx, V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE);
+ if (vb2_is_busy(peer_vq))
+ return -EBUSY;
+
+ ret = avd_try_output_fmt(file, priv, f);
+ if (ret)
+ return ret;
+
+ desc = avd_find_coded_fmt_desc(ctx, f->fmt.pix_mp.pixelformat);
+ if (!desc)
+ return -EINVAL;
+ ctx->coded_fmt_desc = desc;
+ ctx->coded_fmt = *f;
+
+ /*
+ * Current decoded format might have become invalid with newly
+ * selected codec, so reset it to default just to be safe and
+ * keep internal driver state sane. User is mandated to set
+ * the decoded format again after we return, so we don't need
+ * anything smarter.
+ *
+ * Note that this will propagates any size changes to the decoded
+ * format.
+ */
+ ctx->image_fmt = AVD_IMG_FMT_ANY;
+ avd_reset_decoded_fmt(ctx);
+
+ /* Propagate colorspace information to capture. */
+ cap_fmt = &ctx->decoded_fmt;
+ cap_fmt->fmt.pix_mp.colorspace = f->fmt.pix_mp.colorspace;
+ cap_fmt->fmt.pix_mp.xfer_func = f->fmt.pix_mp.xfer_func;
+ cap_fmt->fmt.pix_mp.ycbcr_enc = f->fmt.pix_mp.ycbcr_enc;
+ cap_fmt->fmt.pix_mp.quantization = f->fmt.pix_mp.quantization;
+
+ /* Enable format specific queue features */
+ vq->subsystem_flags |= desc->subsystem_flags;
+
+ return 0;
+}
+
+static int avd_g_output_fmt(struct file *file, void *priv,
+ struct v4l2_format *f)
+{
+ struct avd_ctx *ctx = file_to_ctx(file);
+
+ *f = ctx->coded_fmt;
+ return 0;
+}
+
+static int avd_g_capture_fmt(struct file *file, void *priv,
+ struct v4l2_format *f)
+{
+ struct avd_ctx *ctx = file_to_ctx(file);
+
+ *f = ctx->decoded_fmt;
+ return 0;
+}
+
+static int avd_enum_output_fmt(struct file *file, void *priv,
+ struct v4l2_fmtdesc *f)
+{
+ struct avd_ctx *ctx = file_to_ctx(file);
+ const struct avd_coded_fmt_desc *desc;
+
+ desc = avd_enum_coded_fmt_desc(ctx, f->index);
+ if (!desc)
+ return -EINVAL;
+
+ f->pixelformat = desc->fourcc;
+ return 0;
+}
+
+static int avd_enum_capture_fmt(struct file *file, void *priv,
+ struct v4l2_fmtdesc *f)
+{
+ struct avd_ctx *ctx = file_to_ctx(file);
+ u32 fourcc;
+
+ fourcc = avd_enum_decoded_fmt(ctx, f->index, ctx->image_fmt);
+ if (!fourcc)
+ return -EINVAL;
+
+ f->pixelformat = fourcc;
+ return 0;
+}
+
+const struct v4l2_ioctl_ops avd_ioctl_ops = {
+ .vidioc_querycap = avd_querycap,
+ .vidioc_enum_framesizes = avd_enum_framesizes,
+
+ .vidioc_try_fmt_vid_cap_mplane = avd_try_capture_fmt,
+ .vidioc_try_fmt_vid_out_mplane = avd_try_output_fmt,
+ .vidioc_s_fmt_vid_out_mplane = avd_s_output_fmt,
+ .vidioc_s_fmt_vid_cap_mplane = avd_s_capture_fmt,
+ .vidioc_g_fmt_vid_out_mplane = avd_g_output_fmt,
+ .vidioc_g_fmt_vid_cap_mplane = avd_g_capture_fmt,
+ .vidioc_enum_fmt_vid_out = avd_enum_output_fmt,
+ .vidioc_enum_fmt_vid_cap = avd_enum_capture_fmt,
+
+ .vidioc_reqbufs = v4l2_m2m_ioctl_reqbufs,
+ .vidioc_querybuf = v4l2_m2m_ioctl_querybuf,
+ .vidioc_qbuf = v4l2_m2m_ioctl_qbuf,
+ .vidioc_dqbuf = v4l2_m2m_ioctl_dqbuf,
+ .vidioc_prepare_buf = v4l2_m2m_ioctl_prepare_buf,
+ .vidioc_create_bufs = v4l2_m2m_ioctl_create_bufs,
+ .vidioc_expbuf = v4l2_m2m_ioctl_expbuf,
+
+ .vidioc_subscribe_event = v4l2_ctrl_subscribe_event,
+ .vidioc_unsubscribe_event = v4l2_event_unsubscribe,
+
+ .vidioc_streamon = v4l2_m2m_ioctl_streamon,
+ .vidioc_streamoff = v4l2_m2m_ioctl_streamoff,
+
+ .vidioc_decoder_cmd = v4l2_m2m_ioctl_stateless_decoder_cmd,
+ .vidioc_try_decoder_cmd = v4l2_m2m_ioctl_stateless_try_decoder_cmd,
+};
+
+static int avd_queue_setup(struct vb2_queue *vq, unsigned int *num_buffers,
+ unsigned int *num_planes, unsigned int sizes[],
+ struct device *alloc_devs[])
+{
+ struct avd_ctx *ctx = vb2_get_drv_priv(vq);
+ struct v4l2_format *f;
+ unsigned int i;
+
+ if (V4L2_TYPE_IS_OUTPUT(vq->type))
+ f = &ctx->coded_fmt;
+ else
+ f = &ctx->decoded_fmt;
+
+ if (*num_planes) {
+ if (*num_planes != f->fmt.pix_mp.num_planes)
+ return -EINVAL;
+
+ for (i = 0; i < f->fmt.pix_mp.num_planes; i++) {
+ if (sizes[i] < f->fmt.pix_mp.plane_fmt[i].sizeimage)
+ return -EINVAL;
+ }
+ } else {
+ *num_planes = f->fmt.pix_mp.num_planes;
+ for (i = 0; i < f->fmt.pix_mp.num_planes; i++)
+ sizes[i] = f->fmt.pix_mp.plane_fmt[i].sizeimage;
+ }
+
+ return 0;
+}
+
+static int avd_buf_prepare(struct vb2_buffer *vb)
+{
+ struct vb2_queue *vq = vb->vb2_queue;
+ struct avd_ctx *ctx = vb2_get_drv_priv(vq);
+ struct v4l2_format *f;
+ unsigned int i;
+
+ if (V4L2_TYPE_IS_OUTPUT(vq->type))
+ f = &ctx->coded_fmt;
+ else
+ f = &ctx->decoded_fmt;
+
+ for (i = 0; i < f->fmt.pix_mp.num_planes; ++i) {
+ u32 sizeimage = f->fmt.pix_mp.plane_fmt[i].sizeimage;
+
+ if (vb2_plane_size(vb, i) < sizeimage)
+ return -EINVAL;
+ }
+
+ /*
+ * Buffer's bytesused must be written by driver for CAPTURE buffers.
+ * (for OUTPUT buffers, if userspace passes 0 bytesused, v4l2-core sets
+ * it to buffer length).
+ */
+ if (V4L2_TYPE_IS_CAPTURE(vq->type))
+ vb2_set_plane_payload(vb, 0,
+ f->fmt.pix_mp.plane_fmt[0].sizeimage);
+
+ return 0;
+}
+
+static void avd_buf_queue(struct vb2_buffer *vb)
+{
+ struct avd_ctx *ctx = vb2_get_drv_priv(vb->vb2_queue);
+ struct vb2_v4l2_buffer *vbuf = to_vb2_v4l2_buffer(vb);
+
+ v4l2_m2m_buf_queue(ctx->fh.m2m_ctx, vbuf);
+}
+
+static int avd_buf_out_validate(struct vb2_buffer *vb)
+{
+ struct vb2_v4l2_buffer *vbuf = to_vb2_v4l2_buffer(vb);
+
+ vbuf->field = V4L2_FIELD_NONE;
+ return 0;
+}
+
+static void avd_buf_request_complete(struct vb2_buffer *vb)
+{
+ struct avd_ctx *ctx = vb2_get_drv_priv(vb->vb2_queue);
+
+ v4l2_ctrl_request_complete(vb->req_obj.req, &ctx->ctrl_hdl);
+}
+
+static int avd_start_streaming(struct vb2_queue *q, unsigned int count)
+{
+ struct avd_ctx *ctx = vb2_get_drv_priv(q);
+ const struct avd_coded_fmt_desc *desc;
+ int ret;
+
+ if (V4L2_TYPE_IS_CAPTURE(q->type))
+ return 0;
+
+ desc = ctx->coded_fmt_desc;
+ if (WARN_ON(!desc))
+ return -EINVAL;
+
+ if (desc->ops->start) {
+ ret = desc->ops->start(ctx);
+ if (ret)
+ return ret;
+ }
+
+ return 0;
+}
+
+static void avd_queue_cleanup(struct vb2_queue *vq, u32 state)
+{
+ struct avd_ctx *ctx = vb2_get_drv_priv(vq);
+
+ while (true) {
+ struct vb2_v4l2_buffer *vbuf;
+
+ if (V4L2_TYPE_IS_OUTPUT(vq->type))
+ vbuf = v4l2_m2m_src_buf_remove(ctx->fh.m2m_ctx);
+ else
+ vbuf = v4l2_m2m_dst_buf_remove(ctx->fh.m2m_ctx);
+
+ if (!vbuf)
+ break;
+
+ v4l2_ctrl_request_complete(vbuf->vb2_buf.req_obj.req,
+ &ctx->ctrl_hdl);
+ v4l2_m2m_buf_done(vbuf, state);
+ }
+}
+
+static void avd_stop_streaming(struct vb2_queue *q)
+{
+ struct avd_ctx *ctx = vb2_get_drv_priv(q);
+
+ if (V4L2_TYPE_IS_OUTPUT(q->type)) {
+ const struct avd_coded_fmt_desc *desc = ctx->coded_fmt_desc;
+
+ if (WARN_ON(!desc))
+ return;
+
+ if (desc->ops->stop)
+ desc->ops->stop(ctx);
+ }
+
+ avd_queue_cleanup(q, VB2_BUF_STATE_ERROR);
+}
+
+const struct vb2_ops avd_queue_ops = {
+ .queue_setup = avd_queue_setup,
+ .buf_prepare = avd_buf_prepare,
+ .buf_queue = avd_buf_queue,
+ .buf_out_validate = avd_buf_out_validate,
+ .buf_request_complete = avd_buf_request_complete,
+ .start_streaming = avd_start_streaming,
+ .stop_streaming = avd_stop_streaming,
+};
+
+void avd_job_finish_no_pm(struct avd_ctx *ctx, enum vb2_buffer_state result)
+{
+ if (ctx->coded_fmt_desc->ops->done) {
+ struct vb2_v4l2_buffer *src_buf, *dst_buf;
+
+ src_buf = v4l2_m2m_next_src_buf(ctx->fh.m2m_ctx);
+ dst_buf = v4l2_m2m_next_dst_buf(ctx->fh.m2m_ctx);
+ ctx->coded_fmt_desc->ops->done(ctx, src_buf, dst_buf, result);
+ }
+
+ v4l2_m2m_buf_done_and_job_finish(ctx->dev->m2m_dev, ctx->fh.m2m_ctx,
+ result);
+}
+
+void avd_job_finish(struct avd_ctx *ctx, enum vb2_buffer_state result)
+{
+ struct avd_dev *avd = ctx->dev;
+
+ pm_runtime_put_autosuspend(avd->dev);
+ avd_job_finish_no_pm(ctx, result);
+}
+
+void avd_run_preamble(struct avd_ctx *ctx, struct avd_run *run)
+{
+ struct media_request *src_req;
+ struct avd_decoded_buffer *dst;
+
+ memset(run, 0, sizeof(*run));
+
+ run->bufs.src = v4l2_m2m_next_src_buf(ctx->fh.m2m_ctx);
+ run->bufs.dst = v4l2_m2m_next_dst_buf(ctx->fh.m2m_ctx);
+
+ run->coded_in =
+ vb2_dma_contig_plane_dma_addr(&run->bufs.src->vb2_buf, 0);
+ run->y_out = vb2_dma_contig_plane_dma_addr(&run->bufs.dst->vb2_buf, 0);
+ run->uv_out = run->y_out +
+ ctx->decoded_fmt.fmt.pix_mp.plane_fmt[0].bytesperline *
+ ctx->decoded_fmt.fmt.pix_mp.height;
+
+ run->comp_out = run->y_out + ctx->comp.start_offset;
+ ctx->decomp = !is_interchange(ctx->decoded_fmt.fmt.pix_mp.pixelformat);
+
+ dst = vb2_to_avd_decoded_buf(&run->bufs.dst->vb2_buf);
+ memcpy(&dst->comp, &ctx->comp, sizeof(ctx->comp));
+
+ /* Apply request(s) controls if needed. */
+ src_req = run->bufs.src->vb2_buf.req_obj.req;
+ if (src_req)
+ v4l2_ctrl_request_setup(src_req, &ctx->ctrl_hdl);
+
+ v4l2_m2m_buf_copy_metadata(run->bufs.src, run->bufs.dst);
+}
+
+void avd_run_postamble(struct avd_ctx *ctx, struct avd_run *run)
+{
+ struct media_request *src_req = run->bufs.src->vb2_buf.req_obj.req;
+
+ if (src_req)
+ v4l2_ctrl_request_complete(src_req, &ctx->ctrl_hdl);
+}
+
+static int avd_add_ctrls(struct avd_ctx *ctx, const struct avd_ctrls *ctrls)
+{
+ unsigned int i;
+
+ for (i = 0; i < ctrls->num_ctrls; i++) {
+ const struct v4l2_ctrl_config *cfg = &ctrls->ctrls[i].cfg;
+
+ v4l2_ctrl_new_custom(&ctx->ctrl_hdl, cfg, ctx);
+ if (ctx->ctrl_hdl.error)
+ return ctx->ctrl_hdl.error;
+ }
+
+ return 0;
+}
+
+int avd_init_ctrls(struct avd_ctx *ctx)
+{
+ unsigned int i, nctrls = 0;
+ int ret;
+
+ for (i = 0; i < ARRAY_SIZE(avd_coded_fmts); i++)
+ if (avd_is_capable(ctx, avd_coded_fmts[i].capability))
+ nctrls += avd_coded_fmts[i].ctrls->num_ctrls;
+
+ v4l2_ctrl_handler_init(&ctx->ctrl_hdl, nctrls);
+
+ for (i = 0; i < ARRAY_SIZE(avd_coded_fmts); i++) {
+ if (avd_is_capable(ctx, avd_coded_fmts[i].capability)) {
+ ret = avd_add_ctrls(ctx, avd_coded_fmts[i].ctrls);
+ if (ret)
+ goto err_free_handler;
+ }
+ }
+
+ ret = v4l2_ctrl_handler_setup(&ctx->ctrl_hdl);
+ if (ret)
+ goto err_free_handler;
+
+ ctx->fh.ctrl_handler = &ctx->ctrl_hdl;
+ return 0;
+
+err_free_handler:
+ v4l2_ctrl_handler_free(&ctx->ctrl_hdl);
+ return ret;
+}
diff --git a/drivers/media/platform/apple/avd/avd.h b/drivers/media/platform/apple/avd/avd.h
new file mode 100644
index 000000000000..980ecce4f3df
--- /dev/null
+++ b/drivers/media/platform/apple/avd/avd.h
@@ -0,0 +1,267 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Apple Video Decoder driver
+ *
+ * Copyright (C) 2026 The Asahi Linux Contributors
+ * Copyright (C) 2026 Sofus Forstreuter <sofus.c@icloud.com>
+ *
+ * Based on rkvdec driver by Collabora, Ltd.
+ * Copyright (C) 2019 Collabora, Ltd.
+ * Based on rkvdec driver by Google LLC. (Tomasz Figa <tfiga@chromium.org>)
+ * Based on s5p-mfc driver by Samsung Electronics Co., Ltd.
+ * Copyright (C) 2011 Samsung Electronics Co., Ltd.
+ */
+
+#ifndef AVD_H_
+#define AVD_H_
+
+#include <linux/platform_device.h>
+#include <linux/firmware.h>
+#include <linux/iommu.h>
+#include <linux/slab.h>
+
+#include <media/v4l2-ctrls.h>
+#include <media/v4l2-device.h>
+#include <media/v4l2-mem2mem.h>
+#include <media/v4l2-ioctl.h>
+
+#define AVD_CAPABILITY_HEVC BIT(0)
+#define AVD_CAPABILITY_H264 BIT(1)
+#define AVD_CAPABILITY_VP9 BIT(2)
+#define AVD_CAPABILITY_AV1 BIT(3)
+
+/* Shifts addresses right and bytesperline?? */
+#define AVD_QUIRK_LSR BIT(0)
+#define AVD_QUIRK_NO_PIPE_STATE BIT(1)
+
+#define AVD_REG_RUN_CTRL 0x08
+#define AVD_RUN_CTRL_UNK_RUN BIT(0)
+#define AVD_RUN_CTRL_UNK_STOP (BIT(1) | BIT(2) | BIT(3))
+
+#define AVD_REG_MBOX_IRQ_ENABLE 0x48
+#define AVD_REG_MBOX_IRQ_CLR 0x4c
+#define AVD_MBOX0_EMPTY BIT(0)
+#define AVD_MBOX0_NOT_EMPTY BIT(1)
+
+#define AVD_REG_MBOX0_STATUS 0x50
+#define AVD_REG_MBOX0_RETRIEVE 0x58
+#define AVD_REG_MBOX1_STATUS 0x5c
+#define AVD_REG_MBOX1_SUBMIT 0x60
+#define AVD_MBOX_ENABLE BIT(0)
+
+#define AVD_REG_FLAG0_SET 0x90
+#define AVD_REG_FLAG0_CLR 0x98
+
+/* size in number of words (u32) minus one */
+#define AVD_PIODMA_CMD_SIZE(v) FIELD_PREP(GENMASK(31, 18), v)
+#define AVD_PIODMA_CMD_DEST(v) (GENMASK(17, 2) & (v))
+#define AVD_PIODMA_CMD_CONST BIT(0)
+
+/*
+ * AVD needs most addresses to be aligned to 256
+ * the only exception are the compressed buffers, they are aligned to 128
+ * instead
+ */
+#define AVD_ALIGN 256
+/* hevc, with a b slice where all references are active and weights are sent */
+#define AVD_MAX_INST 512
+
+struct avd_ctx;
+struct avd_dev;
+
+/* Matches hdr mode (inst stream) and register layout */
+enum avd_codec {
+ AVD_CODEC_HEVC = 0,
+ AVD_CODEC_H264 = 1,
+ AVD_CODEC_VP9 = 2,
+ AVD_CODEC_AV1 = 3
+};
+
+struct avd_run {
+ struct {
+ struct vb2_v4l2_buffer *src; /* OUTPUT coded */
+ struct vb2_v4l2_buffer *dst; /* CAPTURE decoded */
+ } bufs;
+
+ dma_addr_t coded_in;
+ dma_addr_t y_out;
+ dma_addr_t uv_out;
+ dma_addr_t comp_out;
+};
+
+struct avd_ctrl_desc {
+ struct v4l2_ctrl_config cfg;
+};
+
+struct avd_ctrls {
+ const struct avd_ctrl_desc *ctrls;
+ unsigned int num_ctrls;
+};
+
+struct avd_comp {
+ u32 size;
+ /* offset to start of compressed data */
+ size_t start_offset;
+ /* relative offsets to start */
+ u32 offsets[4];
+};
+
+struct avd_decoded_buffer {
+ /* Must be the first field in this struct. */
+ struct v4l2_m2m_buffer base;
+ struct avd_comp comp;
+};
+
+static inline struct avd_decoded_buffer *
+vb2_to_avd_decoded_buf(struct vb2_buffer *buf)
+{
+ return container_of(buf, struct avd_decoded_buffer, base.vb.vb2_buf);
+}
+
+struct avd_decoded_buffer *avd_get_ref_buf(struct avd_ctx *ctx,
+ struct vb2_v4l2_buffer *dst,
+ u64 timestamp);
+
+struct avd_coded_fmt_ops {
+ void (*adjust_decoded_fmt)(struct avd_ctx *ctx,
+ struct v4l2_pix_format_mplane *pix_mp);
+ int (*start)(struct avd_ctx *ctx);
+ void (*stop)(struct avd_ctx *ctx);
+ int (*run)(struct avd_ctx *ctx);
+ void (*done)(struct avd_ctx *ctx, struct vb2_v4l2_buffer *src_buf,
+ struct vb2_v4l2_buffer *dst_buf,
+ enum vb2_buffer_state result);
+ int (*try_ctrl)(struct avd_ctx *ctx, struct v4l2_ctrl *ctrl);
+ enum avd_image_fmt (*get_image_fmt)(struct avd_ctx *ctx,
+ struct v4l2_ctrl *ctrl);
+};
+
+enum avd_image_fmt {
+ AVD_IMG_FMT_ANY = 0,
+ AVD_IMG_FMT_420_8BIT,
+ AVD_IMG_FMT_420_10BIT,
+ AVD_IMG_FMT_422_8BIT,
+ AVD_IMG_FMT_422_10BIT,
+};
+
+struct avd_decoded_fmt_desc {
+ u32 fourcc;
+ enum avd_image_fmt image_fmt;
+};
+
+struct avd_coded_fmt_desc {
+ u32 fourcc;
+ struct v4l2_frmsize_stepwise frmsize;
+ const struct avd_ctrls *ctrls;
+ const struct avd_coded_fmt_ops *ops;
+ u32 subsystem_flags;
+ unsigned int capability;
+};
+
+struct avd_variant {
+ unsigned int capabilities;
+ const char *fw_name;
+ unsigned char revision; /* the same as the device tree */
+ unsigned int quirks;
+};
+
+struct avd_dev {
+ struct device *dev;
+ struct v4l2_device v4l2_dev;
+ struct media_device mdev;
+ struct video_device vdev;
+ struct v4l2_m2m_dev *m2m_dev;
+ struct platform_device *pdev;
+ const struct firmware *fw; /* fw is lost on suspend */
+ void __iomem *piodma;
+ void __iomem *code;
+ void __iomem *sram;
+ void __iomem *mbox;
+ void __iomem *ctrl;
+ u32 sram_start;
+ struct iommu_domain *domain;
+ struct iommu_domain *empty_domain;
+ struct mutex vdev_lock; /* serializes ioctls */
+ struct reset_control *rstc;
+ const struct avd_variant *variant;
+};
+
+struct avd_segment {
+ u32 piodma_cmd;
+ u32 num;
+ u32 instructions[AVD_MAX_INST];
+};
+
+struct avd_buf {
+ void *cpu;
+ dma_addr_t addr;
+ size_t size;
+};
+
+struct avd_job {
+ enum avd_codec codec;
+ int dest;
+ size_t num;
+ size_t num_alloc;
+ size_t num_submit;
+ struct avd_segment *segments;
+ struct avd_buf buf;
+};
+
+struct avd_ctx {
+ struct v4l2_fh fh;
+ struct avd_dev *dev;
+ struct v4l2_format coded_fmt;
+ struct v4l2_format decoded_fmt;
+ const struct avd_coded_fmt_desc *coded_fmt_desc;
+ struct v4l2_ctrl_handler ctrl_hdl;
+ enum avd_image_fmt image_fmt;
+ bool decomp;
+ struct delayed_work watchdog_work;
+ void *priv;
+ struct avd_comp comp;
+ struct avd_job job;
+ struct avd_buf inst;
+ struct avd_buf pipe_state;
+};
+
+int avd_end_segment(struct avd_ctx *ctx, bool update_submit);
+int avd_init_job(struct avd_ctx *ctx, enum avd_codec codec, size_t segments);
+int avd_submit_job(struct avd_ctx *ctx);
+
+int avd_buf_alloc(struct avd_dev *avd, struct avd_buf *buf, size_t size);
+void avd_buf_free(struct avd_dev *avd, struct avd_buf *buf);
+
+void avd_reset_coded_fmt(struct avd_ctx *ctx);
+void avd_reset_decoded_fmt(struct avd_ctx *ctx);
+int avd_init_ctrls(struct avd_ctx *ctx);
+
+void avd_job_finish_no_pm(struct avd_ctx *ctx, enum vb2_buffer_state result);
+void avd_job_finish(struct avd_ctx *ctx, enum vb2_buffer_state result);
+
+void avd_run_preamble(struct avd_ctx *ctx, struct avd_run *run);
+void avd_run_postamble(struct avd_ctx *ctx, struct avd_run *run);
+
+extern const struct v4l2_ctrl_ops avd_ctrl_ops;
+extern const struct v4l2_ioctl_ops avd_ioctl_ops;
+extern const struct vb2_ops avd_queue_ops;
+
+static inline u32 fmt_height(struct avd_ctx *ctx)
+{
+ return ctx->coded_fmt.fmt.pix_mp.height;
+}
+
+static inline u32 fmt_width(struct avd_ctx *ctx)
+{
+ return ctx->coded_fmt.fmt.pix_mp.width;
+}
+
+void fill_comp(struct avd_comp *comp, enum avd_image_fmt image_fmt, u32 width,
+ u32 height);
+
+static inline struct avd_ctx *file_to_ctx(struct file *filp)
+{
+ return container_of(file_to_v4l2_fh(filp), struct avd_ctx, fh);
+}
+
+#endif /* AVD_H_ */
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 07/17] media: apple: avd: add h264 support
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (5 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 06/17] media: apple: add avd driver Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 08/17] media: apple: avd: add vp9 support Sofus Forstreuter
` (9 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
The fluster score is 77/135 for JVT-AVC_V1 and 42/69 for JVT-FR-EXT.
While there are no unexpected test cases. FM1_FT_E, SP1_BT_A and
sp2_bt_b still have some unsupported features that are not rejected,
which causes the hardware to fault.
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
drivers/media/platform/apple/avd/Kconfig | 1 +
drivers/media/platform/apple/avd/Makefile | 2 +-
drivers/media/platform/apple/avd/avd-h264.c | 862 ++++++++++++++++++++++++++++
drivers/media/platform/apple/avd/avd-v4l2.c | 72 ++-
drivers/media/platform/apple/avd/avd.h | 2 +
5 files changed, 937 insertions(+), 2 deletions(-)
diff --git a/drivers/media/platform/apple/avd/Kconfig b/drivers/media/platform/apple/avd/Kconfig
index 68efa2df368a..cad7cfa72451 100644
--- a/drivers/media/platform/apple/avd/Kconfig
+++ b/drivers/media/platform/apple/avd/Kconfig
@@ -9,6 +9,7 @@ config VIDEO_APPLE_AVD
depends on V4L_PLATFORM_DRIVERS
select V4L2_MEM2MEM_DEV
select VIDEOBUF2_DMA_CONTIG
+ select V4L2_H264
help
Support for hardware video decoding on Apple Silicon devices using
the Apple Video Decoder (AVD).
diff --git a/drivers/media/platform/apple/avd/Makefile b/drivers/media/platform/apple/avd/Makefile
index b2f6736a790e..482d43f8ea8e 100644
--- a/drivers/media/platform/apple/avd/Makefile
+++ b/drivers/media/platform/apple/avd/Makefile
@@ -1,4 +1,4 @@
# SPDX-License-Identifier: GPL-2.0-only
-apple-avd-y := avd-drv.o avd-v4l2.o
+apple-avd-y := avd-drv.o avd-v4l2.o avd-h264.o
obj-$(CONFIG_VIDEO_APPLE_AVD) += apple-avd.o
diff --git a/drivers/media/platform/apple/avd/avd-h264.c b/drivers/media/platform/apple/avd/avd-h264.c
new file mode 100644
index 000000000000..a6acf4b2a030
--- /dev/null
+++ b/drivers/media/platform/apple/avd/avd-h264.c
@@ -0,0 +1,862 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Apple Video Decoder H264 driver
+ *
+ * Copyright (C) 2026 The Asahi Linux Contributors
+ * Copyright (C) 2026 Sofus Forstreuter <sofus.c@icloud.com>
+ * Copyright (C) 2023 Eileen Yoon <eyn@gmx.com>
+ *
+ * Copyright (c) 2014 Rockchip Electronics Co., Ltd.
+ * Hertz Wong <hertz.wong@rock-chips.com>
+ * Herman Chen <herman.chen@rock-chips.com>
+ *
+ * Copyright (C) 2014 Google, Inc.
+ * Tomasz Figa <tfiga@chromium.org>
+ */
+
+#include <linux/dev_printk.h>
+
+#include <media/v4l2-h264.h>
+#include <media/videobuf2-dma-contig.h>
+
+#include "avd.h"
+#include "avd-inst.h"
+
+#define H264_SCL_DIMS (0x1000000 | ((64 / 4) << 5) | (((16 / 4) << 5) - 1))
+
+#define h264_FLAG_CI_PRED(v) FIELD_PREP(BIT(19), !!(v))
+#define H264_FLAG_ENTROPY_CODING_MODE(v) FIELD_PREP(BIT(20), !!(v))
+#define H264_FLAG_NOT_IDR(v) FIELD_PREP(BIT(21), !!(v))
+
+#define H264_TRANSFORM_8X8_MODE(v) FIELD_PREP(BIT(7), !!(v))
+
+struct avd_h264_run {
+ struct avd_run base;
+
+ const struct v4l2_ctrl_h264_decode_params *decode_params;
+ const struct v4l2_ctrl_h264_sps *sps;
+ const struct v4l2_ctrl_h264_pps *pps;
+ const struct v4l2_ctrl_h264_slice_params *slice_params;
+
+ const struct v4l2_ctrl_h264_scaling_matrix *scaling_matrix;
+ const struct v4l2_ctrl_h264_pred_weights *pred_weights;
+
+ struct run_addr {
+ dma_addr_t mv_color;
+ } addresses;
+
+ s32 cur_poc;
+ u8 num_valid;
+};
+
+/* state */
+struct avd_h264_ctx {
+ struct avd_h264_reflists {
+ struct v4l2_h264_reference p[V4L2_H264_REF_LIST_LEN];
+ struct v4l2_h264_reference b0[V4L2_H264_REF_LIST_LEN];
+ struct v4l2_h264_reference b1[V4L2_H264_REF_LIST_LEN];
+ } reflists;
+
+ struct avd_buf *slices;
+ struct avd_buf *active_slice;
+ size_t slice_num;
+ size_t alloc_slice_num;
+
+ struct avd_h264_bufs {
+ struct avd_buf above_info;
+ struct avd_buf lf_above_info;
+ struct avd_buf lf_above;
+ struct avd_buf ip_above;
+ struct avd_buf mv_above_info;
+ } bufs;
+};
+
+/*
+ * use first_mb_in_slice instead of fh.m2m_ctx->new_frame to detect at new
+ * frame since many clients dont submit propper timestamps
+ * We only support slices in raster order anyway
+ */
+#define is_new_frame(sl) ((sl)->first_mb_in_slice == 0)
+
+/* scaling matrix */
+static const u32 default_8x8_intra[] = {
+ 0x060a0d10, 0x0a0b1012, 0x0d101217, 0x10121719, 0x1217191b, 0x17191b1d,
+ 0x191b1d1f, 0x1b1d1f21, 0x1217191b, 0x17191b1d, 0x191b1d1f, 0x1b1d1f21,
+ 0x1d1f2124, 0x1f212426, 0x21242628, 0x2426282a,
+};
+
+static const u32 default_8x8_inter[] = {
+ 0x090d0f11, 0x0d0d1113, 0x0f111315, 0x11131516, 0x13151618, 0x15161819,
+ 0x1618191b, 0x18191b1c, 0x13151618, 0x15161819, 0x1618191b, 0x18191b1c,
+ 0x191b1c1e, 0x1b1c1e20, 0x1c1e2021, 0x1e202123,
+};
+
+static inline u32 mv_color_size(u32 w, u32 h)
+{
+ return (DIV_ROUND_UP(w, 16) + 1) * (DIV_ROUND_UP(h, 16) + 1) * 64;
+}
+
+/* sorry for the formatting */
+
+static void stream_refs(struct avd_ctx *ctx, struct avd_h264_run *run)
+{
+ const struct v4l2_ctrl_h264_decode_params *decode = run->decode_params;
+ const struct v4l2_h264_dpb_entry *dpb = decode->dpb;
+ struct avd_h264_ctx *h264_ctx = ctx->priv;
+ struct avd_decoded_buffer *dst, *ref;
+ dma_addr_t addr;
+
+ dst = vb2_to_avd_decoded_buf(&run->base.bufs.dst->vb2_buf);
+
+ push(0, "");
+ pusha(h264_ctx->bufs.mv_above_info.addr, "mv_above_info", 7);
+ pusha(run->addresses.mv_color, "mv_color", 0);
+
+ push(0, "");
+ push(0, "");
+ push(0, "");
+ push(0, "");
+
+ for (int i = 0; i < ARRAY_SIZE(decode->dpb); i++) {
+ if (!(dpb[i].flags & V4L2_H264_DPB_ENTRY_FLAG_VALID))
+ continue;
+
+ ref = avd_get_ref_buf(ctx, &dst->base.vb, dpb[i].reference_ts);
+
+ addr = vb2_dma_contig_plane_dma_addr(&ref->base.vb.vb2_buf, 0) +
+ ref->comp.start_offset;
+
+ push(AVD_REF_NUM(run->num_valid - 1) | AVD_REF_FLAG_CONST |
+ AVD_REF_FLAG_LONG(
+ dpb[i].flags &
+ V4L2_H264_DPB_ENTRY_FLAG_LONG_TERM) |
+ AVD_REF_DELTA_POC(run->cur_poc -
+ dpb[i].top_field_order_cnt),
+ "hdr_d0_ref_hdr");
+
+ push_comp(ctx, addr, ctx->comp.offsets);
+ }
+}
+
+static void stream_scaling(struct avd_ctx *ctx, struct avd_h264_run *run)
+{
+ const struct v4l2_ctrl_h264_pps *pps = run->pps;
+ const struct v4l2_ctrl_h264_scaling_matrix *scaling =
+ run->scaling_matrix;
+
+ push(H264_SCL_DIMS, "hdr_4c_pic_scaling_list_dims");
+
+ for (int i = 0; i < 6; i++)
+ for (int j = 0; j < 16; j += 4)
+ push(AVD_SCALING_I3(
+ scaling->scaling_list_4x4[i][j + 0]) |
+ AVD_SCALING_I2(
+ scaling->scaling_list_4x4[i][j + 1]) |
+ AVD_SCALING_I1(
+ scaling->scaling_list_4x4[i][j + 2]) |
+ AVD_SCALING_I0(
+ scaling->scaling_list_4x4[i][j + 3]),
+ "scl_46c_pic_scaling_matrix_4x4");
+
+ /* Instead of 8x8 raster scan order avd expects 4 4x4 subblocks */
+ static const u8 map[16] = {
+ 0, 8, 16, 24, /* top left */
+ 4, 12, 20, 28, /* top right */
+ 32, 40, 48, 56, /* bottom left */
+ 36, 44, 52, 60, /* bottom right */
+ };
+
+ /* 7.3.2.2, only matrix 0 and 1 are used if chroma_format_idc < 3 */
+ if (pps->flags & V4L2_H264_PPS_FLAG_TRANSFORM_8X8_MODE) {
+ for (int i = 0; i < 2; i++)
+ for (int j = 0; j < 16; j++)
+ push(AVD_SCALING_I3(scaling->scaling_list_8x8
+ [i][map[j] + 0]) |
+ AVD_SCALING_I2(
+ scaling->scaling_list_8x8
+ [i][map[j] + 1]) |
+ AVD_SCALING_I1(
+ scaling->scaling_list_8x8
+ [i][map[j] + 2]) |
+ AVD_SCALING_I0(
+ scaling->scaling_list_8x8
+ [i][map[j] + 3]),
+ "scl_4cc_pic_scaling_matrix_8x8");
+
+ } else {
+ for (int i = 0; i < ARRAY_SIZE(default_8x8_intra); i++)
+ push(default_8x8_intra[i],
+ "scl_4cc_pic_scaling_matrix_8x8");
+ for (int i = 0; i < ARRAY_SIZE(default_8x8_inter); i++)
+ push(default_8x8_inter[i],
+ "scl_4cc_pic_scaling_matrix_8x8");
+ }
+}
+
+static void stream_hdr(struct avd_ctx *ctx, struct avd_h264_run *run)
+{
+ const struct v4l2_ctrl_h264_decode_params *decode = run->decode_params;
+ const struct v4l2_ctrl_h264_sps *sps = run->sps;
+ const struct v4l2_ctrl_h264_pps *pps = run->pps;
+ struct avd_dev *avd = ctx->dev;
+ struct avd_h264_ctx *h264_ctx = ctx->priv;
+ u32 bytesperline;
+ u32 width = (sps->pic_width_in_mbs_minus1 + 1) * 16;
+ u32 height = (sps->pic_height_in_map_units_minus1 + 1) * 16;
+
+ push(AVD_OP_HDR | AVD_OP_HDR_FLAG_DECOMP(ctx->decomp) |
+ AVD_OP_HDR_FLAG_INTRA(
+ decode->flags &
+ V4L2_H264_DECODE_PARAM_FLAG_IDR_PIC) |
+ AVD_OP_HDR_CONST |
+ AVD_OP_HDR_FLAG_PIPE_STATE(
+ !(avd->variant->quirks & AVD_QUIRK_NO_PIPE_STATE)),
+ "hdr_34_start_hdr");
+
+ push(AVD_HDR_CODEC_MODE(AVD_CODEC_H264), "hdr_38_mode");
+
+ push(AVD_HDR_HEIGHT(height - 1) | AVD_HDR_WIDTH(width - 1),
+ "hdr_3c_height_width");
+
+ push(0, "hdr_40_zero");
+
+ push(AVD_HDR_HEIGHT((height - 1) >> 3) |
+ AVD_HDR_WIDTH((width - 1) >> 3),
+ "hdr_28_height_width_shift3");
+
+ push(AVD_HDR_COMMON_CHROMA_FORMAT(sps->chroma_format_idc) |
+ AVD_HDR_COMMON_BIT_DEPTH_L(sps->bit_depth_luma_minus8) |
+ AVD_HDR_COMMON_BIT_DEPTH_C(sps->bit_depth_chroma_minus8) |
+ AVD_HDR_COMMON_MIN_LUMA_CBS(1) |
+ AVD_HDR_COMMON_LUMA_CBS(1) |
+ H264_TRANSFORM_8X8_MODE(
+ pps->flags &
+ V4L2_H264_PPS_FLAG_TRANSFORM_8X8_MODE) |
+ AVD_HDR_COMMON_FLAG0(
+ sps->flags &
+ V4L2_H264_SPS_FLAG_DIRECT_8X8_INFERENCE),
+ "hdr_2c_sps_param");
+
+ push(H264_FLAG_ENTROPY_CODING_MODE(
+ pps->flags & V4L2_H264_PPS_FLAG_ENTROPY_CODING_MODE) |
+ H264_FLAG_NOT_IDR(!(decode->flags &
+ V4L2_H264_DECODE_PARAM_FLAG_IDR_PIC)) |
+ h264_FLAG_CI_PRED(
+ pps->flags &
+ V4L2_H264_PPS_FLAG_CONSTRAINED_INTRA_PRED),
+ "hdr_44_flags");
+
+ push(AVD_HDR_H26X_QP_OFFSET_CB(pps->chroma_qp_index_offset) |
+ AVD_HDR_H26X_QP_OFFSET_CR(
+ pps->second_chroma_qp_index_offset),
+ "hdr_48_chroma_qp_index_offset");
+
+ push(AVD_HDR_FEAT_H26X | AVD_HDR_FEAT_COMMON | AVD_HDR_FEAT_H264 |
+ AVD_HDR_FEAT_PIPE_STATE_EN(
+ !(avd->variant->quirks & AVD_QUIRK_NO_PIPE_STATE)),
+ "hdr_58_const_3a");
+
+ push(0, "");
+ push(0, "");
+
+ if (avd->variant->revision == 3)
+ push(0, "zero");
+
+ pusha(h264_ctx->bufs.above_info.addr, "hdr_9c_pps_tile_addr_lsb8", 0);
+
+ push(0, "");
+ push(0, "");
+
+ if (avd->variant->revision == 3)
+ push(0, "zero");
+ else if (!(avd->variant->quirks & AVD_QUIRK_NO_PIPE_STATE))
+ pusha(ctx->pipe_state.addr, "pipe_state", 0);
+
+ pusha(h264_ctx->bufs.ip_above.addr, "ip_above", 1);
+ pusha(h264_ctx->bufs.lf_above.addr, "lf_above", 2);
+ pusha(h264_ctx->bufs.lf_above_info.addr, "lf_above_info", 3);
+ push(0, "");
+
+ push_comp(ctx, run->base.comp_out, ctx->comp.offsets);
+
+ bytesperline = ctx->decoded_fmt.fmt.pix_mp.plane_fmt[0].bytesperline;
+ if (avd->variant->quirks & AVD_QUIRK_LSR)
+ bytesperline = bytesperline >> 4;
+
+ pusha(run->base.y_out, "hdr_210_y_addr_lsb8", 0);
+ push(bytesperline, "hdr_218_width_align");
+ pusha(run->base.uv_out, "hdr_214_uv_addr_lsb8", 0);
+ push(bytesperline, "hdr_21c_width_align");
+
+ push(0, "cm3_mark_end_section");
+ push(((height - 1) << 16) | (width - 1), "hdr_54_height_width");
+
+ if (!(decode->flags & V4L2_H264_DECODE_PARAM_FLAG_IDR_PIC))
+ stream_refs(ctx, run);
+
+ if (pps->flags & V4L2_H264_PPS_FLAG_SCALING_MATRIX_PRESENT)
+ stream_scaling(ctx, run);
+ else
+ push(0, "cm3_mark_end_section_scl");
+}
+
+#define DEFAULT_WEIGHT_DENOM \
+ (AVD_OP_WEIGHTS_HDR_LUMA(5) | AVD_OP_WEIGHTS_HDR_CHROMA(5))
+
+static void stream_weights(struct avd_ctx *ctx, struct avd_h264_run *run)
+{
+ int luma_denom, chroma_denom;
+ struct v4l2_h264_weight_factors factors;
+ const struct v4l2_ctrl_h264_pred_weights *weights = run->pred_weights;
+ const struct v4l2_ctrl_h264_pps *pps = run->pps;
+ const struct v4l2_ctrl_h264_slice_params *sl = run->slice_params;
+
+ bool pred_weight_req = V4L2_H264_CTRL_PRED_WEIGHTS_REQUIRED(pps, sl);
+ bool default_weights = pps->weighted_bipred_idc == 2 &&
+ !pred_weight_req;
+
+ push(AVD_OP_WEIGHTS_HDR | AVD_OP_WEIGHTS_HDR_FLAG1(default_weights) |
+ AVD_OP_WEIGHTS_HDR_FLAG0(pred_weight_req) |
+ AVD_OP_WEIGHTS_HDR_LUMA(
+ !default_weights ?
+ weights->luma_log2_weight_denom :
+ 0) |
+ AVD_OP_WEIGHTS_HDR_CHROMA(
+ !default_weights ?
+ weights->chroma_log2_weight_denom :
+ 0) |
+ (default_weights ? DEFAULT_WEIGHT_DENOM : 0),
+ "slc_76c_cmd_weights_denom");
+
+ if (!pred_weight_req)
+ return;
+
+ luma_denom = 1 << weights->luma_log2_weight_denom;
+ chroma_denom = 1 << weights->chroma_log2_weight_denom;
+
+ for (int y = 0; y < 2; y++) {
+ if (y == 1 && sl->slice_type != V4L2_H264_SLICE_TYPE_B)
+ break;
+
+ factors = weights->weight_factors[y];
+ int to = y == 0 ? sl->num_ref_idx_l0_active_minus1 :
+ sl->num_ref_idx_l1_active_minus1;
+ for (int i = 0; i < to + 1; i++) {
+ /*
+ * AVD only expects offsets/weights if they are not
+ * the default ones, otherwise we get artifacts
+ */
+ if (factors.luma_weight[i] != luma_denom ||
+ factors.luma_offset[i] != 0) {
+ push(AVD_OP_WEIGHTS | AVD_OP_WEIGHTS_IDENT(1) |
+ AVD_OP_WEIGHTS_LIST_IDX(y) |
+ AVD_OP_WEIGHTS_INDEX(i) |
+ AVD_OP_WEIGHTS_WEIGHT(
+ factors.luma_weight[i]),
+ "slc_luma_weights");
+ push(AVD_OP_OFFSETS |
+ AVD_OP_OFFSETS_OFFSET(
+ factors.luma_offset[i]),
+ "slc_luma_offsets");
+ }
+
+ if (factors.chroma_weight[i][0] != chroma_denom ||
+ factors.chroma_offset[i][0] != 0 ||
+ factors.chroma_weight[i][1] != chroma_denom ||
+ factors.chroma_offset[i][1] != 0) {
+ push(AVD_OP_WEIGHTS | AVD_OP_WEIGHTS_IDENT(2) |
+ AVD_OP_WEIGHTS_LIST_IDX(y) |
+ AVD_OP_WEIGHTS_INDEX(i) |
+ AVD_OP_WEIGHTS_WEIGHT(
+ factors.chroma_weight[i][0]),
+ "slc_chroma_weights[0]");
+ push(AVD_OP_OFFSETS |
+ AVD_OP_OFFSETS_OFFSET(
+ factors.chroma_offset[i][0]),
+ "slc_chroma_offsets[0]");
+ push(AVD_OP_WEIGHTS | AVD_OP_WEIGHTS_IDENT(3) |
+ AVD_OP_WEIGHTS_LIST_IDX(y) |
+ AVD_OP_WEIGHTS_INDEX(i) |
+ AVD_OP_WEIGHTS_WEIGHT(
+ factors.chroma_weight[i][1]),
+ "slc_chroma_weights[1]");
+ push(AVD_OP_OFFSETS |
+ AVD_OP_OFFSETS_OFFSET(
+ factors.chroma_offset[i][1]),
+ "slc_chroma_offsets[1]");
+ }
+ }
+ }
+}
+
+static void stream_slice(struct avd_ctx *ctx, struct avd_h264_run *run)
+{
+ const struct v4l2_ctrl_h264_decode_params *decode = run->decode_params;
+ const struct v4l2_ctrl_h264_pps *pps = run->pps;
+ const struct v4l2_ctrl_h264_sps *sps = run->sps;
+ const struct v4l2_ctrl_h264_slice_params *sl = run->slice_params;
+ struct avd_h264_ctx *h264_ctx = ctx->priv;
+ u32 payload_len = h264_ctx->active_slice->size;
+ bool en_mode = (pps->flags & V4L2_H264_PPS_FLAG_ENTROPY_CODING_MODE) ==
+ 0;
+ const u8 *data = h264_ctx->active_slice->cpu;
+ u32 min_off = (sl->header_bit_size + (en_mode ? 0 : 7)) / 8;
+ u32 off = 2;
+ u32 num_ref_idx_active, bytes_read = 2;
+ dma_addr_t coded_in, mv_color_addr;
+ struct avd_decoded_buffer *dst, *ref;
+
+ dst = vb2_to_avd_decoded_buf(&run->base.bufs.dst->vb2_buf);
+
+ if (payload_len < min_off)
+ return;
+
+ /* include emulation byte in offset to slice header */
+ while (bytes_read < min_off && off < payload_len) {
+ if (data[off - 2] != 0x00 || data[off - 1] != 0x00 ||
+ data[off] != 0x03)
+ bytes_read++;
+ off++;
+ }
+
+ coded_in = h264_ctx->active_slice->addr + off;
+
+ push(AVD_OP_CODED_DATA |
+ AVD_OP_CODED_DATA_BIT_OFF(
+ en_mode ? (sl->header_bit_size % 8) : 0) |
+ AVD_OP_CODED_IN_HI(coded_in),
+ "slc_a7c_cmd_set_coded_slice");
+ push(AVD_OP_CODED_IN_LO(coded_in), "slc_a84_slice_addr_low");
+ push(payload_len - off, "slc_a88_slice_hdr_size");
+
+ push(AVD_OP_SL_LOC |
+ AVD_OP_SL_LOC_Y(sl->first_mb_in_slice /
+ (sps->pic_width_in_mbs_minus1 + 1)) |
+ AVD_OP_SL_LOC_X(sl->first_mb_in_slice %
+ (sps->pic_width_in_mbs_minus1 + 1)),
+ "cm3_cmd_exec_mb_vp");
+
+ push(AVD_OP_QP | AVD_OP_QP_VAL(26 + pps->pic_init_qp_minus26 +
+ sl->slice_qp_delta),
+ "slc_a70_cmd_quant_param");
+
+ push(AVD_OP_DBLK |
+ AVD_OP_DBLK_FLAG_FULL_EN(
+ sl->disable_deblocking_filter_idc == 0) |
+ AVD_OP_DBLK_FLAG_EN(sl->disable_deblocking_filter_idc !=
+ 1) |
+ AVD_OP_DBLK_OFF1(sl->slice_beta_offset_div2) |
+ AVD_OP_DBLK_OFF0(sl->slice_alpha_c0_offset_div2),
+ "slc_a74_cmd_deblocking_filter");
+
+ if (sl->slice_type == V4L2_H264_SLICE_TYPE_P ||
+ sl->slice_type == V4L2_H264_SLICE_TYPE_B) {
+ num_ref_idx_active = sl->num_ref_idx_l0_active_minus1 + 1;
+ for (u32 i = 0; i < num_ref_idx_active; i++)
+ push(AVD_OP_REF | AVD_OP_REF_LIST_IDX(0) |
+ AVD_OP_REF_LOOP_IDX(i) |
+ AVD_OP_REF_DBP_IDX(
+ sl->ref_pic_list0[i].index),
+ "slc_6e8_cmd_ref_list_0");
+
+ if (sl->slice_type == V4L2_H264_SLICE_TYPE_B) {
+ num_ref_idx_active =
+ sl->num_ref_idx_l1_active_minus1 + 1;
+ for (u32 i = 0; i < num_ref_idx_active; i++)
+ push(AVD_OP_REF | AVD_OP_REF_LIST_IDX(1) |
+ AVD_OP_REF_LOOP_IDX(i) |
+ AVD_OP_REF_DBP_IDX(
+ sl->ref_pic_list1[i].index),
+ "slc_6e8_cmd_ref_list_0");
+ }
+ stream_weights(ctx, run);
+ }
+
+ if (sl->first_mb_in_slice == 0) {
+ push(AVD_OP_SL_DIM_START, "cm3_cmd_set_mb_dims");
+ push(AVD_SL_DIM_END_Y(sps->pic_height_in_map_units_minus1) |
+ AVD_SL_DIM_END_X(sps->pic_width_in_mbs_minus1),
+ "cm3_set_mb_dims");
+ }
+
+ push(AVD_OP_SL_REF | AVD_OP_SL_REF_FLAG_CABAC(sl->cabac_init_idc == 1) |
+ AVD_OP_SL_REF_FLAG1(sl->cabac_init_idc == 2) |
+ AVD_OP_SL_REF_FLAG2(
+ !(sl->flags &
+ V4L2_H264_SLICE_FLAG_DIRECT_SPATIAL_MV_PRED)) |
+ AVD_OP_SL_REF_NUM_L0(sl->num_ref_idx_l0_active_minus1) |
+ AVD_OP_SL_REF_NUM_L1(sl->num_ref_idx_l1_active_minus1) |
+ AVD_OP_SL_REF_SLICE_P(sl->slice_type ==
+ V4L2_H264_SLICE_TYPE_P) |
+ AVD_OP_SL_REF_SLICE_I(sl->slice_type ==
+ V4L2_H264_SLICE_TYPE_I) |
+ AVD_OP_SL_REF_SLICE_B(sl->slice_type ==
+ V4L2_H264_SLICE_TYPE_B)
+
+ ,
+ "slc_6e4_cmd_ref_type");
+
+ if (sl->slice_type == V4L2_H264_SLICE_TYPE_B) {
+
+ /* sl->ref_pic_list1[0].index < ARRAY_SIZE(decode->dpb) */
+ /* bidirectional reference of previous mv */
+ ref = avd_get_ref_buf(
+ ctx, &dst->base.vb,
+ decode->dpb[sl->ref_pic_list1[0].index].reference_ts);
+
+ mv_color_addr =
+ vb2_dma_contig_plane_dma_addr(&ref->base.vb.vb2_buf,
+ 0) +
+ (ref->base.vb.vb2_buf.planes[0].length -
+ mv_color_size(fmt_width(ctx), fmt_height(ctx)));
+
+ pusha(mv_color_addr, "slc_a78_sps_tile_addr2_lsb8", 0);
+ }
+}
+
+static int avd_h264_alloc_bufs(struct avd_ctx *ctx)
+{
+ struct avd_dev *dev = ctx->dev;
+ struct avd_h264_ctx *h264_ctx = ctx->priv;
+ int ret, w, bit_depth, mb;
+
+ w = fmt_width(ctx);
+ bit_depth = (ctx->image_fmt == AVD_IMG_FMT_420_10BIT ||
+ ctx->image_fmt == AVD_IMG_FMT_422_10BIT) ?
+ 10 :
+ 8;
+
+ mb = DIV_ROUND_UP(w, 16);
+
+ ret = avd_buf_alloc(dev, &h264_ctx->bufs.above_info, mb * 20);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(dev, &h264_ctx->bufs.ip_above, bit_depth * 4 * mb);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(dev, &h264_ctx->bufs.lf_above,
+ bit_depth * 4 * 4 * mb);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(dev, &h264_ctx->bufs.lf_above_info, 32 * mb);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(dev, &h264_ctx->bufs.mv_above_info, 32 * mb);
+ if (ret)
+ return ret;
+
+ return 0;
+}
+
+static int avd_h264_validate_pps(struct avd_ctx *ctx,
+ const struct v4l2_ctrl_h264_pps *pps)
+{
+ if (pps->num_slice_groups_minus1 != 0) {
+ dev_err(ctx->dev->dev, "pps->num_slice_groups_minus1 != 0");
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
+static int avd_h264_validate_sps(struct avd_ctx *ctx,
+ const struct v4l2_ctrl_h264_sps *sps)
+{
+ if (sps->chroma_format_idc > 2)
+ /* Only 4:0:0, 4:2:0 and 4:2:2 are supported */
+ return -EINVAL;
+ if (sps->bit_depth_luma_minus8 != sps->bit_depth_chroma_minus8)
+ /* Luma and chroma bit depth mismatch */
+ return -EINVAL;
+ if (!(sps->flags & V4L2_H264_SPS_FLAG_FRAME_MBS_ONLY))
+ /* no interlaced support */
+ return -EINVAL;
+
+ return 0;
+}
+
+static void avd_h264_stop(struct avd_ctx *ctx)
+{
+ struct avd_h264_ctx *h264_ctx = ctx->priv;
+ struct avd_dev *dev = ctx->dev;
+ int i;
+
+ if (!h264_ctx)
+ return;
+
+ avd_buf_free(dev, &h264_ctx->bufs.above_info);
+ avd_buf_free(dev, &h264_ctx->bufs.lf_above_info);
+ avd_buf_free(dev, &h264_ctx->bufs.lf_above);
+ avd_buf_free(dev, &h264_ctx->bufs.ip_above);
+ avd_buf_free(dev, &h264_ctx->bufs.mv_above_info);
+
+ for (i = 0; i < h264_ctx->slice_num; i++)
+ avd_buf_free(dev, &h264_ctx->slices[i]);
+
+ kfree(h264_ctx->slices);
+ kfree(h264_ctx);
+}
+
+static int avd_h264_start(struct avd_ctx *ctx)
+{
+ struct avd_h264_ctx *h264_ctx;
+ struct v4l2_ctrl *ctrl;
+ int ret;
+
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl, V4L2_CID_STATELESS_H264_SPS);
+ if (!ctrl)
+ return -EINVAL;
+
+ ret = avd_h264_validate_sps(ctx, ctrl->p_new.p_h264_sps);
+ if (ret)
+ return ret;
+
+ h264_ctx = kzalloc_obj(*h264_ctx, GFP_KERNEL);
+ if (!h264_ctx)
+ return -ENOMEM;
+
+ ctx->priv = h264_ctx;
+
+ /* assume one slice per row */
+ h264_ctx->alloc_slice_num =
+ ctrl->p_new.p_h264_sps->pic_height_in_map_units_minus1 + 1;
+ h264_ctx->slices = kzalloc_objs(*h264_ctx->slices,
+ h264_ctx->alloc_slice_num, GFP_KERNEL);
+ if (!h264_ctx->slices)
+ goto err_free_ctx;
+
+ ret = avd_h264_alloc_bufs(ctx);
+ if (ret)
+ goto err_free_ctx;
+
+ return 0;
+
+err_free_ctx:
+ avd_h264_stop(ctx);
+ ctx->priv = NULL;
+ return ret;
+}
+
+static void avd_h264_run_preamble(struct avd_ctx *ctx, struct avd_h264_run *run)
+{
+ struct v4l2_ctrl *ctrl;
+ u32 dst_len, mv_color_len;
+
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_H264_DECODE_PARAMS);
+ run->decode_params = ctrl ? ctrl->p_cur.p : NULL;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl, V4L2_CID_STATELESS_H264_SPS);
+ run->sps = ctrl ? ctrl->p_cur.p : NULL;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl, V4L2_CID_STATELESS_H264_PPS);
+ run->pps = ctrl ? ctrl->p_cur.p : NULL;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_H264_SLICE_PARAMS);
+ run->slice_params = ctrl ? ctrl->p_cur.p : NULL;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_H264_SCALING_MATRIX);
+ run->scaling_matrix = ctrl ? ctrl->p_cur.p : NULL;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_H264_PRED_WEIGHTS);
+ run->pred_weights = ctrl ? ctrl->p_cur.p : NULL;
+
+ avd_run_preamble(ctx, &run->base);
+
+ dst_len = run->base.bufs.dst->vb2_buf.planes[0].length;
+
+ mv_color_len = mv_color_size(fmt_width(ctx), fmt_height(ctx));
+
+ run->addresses.mv_color = run->base.y_out + (dst_len - mv_color_len);
+}
+
+static int avd_h264_realloc_slices(struct avd_ctx *ctx)
+{
+ struct avd_h264_ctx *h264_ctx = ctx->priv;
+ struct avd_job *job = &ctx->job;
+ struct avd_buf tmp_buf = {};
+ size_t alloc_slice_num;
+ void *tmp;
+ int ret;
+
+ tmp = h264_ctx->slices;
+ alloc_slice_num = (h264_ctx->alloc_slice_num * 3) / 2;
+ h264_ctx->slices =
+ kzalloc_objs(*h264_ctx->slices, alloc_slice_num, GFP_KERNEL);
+ if (!h264_ctx->slices) {
+ /* make deallocating a little easier */
+ h264_ctx->slices = tmp;
+ return -ENOMEM;
+ }
+
+ h264_ctx->alloc_slice_num = alloc_slice_num;
+ memcpy(h264_ctx->slices, tmp,
+ sizeof(*h264_ctx->slices) * h264_ctx->slice_num);
+ kfree(tmp);
+
+ /* we could already have the correct size */
+ if ((alloc_slice_num + 1) * sizeof(*job->segments) < job->buf.size)
+ return 0;
+
+ ret = avd_buf_alloc(ctx->dev, &tmp_buf,
+ (alloc_slice_num + 1) * sizeof(*job->segments));
+ if (ret)
+ return -ENOMEM;
+
+ job->segments = tmp_buf.cpu;
+ job->num_alloc = (alloc_slice_num + 1);
+ memset(tmp_buf.cpu, 0, tmp_buf.size);
+ memcpy(job->segments, job->buf.cpu,
+ sizeof(*job->segments) * (job->num + 1));
+
+ avd_buf_free(ctx->dev, &job->buf);
+ memcpy(&job->buf, &tmp_buf, sizeof(tmp_buf));
+ return 0;
+}
+
+static int avd_h264_run(struct avd_ctx *ctx)
+{
+ struct avd_h264_ctx *h264_ctx = ctx->priv;
+ struct v4l2_h264_reflist_builder reflist_builder;
+ struct avd_h264_run run;
+ struct vb2_v4l2_buffer *src;
+ void *slice;
+ int ret;
+
+ avd_h264_run_preamble(ctx, &run);
+
+ src = v4l2_m2m_next_src_buf(ctx->fh.m2m_ctx);
+ slice = vb2_plane_vaddr(&src->vb2_buf, 0);
+ if (!slice) {
+ dev_err_ratelimited(ctx->dev->dev, "could not get slice vaddr");
+ ret = -EINVAL;
+ goto postamble;
+ }
+
+ if (ctx->job.segments &&
+ h264_ctx->slice_num >= h264_ctx->alloc_slice_num) {
+ ret = avd_h264_realloc_slices(ctx);
+ if (ret)
+ goto postamble;
+ }
+
+ h264_ctx->active_slice = &h264_ctx->slices[h264_ctx->slice_num];
+ ret = avd_buf_alloc(ctx->dev, h264_ctx->active_slice,
+ vb2_get_plane_payload(&src->vb2_buf, 0));
+ if (ret)
+ goto postamble;
+
+ memcpy(h264_ctx->active_slice->cpu, slice,
+ h264_ctx->active_slice->size);
+ h264_ctx->slice_num++;
+
+ /* Build the P/B{0,1} ref lists. */
+ v4l2_h264_init_reflist_builder(&reflist_builder, run.decode_params,
+ run.sps, run.decode_params->dpb);
+
+ run.num_valid = reflist_builder.num_valid;
+ run.cur_poc = reflist_builder.cur_pic_order_count;
+
+ v4l2_h264_build_p_ref_list(&reflist_builder, h264_ctx->reflists.p);
+ v4l2_h264_build_b_ref_lists(&reflist_builder, h264_ctx->reflists.b0,
+ h264_ctx->reflists.b1);
+
+
+ if (is_new_frame(run.slice_params)) {
+ ret = avd_init_job(ctx, AVD_CODEC_H264,
+ h264_ctx->alloc_slice_num + 1);
+ if (ret)
+ goto postamble;
+ stream_hdr(ctx, &run);
+ /* h264 only submits once */
+ avd_end_segment(ctx, true);
+ }
+
+ if (!ctx->job.segments) {
+ ret = -EINVAL;
+ goto postamble;
+ }
+
+ stream_slice(ctx, &run);
+ avd_end_segment(ctx, false);
+
+ if (run.base.bufs.src->flags & V4L2_BUF_FLAG_M2M_HOLD_CAPTURE_BUF) {
+ avd_run_postamble(ctx, &run.base);
+ avd_job_finish(ctx, VB2_BUF_STATE_DONE);
+ return 0;
+ }
+
+ ret = avd_submit_job(ctx);
+
+postamble:
+ avd_run_postamble(ctx, &run.base);
+ return ret;
+}
+
+static void avd_h264_done(struct avd_ctx *ctx, struct vb2_v4l2_buffer *src_buf,
+ struct vb2_v4l2_buffer *dst_buf,
+ enum vb2_buffer_state result)
+{
+ struct avd_dev *avd = ctx->dev;
+ struct avd_h264_ctx *h264_ctx = ctx->priv;
+ int i;
+
+ if (!(src_buf->flags & V4L2_BUF_FLAG_M2M_HOLD_CAPTURE_BUF)) {
+ for (i = 0; i < h264_ctx->slice_num; i++)
+ avd_buf_free(avd, &h264_ctx->slices[i]);
+ h264_ctx->slice_num = 0;
+ ctx->job.segments = NULL;
+ }
+}
+
+static enum avd_image_fmt avd_h264_get_image_fmt(struct avd_ctx *ctx,
+ struct v4l2_ctrl *ctrl)
+{
+ const struct v4l2_ctrl_h264_sps *sps = ctrl->p_new.p_h264_sps;
+
+ if (ctrl->id != V4L2_CID_STATELESS_H264_SPS)
+ return AVD_IMG_FMT_ANY;
+
+ if (sps->bit_depth_luma_minus8 == 0) {
+ if (sps->chroma_format_idc == 2)
+ return AVD_IMG_FMT_422_8BIT;
+ else
+ return AVD_IMG_FMT_420_8BIT;
+ } else if (sps->bit_depth_luma_minus8 == 2) {
+ if (sps->chroma_format_idc == 2)
+ return AVD_IMG_FMT_422_10BIT;
+ else
+ return AVD_IMG_FMT_420_10BIT;
+ }
+
+ return AVD_IMG_FMT_ANY;
+}
+
+static void avd_h264_adjust_decoded_fmt(struct avd_ctx *ctx,
+ struct v4l2_pix_format_mplane *pix_mp)
+{
+ pix_mp->plane_fmt[0].sizeimage +=
+ mv_color_size(pix_mp->width, pix_mp->height);
+}
+
+static int avd_h264_try_ctrl(struct avd_ctx *ctx, struct v4l2_ctrl *ctrl)
+{
+ if (ctrl->id == V4L2_CID_STATELESS_H264_SPS)
+ return avd_h264_validate_sps(ctx, ctrl->p_new.p_h264_sps);
+ if (ctrl->id == V4L2_CID_STATELESS_H264_PPS)
+ return avd_h264_validate_pps(ctx, ctrl->p_new.p_h264_pps);
+
+ return 0;
+}
+
+const struct avd_coded_fmt_ops avd_h264_fmt_ops = {
+ .adjust_decoded_fmt = avd_h264_adjust_decoded_fmt,
+ .start = avd_h264_start,
+ .stop = avd_h264_stop,
+ .done = avd_h264_done,
+ .run = avd_h264_run,
+ .try_ctrl = avd_h264_try_ctrl,
+ .get_image_fmt = avd_h264_get_image_fmt,
+};
diff --git a/drivers/media/platform/apple/avd/avd-v4l2.c b/drivers/media/platform/apple/avd/avd-v4l2.c
index 88079da5ee27..e6b9b5f79b4c 100644
--- a/drivers/media/platform/apple/avd/avd-v4l2.c
+++ b/drivers/media/platform/apple/avd/avd-v4l2.c
@@ -201,7 +201,77 @@ const struct v4l2_ctrl_ops avd_ctrl_ops = {
.s_ctrl = avd_s_ctrl,
};
-static const struct avd_coded_fmt_desc avd_coded_fmts[] = {};
+static const struct avd_ctrl_desc avd_h264_ctrl_descs[] = {
+ {
+ .cfg.id = V4L2_CID_STATELESS_H264_DECODE_PARAMS,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_H264_SPS,
+ .cfg.ops = &avd_ctrl_ops,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_H264_PPS,
+ .cfg.ops = &avd_ctrl_ops,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_H264_SCALING_MATRIX,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_H264_PRED_WEIGHTS,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_H264_SLICE_PARAMS,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_H264_DECODE_MODE,
+ .cfg.min = V4L2_STATELESS_H264_DECODE_MODE_SLICE_BASED,
+ .cfg.max = V4L2_STATELESS_H264_DECODE_MODE_SLICE_BASED,
+ .cfg.def = V4L2_STATELESS_H264_DECODE_MODE_SLICE_BASED,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_H264_START_CODE,
+ .cfg.min = V4L2_STATELESS_H264_START_CODE_NONE,
+ .cfg.max = V4L2_STATELESS_H264_START_CODE_NONE,
+ .cfg.def = V4L2_STATELESS_H264_START_CODE_NONE,
+ },
+ {
+ .cfg.id = V4L2_CID_MPEG_VIDEO_H264_PROFILE,
+ .cfg.min = V4L2_MPEG_VIDEO_H264_PROFILE_CONSTRAINED_BASELINE,
+ .cfg.max = V4L2_MPEG_VIDEO_H264_PROFILE_HIGH_422_INTRA,
+ .cfg.menu_skip_mask =
+ BIT(V4L2_MPEG_VIDEO_H264_PROFILE_EXTENDED) |
+ BIT(V4L2_MPEG_VIDEO_H264_PROFILE_HIGH_444_PREDICTIVE),
+ .cfg.def = V4L2_MPEG_VIDEO_H264_PROFILE_MAIN,
+ },
+ {
+ .cfg.id = V4L2_CID_MPEG_VIDEO_H264_LEVEL,
+ .cfg.min = V4L2_MPEG_VIDEO_H264_LEVEL_1_0,
+ .cfg.max = V4L2_MPEG_VIDEO_H264_LEVEL_5_1,
+ },
+};
+
+static const struct avd_ctrls avd_h264_ctrls = {
+ .ctrls = avd_h264_ctrl_descs,
+ .num_ctrls = ARRAY_SIZE(avd_h264_ctrl_descs),
+};
+
+static const struct avd_coded_fmt_desc avd_coded_fmts[] = {
+ {
+ .fourcc = V4L2_PIX_FMT_H264_SLICE,
+ .frmsize = {
+ .min_width = 64,
+ .max_width = 16384,
+ .step_width = 64,
+ .min_height = 64,
+ .max_height = 16384,
+ .step_height = 16,
+ },
+ .ctrls = &avd_h264_ctrls,
+ .ops = &avd_h264_fmt_ops,
+ .subsystem_flags = VB2_V4L2_FL_SUPPORTS_M2M_HOLD_CAPTURE_BUF,
+ .capability = AVD_CAPABILITY_H264,
+ },
+};
static bool avd_is_capable(struct avd_ctx *ctx, unsigned int capability)
{
diff --git a/drivers/media/platform/apple/avd/avd.h b/drivers/media/platform/apple/avd/avd.h
index 980ecce4f3df..a6e4154d2fac 100644
--- a/drivers/media/platform/apple/avd/avd.h
+++ b/drivers/media/platform/apple/avd/avd.h
@@ -242,6 +242,8 @@ void avd_job_finish(struct avd_ctx *ctx, enum vb2_buffer_state result);
void avd_run_preamble(struct avd_ctx *ctx, struct avd_run *run);
void avd_run_postamble(struct avd_ctx *ctx, struct avd_run *run);
+extern const struct avd_coded_fmt_ops avd_h264_fmt_ops;
+
extern const struct v4l2_ctrl_ops avd_ctrl_ops;
extern const struct v4l2_ioctl_ops avd_ioctl_ops;
extern const struct vb2_ops avd_queue_ops;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 08/17] media: apple: avd: add vp9 support
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (6 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 07/17] media: apple: avd: add h264 support Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 09/17] media: apple: avd: add hevc support Sofus Forstreuter
` (8 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
The fluster scores is 216/305 for VP9-TEST-VECTORS and 1/6 for
VP9-TEST-VECTORS-HIGH.
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
drivers/media/platform/apple/avd/Kconfig | 1 +
drivers/media/platform/apple/avd/Makefile | 2 +-
drivers/media/platform/apple/avd/avd-v4l2.c | 37 +
drivers/media/platform/apple/avd/avd-vp9.c | 1054 +++++++++++++++++++++++++++
drivers/media/platform/apple/avd/avd.h | 10 +
5 files changed, 1103 insertions(+), 1 deletion(-)
diff --git a/drivers/media/platform/apple/avd/Kconfig b/drivers/media/platform/apple/avd/Kconfig
index cad7cfa72451..311a121f90f5 100644
--- a/drivers/media/platform/apple/avd/Kconfig
+++ b/drivers/media/platform/apple/avd/Kconfig
@@ -10,6 +10,7 @@ config VIDEO_APPLE_AVD
select V4L2_MEM2MEM_DEV
select VIDEOBUF2_DMA_CONTIG
select V4L2_H264
+ select V4L2_VP9
help
Support for hardware video decoding on Apple Silicon devices using
the Apple Video Decoder (AVD).
diff --git a/drivers/media/platform/apple/avd/Makefile b/drivers/media/platform/apple/avd/Makefile
index 482d43f8ea8e..26763a13edb2 100644
--- a/drivers/media/platform/apple/avd/Makefile
+++ b/drivers/media/platform/apple/avd/Makefile
@@ -1,4 +1,4 @@
# SPDX-License-Identifier: GPL-2.0-only
-apple-avd-y := avd-drv.o avd-v4l2.o avd-h264.o
+apple-avd-y := avd-drv.o avd-v4l2.o avd-h264.o avd-vp9.o
obj-$(CONFIG_VIDEO_APPLE_AVD) += apple-avd.o
diff --git a/drivers/media/platform/apple/avd/avd-v4l2.c b/drivers/media/platform/apple/avd/avd-v4l2.c
index e6b9b5f79b4c..c7958cf65ebd 100644
--- a/drivers/media/platform/apple/avd/avd-v4l2.c
+++ b/drivers/media/platform/apple/avd/avd-v4l2.c
@@ -255,6 +255,29 @@ static const struct avd_ctrls avd_h264_ctrls = {
.num_ctrls = ARRAY_SIZE(avd_h264_ctrl_descs),
};
+static const struct avd_ctrl_desc avd_vp9_ctrl_descs[] = {
+ {
+ .cfg.id = V4L2_CID_STATELESS_VP9_FRAME,
+ .cfg.ops = &avd_ctrl_ops,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_VP9_COMPRESSED_HDR,
+ },
+ {
+ .cfg.id = V4L2_CID_MPEG_VIDEO_VP9_PROFILE,
+ .cfg.min = V4L2_MPEG_VIDEO_VP9_PROFILE_0,
+ .cfg.max = V4L2_MPEG_VIDEO_VP9_PROFILE_2,
+ .cfg.menu_skip_mask =
+ BIT(V4L2_MPEG_VIDEO_VP9_PROFILE_1),
+ .cfg.def = V4L2_MPEG_VIDEO_VP9_PROFILE_0,
+ },
+};
+
+static const struct avd_ctrls avd_vp9_ctrls = {
+ .ctrls = avd_vp9_ctrl_descs,
+ .num_ctrls = ARRAY_SIZE(avd_vp9_ctrl_descs),
+};
+
static const struct avd_coded_fmt_desc avd_coded_fmts[] = {
{
.fourcc = V4L2_PIX_FMT_H264_SLICE,
@@ -271,6 +294,20 @@ static const struct avd_coded_fmt_desc avd_coded_fmts[] = {
.subsystem_flags = VB2_V4L2_FL_SUPPORTS_M2M_HOLD_CAPTURE_BUF,
.capability = AVD_CAPABILITY_H264,
},
+ {
+ .fourcc = V4L2_PIX_FMT_VP9_FRAME,
+ .frmsize = {
+ .min_width = 64,
+ .max_width = 16384,
+ .step_width = 64,
+ .min_height = 64,
+ .max_height = 16384,
+ .step_height = 16,
+ },
+ .ctrls = &avd_vp9_ctrls,
+ .ops = &avd_vp9_fmt_ops,
+ .capability = AVD_CAPABILITY_VP9,
+ },
};
static bool avd_is_capable(struct avd_ctx *ctx, unsigned int capability)
diff --git a/drivers/media/platform/apple/avd/avd-vp9.c b/drivers/media/platform/apple/avd/avd-vp9.c
new file mode 100644
index 000000000000..f7c0ee6b7012
--- /dev/null
+++ b/drivers/media/platform/apple/avd/avd-vp9.c
@@ -0,0 +1,1054 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Apple Video Decoder VP9 driver
+ *
+ * Copyright (C) 2026 The Asahi Linux Contributors
+ * Copyright (C) 2026 Sofus Forstreuter <sofus.c@icloud.com>
+ * Copyright (C) 2023 Eileen Yoon <eyn@gmx.com>
+ *
+ * Copyright (C) 2019 Collabora, Ltd.
+ * Boris Brezillon <boris.brezillon@collabora.com>
+ * Copyright (C) 2021 Collabora, Ltd.
+ * Andrzej Pietrasiewicz <andrzej.p@collabora.com>
+ *
+ * Copyright (C) 2016 Rockchip Electronics Co., Ltd.
+ * Alpha Lin <Alpha.Lin@rock-chips.com>
+ */
+
+#include <linux/dev_printk.h>
+#include <linux/unaligned.h>
+
+#include <media/v4l2-vp9.h>
+#include <media/videobuf2-dma-contig.h>
+
+#include "avd.h"
+#include "avd-inst.h"
+
+#define VP9_Q_IDX(v) FIELD_PREP(GENMASK(31, 15), v)
+#define VP9_Q_DC_Y(v) FIELD_PREP(GENMASK(14, 10), v)
+#define VP9_Q_DC_UV(v) FIELD_PREP(GENMASK(9, 5), v)
+#define VP9_Q_AC_UV(v) FIELD_PREP(GENMASK(4, 0), v)
+
+#define VP9_LF_SHARPNESS(v) FIELD_PREP(GENMASK(31, 28), v)
+#define VP9_LF_REF0(v) FIELD_PREP(GENMASK(27, 21), v)
+#define VP9_LF_REF1(v) FIELD_PREP(GENMASK(20, 14), v)
+#define VP9_LF_REF2(v) FIELD_PREP(GENMASK(13, 7), v)
+#define VP9_LF_REF3(v) FIELD_PREP(GENMASK(6, 0), v)
+
+#define VP9_LF_LV(v) FIELD_PREP(GENMASK(31, 14), v)
+#define VP9_LF_MODE0(v) FIELD_PREP(GENMASK(13, 7), v)
+#define VP9_LF_MODE1(v) FIELD_PREP(GENMASK(6, 0), v)
+
+#define VP9_FEAT_LVL_ALT_Q_EN(v) FIELD_PREP(BIT(21), !!(v))
+#define VP9_FEAT_LVL_ALT_Q(v) FIELD_PREP(GENMASK(20, 12), v)
+#define VP9_FEAT_LVL_ALT_L_EN(v) FIELD_PREP(BIT(11), !!(v))
+#define VP9_FEAT_LVL_ALT_L(v) FIELD_PREP(GENMASK(10, 4), v)
+#define VP9_FEAT_LVL_REF_FRAME_EN(v) FIELD_PREP(BIT(3), !!(v))
+#define VP9_FEAT_LVL_REF_FRAME(v) FIELD_PREP(GENMASK(2, 1), v)
+#define VP9_FEAT_LVL_SKIP_EN(v) FIELD_PREP(BIT(0), !!(v))
+#define VP9_FEAT_LVL_SKIP(v)
+
+#define VP9_REF_SEL_LAST(v) FIELD_PREP(GENMASK(2, 0), v)
+#define VP9_REF_BIAS_LAST(v) FIELD_PREP(BIT(3), !!(v))
+#define VP9_REF_SEL_GOLDEN(v) FIELD_PREP(GENMASK(6, 4), v)
+#define VP9_REF_BIAS_GOLDEN(v) FIELD_PREP(BIT(7), !!(v))
+#define VP9_REF_SEL_ALT(v) FIELD_PREP(GENMASK(10, 8), v)
+#define VP9_REF_BIAS_ALT(v) FIELD_PREP(BIT(11), !!(v))
+#define VP9_REFERENCE_MODE(v) FIELD_PREP(GENMASK(13, 12), v)
+#define VP9_PARALLEL_DEC(v) FIELD_PREP(BIT(14), !!(v))
+#define VP9_REFRESH_CTX(v) FIELD_PREP(BIT(15), !!(v))
+#define VP9_INTERP_FILTER(v) FIELD_PREP(GENMASK(18, 16), v)
+#define VP9_HIGH_PREC_MV(v) FIELD_PREP(BIT(19), !!(v))
+#define VP9_ERR_RES(v) FIELD_PREP(BIT(20), !!(v))
+#define VP9_HAS_REF(v) FIELD_PREP(BIT(21), !!(v))
+#define VP9_SEG_ABS(v) FIELD_PREP(BIT(22), !!(v))
+#define VP9_SEG_UDATA_TEMP(v) FIELD_PREP(BIT(23), !!(v))
+#define VP9_SEG_UPDATE_MAP(v) FIELD_PREP(BIT(24), !!(v))
+#define VP9_SEG_ENABLED(v) FIELD_PREP(BIT(25), !!(v))
+#define VP9_SEG_RESET(v) FIELD_PREP(BIT(26), !!(v))
+
+#define VP9_MAX_TILE_COLS (1 << 6)
+#define VP9_REF_SCALE_SHIFT 14
+#define VP9_LAST_FRAME 1
+#define VP9_GOLDEN_FRAME 2
+#define VP9_ALTREF_FRAME 3
+
+struct avd_vp9_seg_probs {
+ u8 tree_probs[7];
+ u8 pred_probs[3];
+};
+
+struct avd_vp9_probs {
+ struct avd_vp9_seg_probs seg;
+ u8 tx8[2][1];
+ u8 tx16[2][2];
+ u8 tx32[2][3];
+ /* [4][2][2][k=6][(k == 0) ? 3 : 6][3] */
+ u8 coef[1584];
+ u8 skip[3];
+ u8 inter_mode[7][3];
+ u8 interp_filter[4][2];
+ u8 is_inter[4];
+ u8 comp_mode[5];
+ u8 single_ref[5][2];
+ u8 comp_ref[5];
+ u8 y_mode[4][9];
+ u8 uv_mode[10][9];
+ u8 partition[16][3];
+ u8 joint[3];
+ struct mv_comp {
+ u8 sign;
+ u8 classes[10];
+ u8 class0_bit;
+ u8 bits[10];
+ } mv_comp[2];
+ struct mv_fr {
+ u8 class0_fr[2][3];
+ u8 fr[3];
+ } mv_fr[2];
+ struct mv_hp {
+ u8 class0_hp;
+ u8 hp;
+ } mv_hp[2];
+};
+
+struct avd_vp9_frame_symbol_counts {
+ u32 padding;
+ u32 tx8p[2][2];
+ u32 tx16p[2][3];
+ u32 tx32p[2][4];
+ /* [4][2][2][k=6][(k == 0) ? 3 : 6] */
+ u32 eob_0[528];
+ /*
+ * struct ref_cnt { u32 coef[3]; u32 eob_1; }
+ * struct ref_cnt cef_counts[4][2][2][k=6][(k == 0) ? 3 : 6];
+ */
+ u32 ref_cnt[2112];
+ u32 skip[3][2];
+ u32 mv_mode[7][4];
+ u32 filter[4][3];
+ u32 intra_inter[4][2];
+ u32 comp[5][2];
+ u32 single_ref[5][2][2];
+ u32 comp_ref[5][2];
+ u32 y_mode[4][10];
+ u32 uv_mode[10][10];
+ u32 partition[16][4];
+ u32 mv_joint[4];
+ struct mv_comp_ctn {
+ u32 sign[2];
+ u32 classes[11];
+ u32 class0[2];
+ u32 bits[10][2];
+ } mv_comp[2];
+ struct mv_fr_cnt {
+ u32 class0_fr[2][4];
+ u32 fr[4];
+ } mv_fr[2];
+ struct mv_hp_cnt {
+ u32 class0_hp[2];
+ u32 hp[2];
+ } mv_hp[2];
+};
+
+struct avd_vp9_frame_info {
+ u32 valid : 1;
+ u32 segmapid : 1;
+ u32 frame_context_idx : 2;
+ u32 reference_mode : 2;
+ u32 tx_mode : 3;
+ u32 interpolation_filter : 3;
+ u32 flags;
+ u64 timestamp;
+ struct v4l2_vp9_segmentation seg;
+ struct v4l2_vp9_loop_filter lf;
+};
+
+struct avd_vp9_run {
+ struct avd_run base;
+
+ const struct v4l2_ctrl_vp9_frame *decode_params;
+ const struct v4l2_ctrl_vp9_compressed_hdr *prob_updates;
+};
+
+struct avd_vp9_ctx {
+ struct v4l2_vp9_frame_symbol_counts cnts;
+ struct v4l2_vp9_frame_context probability_tables;
+ struct v4l2_vp9_frame_context frame_context[4];
+ struct avd_vp9_bufs {
+ struct avd_buf az_left;
+ struct avd_buf above_info;
+ struct avd_buf ip_above;
+ struct avd_buf lf_above;
+ struct avd_buf lf_left_info;
+ struct avd_buf lf_left;
+ struct avd_buf color;
+ struct avd_buf seg;
+ struct avd_buf counts;
+ struct avd_buf probs;
+ } bufs;
+ struct avd_vp9_frame_info cur;
+ struct avd_vp9_frame_info last;
+};
+
+static void set_refs(struct avd_ctx *ctx, struct avd_vp9_run *run)
+{
+ const struct v4l2_ctrl_vp9_frame *frame = run->decode_params;
+ struct avd_decoded_buffer *dst, *ref_buf[4];
+ dma_addr_t addr;
+ int xscale, yscale;
+
+ dst = vb2_to_avd_decoded_buf(&run->base.bufs.dst->vb2_buf);
+
+ ref_buf[0] = avd_get_ref_buf(ctx, &dst->base.vb, frame->last_frame_ts);
+ ref_buf[1] =
+ avd_get_ref_buf(ctx, &dst->base.vb, frame->golden_frame_ts);
+ ref_buf[2] = avd_get_ref_buf(ctx, &dst->base.vb, frame->alt_frame_ts);
+
+ push(0, "");
+ push(0, "");
+ push(0, "");
+
+ for (int i = 0; i < V4L2_VP9_NUM_FRAME_CTX - 1; i++) {
+ addr = vb2_dma_contig_plane_dma_addr(
+ &ref_buf[i]->base.vb.vb2_buf, 0) +
+ ref_buf[i]->comp.start_offset;
+
+ push(AVD_REF_FLAG_CONST, "hdr_9c_ref_100");
+ push(AVD_HDR_HEIGHT(ref_buf[i]->vp9.height - 1) |
+ AVD_HDR_WIDTH(ref_buf[i]->vp9.width - 1),
+ "hdr_70_ref_height_width");
+
+ xscale = (ref_buf[i]->vp9.width << VP9_REF_SCALE_SHIFT) /
+ dst->vp9.width;
+ yscale = (ref_buf[i]->vp9.height << VP9_REF_SCALE_SHIFT) /
+ dst->vp9.height;
+ push(AVD_HDR_HEIGHT(xscale) | AVD_HDR_WIDTH(yscale),
+ "ref_scale");
+
+ push_comp(ctx, addr, ref_buf[i]->comp.offsets);
+ }
+}
+
+static u32 make_flags1(struct avd_ctx *ctx, struct avd_vp9_run *run)
+{
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ const struct v4l2_ctrl_vp9_frame *frame = run->decode_params;
+ struct avd_decoded_buffer *dst, *last, *golden, *alt;
+
+ dst = vb2_to_avd_decoded_buf(&run->base.bufs.dst->vb2_buf);
+ last = avd_get_ref_buf(ctx, &dst->base.vb, frame->last_frame_ts);
+ golden = avd_get_ref_buf(ctx, &dst->base.vb, frame->golden_frame_ts);
+ alt = avd_get_ref_buf(ctx, &dst->base.vb, frame->alt_frame_ts);
+ u32 flags;
+
+ /*
+ * i have only seen V4L2_VP9_SIGN_BIAS_ALT, im simply guessing the
+ * same applies for last and golden
+ */
+ flags = VP9_REF_SEL_LAST(VP9_LAST_FRAME);
+
+ flags |= VP9_REF_BIAS_LAST(frame->ref_frame_sign_bias &
+ V4L2_VP9_SIGN_BIAS_LAST);
+ flags |= VP9_REF_SEL_GOLDEN(golden == last ? VP9_LAST_FRAME :
+ VP9_GOLDEN_FRAME);
+ flags |= VP9_REF_BIAS_GOLDEN(frame->ref_frame_sign_bias &
+ V4L2_VP9_SIGN_BIAS_GOLDEN);
+ flags |= VP9_REF_SEL_ALT(alt != golden ? VP9_ALTREF_FRAME :
+ alt != last ? VP9_GOLDEN_FRAME :
+ VP9_LAST_FRAME);
+ flags |= VP9_REF_BIAS_ALT(frame->ref_frame_sign_bias &
+ V4L2_VP9_SIGN_BIAS_ALT);
+
+ flags |= VP9_REFERENCE_MODE(frame->reference_mode);
+
+ flags |= VP9_PARALLEL_DEC(frame->flags &
+ V4L2_VP9_FRAME_FLAG_PARALLEL_DEC_MODE);
+ flags |= VP9_REFRESH_CTX(frame->flags &
+ V4L2_VP9_FRAME_FLAG_REFRESH_FRAME_CTX);
+ flags |= VP9_INTERP_FILTER(frame->interpolation_filter);
+ flags |= VP9_HIGH_PREC_MV(frame->flags &
+ V4L2_VP9_FRAME_FLAG_ALLOW_HIGH_PREC_MV);
+
+ flags |=
+ VP9_ERR_RES(frame->flags & V4L2_VP9_FRAME_FLAG_ERROR_RESILIENT);
+
+ flags |= VP9_SEG_ABS(frame->seg.flags &
+ V4L2_VP9_SEGMENTATION_FLAG_ABS_OR_DELTA_UPDATE);
+ flags |= VP9_SEG_UDATA_TEMP(
+ frame->seg.flags & V4L2_VP9_SEGMENTATION_FLAG_UPDATE_DATA &&
+ frame->seg.flags & V4L2_VP9_SEGMENTATION_FLAG_TEMPORAL_UPDATE);
+ flags |= VP9_SEG_UPDATE_MAP(frame->seg.flags &
+ V4L2_VP9_SEGMENTATION_FLAG_UPDATE_MAP);
+ flags |= VP9_SEG_ENABLED(frame->seg.flags &
+ V4L2_VP9_SEGMENTATION_FLAG_ENABLED);
+
+ flags |= VP9_HAS_REF(
+ !(frame->flags & V4L2_VP9_FRAME_FLAG_ERROR_RESILIENT) &&
+ !(frame->flags & V4L2_VP9_FRAME_FLAG_KEY_FRAME) &&
+ vp9_ctx->last.valid &&
+ vp9_ctx->last.flags & V4L2_VP9_FRAME_FLAG_SHOW_FRAME &&
+ !(vp9_ctx->last.flags & V4L2_VP9_FRAME_FLAG_KEY_FRAME));
+
+ flags |= VP9_SEG_RESET(
+ frame->seg.flags & V4L2_VP9_SEGMENTATION_FLAG_ENABLED &&
+ (frame->flags & V4L2_VP9_FRAME_FLAG_KEY_FRAME ||
+ (vp9_ctx->last.valid &&
+ !(vp9_ctx->last.seg.flags &
+ V4L2_VP9_SEGMENTATION_FLAG_ENABLED) &&
+ vp9_ctx->last.flags & (V4L2_VP9_FRAME_FLAG_KEY_FRAME |
+ V4L2_VP9_FRAME_FLAG_INTRA_ONLY))));
+
+ return flags;
+}
+
+static u32 seg_features(struct avd_ctx *ctx, struct avd_vp9_run *run,
+ unsigned int segid)
+{
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ const struct v4l2_vp9_segmentation *seg = &vp9_ctx->cur.seg;
+ u32 feature_val = 0;
+ int feature_id = 0;
+
+ feature_id = V4L2_VP9_SEG_LVL_ALT_Q;
+ feature_val |= VP9_FEAT_LVL_ALT_Q_EN(v4l2_vp9_seg_feat_enabled(
+ seg->feature_enabled, feature_id, segid));
+ feature_val |= VP9_FEAT_LVL_ALT_Q(seg->feature_data[segid][feature_id]);
+
+ feature_id = V4L2_VP9_SEG_LVL_ALT_L;
+ feature_val |= VP9_FEAT_LVL_ALT_L_EN(v4l2_vp9_seg_feat_enabled(
+ seg->feature_enabled, feature_id, segid));
+ feature_val |= VP9_FEAT_LVL_ALT_L(seg->feature_data[segid][feature_id]);
+
+ feature_id = V4L2_VP9_SEG_LVL_REF_FRAME;
+ feature_val |= VP9_FEAT_LVL_REF_FRAME_EN(v4l2_vp9_seg_feat_enabled(
+ seg->feature_enabled, feature_id, segid));
+ feature_val |=
+ VP9_FEAT_LVL_REF_FRAME(seg->feature_data[segid][feature_id]);
+
+ feature_id = V4L2_VP9_SEG_LVL_SKIP;
+ feature_val |= VP9_FEAT_LVL_SKIP_EN(v4l2_vp9_seg_feat_enabled(
+ seg->feature_enabled, feature_id, segid));
+
+ return feature_val;
+}
+
+static void set_header(struct avd_ctx *ctx, struct avd_vp9_run *run)
+{
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ const struct v4l2_ctrl_vp9_compressed_hdr *prob_updates =
+ run->prob_updates;
+ const struct v4l2_ctrl_vp9_frame *frame = run->decode_params;
+ struct avd_dev *avd = ctx->dev;
+ u32 bytesperline;
+
+ bool intra_only = !!(frame->flags & (V4L2_VP9_FRAME_FLAG_KEY_FRAME |
+ V4L2_VP9_FRAME_FLAG_INTRA_ONLY));
+
+ push(AVD_OP_HDR | AVD_OP_HDR_FLAG_DECOMP(ctx->decomp) |
+ AVD_OP_HDR_FLAG_INTRA(intra_only) | AVD_OP_HDR_CONST |
+ AVD_OP_HDR_FLAG_PIPE_STATE(
+ !(avd->variant->quirks & AVD_QUIRK_NO_PIPE_STATE)),
+ "hdr_34_start_hdr");
+
+ push(AVD_HDR_CODEC_MODE(AVD_CODEC_VP9), "hdr_38_mode");
+
+ push(AVD_HDR_HEIGHT(frame->frame_height_minus_1) |
+ AVD_HDR_WIDTH(frame->frame_width_minus_1),
+ "hdr_28_height_width_shift3");
+ push(0, "");
+ push(AVD_HDR_HEIGHT(frame->frame_height_minus_1) |
+ AVD_HDR_WIDTH(frame->frame_width_minus_1),
+ "hdr_38_height_width_shift3");
+
+ push(AVD_HDR_COMMON_CHROMA_FORMAT(1) |
+ AVD_HDR_COMMON_BIT_DEPTH_C(frame->bit_depth - 8) |
+ AVD_HDR_COMMON_BIT_DEPTH_L(frame->bit_depth - 8) |
+ AVD_HDR_COMMON_LUMA_CBS(3) |
+ AVD_HDR_COMMON_LUMA_TBS(min(prob_updates->tx_mode, 3)) |
+ AVD_HDR_COMMON_FLAG0(prob_updates->tx_mode &
+ V4L2_VP9_TX_MODE_SELECT),
+ "hdr_2c_txfm_mode");
+
+ push(make_flags1(ctx, run), "hdr_40_flags1_pt1");
+
+ for (int i = 0; i < 8; i++)
+ push(seg_features(ctx, run, i), "seg");
+
+ push(AVD_HDR_FEAT_VP9, "unk_const");
+ push(0, "");
+ push(0, "");
+
+ pusha(vp9_ctx->bufs.counts.addr, "counts", 0);
+ pusha(vp9_ctx->bufs.probs.addr, "probs", 0);
+ pusha(vp9_ctx->bufs.above_info.addr, "above_info", 0);
+ pusha(vp9_ctx->bufs.seg.addr, "seg", 1);
+ pusha(vp9_ctx->bufs.seg.addr, "seg", 2);
+ pusha(vp9_ctx->bufs.color.addr, "color", 3);
+ pusha(vp9_ctx->bufs.color.addr, "color", 4);
+
+ push(VP9_Q_IDX(frame->quant.base_q_idx) |
+ VP9_Q_DC_Y(frame->quant.delta_q_y_dc) |
+ VP9_Q_DC_UV(frame->quant.delta_q_uv_dc) |
+ VP9_Q_AC_UV(frame->quant.delta_q_uv_ac),
+ "hdr_4c_base_q_idx");
+ /* filter related flags? */
+ push(VP9_LF_SHARPNESS(frame->lf.sharpness) |
+ (frame->lf.flags &
+ V4L2_VP9_LOOP_FILTER_FLAG_DELTA_ENABLED ?
+ VP9_LF_REF0(frame->lf.ref_deltas[0]) |
+ VP9_LF_REF1(frame->lf.ref_deltas[1]) |
+ VP9_LF_REF2(frame->lf.ref_deltas[2]) |
+ VP9_LF_REF3(frame->lf.ref_deltas[3]) :
+ 0),
+ "hdr_44_flags1_pt2");
+
+ push(VP9_LF_LV(frame->lf.level) |
+ (frame->lf.flags &
+ V4L2_VP9_LOOP_FILTER_FLAG_DELTA_ENABLED ?
+ VP9_LF_MODE0(frame->lf.mode_deltas[0]) |
+ VP9_LF_MODE1(frame->lf.mode_deltas[1]) :
+ 0),
+ "hdr_48_loop_filter_level");
+
+ push(0, "");
+ push(0, "");
+
+ if (avd->variant->revision == 3)
+ push(0, "");
+ if (!(avd->variant->quirks & AVD_QUIRK_NO_PIPE_STATE))
+ pusha(ctx->pipe_state.addr, "pipe_state", 0);
+
+ pusha(vp9_ctx->bufs.ip_above.addr, "ip_above", 0);
+ pusha(vp9_ctx->bufs.lf_above.addr, "lf_above", 0);
+ /* no lf_above_info? */
+ pusha((u64)0, "hdr_e8_sps0_tile_addr_lsb8", 0);
+ pusha(vp9_ctx->bufs.lf_left.addr, "lf_left", 0);
+ pusha(vp9_ctx->bufs.lf_left_info.addr, "lf_left_info", 0);
+ pusha(vp9_ctx->bufs.az_left.addr, "az_left", 0);
+
+ push(0, "");
+
+ push_comp(ctx, run->base.comp_out, ctx->comp.offsets);
+
+ pusha((u64)0, "packed_fmt_scratch", 0);
+
+ bytesperline = ctx->decoded_fmt.fmt.pix_mp.plane_fmt[0].bytesperline;
+ if (avd->variant->quirks & AVD_QUIRK_LSR)
+ bytesperline = bytesperline >> 4;
+
+ pusha(run->base.y_out, "hdr_168_y_addr_lsb8", 0);
+ push(bytesperline, "hdr_170_width_align");
+ pusha(run->base.uv_out, "hdr_16c_uv_addr_lsb8", 0);
+ push(bytesperline, "hdr_174_width_align");
+ push(0, "");
+ push(AVD_HDR_HEIGHT(frame->frame_height_minus_1) |
+ AVD_HDR_WIDTH(frame->frame_width_minus_1),
+ "cm3_height_width");
+
+ if (!(intra_only))
+ set_refs(ctx, run);
+}
+
+static int set_tiles(struct avd_ctx *ctx, struct avd_vp9_run *run)
+{
+ const struct v4l2_ctrl_vp9_frame *frame = run->decode_params;
+ struct vb2_v4l2_buffer *src = v4l2_m2m_next_src_buf(ctx->fh.m2m_ctx);
+ const u8 *data = vb2_plane_vaddr(&src->vb2_buf, 0);
+ unsigned long plane_payload = vb2_get_plane_payload(&src->vb2_buf, 0);
+ u32 offset =
+ frame->uncompressed_header_size + frame->compressed_header_size;
+ u32 num_tile_rows = 1 << frame->tile_rows_log2;
+ u32 num_tile_cols = 1 << frame->tile_cols_log2;
+ u32 tile_size, size;
+ dma_addr_t coded_in;
+
+ if (plane_payload < offset)
+ return -EINVAL;
+ size = plane_payload - offset;
+
+ /* 6.2.6 Compute image size syntax */
+ u32 sb_64_cols = (((frame->frame_width_minus_1 + 8) >> 3) + 7) >> 3;
+ u32 sb_64_rows = (((frame->frame_height_minus_1 + 8) >> 3) + 7) >> 3;
+
+ for (int row = 0; row < num_tile_rows; row++)
+ for (int col = 0; col < num_tile_cols; col++) {
+ if (row == num_tile_rows - 1 &&
+ col == num_tile_cols - 1) {
+ tile_size = size;
+ } else {
+ if (offset > plane_payload)
+ return -EINVAL;
+ tile_size = get_unaligned_be32(&data[offset]);
+ if (tile_size > size - 4)
+ return -EINVAL;
+ offset += 4;
+ size -= 4;
+ }
+
+ coded_in = run->base.coded_in + offset;
+ push(AVD_OP_CODED_DATA | AVD_OP_CODED_IN_HI(coded_in),
+ "cm3_cmd_set_slice_data");
+ push(AVD_OP_CODED_IN_LO(coded_in), "coded_in");
+ push(tile_size, "til_ab8_tile_size");
+ push(AVD_OP_SL_DIM_START |
+ AVD_OP_SL_DIM_START_Y((row * sb_64_rows) /
+ num_tile_rows) |
+ AVD_OP_SL_DIM_START_X((col * sb_64_cols) /
+ num_tile_cols),
+ "i");
+
+ push(AVD_SL_DIM_END_COL(col) |
+ AVD_SL_DIM_END_Y(((row + 1) * sb_64_rows) /
+ num_tile_rows - 1) |
+ AVD_SL_DIM_END_X(((col + 1) * sb_64_cols) /
+ num_tile_cols - 1),
+ "til_ac0_tile_dims");
+
+ offset += tile_size;
+ size -= tile_size;
+ avd_end_segment(ctx, true);
+ }
+
+ return 0;
+}
+
+static void update_dec_buf_info(struct avd_decoded_buffer *buf,
+ const struct v4l2_ctrl_vp9_frame *dec_params)
+{
+ buf->vp9.width = dec_params->frame_width_minus_1 + 1;
+ buf->vp9.height = dec_params->frame_height_minus_1 + 1;
+ buf->vp9.bit_depth = dec_params->bit_depth;
+}
+
+static void update_ctx_cur_info(struct avd_vp9_ctx *vp9_ctx,
+ struct avd_decoded_buffer *buf,
+ const struct v4l2_ctrl_vp9_frame *dec_params)
+{
+ vp9_ctx->cur.valid = true;
+ vp9_ctx->cur.reference_mode = dec_params->reference_mode;
+ vp9_ctx->cur.interpolation_filter = dec_params->interpolation_filter;
+ vp9_ctx->cur.flags = dec_params->flags;
+ vp9_ctx->cur.timestamp = buf->base.vb.vb2_buf.timestamp;
+ vp9_ctx->cur.seg = dec_params->seg;
+ vp9_ctx->cur.lf = dec_params->lf;
+}
+
+static void update_ctx_last_info(struct avd_vp9_ctx *vp9_ctx)
+{
+ vp9_ctx->last = vp9_ctx->cur;
+}
+
+static void copy_vp9_frame_mv(struct avd_vp9_probs *avd_probs,
+ const struct v4l2_vp9_frame_context *probs)
+{
+ memcpy(avd_probs->joint, probs->mv.joint, sizeof(avd_probs->joint));
+ for (int i = 0; i < 2; i++) {
+ avd_probs->mv_comp[i].sign = probs->mv.sign[i];
+ memcpy(avd_probs->mv_comp[i].bits, probs->mv.bits[i],
+ sizeof(avd_probs->mv_comp[i].bits));
+ avd_probs->mv_comp[i].class0_bit = probs->mv.class0_bit[i];
+ memcpy(avd_probs->mv_comp[i].classes, probs->mv.classes[i],
+ sizeof(avd_probs->mv_comp[i].bits));
+
+ memcpy(avd_probs->mv_fr[i].class0_fr, probs->mv.class0_fr[i],
+ sizeof(avd_probs->mv_fr[i].class0_fr));
+ memcpy(avd_probs->mv_fr[i].fr, probs->mv.fr[i],
+ sizeof(avd_probs->mv_fr[i].fr));
+
+ avd_probs->mv_hp[i].class0_hp = probs->mv.class0_hp[i];
+ avd_probs->mv_hp[i].hp = probs->mv.hp[i];
+ }
+}
+
+static void init_probs(struct avd_ctx *ctx, const struct avd_vp9_run *run)
+{
+ const struct v4l2_ctrl_vp9_frame *dec_params;
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ struct avd_vp9_probs *avd_probs = vp9_ctx->bufs.probs.cpu;
+ const struct v4l2_vp9_segmentation *seg;
+ const struct v4l2_vp9_frame_context *probs;
+ bool intra_only;
+ int count = 0;
+
+ dec_params = run->decode_params;
+ probs = &vp9_ctx->probability_tables;
+ seg = &dec_params->seg;
+
+ memset(avd_probs, 0, sizeof(*avd_probs));
+
+ intra_only = !!(dec_params->flags & (V4L2_VP9_FRAME_FLAG_KEY_FRAME |
+ V4L2_VP9_FRAME_FLAG_INTRA_ONLY));
+
+ memcpy(avd_probs->tx8, probs->tx8, sizeof(avd_probs->tx8));
+ memcpy(avd_probs->tx16, probs->tx16, sizeof(avd_probs->tx16));
+ memcpy(avd_probs->tx32, probs->tx32, sizeof(avd_probs->tx32));
+ memcpy(avd_probs->skip, probs->skip, sizeof(avd_probs->skip));
+ memcpy(avd_probs->inter_mode, probs->inter_mode,
+ sizeof(avd_probs->inter_mode));
+ memcpy(avd_probs->interp_filter, probs->interp_filter,
+ sizeof(avd_probs->interp_filter));
+ memcpy(avd_probs->is_inter, probs->is_inter,
+ sizeof(avd_probs->is_inter));
+ memcpy(avd_probs->comp_mode, probs->comp_mode,
+ sizeof(avd_probs->comp_mode));
+ memcpy(avd_probs->single_ref, probs->single_ref,
+ sizeof(avd_probs->single_ref));
+ memcpy(avd_probs->comp_ref, probs->comp_ref,
+ sizeof(avd_probs->comp_ref));
+ memcpy(avd_probs->y_mode, probs->y_mode, sizeof(avd_probs->y_mode));
+
+ memcpy(avd_probs->partition,
+ intra_only ? v4l2_vp9_kf_partition_probs : probs->partition,
+ sizeof(avd_probs->partition));
+ memcpy(avd_probs->uv_mode,
+ intra_only ? v4l2_vp9_kf_uv_mode_prob : probs->uv_mode,
+ sizeof(avd_probs->uv_mode));
+
+ copy_vp9_frame_mv(avd_probs, probs);
+
+ for (int t = 0; t < 4; t++)
+ for (int i = 0; i < 2; i++)
+ for (int j = 0; j < 2; j++)
+ for (int k = 0; k < 6; k++) {
+ int max_l = (k == 0) ? 3 : 6;
+ for (int l = 0; l < max_l; l++) {
+ for (int n = 0; n < 3; n++)
+ avd_probs->coef[count++] =
+ probs->coef[t][i][j][k][l][n];
+ }
+ }
+
+ memcpy(avd_probs->seg.pred_probs, seg->pred_probs,
+ sizeof(avd_probs->seg.pred_probs));
+ memcpy(avd_probs->seg.tree_probs, seg->tree_probs,
+ sizeof(avd_probs->seg.tree_probs));
+}
+
+static int validate_dec_params(struct avd_ctx *ctx,
+ const struct v4l2_ctrl_vp9_frame *dec_params)
+{
+ unsigned int aligned_width, aligned_height;
+
+ if (dec_params->bit_depth > 10)
+ /* not implemented */
+ return -EINVAL;
+
+ if (dec_params->profile == 1 || dec_params->profile > 2)
+ return -EINVAL;
+
+ if (dec_params->frame_height_minus_1 + 1 < 64 ||
+ dec_params->frame_width_minus_1 + 1 < 64)
+ return -EINVAL;
+
+ aligned_width = round_up(dec_params->frame_width_minus_1 + 1, 64);
+ aligned_height = round_up(dec_params->frame_height_minus_1 + 1, 16);
+
+ /*
+ * Userspace should update the capture/decoded format when the
+ * resolution changes.
+ */
+ if (aligned_width != ctx->decoded_fmt.fmt.pix_mp.width ||
+ aligned_height != ctx->decoded_fmt.fmt.pix_mp.height) {
+ dev_err(ctx->dev->dev,
+ "unexpected bitstream resolution %dx%d\n",
+ aligned_width, aligned_height);
+ return -EINVAL;
+ }
+
+ return 0;
+}
+
+static int avd_vp9_alloc_bufs(struct avd_ctx *ctx)
+{
+ struct avd_dev *avd = ctx->dev;
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ int ret, w, h, bit_depth;
+
+ w = fmt_width(ctx);
+ h = fmt_height(ctx);
+ bit_depth = (ctx->image_fmt == AVD_IMG_FMT_420_10BIT ||
+ ctx->image_fmt == AVD_IMG_FMT_422_10BIT) ?
+ 10 :
+ 8;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.probs,
+ sizeof(struct avd_vp9_probs));
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.counts,
+ sizeof(struct avd_vp9_frame_symbol_counts));
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.color,
+ DIV_ROUND_UP(w, 16) * DIV_ROUND_UP(h, 64) * 144);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.ip_above,
+ DIV_ROUND_UP(w, 16) * 4 * bit_depth +
+ (VP9_MAX_TILE_COLS - 1) * 128);
+ if (ret)
+ return ret;
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.lf_above,
+ DIV_ROUND_UP(w, 8) * 16 * bit_depth);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.seg, DIV_ROUND_UP(w, 8) * 24);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.lf_left,
+ DIV_ROUND_UP(h, 8) * 16 * bit_depth);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.lf_left_info,
+ DIV_ROUND_UP(h, 64) * 16);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.az_left,
+ bit_depth * 36 * DIV_ROUND_UP(h, 8) +
+ bit_depth * 18 * DIV_ROUND_UP(h, 16));
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &vp9_ctx->bufs.above_info,
+ DIV_ROUND_UP(w, 64) * 288 +
+ (VP9_MAX_TILE_COLS - 1) * 128);
+ if (ret)
+ return ret;
+
+ return 0;
+}
+
+static int avd_vp9_run_preamble(struct avd_ctx *ctx, struct avd_vp9_run *run)
+{
+ struct v4l2_ctrl *ctrl;
+ const struct v4l2_ctrl_vp9_frame *dec_params;
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ unsigned int fctx_idx;
+ int ret;
+
+ avd_run_preamble(ctx, &run->base);
+
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_VP9_COMPRESSED_HDR);
+ if (WARN_ON(!ctrl))
+ return -EINVAL;
+ run->prob_updates = ctrl->p_cur.p;
+
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl, V4L2_CID_STATELESS_VP9_FRAME);
+ if (WARN_ON(!ctrl))
+ return -EINVAL;
+ dec_params = ctrl->p_cur.p;
+
+ ret = validate_dec_params(ctx, dec_params);
+ if (ret)
+ return ret;
+
+ run->decode_params = dec_params;
+
+ vp9_ctx->cur.tx_mode = run->prob_updates->tx_mode;
+
+ fctx_idx = v4l2_vp9_reset_frame_ctx(dec_params, vp9_ctx->frame_context);
+ vp9_ctx->cur.frame_context_idx = fctx_idx;
+
+ vp9_ctx->probability_tables = vp9_ctx->frame_context[fctx_idx];
+ v4l2_vp9_fw_update_probs(&vp9_ctx->probability_tables,
+ run->prob_updates, dec_params);
+
+ return 0;
+}
+
+static int avd_vp9_run(struct avd_ctx *ctx)
+{
+ struct avd_vp9_run run;
+ struct avd_vp9_ctx *vp9_ctx;
+ struct avd_decoded_buffer *dst;
+ int ret;
+
+ ret = avd_vp9_run_preamble(ctx, &run);
+ if (ret)
+ goto postamble;
+
+ ret = avd_init_job(
+ ctx, AVD_CODEC_VP9,
+ (1 << run.decode_params->tile_rows_log2) *
+ (1 << run.decode_params->tile_cols_log2) +
+ 1);
+ if (ret)
+ goto postamble;
+
+ init_probs(ctx, &run);
+
+ vp9_ctx = ctx->priv;
+ dst = vb2_to_avd_decoded_buf(&run.base.bufs.dst->vb2_buf);
+ update_dec_buf_info(dst, run.decode_params);
+ update_ctx_cur_info(vp9_ctx, dst, run.decode_params);
+
+ set_header(ctx, &run);
+ avd_end_segment(ctx, false);
+ ret = set_tiles(ctx, &run);
+ if (ret)
+ goto postamble;
+
+ ret = avd_submit_job(ctx);
+
+postamble:
+ avd_run_postamble(ctx, &run.base);
+ return ret;
+}
+
+#define copy_tx_and_skip(p1, p2) \
+ do { \
+ memcpy((p1)->tx8, (p2)->tx8, sizeof((p1)->tx8)); \
+ memcpy((p1)->tx16, (p2)->tx16, sizeof((p1)->tx16)); \
+ memcpy((p1)->tx32, (p2)->tx32, sizeof((p1)->tx32)); \
+ memcpy((p1)->skip, (p2)->skip, sizeof((p1)->skip)); \
+ } while (0)
+
+static void avd_vp9_done(struct avd_ctx *ctx, struct vb2_v4l2_buffer *src_buf,
+ struct vb2_v4l2_buffer *dst_buf,
+ enum vb2_buffer_state result)
+{
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ unsigned int fctx_idx;
+
+ /* v4l2-specific stuff */
+ if (result == VB2_BUF_STATE_ERROR)
+ goto out_update_last;
+
+ /*
+ * vp9 stuff
+ *
+ * 6.1.2 refresh_probs()
+ *
+ * In the spec a complementary condition goes last in 6.1.2 refresh_probs(),
+ * but it makes no sense to perform all the activities from the first "if"
+ * there if we actually are not refreshing the frame context. On top of that,
+ * because of 6.2 uncompressed_header() whenever error_resilient_mode == 1,
+ * refresh_frame_context == 0. Consequently, if we don't jump to out_update_last
+ * it means error_resilient_mode must be 0.
+ */
+ if (!(vp9_ctx->cur.flags & V4L2_VP9_FRAME_FLAG_REFRESH_FRAME_CTX))
+ goto out_update_last;
+
+ fctx_idx = vp9_ctx->cur.frame_context_idx;
+
+ if (!(vp9_ctx->cur.flags & V4L2_VP9_FRAME_FLAG_PARALLEL_DEC_MODE)) {
+ /* error_resilient_mode == 0 && frame_parallel_decoding_mode == 0 */
+ struct v4l2_vp9_frame_context *probs =
+ &vp9_ctx->probability_tables;
+ bool frame_is_intra = vp9_ctx->cur.flags &
+ (V4L2_VP9_FRAME_FLAG_KEY_FRAME |
+ V4L2_VP9_FRAME_FLAG_INTRA_ONLY);
+ struct tx_and_skip {
+ u8 tx8[2][1];
+ u8 tx16[2][2];
+ u8 tx32[2][3];
+ u8 skip[3];
+ } _tx_skip, *tx_skip = &_tx_skip;
+ struct v4l2_vp9_frame_symbol_counts *counts;
+
+ /* buffer the forward-updated TX and skip probs */
+ if (frame_is_intra)
+ copy_tx_and_skip(tx_skip, probs);
+
+ /* 6.1.2 refresh_probs(): load_probs() and load_probs2() */
+ *probs = vp9_ctx->frame_context[fctx_idx];
+
+ /* if FrameIsIntra then undo the effect of load_probs2() */
+ if (frame_is_intra)
+ copy_tx_and_skip(probs, tx_skip);
+
+ counts = &vp9_ctx->cnts;
+
+ v4l2_vp9_adapt_coef_probs(
+ probs, counts,
+ !vp9_ctx->last.valid ||
+ vp9_ctx->last.flags &
+ V4L2_VP9_FRAME_FLAG_KEY_FRAME,
+ frame_is_intra);
+ if (!frame_is_intra) {
+ const struct avd_vp9_frame_symbol_counts *cnts;
+ int i;
+ u32 tx16p[2][4];
+ u32 sign[2][2];
+ u32 classes[2][11];
+ u32 class0[2][2];
+ u32 bits[2][10][2];
+ u32 class0_fp[2][2][4];
+ u32 fp[2][4];
+ u32 class0_hp[2][2];
+ u32 hp[2][2];
+
+ cnts = vp9_ctx->bufs.counts.cpu;
+
+ for (i = 0; i < ARRAY_SIZE(cnts->tx16p); ++i)
+ memcpy(tx16p[i], cnts->tx16p[i],
+ sizeof(cnts->tx16p[0]));
+
+ for (i = 0; i < 2; i++) {
+ memcpy(sign[i], cnts->mv_comp[i].sign,
+ sizeof(sign[0]));
+ memcpy(classes[i], cnts->mv_comp[i].classes,
+ sizeof(classes[0]));
+ memcpy(class0[i], cnts->mv_comp[i].class0,
+ sizeof(class0[0]));
+ memcpy(bits[i], cnts->mv_comp[i].bits,
+ sizeof(bits[0]));
+ memcpy(class0_fp[i], cnts->mv_fr[i].class0_fr,
+ sizeof(class0_fp[0]));
+ memcpy(fp[i], cnts->mv_fr[i].fr, sizeof(fp[0]));
+ memcpy(class0_hp[i], cnts->mv_hp[i].class0_hp,
+ sizeof(class0_hp[0]));
+ memcpy(hp[i], cnts->mv_hp[i].hp, sizeof(hp[0]));
+ }
+
+ counts->tx16p = &tx16p;
+ counts->sign = &sign;
+ counts->classes = &classes;
+ counts->class0 = &class0;
+ counts->bits = &bits;
+ counts->class0_fp = &class0_fp;
+ counts->fp = &fp;
+ counts->class0_hp = &class0_hp;
+ counts->hp = &hp;
+
+ /* load_probs2() already done */
+ v4l2_vp9_adapt_noncoef_probs(
+ &vp9_ctx->probability_tables, counts,
+ vp9_ctx->cur.reference_mode,
+ vp9_ctx->cur.interpolation_filter,
+ vp9_ctx->cur.tx_mode, vp9_ctx->cur.flags);
+ }
+ }
+
+ /* 6.1.2 refresh_probs(): save_probs(fctx_idx) */
+ vp9_ctx->frame_context[fctx_idx] = vp9_ctx->probability_tables;
+
+out_update_last:
+ update_ctx_last_info(vp9_ctx);
+}
+
+static noinline_for_stack void
+avd_init_v4l2_vp9_count_tbl(struct avd_ctx *ctx)
+{
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ struct avd_vp9_frame_symbol_counts *cnts = vp9_ctx->bufs.counts.cpu;
+ int coeff_cnts = 0, eob_cnts = 0;
+
+ vp9_ctx->cnts.intra_inter = &cnts->intra_inter;
+ vp9_ctx->cnts.y_mode = &cnts->y_mode;
+ vp9_ctx->cnts.uv_mode = &cnts->uv_mode;
+ vp9_ctx->cnts.comp = &cnts->comp;
+ vp9_ctx->cnts.comp_ref = &cnts->comp_ref;
+ vp9_ctx->cnts.single_ref = &cnts->single_ref;
+ vp9_ctx->cnts.filter = &cnts->filter;
+ vp9_ctx->cnts.mv_mode = &cnts->mv_mode;
+ vp9_ctx->cnts.mv_joint = &cnts->mv_joint;
+ /* all of the mv use a different structure, so they must all be copied */
+ vp9_ctx->cnts.tx8p = &cnts->tx8p;
+ /*
+ * AVD also uses "u32 tx16p[2][3]" instead of "u32 tx16p[2][4]", so
+ * they must also be copied
+ */
+ vp9_ctx->cnts.tx32p = &cnts->tx32p;
+ vp9_ctx->cnts.partition = &cnts->partition;
+ vp9_ctx->cnts.skip = &cnts->skip;
+
+ for (int t = 0; t < 4; t++)
+ for (int i = 0; i < 2; i++)
+ for (int j = 0; j < 2; j++)
+ for (int k = 0; k < 6; k++) {
+ int max_l = (k == 0) ? 3 : 6;
+ for (int l = 0; l < max_l; l++) {
+ vp9_ctx->cnts.coeff[t][i][j][k][l] =
+ (u32 (*)[3])&cnts->ref_cnt[coeff_cnts];
+ coeff_cnts += 3;
+ vp9_ctx->cnts.eob[t][i][j][k][l][0] =
+ &cnts->eob_0[eob_cnts++];
+ vp9_ctx->cnts.eob[t][i][j][k][l][1] =
+ &cnts->ref_cnt[coeff_cnts++];
+ }
+ }
+}
+
+static void avd_vp9_stop(struct avd_ctx *ctx)
+{
+ struct avd_vp9_ctx *vp9_ctx = ctx->priv;
+ struct avd_dev *avd = ctx->dev;
+
+ avd_buf_free(avd, &vp9_ctx->bufs.probs);
+ avd_buf_free(avd, &vp9_ctx->bufs.counts);
+ avd_buf_free(avd, &vp9_ctx->bufs.seg);
+ avd_buf_free(avd, &vp9_ctx->bufs.color);
+ avd_buf_free(avd, &vp9_ctx->bufs.ip_above);
+ avd_buf_free(avd, &vp9_ctx->bufs.lf_above);
+ avd_buf_free(avd, &vp9_ctx->bufs.above_info);
+ avd_buf_free(avd, &vp9_ctx->bufs.az_left);
+ avd_buf_free(avd, &vp9_ctx->bufs.lf_left_info);
+ avd_buf_free(avd, &vp9_ctx->bufs.lf_left);
+
+ kfree(vp9_ctx);
+}
+
+static int avd_vp9_start(struct avd_ctx *ctx)
+{
+ struct avd_vp9_ctx *vp9_ctx;
+ int ret;
+
+ vp9_ctx = kzalloc_obj(*vp9_ctx, GFP_KERNEL);
+ if (!vp9_ctx)
+ return -ENOMEM;
+
+ ctx->priv = vp9_ctx;
+ ret = avd_vp9_alloc_bufs(ctx);
+ if (ret)
+ goto err_free_ctx;
+
+ avd_init_v4l2_vp9_count_tbl(ctx);
+
+ return 0;
+
+err_free_ctx:
+ avd_vp9_stop(ctx);
+ ctx->priv = NULL;
+ return ret;
+}
+
+static enum avd_image_fmt avd_vp9_get_image_fmt(struct avd_ctx *ctx,
+ struct v4l2_ctrl *ctrl)
+{
+#define BIT_DEPTH(chroma) \
+ (frame->bit_depth == 8 ? AVD_IMG_FMT_##chroma##_8BIT : \
+ AVD_IMG_FMT_##chroma##_10BIT)
+ const struct v4l2_ctrl_vp9_frame *frame = ctrl->p_new.p_vp9_frame;
+
+ if (ctrl->id != V4L2_CID_STATELESS_VP9_FRAME)
+ return AVD_IMG_FMT_ANY;
+
+ /* 7.2.2 Color config semantics */
+ if (frame->flags & V4L2_VP9_FRAME_FLAG_X_SUBSAMPLING) {
+ if (frame->flags & V4L2_VP9_FRAME_FLAG_Y_SUBSAMPLING)
+ return BIT_DEPTH(420);
+ else
+ return BIT_DEPTH(422);
+ }
+
+ return AVD_IMG_FMT_ANY;
+#undef BIT_DEPTH
+}
+
+const struct avd_coded_fmt_ops avd_vp9_fmt_ops = {
+ .start = avd_vp9_start,
+ .stop = avd_vp9_stop,
+ .run = avd_vp9_run,
+ .done = avd_vp9_done,
+ .get_image_fmt = avd_vp9_get_image_fmt,
+};
diff --git a/drivers/media/platform/apple/avd/avd.h b/drivers/media/platform/apple/avd/avd.h
index a6e4154d2fac..bd6177ab1eb2 100644
--- a/drivers/media/platform/apple/avd/avd.h
+++ b/drivers/media/platform/apple/avd/avd.h
@@ -98,6 +98,12 @@ struct avd_ctrls {
unsigned int num_ctrls;
};
+struct avd_vp9_decoded_buffer_info {
+ unsigned short width;
+ unsigned short height;
+ unsigned int bit_depth : 4;
+};
+
struct avd_comp {
u32 size;
/* offset to start of compressed data */
@@ -110,6 +116,9 @@ struct avd_decoded_buffer {
/* Must be the first field in this struct. */
struct v4l2_m2m_buffer base;
struct avd_comp comp;
+ union {
+ struct avd_vp9_decoded_buffer_info vp9;
+ };
};
static inline struct avd_decoded_buffer *
@@ -243,6 +252,7 @@ void avd_run_preamble(struct avd_ctx *ctx, struct avd_run *run);
void avd_run_postamble(struct avd_ctx *ctx, struct avd_run *run);
extern const struct avd_coded_fmt_ops avd_h264_fmt_ops;
+extern const struct avd_coded_fmt_ops avd_vp9_fmt_ops;
extern const struct v4l2_ctrl_ops avd_ctrl_ops;
extern const struct v4l2_ioctl_ops avd_ioctl_ops;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 09/17] media: apple: avd: add hevc support
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (7 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 08/17] media: apple: avd: add vp9 support Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 11/17] arm64: dts: apple: t8103: add avd nodes Sofus Forstreuter
` (7 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
The fluster score is 143/147 for JCT-VC-HEVC_V1.
RPS_E_qualcomm_5, WPP_D_ericsson_MAIN_2 and WPP_D_ericsson_MAIN10_2 are
failing, but should be supported.
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
drivers/media/platform/apple/avd/Makefile | 2 +-
drivers/media/platform/apple/avd/avd-hevc.c | 1340 +++++++++++++++++++++++++++
drivers/media/platform/apple/avd/avd-v4l2.c | 72 ++
drivers/media/platform/apple/avd/avd.h | 6 +
4 files changed, 1419 insertions(+), 1 deletion(-)
diff --git a/drivers/media/platform/apple/avd/Makefile b/drivers/media/platform/apple/avd/Makefile
index 26763a13edb2..903b86c0fe5c 100644
--- a/drivers/media/platform/apple/avd/Makefile
+++ b/drivers/media/platform/apple/avd/Makefile
@@ -1,4 +1,4 @@
# SPDX-License-Identifier: GPL-2.0-only
-apple-avd-y := avd-drv.o avd-v4l2.o avd-h264.o avd-vp9.o
+apple-avd-y := avd-drv.o avd-v4l2.o avd-h264.o avd-vp9.o avd-hevc.o
obj-$(CONFIG_VIDEO_APPLE_AVD) += apple-avd.o
diff --git a/drivers/media/platform/apple/avd/avd-hevc.c b/drivers/media/platform/apple/avd/avd-hevc.c
new file mode 100644
index 000000000000..d90c7770006b
--- /dev/null
+++ b/drivers/media/platform/apple/avd/avd-hevc.c
@@ -0,0 +1,1340 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Apple Video Decoder HEVC driver
+ *
+ * Copyright (C) 2026 The Asahi Linux Contributors
+ * Copyright (C) 2026 Sofus Forstreuter <sofus.c@icloud.com>
+ * Copyright (C) 2023 Eileen Yoon <eyn@gmx.com>
+ *
+ * Copyright (c) 2014 Rockchip Electronics Co., Ltd.
+ * Hertz Wong <hertz.wong@rock-chips.com>
+ * Herman Chen <herman.chen@rock-chips.com>
+ *
+ * Copyright (C) 2014 Google, Inc.
+ * Tomasz Figa <tfiga@chromium.org>
+ */
+
+#include <linux/dev_printk.h>
+
+#include <media/videobuf2-dma-contig.h>
+
+#include "avd.h"
+#include "avd-inst.h"
+
+#define NEW_TILE_ID BIT(0)
+#define NEW_SLICE BIT(1)
+
+#define HEVC_TR_INTRA(v) FIELD_PREP(GENMASK(4, 1), v)
+#define HEVC_TR_INTER(v) FIELD_PREP(GENMASK(7, 4), v)
+
+#define HEVC_PCM_EN(v) FIELD_PREP(BIT(12), !!(v))
+#define HEVC_PCM_BD_LUMA(v) FIELD_PREP(GENMASK(11, 8), v)
+#define HEVC_PCM_BD_CHROMA(v) FIELD_PREP(GENMASK(7, 4), v)
+#define HEVC_PCM_CB_MIN_LUMA(v) FIELD_PREP(GENMASK(3, 2), v)
+#define HEVC_PCM_CB_LUMA(v) FIELD_PREP(GENMASK(1, 0), v)
+
+#define HEVC_UNK_ISM_EN(v) FIELD_PREP(BIT(9), !!(v))
+#define HEVC_UNK_FLAG FIELD_PREP(BIT(3), 1)
+
+#define HEVC_CTB_SIZE(v) FIELD_PREP(GENMASK(8, 3), v)
+#define HEVC_MERGE_LV(v) FIELD_PREP(GENMASK(11, 9), v)
+#define HEVC_FLAG_ENTROPY_EN(v) FIELD_PREP(BIT(12), !!(v))
+#define HEVC_FLAG_TILES_EN(v) FIELD_PREP(BIT(13), !!(v))
+#define HEVC_FLAG_TQB_EN(v) FIELD_PREP(BIT(14), !!(v))
+#define HEVC_CU_QP_DD(v) FIELD_PREP(GENMASK(16, 15), v)
+#define HEVC_FLAG_CU_QP_EN(v) FIELD_PREP(BIT(17), !!(v))
+#define HEVC_FLAG_TSKIP_EN(v) FIELD_PREP(BIT(18), !!(v))
+#define HEVC_FLAG_CI_PRED(v) FIELD_PREP(BIT(19), !!(v))
+#define HEVC_FLAG_SDH_EN(v) FIELD_PREP(BIT(20), !!(v))
+#define HEVC_FLAG_TMVP_EN(v) FIELD_PREP(BIT(21), !!(v))
+
+#define HEVC_SCL_DIMS 0x127ffff
+
+/* this is for level 7.1 */
+#define HEVC_MAX_TILE_COLS 40
+#define HEVC_MAX_TILE_ROWS 44
+
+static inline u32 mv_color_size(u32 w, u32 h)
+{
+ /* this will waste some memory when max cu size != 64 */
+ return DIV_ROUND_UP(w, 64) * DIV_ROUND_UP(h, 64) * 256;
+}
+
+struct avd_hevc_tile_info {
+ u16 col_width[HEVC_MAX_TILE_COLS];
+ u16 row_height[HEVC_MAX_TILE_ROWS];
+ u32 col_bd[HEVC_MAX_TILE_COLS + 1];
+ u32 row_bd[HEVC_MAX_TILE_ROWS + 1];
+ u32 *ctb_addr_rs_to_ts;
+ u32 *tile_ids;
+};
+
+struct avd_hevc_run {
+ struct avd_run base;
+ const struct v4l2_ctrl_hevc_decode_params *decode;
+ const struct v4l2_ctrl_hevc_sps *sps;
+ const struct v4l2_ctrl_hevc_pps *pps;
+ const struct v4l2_ctrl_hevc_scaling_matrix *scaling_matrix;
+ const struct v4l2_ctrl_hevc_slice_params *(*sl);
+ const u32 *(*entry_point_offsets);
+ int num_entry_point_offsets;
+ int num_slices;
+ struct run_addr {
+ dma_addr_t mv_color;
+ } addresses;
+ struct avd_hevc_tile_info tile_info;
+};
+
+struct avd_hevc_ctx {
+ struct avd_h264_bufs {
+ struct avd_buf mv_above_info;
+ struct avd_buf az_above;
+ struct avd_buf ip_above;
+ struct avd_buf lf_above;
+ struct avd_buf lf_above_info;
+ struct avd_buf lf_left;
+ struct avd_buf lf_left_info;
+ struct avd_buf sw_left;
+ } bufs;
+};
+
+static void stream_refs(struct avd_ctx *ctx, struct avd_hevc_run *run)
+{
+ const struct v4l2_ctrl_hevc_decode_params *decode = run->decode;
+ const struct v4l2_ctrl_hevc_slice_params *sl = &(*run->sl)[0];
+ const struct v4l2_hevc_dpb_entry *dpb;
+ struct avd_hevc_ctx *hevc_ctx = ctx->priv;
+ struct avd_decoded_buffer *dst, *ref_buf;
+
+ dst = vb2_to_avd_decoded_buf(&run->base.bufs.dst->vb2_buf);
+
+ push(0, "");
+ pusha(hevc_ctx->bufs.mv_above_info.addr, "mv_above_info", 7);
+ pusha(run->addresses.mv_color, "mv_color", 0);
+
+ push(0, "");
+ push(0, "");
+ push(0, "");
+ push(0, "");
+
+ for (int i = 0; i < decode->num_active_dpb_entries; i++) {
+ dpb = &decode->dpb[i];
+
+ ref_buf = avd_get_ref_buf(ctx, &dst->base.vb, dpb->timestamp);
+
+ dma_addr_t comp_addr = vb2_dma_contig_plane_dma_addr(
+ &ref_buf->base.vb.vb2_buf, 0) +
+ ref_buf->comp.start_offset;
+
+ push(AVD_REF_NUM(decode->num_active_dpb_entries - 1) |
+ AVD_REF_FLAG_CONST |
+ AVD_REF_FLAG_LONG(
+ dpb->flags &
+ V4L2_HEVC_DPB_ENTRY_LONG_TERM_REFERENCE) |
+ AVD_REF_DELTA_POC(sl->slice_pic_order_cnt -
+ dpb->pic_order_cnt_val),
+ "hdr_d0_ref_hdr");
+
+ push_comp(ctx, comp_addr, ref_buf->comp.offsets);
+ }
+}
+
+static void set_scaling_lists(struct avd_ctx *ctx, struct avd_hevc_run *run)
+{
+ const struct v4l2_ctrl_hevc_scaling_matrix *s = run->scaling_matrix;
+ int i, j, k;
+
+ u8 (*dc_16x16)[3] = (u8(*)[3])s->scaling_list_dc_coef_16x16;
+ u8 (*sc_4x4)[4][4] = (u8(*)[4][4])s->scaling_list_4x4;
+ u8 (*sc_8x8)[2][4][8] = (u8(*)[2][4][8])s->scaling_list_8x8;
+ u8 (*sc_16x16)[2][4][8] = (u8(*)[2][4][8])s->scaling_list_16x16;
+ u8 (*sc_32x32)[2][4][8] = (u8(*)[2][4][8])s->scaling_list_32x32;
+
+ /*
+ * this should presumably be how many of each are enabled?
+ * or some other scaling related thing
+ */
+ push(HEVC_SCL_DIMS, "hdr_7c_pps_scl_dims");
+
+ for (i = 0; i < 2; i++)
+ push(AVD_SCALING_I2(dc_16x16[i][0]) |
+ AVD_SCALING_I1(dc_16x16[i][1]) |
+ AVD_SCALING_I0(dc_16x16[i][2]),
+ "dc_16x16");
+
+ for (i = 0; i < 2; i++)
+ push(AVD_SCALING_I2(s->scaling_list_dc_coef_32x32[i]),
+ "dc_32x32");
+
+ /* transposed in stride 4 */
+ for (i = 0; i < 6; i++)
+ for (j = 0; j < 4; j++)
+ push(AVD_SCALING_I3(sc_4x4[i][0][j]) |
+ AVD_SCALING_I2(sc_4x4[i][1][j]) |
+ AVD_SCALING_I1(sc_4x4[i][2][j]) |
+ AVD_SCALING_I0(sc_4x4[i][3][j]),
+ "scaling_4x4");
+
+ /* transposed in stride 8 */
+ for (i = 0; i < 6; i++)
+ for (j = 0; j < 2; j++)
+ for (k = 0; k < 8; k++)
+ push(AVD_SCALING_I3(sc_8x8[i][j][0][k]) |
+ AVD_SCALING_I2(
+ sc_8x8[i][j][1][k]) |
+ AVD_SCALING_I1(
+ sc_8x8[i][j][2][k]) |
+ AVD_SCALING_I0(sc_8x8[i][j][3][k]),
+ "scaling_8x8");
+
+ for (i = 0; i < 6; i++)
+ for (j = 0; j < 2; j++)
+ for (k = 0; k < 8; k++)
+ push(AVD_SCALING_I3(sc_16x16[i][j][0][k]) |
+ AVD_SCALING_I2(
+ sc_16x16[i][j][1][k]) |
+ AVD_SCALING_I1(
+ sc_16x16[i][j][2][k]) |
+ AVD_SCALING_I0(
+ sc_16x16[i][j][3][k]),
+ "scaling_16x16");
+
+ for (i = 0; i < 2; i++)
+ for (j = 0; j < 2; j++)
+ for (k = 0; k < 8; k++)
+ push(AVD_SCALING_I3(sc_32x32[i][j][0][k]) |
+ AVD_SCALING_I2(
+ sc_32x32[i][j][1][k]) |
+ AVD_SCALING_I1(
+ sc_32x32[i][j][2][k]) |
+ AVD_SCALING_I0(
+ sc_32x32[i][j][3][k]),
+ "scaling_32x32");
+}
+
+static void hevc_set_flags(struct avd_ctx *ctx, struct avd_hevc_run *run)
+{
+ const struct v4l2_ctrl_hevc_decode_params *decode = run->decode;
+ const struct v4l2_ctrl_hevc_sps *sps = run->sps;
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+
+ u32 log2_ctb_size = ((sps->log2_min_luma_coding_block_size_minus3) +
+ sps->log2_diff_max_min_luma_coding_block_size);
+
+ push(HEVC_PCM_EN(sps->flags & V4L2_HEVC_SPS_FLAG_PCM_ENABLED) |
+ HEVC_PCM_BD_LUMA(sps->pcm_sample_bit_depth_luma_minus1) |
+ HEVC_PCM_BD_CHROMA(
+ sps->pcm_sample_bit_depth_chroma_minus1) |
+ HEVC_PCM_CB_MIN_LUMA(
+ sps->log2_min_pcm_luma_coding_block_size_minus3) |
+ HEVC_PCM_CB_LUMA(
+ sps->log2_diff_max_min_pcm_luma_coding_block_size +
+ sps->log2_min_pcm_luma_coding_block_size_minus3),
+ "hdr_30_sps_pcm");
+
+ /*
+ * RExt sets a few new flags here
+ */
+ push(HEVC_UNK_FLAG |
+ HEVC_UNK_ISM_EN(
+ sps->flags &
+ V4L2_HEVC_SPS_FLAG_STRONG_INTRA_SMOOTHING_ENABLED),
+ "hdr_34_sps_flags");
+
+ push(HEVC_CTB_SIZE(log2_ctb_size) |
+ HEVC_MERGE_LV(pps->log2_parallel_merge_level_minus2) |
+ HEVC_FLAG_ENTROPY_EN(
+ pps->flags &
+ V4L2_HEVC_PPS_FLAG_ENTROPY_CODING_SYNC_ENABLED) |
+ HEVC_FLAG_TILES_EN(pps->flags &
+ V4L2_HEVC_PPS_FLAG_TILES_ENABLED) |
+ HEVC_FLAG_TQB_EN(
+ pps->flags &
+ V4L2_HEVC_PPS_FLAG_TRANSQUANT_BYPASS_ENABLED) |
+ HEVC_CU_QP_DD(log2_ctb_size -
+ pps->diff_cu_qp_delta_depth) |
+ HEVC_FLAG_CU_QP_EN(
+ pps->flags &
+ V4L2_HEVC_PPS_FLAG_CU_QP_DELTA_ENABLED) |
+ HEVC_FLAG_TSKIP_EN(
+ pps->flags &
+ V4L2_HEVC_PPS_FLAG_TRANSFORM_SKIP_ENABLED) |
+ HEVC_FLAG_CI_PRED(
+ pps->flags &
+ V4L2_HEVC_PPS_FLAG_CONSTRAINED_INTRA_PRED) |
+ HEVC_FLAG_SDH_EN(
+ pps->flags &
+ V4L2_HEVC_PPS_FLAG_SIGN_DATA_HIDING_ENABLED) |
+ HEVC_FLAG_TMVP_EN(
+ !(decode->flags &
+ V4L2_HEVC_DECODE_PARAM_FLAG_IDR_PIC) &&
+ sps->flags &
+ V4L2_HEVC_SPS_FLAG_SPS_TEMPORAL_MVP_ENABLED),
+ "hdr_5c_pps_flags");
+
+ push(AVD_HDR_H26X_QP_OFFSET_CB(pps->pps_cb_qp_offset) |
+ AVD_HDR_H26X_QP_OFFSET_CR(pps->pps_cr_qp_offset),
+ "hdr_60_pps_qp");
+
+ push(0, "hdr_64_zero");
+ push(0, "hdr_68_zero");
+ push(0, "hdr_6c_zero");
+ push(0, "hdr_70_zero");
+ push(0, "hdr_74_zero");
+ push(0, "hdr_78_zero");
+}
+
+static void set_header(struct avd_ctx *ctx, struct avd_hevc_run *run)
+{
+ const struct v4l2_ctrl_hevc_sps *sps = run->sps;
+ struct avd_dev *avd = ctx->dev;
+ struct avd_hevc_ctx *hevc_ctx = ctx->priv;
+ u32 bytesperline;
+ u32 width = sps->pic_width_in_luma_samples;
+ u32 height = sps->pic_height_in_luma_samples;
+
+ bool is_intra = (*run->sl)[0].slice_type == V4L2_HEVC_SLICE_TYPE_I;
+
+ push(AVD_OP_HDR | AVD_OP_HDR_FLAG_DECOMP(ctx->decomp) |
+ AVD_OP_HDR_FLAG_INTRA(is_intra) | AVD_OP_HDR_CONST |
+ AVD_OP_HDR_FLAG_PIPE_STATE(
+ !(avd->variant->quirks & AVD_QUIRK_NO_PIPE_STATE)),
+ "hdr_34_start_hdr");
+
+ push(AVD_HDR_CODEC_MODE(AVD_CODEC_HEVC), "hdr_50_mode");
+ push(AVD_HDR_HEIGHT(height - 1) | AVD_HDR_WIDTH(width - 1),
+ "hdr_54_height_width");
+ push(0, "hdr_58_pixfmt_zero");
+
+ push(AVD_HDR_HEIGHT((height - 1) >> 3) |
+ AVD_HDR_WIDTH((width - 1) >> 3),
+ "hdr_28_height_width_shift3");
+
+ push(AVD_HDR_COMMON_CHROMA_FORMAT(sps->chroma_format_idc) |
+ AVD_HDR_COMMON_BIT_DEPTH_C(sps->bit_depth_chroma_minus8) |
+ AVD_HDR_COMMON_BIT_DEPTH_L(sps->bit_depth_luma_minus8) |
+ AVD_HDR_COMMON_MIN_LUMA_CBS(
+ sps->log2_min_luma_coding_block_size_minus3) |
+ AVD_HDR_COMMON_LUMA_CBS(
+ sps->log2_diff_max_min_luma_coding_block_size +
+ sps->log2_min_luma_coding_block_size_minus3) |
+ AVD_HDR_COMMON_MIN_LUMA_TBS(
+ sps->log2_min_luma_transform_block_size_minus2) |
+ AVD_HDR_COMMON_LUMA_TBS(
+ sps->log2_diff_max_min_luma_transform_block_size +
+ + sps->log2_min_luma_transform_block_size_minus2) |
+ HEVC_TR_INTER(sps->max_transform_hierarchy_depth_inter) |
+ HEVC_TR_INTRA(sps->max_transform_hierarchy_depth_intra) |
+ AVD_HDR_COMMON_FLAG0(sps->flags &
+ V4L2_HEVC_SPS_FLAG_AMP_ENABLED),
+ "hdr_2c_sps_txfm");
+
+ hevc_set_flags(ctx, run);
+
+ push(AVD_HDR_FEAT_H26X | AVD_HDR_FEAT_COMMON |
+ AVD_HDR_FEAT_PIPE_STATE_EN(
+ !(avd->variant->quirks & AVD_QUIRK_NO_PIPE_STATE)),
+ "hdr_98_const_30");
+
+ push(0, "");
+ push(0, "");
+
+ if (avd->variant->revision == 3)
+ push(0, "");
+
+ push(0, "");
+ push(0, "");
+
+ if (avd->variant->revision == 3)
+ push(0, "");
+ else if (!(avd->variant->quirks & AVD_QUIRK_NO_PIPE_STATE))
+ pusha(ctx->pipe_state.addr, "pipe_state", 0);
+
+ pusha(hevc_ctx->bufs.ip_above.addr, "ip_above", 0);
+ pusha(hevc_ctx->bufs.lf_above.addr, "lf_above", 1);
+ pusha(hevc_ctx->bufs.lf_above_info.addr, "lf_above_info", 2);
+ pusha(hevc_ctx->bufs.lf_left.addr, "lf_left", 3);
+ pusha(hevc_ctx->bufs.lf_left_info.addr, "lf_left_info", 4);
+ pusha(hevc_ctx->bufs.az_above.addr, "az_above", 8);
+ pusha(hevc_ctx->bufs.sw_left.addr, "sw_left", 9);
+
+ push(0, "");
+
+ push_comp(ctx, run->base.comp_out, ctx->comp.offsets);
+
+ pusha((u64)0, "packed_fmt_scratch", 0);
+
+ bytesperline = ctx->decoded_fmt.fmt.pix_mp.plane_fmt[0].bytesperline;
+ if (avd->variant->quirks & AVD_QUIRK_LSR)
+ bytesperline = bytesperline >> 4;
+
+ pusha(run->base.y_out, "y_out", 0);
+ push(bytesperline, "y_bpl");
+ pusha(run->base.uv_out, "uv_out", 0);
+ push(bytesperline, "uv_bpl");
+ push(0, "");
+ push(AVD_HDR_HEIGHT(height - 1) | AVD_HDR_WIDTH(width - 1),
+ "hdr_54_height_width");
+
+ if (!is_intra)
+ stream_refs(ctx, run);
+
+ if ((sps->flags & V4L2_HEVC_SPS_FLAG_SCALING_LIST_ENABLED))
+ set_scaling_lists(ctx, run);
+ else
+ push(0, "");
+}
+
+static void stream_weights(struct avd_ctx *ctx, struct avd_hevc_run *run,
+ const struct v4l2_ctrl_hevc_slice_params *sl)
+{
+ int luma_weight_denom, chroma_weight_denom;
+ u8 chroma_log2_weight_denom;
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ const struct v4l2_hevc_pred_weight_table *pred = &sl->pred_weight_table;
+ bool has_luma_weights =
+ ((pps->flags & V4L2_HEVC_PPS_FLAG_WEIGHTED_PRED) &&
+ sl->slice_type == V4L2_HEVC_SLICE_TYPE_P) ||
+ ((pps->flags & V4L2_HEVC_PPS_FLAG_WEIGHTED_BIPRED) &&
+ sl->slice_type == V4L2_HEVC_SLICE_TYPE_B);
+
+ const s8(*delta_chroma_weights)[2];
+ const s8(*chroma_offsets)[2];
+ const s8 *delta_luma_weights;
+ const s8 *luma_offsets;
+
+ if (!has_luma_weights) {
+ push(AVD_OP_WEIGHTS_HDR, "slc_76c_cmd_weights_denom");
+ return;
+ }
+
+ chroma_log2_weight_denom = pred->luma_log2_weight_denom +
+ pred->delta_chroma_log2_weight_denom;
+
+ push(AVD_OP_WEIGHTS_HDR | AVD_OP_WEIGHTS_HDR_FLAG1(!has_luma_weights) |
+ AVD_OP_WEIGHTS_HDR_FLAG0(has_luma_weights) |
+ AVD_OP_WEIGHTS_HDR_LUMA(pred->luma_log2_weight_denom) |
+ AVD_OP_WEIGHTS_HDR_CHROMA(chroma_log2_weight_denom),
+ "slc_76c_cmd_weights_denom");
+
+ luma_weight_denom = 1 << pred->luma_log2_weight_denom;
+ chroma_weight_denom = 1 << chroma_log2_weight_denom;
+
+ /* comes from 7.4.7.3 */
+
+ for (int y = 0; y < 2; y++) {
+ if (y == 1 && sl->slice_type != V4L2_HEVC_SLICE_TYPE_B)
+ break;
+ luma_offsets = y == 0 ? pred->luma_offset_l0 :
+ pred->luma_offset_l1;
+ delta_luma_weights =
+ (s8(*))(y == 0 ? pred->delta_luma_weight_l0 :
+ pred->delta_luma_weight_l1);
+ chroma_offsets = (s8(*)[2])(y == 0 ? pred->chroma_offset_l0 :
+ pred->chroma_offset_l1);
+ delta_chroma_weights =
+ (s8(*)[2])(y == 0 ? pred->delta_chroma_weight_l0 :
+ pred->delta_chroma_weight_l1);
+
+ int to = y == 0 ? sl->num_ref_idx_l0_active_minus1 :
+ sl->num_ref_idx_l1_active_minus1;
+ for (int i = 0; i < to + 1; i++) {
+ if (delta_luma_weights[i] != 0 ||
+ luma_offsets[i] != 0) {
+ push(AVD_OP_WEIGHTS | AVD_OP_WEIGHTS_IDENT(1) |
+ AVD_OP_WEIGHTS_LIST_IDX(y) |
+ AVD_OP_WEIGHTS_INDEX(i) |
+ AVD_OP_WEIGHTS_WEIGHT(
+ delta_luma_weights[i] +
+ luma_weight_denom),
+ "slc_luma_weights");
+ push(AVD_OP_OFFSETS | AVD_OP_OFFSETS_OFFSET(
+ luma_offsets[i]),
+ "slc_luma_offsets");
+ }
+
+ if (delta_chroma_weights[i][0] != 0 ||
+ chroma_offsets[i][0] != 0 ||
+ delta_chroma_weights[i][1] != 0 ||
+ chroma_offsets[i][1] != 0) {
+ push(AVD_OP_WEIGHTS | AVD_OP_WEIGHTS_IDENT(2) |
+ AVD_OP_WEIGHTS_LIST_IDX(y) |
+ AVD_OP_WEIGHTS_INDEX(i) |
+ AVD_OP_WEIGHTS_WEIGHT(
+ delta_chroma_weights[i][0] +
+ chroma_weight_denom),
+ "slc_chroma_weights[0]");
+ push(AVD_OP_OFFSETS |
+ AVD_OP_OFFSETS_OFFSET(
+ chroma_offsets[i][0]),
+ "slc_chroma_offsets[0]");
+ push(AVD_OP_WEIGHTS | AVD_OP_WEIGHTS_IDENT(3) |
+ AVD_OP_WEIGHTS_LIST_IDX(y) |
+ AVD_OP_WEIGHTS_INDEX(i) |
+ AVD_OP_WEIGHTS_WEIGHT(
+ delta_chroma_weights[i][1] +
+ chroma_weight_denom),
+ "slc_chroma_weights[1]");
+ push(AVD_OP_OFFSETS |
+ AVD_OP_OFFSETS_OFFSET(
+ chroma_offsets[i][1]),
+ "slc_chroma_offsets[1]");
+ }
+ }
+ }
+}
+
+static void stream_slice_dqtblk(struct avd_ctx *ctx, struct avd_hevc_run *run,
+ const struct v4l2_ctrl_hevc_slice_params *sl)
+{
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ const struct v4l2_ctrl_hevc_sps *sps = run->sps;
+
+ push(AVD_OP_QP |
+ AVD_OP_QP_VAL(pps->init_qp_minus26 + 26 +
+ sl->slice_qp_delta) |
+ AVD_OP_QP_CB_OFF(pps->pps_cb_qp_offset +
+ sl->slice_cb_qp_offset) |
+ AVD_OP_QP_CR_OFF(pps->pps_cr_qp_offset +
+ sl->slice_cr_qp_offset),
+ "slc_bcc_cmd_quantization");
+
+ push(AVD_OP_DBLK |
+ AVD_OP_DBLK_FLAG_SAO_CHROMA(
+ sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_SLICE_SAO_CHROMA) |
+ AVD_OP_DBLK_FLAG_SAO_LUMA(
+ sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_SLICE_SAO_LUMA) |
+ AVD_OP_DBLK_OFF0(sl->slice_tc_offset_div2) |
+ AVD_OP_DBLK_OFF1(sl->slice_beta_offset_div2) |
+ /* i wonder what this should actually be */
+ AVD_OP_DBLK_FLAG_EN(
+ !(sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_SLICE_DEBLOCKING_FILTER_DISABLED) &&
+ (sps->flags &
+ V4L2_HEVC_SPS_FLAG_STRONG_INTRA_SMOOTHING_ENABLED ||
+ sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_SLICE_SAO_LUMA ||
+ pps->flags &
+ (V4L2_HEVC_PPS_FLAG_DEBLOCKING_FILTER_OVERRIDE_ENABLED |
+ V4L2_HEVC_PPS_FLAG_DEBLOCKING_FILTER_CONTROL_PRESENT))) |
+ AVD_OP_DBLK_FLAG_FULL_EN(
+ sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_SLICE_LOOP_FILTER_ACROSS_SLICES_ENABLED) |
+ AVD_OP_DBLK_FLAG_TILES_EN(
+ !(pps->flags & V4L2_HEVC_PPS_FLAG_TILES_ENABLED) ||
+ (pps->flags &
+ V4L2_HEVC_PPS_FLAG_LOOP_FILTER_ACROSS_TILES_ENABLED)) |
+ AVD_OP_DBLK_FLAG_PCM_EN(
+ (sps->flags & V4L2_HEVC_SPS_FLAG_PCM_ENABLED) &&
+ !(sps->flags &
+ V4L2_HEVC_SPS_FLAG_PCM_LOOP_FILTER_DISABLED)),
+ "slc_bd0_cmd_deblocking_filter");
+
+ if (sl->slice_type == V4L2_HEVC_SLICE_TYPE_B ||
+ sl->slice_type == V4L2_HEVC_SLICE_TYPE_P) {
+ for (int i = 0; i < sl->num_ref_idx_l0_active_minus1 + 1; i++)
+ push(AVD_OP_REF | AVD_OP_REF_LIST_IDX(0) |
+ AVD_OP_REF_LOOP_IDX(i) |
+ AVD_OP_REF_DBP_IDX(sl->ref_idx_l0[i]),
+ "reference_frames_l0");
+ if (sl->slice_type == V4L2_HEVC_SLICE_TYPE_B)
+ for (int i = 0;
+ i < sl->num_ref_idx_l1_active_minus1 + 1; i++)
+ push(AVD_OP_REF | AVD_OP_REF_LIST_IDX(1) |
+ AVD_OP_REF_LOOP_IDX(i) |
+ AVD_OP_REF_DBP_IDX(
+ sl->ref_idx_l1[i]),
+ "reference_frames_l1");
+
+ stream_weights(ctx, run, sl);
+ }
+}
+
+static void stream_slice_mv(struct avd_ctx *ctx, struct avd_hevc_run *run,
+ const struct v4l2_ctrl_hevc_slice_params *sl,
+ bool is_first)
+{
+ const struct v4l2_ctrl_hevc_decode_params *decode = run->decode;
+ struct avd_decoded_buffer *dst, *ref;
+ bool ref_valid;
+ const u8 *ref_list;
+ int ref_idx = 0;
+
+ ref_list = sl->slice_type == V4L2_HEVC_SLICE_TYPE_P ? sl->ref_idx_l0 :
+ sl->flags & V4L2_HEVC_SLICE_PARAMS_FLAG_COLLOCATED_FROM_L0 ?
+ sl->ref_idx_l0 :
+ sl->ref_idx_l1;
+
+ if (sl->slice_type == V4L2_HEVC_SLICE_TYPE_I) {
+ push(AVD_OP_SL_REF |
+ AVD_OP_SL_REF_SLICE_I(sl->slice_type ==
+ V4L2_HEVC_SLICE_TYPE_I),
+ "slc_a8c_cmd_ref_type");
+ return;
+ }
+ /* bidirectional prediction */
+ if (sl->collocated_ref_idx < V4L2_HEVC_DPB_ENTRIES_NUM_MAX &&
+ ref_list[sl->collocated_ref_idx] < V4L2_HEVC_DPB_ENTRIES_NUM_MAX)
+ ref_idx = ref_list[sl->collocated_ref_idx];
+
+ dst = vb2_to_avd_decoded_buf(&run->base.bufs.dst->vb2_buf);
+ ref = avd_get_ref_buf(ctx, &dst->base.vb,
+ decode->dpb[ref_idx].timestamp);
+
+ ref_valid = !(sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_DEPENDENT_SLICE_SEGMENT) &&
+ (sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_SLICE_TEMPORAL_MVP_ENABLED) &&
+ is_first && !ref->hevc.is_intra;
+
+ push(AVD_OP_SL_REF |
+ AVD_OP_SL_REF_MAX_MERGE(
+ 5 - sl->five_minus_max_num_merge_cand) |
+ AVD_OP_SL_REF_FLAG_CABAC(
+ sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_CABAC_INIT) |
+ AVD_OP_SL_REF_FLAG0(
+ (sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_SLICE_TEMPORAL_MVP_ENABLED) &&
+ !(sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_COLLOCATED_FROM_L0)) |
+ AVD_OP_SL_REF_FLAG1(
+ !(sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_MVD_L1_ZERO)) |
+ AVD_OP_SL_REF_FLAG2(
+ (sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_SLICE_TEMPORAL_MVP_ENABLED) ||
+ (sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_DEPENDENT_SLICE_SEGMENT)) |
+ AVD_OP_SL_REF_NUM_L0(sl->num_ref_idx_l0_active_minus1) |
+ AVD_OP_SL_REF_NUM_L1(sl->num_ref_idx_l1_active_minus1) |
+ AVD_OP_SL_REF_SLICE_P(sl->slice_type ==
+ V4L2_HEVC_SLICE_TYPE_P) |
+ AVD_OP_SL_REF_SLICE_B(ref_valid),
+ "slc_a8c_cmd_ref_type");
+
+ if (ref_valid) {
+ dma_addr_t mv_color_addr =
+ vb2_dma_contig_plane_dma_addr(&ref->base.vb.vb2_buf,
+ 0) +
+ (ref->base.vb.planes[0].length -
+ mv_color_size(fmt_width(ctx), fmt_height(ctx)));
+ pusha(mv_color_addr, "slc_bd4_sps_tile_addr2_lsb8",
+ decode->dpb[ref_list[sl->collocated_ref_idx]]
+ .pic_order_cnt_val);
+ }
+}
+
+static void set_slice(struct avd_ctx *ctx, struct avd_hevc_run *run,
+ const struct v4l2_ctrl_hevc_slice_params *sl, u32 size,
+ u32 offset, u32 flags)
+{
+ dma_addr_t coded_in =
+ run->base.coded_in + offset + sl->data_byte_offset;
+ push(AVD_OP_CODED_DATA | flags | AVD_OP_CODED_IN_HI(coded_in),
+ "cm3_cmd_set_coded_slice");
+ push(AVD_OP_CODED_IN_LO(coded_in), "slc_bd8_slice_addr");
+ push(size, "slc_bdc_slice_size");
+}
+
+static int submit_slice_segment(struct avd_ctx *ctx, struct avd_hevc_run *run,
+ const struct v4l2_ctrl_hevc_slice_params *sl,
+ int row, int col, u32 col_bd[23],
+ u32 row_bd[23], u32 pic_in_cts_width,
+ u32 pic_in_cts_height, bool first_slice,
+ bool hflip, bool vflip, u32 coded_flags,
+ u32 last_tile_block)
+{
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ u32 tb_x, tb_y, tile_block, tile_boundary;
+
+ if (coded_flags & NEW_SLICE) {
+ tb_x = sl->slice_segment_addr % pic_in_cts_width;
+ tb_y = sl->slice_segment_addr / pic_in_cts_width;
+
+ tile_block = AVD_OP_SL_LOC_Y(tb_y) | AVD_OP_SL_LOC_X(tb_x);
+
+ if (!(sl->flags &
+ V4L2_HEVC_SLICE_PARAMS_FLAG_DEPENDENT_SLICE_SEGMENT))
+ last_tile_block = tile_block;
+
+ /*
+ * tile block start
+ * CABAC window
+ */
+ push(AVD_OP_SL_LOC | last_tile_block, "cm3_cmd_set_cabac_xy");
+
+ stream_slice_dqtblk(ctx, run, sl);
+ } else {
+ tile_boundary = AVD_OP_SL_LOC_Y(row_bd[row]) |
+ AVD_OP_SL_LOC_X(col_bd[col]);
+ }
+
+ if (coded_flags & NEW_TILE_ID) {
+ push(AVD_OP_SL_DIM_START |
+ (coded_flags & NEW_SLICE ? tile_block :
+ tile_boundary),
+ "cm3_cmd_set_ctb_xy");
+
+ /* tile boundary end */
+ if (pps->flags & V4L2_HEVC_PPS_FLAG_TILES_ENABLED)
+ push(AVD_SL_DIM_END_ROW((hflip ? 4 : 0) |
+ (vflip ? 8 : 0)) |
+ AVD_SL_DIM_END_COL(col) |
+ AVD_SL_DIM_END_Y(row_bd[row + 1] - 1) |
+ AVD_SL_DIM_END_X(col_bd[col + 1] - 1),
+ "cm3_set_ctb_xy");
+ else /* first slice, one CTB */
+ push(AVD_SL_DIM_END_Y(pic_in_cts_height - 1) |
+ AVD_SL_DIM_END_X(pic_in_cts_width - 1),
+ "cm3_set_ctb_xy");
+ }
+
+ if (coded_flags & NEW_SLICE)
+ stream_slice_mv(ctx, run, sl, first_slice);
+
+ /* current tile block / boundary ?? */
+ /* Unlike entropy, motion vector window resets every time */
+ push(AVD_SL_DIM_END_COL(1) |
+ (coded_flags & NEW_SLICE ? tile_block : tile_boundary),
+ "cm3_set_mv_xy");
+
+ return last_tile_block;
+}
+
+static void compute_tiles_uniform(struct avd_hevc_run *run,
+ u16 log2_min_cb_size, u16 width, u16 height,
+ u32 pic_in_cts_width, u32 pic_in_cts_height,
+ u16 *column_width, u16 *row_height)
+{
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ int i;
+
+ for (i = 0; i < pps->num_tile_columns_minus1 + 1; i++)
+ column_width[i] = ((i + 1) * pic_in_cts_width) /
+ (pps->num_tile_columns_minus1 + 1) -
+ (i * pic_in_cts_width) /
+ (pps->num_tile_columns_minus1 + 1);
+
+ for (i = 0; i < pps->num_tile_rows_minus1 + 1; i++)
+ row_height[i] = ((i + 1) * pic_in_cts_height) /
+ (pps->num_tile_rows_minus1 + 1) -
+ (i * pic_in_cts_height) /
+ (pps->num_tile_rows_minus1 + 1);
+}
+
+static void compute_tiles_non_uniform(struct avd_hevc_run *run,
+ u16 log2_min_cb_size, u16 width,
+ u16 height, u32 pic_in_cts_width,
+ u32 pic_in_cts_height, u16 *column_width,
+ u16 *row_height)
+{
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ u32 sum = 0;
+ int i;
+
+ for (i = 0; i < pps->num_tile_columns_minus1; i++) {
+ column_width[i] = MIN(pps->column_width_minus1[i] + 1,
+ pic_in_cts_width);
+ sum += column_width[i];
+ }
+ column_width[i] = pic_in_cts_width - sum;
+
+ sum = 0;
+ for (i = 0; i < pps->num_tile_rows_minus1; i++) {
+ row_height[i] = MIN(pps->row_height_minus1[i] + 1,
+ pic_in_cts_height);
+ sum += row_height[i];
+ }
+ row_height[i] = pic_in_cts_height - sum;
+}
+
+static void compute_bd(struct avd_hevc_run *run, u32 *col_bd, u32 *row_bd,
+ u16 *column_width, u16 *row_height)
+{
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ int i;
+
+ for (col_bd[0] = 0, i = 0; i <= pps->num_tile_columns_minus1; i++)
+ col_bd[i + 1] = col_bd[i] + column_width[i];
+
+ for (row_bd[0] = 0, i = 0; i <= pps->num_tile_rows_minus1; i++)
+ row_bd[i + 1] = row_bd[i] + row_height[i];
+}
+
+static void compute_rs_to_ts(struct avd_hevc_run *run, u32 pic_in_ctbs_size,
+ u32 pic_in_ctbs_width, u32 *col_bd, u32 *row_bd,
+ u16 *col_width, u16 *row_height,
+ u32 *ctb_addr_rs_to_ts)
+{
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ int i, j, tb_x, tb_y, tile_x = 0, tile_y = 0;
+ u32 ctb_addr_rs;
+
+ for (ctb_addr_rs = 0; ctb_addr_rs < pic_in_ctbs_size; ctb_addr_rs++) {
+ tb_x = ctb_addr_rs % pic_in_ctbs_width;
+ tb_y = ctb_addr_rs / pic_in_ctbs_width;
+ for (i = 0; i <= pps->num_tile_columns_minus1; i++)
+ if (tb_x >= col_bd[i])
+ tile_x = i;
+ for (j = 0; j <= pps->num_tile_rows_minus1; j++)
+ if (tb_y >= row_bd[j])
+ tile_y = j;
+ ctb_addr_rs_to_ts[ctb_addr_rs] = 0;
+ for (i = 0; i < tile_x; i++)
+ ctb_addr_rs_to_ts[ctb_addr_rs] +=
+ row_height[tile_y] * col_width[i];
+ for (j = 0; j < tile_y; j++)
+ ctb_addr_rs_to_ts[ctb_addr_rs] +=
+ pic_in_ctbs_width * row_height[j];
+ ctb_addr_rs_to_ts[ctb_addr_rs] +=
+ (tb_y - row_bd[tile_y]) * col_width[tile_x] + tb_x -
+ col_bd[tile_x];
+ }
+}
+
+static void compute_tile_ids(struct avd_hevc_run *run, u32 pic_in_ctbs_width,
+ u32 *col_bd, u32 *row_bd, u32 *ctb_addr_rs_to_ts,
+ u32 *tile_ids)
+{
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ int j, i, y, x;
+ u32 tile_idx;
+
+ for (j = 0, tile_idx = 0; j <= pps->num_tile_rows_minus1; j++)
+ for (i = 0; i <= pps->num_tile_columns_minus1; i++, tile_idx++)
+ for (y = row_bd[j]; y < row_bd[j + 1]; y++)
+ for (x = col_bd[i]; x < col_bd[i + 1]; x++)
+ tile_ids[ctb_addr_rs_to_ts
+ [y * pic_in_ctbs_width +
+ x]] = tile_idx;
+}
+
+struct sl_ctx {
+ u32 ctx_col;
+ u32 ctx_row;
+ s32 q1_col;
+ s32 q1_row;
+};
+
+static void stream_slices(struct avd_ctx *ctx, struct avd_hevc_run *run)
+{
+ const struct v4l2_ctrl_hevc_sps *sps = run->sps;
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ const struct v4l2_ctrl_hevc_slice_params *sl;
+ struct avd_hevc_tile_info *tile_info = &run->tile_info;
+ bool tiles_enabled, first_slice, first_segment;
+ bool hflip, vflip;
+ int slice_segment_offset, entry_point_idx = 0, pos = 0, offset = 0;
+ int row, col, i, s, to;
+ int slice_flag, size, new_offset;
+ int tile_id, last_tile_id;
+ u16 log2_min_cb_size, width, height;
+ u32 max_cu_width, pic_in_ctbs_width, pic_in_ctbs_height;
+ u32 pic_in_ctbs_size, num_cols, last_tile_block = 0;
+ struct sl_ctx last = {
+ .q1_col = -1,
+ .q1_row = -1,
+ };
+
+ width = sps->pic_width_in_luma_samples;
+ height = sps->pic_height_in_luma_samples;
+
+ tiles_enabled = !!(pps->flags & V4L2_HEVC_PPS_FLAG_TILES_ENABLED);
+
+ log2_min_cb_size = sps->log2_min_luma_coding_block_size_minus3 + 3;
+
+ num_cols = pps->num_tile_columns_minus1 + 1;
+
+ max_cu_width = 1 << (sps->log2_diff_max_min_luma_coding_block_size +
+ log2_min_cb_size);
+ pic_in_ctbs_width = (width + max_cu_width - 1) / max_cu_width;
+ pic_in_ctbs_height = (height + max_cu_width - 1) / max_cu_width;
+ pic_in_ctbs_size = pic_in_ctbs_width * pic_in_ctbs_height;
+
+ for (s = 0; s < run->num_slices; s++) {
+ sl = &(*run->sl)[s];
+ slice_segment_offset = 0;
+
+ /* entry_points are not needed for WPP */
+ to = tiles_enabled ? sl->num_entry_point_offsets + 1 : 1;
+
+ for (i = 0; i < to; i++) {
+ first_segment = i == 0;
+ first_slice = s == 0;
+
+ if (tiles_enabled && to > 1) {
+ if (i < sl->num_entry_point_offsets) {
+ size = (*run->entry_point_offsets)
+ [entry_point_idx++];
+ new_offset = size;
+ } else {
+ size = sl->bit_size / 8 -
+ sl->data_byte_offset -
+ slice_segment_offset;
+ new_offset =
+ size + sl->data_byte_offset;
+ }
+ } else {
+ size = (sl->bit_size) / 8 -
+ sl->data_byte_offset;
+ new_offset = size + sl->data_byte_offset;
+ }
+
+ if (sl->slice_segment_addr > pic_in_ctbs_size)
+ return;
+
+ tile_id = tile_info->tile_ids
+ [tile_info->ctb_addr_rs_to_ts
+ [sl->slice_segment_addr]];
+ last_tile_id =
+ first_slice ?
+ -1 :
+ tile_info->tile_ids
+ [tile_info->ctb_addr_rs_to_ts
+ [(*run->sl)[s - 1]
+ .slice_segment_addr]];
+
+ slice_flag = 0;
+ if (first_slice || tile_id != last_tile_id ||
+ !first_segment)
+ slice_flag |= NEW_TILE_ID;
+
+ if (first_segment)
+ slice_flag |= NEW_SLICE;
+
+ vflip = false;
+ hflip = false;
+
+ row = pos / num_cols;
+ col = pos % num_cols;
+
+ /*
+ * in the JCT-VC-HEVC_V1 tests only TILES_B_Cisco_1
+ * seems affected and it only seems to need hflip.
+ *
+ * If there is a spec equivalent or a better way to
+ * represent this, it could not be found.
+ */
+ if (slice_flag & NEW_TILE_ID) {
+ if ((col >= last.ctx_col &&
+ row > last.ctx_row) ||
+ (col <= last.q1_col && row > last.q1_row))
+ vflip = true;
+
+ if (!(slice_flag & NEW_SLICE)) {
+ if (row && row == last.ctx_row + 1) {
+ hflip = true;
+ if (!vflip) {
+ last.q1_row = row;
+ last.q1_col = col;
+ }
+ }
+ } else {
+ last.ctx_row = row;
+ last.ctx_col = col;
+ last.q1_row = -1;
+ last.q1_col = -1;
+ }
+ }
+
+ set_slice(
+ ctx, run, sl, size,
+ offset + slice_segment_offset,
+ /* makes WPP more reliable? But not really ?? */
+ (sl->flags & V4L2_HEVC_SLICE_PARAMS_FLAG_DEPENDENT_SLICE_SEGMENT ?
+ slice_flag & ~NEW_SLICE :
+ slice_flag)
+ << 13);
+
+ last_tile_block = submit_slice_segment(
+ ctx, run, sl, row, col, tile_info->col_bd,
+ tile_info->row_bd, pic_in_ctbs_width,
+ pic_in_ctbs_height, first_slice, hflip, vflip,
+ slice_flag, last_tile_block);
+
+ if (slice_flag & NEW_TILE_ID)
+ pos++;
+ avd_end_segment(ctx, slice_flag & NEW_TILE_ID);
+
+ slice_segment_offset += new_offset;
+ }
+ offset += sl->bit_size / 8;
+ }
+}
+
+static void update_dec_buf_info(struct avd_decoded_buffer *buf,
+ const struct v4l2_ctrl_hevc_slice_params *sl)
+{
+ buf->hevc.is_intra = sl->slice_type == V4L2_HEVC_SLICE_TYPE_I;
+}
+
+static void avd_hevc_adjust_decoded_fmt(struct avd_ctx *ctx,
+ struct v4l2_pix_format_mplane *pix_mp)
+{
+ pix_mp->plane_fmt[0].sizeimage +=
+ mv_color_size(pix_mp->width, pix_mp->height);
+}
+
+static enum avd_image_fmt avd_hevc_get_image_fmt(struct avd_ctx *ctx,
+ struct v4l2_ctrl *ctrl)
+{
+ const struct v4l2_ctrl_hevc_sps *sps = ctrl->p_new.p_hevc_sps;
+
+ if (ctrl->id != V4L2_CID_STATELESS_HEVC_SPS)
+ return AVD_IMG_FMT_ANY;
+
+ /*
+ * TODO: we can do up to 4:4:4 12 bit
+ * not sure if v4l2 supports RExt
+ */
+
+ if (sps->bit_depth_luma_minus8 == 0) {
+ if (sps->chroma_format_idc == 2)
+ return AVD_IMG_FMT_422_8BIT;
+ else
+ return AVD_IMG_FMT_420_8BIT;
+ } else if (sps->bit_depth_luma_minus8 == 2) {
+ if (sps->chroma_format_idc == 2)
+ return AVD_IMG_FMT_422_10BIT;
+ else
+ return AVD_IMG_FMT_420_10BIT;
+ }
+
+ return AVD_IMG_FMT_ANY;
+}
+
+static int avd_hevc_validate_sps(struct avd_ctx *ctx,
+ const struct v4l2_ctrl_hevc_sps *sps)
+{
+ if (sps->pic_width_in_luma_samples > ctx->coded_fmt.fmt.pix_mp.width ||
+ sps->pic_height_in_luma_samples > ctx->coded_fmt.fmt.pix_mp.height)
+ return -EINVAL;
+
+ if (sps->bit_depth_chroma_minus8 != sps->bit_depth_luma_minus8)
+ return -EINVAL;
+
+ if (sps->bit_depth_luma_minus8 > 2)
+ /* not implemented */
+ return -EINVAL;
+
+ return 0;
+}
+
+static int avd_hevc_validate_pps(struct avd_ctx *ctx,
+ const struct v4l2_ctrl_hevc_pps *pps)
+{
+ if (pps->num_tile_columns_minus1 + 1 > HEVC_MAX_TILE_COLS ||
+ pps->num_tile_rows_minus1 + 1 > HEVC_MAX_TILE_ROWS)
+ return -EINVAL;
+
+ return 0;
+}
+
+static int avd_hevc_start(struct avd_ctx *ctx)
+{
+ struct avd_hevc_ctx *hevc_ctx;
+
+ hevc_ctx = kzalloc_obj(*hevc_ctx, GFP_KERNEL);
+ if (!hevc_ctx)
+ return -ENOMEM;
+
+ ctx->priv = hevc_ctx;
+
+ return 0;
+}
+
+static int avd_hevc_alloc_scratch(struct avd_ctx *ctx, struct avd_hevc_run *run)
+{
+ struct avd_hevc_ctx *hevc_ctx = ctx->priv;
+ struct avd_dev *avd = ctx->dev;
+ const struct v4l2_ctrl_hevc_sps *sps = run->sps;
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ struct avd_hevc_tile_info *tile_info = &run->tile_info;
+ u32 i, w, bit_depth, max_col = 0, max_row = 0, cols, rows;
+ u32 log2_min_cb_size, max_cu_width;
+ int ret;
+
+ cols = pps->num_tile_columns_minus1 + 1;
+ rows = pps->num_tile_rows_minus1 + 1;
+ bit_depth = sps->bit_depth_luma_minus8 + 8;
+ w = sps->pic_width_in_luma_samples;
+
+ log2_min_cb_size = sps->log2_min_luma_coding_block_size_minus3 + 3;
+ max_cu_width = 1 << (sps->log2_diff_max_min_luma_coding_block_size +
+ log2_min_cb_size);
+ if (!max_cu_width)
+ return -EINVAL;
+
+ for (i = 0; i < cols; i++)
+ max_col = tile_info->col_width[i] > max_col ?
+ tile_info->col_width[i] :
+ max_col;
+
+ for (i = 0; i < rows; i++)
+ max_row = tile_info->row_height[i] > max_row ?
+ tile_info->row_height[i] :
+ max_row;
+
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.ip_above,
+ ((max_col * max_cu_width) / 4) * bit_depth);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.mv_above_info,
+ ((max_col * max_cu_width) / 16) * 20);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.lf_above,
+ DIV_ROUND_UP(w, 16) * bit_depth * (10 + 6) +
+ (cols - 1) * 256);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.lf_above_info,
+ DIV_ROUND_UP(w + 7, 16) * 36 + cols * 128);
+ if (ret)
+ return ret;
+
+ if (pps->flags & V4L2_HEVC_PPS_FLAG_TILES_ENABLED) {
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.lf_left,
+ ((max_row * max_cu_width) / 4) * 36 +
+ 144 /* why? */);
+ if (ret)
+ return ret;
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.lf_left_info,
+ ((max_row * max_cu_width) / 4) * 9);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.az_above,
+ /* not sure if its cols or rows */
+ DIV_ROUND_UP(w, 64) * 144 +
+ (cols - 1) * 128);
+ if (ret)
+ return ret;
+
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.sw_left,
+ ((max_row * max_cu_width) / 4) * 216);
+ if (ret)
+ return ret;
+ } else {
+ ret = avd_buf_alloc(avd, &hevc_ctx->bufs.az_above,
+ DIV_ROUND_UP(w + 7, 16) * 4 * bit_depth);
+ if (ret)
+ return ret;
+ }
+
+ return 0;
+}
+
+static int avd_hevc_compute_tiles(struct avd_ctx *ctx, struct avd_hevc_run *run)
+{
+ const struct v4l2_ctrl_hevc_sps *sps = run->sps;
+ const struct v4l2_ctrl_hevc_pps *pps = run->pps;
+ bool tiles_enabled;
+ struct avd_hevc_tile_info *tile_info = &run->tile_info;
+ u16 log2_min_cb_size, width, height;
+ u32 max_cu_width, pic_in_ctbs_width, pic_in_ctbs_height,
+ pic_in_ctbs_size;
+
+ width = sps->pic_width_in_luma_samples;
+ height = sps->pic_height_in_luma_samples;
+
+ log2_min_cb_size = sps->log2_min_luma_coding_block_size_minus3 + 3;
+
+ max_cu_width = 1 << (sps->log2_diff_max_min_luma_coding_block_size +
+ log2_min_cb_size);
+ pic_in_ctbs_width = (width + max_cu_width - 1) / max_cu_width;
+ pic_in_ctbs_height = (height + max_cu_width - 1) / max_cu_width;
+ pic_in_ctbs_size = pic_in_ctbs_height * pic_in_ctbs_width;
+
+ tiles_enabled = !!(pps->flags & V4L2_HEVC_PPS_FLAG_TILES_ENABLED);
+
+ tile_info->ctb_addr_rs_to_ts = kzalloc_objs(
+ *tile_info->ctb_addr_rs_to_ts, pic_in_ctbs_size, GFP_KERNEL);
+ if (!tile_info->ctb_addr_rs_to_ts)
+ return -ENOMEM;
+
+ tile_info->tile_ids = kzalloc_objs(*tile_info->tile_ids,
+ pic_in_ctbs_size, GFP_KERNEL);
+ if (!tile_info->tile_ids)
+ return -ENOMEM;
+
+ if (tiles_enabled) {
+ if (pps->flags & V4L2_HEVC_PPS_FLAG_UNIFORM_SPACING) {
+ compute_tiles_uniform(run, log2_min_cb_size, width,
+ height, pic_in_ctbs_width,
+ pic_in_ctbs_height,
+ tile_info->col_width,
+ tile_info->row_height);
+ } else {
+ compute_tiles_non_uniform(run, log2_min_cb_size, width,
+ height, pic_in_ctbs_width,
+ pic_in_ctbs_height,
+ tile_info->col_width,
+ tile_info->row_height);
+ }
+
+ compute_bd(run, tile_info->col_bd, tile_info->row_bd,
+ tile_info->col_width, tile_info->row_height);
+
+ /*
+ * 6.5.1 CTB raster and tile scanning conversion process
+ * (6-7) and (6-9)
+ *
+ * we really only need to know if tileidx != last tileidx
+ * so maybe there is a better way
+ */
+ compute_rs_to_ts(run, pic_in_ctbs_size, pic_in_ctbs_width,
+ tile_info->col_bd, tile_info->row_bd,
+ tile_info->col_width, tile_info->row_height,
+ tile_info->ctb_addr_rs_to_ts);
+ compute_tile_ids(run, pic_in_ctbs_width, tile_info->col_bd,
+ tile_info->row_bd,
+ tile_info->ctb_addr_rs_to_ts,
+ tile_info->tile_ids);
+ } else {
+ tile_info->col_width[0] =
+ (width + max_cu_width - 1) / max_cu_width;
+ tile_info->row_height[0] =
+ (height + max_cu_width - 1) / max_cu_width;
+ }
+
+ return 0;
+}
+
+static void avd_hevc_stop(struct avd_ctx *ctx)
+{
+ struct avd_hevc_ctx *hevc_ctx = ctx->priv;
+ struct avd_dev *avd = ctx->dev;
+
+ if (!hevc_ctx)
+ return;
+
+ avd_buf_free(avd, &hevc_ctx->bufs.mv_above_info);
+ avd_buf_free(avd, &hevc_ctx->bufs.az_above);
+ avd_buf_free(avd, &hevc_ctx->bufs.ip_above);
+ avd_buf_free(avd, &hevc_ctx->bufs.lf_above);
+ avd_buf_free(avd, &hevc_ctx->bufs.lf_above_info);
+ avd_buf_free(avd, &hevc_ctx->bufs.lf_left);
+ avd_buf_free(avd, &hevc_ctx->bufs.lf_left_info);
+ avd_buf_free(avd, &hevc_ctx->bufs.sw_left);
+
+ kfree(hevc_ctx);
+}
+
+static int avd_hevc_run_preamble(struct avd_ctx *ctx, struct avd_hevc_run *run)
+{
+ struct v4l2_ctrl *ctrl;
+ u32 dst_len, mv_color_len;
+ int i, sum = 0;
+
+ avd_run_preamble(ctx, &run->base);
+
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_HEVC_DECODE_PARAMS);
+ run->decode = ctrl ? ctrl->p_cur.p : NULL;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl, V4L2_CID_STATELESS_HEVC_SPS);
+ run->sps = ctrl ? ctrl->p_cur.p : NULL;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl, V4L2_CID_STATELESS_HEVC_PPS);
+ run->pps = ctrl ? ctrl->p_cur.p : NULL;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_HEVC_SCALING_MATRIX);
+ run->scaling_matrix = ctrl ? ctrl->p_cur.p : NULL;
+
+ /*
+ * save a pointer to the pointer in case the arrays are resized (and
+ * the pointer is updated)
+ */
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_HEVC_ENTRY_POINT_OFFSETS);
+ run->entry_point_offsets = ctrl ? (const u32 **)&ctrl->p_cur.p : NULL;
+ run->num_entry_point_offsets = ctrl ? ctrl->elems : 0;
+ ctrl = v4l2_ctrl_find(&ctx->ctrl_hdl,
+ V4L2_CID_STATELESS_HEVC_SLICE_PARAMS);
+ run->sl = ctrl ?
+ (const struct v4l2_ctrl_hevc_slice_params **)&ctrl->p_cur.p
+ : NULL;
+ run->num_slices = ctrl ? ctrl->elems : 0;
+ if (!run->sl || !run->num_slices)
+ return -EINVAL;
+
+ for (i = 0; i < run->num_slices; i++)
+ sum += (*run->sl)[i].num_entry_point_offsets;
+
+ if (sum && sum != run->num_entry_point_offsets)
+ return -EINVAL;
+
+ dst_len = run->base.bufs.dst->vb2_buf.planes[0].length;
+
+ mv_color_len = mv_color_size(fmt_width(ctx), fmt_height(ctx));
+
+ run->addresses.mv_color = run->base.y_out + (dst_len - mv_color_len);
+ return 0;
+}
+
+static int avd_hevc_run(struct avd_ctx *ctx)
+{
+ struct avd_hevc_run run = {};
+ struct avd_decoded_buffer *dst;
+ int ret;
+
+ ret = avd_hevc_run_preamble(ctx, &run);
+ if (ret)
+ goto postamble;
+
+ dst = vb2_to_avd_decoded_buf(&run.base.bufs.dst->vb2_buf);
+ update_dec_buf_info(dst, &(*run.sl)[0]);
+
+ ret = avd_hevc_compute_tiles(ctx, &run);
+ if (ret)
+ goto done;
+
+ ret = avd_hevc_alloc_scratch(ctx, &run);
+ if (ret)
+ goto done;
+
+ ret = avd_init_job(ctx, AVD_CODEC_HEVC,
+ run.num_slices + run.num_entry_point_offsets + 1);
+ if (ret)
+ goto done;
+
+ set_header(ctx, &run);
+ avd_end_segment(ctx, false);
+
+ stream_slices(ctx, &run);
+
+ ret = avd_submit_job(ctx);
+
+done:
+ kfree(run.tile_info.ctb_addr_rs_to_ts);
+ kfree(run.tile_info.tile_ids);
+
+postamble:
+ avd_run_postamble(ctx, &run.base);
+ return ret;
+}
+
+static int avd_hevc_try_ctrl(struct avd_ctx *ctx, struct v4l2_ctrl *ctrl)
+{
+ if (ctrl->id == V4L2_CID_STATELESS_HEVC_SPS)
+ return avd_hevc_validate_sps(ctx, ctrl->p_new.p_hevc_sps);
+
+ if (ctrl->id == V4L2_CID_STATELESS_HEVC_PPS)
+ return avd_hevc_validate_pps(ctx, ctrl->p_new.p_hevc_pps);
+
+ return 0;
+}
+
+const struct avd_coded_fmt_ops avd_hevc_fmt_ops = {
+ .adjust_decoded_fmt = avd_hevc_adjust_decoded_fmt,
+ .start = avd_hevc_start,
+ .stop = avd_hevc_stop,
+ .run = avd_hevc_run,
+ .try_ctrl = avd_hevc_try_ctrl,
+ .get_image_fmt = avd_hevc_get_image_fmt,
+};
diff --git a/drivers/media/platform/apple/avd/avd-v4l2.c b/drivers/media/platform/apple/avd/avd-v4l2.c
index c7958cf65ebd..70cabae8e66c 100644
--- a/drivers/media/platform/apple/avd/avd-v4l2.c
+++ b/drivers/media/platform/apple/avd/avd-v4l2.c
@@ -201,6 +201,64 @@ const struct v4l2_ctrl_ops avd_ctrl_ops = {
.s_ctrl = avd_s_ctrl,
};
+static const struct avd_ctrl_desc avd_hevc_ctrl_descs[] = {
+ {
+ .cfg.id = V4L2_CID_STATELESS_HEVC_SLICE_PARAMS,
+ .cfg.flags = V4L2_CTRL_FLAG_DYNAMIC_ARRAY,
+ .cfg.type = V4L2_CTRL_TYPE_HEVC_SLICE_PARAMS,
+ .cfg.dims = { 1800 },
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_HEVC_SPS,
+ .cfg.ops = &avd_ctrl_ops,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_HEVC_PPS,
+ .cfg.ops = &avd_ctrl_ops,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_HEVC_SCALING_MATRIX,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_HEVC_DECODE_PARAMS,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_HEVC_DECODE_MODE,
+ .cfg.min = V4L2_STATELESS_HEVC_DECODE_MODE_FRAME_BASED,
+ .cfg.max = V4L2_STATELESS_HEVC_DECODE_MODE_FRAME_BASED,
+ .cfg.def = V4L2_STATELESS_HEVC_DECODE_MODE_FRAME_BASED,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_HEVC_START_CODE,
+ .cfg.min = V4L2_STATELESS_HEVC_START_CODE_NONE,
+ .cfg.def = V4L2_STATELESS_HEVC_START_CODE_NONE,
+ .cfg.max = V4L2_STATELESS_HEVC_START_CODE_NONE,
+ },
+ {
+ .cfg.id = V4L2_CID_STATELESS_HEVC_ENTRY_POINT_OFFSETS,
+ .cfg.flags = V4L2_CTRL_FLAG_DYNAMIC_ARRAY,
+ .cfg.dims = { 1760 },
+ .cfg.max = 0xffffffff,
+ .cfg.step = 1,
+ },
+ {
+ .cfg.id = V4L2_CID_MPEG_VIDEO_HEVC_PROFILE,
+ .cfg.min = V4L2_MPEG_VIDEO_HEVC_PROFILE_MAIN,
+ .cfg.max = V4L2_MPEG_VIDEO_HEVC_PROFILE_MAIN_10,
+ .cfg.def = V4L2_MPEG_VIDEO_HEVC_PROFILE_MAIN,
+ },
+ {
+ .cfg.id = V4L2_CID_MPEG_VIDEO_HEVC_LEVEL,
+ .cfg.min = V4L2_MPEG_VIDEO_HEVC_LEVEL_1,
+ .cfg.max = V4L2_MPEG_VIDEO_HEVC_LEVEL_5_1,
+ },
+};
+
+static const struct avd_ctrls avd_hevc_ctrls = {
+ .ctrls = avd_hevc_ctrl_descs,
+ .num_ctrls = ARRAY_SIZE(avd_hevc_ctrl_descs),
+};
+
static const struct avd_ctrl_desc avd_h264_ctrl_descs[] = {
{
.cfg.id = V4L2_CID_STATELESS_H264_DECODE_PARAMS,
@@ -279,6 +337,20 @@ static const struct avd_ctrls avd_vp9_ctrls = {
};
static const struct avd_coded_fmt_desc avd_coded_fmts[] = {
+ {
+ .fourcc = V4L2_PIX_FMT_HEVC_SLICE,
+ .frmsize = {
+ .min_width = 64,
+ .max_width = 16384,
+ .step_width = 64,
+ .min_height = 64,
+ .max_height = 16384,
+ .step_height = 16,
+ },
+ .ctrls = &avd_hevc_ctrls,
+ .ops = &avd_hevc_fmt_ops,
+ .capability = AVD_CAPABILITY_HEVC,
+ },
{
.fourcc = V4L2_PIX_FMT_H264_SLICE,
.frmsize = {
diff --git a/drivers/media/platform/apple/avd/avd.h b/drivers/media/platform/apple/avd/avd.h
index bd6177ab1eb2..828d25f85902 100644
--- a/drivers/media/platform/apple/avd/avd.h
+++ b/drivers/media/platform/apple/avd/avd.h
@@ -104,6 +104,10 @@ struct avd_vp9_decoded_buffer_info {
unsigned int bit_depth : 4;
};
+struct avd_hevc_decoded_buffer_info {
+ bool is_intra;
+};
+
struct avd_comp {
u32 size;
/* offset to start of compressed data */
@@ -118,6 +122,7 @@ struct avd_decoded_buffer {
struct avd_comp comp;
union {
struct avd_vp9_decoded_buffer_info vp9;
+ struct avd_hevc_decoded_buffer_info hevc;
};
};
@@ -252,6 +257,7 @@ void avd_run_preamble(struct avd_ctx *ctx, struct avd_run *run);
void avd_run_postamble(struct avd_ctx *ctx, struct avd_run *run);
extern const struct avd_coded_fmt_ops avd_h264_fmt_ops;
+extern const struct avd_coded_fmt_ops avd_hevc_fmt_ops;
extern const struct avd_coded_fmt_ops avd_vp9_fmt_ops;
extern const struct v4l2_ctrl_ops avd_ctrl_ops;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 11/17] arm64: dts: apple: t8103: add avd nodes
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (8 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 09/17] media: apple: avd: add hevc support Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 12/17] arm64: dts: apple: t8112: " Sofus Forstreuter
` (6 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
arch/arm64/boot/dts/apple/t8103.dtsi | 32 ++++++++++++++++++++++++++++++++
1 file changed, 32 insertions(+)
diff --git a/arch/arm64/boot/dts/apple/t8103.dtsi b/arch/arm64/boot/dts/apple/t8103.dtsi
index 5b68bf0df4e6..530290e7cb32 100644
--- a/arch/arm64/boot/dts/apple/t8103.dtsi
+++ b/arch/arm64/boot/dts/apple/t8103.dtsi
@@ -984,6 +984,38 @@ pinctrl_aop: pinctrl@24a820000 {
<AIC_IRQ 274 IRQ_TYPE_LEVEL_HIGH>;
};
+ avd_dart: iommu@269010000 {
+ compatible = "apple,t8103-dart";
+ reg = <0x2 0x69010000 0x0 0x4000>;
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 547 IRQ_TYPE_LEVEL_HIGH>;
+ #iommu-cells = <1>;
+ power-domains = <&ps_avd_sys>;
+ };
+
+ video-codec@269070000 {
+ compatible = "apple,t8103-avd";
+ reg = <0x2 0x69070000 0x0 0x4000>,
+ <0x2 0x69080000 0x0 0xc000>,
+ <0x2 0x6908c000 0x0 0xc000>,
+ <0x2 0x69098000 0x0 0x4000>,
+ <0x2 0x69100000 0x0 0x10000>;
+ reg-names = "piodma", "code", "sram", "mbox", "ctrl";
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 540 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 541 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 542 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 543 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 544 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 545 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 546 IRQ_TYPE_LEVEL_HIGH>;
+ interrupt-names = "mbox0", "mbox1", "mbox2", "mbox3",
+ "flag0", "flag1", "piodma";
+ iommus = <&avd_dart 0>, <&avd_dart 1>;
+ power-domains = <&ps_avd_sys>;
+ resets = <&ps_avd_sys>;
+ };
+
ans_mbox: mbox@277408000 {
compatible = "apple,t8103-asc-mailbox", "apple,asc-mailbox-v4";
reg = <0x2 0x77408000 0x0 0x4000>;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 12/17] arm64: dts: apple: t8112: add avd nodes
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (9 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 11/17] arm64: dts: apple: t8103: add avd nodes Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 13/17] arm64: dts: apple: t8122: " Sofus Forstreuter
` (5 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
arch/arm64/boot/dts/apple/t8112.dtsi | 32 ++++++++++++++++++++++++++++++++
1 file changed, 32 insertions(+)
diff --git a/arch/arm64/boot/dts/apple/t8112.dtsi b/arch/arm64/boot/dts/apple/t8112.dtsi
index ec248ca052cb..48ab71565eab 100644
--- a/arch/arm64/boot/dts/apple/t8112.dtsi
+++ b/arch/arm64/boot/dts/apple/t8112.dtsi
@@ -987,6 +987,38 @@ pinctrl_aop: pinctrl@24a820000 {
<AIC_IRQ 307 IRQ_TYPE_LEVEL_HIGH>;
};
+ avd_dart: iommu@269010000 {
+ compatible = "apple,t8110-dart";
+ reg = <0x2 0x69010000 0x0 0x4000>;
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 674 IRQ_TYPE_LEVEL_HIGH>;
+ #iommu-cells = <1>;
+ power-domains = <&ps_avd_sys>;
+ };
+
+ video-codec@269070000 {
+ compatible = "apple,t8112-avd";
+ reg = <0x2 0x69070000 0x0 0x4000>,
+ <0x2 0x69080000 0x0 0x10000>,
+ <0x2 0x69090000 0x0 0x10000>,
+ <0x2 0x690a0000 0x0 0x4000>,
+ <0x2 0x69100000 0x0 0x10000>;
+ reg-names = "piodma", "code", "sram", "mbox", "ctrl";
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 667 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 668 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 669 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 670 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 671 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 672 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 673 IRQ_TYPE_LEVEL_HIGH>;
+ interrupt-names = "mbox0", "mbox1", "mbox2", "mbox3",
+ "flag0", "flag1", "piodma";
+ iommus = <&avd_dart 0>, <&avd_dart 1>;
+ power-domains = <&ps_avd_sys>;
+ resets = <&ps_avd_sys>;
+ };
+
ans_mbox: mbox@277408000 {
compatible = "apple,t8112-asc-mailbox", "apple,asc-mailbox-v4";
reg = <0x2 0x77408000 0x0 0x4000>;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 13/17] arm64: dts: apple: t8122: add avd nodes
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (10 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 12/17] arm64: dts: apple: t8112: " Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 14/17] arm64: dts: apple: t600x: " Sofus Forstreuter
` (4 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
arch/arm64/boot/dts/apple/t8122.dtsi | 32 ++++++++++++++++++++++++++++++++
1 file changed, 32 insertions(+)
diff --git a/arch/arm64/boot/dts/apple/t8122.dtsi b/arch/arm64/boot/dts/apple/t8122.dtsi
index 1ee61c5b3409..25e47755e095 100644
--- a/arch/arm64/boot/dts/apple/t8122.dtsi
+++ b/arch/arm64/boot/dts/apple/t8122.dtsi
@@ -186,6 +186,38 @@ soc {
/* Required to get >32-bit DMA via DARTs */
dma-ranges = <0 0 0 0 0xffffffff 0xffffc000>;
+ avd_dart: iommu@289010000 {
+ compatible = "apple,t8122-dart", "apple,t8110-dart";
+ reg = <0x2 0x89010000 0x0 0x4000>;
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 686 IRQ_TYPE_LEVEL_HIGH>;
+ #iommu-cells = <1>;
+ power-domains = <&ps_avd_sys>;
+ };
+
+ video-codec@289070000 {
+ compatible = "apple,t8122-avd";
+ reg = <0x2 0x89070000 0x0 0x4000>,
+ <0x2 0x89080000 0x0 0x10000>,
+ <0x2 0x89090000 0x0 0x14000>,
+ <0x2 0x890a4000 0x0 0x4000>,
+ <0x2 0x89100000 0x0 0x10000>;
+ reg-names = "piodma", "code", "sram", "mbox", "ctrl";
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 679 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 680 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 681 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 682 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 683 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 684 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 685 IRQ_TYPE_LEVEL_HIGH>;
+ interrupt-names = "mbox0", "mbox1", "mbox2", "mbox3",
+ "flag0", "flag1", "piodma";
+ iommus = <&avd_dart 0>, <&avd_dart 1>;
+ power-domains = <&ps_avd_sys>;
+ resets = <&ps_avd_sys>;
+ };
+
i2c0: i2c@2a1010000 {
compatible = "apple,t8122-i2c", "apple,t8103-i2c";
reg = <0x2 0xa1010000 0x0 0x4000>;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 14/17] arm64: dts: apple: t600x: add avd nodes
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (11 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 13/17] arm64: dts: apple: t8122: " Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 15/17] arm64: dts: apple: t602x: " Sofus Forstreuter
` (3 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
arch/arm64/boot/dts/apple/t600x-dieX.dtsi | 32 +++++++++++++++++++++++++++++++
1 file changed, 32 insertions(+)
diff --git a/arch/arm64/boot/dts/apple/t600x-dieX.dtsi b/arch/arm64/boot/dts/apple/t600x-dieX.dtsi
index 9676d5127039..c3edcb78d016 100644
--- a/arch/arm64/boot/dts/apple/t600x-dieX.dtsi
+++ b/arch/arm64/boot/dts/apple/t600x-dieX.dtsi
@@ -24,6 +24,38 @@ DIE_NODE(cpufreq_p1): cpufreq@212e20000 {
#performance-domain-cells = <0>;
};
+ DIE_NODE(avd_dart): iommu@287010000 {
+ compatible = "apple,t8110-dart";
+ reg = <0x2 0x87010000 0x0 0x4000>;
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ DIE_NO 1018 IRQ_TYPE_LEVEL_HIGH>;
+ #iommu-cells = <1>;
+ power-domains = <&DIE_NODE(ps_avd_sys)>;
+ };
+
+ video-codec@287070000 {
+ compatible = "apple,t6000-avd";
+ reg = <0x2 0x87070000 0x0 0x4000>,
+ <0x2 0x87080000 0x0 0x10000>,
+ <0x2 0x87090000 0x0 0x10000>,
+ <0x2 0x870a0000 0x0 0x4000>,
+ <0x2 0x87100000 0x0 0x10000>;
+ reg-names = "piodma", "code", "sram", "mbox", "ctrl";
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ DIE_NO 1011 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1012 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1013 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1014 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1015 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1016 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1017 IRQ_TYPE_LEVEL_HIGH>;
+ interrupt-names = "mbox0", "mbox1", "mbox2", "mbox3",
+ "flag0", "flag1", "piodma";
+ iommus = <&DIE_NODE(avd_dart) 0>, <&DIE_NODE(avd_dart) 1>;
+ power-domains = <&DIE_NODE(ps_avd_sys)>;
+ resets = <&DIE_NODE(ps_avd_sys)>;
+ };
+
DIE_NODE(pmgr): power-management@28e080000 {
compatible = "apple,t6000-pmgr", "apple,pmgr", "syscon", "simple-mfd";
#address-cells = <1>;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 15/17] arm64: dts: apple: t602x: add avd nodes
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (12 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 14/17] arm64: dts: apple: t600x: " Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 16/17] arm64: dts: apple: t6030: " Sofus Forstreuter
` (2 subsequent siblings)
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
arch/arm64/boot/dts/apple/t602x-dieX.dtsi | 32 +++++++++++++++++++++++++++++++
1 file changed, 32 insertions(+)
diff --git a/arch/arm64/boot/dts/apple/t602x-dieX.dtsi b/arch/arm64/boot/dts/apple/t602x-dieX.dtsi
index ae3d535c5acb..5f744308a6c5 100644
--- a/arch/arm64/boot/dts/apple/t602x-dieX.dtsi
+++ b/arch/arm64/boot/dts/apple/t602x-dieX.dtsi
@@ -23,6 +23,38 @@ DIE_NODE(cpufreq_p1): cpufreq@212e20000 {
#performance-domain-cells = <0>;
};
+ DIE_NODE(avd_dart): iommu@287010000 {
+ compatible = "apple,t6020-dart", "apple,t8110-dart";
+ reg = <0x2 0x87010000 0x0 0x4000>;
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ DIE_NO 1095 IRQ_TYPE_LEVEL_HIGH>;
+ #iommu-cells = <1>;
+ power-domains = <&DIE_NODE(ps_avd_sys)>;
+ };
+
+ video-codec@287070000 {
+ compatible = "apple,t6020-avd";
+ reg = <0x2 0x87070000 0x0 0x4000>,
+ <0x2 0x87080000 0x0 0x10000>,
+ <0x2 0x87090000 0x0 0x10000>,
+ <0x2 0x870a0000 0x0 0x4000>,
+ <0x2 0x87100000 0x0 0x10000>;
+ reg-names = "piodma", "code", "sram", "mbox", "ctrl";
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ DIE_NO 1088 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1089 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1090 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1091 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1092 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1093 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1094 IRQ_TYPE_LEVEL_HIGH>;
+ interrupt-names = "mbox0", "mbox1", "mbox2", "mbox3",
+ "flag0", "flag1", "piodma";
+ iommus = <&DIE_NODE(avd_dart) 0>, <&DIE_NODE(avd_dart) 1>;
+ power-domains = <&DIE_NODE(ps_avd_sys)>;
+ resets = <&DIE_NODE(ps_avd_sys)>;
+ };
+
DIE_NODE(pmgr): power-management@28e080000 {
compatible = "apple,t6020-pmgr", "apple,t8103-pmgr", "syscon", "simple-mfd";
#address-cells = <1>;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 16/17] arm64: dts: apple: t6030: add avd nodes
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (13 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 15/17] arm64: dts: apple: t602x: " Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-09-26 13:14 ` [PATCH v2 17/17] arm64: dts: apple: t6031: " Sofus Forstreuter
2026-10-05 22:29 ` [PATCH v2 00/17] media: apple: add avd driver Neal Gompa
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
arch/arm64/boot/dts/apple/t6030.dtsi | 32 ++++++++++++++++++++++++++++++++
1 file changed, 32 insertions(+)
diff --git a/arch/arm64/boot/dts/apple/t6030.dtsi b/arch/arm64/boot/dts/apple/t6030.dtsi
index dda9568af11f..805042bd9270 100644
--- a/arch/arm64/boot/dts/apple/t6030.dtsi
+++ b/arch/arm64/boot/dts/apple/t6030.dtsi
@@ -361,6 +361,38 @@ pmgr_gfx: power-management@290e80000 {
/* child nodes are added in t6030-pmgr.dtsi */
};
+ avd_dart: iommu@30b010000 {
+ compatible = "apple,t6030-dart", "apple,t8110-dart";
+ reg = <0x3 0x0b010000 0x0 0x4000>;
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 795 IRQ_TYPE_LEVEL_HIGH>;
+ #iommu-cells = <1>;
+ power-domains = <&ps_avd_sys>;
+ };
+
+ video-codec@30b070000 {
+ compatible = "apple,t6030-avd", "apple,t8122-avd";
+ reg = <0x3 0x0b070000 0x0 0x4000>,
+ <0x3 0x0b080000 0x0 0x10000>,
+ <0x3 0x0b090000 0x0 0x14000>,
+ <0x3 0x0b0a4000 0x0 0x4000>,
+ <0x3 0x0b100000 0x0 0x10000>;
+ reg-names = "piodma", "code", "sram", "mbox", "ctrl";
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ 788 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 789 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 790 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 791 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 792 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 793 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ 794 IRQ_TYPE_LEVEL_HIGH>;
+ interrupt-names = "mbox0", "mbox1", "mbox2", "mbox3",
+ "flag0", "flag1", "piodma";
+ iommus = <&avd_dart 0>, <&avd_dart 1>;
+ power-domains = <&ps_avd_sys>;
+ resets = <&ps_avd_sys>;
+ };
+
pinctrl_ap: pinctrl@347100000 {
compatible = "apple,t6030-pinctrl", "apple,t8103-pinctrl";
reg = <0x3 0x47100000 0x0 0x4000>;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* [PATCH v2 17/17] arm64: dts: apple: t6031: add avd nodes
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (14 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 16/17] arm64: dts: apple: t6030: " Sofus Forstreuter
@ 2026-09-26 13:14 ` Sofus Forstreuter
2026-10-05 22:29 ` [PATCH v2 00/17] media: apple: add avd driver Neal Gompa
16 siblings, 0 replies; 19+ messages in thread
From: Sofus Forstreuter @ 2026-09-26 13:14 UTC (permalink / raw)
To: Sven Peter, Janne Grunau, Neal Gompa, Mauro Carvalho Chehab,
Rob Herring, Krzysztof Kozlowski, Conor Dooley, Sofus Forstreuter,
Philipp Zabel, Heiko Stuebner
Cc: asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
---
arch/arm64/boot/dts/apple/t6031-dieX.dtsi | 32 +++++++++++++++++++++++++++++++
1 file changed, 32 insertions(+)
diff --git a/arch/arm64/boot/dts/apple/t6031-dieX.dtsi b/arch/arm64/boot/dts/apple/t6031-dieX.dtsi
index 2e9d3d2e6dea..3b21a7aae409 100644
--- a/arch/arm64/boot/dts/apple/t6031-dieX.dtsi
+++ b/arch/arm64/boot/dts/apple/t6031-dieX.dtsi
@@ -73,6 +73,38 @@ DIE_NODE(pinctrl_aop): pinctrl@2a8824000 {
<AIC_IRQ DIE_NO 601 IRQ_TYPE_LEVEL_HIGH>;
};
+ DIE_NODE(avd_dart): iommu@2cd010000 {
+ compatible = "apple,t6031-dart", "apple,t8110-dart";
+ reg = <0x2 0xcd010000 0x0 0x4000>;
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ DIE_NO 1192 IRQ_TYPE_LEVEL_HIGH>;
+ #iommu-cells = <1>;
+ power-domains = <&DIE_NODE(ps_avd_sys)>;
+ };
+
+ video-codec@2cd070000 {
+ compatible = "apple,t6031-avd", "apple,t8122-avd";
+ reg = <0x2 0xcd070000 0x0 0x4000>,
+ <0x2 0xcd080000 0x0 0x10000>,
+ <0x2 0xcd090000 0x0 0x14000>,
+ <0x2 0xcd0a4000 0x0 0x4000>,
+ <0x2 0xcd100000 0x0 0x10000>;
+ reg-names = "piodma", "code", "sram", "mbox", "ctrl";
+ interrupt-parent = <&aic>;
+ interrupts = <AIC_IRQ DIE_NO 1185 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1186 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1187 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1188 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1189 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1190 IRQ_TYPE_LEVEL_HIGH>,
+ <AIC_IRQ DIE_NO 1191 IRQ_TYPE_LEVEL_HIGH>;
+ interrupt-names = "mbox0", "mbox1", "mbox2", "mbox3",
+ "flag0", "flag1", "piodma";
+ iommus = <&DIE_NODE(avd_dart) 0>, <&DIE_NODE(avd_dart) 1>;
+ power-domains = <&DIE_NODE(ps_avd_sys)>;
+ resets = <&DIE_NODE(ps_avd_sys)>;
+ };
+
DIE_NODE(pinctrl_ap): pinctrl@2b3000000 {
compatible = "apple,t6031-pinctrl", "apple,t8103-pinctrl";
reg = <0x2 0xb3000000 0x0 0x4000>;
--
2.55.0
^ permalink raw reply related [flat|nested] 19+ messages in thread* Re: [PATCH v2 00/17] media: apple: add avd driver
2026-09-26 13:14 [PATCH v2 00/17] media: apple: add avd driver Sofus Forstreuter
` (15 preceding siblings ...)
2026-09-26 13:14 ` [PATCH v2 17/17] arm64: dts: apple: t6031: " Sofus Forstreuter
@ 2026-10-05 22:29 ` Neal Gompa
16 siblings, 0 replies; 19+ messages in thread
From: Neal Gompa @ 2026-10-05 22:29 UTC (permalink / raw)
To: Sofus Forstreuter
Cc: Sven Peter, Janne Grunau, Mauro Carvalho Chehab, Rob Herring,
Krzysztof Kozlowski, Conor Dooley, Philipp Zabel, Heiko Stuebner,
asahi, linux-arm-kernel, linux-media, devicetree, linux-kernel,
linux-rockchip
On Sat, Sep 26, 2026 at 9:16 AM Sofus Forstreuter <sofus.c@icloud.com> wrote:
>
> Hi,
>
> This series enables hardware video decoding on Apple Silicon M1, M2 and
> M3 SoCs. It adds support for all the codecs AVD (Apple Video Decoder)
> supports, which is H264, H265, VP9 and AV1 (only on M3+).
>
> The AVD is a bit unique both as a video decoder and an IP block on Apple
> Silicon SoCs. Instead of following the normal "a register for each
> parameter", AVD is programmed through one single register. This means
> the order of the writes are important, but its also important to ensure
> we dont write faster than the hardware can process. To solve this, we
> use the included Cortex-M3 co-processor. The CM3 accepts unsigned code,
> which we take advantage off [1].
> The driver prepares the instructions and writes them into segments,
> which are submitted to the CM3. The CM3 then programs the hardware and
> notifies the AP when the job is done or if an error occurred.
>
> The driver implements the V4L2 M2M stateless decoder API. The driver's
> framework is based very heavy on rkvdecs driver. It therefore comes as
> no suprise that `v4l2-compliance` reports no errors:
> Total for avd device /dev/video0: 49, Succeeded: 49, Failed: 0, Warnings: 0
>
> The fluster scores are:
> JVT-AVC_V1: 77/135
> JVT-FR-EXT: 42/69
> VP9-TEST-VECTORS: 216/305
> VP9-TEST-VECTORS-HIGH: 1/6
> JCT-VC-HEVC_V1: 143/147
> AV1-TEST-VECTORS: 238/242
>
> Since not all unsupported H264 streams are rejected, the scores
> depend on the reset working, this only works with a patch to the
> apple-dart driver so that the IOMMU attach/detach trick works.
>
> All codecs support 4:2:2, 4:2:0, in 10 and 8 bit. By default, AVD
> outputs in NV12, NV16, P010 and P210. Support for 4:4:4 and 12 bit
> formats has not been added yet.
>
> Internally the AVD uses an Apple specific compressed and tiled format
> called Interchange to store reference image data. This is stored after
> image data across all codecs. Apple Interchange is used for sharing
> buffers with other hardware blocks (GPU, display controller), support
> for this is on its way [2].
> The V4L2 pixel formats have been added to help with calculating sizes.
>
> This driver has been tested on most Apple Silicon SoCs in our downstream
> tree [3]. Its expected this driver will work on M4, M5, M6 and the Neo,
> with firmware already being written for all except the M6 SoCs [4].
>
> The groundwork for this was laid many years ago by Jamie, R and Eileen,
> who through a combined effort managed to reverse engineer many core
> parts of the AVD block.
> This driver would not have been possible without the work done by
> Eileen, who wrote initial support for H264, H265, VP9 and excellent
> tooling [5], allowing me to reverse engineer AV1 and make the driver
> conformant.
>
> [1] https://github.com/sofus13/avd-fw/tree/avd-next-fw
> [2] https://patch.msgid.link/20260917-apple-interchange-modifier-v1-1-875294689c48@gmail.com
> [3] https://github.com/AsahiLinux/linux
> [4] https://github.com/AsahiLinux/avd-fw
> [5] https://github.com/eiln/avd/tree/main
>
> Signed-off-by: Sofus Forstreuter <sofus.c@icloud.com>
> ---
> Changes in v2:
> - Patch order
> - Fixed dt-bindings by constraining all properties
> - Fixed dts, add interrupts, iommus and regs. Drop labels, fix names
> - Add Apple Interchange format
> - Add more av1 and vp9 validation to v4l2-ctrls
> - Validate more codec parameters
> - Use the new CM3 firmware (drop avd-hw, drop submission loops)
> - Updated cover to reflect changes
> - Link to v1: https://patch.msgid.link/20260918-avd-v1-0-49977931f455@icloud.com
>
> To: Sven Peter <sven@kernel.org>
> To: Janne Grunau <j@jannau.net>
> To: Neal Gompa <neal@gompa.dev>
> To: Mauro Carvalho Chehab <mchehab@kernel.org>
> To: Rob Herring <robh@kernel.org>
> To: Krzysztof Kozlowski <krzk+dt@kernel.org>
> To: Conor Dooley <conor+dt@kernel.org>
> To: Sofus Forstreuter <sofus.c@icloud.com>
> To: Philipp Zabel <p.zabel@pengutronix.de>
> To: Heiko Stuebner <heiko@sntech.de>
> Cc: asahi@lists.linux.dev
> Cc: linux-arm-kernel@lists.infradead.org
> Cc: linux-media@vger.kernel.org
> Cc: devicetree@vger.kernel.org
> Cc: linux-kernel@vger.kernel.org
> Cc: linux-rockchip@lists.infradead.org
>
> ---
> Sofus Forstreuter (17):
> dt-bindings: media: add apple,avd
> media: v4l2: Add P210 pixel format
> media: v4l2: Add Apple interchange pixel formats
> media: v4l2-ctrls: validate av1 tile info
> media: v4l2-ctrls: validate vp9 tile_rows_log2
> media: apple: add avd driver
> media: apple: avd: add h264 support
> media: apple: avd: add vp9 support
> media: apple: avd: add hevc support
> media: apple: avd: add av1 support
> arm64: dts: apple: t8103: add avd nodes
> arm64: dts: apple: t8112: add avd nodes
> arm64: dts: apple: t8122: add avd nodes
> arm64: dts: apple: t600x: add avd nodes
> arm64: dts: apple: t602x: add avd nodes
> arm64: dts: apple: t6030: add avd nodes
> arm64: dts: apple: t6031: add avd nodes
>
> .../devicetree/bindings/media/apple,avd.yaml | 127 +
> .../userspace-api/media/v4l/pixfmt-yuv-planar.rst | 14 +-
> MAINTAINERS | 2 +
> arch/arm64/boot/dts/apple/t600x-dieX.dtsi | 32 +
> arch/arm64/boot/dts/apple/t602x-dieX.dtsi | 32 +
> arch/arm64/boot/dts/apple/t6030.dtsi | 32 +
> arch/arm64/boot/dts/apple/t6031-dieX.dtsi | 32 +
> arch/arm64/boot/dts/apple/t8103.dtsi | 32 +
> arch/arm64/boot/dts/apple/t8112.dtsi | 32 +
> arch/arm64/boot/dts/apple/t8122.dtsi | 32 +
> drivers/media/platform/Kconfig | 1 +
> drivers/media/platform/Makefile | 1 +
> drivers/media/platform/apple/Kconfig | 5 +
> drivers/media/platform/apple/Makefile | 3 +
> drivers/media/platform/apple/avd/Kconfig | 18 +
> drivers/media/platform/apple/avd/Makefile | 5 +
> .../media/platform/apple/avd/avd-av1-entropymode.c | 4698 ++++++++++++++++++++
> .../media/platform/apple/avd/avd-av1-entropymode.h | 284 ++
> drivers/media/platform/apple/avd/avd-av1.c | 1432 ++++++
> drivers/media/platform/apple/avd/avd-drv.c | 776 ++++
> drivers/media/platform/apple/avd/avd-h264.c | 862 ++++
> drivers/media/platform/apple/avd/avd-hevc.c | 1340 ++++++
> drivers/media/platform/apple/avd/avd-inst.h | 210 +
> drivers/media/platform/apple/avd/avd-v4l2.c | 992 +++++
> drivers/media/platform/apple/avd/avd-vp9.c | 1054 +++++
> drivers/media/platform/apple/avd/avd.h | 299 ++
> drivers/media/v4l2-core/v4l2-common.c | 8 +
> drivers/media/v4l2-core/v4l2-ctrls-core.c | 17 +-
> drivers/media/v4l2-core/v4l2-ioctl.c | 7 +
> include/uapi/linux/v4l2-controls.h | 2 +-
> include/uapi/linux/videodev2.h | 9 +
> 31 files changed, 12386 insertions(+), 4 deletions(-)
> ---
> base-commit: 4efa625593e1e9c76346232413997556dea1d89f
> change-id: 20260917-avd-7ee5474b468c
>
Series LGTM.
Reviewed-by: Neal Gompa <neal@gompa.dev>
--
真実はいつも一つ!/ Always, there's only one truth!
^ permalink raw reply [flat|nested] 19+ messages in thread