* [PATCH v3 0/6] target/riscv: Add SiFive Xsfvqmaccdod/Xsfvqmaccqoq int8 matmul extensions
@ 2026-08-30 16:03 Max Chou
2026-08-30 16:03 ` [PATCH v3 1/6] target/riscv: Add cfg property for SiFive " Max Chou
` (5 more replies)
0 siblings, 6 replies; 9+ messages in thread
From: Max Chou @ 2026-08-30 16:03 UTC (permalink / raw)
To: qemu-devel, qemu-riscv
Cc: Palmer Dabbelt, Alistair Francis, Daniel Henrique Barboza,
Max Chou, Weiwei Li, Liu Zhiwei, Chao Liu
This patchset adds SiFive's custom int8 matrix-multiply vector
extensions Xsfvqmaccdod and Xsfvqmaccqoq. Each extension provides 4
signedness variants (unsigned/unsigned, signed/signed,
unsigned/signed, signed/unsigned) for a total of 8 instructions:
sf.vqmacc{u,,us,su}.{4x8x4,2x8x2}.
The toolchains support Xsfvqmaccdod/Xsfvqmaccqoq extensions:
- LLVM 18.1.0
- llvm/llvm-project#68295, commit d1985e3d1fed, merged 2023-10-20.
https://github.com/llvm/llvm-project/pull/68295
- GCC 15
- commit 139bd3198a73 ("RISC-V: Add the mini support for SiFive extensions"), 2024-11-17.
References:
- SiFive Int8 Matrix Multiplication Extensions Specification
https://www.sifive.com/document-file/sifive-int8-matrix-multiplication-extensions-specification
Changes since v2:<20260812174417.2468981-1-max.chou@sifive.com>
- Rebase to riscv-to-apply.next branch (commit aee18ec)
- Modified patch 5 to align the new disas flow
Changes since v1:<20260721122034.1567146-1-max.chou@sifive.com>
- Rebase to master branch (commit 84f0721)
- Reformat riscv_multi_ext_implied_rules to list one rule per
line before adding XSFVQMACCDOD_IMPLIED/XSFVQMACCQOQ_IMPLIED, so
future additions to the array are a pure append instead of
reflowing whatever line they land on. Split out into its own
patch. (Thanks for Daniel Henrique Barboza's suggestion)
- Move the SiFive int8 matmul helper out of the shared
vector_helper.c and into a new target/riscv/tcg/xsf_helper.c,
so vendor-specific code no longer lives in common vector-helper
code. (Thanks for Daniel Henrique Barboza's suggestion)
rnax
Frank Chang (2):
target/riscv/cpu: Add SiFive custom int8 matmul extension implied
rules
target/riscv: rvv: Add SiFive custom int8 matmul instructions
Jason Chien (1):
disas/riscv: Add disassembler support for Xsfvqmaccdod/Xsfvqmaccqoq
Max Chou (3):
target/riscv: Add cfg property for SiFive int8 matmul extensions
target/riscv/cpu: List one implied-ext rule per line
tests/tcg/riscv64: Add tests for SiFive int8 matmul extensions
MAINTAINERS | 10 +
disas/meson.build | 3 +-
disas/riscv-xsf-op.c.inc | 11 +
disas/riscv-xsf.c | 49 ++++
disas/riscv-xsf.h | 16 ++
disas/riscv.c | 2 +
disas/riscv.h | 1 +
target/riscv/cpu.c | 80 ++++--
target/riscv/cpu_cfg.h | 5 +
target/riscv/cpu_cfg_fields.h.inc | 2 +
target/riscv/helper.h | 10 +
target/riscv/meson.build | 1 +
target/riscv/tcg/insn_trans/trans_xsf.c.inc | 98 ++++++++
target/riscv/tcg/meson.build | 1 +
target/riscv/tcg/tcg-cpu.c | 16 ++
target/riscv/tcg/translate.c | 3 +
target/riscv/tcg/xsf_helper.c | 87 +++++++
target/riscv/xsf.decode | 30 +++
tests/tcg/riscv64/Makefile.softmmu-target | 7 +
tests/tcg/riscv64/test-xsfvqmacc.S | 260 ++++++++++++++++++++
20 files changed, 677 insertions(+), 15 deletions(-)
create mode 100644 disas/riscv-xsf-op.c.inc
create mode 100644 disas/riscv-xsf.c
create mode 100644 disas/riscv-xsf.h
create mode 100644 target/riscv/tcg/insn_trans/trans_xsf.c.inc
create mode 100644 target/riscv/tcg/xsf_helper.c
create mode 100644 target/riscv/xsf.decode
create mode 100644 tests/tcg/riscv64/test-xsfvqmacc.S
--
2.43.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH v3 1/6] target/riscv: Add cfg property for SiFive int8 matmul extensions
2026-08-30 16:03 [PATCH v3 0/6] target/riscv: Add SiFive Xsfvqmaccdod/Xsfvqmaccqoq int8 matmul extensions Max Chou
@ 2026-08-30 16:03 ` Max Chou
2026-08-30 16:03 ` [PATCH v3 2/6] target/riscv/cpu: List one implied-ext rule per line Max Chou
` (4 subsequent siblings)
5 siblings, 0 replies; 9+ messages in thread
From: Max Chou @ 2026-08-30 16:03 UTC (permalink / raw)
To: qemu-devel, qemu-riscv
Cc: Palmer Dabbelt, Alistair Francis, Daniel Henrique Barboza,
Max Chou, Weiwei Li, Liu Zhiwei, Chao Liu, Alistair Francis
Add the CPU config properties for SiFive's custom int8 matrix
multiply vector extensions:
- Xsfvqmaccdod: 2x8x2 int8 matrix-multiply-accumulate operations
- Xsfvqmaccqoq: 4x8x4 int8 matrix-multiply-accumulate operations
All instructions of these extensions require the Zve32x extension
to be present whenever either extension is enabled. Validate this
in riscv_cpu_validate_vendor_ext() alongside the existing vector
extension validation in riscv_cpu_validate_set_extensions().
Reviewed-by: Daniel Henrique Barboza <daniel.barboza@oss.qualcomm.com>
Reviewed-by: Alistair Francis <alistair.francis@wdc.com>
Signed-off-by: Max Chou <max.chou@sifive.com>
---
target/riscv/cpu.c | 2 ++
target/riscv/cpu_cfg_fields.h.inc | 2 ++
target/riscv/tcg/tcg-cpu.c | 16 ++++++++++++++++
3 files changed, 20 insertions(+)
diff --git a/target/riscv/cpu.c b/target/riscv/cpu.c
index 61109672d60..b4e34daf8ec 100644
--- a/target/riscv/cpu.c
+++ b/target/riscv/cpu.c
@@ -299,6 +299,8 @@ const RISCVIsaExtData isa_edata_arr[] = {
ISA_EXT_DATA_ENTRY(xmipscbop, PRIV_VERSION_1_12_0, ext_xmipscbop),
ISA_EXT_DATA_ENTRY(xmipscmov, PRIV_VERSION_1_12_0, ext_xmipscmov),
ISA_EXT_DATA_ENTRY(xmipslsp, PRIV_VERSION_1_12_0, ext_xmipslsp),
+ ISA_EXT_DATA_ENTRY(xsfvqmaccdod, PRIV_VERSION_1_13_0, ext_xsfvqmaccdod),
+ ISA_EXT_DATA_ENTRY(xsfvqmaccqoq, PRIV_VERSION_1_13_0, ext_xsfvqmaccqoq),
ISA_EXT_DATA_ENTRY(xtheadba, PRIV_VERSION_1_11_0, ext_xtheadba),
ISA_EXT_DATA_ENTRY(xtheadbb, PRIV_VERSION_1_11_0, ext_xtheadbb),
ISA_EXT_DATA_ENTRY(xtheadbs, PRIV_VERSION_1_11_0, ext_xtheadbs),
diff --git a/target/riscv/cpu_cfg_fields.h.inc b/target/riscv/cpu_cfg_fields.h.inc
index f91c780299c..0f3bda1331c 100644
--- a/target/riscv/cpu_cfg_fields.h.inc
+++ b/target/riscv/cpu_cfg_fields.h.inc
@@ -159,6 +159,8 @@ BOOL_FIELD(ext_xmipscbop)
BOOL_FIELD(ext_xmipscmov)
BOOL_FIELD(ext_xmipslsp)
BOOL_FIELD(ext_xlrbr)
+BOOL_FIELD(ext_xsfvqmaccdod)
+BOOL_FIELD(ext_xsfvqmaccqoq)
BOOL_FIELD(big_endian)
BOOL_FIELD(mmu)
diff --git a/target/riscv/tcg/tcg-cpu.c b/target/riscv/tcg/tcg-cpu.c
index 54f1b66216c..bdb6dc6ab3f 100644
--- a/target/riscv/tcg/tcg-cpu.c
+++ b/target/riscv/tcg/tcg-cpu.c
@@ -418,6 +418,16 @@ static void riscv_cpu_validate_v(CPURISCVState *env, RISCVCPUConfig *cfg,
}
}
+static void riscv_cpu_validate_vendor_ext(RISCVCPU *cpu, Error **errp)
+{
+ if ((cpu->cfg.ext_xsfvqmaccdod || cpu->cfg.ext_xsfvqmaccqoq) &&
+ !cpu->cfg.ext_zve32x) {
+ error_setg(errp, "Xsfvqmaccdod/Xsfvqmaccqoq extensions require "
+ "Zve32x extension");
+ return;
+ }
+}
+
static void riscv_cpu_disable_priv_spec_isa_exts(RISCVCPU *cpu)
{
CPURISCVState *env = &cpu->env;
@@ -795,6 +805,12 @@ void riscv_cpu_validate_set_extensions(RISCVCPU *cpu, Error **errp)
return;
}
+ riscv_cpu_validate_vendor_ext(cpu, &local_err);
+ if (local_err != NULL) {
+ error_propagate(errp, local_err);
+ return;
+ }
+
if (mcc->def->misa_mxl_max == MXL_RV32 && cpu->cfg.ext_svukte) {
error_setg(errp, "svukte is not supported for RV32");
return;
--
2.43.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v3 2/6] target/riscv/cpu: List one implied-ext rule per line
2026-08-30 16:03 [PATCH v3 0/6] target/riscv: Add SiFive Xsfvqmaccdod/Xsfvqmaccqoq int8 matmul extensions Max Chou
2026-08-30 16:03 ` [PATCH v3 1/6] target/riscv: Add cfg property for SiFive " Max Chou
@ 2026-08-30 16:03 ` Max Chou
2026-08-30 16:03 ` [PATCH v3 3/6] target/riscv/cpu: Add SiFive custom int8 matmul extension implied rules Max Chou
` (3 subsequent siblings)
5 siblings, 0 replies; 9+ messages in thread
From: Max Chou @ 2026-08-30 16:03 UTC (permalink / raw)
To: qemu-devel, qemu-riscv
Cc: Palmer Dabbelt, Alistair Francis, Daniel Henrique Barboza,
Max Chou, Weiwei Li, Liu Zhiwei, Chao Liu, Alistair Francis
Every entry added to riscv_multi_ext_implied_rules[] currently has to
reflow whatever line it lands on, since entries are packed several
per line. List one rule per line instead so future additions are a
pure append.
Suggested-by: Daniel Henrique Barboza <daniel.barboza@oss.qualcomm.com>
Reviewed-by: Daniel Henrique Barboza <daniel.barboza@oss.qualcomm.com>
Reviewed-by: Alistair Francis <alistair.francis@wdc.com>
Signed-off-by: Max Chou <max.chou@sifive.com>
---
target/riscv/cpu.c | 58 +++++++++++++++++++++++++++++++++++-----------
1 file changed, 44 insertions(+), 14 deletions(-)
diff --git a/target/riscv/cpu.c b/target/riscv/cpu.c
index b4e34daf8ec..af61a1ab48d 100644
--- a/target/riscv/cpu.c
+++ b/target/riscv/cpu.c
@@ -2909,20 +2909,50 @@ RISCVCPUImpliedExtsRule *riscv_misa_ext_implied_rules[] = {
};
RISCVCPUImpliedExtsRule *riscv_multi_ext_implied_rules[] = {
- &ZCB_IMPLIED, &ZCD_IMPLIED, &ZCE_IMPLIED,
- &ZCF_IMPLIED, &ZCMP_IMPLIED, &ZCMT_IMPLIED,
- &ZDINX_IMPLIED, &ZFA_IMPLIED, &ZFBFMIN_IMPLIED,
- &ZFH_IMPLIED, &ZFHMIN_IMPLIED, &ZFINX_IMPLIED,
- &ZHINX_IMPLIED, &ZHINXMIN_IMPLIED, &ZICNTR_IMPLIED,
- &ZIHPM_IMPLIED, &ZK_IMPLIED, &ZKN_IMPLIED,
- &ZKS_IMPLIED, &ZVBB_IMPLIED, &ZVE32F_IMPLIED,
- &ZVE32X_IMPLIED, &ZVE64D_IMPLIED, &ZVE64F_IMPLIED, &ZVE64X_IMPLIED,
- &ZVFBFA_IMPLIED, &ZVFBFMIN_IMPLIED, &ZVFBFWMA_IMPLIED,
- &ZVFH_IMPLIED, &ZVFHMIN_IMPLIED, &ZVKN_IMPLIED,
- &ZVKNC_IMPLIED, &ZVKNG_IMPLIED, &ZVKNHB_IMPLIED,
- &ZVKS_IMPLIED, &ZVKSC_IMPLIED, &ZVKSG_IMPLIED, &SHA_IMPLIED,
- &SSCFG_IMPLIED, &SUPM_IMPLIED, &SSPM_IMPLIED, &SMCTR_IMPLIED,
- &SSCTR_IMPLIED, &SSSTATEEN_IMPLIED,
+ &ZCB_IMPLIED,
+ &ZCD_IMPLIED,
+ &ZCE_IMPLIED,
+ &ZCF_IMPLIED,
+ &ZCMP_IMPLIED,
+ &ZCMT_IMPLIED,
+ &ZDINX_IMPLIED,
+ &ZFA_IMPLIED,
+ &ZFBFMIN_IMPLIED,
+ &ZFH_IMPLIED,
+ &ZFHMIN_IMPLIED,
+ &ZFINX_IMPLIED,
+ &ZHINX_IMPLIED,
+ &ZHINXMIN_IMPLIED,
+ &ZICNTR_IMPLIED,
+ &ZIHPM_IMPLIED,
+ &ZK_IMPLIED,
+ &ZKN_IMPLIED,
+ &ZKS_IMPLIED,
+ &ZVBB_IMPLIED,
+ &ZVE32F_IMPLIED,
+ &ZVE32X_IMPLIED,
+ &ZVE64D_IMPLIED,
+ &ZVE64F_IMPLIED,
+ &ZVE64X_IMPLIED,
+ &ZVFBFA_IMPLIED,
+ &ZVFBFMIN_IMPLIED,
+ &ZVFBFWMA_IMPLIED,
+ &ZVFH_IMPLIED,
+ &ZVFHMIN_IMPLIED,
+ &ZVKN_IMPLIED,
+ &ZVKNC_IMPLIED,
+ &ZVKNG_IMPLIED,
+ &ZVKNHB_IMPLIED,
+ &ZVKS_IMPLIED,
+ &ZVKSC_IMPLIED,
+ &ZVKSG_IMPLIED,
+ &SHA_IMPLIED,
+ &SSCFG_IMPLIED,
+ &SUPM_IMPLIED,
+ &SSPM_IMPLIED,
+ &SMCTR_IMPLIED,
+ &SSCTR_IMPLIED,
+ &SSSTATEEN_IMPLIED,
NULL
};
--
2.43.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v3 3/6] target/riscv/cpu: Add SiFive custom int8 matmul extension implied rules
2026-08-30 16:03 [PATCH v3 0/6] target/riscv: Add SiFive Xsfvqmaccdod/Xsfvqmaccqoq int8 matmul extensions Max Chou
2026-08-30 16:03 ` [PATCH v3 1/6] target/riscv: Add cfg property for SiFive " Max Chou
2026-08-30 16:03 ` [PATCH v3 2/6] target/riscv/cpu: List one implied-ext rule per line Max Chou
@ 2026-08-30 16:03 ` Max Chou
2026-08-30 16:03 ` [PATCH v3 4/6] target/riscv: rvv: Add SiFive custom int8 matmul instructions Max Chou
` (2 subsequent siblings)
5 siblings, 0 replies; 9+ messages in thread
From: Max Chou @ 2026-08-30 16:03 UTC (permalink / raw)
To: qemu-devel, qemu-riscv
Cc: Palmer Dabbelt, Alistair Francis, Daniel Henrique Barboza,
Max Chou, Weiwei Li, Liu Zhiwei, Chao Liu, Frank Chang,
Alistair Francis
From: Frank Chang <frank.chang@sifive.com>
Add SiFive custom int8 matmul extension implied rules to enable the
implied extensions of SiFive custom int8 matmul extension recursively.
Signed-off-by: Frank Chang <frank.chang@sifive.com>
Reviewed-by: Daniel Henrique Barboza <daniel.barboza@oss.qualcomm.com>
Reviewed-by: Alistair Francis <alistair.francis@wdc.com>
Signed-off-by: Max Chou <max.chou@sifive.com>
---
target/riscv/cpu.c | 20 ++++++++++++++++++++
1 file changed, 20 insertions(+)
diff --git a/target/riscv/cpu.c b/target/riscv/cpu.c
index af61a1ab48d..4263b7de241 100644
--- a/target/riscv/cpu.c
+++ b/target/riscv/cpu.c
@@ -2902,6 +2902,24 @@ static RISCVCPUImpliedExtsRule ZVFBFA_IMPLIED = {
},
};
+static RISCVCPUImpliedExtsRule XSFVQMACCDOD_IMPLIED = {
+ .ext = CPU_CFG_OFFSET(ext_xsfvqmaccdod),
+ .implied_multi_exts = {
+ CPU_CFG_OFFSET(ext_zve32x),
+
+ RISCV_IMPLIED_EXTS_RULE_END
+ },
+};
+
+static RISCVCPUImpliedExtsRule XSFVQMACCQOQ_IMPLIED = {
+ .ext = CPU_CFG_OFFSET(ext_xsfvqmaccqoq),
+ .implied_multi_exts = {
+ CPU_CFG_OFFSET(ext_zve32x),
+
+ RISCV_IMPLIED_EXTS_RULE_END
+ },
+};
+
RISCVCPUImpliedExtsRule *riscv_misa_ext_implied_rules[] = {
&RVA_IMPLIED, &RVD_IMPLIED, &RVF_IMPLIED,
&RVM_IMPLIED, &RVV_IMPLIED, &RVG_IMPLIED,
@@ -2953,6 +2971,8 @@ RISCVCPUImpliedExtsRule *riscv_multi_ext_implied_rules[] = {
&SMCTR_IMPLIED,
&SSCTR_IMPLIED,
&SSSTATEEN_IMPLIED,
+ &XSFVQMACCDOD_IMPLIED,
+ &XSFVQMACCQOQ_IMPLIED,
NULL
};
--
2.43.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v3 4/6] target/riscv: rvv: Add SiFive custom int8 matmul instructions
2026-08-30 16:03 [PATCH v3 0/6] target/riscv: Add SiFive Xsfvqmaccdod/Xsfvqmaccqoq int8 matmul extensions Max Chou
` (2 preceding siblings ...)
2026-08-30 16:03 ` [PATCH v3 3/6] target/riscv/cpu: Add SiFive custom int8 matmul extension implied rules Max Chou
@ 2026-08-30 16:03 ` Max Chou
2026-08-30 16:03 ` [PATCH v3 5/6] disas/riscv: Add disassembler support for Xsfvqmaccdod/Xsfvqmaccqoq Max Chou
2026-08-30 16:03 ` [PATCH v3 6/6] tests/tcg/riscv64: Add tests for SiFive int8 matmul extensions Max Chou
5 siblings, 0 replies; 9+ messages in thread
From: Max Chou @ 2026-08-30 16:03 UTC (permalink / raw)
To: qemu-devel, qemu-riscv
Cc: Palmer Dabbelt, Alistair Francis, Daniel Henrique Barboza,
Max Chou, Weiwei Li, Liu Zhiwei, Chao Liu, Frank Chang,
Alistair Francis
From: Frank Chang <frank.chang@sifive.com>
Add the 8 SiFive custom int8 matrix-multiply vector instructions:
sf.vqmacc{u,,us,su}.4x8x4 and sf.vqmacc{u,,us,su}.2x8x2. Each name
suffix encodes the signedness of vs1/vs2.
The 4x8x4 forms multiply-accumulate a 4x8 by 8x4 int8 tile into a
4x4 int32 result; the 2x8x2 forms use a 2x8 by 8x2 tile producing a
2x2 int32 result. Both Xsfvqmaccqoq/Xsfvqmaccdod extensions are
gated on vlenb >= 32, sew == 8 and vm == 1, per the SiFive Int8
Matrix Multiplication Extensions Specification.
Signed-off-by: Frank Chang <frank.chang@sifive.com>
Reviewed-by: Daniel Henrique Barboza <daniel.barboza@oss.qualcomm.com>
Acked-by: Alistair Francis <alistair.francis@wdc.com>
Signed-off-by: Max Chou <max.chou@sifive.com>
---
MAINTAINERS | 8 ++
target/riscv/cpu_cfg.h | 5 ++
target/riscv/helper.h | 10 +++
target/riscv/meson.build | 1 +
target/riscv/tcg/insn_trans/trans_xsf.c.inc | 98 +++++++++++++++++++++
target/riscv/tcg/meson.build | 1 +
target/riscv/tcg/translate.c | 3 +
target/riscv/tcg/xsf_helper.c | 87 ++++++++++++++++++
target/riscv/xsf.decode | 30 +++++++
9 files changed, 243 insertions(+)
create mode 100644 target/riscv/tcg/insn_trans/trans_xsf.c.inc
create mode 100644 target/riscv/tcg/xsf_helper.c
create mode 100644 target/riscv/xsf.decode
diff --git a/MAINTAINERS b/MAINTAINERS
index 2a0ee348c4e..2e8dace19ff 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -398,6 +398,14 @@ F: target/riscv/XVentanaCondOps.decode
F: target/riscv/insn_trans/trans_xventanacondops.c.inc
F: disas/riscv-xventana*
+RISC-V SiFive (Xsf*) extensions
+M: Max Chou <max.chou@sifive.com>
+L: qemu-riscv@nongnu.org
+S: Supported
+F: target/riscv/xsf.decode
+F: target/riscv/tcg/insn_trans/trans_xsf.c.inc
+F: target/riscv/tcg/xsf_helper.c
+
RENESAS RX CPUs
R: Yoshinori Sato <yoshinori.sato@nifty.com>
S: Orphan
diff --git a/target/riscv/cpu_cfg.h b/target/riscv/cpu_cfg.h
index 211d0708ba4..d6db1cfb7cc 100644
--- a/target/riscv/cpu_cfg.h
+++ b/target/riscv/cpu_cfg.h
@@ -51,6 +51,11 @@ static inline bool has_xthead_p(const RISCVCPUConfig *cfg)
cfg->ext_xtheadmempair || cfg->ext_xtheadsync;
}
+static inline bool has_xsf_p(const RISCVCPUConfig *cfg)
+{
+ return cfg->ext_xsfvqmaccdod || cfg->ext_xsfvqmaccqoq;
+}
+
#define MATERIALISE_EXT_PREDICATE(ext) \
static inline bool has_ ## ext ## _p(const RISCVCPUConfig *cfg) \
{ \
diff --git a/target/riscv/helper.h b/target/riscv/helper.h
index 4fc2d3a1559..8033a353328 100644
--- a/target/riscv/helper.h
+++ b/target/riscv/helper.h
@@ -1357,3 +1357,13 @@ DEF_HELPER_1(ssamoswap_disabled, void, env)
/* Zalrsc SC write probe */
DEF_HELPER_FLAGS_3(sc_probe_write, TCG_CALL_NO_WG, void, env, tl, tl)
+
+/* SiFive Custom int8 Matrix-Multiply */
+DEF_HELPER_5(sf_vqmaccu_4x8x4, void, ptr, ptr, ptr, env, i32)
+DEF_HELPER_5(sf_vqmacc_4x8x4, void, ptr, ptr, ptr, env, i32)
+DEF_HELPER_5(sf_vqmaccus_4x8x4, void, ptr, ptr, ptr, env, i32)
+DEF_HELPER_5(sf_vqmaccsu_4x8x4, void, ptr, ptr, ptr, env, i32)
+DEF_HELPER_5(sf_vqmaccu_2x8x2, void, ptr, ptr, ptr, env, i32)
+DEF_HELPER_5(sf_vqmacc_2x8x2, void, ptr, ptr, ptr, env, i32)
+DEF_HELPER_5(sf_vqmaccus_2x8x2, void, ptr, ptr, ptr, env, i32)
+DEF_HELPER_5(sf_vqmaccsu_2x8x2, void, ptr, ptr, ptr, env, i32)
diff --git a/target/riscv/meson.build b/target/riscv/meson.build
index 42d0f6d538a..c06526adb26 100644
--- a/target/riscv/meson.build
+++ b/target/riscv/meson.build
@@ -6,6 +6,7 @@ gen = [
decodetree.process('XVentanaCondOps.decode', extra_args: '--static-decode=decode_XVentanaCodeOps'),
decodetree.process('xmips.decode', extra_args: '--static-decode=decode_xmips'),
decodetree.process('xlrbr.decode', extra_args: '--static-decode=decode_xlrbr'),
+ decodetree.process('xsf.decode', extra_args: '--static-decode=decode_xsf'),
]
riscv_ss = ss.source_set()
diff --git a/target/riscv/tcg/insn_trans/trans_xsf.c.inc b/target/riscv/tcg/insn_trans/trans_xsf.c.inc
new file mode 100644
index 00000000000..1677352689f
--- /dev/null
+++ b/target/riscv/tcg/insn_trans/trans_xsf.c.inc
@@ -0,0 +1,98 @@
+/*
+ * RISC-V translation routines for the SiFive vendor extensions (xsf*)
+ *
+ * Copyright (c) 2023 SiFive, Inc.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+
+/*
+ * SiFive Xsfvqmaccdod/Xsfvqmaccqoq custom int8 matrix-multiply extensions
+ */
+static bool sf_int8_matmul_check(DisasContext *s, arg_rmrr *a)
+{
+ return require_rvv(s) &&
+ vext_check_isa_ill(s) &&
+ s->vstart_eq_zero &&
+ (s->cfg_ptr->vlenb >= 32) &&
+ (s->sew == MO_8) &&
+ (a->vm == 1);
+}
+
+static bool sf_int8_matmul_4x8x4_check(DisasContext *s, arg_rmrr *a)
+{
+ /*
+ * vd has EMUL=2*LMUL
+ * vs2 has EMUL=LMUL
+ * vs1 has EMUL=1
+ * vd must not overlap vs1
+ */
+ return sf_int8_matmul_check(s, a) &&
+ (s->cfg_ptr->ext_xsfvqmaccqoq) &&
+ (s->lmul <= 2) &&
+ require_align(a->rd, s->lmul + 1) &&
+ require_align(a->rs2, s->lmul) &&
+ require_align(a->rs1, 0) &&
+ require_noover(a->rd, s->lmul + 1, a->rs2, s->lmul) &&
+ !is_overlapped(a->rd, 1 << MAX(s->lmul + 1, 0), a->rs1, 1);
+}
+
+static bool sf_int8_matmul_2x8x2_check(DisasContext *s, arg_rmrr *a)
+{
+ /*
+ * vd has EMUL=LMUL
+ * vs2 has EMUL=LMUL
+ * vs1 has EMUL=1
+ * vd must not overlap vs1
+ */
+ return sf_int8_matmul_check(s, a) &&
+ (s->cfg_ptr->ext_xsfvqmaccdod) &&
+ require_align(a->rd, s->lmul) &&
+ require_align(a->rs2, s->lmul) &&
+ require_align(a->rs1, 0) &&
+ !is_overlapped(a->rd, 1 << MAX(s->lmul, 0), a->rs1, 1);
+}
+
+static bool sf_int8_matmul_op(DisasContext *s, arg_rmrr *a, uint8_t seq)
+{
+ static gen_helper_gvec_3_ptr * const fns[8] = {
+ gen_helper_sf_vqmaccu_4x8x4, gen_helper_sf_vqmacc_4x8x4,
+ gen_helper_sf_vqmaccus_4x8x4, gen_helper_sf_vqmaccsu_4x8x4,
+ gen_helper_sf_vqmaccu_2x8x2, gen_helper_sf_vqmacc_2x8x2,
+ gen_helper_sf_vqmaccus_2x8x2, gen_helper_sf_vqmaccsu_2x8x2,
+ };
+
+ /*
+ * The helper raises an illegal-instruction exception when vl is not a
+ * multiple of the tile size; save the opcode so mtval/stval report the
+ * faulting instruction if that exception is thrown.
+ */
+ decode_save_opc(s, 0);
+
+ tcg_gen_gvec_3_ptr(vreg_ofs(s, a->rd), vreg_ofs(s, a->rs1),
+ vreg_ofs(s, a->rs2), tcg_env,
+ s->cfg_ptr->vlenb, s->cfg_ptr->vlenb, 0, fns[seq]);
+
+ finalize_rvv_inst(s);
+
+ return true;
+}
+
+#define GEN_SF_INT8_MATMUL_TRANS(NAME, CHECK, SEQ) \
+static bool trans_##NAME(DisasContext *s, arg_rmrr *a) \
+{ \
+ if (CHECK(s, a)) { \
+ return sf_int8_matmul_op(s, a, SEQ); \
+ } \
+ return false; \
+}
+
+GEN_SF_INT8_MATMUL_TRANS(sf_vqmaccu_4x8x4, sf_int8_matmul_4x8x4_check, 0)
+GEN_SF_INT8_MATMUL_TRANS(sf_vqmacc_4x8x4, sf_int8_matmul_4x8x4_check, 1)
+GEN_SF_INT8_MATMUL_TRANS(sf_vqmaccus_4x8x4, sf_int8_matmul_4x8x4_check, 2)
+GEN_SF_INT8_MATMUL_TRANS(sf_vqmaccsu_4x8x4, sf_int8_matmul_4x8x4_check, 3)
+GEN_SF_INT8_MATMUL_TRANS(sf_vqmaccu_2x8x2, sf_int8_matmul_2x8x2_check, 4)
+GEN_SF_INT8_MATMUL_TRANS(sf_vqmacc_2x8x2, sf_int8_matmul_2x8x2_check, 5)
+GEN_SF_INT8_MATMUL_TRANS(sf_vqmaccus_2x8x2, sf_int8_matmul_2x8x2_check, 6)
+GEN_SF_INT8_MATMUL_TRANS(sf_vqmaccsu_2x8x2, sf_int8_matmul_2x8x2_check, 7)
diff --git a/target/riscv/tcg/meson.build b/target/riscv/tcg/meson.build
index a05ab642f41..86805b30cbc 100644
--- a/target/riscv/tcg/meson.build
+++ b/target/riscv/tcg/meson.build
@@ -15,6 +15,7 @@ riscv_ss.add(files(
'vcrypto_helper.c',
'vector_helper.c',
'vector_internals.c',
+ 'xsf_helper.c',
'zce_helper.c'))
diff --git a/target/riscv/tcg/translate.c b/target/riscv/tcg/translate.c
index 9684dbe7528..41e3dd2fe21 100644
--- a/target/riscv/tcg/translate.c
+++ b/target/riscv/tcg/translate.c
@@ -1216,10 +1216,12 @@ static uint32_t opcode_at(DisasContextBase *dcbase, target_ulong pc)
#include "decode-xthead.c.inc"
#include "decode-xmips.c.inc"
#include "decode-xlrbr.c.inc"
+#include "decode-xsf.c.inc"
#include "insn_trans/trans_xthead.c.inc"
#include "insn_trans/trans_xventanacondops.c.inc"
#include "insn_trans/trans_xmips.c.inc"
#include "insn_trans/trans_xlrbr.c.inc"
+#include "insn_trans/trans_xsf.c.inc"
/* Include the auto-generated decoder for 16 bit insn */
#include "decode-insn16.c.inc"
@@ -1240,6 +1242,7 @@ const RISCVDecoder decoder_table[] = {
{ has_xthead_p, decode_xthead},
{ has_XVentanaCondOps_p, decode_XVentanaCodeOps},
{ has_xlrbr_p, decode_xlrbr},
+ { has_xsf_p, decode_xsf },
};
const size_t decoder_table_size = ARRAY_SIZE(decoder_table);
diff --git a/target/riscv/tcg/xsf_helper.c b/target/riscv/tcg/xsf_helper.c
new file mode 100644
index 00000000000..6e953d12493
--- /dev/null
+++ b/target/riscv/tcg/xsf_helper.c
@@ -0,0 +1,87 @@
+/*
+ * RISC-V translation helpers for the SiFive vendor extensions (xsf*)
+ *
+ * Copyright (c) 2023 SiFive, Inc.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#include "qemu/osdep.h"
+#include "cpu.h"
+#include "exec/helper-proto.h"
+#include "internals.h"
+#include "vector_internals.h"
+
+#define QOP_SUU_B int32_t, uint8_t, uint8_t, int32_t, int32_t
+#define QOP_SUS_B int32_t, uint8_t, int8_t, int32_t, int32_t
+#define QOP_SSU_B int32_t, int8_t, uint8_t, int32_t, int32_t
+#define QOP_SSS_B int32_t, int8_t, int8_t, int32_t, int32_t
+
+/* SiFive Custom int8 Matrix-Multiply */
+/*
+ * vd may overlap vs2, we need to allocate an additional vd array
+ * to save temporary results of vd and write them back at the end.
+ */
+#define GEN_VEXT_SF_INT8_MATMUL(NAME, TD, T1, T2, TX1, TX2, \
+ HD, HS1, HS2, ROWS, COLS, TILE_SIZE) \
+void HELPER(NAME)(void *vd, void *vs1, void *vs2, \
+ CPURISCVState *env, uint32_t desc) \
+{ \
+ int it, il, in, im, ivd, ivs1, ivs2; \
+ TD *vds; \
+ \
+ if (env->vl % TILE_SIZE) { \
+ riscv_raise_exception(env, RISCV_EXCP_ILLEGAL_INST, GETPC()); \
+ return; \
+ } \
+ \
+ VSTART_CHECK_EARLY_EXIT(env, env->vl); \
+ \
+ vds = g_malloc0(sizeof(TD) * \
+ ROWS * ROWS * (env->vl / TILE_SIZE)); \
+ \
+ for (it = 0; it < (env->vl / TILE_SIZE); it++) { \
+ for (il = 0; il < ROWS; il++) { \
+ for (in = 0; in < ROWS; in++) { \
+ ivd = ROWS * ROWS * it + ROWS * il + in; \
+ vds[ivd] = *((TD *)vd + HD(ivd)); \
+ for (im = 0; im < COLS; im++) { \
+ ivs1 = il * COLS + im; \
+ ivs2 = TILE_SIZE * it + im * ROWS + in; \
+ T1 s1 = *((T1 *)vs1 + HS1(ivs1)); \
+ T2 s2 = *((T2 *)vs2 + HS2(ivs2)); \
+ vds[ivd] += (TX1)s1 * (TX2)s2; \
+ } \
+ } \
+ } \
+ } \
+ \
+ for (it = 0; it < (env->vl / TILE_SIZE); it++) { \
+ for (il = 0; il < ROWS; il++) { \
+ for (in = 0; in < ROWS; in++) { \
+ ivd = ROWS * ROWS * it + ROWS * il + in; \
+ *((TD *)vd + HD(ivd)) = vds[ivd]; \
+ } \
+ } \
+ } \
+ \
+ env->vstart = 0; \
+ g_free(vds); \
+}
+
+RVVCALL(GEN_VEXT_SF_INT8_MATMUL, sf_vqmaccu_4x8x4, QOP_SUU_B, H4, H1, H1,
+ 4, 8, 32)
+RVVCALL(GEN_VEXT_SF_INT8_MATMUL, sf_vqmacc_4x8x4, QOP_SSS_B, H4, H1, H1,
+ 4, 8, 32)
+RVVCALL(GEN_VEXT_SF_INT8_MATMUL, sf_vqmaccus_4x8x4, QOP_SUS_B, H4, H1, H1,
+ 4, 8, 32)
+RVVCALL(GEN_VEXT_SF_INT8_MATMUL, sf_vqmaccsu_4x8x4, QOP_SSU_B, H4, H1, H1,
+ 4, 8, 32)
+RVVCALL(GEN_VEXT_SF_INT8_MATMUL, sf_vqmaccu_2x8x2, QOP_SUU_B, H4, H1, H1,
+ 2, 8, 16)
+RVVCALL(GEN_VEXT_SF_INT8_MATMUL, sf_vqmacc_2x8x2, QOP_SSS_B, H4, H1, H1,
+ 2, 8, 16)
+RVVCALL(GEN_VEXT_SF_INT8_MATMUL, sf_vqmaccus_2x8x2, QOP_SUS_B, H4, H1, H1,
+ 2, 8, 16)
+RVVCALL(GEN_VEXT_SF_INT8_MATMUL, sf_vqmaccsu_2x8x2, QOP_SSU_B, H4, H1, H1,
+ 2, 8, 16)
diff --git a/target/riscv/xsf.decode b/target/riscv/xsf.decode
new file mode 100644
index 00000000000..bb585046ab1
--- /dev/null
+++ b/target/riscv/xsf.decode
@@ -0,0 +1,30 @@
+#
+# RISC-V translation routines for the SiFive vendor extensions
+#
+# Copyright (c) 2023 SiFive, Inc.
+#
+# SPDX-License-Identifier: GPL-2.0-or-later
+
+# Fields:
+%rs2 20:5
+%rs1 15:5
+%rd 7:5
+%vm 25:1
+
+# Argument sets:
+&rmrr vm rd rs1 rs2 !extern
+
+# Formats:
+@r_vm_1 ...... . ..... ..... ... ..... ....... &rmrr vm=1 %rs2 %rs1 %rd
+
+# *** Xsfvqmaccqoq: SiFive custom int8 matrix-multiply (4x8x4 tile) ***
+sf_vqmaccu_4x8x4 111100 1 ..... ..... 010 ..... 1011011 @r_vm_1
+sf_vqmacc_4x8x4 111101 1 ..... ..... 010 ..... 1011011 @r_vm_1
+sf_vqmaccus_4x8x4 111110 1 ..... ..... 010 ..... 1011011 @r_vm_1
+sf_vqmaccsu_4x8x4 111111 1 ..... ..... 010 ..... 1011011 @r_vm_1
+
+# *** Xsfvqmaccdod: SiFive custom int8 matrix-multiply (2x8x2 tile) ***
+sf_vqmaccu_2x8x2 101100 1 ..... ..... 010 ..... 1011011 @r_vm_1
+sf_vqmacc_2x8x2 101101 1 ..... ..... 010 ..... 1011011 @r_vm_1
+sf_vqmaccus_2x8x2 101110 1 ..... ..... 010 ..... 1011011 @r_vm_1
+sf_vqmaccsu_2x8x2 101111 1 ..... ..... 010 ..... 1011011 @r_vm_1
--
2.43.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v3 5/6] disas/riscv: Add disassembler support for Xsfvqmaccdod/Xsfvqmaccqoq
2026-08-30 16:03 [PATCH v3 0/6] target/riscv: Add SiFive Xsfvqmaccdod/Xsfvqmaccqoq int8 matmul extensions Max Chou
` (3 preceding siblings ...)
2026-08-30 16:03 ` [PATCH v3 4/6] target/riscv: rvv: Add SiFive custom int8 matmul instructions Max Chou
@ 2026-08-30 16:03 ` Max Chou
2026-09-07 4:12 ` Alistair Francis
2026-08-30 16:03 ` [PATCH v3 6/6] tests/tcg/riscv64: Add tests for SiFive int8 matmul extensions Max Chou
5 siblings, 1 reply; 9+ messages in thread
From: Max Chou @ 2026-08-30 16:03 UTC (permalink / raw)
To: qemu-devel, qemu-riscv
Cc: Palmer Dabbelt, Alistair Francis, Daniel Henrique Barboza,
Max Chou, Weiwei Li, Liu Zhiwei, Chao Liu, Jason Chien
From: Jason Chien <jason.chien@sifive.com>
Add disassembler support for all SiFive custom int8 matrix-multiply
instructions: sf.vqmacc{u,,us,su}.4x8x4 and sf.vqmacc{u,,us,su}.2x8x2.
These instructions share opcode 0b1011011 and funct3 0b010, and are
distinguished by funct6: 60-63 for the 4x8x4 (Xsfvqmaccqoq) tile and
44-47 for the 2x8x2 (Xsfvqmaccdod) tile.
Signed-off-by: Jason Chien <jason.chien@sifive.com>
Reviewed-by: Daniel Henrique Barboza <daniel.barboza@oss.qualcomm.com>
Signed-off-by: Max Chou <max.chou@sifive.com>
---
MAINTAINERS | 1 +
disas/meson.build | 3 ++-
disas/riscv-xsf-op.c.inc | 11 +++++++++
disas/riscv-xsf.c | 49 ++++++++++++++++++++++++++++++++++++++++
disas/riscv-xsf.h | 16 +++++++++++++
disas/riscv.c | 2 ++
disas/riscv.h | 1 +
7 files changed, 82 insertions(+), 1 deletion(-)
create mode 100644 disas/riscv-xsf-op.c.inc
create mode 100644 disas/riscv-xsf.c
create mode 100644 disas/riscv-xsf.h
diff --git a/MAINTAINERS b/MAINTAINERS
index 2e8dace19ff..629a238c925 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -405,6 +405,7 @@ S: Supported
F: target/riscv/xsf.decode
F: target/riscv/tcg/insn_trans/trans_xsf.c.inc
F: target/riscv/tcg/xsf_helper.c
+F: disas/riscv-xsf*
RENESAS RX CPUs
R: Yoshinori Sato <yoshinori.sato@nifty.com>
diff --git a/disas/meson.build b/disas/meson.build
index 42977a1f74d..91357162682 100644
--- a/disas/meson.build
+++ b/disas/meson.build
@@ -8,7 +8,8 @@ common_ss.add(when: 'CONFIG_RISCV_DIS', if_true: files(
'riscv.c',
'riscv-xthead.c',
'riscv-xventana.c',
- 'riscv-xlrbr.c'
+ 'riscv-xlrbr.c',
+ 'riscv-xsf.c'
))
common_ss.add(when: 'CONFIG_SH4_DIS', if_true: files('sh4.c'))
common_ss.add(when: 'CONFIG_SPARC_DIS', if_true: files('sparc.c'))
diff --git a/disas/riscv-xsf-op.c.inc b/disas/riscv-xsf-op.c.inc
new file mode 100644
index 00000000000..c8daa7ebef8
--- /dev/null
+++ b/disas/riscv-xsf-op.c.inc
@@ -0,0 +1,11 @@
+/* Xsfvqmaccqoq */
+OP(sf_vqmaccu_4x8x4, "sf.vqmaccu.4x8x4", rv_codec_v_r, rv_fmt_vd_vs1_vs2)
+OP(sf_vqmacc_4x8x4, "sf.vqmacc.4x8x4", rv_codec_v_r, rv_fmt_vd_vs1_vs2)
+OP(sf_vqmaccus_4x8x4, "sf.vqmaccus.4x8x4", rv_codec_v_r, rv_fmt_vd_vs1_vs2)
+OP(sf_vqmaccsu_4x8x4, "sf.vqmaccsu.4x8x4", rv_codec_v_r, rv_fmt_vd_vs1_vs2)
+
+/* Xsfvqmaccdod */
+OP(sf_vqmaccu_2x8x2, "sf.vqmaccu.2x8x2", rv_codec_v_r, rv_fmt_vd_vs1_vs2)
+OP(sf_vqmacc_2x8x2, "sf.vqmacc.2x8x2", rv_codec_v_r, rv_fmt_vd_vs1_vs2)
+OP(sf_vqmaccus_2x8x2, "sf.vqmaccus.2x8x2", rv_codec_v_r, rv_fmt_vd_vs1_vs2)
+OP(sf_vqmaccsu_2x8x2, "sf.vqmaccsu.2x8x2", rv_codec_v_r, rv_fmt_vd_vs1_vs2)
diff --git a/disas/riscv-xsf.c b/disas/riscv-xsf.c
new file mode 100644
index 00000000000..745fcc8efda
--- /dev/null
+++ b/disas/riscv-xsf.c
@@ -0,0 +1,49 @@
+/*
+ * QEMU RISC-V Disassembler for xsf (SiFive vendor extensions).
+ *
+ * Copyright (c) 2023 SiFive, Inc.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#include "qemu/osdep.h"
+#include "disas/riscv.h"
+#include "disas/riscv-xsf.h"
+
+#define OP(N, ...) static const rv_opcode_data op_##N = { __VA_ARGS__ };
+#include "riscv-xsf-op.c.inc"
+#undef OP
+
+const rv_opcode_data *decode_xsf(rv_decode *dec, rv_isa isa)
+{
+ rv_inst inst = dec->inst;
+
+ switch ((inst >> 0) & 0b1111111) {
+ case 0b1011011:
+ switch ((inst >> 12) & 0b111) {
+ case 0b010:
+ switch ((inst >> 26) & 0b111111) {
+ case 0b101100:
+ return &op_sf_vqmaccu_2x8x2;
+ case 0b101101:
+ return &op_sf_vqmacc_2x8x2;
+ case 0b101110:
+ return &op_sf_vqmaccus_2x8x2;
+ case 0b101111:
+ return &op_sf_vqmaccsu_2x8x2;
+ case 0b111100:
+ return &op_sf_vqmaccu_4x8x4;
+ case 0b111101:
+ return &op_sf_vqmacc_4x8x4;
+ case 0b111110:
+ return &op_sf_vqmaccus_4x8x4;
+ case 0b111111:
+ return &op_sf_vqmaccsu_4x8x4;
+ }
+ break;
+ }
+ break;
+ }
+
+ return NULL;
+}
diff --git a/disas/riscv-xsf.h b/disas/riscv-xsf.h
new file mode 100644
index 00000000000..b15b7aa97c8
--- /dev/null
+++ b/disas/riscv-xsf.h
@@ -0,0 +1,16 @@
+/*
+ * QEMU disassembler -- RISC-V specific header (xsf*).
+ *
+ * Copyright (c) 2023 SiFive, Inc.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+#ifndef DISAS_RISCV_XSF_H
+#define DISAS_RISCV_XSF_H
+
+#include "disas/riscv.h"
+
+const rv_opcode_data *decode_xsf(rv_decode *, rv_isa);
+
+#endif /* DISAS_RISCV_XSF_H */
diff --git a/disas/riscv.c b/disas/riscv.c
index e85dbac0351..60d934744de 100644
--- a/disas/riscv.c
+++ b/disas/riscv.c
@@ -27,6 +27,7 @@
#include "disas/riscv-xthead.h"
#include "disas/riscv-xventana.h"
#include "disas/riscv-xlrbr.h"
+#include "disas/riscv-xsf.h"
/* register names */
@@ -3457,6 +3458,7 @@ static GString *disasm_inst(rv_isa isa, uint64_t pc, rv_inst inst,
{ has_xtheadsync_p, decode_xtheadsync },
{ has_XVentanaCondOps_p, decode_xventanacondops },
{ has_xlrbr_p, decode_xlrbr },
+ { has_xsf_p, decode_xsf },
};
for (size_t i = 0; i < ARRAY_SIZE(decoders); i++) {
diff --git a/disas/riscv.h b/disas/riscv.h
index b0269bb709f..4f1ffc4ae02 100644
--- a/disas/riscv.h
+++ b/disas/riscv.h
@@ -262,6 +262,7 @@ typedef struct {
#define rv_fmt_vd_vs2_imm_vm "O\tD,F,im"
#define rv_fmt_vd_vs2_uimm "O\tD,F,u"
#define rv_fmt_vd_vs2_uimm_vm "O\tD,F,um"
+#define rv_fmt_vd_vs1_vs2 "O\tD,E,F"
#define rv_fmt_vd_vs1_vs2_vm "O\tD,E,Fm"
#define rv_fmt_vd_rs1_vs2_vm "O\tD,1,Fm"
#define rv_fmt_vd_fs1_vs2_vm "O\tD,4,Fm"
--
2.43.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* [PATCH v3 6/6] tests/tcg/riscv64: Add tests for SiFive int8 matmul extensions
2026-08-30 16:03 [PATCH v3 0/6] target/riscv: Add SiFive Xsfvqmaccdod/Xsfvqmaccqoq int8 matmul extensions Max Chou
` (4 preceding siblings ...)
2026-08-30 16:03 ` [PATCH v3 5/6] disas/riscv: Add disassembler support for Xsfvqmaccdod/Xsfvqmaccqoq Max Chou
@ 2026-08-30 16:03 ` Max Chou
2026-09-07 4:22 ` Alistair Francis
5 siblings, 1 reply; 9+ messages in thread
From: Max Chou @ 2026-08-30 16:03 UTC (permalink / raw)
To: qemu-devel, qemu-riscv
Cc: Palmer Dabbelt, Alistair Francis, Daniel Henrique Barboza,
Max Chou, Weiwei Li, Liu Zhiwei, Chao Liu, Alistair Francis
Add a test that exercises all SiFive int8 matrix-multiply instructions
from the Xsfvqmaccqoq and Xsfvqmaccdod extensions.
For each tile shape the test covers the four signedness variants
(sf.vqmaccu / sf.vqmacc / sf.vqmaccus / sf.vqmaccsu) and the
reserved-encoding path: with vl set to a value that is not a multiple
of the tile size the helper raises an illegal-instruction exception.
Reviewed-by: Daniel Henrique Barboza <daniel.barboza@oss.qualcomm.com>
Signed-off-by: Max Chou <max.chou@sifive.com>
---
MAINTAINERS | 1 +
tests/tcg/riscv64/Makefile.softmmu-target | 7 +
tests/tcg/riscv64/test-xsfvqmacc.S | 260 ++++++++++++++++++++++
3 files changed, 268 insertions(+)
create mode 100644 tests/tcg/riscv64/test-xsfvqmacc.S
diff --git a/MAINTAINERS b/MAINTAINERS
index 629a238c925..5218bda63c6 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -406,6 +406,7 @@ F: target/riscv/xsf.decode
F: target/riscv/tcg/insn_trans/trans_xsf.c.inc
F: target/riscv/tcg/xsf_helper.c
F: disas/riscv-xsf*
+F: tests/tcg/riscv64/test-xsfvqmacc.S
RENESAS RX CPUs
R: Yoshinori Sato <yoshinori.sato@nifty.com>
diff --git a/tests/tcg/riscv64/Makefile.softmmu-target b/tests/tcg/riscv64/Makefile.softmmu-target
index 6a219c306c4..c8ffa5ed765 100644
--- a/tests/tcg/riscv64/Makefile.softmmu-target
+++ b/tests/tcg/riscv64/Makefile.softmmu-target
@@ -71,5 +71,12 @@ EXTRA_RUNS += run-test-misa-w
run-test-misa-w: test-misa-w
$(call run-test, $<, $(QEMU) -cpu rv64$(comma)x-misa-w=true$(comma)c=true$(comma)v=true $(QEMU_OPTS)$<)
+# Xsfvqmaccdod/Xsfvqmaccqoq: the SiFive int8 matmul instructions are
+# reserved unless VLEN >= 256, so selects vlen=256 along with both extensions.
+EXTRA_RUNS += run-test-xsfvqmacc
+XSFVQMACC_CPU = rv64$(comma)v=true$(comma)vlen=256$(comma)xsfvqmaccdod=true$(comma)xsfvqmaccqoq=true
+run-test-xsfvqmacc: test-xsfvqmacc
+ $(call run-test, $<, $(QEMU) -cpu $(XSFVQMACC_CPU) $(QEMU_OPTS)$<)
+
# We don't currently support the multiarch system tests
undefine MULTIARCH_TESTS
diff --git a/tests/tcg/riscv64/test-xsfvqmacc.S b/tests/tcg/riscv64/test-xsfvqmacc.S
new file mode 100644
index 00000000000..a116712e502
--- /dev/null
+++ b/tests/tcg/riscv64/test-xsfvqmacc.S
@@ -0,0 +1,260 @@
+/*
+ * Test the SiFive Xsfvqmaccqoq/Xsfvqmaccdod custom int8 matrix-multiply
+ * extensions.
+ *
+ * The eight instructions widen 8-bit integer inputs to 32 bits and perform
+ * a matrix-multiply-accumulate into a vector of int32 accumulators:
+ *
+ * C[j] += A * B[j]
+ *
+ * Xsfvqmaccqoq (4x8x4): A is 4x8, B[j] is 8x4, C[j] is 4x4 int32. One matrix
+ * operation is performed per 32 elements of vl (TILE_SIZE = 32).
+ * Xsfvqmaccdod (2x8x2): A is 2x8, B[j] is 8x2, C[j] is 2x2 int32. One matrix
+ * operation is performed per 16 elements of vl (TILE_SIZE = 16).
+ *
+ * The input tiles below contain bytes >= 0x80 so that each signedness
+ * variant produces a distinct result; the expected int32 tiles were computed
+ * with an independent reference model of the specification. The accumulators
+ * are cleared before every operation, so the expected value is the pure
+ * matrix product.
+ *
+ * Test exits via semihosting with status 0 on success, or the 1-based index
+ * of the first failing operation.
+ *
+ * SPDX-License-Identifier: GPL-2.0-or-later
+ */
+
+ .option arch, +v
+ .option norvc
+
+ .text
+
+/*
+ * Run one 4x8x4 operation: load the 32-byte A and B tiles, clear the 4x4
+ * int32 accumulator (vd = v4:v5, EMUL = 2), execute the op, then compare the
+ * 16 int32 results against the expected tile.
+ */
+.macro RUN_QOQ id, word, exp
+ li a2, 32
+ vsetvli t0, a2, e8, m1, ta, ma
+ la t1, a_qoq
+ vle8.v v1, (t1)
+ la t1, b_qoq
+ vle8.v v2, (t1)
+ vsetvli t0, x0, e32, m2, ta, ma
+ vmv.v.i v4, 0
+ li a2, 32
+ vsetvli t0, a2, e8, m1, ta, ma
+ .word \word
+ vsetvli t0, x0, e32, m2, ta, ma
+ la t1, result
+ vse32.v v4, (t1)
+ la a0, result
+ la a1, \exp
+ li a2, 16
+ li t6, \id
+ call check_words
+.endm
+
+/*
+ * Run one 2x8x2 operation: load the 16-byte A and B tiles, clear the 2x2
+ * int32 accumulator (vd = v4, EMUL = 1), execute the op, then compare the
+ * 4 int32 results against the expected tile.
+ */
+.macro RUN_DOD id, word, exp
+ li a2, 16
+ vsetvli t0, a2, e8, m1, ta, ma
+ la t1, a_dod
+ vle8.v v1, (t1)
+ la t1, b_dod
+ vle8.v v2, (t1)
+ vsetvli t0, x0, e32, m1, ta, ma
+ vmv.v.i v4, 0
+ li a2, 16
+ vsetvli t0, a2, e8, m1, ta, ma
+ .word \word
+ vsetvli t0, x0, e32, m1, ta, ma
+ la t1, result
+ vse32.v v4, (t1)
+ la a0, result
+ la a1, \exp
+ li a2, 4
+ li t6, \id
+ call check_words
+.endm
+
+ .global _start
+_start:
+ /* Enable the vector unit (mstatus.VS = Initial). */
+ li t0, 1 << 9
+ csrs mstatus, t0
+
+ /* Route synchronous traps to trap_handler (mtvec direct mode). */
+ la t0, trap_handler
+ csrw mtvec, t0
+
+ RUN_QOQ 1, 0xF220A25B, exp_qoq_u /* sf.vqmaccu.4x8x4 */
+ RUN_QOQ 2, 0xF620A25B, exp_qoq_s /* sf.vqmacc.4x8x4 */
+ RUN_QOQ 3, 0xFA20A25B, exp_qoq_us /* sf.vqmaccus.4x8x4 */
+ RUN_QOQ 4, 0xFE20A25B, exp_qoq_su /* sf.vqmaccsu.4x8x4 */
+ RUN_DOD 5, 0xB220A25B, exp_dod_u /* sf.vqmaccu.2x8x2 */
+ RUN_DOD 6, 0xB620A25B, exp_dod_s /* sf.vqmacc.2x8x2 */
+ RUN_DOD 7, 0xBA20A25B, exp_dod_us /* sf.vqmaccus.2x8x2 */
+ RUN_DOD 8, 0xBE20A25B, exp_dod_su /* sf.vqmaccsu.2x8x2 */
+
+ /*
+ * Illegal-instruction reporting test (op 9).
+ *
+ * When vl is not a multiple of the tile size the instruction is
+ * reserved and the helper raises an illegal-instruction exception at
+ * runtime.
+ */
+ la t1, trap_mtval
+ sd zero, 0(t1)
+ la t1, trap_mcause
+ sd zero, 0(t1)
+ /* vl = 16 is not a multiple of the 4x8x4 tile size (32 elements). */
+ li a2, 16
+ vsetvli t0, a2, e8, m1, ta, ma
+ .word 0xF220A25B /* sf.vqmaccu.4x8x4, traps */
+ /* The handler skips the faulting insn, so control resumes here. */
+ li a0, 9 /* fail code for this check */
+ la t1, trap_mcause
+ ld t2, 0(t1)
+ li t3, 2 /* RISCV_EXCP_ILLEGAL_INST */
+ bne t2, t3, _exit
+ la t1, trap_mtval
+ ld t2, 0(t1)
+ li t3, 0xF220A25B
+ bne t2, t3, _exit
+
+ li a0, 0
+ j _exit
+
+/*
+ * Machine-mode trap handler: record mcause and mtval, then advance mepc
+ * past the 4-byte faulting instruction and return.
+ */
+ .balign 4
+trap_handler:
+ csrr t4, mcause
+ la t5, trap_mcause
+ sd t4, 0(t5)
+ csrr t4, mtval
+ la t5, trap_mtval
+ sd t4, 0(t5)
+ csrr t4, mepc
+ addi t4, t4, 4
+ csrw mepc, t4
+ mret
+
+/*
+ * check_words(a0 = result, a1 = expected, a2 = word count).
+ * Returns to the caller if every word matches; otherwise exits with the
+ * operation index held in t6.
+ */
+check_words:
+1: beqz a2, 2f
+ lw t2, 0(a0)
+ lw t3, 0(a1)
+ bne t2, t3, 3f
+ addi a0, a0, 4
+ addi a1, a1, 4
+ addi a2, a2, -1
+ j 1b
+2: ret
+3: mv a0, t6
+ j _exit
+
+/* Exit through the semihosting SYS_EXIT_EXTENDED call with a0 as the code. */
+_exit:
+ la a1, semiargs
+ li t0, 0x20026 /* ADP_Stopped_ApplicationExit */
+ sd t0, 0(a1)
+ sd a0, 8(a1)
+ li a0, 0x20 /* TARGET_SYS_EXIT_EXTENDED */
+ .balign 16
+ slli zero, zero, 0x1f
+ ebreak
+ srai zero, zero, 0x7
+ j .
+
+ .data
+ .balign 8
+semiargs:
+ .space 16
+result:
+ .space 64
+trap_mcause:
+ .space 8
+trap_mtval:
+ .space 8
+
+/*
+ * A (4x8, row-major) in bytes [0, 32) of vs1.
+ */
+a_qoq:
+ .byte 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x80
+ .byte 0xFF, 0x01, 0x00, 0x02, 0x7F, 0x03, 0x04, 0x05
+ .byte 0x10, 0x20, 0x30, 0x40, 0x50, 0x60, 0x70, 0x08
+ .byte 0x00, 0x00, 0x81, 0x00, 0x00, 0x02, 0x00, 0x03
+/*
+ * B (8x4, row-major) in bytes [0, 32) of vs2.
+ */
+b_qoq:
+ .byte 0x01, 0x02, 0x03, 0x04
+ .byte 0x05, 0x06, 0x07, 0x08
+ .byte 0x80, 0x7F, 0xFF, 0x01
+ .byte 0x00, 0x10, 0x20, 0x30
+ .byte 0x02, 0x02, 0x02, 0x02
+ .byte 0xFE, 0x01, 0x00, 0x7F
+ .byte 0x03, 0x00, 0x81, 0x04
+ .byte 0x40, 0x50, 0x60, 0x70
+/*
+ * A (2x8, row-major) in bytes [0, 16) of vs1.
+ */
+a_dod:
+ .byte 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x80
+ .byte 0xFF, 0x7F, 0x00, 0x02, 0x81, 0x03, 0x04, 0x05
+/*
+ * B (8x2, row-major) in bytes [0, 16) of vs2.
+ */
+b_dod:
+ .byte 0x01, 0x02
+ .byte 0x80, 0x7F
+ .byte 0x03, 0xFF
+ .byte 0x00, 0x10
+ .byte 0x02, 0xFE
+ .byte 0x7F, 0x01
+ .byte 0x81, 0x04
+ .byte 0x40, 0x50
+
+ .balign 4
+exp_qoq_u:
+ .word 0x0000279e, 0x000029db, 0x0000371f, 0x00003bf7
+ .word 0x00000648, 0x000004b5, 0x00000826, 0x0000091f
+ .word 0x00007be0, 0x00002030, 0x000074f0, 0x000042f0
+ .word 0x0000433c, 0x000040f1, 0x0000819f, 0x000002cf
+exp_qoq_s:
+ .word 0xffffde9e, 0xffffd9db, 0xffffcd1f, 0xffffcbf7
+ .word 0x00000248, 0x000002b5, 0x00000126, 0x0000051f
+ .word 0xffffebe0, 0x00002030, 0xffffd4f0, 0x000042f0
+ .word 0x0000403c, 0xffffc1f1, 0x0000019f, 0x000001cf
+exp_qoq_us:
+ .word 0x00001e9e, 0x000029db, 0x00002d1f, 0x00003bf7
+ .word 0x00000348, 0x000004b5, 0x00000426, 0x0000091f
+ .word 0xffffebe0, 0x00002030, 0xffffd4f0, 0x000042f0
+ .word 0xffffc03c, 0x000040f1, 0x0000009f, 0x000002cf
+exp_qoq_su:
+ .word 0xffffe79e, 0xffffd9db, 0xffffd71f, 0xffffcbf7
+ .word 0x00000548, 0x000002b5, 0x00000526, 0x0000051f
+ .word 0x00007be0, 0x00002030, 0x000074f0, 0x000042f0
+ .word 0xffffc33c, 0xffffc1f1, 0xffff829f, 0x000001cf
+exp_dod_u:
+ .word 0x00002795, 0x00003155, 0x00004642, 0x0000c2c0
+exp_dod_s:
+ .word 0xffffde95, 0xffffd955, 0xffffc042, 0x000041c0
+exp_dod_us:
+ .word 0x00001e95, 0x00002955, 0xffffc342, 0x000041c0
+exp_dod_su:
+ .word 0xffffe795, 0xffffe155, 0x00004342, 0xffffc2c0
--
2.43.0
^ permalink raw reply related [flat|nested] 9+ messages in thread
* Re: [PATCH v3 5/6] disas/riscv: Add disassembler support for Xsfvqmaccdod/Xsfvqmaccqoq
2026-08-30 16:03 ` [PATCH v3 5/6] disas/riscv: Add disassembler support for Xsfvqmaccdod/Xsfvqmaccqoq Max Chou
@ 2026-09-07 4:12 ` Alistair Francis
0 siblings, 0 replies; 9+ messages in thread
From: Alistair Francis @ 2026-09-07 4:12 UTC (permalink / raw)
To: max.chou@sifive.com, qemu-riscv@nongnu.org, qemu-devel@nongnu.org
Cc: palmer@dabbelt.com, daniel.barboza@oss.qualcomm.com,
chao.liu@processmission.com, jason.chien@sifive.com,
liwei1518@gmail.com, zhiwei_liu@linux.alibaba.com
On Mon, 2026-08-31 at 00:03 +0800, Max Chou wrote:
> From: Jason Chien <jason.chien@sifive.com>
>
> Add disassembler support for all SiFive custom int8 matrix-multiply
> instructions: sf.vqmacc{u,,us,su}.4x8x4 and
> sf.vqmacc{u,,us,su}.2x8x2.
>
> These instructions share opcode 0b1011011 and funct3 0b010, and are
> distinguished by funct6: 60-63 for the 4x8x4 (Xsfvqmaccqoq) tile and
> 44-47 for the 2x8x2 (Xsfvqmaccdod) tile.
>
> Signed-off-by: Jason Chien <jason.chien@sifive.com>
> Reviewed-by: Daniel Henrique Barboza
> <daniel.barboza@oss.qualcomm.com>
> Signed-off-by: Max Chou <max.chou@sifive.com>
Acked-by: Alistair Francis <alistair.francis@wdc.com>
Alistair
> ---
> MAINTAINERS | 1 +
> disas/meson.build | 3 ++-
> disas/riscv-xsf-op.c.inc | 11 +++++++++
> disas/riscv-xsf.c | 49
> ++++++++++++++++++++++++++++++++++++++++
> disas/riscv-xsf.h | 16 +++++++++++++
> disas/riscv.c | 2 ++
> disas/riscv.h | 1 +
> 7 files changed, 82 insertions(+), 1 deletion(-)
> create mode 100644 disas/riscv-xsf-op.c.inc
> create mode 100644 disas/riscv-xsf.c
> create mode 100644 disas/riscv-xsf.h
>
> diff --git a/MAINTAINERS b/MAINTAINERS
> index 2e8dace19ff..629a238c925 100644
> --- a/MAINTAINERS
> +++ b/MAINTAINERS
> @@ -405,6 +405,7 @@ S: Supported
> F: target/riscv/xsf.decode
> F: target/riscv/tcg/insn_trans/trans_xsf.c.inc
> F: target/riscv/tcg/xsf_helper.c
> +F: disas/riscv-xsf*
>
> RENESAS RX CPUs
> R: Yoshinori Sato <yoshinori.sato@nifty.com>
> diff --git a/disas/meson.build b/disas/meson.build
> index 42977a1f74d..91357162682 100644
> --- a/disas/meson.build
> +++ b/disas/meson.build
> @@ -8,7 +8,8 @@ common_ss.add(when: 'CONFIG_RISCV_DIS', if_true:
> files(
> 'riscv.c',
> 'riscv-xthead.c',
> 'riscv-xventana.c',
> - 'riscv-xlrbr.c'
> + 'riscv-xlrbr.c',
> + 'riscv-xsf.c'
> ))
> common_ss.add(when: 'CONFIG_SH4_DIS', if_true: files('sh4.c'))
> common_ss.add(when: 'CONFIG_SPARC_DIS', if_true: files('sparc.c'))
> diff --git a/disas/riscv-xsf-op.c.inc b/disas/riscv-xsf-op.c.inc
> new file mode 100644
> index 00000000000..c8daa7ebef8
> --- /dev/null
> +++ b/disas/riscv-xsf-op.c.inc
> @@ -0,0 +1,11 @@
> +/* Xsfvqmaccqoq */
> +OP(sf_vqmaccu_4x8x4, "sf.vqmaccu.4x8x4", rv_codec_v_r,
> rv_fmt_vd_vs1_vs2)
> +OP(sf_vqmacc_4x8x4, "sf.vqmacc.4x8x4", rv_codec_v_r,
> rv_fmt_vd_vs1_vs2)
> +OP(sf_vqmaccus_4x8x4, "sf.vqmaccus.4x8x4", rv_codec_v_r,
> rv_fmt_vd_vs1_vs2)
> +OP(sf_vqmaccsu_4x8x4, "sf.vqmaccsu.4x8x4", rv_codec_v_r,
> rv_fmt_vd_vs1_vs2)
> +
> +/* Xsfvqmaccdod */
> +OP(sf_vqmaccu_2x8x2, "sf.vqmaccu.2x8x2", rv_codec_v_r,
> rv_fmt_vd_vs1_vs2)
> +OP(sf_vqmacc_2x8x2, "sf.vqmacc.2x8x2", rv_codec_v_r,
> rv_fmt_vd_vs1_vs2)
> +OP(sf_vqmaccus_2x8x2, "sf.vqmaccus.2x8x2", rv_codec_v_r,
> rv_fmt_vd_vs1_vs2)
> +OP(sf_vqmaccsu_2x8x2, "sf.vqmaccsu.2x8x2", rv_codec_v_r,
> rv_fmt_vd_vs1_vs2)
> diff --git a/disas/riscv-xsf.c b/disas/riscv-xsf.c
> new file mode 100644
> index 00000000000..745fcc8efda
> --- /dev/null
> +++ b/disas/riscv-xsf.c
> @@ -0,0 +1,49 @@
> +/*
> + * QEMU RISC-V Disassembler for xsf (SiFive vendor extensions).
> + *
> + * Copyright (c) 2023 SiFive, Inc.
> + *
> + * SPDX-License-Identifier: GPL-2.0-or-later
> + */
> +
> +#include "qemu/osdep.h"
> +#include "disas/riscv.h"
> +#include "disas/riscv-xsf.h"
> +
> +#define OP(N, ...) static const rv_opcode_data op_##N = {
> __VA_ARGS__ };
> +#include "riscv-xsf-op.c.inc"
> +#undef OP
> +
> +const rv_opcode_data *decode_xsf(rv_decode *dec, rv_isa isa)
> +{
> + rv_inst inst = dec->inst;
> +
> + switch ((inst >> 0) & 0b1111111) {
> + case 0b1011011:
> + switch ((inst >> 12) & 0b111) {
> + case 0b010:
> + switch ((inst >> 26) & 0b111111) {
> + case 0b101100:
> + return &op_sf_vqmaccu_2x8x2;
> + case 0b101101:
> + return &op_sf_vqmacc_2x8x2;
> + case 0b101110:
> + return &op_sf_vqmaccus_2x8x2;
> + case 0b101111:
> + return &op_sf_vqmaccsu_2x8x2;
> + case 0b111100:
> + return &op_sf_vqmaccu_4x8x4;
> + case 0b111101:
> + return &op_sf_vqmacc_4x8x4;
> + case 0b111110:
> + return &op_sf_vqmaccus_4x8x4;
> + case 0b111111:
> + return &op_sf_vqmaccsu_4x8x4;
> + }
> + break;
> + }
> + break;
> + }
> +
> + return NULL;
> +}
> diff --git a/disas/riscv-xsf.h b/disas/riscv-xsf.h
> new file mode 100644
> index 00000000000..b15b7aa97c8
> --- /dev/null
> +++ b/disas/riscv-xsf.h
> @@ -0,0 +1,16 @@
> +/*
> + * QEMU disassembler -- RISC-V specific header (xsf*).
> + *
> + * Copyright (c) 2023 SiFive, Inc.
> + *
> + * SPDX-License-Identifier: GPL-2.0-or-later
> + */
> +
> +#ifndef DISAS_RISCV_XSF_H
> +#define DISAS_RISCV_XSF_H
> +
> +#include "disas/riscv.h"
> +
> +const rv_opcode_data *decode_xsf(rv_decode *, rv_isa);
> +
> +#endif /* DISAS_RISCV_XSF_H */
> diff --git a/disas/riscv.c b/disas/riscv.c
> index e85dbac0351..60d934744de 100644
> --- a/disas/riscv.c
> +++ b/disas/riscv.c
> @@ -27,6 +27,7 @@
> #include "disas/riscv-xthead.h"
> #include "disas/riscv-xventana.h"
> #include "disas/riscv-xlrbr.h"
> +#include "disas/riscv-xsf.h"
>
> /* register names */
>
> @@ -3457,6 +3458,7 @@ static GString *disasm_inst(rv_isa isa,
> uint64_t pc, rv_inst inst,
> { has_xtheadsync_p, decode_xtheadsync },
> { has_XVentanaCondOps_p, decode_xventanacondops },
> { has_xlrbr_p, decode_xlrbr },
> + { has_xsf_p, decode_xsf },
> };
>
> for (size_t i = 0; i < ARRAY_SIZE(decoders); i++) {
> diff --git a/disas/riscv.h b/disas/riscv.h
> index b0269bb709f..4f1ffc4ae02 100644
> --- a/disas/riscv.h
> +++ b/disas/riscv.h
> @@ -262,6 +262,7 @@ typedef struct {
> #define rv_fmt_vd_vs2_imm_vm "O\tD,F,im"
> #define rv_fmt_vd_vs2_uimm "O\tD,F,u"
> #define rv_fmt_vd_vs2_uimm_vm "O\tD,F,um"
> +#define rv_fmt_vd_vs1_vs2 "O\tD,E,F"
> #define rv_fmt_vd_vs1_vs2_vm "O\tD,E,Fm"
> #define rv_fmt_vd_rs1_vs2_vm "O\tD,1,Fm"
> #define rv_fmt_vd_fs1_vs2_vm "O\tD,4,Fm"
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH v3 6/6] tests/tcg/riscv64: Add tests for SiFive int8 matmul extensions
2026-08-30 16:03 ` [PATCH v3 6/6] tests/tcg/riscv64: Add tests for SiFive int8 matmul extensions Max Chou
@ 2026-09-07 4:22 ` Alistair Francis
0 siblings, 0 replies; 9+ messages in thread
From: Alistair Francis @ 2026-09-07 4:22 UTC (permalink / raw)
To: max.chou@sifive.com, qemu-riscv@nongnu.org, qemu-devel@nongnu.org
Cc: palmer@dabbelt.com, daniel.barboza@oss.qualcomm.com,
chao.liu@processmission.com, liwei1518@gmail.com,
zhiwei_liu@linux.alibaba.com
On Mon, 2026-08-31 at 00:03 +0800, Max Chou wrote:
> Add a test that exercises all SiFive int8 matrix-multiply
> instructions
> from the Xsfvqmaccqoq and Xsfvqmaccdod extensions.
>
> For each tile shape the test covers the four signedness variants
> (sf.vqmaccu / sf.vqmacc / sf.vqmaccus / sf.vqmaccsu) and the
> reserved-encoding path: with vl set to a value that is not a multiple
> of the tile size the helper raises an illegal-instruction exception.
>
> Reviewed-by: Daniel Henrique Barboza
> <daniel.barboza@oss.qualcomm.com>
> Signed-off-by: Max Chou <max.chou@sifive.com>
Acked-by: Alistair Francis <alistair.francis@wdc.com>
This will need a rebase on the new meson build system though
Alistair
> ---
> MAINTAINERS | 1 +
> tests/tcg/riscv64/Makefile.softmmu-target | 7 +
> tests/tcg/riscv64/test-xsfvqmacc.S | 260
> ++++++++++++++++++++++
> 3 files changed, 268 insertions(+)
> create mode 100644 tests/tcg/riscv64/test-xsfvqmacc.S
>
> diff --git a/MAINTAINERS b/MAINTAINERS
> index 629a238c925..5218bda63c6 100644
> --- a/MAINTAINERS
> +++ b/MAINTAINERS
> @@ -406,6 +406,7 @@ F: target/riscv/xsf.decode
> F: target/riscv/tcg/insn_trans/trans_xsf.c.inc
> F: target/riscv/tcg/xsf_helper.c
> F: disas/riscv-xsf*
> +F: tests/tcg/riscv64/test-xsfvqmacc.S
>
> RENESAS RX CPUs
> R: Yoshinori Sato <yoshinori.sato@nifty.com>
> diff --git a/tests/tcg/riscv64/Makefile.softmmu-target
> b/tests/tcg/riscv64/Makefile.softmmu-target
> index 6a219c306c4..c8ffa5ed765 100644
> --- a/tests/tcg/riscv64/Makefile.softmmu-target
> +++ b/tests/tcg/riscv64/Makefile.softmmu-target
> @@ -71,5 +71,12 @@ EXTRA_RUNS += run-test-misa-w
> run-test-misa-w: test-misa-w
> $(call run-test, $<, $(QEMU) -cpu rv64$(comma)x-misa-
> w=true$(comma)c=true$(comma)v=true $(QEMU_OPTS)$<)
>
> +# Xsfvqmaccdod/Xsfvqmaccqoq: the SiFive int8 matmul instructions are
> +# reserved unless VLEN >= 256, so selects vlen=256 along with both
> extensions.
> +EXTRA_RUNS += run-test-xsfvqmacc
> +XSFVQMACC_CPU =
> rv64$(comma)v=true$(comma)vlen=256$(comma)xsfvqmaccdod=true$(comma)xs
> fvqmaccqoq=true
> +run-test-xsfvqmacc: test-xsfvqmacc
> + $(call run-test, $<, $(QEMU) -cpu $(XSFVQMACC_CPU)
> $(QEMU_OPTS)$<)
> +
> # We don't currently support the multiarch system tests
> undefine MULTIARCH_TESTS
> diff --git a/tests/tcg/riscv64/test-xsfvqmacc.S
> b/tests/tcg/riscv64/test-xsfvqmacc.S
> new file mode 100644
> index 00000000000..a116712e502
> --- /dev/null
> +++ b/tests/tcg/riscv64/test-xsfvqmacc.S
> @@ -0,0 +1,260 @@
> +/*
> + * Test the SiFive Xsfvqmaccqoq/Xsfvqmaccdod custom int8 matrix-
> multiply
> + * extensions.
> + *
> + * The eight instructions widen 8-bit integer inputs to 32 bits and
> perform
> + * a matrix-multiply-accumulate into a vector of int32 accumulators:
> + *
> + * C[j] += A * B[j]
> + *
> + * Xsfvqmaccqoq (4x8x4): A is 4x8, B[j] is 8x4, C[j] is 4x4 int32.
> One matrix
> + * operation is performed per 32 elements of vl (TILE_SIZE = 32).
> + * Xsfvqmaccdod (2x8x2): A is 2x8, B[j] is 8x2, C[j] is 2x2 int32.
> One matrix
> + * operation is performed per 16 elements of vl (TILE_SIZE = 16).
> + *
> + * The input tiles below contain bytes >= 0x80 so that each
> signedness
> + * variant produces a distinct result; the expected int32 tiles were
> computed
> + * with an independent reference model of the specification. The
> accumulators
> + * are cleared before every operation, so the expected value is the
> pure
> + * matrix product.
> + *
> + * Test exits via semihosting with status 0 on success, or the 1-
> based index
> + * of the first failing operation.
> + *
> + * SPDX-License-Identifier: GPL-2.0-or-later
> + */
> +
> + .option arch, +v
> + .option norvc
> +
> + .text
> +
> +/*
> + * Run one 4x8x4 operation: load the 32-byte A and B tiles, clear
> the 4x4
> + * int32 accumulator (vd = v4:v5, EMUL = 2), execute the op, then
> compare the
> + * 16 int32 results against the expected tile.
> + */
> +.macro RUN_QOQ id, word, exp
> + li a2, 32
> + vsetvli t0, a2, e8, m1, ta, ma
> + la t1, a_qoq
> + vle8.v v1, (t1)
> + la t1, b_qoq
> + vle8.v v2, (t1)
> + vsetvli t0, x0, e32, m2, ta, ma
> + vmv.v.i v4, 0
> + li a2, 32
> + vsetvli t0, a2, e8, m1, ta, ma
> + .word \word
> + vsetvli t0, x0, e32, m2, ta, ma
> + la t1, result
> + vse32.v v4, (t1)
> + la a0, result
> + la a1, \exp
> + li a2, 16
> + li t6, \id
> + call check_words
> +.endm
> +
> +/*
> + * Run one 2x8x2 operation: load the 16-byte A and B tiles, clear
> the 2x2
> + * int32 accumulator (vd = v4, EMUL = 1), execute the op, then
> compare the
> + * 4 int32 results against the expected tile.
> + */
> +.macro RUN_DOD id, word, exp
> + li a2, 16
> + vsetvli t0, a2, e8, m1, ta, ma
> + la t1, a_dod
> + vle8.v v1, (t1)
> + la t1, b_dod
> + vle8.v v2, (t1)
> + vsetvli t0, x0, e32, m1, ta, ma
> + vmv.v.i v4, 0
> + li a2, 16
> + vsetvli t0, a2, e8, m1, ta, ma
> + .word \word
> + vsetvli t0, x0, e32, m1, ta, ma
> + la t1, result
> + vse32.v v4, (t1)
> + la a0, result
> + la a1, \exp
> + li a2, 4
> + li t6, \id
> + call check_words
> +.endm
> +
> + .global _start
> +_start:
> + /* Enable the vector unit (mstatus.VS = Initial). */
> + li t0, 1 << 9
> + csrs mstatus, t0
> +
> + /* Route synchronous traps to trap_handler (mtvec direct
> mode). */
> + la t0, trap_handler
> + csrw mtvec, t0
> +
> + RUN_QOQ 1, 0xF220A25B, exp_qoq_u /*
> sf.vqmaccu.4x8x4 */
> + RUN_QOQ 2, 0xF620A25B, exp_qoq_s /*
> sf.vqmacc.4x8x4 */
> + RUN_QOQ 3, 0xFA20A25B, exp_qoq_us /*
> sf.vqmaccus.4x8x4 */
> + RUN_QOQ 4, 0xFE20A25B, exp_qoq_su /*
> sf.vqmaccsu.4x8x4 */
> + RUN_DOD 5, 0xB220A25B, exp_dod_u /*
> sf.vqmaccu.2x8x2 */
> + RUN_DOD 6, 0xB620A25B, exp_dod_s /*
> sf.vqmacc.2x8x2 */
> + RUN_DOD 7, 0xBA20A25B, exp_dod_us /*
> sf.vqmaccus.2x8x2 */
> + RUN_DOD 8, 0xBE20A25B, exp_dod_su /*
> sf.vqmaccsu.2x8x2 */
> +
> + /*
> + * Illegal-instruction reporting test (op 9).
> + *
> + * When vl is not a multiple of the tile size the
> instruction is
> + * reserved and the helper raises an illegal-instruction
> exception at
> + * runtime.
> + */
> + la t1, trap_mtval
> + sd zero, 0(t1)
> + la t1, trap_mcause
> + sd zero, 0(t1)
> + /* vl = 16 is not a multiple of the 4x8x4 tile size (32
> elements). */
> + li a2, 16
> + vsetvli t0, a2, e8, m1, ta, ma
> + .word 0xF220A25B /* sf.vqmaccu.4x8x4, traps
> */
> + /* The handler skips the faulting insn, so control resumes
> here. */
> + li a0, 9 /* fail code for this check
> */
> + la t1, trap_mcause
> + ld t2, 0(t1)
> + li t3, 2 /* RISCV_EXCP_ILLEGAL_INST
> */
> + bne t2, t3, _exit
> + la t1, trap_mtval
> + ld t2, 0(t1)
> + li t3, 0xF220A25B
> + bne t2, t3, _exit
> +
> + li a0, 0
> + j _exit
> +
> +/*
> + * Machine-mode trap handler: record mcause and mtval, then advance
> mepc
> + * past the 4-byte faulting instruction and return.
> + */
> + .balign 4
> +trap_handler:
> + csrr t4, mcause
> + la t5, trap_mcause
> + sd t4, 0(t5)
> + csrr t4, mtval
> + la t5, trap_mtval
> + sd t4, 0(t5)
> + csrr t4, mepc
> + addi t4, t4, 4
> + csrw mepc, t4
> + mret
> +
> +/*
> + * check_words(a0 = result, a1 = expected, a2 = word count).
> + * Returns to the caller if every word matches; otherwise exits with
> the
> + * operation index held in t6.
> + */
> +check_words:
> +1: beqz a2, 2f
> + lw t2, 0(a0)
> + lw t3, 0(a1)
> + bne t2, t3, 3f
> + addi a0, a0, 4
> + addi a1, a1, 4
> + addi a2, a2, -1
> + j 1b
> +2: ret
> +3: mv a0, t6
> + j _exit
> +
> +/* Exit through the semihosting SYS_EXIT_EXTENDED call with a0 as
> the code. */
> +_exit:
> + la a1, semiargs
> + li t0, 0x20026 /* ADP_Stopped_ApplicationExit */
> + sd t0, 0(a1)
> + sd a0, 8(a1)
> + li a0, 0x20 /* TARGET_SYS_EXIT_EXTENDED */
> + .balign 16
> + slli zero, zero, 0x1f
> + ebreak
> + srai zero, zero, 0x7
> + j .
> +
> + .data
> + .balign 8
> +semiargs:
> + .space 16
> +result:
> + .space 64
> +trap_mcause:
> + .space 8
> +trap_mtval:
> + .space 8
> +
> +/*
> + * A (4x8, row-major) in bytes [0, 32) of vs1.
> + */
> +a_qoq:
> + .byte 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x80
> + .byte 0xFF, 0x01, 0x00, 0x02, 0x7F, 0x03, 0x04, 0x05
> + .byte 0x10, 0x20, 0x30, 0x40, 0x50, 0x60, 0x70, 0x08
> + .byte 0x00, 0x00, 0x81, 0x00, 0x00, 0x02, 0x00, 0x03
> +/*
> + * B (8x4, row-major) in bytes [0, 32) of vs2.
> + */
> +b_qoq:
> + .byte 0x01, 0x02, 0x03, 0x04
> + .byte 0x05, 0x06, 0x07, 0x08
> + .byte 0x80, 0x7F, 0xFF, 0x01
> + .byte 0x00, 0x10, 0x20, 0x30
> + .byte 0x02, 0x02, 0x02, 0x02
> + .byte 0xFE, 0x01, 0x00, 0x7F
> + .byte 0x03, 0x00, 0x81, 0x04
> + .byte 0x40, 0x50, 0x60, 0x70
> +/*
> + * A (2x8, row-major) in bytes [0, 16) of vs1.
> + */
> +a_dod:
> + .byte 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x80
> + .byte 0xFF, 0x7F, 0x00, 0x02, 0x81, 0x03, 0x04, 0x05
> +/*
> + * B (8x2, row-major) in bytes [0, 16) of vs2.
> + */
> +b_dod:
> + .byte 0x01, 0x02
> + .byte 0x80, 0x7F
> + .byte 0x03, 0xFF
> + .byte 0x00, 0x10
> + .byte 0x02, 0xFE
> + .byte 0x7F, 0x01
> + .byte 0x81, 0x04
> + .byte 0x40, 0x50
> +
> + .balign 4
> +exp_qoq_u:
> + .word 0x0000279e, 0x000029db, 0x0000371f, 0x00003bf7
> + .word 0x00000648, 0x000004b5, 0x00000826, 0x0000091f
> + .word 0x00007be0, 0x00002030, 0x000074f0, 0x000042f0
> + .word 0x0000433c, 0x000040f1, 0x0000819f, 0x000002cf
> +exp_qoq_s:
> + .word 0xffffde9e, 0xffffd9db, 0xffffcd1f, 0xffffcbf7
> + .word 0x00000248, 0x000002b5, 0x00000126, 0x0000051f
> + .word 0xffffebe0, 0x00002030, 0xffffd4f0, 0x000042f0
> + .word 0x0000403c, 0xffffc1f1, 0x0000019f, 0x000001cf
> +exp_qoq_us:
> + .word 0x00001e9e, 0x000029db, 0x00002d1f, 0x00003bf7
> + .word 0x00000348, 0x000004b5, 0x00000426, 0x0000091f
> + .word 0xffffebe0, 0x00002030, 0xffffd4f0, 0x000042f0
> + .word 0xffffc03c, 0x000040f1, 0x0000009f, 0x000002cf
> +exp_qoq_su:
> + .word 0xffffe79e, 0xffffd9db, 0xffffd71f, 0xffffcbf7
> + .word 0x00000548, 0x000002b5, 0x00000526, 0x0000051f
> + .word 0x00007be0, 0x00002030, 0x000074f0, 0x000042f0
> + .word 0xffffc33c, 0xffffc1f1, 0xffff829f, 0x000001cf
> +exp_dod_u:
> + .word 0x00002795, 0x00003155, 0x00004642, 0x0000c2c0
> +exp_dod_s:
> + .word 0xffffde95, 0xffffd955, 0xffffc042, 0x000041c0
> +exp_dod_us:
> + .word 0x00001e95, 0x00002955, 0xffffc342, 0x000041c0
> +exp_dod_su:
> + .word 0xffffe795, 0xffffe155, 0x00004342, 0xffffc2c0
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-09-07 4:23 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-30 16:03 [PATCH v3 0/6] target/riscv: Add SiFive Xsfvqmaccdod/Xsfvqmaccqoq int8 matmul extensions Max Chou
2026-08-30 16:03 ` [PATCH v3 1/6] target/riscv: Add cfg property for SiFive " Max Chou
2026-08-30 16:03 ` [PATCH v3 2/6] target/riscv/cpu: List one implied-ext rule per line Max Chou
2026-08-30 16:03 ` [PATCH v3 3/6] target/riscv/cpu: Add SiFive custom int8 matmul extension implied rules Max Chou
2026-08-30 16:03 ` [PATCH v3 4/6] target/riscv: rvv: Add SiFive custom int8 matmul instructions Max Chou
2026-08-30 16:03 ` [PATCH v3 5/6] disas/riscv: Add disassembler support for Xsfvqmaccdod/Xsfvqmaccqoq Max Chou
2026-09-07 4:12 ` Alistair Francis
2026-08-30 16:03 ` [PATCH v3 6/6] tests/tcg/riscv64: Add tests for SiFive int8 matmul extensions Max Chou
2026-09-07 4:22 ` Alistair Francis
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.