* [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB
@ 2025-08-09 23:42 Richard Henderson
2025-08-09 23:42 ` [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension Richard Henderson
` (3 more replies)
0 siblings, 4 replies; 7+ messages in thread
From: Richard Henderson @ 2025-08-09 23:42 UTC (permalink / raw)
To: qemu-devel
x86 doesn't directly support 8-bit vector shifts, so we have
some 2 to 5 insn expansions. With VGF2P8AFFINEQB, we can do
it in 1 insn, plus a (possibly shared) constant load.
r~
Richard Henderson (3):
cpuinfo/i386: Detect GFNI as an AVX extension
tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec
tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts
host/include/i386/host/cpuinfo.h | 1 +
include/qemu/cpuid.h | 3 ++
util/cpuinfo-i386.c | 1 +
tcg/i386/tcg-target-opc.h.inc | 1 +
tcg/i386/tcg-target.c.inc | 81 ++++++++++++++++++++++++++++++--
5 files changed, 83 insertions(+), 4 deletions(-)
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread* [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension 2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson @ 2025-08-09 23:42 ` Richard Henderson 2025-08-09 23:42 ` [PATCH 2/3] tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec Richard Henderson ` (2 subsequent siblings) 3 siblings, 0 replies; 7+ messages in thread From: Richard Henderson @ 2025-08-09 23:42 UTC (permalink / raw) To: qemu-devel We won't use the SSE GFNI instructions, so delay detection until we know AVX is present. Signed-off-by: Richard Henderson <richard.henderson@linaro.org> --- host/include/i386/host/cpuinfo.h | 1 + include/qemu/cpuid.h | 3 +++ util/cpuinfo-i386.c | 1 + 3 files changed, 5 insertions(+) diff --git a/host/include/i386/host/cpuinfo.h b/host/include/i386/host/cpuinfo.h index 9541a64da6..93d029d499 100644 --- a/host/include/i386/host/cpuinfo.h +++ b/host/include/i386/host/cpuinfo.h @@ -27,6 +27,7 @@ #define CPUINFO_ATOMIC_VMOVDQU (1u << 17) #define CPUINFO_AES (1u << 18) #define CPUINFO_PCLMUL (1u << 19) +#define CPUINFO_GFNI (1u << 20) /* Initialized with a constructor. */ extern unsigned cpuinfo; diff --git a/include/qemu/cpuid.h b/include/qemu/cpuid.h index b11161555b..f8351b80b0 100644 --- a/include/qemu/cpuid.h +++ b/include/qemu/cpuid.h @@ -68,6 +68,9 @@ #ifndef bit_AVX512VBMI2 #define bit_AVX512VBMI2 (1 << 6) #endif +#ifndef bit_GNFI +#define bit_GNFI (1 << 8) +#endif /* Leaf 0x80000001, %ecx */ #ifndef bit_LZCNT diff --git a/util/cpuinfo-i386.c b/util/cpuinfo-i386.c index c8c8a1b370..f4c5b6ff40 100644 --- a/util/cpuinfo-i386.c +++ b/util/cpuinfo-i386.c @@ -50,6 +50,7 @@ unsigned __attribute__((constructor)) cpuinfo_init(void) if ((bv & 6) == 6) { info |= CPUINFO_AVX1; info |= (b7 & bit_AVX2 ? CPUINFO_AVX2 : 0); + info |= (c7 & bit_GFNI ? CPUINFO_GFNI : 0); if ((bv & 0xe0) == 0xe0) { info |= (b7 & bit_AVX512F ? CPUINFO_AVX512F : 0); -- 2.43.0 ^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH 2/3] tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec 2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson 2025-08-09 23:42 ` [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension Richard Henderson @ 2025-08-09 23:42 ` Richard Henderson 2025-08-09 23:42 ` [PATCH 3/3] tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts Richard Henderson 2025-08-27 7:26 ` [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson 3 siblings, 0 replies; 7+ messages in thread From: Richard Henderson @ 2025-08-09 23:42 UTC (permalink / raw) To: qemu-devel Add a backend-specific opcode for expanding the GFNI vgf2p8affineqb instruction, which we can use for expanding 8-bit immediate shifts and rotates. Signed-off-by: Richard Henderson <richard.henderson@linaro.org> --- tcg/i386/tcg-target-opc.h.inc | 1 + tcg/i386/tcg-target.c.inc | 6 ++++++ 2 files changed, 7 insertions(+) diff --git a/tcg/i386/tcg-target-opc.h.inc b/tcg/i386/tcg-target-opc.h.inc index 8cc0dbaeaf..8a5cb34dbe 100644 --- a/tcg/i386/tcg-target-opc.h.inc +++ b/tcg/i386/tcg-target-opc.h.inc @@ -35,3 +35,4 @@ DEF(x86_punpckh_vec, 1, 2, 0, TCG_OPF_VECTOR) DEF(x86_vpshldi_vec, 1, 2, 1, TCG_OPF_VECTOR) DEF(x86_vpshldv_vec, 1, 3, 0, TCG_OPF_VECTOR) DEF(x86_vpshrdv_vec, 1, 3, 0, TCG_OPF_VECTOR) +DEF(x86_vgf2p8affineqb_vec, 1, 2, 1, TCG_OPF_VECTOR) diff --git a/tcg/i386/tcg-target.c.inc b/tcg/i386/tcg-target.c.inc index 088c6c9264..9dd588fc41 100644 --- a/tcg/i386/tcg-target.c.inc +++ b/tcg/i386/tcg-target.c.inc @@ -451,6 +451,7 @@ static bool tcg_target_const_match(int64_t val, int ct, #define OPC_VPBROADCASTW (0x79 | P_EXT38 | P_DATA16) #define OPC_VPBROADCASTD (0x58 | P_EXT38 | P_DATA16) #define OPC_VPBROADCASTQ (0x59 | P_EXT38 | P_DATA16) +#define OPC_VGF2P8AFFINEQB (0xce | P_EXT3A | P_DATA16 | P_VEXW) #define OPC_VPMOVM2B (0x28 | P_EXT38 | P_SIMDF3 | P_EVEX) #define OPC_VPMOVM2W (0x28 | P_EXT38 | P_SIMDF3 | P_VEXW | P_EVEX) #define OPC_VPMOVM2D (0x38 | P_EXT38 | P_SIMDF3 | P_EVEX) @@ -4084,6 +4085,10 @@ static void tcg_out_vec_op(TCGContext *s, TCGOpcode opc, insn = vpshldi_insn[vece]; sub = args[3]; goto gen_simd_imm8; + case INDEX_op_x86_vgf2p8affineqb_vec: + insn = OPC_VGF2P8AFFINEQB; + sub = args[3]; + goto gen_simd_imm8; case INDEX_op_not_vec: insn = OPC_VPTERNLOGQ; @@ -4188,6 +4193,7 @@ tcg_target_op_def(TCGOpcode op, TCGType type, unsigned flags) case INDEX_op_x86_punpckl_vec: case INDEX_op_x86_punpckh_vec: case INDEX_op_x86_vpshldi_vec: + case INDEX_op_x86_vgf2p8affineqb_vec: #if TCG_TARGET_REG_BITS == 32 case INDEX_op_dup2_vec: #endif -- 2.43.0 ^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH 3/3] tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts 2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson 2025-08-09 23:42 ` [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension Richard Henderson 2025-08-09 23:42 ` [PATCH 2/3] tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec Richard Henderson @ 2025-08-09 23:42 ` Richard Henderson 2025-08-27 7:26 ` [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson 3 siblings, 0 replies; 7+ messages in thread From: Richard Henderson @ 2025-08-09 23:42 UTC (permalink / raw) To: qemu-devel A constant matrix can describe the movement of the 8 bits, so these shifts can be performed with one instruction. Logic courtesy of Andi Kleen <ak@linux.intel.com>: https://gcc.gnu.org/pipermail/gcc-patches/2025-August/691624.html Signed-off-by: Richard Henderson <richard.henderson@linaro.org> --- tcg/i386/tcg-target.c.inc | 75 ++++++++++++++++++++++++++++++++++++--- 1 file changed, 71 insertions(+), 4 deletions(-) diff --git a/tcg/i386/tcg-target.c.inc b/tcg/i386/tcg-target.c.inc index 9dd588fc41..fb76724941 100644 --- a/tcg/i386/tcg-target.c.inc +++ b/tcg/i386/tcg-target.c.inc @@ -4342,12 +4342,46 @@ int tcg_can_emit_vec_op(TCGOpcode opc, TCGType type, unsigned vece) } } +static void gen_vgf2p8affineqb0(TCGType type, TCGv_vec v0, + TCGv_vec v1, uint64_t matrix) +{ + vec_gen_4(INDEX_op_x86_vgf2p8affineqb_vec, type, MO_8, + tcgv_vec_arg(v0), tcgv_vec_arg(v1), + tcgv_vec_arg(tcg_constant_vec(type, MO_64, matrix)), 0); +} + static void expand_vec_shi(TCGType type, unsigned vece, bool right, TCGv_vec v0, TCGv_vec v1, TCGArg imm) { + static const uint64_t gf2_shi[2][8] = { + /* left shift */ + { 0, + 0x0001020408102040ull, + 0x0000010204081020ull, + 0x0000000102040810ull, + 0x0000000001020408ull, + 0x0000000000010204ull, + 0x0000000000000102ull, + 0x0000000000000001ull }, + /* right shift */ + { 0, + 0x0204081020408000ull, + 0x0408102040800000ull, + 0x0810204080000000ull, + 0x1020408000000000ull, + 0x2040800000000000ull, + 0x4080000000000000ull, + 0x8000000000000000ull } + }; uint8_t mask; tcg_debug_assert(vece == MO_8); + + if (cpuinfo & CPUINFO_GFNI) { + gen_vgf2p8affineqb0(type, v0, v1, gf2_shi[right][imm]); + return; + } + if (right) { mask = 0xff >> imm; tcg_gen_shri_vec(MO_16, v0, v1, imm); @@ -4361,10 +4395,25 @@ static void expand_vec_shi(TCGType type, unsigned vece, bool right, static void expand_vec_sari(TCGType type, unsigned vece, TCGv_vec v0, TCGv_vec v1, TCGArg imm) { + static const uint64_t gf2_sar[8] = { + 0, + 0x0204081020408080ull, + 0x0408102040808080ull, + 0x0810204080808080ull, + 0x1020408080808080ull, + 0x2040808080808080ull, + 0x4080808080808080ull, + 0x8080808080808080ull, + }; TCGv_vec t1, t2; switch (vece) { case MO_8: + if (cpuinfo & CPUINFO_GFNI) { + gen_vgf2p8affineqb0(type, v0, v1, gf2_sar[imm]); + break; + } + /* Unpack to 16-bit, shift, and repack. */ t1 = tcg_temp_new_vec(type); t2 = tcg_temp_new_vec(type); @@ -4416,12 +4465,30 @@ static void expand_vec_sari(TCGType type, unsigned vece, static void expand_vec_rotli(TCGType type, unsigned vece, TCGv_vec v0, TCGv_vec v1, TCGArg imm) { + static const uint64_t gf2_rol[8] = { + 0, + 0x8001020408102040ull, + 0x4080010204081020ull, + 0x2040800102040810ull, + 0x1020408001020408ull, + 0x0810204080010204ull, + 0x0408102040800102ull, + 0x0204081020408001ull, + }; TCGv_vec t; - if (vece != MO_8 && have_avx512vbmi2) { - vec_gen_4(INDEX_op_x86_vpshldi_vec, type, vece, - tcgv_vec_arg(v0), tcgv_vec_arg(v1), tcgv_vec_arg(v1), imm); - return; + if (vece == MO_8) { + if (cpuinfo & CPUINFO_GFNI) { + gen_vgf2p8affineqb0(type, v0, v1, gf2_rol[imm]); + return; + } + } else { + if (have_avx512vbmi2) { + vec_gen_4(INDEX_op_x86_vpshldi_vec, type, vece, + tcgv_vec_arg(v0), tcgv_vec_arg(v1), + tcgv_vec_arg(v1), imm); + return; + } } t = tcg_temp_new_vec(type); -- 2.43.0 ^ permalink raw reply related [flat|nested] 7+ messages in thread
* Re: [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB 2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson ` (2 preceding siblings ...) 2025-08-09 23:42 ` [PATCH 3/3] tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts Richard Henderson @ 2025-08-27 7:26 ` Richard Henderson 2025-08-27 8:12 ` Paolo Bonzini 3 siblings, 1 reply; 7+ messages in thread From: Richard Henderson @ 2025-08-27 7:26 UTC (permalink / raw) To: qemu-devel On 8/10/25 09:42, Richard Henderson wrote: > x86 doesn't directly support 8-bit vector shifts, so we have > some 2 to 5 insn expansions. With VGF2P8AFFINEQB, we can do > it in 1 insn, plus a (possibly shared) constant load. > > > r~ > > > Richard Henderson (3): > cpuinfo/i386: Detect GFNI as an AVX extension > tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec > tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts > > host/include/i386/host/cpuinfo.h | 1 + > include/qemu/cpuid.h | 3 ++ > util/cpuinfo-i386.c | 1 + > tcg/i386/tcg-target-opc.h.inc | 1 + > tcg/i386/tcg-target.c.inc | 81 ++++++++++++++++++++++++++++++-- > 5 files changed, 83 insertions(+), 4 deletions(-) > Ping. r~ ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB 2025-08-27 7:26 ` [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson @ 2025-08-27 8:12 ` Paolo Bonzini 2025-08-27 9:34 ` Richard Henderson 0 siblings, 1 reply; 7+ messages in thread From: Paolo Bonzini @ 2025-08-27 8:12 UTC (permalink / raw) To: Richard Henderson, qemu-devel On 8/27/25 09:26, Richard Henderson wrote: > On 8/10/25 09:42, Richard Henderson wrote: >> x86 doesn't directly support 8-bit vector shifts, so we have >> some 2 to 5 insn expansions. With VGF2P8AFFINEQB, we can do >> it in 1 insn, plus a (possibly shared) constant load. >> >> >> r~ >> >> >> Richard Henderson (3): >> cpuinfo/i386: Detect GFNI as an AVX extension >> tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec >> tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts >> >> host/include/i386/host/cpuinfo.h | 1 + >> include/qemu/cpuid.h | 3 ++ >> util/cpuinfo-i386.c | 1 + >> tcg/i386/tcg-target-opc.h.inc | 1 + >> tcg/i386/tcg-target.c.inc | 81 ++++++++++++++++++++++++++++++-- >> 5 files changed, 83 insertions(+), 4 deletions(-) >> > > Ping. I don't know the target-independent part of TCG, but arithmetic right shift by 7 probably should keep using pcmpgtb? There's also a typo in patch 1 (s/NF/FN/): +#ifndef bit_GNFI +#define bit_GNFI (1 << 8) +#endif Paolo ^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB 2025-08-27 8:12 ` Paolo Bonzini @ 2025-08-27 9:34 ` Richard Henderson 0 siblings, 0 replies; 7+ messages in thread From: Richard Henderson @ 2025-08-27 9:34 UTC (permalink / raw) To: Paolo Bonzini, qemu-devel On 8/27/25 18:12, Paolo Bonzini wrote: > On 8/27/25 09:26, Richard Henderson wrote: >> On 8/10/25 09:42, Richard Henderson wrote: >>> x86 doesn't directly support 8-bit vector shifts, so we have >>> some 2 to 5 insn expansions. With VGF2P8AFFINEQB, we can do >>> it in 1 insn, plus a (possibly shared) constant load. >>> >>> >>> r~ >>> >>> >>> Richard Henderson (3): >>> cpuinfo/i386: Detect GFNI as an AVX extension >>> tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec >>> tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts >>> >>> host/include/i386/host/cpuinfo.h | 1 + >>> include/qemu/cpuid.h | 3 ++ >>> util/cpuinfo-i386.c | 1 + >>> tcg/i386/tcg-target-opc.h.inc | 1 + >>> tcg/i386/tcg-target.c.inc | 81 ++++++++++++++++++++++++++++++-- >>> 5 files changed, 83 insertions(+), 4 deletions(-) >>> >> >> Ping. > > I don't know the target-independent part of TCG, but arithmetic right shift by 7 probably > should keep using pcmpgtb? There is no such target independent expansion. But it's a good idea to add to x86, at least for the two sizes that don't directly support such shift. r~ ^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2025-08-27 9:36 UTC | newest] Thread overview: 7+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson 2025-08-09 23:42 ` [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension Richard Henderson 2025-08-09 23:42 ` [PATCH 2/3] tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec Richard Henderson 2025-08-09 23:42 ` [PATCH 3/3] tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts Richard Henderson 2025-08-27 7:26 ` [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson 2025-08-27 8:12 ` Paolo Bonzini 2025-08-27 9:34 ` Richard Henderson
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.