* [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB
@ 2025-08-09 23:42 Richard Henderson
2025-08-09 23:42 ` [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension Richard Henderson
` (3 more replies)
0 siblings, 4 replies; 7+ messages in thread
From: Richard Henderson @ 2025-08-09 23:42 UTC (permalink / raw)
To: qemu-devel
x86 doesn't directly support 8-bit vector shifts, so we have
some 2 to 5 insn expansions. With VGF2P8AFFINEQB, we can do
it in 1 insn, plus a (possibly shared) constant load.
r~
Richard Henderson (3):
cpuinfo/i386: Detect GFNI as an AVX extension
tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec
tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts
host/include/i386/host/cpuinfo.h | 1 +
include/qemu/cpuid.h | 3 ++
util/cpuinfo-i386.c | 1 +
tcg/i386/tcg-target-opc.h.inc | 1 +
tcg/i386/tcg-target.c.inc | 81 ++++++++++++++++++++++++++++++--
5 files changed, 83 insertions(+), 4 deletions(-)
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension
2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
@ 2025-08-09 23:42 ` Richard Henderson
2025-08-09 23:42 ` [PATCH 2/3] tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec Richard Henderson
` (2 subsequent siblings)
3 siblings, 0 replies; 7+ messages in thread
From: Richard Henderson @ 2025-08-09 23:42 UTC (permalink / raw)
To: qemu-devel
We won't use the SSE GFNI instructions, so delay
detection until we know AVX is present.
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
---
host/include/i386/host/cpuinfo.h | 1 +
include/qemu/cpuid.h | 3 +++
util/cpuinfo-i386.c | 1 +
3 files changed, 5 insertions(+)
diff --git a/host/include/i386/host/cpuinfo.h b/host/include/i386/host/cpuinfo.h
index 9541a64da6..93d029d499 100644
--- a/host/include/i386/host/cpuinfo.h
+++ b/host/include/i386/host/cpuinfo.h
@@ -27,6 +27,7 @@
#define CPUINFO_ATOMIC_VMOVDQU (1u << 17)
#define CPUINFO_AES (1u << 18)
#define CPUINFO_PCLMUL (1u << 19)
+#define CPUINFO_GFNI (1u << 20)
/* Initialized with a constructor. */
extern unsigned cpuinfo;
diff --git a/include/qemu/cpuid.h b/include/qemu/cpuid.h
index b11161555b..f8351b80b0 100644
--- a/include/qemu/cpuid.h
+++ b/include/qemu/cpuid.h
@@ -68,6 +68,9 @@
#ifndef bit_AVX512VBMI2
#define bit_AVX512VBMI2 (1 << 6)
#endif
+#ifndef bit_GNFI
+#define bit_GNFI (1 << 8)
+#endif
/* Leaf 0x80000001, %ecx */
#ifndef bit_LZCNT
diff --git a/util/cpuinfo-i386.c b/util/cpuinfo-i386.c
index c8c8a1b370..f4c5b6ff40 100644
--- a/util/cpuinfo-i386.c
+++ b/util/cpuinfo-i386.c
@@ -50,6 +50,7 @@ unsigned __attribute__((constructor)) cpuinfo_init(void)
if ((bv & 6) == 6) {
info |= CPUINFO_AVX1;
info |= (b7 & bit_AVX2 ? CPUINFO_AVX2 : 0);
+ info |= (c7 & bit_GFNI ? CPUINFO_GFNI : 0);
if ((bv & 0xe0) == 0xe0) {
info |= (b7 & bit_AVX512F ? CPUINFO_AVX512F : 0);
--
2.43.0
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH 2/3] tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec
2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
2025-08-09 23:42 ` [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension Richard Henderson
@ 2025-08-09 23:42 ` Richard Henderson
2025-08-09 23:42 ` [PATCH 3/3] tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts Richard Henderson
2025-08-27 7:26 ` [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
3 siblings, 0 replies; 7+ messages in thread
From: Richard Henderson @ 2025-08-09 23:42 UTC (permalink / raw)
To: qemu-devel
Add a backend-specific opcode for expanding the
GFNI vgf2p8affineqb instruction, which we can use
for expanding 8-bit immediate shifts and rotates.
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
---
tcg/i386/tcg-target-opc.h.inc | 1 +
tcg/i386/tcg-target.c.inc | 6 ++++++
2 files changed, 7 insertions(+)
diff --git a/tcg/i386/tcg-target-opc.h.inc b/tcg/i386/tcg-target-opc.h.inc
index 8cc0dbaeaf..8a5cb34dbe 100644
--- a/tcg/i386/tcg-target-opc.h.inc
+++ b/tcg/i386/tcg-target-opc.h.inc
@@ -35,3 +35,4 @@ DEF(x86_punpckh_vec, 1, 2, 0, TCG_OPF_VECTOR)
DEF(x86_vpshldi_vec, 1, 2, 1, TCG_OPF_VECTOR)
DEF(x86_vpshldv_vec, 1, 3, 0, TCG_OPF_VECTOR)
DEF(x86_vpshrdv_vec, 1, 3, 0, TCG_OPF_VECTOR)
+DEF(x86_vgf2p8affineqb_vec, 1, 2, 1, TCG_OPF_VECTOR)
diff --git a/tcg/i386/tcg-target.c.inc b/tcg/i386/tcg-target.c.inc
index 088c6c9264..9dd588fc41 100644
--- a/tcg/i386/tcg-target.c.inc
+++ b/tcg/i386/tcg-target.c.inc
@@ -451,6 +451,7 @@ static bool tcg_target_const_match(int64_t val, int ct,
#define OPC_VPBROADCASTW (0x79 | P_EXT38 | P_DATA16)
#define OPC_VPBROADCASTD (0x58 | P_EXT38 | P_DATA16)
#define OPC_VPBROADCASTQ (0x59 | P_EXT38 | P_DATA16)
+#define OPC_VGF2P8AFFINEQB (0xce | P_EXT3A | P_DATA16 | P_VEXW)
#define OPC_VPMOVM2B (0x28 | P_EXT38 | P_SIMDF3 | P_EVEX)
#define OPC_VPMOVM2W (0x28 | P_EXT38 | P_SIMDF3 | P_VEXW | P_EVEX)
#define OPC_VPMOVM2D (0x38 | P_EXT38 | P_SIMDF3 | P_EVEX)
@@ -4084,6 +4085,10 @@ static void tcg_out_vec_op(TCGContext *s, TCGOpcode opc,
insn = vpshldi_insn[vece];
sub = args[3];
goto gen_simd_imm8;
+ case INDEX_op_x86_vgf2p8affineqb_vec:
+ insn = OPC_VGF2P8AFFINEQB;
+ sub = args[3];
+ goto gen_simd_imm8;
case INDEX_op_not_vec:
insn = OPC_VPTERNLOGQ;
@@ -4188,6 +4193,7 @@ tcg_target_op_def(TCGOpcode op, TCGType type, unsigned flags)
case INDEX_op_x86_punpckl_vec:
case INDEX_op_x86_punpckh_vec:
case INDEX_op_x86_vpshldi_vec:
+ case INDEX_op_x86_vgf2p8affineqb_vec:
#if TCG_TARGET_REG_BITS == 32
case INDEX_op_dup2_vec:
#endif
--
2.43.0
^ permalink raw reply related [flat|nested] 7+ messages in thread
* [PATCH 3/3] tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts
2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
2025-08-09 23:42 ` [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension Richard Henderson
2025-08-09 23:42 ` [PATCH 2/3] tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec Richard Henderson
@ 2025-08-09 23:42 ` Richard Henderson
2025-08-27 7:26 ` [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
3 siblings, 0 replies; 7+ messages in thread
From: Richard Henderson @ 2025-08-09 23:42 UTC (permalink / raw)
To: qemu-devel
A constant matrix can describe the movement of the 8 bits,
so these shifts can be performed with one instruction.
Logic courtesy of Andi Kleen <ak@linux.intel.com>:
https://gcc.gnu.org/pipermail/gcc-patches/2025-August/691624.html
Signed-off-by: Richard Henderson <richard.henderson@linaro.org>
---
tcg/i386/tcg-target.c.inc | 75 ++++++++++++++++++++++++++++++++++++---
1 file changed, 71 insertions(+), 4 deletions(-)
diff --git a/tcg/i386/tcg-target.c.inc b/tcg/i386/tcg-target.c.inc
index 9dd588fc41..fb76724941 100644
--- a/tcg/i386/tcg-target.c.inc
+++ b/tcg/i386/tcg-target.c.inc
@@ -4342,12 +4342,46 @@ int tcg_can_emit_vec_op(TCGOpcode opc, TCGType type, unsigned vece)
}
}
+static void gen_vgf2p8affineqb0(TCGType type, TCGv_vec v0,
+ TCGv_vec v1, uint64_t matrix)
+{
+ vec_gen_4(INDEX_op_x86_vgf2p8affineqb_vec, type, MO_8,
+ tcgv_vec_arg(v0), tcgv_vec_arg(v1),
+ tcgv_vec_arg(tcg_constant_vec(type, MO_64, matrix)), 0);
+}
+
static void expand_vec_shi(TCGType type, unsigned vece, bool right,
TCGv_vec v0, TCGv_vec v1, TCGArg imm)
{
+ static const uint64_t gf2_shi[2][8] = {
+ /* left shift */
+ { 0,
+ 0x0001020408102040ull,
+ 0x0000010204081020ull,
+ 0x0000000102040810ull,
+ 0x0000000001020408ull,
+ 0x0000000000010204ull,
+ 0x0000000000000102ull,
+ 0x0000000000000001ull },
+ /* right shift */
+ { 0,
+ 0x0204081020408000ull,
+ 0x0408102040800000ull,
+ 0x0810204080000000ull,
+ 0x1020408000000000ull,
+ 0x2040800000000000ull,
+ 0x4080000000000000ull,
+ 0x8000000000000000ull }
+ };
uint8_t mask;
tcg_debug_assert(vece == MO_8);
+
+ if (cpuinfo & CPUINFO_GFNI) {
+ gen_vgf2p8affineqb0(type, v0, v1, gf2_shi[right][imm]);
+ return;
+ }
+
if (right) {
mask = 0xff >> imm;
tcg_gen_shri_vec(MO_16, v0, v1, imm);
@@ -4361,10 +4395,25 @@ static void expand_vec_shi(TCGType type, unsigned vece, bool right,
static void expand_vec_sari(TCGType type, unsigned vece,
TCGv_vec v0, TCGv_vec v1, TCGArg imm)
{
+ static const uint64_t gf2_sar[8] = {
+ 0,
+ 0x0204081020408080ull,
+ 0x0408102040808080ull,
+ 0x0810204080808080ull,
+ 0x1020408080808080ull,
+ 0x2040808080808080ull,
+ 0x4080808080808080ull,
+ 0x8080808080808080ull,
+ };
TCGv_vec t1, t2;
switch (vece) {
case MO_8:
+ if (cpuinfo & CPUINFO_GFNI) {
+ gen_vgf2p8affineqb0(type, v0, v1, gf2_sar[imm]);
+ break;
+ }
+
/* Unpack to 16-bit, shift, and repack. */
t1 = tcg_temp_new_vec(type);
t2 = tcg_temp_new_vec(type);
@@ -4416,12 +4465,30 @@ static void expand_vec_sari(TCGType type, unsigned vece,
static void expand_vec_rotli(TCGType type, unsigned vece,
TCGv_vec v0, TCGv_vec v1, TCGArg imm)
{
+ static const uint64_t gf2_rol[8] = {
+ 0,
+ 0x8001020408102040ull,
+ 0x4080010204081020ull,
+ 0x2040800102040810ull,
+ 0x1020408001020408ull,
+ 0x0810204080010204ull,
+ 0x0408102040800102ull,
+ 0x0204081020408001ull,
+ };
TCGv_vec t;
- if (vece != MO_8 && have_avx512vbmi2) {
- vec_gen_4(INDEX_op_x86_vpshldi_vec, type, vece,
- tcgv_vec_arg(v0), tcgv_vec_arg(v1), tcgv_vec_arg(v1), imm);
- return;
+ if (vece == MO_8) {
+ if (cpuinfo & CPUINFO_GFNI) {
+ gen_vgf2p8affineqb0(type, v0, v1, gf2_rol[imm]);
+ return;
+ }
+ } else {
+ if (have_avx512vbmi2) {
+ vec_gen_4(INDEX_op_x86_vpshldi_vec, type, vece,
+ tcgv_vec_arg(v0), tcgv_vec_arg(v1),
+ tcgv_vec_arg(v1), imm);
+ return;
+ }
}
t = tcg_temp_new_vec(type);
--
2.43.0
^ permalink raw reply related [flat|nested] 7+ messages in thread
* Re: [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB
2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
` (2 preceding siblings ...)
2025-08-09 23:42 ` [PATCH 3/3] tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts Richard Henderson
@ 2025-08-27 7:26 ` Richard Henderson
2025-08-27 8:12 ` Paolo Bonzini
3 siblings, 1 reply; 7+ messages in thread
From: Richard Henderson @ 2025-08-27 7:26 UTC (permalink / raw)
To: qemu-devel
On 8/10/25 09:42, Richard Henderson wrote:
> x86 doesn't directly support 8-bit vector shifts, so we have
> some 2 to 5 insn expansions. With VGF2P8AFFINEQB, we can do
> it in 1 insn, plus a (possibly shared) constant load.
>
>
> r~
>
>
> Richard Henderson (3):
> cpuinfo/i386: Detect GFNI as an AVX extension
> tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec
> tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts
>
> host/include/i386/host/cpuinfo.h | 1 +
> include/qemu/cpuid.h | 3 ++
> util/cpuinfo-i386.c | 1 +
> tcg/i386/tcg-target-opc.h.inc | 1 +
> tcg/i386/tcg-target.c.inc | 81 ++++++++++++++++++++++++++++++--
> 5 files changed, 83 insertions(+), 4 deletions(-)
>
Ping.
r~
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB
2025-08-27 7:26 ` [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
@ 2025-08-27 8:12 ` Paolo Bonzini
2025-08-27 9:34 ` Richard Henderson
0 siblings, 1 reply; 7+ messages in thread
From: Paolo Bonzini @ 2025-08-27 8:12 UTC (permalink / raw)
To: Richard Henderson, qemu-devel
On 8/27/25 09:26, Richard Henderson wrote:
> On 8/10/25 09:42, Richard Henderson wrote:
>> x86 doesn't directly support 8-bit vector shifts, so we have
>> some 2 to 5 insn expansions. With VGF2P8AFFINEQB, we can do
>> it in 1 insn, plus a (possibly shared) constant load.
>>
>>
>> r~
>>
>>
>> Richard Henderson (3):
>> cpuinfo/i386: Detect GFNI as an AVX extension
>> tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec
>> tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts
>>
>> host/include/i386/host/cpuinfo.h | 1 +
>> include/qemu/cpuid.h | 3 ++
>> util/cpuinfo-i386.c | 1 +
>> tcg/i386/tcg-target-opc.h.inc | 1 +
>> tcg/i386/tcg-target.c.inc | 81 ++++++++++++++++++++++++++++++--
>> 5 files changed, 83 insertions(+), 4 deletions(-)
>>
>
> Ping.
I don't know the target-independent part of TCG, but arithmetic right
shift by 7 probably should keep using pcmpgtb?
There's also a typo in patch 1 (s/NF/FN/):
+#ifndef bit_GNFI
+#define bit_GNFI (1 << 8)
+#endif
Paolo
^ permalink raw reply [flat|nested] 7+ messages in thread
* Re: [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB
2025-08-27 8:12 ` Paolo Bonzini
@ 2025-08-27 9:34 ` Richard Henderson
0 siblings, 0 replies; 7+ messages in thread
From: Richard Henderson @ 2025-08-27 9:34 UTC (permalink / raw)
To: Paolo Bonzini, qemu-devel
On 8/27/25 18:12, Paolo Bonzini wrote:
> On 8/27/25 09:26, Richard Henderson wrote:
>> On 8/10/25 09:42, Richard Henderson wrote:
>>> x86 doesn't directly support 8-bit vector shifts, so we have
>>> some 2 to 5 insn expansions. With VGF2P8AFFINEQB, we can do
>>> it in 1 insn, plus a (possibly shared) constant load.
>>>
>>>
>>> r~
>>>
>>>
>>> Richard Henderson (3):
>>> cpuinfo/i386: Detect GFNI as an AVX extension
>>> tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec
>>> tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts
>>>
>>> host/include/i386/host/cpuinfo.h | 1 +
>>> include/qemu/cpuid.h | 3 ++
>>> util/cpuinfo-i386.c | 1 +
>>> tcg/i386/tcg-target-opc.h.inc | 1 +
>>> tcg/i386/tcg-target.c.inc | 81 ++++++++++++++++++++++++++++++--
>>> 5 files changed, 83 insertions(+), 4 deletions(-)
>>>
>>
>> Ping.
>
> I don't know the target-independent part of TCG, but arithmetic right shift by 7 probably
> should keep using pcmpgtb?
There is no such target independent expansion. But it's a good idea to add to x86, at
least for the two sizes that don't directly support such shift.
r~
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2025-08-27 9:36 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2025-08-09 23:42 [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
2025-08-09 23:42 ` [PATCH 1/3] cpuinfo/i386: Detect GFNI as an AVX extension Richard Henderson
2025-08-09 23:42 ` [PATCH 2/3] tcg/i386: Add INDEX_op_x86_vgf2p8affineqb_vec Richard Henderson
2025-08-09 23:42 ` [PATCH 3/3] tcg/i386: Use vgf2p8affineqb for MO_8 vector shifts Richard Henderson
2025-08-27 7:26 ` [PATCH 0/3] tcg/i386: Improve 8-bit shifts with VGF2P8AFFINEQB Richard Henderson
2025-08-27 8:12 ` Paolo Bonzini
2025-08-27 9:34 ` Richard Henderson
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.