> -----Original Message-----
> From: Jakub Jelinek <[email protected]>
> Sent: Sunday, September 20, 2026 2:53 AM
> To: Uros Bizjak <[email protected]>; Hongtao Liu <[email protected]>; Liu,
> Hongtao <[email protected]>
> Cc: [email protected]; Andi Kleen <[email protected]>
> Subject: [PATCH] i386: Only support GFNI V32QImode rotates with AVX2, not
> just AVX [PR127498]
>
> Hi!
>
> The following testcase ICEs, because the middle-end doesn't like a target to
> announce rot{l,r}v32qi3 optab availability only with const_int_operand rotate
> with non-working fallback.
> If the rotate optab fails because of unsatisfiable predicates, the middle-end
> will try to expand the rotate as bitwise or of 2 shifts, but the problem with
> TARGET_GFNI && TARGET_AVX && !TARGET_AVX2 is that the V32QImode
> shifts aren't supported either, one needs AVX2 for that.
> I believe there are no CPUs with GFNI and AVX but without AVX2, I think there
> are just CPUs with both GFNI and AVX/AVX2 or with GFNI and no AVX, so the
> following patch just disables the pattern for the TARGET_GFNI &&
> TARGET_AVX combination and just requires TARGET_GFNI && TARGET_AVX2
> instead. It is true that rotate can be expanded with constant rotate count
> even for V32QImode, but at the expense of ICEing for non-constant rotate
> count.
> The alternative would be to extend the pattern to provide a fallback for
> TARGET_GFNI && TARGET_AVX && !TARGET_AVX2 when the last operand is
> not const_int_operand, it can be handled as 4 V16QImode shifts, combining
> stuff together, ...
>
> Bootstrapped/regtested on x86_64-linux and i686-linux, ok for trunk?
Ok.
>
> 2026-09-19 Jakub Jelinek <[email protected]>
>
> PR target/127498
> * config/i386/sse.md (<insn><mode>3
> any_rotate:VI1_AVX512_3264): Add
> TARGET_AVX2 to condition.
>
> * gcc.target/i386/gfni-pr127498.c: New test.
> * gcc.target/i386/shift-gf2p8affine-5.c: Use -mavx2 in dg-options
> instead of -mavx. Expect 46 insns instead of 31.
>
> --- a/gcc/config/i386/sse.md 2026-09-15 12:01:51.693090054 +0200
> +++ b/gcc/config/i386/sse.md 2026-09-19 09:52:47.935239962 +0200
> @@ -28458,7 +28458,7 @@ (define_expand "<insn><mode>3"
> (any_rotate:VI1_AVX512_3264
> (match_operand:VI1_AVX512_3264 1 "register_operand")
> (match_operand:SI 2 "const_int_operand")))]
> - "TARGET_GFNI"
> + "TARGET_GFNI && TARGET_AVX2"
> {
> rtx matrix = ix86_vgf2p8affine_shift_matrix (operands[0], operands[2],
> <CODE>);
> emit_insn (gen_vgf2p8affineqb_<mode> (operands[0], operands[1], matrix,
> --- a/gcc/testsuite/gcc.target/i386/gfni-pr127498.c 2026-09-19
> 09:49:43.267818583 +0200
> +++ b/gcc/testsuite/gcc.target/i386/gfni-pr127498.c 2026-09-19
> 09:45:51.120060191 +0200
> @@ -0,0 +1,17 @@
> +/* PR target/127498 */
> +/* { dg-do compile } */
> +/* { dg-options "-O2 -mgfni -mavx -mno-avx2" } */
> +
> +typedef unsigned char V __attribute__ ((vector_size (32)));
> +
> +V
> +foo (V x, int n)
> +{
> + return (x >> n) | (x << (8 - n));
> +}
> +
> +V
> +bar (V x, int n)
> +{
> + return (x << n) | (x >> (8 - n));
> +}
> --- a/gcc/testsuite/gcc.target/i386/shift-gf2p8affine-5.c 2025-09-01
> 18:50:42.440203466 +0200
> +++ b/gcc/testsuite/gcc.target/i386/shift-gf2p8affine-5.c 2026-09-19
> 19:19:43.456530785 +0200
> @@ -1,5 +1,5 @@
> /* { dg-do compile } */
> -/* { dg-options "-mgfni -mavx -O3 -Wno-shift-count-negative -march=x86-64
> -mtune=generic" } */
> -/* { dg-final { scan-assembler-times "vgf2p8affineqb" 31 } } */
> +/* { dg-options "-mgfni -mavx2 -O3 -Wno-shift-count-negative
> +-march=x86-64 -mtune=generic" } */
> +/* { dg-final { scan-assembler-times "vgf2p8affineqb" 46 } } */
>
> #include "shift-gf2p8affine-2.c"
>
> Jakub