Hi!

The following testcase ICEs, because the middle-end doesn't like a target
to announce rot{l,r}v32qi3 optab availability only with const_int_operand
rotate with non-working fallback.
If the rotate optab fails because of unsatisfiable predicates, the
middle-end will try to expand the rotate as bitwise or of 2 shifts, but
the problem with TARGET_GFNI && TARGET_AVX && !TARGET_AVX2 is that
the V32QImode shifts aren't supported either, one needs AVX2 for that.
I believe there are no CPUs with GFNI and AVX but without AVX2, I think
there are just CPUs with both GFNI and AVX/AVX2 or with GFNI and no AVX,
so the following patch just disables the pattern for the
TARGET_GFNI && TARGET_AVX combination and just requires
TARGET_GFNI && TARGET_AVX2 instead.  It is true that rotate can be expanded
with constant rotate count even for V32QImode, but at the expense of ICEing for
non-constant rotate count.
The alternative would be to extend the pattern to provide a fallback for
TARGET_GFNI && TARGET_AVX && !TARGET_AVX2 when the last operand is not
const_int_operand, it can be handled as 4 V16QImode shifts, combining stuff
together, ...

Bootstrapped/regtested on x86_64-linux and i686-linux, ok for trunk?

2026-09-19  Jakub Jelinek  <[email protected]>

        PR target/127498
        * config/i386/sse.md (<insn><mode>3 any_rotate:VI1_AVX512_3264): Add
        TARGET_AVX2 to condition.

        * gcc.target/i386/gfni-pr127498.c: New test.
        * gcc.target/i386/shift-gf2p8affine-5.c: Use -mavx2 in dg-options
        instead of -mavx.  Expect 46 insns instead of 31.

--- a/gcc/config/i386/sse.md    2026-09-15 12:01:51.693090054 +0200
+++ b/gcc/config/i386/sse.md    2026-09-19 09:52:47.935239962 +0200
@@ -28458,7 +28458,7 @@ (define_expand "<insn><mode>3"
        (any_rotate:VI1_AVX512_3264
          (match_operand:VI1_AVX512_3264 1 "register_operand")
          (match_operand:SI 2 "const_int_operand")))]
-  "TARGET_GFNI"
+  "TARGET_GFNI && TARGET_AVX2"
 {
   rtx matrix = ix86_vgf2p8affine_shift_matrix (operands[0], operands[2], 
<CODE>);
   emit_insn (gen_vgf2p8affineqb_<mode> (operands[0], operands[1], matrix,
--- a/gcc/testsuite/gcc.target/i386/gfni-pr127498.c     2026-09-19 
09:49:43.267818583 +0200
+++ b/gcc/testsuite/gcc.target/i386/gfni-pr127498.c     2026-09-19 
09:45:51.120060191 +0200
@@ -0,0 +1,17 @@
+/* PR target/127498 */
+/* { dg-do compile } */
+/* { dg-options "-O2 -mgfni -mavx -mno-avx2" } */
+
+typedef unsigned char V __attribute__ ((vector_size (32)));
+
+V
+foo (V x, int n)
+{
+  return (x >> n) | (x << (8 - n));
+}
+
+V
+bar (V x, int n)
+{
+  return (x << n) | (x >> (8 - n));
+}
--- a/gcc/testsuite/gcc.target/i386/shift-gf2p8affine-5.c       2025-09-01 
18:50:42.440203466 +0200
+++ b/gcc/testsuite/gcc.target/i386/shift-gf2p8affine-5.c       2026-09-19 
19:19:43.456530785 +0200
@@ -1,5 +1,5 @@
 /* { dg-do compile } */
-/* { dg-options "-mgfni -mavx -O3 -Wno-shift-count-negative -march=x86-64 
-mtune=generic" } */
-/* { dg-final { scan-assembler-times "vgf2p8affineqb" 31 } } */
+/* { dg-options "-mgfni -mavx2 -O3 -Wno-shift-count-negative -march=x86-64 
-mtune=generic" } */
+/* { dg-final { scan-assembler-times "vgf2p8affineqb" 46 } } */
 
 #include "shift-gf2p8affine-2.c"

        Jakub

Reply via email to