> On 12 Aug 2026, at 16:45, Tamar Christina <[email protected]> wrote: > >> -----Original Message----- >> From: [email protected] <[email protected]> >> Sent: 12 August 2026 12:44 >> To: [email protected] >> Cc: Tamar Christina <[email protected]>; Wilco Dijkstra >> <[email protected]>; Kyrylo Tkachov <[email protected]> >> Subject: [PATCH 2/2] aarch64: use [SU]DOT for the byte to word step of a >> widening sum >> >> From: Kyrylo Tkachov <[email protected]> >> >> A widening sum from bytes into 64-bit elements spends two [SU]ADDLP >> getting from bytes to words. With dot product that step is a single >> [SU]DOT against a vector of ones, which is what the byte to word expander >> already does for a 4x reduction. Each 32-bit element then holds the sum >> of four input elements, at most 4 * 255 unsigned and within -512 to 508 >> signed, so no sum can overflow. >> >> Move the dot product step into aarch64_expand_reduc_widen_sum, so that >> any >> chain that passes through a byte to word step uses it. The only shape that >> gains is V2DI <- V16QI, because the other shapes either do not start from >> bytes or already stop at 32-bit elements: >> >> V8HI <- V16QI [SU]ADALP word elements would be too wide >> V4SI <- V8HI [SU]ADALP not a byte source >> V2DI <- V4SI [SU]ADALP not a byte source >> V2SI <- V8QI [SU]DOT unchanged >> V4SI <- V16QI [SU]DOT unchanged >> V2DI <- V8HI [SU]ADDLP + [SU]ADALP not a byte source >> V2DI <- V16QI [SU]DOT + [SU]ADALP new >> >> Without dot product every shape keeps the pairwise chain. >> >> For a sum of unsigned char into long the inner loop changes from >> >> ldr q31, [x2], 16 >> uaddlp v31.8h, v31.16b >> uaddlp v31.4s, v31.8h >> uadalp v30.2d, v31.4s >> >> to >> >> ldr q29, [x2], 16 >> movi v31.4s, 0 >> udot v31.4s, v29.16b, v27.16b >> uadalp v30.2d, v31.4s >> >> with the vector of ones in v27 hoisted out of the loop. The instruction >> count is unchanged but the vector work is spread better. >> A sum of unsigned char into long runs about 24% faster at >> -march=armv8.2-a+dotprod, and about 20% faster with an L1 resident >> working >> set at -mcpu=neoverse-v2, where the vectorizer unrolls the loop by four. >> >> Bootstrapped and tested on aarch64-none-linux-gnu. >> Ok for trunk? > > This looks good to me, however do we not provide this optab for SVE? > > https://godbolt.org/z/bxas763ze seems like we don't. I was expecting > it to do a widening load from b -> h and then use udot from h -> d. > > instead of widening from b -> d, as that's a much lower VF.. > > Would you mind checking? >
I think we’d need either vect_recog_widen_sum_pattern to try an intermediate halfword type (it currently looks through the promotion to the original byte type and checks only a direct byte-to-doubleword optab) Or we’d need to implement a VNx2DI <- VNx8QI widening sum expander for SVE. I can look at either as a follow-up, but I’m guessing they are beyond the scope of this patch. Thanks, Kyrill > Thanks, > Tamar > >> Thanks, >> Kyrill >> >> gcc/ChangeLog: >> >> * config/aarch64/aarch64-simd.md >> (reduc_widen_<su>sum<mode><vsi2qi>3): >> Expand through aarch64_expand_reduc_widen_sum. >> * config/aarch64/aarch64.cc (aarch64_expand_reduc_widen_sum): >> Use >> [SU]DOT for a step from byte to word elements. >> >> gcc/testsuite/ChangeLog: >> >> * gcc.target/aarch64/widen_sum_pairwise_2.c: Cover every widening >> sum shape and check the dot product sequences. >> * gcc.target/aarch64/widen_sum_pairwise_3.c: New test. >> >> Signed-off-by: Kyrylo Tkachov <[email protected]> >> --- >> gcc/config/aarch64/aarch64-simd.md | 16 +--- >> gcc/config/aarch64/aarch64.cc | 41 ++++++++- >> .../gcc.target/aarch64/widen_sum_pairwise_2.c | 71 ++++++++++------ >> .../gcc.target/aarch64/widen_sum_pairwise_3.c | 83 >> +++++++++++++++++++ >> 4 files changed, 170 insertions(+), 41 deletions(-) >> create mode 100644 >> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c >> >> diff --git a/gcc/config/aarch64/aarch64-simd.md >> b/gcc/config/aarch64/aarch64-simd.md >> index 43be461ef04..6650cbb5f7e 100644 >> --- a/gcc/config/aarch64/aarch64-simd.md >> +++ b/gcc/config/aarch64/aarch64-simd.md >> @@ -5320,10 +5320,7 @@ >> DONE; >> }) >> >> -;; A widening sum reduction that quarters the lane count. With dot product >> -;; this is one [SU]DOT with a vector of ones, i.e. += a becomes += (a * 1). >> -;; Otherwise it is a pairwise widening add feeding a pairwise widening >> -;; accumulate. >> +;; A widening sum reduction that quarters the lane count. >> (define_expand "reduc_widen_<su>sum<mode><vsi2qi>3" >> [(set (match_operand:VS 0 "register_operand") >> (plus:VS (ANY_EXTEND:VS >> @@ -5331,15 +5328,8 @@ >> (match_operand:VS 2 "register_operand")))] >> "TARGET_SIMD" >> { >> - if (TARGET_DOTPROD) >> - { >> - rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); >> - emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0], >> operands[1], >> - ones, operands[2])); >> - } >> - else >> - aarch64_expand_reduc_widen_sum (operands[0], operands[2], >> operands[1], >> - <CODE>); >> + aarch64_expand_reduc_widen_sum (operands[0], operands[2], >> operands[1], >> + <CODE>); >> DONE; >> } >> ) >> diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc >> index c1d57ca3964..fdffb13ad22 100644 >> --- a/gcc/config/aarch64/aarch64.cc >> +++ b/gcc/config/aarch64/aarch64.cc >> @@ -26331,8 +26331,9 @@ aarch64_expand_vector_init (rtx target, rtx vals) >> Advanced SIMD vector SRC holds an even multiple of the number of lanes >> of the accumulator ACC and of the result DEST. EXTEND_CODE is >> SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form. >> - Halve the lane count with [SU]ADDLP until a single pairwise step is >> - left, then accumulate into ACC with [SU]ADALP. */ >> + Quarter the lane count of a vector of bytes with a [SU]DOT against a >> + vector of ones where that is available, halve it with [SU]ADDLP until a >> + single pairwise step is left, then accumulate into ACC with [SU]ADALP. >> */ >> >> void >> aarch64_expand_reduc_widen_sum (rtx dest, rtx acc, rtx src, >> @@ -26340,7 +26341,41 @@ aarch64_expand_reduc_widen_sum (rtx dest, >> rtx acc, rtx src, >> { >> unsigned int dest_nunits = GET_MODE_NUNITS (GET_MODE >> (dest)).to_constant (); >> machine_mode mode = GET_MODE (src); >> - gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * 2) >> == 0); >> + unsigned int nunits = GET_MODE_NUNITS (mode).to_constant (); >> + gcc_assert (nunits % (dest_nunits * 2) == 0); >> + >> + /* [SU]DOT against a vector of ones turns += a into += (a * 1), which >> + sums four bytes into each 32-bit element and so covers two halving >> + steps in one operation. The widest intermediate is 4 * 255, so no >> + product sum can overflow. Only a step from bytes to words qualifies, >> + and only if the accumulator is at least that wide. */ >> + if (TARGET_DOTPROD >> + && GET_MODE_INNER (mode) == QImode >> + && nunits >= dest_nunits * 4) >> + { >> + machine_mode sum_mode >> + = related_vector_mode (mode, SImode, nunits / 4).require (); >> + convert_optab dot = (extend_code == SIGN_EXTEND >> + ? sdot_prod_optab : udot_prod_optab); >> + insn_code icode = convert_optab_handler (dot, sum_mode, mode); >> + rtx ones = force_reg (mode, CONST1_RTX (mode)); >> + >> + /* A dot product that already reaches the element width of DEST >> + accumulates into ACC itself, otherwise it starts from zero and the >> + remaining steps carry its result into ACC. */ >> + if (sum_mode == GET_MODE (dest)) >> + { >> + emit_insn (GEN_FCN (icode) (dest, src, ones, acc)); >> + return; >> + } >> + >> + rtx tmp = gen_reg_rtx (sum_mode); >> + emit_insn (GEN_FCN (icode) (tmp, src, ones, >> + force_reg (sum_mode, >> + CONST0_RTX (sum_mode)))); >> + src = tmp; >> + mode = sum_mode; >> + } >> >> while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2) >> { >> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >> index 01537deeb9f..9b3ba07637f 100644 >> --- a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >> @@ -1,29 +1,50 @@ >> /* { dg-do compile } */ >> /* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec- >> preference=asimd-only --param vect-epilogues-nomask=0" } */ >> >> -/* With dot product a 4x widening sum stays a single [SU]DOT, while a >> - sum into 64-bit elements uses the pairwise widening instructions. */ >> - >> -int >> -sum_u8_i (const unsigned char *a, long n) >> -{ >> - int s = 0; >> - for (long i = 0; i < n; i++) >> - s += a[i]; >> - return s; >> -} >> - >> -long >> -sum_u8_l (const unsigned char *a, long n) >> -{ >> - long s = 0; >> - for (long i = 0; i < n; i++) >> - s += a[i]; >> - return s; >> -} >> - >> -/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, >> v[0- >> 9]+\.16b\n} 1 } } */ >> -/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n} >> 1 } } */ >> +/* With dot product every widening sum that passes through a byte to word >> + step uses one [SU]DOT for that step. A step that starts or ends >> + somewhere else still uses the pairwise widening instructions. */ >> + >> +#define DEF(NAME, ITYPE, OTYPE) \ >> + OTYPE NAME (const ITYPE *a, long n) \ >> + { \ >> + OTYPE s = 0; \ >> + for (long i = 0; i < n; i++) \ >> + s += a[i]; \ >> + return s; \ >> + } >> + >> +/* 2x, no dot product: the result elements are too narrow. */ >> +DEF (sum_u8_h, unsigned char, unsigned short) >> +DEF (sum_i8_h, signed char, short) >> +DEF (sum_u16_i, unsigned short, int) >> +DEF (sum_i16_i, short, int) >> +DEF (sum_u32_l, unsigned int, long) >> +DEF (sum_i32_l, int, long) >> + >> +/* 4x from bytes: one dot product. */ >> +DEF (sum_u8_i, unsigned char, int) >> +DEF (sum_i8_i, signed char, int) >> + >> +/* 4x from halfwords: no dot product for that element size. */ >> +DEF (sum_u16_l, unsigned short, long) >> +DEF (sum_i16_l, short, long) >> + >> +/* 8x from bytes: a dot product followed by one pairwise accumulate. */ >> +DEF (sum_u8_l, unsigned char, long) >> +DEF (sum_i8_l, signed char, long) >> + >> +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, >> v[0- >> 9]+\.16b\n} 2 } } */ >> +/* { dg-final { scan-assembler-times {\tsdot\tv[0-9]+\.4s, v[0-9]+\.16b, >> v[0- >> 9]+\.16b\n} 2 } } */ >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.8h, v[0-9]+\.16b\n} >> 1 } } */ >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.8h, v[0-9]+\.16b\n} >> 1 } } */ >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} >> 1 >> } } */ >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} >> 1 >> } } */ >> /* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} 1 >> } } */ >> -/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} >> 1 >> } } */ >> -/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */ >> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} >> 1 >> } } */ >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} >> 3 >> } } */ >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} >> 3 >> } } */ >> + >> +/* The byte to halfword step is what the dot product replaces. */ >> +/* { dg-final { scan-assembler-not {\t[su]addlp\tv[0-9]+\.8h, v[0- >> 9]+\.16b\n} } } */ >> +/* { dg-final { scan-assembler-not {\t[su]addw2?\t} } } */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c >> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c >> new file mode 100644 >> index 00000000000..d2eb8152722 >> --- /dev/null >> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_3.c >> @@ -0,0 +1,83 @@ >> +/* { dg-do run } */ >> +/* { dg-require-effective-target arm_v8_2a_dotprod_neon_hw } */ >> +/* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec- >> preference=asimd-only" } */ >> + >> +/* Both expansions of a widening sum reduction, with and without dot >> + product, must agree with a scalar sum for every narrow to wide type >> + pair. The accumulators are unsigned so that overflow wraps. */ >> + >> +#define TYPES(X) \ >> + X (u8_h, unsigned char, unsigned short) \ >> + X (i8_h, signed char, unsigned short) \ >> + X (u8_i, unsigned char, unsigned int) \ >> + X (i8_i, signed char, unsigned int) \ >> + X (u8_l, unsigned char, unsigned long) \ >> + X (i8_l, signed char, unsigned long) \ >> + X (u16_i, unsigned short, unsigned int) \ >> + X (i16_i, short, unsigned int) \ >> + X (u16_l, unsigned short, unsigned long) \ >> + X (i16_l, short, unsigned long) \ >> + X (u32_l, unsigned int, unsigned long) \ >> + X (i32_l, int, unsigned long) >> + >> +#define SUM(PREFIX, NAME, ITYPE, OTYPE) \ >> + __attribute__ ((noipa)) \ >> + OTYPE PREFIX##_##NAME (const ITYPE *a, int n) \ >> + { \ >> + OTYPE s = 0; \ >> + for (int i = 0; i < n; i++) \ >> + s += a[i]; \ >> + return s; \ >> + } >> + >> +#define DOT(NAME, ITYPE, OTYPE) SUM (dot, NAME, ITYPE, OTYPE) >> +#define NODOT(NAME, ITYPE, OTYPE) SUM (nodot, NAME, ITYPE, OTYPE) >> + >> +/* A volatile accumulator keeps this loop scalar. */ >> +#define REF(NAME, ITYPE, OTYPE) \ >> + __attribute__ ((noipa)) \ >> + OTYPE ref_##NAME (const ITYPE *a, int n) \ >> + { \ >> + volatile OTYPE s = 0; \ >> + for (int i = 0; i < n; i++) \ >> + s = s + a[i]; \ >> + return s; \ >> + } >> + >> +TYPES (DOT) >> +TYPES (REF) >> + >> +#pragma GCC push_options >> +#pragma GCC target ("+nodotprod") >> +TYPES (NODOT) >> +#pragma GCC pop_options >> + >> +#define BYTES 8192 >> +static unsigned char buf[BYTES] __attribute__ ((aligned (64))); >> + >> +#define CHECK(NAME, ITYPE, OTYPE) \ >> + { \ >> + const ITYPE *p = (const ITYPE *) (buf + off); \ >> + OTYPE want = ref_##NAME (p, n); \ >> + if (dot_##NAME (p, n) != want || nodot_##NAME (p, n) != want) \ >> + __builtin_abort (); \ >> + } >> + >> +int >> +main (void) >> +{ >> + unsigned long x = 1; >> + for (int i = 0; i < BYTES; i++) >> + { >> + x = x * 6364136223846793005UL + 1442695040888963407UL; >> + buf[i] = x >> 40; >> + } >> + >> + for (int off = 0; off < 8; off += 4) >> + for (int n = 0; n <= 260; n++) >> + { >> + TYPES (CHECK) >> + } >> + >> + return 0; >> +} >> -- >> 2.50.1 (Apple Git-155)
