Hi Tamar, > On 4 Aug 2026, at 14:52, Tamar Christina <[email protected]> wrote: > > Hi Kyrill, > >> -----Original Message----- >> From: [email protected] <[email protected]> >> Sent: 04 August 2026 13:39 >> To: [email protected] >> Cc: Tamar Christina <[email protected]>; [email protected]; Kyrylo >> Tkachov <[email protected]> >> Subject: [PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum >> reductions >> >> From: Kyrylo Tkachov <[email protected]> >> >> The Advanced SIMD widen_[su]sum optabs only cover a single widening step, >> expanded as a dependent <su>addw + <su>addw2 pair, plus a 4x form that >> requires dot product. A reduction into an accumulator that is more than >> twice as wide as the data therefore has to extend the input explicitly and >> then issue one widening add per half vector. Summing bytes into a 64-bit >> accumulator costs fifteen SIMD operations per 16 bytes of input. >> >> [SU]ADDLP and [SU]ADALP add adjacent lane pairs into the next wider >> element, so a chain of them expresses any power-of-two widening sum >> reduction in one operation per step. The regrouping is exact because the >> sum of two elements always fits in the doubled element width, and the >> grouping of lanes inside a reduction accumulator is already unconstrained >> for WIDEN_SUM_EXPR, which the existing dot product based 4x expander also >> relies on. >> >> Expand the 2x forms as a single [SU]ADALP, add the missing V4SI <- V16QI >> and V2SI <- V8QI forms for !TARGET_DOTPROD, and add the V2DI <- V8HI and >> V2DI <- V16QI forms that no expander covered. All of them are built by >> aarch64_expand_widen_sum, which halves the lane count with [SU]ADDLP >> until >> one pairwise step remains and then accumulates with [SU]ADALP. >> >> For a sum of unsigned char into long the inner loop changes from >> >> ldr q30, [x1], 16 >> zip1 v28.16b, v30.16b, v29.16b >> zip2 v30.16b, v30.16b, v29.16b >> zip1 v26.8h, v28.8h, v29.8h >> zip2 v28.8h, v28.8h, v29.8h >> zip1 v27.8h, v30.8h, v29.8h >> zip2 v30.8h, v30.8h, v29.8h >> uaddw v31.2d, v31.2d, v26.2s >> uaddw2 v31.2d, v31.2d, v26.4s >> ... (six more uaddw/uaddw2) >> >> to >> >> ldr q31, [x1], 16 >> uaddlp v31.8h, v31.16b >> uaddlp v31.4s, v31.8h >> uadalp v30.2d, v31.4s >> > > Have you considered considered even for the b to d case to use dotprod > for the b -> s and then uadalp for the final 4 to d? or is the above codegen > the fallback for when !dotprod? wasn't quite clear.. >
It’s not a fallback as the dotprod path exists only for b -> s. Using dotprod here is interesting, I’ll try it out. Thanks for the suggestion! Kyrill > That should still shave off 1 cycle. > > Thanks, > Tamar > >> and for a sum of int into long the saddw/saddw2 pair becomes one sadalp. >> On a Neoverse V2 core with an L1 resident working set this cuts the time of >> the >> byte loop by about 88% and of the int loop by about 68%. >> >> Bootstrapped and tested on aarch64-none-linux-gnu. >> Ok for trunk? >> Thanks, >> Kyrill >> >> gcc/ChangeLog: >> >> * config/aarch64/aarch64-protos.h (aarch64_expand_widen_sum): >> Declare. >> * config/aarch64/aarch64.cc (aarch64_expand_widen_sum): New >> function. >> * config/aarch64/aarch64-simd.md (aarch64_<su>adalp<mode>): >> Rename >> to ... >> (@aarch64_<su>adalp<mode>): ... this. >> (widen_ssum<Vdblw><mode>3, widen_usum<Vdblw><mode>3): >> Replace by ... >> (widen_<su>sum<Vdblw><mode>3): ... this. Expand to [SU]ADALP. >> (widen_ssum<mode><vsi2qi>3, widen_usum<mode><vsi2qi>3): >> Replace >> by ... >> (widen_<su>sum<mode><vsi2qi>3): ... this. Handle >> !TARGET_DOTPROD. >> (widen_<su>sumv2di<mode>3): New expander. >> * config/aarch64/iterators.md (VQ_BH): New mode iterator. >> >> gcc/testsuite/ChangeLog: >> >> * gcc.target/aarch64/pr122069_1.c: Update expected output. >> * gcc.target/aarch64/pr122069_3.c: Likewise. >> * gcc.target/aarch64/saddw-1.c: Renamed to... >> * gcc.target/aarch64/sadalp-1.c: ...this. Update expected output. >> * gcc.target/aarch64/saddw-2.c: Renamed to... >> * gcc.target/aarch64/sadalp-2.c: ...this. Update expected output. >> * gcc.target/aarch64/uaddw-1.c: Renamed to... >> * gcc.target/aarch64/uadalp-1.c: ...this. Update expected output. >> * gcc.target/aarch64/uaddw-2.c: Renamed to... >> * gcc.target/aarch64/uadalp-2.c: ...this. Update expected output. >> * gcc.target/aarch64/uaddw-3.c: Renamed to... >> * gcc.target/aarch64/uadalp-3.c: ...this. Update expected output. >> * gcc.target/aarch64/widen_sum_pairwise_1.c: New test. >> * gcc.target/aarch64/widen_sum_pairwise_2.c: New test. >> >> Signed-off-by: Kyrylo Tkachov <[email protected]> >> --- >> gcc/config/aarch64/aarch64-protos.h | 1 + >> gcc/config/aarch64/aarch64-simd.md | 78 ++++++++----------- >> gcc/config/aarch64/aarch64.cc | 27 +++++++ >> gcc/config/aarch64/iterators.md | 4 + >> gcc/testsuite/gcc.target/aarch64/pr122069_1.c | 11 +-- >> gcc/testsuite/gcc.target/aarch64/pr122069_3.c | 3 +- >> .../aarch64/{saddw-1.c => sadalp-1.c} | 3 +- >> .../aarch64/{saddw-2.c => sadalp-2.c} | 3 +- >> .../aarch64/{uaddw-1.c => uadalp-1.c} | 3 +- >> .../aarch64/{uaddw-2.c => uadalp-2.c} | 3 +- >> .../aarch64/{uaddw-3.c => uadalp-3.c} | 3 +- >> .../gcc.target/aarch64/widen_sum_pairwise_1.c | 39 ++++++++++ >> .../gcc.target/aarch64/widen_sum_pairwise_2.c | 29 +++++++ >> 13 files changed, 141 insertions(+), 66 deletions(-) >> rename gcc/testsuite/gcc.target/aarch64/{saddw-1.c => sadalp-1.c} (74%) >> rename gcc/testsuite/gcc.target/aarch64/{saddw-2.c => sadalp-2.c} (74%) >> rename gcc/testsuite/gcc.target/aarch64/{uaddw-1.c => uadalp-1.c} (75%) >> rename gcc/testsuite/gcc.target/aarch64/{uaddw-2.c => uadalp-2.c} (75%) >> rename gcc/testsuite/gcc.target/aarch64/{uaddw-3.c => uadalp-3.c} (74%) >> create mode 100644 >> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c >> create mode 100644 >> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >> >> diff --git a/gcc/config/aarch64/aarch64-protos.h >> b/gcc/config/aarch64/aarch64-protos.h >> index bcc833cfaa1..9303f12f80c 100644 >> --- a/gcc/config/aarch64/aarch64-protos.h >> +++ b/gcc/config/aarch64/aarch64-protos.h >> @@ -1066,6 +1066,7 @@ void aarch64_emit_sve_pred_vec_duplicate >> (machine_mode, rtx, rtx); >> void aarch64_expand_prologue (void); >> void aarch64_decompose_vec_struct_index (machine_mode, rtx *, rtx *, >> bool); >> void aarch64_expand_vector_init (rtx, rtx); >> +void aarch64_expand_widen_sum (rtx, rtx, rtx, rtx_code); >> void aarch64_sve_expand_vector_init_subvector (rtx, rtx); >> void aarch64_sve_expand_vector_init (rtx, rtx); >> void aarch64_init_cumulative_args (CUMULATIVE_ARGS *, const_tree, rtx, >> diff --git a/gcc/config/aarch64/aarch64-simd.md >> b/gcc/config/aarch64/aarch64-simd.md >> index 433f16052bf..d119ac17352 100644 >> --- a/gcc/config/aarch64/aarch64-simd.md >> +++ b/gcc/config/aarch64/aarch64-simd.md >> @@ -1182,7 +1182,7 @@ >> } >> ) >> >> -(define_expand "aarch64_<su>adalp<mode>" >> +(define_expand "@aarch64_<su>adalp<mode>" >> [(set (match_operand:<VDBLW> 0 "register_operand") >> (plus:<VDBLW> >> (plus:<VDBLW> >> @@ -5283,19 +5283,17 @@ >> >> ;; <su><addsub>w<q>. >> >> -(define_expand "widen_ssum<Vdblw><mode>3" >> +;; A widening sum reduction that halves the lane count is a single pairwise >> +;; widening accumulate. >> +(define_expand "widen_<su>sum<Vdblw><mode>3" >> [(set (match_operand:<VDBLW> 0 "register_operand") >> - (plus:<VDBLW> (sign_extend:<VDBLW> >> - (match_operand:VQW 1 "register_operand")) >> + (plus:<VDBLW> (ANY_EXTEND:<VDBLW> >> + (match_operand:VQW 1 "register_operand")) >> (match_operand:<VDBLW> 2 "register_operand")))] >> "TARGET_SIMD" >> { >> - rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false); >> - rtx temp = gen_reg_rtx (GET_MODE (operands[0])); >> - >> - emit_insn (gen_aarch64_saddw<mode>_internal (temp, operands[2], >> - operands[1], p)); >> - emit_insn (gen_aarch64_saddw2<mode> (operands[0], temp, >> operands[1])); >> + emit_insn (gen_aarch64_<su>adalp<mode> (operands[0], operands[2], >> + operands[1])); >> DONE; >> } >> ) >> @@ -5311,23 +5309,6 @@ >> DONE; >> }) >> >> -(define_expand "widen_usum<Vdblw><mode>3" >> - [(set (match_operand:<VDBLW> 0 "register_operand") >> - (plus:<VDBLW> (zero_extend:<VDBLW> >> - (match_operand:VQW 1 "register_operand")) >> - (match_operand:<VDBLW> 2 "register_operand")))] >> - "TARGET_SIMD" >> - { >> - rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false); >> - rtx temp = gen_reg_rtx (GET_MODE (operands[0])); >> - >> - emit_insn (gen_aarch64_uaddw<mode>_internal (temp, operands[2], >> - operands[1], p)); >> - emit_insn (gen_aarch64_uaddw2<mode> (operands[0], temp, >> operands[1])); >> - DONE; >> - } >> -) >> - >> (define_expand "widen_usum<Vwide><mode>3" >> [(set (match_operand:<VWIDE> 0 "register_operand") >> (plus:<VWIDE> (zero_extend:<VWIDE> >> @@ -5339,38 +5320,43 @@ >> DONE; >> }) >> >> -(define_expand "widen_ssum<mode><vsi2qi>3" >> +;; A widening sum reduction that quarters the lane count. With dot product >> +;; this is one [SU]DOT with a vector of ones, i.e. += a becomes += (a * 1). >> +;; Otherwise it is a pairwise widening add feeding a pairwise widening >> +;; accumulate. >> +(define_expand "widen_<su>sum<mode><vsi2qi>3" >> [(set (match_operand:VS 0 "register_operand") >> - (plus:VS (sign_extend:VS >> + (plus:VS (ANY_EXTEND:VS >> (match_operand:<VSI2QI> 1 "register_operand")) >> (match_operand:VS 2 "register_operand")))] >> - "TARGET_DOTPROD" >> + "TARGET_SIMD" >> { >> - rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); >> - emit_insn (gen_sdot_prod<mode><vsi2qi> (operands[0], operands[1], >> ones, >> - operands[2])); >> + if (TARGET_DOTPROD) >> + { >> + rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); >> + emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0], >> operands[1], >> + ones, operands[2])); >> + } >> + else >> + aarch64_expand_widen_sum (operands[0], operands[2], operands[1], >> <CODE>); >> DONE; >> } >> ) >> >> -;; Use dot product to perform double widening sum reductions by >> -;; changing += a into += (a * 1). i.e. we seed the multiplication with 1. >> -(define_expand "widen_usum<mode><vsi2qi>3" >> - [(set (match_operand:VS 0 "register_operand") >> - (plus:VS (zero_extend:VS >> - (match_operand:<VSI2QI> 1 "register_operand")) >> - (match_operand:VS 2 "register_operand")))] >> - "TARGET_DOTPROD" >> +;; Widening sum reductions into 64-bit elements. These need two or three >> +;; pairwise widening steps. >> +(define_expand "widen_<su>sumv2di<mode>3" >> + [(set (match_operand:V2DI 0 "register_operand") >> + (plus:V2DI (ANY_EXTEND:V2DI >> + (match_operand:VQ_BH 1 "register_operand")) >> + (match_operand:V2DI 2 "register_operand")))] >> + "TARGET_SIMD" >> { >> - rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); >> - emit_insn (gen_udot_prod<mode><vsi2qi> (operands[0], operands[1], >> ones, >> - operands[2])); >> + aarch64_expand_widen_sum (operands[0], operands[2], operands[1], >> <CODE>); >> DONE; >> } >> ) >> >> -;; Use dot product to perform double widening sum reductions by >> -;; changing += a into += (a * 1). i.e. we seed the multiplication with 1. >> (define_insn "aarch64_<ANY_EXTEND:su>subw<mode>" >> [(set (match_operand:<VWIDE> 0 "register_operand" "=w") >> (minus:<VWIDE> (match_operand:<VWIDE> 1 "register_operand" >> "w") >> diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc >> index d19ca305d82..628e94e8c40 100644 >> --- a/gcc/config/aarch64/aarch64.cc >> +++ b/gcc/config/aarch64/aarch64.cc >> @@ -26327,6 +26327,33 @@ aarch64_expand_vector_init (rtx target, rtx >> vals) >> emit_insn (seq_total_cost < fallback_seq_cost ? seq : fallback_seq); >> } >> >> +/* Expand the widening sum reduction DEST = ACC + (WIDE) SRC, where the >> + Advanced SIMD vector SRC holds an even multiple of the number of lanes >> + of the accumulator ACC and of the result DEST. EXTEND_CODE is >> + SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form. >> + Halve the lane count with [SU]ADDLP until a single pairwise step is >> + left, then accumulate into ACC with [SU]ADALP. */ >> + >> +void >> +aarch64_expand_widen_sum (rtx dest, rtx acc, rtx src, rtx_code extend_code) >> +{ >> + unsigned int dest_nunits = GET_MODE_NUNITS (GET_MODE >> (dest)).to_constant (); >> + machine_mode mode = GET_MODE (src); >> + gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * 2) >> == 0); >> + >> + while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2) >> + { >> + insn_code icode = code_for_aarch64_addlp (extend_code, mode); >> + mode = insn_data[icode].operand[0].mode; >> + rtx tmp = gen_reg_rtx (mode); >> + emit_insn (GEN_FCN (icode) (tmp, src)); >> + src = tmp; >> + } >> + >> + emit_insn (GEN_FCN (code_for_aarch64_adalp (extend_code, mode)) >> (dest, acc, >> + src)); >> +} >> + >> /* Emit RTL corresponding to: >> insr TARGET, ELEM. */ >> >> diff --git a/gcc/config/aarch64/iterators.md >> b/gcc/config/aarch64/iterators.md >> index 0d319751430..6a8c93cce37 100644 >> --- a/gcc/config/aarch64/iterators.md >> +++ b/gcc/config/aarch64/iterators.md >> @@ -313,6 +313,10 @@ >> ;; All quad integer widen-able modes. >> (define_mode_iterator VQW [V16QI V8HI V4SI]) >> >> +;; Quad integer modes that reach 64-bit elements through more than one >> +;; pairwise widening step. >> +(define_mode_iterator VQ_BH [V16QI V8HI]) >> + >> ;; Double vector modes for combines. >> (define_mode_iterator VDC [V8QI V4HI V4BF V4HF V2SI V2SF DI DF]) >> >> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c >> b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c >> index b2f973261ea..d99b5493ade 100644 >> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c >> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c >> @@ -10,12 +10,8 @@ inline char char_abs(char i) { >> ** foo_int: >> ** ... >> ** sub v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b >> -** zip1 v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b >> -** zip2 v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b >> -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h >> -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h >> -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h >> -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h >> +** uaddlp v[0-9]+.8h, v[0-9]+.16b >> +** uadalp v[0-9]+.4s, v[0-9]+.8h >> ** ... >> */ >> int foo_int(unsigned char *x, unsigned char * restrict y) { >> @@ -29,8 +25,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) { >> ** foo2_int: >> ** ... >> ** add v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h >> -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h >> -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h >> +** uadalp v[0-9]+.4s, v[0-9]+.8h >> ** ... >> */ >> int foo2_int(unsigned short *x, unsigned short * restrict y) { >> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c >> b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c >> index 0e832c43032..f29fc2b2ed4 100644 >> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c >> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c >> @@ -24,8 +24,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) { >> ** foo2_int: >> ** ... >> ** add v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h >> -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h >> -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h >> +** uadalp v[0-9]+.4s, v[0-9]+.8h >> ** ... >> */ >> int foo2_int(unsigned short *x, unsigned short * restrict y) { >> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-1.c >> b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c >> similarity index 74% >> rename from gcc/testsuite/gcc.target/aarch64/saddw-1.c >> rename to gcc/testsuite/gcc.target/aarch64/sadalp-1.c >> index f8871209b8a..61f9633f1a0 100644 >> --- a/gcc/testsuite/gcc.target/aarch64/saddw-1.c >> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c >> @@ -14,5 +14,4 @@ t6(int len, void * dummy, short * __restrict x) >> return result; >> } >> >> -/* { dg-final { scan-assembler "saddw" } } */ >> -/* { dg-final { scan-assembler "saddw2" } } */ >> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-2.c >> b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c >> similarity index 74% >> rename from gcc/testsuite/gcc.target/aarch64/saddw-2.c >> rename to gcc/testsuite/gcc.target/aarch64/sadalp-2.c >> index b9fc442a2f7..873fda2e1ea 100644 >> --- a/gcc/testsuite/gcc.target/aarch64/saddw-2.c >> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c >> @@ -14,5 +14,4 @@ t6(int len, void * dummy, int * __restrict x) >> return result; >> } >> >> -/* { dg-final { scan-assembler "saddw" } } */ >> -/* { dg-final { scan-assembler "saddw2" } } */ >> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s} } } */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c >> b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c >> similarity index 75% >> rename from gcc/testsuite/gcc.target/aarch64/uaddw-1.c >> rename to gcc/testsuite/gcc.target/aarch64/uadalp-1.c >> index 14dff87d7f0..c4034384aae 100644 >> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c >> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c >> @@ -14,5 +14,4 @@ t6(int len, void * dummy, unsigned short * __restrict x) >> return result; >> } >> >> -/* { dg-final { scan-assembler "uaddw" } } */ >> -/* { dg-final { scan-assembler "uaddw2" } } */ >> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c >> b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c >> similarity index 75% >> rename from gcc/testsuite/gcc.target/aarch64/uaddw-2.c >> rename to gcc/testsuite/gcc.target/aarch64/uadalp-2.c >> index 79d0d094fc3..395d36c7c00 100644 >> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c >> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c >> @@ -14,6 +14,5 @@ t6(int len, void * dummy, unsigned short * __restrict x) >> return result; >> } >> >> -/* { dg-final { scan-assembler "uaddw" } } */ >> -/* { dg-final { scan-assembler "uaddw2" } } */ >> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */ >> >> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c >> b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c >> similarity index 74% >> rename from gcc/testsuite/gcc.target/aarch64/uaddw-3.c >> rename to gcc/testsuite/gcc.target/aarch64/uadalp-3.c >> index 39cbd6b6cc2..5fdb1639ab8 100644 >> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c >> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c >> @@ -14,5 +14,4 @@ t6(int len, void * dummy, char * __restrict x) >> return result; >> } >> >> -/* { dg-final { scan-assembler "uaddw" } } */ >> -/* { dg-final { scan-assembler "uaddw2" } } */ >> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.8h, v[0-9]+\.16b} } } */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c >> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c >> new file mode 100644 >> index 00000000000..0aec0bf81c8 >> --- /dev/null >> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c >> @@ -0,0 +1,39 @@ >> +/* { dg-do compile } */ >> +/* { dg-options "-O3 -march=armv8-a -mautovec-preference=asimd-only -- >> param vect-epilogues-nomask=0" } */ >> + >> +/* Widening sum reductions should use the pairwise widening add and >> + accumulate instructions rather than a chain of extensions feeding >> + [SU]ADDW pairs. */ >> + >> +#define DEF(NAME, ITYPE, OTYPE) \ >> + OTYPE NAME (const ITYPE *a, long n) \ >> + { \ >> + OTYPE s = 0; \ >> + for (long i = 0; i < n; i++) \ >> + s += a[i]; \ >> + return s; \ >> + } >> + >> +DEF (sum_u8_l, unsigned char, long) >> +DEF (sum_i8_l, signed char, long) >> +DEF (sum_u16_l, unsigned short, long) >> +DEF (sum_i16_l, short, long) >> +DEF (sum_u32_l, unsigned int, long) >> +DEF (sum_i32_l, int, long) >> +DEF (sum_u8_i, unsigned char, int) >> +DEF (sum_i8_i, signed char, int) >> +DEF (sum_u16_i, unsigned short, int) >> +DEF (sum_i16_i, short, int) >> + >> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n} >> 2 } } */ >> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n} >> 2 } } */ >> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} >> 2 >> } } */ >> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} >> 2 >> } } */ >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} >> 3 >> } } */ >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} >> 3 >> } } */ >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} >> 2 >> } } */ >> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h\n} >> 2 >> } } */ >> + >> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */ >> +/* { dg-final { scan-assembler-not {\tsaddw2?\t} } } */ >> +/* { dg-final { scan-assembler-not {\tzip1\t} } } */ >> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >> new file mode 100644 >> index 00000000000..01537deeb9f >> --- /dev/null >> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >> @@ -0,0 +1,29 @@ >> +/* { dg-do compile } */ >> +/* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec- >> preference=asimd-only --param vect-epilogues-nomask=0" } */ >> + >> +/* With dot product a 4x widening sum stays a single [SU]DOT, while a >> + sum into 64-bit elements uses the pairwise widening instructions. */ >> + >> +int >> +sum_u8_i (const unsigned char *a, long n) >> +{ >> + int s = 0; >> + for (long i = 0; i < n; i++) >> + s += a[i]; >> + return s; >> +} >> + >> +long >> +sum_u8_l (const unsigned char *a, long n) >> +{ >> + long s = 0; >> + for (long i = 0; i < n; i++) >> + s += a[i]; >> + return s; >> +} >> + >> +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, >> v[0- >> 9]+\.16b\n} 1 } } */ >> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, v[0-9]+\.16b\n} >> 1 } } */ >> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, v[0-9]+\.8h\n} >> 1 >> } } */ >> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, v[0-9]+\.4s\n} >> 1 >> } } */ >> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */ >> -- >> 2.50.1 (Apple Git-155)
