> On 4 Aug 2026, at 15:26, Kyrylo Tkachov <[email protected]> wrote: > > Hi Tamar, > >> On 4 Aug 2026, at 14:52, Tamar Christina <[email protected]> wrote: >> >> Hi Kyrill, >> >>> -----Original Message----- >>> From: [email protected] <[email protected]> >>> Sent: 04 August 2026 13:39 >>> To: [email protected] >>> Cc: Tamar Christina <[email protected]>; [email protected]; Kyrylo >>> Tkachov <[email protected]> >>> Subject: [PATCH 2/2] aarch64: use [SU]ADDLP/[SU]ADALP for widening sum >>> reductions >>> >>> From: Kyrylo Tkachov <[email protected]> >>> >>> The Advanced SIMD widen_[su]sum optabs only cover a single widening step, >>> expanded as a dependent <su>addw + <su>addw2 pair, plus a 4x form that >>> requires dot product. A reduction into an accumulator that is more than >>> twice as wide as the data therefore has to extend the input explicitly and >>> then issue one widening add per half vector. Summing bytes into a 64-bit >>> accumulator costs fifteen SIMD operations per 16 bytes of input. >>> >>> [SU]ADDLP and [SU]ADALP add adjacent lane pairs into the next wider >>> element, so a chain of them expresses any power-of-two widening sum >>> reduction in one operation per step. The regrouping is exact because the >>> sum of two elements always fits in the doubled element width, and the >>> grouping of lanes inside a reduction accumulator is already unconstrained >>> for WIDEN_SUM_EXPR, which the existing dot product based 4x expander also >>> relies on. >>> >>> Expand the 2x forms as a single [SU]ADALP, add the missing V4SI <- V16QI >>> and V2SI <- V8QI forms for !TARGET_DOTPROD, and add the V2DI <- V8HI and >>> V2DI <- V16QI forms that no expander covered. All of them are built by >>> aarch64_expand_widen_sum, which halves the lane count with [SU]ADDLP >>> until >>> one pairwise step remains and then accumulates with [SU]ADALP. >>> >>> For a sum of unsigned char into long the inner loop changes from >>> >>> ldr q30, [x1], 16 >>> zip1 v28.16b, v30.16b, v29.16b >>> zip2 v30.16b, v30.16b, v29.16b >>> zip1 v26.8h, v28.8h, v29.8h >>> zip2 v28.8h, v28.8h, v29.8h >>> zip1 v27.8h, v30.8h, v29.8h >>> zip2 v30.8h, v30.8h, v29.8h >>> uaddw v31.2d, v31.2d, v26.2s >>> uaddw2 v31.2d, v31.2d, v26.4s >>> ... (six more uaddw/uaddw2) >>> >>> to >>> >>> ldr q31, [x1], 16 >>> uaddlp v31.8h, v31.16b >>> uaddlp v31.4s, v31.8h >>> uadalp v30.2d, v31.4s >>> >> >> Have you considered considered even for the b to d case to use dotprod >> for the b -> s and then uadalp for the final 4 to d? or is the above codegen >> the fallback for when !dotprod? wasn't quite clear.. >> > > It’s not a fallback as the dotprod path exists only for b -> s. > Using dotprod here is interesting, I’ll try it out. > Thanks for the suggestion!
It works well and now that the optab renaming has landed I’ve sent out two patches to implement these expansions. Thanks, Kyrill > Kyrill > >> That should still shave off 1 cycle. >> >> Thanks, >> Tamar >> >>> and for a sum of int into long the saddw/saddw2 pair becomes one sadalp. >>> On a Neoverse V2 core with an L1 resident working set this cuts the time of >>> the >>> byte loop by about 88% and of the int loop by about 68%. >>> >>> Bootstrapped and tested on aarch64-none-linux-gnu. >>> Ok for trunk? >>> Thanks, >>> Kyrill >>> >>> gcc/ChangeLog: >>> >>> * config/aarch64/aarch64-protos.h (aarch64_expand_widen_sum): >>> Declare. >>> * config/aarch64/aarch64.cc (aarch64_expand_widen_sum): New >>> function. >>> * config/aarch64/aarch64-simd.md (aarch64_<su>adalp<mode>): >>> Rename >>> to ... >>> (@aarch64_<su>adalp<mode>): ... this. >>> (widen_ssum<Vdblw><mode>3, widen_usum<Vdblw><mode>3): >>> Replace by ... >>> (widen_<su>sum<Vdblw><mode>3): ... this. Expand to [SU]ADALP. >>> (widen_ssum<mode><vsi2qi>3, widen_usum<mode><vsi2qi>3): >>> Replace >>> by ... >>> (widen_<su>sum<mode><vsi2qi>3): ... this. Handle >>> !TARGET_DOTPROD. >>> (widen_<su>sumv2di<mode>3): New expander. >>> * config/aarch64/iterators.md (VQ_BH): New mode iterator. >>> >>> gcc/testsuite/ChangeLog: >>> >>> * gcc.target/aarch64/pr122069_1.c: Update expected output. >>> * gcc.target/aarch64/pr122069_3.c: Likewise. >>> * gcc.target/aarch64/saddw-1.c: Renamed to... >>> * gcc.target/aarch64/sadalp-1.c: ...this. Update expected output. >>> * gcc.target/aarch64/saddw-2.c: Renamed to... >>> * gcc.target/aarch64/sadalp-2.c: ...this. Update expected output. >>> * gcc.target/aarch64/uaddw-1.c: Renamed to... >>> * gcc.target/aarch64/uadalp-1.c: ...this. Update expected output. >>> * gcc.target/aarch64/uaddw-2.c: Renamed to... >>> * gcc.target/aarch64/uadalp-2.c: ...this. Update expected output. >>> * gcc.target/aarch64/uaddw-3.c: Renamed to... >>> * gcc.target/aarch64/uadalp-3.c: ...this. Update expected output. >>> * gcc.target/aarch64/widen_sum_pairwise_1.c: New test. >>> * gcc.target/aarch64/widen_sum_pairwise_2.c: New test. >>> >>> Signed-off-by: Kyrylo Tkachov <[email protected]> >>> --- >>> gcc/config/aarch64/aarch64-protos.h | 1 + >>> gcc/config/aarch64/aarch64-simd.md | 78 ++++++++----------- >>> gcc/config/aarch64/aarch64.cc | 27 +++++++ >>> gcc/config/aarch64/iterators.md | 4 + >>> gcc/testsuite/gcc.target/aarch64/pr122069_1.c | 11 +-- >>> gcc/testsuite/gcc.target/aarch64/pr122069_3.c | 3 +- >>> .../aarch64/{saddw-1.c => sadalp-1.c} | 3 +- >>> .../aarch64/{saddw-2.c => sadalp-2.c} | 3 +- >>> .../aarch64/{uaddw-1.c => uadalp-1.c} | 3 +- >>> .../aarch64/{uaddw-2.c => uadalp-2.c} | 3 +- >>> .../aarch64/{uaddw-3.c => uadalp-3.c} | 3 +- >>> .../gcc.target/aarch64/widen_sum_pairwise_1.c | 39 ++++++++++ >>> .../gcc.target/aarch64/widen_sum_pairwise_2.c | 29 +++++++ >>> 13 files changed, 141 insertions(+), 66 deletions(-) >>> rename gcc/testsuite/gcc.target/aarch64/{saddw-1.c => sadalp-1.c} (74%) >>> rename gcc/testsuite/gcc.target/aarch64/{saddw-2.c => sadalp-2.c} (74%) >>> rename gcc/testsuite/gcc.target/aarch64/{uaddw-1.c => uadalp-1.c} (75%) >>> rename gcc/testsuite/gcc.target/aarch64/{uaddw-2.c => uadalp-2.c} (75%) >>> rename gcc/testsuite/gcc.target/aarch64/{uaddw-3.c => uadalp-3.c} (74%) >>> create mode 100644 >>> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c >>> create mode 100644 >>> gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >>> >>> diff --git a/gcc/config/aarch64/aarch64-protos.h >>> b/gcc/config/aarch64/aarch64-protos.h >>> index bcc833cfaa1..9303f12f80c 100644 >>> --- a/gcc/config/aarch64/aarch64-protos.h >>> +++ b/gcc/config/aarch64/aarch64-protos.h >>> @@ -1066,6 +1066,7 @@ void aarch64_emit_sve_pred_vec_duplicate >>> (machine_mode, rtx, rtx); >>> void aarch64_expand_prologue (void); >>> void aarch64_decompose_vec_struct_index (machine_mode, rtx *, rtx *, >>> bool); >>> void aarch64_expand_vector_init (rtx, rtx); >>> +void aarch64_expand_widen_sum (rtx, rtx, rtx, rtx_code); >>> void aarch64_sve_expand_vector_init_subvector (rtx, rtx); >>> void aarch64_sve_expand_vector_init (rtx, rtx); >>> void aarch64_init_cumulative_args (CUMULATIVE_ARGS *, const_tree, rtx, >>> diff --git a/gcc/config/aarch64/aarch64-simd.md >>> b/gcc/config/aarch64/aarch64-simd.md >>> index 433f16052bf..d119ac17352 100644 >>> --- a/gcc/config/aarch64/aarch64-simd.md >>> +++ b/gcc/config/aarch64/aarch64-simd.md >>> @@ -1182,7 +1182,7 @@ >>> } >>> ) >>> >>> -(define_expand "aarch64_<su>adalp<mode>" >>> +(define_expand "@aarch64_<su>adalp<mode>" >>> [(set (match_operand:<VDBLW> 0 "register_operand") >>> (plus:<VDBLW> >>> (plus:<VDBLW> >>> @@ -5283,19 +5283,17 @@ >>> >>> ;; <su><addsub>w<q>. >>> >>> -(define_expand "widen_ssum<Vdblw><mode>3" >>> +;; A widening sum reduction that halves the lane count is a single pairwise >>> +;; widening accumulate. >>> +(define_expand "widen_<su>sum<Vdblw><mode>3" >>> [(set (match_operand:<VDBLW> 0 "register_operand") >>> - (plus:<VDBLW> (sign_extend:<VDBLW> >>> - (match_operand:VQW 1 "register_operand")) >>> + (plus:<VDBLW> (ANY_EXTEND:<VDBLW> >>> + (match_operand:VQW 1 "register_operand")) >>> (match_operand:<VDBLW> 2 "register_operand")))] >>> "TARGET_SIMD" >>> { >>> - rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false); >>> - rtx temp = gen_reg_rtx (GET_MODE (operands[0])); >>> - >>> - emit_insn (gen_aarch64_saddw<mode>_internal (temp, operands[2], >>> - operands[1], p)); >>> - emit_insn (gen_aarch64_saddw2<mode> (operands[0], temp, >>> operands[1])); >>> + emit_insn (gen_aarch64_<su>adalp<mode> (operands[0], operands[2], >>> + operands[1])); >>> DONE; >>> } >>> ) >>> @@ -5311,23 +5309,6 @@ >>> DONE; >>> }) >>> >>> -(define_expand "widen_usum<Vdblw><mode>3" >>> - [(set (match_operand:<VDBLW> 0 "register_operand") >>> - (plus:<VDBLW> (zero_extend:<VDBLW> >>> - (match_operand:VQW 1 "register_operand")) >>> - (match_operand:<VDBLW> 2 "register_operand")))] >>> - "TARGET_SIMD" >>> - { >>> - rtx p = aarch64_simd_vect_par_cnst_half (<MODE>mode, <nunits>, false); >>> - rtx temp = gen_reg_rtx (GET_MODE (operands[0])); >>> - >>> - emit_insn (gen_aarch64_uaddw<mode>_internal (temp, operands[2], >>> - operands[1], p)); >>> - emit_insn (gen_aarch64_uaddw2<mode> (operands[0], temp, >>> operands[1])); >>> - DONE; >>> - } >>> -) >>> - >>> (define_expand "widen_usum<Vwide><mode>3" >>> [(set (match_operand:<VWIDE> 0 "register_operand") >>> (plus:<VWIDE> (zero_extend:<VWIDE> >>> @@ -5339,38 +5320,43 @@ >>> DONE; >>> }) >>> >>> -(define_expand "widen_ssum<mode><vsi2qi>3" >>> +;; A widening sum reduction that quarters the lane count. With dot product >>> +;; this is one [SU]DOT with a vector of ones, i.e. += a becomes += (a * 1). >>> +;; Otherwise it is a pairwise widening add feeding a pairwise widening >>> +;; accumulate. >>> +(define_expand "widen_<su>sum<mode><vsi2qi>3" >>> [(set (match_operand:VS 0 "register_operand") >>> - (plus:VS (sign_extend:VS >>> + (plus:VS (ANY_EXTEND:VS >>> (match_operand:<VSI2QI> 1 "register_operand")) >>> (match_operand:VS 2 "register_operand")))] >>> - "TARGET_DOTPROD" >>> + "TARGET_SIMD" >>> { >>> - rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); >>> - emit_insn (gen_sdot_prod<mode><vsi2qi> (operands[0], operands[1], >>> ones, >>> - operands[2])); >>> + if (TARGET_DOTPROD) >>> + { >>> + rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); >>> + emit_insn (gen_<su>dot_prod<mode><vsi2qi> (operands[0], >>> operands[1], >>> + ones, operands[2])); >>> + } >>> + else >>> + aarch64_expand_widen_sum (operands[0], operands[2], operands[1], >>> <CODE>); >>> DONE; >>> } >>> ) >>> >>> -;; Use dot product to perform double widening sum reductions by >>> -;; changing += a into += (a * 1). i.e. we seed the multiplication with 1. >>> -(define_expand "widen_usum<mode><vsi2qi>3" >>> - [(set (match_operand:VS 0 "register_operand") >>> - (plus:VS (zero_extend:VS >>> - (match_operand:<VSI2QI> 1 "register_operand")) >>> - (match_operand:VS 2 "register_operand")))] >>> - "TARGET_DOTPROD" >>> +;; Widening sum reductions into 64-bit elements. These need two or three >>> +;; pairwise widening steps. >>> +(define_expand "widen_<su>sumv2di<mode>3" >>> + [(set (match_operand:V2DI 0 "register_operand") >>> + (plus:V2DI (ANY_EXTEND:V2DI >>> + (match_operand:VQ_BH 1 "register_operand")) >>> + (match_operand:V2DI 2 "register_operand")))] >>> + "TARGET_SIMD" >>> { >>> - rtx ones = force_reg (<VSI2QI>mode, CONST1_RTX (<VSI2QI>mode)); >>> - emit_insn (gen_udot_prod<mode><vsi2qi> (operands[0], operands[1], >>> ones, >>> - operands[2])); >>> + aarch64_expand_widen_sum (operands[0], operands[2], operands[1], >>> <CODE>); >>> DONE; >>> } >>> ) >>> >>> -;; Use dot product to perform double widening sum reductions by >>> -;; changing += a into += (a * 1). i.e. we seed the multiplication with 1. >>> (define_insn "aarch64_<ANY_EXTEND:su>subw<mode>" >>> [(set (match_operand:<VWIDE> 0 "register_operand" "=w") >>> (minus:<VWIDE> (match_operand:<VWIDE> 1 "register_operand" >>> "w") >>> diff --git a/gcc/config/aarch64/aarch64.cc b/gcc/config/aarch64/aarch64.cc >>> index d19ca305d82..628e94e8c40 100644 >>> --- a/gcc/config/aarch64/aarch64.cc >>> +++ b/gcc/config/aarch64/aarch64.cc >>> @@ -26327,6 +26327,33 @@ aarch64_expand_vector_init (rtx target, rtx >>> vals) >>> emit_insn (seq_total_cost < fallback_seq_cost ? seq : fallback_seq); >>> } >>> >>> +/* Expand the widening sum reduction DEST = ACC + (WIDE) SRC, where the >>> + Advanced SIMD vector SRC holds an even multiple of the number of lanes >>> + of the accumulator ACC and of the result DEST. EXTEND_CODE is >>> + SIGN_EXTEND or ZERO_EXTEND and selects the signed or unsigned form. >>> + Halve the lane count with [SU]ADDLP until a single pairwise step is >>> + left, then accumulate into ACC with [SU]ADALP. */ >>> + >>> +void >>> +aarch64_expand_widen_sum (rtx dest, rtx acc, rtx src, rtx_code extend_code) >>> +{ >>> + unsigned int dest_nunits = GET_MODE_NUNITS (GET_MODE >>> (dest)).to_constant (); >>> + machine_mode mode = GET_MODE (src); >>> + gcc_assert (GET_MODE_NUNITS (mode).to_constant () % (dest_nunits * 2) >>> == 0); >>> + >>> + while (GET_MODE_NUNITS (mode).to_constant () > dest_nunits * 2) >>> + { >>> + insn_code icode = code_for_aarch64_addlp (extend_code, mode); >>> + mode = insn_data[icode].operand[0].mode; >>> + rtx tmp = gen_reg_rtx (mode); >>> + emit_insn (GEN_FCN (icode) (tmp, src)); >>> + src = tmp; >>> + } >>> + >>> + emit_insn (GEN_FCN (code_for_aarch64_adalp (extend_code, mode)) >>> (dest, acc, >>> + src)); >>> +} >>> + >>> /* Emit RTL corresponding to: >>> insr TARGET, ELEM. */ >>> >>> diff --git a/gcc/config/aarch64/iterators.md >>> b/gcc/config/aarch64/iterators.md >>> index 0d319751430..6a8c93cce37 100644 >>> --- a/gcc/config/aarch64/iterators.md >>> +++ b/gcc/config/aarch64/iterators.md >>> @@ -313,6 +313,10 @@ >>> ;; All quad integer widen-able modes. >>> (define_mode_iterator VQW [V16QI V8HI V4SI]) >>> >>> +;; Quad integer modes that reach 64-bit elements through more than one >>> +;; pairwise widening step. >>> +(define_mode_iterator VQ_BH [V16QI V8HI]) >>> + >>> ;; Double vector modes for combines. >>> (define_mode_iterator VDC [V8QI V4HI V4BF V4HF V2SI V2SF DI DF]) >>> >>> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c >>> b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c >>> index b2f973261ea..d99b5493ade 100644 >>> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_1.c >>> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_1.c >>> @@ -10,12 +10,8 @@ inline char char_abs(char i) { >>> ** foo_int: >>> ** ... >>> ** sub v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b >>> -** zip1 v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b >>> -** zip2 v[0-9]+.16b, v[0-9]+.16b, v[0-9]+.16b >>> -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h >>> -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h >>> -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h >>> -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h >>> +** uaddlp v[0-9]+.8h, v[0-9]+.16b >>> +** uadalp v[0-9]+.4s, v[0-9]+.8h >>> ** ... >>> */ >>> int foo_int(unsigned char *x, unsigned char * restrict y) { >>> @@ -29,8 +25,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) >>> { >>> ** foo2_int: >>> ** ... >>> ** add v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h >>> -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h >>> -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h >>> +** uadalp v[0-9]+.4s, v[0-9]+.8h >>> ** ... >>> */ >>> int foo2_int(unsigned short *x, unsigned short * restrict y) { >>> diff --git a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c >>> b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c >>> index 0e832c43032..f29fc2b2ed4 100644 >>> --- a/gcc/testsuite/gcc.target/aarch64/pr122069_3.c >>> +++ b/gcc/testsuite/gcc.target/aarch64/pr122069_3.c >>> @@ -24,8 +24,7 @@ int foo_int(unsigned char *x, unsigned char * restrict y) >>> { >>> ** foo2_int: >>> ** ... >>> ** add v[0-9]+.8h, v[0-9]+.8h, v[0-9]+.8h >>> -** uaddw v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.4h >>> -** uaddw2 v[0-9]+.4s, v[0-9]+.4s, v[0-9]+.8h >>> +** uadalp v[0-9]+.4s, v[0-9]+.8h >>> ** ... >>> */ >>> int foo2_int(unsigned short *x, unsigned short * restrict y) { >>> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-1.c >>> b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c >>> similarity index 74% >>> rename from gcc/testsuite/gcc.target/aarch64/saddw-1.c >>> rename to gcc/testsuite/gcc.target/aarch64/sadalp-1.c >>> index f8871209b8a..61f9633f1a0 100644 >>> --- a/gcc/testsuite/gcc.target/aarch64/saddw-1.c >>> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-1.c >>> @@ -14,5 +14,4 @@ t6(int len, void * dummy, short * __restrict x) >>> return result; >>> } >>> >>> -/* { dg-final { scan-assembler "saddw" } } */ >>> -/* { dg-final { scan-assembler "saddw2" } } */ >>> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */ >>> diff --git a/gcc/testsuite/gcc.target/aarch64/saddw-2.c >>> b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c >>> similarity index 74% >>> rename from gcc/testsuite/gcc.target/aarch64/saddw-2.c >>> rename to gcc/testsuite/gcc.target/aarch64/sadalp-2.c >>> index b9fc442a2f7..873fda2e1ea 100644 >>> --- a/gcc/testsuite/gcc.target/aarch64/saddw-2.c >>> +++ b/gcc/testsuite/gcc.target/aarch64/sadalp-2.c >>> @@ -14,5 +14,4 @@ t6(int len, void * dummy, int * __restrict x) >>> return result; >>> } >>> >>> -/* { dg-final { scan-assembler "saddw" } } */ >>> -/* { dg-final { scan-assembler "saddw2" } } */ >>> +/* { dg-final { scan-assembler {\tsadalp\tv[0-9]+\.2d, v[0-9]+\.4s} } } */ >>> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c >>> b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c >>> similarity index 75% >>> rename from gcc/testsuite/gcc.target/aarch64/uaddw-1.c >>> rename to gcc/testsuite/gcc.target/aarch64/uadalp-1.c >>> index 14dff87d7f0..c4034384aae 100644 >>> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-1.c >>> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-1.c >>> @@ -14,5 +14,4 @@ t6(int len, void * dummy, unsigned short * __restrict x) >>> return result; >>> } >>> >>> -/* { dg-final { scan-assembler "uaddw" } } */ >>> -/* { dg-final { scan-assembler "uaddw2" } } */ >>> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */ >>> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c >>> b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c >>> similarity index 75% >>> rename from gcc/testsuite/gcc.target/aarch64/uaddw-2.c >>> rename to gcc/testsuite/gcc.target/aarch64/uadalp-2.c >>> index 79d0d094fc3..395d36c7c00 100644 >>> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-2.c >>> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-2.c >>> @@ -14,6 +14,5 @@ t6(int len, void * dummy, unsigned short * __restrict x) >>> return result; >>> } >>> >>> -/* { dg-final { scan-assembler "uaddw" } } */ >>> -/* { dg-final { scan-assembler "uaddw2" } } */ >>> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.4s, v[0-9]+\.8h} } } */ >>> >>> diff --git a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c >>> b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c >>> similarity index 74% >>> rename from gcc/testsuite/gcc.target/aarch64/uaddw-3.c >>> rename to gcc/testsuite/gcc.target/aarch64/uadalp-3.c >>> index 39cbd6b6cc2..5fdb1639ab8 100644 >>> --- a/gcc/testsuite/gcc.target/aarch64/uaddw-3.c >>> +++ b/gcc/testsuite/gcc.target/aarch64/uadalp-3.c >>> @@ -14,5 +14,4 @@ t6(int len, void * dummy, char * __restrict x) >>> return result; >>> } >>> >>> -/* { dg-final { scan-assembler "uaddw" } } */ >>> -/* { dg-final { scan-assembler "uaddw2" } } */ >>> +/* { dg-final { scan-assembler {\tuadalp\tv[0-9]+\.8h, v[0-9]+\.16b} } } */ >>> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c >>> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c >>> new file mode 100644 >>> index 00000000000..0aec0bf81c8 >>> --- /dev/null >>> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_1.c >>> @@ -0,0 +1,39 @@ >>> +/* { dg-do compile } */ >>> +/* { dg-options "-O3 -march=armv8-a -mautovec-preference=asimd-only -- >>> param vect-epilogues-nomask=0" } */ >>> + >>> +/* Widening sum reductions should use the pairwise widening add and >>> + accumulate instructions rather than a chain of extensions feeding >>> + [SU]ADDW pairs. */ >>> + >>> +#define DEF(NAME, ITYPE, OTYPE) \ >>> + OTYPE NAME (const ITYPE *a, long n) \ >>> + { \ >>> + OTYPE s = 0; \ >>> + for (long i = 0; i < n; i++) \ >>> + s += a[i]; \ >>> + return s; \ >>> + } >>> + >>> +DEF (sum_u8_l, unsigned char, long) >>> +DEF (sum_i8_l, signed char, long) >>> +DEF (sum_u16_l, unsigned short, long) >>> +DEF (sum_i16_l, short, long) >>> +DEF (sum_u32_l, unsigned int, long) >>> +DEF (sum_i32_l, int, long) >>> +DEF (sum_u8_i, unsigned char, int) >>> +DEF (sum_i8_i, signed char, int) >>> +DEF (sum_u16_i, unsigned short, int) >>> +DEF (sum_i16_i, short, int) >>> + >>> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, >>> v[0-9]+\.16b\n} >>> 2 } } */ >>> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.8h, >>> v[0-9]+\.16b\n} >>> 2 } } */ >>> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, >>> v[0-9]+\.8h\n} 2 >>> } } */ >>> +/* { dg-final { scan-assembler-times {\tsaddlp\tv[0-9]+\.4s, >>> v[0-9]+\.8h\n} 2 >>> } } */ >>> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, >>> v[0-9]+\.4s\n} 3 >>> } } */ >>> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.2d, >>> v[0-9]+\.4s\n} 3 >>> } } */ >>> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.4s, >>> v[0-9]+\.8h\n} 2 >>> } } */ >>> +/* { dg-final { scan-assembler-times {\tsadalp\tv[0-9]+\.4s, >>> v[0-9]+\.8h\n} 2 >>> } } */ >>> + >>> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */ >>> +/* { dg-final { scan-assembler-not {\tsaddw2?\t} } } */ >>> +/* { dg-final { scan-assembler-not {\tzip1\t} } } */ >>> diff --git a/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >>> b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >>> new file mode 100644 >>> index 00000000000..01537deeb9f >>> --- /dev/null >>> +++ b/gcc/testsuite/gcc.target/aarch64/widen_sum_pairwise_2.c >>> @@ -0,0 +1,29 @@ >>> +/* { dg-do compile } */ >>> +/* { dg-options "-O3 -march=armv8.2-a+dotprod -mautovec- >>> preference=asimd-only --param vect-epilogues-nomask=0" } */ >>> + >>> +/* With dot product a 4x widening sum stays a single [SU]DOT, while a >>> + sum into 64-bit elements uses the pairwise widening instructions. */ >>> + >>> +int >>> +sum_u8_i (const unsigned char *a, long n) >>> +{ >>> + int s = 0; >>> + for (long i = 0; i < n; i++) >>> + s += a[i]; >>> + return s; >>> +} >>> + >>> +long >>> +sum_u8_l (const unsigned char *a, long n) >>> +{ >>> + long s = 0; >>> + for (long i = 0; i < n; i++) >>> + s += a[i]; >>> + return s; >>> +} >>> + >>> +/* { dg-final { scan-assembler-times {\tudot\tv[0-9]+\.4s, v[0-9]+\.16b, >>> v[0- >>> 9]+\.16b\n} 1 } } */ >>> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.8h, >>> v[0-9]+\.16b\n} >>> 1 } } */ >>> +/* { dg-final { scan-assembler-times {\tuaddlp\tv[0-9]+\.4s, >>> v[0-9]+\.8h\n} 1 >>> } } */ >>> +/* { dg-final { scan-assembler-times {\tuadalp\tv[0-9]+\.2d, >>> v[0-9]+\.4s\n} 1 >>> } } */ >>> +/* { dg-final { scan-assembler-not {\tuaddw2?\t} } } */ >>> -- >>> 2.50.1 (Apple Git-155)
