> On 5 Aug 2026, at 14:41, Richard Biener <[email protected]> wrote:
> 
> On Wed, 5 Aug 2026, Kyrylo Tkachov wrote:
> 
>> 
>> 
>>> On 4 Aug 2026, at 16:19, Tamar Christina <[email protected]> wrote:
>>> 
>>>> -----Original Message-----
>>>> From: Robin Dapp <[email protected]>
>>>> Sent: 04 August 2026 15:07
>>>> To: Richard Biener <[email protected]>; Kyrylo Tkachov
>>>> <[email protected]>
>>>> Cc: [email protected]; Tamar Christina <[email protected]>;
>>>> Robin Dapp <[email protected]>
>>>> Subject: Re: [PATCH 1/2] doc: describe the lane freedom of the widen_ssum
>>>> and widen_usum patterns
>>>> 
>>>>> I wonder if this is a step in the correct direction, see PR67612 where
>>>>> we want to apply a SLP pattern to exercise the actual lane mapping
>>>>> done.
>>>> 
>>>>> Do we want to have separate optabs for this?  IMO what you document
>>>>> would be better named reduc_widen_ssum_optab (as opposed to
>>>>> reduc_widen_ssum_scal_optab or widen_ssum_optab).  Where the _scal
>>>>> variant would be expected to accumulate to a scalar (or lane zero),
>>>>> the reduc_*_optab would be free in how to accumulate lanes and
>>>>> the widen_ssum_optab would lay out exactly which source lanes are
>>>>> summed to which destination lanes (I hope ISAs have matching behavior
>>>>> here).
>>>> 
>>>> I would be very much in favor of a separate optab rather than narrowing
>>>> widen_[us]sum's scope.  You could argue that sum vs plus is already an
>>>> inconsistency, though.
>>>> Last year I experimented with making widen_[us]sum available to RVV.
>>>> That necessitated a few changes in the vectorizer (we don't expect
>>>> "SLP-style" "number of lanes stays the same" everywhere) but it's not
>>>> that big of a change.
>>>> 
>>>> Of course the general question remains where the widening should
>>>> actually be recognized.  For riscv it works quite well doing it "late".
>>>> 
>>> 
>>> I think the widening and other similar optimizations need to stay in the 
>>> vectorizer
>>> because for masked based targets like SVE if we generate an unpacking then
>>> the predicate needs to be unpacked too.
>>> 
>>> That means any "late" detection has to clean up the predicate operations too
>>> which usually brings the chain out of what combine supports.  But also by
>>> recognizing it In the vectorizer it allows us to cost alternate strategies 
>>> like
>>> unpacking or using sparse vectors.
>>> 
>>> So I don't think we can move this past costing as it has a significant 
>>> effect on
>>> Codegen.
>>> 
>> 
>> So as I’m not dealing day-to-day with the vectorizer, what would you 
>> recommend as the way forward here?
>> Rename the optabs currently used for lane-agnostic reduction to 
>> reduc_widen_[us]sum_optab and implement those for aarch64 as in patch 2/2 
>> and document the widen_[us]sum optabs as lane-preserving?
> 
> naming consistency would be nice indeed.  Note WIDEN_SUM_EXPR
> and related should also be unambiguously documented (I'm fine with
> chosing lane-agnostic interpretations for them - in fact they should
> all cease to exist and be replaced with internal functions only).

Something like:
https://gcc.gnu.org/pipermail/gcc-patches/2026-August/727035.html ?
Thanks,
Kyrill


> 
>> Thanks,
>> Kyrill
>> 
>>> Thanks,
>>> Tamar
>>> 
>>>> --
>>>> Regards
>>>> Robin
>> 
>> 
>> 
> 
> -- 
> Richard Biener <[email protected]>
> SUSE Software Solutions Germany GmbH,
> Frankenstrasse 146, 90461 Nuernberg, Germany;
> GF: Jochen Jaser, Andrew McDonald, Abhinav Puri; (HRB 36809, AG Nuernberg)


Reply via email to